# Faces and identity — SCRFD and MobileFaceNet Spec for **S14**, the spike that decides whether §3.9.1 is buildable, and the build that follows it. FR-CULL-8 … FR-CULL-12 specify the subsystem in terms of "a 512-dimension embedding from a stated model" and stop there, deliberately: [catalog.md §10.4](catalog.md) records that the model was left abstract because D13 was open. This document names the two models, fixes the pre- and post-processing they need, specifies the calibration and clustering that sit on top, and states what S14 has to measure before any of it is trusted. **It does not close the licensing half of D13.** §2 is the reason, and it comes first because [segmentation.md §7](segmentation.md) established the precedent that reading the grant is cheaper than discovering it at packaging time. --- ## 1. The two models, and why these two **Detector — SCRFD.** *Sample and Computation Redistribution for Efficient Face Detection* (Guo et al., ICLR 2022). A single-stage anchor-based detector whose contribution is where the FLOPs go, not a new architecture: the shallow stages carry more capacity because that is where small faces are decided. It emits, per face, a box, a confidence, and **five landmarks in the same forward pass** — which is the property that matters here, because FR-CULL-8 requires landmarks and because §4's alignment step is not optional. A detector without landmarks would need a second network to supply them. The `500M` variant is the one to ship: 500 MFLOPs at VGA, against 10 GFLOPs for `10G`. Face indexing is a background sweep over a whole library on a phone as well as a desktop (NFR-RES-2), so the cheap variant's accuracy loss on tiny faces is the right trade — a face too small for `500M` to find is also too small for §5 to embed usefully. **Embedder — MobileFaceNet trained with ArcFace loss ("MBF").** ~1M parameters, ~0.45 GFLOPs for one 112×112 crop, 512-d output. ArcFace's additive angular margin is what makes the output space one where *cosine similarity means something* — the loss explicitly optimises angular separation between identities, which is what §6's calibration then has to convert into a probability. **Why not the larger ResNet50 embedder** (`w600k_r50`, the other half of InsightFace's `buffalo_l`): because it is not measurably better here, and it is 13× the file. §1.1's reference project fitted the same Platt calibration (§8) to all four candidates over a 2,418-identity gallery, which is a clean read on the embedding space alone with no tracking or matching logic in it: | Model | File | Steepness `a` | Boundary at P=0.5 | Held-out macro F1 | |---|---|---|---|---| | LVFace-B (Glint360K) | 455 MB | 17.7 | sim 0.228 | 67.4% | | **ArcFace w600k-MBF** | **13 MB** | **16.2** | **sim 0.267** | **64.4%** | | ArcFace w600k-R50 | 174 MB | 15.4 | sim 0.301 | — | | ArcFace R18 | 48 MB | 15.3 | sim 0.309 | 63.1% | **MBF's calibration curve is steeper than R50's and R18's**, separating same-identity from different-identity pairs more confidently and at a lower boundary, at a fraction of the size. LVFace-B is genuinely better and is 35× the file — a defensible choice for a desktop-only feature and not one for a subsystem that has to index a library on a phone (NFR-RES-2). Two caveats on that table, both from the source: the R50 gallery was built with ~30% fewer reference images per identity, which is why it carries no F1 figure — the calibration row does not depend on the image count and is comparable, the benchmark row would not have been. And the F1 column measures a *film-cast identification* task, not this one. It ranks the models; it does not predict DarkRoom's accuracy. **Both are pure convolutional graphs**, which matters for §3: the runtime is tract, and tract's operator coverage is the thing that decides whether a graph runs at all. ### 1.1 A working reference implementation exists `../scene-actor-extraction` is a C++ pipeline by the same author that runs **exactly this model pair** — SCRFD-500MF then ArcFace over 5-point-aligned 112×112 crops — over feature films, with a fitted Platt calibration and a scored benchmark behind it. It is **MIT licensed**, so porting from it into GPL-3.0-or-later is straightforward and needs attribution, not permission. That changes what S14 is. The open questions are no longer "does this pair work" and "what are the magic numbers" — they are *does tract load these graphs* (§12 M1, and §4.1 says why that is in genuine doubt) and *does the accuracy hold on family snapshots rather than film frames*. The pre-processing constants, the decode layout, and the calibration algorithm are all readable rather than rediscoverable, and several of them are not what the published Python would lead you to write. | What | Reference | Ports to | |---|---|---| | ArcFace 5-point template and warp | `src/face_utils.hpp` | `align.rs` (§5) | | SCRFD pre-process, decode, NMS | `src/backends/ort_backend.cpp` `SCRFDDecoder` | `detect.rs` (§4) | | ArcFace pre-process and L2 normalise | same file, `ArcFaceEmbedder` | `embed.rs` (§6) | | Platt fit over a similarity histogram | `src/gallery/gallery_calibration.hpp` | `calibrate.rs` (§8) | | Model-free unit tests for both | `tests/test_face_utils.cpp`, `tests/test_calibration.cpp` | the §3 feature split | That last row is worth noticing: the reference already separates the geometry and arithmetic from the inference well enough to unit-test them with no model on the machine. §3's `inference` feature flag is the same boundary, enforced by Cargo rather than by discipline. **What does not port.** The reference matches faces against a *known gallery* of named identities; DarkRoom clusters an unlabelled library (§9). And it is a video pipeline, so it has a tracker, a `max_faces` cap, and frame-to-frame identity annealing that have no analogue here. --- ## 2. Licensing — the open half of D13 The architectures are published research. **The weights everyone actually uses are not redistributable by this project.** InsightFace's code is MIT. Its *pretrained models* — `buffalo_l`, `buffalo_s`, `buffalo_sc`, and every `det_*` and `w600k_*` checkpoint inside them — carry a **non-commercial research-only** grant, stated in the model-zoo README and applying to both the auto-downloaded and the manually downloaded artefacts. `w600k_mbf` is trained on WebFace600K, whose own terms are research-only as well, so the restriction has two independent sources rather than one that might be renegotiated. These are the exact artefacts in question — §1.1's reference project fetches them from InsightFace's own release page, which is where the grant attaches: | File | Source | |---|---| | `det_500m.onnx` (2.5 MB) | `insightface/releases/download/v0.7/buffalo_sc.zip` | | `w600k_mbf.onnx` (13 MB) | `insightface/releases/download/v0.7/buffalo_s.zip` | DarkRoom is GPL-3.0-or-later. A non-commercial field-of-use restriction is not a GPL-compatible grant, so: - these weights **cannot be committed to this repository**, the way `yolo26n-seg.onnx` is (D14); - they cannot ship inside an APK, a Flatpak, or an F-Droid build; - and the restriction binds the *user*, not only the project — a professional photographer using DarkRoom commercially is outside the grant even if they fetched the file themselves. That last point is the one that is easy to miss and dishonest to leave implicit. It is why §2.2's recommended route puts the licence text in front of the user rather than in a footnote. ### 2.1 The routes, priced | Route | What ships | Cost | |---|---|---| | **A — commit the weights** | Everything works out of the box, one `git lfs pull` | **Not available.** Licence-incompatible with GPL-3.0 and with all three distribution channels (NFR-COMPAT-2). Listed only so it is visibly rejected rather than quietly assumed. | | **B — permissively licensed weights** | Same, legally | Requires weights that do not exist off the shelf in this architecture pair. Either found (§2.3) or trained, and training an ArcFace embedder needs a face dataset whose own terms permit redistribution — the harder half of the problem. | | **C — the user fetches them** | The app ships the *code*, not the weights; face indexing is off until the user obtains a model | Works today, keeps the repository clean, and is honest. Costs a first-run flow, an F-Droid anti-feature declaration, and a licence notice the user has to read. | **Recommendation: C now, B when it becomes possible.** §1.1's project already works this way — a `scripts/download_models.sh` that fetches the weights rather than a repository that carries them — so route C is a practised pattern here rather than a theory. Two things DarkRoom must do that the reference does not: pin a checksum for **every** model (the reference verifies only one of the four), and put the licence in front of the user, because a photo editor's users are not all researchers. The schema already forces this to be a survivable choice — `faces.model_id` (catalog.md §10.1) exists precisely so that a model change is detectable and re-indexable rather than silently poisoning every similarity in the library. Under C, swapping in a permissive model later is a new `model_id` and a re-index, not a migration. ### 2.2 What route C requires of the code - **The weights are never a build input.** `dr-face` (§3) takes bytes; it has no `embedded-model` feature and no `models/` directory. This is the one structural difference from `dr-segment`, and it is deliberate — a feature flag that *could* embed weights is a feature flag someone eventually turns on in a packaging script. - **The fetch does not live in `dr-face`.** It lives in the app layer, so the inference crate keeps no network dependency at all. §11 makes that a checkable property rather than a convention. - **The licence is shown, not linked.** Before the first download, the app states in plain language that the weights are research-only, that commercial use is outside the grant, and who the grantor is. NFR-SEC-5's last bullet already requires the models to be named, versioned and licensed in the about screen; this is the same obligation moved to the moment where it can still change a decision. - **A checksum is pinned.** The app verifies the digest of what it fetched against a value compiled in, and refuses a mismatch. An unverified model file is an arbitrary graph executed over the user's family photographs. - **Indexing stays off until a model is present and verified**, and disabling it deletes nothing — that is a separate control (NFR-SEC-5, §11). ### 2.3 Before S14 writes any code Three questions to answer by reading, in this order, and to record with the date they were checked — the same discipline [segmentation.md §7](segmentation.md) applied: 1. **Is there a permissively licensed SCRFD export?** Third-party ONNX exports of SCRFD are plentiful; a re-export inherits the original weights' grant, so the question is whether anyone has *trained* SCRFD on a redistributable dataset, not whether they have converted the InsightFace one. 2. **Is there a permissively licensed 512-d MobileFaceNet/ArcFace checkpoint?** Candidates worth reading the terms on rather than assuming: EdgeFace (Idiap), the ONNX Model Zoo's ArcFace entry (MIT, but a ResNet100 at ~250 MB — licence-clean and the wrong size), and any MobileFaceNet trained on a dataset with distribution terms. 3. **What does a permissive *detector* alternative cost in accuracy?** YuNet (OpenCV Zoo) is 232 KB, emits five landmarks, and comes from a permissively licensed repository. If it is close enough, the detector half of the licence problem disappears and only the embedder remains — a materially better position than either half alone. Question 3 is nearly free to answer: `face_detection_yunet_2023mar.onnx` is **already sitting in §1.1's `models/` directory**, fetched from `opencv/opencv_zoo`. It needs its own decode path (its output layout is different, which is why the reference's `SCRFDDecoder` explicitly rejects it at load rather than silently misreading it), and then it is one more row in §12's table. --- ## 3. Crate shape — `core/dr-face` A new workspace member, modelled on `dr-segment` and for the same reason: everything that reasons about faces is testable with no GPU adapter and no model present (ARCH §6.5a). ``` core/dr-face/ src/lib.rs FaceError, re-exports src/detect.rs SCRFD: preprocess, decode, NMS src/align.rs 5-point similarity transform → 112×112 crop src/embed.rs MBF: preprocess, forward, L2 normalise src/calibrate.rs cosine → P(same person) (FR-CULL-9) src/cluster.rs constrained agglomeration (FR-CULL-10) ``` ```toml [features] # No `embedded-model`. §2.2 — the weights are not a build input, ever. default = [] # The ONNX runtime. Off by default so `calibrate` and `cluster` — which are # arithmetic over embeddings and have no model in them — stay testable in a # build that carries no inference engine at all. inference = ["dep:ort", "dep:ort-tract", "dep:ndarray"] ``` The `ort` + `ort-tract` pairing is settled by D13's 2026-08-21 update and already in the workspace manifest: `ort`'s API, tract's pure-Rust engine, no C under the NDK. **The split between `inference` and the rest is load-bearing.** Calibration and clustering are where the subsystem's accuracy actually lives, they are pure functions of embeddings and labels, and they must be testable against synthetic embedding sets with no weights on the machine. A test suite that needs a research-licensed download to run is a test suite that does not run in CI. ### 3.1 The API ```rust /// A loaded detector. Fixed input shape — see §4. pub struct Detector { /* session, input edge */ } impl Detector { pub fn from_bytes(onnx: &[u8], id: ModelId) -> Result; /// `rgb` is f32 0..=1, row-major, three per pixel. pub fn detect(&self, rgb: &[f32], w: usize, h: usize, opts: &DetectOptions) -> Result, FaceError>; } pub struct Detection { /// Normalised to the image's long edge (catalog.md §10.1). pub bbox: (f32, f32, f32, f32), /// Five points, same normalisation, in the model's own order (§5). pub landmarks: [(f32, f32); 5], pub confidence: f32, } pub struct Embedder { /* session */ } impl Embedder { pub fn from_bytes(onnx: &[u8], id: ModelId) -> Result; /// Takes an *aligned* 112×112 crop. Alignment is `align::warp`; passing a /// raw bbox crop here is the silent-accuracy-loss bug §5 exists to prevent. pub fn embed(&self, aligned: &Aligned112) -> Result; } /// L2-normalised, 512-d. Carries its model id so a comparison across models /// is a type error rather than a plausible-looking number (catalog.md §10.1). pub struct Embedding { pub model: ModelId, pub v: [f32; 512] } ``` `Aligned112` is a newtype over the pixel buffer that only `align::warp` can construct. That is the whole defence against §5's failure mode, and it costs nothing. --- ## 4. Detection ### 4.1 Fixed input, and how the image is fitted to it **`det_500m.onnx` has a dynamic H/W input, and this is the largest single risk in the document.** §1.1's reference had to use ONNX Runtime rather than OpenCV's `dnn` module precisely because OpenCV could not load SCRFD's dynamic `Shape` nodes. tract is in the same family of problem: `dr-segment` exists in its current shape because tract failed shape inference on YOLO's dynamic export, which is why `tools/export-seg-model.sh` passes `dynamic=False` and why `semantic.rs` has a fixed `INPUT_EDGE`. So the graph almost certainly needs its input dimensions frozen before tract will touch it. That is a known, cheap operation — `onnxruntime.tools.make_dynamic_shape_fixed` rewrites the declared dims without retraining or re-exporting from PyTorch — but it is a step, it produces an artefact that must itself be reproducible (a `tools/fix-face-model-shapes.sh`, in the spirit of the existing export script), and **it must be tried before anything else in this document is scheduled.** §12's M1. Fixed at **640**, with 320 available as a faster, blinder option to measure. **Letterbox.** Scale by `min(640/w, 640/h)` preserving aspect, then paste into a 640×640 canvas. The reference centres the image and fills the margin with grey `114`, inverting with `x_src = (x_model − pad_x) / scale`. What matters is not *where* the padding goes but that the forward and inverse agree and that the fill value is treated as part of the contract: a mismatch between them offsets every box and landmark the model returns by the padding, which yields detections that look plausible, landmarks that align badly, and an embedding-quality loss that only shows up as poor clustering three stages later. Port the reference's convention rather than inventing one, and assert it with a round-trip test. Normalisation is `(x·255 − 127.5) / 128`, RGB, NCHW — note `/128`, not `/127.5`; see §6. ### 4.2 Outputs, and decoding them **Nine tensors for three strides `{8, 16, 32}`, twelve for four `{8, 16, 32, 64}`.** Which one a given export produces is discovered at load — `strides = output_count / 3` — not assumed, because both variants exist and hardcoding three silently ignores the largest faces in a twelve-output model. With two anchors per location and `N_s = (640/s)²` locations: | Tensor | Shape | Meaning | |---|---|---| | `score_s` | `[N_s·2, 1]` | sigmoid already applied inside the graph | | `bbox_s` | `[N_s·2, 4]` | distances left, top, right, bottom **in units of the stride** | | `kps_s` | `[N_s·2, 10]` | five `(dx, dy)` offsets, same units | Anchor centres are `(x·s, y·s)` for each grid cell, repeated once per anchor. Decoding is therefore ``` x1 = cx − l·s x2 = cx + r·s kx_i = cx + dx_i·s y1 = cy − t·s y2 = cy + b·s ky_i = cy + dy_i·s ``` then divide by the letterbox scale to return to source pixels, then normalise by the long edge before storage. Flat index within a stride is `(row · fw + col) · 2 + anchor`. **A shape assertion at load time, not a decode-time surprise.** `dr-segment` already learned this — `SegmentError::OutputShape` exists because a different export of the same model produces confidently wrong results otherwise. The reference implements exactly this and it is worth porting verbatim: output count divisible by three and between 9 and 12, then the **last dimension of each group checked against `{1, 4, 10}`**. That single check is what catches a YuNet file passed where an SCRFD one was meant — YuNet also has twelve outputs, so the count alone does not distinguish them, and the failure without the check is not an error but a page of plausible garbage. ### 4.3 Thresholds Score ≥ **0.5**, NMS IoU **0.4**, plain greedy NMS per image across all strides together. Deliberately *not* the low threshold `dr-segment` chose. There, a false positive costs one spurious entry in a list the user is picking from. Here it costs an entry in the People view that the user has to reject, in a library with thousands of images — and worse, a garbage embedding that participates in clustering and can bridge two real clusters into one. False negatives are recoverable by a later re-index with a better model; a polluted cluster graph is not, once the user has confirmed faces inside it. A minimum box size of **40 px on the source proxy** is applied on top — the reference's figure, chosen there for the same reason it holds here: below it there is not enough face left to align reliably, and §7's table shows what the embedder actually receives at that size. §12's M4 is what moves this number, if measurement says it should move. **One thing not to port:** the reference caps detections at ten per frame, largest first. That is right for a film frame, where the extras in the background are noise. It is wrong for a photo library, where a group shot with thirty faces in it is precisely the picture the user wants indexed. No cap; the min-size floor is the only filter. --- ## 5. Alignment — the step that is silently wrong when skipped ArcFace embeddings are trained on faces warped to a canonical 112×112 arrangement. Feeding the model a plain bounding-box crop *works* — it produces 512 numbers, they are unit-norm, and cosine similarities between them look entirely reasonable. They are just much worse, and nothing in the system reports it. The transform is a **similarity transform** — rotation, uniform scale, translation, four degrees of freedom — from the five detected landmarks to this template, which is the arrangement the weights were trained against: ``` (38.2946, 51.6963) subject's right eye ─┐ image-left of centre (73.5318, 51.5014) subject's left eye ─┘ (56.0252, 71.7366) nose tip (41.5493, 92.3655) subject's right mouth corner (70.7299, 92.2041) subject's left mouth corner ``` Similarity, not affine: a full affine fit to five noisy points will happily shear a face, and shear is exactly the deformation the embedder was never shown. **The naming is a trap and the order is not.** Point 0 sits at x=38 on a 112-wide canvas — left of centre *in the image*, which is the subject's **right** eye. Both namings are in circulation and they are opposite. What matters is that SCRFD emits its five points in this same order, so the correct amount of reordering between detector and template is **none**; the reference states this explicitly in `types.hpp` for the benefit of whoever next reads it and doubts it. A future detector with a different order carries its own permutation next to its `model_id`, rather than this file growing an assumption. **One divergence from the reference to settle by measurement.** It fits the transform with OpenCV's `estimateAffinePartial2D` under **RANSAC** at a 3-pixel threshold, and drops the detection when the fit fails. RANSAC over five points is a strange fit — the minimal sample for a similarity is two, so it can discard landmarks it judges outliers and solve from a subset, which on a profile face is as likely to be the correct geometry as the wrong one. InsightFace's own pipeline uses a plain least-squares (Umeyama) similarity over all five points, which cannot silently drop anything and is the one to write first. §12's M5 measures whether the choice matters; the degenerate-input path still has to exist either way, because collinear landmarks do occur. **A trick worth keeping.** The reference retries detection on a frame where nothing was found, after replicate-padding by 25% and applying CLAHE to the L channel. Padding gives a face at the very edge a context the detector needs; CLAHE rescues underexposed frames. Both are plausible on scanned negatives and backlit photographs — worth trying at low priority, since a retry doubles the cost of exactly the images that yielded nothing. Sampling is bilinear from the *source proxy*, in one step — never crop-then-warp, which resamples twice and throws away detail the warp could have used. Pixels falling outside the source are black. **The landmark order must match the template order.** The template above is written in the detector's own output order; if a future detector emits them differently, the template is reordered with it and that mapping belongs next to the model id, not compiled in as an assumption. *Test:* warp a synthetic image with a known rotation and scale, and assert the five points land on the template within a fraction of a pixel. This is testable with no model present, which is why `align` sits outside the `inference` feature. --- ## 6. Embedding 112×112 **RGB** — the crop comes out of the warp in whatever order the source was in, and the reference converts BGR→RGB explicitly before the blob rather than relying on a flag, which is the readable way to do it. Normalisation is `(x·255 − 127.5) / 128`. **Note `/128`, not `/127.5`:** InsightFace's published Python uses `1.0/127.5` for the recognition model, the reference uses `1.0/128` for both models, and every measured number in §1's table was produced with `/128`. The difference is 0.4% of scale and almost certainly immaterial, but "almost certainly" is not a reason to pick silently — write `/128` to match the numbers we have, and settle it with one back-to-back run in §12. Output is 512 floats; **L2-normalise before storing**, so every downstream comparison is a dot product and no code path has to remember to normalise. The reference clamps the norm at 1e-6 before dividing, which costs nothing and removes a NaN path. Some exports of these graphs are fp16 in, fp16 out. The reference detects this from the graph's declared element type rather than from the filename; worth porting, because the alternative failure is a silent garbage tensor. Storage is `512 × f16` (catalog.md §10.1) — 1 KB per face, 25 MB for a 25,000-face library. The f16 round-trip perturbs a unit vector by ~1e-3 in cosine, three orders of magnitude below the separation between a match and a non-match, and the halving matters because these rows are the ones NFR-SEC-5 contemplates optionally syncing. Re-normalise on load after the f16 widen. It is one pass over 512 floats and it removes a class of drift that is otherwise invisible. --- ## 7. Which pixels the pipeline actually sees FR-CULL-8 pins indexing to the FR-CULL-2 ladder: **the proxy tier, never a full decode.** The relevant tier is `ThumbSize::Large` — 1024 px on the long edge (`dr-thumbs`). That has a consequence worth stating in numbers rather than discovering in a clustering report. On a 1024 px proxy: | Face size in frame | Pixels across | What §6 receives | |---|---|---| | A portrait, face fills a third of the frame | ~340 | Downsampled to 112. Ideal. | | Two people, half-length | ~120 | Roughly native. Good. | | A group of eight | ~50 | **Upsampled** to 112. Degraded, and usable. | | A figure in a landscape | ~20 | At or below §4.3's floor. Rejected. | So `faces` records one column beyond catalog.md §10.1's schema: ```sql ALTER TABLE faces ADD COLUMN crop_px INTEGER NOT NULL; -- source pixels across the aligned crop ``` `crop_px` earns its place three times over. It is the honest quality signal for the UI; it is a **feature in §8's calibration**, which FR-CULL-9 explicitly demands ("a raw cosine means something different for every model, every population, and *every face size*"); and it is what a future higher-resolution re-embedding pass would select on, so that pass becomes a query rather than a re-index of everything. **No RAW decode is added.** Where the Large proxy is missing, the job enqueues a `Thumbnail` job at background priority and re-queues itself, exactly as FR-CULL-8 requires. --- ## 8. Calibration — cosine to probability FR-CULL-9 makes this a hard requirement: no code path may threshold a bare cosine, every threshold in the subsystem is stated as a probability, and the fit is per library and reports its own validity. This section is how that is obtained, and the interesting part is where the training pairs come from when the user has labelled nothing yet. ### 8.1 Where the pairs come from **Negatives are free and abundant.** Two faces detected in *the same photograph* are almost never the same person. That gives every multi-face image in the library a full set of negative pairs at no labelling cost — and they are *hard* negatives, drawn from the same camera, lighting, and processing, which is exactly the population where a threshold tuned on easy negatives fails. The known exceptions — mirrors, photographs of photographs, collages, a spliced panorama — are rare enough to be noise at this scale and worth naming so the exception is not mistaken for a bug later. **Positives, in order of trustworthiness:** 1. **User confirmations** (FR-CULL-10). Every pair of faces confirmed to the same person. The gold standard, and empty on day one. 2. **Burst siblings.** FR-CULL-5 already groups bursts. Two faces in adjacent frames of one burst, in nearly the same position, are near-certainly the same person. Free, requires no labelling, and available immediately on any library with continuous-shooting frames in it. **Their purity is an S14 measurement** (§12), not an assumption — if bursts turn out to be dirtier than expected, this source is dropped and the calibration simply stays invalid for longer. 3. Nothing else. Bootstrapping positives from high cosine similarity is circular — it fits the calibration to the belief it was supposed to test. ### 8.2 The fit §1.1's reference already implements this and its shape should be ported rather than reinvented: ``` P(same | cos) = σ(a·cos + b + log_prior_odds) ``` Four details in it are the difference between working and nearly working. **Fit against a histogram, not against pairs.** A 25,000-face library has 3×10⁸ pairs; no gradient descent is running over that. The reference buckets every pair into **200 bins over cos ∈ [−1, 1]**, carrying a positive and a negative count per bin, and fits the two parameters against the per-bin counts. The cost becomes the similarity matrix plus 200 numbers, and the matrix is a single GEMM (§9). This is the trick that makes a per-library fit affordable at all, and it is not obvious from the outside. **Balance the classes explicitly.** `w_pos = total/(2·n_pos)`, `w_neg = total/(2·n_neg)`. Negatives outnumber positives by orders of magnitude; an unweighted fit produces a well-shaped curve sitting at the wrong height — precisely the "plausible number all the way to the user interface" failure FR-CULL-9 describes. **The base rate is a runtime argument, not part of the fit.** `log_prior_odds` is added at evaluation time, so the balanced fit is stored once and the prior varies per query — the odds that two faces drawn from a 40-image holiday album match are not the odds for a 40,000-image archive. Its inverse, `boundary_at(p)`, gives the cosine at which the probability crosses `p`, which is what turns §9's "merge above 0.9" into an actual comparison. Baking a prior into `a` and `b` would need a refit per context and would make the stored parameters mean something different depending on where they came from. **Deduplicate before pairing.** Near-identical embeddings (cos > 1 − 1e-7) are the same photograph counted twice; the reference drops them per identity first. In a photo library the equivalent is duplicates and virtual copies (FR-CAT-11), and leaving them in stacks the positive histogram at cos ≈ 1 with pairs that teach the fit nothing about hard cases. **A third feature this design adds.** The reference fits on cosine alone; DarkRoom should carry `crop_px` too: ``` logit P(same) = w₀ + w₁·cos + w₂·log₂(min(crop_px_a, crop_px_b)) ``` FR-CULL-9 explicitly names face size as an axis along which an uncalibrated similarity misbehaves, §7 shows a real library spans 40 px to 340 px of face, and `crop_px` is already in hand. The *minimum* of the pair, because a comparison is only as good as its worse crop. It generalises the histogram to a small 2-D grid of bins, which changes nothing structural. If M7 says the extra feature buys nothing, drop it and match the reference exactly — but a film pipeline working from broadcast frames had far less size variation to explain than this one does. ### 8.3 Validity, and saying so Stored per catalog.md §10.3: parameters, a validity flag, and a hash of the face set fitted from. Invalid when fewer than **200 positive pairs** or **2,000 negative pairs** are available, or when the reliability check fails. Deliberately far stricter than the reference's floor of two positives and one negative. That floor is reasonable there: its pairs come from a curated gallery of labelled reference portraits, where a positive pair is trustworthy by construction. Here the positives are bootstrapped from bursts and early confirmations (§8.1) and the whole risk is fitting confidently to a handful of them. The reference also refuses to draw positives from an identity with fewer than five distinct embeddings, letting it contribute negatives only — the same asymmetry applies to a thinly-confirmed person and is worth keeping. In that state the UI says confidence is unavailable and the People view still works — clustering falls back to a documented default operating point, labelled in the interface as an untuned default, and no probability is displayed. That is FR-CULL-9's requirement read literally: not presenting an untuned default *as though it were measured*. Refit is triggered by the same debounce as clustering (§9), and when the confirmed-pair count grows materially. *Acceptance (FR-CULL-9):* a reliability diagram over ten probability bins, on a held-out labelled split, with observed match rate within a stated tolerance of the predicted probability in each populated bin. A single accuracy figure is not an answer to this requirement. --- ## 9. Clustering Not a job kind, for the reason catalog.md §10.2 gives: it is a whole-library operation with no natural `subject_id`. It runs as a debounced library pass when detection has been idle and the face count has moved materially. **The graph.** For each face, its `k = 20` nearest neighbours by cosine, then each candidate edge scored through §8 to a probability. Brute force is honest arithmetic here, and §1.1's project is the evidence rather than the estimate: it computes the full similarity matrix as **one GEMM** (`E · Eᵀ` over L2-normalised rows) at gallery scale, on CPU, and instruments it well enough to have compared CPU against OpenCL and found the question worth asking rather than urgent. For DarkRoom that is a 25,000 × 512 matrix at 25 MB in and ~2.5 GB out — so the matrix is computed **in row blocks**, with each block reduced to its top-k and its histogram contribution before the next is started, and never materialised whole. Once, in the background, after an indexing sweep that took an hour. An approximate index is an optimisation to reach for when §12 says it is needed, not before. **Constraints, not just thresholds:** - **Cannot-link on co-occurrence.** Two faces in the same image are never merged. This is the same observation §8.1 mines for negatives, used here as a hard constraint, and it is the single cheapest defence against the over-merging FR-CULL-10 warns about. - **Confirmed faces are anchors.** A confirmation is user data (FR-CULL-12) and clustering never moves it. Two clusters each containing confirmations of *different* people cannot merge; a cluster containing confirmations of one person absorbs suggestions but never reassigns the confirmed. **The algorithm.** Constrained average-link agglomeration over the probability graph, merging while the average pairwise probability exceeds **0.9** and no cannot-link is violated. Average-link rather than single-link because single-link chains — one bad edge welds two identities together, which is the documented way face clustering fails on families. **Incremental by default.** A new face joins the existing cluster whose average probability against it is highest, if that clears the threshold; otherwise it starts an unnamed group. A full re-agglomeration happens on model change, on recalibration, and on explicit user request. Suggestions are recomputed freely; confirmations survive all of it (FR-CULL-10). **Splitting** re-agglomerates within one person at a raised threshold and offers the resulting groups as the split. FR-CULL-10 requires splitting to be as easy as merging, and a split that hands the user a pile of loose faces to re-sort is not that. --- ## 10. Catalog and jobs **Schema** is catalog.md §10.1's v5 migration, plus `faces.crop_px` (§7) and the calibration table (§8.3). Nothing else changes. **One new job kind:** ```rust /// Detect and embed faces in one image, from its proxy (FR-CULL-8). DetectFaces = 8, ``` Background priority, coalesced per `image_id`, interruptible, resumable — it inherits FR-CAT-3's properties with no exceptions, which is what FR-CULL-8's acceptance criterion is about. It is not a network job. One job does detection *and* embedding for every face in the image, rather than splitting them. Splitting would double the queue's row count for no benefit: the proxy is already decoded and in memory, and the natural unit of resumable work is one photograph. **Selector term** (FR-CULL-11): ```rust Person { id: PersonId, include_suggested: bool }, // defaults to false ``` Confirmed-only by default, so a saved smart collection does not silently change membership when a later indexing pass revises a guess. --- ## 11. Privacy obligations that constrain the code's shape NFR-SEC-5 is not a policy to remember; it is a set of properties the code either has structurally or does not have at all. The checkable ones: - **`dr-face` has no network dependency.** No `reqwest`, no `ureq`, transitively. Worth a CI check over `cargo tree`, alongside the existing lints — the crate that holds the embeddings should be provably unable to send them anywhere. - **The model fetch (§2.2) is in the app layer**, which is why the API in §3.1 takes bytes. - **The diagnostics bundle is an allowlist** (NFR-OPS-1), so `faces`, `face_person` and the calibration table are excluded by not being named, and a future table cannot become uploadable by existing. - **Delete-all is one transaction and one control**: faces, links, people, calibration, and the cached crops if any. Separately, a switch that stops indexing so no such data is produced. The two are different actions and are not collapsed into one. - **The about screen reads the model metadata from the loader** — id, version, licence — rather than from a hardcoded string that will drift from what is actually running. --- ## 12. What S14 measures The corpus is a real personal library — the ~2,000-image sample S14 already specifies — with a hand-labelled identity subset. Fixed in advance, so the result is a measurement rather than an argument: | # | Measure | Why it decides something | |---|---|---| | **M1** | **Does tract load both graphs?** As shipped, then with the input dims frozen | Go/no-go, and it is *first*. `det_500m.onnx` has a **dynamic H/W input** (§4.1), which is the exact thing tract failed on for YOLO and which OpenCV's `dnn` could not load either — so the as-shipped answer is expected to be no, and the real question is whether `make_dynamic_shape_fixed` is enough or whether a PyTorch re-export is needed. Costs an afternoon; everything below is void without it. | | **M2** | Detection latency per image at 640, and at 320, on the reference desktop and one Android device | Whether a 17k library indexes in a background sweep or a weekend. The one measured datapoint we have — YOLO26n-seg, ~470 ms at 640×640 in tract — suggests SCRFD-500M lands well under it, but tract is not ORT and an extrapolation is not a measurement. | | **M3** | Embedding latency per face, and faces per image in a real library | The multiplier on M2. A library averaging 1.5 faces per image at 100 ms per face is an hour for 25k faces; at 400 ms it is four. | | **M4** | Detection recall against hand-labelled faces, bucketed by `crop_px` | Where §4.3's floor should actually sit, and what proportion of a real library's faces are in the degraded bucket §7 predicts. | | **M5** | **Aligned versus unaligned embeddings**, same corpus, same clustering. Then two A/Bs that cost one run each: least-squares versus RANSAC alignment (§5), and `/128` versus `/127.5` normalisation (§6) | Quantifies §5. If the gap is small the alignment code is still correct; if it is large, this is the measurement that stops someone "simplifying" it later. The two A/Bs close divergences between the reference and the published Python that are currently settled by assertion. | | **M6** | Cluster purity and completeness against the labelled subset | Whether the subsystem is worth building at all. §1's F1 figures rank the models on *film frames against a known cast*; nothing yet says how MBF behaves on family snapshots it must cluster blind. | | **M7** | Calibration reliability, and **the sample size at which the fit becomes valid** | FR-CULL-9's acceptance criterion, plus the practical question of whether a 2,000-image library ever gets a valid fit or whether §8.3's "unavailable" is the normal state. | | **M8** | Burst-derived positive-pair purity (§8.1) | Whether the free positives are usable or the fit waits for user confirmations. | | **M9** | YuNet as a drop-in detector: M4 and M6, re-run | §2.3's question 3, and the model file is already on disk. If the answer is "close enough", half the licence problem disappears. Needs its own decode path — its outputs are not SCRFD's. | | **M10** | Peak RSS during an indexing sweep | NFR-RES-2. Two loaded graphs plus a proxy plus a batch of crops, on a phone. | ### 12.1 M1 result — **PASS, conditionally** · 2026-08-26 Measured, not extrapolated. Both InsightFace graphs **fail to load in tract as shipped**, exactly as §4.1 predicted and for the reason it gave: ``` scrfd_500m_bnkps.onnx Translating node #0 "input.1" Source ToTypedTranslator arcface_w600k_mbf.onnx Failed analyse for node #139 "Conv_0" ConvHir ``` Both **load cleanly once their input dimensions are pinned** — SCRFD's unnamed H/W to 640, ArcFace's `None` batch to 1 — by `tools/fix-face-model-shapes.sh`, which rewrites the declared dims and touches no weights. The frozen SCRFD reports the layout §4.2 specifies, which is the second half of the answer: nine outputs, three strides, last dims 1/4/10, and `12800 = 80 × 80 × 2` confirming two anchors per location at stride 8. Two things worth carrying forward: **SCRFD's outputs were already static.** The export was made at 640 and only its input forgot to say so, so pinning to 640 is not a choice this project is making — it is the shape the graph was always going to run at. §12's "320 as a faster option" would need a different export, not a different flag. **YuNet loads with no intervention at all**, at a fixed `[1, 3, 640, 640]`, with twelve outputs in three strides — `cls`/`obj`/`bbox`/`kps`, which is a *different layout* from SCRFD's and confirms why §4.2's load-time check has to look at shapes rather than count outputs. Combined with its permissive licence (§2.3) that makes M9 more interesting than it looked: the permissive detector is also the one with no shape-fixing step in front of it. ### 12.2 First end-to-end run · 2026-08-26 The Rust port produces the separation it is supposed to. Three distinct portraits of one identity against two of another, from §1.1's labelled gallery: | Pair | Cosine | |---|---| | Same identity, different photographs | **0.596** | | Same identity, byte-identical duplicate files | 1.000 | | Different identities | **0.049 – 0.050** | Against the reference's fitted MBF boundary of cos 0.267 (§1), 0.596 and 0.05 fall either side with room to spare — which is the check that the port's pre-processing, letterbox inversion and alignment are right, since any of them being wrong degrades the same-identity number first. The duplicate row is not a curiosity: several files in that gallery are byte-identical under different names, which is exactly the case §8.2's dedup step exists for, and it would otherwise stack the positive histogram at cos ≈ 1 with pairs that teach the fit nothing. **M2/M3, provisionally, on the reference desktop:** detection **~1.0–1.4 s** per image, embedding **~160–280 ms** per face, model load 80–90 ms for the pair. Slower than the extrapolation in §12's M2 row, which guessed SCRFD-500M would come in under YOLO26n-seg's 470 ms — tract is not ORT, and this is what it costs. A 17k-image library at ~1.3 s plus ~1.5 faces each is on the order of **7 hours** of background indexing. That is survivable for a resumable, preempted background job (FR-CULL-8) and it is not survivable on a phone, so NFR-RES-2 needs its own measurement rather than an extrapolation from this one. Not yet profiled: how much of the detection second is tract and how much is the scalar letterbox resample in front of it. --- ## 13. Order 1. **M1** — load both graphs in tract, freezing the input dims if needed (§4.1). Go/no-go, and it now comes first: it is an afternoon, and §2.3's reading is wasted effort if the graphs will not run at all. This is a change from the S14 brief's ordering, made because §1.1 removed the uncertainty that justified reading first — the models are known to work, so the open question is the runtime, not the choice. 2. **Licence reading** (§2.3), in parallel with 3. Before anything *ships*, per D13. 3. `dr-face` skeleton, `detect` + `align` + `embed`, ported from §1.1's table, with an example binary that draws boxes and landmarks on a JPEG — the same shape as `dr-segment`'s `examples/detect.rs`, and for the same reason: the thing worth looking at is whether the landmarks land on a real photograph. Port `tests/test_face_utils.cpp`'s cases first; they are model-free and they fail loudly on exactly the mistakes §5 describes. 4. M2–M5 on the corpus. Alignment is validated here, before anything depends on embedding quality. 5. `calibrate` + `cluster` against the labelled subset. M6–M8. The Platt fit ports from `gallery_calibration.hpp`; the pair *sourcing* (§8.1) is new and is the part to get wrong. 6. The v5 migration, the `DetectFaces` job, the debounced clustering pass. 7. UI: People view, confirm and reject, merge and split, the `Person` selector term. 8. The route-C first-run flow, the about-screen entry, and the delete-all control — with the feature shipped disabled until they exist. Steps 1–5 are the spike. Steps 6–8 are the build, and they are only justified if §12 says so. --- ## 14. What this does not settle - **Whether a permissive model pair exists.** §2.3. If it does, route B replaces route C and step 8's first-run flow shrinks to nothing. - **Whether tract runs these graphs at all.** §12 M1, and §4.1 says why the answer is in real doubt. Every other line of this document is conditional on it. - **Faces in trashed images.** catalog.md §10.4's open question, unchanged: probably excluded from suggestions but not deleted, so a restore does not re-index. - **Whether embeddings sync.** NFR-SEC-5 permits it, opt-in, separately consented. Nothing here depends on the answer. - **Re-embedding at higher resolution.** §7's `crop_px` makes it a query rather than a full re-index, but whether it is worth doing is an M4 question. - **Approximate nearest neighbours.** §9 says brute force until measured otherwise. A 100k-face library is where this stops being true. --- ## 15. Register entries **D13** — *face inference runtime and model licensing* · the runtime half stays answered (`ort` + `ort-tract`, unchanged since 2026-08-21). **The licensing half is answered conditionally by §2:** route C — ship the code, not the weights — unblocks the build without breaking GPL-3.0 or any distribution channel, and §2.3's reading may yet convert it to route B. Closing D13 outright waits on that reading. **D17** — *face model pair and distribution route* · **PROPOSED**. SCRFD-500M for detection, MobileFaceNet/ArcFace for embedding, weights obtained by the user rather than redistributed, with a pinned checksum and a licence notice shown before the first fetch. The alternative considered and rejected is committing the InsightFace weights (§2.1 route A), which is not available at any price. The model *pair* is better evidenced than a proposal usually is — §1's table is a measured comparison over a 2,418-identity gallery, not a literature reading — so the open half of this decision is the route, not the pair. Relates to: D13, D14, NFR-COMPAT-2, FR-CULL-8. **D18** — *porting from `scene-actor-extraction`* · **PROPOSED, and the easy half of a decision**. That project is MIT and by this project's author, so MIT-into-GPL-3.0-or-later is a compatible one-way combination requiring attribution, not permission. Ported files carry a header naming the origin. Worth recording because "we already have this working in another language" is exactly the provenance that goes undocumented and then cannot be answered three years later. **S14** — *face pipeline in Rust* · scope sharpened, and **substantially de-risked**, by this document. §12 replaces "measure per-image latency, cluster purity and calibration convergence" with ten specific measures. Two changes to the brief itself: its instruction to resolve the licence question *before writing any of it* is relaxed to "before shipping any of it", because M1 is an afternoon and is worth knowing first (§13); and its central question — whether the pipeline works at all — is largely answered in advance by §1.1, leaving the runtime and the domain shift from film frames to family snapshots as the real unknowns. --- ## 16. Requirements touched | ID | How this document addresses it | |---|---| | FR-CULL-8 | §4 detection, §6 embedding, §7 the proxy tier and its consequences, §10 the `DetectFaces` job | | FR-CULL-9 | §8 — pairs, fit, validity, and the reliability-diagram acceptance test | | FR-CULL-10 | §9 constrained agglomeration, confirmations as anchors, split by re-agglomeration | | FR-CULL-11 | §10 the `Person` selector term, confirmed-only by default | | FR-CULL-12 | §10 schema unchanged from catalog.md §10.1: embeddings derived, names to the sidecar | | NFR-SEC-5 | §11 — the obligations restated as structural properties, one of them CI-checkable | | NFR-COMPAT-2 | §2 — why the obvious weights cannot ship, and what does instead | | NFR-RES-2 | §1 the cheap model pair, §12 M2/M3/M10 on a phone | | NFR-ARCH-2 | §10 background priority, preempted by visible work |