Every face stored before its quality was kept holds a unit vector, and V14 forgot the run marker of each image holding one so that the next sweep would look again. Looking again meant detecting again: a whole re-detection per image, with every suggestion on it thrown away and the confirmations carried across by box overlap, to recover one number. The sweep now has a measuring pass between the proxy repair and the un-indexed images. It lists every image holding an unmeasured face, fetches the original once, warps each stored face from the landmarks it already has, embeds it, and writes the raw vector and its length over the old row. Ids, boxes and identities are untouched; the marker is re-written fresh so the sync exports the measured vectors. A face whose landmarks no longer make a warp is dropped, as detection would have refused to store it. `faces_unindexed` leaves those images to the measuring pass, so the V14 deletion no longer costs a second detection.
74 KiB
Faces and identity — SCRFD and MobileFaceNet
Spec for S14, the spike that decides whether §3.9.1 is buildable, and the build that follows it.
FR-CULL-8 … FR-CULL-12 specify the subsystem in terms of "a 512-dimension embedding from a stated model" and stop there, deliberately: catalog.md §10.4 records that the model was left abstract because D13 was open. This document names the two models, fixes the pre- and post-processing they need, specifies the calibration and clustering that sit on top, and states what S14 has to measure before any of it is trusted.
It does not close the licensing half of D13. §2 is the reason, and it comes first because segmentation.md §7 established the precedent that reading the grant is cheaper than discovering it at packaging time.
1. The two models, and why these two
Detector — SCRFD. Sample and Computation Redistribution for Efficient Face Detection (Guo et al., ICLR 2022). A single-stage anchor-based detector whose contribution is where the FLOPs go, not a new architecture: the shallow stages carry more capacity because that is where small faces are decided. It emits, per face, a box, a confidence, and five landmarks in the same forward pass — which is the property that matters here, because FR-CULL-8 requires landmarks and because §4's alignment step is not optional. A detector without landmarks would need a second network to supply them.
The 500M variant is the one to ship: 500 MFLOPs at VGA, against 10 GFLOPs for 10G. Face indexing
is a background sweep over a whole library on a phone as well as a desktop (NFR-RES-2), so the cheap
variant's accuracy loss on tiny faces is the right trade — a face too small for 500M to find is
also too small for §5 to embed usefully.
Embedder — MobileFaceNet trained with ArcFace loss ("MBF"). ~1M parameters, ~0.45 GFLOPs for one 112×112 crop, 512-d output. ArcFace's additive angular margin is what makes the output space one where cosine similarity means something — the loss explicitly optimises angular separation between identities, which is what §6's calibration then has to convert into a probability.
Why not the larger ResNet50 embedder (w600k_r50, the other half of InsightFace's buffalo_l):
because it is not measurably better here, and it is 13× the file. §1.1's reference project fitted the
same Platt calibration (§8) to all four candidates over a 2,418-identity gallery, which is a clean
read on the embedding space alone with no tracking or matching logic in it:
| Model | File | Steepness a |
Boundary at P=0.5 | Held-out macro F1 |
|---|---|---|---|---|
| LVFace-B (Glint360K) | 455 MB | 17.7 | sim 0.228 | 67.4% |
| ArcFace w600k-MBF | 13 MB | 16.2 | sim 0.267 | 64.4% |
| ArcFace w600k-R50 | 174 MB | 15.4 | sim 0.301 | — |
| ArcFace R18 | 48 MB | 15.3 | sim 0.309 | 63.1% |
MBF's calibration curve is steeper than R50's and R18's, separating same-identity from different-identity pairs more confidently and at a lower boundary, at a fraction of the size. LVFace-B is genuinely better and is 35× the file — a defensible choice for a desktop-only feature and not one for a subsystem that has to index a library on a phone (NFR-RES-2).
Two caveats on that table, both from the source: the R50 gallery was built with ~30% fewer reference images per identity, which is why it carries no F1 figure — the calibration row does not depend on the image count and is comparable, the benchmark row would not have been. And the F1 column measures a film-cast identification task, not this one. It ranks the models; it does not predict DarkRoom's accuracy.
Both are pure convolutional graphs, which matters for §3: the runtime is tract, and tract's operator coverage is the thing that decides whether a graph runs at all.
1.1 A working reference implementation exists
../scene-actor-extraction is a C++ pipeline by the same author that runs exactly this model
pair — SCRFD-500MF then ArcFace over 5-point-aligned 112×112 crops — over feature films, with a
fitted Platt calibration and a scored benchmark behind it. It is MIT licensed, so porting from it
into GPL-3.0-or-later is straightforward and needs attribution, not permission.
That changes what S14 is. The open questions are no longer "does this pair work" and "what are the magic numbers" — they are does tract load these graphs (§12 M1, and §4.1 says why that is in genuine doubt) and does the accuracy hold on family snapshots rather than film frames. The pre-processing constants, the decode layout, and the calibration algorithm are all readable rather than rediscoverable, and several of them are not what the published Python would lead you to write.
| What | Reference | Ports to |
|---|---|---|
| ArcFace 5-point template and warp | src/face_utils.hpp |
align.rs (§5) |
| SCRFD pre-process, decode, NMS | src/backends/ort_backend.cpp SCRFDDecoder |
detect.rs (§4) |
| ArcFace pre-process and L2 normalise | same file, ArcFaceEmbedder |
embed.rs (§6) |
| Platt fit over a similarity histogram | src/gallery/gallery_calibration.hpp |
calibrate.rs (§8) |
| Model-free unit tests for both | tests/test_face_utils.cpp, tests/test_calibration.cpp |
the §3 feature split |
That last row is worth noticing: the reference already separates the geometry and arithmetic from the
inference well enough to unit-test them with no model on the machine. §3's inference feature flag is
the same boundary, enforced by Cargo rather than by discipline.
What does not port. The reference matches faces against a known gallery of named identities;
DarkRoom clusters an unlabelled library (§9). And it is a video pipeline, so it has a tracker, a
max_faces cap, and frame-to-frame identity annealing that have no analogue here.
2. Licensing — the open half of D13
The architectures are published research. The weights everyone actually uses are not redistributable by this project.
InsightFace's code is MIT. Its pretrained models — buffalo_l, buffalo_s, buffalo_sc, and
every det_* and w600k_* checkpoint inside them — carry a non-commercial research-only grant,
stated in the model-zoo README and applying to both the auto-downloaded and the manually downloaded
artefacts. w600k_mbf is trained on WebFace600K, whose own terms are research-only as well, so the
restriction has two independent sources rather than one that might be renegotiated.
These are the exact artefacts in question — §1.1's reference project fetches them from InsightFace's own release page, which is where the grant attaches:
| File | Source |
|---|---|
det_500m.onnx (2.5 MB) |
insightface/releases/download/v0.7/buffalo_sc.zip |
w600k_mbf.onnx (13 MB) |
insightface/releases/download/v0.7/buffalo_s.zip |
DarkRoom is GPL-3.0-or-later. A non-commercial field-of-use restriction is not a GPL-compatible grant, so:
- these weights are not redistributable under this project's licence, so they cannot go into a published build — an APK, a Flatpak, or an F-Droid entry — and could not go into a repository intended to be one (§2.2a records what was actually decided here, and why the two differ);
- and the restriction binds the user, not only the project — a professional photographer using DarkRoom commercially is outside the grant even if they fetched the file themselves.
That last point is the one that is easy to miss and dishonest to leave implicit. It is why §2.2's recommended route puts the licence text in front of the user rather than in a footnote.
2.1 The routes, priced
| Route | What ships | Cost |
|---|---|---|
| A — commit the weights | Everything works out of the box, one git lfs pull |
Not available. Licence-incompatible with GPL-3.0 and with all three distribution channels (NFR-COMPAT-2). Listed only so it is visibly rejected rather than quietly assumed. |
| B — permissively licensed weights | Same, legally | Requires weights that do not exist off the shelf in this architecture pair. Either found (§2.3) or trained, and training an ArcFace embedder needs a face dataset whose own terms permit redistribution — the harder half of the problem. |
| C — the user fetches them | The app ships the code, not the weights; face indexing is off until the user obtains a model | Works today, keeps the repository clean, and is honest. Costs a first-run flow, an F-Droid anti-feature declaration, and a licence notice the user has to read. |
Recommendation: C now, B when it becomes possible. §1.1's project already works this way — a
scripts/download_models.sh that fetches the weights rather than a repository that carries them —
so route C is a practised pattern here rather than a theory. Two things DarkRoom must do that the
reference does not: pin a checksum for every model (the reference verifies only one of the four),
and put the licence in front of the user, because a photo editor's users are not all researchers.
The schema already forces this to be a survivable choice — faces.model_id (catalog.md §10.1) exists precisely so that a model change is
detectable and re-indexable rather than silently poisoning every similarity in the library. Under C,
swapping in a permissive model later is a new model_id and a re-index, not a migration.
2.2 What route C requires of the code
- The weights are never a build input.
dr-face(§3) takes bytes; it has noembedded-modelfeature and nomodels/directory. This is the one structural difference fromdr-segment, and it is deliberate — a feature flag that could embed weights is a feature flag someone eventually turns on in a packaging script. - The fetch does not live in
dr-face. It lives in the app layer, so the inference crate keeps no network dependency at all. §11 makes that a checkable property rather than a convention. - The licence is shown, not linked. Before the first download, the app states in plain language that the weights are research-only, that commercial use is outside the grant, and who the grantor is. NFR-SEC-5's last bullet already requires the models to be named, versioned and licensed in the about screen; this is the same obligation moved to the moment where it can still change a decision.
- A checksum is pinned. The app verifies the digest of what it fetched against a value compiled in, and refuses a mismatch. An unverified model file is an arbitrary graph executed over the user's family photographs.
- Indexing stays off until a model is present and verified, and disabling it deletes nothing — that is a separate control (NFR-SEC-5, §11).
2.2a Android, which route C forgot
Route C says "the user obtains the model and the app loads it". On a desktop that is a real gesture:
the files go in ~/.local/share/darkroom/models/ and face indexing starts working. On Android that
gesture does not exist. internal_data_path is app-private, adb shell run-as needs a debuggable
build, there is no picker and no fetch in the app, and so a phone could not acquire a model by any
means at all. Face indexing was not "off until the user supplies weights" there; it was off, full
stop, and the settings page said so on every launch with no action available behind the message.
Decision, 2026-08-27: the shape-fixed pair is committed to LFS at
apps/darkroom-android/android/assets/models/, and assemble-apk.sh bundles it into the APK.
android_main unpacks it to the shared models directory on first launch, before anything asks
whether a model is present. This is a personal project on a private Gitea; the grant restricts
redistribution, and a private repository and a self-installed APK are not that.
What that does and does not settle:
- It does not make §2.1's route A available. The moment this project publishes — an F-Droid entry, a release APK, a Flatpak — these files come back out and route C's unbuilt half (the fetch, the licence screen, the pinned checksum) has to exist first. §2.1's table stands as the answer for a published build; this is the answer for the author's own phone.
- It does not make the weights a build input.
dr-facestill has nomodels/directory and noembedded-modelfeature, and nothing in the cargo build reads these files — §2.2's first bullet guards against a flag someone flips in a packaging script, and that remains guarded. The APK assembly step copies two files; it is the only thing in the tree that knows they exist. - It does not bind only the project. §2's last bullet is unchanged and is the one with teeth: the research-only grant restricts the user, so commercial photography with a DarkRoom that has these weights in it is outside the grant no matter who put them there.
The in-app fetch, the licence screen, and the pinned checksum that §2.2 specifies are still unbuilt, on every platform. When they are built, this becomes redundant and the assets directory empties.
2.3 Before S14 writes any code
Three questions to answer by reading, in this order, and to record with the date they were checked — the same discipline segmentation.md §7 applied:
- Is there a permissively licensed SCRFD export? Third-party ONNX exports of SCRFD are plentiful; a re-export inherits the original weights' grant, so the question is whether anyone has trained SCRFD on a redistributable dataset, not whether they have converted the InsightFace one.
- Is there a permissively licensed 512-d MobileFaceNet/ArcFace checkpoint? Candidates worth reading the terms on rather than assuming: EdgeFace (Idiap), the ONNX Model Zoo's ArcFace entry (MIT, but a ResNet100 at ~250 MB — licence-clean and the wrong size), and any MobileFaceNet trained on a dataset with distribution terms.
- What does a permissive detector alternative cost in accuracy? YuNet (OpenCV Zoo) is 232 KB, emits five landmarks, and comes from a permissively licensed repository. If it is close enough, the detector half of the licence problem disappears and only the embedder remains — a materially better position than either half alone.
Question 3 is nearly free to answer: face_detection_yunet_2023mar.onnx is already sitting in
§1.1's models/ directory, fetched from opencv/opencv_zoo. It needs its own decode path (its
output layout is different, which is why the reference's SCRFDDecoder explicitly rejects it at load
rather than silently misreading it), and then it is one more row in §12's table.
3. Crate shape — core/dr-face
A new workspace member, modelled on dr-segment and for the same reason: everything that reasons
about faces is testable with no GPU adapter and no model present (ARCH §6.5a).
core/dr-face/
src/lib.rs FaceError, re-exports
src/detect.rs SCRFD: preprocess, decode, NMS
src/align.rs 5-point similarity transform → 112×112 crop
src/embed.rs MBF: preprocess, forward, L2 normalise
src/calibrate.rs cosine → P(same person) (FR-CULL-9)
src/cluster.rs constrained agglomeration (FR-CULL-10)
src/assign.rs which person, and how sure (§9.1, FR-CULL-9)
[features]
# No `embedded-model`. §2.2 — the weights are not a build input, ever.
default = []
# The ONNX runtime. Off by default so `calibrate` and `cluster` — which are
# arithmetic over embeddings and have no model in them — stay testable in a
# build that carries no inference engine at all.
inference = ["dep:ort", "dep:ort-tract", "dep:ndarray"]
The ort + ort-tract pairing is settled by D13's 2026-08-21 update and already in the workspace
manifest: ort's API, tract's pure-Rust engine, no C under the NDK.
The split between inference and the rest is load-bearing. Calibration and clustering are where
the subsystem's accuracy actually lives, they are pure functions of embeddings and labels, and they
must be testable against synthetic embedding sets with no weights on the machine. A test suite that
needs a research-licensed download to run is a test suite that does not run in CI.
3.1 The API
/// A loaded detector. Fixed input shape — see §4.
pub struct Detector { /* session, input edge */ }
impl Detector {
pub fn from_bytes(onnx: &[u8], id: ModelId) -> Result<Self, FaceError>;
/// `rgb` is f32 0..=1, row-major, three per pixel.
pub fn detect(&self, rgb: &[f32], w: usize, h: usize, opts: &DetectOptions)
-> Result<Vec<Detection>, FaceError>;
}
pub struct Detection {
/// Normalised to the image's long edge (catalog.md §10.1).
pub bbox: (f32, f32, f32, f32),
/// Five points, same normalisation, in the model's own order (§5).
pub landmarks: [(f32, f32); 5],
pub confidence: f32,
}
pub struct Embedder { /* session */ }
impl Embedder {
pub fn from_bytes(onnx: &[u8], id: ModelId) -> Result<Self, FaceError>;
/// Takes an *aligned* 112×112 crop. Alignment is `align::warp`; passing a
/// raw bbox crop here is the silent-accuracy-loss bug §5 exists to prevent.
pub fn embed(&self, aligned: &Aligned112) -> Result<Embedding, FaceError>;
}
/// L2-normalised, 512-d. Carries its model id so a comparison across models
/// is a type error rather than a plausible-looking number (catalog.md §10.1).
pub struct Embedding { pub model: ModelId, pub v: [f32; 512] }
Aligned112 is a newtype over the pixel buffer that only align::warp can construct. That is the
whole defence against §5's failure mode, and it costs nothing.
4. Detection
4.1 Fixed input, and how the image is fitted to it
det_500m.onnx has a dynamic H/W input, and this is the largest single risk in the document.
§1.1's reference had to use ONNX Runtime rather than OpenCV's dnn module precisely because OpenCV
could not load SCRFD's dynamic Shape nodes. tract is in the same family of problem: dr-segment
exists in its current shape because tract failed shape inference on YOLO's dynamic export, which is
why tools/export-seg-model.sh passes dynamic=False and why semantic.rs has a fixed
INPUT_EDGE.
So the graph almost certainly needs its input dimensions frozen before tract will touch it. That is a
known, cheap operation — onnxruntime.tools.make_dynamic_shape_fixed rewrites the declared dims
without retraining or re-exporting from PyTorch — but it is a step, it produces an artefact that must
itself be reproducible (a tools/fix-face-model-shapes.sh, in the spirit of the existing export
script), and it must be tried before anything else in this document is scheduled. §12's M1.
Fixed at 640, with 320 available as a faster, blinder option to measure.
Letterbox. Scale by min(640/w, 640/h) preserving aspect, then paste into a 640×640 canvas. The
reference centres the image and fills the margin with grey 114, inverting with
x_src = (x_model − pad_x) / scale. What matters is not where the padding goes but that the
forward and inverse agree and that the fill value is treated as part of the contract: a mismatch
between them offsets every box and landmark the model returns by the padding, which yields detections
that look plausible, landmarks that align badly, and an embedding-quality loss that only shows up as
poor clustering three stages later. Port the reference's convention rather than inventing one, and
assert it with a round-trip test.
Normalisation is (x·255 − 127.5) / 128, RGB, NCHW — note /128, not /127.5; see §6.
4.2 Outputs, and decoding them
Nine tensors for three strides {8, 16, 32}, twelve for four {8, 16, 32, 64}. Which one a
given export produces is discovered at load — strides = output_count / 3 — not assumed, because
both variants exist and hardcoding three silently ignores the largest faces in a twelve-output model.
With two anchors per location and N_s = (640/s)² locations:
| Tensor | Shape | Meaning |
|---|---|---|
score_s |
[N_s·2, 1] |
sigmoid already applied inside the graph |
bbox_s |
[N_s·2, 4] |
distances left, top, right, bottom in units of the stride |
kps_s |
[N_s·2, 10] |
five (dx, dy) offsets, same units |
Anchor centres are (x·s, y·s) for each grid cell, repeated once per anchor. Decoding is therefore
x1 = cx − l·s x2 = cx + r·s kx_i = cx + dx_i·s
y1 = cy − t·s y2 = cy + b·s ky_i = cy + dy_i·s
then divide by the letterbox scale to return to source pixels, then normalise by the long edge before storage.
Flat index within a stride is (row · fw + col) · 2 + anchor.
A shape assertion at load time, not a decode-time surprise. dr-segment already learned this —
SegmentError::OutputShape exists because a different export of the same model produces confidently
wrong results otherwise. The reference implements exactly this and it is worth porting verbatim:
output count divisible by three and between 9 and 12, then the last dimension of each group checked
against {1, 4, 10}. That single check is what catches a YuNet file passed where an SCRFD one was
meant — YuNet also has twelve outputs, so the count alone does not distinguish them, and the failure
without the check is not an error but a page of plausible garbage.
4.3 Thresholds
Score ≥ 0.5, NMS IoU 0.4, plain greedy NMS per image across all strides together.
Deliberately not the low threshold dr-segment chose. There, a false positive costs one spurious
entry in a list the user is picking from. Here it costs an entry in the People view that the user has
to reject, in a library with thousands of images — and worse, a garbage embedding that participates
in clustering and can bridge two real clusters into one. False negatives are recoverable by a later
re-index with a better model; a polluted cluster graph is not, once the user has confirmed faces
inside it.
A minimum box size of 40 px on the source proxy is applied on top — the reference's figure, chosen there for the same reason it holds here: below it there is not enough face left to align reliably, and §7's table shows what the embedder actually receives at that size. §12's M4 is what moves this number, if measurement says it should move.
One thing not to port: the reference caps detections at ten per frame, largest first. That is right for a film frame, where the extras in the background are noise. It is wrong for a photo library, where a group shot with thirty faces in it is precisely the picture the user wants indexed. No cap; the min-size floor is the only filter.
5. Alignment — the step that is silently wrong when skipped
ArcFace embeddings are trained on faces warped to a canonical 112×112 arrangement. Feeding the model a plain bounding-box crop works — it produces 512 numbers, they are unit-norm, and cosine similarities between them look entirely reasonable. They are just much worse, and nothing in the system reports it.
The transform is a similarity transform — rotation, uniform scale, translation, four degrees of freedom — from the five detected landmarks to this template, which is the arrangement the weights were trained against:
(38.2946, 51.6963) subject's right eye ─┐ image-left of centre
(73.5318, 51.5014) subject's left eye ─┘
(56.0252, 71.7366) nose tip
(41.5493, 92.3655) subject's right mouth corner
(70.7299, 92.2041) subject's left mouth corner
Similarity, not affine: a full affine fit to five noisy points will happily shear a face, and shear is exactly the deformation the embedder was never shown.
The naming is a trap and the order is not. Point 0 sits at x=38 on a 112-wide canvas — left of
centre in the image, which is the subject's right eye. Both namings are in circulation and they
are opposite. What matters is that SCRFD emits its five points in this same order, so the correct
amount of reordering between detector and template is none; the reference states this explicitly
in types.hpp for the benefit of whoever next reads it and doubts it. A future detector with a
different order carries its own permutation next to its model_id, rather than this file growing an
assumption.
One divergence from the reference to settle by measurement. It fits the transform with OpenCV's
estimateAffinePartial2D under RANSAC at a 3-pixel threshold, and drops the detection when the
fit fails. RANSAC over five points is a strange fit — the minimal sample for a similarity is two, so
it can discard landmarks it judges outliers and solve from a subset, which on a profile face is as
likely to be the correct geometry as the wrong one. InsightFace's own pipeline uses a plain
least-squares (Umeyama) similarity over all five points, which cannot silently drop anything and is
the one to write first. §12's M5 measures whether the choice matters; the degenerate-input path still
has to exist either way, because collinear landmarks do occur.
A trick worth keeping. The reference retries detection on a frame where nothing was found, after replicate-padding by 25% and applying CLAHE to the L channel. Padding gives a face at the very edge a context the detector needs; CLAHE rescues underexposed frames. Both are plausible on scanned negatives and backlit photographs — worth trying at low priority, since a retry doubles the cost of exactly the images that yielded nothing.
Sampling is bilinear from the native render, in one step — never crop-then-warp, which resamples twice and throws away detail the warp could have used, and never from the downscaled buffer the detector was given. The landmarks arrive in detector-input coordinates and are scaled back to native before the warp reads a single pixel; §7 is why. Pixels falling outside the source are black.
The landmark order must match the template order. The template above is written in the detector's own output order; if a future detector emits them differently, the template is reordered with it and that mapping belongs next to the model id, not compiled in as an assumption.
Test: warp a synthetic image with a known rotation and scale, and assert the five points land on
the template within a fraction of a pixel. This is testable with no model present, which is why
align sits outside the inference feature.
6. Embedding
112×112 RGB — the crop comes out of the warp in whatever order the source was in, and the reference converts BGR→RGB explicitly before the blob rather than relying on a flag, which is the readable way to do it.
Normalisation is (x·255 − 127.5) / 128. Note /128, not /127.5: InsightFace's published
Python uses 1.0/127.5 for the recognition model, the reference uses 1.0/128 for both models, and
every measured number in §1's table was produced with /128. The difference is 0.4% of scale and
almost certainly immaterial, but "almost certainly" is not a reason to pick silently — write /128 to
match the numbers we have, and settle it with one back-to-back run in §12.
Output is 512 floats. Every downstream comparison is a dot product over the unit vector, and the reference L2-normalises before storing so that no code path has to remember to. The reference clamps the norm at 1e-6 before dividing, which costs nothing and removes a NaN path.
Keep the length. The norm the normalisation divides out is not noise. ArcFace trains the
direction of its output and nothing else, and the magnitude it leaves behind grows with how much of a
face the model could make out — MagFace (Meng et al., CVPR 2021) made that the training objective,
and the plain ArcFace heads this crate runs already show it, weaker but usable. A blur, an occlusion,
a hard profile or a badly lit crop comes out short. On the reference library w600k_mbf's norms run
from about 8 on a blur to the high 20s on a clean portrait.
So the store holds the raw vector, not the unit one — dr_face::Embedded::to_f16_bytes — and
readers re-normalise on load, which they had to do anyway (below). f16 keeps the same three figures
of a component whatever the vector's length, so this costs nothing in precision. The length is also
kept beside the blob as faces.quality, for the readers that never load the vector (the People
screen), and it is NULL for a face stored as a unit vector before this — a unit vector reads as a
length of one, and one is not "unmeasured".
What the number does is in §9: a face whose quality is under MIN_GALLERY_QUALITY = 14 is still
placed, but is never what another face is compared against. The screen shows it as "Quality 17.3",
dimmed below the floor, so a user asking why a group did not gather the rest of a person can see that
none of its members can vouch for anyone.
Some exports of these graphs are fp16 in, fp16 out. The reference detects this from the graph's declared element type rather than from the filename; worth porting, because the alternative failure is a silent garbage tensor.
Storage is 512 × f16 (catalog.md §10.1) — 1 KB per face, 25 MB for a 25,000-face library. The f16
round-trip perturbs a unit vector by ~1e-3 in cosine, three orders of magnitude below the separation
between a match and a non-match, and the halving matters because these rows are the ones NFR-SEC-5
contemplates optionally syncing.
Re-normalise on load after the f16 widen. It is one pass over 512 floats and it removes a class of drift that is otherwise invisible — and, since the blob is raw, it is what turns the stored vector back into the unit one every comparison expects.
7. Which pixels the pipeline actually sees
This section previously specified the proxy tier — ThumbSize::Large, 1024 px — as the buffer
both detection and cropping read. That was wrong, and §7b is the measurement that says so.
FR-CULL-8 now separates the two, because they want opposite things:
| Stage | Resolution | Why |
|---|---|---|
| Source render | native | The only stage where more pixels exist to be had |
| Detector input | downscaled to ~640 | §4.1 letterboxes to 640×640 regardless; more is wasted CPU |
| Crop + align | sampled from the native render | The 112×112 is fixed, so this decides whether it holds real pixels |
| Embedding | 112×112 | §6 |
The detector's indifference to resolution is the whole reason the split works. §4.1 fixes its input at 640×640 and letterboxes whatever arrives, so a face occupying 2% of the frame presents at 12 px to the model whether the buffer handed over is 1024 px or 6000 px. Feeding it native pixels buys nothing. Feeding the crop native pixels buys everything, because §5's warp is the one place where source resolution converts directly into embedding quality.
What the crop receives, by face size, from a native render of a 24 MP frame (~6000 px long edge):
| Face size in frame | Pixels across, native | Pixels across, 1024 proxy | What §6 receives |
|---|---|---|---|
| A portrait, face fills a third of the frame | ~2000 | ~340 | Downsampled. Ideal either way. |
| Two people, half-length | ~700 | ~120 | Native: comfortable. Proxy: marginal. |
| A group of eight | ~290 | ~50 | Native: real pixels. Proxy: upsampled 2.2×. |
| A figure in a landscape | ~120 | ~20 | Native: usable. Proxy: below §4.3's floor. |
So faces records one column beyond catalog.md §10.1's schema:
ALTER TABLE faces ADD COLUMN crop_px INTEGER NOT NULL; -- source pixels across the aligned crop
crop_px earns its place four times over. It is the honest quality signal for the UI; it is a
feature in §8's calibration, which FR-CULL-9 explicitly demands ("a raw cosine means something
different for every model, every population, and every face size"); it is what a
higher-resolution re-embedding pass selects on, so that pass is a query rather than a re-index of
everything; and it is the only way to audit whether the rule above is actually being followed —
which is how §7b found that it was not.
7a. Recording that detection ran
The spec above assumed the faces table could answer "has this image been indexed". It cannot,
and the difference is the one that decides whether a background pass ever finishes.
A photograph with no face in it produces no rows. So does one that has never been looked at. Asking
faces therefore re-queues every landscape, still life and document scan on every pass, for ever —
and in a personal library that is most of it. Measured on the reference library: of the first 110
images indexed, 64 contain no face at all.
So face_index records the run: one row per (image, model) carrying the timestamp, the number of
faces found — zero is the interesting value — and the long edge of the proxy it read. Keyed on the
model, so a model change puts every image back in the queue without anyone having to remember to
clear anything.
Three things fall out of it that were not otherwise available:
- A coverage figure. "4,812 of 5,000 indexed" is what a user wants to see; counting face rows can only ever report how many faces exist, which is a different number that never reaches the image count.
- A reason for the ones outstanding. The audit splits them by whether a proxy exists, because waiting on the thumbnail sweep and waiting on face indexing are different problems and only one of them is fixed by running this again. On the reference library the first check reported 110 ready and 23,417 awaiting a proxy — which is the real state of that library, and not something the face subsystem can do anything about.
- Something to sync. §14's shards carry the marker with the faces, so an adopted image is not re-detected on the receiving device.
7b. What the proxy tier actually cost, measured
The table in §7 predicted upsampling for small faces and called it "degraded, and usable". On the
reference library of 23,531 images it was not the edge case that description implies. Reading
crop_px across the 18,671 faces stored under the proxy-tier implementation:
| Source pixels across the aligned crop | Faces | Share |
|---|---|---|
| <56 (upsampled more than 2×) | 314 | 1.7% |
| 56–111 (upsampled) | 8,505 | 45.6% |
| 112–223 (roughly native) | 6,313 | 33.8% |
| ≥224 (downsampled — ideal) | 3,539 | 19.0% |
47.3% of every face in the library was upsampled to reach the embedder, with crop_px as low
as 34 — a 3.3× enlargement — against a mean of 178. An upsampled crop does not fail loudly. It
produces a confident 512-d embedding describing detail that was interpolated rather than
photographed, and the damage appears three stages later as clusters that will not separate.
Measured again, against the implementation
The figures above are read out of a catalog after the fact, so they describe what the old code did
rather than what the new code does. examples/face_native.rs renders one file and indexes it both
ways, so the difference can be attributed to the resolution and nothing else. Fourteen originals
from the reference library — 5472×3648 Canon CR2 and DNG — each rendered once and indexed twice:
| Path | Faces found | Mean crop_px |
|---|---|---|
| Native (this specification) | 9 | 287 |
| Everything from a 1024 proxy | 5 | 75 |
Crops 3.8× larger, and on the right side of the line that matters: 75 px is below [ALIGNED_EDGE]
so the proxy path was upsampling into the embedder on average, where the native path downsamples
into it.
Sample of fourteen images and nine faces. Enough to show the direction and to catch a wrong landmark scaling, which is what it was written for; not enough to quote a ratio as the library-wide figure. §12's M4 is still where the recall curve gets established.
A second effect, recorded here because it was measured and because the mechanism is not
established. Grouping the same runs by face_index.source_edge — the buffer detection ran
against — gives 0.078 faces per image at 1024 or below, against 1.82 at 2048 or better. Controlled
for file type and size (1,592 DNGs averaging 21.0 MB against 7,724 averaging 23.3 MB, same library,
same cameras), so it is not a composition artefact. But it cannot be a matter of the detector
seeing fewer pixels, since §4.1 letterboxes both to 640: a 1024 buffer and a 3072 buffer present
the same face at the same size to the model. The likeliest explanation is that the 1024 proxy is
itself a downscale of a larger preview, so the detector sees a twice-resampled image where the
larger buffer is resampled once — but that is a hypothesis, not a finding, and §12's M4 is where it
should be settled. The crop measurement above stands on its own and does not depend on it.
The A/B is evidence for that hypothesis without settling it. Native found nine faces where the 1024 path found five, on four files where the proxy path found none at all — so detector input does affect recall, which the letterbox says it should not. The two paths differ in their resampling as well as their size (one box filter and one letterbox against a downscale and a letterbox), and this experiment does not separate those. Isolating it means holding the chain fixed and varying only the buffer, which is M4's job.
8. Calibration — cosine to probability
FR-CULL-9 makes this a hard requirement: no code path may threshold a bare cosine, every threshold in the subsystem is stated as a probability, and the fit is per library and reports its own validity. This section is how that is obtained, and the interesting part is where the training pairs come from when the user has labelled nothing yet.
8.1 Where the pairs come from
Negatives are free and abundant. Two faces detected in the same photograph are almost never the same person. That gives every multi-face image in the library a full set of negative pairs at no labelling cost — and they are hard negatives, drawn from the same camera, lighting, and processing, which is exactly the population where a threshold tuned on easy negatives fails. The known exceptions — mirrors, photographs of photographs, collages, a spliced panorama — are rare enough to be noise at this scale and worth naming so the exception is not mistaken for a bug later.
Positives, in order of trustworthiness:
- User confirmations (FR-CULL-10). Every pair of faces confirmed to the same person. The gold standard, and empty on day one.
- Burst siblings. FR-CULL-5 already groups bursts. Two faces in adjacent frames of one burst, in nearly the same position, are near-certainly the same person. Free, requires no labelling, and available immediately on any library with continuous-shooting frames in it. Their purity is an S14 measurement (§12), not an assumption — if bursts turn out to be dirtier than expected, this source is dropped and the calibration simply stays invalid for longer.
- Nothing else. Bootstrapping positives from high cosine similarity is circular — it fits the calibration to the belief it was supposed to test.
8.2 The fit
§1.1's reference already implements this and its shape should be ported rather than reinvented:
P(same | cos) = σ(a·cos + b + log_prior_odds)
Four details in it are the difference between working and nearly working.
Fit against a histogram, not against pairs. A 25,000-face library has 3×10⁸ pairs; no gradient descent is running over that. The reference buckets every pair into 200 bins over cos ∈ [−1, 1], carrying a positive and a negative count per bin, and fits the two parameters against the per-bin counts. The cost becomes the similarity matrix plus 200 numbers, and the matrix is a single GEMM (§9). This is the trick that makes a per-library fit affordable at all, and it is not obvious from the outside.
Balance the classes explicitly. w_pos = total/(2·n_pos), w_neg = total/(2·n_neg). Negatives
outnumber positives by orders of magnitude; an unweighted fit produces a well-shaped curve sitting at
the wrong height — precisely the "plausible number all the way to the user interface" failure
FR-CULL-9 describes.
The base rate is a runtime argument, not part of the fit. log_prior_odds is added at evaluation
time, so the balanced fit is stored once and the prior varies per query — the odds that two faces
drawn from a 40-image holiday album match are not the odds for a 40,000-image archive. Its inverse,
boundary_at(p), gives the cosine at which the probability crosses p, which is what turns §9's
"merge above 0.9" into an actual comparison. Baking a prior into a and b would need a refit per
context and would make the stored parameters mean something different depending on where they came
from.
Deduplicate before pairing. Near-identical embeddings (cos > 1 − 1e-7) are the same photograph counted twice; the reference drops them per identity first. In a photo library the equivalent is duplicates and virtual copies (FR-CAT-11), and leaving them in stacks the positive histogram at cos ≈ 1 with pairs that teach the fit nothing about hard cases.
A third feature this design adds. The reference fits on cosine alone; DarkRoom should carry
crop_px too:
logit P(same) = w₀ + w₁·cos + w₂·log₂(min(crop_px_a, crop_px_b))
FR-CULL-9 explicitly names face size as an axis along which an uncalibrated similarity misbehaves,
§7 shows a real library spans 40 px to 340 px of face, and crop_px is already in hand. The
minimum of the pair, because a comparison is only as good as its worse crop. It generalises the
histogram to a small 2-D grid of bins, which changes nothing structural. If M7 says the extra feature
buys nothing, drop it and match the reference exactly — but a film pipeline working from broadcast
frames had far less size variation to explain than this one does.
8.3 Validity, and saying so
Stored per catalog.md §10.3: parameters, a validity flag, and a hash of the face set fitted from.
Invalid when fewer than 200 positive pairs or 2,000 negative pairs are available, or when the reliability check fails.
Deliberately far stricter than the reference's floor of two positives and one negative. That floor is reasonable there: its pairs come from a curated gallery of labelled reference portraits, where a positive pair is trustworthy by construction. Here the positives are bootstrapped from bursts and early confirmations (§8.1) and the whole risk is fitting confidently to a handful of them. The reference also refuses to draw positives from an identity with fewer than five distinct embeddings, letting it contribute negatives only — the same asymmetry applies to a thinly-confirmed person and is worth keeping. In that state both clustering and the displayed confidences fall back to the reference implementation's fitted curve — a documented operating point, not an invention — and the People screen says so once, above the grid, rather than blanking every percentage. That is FR-CULL-9's requirement read as written: what may not happen is an untuned default presented as though it were measured on this library.
Blanking them was the first reading, and it was wrong in a way worth recording. A young library has no fit; a fit needs confirmations; confirmations are made on a screen the user ranks by confidence. Withholding the confidence until the fit exists is a deadlock in which the normal state of the feature is its degraded one.
Refit is triggered by the same debounce as clustering (§9), and when the confirmed-pair count grows materially.
Acceptance (FR-CULL-9): a reliability diagram over ten probability bins, on a held-out labelled split, with observed match rate within a stated tolerance of the predicted probability in each populated bin. A single accuracy figure is not an answer to this requirement.
9. Clustering
Not a job kind, for the reason catalog.md §10.2 gives: it is a whole-library operation with no
natural subject_id. It runs as a debounced library pass when detection has been idle and the face
count has moved materially.
The graph. For each face, its k = 20 nearest neighbours by cosine, then each candidate edge
scored through §8 to a probability.
Brute force is honest arithmetic here, and §1.1's project is the evidence rather than the estimate: it
computes the full similarity matrix as one GEMM (E · Eᵀ over L2-normalised rows) at gallery
scale, on CPU, and instruments it well enough to have compared CPU against OpenCL and found the
question worth asking rather than urgent. For DarkRoom that is a 25,000 × 512 matrix at 25 MB in and
~2.5 GB out — so the matrix is computed in row blocks, with each block reduced to its top-k and
its histogram contribution before the next is started, and never materialised whole. Once, in the
background, after an indexing sweep that took an hour. An approximate index is an optimisation to
reach for when §12 says it is needed, not before.
The kernel is most of the cost, and it was running at a tenth of the machine. The scan is
O(n²) dot products and nothing else, so its speed is the subsystem's speed. Two things were
wrong with the first one, both measured over a real 18,143-face library on a twenty-core desktop:
| scan | GFLOP/s | |
|---|---|---|
a row against every other row, &[Vec<f32>] |
4.64 s | 36 |
| tiled on the column side too, one flat buffer | 2.81 s | 60 |
| plus AVX2 + FMA | 0.86 s | 195 |
The first is memory: walking the whole embedding array once per row moves ~336 GB for that library,
where a column tile that fits in L2 is read once per tile of rows. The second is that the workspace
builds for baseline x86-64 — SSE2, no FMA — and the portable loop was not being vectorised into
even that, at 0.7 flops per cycle.
So the dot product is chosen per machine: AVX2 + FMA where is_x86_feature_detected! finds it,
NEON unconditionally on aarch64 — Advanced SIMD is in that baseline, so every Android device the
app builds for has it, and the explicit vfmaq matters because LLVM will not fuse a multiply and an
add on its own. The portable loop remains the definition the others are tested against. All three
produce the same 1,531,969 pairs.
The NEON kernel is the one a desktop cargo test never executes, so
tools/face-tests-on-device.sh runs the suite on an attached device: dr-face carries no weights
and touches no display, so its tests are a plain ARM64 binary that runs under adb shell with
nothing installed. Worth running whenever the kernels change.
Where a regroup's time actually goes, on that library, because the answer moved twice while it
was being looked at. Measured with cargo run --release -p dr-catalog --example face_confidence -- CATALOG --full, on the reference desktop and on a Honor tablet (ROD2-W09, aarch64):
| desktop, before | desktop | tablet | |
|---|---|---|---|
| scan | 4.60 s | 0.96 s | 2.61 s |
| agglomerate | 4.84 s | 1.69 s | 2.16 s |
| score | 0.23 s | 0.26 s | 0.40 s |
| total | 10.0 s | 3.1 s | 5.2 s |
The scan came down by the kernel work above. The agglomeration was not, as it looked, an irreducible
sequential heap walk: 3.06 s of it was every component scanning the whole pair list for the pairs
that were its own — 382 million set lookups to place 804,499 pairs — which the union-find can do in
one pass while it is finding the components anyway. Another 0.54 s was Engine::cross computing dot
products with the portable loop while the scan beside it used the machine's SIMD.
The two architectures agree exactly, which is worth more than either column: the same 1,531,969 evidence pairs, the same 2,518 groups holding the same 16,246 faces, and the same reliability table, from AVX2 on the desktop and NEON on the tablet. That is the cross-kernel check the unit test can only approximate.
The tablet is where a GPU GEMM would pay. Its scan is half the pass, against under a third on the desktop — twenty cores of AVX2 pull ahead of a tablet's NEON far more than the merge engine's single-threaded hashing does. So a perfect GEMM is worth about 2× a regroup there and about 1.5× here, and it is the phone and tablet story that should decide whether it gets built.
Constraints, not just thresholds:
- Cannot-link on co-occurrence. Two faces in the same image are never merged. This is the same observation §8.1 mines for negatives, used here as a hard constraint, and it is the single cheapest defence against the over-merging FR-CULL-10 warns about.
- Confirmed faces are anchors. A confirmation is user data (FR-CULL-12) and clustering never moves it. Two clusters each containing confirmations of different people cannot merge; a cluster containing confirmations of one person absorbs suggestions but never reassigns the confirmed.
- A short embedding is never a reference. The length of the raw vector is the model's own
reading of the crop (§6), and a short one sits near the middle of the sphere, matching a little of
everybody — one of those in a group is a bridge to the next group over. So the population is
split: faces at or above
MIN_GALLERY_QUALITYare the gallery and cluster as described below; faces under it are probes, each measured against the finished groups and placed in the one it fits by the same average-link rule under the same two constraints — but measured against gallery members only, never against another probe, and once placed never part of what the next face is measured against. Two probes are never paired at all, andneighboursdrops those pairs before anything downstream sees them. A probe's confidence (§9.1) is computed from the references it matched; a reference's confidence hears nothing from a probe. A face whose quality was never recorded is admitted to the gallery — a rule that cannot be checked admits rather than excludes — and the next indexing pass measures it:faces_unmeasuredlists every image holding one, and each such face is embedded again from the native render with the landmarks it already has, the raw vector written over the old one and its id, box and identity untouched (faces::record_measurements). No detector runs and no suggestion is lost — the cost is the original fetched once more, since the length exists only at the moment of embedding.
The algorithm. Constrained average-link agglomeration over the probability graph, merging while the average pairwise probability exceeds 0.9 and no cannot-link is violated. Average-link rather than single-link because single-link chains — one bad edge welds two identities together, which is the documented way face clustering fails on families.
Incremental by default. A new face joins the existing cluster whose average probability against it is highest, if that clears the threshold; otherwise it starts an unnamed group. A full re-agglomeration happens on model change, on recalibration, and on explicit user request. Suggestions are recomputed freely; confirmations survive all of it (FR-CULL-10).
Splitting re-agglomerates within one person at a raised threshold and offers the resulting groups as the split. FR-CULL-10 requires splitting to be as easy as merging, and a split that hands the user a pile of loose faces to re-sort is not that.
9.1 The number beside a suggestion
Which person a face belongs to and how sure that is are different questions, and the second one is not answered by the pairwise probabilities that settled the first.
The first implementation answered it with the mean calibrated probability between the face and the rest of its group, and that measures the wrong thing twice. It punishes coverage: a person with two hundred faces across fifteen years is supposed to have members a new photograph is orthogonal to, so the better someone is photographed the worse their suggestions score. And it never asks who else the face could be — a face matching Anna at 0.95 and nobody else, and one matching Anna at 0.95 and her sister at 0.93, come out identical, when the second is the only one the user needs to look at.
Two questions, so two factors, multiplied:
evidence(P) = Σ of the top n of { P(same | this face, f) : f ∈ P } n = 10
coherence = evidence(own) / how many of the top n there were
uniqueness = evidence(own) / (evidence(own) + Σ evidence(named rivals))
confidence = coherence × uniqueness
Coherence is the old mean with a cap on it, and the cap is the whole fix: the two hundred faces a given photograph is legitimately orthogonal to stop counting against it. Uniqueness is the competition, and it is what makes an ambiguous face read as ambiguous — two identities matching equally well land at 0.5 each, which is the truth about a sibling.
Only named people compete, and they compete per person. This is the part that had to be measured rather than reasoned about. Normalising across every group made the number useless on a real 18,000-face library — median suggestion 21%, four in five under half — because clustering leaves one person spread across many groups, so a face competes against itself. Counting only the groups the user has ruled on — one holding a confirmation, a name, or an ignore — fixed most of it; counting them per person rather than per group fixed the rest, since one person is left in several anchored groups for the same reason.
Rivals are gathered below the merge threshold, down to even odds: a named person who matches at 0.6 will never be merged into but is exactly the competition a suggestion should be discounted for. The floor matters in both directions — summing the near-orthogonal pairs instead of dropping them lets fifty identities' worth of upper-tail noise outweigh one real match, which on the same library moved the median stated confidence from 100% to 31%.
Measured (cargo run --release -p dr-catalog --example face_confidence), leave-one-out over that
library's 2,702 confirmations across 54 named people:
| share | old mean | |
|---|---|---|
| right person picked | 99.33% | 99.15% |
| stated for the right person, median | 99.3% | 90.4% |
| stated for the right person, p10 | 79.2% | 68.0% |
The reliability table is monotone and errs low: 100% correct wherever it states 80% or more, 84% correct where it states under half. Understating is the safe direction for a screen whose purpose is deciding what to look at first, but the low bands are not calibrated and should not be read as though they were — and the leave-one-out task asks which of these people, never is it any of them, so it cannot speak to a stranger at all.
It is not a merge threshold and must not become one. Uniqueness is relative, so a library with one named person would hand every stray face a 1. "Is this the same person at all" stays §8's question, and coherence is the half of the product that carries it.
9.2 The two numbers the user is allowed to move · 2026-08-29
The merge probability and the smallest group the pass will call a person are FaceSettings in
dr-types, edited from the People screen and saved per device beside the cache budgets. They were
constants: dr_face::DEFAULT_MERGE_PROBABILITY and a bare < 2 in dr_ui::faces::recluster.
Why they had to become settings. The default was tuned on one library — the table in
dr_face::cluster's doc comment is 1,813 faces of one photographer's family — and the quantity it
optimises is a property of the population, not of the model. A library of one household at close
family resemblance and a library of two thousand strangers at a wedding want different answers, and
neither of them is the reference library. The doc comment already conceded the point ("this is a
default, not a constant of nature") and pointed at face_index --tune as the way to find a better
one; a photographer does not have a terminal.
Why moving them is safe, and why that is the reason there is no confirmation on it. A regroup writes only the suggested half. Confirmations, names and ignores enter as anchors and come back unchanged (FR-CULL-10), so the pass is re-runnable by construction and a dial the user can move is just that property being used. The smallest-group rule is applied only to groups the system invented: a group carrying a person — confirmed, named or set aside — survives it whatever its size, because a display preference does not overrule a judgement (FR-CULL-12).
Withdrawal, which the setting does not work without. Raising the smallest group stops the pass
creating small groups; it does not by itself remove the ones a previous pass made, because those
still hold their suggestions, so they are not empty, so prune_empty_unnamed leaves them. The pass
therefore now releases every unanchored face it did not place — faces::unassign — before pruning.
Without that step the control appears to do nothing until the library is reindexed.
The preview. dr_ui::faces::preview_grouping runs the same population through
dr_face::cluster and reports groups, faces grouped and largest group without opening a
transaction. It is face_index --tune's row for one setting, on the user's own library, on a worker
thread. The line leads with the group count because that is the number that says which side of
the right setting you are on: it climbs as fragments are gathered into people and falls as separate
people start being welded together, while the grouped-face count rises straight through both.
10. Catalog and jobs
Schema is catalog.md §10.1's v5 migration, plus faces.crop_px (§7) and the calibration table
(§8.3). Nothing else changes.
One new job kind:
/// Detect and embed faces in one image, from its proxy (FR-CULL-8).
DetectFaces = 8,
Background priority, coalesced per image_id, interruptible, resumable — it inherits FR-CAT-3's
properties with no exceptions, which is what FR-CULL-8's acceptance criterion is about. It is not a
network job.
One job does detection and embedding for every face in the image, rather than splitting them. Splitting would double the queue's row count for no benefit: the proxy is already decoded and in memory, and the natural unit of resumable work is one photograph.
Selector term (FR-CULL-11):
Person { id: PersonId, include_suggested: bool }, // defaults to false
Confirmed-only by default, so a saved smart collection does not silently change membership when a later indexing pass revises a guess.
11. Privacy obligations that constrain the code's shape
NFR-SEC-5 is not a policy to remember; it is a set of properties the code either has structurally or does not have at all. The checkable ones:
dr-facehas no network dependency. Noreqwest, noureq, transitively. Worth a CI check overcargo tree, alongside the existing lints — the crate that holds the embeddings should be provably unable to send them anywhere.- The model fetch (§2.2) is in the app layer, which is why the API in §3.1 takes bytes.
- The diagnostics bundle is an allowlist (NFR-OPS-1), so
faces,face_personand the calibration table are excluded by not being named, and a future table cannot become uploadable by existing. - Delete-all is one transaction and one control: faces, links, people, calibration, and the cached crops if any. Separately, a switch that stops indexing so no such data is produced. The two are different actions and are not collapsed into one.
- The about screen reads the model metadata from the loader — id, version, licence — rather than from a hardcoded string that will drift from what is actually running.
12. What S14 measures
The corpus is a real personal library — the ~2,000-image sample S14 already specifies — with a hand-labelled identity subset. Fixed in advance, so the result is a measurement rather than an argument:
| # | Measure | Why it decides something |
|---|---|---|
| M1 | Does tract load both graphs? As shipped, then with the input dims frozen | Go/no-go, and it is first. det_500m.onnx has a dynamic H/W input (§4.1), which is the exact thing tract failed on for YOLO and which OpenCV's dnn could not load either — so the as-shipped answer is expected to be no, and the real question is whether make_dynamic_shape_fixed is enough or whether a PyTorch re-export is needed. Costs an afternoon; everything below is void without it. |
| M2 | Detection latency per image at 640, and at 320, on the reference desktop and one Android device | Whether a 17k library indexes in a background sweep or a weekend. The one measured datapoint we have — YOLO26n-seg, ~470 ms at 640×640 in tract — suggests SCRFD-500M lands well under it, but tract is not ORT and an extrapolation is not a measurement. |
| M3 | Embedding latency per face, and faces per image in a real library | The multiplier on M2. A library averaging 1.5 faces per image at 100 ms per face is an hour for 25k faces; at 400 ms it is four. |
| M4 | Detection recall against hand-labelled faces, bucketed by crop_px |
Where §4.3's floor should actually sit, and what proportion of a real library's faces are in the degraded bucket §7 predicts. |
| M5 | Aligned versus unaligned embeddings, same corpus, same clustering. Then two A/Bs that cost one run each: least-squares versus RANSAC alignment (§5), and /128 versus /127.5 normalisation (§6) |
Quantifies §5. If the gap is small the alignment code is still correct; if it is large, this is the measurement that stops someone "simplifying" it later. The two A/Bs close divergences between the reference and the published Python that are currently settled by assertion. |
| M6 | Cluster purity and completeness against the labelled subset | Whether the subsystem is worth building at all. §1's F1 figures rank the models on film frames against a known cast; nothing yet says how MBF behaves on family snapshots it must cluster blind. |
| M7 | Calibration reliability, and the sample size at which the fit becomes valid | FR-CULL-9's acceptance criterion, plus the practical question of whether a 2,000-image library ever gets a valid fit or whether §8.3's "unavailable" is the normal state. |
| M8 | Burst-derived positive-pair purity (§8.1) | Whether the free positives are usable or the fit waits for user confirmations. |
| M9 | YuNet as a drop-in detector: M4 and M6, re-run | §2.3's question 3, and the model file is already on disk. If the answer is "close enough", half the licence problem disappears. Needs its own decode path — its outputs are not SCRFD's. |
| M10 | Peak RSS during an indexing sweep | NFR-RES-2. Two loaded graphs plus a proxy plus a batch of crops, on a phone. |
12.1 M1 result — PASS, conditionally · 2026-08-26
Measured, not extrapolated. Both InsightFace graphs fail to load in tract as shipped, exactly as §4.1 predicted and for the reason it gave:
scrfd_500m_bnkps.onnx Translating node #0 "input.1" Source ToTypedTranslator
arcface_w600k_mbf.onnx Failed analyse for node #139 "Conv_0" ConvHir
Both load cleanly once their input dimensions are pinned — SCRFD's unnamed H/W to 640, ArcFace's
None batch to 1 — by tools/fix-face-model-shapes.sh, which rewrites the declared dims and touches
no weights. The frozen SCRFD reports the layout §4.2 specifies, which is the second half of the
answer: nine outputs, three strides, last dims 1/4/10, and 12800 = 80 × 80 × 2 confirming two
anchors per location at stride 8.
Two things worth carrying forward:
SCRFD's outputs were already static. The export was made at 640 and only its input forgot to say so, so pinning to 640 is not a choice this project is making — it is the shape the graph was always going to run at. §12's "320 as a faster option" would need a different export, not a different flag.
YuNet loads with no intervention at all, at a fixed [1, 3, 640, 640], with twelve outputs in
three strides — cls/obj/bbox/kps, which is a different layout from SCRFD's and confirms why
§4.2's load-time check has to look at shapes rather than count outputs. Combined with its permissive
licence (§2.3) that makes M9 more interesting than it looked: the permissive detector is also the one
with no shape-fixing step in front of it.
12.2 First end-to-end run · 2026-08-26
The Rust port produces the separation it is supposed to. Three distinct portraits of one identity against two of another, from §1.1's labelled gallery:
| Pair | Cosine |
|---|---|
| Same identity, different photographs | 0.596 |
| Same identity, byte-identical duplicate files | 1.000 |
| Different identities | 0.049 – 0.050 |
Against the reference's fitted MBF boundary of cos 0.267 (§1), 0.596 and 0.05 fall either side with room to spare — which is the check that the port's pre-processing, letterbox inversion and alignment are right, since any of them being wrong degrades the same-identity number first.
The duplicate row is not a curiosity: several files in that gallery are byte-identical under different names, which is exactly the case §8.2's dedup step exists for, and it would otherwise stack the positive histogram at cos ≈ 1 with pairs that teach the fit nothing.
M2/M3 in a debug build: detection ~1.0–1.4 s per image, embedding ~160–280 ms per face.
M2/M3 in release, over a real library — the number that counts: 3.5 images/second, end to end, including the JPEG decode and the catalog write. 110 images with 125 faces in 30 seconds on the reference desktop. That is roughly 4× the debug figure, and it moves a 23.5k-image library from the "seven hours" the debug numbers implied to about 110 minutes.
Worth stating plainly because the debug measurement was nearly a wrong conclusion: it was on the edge of making the pure-Rust runtime look unaffordable for a large library, and it was measuring the profile rather than the pipeline. Any future timing of this subsystem is a release timing.
Still unmeasured: the same pass on a phone (NFR-RES-2), which does not follow from this one.
13. Order
- M1 — load both graphs in tract, freezing the input dims if needed (§4.1). Go/no-go, and it now comes first: it is an afternoon, and §2.3's reading is wasted effort if the graphs will not run at all. This is a change from the S14 brief's ordering, made because §1.1 removed the uncertainty that justified reading first — the models are known to work, so the open question is the runtime, not the choice.
- Licence reading (§2.3), in parallel with 3. Before anything ships, per D13.
dr-faceskeleton,detect+align+embed, ported from §1.1's table, with an example binary that draws boxes and landmarks on a JPEG — the same shape asdr-segment'sexamples/detect.rs, and for the same reason: the thing worth looking at is whether the landmarks land on a real photograph. Porttests/test_face_utils.cpp's cases first; they are model-free and they fail loudly on exactly the mistakes §5 describes.- M2–M5 on the corpus. Alignment is validated here, before anything depends on embedding quality.
calibrate+clusteragainst the labelled subset. M6–M8. The Platt fit ports fromgallery_calibration.hpp; the pair sourcing (§8.1) is new and is the part to get wrong.- The v5 migration, the
DetectFacesjob, the debounced clustering pass. - UI: People view, confirm and reject, merge and split, the
Personselector term. - The route-C first-run flow, the about-screen entry, and the delete-all control — with the feature shipped disabled until they exist.
Steps 1–5 are the spike. Steps 6–8 are the build, and they are only justified if §12 says so.
13.1 Where this had got to · 2026-08-26
- 1 — done. M1 passes conditionally; §12.1.
- 2 — outstanding. §2.3's three questions are unanswered and gate shipping, not building.
- 3 — done.
core/dr-face:align,detect,embed,embedding, andexamples/faces.rs. §12.2 is its first end-to-end run. - 4 — partly. M2/M3 have provisional numbers; M4/M5 need the labelled corpus.
- 5 — built, not yet measured.
calibrateandclusterexist with 34 model-free tests behind them; M6–M8 are a run over a real library, which is what the corpus in §1.1'simages/is for. - 6 — done for storage. Catalog schema v8 —
people,faces,face_person,face_person_rejected,face_calibration— with the identity operations FR-CULL-10 requires. TheDetectFacesjob kind and the debounced clustering pass are not wired yet. - 7 — done. The Identity screen: a third top-level mode beside library and develop, reached from the library header. People rail, face grid with per-face confirm/reject, rename in place, confirm all, split off a multi-selection, regroup, and the NFR-SEC-5 delete-everything control.
- 8 — not started. The route-C first-run flow: no model fetch, no checksum pin, no licence
notice. The screen says "No face model installed" and stops, which is honest but is not the
feature. The models go in
<catalog dir>/models/as the shape-fixed exports, by hand for now. - 9 — done, and not previously in this plan. The run marker (§10a), the coverage audit, the
face_indexbatch job, and cross-device sync of face shards (§14).
The screen taught the design one thing worth recording. Splitting has to reject before it confirms. Moving faces to a new person is not enough on its own: the next clustering pass sees a face that still looks like the person it left, suggests it back, and the user's correction becomes an argument they keep having. §9's cannot-link constraint handles co-occurrence; this is the same idea applied to a judgement the user made by hand.
Two things the build changed in this document's own design:
crop_px reached the schema as §7 argued it should, and the calibration carries a w_size term
for it.
A rejection table was added, which §8 and §9 did not contemplate. Rejection is not the absence of an assignment: without storing it, the next clustering pass re-suggests exactly the face the user just pushed away. It is user data in the same sense a confirmation is (FR-CULL-12), just negative.
14. What this does not settle
-
Whether a permissive model pair exists. §2.3. If it does, route B replaces route C and step 8's first-run flow shrinks to nothing.
-
Whether tract runs these graphs at all. §12 M1, and §4.1 says why the answer is in real doubt. Every other line of this document is conditional on it.
-
Faces in trashed images. catalog.md §10.4's open question, unchanged: probably excluded from suggestions but not deleted, so a restore does not re-index.
-
Whether embeddings sync.Answered: they do, as sealed shards (dr_catalog::face_shard). Indexing is hours of CPU and its result is byte-identical on every device, so paying for it once per account rather than once per device is the whole argument. §12.2's 3.5 images/second is also what makes the case: it is fast enough to be worth doing and slow enough to be worth not repeating.Shards rather than the catalog snapshot, which is the design decision worth recording. The snapshot uploads whole on every sync, and a fully indexed 23.5k library carries ~30 MB of embeddings — exactly the cost
dr_thumbs's 25 MB cap exists to bound. So the split follows the one already in the tree: bulk immutable data in sealed shards, small mutable data in the snapshot. Faces, landmarks, embeddings and run markers shard; people, names and assignments ride the catalog and merge by uuid. The cap is imported fromdr_thumbsrather than restated, because it is a statement about transfer cost and two copies of it would drift.Everything is keyed on
oc:fileid, neverimage_id: a row id means nothing on another device. -
Re-embedding at higher resolution. §7's
crop_pxmakes it a query rather than a full re-index, but whether it is worth doing is an M4 question. -
Approximate nearest neighbours. §9 says brute force until measured otherwise. A 100k-face library is where this stops being true.
15. Register entries
D13 — face inference runtime and model licensing · the runtime half stays answered (ort +
ort-tract, unchanged since 2026-08-21). The licensing half is answered conditionally by §2:
route C — ship the code, not the weights — unblocks the build without breaking GPL-3.0 or any
distribution channel, and §2.3's reading may yet convert it to route B. Closing D13 outright waits on
that reading.
D17 — face model pair and distribution route · PROPOSED. SCRFD-500M for detection, MobileFaceNet/ArcFace for embedding, weights obtained by the user rather than redistributed, with a pinned checksum and a licence notice shown before the first fetch. The alternative considered and rejected is committing the InsightFace weights (§2.1 route A), which is not available at any price. The model pair is better evidenced than a proposal usually is — §1's table is a measured comparison over a 2,418-identity gallery, not a literature reading — so the open half of this decision is the route, not the pair. Relates to: D13, D14, NFR-COMPAT-2, FR-CULL-8.
D18 — porting from scene-actor-extraction · PROPOSED, and the easy half of a decision.
That project is MIT and by this project's author, so MIT-into-GPL-3.0-or-later is a compatible
one-way combination requiring attribution, not permission. Ported files carry a header naming the
origin. Worth recording because "we already have this working in another language" is exactly the
provenance that goes undocumented and then cannot be answered three years later.
S14 — face pipeline in Rust · scope sharpened, and substantially de-risked, by this document. §12 replaces "measure per-image latency, cluster purity and calibration convergence" with ten specific measures. Two changes to the brief itself: its instruction to resolve the licence question before writing any of it is relaxed to "before shipping any of it", because M1 is an afternoon and is worth knowing first (§13); and its central question — whether the pipeline works at all — is largely answered in advance by §1.1, leaving the runtime and the domain shift from film frames to family snapshots as the real unknowns.
16. Requirements touched
| ID | How this document addresses it |
|---|---|
| FR-CULL-8 | §4 detection, §6 embedding, §7 the proxy tier and its consequences, §10 the DetectFaces job |
| FR-CULL-9 | §8 — pairs, fit, validity, and the reliability-diagram acceptance test |
| FR-CULL-10 | §9 constrained agglomeration, confirmations as anchors, split by re-agglomeration |
| FR-CULL-11 | §10 the Person selector term, confirmed-only by default |
| FR-CULL-12 | §10 schema unchanged from catalog.md §10.1: embeddings derived, names to the sidecar |
| NFR-SEC-5 | §11 — the obligations restated as structural properties, one of them CI-checkable |
| NFR-COMPAT-2 | §2 — why the obvious weights cannot ship, and what does instead |
| NFR-RES-2 | §1 the cheap model pair, §12 M2/M3/M10 on a phone |
| NFR-ARCH-2 | §10 background priority, preempted by visible work |