# Faces and identity — SCRFD and MobileFaceNet Spec for **S14**, the spike that decides whether §3.9.1 is buildable, and the build that follows it. FR-CULL-8 … FR-CULL-12 specify the subsystem in terms of "a 512-dimension embedding from a stated model" and stop there, deliberately: [catalog.md §10.4](catalog.md) records that the model was left abstract because D13 was open. This document names the two models, fixes the pre- and post-processing they need, specifies the calibration and clustering that sit on top, and states what S14 has to measure before any of it is trusted. **It does not close the licensing half of D13.** §2 is the reason, and it comes first because [segmentation.md §7](segmentation.md) established the precedent that reading the grant is cheaper than discovering it at packaging time. --- ## 1. The two models, and why these two **Detector — SCRFD.** *Sample and Computation Redistribution for Efficient Face Detection* (Guo et al., ICLR 2022). A single-stage anchor-based detector whose contribution is where the FLOPs go, not a new architecture: the shallow stages carry more capacity because that is where small faces are decided. It emits, per face, a box, a confidence, and **five landmarks in the same forward pass** — which is the property that matters here, because FR-CULL-8 requires landmarks and because §4's alignment step is not optional. A detector without landmarks would need a second network to supply them. The `500M` variant is the one to ship: 500 MFLOPs at VGA, against 10 GFLOPs for `10G`. Face indexing is a background sweep over a whole library on a phone as well as a desktop (NFR-RES-2), so the cheap variant's accuracy loss on tiny faces is the right trade — a face too small for `500M` to find is also too small for §5 to embed usefully. **Embedder — MobileFaceNet trained with ArcFace loss ("MBF").** ~1M parameters, ~0.45 GFLOPs for one 112×112 crop, 512-d output. ArcFace's additive angular margin is what makes the output space one where *cosine similarity means something* — the loss explicitly optimises angular separation between identities, which is what §6's calibration then has to convert into a probability. **Why not the larger ResNet50 embedder** (`w600k_r50`, the other half of InsightFace's `buffalo_l`): because it is not measurably better here, and it is 13× the file. §1.1's reference project fitted the same Platt calibration (§8) to all four candidates over a 2,418-identity gallery, which is a clean read on the embedding space alone with no tracking or matching logic in it: | Model | File | Steepness `a` | Boundary at P=0.5 | Held-out macro F1 | |---|---|---|---|---| | LVFace-B (Glint360K) | 455 MB | 17.7 | sim 0.228 | 67.4% | | **ArcFace w600k-MBF** | **13 MB** | **16.2** | **sim 0.267** | **64.4%** | | ArcFace w600k-R50 | 174 MB | 15.4 | sim 0.301 | — | | ArcFace R18 | 48 MB | 15.3 | sim 0.309 | 63.1% | **MBF's calibration curve is steeper than R50's and R18's**, separating same-identity from different-identity pairs more confidently and at a lower boundary, at a fraction of the size. LVFace-B is genuinely better and is 35× the file — a defensible choice for a desktop-only feature and not one for a subsystem that has to index a library on a phone (NFR-RES-2). Two caveats on that table, both from the source: the R50 gallery was built with ~30% fewer reference images per identity, which is why it carries no F1 figure — the calibration row does not depend on the image count and is comparable, the benchmark row would not have been. And the F1 column measures a *film-cast identification* task, not this one. It ranks the models; it does not predict DarkRoom's accuracy. **Both are pure convolutional graphs**, which matters for §3: the runtime is tract, and tract's operator coverage is the thing that decides whether a graph runs at all. ### 1.1 A working reference implementation exists `../scene-actor-extraction` is a C++ pipeline by the same author that runs **exactly this model pair** — SCRFD-500MF then ArcFace over 5-point-aligned 112×112 crops — over feature films, with a fitted Platt calibration and a scored benchmark behind it. It is **MIT licensed**, so porting from it into GPL-3.0-or-later is straightforward and needs attribution, not permission. That changes what S14 is. The open questions are no longer "does this pair work" and "what are the magic numbers" — they are *does tract load these graphs* (§12 M1, and §4.1 says why that is in genuine doubt) and *does the accuracy hold on family snapshots rather than film frames*. The pre-processing constants, the decode layout, and the calibration algorithm are all readable rather than rediscoverable, and several of them are not what the published Python would lead you to write. | What | Reference | Ports to | |---|---|---| | ArcFace 5-point template and warp | `src/face_utils.hpp` | `align.rs` (§5) | | SCRFD pre-process, decode, NMS | `src/backends/ort_backend.cpp` `SCRFDDecoder` | `detect.rs` (§4) | | ArcFace pre-process and L2 normalise | same file, `ArcFaceEmbedder` | `embed.rs` (§6) | | Platt fit over a similarity histogram | `src/gallery/gallery_calibration.hpp` | `calibrate.rs` (§8) | | Model-free unit tests for both | `tests/test_face_utils.cpp`, `tests/test_calibration.cpp` | the §3 feature split | That last row is worth noticing: the reference already separates the geometry and arithmetic from the inference well enough to unit-test them with no model on the machine. §3's `inference` feature flag is the same boundary, enforced by Cargo rather than by discipline. **What does not port.** The reference matches faces against a *known gallery* of named identities; DarkRoom clusters an unlabelled library (§9). And it is a video pipeline, so it has a tracker, a `max_faces` cap, and frame-to-frame identity annealing that have no analogue here. --- ## 2. Licensing — the open half of D13 The architectures are published research. **The weights everyone actually uses are not redistributable by this project.** InsightFace's code is MIT. Its *pretrained models* — `buffalo_l`, `buffalo_s`, `buffalo_sc`, and every `det_*` and `w600k_*` checkpoint inside them — carry a **non-commercial research-only** grant, stated in the model-zoo README and applying to both the auto-downloaded and the manually downloaded artefacts. `w600k_mbf` is trained on WebFace600K, whose own terms are research-only as well, so the restriction has two independent sources rather than one that might be renegotiated. These are the exact artefacts in question — §1.1's reference project fetches them from InsightFace's own release page, which is where the grant attaches: | File | Source | |---|---| | `det_500m.onnx` (2.5 MB) | `insightface/releases/download/v0.7/buffalo_sc.zip` | | `w600k_mbf.onnx` (13 MB) | `insightface/releases/download/v0.7/buffalo_s.zip` | DarkRoom is GPL-3.0-or-later. A non-commercial field-of-use restriction is not a GPL-compatible grant, so: - these weights are **not redistributable under this project's licence**, so they cannot go into a *published* build — an APK, a Flatpak, or an F-Droid entry — and could not go into a repository intended to be one (§2.2a records what was actually decided here, and why the two differ); - and the restriction binds the *user*, not only the project — a professional photographer using DarkRoom commercially is outside the grant even if they fetched the file themselves. That last point is the one that is easy to miss and dishonest to leave implicit. It is why §2.2's recommended route puts the licence text in front of the user rather than in a footnote. ### 2.1 The routes, priced | Route | What ships | Cost | |---|---|---| | **A — commit the weights** | Everything works out of the box, one `git lfs pull` | **Not available.** Licence-incompatible with GPL-3.0 and with all three distribution channels (NFR-COMPAT-2). Listed only so it is visibly rejected rather than quietly assumed. | | **B — permissively licensed weights** | Same, legally | Requires weights that do not exist off the shelf in this architecture pair. Either found (§2.3) or trained, and training an ArcFace embedder needs a face dataset whose own terms permit redistribution — the harder half of the problem. | | **C — the user fetches them** | The app ships the *code*, not the weights; face indexing is off until the user obtains a model | Works today, keeps the repository clean, and is honest. Costs a first-run flow, an F-Droid anti-feature declaration, and a licence notice the user has to read. | **Recommendation: C now, B when it becomes possible.** §1.1's project already works this way — a `scripts/download_models.sh` that fetches the weights rather than a repository that carries them — so route C is a practised pattern here rather than a theory. Two things DarkRoom must do that the reference does not: pin a checksum for **every** model (the reference verifies only one of the four), and put the licence in front of the user, because a photo editor's users are not all researchers. The schema already forces this to be a survivable choice — `faces.model_id` (catalog.md §10.1) exists precisely so that a model change is detectable and re-indexable rather than silently poisoning every similarity in the library. Under C, swapping in a permissive model later is a new `model_id` and a re-index, not a migration. ### 2.2 What route C requires of the code - **The weights are never a build input.** `dr-face` (§3) takes bytes; it has no `embedded-model` feature and no `models/` directory. This is the one structural difference from `dr-segment`, and it is deliberate — a feature flag that *could* embed weights is a feature flag someone eventually turns on in a packaging script. - **The fetch does not live in `dr-face`.** It lives in the app layer, so the inference crate keeps no network dependency at all. §11 makes that a checkable property rather than a convention. - **The licence is shown, not linked.** Before the first download, the app states in plain language that the weights are research-only, that commercial use is outside the grant, and who the grantor is. NFR-SEC-5's last bullet already requires the models to be named, versioned and licensed in the about screen; this is the same obligation moved to the moment where it can still change a decision. - **A checksum is pinned.** The app verifies the digest of what it fetched against a value compiled in, and refuses a mismatch. An unverified model file is an arbitrary graph executed over the user's family photographs. - **Indexing stays off until a model is present and verified**, and disabling it deletes nothing — that is a separate control (NFR-SEC-5, §11). ### 2.2a Android, which route C forgot Route C says "the user obtains the model and the app loads it". On a desktop that is a real gesture: the files go in `~/.local/share/darkroom/models/` and face indexing starts working. **On Android that gesture does not exist.** `internal_data_path` is app-private, `adb shell run-as` needs a debuggable build, there is no picker and no fetch in the app, and so a phone could not acquire a model by any means at all. Face indexing was not "off until the user supplies weights" there; it was off, full stop, and the settings page said so on every launch with no action available behind the message. **Decision, 2026-08-27: the shape-fixed pair is committed to LFS at `apps/darkroom-android/android/assets/models/`, and `assemble-apk.sh` bundles it into the APK.** `android_main` unpacks it to the shared models directory on first launch, before anything asks whether a model is present. This is a personal project on a private Gitea; the grant restricts redistribution, and a private repository and a self-installed APK are not that. What that does and does not settle: - **It does not make §2.1's route A available.** The moment this project publishes — an F-Droid entry, a release APK, a Flatpak — these files come back out and route C's unbuilt half (the fetch, the licence screen, the pinned checksum) has to exist first. §2.1's table stands as the answer for a *published* build; this is the answer for the author's own phone. - **It does not make the weights a build input.** `dr-face` still has no `models/` directory and no `embedded-model` feature, and nothing in the cargo build reads these files — §2.2's first bullet guards against a flag someone flips in a packaging script, and that remains guarded. The APK assembly step copies two files; it is the only thing in the tree that knows they exist. - **It does not bind only the project.** §2's last bullet is unchanged and is the one with teeth: the research-only grant restricts the *user*, so commercial photography with a DarkRoom that has these weights in it is outside the grant no matter who put them there. The in-app fetch, the licence screen, and the pinned checksum that §2.2 specifies are still unbuilt, on every platform. When they are built, this becomes redundant and the assets directory empties. ### 2.3 Before S14 writes any code Three questions to answer by reading, in this order, and to record with the date they were checked — the same discipline [segmentation.md §7](segmentation.md) applied: 1. **Is there a permissively licensed SCRFD export?** Third-party ONNX exports of SCRFD are plentiful; a re-export inherits the original weights' grant, so the question is whether anyone has *trained* SCRFD on a redistributable dataset, not whether they have converted the InsightFace one. 2. **Is there a permissively licensed 512-d MobileFaceNet/ArcFace checkpoint?** Candidates worth reading the terms on rather than assuming: EdgeFace (Idiap), the ONNX Model Zoo's ArcFace entry (MIT, but a ResNet100 at ~250 MB — licence-clean and the wrong size), and any MobileFaceNet trained on a dataset with distribution terms. 3. **What does a permissive *detector* alternative cost in accuracy?** YuNet (OpenCV Zoo) is 232 KB, emits five landmarks, and comes from a permissively licensed repository. If it is close enough, the detector half of the licence problem disappears and only the embedder remains — a materially better position than either half alone. Question 3 is nearly free to answer: `face_detection_yunet_2023mar.onnx` is **already sitting in §1.1's `models/` directory**, fetched from `opencv/opencv_zoo`. It needs its own decode path (its output layout is different, which is why the reference's `SCRFDDecoder` explicitly rejects it at load rather than silently misreading it), and then it is one more row in §12's table. --- ## 3. Crate shape — `core/dr-face` A new workspace member, modelled on `dr-segment` and for the same reason: everything that reasons about faces is testable with no GPU adapter and no model present (ARCH §6.5a). ``` core/dr-face/ src/lib.rs FaceError, re-exports src/detect.rs SCRFD: preprocess, decode, NMS src/align.rs 5-point similarity transform → 112×112 crop src/embed.rs MBF: preprocess, forward, L2 normalise src/calibrate.rs cosine → P(same person) (FR-CULL-9) src/cluster.rs constrained agglomeration (FR-CULL-10) src/assign.rs which person, and how sure (§9.1, FR-CULL-9) ``` ```toml [features] # No `embedded-model`. §2.2 — the weights are not a build input, ever. default = [] # The ONNX runtime. Off by default so `calibrate` and `cluster` — which are # arithmetic over embeddings and have no model in them — stay testable in a # build that carries no inference engine at all. inference = ["dep:ort", "dep:ort-tract", "dep:ndarray"] ``` The `ort` + `ort-tract` pairing is settled by D13's 2026-08-21 update and already in the workspace manifest: `ort`'s API, tract's pure-Rust engine, no C under the NDK. **The split between `inference` and the rest is load-bearing.** Calibration and clustering are where the subsystem's accuracy actually lives, they are pure functions of embeddings and labels, and they must be testable against synthetic embedding sets with no weights on the machine. A test suite that needs a research-licensed download to run is a test suite that does not run in CI. ### 3.1 The API ```rust /// A loaded detector. Fixed input shape — see §4. pub struct Detector { /* session, input edge */ } impl Detector { pub fn from_bytes(onnx: &[u8], id: ModelId) -> Result; /// `rgb` is f32 0..=1, row-major, three per pixel. pub fn detect(&self, rgb: &[f32], w: usize, h: usize, opts: &DetectOptions) -> Result, FaceError>; } pub struct Detection { /// Normalised to the image's long edge (catalog.md §10.1). pub bbox: (f32, f32, f32, f32), /// Five points, same normalisation, in the model's own order (§5). pub landmarks: [(f32, f32); 5], pub confidence: f32, } pub struct Embedder { /* session */ } impl Embedder { pub fn from_bytes(onnx: &[u8], id: ModelId) -> Result; /// Takes an *aligned* 112×112 crop. Alignment is `align::warp`; passing a /// raw bbox crop here is the silent-accuracy-loss bug §5 exists to prevent. pub fn embed(&self, aligned: &Aligned112) -> Result; } /// L2-normalised, 512-d. Carries its model id so a comparison across models /// is a type error rather than a plausible-looking number (catalog.md §10.1). pub struct Embedding { pub model: ModelId, pub v: [f32; 512] } ``` `Aligned112` is a newtype over the pixel buffer that only `align::warp` can construct. That is the whole defence against §5's failure mode, and it costs nothing. --- ## 4. Detection ### 4.1 Fixed input, and how the image is fitted to it **`det_500m.onnx` has a dynamic H/W input, and this is the largest single risk in the document.** §1.1's reference had to use ONNX Runtime rather than OpenCV's `dnn` module precisely because OpenCV could not load SCRFD's dynamic `Shape` nodes. tract is in the same family of problem: `dr-segment` exists in its current shape because tract failed shape inference on YOLO's dynamic export, which is why `tools/export-seg-model.sh` passes `dynamic=False` and why `semantic.rs` has a fixed `INPUT_EDGE`. So the graph almost certainly needs its input dimensions frozen before tract will touch it. That is a known, cheap operation — `onnxruntime.tools.make_dynamic_shape_fixed` rewrites the declared dims without retraining or re-exporting from PyTorch — but it is a step, it produces an artefact that must itself be reproducible (a `tools/fix-face-model-shapes.sh`, in the spirit of the existing export script), and **it must be tried before anything else in this document is scheduled.** §12's M1. Fixed at **640**, with 320 available as a faster, blinder option to measure. **Letterbox.** Scale by `min(640/w, 640/h)` preserving aspect, then paste into a 640×640 canvas. The reference centres the image and fills the margin with grey `114`, inverting with `x_src = (x_model − pad_x) / scale`. What matters is not *where* the padding goes but that the forward and inverse agree and that the fill value is treated as part of the contract: a mismatch between them offsets every box and landmark the model returns by the padding, which yields detections that look plausible, landmarks that align badly, and an embedding-quality loss that only shows up as poor clustering three stages later. Port the reference's convention rather than inventing one, and assert it with a round-trip test. Normalisation is `(x·255 − 127.5) / 128`, RGB, NCHW — note `/128`, not `/127.5`; see §6. ### 4.2 Outputs, and decoding them **Nine tensors for three strides `{8, 16, 32}`, twelve for four `{8, 16, 32, 64}`.** Which one a given export produces is discovered at load — `strides = output_count / 3` — not assumed, because both variants exist and hardcoding three silently ignores the largest faces in a twelve-output model. With two anchors per location and `N_s = (640/s)²` locations: | Tensor | Shape | Meaning | |---|---|---| | `score_s` | `[N_s·2, 1]` | sigmoid already applied inside the graph | | `bbox_s` | `[N_s·2, 4]` | distances left, top, right, bottom **in units of the stride** | | `kps_s` | `[N_s·2, 10]` | five `(dx, dy)` offsets, same units | Anchor centres are `(x·s, y·s)` for each grid cell, repeated once per anchor. Decoding is therefore ``` x1 = cx − l·s x2 = cx + r·s kx_i = cx + dx_i·s y1 = cy − t·s y2 = cy + b·s ky_i = cy + dy_i·s ``` then divide by the letterbox scale to return to source pixels, then normalise by the long edge before storage. Flat index within a stride is `(row · fw + col) · 2 + anchor`. **A shape assertion at load time, not a decode-time surprise.** `dr-segment` already learned this — `SegmentError::OutputShape` exists because a different export of the same model produces confidently wrong results otherwise. The reference implements exactly this and it is worth porting verbatim: output count divisible by three and between 9 and 12, then the **last dimension of each group checked against `{1, 4, 10}`**. That single check is what catches a YuNet file passed where an SCRFD one was meant — YuNet also has twelve outputs, so the count alone does not distinguish them, and the failure without the check is not an error but a page of plausible garbage. ### 4.3 Thresholds Score ≥ **0.5**, NMS IoU **0.4**, plain greedy NMS per image across all strides together. Deliberately *not* the low threshold `dr-segment` chose. There, a false positive costs one spurious entry in a list the user is picking from. Here it costs an entry in the People view that the user has to reject, in a library with thousands of images — and worse, a garbage embedding that participates in clustering and can bridge two real clusters into one. False negatives are recoverable by a later re-index with a better model; a polluted cluster graph is not, once the user has confirmed faces inside it. A minimum box size of **40 px on the source proxy** is applied on top — the reference's figure, chosen there for the same reason it holds here: below it there is not enough face left to align reliably, and §7's table shows what the embedder actually receives at that size. §12's M4 is what moves this number, if measurement says it should move. **One thing not to port:** the reference caps detections at ten per frame, largest first. That is right for a film frame, where the extras in the background are noise. It is wrong for a photo library, where a group shot with thirty faces in it is precisely the picture the user wants indexed. No cap; the min-size floor is the only filter. --- ## 5. Alignment — the step that is silently wrong when skipped ArcFace embeddings are trained on faces warped to a canonical 112×112 arrangement. Feeding the model a plain bounding-box crop *works* — it produces 512 numbers, they are unit-norm, and cosine similarities between them look entirely reasonable. They are just much worse, and nothing in the system reports it. The transform is a **similarity transform** — rotation, uniform scale, translation, four degrees of freedom — from the five detected landmarks to this template, which is the arrangement the weights were trained against: ``` (38.2946, 51.6963) subject's right eye ─┐ image-left of centre (73.5318, 51.5014) subject's left eye ─┘ (56.0252, 71.7366) nose tip (41.5493, 92.3655) subject's right mouth corner (70.7299, 92.2041) subject's left mouth corner ``` Similarity, not affine: a full affine fit to five noisy points will happily shear a face, and shear is exactly the deformation the embedder was never shown. **The naming is a trap and the order is not.** Point 0 sits at x=38 on a 112-wide canvas — left of centre *in the image*, which is the subject's **right** eye. Both namings are in circulation and they are opposite. What matters is that SCRFD emits its five points in this same order, so the correct amount of reordering between detector and template is **none**; the reference states this explicitly in `types.hpp` for the benefit of whoever next reads it and doubts it. A future detector with a different order carries its own permutation next to its `model_id`, rather than this file growing an assumption. **One divergence from the reference to settle by measurement.** It fits the transform with OpenCV's `estimateAffinePartial2D` under **RANSAC** at a 3-pixel threshold, and drops the detection when the fit fails. RANSAC over five points is a strange fit — the minimal sample for a similarity is two, so it can discard landmarks it judges outliers and solve from a subset, which on a profile face is as likely to be the correct geometry as the wrong one. InsightFace's own pipeline uses a plain least-squares (Umeyama) similarity over all five points, which cannot silently drop anything and is the one to write first. §12's M5 measures whether the choice matters; the degenerate-input path still has to exist either way, because collinear landmarks do occur. **A trick worth keeping.** The reference retries detection on a frame where nothing was found, after replicate-padding by 25% and applying CLAHE to the L channel. Padding gives a face at the very edge a context the detector needs; CLAHE rescues underexposed frames. Both are plausible on scanned negatives and backlit photographs — worth trying at low priority, since a retry doubles the cost of exactly the images that yielded nothing. Sampling is bilinear from the **native render**, in one step — never crop-then-warp, which resamples twice and throws away detail the warp could have used, and never from the downscaled buffer the detector was given. The landmarks arrive in detector-input coordinates and are scaled back to native before the warp reads a single pixel; §7 is why. Pixels falling outside the source are black. **The landmark order must match the template order.** The template above is written in the detector's own output order; if a future detector emits them differently, the template is reordered with it and that mapping belongs next to the model id, not compiled in as an assumption. *Test:* warp a synthetic image with a known rotation and scale, and assert the five points land on the template within a fraction of a pixel. This is testable with no model present, which is why `align` sits outside the `inference` feature. --- ## 6. Embedding 112×112 **RGB** — the crop comes out of the warp in whatever order the source was in, and the reference converts BGR→RGB explicitly before the blob rather than relying on a flag, which is the readable way to do it. Normalisation is `(x·255 − 127.5) / 128`. **Note `/128`, not `/127.5`:** InsightFace's published Python uses `1.0/127.5` for the recognition model, the reference uses `1.0/128` for both models, and every measured number in §1's table was produced with `/128`. The difference is 0.4% of scale and almost certainly immaterial, but "almost certainly" is not a reason to pick silently — write `/128` to match the numbers we have, and settle it with one back-to-back run in §12. Output is 512 floats. Every downstream comparison is a dot product over the **unit** vector, and the reference L2-normalises before storing so that no code path has to remember to. The reference clamps the norm at 1e-6 before dividing, which costs nothing and removes a NaN path. **Keep the length.** The norm the normalisation divides out is not noise. ArcFace trains the direction of its output and nothing else, and the magnitude it leaves behind grows with how much of a face the model could make out — MagFace (Meng et al., CVPR 2021) made that the training objective, and the plain ArcFace heads this crate runs already show it, weaker but usable. A blur, an occlusion, a hard profile or a badly lit crop comes out short. On the reference library `w600k_mbf`'s norms run from about 8 on a blur to the high 20s on a clean portrait. So the store holds the **raw** vector, not the unit one — `dr_face::Embedded::to_f16_bytes` — and readers re-normalise on load, which they had to do anyway (below). f16 keeps the same three figures of a component whatever the vector's length, so this costs nothing in precision. The length is also kept beside the blob as `faces.quality`, for the readers that never load the vector (the People screen), and it is `NULL` for a face stored as a unit vector before this — a unit vector reads as a length of one, and one is not "unmeasured". What the number does is in §9: a face whose quality is under **`MIN_GALLERY_QUALITY` = 14** is still placed, but is never what another face is compared *against*. The screen shows it as "Quality 17.3", dimmed below the floor, so a user asking why a group did not gather the rest of a person can see that none of its members can vouch for anyone. Some exports of these graphs are fp16 in, fp16 out. The reference detects this from the graph's declared element type rather than from the filename; worth porting, because the alternative failure is a silent garbage tensor. Storage is `512 × f16` (catalog.md §10.1) — 1 KB per face, 25 MB for a 25,000-face library. The f16 round-trip perturbs a unit vector by ~1e-3 in cosine, three orders of magnitude below the separation between a match and a non-match, and the halving matters because these rows are the ones NFR-SEC-5 contemplates optionally syncing. Re-normalise on load after the f16 widen. It is one pass over 512 floats and it removes a class of drift that is otherwise invisible — and, since the blob is raw, it is what turns the stored vector back into the unit one every comparison expects. --- ## 7. Which pixels the pipeline actually sees **This section previously specified the proxy tier — `ThumbSize::Large`, 1024 px — as the buffer both detection and cropping read. That was wrong, and §7b is the measurement that says so.** FR-CULL-8 now separates the two, because they want opposite things: | Stage | Resolution | Why | |---|---|---| | Source render | **native** | The only stage where more pixels exist to be had | | Detector input | downscaled to ~640 | §4.1 letterboxes to 640×640 regardless; more is wasted CPU | | Crop + align | **sampled from the native render** | The 112×112 is fixed, so this decides whether it holds real pixels | | Embedding | 112×112 | §6 | The detector's indifference to resolution is the whole reason the split works. §4.1 fixes its input at 640×640 and letterboxes whatever arrives, so a face occupying 2% of the frame presents at 12 px to the model whether the buffer handed over is 1024 px or 6000 px. Feeding it native pixels buys nothing. **Feeding the *crop* native pixels buys everything**, because §5's warp is the one place where source resolution converts directly into embedding quality. What the crop receives, by face size, from a native render of a 24 MP frame (~6000 px long edge): | Face size in frame | Pixels across, native | Pixels across, 1024 proxy | What §6 receives | |---|---|---|---| | A portrait, face fills a third of the frame | ~2000 | ~340 | Downsampled. Ideal either way. | | Two people, half-length | ~700 | ~120 | Native: comfortable. Proxy: marginal. | | A group of eight | ~290 | ~50 | Native: real pixels. Proxy: **upsampled 2.2×**. | | A figure in a landscape | ~120 | ~20 | Native: usable. Proxy: below §4.3's floor. | So `faces` records one column beyond catalog.md §10.1's schema: ```sql ALTER TABLE faces ADD COLUMN crop_px INTEGER NOT NULL; -- source pixels across the aligned crop ``` `crop_px` earns its place four times over. It is the honest quality signal for the UI; it is a **feature in §8's calibration**, which FR-CULL-9 explicitly demands ("a raw cosine means something different for every model, every population, and *every face size*"); it is what a higher-resolution re-embedding pass selects on, so that pass is a query rather than a re-index of everything; and it is the only way to audit whether the rule above is actually being followed — which is how §7b found that it was not. --- ## 7a. Recording that detection *ran* The spec above assumed the `faces` table could answer "has this image been indexed". **It cannot**, and the difference is the one that decides whether a background pass ever finishes. A photograph with no face in it produces no rows. So does one that has never been looked at. Asking `faces` therefore re-queues every landscape, still life and document scan on every pass, for ever — and in a personal library that is most of it. Measured on the reference library: of the first 110 images indexed, **64 contain no face at all**. So `face_index` records the *run*: one row per `(image, model)` carrying the timestamp, the number of faces found — zero is the interesting value — and the long edge of the proxy it read. Keyed on the model, so an *embedder* change puts every image back in the queue without anyone having to remember to clear anything; a detector change in front of the same embedder does not (§12.3). Three things fall out of it that were not otherwise available: - **A coverage figure.** "4,812 of 5,000 indexed" is what a user wants to see; counting face rows can only ever report how many faces exist, which is a different number that never reaches the image count. - **A reason for the ones outstanding.** The audit splits them by whether a proxy exists, because *waiting on the thumbnail sweep* and *waiting on face indexing* are different problems and only one of them is fixed by running this again. On the reference library the first check reported 110 ready and 23,417 awaiting a proxy — which is the real state of that library, and not something the face subsystem can do anything about. - **Something to sync.** §14's shards carry the marker with the faces, so an adopted image is not re-detected on the receiving device. --- ## 7b. What the proxy tier actually cost, measured The table in §7 predicted upsampling for small faces and called it "degraded, and usable". On the reference library of 23,531 images it was not the edge case that description implies. Reading `crop_px` across the 18,671 faces stored under the proxy-tier implementation: | Source pixels across the aligned crop | Faces | Share | |---|---|---| | <56 (upsampled more than 2×) | 314 | 1.7% | | 56–111 (upsampled) | 8,505 | 45.6% | | 112–223 (roughly native) | 6,313 | 33.8% | | ≥224 (downsampled — ideal) | 3,539 | 19.0% | **47.3% of every face in the library was upsampled to reach the embedder**, with `crop_px` as low as 34 — a 3.3× enlargement — against a mean of 178. An upsampled crop does not fail loudly. It produces a confident 512-d embedding describing detail that was interpolated rather than photographed, and the damage appears three stages later as clusters that will not separate. ### Measured again, against the implementation The figures above are read out of a catalog after the fact, so they describe what the old code did rather than what the new code does. `examples/face_native.rs` renders one file and indexes it both ways, so the difference can be attributed to the resolution and nothing else. Fourteen originals from the reference library — 5472×3648 Canon CR2 and DNG — each rendered once and indexed twice: | Path | Faces found | Mean `crop_px` | |---|---|---| | Native (this specification) | 9 | **287** | | Everything from a 1024 proxy | 5 | **75** | Crops 3.8× larger, and on the right side of the line that matters: 75 px is *below* [`ALIGNED_EDGE`] so the proxy path was upsampling into the embedder on average, where the native path downsamples into it. **Sample of fourteen images and nine faces.** Enough to show the direction and to catch a wrong landmark scaling, which is what it was written for; not enough to quote a ratio as the library-wide figure. §12's M4 is still where the recall curve gets established. A second effect, recorded here because it was measured and because the mechanism is **not** established. Grouping the same runs by `face_index.source_edge` — the buffer detection ran against — gives 0.078 faces per image at 1024 or below, against 1.82 at 2048 or better. Controlled for file type and size (1,592 DNGs averaging 21.0 MB against 7,724 averaging 23.3 MB, same library, same cameras), so it is not a composition artefact. But it cannot be a matter of the detector seeing fewer pixels, since §4.1 letterboxes both to 640: a 1024 buffer and a 3072 buffer present the same face at the same size to the model. The likeliest explanation is that the 1024 proxy is itself a downscale of a larger preview, so the detector sees a twice-resampled image where the larger buffer is resampled once — but that is a hypothesis, not a finding, and §12's M4 is where it should be settled. **The crop measurement above stands on its own and does not depend on it.** The A/B is evidence for that hypothesis without settling it. Native found nine faces where the 1024 path found five, on four files where the proxy path found none at all — so detector input *does* affect recall, which the letterbox says it should not. The two paths differ in their resampling as well as their size (one box filter and one letterbox against a downscale and a letterbox), and this experiment does not separate those. Isolating it means holding the chain fixed and varying only the buffer, which is M4's job. --- ## 8. Calibration — cosine to probability FR-CULL-9 makes this a hard requirement: no code path may threshold a bare cosine, every threshold in the subsystem is stated as a probability, and the fit is per library and reports its own validity. This section is how that is obtained, and the interesting part is where the training pairs come from when the user has labelled nothing yet. ### 8.1 Where the pairs come from **Negatives are free and abundant.** Two faces detected in *the same photograph* are almost never the same person. That gives every multi-face image in the library a full set of negative pairs at no labelling cost — and they are *hard* negatives, drawn from the same camera, lighting, and processing, which is exactly the population where a threshold tuned on easy negatives fails. The known exceptions — mirrors, photographs of photographs, collages, a spliced panorama — are rare enough to be noise at this scale and worth naming so the exception is not mistaken for a bug later. **Positives, in order of trustworthiness:** 1. **User confirmations** (FR-CULL-10). Every pair of faces confirmed to the same person. The gold standard, and empty on day one. 2. **Burst siblings.** FR-CULL-5 already groups bursts. Two faces in adjacent frames of one burst, in nearly the same position, are near-certainly the same person. Free, requires no labelling, and available immediately on any library with continuous-shooting frames in it. **Their purity is an S14 measurement** (§12), not an assumption — if bursts turn out to be dirtier than expected, this source is dropped and the calibration simply stays invalid for longer. 3. Nothing else. Bootstrapping positives from high cosine similarity is circular — it fits the calibration to the belief it was supposed to test. ### 8.2 The fit §1.1's reference already implements this and its shape should be ported rather than reinvented: ``` P(same | cos) = σ(a·cos + b + log_prior_odds) ``` Four details in it are the difference between working and nearly working. **Fit against a histogram, not against pairs.** A 25,000-face library has 3×10⁸ pairs; no gradient descent is running over that. The reference buckets every pair into **200 bins over cos ∈ [−1, 1]**, carrying a positive and a negative count per bin, and fits the two parameters against the per-bin counts. The cost becomes the similarity matrix plus 200 numbers, and the matrix is a single GEMM (§9). This is the trick that makes a per-library fit affordable at all, and it is not obvious from the outside. **Balance the classes explicitly.** `w_pos = total/(2·n_pos)`, `w_neg = total/(2·n_neg)`. Negatives outnumber positives by orders of magnitude; an unweighted fit produces a well-shaped curve sitting at the wrong height — precisely the "plausible number all the way to the user interface" failure FR-CULL-9 describes. **The base rate is a runtime argument, not part of the fit.** `log_prior_odds` is added at evaluation time, so the balanced fit is stored once and the prior varies per query — the odds that two faces drawn from a 40-image holiday album match are not the odds for a 40,000-image archive. Its inverse, `boundary_at(p)`, gives the cosine at which the probability crosses `p`, which is what turns §9's "merge above 0.9" into an actual comparison. Baking a prior into `a` and `b` would need a refit per context and would make the stored parameters mean something different depending on where they came from. **Deduplicate before pairing.** Near-identical embeddings (cos > 1 − 1e-7) are the same photograph counted twice; the reference drops them per identity first. In a photo library the equivalent is duplicates and virtual copies (FR-CAT-11), and leaving them in stacks the positive histogram at cos ≈ 1 with pairs that teach the fit nothing about hard cases. **A third feature this design adds.** The reference fits on cosine alone; DarkRoom should carry `crop_px` too: ``` logit P(same) = w₀ + w₁·cos + w₂·log₂(min(crop_px_a, crop_px_b)) ``` FR-CULL-9 explicitly names face size as an axis along which an uncalibrated similarity misbehaves, §7 shows a real library spans 40 px to 340 px of face, and `crop_px` is already in hand. The *minimum* of the pair, because a comparison is only as good as its worse crop. It generalises the histogram to a small 2-D grid of bins, which changes nothing structural. If M7 says the extra feature buys nothing, drop it and match the reference exactly — but a film pipeline working from broadcast frames had far less size variation to explain than this one does. ### 8.3 Validity, and saying so Stored per catalog.md §10.3: parameters, a validity flag, and a hash of the face set fitted from. Invalid when fewer than **200 positive pairs** or **2,000 negative pairs** are available, or when the reliability check fails. Deliberately far stricter than the reference's floor of two positives and one negative. That floor is reasonable there: its pairs come from a curated gallery of labelled reference portraits, where a positive pair is trustworthy by construction. Here the positives are bootstrapped from bursts and early confirmations (§8.1) and the whole risk is fitting confidently to a handful of them. The reference also refuses to draw positives from an identity with fewer than five distinct embeddings, letting it contribute negatives only — the same asymmetry applies to a thinly-confirmed person and is worth keeping. In that state both clustering and the displayed confidences fall back to the reference implementation's fitted curve — a documented operating point, not an invention — and the People screen says so once, above the grid, rather than blanking every percentage. That is FR-CULL-9's requirement read as written: what may not happen is an untuned default presented *as though it were measured* on this library. Blanking them was the first reading, and it was wrong in a way worth recording. A young library has no fit; a fit needs confirmations; confirmations are made on a screen the user ranks by confidence. Withholding the confidence until the fit exists is a deadlock in which the normal state of the feature is its degraded one. Refit is triggered by the same debounce as clustering (§9), and when the confirmed-pair count grows materially. *Acceptance (FR-CULL-9):* a reliability diagram over ten probability bins, on a held-out labelled split, with observed match rate within a stated tolerance of the predicted probability in each populated bin. A single accuracy figure is not an answer to this requirement. --- ## 9. Clustering Not a job kind, for the reason catalog.md §10.2 gives: it is a whole-library operation with no natural `subject_id`. It runs as a debounced library pass when detection has been idle and the face count has moved materially. **The graph.** For each face, its `k = 20` nearest neighbours by cosine, then each candidate edge scored through §8 to a probability. Brute force is honest arithmetic here, and §1.1's project is the evidence rather than the estimate: it computes the full similarity matrix as **one GEMM** (`E · Eᵀ` over L2-normalised rows) at gallery scale, on CPU, and instruments it well enough to have compared CPU against OpenCL and found the question worth asking rather than urgent. For DarkRoom that is a 25,000 × 512 matrix at 25 MB in and ~2.5 GB out — so the matrix is computed **in row blocks**, with each block reduced to its top-k and its histogram contribution before the next is started, and never materialised whole. Once, in the background, after an indexing sweep that took an hour. An approximate index is an optimisation to reach for when §12 says it is needed, not before. **The kernel is most of the cost, and it was running at a tenth of the machine.** The scan is `O(n²)` dot products and nothing else, so its speed *is* the subsystem's speed. Two things were wrong with the first one, both measured over a real 18,143-face library on a twenty-core desktop: | | scan | GFLOP/s | |---|---|---| | a row against every other row, `&[Vec]` | 4.64 s | 36 | | tiled on the column side too, one flat buffer | 2.81 s | 60 | | **plus AVX2 + FMA** | **0.86 s** | **195** | The first is memory: walking the whole embedding array once per row moves ~336 GB for that library, where a column tile that fits in L2 is read once per *tile of rows*. The second is that the workspace builds for baseline `x86-64` — SSE2, no FMA — and the portable loop was not being vectorised into even that, at 0.7 flops per cycle. So the dot product is chosen per machine: AVX2 + FMA where `is_x86_feature_detected!` finds it, **NEON unconditionally on aarch64** — Advanced SIMD is in that baseline, so every Android device the app builds for has it, and the explicit `vfmaq` matters because LLVM will not fuse a multiply and an add on its own. The portable loop remains the definition the others are tested against. All three produce the same 1,531,969 pairs. The NEON kernel is the one a desktop `cargo test` never executes, so `tools/face-tests-on-device.sh` runs the suite on an attached device: `dr-face` carries no weights and touches no display, so its tests are a plain ARM64 binary that runs under `adb shell` with nothing installed. Worth running whenever the kernels change. **Where a regroup's time actually goes**, on that library, because the answer moved twice while it was being looked at. Measured with `cargo run --release -p dr-catalog --example face_confidence -- CATALOG --full`, on the reference desktop and on a Honor tablet (ROD2-W09, aarch64): | | desktop, before | desktop | **tablet** | |---|---|---|---| | scan | 4.60 s | 0.96 s | **2.61 s** | | agglomerate | 4.84 s | 1.69 s | **2.16 s** | | score | 0.23 s | 0.26 s | **0.40 s** | | **total** | **10.0 s** | **3.1 s** | **5.2 s** | The scan came down by the kernel work above. The agglomeration was not, as it looked, an irreducible sequential heap walk: 3.06 s of it was every component scanning the *whole* pair list for the pairs that were its own — 382 million set lookups to place 804,499 pairs — which the union-find can do in one pass while it is finding the components anyway. Another 0.54 s was `Engine::cross` computing dot products with the portable loop while the scan beside it used the machine's SIMD. **The two architectures agree exactly**, which is worth more than either column: the same 1,531,969 evidence pairs, the same 2,518 groups holding the same 16,246 faces, and the same reliability table, from AVX2 on the desktop and NEON on the tablet. That is the cross-kernel check the unit test can only approximate. **The tablet is where a GPU GEMM would pay.** Its scan is *half* the pass, against under a third on the desktop — twenty cores of AVX2 pull ahead of a tablet's NEON far more than the merge engine's single-threaded hashing does. So a perfect GEMM is worth about 2× a regroup there and about 1.5× here, and it is the phone and tablet story that should decide whether it gets built. **Constraints, not just thresholds:** - **Cannot-link on co-occurrence.** Two faces in the same image are never merged. This is the same observation §8.1 mines for negatives, used here as a hard constraint, and it is the single cheapest defence against the over-merging FR-CULL-10 warns about. - **Confirmed faces are anchors.** A confirmation is user data (FR-CULL-12) and clustering never moves it. Two clusters each containing confirmations of *different* people cannot merge; a cluster containing confirmations of one person absorbs suggestions but never reassigns the confirmed. - **A short embedding is never a reference.** The length of the raw vector is the model's own reading of the crop (§6), and a short one sits near the middle of the sphere, matching a little of everybody — one of those in a group is a bridge to the next group over. So the population is split: faces at or above `MIN_GALLERY_QUALITY` are the **gallery** and cluster as described below; faces under it are **probes**, each measured against the finished groups and placed in the one it fits by the same average-link rule under the same two constraints — but measured against gallery members only, never against another probe, and once placed never part of what the next face is measured against. Two probes are never paired at all, and `neighbours` drops those pairs before anything downstream sees them. A probe's confidence (§9.1) is computed from the references it matched; a reference's confidence hears nothing from a probe. A face whose quality was never recorded is admitted to the gallery — a rule that cannot be checked admits rather than excludes — and the next indexing pass **measures** it: the `face-quality` repair (§18.1) lists every image holding one, and each such face is embedded again from the native render with the landmarks it already has, the raw vector written over the old one and its id, box and identity untouched (`faces::record_updates`). No detector runs and no suggestion is lost — the cost is the original fetched once more, since the length exists only at the moment of embedding. - **A person is stood for by their references.** Every face the user has ruled on is an anchor, and the scan is exhaustive, so a person with 750 confirmations would cost 750 comparisons against every other face — and the cost of a library would grow with how well it was named. Instead each person enters through at most `MAX_REFERENCES` (100) of their anchored faces, chosen by `dr_face::references`: those whose raw embedding is at least `MIN_REFERENCE_QUALITY` (15) long, and among them the set spanning the greatest volume — greedy max-determinant, the longest vector first and then, at each step, the face with the largest component orthogonal to the chosen so far. Thirty frames from one afternoon contribute one reference; the single profile shot is taken early. The faces not chosen keep their confirmations and are not touched by the pass; they are simply not compared. **The algorithm.** Constrained average-link agglomeration over the probability graph, merging while the average pairwise probability exceeds **0.9** and no cannot-link is violated. Average-link rather than single-link because single-link chains — one bad edge welds two identities together, which is the documented way face clustering fails on families. **Incremental by default.** A new face joins the existing cluster whose average probability against it is highest, if that clears the threshold; otherwise it starts an unnamed group. A full re-agglomeration happens on model change, on recalibration, and on explicit user request. Suggestions are recomputed freely; confirmations survive all of it (FR-CULL-10). **Splitting** re-agglomerates within one person at a raised threshold and offers the resulting groups as the split. FR-CULL-10 requires splitting to be as easy as merging, and a split that hands the user a pile of loose faces to re-sort is not that. ### 9.1 The number beside a suggestion Which person a face belongs to and how sure that is are **different questions**, and the second one is not answered by the pairwise probabilities that settled the first. The first implementation answered it with the mean calibrated probability between the face and the rest of its group, and that measures the wrong thing twice. It punishes coverage: a person with two hundred faces across fifteen years is *supposed* to have members a new photograph is orthogonal to, so the better someone is photographed the worse their suggestions score. And it never asks who else the face could be — a face matching Anna at 0.95 and nobody else, and one matching Anna at 0.95 and her sister at 0.93, come out identical, when the second is the only one the user needs to look at. Two questions, so two factors, multiplied: ``` evidence(P) = Σ of the top n of { P(same | this face, f) : f ∈ P } n = 10 coherence = evidence(own) / how many of the top n there were uniqueness = evidence(own) / (evidence(own) + Σ evidence(named rivals)) confidence = coherence × uniqueness ``` **Coherence** is the old mean with a cap on it, and the cap is the whole fix: the two hundred faces a given photograph is legitimately orthogonal to stop counting against it. **Uniqueness** is the competition, and it is what makes an ambiguous face read as ambiguous — two identities matching equally well land at 0.5 each, which is the truth about a sibling. **Only named people compete, and they compete per person.** This is the part that had to be measured rather than reasoned about. Normalising across *every* group made the number useless on a real 18,000-face library — median suggestion 21%, four in five under half — because clustering leaves one person spread across many groups, so a face competes against itself. Counting only the groups the user has ruled on — one holding a confirmation, a name, or an ignore — fixed most of it; counting them **per person** rather than per group fixed the rest, since one person is left in several anchored groups for the same reason. **Rivals are gathered below the merge threshold**, down to even odds: a named person who matches at 0.6 will never be merged into but is exactly the competition a suggestion should be discounted for. The floor matters in both directions — summing the near-orthogonal pairs instead of dropping them lets fifty identities' worth of upper-tail noise outweigh one real match, which on the same library moved the median stated confidence from 100% to 31%. *Measured (`cargo run --release -p dr-catalog --example face_confidence`), leave-one-out over that library's 2,702 confirmations across 54 named people:* | | share | old mean | |---|---|---| | right person picked | **99.33%** | 99.15% | | stated for the right person, median | **99.3%** | 90.4% | | stated for the right person, p10 | **79.2%** | 68.0% | The reliability table is monotone and **errs low**: 100% correct wherever it states 80% or more, 84% correct where it states under half. Understating is the safe direction for a screen whose purpose is deciding what to look at first, but the low bands are not calibrated and should not be read as though they were — and the leave-one-out task asks *which of these people*, never *is it any of them*, so it cannot speak to a stranger at all. It is **not** a merge threshold and must not become one. Uniqueness is relative, so a library with one named person would hand every stray face a 1. "Is this the same person at all" stays §8's question, and coherence is the half of the product that carries it. ### 9.2 The two numbers the user is allowed to move · 2026-08-29 The merge probability and the smallest group the pass will call a person are `FaceSettings` in `dr-types`, edited from the People screen and saved per device beside the cache budgets. They were constants: `dr_face::DEFAULT_MERGE_PROBABILITY` and a bare `< 2` in `dr_ui::faces::recluster`. **Why they had to become settings.** The default was tuned on one library — the table in `dr_face::cluster`'s doc comment is 1,813 faces of one photographer's family — and the quantity it optimises is a property of the population, not of the model. A library of one household at close family resemblance and a library of two thousand strangers at a wedding want different answers, and neither of them is the reference library. The doc comment already conceded the point ("this is a *default*, not a constant of nature") and pointed at `face_index --tune` as the way to find a better one; a photographer does not have a terminal. **Why moving them is safe, and why that is the reason there is no confirmation on it.** A regroup writes only the *suggested* half. Confirmations, names and ignores enter as anchors and come back unchanged (FR-CULL-10), so the pass is re-runnable by construction and a dial the user can move is just that property being used. The smallest-group rule is applied only to groups the system invented: a group carrying a person — confirmed, named or set aside — survives it whatever its size, because a display preference does not overrule a judgement (FR-CULL-12). **Withdrawal, which the setting does not work without.** Raising the smallest group stops the pass *creating* small groups; it does not by itself remove the ones a previous pass made, because those still hold their suggestions, so they are not empty, so `prune_empty_unnamed` leaves them. The pass therefore now releases every unanchored face it did not place — `faces::unassign` — before pruning. Without that step the control appears to do nothing until the library is reindexed. **The preview.** `dr_ui::faces::preview_grouping` runs the same population through `dr_face::cluster` and reports groups, faces grouped and largest group without opening a transaction. It is `face_index --tune`'s row for one setting, on the user's own library, on a worker thread. The line leads with the **group count** because that is the number that says which side of the right setting you are on: it climbs as fragments are gathered into people and falls as separate people start being welded together, while the grouped-face count rises straight through both. --- ## 10. Catalog and jobs **Schema** is catalog.md §10.1's v5 migration, plus `faces.crop_px` (§7) and the calibration table (§8.3). Nothing else changes. **One new job kind:** ```rust /// Detect and embed faces in one image, from its proxy (FR-CULL-8). DetectFaces = 8, ``` Background priority, coalesced per `image_id`, interruptible, resumable — it inherits FR-CAT-3's properties with no exceptions, which is what FR-CULL-8's acceptance criterion is about. It is not a network job. One job does detection *and* embedding for every face in the image, rather than splitting them. Splitting would double the queue's row count for no benefit: the proxy is already decoded and in memory, and the natural unit of resumable work is one photograph. **Selector term** (FR-CULL-11): ```rust Person { id: PersonId, include_suggested: bool }, // defaults to false ``` Confirmed-only by default, so a saved smart collection does not silently change membership when a later indexing pass revises a guess. --- ## 11. Privacy obligations that constrain the code's shape NFR-SEC-5 is not a policy to remember; it is a set of properties the code either has structurally or does not have at all. The checkable ones: - **`dr-face` has no network dependency.** No `reqwest`, no `ureq`, transitively. Worth a CI check over `cargo tree`, alongside the existing lints — the crate that holds the embeddings should be provably unable to send them anywhere. - **The model fetch (§2.2) is in the app layer**, which is why the API in §3.1 takes bytes. - **The diagnostics bundle is an allowlist** (NFR-OPS-1), so `faces`, `face_person` and the calibration table are excluded by not being named, and a future table cannot become uploadable by existing. - **Delete-all is one transaction and one control**: faces, links, people, calibration, and the cached crops if any. Separately, a switch that stops indexing so no such data is produced. The two are different actions and are not collapsed into one. - **The about screen reads the model metadata from the loader** — id, version, licence — rather than from a hardcoded string that will drift from what is actually running. --- ## 12. What S14 measures The corpus is a real personal library — the ~2,000-image sample S14 already specifies — with a hand-labelled identity subset. Fixed in advance, so the result is a measurement rather than an argument: | # | Measure | Why it decides something | |---|---|---| | **M1** | **Does tract load both graphs?** As shipped, then with the input dims frozen | Go/no-go, and it is *first*. `det_500m.onnx` has a **dynamic H/W input** (§4.1), which is the exact thing tract failed on for YOLO and which OpenCV's `dnn` could not load either — so the as-shipped answer is expected to be no, and the real question is whether `make_dynamic_shape_fixed` is enough or whether a PyTorch re-export is needed. Costs an afternoon; everything below is void without it. | | **M2** | Detection latency per image at 640, and at 320, on the reference desktop and one Android device | Whether a 17k library indexes in a background sweep or a weekend. The one measured datapoint we have — YOLO26n-seg, ~470 ms at 640×640 in tract — suggests SCRFD-500M lands well under it, but tract is not ORT and an extrapolation is not a measurement. | | **M3** | Embedding latency per face, and faces per image in a real library | The multiplier on M2. A library averaging 1.5 faces per image at 100 ms per face is an hour for 25k faces; at 400 ms it is four. | | **M4** | Detection recall against hand-labelled faces, bucketed by `crop_px` | Where §4.3's floor should actually sit, and what proportion of a real library's faces are in the degraded bucket §7 predicts. | | **M5** | **Aligned versus unaligned embeddings**, same corpus, same clustering. Then two A/Bs that cost one run each: least-squares versus RANSAC alignment (§5), and `/128` versus `/127.5` normalisation (§6) | Quantifies §5. If the gap is small the alignment code is still correct; if it is large, this is the measurement that stops someone "simplifying" it later. The two A/Bs close divergences between the reference and the published Python that are currently settled by assertion. | | **M6** | Cluster purity and completeness against the labelled subset | Whether the subsystem is worth building at all. §1's F1 figures rank the models on *film frames against a known cast*; nothing yet says how MBF behaves on family snapshots it must cluster blind. | | **M7** | Calibration reliability, and **the sample size at which the fit becomes valid** | FR-CULL-9's acceptance criterion, plus the practical question of whether a 2,000-image library ever gets a valid fit or whether §8.3's "unavailable" is the normal state. | | **M8** | Burst-derived positive-pair purity (§8.1) | Whether the free positives are usable or the fit waits for user confirmations. | | **M9** | YuNet as a drop-in detector: M4 and M6, re-run | §2.3's question 3, and the model file is already on disk. If the answer is "close enough", half the licence problem disappears. Needs its own decode path — its outputs are not SCRFD's. | | **M10** | Peak RSS during an indexing sweep | NFR-RES-2. Two loaded graphs plus a proxy plus a batch of crops, on a phone. | ### 12.1 M1 result — **PASS, conditionally** · 2026-08-26 Measured, not extrapolated. Both InsightFace graphs **fail to load in tract as shipped**, exactly as §4.1 predicted and for the reason it gave: ``` scrfd_500m_bnkps.onnx Translating node #0 "input.1" Source ToTypedTranslator arcface_w600k_mbf.onnx Failed analyse for node #139 "Conv_0" ConvHir ``` Both **load cleanly once their input dimensions are pinned** — SCRFD's unnamed H/W to 640, ArcFace's `None` batch to 1 — by `tools/fix-face-model-shapes.sh`, which rewrites the declared dims and touches no weights. The frozen SCRFD reports the layout §4.2 specifies, which is the second half of the answer: nine outputs, three strides, last dims 1/4/10, and `12800 = 80 × 80 × 2` confirming two anchors per location at stride 8. Two things worth carrying forward: **SCRFD's outputs were already static.** The export was made at 640 and only its input forgot to say so, so pinning to 640 is not a choice this project is making — it is the shape the graph was always going to run at. §12's "320 as a faster option" would need a different export, not a different flag. **YuNet loads with no intervention at all**, at a fixed `[1, 3, 640, 640]`, with twelve outputs in three strides — `cls`/`obj`/`bbox`/`kps`, which is a *different layout* from SCRFD's and confirms why §4.2's load-time check has to look at shapes rather than count outputs. Combined with its permissive licence (§2.3) that makes M9 more interesting than it looked: the permissive detector is also the one with no shape-fixing step in front of it. ### 12.2 First end-to-end run · 2026-08-26 The Rust port produces the separation it is supposed to. Three distinct portraits of one identity against two of another, from §1.1's labelled gallery: | Pair | Cosine | |---|---| | Same identity, different photographs | **0.596** | | Same identity, byte-identical duplicate files | 1.000 | | Different identities | **0.049 – 0.050** | Against the reference's fitted MBF boundary of cos 0.267 (§1), 0.596 and 0.05 fall either side with room to spare — which is the check that the port's pre-processing, letterbox inversion and alignment are right, since any of them being wrong degrades the same-identity number first. The duplicate row is not a curiosity: several files in that gallery are byte-identical under different names, which is exactly the case §8.2's dedup step exists for, and it would otherwise stack the positive histogram at cos ≈ 1 with pairs that teach the fit nothing. **M2/M3 in a debug build:** detection ~1.0–1.4 s per image, embedding ~160–280 ms per face. **M2/M3 in release, over a real library — the number that counts: 3.5 images/second**, end to end, including the JPEG decode and the catalog write. 110 images with 125 faces in 30 seconds on the reference desktop. That is roughly **4× the debug figure**, and it moves a 23.5k-image library from the "seven hours" the debug numbers implied to about **110 minutes**. Worth stating plainly because the debug measurement was nearly a wrong conclusion: it was on the edge of making the pure-Rust runtime look unaffordable for a large library, and it was measuring the profile rather than the pipeline. Any future timing of this subsystem is a release timing. Still unmeasured: the same pass on a phone (NFR-RES-2), which does not follow from this one. ### 12.3 SCRFD-2.5G and 10G against 500M · 2026-09-11 §1 chose `500M` on cost; the recall it gives up was never measured. `examples/face_detectors` in `dr-ui` runs several detectors over the same 400 proxies, spaced evenly through the reference library, matches boxes at IoU ≥ 0.5 against the first, and writes contact sheets of the disagreements — because a count of extra faces says nothing until someone has looked at whether they are faces. Both size gates off, confidence at the production 0.5, release build, reference desktop. Sizes are the box's shorter edge in proxy pixels; the network sees 0.625 of that. | Detector | File | Mean ms | Faces | <16 | 16–32 | 32–64 | ≥64 | |---|---|---|---|---|---|---|---| | `scrfd_500m` | 2.5 MB | 157 | 1,490 | 153 | 730 | 399 | 208 | | `scrfd_2.5g` | 3.3 MB | 177 | 1,698 | 268 | 801 | 418 | 211 | | `scrfd_10g` | 17 MB | 488 | 1,896 | 357 | 886 | 438 | 215 | | Against `500M` | Both | Candidate only | Baseline only | |---|---|---|---| | `2.5G` | 1,404 | **294** (242 of them under 32 px) | 86 | | `10G` | 1,419 | **477** (411 under 32 px) | 71 | Two things the sheets settled that the counts cannot: - **The candidates' extras are faces.** The hundred smallest `2.5G`-only tiles are people — soft, small, a few motion-blurred, and almost none of them anything else. These are the group shots and the figures in the background that §7's table predicted `500M` would lose at 640. - **`500M`'s "extras" are mostly not.** Of the 86 faces only `500M` found, the sheet shows the same dog a dozen times, a stop sign, a wheel, two hands, the backs of several heads and one face upside down. The larger models did not miss these; they declined them. So `500M` is paying twice — for the faces it cannot find *and* for the non-faces it embeds, which is exactly the garbage-embedding-bridges-two-clusters failure `DetectOptions::confidence` is set against. **`2.5G` is the right detector.** 12% more time for 14% more faces and a cleaner set, in a file 0.8 MB larger; `10G` finds a further 12% for 3.1× the time, which is a desktop-only price and this is not a desktop-only feature (NFR-RES-2). It loads in tract with the same fix as `500M` (`--input input.1=1,3,640,640`; its outputs are declared dynamic and tract infers them) and decodes through the same nine-output path unchanged. Same licence, same `buffalo_m` release page. **Which detector runs is a setting** — `FaceSettings::detector`, per device, on the settings page beside the indexing button as Fast / Balanced / Thorough. All three files ship. Each detector writes its own `faces.model_id` (`w600k_mbf` for `500M`, unchanged; `scrfd_2.5g+w600k_mbf` and `scrfd_10g+w600k_mbf` for the others), so which pipeline drew a box is always on record. The default stays `500M` so that an upgrade changes nothing until the user chooses; the recommendation is `2.5G`. **One population per embedder, not one per detector · 2026-09-19.** The first cut of the setting keyed every reader on the full id — the clustering pass, the coverage figure, the sweep's work list, the shard export and import, and the sync merge's face matching — on the theory that a detector change is a model change. Measured on the reference library it was a disaster: choosing Thorough on both devices restarted coverage at 1,834 of 19,140, the People screen showed only the faces the new pipeline had reached, the desktop's 3,583 confirmations under the old id could not reach the tablet because the merge demanded the same id on both sides, and each device faced a ~400 GB re-fetch before the library looked whole again. The embedder is `w600k_mbf` in every variant; its vectors are one space, and the detector only decides where the boxes are. So every reader now keys on the embedder half of the id (`faces::embedder_of`, and `embedder_sql` for the queries): all three detectors are one population, and changing between them empties nothing. `record_detections` is unchanged — an image holds one pipeline's faces at a time, and a re-detection carries identities across by box overlap and embedding (§18) — and it is where the generations meet. The merge's `match_faces` matches within an embedder for the same reason. The shards travel every generation, each under its own id, and a peer adopts whichever it is sent. What a stronger choice still does is queue the images a weaker detector indexed for re-detection (`FaceDetector::supersedes`), after the ones nothing has indexed and never downwards, so a tablet on Fast keeps the desktop's Thorough faces rather than replacing them with fewer. The calibration (§8) is keyed on the embedder too: it is a fit over the similarity space, and that space did not change. --- ## 13. Order 1. **M1** — load both graphs in tract, freezing the input dims if needed (§4.1). Go/no-go, and it now comes first: it is an afternoon, and §2.3's reading is wasted effort if the graphs will not run at all. This is a change from the S14 brief's ordering, made because §1.1 removed the uncertainty that justified reading first — the models are known to work, so the open question is the runtime, not the choice. 2. **Licence reading** (§2.3), in parallel with 3. Before anything *ships*, per D13. 3. `dr-face` skeleton, `detect` + `align` + `embed`, ported from §1.1's table, with an example binary that draws boxes and landmarks on a JPEG — the same shape as `dr-segment`'s `examples/detect.rs`, and for the same reason: the thing worth looking at is whether the landmarks land on a real photograph. Port `tests/test_face_utils.cpp`'s cases first; they are model-free and they fail loudly on exactly the mistakes §5 describes. 4. M2–M5 on the corpus. Alignment is validated here, before anything depends on embedding quality. 5. `calibrate` + `cluster` against the labelled subset. M6–M8. The Platt fit ports from `gallery_calibration.hpp`; the pair *sourcing* (§8.1) is new and is the part to get wrong. 6. The v5 migration, the `DetectFaces` job, the debounced clustering pass. 7. UI: People view, confirm and reject, merge and split, the `Person` selector term. 8. The route-C first-run flow, the about-screen entry, and the delete-all control — with the feature shipped disabled until they exist. Steps 1–5 are the spike. Steps 6–8 are the build, and they are only justified if §12 says so. ### 13.1 Where this had got to · 2026-08-26 - **1 — done.** M1 passes conditionally; §12.1. - **2 — outstanding.** §2.3's three questions are unanswered and gate *shipping*, not building. - **3 — done.** `core/dr-face`: `align`, `detect`, `embed`, `embedding`, and `examples/faces.rs`. §12.2 is its first end-to-end run. - **4 — partly.** M2/M3 have provisional numbers; M4/M5 need the labelled corpus. - **5 — built, not yet measured.** `calibrate` and `cluster` exist with 34 model-free tests behind them; M6–M8 are a run over a real library, which is what the corpus in §1.1's `images/` is for. - **6 — done for storage.** Catalog schema v8 — `people`, `faces`, `face_person`, `face_person_rejected`, `face_calibration` — with the identity operations FR-CULL-10 requires. The `DetectFaces` job kind and the debounced clustering pass are not wired yet. - **7 — done.** The Identity screen: a third top-level mode beside library and develop, reached from the library header. People rail, face grid with per-face confirm/reject, rename in place, confirm all, split off a multi-selection, regroup, and the NFR-SEC-5 delete-everything control. - **8 — not started.** The route-C first-run flow: no model fetch, no checksum pin, no licence notice. The screen says "No face model installed" and stops, which is honest but is not the feature. The models go in `/models/` as the shape-fixed exports, by hand for now. - **9 — done, and not previously in this plan.** The run marker (§10a), the coverage audit, the `face_index` batch job, and cross-device sync of face shards (§14). The screen taught the design one thing worth recording. **Splitting has to reject before it confirms.** Moving faces to a new person is not enough on its own: the next clustering pass sees a face that still looks like the person it left, suggests it back, and the user's correction becomes an argument they keep having. §9's cannot-link constraint handles co-occurrence; this is the same idea applied to a judgement the user made by hand. Two things the build changed in this document's own design: **`crop_px` reached the schema** as §7 argued it should, and the calibration carries a `w_size` term for it. **A rejection table was added**, which §8 and §9 did not contemplate. Rejection is not the absence of an assignment: without storing it, the next clustering pass re-suggests exactly the face the user just pushed away. It is user data in the same sense a confirmation is (FR-CULL-12), just negative. --- ## 14. What this does not settle - **Whether a permissive model pair exists.** §2.3. If it does, route B replaces route C and step 8's first-run flow shrinks to nothing. - **Whether tract runs these graphs at all.** §12 M1, and §4.1 says why the answer is in real doubt. Every other line of this document is conditional on it. - **Faces in trashed images.** catalog.md §10.4's open question, unchanged: probably excluded from suggestions but not deleted, so a restore does not re-index. - ~~**Whether embeddings sync.**~~ **Answered: they do**, as sealed shards (`dr_catalog::face_shard`). Indexing is hours of CPU and its result is byte-identical on every device, so paying for it once per account rather than once per device is the whole argument. §12.2's 3.5 images/second is also what makes the case: it is fast enough to be worth doing and slow enough to be worth not repeating. **Shards rather than the catalog snapshot**, which is the design decision worth recording. The snapshot uploads whole on every sync, and a fully indexed 23.5k library carries ~30 MB of embeddings — exactly the cost `dr_thumbs`'s 25 MB cap exists to bound. So the split follows the one already in the tree: bulk immutable data in sealed shards, small mutable data in the snapshot. Faces, landmarks, embeddings and run markers shard; people, names and assignments ride the catalog and merge by uuid. The cap is *imported* from `dr_thumbs` rather than restated, because it is a statement about transfer cost and two copies of it would drift. Everything is keyed on `oc:fileid`, never `image_id`: a row id means nothing on another device. - **Re-embedding at higher resolution.** §7's `crop_px` makes it a query rather than a full re-index, but whether it is worth doing is an M4 question. - **Approximate nearest neighbours.** §9 says brute force until measured otherwise. A 100k-face library is where this stops being true. --- ## 15. Register entries **D13** — *face inference runtime and model licensing* · the runtime half stays answered (`ort` + `ort-tract`, unchanged since 2026-08-21). **The licensing half is answered conditionally by §2:** route C — ship the code, not the weights — unblocks the build without breaking GPL-3.0 or any distribution channel, and §2.3's reading may yet convert it to route B. Closing D13 outright waits on that reading. **D17** — *face model pair and distribution route* · **PROPOSED**. SCRFD-500M for detection, MobileFaceNet/ArcFace for embedding, weights obtained by the user rather than redistributed, with a pinned checksum and a licence notice shown before the first fetch. The alternative considered and rejected is committing the InsightFace weights (§2.1 route A), which is not available at any price. The model *pair* is better evidenced than a proposal usually is — §1's table is a measured comparison over a 2,418-identity gallery, not a literature reading — so the open half of this decision is the route, not the pair. Relates to: D13, D14, NFR-COMPAT-2, FR-CULL-8. **D18** — *porting from `scene-actor-extraction`* · **PROPOSED, and the easy half of a decision**. That project is MIT and by this project's author, so MIT-into-GPL-3.0-or-later is a compatible one-way combination requiring attribution, not permission. Ported files carry a header naming the origin. Worth recording because "we already have this working in another language" is exactly the provenance that goes undocumented and then cannot be answered three years later. **S14** — *face pipeline in Rust* · scope sharpened, and **substantially de-risked**, by this document. §12 replaces "measure per-image latency, cluster purity and calibration convergence" with ten specific measures. Two changes to the brief itself: its instruction to resolve the licence question *before writing any of it* is relaxed to "before shipping any of it", because M1 is an afternoon and is worth knowing first (§13); and its central question — whether the pipeline works at all — is largely answered in advance by §1.1, leaving the runtime and the domain shift from film frames to family snapshots as the real unknowns. --- ## 16. Requirements touched | ID | How this document addresses it | |---|---| | FR-CULL-8 | §4 detection, §6 embedding, §7 the proxy tier and its consequences, §10 the `DetectFaces` job, §18 the re-index | | FR-CULL-9 | §8 — pairs, fit, validity, and the reliability-diagram acceptance test | | FR-CULL-10 | §9 constrained agglomeration, confirmations as anchors, split by re-agglomeration; §18 what a re-index carries across | | FR-CULL-11 | §10 the `Person` selector term, confirmed-only by default | | FR-CULL-12 | §10 schema unchanged from catalog.md §10.1: embeddings derived, names to the sidecar | | NFR-SEC-5 | §11 — the obligations restated as structural properties, one of them CI-checkable | | NFR-COMPAT-2 | §2 — why the obvious weights cannot ship, and what does instead | | NFR-RES-2 | §1 the cheap model pair, §12 M2/M3/M10 on a phone | | NFR-ARCH-2 | §10 background priority, preempted by visible work | | FR-CULL-8a | §17 — eye state and sunglasses: the models, the crops, the readability rule, the measurements | | FR-CULL-13 | §17.3 — the reading is evidence: a chip that narrows and a badge that explains, and nothing that writes a judgement | --- ## 17. Eyes and sunglasses · 2026-09-19 FR-CULL-8a's eye state, under FR-CULL-13's rule. Three more models run over every face the pipeline already aligns, and what they produce is a **filter term** — "eyes open", on the people filter — and a badge on the People screen. Nothing acts on it: FR-CULL-13 says the signal is evidence, shown and filterable, and never holds the pen, and §3.9.1's exclusion of blink *detection* was re-read the same day as the exclusion of blink *selection* it always was. The chip narrows the grid the way a star count does, and rates nothing. Head pose, the other half of FR-CULL-8a, is not built; §17.3 says what stands in for it. ### 17.1 The models, and why these | | 2d106det | OCEC | SGC | |---|---|---|---| | Answers | 106 landmarks, ten round each eye's lids | P(this eye is open) | P(this head wears sunglasses) | | Input | 192² RGB 0..255, the detector box at 1.5× | one eye, 40×24 RGB, `x/255` | one head, 48×48 RGB, `x/255` | | Shipped | `buffalo_l`'s, 4.8 MB | S, 483 KB, F1 0.9943 on its own split | L, 6.1 MB, F1 0.9554 on its own split | | Licence | InsightFace's research-only grant, like the pair (§2.2a) | MIT, code and weights | MIT, code and weights | | Training data | InsightFace's | *Open and Closed Eyes* (ODC-By 1.0) + Wholebody34 crops (Apache 2.0) | **not stated** — recorded in `models/face/README.md` | | Cost in tract | ~24 ms per face | ~6 ms per eye | ~6 ms per framing, two framings | The two classifiers are from Katsuya Hyodo's ultra-lightweight series — the same author as the whole-body detector §1.1's reference pipeline uses — and are the first weights in `models/face/` that do not come out when the project publishes. The landmark model is under the grant the pair already carries; it was chosen over two permissively licensed alternatives on a measurement (§17.2) after the decision that this project will not be commercial, which is what §2.2a already records for the pair. All three load in tract as shipped, dynamic batch and all — the first graphs in this subsystem to do so — and are pinned to a batch of 1 by `tools/fix-face-model-shapes.sh` anyway, because a graph the engine *analyses* and a graph it has been *measured running* are different claims, and the embedder's precedent is the safer one. Six milliseconds per classifier call against 0.4 in the reference README is tract's per-call overhead on a graph this small; the whole reading is under 60 ms per face beside an embedding at 160 ms and a native decode in seconds. ### 17.2 Where the eye box comes from **The eye classifier was trained on a whole-body detector's eye boxes, and this pipeline has no eye boxes.** It has five landmarks, and SCRFD's eye point is loose: it is one of five points that place a face, not an eye centre, and on a turned or smiling head the eye sat in a corner of a window centred on it. Everything below was measured on 60 proxies from the reference library with 25 plainly open-eyed faces labelled by hand (`examples/eyes.rs --dump`, then a contact sheet), and the count that matters is how many of those 25 the classifier read as open in both eyes. **A window on the SCRFD point: 19 of 25.** Windows from 20×10 to 34×17 template units all gave 19–20; smaller lost eyes. Two model-free ways of re-centring the window were then tried and both lost eyes: the darkest blob near the landmark is the inner corner's shadow or the lash line (19 → 15), and the most contrasty window is the one that takes in the edge of the nose (19 → 9). The landmark as SCRFD gives it beats either. **A box from a landmark model's lid contour: 22 of 25.** Three models were run over the same faces, each fed the crop its reference code feeds it, and the eye box cut as the bounding box of the lid points grown by a margin: | model | points | input | tract | per face | open at margin 0.1 | |---|---|---|---|---|---| | MediaPipe Face Mesh V2 (Apache 2.0) | 478, with z | 256² | loads | ~36 ms | 22 | | PIPNet, PINTO's irnet18 export (WFLW, research-only) | 68 | 256² | loads | ~98 ms | 20 | | **InsightFace 2d106det** | **106** | **192²** | **loads** | **~24 ms** | **22** | The margin was swept on the two that tied: 22 at 0 and 0.1, 18 at 0.4, 14 at 0.6 — the training crops were tight detector boxes, and a tight box is what the classifier wants (`EYE_BOX_MARGIN`). 2d106det ships: it tied the best, costs the least, and is under a grant the project has already accepted. Face Mesh would be the choice if that changed; it also gives z and an iris, neither of which this needs yet. The box is cut **upright from the native render**, not through the face's alignment — the training crops were detector boxes, and the contour already says where the eye is on a tilted head (`align::eye_patch`). A shut eye's contour has no height and is given an open eye's (`EYE_BOX_MIN_ASPECT`), so the classifier sees the same framing either way. ### 17.3 Not asking what cannot be answered The 25 open faces were never the real problem. The real problem was the faces that were *not* open-eyed by the classifier's account and were not blinks either, and on the reference sample they were the commonest wrong answer of all: **a soft eye reads as closed.** A face small enough that its eye was seven pixels wide, a motion-blurred face, a face from a 1024 proxy where the native render should have been — each produced a confident "closed" from a classifier shown a smear. The same failure the face's own sharpness gate exists for (§4.3), one stage down, where the face's gate cannot see it: a face sharp enough to embed can hold an eye too soft to read, because the eye is a fortieth of it. So the reading is **seven numbers, not a verdict** — per eye P(open), the source pixels across its box and the sharpness of the patch the classifier saw; and P(sunglasses) — stored as such (`faces.eye_right`, `faces.eye_right_px`, `faces.eye_right_sharp`, likewise `eye_left`, and `faces.sunglasses`, schema V16), and the verdict is a rule with thresholds in it, `dr_face::eyes::EyeReading::state`, the only place the thresholds live: ``` sunglasses ≥ 0.5 → Sunglasses (whatever the eyes said) an eye is readable when px ≥ 12 and sharpness ≥ 0.02 and px ≥ 0.6 × the other eye's px no readable eye → Unreadable a readable eye < 0.5 → Closed (a blink, or a wink) otherwise → Open ``` **Sunglasses take precedence** because the eye classifier answers confidently over dark glass: over a woman in sunglasses on the reference library it read her right eye 0.97 open. **The pixel floor** is where the classifier's own training stopped — its reference footage averaged 15–21 pixels an eye. **The sharpness floor** is the face's measure over the patch, set where the sample's open eyes were being called closed: the open set ran from 0.019 (a lens reflection) to 5.4, the unreadable ones under 0.02 with the pixels to match. **The width ratio** is the profile: a landmark model's contour for the far eye of a turned head collapses towards the nose. On the twenty native renders of §17.4, profiles put the far eye at 0.02–0.43 of the near one's width, two three-quarter faces whose far eye was reading closed sat at 0.54, and every face looking at the camera sat at 0.78 or more — a shut eye's box keeps its width, so a wink is not mistaken for a turn. 0.6 splits the gap. An eye that fails any of the three is not asked, the near eye still decides, and a face with no readable eye is *unclear* — which is not a blink, and not open, and which no filter drops. **The two eyes are kept apart** rather than averaged, because a wink averages to 0.5 — the one value that says the least — and "eyes open" means every eye that could be read. With the rule in place, the same 60 proxies read: 33 open, 14 closed, 24 sunglasses, 14 unclear. Of the 25 labelled open faces, 22 open, 2 unclear (eye boxes of 7 and 12 pixels on a child's face), 1 closed — a squinting smile whose contour collapsed to eleven pixels, which the classifier is not wrong to call narrow. The 14 closed are downcast eyes, laughs, two sunglasses the head classifier missed, and the squint. A face 141 pixels across the eye but motion-blurred to a sharpness of 0.016 reads *unclear* where it read *closed* before, which is the change this section is for. **The filter drops only *Closed*.** `RatingFilter::eyes_open` compiles the rule above into a predicate on the face row, ANDed into the chosen people's face subquery, so "Anna, eyes open" asks about Anna's face and not about Bob blinking beside her. The chip is offered only while someone is chosen and goes when the last person does — without a name in front of it, it would be a verdict on everyone in the frame. The predicate still handles the empty case, as `NOT EXISTS` over every face, for a filter arriving by another route; a landscape passes because there is no one in it to have blinked. Sunglasses pass. Unclear passes. Never read passes — that last is what keeps an old library from emptying its grid the moment the chip is pressed: until the measuring pass has run, the honest answer is "everything". A test drives the same five readings through the SQL and through `state()` and requires the two to agree, so the badge and the grid cannot say different things. **And an index, learned the slow way.** The people filter was always served from a covering index on `faces(image_id)`; the moment its subquery read the eye columns it had to read the face *row*, and `ALTER TABLE ADD COLUMN` had put those seven floats after the embedding and the crop blob — six kilobytes to reach every one. One count took 24 seconds on the reference library, thirteen of them system time. `faces_eyes` (schema V17) covers the subquery again: five milliseconds. ### 17.4 What the sample says about accuracy, and what it does not The 60-proxy sample above was run at proxy resolution, where the production pass reads the native render; the eye box on a 200-pixel face is 40 source pixels from the proxy and 240 from the original. So the shipped configuration was also run over **twenty native renders** from the reference library — a wedding burst of six frames with six or seven faces each, and a dozen singles — exported by `face_native --export` and read by `examples/eyes.rs --dump`, 62 faces in all: 20 open, 28 closed, 12 sunglasses, 2 unclear before the width ratio was moved (below). Read off the contact sheet, face by face: the one real blink in the set (`7884.dng`, a man with his eyes shut) is *closed*; the laughing faces with their eyes screwed shut are *closed*, which a photographer would call right; the downcast faces are *closed*, which is arguable; the profiles are judged on the near eye and mostly *open*, which the SCRFD-point pass could not do. Two faces were wrong: three-quarter views whose far eye's box came to 0.54 of the near one's and read closed over a cheek, which is what moved the width ratio from 0.45 to 0.6. One is beyond any floor: a face with a porcelain pot held over the eyes, whose contour is a guess and whose "eyes" are sharp white china. Two things no sample so far can say. None holds more than one real blink, so the precision of *Closed* is not measured — every closed verdict on the two sheets but the occlusion and the sunglasses misses is a narrowed or shut eye rather than a wrong one, but that is a reading of a contact sheet, not a number. And the floors were set on a few dozen faces. Both are M11. ### 17.4a What is kept per face The box, the five landmarks, `crop_px`, the embedding and the crop were already there. The eye pass adds the seven numbers of §17.3 and the **106 dense landmarks** it read the eye boxes from, packed as 16-bit fixed point over the frame — 424 bytes a face, a seventh of a pixel on a 6000-pixel frame (`dr_face::Landmarks::to_packed_bytes`, schema V18). Kept for the reason the embedding is kept: it cost a fetch of the original and a model run, and the next per-face pass — head pose, expression, whatever FR-CULL-8a grows — should run from the catalog. `f16` would have been the same size and worse: three figures near 1.0 is six pixels at that scale. ### 17.5 The measuring pass, and shards A face indexed before the models existed, or on a device without them, has no reading. The `face-eyes` repair (§18.1; on the day this was written, the sweep's measuring pass — the one V14 built to re-embed faces stored as unit vectors) lists those faces, on a device that has the models, and reads their eyes from the same native render with the box and landmarks already stored. No detector runs and no identity moves. A device *without* the models has no such repair, or it would fetch every original in the library to do nothing to it; the repair's predicate is the one both the count and the work list use, so the pass converges. The People screen's coverage line counts these faces as work to read and keeps the button while any remain — reading **Read eye state** once detection is complete and only readings are left, which is the state an already-indexed library is in the day the models arrive. Shards carry the seven columns beside `quality`. A peer's faces without a reading are **adopted**, unlike a peer's faces without a quality (§14): the measuring pass finds this work by the NULL and not by the run marker, so adoption costs the reading nothing, and a peer with no eye models may be the only device that has done the detection at all. ### 17.6 Still to measure | # | Measure | Why it decides something | |---|---|---| | **M11** | Open-eye recall and blink precision from the **native** pass with the shipped configuration, on a labelled sample that contains real blinks — a burst with one in it is enough | §17.3's figures are proxy figures from 25 faces, with no precision beside them; this is the number FR-CULL-13's acceptance clause asks for, and where the two readability floors get set on more than 25 faces | | **M12** | Whether the remaining open-eye failure — a lens reflection over an open eye — moves with the sharpness floor, or needs the eye classifier told about spectacles | If the latter, the fix is a classifier trained closer to this domain, and that is a different decision | | **M13** | Sunglasses recall on more than twelve faces, and the false-positive rate on caps and clear glasses | The two sunglasses the head classifier missed on the sample became false blinks; a library of skiers would say whether that is two faces or a class | --- ## 18. The completeness job, and what a re-index carries across · 2026-09-19 The reference library on the day this was written: 17,762 faces under the bare `w600k_mbf` id — found by the fast detector on 1024 px proxies, stored as unit vectors, no quality, no crop on 4,144 of them, no eye reading, no dense landmarks — beside 1,177 under `scrfd_10g+w600k_mbf` from the native pass, and 12,217 images the fast detector examined and found nothing in. Over those faces: 3,778 confirmations, 13,011 suggestions, 77 rejections, and 17,276 people rows. Every one of those gaps was, until now, its own pass: V14's measuring pass for the quality, §17.5's for the eyes, the sweep's proxy repair, the sweep's detector upgrade, and a re-index that did not exist. Adding a per-face field meant adding a pass, with its own work list, its own count and its own idea of done. ### 18.1 One job over a registry `dr_ui::repairs` replaces them with one job over a **registry**. A `Repair` names one thing a catalog record can lack — the predicate that says which images still owe it, the input its handler needs (the file's header, the whole original, or a native render), the handler that fills it, and, where there is one, what to record for an image that can never be done. The job unions the predicates into one work list, fetches each image once at the most any claimant asks for, renders it at most once, and runs every handler whose predicate that image still matches — checked again before each, because one handler's write satisfies the next's (a detection writes every field a per-face handler would fill). The registry today: | Repair | Owed by | Input | Handler | |---|---|---|---| | `face-proxy` (sweep only) | images holding faces whose 1024 px proxy is not in the store | native render | detect again, write the proxy | | `face-quality` | faces with `quality IS NULL` | native render | warp from the stored landmarks, embed, write the raw vector and its length (reads eyes on the same warp where it can) | | `face-eyes` | faces with no eye reading or no dense landmarks, on a device with the eye models | native render | read the eyes and dense landmarks from the stored box and landmarks | | `face-crop` | faces with `crop IS NULL` | native render | cut the crop from the frame | | `face-detection` | sweep: images with no marker under the embedder and no faces; re-index: images with no marker under the **chosen detector**, either spelling | native render | detect, embed, replace the faces, carry identities across (§18.2) | | `face-upgrade` (sweep only) | images whose marker is a weaker detector's | native render | as `face-detection` | | `metadata` | `images.metadata_state < 2` | header | EXIF to the catalog, the dateless marked examined | A repair's predicate is the *only* definition of its work. The count the settings page shows (`faces::audit`, per repair), the list the job fetches and the check before each handler are one predicate, so a record the count reports is one the job fetches and one the handler fills, and the job ends. This is why an eye reading that cannot be cut is not a criterion — such a face stays unread however often it is detected, and listing it would fetch its original on every press — and why a degenerate face is dropped rather than left. It is also why the registry is cut to what the device can do (`Capabilities`): a device without the eye models has no `face-eyes` entry, rather than an entry it skips, because an entry is a count and a set of originals to fetch. The registry is ordered, and the order is the work list's: an image only the last repair claims comes after one the first does, which is what puts a few hundred proxy repairs ahead of twenty thousand un-indexed images. Within one image the same order runs the handlers, detection before the per-face repairs, since detection fills what they would. Adding a field is one entry: a predicate over `faces f` or `images i`, and a handler that fills it from `Fetched`. `metadata` is in the table to say that this is not a face job — the same machinery carries a capture date, and could carry a thumbnail, a perceptual hash or a head pose. ### 18.1a The two scopes Both buttons on the settings page run the job; they differ in one predicate. **Index faces** (`Scope::Outstanding`) converges on *coverage* — has anything examined this image — and treats a face a weaker detector found on a proxy as found, which is the right question for a pass that must not fetch the library twice. **Re-index every face** (`Scope::Reindex`) converges on *provenance*: `face-detection` claims every image with no marker under the chosen detector, in either of its forms (`FaceDetector::model_ids`, so a desktop running it in f32 and a tablet on the Hexagon in int8 do not re-index each other's work), and a marker saying a weaker one looked is not that. This is the one place in the subsystem keyed on the exact detector rather than the embedder. Convergent all the same: an image the job has been through leaves the list, a kill costs the images in flight, and a second press resumes. An original over the fetch budget is skipped without being fetched. Under the sweep, detection records an examination that found nothing — the honest record for an image that cannot be examined, and what stops the half gigabyte being spent once per sweep. Under the re-index, and under every repair over records that already exist, it is left exactly as it was: a re-detection with nothing found would delete the faces, and "cannot fetch" is not "no faces". ### 18.2 What is carried across `dr_catalog::faces::record_detections` replaces an image's faces and carries identities onto the new ones. Before this section it carried confirmations only, by box overlap above 0.5 IoU, and a re-detection of the library above would have left 13,011 suggestions and 77 rejections on the floor — correct by FR-CULL-12's letter, since suggestions are derived data, and a People screen emptied to strangers by the user's own button. Now every old face is read before the delete — box, vector, assignment, rejections — and matched to the new faces one-to-one, best pair first. A pair qualifies when the boxes **overlap at all** and either the overlap alone says so (IoU above 0.5, the old rule) or the embeddings do (cosine above `SAME_FACE_COSINE` = 0.45, the reference library's P≈0.95 line from §9's table). The embedding route is for the box a low-resolution pass drew badly enough that overlap alone would not claim it; the vector is also what breaks the tie in a group photograph, where two neighbouring faces overlap both new boxes. Overlap is required on both routes because the same vector elsewhere in the frame — a mirror, a print on the wall — is not the same face and must not take its name. Onto the matched face go the assignment as it was, confirmed or suggested with its probability, and every rejection. It is a match, not an update in place, and that is why the per-face repairs exist beside detection: where nothing about a face but one field needs doing, `record_updates` keeps the id and there is nothing to judge. Since #77 (0.18.0) the merge's `match_faces` answers it too, within a photograph's `file_id` and one embedder: box IoU ≥ 0.5, unique on both sides, first; then, only for photographs where a remote face is left over and a local face is free, embedding cosine ≥ 0.7, mutual best, with a lead of ≥ 0.2 over the runner-up on both sides. A box match is never overruled by a low cosine (about 150 genuine cross-device pairs of tiny faces score below 0.45). On the reference desktop/tablet pair this recovers 20 of 631 unmatched faces with no false matches; the rest are faces one device alone found. The merge also keeps one person to one face per photograph: an incoming assignment is refused when another local face already holds that person, unless it is a remote confirmation over a local suggestion, which moves the suggestion. Refusals are counted in `faces_one_per_photograph`. The threshold differs from `SAME_FACE_COSINE` (0.45) above on purpose: re-detection additionally requires the boxes to overlap, while the merge's embedding route exists for boxes that don't. ## 19. Deduplicating people · 2026-09-26 `dr_catalog::dedup_people::run` runs after every successful sync merge (`sync::merge_remote`, on the sync worker), in one transaction, and logs one `dedup:` line (#78). **People.** Named people with the same name, trimmed and case-folded, merge into the one with the most confirmed faces (ties go to the smaller uuid) when every shared embedder's confirmed-face centroids agree at cosine ≥ 0.7 (distance < 0.3). Each side needs at least two confirmed faces to compare; a namesake holding no faces merges outright; a face confirmed as one and rejected as the other keeps them apart; unnamed and set-aside people are never touched. On the reference library the same-person centroid median is 0.91, and different named people have a 99.9th percentile of 0.41. **Faces.** Two faces in the same image and embedder with IoU ≥ 0.5 and cosine ≥ 0.7 are one: the job keeps the stronger detector's face (`FaceDetector::outranks`), then the confirmed one, then the lower id, and it takes both faces' assignment and rejections. **Propagation.** The merge is `faces::merge_people`, whose `merged_into` redirect a 0.17.0 peer already honours, so an older device never resurrects the duplicate. The job also follows redirects left by earlier manual merges, moving this device's own assignments onto the person kept, and breaks a mutual redirect at the smaller uuid, which every device computes alike. A merge now also carries the merged-away person's rejections to the person kept.