catalog.md §8.2 still said confirmed and rejected faces were matched to local faces "by box", and that the merge reads only a remote face's box and model; faces.md said `match_faces` "still matches by overlap alone across devices". Since #77 the match falls back to embeddings, on the photographs where a box leaves a remote face over, and since #78 `dedup_people` folds people of one name whose faces agree after every sync. Both now say so, with the thresholds and the reason a less decisive pair stays unmatched, taken from the code's own documentation.
1612 lines
104 KiB
Markdown
1612 lines
104 KiB
Markdown
# Faces and identity — SCRFD and MobileFaceNet
|
||
|
||
Spec for **S14**, the spike that decides whether §3.9.1 is buildable, and the build that follows it.
|
||
|
||
FR-CULL-8 … FR-CULL-12 specify the subsystem in terms of "a 512-dimension embedding from a stated
|
||
model" and stop there, deliberately: [catalog.md §10.4](catalog.md) records that the model was left
|
||
abstract because D13 was open. This document names the two models, fixes the pre- and
|
||
post-processing they need, specifies the calibration and clustering that sit on top, and states what
|
||
S14 has to measure before any of it is trusted.
|
||
|
||
**It does not close the licensing half of D13.** §2 is the reason, and it comes first because
|
||
[segmentation.md §7](segmentation.md) established the precedent that reading the grant is cheaper
|
||
than discovering it at packaging time.
|
||
|
||
---
|
||
|
||
## 1. The two models, and why these two
|
||
|
||
**Detector — SCRFD.** *Sample and Computation Redistribution for Efficient Face Detection* (Guo et
|
||
al., ICLR 2022). A single-stage anchor-based detector whose contribution is where the FLOPs go, not
|
||
a new architecture: the shallow stages carry more capacity because that is where small faces are
|
||
decided. It emits, per face, a box, a confidence, and **five landmarks in the same forward pass** —
|
||
which is the property that matters here, because FR-CULL-8 requires landmarks and because §4's
|
||
alignment step is not optional. A detector without landmarks would need a second network to supply
|
||
them.
|
||
|
||
The `500M` variant is the one to ship: 500 MFLOPs at VGA, against 10 GFLOPs for `10G`. Face indexing
|
||
is a background sweep over a whole library on a phone as well as a desktop (NFR-RES-2), so the cheap
|
||
variant's accuracy loss on tiny faces is the right trade — a face too small for `500M` to find is
|
||
also too small for §5 to embed usefully.
|
||
|
||
**Embedder — MobileFaceNet trained with ArcFace loss ("MBF").** ~1M parameters, ~0.45 GFLOPs for one
|
||
112×112 crop, 512-d output. ArcFace's additive angular margin is what makes the output space one
|
||
where *cosine similarity means something* — the loss explicitly optimises angular separation between
|
||
identities, which is what §6's calibration then has to convert into a probability.
|
||
|
||
**Why not the larger ResNet50 embedder** (`w600k_r50`, the other half of InsightFace's `buffalo_l`):
|
||
because it is not measurably better here, and it is 13× the file. §1.1's reference project fitted the
|
||
same Platt calibration (§8) to all four candidates over a 2,418-identity gallery, which is a clean
|
||
read on the embedding space alone with no tracking or matching logic in it:
|
||
|
||
| Model | File | Steepness `a` | Boundary at P=0.5 | Held-out macro F1 |
|
||
|---|---|---|---|---|
|
||
| LVFace-B (Glint360K) | 455 MB | 17.7 | sim 0.228 | 67.4% |
|
||
| **ArcFace w600k-MBF** | **13 MB** | **16.2** | **sim 0.267** | **64.4%** |
|
||
| ArcFace w600k-R50 | 174 MB | 15.4 | sim 0.301 | — |
|
||
| ArcFace R18 | 48 MB | 15.3 | sim 0.309 | 63.1% |
|
||
|
||
**MBF's calibration curve is steeper than R50's and R18's**, separating same-identity from
|
||
different-identity pairs more confidently and at a lower boundary, at a fraction of the size. LVFace-B
|
||
is genuinely better and is 35× the file — a defensible choice for a desktop-only feature and not one
|
||
for a subsystem that has to index a library on a phone (NFR-RES-2).
|
||
|
||
Two caveats on that table, both from the source: the R50 gallery was built with ~30% fewer reference
|
||
images per identity, which is why it carries no F1 figure — the calibration row does not depend on the
|
||
image count and is comparable, the benchmark row would not have been. And the F1 column measures a
|
||
*film-cast identification* task, not this one. It ranks the models; it does not predict DarkRoom's
|
||
accuracy.
|
||
|
||
**Both are pure convolutional graphs**, which matters for §3: the runtime is tract, and tract's
|
||
operator coverage is the thing that decides whether a graph runs at all.
|
||
|
||
### 1.1 A working reference implementation exists
|
||
|
||
`../scene-actor-extraction` is a C++ pipeline by the same author that runs **exactly this model
|
||
pair** — SCRFD-500MF then ArcFace over 5-point-aligned 112×112 crops — over feature films, with a
|
||
fitted Platt calibration and a scored benchmark behind it. It is **MIT licensed**, so porting from it
|
||
into GPL-3.0-or-later is straightforward and needs attribution, not permission.
|
||
|
||
That changes what S14 is. The open questions are no longer "does this pair work" and "what are the
|
||
magic numbers" — they are *does tract load these graphs* (§12 M1, and §4.1 says why that is in
|
||
genuine doubt) and *does the accuracy hold on family snapshots rather than film frames*. The
|
||
pre-processing constants, the decode layout, and the calibration algorithm are all readable rather
|
||
than rediscoverable, and several of them are not what the published Python would lead you to write.
|
||
|
||
| What | Reference | Ports to |
|
||
|---|---|---|
|
||
| ArcFace 5-point template and warp | `src/face_utils.hpp` | `align.rs` (§5) |
|
||
| SCRFD pre-process, decode, NMS | `src/backends/ort_backend.cpp` `SCRFDDecoder` | `detect.rs` (§4) |
|
||
| ArcFace pre-process and L2 normalise | same file, `ArcFaceEmbedder` | `embed.rs` (§6) |
|
||
| Platt fit over a similarity histogram | `src/gallery/gallery_calibration.hpp` | `calibrate.rs` (§8) |
|
||
| Model-free unit tests for both | `tests/test_face_utils.cpp`, `tests/test_calibration.cpp` | the §3 feature split |
|
||
|
||
That last row is worth noticing: the reference already separates the geometry and arithmetic from the
|
||
inference well enough to unit-test them with no model on the machine. §3's `inference` feature flag is
|
||
the same boundary, enforced by Cargo rather than by discipline.
|
||
|
||
**What does not port.** The reference matches faces against a *known gallery* of named identities;
|
||
DarkRoom clusters an unlabelled library (§9). And it is a video pipeline, so it has a tracker, a
|
||
`max_faces` cap, and frame-to-frame identity annealing that have no analogue here.
|
||
|
||
---
|
||
|
||
## 2. Licensing — the open half of D13
|
||
|
||
The architectures are published research. **The weights everyone actually uses are not
|
||
redistributable by this project.**
|
||
|
||
InsightFace's code is MIT. Its *pretrained models* — `buffalo_l`, `buffalo_s`, `buffalo_sc`, and
|
||
every `det_*` and `w600k_*` checkpoint inside them — carry a **non-commercial research-only** grant,
|
||
stated in the model-zoo README and applying to both the auto-downloaded and the manually downloaded
|
||
artefacts. `w600k_mbf` is trained on WebFace600K, whose own terms are research-only as well, so the
|
||
restriction has two independent sources rather than one that might be renegotiated.
|
||
|
||
These are the exact artefacts in question — §1.1's reference project fetches them from InsightFace's
|
||
own release page, which is where the grant attaches:
|
||
|
||
| File | Source |
|
||
|---|---|
|
||
| `det_500m.onnx` (2.5 MB) | `insightface/releases/download/v0.7/buffalo_sc.zip` |
|
||
| `w600k_mbf.onnx` (13 MB) | `insightface/releases/download/v0.7/buffalo_s.zip` |
|
||
|
||
DarkRoom is GPL-3.0-or-later. A non-commercial field-of-use restriction is not a GPL-compatible
|
||
grant, so:
|
||
|
||
- these weights are **not redistributable under this project's licence**, so they cannot go into a
|
||
*published* build — an APK, a Flatpak, or an F-Droid entry — and could not go into a repository
|
||
intended to be one (§2.2a records what was actually decided here, and why the two differ);
|
||
- and the restriction binds the *user*, not only the project — a professional photographer using
|
||
DarkRoom commercially is outside the grant even if they fetched the file themselves.
|
||
|
||
That last point is the one that is easy to miss and dishonest to leave implicit. It is why §2.2's
|
||
recommended route puts the licence text in front of the user rather than in a footnote.
|
||
|
||
### 2.1 The routes, priced
|
||
|
||
| Route | What ships | Cost |
|
||
|---|---|---|
|
||
| **A — commit the weights** | Everything works out of the box, one `git lfs pull` | **Not available.** Licence-incompatible with GPL-3.0 and with all three distribution channels (NFR-COMPAT-2). Listed only so it is visibly rejected rather than quietly assumed. |
|
||
| **B — permissively licensed weights** | Same, legally | Requires weights that do not exist off the shelf in this architecture pair. Either found (§2.3) or trained, and training an ArcFace embedder needs a face dataset whose own terms permit redistribution — the harder half of the problem. |
|
||
| **C — the user fetches them** | The app ships the *code*, not the weights; face indexing is off until the user obtains a model | Works today, keeps the repository clean, and is honest. Costs a first-run flow, an F-Droid anti-feature declaration, and a licence notice the user has to read. |
|
||
|
||
**Recommendation: C now, B when it becomes possible.** §1.1's project already works this way — a
|
||
`scripts/download_models.sh` that fetches the weights rather than a repository that carries them —
|
||
so route C is a practised pattern here rather than a theory. Two things DarkRoom must do that the
|
||
reference does not: pin a checksum for **every** model (the reference verifies only one of the four),
|
||
and put the licence in front of the user, because a photo editor's users are not all researchers.
|
||
|
||
The schema already forces this to be a survivable choice — `faces.model_id` (catalog.md §10.1) exists precisely so that a model change is
|
||
detectable and re-indexable rather than silently poisoning every similarity in the library. Under C,
|
||
swapping in a permissive model later is a new `model_id` and a re-index, not a migration.
|
||
|
||
### 2.2 What route C requires of the code
|
||
|
||
- **The weights are never a build input.** `dr-face` (§3) takes bytes; it has no `embedded-model`
|
||
feature and no `models/` directory. This is the one structural difference from `dr-segment`, and
|
||
it is deliberate — a feature flag that *could* embed weights is a feature flag someone eventually
|
||
turns on in a packaging script.
|
||
- **The fetch does not live in `dr-face`.** It lives in the app layer, so the inference crate keeps
|
||
no network dependency at all. §11 makes that a checkable property rather than a convention.
|
||
- **The licence is shown, not linked.** Before the first download, the app states in plain language
|
||
that the weights are research-only, that commercial use is outside the grant, and who the grantor
|
||
is. NFR-SEC-5's last bullet already requires the models to be named, versioned and licensed in the
|
||
about screen; this is the same obligation moved to the moment where it can still change a
|
||
decision.
|
||
- **A checksum is pinned.** The app verifies the digest of what it fetched against a value compiled
|
||
in, and refuses a mismatch. An unverified model file is an arbitrary graph executed over the
|
||
user's family photographs.
|
||
- **Indexing stays off until a model is present and verified**, and disabling it deletes nothing —
|
||
that is a separate control (NFR-SEC-5, §11).
|
||
|
||
### 2.2a Android, which route C forgot
|
||
|
||
Route C says "the user obtains the model and the app loads it". On a desktop that is a real gesture:
|
||
the files go in `~/.local/share/darkroom/models/` and face indexing starts working. **On Android that
|
||
gesture does not exist.** `internal_data_path` is app-private, `adb shell run-as` needs a debuggable
|
||
build, there is no picker and no fetch in the app, and so a phone could not acquire a model by any
|
||
means at all. Face indexing was not "off until the user supplies weights" there; it was off, full
|
||
stop, and the settings page said so on every launch with no action available behind the message.
|
||
|
||
**Decision, 2026-08-27: the shape-fixed pair is committed to LFS at
|
||
`apps/darkroom-android/android/assets/models/`, and `assemble-apk.sh` bundles it into the APK.**
|
||
`android_main` unpacks it to the shared models directory on first launch, before anything asks
|
||
whether a model is present. This is a personal project on a private Gitea; the grant restricts
|
||
redistribution, and a private repository and a self-installed APK are not that.
|
||
|
||
What that does and does not settle:
|
||
|
||
- **It does not make §2.1's route A available.** The moment this project publishes — an F-Droid
|
||
entry, a release APK, a Flatpak — these files come back out and route C's unbuilt half (the fetch,
|
||
the licence screen, the pinned checksum) has to exist first. §2.1's table stands as the answer for
|
||
a *published* build; this is the answer for the author's own phone.
|
||
- **It does not make the weights a build input.** `dr-face` still has no `models/` directory and no
|
||
`embedded-model` feature, and nothing in the cargo build reads these files — §2.2's first bullet
|
||
guards against a flag someone flips in a packaging script, and that remains guarded. The APK
|
||
assembly step copies two files; it is the only thing in the tree that knows they exist.
|
||
- **It does not bind only the project.** §2's last bullet is unchanged and is the one with teeth: the
|
||
research-only grant restricts the *user*, so commercial photography with a DarkRoom that has these
|
||
weights in it is outside the grant no matter who put them there.
|
||
|
||
The in-app fetch, the licence screen, and the pinned checksum that §2.2 specifies are still unbuilt,
|
||
on every platform. When they are built, this becomes redundant and the assets directory empties.
|
||
|
||
### 2.3 Before S14 writes any code
|
||
|
||
Three questions to answer by reading, in this order, and to record with the date they were checked —
|
||
the same discipline [segmentation.md §7](segmentation.md) applied:
|
||
|
||
1. **Is there a permissively licensed SCRFD export?** Third-party ONNX exports of SCRFD are
|
||
plentiful; a re-export inherits the original weights' grant, so the question is whether anyone has
|
||
*trained* SCRFD on a redistributable dataset, not whether they have converted the InsightFace one.
|
||
2. **Is there a permissively licensed 512-d MobileFaceNet/ArcFace checkpoint?** Candidates worth
|
||
reading the terms on rather than assuming: EdgeFace (Idiap), the ONNX Model Zoo's ArcFace entry
|
||
(MIT, but a ResNet100 at ~250 MB — licence-clean and the wrong size), and any MobileFaceNet
|
||
trained on a dataset with distribution terms.
|
||
3. **What does a permissive *detector* alternative cost in accuracy?** YuNet (OpenCV Zoo) is 232 KB,
|
||
emits five landmarks, and comes from a permissively licensed repository. If it is close enough,
|
||
the detector half of the licence problem disappears and only the embedder remains — a materially
|
||
better position than either half alone.
|
||
|
||
Question 3 is nearly free to answer: `face_detection_yunet_2023mar.onnx` is **already sitting in
|
||
§1.1's `models/` directory**, fetched from `opencv/opencv_zoo`. It needs its own decode path (its
|
||
output layout is different, which is why the reference's `SCRFDDecoder` explicitly rejects it at load
|
||
rather than silently misreading it), and then it is one more row in §12's table.
|
||
|
||
---
|
||
|
||
## 3. Crate shape — `core/dr-face`
|
||
|
||
A new workspace member, modelled on `dr-segment` and for the same reason: everything that reasons
|
||
about faces is testable with no GPU adapter and no model present (ARCH §6.5a).
|
||
|
||
```
|
||
core/dr-face/
|
||
src/lib.rs FaceError, re-exports
|
||
src/detect.rs SCRFD: preprocess, decode, NMS
|
||
src/align.rs 5-point similarity transform → 112×112 crop
|
||
src/embed.rs MBF: preprocess, forward, L2 normalise
|
||
src/calibrate.rs cosine → P(same person) (FR-CULL-9)
|
||
src/cluster.rs constrained agglomeration (FR-CULL-10)
|
||
src/assign.rs which person, and how sure (§9.1, FR-CULL-9)
|
||
```
|
||
|
||
```toml
|
||
[features]
|
||
# No `embedded-model`. §2.2 — the weights are not a build input, ever.
|
||
default = []
|
||
# The ONNX runtime. Off by default so `calibrate` and `cluster` — which are
|
||
# arithmetic over embeddings and have no model in them — stay testable in a
|
||
# build that carries no inference engine at all.
|
||
inference = ["dep:ort", "dep:ort-tract", "dep:ndarray"]
|
||
```
|
||
|
||
The `ort` + `ort-tract` pairing is settled by D13's 2026-08-21 update and already in the workspace
|
||
manifest: `ort`'s API, tract's pure-Rust engine, no C under the NDK.
|
||
|
||
**The split between `inference` and the rest is load-bearing.** Calibration and clustering are where
|
||
the subsystem's accuracy actually lives, they are pure functions of embeddings and labels, and they
|
||
must be testable against synthetic embedding sets with no weights on the machine. A test suite that
|
||
needs a research-licensed download to run is a test suite that does not run in CI.
|
||
|
||
### 3.1 The API
|
||
|
||
```rust
|
||
/// A loaded detector. Fixed input shape — see §4.
|
||
pub struct Detector { /* session, input edge */ }
|
||
|
||
impl Detector {
|
||
pub fn from_bytes(onnx: &[u8], id: ModelId) -> Result<Self, FaceError>;
|
||
/// `rgb` is f32 0..=1, row-major, three per pixel.
|
||
pub fn detect(&self, rgb: &[f32], w: usize, h: usize, opts: &DetectOptions)
|
||
-> Result<Vec<Detection>, FaceError>;
|
||
}
|
||
|
||
pub struct Detection {
|
||
/// Normalised to the image's long edge (catalog.md §10.1).
|
||
pub bbox: (f32, f32, f32, f32),
|
||
/// Five points, same normalisation, in the model's own order (§5).
|
||
pub landmarks: [(f32, f32); 5],
|
||
pub confidence: f32,
|
||
}
|
||
|
||
pub struct Embedder { /* session */ }
|
||
|
||
impl Embedder {
|
||
pub fn from_bytes(onnx: &[u8], id: ModelId) -> Result<Self, FaceError>;
|
||
/// Takes an *aligned* 112×112 crop. Alignment is `align::warp`; passing a
|
||
/// raw bbox crop here is the silent-accuracy-loss bug §5 exists to prevent.
|
||
pub fn embed(&self, aligned: &Aligned112) -> Result<Embedding, FaceError>;
|
||
}
|
||
|
||
/// L2-normalised, 512-d. Carries its model id so a comparison across models
|
||
/// is a type error rather than a plausible-looking number (catalog.md §10.1).
|
||
pub struct Embedding { pub model: ModelId, pub v: [f32; 512] }
|
||
```
|
||
|
||
`Aligned112` is a newtype over the pixel buffer that only `align::warp` can construct. That is the
|
||
whole defence against §5's failure mode, and it costs nothing.
|
||
|
||
---
|
||
|
||
## 4. Detection
|
||
|
||
### 4.1 Fixed input, and how the image is fitted to it
|
||
|
||
**`det_500m.onnx` has a dynamic H/W input, and this is the largest single risk in the document.**
|
||
§1.1's reference had to use ONNX Runtime rather than OpenCV's `dnn` module precisely because OpenCV
|
||
could not load SCRFD's dynamic `Shape` nodes. tract is in the same family of problem: `dr-segment`
|
||
exists in its current shape because tract failed shape inference on YOLO's dynamic export, which is
|
||
why `tools/export-seg-model.sh` passes `dynamic=False` and why `semantic.rs` has a fixed
|
||
`INPUT_EDGE`.
|
||
|
||
So the graph almost certainly needs its input dimensions frozen before tract will touch it. That is a
|
||
known, cheap operation — `onnxruntime.tools.make_dynamic_shape_fixed` rewrites the declared dims
|
||
without retraining or re-exporting from PyTorch — but it is a step, it produces an artefact that must
|
||
itself be reproducible (a `tools/fix-face-model-shapes.sh`, in the spirit of the existing export
|
||
script), and **it must be tried before anything else in this document is scheduled.** §12's M1.
|
||
|
||
Fixed at **640**, with 320 available as a faster, blinder option to measure.
|
||
|
||
**Letterbox.** Scale by `min(640/w, 640/h)` preserving aspect, then paste into a 640×640 canvas. The
|
||
reference centres the image and fills the margin with grey `114`, inverting with
|
||
`x_src = (x_model − pad_x) / scale`. What matters is not *where* the padding goes but that the
|
||
forward and inverse agree and that the fill value is treated as part of the contract: a mismatch
|
||
between them offsets every box and landmark the model returns by the padding, which yields detections
|
||
that look plausible, landmarks that align badly, and an embedding-quality loss that only shows up as
|
||
poor clustering three stages later. Port the reference's convention rather than inventing one, and
|
||
assert it with a round-trip test.
|
||
|
||
Normalisation is `(x·255 − 127.5) / 128`, RGB, NCHW — note `/128`, not `/127.5`; see §6.
|
||
|
||
### 4.2 Outputs, and decoding them
|
||
|
||
**Nine tensors for three strides `{8, 16, 32}`, twelve for four `{8, 16, 32, 64}`.** Which one a
|
||
given export produces is discovered at load — `strides = output_count / 3` — not assumed, because
|
||
both variants exist and hardcoding three silently ignores the largest faces in a twelve-output model.
|
||
With two anchors per location and `N_s = (640/s)²` locations:
|
||
|
||
| Tensor | Shape | Meaning |
|
||
|---|---|---|
|
||
| `score_s` | `[N_s·2, 1]` | sigmoid already applied inside the graph |
|
||
| `bbox_s` | `[N_s·2, 4]` | distances left, top, right, bottom **in units of the stride** |
|
||
| `kps_s` | `[N_s·2, 10]` | five `(dx, dy)` offsets, same units |
|
||
|
||
Anchor centres are `(x·s, y·s)` for each grid cell, repeated once per anchor. Decoding is therefore
|
||
|
||
```
|
||
x1 = cx − l·s x2 = cx + r·s kx_i = cx + dx_i·s
|
||
y1 = cy − t·s y2 = cy + b·s ky_i = cy + dy_i·s
|
||
```
|
||
|
||
then divide by the letterbox scale to return to source pixels, then normalise by the long edge before
|
||
storage.
|
||
|
||
Flat index within a stride is `(row · fw + col) · 2 + anchor`.
|
||
|
||
**A shape assertion at load time, not a decode-time surprise.** `dr-segment` already learned this —
|
||
`SegmentError::OutputShape` exists because a different export of the same model produces confidently
|
||
wrong results otherwise. The reference implements exactly this and it is worth porting verbatim:
|
||
output count divisible by three and between 9 and 12, then the **last dimension of each group checked
|
||
against `{1, 4, 10}`**. That single check is what catches a YuNet file passed where an SCRFD one was
|
||
meant — YuNet also has twelve outputs, so the count alone does not distinguish them, and the failure
|
||
without the check is not an error but a page of plausible garbage.
|
||
|
||
### 4.3 Thresholds
|
||
|
||
Score ≥ **0.5**, NMS IoU **0.4**, plain greedy NMS per image across all strides together.
|
||
|
||
Deliberately *not* the low threshold `dr-segment` chose. There, a false positive costs one spurious
|
||
entry in a list the user is picking from. Here it costs an entry in the People view that the user has
|
||
to reject, in a library with thousands of images — and worse, a garbage embedding that participates
|
||
in clustering and can bridge two real clusters into one. False negatives are recoverable by a later
|
||
re-index with a better model; a polluted cluster graph is not, once the user has confirmed faces
|
||
inside it.
|
||
|
||
A minimum box size of **40 px on the source proxy** is applied on top — the reference's figure,
|
||
chosen there for the same reason it holds here: below it there is not enough face left to align
|
||
reliably, and §7's table shows what the embedder actually receives at that size. §12's M4 is what
|
||
moves this number, if measurement says it should move.
|
||
|
||
**One thing not to port:** the reference caps detections at ten per frame, largest first. That is
|
||
right for a film frame, where the extras in the background are noise. It is wrong for a photo
|
||
library, where a group shot with thirty faces in it is precisely the picture the user wants indexed.
|
||
No cap; the min-size floor is the only filter.
|
||
|
||
---
|
||
|
||
## 5. Alignment — the step that is silently wrong when skipped
|
||
|
||
ArcFace embeddings are trained on faces warped to a canonical 112×112 arrangement. Feeding the model
|
||
a plain bounding-box crop *works* — it produces 512 numbers, they are unit-norm, and cosine
|
||
similarities between them look entirely reasonable. They are just much worse, and nothing in the
|
||
system reports it.
|
||
|
||
The transform is a **similarity transform** — rotation, uniform scale, translation, four degrees of
|
||
freedom — from the five detected landmarks to this template, which is the arrangement the weights were
|
||
trained against:
|
||
|
||
```
|
||
(38.2946, 51.6963) subject's right eye ─┐ image-left of centre
|
||
(73.5318, 51.5014) subject's left eye ─┘
|
||
(56.0252, 71.7366) nose tip
|
||
(41.5493, 92.3655) subject's right mouth corner
|
||
(70.7299, 92.2041) subject's left mouth corner
|
||
```
|
||
|
||
Similarity, not affine: a full affine fit to five noisy points will happily shear a face, and shear is
|
||
exactly the deformation the embedder was never shown.
|
||
|
||
**The naming is a trap and the order is not.** Point 0 sits at x=38 on a 112-wide canvas — left of
|
||
centre *in the image*, which is the subject's **right** eye. Both namings are in circulation and they
|
||
are opposite. What matters is that SCRFD emits its five points in this same order, so the correct
|
||
amount of reordering between detector and template is **none**; the reference states this explicitly
|
||
in `types.hpp` for the benefit of whoever next reads it and doubts it. A future detector with a
|
||
different order carries its own permutation next to its `model_id`, rather than this file growing an
|
||
assumption.
|
||
|
||
**One divergence from the reference to settle by measurement.** It fits the transform with OpenCV's
|
||
`estimateAffinePartial2D` under **RANSAC** at a 3-pixel threshold, and drops the detection when the
|
||
fit fails. RANSAC over five points is a strange fit — the minimal sample for a similarity is two, so
|
||
it can discard landmarks it judges outliers and solve from a subset, which on a profile face is as
|
||
likely to be the correct geometry as the wrong one. InsightFace's own pipeline uses a plain
|
||
least-squares (Umeyama) similarity over all five points, which cannot silently drop anything and is
|
||
the one to write first. §12's M5 measures whether the choice matters; the degenerate-input path still
|
||
has to exist either way, because collinear landmarks do occur.
|
||
|
||
**A trick worth keeping.** The reference retries detection on a frame where nothing was found, after
|
||
replicate-padding by 25% and applying CLAHE to the L channel. Padding gives a face at the very edge a
|
||
context the detector needs; CLAHE rescues underexposed frames. Both are plausible on scanned negatives
|
||
and backlit photographs — worth trying at low priority, since a retry doubles the cost of exactly the
|
||
images that yielded nothing.
|
||
|
||
Sampling is bilinear from the **native render**, in one step — never crop-then-warp, which resamples
|
||
twice and throws away detail the warp could have used, and never from the downscaled buffer the
|
||
detector was given. The landmarks arrive in detector-input coordinates and are scaled back to
|
||
native before the warp reads a single pixel; §7 is why. Pixels falling outside the source are black.
|
||
|
||
**The landmark order must match the template order.** The template above is written in the detector's
|
||
own output order; if a future detector emits them differently, the template is reordered with it and
|
||
that mapping belongs next to the model id, not compiled in as an assumption.
|
||
|
||
*Test:* warp a synthetic image with a known rotation and scale, and assert the five points land on
|
||
the template within a fraction of a pixel. This is testable with no model present, which is why
|
||
`align` sits outside the `inference` feature.
|
||
|
||
---
|
||
|
||
## 6. Embedding
|
||
|
||
112×112 **RGB** — the crop comes out of the warp in whatever order the source was in, and the
|
||
reference converts BGR→RGB explicitly before the blob rather than relying on a flag, which is the
|
||
readable way to do it.
|
||
|
||
Normalisation is `(x·255 − 127.5) / 128`. **Note `/128`, not `/127.5`:** InsightFace's published
|
||
Python uses `1.0/127.5` for the recognition model, the reference uses `1.0/128` for both models, and
|
||
every measured number in §1's table was produced with `/128`. The difference is 0.4% of scale and
|
||
almost certainly immaterial, but "almost certainly" is not a reason to pick silently — write `/128` to
|
||
match the numbers we have, and settle it with one back-to-back run in §12.
|
||
|
||
Output is 512 floats. Every downstream comparison is a dot product over the **unit** vector, and the
|
||
reference L2-normalises before storing so that no code path has to remember to. The reference clamps
|
||
the norm at 1e-6 before dividing, which costs nothing and removes a NaN path.
|
||
|
||
**Keep the length.** The norm the normalisation divides out is not noise. ArcFace trains the
|
||
direction of its output and nothing else, and the magnitude it leaves behind grows with how much of a
|
||
face the model could make out — MagFace (Meng et al., CVPR 2021) made that the training objective,
|
||
and the plain ArcFace heads this crate runs already show it, weaker but usable. A blur, an occlusion,
|
||
a hard profile or a badly lit crop comes out short. On the reference library `w600k_mbf`'s norms run
|
||
from about 8 on a blur to the high 20s on a clean portrait.
|
||
|
||
So the store holds the **raw** vector, not the unit one — `dr_face::Embedded::to_f16_bytes` — and
|
||
readers re-normalise on load, which they had to do anyway (below). f16 keeps the same three figures
|
||
of a component whatever the vector's length, so this costs nothing in precision. The length is also
|
||
kept beside the blob as `faces.quality`, for the readers that never load the vector (the People
|
||
screen), and it is `NULL` for a face stored as a unit vector before this — a unit vector reads as a
|
||
length of one, and one is not "unmeasured".
|
||
|
||
What the number does is in §9: a face whose quality is under **`MIN_GALLERY_QUALITY` = 14** is still
|
||
placed, but is never what another face is compared *against*. The screen shows it as "Quality 17.3",
|
||
dimmed below the floor, so a user asking why a group did not gather the rest of a person can see that
|
||
none of its members can vouch for anyone.
|
||
|
||
Some exports of these graphs are fp16 in, fp16 out. The reference detects this from the graph's
|
||
declared element type rather than from the filename; worth porting, because the alternative failure is
|
||
a silent garbage tensor.
|
||
|
||
Storage is `512 × f16` (catalog.md §10.1) — 1 KB per face, 25 MB for a 25,000-face library. The f16
|
||
round-trip perturbs a unit vector by ~1e-3 in cosine, three orders of magnitude below the separation
|
||
between a match and a non-match, and the halving matters because these rows are the ones NFR-SEC-5
|
||
contemplates optionally syncing.
|
||
|
||
Re-normalise on load after the f16 widen. It is one pass over 512 floats and it removes a class of
|
||
drift that is otherwise invisible — and, since the blob is raw, it is what turns the stored vector
|
||
back into the unit one every comparison expects.
|
||
|
||
---
|
||
|
||
## 7. Which pixels the pipeline actually sees
|
||
|
||
**This section previously specified the proxy tier — `ThumbSize::Large`, 1024 px — as the buffer
|
||
both detection and cropping read. That was wrong, and §7b is the measurement that says so.**
|
||
FR-CULL-8 now separates the two, because they want opposite things:
|
||
|
||
| Stage | Resolution | Why |
|
||
|---|---|---|
|
||
| Source render | **native** | The only stage where more pixels exist to be had |
|
||
| Detector input | downscaled to ~640 | §4.1 letterboxes to 640×640 regardless; more is wasted CPU |
|
||
| Crop + align | **sampled from the native render** | The 112×112 is fixed, so this decides whether it holds real pixels |
|
||
| Embedding | 112×112 | §6 |
|
||
|
||
The detector's indifference to resolution is the whole reason the split works. §4.1 fixes its input
|
||
at 640×640 and letterboxes whatever arrives, so a face occupying 2% of the frame presents at 12 px
|
||
to the model whether the buffer handed over is 1024 px or 6000 px. Feeding it native pixels buys
|
||
nothing. **Feeding the *crop* native pixels buys everything**, because §5's warp is the one place
|
||
where source resolution converts directly into embedding quality.
|
||
|
||
What the crop receives, by face size, from a native render of a 24 MP frame (~6000 px long edge):
|
||
|
||
| Face size in frame | Pixels across, native | Pixels across, 1024 proxy | What §6 receives |
|
||
|---|---|---|---|
|
||
| A portrait, face fills a third of the frame | ~2000 | ~340 | Downsampled. Ideal either way. |
|
||
| Two people, half-length | ~700 | ~120 | Native: comfortable. Proxy: marginal. |
|
||
| A group of eight | ~290 | ~50 | Native: real pixels. Proxy: **upsampled 2.2×**. |
|
||
| A figure in a landscape | ~120 | ~20 | Native: usable. Proxy: below §4.3's floor. |
|
||
|
||
So `faces` records one column beyond catalog.md §10.1's schema:
|
||
|
||
```sql
|
||
ALTER TABLE faces ADD COLUMN crop_px INTEGER NOT NULL; -- source pixels across the aligned crop
|
||
```
|
||
|
||
`crop_px` earns its place four times over. It is the honest quality signal for the UI; it is a
|
||
**feature in §8's calibration**, which FR-CULL-9 explicitly demands ("a raw cosine means something
|
||
different for every model, every population, and *every face size*"); it is what a
|
||
higher-resolution re-embedding pass selects on, so that pass is a query rather than a re-index of
|
||
everything; and it is the only way to audit whether the rule above is actually being followed —
|
||
which is how §7b found that it was not.
|
||
|
||
---
|
||
|
||
## 7a. Recording that detection *ran*
|
||
|
||
The spec above assumed the `faces` table could answer "has this image been indexed". **It cannot**,
|
||
and the difference is the one that decides whether a background pass ever finishes.
|
||
|
||
A photograph with no face in it produces no rows. So does one that has never been looked at. Asking
|
||
`faces` therefore re-queues every landscape, still life and document scan on every pass, for ever —
|
||
and in a personal library that is most of it. Measured on the reference library: of the first 110
|
||
images indexed, **64 contain no face at all**.
|
||
|
||
So `face_index` records the *run*: one row per `(image, model)` carrying the timestamp, the number of
|
||
faces found — zero is the interesting value — and the long edge of the proxy it read. Keyed on the
|
||
model, so an *embedder* change puts every image back in the queue without anyone having to remember
|
||
to clear anything; a detector change in front of the same embedder does not (§12.3).
|
||
|
||
Three things fall out of it that were not otherwise available:
|
||
|
||
- **A coverage figure.** "4,812 of 5,000 indexed" is what a user wants to see; counting face rows can
|
||
only ever report how many faces exist, which is a different number that never reaches the image
|
||
count.
|
||
- **A reason for the ones outstanding.** The audit splits them by whether a proxy exists, because
|
||
*waiting on the thumbnail sweep* and *waiting on face indexing* are different problems and only one
|
||
of them is fixed by running this again. On the reference library the first check reported 110 ready
|
||
and 23,417 awaiting a proxy — which is the real state of that library, and not something the face
|
||
subsystem can do anything about.
|
||
- **Something to sync.** §14's shards carry the marker with the faces, so an adopted image is not
|
||
re-detected on the receiving device.
|
||
|
||
---
|
||
|
||
## 7b. What the proxy tier actually cost, measured
|
||
|
||
The table in §7 predicted upsampling for small faces and called it "degraded, and usable". On the
|
||
reference library of 23,531 images it was not the edge case that description implies. Reading
|
||
`crop_px` across the 18,671 faces stored under the proxy-tier implementation:
|
||
|
||
| Source pixels across the aligned crop | Faces | Share |
|
||
|---|---|---|
|
||
| <56 (upsampled more than 2×) | 314 | 1.7% |
|
||
| 56–111 (upsampled) | 8,505 | 45.6% |
|
||
| 112–223 (roughly native) | 6,313 | 33.8% |
|
||
| ≥224 (downsampled — ideal) | 3,539 | 19.0% |
|
||
|
||
**47.3% of every face in the library was upsampled to reach the embedder**, with `crop_px` as low
|
||
as 34 — a 3.3× enlargement — against a mean of 178. An upsampled crop does not fail loudly. It
|
||
produces a confident 512-d embedding describing detail that was interpolated rather than
|
||
photographed, and the damage appears three stages later as clusters that will not separate.
|
||
|
||
### Measured again, against the implementation
|
||
|
||
The figures above are read out of a catalog after the fact, so they describe what the old code did
|
||
rather than what the new code does. `examples/face_native.rs` renders one file and indexes it both
|
||
ways, so the difference can be attributed to the resolution and nothing else. Fourteen originals
|
||
from the reference library — 5472×3648 Canon CR2 and DNG — each rendered once and indexed twice:
|
||
|
||
| Path | Faces found | Mean `crop_px` |
|
||
|---|---|---|
|
||
| Native (this specification) | 9 | **287** |
|
||
| Everything from a 1024 proxy | 5 | **75** |
|
||
|
||
Crops 3.8× larger, and on the right side of the line that matters: 75 px is *below* [`ALIGNED_EDGE`]
|
||
so the proxy path was upsampling into the embedder on average, where the native path downsamples
|
||
into it.
|
||
|
||
**Sample of fourteen images and nine faces.** Enough to show the direction and to catch a wrong
|
||
landmark scaling, which is what it was written for; not enough to quote a ratio as the library-wide
|
||
figure. §12's M4 is still where the recall curve gets established.
|
||
|
||
A second effect, recorded here because it was measured and because the mechanism is **not**
|
||
established. Grouping the same runs by `face_index.source_edge` — the buffer detection ran
|
||
against — gives 0.078 faces per image at 1024 or below, against 1.82 at 2048 or better. Controlled
|
||
for file type and size (1,592 DNGs averaging 21.0 MB against 7,724 averaging 23.3 MB, same library,
|
||
same cameras), so it is not a composition artefact. But it cannot be a matter of the detector
|
||
seeing fewer pixels, since §4.1 letterboxes both to 640: a 1024 buffer and a 3072 buffer present
|
||
the same face at the same size to the model. The likeliest explanation is that the 1024 proxy is
|
||
itself a downscale of a larger preview, so the detector sees a twice-resampled image where the
|
||
larger buffer is resampled once — but that is a hypothesis, not a finding, and §12's M4 is where it
|
||
should be settled. **The crop measurement above stands on its own and does not depend on it.**
|
||
|
||
The A/B is evidence for that hypothesis without settling it. Native found nine faces where the
|
||
1024 path found five, on four files where the proxy path found none at all — so detector input
|
||
*does* affect recall, which the letterbox says it should not. The two paths differ in their
|
||
resampling as well as their size (one box filter and one letterbox against a downscale and a
|
||
letterbox), and this experiment does not separate those. Isolating it means holding the chain fixed
|
||
and varying only the buffer, which is M4's job.
|
||
|
||
---
|
||
|
||
## 8. Calibration — cosine to probability
|
||
|
||
FR-CULL-9 makes this a hard requirement: no code path may threshold a bare cosine, every threshold
|
||
in the subsystem is stated as a probability, and the fit is per library and reports its own validity.
|
||
This section is how that is obtained, and the interesting part is where the training pairs come from
|
||
when the user has labelled nothing yet.
|
||
|
||
### 8.1 Where the pairs come from
|
||
|
||
**Negatives are free and abundant.** Two faces detected in *the same photograph* are almost never the
|
||
same person. That gives every multi-face image in the library a full set of negative pairs at no
|
||
labelling cost — and they are *hard* negatives, drawn from the same camera, lighting, and processing,
|
||
which is exactly the population where a threshold tuned on easy negatives fails. The known exceptions
|
||
— mirrors, photographs of photographs, collages, a spliced panorama — are rare enough to be noise at
|
||
this scale and worth naming so the exception is not mistaken for a bug later.
|
||
|
||
**Positives, in order of trustworthiness:**
|
||
|
||
1. **User confirmations** (FR-CULL-10). Every pair of faces confirmed to the same person. The gold
|
||
standard, and empty on day one.
|
||
2. **Burst siblings.** FR-CULL-5 already groups bursts. Two faces in adjacent frames of one burst,
|
||
in nearly the same position, are near-certainly the same person. Free, requires no labelling, and
|
||
available immediately on any library with continuous-shooting frames in it. **Their purity is an
|
||
S14 measurement** (§12), not an assumption — if bursts turn out to be dirtier than expected, this
|
||
source is dropped and the calibration simply stays invalid for longer.
|
||
3. Nothing else. Bootstrapping positives from high cosine similarity is circular — it fits the
|
||
calibration to the belief it was supposed to test.
|
||
|
||
### 8.2 The fit
|
||
|
||
§1.1's reference already implements this and its shape should be ported rather than reinvented:
|
||
|
||
```
|
||
P(same | cos) = σ(a·cos + b + log_prior_odds)
|
||
```
|
||
|
||
Four details in it are the difference between working and nearly working.
|
||
|
||
**Fit against a histogram, not against pairs.** A 25,000-face library has 3×10⁸ pairs; no gradient
|
||
descent is running over that. The reference buckets every pair into **200 bins over cos ∈ [−1, 1]**,
|
||
carrying a positive and a negative count per bin, and fits the two parameters against the per-bin
|
||
counts. The cost becomes the similarity matrix plus 200 numbers, and the matrix is a single GEMM
|
||
(§9). This is the trick that makes a per-library fit affordable at all, and it is not obvious from
|
||
the outside.
|
||
|
||
**Balance the classes explicitly.** `w_pos = total/(2·n_pos)`, `w_neg = total/(2·n_neg)`. Negatives
|
||
outnumber positives by orders of magnitude; an unweighted fit produces a well-shaped curve sitting at
|
||
the wrong height — precisely the "plausible number all the way to the user interface" failure
|
||
FR-CULL-9 describes.
|
||
|
||
**The base rate is a runtime argument, not part of the fit.** `log_prior_odds` is added at evaluation
|
||
time, so the balanced fit is stored once and the prior varies per query — the odds that two faces
|
||
drawn from a 40-image holiday album match are not the odds for a 40,000-image archive. Its inverse,
|
||
`boundary_at(p)`, gives the cosine at which the probability crosses `p`, which is what turns §9's
|
||
"merge above 0.9" into an actual comparison. Baking a prior into `a` and `b` would need a refit per
|
||
context and would make the stored parameters mean something different depending on where they came
|
||
from.
|
||
|
||
**Deduplicate before pairing.** Near-identical embeddings (cos > 1 − 1e-7) are the same photograph
|
||
counted twice; the reference drops them per identity first. In a photo library the equivalent is
|
||
duplicates and virtual copies (FR-CAT-11), and leaving them in stacks the positive histogram at
|
||
cos ≈ 1 with pairs that teach the fit nothing about hard cases.
|
||
|
||
**A third feature this design adds.** The reference fits on cosine alone; DarkRoom should carry
|
||
`crop_px` too:
|
||
|
||
```
|
||
logit P(same) = w₀ + w₁·cos + w₂·log₂(min(crop_px_a, crop_px_b))
|
||
```
|
||
|
||
FR-CULL-9 explicitly names face size as an axis along which an uncalibrated similarity misbehaves,
|
||
§7 shows a real library spans 40 px to 340 px of face, and `crop_px` is already in hand. The
|
||
*minimum* of the pair, because a comparison is only as good as its worse crop. It generalises the
|
||
histogram to a small 2-D grid of bins, which changes nothing structural. If M7 says the extra feature
|
||
buys nothing, drop it and match the reference exactly — but a film pipeline working from broadcast
|
||
frames had far less size variation to explain than this one does.
|
||
|
||
### 8.3 Validity, and saying so
|
||
|
||
Stored per catalog.md §10.3: parameters, a validity flag, and a hash of the face set fitted from.
|
||
|
||
Invalid when fewer than **200 positive pairs** or **2,000 negative pairs** are available, or when the
|
||
reliability check fails.
|
||
|
||
Deliberately far stricter than the reference's floor of two positives and one negative. That floor is
|
||
reasonable there: its pairs come from a curated gallery of labelled reference portraits, where a
|
||
positive pair is trustworthy by construction. Here the positives are bootstrapped from bursts and
|
||
early confirmations (§8.1) and the whole risk is fitting confidently to a handful of them. The
|
||
reference also refuses to draw positives from an identity with fewer than five distinct embeddings,
|
||
letting it contribute negatives only — the same asymmetry applies to a thinly-confirmed person and is
|
||
worth keeping. In that state both clustering and the displayed confidences fall back to the
|
||
reference implementation's fitted curve — a documented operating point, not an invention — and the
|
||
People screen says so once, above the grid, rather than blanking every percentage. That is
|
||
FR-CULL-9's requirement read as written: what may not happen is an untuned default presented *as
|
||
though it were measured* on this library.
|
||
|
||
Blanking them was the first reading, and it was wrong in a way worth recording. A young library has
|
||
no fit; a fit needs confirmations; confirmations are made on a screen the user ranks by confidence.
|
||
Withholding the confidence until the fit exists is a deadlock in which the normal state of the
|
||
feature is its degraded one.
|
||
|
||
Refit is triggered by the same debounce as clustering (§9), and when the confirmed-pair count grows
|
||
materially.
|
||
|
||
*Acceptance (FR-CULL-9):* a reliability diagram over ten probability bins, on a held-out labelled
|
||
split, with observed match rate within a stated tolerance of the predicted probability in each
|
||
populated bin. A single accuracy figure is not an answer to this requirement.
|
||
|
||
---
|
||
|
||
## 9. Clustering
|
||
|
||
Not a job kind, for the reason catalog.md §10.2 gives: it is a whole-library operation with no
|
||
natural `subject_id`. It runs as a debounced library pass when detection has been idle and the face
|
||
count has moved materially.
|
||
|
||
**The graph.** For each face, its `k = 20` nearest neighbours by cosine, then each candidate edge
|
||
scored through §8 to a probability.
|
||
|
||
Brute force is honest arithmetic here, and §1.1's project is the evidence rather than the estimate: it
|
||
computes the full similarity matrix as **one GEMM** (`E · Eᵀ` over L2-normalised rows) at gallery
|
||
scale, on CPU, and instruments it well enough to have compared CPU against OpenCL and found the
|
||
question worth asking rather than urgent. For DarkRoom that is a 25,000 × 512 matrix at 25 MB in and
|
||
~2.5 GB out — so the matrix is computed **in row blocks**, with each block reduced to its top-k and
|
||
its histogram contribution before the next is started, and never materialised whole. Once, in the
|
||
background, after an indexing sweep that took an hour. An approximate index is an optimisation to
|
||
reach for when §12 says it is needed, not before.
|
||
|
||
**The kernel is most of the cost, and it was running at a tenth of the machine.** The scan is
|
||
`O(n²)` dot products and nothing else, so its speed *is* the subsystem's speed. Two things were
|
||
wrong with the first one, both measured over a real 18,143-face library on a twenty-core desktop:
|
||
|
||
| | scan | GFLOP/s |
|
||
|---|---|---|
|
||
| a row against every other row, `&[Vec<f32>]` | 4.64 s | 36 |
|
||
| tiled on the column side too, one flat buffer | 2.81 s | 60 |
|
||
| **plus AVX2 + FMA** | **0.86 s** | **195** |
|
||
|
||
The first is memory: walking the whole embedding array once per row moves ~336 GB for that library,
|
||
where a column tile that fits in L2 is read once per *tile of rows*. The second is that the workspace
|
||
builds for baseline `x86-64` — SSE2, no FMA — and the portable loop was not being vectorised into
|
||
even that, at 0.7 flops per cycle.
|
||
|
||
So the dot product is chosen per machine: AVX2 + FMA where `is_x86_feature_detected!` finds it,
|
||
**NEON unconditionally on aarch64** — Advanced SIMD is in that baseline, so every Android device the
|
||
app builds for has it, and the explicit `vfmaq` matters because LLVM will not fuse a multiply and an
|
||
add on its own. The portable loop remains the definition the others are tested against. All three
|
||
produce the same 1,531,969 pairs.
|
||
|
||
The NEON kernel is the one a desktop `cargo test` never executes, so
|
||
`tools/face-tests-on-device.sh` runs the suite on an attached device: `dr-face` carries no weights
|
||
and touches no display, so its tests are a plain ARM64 binary that runs under `adb shell` with
|
||
nothing installed. Worth running whenever the kernels change.
|
||
|
||
**Where a regroup's time actually goes**, on that library, because the answer moved twice while it
|
||
was being looked at. Measured with `cargo run --release -p dr-catalog --example face_confidence --
|
||
CATALOG --full`, on the reference desktop and on a Honor tablet (ROD2-W09, aarch64):
|
||
|
||
| | desktop, before | desktop | **tablet** |
|
||
|---|---|---|---|
|
||
| scan | 4.60 s | 0.96 s | **2.61 s** |
|
||
| agglomerate | 4.84 s | 1.69 s | **2.16 s** |
|
||
| score | 0.23 s | 0.26 s | **0.40 s** |
|
||
| **total** | **10.0 s** | **3.1 s** | **5.2 s** |
|
||
|
||
The scan came down by the kernel work above. The agglomeration was not, as it looked, an irreducible
|
||
sequential heap walk: 3.06 s of it was every component scanning the *whole* pair list for the pairs
|
||
that were its own — 382 million set lookups to place 804,499 pairs — which the union-find can do in
|
||
one pass while it is finding the components anyway. Another 0.54 s was `Engine::cross` computing dot
|
||
products with the portable loop while the scan beside it used the machine's SIMD.
|
||
|
||
**The two architectures agree exactly**, which is worth more than either column: the same 1,531,969
|
||
evidence pairs, the same 2,518 groups holding the same 16,246 faces, and the same reliability table,
|
||
from AVX2 on the desktop and NEON on the tablet. That is the cross-kernel check the unit test can
|
||
only approximate.
|
||
|
||
**The tablet is where a GPU GEMM would pay.** Its scan is *half* the pass, against under a third on
|
||
the desktop — twenty cores of AVX2 pull ahead of a tablet's NEON far more than the merge engine's
|
||
single-threaded hashing does. So a perfect GEMM is worth about 2× a regroup there and about 1.5×
|
||
here, and it is the phone and tablet story that should decide whether it gets built.
|
||
|
||
**Constraints, not just thresholds:**
|
||
|
||
- **Cannot-link on co-occurrence.** Two faces in the same image are never merged. This is the same
|
||
observation §8.1 mines for negatives, used here as a hard constraint, and it is the single
|
||
cheapest defence against the over-merging FR-CULL-10 warns about.
|
||
- **Confirmed faces are anchors.** A confirmation is user data (FR-CULL-12) and clustering never
|
||
moves it. Two clusters each containing confirmations of *different* people cannot merge; a cluster
|
||
containing confirmations of one person absorbs suggestions but never reassigns the confirmed.
|
||
- **A short embedding is never a reference.** The length of the raw vector is the model's own
|
||
reading of the crop (§6), and a short one sits near the middle of the sphere, matching a little of
|
||
everybody — one of those in a group is a bridge to the next group over. So the population is
|
||
split: faces at or above `MIN_GALLERY_QUALITY` are the **gallery** and cluster as described below;
|
||
faces under it are **probes**, each measured against the finished groups and placed in the one it
|
||
fits by the same average-link rule under the same two constraints — but measured against gallery
|
||
members only, never against another probe, and once placed never part of what the next face is
|
||
measured against. Two probes are never paired at all, and `neighbours` drops those pairs before
|
||
anything downstream sees them. A probe's confidence (§9.1) is computed from the references it
|
||
matched; a reference's confidence hears nothing from a probe. A face whose quality was never
|
||
recorded is admitted to the gallery — a rule that cannot be checked admits rather than excludes —
|
||
and the next indexing pass **measures** it: the `face-quality` repair (§18.1) lists every image
|
||
holding one, and each such face is embedded again from the native render with the landmarks it
|
||
already has, the raw vector written over the old one and its id, box and identity untouched
|
||
(`faces::record_updates`). No detector runs and no suggestion is lost — the cost is the
|
||
original fetched once more, since the length exists only at the moment of embedding.
|
||
- **A person is stood for by their references.** Every face the user has ruled on is an anchor,
|
||
and the scan is exhaustive, so a person with 750 confirmations would cost 750 comparisons
|
||
against every other face — and the cost of a library would grow with how well it was named.
|
||
Instead each person enters through at most `MAX_REFERENCES` (100) of their anchored faces,
|
||
chosen by `dr_face::references`: those whose raw embedding is at least `MIN_REFERENCE_QUALITY`
|
||
(15) long, and among them the set spanning the greatest volume — greedy max-determinant, the
|
||
longest vector first and then, at each step, the face with the largest component orthogonal to
|
||
the chosen so far. Thirty frames from one afternoon contribute one reference; the single profile
|
||
shot is taken early. The faces not chosen keep their confirmations and are not touched by the
|
||
pass; they are simply not compared.
|
||
|
||
**The algorithm.** Constrained average-link agglomeration over the probability graph, merging while
|
||
the average pairwise probability exceeds **0.9** and no cannot-link is violated. Average-link rather
|
||
than single-link because single-link chains — one bad edge welds two identities together, which is
|
||
the documented way face clustering fails on families.
|
||
|
||
**Incremental by default.** A new face joins the existing cluster whose average probability against
|
||
it is highest, if that clears the threshold; otherwise it starts an unnamed group. A full
|
||
re-agglomeration happens on model change, on recalibration, and on explicit user request. Suggestions
|
||
are recomputed freely; confirmations survive all of it (FR-CULL-10).
|
||
|
||
**Splitting** re-agglomerates within one person at a raised threshold and offers the resulting groups
|
||
as the split. FR-CULL-10 requires splitting to be as easy as merging, and a split that hands the user
|
||
a pile of loose faces to re-sort is not that.
|
||
|
||
### 9.1 The number beside a suggestion
|
||
|
||
Which person a face belongs to and how sure that is are **different questions**, and the second one
|
||
is not answered by the pairwise probabilities that settled the first.
|
||
|
||
The first implementation answered it with the mean calibrated probability between the face and the
|
||
rest of its group, and that measures the wrong thing twice. It punishes coverage: a person with two
|
||
hundred faces across fifteen years is *supposed* to have members a new photograph is orthogonal to,
|
||
so the better someone is photographed the worse their suggestions score. And it never asks who else
|
||
the face could be — a face matching Anna at 0.95 and nobody else, and one matching Anna at 0.95 and
|
||
her sister at 0.93, come out identical, when the second is the only one the user needs to look at.
|
||
|
||
Two questions, so two factors, multiplied:
|
||
|
||
```
|
||
evidence(P) = Σ of the top n of { P(same | this face, f) : f ∈ P } n = 10
|
||
coherence = evidence(own) / how many of the top n there were
|
||
uniqueness = evidence(own) / (evidence(own) + Σ evidence(named rivals))
|
||
confidence = coherence × uniqueness
|
||
```
|
||
|
||
**Coherence** is the old mean with a cap on it, and the cap is the whole fix: the two hundred faces a
|
||
given photograph is legitimately orthogonal to stop counting against it. **Uniqueness** is the
|
||
competition, and it is what makes an ambiguous face read as ambiguous — two identities matching
|
||
equally well land at 0.5 each, which is the truth about a sibling.
|
||
|
||
**Only named people compete, and they compete per person.** This is the part that had to be measured
|
||
rather than reasoned about. Normalising across *every* group made the number useless on a real
|
||
18,000-face library — median suggestion 21%, four in five under half — because clustering leaves one
|
||
person spread across many groups, so a face competes against itself. Counting only the groups the user has
|
||
ruled on — one holding a confirmation, a name, or an ignore — fixed most of it; counting them **per
|
||
person** rather than per group fixed the rest, since one person is left in several anchored groups
|
||
for the same reason.
|
||
|
||
**Rivals are gathered below the merge threshold**, down to even odds: a named person who matches at
|
||
0.6 will never be merged into but is exactly the competition a suggestion should be discounted for.
|
||
The floor matters in both directions — summing the near-orthogonal pairs instead of dropping them
|
||
lets fifty identities' worth of upper-tail noise outweigh one real match, which on the same library
|
||
moved the median stated confidence from 100% to 31%.
|
||
|
||
*Measured (`cargo run --release -p dr-catalog --example face_confidence`), leave-one-out over that
|
||
library's 2,702 confirmations across 54 named people:*
|
||
|
||
| | share | old mean |
|
||
|---|---|---|
|
||
| right person picked | **99.33%** | 99.15% |
|
||
| stated for the right person, median | **99.3%** | 90.4% |
|
||
| stated for the right person, p10 | **79.2%** | 68.0% |
|
||
|
||
The reliability table is monotone and **errs low**: 100% correct wherever it states 80% or more, 84%
|
||
correct where it states under half. Understating is the safe direction for a screen whose purpose is
|
||
deciding what to look at first, but the low bands are not calibrated and should not be read as
|
||
though they were — and the leave-one-out task asks *which of these people*, never *is it any of
|
||
them*, so it cannot speak to a stranger at all.
|
||
|
||
It is **not** a merge threshold and must not become one. Uniqueness is relative, so a library with
|
||
one named person would hand every stray face a 1. "Is this the same person at all" stays §8's
|
||
question, and coherence is the half of the product that carries it.
|
||
|
||
### 9.2 The two numbers the user is allowed to move · 2026-08-29
|
||
|
||
The merge probability and the smallest group the pass will call a person are `FaceSettings` in
|
||
`dr-types`, edited from the People screen and saved per device beside the cache budgets. They were
|
||
constants: `dr_face::DEFAULT_MERGE_PROBABILITY` and a bare `< 2` in `dr_ui::faces::recluster`.
|
||
|
||
**Why they had to become settings.** The default was tuned on one library — the table in
|
||
`dr_face::cluster`'s doc comment is 1,813 faces of one photographer's family — and the quantity it
|
||
optimises is a property of the population, not of the model. A library of one household at close
|
||
family resemblance and a library of two thousand strangers at a wedding want different answers, and
|
||
neither of them is the reference library. The doc comment already conceded the point ("this is a
|
||
*default*, not a constant of nature") and pointed at `face_index --tune` as the way to find a better
|
||
one; a photographer does not have a terminal.
|
||
|
||
**Why moving them is safe, and why that is the reason there is no confirmation on it.** A regroup
|
||
writes only the *suggested* half. Confirmations, names and ignores enter as anchors and come back
|
||
unchanged (FR-CULL-10), so the pass is re-runnable by construction and a dial the user can move is
|
||
just that property being used. The smallest-group rule is applied only to groups the system invented:
|
||
a group carrying a person — confirmed, named or set aside — survives it whatever its size, because a
|
||
display preference does not overrule a judgement (FR-CULL-12).
|
||
|
||
**Withdrawal, which the setting does not work without.** Raising the smallest group stops the pass
|
||
*creating* small groups; it does not by itself remove the ones a previous pass made, because those
|
||
still hold their suggestions, so they are not empty, so `prune_empty_unnamed` leaves them. The pass
|
||
therefore now releases every unanchored face it did not place — `faces::unassign` — before pruning.
|
||
Without that step the control appears to do nothing until the library is reindexed.
|
||
|
||
**The preview.** `dr_ui::faces::preview_grouping` runs the same population through
|
||
`dr_face::cluster` and reports groups, faces grouped and largest group without opening a
|
||
transaction. It is `face_index --tune`'s row for one setting, on the user's own library, on a worker
|
||
thread. The line leads with the **group count** because that is the number that says which side of
|
||
the right setting you are on: it climbs as fragments are gathered into people and falls as separate
|
||
people start being welded together, while the grouped-face count rises straight through both.
|
||
|
||
---
|
||
|
||
## 10. Catalog and jobs
|
||
|
||
**Schema** is catalog.md §10.1's v5 migration, plus `faces.crop_px` (§7) and the calibration table
|
||
(§8.3). Nothing else changes.
|
||
|
||
**One new job kind:**
|
||
|
||
```rust
|
||
/// Detect and embed faces in one image, from its proxy (FR-CULL-8).
|
||
DetectFaces = 8,
|
||
```
|
||
|
||
Background priority, coalesced per `image_id`, interruptible, resumable — it inherits FR-CAT-3's
|
||
properties with no exceptions, which is what FR-CULL-8's acceptance criterion is about. It is not a
|
||
network job.
|
||
|
||
One job does detection *and* embedding for every face in the image, rather than splitting them.
|
||
Splitting would double the queue's row count for no benefit: the proxy is already decoded and in
|
||
memory, and the natural unit of resumable work is one photograph.
|
||
|
||
**Selector term** (FR-CULL-11):
|
||
|
||
```rust
|
||
Person { id: PersonId, include_suggested: bool }, // defaults to false
|
||
```
|
||
|
||
Confirmed-only by default, so a saved smart collection does not silently change membership when a
|
||
later indexing pass revises a guess.
|
||
|
||
---
|
||
|
||
## 11. Privacy obligations that constrain the code's shape
|
||
|
||
NFR-SEC-5 is not a policy to remember; it is a set of properties the code either has structurally or
|
||
does not have at all. The checkable ones:
|
||
|
||
- **`dr-face` has no network dependency.** No `reqwest`, no `ureq`, transitively. Worth a CI check
|
||
over `cargo tree`, alongside the existing lints — the crate that holds the embeddings should be
|
||
provably unable to send them anywhere.
|
||
- **The model fetch (§2.2) is in the app layer**, which is why the API in §3.1 takes bytes.
|
||
- **The diagnostics bundle is an allowlist** (NFR-OPS-1), so `faces`, `face_person` and the
|
||
calibration table are excluded by not being named, and a future table cannot become uploadable by
|
||
existing.
|
||
- **Delete-all is one transaction and one control**: faces, links, people, calibration, and the
|
||
cached crops if any. Separately, a switch that stops indexing so no such data is produced. The two
|
||
are different actions and are not collapsed into one.
|
||
- **The about screen reads the model metadata from the loader** — id, version, licence — rather than
|
||
from a hardcoded string that will drift from what is actually running.
|
||
|
||
---
|
||
|
||
## 12. What S14 measures
|
||
|
||
The corpus is a real personal library — the ~2,000-image sample S14 already specifies — with a
|
||
hand-labelled identity subset. Fixed in advance, so the result is a measurement rather than an
|
||
argument:
|
||
|
||
| # | Measure | Why it decides something |
|
||
|---|---|---|
|
||
| **M1** | **Does tract load both graphs?** As shipped, then with the input dims frozen | Go/no-go, and it is *first*. `det_500m.onnx` has a **dynamic H/W input** (§4.1), which is the exact thing tract failed on for YOLO and which OpenCV's `dnn` could not load either — so the as-shipped answer is expected to be no, and the real question is whether `make_dynamic_shape_fixed` is enough or whether a PyTorch re-export is needed. Costs an afternoon; everything below is void without it. |
|
||
| **M2** | Detection latency per image at 640, and at 320, on the reference desktop and one Android device | Whether a 17k library indexes in a background sweep or a weekend. The one measured datapoint we have — YOLO26n-seg, ~470 ms at 640×640 in tract — suggests SCRFD-500M lands well under it, but tract is not ORT and an extrapolation is not a measurement. |
|
||
| **M3** | Embedding latency per face, and faces per image in a real library | The multiplier on M2. A library averaging 1.5 faces per image at 100 ms per face is an hour for 25k faces; at 400 ms it is four. |
|
||
| **M4** | Detection recall against hand-labelled faces, bucketed by `crop_px` | Where §4.3's floor should actually sit, and what proportion of a real library's faces are in the degraded bucket §7 predicts. |
|
||
| **M5** | **Aligned versus unaligned embeddings**, same corpus, same clustering. Then two A/Bs that cost one run each: least-squares versus RANSAC alignment (§5), and `/128` versus `/127.5` normalisation (§6) | Quantifies §5. If the gap is small the alignment code is still correct; if it is large, this is the measurement that stops someone "simplifying" it later. The two A/Bs close divergences between the reference and the published Python that are currently settled by assertion. |
|
||
| **M6** | Cluster purity and completeness against the labelled subset | Whether the subsystem is worth building at all. §1's F1 figures rank the models on *film frames against a known cast*; nothing yet says how MBF behaves on family snapshots it must cluster blind. |
|
||
| **M7** | Calibration reliability, and **the sample size at which the fit becomes valid** | FR-CULL-9's acceptance criterion, plus the practical question of whether a 2,000-image library ever gets a valid fit or whether §8.3's "unavailable" is the normal state. |
|
||
| **M8** | Burst-derived positive-pair purity (§8.1) | Whether the free positives are usable or the fit waits for user confirmations. |
|
||
| **M9** | YuNet as a drop-in detector: M4 and M6, re-run | §2.3's question 3, and the model file is already on disk. If the answer is "close enough", half the licence problem disappears. Needs its own decode path — its outputs are not SCRFD's. |
|
||
| **M10** | Peak RSS during an indexing sweep | NFR-RES-2. Two loaded graphs plus a proxy plus a batch of crops, on a phone. |
|
||
|
||
### 12.1 M1 result — **PASS, conditionally** · 2026-08-26
|
||
|
||
Measured, not extrapolated. Both InsightFace graphs **fail to load in tract as shipped**, exactly as
|
||
§4.1 predicted and for the reason it gave:
|
||
|
||
```
|
||
scrfd_500m_bnkps.onnx Translating node #0 "input.1" Source ToTypedTranslator
|
||
arcface_w600k_mbf.onnx Failed analyse for node #139 "Conv_0" ConvHir
|
||
```
|
||
|
||
Both **load cleanly once their input dimensions are pinned** — SCRFD's unnamed H/W to 640, ArcFace's
|
||
`None` batch to 1 — by `tools/fix-face-model-shapes.sh`, which rewrites the declared dims and touches
|
||
no weights. The frozen SCRFD reports the layout §4.2 specifies, which is the second half of the
|
||
answer: nine outputs, three strides, last dims 1/4/10, and `12800 = 80 × 80 × 2` confirming two
|
||
anchors per location at stride 8.
|
||
|
||
Two things worth carrying forward:
|
||
|
||
**SCRFD's outputs were already static.** The export was made at 640 and only its input forgot to say
|
||
so, so pinning to 640 is not a choice this project is making — it is the shape the graph was always
|
||
going to run at. §12's "320 as a faster option" would need a different export, not a different flag.
|
||
|
||
**YuNet loads with no intervention at all**, at a fixed `[1, 3, 640, 640]`, with twelve outputs in
|
||
three strides — `cls`/`obj`/`bbox`/`kps`, which is a *different layout* from SCRFD's and confirms why
|
||
§4.2's load-time check has to look at shapes rather than count outputs. Combined with its permissive
|
||
licence (§2.3) that makes M9 more interesting than it looked: the permissive detector is also the one
|
||
with no shape-fixing step in front of it.
|
||
|
||
|
||
### 12.2 First end-to-end run · 2026-08-26
|
||
|
||
The Rust port produces the separation it is supposed to. Three distinct portraits of one identity
|
||
against two of another, from §1.1's labelled gallery:
|
||
|
||
| Pair | Cosine |
|
||
|---|---|
|
||
| Same identity, different photographs | **0.596** |
|
||
| Same identity, byte-identical duplicate files | 1.000 |
|
||
| Different identities | **0.049 – 0.050** |
|
||
|
||
Against the reference's fitted MBF boundary of cos 0.267 (§1), 0.596 and 0.05 fall either side with
|
||
room to spare — which is the check that the port's pre-processing, letterbox inversion and alignment
|
||
are right, since any of them being wrong degrades the same-identity number first.
|
||
|
||
The duplicate row is not a curiosity: several files in that gallery are byte-identical under
|
||
different names, which is exactly the case §8.2's dedup step exists for, and it would otherwise stack
|
||
the positive histogram at cos ≈ 1 with pairs that teach the fit nothing.
|
||
|
||
**M2/M3 in a debug build:** detection ~1.0–1.4 s per image, embedding ~160–280 ms per face.
|
||
|
||
**M2/M3 in release, over a real library — the number that counts: 3.5 images/second**, end to end,
|
||
including the JPEG decode and the catalog write. 110 images with 125 faces in 30 seconds on the
|
||
reference desktop. That is roughly **4× the debug figure**, and it moves a 23.5k-image library from
|
||
the "seven hours" the debug numbers implied to about **110 minutes**.
|
||
|
||
Worth stating plainly because the debug measurement was nearly a wrong conclusion: it was on the
|
||
edge of making the pure-Rust runtime look unaffordable for a large library, and it was measuring the
|
||
profile rather than the pipeline. Any future timing of this subsystem is a release timing.
|
||
|
||
Still unmeasured: the same pass on a phone (NFR-RES-2), which does not follow from this one.
|
||
|
||
### 12.3 SCRFD-2.5G and 10G against 500M · 2026-09-11
|
||
|
||
§1 chose `500M` on cost; the recall it gives up was never measured. `examples/face_detectors`
|
||
in `dr-ui` runs several detectors over the same 400 proxies, spaced evenly through the reference
|
||
library, matches boxes at IoU ≥ 0.5 against the first, and writes contact sheets of the
|
||
disagreements — because a count of extra faces says nothing until someone has looked at whether
|
||
they are faces. Both size gates off, confidence at the production 0.5, release build, reference
|
||
desktop. Sizes are the box's shorter edge in proxy pixels; the network sees 0.625 of that.
|
||
|
||
| Detector | File | Mean ms | Faces | <16 | 16–32 | 32–64 | ≥64 |
|
||
|---|---|---|---|---|---|---|---|
|
||
| `scrfd_500m` | 2.5 MB | 157 | 1,490 | 153 | 730 | 399 | 208 |
|
||
| `scrfd_2.5g` | 3.3 MB | 177 | 1,698 | 268 | 801 | 418 | 211 |
|
||
| `scrfd_10g` | 17 MB | 488 | 1,896 | 357 | 886 | 438 | 215 |
|
||
|
||
| Against `500M` | Both | Candidate only | Baseline only |
|
||
|---|---|---|---|
|
||
| `2.5G` | 1,404 | **294** (242 of them under 32 px) | 86 |
|
||
| `10G` | 1,419 | **477** (411 under 32 px) | 71 |
|
||
|
||
Two things the sheets settled that the counts cannot:
|
||
|
||
- **The candidates' extras are faces.** The hundred smallest `2.5G`-only tiles are people —
|
||
soft, small, a few motion-blurred, and almost none of them anything else. These are the group
|
||
shots and the figures in the background that §7's table predicted `500M` would lose at 640.
|
||
- **`500M`'s "extras" are mostly not.** Of the 86 faces only `500M` found, the sheet shows the
|
||
same dog a dozen times, a stop sign, a wheel, two hands, the backs of several heads and one face
|
||
upside down. The larger models did not miss these; they declined them. So `500M` is paying
|
||
twice — for the faces it cannot find *and* for the non-faces it embeds, which is exactly the
|
||
garbage-embedding-bridges-two-clusters failure `DetectOptions::confidence` is set against.
|
||
|
||
**`2.5G` is the right detector.** 12% more time for 14% more faces and a cleaner set, in a file
|
||
0.8 MB larger; `10G` finds a further 12% for 3.1× the time, which is a desktop-only price and this
|
||
is not a desktop-only feature (NFR-RES-2). It loads in tract with the same fix as `500M`
|
||
(`--input input.1=1,3,640,640`; its outputs are declared dynamic and tract infers them) and
|
||
decodes through the same nine-output path unchanged. Same licence, same `buffalo_m` release page.
|
||
|
||
**Which detector runs is a setting** — `FaceSettings::detector`, per device, on the settings page
|
||
beside the indexing button as Fast / Balanced / Thorough. All three files ship. Each detector
|
||
writes its own `faces.model_id` (`w600k_mbf` for `500M`, unchanged; `scrfd_2.5g+w600k_mbf` and
|
||
`scrfd_10g+w600k_mbf` for the others), so which pipeline drew a box is always on record. The
|
||
default stays `500M` so that an upgrade changes nothing until the user chooses; the recommendation
|
||
is `2.5G`.
|
||
|
||
**One population per embedder, not one per detector · 2026-09-19.** The first cut of the setting
|
||
keyed every reader on the full id — the clustering pass, the coverage figure, the sweep's work
|
||
list, the shard export and import, and the sync merge's face matching — on the theory that a
|
||
detector change is a model change. Measured on the reference library it was a disaster: choosing
|
||
Thorough on both devices restarted coverage at 1,834 of 19,140, the People screen showed only the
|
||
faces the new pipeline had reached, the desktop's 3,583 confirmations under the old id could not
|
||
reach the tablet because the merge demanded the same id on both sides, and each device faced a
|
||
~400 GB re-fetch before the library looked whole again. The embedder is `w600k_mbf` in every
|
||
variant; its vectors are one space, and the detector only decides where the boxes are.
|
||
|
||
So every reader now keys on the embedder half of the id (`faces::embedder_of`, and
|
||
`embedder_sql` for the queries): all three detectors are one population, and changing between
|
||
them empties nothing. `record_detections` is unchanged — an image holds one pipeline's faces at a
|
||
time, and a re-detection carries identities across by box overlap and embedding (§18) — and it is
|
||
where the generations meet. The merge's `match_faces` matches within an embedder for the same reason. The
|
||
shards travel every generation, each under its own id, and a peer adopts whichever it is sent.
|
||
What a stronger choice still does is queue the images a weaker detector indexed for re-detection
|
||
(`FaceDetector::supersedes`), after the ones nothing has indexed and never downwards, so a tablet
|
||
on Fast keeps the desktop's Thorough faces rather than replacing them with fewer. The calibration
|
||
(§8) is keyed on the embedder too: it is a fit over the similarity space, and that space did not
|
||
change.
|
||
|
||
---
|
||
|
||
## 13. Order
|
||
|
||
1. **M1** — load both graphs in tract, freezing the input dims if needed (§4.1). Go/no-go, and it
|
||
now comes first: it is an afternoon, and §2.3's reading is wasted effort if the graphs will not
|
||
run at all. This is a change from the S14 brief's ordering, made because §1.1 removed the
|
||
uncertainty that justified reading first — the models are known to work, so the open question is
|
||
the runtime, not the choice.
|
||
2. **Licence reading** (§2.3), in parallel with 3. Before anything *ships*, per D13.
|
||
3. `dr-face` skeleton, `detect` + `align` + `embed`, ported from §1.1's table, with an example binary
|
||
that draws boxes and landmarks on a JPEG — the same shape as `dr-segment`'s `examples/detect.rs`,
|
||
and for the same reason: the thing worth looking at is whether the landmarks land on a real
|
||
photograph. Port `tests/test_face_utils.cpp`'s cases first; they are model-free and they fail
|
||
loudly on exactly the mistakes §5 describes.
|
||
4. M2–M5 on the corpus. Alignment is validated here, before anything depends on embedding quality.
|
||
5. `calibrate` + `cluster` against the labelled subset. M6–M8. The Platt fit ports from
|
||
`gallery_calibration.hpp`; the pair *sourcing* (§8.1) is new and is the part to get wrong.
|
||
6. The v5 migration, the `DetectFaces` job, the debounced clustering pass.
|
||
7. UI: People view, confirm and reject, merge and split, the `Person` selector term.
|
||
8. The route-C first-run flow, the about-screen entry, and the delete-all control — with the feature
|
||
shipped disabled until they exist.
|
||
|
||
Steps 1–5 are the spike. Steps 6–8 are the build, and they are only justified if §12 says so.
|
||
|
||
### 13.1 Where this had got to · 2026-08-26
|
||
|
||
- **1 — done.** M1 passes conditionally; §12.1.
|
||
- **2 — outstanding.** §2.3's three questions are unanswered and gate *shipping*, not building.
|
||
- **3 — done.** `core/dr-face`: `align`, `detect`, `embed`, `embedding`, and `examples/faces.rs`.
|
||
§12.2 is its first end-to-end run.
|
||
- **4 — partly.** M2/M3 have provisional numbers; M4/M5 need the labelled corpus.
|
||
- **5 — built, not yet measured.** `calibrate` and `cluster` exist with 34 model-free tests behind
|
||
them; M6–M8 are a run over a real library, which is what the corpus in §1.1's `images/` is for.
|
||
- **6 — done for storage.** Catalog schema v8 — `people`, `faces`, `face_person`,
|
||
`face_person_rejected`, `face_calibration` — with the identity operations FR-CULL-10 requires.
|
||
The `DetectFaces` job kind and the debounced clustering pass are not wired yet.
|
||
- **7 — done.** The Identity screen: a third top-level mode beside library and develop, reached from
|
||
the library header. People rail, face grid with per-face confirm/reject, rename in place, confirm
|
||
all, split off a multi-selection, regroup, and the NFR-SEC-5 delete-everything control.
|
||
- **8 — not started.** The route-C first-run flow: no model fetch, no checksum pin, no licence
|
||
notice. The screen says "No face model installed" and stops, which is honest but is not the
|
||
feature. The models go in `<catalog dir>/models/` as the shape-fixed exports, by hand for now.
|
||
- **9 — done, and not previously in this plan.** The run marker (§10a), the coverage audit, the
|
||
`face_index` batch job, and cross-device sync of face shards (§14).
|
||
|
||
The screen taught the design one thing worth recording. **Splitting has to reject before it
|
||
confirms.** Moving faces to a new person is not enough on its own: the next clustering pass sees a
|
||
face that still looks like the person it left, suggests it back, and the user's correction becomes an
|
||
argument they keep having. §9's cannot-link constraint handles co-occurrence; this is the same idea
|
||
applied to a judgement the user made by hand.
|
||
|
||
Two things the build changed in this document's own design:
|
||
|
||
**`crop_px` reached the schema** as §7 argued it should, and the calibration carries a `w_size` term
|
||
for it.
|
||
|
||
**A rejection table was added**, which §8 and §9 did not contemplate. Rejection is not the absence of
|
||
an assignment: without storing it, the next clustering pass re-suggests exactly the face the user
|
||
just pushed away. It is user data in the same sense a confirmation is (FR-CULL-12), just negative.
|
||
|
||
---
|
||
|
||
## 14. What this does not settle
|
||
|
||
- **Whether a permissive model pair exists.** §2.3. If it does, route B replaces route C and step 8's
|
||
first-run flow shrinks to nothing.
|
||
- **Whether tract runs these graphs at all.** §12 M1, and §4.1 says why the answer is in real doubt.
|
||
Every other line of this document is conditional on it.
|
||
- **Faces in trashed images.** catalog.md §10.4's open question, unchanged: probably excluded from
|
||
suggestions but not deleted, so a restore does not re-index.
|
||
- ~~**Whether embeddings sync.**~~ **Answered: they do**, as sealed shards
|
||
(`dr_catalog::face_shard`). Indexing is hours of CPU and its result is byte-identical on every
|
||
device, so paying for it once per account rather than once per device is the whole argument.
|
||
§12.2's 3.5 images/second is also what makes the case: it is fast enough to be worth doing and slow
|
||
enough to be worth not repeating.
|
||
|
||
**Shards rather than the catalog snapshot**, which is the design decision worth recording. The
|
||
snapshot uploads whole on every sync, and a fully indexed 23.5k library carries ~30 MB of
|
||
embeddings — exactly the cost `dr_thumbs`'s 25 MB cap exists to bound. So the split follows the one
|
||
already in the tree: bulk immutable data in sealed shards, small mutable data in the snapshot.
|
||
Faces, landmarks, embeddings and run markers shard; people, names and assignments ride the catalog
|
||
and merge by uuid. The cap is *imported* from `dr_thumbs` rather than restated, because it is a
|
||
statement about transfer cost and two copies of it would drift.
|
||
|
||
Everything is keyed on `oc:fileid`, never `image_id`: a row id means nothing on another device.
|
||
- **Re-embedding at higher resolution.** §7's `crop_px` makes it a query rather than a full re-index,
|
||
but whether it is worth doing is an M4 question.
|
||
- **Approximate nearest neighbours.** §9 says brute force until measured otherwise. A 100k-face
|
||
library is where this stops being true.
|
||
|
||
---
|
||
|
||
## 15. Register entries
|
||
|
||
**D13** — *face inference runtime and model licensing* · the runtime half stays answered (`ort` +
|
||
`ort-tract`, unchanged since 2026-08-21). **The licensing half is answered conditionally by §2:**
|
||
route C — ship the code, not the weights — unblocks the build without breaking GPL-3.0 or any
|
||
distribution channel, and §2.3's reading may yet convert it to route B. Closing D13 outright waits on
|
||
that reading.
|
||
|
||
**D17** — *face model pair and distribution route* · **PROPOSED**. SCRFD-500M for detection,
|
||
MobileFaceNet/ArcFace for embedding, weights obtained by the user rather than redistributed, with a
|
||
pinned checksum and a licence notice shown before the first fetch. The alternative considered and
|
||
rejected is committing the InsightFace weights (§2.1 route A), which is not available at any price.
|
||
The model *pair* is better evidenced than a proposal usually is — §1's table is a measured comparison
|
||
over a 2,418-identity gallery, not a literature reading — so the open half of this decision is the
|
||
route, not the pair. Relates to: D13, D14, NFR-COMPAT-2, FR-CULL-8.
|
||
|
||
**D18** — *porting from `scene-actor-extraction`* · **PROPOSED, and the easy half of a decision**.
|
||
That project is MIT and by this project's author, so MIT-into-GPL-3.0-or-later is a compatible
|
||
one-way combination requiring attribution, not permission. Ported files carry a header naming the
|
||
origin. Worth recording because "we already have this working in another language" is exactly the
|
||
provenance that goes undocumented and then cannot be answered three years later.
|
||
|
||
**S14** — *face pipeline in Rust* · scope sharpened, and **substantially de-risked**, by this
|
||
document. §12 replaces "measure per-image latency, cluster purity and calibration convergence" with
|
||
ten specific measures. Two changes to the brief itself: its instruction to resolve the licence
|
||
question *before writing any of it* is relaxed to "before shipping any of it", because M1 is an
|
||
afternoon and is worth knowing first (§13); and its central question — whether the pipeline works at
|
||
all — is largely answered in advance by §1.1, leaving the runtime and the domain shift from film
|
||
frames to family snapshots as the real unknowns.
|
||
|
||
---
|
||
|
||
## 16. Requirements touched
|
||
|
||
| ID | How this document addresses it |
|
||
|---|---|
|
||
| FR-CULL-8 | §4 detection, §6 embedding, §7 the proxy tier and its consequences, §10 the `DetectFaces` job, §18 the re-index |
|
||
| FR-CULL-9 | §8 — pairs, fit, validity, and the reliability-diagram acceptance test |
|
||
| FR-CULL-10 | §9 constrained agglomeration, confirmations as anchors, split by re-agglomeration; §18 what a re-index carries across |
|
||
| FR-CULL-11 | §10 the `Person` selector term, confirmed-only by default |
|
||
| FR-CULL-12 | §10 schema unchanged from catalog.md §10.1: embeddings derived, names to the sidecar |
|
||
| NFR-SEC-5 | §11 — the obligations restated as structural properties, one of them CI-checkable |
|
||
| NFR-COMPAT-2 | §2 — why the obvious weights cannot ship, and what does instead |
|
||
| NFR-RES-2 | §1 the cheap model pair, §12 M2/M3/M10 on a phone |
|
||
| NFR-ARCH-2 | §10 background priority, preempted by visible work |
|
||
| FR-CULL-8a | §17 — eye state and sunglasses: the models, the crops, the readability rule, the measurements |
|
||
| FR-CULL-13 | §17.3 — the reading is evidence: a chip that narrows and a badge that explains, and nothing that writes a judgement |
|
||
|
||
---
|
||
|
||
## 17. Eyes and sunglasses · 2026-09-19
|
||
|
||
FR-CULL-8a's eye state, under FR-CULL-13's rule. Three more models run over every face the
|
||
pipeline already aligns, and what they produce is a **filter term** — "eyes open", on the people
|
||
filter — and a badge on the People screen. Nothing acts on it: FR-CULL-13 says the signal is
|
||
evidence, shown and filterable, and never holds the pen, and §3.9.1's exclusion of blink
|
||
*detection* was re-read the same day as the exclusion of blink *selection* it always was. The chip
|
||
narrows the grid the way a star count does, and rates nothing. Head pose, the other half of
|
||
FR-CULL-8a, is not built; §17.3 says what stands in for it.
|
||
|
||
### 17.1 The models, and why these
|
||
|
||
| | 2d106det | OCEC | SGC |
|
||
|---|---|---|---|
|
||
| Answers | 106 landmarks, ten round each eye's lids | P(this eye is open) | P(this head wears sunglasses) |
|
||
| Input | 192² RGB 0..255, the detector box at 1.5× | one eye, 40×24 RGB, `x/255` | one head, 48×48 RGB, `x/255` |
|
||
| Shipped | `buffalo_l`'s, 4.8 MB | S, 483 KB, F1 0.9943 on its own split | L, 6.1 MB, F1 0.9554 on its own split |
|
||
| Licence | InsightFace's research-only grant, like the pair (§2.2a) | MIT, code and weights | MIT, code and weights |
|
||
| Training data | InsightFace's | *Open and Closed Eyes* (ODC-By 1.0) + Wholebody34 crops (Apache 2.0) | **not stated** — recorded in `models/face/README.md` |
|
||
| Cost in tract | ~24 ms per face | ~6 ms per eye | ~6 ms per framing, two framings |
|
||
|
||
The two classifiers are from Katsuya Hyodo's ultra-lightweight series — the same author as the
|
||
whole-body detector §1.1's reference pipeline uses — and are the first weights in `models/face/`
|
||
that do not come out when the project publishes. The landmark model is under the grant the pair
|
||
already carries; it was chosen over two permissively licensed alternatives on a measurement
|
||
(§17.2) after the decision that this project will not be commercial, which is what §2.2a already
|
||
records for the pair.
|
||
|
||
All three load in tract as shipped, dynamic batch and all — the first graphs in this subsystem to
|
||
do so — and are pinned to a batch of 1 by `tools/fix-face-model-shapes.sh` anyway, because a graph
|
||
the engine *analyses* and a graph it has been *measured running* are different claims, and the
|
||
embedder's precedent is the safer one. Six milliseconds per classifier call against 0.4 in the
|
||
reference README is tract's per-call overhead on a graph this small; the whole reading is under
|
||
60 ms per face beside an embedding at 160 ms and a native decode in seconds.
|
||
|
||
### 17.2 Where the eye box comes from
|
||
|
||
**The eye classifier was trained on a whole-body detector's eye boxes, and this pipeline has no
|
||
eye boxes.** It has five landmarks, and SCRFD's eye point is loose: it is one of five points that
|
||
place a face, not an eye centre, and on a turned or smiling head the eye sat in a corner of a
|
||
window centred on it. Everything below was measured on 60 proxies from the reference library with
|
||
25 plainly open-eyed faces labelled by hand (`examples/eyes.rs --dump`, then a contact sheet), and
|
||
the count that matters is how many of those 25 the classifier read as open in both eyes.
|
||
|
||
**A window on the SCRFD point: 19 of 25.** Windows from 20×10 to 34×17 template units all gave
|
||
19–20; smaller lost eyes. Two model-free ways of re-centring the window were then tried and both
|
||
lost eyes: the darkest blob near the landmark is the inner corner's shadow or the lash line
|
||
(19 → 15), and the most contrasty window is the one that takes in the edge of the nose (19 → 9).
|
||
The landmark as SCRFD gives it beats either.
|
||
|
||
**A box from a landmark model's lid contour: 22 of 25.** Three models were run over the same
|
||
faces, each fed the crop its reference code feeds it, and the eye box cut as the bounding box of
|
||
the lid points grown by a margin:
|
||
|
||
| model | points | input | tract | per face | open at margin 0.1 |
|
||
|---|---|---|---|---|---|
|
||
| MediaPipe Face Mesh V2 (Apache 2.0) | 478, with z | 256² | loads | ~36 ms | 22 |
|
||
| PIPNet, PINTO's irnet18 export (WFLW, research-only) | 68 | 256² | loads | ~98 ms | 20 |
|
||
| **InsightFace 2d106det** | **106** | **192²** | **loads** | **~24 ms** | **22** |
|
||
|
||
The margin was swept on the two that tied: 22 at 0 and 0.1, 18 at 0.4, 14 at 0.6 — the training
|
||
crops were tight detector boxes, and a tight box is what the classifier wants (`EYE_BOX_MARGIN`).
|
||
2d106det ships: it tied the best, costs the least, and is under a grant the project has already
|
||
accepted. Face Mesh would be the choice if that changed; it also gives z and an iris, neither of
|
||
which this needs yet.
|
||
|
||
The box is cut **upright from the native render**, not through the face's alignment — the
|
||
training crops were detector boxes, and the contour already says where the eye is on a tilted
|
||
head (`align::eye_patch`). A shut eye's contour has no height and is given an open eye's
|
||
(`EYE_BOX_MIN_ASPECT`), so the classifier sees the same framing either way.
|
||
|
||
### 17.3 Not asking what cannot be answered
|
||
|
||
The 25 open faces were never the real problem. The real problem was the faces that were *not*
|
||
open-eyed by the classifier's account and were not blinks either, and on the reference sample they
|
||
were the commonest wrong answer of all: **a soft eye reads as closed.** A face small enough that
|
||
its eye was seven pixels wide, a motion-blurred face, a face from a 1024 proxy where the native
|
||
render should have been — each produced a confident "closed" from a classifier shown a smear. The
|
||
same failure the face's own sharpness gate exists for (§4.3), one stage down, where the face's gate
|
||
cannot see it: a face sharp enough to embed can hold an eye too soft to read, because the eye is a
|
||
fortieth of it.
|
||
|
||
So the reading is **seven numbers, not a verdict** — per eye P(open), the source pixels across its
|
||
box and the sharpness of the patch the classifier saw; and P(sunglasses) — stored as such
|
||
(`faces.eye_right`, `faces.eye_right_px`, `faces.eye_right_sharp`, likewise `eye_left`, and
|
||
`faces.sunglasses`, schema V16), and the verdict is a rule with thresholds in it,
|
||
`dr_face::eyes::EyeReading::state`, the only place the thresholds live:
|
||
|
||
```
|
||
sunglasses ≥ 0.5 → Sunglasses (whatever the eyes said)
|
||
an eye is readable when px ≥ 12
|
||
and sharpness ≥ 0.02
|
||
and px ≥ 0.6 × the other eye's px
|
||
no readable eye → Unreadable
|
||
a readable eye < 0.5 → Closed (a blink, or a wink)
|
||
otherwise → Open
|
||
```
|
||
|
||
**Sunglasses take precedence** because the eye classifier answers confidently over dark glass:
|
||
over a woman in sunglasses on the reference library it read her right eye 0.97 open. **The
|
||
pixel floor** is where the classifier's own training stopped — its reference footage averaged
|
||
15–21 pixels an eye. **The sharpness floor** is the face's measure over the patch, set where the
|
||
sample's open eyes were being called closed: the open set ran from 0.019 (a lens reflection) to
|
||
5.4, the unreadable ones under 0.02 with the pixels to match. **The width ratio** is the profile:
|
||
a landmark model's contour for the far eye of a turned head collapses towards the nose. On the
|
||
twenty native renders of §17.4, profiles put the far eye at 0.02–0.43 of the near one's width,
|
||
two three-quarter faces whose far eye was reading closed sat at 0.54, and every face looking at
|
||
the camera sat at 0.78 or more — a shut eye's box keeps its width, so a wink is not mistaken for
|
||
a turn. 0.6 splits the gap. An eye that fails any of the three is not asked, the near eye still
|
||
decides, and a face with no readable eye is *unclear* — which is not a blink, and not open, and
|
||
which no filter drops.
|
||
|
||
**The two eyes are kept apart** rather than averaged, because a wink averages to 0.5 — the one
|
||
value that says the least — and "eyes open" means every eye that could be read.
|
||
|
||
With the rule in place, the same 60 proxies read: 33 open, 14 closed, 24 sunglasses, 14 unclear.
|
||
Of the 25 labelled open faces, 22 open, 2 unclear (eye boxes of 7 and 12 pixels on a child's
|
||
face), 1 closed — a squinting smile whose contour collapsed to eleven pixels, which the classifier
|
||
is not wrong to call narrow. The 14 closed are downcast eyes, laughs, two sunglasses the head
|
||
classifier missed, and the squint. A face 141 pixels across the eye but motion-blurred to a
|
||
sharpness of 0.016 reads *unclear* where it read *closed* before, which is the change this
|
||
section is for.
|
||
|
||
**The filter drops only *Closed*.** `RatingFilter::eyes_open` compiles the rule above into a
|
||
predicate on the face row, ANDed into the chosen people's face subquery, so "Anna, eyes open" asks
|
||
about Anna's face and not about Bob blinking beside her. The chip is offered only while someone is
|
||
chosen and goes when the last person does — without a name in front of it, it would be a verdict
|
||
on everyone in the frame. The predicate still handles the empty case, as `NOT EXISTS` over every
|
||
face, for a filter arriving by another route; a landscape passes because there is no one in it to
|
||
have blinked. Sunglasses pass. Unclear passes. Never read passes — that last is what keeps an old
|
||
library from emptying its grid the moment the chip is pressed: until the measuring pass has run,
|
||
the honest answer is "everything". A test drives the same five readings through the SQL and
|
||
through `state()` and requires the two to agree, so the badge and the grid cannot say different
|
||
things.
|
||
|
||
**And an index, learned the slow way.** The people filter was always served from a covering index
|
||
on `faces(image_id)`; the moment its subquery read the eye columns it had to read the face *row*,
|
||
and `ALTER TABLE ADD COLUMN` had put those seven floats after the embedding and the crop blob —
|
||
six kilobytes to reach every one. One count took 24 seconds on the reference library, thirteen of
|
||
them system time. `faces_eyes` (schema V17) covers the subquery again: five milliseconds.
|
||
|
||
### 17.4 What the sample says about accuracy, and what it does not
|
||
|
||
The 60-proxy sample above was run at proxy resolution, where the production pass reads the native
|
||
render; the eye box on a 200-pixel face is 40 source pixels from the proxy and 240 from the
|
||
original. So the shipped configuration was also run over **twenty native renders** from the
|
||
reference library — a wedding burst of six frames with six or seven faces each, and a dozen
|
||
singles — exported by `face_native --export` and read by `examples/eyes.rs --dump`, 62 faces in
|
||
all: 20 open, 28 closed, 12 sunglasses, 2 unclear before the width ratio was moved (below).
|
||
|
||
Read off the contact sheet, face by face: the one real blink in the set (`7884.dng`, a man with
|
||
his eyes shut) is *closed*; the laughing faces with their eyes screwed shut are *closed*, which a
|
||
photographer would call right; the downcast faces are *closed*, which is arguable; the profiles
|
||
are judged on the near eye and mostly *open*, which the SCRFD-point pass could not do. Two
|
||
faces were wrong: three-quarter views whose far eye's box came to 0.54 of the near one's and read
|
||
closed over a cheek, which is what moved the width ratio from 0.45 to 0.6. One is beyond any
|
||
floor: a face with a porcelain pot held over the eyes, whose contour is a guess and whose "eyes"
|
||
are sharp white china.
|
||
|
||
Two things no sample so far can say. None holds more than one real blink, so the precision of
|
||
*Closed* is not measured — every closed verdict on the two sheets but the occlusion and the
|
||
sunglasses misses is a narrowed or shut eye rather than a wrong one, but that is a reading of a
|
||
contact sheet, not a number. And the floors were set on a few dozen faces. Both are M11.
|
||
|
||
### 17.4a What is kept per face
|
||
|
||
The box, the five landmarks, `crop_px`, the embedding and the crop were already there. The eye
|
||
pass adds the seven numbers of §17.3 and the **106 dense landmarks** it read the eye boxes from,
|
||
packed as 16-bit fixed point over the frame — 424 bytes a face, a seventh of a pixel on a
|
||
6000-pixel frame (`dr_face::Landmarks::to_packed_bytes`, schema V18). Kept for the reason the
|
||
embedding is kept: it cost a fetch of the original and a model run, and the next per-face pass —
|
||
head pose, expression, whatever FR-CULL-8a grows — should run from the catalog. `f16` would have
|
||
been the same size and worse: three figures near 1.0 is six pixels at that scale.
|
||
|
||
### 17.5 The measuring pass, and shards
|
||
|
||
A face indexed before the models existed, or on a device without them, has no reading. The
|
||
`face-eyes` repair (§18.1; on the day this was written, the sweep's measuring pass — the one V14
|
||
built to re-embed faces stored as unit vectors) lists those faces, on a device that has the models,
|
||
and reads their eyes from the same native render with the box and landmarks already stored. No
|
||
detector runs and no identity moves. A device *without* the models has no such repair, or it would
|
||
fetch every original in the library to do nothing to it; the repair's predicate is the one both the
|
||
count and the work list use, so the pass converges. The People screen's coverage line counts
|
||
these faces as work to read and keeps the button while any remain — reading **Read eye state**
|
||
once detection is complete and only readings are left, which is the state an already-indexed
|
||
library is in the day the models arrive.
|
||
|
||
Shards carry the seven columns beside `quality`. A peer's faces without a reading are **adopted**,
|
||
unlike a peer's faces without a quality (§14): the measuring pass finds this work by the NULL and
|
||
not by the run marker, so adoption costs the reading nothing, and a peer with no eye models may be
|
||
the only device that has done the detection at all.
|
||
|
||
### 17.6 Still to measure
|
||
|
||
| # | Measure | Why it decides something |
|
||
|---|---|---|
|
||
| **M11** | Open-eye recall and blink precision from the **native** pass with the shipped configuration, on a labelled sample that contains real blinks — a burst with one in it is enough | §17.3's figures are proxy figures from 25 faces, with no precision beside them; this is the number FR-CULL-13's acceptance clause asks for, and where the two readability floors get set on more than 25 faces |
|
||
| **M12** | Whether the remaining open-eye failure — a lens reflection over an open eye — moves with the sharpness floor, or needs the eye classifier told about spectacles | If the latter, the fix is a classifier trained closer to this domain, and that is a different decision |
|
||
| **M13** | Sunglasses recall on more than twelve faces, and the false-positive rate on caps and clear glasses | The two sunglasses the head classifier missed on the sample became false blinks; a library of skiers would say whether that is two faces or a class |
|
||
|
||
---
|
||
|
||
## 18. The completeness job, and what a re-index carries across · 2026-09-19
|
||
|
||
The reference library on the day this was written: 17,762 faces under the bare `w600k_mbf` id —
|
||
found by the fast detector on 1024 px proxies, stored as unit vectors, no quality, no crop on 4,144
|
||
of them, no eye reading, no dense landmarks — beside 1,177 under `scrfd_10g+w600k_mbf` from the
|
||
native pass, and 12,217 images the fast detector examined and found nothing in. Over those faces:
|
||
3,778 confirmations, 13,011 suggestions, 77 rejections, and 17,276 people rows. Every one of those
|
||
gaps was, until now, its own pass: V14's measuring pass for the quality, §17.5's for the eyes, the
|
||
sweep's proxy repair, the sweep's detector upgrade, and a re-index that did not exist. Adding a
|
||
per-face field meant adding a pass, with its own work list, its own count and its own idea of done.
|
||
|
||
### 18.1 One job over a registry
|
||
|
||
`dr_ui::repairs` replaces them with one job over a **registry**. A `Repair` names one thing a catalog
|
||
record can lack — the predicate that says which images still owe it, the input its handler needs
|
||
(the file's header, the whole original, or a native render), the handler that fills it, and, where
|
||
there is one, what to record for an image that can never be done. The job unions the predicates
|
||
into one work list, fetches each image once at the most any claimant asks for, renders it at most
|
||
once, and runs every handler whose predicate that image still matches — checked again before each,
|
||
because one handler's write satisfies the next's (a detection writes every field a per-face handler
|
||
would fill). The registry today:
|
||
|
||
| Repair | Owed by | Input | Handler |
|
||
|---|---|---|---|
|
||
| `face-proxy` (sweep only) | images holding faces whose 1024 px proxy is not in the store | native render | detect again, write the proxy |
|
||
| `face-quality` | faces with `quality IS NULL` | native render | warp from the stored landmarks, embed, write the raw vector and its length (reads eyes on the same warp where it can) |
|
||
| `face-eyes` | faces with no eye reading or no dense landmarks, on a device with the eye models | native render | read the eyes and dense landmarks from the stored box and landmarks |
|
||
| `face-crop` | faces with `crop IS NULL` | native render | cut the crop from the frame |
|
||
| `face-detection` | sweep: images with no marker under the embedder and no faces; re-index: images with no marker under the **chosen detector**, either spelling | native render | detect, embed, replace the faces, carry identities across (§18.2) |
|
||
| `face-upgrade` (sweep only) | images whose marker is a weaker detector's | native render | as `face-detection` |
|
||
| `metadata` | `images.metadata_state < 2` | header | EXIF to the catalog, the dateless marked examined |
|
||
|
||
A repair's predicate is the *only* definition of its work. The count the settings page shows
|
||
(`faces::audit`, per repair), the list the job fetches and the check before each handler are one
|
||
predicate, so a record the count reports is one the job fetches and one the handler fills, and the
|
||
job ends. This is why an eye reading that cannot be cut is not a criterion — such a face stays
|
||
unread however often it is detected, and listing it would fetch its original on every press — and
|
||
why a degenerate face is dropped rather than left. It is also why the registry is cut to what the
|
||
device can do (`Capabilities`): a device without the eye models has no `face-eyes` entry, rather
|
||
than an entry it skips, because an entry is a count and a set of originals to fetch.
|
||
|
||
The registry is ordered, and the order is the work list's: an image only the last repair claims
|
||
comes after one the first does, which is what puts a few hundred proxy repairs ahead of twenty
|
||
thousand un-indexed images. Within one image the same order runs the handlers, detection before the
|
||
per-face repairs, since detection fills what they would.
|
||
|
||
Adding a field is one entry: a predicate over `faces f` or `images i`, and a handler that fills it
|
||
from `Fetched`. `metadata` is in the table to say that this is not a face job — the same machinery
|
||
carries a capture date, and could carry a thumbnail, a perceptual hash or a head pose.
|
||
|
||
### 18.1a The two scopes
|
||
|
||
Both buttons on the settings page run the job; they differ in one predicate. **Index faces**
|
||
(`Scope::Outstanding`) converges on *coverage* — has anything examined this image — and treats a
|
||
face a weaker detector found on a proxy as found, which is the right question for a pass that must
|
||
not fetch the library twice. **Re-index every face** (`Scope::Reindex`) converges on *provenance*:
|
||
`face-detection` claims every image with no marker under the chosen detector, in either of its
|
||
forms (`FaceDetector::model_ids`, so a desktop running it in f32 and a tablet on the Hexagon in int8
|
||
do not re-index each other's work), and a marker saying a weaker one looked is not that. This is
|
||
the one place in the subsystem keyed on the exact detector rather than the embedder. Convergent all
|
||
the same: an image the job has been through leaves the list, a kill costs the images in flight, and
|
||
a second press resumes.
|
||
|
||
An original over the fetch budget is skipped without being fetched. Under the sweep, detection
|
||
records an examination that found nothing — the honest record for an image that cannot be
|
||
examined, and what stops the half gigabyte being spent once per sweep. Under the re-index, and
|
||
under every repair over records that already exist, it is left exactly as it was: a re-detection
|
||
with nothing found would delete the faces, and "cannot fetch" is not "no faces".
|
||
|
||
### 18.2 What is carried across
|
||
|
||
`dr_catalog::faces::record_detections` replaces an image's faces and carries identities onto the
|
||
new ones. Before this section it carried confirmations only, by box overlap above 0.5 IoU, and a
|
||
re-detection of the library above would have left 13,011 suggestions and 77 rejections on the
|
||
floor — correct by FR-CULL-12's letter, since suggestions are derived data, and a People screen
|
||
emptied to strangers by the user's own button.
|
||
|
||
Now every old face is read before the delete — box, vector, assignment, rejections — and matched to
|
||
the new faces one-to-one, best pair first. A pair qualifies when the boxes **overlap at all** and
|
||
either the overlap alone says so (IoU above 0.5, the old rule) or the embeddings do (cosine above
|
||
`SAME_FACE_COSINE` = 0.45, the reference library's P≈0.95 line from §9's table). The embedding route
|
||
is for the box a low-resolution pass drew badly enough that overlap alone would not claim it; the
|
||
vector is also what breaks the tie in a group photograph, where two neighbouring faces overlap both
|
||
new boxes. Overlap is required on both routes because the same vector elsewhere in the frame — a
|
||
mirror, a print on the wall — is not the same face and must not take its name. Onto the matched
|
||
face go the assignment as it was, confirmed or suggested with its probability, and every
|
||
rejection.
|
||
|
||
It is a match, not an update in place, and that is why the per-face repairs exist beside
|
||
detection: where nothing about a face but one field needs doing, `record_updates` keeps the id and
|
||
there is nothing to judge.
|
||
|
||
Since #77 (0.18.0) the merge's `match_faces` answers it too, within a photograph's `file_id` and
|
||
one embedder: box IoU ≥ 0.5, unique on both sides, first; then, only for photographs where a remote
|
||
face is left over and a local face is free, embedding cosine ≥ 0.7, mutual best, with a lead of
|
||
≥ 0.2 over the runner-up on both sides. A box match is never overruled by a low cosine (about 150
|
||
genuine cross-device pairs of tiny faces score below 0.45). On the reference desktop/tablet pair this
|
||
recovers 20 of 631 unmatched faces with no false matches; the rest are faces one device alone found.
|
||
The merge also keeps one person to one face per photograph: an incoming assignment is refused when
|
||
another local face already holds that person, unless it is a remote confirmation over a local
|
||
suggestion, which moves the suggestion. Refusals are counted in `faces_one_per_photograph`.
|
||
|
||
The threshold differs from `SAME_FACE_COSINE` (0.45) above on purpose: re-detection additionally
|
||
requires the boxes to overlap, while the merge's embedding route exists for boxes that don't.
|
||
|
||
## 19. Deduplicating people · 2026-09-26
|
||
|
||
`dr_catalog::dedup_people::run` runs after every successful sync merge (`sync::merge_remote`, on the
|
||
sync worker), in one transaction, and logs one `dedup:` line (#78).
|
||
|
||
**People.** Named people with the same name, trimmed and case-folded, merge into the one with the
|
||
most confirmed faces (ties go to the smaller uuid) when every shared embedder's confirmed-face
|
||
centroids agree at cosine ≥ 0.7 (distance < 0.3). Each side needs at least two confirmed faces to
|
||
compare; a namesake holding no faces merges outright; a face confirmed as one and rejected as the
|
||
other keeps them apart; unnamed and set-aside people are never touched. On the reference library the
|
||
same-person centroid median is 0.91, and different named people have a 99.9th percentile of 0.41.
|
||
|
||
**Faces.** Two faces in the same image and embedder with IoU ≥ 0.5 and cosine ≥ 0.7 are one: the
|
||
job keeps the stronger detector's face (`FaceDetector::outranks`), then the confirmed one, then the
|
||
lower id, and it takes both faces' assignment and rejections.
|
||
|
||
**Propagation.** The merge is `faces::merge_people`, whose `merged_into` redirect a 0.17.0 peer
|
||
already honours, so an older device never resurrects the duplicate. The job also follows redirects
|
||
left by earlier manual merges, moving this device's own assignments onto the person kept, and
|
||
breaks a mutual redirect at the smaller uuid, which every device computes alike. A merge now also
|
||
carries the merged-away person's rejections to the person kept.
|