One completeness job over a registry of repairs, and a re-index button

A library's records are never all complete at once. A face found before
its quality was kept has no quality; one found before the eye models
existed has no reading; one adopted from a peer's shard has no crop; an
image the fast detector examined on a 1024 px proxy has boxes the current
detector would not have drawn; an image the scan stat'ed has no capture
date. On the reference library that is 17,762 faces under the bare
w600k_mbf id with no quality, no reading and no dense landmarks, 4,144 of
them without a crop, beside 12,217 images the fast detector examined and
found nothing in. Every one of those gaps was its own pass — V14's
measuring pass, §17.5's eye pass, the sweep's proxy repair, the sweep's
detector upgrade — with its own work list, its own count and its own idea
of done, and adding a per-face field meant adding a pass. There was no
pass at all for the case the library is actually in: boxes and landmarks
drawn by a weaker detector on a proxy, which every later per-face pass
would have read from.

dr_ui::repairs replaces them with one job over a registry. A Repair names
one thing a record can lack — the predicate that says which images still
owe it, the input its handler needs (a header, the original, or a native
render), the handler, and what to record for an image that can never be
done. The job unions the predicates into one work list, fetches each
image once at the most any claimant asks for, renders it at most once,
and runs every handler whose predicate that image still matches, checked
again before each because a detection writes every field a per-face
handler would fill. The registry today: face-proxy, face-quality,
face-eyes, face-crop, face-detection, face-upgrade, metadata — the last
there to say that this is not a face job. Adding a field is one entry.

A repair's predicate is the only definition of its work: the count the
settings page shows, the list the job fetches and the check before its
handler run are one predicate, so the job converges. That is why the
registry is cut to what the device can do rather than listing what it
skips — an entry is a count and a set of originals to fetch — and why an
eye reading that cannot be cut is not a criterion.

The catalog side is generic to match: record_updates writes whichever
fields a FaceUpdate carries and re-marks the image so the shards export
it; faces_needing and count_needing answer a predicate the caller
supplies, replacing the measuring pass's three special cases.

Two buttons on the settings page run the job and differ in one
predicate. "Index faces" converges on coverage: has anything examined
this image. "Re-index every face" converges on provenance: face-detection
claims every image with no marker under the chosen detector, in either
of its forms (FaceDetector::model_ids, so a desktop in f32 and a tablet
on the Hexagon do not re-index each other's work), and a marker saying a
weaker one looked is not that. An original over the fetch budget is left
exactly as it was under the re-index, where the sweep marks it examined:
a re-detection with nothing found would delete the faces, and "cannot
fetch" is not "no faces".
This commit is contained in:
2026-09-19 18:52:13 +02:00
parent 2a4ac0ed3d
commit 5c00942b84
15 changed files with 2201 additions and 1450 deletions
+111 -14
View File
@@ -816,10 +816,10 @@ here, and it is the phone and tablet story that should decide whether it gets bu
anything downstream sees them. A probe's confidence (§9.1) is computed from the references it
matched; a reference's confidence hears nothing from a probe. A face whose quality was never
recorded is admitted to the gallery — a rule that cannot be checked admits rather than excludes —
and the next indexing pass **measures** it: `faces_unmeasured` lists every image holding one, and
each such face is embedded again from the native render with the landmarks it already has, the raw
vector written over the old one and its id, box and identity untouched
(`faces::record_measurements`). No detector runs and no suggestion is lost — the cost is the
and the next indexing pass **measures** it: the `face-quality` repair (§18.1) lists every image
holding one, and each such face is embedded again from the native render with the landmarks it
already has, the raw vector written over the old one and its id, box and identity untouched
(`faces::record_updates`). No detector runs and no suggestion is lost — the cost is the
original fetched once more, since the length exists only at the moment of embedding.
**The algorithm.** Constrained average-link agglomeration over the probability graph, merging while
@@ -1119,8 +1119,8 @@ variant; its vectors are one space, and the detector only decides where the boxe
So every reader now keys on the embedder half of the id (`faces::embedder_of`, and
`embedder_sql` for the queries): all three detectors are one population, and changing between
them empties nothing. `record_detections` is unchanged — an image holds one pipeline's faces at a
time, and a re-detection carries confirmations across by box overlap — and it is where the
generations meet. The merge's `match_faces` matches within an embedder for the same reason. The
time, and a re-detection carries identities across by box overlap and embedding (§18) — and it is
where the generations meet. The merge's `match_faces` matches within an embedder for the same reason. The
shards travel every generation, each under its own id, and a peer adopts whichever it is sent.
What a stronger choice still does is queue the images a weaker detector indexed for re-detection
(`FaceDetector::supersedes`), after the ones nothing has indexed and never downwards, so a tablet
@@ -1257,9 +1257,9 @@ frames to family snapshots as the real unknowns.
| ID | How this document addresses it |
|---|---|
| FR-CULL-8 | §4 detection, §6 embedding, §7 the proxy tier and its consequences, §10 the `DetectFaces` job |
| FR-CULL-8 | §4 detection, §6 embedding, §7 the proxy tier and its consequences, §10 the `DetectFaces` job, §18 the re-index |
| FR-CULL-9 | §8 — pairs, fit, validity, and the reliability-diagram acceptance test |
| FR-CULL-10 | §9 constrained agglomeration, confirmations as anchors, split by re-agglomeration |
| FR-CULL-10 | §9 constrained agglomeration, confirmations as anchors, split by re-agglomeration; §18 what a re-index carries across |
| FR-CULL-11 | §10 the `Person` selector term, confirmed-only by default |
| FR-CULL-12 | §10 schema unchanged from catalog.md §10.1: embeddings derived, names to the sidecar |
| NFR-SEC-5 | §11 — the obligations restated as structural properties, one of them CI-checkable |
@@ -1447,12 +1447,12 @@ been the same size and worse: three figures near 1.0 is six pixels at that scale
### 17.5 The measuring pass, and shards
A face indexed before the models existed, or on a device without them, has no reading. The sweep's
measuring pass — the one V14 built to re-embed faces stored as unit vectors — lists those faces
too, on a device that has the models, and reads their eyes from the same native render with the
box and landmarks already stored (`dr_ui::faces::measure_native`). No detector runs and no identity
moves. A device *without* the models does not list them, or it would fetch every original in the
library to do nothing to it; `dr_catalog::faces::unmeasured_sql` is the one predicate both the
A face indexed before the models existed, or on a device without them, has no reading. The
`face-eyes` repair (§18.1; on the day this was written, the sweep's measuring pass — the one V14
built to re-embed faces stored as unit vectors) lists those faces, on a device that has the models,
and reads their eyes from the same native render with the box and landmarks already stored. No
detector runs and no identity moves. A device *without* the models has no such repair, or it would
fetch every original in the library to do nothing to it; the repair's predicate is the one both the
count and the work list use, so the pass converges. The People screen's coverage line counts
these faces as work to read and keeps the button while any remain — reading **Read eye state**
once detection is complete and only readings are left, which is the state an already-indexed
@@ -1470,3 +1470,100 @@ the only device that has done the detection at all.
| **M11** | Open-eye recall and blink precision from the **native** pass with the shipped configuration, on a labelled sample that contains real blinks — a burst with one in it is enough | §17.3's figures are proxy figures from 25 faces, with no precision beside them; this is the number FR-CULL-13's acceptance clause asks for, and where the two readability floors get set on more than 25 faces |
| **M12** | Whether the remaining open-eye failure — a lens reflection over an open eye — moves with the sharpness floor, or needs the eye classifier told about spectacles | If the latter, the fix is a classifier trained closer to this domain, and that is a different decision |
| **M13** | Sunglasses recall on more than twelve faces, and the false-positive rate on caps and clear glasses | The two sunglasses the head classifier missed on the sample became false blinks; a library of skiers would say whether that is two faces or a class |
---
## 18. The completeness job, and what a re-index carries across · 2026-09-19
The reference library on the day this was written: 17,762 faces under the bare `w600k_mbf` id —
found by the fast detector on 1024 px proxies, stored as unit vectors, no quality, no crop on 4,144
of them, no eye reading, no dense landmarks — beside 1,177 under `scrfd_10g+w600k_mbf` from the
native pass, and 12,217 images the fast detector examined and found nothing in. Over those faces:
3,778 confirmations, 13,011 suggestions, 77 rejections, and 17,276 people rows. Every one of those
gaps was, until now, its own pass: V14's measuring pass for the quality, §17.5's for the eyes, the
sweep's proxy repair, the sweep's detector upgrade, and a re-index that did not exist. Adding a
per-face field meant adding a pass, with its own work list, its own count and its own idea of done.
### 18.1 One job over a registry
`dr_ui::repairs` replaces them with one job over a **registry**. A `Repair` names one thing a catalog
record can lack — the predicate that says which images still owe it, the input its handler needs
(the file's header, the whole original, or a native render), the handler that fills it, and, where
there is one, what to record for an image that can never be done. The job unions the predicates
into one work list, fetches each image once at the most any claimant asks for, renders it at most
once, and runs every handler whose predicate that image still matches — checked again before each,
because one handler's write satisfies the next's (a detection writes every field a per-face handler
would fill). The registry today:
| Repair | Owed by | Input | Handler |
|---|---|---|---|
| `face-proxy` (sweep only) | images holding faces whose 1024 px proxy is not in the store | native render | detect again, write the proxy |
| `face-quality` | faces with `quality IS NULL` | native render | warp from the stored landmarks, embed, write the raw vector and its length (reads eyes on the same warp where it can) |
| `face-eyes` | faces with no eye reading or no dense landmarks, on a device with the eye models | native render | read the eyes and dense landmarks from the stored box and landmarks |
| `face-crop` | faces with `crop IS NULL` | native render | cut the crop from the frame |
| `face-detection` | sweep: images with no marker under the embedder and no faces; re-index: images with no marker under the **chosen detector**, either spelling | native render | detect, embed, replace the faces, carry identities across (§18.2) |
| `face-upgrade` (sweep only) | images whose marker is a weaker detector's | native render | as `face-detection` |
| `metadata` | `images.metadata_state < 2` | header | EXIF to the catalog, the dateless marked examined |
A repair's predicate is the *only* definition of its work. The count the settings page shows
(`faces::audit`, per repair), the list the job fetches and the check before each handler are one
predicate, so a record the count reports is one the job fetches and one the handler fills, and the
job ends. This is why an eye reading that cannot be cut is not a criterion — such a face stays
unread however often it is detected, and listing it would fetch its original on every press — and
why a degenerate face is dropped rather than left. It is also why the registry is cut to what the
device can do (`Capabilities`): a device without the eye models has no `face-eyes` entry, rather
than an entry it skips, because an entry is a count and a set of originals to fetch.
The registry is ordered, and the order is the work list's: an image only the last repair claims
comes after one the first does, which is what puts a few hundred proxy repairs ahead of twenty
thousand un-indexed images. Within one image the same order runs the handlers, detection before the
per-face repairs, since detection fills what they would.
Adding a field is one entry: a predicate over `faces f` or `images i`, and a handler that fills it
from `Fetched`. `metadata` is in the table to say that this is not a face job — the same machinery
carries a capture date, and could carry a thumbnail, a perceptual hash or a head pose.
### 18.1a The two scopes
Both buttons on the settings page run the job; they differ in one predicate. **Index faces**
(`Scope::Outstanding`) converges on *coverage* — has anything examined this image — and treats a
face a weaker detector found on a proxy as found, which is the right question for a pass that must
not fetch the library twice. **Re-index every face** (`Scope::Reindex`) converges on *provenance*:
`face-detection` claims every image with no marker under the chosen detector, in either of its
forms (`FaceDetector::model_ids`, so a desktop running it in f32 and a tablet on the Hexagon in int8
do not re-index each other's work), and a marker saying a weaker one looked is not that. This is
the one place in the subsystem keyed on the exact detector rather than the embedder. Convergent all
the same: an image the job has been through leaves the list, a kill costs the images in flight, and
a second press resumes.
An original over the fetch budget is skipped without being fetched. Under the sweep, detection
records an examination that found nothing — the honest record for an image that cannot be
examined, and what stops the half gigabyte being spent once per sweep. Under the re-index, and
under every repair over records that already exist, it is left exactly as it was: a re-detection
with nothing found would delete the faces, and "cannot fetch" is not "no faces".
### 18.2 What is carried across
`dr_catalog::faces::record_detections` replaces an image's faces and carries identities onto the
new ones. Before this section it carried confirmations only, by box overlap above 0.5 IoU, and a
re-detection of the library above would have left 13,011 suggestions and 77 rejections on the
floor — correct by FR-CULL-12's letter, since suggestions are derived data, and a People screen
emptied to strangers by the user's own button.
Now every old face is read before the delete — box, vector, assignment, rejections — and matched to
the new faces one-to-one, best pair first. A pair qualifies when the boxes **overlap at all** and
either the overlap alone says so (IoU above 0.5, the old rule) or the embeddings do (cosine above
`SAME_FACE_COSINE` = 0.45, the reference library's P≈0.95 line from §9's table). The embedding route
is for the box a low-resolution pass drew badly enough that overlap alone would not claim it; the
vector is also what breaks the tie in a group photograph, where two neighbouring faces overlap both
new boxes. Overlap is required on both routes because the same vector elsewhere in the frame — a
mirror, a print on the wall — is not the same face and must not take its name. Onto the matched
face go the assignment as it was, confirmed or suggested with its probability, and every
rejection.
It is a match, not an update in place, and that is why the per-face repairs exist beside
detection: where nothing about a face but one field needs doing, `record_updates` keeps the id and
there is nothing to judge.
The merge's `match_faces` still matches by overlap alone across devices. It is the same question,
and the same answer would serve it; it is not changed here.
+13 -13
View File
File diff suppressed because one or more lines are too long