One completeness job over a registry of repairs, and a re-index button
A library's records are never all complete at once. A face found before its quality was kept has no quality; one found before the eye models existed has no reading; one adopted from a peer's shard has no crop; an image the fast detector examined on a 1024 px proxy has boxes the current detector would not have drawn; an image the scan stat'ed has no capture date. On the reference library that is 17,762 faces under the bare w600k_mbf id with no quality, no reading and no dense landmarks, 4,144 of them without a crop, beside 12,217 images the fast detector examined and found nothing in. Every one of those gaps was its own pass — V14's measuring pass, §17.5's eye pass, the sweep's proxy repair, the sweep's detector upgrade — with its own work list, its own count and its own idea of done, and adding a per-face field meant adding a pass. There was no pass at all for the case the library is actually in: boxes and landmarks drawn by a weaker detector on a proxy, which every later per-face pass would have read from. dr_ui::repairs replaces them with one job over a registry. A Repair names one thing a record can lack — the predicate that says which images still owe it, the input its handler needs (a header, the original, or a native render), the handler, and what to record for an image that can never be done. The job unions the predicates into one work list, fetches each image once at the most any claimant asks for, renders it at most once, and runs every handler whose predicate that image still matches, checked again before each because a detection writes every field a per-face handler would fill. The registry today: face-proxy, face-quality, face-eyes, face-crop, face-detection, face-upgrade, metadata — the last there to say that this is not a face job. Adding a field is one entry. A repair's predicate is the only definition of its work: the count the settings page shows, the list the job fetches and the check before its handler run are one predicate, so the job converges. That is why the registry is cut to what the device can do rather than listing what it skips — an entry is a count and a set of originals to fetch — and why an eye reading that cannot be cut is not a criterion. The catalog side is generic to match: record_updates writes whichever fields a FaceUpdate carries and re-marks the image so the shards export it; faces_needing and count_needing answer a predicate the caller supplies, replacing the measuring pass's three special cases. Two buttons on the settings page run the job and differ in one predicate. "Index faces" converges on coverage: has anything examined this image. "Re-index every face" converges on provenance: face-detection claims every image with no marker under the chosen detector, in either of its forms (FaceDetector::model_ids, so a desktop in f32 and a tablet on the Hexagon do not re-index each other's work), and a marker saying a weaker one looked is not that. An original over the fetch budget is left exactly as it was under the re-index, where the sweep marks it examined: a re-detection with nothing found would delete the faces, and "cannot fetch" is not "no faces".
This commit is contained in:
+111
-14
@@ -816,10 +816,10 @@ here, and it is the phone and tablet story that should decide whether it gets bu
|
||||
anything downstream sees them. A probe's confidence (§9.1) is computed from the references it
|
||||
matched; a reference's confidence hears nothing from a probe. A face whose quality was never
|
||||
recorded is admitted to the gallery — a rule that cannot be checked admits rather than excludes —
|
||||
and the next indexing pass **measures** it: `faces_unmeasured` lists every image holding one, and
|
||||
each such face is embedded again from the native render with the landmarks it already has, the raw
|
||||
vector written over the old one and its id, box and identity untouched
|
||||
(`faces::record_measurements`). No detector runs and no suggestion is lost — the cost is the
|
||||
and the next indexing pass **measures** it: the `face-quality` repair (§18.1) lists every image
|
||||
holding one, and each such face is embedded again from the native render with the landmarks it
|
||||
already has, the raw vector written over the old one and its id, box and identity untouched
|
||||
(`faces::record_updates`). No detector runs and no suggestion is lost — the cost is the
|
||||
original fetched once more, since the length exists only at the moment of embedding.
|
||||
|
||||
**The algorithm.** Constrained average-link agglomeration over the probability graph, merging while
|
||||
@@ -1119,8 +1119,8 @@ variant; its vectors are one space, and the detector only decides where the boxe
|
||||
So every reader now keys on the embedder half of the id (`faces::embedder_of`, and
|
||||
`embedder_sql` for the queries): all three detectors are one population, and changing between
|
||||
them empties nothing. `record_detections` is unchanged — an image holds one pipeline's faces at a
|
||||
time, and a re-detection carries confirmations across by box overlap — and it is where the
|
||||
generations meet. The merge's `match_faces` matches within an embedder for the same reason. The
|
||||
time, and a re-detection carries identities across by box overlap and embedding (§18) — and it is
|
||||
where the generations meet. The merge's `match_faces` matches within an embedder for the same reason. The
|
||||
shards travel every generation, each under its own id, and a peer adopts whichever it is sent.
|
||||
What a stronger choice still does is queue the images a weaker detector indexed for re-detection
|
||||
(`FaceDetector::supersedes`), after the ones nothing has indexed and never downwards, so a tablet
|
||||
@@ -1257,9 +1257,9 @@ frames to family snapshots as the real unknowns.
|
||||
|
||||
| ID | How this document addresses it |
|
||||
|---|---|
|
||||
| FR-CULL-8 | §4 detection, §6 embedding, §7 the proxy tier and its consequences, §10 the `DetectFaces` job |
|
||||
| FR-CULL-8 | §4 detection, §6 embedding, §7 the proxy tier and its consequences, §10 the `DetectFaces` job, §18 the re-index |
|
||||
| FR-CULL-9 | §8 — pairs, fit, validity, and the reliability-diagram acceptance test |
|
||||
| FR-CULL-10 | §9 constrained agglomeration, confirmations as anchors, split by re-agglomeration |
|
||||
| FR-CULL-10 | §9 constrained agglomeration, confirmations as anchors, split by re-agglomeration; §18 what a re-index carries across |
|
||||
| FR-CULL-11 | §10 the `Person` selector term, confirmed-only by default |
|
||||
| FR-CULL-12 | §10 schema unchanged from catalog.md §10.1: embeddings derived, names to the sidecar |
|
||||
| NFR-SEC-5 | §11 — the obligations restated as structural properties, one of them CI-checkable |
|
||||
@@ -1447,12 +1447,12 @@ been the same size and worse: three figures near 1.0 is six pixels at that scale
|
||||
|
||||
### 17.5 The measuring pass, and shards
|
||||
|
||||
A face indexed before the models existed, or on a device without them, has no reading. The sweep's
|
||||
measuring pass — the one V14 built to re-embed faces stored as unit vectors — lists those faces
|
||||
too, on a device that has the models, and reads their eyes from the same native render with the
|
||||
box and landmarks already stored (`dr_ui::faces::measure_native`). No detector runs and no identity
|
||||
moves. A device *without* the models does not list them, or it would fetch every original in the
|
||||
library to do nothing to it; `dr_catalog::faces::unmeasured_sql` is the one predicate both the
|
||||
A face indexed before the models existed, or on a device without them, has no reading. The
|
||||
`face-eyes` repair (§18.1; on the day this was written, the sweep's measuring pass — the one V14
|
||||
built to re-embed faces stored as unit vectors) lists those faces, on a device that has the models,
|
||||
and reads their eyes from the same native render with the box and landmarks already stored. No
|
||||
detector runs and no identity moves. A device *without* the models has no such repair, or it would
|
||||
fetch every original in the library to do nothing to it; the repair's predicate is the one both the
|
||||
count and the work list use, so the pass converges. The People screen's coverage line counts
|
||||
these faces as work to read and keeps the button while any remain — reading **Read eye state**
|
||||
once detection is complete and only readings are left, which is the state an already-indexed
|
||||
@@ -1470,3 +1470,100 @@ the only device that has done the detection at all.
|
||||
| **M11** | Open-eye recall and blink precision from the **native** pass with the shipped configuration, on a labelled sample that contains real blinks — a burst with one in it is enough | §17.3's figures are proxy figures from 25 faces, with no precision beside them; this is the number FR-CULL-13's acceptance clause asks for, and where the two readability floors get set on more than 25 faces |
|
||||
| **M12** | Whether the remaining open-eye failure — a lens reflection over an open eye — moves with the sharpness floor, or needs the eye classifier told about spectacles | If the latter, the fix is a classifier trained closer to this domain, and that is a different decision |
|
||||
| **M13** | Sunglasses recall on more than twelve faces, and the false-positive rate on caps and clear glasses | The two sunglasses the head classifier missed on the sample became false blinks; a library of skiers would say whether that is two faces or a class |
|
||||
|
||||
---
|
||||
|
||||
## 18. The completeness job, and what a re-index carries across · 2026-09-19
|
||||
|
||||
The reference library on the day this was written: 17,762 faces under the bare `w600k_mbf` id —
|
||||
found by the fast detector on 1024 px proxies, stored as unit vectors, no quality, no crop on 4,144
|
||||
of them, no eye reading, no dense landmarks — beside 1,177 under `scrfd_10g+w600k_mbf` from the
|
||||
native pass, and 12,217 images the fast detector examined and found nothing in. Over those faces:
|
||||
3,778 confirmations, 13,011 suggestions, 77 rejections, and 17,276 people rows. Every one of those
|
||||
gaps was, until now, its own pass: V14's measuring pass for the quality, §17.5's for the eyes, the
|
||||
sweep's proxy repair, the sweep's detector upgrade, and a re-index that did not exist. Adding a
|
||||
per-face field meant adding a pass, with its own work list, its own count and its own idea of done.
|
||||
|
||||
### 18.1 One job over a registry
|
||||
|
||||
`dr_ui::repairs` replaces them with one job over a **registry**. A `Repair` names one thing a catalog
|
||||
record can lack — the predicate that says which images still owe it, the input its handler needs
|
||||
(the file's header, the whole original, or a native render), the handler that fills it, and, where
|
||||
there is one, what to record for an image that can never be done. The job unions the predicates
|
||||
into one work list, fetches each image once at the most any claimant asks for, renders it at most
|
||||
once, and runs every handler whose predicate that image still matches — checked again before each,
|
||||
because one handler's write satisfies the next's (a detection writes every field a per-face handler
|
||||
would fill). The registry today:
|
||||
|
||||
| Repair | Owed by | Input | Handler |
|
||||
|---|---|---|---|
|
||||
| `face-proxy` (sweep only) | images holding faces whose 1024 px proxy is not in the store | native render | detect again, write the proxy |
|
||||
| `face-quality` | faces with `quality IS NULL` | native render | warp from the stored landmarks, embed, write the raw vector and its length (reads eyes on the same warp where it can) |
|
||||
| `face-eyes` | faces with no eye reading or no dense landmarks, on a device with the eye models | native render | read the eyes and dense landmarks from the stored box and landmarks |
|
||||
| `face-crop` | faces with `crop IS NULL` | native render | cut the crop from the frame |
|
||||
| `face-detection` | sweep: images with no marker under the embedder and no faces; re-index: images with no marker under the **chosen detector**, either spelling | native render | detect, embed, replace the faces, carry identities across (§18.2) |
|
||||
| `face-upgrade` (sweep only) | images whose marker is a weaker detector's | native render | as `face-detection` |
|
||||
| `metadata` | `images.metadata_state < 2` | header | EXIF to the catalog, the dateless marked examined |
|
||||
|
||||
A repair's predicate is the *only* definition of its work. The count the settings page shows
|
||||
(`faces::audit`, per repair), the list the job fetches and the check before each handler are one
|
||||
predicate, so a record the count reports is one the job fetches and one the handler fills, and the
|
||||
job ends. This is why an eye reading that cannot be cut is not a criterion — such a face stays
|
||||
unread however often it is detected, and listing it would fetch its original on every press — and
|
||||
why a degenerate face is dropped rather than left. It is also why the registry is cut to what the
|
||||
device can do (`Capabilities`): a device without the eye models has no `face-eyes` entry, rather
|
||||
than an entry it skips, because an entry is a count and a set of originals to fetch.
|
||||
|
||||
The registry is ordered, and the order is the work list's: an image only the last repair claims
|
||||
comes after one the first does, which is what puts a few hundred proxy repairs ahead of twenty
|
||||
thousand un-indexed images. Within one image the same order runs the handlers, detection before the
|
||||
per-face repairs, since detection fills what they would.
|
||||
|
||||
Adding a field is one entry: a predicate over `faces f` or `images i`, and a handler that fills it
|
||||
from `Fetched`. `metadata` is in the table to say that this is not a face job — the same machinery
|
||||
carries a capture date, and could carry a thumbnail, a perceptual hash or a head pose.
|
||||
|
||||
### 18.1a The two scopes
|
||||
|
||||
Both buttons on the settings page run the job; they differ in one predicate. **Index faces**
|
||||
(`Scope::Outstanding`) converges on *coverage* — has anything examined this image — and treats a
|
||||
face a weaker detector found on a proxy as found, which is the right question for a pass that must
|
||||
not fetch the library twice. **Re-index every face** (`Scope::Reindex`) converges on *provenance*:
|
||||
`face-detection` claims every image with no marker under the chosen detector, in either of its
|
||||
forms (`FaceDetector::model_ids`, so a desktop running it in f32 and a tablet on the Hexagon in int8
|
||||
do not re-index each other's work), and a marker saying a weaker one looked is not that. This is
|
||||
the one place in the subsystem keyed on the exact detector rather than the embedder. Convergent all
|
||||
the same: an image the job has been through leaves the list, a kill costs the images in flight, and
|
||||
a second press resumes.
|
||||
|
||||
An original over the fetch budget is skipped without being fetched. Under the sweep, detection
|
||||
records an examination that found nothing — the honest record for an image that cannot be
|
||||
examined, and what stops the half gigabyte being spent once per sweep. Under the re-index, and
|
||||
under every repair over records that already exist, it is left exactly as it was: a re-detection
|
||||
with nothing found would delete the faces, and "cannot fetch" is not "no faces".
|
||||
|
||||
### 18.2 What is carried across
|
||||
|
||||
`dr_catalog::faces::record_detections` replaces an image's faces and carries identities onto the
|
||||
new ones. Before this section it carried confirmations only, by box overlap above 0.5 IoU, and a
|
||||
re-detection of the library above would have left 13,011 suggestions and 77 rejections on the
|
||||
floor — correct by FR-CULL-12's letter, since suggestions are derived data, and a People screen
|
||||
emptied to strangers by the user's own button.
|
||||
|
||||
Now every old face is read before the delete — box, vector, assignment, rejections — and matched to
|
||||
the new faces one-to-one, best pair first. A pair qualifies when the boxes **overlap at all** and
|
||||
either the overlap alone says so (IoU above 0.5, the old rule) or the embeddings do (cosine above
|
||||
`SAME_FACE_COSINE` = 0.45, the reference library's P≈0.95 line from §9's table). The embedding route
|
||||
is for the box a low-resolution pass drew badly enough that overlap alone would not claim it; the
|
||||
vector is also what breaks the tie in a group photograph, where two neighbouring faces overlap both
|
||||
new boxes. Overlap is required on both routes because the same vector elsewhere in the frame — a
|
||||
mirror, a print on the wall — is not the same face and must not take its name. Onto the matched
|
||||
face go the assignment as it was, confirmed or suggested with its probability, and every
|
||||
rejection.
|
||||
|
||||
It is a match, not an update in place, and that is why the per-face repairs exist beside
|
||||
detection: where nothing about a face but one field needs doing, `record_updates` keeps the id and
|
||||
there is nothing to judge.
|
||||
|
||||
The merge's `match_faces` still matches by overlap alone across devices. It is the same question,
|
||||
and the same answer would serve it; it is not changed here.
|
||||
|
||||
Reference in New Issue
Block a user