Fuse every detector's faces into one population per embedder
Choosing "Thorough" made the library look empty. The detector setting writes under its own faces.model_id, and every reader of "the faces" keyed on that exact id: the clustering pass, the coverage figure, the sweep's work list, the shard export and import, and the sync merge's face matching. On the reference library that restarted coverage at 1,834 of 19,140, drew a People rail of 36 faces for a person with 520, queued a ~400 GB re-fetch on each device, and stranded the desktop's 3,583 confirmations under the old id: the tablet held the same faces under the new one and the merge refused to match them. Same photograph, same box, same embedder, two ids — that is one face, not two libraries. The embedder half of the id is now the key. embedder_of and embedder_sql give it to every query; writes keep the full id, so which detector drew a box stays on record. record_detections is unchanged and is where the generations meet: an image holds one pipeline's faces at a time, and a re-detection carries confirmations across by box overlap. The merge's match_faces applies the same rule within an embedder. The calibration is keyed on the embedder too, since the similarity space did not change. Shards travel every generation, each under its own id, and a peer adopts whichever it is sent — including a stronger detector's pass over an image it indexed itself with a weaker one, which is the re-detection its own sweep would otherwise queue, already done. Never downwards: a tablet on Fast keeps the desktop's Thorough faces. The sweep gains the same tail — images a weaker detector indexed, after the ones nothing has — driven by FaceDetector::supersedes, so choosing a stronger detector still improves the library over time without first making it disappear.
This commit is contained in:
+29
-10
@@ -540,8 +540,8 @@ images indexed, **64 contain no face at all**.
|
||||
|
||||
So `face_index` records the *run*: one row per `(image, model)` carrying the timestamp, the number of
|
||||
faces found — zero is the interesting value — and the long edge of the proxy it read. Keyed on the
|
||||
model, so a model change puts every image back in the queue without anyone having to remember to
|
||||
clear anything.
|
||||
model, so an *embedder* change puts every image back in the queue without anyone having to remember
|
||||
to clear anything; a detector change in front of the same embedder does not (§12.3).
|
||||
|
||||
Three things fall out of it that were not otherwise available:
|
||||
|
||||
@@ -1100,14 +1100,33 @@ is not a desktop-only feature (NFR-RES-2). It loads in tract with the same fix a
|
||||
decodes through the same nine-output path unchanged. Same licence, same `buffalo_m` release page.
|
||||
|
||||
**Which detector runs is a setting** — `FaceSettings::detector`, per device, on the settings page
|
||||
beside the indexing button as Fast / Balanced / Thorough. All three files ship. Each detector is
|
||||
its own `faces.model_id` (`w600k_mbf` for `500M`, unchanged, so nothing already indexed is
|
||||
disturbed; `scrfd_2.5g+w600k_mbf` and `scrfd_10g+w600k_mbf` for the others), which is the
|
||||
mechanism §2.1 always intended for a model change: the coverage figure restarts at zero under the
|
||||
new id, the sweep re-detects, `record_detections` carries confirmed names across by box overlap and
|
||||
drops the previous pipeline's marker for each image it revisits, and the sync shards are keyed by
|
||||
the same id so a peer on another setting neither adopts nor pollutes them. The default stays `500M`
|
||||
so that an upgrade changes nothing until the user chooses; the recommendation is `2.5G`.
|
||||
beside the indexing button as Fast / Balanced / Thorough. All three files ship. Each detector
|
||||
writes its own `faces.model_id` (`w600k_mbf` for `500M`, unchanged; `scrfd_2.5g+w600k_mbf` and
|
||||
`scrfd_10g+w600k_mbf` for the others), so which pipeline drew a box is always on record. The
|
||||
default stays `500M` so that an upgrade changes nothing until the user chooses; the recommendation
|
||||
is `2.5G`.
|
||||
|
||||
**One population per embedder, not one per detector · 2026-09-19.** The first cut of the setting
|
||||
keyed every reader on the full id — the clustering pass, the coverage figure, the sweep's work
|
||||
list, the shard export and import, and the sync merge's face matching — on the theory that a
|
||||
detector change is a model change. Measured on the reference library it was a disaster: choosing
|
||||
Thorough on both devices restarted coverage at 1,834 of 19,140, the People screen showed only the
|
||||
faces the new pipeline had reached, the desktop's 3,583 confirmations under the old id could not
|
||||
reach the tablet because the merge demanded the same id on both sides, and each device faced a
|
||||
~400 GB re-fetch before the library looked whole again. The embedder is `w600k_mbf` in every
|
||||
variant; its vectors are one space, and the detector only decides where the boxes are.
|
||||
|
||||
So every reader now keys on the embedder half of the id (`faces::embedder_of`, and
|
||||
`embedder_sql` for the queries): all three detectors are one population, and changing between
|
||||
them empties nothing. `record_detections` is unchanged — an image holds one pipeline's faces at a
|
||||
time, and a re-detection carries confirmations across by box overlap — and it is where the
|
||||
generations meet. The merge's `match_faces` matches within an embedder for the same reason. The
|
||||
shards travel every generation, each under its own id, and a peer adopts whichever it is sent.
|
||||
What a stronger choice still does is queue the images a weaker detector indexed for re-detection
|
||||
(`FaceDetector::supersedes`), after the ones nothing has indexed and never downwards, so a tablet
|
||||
on Fast keeps the desktop's Thorough faces rather than replacing them with fewer. The calibration
|
||||
(§8) is keyed on the embedder too: it is a fit over the similarity space, and that space did not
|
||||
change.
|
||||
|
||||
---
|
||||
|
||||
|
||||
+28
-28
File diff suppressed because one or more lines are too long
Reference in New Issue
Block a user