Record that face detection has run, not just what it found

An image with no faces in it was indistinguishable from one that had
never been looked at, so every landscape, still life and document scan in
the library was re-detected on every pass, for ever. In a real library
that is most of it: on the 23,527-image test library, 64 of the first 110
images indexed contain no face at all.

Schema v9 adds face_index, a run marker per (image, model) carrying the
face count and the proxy edge it read. Keyed on the model, so a model
change puts every image back in the queue by itself.

That makes a coverage figure possible, which is the thing a user actually
wants to see. The audit also splits the outstanding set by whether a
proxy exists, because 23,417 awaiting a proxy and 110 ready to index are
different problems, and telling the user to run indexing again would not
fix the first.

The Identity screen gains Index faces, Stop, and the coverage line.
examples/face_index.rs is the same check and sweep without a window,
which is the right shape for an overnight pass.

Measured on the real library in release: 3.5 images/second, 110 images
and 125 faces in 30 seconds, and a second run correctly finds nothing
left to do.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-26 22:13:41 +02:00
co-authored by Claude Opus 5
parent 3e607222c6
commit 26a1eb7e28
14 changed files with 971 additions and 84 deletions
+12 -8
View File
@@ -726,14 +726,18 @@ The duplicate row is not a curiosity: several files in that gallery are byte-ide
different names, which is exactly the case §8.2's dedup step exists for, and it would otherwise stack
the positive histogram at cos ≈ 1 with pairs that teach the fit nothing.
**M2/M3, provisionally, on the reference desktop:** detection **~1.0–1.4 s** per image, embedding
**~160–280 ms** per face, model load 80–90 ms for the pair. Slower than the extrapolation in §12's M2
row, which guessed SCRFD-500M would come in under YOLO26n-seg's 470 ms — tract is not ORT, and this
is what it costs. A 17k-image library at ~1.3 s plus ~1.5 faces each is on the order of **7 hours** of
background indexing. That is survivable for a resumable, preempted background job (FR-CULL-8) and it
is not survivable on a phone, so NFR-RES-2 needs its own measurement rather than an extrapolation
from this one. Not yet profiled: how much of the detection second is tract and how much is the
scalar letterbox resample in front of it.
**M2/M3 in a debug build:** detection ~1.0–1.4 s per image, embedding ~160–280 ms per face.
**M2/M3 in release, over a real library — the number that counts: 3.5 images/second**, end to end,
including the JPEG decode and the catalog write. 110 images with 125 faces in 30 seconds on the
reference desktop. That is roughly **4× the debug figure**, and it moves a 23.5k-image library from
the "seven hours" the debug numbers implied to about **110 minutes**.
Worth stating plainly because the debug measurement was nearly a wrong conclusion: it was on the
edge of making the pure-Rust runtime look unaffordable for a large library, and it was measuring the
profile rather than the pipeline. Any future timing of this subsystem is a release timing.
Still unmeasured: the same pass on a phone (NFR-RES-2), which does not follow from this one.
---