The embedder's raw output has a length, and the length is a reading of
how recognisable the crop was: a blur, an occlusion or a hard profile
comes out short. Normalising threw it away. A short vector sits near
the middle of the sphere and matches a little of everyone, which is how
one bad crop bridges two people in a grouping pass.
So the length is kept — the store now holds the raw vector, re-normalised
on load, with the length beside it as `faces.quality` — and a face under
MIN_GALLERY_QUALITY (14) is a probe: measured against the gallery and
placed where it fits, but never what another face is measured against.
Two probes are never paired, and a probe is nobody's evidence for a
confidence. The People screen shows the number as "Quality 17.3", dimmed
below the floor.
Faces indexed before this stored unit vectors and have no reading; they
are admitted to the gallery, and schema V14 forgets the run marker of
every image holding one so the next indexing pass measures them. A
peer's unmeasured shard faces are not adopted, or a sync would write
that marker back.
The GPU question needed a number nobody had: how a regroup divides on
the hardware whose CPU is weakest. dr-face carries no weights and
touches no display, and dr-catalog's example needs only a catalog file,
so both run under adb shell against a copy of a real library.
On the same 18,143 faces — desktop against the tablet — scan 0.96s /
2.61s, agglomerate 1.69s / 2.16s, score 0.26s / 0.40s. The scan is half
the pass on the tablet and under a third on the desktop, because twenty
cores of AVX2 pull ahead of NEON much further than the merge engine's
single-threaded hashing does. So a GPU GEMM is worth roughly 2× a
regroup on the tablet and 1.5× here, and it is the tablet that should
decide whether it is built.
The two architectures agree exactly: the same 1,531,969 evidence pairs,
the same 2,518 groups holding the same 16,246 faces, the same
reliability table. That is a better check on the NEON kernel than the
unit test can be.
Two instruments, both read-only: the example now prints its phases, and
dr-face gains scan_bench, which needs no library at all and so can
answer "how fast is this machine" on a device with nothing on it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>