Keep each face's quality, and never compare against a poor one

The embedder's raw output has a length, and the length is a reading of
how recognisable the crop was: a blur, an occlusion or a hard profile
comes out short. Normalising threw it away. A short vector sits near
the middle of the sphere and matches a little of everyone, which is how
one bad crop bridges two people in a grouping pass.

So the length is kept — the store now holds the raw vector, re-normalised
on load, with the length beside it as `faces.quality` — and a face under
MIN_GALLERY_QUALITY (14) is a probe: measured against the gallery and
placed where it fits, but never what another face is measured against.
Two probes are never paired, and a probe is nobody's evidence for a
confidence. The People screen shows the number as "Quality 17.3", dimmed
below the floor.

Faces indexed before this stored unit vectors and have no reading; they
are admitted to the gallery, and schema V14 forgets the run marker of
every image holding one so the next indexing pass measures them. A
peer's unmeasured shard faces are not adopted, or a sync would write
that marker back.
This commit is contained in:
2026-09-11 21:50:12 +02:00
parent a87139b838
commit 8b3abdb787
22 changed files with 1070 additions and 149 deletions
+37 -4
View File
@@ -447,9 +447,28 @@ every measured number in §1's table was produced with `/128`. The difference is
almost certainly immaterial, but "almost certainly" is not a reason to pick silently — write `/128` to
match the numbers we have, and settle it with one back-to-back run in §12.
Output is 512 floats; **L2-normalise before storing**, so every downstream comparison is a dot product
and no code path has to remember to normalise. The reference clamps the norm at 1e-6 before dividing,
which costs nothing and removes a NaN path.
Output is 512 floats. Every downstream comparison is a dot product over the **unit** vector, and the
reference L2-normalises before storing so that no code path has to remember to. The reference clamps
the norm at 1e-6 before dividing, which costs nothing and removes a NaN path.
**Keep the length.** The norm the normalisation divides out is not noise. ArcFace trains the
direction of its output and nothing else, and the magnitude it leaves behind grows with how much of a
face the model could make out — MagFace (Meng et al., CVPR 2021) made that the training objective,
and the plain ArcFace heads this crate runs already show it, weaker but usable. A blur, an occlusion,
a hard profile or a badly lit crop comes out short. On the reference library `w600k_mbf`'s norms run
from about 8 on a blur to the high 20s on a clean portrait.
So the store holds the **raw** vector, not the unit one — `dr_face::Embedded::to_f16_bytes` — and
readers re-normalise on load, which they had to do anyway (below). f16 keeps the same three figures
of a component whatever the vector's length, so this costs nothing in precision. The length is also
kept beside the blob as `faces.quality`, for the readers that never load the vector (the People
screen), and it is `NULL` for a face stored as a unit vector before this — a unit vector reads as a
length of one, and one is not "unmeasured".
What the number does is in §9: a face whose quality is under **`MIN_GALLERY_QUALITY` = 14** is still
placed, but is never what another face is compared *against*. The screen shows it as "Quality 17.3",
dimmed below the floor, so a user asking why a group did not gather the rest of a person can see that
none of its members can vouch for anyone.
Some exports of these graphs are fp16 in, fp16 out. The reference detects this from the graph's
declared element type rather than from the filename; worth porting, because the alternative failure is
@@ -461,7 +480,8 @@ between a match and a non-match, and the halving matters because these rows are
contemplates optionally syncing.
Re-normalise on load after the f16 widen. It is one pass over 512 floats and it removes a class of
drift that is otherwise invisible.
drift that is otherwise invisible — and, since the blob is raw, it is what turns the stored vector
back into the unit one every comparison expects.
---
@@ -785,6 +805,19 @@ here, and it is the phone and tablet story that should decide whether it gets bu
- **Confirmed faces are anchors.** A confirmation is user data (FR-CULL-12) and clustering never
moves it. Two clusters each containing confirmations of *different* people cannot merge; a cluster
containing confirmations of one person absorbs suggestions but never reassigns the confirmed.
- **A short embedding is never a reference.** The length of the raw vector is the model's own
reading of the crop (§6), and a short one sits near the middle of the sphere, matching a little of
everybody — one of those in a group is a bridge to the next group over. So the population is
split: faces at or above `MIN_GALLERY_QUALITY` are the **gallery** and cluster as described below;
faces under it are **probes**, each measured against the finished groups and placed in the one it
fits by the same average-link rule under the same two constraints — but measured against gallery
members only, never against another probe, and once placed never part of what the next face is
measured against. Two probes are never paired at all, and `neighbours` drops those pairs before
anything downstream sees them. A probe's confidence (§9.1) is computed from the references it
matched; a reference's confidence hears nothing from a probe. A face whose quality was never
recorded is admitted to the gallery — a rule that cannot be checked admits rather than excludes —
and schema V14 forgets the run marker of every image holding one, so the next indexing pass
measures it.
**The algorithm.** Constrained average-link agglomeration over the probability graph, merging while
the average pairwise probability exceeds **0.9** and no cannot-link is violated. Average-link rather