Keep each face's quality, and never compare against a poor one
The embedder's raw output has a length, and the length is a reading of how recognisable the crop was: a blur, an occlusion or a hard profile comes out short. Normalising threw it away. A short vector sits near the middle of the sphere and matches a little of everyone, which is how one bad crop bridges two people in a grouping pass. So the length is kept — the store now holds the raw vector, re-normalised on load, with the length beside it as `faces.quality` — and a face under MIN_GALLERY_QUALITY (14) is a probe: measured against the gallery and placed where it fits, but never what another face is measured against. Two probes are never paired, and a probe is nobody's evidence for a confidence. The People screen shows the number as "Quality 17.3", dimmed below the floor. Faces indexed before this stored unit vectors and have no reading; they are admitted to the gallery, and schema V14 forgets the run marker of every image holding one so the next indexing pass measures them. A peer's unmeasured shard faces are not adopted, or a sync would write that marker back.
This commit is contained in:
+6
-1
@@ -742,7 +742,12 @@ CREATE TABLE faces (
|
||||
x REAL NOT NULL, y REAL NOT NULL, w REAL NOT NULL, h REAL NOT NULL,
|
||||
landmarks BLOB, -- 5 × (x, y) f32, the alignment input
|
||||
detector_confidence REAL NOT NULL,
|
||||
embedding BLOB NOT NULL, -- 512 × f16, L2-normalised
|
||||
embedding BLOB NOT NULL, -- 512 × f16, the raw model output; re-normalised on load
|
||||
-- Length of that vector: the model's own reading of how recognisable the
|
||||
-- crop was, and the gate on whether this face may be compared *against*
|
||||
-- (faces.md §6, §9). NULL for a face stored as a unit vector before it
|
||||
-- was kept.
|
||||
quality REAL,
|
||||
-- Which model produced this. An embedding is only comparable to others
|
||||
-- from the same model; mixing them silently yields nonsense similarities.
|
||||
model_id TEXT NOT NULL,
|
||||
|
||||
+37
-4
@@ -447,9 +447,28 @@ every measured number in §1's table was produced with `/128`. The difference is
|
||||
almost certainly immaterial, but "almost certainly" is not a reason to pick silently — write `/128` to
|
||||
match the numbers we have, and settle it with one back-to-back run in §12.
|
||||
|
||||
Output is 512 floats; **L2-normalise before storing**, so every downstream comparison is a dot product
|
||||
and no code path has to remember to normalise. The reference clamps the norm at 1e-6 before dividing,
|
||||
which costs nothing and removes a NaN path.
|
||||
Output is 512 floats. Every downstream comparison is a dot product over the **unit** vector, and the
|
||||
reference L2-normalises before storing so that no code path has to remember to. The reference clamps
|
||||
the norm at 1e-6 before dividing, which costs nothing and removes a NaN path.
|
||||
|
||||
**Keep the length.** The norm the normalisation divides out is not noise. ArcFace trains the
|
||||
direction of its output and nothing else, and the magnitude it leaves behind grows with how much of a
|
||||
face the model could make out — MagFace (Meng et al., CVPR 2021) made that the training objective,
|
||||
and the plain ArcFace heads this crate runs already show it, weaker but usable. A blur, an occlusion,
|
||||
a hard profile or a badly lit crop comes out short. On the reference library `w600k_mbf`'s norms run
|
||||
from about 8 on a blur to the high 20s on a clean portrait.
|
||||
|
||||
So the store holds the **raw** vector, not the unit one — `dr_face::Embedded::to_f16_bytes` — and
|
||||
readers re-normalise on load, which they had to do anyway (below). f16 keeps the same three figures
|
||||
of a component whatever the vector's length, so this costs nothing in precision. The length is also
|
||||
kept beside the blob as `faces.quality`, for the readers that never load the vector (the People
|
||||
screen), and it is `NULL` for a face stored as a unit vector before this — a unit vector reads as a
|
||||
length of one, and one is not "unmeasured".
|
||||
|
||||
What the number does is in §9: a face whose quality is under **`MIN_GALLERY_QUALITY` = 14** is still
|
||||
placed, but is never what another face is compared *against*. The screen shows it as "Quality 17.3",
|
||||
dimmed below the floor, so a user asking why a group did not gather the rest of a person can see that
|
||||
none of its members can vouch for anyone.
|
||||
|
||||
Some exports of these graphs are fp16 in, fp16 out. The reference detects this from the graph's
|
||||
declared element type rather than from the filename; worth porting, because the alternative failure is
|
||||
@@ -461,7 +480,8 @@ between a match and a non-match, and the halving matters because these rows are
|
||||
contemplates optionally syncing.
|
||||
|
||||
Re-normalise on load after the f16 widen. It is one pass over 512 floats and it removes a class of
|
||||
drift that is otherwise invisible.
|
||||
drift that is otherwise invisible — and, since the blob is raw, it is what turns the stored vector
|
||||
back into the unit one every comparison expects.
|
||||
|
||||
---
|
||||
|
||||
@@ -785,6 +805,19 @@ here, and it is the phone and tablet story that should decide whether it gets bu
|
||||
- **Confirmed faces are anchors.** A confirmation is user data (FR-CULL-12) and clustering never
|
||||
moves it. Two clusters each containing confirmations of *different* people cannot merge; a cluster
|
||||
containing confirmations of one person absorbs suggestions but never reassigns the confirmed.
|
||||
- **A short embedding is never a reference.** The length of the raw vector is the model's own
|
||||
reading of the crop (§6), and a short one sits near the middle of the sphere, matching a little of
|
||||
everybody — one of those in a group is a bridge to the next group over. So the population is
|
||||
split: faces at or above `MIN_GALLERY_QUALITY` are the **gallery** and cluster as described below;
|
||||
faces under it are **probes**, each measured against the finished groups and placed in the one it
|
||||
fits by the same average-link rule under the same two constraints — but measured against gallery
|
||||
members only, never against another probe, and once placed never part of what the next face is
|
||||
measured against. Two probes are never paired at all, and `neighbours` drops those pairs before
|
||||
anything downstream sees them. A probe's confidence (§9.1) is computed from the references it
|
||||
matched; a reference's confidence hears nothing from a probe. A face whose quality was never
|
||||
recorded is admitted to the gallery — a rule that cannot be checked admits rather than excludes —
|
||||
and schema V14 forgets the run marker of every image holding one, so the next indexing pass
|
||||
measures it.
|
||||
|
||||
**The algorithm.** Constrained average-link agglomeration over the probability graph, merging while
|
||||
the average pairwise probability exceeds **0.9** and no cannot-link is violated. Average-link rather
|
||||
|
||||
+4
-4
@@ -197,7 +197,7 @@ Double-click is what a file manager and a Lightroom panel use for the same thing
|
||||
|
||||
Grouping over-merges on siblings, on parents and children, and on the same person a decade apart, so splitting is as prominent as merging. A tool that can only merge makes its own errors permanent.
|
||||
|
||||
<sub>`ui/dr-ui/ui/identity.slint:130`</sub>
|
||||
<sub>`ui/dr-ui/ui/identity.slint:161`</sub>
|
||||
|
||||
### Rule on a suggested face
|
||||
|
||||
@@ -206,7 +206,7 @@ Grouping over-merges on siblings, on parents and children, and on the same perso
|
||||
|
||||
A face is either the system's guess or the user's judgement, and the two are never conflated. A rejection is remembered, so the face is not suggested for that person again. The gesture note above is the whole label: a tick and a cross are only "confirm" and "reject" to someone who can see the suggestion they sit beside, and `IconButton`'s fallback would announce them as "check" and "cross" — two icon names that say nothing about which person is being ruled on.
|
||||
|
||||
<sub>`ui/dr-ui/ui/identity.slint:150`</sub>
|
||||
<sub>`ui/dr-ui/ui/identity.slint:181`</sub>
|
||||
|
||||
### See a person's photographs
|
||||
|
||||
@@ -215,7 +215,7 @@ A face is either the system's guess or the user's judgement, and the two are nev
|
||||
|
||||
This is the point of having identified anybody. Without it the screen is a filing cabinet with no drawer handles.
|
||||
|
||||
<sub>`ui/dr-ui/ui/identity.slint:571`</sub>
|
||||
<sub>`ui/dr-ui/ui/identity.slint:602`</sub>
|
||||
|
||||
### Change how faces are grouped
|
||||
|
||||
@@ -224,7 +224,7 @@ This is the point of having identified anybody. Without it the screen is a filin
|
||||
|
||||
The right match confidence is a property of your library, not of the model. "What would this do?" answers for this library without writing anything; names, confirmations and the groups you have set aside are kept whatever the dials say.
|
||||
|
||||
<sub>`ui/dr-ui/ui/identity.slint:608`</sub>
|
||||
<sub>`ui/dr-ui/ui/identity.slint:639`</sub>
|
||||
|
||||
## Library grid
|
||||
|
||||
|
||||
+21
-21
File diff suppressed because one or more lines are too long
Reference in New Issue
Block a user