Score a suggestion against the people the user has named
The number beside a suggestion was the mean calibrated probability between the face and the rest of its group, which measures the wrong thing twice. It punishes coverage: a person with two hundred faces over fifteen years is *meant* to have members a given photograph is orthogonal to, so a correct suggestion onto a well-photographed person scored low for being well photographed. And it never asked who else the face might be — a face matching Anna at 0.95 and nobody else, and one matching Anna at 0.95 and her sister at 0.93, came out identical, when the second is the only one worth the user's attention. dr_face::assign answers both, and multiplies them: the mean of the best ten calibrated matches into the identity (the old mean, capped, which is what stops coverage counting against it), times that identity's share of the evidence against every *named* rival. Only named people compete, and per person rather than per group. Both halves of that had to be measured on a real 18,000-face library rather than reasoned about. Normalising across every group made the number useless — median suggestion 21%, four in five under half — because clustering leaves one person spread over many groups, so a face competed against itself; and keying rivals by group left Catherine competing with Catherine, median 39%. Per named person: median 99.5%. Rivals are gathered below the merge threshold, down to even odds: a named person matching at 0.6 will never be merged into but is exactly the competition to discount for. That would be a second similarity scan, the expensive half of regrouping a library, so cluster_scored scans once at the looser floor and hands the merge engine the subset at or above the threshold — pair for pair what it would have scanned for itself, held to that by a test. Leave-one-out over that library's 2,702 confirmations across 54 named people: 99.33% of faces placed on the right person against the old mean's 99.15%, and the number shown for the right person moves from a median of 90.4% to 99.3%. It errs low — 100% correct wherever it states 80% or more — which is the safe direction, and docs/faces.md §9.1 says plainly that the low bands are not calibrated. The example that measures it comes too: this is a claim about a library's numbers, and nobody should have to take it on faith. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -228,6 +228,7 @@ core/dr-face/
|
||||
src/embed.rs MBF: preprocess, forward, L2 normalise
|
||||
src/calibrate.rs cosine → P(same person) (FR-CULL-9)
|
||||
src/cluster.rs constrained agglomeration (FR-CULL-10)
|
||||
src/assign.rs which person, and how sure (§9.1, FR-CULL-9)
|
||||
```
|
||||
|
||||
```toml
|
||||
@@ -675,6 +676,64 @@ are recomputed freely; confirmations survive all of it (FR-CULL-10).
|
||||
as the split. FR-CULL-10 requires splitting to be as easy as merging, and a split that hands the user
|
||||
a pile of loose faces to re-sort is not that.
|
||||
|
||||
### 9.1 The number beside a suggestion
|
||||
|
||||
Which person a face belongs to and how sure that is are **different questions**, and the second one
|
||||
is not answered by the pairwise probabilities that settled the first.
|
||||
|
||||
The first implementation answered it with the mean calibrated probability between the face and the
|
||||
rest of its group, and that measures the wrong thing twice. It punishes coverage: a person with two
|
||||
hundred faces across fifteen years is *supposed* to have members a new photograph is orthogonal to,
|
||||
so the better someone is photographed the worse their suggestions score. And it never asks who else
|
||||
the face could be — a face matching Anna at 0.95 and nobody else, and one matching Anna at 0.95 and
|
||||
her sister at 0.93, come out identical, when the second is the only one the user needs to look at.
|
||||
|
||||
Two questions, so two factors, multiplied:
|
||||
|
||||
```
|
||||
evidence(P) = Σ of the top n of { P(same | this face, f) : f ∈ P } n = 10
|
||||
coherence = evidence(own) / how many of the top n there were
|
||||
uniqueness = evidence(own) / (evidence(own) + Σ evidence(named rivals))
|
||||
confidence = coherence × uniqueness
|
||||
```
|
||||
|
||||
**Coherence** is the old mean with a cap on it, and the cap is the whole fix: the two hundred faces a
|
||||
given photograph is legitimately orthogonal to stop counting against it. **Uniqueness** is the
|
||||
competition, and it is what makes an ambiguous face read as ambiguous — two identities matching
|
||||
equally well land at 0.5 each, which is the truth about a sibling.
|
||||
|
||||
**Only named people compete, and they compete per person.** This is the part that had to be measured
|
||||
rather than reasoned about. Normalising across *every* group made the number useless on a real
|
||||
18,000-face library — median suggestion 21%, four in five under half — because clustering leaves one
|
||||
person spread across many groups, so a face competes against itself. Counting only groups holding a
|
||||
confirmation fixed most of it; counting them **per person** rather than per group fixed the rest,
|
||||
since a named person is left in several anchored groups for the same reason.
|
||||
|
||||
**Rivals are gathered below the merge threshold**, down to even odds: a named person who matches at
|
||||
0.6 will never be merged into but is exactly the competition a suggestion should be discounted for.
|
||||
The floor matters in both directions — summing the near-orthogonal pairs instead of dropping them
|
||||
lets fifty identities' worth of upper-tail noise outweigh one real match, which on the same library
|
||||
moved the median stated confidence from 100% to 31%.
|
||||
|
||||
*Measured (`cargo run --release -p dr-catalog --example face_confidence`), leave-one-out over that
|
||||
library's 2,702 confirmations across 54 named people:*
|
||||
|
||||
| | share | old mean |
|
||||
|---|---|---|
|
||||
| right person picked | **99.33%** | 99.15% |
|
||||
| stated for the right person, median | **99.3%** | 90.4% |
|
||||
| stated for the right person, p10 | **79.2%** | 68.0% |
|
||||
|
||||
The reliability table is monotone and **errs low**: 100% correct wherever it states 80% or more, 84%
|
||||
correct where it states under half. Understating is the safe direction for a screen whose purpose is
|
||||
deciding what to look at first, but the low bands are not calibrated and should not be read as
|
||||
though they were — and the leave-one-out task asks *which of these people*, never *is it any of
|
||||
them*, so it cannot speak to a stranger at all.
|
||||
|
||||
It is **not** a merge threshold and must not become one. Uniqueness is relative, so a library with
|
||||
one named person would hand every stray face a 1. "Is this the same person at all" stays §8's
|
||||
question, and coherence is the half of the product that carries it.
|
||||
|
||||
---
|
||||
|
||||
## 10. Catalog and jobs
|
||||
|
||||
@@ -9,7 +9,7 @@ Denominators are parsed from [`requirements.md`](requirements.md) at run time, n
|
||||
|
||||
| Metric | Value |
|
||||
|---|---|
|
||||
| Source files scanned | 279 |
|
||||
| Source files scanned | 280 |
|
||||
| TRACES tags found | 812 |
|
||||
| Requirements defined | 177 |
|
||||
| Requirements covered | 106 |
|
||||
|
||||
Reference in New Issue
Block a user