assign's denominator is the identities the user has ruled on, and it reads Cluster::person to find them. Master's "Let a name hold a group together" widened what sets that field: a confirmation, a name, or an ignore, where before it was a confirmation alone. The behaviour is right either way — a named person is exactly the identity a suggestion should be discounted against — but the module note and faces.md §9.1 both said "a confirmation", which is now too narrow. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
363 lines
15 KiB
Rust
363 lines
15 KiB
Rust
//! TRACES: FR-CULL-9 | FR-CULL-10
|
||
//! How sure a *suggestion* is: an identity's share of the evidence for a face.
|
||
//!
|
||
//! [`crate::cluster`] decides which people exist; this decides what number to
|
||
//! put beside "we think this is Anna". They are not the same question, and the
|
||
//! answer to the second used to be a by-product of the first — the mean
|
||
//! calibrated probability between a face and *every* other member of its group.
|
||
//!
|
||
//! # Why a mean over the group is the wrong number
|
||
//!
|
||
//! It measures the wrong thing twice over.
|
||
//!
|
||
//! **It punishes large, well-photographed people.** Anna has two hundred faces
|
||
//! spanning fifteen years; a new photograph of her matches thirty of them
|
||
//! strongly and is near-orthogonal to the rest, because a face at 8 and a face
|
||
//! at 23 genuinely are. The mean lands around 0.2 and the interface reports a
|
||
//! correct suggestion as a doubtful one. The better a person is covered, the
|
||
//! worse their confidences get, which is exactly backwards.
|
||
//!
|
||
//! **It never asks who else it could be.** A face that matches Anna at 0.95 and
|
||
//! matches nobody else at all, and a face that matches Anna at 0.95 *and her
|
||
//! sister at 0.93*, are the same number under a within-group mean. The second
|
||
//! is the one the user actually needs to look at, and it was indistinguishable
|
||
//! from the first.
|
||
//!
|
||
//! # Coherence, times uniqueness
|
||
//!
|
||
//! Two questions, and the number is their product because they are genuinely
|
||
//! independent: *is this the same person at all*, and *of the people we know,
|
||
//! is it uniquely this one*.
|
||
//!
|
||
//! ```text
|
||
//! evidence(P) = Σ of the top n of { P(same | this face, f) : f ∈ P }
|
||
//! coherence = evidence(own) / (however many of the top n there were)
|
||
//! uniqueness = evidence(own) / (evidence(own) + Σ evidence(named rivals))
|
||
//! confidence = coherence × uniqueness
|
||
//! ```
|
||
//!
|
||
//! **Coherence** is a mean, like the old number, but over the face's best
|
||
//! [`TOP_MATCHES`] matches into the identity rather than over all of them. That single cap is what
|
||
//! stops a well-photographed person scoring worse than a thin one: the two
|
||
//! hundred faces a given photograph is legitimately orthogonal to no longer
|
||
//! count against it.
|
||
//!
|
||
//! **Uniqueness** is the competition. A sole strong match leaves it at 1 and
|
||
//! the confidence is the coherence; two identities matching equally well pull
|
||
//! it to 0.5 each, and the screen has told the user the truth, which is that
|
||
//! this face is a coin toss between two people.
|
||
//!
|
||
//! # Only the people the user has named compete
|
||
//!
|
||
//! Measured on a real 18,000-face library, normalising across *every* group
|
||
//! made the number useless: the median suggestion read 21% and four in five
|
||
//! read under half. The cause is not a bug in the arithmetic but a fact about
|
||
//! clustering — one person is spread across many groups, since the pairs that
|
||
//! would have joined them are the ones that fell short of the merge threshold.
|
||
//! Normalising over groups therefore makes a face compete against *itself*,
|
||
//! and the better covered the person, the more fragments there are to lose to.
|
||
//!
|
||
//! A fragment is not a rival. An identity the user has actually asserted is, so
|
||
//! the denominator counts only the people they have ruled on — a group carries
|
||
//! a [`Cluster::person`] when it holds a confirmation, a name, or an ignore —
|
||
//! and counts them **per person, not per group**, since one person is left in
|
||
//! several anchored groups for the same reason. Keying it by group had
|
||
//! Catherine competing with Catherine and put the median suggestion onto a
|
||
//! named person at 39%; keying it by person put it at 99.5%.
|
||
//!
|
||
//! Leave-one-out over that library's 2,702 confirmations across 54 named
|
||
//! people, this is the regime where the number is worth having: 99.3% of faces
|
||
//! are placed on the right person (the old mean managed 99.15%), the stated
|
||
//! percentage is monotone in being right, and it errs low — 100% correct
|
||
//! wherever it states 80% or more, 84% correct where it states under half.
|
||
//! Understating is the safe direction for a screen whose whole purpose is
|
||
//! deciding what to look at first, but it is *not* calibrated in the low bands
|
||
//! and should not be read as though it were.
|
||
//!
|
||
//! Two faces of one unnamed group cannot be told from two fragments of one
|
||
//! person by similarity alone; that is exactly why clustering stopped where it
|
||
//! did. So the module does not pretend to: where nobody is named, uniqueness is
|
||
//! 1 and the number falls back to plain coherence.
|
||
//!
|
||
//! # What it does not do
|
||
//!
|
||
//! It is not a merge threshold and must not become one. Clustering keeps
|
||
//! deciding on the pairwise calibrated probability: uniqueness is *relative*,
|
||
//! so a library with one named person in it would hand every stray face a
|
||
//! uniqueness of 1. The absolute question ("is this the same person at all")
|
||
//! and the comparative one ("of the people we know, which") are different, and
|
||
//! the product is what keeps both in the answer.
|
||
|
||
use std::collections::BTreeMap;
|
||
|
||
use crate::cluster::Cluster;
|
||
use crate::neighbours::Pair;
|
||
|
||
/// Who the evidence is for.
|
||
///
|
||
/// Ordered rather than hashed so the sums below are reproducible; the ordering
|
||
/// itself carries no meaning.
|
||
#[derive(Debug, Clone, Copy, PartialEq, Eq, PartialOrd, Ord)]
|
||
enum Identity {
|
||
/// A person the user has confirmed a face onto. Every group anchored to
|
||
/// them is the same identity, however many of them the clusterer left.
|
||
Person(u64),
|
||
/// A group nobody has ruled on. It stands for itself and competes with
|
||
/// nothing.
|
||
Group(usize),
|
||
}
|
||
|
||
/// How many of an identity's best matches count as its evidence.
|
||
///
|
||
/// The cap is the whole reason the sum works: uncapped, evidence would grow
|
||
/// with a person's face count and the largest group in the library would win
|
||
/// every contest. Ten is enough that a person photographed from several angles
|
||
/// contributes more than one lucky frame, and small enough that the hundred
|
||
/// mediocre matches inside a well-covered identity cannot add up to a strong
|
||
/// one. It is a starting point, not a measured optimum — M7's corpus is where
|
||
/// it would be tuned.
|
||
pub const TOP_MATCHES: usize = 10;
|
||
|
||
/// The weakest match that counts as evidence for an identity.
|
||
///
|
||
/// Rivals are half the point of this module, so the evidence scan has to reach
|
||
/// *below* the merge threshold — a named person who matches at 0.6 will never
|
||
/// be merged into but is precisely the competition a suggestion should be
|
||
/// discounted for. Even odds is the natural floor: below it a pair is more
|
||
/// likely different people than the same, and it is not a small distinction —
|
||
/// summing the near-orthogonal pairs instead of dropping them lets fifty
|
||
/// identities' worth of noise, each contributing its *upper tail*, outweigh one
|
||
/// real match. Measured on a real library, that alone moved the median stated
|
||
/// confidence from 100% to 31%.
|
||
pub const RIVAL_FLOOR: f32 = 0.5;
|
||
|
||
/// How confident each face's placement is, indexed like the face slice the
|
||
/// clusters came from.
|
||
///
|
||
/// A face in no group, or one with no evidence for anybody, scores 0.
|
||
///
|
||
/// `pairs` must be the *evidence* list — scanned at [`RIVAL_FLOOR`], not at the
|
||
/// merge threshold. Passing the merge list still works but silently removes
|
||
/// every rival weaker than a merge, which is most of them, and every uniqueness
|
||
/// collapses to 1.
|
||
pub fn identity_shares(faces: usize, clusters: &[Cluster], pairs: &[Pair], top: usize) -> Vec<f32> {
|
||
// An identity is a *person*, not a group. One person routinely holds
|
||
// several anchored groups — the same reason they hold several unnamed ones
|
||
// — and keying this by group had Catherine competing with Catherine, which
|
||
// on the library it was measured against put the median suggestion onto a
|
||
// named person at 39%.
|
||
let key_of: Vec<Identity> = clusters
|
||
.iter()
|
||
.enumerate()
|
||
.map(|(g, c)| match c.person {
|
||
Some(p) => Identity::Person(p),
|
||
None => Identity::Group(g),
|
||
})
|
||
.collect();
|
||
|
||
let mut group_of = vec![usize::MAX; faces];
|
||
for (g, c) in clusters.iter().enumerate() {
|
||
for &m in &c.members {
|
||
if m < faces {
|
||
group_of[m] = g;
|
||
}
|
||
}
|
||
}
|
||
|
||
// Ordered, not hashed: the numbers are sums of floats over these buckets
|
||
// and this module inherits [`crate::cluster`]'s promise that the same input
|
||
// yields the same output, bit for bit.
|
||
let mut evidence: Vec<BTreeMap<Identity, Vec<f32>>> = vec![BTreeMap::new(); faces];
|
||
for p in pairs {
|
||
if p.i >= faces || p.j >= faces {
|
||
continue;
|
||
}
|
||
// A pair is evidence in both directions: j's identity hears about i,
|
||
// and i's identity hears about j. The pair list holds each unordered
|
||
// pair once, so both have to be recorded here.
|
||
let (gi, gj) = (group_of[p.i], group_of[p.j]);
|
||
if gj != usize::MAX {
|
||
evidence[p.i]
|
||
.entry(key_of[gj])
|
||
.or_default()
|
||
.push(p.probability);
|
||
}
|
||
if gi != usize::MAX {
|
||
evidence[p.j]
|
||
.entry(key_of[gi])
|
||
.or_default()
|
||
.push(p.probability);
|
||
}
|
||
}
|
||
|
||
let mut out = vec![0.0; faces];
|
||
for (i, buckets) in evidence.iter_mut().enumerate() {
|
||
let mine = group_of[i];
|
||
if mine == usize::MAX {
|
||
continue;
|
||
}
|
||
let mine = key_of[mine];
|
||
let mut coherence = 0.0;
|
||
let mut ours = 0.0;
|
||
let mut rivals = 0.0;
|
||
for (&who, probabilities) in buckets.iter_mut() {
|
||
// Descending, and the ties broken by nothing: equal probabilities
|
||
// sum the same whichever order they land in.
|
||
probabilities.sort_by(|a, b| b.total_cmp(a));
|
||
let counted = probabilities.len().min(top);
|
||
let score: f32 = probabilities.iter().take(top).sum();
|
||
if who == mine {
|
||
ours = score;
|
||
coherence = score / counted as f32;
|
||
} else if matches!(who, Identity::Person(_)) {
|
||
// Only an identity the user has ruled on competes. A group
|
||
// nobody has ruled on and that matches this face is far more
|
||
// likely to be another fragment of the same person than a
|
||
// different one — see the module note, and the library it was
|
||
// measured on.
|
||
rivals += score;
|
||
}
|
||
}
|
||
let total = ours + rivals;
|
||
// No evidence at all: a face anchored into a group it has no measured
|
||
// similarity to. Nothing honest to report, so nothing is claimed.
|
||
out[i] = if total > 0.0 {
|
||
coherence * (ours / total)
|
||
} else {
|
||
0.0
|
||
};
|
||
}
|
||
out
|
||
}
|
||
|
||
#[cfg(test)]
|
||
mod tests {
|
||
use super::*;
|
||
|
||
/// An unnamed group: nobody has ruled on it, so it competes with nothing.
|
||
fn cluster(members: &[usize]) -> Cluster {
|
||
Cluster {
|
||
members: members.to_vec(),
|
||
person: None,
|
||
}
|
||
}
|
||
|
||
/// A group the user has confirmed a face onto — an identity, and therefore
|
||
/// a rival.
|
||
fn named(members: &[usize], person: u64) -> Cluster {
|
||
Cluster {
|
||
members: members.to_vec(),
|
||
person: Some(person),
|
||
}
|
||
}
|
||
|
||
fn pair(i: usize, j: usize, probability: f32) -> Pair {
|
||
Pair { i, j, probability }
|
||
}
|
||
|
||
/// The failure the module exists to fix: face 0 matches its own group's
|
||
/// three members strongly, and the group has forty more it is unrelated to.
|
||
/// The old within-group mean reported ~0.07 for this.
|
||
#[test]
|
||
fn a_large_group_does_not_dilute_a_strong_match() {
|
||
let members: Vec<usize> = (0..44).collect();
|
||
let clusters = vec![cluster(&members)];
|
||
let pairs = vec![pair(0, 1, 0.99), pair(0, 2, 0.97), pair(0, 3, 0.95)];
|
||
|
||
let shares = identity_shares(44, &clusters, &pairs, TOP_MATCHES);
|
||
assert!(
|
||
(shares[0] - 0.97).abs() < 1e-6,
|
||
"the mean of its three real matches, undiluted: {}",
|
||
shares[0]
|
||
);
|
||
}
|
||
|
||
/// Two named people matching equally well is a coin toss, and saying so is
|
||
/// the point — this is the sibling case FR-CULL-10 warns about.
|
||
#[test]
|
||
fn an_ambiguous_face_splits_its_confidence_between_the_rivals() {
|
||
let clusters = vec![named(&[0, 1, 2], 1), named(&[3, 4], 2)];
|
||
let pairs = vec![
|
||
pair(0, 1, 0.90),
|
||
pair(0, 2, 0.90),
|
||
pair(0, 3, 0.90),
|
||
pair(0, 4, 0.90),
|
||
];
|
||
|
||
let shares = identity_shares(5, &clusters, &pairs, TOP_MATCHES);
|
||
// Coherent at 0.90, and only half of the evidence is its own.
|
||
assert!(
|
||
(shares[0] - 0.45).abs() < 1e-6,
|
||
"even evidence both ways: {}",
|
||
shares[0]
|
||
);
|
||
}
|
||
|
||
/// A rival below the merge threshold still has to count, which is why the
|
||
/// evidence scan reaches down to [`RIVAL_FLOOR`].
|
||
#[test]
|
||
fn a_rival_too_weak_to_merge_still_lowers_the_confidence() {
|
||
let clusters = vec![named(&[0, 1], 1), named(&[2, 3], 2)];
|
||
let sure = identity_shares(4, &clusters, &[pair(0, 1, 0.95)], TOP_MATCHES);
|
||
let contested = identity_shares(
|
||
4,
|
||
&clusters,
|
||
&[pair(0, 1, 0.95), pair(0, 2, 0.60)],
|
||
TOP_MATCHES,
|
||
);
|
||
|
||
assert_eq!(sure[0], 0.95, "nobody else to be: its coherence stands");
|
||
assert!(
|
||
contested[0] < 0.59 && contested[0] > 0.57,
|
||
"0.95 coherent, but 0.95 against 0.60: {}",
|
||
contested[0]
|
||
);
|
||
}
|
||
|
||
/// A fragment of the same person is not a rival. Measured on a real
|
||
/// library, counting unnamed groups as competition put four suggestions in
|
||
/// five under half — see the module note.
|
||
#[test]
|
||
fn an_unnamed_group_is_not_treated_as_competition() {
|
||
let clusters = vec![cluster(&[0, 1]), cluster(&[2, 3])];
|
||
let shares = identity_shares(
|
||
4,
|
||
&clusters,
|
||
&[pair(0, 1, 0.95), pair(0, 2, 0.90)],
|
||
TOP_MATCHES,
|
||
);
|
||
assert_eq!(
|
||
shares[0], 0.95,
|
||
"an unnamed group took evidence off a suggestion"
|
||
);
|
||
}
|
||
|
||
/// The cap, doing its job: an identity with fifty mediocre matches must not
|
||
/// beat one with ten strong ones on volume alone.
|
||
#[test]
|
||
fn evidence_is_capped_so_the_biggest_group_cannot_win_on_volume() {
|
||
let small: Vec<usize> = (0..11).collect();
|
||
let large: Vec<usize> = (11..62).collect();
|
||
let clusters = vec![named(&small, 1), named(&large, 2)];
|
||
|
||
let mut pairs: Vec<Pair> = (1..11).map(|j| pair(0, j, 0.90)).collect();
|
||
pairs.extend((11..62).map(|j| pair(0, j, 0.55)));
|
||
|
||
let shares = identity_shares(62, &clusters, &pairs, TOP_MATCHES);
|
||
// Ten at 0.90 against ten at 0.55 — not fifty-one at 0.55.
|
||
assert!(
|
||
(shares[0] - 0.90 * (9.0 / 14.5)).abs() < 1e-5,
|
||
"capped at ten either side: {}",
|
||
shares[0]
|
||
);
|
||
}
|
||
|
||
/// A face nothing has any evidence about claims nothing.
|
||
#[test]
|
||
fn a_face_with_no_evidence_reports_no_confidence() {
|
||
let clusters = vec![cluster(&[0, 1])];
|
||
let shares = identity_shares(2, &clusters, &[], TOP_MATCHES);
|
||
assert_eq!(shares, vec![0.0, 0.0]);
|
||
}
|
||
}
|