Score a suggestion against the people the user has named
The number beside a suggestion was the mean calibrated probability between the face and the rest of its group, which measures the wrong thing twice. It punishes coverage: a person with two hundred faces over fifteen years is *meant* to have members a given photograph is orthogonal to, so a correct suggestion onto a well-photographed person scored low for being well photographed. And it never asked who else the face might be — a face matching Anna at 0.95 and nobody else, and one matching Anna at 0.95 and her sister at 0.93, came out identical, when the second is the only one worth the user's attention. dr_face::assign answers both, and multiplies them: the mean of the best ten calibrated matches into the identity (the old mean, capped, which is what stops coverage counting against it), times that identity's share of the evidence against every *named* rival. Only named people compete, and per person rather than per group. Both halves of that had to be measured on a real 18,000-face library rather than reasoned about. Normalising across every group made the number useless — median suggestion 21%, four in five under half — because clustering leaves one person spread over many groups, so a face competed against itself; and keying rivals by group left Catherine competing with Catherine, median 39%. Per named person: median 99.5%. Rivals are gathered below the merge threshold, down to even odds: a named person matching at 0.6 will never be merged into but is exactly the competition to discount for. That would be a second similarity scan, the expensive half of regrouping a library, so cluster_scored scans once at the looser floor and hands the merge engine the subset at or above the threshold — pair for pair what it would have scanned for itself, held to that by a test. Leave-one-out over that library's 2,702 confirmations across 54 named people: 99.33% of faces placed on the right person against the old mean's 99.15%, and the number shown for the right person moves from a median of 90.4% to 99.3%. It errs low — 100% correct wherever it states 80% or more — which is the safe direction, and docs/faces.md §9.1 says plainly that the low bands are not calibrated. The example that measures it comes too: this is a claim about a library's numbers, and nobody should have to take it on faith. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
+160
-12
@@ -148,22 +148,110 @@ pub fn cluster(faces: &[Candidate], cal: &Calibration, min_probability: f32) ->
|
||||
return Vec::new();
|
||||
}
|
||||
|
||||
let embeddings: Vec<Vec<f32>> = faces.iter().map(|f| f.embedding.clone()).collect();
|
||||
let crop_px: Vec<f32> = faces.iter().map(|f| f.crop_px).collect();
|
||||
let images: Vec<u64> = faces.iter().map(|f| f.image).collect();
|
||||
let view = Faces {
|
||||
embeddings: &embeddings,
|
||||
crop_px: &crop_px,
|
||||
images: &images,
|
||||
};
|
||||
|
||||
let columns = Columns::of(faces);
|
||||
// Every pair that could ever contribute to a merge. See the module note on
|
||||
// why nothing outside this list can matter.
|
||||
let pairs = neighbours::above_threshold(&view, cal, min_probability);
|
||||
let pairs = neighbours::above_threshold(&columns.view(), cal, min_probability);
|
||||
build(faces, cal, min_probability, &pairs)
|
||||
}
|
||||
|
||||
/// Groups, and how confident each face's placement is.
|
||||
///
|
||||
/// The second half is [`crate::assign`]'s share, not the pairwise probability
|
||||
/// that put the face in the group — see that module for why the two are
|
||||
/// different questions.
|
||||
#[derive(Debug, Clone, PartialEq)]
|
||||
pub struct Grouping {
|
||||
pub clusters: Vec<Cluster>,
|
||||
/// Indexed like the input faces. 0 for a face in no group.
|
||||
pub confidence: Vec<f32>,
|
||||
}
|
||||
|
||||
/// Group faces into people, and score each placement against its rivals.
|
||||
///
|
||||
/// What a caller writing suggestions into a catalog wants: [`cluster`] answers
|
||||
/// *which person*, this answers *and how sure*.
|
||||
///
|
||||
/// One scan, two thresholds. The similarity scan is the expensive part of the
|
||||
/// whole subsystem and running it twice — once to merge, once to find rivals —
|
||||
/// would double the cost of regrouping a library. So it runs once at the looser
|
||||
/// of the two floors, and the merge engine takes the subset at or above
|
||||
/// `min_probability`. That subset is identical, pair for pair and in the same
|
||||
/// order, to what a scan at `min_probability` would have produced, so grouping
|
||||
/// is unchanged by scoring being asked for: `scoring_does_not_change_the_
|
||||
/// groups` holds it to that.
|
||||
pub fn cluster_scored(faces: &[Candidate], cal: &Calibration, min_probability: f32) -> Grouping {
|
||||
if faces.is_empty() {
|
||||
return Grouping {
|
||||
clusters: Vec::new(),
|
||||
confidence: Vec::new(),
|
||||
};
|
||||
}
|
||||
|
||||
let columns = Columns::of(faces);
|
||||
let evidence = neighbours::above_threshold(
|
||||
&columns.view(),
|
||||
cal,
|
||||
min_probability.min(crate::assign::RIVAL_FLOOR),
|
||||
);
|
||||
let merges: Vec<neighbours::Pair> = evidence
|
||||
.iter()
|
||||
.copied()
|
||||
.filter(|p| p.probability >= min_probability)
|
||||
.collect();
|
||||
|
||||
let clusters = build(faces, cal, min_probability, &merges);
|
||||
let confidence = crate::assign::identity_shares(
|
||||
faces.len(),
|
||||
&clusters,
|
||||
&evidence,
|
||||
crate::assign::TOP_MATCHES,
|
||||
);
|
||||
Grouping {
|
||||
clusters,
|
||||
confidence,
|
||||
}
|
||||
}
|
||||
|
||||
/// The three arrays [`neighbours::Faces`] borrows, owned.
|
||||
///
|
||||
/// [`neighbours`] takes parallel slices rather than candidates on purpose — it
|
||||
/// has no business knowing what a person is — so somebody has to hold the
|
||||
/// columns. Both entry points do, identically, which is the only reason this is
|
||||
/// a type and not three locals.
|
||||
struct Columns {
|
||||
embeddings: Vec<Vec<f32>>,
|
||||
crop_px: Vec<f32>,
|
||||
images: Vec<u64>,
|
||||
}
|
||||
|
||||
impl Columns {
|
||||
fn of(faces: &[Candidate]) -> Self {
|
||||
Self {
|
||||
embeddings: faces.iter().map(|f| f.embedding.clone()).collect(),
|
||||
crop_px: faces.iter().map(|f| f.crop_px).collect(),
|
||||
images: faces.iter().map(|f| f.image).collect(),
|
||||
}
|
||||
}
|
||||
|
||||
fn view(&self) -> Faces<'_> {
|
||||
Faces {
|
||||
embeddings: &self.embeddings,
|
||||
crop_px: &self.crop_px,
|
||||
images: &self.images,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
fn build(
|
||||
faces: &[Candidate],
|
||||
cal: &Calibration,
|
||||
min_probability: f32,
|
||||
pairs: &[neighbours::Pair],
|
||||
) -> Vec<Cluster> {
|
||||
let mut engine = Engine::new(faces, cal, min_probability);
|
||||
for component in components(faces.len(), &pairs) {
|
||||
engine.agglomerate(&component, &pairs);
|
||||
for component in components(faces.len(), pairs) {
|
||||
engine.agglomerate(&component, pairs);
|
||||
}
|
||||
engine.finish()
|
||||
}
|
||||
@@ -663,6 +751,66 @@ mod tests {
|
||||
assert_eq!(out.len(), 2, "clustering overrode two user confirmations");
|
||||
}
|
||||
|
||||
/// A face at a chosen cosine to identity 0 *and* to identity 1 at once —
|
||||
/// the sibling geometry, which [`at_cosine`]'s per-identity subspaces
|
||||
/// cannot express.
|
||||
fn contested(to_first: f32, to_second: f32) -> Vec<f32> {
|
||||
let mut v = vec![0.0_f32; EMBEDDING_DIM];
|
||||
v[0] = to_first;
|
||||
v[2] = to_second;
|
||||
v[4] = (1.0 - to_first * to_first - to_second * to_second)
|
||||
.max(0.0)
|
||||
.sqrt();
|
||||
v
|
||||
}
|
||||
|
||||
/// The population the scoring tests share: two faces of one person, a
|
||||
/// stranger, and a face that matches the person well and the stranger
|
||||
/// weakly — weakly enough that it will never merge with them, which is
|
||||
/// exactly the rival a within-group score cannot see.
|
||||
fn with_a_rival() -> Vec<Candidate> {
|
||||
let mut x = candidate(4, 13, 0, 0.0);
|
||||
x.embedding = contested(0.45, 0.37);
|
||||
// The stranger is *named*: only an identity the user has asserted
|
||||
// competes for a face (crate::assign).
|
||||
let mut stranger = candidate(3, 12, 1, 1.0);
|
||||
stranger.confirmed_person = Some(7);
|
||||
vec![
|
||||
candidate(1, 10, 0, 1.0),
|
||||
candidate(2, 11, 0, 1.0),
|
||||
stranger,
|
||||
x,
|
||||
]
|
||||
}
|
||||
|
||||
/// Asking for confidences must not move a single face. The scan runs at a
|
||||
/// looser floor to find rivals, and the merge engine has to see exactly the
|
||||
/// pairs it would have seen without them.
|
||||
#[test]
|
||||
fn scoring_does_not_change_the_groups() {
|
||||
let faces = with_a_rival();
|
||||
let plain = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY);
|
||||
let scored = cluster_scored(&faces, &cal(), DEFAULT_MERGE_PROBABILITY);
|
||||
assert_eq!(plain, scored.clusters);
|
||||
}
|
||||
|
||||
/// The number the user is shown answers "which of these people", so a
|
||||
/// second claimant has to lower it even when it is too weak to merge.
|
||||
#[test]
|
||||
fn a_face_two_identities_could_claim_is_reported_as_less_certain() {
|
||||
let faces = with_a_rival();
|
||||
let scored = cluster_scored(&faces, &cal(), DEFAULT_MERGE_PROBABILITY);
|
||||
|
||||
let alone = cluster_scored(&faces[..2], &cal(), DEFAULT_MERGE_PROBABILITY);
|
||||
assert!(alone.confidence[0] > 0.99, "nobody else to be");
|
||||
|
||||
let contested = scored.confidence[3];
|
||||
assert!(
|
||||
(0.6..0.85).contains(&contested),
|
||||
"a face with a second claimant: {contested}"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_suggestion_joins_the_person_its_group_is_anchored_to() {
|
||||
let mut anchor = candidate(1, 10, 0, 1.0);
|
||||
|
||||
Reference in New Issue
Block a user