Score a suggestion against the people the user has named
The number beside a suggestion was the mean calibrated probability between the face and the rest of its group, which measures the wrong thing twice. It punishes coverage: a person with two hundred faces over fifteen years is *meant* to have members a given photograph is orthogonal to, so a correct suggestion onto a well-photographed person scored low for being well photographed. And it never asked who else the face might be — a face matching Anna at 0.95 and nobody else, and one matching Anna at 0.95 and her sister at 0.93, came out identical, when the second is the only one worth the user's attention. dr_face::assign answers both, and multiplies them: the mean of the best ten calibrated matches into the identity (the old mean, capped, which is what stops coverage counting against it), times that identity's share of the evidence against every *named* rival. Only named people compete, and per person rather than per group. Both halves of that had to be measured on a real 18,000-face library rather than reasoned about. Normalising across every group made the number useless — median suggestion 21%, four in five under half — because clustering leaves one person spread over many groups, so a face competed against itself; and keying rivals by group left Catherine competing with Catherine, median 39%. Per named person: median 99.5%. Rivals are gathered below the merge threshold, down to even odds: a named person matching at 0.6 will never be merged into but is exactly the competition to discount for. That would be a second similarity scan, the expensive half of regrouping a library, so cluster_scored scans once at the looser floor and hands the merge engine the subset at or above the threshold — pair for pair what it would have scanned for itself, held to that by a test. Leave-one-out over that library's 2,702 confirmations across 54 named people: 99.33% of faces placed on the right person against the old mean's 99.15%, and the number shown for the right person moves from a median of 90.4% to 99.3%. It errs low — 100% correct wherever it states 80% or more — which is the safe direction, and docs/faces.md §9.1 says plainly that the low bands are not calibrated. The example that measures it comes too: this is a claim about a library's numbers, and nobody should have to take it on faith. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,373 @@
|
||||
//! What the suggestion confidence would say about a real library.
|
||||
//!
|
||||
//! cargo run --release -p dr-catalog --example face_confidence -- CATALOG.sqlite [--full]
|
||||
//!
|
||||
//! Read-only: it writes nothing to the catalog, so it can be pointed at a copy
|
||||
//! of a live library and re-run at will.
|
||||
//!
|
||||
//! # What it measures
|
||||
//!
|
||||
//! The user's own confirmations are the only ground truth a library has, so
|
||||
//! the evaluation is leave-one-out over them: hide one confirmed face, ask the
|
||||
//! scorer which of the confirmed identities it belongs to, and compare with
|
||||
//! what the user said. Faces from the same photograph are excluded exactly as
|
||||
//! the clusterer excludes them, so nothing is scored against a co-occurrence
|
||||
//! that would never have been allowed to merge.
|
||||
//!
|
||||
//! Two numbers are compared on that task: the share (`dr_face::assign`) and the
|
||||
//! mean-within-group figure it replaced. Accuracy says which one picks the
|
||||
//! right person; the reliability table says whether the percentage the user is
|
||||
//! shown means what it claims — which is the question FR-CULL-9 exists for.
|
||||
|
||||
use std::collections::HashMap;
|
||||
|
||||
use dr_catalog::faces::{self, PersonId};
|
||||
use dr_catalog::Catalog;
|
||||
|
||||
const MODEL_ID: &str = "w600k_mbf";
|
||||
const TOP: usize = 10;
|
||||
|
||||
struct Known {
|
||||
image: u64,
|
||||
person: PersonId,
|
||||
embedding: Vec<f32>,
|
||||
crop_px: f32,
|
||||
}
|
||||
|
||||
fn main() {
|
||||
let args: Vec<String> = std::env::args().skip(1).collect();
|
||||
let Some(path) = args.first() else {
|
||||
eprintln!("usage: face_confidence CATALOG.sqlite [--full]");
|
||||
std::process::exit(2);
|
||||
};
|
||||
|
||||
let catalog = Catalog::open(std::path::Path::new(path)).expect("open catalog");
|
||||
let conn = catalog.connection();
|
||||
|
||||
let cal = match faces::calibration(conn, MODEL_ID) {
|
||||
Ok(Some((c, _))) => c,
|
||||
_ => dr_face::Calibration::default(),
|
||||
};
|
||||
println!(
|
||||
"calibration: a={:.2} b={:.2} w_size={:.3} valid={} (P=0.5 at cosine {:.3})",
|
||||
cal.a,
|
||||
cal.b,
|
||||
cal.w_size,
|
||||
cal.valid,
|
||||
cal.boundary_at(0.5, 150.0, 0.0)
|
||||
);
|
||||
|
||||
let model = dr_face::ModelId::new(MODEL_ID.to_string());
|
||||
let stored = faces::embeddings(conn, MODEL_ID).expect("embeddings");
|
||||
let mut embedding_of = HashMap::new();
|
||||
for (id, image, blob, crop_px) in stored {
|
||||
if let Some(e) = dr_face::Embedding::from_f16_bytes(model.clone(), &blob) {
|
||||
embedding_of.insert(id, (image.0, e.v.to_vec(), crop_px));
|
||||
}
|
||||
}
|
||||
println!("faces with embeddings: {}", embedding_of.len());
|
||||
|
||||
// The ground truth: every confirmed face, under the person the user put it
|
||||
// on. Identities with a single confirmation are dropped — leaving one out
|
||||
// leaves that identity with no evidence at all, so they measure nothing.
|
||||
let people = faces::people(conn).expect("people");
|
||||
let mut known: Vec<Known> = Vec::new();
|
||||
let mut identities = 0usize;
|
||||
for p in &people {
|
||||
if p.confirmed_faces < 2 {
|
||||
continue;
|
||||
}
|
||||
let mut mine = Vec::new();
|
||||
for f in faces::for_person(conn, p.id, false).expect("faces") {
|
||||
if !f.confirmed {
|
||||
continue;
|
||||
}
|
||||
if let Some((image, embedding, crop_px)) = embedding_of.get(&f.id) {
|
||||
mine.push(Known {
|
||||
image: *image,
|
||||
person: p.id,
|
||||
embedding: embedding.clone(),
|
||||
crop_px: *crop_px,
|
||||
});
|
||||
}
|
||||
}
|
||||
if mine.len() >= 2 {
|
||||
identities += 1;
|
||||
known.extend(mine);
|
||||
}
|
||||
}
|
||||
println!(
|
||||
"ground truth: {} confirmed faces across {identities} identities\n",
|
||||
known.len()
|
||||
);
|
||||
if known.len() < 2 {
|
||||
println!("not enough confirmations to evaluate.");
|
||||
return;
|
||||
}
|
||||
|
||||
let mut share_right = 0usize;
|
||||
let mut mean_right = 0usize;
|
||||
// (share of the winner, was the winner correct)
|
||||
let mut reliability: Vec<(f32, bool)> = Vec::with_capacity(known.len());
|
||||
// What each scorer would have *displayed* for the correct answer.
|
||||
let mut shown_share = Vec::with_capacity(known.len());
|
||||
let mut shown_mean = Vec::with_capacity(known.len());
|
||||
|
||||
for (i, me) in known.iter().enumerate() {
|
||||
let mut per_person: HashMap<PersonId, Vec<f32>> = HashMap::new();
|
||||
// The mean baseline is the old code's, which had no floor: it averaged
|
||||
// over every member of the group.
|
||||
let mut all_person: HashMap<PersonId, Vec<f32>> = HashMap::new();
|
||||
for (j, them) in known.iter().enumerate() {
|
||||
if i == j || me.image == them.image {
|
||||
continue;
|
||||
}
|
||||
let cos: f32 = me
|
||||
.embedding
|
||||
.iter()
|
||||
.zip(&them.embedding)
|
||||
.map(|(a, b)| a * b)
|
||||
.sum();
|
||||
// The floor the real scorer sees: `cluster_scored` scans at
|
||||
// `RIVAL_FLOOR` and `identity_shares` never learns about a pair
|
||||
// below it. Summing the near-orthogonal ones here instead of
|
||||
// dropping them is not a stricter test, it is a different
|
||||
// function — fifty identities contributing their *upper tail* of
|
||||
// noise outweigh one contributing a real match.
|
||||
let probability = cal.probability(cos, me.crop_px.min(them.crop_px), 0.0);
|
||||
if probability >= dr_face::RIVAL_FLOOR {
|
||||
per_person.entry(them.person).or_default().push(probability);
|
||||
}
|
||||
all_person.entry(them.person).or_default().push(probability);
|
||||
}
|
||||
|
||||
// The share: sum of the best TOP matches per identity, normalised.
|
||||
// Evidence, and the coherence that goes with it: the sum of the best
|
||||
// TOP matches, and their mean. dr_face::assign shows the product of
|
||||
// that mean and the identity's share of the total.
|
||||
let mut evidence: Vec<(PersonId, f32, f32)> = per_person
|
||||
.iter()
|
||||
.map(|(&p, probabilities)| {
|
||||
let mut v = probabilities.clone();
|
||||
v.sort_by(|a, b| b.total_cmp(a));
|
||||
let counted = v.len().min(TOP);
|
||||
let sum = v.iter().take(TOP).sum::<f32>();
|
||||
(p, sum, sum / counted as f32)
|
||||
})
|
||||
.collect();
|
||||
let total: f32 = evidence.iter().map(|(_, s, _)| *s).sum();
|
||||
evidence.sort_by(|a, b| b.1.total_cmp(&a.1));
|
||||
|
||||
// The number it replaced: the mean over every member of the identity.
|
||||
let mut means: Vec<(PersonId, f32)> = all_person
|
||||
.iter()
|
||||
.map(|(&p, probabilities)| {
|
||||
(
|
||||
p,
|
||||
probabilities.iter().sum::<f32>() / probabilities.len() as f32,
|
||||
)
|
||||
})
|
||||
.collect();
|
||||
means.sort_by(|a, b| b.1.total_cmp(&a.1));
|
||||
|
||||
if let (Some(&(winner, score, coherence)), true) = (evidence.first(), total > 0.0) {
|
||||
let correct = winner == me.person;
|
||||
share_right += correct as usize;
|
||||
// What the screen would say about the identity it picked.
|
||||
reliability.push((coherence * score / total, correct));
|
||||
let ours = evidence
|
||||
.iter()
|
||||
.find(|(p, _, _)| *p == me.person)
|
||||
.map(|(_, s, c)| c * s / total)
|
||||
.unwrap_or(0.0);
|
||||
shown_share.push(ours);
|
||||
}
|
||||
if let Some(&(winner, _)) = means.first() {
|
||||
mean_right += (winner == me.person) as usize;
|
||||
shown_mean.push(
|
||||
means
|
||||
.iter()
|
||||
.find(|(p, _)| *p == me.person)
|
||||
.map(|(_, s)| *s)
|
||||
.unwrap_or(0.0),
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
let n = known.len() as f64;
|
||||
println!("which identity does this face belong to? (leave-one-out, top-1)");
|
||||
println!(
|
||||
" share of evidence {:>6.2}% ({share_right}/{})",
|
||||
100.0 * share_right as f64 / n,
|
||||
known.len()
|
||||
);
|
||||
println!(
|
||||
" mean within group {:>6.2}% ({mean_right}/{})\n",
|
||||
100.0 * mean_right as f64 / n,
|
||||
known.len()
|
||||
);
|
||||
|
||||
println!("what the screen would show for the answer the user gave:");
|
||||
band(" share ", &shown_share);
|
||||
band(" mean ", &shown_mean);
|
||||
|
||||
println!("\nreliability of the share — is a stated {{n}}% right {{n}}% of the time?");
|
||||
println!(
|
||||
" {:>12} {:>7} {:>9} {:>8}",
|
||||
"stated", "faces", "correct", "gap"
|
||||
);
|
||||
for (lo, hi) in [
|
||||
(0.0, 0.5),
|
||||
(0.5, 0.6),
|
||||
(0.6, 0.7),
|
||||
(0.7, 0.8),
|
||||
(0.8, 0.9),
|
||||
(0.9, 0.95),
|
||||
(0.95, 1.001),
|
||||
] {
|
||||
let bucket: Vec<bool> = reliability
|
||||
.iter()
|
||||
.filter(|(s, _)| *s >= lo && *s < hi)
|
||||
.map(|(_, c)| *c)
|
||||
.collect();
|
||||
if bucket.is_empty() {
|
||||
continue;
|
||||
}
|
||||
let observed = bucket.iter().filter(|c| **c).count() as f64 / bucket.len() as f64;
|
||||
let stated = reliability
|
||||
.iter()
|
||||
.filter(|(s, _)| *s >= lo && *s < hi)
|
||||
.map(|(s, _)| *s as f64)
|
||||
.sum::<f64>()
|
||||
/ bucket.len() as f64;
|
||||
println!(
|
||||
" {:>5.0}–{:>3.0}% {:>9} {:>8.1}% {:>+7.1}",
|
||||
lo * 100.0,
|
||||
hi.min(1.0) * 100.0,
|
||||
bucket.len(),
|
||||
100.0 * observed,
|
||||
100.0 * (observed - stated)
|
||||
);
|
||||
}
|
||||
|
||||
if args.iter().any(|a| a == "--full") {
|
||||
// The confirmations go in as anchors, exactly as `recluster` sends
|
||||
// them: they are what makes a group a named identity, and therefore
|
||||
// what makes it a rival.
|
||||
let mut confirmed = HashMap::new();
|
||||
for p in &people {
|
||||
for f in faces::for_person(conn, p.id, false).expect("faces") {
|
||||
if f.confirmed {
|
||||
confirmed.insert(f.id, p.id.0);
|
||||
}
|
||||
}
|
||||
}
|
||||
full_library(&embedding_of, &confirmed, &cal);
|
||||
}
|
||||
}
|
||||
|
||||
/// Where a set of confidences actually falls.
|
||||
fn band(label: &str, v: &[f32]) {
|
||||
if v.is_empty() {
|
||||
return;
|
||||
}
|
||||
let mut s = v.to_vec();
|
||||
s.sort_by(|a, b| a.total_cmp(b));
|
||||
let pct = |q: f64| s[((s.len() - 1) as f64 * q) as usize];
|
||||
let mean = s.iter().sum::<f32>() / s.len() as f32;
|
||||
println!(
|
||||
"{label} median {:>5.1}% mean {:>5.1}% p10 {:>5.1}% p90 {:>5.1}% under 50%: {:>5.1}%",
|
||||
100.0 * pct(0.5),
|
||||
100.0 * mean,
|
||||
100.0 * pct(0.10),
|
||||
100.0 * pct(0.90),
|
||||
100.0 * s.iter().filter(|x| **x < 0.5).count() as f32 / s.len() as f32
|
||||
);
|
||||
}
|
||||
|
||||
/// The whole library through the real clusterer, for the numbers it would
|
||||
/// actually write.
|
||||
fn full_library(
|
||||
embedding_of: &HashMap<faces::FaceId, (u64, Vec<f32>, f32)>,
|
||||
confirmed: &HashMap<faces::FaceId, u64>,
|
||||
cal: &dr_face::Calibration,
|
||||
) {
|
||||
let mut candidates: Vec<dr_face::Candidate> = embedding_of
|
||||
.iter()
|
||||
.map(|(id, (image, embedding, crop_px))| dr_face::Candidate {
|
||||
face: id.0,
|
||||
image: *image,
|
||||
embedding: embedding.clone(),
|
||||
crop_px: *crop_px,
|
||||
confirmed_person: confirmed.get(id).copied(),
|
||||
})
|
||||
.collect();
|
||||
candidates.sort_by_key(|c| c.face);
|
||||
|
||||
println!("\nthe whole library, at the default merge probability:");
|
||||
let start = std::time::Instant::now();
|
||||
let grouping = dr_face::cluster_scored(&candidates, cal, dr_face::DEFAULT_MERGE_PROBABILITY);
|
||||
let real: Vec<_> = grouping
|
||||
.clusters
|
||||
.iter()
|
||||
.filter(|c| c.members.len() >= 2)
|
||||
.collect();
|
||||
let grouped: usize = real.iter().map(|c| c.members.len()).sum();
|
||||
println!(
|
||||
" {} face(s) → {} group(s) of two or more, holding {grouped} faces ({:.0}%), in {:.1}s",
|
||||
candidates.len(),
|
||||
real.len(),
|
||||
100.0 * grouped as f64 / candidates.len() as f64,
|
||||
start.elapsed().as_secs_f64()
|
||||
);
|
||||
|
||||
let named: Vec<_> = real.iter().filter(|c| c.person.is_some()).collect();
|
||||
println!(
|
||||
" {} of those group(s) carry a confirmation, holding {} faces",
|
||||
named.len(),
|
||||
named.iter().map(|c| c.members.len()).sum::<usize>()
|
||||
);
|
||||
|
||||
let shown: Vec<f32> = real
|
||||
.iter()
|
||||
.flat_map(|c| c.members.iter().map(|&m| grouping.confidence[m]))
|
||||
.collect();
|
||||
band(" new, all groups ", &shown);
|
||||
let onto_people: Vec<f32> = named
|
||||
.iter()
|
||||
.flat_map(|c| c.members.iter().map(|&m| grouping.confidence[m]))
|
||||
.collect();
|
||||
band(" new, onto a person", &onto_people);
|
||||
|
||||
// The number the old code would have written for the same grouping.
|
||||
let means: Vec<f32> = real
|
||||
.iter()
|
||||
.flat_map(|c| {
|
||||
c.members.iter().map(|&m| {
|
||||
let me = &candidates[m];
|
||||
let mut sum = 0.0;
|
||||
let mut n = 0.0;
|
||||
for &other in &c.members {
|
||||
if other == m {
|
||||
continue;
|
||||
}
|
||||
let them = &candidates[other];
|
||||
let cos: f32 = me
|
||||
.embedding
|
||||
.iter()
|
||||
.zip(&them.embedding)
|
||||
.map(|(a, b)| a * b)
|
||||
.sum();
|
||||
sum += cal.probability(cos, me.crop_px.min(them.crop_px), 0.0);
|
||||
n += 1.0;
|
||||
}
|
||||
if n == 0.0 {
|
||||
1.0
|
||||
} else {
|
||||
sum / n
|
||||
}
|
||||
})
|
||||
})
|
||||
.collect();
|
||||
band(" old, all groups ", &means);
|
||||
}
|
||||
Reference in New Issue
Block a user