Files
DarkRoom/core/dr-face/src/cluster.rs
T
dtourolleandClaude Opus 5 ebb7d3cf5c Score a suggestion against the people the user has named
The number beside a suggestion was the mean calibrated probability
between the face and the rest of its group, which measures the wrong
thing twice. It punishes coverage: a person with two hundred faces over
fifteen years is *meant* to have members a given photograph is
orthogonal to, so a correct suggestion onto a well-photographed person
scored low for being well photographed. And it never asked who else the
face might be — a face matching Anna at 0.95 and nobody else, and one
matching Anna at 0.95 and her sister at 0.93, came out identical, when
the second is the only one worth the user's attention.

dr_face::assign answers both, and multiplies them: the mean of the best
ten calibrated matches into the identity (the old mean, capped, which is
what stops coverage counting against it), times that identity's share of
the evidence against every *named* rival.

Only named people compete, and per person rather than per group. Both
halves of that had to be measured on a real 18,000-face library rather
than reasoned about. Normalising across every group made the number
useless — median suggestion 21%, four in five under half — because
clustering leaves one person spread over many groups, so a face competed
against itself; and keying rivals by group left Catherine competing with
Catherine, median 39%. Per named person: median 99.5%.

Rivals are gathered below the merge threshold, down to even odds: a
named person matching at 0.6 will never be merged into but is exactly
the competition to discount for. That would be a second similarity scan,
the expensive half of regrouping a library, so cluster_scored scans once
at the looser floor and hands the merge engine the subset at or above
the threshold — pair for pair what it would have scanned for itself,
held to that by a test.

Leave-one-out over that library's 2,702 confirmations across 54 named
people: 99.33% of faces placed on the right person against the old
mean's 99.15%, and the number shown for the right person moves from a
median of 90.4% to 99.3%. It errs low — 100% correct wherever it states
80% or more — which is the safe direction, and docs/faces.md §9.1 says
plainly that the low bands are not calibrated.

The example that measures it comes too: this is a claim about a
library's numbers, and nobody should have to take it on faith.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 10:44:01 +02:00

1140 lines
41 KiB
Rust
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
//! Grouping faces into people (docs/faces.md §9, FR-CULL-10).
//!
//! Model-free: this is arithmetic over embeddings, and it is where the
//! subsystem's accuracy actually lives, so it is testable with no weights on
//! the machine.
//!
//! # Constraints, not just a threshold
//!
//! FR-CULL-10 warns that clustering will over-merge on siblings, on parents and
//! children, and on the same person a decade apart. Two structural defences,
//! both cheaper than a better threshold:
//!
//! **Cannot-link on co-occurrence.** Two faces in the same photograph are never
//! merged. It is the same observation [`crate::calibrate`] mines for free
//! negatives, used here as a hard constraint, and it is the single cheapest
//! defence against over-merging that exists.
//!
//! **Confirmed faces are anchors.** A confirmation is user data (FR-CULL-12)
//! and clustering never moves it. Two groups holding confirmations of
//! *different* people cannot merge, whatever their similarity says.
//!
//! # Average link, not single link
//!
//! Single-link chains: one bad edge welds two identities together, and it is
//! the documented way face clustering fails on families. Average link asks
//! whether the *groups* are similar, which one outlier cannot force.
//!
//! # How it runs, and why the obvious way does not
//!
//! The first implementation of this was the textbook one: compute every
//! pairwise cosine, then repeatedly scan all live group pairs, score each with
//! average link, and merge the best. It is correct, it is twenty lines, and on
//! a real library it does not finish.
//!
//! The reason is that the scan is inside the loop. Each merge rescans every
//! surviving pair — `O(g²)` of them — and each score is recomputed from
//! scratch over every cross pair, `O(|A|·|B|)`. With 1,813 faces that is
//! roughly 1.6 million pair scores per merge and some 700 merges to do; the
//! window simply stops responding, which is what a user reports as "Regroup is
//! broken". At 25,000 faces it is not slow, it is impossible.
//!
//! Three changes, none of which alter the answer:
//!
//! **Only above-threshold pairs can ever matter.** An average that reaches the
//! threshold must have at least one term at or above it, so two groups with no
//! qualifying pair between them can never merge — not now and not after any
//! sequence of merges, since merging only adds terms. [`crate::neighbours`]
//! produces exactly that sparse pair list, and everything below works on it.
//! On the library above it is 7,875 pairs rather than 1.6 million.
//!
//! **Merges cannot cross components.** Groups only ever merge along those
//! pairs, so the connected components of that graph are independent problems.
//! A library of four hundred people becomes four hundred small agglomerations
//! instead of one large one, and the quadratic term is paid per component.
//!
//! **Average link is additive.** `sum(A ∪ B, C) = sum(A, C) + sum(B, C)`, so a
//! merged group's scores follow from the two it came from by addition — the
//! Lance-Williams update. Kept as running `(sum, count)` per adjacent pair, a
//! score costs one division instead of a nested loop, and a binary heap with
//! lazy invalidation replaces the rescan.
//!
//! The output is unchanged, deliberately and testably so: `the_fast_engine_
//! agrees_with_the_reference` runs both over the same population and asserts
//! the clusters are identical.
use std::cmp::Ordering;
use std::collections::{BinaryHeap, HashMap, HashSet};
use crate::calibrate::Calibration;
use crate::neighbours::{self, Faces};
/// Probability above which two groups are judged the same person.
///
/// Stated as a probability and not a cosine, because FR-CULL-9 forbids
/// thresholding a bare similarity anywhere in this subsystem.
///
/// # Why 0.80
///
/// It was 0.90, and 0.90 left most of a real library ungrouped. Measured over
/// the 1,813-face reference library, with `dr-ui`'s `face_index --tune`:
///
/// | P | cosine | groups | faces grouped | largest group |
/// |---|---|---|---|---|
/// | 0.95 | 0.449 | 311 | 62% | 51 |
/// | 0.90 | 0.403 | 316 | 67% | 51 |
/// | 0.85 | 0.374 | 318 | 70% | 57 |
/// | **0.80** | **0.353** | **328** | **74%** | **69** |
/// | 0.75 | 0.335 | 327 | 77% | 69 |
/// | 0.70 | 0.319 | 326 | 79% | 81 |
/// | 0.50 | 0.267 | 303 | 85% | 90 |
///
/// The count of *groups* is the signal, not the count of grouped faces. Loosen
/// from 0.95 and it climbs: real people are being assembled out of fragments.
/// It peaks at 0.80 and then falls, and a falling group count while the grouped
/// faces keep rising is the shape of over-merging — separate identities being
/// welded, which is the failure FR-CULL-10 warns about and the one the user
/// cannot easily undo by hand.
///
/// So: the loosest setting that is still building people rather than melting
/// them together. A third more of the library gets grouped than at 0.90, and
/// the largest group grows by eighteen faces rather than by forty.
///
/// This is a *default*, not a constant of nature — the numbers above are one
/// library, and `--tune` reruns the table on any other.
pub const DEFAULT_MERGE_PROBABILITY: f32 = 0.80;
/// A face presented to the clusterer.
///
/// Ids are opaque `u64`s rather than catalog types: this crate has no business
/// knowing what a `FaceId` means, and the caller does the translation.
#[derive(Debug, Clone)]
pub struct Candidate {
pub face: u64,
/// Which photograph it came from — the cannot-link key.
pub image: u64,
/// L2-normalised, `EMBEDDING_DIM` long.
pub embedding: Vec<f32>,
/// Source pixels across the aligned crop, for the calibration's size term.
pub crop_px: f32,
/// The person this face is *confirmed* to be, if any.
///
/// Suggestions are deliberately not passed here. They are this function's
/// own previous output, and feeding them back in would let a guess harden
/// into a fact across successive passes.
pub confirmed_person: Option<u64>,
}
/// One group of faces the clusterer believes are one person.
#[derive(Debug, Clone, PartialEq)]
pub struct Cluster {
/// Indices into the input slice.
pub members: Vec<usize>,
/// The person this group is already known to be, from its anchors.
///
/// `Some` means the group contains confirmed faces and the suggestions in
/// it attach to that existing person. `None` is a new unnamed group.
pub person: Option<u64>,
}
/// Group faces into people.
///
/// `min_probability` is compared against the calibrated average-link
/// probability between two groups. Deterministic: the same input yields the
/// same clusters, because the merge order is by score with the index pair as
/// the tiebreak.
pub fn cluster(faces: &[Candidate], cal: &Calibration, min_probability: f32) -> Vec<Cluster> {
if faces.is_empty() {
return Vec::new();
}
let columns = Columns::of(faces);
// Every pair that could ever contribute to a merge. See the module note on
// why nothing outside this list can matter.
let pairs = neighbours::above_threshold(&columns.view(), cal, min_probability);
build(faces, cal, min_probability, &pairs)
}
/// Groups, and how confident each face's placement is.
///
/// The second half is [`crate::assign`]'s share, not the pairwise probability
/// that put the face in the group — see that module for why the two are
/// different questions.
#[derive(Debug, Clone, PartialEq)]
pub struct Grouping {
pub clusters: Vec<Cluster>,
/// Indexed like the input faces. 0 for a face in no group.
pub confidence: Vec<f32>,
}
/// Group faces into people, and score each placement against its rivals.
///
/// What a caller writing suggestions into a catalog wants: [`cluster`] answers
/// *which person*, this answers *and how sure*.
///
/// One scan, two thresholds. The similarity scan is the expensive part of the
/// whole subsystem and running it twice — once to merge, once to find rivals —
/// would double the cost of regrouping a library. So it runs once at the looser
/// of the two floors, and the merge engine takes the subset at or above
/// `min_probability`. That subset is identical, pair for pair and in the same
/// order, to what a scan at `min_probability` would have produced, so grouping
/// is unchanged by scoring being asked for: `scoring_does_not_change_the_
/// groups` holds it to that.
pub fn cluster_scored(faces: &[Candidate], cal: &Calibration, min_probability: f32) -> Grouping {
if faces.is_empty() {
return Grouping {
clusters: Vec::new(),
confidence: Vec::new(),
};
}
let columns = Columns::of(faces);
let evidence = neighbours::above_threshold(
&columns.view(),
cal,
min_probability.min(crate::assign::RIVAL_FLOOR),
);
let merges: Vec<neighbours::Pair> = evidence
.iter()
.copied()
.filter(|p| p.probability >= min_probability)
.collect();
let clusters = build(faces, cal, min_probability, &merges);
let confidence = crate::assign::identity_shares(
faces.len(),
&clusters,
&evidence,
crate::assign::TOP_MATCHES,
);
Grouping {
clusters,
confidence,
}
}
/// The three arrays [`neighbours::Faces`] borrows, owned.
///
/// [`neighbours`] takes parallel slices rather than candidates on purpose — it
/// has no business knowing what a person is — so somebody has to hold the
/// columns. Both entry points do, identically, which is the only reason this is
/// a type and not three locals.
struct Columns {
embeddings: Vec<Vec<f32>>,
crop_px: Vec<f32>,
images: Vec<u64>,
}
impl Columns {
fn of(faces: &[Candidate]) -> Self {
Self {
embeddings: faces.iter().map(|f| f.embedding.clone()).collect(),
crop_px: faces.iter().map(|f| f.crop_px).collect(),
images: faces.iter().map(|f| f.image).collect(),
}
}
fn view(&self) -> Faces<'_> {
Faces {
embeddings: &self.embeddings,
crop_px: &self.crop_px,
images: &self.images,
}
}
}
fn build(
faces: &[Candidate],
cal: &Calibration,
min_probability: f32,
pairs: &[neighbours::Pair],
) -> Vec<Cluster> {
let mut engine = Engine::new(faces, cal, min_probability);
for component in components(faces.len(), pairs) {
engine.agglomerate(&component, pairs);
}
engine.finish()
}
/// Split one person's faces into the groups a raised threshold separates them
/// into.
///
/// FR-CULL-10 requires splitting to be as easy as merging, and a split that
/// hands the user a pile of loose faces to re-sort is not that. This re-runs
/// the same agglomeration at a stricter probability so the user is offered
/// coherent sub-groups to pull apart.
///
/// Anchors are ignored here on purpose: every face in the input is already
/// nominally the same person, so honouring the anchors would refuse to split
/// anything.
pub fn split(faces: &[Candidate], cal: &Calibration, min_probability: f32) -> Vec<Cluster> {
let anchorless: Vec<Candidate> = faces
.iter()
.cloned()
.map(|mut f| {
f.confirmed_person = None;
f
})
.collect();
cluster(&anchorless, cal, min_probability)
}
// ── the merge engine ──────────────────────────────────────────────────────
/// One live group, identified throughout by the index of its lowest member.
#[derive(Debug)]
struct Group {
/// Ascending, always — [`Engine::cross`] sums in this order, and a stable
/// order is what makes the floating-point total reproducible.
members: Vec<usize>,
images: HashSet<u64>,
person: Option<u64>,
alive: bool,
/// Bumped on every merge, so heap entries naming an older state can be
/// recognised and dropped instead of acted on.
version: u64,
}
/// Running average-link state for one adjacent pair of groups.
///
/// `sum` is over **every** cross pair, not only the above-threshold ones —
/// average link is an average over all of them, and counting only the
/// qualifying pairs would report a similarity no group actually has.
#[derive(Debug, Clone, Copy)]
struct Link {
sum: f64,
count: f64,
}
impl Link {
fn probability(&self) -> f32 {
if self.count == 0.0 {
0.0
} else {
(self.sum / self.count) as f32
}
}
}
/// A candidate merge, waiting in the heap.
#[derive(Debug, Clone, Copy)]
struct Pending {
probability: f32,
a: usize,
b: usize,
/// Group versions when this was pushed. A mismatch on pop means a merge
/// has happened since and a fresher entry for this pair is already queued.
va: u64,
vb: u64,
}
impl PartialEq for Pending {
fn eq(&self, other: &Self) -> bool {
self.cmp(other) == Ordering::Equal
}
}
impl Eq for Pending {}
impl PartialOrd for Pending {
fn partial_cmp(&self, other: &Self) -> Option<Ordering> {
Some(self.cmp(other))
}
}
impl Ord for Pending {
/// Greatest pops first, so: highest probability, and on a tie the lowest
/// index pair. That tiebreak is not cosmetic — it is what the old
/// ascending scan did, and it is the whole of the determinism guarantee.
fn cmp(&self, other: &Self) -> Ordering {
self.probability
.total_cmp(&other.probability)
.then_with(|| other.a.cmp(&self.a))
.then_with(|| other.b.cmp(&self.b))
}
}
struct Engine<'a> {
faces: &'a [Candidate],
cal: &'a Calibration,
min_probability: f32,
groups: Vec<Group>,
links: HashMap<(usize, usize), Link>,
/// Adjacency, as group ids. Kept alongside `links` so a merge can find
/// everything it has to update without scanning the whole map.
adjacent: Vec<HashSet<usize>>,
}
impl<'a> Engine<'a> {
fn new(faces: &'a [Candidate], cal: &'a Calibration, min_probability: f32) -> Self {
let groups = faces
.iter()
.enumerate()
.map(|(i, f)| Group {
members: vec![i],
images: HashSet::from([f.image]),
person: f.confirmed_person,
alive: true,
version: 0,
})
.collect();
Self {
faces,
cal,
min_probability,
groups,
links: HashMap::new(),
adjacent: vec![HashSet::new(); faces.len()],
}
}
/// Agglomerate one connected component to exhaustion.
fn agglomerate(&mut self, component: &[usize], pairs: &[neighbours::Pair]) {
if component.len() < 2 {
return;
}
let members: HashSet<usize> = component.iter().copied().collect();
let mut heap = BinaryHeap::new();
for p in pairs.iter().filter(|p| members.contains(&p.i)) {
self.links.insert(
key(p.i, p.j),
Link {
sum: p.probability as f64,
count: 1.0,
},
);
self.adjacent[p.i].insert(p.j);
self.adjacent[p.j].insert(p.i);
heap.push(Pending {
probability: p.probability,
a: p.i.min(p.j),
b: p.i.max(p.j),
va: 0,
vb: 0,
});
}
while let Some(top) = heap.pop() {
let Pending {
probability,
a,
b,
va,
vb,
} = top;
// Stale: one side has merged since this was queued, and the
// replacement entry is already in the heap.
if !self.groups[a].alive
|| !self.groups[b].alive
|| self.groups[a].version != va
|| self.groups[b].version != vb
{
continue;
}
// The heap is ordered by probability, so the first entry below the
// bar means nothing left in this component can reach it.
if probability < self.min_probability {
break;
}
if !self.can_link(a, b) {
// Never becomes possible again: images only accumulate and an
// anchor is never given up, so drop the pair for good.
self.unlink(a, b);
continue;
}
self.merge(a, b, &mut heap);
}
}
/// Whether two groups are allowed to merge at all, before similarity is
/// asked.
fn can_link(&self, a: usize, b: usize) -> bool {
let (ga, gb) = (&self.groups[a], &self.groups[b]);
// Two confirmations of different people. The user has said these are
// not the same person, and no similarity overrides that.
if let (Some(pa), Some(pb)) = (ga.person, gb.person) {
if pa != pb {
return false;
}
}
// Co-occurrence: a photograph containing a face from each group means
// the two faces are in the same frame, so they are not the same person.
ga.images.is_disjoint(&gb.images)
}
/// Fold `b` into `a` and re-score everything that touched either.
fn merge(&mut self, a: usize, b: usize, heap: &mut BinaryHeap<Pending>) {
// Sorted, and deduplicated by the set: `a` and `b` may share
// neighbours, and each must be visited once. Sorting is what keeps the
// floating-point sums identical from run to run.
let mut touched: Vec<usize> = self.adjacent[a]
.union(&self.adjacent[b])
.copied()
.filter(|&c| c != a && c != b && self.groups[c].alive)
.collect();
touched.sort_unstable();
// Take the pair sums before the groups change underneath them.
let carried: Vec<(usize, Option<Link>, Option<Link>)> = touched
.iter()
.map(|&c| {
(
c,
self.links.get(&key(a, c)).copied(),
self.links.get(&key(b, c)).copied(),
)
})
.collect();
// a's own members, before b's are folded in. A missing a-side sum has
// to be computed over these and not over the merged list, or b's
// contribution would be counted twice.
let a_members = self.groups[a].members.clone();
// Absorb b into a.
let taken = std::mem::replace(
&mut self.groups[b],
Group {
members: Vec::new(),
images: HashSet::new(),
person: None,
alive: false,
version: 0,
},
);
{
let ga = &mut self.groups[a];
ga.members.extend(taken.members.iter().copied());
ga.members.sort_unstable();
ga.images.extend(taken.images.iter().copied());
// At most one side carries a person: `can_link` refuses a merge of
// two groups anchored to different people, so this cannot silently
// discard one of them.
ga.person = ga.person.or(taken.person);
ga.version += 1;
}
// b's own links are gone with it.
for c in self.adjacent[b].clone() {
self.links.remove(&key(b, c));
self.adjacent[c].remove(&b);
}
self.adjacent[b].clear();
self.links.remove(&key(a, b));
self.adjacent[a].remove(&b);
for (c, from_a, from_b) in carried {
// Dropping a pair the constraints now forbid saves computing a
// score for a merge that can never happen — which for a newly
// adjacent side is a real cost, not a bookkeeping one.
if !self.can_link(a, c) {
self.unlink(a, c);
continue;
}
// A side with no stored link was not adjacent before, so its cross
// pairs were all below threshold and were never summed. They still
// belong in the average, so they are computed now — once, after
// which the additive update carries them forward.
let from_a = from_a.unwrap_or_else(|| self.cross(&a_members, c));
let from_b = from_b.unwrap_or_else(|| self.cross(&taken.members, c));
let merged = Link {
sum: from_a.sum + from_b.sum,
count: from_a.count + from_b.count,
};
self.links.insert(key(a, c), merged);
self.adjacent[a].insert(c);
self.adjacent[c].insert(a);
heap.push(Pending {
probability: merged.probability(),
a: a.min(c),
b: a.max(c),
va: self.groups[a.min(c)].version,
vb: self.groups[a.max(c)].version,
});
}
}
/// Exact `(sum, count)` over every cross pair between a member list and a
/// group.
///
/// The one place a dot product is still computed during agglomeration, and
/// it happens only when two groups become adjacent through a third — at
/// which point their sub-threshold pairs, never summed because they were
/// never interesting, have to be accounted for.
fn cross(&self, members: &[usize], group: usize) -> Link {
let mut sum = 0.0_f64;
let mut count = 0.0_f64;
for &i in members {
for &j in &self.groups[group].members {
let cos = neighbours::dot(&self.faces[i].embedding, &self.faces[j].embedding);
let min_crop = self.faces[i].crop_px.min(self.faces[j].crop_px);
sum += self.cal.probability(cos, min_crop, 0.0) as f64;
count += 1.0;
}
}
Link { sum, count }
}
fn unlink(&mut self, a: usize, b: usize) {
self.links.remove(&key(a, b));
self.adjacent[a].remove(&b);
self.adjacent[b].remove(&a);
}
fn finish(self) -> Vec<Cluster> {
let mut out: Vec<Cluster> = self
.groups
.into_iter()
.filter(|g| g.alive)
.map(|g| Cluster {
members: g.members,
person: g.person,
})
.collect();
// Largest first: the People view shows the best-evidenced groups at the
// top.
out.sort_by(|x, y| {
y.members
.len()
.cmp(&x.members.len())
.then(x.members[0].cmp(&y.members[0]))
});
out
}
}
fn key(a: usize, b: usize) -> (usize, usize) {
if a < b {
(a, b)
} else {
(b, a)
}
}
/// Connected components of the above-threshold graph.
///
/// Faces in different components can never end up in one group, so each is a
/// separate and much smaller agglomeration. Returned with the members of each
/// component ascending, and the components themselves in order of their lowest
/// member — the determinism the merge order inherits.
fn components(n: usize, pairs: &[neighbours::Pair]) -> Vec<Vec<usize>> {
let mut parent: Vec<usize> = (0..n).collect();
fn find(parent: &mut [usize], mut x: usize) -> usize {
while parent[x] != x {
// Path halving: keeps the tree flat without a second pass.
parent[x] = parent[parent[x]];
x = parent[x];
}
x
}
for p in pairs {
let (ra, rb) = (find(&mut parent, p.i), find(&mut parent, p.j));
if ra != rb {
// Lowest root wins, so the representative of a component is
// reproducible rather than an artefact of union order.
let (lo, hi) = if ra < rb { (ra, rb) } else { (rb, ra) };
parent[hi] = lo;
}
}
let mut by_root: HashMap<usize, Vec<usize>> = HashMap::new();
for i in 0..n {
let r = find(&mut parent, i);
by_root.entry(r).or_default().push(i);
}
let mut out: Vec<Vec<usize>> = by_root.into_values().filter(|c| c.len() > 1).collect();
out.sort_unstable_by_key(|c| c[0]);
out
}
#[cfg(test)]
mod tests {
use super::*;
use crate::embedding::EMBEDDING_DIM;
/// An embedding a known cosine away from a base direction, built by mixing
/// two orthogonal unit vectors. Lets a test state "these two faces are 0.7
/// similar" and have it be exactly true.
fn at_cosine(identity: usize, cosine: f32) -> Vec<f32> {
let mut v = vec![0.0_f32; EMBEDDING_DIM];
let base = identity * 2;
let perp = identity * 2 + 1;
v[base] = cosine;
v[perp] = (1.0 - cosine * cosine).max(0.0).sqrt();
v
}
fn candidate(face: u64, image: u64, identity: usize, cosine: f32) -> Candidate {
Candidate {
face,
image,
embedding: at_cosine(identity, cosine),
crop_px: 150.0,
confirmed_person: None,
}
}
/// A calibration steep enough that the test's cosines are unambiguous:
/// 0.6 is near-certain, 0.1 is near-impossible.
fn cal() -> Calibration {
Calibration {
a: 30.0,
b: -30.0 * 0.35,
w_size: 0.0,
valid: true,
positive_pairs: 1000,
negative_pairs: 10_000,
}
}
#[test]
fn no_faces_makes_no_clusters() {
assert!(cluster(&[], &cal(), DEFAULT_MERGE_PROBABILITY).is_empty());
}
#[test]
fn similar_faces_from_different_photographs_group_together() {
let faces = vec![
candidate(1, 10, 0, 1.0),
candidate(2, 11, 0, 0.95),
candidate(3, 12, 0, 0.92),
];
let out = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY);
assert_eq!(out.len(), 1);
assert_eq!(out[0].members, vec![0, 1, 2]);
}
#[test]
fn dissimilar_faces_stay_apart() {
let faces = vec![candidate(1, 10, 0, 1.0), candidate(2, 11, 1, 1.0)];
let out = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY);
assert_eq!(out.len(), 2);
}
/// The cheapest defence against over-merging: two faces in one frame are
/// not the same person however similar the model finds them.
#[test]
fn two_faces_in_one_photograph_never_merge() {
// Identical embeddings — siblings, or a model that cannot tell them
// apart — but both in image 10.
let faces = vec![candidate(1, 10, 0, 1.0), candidate(2, 10, 0, 1.0)];
let out = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY);
assert_eq!(out.len(), 2, "co-occurring faces were merged");
}
/// And the constraint has to survive transitively: once a group holds a
/// face from image 10, no other group holding one from image 10 may join
/// it, even indirectly.
#[test]
fn the_co_occurrence_constraint_propagates_through_a_group() {
let faces = vec![
candidate(1, 10, 0, 1.0), // A, in the group photo
candidate(2, 10, 0, 1.0), // B, in the same group photo
candidate(3, 11, 0, 0.99), // A again, alone
];
let out = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY);
assert_eq!(out.len(), 2);
// Whichever of A/B absorbed face 2, the other stays out.
assert!(out.iter().any(|c| c.members.len() == 2));
assert!(out.iter().any(|c| c.members.len() == 1));
}
/// FR-CULL-10: a confirmation is user data and no inference overrides it.
#[test]
fn groups_confirmed_as_different_people_do_not_merge() {
let mut a = candidate(1, 10, 0, 1.0);
let mut b = candidate(2, 11, 0, 1.0);
a.confirmed_person = Some(100);
b.confirmed_person = Some(200);
let out = cluster(&[a, b], &cal(), DEFAULT_MERGE_PROBABILITY);
assert_eq!(out.len(), 2, "clustering overrode two user confirmations");
}
/// A face at a chosen cosine to identity 0 *and* to identity 1 at once —
/// the sibling geometry, which [`at_cosine`]'s per-identity subspaces
/// cannot express.
fn contested(to_first: f32, to_second: f32) -> Vec<f32> {
let mut v = vec![0.0_f32; EMBEDDING_DIM];
v[0] = to_first;
v[2] = to_second;
v[4] = (1.0 - to_first * to_first - to_second * to_second)
.max(0.0)
.sqrt();
v
}
/// The population the scoring tests share: two faces of one person, a
/// stranger, and a face that matches the person well and the stranger
/// weakly — weakly enough that it will never merge with them, which is
/// exactly the rival a within-group score cannot see.
fn with_a_rival() -> Vec<Candidate> {
let mut x = candidate(4, 13, 0, 0.0);
x.embedding = contested(0.45, 0.37);
// The stranger is *named*: only an identity the user has asserted
// competes for a face (crate::assign).
let mut stranger = candidate(3, 12, 1, 1.0);
stranger.confirmed_person = Some(7);
vec![
candidate(1, 10, 0, 1.0),
candidate(2, 11, 0, 1.0),
stranger,
x,
]
}
/// Asking for confidences must not move a single face. The scan runs at a
/// looser floor to find rivals, and the merge engine has to see exactly the
/// pairs it would have seen without them.
#[test]
fn scoring_does_not_change_the_groups() {
let faces = with_a_rival();
let plain = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY);
let scored = cluster_scored(&faces, &cal(), DEFAULT_MERGE_PROBABILITY);
assert_eq!(plain, scored.clusters);
}
/// The number the user is shown answers "which of these people", so a
/// second claimant has to lower it even when it is too weak to merge.
#[test]
fn a_face_two_identities_could_claim_is_reported_as_less_certain() {
let faces = with_a_rival();
let scored = cluster_scored(&faces, &cal(), DEFAULT_MERGE_PROBABILITY);
let alone = cluster_scored(&faces[..2], &cal(), DEFAULT_MERGE_PROBABILITY);
assert!(alone.confidence[0] > 0.99, "nobody else to be");
let contested = scored.confidence[3];
assert!(
(0.6..0.85).contains(&contested),
"a face with a second claimant: {contested}"
);
}
#[test]
fn a_suggestion_joins_the_person_its_group_is_anchored_to() {
let mut anchor = candidate(1, 10, 0, 1.0);
anchor.confirmed_person = Some(42);
let loose = candidate(2, 11, 0, 0.95);
let out = cluster(&[anchor, loose], &cal(), DEFAULT_MERGE_PROBABILITY);
assert_eq!(out.len(), 1);
assert_eq!(out[0].person, Some(42));
assert_eq!(out[0].members.len(), 2);
}
#[test]
fn an_unanchored_group_is_a_new_unnamed_person() {
let out = cluster(
&[candidate(1, 10, 0, 1.0), candidate(2, 11, 0, 0.95)],
&cal(),
DEFAULT_MERGE_PROBABILITY,
);
assert_eq!(out[0].person, None);
}
/// Average link rather than single link: one strong edge must not weld two
/// otherwise-dissimilar groups together. This is the family failure mode
/// FR-CULL-10 names.
#[test]
fn one_strong_edge_does_not_chain_two_groups_together() {
// Two tight pairs, with a single borderline link between them.
let faces = vec![
candidate(1, 10, 0, 1.00),
candidate(2, 11, 0, 0.99),
candidate(3, 12, 0, 0.42),
candidate(4, 13, 0, 0.40),
];
let out = cluster(&faces, &cal(), 0.99);
assert!(
out.len() >= 2,
"single-link chaining merged everything into {} cluster(s)",
out.len()
);
}
#[test]
fn clustering_is_deterministic() {
let faces = vec![
candidate(1, 10, 0, 1.0),
candidate(2, 11, 0, 0.96),
candidate(3, 12, 1, 1.0),
candidate(4, 13, 1, 0.97),
candidate(5, 14, 0, 0.94),
];
let a = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY);
let b = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY);
assert_eq!(a, b);
}
#[test]
fn clusters_come_back_largest_first() {
let faces = vec![
candidate(1, 10, 1, 1.0),
candidate(2, 11, 0, 1.0),
candidate(3, 12, 0, 0.97),
candidate(4, 13, 0, 0.95),
];
let out = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY);
assert_eq!(out[0].members.len(), 3);
assert_eq!(out[1].members.len(), 1);
}
/// Splitting is the inverse operation and must actually separate a group
/// that a looser threshold had merged.
#[test]
fn split_separates_a_group_that_a_looser_threshold_merged() {
let faces = vec![
candidate(1, 10, 0, 1.00),
candidate(2, 11, 0, 0.98),
candidate(3, 12, 0, 0.45),
candidate(4, 13, 0, 0.43),
];
// Loose: one person.
assert_eq!(cluster(&faces, &cal(), 0.5).len(), 1);
// Strict: the two sub-groups the user wants offered.
let parts = split(&faces, &cal(), 0.999);
assert!(parts.len() >= 2, "split produced {} group(s)", parts.len());
}
#[test]
fn split_ignores_the_anchor_so_a_mislabelled_person_can_be_taken_apart() {
let mut a = candidate(1, 10, 0, 1.0);
let mut b = candidate(2, 11, 0, 0.40);
a.confirmed_person = Some(7);
b.confirmed_person = Some(7);
let parts = split(&[a, b], &cal(), 0.99);
assert_eq!(parts.len(), 2);
}
/// The size term earns its place: the same cosine between two thumbnail-
/// sized faces should be less convincing than between two large ones.
#[test]
fn the_face_size_term_moves_the_probability() {
let sized = Calibration {
w_size: 0.5,
b: -30.0 * 0.35 - 0.5 * 7.0,
..cal()
};
let big = sized.probability(0.5, 300.0, 0.0);
let small = sized.probability(0.5, 40.0, 0.0);
assert!(big > small, "big {big} should beat small {small}");
}
// ── the fast engine against the obvious one ───────────────────────────
/// The original implementation, kept as the oracle.
///
/// Deliberately the naive version this module replaced: rescan every live
/// pair, score it from scratch over all cross pairs, merge the best,
/// repeat. It is the definition of the answer, and the only thing the
/// rewrite was allowed to change is how long it takes to get there.
fn reference(faces: &[Candidate], cal: &Calibration, min_probability: f32) -> Vec<Cluster> {
#[derive(Clone)]
struct G {
members: Vec<usize>,
images: HashSet<u64>,
person: Option<u64>,
alive: bool,
}
if faces.is_empty() {
return Vec::new();
}
let n = faces.len();
let mut groups: Vec<G> = faces
.iter()
.enumerate()
.map(|(i, f)| G {
members: vec![i],
images: HashSet::from([f.image]),
person: f.confirmed_person,
alive: true,
})
.collect();
let mut cos = vec![0.0_f32; n * n];
for i in 0..n {
for j in i + 1..n {
let c = neighbours::dot(&faces[i].embedding, &faces[j].embedding);
cos[i * n + j] = c;
cos[j * n + i] = c;
}
}
let linkable = |a: &G, b: &G| {
if let (Some(pa), Some(pb)) = (a.person, b.person) {
if pa != pb {
return false;
}
}
a.images.is_disjoint(&b.images)
};
let average = |a: &G, b: &G| {
let mut sum = 0.0_f32;
let mut count = 0.0_f32;
for &i in &a.members {
for &j in &b.members {
let min_crop = faces[i].crop_px.min(faces[j].crop_px);
sum += cal.probability(cos[i * n + j], min_crop, 0.0);
count += 1.0;
}
}
if count == 0.0 {
0.0
} else {
sum / count
}
};
loop {
let mut best: Option<(f32, usize, usize)> = None;
for a in 0..n {
if !groups[a].alive {
continue;
}
for b in a + 1..n {
if !groups[b].alive || !linkable(&groups[a], &groups[b]) {
continue;
}
let p = average(&groups[a], &groups[b]);
if p >= min_probability && best.is_none_or(|(bp, _, _)| p > bp) {
best = Some((p, a, b));
}
}
}
let Some((_, a, b)) = best else { break };
let taken = groups[b].clone();
groups[b].alive = false;
groups[a].members.extend(taken.members);
groups[a].images.extend(taken.images);
groups[a].person = groups[a].person.or(taken.person);
}
let mut out: Vec<Cluster> = groups
.into_iter()
.filter(|g| g.alive)
.map(|mut g| {
g.members.sort_unstable();
Cluster {
members: g.members,
person: g.person,
}
})
.collect();
out.sort_by(|x, y| {
y.members
.len()
.cmp(&x.members.len())
.then(x.members[0].cmp(&y.members[0]))
});
out
}
/// `people` identities of `per` faces, each face in its own photograph,
/// spread either side of the threshold so the population has genuine
/// near-misses rather than obvious answers.
fn population(people: usize, per: usize) -> Vec<Candidate> {
let mut out = Vec::new();
let mut image = 0u64;
for p in 0..people {
for m in 0..per {
// Walks down through the merge boundary as m grows, so some
// members join their group and some do not.
let cosine = 1.0 - (m as f32) * 0.035;
out.push(Candidate {
face: out.len() as u64,
image,
embedding: at_cosine(p, cosine),
crop_px: 60.0 + ((out.len() % 11) as f32) * 25.0,
confirmed_person: None,
});
image += 1;
}
}
out
}
/// The point of the rewrite: same clusters, less work. A disagreement here
/// is the rewrite being wrong, not the reference being slow.
#[test]
fn the_fast_engine_agrees_with_the_reference() {
let faces = population(40, 6);
assert_eq!(
cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY),
reference(&faces, &cal(), DEFAULT_MERGE_PROBABILITY),
);
}
/// The two structural constraints are the ones a sparse graph could
/// plausibly break, so they get their own comparison with anchors and
/// co-occurrence in play.
#[test]
fn the_fast_engine_agrees_with_the_reference_under_constraints() {
let mut faces = population(30, 6);
// Some faces share a photograph, so cannot-link has to propagate
// through groups that formed for other reasons.
for i in (0..faces.len()).step_by(7) {
faces[i].image = 900 + (i as u64 % 4);
}
// And some carry confirmations, including two of different people that
// must never be brought together.
for (n, i) in (0..faces.len()).step_by(11).enumerate() {
faces[i].confirmed_person = Some(1 + (n as u64 % 3));
}
assert_eq!(
cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY),
reference(&faces, &cal(), DEFAULT_MERGE_PROBABILITY),
);
}
/// The size term makes the merge boundary depend on the pair, which is the
/// case the sparse pre-filter has to be built carefully to preserve.
#[test]
fn the_fast_engine_agrees_with_the_reference_with_a_size_term() {
let faces = population(30, 6);
let sized = Calibration {
w_size: 0.4,
b: -30.0 * 0.35 - 0.4 * 7.0,
..cal()
};
assert_eq!(
cluster(&faces, &sized, DEFAULT_MERGE_PROBABILITY),
reference(&faces, &sized, DEFAULT_MERGE_PROBABILITY),
);
}
/// Determinism has to hold at a size where the indexed neighbour search is
/// in play, not just on the handful of faces the small cases use.
#[test]
fn clustering_is_deterministic_at_scale() {
let faces = population(200, 6);
assert!(
faces.len() > 1024,
"population is below the indexing cutoff"
);
assert_eq!(
cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY),
cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY),
);
}
/// A face that matches nobody is left alone rather than being swept into
/// the nearest group, and costs nothing to establish — it is in no
/// component at all.
#[test]
fn a_face_matching_nothing_stays_on_its_own() {
let mut faces = population(5, 4);
faces.push(Candidate {
face: 999,
image: 5_000,
embedding: at_cosine(200, 1.0),
crop_px: 150.0,
confirmed_person: None,
});
let out = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY);
let last = faces.len() - 1;
assert!(
out.iter().any(|c| c.members == vec![last]),
"the outlier was absorbed"
);
}
}