//! Grouping faces into people (docs/faces.md §9, FR-CULL-10). //! //! Model-free: this is arithmetic over embeddings, and it is where the //! subsystem's accuracy actually lives, so it is testable with no weights on //! the machine. //! //! # Constraints, not just a threshold //! //! FR-CULL-10 warns that clustering will over-merge on siblings, on parents and //! children, and on the same person a decade apart. Two structural defences, //! both cheaper than a better threshold: //! //! **Cannot-link on co-occurrence.** Two faces in the same photograph are never //! merged. It is the same observation [`crate::calibrate`] mines for free //! negatives, used here as a hard constraint, and it is the single cheapest //! defence against over-merging that exists. //! //! **Confirmed faces are anchors.** A confirmation is user data (FR-CULL-12) //! and clustering never moves it. Two groups holding confirmations of //! *different* people cannot merge, whatever their similarity says. //! //! # Average link, not single link //! //! Single-link chains: one bad edge welds two identities together, and it is //! the documented way face clustering fails on families. Average link asks //! whether the *groups* are similar, which one outlier cannot force. //! //! # How it runs, and why the obvious way does not //! //! The first implementation of this was the textbook one: compute every //! pairwise cosine, then repeatedly scan all live group pairs, score each with //! average link, and merge the best. It is correct, it is twenty lines, and on //! a real library it does not finish. //! //! The reason is that the scan is inside the loop. Each merge rescans every //! surviving pair — `O(g²)` of them — and each score is recomputed from //! scratch over every cross pair, `O(|A|·|B|)`. With 1,813 faces that is //! roughly 1.6 million pair scores per merge and some 700 merges to do; the //! window simply stops responding, which is what a user reports as "Regroup is //! broken". At 25,000 faces it is not slow, it is impossible. //! //! Three changes, none of which alter the answer: //! //! **Only above-threshold pairs can ever matter.** An average that reaches the //! threshold must have at least one term at or above it, so two groups with no //! qualifying pair between them can never merge — not now and not after any //! sequence of merges, since merging only adds terms. [`crate::neighbours`] //! produces exactly that sparse pair list, and everything below works on it. //! On the library above it is 7,875 pairs rather than 1.6 million. //! //! **Merges cannot cross components.** Groups only ever merge along those //! pairs, so the connected components of that graph are independent problems. //! A library of four hundred people becomes four hundred small agglomerations //! instead of one large one, and the quadratic term is paid per component. //! //! **Average link is additive.** `sum(A ∪ B, C) = sum(A, C) + sum(B, C)`, so a //! merged group's scores follow from the two it came from by addition — the //! Lance-Williams update. Kept as running `(sum, count)` per adjacent pair, a //! score costs one division instead of a nested loop, and a binary heap with //! lazy invalidation replaces the rescan. //! //! The output is unchanged, deliberately and testably so: `the_fast_engine_ //! agrees_with_the_reference` runs both over the same population and asserts //! the clusters are identical. use std::cmp::Ordering; use std::collections::{BinaryHeap, HashMap, HashSet}; use crate::calibrate::Calibration; use crate::neighbours::{self, Faces}; /// Probability above which two groups are judged the same person. /// /// Stated as a probability and not a cosine, because FR-CULL-9 forbids /// thresholding a bare similarity anywhere in this subsystem. /// /// # Why 0.80 /// /// It was 0.90, and 0.90 left most of a real library ungrouped. Measured over /// the 1,813-face reference library, with `dr-ui`'s `face_index --tune`: /// /// | P | cosine | groups | faces grouped | largest group | /// |---|---|---|---|---| /// | 0.95 | 0.449 | 311 | 62% | 51 | /// | 0.90 | 0.403 | 316 | 67% | 51 | /// | 0.85 | 0.374 | 318 | 70% | 57 | /// | **0.80** | **0.353** | **328** | **74%** | **69** | /// | 0.75 | 0.335 | 327 | 77% | 69 | /// | 0.70 | 0.319 | 326 | 79% | 81 | /// | 0.50 | 0.267 | 303 | 85% | 90 | /// /// The count of *groups* is the signal, not the count of grouped faces. Loosen /// from 0.95 and it climbs: real people are being assembled out of fragments. /// It peaks at 0.80 and then falls, and a falling group count while the grouped /// faces keep rising is the shape of over-merging — separate identities being /// welded, which is the failure FR-CULL-10 warns about and the one the user /// cannot easily undo by hand. /// /// So: the loosest setting that is still building people rather than melting /// them together. A third more of the library gets grouped than at 0.90, and /// the largest group grows by eighteen faces rather than by forty. /// /// This is a *default*, not a constant of nature — the numbers above are one /// library, and `--tune` reruns the table on any other. pub const DEFAULT_MERGE_PROBABILITY: f32 = 0.80; /// A face presented to the clusterer. /// /// Ids are opaque `u64`s rather than catalog types: this crate has no business /// knowing what a `FaceId` means, and the caller does the translation. #[derive(Debug, Clone)] pub struct Candidate { pub face: u64, /// Which photograph it came from — the cannot-link key. pub image: u64, /// L2-normalised, `EMBEDDING_DIM` long. pub embedding: Vec, /// Source pixels across the aligned crop, for the calibration's size term. pub crop_px: f32, /// The person this face is *confirmed* to be, if any. /// /// Suggestions are deliberately not passed here. They are this function's /// own previous output, and feeding them back in would let a guess harden /// into a fact across successive passes. pub confirmed_person: Option, } /// One group of faces the clusterer believes are one person. #[derive(Debug, Clone, PartialEq)] pub struct Cluster { /// Indices into the input slice. pub members: Vec, /// The person this group is already known to be, from its anchors. /// /// `Some` means the group contains confirmed faces and the suggestions in /// it attach to that existing person. `None` is a new unnamed group. pub person: Option, } /// Group faces into people. /// /// `min_probability` is compared against the calibrated average-link /// probability between two groups. Deterministic: the same input yields the /// same clusters, because the merge order is by score with the index pair as /// the tiebreak. pub fn cluster(faces: &[Candidate], cal: &Calibration, min_probability: f32) -> Vec { if faces.is_empty() { return Vec::new(); } let columns = Columns::of(faces); // Every pair that could ever contribute to a merge. See the module note on // why nothing outside this list can matter. let pairs = neighbours::above_threshold(&columns.view(), cal, min_probability); build(faces, cal, min_probability, &pairs) } /// Groups, and how confident each face's placement is. /// /// The second half is [`crate::assign`]'s share, not the pairwise probability /// that put the face in the group — see that module for why the two are /// different questions. #[derive(Debug, Clone, PartialEq)] pub struct Grouping { pub clusters: Vec, /// Indexed like the input faces. 0 for a face in no group. pub confidence: Vec, } /// Group faces into people, and score each placement against its rivals. /// /// What a caller writing suggestions into a catalog wants: [`cluster`] answers /// *which person*, this answers *and how sure*. /// /// One scan, two thresholds. The similarity scan is the expensive part of the /// whole subsystem and running it twice — once to merge, once to find rivals — /// would double the cost of regrouping a library. So it runs once at the looser /// of the two floors, and the merge engine takes the subset at or above /// `min_probability`. That subset is identical, pair for pair and in the same /// order, to what a scan at `min_probability` would have produced, so grouping /// is unchanged by scoring being asked for: `scoring_does_not_change_the_ /// groups` holds it to that. pub fn cluster_scored(faces: &[Candidate], cal: &Calibration, min_probability: f32) -> Grouping { if faces.is_empty() { return Grouping { clusters: Vec::new(), confidence: Vec::new(), }; } let columns = Columns::of(faces); let evidence = neighbours::above_threshold( &columns.view(), cal, min_probability.min(crate::assign::RIVAL_FLOOR), ); let merges: Vec = evidence .iter() .copied() .filter(|p| p.probability >= min_probability) .collect(); let clusters = build(faces, cal, min_probability, &merges); let confidence = crate::assign::identity_shares( faces.len(), &clusters, &evidence, crate::assign::TOP_MATCHES, ); Grouping { clusters, confidence, } } /// The three arrays [`neighbours::Faces`] borrows, owned. /// /// [`neighbours`] takes parallel slices rather than candidates on purpose — it /// has no business knowing what a person is — so somebody has to hold the /// columns. Both entry points do, identically, which is the only reason this is /// a type and not three locals. struct Columns { /// Every embedding end to end — see [`Faces::embeddings`] for why flat. embeddings: Vec, dim: usize, crop_px: Vec, images: Vec, } impl Columns { fn of(faces: &[Candidate]) -> Self { // Ragged input would index the wrong row for every face after the odd // one, so the widest wins and short rows are padded with zeros: a // zero-padded row scores lower against everything, which is the safe // direction. It does not happen — one model, one dimension — and it is // handled rather than trusted because the failure would be silent. let dim = faces.iter().map(|f| f.embedding.len()).max().unwrap_or(0); let mut embeddings = Vec::with_capacity(faces.len() * dim); for f in faces { embeddings.extend_from_slice(&f.embedding); embeddings.resize(embeddings.len() + dim - f.embedding.len(), 0.0); } Self { embeddings, dim, crop_px: faces.iter().map(|f| f.crop_px).collect(), images: faces.iter().map(|f| f.image).collect(), } } fn view(&self) -> Faces<'_> { Faces { embeddings: &self.embeddings, dim: self.dim, crop_px: &self.crop_px, images: &self.images, } } } fn build( faces: &[Candidate], cal: &Calibration, min_probability: f32, pairs: &[neighbours::Pair], ) -> Vec { let mut engine = Engine::new(faces, cal, min_probability); let parts = components(faces.len(), pairs); for (component, edges) in parts.members.iter().zip(&parts.edges) { engine.agglomerate(component, edges); } engine.finish() } /// Split one person's faces into the groups a raised threshold separates them /// into. /// /// FR-CULL-10 requires splitting to be as easy as merging, and a split that /// hands the user a pile of loose faces to re-sort is not that. This re-runs /// the same agglomeration at a stricter probability so the user is offered /// coherent sub-groups to pull apart. /// /// Anchors are ignored here on purpose: every face in the input is already /// nominally the same person, so honouring the anchors would refuse to split /// anything. pub fn split(faces: &[Candidate], cal: &Calibration, min_probability: f32) -> Vec { let anchorless: Vec = faces .iter() .cloned() .map(|mut f| { f.confirmed_person = None; f }) .collect(); cluster(&anchorless, cal, min_probability) } // ── the merge engine ────────────────────────────────────────────────────── /// One live group, identified throughout by the index of its lowest member. #[derive(Debug)] struct Group { /// Ascending, always — [`Engine::cross`] sums in this order, and a stable /// order is what makes the floating-point total reproducible. members: Vec, images: HashSet, person: Option, alive: bool, /// Bumped on every merge, so heap entries naming an older state can be /// recognised and dropped instead of acted on. version: u64, } /// Running average-link state for one adjacent pair of groups. /// /// `sum` is over **every** cross pair, not only the above-threshold ones — /// average link is an average over all of them, and counting only the /// qualifying pairs would report a similarity no group actually has. #[derive(Debug, Clone, Copy)] struct Link { sum: f64, count: f64, } impl Link { fn probability(&self) -> f32 { if self.count == 0.0 { 0.0 } else { (self.sum / self.count) as f32 } } } /// A candidate merge, waiting in the heap. #[derive(Debug, Clone, Copy)] struct Pending { probability: f32, a: usize, b: usize, /// Group versions when this was pushed. A mismatch on pop means a merge /// has happened since and a fresher entry for this pair is already queued. va: u64, vb: u64, } impl PartialEq for Pending { fn eq(&self, other: &Self) -> bool { self.cmp(other) == Ordering::Equal } } impl Eq for Pending {} impl PartialOrd for Pending { fn partial_cmp(&self, other: &Self) -> Option { Some(self.cmp(other)) } } impl Ord for Pending { /// Greatest pops first, so: highest probability, and on a tie the lowest /// index pair. That tiebreak is not cosmetic — it is what the old /// ascending scan did, and it is the whole of the determinism guarantee. fn cmp(&self, other: &Self) -> Ordering { self.probability .total_cmp(&other.probability) .then_with(|| other.a.cmp(&self.a)) .then_with(|| other.b.cmp(&self.b)) } } struct Engine<'a> { faces: &'a [Candidate], cal: &'a Calibration, /// The same kernel [`neighbours`] scans with. [`Engine::cross`] is the one /// place a dot product is computed during agglomeration, and it was using /// the portable loop while the scan beside it had the machine's SIMD. dot: neighbours::DotFn, min_probability: f32, groups: Vec, links: HashMap<(usize, usize), Link>, /// Adjacency, as group ids. Kept alongside `links` so a merge can find /// everything it has to update without scanning the whole map. adjacent: Vec>, } impl<'a> Engine<'a> { fn new(faces: &'a [Candidate], cal: &'a Calibration, min_probability: f32) -> Self { let groups = faces .iter() .enumerate() .map(|(i, f)| Group { members: vec![i], images: HashSet::from([f.image]), person: f.confirmed_person, alive: true, version: 0, }) .collect(); Self { faces, cal, dot: neighbours::fastest_dot(), min_probability, groups, links: HashMap::new(), adjacent: vec![HashSet::new(); faces.len()], } } /// Agglomerate one connected component to exhaustion. fn agglomerate(&mut self, component: &[usize], pairs: &[neighbours::Pair]) { if component.len() < 2 { return; } let mut heap = BinaryHeap::with_capacity(pairs.len()); for p in pairs { self.links.insert( key(p.i, p.j), Link { sum: p.probability as f64, count: 1.0, }, ); self.adjacent[p.i].insert(p.j); self.adjacent[p.j].insert(p.i); heap.push(Pending { probability: p.probability, a: p.i.min(p.j), b: p.i.max(p.j), va: 0, vb: 0, }); } while let Some(top) = heap.pop() { let Pending { probability, a, b, va, vb, } = top; // Stale: one side has merged since this was queued, and the // replacement entry is already in the heap. if !self.groups[a].alive || !self.groups[b].alive || self.groups[a].version != va || self.groups[b].version != vb { continue; } // The heap is ordered by probability, so the first entry below the // bar means nothing left in this component can reach it. if probability < self.min_probability { break; } if !self.can_link(a, b) { // Never becomes possible again: images only accumulate and an // anchor is never given up, so drop the pair for good. self.unlink(a, b); continue; } self.merge(a, b, &mut heap); } } /// Whether two groups are allowed to merge at all, before similarity is /// asked. fn can_link(&self, a: usize, b: usize) -> bool { let (ga, gb) = (&self.groups[a], &self.groups[b]); // Two confirmations of different people. The user has said these are // not the same person, and no similarity overrides that. if let (Some(pa), Some(pb)) = (ga.person, gb.person) { if pa != pb { return false; } } // Co-occurrence: a photograph containing a face from each group means // the two faces are in the same frame, so they are not the same person. ga.images.is_disjoint(&gb.images) } /// Fold `b` into `a` and re-score everything that touched either. fn merge(&mut self, a: usize, b: usize, heap: &mut BinaryHeap) { // Sorted, and deduplicated by the set: `a` and `b` may share // neighbours, and each must be visited once. Sorting is what keeps the // floating-point sums identical from run to run. let mut touched: Vec = self.adjacent[a] .union(&self.adjacent[b]) .copied() .filter(|&c| c != a && c != b && self.groups[c].alive) .collect(); touched.sort_unstable(); // Take the pair sums before the groups change underneath them. let carried: Vec<(usize, Option, Option)> = touched .iter() .map(|&c| { ( c, self.links.get(&key(a, c)).copied(), self.links.get(&key(b, c)).copied(), ) }) .collect(); // a's own members, before b's are folded in. A missing a-side sum has // to be computed over these and not over the merged list, or b's // contribution would be counted twice. let a_members = self.groups[a].members.clone(); // Absorb b into a. let taken = std::mem::replace( &mut self.groups[b], Group { members: Vec::new(), images: HashSet::new(), person: None, alive: false, version: 0, }, ); { let ga = &mut self.groups[a]; ga.members.extend(taken.members.iter().copied()); ga.members.sort_unstable(); ga.images.extend(taken.images.iter().copied()); // At most one side carries a person: `can_link` refuses a merge of // two groups anchored to different people, so this cannot silently // discard one of them. ga.person = ga.person.or(taken.person); ga.version += 1; } // b's own links are gone with it. for c in self.adjacent[b].clone() { self.links.remove(&key(b, c)); self.adjacent[c].remove(&b); } self.adjacent[b].clear(); self.links.remove(&key(a, b)); self.adjacent[a].remove(&b); for (c, from_a, from_b) in carried { // Dropping a pair the constraints now forbid saves computing a // score for a merge that can never happen — which for a newly // adjacent side is a real cost, not a bookkeeping one. if !self.can_link(a, c) { self.unlink(a, c); continue; } // A side with no stored link was not adjacent before, so its cross // pairs were all below threshold and were never summed. They still // belong in the average, so they are computed now — once, after // which the additive update carries them forward. let from_a = from_a.unwrap_or_else(|| self.cross(&a_members, c)); let from_b = from_b.unwrap_or_else(|| self.cross(&taken.members, c)); let merged = Link { sum: from_a.sum + from_b.sum, count: from_a.count + from_b.count, }; self.links.insert(key(a, c), merged); self.adjacent[a].insert(c); self.adjacent[c].insert(a); heap.push(Pending { probability: merged.probability(), a: a.min(c), b: a.max(c), va: self.groups[a.min(c)].version, vb: self.groups[a.max(c)].version, }); } } /// Exact `(sum, count)` over every cross pair between a member list and a /// group. /// /// The one place a dot product is still computed during agglomeration, and /// it happens only when two groups become adjacent through a third — at /// which point their sub-threshold pairs, never summed because they were /// never interesting, have to be accounted for. fn cross(&self, members: &[usize], group: usize) -> Link { let mut sum = 0.0_f64; let mut count = 0.0_f64; for &i in members { for &j in &self.groups[group].members { let cos = (self.dot)(&self.faces[i].embedding, &self.faces[j].embedding); let min_crop = self.faces[i].crop_px.min(self.faces[j].crop_px); sum += self.cal.probability(cos, min_crop, 0.0) as f64; count += 1.0; } } Link { sum, count } } fn unlink(&mut self, a: usize, b: usize) { self.links.remove(&key(a, b)); self.adjacent[a].remove(&b); self.adjacent[b].remove(&a); } fn finish(self) -> Vec { let mut out: Vec = self .groups .into_iter() .filter(|g| g.alive) .map(|g| Cluster { members: g.members, person: g.person, }) .collect(); // Largest first: the People view shows the best-evidenced groups at the // top. out.sort_by(|x, y| { y.members .len() .cmp(&x.members.len()) .then(x.members[0].cmp(&y.members[0])) }); out } } fn key(a: usize, b: usize) -> (usize, usize) { if a < b { (a, b) } else { (b, a) } } /// Connected components of the above-threshold graph. /// /// Faces in different components can never end up in one group, so each is a /// separate and much smaller agglomeration. Returned with the members of each /// component ascending, and the components themselves in order of their lowest /// member — the determinism the merge order inherits. fn components(n: usize, pairs: &[neighbours::Pair]) -> Components { let mut parent: Vec = (0..n).collect(); fn find(parent: &mut [usize], mut x: usize) -> usize { while parent[x] != x { // Path halving: keeps the tree flat without a second pass. parent[x] = parent[parent[x]]; x = parent[x]; } x } for p in pairs { let (ra, rb) = (find(&mut parent, p.i), find(&mut parent, p.j)); if ra != rb { // Lowest root wins, so the representative of a component is // reproducible rather than an artefact of union order. let (lo, hi) = if ra < rb { (ra, rb) } else { (rb, ra) }; parent[hi] = lo; } } let mut by_root: HashMap> = HashMap::new(); for i in 0..n { let r = find(&mut parent, i); by_root.entry(r).or_default().push(i); } let mut members: Vec> = by_root.into_values().filter(|c| c.len() > 1).collect(); members.sort_unstable_by_key(|c| c[0]); // Each pair filed under its component, in one pass. // // **This is not bookkeeping, it is the cost of the whole stage.** Each // component used to scan the entire pair list for the ones that were its // own: 475 components against 804,499 pairs on the reference library, 382 // million set lookups, and 3.06 s of a 5.93 s regroup spent before a single // merge was considered. A pair can only ever join two faces of one // component — that is what a component *is* — so one pass over the list // places every pair exactly where it is needed. let mut slot_of = vec![usize::MAX; n]; for (slot, c) in members.iter().enumerate() { for &m in c { slot_of[m] = slot; } } let mut edges: Vec> = vec![Vec::new(); members.len()]; for p in pairs { // Ordering within a component follows the global order, which is what // the seeded heap's tiebreak — and so the determinism promise — rests // on. if let Some(slot) = slot_of.get(p.i).copied().filter(|&s| s != usize::MAX) { edges[slot].push(*p); } } Components { members, edges } } /// The connected components of the pair graph, each with the pairs inside it. /// /// The two halves are parallel: `members[k]` and `edges[k]` describe the same /// component. struct Components { members: Vec>, edges: Vec>, } #[cfg(test)] mod tests { use super::*; use crate::embedding::EMBEDDING_DIM; /// An embedding a known cosine away from a base direction, built by mixing /// two orthogonal unit vectors. Lets a test state "these two faces are 0.7 /// similar" and have it be exactly true. fn at_cosine(identity: usize, cosine: f32) -> Vec { let mut v = vec![0.0_f32; EMBEDDING_DIM]; let base = identity * 2; let perp = identity * 2 + 1; v[base] = cosine; v[perp] = (1.0 - cosine * cosine).max(0.0).sqrt(); v } fn candidate(face: u64, image: u64, identity: usize, cosine: f32) -> Candidate { Candidate { face, image, embedding: at_cosine(identity, cosine), crop_px: 150.0, confirmed_person: None, } } /// A calibration steep enough that the test's cosines are unambiguous: /// 0.6 is near-certain, 0.1 is near-impossible. fn cal() -> Calibration { Calibration { a: 30.0, b: -30.0 * 0.35, w_size: 0.0, valid: true, positive_pairs: 1000, negative_pairs: 10_000, } } #[test] fn no_faces_makes_no_clusters() { assert!(cluster(&[], &cal(), DEFAULT_MERGE_PROBABILITY).is_empty()); } #[test] fn similar_faces_from_different_photographs_group_together() { let faces = vec![ candidate(1, 10, 0, 1.0), candidate(2, 11, 0, 0.95), candidate(3, 12, 0, 0.92), ]; let out = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY); assert_eq!(out.len(), 1); assert_eq!(out[0].members, vec![0, 1, 2]); } #[test] fn dissimilar_faces_stay_apart() { let faces = vec![candidate(1, 10, 0, 1.0), candidate(2, 11, 1, 1.0)]; let out = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY); assert_eq!(out.len(), 2); } /// The cheapest defence against over-merging: two faces in one frame are /// not the same person however similar the model finds them. #[test] fn two_faces_in_one_photograph_never_merge() { // Identical embeddings — siblings, or a model that cannot tell them // apart — but both in image 10. let faces = vec![candidate(1, 10, 0, 1.0), candidate(2, 10, 0, 1.0)]; let out = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY); assert_eq!(out.len(), 2, "co-occurring faces were merged"); } /// And the constraint has to survive transitively: once a group holds a /// face from image 10, no other group holding one from image 10 may join /// it, even indirectly. #[test] fn the_co_occurrence_constraint_propagates_through_a_group() { let faces = vec![ candidate(1, 10, 0, 1.0), // A, in the group photo candidate(2, 10, 0, 1.0), // B, in the same group photo candidate(3, 11, 0, 0.99), // A again, alone ]; let out = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY); assert_eq!(out.len(), 2); // Whichever of A/B absorbed face 2, the other stays out. assert!(out.iter().any(|c| c.members.len() == 2)); assert!(out.iter().any(|c| c.members.len() == 1)); } /// FR-CULL-10: a confirmation is user data and no inference overrides it. #[test] fn groups_confirmed_as_different_people_do_not_merge() { let mut a = candidate(1, 10, 0, 1.0); let mut b = candidate(2, 11, 0, 1.0); a.confirmed_person = Some(100); b.confirmed_person = Some(200); let out = cluster(&[a, b], &cal(), DEFAULT_MERGE_PROBABILITY); assert_eq!(out.len(), 2, "clustering overrode two user confirmations"); } /// A face at a chosen cosine to identity 0 *and* to identity 1 at once — /// the sibling geometry, which [`at_cosine`]'s per-identity subspaces /// cannot express. fn contested(to_first: f32, to_second: f32) -> Vec { let mut v = vec![0.0_f32; EMBEDDING_DIM]; v[0] = to_first; v[2] = to_second; v[4] = (1.0 - to_first * to_first - to_second * to_second) .max(0.0) .sqrt(); v } /// The population the scoring tests share: two faces of one person, a /// stranger, and a face that matches the person well and the stranger /// weakly — weakly enough that it will never merge with them, which is /// exactly the rival a within-group score cannot see. fn with_a_rival() -> Vec { let mut x = candidate(4, 13, 0, 0.0); x.embedding = contested(0.45, 0.37); // The stranger is *named*: only an identity the user has asserted // competes for a face (crate::assign). let mut stranger = candidate(3, 12, 1, 1.0); stranger.confirmed_person = Some(7); vec![ candidate(1, 10, 0, 1.0), candidate(2, 11, 0, 1.0), stranger, x, ] } /// Asking for confidences must not move a single face. The scan runs at a /// looser floor to find rivals, and the merge engine has to see exactly the /// pairs it would have seen without them. #[test] fn scoring_does_not_change_the_groups() { let faces = with_a_rival(); let plain = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY); let scored = cluster_scored(&faces, &cal(), DEFAULT_MERGE_PROBABILITY); assert_eq!(plain, scored.clusters); } /// The number the user is shown answers "which of these people", so a /// second claimant has to lower it even when it is too weak to merge. #[test] fn a_face_two_identities_could_claim_is_reported_as_less_certain() { let faces = with_a_rival(); let scored = cluster_scored(&faces, &cal(), DEFAULT_MERGE_PROBABILITY); let alone = cluster_scored(&faces[..2], &cal(), DEFAULT_MERGE_PROBABILITY); assert!(alone.confidence[0] > 0.99, "nobody else to be"); let contested = scored.confidence[3]; assert!( (0.6..0.85).contains(&contested), "a face with a second claimant: {contested}" ); } #[test] fn a_suggestion_joins_the_person_its_group_is_anchored_to() { let mut anchor = candidate(1, 10, 0, 1.0); anchor.confirmed_person = Some(42); let loose = candidate(2, 11, 0, 0.95); let out = cluster(&[anchor, loose], &cal(), DEFAULT_MERGE_PROBABILITY); assert_eq!(out.len(), 1); assert_eq!(out[0].person, Some(42)); assert_eq!(out[0].members.len(), 2); } #[test] fn an_unanchored_group_is_a_new_unnamed_person() { let out = cluster( &[candidate(1, 10, 0, 1.0), candidate(2, 11, 0, 0.95)], &cal(), DEFAULT_MERGE_PROBABILITY, ); assert_eq!(out[0].person, None); } /// Average link rather than single link: one strong edge must not weld two /// otherwise-dissimilar groups together. This is the family failure mode /// FR-CULL-10 names. #[test] fn one_strong_edge_does_not_chain_two_groups_together() { // Two tight pairs, with a single borderline link between them. let faces = vec![ candidate(1, 10, 0, 1.00), candidate(2, 11, 0, 0.99), candidate(3, 12, 0, 0.42), candidate(4, 13, 0, 0.40), ]; let out = cluster(&faces, &cal(), 0.99); assert!( out.len() >= 2, "single-link chaining merged everything into {} cluster(s)", out.len() ); } #[test] fn clustering_is_deterministic() { let faces = vec![ candidate(1, 10, 0, 1.0), candidate(2, 11, 0, 0.96), candidate(3, 12, 1, 1.0), candidate(4, 13, 1, 0.97), candidate(5, 14, 0, 0.94), ]; let a = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY); let b = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY); assert_eq!(a, b); } #[test] fn clusters_come_back_largest_first() { let faces = vec![ candidate(1, 10, 1, 1.0), candidate(2, 11, 0, 1.0), candidate(3, 12, 0, 0.97), candidate(4, 13, 0, 0.95), ]; let out = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY); assert_eq!(out[0].members.len(), 3); assert_eq!(out[1].members.len(), 1); } /// Splitting is the inverse operation and must actually separate a group /// that a looser threshold had merged. #[test] fn split_separates_a_group_that_a_looser_threshold_merged() { let faces = vec![ candidate(1, 10, 0, 1.00), candidate(2, 11, 0, 0.98), candidate(3, 12, 0, 0.45), candidate(4, 13, 0, 0.43), ]; // Loose: one person. assert_eq!(cluster(&faces, &cal(), 0.5).len(), 1); // Strict: the two sub-groups the user wants offered. let parts = split(&faces, &cal(), 0.999); assert!(parts.len() >= 2, "split produced {} group(s)", parts.len()); } #[test] fn split_ignores_the_anchor_so_a_mislabelled_person_can_be_taken_apart() { let mut a = candidate(1, 10, 0, 1.0); let mut b = candidate(2, 11, 0, 0.40); a.confirmed_person = Some(7); b.confirmed_person = Some(7); let parts = split(&[a, b], &cal(), 0.99); assert_eq!(parts.len(), 2); } /// The size term earns its place: the same cosine between two thumbnail- /// sized faces should be less convincing than between two large ones. #[test] fn the_face_size_term_moves_the_probability() { let sized = Calibration { w_size: 0.5, b: -30.0 * 0.35 - 0.5 * 7.0, ..cal() }; let big = sized.probability(0.5, 300.0, 0.0); let small = sized.probability(0.5, 40.0, 0.0); assert!(big > small, "big {big} should beat small {small}"); } // ── the fast engine against the obvious one ─────────────────────────── /// The original implementation, kept as the oracle. /// /// Deliberately the naive version this module replaced: rescan every live /// pair, score it from scratch over all cross pairs, merge the best, /// repeat. It is the definition of the answer, and the only thing the /// rewrite was allowed to change is how long it takes to get there. fn reference(faces: &[Candidate], cal: &Calibration, min_probability: f32) -> Vec { #[derive(Clone)] struct G { members: Vec, images: HashSet, person: Option, alive: bool, } if faces.is_empty() { return Vec::new(); } let n = faces.len(); let mut groups: Vec = faces .iter() .enumerate() .map(|(i, f)| G { members: vec![i], images: HashSet::from([f.image]), person: f.confirmed_person, alive: true, }) .collect(); let mut cos = vec![0.0_f32; n * n]; for i in 0..n { for j in i + 1..n { let c = neighbours::dot(&faces[i].embedding, &faces[j].embedding); cos[i * n + j] = c; cos[j * n + i] = c; } } let linkable = |a: &G, b: &G| { if let (Some(pa), Some(pb)) = (a.person, b.person) { if pa != pb { return false; } } a.images.is_disjoint(&b.images) }; let average = |a: &G, b: &G| { let mut sum = 0.0_f32; let mut count = 0.0_f32; for &i in &a.members { for &j in &b.members { let min_crop = faces[i].crop_px.min(faces[j].crop_px); sum += cal.probability(cos[i * n + j], min_crop, 0.0); count += 1.0; } } if count == 0.0 { 0.0 } else { sum / count } }; loop { let mut best: Option<(f32, usize, usize)> = None; for a in 0..n { if !groups[a].alive { continue; } for b in a + 1..n { if !groups[b].alive || !linkable(&groups[a], &groups[b]) { continue; } let p = average(&groups[a], &groups[b]); if p >= min_probability && best.is_none_or(|(bp, _, _)| p > bp) { best = Some((p, a, b)); } } } let Some((_, a, b)) = best else { break }; let taken = groups[b].clone(); groups[b].alive = false; groups[a].members.extend(taken.members); groups[a].images.extend(taken.images); groups[a].person = groups[a].person.or(taken.person); } let mut out: Vec = groups .into_iter() .filter(|g| g.alive) .map(|mut g| { g.members.sort_unstable(); Cluster { members: g.members, person: g.person, } }) .collect(); out.sort_by(|x, y| { y.members .len() .cmp(&x.members.len()) .then(x.members[0].cmp(&y.members[0])) }); out } /// `people` identities of `per` faces, each face in its own photograph, /// spread either side of the threshold so the population has genuine /// near-misses rather than obvious answers. fn population(people: usize, per: usize) -> Vec { let mut out = Vec::new(); let mut image = 0u64; for p in 0..people { for m in 0..per { // Walks down through the merge boundary as m grows, so some // members join their group and some do not. let cosine = 1.0 - (m as f32) * 0.035; out.push(Candidate { face: out.len() as u64, image, embedding: at_cosine(p, cosine), crop_px: 60.0 + ((out.len() % 11) as f32) * 25.0, confirmed_person: None, }); image += 1; } } out } /// The point of the rewrite: same clusters, less work. A disagreement here /// is the rewrite being wrong, not the reference being slow. #[test] fn the_fast_engine_agrees_with_the_reference() { let faces = population(40, 6); assert_eq!( cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY), reference(&faces, &cal(), DEFAULT_MERGE_PROBABILITY), ); } /// The two structural constraints are the ones a sparse graph could /// plausibly break, so they get their own comparison with anchors and /// co-occurrence in play. #[test] fn the_fast_engine_agrees_with_the_reference_under_constraints() { let mut faces = population(30, 6); // Some faces share a photograph, so cannot-link has to propagate // through groups that formed for other reasons. for i in (0..faces.len()).step_by(7) { faces[i].image = 900 + (i as u64 % 4); } // And some carry confirmations, including two of different people that // must never be brought together. for (n, i) in (0..faces.len()).step_by(11).enumerate() { faces[i].confirmed_person = Some(1 + (n as u64 % 3)); } assert_eq!( cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY), reference(&faces, &cal(), DEFAULT_MERGE_PROBABILITY), ); } /// The size term makes the merge boundary depend on the pair, which is the /// case the sparse pre-filter has to be built carefully to preserve. #[test] fn the_fast_engine_agrees_with_the_reference_with_a_size_term() { let faces = population(30, 6); let sized = Calibration { w_size: 0.4, b: -30.0 * 0.35 - 0.4 * 7.0, ..cal() }; assert_eq!( cluster(&faces, &sized, DEFAULT_MERGE_PROBABILITY), reference(&faces, &sized, DEFAULT_MERGE_PROBABILITY), ); } /// Determinism has to hold at a size where the indexed neighbour search is /// in play, not just on the handful of faces the small cases use. #[test] fn clustering_is_deterministic_at_scale() { let faces = population(200, 6); assert!( faces.len() > 1024, "population is below the indexing cutoff" ); assert_eq!( cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY), cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY), ); } /// A face that matches nobody is left alone rather than being swept into /// the nearest group, and costs nothing to establish — it is in no /// component at all. #[test] fn a_face_matching_nothing_stays_on_its_own() { let mut faces = population(5, 4); faces.push(Candidate { face: 999, image: 5_000, embedding: at_cosine(200, 1.0), crop_px: 150.0, confirmed_person: None, }); let out = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY); let last = faces.len() - 1; assert!( out.iter().any(|c| c.members == vec![last]), "the outlier was absorbed" ); } }