//! Grouping faces into people (docs/dev/faces.md §9, FR-CULL-10). //! //! Model-free: this is arithmetic over embeddings, and it is where the //! subsystem's accuracy actually lives, so it is testable with no weights on //! the machine. //! //! # Constraints, not just a threshold //! //! FR-CULL-10 warns that clustering will over-merge on siblings, on parents and //! children, and on the same person a decade apart. Two structural defences, //! both cheaper than a better threshold: //! //! **Cannot-link on co-occurrence.** Two faces in the same photograph are never //! merged. It is the same observation [`crate::calibrate`] mines for free //! negatives, used here as a hard constraint, and it is the single cheapest //! defence against over-merging that exists. //! //! **Confirmed faces are anchors.** A confirmation is user data (FR-CULL-12) //! and clustering never moves it. Two groups holding confirmations of //! *different* people cannot merge, whatever their similarity says. //! //! # The gallery, and the faces that are only ever compared against it //! //! A third defence, and the cheapest of all: **a short embedding is never a //! reference.** The length of the raw vector is the model's own reading of //! how recognisable the crop was ([`crate::embedding::MIN_GALLERY_QUALITY`]), //! and a short one sits near the centre of the sphere, matching a little of //! everybody. One of those in a group is a bridge to the next group over. //! //! So the population is split. Faces at or above the floor are the //! **gallery**, and they cluster exactly as described below. Faces under it //! are **probes**: each is measured against the finished groups and joins the //! one it fits, by the same average-link rule and under the same constraints //! — but it is measured against the gallery members only, never against //! another probe, and once placed it is never part of what the next face is //! measured against. A blurred photograph of a known person is still named; //! it just cannot vouch for anyone else. //! //! # Average link, not single link //! //! Single-link chains: one bad edge welds two identities together, and it is //! the documented way face clustering fails on families. Average link asks //! whether the *groups* are similar, which one outlier cannot force. //! //! # How it runs, and why the obvious way does not //! //! The first implementation of this was the textbook one: compute every //! pairwise cosine, then repeatedly scan all live group pairs, score each with //! average link, and merge the best. It is correct, it is twenty lines, and on //! a real library it does not finish. //! //! The reason is that the scan is inside the loop. Each merge rescans every //! surviving pair — `O(g²)` of them — and each score is recomputed from //! scratch over every cross pair, `O(|A|·|B|)`. With 1,813 faces that is //! roughly 1.6 million pair scores per merge and some 700 merges to do; the //! window simply stops responding, which is what a user reports as "Regroup is //! broken". At 25,000 faces it is not slow, it is impossible. //! //! Three changes, none of which alter the answer: //! //! **Only above-threshold pairs can ever matter.** An average that reaches the //! threshold must have at least one term at or above it, so two groups with no //! qualifying pair between them can never merge — not now and not after any //! sequence of merges, since merging only adds terms. [`crate::neighbours`] //! produces exactly that sparse pair list, and everything below works on it. //! On the library above it is 7,875 pairs rather than 1.6 million. //! //! **Merges cannot cross components.** Groups only ever merge along those //! pairs, so the connected components of that graph are independent problems. //! A library of four hundred people becomes four hundred small agglomerations //! instead of one large one, and the quadratic term is paid per component. //! //! **Average link is additive.** `sum(A ∪ B, C) = sum(A, C) + sum(B, C)`, so a //! merged group's scores follow from the two it came from by addition — the //! Lance-Williams update. Kept as running `(sum, count)` per adjacent pair, a //! score costs one division instead of a nested loop, and a binary heap with //! lazy invalidation replaces the rescan. //! //! The output is unchanged, deliberately and testably so: `the_fast_engine_ //! agrees_with_the_reference` runs both over the same population and asserts //! the clusters are identical. use std::cmp::Ordering; use std::collections::{BinaryHeap, HashMap, HashSet}; use crate::calibrate::Calibration; use crate::neighbours::{self, Faces}; /// Probability above which two groups are judged the same person. /// /// Stated as a probability and not a cosine, because FR-CULL-9 forbids /// thresholding a bare similarity anywhere in this subsystem. /// /// # Why 0.80 /// /// It was 0.90, and 0.90 left most of a real library ungrouped. Measured over /// the 1,813-face reference library, with `dr-ui`'s `face_index --tune`: /// /// | P | cosine | groups | faces grouped | largest group | /// |---|---|---|---|---| /// | 0.95 | 0.449 | 311 | 62% | 51 | /// | 0.90 | 0.403 | 316 | 67% | 51 | /// | 0.85 | 0.374 | 318 | 70% | 57 | /// | **0.80** | **0.353** | **328** | **74%** | **69** | /// | 0.75 | 0.335 | 327 | 77% | 69 | /// | 0.70 | 0.319 | 326 | 79% | 81 | /// | 0.50 | 0.267 | 303 | 85% | 90 | /// /// The count of *groups* is the signal, not the count of grouped faces. Loosen /// from 0.95 and it climbs: real people are being assembled out of fragments. /// It peaks at 0.80 and then falls, and a falling group count while the grouped /// faces keep rising is the shape of over-merging — separate identities being /// welded, which is the failure FR-CULL-10 warns about and the one the user /// cannot easily undo by hand. /// /// So: the loosest setting that is still building people rather than melting /// them together. A third more of the library gets grouped than at 0.90, and /// the largest group grows by eighteen faces rather than by forty. /// /// This is a *default*, not a constant of nature — the numbers above are one /// library, and `--tune` reruns the table on any other. pub const DEFAULT_MERGE_PROBABILITY: f32 = 0.80; /// A face presented to the clusterer. /// /// Ids are opaque `u64`s rather than catalog types: this crate has no business /// knowing what a `FaceId` means, and the caller does the translation. #[derive(Debug, Clone)] pub struct Candidate { pub face: u64, /// Which photograph it came from — the cannot-link key. pub image: u64, /// L2-normalised, `EMBEDDING_DIM` long. pub embedding: Vec, /// Source pixels across the aligned crop, for the calibration's size term. pub crop_px: f32, /// Length of the raw embedding, where it was recorded /// ([`crate::embedding::MIN_GALLERY_QUALITY`]). `None` for a face indexed /// before it was kept, which is admitted to the gallery — see /// [`Candidate::in_gallery`]. pub quality: Option, /// The person this face is *confirmed* to be, if any. /// /// Suggestions are deliberately not passed here. They are this function's /// own previous output, and feeding them back in would let a guess harden /// into a fact across successive passes. pub confirmed_person: Option, } impl Candidate { /// Whether this face may be compared *against*, as well as compared. pub fn in_gallery(&self) -> bool { crate::embedding::in_gallery(self.quality) } } /// One group of faces the clusterer believes are one person. #[derive(Debug, Clone, PartialEq)] pub struct Cluster { /// Indices into the input slice. pub members: Vec, /// The person this group is already known to be, from its anchors. /// /// `Some` means the group contains confirmed faces and the suggestions in /// it attach to that existing person. `None` is a new unnamed group. pub person: Option, } /// Group faces into people. /// /// `min_probability` is compared against the calibrated average-link /// probability between two groups. Deterministic: the same input yields the /// same clusters, because the merge order is by score with the index pair as /// the tiebreak. pub fn cluster(faces: &[Candidate], cal: &Calibration, min_probability: f32) -> Vec { if faces.is_empty() { return Vec::new(); } let columns = Columns::of(faces); // Every pair that could ever contribute to a merge. See the module note on // why nothing outside this list can matter. let pairs = neighbours::above_threshold(&columns.view(), cal, min_probability); build(faces, cal, min_probability, &pairs) } /// Groups, and how confident each face's placement is. /// /// The second half is [`crate::assign`]'s share, not the pairwise probability /// that put the face in the group — see that module for why the two are /// different questions. #[derive(Debug, Clone, PartialEq)] pub struct Grouping { pub clusters: Vec, /// Indexed like the input faces. 0 for a face in no group. pub confidence: Vec, } /// Group faces into people, and score each placement against its rivals. /// /// What a caller writing suggestions into a catalog wants: [`cluster`] answers /// *which person*, this answers *and how sure*. /// /// One scan, two thresholds. The similarity scan is the expensive part of the /// whole subsystem and running it twice — once to merge, once to find rivals — /// would double the cost of regrouping a library. So it runs once at the looser /// of the two floors, and the merge engine takes the subset at or above /// `min_probability`. That subset is identical, pair for pair and in the same /// order, to what a scan at `min_probability` would have produced, so grouping /// is unchanged by scoring being asked for: `scoring_does_not_change_the_ /// groups` holds it to that. pub fn cluster_scored(faces: &[Candidate], cal: &Calibration, min_probability: f32) -> Grouping { if faces.is_empty() { return Grouping { clusters: Vec::new(), confidence: Vec::new(), }; } let columns = Columns::of(faces); let evidence = neighbours::above_threshold( &columns.view(), cal, min_probability.min(crate::assign::RIVAL_FLOOR), ); let merges: Vec = evidence .iter() .copied() .filter(|p| p.probability >= min_probability) .collect(); let clusters = build(faces, cal, min_probability, &merges); let confidence = crate::assign::identity_shares( &columns.gallery, &clusters, &evidence, crate::assign::TOP_MATCHES, ); Grouping { clusters, confidence, } } /// The three arrays [`neighbours::Faces`] borrows, owned. /// /// [`neighbours`] takes parallel slices rather than candidates on purpose — it /// has no business knowing what a person is — so somebody has to hold the /// columns. Both entry points do, identically, which is the only reason this is /// a type and not three locals. struct Columns { /// Every embedding end to end — see [`Faces::embeddings`] for why flat. embeddings: Vec, dim: usize, crop_px: Vec, images: Vec, gallery: Vec, } impl Columns { fn of(faces: &[Candidate]) -> Self { // Ragged input would index the wrong row for every face after the odd // one, so the widest wins and short rows are padded with zeros: a // zero-padded row scores lower against everything, which is the safe // direction. It does not happen — one model, one dimension — and it is // handled rather than trusted because the failure would be silent. let dim = faces.iter().map(|f| f.embedding.len()).max().unwrap_or(0); let mut embeddings = Vec::with_capacity(faces.len() * dim); for f in faces { embeddings.extend_from_slice(&f.embedding); embeddings.resize(embeddings.len() + dim - f.embedding.len(), 0.0); } Self { embeddings, dim, crop_px: faces.iter().map(|f| f.crop_px).collect(), images: faces.iter().map(|f| f.image).collect(), gallery: faces.iter().map(Candidate::in_gallery).collect(), } } fn view(&self) -> Faces<'_> { Faces { embeddings: &self.embeddings, dim: self.dim, crop_px: &self.crop_px, images: &self.images, gallery: &self.gallery, } } } /// Agglomerate the gallery over its pairs, then place the probes. /// /// `pairs` is what [`neighbours::above_threshold`] returned: every pair has a /// gallery side, but a pair with a probe on the other side is not a merge — /// it is the evidence [`place_probes`] works from. Only the gallery-to-gallery /// pairs reach the engine, so a probe enters it as a singleton with no edges /// and comes out exactly as it went in. fn build( faces: &[Candidate], cal: &Calibration, min_probability: f32, pairs: &[neighbours::Pair], ) -> Vec { let gallery: Vec = faces.iter().map(Candidate::in_gallery).collect(); let (merges, probe_pairs): (Vec<_>, Vec<_>) = pairs .iter() .copied() .partition(|p| gallery[p.i] && gallery[p.j]); let mut engine = Engine::new(faces, cal, min_probability); let parts = components(faces.len(), &merges); for (component, edges) in parts.members.iter().zip(&parts.edges) { engine.agglomerate(component, edges); } let dot = engine.dot; let clusters = engine.finish(); if probe_pairs.is_empty() { return clusters; } place_probes( faces, cal, min_probability, dot, &gallery, clusters, &probe_pairs, ) } /// Put each probe into the finished group it fits, or leave it alone. /// /// The same decision the engine makes for a singleton — average link over the /// group, at or above `min_probability`, subject to [`Engine::can_link`]'s two /// constraints — with one difference that is the whole point: the average is /// over the group's **gallery** members. A probe already placed is not part of /// what the next one is measured against, so a run of short vectors cannot /// pull each other in one after another. /// /// Probes are placed in index order and each placement is final, which is /// what keeps this deterministic. The group a probe joins gains its /// photograph, so a second face from the same frame cannot follow it — the /// co-occurrence rule, applied exactly as the engine applies it. fn place_probes( faces: &[Candidate], cal: &Calibration, min_probability: f32, dot: neighbours::DotFn, gallery: &[bool], mut clusters: Vec, probe_pairs: &[neighbours::Pair], ) -> Vec { // Where each face sits, and what each group's photographs and gallery // members are. The probe's own singleton is here too, and is dropped once // it has moved. let mut group_of = vec![usize::MAX; faces.len()]; for (g, c) in clusters.iter().enumerate() { for &m in &c.members { group_of[m] = g; } } let mut images: Vec> = clusters .iter() .map(|c| c.members.iter().map(|&m| faces[m].image).collect()) .collect(); let references: Vec> = clusters .iter() .map(|c| c.members.iter().copied().filter(|&m| gallery[m]).collect()) .collect(); // Which groups each probe has any above-threshold pair into. Only those // can average above the threshold — the argument the module note makes // for the engine holds here unchanged. let mut candidates: Vec> = vec![Vec::new(); faces.len()]; for p in probe_pairs { let (probe, reference) = if gallery[p.i] { (p.j, p.i) } else { (p.i, p.j) }; candidates[probe].push(group_of[reference]); } let mut moved: Vec = Vec::new(); for probe in 0..faces.len() { if gallery[probe] || candidates[probe].is_empty() { continue; } let mut groups = std::mem::take(&mut candidates[probe]); groups.sort_unstable(); groups.dedup(); let face = &faces[probe]; let mut best: Option<(f32, usize)> = None; for g in groups { let target = &clusters[g]; if let (Some(mine), Some(theirs)) = (face.confirmed_person, target.person) { if mine != theirs { continue; } } if images[g].contains(&face.image) { continue; } let (mut sum, mut count) = (0.0_f64, 0.0_f64); for &r in &references[g] { let cos = dot(&face.embedding, &faces[r].embedding); let min_crop = face.crop_px.min(faces[r].crop_px); sum += cal.probability(cos, min_crop, 0.0) as f64; count += 1.0; } if count == 0.0 { continue; } let p = (sum / count) as f32; // Strictly better wins; on a tie the lowest group index, which is // the engine's own tiebreak. if p >= min_probability && best.is_none_or(|(bp, _)| p > bp) { best = Some((p, g)); } } let Some((_, g)) = best else { continue }; let own = group_of[probe]; clusters[g].members.push(probe); clusters[g].members.sort_unstable(); clusters[g].person = clusters[g].person.or(face.confirmed_person); images[g].insert(face.image); group_of[probe] = g; moved.push(own); } if moved.is_empty() { return clusters; } // The singletons the probes left behind, then the order `Engine::finish` // promises: largest first, lowest member first among equals. let mut vacated = vec![false; clusters.len()]; for g in moved { vacated[g] = true; } let mut out: Vec = clusters .into_iter() .zip(vacated) .filter(|(_, gone)| !gone) .map(|(c, _)| c) .collect(); out.sort_by(|x, y| { y.members .len() .cmp(&x.members.len()) .then(x.members[0].cmp(&y.members[0])) }); out } /// Split one person's faces into the groups a raised threshold separates them /// into. /// /// FR-CULL-10 requires splitting to be as easy as merging, and a split that /// hands the user a pile of loose faces to re-sort is not that. This re-runs /// the same agglomeration at a stricter probability so the user is offered /// coherent sub-groups to pull apart. /// /// Anchors are ignored here on purpose: every face in the input is already /// nominally the same person, so honouring the anchors would refuse to split /// anything. pub fn split(faces: &[Candidate], cal: &Calibration, min_probability: f32) -> Vec { let anchorless: Vec = faces .iter() .cloned() .map(|mut f| { f.confirmed_person = None; f }) .collect(); cluster(&anchorless, cal, min_probability) } // ── the merge engine ────────────────────────────────────────────────────── /// One live group, identified throughout by the index of its lowest member. #[derive(Debug)] struct Group { /// Ascending, always — [`Engine::cross`] sums in this order, and a stable /// order is what makes the floating-point total reproducible. members: Vec, images: HashSet, person: Option, alive: bool, /// Bumped on every merge, so heap entries naming an older state can be /// recognised and dropped instead of acted on. version: u64, } /// Running average-link state for one adjacent pair of groups. /// /// `sum` is over **every** cross pair, not only the above-threshold ones — /// average link is an average over all of them, and counting only the /// qualifying pairs would report a similarity no group actually has. #[derive(Debug, Clone, Copy)] struct Link { sum: f64, count: f64, } impl Link { fn probability(&self) -> f32 { if self.count == 0.0 { 0.0 } else { (self.sum / self.count) as f32 } } } /// A candidate merge, waiting in the heap. #[derive(Debug, Clone, Copy)] struct Pending { probability: f32, a: usize, b: usize, /// Group versions when this was pushed. A mismatch on pop means a merge /// has happened since and a fresher entry for this pair is already queued. va: u64, vb: u64, } impl PartialEq for Pending { fn eq(&self, other: &Self) -> bool { self.cmp(other) == Ordering::Equal } } impl Eq for Pending {} impl PartialOrd for Pending { fn partial_cmp(&self, other: &Self) -> Option { Some(self.cmp(other)) } } impl Ord for Pending { /// Greatest pops first, so: highest probability, and on a tie the lowest /// index pair. That tiebreak is not cosmetic — it is what the old /// ascending scan did, and it is the whole of the determinism guarantee. fn cmp(&self, other: &Self) -> Ordering { self.probability .total_cmp(&other.probability) .then_with(|| other.a.cmp(&self.a)) .then_with(|| other.b.cmp(&self.b)) } } struct Engine<'a> { faces: &'a [Candidate], cal: &'a Calibration, /// The same kernel [`neighbours`] scans with. [`Engine::cross`] is the one /// place a dot product is computed during agglomeration, and it was using /// the portable loop while the scan beside it had the machine's SIMD. dot: neighbours::DotFn, min_probability: f32, groups: Vec, links: HashMap<(usize, usize), Link>, /// Adjacency, as group ids. Kept alongside `links` so a merge can find /// everything it has to update without scanning the whole map. adjacent: Vec>, } impl<'a> Engine<'a> { fn new(faces: &'a [Candidate], cal: &'a Calibration, min_probability: f32) -> Self { let groups = faces .iter() .enumerate() .map(|(i, f)| Group { members: vec![i], images: HashSet::from([f.image]), person: f.confirmed_person, alive: true, version: 0, }) .collect(); Self { faces, cal, dot: neighbours::fastest_dot(), min_probability, groups, links: HashMap::new(), adjacent: vec![HashSet::new(); faces.len()], } } /// Agglomerate one connected component to exhaustion. fn agglomerate(&mut self, component: &[usize], pairs: &[neighbours::Pair]) { if component.len() < 2 { return; } let mut heap = BinaryHeap::with_capacity(pairs.len()); for p in pairs { self.links.insert( key(p.i, p.j), Link { sum: p.probability as f64, count: 1.0, }, ); self.adjacent[p.i].insert(p.j); self.adjacent[p.j].insert(p.i); heap.push(Pending { probability: p.probability, a: p.i.min(p.j), b: p.i.max(p.j), va: 0, vb: 0, }); } while let Some(top) = heap.pop() { let Pending { probability, a, b, va, vb, } = top; // Stale: one side has merged since this was queued, and the // replacement entry is already in the heap. if !self.groups[a].alive || !self.groups[b].alive || self.groups[a].version != va || self.groups[b].version != vb { continue; } // The heap is ordered by probability, so the first entry below the // bar means nothing left in this component can reach it. if probability < self.min_probability { break; } if !self.can_link(a, b) { // Never becomes possible again: images only accumulate and an // anchor is never given up, so drop the pair for good. self.unlink(a, b); continue; } self.merge(a, b, &mut heap); } } /// Whether two groups are allowed to merge at all, before similarity is /// asked. fn can_link(&self, a: usize, b: usize) -> bool { let (ga, gb) = (&self.groups[a], &self.groups[b]); // Two confirmations of different people. The user has said these are // not the same person, and no similarity overrides that. if let (Some(pa), Some(pb)) = (ga.person, gb.person) { if pa != pb { return false; } } // Co-occurrence: a photograph containing a face from each group means // the two faces are in the same frame, so they are not the same person. ga.images.is_disjoint(&gb.images) } /// Fold `b` into `a` and re-score everything that touched either. fn merge(&mut self, a: usize, b: usize, heap: &mut BinaryHeap) { // Sorted, and deduplicated by the set: `a` and `b` may share // neighbours, and each must be visited once. Sorting is what keeps the // floating-point sums identical from run to run. let mut touched: Vec = self.adjacent[a] .union(&self.adjacent[b]) .copied() .filter(|&c| c != a && c != b && self.groups[c].alive) .collect(); touched.sort_unstable(); // Take the pair sums before the groups change underneath them. let carried: Vec<(usize, Option, Option)> = touched .iter() .map(|&c| { ( c, self.links.get(&key(a, c)).copied(), self.links.get(&key(b, c)).copied(), ) }) .collect(); // a's own members, before b's are folded in. A missing a-side sum has // to be computed over these and not over the merged list, or b's // contribution would be counted twice. let a_members = self.groups[a].members.clone(); // Absorb b into a. let taken = std::mem::replace( &mut self.groups[b], Group { members: Vec::new(), images: HashSet::new(), person: None, alive: false, version: 0, }, ); { let ga = &mut self.groups[a]; ga.members.extend(taken.members.iter().copied()); ga.members.sort_unstable(); ga.images.extend(taken.images.iter().copied()); // At most one side carries a person: `can_link` refuses a merge of // two groups anchored to different people, so this cannot silently // discard one of them. ga.person = ga.person.or(taken.person); ga.version += 1; } // b's own links are gone with it. for c in self.adjacent[b].clone() { self.links.remove(&key(b, c)); self.adjacent[c].remove(&b); } self.adjacent[b].clear(); self.links.remove(&key(a, b)); self.adjacent[a].remove(&b); for (c, from_a, from_b) in carried { // Dropping a pair the constraints now forbid saves computing a // score for a merge that can never happen — which for a newly // adjacent side is a real cost, not a bookkeeping one. if !self.can_link(a, c) { self.unlink(a, c); continue; } // A side with no stored link was not adjacent before, so its cross // pairs were all below threshold and were never summed. They still // belong in the average, so they are computed now — once, after // which the additive update carries them forward. let from_a = from_a.unwrap_or_else(|| self.cross(&a_members, c)); let from_b = from_b.unwrap_or_else(|| self.cross(&taken.members, c)); let merged = Link { sum: from_a.sum + from_b.sum, count: from_a.count + from_b.count, }; self.links.insert(key(a, c), merged); self.adjacent[a].insert(c); self.adjacent[c].insert(a); heap.push(Pending { probability: merged.probability(), a: a.min(c), b: a.max(c), va: self.groups[a.min(c)].version, vb: self.groups[a.max(c)].version, }); } } /// Exact `(sum, count)` over every cross pair between a member list and a /// group. /// /// The one place a dot product is still computed during agglomeration, and /// it happens only when two groups become adjacent through a third — at /// which point their sub-threshold pairs, never summed because they were /// never interesting, have to be accounted for. fn cross(&self, members: &[usize], group: usize) -> Link { let mut sum = 0.0_f64; let mut count = 0.0_f64; for &i in members { for &j in &self.groups[group].members { let cos = (self.dot)(&self.faces[i].embedding, &self.faces[j].embedding); let min_crop = self.faces[i].crop_px.min(self.faces[j].crop_px); sum += self.cal.probability(cos, min_crop, 0.0) as f64; count += 1.0; } } Link { sum, count } } fn unlink(&mut self, a: usize, b: usize) { self.links.remove(&key(a, b)); self.adjacent[a].remove(&b); self.adjacent[b].remove(&a); } fn finish(self) -> Vec { let mut out: Vec = self .groups .into_iter() .filter(|g| g.alive) .map(|g| Cluster { members: g.members, person: g.person, }) .collect(); // Largest first: the People view shows the best-evidenced groups at the // top. out.sort_by(|x, y| { y.members .len() .cmp(&x.members.len()) .then(x.members[0].cmp(&y.members[0])) }); out } } fn key(a: usize, b: usize) -> (usize, usize) { if a < b { (a, b) } else { (b, a) } } /// Connected components of the above-threshold graph. /// /// Faces in different components can never end up in one group, so each is a /// separate and much smaller agglomeration. Returned with the members of each /// component ascending, and the components themselves in order of their lowest /// member — the determinism the merge order inherits. fn components(n: usize, pairs: &[neighbours::Pair]) -> Components { let mut parent: Vec = (0..n).collect(); fn find(parent: &mut [usize], mut x: usize) -> usize { while parent[x] != x { // Path halving: keeps the tree flat without a second pass. parent[x] = parent[parent[x]]; x = parent[x]; } x } for p in pairs { let (ra, rb) = (find(&mut parent, p.i), find(&mut parent, p.j)); if ra != rb { // Lowest root wins, so the representative of a component is // reproducible rather than an artefact of union order. let (lo, hi) = if ra < rb { (ra, rb) } else { (rb, ra) }; parent[hi] = lo; } } let mut by_root: HashMap> = HashMap::new(); for i in 0..n { let r = find(&mut parent, i); by_root.entry(r).or_default().push(i); } let mut members: Vec> = by_root.into_values().filter(|c| c.len() > 1).collect(); members.sort_unstable_by_key(|c| c[0]); // Each pair filed under its component, in one pass. // // **This is not bookkeeping, it is the cost of the whole stage.** Each // component used to scan the entire pair list for the ones that were its // own: 475 components against 804,499 pairs on the reference library, 382 // million set lookups, and 3.06 s of a 5.93 s regroup spent before a single // merge was considered. A pair can only ever join two faces of one // component — that is what a component *is* — so one pass over the list // places every pair exactly where it is needed. let mut slot_of = vec![usize::MAX; n]; for (slot, c) in members.iter().enumerate() { for &m in c { slot_of[m] = slot; } } let mut edges: Vec> = vec![Vec::new(); members.len()]; for p in pairs { // Ordering within a component follows the global order, which is what // the seeded heap's tiebreak — and so the determinism promise — rests // on. if let Some(slot) = slot_of.get(p.i).copied().filter(|&s| s != usize::MAX) { edges[slot].push(*p); } } Components { members, edges } } /// The connected components of the pair graph, each with the pairs inside it. /// /// The two halves are parallel: `members[k]` and `edges[k]` describe the same /// component. struct Components { members: Vec>, edges: Vec>, } #[cfg(test)] mod tests { use super::*; use crate::embedding::EMBEDDING_DIM; /// An embedding a known cosine away from a base direction, built by mixing /// two orthogonal unit vectors. Lets a test state "these two faces are 0.7 /// similar" and have it be exactly true. fn at_cosine(identity: usize, cosine: f32) -> Vec { let mut v = vec![0.0_f32; EMBEDDING_DIM]; let base = identity * 2; let perp = identity * 2 + 1; v[base] = cosine; v[perp] = (1.0 - cosine * cosine).max(0.0).sqrt(); v } fn candidate(face: u64, image: u64, identity: usize, cosine: f32) -> Candidate { Candidate { face, image, embedding: at_cosine(identity, cosine), crop_px: 150.0, quality: None, confirmed_person: None, } } /// A face too short to be a reference: compared, never compared against. fn probe(face: u64, image: u64, identity: usize, cosine: f32) -> Candidate { Candidate { quality: Some(crate::embedding::MIN_GALLERY_QUALITY - 5.0), ..candidate(face, image, identity, cosine) } } /// A calibration steep enough that the test's cosines are unambiguous: /// 0.6 is near-certain, 0.1 is near-impossible. fn cal() -> Calibration { Calibration { a: 30.0, b: -30.0 * 0.35, w_size: 0.0, valid: true, positive_pairs: 1000, negative_pairs: 10_000, } } #[test] fn no_faces_makes_no_clusters() { assert!(cluster(&[], &cal(), DEFAULT_MERGE_PROBABILITY).is_empty()); } #[test] fn similar_faces_from_different_photographs_group_together() { let faces = vec![ candidate(1, 10, 0, 1.0), candidate(2, 11, 0, 0.95), candidate(3, 12, 0, 0.92), ]; let out = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY); assert_eq!(out.len(), 1); assert_eq!(out[0].members, vec![0, 1, 2]); } #[test] fn dissimilar_faces_stay_apart() { let faces = vec![candidate(1, 10, 0, 1.0), candidate(2, 11, 1, 1.0)]; let out = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY); assert_eq!(out.len(), 2); } /// The cheapest defence against over-merging: two faces in one frame are /// not the same person however similar the model finds them. #[test] fn two_faces_in_one_photograph_never_merge() { // Identical embeddings — siblings, or a model that cannot tell them // apart — but both in image 10. let faces = vec![candidate(1, 10, 0, 1.0), candidate(2, 10, 0, 1.0)]; let out = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY); assert_eq!(out.len(), 2, "co-occurring faces were merged"); } /// And the constraint has to survive transitively: once a group holds a /// face from image 10, no other group holding one from image 10 may join /// it, even indirectly. #[test] fn the_co_occurrence_constraint_propagates_through_a_group() { let faces = vec![ candidate(1, 10, 0, 1.0), // A, in the group photo candidate(2, 10, 0, 1.0), // B, in the same group photo candidate(3, 11, 0, 0.99), // A again, alone ]; let out = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY); assert_eq!(out.len(), 2); // Whichever of A/B absorbed face 2, the other stays out. assert!(out.iter().any(|c| c.members.len() == 2)); assert!(out.iter().any(|c| c.members.len() == 1)); } /// FR-CULL-10: a confirmation is user data and no inference overrides it. #[test] fn groups_confirmed_as_different_people_do_not_merge() { let mut a = candidate(1, 10, 0, 1.0); let mut b = candidate(2, 11, 0, 1.0); a.confirmed_person = Some(100); b.confirmed_person = Some(200); let out = cluster(&[a, b], &cal(), DEFAULT_MERGE_PROBABILITY); assert_eq!(out.len(), 2, "clustering overrode two user confirmations"); } /// A face at a chosen cosine to identity 0 *and* to identity 1 at once — /// the sibling geometry, which [`at_cosine`]'s per-identity subspaces /// cannot express. fn contested(to_first: f32, to_second: f32) -> Vec { let mut v = vec![0.0_f32; EMBEDDING_DIM]; v[0] = to_first; v[2] = to_second; v[4] = (1.0 - to_first * to_first - to_second * to_second) .max(0.0) .sqrt(); v } /// The population the scoring tests share: two faces of one person, a /// stranger, and a face that matches the person well and the stranger /// weakly — weakly enough that it will never merge with them, which is /// exactly the rival a within-group score cannot see. fn with_a_rival() -> Vec { let mut x = candidate(4, 13, 0, 0.0); x.embedding = contested(0.45, 0.37); // The stranger is *named*: only an identity the user has asserted // competes for a face (crate::assign). let mut stranger = candidate(3, 12, 1, 1.0); stranger.confirmed_person = Some(7); vec![ candidate(1, 10, 0, 1.0), candidate(2, 11, 0, 1.0), stranger, x, ] } /// Asking for confidences must not move a single face. The scan runs at a /// looser floor to find rivals, and the merge engine has to see exactly the /// pairs it would have seen without them. #[test] fn scoring_does_not_change_the_groups() { let faces = with_a_rival(); let plain = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY); let scored = cluster_scored(&faces, &cal(), DEFAULT_MERGE_PROBABILITY); assert_eq!(plain, scored.clusters); } /// The number the user is shown answers "which of these people", so a /// second claimant has to lower it even when it is too weak to merge. #[test] fn a_face_two_identities_could_claim_is_reported_as_less_certain() { let faces = with_a_rival(); let scored = cluster_scored(&faces, &cal(), DEFAULT_MERGE_PROBABILITY); let alone = cluster_scored(&faces[..2], &cal(), DEFAULT_MERGE_PROBABILITY); assert!(alone.confidence[0] > 0.99, "nobody else to be"); let contested = scored.confidence[3]; assert!( (0.6..0.85).contains(&contested), "a face with a second claimant: {contested}" ); } #[test] fn a_suggestion_joins_the_person_its_group_is_anchored_to() { let mut anchor = candidate(1, 10, 0, 1.0); anchor.confirmed_person = Some(42); let loose = candidate(2, 11, 0, 0.95); let out = cluster(&[anchor, loose], &cal(), DEFAULT_MERGE_PROBABILITY); assert_eq!(out.len(), 1); assert_eq!(out[0].person, Some(42)); assert_eq!(out[0].members.len(), 2); } #[test] fn an_unanchored_group_is_a_new_unnamed_person() { let out = cluster( &[candidate(1, 10, 0, 1.0), candidate(2, 11, 0, 0.95)], &cal(), DEFAULT_MERGE_PROBABILITY, ); assert_eq!(out[0].person, None); } /// Average link rather than single link: one strong edge must not weld two /// otherwise-dissimilar groups together. This is the family failure mode /// FR-CULL-10 names. #[test] fn one_strong_edge_does_not_chain_two_groups_together() { // Two tight pairs, with a single borderline link between them. let faces = vec![ candidate(1, 10, 0, 1.00), candidate(2, 11, 0, 0.99), candidate(3, 12, 0, 0.42), candidate(4, 13, 0, 0.40), ]; let out = cluster(&faces, &cal(), 0.99); assert!( out.len() >= 2, "single-link chaining merged everything into {} cluster(s)", out.len() ); } #[test] fn clustering_is_deterministic() { let faces = vec![ candidate(1, 10, 0, 1.0), candidate(2, 11, 0, 0.96), candidate(3, 12, 1, 1.0), candidate(4, 13, 1, 0.97), candidate(5, 14, 0, 0.94), ]; let a = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY); let b = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY); assert_eq!(a, b); } #[test] fn clusters_come_back_largest_first() { let faces = vec![ candidate(1, 10, 1, 1.0), candidate(2, 11, 0, 1.0), candidate(3, 12, 0, 0.97), candidate(4, 13, 0, 0.95), ]; let out = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY); assert_eq!(out[0].members.len(), 3); assert_eq!(out[1].members.len(), 1); } /// Splitting is the inverse operation and must actually separate a group /// that a looser threshold had merged. #[test] fn split_separates_a_group_that_a_looser_threshold_merged() { let faces = vec![ candidate(1, 10, 0, 1.00), candidate(2, 11, 0, 0.98), candidate(3, 12, 0, 0.45), candidate(4, 13, 0, 0.43), ]; // Loose: one person. assert_eq!(cluster(&faces, &cal(), 0.5).len(), 1); // Strict: the two sub-groups the user wants offered. let parts = split(&faces, &cal(), 0.999); assert!(parts.len() >= 2, "split produced {} group(s)", parts.len()); } #[test] fn split_ignores_the_anchor_so_a_mislabelled_person_can_be_taken_apart() { let mut a = candidate(1, 10, 0, 1.0); let mut b = candidate(2, 11, 0, 0.40); a.confirmed_person = Some(7); b.confirmed_person = Some(7); let parts = split(&[a, b], &cal(), 0.99); assert_eq!(parts.len(), 2); } /// The size term earns its place: the same cosine between two thumbnail- /// sized faces should be less convincing than between two large ones. #[test] fn the_face_size_term_moves_the_probability() { let sized = Calibration { w_size: 0.5, b: -30.0 * 0.35 - 0.5 * 7.0, ..cal() }; let big = sized.probability(0.5, 300.0, 0.0); let small = sized.probability(0.5, 40.0, 0.0); assert!(big > small, "big {big} should beat small {small}"); } // ── the fast engine against the obvious one ─────────────────────────── /// The original implementation, kept as the oracle. /// /// Deliberately the naive version this module replaced: rescan every live /// pair, score it from scratch over all cross pairs, merge the best, /// repeat. It is the definition of the answer, and the only thing the /// rewrite was allowed to change is how long it takes to get there. fn reference(faces: &[Candidate], cal: &Calibration, min_probability: f32) -> Vec { #[derive(Clone)] struct G { members: Vec, images: HashSet, person: Option, alive: bool, } if faces.is_empty() { return Vec::new(); } let n = faces.len(); let mut groups: Vec = faces .iter() .enumerate() .map(|(i, f)| G { members: vec![i], images: HashSet::from([f.image]), person: f.confirmed_person, alive: true, }) .collect(); let mut cos = vec![0.0_f32; n * n]; for i in 0..n { for j in i + 1..n { let c = neighbours::dot(&faces[i].embedding, &faces[j].embedding); cos[i * n + j] = c; cos[j * n + i] = c; } } let linkable = |a: &G, b: &G| { if let (Some(pa), Some(pb)) = (a.person, b.person) { if pa != pb { return false; } } a.images.is_disjoint(&b.images) }; let average = |a: &G, b: &G| { let mut sum = 0.0_f32; let mut count = 0.0_f32; for &i in &a.members { for &j in &b.members { let min_crop = faces[i].crop_px.min(faces[j].crop_px); sum += cal.probability(cos[i * n + j], min_crop, 0.0); count += 1.0; } } if count == 0.0 { 0.0 } else { sum / count } }; loop { let mut best: Option<(f32, usize, usize)> = None; for a in 0..n { if !groups[a].alive { continue; } for b in a + 1..n { if !groups[b].alive || !linkable(&groups[a], &groups[b]) { continue; } let p = average(&groups[a], &groups[b]); if p >= min_probability && best.is_none_or(|(bp, _, _)| p > bp) { best = Some((p, a, b)); } } } let Some((_, a, b)) = best else { break }; let taken = groups[b].clone(); groups[b].alive = false; groups[a].members.extend(taken.members); groups[a].images.extend(taken.images); groups[a].person = groups[a].person.or(taken.person); } let mut out: Vec = groups .into_iter() .filter(|g| g.alive) .map(|mut g| { g.members.sort_unstable(); Cluster { members: g.members, person: g.person, } }) .collect(); out.sort_by(|x, y| { y.members .len() .cmp(&x.members.len()) .then(x.members[0].cmp(&y.members[0])) }); out } /// `people` identities of `per` faces, each face in its own photograph, /// spread either side of the threshold so the population has genuine /// near-misses rather than obvious answers. fn population(people: usize, per: usize) -> Vec { let mut out = Vec::new(); let mut image = 0u64; for p in 0..people { for m in 0..per { // Walks down through the merge boundary as m grows, so some // members join their group and some do not. let cosine = 1.0 - (m as f32) * 0.035; out.push(Candidate { face: out.len() as u64, image, embedding: at_cosine(p, cosine), crop_px: 60.0 + ((out.len() % 11) as f32) * 25.0, quality: None, confirmed_person: None, }); image += 1; } } out } /// The point of the rewrite: same clusters, less work. A disagreement here /// is the rewrite being wrong, not the reference being slow. #[test] fn the_fast_engine_agrees_with_the_reference() { let faces = population(40, 6); assert_eq!( cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY), reference(&faces, &cal(), DEFAULT_MERGE_PROBABILITY), ); } /// The two structural constraints are the ones a sparse graph could /// plausibly break, so they get their own comparison with anchors and /// co-occurrence in play. #[test] fn the_fast_engine_agrees_with_the_reference_under_constraints() { let mut faces = population(30, 6); // Some faces share a photograph, so cannot-link has to propagate // through groups that formed for other reasons. for i in (0..faces.len()).step_by(7) { faces[i].image = 900 + (i as u64 % 4); } // And some carry confirmations, including two of different people that // must never be brought together. for (n, i) in (0..faces.len()).step_by(11).enumerate() { faces[i].confirmed_person = Some(1 + (n as u64 % 3)); } assert_eq!( cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY), reference(&faces, &cal(), DEFAULT_MERGE_PROBABILITY), ); } /// The size term makes the merge boundary depend on the pair, which is the /// case the sparse pre-filter has to be built carefully to preserve. #[test] fn the_fast_engine_agrees_with_the_reference_with_a_size_term() { let faces = population(30, 6); let sized = Calibration { w_size: 0.4, b: -30.0 * 0.35 - 0.4 * 7.0, ..cal() }; assert_eq!( cluster(&faces, &sized, DEFAULT_MERGE_PROBABILITY), reference(&faces, &sized, DEFAULT_MERGE_PROBABILITY), ); } /// Determinism has to hold at a size where the indexed neighbour search is /// in play, not just on the handful of faces the small cases use. #[test] fn clustering_is_deterministic_at_scale() { let faces = population(200, 6); assert!( faces.len() > 1024, "population is below the indexing cutoff" ); assert_eq!( cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY), cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY), ); } /// A face that matches nobody is left alone rather than being swept into /// the nearest group, and costs nothing to establish — it is in no /// component at all. #[test] fn a_face_matching_nothing_stays_on_its_own() { let mut faces = population(5, 4); faces.push(Candidate { face: 999, image: 5_000, embedding: at_cosine(200, 1.0), crop_px: 150.0, quality: None, confirmed_person: None, }); let out = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY); let last = faces.len() - 1; assert!( out.iter().any(|c| c.members == vec![last]), "the outlier was absorbed" ); } // ── the gallery ─────────────────────────────────────────────────────── /// A short vector is still somebody: it joins the group it matches. #[test] fn a_probe_joins_the_group_it_matches() { let faces = vec![ candidate(1, 10, 0, 1.0), candidate(2, 11, 0, 0.95), probe(3, 12, 0, 0.92), ]; let out = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY); assert_eq!(out.len(), 1); assert_eq!(out[0].members, vec![0, 1, 2]); } /// Two short vectors that resemble each other are noise agreeing with /// noise, and there is nothing in the gallery for either to be measured /// against. #[test] fn two_probes_are_never_grouped_with_each_other() { let faces = vec![probe(1, 10, 0, 1.0), probe(2, 11, 0, 0.98)]; let out = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY); assert_eq!(out.len(), 2, "two probes were grouped: {out:?}"); } /// The point of measuring against the gallery only: a probe that has been /// placed is not a stepping stone for the next one. #[test] fn a_placed_probe_is_not_what_the_next_probe_is_measured_against() { let mut first = probe(2, 11, 0, 0.6); // 0.6 along identity 0 and 0.8 along its perpendicular: near enough to // the reference to join it, and much nearer to the face below. first.embedding = at_cosine(0, 0.6); let mut second = probe(3, 12, 0, 0.0); second.embedding = at_cosine(0, 0.0); let faces = vec![candidate(1, 10, 0, 1.0), first, second]; let out = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY); let group = out.iter().find(|c| c.members.contains(&0)).unwrap(); assert_eq!( group.members, vec![0, 1], "the first probe should have joined" ); assert!( out.iter().any(|c| c.members == vec![2]), "the second probe reached the group through the first: {out:?}" ); } /// A confirmation on a probe is still the user's word: the group it joins /// becomes that person, and a group already someone else's is closed to it. #[test] fn a_probe_carries_its_confirmation_and_respects_others() { let mut anchored = probe(3, 12, 0, 0.92); anchored.confirmed_person = Some(7); let faces = vec![ candidate(1, 10, 0, 1.0), candidate(2, 11, 0, 0.95), anchored, ]; let out = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY); assert_eq!(out.len(), 1); assert_eq!(out[0].person, Some(7)); let mut theirs = candidate(1, 10, 0, 1.0); theirs.confirmed_person = Some(8); let faces = vec![theirs, candidate(2, 11, 0, 0.95), { let mut a = probe(3, 12, 0, 0.92); a.confirmed_person = Some(7); a }]; let out = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY); assert!( out.iter() .any(|c| c.members == vec![2] && c.person == Some(7)), "a probe confirmed as one person joined another's group: {out:?}" ); } /// The co-occurrence rule follows a probe in: once it has joined, its /// photograph is the group's. #[test] fn a_probe_cannot_join_a_group_holding_a_face_from_its_own_photograph() { let faces = vec![ candidate(1, 10, 0, 1.0), candidate(2, 11, 0, 0.95), probe(3, 10, 0, 0.92), ]; let out = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY); assert!(out.iter().any(|c| c.members == vec![2]), "{out:?}"); } /// A probe's placement is scored like anyone else's, from the references /// it matched — and the references' own scores do not hear from it. #[test] fn a_probe_is_scored_but_is_not_evidence() { let gallery_only = vec![candidate(1, 10, 0, 1.0), candidate(2, 11, 0, 0.95)]; let without = cluster_scored(&gallery_only, &cal(), DEFAULT_MERGE_PROBABILITY); let mut with_probe = gallery_only.clone(); with_probe.push(probe(3, 12, 0, 0.99)); let with = cluster_scored(&with_probe, &cal(), DEFAULT_MERGE_PROBABILITY); assert_eq!(with.clusters[0].members, vec![0, 1, 2]); assert!(with.confidence[2] > 0.9, "{}", with.confidence[2]); assert_eq!( &with.confidence[..2], &without.confidence[..], "a probe changed what the references were sure of" ); } }