Cluster faces into people, and calibrate what a similarity means
FR-CULL-9 forbids thresholding a bare cosine anywhere in the subsystem, so calibrate fits P(same person) per library and reports whether the fit is trustworthy. Two details carry most of the weight. The fit runs against a 200-bin histogram rather than a pair list: a 25,000-face library has ~3e8 pairs and no gradient descent is running over that. And a fresh library has no valid calibration, because the positives have to come from user confirmations or burst siblings -- bootstrapping them from high cosine would fit the calibration to the belief it was supposed to test. Clustering defends against the over-merging FR-CULL-10 warns about with constraints rather than a better threshold: two faces in one photograph never merge, and two groups confirmed as different people never merge. Average link rather than single link, so one strong edge cannot weld two families together. Calibration is defined once, in dr-face, and dr-catalog re-exports it. Two implementations of one probability model is exactly how a number comes to mean the wrong thing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,436 @@
|
||||
//! Cosine to probability (docs/faces.md §8, FR-CULL-9).
|
||||
//!
|
||||
//! FR-CULL-9 is a hard requirement rather than an implementation detail: no
|
||||
//! code path may threshold a bare cosine, every threshold in the subsystem is
|
||||
//! stated as a probability, and the fit is per library and reports its own
|
||||
//! validity. The failure it guards against is invisible — a raw cosine means
|
||||
//! something different for every model, every population and every face size,
|
||||
//! and an uncalibrated similarity still *looks* like a plausible number all the
|
||||
//! way to the user interface.
|
||||
//!
|
||||
//! Model-free, so the part of this subsystem most likely to be subtly wrong is
|
||||
//! testable on synthetic embeddings with no weights on the machine.
|
||||
//!
|
||||
//! # Where the training pairs come from
|
||||
//!
|
||||
//! **Negatives are free and abundant.** Two faces detected in *the same
|
||||
//! photograph* are almost never the same person, which hands every multi-face
|
||||
//! image in the library a full set of negative pairs at no labelling cost — and
|
||||
//! they are *hard* negatives, from the same camera, lighting and processing,
|
||||
//! which is exactly the population where a threshold tuned on easy negatives
|
||||
//! fails. The exceptions (mirrors, photographs of photographs, collages) are
|
||||
//! rare enough to be noise at this scale.
|
||||
//!
|
||||
//! **Positives have to be earned.** In order of trustworthiness: pairs the user
|
||||
//! has confirmed onto one person; then burst siblings, since FR-CULL-5 already
|
||||
//! groups bursts and two faces in adjacent frames are near-certainly the same
|
||||
//! person. Nothing else — bootstrapping positives from high cosine is circular,
|
||||
//! fitting the calibration to the belief it was supposed to test.
|
||||
//!
|
||||
//! Which is why a fresh library has **no valid calibration**, and says so.
|
||||
|
||||
/// Bins over cosine ∈ [-1, 1].
|
||||
///
|
||||
/// 200 is the reference implementation's figure and the resolution is not
|
||||
/// critical; what matters is that there *is* a histogram. See [`Pairs`].
|
||||
const BINS: usize = 200;
|
||||
|
||||
/// Minimum evidence before a fit is trusted.
|
||||
///
|
||||
/// Far stricter than the reference implementation's floor of two positives and
|
||||
/// one negative. That floor is reasonable there: its pairs come from a curated
|
||||
/// gallery of labelled reference portraits, where a positive pair is
|
||||
/// trustworthy by construction. Here the positives are bootstrapped from bursts
|
||||
/// and a handful of early confirmations, and the whole risk is fitting
|
||||
/// confidently to too few of them.
|
||||
pub const MIN_POSITIVE_PAIRS: u64 = 200;
|
||||
pub const MIN_NEGATIVE_PAIRS: u64 = 2_000;
|
||||
|
||||
/// A fitted `P(same person | cosine, face size)`.
|
||||
///
|
||||
/// The single definition of what a similarity means in this subsystem. The
|
||||
/// catalog stores its parameters; nothing re-implements the sigmoid.
|
||||
#[derive(Debug, Clone, Copy, PartialEq)]
|
||||
pub struct Calibration {
|
||||
pub a: f32,
|
||||
pub b: f32,
|
||||
/// Weight on `log2(min crop_px)` — the face-size term FR-CULL-9 asks for.
|
||||
pub w_size: f32,
|
||||
/// Whether there was enough evidence to trust the fit.
|
||||
///
|
||||
/// When false the UI says confidence is unavailable. It does **not** present
|
||||
/// an untuned default as though it were measured, which is the distinction
|
||||
/// FR-CULL-9 spends a paragraph on.
|
||||
pub valid: bool,
|
||||
pub positive_pairs: u64,
|
||||
pub negative_pairs: u64,
|
||||
}
|
||||
|
||||
impl Default for Calibration {
|
||||
/// The reference implementation's fitted MBF curve (docs/faces.md §1):
|
||||
/// steepness 16.2, P=0.5 at cosine 0.267.
|
||||
///
|
||||
/// **`valid` is false**, and that is the point. This exists so an
|
||||
/// un-calibrated library has a documented operating point to cluster at
|
||||
/// rather than no behaviour at all — but nothing may show its output as a
|
||||
/// measured confidence.
|
||||
fn default() -> Self {
|
||||
Self {
|
||||
a: 16.2,
|
||||
b: -16.2 * 0.267,
|
||||
w_size: 0.0,
|
||||
valid: false,
|
||||
positive_pairs: 0,
|
||||
negative_pairs: 0,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
impl Calibration {
|
||||
/// P(same person), shifted by a base-rate prior.
|
||||
///
|
||||
/// `log_prior_odds` is applied at evaluation rather than folded into the
|
||||
/// fit, so one stored calibration serves every context: the odds that two
|
||||
/// faces in a 40-image album match are not the odds in a 40,000-image
|
||||
/// archive. Folding a prior in would need a refit per context and would
|
||||
/// make the stored parameters mean different things depending on where they
|
||||
/// came from.
|
||||
pub fn probability(&self, cosine: f32, min_crop_px: f32, log_prior_odds: f32) -> f32 {
|
||||
sigmoid(self.logit(cosine, min_crop_px) + log_prior_odds)
|
||||
}
|
||||
|
||||
fn logit(&self, cosine: f32, min_crop_px: f32) -> f32 {
|
||||
self.a * cosine + self.b + self.w_size * min_crop_px.max(1.0).log2()
|
||||
}
|
||||
|
||||
/// The cosine at which [`Calibration::probability`] crosses `p`.
|
||||
///
|
||||
/// What turns "merge above 0.9" into one comparison against a stored
|
||||
/// similarity, rather than a sigmoid evaluated per candidate edge.
|
||||
pub fn boundary_at(&self, p: f32, min_crop_px: f32, log_prior_odds: f32) -> f32 {
|
||||
((p / (1.0 - p)).ln()
|
||||
- self.b
|
||||
- self.w_size * min_crop_px.max(1.0).log2()
|
||||
- log_prior_odds)
|
||||
/ self.a
|
||||
}
|
||||
}
|
||||
|
||||
fn sigmoid(z: f32) -> f32 {
|
||||
// Branch on the sign so neither tail overflows: exp(-z) for large positive
|
||||
// z, exp(z) for large negative.
|
||||
if z >= 0.0 {
|
||||
1.0 / (1.0 + (-z).exp())
|
||||
} else {
|
||||
let e = z.exp();
|
||||
e / (1.0 + e)
|
||||
}
|
||||
}
|
||||
|
||||
/// Accumulated pair evidence, as a histogram rather than a list.
|
||||
///
|
||||
/// # Why a histogram
|
||||
///
|
||||
/// A 25,000-face library has ~3×10⁸ pairs and no gradient descent is running
|
||||
/// over that. Bucketing them costs 200 counters per class and reduces the fit
|
||||
/// to two parameters against per-bin totals; the expensive part becomes the
|
||||
/// similarity matrix, which is one blocked GEMM. This is the trick that makes a
|
||||
/// per-library fit affordable at all, and it is not obvious from outside.
|
||||
#[derive(Debug, Clone)]
|
||||
pub struct Pairs {
|
||||
positive: Vec<f64>,
|
||||
negative: Vec<f64>,
|
||||
n_pos: u64,
|
||||
n_neg: u64,
|
||||
}
|
||||
|
||||
impl Default for Pairs {
|
||||
fn default() -> Self {
|
||||
Self::new()
|
||||
}
|
||||
}
|
||||
|
||||
impl Pairs {
|
||||
pub fn new() -> Self {
|
||||
Self {
|
||||
positive: vec![0.0; BINS],
|
||||
negative: vec![0.0; BINS],
|
||||
n_pos: 0,
|
||||
n_neg: 0,
|
||||
}
|
||||
}
|
||||
|
||||
/// Record a pair known to be the same person.
|
||||
pub fn push_positive(&mut self, cosine: f32) {
|
||||
self.positive[bin(cosine)] += 1.0;
|
||||
self.n_pos += 1;
|
||||
}
|
||||
|
||||
/// Record a pair known to be different people.
|
||||
pub fn push_negative(&mut self, cosine: f32) {
|
||||
self.negative[bin(cosine)] += 1.0;
|
||||
self.n_neg += 1;
|
||||
}
|
||||
|
||||
pub fn positives(&self) -> u64 {
|
||||
self.n_pos
|
||||
}
|
||||
pub fn negatives(&self) -> u64 {
|
||||
self.n_neg
|
||||
}
|
||||
|
||||
/// Fit `P(same) = σ(a·cos + b)` by weighted logistic regression.
|
||||
///
|
||||
/// Class weights are explicit because negatives outnumber positives by
|
||||
/// orders of magnitude, and an unweighted fit produces a well-shaped curve
|
||||
/// sitting at the wrong height — precisely the "plausible number all the
|
||||
/// way to the user interface" failure FR-CULL-9 describes.
|
||||
///
|
||||
/// Returns a calibration with `valid` set only if there was enough
|
||||
/// evidence; the parameters are filled in either way so a caller with no
|
||||
/// better option can still cluster at a documented operating point.
|
||||
pub fn fit(&self) -> Calibration {
|
||||
let base = Calibration {
|
||||
positive_pairs: self.n_pos,
|
||||
negative_pairs: self.n_neg,
|
||||
..Calibration::default()
|
||||
};
|
||||
if self.n_pos < MIN_POSITIVE_PAIRS || self.n_neg < MIN_NEGATIVE_PAIRS {
|
||||
return base;
|
||||
}
|
||||
|
||||
let total = self.n_pos as f64 + self.n_neg as f64;
|
||||
let w_pos = total / (2.0 * self.n_pos as f64);
|
||||
let w_neg = total / (2.0 * self.n_neg as f64);
|
||||
|
||||
// Start from the reference's fitted MBF curve rather than from zero:
|
||||
// it is the right order of magnitude for every model in this family,
|
||||
// so descent converges in far fewer steps and cannot wander into a
|
||||
// sign-flipped solution on thin evidence.
|
||||
let mut a = base.a as f64;
|
||||
let mut b = base.b as f64;
|
||||
const LR: f64 = 0.05;
|
||||
const MAX_ITER: usize = 20_000;
|
||||
const TOL: f64 = 1e-7;
|
||||
|
||||
for _ in 0..MAX_ITER {
|
||||
let (mut da, mut db) = (0.0, 0.0);
|
||||
for i in 0..BINS {
|
||||
let x = bin_centre(i) as f64;
|
||||
let s = 1.0 / (1.0 + (-(a * x + b)).exp());
|
||||
if self.positive[i] > 0.0 {
|
||||
let e = (s - 1.0) * w_pos * self.positive[i];
|
||||
da += e * x;
|
||||
db += e;
|
||||
}
|
||||
if self.negative[i] > 0.0 {
|
||||
let e = s * w_neg * self.negative[i];
|
||||
da += e * x;
|
||||
db += e;
|
||||
}
|
||||
}
|
||||
da /= total;
|
||||
db /= total;
|
||||
a -= LR * da;
|
||||
b -= LR * db;
|
||||
if da * da + db * db < TOL * TOL {
|
||||
break;
|
||||
}
|
||||
}
|
||||
|
||||
Calibration {
|
||||
a: a as f32,
|
||||
b: b as f32,
|
||||
w_size: 0.0,
|
||||
valid: true,
|
||||
positive_pairs: self.n_pos,
|
||||
negative_pairs: self.n_neg,
|
||||
}
|
||||
}
|
||||
|
||||
/// How well the fit predicts the evidence, as a reliability diagram.
|
||||
///
|
||||
/// FR-CULL-9's acceptance criterion is exactly this and not a single
|
||||
/// accuracy figure: for each populated probability band, the observed match
|
||||
/// rate against the predicted one. Returned rather than asserted so the
|
||||
/// caller can show it, log it, or fail a test on it.
|
||||
pub fn reliability(&self, cal: &Calibration, bands: usize) -> Vec<ReliabilityBand> {
|
||||
let mut out = vec![
|
||||
ReliabilityBand {
|
||||
predicted: 0.0,
|
||||
observed: 0.0,
|
||||
count: 0
|
||||
};
|
||||
bands
|
||||
];
|
||||
let mut sum_pred = vec![0.0_f64; bands];
|
||||
for i in 0..BINS {
|
||||
let n_pos = self.positive[i];
|
||||
let n_neg = self.negative[i];
|
||||
if n_pos + n_neg == 0.0 {
|
||||
continue;
|
||||
}
|
||||
let p = cal.probability(bin_centre(i), 112.0, 0.0) as f64;
|
||||
let band = ((p * bands as f64) as usize).min(bands - 1);
|
||||
sum_pred[band] += p * (n_pos + n_neg);
|
||||
out[band].observed += n_pos as f32;
|
||||
out[band].count += (n_pos + n_neg) as u64;
|
||||
}
|
||||
for (band, o) in out.iter_mut().enumerate() {
|
||||
if o.count > 0 {
|
||||
o.predicted = (sum_pred[band] / o.count as f64) as f32;
|
||||
o.observed /= o.count as f32;
|
||||
}
|
||||
}
|
||||
out
|
||||
}
|
||||
}
|
||||
|
||||
/// One row of a reliability diagram.
|
||||
#[derive(Debug, Clone, Copy, PartialEq)]
|
||||
pub struct ReliabilityBand {
|
||||
/// Mean probability the calibration predicted for pairs in this band.
|
||||
pub predicted: f32,
|
||||
/// Fraction of them that were actually the same person.
|
||||
pub observed: f32,
|
||||
pub count: u64,
|
||||
}
|
||||
|
||||
fn bin(cosine: f32) -> usize {
|
||||
let width = 2.0 / BINS as f32;
|
||||
(((cosine + 1.0) / width) as usize).min(BINS - 1)
|
||||
}
|
||||
|
||||
fn bin_centre(i: usize) -> f32 {
|
||||
let width = 2.0 / BINS as f32;
|
||||
-1.0 + (i as f32 + 0.5) * width
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
/// Synthesise pairs from two well-separated cosine distributions, the way
|
||||
/// a real embedding space behaves: positives near 0.6, negatives near 0.05
|
||||
/// — the numbers our own end-to-end run actually produced.
|
||||
fn realistic_pairs(n_pos: u64, n_neg: u64) -> Pairs {
|
||||
let mut p = Pairs::new();
|
||||
let mut s = 12345_u32;
|
||||
let mut rand = move || {
|
||||
s = s.wrapping_mul(1_664_525).wrapping_add(1_013_904_223);
|
||||
(s >> 8) as f32 / (1u32 << 24) as f32
|
||||
};
|
||||
for _ in 0..n_pos {
|
||||
// ~N(0.60, 0.12), by summing uniforms.
|
||||
let g = (0..4).map(|_| rand()).sum::<f32>() / 4.0 - 0.5;
|
||||
p.push_positive((0.60 + g * 0.48).clamp(-1.0, 1.0));
|
||||
}
|
||||
for _ in 0..n_neg {
|
||||
let g = (0..4).map(|_| rand()).sum::<f32>() / 4.0 - 0.5;
|
||||
p.push_negative((0.05 + g * 0.32).clamp(-1.0, 1.0));
|
||||
}
|
||||
p
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn an_empty_library_has_no_valid_calibration() {
|
||||
let cal = Pairs::new().fit();
|
||||
assert!(!cal.valid, "a fit with no evidence must not claim validity");
|
||||
assert_eq!(cal.positive_pairs, 0);
|
||||
}
|
||||
|
||||
/// The exact case FR-CULL-9 legislates: enough negatives, too few
|
||||
/// positives. The answer is "unavailable", not a plausible-looking curve.
|
||||
#[test]
|
||||
fn too_few_positives_is_invalid_however_many_negatives_there_are() {
|
||||
let p = realistic_pairs(MIN_POSITIVE_PAIRS - 1, MIN_NEGATIVE_PAIRS * 10);
|
||||
assert!(!p.fit().valid);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn too_few_negatives_is_invalid_too() {
|
||||
let p = realistic_pairs(MIN_POSITIVE_PAIRS * 10, MIN_NEGATIVE_PAIRS - 1);
|
||||
assert!(!p.fit().valid);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_well_separated_library_fits_a_usable_curve() {
|
||||
let cal = realistic_pairs(2_000, 40_000).fit();
|
||||
assert!(cal.valid);
|
||||
assert!(cal.a > 0.0, "steepness must be positive: {}", cal.a);
|
||||
|
||||
// The decision boundary lands between the two populations.
|
||||
let boundary = cal.boundary_at(0.5, 112.0, 0.0);
|
||||
assert!(
|
||||
boundary > 0.05 && boundary < 0.60,
|
||||
"boundary {boundary} is not between the negative and positive modes"
|
||||
);
|
||||
|
||||
// And the measured cosines from the real end-to-end run fall the
|
||||
// right side of it.
|
||||
assert!(cal.probability(0.596, 200.0, 0.0) > 0.9);
|
||||
assert!(cal.probability(0.050, 200.0, 0.0) < 0.1);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn probability_and_boundary_are_inverses() {
|
||||
let cal = realistic_pairs(2_000, 40_000).fit();
|
||||
for &p in &[0.1_f32, 0.5, 0.9, 0.99] {
|
||||
let cos = cal.boundary_at(p, 112.0, 0.0);
|
||||
assert!((cal.probability(cos, 112.0, 0.0) - p).abs() < 1e-3);
|
||||
}
|
||||
}
|
||||
|
||||
/// A base rate shifts the answer without a refit — the property that lets
|
||||
/// one stored calibration serve a small album and a large archive.
|
||||
#[test]
|
||||
fn a_prior_moves_the_boundary_in_the_right_direction() {
|
||||
let cal = realistic_pairs(2_000, 40_000).fit();
|
||||
let neutral = cal.probability(0.4, 112.0, 0.0);
|
||||
let pessimistic = cal.probability(0.4, 112.0, -2.0);
|
||||
let optimistic = cal.probability(0.4, 112.0, 2.0);
|
||||
assert!(pessimistic < neutral && neutral < optimistic);
|
||||
}
|
||||
|
||||
/// FR-CULL-9's acceptance criterion, run against the fit's own evidence:
|
||||
/// in every populated band, the stated probability should track the
|
||||
/// observed match rate.
|
||||
#[test]
|
||||
fn the_fit_is_reliable_on_the_evidence_it_was_fitted_to() {
|
||||
let pairs = realistic_pairs(4_000, 40_000);
|
||||
let cal = pairs.fit();
|
||||
let bands = pairs.reliability(&cal, 10);
|
||||
|
||||
let mut checked = 0;
|
||||
for b in &bands {
|
||||
// Thinly populated bands are noise, not evidence.
|
||||
if b.count < 200 {
|
||||
continue;
|
||||
}
|
||||
checked += 1;
|
||||
assert!(
|
||||
(b.predicted - b.observed).abs() < 0.15,
|
||||
"band predicted {:.3} but observed {:.3} over {} pairs",
|
||||
b.predicted,
|
||||
b.observed,
|
||||
b.count
|
||||
);
|
||||
}
|
||||
assert!(checked >= 2, "only {checked} bands had enough pairs to check");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn the_default_curve_is_the_references_and_is_not_marked_valid() {
|
||||
let cal = Calibration::default();
|
||||
assert!(!cal.valid);
|
||||
assert!((cal.boundary_at(0.5, 112.0, 0.0) - 0.267).abs() < 1e-3);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn bins_cover_the_cosine_range_without_overflowing() {
|
||||
assert_eq!(bin(-1.0), 0);
|
||||
assert_eq!(bin(1.0), BINS - 1);
|
||||
assert_eq!(bin(2.0), BINS - 1, "an out-of-range cosine must not panic");
|
||||
assert!((bin_centre(bin(0.5)) - 0.5).abs() < 0.01);
|
||||
}
|
||||
}
|
||||
Reference in New Issue
Block a user