`calibrate.rs` claimed its positive pairs came from user confirmations "then burst siblings, since FR-CULL-5 already groups bursts", and repeated it beside the pair floor: "the positives are bootstrapped from bursts and a handful of early confirmations". Neither is true and neither ever has been. Nothing in the workspace pushes a burst pair into a `Pairs`; nothing pushes any pair at all outside this file's own tests. The sentence was written while both halves of docs/faces.md §8.1 were being planned together, describing a source that was going to exist, and it has read since as a description of what the code does. The distinction matters more here than in most comments, because this file is the one place in the subsystem allowed to say what a similarity *means*. A reader who believes the fit is drawing on bursts believes a young library is gathering positives on its own, which is precisely the opposite of the state FR-CULL-9 legislates for — a library with no fit, no valid calibration, and a reference curve it must not present as a measurement of itself. The floor of 200 positive pairs looks arbitrary under the wrong story and obvious under the right one: confirmations arrive one at a time, from a person. So the comment now says what is here — a positive is a pair confirmed onto one person, and there is no second source — and keeps the burst idea where it belongs, as §8.1's proposal, with the two reasons it is not in the code: this crate is handed cosines and cannot see a catalog, and the purity of a burst pair is a thing to measure before it is a thing to trust. No behaviour changes; the arithmetic is untouched.
453 lines
17 KiB
Rust
453 lines
17 KiB
Rust
//! Cosine to probability (docs/faces.md §8, FR-CULL-9).
|
||
//!
|
||
//! FR-CULL-9 is a hard requirement rather than an implementation detail: no
|
||
//! code path may threshold a bare cosine, every threshold in the subsystem is
|
||
//! stated as a probability, and the fit is per library and reports its own
|
||
//! validity. The failure it guards against is invisible — a raw cosine means
|
||
//! something different for every model, every population and every face size,
|
||
//! and an uncalibrated similarity still *looks* like a plausible number all the
|
||
//! way to the user interface.
|
||
//!
|
||
//! Model-free, so the part of this subsystem most likely to be subtly wrong is
|
||
//! testable on synthetic embeddings with no weights on the machine.
|
||
//!
|
||
//! # Where the training pairs come from
|
||
//!
|
||
//! **Negatives are free and abundant.** Two faces detected in *the same
|
||
//! photograph* are almost never the same person, which hands every multi-face
|
||
//! image in the library a full set of negative pairs at no labelling cost — and
|
||
//! they are *hard* negatives, from the same camera, lighting and processing,
|
||
//! which is exactly the population where a threshold tuned on easy negatives
|
||
//! fails. The exceptions (mirrors, photographs of photographs, collages) are
|
||
//! rare enough to be noise at this scale.
|
||
//!
|
||
//! **Positives have to be earned.** A positive is a pair of faces the user has
|
||
//! confirmed onto one person (FR-CULL-10), and there is no second source: the
|
||
//! only labelling this subsystem has is the labelling somebody did by hand.
|
||
//! Bootstrapping positives from a high cosine is circular — it fits the
|
||
//! calibration to the belief it was supposed to test — and that is the whole
|
||
//! of the alternative.
|
||
//!
|
||
//! docs/faces.md §8.1 names one more that would cost no labelling at all: two
|
||
//! faces in adjacent frames of one burst are near-certainly the same person,
|
||
//! and FR-CULL-5's grouping is sitting there. Nothing draws on it. This crate
|
||
//! cannot see a catalog, let alone the bursts in one — it is handed cosines by
|
||
//! whoever assembled the pair — and no caller does that assembly on its behalf
|
||
//! yet. Until one does, and until the purity of a burst pair is *measured*
|
||
//! rather than assumed, the positives are the confirmations and nothing else.
|
||
//!
|
||
//! Which is why a fresh library has **no valid calibration** — no fit of its
|
||
//! own — and says so. It is not left without a curve: it uses the reference
|
||
//! implementation's fitted one ([`Calibration::default`]), which is a published
|
||
//! operating point rather than an invention, and the interface reports which of
|
||
//! the two it is speaking from.
|
||
|
||
/// Bins over cosine ∈ [-1, 1].
|
||
///
|
||
/// 200 is the reference implementation's figure and the resolution is not
|
||
/// critical; what matters is that there *is* a histogram. See [`Pairs`].
|
||
const BINS: usize = 200;
|
||
|
||
/// Minimum evidence before a fit is trusted.
|
||
///
|
||
/// Far stricter than the reference implementation's floor of two positives and
|
||
/// one negative. That floor is reasonable there: its pairs come from a curated
|
||
/// gallery of labelled reference portraits, where a positive pair is
|
||
/// trustworthy by construction. Here every positive is a pair somebody
|
||
/// confirmed while working through a young library's suggestions — a handful,
|
||
/// arriving slowly — and the whole risk is fitting confidently to too few of
|
||
/// them.
|
||
pub const MIN_POSITIVE_PAIRS: u64 = 200;
|
||
pub const MIN_NEGATIVE_PAIRS: u64 = 2_000;
|
||
|
||
/// A fitted `P(same person | cosine, face size)`.
|
||
///
|
||
/// The single definition of what a similarity means in this subsystem. The
|
||
/// catalog stores its parameters; nothing re-implements the sigmoid.
|
||
#[derive(Debug, Clone, Copy, PartialEq)]
|
||
pub struct Calibration {
|
||
pub a: f32,
|
||
pub b: f32,
|
||
/// Weight on `log2(min crop_px)` — the face-size term FR-CULL-9 asks for.
|
||
pub w_size: f32,
|
||
/// Whether there was enough evidence to fit this library's own curve.
|
||
///
|
||
/// False means the numbers came from the built-in reference curve, and the
|
||
/// UI says so — once, at the screen level. It does **not** mean confidences
|
||
/// are withheld: FR-CULL-9's distinction is between presenting an untuned
|
||
/// default *as though it were measured* and presenting it as what it is,
|
||
/// and only the first is forbidden.
|
||
pub valid: bool,
|
||
pub positive_pairs: u64,
|
||
pub negative_pairs: u64,
|
||
}
|
||
|
||
impl Default for Calibration {
|
||
/// The reference implementation's fitted MBF curve (docs/faces.md §1):
|
||
/// steepness 16.2, P=0.5 at cosine 0.267.
|
||
///
|
||
/// **`valid` is false**, and that is the point. It is a documented
|
||
/// operating point rather than an invented one, so a library with no fit of
|
||
/// its own can both cluster and quote a probability from it — what it may
|
||
/// not do is call that probability a measurement of *this* library.
|
||
fn default() -> Self {
|
||
Self {
|
||
a: 16.2,
|
||
b: -16.2 * 0.267,
|
||
w_size: 0.0,
|
||
valid: false,
|
||
positive_pairs: 0,
|
||
negative_pairs: 0,
|
||
}
|
||
}
|
||
}
|
||
|
||
impl Calibration {
|
||
/// P(same person), shifted by a base-rate prior.
|
||
///
|
||
/// `log_prior_odds` is applied at evaluation rather than folded into the
|
||
/// fit, so one stored calibration serves every context: the odds that two
|
||
/// faces in a 40-image album match are not the odds in a 40,000-image
|
||
/// archive. Folding a prior in would need a refit per context and would
|
||
/// make the stored parameters mean different things depending on where they
|
||
/// came from.
|
||
pub fn probability(&self, cosine: f32, min_crop_px: f32, log_prior_odds: f32) -> f32 {
|
||
sigmoid(self.logit(cosine, min_crop_px) + log_prior_odds)
|
||
}
|
||
|
||
fn logit(&self, cosine: f32, min_crop_px: f32) -> f32 {
|
||
self.a * cosine + self.b + self.w_size * min_crop_px.max(1.0).log2()
|
||
}
|
||
|
||
/// The cosine at which [`Calibration::probability`] crosses `p`.
|
||
///
|
||
/// What turns "merge above 0.9" into one comparison against a stored
|
||
/// similarity, rather than a sigmoid evaluated per candidate edge.
|
||
pub fn boundary_at(&self, p: f32, min_crop_px: f32, log_prior_odds: f32) -> f32 {
|
||
((p / (1.0 - p)).ln() - self.b - self.w_size * min_crop_px.max(1.0).log2() - log_prior_odds)
|
||
/ self.a
|
||
}
|
||
}
|
||
|
||
fn sigmoid(z: f32) -> f32 {
|
||
// Branch on the sign so neither tail overflows: exp(-z) for large positive
|
||
// z, exp(z) for large negative.
|
||
if z >= 0.0 {
|
||
1.0 / (1.0 + (-z).exp())
|
||
} else {
|
||
let e = z.exp();
|
||
e / (1.0 + e)
|
||
}
|
||
}
|
||
|
||
/// Accumulated pair evidence, as a histogram rather than a list.
|
||
///
|
||
/// # Why a histogram
|
||
///
|
||
/// A 25,000-face library has ~3×10⁸ pairs and no gradient descent is running
|
||
/// over that. Bucketing them costs 200 counters per class and reduces the fit
|
||
/// to two parameters against per-bin totals; the expensive part becomes the
|
||
/// similarity matrix, which is one blocked GEMM. This is the trick that makes a
|
||
/// per-library fit affordable at all, and it is not obvious from outside.
|
||
#[derive(Debug, Clone)]
|
||
pub struct Pairs {
|
||
positive: Vec<f64>,
|
||
negative: Vec<f64>,
|
||
n_pos: u64,
|
||
n_neg: u64,
|
||
}
|
||
|
||
impl Default for Pairs {
|
||
fn default() -> Self {
|
||
Self::new()
|
||
}
|
||
}
|
||
|
||
impl Pairs {
|
||
pub fn new() -> Self {
|
||
Self {
|
||
positive: vec![0.0; BINS],
|
||
negative: vec![0.0; BINS],
|
||
n_pos: 0,
|
||
n_neg: 0,
|
||
}
|
||
}
|
||
|
||
/// Record a pair known to be the same person.
|
||
pub fn push_positive(&mut self, cosine: f32) {
|
||
self.positive[bin(cosine)] += 1.0;
|
||
self.n_pos += 1;
|
||
}
|
||
|
||
/// Record a pair known to be different people.
|
||
pub fn push_negative(&mut self, cosine: f32) {
|
||
self.negative[bin(cosine)] += 1.0;
|
||
self.n_neg += 1;
|
||
}
|
||
|
||
pub fn positives(&self) -> u64 {
|
||
self.n_pos
|
||
}
|
||
pub fn negatives(&self) -> u64 {
|
||
self.n_neg
|
||
}
|
||
|
||
/// Fit `P(same) = σ(a·cos + b)` by weighted logistic regression.
|
||
///
|
||
/// Class weights are explicit because negatives outnumber positives by
|
||
/// orders of magnitude, and an unweighted fit produces a well-shaped curve
|
||
/// sitting at the wrong height — precisely the "plausible number all the
|
||
/// way to the user interface" failure FR-CULL-9 describes.
|
||
///
|
||
/// Returns a calibration with `valid` set only if there was enough
|
||
/// evidence; the parameters are filled in either way so a caller with no
|
||
/// better option can still cluster at a documented operating point.
|
||
pub fn fit(&self) -> Calibration {
|
||
let base = Calibration {
|
||
positive_pairs: self.n_pos,
|
||
negative_pairs: self.n_neg,
|
||
..Calibration::default()
|
||
};
|
||
if self.n_pos < MIN_POSITIVE_PAIRS || self.n_neg < MIN_NEGATIVE_PAIRS {
|
||
return base;
|
||
}
|
||
|
||
let total = self.n_pos as f64 + self.n_neg as f64;
|
||
let w_pos = total / (2.0 * self.n_pos as f64);
|
||
let w_neg = total / (2.0 * self.n_neg as f64);
|
||
|
||
// Start from the reference's fitted MBF curve rather than from zero:
|
||
// it is the right order of magnitude for every model in this family,
|
||
// so descent converges in far fewer steps and cannot wander into a
|
||
// sign-flipped solution on thin evidence.
|
||
let mut a = base.a as f64;
|
||
let mut b = base.b as f64;
|
||
const LR: f64 = 0.05;
|
||
const MAX_ITER: usize = 20_000;
|
||
const TOL: f64 = 1e-7;
|
||
|
||
for _ in 0..MAX_ITER {
|
||
let (mut da, mut db) = (0.0, 0.0);
|
||
for i in 0..BINS {
|
||
let x = bin_centre(i) as f64;
|
||
let s = 1.0 / (1.0 + (-(a * x + b)).exp());
|
||
if self.positive[i] > 0.0 {
|
||
let e = (s - 1.0) * w_pos * self.positive[i];
|
||
da += e * x;
|
||
db += e;
|
||
}
|
||
if self.negative[i] > 0.0 {
|
||
let e = s * w_neg * self.negative[i];
|
||
da += e * x;
|
||
db += e;
|
||
}
|
||
}
|
||
da /= total;
|
||
db /= total;
|
||
a -= LR * da;
|
||
b -= LR * db;
|
||
if da * da + db * db < TOL * TOL {
|
||
break;
|
||
}
|
||
}
|
||
|
||
Calibration {
|
||
a: a as f32,
|
||
b: b as f32,
|
||
w_size: 0.0,
|
||
valid: true,
|
||
positive_pairs: self.n_pos,
|
||
negative_pairs: self.n_neg,
|
||
}
|
||
}
|
||
|
||
/// How well the fit predicts the evidence, as a reliability diagram.
|
||
///
|
||
/// FR-CULL-9's acceptance criterion is exactly this and not a single
|
||
/// accuracy figure: for each populated probability band, the observed match
|
||
/// rate against the predicted one. Returned rather than asserted so the
|
||
/// caller can show it, log it, or fail a test on it.
|
||
pub fn reliability(&self, cal: &Calibration, bands: usize) -> Vec<ReliabilityBand> {
|
||
let mut out = vec![
|
||
ReliabilityBand {
|
||
predicted: 0.0,
|
||
observed: 0.0,
|
||
count: 0
|
||
};
|
||
bands
|
||
];
|
||
let mut sum_pred = vec![0.0_f64; bands];
|
||
for i in 0..BINS {
|
||
let n_pos = self.positive[i];
|
||
let n_neg = self.negative[i];
|
||
if n_pos + n_neg == 0.0 {
|
||
continue;
|
||
}
|
||
let p = cal.probability(bin_centre(i), 112.0, 0.0) as f64;
|
||
let band = ((p * bands as f64) as usize).min(bands - 1);
|
||
sum_pred[band] += p * (n_pos + n_neg);
|
||
out[band].observed += n_pos as f32;
|
||
out[band].count += (n_pos + n_neg) as u64;
|
||
}
|
||
for (band, o) in out.iter_mut().enumerate() {
|
||
if o.count > 0 {
|
||
o.predicted = (sum_pred[band] / o.count as f64) as f32;
|
||
o.observed /= o.count as f32;
|
||
}
|
||
}
|
||
out
|
||
}
|
||
}
|
||
|
||
/// One row of a reliability diagram.
|
||
#[derive(Debug, Clone, Copy, PartialEq)]
|
||
pub struct ReliabilityBand {
|
||
/// Mean probability the calibration predicted for pairs in this band.
|
||
pub predicted: f32,
|
||
/// Fraction of them that were actually the same person.
|
||
pub observed: f32,
|
||
pub count: u64,
|
||
}
|
||
|
||
fn bin(cosine: f32) -> usize {
|
||
let width = 2.0 / BINS as f32;
|
||
(((cosine + 1.0) / width) as usize).min(BINS - 1)
|
||
}
|
||
|
||
fn bin_centre(i: usize) -> f32 {
|
||
let width = 2.0 / BINS as f32;
|
||
-1.0 + (i as f32 + 0.5) * width
|
||
}
|
||
|
||
#[cfg(test)]
|
||
mod tests {
|
||
use super::*;
|
||
|
||
/// Synthesise pairs from two well-separated cosine distributions, the way
|
||
/// a real embedding space behaves: positives near 0.6, negatives near 0.05
|
||
/// — the numbers our own end-to-end run actually produced.
|
||
fn realistic_pairs(n_pos: u64, n_neg: u64) -> Pairs {
|
||
let mut p = Pairs::new();
|
||
let mut s = 12345_u32;
|
||
let mut rand = move || {
|
||
s = s.wrapping_mul(1_664_525).wrapping_add(1_013_904_223);
|
||
(s >> 8) as f32 / (1u32 << 24) as f32
|
||
};
|
||
for _ in 0..n_pos {
|
||
// ~N(0.60, 0.12), by summing uniforms.
|
||
let g = (0..4).map(|_| rand()).sum::<f32>() / 4.0 - 0.5;
|
||
p.push_positive((0.60 + g * 0.48).clamp(-1.0, 1.0));
|
||
}
|
||
for _ in 0..n_neg {
|
||
let g = (0..4).map(|_| rand()).sum::<f32>() / 4.0 - 0.5;
|
||
p.push_negative((0.05 + g * 0.32).clamp(-1.0, 1.0));
|
||
}
|
||
p
|
||
}
|
||
|
||
#[test]
|
||
fn an_empty_library_has_no_valid_calibration() {
|
||
let cal = Pairs::new().fit();
|
||
assert!(!cal.valid, "a fit with no evidence must not claim validity");
|
||
assert_eq!(cal.positive_pairs, 0);
|
||
}
|
||
|
||
/// The exact case FR-CULL-9 legislates: enough negatives, too few
|
||
/// positives. The answer is "unavailable", not a plausible-looking curve.
|
||
#[test]
|
||
fn too_few_positives_is_invalid_however_many_negatives_there_are() {
|
||
let p = realistic_pairs(MIN_POSITIVE_PAIRS - 1, MIN_NEGATIVE_PAIRS * 10);
|
||
assert!(!p.fit().valid);
|
||
}
|
||
|
||
#[test]
|
||
fn too_few_negatives_is_invalid_too() {
|
||
let p = realistic_pairs(MIN_POSITIVE_PAIRS * 10, MIN_NEGATIVE_PAIRS - 1);
|
||
assert!(!p.fit().valid);
|
||
}
|
||
|
||
#[test]
|
||
fn a_well_separated_library_fits_a_usable_curve() {
|
||
let cal = realistic_pairs(2_000, 40_000).fit();
|
||
assert!(cal.valid);
|
||
assert!(cal.a > 0.0, "steepness must be positive: {}", cal.a);
|
||
|
||
// The decision boundary lands between the two populations.
|
||
let boundary = cal.boundary_at(0.5, 112.0, 0.0);
|
||
assert!(
|
||
boundary > 0.05 && boundary < 0.60,
|
||
"boundary {boundary} is not between the negative and positive modes"
|
||
);
|
||
|
||
// And the measured cosines from the real end-to-end run fall the
|
||
// right side of it.
|
||
assert!(cal.probability(0.596, 200.0, 0.0) > 0.9);
|
||
assert!(cal.probability(0.050, 200.0, 0.0) < 0.1);
|
||
}
|
||
|
||
#[test]
|
||
fn probability_and_boundary_are_inverses() {
|
||
let cal = realistic_pairs(2_000, 40_000).fit();
|
||
for &p in &[0.1_f32, 0.5, 0.9, 0.99] {
|
||
let cos = cal.boundary_at(p, 112.0, 0.0);
|
||
assert!((cal.probability(cos, 112.0, 0.0) - p).abs() < 1e-3);
|
||
}
|
||
}
|
||
|
||
/// A base rate shifts the answer without a refit — the property that lets
|
||
/// one stored calibration serve a small album and a large archive.
|
||
#[test]
|
||
fn a_prior_moves_the_boundary_in_the_right_direction() {
|
||
let cal = realistic_pairs(2_000, 40_000).fit();
|
||
let neutral = cal.probability(0.4, 112.0, 0.0);
|
||
let pessimistic = cal.probability(0.4, 112.0, -2.0);
|
||
let optimistic = cal.probability(0.4, 112.0, 2.0);
|
||
assert!(pessimistic < neutral && neutral < optimistic);
|
||
}
|
||
|
||
/// FR-CULL-9's acceptance criterion, run against the fit's own evidence:
|
||
/// in every populated band, the stated probability should track the
|
||
/// observed match rate.
|
||
#[test]
|
||
fn the_fit_is_reliable_on_the_evidence_it_was_fitted_to() {
|
||
let pairs = realistic_pairs(4_000, 40_000);
|
||
let cal = pairs.fit();
|
||
let bands = pairs.reliability(&cal, 10);
|
||
|
||
let mut checked = 0;
|
||
for b in &bands {
|
||
// Thinly populated bands are noise, not evidence.
|
||
if b.count < 200 {
|
||
continue;
|
||
}
|
||
checked += 1;
|
||
assert!(
|
||
(b.predicted - b.observed).abs() < 0.15,
|
||
"band predicted {:.3} but observed {:.3} over {} pairs",
|
||
b.predicted,
|
||
b.observed,
|
||
b.count
|
||
);
|
||
}
|
||
assert!(
|
||
checked >= 2,
|
||
"only {checked} bands had enough pairs to check"
|
||
);
|
||
}
|
||
|
||
#[test]
|
||
fn the_default_curve_is_the_references_and_is_not_marked_valid() {
|
||
let cal = Calibration::default();
|
||
assert!(!cal.valid);
|
||
assert!((cal.boundary_at(0.5, 112.0, 0.0) - 0.267).abs() < 1e-3);
|
||
}
|
||
|
||
#[test]
|
||
fn bins_cover_the_cosine_range_without_overflowing() {
|
||
assert_eq!(bin(-1.0), 0);
|
||
assert_eq!(bin(1.0), BINS - 1);
|
||
assert_eq!(bin(2.0), BINS - 1, "an out-of-range cosine must not panic");
|
||
assert!((bin_centre(bin(0.5)) - 0.5).abs() < 0.01);
|
||
}
|
||
}
|