Files
DarkRoom/core/dr-face/src/calibrate.rs
T
dtourolle 43652bd613 Say where the positives really come from, not where they were going to
`calibrate.rs` claimed its positive pairs came from user confirmations "then
burst siblings, since FR-CULL-5 already groups bursts", and repeated it beside
the pair floor: "the positives are bootstrapped from bursts and a handful of
early confirmations". Neither is true and neither ever has been. Nothing in
the workspace pushes a burst pair into a `Pairs`; nothing pushes any pair at
all outside this file's own tests. The sentence was written while both halves
of docs/faces.md §8.1 were being planned together, describing a source that
was going to exist, and it has read since as a description of what the code
does.

The distinction matters more here than in most comments, because this file is
the one place in the subsystem allowed to say what a similarity *means*. A
reader who believes the fit is drawing on bursts believes a young library is
gathering positives on its own, which is precisely the opposite of the state
FR-CULL-9 legislates for — a library with no fit, no valid calibration, and a
reference curve it must not present as a measurement of itself. The floor of
200 positive pairs looks arbitrary under the wrong story and obvious under the
right one: confirmations arrive one at a time, from a person.

So the comment now says what is here — a positive is a pair confirmed onto one
person, and there is no second source — and keeps the burst idea where it
belongs, as §8.1's proposal, with the two reasons it is not in the code: this
crate is handed cosines and cannot see a catalog, and the purity of a burst
pair is a thing to measure before it is a thing to trust.

No behaviour changes; the arithmetic is untouched.
2026-09-06 19:01:50 +02:00

453 lines
17 KiB
Rust
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
//! Cosine to probability (docs/faces.md §8, FR-CULL-9).
//!
//! FR-CULL-9 is a hard requirement rather than an implementation detail: no
//! code path may threshold a bare cosine, every threshold in the subsystem is
//! stated as a probability, and the fit is per library and reports its own
//! validity. The failure it guards against is invisible — a raw cosine means
//! something different for every model, every population and every face size,
//! and an uncalibrated similarity still *looks* like a plausible number all the
//! way to the user interface.
//!
//! Model-free, so the part of this subsystem most likely to be subtly wrong is
//! testable on synthetic embeddings with no weights on the machine.
//!
//! # Where the training pairs come from
//!
//! **Negatives are free and abundant.** Two faces detected in *the same
//! photograph* are almost never the same person, which hands every multi-face
//! image in the library a full set of negative pairs at no labelling cost — and
//! they are *hard* negatives, from the same camera, lighting and processing,
//! which is exactly the population where a threshold tuned on easy negatives
//! fails. The exceptions (mirrors, photographs of photographs, collages) are
//! rare enough to be noise at this scale.
//!
//! **Positives have to be earned.** A positive is a pair of faces the user has
//! confirmed onto one person (FR-CULL-10), and there is no second source: the
//! only labelling this subsystem has is the labelling somebody did by hand.
//! Bootstrapping positives from a high cosine is circular — it fits the
//! calibration to the belief it was supposed to test — and that is the whole
//! of the alternative.
//!
//! docs/faces.md §8.1 names one more that would cost no labelling at all: two
//! faces in adjacent frames of one burst are near-certainly the same person,
//! and FR-CULL-5's grouping is sitting there. Nothing draws on it. This crate
//! cannot see a catalog, let alone the bursts in one — it is handed cosines by
//! whoever assembled the pair — and no caller does that assembly on its behalf
//! yet. Until one does, and until the purity of a burst pair is *measured*
//! rather than assumed, the positives are the confirmations and nothing else.
//!
//! Which is why a fresh library has **no valid calibration** — no fit of its
//! own — and says so. It is not left without a curve: it uses the reference
//! implementation's fitted one ([`Calibration::default`]), which is a published
//! operating point rather than an invention, and the interface reports which of
//! the two it is speaking from.
/// Bins over cosine ∈ [-1, 1].
///
/// 200 is the reference implementation's figure and the resolution is not
/// critical; what matters is that there *is* a histogram. See [`Pairs`].
const BINS: usize = 200;
/// Minimum evidence before a fit is trusted.
///
/// Far stricter than the reference implementation's floor of two positives and
/// one negative. That floor is reasonable there: its pairs come from a curated
/// gallery of labelled reference portraits, where a positive pair is
/// trustworthy by construction. Here every positive is a pair somebody
/// confirmed while working through a young library's suggestions — a handful,
/// arriving slowly — and the whole risk is fitting confidently to too few of
/// them.
pub const MIN_POSITIVE_PAIRS: u64 = 200;
pub const MIN_NEGATIVE_PAIRS: u64 = 2_000;
/// A fitted `P(same person | cosine, face size)`.
///
/// The single definition of what a similarity means in this subsystem. The
/// catalog stores its parameters; nothing re-implements the sigmoid.
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct Calibration {
pub a: f32,
pub b: f32,
/// Weight on `log2(min crop_px)` — the face-size term FR-CULL-9 asks for.
pub w_size: f32,
/// Whether there was enough evidence to fit this library's own curve.
///
/// False means the numbers came from the built-in reference curve, and the
/// UI says so — once, at the screen level. It does **not** mean confidences
/// are withheld: FR-CULL-9's distinction is between presenting an untuned
/// default *as though it were measured* and presenting it as what it is,
/// and only the first is forbidden.
pub valid: bool,
pub positive_pairs: u64,
pub negative_pairs: u64,
}
impl Default for Calibration {
/// The reference implementation's fitted MBF curve (docs/faces.md §1):
/// steepness 16.2, P=0.5 at cosine 0.267.
///
/// **`valid` is false**, and that is the point. It is a documented
/// operating point rather than an invented one, so a library with no fit of
/// its own can both cluster and quote a probability from it — what it may
/// not do is call that probability a measurement of *this* library.
fn default() -> Self {
Self {
a: 16.2,
b: -16.2 * 0.267,
w_size: 0.0,
valid: false,
positive_pairs: 0,
negative_pairs: 0,
}
}
}
impl Calibration {
/// P(same person), shifted by a base-rate prior.
///
/// `log_prior_odds` is applied at evaluation rather than folded into the
/// fit, so one stored calibration serves every context: the odds that two
/// faces in a 40-image album match are not the odds in a 40,000-image
/// archive. Folding a prior in would need a refit per context and would
/// make the stored parameters mean different things depending on where they
/// came from.
pub fn probability(&self, cosine: f32, min_crop_px: f32, log_prior_odds: f32) -> f32 {
sigmoid(self.logit(cosine, min_crop_px) + log_prior_odds)
}
fn logit(&self, cosine: f32, min_crop_px: f32) -> f32 {
self.a * cosine + self.b + self.w_size * min_crop_px.max(1.0).log2()
}
/// The cosine at which [`Calibration::probability`] crosses `p`.
///
/// What turns "merge above 0.9" into one comparison against a stored
/// similarity, rather than a sigmoid evaluated per candidate edge.
pub fn boundary_at(&self, p: f32, min_crop_px: f32, log_prior_odds: f32) -> f32 {
((p / (1.0 - p)).ln() - self.b - self.w_size * min_crop_px.max(1.0).log2() - log_prior_odds)
/ self.a
}
}
fn sigmoid(z: f32) -> f32 {
// Branch on the sign so neither tail overflows: exp(-z) for large positive
// z, exp(z) for large negative.
if z >= 0.0 {
1.0 / (1.0 + (-z).exp())
} else {
let e = z.exp();
e / (1.0 + e)
}
}
/// Accumulated pair evidence, as a histogram rather than a list.
///
/// # Why a histogram
///
/// A 25,000-face library has ~3×10⁸ pairs and no gradient descent is running
/// over that. Bucketing them costs 200 counters per class and reduces the fit
/// to two parameters against per-bin totals; the expensive part becomes the
/// similarity matrix, which is one blocked GEMM. This is the trick that makes a
/// per-library fit affordable at all, and it is not obvious from outside.
#[derive(Debug, Clone)]
pub struct Pairs {
positive: Vec<f64>,
negative: Vec<f64>,
n_pos: u64,
n_neg: u64,
}
impl Default for Pairs {
fn default() -> Self {
Self::new()
}
}
impl Pairs {
pub fn new() -> Self {
Self {
positive: vec![0.0; BINS],
negative: vec![0.0; BINS],
n_pos: 0,
n_neg: 0,
}
}
/// Record a pair known to be the same person.
pub fn push_positive(&mut self, cosine: f32) {
self.positive[bin(cosine)] += 1.0;
self.n_pos += 1;
}
/// Record a pair known to be different people.
pub fn push_negative(&mut self, cosine: f32) {
self.negative[bin(cosine)] += 1.0;
self.n_neg += 1;
}
pub fn positives(&self) -> u64 {
self.n_pos
}
pub fn negatives(&self) -> u64 {
self.n_neg
}
/// Fit `P(same) = σ(a·cos + b)` by weighted logistic regression.
///
/// Class weights are explicit because negatives outnumber positives by
/// orders of magnitude, and an unweighted fit produces a well-shaped curve
/// sitting at the wrong height — precisely the "plausible number all the
/// way to the user interface" failure FR-CULL-9 describes.
///
/// Returns a calibration with `valid` set only if there was enough
/// evidence; the parameters are filled in either way so a caller with no
/// better option can still cluster at a documented operating point.
pub fn fit(&self) -> Calibration {
let base = Calibration {
positive_pairs: self.n_pos,
negative_pairs: self.n_neg,
..Calibration::default()
};
if self.n_pos < MIN_POSITIVE_PAIRS || self.n_neg < MIN_NEGATIVE_PAIRS {
return base;
}
let total = self.n_pos as f64 + self.n_neg as f64;
let w_pos = total / (2.0 * self.n_pos as f64);
let w_neg = total / (2.0 * self.n_neg as f64);
// Start from the reference's fitted MBF curve rather than from zero:
// it is the right order of magnitude for every model in this family,
// so descent converges in far fewer steps and cannot wander into a
// sign-flipped solution on thin evidence.
let mut a = base.a as f64;
let mut b = base.b as f64;
const LR: f64 = 0.05;
const MAX_ITER: usize = 20_000;
const TOL: f64 = 1e-7;
for _ in 0..MAX_ITER {
let (mut da, mut db) = (0.0, 0.0);
for i in 0..BINS {
let x = bin_centre(i) as f64;
let s = 1.0 / (1.0 + (-(a * x + b)).exp());
if self.positive[i] > 0.0 {
let e = (s - 1.0) * w_pos * self.positive[i];
da += e * x;
db += e;
}
if self.negative[i] > 0.0 {
let e = s * w_neg * self.negative[i];
da += e * x;
db += e;
}
}
da /= total;
db /= total;
a -= LR * da;
b -= LR * db;
if da * da + db * db < TOL * TOL {
break;
}
}
Calibration {
a: a as f32,
b: b as f32,
w_size: 0.0,
valid: true,
positive_pairs: self.n_pos,
negative_pairs: self.n_neg,
}
}
/// How well the fit predicts the evidence, as a reliability diagram.
///
/// FR-CULL-9's acceptance criterion is exactly this and not a single
/// accuracy figure: for each populated probability band, the observed match
/// rate against the predicted one. Returned rather than asserted so the
/// caller can show it, log it, or fail a test on it.
pub fn reliability(&self, cal: &Calibration, bands: usize) -> Vec<ReliabilityBand> {
let mut out = vec![
ReliabilityBand {
predicted: 0.0,
observed: 0.0,
count: 0
};
bands
];
let mut sum_pred = vec![0.0_f64; bands];
for i in 0..BINS {
let n_pos = self.positive[i];
let n_neg = self.negative[i];
if n_pos + n_neg == 0.0 {
continue;
}
let p = cal.probability(bin_centre(i), 112.0, 0.0) as f64;
let band = ((p * bands as f64) as usize).min(bands - 1);
sum_pred[band] += p * (n_pos + n_neg);
out[band].observed += n_pos as f32;
out[band].count += (n_pos + n_neg) as u64;
}
for (band, o) in out.iter_mut().enumerate() {
if o.count > 0 {
o.predicted = (sum_pred[band] / o.count as f64) as f32;
o.observed /= o.count as f32;
}
}
out
}
}
/// One row of a reliability diagram.
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct ReliabilityBand {
/// Mean probability the calibration predicted for pairs in this band.
pub predicted: f32,
/// Fraction of them that were actually the same person.
pub observed: f32,
pub count: u64,
}
fn bin(cosine: f32) -> usize {
let width = 2.0 / BINS as f32;
(((cosine + 1.0) / width) as usize).min(BINS - 1)
}
fn bin_centre(i: usize) -> f32 {
let width = 2.0 / BINS as f32;
-1.0 + (i as f32 + 0.5) * width
}
#[cfg(test)]
mod tests {
use super::*;
/// Synthesise pairs from two well-separated cosine distributions, the way
/// a real embedding space behaves: positives near 0.6, negatives near 0.05
/// — the numbers our own end-to-end run actually produced.
fn realistic_pairs(n_pos: u64, n_neg: u64) -> Pairs {
let mut p = Pairs::new();
let mut s = 12345_u32;
let mut rand = move || {
s = s.wrapping_mul(1_664_525).wrapping_add(1_013_904_223);
(s >> 8) as f32 / (1u32 << 24) as f32
};
for _ in 0..n_pos {
// ~N(0.60, 0.12), by summing uniforms.
let g = (0..4).map(|_| rand()).sum::<f32>() / 4.0 - 0.5;
p.push_positive((0.60 + g * 0.48).clamp(-1.0, 1.0));
}
for _ in 0..n_neg {
let g = (0..4).map(|_| rand()).sum::<f32>() / 4.0 - 0.5;
p.push_negative((0.05 + g * 0.32).clamp(-1.0, 1.0));
}
p
}
#[test]
fn an_empty_library_has_no_valid_calibration() {
let cal = Pairs::new().fit();
assert!(!cal.valid, "a fit with no evidence must not claim validity");
assert_eq!(cal.positive_pairs, 0);
}
/// The exact case FR-CULL-9 legislates: enough negatives, too few
/// positives. The answer is "unavailable", not a plausible-looking curve.
#[test]
fn too_few_positives_is_invalid_however_many_negatives_there_are() {
let p = realistic_pairs(MIN_POSITIVE_PAIRS - 1, MIN_NEGATIVE_PAIRS * 10);
assert!(!p.fit().valid);
}
#[test]
fn too_few_negatives_is_invalid_too() {
let p = realistic_pairs(MIN_POSITIVE_PAIRS * 10, MIN_NEGATIVE_PAIRS - 1);
assert!(!p.fit().valid);
}
#[test]
fn a_well_separated_library_fits_a_usable_curve() {
let cal = realistic_pairs(2_000, 40_000).fit();
assert!(cal.valid);
assert!(cal.a > 0.0, "steepness must be positive: {}", cal.a);
// The decision boundary lands between the two populations.
let boundary = cal.boundary_at(0.5, 112.0, 0.0);
assert!(
boundary > 0.05 && boundary < 0.60,
"boundary {boundary} is not between the negative and positive modes"
);
// And the measured cosines from the real end-to-end run fall the
// right side of it.
assert!(cal.probability(0.596, 200.0, 0.0) > 0.9);
assert!(cal.probability(0.050, 200.0, 0.0) < 0.1);
}
#[test]
fn probability_and_boundary_are_inverses() {
let cal = realistic_pairs(2_000, 40_000).fit();
for &p in &[0.1_f32, 0.5, 0.9, 0.99] {
let cos = cal.boundary_at(p, 112.0, 0.0);
assert!((cal.probability(cos, 112.0, 0.0) - p).abs() < 1e-3);
}
}
/// A base rate shifts the answer without a refit — the property that lets
/// one stored calibration serve a small album and a large archive.
#[test]
fn a_prior_moves_the_boundary_in_the_right_direction() {
let cal = realistic_pairs(2_000, 40_000).fit();
let neutral = cal.probability(0.4, 112.0, 0.0);
let pessimistic = cal.probability(0.4, 112.0, -2.0);
let optimistic = cal.probability(0.4, 112.0, 2.0);
assert!(pessimistic < neutral && neutral < optimistic);
}
/// FR-CULL-9's acceptance criterion, run against the fit's own evidence:
/// in every populated band, the stated probability should track the
/// observed match rate.
#[test]
fn the_fit_is_reliable_on_the_evidence_it_was_fitted_to() {
let pairs = realistic_pairs(4_000, 40_000);
let cal = pairs.fit();
let bands = pairs.reliability(&cal, 10);
let mut checked = 0;
for b in &bands {
// Thinly populated bands are noise, not evidence.
if b.count < 200 {
continue;
}
checked += 1;
assert!(
(b.predicted - b.observed).abs() < 0.15,
"band predicted {:.3} but observed {:.3} over {} pairs",
b.predicted,
b.observed,
b.count
);
}
assert!(
checked >= 2,
"only {checked} bands had enough pairs to check"
);
}
#[test]
fn the_default_curve_is_the_references_and_is_not_marked_valid() {
let cal = Calibration::default();
assert!(!cal.valid);
assert!((cal.boundary_at(0.5, 112.0, 0.0) - 0.267).abs() < 1e-3);
}
#[test]
fn bins_cover_the_cosine_range_without_overflowing() {
assert_eq!(bin(-1.0), 0);
assert_eq!(bin(1.0), BINS - 1);
assert_eq!(bin(2.0), BINS - 1, "an out-of-range cosine must not panic");
assert!((bin_centre(bin(0.5)) - 0.5).abs() < 0.01);
}
}