Group people at the threshold the library actually supports

0.90 left a third of the reference library ungrouped: 1,213 of 1,813 faces in a
group, and the rest sitting alone in a screen that had nothing to offer for
them.

"Is 0.90 too tight" is not answerable from the number. It is a probability, and
which cosine it lands on depends on the calibration — so the first half of this
is a way to ask the question properly. `face_index --tune` runs the real
clusterer over the real embeddings at ten thresholds and prints what each one
produces. It writes nothing; comparing thresholds by applying them would have
each one pollute the next.

On the reference library:

      P   cosine   groups  grouped  largest
   0.95    0.449      311      62%       51
   0.90    0.403      316      67%       51
   0.85    0.374      318      70%       57
   0.80    0.353      328      74%       69
   0.75    0.335      327      77%       69
   0.70    0.319      326      79%       81
   0.50    0.267      303      85%       90

The count of *groups* is the signal, not the count of grouped faces. Loosening
from 0.95 makes it climb: real people are being assembled out of fragments. It
peaks at 0.80 and then falls — and a falling group count while the grouped faces
keep rising is the shape of over-merging, separate identities being welded
together. That is the FR-CULL-10 failure, and the one the user cannot undo by
hand.

So 0.80: the loosest setting still building people rather than melting them
together. A third more of the library gets grouped than at 0.90, and the largest
group grows by eighteen faces rather than by forty.

The table is one library, and the doc comment says so — `--tune` reruns it on
any other.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-27 21:17:16 +02:00
co-authored by Claude Opus 5
parent c10dca984f
commit 7275c020d7
2 changed files with 126 additions and 2 deletions
+30 -1
View File
@@ -73,7 +73,36 @@ use crate::neighbours::{self, Faces};
///
/// Stated as a probability and not a cosine, because FR-CULL-9 forbids
/// thresholding a bare similarity anywhere in this subsystem.
pub const DEFAULT_MERGE_PROBABILITY: f32 = 0.9;
///
/// # Why 0.80
///
/// It was 0.90, and 0.90 left most of a real library ungrouped. Measured over
/// the 1,813-face reference library, with `dr-ui`'s `face_index --tune`:
///
/// | P | cosine | groups | faces grouped | largest group |
/// |---|---|---|---|---|
/// | 0.95 | 0.449 | 311 | 62% | 51 |
/// | 0.90 | 0.403 | 316 | 67% | 51 |
/// | 0.85 | 0.374 | 318 | 70% | 57 |
/// | **0.80** | **0.353** | **328** | **74%** | **69** |
/// | 0.75 | 0.335 | 327 | 77% | 69 |
/// | 0.70 | 0.319 | 326 | 79% | 81 |
/// | 0.50 | 0.267 | 303 | 85% | 90 |
///
/// The count of *groups* is the signal, not the count of grouped faces. Loosen
/// from 0.95 and it climbs: real people are being assembled out of fragments.
/// It peaks at 0.80 and then falls, and a falling group count while the grouped
/// faces keep rising is the shape of over-merging — separate identities being
/// welded, which is the failure FR-CULL-10 warns about and the one the user
/// cannot easily undo by hand.
///
/// So: the loosest setting that is still building people rather than melting
/// them together. A third more of the library gets grouped than at 0.90, and
/// the largest group grows by eighteen faces rather than by forty.
///
/// This is a *default*, not a constant of nature — the numbers above are one
/// library, and `--tune` reruns the table on any other.
pub const DEFAULT_MERGE_PROBABILITY: f32 = 0.80;
/// A face presented to the clusterer.
///
+96 -1
View File
@@ -35,7 +35,8 @@ fn main() {
eprintln!(
"usage: face_index CATALOG.db THUMBS_DIR [--run DETECTOR.onnx EMBEDDER.onnx]\n\
\n\
With no --run this only reports; nothing is written."
With no --run this only reports; nothing is written.\n\
--cluster groups what is indexed; --tune compares thresholds without writing."
);
std::process::exit(2);
}
@@ -99,6 +100,14 @@ fn main() {
return;
}
// Tuning, and deliberately read-only: it answers "what would this
// threshold do to my library" without writing a single suggestion, which
// is the only way to compare several without each one polluting the next.
if args.iter().any(|a| a == "--tune") {
tune_thresholds(&catalog);
return;
}
let run = args.iter().position(|a| a == "--run");
let Some(i) = run else {
if audit.coverage.is_complete() {
@@ -200,3 +209,89 @@ fn report_people(catalog: &Catalog) {
println!(" … and {} more", people.len() - 30);
}
}
/// What several merge thresholds would each do to this library.
///
/// The default 0.9 is a *probability*, and the cosine it lands on depends on
/// the calibration — so "is 0.9 too tight" is not a question anyone can answer
/// from the number alone. This runs the real clusterer over the real
/// embeddings at a range of thresholds and prints what each one produces, which
/// is the only honest way to choose.
///
/// Nothing is written. Run it, read the table, then pass the number you want.
fn tune_thresholds(catalog: &Catalog) {
use dr_catalog::faces;
let conn = catalog.connection();
let cal = match faces::calibration(conn, MODEL_ID) {
Ok(Some((c, _))) => c,
_ => dr_face::Calibration::default(),
};
let stored = match faces::embeddings(conn, MODEL_ID) {
Ok(s) => s,
Err(e) => {
eprintln!("cannot read embeddings: {e}");
std::process::exit(1);
}
};
if stored.is_empty() {
println!("no faces indexed yet — nothing to tune.");
return;
}
let model = dr_face::ModelId::new(MODEL_ID.to_string());
let mut candidates = Vec::with_capacity(stored.len());
for (face_id, image_id, blob, crop_px) in stored {
let Some(emb) = dr_face::Embedding::from_f16_bytes(model.clone(), &blob) else {
continue;
};
candidates.push(dr_face::Candidate {
face: face_id.0,
image: image_id.0,
embedding: emb.v.to_vec(),
crop_px,
confirmed_person: None,
});
}
println!(
"\n{} face(s), calibration valid: {}",
candidates.len(),
cal.valid
);
println!(
"\n{:>6} {:>7} {:>7} {:>7} {:>7} {:>8} {:>7}",
"P", "cosine", "groups", "grouped", "largest", "in groups", "time"
);
println!("{}", "-".repeat(60));
for p in [0.99_f32, 0.97, 0.95, 0.9, 0.85, 0.8, 0.75, 0.7, 0.6, 0.5] {
let start = std::time::Instant::now();
let clusters = dr_face::cluster(&candidates, &cal, p);
let elapsed = start.elapsed();
// A group of one is not a person, and `recluster` discards those, so
// the interesting figures count only the real groups.
let real: Vec<_> = clusters.iter().filter(|c| c.members.len() >= 2).collect();
let grouped: usize = real.iter().map(|c| c.members.len()).sum();
let largest = real.first().map(|c| c.members.len()).unwrap_or(0);
println!(
"{p:>6.2} {:>7.3} {:>7} {:>7} {:>7} {:>7.0}% {:>6.2}s",
cal.boundary_at(p, 150.0, 0.0),
real.len(),
grouped,
largest,
100.0 * grouped as f64 / candidates.len() as f64,
elapsed.as_secs_f64(),
);
}
println!(
"\nA looser threshold makes bigger groups and merges people who are not\n\
the same; a tighter one splits one person across several. The largest\n\
group is the tell: when it starts growing much faster than the rest,\n\
identities are being welded together."
);
}