Stop indexing faces too small or too blurred to be anyone
The library was storing faces at 52 source pixels and embedding whatever came
back. There was a size floor, but it was 40 pixels on the *bounding box*, and
there was no blur gate at all — so a subject walking through a half-second
exposure detected confidently, aligned cleanly, and produced a perfectly
ordinary-looking 512-vector. Nothing downstream can tell that apart from a real
face, and because blurs resemble each other more than they resemble the people
they were, they cluster together and weld unrelated identities into one group.
Two floors, both measured rather than guessed. `face_index --quality` runs the
detector over real proxies with both gates disabled and prints the distribution;
over 1,503 faces in 600 images of the reference library:
percentile crop px sharpness
1% 16 0.0006
25% 23 0.0025
50% 38 0.0071
75% 76 0.0284
99% 352 0.4282
The median face in a personal library is 38 pixels. Most of what the detector
finds is background: people across a square, a face on a poster, a stranger at
the next table. They are real detections and useless identifications.
**Size, on the crop rather than the box.** "At least 64x64" has to mean the
pixels the *embedder* sees, and the box is not that — the ArcFace template
reaches past it for forehead and chin, so the aligned crop spans roughly 1.3x
the box's shorter edge. The floor is therefore `min_source_px` on the aligned
crop, applied after the warp fixes the scale, and `min_face_px` drops to 48 as
what it always really was: a cheap pre-filter set low enough that it cannot
reject a face the real floor would have kept.
**Sharpness.** Variance of the Laplacian divided by the variance of the luma it
was taken over. The division is the part that matters: raw Laplacian variance
scales with contrast, so a threshold on it would quietly discard every backlit
portrait in the library. The ratio asks how much of the crop's variation is
edges rather than broad gradients, and is invariant to exposure.
What each pair removes, cumulatively, of everything the detector finds:
min crop min sharp size cut blur cut kept
64 0.000 70% 0% 30%
64 0.010 70% 3% 27%
64 0.020 70% 7% 23%
80 0.010 76% 2% 21%
64 and 0.020. The size floor does most of the work, and the blur floor removing
only 7% on top of it is the point rather than a disappointment: at 64 pixels
most faces are already sharp, and what it takes out is the large-but-soft one —
precisely the face that would otherwise contribute a confident, wrong embedding.
The two gates are not independent and the doc comments say so: a face under 112
pixels was upsampled to reach the embedder, and upsampling invents no edges, so
small faces score low on sharpness even when the original was crisp. That is why
`--quality` prints them together.
**This will re-index.** Around 70% of what the current settings store falls below
the new floors — faces between 20 and 40 pixels that nobody could identify. The
People screen gets shorter and every group in it gets better.
66 dr-face tests pass, including that a blurred crop scores below a sharp one,
that halving the contrast does not move the score, and that an upsampled face
scores below the same face at full size.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -36,7 +36,8 @@ fn main() {
|
||||
"usage: face_index CATALOG.db THUMBS_DIR [--run DETECTOR.onnx EMBEDDER.onnx]\n\
|
||||
\n\
|
||||
With no --run this only reports; nothing is written.\n\
|
||||
--cluster groups what is indexed; --tune compares thresholds without writing."
|
||||
--cluster groups what is indexed; --tune compares thresholds without writing.\n\
|
||||
--quality DET.onnx EMB.onnx reports face size and sharpness, also without writing."
|
||||
);
|
||||
std::process::exit(2);
|
||||
}
|
||||
@@ -108,6 +109,25 @@ fn main() {
|
||||
return;
|
||||
}
|
||||
|
||||
// Quality tuning, and read-only like `--tune`. Runs the real detector over
|
||||
// real proxies with **both gates disabled**, so the distribution it prints
|
||||
// is of everything the detector finds rather than of what survives the
|
||||
// current settings — which is the only way to see what a threshold would
|
||||
// actually remove.
|
||||
if let Some(i) = args.iter().position(|a| a == "--quality") {
|
||||
let (Some(detector), Some(embedder)) = (args.get(i + 1), args.get(i + 2)) else {
|
||||
eprintln!("--quality needs both a detector and an embedder");
|
||||
std::process::exit(2);
|
||||
};
|
||||
report_quality(
|
||||
&catalog,
|
||||
&store,
|
||||
std::path::Path::new(detector),
|
||||
std::path::Path::new(embedder),
|
||||
);
|
||||
return;
|
||||
}
|
||||
|
||||
let run = args.iter().position(|a| a == "--run");
|
||||
let Some(i) = run else {
|
||||
if audit.coverage.is_complete() {
|
||||
@@ -295,3 +315,182 @@ fn tune_thresholds(catalog: &Catalog) {
|
||||
identities are being welded together."
|
||||
);
|
||||
}
|
||||
|
||||
/// How many images to sample for the quality report.
|
||||
///
|
||||
/// Enough for the distribution to settle, few enough to finish while the user
|
||||
/// is watching: detection is ~100ms an image, so this is a couple of minutes.
|
||||
const QUALITY_SAMPLE: usize = 600;
|
||||
|
||||
/// What the detector finds, before either quality gate is applied.
|
||||
///
|
||||
/// The two floors — face size and sharpness — are not independent: a face
|
||||
/// smaller than the embedder's 112-pixel input was upsampled to reach it, and
|
||||
/// upsampling invents no edges, so small faces score low on sharpness even when
|
||||
/// the original was crisp. Choosing either number without seeing the other is
|
||||
/// how you end up with one gate doing nothing and the other doing too much.
|
||||
///
|
||||
/// So this prints them together, over the real library, with nothing filtered.
|
||||
fn report_quality(
|
||||
catalog: &Catalog,
|
||||
store: &ThumbStore,
|
||||
detector: &std::path::Path,
|
||||
embedder: &std::path::Path,
|
||||
) {
|
||||
let mut det = match dr_face::Detector::from_path(detector) {
|
||||
Ok(d) => d,
|
||||
Err(e) => {
|
||||
eprintln!("cannot load the detector: {e}");
|
||||
std::process::exit(1);
|
||||
}
|
||||
};
|
||||
// Loaded but unused: the point is to fail here, before a two-minute scan,
|
||||
// if the pair the user passed is not the pair indexing would use.
|
||||
if let Err(e) =
|
||||
dr_face::Embedder::from_path(embedder, dr_face::ModelId::new(MODEL_ID.to_string()))
|
||||
{
|
||||
eprintln!("cannot load the embedder: {e}");
|
||||
std::process::exit(1);
|
||||
}
|
||||
|
||||
// Everything the detector can find: no size floor, no sharpness floor.
|
||||
let options = dr_face::DetectOptions {
|
||||
min_face_px: 0.0,
|
||||
min_source_px: 0.0,
|
||||
min_sharpness: 0.0,
|
||||
..Default::default()
|
||||
};
|
||||
|
||||
let mut stmt = match catalog.connection().prepare(
|
||||
"SELECT r.file_id FROM remote r
|
||||
JOIN images i ON i.id = r.image_id
|
||||
WHERE r.file_id IS NOT NULL AND i.trashed_at IS NULL
|
||||
ORDER BY i.id",
|
||||
) {
|
||||
Ok(s) => s,
|
||||
Err(e) => {
|
||||
eprintln!("cannot list images: {e}");
|
||||
std::process::exit(1);
|
||||
}
|
||||
};
|
||||
let file_ids: Vec<u64> = stmt
|
||||
.query_map([], |r| r.get::<_, i64>(0))
|
||||
.into_iter()
|
||||
.flatten()
|
||||
.filter_map(Result::ok)
|
||||
.map(|v| v as u64)
|
||||
.filter(|id| store.contains(*id, faces::FACE_TIER))
|
||||
.take(QUALITY_SAMPLE)
|
||||
.collect();
|
||||
|
||||
if file_ids.is_empty() {
|
||||
println!("no proxies on disk to measure — browse the library first.");
|
||||
return;
|
||||
}
|
||||
println!("\nmeasuring {} image(s)…", file_ids.len());
|
||||
|
||||
// (source_px, sharpness) per detected face.
|
||||
let mut found: Vec<(f32, f32)> = Vec::new();
|
||||
let mut images = 0usize;
|
||||
for id in &file_ids {
|
||||
let Ok(Some(thumb)) = store.get(*id, faces::FACE_TIER) else {
|
||||
continue;
|
||||
};
|
||||
let Ok((w, h, rgba)) = dr_thumbs::codec::decode_rgba(&thumb.bytes) else {
|
||||
continue;
|
||||
};
|
||||
let rgb: Vec<f32> = rgba
|
||||
.chunks_exact(4)
|
||||
.flat_map(|p| {
|
||||
[
|
||||
p[0] as f32 / 255.0,
|
||||
p[1] as f32 / 255.0,
|
||||
p[2] as f32 / 255.0,
|
||||
]
|
||||
})
|
||||
.collect();
|
||||
|
||||
let Ok(dets) = det.detect(&rgb, w as usize, h as usize, &options) else {
|
||||
continue;
|
||||
};
|
||||
images += 1;
|
||||
for d in &dets {
|
||||
if let Some(a) = dr_face::warp(&rgb, w as usize, h as usize, &d.landmarks) {
|
||||
found.push((a.source_px(), a.sharpness()));
|
||||
}
|
||||
}
|
||||
if images.is_multiple_of(50) {
|
||||
println!(
|
||||
" {images}/{} images, {} face(s)",
|
||||
file_ids.len(),
|
||||
found.len()
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
if found.is_empty() {
|
||||
println!("no faces found in the sample.");
|
||||
return;
|
||||
}
|
||||
|
||||
let pct = |v: &mut Vec<f32>, p: f64| -> f32 {
|
||||
v.sort_by(|a, b| a.total_cmp(b));
|
||||
v[(((v.len() - 1) as f64) * p) as usize]
|
||||
};
|
||||
let mut sizes: Vec<f32> = found.iter().map(|f| f.0).collect();
|
||||
let mut sharps: Vec<f32> = found.iter().map(|f| f.1).collect();
|
||||
|
||||
println!("\n{} face(s) in {images} image(s)\n", found.len());
|
||||
println!(
|
||||
"{:>12} {:>8} {:>10}",
|
||||
"percentile", "size px", "sharpness"
|
||||
);
|
||||
println!("{}", "-".repeat(34));
|
||||
for p in [0.01, 0.05, 0.10, 0.25, 0.50, 0.75, 0.90, 0.99] {
|
||||
println!(
|
||||
"{:>11.0}% {:>8.0} {:>10.4}",
|
||||
p * 100.0,
|
||||
pct(&mut sizes, p),
|
||||
pct(&mut sharps, p)
|
||||
);
|
||||
}
|
||||
|
||||
// What each candidate pair would remove. Cumulative, because the gates are
|
||||
// applied together and their overlap is the whole question.
|
||||
println!(
|
||||
"\n{:>8} {:>10} {:>9} {:>9} {:>9}",
|
||||
"min crop", "min sharp", "size cut", "blur cut", "kept"
|
||||
);
|
||||
println!("{}", "-".repeat(52));
|
||||
for (min_px, min_sharp) in [
|
||||
(0.0_f32, 0.0_f32),
|
||||
(64.0, 0.0),
|
||||
(0.0, 0.010),
|
||||
(64.0, 0.005),
|
||||
(64.0, 0.010),
|
||||
(64.0, 0.020),
|
||||
(80.0, 0.010),
|
||||
(96.0, 0.010),
|
||||
] {
|
||||
let by_size = found.iter().filter(|f| f.0 < min_px).count();
|
||||
let by_blur = found
|
||||
.iter()
|
||||
.filter(|f| f.0 >= min_px && f.1 < min_sharp)
|
||||
.count();
|
||||
let kept = found.len() - by_size - by_blur;
|
||||
println!(
|
||||
"{min_px:>8.0} {min_sharp:>10.3} {:>8.0}% {:>8.0}% {:>8.0}%",
|
||||
100.0 * by_size as f64 / found.len() as f64,
|
||||
100.0 * by_blur as f64 / found.len() as f64,
|
||||
100.0 * kept as f64 / found.len() as f64,
|
||||
);
|
||||
}
|
||||
|
||||
println!(
|
||||
"\n`size cut` is what the size floor removes; `blur cut` is what the\n\
|
||||
sharpness floor removes *of what the size floor left*, so the two\n\
|
||||
columns do not double-count. A sharpness floor that cuts almost\n\
|
||||
nothing once the size floor is in place is a floor that is not\n\
|
||||
earning its place."
|
||||
);
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user