Stop indexing faces too small or too blurred to be anyone
The library was storing faces at 52 source pixels and embedding whatever came
back. There was a size floor, but it was 40 pixels on the *bounding box*, and
there was no blur gate at all — so a subject walking through a half-second
exposure detected confidently, aligned cleanly, and produced a perfectly
ordinary-looking 512-vector. Nothing downstream can tell that apart from a real
face, and because blurs resemble each other more than they resemble the people
they were, they cluster together and weld unrelated identities into one group.
Two floors, both measured rather than guessed. `face_index --quality` runs the
detector over real proxies with both gates disabled and prints the distribution;
over 1,503 faces in 600 images of the reference library:
percentile crop px sharpness
1% 16 0.0006
25% 23 0.0025
50% 38 0.0071
75% 76 0.0284
99% 352 0.4282
The median face in a personal library is 38 pixels. Most of what the detector
finds is background: people across a square, a face on a poster, a stranger at
the next table. They are real detections and useless identifications.
**Size, on the crop rather than the box.** "At least 64x64" has to mean the
pixels the *embedder* sees, and the box is not that — the ArcFace template
reaches past it for forehead and chin, so the aligned crop spans roughly 1.3x
the box's shorter edge. The floor is therefore `min_source_px` on the aligned
crop, applied after the warp fixes the scale, and `min_face_px` drops to 48 as
what it always really was: a cheap pre-filter set low enough that it cannot
reject a face the real floor would have kept.
**Sharpness.** Variance of the Laplacian divided by the variance of the luma it
was taken over. The division is the part that matters: raw Laplacian variance
scales with contrast, so a threshold on it would quietly discard every backlit
portrait in the library. The ratio asks how much of the crop's variation is
edges rather than broad gradients, and is invariant to exposure.
What each pair removes, cumulatively, of everything the detector finds:
min crop min sharp size cut blur cut kept
64 0.000 70% 0% 30%
64 0.010 70% 3% 27%
64 0.020 70% 7% 23%
80 0.010 76% 2% 21%
64 and 0.020. The size floor does most of the work, and the blur floor removing
only 7% on top of it is the point rather than a disappointment: at 64 pixels
most faces are already sharp, and what it takes out is the large-but-soft one —
precisely the face that would otherwise contribute a confident, wrong embedding.
The two gates are not independent and the doc comments say so: a face under 112
pixels was upsampled to reach the embedder, and upsampling invents no edges, so
small faces score low on sharpness even when the original was crisp. That is why
`--quality` prints them together.
**This will re-index.** Around 70% of what the current settings store falls below
the new floors — faces between 20 and 40 pixels that nobody could identify. The
People screen gets shorter and every group in it gets better.
66 dr-face tests pass, including that a blurred crop scores below a sharp one,
that halving the contrast does not move the score, and that an upsampled face
scores below the same face at full size.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -74,6 +74,92 @@ impl Aligned112 {
|
||||
pub fn source_px(&self) -> f32 {
|
||||
self.source_px
|
||||
}
|
||||
|
||||
/// How sharp the face the embedder is about to see actually is.
|
||||
///
|
||||
/// # Why size is not enough
|
||||
///
|
||||
/// A face can be large and useless. A subject walking through a half-second
|
||||
/// exposure, a frame focused on the person behind them, a hand-held shot at
|
||||
/// 1/15 — all yield a big box, a confident detection and five landmarks in
|
||||
/// plausible places. The embedding that comes back is not *wrong* in any
|
||||
/// way the system can see: it is unit-norm and its cosines look ordinary.
|
||||
/// It is simply an embedding of a blur, and blurs resemble each other more
|
||||
/// than they resemble the people they were, so they cluster together and
|
||||
/// bridge identities that have nothing to do with one another.
|
||||
///
|
||||
/// That is the failure this exists to prevent, and it is the same class of
|
||||
/// fault as the unaligned-crop one the [`Aligned112`] newtype guards
|
||||
/// against: plausible output, no error, worse results, nothing reported.
|
||||
///
|
||||
/// # The measure
|
||||
///
|
||||
/// Variance of the Laplacian — the standard blur metric — **divided by the
|
||||
/// variance of the luma it was taken over**. The division is what makes it
|
||||
/// usable here. Raw Laplacian variance scales with contrast, so a sharp
|
||||
/// face in flat, hazy or backlit light scores like a blurred one in hard
|
||||
/// light, and a threshold on it would quietly throw away every face shot
|
||||
/// against a bright sky. The ratio asks the question that actually matters
|
||||
/// — *how much of this crop's variation is edges rather than broad
|
||||
/// gradients* — and is invariant to exposure and contrast.
|
||||
///
|
||||
/// Computed on luma over the interior, so the 3x3 kernel never needs a
|
||||
/// border rule. Returns 0.0 for a crop with no variation at all, which is
|
||||
/// a flat patch and correctly unusable rather than infinitely sharp.
|
||||
///
|
||||
/// # This is not independent of size
|
||||
///
|
||||
/// A face smaller than 112 pixels was *upsampled* to reach the embedder,
|
||||
/// and upsampling invents no edges — so a small face scores low here even
|
||||
/// when the original was perfectly sharp. That is not a flaw to correct: it
|
||||
/// is the honest statement that the embedder is looking at a soft image.
|
||||
/// The size floor and this one overlap deliberately, and
|
||||
/// `face_index --quality` prints the joint distribution so the two are
|
||||
/// chosen together rather than each in ignorance of the other.
|
||||
pub fn sharpness(&self) -> f32 {
|
||||
let e = ALIGNED_EDGE;
|
||||
let luma: Vec<f32> = self
|
||||
.pixels
|
||||
.chunks_exact(3)
|
||||
.map(|p| 0.2126 * p[0] + 0.7152 * p[1] + 0.0722 * p[2])
|
||||
.collect();
|
||||
|
||||
let (mut lap_sum, mut lap_sq) = (0.0_f64, 0.0_f64);
|
||||
let (mut lum_sum, mut lum_sq) = (0.0_f64, 0.0_f64);
|
||||
let mut n = 0.0_f64;
|
||||
|
||||
for y in 1..e - 1 {
|
||||
for x in 1..e - 1 {
|
||||
let i = y * e + x;
|
||||
// Four-neighbour Laplacian. The 8-neighbour form is more
|
||||
// sensitive to diagonal detail and also to noise, which on a
|
||||
// high-ISO frame is exactly the thing that must not read as
|
||||
// sharpness.
|
||||
let lap = 4.0 * luma[i] - luma[i - 1] - luma[i + 1] - luma[i - e] - luma[i + e];
|
||||
let lap = lap as f64;
|
||||
lap_sum += lap;
|
||||
lap_sq += lap * lap;
|
||||
|
||||
let l = luma[i] as f64;
|
||||
lum_sum += l;
|
||||
lum_sq += l * l;
|
||||
n += 1.0;
|
||||
}
|
||||
}
|
||||
|
||||
if n == 0.0 {
|
||||
return 0.0;
|
||||
}
|
||||
let lap_var = (lap_sq / n - (lap_sum / n).powi(2)).max(0.0);
|
||||
let lum_var = (lum_sq / n - (lum_sum / n).powi(2)).max(0.0);
|
||||
|
||||
// A crop with no luma variation has no edges to find either, so the
|
||||
// ratio is 0/0. Zero is the right answer: nothing there is a face.
|
||||
if lum_var <= 1e-9 {
|
||||
return 0.0;
|
||||
}
|
||||
(lap_var / lum_var) as f32
|
||||
}
|
||||
}
|
||||
|
||||
/// A similarity transform: rotation, uniform scale, translation.
|
||||
@@ -357,4 +443,132 @@ mod tests {
|
||||
let a = warp(&rgb, 32, 32, &lm).unwrap();
|
||||
assert!(a.pixels().iter().all(|&v| v == 0.0));
|
||||
}
|
||||
|
||||
// ── sharpness ─────────────────────────────────────────────────────────
|
||||
|
||||
/// An image of `edge` square, filled by `f(x, y) -> luma`.
|
||||
fn image(edge: usize, f: impl Fn(usize, usize) -> f32) -> Vec<f32> {
|
||||
let mut v = Vec::with_capacity(edge * edge * 3);
|
||||
for y in 0..edge {
|
||||
for x in 0..edge {
|
||||
let l = f(x, y);
|
||||
v.extend_from_slice(&[l, l, l]);
|
||||
}
|
||||
}
|
||||
v
|
||||
}
|
||||
|
||||
/// One box-blur pass, which is enough to move the metric a long way.
|
||||
fn blur(rgb: &[f32], edge: usize) -> Vec<f32> {
|
||||
let mut out = rgb.to_vec();
|
||||
for y in 1..edge - 1 {
|
||||
for x in 1..edge - 1 {
|
||||
for c in 0..3 {
|
||||
let mut sum = 0.0;
|
||||
for dy in -1isize..=1 {
|
||||
for dx in -1isize..=1 {
|
||||
let i = (((y as isize + dy) as usize) * edge
|
||||
+ ((x as isize + dx) as usize))
|
||||
* 3
|
||||
+ c;
|
||||
sum += rgb[i];
|
||||
}
|
||||
}
|
||||
out[(y * edge + x) * 3 + c] = sum / 9.0;
|
||||
}
|
||||
}
|
||||
}
|
||||
out
|
||||
}
|
||||
|
||||
/// Landmarks placing the template into a larger image at scale 1, so the
|
||||
/// warp resamples one-to-one and the metric sees the source detail.
|
||||
fn centred(edge: usize) -> [(f32, f32); 5] {
|
||||
let off = (edge as f32 - ALIGNED_EDGE as f32) / 2.0;
|
||||
shifted_scaled(1.0, off, off, 0.0)
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_blurred_face_scores_lower_than_a_sharp_one() {
|
||||
let edge = 200;
|
||||
let sharp = image(
|
||||
edge,
|
||||
|x, y| if (x / 3 + y / 3) % 2 == 0 { 0.9 } else { 0.1 },
|
||||
);
|
||||
let soft = blur(&blur(&sharp, edge), edge);
|
||||
|
||||
let a = warp(&sharp, edge, edge, ¢red(edge))
|
||||
.unwrap()
|
||||
.sharpness();
|
||||
let b = warp(&soft, edge, edge, ¢red(edge)).unwrap().sharpness();
|
||||
assert!(a > b * 2.0, "sharp {a} should clearly beat blurred {b}");
|
||||
}
|
||||
|
||||
/// The reason for dividing by luma variance. A sharp face photographed
|
||||
/// against a bright sky is low-contrast, and a raw Laplacian variance would
|
||||
/// reject it as blurred — which would quietly throw away every backlit
|
||||
/// portrait in the library.
|
||||
#[test]
|
||||
fn sharpness_survives_the_contrast_being_halved() {
|
||||
let edge = 200;
|
||||
let full = image(
|
||||
edge,
|
||||
|x, y| if (x / 3 + y / 3) % 2 == 0 { 0.9 } else { 0.1 },
|
||||
);
|
||||
// Same detail, half the contrast, lifted so it does not clip.
|
||||
let flat = image(
|
||||
edge,
|
||||
|x, y| {
|
||||
if (x / 3 + y / 3) % 2 == 0 {
|
||||
0.55
|
||||
} else {
|
||||
0.45
|
||||
}
|
||||
},
|
||||
);
|
||||
|
||||
let a = warp(&full, edge, edge, ¢red(edge)).unwrap().sharpness();
|
||||
let b = warp(&flat, edge, edge, ¢red(edge)).unwrap().sharpness();
|
||||
let ratio = a / b;
|
||||
assert!(
|
||||
(0.5..2.0).contains(&ratio),
|
||||
"contrast changed the score {ratio}x ({a} vs {b})"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_flat_crop_has_no_sharpness() {
|
||||
let edge = 200;
|
||||
let flat = image(edge, |_, _| 0.5);
|
||||
assert_eq!(
|
||||
warp(&flat, edge, edge, ¢red(edge)).unwrap().sharpness(),
|
||||
0.0
|
||||
);
|
||||
}
|
||||
|
||||
/// Upsampling invents no detail, so a face that had to be stretched to
|
||||
/// reach the embedder scores lower than the same face at full size. That
|
||||
/// overlap with the size floor is deliberate and documented; this pins it
|
||||
/// so a future change cannot quietly remove it.
|
||||
#[test]
|
||||
fn an_upsampled_face_scores_lower_than_the_same_face_at_full_size() {
|
||||
let edge = 200;
|
||||
let src = image(
|
||||
edge,
|
||||
|x, y| if (x / 3 + y / 3) % 2 == 0 { 0.9 } else { 0.1 },
|
||||
);
|
||||
|
||||
let full = warp(&src, edge, edge, ¢red(edge)).unwrap();
|
||||
// Half scale: the crop spans 56 source pixels and is stretched to 112.
|
||||
let off = (edge as f32 - ALIGNED_EDGE as f32 / 2.0) / 2.0;
|
||||
let small = warp(&src, edge, edge, &shifted_scaled(0.5, off, off, 0.0)).unwrap();
|
||||
|
||||
assert!(small.source_px() < full.source_px());
|
||||
assert!(
|
||||
small.sharpness() < full.sharpness(),
|
||||
"upsampled {} should be softer than full {}",
|
||||
small.sharpness(),
|
||||
full.sharpness()
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user