Turn face boxes back into sensor space before matching regions
Faces are found on the thumbnail, which is cached the right way up -- the grid would lie on its side otherwise. Segmentation runs on a proxy rendered through a neutral edit graph, which carries no orientation and is therefore in sensor order. For anything shot in portrait the two differ by a quarter turn, so a face and the person containing it were being compared in spaces 90 degrees apart: no match, or worse, a match against somebody else's region. The transform goes on the face rather than on the proxy. Instance masks are defined in the proxy's space and sampled long afterwards, so turning that space would be a far larger change than naming a region warrants. Also two things the first screenshot of the running app showed that no test would have: 110 of 23,528 displayed as "0%", which reads as the feature having done nothing. One decimal below ten percent, and a floor so real progress never shows as none. The rail picked some near-black covers, because the largest face in a group is often the nearest one in a badly lit frame and a black square beside a name identifies nobody. It now cuts the best few and takes the first legible one, falling back to the largest when a person's every photograph is dark -- which happens, and showing it beats showing nothing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
+91
-14
@@ -205,22 +205,70 @@ pub fn load_cover(
|
||||
store: &ThumbStore,
|
||||
person: PersonId,
|
||||
) -> Result<Option<FaceCrop>, dr_catalog::CatalogError> {
|
||||
let rows = faces::for_person(catalog.connection(), person, true)?;
|
||||
let Some(best) = rows
|
||||
.into_iter()
|
||||
.max_by(|a, b| {
|
||||
a.confirmed
|
||||
.cmp(&b.confirmed)
|
||||
.then(a.crop_px.total_cmp(&b.crop_px))
|
||||
})
|
||||
else {
|
||||
let mut rows = faces::for_person(catalog.connection(), person, true)?;
|
||||
if rows.is_empty() {
|
||||
return Ok(None);
|
||||
};
|
||||
}
|
||||
// Confirmed first, then largest.
|
||||
rows.sort_by(|a, b| {
|
||||
b.confirmed
|
||||
.cmp(&a.confirmed)
|
||||
.then(b.crop_px.total_cmp(&a.crop_px))
|
||||
});
|
||||
|
||||
let Some((w, h, rgba)) = decode_proxy(catalog, store, best.image_id) else {
|
||||
return Ok(None);
|
||||
};
|
||||
Ok(crop_face(&rgba, w, h, &best, COVER_CROP_EDGE))
|
||||
// Cut the best few and take the first legible one.
|
||||
//
|
||||
// Size alone picked some very dark crops on the reference library: the
|
||||
// largest face in a group is often the one nearest the camera in a badly
|
||||
// lit frame, and a black square beside a name identifies nobody. Bounded
|
||||
// at a handful because each candidate costs a JPEG decode, and the list is
|
||||
// already in preference order — so this gives up size only when the
|
||||
// preferred face is genuinely too dark to recognise.
|
||||
let mut fallback: Option<FaceCrop> = None;
|
||||
for face in rows.iter().take(COVER_CANDIDATES) {
|
||||
let Some((w, h, rgba)) = decode_proxy(catalog, store, face.image_id) else {
|
||||
continue;
|
||||
};
|
||||
let Some(crop) = crop_face(&rgba, w, h, face, COVER_CROP_EDGE) else {
|
||||
continue;
|
||||
};
|
||||
if mean_luma(&crop) >= MIN_COVER_LUMA {
|
||||
return Ok(Some(crop));
|
||||
}
|
||||
fallback.get_or_insert(crop);
|
||||
}
|
||||
// Everything this person has is dark. Their largest face is still the best
|
||||
// answer available, and showing it beats showing nothing.
|
||||
Ok(fallback)
|
||||
}
|
||||
|
||||
/// How many faces to cut before settling for the largest.
|
||||
const COVER_CANDIDATES: usize = 4;
|
||||
|
||||
/// Mean luma a cover must reach to be preferred over a larger, darker one.
|
||||
///
|
||||
/// Low: this rejects the near-black, not the moody. A crop at 0.18 is a
|
||||
/// legible face in a dim room; one at 0.05 is a silhouette.
|
||||
const MIN_COVER_LUMA: f32 = 0.18;
|
||||
|
||||
/// Rec. 709 luma, averaged over the crop, ignoring transparent margin.
|
||||
fn mean_luma(crop: &FaceCrop) -> f32 {
|
||||
let mut sum = 0.0_f32;
|
||||
let mut n = 0_u32;
|
||||
for px in crop.rgba.chunks_exact(4) {
|
||||
// Out-of-frame margin is transparent and black; counting it would make
|
||||
// every edge-of-frame face look darker than it is.
|
||||
if px[3] == 0 {
|
||||
continue;
|
||||
}
|
||||
sum += (0.2126 * px[0] as f32 + 0.7152 * px[1] as f32 + 0.0722 * px[2] as f32) / 255.0;
|
||||
n += 1;
|
||||
}
|
||||
if n == 0 {
|
||||
0.0
|
||||
} else {
|
||||
sum / n as f32
|
||||
}
|
||||
}
|
||||
|
||||
fn decode_proxy(
|
||||
@@ -754,6 +802,35 @@ mod tests {
|
||||
assert!(named_boxes_for_image(&c, img, 1024, 683).unwrap().is_empty());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn mean_luma_ignores_the_transparent_margin() {
|
||||
// Half opaque white, half transparent black. Counting the margin would
|
||||
// report 0.5; ignoring it reports 1.0, which is what the face is.
|
||||
let mut rgba = vec![0u8; 4 * 4 * 4];
|
||||
for (i, px) in rgba.chunks_exact_mut(4).enumerate() {
|
||||
if i < 8 {
|
||||
px.copy_from_slice(&[255, 255, 255, 255]);
|
||||
}
|
||||
}
|
||||
let crop = FaceCrop {
|
||||
width: 4,
|
||||
height: 4,
|
||||
rgba,
|
||||
};
|
||||
assert!((mean_luma(&crop) - 1.0).abs() < 1e-3, "{}", mean_luma(&crop));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_dark_crop_falls_below_the_cover_threshold_and_a_lit_one_clears_it() {
|
||||
let solid = |v: u8| FaceCrop {
|
||||
width: 2,
|
||||
height: 2,
|
||||
rgba: vec![v, v, v, 255, v, v, v, 255, v, v, v, 255, v, v, v, 255],
|
||||
};
|
||||
assert!(mean_luma(&solid(10)) < MIN_COVER_LUMA);
|
||||
assert!(mean_luma(&solid(120)) >= MIN_COVER_LUMA);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn deleting_everything_empties_the_screen() {
|
||||
let c = catalog();
|
||||
|
||||
Reference in New Issue
Block a user