Turn face boxes back into sensor space before matching regions
Faces are found on the thumbnail, which is cached the right way up -- the grid would lie on its side otherwise. Segmentation runs on a proxy rendered through a neutral edit graph, which carries no orientation and is therefore in sensor order. For anything shot in portrait the two differ by a quarter turn, so a face and the person containing it were being compared in spaces 90 degrees apart: no match, or worse, a match against somebody else's region. The transform goes on the face rather than on the proxy. Instance masks are defined in the proxy's space and sampled long afterwards, so turning that space would be a far larger change than naming a region warrants. Also two things the first screenshot of the running app showed that no test would have: 110 of 23,528 displayed as "0%", which reads as the feature having done nothing. One decimal below ten percent, and a floor so real progress never shows as none. The rail picked some near-black covers, because the largest face in a group is often the nearest one in a badly lit frame and a black square beside a name identifies nobody. It now cuts the best few and takes the first legible one, falling back to the largest when a person's every photograph is dark -- which happens, and showing it beats showing nothing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -282,3 +282,160 @@ mod tests {
|
||||
assert_eq!(instances[0].class, "person");
|
||||
}
|
||||
}
|
||||
|
||||
/// Map a box from the **displayed** (upright) image back into the **stored**
|
||||
/// (sensor) one.
|
||||
///
|
||||
/// # Why this is needed at all
|
||||
///
|
||||
/// Faces are found on the thumbnail, which is cached the right way up — the
|
||||
/// grid would lie on its side otherwise. Segmentation runs on a proxy rendered
|
||||
/// through a *neutral* edit graph, which carries no orientation, so it is in
|
||||
/// sensor order. For any photograph shot in portrait the two spaces differ by a
|
||||
/// quarter turn, and matching a face against an instance without undoing that
|
||||
/// finds nothing — or worse, finds the wrong person, since a rotated box can
|
||||
/// still land inside some other instance.
|
||||
///
|
||||
/// The transform is applied to the face rather than to the segmentation proxy
|
||||
/// on purpose. Instance masks are defined in the proxy's space and sampled long
|
||||
/// afterwards; turning that space would be a far larger change than naming a
|
||||
/// region warrants.
|
||||
///
|
||||
/// `displayed` is the size of the upright image in the same units as `bbox`.
|
||||
/// Orientation is `(quarter_turns clockwise, flip_h, flip_v)`, applied by the
|
||||
/// renderer in that order — so undoing it means undoing the flips first.
|
||||
pub fn to_sensor_space(
|
||||
bbox: (f32, f32, f32, f32),
|
||||
displayed: (f32, f32),
|
||||
quarter_turns: u8,
|
||||
flip_h: bool,
|
||||
flip_v: bool,
|
||||
) -> (f32, f32, f32, f32) {
|
||||
let (dw, dh) = displayed;
|
||||
let (mut x0, mut y0, mut x1, mut y1) = bbox;
|
||||
|
||||
// Undo the mirrors, which the renderer applied last.
|
||||
if flip_h {
|
||||
let (a, b) = (dw - x1, dw - x0);
|
||||
x0 = a;
|
||||
x1 = b;
|
||||
}
|
||||
if flip_v {
|
||||
let (a, b) = (dh - y1, dh - y0);
|
||||
y0 = a;
|
||||
y1 = b;
|
||||
}
|
||||
|
||||
// Undo the turn. Each step rotates the box a quarter turn anticlockwise
|
||||
// within the frame it currently occupies, swapping the frame's extents as
|
||||
// it goes — which is why `w` and `h` are tracked rather than assumed.
|
||||
let (mut w, mut h) = (dw, dh);
|
||||
for _ in 0..(quarter_turns % 4) {
|
||||
// Clockwise forward is (x, y) -> (h_before - y, x); anticlockwise back
|
||||
// is (x, y) -> (y, w - x).
|
||||
let (nx0, ny0) = (y0, w - x1);
|
||||
let (nx1, ny1) = (y1, w - x0);
|
||||
x0 = nx0;
|
||||
y0 = ny0;
|
||||
x1 = nx1;
|
||||
y1 = ny1;
|
||||
std::mem::swap(&mut w, &mut h);
|
||||
}
|
||||
|
||||
(x0, y0, x1, y1)
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod orientation_tests {
|
||||
use super::*;
|
||||
|
||||
/// A landscape frame with a face near the top left.
|
||||
const DISPLAYED: (f32, f32) = (1000.0, 600.0);
|
||||
const FACE: (f32, f32, f32, f32) = (100.0, 50.0, 200.0, 150.0);
|
||||
|
||||
#[test]
|
||||
fn an_upright_image_needs_no_transform() {
|
||||
assert_eq!(to_sensor_space(FACE, DISPLAYED, 0, false, false), FACE);
|
||||
}
|
||||
|
||||
/// The case that motivated this: a portrait photograph, stored sideways
|
||||
/// and displayed with one clockwise quarter turn.
|
||||
#[test]
|
||||
fn a_quarter_turn_round_trips() {
|
||||
let sensor = to_sensor_space(FACE, DISPLAYED, 1, false, false);
|
||||
// Sensor frame is 600 × 1000 — the displayed extents swapped.
|
||||
assert!(sensor.0 >= 0.0 && sensor.2 <= 600.0, "{sensor:?}");
|
||||
assert!(sensor.1 >= 0.0 && sensor.3 <= 1000.0, "{sensor:?}");
|
||||
// And the box keeps its size, only turned.
|
||||
let (w, h) = (sensor.2 - sensor.0, sensor.3 - sensor.1);
|
||||
assert!((w - 100.0).abs() < 1e-3, "width {w}");
|
||||
assert!((h - 100.0).abs() < 1e-3, "height {h}");
|
||||
}
|
||||
|
||||
/// Four quarter turns is the identity, which is the cheapest possible
|
||||
/// check that the rotation step is self-consistent.
|
||||
#[test]
|
||||
fn four_quarter_turns_return_the_original() {
|
||||
let mut b = FACE;
|
||||
let mut frame = DISPLAYED;
|
||||
for _ in 0..4 {
|
||||
b = to_sensor_space(b, frame, 1, false, false);
|
||||
frame = (frame.1, frame.0);
|
||||
}
|
||||
assert!((b.0 - FACE.0).abs() < 1e-3, "{b:?}");
|
||||
assert!((b.1 - FACE.1).abs() < 1e-3, "{b:?}");
|
||||
assert!((b.2 - FACE.2).abs() < 1e-3, "{b:?}");
|
||||
assert!((b.3 - FACE.3).abs() < 1e-3, "{b:?}");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_horizontal_mirror_reflects_across_the_width() {
|
||||
let s = to_sensor_space(FACE, DISPLAYED, 0, true, false);
|
||||
assert_eq!(s, (800.0, 50.0, 900.0, 150.0));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_vertical_mirror_reflects_across_the_height() {
|
||||
let s = to_sensor_space(FACE, DISPLAYED, 0, false, true);
|
||||
assert_eq!(s, (100.0, 450.0, 200.0, 550.0));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_half_turn_maps_a_corner_to_the_opposite_corner() {
|
||||
let corner = (0.0, 0.0, 100.0, 100.0);
|
||||
let s = to_sensor_space(corner, DISPLAYED, 2, false, false);
|
||||
assert!((s.0 - 900.0).abs() < 1e-3, "{s:?}");
|
||||
assert!((s.1 - 500.0).abs() < 1e-3, "{s:?}");
|
||||
}
|
||||
|
||||
/// The boxes must stay well-formed whatever the transform: `x0 <= x1` and
|
||||
/// `y0 <= y1`, or every containment test downstream silently returns zero.
|
||||
#[test]
|
||||
fn every_orientation_produces_a_well_formed_box() {
|
||||
for turns in 0..4u8 {
|
||||
for &fh in &[false, true] {
|
||||
for &fv in &[false, true] {
|
||||
let s = to_sensor_space(FACE, DISPLAYED, turns, fh, fv);
|
||||
assert!(s.0 <= s.2, "turns={turns} fh={fh} fv={fv}: {s:?}");
|
||||
assert!(s.1 <= s.3, "turns={turns} fh={fh} fv={fv}: {s:?}");
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// A face that was inside the frame must stay inside it, whichever way the
|
||||
/// frame is turned.
|
||||
#[test]
|
||||
fn a_face_inside_the_frame_stays_inside_it() {
|
||||
for turns in 0..4u8 {
|
||||
let s = to_sensor_space(FACE, DISPLAYED, turns, false, false);
|
||||
let (fw, fh) = if turns % 2 == 1 {
|
||||
(DISPLAYED.1, DISPLAYED.0)
|
||||
} else {
|
||||
DISPLAYED
|
||||
};
|
||||
assert!(s.0 >= -1e-3 && s.2 <= fw + 1e-3, "turns={turns}: {s:?}");
|
||||
assert!(s.1 >= -1e-3 && s.3 <= fh + 1e-3, "turns={turns}: {s:?}");
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user