Index faces from the native render, not from a preview of it
Implements the FR-CULL-8 written two commits ago. The sweep fetched the JPEG preview embedded in each RAW and used that one buffer for both detection and the crop; it now fetches the original, renders it through the same path export uses, reduces that for the detector, and warps the crop back out of the native frame. Three pieces, and each exists for a reason worth stating. dr_face::Pixels lets the warp sample 8-bit RGBA directly. A 24 MP native frame is 96 MB as RGBA and 288 MB converted to the f32 RGB align.rs was written against, and the warp reads about forty thousand pixels out of it. Converting the whole frame to sample 0.2% of it is NFR-RES-2's budget spent on a copy, per image, for a whole library. The variant costs one branch per sample and a test asserts both layouts produce identical crops. The detector gets a box-filtered reduction to 1600px, not the native frame and not a point-sampled one. Averaging rather than sampling because the detector's job is finding small faces and decimation is precisely the operation that removes them: at 4x, fifteen of every sixteen pixels are discarded and a 40px face survives or not depending on where it falls relative to the sample grid. 1600 rather than 640 leaves the letterbox a mild 2.5x rather than a 9x, and bounds the f32 buffer at 20 MB. Landmarks come back in the reduction's coordinates and are scaled to native in one place before any crop pixel is read. This is the failure mode that would not announce itself -- unscaled landmarks put every crop near the top-left corner, which yields faces of something else, cleanly embedded and confidently clustered. The sweep fetches SWEEP_LANES-wide and renders sequentially. Not a placeholder for a parallel version: there is one GPU, so concurrent renders queue on it regardless, and each materialises a native frame. Overlapping them would multiply the one allocation that threatens the memory budget while buying parallelism that does not exist. The chunk drops from 96 to 6 for the same reason -- 96 held 8 MB previews, this holds whole RAWs. The stored edit is deliberately not applied, which is where this departs from export::render_from_library. Face geometry is normalised to the frame, so indexing a cropped render would record boxes against a frame that changes whenever the user changes their mind, and every stored box would quietly become wrong. Orientation is applied: that is a fact about the file rather than an edit. examples/face_native.rs renders one file and indexes it both ways, so the claim behind all of this can be checked against photographs rather than re-read out of the catalog it came from. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
+91
-10
@@ -285,7 +285,67 @@ pub fn warp(
|
||||
height: usize,
|
||||
landmarks: &[(f32, f32); 5],
|
||||
) -> Option<Aligned112> {
|
||||
if rgb.len() != width * height * 3 {
|
||||
warp_pixels(Pixels::RgbF32(rgb), width, height, landmarks)
|
||||
}
|
||||
|
||||
/// TRACES: FR-CULL-8
|
||||
/// What the warp may sample, in whichever layout the caller already holds.
|
||||
///
|
||||
/// # Why the 8-bit variant exists
|
||||
///
|
||||
/// FR-CULL-8 requires the crop to come from the **native** render, and a native
|
||||
/// render is large: a 24 MP frame is 96 MB as `RGBA8` and 288 MB converted to
|
||||
/// the `f32` RGB this module was originally written against. Converting the
|
||||
/// whole frame to sample 112×112 from it is three hundred megabytes allocated
|
||||
/// to read about forty thousand pixels, per image, on a pass that runs over a
|
||||
/// whole library — and on Android it is NFR-RES-2's budget spent outright.
|
||||
///
|
||||
/// So the warp reads whatever the caller has instead. It touches so few pixels
|
||||
/// that the per-sample conversion is free, and the buffer never has to be
|
||||
/// duplicated in another layout.
|
||||
#[derive(Debug, Clone, Copy)]
|
||||
pub enum Pixels<'a> {
|
||||
/// Tightly packed `f32` RGB in `0.0..=1.0`, row-major.
|
||||
RgbF32(&'a [f32]),
|
||||
/// Tightly packed 8-bit RGBA, row-major. Alpha is ignored: a face crop has
|
||||
/// no use for it and carrying it would change what the embedder receives.
|
||||
Rgba8(&'a [u8]),
|
||||
}
|
||||
|
||||
impl Pixels<'_> {
|
||||
/// Whether the buffer is the size `width × height` implies.
|
||||
fn fits(&self, width: usize, height: usize) -> bool {
|
||||
match self {
|
||||
Pixels::RgbF32(v) => v.len() == width * height * 3,
|
||||
Pixels::Rgba8(v) => v.len() == width * height * 4,
|
||||
}
|
||||
}
|
||||
|
||||
/// One channel of one pixel, as `0.0..=1.0`. Outside the buffer reads black.
|
||||
///
|
||||
/// Public because the face *crop* stored for the People screen is cut from
|
||||
/// the same buffer by the same caller, and it should not need a second
|
||||
/// copy of this to do it.
|
||||
pub fn channel(&self, w: usize, h: usize, x: isize, y: isize, c: usize) -> f32 {
|
||||
if x < 0 || y < 0 || x >= w as isize || y >= h as isize {
|
||||
return 0.0;
|
||||
}
|
||||
let i = y as usize * w + x as usize;
|
||||
match self {
|
||||
Pixels::RgbF32(v) => v[i * 3 + c],
|
||||
Pixels::Rgba8(v) => v[i * 4 + c] as f32 / 255.0,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// [`warp`], over any layout [`Pixels`] describes.
|
||||
pub fn warp_pixels(
|
||||
px: Pixels<'_>,
|
||||
width: usize,
|
||||
height: usize,
|
||||
landmarks: &[(f32, f32); 5],
|
||||
) -> Option<Aligned112> {
|
||||
if !px.fits(width, height) {
|
||||
return None;
|
||||
}
|
||||
let m = fit_similarity(landmarks, &ARCFACE_TEMPLATE)?;
|
||||
@@ -300,7 +360,7 @@ pub fn warp(
|
||||
let (x, y) = m.invert(u as f32 + 0.5, v as f32 + 0.5);
|
||||
let (x, y) = (x - 0.5, y - 0.5);
|
||||
let out = (v * e + u) * 3;
|
||||
sample_bilinear(rgb, width, height, x, y, &mut pixels[out..out + 3]);
|
||||
sample_bilinear(px, width, height, x, y, &mut pixels[out..out + 3]);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -312,7 +372,7 @@ pub fn warp(
|
||||
})
|
||||
}
|
||||
|
||||
fn sample_bilinear(rgb: &[f32], w: usize, h: usize, x: f32, y: f32, out: &mut [f32]) {
|
||||
fn sample_bilinear(px: Pixels<'_>, w: usize, h: usize, x: f32, y: f32, out: &mut [f32]) {
|
||||
let x0 = x.floor();
|
||||
let y0 = y.floor();
|
||||
let fx = x - x0;
|
||||
@@ -321,13 +381,7 @@ fn sample_bilinear(rgb: &[f32], w: usize, h: usize, x: f32, y: f32, out: &mut [f
|
||||
let y0 = y0 as isize;
|
||||
|
||||
for (c, o) in out.iter_mut().enumerate() {
|
||||
let get = |xi: isize, yi: isize| -> f32 {
|
||||
if xi < 0 || yi < 0 || xi >= w as isize || yi >= h as isize {
|
||||
0.0
|
||||
} else {
|
||||
rgb[(yi as usize * w + xi as usize) * 3 + c]
|
||||
}
|
||||
};
|
||||
let get = |xi: isize, yi: isize| -> f32 { px.channel(w, h, xi, yi, c) };
|
||||
let top = get(x0, y0) * (1.0 - fx) + get(x0 + 1, y0) * fx;
|
||||
let bot = get(x0, y0 + 1) * (1.0 - fx) + get(x0 + 1, y0 + 1) * fx;
|
||||
*o = top * (1.0 - fy) + bot * fy;
|
||||
@@ -508,6 +562,33 @@ mod tests {
|
||||
/// against a bright sky is low-contrast, and a raw Laplacian variance would
|
||||
/// reject it as blurred — which would quietly throw away every backlit
|
||||
/// portrait in the library.
|
||||
#[test]
|
||||
fn both_pixel_layouts_warp_to_the_same_crop() {
|
||||
// The 8-bit path exists so a native render need not be converted to
|
||||
// f32 whole; it has to agree with the path it replaces to within the
|
||||
// quantisation it introduces.
|
||||
let (w, h) = (64usize, 64usize);
|
||||
let mut rgba = vec![0u8; w * h * 4];
|
||||
let mut rgb = vec![0.0f32; w * h * 3];
|
||||
for y in 0..h {
|
||||
for x in 0..w {
|
||||
let v = [(x * 4 % 256) as u8, (y * 4 % 256) as u8, ((x + y) % 256) as u8];
|
||||
for c in 0..3 {
|
||||
rgba[(y * w + x) * 4 + c] = v[c];
|
||||
rgb[(y * w + x) * 3 + c] = v[c] as f32 / 255.0;
|
||||
}
|
||||
rgba[(y * w + x) * 4 + 3] = 255;
|
||||
}
|
||||
}
|
||||
let lm = shifted_scaled(0.35, 32.0, 32.0, 0.2);
|
||||
let a = warp_pixels(Pixels::RgbF32(&rgb), w, h, &lm).unwrap();
|
||||
let b = warp_pixels(Pixels::Rgba8(&rgba), w, h, &lm).unwrap();
|
||||
assert_eq!(a.source_px(), b.source_px());
|
||||
for (x, y) in a.pixels().iter().zip(b.pixels()) {
|
||||
assert!((x - y).abs() < 1e-6, "{x} vs {y}");
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn sharpness_survives_the_contrast_being_halved() {
|
||||
let edge = 200;
|
||||
|
||||
@@ -64,7 +64,7 @@ pub mod neighbours;
|
||||
/// from the wrong tier.
|
||||
pub const MIN_CROP_EDGE: u32 = 1025;
|
||||
|
||||
pub use align::{warp, Aligned112, Similarity, ALIGNED_EDGE, ARCFACE_TEMPLATE};
|
||||
pub use align::{warp, warp_pixels, Aligned112, Pixels, Similarity, ALIGNED_EDGE, ARCFACE_TEMPLATE};
|
||||
pub use assign::{identity_shares, RIVAL_FLOOR, TOP_MATCHES};
|
||||
pub use calibrate::{Calibration, Pairs, ReliabilityBand};
|
||||
pub use cluster::{
|
||||
|
||||
Reference in New Issue
Block a user