Index faces from the native render, not from a preview of it
Implements the FR-CULL-8 written two commits ago. The sweep fetched the JPEG preview embedded in each RAW and used that one buffer for both detection and the crop; it now fetches the original, renders it through the same path export uses, reduces that for the detector, and warps the crop back out of the native frame. Three pieces, and each exists for a reason worth stating. dr_face::Pixels lets the warp sample 8-bit RGBA directly. A 24 MP native frame is 96 MB as RGBA and 288 MB converted to the f32 RGB align.rs was written against, and the warp reads about forty thousand pixels out of it. Converting the whole frame to sample 0.2% of it is NFR-RES-2's budget spent on a copy, per image, for a whole library. The variant costs one branch per sample and a test asserts both layouts produce identical crops. The detector gets a box-filtered reduction to 1600px, not the native frame and not a point-sampled one. Averaging rather than sampling because the detector's job is finding small faces and decimation is precisely the operation that removes them: at 4x, fifteen of every sixteen pixels are discarded and a 40px face survives or not depending on where it falls relative to the sample grid. 1600 rather than 640 leaves the letterbox a mild 2.5x rather than a 9x, and bounds the f32 buffer at 20 MB. Landmarks come back in the reduction's coordinates and are scaled to native in one place before any crop pixel is read. This is the failure mode that would not announce itself -- unscaled landmarks put every crop near the top-left corner, which yields faces of something else, cleanly embedded and confidently clustered. The sweep fetches SWEEP_LANES-wide and renders sequentially. Not a placeholder for a parallel version: there is one GPU, so concurrent renders queue on it regardless, and each materialises a native frame. Overlapping them would multiply the one allocation that threatens the memory budget while buying parallelism that does not exist. The chunk drops from 96 to 6 for the same reason -- 96 held 8 MB previews, this holds whole RAWs. The stored edit is deliberately not applied, which is where this departs from export::render_from_library. Face geometry is normalised to the frame, so indexing a cropped render would record boxes against a frame that changes whenever the user changes their mind, and every stored box would quietly become wrong. Orientation is applied: that is a fact about the file rather than an edit. examples/face_native.rs renders one file and indexes it both ways, so the claim behind all of this can be checked against photographs rather than re-read out of the catalog it came from. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -327,6 +327,253 @@ pub fn index_proxy(
|
||||
Ok(out)
|
||||
}
|
||||
|
||||
/// TRACES: FR-CULL-8
|
||||
/// Long edge the detector's input is reduced to, in pixels.
|
||||
///
|
||||
/// Detection is indifferent above this — `dr_face::detect` letterboxes into a
|
||||
/// fixed 640×640 whatever it receives, so a face is the same size to the model
|
||||
/// from a 1600px buffer as from a 6000px one (faces.md §4.1). What the
|
||||
/// reduction buys is the aliasing that a single bilinear step from native to
|
||||
/// 640 would introduce: a 9× decimation samples one pixel in nine and drops
|
||||
/// small faces into the gaps between samples. Box-filtering to 1600 first, and
|
||||
/// letting the letterbox take the remaining 2.5×, keeps every source pixel in
|
||||
/// the average.
|
||||
///
|
||||
/// It also bounds the buffer that has to exist as `f32`: 1600×1067 is 20 MB
|
||||
/// against 288 MB for a native 24 MP frame.
|
||||
const DETECT_EDGE: usize = 1600;
|
||||
|
||||
/// TRACES: FR-CULL-8 | NFR-RES-2
|
||||
/// Detect and embed every face in one **native-resolution** render.
|
||||
///
|
||||
/// The pass FR-CULL-8 specifies, and the two resolutions in it are the whole
|
||||
/// point:
|
||||
///
|
||||
/// - the **detector** is handed a box-filtered reduction, because it discards
|
||||
/// anything above its own 640px input anyway;
|
||||
/// - the **crop** is warped out of the native buffer, because that is the one
|
||||
/// place source resolution becomes embedding quality.
|
||||
///
|
||||
/// Boxes and landmarks come back in the reduction's coordinates and are scaled
|
||||
/// to native before a single crop pixel is read. Getting that scaling wrong
|
||||
/// does not fail loudly — it yields faces in plausible-looking places with
|
||||
/// crops taken from beside them — so it is one multiply applied in one place
|
||||
/// rather than at each use.
|
||||
///
|
||||
/// `rgba` is the native render, tightly packed 8-bit RGBA. It is never
|
||||
/// converted wholesale: [`dr_face::Pixels`] samples it where the warp and the
|
||||
/// stored crop actually touch it.
|
||||
pub fn index_native(
|
||||
detector: &mut Detector,
|
||||
embedder: &mut Embedder,
|
||||
rgba: &[u8],
|
||||
width: usize,
|
||||
height: usize,
|
||||
options: &DetectOptions,
|
||||
) -> Result<Vec<DetectedFace>, dr_face::FaceError> {
|
||||
if !index_native_shape_ok(rgba, width, height) {
|
||||
return Ok(Vec::new());
|
||||
}
|
||||
let long_edge = width.max(height) as f32;
|
||||
let native = dr_face::Pixels::Rgba8(rgba);
|
||||
|
||||
// One reduction, reused for every face on the image.
|
||||
let (dw, dh, small) = reduce_for_detection(rgba, width, height);
|
||||
if dw == 0 || dh == 0 {
|
||||
return Ok(Vec::new());
|
||||
}
|
||||
let dets = detector.detect(&small, dw, dh, options)?;
|
||||
|
||||
// Detector coordinates to native. Separate factors rather than one, because
|
||||
// the reduction rounds each axis independently and assuming they match puts
|
||||
// every landmark a fraction of a face off on the shorter one.
|
||||
let (sx, sy) = (width as f32 / dw as f32, height as f32 / dh as f32);
|
||||
|
||||
let mut out = Vec::with_capacity(dets.len());
|
||||
for d in &dets {
|
||||
let landmarks = scale_landmarks(&d.landmarks, sx, sy);
|
||||
let Some(aligned) = dr_face::warp_pixels(native, width, height, &landmarks) else {
|
||||
log::debug!("face with degenerate landmarks skipped");
|
||||
continue;
|
||||
};
|
||||
|
||||
// Both gates read the aligned crop, so they mean what they say only now
|
||||
// that the crop comes from native pixels — `source_px` is the real
|
||||
// count, not the proxy's idea of it. See `index_proxy` for why each
|
||||
// one is here rather than in the detector.
|
||||
if aligned.source_px() < options.min_source_px {
|
||||
log::debug!(
|
||||
"face skipped: {:.0} source px below {:.0}",
|
||||
aligned.source_px(),
|
||||
options.min_source_px
|
||||
);
|
||||
continue;
|
||||
}
|
||||
let sharpness = aligned.sharpness();
|
||||
if sharpness < options.min_sharpness {
|
||||
log::debug!(
|
||||
"face at {:.0}px skipped: sharpness {sharpness:.4} below {:.4}",
|
||||
aligned.source_px(),
|
||||
options.min_sharpness
|
||||
);
|
||||
continue;
|
||||
}
|
||||
|
||||
let embedding = embedder.embed(&aligned)?;
|
||||
let (bx, by) = (d.bbox.0 * sx, d.bbox.1 * sy);
|
||||
let (bw, bh) = (d.width() * sx, d.height() * sy);
|
||||
|
||||
out.push(DetectedFace {
|
||||
x: bx / long_edge,
|
||||
y: by / long_edge,
|
||||
w: bw / long_edge,
|
||||
h: bh / long_edge,
|
||||
landmarks: normalise_landmarks(&landmarks, long_edge),
|
||||
confidence: d.confidence,
|
||||
embedding: embedding.to_f16_bytes(),
|
||||
crop_px: aligned.source_px(),
|
||||
model_id: embedder.model().as_str().to_string(),
|
||||
crop: cut_crop_native(native, width, height, (bx, by, bw, bh)).unwrap_or_default(),
|
||||
});
|
||||
}
|
||||
|
||||
Ok(out)
|
||||
}
|
||||
|
||||
/// Whether a buffer is the native frame it claims to be.
|
||||
///
|
||||
/// Checked before anything expensive, and before the detector above all: a
|
||||
/// mismatched buffer would otherwise be read past its end by the reduction.
|
||||
fn index_native_shape_ok(rgba: &[u8], width: usize, height: usize) -> bool {
|
||||
width > 0 && height > 0 && rgba.len() == width * height * 4
|
||||
}
|
||||
|
||||
/// Landmarks from the detector's reduction into native coordinates.
|
||||
fn scale_landmarks(lm: &[(f32, f32); 5], sx: f32, sy: f32) -> [(f32, f32); 5] {
|
||||
let mut out = [(0.0_f32, 0.0_f32); 5];
|
||||
for (o, &(x, y)) in out.iter_mut().zip(lm.iter()) {
|
||||
*o = (x * sx, y * sy);
|
||||
}
|
||||
out
|
||||
}
|
||||
|
||||
/// Box-filter a native RGBA render down to [`DETECT_EDGE`], as packed `f32` RGB.
|
||||
///
|
||||
/// Averaging rather than sampling. The detector's job is to find small faces,
|
||||
/// and a point-sampled reduction is exactly the operation that removes them:
|
||||
/// at a 4× decimation fifteen of every sixteen pixels are discarded, and a
|
||||
/// 40px face survives or not depending on where it happens to sit relative to
|
||||
/// the sample grid.
|
||||
fn reduce_for_detection(rgba: &[u8], width: usize, height: usize) -> (usize, usize, Vec<f32>) {
|
||||
reduce_to(rgba, width, height, DETECT_EDGE)
|
||||
}
|
||||
|
||||
/// [`reduce_for_detection`], to a stated long edge.
|
||||
///
|
||||
/// Split out so the averaging can be tested at a size a test can reason about
|
||||
/// by hand, rather than by building a 1600px fixture.
|
||||
fn reduce_to(
|
||||
rgba: &[u8],
|
||||
width: usize,
|
||||
height: usize,
|
||||
target: usize,
|
||||
) -> (usize, usize, Vec<f32>) {
|
||||
let long = width.max(height);
|
||||
if long <= target {
|
||||
// Already small enough; convert without resampling rather than round
|
||||
// -tripping through a scale of 1.
|
||||
let mut out = vec![0.0_f32; width * height * 3];
|
||||
for (i, px) in rgba.chunks_exact(4).enumerate() {
|
||||
for c in 0..3 {
|
||||
out[i * 3 + c] = px[c] as f32 / 255.0;
|
||||
}
|
||||
}
|
||||
return (width, height, out);
|
||||
}
|
||||
|
||||
let scale = target as f32 / long as f32;
|
||||
let (dw, dh) = (
|
||||
((width as f32 * scale).round() as usize).max(1),
|
||||
((height as f32 * scale).round() as usize).max(1),
|
||||
);
|
||||
let mut out = vec![0.0_f32; dw * dh * 3];
|
||||
for oy in 0..dh {
|
||||
// Source span of this destination row, as a half-open range, so
|
||||
// adjacent rows tile the source exactly and no row is counted twice.
|
||||
let y0 = oy * height / dh;
|
||||
let y1 = (((oy + 1) * height) / dh).max(y0 + 1).min(height);
|
||||
for ox in 0..dw {
|
||||
let x0 = ox * width / dw;
|
||||
let x1 = (((ox + 1) * width) / dw).max(x0 + 1).min(width);
|
||||
|
||||
let mut acc = [0.0_f32; 3];
|
||||
let mut n = 0.0_f32;
|
||||
for y in y0..y1 {
|
||||
for x in x0..x1 {
|
||||
let i = (y * width + x) * 4;
|
||||
for (c, a) in acc.iter_mut().enumerate() {
|
||||
*a += rgba[i + c] as f32;
|
||||
}
|
||||
n += 1.0;
|
||||
}
|
||||
}
|
||||
let o = (oy * dw + ox) * 3;
|
||||
for c in 0..3 {
|
||||
out[o + c] = acc[c] / n / 255.0;
|
||||
}
|
||||
}
|
||||
}
|
||||
(dw, dh, out)
|
||||
}
|
||||
|
||||
/// The stored face thumbnail, cut from the native buffer.
|
||||
///
|
||||
/// Same framing as [`cut_crop`] and deliberately a separate function rather
|
||||
/// than a generalisation of it: this one takes a box already in native
|
||||
/// coordinates, and blurring that distinction is how a crop ends up sampled
|
||||
/// from the wrong scale.
|
||||
fn cut_crop_native(
|
||||
px: dr_face::Pixels<'_>,
|
||||
width: usize,
|
||||
height: usize,
|
||||
bbox: (f32, f32, f32, f32),
|
||||
) -> Option<Vec<u8>> {
|
||||
if width == 0 || height == 0 {
|
||||
return None;
|
||||
}
|
||||
let (bx, by, bw, bh) = bbox;
|
||||
let cx = bx + bw * 0.5;
|
||||
let cy = by + bh * 0.5;
|
||||
let half = bw.max(bh) * 0.5 * (1.0 + CROP_MARGIN);
|
||||
if !(half.is_finite() && half > 0.5 && cx.is_finite() && cy.is_finite()) {
|
||||
return None;
|
||||
}
|
||||
|
||||
let edge = STORED_CROP_EDGE;
|
||||
let mut out = vec![0u8; (edge * edge * 4) as usize];
|
||||
let step = (half * 2.0) / edge as f32;
|
||||
for oy in 0..edge {
|
||||
let sy = cy - half + (oy as f32 + 0.5) * step;
|
||||
for ox in 0..edge {
|
||||
let sx = cx - half + (ox as f32 + 0.5) * step;
|
||||
let o = ((oy * edge + ox) * 4) as usize;
|
||||
for c in 0..3 {
|
||||
let v = px.channel(width, height, sx as isize, sy as isize, c);
|
||||
out[o + c] = (v.clamp(0.0, 1.0) * 255.0).round() as u8;
|
||||
}
|
||||
out[o + 3] = 255;
|
||||
}
|
||||
}
|
||||
|
||||
match dr_thumbs::encode_rgba(edge, edge, &out) {
|
||||
Ok(bytes) => Some(bytes),
|
||||
Err(e) => {
|
||||
log::debug!("encoding a face crop: {e}");
|
||||
None
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
fn normalise_landmarks(lm: &[(f32, f32); 5], long_edge: f32) -> [(f32, f32); 5] {
|
||||
let mut out = [(0.0_f32, 0.0_f32); 5];
|
||||
for (o, &(x, y)) in out.iter_mut().zip(lm.iter()) {
|
||||
@@ -1084,6 +1331,69 @@ fn rgba_to_rgb_f32(rgba: &[u8]) -> Vec<f32> {
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
#[test]
|
||||
fn the_detector_reduction_averages_rather_than_samples() {
|
||||
// Four source pixels per destination pixel, with one bright pixel in
|
||||
// each group. A point sampler returns either 255 or 0 depending on
|
||||
// which corner it lands on; the average is the same every time, and
|
||||
// that stability is what stops a small face vanishing on a grid
|
||||
// alignment it has no control over.
|
||||
let (w, h) = (4usize, 4usize);
|
||||
let mut rgba = vec![0u8; w * h * 4];
|
||||
for y in 0..h {
|
||||
for x in 0..w {
|
||||
let i = (y * w + x) * 4;
|
||||
let bright = x % 2 == 0 && y % 2 == 0;
|
||||
for c in 0..3 {
|
||||
rgba[i + c] = if bright { 255 } else { 0 };
|
||||
}
|
||||
rgba[i + 3] = 255;
|
||||
}
|
||||
}
|
||||
// Force a 2x reduction regardless of DETECT_EDGE by reducing by hand
|
||||
// through the same helper on a source larger than the target.
|
||||
let (dw, dh, out) = reduce_to(&rgba, w, h, 2);
|
||||
assert_eq!((dw, dh), (2, 2));
|
||||
for v in out.chunks_exact(3) {
|
||||
assert!(
|
||||
(v[0] - 0.25).abs() < 1e-6,
|
||||
"each destination pixel averages one bright of four, got {}",
|
||||
v[0]
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn the_reduction_keeps_the_aspect_ratio_it_was_given() {
|
||||
let (w, h) = (400usize, 100usize);
|
||||
let rgba = vec![128u8; w * h * 4];
|
||||
let (dw, dh, out) = reduce_to(&rgba, w, h, 100);
|
||||
assert_eq!(dw, 100);
|
||||
assert_eq!(dh, 25);
|
||||
assert_eq!(out.len(), dw * dh * 3);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn landmarks_scale_back_to_the_native_frame() {
|
||||
// The failure this guards is silent: landmarks left in the detector's
|
||||
// coordinates put every crop near the top-left corner of the frame,
|
||||
// which produces faces that look like faces of something else.
|
||||
let lm = [(10.0, 20.0), (30.0, 40.0), (50.0, 60.0), (70.0, 80.0), (90.0, 100.0)];
|
||||
let out = scale_landmarks(&lm, 4.0, 2.0);
|
||||
assert_eq!(out[0], (40.0, 40.0));
|
||||
assert_eq!(out[4], (360.0, 200.0));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_buffer_that_is_not_the_stated_size_indexes_nothing() {
|
||||
// No model is loaded here, so reaching the detector would panic. The
|
||||
// point is that it does not: the shape check comes first.
|
||||
let rgba = vec![0u8; 10];
|
||||
assert!(!index_native_shape_ok(&rgba, 100, 100));
|
||||
assert!(!index_native_shape_ok(&rgba, 0, 0));
|
||||
assert!(index_native_shape_ok(&vec![0u8; 4 * 100 * 100], 100, 100));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn an_audit_summary_names_both_kinds_of_outstanding() {
|
||||
let a = IndexAudit {
|
||||
|
||||
Reference in New Issue
Block a user