Index faces from the native render, not from a preview of it

Implements the FR-CULL-8 written two commits ago. The sweep fetched the
JPEG preview embedded in each RAW and used that one buffer for both
detection and the crop; it now fetches the original, renders it through
the same path export uses, reduces that for the detector, and warps the
crop back out of the native frame.

Three pieces, and each exists for a reason worth stating.

dr_face::Pixels lets the warp sample 8-bit RGBA directly. A 24 MP native
frame is 96 MB as RGBA and 288 MB converted to the f32 RGB align.rs was
written against, and the warp reads about forty thousand pixels out of
it. Converting the whole frame to sample 0.2% of it is NFR-RES-2's
budget spent on a copy, per image, for a whole library. The variant
costs one branch per sample and a test asserts both layouts produce
identical crops.

The detector gets a box-filtered reduction to 1600px, not the native
frame and not a point-sampled one. Averaging rather than sampling
because the detector's job is finding small faces and decimation is
precisely the operation that removes them: at 4x, fifteen of every
sixteen pixels are discarded and a 40px face survives or not depending
on where it falls relative to the sample grid. 1600 rather than 640
leaves the letterbox a mild 2.5x rather than a 9x, and bounds the f32
buffer at 20 MB.

Landmarks come back in the reduction's coordinates and are scaled to
native in one place before any crop pixel is read. This is the failure
mode that would not announce itself -- unscaled landmarks put every crop
near the top-left corner, which yields faces of something else, cleanly
embedded and confidently clustered.

The sweep fetches SWEEP_LANES-wide and renders sequentially. Not a
placeholder for a parallel version: there is one GPU, so concurrent
renders queue on it regardless, and each materialises a native frame.
Overlapping them would multiply the one allocation that threatens the
memory budget while buying parallelism that does not exist. The chunk
drops from 96 to 6 for the same reason -- 96 held 8 MB previews, this
holds whole RAWs.

The stored edit is deliberately not applied, which is where this departs
from export::render_from_library. Face geometry is normalised to the
frame, so indexing a cropped render would record boxes against a frame
that changes whenever the user changes their mind, and every stored box
would quietly become wrong. Orientation is applied: that is a fact about
the file rather than an edit.

examples/face_native.rs renders one file and indexes it both ways, so
the claim behind all of this can be checked against photographs rather
than re-read out of the catalog it came from.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-30 19:41:30 +02:00
co-authored by Claude Opus 5
parent 9ddc1273c0
commit 4af3b93dfa
8 changed files with 898 additions and 270 deletions
+91 -10
View File
@@ -285,7 +285,67 @@ pub fn warp(
height: usize, height: usize,
landmarks: &[(f32, f32); 5], landmarks: &[(f32, f32); 5],
) -> Option<Aligned112> { ) -> Option<Aligned112> {
if rgb.len() != width * height * 3 { warp_pixels(Pixels::RgbF32(rgb), width, height, landmarks)
}
/// TRACES: FR-CULL-8
/// What the warp may sample, in whichever layout the caller already holds.
///
/// # Why the 8-bit variant exists
///
/// FR-CULL-8 requires the crop to come from the **native** render, and a native
/// render is large: a 24 MP frame is 96 MB as `RGBA8` and 288 MB converted to
/// the `f32` RGB this module was originally written against. Converting the
/// whole frame to sample 112×112 from it is three hundred megabytes allocated
/// to read about forty thousand pixels, per image, on a pass that runs over a
/// whole library — and on Android it is NFR-RES-2's budget spent outright.
///
/// So the warp reads whatever the caller has instead. It touches so few pixels
/// that the per-sample conversion is free, and the buffer never has to be
/// duplicated in another layout.
#[derive(Debug, Clone, Copy)]
pub enum Pixels<'a> {
/// Tightly packed `f32` RGB in `0.0..=1.0`, row-major.
RgbF32(&'a [f32]),
/// Tightly packed 8-bit RGBA, row-major. Alpha is ignored: a face crop has
/// no use for it and carrying it would change what the embedder receives.
Rgba8(&'a [u8]),
}
impl Pixels<'_> {
/// Whether the buffer is the size `width × height` implies.
fn fits(&self, width: usize, height: usize) -> bool {
match self {
Pixels::RgbF32(v) => v.len() == width * height * 3,
Pixels::Rgba8(v) => v.len() == width * height * 4,
}
}
/// One channel of one pixel, as `0.0..=1.0`. Outside the buffer reads black.
///
/// Public because the face *crop* stored for the People screen is cut from
/// the same buffer by the same caller, and it should not need a second
/// copy of this to do it.
pub fn channel(&self, w: usize, h: usize, x: isize, y: isize, c: usize) -> f32 {
if x < 0 || y < 0 || x >= w as isize || y >= h as isize {
return 0.0;
}
let i = y as usize * w + x as usize;
match self {
Pixels::RgbF32(v) => v[i * 3 + c],
Pixels::Rgba8(v) => v[i * 4 + c] as f32 / 255.0,
}
}
}
/// [`warp`], over any layout [`Pixels`] describes.
pub fn warp_pixels(
px: Pixels<'_>,
width: usize,
height: usize,
landmarks: &[(f32, f32); 5],
) -> Option<Aligned112> {
if !px.fits(width, height) {
return None; return None;
} }
let m = fit_similarity(landmarks, &ARCFACE_TEMPLATE)?; let m = fit_similarity(landmarks, &ARCFACE_TEMPLATE)?;
@@ -300,7 +360,7 @@ pub fn warp(
let (x, y) = m.invert(u as f32 + 0.5, v as f32 + 0.5); let (x, y) = m.invert(u as f32 + 0.5, v as f32 + 0.5);
let (x, y) = (x - 0.5, y - 0.5); let (x, y) = (x - 0.5, y - 0.5);
let out = (v * e + u) * 3; let out = (v * e + u) * 3;
sample_bilinear(rgb, width, height, x, y, &mut pixels[out..out + 3]); sample_bilinear(px, width, height, x, y, &mut pixels[out..out + 3]);
} }
} }
@@ -312,7 +372,7 @@ pub fn warp(
}) })
} }
fn sample_bilinear(rgb: &[f32], w: usize, h: usize, x: f32, y: f32, out: &mut [f32]) { fn sample_bilinear(px: Pixels<'_>, w: usize, h: usize, x: f32, y: f32, out: &mut [f32]) {
let x0 = x.floor(); let x0 = x.floor();
let y0 = y.floor(); let y0 = y.floor();
let fx = x - x0; let fx = x - x0;
@@ -321,13 +381,7 @@ fn sample_bilinear(rgb: &[f32], w: usize, h: usize, x: f32, y: f32, out: &mut [f
let y0 = y0 as isize; let y0 = y0 as isize;
for (c, o) in out.iter_mut().enumerate() { for (c, o) in out.iter_mut().enumerate() {
let get = |xi: isize, yi: isize| -> f32 { let get = |xi: isize, yi: isize| -> f32 { px.channel(w, h, xi, yi, c) };
if xi < 0 || yi < 0 || xi >= w as isize || yi >= h as isize {
0.0
} else {
rgb[(yi as usize * w + xi as usize) * 3 + c]
}
};
let top = get(x0, y0) * (1.0 - fx) + get(x0 + 1, y0) * fx; let top = get(x0, y0) * (1.0 - fx) + get(x0 + 1, y0) * fx;
let bot = get(x0, y0 + 1) * (1.0 - fx) + get(x0 + 1, y0 + 1) * fx; let bot = get(x0, y0 + 1) * (1.0 - fx) + get(x0 + 1, y0 + 1) * fx;
*o = top * (1.0 - fy) + bot * fy; *o = top * (1.0 - fy) + bot * fy;
@@ -508,6 +562,33 @@ mod tests {
/// against a bright sky is low-contrast, and a raw Laplacian variance would /// against a bright sky is low-contrast, and a raw Laplacian variance would
/// reject it as blurred — which would quietly throw away every backlit /// reject it as blurred — which would quietly throw away every backlit
/// portrait in the library. /// portrait in the library.
#[test]
fn both_pixel_layouts_warp_to_the_same_crop() {
// The 8-bit path exists so a native render need not be converted to
// f32 whole; it has to agree with the path it replaces to within the
// quantisation it introduces.
let (w, h) = (64usize, 64usize);
let mut rgba = vec![0u8; w * h * 4];
let mut rgb = vec![0.0f32; w * h * 3];
for y in 0..h {
for x in 0..w {
let v = [(x * 4 % 256) as u8, (y * 4 % 256) as u8, ((x + y) % 256) as u8];
for c in 0..3 {
rgba[(y * w + x) * 4 + c] = v[c];
rgb[(y * w + x) * 3 + c] = v[c] as f32 / 255.0;
}
rgba[(y * w + x) * 4 + 3] = 255;
}
}
let lm = shifted_scaled(0.35, 32.0, 32.0, 0.2);
let a = warp_pixels(Pixels::RgbF32(&rgb), w, h, &lm).unwrap();
let b = warp_pixels(Pixels::Rgba8(&rgba), w, h, &lm).unwrap();
assert_eq!(a.source_px(), b.source_px());
for (x, y) in a.pixels().iter().zip(b.pixels()) {
assert!((x - y).abs() < 1e-6, "{x} vs {y}");
}
}
#[test] #[test]
fn sharpness_survives_the_contrast_being_halved() { fn sharpness_survives_the_contrast_being_halved() {
let edge = 200; let edge = 200;
+1 -1
View File
@@ -64,7 +64,7 @@ pub mod neighbours;
/// from the wrong tier. /// from the wrong tier.
pub const MIN_CROP_EDGE: u32 = 1025; pub const MIN_CROP_EDGE: u32 = 1025;
pub use align::{warp, Aligned112, Similarity, ALIGNED_EDGE, ARCFACE_TEMPLATE}; pub use align::{warp, warp_pixels, Aligned112, Pixels, Similarity, ALIGNED_EDGE, ARCFACE_TEMPLATE};
pub use assign::{identity_shares, RIVAL_FLOOR, TOP_MATCHES}; pub use assign::{identity_shares, RIVAL_FLOOR, TOP_MATCHES};
pub use calibrate::{Calibration, Pairs, ReliabilityBand}; pub use calibrate::{Calibration, Pairs, ReliabilityBand};
pub use cluster::{ pub use cluster::{
+53 -53
View File
File diff suppressed because one or more lines are too long
+144
View File
@@ -0,0 +1,144 @@
//! TRACES: FR-CULL-8 | FR-EXP-9
//! Measure what indexing at native resolution is actually worth.
//!
//! cargo run -p dr-ui --example face_native -- DET.onnx EMB.onnx FILE [FILE…]
//!
//! Renders each file once at native resolution, then indexes it twice: the way
//! FR-CULL-8 now specifies, and the way it used to be done — everything, both
//! stages, from a 1024px proxy. Prints the faces found and the `crop_px` each
//! run gave the embedder.
//!
//! It exists because the case for the change was made from `crop_px` readings
//! taken out of a catalog after the fact. That is evidence about what happened;
//! this is evidence about what the new code does, on the same photographs, with
//! nothing between the two runs but the resolution.
//!
//! The models must have had their input dims frozen first; see
//! `tools/fix-face-model-shapes.sh`.
use std::path::PathBuf;
const MODEL_ID: &str = "w600k_mbf";
fn main() {
env_logger::init();
let args: Vec<String> = std::env::args().skip(1).collect();
if args.len() < 3 {
eprintln!("usage: face_native DETECTOR.onnx EMBEDDER.onnx FILE [FILE…]");
std::process::exit(2);
}
let (detector_model, embedder_model) = (PathBuf::from(&args[0]), PathBuf::from(&args[1]));
let Some(gpu) = pollster::block_on(dr_gpu::GpuContext::new_headless()).ok() else {
eprintln!("no GPU adapter; a native render needs one");
std::process::exit(1);
};
let mut detector = match dr_face::Detector::from_path(&detector_model) {
Ok(d) => d,
Err(e) => {
eprintln!("detector: {e}");
std::process::exit(1);
}
};
let mut embedder =
match dr_face::Embedder::from_path(&embedder_model, dr_face::ModelId::new(MODEL_ID)) {
Ok(e) => e,
Err(e) => {
eprintln!("embedder: {e}");
std::process::exit(1);
}
};
let options = dr_face::DetectOptions::default();
println!(
"{:<28} {:>11} {:>17} {:>17}",
"file", "native", "native faces/px", "1024 faces/px"
);
let (mut n_native, mut n_proxy) = (0usize, 0usize);
let (mut px_native, mut px_proxy) = (0.0f32, 0.0f32);
for path in &args[2..] {
let bytes = match std::fs::read(path) {
Ok(b) => b,
Err(e) => {
println!("{path}: cannot read: {e}");
continue;
}
};
let frame = match dr_ui::render_native(&gpu, &bytes) {
Ok(f) => f,
Err(e) => {
println!("{path}: cannot render: {e}");
continue;
}
};
let (w, h) = (frame.width as usize, frame.height as usize);
let native = dr_ui::faces::index_native(
&mut detector,
&mut embedder,
&frame.rgba,
w,
h,
&options,
)
.unwrap_or_default();
// The old path, reproduced exactly: one buffer at 1024, used for both
// detection and the crop.
let mut small = dr_decode::Preview {
width: frame.width,
height: frame.height,
rgba: frame.rgba.clone(),
};
small.downscale_to(dr_thumbs::ThumbSize::Large.edge());
let proxy = dr_ui::faces::index_preview(&mut detector, &mut embedder, &small, &options)
.map(|(f, _)| f)
.unwrap_or_default();
let mean = |v: &[dr_catalog::faces::DetectedFace]| {
if v.is_empty() {
0.0
} else {
v.iter().map(|f| f.crop_px).sum::<f32>() / v.len() as f32
}
};
let name = std::path::Path::new(path)
.file_name()
.map(|s| s.to_string_lossy().into_owned())
.unwrap_or_else(|| path.clone());
println!(
"{:<28} {:>11} {:>10} /{:>5.0} {:>10} /{:>5.0}",
name,
format!("{}x{}", frame.width, frame.height),
native.len(),
mean(&native),
proxy.len(),
mean(&proxy),
);
n_native += native.len();
n_proxy += proxy.len();
px_native += native.iter().map(|f| f.crop_px).sum::<f32>();
px_proxy += proxy.iter().map(|f| f.crop_px).sum::<f32>();
}
println!(
"\ntotal: native {n_native} face(s), mean crop {:.0}px | \
1024 proxy {n_proxy} face(s), mean crop {:.0}px",
if n_native == 0 {
0.0
} else {
px_native / n_native as f32
},
if n_proxy == 0 {
0.0
} else {
px_proxy / n_proxy as f32
},
);
}
+310
View File
@@ -327,6 +327,253 @@ pub fn index_proxy(
Ok(out) Ok(out)
} }
/// TRACES: FR-CULL-8
/// Long edge the detector's input is reduced to, in pixels.
///
/// Detection is indifferent above this — `dr_face::detect` letterboxes into a
/// fixed 640×640 whatever it receives, so a face is the same size to the model
/// from a 1600px buffer as from a 6000px one (faces.md §4.1). What the
/// reduction buys is the aliasing that a single bilinear step from native to
/// 640 would introduce: a 9× decimation samples one pixel in nine and drops
/// small faces into the gaps between samples. Box-filtering to 1600 first, and
/// letting the letterbox take the remaining 2.5×, keeps every source pixel in
/// the average.
///
/// It also bounds the buffer that has to exist as `f32`: 1600×1067 is 20 MB
/// against 288 MB for a native 24 MP frame.
const DETECT_EDGE: usize = 1600;
/// TRACES: FR-CULL-8 | NFR-RES-2
/// Detect and embed every face in one **native-resolution** render.
///
/// The pass FR-CULL-8 specifies, and the two resolutions in it are the whole
/// point:
///
/// - the **detector** is handed a box-filtered reduction, because it discards
/// anything above its own 640px input anyway;
/// - the **crop** is warped out of the native buffer, because that is the one
/// place source resolution becomes embedding quality.
///
/// Boxes and landmarks come back in the reduction's coordinates and are scaled
/// to native before a single crop pixel is read. Getting that scaling wrong
/// does not fail loudly — it yields faces in plausible-looking places with
/// crops taken from beside them — so it is one multiply applied in one place
/// rather than at each use.
///
/// `rgba` is the native render, tightly packed 8-bit RGBA. It is never
/// converted wholesale: [`dr_face::Pixels`] samples it where the warp and the
/// stored crop actually touch it.
pub fn index_native(
detector: &mut Detector,
embedder: &mut Embedder,
rgba: &[u8],
width: usize,
height: usize,
options: &DetectOptions,
) -> Result<Vec<DetectedFace>, dr_face::FaceError> {
if !index_native_shape_ok(rgba, width, height) {
return Ok(Vec::new());
}
let long_edge = width.max(height) as f32;
let native = dr_face::Pixels::Rgba8(rgba);
// One reduction, reused for every face on the image.
let (dw, dh, small) = reduce_for_detection(rgba, width, height);
if dw == 0 || dh == 0 {
return Ok(Vec::new());
}
let dets = detector.detect(&small, dw, dh, options)?;
// Detector coordinates to native. Separate factors rather than one, because
// the reduction rounds each axis independently and assuming they match puts
// every landmark a fraction of a face off on the shorter one.
let (sx, sy) = (width as f32 / dw as f32, height as f32 / dh as f32);
let mut out = Vec::with_capacity(dets.len());
for d in &dets {
let landmarks = scale_landmarks(&d.landmarks, sx, sy);
let Some(aligned) = dr_face::warp_pixels(native, width, height, &landmarks) else {
log::debug!("face with degenerate landmarks skipped");
continue;
};
// Both gates read the aligned crop, so they mean what they say only now
// that the crop comes from native pixels — `source_px` is the real
// count, not the proxy's idea of it. See `index_proxy` for why each
// one is here rather than in the detector.
if aligned.source_px() < options.min_source_px {
log::debug!(
"face skipped: {:.0} source px below {:.0}",
aligned.source_px(),
options.min_source_px
);
continue;
}
let sharpness = aligned.sharpness();
if sharpness < options.min_sharpness {
log::debug!(
"face at {:.0}px skipped: sharpness {sharpness:.4} below {:.4}",
aligned.source_px(),
options.min_sharpness
);
continue;
}
let embedding = embedder.embed(&aligned)?;
let (bx, by) = (d.bbox.0 * sx, d.bbox.1 * sy);
let (bw, bh) = (d.width() * sx, d.height() * sy);
out.push(DetectedFace {
x: bx / long_edge,
y: by / long_edge,
w: bw / long_edge,
h: bh / long_edge,
landmarks: normalise_landmarks(&landmarks, long_edge),
confidence: d.confidence,
embedding: embedding.to_f16_bytes(),
crop_px: aligned.source_px(),
model_id: embedder.model().as_str().to_string(),
crop: cut_crop_native(native, width, height, (bx, by, bw, bh)).unwrap_or_default(),
});
}
Ok(out)
}
/// Whether a buffer is the native frame it claims to be.
///
/// Checked before anything expensive, and before the detector above all: a
/// mismatched buffer would otherwise be read past its end by the reduction.
fn index_native_shape_ok(rgba: &[u8], width: usize, height: usize) -> bool {
width > 0 && height > 0 && rgba.len() == width * height * 4
}
/// Landmarks from the detector's reduction into native coordinates.
fn scale_landmarks(lm: &[(f32, f32); 5], sx: f32, sy: f32) -> [(f32, f32); 5] {
let mut out = [(0.0_f32, 0.0_f32); 5];
for (o, &(x, y)) in out.iter_mut().zip(lm.iter()) {
*o = (x * sx, y * sy);
}
out
}
/// Box-filter a native RGBA render down to [`DETECT_EDGE`], as packed `f32` RGB.
///
/// Averaging rather than sampling. The detector's job is to find small faces,
/// and a point-sampled reduction is exactly the operation that removes them:
/// at a 4× decimation fifteen of every sixteen pixels are discarded, and a
/// 40px face survives or not depending on where it happens to sit relative to
/// the sample grid.
fn reduce_for_detection(rgba: &[u8], width: usize, height: usize) -> (usize, usize, Vec<f32>) {
reduce_to(rgba, width, height, DETECT_EDGE)
}
/// [`reduce_for_detection`], to a stated long edge.
///
/// Split out so the averaging can be tested at a size a test can reason about
/// by hand, rather than by building a 1600px fixture.
fn reduce_to(
rgba: &[u8],
width: usize,
height: usize,
target: usize,
) -> (usize, usize, Vec<f32>) {
let long = width.max(height);
if long <= target {
// Already small enough; convert without resampling rather than round
// -tripping through a scale of 1.
let mut out = vec![0.0_f32; width * height * 3];
for (i, px) in rgba.chunks_exact(4).enumerate() {
for c in 0..3 {
out[i * 3 + c] = px[c] as f32 / 255.0;
}
}
return (width, height, out);
}
let scale = target as f32 / long as f32;
let (dw, dh) = (
((width as f32 * scale).round() as usize).max(1),
((height as f32 * scale).round() as usize).max(1),
);
let mut out = vec![0.0_f32; dw * dh * 3];
for oy in 0..dh {
// Source span of this destination row, as a half-open range, so
// adjacent rows tile the source exactly and no row is counted twice.
let y0 = oy * height / dh;
let y1 = (((oy + 1) * height) / dh).max(y0 + 1).min(height);
for ox in 0..dw {
let x0 = ox * width / dw;
let x1 = (((ox + 1) * width) / dw).max(x0 + 1).min(width);
let mut acc = [0.0_f32; 3];
let mut n = 0.0_f32;
for y in y0..y1 {
for x in x0..x1 {
let i = (y * width + x) * 4;
for (c, a) in acc.iter_mut().enumerate() {
*a += rgba[i + c] as f32;
}
n += 1.0;
}
}
let o = (oy * dw + ox) * 3;
for c in 0..3 {
out[o + c] = acc[c] / n / 255.0;
}
}
}
(dw, dh, out)
}
/// The stored face thumbnail, cut from the native buffer.
///
/// Same framing as [`cut_crop`] and deliberately a separate function rather
/// than a generalisation of it: this one takes a box already in native
/// coordinates, and blurring that distinction is how a crop ends up sampled
/// from the wrong scale.
fn cut_crop_native(
px: dr_face::Pixels<'_>,
width: usize,
height: usize,
bbox: (f32, f32, f32, f32),
) -> Option<Vec<u8>> {
if width == 0 || height == 0 {
return None;
}
let (bx, by, bw, bh) = bbox;
let cx = bx + bw * 0.5;
let cy = by + bh * 0.5;
let half = bw.max(bh) * 0.5 * (1.0 + CROP_MARGIN);
if !(half.is_finite() && half > 0.5 && cx.is_finite() && cy.is_finite()) {
return None;
}
let edge = STORED_CROP_EDGE;
let mut out = vec![0u8; (edge * edge * 4) as usize];
let step = (half * 2.0) / edge as f32;
for oy in 0..edge {
let sy = cy - half + (oy as f32 + 0.5) * step;
for ox in 0..edge {
let sx = cx - half + (ox as f32 + 0.5) * step;
let o = ((oy * edge + ox) * 4) as usize;
for c in 0..3 {
let v = px.channel(width, height, sx as isize, sy as isize, c);
out[o + c] = (v.clamp(0.0, 1.0) * 255.0).round() as u8;
}
out[o + 3] = 255;
}
}
match dr_thumbs::encode_rgba(edge, edge, &out) {
Ok(bytes) => Some(bytes),
Err(e) => {
log::debug!("encoding a face crop: {e}");
None
}
}
}
fn normalise_landmarks(lm: &[(f32, f32); 5], long_edge: f32) -> [(f32, f32); 5] { fn normalise_landmarks(lm: &[(f32, f32); 5], long_edge: f32) -> [(f32, f32); 5] {
let mut out = [(0.0_f32, 0.0_f32); 5]; let mut out = [(0.0_f32, 0.0_f32); 5];
for (o, &(x, y)) in out.iter_mut().zip(lm.iter()) { for (o, &(x, y)) in out.iter_mut().zip(lm.iter()) {
@@ -1084,6 +1331,69 @@ fn rgba_to_rgb_f32(rgba: &[u8]) -> Vec<f32> {
mod tests { mod tests {
use super::*; use super::*;
#[test]
fn the_detector_reduction_averages_rather_than_samples() {
// Four source pixels per destination pixel, with one bright pixel in
// each group. A point sampler returns either 255 or 0 depending on
// which corner it lands on; the average is the same every time, and
// that stability is what stops a small face vanishing on a grid
// alignment it has no control over.
let (w, h) = (4usize, 4usize);
let mut rgba = vec![0u8; w * h * 4];
for y in 0..h {
for x in 0..w {
let i = (y * w + x) * 4;
let bright = x % 2 == 0 && y % 2 == 0;
for c in 0..3 {
rgba[i + c] = if bright { 255 } else { 0 };
}
rgba[i + 3] = 255;
}
}
// Force a 2x reduction regardless of DETECT_EDGE by reducing by hand
// through the same helper on a source larger than the target.
let (dw, dh, out) = reduce_to(&rgba, w, h, 2);
assert_eq!((dw, dh), (2, 2));
for v in out.chunks_exact(3) {
assert!(
(v[0] - 0.25).abs() < 1e-6,
"each destination pixel averages one bright of four, got {}",
v[0]
);
}
}
#[test]
fn the_reduction_keeps_the_aspect_ratio_it_was_given() {
let (w, h) = (400usize, 100usize);
let rgba = vec![128u8; w * h * 4];
let (dw, dh, out) = reduce_to(&rgba, w, h, 100);
assert_eq!(dw, 100);
assert_eq!(dh, 25);
assert_eq!(out.len(), dw * dh * 3);
}
#[test]
fn landmarks_scale_back_to_the_native_frame() {
// The failure this guards is silent: landmarks left in the detector's
// coordinates put every crop near the top-left corner of the frame,
// which produces faces that look like faces of something else.
let lm = [(10.0, 20.0), (30.0, 40.0), (50.0, 60.0), (70.0, 80.0), (90.0, 100.0)];
let out = scale_landmarks(&lm, 4.0, 2.0);
assert_eq!(out[0], (40.0, 40.0));
assert_eq!(out[4], (360.0, 200.0));
}
#[test]
fn a_buffer_that_is_not_the_stated_size_indexes_nothing() {
// No model is loaded here, so reaching the detector would panic. The
// point is that it does not: the shape check comes first.
let rgba = vec![0u8; 10];
assert!(!index_native_shape_ok(&rgba, 100, 100));
assert!(!index_native_shape_ok(&rgba, 0, 0));
assert!(index_native_shape_ok(&vec![0u8; 4 * 100 * 100], 100, 100));
}
#[test] #[test]
fn an_audit_summary_names_both_kinds_of_outstanding() { fn an_audit_summary_names_both_kinds_of_outstanding() {
let a = IndexAudit { let a = IndexAudit {
+15
View File
@@ -537,11 +537,18 @@ pub fn wire<S, M, P>(
store: S, store: S,
models: M, models: M,
paths: P, paths: P,
// TRACES: FR-CULL-8
// `None` on a build with no adapter. Indexing needs a native render and a
// native render needs the GPU, so the button reports that the same way it
// reports a missing model rather than starting a pass that cannot produce
// anything.
gpu: Option<dr_gpu::GpuContext>,
) where ) where
S: Fn() -> Option<Rc<ThumbStore>> + 'static, S: Fn() -> Option<Rc<ThumbStore>> + 'static,
M: Fn() -> Option<ModelPaths> + 'static, M: Fn() -> Option<ModelPaths> + 'static,
P: Fn() -> Option<SweepPaths> + 'static, P: Fn() -> Option<SweepPaths> + 'static,
{ {
let gpu = Rc::new(gpu);
let store: Rc<dyn Fn() -> Option<Rc<ThumbStore>>> = Rc::new(store); let store: Rc<dyn Fn() -> Option<Rc<ThumbStore>>> = Rc::new(store);
let models: Rc<dyn Fn() -> Option<ModelPaths>> = Rc::new(models); let models: Rc<dyn Fn() -> Option<ModelPaths>> = Rc::new(models);
let paths: Rc<dyn Fn() -> Option<SweepPaths>> = Rc::new(paths); let paths: Rc<dyn Fn() -> Option<SweepPaths>> = Rc::new(paths);
@@ -1026,6 +1033,7 @@ pub fn wire<S, M, P>(
let store = store.clone(); let store = store.clone();
let models = models.clone(); let models = models.clone();
let paths = paths.clone(); let paths = paths.clone();
let gpu = gpu.clone();
window.on_identity_index(move || { window.on_identity_index(move || {
let Some(w) = weak.upgrade() else { return }; let Some(w) = weak.upgrade() else { return };
if ctl.sweep.borrow().is_some() { if ctl.sweep.borrow().is_some() {
@@ -1038,6 +1046,12 @@ pub fn wire<S, M, P>(
let Some((conn, catalog_path, store_dir)) = paths() else { let Some((conn, catalog_path, store_dir)) = paths() else {
return; return;
}; };
// Same reporting path as a missing model: both mean the pass
// cannot run, and both are states a fresh install can be in.
let Some(gpu) = gpu.as_ref().clone() else {
w.set_identity_model_missing(true);
return;
};
ctl.progress.set((0, 0)); ctl.progress.set((0, 0));
ctl.faces_found.set(0); ctl.faces_found.set(0);
@@ -1052,6 +1066,7 @@ pub fn wire<S, M, P>(
embedder, embedder,
MODEL_ID.to_string(), MODEL_ID.to_string(),
dr_face::DetectOptions::default(), dr_face::DetectOptions::default(),
gpu,
)); ));
w.set_identity_indexing(true); w.set_identity_indexing(true);
w.set_identity_indexing_status("looking for images to index…".into()); w.set_identity_indexing_status("looking for images to index…".into());
+6
View File
@@ -27,6 +27,7 @@ mod develop;
mod display_ui; mod display_ui;
mod export; mod export;
pub mod faces; pub mod faces;
pub use library::render_native;
// Generated from the `GESTURE:` comments beside the code that implements each // Generated from the `GESTURE:` comments beside the code that implements each
// one — see `tools/traceability`. Regenerate with // one — see `tools/traceability`. Regenerate with
// `cargo run -p traceability -- gestures`; CI fails if it has drifted. // `cargo run -p traceability -- gestures`; CI fails if it has drifted.
@@ -1208,6 +1209,11 @@ pub fn run(paths: Vec<PathBuf>) -> Result<()> {
Some((conn, catalog, thumbs)) Some((conn, catalog, thumbs))
} }
}, },
// TRACES: FR-CULL-8
// The same device the develop view draws with. Face indexing
// renders natively through FR-EXP-9's path, so it needs a GPU
// for the same reason export does.
gpu.clone(),
); );
} }
+278 -206
View File
@@ -3007,17 +3007,6 @@ fn faces_unindexed(
Ok(rows) Ok(rows)
} }
/// One image the face sweep got through, on its way back from a fetch lane.
///
/// The catalog id, the faces found, the long edge they were normalised
/// against, and the proxy to keep — `None` where the image held no face, since
/// nothing will ever ask to crop one out of it.
type IndexedImage = (
i64,
Vec<dr_catalog::faces::DetectedFace>,
u32,
Option<(u64, dr_thumbs::Thumbnail)>,
);
/// Images whose faces have nothing left to be cut out of. /// Images whose faces have nothing left to be cut out of.
/// ///
@@ -3066,40 +3055,47 @@ fn faces_without_proxy(
Ok(rows) Ok(rows)
} }
/// TRACES: FR-CULL-8 | NFR-ARCH-2 /// TRACES: FR-CULL-8 | FR-EXP-9 | NFR-ARCH-2 | NFR-RES-2
/// Index faces across the **whole** library, fetching what it needs. /// Index faces across the **whole** library, at native resolution.
/// ///
/// # Why this replaced a pass that read the thumbnail store /// # The two resolutions, and why they are not one
/// ///
/// The previous sweep filtered its work list down to images that already had a /// FR-CULL-8 asks for a native render, a *reduction* for the detector, and the
/// `ThumbSize::Large` proxy on disk, on the reading that FR-CULL-8 keeps face /// crop taken back out of the native buffer. That is not three sizes for the
/// indexing off the network. It does not: it keeps indexing off the *full /// sake of it. The detector letterboxes whatever it is handed into a fixed
/// decode*, and says plainly that "where no proxy exists, the job requests one /// 640×640, so above that its input resolution decides nothing and paying for
/// at background priority". Nothing filled the large class for a whole library /// it is waste; `align::warp` produces the fixed 112×112 ArcFace sees, so
/// — `SWEEP_THUMB_SIZE` is deliberately `Grid` — so the pass could only ever /// *its* input resolution decides everything and economising there is a
/// reach photographs the user had personally zoomed into. Measured on the /// silent loss. The two stages want opposite things, and a single buffer
/// reference library: 220 images indexed out of 23,529. /// serving both is how this pass previously came to store 47% of the
/// reference library's faces upsampled (faces.md §7b).
/// ///
/// So this fetches, by exactly the two-stage route the thumbnail sweep uses: /// # Why this replaced a pass that read an embedded preview
/// the header, then the located preview's own byte range (FR-NC-3). No full
/// file is pulled and no RAW is decoded — an embedded preview is a JPEG.
/// ///
/// # Why it does not simply ride the thumbnail sweep /// The previous version range-fetched the JPEG preview embedded in each RAW
/// (FR-NC-3) and used it for both stages. It was cheap and it was the tier
/// FR-CULL-8 named at the time. On the reference library that preview tops out
/// at 3072 px against a ~6000 px sensor, which is what put those 8,505 faces
/// below the embedder's 112 px with no way to tell from the catalog that
/// anything was wrong.
/// ///
/// It could, and it would be free: that pass already fetches the largest /// Before that it filtered its work list to images with a `ThumbSize::Large`
/// preview, decodes it, and downscales it to 256. Riding it is the right shape /// proxy already on disk, which nothing filled for a whole library, so it
/// for *new* images and is the obvious next step. It cannot serve this /// reached 220 images out of 23,529.
/// operation, though, because the thumbnail sweep's work list is what the store
/// does not have — so every image already thumbnailed, which for an established
/// library is most of them, would never come back past the detector.
/// ///
/// # Cost, stated plainly /// # Cost, stated plainly
/// ///
/// One preview fetch per un-indexed image, capped at [`MAX_PREVIEW_BYTES`]. /// **One whole original per un-indexed image, and one full render.** On the
/// That is the same transfer the thumbnail sweep pays per image, paid a second /// reference library that is 412 GB and roughly a hundred minutes of decode —
/// time because the first one kept only 256 px. Resumable by construction: the /// a different order of thing from the byte ranges this used to pay, which is
/// work list is what the catalog has no `face_index` row for, so a kill costs /// why FR-CULL-8 makes a whole-library pass a transfer under FR-NC-6 rather
/// the images in flight and nothing else. /// than something that may start on its own.
///
/// Nothing is kept that was not already wanted: the original is borrowed and
/// given back (ARCH §9.0a), and the only thing written per image is the
/// 1024 px proxy the People screen crops from, and only where a face was
/// found. Resumable by construction — the work list is what the catalog has no
/// `face_index` row for, so a kill costs the images in flight and nothing else.
#[allow(clippy::too_many_arguments)] #[allow(clippy::too_many_arguments)]
pub fn spawn_face_sweep( pub fn spawn_face_sweep(
conn: Connection, conn: Connection,
@@ -3109,6 +3105,7 @@ pub fn spawn_face_sweep(
embedder_model: PathBuf, embedder_model: PathBuf,
model_id: String, model_id: String,
options: dr_face::DetectOptions, options: dr_face::DetectOptions,
gpu: dr_gpu::GpuContext,
) -> Receiver<crate::faces::FaceSweepMessage> { ) -> Receiver<crate::faces::FaceSweepMessage> {
use crate::faces::FaceSweepMessage; use crate::faces::FaceSweepMessage;
let (tx, rx) = std::sync::mpsc::channel(); let (tx, rx) = std::sync::mpsc::channel();
@@ -3239,16 +3236,7 @@ pub fn spawn_face_sweep(
} }
}; };
// One detector and one embedder for every lane. let (mut images, mut found, mut failed) = (0usize, 0usize, 0usize);
//
// The lanes are concurrent futures on a single thread, not threads,
// so they interleave only at await points — and inference contains
// none. A `RefCell` borrow therefore never overlaps another, and
// the alternative, a pair per lane, would be ~16 MB of weights
// duplicated for no parallelism at all.
let models = std::cell::RefCell::new((&mut detector, &mut embedder));
let (mut done, mut images, mut found, mut failed) = (0usize, 0usize, 0usize, 0usize);
let mut offline = false; let mut offline = false;
// TRACES: FR-NC-6c // TRACES: FR-NC-6c
@@ -3258,180 +3246,152 @@ pub fn spawn_face_sweep(
// the later one finished. Each gives its own back (ARCH §9.0a). // the later one finished. Each gives its own back (ARCH §9.0a).
let pool = dr_sync_folder::BorrowPool::new(); let pool = dr_sync_folder::BorrowPool::new();
for chunk in wanted.chunks(SWEEP_CHUNK) { // TRACES: NFR-RES-2
let lanes: Vec<Vec<&ThumbnailRequest>> = (0..SWEEP_LANES) // **Fetch wide, render narrow.** The chunk is one original per
.map(|lane| chunk.iter().skip(lane).step_by(SWEEP_LANES).collect()) // lane and not the 96 the preview sweep used, because the two
.collect(); // passes hold different things: that one kept an 8 MB preview per
// image, this one keeps whole RAWs. Six at ~23 MB is a working set
let results = futures_join_all(lanes.into_iter().map(|lane| { // a phone can carry; ninety-six is not.
//
// The render is then sequential, and that is not a limitation to
// be optimised away later. There is one GPU, so concurrent renders
// would queue on it anyway, and each one materialises a native
// frame -- 96 MB for a 24 MP photograph. Overlapping them would
// multiply the one allocation that actually threatens the budget
// while buying no parallelism that exists.
for chunk in wanted.chunks(SWEEP_LANES) {
let fetched = futures_join_all(chunk.iter().map(|req| {
let backend = &*backend; let backend = &*backend;
let models = &models;
let options = &options;
let pool = &pool; let pool = &pool;
async move { async move {
let mut indexed: Vec<IndexedImage> = Vec::new(); let held = match pool.borrow(backend, &RemotePath::new(&req.path)).await {
let mut discard = Vec::new(); Ok(h) => h,
let mut attempted = 0usize; Err(e) if e.indicates_offline() => {
let mut failed = 0usize; log::info!("face sweep: {e}");
let mut offline = false; return (req, Err(FetchOutcome::Offline));
for req in lane {
attempted += 1;
let _held =
match pool.borrow(backend, &RemotePath::new(&req.path)).await {
Ok(h) => h,
Err(e) if e.indicates_offline() => {
log::info!("face sweep: {e}");
attempted -= 1;
offline = true;
break;
}
Err(e) => {
log::debug!("face sweep: {}: {e}", req.path);
failed += 1;
continue;
}
};
match fetch_preview(backend, req, &mut discard).await {
PreviewOutcome::Ready(mut preview) => {
// No await inside this borrow — see the
// note where `models` is built.
let found = {
let mut m = models.borrow_mut();
let (det, emb) = &mut *m;
crate::faces::index_preview(det, emb, &preview, options)
};
match found {
Ok((faces, edge)) => {
// Keep the proxy only where there
// is a face to cut out of it. Two
// thirds of a personal library is
// landscapes and documents
// (docs/faces.md §7a), and those
// never need a crop — so this fills
// the large class for the images
// the People screen will actually
// ask about and leaves the rest
// alone, rather than paying the
// whole-library cost
// `SWEEP_THUMB_SIZE` avoids.
//
// Downscaled only now: detection
// needed the full buffer, and this
// is the last use of it.
let keep = match (faces.is_empty(), req.file_id) {
(false, Some(file_id)) => {
preview.downscale_to(
dr_thumbs::ThumbSize::Large.edge(),
);
encode_preview(file_id, &preview)
.map(|t| (file_id, t))
}
_ => None,
};
indexed.push((req.image_id, faces, edge, keep));
}
Err(e) => {
log::debug!("face sweep: {}: {e}", req.path);
failed += 1;
}
}
}
PreviewOutcome::Unavailable(reason) => {
log::debug!("face sweep: {}: {reason}", req.path);
failed += 1;
}
PreviewOutcome::Offline(reason) => {
log::info!("face sweep: server unreachable: {reason}");
attempted -= 1;
offline = true;
break;
}
} }
} Err(e) => {
(indexed, attempted, failed, offline) log::debug!("face sweep: {}: {e}", req.path);
return (req, Err(FetchOutcome::Failed));
}
};
// TRACES: FR-CULL-8
// The whole file, not FR-NC-3's byte range. The
// requirement now asks for a native render and there
// is no native render without the original -- which
// is why the pass is a transfer under FR-NC-6 and says
// so before it starts.
let id = RemoteId::Path(RemotePath::new(&req.path));
let got = match backend.get(&id, None).await {
Ok(b) => Ok(b),
Err(e) if e.indicates_offline() => {
log::info!("face sweep: server unreachable: {e}");
Err(FetchOutcome::Offline)
}
Err(e) => {
log::debug!("face sweep: {}: {e}", req.path);
Err(FetchOutcome::Failed)
}
};
drop(held);
(req, got)
} }
})) }))
.await; .await;
for (indexed, attempted, lane_failed, lane_offline) in results { let mut lane_failed = 0usize;
done += attempted; for (req, got) in fetched {
failed += lane_failed; let bytes = match got {
offline |= lane_offline; Ok(b) => b,
// Before the indexed ones, so a screen watching this sees Err(FetchOutcome::Offline) => {
// the count move for work that produced nothing. Images offline = true;
// refused for size are *not* reported here: they are in continue;
// `indexed` below, having earned a marker, and counting
// them twice would run the progress figure past its total.
if lane_failed > 0
&& tx
.send(FaceSweepMessage::Failed {
images: lane_failed,
})
.is_err()
{
log::info!("face sweep: cancelled after {images} image(s)");
pool.release_all(&*backend).await;
return;
}
for (image_id, faces, edge, keep) in indexed {
// Before the detections, so a kill between the two
// leaves a proxy with no faces recorded — which the
// next pass simply re-indexes — rather than faces with
// no proxy, which is the state that draws an empty
// grid and cannot repair itself.
if let Some((file_id, thumb)) = keep {
store_thumbnail(
&mut store,
file_id,
dr_thumbs::ThumbSize::Large,
&thumb,
);
} }
// Written per image, including the ones with no face in Err(FetchOutcome::Failed) => {
// them: `face_index` records that detection *ran*, and lane_failed += 1;
// zero is its most valuable value — without the row, continue;
// every landscape and document scan returns on the next }
// pass, for ever (docs/faces.md §7a). };
match dr_catalog::faces::record_detections(
catalog.connection(), match index_one_native(
dr_types::ImageId(image_id as u64), &gpu,
&model_id, &mut detector,
edge, &mut embedder,
&faces, &bytes,
) { &options,
Ok(_) => { ) {
images += 1; Ok((faces, edge, proxy)) => {
found += faces.len(); // Before the detections, so a kill between the two
if tx // leaves a proxy with no faces recorded -- which
.send(FaceSweepMessage::Indexed { // the next pass simply re-indexes -- rather than
image: dr_types::ImageId(image_id as u64), // faces with no proxy, which is the state that
faces: faces.len(), // draws an empty grid and cannot repair itself.
}) if let (Some(file_id), Some(thumb)) = (req.file_id, proxy) {
.is_err() store_thumbnail(
{ &mut store,
// Receiver dropped: the screen closed, or file_id,
// the user pressed Stop. Everything written dr_thumbs::ThumbSize::Large,
// so far stays written — and everything &thumb,
// borrowed is given back. A cancelled pass );
// that kept the library hydrated would be }
// the worst of both: the disk spent and match dr_catalog::faces::record_detections(
// the work abandoned. catalog.connection(),
log::info!("face sweep: cancelled after {images} image(s)"); dr_types::ImageId(req.image_id as u64),
pool.release_all(&*backend).await; &model_id,
return; edge,
&faces,
) {
Ok(_) => {
images += 1;
found += faces.len();
if tx
.send(FaceSweepMessage::Indexed {
image: dr_types::ImageId(req.image_id as u64),
faces: faces.len(),
})
.is_err()
{
// Receiver dropped: the screen closed,
// or the user pressed Stop. Everything
// written so far stays written -- and
// everything borrowed is given back.
log::info!(
"face sweep: cancelled after {images} image(s)"
);
pool.release_all(&*backend).await;
return;
}
}
Err(e) => {
log::warn!(
"face sweep: storing faces for {}: {e}",
req.image_id
);
lane_failed += 1;
} }
} }
Err(e) => { }
log::warn!("face sweep: storing faces for {image_id}: {e}"); Err(e) => {
failed += 1; log::debug!("face sweep: {}: {e}", req.path);
} lane_failed += 1;
} }
} }
} }
let _ = done; failed += lane_failed;
if lane_failed > 0
&& tx
.send(FaceSweepMessage::Failed {
images: lane_failed,
})
.is_err()
{
log::info!("face sweep: cancelled after {images} image(s)");
pool.release_all(&*backend).await;
return;
}
if offline { if offline {
break; break;
} }
@@ -3460,6 +3420,118 @@ pub fn spawn_face_sweep(
rx rx
} }
/// TRACES: FR-CULL-8 | FR-EXP-9
/// Open one original for a native render, orientation applied and nothing else.
///
/// The half of [`index_one_native`] that has nothing to do with faces, exposed
/// because measuring what this pass is worth means rendering the same file two
/// ways and comparing the crops — see `examples/face_native.rs`. A tool that
/// had to reimplement the render would be measuring its own reimplementation.
pub fn render_native(
gpu: &dr_gpu::GpuContext,
bytes: &[u8],
) -> Result<dr_export::Frame, String> {
open_native(gpu, bytes)?.render_for_export(dr_types::ColourSpace::Srgb)
}
/// The session behind [`render_native`], kept private because `DevelopSession`
/// is. The indexing path needs the session itself rather than just its frame:
/// it renders the People screen's proxy from the same open session rather than
/// opening the file twice.
fn open_native(
gpu: &dr_gpu::GpuContext,
bytes: &[u8],
) -> Result<crate::develop::DevelopSession, String> {
let orientation = dr_decode::metadata(bytes)
.ok()
.and_then(|m| m.orientation)
.unwrap_or_default();
crate::open_session(gpu, bytes, orientation)
}
/// Why one original did not arrive.
///
/// Named rather than a bool because the two mean opposite things to the loop:
/// one image failing is one image, and the server going away means nothing
/// after it would have worked either.
enum FetchOutcome {
Failed,
Offline,
}
/// TRACES: FR-CULL-8 | FR-EXP-9
/// Render one original at native resolution and index the faces in it.
///
/// The FR-CULL-8 pipeline end to end, for one photograph: decode, render
/// through the same path export uses, detect on a reduction, crop from the
/// native frame. Returns the faces, the native long edge that went into the
/// run marker, and the 1024px proxy the People screen later cuts thumbnails
/// from.
///
/// # Why the stored edit is not applied
///
/// `export::render_from_library` fetches the sidecar and applies it, because
/// an export is of the photograph the user has made. This is not: the face
/// geometry stored in the catalog is normalised to the frame, so applying a
/// crop would record faces against a frame that changes whenever the user
/// changes their mind, and every stored box would silently become wrong. What
/// is applied is the orientation, which is a fact about the file rather than
/// an edit.
fn index_one_native(
gpu: &dr_gpu::GpuContext,
detector: &mut dr_face::Detector,
embedder: &mut dr_face::Embedder,
bytes: &[u8],
options: &dr_face::DetectOptions,
) -> Result<(Vec<dr_catalog::faces::DetectedFace>, u32, Option<dr_thumbs::Thumbnail>), String> {
let mut session = open_native(gpu, bytes)?;
let frame = session.render_for_export(dr_types::ColourSpace::Srgb)?;
let edge = frame.width.max(frame.height);
let faces = crate::faces::index_native(
detector,
embedder,
&frame.rgba,
frame.width as usize,
frame.height as usize,
options,
)
.map_err(|e| e.to_string())?;
// Only where there is a face to cut out of it. Two thirds of a personal
// library is landscapes and documents, and those never need a crop -- so
// this fills the large class for the images the People screen will
// actually ask about and leaves the rest alone.
//
// Rendered rather than downscaled from the frame in hand: the session is
// still open and `render_thumbnail` is the path the grid's own thumbnails
// take, so the proxy this writes is the one the store would have had
// anyway.
let proxy = if faces.is_empty() {
None
} else {
match session.render_thumbnail(dr_thumbs::ThumbSize::Large.edge()) {
Ok((w, h, rgba)) => match dr_thumbs::encode_rgba(w, h, &rgba) {
Ok(bytes) => Some(dr_thumbs::Thumbnail {
width: w,
height: h,
bytes,
}),
Err(e) => {
log::debug!("encoding a face proxy: {e}");
None
}
},
Err(e) => {
log::debug!("rendering a face proxy: {e}");
None
}
}
};
Ok((faces, edge, proxy))
}
/// The class the whole-library pass fills. /// The class the whole-library pass fills.
/// ///
/// Grid only, deliberately. The large class is four times the transfer for a /// Grid only, deliberately. The large class is four times the transfer for a