Index faces from the native render, not from a preview of it

Implements the FR-CULL-8 written two commits ago. The sweep fetched the
JPEG preview embedded in each RAW and used that one buffer for both
detection and the crop; it now fetches the original, renders it through
the same path export uses, reduces that for the detector, and warps the
crop back out of the native frame.

Three pieces, and each exists for a reason worth stating.

dr_face::Pixels lets the warp sample 8-bit RGBA directly. A 24 MP native
frame is 96 MB as RGBA and 288 MB converted to the f32 RGB align.rs was
written against, and the warp reads about forty thousand pixels out of
it. Converting the whole frame to sample 0.2% of it is NFR-RES-2's
budget spent on a copy, per image, for a whole library. The variant
costs one branch per sample and a test asserts both layouts produce
identical crops.

The detector gets a box-filtered reduction to 1600px, not the native
frame and not a point-sampled one. Averaging rather than sampling
because the detector's job is finding small faces and decimation is
precisely the operation that removes them: at 4x, fifteen of every
sixteen pixels are discarded and a 40px face survives or not depending
on where it falls relative to the sample grid. 1600 rather than 640
leaves the letterbox a mild 2.5x rather than a 9x, and bounds the f32
buffer at 20 MB.

Landmarks come back in the reduction's coordinates and are scaled to
native in one place before any crop pixel is read. This is the failure
mode that would not announce itself -- unscaled landmarks put every crop
near the top-left corner, which yields faces of something else, cleanly
embedded and confidently clustered.

The sweep fetches SWEEP_LANES-wide and renders sequentially. Not a
placeholder for a parallel version: there is one GPU, so concurrent
renders queue on it regardless, and each materialises a native frame.
Overlapping them would multiply the one allocation that threatens the
memory budget while buying parallelism that does not exist. The chunk
drops from 96 to 6 for the same reason -- 96 held 8 MB previews, this
holds whole RAWs.

The stored edit is deliberately not applied, which is where this departs
from export::render_from_library. Face geometry is normalised to the
frame, so indexing a cropped render would record boxes against a frame
that changes whenever the user changes their mind, and every stored box
would quietly become wrong. Orientation is applied: that is a fact about
the file rather than an edit.

examples/face_native.rs renders one file and indexes it both ways, so
the claim behind all of this can be checked against photographs rather
than re-read out of the catalog it came from.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-30 19:41:30 +02:00
co-authored by Claude Opus 5
parent 9ddc1273c0
commit 4af3b93dfa
8 changed files with 898 additions and 270 deletions
+144
View File
@@ -0,0 +1,144 @@
//! TRACES: FR-CULL-8 | FR-EXP-9
//! Measure what indexing at native resolution is actually worth.
//!
//! cargo run -p dr-ui --example face_native -- DET.onnx EMB.onnx FILE [FILE…]
//!
//! Renders each file once at native resolution, then indexes it twice: the way
//! FR-CULL-8 now specifies, and the way it used to be done — everything, both
//! stages, from a 1024px proxy. Prints the faces found and the `crop_px` each
//! run gave the embedder.
//!
//! It exists because the case for the change was made from `crop_px` readings
//! taken out of a catalog after the fact. That is evidence about what happened;
//! this is evidence about what the new code does, on the same photographs, with
//! nothing between the two runs but the resolution.
//!
//! The models must have had their input dims frozen first; see
//! `tools/fix-face-model-shapes.sh`.
use std::path::PathBuf;
const MODEL_ID: &str = "w600k_mbf";
fn main() {
env_logger::init();
let args: Vec<String> = std::env::args().skip(1).collect();
if args.len() < 3 {
eprintln!("usage: face_native DETECTOR.onnx EMBEDDER.onnx FILE [FILE…]");
std::process::exit(2);
}
let (detector_model, embedder_model) = (PathBuf::from(&args[0]), PathBuf::from(&args[1]));
let Some(gpu) = pollster::block_on(dr_gpu::GpuContext::new_headless()).ok() else {
eprintln!("no GPU adapter; a native render needs one");
std::process::exit(1);
};
let mut detector = match dr_face::Detector::from_path(&detector_model) {
Ok(d) => d,
Err(e) => {
eprintln!("detector: {e}");
std::process::exit(1);
}
};
let mut embedder =
match dr_face::Embedder::from_path(&embedder_model, dr_face::ModelId::new(MODEL_ID)) {
Ok(e) => e,
Err(e) => {
eprintln!("embedder: {e}");
std::process::exit(1);
}
};
let options = dr_face::DetectOptions::default();
println!(
"{:<28} {:>11} {:>17} {:>17}",
"file", "native", "native faces/px", "1024 faces/px"
);
let (mut n_native, mut n_proxy) = (0usize, 0usize);
let (mut px_native, mut px_proxy) = (0.0f32, 0.0f32);
for path in &args[2..] {
let bytes = match std::fs::read(path) {
Ok(b) => b,
Err(e) => {
println!("{path}: cannot read: {e}");
continue;
}
};
let frame = match dr_ui::render_native(&gpu, &bytes) {
Ok(f) => f,
Err(e) => {
println!("{path}: cannot render: {e}");
continue;
}
};
let (w, h) = (frame.width as usize, frame.height as usize);
let native = dr_ui::faces::index_native(
&mut detector,
&mut embedder,
&frame.rgba,
w,
h,
&options,
)
.unwrap_or_default();
// The old path, reproduced exactly: one buffer at 1024, used for both
// detection and the crop.
let mut small = dr_decode::Preview {
width: frame.width,
height: frame.height,
rgba: frame.rgba.clone(),
};
small.downscale_to(dr_thumbs::ThumbSize::Large.edge());
let proxy = dr_ui::faces::index_preview(&mut detector, &mut embedder, &small, &options)
.map(|(f, _)| f)
.unwrap_or_default();
let mean = |v: &[dr_catalog::faces::DetectedFace]| {
if v.is_empty() {
0.0
} else {
v.iter().map(|f| f.crop_px).sum::<f32>() / v.len() as f32
}
};
let name = std::path::Path::new(path)
.file_name()
.map(|s| s.to_string_lossy().into_owned())
.unwrap_or_else(|| path.clone());
println!(
"{:<28} {:>11} {:>10} /{:>5.0} {:>10} /{:>5.0}",
name,
format!("{}x{}", frame.width, frame.height),
native.len(),
mean(&native),
proxy.len(),
mean(&proxy),
);
n_native += native.len();
n_proxy += proxy.len();
px_native += native.iter().map(|f| f.crop_px).sum::<f32>();
px_proxy += proxy.iter().map(|f| f.crop_px).sum::<f32>();
}
println!(
"\ntotal: native {n_native} face(s), mean crop {:.0}px | \
1024 proxy {n_proxy} face(s), mean crop {:.0}px",
if n_native == 0 {
0.0
} else {
px_native / n_native as f32
},
if n_proxy == 0 {
0.0
} else {
px_proxy / n_proxy as f32
},
);
}