Index faces from the native render, not from a preview of it

Implements the FR-CULL-8 written two commits ago. The sweep fetched the
JPEG preview embedded in each RAW and used that one buffer for both
detection and the crop; it now fetches the original, renders it through
the same path export uses, reduces that for the detector, and warps the
crop back out of the native frame.

Three pieces, and each exists for a reason worth stating.

dr_face::Pixels lets the warp sample 8-bit RGBA directly. A 24 MP native
frame is 96 MB as RGBA and 288 MB converted to the f32 RGB align.rs was
written against, and the warp reads about forty thousand pixels out of
it. Converting the whole frame to sample 0.2% of it is NFR-RES-2's
budget spent on a copy, per image, for a whole library. The variant
costs one branch per sample and a test asserts both layouts produce
identical crops.

The detector gets a box-filtered reduction to 1600px, not the native
frame and not a point-sampled one. Averaging rather than sampling
because the detector's job is finding small faces and decimation is
precisely the operation that removes them: at 4x, fifteen of every
sixteen pixels are discarded and a 40px face survives or not depending
on where it falls relative to the sample grid. 1600 rather than 640
leaves the letterbox a mild 2.5x rather than a 9x, and bounds the f32
buffer at 20 MB.

Landmarks come back in the reduction's coordinates and are scaled to
native in one place before any crop pixel is read. This is the failure
mode that would not announce itself -- unscaled landmarks put every crop
near the top-left corner, which yields faces of something else, cleanly
embedded and confidently clustered.

The sweep fetches SWEEP_LANES-wide and renders sequentially. Not a
placeholder for a parallel version: there is one GPU, so concurrent
renders queue on it regardless, and each materialises a native frame.
Overlapping them would multiply the one allocation that threatens the
memory budget while buying parallelism that does not exist. The chunk
drops from 96 to 6 for the same reason -- 96 held 8 MB previews, this
holds whole RAWs.

The stored edit is deliberately not applied, which is where this departs
from export::render_from_library. Face geometry is normalised to the
frame, so indexing a cropped render would record boxes against a frame
that changes whenever the user changes their mind, and every stored box
would quietly become wrong. Orientation is applied: that is a fact about
the file rather than an edit.

examples/face_native.rs renders one file and indexes it both ways, so
the claim behind all of this can be checked against photographs rather
than re-read out of the catalog it came from.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-30 19:41:30 +02:00
co-authored by Claude Opus 5
parent 9ddc1273c0
commit 4af3b93dfa
8 changed files with 898 additions and 270 deletions
+278 -206
View File
@@ -3007,17 +3007,6 @@ fn faces_unindexed(
Ok(rows)
}
/// One image the face sweep got through, on its way back from a fetch lane.
///
/// The catalog id, the faces found, the long edge they were normalised
/// against, and the proxy to keep — `None` where the image held no face, since
/// nothing will ever ask to crop one out of it.
type IndexedImage = (
i64,
Vec<dr_catalog::faces::DetectedFace>,
u32,
Option<(u64, dr_thumbs::Thumbnail)>,
);
/// Images whose faces have nothing left to be cut out of.
///
@@ -3066,40 +3055,47 @@ fn faces_without_proxy(
Ok(rows)
}
/// TRACES: FR-CULL-8 | NFR-ARCH-2
/// Index faces across the **whole** library, fetching what it needs.
/// TRACES: FR-CULL-8 | FR-EXP-9 | NFR-ARCH-2 | NFR-RES-2
/// Index faces across the **whole** library, at native resolution.
///
/// # Why this replaced a pass that read the thumbnail store
/// # The two resolutions, and why they are not one
///
/// The previous sweep filtered its work list down to images that already had a
/// `ThumbSize::Large` proxy on disk, on the reading that FR-CULL-8 keeps face
/// indexing off the network. It does not: it keeps indexing off the *full
/// decode*, and says plainly that "where no proxy exists, the job requests one
/// at background priority". Nothing filled the large class for a whole library
/// — `SWEEP_THUMB_SIZE` is deliberately `Grid` — so the pass could only ever
/// reach photographs the user had personally zoomed into. Measured on the
/// reference library: 220 images indexed out of 23,529.
/// FR-CULL-8 asks for a native render, a *reduction* for the detector, and the
/// crop taken back out of the native buffer. That is not three sizes for the
/// sake of it. The detector letterboxes whatever it is handed into a fixed
/// 640×640, so above that its input resolution decides nothing and paying for
/// it is waste; `align::warp` produces the fixed 112×112 ArcFace sees, so
/// *its* input resolution decides everything and economising there is a
/// silent loss. The two stages want opposite things, and a single buffer
/// serving both is how this pass previously came to store 47% of the
/// reference library's faces upsampled (faces.md §7b).
///
/// So this fetches, by exactly the two-stage route the thumbnail sweep uses:
/// the header, then the located preview's own byte range (FR-NC-3). No full
/// file is pulled and no RAW is decoded — an embedded preview is a JPEG.
/// # Why this replaced a pass that read an embedded preview
///
/// # Why it does not simply ride the thumbnail sweep
/// The previous version range-fetched the JPEG preview embedded in each RAW
/// (FR-NC-3) and used it for both stages. It was cheap and it was the tier
/// FR-CULL-8 named at the time. On the reference library that preview tops out
/// at 3072 px against a ~6000 px sensor, which is what put those 8,505 faces
/// below the embedder's 112 px with no way to tell from the catalog that
/// anything was wrong.
///
/// It could, and it would be free: that pass already fetches the largest
/// preview, decodes it, and downscales it to 256. Riding it is the right shape
/// for *new* images and is the obvious next step. It cannot serve this
/// operation, though, because the thumbnail sweep's work list is what the store
/// does not have — so every image already thumbnailed, which for an established
/// library is most of them, would never come back past the detector.
/// Before that it filtered its work list to images with a `ThumbSize::Large`
/// proxy already on disk, which nothing filled for a whole library, so it
/// reached 220 images out of 23,529.
///
/// # Cost, stated plainly
///
/// One preview fetch per un-indexed image, capped at [`MAX_PREVIEW_BYTES`].
/// That is the same transfer the thumbnail sweep pays per image, paid a second
/// time because the first one kept only 256 px. Resumable by construction: the
/// work list is what the catalog has no `face_index` row for, so a kill costs
/// the images in flight and nothing else.
/// **One whole original per un-indexed image, and one full render.** On the
/// reference library that is 412 GB and roughly a hundred minutes of decode —
/// a different order of thing from the byte ranges this used to pay, which is
/// why FR-CULL-8 makes a whole-library pass a transfer under FR-NC-6 rather
/// than something that may start on its own.
///
/// Nothing is kept that was not already wanted: the original is borrowed and
/// given back (ARCH §9.0a), and the only thing written per image is the
/// 1024 px proxy the People screen crops from, and only where a face was
/// found. Resumable by construction — the work list is what the catalog has no
/// `face_index` row for, so a kill costs the images in flight and nothing else.
#[allow(clippy::too_many_arguments)]
pub fn spawn_face_sweep(
conn: Connection,
@@ -3109,6 +3105,7 @@ pub fn spawn_face_sweep(
embedder_model: PathBuf,
model_id: String,
options: dr_face::DetectOptions,
gpu: dr_gpu::GpuContext,
) -> Receiver<crate::faces::FaceSweepMessage> {
use crate::faces::FaceSweepMessage;
let (tx, rx) = std::sync::mpsc::channel();
@@ -3239,16 +3236,7 @@ pub fn spawn_face_sweep(
}
};
// One detector and one embedder for every lane.
//
// The lanes are concurrent futures on a single thread, not threads,
// so they interleave only at await points — and inference contains
// none. A `RefCell` borrow therefore never overlaps another, and
// the alternative, a pair per lane, would be ~16 MB of weights
// duplicated for no parallelism at all.
let models = std::cell::RefCell::new((&mut detector, &mut embedder));
let (mut done, mut images, mut found, mut failed) = (0usize, 0usize, 0usize, 0usize);
let (mut images, mut found, mut failed) = (0usize, 0usize, 0usize);
let mut offline = false;
// TRACES: FR-NC-6c
@@ -3258,180 +3246,152 @@ pub fn spawn_face_sweep(
// the later one finished. Each gives its own back (ARCH §9.0a).
let pool = dr_sync_folder::BorrowPool::new();
for chunk in wanted.chunks(SWEEP_CHUNK) {
let lanes: Vec<Vec<&ThumbnailRequest>> = (0..SWEEP_LANES)
.map(|lane| chunk.iter().skip(lane).step_by(SWEEP_LANES).collect())
.collect();
let results = futures_join_all(lanes.into_iter().map(|lane| {
// TRACES: NFR-RES-2
// **Fetch wide, render narrow.** The chunk is one original per
// lane and not the 96 the preview sweep used, because the two
// passes hold different things: that one kept an 8 MB preview per
// image, this one keeps whole RAWs. Six at ~23 MB is a working set
// a phone can carry; ninety-six is not.
//
// The render is then sequential, and that is not a limitation to
// be optimised away later. There is one GPU, so concurrent renders
// would queue on it anyway, and each one materialises a native
// frame -- 96 MB for a 24 MP photograph. Overlapping them would
// multiply the one allocation that actually threatens the budget
// while buying no parallelism that exists.
for chunk in wanted.chunks(SWEEP_LANES) {
let fetched = futures_join_all(chunk.iter().map(|req| {
let backend = &*backend;
let models = &models;
let options = &options;
let pool = &pool;
async move {
let mut indexed: Vec<IndexedImage> = Vec::new();
let mut discard = Vec::new();
let mut attempted = 0usize;
let mut failed = 0usize;
let mut offline = false;
for req in lane {
attempted += 1;
let _held =
match pool.borrow(backend, &RemotePath::new(&req.path)).await {
Ok(h) => h,
Err(e) if e.indicates_offline() => {
log::info!("face sweep: {e}");
attempted -= 1;
offline = true;
break;
}
Err(e) => {
log::debug!("face sweep: {}: {e}", req.path);
failed += 1;
continue;
}
};
match fetch_preview(backend, req, &mut discard).await {
PreviewOutcome::Ready(mut preview) => {
// No await inside this borrow — see the
// note where `models` is built.
let found = {
let mut m = models.borrow_mut();
let (det, emb) = &mut *m;
crate::faces::index_preview(det, emb, &preview, options)
};
match found {
Ok((faces, edge)) => {
// Keep the proxy only where there
// is a face to cut out of it. Two
// thirds of a personal library is
// landscapes and documents
// (docs/faces.md §7a), and those
// never need a crop — so this fills
// the large class for the images
// the People screen will actually
// ask about and leaves the rest
// alone, rather than paying the
// whole-library cost
// `SWEEP_THUMB_SIZE` avoids.
//
// Downscaled only now: detection
// needed the full buffer, and this
// is the last use of it.
let keep = match (faces.is_empty(), req.file_id) {
(false, Some(file_id)) => {
preview.downscale_to(
dr_thumbs::ThumbSize::Large.edge(),
);
encode_preview(file_id, &preview)
.map(|t| (file_id, t))
}
_ => None,
};
indexed.push((req.image_id, faces, edge, keep));
}
Err(e) => {
log::debug!("face sweep: {}: {e}", req.path);
failed += 1;
}
}
}
PreviewOutcome::Unavailable(reason) => {
log::debug!("face sweep: {}: {reason}", req.path);
failed += 1;
}
PreviewOutcome::Offline(reason) => {
log::info!("face sweep: server unreachable: {reason}");
attempted -= 1;
offline = true;
break;
}
let held = match pool.borrow(backend, &RemotePath::new(&req.path)).await {
Ok(h) => h,
Err(e) if e.indicates_offline() => {
log::info!("face sweep: {e}");
return (req, Err(FetchOutcome::Offline));
}
}
(indexed, attempted, failed, offline)
Err(e) => {
log::debug!("face sweep: {}: {e}", req.path);
return (req, Err(FetchOutcome::Failed));
}
};
// TRACES: FR-CULL-8
// The whole file, not FR-NC-3's byte range. The
// requirement now asks for a native render and there
// is no native render without the original -- which
// is why the pass is a transfer under FR-NC-6 and says
// so before it starts.
let id = RemoteId::Path(RemotePath::new(&req.path));
let got = match backend.get(&id, None).await {
Ok(b) => Ok(b),
Err(e) if e.indicates_offline() => {
log::info!("face sweep: server unreachable: {e}");
Err(FetchOutcome::Offline)
}
Err(e) => {
log::debug!("face sweep: {}: {e}", req.path);
Err(FetchOutcome::Failed)
}
};
drop(held);
(req, got)
}
}))
.await;
for (indexed, attempted, lane_failed, lane_offline) in results {
done += attempted;
failed += lane_failed;
offline |= lane_offline;
// Before the indexed ones, so a screen watching this sees
// the count move for work that produced nothing. Images
// refused for size are *not* reported here: they are in
// `indexed` below, having earned a marker, and counting
// them twice would run the progress figure past its total.
if lane_failed > 0
&& tx
.send(FaceSweepMessage::Failed {
images: lane_failed,
})
.is_err()
{
log::info!("face sweep: cancelled after {images} image(s)");
pool.release_all(&*backend).await;
return;
}
for (image_id, faces, edge, keep) in indexed {
// Before the detections, so a kill between the two
// leaves a proxy with no faces recorded — which the
// next pass simply re-indexes — rather than faces with
// no proxy, which is the state that draws an empty
// grid and cannot repair itself.
if let Some((file_id, thumb)) = keep {
store_thumbnail(
&mut store,
file_id,
dr_thumbs::ThumbSize::Large,
&thumb,
);
let mut lane_failed = 0usize;
for (req, got) in fetched {
let bytes = match got {
Ok(b) => b,
Err(FetchOutcome::Offline) => {
offline = true;
continue;
}
// Written per image, including the ones with no face in
// them: `face_index` records that detection *ran*, and
// zero is its most valuable value — without the row,
// every landscape and document scan returns on the next
// pass, for ever (docs/faces.md §7a).
match dr_catalog::faces::record_detections(
catalog.connection(),
dr_types::ImageId(image_id as u64),
&model_id,
edge,
&faces,
) {
Ok(_) => {
images += 1;
found += faces.len();
if tx
.send(FaceSweepMessage::Indexed {
image: dr_types::ImageId(image_id as u64),
faces: faces.len(),
})
.is_err()
{
// Receiver dropped: the screen closed, or
// the user pressed Stop. Everything written
// so far stays written — and everything
// borrowed is given back. A cancelled pass
// that kept the library hydrated would be
// the worst of both: the disk spent and
// the work abandoned.
log::info!("face sweep: cancelled after {images} image(s)");
pool.release_all(&*backend).await;
return;
Err(FetchOutcome::Failed) => {
lane_failed += 1;
continue;
}
};
match index_one_native(
&gpu,
&mut detector,
&mut embedder,
&bytes,
&options,
) {
Ok((faces, edge, proxy)) => {
// Before the detections, so a kill between the two
// leaves a proxy with no faces recorded -- which
// the next pass simply re-indexes -- rather than
// faces with no proxy, which is the state that
// draws an empty grid and cannot repair itself.
if let (Some(file_id), Some(thumb)) = (req.file_id, proxy) {
store_thumbnail(
&mut store,
file_id,
dr_thumbs::ThumbSize::Large,
&thumb,
);
}
match dr_catalog::faces::record_detections(
catalog.connection(),
dr_types::ImageId(req.image_id as u64),
&model_id,
edge,
&faces,
) {
Ok(_) => {
images += 1;
found += faces.len();
if tx
.send(FaceSweepMessage::Indexed {
image: dr_types::ImageId(req.image_id as u64),
faces: faces.len(),
})
.is_err()
{
// Receiver dropped: the screen closed,
// or the user pressed Stop. Everything
// written so far stays written -- and
// everything borrowed is given back.
log::info!(
"face sweep: cancelled after {images} image(s)"
);
pool.release_all(&*backend).await;
return;
}
}
Err(e) => {
log::warn!(
"face sweep: storing faces for {}: {e}",
req.image_id
);
lane_failed += 1;
}
}
Err(e) => {
log::warn!("face sweep: storing faces for {image_id}: {e}");
failed += 1;
}
}
Err(e) => {
log::debug!("face sweep: {}: {e}", req.path);
lane_failed += 1;
}
}
}
let _ = done;
failed += lane_failed;
if lane_failed > 0
&& tx
.send(FaceSweepMessage::Failed {
images: lane_failed,
})
.is_err()
{
log::info!("face sweep: cancelled after {images} image(s)");
pool.release_all(&*backend).await;
return;
}
if offline {
break;
}
@@ -3460,6 +3420,118 @@ pub fn spawn_face_sweep(
rx
}
/// TRACES: FR-CULL-8 | FR-EXP-9
/// Open one original for a native render, orientation applied and nothing else.
///
/// The half of [`index_one_native`] that has nothing to do with faces, exposed
/// because measuring what this pass is worth means rendering the same file two
/// ways and comparing the crops — see `examples/face_native.rs`. A tool that
/// had to reimplement the render would be measuring its own reimplementation.
pub fn render_native(
gpu: &dr_gpu::GpuContext,
bytes: &[u8],
) -> Result<dr_export::Frame, String> {
open_native(gpu, bytes)?.render_for_export(dr_types::ColourSpace::Srgb)
}
/// The session behind [`render_native`], kept private because `DevelopSession`
/// is. The indexing path needs the session itself rather than just its frame:
/// it renders the People screen's proxy from the same open session rather than
/// opening the file twice.
fn open_native(
gpu: &dr_gpu::GpuContext,
bytes: &[u8],
) -> Result<crate::develop::DevelopSession, String> {
let orientation = dr_decode::metadata(bytes)
.ok()
.and_then(|m| m.orientation)
.unwrap_or_default();
crate::open_session(gpu, bytes, orientation)
}
/// Why one original did not arrive.
///
/// Named rather than a bool because the two mean opposite things to the loop:
/// one image failing is one image, and the server going away means nothing
/// after it would have worked either.
enum FetchOutcome {
Failed,
Offline,
}
/// TRACES: FR-CULL-8 | FR-EXP-9
/// Render one original at native resolution and index the faces in it.
///
/// The FR-CULL-8 pipeline end to end, for one photograph: decode, render
/// through the same path export uses, detect on a reduction, crop from the
/// native frame. Returns the faces, the native long edge that went into the
/// run marker, and the 1024px proxy the People screen later cuts thumbnails
/// from.
///
/// # Why the stored edit is not applied
///
/// `export::render_from_library` fetches the sidecar and applies it, because
/// an export is of the photograph the user has made. This is not: the face
/// geometry stored in the catalog is normalised to the frame, so applying a
/// crop would record faces against a frame that changes whenever the user
/// changes their mind, and every stored box would silently become wrong. What
/// is applied is the orientation, which is a fact about the file rather than
/// an edit.
fn index_one_native(
gpu: &dr_gpu::GpuContext,
detector: &mut dr_face::Detector,
embedder: &mut dr_face::Embedder,
bytes: &[u8],
options: &dr_face::DetectOptions,
) -> Result<(Vec<dr_catalog::faces::DetectedFace>, u32, Option<dr_thumbs::Thumbnail>), String> {
let mut session = open_native(gpu, bytes)?;
let frame = session.render_for_export(dr_types::ColourSpace::Srgb)?;
let edge = frame.width.max(frame.height);
let faces = crate::faces::index_native(
detector,
embedder,
&frame.rgba,
frame.width as usize,
frame.height as usize,
options,
)
.map_err(|e| e.to_string())?;
// Only where there is a face to cut out of it. Two thirds of a personal
// library is landscapes and documents, and those never need a crop -- so
// this fills the large class for the images the People screen will
// actually ask about and leaves the rest alone.
//
// Rendered rather than downscaled from the frame in hand: the session is
// still open and `render_thumbnail` is the path the grid's own thumbnails
// take, so the proxy this writes is the one the store would have had
// anyway.
let proxy = if faces.is_empty() {
None
} else {
match session.render_thumbnail(dr_thumbs::ThumbSize::Large.edge()) {
Ok((w, h, rgba)) => match dr_thumbs::encode_rgba(w, h, &rgba) {
Ok(bytes) => Some(dr_thumbs::Thumbnail {
width: w,
height: h,
bytes,
}),
Err(e) => {
log::debug!("encoding a face proxy: {e}");
None
}
},
Err(e) => {
log::debug!("rendering a face proxy: {e}");
None
}
}
};
Ok((faces, edge, proxy))
}
/// The class the whole-library pass fills.
///
/// Grid only, deliberately. The large class is four times the transfer for a