Index the whole library by fetching what it has not seen

"Index faces in the whole library" could not. Its work list was intersected
with the thumbnail store at `ThumbSize::Large`, and nothing fills that class for
a whole library — `SWEEP_THUMB_SIZE` is deliberately `Grid`, because the large
class is ~860 MB of shards against ~200 MB and every syncing device pays it. So
the only images with a large proxy were the ones the user had personally zoomed
into or opened in the loupe. On this library that was 220 of 23,529.

The comment defending it misread the requirement:

    // Requesting one here would put face indexing on the network path,
    // which FR-CULL-8 explicitly keeps it off.

FR-CULL-8 keeps indexing off the **full decode**, not the network, and then says
the opposite in the same paragraph: "where no proxy exists, the job requests one
at background priority rather than decoding inline". faces.md §7 repeats it.
Neither was implemented.

So the pass fetches. Same two-stage route the thumbnail sweep uses — the header,
then the located preview's own byte range (FR-NC-3) — so no whole file is pulled
and no RAW is decoded, because an embedded preview is a JPEG. The work list is
now every visible image with no `face_index` row for the model: 23,308 here,
against nearly none before.

**It indexes at the resolution the preview actually has**, not the 1024 the old
tier would have given. `locate_preview` already picks the largest embedded
preview, and the thumbnail sweep was decoding it and throwing the detail away at
`downscale_to(256)`. A face 2% across the frame is 5 px on a grid thumbnail and
~61 px at the cap here — and 112 is what the embedder samples, so this is the
difference between an upsampled crop and a real one. `crop_px` records which,
per face, as §7 intended.

Capped at 3072 rather than truly full: `index_proxy` needs packed `f32` RGB at
12 bytes a pixel, so a 24 MP frame is ~288 MB and the fetch lanes hold one each.
The constant is named and sits next to the reason.

Orientation is applied **before** detection, not after downscaling. That costs a
permutation of a larger buffer — ~15 ms against a ~150 ms decode — and buys the
entire class of bug this codebase keeps having: detection then runs on the
photograph rather than the sensor, so every box and landmark is already in the
space the catalog stores and the overlay draws, with no second mapping to get
backwards.

One detector and one embedder serve every lane. The lanes are concurrent futures
on a single thread, not threads, and inference contains no await, so a `RefCell`
borrow never overlaps another — a pair per lane would duplicate ~16 MB of
weights for no parallelism.

Images with no face in them are recorded too. `face_index` records that
detection *ran*, and zero is its most valuable value: without the row every
landscape and document scan returns on every pass, for ever, and in a personal
library that is most of it (§7a).

The old store-only pass survives as `spawn_store_face_sweep` for
`examples/face_index.rs`, which indexes a local store with no network. The
settings copy no longer claims indexing reads "the photographs already
thumbnailed above", and the audit line says "to fetch" rather than "awaiting a
proxy", which had become a blocker that no longer blocks.

Verified against the real catalog: the new work list returns 23,308 where the
old one returned effectively nothing. 469 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-27 18:51:53 +02:00
co-authored by Claude Opus 5
parent 57c0cc0d35
commit c8c6368542
8 changed files with 500 additions and 64 deletions
+374 -1
View File
@@ -39,6 +39,17 @@ use dr_types::FormatFilter;
/// thumbnail needs nothing like that much detail.
const MAX_PREVIEW_BYTES: u64 = 8 * 1024 * 1024;
/// Longest edge face indexing works at.
///
/// Not the full preview: `index_proxy` needs packed `f32` RGB, which is 12
/// bytes a pixel, so a 24 MP frame would be ~288 MB and the fetch lanes hold
/// one each. 3072 costs ~75 MB at the same moment and still puts a face 2%
/// across the frame at ~61 source pixels, against 5 on a grid thumbnail.
///
/// Raise it if the embedder is ever given a larger input than 112: it is the
/// resolution the *crop* is sampled from, so it bounds face quality directly.
const FACE_SOURCE_EDGE: u32 = 3072;
/// Progress and results from the scan worker.
#[derive(Debug)]
pub enum ScanMessage {
@@ -1290,6 +1301,15 @@ pub struct ThumbnailRequest {
/// Whether this image still needs its EXIF read. Where false the header is
/// still fetched — the preview needs it — but nothing is parsed or written.
pub needs_metadata: bool,
/// Keep the preview at the resolution it was decoded at, ignoring
/// `thumb_size`.
///
/// For face indexing, which wants the pixels a thumbnail throws away: a
/// face 2% across the frame is 5 px on a grid thumbnail and 120 px on the
/// embedded preview, and 112 is what the embedder samples. Capped by
/// [`FACE_SOURCE_EDGE`] rather than truly unbounded, because a 24 MP buffer
/// converted to `f32` RGB is ~288 MB and several lanes hold one at once.
pub full_resolution: bool,
}
/// Why a full fetch failed, keeping the one bit the UI cannot re-derive.
@@ -2053,7 +2073,13 @@ async fn fetch_preview(
Ok(p) => p,
Err(e) => return PreviewOutcome::Unavailable(e.to_string()),
};
preview.downscale_to(req.thumb_size.edge());
// Face indexing keeps the detail; every other caller is filling a cell of a
// known size and the full preview is waste from here on.
preview.downscale_to(if req.full_resolution {
FACE_SOURCE_EDGE
} else {
req.thumb_size.edge()
});
// Turn it the right way up before it is measured, cached or shown. An
// embedded preview is written in the sensor's orientation, so without this
// every frame shot in portrait lies on its side in the grid — and, because
@@ -2061,6 +2087,14 @@ async fn fetch_preview(
//
// After the downscale, so the permutation moves thumbnail-sized bytes
// rather than the full preview's.
//
// **Doing it here is what keeps face geometry honest.** Detection runs on
// whatever this returns, so returning the photograph rather than the sensor
// means every box and landmark is already in the space the catalog stores
// and the develop overlay draws — no second mapping to get backwards, which
// is the one orientation bug this codebase keeps having. It costs a
// permutation of a larger buffer for the face path; that is ~15 ms against
// a decode of ~150 ms, and it buys the whole class of bug.
preview.apply_orientation(orientation);
PreviewOutcome::Ready(preview)
@@ -2565,6 +2599,7 @@ fn next_outstanding(
file_id: r.get::<_, Option<i64>>(2)?.map(|v| v as u64),
size: r.get::<_, Option<i64>>(3)?.unwrap_or(0) as u64,
needs_metadata: true,
full_resolution: false,
})
})?
.collect::<Result<Vec<_>, _>>()?;
@@ -2602,6 +2637,305 @@ fn flush_sweep(catalog: &Catalog, found: &mut Vec<MetadataFound>) {
}
/// TRACES: FR-CAT-3 | FR-NC-3 | NFR-RES-4
/// TRACES: FR-CULL-8
/// Every visible image this model has not been run over.
///
/// **No thumbnail-store filter.** The pass this feeds fetches its own pixels,
/// so an image with no proxy is work to be done rather than work to be skipped
/// — which is the whole difference between indexing a library and indexing the
/// fraction of it that has been browsed.
fn faces_unindexed(
catalog: &Catalog,
model_id: &str,
) -> Result<Vec<ThumbnailRequest>, dr_catalog::CatalogError> {
let mut stmt = catalog.connection().prepare(&format!(
"SELECT i.id, i.source_ref, r.file_id, i.file_size
FROM images i
JOIN remote r ON r.image_id = i.id
WHERE r.file_id IS NOT NULL AND {VISIBLE}
AND NOT EXISTS (
SELECT 1 FROM face_index fi
WHERE fi.image_id = i.id AND fi.model_id = ?1
)
ORDER BY i.id"
))?;
let rows = stmt
.query_map([model_id], |r| {
Ok(ThumbnailRequest {
// Face indexing wants the detail a thumbnail discards.
full_resolution: true,
// Named because the field must say something; ignored, because
// `full_resolution` overrides it.
thumb_size: dr_thumbs::ThumbSize::Large,
// No grid cell waits on this.
row: 0,
image_id: r.get(0)?,
path: r.get(1)?,
file_id: r.get::<_, Option<i64>>(2)?.map(|v| v as u64),
size: r.get::<_, Option<i64>>(3)?.unwrap_or(0) as u64,
// The thumbnail sweep owns dating. Reading EXIF here would
// write the same rows from a second pass for no gain.
needs_metadata: false,
})
})?
.collect::<Result<Vec<_>, _>>()?;
Ok(rows)
}
/// TRACES: FR-CULL-8 | NFR-ARCH-2
/// Index faces across the **whole** library, fetching what it needs.
///
/// # Why this replaced a pass that read the thumbnail store
///
/// The previous sweep filtered its work list down to images that already had a
/// `ThumbSize::Large` proxy on disk, on the reading that FR-CULL-8 keeps face
/// indexing off the network. It does not: it keeps indexing off the *full
/// decode*, and says plainly that "where no proxy exists, the job requests one
/// at background priority". Nothing filled the large class for a whole library
/// — `SWEEP_THUMB_SIZE` is deliberately `Grid` — so the pass could only ever
/// reach photographs the user had personally zoomed into. Measured on the
/// reference library: 220 images indexed out of 23,529.
///
/// So this fetches, by exactly the two-stage route the thumbnail sweep uses:
/// the header, then the located preview's own byte range (FR-NC-3). No full
/// file is pulled and no RAW is decoded — an embedded preview is a JPEG.
///
/// # Why it does not simply ride the thumbnail sweep
///
/// It could, and it would be free: that pass already fetches the largest
/// preview, decodes it, and downscales it to 256. Riding it is the right shape
/// for *new* images and is the obvious next step. It cannot serve this
/// operation, though, because the thumbnail sweep's work list is what the store
/// does not have — so every image already thumbnailed, which for an established
/// library is most of them, would never come back past the detector.
///
/// # Cost, stated plainly
///
/// One preview fetch per un-indexed image, capped at [`MAX_PREVIEW_BYTES`].
/// That is the same transfer the thumbnail sweep pays per image, paid a second
/// time because the first one kept only 256 px. Resumable by construction: the
/// work list is what the catalog has no `face_index` row for, so a kill costs
/// the images in flight and nothing else.
#[allow(clippy::too_many_arguments)]
pub fn spawn_face_sweep(
creds: AppCredentials,
user_id: String,
catalog_path: PathBuf,
detector_model: PathBuf,
embedder_model: PathBuf,
model_id: String,
options: dr_face::DetectOptions,
) -> Receiver<crate::faces::FaceSweepMessage> {
use crate::faces::FaceSweepMessage;
let (tx, rx) = std::sync::mpsc::channel();
std::thread::spawn(move || {
let finish_empty = |tx: &Sender<FaceSweepMessage>| {
let _ = tx.send(FaceSweepMessage::Finished {
images: 0,
faces: 0,
failed: 0,
});
};
let catalog = match Catalog::open(&catalog_path) {
Ok(c) => c,
Err(e) => {
log::warn!("face sweep: cannot open catalog: {e}");
finish_empty(&tx);
return;
}
};
// Models before the work list: they are the expensive failure, and
// listing twenty thousand images before discovering the weights are
// missing helps nobody. A library with no model installed takes this
// path, so it is a quiet return rather than an error.
let mut detector = match dr_face::Detector::from_path(&detector_model) {
Ok(d) => d,
Err(e) => {
log::warn!("face sweep: cannot load the detector: {e}");
finish_empty(&tx);
return;
}
};
let mut embedder = match dr_face::Embedder::from_path(
&embedder_model,
dr_face::ModelId::new(model_id.clone()),
) {
Ok(e) => e,
Err(e) => {
log::warn!("face sweep: cannot load the embedder: {e}");
finish_empty(&tx);
return;
}
};
let wanted = match faces_unindexed(&catalog, &model_id) {
Ok(w) => w,
Err(e) => {
log::warn!("face sweep: {e}");
finish_empty(&tx);
return;
}
};
let total = wanted.len();
if total == 0 {
log::info!("face sweep: every image has been through this model");
finish_empty(&tx);
return;
}
log::info!("face sweep: {total} image(s) to index");
if tx.send(FaceSweepMessage::Total(total)).is_err() {
return;
}
let rt = match crate::net_runtime::build() {
Ok(rt) => rt,
Err(e) => {
log::warn!("face sweep: no runtime: {e}");
finish_empty(&tx);
return;
}
};
rt.block_on(async {
let backend = match NextcloudBackend::new(&creds, &user_id) {
Ok(b) => b,
Err(e) => {
log::warn!("face sweep: {e}");
finish_empty(&tx);
return;
}
};
// One detector and one embedder for every lane.
//
// The lanes are concurrent futures on a single thread, not threads,
// so they interleave only at await points — and inference contains
// none. A `RefCell` borrow therefore never overlaps another, and
// the alternative, a pair per lane, would be ~16 MB of weights
// duplicated for no parallelism at all.
let models = std::cell::RefCell::new((&mut detector, &mut embedder));
let (mut done, mut images, mut found, mut failed) = (0usize, 0usize, 0usize, 0usize);
let mut offline = false;
for chunk in wanted.chunks(SWEEP_CHUNK) {
let lanes: Vec<Vec<&ThumbnailRequest>> = (0..SWEEP_LANES)
.map(|lane| chunk.iter().skip(lane).step_by(SWEEP_LANES).collect())
.collect();
let results = futures_join_all(lanes.into_iter().map(|lane| {
let backend = &backend;
let models = &models;
let options = &options;
async move {
let mut indexed: Vec<(i64, Vec<dr_catalog::faces::DetectedFace>, u32)> =
Vec::new();
let mut discard = Vec::new();
let mut attempted = 0usize;
let mut failed = 0usize;
let mut offline = false;
for req in lane {
attempted += 1;
match fetch_preview(backend, req, &mut discard).await {
PreviewOutcome::Ready(preview) => {
// No await inside this borrow — see the
// note where `models` is built.
let mut m = models.borrow_mut();
let (det, emb) = &mut *m;
match crate::faces::index_preview(det, emb, &preview, options) {
Ok((faces, edge)) => {
indexed.push((req.image_id, faces, edge))
}
Err(e) => {
log::debug!("face sweep: {}: {e}", req.path);
failed += 1;
}
}
}
PreviewOutcome::Unavailable(reason) => {
log::debug!("face sweep: {}: {reason}", req.path);
failed += 1;
}
PreviewOutcome::Offline(reason) => {
log::info!("face sweep: server unreachable: {reason}");
attempted -= 1;
offline = true;
break;
}
}
}
(indexed, attempted, failed, offline)
}
}))
.await;
for (indexed, attempted, lane_failed, lane_offline) in results {
done += attempted;
failed += lane_failed;
offline |= lane_offline;
for (image_id, faces, edge) in indexed {
// Written per image, including the ones with no face in
// them: `face_index` records that detection *ran*, and
// zero is its most valuable value — without the row,
// every landscape and document scan returns on the next
// pass, for ever (docs/faces.md §7a).
match dr_catalog::faces::record_detections(
catalog.connection(),
dr_types::ImageId(image_id as u64),
&model_id,
edge,
&faces,
) {
Ok(_) => {
images += 1;
found += faces.len();
if tx
.send(FaceSweepMessage::Indexed {
image: dr_types::ImageId(image_id as u64),
faces: faces.len(),
})
.is_err()
{
// Receiver dropped: the screen closed, or
// the user pressed Stop. Everything written
// so far stays written.
log::info!("face sweep: cancelled after {images} image(s)");
return;
}
}
Err(e) => {
log::warn!("face sweep: storing faces for {image_id}: {e}");
failed += 1;
}
}
}
}
let _ = done;
if offline {
break;
}
}
log::info!(
"face sweep: {found} face(s) across {images} image(s), {failed} failed{}",
if offline { ", server went away" } else { "" }
);
let _ = tx.send(FaceSweepMessage::Finished {
images,
faces: found,
failed,
});
});
});
rx
}
/// The class the whole-library pass fills.
///
/// Grid only, deliberately. The large class is four times the transfer for a
@@ -2875,6 +3209,7 @@ fn thumbnails_outstanding(
// The header this fetch reads is the one EXIF lives in, so an
// undated image is dated on the way past for nothing.
needs_metadata: r.get::<_, i64>(4)? < 2,
full_resolution: false,
})
})?
.filter_map(Result::ok)
@@ -4082,6 +4417,44 @@ mod tests {
.collect()
}
/// The regression this whole pass exists for.
///
/// The previous work list intersected with the thumbnail store, so a
/// library nobody had zoomed into produced an empty one and "index the
/// whole library" indexed nothing. Nothing here puts a proxy on disk.
#[test]
fn every_unindexed_image_is_work_even_with_no_proxy_anywhere() {
let catalog = with_images(10);
let wanted = faces_unindexed(&catalog, "w600k_mbf").unwrap();
assert_eq!(
wanted.len(),
10,
"a library with no thumbnails is still work"
);
assert!(
wanted.iter().all(|r| r.full_resolution),
"the pass must keep the detail a thumbnail would throw away"
);
}
/// `face_index` records that detection *ran*, so an image with no face in
/// it must not come back on the next pass — otherwise a personal library,
/// which is mostly landscapes and documents, never finishes.
#[test]
fn an_image_already_run_over_is_not_work_again() {
let catalog = with_images(3);
let ids = image_ids(&catalog);
dr_catalog::faces::record_detections(catalog.connection(), ids[0], "w600k_mbf", 1024, &[])
.unwrap();
let wanted = faces_unindexed(&catalog, "w600k_mbf").unwrap();
assert_eq!(wanted.len(), 2);
assert!(!wanted.iter().any(|r| r.image_id == ids[0].0 as i64));
// A different model has seen none of them.
assert_eq!(faces_unindexed(&catalog, "other").unwrap().len(), 3);
}
#[test]
fn a_scoped_grid_shows_only_that_collections_images() {
use dr_catalog::collections::{self as coll, CollectionKind};