Split library.rs into library/ by area of behaviour

library.rs was 7,729 lines wiring together everything "open a remote
library" touches: scanning, pulling other devices' judgements out of
sidecars found along the way, writing local edits back out to the
sidecar outbox, pushing/reloading XMP by hand, fetching and prefetching
thumbnails and originals, generating thumbnails locally, the metadata
and thumbnail background sweeps, on-disk paths for the catalog and
model files, and reading the grid's cells, spans and rating filter.
Same motivation as the develop.rs split (docs/dev/code-health.md CH-1):
a pure, no-behaviour-change move into one file per area, each under
about 1,500 lines.

Tracing actual call sites rather than trusting the file's physical
layout mattered here: `persist`, `load_folder_etags`, `pull_sidecars`,
`load_sidecar_etags`, `record_sidecar_read` and `apply_judgement` sit
textually beside the XMP push/reload functions but are called only
from `run_scan` (pulling a device's own past judgements out of the
sidecars a scan just walked), so they went to scan.rs and not xmp.rs.
`cells` came out at over 1,800 lines once its tests moved with it and
split further into cells.rs (windowed reads, trash, ordinals) and
spans.rs (collection scope, manual reordering, the capture-time
histogram) -- ten submodules rather than the nine first planned.

Previously-private items reached from a sibling module became
`pub(super)`, narrower than the whole-crate reachability one file gave
them. Tests moved with the code they test; the two test fixtures used
across more than one file (`scanned`, and develop.rs's
`session_with_a_left_half_subject` in the matching commit) joined the
shared `test_support` module alongside the existing `entry`/
`with_images`/`image_ids` helpers. `mod.rs` re-exports every module's
public items under `library::`, including the `pub(crate)`
`test_support` module `repairs.rs` reads its fixtures from, so no file
outside `library` needed a change.

The previous commit split develop.rs the same way; taken alone it left
dr-ui without library.rs, so that intermediate commit does not build on
its own. This one restores it.
This commit is contained in:
2026-09-20 18:21:43 +02:00
parent 050c2c9d16
commit a1d511fd4b
12 changed files with 7895 additions and 7729 deletions
+616
View File
@@ -0,0 +1,616 @@
//! Generating thumbnails locally from a decoded preview or original, and
//! the metadata that comes along for the ride.
#[cfg(test)]
use dr_catalog::Catalog;
use dr_sync::{Connection, RemoteBackend, RemoteId, RemotePath};
use dr_thumbs::ThumbStore;
use std::path::PathBuf;
use std::sync::mpsc::Receiver;
use super::sweep::{flush_metadata, read_metadata_only, MetadataFound};
#[cfg(test)]
use super::sweep::{thumbnails_outstanding, SWEEP_THUMB_SIZE};
use super::thumbnails_fetch::{ThumbnailRequest, MAX_PREVIEW_BYTES};
/// Longest edge face indexing works at.
///
/// Not the full preview: `index_proxy` needs packed `f32` RGB, which is 12
/// bytes a pixel, so a 24 MP frame would be ~288 MB and the fetch lanes hold
/// one each. 3072 costs ~75 MB at the same moment and still puts a face 2%
/// across the frame at ~61 source pixels, against 5 on a grid thumbnail.
///
/// Raise it if the embedder is ever given a larger input than 112: it is the
/// resolution the *crop* is sampled from, so it bounds face quality directly.
pub(super) const FACE_SOURCE_EDGE: u32 = 3072;
/// One decoded thumbnail, ready for the grid.
#[derive(Debug)]
pub struct ThumbnailReady {
/// Index into the grid model this belongs to.
pub row: usize,
pub width: u32,
pub height: u32,
pub rgba: Vec<u8>,
/// Whether these pixels came off local disk rather than the server.
///
/// The grid paints both identically, so this exists solely for
/// reachability: a store hit is evidence about the *cache*, not the
/// network, and treating one as proof of connectivity clears offline mode
/// before a single request has been attempted.
pub from_cache: bool,
}
/// Messages from the thumbnail worker.
#[derive(Debug)]
pub enum ThumbnailMessage {
Ready(Box<ThumbnailReady>),
/// No preview could be extracted. The cell stays a placeholder rather than
/// silently retrying forever.
Unavailable {
row: usize,
reason: String,
},
/// How the batch split between the store and the network.
///
/// Sent once, before any fetch. Without it there is no way to tell a
/// working cache from a broken one — both fill the grid, one just costs
/// nothing.
Plan {
cached: usize,
fetching: usize,
dating: usize,
},
/// One header-only date read is starting.
///
/// Reported separately from thumbnail progress: this work produces no
/// visible cell, so without it the window looks idle while it runs.
DateProgress,
/// Capture dates were written to the catalog.
///
/// The timeline is rebuilt on this rather than per image — a histogram
/// that redrew 120 times during a batch would flicker for no benefit.
DatesRecorded(usize),
/// TRACES: FR-CAT-9
/// The server could not be reached while filling this batch.
///
/// Distinct from a run of [`Unavailable`](Self::Unavailable): those are
/// per-image verdicts ("this file has no extractable preview") and leave
/// the rest of the library alone, where this is a statement about the
/// connection. Sent at most once per batch, because a dropped connection
/// produces one of these per *cell* otherwise and the banner would be
/// rewritten sixty times.
Offline {
reason: String,
},
}
/// Serve thumbnails for a set of rows: store first, network second.
///
/// The store is consulted before any request goes out, so a second launch —
/// or a second device that synced the shards — fills the grid with no transfer
/// at all. Only genuine misses reach the network.
pub fn spawn_thumbnails(
conn: Connection,
wanted: Vec<ThumbnailRequest>,
store_dir: PathBuf,
catalog_path: PathBuf,
) -> Receiver<ThumbnailMessage> {
let (tx, rx) = std::sync::mpsc::channel();
std::thread::spawn(move || {
let mut store = match ThumbStore::open(&store_dir) {
Ok(s) => Some(s),
Err(e) => {
// A broken store costs speed, never correctness — every
// thumbnail can still be fetched.
log::warn!("thumbnail store unavailable, fetching everything: {e}");
None
}
};
// Split the batch before delivering anything, so the plan can be
// reported first and the UI knows the shape of the work up front.
// Decoding happens here rather than in the split, because a corrupt
// blob turns a hit into a miss.
let mut hits = Vec::new();
let mut to_fetch = Vec::new();
// Images whose thumbnail is cached but whose date is still unknown.
//
// These need a header read even though no pixels are wanted. Without
// this pass an image is dated *only* on the one visit that produced
// its thumbnail — so a library browsed once before the EXIF code
// existed, or synced from another device's shards, stays permanently
// undated and never appears on the timeline.
let mut metadata_only = Vec::new();
for req in wanted {
let stored = req
.file_id
.zip(store.as_ref())
.and_then(|(id, s)| s.get(id, req.thumb_size).ok().flatten());
match stored.map(|t| dr_thumbs::decode_rgba(&t.bytes)) {
Some(Ok((width, height, rgba))) => {
if req.needs_metadata {
metadata_only.push(req.clone());
}
hits.push(ThumbnailReady {
row: req.row,
width,
height,
rgba,
from_cache: true,
});
}
// A corrupt stored blob is a miss, not a failure.
Some(Err(e)) => {
log::debug!("stored thumbnail unreadable, refetching: {e}");
to_fetch.push(req);
}
None => to_fetch.push(req),
}
}
log::info!(
"thumbnails: {} from store, {} to fetch{}",
hits.len(),
to_fetch.len(),
if metadata_only.is_empty() {
String::new()
} else {
format!(" · {} dates to read", metadata_only.len())
}
);
if tx
.send(ThumbnailMessage::Plan {
cached: hits.len(),
fetching: to_fetch.len(),
dating: metadata_only.len(),
})
.is_err()
{
return;
}
for hit in hits {
if tx.send(ThumbnailMessage::Ready(Box::new(hit))).is_err() {
return;
}
}
if to_fetch.is_empty() && metadata_only.is_empty() {
return;
}
let rt = match crate::net_runtime::build() {
Ok(rt) => rt,
Err(e) => {
for req in &to_fetch {
let _ = tx.send(ThumbnailMessage::Unavailable {
row: req.row,
reason: e.to_string(),
});
}
return;
}
};
rt.block_on(async {
let backend = match crate::remote::connect(&conn) {
Ok(b) => b,
Err(e) => {
for req in &to_fetch {
let _ = tx.send(ThumbnailMessage::Unavailable {
row: req.row,
reason: e.to_string(),
});
}
return;
}
};
// Batched rather than written per image: one transaction per
// batch instead of 120, and the grid does not need each date the
// instant it is read.
let mut found = Vec::new();
// Set when the server proves unreachable, which abandons the rest
// of the batch. The remaining cells would each take a full timeout
// to reach the same conclusion — on a 120-cell window, minutes of
// the grid appearing to load against a server that is not there.
let mut offline = false;
for req in to_fetch {
let msg = fetch_one(&*backend, store.as_mut(), &req, &mut found).await;
offline = matches!(msg, ThumbnailMessage::Offline { .. });
// A closed channel means the window went away mid-fetch.
if tx.send(msg).is_err() || offline {
break;
}
}
// Dates for images whose pixels were already cached. Header only —
// no preview range, no decode.
//
// Flushed in chunks rather than once at the end: 119 sequential
// header fetches take tens of seconds, and a single write at the
// finish loses every one of them if the window closes first. It
// also lets the timeline appear while the rest are still arriving.
//
// Skipped entirely when the connection has already failed: these
// are network reads too, and there is nothing left to read from.
const FLUSH_EVERY: usize = 16;
if !offline {
log::info!("reading dates for {} image(s)", metadata_only.len());
for req in metadata_only {
if tx.send(ThumbnailMessage::DateProgress).is_err() {
break;
}
read_metadata_only(&*backend, &req, &mut found).await;
if found.len() >= FLUSH_EVERY {
flush_metadata(&catalog_path, &mut found, &tx);
}
}
}
// Always flushed, even when the batch was abandoned: whatever was
// read before the connection died is still true, and discarding it
// would mean re-fetching those headers next time.
flush_metadata(&catalog_path, &mut found, &tx);
});
});
rx
}
pub(super) async fn fetch_one(
backend: &dyn RemoteBackend,
store: Option<&mut ThumbStore>,
req: &ThumbnailRequest,
found_metadata: &mut Vec<MetadataFound>,
) -> ThumbnailMessage {
let preview = match fetch_preview(backend, req, found_metadata).await {
PreviewOutcome::Ready(p) => p,
PreviewOutcome::Unavailable(reason) => {
return ThumbnailMessage::Unavailable {
row: req.row,
reason,
}
}
PreviewOutcome::Offline(reason) => return ThumbnailMessage::Offline { reason },
};
// Persist for next time, and for every other client that syncs the shard.
// A store failure is logged and dropped: the pixels are already in hand,
// and refusing to display them because they could not be cached would be
// the wrong trade.
if let (Some(store), Some(file_id)) = (store, req.file_id) {
if let Some(thumb) = encode_preview(file_id, &preview) {
store_thumbnail(store, file_id, req.thumb_size, &thumb);
}
}
ThumbnailMessage::Ready(Box::new(ThumbnailReady {
row: req.row,
width: preview.width,
height: preview.height,
rgba: preview.rgba,
from_cache: false,
}))
}
/// What one fetch produced.
///
/// Separate from [`ThumbnailMessage`] because not every caller has a grid row
/// to report against or a store to write through. The whole-library pass
/// ([`spawn_thumbnail_sweep`]) fetches on several lanes at once and stores the
/// results on the one thread that owns the store, so it needs the pixels
/// *before* anything is written or addressed to a cell.
pub(super) enum PreviewOutcome {
Ready(dr_decode::Preview),
/// This image has no usable preview. The batch continues past it.
Unavailable(String),
/// The server is unreachable, so nothing after this would succeed either.
Offline(String),
}
/// Fetch a preview in two stages: header, then the exact preview range.
///
/// This is what FR-NC-3 specifies, and the single-stage version it replaces
/// was wrong in a way that looked like corruption: fetching a fixed prefix cut
/// the embedded JPEG partway through, and decoders render a truncated JPEG as
/// the top fraction of the frame rather than reporting an error.
pub(super) async fn fetch_preview(
backend: &dyn RemoteBackend,
req: &ThumbnailRequest,
found_metadata: &mut Vec<MetadataFound>,
) -> PreviewOutcome {
let id = RemoteId::Path(RemotePath::new(&req.path));
// A connection failure is not this image's verdict. Reported as such so
// the caller can stop the batch rather than marking sixty cells
// individually unpreviewable over one dropped connection — a state the
// grid would then keep until something forced a reload.
let classify = |e: dr_sync::RemoteError| {
if e.indicates_offline() {
PreviewOutcome::Offline(e.to_string())
} else {
PreviewOutcome::Unavailable(e.to_string())
}
};
// Stage one: the header, enough to parse the container's IFDs.
let header = match backend.get(&id, Some(0..dr_decode::HEADER_BYTES)).await {
Ok(b) => b,
Err(e) => return classify(e),
};
// The same bytes carry EXIF. Reading it here is free — the alternative is
// a second 256 KB fetch per image over the whole library.
if req.needs_metadata {
collect_metadata(backend, &id, &header, req, found_metadata).await;
}
// Read unconditionally, unlike the rest of the EXIF above: `needs_metadata`
// is false once an image has been catalogued, but a thumbnail can still be
// regenerated long after that — a cleared cache, a new size — and a
// thumbnail that came out upright the first time must come out upright
// every time. This is a header walk, not a decode; see `dr_decode::orientation`.
let orientation = dr_decode::orientation(&header).unwrap_or_default();
// A plain JPEG is its own preview; anything else needs locating.
let bytes = if header.starts_with(&[0xFF, 0xD8, 0xFF]) {
match backend.get(&id, None).await {
Ok(b) => b,
Err(e) => return classify(e),
}
} else {
let Some(loc) = dr_decode::locate_preview(&header, req.size) else {
// No locatable preview. Declining beats fetching the whole file:
// that is the 370 GB path FR-NC-3 exists to avoid.
return PreviewOutcome::Unavailable("no locatable embedded preview".into());
};
if loc.len() > MAX_PREVIEW_BYTES {
return PreviewOutcome::Unavailable(format!(
"preview is {} bytes, too large",
loc.len()
));
}
// Stage two: exactly the preview's bytes.
match backend.get(&id, Some(loc.range.clone())).await {
Ok(b) => b,
Err(e) => return classify(e),
}
};
// Verify before decoding. A truncated JPEG decodes "successfully" into a
// partial frame, so without this the broken result reaches the cache and
// the screen looking like a corrupt file.
if !dr_decode::is_complete_jpeg(&bytes) {
return PreviewOutcome::Unavailable("preview bytes are incomplete".into());
}
// Decode on the worker, never the UI thread.
let mut preview = match dr_decode::decode_jpeg(&bytes) {
Ok(p) => p,
Err(e) => return PreviewOutcome::Unavailable(e.to_string()),
};
// Face indexing keeps the detail; every other caller is filling a cell of a
// known size and the full preview is waste from here on.
preview.downscale_to(if req.full_resolution {
FACE_SOURCE_EDGE
} else {
req.thumb_size.edge()
});
// Turn it the right way up before it is measured, cached or shown. An
// embedded preview is written in the sensor's orientation, so without this
// every frame shot in portrait lies on its side in the grid — and, because
// the cache is keyed by file and size alone, stays that way.
//
// After the downscale, so the permutation moves thumbnail-sized bytes
// rather than the full preview's.
//
// **Doing it here is what keeps face geometry honest.** Detection runs on
// whatever this returns, so returning the photograph rather than the sensor
// means every box and landmark is already in the space the catalog stores
// and the develop overlay draws — no second mapping to get backwards, which
// is the one orientation bug this codebase keeps having. It costs a
// permutation of a larger buffer for the face path; that is ~15 ms against
// a decode of ~150 ms, and it buys the whole class of bug.
preview.apply_orientation(orientation);
PreviewOutcome::Ready(preview)
}
/// Compress a decoded preview to what the store holds.
///
/// Split from the write so the whole-library pass can do it on the lane that
/// fetched the image: encoding is the one part of storing a thumbnail that
/// costs CPU rather than the store's lock, and it turns a 256 KB RGBA buffer
/// into ~20 KB before the chunk is handed to the single thread that owns the
/// store.
pub(super) fn encode_preview(
file_id: u64,
preview: &dr_decode::Preview,
) -> Option<dr_thumbs::Thumbnail> {
match dr_thumbs::encode_rgba(preview.width, preview.height, &preview.rgba) {
Ok(bytes) => Some(dr_thumbs::Thumbnail {
width: preview.width,
height: preview.height,
bytes,
}),
Err(e) => {
log::debug!("encoding thumbnail {file_id}: {e}");
None
}
}
}
/// Put a thumbnail in the store, logging rather than failing.
///
/// A store failure costs a re-fetch next time and nothing else — the pixels
/// are already in hand, and the caller has something to show or count either
/// way (ARCH §6.12: the store is derived, never authoritative).
pub(crate) fn store_thumbnail(
store: &mut ThumbStore,
file_id: u64,
size: dr_thumbs::ThumbSize,
thumb: &dr_thumbs::Thumbnail,
) -> bool {
match store.put(file_id, size, thumb) {
Ok(_) => true,
Err(e) => {
log::debug!("storing thumbnail {file_id}: {e}");
false
}
}
}
/// Parse EXIF out of a header and record it.
///
/// Shared by both paths — the thumbnail fetch, which gets the header anyway,
/// and the header-only pass for images whose pixels were already cached.
pub(super) async fn collect_metadata(
backend: &dyn RemoteBackend,
id: &RemoteId,
header: &[u8],
req: &ThumbnailRequest,
out: &mut Vec<MetadataFound>,
) {
let md = match dr_decode::metadata(header) {
Ok(md) => md,
Err(first) => {
// A file whose IFDs follow its pixels — the linear DNG a merge
// writes — has nothing for the decoder in its head but a
// pointer. Its structure is a few kilobytes at the end; fetch
// that and read the two ranges together, rather than leave the
// composite undated at the end of the grid.
let Some(at) = dr_decode::trailing_ifd(header).filter(|at| *at < req.size) else {
return;
};
match backend.get(id, Some(at..req.size)).await {
Ok(tail) => match dr_decode::metadata_split(header, &tail, at) {
Ok(md) => md,
Err(e) => {
log::debug!("metadata: {}: head {first}; head and tail {e}", req.path);
return;
}
},
Err(e) => {
log::debug!("metadata: {}: tail not fetched: {e}", req.path);
return;
}
}
}
};
out.push(MetadataFound {
image_id: req.image_id,
captured_at: md.captured_at,
captured_offset: md.captured_offset,
camera: camera_label(md.make.as_deref(), md.model.as_deref()),
lens: md.lens.map(|l| l.trim().to_string()),
iso: md.iso,
});
}
/// TRACES: FR-CAT-11
/// The camera string the catalog stores, from an EXIF make and model.
///
/// One definition rather than one per caller, because an import's duplicate
/// check compares against what a scan wrote (`dr_catalog::dedup`). Two
/// spellings of the same body would not fail loudly — they would silently
/// disable the cheap tier, and every re-inserted card would transfer in full
/// before the digest caught it.
pub fn camera_label(make: Option<&str>, model: Option<&str>) -> Option<String> {
match (make, model) {
// Bodies repeat the make inside the model ("Canon EOS 6D"), so
// joining unconditionally yields "Canon Canon EOS 6D".
(Some(make), Some(model)) if model.starts_with(make) => Some(model.trim().to_string()),
(Some(make), Some(model)) => Some(format!("{} {}", make.trim(), model.trim())),
(None, Some(model)) => Some(model.trim().to_string()),
_ => None,
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn the_thumbnail_pass_asks_only_for_what_is_missing() {
// The work list is the whole point of the pass being resumable and of
// it being safe to press twice: it is derived from what the store
// lacks, not from a flag in the catalog. Three things it must respect
// — a thumbnail already stored, a trashed or shadowed image, and an
// image with no `oc:fileid`, which the store cannot key on at all.
let catalog = Catalog::in_memory().unwrap();
let c = catalog.connection();
c.execute(
"INSERT INTO roots(id, kind, label) VALUES (1, 'remote', 'lib')",
[],
)
.unwrap();
// (id, name, file_id, shadowed_by, trashed_at)
for (id, name, file_id, shadow, trashed) in [
(1i64, "a.CR2", Some(11i64), None, None),
(2, "b.CR2", Some(22), None, None),
(3, "b.JPG", Some(33), Some(2i64), None),
(4, "c.CR2", Some(44), None, Some(1000i64)),
// Scanned without a file id: nothing to key the store on.
(5, "d.CR2", None, None, None),
] {
c.execute(
"INSERT INTO images(id, root_id, source_ref, shadowed_by, trashed_at, added_at)
VALUES (?1, 1, ?2, ?3, ?4, 0)",
rusqlite::params![id, name, shadow, trashed],
)
.unwrap();
if let Some(file_id) = file_id {
c.execute(
"INSERT INTO remote(image_id, file_id) VALUES (?1, ?2)",
rusqlite::params![id, file_id],
)
.unwrap();
}
}
let dir = std::env::temp_dir().join(format!("dr-thumb-sweep-{}", std::process::id()));
let _ = std::fs::remove_dir_all(&dir);
std::fs::create_dir_all(&dir).unwrap();
let mut store = ThumbStore::open(&dir).unwrap();
let all = thumbnails_outstanding(&catalog, &store).unwrap();
let names: Vec<&str> = all.iter().map(|r| r.path.as_str()).collect();
assert_eq!(
names,
vec!["a.CR2", "b.CR2"],
"shadowed, trashed and file-id-less images are not work this pass can do"
);
// Store one, and it drops out — this is what stops a second run
// re-fetching an hour of previews.
store
.put(
11,
SWEEP_THUMB_SIZE,
&dr_thumbs::Thumbnail {
width: 4,
height: 4,
bytes: vec![0xFF, 0xD8, 0xFF, 0xD9],
},
)
.unwrap();
let rest = thumbnails_outstanding(&catalog, &store).unwrap();
let names: Vec<&str> = rest.iter().map(|r| r.path.as_str()).collect();
assert_eq!(names, vec!["b.CR2"]);
// The large class is a different key, so filling the grid class does
// not make the pass think the library is done at another size.
assert!(!store.contains(11, dr_thumbs::ThumbSize::Large));
let _ = std::fs::remove_dir_all(&dir);
}
}