Thumbnail the whole library on request, and send the shards up

The grid fetches a preview only for cells that are actually browsed, which is
the right posture over a link that must not be saturated to show one screen
(FR-NC-3). The cost is that the thumbnail store ends up holding the fraction of
the library someone happened to scroll past — and that store is the one derived
thing worth syncing, since a second device that downloads the shards gets a
full grid without touching a single RAW. So the complete set is worth an hour
of range fetches paid once, deliberately, on a machine that can afford it.

That is what this adds: a pass over every visible image with an `oc:fileid`,
launched from the settings page and reported in the activity register like
every other background job. The work list is what the store lacks rather than a
flag in the catalog, so it is resumable by construction and safe to press
twice.

Lane-parallel like the metadata sweep, but it does what that sweep declined to.
The store is `&mut` and cannot cross lanes, which is why dating the library
skips thumbnails entirely; here the lanes fetch, decode and *encode*, and only
the ~20 KB result crosses back to the one thread that owns the store and writes
the chunk. The parallelism is real and the single-writer rule is not bent.

Grid class only. The large class is ~860 MB of shards against ~200 MB on the
reference library, paid by every device that syncs them; a photograph looked at
closely still gets its large thumbnail from the interactive path.

Dates come free — the header a preview needs is the header EXIF lives in — and
the pass ends by pushing the shards to the server. A filled store that never
leaves this device would be most of the cost for none of the point.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-22 19:34:59 +02:00
co-authored by Claude Opus 5
parent 28325af448
commit d91f1ec277
3 changed files with 681 additions and 42 deletions
+465 -42
View File
@@ -1845,34 +1845,78 @@ pub fn spawn_thumbnails(
rx
}
/// Fetch a preview in two stages: header, then the exact preview range.
///
/// This is what FR-NC-3 specifies, and the single-stage version it replaces
/// was wrong in a way that looked like corruption: fetching a fixed prefix cut
/// the embedded JPEG partway through, and decoders render a truncated JPEG as
/// the top fraction of the frame rather than reporting an error.
async fn fetch_one(
backend: &NextcloudBackend,
store: Option<&mut ThumbStore>,
req: &ThumbnailRequest,
found_metadata: &mut Vec<MetadataFound>,
) -> ThumbnailMessage {
let id = RemoteId::Path(RemotePath::new(&req.path));
let fail = |reason: String| ThumbnailMessage::Unavailable {
row: req.row,
reason,
let preview = match fetch_preview(backend, req, found_metadata).await {
PreviewOutcome::Ready(p) => p,
PreviewOutcome::Unavailable(reason) => {
return ThumbnailMessage::Unavailable {
row: req.row,
reason,
}
}
PreviewOutcome::Offline(reason) => return ThumbnailMessage::Offline { reason },
};
// Persist for next time, and for every other client that syncs the shard.
// A store failure is logged and dropped: the pixels are already in hand,
// and refusing to display them because they could not be cached would be
// the wrong trade.
if let (Some(store), Some(file_id)) = (store, req.file_id) {
if let Some(thumb) = encode_preview(file_id, &preview) {
store_thumbnail(store, file_id, req.thumb_size, &thumb);
}
}
ThumbnailMessage::Ready(Box::new(ThumbnailReady {
row: req.row,
width: preview.width,
height: preview.height,
rgba: preview.rgba,
from_cache: false,
}))
}
/// What one fetch produced.
///
/// Separate from [`ThumbnailMessage`] because not every caller has a grid row
/// to report against or a store to write through. The whole-library pass
/// ([`spawn_thumbnail_sweep`]) fetches on several lanes at once and stores the
/// results on the one thread that owns the store, so it needs the pixels
/// *before* anything is written or addressed to a cell.
enum PreviewOutcome {
Ready(dr_decode::Preview),
/// This image has no usable preview. The batch continues past it.
Unavailable(String),
/// The server is unreachable, so nothing after this would succeed either.
Offline(String),
}
/// Fetch a preview in two stages: header, then the exact preview range.
///
/// This is what FR-NC-3 specifies, and the single-stage version it replaces
/// was wrong in a way that looked like corruption: fetching a fixed prefix cut
/// the embedded JPEG partway through, and decoders render a truncated JPEG as
/// the top fraction of the frame rather than reporting an error.
async fn fetch_preview(
backend: &NextcloudBackend,
req: &ThumbnailRequest,
found_metadata: &mut Vec<MetadataFound>,
) -> PreviewOutcome {
let id = RemoteId::Path(RemotePath::new(&req.path));
// A connection failure is not this image's verdict. Reported as such so
// the caller can stop the batch rather than marking sixty cells
// individually unpreviewable over one dropped connection — a state the
// grid would then keep until something forced a reload.
let classify = |e: dr_sync::RemoteError| {
if e.indicates_offline() {
ThumbnailMessage::Offline {
reason: e.to_string(),
}
PreviewOutcome::Offline(e.to_string())
} else {
fail(e.to_string())
PreviewOutcome::Unavailable(e.to_string())
}
};
@@ -1905,10 +1949,13 @@ async fn fetch_one(
let Some(loc) = dr_decode::locate_preview(&header, req.size) else {
// No locatable preview. Declining beats fetching the whole file:
// that is the 370 GB path FR-NC-3 exists to avoid.
return fail("no locatable embedded preview".into());
return PreviewOutcome::Unavailable("no locatable embedded preview".into());
};
if loc.len() > MAX_PREVIEW_BYTES {
return fail(format!("preview is {} bytes, too large", loc.len()));
return PreviewOutcome::Unavailable(format!(
"preview is {} bytes, too large",
loc.len()
));
}
// Stage two: exactly the preview's bytes.
@@ -1922,13 +1969,13 @@ async fn fetch_one(
// partial frame, so without this the broken result reaches the cache and
// the screen looking like a corrupt file.
if !dr_decode::is_complete_jpeg(&bytes) {
return fail("preview bytes are incomplete".into());
return PreviewOutcome::Unavailable("preview bytes are incomplete".into());
}
// Decode on the worker, never the UI thread.
let mut preview = match dr_decode::decode_jpeg(&bytes) {
Ok(p) => p,
Err(e) => return fail(e.to_string()),
Err(e) => return PreviewOutcome::Unavailable(e.to_string()),
};
preview.downscale_to(req.thumb_size.edge());
// Turn it the right way up before it is measured, cached or shown. An
@@ -1940,33 +1987,48 @@ async fn fetch_one(
// rather than the full preview's.
preview.apply_orientation(orientation);
// Persist for next time, and for every other client that syncs the shard.
// A store failure is logged and dropped: the pixels are already in hand,
// and refusing to display them because they could not be cached would be
// the wrong trade.
if let (Some(store), Some(file_id)) = (store, req.file_id) {
match dr_thumbs::encode_rgba(preview.width, preview.height, &preview.rgba) {
Ok(encoded) => {
let thumb = dr_thumbs::Thumbnail {
width: preview.width,
height: preview.height,
bytes: encoded,
};
if let Err(e) = store.put(file_id, req.thumb_size, &thumb) {
log::debug!("storing thumbnail {file_id}: {e}");
}
}
Err(e) => log::debug!("encoding thumbnail {file_id}: {e}"),
PreviewOutcome::Ready(preview)
}
/// Compress a decoded preview to what the store holds.
///
/// Split from the write so the whole-library pass can do it on the lane that
/// fetched the image: encoding is the one part of storing a thumbnail that
/// costs CPU rather than the store's lock, and it turns a 256 KB RGBA buffer
/// into ~20 KB before the chunk is handed to the single thread that owns the
/// store.
fn encode_preview(file_id: u64, preview: &dr_decode::Preview) -> Option<dr_thumbs::Thumbnail> {
match dr_thumbs::encode_rgba(preview.width, preview.height, &preview.rgba) {
Ok(bytes) => Some(dr_thumbs::Thumbnail {
width: preview.width,
height: preview.height,
bytes,
}),
Err(e) => {
log::debug!("encoding thumbnail {file_id}: {e}");
None
}
}
}
ThumbnailMessage::Ready(Box::new(ThumbnailReady {
row: req.row,
width: preview.width,
height: preview.height,
rgba: preview.rgba,
from_cache: false,
}))
/// Put a thumbnail in the store, logging rather than failing.
///
/// A store failure costs a re-fetch next time and nothing else — the pixels
/// are already in hand, and the caller has something to show or count either
/// way (ARCH §6.12: the store is derived, never authoritative).
fn store_thumbnail(
store: &mut ThumbStore,
file_id: u64,
size: dr_thumbs::ThumbSize,
thumb: &dr_thumbs::Thumbnail,
) -> bool {
match store.put(file_id, size, thumb) {
Ok(_) => true,
Err(e) => {
log::debug!("storing thumbnail {file_id}: {e}");
false
}
}
}
/// Parse EXIF out of a header and record it.
@@ -2463,6 +2525,291 @@ fn flush_sweep(catalog: &Catalog, found: &mut Vec<MetadataFound>) {
found.clear();
}
/// TRACES: FR-CAT-3 | FR-NC-3 | NFR-RES-4
/// The class the whole-library pass fills.
///
/// Grid only, deliberately. The large class is four times the transfer for a
/// detail only a zoomed cell or the loupe asks for — on the reference library
/// that is ~200 MB of shards against ~860 MB, paid by *every* device that
/// syncs them (see [`dr_thumbs::ThumbSize`]). A photograph actually looked at
/// closely still gets its large thumbnail from the interactive path.
const SWEEP_THUMB_SIZE: dr_thumbs::ThumbSize = dr_thumbs::ThumbSize::Grid;
/// Progress from the whole-library thumbnail pass.
#[derive(Debug)]
pub enum ThumbSweepMessage {
/// How many images still lack a thumbnail, counted once at the start.
Total(usize),
/// Another chunk finished. Carries cumulative counts.
Progress { done: usize, stored: usize },
Finished {
stored: usize,
failed: usize,
/// Stopped early because the server stopped answering. The pass is
/// resumable, so this is "come back later", not a failure.
offline: bool,
},
}
/// TRACES: FR-CAT-3 | FR-NC-3 | FR-NC-7
/// Thumbnail **every** image in the library, not just the ones browsed.
///
/// # Why this exists next to the grid's own fetching
///
/// The interactive path fills cells as they are scrolled past, which is the
/// right posture for a remote library (FR-NC-3) and the wrong one for handing
/// the result to a second device: a tablet that syncs the shards inherits only
/// the fraction of the library its sibling happened to look at. This is the
/// deliberate, user-launched version of the same work — an hour of range
/// fetches paid once, on the machine that can afford it, so every other client
/// gets a full grid for the cost of a few hundred MB (`derived_sync`).
///
/// # Shape, and why it borrows the metadata sweep's
///
/// Chunked and lane-parallel exactly as [`spawn_sweep`] is, for the same
/// reason: each image is ~0.6 s of round-trip latency and almost no
/// bandwidth, so the sequential version spends its life waiting. What differs
/// is the store — it is `&mut` and cannot be shared across lanes, which is why
/// the metadata sweep skips thumbnails entirely. Here the lanes fetch, decode
/// and *encode*, and only the ~20 KB result crosses back to this thread, which
/// owns the store and writes the chunk in one go. So the parallelism is real
/// and the single-writer rule is never bent.
///
/// Dates arrive free: the header a preview needs is the header EXIF lives in,
/// so an image this pass reaches is dated on the same fetch rather than
/// costing a second one.
///
/// Resumable by construction — the work list is what the store does not have,
/// so a kill costs the chunk in flight and nothing more. An image with no
/// locatable preview is retried on a later run; it is one header fetch, and
/// the alternative is a second piece of state that has to be invalidated when
/// a file is replaced.
pub fn spawn_thumbnail_sweep(
creds: AppCredentials,
user_id: String,
catalog_path: PathBuf,
store_dir: PathBuf,
) -> Receiver<ThumbSweepMessage> {
let (tx, rx) = std::sync::mpsc::channel();
std::thread::spawn(move || {
let finish_empty = |tx: &Sender<ThumbSweepMessage>| {
let _ = tx.send(ThumbSweepMessage::Finished {
stored: 0,
failed: 0,
offline: false,
});
};
let catalog = match Catalog::open(&catalog_path) {
Ok(c) => c,
Err(e) => {
log::warn!(
"thumbnail sweep: cannot open catalog at {}: {e}",
catalog_path.display()
);
finish_empty(&tx);
return;
}
};
// Unlike the grid's fetch, which carries on without a store and simply
// shows what it downloaded, a store that will not open ends this: the
// pass exists to fill it, and running an hour of transfers with
// nowhere to put them would be worse than not starting.
let mut store = match ThumbStore::open(&store_dir) {
Ok(s) => s,
Err(e) => {
log::warn!(
"thumbnail sweep: cannot open the thumbnail store at {}: {e}",
store_dir.display()
);
finish_empty(&tx);
return;
}
};
let wanted = match thumbnails_outstanding(&catalog, &store) {
Ok(w) => w,
Err(e) => {
log::warn!("thumbnail sweep: {e}");
finish_empty(&tx);
return;
}
};
let total = wanted.len();
if total == 0 {
log::info!("thumbnail sweep: every image already has a thumbnail");
finish_empty(&tx);
return;
}
log::info!("thumbnail sweep: {total} image(s) need a thumbnail");
if tx.send(ThumbSweepMessage::Total(total)).is_err() {
return;
}
let rt = match crate::net_runtime::build() {
Ok(rt) => rt,
Err(e) => {
log::warn!("thumbnail sweep: no runtime: {e}");
finish_empty(&tx);
return;
}
};
rt.block_on(async {
let backend = match NextcloudBackend::new(&creds, &user_id) {
Ok(b) => b,
Err(e) => {
log::warn!("thumbnail sweep: {e}");
finish_empty(&tx);
return;
}
};
let (mut done, mut stored, mut failed) = (0usize, 0usize, 0usize);
let mut offline = false;
let mut found = Vec::new();
for chunk in wanted.chunks(SWEEP_CHUNK) {
// Each lane owns a disjoint slice and its own output, so
// nothing is shared and no lock is needed. The store is not
// touched here — see the note on the function.
let lanes: Vec<Vec<&ThumbnailRequest>> = (0..SWEEP_LANES)
.map(|lane| chunk.iter().skip(lane).step_by(SWEEP_LANES).collect())
.collect();
let results = futures_join_all(lanes.into_iter().map(|lane| {
let backend = &backend;
async move {
let mut made: Vec<(u64, dr_thumbs::Thumbnail)> = Vec::new();
let mut found = Vec::new();
let mut attempted = 0usize;
let mut failed = 0usize;
let mut offline = false;
for req in lane {
// Enforced by the query, which joins `remote`: an
// image with no file id has nothing to key the
// store on and is not a candidate.
let Some(file_id) = req.file_id else { continue };
attempted += 1;
match fetch_preview(backend, req, &mut found).await {
PreviewOutcome::Ready(preview) => {
match encode_preview(file_id, &preview) {
Some(thumb) => made.push((file_id, thumb)),
None => failed += 1,
}
}
PreviewOutcome::Unavailable(reason) => {
log::debug!("thumbnail sweep: {}: {reason}", req.path);
failed += 1;
}
// Nothing after this would reach the server
// either, so the lane stops rather than
// spending a timeout per remaining image.
PreviewOutcome::Offline(reason) => {
log::info!("thumbnail sweep: server unreachable: {reason}");
attempted -= 1;
offline = true;
break;
}
}
}
(made, found, attempted, failed, offline)
}
}))
.await;
for (made, lane_found, attempted, lane_failed, lane_offline) in results {
done += attempted;
failed += lane_failed;
offline |= lane_offline;
found.extend(lane_found);
for (file_id, thumb) in made {
if store_thumbnail(&mut store, file_id, SWEEP_THUMB_SIZE, &thumb) {
stored += 1;
} else {
failed += 1;
}
}
}
// Committed per chunk rather than at the end, so a kill keeps
// every date read so far — the same bargain the metadata sweep
// makes, and for the same reason.
flush_sweep(&catalog, &mut found);
if tx
.send(ThumbSweepMessage::Progress { done, stored })
.is_err()
{
return;
}
if offline {
break;
}
}
flush_sweep(&catalog, &mut found);
log::info!("thumbnail sweep: {stored} stored, {failed} without a usable preview");
let _ = tx.send(ThumbSweepMessage::Finished {
stored,
failed,
offline,
});
});
});
rx
}
/// Every visible image on the server that the store has no grid thumbnail for.
///
/// Joined against `remote` rather than left-joined: the store is keyed on
/// Nextcloud's `oc:fileid` (FR-NC-5), so an image the scan recorded without
/// one cannot be stored and is not work this pass can do.
///
/// The whole list is built up front rather than re-queried per chunk, unlike
/// the metadata sweep: "does the store have this" is answered by the store's
/// index, which this thread is also the one writing, so a stale list is not a
/// risk the way a concurrently-dating grid made it one there.
fn thumbnails_outstanding(
catalog: &Catalog,
store: &ThumbStore,
) -> Result<Vec<ThumbnailRequest>, dr_catalog::CatalogError> {
let mut stmt = catalog.connection().prepare(&format!(
"SELECT i.id, i.source_ref, r.file_id, i.file_size, i.metadata_state
FROM images i
JOIN remote r ON r.image_id = i.id
WHERE r.file_id IS NOT NULL AND {VISIBLE}
ORDER BY i.id"
))?;
let rows = stmt
.query_map([], |r| {
let file_id = r.get::<_, Option<i64>>(2)?.map(|v| v as u64);
Ok(ThumbnailRequest {
thumb_size: SWEEP_THUMB_SIZE,
// No grid cell is waiting on this, so nothing consumes the row.
row: 0,
image_id: r.get(0)?,
path: r.get(1)?,
file_id,
size: r.get::<_, Option<i64>>(3)?.unwrap_or(0) as u64,
// The header this fetch reads is the one EXIF lives in, so an
// undated image is dated on the way past for nothing.
needs_metadata: r.get::<_, i64>(4)? < 2,
})
})?
.filter_map(Result::ok)
.filter(|req| {
req.file_id
.is_some_and(|id| !store.contains(id, SWEEP_THUMB_SIZE))
})
.collect();
Ok(rows)
}
/// Where an account's thumbnail shards live.
///
/// Beside the catalog rather than in the cache directory: these sync to the
@@ -2965,6 +3312,82 @@ mod tests {
}
}
#[test]
fn the_thumbnail_pass_asks_only_for_what_is_missing() {
// The work list is the whole point of the pass being resumable and of
// it being safe to press twice: it is derived from what the store
// lacks, not from a flag in the catalog. Three things it must respect
// — a thumbnail already stored, a trashed or shadowed image, and an
// image with no `oc:fileid`, which the store cannot key on at all.
let catalog = Catalog::in_memory().unwrap();
let c = catalog.connection();
c.execute(
"INSERT INTO roots(id, kind, label) VALUES (1, 'remote', 'lib')",
[],
)
.unwrap();
// (id, name, file_id, shadowed_by, trashed_at)
for (id, name, file_id, shadow, trashed) in [
(1i64, "a.CR2", Some(11i64), None, None),
(2, "b.CR2", Some(22), None, None),
(3, "b.JPG", Some(33), Some(2i64), None),
(4, "c.CR2", Some(44), None, Some(1000i64)),
// Scanned without a file id: nothing to key the store on.
(5, "d.CR2", None, None, None),
] {
c.execute(
"INSERT INTO images(id, root_id, source_ref, shadowed_by, trashed_at, added_at)
VALUES (?1, 1, ?2, ?3, ?4, 0)",
rusqlite::params![id, name, shadow, trashed],
)
.unwrap();
if let Some(file_id) = file_id {
c.execute(
"INSERT INTO remote(image_id, file_id) VALUES (?1, ?2)",
rusqlite::params![id, file_id],
)
.unwrap();
}
}
let dir = std::env::temp_dir().join(format!("dr-thumb-sweep-{}", std::process::id()));
let _ = std::fs::remove_dir_all(&dir);
std::fs::create_dir_all(&dir).unwrap();
let mut store = ThumbStore::open(&dir).unwrap();
let all = thumbnails_outstanding(&catalog, &store).unwrap();
let names: Vec<&str> = all.iter().map(|r| r.path.as_str()).collect();
assert_eq!(
names,
vec!["a.CR2", "b.CR2"],
"shadowed, trashed and file-id-less images are not work this pass can do"
);
// Store one, and it drops out — this is what stops a second run
// re-fetching an hour of previews.
store
.put(
11,
SWEEP_THUMB_SIZE,
&dr_thumbs::Thumbnail {
width: 4,
height: 4,
bytes: vec![0xFF, 0xD8, 0xFF, 0xD9],
},
)
.unwrap();
let rest = thumbnails_outstanding(&catalog, &store).unwrap();
let names: Vec<&str> = rest.iter().map(|r| r.path.as_str()).collect();
assert_eq!(names, vec!["b.CR2"]);
// The large class is a different key, so filling the grid class does
// not make the pass think the library is done at another size.
assert!(!store.contains(11, dr_thumbs::ThumbSize::Large));
let _ = std::fs::remove_dir_all(&dir);
}
#[test]
fn catalog_paths_separate_accounts() {
// Two accounts on one machine must not share an index, or one