Run the library from local data when the server is unreachable

Also carries in-flight work that shared these files: the zoom structure-key
fix in the adjust pipeline, nearest-neighbour filtering past 1:1, the
timeline scrub marker correction, the 423-Locked retry in the metadata
sweep, and the thumbnail size-class migration.

# Offline mode (FR-CAT-9)

The app previously assumed the server was reachable and treated its absence
as a series of unrelated per-operation failures. A launch without a
connection produced an empty grid, even with a complete catalog on disk and
every thumbnail already in the shards.

Reachability is now inferred from traffic the app was already making, rather
than probed for. `RemoteError::indicates_offline` draws the line that makes
this possible: a dead connection is offline, a 403 or a 500 is not — the
server answered, so blanking the library over one forbidden file would be a
worse error than the one being reported. `Reachability` turns those outcomes
into a state, so a library browsing happily never issues a probe at all.

Going offline takes one failure, because the user is already experiencing it.
Coming back requires evidence — a completed scan or a fetched thumbnail —
with a capped exponential backoff behind the manual retry, so twelve sweep
lanes failing together do not schedule twelve immediate probes.

What keeps working: the catalog opens even when the scan that normally
provides it failed, so the grid fills from the last successful scan.
Thumbnails come from the shards. Rating, flagging and collecting are catalog
writes that never touched the network. What stops is opening an original that
was never stored locally, and it now says so in those words instead of
reporting "network error: connection refused" over a photograph.

Work that is pure network is refused rather than left to fail slowly: the
metadata sweep, derived sync, and sidecar writes. The sweep would otherwise
spend a timeout per image across the whole library while the progress bar
implied something was happening. Deferring sidecars is a real gap rather than
a hidden one — a rating made offline reaches its sidecar only when that image
is judged again while connected — and it is recorded as such at the call site.

# The "On this device" filter

A chip beside the rating filters, narrowing the grid to images whose original
is held locally. It composes with the rating terms rather than replacing them,
so "five-star frames I can actually edit on this train" is one filter. The
predicate is SQL, like the rating terms and for the same reason: the count in
the header has to agree with the cells drawn.

It reads `image_cache.tier_actual`, which nothing writes yet — the next
commit fills it. Until then the chip honestly reports zero.

`Tier` gains an explicit on-disk encoding. The variants are ordered by
generosity and the derived `Ord` invites reordering them, which would
silently reinterpret every cached row; the round-trip test is what holds the
two in agreement.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-08-11 20:36:50 +02:00
co-authored by Claude Opus 5
parent f6a100863e
commit cd75e5a4c6
18 changed files with 2604 additions and 109 deletions
+375 -44
View File
@@ -37,9 +37,6 @@ use dr_types::FormatFilter;
/// thumbnail needs nothing like that much detail.
const MAX_PREVIEW_BYTES: u64 = 8 * 1024 * 1024;
/// Long edge of a grid thumbnail, in pixels.
const THUMBNAIL_EDGE: u32 = 256;
/// Progress and results from the scan worker.
#[derive(Debug)]
pub enum ScanMessage {
@@ -62,7 +59,16 @@ pub enum ScanMessage {
pruned: usize,
elapsed_ms: u64,
},
Failed(String),
/// The scan could not finish.
///
/// `offline` distinguishes "the server could not be reached" from "the
/// server refused", and it is carried here rather than re-derived because
/// the classification is only possible on the worker side: crossing the
/// channel flattens a [`dr_sync::RemoteError`] into a message, and no
/// amount of string matching on the far side can reliably recover it.
/// Without the flag a dead connection and a bad password produce the same
/// banner, which sends the user to re-enter a credential that was fine.
Failed { message: String, offline: bool },
}
/// One decoded thumbnail, ready for the grid.
@@ -117,6 +123,16 @@ pub enum ThumbnailMessage {
/// The timeline is rebuilt on this rather than per image — a histogram
/// that redrew 120 times during a batch would flicker for no benefit.
DatesRecorded(usize),
/// TRACES: FR-CAT-9
/// The server could not be reached while filling this batch.
///
/// Distinct from a run of [`Unavailable`](Self::Unavailable): those are
/// per-image verdicts ("this file has no extractable preview") and leave
/// the rest of the library alone, where this is a statement about the
/// connection. Sent at most once per batch, because a dropped connection
/// produces one of these per *cell* otherwise and the banner would be
/// rewritten sixty times.
Offline { reason: String },
}
/// TRACES: FR-CAT-15 | FR-CAT-11
@@ -163,12 +179,22 @@ pub struct RatingFilter {
pub unjudged: bool,
/// `None` for no flag constraint, otherwise exactly that flag.
pub flag: Option<dr_types::FlagState>,
/// TRACES: FR-CAT-9
/// Only images whose original is stored on this device.
///
/// Carried here, beside the rating terms, because every query path already
/// threads this one struct: adding a parallel parameter to
/// `read_cells_scoped`, `read_cells_all` and both counts would give four
/// call sites the chance to disagree about what the grid is showing, and
/// the count disagreeing with the cells is the specific bug this type's
/// "filter in SQL" rule exists to prevent.
pub local_only: bool,
}
impl RatingFilter {
/// Whether this narrows anything, so the caller can skip the join.
pub fn is_unfiltered(&self) -> bool {
self.min_rating == 0 && !self.unjudged && self.flag.is_none()
self.min_rating == 0 && !self.unjudged && self.flag.is_none() && !self.local_only
}
/// The SQL predicate, against an `images` aliased as `i`.
@@ -218,6 +244,19 @@ impl RatingFilter {
));
}
if self.local_only {
// `tier_actual`, not `tier_desired`: the question is what is
// *here*, not what a pin has promised will be. An image queued for
// download is exactly the one that cannot be opened yet, so
// showing it under "on this device" would be the wrong answer to
// the only question this filter is asked.
terms.push(format!(
"EXISTS (SELECT 1 FROM image_cache ic
WHERE ic.image_id = i.id AND ic.tier_actual >= {})",
dr_types::Tier::Original.stored()
));
}
if terms.is_empty() {
String::new()
} else {
@@ -467,13 +506,45 @@ pub fn spawn_scan(
std::thread::spawn(move || {
let started = std::time::Instant::now();
if let Err(e) = run_scan(&tx, creds, user_id, root, filter, catalog_path, started) {
let _ = tx.send(ScanMessage::Failed(e));
let _ = tx.send(ScanMessage::Failed {
message: e.message,
offline: e.offline,
});
}
});
rx
}
/// A scan failure that still knows whether it was a connectivity failure.
///
/// The scan crosses a thread boundary, so the typed error cannot travel with
/// it; this carries the one bit that must survive.
struct ScanFailure {
message: String,
offline: bool,
}
impl ScanFailure {
/// A failure that is nothing to do with reachability — local I/O, a
/// runtime that would not start, a catalog that would not open.
fn local(message: impl std::fmt::Display) -> Self {
Self {
message: message.to_string(),
offline: false,
}
}
}
impl From<dr_sync::RemoteError> for ScanFailure {
fn from(e: dr_sync::RemoteError) -> Self {
Self {
offline: e.indicates_offline(),
message: e.to_string(),
}
}
}
fn run_scan(
tx: &Sender<ScanMessage>,
creds: AppCredentials,
@@ -482,19 +553,20 @@ fn run_scan(
filter: FormatFilter,
catalog_path: PathBuf,
started: std::time::Instant,
) -> Result<(), String> {
) -> Result<(), ScanFailure> {
if let Some(dir) = catalog_path.parent() {
std::fs::create_dir_all(dir).map_err(|e| format!("creating {}: {e}", dir.display()))?;
std::fs::create_dir_all(dir)
.map_err(|e| ScanFailure::local(format!("creating {}: {e}", dir.display())))?;
}
let catalog = Catalog::open(&catalog_path).map_err(|e| e.to_string())?;
let catalog = Catalog::open(&catalog_path).map_err(ScanFailure::local)?;
let rt = tokio::runtime::Builder::new_current_thread()
.enable_all()
.build()
.map_err(|e| e.to_string())?;
.map_err(ScanFailure::local)?;
rt.block_on(async {
let backend = NextcloudBackend::new(&creds, &user_id).map_err(|e| e.to_string())?;
let backend = NextcloudBackend::new(&creds, &user_id).map_err(ScanFailure::local)?;
// Stored folder ETags, so an unchanged subtree is skipped whole. On a
// first run this is empty and the walk is complete; on every run after
@@ -514,10 +586,9 @@ fn run_scan(
});
},
)
.await
.map_err(|e| e.to_string())?;
.await?;
persist(&catalog, &root, &result).map_err(|e| e.to_string())?;
persist(&catalog, &root, &result).map_err(ScanFailure::local)?;
// Report what the catalog holds, not what this pass listed. An
// incremental rescan lists only what changed, so its own count is
@@ -704,6 +775,41 @@ pub struct ThumbnailRequest {
pub needs_metadata: bool,
}
/// Why a full fetch failed, keeping the one bit the UI cannot re-derive.
///
/// The same reasoning as [`ScanFailure`]: the typed error cannot cross the
/// channel, and "offline" versus "refused" decides whether develop shows
/// "you are offline — this image is not stored locally" or a real error.
#[derive(Debug)]
pub struct FetchFailure {
pub message: String,
pub offline: bool,
}
impl FetchFailure {
fn local(message: impl std::fmt::Display) -> Self {
Self {
message: message.to_string(),
offline: false,
}
}
}
impl From<dr_sync::RemoteError> for FetchFailure {
fn from(e: dr_sync::RemoteError) -> Self {
Self {
offline: e.indicates_offline(),
message: e.to_string(),
}
}
}
impl std::fmt::Display for FetchFailure {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
f.write_str(&self.message)
}
}
/// Fetch one file in full, for opening it in develop.
///
/// Deliberately *not* the preview path. Browsing fetches a range and decodes
@@ -717,7 +823,7 @@ pub fn spawn_full_fetch(
creds: AppCredentials,
user_id: String,
path: String,
) -> Receiver<Result<Vec<u8>, String>> {
) -> Receiver<Result<Vec<u8>, FetchFailure>> {
let (tx, rx) = std::sync::mpsc::channel();
std::thread::spawn(move || {
@@ -727,7 +833,7 @@ pub fn spawn_full_fetch(
{
Ok(rt) => rt,
Err(e) => {
let _ = tx.send(Err(e.to_string()));
let _ = tx.send(Err(FetchFailure::local(e)));
return;
}
};
@@ -736,16 +842,13 @@ pub fn spawn_full_fetch(
let backend = match NextcloudBackend::new(&creds, &user_id) {
Ok(b) => b,
Err(e) => {
let _ = tx.send(Err(e.to_string()));
let _ = tx.send(Err(FetchFailure::local(e)));
return;
}
};
let id = RemoteId::Path(RemotePath::new(&path));
let got = backend
.get(&id, None)
.await
.map_err(|e| e.to_string());
let got = backend.get(&id, None).await.map_err(FetchFailure::from);
let _ = tx.send(got);
});
});
@@ -886,10 +989,17 @@ pub fn spawn_thumbnails(
// instant it is read.
let mut found = Vec::new();
// Set when the server proves unreachable, which abandons the rest
// of the batch. The remaining cells would each take a full timeout
// to reach the same conclusion — on a 120-cell window, minutes of
// the grid appearing to load against a server that is not there.
let mut offline = false;
for req in to_fetch {
let msg = fetch_one(&backend, store.as_mut(), &req, &mut found).await;
offline = matches!(msg, ThumbnailMessage::Offline { .. });
// A closed channel means the window went away mid-fetch.
if tx.send(msg).is_err() {
if tx.send(msg).is_err() || offline {
break;
}
}
@@ -901,19 +1011,27 @@ pub fn spawn_thumbnails(
// header fetches take tens of seconds, and a single write at the
// finish loses every one of them if the window closes first. It
// also lets the timeline appear while the rest are still arriving.
//
// Skipped entirely when the connection has already failed: these
// are network reads too, and there is nothing left to read from.
const FLUSH_EVERY: usize = 16;
log::info!("reading dates for {} image(s)", metadata_only.len());
for req in metadata_only {
if tx.send(ThumbnailMessage::DateProgress).is_err() {
break;
}
read_metadata_only(&backend, &req, &mut found).await;
if !offline {
log::info!("reading dates for {} image(s)", metadata_only.len());
for req in metadata_only {
if tx.send(ThumbnailMessage::DateProgress).is_err() {
break;
}
read_metadata_only(&backend, &req, &mut found).await;
if found.len() >= FLUSH_EVERY {
flush_metadata(&catalog_path, &mut found, &tx);
if found.len() >= FLUSH_EVERY {
flush_metadata(&catalog_path, &mut found, &tx);
}
}
}
// Always flushed, even when the batch was abandoned: whatever was
// read before the connection died is still true, and discarding it
// would mean re-fetching those headers next time.
flush_metadata(&catalog_path, &mut found, &tx);
});
});
@@ -938,11 +1056,24 @@ async fn fetch_one(
row: req.row,
reason,
};
// A connection failure is not this image's verdict. Reported as such so
// the caller can stop the batch rather than marking sixty cells
// individually unpreviewable over one dropped connection — a state the
// grid would then keep until something forced a reload.
let classify = |e: dr_sync::RemoteError| {
if e.indicates_offline() {
ThumbnailMessage::Offline {
reason: e.to_string(),
}
} else {
fail(e.to_string())
}
};
// Stage one: the header, enough to parse the container's IFDs.
let header = match backend.get(&id, Some(0..dr_decode::HEADER_BYTES)).await {
Ok(b) => b,
Err(e) => return fail(e.to_string()),
Err(e) => return classify(e),
};
// The same bytes carry EXIF. Reading it here is free — the alternative is
@@ -955,7 +1086,7 @@ async fn fetch_one(
let bytes = if header.starts_with(&[0xFF, 0xD8, 0xFF]) {
match backend.get(&id, None).await {
Ok(b) => b,
Err(e) => return fail(e.to_string()),
Err(e) => return classify(e),
}
} else {
let Some(loc) = dr_decode::locate_preview(&header, req.size) else {
@@ -970,7 +1101,7 @@ async fn fetch_one(
// Stage two: exactly the preview's bytes.
match backend.get(&id, Some(loc.range.clone())).await {
Ok(b) => b,
Err(e) => return fail(e.to_string()),
Err(e) => return classify(e),
}
};
@@ -1077,18 +1208,53 @@ fn flush_metadata(
/// One 256 KB header request, no preview range and no decode. This is what
/// gets a library dated when its thumbnails came from the store — including
/// shards synced from another device, which carry pixels but no metadata.
/// Read a header for its date.
///
/// Returns whether the file was **reached**, which the caller needs and cannot
/// otherwise tell: a header that carried no EXIF and a fetch that never
/// happened both leave `found` untouched, and recording the second as "this
/// image has no date" would let one lock mark it dateless for good.
async fn read_metadata_only(
backend: &NextcloudBackend,
req: &ThumbnailRequest,
found: &mut Vec<MetadataFound>,
) {
) -> bool {
let id = RemoteId::Path(RemotePath::new(&req.path));
match backend.get(&id, Some(0..dr_decode::HEADER_BYTES)).await {
Ok(header) => collect_metadata(&header, req, found),
// Not worth surfacing: the cell is already showing its thumbnail, and
// a missing date leaves the image off the timeline rather than broken.
Err(e) => log::debug!("reading date for {}: {e}", req.path),
// Retried, because one failure here is usually a lock rather than a
// verdict. Nextcloud's file locking answers a plain *read* with 423 under
// concurrency, and the identical range succeeds moments later — measured
// against a real server while twelve lanes were running. Without a retry
// those images sit out the whole pass over a lock that lasted a moment.
//
// Bounded and short: a genuinely missing or forbidden file must not cost
// three round trips before the sweep moves on.
const ATTEMPTS: usize = 3;
for attempt in 1..=ATTEMPTS {
match backend.get(&id, Some(0..dr_decode::HEADER_BYTES)).await {
Ok(header) => {
collect_metadata(&header, req, found);
return true;
}
Err(e) if e.is_transient() && attempt < ATTEMPTS => {
// Backing off at all matters more than the exact interval: the
// contention that produced the lock is our own lanes, so any
// pause lets the holder finish.
tokio::time::sleep(std::time::Duration::from_millis(
200 * attempt as u64,
))
.await;
}
Err(e) => {
// Not surfaced: a missing date leaves the image off the
// timeline rather than breaking anything, and the next sweep
// retries it regardless.
log::debug!("reading date for {} ({attempt} attempts): {e}", req.path);
return false;
}
}
}
false
}
/// Write capture metadata read during the thumbnail pass.
@@ -1209,7 +1375,12 @@ const SWEEP_CHUNK: usize = 96;
/// Deliberately bounded rather than unlimited: the grid's own interactive
/// fetches share this server, and a sweep that saturated the connection would
/// make browsing feel broken while it ran.
const SWEEP_LANES: usize = 12;
///
/// **Lowered from twelve after measuring.** Twelve produced 423 Locked on a
/// real server — Nextcloud's file locking answering a plain read under
/// contention we were creating ourselves. Six keeps most of the speedup
/// without provoking it; the retry above covers what still slips through.
const SWEEP_LANES: usize = 6;
/// Date **every** image in the library, not just the ones on screen.
///
@@ -1310,19 +1481,54 @@ pub fn spawn_sweep(
let backend = &backend;
async move {
let mut found = Vec::new();
let mut reached = Vec::new();
for req in lane {
read_metadata_only(backend, req, &mut found).await;
if read_metadata_only(backend, req, &mut found).await {
reached.push(req.image_id);
}
}
found
(found, reached)
}
}))
.await;
let mut found: Vec<MetadataFound> = results.into_iter().flatten().collect();
let mut found = Vec::new();
let mut reached = std::collections::HashSet::new();
for (lane_found, lane_reached) in results {
found.extend(lane_found);
reached.extend(lane_reached);
}
done += chunk.len();
// An image whose header carried no EXIF at all yields nothing
// to `found`, so nothing marks it examined and the next sweep
// fetches it again — for ever. Darktable exports strip
// metadata by default, and 2,188 of them in the reference
// library meant 2,188 pointless round trips per run.
//
// Recorded as examined with no date: the file was read and
// genuinely has none, which is a different state from "not
// looked at yet" and must not be confused with it.
let answered: std::collections::HashSet<i64> =
found.iter().map(|m| m.image_id).collect();
// Only files actually read. One that could not be fetched is
// left alone so the next pass retries it, rather than being
// written off over a lock or a dropped connection.
found.extend(chunk.iter().filter(|r| {
reached.contains(&r.image_id) && !answered.contains(&r.image_id)
}).map(
|r| MetadataFound {
image_id: r.image_id,
captured_at: None,
captured_offset: None,
camera: None,
lens: None,
iso: None,
},
));
dated += found.iter().filter(|m| m.captured_at.is_some()).count();
let read = found.len();
let read = answered.len();
flush_sweep(&catalog, &mut found);
log::info!(
"sweep: {done} done, {dated} dated ({read} read in {:.1}s)",
@@ -1661,6 +1867,27 @@ fn total_images_filtered(
Ok(n as usize)
}
/// TRACES: FR-CAT-9
/// How many visible images have their original stored on this device.
///
/// Whole-library, like the star counts beside it: the chip says what narrowing
/// to it would show, so counting only the current window would make it
/// describe the view it exists to change.
pub fn local_original_count(catalog: &Catalog) -> Result<usize, dr_catalog::CatalogError> {
let n: i64 = catalog.connection().query_row(
&format!(
"SELECT count(*) FROM images i
WHERE {VISIBLE}
AND EXISTS (SELECT 1 FROM image_cache ic
WHERE ic.image_id = i.id AND ic.tier_actual >= {})",
dr_types::Tier::Original.stored()
),
[],
|r| r.get(0),
)?;
Ok(n as usize)
}
/// Total images in the catalog, unfiltered.
///
/// What the scan reports and what the sidebar's "all images" row shows — the
@@ -2342,4 +2569,108 @@ mod tests {
has_preview: false,
}
}
// --- the local-only filter (FR-CAT-9) ---------------------------------
/// Record that an image's original is held locally at `tier`.
fn cache_at(catalog: &Catalog, id: dr_types::ImageId, tier: dr_types::Tier) {
catalog
.connection()
.execute(
"INSERT INTO image_cache (image_id, tier_actual, bytes)
VALUES (?1, ?2, 0)",
rusqlite::params![id.0 as i64, tier.stored()],
)
.unwrap();
}
#[test]
fn local_only_shows_just_the_images_held_here() {
let catalog = with_images(10);
let ids = image_ids(&catalog);
for id in &ids[0..3] {
cache_at(&catalog, *id, dr_types::Tier::Original);
}
let filter = RatingFilter {
local_only: true,
..Default::default()
};
let cells = read_cells_all(&catalog, &filter, 0, 120).unwrap();
assert_eq!(cells.len(), 3);
// The count the header shows must agree with the cells drawn, which is
// the whole reason the predicate lives in SQL rather than in a
// post-filter over the rows.
assert_eq!(total_images_filtered(&catalog, &filter).unwrap(), 3);
assert_eq!(local_original_count(&catalog).unwrap(), 3);
}
#[test]
fn a_cached_preview_is_not_a_local_original() {
// The filter answers "can I open this in develop right now", and a
// preview cannot. Counting it would put images in the offline set that
// fail the moment they are clicked.
let catalog = with_images(5);
let ids = image_ids(&catalog);
cache_at(&catalog, ids[0], dr_types::Tier::Preview);
cache_at(&catalog, ids[1], dr_types::Tier::Original);
let filter = RatingFilter {
local_only: true,
..Default::default()
};
assert_eq!(read_cells_all(&catalog, &filter, 0, 120).unwrap().len(), 1);
assert_eq!(local_original_count(&catalog).unwrap(), 1);
}
#[test]
fn local_only_composes_with_the_rating_filter() {
// "Five-star frames I can actually edit on this train" is one filter,
// not a mode that replaces the others.
let catalog = with_images(6);
let ids = image_ids(&catalog);
for id in &ids[0..4] {
cache_at(&catalog, *id, dr_types::Tier::Original);
}
// Rate two of the cached ones, and one that is not cached.
for id in [ids[0], ids[1], ids[5]] {
dr_catalog::rating::set_rating(catalog.connection(), id, 5).unwrap();
}
let filter = RatingFilter {
min_rating: 5,
local_only: true,
..Default::default()
};
let cells = read_cells_all(&catalog, &filter, 0, 120).unwrap();
assert_eq!(cells.len(), 2, "five-starred AND held locally");
assert_eq!(total_images_filtered(&catalog, &filter).unwrap(), 2);
}
#[test]
fn an_empty_cache_is_not_an_empty_library() {
// The unfiltered grid must not depend on the cache table having rows —
// a library nothing has been downloaded from is still a full library.
let catalog = with_images(4);
assert_eq!(local_original_count(&catalog).unwrap(), 0);
assert_eq!(
read_cells_all(&catalog, &RatingFilter::default(), 0, 120)
.unwrap()
.len(),
4
);
}
#[test]
fn local_only_counts_as_a_narrowing_filter() {
// `is_unfiltered` gates the "filtered" indicator. Reporting this one as
// unfiltered would leave a narrowed grid looking like the whole
// library, which is the state the indicator exists to prevent.
assert!(RatingFilter::default().is_unfiltered());
assert!(!RatingFilter {
local_only: true,
..Default::default()
}
.is_unfiltered());
}
}