Offer the backup, and then the rebuild, when the index turns out to be damaged

NFR-R6 asks for an integrity check at startup and two offers behind it, and
none of it existed. `PRAGMA integrity_check` appeared nowhere in the tree,
`Catalog::open` was `open` → `configure` → `migrate` → `backfill` and nothing
else, and corruption therefore surfaced as whatever rusqlite error the first
unlucky query happened to produce — "database disk image is malformed"
attached to a thumbnail refresh, elided into a 34px banner, over an empty
grid saying "No images found · Check the library folder". Two messages that
disagreed, and no way forward but deleting catalog.sqlite by hand.

The property that makes the second offer real was already here and load-
bearing: the catalog is an index, not a source of truth, rebuildable from
sources plus sidecars (invariant §5.2.4, cited by schema.rs, trash.rs and
lib.rs). And sync.rs already knew how to take a coherent snapshot of a WAL
database. What was missing was the check, the type, and the conversation.

Four pieces:

**The type.** `CatalogError::Corrupt`, and — the part that makes it worth
having — a hand-written `From<rusqlite::Error>` that classifies rather than
wraps. `SQLITE_CORRUPT` and `SQLITE_NOTADB` become `Corrupt` wherever they
arise, so a background job that trips over the damage first reports the same
thing the startup check would have. `SQLITE_IOERR` and `SQLITE_BUSY`
deliberately do not: a dropped network mount is a different problem, and
telling someone to rebuild their index would be a wrong answer delivered
confidently.

**The check.** `Catalog::open_verified`, `quick_check` before the open rather
than after, because opening runs migrations and a damaged catalog with an
intact header would otherwise have structure rewritten on top of structure
that is already wrong. Bound to `open_verified` and not to `open`: the check
reads every page, which is affordable once at startup where a user can answer
a question, and not affordable on the dozens of opens a session's background
tasks make.

**The backup.** NFR-R2's second clause, taken between `configure` and
`migrate` in `Catalog::open`. A migration is the one routine operation that
rewrites table structure, so it is the likeliest way this file becomes
unreadable, and it is the last moment the pre-migration state exists to be
copied. Three generations, through SQLite's backup API after a TRUNCATE
checkpoint — never `fs::copy`, which on a WAL database backs up a state older
than the catalog and possibly torn. A failure to take the copy is logged, not
raised: a full disk must not be what makes a library unopenable.

**The conversation.** The first line of the dialogue is that the photographs
and the edits are safe, before the diagnosis, because that is the question the
user is actually asking. Then the two offers, which are *not* interchangeable
and are not presented as if they were: a restore keeps collections, and a
rebuild cannot, because a manual collection is a set of images assembled by
hand and nothing in the filesystem records it (docs/catalog.md §8.1). The
labels say so, and the rebuild does not take the affirmative styling while a
restore is on the table.

One thing that is a fix rather than a feature: `show_catalog_now` now gates
the scan. `Catalog::open` succeeds on a file whose header survived, so the
scan that used to start immediately afterwards would write folder ETags and
image rows into damaged pages in the seconds while the user was still reading
the question — turning a file that had a backup into one where the backup is
the only copy left.

Restore also deletes the damaged catalog's `-wal` and `-shm`. That step is
easy to leave out and fatal to leave out: a journal belonging to the old file,
sitting beside the new one under the same name, is replayed into it on the
next open. That is not a restore, it is a fresh corruption with the evidence
gone.

Tested by corrupting a fixture catalog — 500 images and a collection, then
every page past the second overwritten — and driving both branches. The
restore is asserted on the collection, because a collection is precisely what
distinguishes the two paths; the rebuild on the damaged file being kept and
the next open producing an empty catalog at the current schema. Plus the
`SQLITE_NOTADB` presentation, a damaged backup being refused rather than
installed, and a v1 catalog whose pre-migration backup comes back reading
v1 rather than v11.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-30 10:34:12 +02:00
co-authored by Claude Opus 5
parent ef07e6ca3e
commit 8eeb9ba0f6
9 changed files with 1295 additions and 11 deletions
+1
View File
@@ -49,6 +49,7 @@ mod net_runtime;
mod peaking;
mod preset_store;
mod presets;
mod recovery_ui;
mod remote;
mod segmentation;
mod settings_store;
+46 -7
View File
@@ -862,7 +862,13 @@ pub fn open(
// Before the scan, not after it: the grid can be filled from disk now and
// the scan is only ever going to add to it.
show_catalog_now(window, &ctl, &path, &coll_ctl);
//
// And it is the gate on the scan, not merely a prelude to it — a damaged
// catalog has a question on screen, and a scan writing into it while that
// question is unanswered is how the last good copy gets destroyed.
if !show_catalog_now(window, &ctl, &path, &coll_ctl) {
return;
}
let rx = library::spawn_scan(
conn.clone(),
@@ -1095,7 +1101,7 @@ fn drain_scan(
/// the same operation: a scan is the only request that both proves the server
/// is reachable and brings the catalog up to date. Keeping them one function
/// is what stops "retry" from quietly becoming a weaker probe than "rescan".
fn start_rescan(
pub(crate) fn start_rescan(
window: &AppWindow,
ctl: &Rc<LibraryController>,
coll_ctl: &Rc<crate::collections_ui::CollectionsController>,
@@ -1916,23 +1922,38 @@ fn schedule_reload(window: &AppWindow, ctl: &Rc<LibraryController>) {
/// state already says "Scanning…", and an error here would contradict a scan
/// that is working perfectly. `Catalog::open` creates the file in that case, so
/// what the grid reads is an empty catalog rather than a failure.
fn show_catalog_now(
/// Returns whether it is safe to go on and scan.
///
/// `false` means the catalog is damaged and the recovery question is up. The
/// caller must not start a scan on that answer: `Catalog::open` succeeds on a
/// file whose header survived, so the scan would write ETags and image rows
/// into damaged pages while the user is still reading the question — turning a
/// file that had a backup into one where the backup is the only copy left.
pub(crate) fn show_catalog_now(
window: &AppWindow,
ctl: &Rc<LibraryController>,
catalog_path: &std::path::Path,
coll_ctl: &Rc<crate::collections_ui::CollectionsController>,
) {
) -> bool {
if ctl.catalog.borrow().is_some() {
return;
return true;
}
let cat = match Catalog::open(catalog_path) {
// Verified rather than plain: this is the once-per-launch moment where a
// full check is affordable and there is a user in front of it who can
// answer the question. See `dr_catalog::recovery` for why it is not on
// every open.
let cat = match Catalog::open_verified(catalog_path) {
Ok(cat) => cat,
Err(dr_catalog::CatalogError::Corrupt { detail }) => {
crate::recovery_ui::offer(window, catalog_path, &detail);
return false;
}
Err(e) => {
// Not surfaced: the scan is the thing that has to work, and it is
// still running. If it fails too, it reports for both of them.
log::info!("no catalog to show before the scan: {e}");
return;
return true;
}
};
@@ -1964,6 +1985,19 @@ fn show_catalog_now(
crate::collections_ui::refresh_tree(window, coll_ctl, &cat);
*ctl.catalog.borrow_mut() = Some(cat);
load_window(window, ctl);
true
}
/// Drop the open catalog, so the next `show_catalog_now` opens the file
/// again rather than returning early.
///
/// Only recovery needs this, and it needs it for a specific reason: the file
/// under that connection has been replaced. A handle to the catalog that was
/// there before is a handle to a file that no longer has a name, and every
/// read through it would return the damaged pages the recovery just moved out
/// of the way.
pub(crate) fn forget_catalog(ctl: &Rc<LibraryController>) {
*ctl.catalog.borrow_mut() = None;
}
/// What one cell of the outgoing model is worth keeping.
@@ -5642,6 +5676,11 @@ pub fn wire<F>(
start_rescan(&w, &ctl, &coll_ctl);
});
}
// Last, and in its own module: the answers to a damaged catalog have
// nothing to do with the library view except that they run before it
// exists.
crate::recovery_ui::wire(window, &ctl, &coll_ctl);
}
/// Reload the grid after the filter changed.
+285
View File
@@ -0,0 +1,285 @@
//! The two offers made when the catalog turns out to be damaged.
//!
//! `dr_catalog::recovery` owns the mechanism — the integrity check, the
//! backups, the restore, setting the damaged file aside. This module owns the
//! *conversation*: what a user is told has happened, which of the two answers
//! are available, and what runs afterwards.
//!
//! # The thing that has to be said first
//!
//! **The photographs are fine, and so are the edits.** A user told that their
//! library database is corrupt will assume they have lost their work, because
//! in every other photo application they would have. Here they have not:
//! sources are read-only to this application (NFR-R4), and ratings, keywords
//! and edit graphs live in sidecars beside the images for every catalogued
//! photograph, whether or not an account exists (FR-CAT-8, invariant §5.2.4).
//! That sentence is the first line of the dialogue, before the diagnosis,
//! because it is the answer to the question the user is actually asking.
//!
//! # Why the two answers are not interchangeable
//!
//! A restore brings back **collections**; a rebuild cannot. Every other thing
//! the catalog holds has authoritative backing outside it, which is what makes
//! a rebuild survivable — but a manual collection is a set of images the user
//! assembled by hand and nothing in the filesystem records it
//! (`docs/catalog.md` §8.1). So the labels say which one loses them, and the
//! rebuild is not given the affirmative styling while a restore is on offer.
//!
//! # Why the scan is held back
//!
//! `library_ui::open` shows the catalog and then starts a scan. On a damaged
//! catalog the scan is actively harmful: a plain `Catalog::open` on a file
//! whose header is intact succeeds, and the scan would then write folder
//! ETags and image rows into damaged pages — turning a recoverable file into
//! one whose backup is the only copy left, and doing it in the seconds while
//! the user is still reading the question. So `show_catalog_now` reports
//! whether it is safe to continue, and this module restarts the scan itself
//! once the file underneath has been replaced.
use std::cell::RefCell;
use std::path::{Path, PathBuf};
use std::rc::Rc;
use dr_catalog::recovery;
use slint::ComponentHandle;
use crate::library_ui::LibraryController;
use crate::AppWindow;
/// What the open question is about.
///
/// A thread-local rather than a field on `LibraryController`, because the
/// question is asked *before* that controller has a catalog and is answered by
/// callbacks wired at startup. Thread-local is sound here for the same reason
/// [`crate::memory`]'s registry is: everything below runs on the Slint event
/// loop thread, which is the only thread that has an `AppWindow` to show it
/// on.
thread_local! {
static PENDING: RefCell<Option<Pending>> = const { RefCell::new(None) };
}
/// The damaged catalog and what can be done about it.
struct Pending {
catalog: PathBuf,
/// Newest first. Empty is the ordinary case on a young install and is not
/// an error — it removes one offer, not both.
backups: Vec<recovery::Backup>,
}
/// Ask what should happen to a damaged catalog.
///
/// Called from `library_ui::show_catalog_now` when the startup integrity check
/// fails. `detail` is what SQLite said, carried through verbatim: a diagnosis
/// the user can quote into a bug report is worth more than a reassurance they
/// cannot check.
pub(crate) fn offer(window: &AppWindow, catalog: &Path, detail: &str) {
let backups = recovery::backups(catalog);
log::error!(
"catalog {} failed its integrity check: {detail} ({} backup(s) available)",
catalog.display(),
backups.len()
);
window.set_recovery_title("This library's index is damaged".into());
window.set_recovery_detail(
// Two facts and their order matters: what is safe, then what is lost.
"Your photographs and your edits are safe — they are in the files \
themselves and in the sidecars beside them. What is damaged is only \
DarkRoom's index of them, which can be rebuilt."
.into(),
);
window.set_recovery_diagnosis(detail.into());
match backups.first() {
Some(newest) => {
window.set_recovery_can_restore(true);
window.set_recovery_restore_label(
format!(
"Restore the backup from {} · keeps your collections",
describe_age(newest.taken_at)
)
.into(),
);
}
None => {
window.set_recovery_can_restore(false);
window.set_recovery_restore_label(slint::SharedString::new());
}
}
window.set_recovery_rebuild_label(
if backups.is_empty() {
// Nothing to compare it against, so the label states the cost
// rather than the difference.
"Rebuild from your photographs · rescans the library"
} else {
"Rebuild from your photographs · loses your collections"
}
.into(),
);
window.set_recovery_busy(false);
// The scan was held back, so the "Scanning…" the grid is showing behind
// this would be a lie the moment the question is dismissed.
window.set_library_scanning(false);
PENDING.with(|p| {
*p.borrow_mut() = Some(Pending {
catalog: catalog.to_path_buf(),
backups,
})
});
}
/// Close the question without answering it.
///
/// Leaves the banner set, because the library genuinely does not work and a
/// dialogue that vanishes leaving no trace of why nothing loads is worse than
/// no dialogue at all.
fn dismiss(window: &AppWindow) {
PENDING.with(|p| *p.borrow_mut() = None);
window.set_recovery_title(slint::SharedString::new());
window.set_library_scanning(false);
window.set_library_error(
"The library index is damaged. Rescan to rebuild it, or restore a backup.".into(),
);
}
/// Install the three answers.
///
/// Called at the end of `library_ui::wire`, which is where every other
/// window-level callback in this area is installed.
pub(crate) fn wire(
window: &AppWindow,
ctl: &Rc<LibraryController>,
coll_ctl: &Rc<crate::collections_ui::CollectionsController>,
) {
{
let weak = window.as_weak();
let ctl = ctl.clone();
let coll = coll_ctl.clone();
window.on_recovery_restore(move || {
let Some(w) = weak.upgrade() else { return };
answer(&w, &ctl, &coll, Answer::Restore);
});
}
{
let weak = window.as_weak();
let ctl = ctl.clone();
let coll = coll_ctl.clone();
window.on_recovery_rebuild(move || {
let Some(w) = weak.upgrade() else { return };
answer(&w, &ctl, &coll, Answer::Rebuild);
});
}
{
let weak = window.as_weak();
window.on_recovery_dismiss(move || {
let Some(w) = weak.upgrade() else { return };
dismiss(&w);
});
}
}
/// Which of the two the user chose.
#[derive(Clone, Copy, PartialEq, Eq)]
enum Answer {
Restore,
Rebuild,
}
/// Carry out an answer, then get the library going again.
///
/// Both answers end the same way — the file under `catalog_path` is one this
/// build can open — so both continue into the same two steps: open the catalog
/// for the grid, and start a scan. A rebuild needs the scan to have anything
/// at all; a restore needs it because the backup is by definition older than
/// the library.
fn answer(
window: &AppWindow,
ctl: &Rc<LibraryController>,
coll_ctl: &Rc<crate::collections_ui::CollectionsController>,
which: Answer,
) {
let Some(pending) = PENDING.with(|p| p.borrow_mut().take()) else {
return;
};
window.set_recovery_busy(true);
// Before the file moves, not after. There is normally no open catalog here
// — `show_catalog_now` returned before storing one — but "normally" is not
// a guarantee worth resting a file rename on, and a connection to a file
// that has just been renamed out from under it reads the damaged pages
// forever.
crate::library_ui::forget_catalog(ctl);
// Synchronous, on the UI thread, and that is a considered choice rather
// than an oversight: this is a file copy of a catalog — tens of megabytes
// at 50k images — at a moment when there is nothing else on screen to
// block, no scan running, and no frame worth keeping smooth. Moving it to
// a worker would buy a spinner and cost the guarantee that nothing else
// touches the file while it is being replaced.
let outcome = match which {
Answer::Restore => match pending.backups.first() {
Some(b) => recovery::restore(&pending.catalog, &b.path),
None => Ok(()),
},
Answer::Rebuild => recovery::set_aside(&pending.catalog).map(|_| ()),
};
if let Err(e) = outcome {
// The question stays up: the *other* answer may still work, and a
// failed restore in particular leaves the rebuild untouched.
log::error!("recovery failed: {e}");
window.set_recovery_busy(false);
window.set_recovery_diagnosis(format!("That did not work: {e}").into());
PENDING.with(|p| *p.borrow_mut() = Some(pending));
return;
}
window.set_recovery_busy(false);
window.set_recovery_title(slint::SharedString::new());
window.set_library_error(slint::SharedString::new());
if crate::library_ui::show_catalog_now(window, ctl, &pending.catalog, coll_ctl) {
crate::library_ui::start_rescan(window, ctl, coll_ctl);
}
}
/// "today", "3 days ago" — enough to choose by, without a date library.
///
/// The user is deciding how much work a restore costs them, and the answer to
/// that is an *age*, not a timestamp: "yesterday" is immediately actionable
/// and "1756512000" is not. Whole days, because an hour's precision would
/// invite a confidence the backup schedule does not earn.
fn describe_age(taken_at: i64) -> String {
let now = std::time::SystemTime::now()
.duration_since(std::time::UNIX_EPOCH)
.map(|d| d.as_secs() as i64)
.unwrap_or(0);
let days = (now - taken_at).max(0) / 86_400;
match days {
0 => "today".to_string(),
1 => "yesterday".to_string(),
d => format!("{d} days ago"),
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn an_age_reads_as_an_age() {
let now = std::time::SystemTime::now()
.duration_since(std::time::UNIX_EPOCH)
.unwrap()
.as_secs() as i64;
assert_eq!(describe_age(now), "today");
assert_eq!(describe_age(now - 86_400), "yesterday");
assert_eq!(describe_age(now - 5 * 86_400), "5 days ago");
// A clock that has gone backwards must not produce "-2 days ago".
assert_eq!(describe_age(now + 86_400), "today");
}
}
+43
View File
@@ -11,6 +11,7 @@ import { GestureRow } from "gestures.slint";
import { Button, PanelHeading, Label, Value, Caption, Panel, EmptyState, ProgressBar, ActivityRow } from "widgets.slint";
import { CollectionsPanel, CollectionRow, OfflinePrompt } from "collections.slint";
import { HistogramPanel, HistogramView } from "histogram.slint";
import { RecoveryPrompt } from "recovery.slint";
import { PresetSheet, ScopeChips, ScopeKind } from "presets.slint";
import { FocusMarks, FocusPanel } from "peaking.slint";
import { SettingsPage } from "settings.slint";
@@ -365,6 +366,21 @@ export component AppWindow inherits Window {
callback offline-prompt-release();
callback offline-prompt-dismiss();
// The question a damaged catalog asks. Same shape as the prompt above and
// for the same reason: an empty title is what closes it, and every word in
// it is composed in Rust, which is the only side that knows what SQLite
// said and which backups exist.
in property <string> recovery-title: "";
in property <string> recovery-detail: "";
in property <string> recovery-diagnosis: "";
in property <string> recovery-restore-label: "";
in property <bool> recovery-can-restore: false;
in property <string> recovery-rebuild-label: "";
in property <bool> recovery-busy: false;
callback recovery-restore();
callback recovery-rebuild();
callback recovery-dismiss();
in property <string> library-root-label: "";
in-out property <[TimelineBar]> library-timeline;
in property <string> library-timeline-label: "";
@@ -1145,6 +1161,14 @@ in property <bool> panel-visible: true;
// would leave the library from behind an open question — the
// view changing underneath a modal, which reads as the app
// having lost its place.
//
// The recovery question is asked first because it is drawn
// over everything, the offline prompt included: Back must
// reach the thing the user can actually see.
if (root.recovery-title != "") {
root.recovery-dismiss();
return accept;
}
if (root.offline-prompt-title != "") {
root.offline-prompt-dismiss();
return accept;
@@ -2664,5 +2688,24 @@ in property <bool> panel-visible: true;
release() => { root.offline-prompt-release(); }
dismiss() => { root.offline-prompt-dismiss(); }
}
// Last, and therefore over everything including the settings page and
// the offline prompt. Not a preference about layering: this is asked
// before the grid exists, and nothing else in the window is about a
// library that can be read.
RecoveryPrompt {
width: 100%;
height: 100%;
title: root.recovery-title;
detail: root.recovery-detail;
diagnosis: root.recovery-diagnosis;
restore-label: root.recovery-restore-label;
can-restore: root.recovery-can-restore;
rebuild-label: root.recovery-rebuild-label;
busy: root.recovery-busy;
restore() => { root.recovery-restore(); }
rebuild() => { root.recovery-rebuild(); }
dismiss() => { root.recovery-dismiss(); }
}
}
}
+145
View File
@@ -0,0 +1,145 @@
// The question asked when the catalog turns out to be damaged.
//
// # Why this is a modal, when almost nothing else here is
//
// The house rule in this interface is to put the consequence in the button's
// label rather than to raise a dialogue — "Export 40", "Empty trash · 128" —
// and a genuine modal is kept for the two cases where the answer commits
// gigabytes. This is the third case, and it earns it for a different reason:
// there is nothing behind it to interact with. The grid cannot be drawn, the
// scan must not run (it would write into the damage), and every control in the
// window is about a library that cannot be read. A banner over an empty grid
// would be a question the user could scroll away from and then wonder why
// nothing worked.
//
// # Why the backdrop does not dismiss it
//
// Every other overlay here closes on a tap outside, and this one deliberately
// does not. A stray tap that loses the two offers leaves the application in a
// state with no way forward and no obvious way back to the question. There is
// a "Leave it for now" button instead, which says what it does.
//
// # Why the destructive answer is not the primary one
//
// A restore keeps the user's collections; a rebuild cannot, because a manual
// collection is a set of images the user assembled by hand and nothing in the
// filesystem records it (docs/catalog.md §8.1). So the two answers are not
// interchangeable, the difference is stated in the button rather than in a
// second dialogue after it, and the rebuild is the plain button even when it
// is the only one available.
import { Theme } from "theme.slint";
import { Button } from "widgets.slint";
export component RecoveryPrompt inherits Rectangle {
/// What went wrong, in the user's terms. Empty closes the prompt — one
/// source for "is this open", rather than a bool that can disagree with
/// the words beside it.
in property <string> title;
/// What is safe and what is not, which is the part that determines whether
/// the next minute is frightening.
in property <string> detail;
/// What SQLite actually said, kept because a bug report needs it and
/// because a diagnosis the user can read is worth more than a reassurance
/// they cannot check.
in property <string> diagnosis;
/// The restore offer, naming the backup's date. Empty when there is no
/// backup to restore from, which is the case a fresh install is in.
in property <string> restore-label;
in property <bool> can-restore: false;
/// The rebuild offer, naming what it costs — a full rescan, and the
/// collections it cannot bring back.
in property <string> rebuild-label;
/// Set while a restore or rebuild is running, so neither can be started
/// twice against the same file.
in property <bool> busy: false;
callback restore();
callback rebuild();
callback dismiss();
visible: root.title != "";
background: #000000E0;
// Swallows everything that misses the card, and answers nothing. See the
// header: losing this by a stray tap leaves nowhere to go.
TouchArea { }
Rectangle {
width: min(460px, parent.width - 2 * Theme.gap-lg);
height: min(card.preferred-height, parent.height - 2 * Theme.gap-lg);
x: (parent.width - self.width) / 2;
y: (parent.height - self.height) / 2;
background: Theme.surface;
border-radius: Theme.radius;
border-width: 1px;
border-color: Theme.rule;
TouchArea { }
card := VerticalLayout {
padding: Theme.gap-lg;
spacing: Theme.gap;
Text {
text: "Recover library";
color: Theme.ink-faint;
font-size: Theme.text-sm;
font-weight: 700;
letter-spacing: 1.2px;
}
Text {
text: root.title;
color: Theme.ink;
font-size: Theme.text-lg;
font-weight: 600;
wrap: word-wrap;
}
Text {
text: root.detail;
color: Theme.ink-dim;
font-size: Theme.text;
wrap: word-wrap;
}
// Wrapped rather than elided: this is the one line a bug report
// needs verbatim, and a truncated SQLite message is no message.
Text {
text: root.diagnosis;
color: Theme.ink-faint;
font-size: Theme.text-sm;
wrap: word-wrap;
}
Rectangle { height: 1px; background: Theme.rule; }
// Stacked, not a row: each label carries what its answer costs —
// a date, a count of photographs — and three of those side by side
// elide away exactly the part that lets the user choose.
if root.can-restore: Button {
text: root.busy ? "Working…" : root.restore-label;
primary: true;
enabled: !root.busy;
clicked => { root.restore(); }
}
Button {
text: root.busy ? "Working…" : root.rebuild-label;
// Primary only when it is the only answer there is. A rebuild
// discards collections, so it does not get the emphasis while
// a restore that keeps them is on the table.
primary: !root.can-restore;
enabled: !root.busy;
clicked => { root.rebuild(); }
}
Button {
text: "Leave it for now";
enabled: !root.busy;
clicked => { root.dismiss(); }
}
}
}
}