Offer the backup, and then the rebuild, when the index turns out to be damaged

NFR-R6 asks for an integrity check at startup and two offers behind it, and
none of it existed. `PRAGMA integrity_check` appeared nowhere in the tree,
`Catalog::open` was `open` → `configure` → `migrate` → `backfill` and nothing
else, and corruption therefore surfaced as whatever rusqlite error the first
unlucky query happened to produce — "database disk image is malformed"
attached to a thumbnail refresh, elided into a 34px banner, over an empty
grid saying "No images found · Check the library folder". Two messages that
disagreed, and no way forward but deleting catalog.sqlite by hand.

The property that makes the second offer real was already here and load-
bearing: the catalog is an index, not a source of truth, rebuildable from
sources plus sidecars (invariant §5.2.4, cited by schema.rs, trash.rs and
lib.rs). And sync.rs already knew how to take a coherent snapshot of a WAL
database. What was missing was the check, the type, and the conversation.

Four pieces:

**The type.** `CatalogError::Corrupt`, and — the part that makes it worth
having — a hand-written `From<rusqlite::Error>` that classifies rather than
wraps. `SQLITE_CORRUPT` and `SQLITE_NOTADB` become `Corrupt` wherever they
arise, so a background job that trips over the damage first reports the same
thing the startup check would have. `SQLITE_IOERR` and `SQLITE_BUSY`
deliberately do not: a dropped network mount is a different problem, and
telling someone to rebuild their index would be a wrong answer delivered
confidently.

**The check.** `Catalog::open_verified`, `quick_check` before the open rather
than after, because opening runs migrations and a damaged catalog with an
intact header would otherwise have structure rewritten on top of structure
that is already wrong. Bound to `open_verified` and not to `open`: the check
reads every page, which is affordable once at startup where a user can answer
a question, and not affordable on the dozens of opens a session's background
tasks make.

**The backup.** NFR-R2's second clause, taken between `configure` and
`migrate` in `Catalog::open`. A migration is the one routine operation that
rewrites table structure, so it is the likeliest way this file becomes
unreadable, and it is the last moment the pre-migration state exists to be
copied. Three generations, through SQLite's backup API after a TRUNCATE
checkpoint — never `fs::copy`, which on a WAL database backs up a state older
than the catalog and possibly torn. A failure to take the copy is logged, not
raised: a full disk must not be what makes a library unopenable.

**The conversation.** The first line of the dialogue is that the photographs
and the edits are safe, before the diagnosis, because that is the question the
user is actually asking. Then the two offers, which are *not* interchangeable
and are not presented as if they were: a restore keeps collections, and a
rebuild cannot, because a manual collection is a set of images assembled by
hand and nothing in the filesystem records it (docs/catalog.md §8.1). The
labels say so, and the rebuild does not take the affirmative styling while a
restore is on the table.

One thing that is a fix rather than a feature: `show_catalog_now` now gates
the scan. `Catalog::open` succeeds on a file whose header survived, so the
scan that used to start immediately afterwards would write folder ETags and
image rows into damaged pages in the seconds while the user was still reading
the question — turning a file that had a backup into one where the backup is
the only copy left.

Restore also deletes the damaged catalog's `-wal` and `-shm`. That step is
easy to leave out and fatal to leave out: a journal belonging to the old file,
sitting beside the new one under the same name, is replayed into it on the
next open. That is not a restore, it is a fresh corruption with the evidence
gone.

Tested by corrupting a fixture catalog — 500 images and a collection, then
every page past the second overwritten — and driving both branches. The
restore is asserted on the collection, because a collection is precisely what
distinguishes the two paths; the rebuild on the damaged file being kept and
the next open producing an empty catalog at the current schema. Plus the
`SQLITE_NOTADB` presentation, a damaged backup being refused rather than
installed, and a v1 catalog whose pre-migration backup comes back reading
v1 rather than v11.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-30 10:34:12 +02:00
co-authored by Claude Opus 5
parent ef07e6ca3e
commit 8eeb9ba0f6
9 changed files with 1295 additions and 11 deletions
+285
View File
@@ -0,0 +1,285 @@
//! The two offers made when the catalog turns out to be damaged.
//!
//! `dr_catalog::recovery` owns the mechanism — the integrity check, the
//! backups, the restore, setting the damaged file aside. This module owns the
//! *conversation*: what a user is told has happened, which of the two answers
//! are available, and what runs afterwards.
//!
//! # The thing that has to be said first
//!
//! **The photographs are fine, and so are the edits.** A user told that their
//! library database is corrupt will assume they have lost their work, because
//! in every other photo application they would have. Here they have not:
//! sources are read-only to this application (NFR-R4), and ratings, keywords
//! and edit graphs live in sidecars beside the images for every catalogued
//! photograph, whether or not an account exists (FR-CAT-8, invariant §5.2.4).
//! That sentence is the first line of the dialogue, before the diagnosis,
//! because it is the answer to the question the user is actually asking.
//!
//! # Why the two answers are not interchangeable
//!
//! A restore brings back **collections**; a rebuild cannot. Every other thing
//! the catalog holds has authoritative backing outside it, which is what makes
//! a rebuild survivable — but a manual collection is a set of images the user
//! assembled by hand and nothing in the filesystem records it
//! (`docs/catalog.md` §8.1). So the labels say which one loses them, and the
//! rebuild is not given the affirmative styling while a restore is on offer.
//!
//! # Why the scan is held back
//!
//! `library_ui::open` shows the catalog and then starts a scan. On a damaged
//! catalog the scan is actively harmful: a plain `Catalog::open` on a file
//! whose header is intact succeeds, and the scan would then write folder
//! ETags and image rows into damaged pages — turning a recoverable file into
//! one whose backup is the only copy left, and doing it in the seconds while
//! the user is still reading the question. So `show_catalog_now` reports
//! whether it is safe to continue, and this module restarts the scan itself
//! once the file underneath has been replaced.
use std::cell::RefCell;
use std::path::{Path, PathBuf};
use std::rc::Rc;
use dr_catalog::recovery;
use slint::ComponentHandle;
use crate::library_ui::LibraryController;
use crate::AppWindow;
/// What the open question is about.
///
/// A thread-local rather than a field on `LibraryController`, because the
/// question is asked *before* that controller has a catalog and is answered by
/// callbacks wired at startup. Thread-local is sound here for the same reason
/// [`crate::memory`]'s registry is: everything below runs on the Slint event
/// loop thread, which is the only thread that has an `AppWindow` to show it
/// on.
thread_local! {
static PENDING: RefCell<Option<Pending>> = const { RefCell::new(None) };
}
/// The damaged catalog and what can be done about it.
struct Pending {
catalog: PathBuf,
/// Newest first. Empty is the ordinary case on a young install and is not
/// an error — it removes one offer, not both.
backups: Vec<recovery::Backup>,
}
/// Ask what should happen to a damaged catalog.
///
/// Called from `library_ui::show_catalog_now` when the startup integrity check
/// fails. `detail` is what SQLite said, carried through verbatim: a diagnosis
/// the user can quote into a bug report is worth more than a reassurance they
/// cannot check.
pub(crate) fn offer(window: &AppWindow, catalog: &Path, detail: &str) {
let backups = recovery::backups(catalog);
log::error!(
"catalog {} failed its integrity check: {detail} ({} backup(s) available)",
catalog.display(),
backups.len()
);
window.set_recovery_title("This library's index is damaged".into());
window.set_recovery_detail(
// Two facts and their order matters: what is safe, then what is lost.
"Your photographs and your edits are safe — they are in the files \
themselves and in the sidecars beside them. What is damaged is only \
DarkRoom's index of them, which can be rebuilt."
.into(),
);
window.set_recovery_diagnosis(detail.into());
match backups.first() {
Some(newest) => {
window.set_recovery_can_restore(true);
window.set_recovery_restore_label(
format!(
"Restore the backup from {} · keeps your collections",
describe_age(newest.taken_at)
)
.into(),
);
}
None => {
window.set_recovery_can_restore(false);
window.set_recovery_restore_label(slint::SharedString::new());
}
}
window.set_recovery_rebuild_label(
if backups.is_empty() {
// Nothing to compare it against, so the label states the cost
// rather than the difference.
"Rebuild from your photographs · rescans the library"
} else {
"Rebuild from your photographs · loses your collections"
}
.into(),
);
window.set_recovery_busy(false);
// The scan was held back, so the "Scanning…" the grid is showing behind
// this would be a lie the moment the question is dismissed.
window.set_library_scanning(false);
PENDING.with(|p| {
*p.borrow_mut() = Some(Pending {
catalog: catalog.to_path_buf(),
backups,
})
});
}
/// Close the question without answering it.
///
/// Leaves the banner set, because the library genuinely does not work and a
/// dialogue that vanishes leaving no trace of why nothing loads is worse than
/// no dialogue at all.
fn dismiss(window: &AppWindow) {
PENDING.with(|p| *p.borrow_mut() = None);
window.set_recovery_title(slint::SharedString::new());
window.set_library_scanning(false);
window.set_library_error(
"The library index is damaged. Rescan to rebuild it, or restore a backup.".into(),
);
}
/// Install the three answers.
///
/// Called at the end of `library_ui::wire`, which is where every other
/// window-level callback in this area is installed.
pub(crate) fn wire(
window: &AppWindow,
ctl: &Rc<LibraryController>,
coll_ctl: &Rc<crate::collections_ui::CollectionsController>,
) {
{
let weak = window.as_weak();
let ctl = ctl.clone();
let coll = coll_ctl.clone();
window.on_recovery_restore(move || {
let Some(w) = weak.upgrade() else { return };
answer(&w, &ctl, &coll, Answer::Restore);
});
}
{
let weak = window.as_weak();
let ctl = ctl.clone();
let coll = coll_ctl.clone();
window.on_recovery_rebuild(move || {
let Some(w) = weak.upgrade() else { return };
answer(&w, &ctl, &coll, Answer::Rebuild);
});
}
{
let weak = window.as_weak();
window.on_recovery_dismiss(move || {
let Some(w) = weak.upgrade() else { return };
dismiss(&w);
});
}
}
/// Which of the two the user chose.
#[derive(Clone, Copy, PartialEq, Eq)]
enum Answer {
Restore,
Rebuild,
}
/// Carry out an answer, then get the library going again.
///
/// Both answers end the same way — the file under `catalog_path` is one this
/// build can open — so both continue into the same two steps: open the catalog
/// for the grid, and start a scan. A rebuild needs the scan to have anything
/// at all; a restore needs it because the backup is by definition older than
/// the library.
fn answer(
window: &AppWindow,
ctl: &Rc<LibraryController>,
coll_ctl: &Rc<crate::collections_ui::CollectionsController>,
which: Answer,
) {
let Some(pending) = PENDING.with(|p| p.borrow_mut().take()) else {
return;
};
window.set_recovery_busy(true);
// Before the file moves, not after. There is normally no open catalog here
// — `show_catalog_now` returned before storing one — but "normally" is not
// a guarantee worth resting a file rename on, and a connection to a file
// that has just been renamed out from under it reads the damaged pages
// forever.
crate::library_ui::forget_catalog(ctl);
// Synchronous, on the UI thread, and that is a considered choice rather
// than an oversight: this is a file copy of a catalog — tens of megabytes
// at 50k images — at a moment when there is nothing else on screen to
// block, no scan running, and no frame worth keeping smooth. Moving it to
// a worker would buy a spinner and cost the guarantee that nothing else
// touches the file while it is being replaced.
let outcome = match which {
Answer::Restore => match pending.backups.first() {
Some(b) => recovery::restore(&pending.catalog, &b.path),
None => Ok(()),
},
Answer::Rebuild => recovery::set_aside(&pending.catalog).map(|_| ()),
};
if let Err(e) = outcome {
// The question stays up: the *other* answer may still work, and a
// failed restore in particular leaves the rebuild untouched.
log::error!("recovery failed: {e}");
window.set_recovery_busy(false);
window.set_recovery_diagnosis(format!("That did not work: {e}").into());
PENDING.with(|p| *p.borrow_mut() = Some(pending));
return;
}
window.set_recovery_busy(false);
window.set_recovery_title(slint::SharedString::new());
window.set_library_error(slint::SharedString::new());
if crate::library_ui::show_catalog_now(window, ctl, &pending.catalog, coll_ctl) {
crate::library_ui::start_rescan(window, ctl, coll_ctl);
}
}
/// "today", "3 days ago" — enough to choose by, without a date library.
///
/// The user is deciding how much work a restore costs them, and the answer to
/// that is an *age*, not a timestamp: "yesterday" is immediately actionable
/// and "1756512000" is not. Whole days, because an hour's precision would
/// invite a confidence the backup schedule does not earn.
fn describe_age(taken_at: i64) -> String {
let now = std::time::SystemTime::now()
.duration_since(std::time::UNIX_EPOCH)
.map(|d| d.as_secs() as i64)
.unwrap_or(0);
let days = (now - taken_at).max(0) / 86_400;
match days {
0 => "today".to_string(),
1 => "yesterday".to_string(),
d => format!("{d} days ago"),
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn an_age_reads_as_an_age() {
let now = std::time::SystemTime::now()
.duration_since(std::time::UNIX_EPOCH)
.unwrap()
.as_secs() as i64;
assert_eq!(describe_age(now), "today");
assert_eq!(describe_age(now - 86_400), "yesterday");
assert_eq!(describe_age(now - 5 * 86_400), "5 days ago");
// A clock that has gone backwards must not produce "-2 days ago".
assert_eq!(describe_age(now + 86_400), "today");
}
}