Files
DarkRoom/core/dr-catalog/src/backfilled.rs
T
dtourolle 02ddce8d80 Stamp the backfill on the newest rows, not only their ids
The backfill stamp read max(id) of images and versions and max(rowid)
of keywords. None of those tables is AUTOINCREMENT, so SQLite hands a
freed newest id out again: empty the trash of the newest photograph and
scan a new one, or let a local folder's walk delete a renamed file's row
and insert the new name in the same pass, and the new image takes the
old id. max(id) does not move, nor does count(*), and when a newer
version elsewhere keeps max(versions.id) still too, the stamp matched
and the open skipped the backfill.

That row is exactly one that needs it. Neither scan path creates the
default version: scan::persist and walk insert the image and leave the
version, the RAW/JPEG pairing and the keyword terms to the next open.
Skipped, the image went without them until the app restarted, so a
rating or a pulled sidecar judgement had no version to land on and a
JPEG beside its RAW showed twice.

The stamp now carries the newest row's content: the newest image's id,
path, added time and whether it has a version; the newest version's id
and image; the newest assignment's rowid, version and word. Whether the
newest image has a version is the part that cannot be fooled - after a
backfill every image has one, and a row that has just taken a freed id
has none - so the two stamps differ even when the same file comes back
at the same id in the same second. Still one statement: three reverse
rowid scans that stop at the first row, and one probe of versions_image.
An open that skips still costs ~1 ms on the reference catalog copy.

This closes the hole in the stamp itself rather than by a forget() at
each delete site, so a delete path added later, or one in another
process, cannot reopen it. Two tests delete the newest image and insert
another at the freed id on a separate connection, with a newer version
elsewhere holding max(versions.id); both fail against the old stamp.
2026-09-26 14:26:10 -04:00

418 lines
17 KiB
Rust

//! TRACES: NFR-P9
//! Which catalog files this process has already backfilled, and as of what.
//!
//! # Why this exists
//!
//! [`crate::schema::backfill`] used to run inside every [`crate::Catalog::open`],
//! and every worker thread opens its own connection. Landing on a photograph
//! in develop opened the catalog five times — the fetch of the original and a
//! cache check per prefetched neighbour — and each open paid the whole
//! backfill: an anti-join of every image against its versions, a pass over
//! every default version's uuid, the unpaired JPEGs and the keyword
//! vocabulary. On the reference library that was ~12 ms an open and ~60 ms of
//! CPU a landing, spent confirming that nothing had changed since the open
//! before.
//!
//! # What makes skipping it safe
//!
//! Everything the backfill repairs is a row some write *added*: an image
//! inserted by a scan or an import has no default version and may be the RAW
//! beside an unpaired JPEG; a version merged or restored from an older build
//! may carry a minted uuid; a keyword assignment merged from a remote may name
//! a word with no term. So the question "is any work owed?" is answered by
//! whether those tables have gained rows since the last backfill, and that is
//! a read of each table's last row — the last page of its b-tree — rather than
//! a scan.
//!
//! # Why the last row, and not only its id
//!
//! None of these tables is `AUTOINCREMENT`, so SQLite hands out the largest
//! rowid plus one, and an id freed by deleting the newest row is handed out
//! again. That is an ordinary sequence, not a contrived one: emptying the
//! trash of the newest photograph and then scanning a new one, or a local
//! folder's walk removing a renamed file's row and inserting the new name in
//! the same pass. `max(id)` does not move, and neither does `count(*)`. And
//! the row that took the id is exactly one that needs the backfill, because
//! neither scan creates default versions — `persist` and the walk insert the
//! image and leave the version, the pairing and the keyword terms to the next
//! open. Skipped, it would go without them until the app restarted: a rating
//! or a keyword with nowhere to land, a JPEG beside its RAW shown twice.
//!
//! So the stamp carries the last row's content as well as its id: the newest
//! image's path, when it was added, and **whether it has a version**; the
//! newest version's image; the newest assignment's word and version. Whether
//! the newest image has a version is the part that cannot be fooled: once the
//! backfill has run, every image has one, and a row that has just taken a
//! freed id has none, so the two stamps differ whatever the path and the time
//! say. The others make the newest version or assignment a different row
//! whenever a different one took its id; one that is the same content at the
//! same id is the same row as far as the backfill is concerned.
//!
//! The [`Stamp`] is those, the schema version, and the file's identity.
//! An open whose stamp matches the one recorded at the last backfill of the
//! same path skips it; anything else runs it. That covers the cases that must
//! run it:
//!
//! - **The first open in a process.** Nothing is recorded yet.
//! - **A migration.** `user_version` is in the stamp, and [`crate::Catalog::open`]
//! also runs the backfill unconditionally whenever `migrate` moved the
//! schema, because that is what the backfill was written for.
//! - **A pulled catalog.** The merge inserts assignments, which moves the
//! stamp; and [`crate::sync::merge_remote`] [`forget`]s the path as well, so
//! the next open backfills even when every incoming row collided.
//! - **A file replaced underneath the path** — a restore from backup, a
//! rebuild, a catalog copied in. On unix the device and inode are in the
//! stamp, and a replacement is a new inode; [`crate::recovery::set_aside`],
//! the first step of both a restore and a rebuild, forgets the path too.
//! - **Another process writing.** The stamp is read from the file, not from
//! anything this process did, so a scan in a second instance moves it just
//! the same.
//!
//! # Why the stamp is taken before the backfill
//!
//! The backfill adds versions and terms itself, so a stamp read afterwards
//! would describe its own writes. Read afterwards it could also describe an
//! image another connection inserted between the backfill's read and the
//! stamp's — and record that image as covered when it was not. Read before,
//! the worst case is the reverse: the backfill's own inserts move the stamp,
//! and the next open runs one more backfill that finds nothing. That costs one
//! redundant pass after a backfill that did real work, and never misses a row.
//!
//! # What it does not see
//!
//! An `UPDATE` that creates work without adding a row. None of this build's
//! writers does: a scan's move of a file is a new `source_ref` and so a new
//! image, and uuids are only rewritten by the backfill itself. Should one
//! appear, the cost is that its repair waits for the next insert or the next
//! start of the app — which is exactly where the backfill ran before it ran on
//! every open.
//!
//! Kept in memory rather than in the catalog on purpose: a row in the file
//! would travel in the sync snapshot and would need a table an older build
//! does not have, and a flag that another device's catalog carried in would
//! say nothing about this one.
use std::collections::HashMap;
use std::path::{Path, PathBuf};
use std::sync::{Mutex, OnceLock};
use rusqlite::Connection;
use crate::error::CatalogError;
/// What a catalog looked like, as far as the backfill cares.
#[derive(Debug, Clone, PartialEq, Eq)]
pub(crate) struct Stamp {
/// Device and inode, so a file swapped in under the same name is a new
/// catalog. `None` where the platform has no such thing.
file: Option<(u64, u64)>,
user_version: i64,
/// The newest image: id, path, when added, and whether it has a version.
last_image: Option<String>,
/// The newest version: id and the image it belongs to.
last_version: Option<String>,
/// The newest keyword assignment: rowid, version and word.
last_keyword: Option<String>,
}
/// The stamp recorded at the last backfill, per catalog file.
fn done() -> &'static Mutex<HashMap<PathBuf, Stamp>> {
static DONE: OnceLock<Mutex<HashMap<PathBuf, Stamp>>> = OnceLock::new();
DONE.get_or_init(Default::default)
}
/// One name per file, whichever spelling of its path the caller used.
fn key(path: &Path) -> PathBuf {
std::fs::canonicalize(path).unwrap_or_else(|_| path.to_path_buf())
}
/// Read the stamp of the catalog behind `conn`, which was opened from `path`.
///
/// One statement: the last row of each of three tables, each found by
/// descending its rowid b-tree to the last page, plus one probe of
/// `versions_image` for the newest image — and a `stat` of the file.
pub(crate) fn stamp(conn: &Connection, path: &Path) -> Result<Stamp, CatalogError> {
let (user_version, last_image, last_version, last_keyword) = conn.query_row(
"SELECT (SELECT user_version FROM pragma_user_version),
(SELECT printf('%d|%d|%d|%s', i.id, i.added_at,
EXISTS (SELECT 1 FROM versions v WHERE v.image_id = i.id),
i.source_ref)
FROM images i ORDER BY i.id DESC LIMIT 1),
(SELECT printf('%d|%d', id, image_id)
FROM versions ORDER BY id DESC LIMIT 1),
(SELECT printf('%d|%d|%s', rowid, version_id, keyword)
FROM keywords ORDER BY rowid DESC LIMIT 1)",
[],
|r| Ok((r.get(0)?, r.get(1)?, r.get(2)?, r.get(3)?)),
)?;
Ok(Stamp {
file: file_identity(path),
user_version,
last_image,
last_version,
last_keyword,
})
}
#[cfg(unix)]
fn file_identity(path: &Path) -> Option<(u64, u64)> {
use std::os::unix::fs::MetadataExt;
std::fs::metadata(path).ok().map(|m| (m.dev(), m.ino()))
}
#[cfg(not(unix))]
fn file_identity(_path: &Path) -> Option<(u64, u64)> {
None
}
/// Whether the catalog at `path` was last backfilled at exactly `stamp`.
pub(crate) fn is_current(path: &Path, stamp: &Stamp) -> bool {
done()
.lock()
.unwrap_or_else(|e| e.into_inner())
.get(&key(path))
== Some(stamp)
}
/// Record that the catalog at `path` has been backfilled as of `stamp`.
pub(crate) fn record(path: &Path, stamp: Stamp) {
done()
.lock()
.unwrap_or_else(|e| e.into_inner())
.insert(key(path), stamp);
}
/// Make the next open of `path` backfill, whatever its stamp says.
///
/// For the writers that know they have changed the catalog wholesale — a
/// merge of a pulled catalog, a restore from backup — so their correctness
/// does not rest on the stamp happening to move.
pub(crate) fn forget(path: &Path) {
done()
.lock()
.unwrap_or_else(|e| e.into_inner())
.remove(&key(path));
}
#[cfg(test)]
mod tests {
use crate::rating::derived_version_uuid;
use crate::Catalog;
use std::path::PathBuf;
/// A catalog file of its own, holding one image the server has named,
/// backfilled and settled.
///
/// Opened three times on the way: to create it; after the image went in,
/// which gives the image its default version; and once more, because that
/// version moved the stamp and the next open runs the one redundant pass
/// the module header describes. After that the stamp stands still.
fn catalog(tag: &str) -> PathBuf {
let dir = std::env::temp_dir().join(format!(
"dr-backfilled-{tag}-{}-{:?}",
std::process::id(),
std::thread::current().id()
));
let _ = std::fs::remove_dir_all(&dir);
std::fs::create_dir_all(&dir).unwrap();
let path = dir.join("catalog.sqlite");
{
let cat = Catalog::open(&path).unwrap();
let c = cat.connection();
c.execute_batch(
"INSERT INTO roots(id, kind, label) VALUES (1, 'remote', 'Photos');
INSERT INTO images(id, root_id, source_ref, added_at)
VALUES (1, 1, 'Photos/a.CR3', 0);
INSERT INTO remote(image_id, file_id) VALUES (1, 77);",
)
.unwrap();
}
assert_eq!(uuid(&path, 1), Some(derived_version_uuid(77)));
assert_eq!(uuid(&path, 1), Some(derived_version_uuid(77)));
path
}
/// The default version's uuid for `image`, read through an ordinary open.
fn uuid(path: &std::path::Path, image: i64) -> Option<String> {
let cat = Catalog::open(path).unwrap();
cat.connection()
.query_row(
"SELECT uuid FROM versions WHERE image_id = ?1 AND is_default = 1",
[image],
|r| r.get(0),
)
.ok()
}
/// Put the one row back into the state the backfill repairs, with an
/// `UPDATE` — which moves none of the stamp's maxima, so only the stamp's
/// other parts or an explicit `forget` can bring the backfill back.
fn unalign(path: &std::path::Path) {
rusqlite::Connection::open(path)
.unwrap()
.execute("UPDATE versions SET uuid = 'minted' WHERE image_id = 1", [])
.unwrap();
}
#[test]
fn an_unchanged_catalog_is_not_backfilled_again() {
let path = catalog("unchanged");
unalign(&path);
assert_eq!(
uuid(&path, 1).as_deref(),
Some("minted"),
"nothing was added since the last backfill, so the open skipped it"
);
}
#[test]
fn an_image_a_scan_added_is_backfilled_on_the_next_open() {
let path = catalog("scanned");
rusqlite::Connection::open(&path)
.unwrap()
.execute(
"INSERT INTO images(id, root_id, source_ref, added_at)
VALUES (2, 1, 'Photos/b.CR3', 0)",
[],
)
.unwrap();
assert!(uuid(&path, 2).is_some(), "the new image got its version");
}
/// The id of a deleted newest row is handed out again, so `max(id)` is
/// the same before and after — the sequence emptying the trash and then
/// scanning makes. The image that took the id still needs its version.
///
/// A virtual copy on the older image holds the newest version id, so
/// the deletion does not move `max(versions.id)` either: nothing the old
/// stamp read changes, which is the case that went unrepaired.
#[test]
fn an_image_that_reuses_a_deleted_id_is_backfilled_on_the_next_open() {
let path = catalog("reused");
rusqlite::Connection::open(&path)
.unwrap()
.execute(
"INSERT INTO images(id, root_id, source_ref, added_at)
VALUES (2, 1, 'Photos/b.CR3', 0)",
[],
)
.unwrap();
assert!(uuid(&path, 2).is_some());
rusqlite::Connection::open(&path)
.unwrap()
.execute(
"INSERT INTO versions(image_id, uuid, name, is_default)
VALUES (1, 'copy', 'Crop', 0)",
[],
)
.unwrap();
// Settle: one open backfills after the new version, one more runs
// the redundant pass and records the stamp that stands.
assert!(uuid(&path, 2).is_some());
assert!(uuid(&path, 2).is_some());
let c = rusqlite::Connection::open(&path).unwrap();
let before: (i64, i64) = c
.query_row(
"SELECT (SELECT max(id) FROM images), (SELECT max(id) FROM versions)",
[],
|r| Ok((r.get(0)?, r.get(1)?)),
)
.unwrap();
c.execute("DELETE FROM images WHERE id = 2", []).unwrap();
c.execute(
"INSERT INTO images(root_id, source_ref, added_at)
VALUES (1, 'Photos/c.CR3', 0)",
[],
)
.unwrap();
let after: (i64, i64) = c
.query_row(
"SELECT (SELECT max(id) FROM images), (SELECT max(id) FROM versions)",
[],
|r| Ok((r.get(0)?, r.get(1)?)),
)
.unwrap();
assert_eq!(before, after, "SQLite handed the freed id out again");
drop(c);
assert!(uuid(&path, 2).is_some(), "the new image got its version");
}
/// The same, with the same file coming back at the same id in the same
/// second: path and time match, and only the missing version tells.
#[test]
fn the_same_file_back_at_the_same_id_is_backfilled_on_the_next_open() {
let path = catalog("returned");
let c = rusqlite::Connection::open(&path).unwrap();
// As above: a newer version on another image keeps the deletion
// from moving `max(versions.id)`.
c.execute_batch(
"INSERT INTO images(id, root_id, source_ref, added_at)
VALUES (0, 1, 'Photos/0.CR3', 0);
INSERT INTO versions(image_id, uuid, name, is_default)
VALUES (0, 'copy', 'Crop', 0);",
)
.unwrap();
drop(c);
assert!(uuid(&path, 1).is_some());
assert!(uuid(&path, 1).is_some());
let c = rusqlite::Connection::open(&path).unwrap();
c.execute_batch(
"DELETE FROM images WHERE id = 1;
INSERT INTO images(id, root_id, source_ref, added_at)
VALUES (1, 1, 'Photos/a.CR3', 0);
INSERT INTO remote(image_id, file_id) VALUES (1, 77);",
)
.unwrap();
drop(c);
assert_eq!(uuid(&path, 1), Some(derived_version_uuid(77)));
}
#[test]
fn the_first_open_after_a_migration_backfills() {
let path = catalog("migrated");
unalign(&path);
rusqlite::Connection::open(&path)
.unwrap()
.pragma_update(None, "user_version", crate::schema::SCHEMA_VERSION - 1)
.unwrap();
assert_eq!(uuid(&path, 1), Some(derived_version_uuid(77)));
}
#[test]
fn the_first_open_after_a_pulled_catalog_backfills() {
let path = catalog("pulled");
// The remote is this catalog as it stands, so every row the merge
// offers collides and nothing in the stamp moves: only the merge
// saying so can make the next open backfill.
let remote = path.with_file_name("remote.sqlite");
rusqlite::Connection::open(&path)
.unwrap()
.execute("VACUUM INTO ?1", [remote.to_string_lossy().as_ref()])
.unwrap();
unalign(&path);
assert_eq!(uuid(&path, 1).as_deref(), Some("minted"));
Catalog::open(&path)
.unwrap()
.merge_remote_catalog(&remote)
.unwrap();
assert_eq!(uuid(&path, 1), Some(derived_version_uuid(77)));
}
#[cfg(unix)]
#[test]
fn a_catalog_replaced_under_the_same_name_backfills() {
let path = catalog("replaced");
unalign(&path);
assert_eq!(uuid(&path, 1).as_deref(), Some("minted"));
// Every connection is closed, so the WAL is folded in and the main
// file is the whole catalog. A copy renamed over it is the same rows
// in a new file — which is what a restore or a copied-in catalog is.
let copy = path.with_file_name("copy.sqlite");
std::fs::copy(&path, &copy).unwrap();
std::fs::rename(&copy, &path).unwrap();
assert_eq!(uuid(&path, 1), Some(derived_version_uuid(77)));
}
}