Keywords are catalog state, and the catalog syncs. Without this, two devices keywording the same library would resolve to whichever synced last, and an afternoon of work would vanish with no sign it had ever happened. The vocabulary merges per row on the rule collections already use: revision first, timestamp only to break a tie, so a device with a skewed clock cannot win by having the wrong idea of the time. Assignments merge as a set union, which is FR-NC-9's principle applied to metadata instead of edit nodes — disjoint work survives on both sides. Three things needed care and are commented where they happen: A deletion travels *by name*, not by identity. Both devices may have minted their own uuid for one word before they ever synced, so deleting by uuid would tombstone a row nothing was assigned to and leave every photograph still carrying the word. The union then refuses to readmit a word a winning tombstone has just removed — without that filter the remote's live assignments would resurrect it on the very same pass. Images are resolved by the server's file id first and the content hash second. Membership has always used the hash alone, but the hash is computed only when import dedup or a reconnect asks for it, which for most libraries is never — so a hash-only union would have quietly done nothing for the ordinary photograph. A word lands on the local default version. Version uuids do not reconcile in the catalog at all: ensure_default_versions mints a fresh one per device, so a uuid-keyed join would have unioned nothing. Removal still does not propagate. That is the trade collection membership already makes, for the same reason — an unwanted keyword is removed again in a second, and a silently lost afternoon is not recoverable at all. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
254 lines
9.0 KiB
Rust
254 lines
9.0 KiB
Rust
//! TRACES: FR-CAT-7 | FR-NC-9 | NFR-R1
|
|
//! Preparing the catalog file for upload, and taking in a remote one.
|
|
//!
|
|
//! # The hazard this module exists to handle
|
|
//!
|
|
//! A WAL-mode SQLite database is not one file. Committed transactions can live
|
|
//! in `catalog.sqlite-wal` with the main file lagging behind, so copying
|
|
//! `catalog.sqlite` alone uploads a **torn snapshot**: internally consistent as
|
|
//! of some older point, missing everything since. Worse, a naive copy taken
|
|
//! while a writer is mid-transaction can be structurally corrupt.
|
|
//!
|
|
//! So an upload never copies the live file. It runs a TRUNCATE checkpoint to
|
|
//! fold the WAL back into the main file, then uses SQLite's own backup API to
|
|
//! take a consistent snapshot — which serialises correctly against concurrent
|
|
//! writers rather than racing them.
|
|
//!
|
|
//! # What is actually synced
|
|
//!
|
|
//! Only the *user's judgements about their library* merge: collections, and the
|
|
//! keyword vocabulary with its assignments (see [`crate::merge`]). The rest of
|
|
//! the catalog is a *local index* of *local* storage — folder mtimes, cache
|
|
//! paths, job rows — and copying another device's version of those in would be
|
|
//! actively wrong. The remote file is read for those two and then discarded.
|
|
//!
|
|
//! This is why the catalog remains disposable in the ARCH §6.12 sense: nothing
|
|
//! here makes the local database authoritative for anything a rebuild could
|
|
//! not recover.
|
|
|
|
use std::path::{Path, PathBuf};
|
|
|
|
use rusqlite::Connection;
|
|
|
|
use crate::error::CatalogError;
|
|
use crate::merge::{self, MergeReport};
|
|
|
|
/// Schema name the downloaded remote catalog is attached under.
|
|
const REMOTE_SCHEMA: &str = "remote_cat";
|
|
|
|
/// Fold the WAL into the main database file.
|
|
///
|
|
/// TRUNCATE rather than PASSIVE: passive checkpointing gives up when a reader
|
|
/// holds the WAL open, which would leave recent commits out of the snapshot
|
|
/// without saying so.
|
|
pub fn checkpoint(conn: &Connection) -> Result<(), CatalogError> {
|
|
conn.pragma_update(None, "wal_checkpoint", "TRUNCATE")?;
|
|
Ok(())
|
|
}
|
|
|
|
/// Write a consistent snapshot of the catalog to `dest`, ready to upload.
|
|
///
|
|
/// Uses the backup API rather than a filesystem copy so the snapshot is
|
|
/// coherent even with writers active. Callers should still prefer a quiet
|
|
/// moment — this competes with background jobs for the write lock.
|
|
pub fn snapshot_for_upload(conn: &Connection, dest: &Path) -> Result<(), CatalogError> {
|
|
checkpoint(conn)?;
|
|
|
|
let mut out = Connection::open(dest)?;
|
|
let backup = rusqlite::backup::Backup::new(conn, &mut out)?;
|
|
// SQLite's own "copy everything" sentinel is -1, but rusqlite asserts a
|
|
// positive page count, so ask for more pages than a catalog will ever
|
|
// have. The effect is the same: one step, no interleaved writers, no
|
|
// progress callback. A 50k-image catalog is tens of megabytes.
|
|
backup.run_to_completion(i32::MAX, std::time::Duration::ZERO, None)?;
|
|
Ok(())
|
|
}
|
|
|
|
/// Whether a downloaded remote catalog is worth merging.
|
|
///
|
|
/// Cheap guard before attaching: a remote written by a newer build may contain
|
|
/// tables and columns this one cannot read, and attempting the merge would
|
|
/// fail mid-transaction rather than declining cleanly.
|
|
pub fn remote_is_mergeable(remote: &Path) -> Result<bool, CatalogError> {
|
|
let conn = Connection::open_with_flags(
|
|
remote,
|
|
rusqlite::OpenFlags::SQLITE_OPEN_READ_ONLY | rusqlite::OpenFlags::SQLITE_OPEN_NO_MUTEX,
|
|
)?;
|
|
let v: i64 = conn.query_row("PRAGMA user_version", [], |r| r.get(0))?;
|
|
Ok(v <= crate::schema::SCHEMA_VERSION)
|
|
}
|
|
|
|
/// Attach a downloaded remote catalog, merge its collections, detach.
|
|
///
|
|
/// The remote file is opened **read-only** — this device never writes to
|
|
/// another device's catalog, it only reads collections out of it.
|
|
pub fn merge_remote(conn: &Connection, remote: &Path) -> Result<MergeReport, CatalogError> {
|
|
if !remote_is_mergeable(remote)? {
|
|
return Err(CatalogError::SchemaTooNew {
|
|
found: -1,
|
|
supported: crate::schema::SCHEMA_VERSION,
|
|
});
|
|
}
|
|
|
|
// Path binds as a parameter; ATTACH accepts one, so a path containing a
|
|
// quote cannot break out into SQL.
|
|
conn.execute(
|
|
&format!("ATTACH DATABASE ?1 AS {REMOTE_SCHEMA}"),
|
|
[remote.to_string_lossy().as_ref()],
|
|
)?;
|
|
|
|
let result = merge::merge_all(conn);
|
|
|
|
// Detach even if the merge failed, or the next attempt errors with
|
|
// "database remote_cat is already in use".
|
|
let detach = conn.execute(&format!("DETACH DATABASE {REMOTE_SCHEMA}"), []);
|
|
if let Err(e) = detach {
|
|
log::warn!("failed to detach remote catalog: {e}");
|
|
}
|
|
|
|
result
|
|
}
|
|
|
|
/// Where the catalog snapshot and the downloaded remote live.
|
|
///
|
|
/// Both are transient working files, not the catalog itself, so they belong in
|
|
/// the cache directory rather than beside the live database.
|
|
#[derive(Debug, Clone)]
|
|
pub struct SyncPaths {
|
|
pub upload_snapshot: PathBuf,
|
|
pub downloaded_remote: PathBuf,
|
|
}
|
|
|
|
impl SyncPaths {
|
|
pub fn in_dir(cache_dir: &Path) -> Self {
|
|
SyncPaths {
|
|
upload_snapshot: cache_dir.join("catalog-upload.sqlite"),
|
|
downloaded_remote: cache_dir.join("catalog-remote.sqlite"),
|
|
}
|
|
}
|
|
}
|
|
|
|
#[cfg(test)]
|
|
mod tests {
|
|
use super::*;
|
|
use crate::schema;
|
|
|
|
fn seeded(path: &Path) -> Connection {
|
|
let c = Connection::open(path).unwrap();
|
|
schema::configure(&c).unwrap();
|
|
schema::migrate(&c).unwrap();
|
|
c
|
|
}
|
|
|
|
#[test]
|
|
fn snapshot_captures_committed_data() {
|
|
let dir = tempdir();
|
|
let live = dir.join("catalog.sqlite");
|
|
let snap = dir.join("snap.sqlite");
|
|
|
|
let c = seeded(&live);
|
|
c.execute(
|
|
"INSERT INTO collections(uuid, name, kind, created, revision, modified)
|
|
VALUES ('u1', 'Iceland', 0, 0, 1, 1)",
|
|
[],
|
|
)
|
|
.unwrap();
|
|
|
|
snapshot_for_upload(&c, &snap).unwrap();
|
|
|
|
// The snapshot must hold the row even though it was written after the
|
|
// database was created — the torn-file failure this guards against.
|
|
let s = Connection::open(&snap).unwrap();
|
|
let name: String = s
|
|
.query_row("SELECT name FROM collections", [], |r| r.get(0))
|
|
.unwrap();
|
|
assert_eq!(name, "Iceland");
|
|
}
|
|
|
|
#[test]
|
|
fn a_remote_from_a_newer_build_is_declined_not_attempted() {
|
|
let dir = tempdir();
|
|
let remote = dir.join("remote.sqlite");
|
|
let r = seeded(&remote);
|
|
r.pragma_update(None, "user_version", schema::SCHEMA_VERSION + 1)
|
|
.unwrap();
|
|
drop(r);
|
|
|
|
assert!(!remote_is_mergeable(&remote).unwrap());
|
|
|
|
let local = seeded(&dir.join("local.sqlite"));
|
|
assert!(matches!(
|
|
merge_remote(&local, &remote),
|
|
Err(CatalogError::SchemaTooNew { .. })
|
|
));
|
|
}
|
|
|
|
#[test]
|
|
fn merge_remote_round_trips_a_collection() {
|
|
let dir = tempdir();
|
|
let remote_path = dir.join("remote.sqlite");
|
|
{
|
|
let r = seeded(&remote_path);
|
|
r.execute(
|
|
"INSERT INTO collections(uuid, name, kind, created, revision, modified)
|
|
VALUES ('u-remote', 'Portugal', 0, 0, 1, 1)",
|
|
[],
|
|
)
|
|
.unwrap();
|
|
checkpoint(&r).unwrap();
|
|
}
|
|
|
|
let local = seeded(&dir.join("local.sqlite"));
|
|
local
|
|
.execute(
|
|
"INSERT INTO collections(uuid, name, kind, created, revision, modified)
|
|
VALUES ('u-local', 'Iceland', 0, 0, 1, 1)",
|
|
[],
|
|
)
|
|
.unwrap();
|
|
|
|
let report = merge_remote(&local, &remote_path).unwrap();
|
|
assert_eq!(report.inserted, 1);
|
|
|
|
let n: i64 = local
|
|
.query_row("SELECT count(*) FROM collections", [], |r| r.get(0))
|
|
.unwrap();
|
|
assert_eq!(n, 2);
|
|
}
|
|
|
|
#[test]
|
|
fn the_remote_can_be_merged_twice_without_attach_conflict() {
|
|
// Detach must happen even on the failure path, or the second attempt
|
|
// errors with "database remote_cat is already in use".
|
|
let dir = tempdir();
|
|
let remote_path = dir.join("remote.sqlite");
|
|
{
|
|
let r = seeded(&remote_path);
|
|
r.execute(
|
|
"INSERT INTO collections(uuid, name, kind, created, revision, modified)
|
|
VALUES ('u-remote', 'Portugal', 0, 0, 1, 1)",
|
|
[],
|
|
)
|
|
.unwrap();
|
|
checkpoint(&r).unwrap();
|
|
}
|
|
let local = seeded(&dir.join("local.sqlite"));
|
|
|
|
merge_remote(&local, &remote_path).unwrap();
|
|
let second = merge_remote(&local, &remote_path).unwrap();
|
|
assert!(!second.local_changed());
|
|
}
|
|
|
|
/// A scratch directory that cleans up with the test.
|
|
fn tempdir() -> PathBuf {
|
|
let base = std::env::temp_dir().join(format!(
|
|
"dr-catalog-test-{}-{:?}",
|
|
std::process::id(),
|
|
std::thread::current().id()
|
|
));
|
|
let _ = std::fs::remove_dir_all(&base);
|
|
std::fs::create_dir_all(&base).unwrap();
|
|
base
|
|
}
|
|
}
|