Files
DarkRoom/core/dr-catalog/src/schema.rs
T
dtourolleandClaude Opus 5 03326242a1
Build and test / Desktop (Linux) (push) Successful in 1h23m26s
Build and test / Android (aarch64) (push) Failing after 2s
Build and test / Layer separation (push) Successful in 50s
Traceability / Requirement traces (push) Failing after 1m9s
Make the CI checks say what they mean, and format the workspace
The Android job's "Verify minimum API level" step has never verified the
minimum API level. It took the first `*.so` anywhere under the target
directory, which is a host proc-macro from debug/deps — an x86-64 object
built by the runner's gcc, whose .comment section cannot mention Android
and so can never contradict the expected value. It now reads the artifact
under the target triple, compares against MIN_API parsed from the
Dockerfile rather than a second copy of the number, and fails on a
mismatch. Both sides are checked non-empty first: two failed parses would
otherwise compare equal and pass, which is the same silent success in a
new costume.

The Android image installs one SDK package per layer and keeps the
output. sdkmanager is a JVM program that aborts when it cannot get memory,
and the single `> /dev/null` step reported that as a bare "exit code 134"
while a retry re-downloaded everything that had already succeeded.

tools/ci-local.sh runs all four jobs — desktop, android, layering,
traceability — against the host toolchain, which is pinned to the same
1.92.0 CI installs. Its matrix check compares regeneration against the
working tree rather than against HEAD: CI starts from a clean checkout, so
git's answer is the right one there and reports every local run stale here.

The rest is rustfmt across the workspace, and the clippy findings that
surfaced once it did: manual_contains in dr-thumbs and collections_ui, a
map iterated as pairs for its keys, an index loop over a slice, and two
runtime assertions on a constant now made at compile time.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 12:02:51 +02:00

757 lines
28 KiB
Rust

//! TRACES: FR-CAT-2 | NFR-R5
//! Schema definition and forward-only migrations.
//!
//! The catalog is an *index*, not a source of truth (ARCH §6.12) — it is
//! deletable and rebuildable from sources plus sidecars. That is what makes
//! migration failure survivable, and why the recovery path is the normal
//! mechanism rather than a last resort.
//!
//! Migrations are forward-only, transactional, and idempotent on retry
//! (NFR-R5). The app refuses to open a catalog newer than it understands
//! rather than corrupting it.
use rusqlite::Connection;
use crate::error::CatalogError;
/// Schema version this build writes and understands.
pub const SCHEMA_VERSION: i64 = 5;
/// Apply migrations up to [`SCHEMA_VERSION`].
///
/// Returns the version migrated from, so callers can log or back up before a
/// real migration (NFR-R2 requires a backup before schema change).
pub fn migrate(conn: &Connection) -> Result<i64, CatalogError> {
let from: i64 = conn.query_row("PRAGMA user_version", [], |r| r.get(0))?;
if from > SCHEMA_VERSION {
return Err(CatalogError::SchemaTooNew {
found: from,
supported: SCHEMA_VERSION,
});
}
if from == SCHEMA_VERSION {
return Ok(from);
}
// Each step runs in its own transaction so a failure leaves the catalog
// at a coherent version rather than half-migrated.
if from < 1 {
let tx = conn.unchecked_transaction()?;
tx.execute_batch(V1)?;
tx.pragma_update(None, "user_version", 1)?;
tx.commit()?;
}
if from < 2 {
let tx = conn.unchecked_transaction()?;
tx.execute_batch(V2)?;
tx.pragma_update(None, "user_version", 2)?;
tx.commit()?;
}
if from < 3 {
let tx = conn.unchecked_transaction()?;
tx.execute_batch(V3)?;
tx.pragma_update(None, "user_version", 3)?;
tx.commit()?;
}
if from < 4 {
let tx = conn.unchecked_transaction()?;
tx.execute_batch(V4)?;
tx.pragma_update(None, "user_version", 4)?;
tx.commit()?;
}
if from < 5 {
let tx = conn.unchecked_transaction()?;
tx.execute_batch(V5)?;
tx.pragma_update(None, "user_version", 5)?;
tx.commit()?;
}
Ok(from)
}
/// Recompute columns a migration added, for rows that predate it.
///
/// A migration adds a column with a default; it cannot know what the value
/// *should* be for the rows already present. Without a backfill those rows are
/// silently partial — present, queryable, and wrong — which is worse than
/// missing, because nothing signals that they need attention.
///
/// Cheap enough to run on every open: each pass is one indexed UPDATE, and
/// re-running it is a no-op once the values are already right.
///
/// Returns how many rows each backfill touched, for logging.
pub fn backfill(conn: &Connection) -> Result<Vec<(&'static str, usize)>, CatalogError> {
let mut out = Vec::new();
// v2: `shadowed_by`. A JPEG sitting beside a RAW of the same name is the
// camera's own rendering of that frame, not a second photograph, so it is
// hidden from the grid, the timeline and the sweep.
let n = pair_raw_and_jpeg(conn)?;
if n > 0 {
out.push(("shadowed_by", n));
}
// v3: every image needs a default version to carry its rating and flag.
// Libraries scanned before ratings existed have images and no versions at
// all, so there was nowhere for a judgement to go — see
// [`crate::rating`]. Backfilled rather than migrated in SQL because the
// UUID per row is the cross-device merge identity and must be generated,
// not derived.
let n = crate::rating::ensure_default_versions(conn)?;
if n > 0 {
out.push(("default_versions", n));
}
Ok(out)
}
/// Connection setup applied on every open, migration or not.
///
/// WAL is required by NFR-R1: it survives power loss without corruption, and
/// it lets a background job write while the grid reads.
pub fn configure(conn: &Connection) -> Result<(), CatalogError> {
conn.pragma_update(None, "journal_mode", "WAL")?;
// NORMAL rather than FULL: with WAL this is durable across process death
// (which is what FR-PLAT-AND-3 cares about) and only risks the last
// transaction on power loss. The catalog is rebuildable; the sidecars are
// not, and they are written separately with their own fsync discipline.
conn.pragma_update(None, "synchronous", "NORMAL")?;
conn.pragma_update(None, "foreign_keys", true)?;
// A scan touching thousands of rows is transient; let SQLite spill to
// memory rather than materialising temp b-trees on disk.
conn.pragma_update(None, "temp_store", "MEMORY")?;
Ok(())
}
/// The v1 schema rewritten to target an attached database.
///
/// Needed because a downloaded remote catalog is `ATTACH`ed under its own
/// schema name before merging, and tests build one from scratch. SQLite has no
/// "create these tables over there" form, so the names are rewritten.
///
/// The rewrite is textual and therefore only as good as the naming discipline
/// in [`V1`]: every `CREATE TABLE`/`CREATE INDEX` must name its object
/// unqualified, which they do.
pub fn v1_for_attached(schema_name: &str) -> String {
V1.replace("CREATE TABLE ", &format!("CREATE TABLE {schema_name}."))
.replace("CREATE INDEX ", &format!("CREATE INDEX {schema_name}."))
.replace(
"CREATE UNIQUE INDEX ",
&format!("CREATE UNIQUE INDEX {schema_name}."),
)
// REFERENCES within an attached schema resolve to that schema already,
// so foreign keys need no rewriting — but the ON clause of an index
// does, and `CREATE INDEX x.name ON table` is the correct form.
}
/// Mark each JPEG that sits beside a RAW of the same name.
///
/// Matched on folder plus stem, case-insensitively. Same-folder is what makes
/// this safe: cameras write the pair side by side, and matching across folders
/// would risk pairing unrelated frames, since camera filenames wrap at
/// IMG_9999 (FR-CAT-11).
///
/// Done in Rust rather than SQL because the comparison needs a filename stem,
/// and SQLite has no such function without enabling `rusqlite/functions` —
/// a dependency feature for one string operation, whose SQL spelling would be
/// unreadable and would mishandle names with no extension.
fn pair_raw_and_jpeg(conn: &Connection) -> Result<usize, CatalogError> {
use std::collections::HashMap;
// (folder, lowercase stem) -> RAW id, built in one pass over the RAWs.
let mut raws: HashMap<(Option<i64>, String), i64> = HashMap::new();
{
let mut stmt = conn.prepare(
"SELECT id, folder_id, source_ref FROM images
WHERE lower(format) IN
('cr2','cr3','nef','arw','raf','rw2','orf','dng')",
)?;
let rows = stmt.query_map([], |r| {
Ok((
r.get::<_, i64>(0)?,
r.get::<_, Option<i64>>(1)?,
r.get::<_, String>(2)?,
))
})?;
for row in rows {
let (id, folder, path) = row?;
raws.insert((folder, stem_of(&path).to_ascii_lowercase()), id);
}
}
if raws.is_empty() {
return Ok(0);
}
let pairs: Vec<(i64, i64)> = {
let mut stmt = conn.prepare(
"SELECT id, folder_id, source_ref FROM images
WHERE lower(format) IN ('jpg','jpeg') AND shadowed_by IS NULL",
)?;
let rows = stmt.query_map([], |r| {
Ok((
r.get::<_, i64>(0)?,
r.get::<_, Option<i64>>(1)?,
r.get::<_, String>(2)?,
))
})?;
rows.filter_map(|row| {
let (id, folder, path) = row.ok()?;
let raw = raws.get(&(folder, stem_of(&path).to_ascii_lowercase()))?;
Some((id, *raw))
})
.collect()
};
let tx = conn.unchecked_transaction()?;
for (jpeg, raw) in &pairs {
tx.execute(
"UPDATE images SET shadowed_by = ?2 WHERE id = ?1",
[jpeg, raw],
)?;
}
tx.commit()?;
Ok(pairs.len())
}
/// A filename without its extension.
///
/// Only the final path component, and only its last dot — a directory
/// containing a dot must not truncate the name.
fn stem_of(path: &str) -> &str {
let name = path.rsplit(['/', ':']).next().unwrap_or(path);
match name.rsplit_once('.') {
Some((stem, _)) if !stem.is_empty() => stem,
_ => name,
}
}
const V5: &str = r#"
-- TRACES: FR-NC-6a | FR-CAT-9 | NFR-RES-4
-- Offline availability: what is kept, why it is kept, and where it lives.
--
-- `pinned` separates a promise from a convenience, and the distinction has to
-- be a *column* rather than something inferred from `pinned_by_rule`. A pin is
-- the user saying "this collection comes with me"; a passively cached original
-- is the app noticing they opened something. Only the second is evictable, so
-- the eviction query has to be able to ask the question directly — and it has
-- to keep answering correctly for an image whose pinning rule was since
-- deleted, which `pinned_by_rule` alone cannot do because it is
-- ON DELETE SET NULL.
ALTER TABLE image_cache ADD COLUMN pinned INTEGER NOT NULL DEFAULT 0;
-- Where the cached original actually is, relative to the cache directory.
-- Relative rather than absolute: the library moves between machines and
-- between an app sandbox and a user directory, and an absolute path baked in
-- at download time would break on every one of those.
ALTER TABLE image_cache ADD COLUMN path TEXT;
-- Eviction reads exactly this: unpinned rows, oldest use first. Partial on
-- `pinned = 0` because pinned rows are never candidates and including them
-- would make the index proportional to the whole library rather than to the
-- passive cache.
CREATE INDEX image_cache_evictable ON image_cache(last_used)
WHERE pinned = 0;
"#;
const V4: &str = r#"
-- TRACES: FR-CAT-15
-- Soft delete. A trashed image is a real file that has been *moved* to a trash
-- folder under the library root, not a row hidden by a flag: the catalog is a
-- rebuildable index (ARCH §6.12), so a flag alone would evaporate the moment
-- the catalog was deleted and every trashed photograph would return.
--
-- `source_ref` follows the file to its new path, because that is where the bytes
-- now are and every fetch resolves through it. `trashed_from` remembers where it
-- came from, which is the only way a restore can put it back — the trash is flat
-- and the original folder structure is not recoverable from the trashed path.
ALTER TABLE images ADD COLUMN trashed_at INTEGER;
ALTER TABLE images ADD COLUMN trashed_from TEXT;
-- Partial: almost no rows are trashed, and the grid's "not trashed" predicate is
-- answered by the absence of an entry rather than by scanning every image.
CREATE INDEX images_trashed ON images(trashed_at) WHERE trashed_at IS NOT NULL;
"#;
const V3: &str = r#"
-- Ratings and flags are read per grid window and counted for the filter bar's
-- histogram, both of which key on the *default* version. Without this the
-- histogram is a full scan of `versions` on every judgement.
--
-- Partial on `is_default`: a virtual copy's rating is real but is never what
-- these two queries ask for, and excluding them keeps the index roughly one
-- entry per image rather than one per version.
CREATE INDEX versions_judgement ON versions(image_id, rating, flag)
WHERE is_default = 1;
"#;
const V2: &str = r#"
-- A JPEG the camera wrote alongside a RAW of the same name is that RAW's own
-- rendering, not a second photograph. Recording *which* RAW shadows it, rather
-- than a bare flag, keeps the relationship usable: the JPEG is a ready-made
-- preview for its RAW, and the pairing can be undone without a rescan.
ALTER TABLE images ADD COLUMN shadowed_by INTEGER REFERENCES images(id) ON DELETE SET NULL;
CREATE INDEX images_shadowed ON images(shadowed_by) WHERE shadowed_by IS NOT NULL;
"#;
const V1: &str = r#"
-- Roots -------------------------------------------------------------------
CREATE TABLE roots (
id INTEGER PRIMARY KEY,
kind TEXT NOT NULL, -- 'local' | 'saf' | 'remote'
grant_blob BLOB, -- SAF persisted permission; NULL on Linux
label TEXT NOT NULL,
last_seen INTEGER,
-- Bumped once per completed scan. Folders record the generation they were
-- reached in; anything older was not reached and no longer exists.
scan_generation INTEGER NOT NULL DEFAULT 0,
-- One row per granted location. Without this, a rescan inserts a second
-- root for the same folder and the library silently fragments across
-- them — images split between roots, and pruning compares against the
-- wrong generation.
UNIQUE(kind, label)
);
-- Folders: the unit of change detection, local and remote alike -----------
CREATE TABLE folders (
id INTEGER PRIMARY KEY,
root_id INTEGER NOT NULL REFERENCES roots(id) ON DELETE CASCADE,
parent_id INTEGER REFERENCES folders(id) ON DELETE CASCADE,
path TEXT NOT NULL,
-- Remote: the propagating ETag that makes a no-op sync one request.
etag TEXT,
-- Local: directory mtime plus direct-entry count. mtime alone misses a
-- paired create+delete inside one timestamp tick; the count narrows that.
mtime INTEGER,
entry_count INTEGER,
scanned_generation INTEGER NOT NULL DEFAULT 0,
UNIQUE(root_id, path)
);
CREATE INDEX folders_parent ON folders(parent_id);
-- Images ------------------------------------------------------------------
CREATE TABLE images (
id INTEGER PRIMARY KEY,
root_id INTEGER NOT NULL REFERENCES roots(id) ON DELETE CASCADE,
folder_id INTEGER REFERENCES folders(id) ON DELETE CASCADE,
source_ref TEXT NOT NULL,
-- Expensive: requires reading the whole file. Computed only when
-- something needs it (import dedup, reconnect-by-hash), never in a scan.
content_hash TEXT,
format TEXT,
w INTEGER,
h INTEGER,
-- UTC seconds. NULL until EXIF is read, or if the file carries none.
captured_at INTEGER,
-- Minutes east of UTC. A photograph's timestamp is local to where it was
-- taken; storing UTC alone makes a Tokyo shoot span two days in Paris.
captured_offset INTEGER,
camera TEXT,
lens TEXT,
iso INTEGER,
aperture REAL,
shutter REAL,
availability INTEGER NOT NULL DEFAULT 0,
file_size INTEGER,
file_mtime INTEGER,
-- 0 = nothing, 1 = stat-only, 2 = full EXIF. The grid is usable at 1.
metadata_state INTEGER NOT NULL DEFAULT 0,
sidecar_mtime INTEGER,
added_at INTEGER NOT NULL,
UNIQUE(root_id, source_ref)
);
CREATE INDEX images_captured ON images(captured_at);
CREATE INDEX images_folder ON images(folder_id);
-- Partial: content_hash is NULL for most rows most of the time, and the
-- non-NULL subset is exactly what reconnect and dedup query.
CREATE INDEX images_hash ON images(content_hash) WHERE content_hash IS NOT NULL;
-- Versions ----------------------------------------------------------------
CREATE TABLE versions (
id INTEGER PRIMARY KEY,
image_id INTEGER NOT NULL REFERENCES images(id) ON DELETE CASCADE,
uuid TEXT NOT NULL UNIQUE,
name TEXT NOT NULL,
is_default INTEGER NOT NULL DEFAULT 0,
graph_hash TEXT,
rating INTEGER NOT NULL DEFAULT 0,
label INTEGER,
flag INTEGER NOT NULL DEFAULT 0
);
CREATE INDEX versions_image ON versions(image_id);
CREATE TABLE keywords (
version_id INTEGER NOT NULL REFERENCES versions(id) ON DELETE CASCADE,
keyword TEXT NOT NULL,
PRIMARY KEY(version_id, keyword)
);
CREATE INDEX keywords_term ON keywords(keyword);
-- Remote mapping ----------------------------------------------------------
CREATE TABLE remote (
image_id INTEGER PRIMARY KEY REFERENCES images(id) ON DELETE CASCADE,
-- oc:fileid — stable across server-side rename and move, so a move is not
-- a re-download of 80 MB.
file_id INTEGER NOT NULL,
etag TEXT,
sync_state INTEGER NOT NULL DEFAULT 0,
remote_path TEXT
);
CREATE UNIQUE INDEX remote_file ON remote(file_id);
-- Collections -------------------------------------------------------------
CREATE TABLE collections (
id INTEGER PRIMARY KEY,
-- Device-independent identity. The integer id is local and collides
-- across devices; the UUID is what a cross-device merge keys on.
uuid TEXT NOT NULL UNIQUE,
name TEXT NOT NULL,
parent_id INTEGER REFERENCES collections(id) ON DELETE CASCADE,
kind INTEGER NOT NULL, -- 0 = manual, 1 = smart
selector_json TEXT, -- smart only
created INTEGER NOT NULL,
-- Monotonic per collection, bumped on every local edit. Merge compares
-- these rather than file mtimes, so a clock-skewed device cannot silently
-- win.
revision INTEGER NOT NULL DEFAULT 1,
modified INTEGER NOT NULL,
-- Tombstone. A deleted collection must outlive its deletion, or a merge
-- with a device that still has it would resurrect it.
deleted INTEGER NOT NULL DEFAULT 0
);
CREATE TABLE collection_members (
collection_id INTEGER NOT NULL REFERENCES collections(id) ON DELETE CASCADE,
image_id INTEGER NOT NULL REFERENCES images(id) ON DELETE CASCADE,
position INTEGER, -- manual ordering; NULL = by capture time
added INTEGER NOT NULL,
PRIMARY KEY(collection_id, image_id)
);
CREATE INDEX members_image ON collection_members(image_id);
-- Cache -------------------------------------------------------------------
CREATE TABLE cache (
id INTEGER PRIMARY KEY,
version_id INTEGER REFERENCES versions(id) ON DELETE CASCADE,
image_id INTEGER REFERENCES images(id) ON DELETE CASCADE,
kind INTEGER NOT NULL, -- thumbnail | proxy | original
resolution INTEGER,
graph_hash TEXT,
path TEXT NOT NULL,
bytes INTEGER NOT NULL,
last_used INTEGER NOT NULL
);
CREATE INDEX cache_lru ON cache(last_used);
CREATE TABLE cache_rules (
id INTEGER PRIMARY KEY,
selector_json TEXT NOT NULL,
tier INTEGER NOT NULL,
priority INTEGER NOT NULL DEFAULT 0,
enabled INTEGER NOT NULL DEFAULT 1
);
CREATE TABLE image_cache (
image_id INTEGER PRIMARY KEY REFERENCES images(id) ON DELETE CASCADE,
tier_actual INTEGER NOT NULL DEFAULT 0,
-- Materialised rather than recomputed, so the grid can draw availability
-- badges without evaluating every rule for every visible cell.
tier_desired INTEGER NOT NULL DEFAULT 0,
bytes INTEGER NOT NULL DEFAULT 0,
last_used INTEGER,
pinned_by_rule INTEGER REFERENCES cache_rules(id) ON DELETE SET NULL
);
-- Jobs --------------------------------------------------------------------
CREATE TABLE jobs (
id INTEGER PRIMARY KEY,
kind INTEGER NOT NULL,
subject_id INTEGER,
priority INTEGER NOT NULL DEFAULT 0,
state INTEGER NOT NULL DEFAULT 0, -- 0=pending 1=running 2=failed
attempts INTEGER NOT NULL DEFAULT 0,
not_before INTEGER NOT NULL DEFAULT 0,
payload TEXT,
last_error TEXT,
-- Coalescing. Enqueueing the same work twice updates one row rather than
-- queueing it twice, which is what makes "enqueue on any change" safe to
-- call liberally.
UNIQUE(kind, subject_id)
);
CREATE INDEX jobs_ready ON jobs(state, priority DESC, not_before);
"#;
#[cfg(test)]
mod tests {
use super::*;
fn mem() -> Connection {
let c = Connection::open_in_memory().unwrap();
configure(&c).unwrap();
c
}
/// How many rows a named backfill touched, ignoring the others.
///
/// Asserting on the whole vector would couple every test to which other
/// backfills happen to exist.
fn backfilled(c: &Connection, what: &str) -> usize {
backfill(c)
.unwrap()
.into_iter()
.find(|(name, _)| *name == what)
.map(|(_, n)| n)
.unwrap_or(0)
}
/// Insert an image and return its id.
fn image(c: &Connection, folder: Option<i64>, name: &str, format: &str) -> i64 {
c.execute(
"INSERT INTO images(root_id, folder_id, source_ref, format, added_at)
VALUES (1, ?1, ?2, ?3, 0)",
rusqlite::params![folder, name, format],
)
.unwrap();
c.last_insert_rowid()
}
fn with_root() -> Connection {
let c = mem();
migrate(&c).unwrap();
c.execute(
"INSERT INTO roots(id, kind, label) VALUES (1, 'remote', 'lib')",
[],
)
.unwrap();
c.execute(
"INSERT INTO folders(id, root_id, path) VALUES (1, 1, 'a'), (2, 1, 'b')",
[],
)
.unwrap();
c
}
#[test]
fn a_jpeg_beside_its_raw_is_shadowed() {
// The camera's own rendering of a frame, not a second photograph.
let c = with_root();
let raw = image(&c, Some(1), "a/IMG_1234.CR2", "cr2");
let jpeg = image(&c, Some(1), "a/IMG_1234.JPG", "jpg");
assert_eq!(backfilled(&c, "shadowed_by"), 1);
let got: Option<i64> = c
.query_row(
"SELECT shadowed_by FROM images WHERE id = ?1",
[jpeg],
|r| r.get(0),
)
.unwrap();
assert_eq!(got, Some(raw));
}
#[test]
fn extension_case_does_not_matter() {
let c = with_root();
image(&c, Some(1), "a/IMG_1.cr2", "cr2");
image(&c, Some(1), "a/img_1.JPG", "jpg");
assert_eq!(backfilled(&c, "shadowed_by"), 1);
}
#[test]
fn a_standalone_jpeg_is_untouched() {
// Scanned film has no RAW sibling and must stay visible — 2,656 of
// them in the reference library.
let c = with_root();
image(&c, Some(1), "a/SCAN_0001.jpg", "jpg");
assert_eq!(backfilled(&c, "shadowed_by"), 0);
}
#[test]
fn a_jpeg_in_a_different_folder_is_not_shadowed() {
// Camera filenames wrap at IMG_9999, so the same stem recurs across
// shoots (FR-CAT-11). Only a same-folder pair is safe to collapse.
let c = with_root();
image(&c, Some(1), "a/IMG_1234.CR2", "cr2");
image(&c, Some(2), "b/IMG_1234.JPG", "jpg");
assert_eq!(backfilled(&c, "shadowed_by"), 0);
}
#[test]
fn a_raw_is_never_shadowed_by_a_jpeg() {
// The relationship is one-way: the RAW is the photograph.
let c = with_root();
let raw = image(&c, Some(1), "a/IMG_1.CR2", "cr2");
image(&c, Some(1), "a/IMG_1.JPG", "jpg");
backfill(&c).unwrap();
let got: Option<i64> = c
.query_row("SELECT shadowed_by FROM images WHERE id = ?1", [raw], |r| {
r.get(0)
})
.unwrap();
assert_eq!(got, None);
}
#[test]
fn backfill_is_idempotent() {
// It runs on every open, so a second pass must find nothing to do.
let c = with_root();
image(&c, Some(1), "a/IMG_1.CR2", "cr2");
image(&c, Some(1), "a/IMG_1.JPG", "jpg");
assert_eq!(backfilled(&c, "shadowed_by"), 1);
assert_eq!(backfilled(&c, "shadowed_by"), 0, "second pass is a no-op");
}
#[test]
fn a_v1_catalog_gains_the_column_and_is_backfilled() {
// The migration case that motivated this: rows already present when a
// column is added are silently partial until something backfills them.
let c = mem();
c.execute_batch(V1).unwrap();
c.pragma_update(None, "user_version", 1).unwrap();
c.execute(
"INSERT INTO roots(id, kind, label) VALUES (1, 'remote', 'lib')",
[],
)
.unwrap();
c.execute(
"INSERT INTO images(root_id, source_ref, format, added_at)
VALUES (1, 'IMG_9.CR2', 'cr2', 0), (1, 'IMG_9.JPG', 'jpg', 0)",
[],
)
.unwrap();
assert_eq!(migrate(&c).unwrap(), 1, "migrated from v1");
assert_eq!(backfilled(&c, "shadowed_by"), 1);
}
#[test]
fn a_v4_catalog_gains_the_pinning_columns() {
// TRACES: FR-NC-6a
// An existing library must not have to be rescanned to gain offline
// pinning. The rows are already there; only the columns are new.
let c = mem();
c.execute_batch(V1).unwrap();
c.execute_batch(V2).unwrap();
c.execute_batch(V3).unwrap();
c.execute_batch(V4).unwrap();
c.pragma_update(None, "user_version", 4).unwrap();
c.execute(
"INSERT INTO roots(id, kind, label) VALUES (1, 'remote', 'lib')",
[],
)
.unwrap();
c.execute(
"INSERT INTO images(id, root_id, source_ref, added_at)
VALUES (7, 1, 'IMG_7.CR2', 0)",
[],
)
.unwrap();
// A cache row written before pinning existed.
c.execute(
"INSERT INTO image_cache(image_id, tier_actual, bytes) VALUES (7, 2, 100)",
[],
)
.unwrap();
assert_eq!(migrate(&c).unwrap(), 4, "migrated from v4");
// The pre-existing row survives, and defaults to unpinned — the safe
// direction, since claiming a pin nobody made would exempt it from
// eviction for ever.
let (pinned, bytes): (i64, i64) = c
.query_row(
"SELECT pinned, bytes FROM image_cache WHERE image_id = 7",
[],
|r| Ok((r.get(0)?, r.get(1)?)),
)
.unwrap();
assert_eq!(pinned, 0);
assert_eq!(bytes, 100, "the existing row is untouched");
}
#[test]
fn stems_ignore_directories_containing_dots() {
assert_eq!(stem_of("2026.08/IMG_1.CR2"), "IMG_1");
assert_eq!(stem_of("IMG_1.CR2"), "IMG_1");
assert_eq!(stem_of("noextension"), "noextension");
// A dotfile is all stem, not an empty name with an extension.
assert_eq!(stem_of(".hidden"), ".hidden");
}
#[test]
fn migrate_creates_schema_at_current_version() {
let c = mem();
assert_eq!(migrate(&c).unwrap(), 0);
let v: i64 = c
.query_row("PRAGMA user_version", [], |r| r.get(0))
.unwrap();
assert_eq!(v, SCHEMA_VERSION);
}
#[test]
fn migrate_is_idempotent() {
let c = mem();
migrate(&c).unwrap();
// Re-running must not error or duplicate anything — NFR-R5 requires
// idempotency on retry, since a migration can be interrupted.
assert_eq!(migrate(&c).unwrap(), SCHEMA_VERSION);
}
#[test]
fn refuses_a_catalog_from_a_newer_build() {
let c = mem();
migrate(&c).unwrap();
c.pragma_update(None, "user_version", SCHEMA_VERSION + 1)
.unwrap();
// Opening it read-write would corrupt data this build cannot
// represent. Refusing is the specified behaviour (NFR-R5).
assert!(matches!(
migrate(&c),
Err(CatalogError::SchemaTooNew { .. })
));
}
#[test]
fn foreign_keys_cascade_from_root_to_image() {
let c = mem();
migrate(&c).unwrap();
c.execute(
"INSERT INTO roots(id, kind, label) VALUES (1, 'local', 'test')",
[],
)
.unwrap();
c.execute(
"INSERT INTO images(id, root_id, source_ref, added_at) VALUES (1, 1, 'a.CR3', 0)",
[],
)
.unwrap();
c.execute("DELETE FROM roots WHERE id = 1", []).unwrap();
let n: i64 = c
.query_row("SELECT count(*) FROM images", [], |r| r.get(0))
.unwrap();
assert_eq!(n, 0, "images must not outlive their root");
}
#[test]
fn job_uniqueness_coalesces_rather_than_duplicating() {
let c = mem();
migrate(&c).unwrap();
for _ in 0..5 {
c.execute(
"INSERT INTO jobs(kind, subject_id, priority) VALUES (1, 42, 0)
ON CONFLICT(kind, subject_id)
DO UPDATE SET priority = max(priority, excluded.priority)",
[],
)
.unwrap();
}
let n: i64 = c
.query_row("SELECT count(*) FROM jobs", [], |r| r.get(0))
.unwrap();
assert_eq!(n, 1, "five enqueues of the same work is one job");
}
}