Files
dtourolle 3c2eacbf3f Find catalog duplicates and fold a group onto one copy in one transaction
The library holds the same RAW in several folders: a dated folder, a
bck/ beside it, a renamed Darktable export tree. dr_catalog::duplicates
is the catalog half of consolidating them (#67).

candidates() is one grouped query over root, camera, capture instant and
size, joined back for the rows; count() is the same grouping under
COUNT. On a copy of the reference catalog (23,582 images) both take
10-35 ms and find 1,836 groups holding 3,379 spare copies.

survivor() prefers a copy outside a backup-looking folder, then one
still named the way the camera named it, then the oldest, then the
lowest id.

consolidate() re-checks the plan against the catalog, merges the copies'
judgements onto the survivor (highest rating, keywords unioned,
collections unioned with the survivor keeping its place, a flag or label
the copies agree on, faces via faces::carry_onto_copy) and records the
copies as trashed, all in one transaction, so a failure part way leaves
the group untouched. preview() runs the same code and rolls it back.

Sameness probes are kept in dedup_probes, created on first use rather
than by a migration: a schema bump would make older builds refuse this
catalog's snapshot at sync. trash::record_trashed_within lets the trash
write share the merge's transaction.
2026-09-26 07:19:13 -04:00

649 lines
24 KiB
Rust

//! TRACES: FR-CAT-15 | NFR-R2
//! Soft delete, restore, and the permanent delete that follows.
//!
//! # Why the trash is a folder and not a flag
//!
//! The catalog is a *rebuildable index* (ARCH §6.12): delete `catalog.sqlite`
//! and it is reconstructed by rescanning sources. A trash implemented as a
//! column alone would therefore not survive its own design — a rebuild would
//! find every trashed file still sitting in the library and re-index it as an
//! ordinary photograph, silently undoing every delete the user had made.
//!
//! So a soft delete **moves the file** into `.darkroom-trash/` under the library
//! root, and the catalog merely records that this happened. The folder is the
//! durable fact; the row is the convenience. Recovering by hand needs no
//! DarkRoom at all, which is the property that matters when the thing being
//! risked is a photograph.
//!
//! `dr_sync::scan::is_excluded` keeps the scanner out of that folder. Without
//! it the next scan re-indexes the trash and the delete comes undone — the two
//! halves are one mechanism and neither works alone.
//!
//! # The two steps
//!
//! **Soft** ([`trash`]) — `MOVE` to the trash folder, record `trashed_at` and
//! the path it came from. Reversible by [`restore`], which is why the original
//! path has to be remembered: the trash is flat, and the folder structure cannot
//! be recovered from the trashed name.
//!
//! **Hard** ([`purge`]) — `DELETE` the file, then delete the row. Irreversible
//! from DarkRoom's side, though the server's own trashbin may still hold it.
//! Ordered file-first deliberately: see [`purge_order`].
//!
//! # What this module does not do
//!
//! It performs no I/O. Every function here records or reads catalog state, and
//! the caller pairs it with the remote operation — because the remote call is
//! async and the catalog is not, and because the *order* of the two is a
//! correctness property that belongs in one visible place rather than buried in
//! a transaction.
use rusqlite::{Connection, OptionalExtension};
use dr_types::ImageId;
use crate::error::CatalogError;
/// Directory holding soft-deleted images, under the library root.
///
/// The same constant `dr_sync::scan` excludes. Duplicated as a `const` here
/// rather than depended upon because `dr-catalog` does not (and should not)
/// depend on `dr-sync`; the pairing is asserted by a test.
pub const TRASH_DIR: &str = ".darkroom-trash";
/// One trashed image, as the trash view lists it.
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct TrashedImage {
pub image_id: ImageId,
/// Where the file is *now* — inside the trash folder.
pub source_ref: String,
/// Where it was before, and where [`restore`] will put it back.
pub trashed_from: String,
/// UTC seconds when it was trashed.
pub trashed_at: i64,
/// `oc:fileid`, preserved across the move. What the thumbnail store keys on,
/// and what makes a restore free rather than a re-download.
pub file_id: Option<u64>,
pub size: u64,
}
/// The path a soft-deleted image should be moved to.
///
/// Flat: the trash is a holding area, not an archive, and mirroring the library
/// tree inside it would mean creating directories on the way to deleting things.
/// The original path is remembered in the catalog instead, which is what
/// [`restore`] reads.
///
/// **Collisions are resolved rather than allowed to overwrite.** Two files named
/// `IMG_0001.CR2` from different folders are different photographs, and a `MOVE`
/// onto an existing name would destroy one of them — the precise failure a trash
/// exists to prevent. The image id disambiguates, and being already unique it
/// needs no retry loop.
pub fn trash_path(root: &str, image: ImageId, original: &str) -> String {
let name = original.rsplit(['/', ':']).next().unwrap_or(original);
let prefix = if root.is_empty() {
String::new()
} else {
format!("{root}/")
};
format!("{prefix}{TRASH_DIR}/{}-{name}", image.0)
}
/// Where a trashed image goes back to.
///
/// The stored original path, verbatim. Returns `None` where the image is not
/// trashed, so a caller cannot restore something that was never deleted.
pub fn restore_path(conn: &Connection, image: ImageId) -> Result<Option<String>, CatalogError> {
let path: Option<String> = conn
.query_row(
"SELECT trashed_from FROM images
WHERE id = ?1 AND trashed_at IS NOT NULL",
[image.0 as i64],
|r| r.get(0),
)
.optional()?
.flatten();
Ok(path)
}
/// Record that images have been moved to the trash.
///
/// Call **after** the move succeeds. Recording first and moving second would
/// leave the catalog claiming a file is trashed while it sits in the library,
/// where the next scan finds it — and since the scan excludes the trash folder,
/// the row would never be corrected.
///
/// `moved` pairs each image with the path it now occupies, which is what
/// [`trash_path`] produced for it.
///
/// Idempotent on `trashed_at`: re-trashing an already-trashed image keeps the
/// *original* timestamp and original path, so a retry after a partial failure
/// cannot rewrite `trashed_from` to a path inside the trash — which would make
/// the image unrestorable.
pub fn record_trashed(
conn: &Connection,
moved: &[(ImageId, String)],
now: i64,
) -> Result<usize, CatalogError> {
if moved.is_empty() {
return Ok(0);
}
let tx = conn.unchecked_transaction()?;
let n = record_trashed_within(&tx, moved, now)?;
tx.commit()?;
Ok(n)
}
/// [`record_trashed`] inside a transaction the caller owns.
///
/// For a caller whose trash is one half of a larger write that must land
/// whole or not at all — consolidating duplicates (`crate::duplicates`)
/// merges a copy's judgements onto the survivor and trashes the copy in one
/// commit. `unchecked_transaction` cannot nest, so this is offered here
/// rather than wrapped from above.
pub fn record_trashed_within(
tx: &Connection,
moved: &[(ImageId, String)],
now: i64,
) -> Result<usize, CatalogError> {
let mut n = 0;
let mut stmt = tx.prepare(
"UPDATE images
SET trashed_from = CASE
WHEN trashed_at IS NULL THEN source_ref
ELSE trashed_from
END,
source_ref = ?2,
trashed_at = coalesce(trashed_at, ?3)
WHERE id = ?1",
)?;
for (image, path) in moved {
n += stmt.execute(rusqlite::params![image.0 as i64, path, now])?;
}
Ok(n)
}
/// Record that images have been moved back out of the trash.
///
/// Call after the move succeeds, for the same reason as [`record_trashed`].
/// Clears both columns: a restored image is an ordinary one, and leaving
/// `trashed_from` set would make the next trash-and-restore cycle restore it to
/// a stale location.
pub fn record_restored(
conn: &Connection,
restored: &[(ImageId, String)],
) -> Result<usize, CatalogError> {
if restored.is_empty() {
return Ok(0);
}
let tx = conn.unchecked_transaction()?;
let mut n = 0;
{
let mut stmt = tx.prepare(
"UPDATE images
SET source_ref = ?2, trashed_at = NULL, trashed_from = NULL
WHERE id = ?1 AND trashed_at IS NOT NULL",
)?;
for (image, path) in restored {
n += stmt.execute(rusqlite::params![image.0 as i64, path])?;
}
}
tx.commit()?;
Ok(n)
}
/// Forget images whose files have been permanently deleted.
///
/// Call **after** the remote delete succeeds — see [`purge_order`].
///
/// Deletes the catalog rows outright rather than tombstoning them. There is
/// nothing to merge: unlike a collection, an image row is derived from a file
/// that no longer exists, so a rescan on another device will not reintroduce it
/// and needs no tombstone to be told so. `ON DELETE CASCADE` takes the versions,
/// keywords, remote mapping and cache rows with it.
///
/// Returns how many rows went.
pub fn forget(conn: &Connection, images: &[ImageId]) -> Result<usize, CatalogError> {
if images.is_empty() {
return Ok(0);
}
let tx = conn.unchecked_transaction()?;
let mut n = 0;
{
let mut stmt = tx.prepare("DELETE FROM images WHERE id = ?1")?;
for image in images {
n += stmt.execute([image.0 as i64])?;
}
}
tx.commit()?;
Ok(n)
}
/// Why the file is deleted before the row.
///
/// Not a function — a note with a name, so the reasoning is findable from the
/// call site.
///
/// **File first, then the row.** If the delete succeeds and the process dies
/// before the row goes, the catalog holds a trashed row whose file is gone; the
/// user sees it in the trash, empties again, gets a `404`, and it is treated as
/// already-deleted (see [`is_already_gone`]). Recoverable, and visible.
///
/// The other order loses the file silently. Dropping the row first and dying
/// before the delete leaves an orphan in `.darkroom-trash/` that nothing in the
/// UI lists, nothing counts, and no scan will ever find — because the scanner
/// excludes that folder. It consumes quota forever and the user has no way to
/// learn it is there.
pub const fn purge_order() {}
/// Whether a delete failure means the file was already gone.
///
/// A `404` on the way to deleting something is success: the goal state is
/// "this file does not exist", and it does not. Treating it as an error would
/// wedge an empty-trash operation on a file the user had removed by hand, and
/// no amount of retrying would clear it.
pub fn is_already_gone(status: Option<u16>) -> bool {
matches!(status, Some(404) | Some(410))
}
/// List what is in the trash, newest first.
///
/// Newest first because the trash is reviewed to undo a recent mistake, not
/// browsed chronologically.
pub fn list(conn: &Connection, limit: usize) -> Result<Vec<TrashedImage>, CatalogError> {
let mut stmt = conn.prepare(
"SELECT i.id, i.source_ref, i.trashed_from, i.trashed_at, r.file_id, i.file_size
FROM images i
LEFT JOIN remote r ON r.image_id = i.id
WHERE i.trashed_at IS NOT NULL
ORDER BY i.trashed_at DESC, i.id DESC
LIMIT ?1",
)?;
let rows = stmt
.query_map([limit as i64], |r| {
let source_ref: String = r.get(1)?;
Ok(TrashedImage {
image_id: ImageId(r.get::<_, i64>(0)? as u64),
// A row with no `trashed_from` predates nothing — it cannot
// happen through this module — but a hand-edited or
// partially-migrated catalog could produce one. Falling back to
// the current path keeps it listed and deletable rather than
// invisible; a restore to the trash folder is a no-op the user
// can see, where a hidden row is not.
trashed_from: r
.get::<_, Option<String>>(2)?
.unwrap_or_else(|| source_ref.clone()),
source_ref,
trashed_at: r.get(3)?,
file_id: r.get::<_, Option<i64>>(4)?.map(|v| v as u64),
size: r.get::<_, Option<i64>>(5)?.unwrap_or(0) as u64,
})
})?
.collect::<Result<Vec<_>, _>>()?;
Ok(rows)
}
/// Every trashed image id, for emptying the whole trash.
///
/// Separate from [`list`] because emptying needs all of them, not a window, and
/// wants no per-row detail.
pub fn all_trashed(conn: &Connection) -> Result<Vec<ImageId>, CatalogError> {
let mut stmt = conn.prepare("SELECT id FROM images WHERE trashed_at IS NOT NULL")?;
let rows = stmt
.query_map([], |r| Ok(ImageId(r.get::<_, i64>(0)? as u64)))?
.collect::<Result<Vec<_>, _>>()?;
Ok(rows)
}
/// How many images are in the trash, and how many bytes they hold.
///
/// The bytes are the point: "empty trash" is a destructive action, and the
/// amount being freed is what tells the user whether they meant it.
pub fn summary(conn: &Connection) -> Result<(usize, u64), CatalogError> {
let (n, bytes): (i64, i64) = conn.query_row(
"SELECT count(*), coalesce(sum(file_size), 0)
FROM images WHERE trashed_at IS NOT NULL",
[],
|r| Ok((r.get(0)?, r.get(1)?)),
)?;
Ok((n as usize, bytes as u64))
}
/// `oc:fileid`s of trashed images, so their thumbnails can be dropped.
///
/// The thumbnail store is keyed on the stable file id and shared with other
/// clients, so a purge that left its entries behind would keep serving previews
/// of photographs that no longer exist — and the shards sync, so it would keep
/// doing so on every other device too.
pub fn file_ids_for(conn: &Connection, images: &[ImageId]) -> Result<Vec<u64>, CatalogError> {
if images.is_empty() {
return Ok(Vec::new());
}
let placeholders = std::iter::repeat_n("?", images.len())
.collect::<Vec<_>>()
.join(",");
let sql = format!("SELECT file_id FROM remote WHERE image_id IN ({placeholders})");
let params: Vec<rusqlite::types::Value> = images
.iter()
.map(|i| rusqlite::types::Value::Integer(i.0 as i64))
.collect();
let mut stmt = conn.prepare(&sql)?;
let rows = stmt
.query_map(rusqlite::params_from_iter(params.iter()), |r| {
Ok(r.get::<_, i64>(0)? as u64)
})?
.collect::<Result<Vec<_>, _>>()?;
Ok(rows)
}
#[cfg(test)]
mod tests {
use super::*;
use crate::Catalog;
fn seeded() -> Catalog {
let cat = Catalog::in_memory().unwrap();
let c = cat.connection();
c.execute(
"INSERT INTO roots(id, kind, label) VALUES (1, 'remote', 'PhotosRaw')",
[],
)
.unwrap();
for i in 1..=4i64 {
c.execute(
"INSERT INTO images(id, root_id, source_ref, file_size, added_at)
VALUES (?1, 1, ?2, ?3, 0)",
rusqlite::params![i, format!("PhotosRaw/2019/IMG_{i:04}.CR2"), 30_000_000 * i],
)
.unwrap();
c.execute(
"INSERT INTO remote(image_id, file_id) VALUES (?1, ?2)",
rusqlite::params![i, 1000 + i],
)
.unwrap();
}
cat
}
fn img(i: u64) -> ImageId {
ImageId(i)
}
/// Trash one image the way the UI does: compute the path, then record.
fn do_trash(cat: &Catalog, i: u64, now: i64) -> String {
let c = cat.connection();
let original: String = c
.query_row(
"SELECT source_ref FROM images WHERE id = ?1",
[i as i64],
|r| r.get(0),
)
.unwrap();
let to = trash_path("PhotosRaw", img(i), &original);
record_trashed(c, &[(img(i), to.clone())], now).unwrap();
to
}
#[test]
fn the_trash_directory_matches_the_one_the_scanner_excludes() {
// These are two constants in two crates that must agree, or the scan
// re-indexes the trash and every soft delete comes undone.
assert_eq!(TRASH_DIR, dr_sync_trash_dir());
}
/// The scanner's constant, quoted rather than imported — `dr-catalog` does
/// not depend on `dr-sync`, and adding that dependency for one string would
/// invert the layering.
fn dr_sync_trash_dir() -> &'static str {
".darkroom-trash"
}
#[test]
fn trashing_moves_the_path_and_remembers_where_it_came_from() {
let cat = seeded();
let c = cat.connection();
do_trash(&cat, 1, 5_000);
let (source, from, at): (String, String, i64) = c
.query_row(
"SELECT source_ref, trashed_from, trashed_at FROM images WHERE id = 1",
[],
|r| Ok((r.get(0)?, r.get(1)?, r.get(2)?)),
)
.unwrap();
// `source_ref` follows the bytes: this is where a fetch must now look.
assert!(source.contains(TRASH_DIR), "{source}");
// And the original is remembered, or a restore has nowhere to go.
assert_eq!(from, "PhotosRaw/2019/IMG_0001.CR2");
assert_eq!(at, 5_000);
}
#[test]
fn the_trash_path_keeps_the_original_filename_recognisable() {
// The user reviewing the trash needs to recognise the photograph; an
// opaque id alone would make the list unreadable.
let p = trash_path("PhotosRaw", img(7), "PhotosRaw/2019/IMG_0042.CR2");
assert!(p.ends_with("IMG_0042.CR2"), "{p}");
assert!(p.starts_with("PhotosRaw/.darkroom-trash/"), "{p}");
}
#[test]
fn two_files_with_the_same_name_do_not_collide_in_the_trash() {
// The failure a trash exists to prevent: a MOVE onto an existing name
// destroys one of two different photographs.
let a = trash_path("PhotosRaw", img(1), "PhotosRaw/2019/IMG_0001.CR2");
let b = trash_path("PhotosRaw", img(2), "PhotosRaw/2024/IMG_0001.CR2");
assert_ne!(a, b);
}
#[test]
fn a_whole_account_root_yields_no_leading_slash() {
// The root is empty when the library is the whole account; a path
// beginning "/" would resolve differently on the server.
let p = trash_path("", img(3), "2019/IMG_0003.CR2");
assert_eq!(p, ".darkroom-trash/3-IMG_0003.CR2");
}
#[test]
fn restoring_puts_the_original_path_back_and_clears_the_flag() {
let cat = seeded();
let c = cat.connection();
do_trash(&cat, 1, 5_000);
let back = restore_path(c, img(1))
.unwrap()
.expect("knows where it came from");
assert_eq!(back, "PhotosRaw/2019/IMG_0001.CR2");
record_restored(c, &[(img(1), back.clone())]).unwrap();
let (source, at): (String, Option<i64>) = c
.query_row(
"SELECT source_ref, trashed_at FROM images WHERE id = 1",
[],
|r| Ok((r.get(0)?, r.get(1)?)),
)
.unwrap();
assert_eq!(source, back);
assert_eq!(at, None, "a restored image is an ordinary one");
assert!(restore_path(c, img(1)).unwrap().is_none());
}
#[test]
fn a_trash_restore_trash_cycle_restores_to_the_right_place_twice() {
// If `trashed_from` were not cleared on restore, the second trash would
// record a stale origin and the second restore would put the file
// somewhere it never was.
let cat = seeded();
let c = cat.connection();
do_trash(&cat, 1, 1_000);
let first = restore_path(c, img(1)).unwrap().unwrap();
record_restored(c, &[(img(1), first.clone())]).unwrap();
do_trash(&cat, 1, 2_000);
let second = restore_path(c, img(1)).unwrap().unwrap();
assert_eq!(
first, second,
"the origin is the library path, not the trash"
);
}
#[test]
fn re_trashing_does_not_overwrite_the_original_path() {
// A retry after a partial failure must not record a trash-folder path as
// the origin — that makes the image unrestorable.
let cat = seeded();
let c = cat.connection();
let to = do_trash(&cat, 1, 1_000);
// Second attempt, as a retry would do.
record_trashed(c, &[(img(1), to)], 9_999).unwrap();
let (from, at): (String, i64) = c
.query_row(
"SELECT trashed_from, trashed_at FROM images WHERE id = 1",
[],
|r| Ok((r.get(0)?, r.get(1)?)),
)
.unwrap();
assert_eq!(from, "PhotosRaw/2019/IMG_0001.CR2");
assert_eq!(at, 1_000, "the original timestamp survives a retry");
}
#[test]
fn restoring_something_that_was_never_trashed_does_nothing() {
let cat = seeded();
let c = cat.connection();
assert!(restore_path(c, img(2)).unwrap().is_none());
assert_eq!(
record_restored(c, &[(img(2), "elsewhere".into())]).unwrap(),
0
);
// And its path is untouched.
let source: String = c
.query_row("SELECT source_ref FROM images WHERE id = 2", [], |r| {
r.get(0)
})
.unwrap();
assert_eq!(source, "PhotosRaw/2019/IMG_0002.CR2");
}
#[test]
fn the_trash_lists_newest_first() {
// Reviewed to undo a recent mistake, not browsed chronologically.
let cat = seeded();
do_trash(&cat, 1, 1_000);
do_trash(&cat, 2, 3_000);
do_trash(&cat, 3, 2_000);
let listed = list(cat.connection(), 100).unwrap();
let order: Vec<u64> = listed.iter().map(|t| t.image_id.0).collect();
assert_eq!(order, vec![2, 3, 1]);
}
#[test]
fn the_trash_list_carries_the_file_id_a_restore_needs() {
// Without it a restore cannot find the thumbnail it already has, and
// re-downloads a preview it is holding.
let cat = seeded();
do_trash(&cat, 1, 1_000);
let listed = list(cat.connection(), 10).unwrap();
assert_eq!(listed[0].file_id, Some(1001));
}
#[test]
fn the_summary_reports_what_emptying_would_free() {
// "Empty trash" is destructive; the size is what tells the user whether
// they meant it.
let cat = seeded();
do_trash(&cat, 1, 1_000);
do_trash(&cat, 2, 1_000);
let (n, bytes) = summary(cat.connection()).unwrap();
assert_eq!(n, 2);
assert_eq!(bytes, 30_000_000 + 60_000_000);
}
#[test]
fn an_empty_trash_summarises_as_zero_rather_than_erroring() {
let cat = seeded();
assert_eq!(summary(cat.connection()).unwrap(), (0, 0));
assert!(all_trashed(cat.connection()).unwrap().is_empty());
}
#[test]
fn purging_removes_the_row_and_everything_hanging_off_it() {
let cat = seeded();
let c = cat.connection();
do_trash(&cat, 1, 1_000);
assert_eq!(forget(c, &[img(1)]).unwrap(), 1);
let n: i64 = c
.query_row("SELECT count(*) FROM images WHERE id = 1", [], |r| r.get(0))
.unwrap();
assert_eq!(n, 0);
// The remote mapping must go too, or a later scan could pair a new file
// with a dead image's id.
let n: i64 = c
.query_row("SELECT count(*) FROM remote WHERE image_id = 1", [], |r| {
r.get(0)
})
.unwrap();
assert_eq!(n, 0, "cascaded");
}
#[test]
fn purging_leaves_untrashed_images_alone() {
let cat = seeded();
let c = cat.connection();
do_trash(&cat, 1, 1_000);
forget(c, &all_trashed(c).unwrap()).unwrap();
let n: i64 = c
.query_row("SELECT count(*) FROM images", [], |r| r.get(0))
.unwrap();
assert_eq!(n, 3, "only the trashed one went");
}
#[test]
fn file_ids_are_collected_so_thumbnails_can_be_dropped() {
// The shards sync to the server; a purge that left them would serve
// previews of deleted photographs on every device.
let cat = seeded();
let c = cat.connection();
do_trash(&cat, 1, 1_000);
do_trash(&cat, 2, 1_000);
let mut ids = file_ids_for(c, &[img(1), img(2)]).unwrap();
ids.sort_unstable();
assert_eq!(ids, vec![1001, 1002]);
}
#[test]
fn a_missing_file_counts_as_already_deleted() {
// Otherwise one file removed by hand wedges every future empty-trash,
// and no amount of retrying clears it.
assert!(is_already_gone(Some(404)));
assert!(is_already_gone(Some(410)));
assert!(!is_already_gone(Some(403)), "a permission failure is real");
assert!(!is_already_gone(Some(500)));
assert!(!is_already_gone(None));
}
#[test]
fn empty_batches_are_no_ops_rather_than_errors() {
// The UI can reach these with nothing selected.
let cat = seeded();
let c = cat.connection();
assert_eq!(record_trashed(c, &[], 0).unwrap(), 0);
assert_eq!(record_restored(c, &[]).unwrap(), 0);
assert_eq!(forget(c, &[]).unwrap(), 0);
assert!(file_ids_for(c, &[]).unwrap().is_empty());
}
}