Files
DarkRoom/core/dr-catalog/src/cache.rs
T
dtourolle 5768100816 Borrow the library to index it, and give it back
The passes that need every photograph's bytes — thumbnails, face
indexing — now borrow each one and release it at the end. On a
placeholder library that is the difference between peak disk being the
working set and being the whole library.

Including on cancellation, which was nearly missed: the face sweep
returns mid-loop when the user presses Stop, and without releasing there
the disk is spent and nothing is delivered for it.

`materialise` now answers whether *it* fetched the content. The pool used
to work that out by listing a file's parent directory — one listing per
file across a library — when the backend already had to `stat` it to
decide whether to ask. One syscall instead of a directory walk, and it
removes the bug class the tests found earlier: a file at the library root
has no `parent()`, so every one of them read as already-downloaded.

**Pinning is the retention control**, and it drives the model the catalog
already had rather than a second one. `tier_desired` is what the user
asked to keep hydrated, `pending_pins` is the resumable work list, and a
pinned collection is never dehydrated for the same reason it was never
evicted. It was in fact *broken* here before: `get` on a stub failed, and
the pin worker logged "one unreadable file must not abandon the whole
pin" and silently did nothing.

Pinned originals on such a library are recorded with `path = NULL`
(`Cache::record_in_place`) rather than copied under `originals/`. Two
reasons, and the second is the important one. A copy would hold every
pinned photograph twice, with the budget able to evict the half that was
not costing the disk. And `release` deletes the file a row names — so a
row that names none cannot delete anything, which puts the one
catastrophic operation out of reach by construction rather than by
remembering not to call it. Deleting a materialised file inside a synced
tree removes the photograph from the server and every other device.

Handing disk back is `spawn_dehydrate`, which asks the client.

Two gaps written down rather than papered over (docs/storage.md §7): a
hydrating pass cannot yet quote its cost, because a stub reports no size;
and the two sweeps hold separate pools, so a library indexed for both
fetches twice.
2026-08-29 09:57:53 +02:00

1037 lines
37 KiB
Rust

//! TRACES: FR-NC-6a | FR-CAT-9 | NFR-RES-4
//! Which originals are kept on this device, and which may be evicted.
//!
//! # Two populations, one table
//!
//! An original ends up here two ways, and conflating them produces exactly the
//! failure the whole feature exists to prevent.
//!
//! **Pinned** originals were asked for. A user pins a collection before a trip
//! and expects those photographs to be there when there is no connection —
//! that is a promise, so pinned rows are never evicted and never counted
//! against the budget. A cap that could silently delete a pinned trip would
//! make pinning worthless, because the user could not rely on it without
//! checking.
//!
//! **Passively cached** originals are a side effect of working: opening an
//! image in develop downloads it, so keeping the bytes costs nothing extra and
//! saves the whole transfer next time. This population is bounded by
//! [`Budget`] and evicted least-recently-used, because it grows without limit
//! otherwise — a day of culling would fill a disk.
//!
//! The two budgets are separate rather than shared. Sharing them means a large
//! pin starves the passive cache, or worse, that browsing evicts a pin.
//!
//! # What this module does and does not own
//!
//! It owns the *bookkeeping*: which images are held, at what tier, how large,
//! when last used, and which are pinned. The bytes are files under a cache
//! directory, and [`store`](Cache::store) writes them; but deciding to
//! download something is the caller's business, because that needs a network
//! and this crate has none.
//!
//! # Why `tier_actual` is the truth
//!
//! `tier_desired` is what a pin asks for; `tier_actual` is what is on disk.
//! Only the second answers "can this be opened right now", which is the
//! question offline mode asks (FR-CAT-9). A pinned image whose download has
//! not run yet is precisely the one that would fail, so it must not report as
//! available.
use std::path::{Path, PathBuf};
use dr_types::{ImageId, Tier};
use rusqlite::{Connection, OptionalExtension as _};
use crate::error::CatalogError;
/// Default ceiling for passively cached originals.
///
/// 1 GB holds roughly 30 full-frame RAWs — a working session's worth, which is
/// what this cache is for. It is deliberately modest: the passive cache is a
/// convenience that should not quietly consume a disk, and a user who wants
/// more kept is better served by pinning, which says so explicitly and is not
/// subject to eviction at all.
pub const DEFAULT_BUDGET_BYTES: u64 = 1024 * 1024 * 1024;
/// How much disk the passive cache may use.
///
/// A newtype rather than a bare `u64` so a byte count cannot be passed where a
/// budget belongs, and to give the "unlimited" case a name — some users have a
/// large disk and would rather never re-download.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub struct Budget(Option<u64>);
impl Default for Budget {
fn default() -> Self {
Self::bytes(DEFAULT_BUDGET_BYTES)
}
}
impl Budget {
pub fn bytes(n: u64) -> Self {
Self(Some(n))
}
/// No ceiling: nothing is ever evicted for space.
pub fn unlimited() -> Self {
Self(None)
}
pub fn limit(self) -> Option<u64> {
self.0
}
/// How much must be freed to fit `used` within this budget.
fn overage(self, used: u64) -> u64 {
self.0.map_or(0, |cap| used.saturating_sub(cap))
}
}
/// What is held for one image.
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct Entry {
pub image: ImageId,
/// What is actually on disk.
pub tier: Tier,
/// What a pin has asked for, which may be ahead of `tier`.
pub desired: Tier,
pub bytes: u64,
/// Unix seconds, or `None` if never read back since being stored.
pub last_used: Option<i64>,
pub pinned: bool,
/// Path relative to the cache directory.
pub path: Option<String>,
}
/// How the cache is currently filled.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Default)]
pub struct Usage {
/// Bytes held by pinned originals. Not subject to the budget.
pub pinned_bytes: u64,
/// Bytes held by passively cached originals. What the budget bounds.
pub passive_bytes: u64,
pub pinned_count: usize,
pub passive_count: usize,
}
impl Usage {
pub fn total_bytes(self) -> u64 {
self.pinned_bytes + self.passive_bytes
}
}
/// The on-disk cache of originals, rooted at a directory.
pub struct Cache {
dir: PathBuf,
budget: Budget,
}
impl Cache {
/// Open a cache rooted at `dir`, creating it if needed.
pub fn open(dir: &Path, budget: Budget) -> Result<Self, CatalogError> {
std::fs::create_dir_all(dir)
.map_err(|e| CatalogError::Io(format!("creating {}: {e}", dir.display())))?;
Ok(Self {
dir: dir.to_path_buf(),
budget,
})
}
pub fn dir(&self) -> &Path {
&self.dir
}
pub fn budget(&self) -> Budget {
self.budget
}
/// Absolute path for a cached original.
///
/// Named by image id rather than by the remote filename: two folders on
/// the server may hold `IMG_0001.CR2`, and a flat cache keyed on the name
/// would have them overwrite each other. The extension is preserved so the
/// decoder's format probe sees what it expects.
fn relative_path(image: ImageId, source_ref: &str) -> String {
let ext = source_ref
.rsplit_once('.')
.map(|(_, e)| e.to_ascii_lowercase())
.filter(|e| {
!e.is_empty() && e.len() <= 8 && e.chars().all(|c| c.is_ascii_alphanumeric())
})
.unwrap_or_else(|| "bin".to_string());
format!("{}.{ext}", image.0)
}
/// Store an original's bytes and record it.
///
/// `pinned` says which population this belongs to. Storing an image that
/// is already present updates it rather than duplicating — the same
/// photograph opened twice is one cache entry, and the second store simply
/// refreshes the bytes and the timestamp.
///
/// Does **not** evict. The caller runs [`enforce`](Self::enforce) once it
/// has finished storing, so a batch of downloads is trimmed once rather
/// than after every file.
pub fn store(
&self,
conn: &Connection,
image: ImageId,
source_ref: &str,
bytes: &[u8],
pinned: bool,
now: i64,
) -> Result<(), CatalogError> {
let rel = Self::relative_path(image, source_ref);
let abs = self.dir.join(&rel);
// Written to a temporary and renamed, so a crash or a dropped
// connection mid-write cannot leave a truncated file that the catalog
// records as a complete original — which would then fail to decode
// with no indication that the *cache* was at fault rather than the
// photograph.
let tmp = abs.with_extension("partial");
std::fs::write(&tmp, bytes)
.map_err(|e| CatalogError::Io(format!("writing {}: {e}", tmp.display())))?;
std::fs::rename(&tmp, &abs)
.map_err(|e| CatalogError::Io(format!("renaming {}: {e}", abs.display())))?;
// `pinned` is OR-ed rather than assigned: an image that was already
// pinned must not be demoted to evictable because it happened to be
// opened in develop, which is a passive store.
conn.execute(
"INSERT INTO image_cache
(image_id, tier_actual, tier_desired, bytes, last_used, pinned, path)
VALUES (?1, ?2, ?2, ?3, ?4, ?5, ?6)
ON CONFLICT(image_id) DO UPDATE SET
tier_actual = ?2,
tier_desired = max(tier_desired, ?2),
bytes = ?3,
last_used = ?4,
pinned = max(pinned, ?5),
path = ?6",
rusqlite::params![
image.0 as i64,
Tier::Original.stored(),
bytes.len() as i64,
now,
i64::from(pinned),
rel,
],
)?;
Ok(())
}
/// TRACES: FR-NC-6c | FR-NC-6a
/// Record an original this cache does **not** own the bytes of.
///
/// The virtual-filesystem case. On a library kept by a sync client the
/// original is materialised *in the library folder itself*, so copying it
/// under `originals/` would hold two copies of every pinned photograph —
/// and the copy would be the one the budget could evict while the real
/// disk cost stayed.
///
/// So the bytes are left where they are and only the bookkeeping is kept.
/// `path` is deliberately `NULL`, which is what makes this safe:
/// [`release`](Self::release) deletes the file a row names, and a row that
/// names none deletes nothing. **That matters more than it sounds.**
/// Deleting a materialised file inside a synced folder does not free a
/// cache — it deletes the photograph, and the client propagates that to
/// the server and to every other device. Handing the disk back is the
/// backend's job (`RemoteBackend::dematerialise`), not this one's.
///
/// `bytes` is what the original occupies where it lies, for the budget and
/// for reporting; pass 0 where it is not known.
pub fn record_in_place(
&self,
conn: &Connection,
image: ImageId,
bytes: u64,
pinned: bool,
now: i64,
) -> Result<(), CatalogError> {
conn.execute(
"INSERT INTO image_cache
(image_id, tier_actual, tier_desired, bytes, last_used, pinned, path)
VALUES (?1, ?2, ?2, ?3, ?4, ?5, NULL)
ON CONFLICT(image_id) DO UPDATE SET
tier_actual = ?2,
tier_desired = max(tier_desired, ?2),
bytes = ?3,
last_used = ?4,
pinned = max(pinned, ?5),
path = NULL",
rusqlite::params![
image.0 as i64,
Tier::Original.stored(),
bytes as i64,
now,
i64::from(pinned),
],
)?;
Ok(())
}
/// Read a cached original back, if it is here.
///
/// Touches `last_used`, which is what makes the eviction order reflect
/// actual use rather than download order. A read that finds the row but
/// not the file repairs the catalog rather than returning bytes it does
/// not have — the two can diverge if a user clears the directory by hand.
pub fn load(
&self,
conn: &Connection,
image: ImageId,
now: i64,
) -> Result<Option<Vec<u8>>, CatalogError> {
let path: Option<String> = conn
.query_row(
"SELECT path FROM image_cache
WHERE image_id = ?1 AND tier_actual >= ?2",
rusqlite::params![image.0 as i64, Tier::Original.stored()],
|r| r.get(0),
)
.ok()
.flatten();
let Some(rel) = path else { return Ok(None) };
let abs = self.dir.join(&rel);
match std::fs::read(&abs) {
Ok(bytes) => {
conn.execute(
"UPDATE image_cache SET last_used = ?2 WHERE image_id = ?1",
rusqlite::params![image.0 as i64, now],
)?;
Ok(Some(bytes))
}
Err(e) => {
// The file is gone but the row says it is here. Believing the
// row would report the image as locally available for ever
// while every open failed.
log::debug!(
"cached original {} missing, forgetting it: {e}",
abs.display()
);
self.forget(conn, &[image])?;
Ok(None)
}
}
}
/// Whether an image's original is on this device.
pub fn holds_original(&self, conn: &Connection, image: ImageId) -> bool {
conn.query_row(
"SELECT 1 FROM image_cache
WHERE image_id = ?1 AND tier_actual >= ?2",
rusqlite::params![image.0 as i64, Tier::Original.stored()],
|_| Ok(()),
)
.is_ok()
}
/// Mark images as pinned, so they are kept regardless of the budget.
///
/// Pinning records the *intent* — `tier_desired` — without downloading
/// anything: the download needs a network, which belongs to the caller.
/// An image already cached passively becomes pinned in place, keeping its
/// bytes rather than re-fetching them.
pub fn pin(&self, conn: &Connection, images: &[ImageId]) -> Result<usize, CatalogError> {
self.set_pinned(conn, images, true)
}
/// Release a pin, returning those images to the evictable population.
///
/// The bytes stay until eviction needs the room. Deleting immediately
/// would make unpinning destructive, when it is meant only to withdraw a
/// guarantee.
pub fn unpin(&self, conn: &Connection, images: &[ImageId]) -> Result<usize, CatalogError> {
self.set_pinned(conn, images, false)
}
/// Release the pin *and* delete the bytes it was holding.
///
/// The destructive half of the pair [`unpin`](Self::unpin) deliberately is
/// not. Unpinning answers "stop promising"; this answers "give me the disk
/// back", which is the question actually being asked when a trip is over
/// and the device is full. Leaving those gigabytes to sit until some future
/// eviction happens to want the room is not an answer to it.
///
/// Nothing is lost that cannot be fetched again: the original lives on the
/// server, and the catalog row, the ratings and the edit graph are all
/// untouched here — they are authoritative and small (FR-NC-6b).
///
/// Returns how many images were released and how many bytes that freed.
/// A file that has already vanished frees nothing and is still counted as
/// released, because the row describing it goes either way.
pub fn release(
&self,
conn: &Connection,
images: &[ImageId],
) -> Result<(usize, u64), CatalogError> {
if images.is_empty() {
return Ok((0, 0));
}
// Read the paths before the rows are rewritten: `forget` clears `path`,
// and a file whose name has been forgotten cannot be deleted.
let mut held = Vec::new();
{
let mut stmt = conn.prepare(
"SELECT bytes, path FROM image_cache
WHERE image_id = ?1 AND path IS NOT NULL",
)?;
for image in images {
if let Some(row) = stmt
.query_row(rusqlite::params![image.0 as i64], |r| {
Ok((r.get::<_, i64>(0)? as u64, r.get::<_, String>(1)?))
})
.optional()?
{
held.push(row);
}
}
}
let mut freed = 0u64;
for (bytes, rel) in &held {
let abs = self.dir.join(rel);
match std::fs::remove_file(&abs) {
Ok(()) => freed += bytes,
// Already gone is the ordinary case after a crash mid-write,
// not a failure: the row still has to go, or the cache accounts
// for space nothing occupies.
Err(e) => log::debug!("releasing {}: {e}", abs.display()),
}
}
// Unpin first, then forget. The other order would leave a pinned row
// claiming an original it no longer has, which `pending_pins` would
// then dutifully download again — the exact opposite of what was asked.
self.set_pinned(conn, images, false)?;
self.forget(conn, images)?;
Ok((images.len(), freed))
}
fn set_pinned(
&self,
conn: &Connection,
images: &[ImageId],
pinned: bool,
) -> Result<usize, CatalogError> {
if images.is_empty() {
return Ok(0);
}
let tx = conn.unchecked_transaction()?;
let mut n = 0;
for image in images {
n += tx.execute(
"INSERT INTO image_cache (image_id, tier_actual, tier_desired, bytes, pinned)
VALUES (?1, ?2, ?3, 0, ?4)
ON CONFLICT(image_id) DO UPDATE SET
pinned = ?4,
-- A pin raises the target; releasing one lowers it back to
-- whatever is actually held, so a released image is not
-- left permanently claiming it wants an original.
tier_desired = CASE WHEN ?4 = 1 THEN ?3 ELSE tier_actual END",
rusqlite::params![
image.0 as i64,
Tier::Metadata.stored(),
Tier::Original.stored(),
i64::from(pinned),
],
)?;
}
tx.commit()?;
Ok(n)
}
/// Images a pin wants but which are not yet downloaded.
///
/// The work list for whatever fetches originals. Ordered by id for a
/// stable, resumable sequence rather than an arbitrary one.
pub fn pending_pins(&self, conn: &Connection) -> Result<Vec<ImageId>, CatalogError> {
let mut stmt = conn.prepare(
"SELECT image_id FROM image_cache
WHERE pinned = 1 AND tier_actual < tier_desired
ORDER BY image_id",
)?;
let rows = stmt
.query_map([], |r| Ok(ImageId(r.get::<_, i64>(0)? as u64)))?
.collect::<Result<Vec<_>, _>>()?;
Ok(rows)
}
/// How full the cache is, split by population.
///
/// Counts only rows that actually hold an original: a pin that has not
/// downloaded yet occupies no disk, and counting its intent would evict
/// real files to make room for bytes that do not exist.
pub fn usage(&self, conn: &Connection) -> Result<Usage, CatalogError> {
let mut stmt = conn.prepare(
"SELECT pinned, count(*), coalesce(sum(bytes), 0)
FROM image_cache
WHERE tier_actual >= ?1
GROUP BY pinned",
)?;
let mut usage = Usage::default();
let rows = stmt.query_map(rusqlite::params![Tier::Original.stored()], |r| {
Ok((
r.get::<_, i64>(0)?,
r.get::<_, i64>(1)?,
r.get::<_, i64>(2)?,
))
})?;
for row in rows {
let (pinned, count, bytes) = row?;
if pinned == 1 {
usage.pinned_count = count as usize;
usage.pinned_bytes = bytes as u64;
} else {
usage.passive_count = count as usize;
usage.passive_bytes = bytes as u64;
}
}
Ok(usage)
}
/// Evict least-recently-used passive entries until the budget is met.
///
/// Returns how many images were dropped. Pinned entries are never
/// candidates, which is the guarantee that makes a pin worth making.
///
/// A row whose file has already vanished is still dropped from the
/// catalog: it frees no disk, but leaving it would let a phantom entry
/// hold the cache permanently over budget and evict real files in its
/// place.
pub fn enforce(&self, conn: &Connection) -> Result<usize, CatalogError> {
let usage = self.usage(conn)?;
let mut over = self.budget.overage(usage.passive_bytes);
if over == 0 {
return Ok(0);
}
// Oldest first. `last_used IS NULL` sorts first deliberately: a row
// that has never been read back is the least valuable thing here.
let mut stmt = conn.prepare(
"SELECT image_id, bytes, path FROM image_cache
WHERE pinned = 0 AND tier_actual >= ?1
ORDER BY last_used IS NULL DESC, last_used ASC",
)?;
let candidates = stmt
.query_map(rusqlite::params![Tier::Original.stored()], |r| {
Ok((
ImageId(r.get::<_, i64>(0)? as u64),
r.get::<_, i64>(1)? as u64,
r.get::<_, Option<String>>(2)?,
))
})?
.collect::<Result<Vec<_>, _>>()?;
let mut evicted = Vec::new();
for (image, bytes, path) in candidates {
if over == 0 {
break;
}
if let Some(rel) = path {
let abs = self.dir.join(rel);
if let Err(e) = std::fs::remove_file(&abs) {
// Already gone is the common case and not a failure; the
// row still has to go, or it accounts for space nothing
// occupies.
log::debug!("evicting {}: {e}", abs.display());
}
}
over = over.saturating_sub(bytes);
evicted.push(image);
}
let n = evicted.len();
self.forget(conn, &evicted)?;
Ok(n)
}
/// Drop cache rows, without touching files.
///
/// The row is reduced to `Metadata` rather than deleted, so a pin recorded
/// against it survives: unpinning is the only thing that should clear a
/// pin, and eviction of the bytes is not unpinning.
fn forget(&self, conn: &Connection, images: &[ImageId]) -> Result<(), CatalogError> {
if images.is_empty() {
return Ok(());
}
let tx = conn.unchecked_transaction()?;
for image in images {
tx.execute(
"UPDATE image_cache
SET tier_actual = ?2, bytes = 0, path = NULL
WHERE image_id = ?1",
rusqlite::params![image.0 as i64, Tier::Metadata.stored()],
)?;
}
tx.commit()?;
Ok(())
}
/// Everything currently held, newest use first. For a cache management view.
pub fn entries(&self, conn: &Connection) -> Result<Vec<Entry>, CatalogError> {
let mut stmt = conn.prepare(
"SELECT image_id, tier_actual, tier_desired, bytes, last_used, pinned, path
FROM image_cache
WHERE tier_actual >= ?1
ORDER BY last_used IS NULL, last_used DESC",
)?;
let rows = stmt
.query_map(rusqlite::params![Tier::Original.stored()], |r| {
Ok(Entry {
image: ImageId(r.get::<_, i64>(0)? as u64),
tier: Tier::from_stored(r.get(1)?),
desired: Tier::from_stored(r.get(2)?),
bytes: r.get::<_, i64>(3)? as u64,
last_used: r.get(4)?,
pinned: r.get::<_, i64>(5)? == 1,
path: r.get(6)?,
})
})?
.collect::<Result<Vec<_>, _>>()?;
Ok(rows)
}
}
#[cfg(test)]
mod tests {
use super::*;
use crate::Catalog;
/// Distinguishes concurrent fixtures. The harness runs tests in parallel,
/// and a shared directory would have one test's eviction delete another's
/// files.
static SEQ: std::sync::atomic::AtomicU64 = std::sync::atomic::AtomicU64::new(0);
/// A scratch directory that is fresh for each call.
fn tempdir() -> PathBuf {
let base = std::env::temp_dir().join(format!(
"dr-cache-test-{}-{}",
std::process::id(),
SEQ.fetch_add(1, std::sync::atomic::Ordering::Relaxed)
));
let _ = std::fs::remove_dir_all(&base);
std::fs::create_dir_all(&base).unwrap();
base
}
/// A catalog with `n` images, and a cache in a scratch directory.
fn fixture(n: usize) -> (Catalog, Cache, PathBuf, Vec<ImageId>) {
fixture_with(n, Budget::bytes(1000))
}
fn fixture_with(n: usize, budget: Budget) -> (Catalog, Cache, PathBuf, Vec<ImageId>) {
let catalog = Catalog::in_memory().unwrap();
catalog
.connection()
.execute(
"INSERT INTO roots (id, kind, label) VALUES (1, 'remote', 'test')",
[],
)
.unwrap();
let mut ids = Vec::new();
for i in 0..n {
catalog
.connection()
.execute(
"INSERT INTO images (root_id, source_ref, added_at)
VALUES (1, ?1, 0)",
rusqlite::params![format!("Photos/img{i:03}.CR2")],
)
.unwrap();
ids.push(ImageId(catalog.connection().last_insert_rowid() as u64));
}
let dir = tempdir();
let cache = Cache::open(&dir, budget).unwrap();
(catalog, cache, dir, ids)
}
#[test]
fn a_stored_original_reads_back() {
let (cat, cache, _dir, ids) = fixture(1);
cache
.store(cat.connection(), ids[0], "a.CR2", b"raw bytes", false, 10)
.unwrap();
assert!(cache.holds_original(cat.connection(), ids[0]));
assert_eq!(
cache.load(cat.connection(), ids[0], 20).unwrap().as_deref(),
Some(&b"raw bytes"[..])
);
}
#[test]
fn an_image_never_stored_is_absent() {
let (cat, cache, _dir, ids) = fixture(1);
assert!(!cache.holds_original(cat.connection(), ids[0]));
assert_eq!(cache.load(cat.connection(), ids[0], 0).unwrap(), None);
}
#[test]
fn eviction_takes_the_least_recently_used_first() {
let (cat, cache, _dir, ids) = fixture(3);
// 400 each against a 1000 budget: storing the third puts it 200 over.
let bytes = vec![0u8; 400];
cache
.store(cat.connection(), ids[0], "a.CR2", &bytes, false, 10)
.unwrap();
cache
.store(cat.connection(), ids[1], "b.CR2", &bytes, false, 20)
.unwrap();
cache
.store(cat.connection(), ids[2], "c.CR2", &bytes, false, 30)
.unwrap();
// Touch the oldest so it is no longer the least recently used.
cache.load(cat.connection(), ids[0], 40).unwrap();
assert_eq!(cache.enforce(cat.connection()).unwrap(), 1);
// ids[1] was the stalest by the time eviction ran.
assert!(!cache.holds_original(cat.connection(), ids[1]));
assert!(cache.holds_original(cat.connection(), ids[0]));
assert!(cache.holds_original(cat.connection(), ids[2]));
}
#[test]
fn a_pinned_original_is_never_evicted() {
// The guarantee the whole feature rests on: a pinned trip must still
// be there after a day of browsing pushes the cache over its cap.
let (cat, cache, _dir, ids) = fixture(3);
let bytes = vec![0u8; 800];
cache
.store(cat.connection(), ids[0], "a.CR2", &bytes, true, 10)
.unwrap();
cache
.store(cat.connection(), ids[1], "b.CR2", &bytes, false, 20)
.unwrap();
cache
.store(cat.connection(), ids[2], "c.CR2", &bytes, false, 30)
.unwrap();
cache.enforce(cat.connection()).unwrap();
assert!(
cache.holds_original(cat.connection(), ids[0]),
"the pinned original survives even though it is the oldest"
);
}
#[test]
fn releasing_a_pin_frees_the_disk_it_was_holding() {
// What "remove the local copies" has to mean. Unpinning alone leaves
// the bytes for a future eviction to notice, which is no answer at all
// to a device that is full now.
let (cat, cache, dir, ids) = fixture(2);
let bytes = vec![0u8; 700];
cache
.store(cat.connection(), ids[0], "a.CR2", &bytes, true, 10)
.unwrap();
cache
.store(cat.connection(), ids[1], "b.CR2", &bytes, true, 20)
.unwrap();
let (released, freed) = cache.release(cat.connection(), &ids).unwrap();
assert_eq!(released, 2);
assert_eq!(freed, 1400);
assert!(!cache.holds_original(cat.connection(), ids[0]));
assert_eq!(cache.usage(cat.connection()).unwrap().pinned_bytes, 0);
// The files themselves, not just the bookkeeping: a row cleared over a
// file still on disk is how a cache comes to hold gigabytes it does not
// know about.
let left: Vec<_> = walk_files(&dir);
assert!(left.is_empty(), "files remain on disk: {left:?}");
}
#[test]
fn a_released_pin_is_not_downloaded_all_over_again() {
// The failure mode of releasing in the wrong order: bytes deleted while
// the row still says an original is wanted, so the next pin fetch pulls
// the whole trip back down.
let (cat, cache, _dir, ids) = fixture(1);
cache
.store(cat.connection(), ids[0], "a.CR2", &[0u8; 100], true, 10)
.unwrap();
cache.release(cat.connection(), &ids).unwrap();
assert!(cache.pending_pins(cat.connection()).unwrap().is_empty());
}
/// Every file under `dir`, for asserting that a release left nothing.
fn walk_files(dir: &Path) -> Vec<PathBuf> {
let mut out = Vec::new();
let Ok(entries) = std::fs::read_dir(dir) else {
return out;
};
for entry in entries.flatten() {
let path = entry.path();
if path.is_dir() {
out.extend(walk_files(&path));
} else {
out.push(path);
}
}
out
}
#[test]
fn pinned_bytes_do_not_count_against_the_budget() {
// Otherwise a large pin starves the passive cache into evicting
// everything, and browsing becomes uncacheable the moment a trip is
// pinned.
let (cat, cache, _dir, ids) = fixture(2);
cache
.store(
cat.connection(),
ids[0],
"a.CR2",
&vec![0u8; 5000],
true,
10,
)
.unwrap();
cache
.store(
cat.connection(),
ids[1],
"b.CR2",
&vec![0u8; 500],
false,
20,
)
.unwrap();
// Pinned use is far past the 1000 budget, but the passive 500 fits.
assert_eq!(cache.enforce(cat.connection()).unwrap(), 0);
assert!(cache.holds_original(cat.connection(), ids[1]));
let usage = cache.usage(cat.connection()).unwrap();
assert_eq!(usage.pinned_bytes, 5000);
assert_eq!(usage.passive_bytes, 500);
}
#[test]
fn an_unlimited_budget_evicts_nothing() {
let (cat, cache, _dir, ids) = fixture_with(2, Budget::unlimited());
for (i, id) in ids.iter().enumerate() {
cache
.store(
cat.connection(),
*id,
"a.CR2",
&vec![0u8; 100_000],
false,
i as i64,
)
.unwrap();
}
assert_eq!(cache.enforce(cat.connection()).unwrap(), 0);
}
#[test]
fn pinning_records_intent_without_bytes() {
// A pin is not a download: it says what should be here, and something
// with a network makes it so.
let (cat, cache, _dir, ids) = fixture(2);
cache.pin(cat.connection(), &ids).unwrap();
assert!(!cache.holds_original(cat.connection(), ids[0]));
assert_eq!(cache.pending_pins(cat.connection()).unwrap(), ids);
assert_eq!(cache.usage(cat.connection()).unwrap().pinned_bytes, 0);
}
#[test]
fn a_downloaded_pin_stops_being_pending() {
let (cat, cache, _dir, ids) = fixture(2);
cache.pin(cat.connection(), &ids).unwrap();
cache
.store(cat.connection(), ids[0], "a.CR2", b"bytes", true, 10)
.unwrap();
assert_eq!(cache.pending_pins(cat.connection()).unwrap(), vec![ids[1]]);
}
#[test]
fn pinning_an_already_cached_image_keeps_its_bytes() {
// Re-downloading something already on disk because the user pinned it
// would be the most visible possible waste.
let (cat, cache, _dir, ids) = fixture(1);
cache
.store(cat.connection(), ids[0], "a.CR2", b"raw bytes", false, 10)
.unwrap();
cache.pin(cat.connection(), &ids).unwrap();
assert!(cache.pending_pins(cat.connection()).unwrap().is_empty());
assert_eq!(
cache.load(cat.connection(), ids[0], 20).unwrap().as_deref(),
Some(&b"raw bytes"[..])
);
assert_eq!(cache.usage(cat.connection()).unwrap().pinned_bytes, 9);
}
#[test]
fn opening_a_pinned_image_does_not_unpin_it() {
// The develop path stores passively. If that overwrote `pinned`, then
// simply *looking at* a pinned photograph would silently make it
// evictable — the pin would decay through use.
let (cat, cache, _dir, ids) = fixture(1);
cache.pin(cat.connection(), &ids).unwrap();
cache
.store(cat.connection(), ids[0], "a.CR2", b"bytes", false, 10)
.unwrap();
let entries = cache.entries(cat.connection()).unwrap();
assert!(entries[0].pinned, "still pinned after a passive store");
}
#[test]
fn unpinning_keeps_the_bytes_but_makes_them_evictable() {
let (cat, cache, _dir, ids) = fixture(2);
cache
.store(cat.connection(), ids[0], "a.CR2", &vec![0u8; 800], true, 10)
.unwrap();
cache.unpin(cat.connection(), &ids[0..1]).unwrap();
// Still here — unpinning withdraws a guarantee, it does not delete.
assert!(cache.holds_original(cat.connection(), ids[0]));
// But now it is a candidate.
cache
.store(
cat.connection(),
ids[1],
"b.CR2",
&vec![0u8; 800],
false,
20,
)
.unwrap();
assert_eq!(cache.enforce(cat.connection()).unwrap(), 1);
assert!(!cache.holds_original(cat.connection(), ids[0]));
}
#[test]
fn a_missing_file_is_forgotten_rather_than_reported_present() {
// A user clearing the cache directory by hand must not leave every
// image claiming to be local while every open fails.
let (cat, cache, dir, ids) = fixture_with(1, Budget::bytes(1000));
cache
.store(cat.connection(), ids[0], "a.CR2", b"bytes", false, 10)
.unwrap();
for entry in std::fs::read_dir(&dir).unwrap() {
std::fs::remove_file(entry.unwrap().path()).unwrap();
}
assert_eq!(cache.load(cat.connection(), ids[0], 20).unwrap(), None);
assert!(!cache.holds_original(cat.connection(), ids[0]));
}
#[test]
fn storing_the_same_image_twice_is_one_entry() {
let (cat, cache, _dir, ids) = fixture(1);
cache
.store(cat.connection(), ids[0], "a.CR2", b"first", false, 10)
.unwrap();
cache
.store(cat.connection(), ids[0], "a.CR2", b"second try", false, 20)
.unwrap();
let usage = cache.usage(cat.connection()).unwrap();
assert_eq!(usage.passive_count, 1);
assert_eq!(usage.passive_bytes, 10, "the later size, not the sum");
assert_eq!(
cache.load(cat.connection(), ids[0], 30).unwrap().as_deref(),
Some(&b"second try"[..])
);
}
#[test]
fn two_images_with_the_same_filename_do_not_collide() {
// `Photos/IMG_0001.CR2` and `Trips/IMG_0001.CR2` are different
// photographs; a cache keyed on the filename would serve one for the
// other, which is the worst failure this cache could have.
let (cat, cache, _dir, ids) = fixture(2);
cache
.store(
cat.connection(),
ids[0],
"Photos/IMG_0001.CR2",
b"first",
false,
10,
)
.unwrap();
cache
.store(
cat.connection(),
ids[1],
"Trips/IMG_0001.CR2",
b"second",
false,
20,
)
.unwrap();
assert_eq!(
cache.load(cat.connection(), ids[0], 30).unwrap().as_deref(),
Some(&b"first"[..])
);
assert_eq!(
cache.load(cat.connection(), ids[1], 30).unwrap().as_deref(),
Some(&b"second"[..])
);
}
#[test]
fn eviction_stops_once_the_budget_is_met() {
// Evicting everything on a small overage would throw away a working
// set to reclaim a few bytes.
let (cat, cache, _dir, ids) = fixture(3);
for (i, id) in ids.iter().enumerate() {
cache
.store(
cat.connection(),
*id,
"a.CR2",
&vec![0u8; 400],
false,
i as i64,
)
.unwrap();
}
// 1200 held against 1000: dropping one 400-byte entry suffices.
assert_eq!(cache.enforce(cat.connection()).unwrap(), 1);
assert_eq!(cache.usage(cat.connection()).unwrap().passive_count, 2);
}
#[test]
fn an_extensionless_source_still_gets_a_path() {
let (cat, cache, _dir, ids) = fixture(1);
cache
.store(
cat.connection(),
ids[0],
"Photos/no-extension",
b"bytes",
false,
10,
)
.unwrap();
assert_eq!(
cache.load(cat.connection(), ids[0], 20).unwrap().as_deref(),
Some(&b"bytes"[..])
);
}
}