The passes that need every photograph's bytes — thumbnails, face indexing — now borrow each one and release it at the end. On a placeholder library that is the difference between peak disk being the working set and being the whole library. Including on cancellation, which was nearly missed: the face sweep returns mid-loop when the user presses Stop, and without releasing there the disk is spent and nothing is delivered for it. `materialise` now answers whether *it* fetched the content. The pool used to work that out by listing a file's parent directory — one listing per file across a library — when the backend already had to `stat` it to decide whether to ask. One syscall instead of a directory walk, and it removes the bug class the tests found earlier: a file at the library root has no `parent()`, so every one of them read as already-downloaded. **Pinning is the retention control**, and it drives the model the catalog already had rather than a second one. `tier_desired` is what the user asked to keep hydrated, `pending_pins` is the resumable work list, and a pinned collection is never dehydrated for the same reason it was never evicted. It was in fact *broken* here before: `get` on a stub failed, and the pin worker logged "one unreadable file must not abandon the whole pin" and silently did nothing. Pinned originals on such a library are recorded with `path = NULL` (`Cache::record_in_place`) rather than copied under `originals/`. Two reasons, and the second is the important one. A copy would hold every pinned photograph twice, with the budget able to evict the half that was not costing the disk. And `release` deletes the file a row names — so a row that names none cannot delete anything, which puts the one catastrophic operation out of reach by construction rather than by remembering not to call it. Deleting a materialised file inside a synced tree removes the photograph from the server and every other device. Handing disk back is `spawn_dehydrate`, which asks the client. Two gaps written down rather than papered over (docs/storage.md §7): a hydrating pass cannot yet quote its cost, because a stub reports no size; and the two sweeps hold separate pools, so a library indexed for both fetches twice.
208 lines
7.3 KiB
Rust
208 lines
7.3 KiB
Rust
// TRACES: FR-NC-6c | FR-NC-6a
|
|
//! Hydrating a file for as long as it is needed, and no longer.
|
|
//!
|
|
//! A pass over a library — thumbnails, face indexing — needs each photograph's
|
|
//! bytes for a moment and never again. On a virtual-filesystem folder those
|
|
//! bytes may not be here, and fetching them is whole-file: hydrating a 17,000
|
|
//! image library to index it would land the entire library on a disk the user
|
|
//! deliberately keeps most of it off (ARCH §9.0).
|
|
//!
|
|
//! So hydration is a **borrow**. Ask for a file, use it, give it back. Peak
|
|
//! disk becomes the working set rather than the library, and the transfer is
|
|
//! paid once for a thumbnail that is then kept for ever — and pushed to the
|
|
//! server for other devices, which never pay it at all.
|
|
//!
|
|
//! # The rule that makes it safe
|
|
//!
|
|
//! **A file is returned to the state it was found in.** If it was already
|
|
//! downloaded — the user pinned it, opened it yesterday, or never uses VFS —
|
|
//! the borrow leaves it downloaded. Only what this pass hydrated is released.
|
|
//! Anything else silently undoes a choice the user made, and "my pinned trip
|
|
//! evaporated after an indexing run" is the kind of failure that makes people
|
|
//! stop trusting the feature.
|
|
//!
|
|
//! # Why it is reference counted
|
|
//!
|
|
//! Lanes run concurrently and two of them meet on the same file: the
|
|
//! thumbnail pass and the face pass want the same RAW. Without counting, the
|
|
//! first to finish dehydrates the file the second is reading. With it, the
|
|
//! transfer is paid once and the release happens when the last borrower is
|
|
//! done.
|
|
|
|
use std::collections::HashMap;
|
|
use std::sync::{Arc, Mutex};
|
|
|
|
use dr_sync::{RemoteBackend, RemoteError, RemoteId, RemotePath};
|
|
|
|
/// What a borrow is holding, per path.
|
|
#[derive(Debug, Default)]
|
|
struct Held {
|
|
/// How many borrowers are using it now.
|
|
borrowers: usize,
|
|
/// Whether *we* brought it here. False means it was already downloaded
|
|
/// and must be left that way.
|
|
ours: bool,
|
|
}
|
|
|
|
/// Tracks what has been hydrated and by whom.
|
|
///
|
|
/// Cheap to clone — every worker holds one and they share the same state.
|
|
#[derive(Clone, Default, Debug)]
|
|
pub struct BorrowPool {
|
|
held: Arc<Mutex<HashMap<RemotePath, Held>>>,
|
|
}
|
|
|
|
/// What a completed borrow did, for reporting a pass's real cost.
|
|
#[derive(Debug, Clone, Copy, Default, PartialEq, Eq)]
|
|
pub struct BorrowStats {
|
|
/// Files that were already here. These cost nothing.
|
|
pub already_local: usize,
|
|
/// Files this pass downloaded.
|
|
pub hydrated: usize,
|
|
/// Files released again afterwards.
|
|
pub released: usize,
|
|
/// Files left downloaded because they were already so.
|
|
pub kept: usize,
|
|
}
|
|
|
|
impl BorrowPool {
|
|
pub fn new() -> Self {
|
|
Self::default()
|
|
}
|
|
|
|
/// Borrow a file's content for the life of the returned guard.
|
|
///
|
|
/// Downloads it if it is a placeholder; does nothing if it is already
|
|
/// here. The guard releases it on drop, but only if this pool hydrated it
|
|
/// and nothing else still holds it.
|
|
///
|
|
/// A backend without [`Materialisation::OnDemand`] short-circuits: the
|
|
/// borrow succeeds and does nothing, so a caller written for a VFS library
|
|
/// runs unchanged against a server or a plain folder.
|
|
///
|
|
/// [`Materialisation::OnDemand`]: dr_sync::Materialisation::OnDemand
|
|
pub async fn borrow<'p>(
|
|
&'p self,
|
|
backend: &dyn RemoteBackend,
|
|
path: &RemotePath,
|
|
) -> Result<Borrowed<'p>, RemoteError> {
|
|
if !backend.capabilities().materialisation.can_materialise() {
|
|
return Ok(Borrowed {
|
|
pool: None,
|
|
path: path.clone(),
|
|
hydrated: false,
|
|
});
|
|
}
|
|
|
|
// Another borrower already has it: join them rather than asking the
|
|
// client a second time.
|
|
{
|
|
let mut held = self.lock();
|
|
if let Some(entry) = held.get_mut(path) {
|
|
entry.borrowers += 1;
|
|
return Ok(Borrowed {
|
|
pool: Some(self),
|
|
path: path.clone(),
|
|
hydrated: false,
|
|
});
|
|
}
|
|
}
|
|
|
|
// The backend answers whether *it* fetched the content, because it had
|
|
// to look before deciding. Determining that here instead would cost a
|
|
// directory listing per file, and getting it wrong in the wrong
|
|
// direction releases a file the user pinned.
|
|
let ours = backend.materialise(&RemoteId::Path(path.clone())).await?;
|
|
|
|
self.lock()
|
|
.insert(path.clone(), Held { borrowers: 1, ours });
|
|
|
|
Ok(Borrowed {
|
|
pool: Some(self),
|
|
path: path.clone(),
|
|
hydrated: ours,
|
|
})
|
|
}
|
|
|
|
/// Release everything this pool still holds that it hydrated.
|
|
///
|
|
/// The end-of-pass sweep. A guard dropped on a panicking worker cannot run
|
|
/// its async release, so the pool is drained deliberately at the end
|
|
/// rather than trusted to unwind cleanly.
|
|
pub async fn release_all(&self, backend: &dyn RemoteBackend) -> BorrowStats {
|
|
let ours: Vec<RemotePath> = {
|
|
let held = self.lock();
|
|
held.iter()
|
|
.filter(|(_, h)| h.ours)
|
|
.map(|(p, _)| p.clone())
|
|
.collect()
|
|
};
|
|
|
|
let mut stats = BorrowStats::default();
|
|
for path in ours {
|
|
match backend.dematerialise(&RemoteId::Path(path.clone())).await {
|
|
Ok(()) => stats.released += 1,
|
|
// Not fatal, and not worth failing a completed pass over: the
|
|
// content stays, which costs disk and loses nothing.
|
|
Err(e) => log::debug!("releasing {path}: {e}"),
|
|
}
|
|
}
|
|
self.lock().clear();
|
|
stats
|
|
}
|
|
|
|
/// How many paths are currently held.
|
|
pub fn held(&self) -> usize {
|
|
self.lock().len()
|
|
}
|
|
|
|
fn lock(&self) -> std::sync::MutexGuard<'_, HashMap<RemotePath, Held>> {
|
|
// A poisoned lock means a worker panicked while holding it. The map is
|
|
// bookkeeping, not a resource — carrying on with it is better than
|
|
// taking the whole pass down.
|
|
self.held.lock().unwrap_or_else(|e| e.into_inner())
|
|
}
|
|
}
|
|
|
|
/// A file held local for as long as this lives.
|
|
///
|
|
/// Dropping it marks the borrow finished. The actual release happens in
|
|
/// [`BorrowPool::release_all`], because dropping cannot await.
|
|
#[derive(Debug)]
|
|
pub struct Borrowed<'p> {
|
|
pool: Option<&'p BorrowPool>,
|
|
path: RemotePath,
|
|
/// Whether this borrow was the one that downloaded it.
|
|
hydrated: bool,
|
|
}
|
|
|
|
impl Borrowed<'_> {
|
|
/// Whether this borrow paid for a download.
|
|
pub fn hydrated(&self) -> bool {
|
|
self.hydrated
|
|
}
|
|
|
|
pub fn path(&self) -> &RemotePath {
|
|
&self.path
|
|
}
|
|
}
|
|
|
|
impl Drop for Borrowed<'_> {
|
|
fn drop(&mut self) {
|
|
let Some(pool) = self.pool else { return };
|
|
let mut held = pool.lock();
|
|
if let Some(entry) = held.get_mut(&self.path) {
|
|
entry.borrowers = entry.borrowers.saturating_sub(1);
|
|
// Left in the map even at zero borrowers: `release_all` needs to
|
|
// know it was ours, and a file wanted again a moment later should
|
|
// not be downloaded twice.
|
|
if entry.borrowers == 0 && !entry.ours {
|
|
held.remove(&self.path);
|
|
}
|
|
}
|
|
}
|
|
}
|
|
|
|
#[cfg(test)]
|
|
mod tests;
|