Treat a placeholder as the photograph, not as a one-byte file

The folder connector was pointed at a Nextcloud VFS tree and got three
things wrong, the first of which loses work.

**A dehydrated sidecar read as absent.** `a.drsc` does not exist when the
client has dehydrated it — only `a.drsc.nextcloud` does — so `get` missed,
`.ok()` swallowed the `NotFound`, and the sidecar writer took that for
"there is no sidecar yet" and wrote a fresh document over the existing
one. Every edit another device had put there went with it. That function's
own doc comment calls this the exact loss the format's unknown-key
preservation exists to prevent.

**A stub was catalogued as a 1-byte image**, and ARCH §9.0 measured this
machine at 121,785 placeholders against 10,267 real files — so a folder
library on a synced tree was ~92% broken rows.

**Identity changed on hydration**, so downloading a photograph looked like
a delete and an add, orphaning its thumbnail and its face rows.

Entries now carry the photograph's own name and a `materialised` flag;
`get` on a stub returns the new `RemoteError::NotMaterialised`, which is
distinct from `NotFound` precisely because the sidecar writer must treat
them differently — it fetches the sidecar and merges, or leaves the entry
queued.

Hydration is a **borrow**. `BorrowPool` records what was on disk before it
asked, so `release_all` dehydrates only what a pass brought and leaves
what the user already had. Reference counted: the thumbnail pass and the
face pass meet on the same RAW, and without counting the first to finish
dehydrates the file the second is reading. A borrow against a plain folder
or a server does nothing, so a pass written for VFS runs everywhere.

Releasing means asking the client to dehydrate and never deleting: a
deletion inside a synced tree propagates to the server and removes the
photograph from every device.

Not a second backend — the capability is per *connection*, not per type,
since the same folder hydrates only while the client runs. The convention
arrives through a detector the registry supplies, so `dr-sync-folder`
still knows nothing about any client's protocol.

ARCH §9.0a records this as an amendment: finding 3 rejected hydration
because it costs 100× a range read, and that comparison assumed a
connector was available. A folder library has none.
This commit is contained in:
2026-08-29 09:57:52 +02:00
parent 6c363cee97
commit c102ba9df2
22 changed files with 1555 additions and 53 deletions
+264
View File
@@ -0,0 +1,264 @@
// TRACES: FR-NC-6c | FR-NC-6a
//! Hydrating a file for as long as it is needed, and no longer.
//!
//! A pass over a library — thumbnails, face indexing — needs each photograph's
//! bytes for a moment and never again. On a virtual-filesystem folder those
//! bytes may not be here, and fetching them is whole-file: hydrating a 17,000
//! image library to index it would land the entire library on a disk the user
//! deliberately keeps most of it off (ARCH §9.0).
//!
//! So hydration is a **borrow**. Ask for a file, use it, give it back. Peak
//! disk becomes the working set rather than the library, and the transfer is
//! paid once for a thumbnail that is then kept for ever — and pushed to the
//! server for other devices, which never pay it at all.
//!
//! # The rule that makes it safe
//!
//! **A file is returned to the state it was found in.** If it was already
//! downloaded — the user pinned it, opened it yesterday, or never uses VFS —
//! the borrow leaves it downloaded. Only what this pass hydrated is released.
//! Anything else silently undoes a choice the user made, and "my pinned trip
//! evaporated after an indexing run" is the kind of failure that makes people
//! stop trusting the feature.
//!
//! # Why it is reference counted
//!
//! Lanes run concurrently and two of them meet on the same file: the
//! thumbnail pass and the face pass want the same RAW. Without counting, the
//! first to finish dehydrates the file the second is reading. With it, the
//! transfer is paid once and the release happens when the last borrower is
//! done.
use std::collections::HashMap;
use std::sync::{Arc, Mutex};
use dr_sync::{RemoteBackend, RemoteError, RemoteId, RemotePath};
/// What a borrow is holding, per path.
#[derive(Debug, Default)]
struct Held {
/// How many borrowers are using it now.
borrowers: usize,
/// Whether *we* brought it here. False means it was already downloaded
/// and must be left that way.
ours: bool,
}
/// Tracks what has been hydrated and by whom.
///
/// Cheap to clone — every worker holds one and they share the same state.
#[derive(Clone, Default, Debug)]
pub struct BorrowPool {
held: Arc<Mutex<HashMap<RemotePath, Held>>>,
}
/// What a completed borrow did, for reporting a pass's real cost.
#[derive(Debug, Clone, Copy, Default, PartialEq, Eq)]
pub struct BorrowStats {
/// Files that were already here. These cost nothing.
pub already_local: usize,
/// Files this pass downloaded.
pub hydrated: usize,
/// Files released again afterwards.
pub released: usize,
/// Files left downloaded because they were already so.
pub kept: usize,
}
impl BorrowPool {
pub fn new() -> Self {
Self::default()
}
/// Borrow a file's content for the life of the returned guard.
///
/// Downloads it if it is a placeholder; does nothing if it is already
/// here. The guard releases it on drop, but only if this pool hydrated it
/// and nothing else still holds it.
///
/// A backend without [`Materialisation::OnDemand`] short-circuits: the
/// borrow succeeds and does nothing, so a caller written for a VFS library
/// runs unchanged against a server or a plain folder.
///
/// [`Materialisation::OnDemand`]: dr_sync::Materialisation::OnDemand
pub async fn borrow<'p>(
&'p self,
backend: &dyn RemoteBackend,
path: &RemotePath,
) -> Result<Borrowed<'p>, RemoteError> {
self.borrow_known(backend, path, None).await
}
/// [`borrow`](Self::borrow) where the caller already knows whether the
/// content is here.
///
/// A sweep reads availability out of the catalog on its way to building a
/// work list, so making the pool re-derive it costs a directory listing
/// per file for an answer already in hand.
///
/// **`None` means "find out", and being wrong is not symmetric.** Claiming
/// a file was already local leaves it downloaded — disk, and nothing else.
/// Claiming it was not releases a file the user may have pinned. So an
/// uncertain caller passes `Some(true)` or `None`, never a guess at
/// `false`.
pub async fn borrow_known<'p>(
&'p self,
backend: &dyn RemoteBackend,
path: &RemotePath,
already_local: Option<bool>,
) -> Result<Borrowed<'p>, RemoteError> {
if !backend.capabilities().materialisation.can_materialise() {
return Ok(Borrowed {
pool: None,
path: path.clone(),
hydrated: false,
});
}
// Another borrower already has it: join them rather than asking the
// client a second time.
{
let mut held = self.lock();
if let Some(entry) = held.get_mut(path) {
entry.borrowers += 1;
return Ok(Borrowed {
pool: Some(self),
path: path.clone(),
hydrated: false,
});
}
}
let id = RemoteId::Path(path.clone());
// What was here *before* we asked. The whole contract rests on this
// being read first: after `materialise` there is no way to tell what
// we brought from what was already there.
let was_local = match already_local {
Some(known) => known,
None => is_materialised(backend, path).await,
};
if !was_local {
backend.materialise(&id).await?;
}
self.lock().insert(
path.clone(),
Held {
borrowers: 1,
ours: !was_local,
},
);
Ok(Borrowed {
pool: Some(self),
path: path.clone(),
hydrated: !was_local,
})
}
/// Release everything this pool still holds that it hydrated.
///
/// The end-of-pass sweep. A guard dropped on a panicking worker cannot run
/// its async release, so the pool is drained deliberately at the end
/// rather than trusted to unwind cleanly.
pub async fn release_all(&self, backend: &dyn RemoteBackend) -> BorrowStats {
let ours: Vec<RemotePath> = {
let held = self.lock();
held.iter()
.filter(|(_, h)| h.ours)
.map(|(p, _)| p.clone())
.collect()
};
let mut stats = BorrowStats::default();
for path in ours {
match backend.dematerialise(&RemoteId::Path(path.clone())).await {
Ok(()) => stats.released += 1,
// Not fatal, and not worth failing a completed pass over: the
// content stays, which costs disk and loses nothing.
Err(e) => log::debug!("releasing {path}: {e}"),
}
}
self.lock().clear();
stats
}
/// How many paths are currently held.
pub fn held(&self) -> usize {
self.lock().len()
}
fn lock(&self) -> std::sync::MutexGuard<'_, HashMap<RemotePath, Held>> {
// A poisoned lock means a worker panicked while holding it. The map is
// bookkeeping, not a resource — carrying on with it is better than
// taking the whole pass down.
self.held.lock().unwrap_or_else(|e| e.into_inner())
}
}
/// A file held local for as long as this lives.
///
/// Dropping it marks the borrow finished. The actual release happens in
/// [`BorrowPool::release_all`], because dropping cannot await.
#[derive(Debug)]
pub struct Borrowed<'p> {
pool: Option<&'p BorrowPool>,
path: RemotePath,
/// Whether this borrow was the one that downloaded it.
hydrated: bool,
}
impl Borrowed<'_> {
/// Whether this borrow paid for a download.
pub fn hydrated(&self) -> bool {
self.hydrated
}
pub fn path(&self) -> &RemotePath {
&self.path
}
}
impl Drop for Borrowed<'_> {
fn drop(&mut self) {
let Some(pool) = self.pool else { return };
let mut held = pool.lock();
if let Some(entry) = held.get_mut(&self.path) {
entry.borrowers = entry.borrowers.saturating_sub(1);
// Left in the map even at zero borrowers: `release_all` needs to
// know it was ours, and a file wanted again a moment later should
// not be downloaded twice.
if entry.borrowers == 0 && !entry.ours {
held.remove(&self.path);
}
}
}
}
/// Whether the backend currently holds this file's content.
///
/// Asked by listing its parent, because the trait has no "stat one object" —
/// and adding one for this would put a method on every backend to serve a case
/// only one of them has.
async fn is_materialised(backend: &dyn RemoteBackend, path: &RemotePath) -> bool {
// `parent()` is `None` for a file directly under the library root, whose
// parent *is* the root — not "no parent". Treating the two alike reported
// every top-level photograph as already downloaded, so nothing was ever
// hydrated and nothing was ever released.
let parent = path.parent().unwrap_or_else(RemotePath::root);
match backend.list(&parent, None).await {
Ok(entries) => entries
.iter()
.find(|e| &e.path == path)
.map(|e| e.materialised)
// Not listed at all: nothing to hydrate, and the caller's own read
// will report the miss properly.
.unwrap_or(true),
Err(_) => true,
}
}
#[cfg(test)]
mod tests;