Treat a placeholder as the photograph, not as a one-byte file

The folder connector was pointed at a Nextcloud VFS tree and got three
things wrong, the first of which loses work.

**A dehydrated sidecar read as absent.** `a.drsc` does not exist when the
client has dehydrated it — only `a.drsc.nextcloud` does — so `get` missed,
`.ok()` swallowed the `NotFound`, and the sidecar writer took that for
"there is no sidecar yet" and wrote a fresh document over the existing
one. Every edit another device had put there went with it. That function's
own doc comment calls this the exact loss the format's unknown-key
preservation exists to prevent.

**A stub was catalogued as a 1-byte image**, and ARCH §9.0 measured this
machine at 121,785 placeholders against 10,267 real files — so a folder
library on a synced tree was ~92% broken rows.

**Identity changed on hydration**, so downloading a photograph looked like
a delete and an add, orphaning its thumbnail and its face rows.

Entries now carry the photograph's own name and a `materialised` flag;
`get` on a stub returns the new `RemoteError::NotMaterialised`, which is
distinct from `NotFound` precisely because the sidecar writer must treat
them differently — it fetches the sidecar and merges, or leaves the entry
queued.

Hydration is a **borrow**. `BorrowPool` records what was on disk before it
asked, so `release_all` dehydrates only what a pass brought and leaves
what the user already had. Reference counted: the thumbnail pass and the
face pass meet on the same RAW, and without counting the first to finish
dehydrates the file the second is reading. A borrow against a plain folder
or a server does nothing, so a pass written for VFS runs everywhere.

Releasing means asking the client to dehydrate and never deleting: a
deletion inside a synced tree propagates to the server and removes the
photograph from every device.

Not a second backend — the capability is per *connection*, not per type,
since the same folder hydrates only while the client runs. The convention
arrives through a detector the registry supplies, so `dr-sync-folder`
still knows nothing about any client's protocol.

ARCH §9.0a records this as an amendment: finding 3 rejected hydration
because it costs 100× a range read, and that comparison assumed a
connector was available. A folder library has none.
This commit is contained in:
2026-08-29 09:57:52 +02:00
parent 6c363cee97
commit c102ba9df2
22 changed files with 1555 additions and 53 deletions
+264
View File
@@ -0,0 +1,264 @@
// TRACES: FR-NC-6c | FR-NC-6a
//! Hydrating a file for as long as it is needed, and no longer.
//!
//! A pass over a library — thumbnails, face indexing — needs each photograph's
//! bytes for a moment and never again. On a virtual-filesystem folder those
//! bytes may not be here, and fetching them is whole-file: hydrating a 17,000
//! image library to index it would land the entire library on a disk the user
//! deliberately keeps most of it off (ARCH §9.0).
//!
//! So hydration is a **borrow**. Ask for a file, use it, give it back. Peak
//! disk becomes the working set rather than the library, and the transfer is
//! paid once for a thumbnail that is then kept for ever — and pushed to the
//! server for other devices, which never pay it at all.
//!
//! # The rule that makes it safe
//!
//! **A file is returned to the state it was found in.** If it was already
//! downloaded — the user pinned it, opened it yesterday, or never uses VFS —
//! the borrow leaves it downloaded. Only what this pass hydrated is released.
//! Anything else silently undoes a choice the user made, and "my pinned trip
//! evaporated after an indexing run" is the kind of failure that makes people
//! stop trusting the feature.
//!
//! # Why it is reference counted
//!
//! Lanes run concurrently and two of them meet on the same file: the
//! thumbnail pass and the face pass want the same RAW. Without counting, the
//! first to finish dehydrates the file the second is reading. With it, the
//! transfer is paid once and the release happens when the last borrower is
//! done.
use std::collections::HashMap;
use std::sync::{Arc, Mutex};
use dr_sync::{RemoteBackend, RemoteError, RemoteId, RemotePath};
/// What a borrow is holding, per path.
#[derive(Debug, Default)]
struct Held {
/// How many borrowers are using it now.
borrowers: usize,
/// Whether *we* brought it here. False means it was already downloaded
/// and must be left that way.
ours: bool,
}
/// Tracks what has been hydrated and by whom.
///
/// Cheap to clone — every worker holds one and they share the same state.
#[derive(Clone, Default, Debug)]
pub struct BorrowPool {
held: Arc<Mutex<HashMap<RemotePath, Held>>>,
}
/// What a completed borrow did, for reporting a pass's real cost.
#[derive(Debug, Clone, Copy, Default, PartialEq, Eq)]
pub struct BorrowStats {
/// Files that were already here. These cost nothing.
pub already_local: usize,
/// Files this pass downloaded.
pub hydrated: usize,
/// Files released again afterwards.
pub released: usize,
/// Files left downloaded because they were already so.
pub kept: usize,
}
impl BorrowPool {
pub fn new() -> Self {
Self::default()
}
/// Borrow a file's content for the life of the returned guard.
///
/// Downloads it if it is a placeholder; does nothing if it is already
/// here. The guard releases it on drop, but only if this pool hydrated it
/// and nothing else still holds it.
///
/// A backend without [`Materialisation::OnDemand`] short-circuits: the
/// borrow succeeds and does nothing, so a caller written for a VFS library
/// runs unchanged against a server or a plain folder.
///
/// [`Materialisation::OnDemand`]: dr_sync::Materialisation::OnDemand
pub async fn borrow<'p>(
&'p self,
backend: &dyn RemoteBackend,
path: &RemotePath,
) -> Result<Borrowed<'p>, RemoteError> {
self.borrow_known(backend, path, None).await
}
/// [`borrow`](Self::borrow) where the caller already knows whether the
/// content is here.
///
/// A sweep reads availability out of the catalog on its way to building a
/// work list, so making the pool re-derive it costs a directory listing
/// per file for an answer already in hand.
///
/// **`None` means "find out", and being wrong is not symmetric.** Claiming
/// a file was already local leaves it downloaded — disk, and nothing else.
/// Claiming it was not releases a file the user may have pinned. So an
/// uncertain caller passes `Some(true)` or `None`, never a guess at
/// `false`.
pub async fn borrow_known<'p>(
&'p self,
backend: &dyn RemoteBackend,
path: &RemotePath,
already_local: Option<bool>,
) -> Result<Borrowed<'p>, RemoteError> {
if !backend.capabilities().materialisation.can_materialise() {
return Ok(Borrowed {
pool: None,
path: path.clone(),
hydrated: false,
});
}
// Another borrower already has it: join them rather than asking the
// client a second time.
{
let mut held = self.lock();
if let Some(entry) = held.get_mut(path) {
entry.borrowers += 1;
return Ok(Borrowed {
pool: Some(self),
path: path.clone(),
hydrated: false,
});
}
}
let id = RemoteId::Path(path.clone());
// What was here *before* we asked. The whole contract rests on this
// being read first: after `materialise` there is no way to tell what
// we brought from what was already there.
let was_local = match already_local {
Some(known) => known,
None => is_materialised(backend, path).await,
};
if !was_local {
backend.materialise(&id).await?;
}
self.lock().insert(
path.clone(),
Held {
borrowers: 1,
ours: !was_local,
},
);
Ok(Borrowed {
pool: Some(self),
path: path.clone(),
hydrated: !was_local,
})
}
/// Release everything this pool still holds that it hydrated.
///
/// The end-of-pass sweep. A guard dropped on a panicking worker cannot run
/// its async release, so the pool is drained deliberately at the end
/// rather than trusted to unwind cleanly.
pub async fn release_all(&self, backend: &dyn RemoteBackend) -> BorrowStats {
let ours: Vec<RemotePath> = {
let held = self.lock();
held.iter()
.filter(|(_, h)| h.ours)
.map(|(p, _)| p.clone())
.collect()
};
let mut stats = BorrowStats::default();
for path in ours {
match backend.dematerialise(&RemoteId::Path(path.clone())).await {
Ok(()) => stats.released += 1,
// Not fatal, and not worth failing a completed pass over: the
// content stays, which costs disk and loses nothing.
Err(e) => log::debug!("releasing {path}: {e}"),
}
}
self.lock().clear();
stats
}
/// How many paths are currently held.
pub fn held(&self) -> usize {
self.lock().len()
}
fn lock(&self) -> std::sync::MutexGuard<'_, HashMap<RemotePath, Held>> {
// A poisoned lock means a worker panicked while holding it. The map is
// bookkeeping, not a resource — carrying on with it is better than
// taking the whole pass down.
self.held.lock().unwrap_or_else(|e| e.into_inner())
}
}
/// A file held local for as long as this lives.
///
/// Dropping it marks the borrow finished. The actual release happens in
/// [`BorrowPool::release_all`], because dropping cannot await.
#[derive(Debug)]
pub struct Borrowed<'p> {
pool: Option<&'p BorrowPool>,
path: RemotePath,
/// Whether this borrow was the one that downloaded it.
hydrated: bool,
}
impl Borrowed<'_> {
/// Whether this borrow paid for a download.
pub fn hydrated(&self) -> bool {
self.hydrated
}
pub fn path(&self) -> &RemotePath {
&self.path
}
}
impl Drop for Borrowed<'_> {
fn drop(&mut self) {
let Some(pool) = self.pool else { return };
let mut held = pool.lock();
if let Some(entry) = held.get_mut(&self.path) {
entry.borrowers = entry.borrowers.saturating_sub(1);
// Left in the map even at zero borrowers: `release_all` needs to
// know it was ours, and a file wanted again a moment later should
// not be downloaded twice.
if entry.borrowers == 0 && !entry.ours {
held.remove(&self.path);
}
}
}
}
/// Whether the backend currently holds this file's content.
///
/// Asked by listing its parent, because the trait has no "stat one object" —
/// and adding one for this would put a method on every backend to serve a case
/// only one of them has.
async fn is_materialised(backend: &dyn RemoteBackend, path: &RemotePath) -> bool {
// `parent()` is `None` for a file directly under the library root, whose
// parent *is* the root — not "no parent". Treating the two alike reported
// every top-level photograph as already downloaded, so nothing was ever
// hydrated and nothing was ever released.
let parent = path.parent().unwrap_or_else(RemotePath::root);
match backend.list(&parent, None).await {
Ok(entries) => entries
.iter()
.find(|e| &e.path == path)
.map(|e| e.materialised)
// Not listed at all: nothing to hydrate, and the caller's own read
// will report the miss properly.
.unwrap_or(true),
Err(_) => true,
}
}
#[cfg(test)]
mod tests;
+242
View File
@@ -0,0 +1,242 @@
//! The borrow contract, against a filesystem and a fake client.
//!
//! The fake stands in for the sync client's socket, not for the filesystem:
//! it renames stubs exactly as suffix-mode VFS does, so everything under test
//! is the real path resolution and the real state tracking.
use super::*;
use crate::{FolderBackend, Vfs};
use std::path::{Path, PathBuf};
use std::sync::atomic::{AtomicUsize, Ordering};
/// A stand-in for a sync client, counting what it was asked to do.
struct FakeClient {
suffix: &'static str,
hydrations: AtomicUsize,
dehydrations: AtomicUsize,
/// When true, refuse to hydrate — the client is running but the server is
/// not reachable.
broken: bool,
}
impl FakeClient {
fn new() -> Arc<Self> {
Arc::new(Self {
suffix: ".nextcloud",
hydrations: AtomicUsize::new(0),
dehydrations: AtomicUsize::new(0),
broken: false,
})
}
fn broken() -> Arc<Self> {
Arc::new(Self {
suffix: ".nextcloud",
hydrations: AtomicUsize::new(0),
dehydrations: AtomicUsize::new(0),
broken: true,
})
}
}
impl Vfs for FakeClient {
fn name(&self) -> &'static str {
"fake"
}
fn is_placeholder(&self, on_disk: &str) -> bool {
on_disk.ends_with(self.suffix)
}
fn real_name<'a>(&self, on_disk: &'a str) -> &'a str {
on_disk.strip_suffix(self.suffix).unwrap_or(on_disk)
}
fn placeholder_name(&self, name: &str) -> std::borrow::Cow<'_, str> {
std::borrow::Cow::Owned(format!("{name}{}", self.suffix))
}
fn can_materialise(&self) -> bool {
true
}
fn materialise(&self, local: &Path) -> Result<(), RemoteError> {
self.hydrations.fetch_add(1, Ordering::SeqCst);
if self.broken {
return Err(RemoteError::Network("no server".into()));
}
// Suffix mode renames rather than filling in place, and writes the
// real content.
let real = PathBuf::from(local.to_string_lossy().strip_suffix(self.suffix).unwrap());
std::fs::write(&real, vec![9u8; 4096]).unwrap();
std::fs::remove_file(local).unwrap();
Ok(())
}
fn dematerialise(&self, local: &Path) -> Result<(), RemoteError> {
self.dehydrations.fetch_add(1, Ordering::SeqCst);
let stub = format!("{}{}", local.display(), self.suffix);
std::fs::write(&stub, [0u8]).unwrap();
std::fs::remove_file(local).unwrap();
Ok(())
}
}
struct Tmp(PathBuf);
impl Tmp {
fn new(name: &str) -> Self {
let d = std::env::temp_dir().join(format!("dr-borrow-{name}"));
let _ = std::fs::remove_dir_all(&d);
std::fs::create_dir_all(&d).unwrap();
Tmp(d)
}
/// A dehydrated photograph.
fn stub(&self, rel: &str) -> &Self {
std::fs::write(self.0.join(format!("{rel}.nextcloud")), [0u8]).unwrap();
self
}
/// One the user already has.
fn real(&self, rel: &str) -> &Self {
std::fs::write(self.0.join(rel), vec![1u8; 2048]).unwrap();
self
}
fn has(&self, rel: &str) -> bool {
self.0.join(rel).is_file()
}
fn backend(&self, vfs: Arc<dyn Vfs>) -> FolderBackend {
FolderBackend::with_vfs(&self.0, vfs).unwrap()
}
}
impl Drop for Tmp {
fn drop(&mut self) {
let _ = std::fs::remove_dir_all(&self.0);
}
}
#[tokio::test]
async fn a_borrowed_placeholder_is_downloaded_and_given_back() {
let t = Tmp::new("cycle");
t.stub("a.CR2");
let client = FakeClient::new();
let b = t.backend(client.clone());
let pool = BorrowPool::new();
let path = RemotePath::new("a.CR2");
{
let held = pool.borrow(&b, &path).await.unwrap();
assert!(held.hydrated(), "this borrow paid for it");
assert!(t.has("a.CR2"), "content is here while borrowed");
assert_eq!(
b.get(&RemoteId::Path(path.clone()), None)
.await
.unwrap()
.len(),
4096
);
}
let stats = pool.release_all(&b).await;
assert_eq!(stats.released, 1);
assert!(!t.has("a.CR2"), "given back");
assert!(t.has("a.CR2.nextcloud"), "a placeholder is left behind");
assert_eq!(client.dehydrations.load(Ordering::SeqCst), 1);
}
#[tokio::test]
async fn a_file_the_user_already_had_is_never_taken_away() {
// The rule the whole design rests on. Silently undoing a pin — or just a
// file someone opened yesterday — after an indexing run is the failure
// that would make people stop trusting this.
let t = Tmp::new("keep");
t.real("pinned.CR2");
let client = FakeClient::new();
let b = t.backend(client.clone());
let pool = BorrowPool::new();
{
let held = pool
.borrow(&b, &RemotePath::new("pinned.CR2"))
.await
.unwrap();
assert!(!held.hydrated(), "nothing was downloaded");
}
let stats = pool.release_all(&b).await;
assert_eq!(stats.released, 0);
assert!(t.has("pinned.CR2"), "still here");
assert_eq!(client.hydrations.load(Ordering::SeqCst), 0);
assert_eq!(client.dehydrations.load(Ordering::SeqCst), 0);
}
#[tokio::test]
async fn two_lanes_wanting_one_file_download_it_once() {
// The thumbnail pass and the face pass meet on the same RAW. Without
// counting, the first to finish dehydrates the file the second is reading.
let t = Tmp::new("shared");
t.stub("a.CR2");
let client = FakeClient::new();
let b = t.backend(client.clone());
let pool = BorrowPool::new();
let path = RemotePath::new("a.CR2");
let first = pool.borrow(&b, &path).await.unwrap();
let second = pool.borrow(&b, &path).await.unwrap();
assert_eq!(client.hydrations.load(Ordering::SeqCst), 1, "paid once");
drop(first);
assert!(t.has("a.CR2"), "still held by the second borrower");
drop(second);
pool.release_all(&b).await;
assert!(!t.has("a.CR2"));
}
#[tokio::test]
async fn a_failed_download_does_not_leave_a_phantom_borrow() {
// The client is up but the server is not. The pass must see the failure
// and the pool must not believe it holds anything.
let t = Tmp::new("failed");
t.stub("a.CR2");
let b = t.backend(FakeClient::broken());
let pool = BorrowPool::new();
let e = pool
.borrow(&b, &RemotePath::new("a.CR2"))
.await
.unwrap_err();
assert!(matches!(e, RemoteError::Network(_)), "{e:?}");
assert_eq!(pool.held(), 0);
assert!(t.has("a.CR2.nextcloud"), "left as it was found");
}
#[tokio::test]
async fn borrowing_against_a_plain_folder_does_nothing_at_all() {
// A caller written for a VFS library must run unchanged elsewhere, or
// every sweep grows two code paths.
let t = Tmp::new("plain");
t.real("a.CR2");
let b = FolderBackend::new(&t.0).unwrap();
let pool = BorrowPool::new();
let held = pool.borrow(&b, &RemotePath::new("a.CR2")).await.unwrap();
assert!(!held.hydrated());
drop(held);
assert_eq!(pool.release_all(&b).await.released, 0);
assert!(t.has("a.CR2"));
}
#[tokio::test]
async fn an_uncertain_caller_keeps_the_file_rather_than_releasing_it() {
// The asymmetry stated on `borrow_known`: claiming "already local" costs
// disk, claiming "not local" can release a pin.
let t = Tmp::new("uncertain");
t.real("a.CR2");
let client = FakeClient::new();
let b = t.backend(client.clone());
let pool = BorrowPool::new();
{
let _held = pool
.borrow_known(&b, &RemotePath::new("a.CR2"), Some(true))
.await
.unwrap();
}
pool.release_all(&b).await;
assert!(t.has("a.CR2"), "kept");
assert_eq!(client.dehydrations.load(Ordering::SeqCst), 0);
}
+238 -14
View File
@@ -50,23 +50,76 @@ use std::io::{Read, Seek, SeekFrom, Write};
use std::ops::Range;
use std::path::{Component, Path, PathBuf};
use std::sync::Arc;
use async_trait::async_trait;
use dr_sync::{
Account, BackendProvider, Capabilities, ChangeDetection, Connection, Cursor, EntryKind,
Precondition, RemoteBackend, RemoteChange, RemoteEntry, RemoteError, RemoteId, RemotePath,
ServerPreviews, SignIn, Validator,
Materialisation, Precondition, RemoteBackend, RemoteChange, RemoteEntry, RemoteError, RemoteId,
RemotePath, ServerPreviews, SignIn, Validator,
};
pub mod borrow;
pub mod vfs;
pub use borrow::{BorrowPool, BorrowStats, Borrowed};
pub use vfs::{NoVfs, Vfs};
/// The id written to [`Account::backend`] for a folder library.
///
/// On-disk configuration: changing it orphans every folder account.
pub const BACKEND_ID: &str = "folder";
/// TRACES: FR-NC-13
/// TRACES: FR-NC-13 | FR-NC-6c
/// Registers the folder connector.
///
/// See [`dr_sync::provider`] for what each method is for.
pub struct FolderProvider;
///
/// # The detector
///
/// This crate knows how to read a directory and nothing about sync clients,
/// so the placeholder convention arrives from outside: whoever registers the
/// provider supplies a function that recognises a synced folder and returns
/// the [`Vfs`] for it. That keeps `dr-sync-folder` free of any client's
/// protocol, and it is what lets one connector serve a plain disk, a Nextcloud
/// tree, and whatever comes next.
///
/// Detection runs per connection because the answer changes: the same
/// directory offers hydration while the client is up and not while it is down.
/// Recognises a placeholder convention in a directory, if any applies.
///
/// Runs per connection rather than once, because the answer changes: the same
/// folder offers hydration while the sync client is up and not while it is
/// down.
pub type VfsDetector = dyn Fn(&Path) -> Option<Arc<dyn Vfs>> + Send + Sync;
#[derive(Default)]
pub struct FolderProvider {
detect_vfs: Option<Box<VfsDetector>>,
}
impl FolderProvider {
/// A folder connector that treats every directory as ordinary.
pub fn new() -> Self {
Self::default()
}
/// A folder connector that recognises placeholder conventions.
pub fn with_vfs_detector(
detect: impl Fn(&Path) -> Option<Arc<dyn Vfs>> + Send + Sync + 'static,
) -> Self {
Self {
detect_vfs: Some(Box::new(detect)),
}
}
fn vfs_for(&self, root: &Path) -> Arc<dyn Vfs> {
self.detect_vfs
.as_ref()
.and_then(|d| d(root))
.unwrap_or_else(|| Arc::new(NoVfs))
}
}
impl BackendProvider for FolderProvider {
fn id(&self) -> &'static str {
@@ -136,18 +189,30 @@ impl BackendProvider for FolderProvider {
}
fn connect(&self, conn: &Connection) -> Result<Box<dyn RemoteBackend>, RemoteError> {
Ok(Box::new(FolderBackend::new(&conn.account.endpoint)?))
let root = Path::new(&conn.account.endpoint);
Ok(Box::new(FolderBackend::with_vfs(root, self.vfs_for(root))?))
}
}
/// TRACES: FR-NC-13 | FR-NC-4
/// TRACES: FR-NC-13 | FR-NC-4 | FR-NC-6c
/// A library rooted at a directory.
#[derive(Debug, Clone)]
#[derive(Clone)]
pub struct FolderBackend {
root: PathBuf,
/// The placeholder convention in force, [`NoVfs`] for an ordinary folder.
vfs: Arc<dyn Vfs>,
caps: Capabilities,
}
impl std::fmt::Debug for FolderBackend {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
f.debug_struct("FolderBackend")
.field("root", &self.root)
.field("vfs", &self.vfs.name())
.finish_non_exhaustive()
}
}
impl FolderBackend {
/// Open the folder at `root`.
///
@@ -156,6 +221,15 @@ impl FolderBackend {
/// [`RemoteError::Network`], which is what puts the app into offline mode
/// and leaves the catalog readable, exactly as a dead server does.
pub fn new(root: impl Into<PathBuf>) -> Result<Self, RemoteError> {
Self::with_vfs(root, Arc::new(NoVfs))
}
/// Open the folder at `root` under a placeholder convention.
///
/// The convention is chosen by the caller rather than sniffed here: the
/// connector that knows how to talk to a given sync client is the one that
/// knows whether it is running (see `dr_sync_nextcloud`).
pub fn with_vfs(root: impl Into<PathBuf>, vfs: Arc<dyn Vfs>) -> Result<Self, RemoteError> {
let root = root.into();
if !root.is_dir() {
return Err(RemoteError::Configuration(format!(
@@ -163,8 +237,19 @@ impl FolderBackend {
root.display()
)));
}
// Reported per connection, not per backend: the same folder offers
// hydration while the client is up and not while it is down, so this
// cannot be a constant of the type (see `vfs`).
let materialisation = if vfs.can_materialise() {
Materialisation::OnDemand
} else if vfs.name() == NoVfs.name() {
Materialisation::Always
} else {
Materialisation::Placeholders
};
Ok(Self {
root,
vfs,
caps: Capabilities {
// A directory's mtime describes its own entry list and nothing
// below it, so there is no propagation to exploit; the engine
@@ -179,6 +264,7 @@ impl FolderBackend {
bulk_upload: false,
conditional_write: true,
server_previews: ServerPreviews::None,
materialisation,
},
})
}
@@ -210,6 +296,30 @@ impl FolderBackend {
Ok(self.root.join(rel))
}
/// Where a photograph's bytes are on disk, and whether they are really
/// there.
///
/// A placeholder lives under a *different* name — suffix-mode VFS renames
/// on hydration rather than filling in place — so every read and write has
/// to look for both. The materialised name is tried first: it is the
/// common case, and the second `stat` is paid only when it misses.
///
/// Returns the path to use and whether it holds real content.
fn locate(&self, path: &RemotePath) -> Result<(PathBuf, bool), RemoteError> {
let direct = self.resolve(path)?;
if self.vfs.name() == NoVfs.name() || direct.exists() {
return Ok((direct, true));
}
let stub = self.resolve(&RemotePath::new(
self.vfs.placeholder_name(path.as_str()).into_owned(),
))?;
if stub.exists() {
return Ok((stub, false));
}
// Neither: genuinely missing. Report the name the caller asked for.
Ok((direct, true))
}
/// The local path a [`RemoteId`] names.
///
/// A stable id here is a hash and nothing can be resolved from it, exactly
@@ -223,6 +333,16 @@ impl FolderBackend {
)),
}
}
/// [`locate`](Self::locate) for an id.
fn locate_id(&self, id: &RemoteId) -> Result<(PathBuf, bool), RemoteError> {
match id {
RemoteId::Path(p) => self.locate(p),
RemoteId::Stable(_) => Err(RemoteError::Unsupported(
"a folder cannot be addressed by id; use RemoteId::Path",
)),
}
}
}
/// Run a filesystem operation off the async worker that asked for it.
@@ -274,6 +394,16 @@ fn map_io(e: std::io::Error, what: &str) -> RemoteError {
}
}
/// How long to wait for a requested download to land.
///
/// Generous, because the file may be tens of megabytes over a domestic
/// connection, and bounded, because a client that has stopped transferring
/// must not wedge a whole pass.
const MATERIALISE_TIMEOUT: std::time::Duration = std::time::Duration::from_secs(300);
/// How often to look for the materialised file while waiting.
const POLL: std::time::Duration = std::time::Duration::from_millis(200);
/// The identity of a file, from its path relative to the library root.
///
/// FNV-1a rather than `DefaultHasher`, whose output is explicitly unstable
@@ -330,6 +460,7 @@ impl RemoteBackend for FolderBackend {
) -> Result<Vec<RemoteEntry>, RemoteError> {
let local = self.resolve(dir)?;
let dir = dir.clone();
let vfs = self.vfs.clone();
blocking(move || {
let read =
std::fs::read_dir(&local).map_err(|e| map_io(e, &local.display().to_string()))?;
@@ -368,7 +499,13 @@ impl RemoteBackend for FolderBackend {
}
};
let path = dir.join(name);
// The photograph's own name, never the stub's. Identity is
// derived from it, so downloading a file must not look like a
// delete and an add — and `source_ref` must match what every
// other device calls the same photograph.
let stub = vfs.is_placeholder(name);
let path = dir.join(vfs.real_name(name));
out.push(RemoteEntry {
id: RemoteId::Stable(identity(&path)),
kind: if meta.is_dir() {
@@ -377,11 +514,15 @@ impl RemoteBackend for FolderBackend {
EntryKind::File
},
validator: validator_of(&meta),
size: meta.len(),
// A stub is one byte and says nothing about what it stands
// for. Reporting that byte count would put a 1-byte
// `file_size` in the catalog for most of the library.
size: if stub { 0 } else { meta.len() },
modified: modified_secs(&meta),
// No renderer behind a folder; previews are extracted
// locally from the file itself.
has_preview: false,
materialised: !stub,
path,
});
}
@@ -408,7 +549,14 @@ impl RemoteBackend for FolderBackend {
}
async fn get(&self, id: &RemoteId, range: Option<Range<u64>>) -> Result<Vec<u8>, RemoteError> {
let local = self.resolve_id(id)?;
let (local, materialised) = self.locate_id(id)?;
if !materialised {
// The one byte in the stub is not the file. Returning it produced
// a sidecar that parsed as empty and a thumbnail that never
// decoded; reporting `NotFound` made the sidecar writer treat an
// existing document as absent and overwrite it.
return Err(RemoteError::NotMaterialised(local.display().to_string()));
}
blocking(move || {
let what = local.display().to_string();
let mut file = std::fs::File::open(&local).map_err(|e| map_io(e, &what))?;
@@ -440,7 +588,15 @@ impl RemoteBackend for FolderBackend {
body: Vec<u8>,
precond: Option<Precondition>,
) -> Result<Validator, RemoteError> {
let local = self.resolve(path)?;
let (local, materialised) = self.locate(path)?;
if !materialised {
// Writing `a.drsc` while `a.drsc.nextcloud` sits beside it creates
// two files for one document and hands the sync client a conflict
// to resolve — in favour of whichever it sees last. The caller
// must materialise it and merge, which is what the typed error is
// for.
return Err(RemoteError::NotMaterialised(local.display().to_string()));
}
blocking(move || {
let what = local.display().to_string();
if let Some(parent) = local.parent() {
@@ -522,7 +678,10 @@ impl RemoteBackend for FolderBackend {
id: &RemoteId,
precond: Option<Precondition>,
) -> Result<(), RemoteError> {
let local = self.resolve_id(id)?;
// Deliberately by whichever name is on disk: deleting a photograph
// means deleting it whether or not its content happens to be here, and
// a stub left behind would be re-listed by the next scan.
let (local, _) = self.locate_id(id)?;
blocking(move || {
let what = local.display().to_string();
let meta = std::fs::symlink_metadata(&local).map_err(|e| map_io(e, &what))?;
@@ -562,8 +721,19 @@ impl RemoteBackend for FolderBackend {
}
async fn move_to(&self, from: &RemoteId, to: &RemotePath) -> Result<(), RemoteError> {
let src = self.resolve_id(from)?;
let dst = self.resolve(to)?;
// Move whichever name exists. Trashing a photograph that is not
// downloaded is a perfectly ordinary thing to do, and it must move the
// stub — renaming a placeholder keeps it a placeholder.
let (src, materialised) = self.locate_id(from)?;
let dst = if materialised {
self.resolve(to)?
} else {
// The destination keeps the placeholder suffix, or the client
// would see a one-byte file appear where a photograph should be.
self.resolve(&RemotePath::new(
self.vfs.placeholder_name(to.as_str()).into_owned(),
))?
};
blocking(move || {
let what = dst.display().to_string();
// Parents first: the trash folder does not exist until the first
@@ -597,6 +767,60 @@ impl RemoteBackend for FolderBackend {
.await
}
/// TRACES: FR-NC-6c
/// Ask the sync client to download a placeholder, and wait for it.
///
/// Suffix-mode VFS *renames* on hydration, so completion is the
/// materialised path appearing — not the stub changing size. Polling the
/// original would wait forever.
async fn materialise(&self, id: &RemoteId) -> Result<(), RemoteError> {
let (local, materialised) = self.locate_id(id)?;
if materialised {
// Already here. Not an error, and not a reason to ask again: the
// borrow pool relies on this being idempotent.
return Ok(());
}
let vfs = self.vfs.clone();
let target = self.resolve_id(id)?;
blocking(move || {
vfs.materialise(&local)?;
// The client acknowledges the command, not the transfer, so this
// waits for the file to appear. A bounded wait: a hydration that
// has not landed in this long is one the caller should be told
// about rather than blocked on for ever — the pass can come back
// to it.
let deadline = std::time::Instant::now() + MATERIALISE_TIMEOUT;
while std::time::Instant::now() < deadline {
if target.is_file() {
return Ok(());
}
std::thread::sleep(POLL);
}
Err(RemoteError::Network(format!(
"{} did not download within {}s",
target.display(),
MATERIALISE_TIMEOUT.as_secs()
)))
})
.await
}
/// TRACES: FR-NC-6c
/// Hand the content back, leaving a placeholder.
///
/// **Never a delete.** In a synced tree removing the file propagates the
/// removal to the server; the client is asked to dehydrate, and if it
/// cannot the content simply stays.
async fn dematerialise(&self, id: &RemoteId) -> Result<(), RemoteError> {
let (local, materialised) = self.locate_id(id)?;
if !materialised {
return Ok(());
}
let vfs = self.vfs.clone();
blocking(move || vfs.dematerialise(&local)).await
}
async fn create_dir(&self, path: &RemotePath) -> Result<(), RemoteError> {
let local = self.resolve(path)?;
blocking(move || {
+164 -4
View File
@@ -482,7 +482,7 @@ async fn an_upload_lands_where_the_engine_places_it() {
#[test]
fn an_endpoint_is_checked_before_an_account_is_written_for_it() {
let t = Tmp::new("provider");
let p = FolderProvider;
let p = FolderProvider::new();
assert!(p.normalise_endpoint(" ").is_err(), "empty");
assert!(p.normalise_endpoint("Pictures").is_err(), "relative");
@@ -504,7 +504,7 @@ fn two_spellings_of_one_folder_become_one_account() {
// Otherwise the same photographs are indexed twice, into two catalogs.
let t = Tmp::new("canonical");
t.file("sub/a.CR2", b"x");
let p = FolderProvider;
let p = FolderProvider::new();
let direct = p
.normalise_endpoint(&t.0.join("sub").to_string_lossy())
.unwrap();
@@ -516,7 +516,7 @@ fn two_spellings_of_one_folder_become_one_account() {
#[test]
fn a_folder_account_needs_no_credential() {
let p = FolderProvider;
let p = FolderProvider::new();
assert_eq!(p.sign_in(), SignIn::EndpointOnly);
assert!(!p.sign_in().needs_secret());
@@ -532,9 +532,169 @@ fn the_registry_opens_a_folder_account() {
// with nothing in between naming this crate.
let t = Tmp::new("registry");
let mut registry = dr_sync::BackendRegistry::new();
registry.register(std::sync::Arc::new(FolderProvider));
registry.register(std::sync::Arc::new(FolderProvider::new()));
let account = Account::new(BACKEND_ID, t.0.to_string_lossy());
let backend = registry.connect(&Connection::new(account, None)).unwrap();
assert_eq!(backend.name(), "Folder");
}
// --- virtual filesystems --------------------------------------------------
//
// A suffix-mode convention, matching the only one Linux supports. The
// behaviour under test is what the *backend* does with it; the borrow cycle
// has its own tests beside the pool.
struct SuffixVfs;
impl Vfs for SuffixVfs {
fn name(&self) -> &'static str {
"suffix"
}
fn is_placeholder(&self, on_disk: &str) -> bool {
on_disk.ends_with(".stub")
}
fn real_name<'a>(&self, on_disk: &'a str) -> &'a str {
on_disk.strip_suffix(".stub").unwrap_or(on_disk)
}
fn placeholder_name(&self, name: &str) -> std::borrow::Cow<'_, str> {
std::borrow::Cow::Owned(format!("{name}.stub"))
}
}
fn with_stubs(t: &Tmp) -> FolderBackend {
FolderBackend::with_vfs(&t.0, std::sync::Arc::new(SuffixVfs)).unwrap()
}
#[tokio::test]
async fn a_placeholder_is_listed_under_the_photographs_own_name() {
// The catalog records this as `source_ref`, and identity is derived from
// it. Reporting the stub's name gives the same photograph two identities
// and a name no other device recognises.
let t = Tmp::new("vfs-name");
t.file("shoot/IMG_0001.CR2.stub", &[0u8]);
let b = with_stubs(&t);
let entries = b.list(&RemotePath::new("shoot"), None).await.unwrap();
assert_eq!(entries[0].path.as_str(), "shoot/IMG_0001.CR2");
assert!(!entries[0].materialised, "the content is not here");
// One byte is not the photograph's size, and putting it in the catalog
// would claim a 30 MB RAW is a single byte.
assert_eq!(entries[0].size, 0, "unknown, not one");
}
#[tokio::test]
async fn identity_survives_a_download() {
// The failure this prevents: downloading a photograph looked like a
// delete and an add, which orphaned its thumbnail and its face rows.
let t = Tmp::new("vfs-identity");
t.file("a.CR2.stub", &[0u8]);
let b = with_stubs(&t);
let before = b.list(&RemotePath::root(), None).await.unwrap()[0]
.id
.clone();
std::fs::remove_file(t.0.join("a.CR2.stub")).unwrap();
std::fs::write(t.0.join("a.CR2"), vec![3u8; 4096]).unwrap();
let after = b.list(&RemotePath::root(), None).await.unwrap()[0]
.id
.clone();
assert_eq!(before, after, "the same photograph throughout");
}
#[tokio::test]
async fn reading_a_placeholder_is_distinguishable_from_a_missing_file() {
// The distinction the sidecar writer depends on: "not here" is fetchable
// and "not found" means create a new one. Conflating them overwrites an
// existing sidecar with a fresh document.
let t = Tmp::new("vfs-read");
t.file("a.drsc.stub", &[0u8]);
let b = with_stubs(&t);
let stub = b
.get(&RemoteId::Path(RemotePath::new("a.drsc")), None)
.await
.unwrap_err();
assert!(matches!(stub, RemoteError::NotMaterialised(_)), "{stub:?}");
let absent = b
.get(&RemoteId::Path(RemotePath::new("nothing.drsc")), None)
.await
.unwrap_err();
assert!(matches!(absent, RemoteError::NotFound(_)), "{absent:?}");
// And emphatically not the stub's one byte, which is what made a
// dehydrated sidecar parse as an empty document.
assert!(!matches!(stub, RemoteError::NotFound(_)));
}
#[tokio::test]
async fn writing_over_a_placeholder_is_refused() {
// Writing `a.drsc` beside `a.drsc.stub` makes two files for one document
// and hands the sync client a conflict it resolves arbitrarily.
let t = Tmp::new("vfs-write");
t.file("a.drsc.stub", &[0u8]);
let b = with_stubs(&t);
let e = b
.put(&RemotePath::new("a.drsc"), b"<new/>".to_vec(), None)
.await
.unwrap_err();
assert!(matches!(e, RemoteError::NotMaterialised(_)), "{e:?}");
assert!(!t.0.join("a.drsc").exists(), "no rival file created");
}
#[tokio::test]
async fn trashing_a_photograph_that_is_not_downloaded_moves_the_placeholder() {
// Culling without downloading is the ordinary way to use a VFS library.
// The stub has to move, and has to stay a stub — leaving it behind means
// the next scan re-lists the image and undoes the delete.
let t = Tmp::new("vfs-trash");
t.file("a.CR2.stub", &[0u8]);
let b = with_stubs(&t);
b.move_to(
&RemoteId::Path(RemotePath::new("a.CR2")),
&RemotePath::new(".darkroom-trash/a.CR2"),
)
.await
.unwrap();
assert!(!t.0.join("a.CR2.stub").exists());
assert!(
t.0.join(".darkroom-trash/a.CR2.stub").is_file(),
"still a stub"
);
}
#[tokio::test]
async fn a_folder_without_a_client_still_lists_and_reads_what_is_there() {
// No hydration available is a degraded mode, not a broken one: the
// materialised half of the library works completely.
let t = Tmp::new("vfs-degraded");
t.file("here.CR2", b"real").file("gone.CR2.stub", &[0u8]);
let b = with_stubs(&t);
assert_eq!(
b.capabilities().materialisation,
dr_sync::Materialisation::Placeholders,
"stubs exist and nothing can fetch them"
);
assert!(!b.capabilities().materialisation.can_materialise());
let got = b
.get(&RemoteId::Path(RemotePath::new("here.CR2")), None)
.await
.unwrap();
assert_eq!(got, b"real");
}
#[test]
fn a_plain_folder_reports_that_everything_it_lists_is_readable() {
let t = Tmp::new("vfs-plain");
assert_eq!(
t.backend().capabilities().materialisation,
dr_sync::Materialisation::Always
);
}
+131
View File
@@ -0,0 +1,131 @@
// TRACES: FR-NC-6c
//! Virtual-filesystem conventions layered over a directory.
//!
//! A sync client in virtual-files mode leaves a *placeholder* where a file is
//! catalogued but not downloaded. The folder is otherwise ordinary, so all of
//! [`FolderBackend`](crate::FolderBackend) applies — only three questions
//! differ, and they are the whole of this trait: what is a placeholder, what
//! is the photograph really called, and can the content be summoned.
//!
//! # Why this is not a separate backend
//!
//! It varies nothing about listing, reading, writing, moving or deleting — a
//! second connector would duplicate every one of those to change a name test.
//! More decisively, **the interesting capability is not a property of the
//! backend at all**: the same folder can materialise on demand while the sync
//! client is running and cannot when it is not, so it has to be computed per
//! connection either way. Registering a `folder-vfs` provider beside `folder`
//! would ask the user to choose between two things that differ by whether a
//! background process happens to be up.
//!
//! # Why the plain case is a `Vfs` too
//!
//! [`NoVfs`] answers "nothing is a placeholder" and refuses to materialise.
//! That keeps one code path through the backend rather than an `Option` tested
//! at every call site, and it is the shape a third convention — Dropbox,
//! OneDrive, macOS FileProvider — slots into.
//!
//! # What is *not* abstracted here
//!
//! Windows and macOS express placeholders in filesystem metadata rather than
//! in the name: a reparse point, or `st_blocks == 0` against a non-zero
//! `st_size`. That form needs a `Metadata` to answer, not a name, and the one
//! convention this project has met needs only a name. Widening the trait for a
//! platform nobody has run this on would be guessing at the shape.
use std::borrow::Cow;
use std::path::Path;
use dr_sync::RemoteError;
/// A placeholder convention, and what can be done about it.
///
/// Implementations are held behind an `Arc` and used from every worker
/// thread.
pub trait Vfs: Send + Sync {
/// A name for logs and the interface. "none", "Nextcloud".
fn name(&self) -> &'static str;
/// Whether a name **on disk** stands for content that is not here.
fn is_placeholder(&self, on_disk: &str) -> bool;
/// The photograph's own name, given whatever is on disk.
///
/// This is what the catalog records and what identity is derived from, so
/// a file keeps one name and one id across being downloaded and released.
/// Reporting the on-disk name instead makes hydration look like a delete
/// and an add.
fn real_name<'a>(&self, on_disk: &'a str) -> &'a str;
/// What a placeholder for `name` would be called on disk.
fn placeholder_name(&self, name: &str) -> Cow<'_, str>;
/// Whether content can actually be summoned right now.
///
/// False where the mechanism is absent — the client is not running, the
/// platform has no socket — which is an ordinary state and not an error.
/// The backend reports [`Materialisation::Placeholders`] rather than
/// [`OnDemand`] when this is false.
///
/// [`Materialisation::Placeholders`]: dr_sync::Materialisation::Placeholders
/// [`OnDemand`]: dr_sync::Materialisation::OnDemand
fn can_materialise(&self) -> bool {
false
}
/// Ask for a placeholder's content. Whole-file and slow.
fn materialise(&self, _local: &Path) -> Result<(), RemoteError> {
Err(RemoteError::Unsupported("this folder has no VFS client"))
}
/// Give the content back, leaving a placeholder.
///
/// **Must not delete.** In a synced tree a deletion propagates to the
/// server and removes the photograph from every device. An implementation
/// that cannot dehydrate returns `Unsupported`.
fn dematerialise(&self, _local: &Path) -> Result<(), RemoteError> {
Err(RemoteError::Unsupported("this folder has no VFS client"))
}
}
/// An ordinary directory: every file is what it appears to be.
#[derive(Debug, Clone, Copy, Default)]
pub struct NoVfs;
impl Vfs for NoVfs {
fn name(&self) -> &'static str {
"none"
}
fn is_placeholder(&self, _on_disk: &str) -> bool {
false
}
fn real_name<'a>(&self, on_disk: &'a str) -> &'a str {
on_disk
}
fn placeholder_name(&self, name: &str) -> Cow<'_, str> {
// Nothing is ever a placeholder here, so the only honest answer is
// the name itself — the backend will look for it, not find a second
// candidate, and report the file missing.
Cow::Owned(name.to_string())
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn a_plain_folder_has_no_placeholders_and_cannot_summon_anything() {
let v = NoVfs;
assert!(
!v.is_placeholder("IMG.CR2.nextcloud"),
"not this folder's convention"
);
assert_eq!(v.real_name("IMG.CR2"), "IMG.CR2");
assert!(!v.can_materialise());
assert!(v.materialise(Path::new("/x")).is_err());
// And it must refuse rather than approximate: deleting a file to
// "dehydrate" it would remove the photograph.
assert!(v.dematerialise(Path::new("/x")).is_err());
}
}