Treat a placeholder as the photograph, not as a one-byte file

The folder connector was pointed at a Nextcloud VFS tree and got three
things wrong, the first of which loses work.

**A dehydrated sidecar read as absent.** `a.drsc` does not exist when the
client has dehydrated it — only `a.drsc.nextcloud` does — so `get` missed,
`.ok()` swallowed the `NotFound`, and the sidecar writer took that for
"there is no sidecar yet" and wrote a fresh document over the existing
one. Every edit another device had put there went with it. That function's
own doc comment calls this the exact loss the format's unknown-key
preservation exists to prevent.

**A stub was catalogued as a 1-byte image**, and ARCH §9.0 measured this
machine at 121,785 placeholders against 10,267 real files — so a folder
library on a synced tree was ~92% broken rows.

**Identity changed on hydration**, so downloading a photograph looked like
a delete and an add, orphaning its thumbnail and its face rows.

Entries now carry the photograph's own name and a `materialised` flag;
`get` on a stub returns the new `RemoteError::NotMaterialised`, which is
distinct from `NotFound` precisely because the sidecar writer must treat
them differently — it fetches the sidecar and merges, or leaves the entry
queued.

Hydration is a **borrow**. `BorrowPool` records what was on disk before it
asked, so `release_all` dehydrates only what a pass brought and leaves
what the user already had. Reference counted: the thumbnail pass and the
face pass meet on the same RAW, and without counting the first to finish
dehydrates the file the second is reading. A borrow against a plain folder
or a server does nothing, so a pass written for VFS runs everywhere.

Releasing means asking the client to dehydrate and never deleting: a
deletion inside a synced tree propagates to the server and removes the
photograph from every device.

Not a second backend — the capability is per *connection*, not per type,
since the same folder hydrates only while the client runs. The convention
arrives through a detector the registry supplies, so `dr-sync-folder`
still knows nothing about any client's protocol.

ARCH §9.0a records this as an amendment: finding 3 rejected hydration
because it costs 100× a range read, and that comparison assumed a
connector was available. A folder library has none.
This commit is contained in:
2026-08-29 09:57:52 +02:00
parent 6c363cee97
commit c102ba9df2
22 changed files with 1555 additions and 53 deletions
+238 -14
View File
@@ -50,23 +50,76 @@ use std::io::{Read, Seek, SeekFrom, Write};
use std::ops::Range;
use std::path::{Component, Path, PathBuf};
use std::sync::Arc;
use async_trait::async_trait;
use dr_sync::{
Account, BackendProvider, Capabilities, ChangeDetection, Connection, Cursor, EntryKind,
Precondition, RemoteBackend, RemoteChange, RemoteEntry, RemoteError, RemoteId, RemotePath,
ServerPreviews, SignIn, Validator,
Materialisation, Precondition, RemoteBackend, RemoteChange, RemoteEntry, RemoteError, RemoteId,
RemotePath, ServerPreviews, SignIn, Validator,
};
pub mod borrow;
pub mod vfs;
pub use borrow::{BorrowPool, BorrowStats, Borrowed};
pub use vfs::{NoVfs, Vfs};
/// The id written to [`Account::backend`] for a folder library.
///
/// On-disk configuration: changing it orphans every folder account.
pub const BACKEND_ID: &str = "folder";
/// TRACES: FR-NC-13
/// TRACES: FR-NC-13 | FR-NC-6c
/// Registers the folder connector.
///
/// See [`dr_sync::provider`] for what each method is for.
pub struct FolderProvider;
///
/// # The detector
///
/// This crate knows how to read a directory and nothing about sync clients,
/// so the placeholder convention arrives from outside: whoever registers the
/// provider supplies a function that recognises a synced folder and returns
/// the [`Vfs`] for it. That keeps `dr-sync-folder` free of any client's
/// protocol, and it is what lets one connector serve a plain disk, a Nextcloud
/// tree, and whatever comes next.
///
/// Detection runs per connection because the answer changes: the same
/// directory offers hydration while the client is up and not while it is down.
/// Recognises a placeholder convention in a directory, if any applies.
///
/// Runs per connection rather than once, because the answer changes: the same
/// folder offers hydration while the sync client is up and not while it is
/// down.
pub type VfsDetector = dyn Fn(&Path) -> Option<Arc<dyn Vfs>> + Send + Sync;
#[derive(Default)]
pub struct FolderProvider {
detect_vfs: Option<Box<VfsDetector>>,
}
impl FolderProvider {
/// A folder connector that treats every directory as ordinary.
pub fn new() -> Self {
Self::default()
}
/// A folder connector that recognises placeholder conventions.
pub fn with_vfs_detector(
detect: impl Fn(&Path) -> Option<Arc<dyn Vfs>> + Send + Sync + 'static,
) -> Self {
Self {
detect_vfs: Some(Box::new(detect)),
}
}
fn vfs_for(&self, root: &Path) -> Arc<dyn Vfs> {
self.detect_vfs
.as_ref()
.and_then(|d| d(root))
.unwrap_or_else(|| Arc::new(NoVfs))
}
}
impl BackendProvider for FolderProvider {
fn id(&self) -> &'static str {
@@ -136,18 +189,30 @@ impl BackendProvider for FolderProvider {
}
fn connect(&self, conn: &Connection) -> Result<Box<dyn RemoteBackend>, RemoteError> {
Ok(Box::new(FolderBackend::new(&conn.account.endpoint)?))
let root = Path::new(&conn.account.endpoint);
Ok(Box::new(FolderBackend::with_vfs(root, self.vfs_for(root))?))
}
}
/// TRACES: FR-NC-13 | FR-NC-4
/// TRACES: FR-NC-13 | FR-NC-4 | FR-NC-6c
/// A library rooted at a directory.
#[derive(Debug, Clone)]
#[derive(Clone)]
pub struct FolderBackend {
root: PathBuf,
/// The placeholder convention in force, [`NoVfs`] for an ordinary folder.
vfs: Arc<dyn Vfs>,
caps: Capabilities,
}
impl std::fmt::Debug for FolderBackend {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
f.debug_struct("FolderBackend")
.field("root", &self.root)
.field("vfs", &self.vfs.name())
.finish_non_exhaustive()
}
}
impl FolderBackend {
/// Open the folder at `root`.
///
@@ -156,6 +221,15 @@ impl FolderBackend {
/// [`RemoteError::Network`], which is what puts the app into offline mode
/// and leaves the catalog readable, exactly as a dead server does.
pub fn new(root: impl Into<PathBuf>) -> Result<Self, RemoteError> {
Self::with_vfs(root, Arc::new(NoVfs))
}
/// Open the folder at `root` under a placeholder convention.
///
/// The convention is chosen by the caller rather than sniffed here: the
/// connector that knows how to talk to a given sync client is the one that
/// knows whether it is running (see `dr_sync_nextcloud`).
pub fn with_vfs(root: impl Into<PathBuf>, vfs: Arc<dyn Vfs>) -> Result<Self, RemoteError> {
let root = root.into();
if !root.is_dir() {
return Err(RemoteError::Configuration(format!(
@@ -163,8 +237,19 @@ impl FolderBackend {
root.display()
)));
}
// Reported per connection, not per backend: the same folder offers
// hydration while the client is up and not while it is down, so this
// cannot be a constant of the type (see `vfs`).
let materialisation = if vfs.can_materialise() {
Materialisation::OnDemand
} else if vfs.name() == NoVfs.name() {
Materialisation::Always
} else {
Materialisation::Placeholders
};
Ok(Self {
root,
vfs,
caps: Capabilities {
// A directory's mtime describes its own entry list and nothing
// below it, so there is no propagation to exploit; the engine
@@ -179,6 +264,7 @@ impl FolderBackend {
bulk_upload: false,
conditional_write: true,
server_previews: ServerPreviews::None,
materialisation,
},
})
}
@@ -210,6 +296,30 @@ impl FolderBackend {
Ok(self.root.join(rel))
}
/// Where a photograph's bytes are on disk, and whether they are really
/// there.
///
/// A placeholder lives under a *different* name — suffix-mode VFS renames
/// on hydration rather than filling in place — so every read and write has
/// to look for both. The materialised name is tried first: it is the
/// common case, and the second `stat` is paid only when it misses.
///
/// Returns the path to use and whether it holds real content.
fn locate(&self, path: &RemotePath) -> Result<(PathBuf, bool), RemoteError> {
let direct = self.resolve(path)?;
if self.vfs.name() == NoVfs.name() || direct.exists() {
return Ok((direct, true));
}
let stub = self.resolve(&RemotePath::new(
self.vfs.placeholder_name(path.as_str()).into_owned(),
))?;
if stub.exists() {
return Ok((stub, false));
}
// Neither: genuinely missing. Report the name the caller asked for.
Ok((direct, true))
}
/// The local path a [`RemoteId`] names.
///
/// A stable id here is a hash and nothing can be resolved from it, exactly
@@ -223,6 +333,16 @@ impl FolderBackend {
)),
}
}
/// [`locate`](Self::locate) for an id.
fn locate_id(&self, id: &RemoteId) -> Result<(PathBuf, bool), RemoteError> {
match id {
RemoteId::Path(p) => self.locate(p),
RemoteId::Stable(_) => Err(RemoteError::Unsupported(
"a folder cannot be addressed by id; use RemoteId::Path",
)),
}
}
}
/// Run a filesystem operation off the async worker that asked for it.
@@ -274,6 +394,16 @@ fn map_io(e: std::io::Error, what: &str) -> RemoteError {
}
}
/// How long to wait for a requested download to land.
///
/// Generous, because the file may be tens of megabytes over a domestic
/// connection, and bounded, because a client that has stopped transferring
/// must not wedge a whole pass.
const MATERIALISE_TIMEOUT: std::time::Duration = std::time::Duration::from_secs(300);
/// How often to look for the materialised file while waiting.
const POLL: std::time::Duration = std::time::Duration::from_millis(200);
/// The identity of a file, from its path relative to the library root.
///
/// FNV-1a rather than `DefaultHasher`, whose output is explicitly unstable
@@ -330,6 +460,7 @@ impl RemoteBackend for FolderBackend {
) -> Result<Vec<RemoteEntry>, RemoteError> {
let local = self.resolve(dir)?;
let dir = dir.clone();
let vfs = self.vfs.clone();
blocking(move || {
let read =
std::fs::read_dir(&local).map_err(|e| map_io(e, &local.display().to_string()))?;
@@ -368,7 +499,13 @@ impl RemoteBackend for FolderBackend {
}
};
let path = dir.join(name);
// The photograph's own name, never the stub's. Identity is
// derived from it, so downloading a file must not look like a
// delete and an add — and `source_ref` must match what every
// other device calls the same photograph.
let stub = vfs.is_placeholder(name);
let path = dir.join(vfs.real_name(name));
out.push(RemoteEntry {
id: RemoteId::Stable(identity(&path)),
kind: if meta.is_dir() {
@@ -377,11 +514,15 @@ impl RemoteBackend for FolderBackend {
EntryKind::File
},
validator: validator_of(&meta),
size: meta.len(),
// A stub is one byte and says nothing about what it stands
// for. Reporting that byte count would put a 1-byte
// `file_size` in the catalog for most of the library.
size: if stub { 0 } else { meta.len() },
modified: modified_secs(&meta),
// No renderer behind a folder; previews are extracted
// locally from the file itself.
has_preview: false,
materialised: !stub,
path,
});
}
@@ -408,7 +549,14 @@ impl RemoteBackend for FolderBackend {
}
async fn get(&self, id: &RemoteId, range: Option<Range<u64>>) -> Result<Vec<u8>, RemoteError> {
let local = self.resolve_id(id)?;
let (local, materialised) = self.locate_id(id)?;
if !materialised {
// The one byte in the stub is not the file. Returning it produced
// a sidecar that parsed as empty and a thumbnail that never
// decoded; reporting `NotFound` made the sidecar writer treat an
// existing document as absent and overwrite it.
return Err(RemoteError::NotMaterialised(local.display().to_string()));
}
blocking(move || {
let what = local.display().to_string();
let mut file = std::fs::File::open(&local).map_err(|e| map_io(e, &what))?;
@@ -440,7 +588,15 @@ impl RemoteBackend for FolderBackend {
body: Vec<u8>,
precond: Option<Precondition>,
) -> Result<Validator, RemoteError> {
let local = self.resolve(path)?;
let (local, materialised) = self.locate(path)?;
if !materialised {
// Writing `a.drsc` while `a.drsc.nextcloud` sits beside it creates
// two files for one document and hands the sync client a conflict
// to resolve — in favour of whichever it sees last. The caller
// must materialise it and merge, which is what the typed error is
// for.
return Err(RemoteError::NotMaterialised(local.display().to_string()));
}
blocking(move || {
let what = local.display().to_string();
if let Some(parent) = local.parent() {
@@ -522,7 +678,10 @@ impl RemoteBackend for FolderBackend {
id: &RemoteId,
precond: Option<Precondition>,
) -> Result<(), RemoteError> {
let local = self.resolve_id(id)?;
// Deliberately by whichever name is on disk: deleting a photograph
// means deleting it whether or not its content happens to be here, and
// a stub left behind would be re-listed by the next scan.
let (local, _) = self.locate_id(id)?;
blocking(move || {
let what = local.display().to_string();
let meta = std::fs::symlink_metadata(&local).map_err(|e| map_io(e, &what))?;
@@ -562,8 +721,19 @@ impl RemoteBackend for FolderBackend {
}
async fn move_to(&self, from: &RemoteId, to: &RemotePath) -> Result<(), RemoteError> {
let src = self.resolve_id(from)?;
let dst = self.resolve(to)?;
// Move whichever name exists. Trashing a photograph that is not
// downloaded is a perfectly ordinary thing to do, and it must move the
// stub — renaming a placeholder keeps it a placeholder.
let (src, materialised) = self.locate_id(from)?;
let dst = if materialised {
self.resolve(to)?
} else {
// The destination keeps the placeholder suffix, or the client
// would see a one-byte file appear where a photograph should be.
self.resolve(&RemotePath::new(
self.vfs.placeholder_name(to.as_str()).into_owned(),
))?
};
blocking(move || {
let what = dst.display().to_string();
// Parents first: the trash folder does not exist until the first
@@ -597,6 +767,60 @@ impl RemoteBackend for FolderBackend {
.await
}
/// TRACES: FR-NC-6c
/// Ask the sync client to download a placeholder, and wait for it.
///
/// Suffix-mode VFS *renames* on hydration, so completion is the
/// materialised path appearing — not the stub changing size. Polling the
/// original would wait forever.
async fn materialise(&self, id: &RemoteId) -> Result<(), RemoteError> {
let (local, materialised) = self.locate_id(id)?;
if materialised {
// Already here. Not an error, and not a reason to ask again: the
// borrow pool relies on this being idempotent.
return Ok(());
}
let vfs = self.vfs.clone();
let target = self.resolve_id(id)?;
blocking(move || {
vfs.materialise(&local)?;
// The client acknowledges the command, not the transfer, so this
// waits for the file to appear. A bounded wait: a hydration that
// has not landed in this long is one the caller should be told
// about rather than blocked on for ever — the pass can come back
// to it.
let deadline = std::time::Instant::now() + MATERIALISE_TIMEOUT;
while std::time::Instant::now() < deadline {
if target.is_file() {
return Ok(());
}
std::thread::sleep(POLL);
}
Err(RemoteError::Network(format!(
"{} did not download within {}s",
target.display(),
MATERIALISE_TIMEOUT.as_secs()
)))
})
.await
}
/// TRACES: FR-NC-6c
/// Hand the content back, leaving a placeholder.
///
/// **Never a delete.** In a synced tree removing the file propagates the
/// removal to the server; the client is asked to dehydrate, and if it
/// cannot the content simply stays.
async fn dematerialise(&self, id: &RemoteId) -> Result<(), RemoteError> {
let (local, materialised) = self.locate_id(id)?;
if !materialised {
return Ok(());
}
let vfs = self.vfs.clone();
blocking(move || vfs.dematerialise(&local)).await
}
async fn create_dir(&self, path: &RemotePath) -> Result<(), RemoteError> {
let local = self.resolve(path)?;
blocking(move || {