Treat a placeholder as the photograph, not as a one-byte file
The folder connector was pointed at a Nextcloud VFS tree and got three things wrong, the first of which loses work. **A dehydrated sidecar read as absent.** `a.drsc` does not exist when the client has dehydrated it — only `a.drsc.nextcloud` does — so `get` missed, `.ok()` swallowed the `NotFound`, and the sidecar writer took that for "there is no sidecar yet" and wrote a fresh document over the existing one. Every edit another device had put there went with it. That function's own doc comment calls this the exact loss the format's unknown-key preservation exists to prevent. **A stub was catalogued as a 1-byte image**, and ARCH §9.0 measured this machine at 121,785 placeholders against 10,267 real files — so a folder library on a synced tree was ~92% broken rows. **Identity changed on hydration**, so downloading a photograph looked like a delete and an add, orphaning its thumbnail and its face rows. Entries now carry the photograph's own name and a `materialised` flag; `get` on a stub returns the new `RemoteError::NotMaterialised`, which is distinct from `NotFound` precisely because the sidecar writer must treat them differently — it fetches the sidecar and merges, or leaves the entry queued. Hydration is a **borrow**. `BorrowPool` records what was on disk before it asked, so `release_all` dehydrates only what a pass brought and leaves what the user already had. Reference counted: the thumbnail pass and the face pass meet on the same RAW, and without counting the first to finish dehydrates the file the second is reading. A borrow against a plain folder or a server does nothing, so a pass written for VFS runs everywhere. Releasing means asking the client to dehydrate and never deleting: a deletion inside a synced tree propagates to the server and removes the photograph from every device. Not a second backend — the capability is per *connection*, not per type, since the same folder hydrates only while the client runs. The convention arrives through a detector the registry supplies, so `dr-sync-folder` still knows nothing about any client's protocol. ARCH §9.0a records this as an amendment: finding 3 rejected hydration because it costs 100× a range read, and that comparison assumed a connector was available. A folder library has none.
This commit is contained in:
+238
-14
@@ -50,23 +50,76 @@ use std::io::{Read, Seek, SeekFrom, Write};
|
||||
use std::ops::Range;
|
||||
use std::path::{Component, Path, PathBuf};
|
||||
|
||||
use std::sync::Arc;
|
||||
|
||||
use async_trait::async_trait;
|
||||
use dr_sync::{
|
||||
Account, BackendProvider, Capabilities, ChangeDetection, Connection, Cursor, EntryKind,
|
||||
Precondition, RemoteBackend, RemoteChange, RemoteEntry, RemoteError, RemoteId, RemotePath,
|
||||
ServerPreviews, SignIn, Validator,
|
||||
Materialisation, Precondition, RemoteBackend, RemoteChange, RemoteEntry, RemoteError, RemoteId,
|
||||
RemotePath, ServerPreviews, SignIn, Validator,
|
||||
};
|
||||
|
||||
pub mod borrow;
|
||||
pub mod vfs;
|
||||
|
||||
pub use borrow::{BorrowPool, BorrowStats, Borrowed};
|
||||
pub use vfs::{NoVfs, Vfs};
|
||||
|
||||
/// The id written to [`Account::backend`] for a folder library.
|
||||
///
|
||||
/// On-disk configuration: changing it orphans every folder account.
|
||||
pub const BACKEND_ID: &str = "folder";
|
||||
|
||||
/// TRACES: FR-NC-13
|
||||
/// TRACES: FR-NC-13 | FR-NC-6c
|
||||
/// Registers the folder connector.
|
||||
///
|
||||
/// See [`dr_sync::provider`] for what each method is for.
|
||||
pub struct FolderProvider;
|
||||
///
|
||||
/// # The detector
|
||||
///
|
||||
/// This crate knows how to read a directory and nothing about sync clients,
|
||||
/// so the placeholder convention arrives from outside: whoever registers the
|
||||
/// provider supplies a function that recognises a synced folder and returns
|
||||
/// the [`Vfs`] for it. That keeps `dr-sync-folder` free of any client's
|
||||
/// protocol, and it is what lets one connector serve a plain disk, a Nextcloud
|
||||
/// tree, and whatever comes next.
|
||||
///
|
||||
/// Detection runs per connection because the answer changes: the same
|
||||
/// directory offers hydration while the client is up and not while it is down.
|
||||
/// Recognises a placeholder convention in a directory, if any applies.
|
||||
///
|
||||
/// Runs per connection rather than once, because the answer changes: the same
|
||||
/// folder offers hydration while the sync client is up and not while it is
|
||||
/// down.
|
||||
pub type VfsDetector = dyn Fn(&Path) -> Option<Arc<dyn Vfs>> + Send + Sync;
|
||||
|
||||
#[derive(Default)]
|
||||
pub struct FolderProvider {
|
||||
detect_vfs: Option<Box<VfsDetector>>,
|
||||
}
|
||||
|
||||
impl FolderProvider {
|
||||
/// A folder connector that treats every directory as ordinary.
|
||||
pub fn new() -> Self {
|
||||
Self::default()
|
||||
}
|
||||
|
||||
/// A folder connector that recognises placeholder conventions.
|
||||
pub fn with_vfs_detector(
|
||||
detect: impl Fn(&Path) -> Option<Arc<dyn Vfs>> + Send + Sync + 'static,
|
||||
) -> Self {
|
||||
Self {
|
||||
detect_vfs: Some(Box::new(detect)),
|
||||
}
|
||||
}
|
||||
|
||||
fn vfs_for(&self, root: &Path) -> Arc<dyn Vfs> {
|
||||
self.detect_vfs
|
||||
.as_ref()
|
||||
.and_then(|d| d(root))
|
||||
.unwrap_or_else(|| Arc::new(NoVfs))
|
||||
}
|
||||
}
|
||||
|
||||
impl BackendProvider for FolderProvider {
|
||||
fn id(&self) -> &'static str {
|
||||
@@ -136,18 +189,30 @@ impl BackendProvider for FolderProvider {
|
||||
}
|
||||
|
||||
fn connect(&self, conn: &Connection) -> Result<Box<dyn RemoteBackend>, RemoteError> {
|
||||
Ok(Box::new(FolderBackend::new(&conn.account.endpoint)?))
|
||||
let root = Path::new(&conn.account.endpoint);
|
||||
Ok(Box::new(FolderBackend::with_vfs(root, self.vfs_for(root))?))
|
||||
}
|
||||
}
|
||||
|
||||
/// TRACES: FR-NC-13 | FR-NC-4
|
||||
/// TRACES: FR-NC-13 | FR-NC-4 | FR-NC-6c
|
||||
/// A library rooted at a directory.
|
||||
#[derive(Debug, Clone)]
|
||||
#[derive(Clone)]
|
||||
pub struct FolderBackend {
|
||||
root: PathBuf,
|
||||
/// The placeholder convention in force, [`NoVfs`] for an ordinary folder.
|
||||
vfs: Arc<dyn Vfs>,
|
||||
caps: Capabilities,
|
||||
}
|
||||
|
||||
impl std::fmt::Debug for FolderBackend {
|
||||
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
|
||||
f.debug_struct("FolderBackend")
|
||||
.field("root", &self.root)
|
||||
.field("vfs", &self.vfs.name())
|
||||
.finish_non_exhaustive()
|
||||
}
|
||||
}
|
||||
|
||||
impl FolderBackend {
|
||||
/// Open the folder at `root`.
|
||||
///
|
||||
@@ -156,6 +221,15 @@ impl FolderBackend {
|
||||
/// [`RemoteError::Network`], which is what puts the app into offline mode
|
||||
/// and leaves the catalog readable, exactly as a dead server does.
|
||||
pub fn new(root: impl Into<PathBuf>) -> Result<Self, RemoteError> {
|
||||
Self::with_vfs(root, Arc::new(NoVfs))
|
||||
}
|
||||
|
||||
/// Open the folder at `root` under a placeholder convention.
|
||||
///
|
||||
/// The convention is chosen by the caller rather than sniffed here: the
|
||||
/// connector that knows how to talk to a given sync client is the one that
|
||||
/// knows whether it is running (see `dr_sync_nextcloud`).
|
||||
pub fn with_vfs(root: impl Into<PathBuf>, vfs: Arc<dyn Vfs>) -> Result<Self, RemoteError> {
|
||||
let root = root.into();
|
||||
if !root.is_dir() {
|
||||
return Err(RemoteError::Configuration(format!(
|
||||
@@ -163,8 +237,19 @@ impl FolderBackend {
|
||||
root.display()
|
||||
)));
|
||||
}
|
||||
// Reported per connection, not per backend: the same folder offers
|
||||
// hydration while the client is up and not while it is down, so this
|
||||
// cannot be a constant of the type (see `vfs`).
|
||||
let materialisation = if vfs.can_materialise() {
|
||||
Materialisation::OnDemand
|
||||
} else if vfs.name() == NoVfs.name() {
|
||||
Materialisation::Always
|
||||
} else {
|
||||
Materialisation::Placeholders
|
||||
};
|
||||
Ok(Self {
|
||||
root,
|
||||
vfs,
|
||||
caps: Capabilities {
|
||||
// A directory's mtime describes its own entry list and nothing
|
||||
// below it, so there is no propagation to exploit; the engine
|
||||
@@ -179,6 +264,7 @@ impl FolderBackend {
|
||||
bulk_upload: false,
|
||||
conditional_write: true,
|
||||
server_previews: ServerPreviews::None,
|
||||
materialisation,
|
||||
},
|
||||
})
|
||||
}
|
||||
@@ -210,6 +296,30 @@ impl FolderBackend {
|
||||
Ok(self.root.join(rel))
|
||||
}
|
||||
|
||||
/// Where a photograph's bytes are on disk, and whether they are really
|
||||
/// there.
|
||||
///
|
||||
/// A placeholder lives under a *different* name — suffix-mode VFS renames
|
||||
/// on hydration rather than filling in place — so every read and write has
|
||||
/// to look for both. The materialised name is tried first: it is the
|
||||
/// common case, and the second `stat` is paid only when it misses.
|
||||
///
|
||||
/// Returns the path to use and whether it holds real content.
|
||||
fn locate(&self, path: &RemotePath) -> Result<(PathBuf, bool), RemoteError> {
|
||||
let direct = self.resolve(path)?;
|
||||
if self.vfs.name() == NoVfs.name() || direct.exists() {
|
||||
return Ok((direct, true));
|
||||
}
|
||||
let stub = self.resolve(&RemotePath::new(
|
||||
self.vfs.placeholder_name(path.as_str()).into_owned(),
|
||||
))?;
|
||||
if stub.exists() {
|
||||
return Ok((stub, false));
|
||||
}
|
||||
// Neither: genuinely missing. Report the name the caller asked for.
|
||||
Ok((direct, true))
|
||||
}
|
||||
|
||||
/// The local path a [`RemoteId`] names.
|
||||
///
|
||||
/// A stable id here is a hash and nothing can be resolved from it, exactly
|
||||
@@ -223,6 +333,16 @@ impl FolderBackend {
|
||||
)),
|
||||
}
|
||||
}
|
||||
|
||||
/// [`locate`](Self::locate) for an id.
|
||||
fn locate_id(&self, id: &RemoteId) -> Result<(PathBuf, bool), RemoteError> {
|
||||
match id {
|
||||
RemoteId::Path(p) => self.locate(p),
|
||||
RemoteId::Stable(_) => Err(RemoteError::Unsupported(
|
||||
"a folder cannot be addressed by id; use RemoteId::Path",
|
||||
)),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Run a filesystem operation off the async worker that asked for it.
|
||||
@@ -274,6 +394,16 @@ fn map_io(e: std::io::Error, what: &str) -> RemoteError {
|
||||
}
|
||||
}
|
||||
|
||||
/// How long to wait for a requested download to land.
|
||||
///
|
||||
/// Generous, because the file may be tens of megabytes over a domestic
|
||||
/// connection, and bounded, because a client that has stopped transferring
|
||||
/// must not wedge a whole pass.
|
||||
const MATERIALISE_TIMEOUT: std::time::Duration = std::time::Duration::from_secs(300);
|
||||
|
||||
/// How often to look for the materialised file while waiting.
|
||||
const POLL: std::time::Duration = std::time::Duration::from_millis(200);
|
||||
|
||||
/// The identity of a file, from its path relative to the library root.
|
||||
///
|
||||
/// FNV-1a rather than `DefaultHasher`, whose output is explicitly unstable
|
||||
@@ -330,6 +460,7 @@ impl RemoteBackend for FolderBackend {
|
||||
) -> Result<Vec<RemoteEntry>, RemoteError> {
|
||||
let local = self.resolve(dir)?;
|
||||
let dir = dir.clone();
|
||||
let vfs = self.vfs.clone();
|
||||
blocking(move || {
|
||||
let read =
|
||||
std::fs::read_dir(&local).map_err(|e| map_io(e, &local.display().to_string()))?;
|
||||
@@ -368,7 +499,13 @@ impl RemoteBackend for FolderBackend {
|
||||
}
|
||||
};
|
||||
|
||||
let path = dir.join(name);
|
||||
// The photograph's own name, never the stub's. Identity is
|
||||
// derived from it, so downloading a file must not look like a
|
||||
// delete and an add — and `source_ref` must match what every
|
||||
// other device calls the same photograph.
|
||||
let stub = vfs.is_placeholder(name);
|
||||
let path = dir.join(vfs.real_name(name));
|
||||
|
||||
out.push(RemoteEntry {
|
||||
id: RemoteId::Stable(identity(&path)),
|
||||
kind: if meta.is_dir() {
|
||||
@@ -377,11 +514,15 @@ impl RemoteBackend for FolderBackend {
|
||||
EntryKind::File
|
||||
},
|
||||
validator: validator_of(&meta),
|
||||
size: meta.len(),
|
||||
// A stub is one byte and says nothing about what it stands
|
||||
// for. Reporting that byte count would put a 1-byte
|
||||
// `file_size` in the catalog for most of the library.
|
||||
size: if stub { 0 } else { meta.len() },
|
||||
modified: modified_secs(&meta),
|
||||
// No renderer behind a folder; previews are extracted
|
||||
// locally from the file itself.
|
||||
has_preview: false,
|
||||
materialised: !stub,
|
||||
path,
|
||||
});
|
||||
}
|
||||
@@ -408,7 +549,14 @@ impl RemoteBackend for FolderBackend {
|
||||
}
|
||||
|
||||
async fn get(&self, id: &RemoteId, range: Option<Range<u64>>) -> Result<Vec<u8>, RemoteError> {
|
||||
let local = self.resolve_id(id)?;
|
||||
let (local, materialised) = self.locate_id(id)?;
|
||||
if !materialised {
|
||||
// The one byte in the stub is not the file. Returning it produced
|
||||
// a sidecar that parsed as empty and a thumbnail that never
|
||||
// decoded; reporting `NotFound` made the sidecar writer treat an
|
||||
// existing document as absent and overwrite it.
|
||||
return Err(RemoteError::NotMaterialised(local.display().to_string()));
|
||||
}
|
||||
blocking(move || {
|
||||
let what = local.display().to_string();
|
||||
let mut file = std::fs::File::open(&local).map_err(|e| map_io(e, &what))?;
|
||||
@@ -440,7 +588,15 @@ impl RemoteBackend for FolderBackend {
|
||||
body: Vec<u8>,
|
||||
precond: Option<Precondition>,
|
||||
) -> Result<Validator, RemoteError> {
|
||||
let local = self.resolve(path)?;
|
||||
let (local, materialised) = self.locate(path)?;
|
||||
if !materialised {
|
||||
// Writing `a.drsc` while `a.drsc.nextcloud` sits beside it creates
|
||||
// two files for one document and hands the sync client a conflict
|
||||
// to resolve — in favour of whichever it sees last. The caller
|
||||
// must materialise it and merge, which is what the typed error is
|
||||
// for.
|
||||
return Err(RemoteError::NotMaterialised(local.display().to_string()));
|
||||
}
|
||||
blocking(move || {
|
||||
let what = local.display().to_string();
|
||||
if let Some(parent) = local.parent() {
|
||||
@@ -522,7 +678,10 @@ impl RemoteBackend for FolderBackend {
|
||||
id: &RemoteId,
|
||||
precond: Option<Precondition>,
|
||||
) -> Result<(), RemoteError> {
|
||||
let local = self.resolve_id(id)?;
|
||||
// Deliberately by whichever name is on disk: deleting a photograph
|
||||
// means deleting it whether or not its content happens to be here, and
|
||||
// a stub left behind would be re-listed by the next scan.
|
||||
let (local, _) = self.locate_id(id)?;
|
||||
blocking(move || {
|
||||
let what = local.display().to_string();
|
||||
let meta = std::fs::symlink_metadata(&local).map_err(|e| map_io(e, &what))?;
|
||||
@@ -562,8 +721,19 @@ impl RemoteBackend for FolderBackend {
|
||||
}
|
||||
|
||||
async fn move_to(&self, from: &RemoteId, to: &RemotePath) -> Result<(), RemoteError> {
|
||||
let src = self.resolve_id(from)?;
|
||||
let dst = self.resolve(to)?;
|
||||
// Move whichever name exists. Trashing a photograph that is not
|
||||
// downloaded is a perfectly ordinary thing to do, and it must move the
|
||||
// stub — renaming a placeholder keeps it a placeholder.
|
||||
let (src, materialised) = self.locate_id(from)?;
|
||||
let dst = if materialised {
|
||||
self.resolve(to)?
|
||||
} else {
|
||||
// The destination keeps the placeholder suffix, or the client
|
||||
// would see a one-byte file appear where a photograph should be.
|
||||
self.resolve(&RemotePath::new(
|
||||
self.vfs.placeholder_name(to.as_str()).into_owned(),
|
||||
))?
|
||||
};
|
||||
blocking(move || {
|
||||
let what = dst.display().to_string();
|
||||
// Parents first: the trash folder does not exist until the first
|
||||
@@ -597,6 +767,60 @@ impl RemoteBackend for FolderBackend {
|
||||
.await
|
||||
}
|
||||
|
||||
/// TRACES: FR-NC-6c
|
||||
/// Ask the sync client to download a placeholder, and wait for it.
|
||||
///
|
||||
/// Suffix-mode VFS *renames* on hydration, so completion is the
|
||||
/// materialised path appearing — not the stub changing size. Polling the
|
||||
/// original would wait forever.
|
||||
async fn materialise(&self, id: &RemoteId) -> Result<(), RemoteError> {
|
||||
let (local, materialised) = self.locate_id(id)?;
|
||||
if materialised {
|
||||
// Already here. Not an error, and not a reason to ask again: the
|
||||
// borrow pool relies on this being idempotent.
|
||||
return Ok(());
|
||||
}
|
||||
let vfs = self.vfs.clone();
|
||||
let target = self.resolve_id(id)?;
|
||||
blocking(move || {
|
||||
vfs.materialise(&local)?;
|
||||
|
||||
// The client acknowledges the command, not the transfer, so this
|
||||
// waits for the file to appear. A bounded wait: a hydration that
|
||||
// has not landed in this long is one the caller should be told
|
||||
// about rather than blocked on for ever — the pass can come back
|
||||
// to it.
|
||||
let deadline = std::time::Instant::now() + MATERIALISE_TIMEOUT;
|
||||
while std::time::Instant::now() < deadline {
|
||||
if target.is_file() {
|
||||
return Ok(());
|
||||
}
|
||||
std::thread::sleep(POLL);
|
||||
}
|
||||
Err(RemoteError::Network(format!(
|
||||
"{} did not download within {}s",
|
||||
target.display(),
|
||||
MATERIALISE_TIMEOUT.as_secs()
|
||||
)))
|
||||
})
|
||||
.await
|
||||
}
|
||||
|
||||
/// TRACES: FR-NC-6c
|
||||
/// Hand the content back, leaving a placeholder.
|
||||
///
|
||||
/// **Never a delete.** In a synced tree removing the file propagates the
|
||||
/// removal to the server; the client is asked to dehydrate, and if it
|
||||
/// cannot the content simply stays.
|
||||
async fn dematerialise(&self, id: &RemoteId) -> Result<(), RemoteError> {
|
||||
let (local, materialised) = self.locate_id(id)?;
|
||||
if !materialised {
|
||||
return Ok(());
|
||||
}
|
||||
let vfs = self.vfs.clone();
|
||||
blocking(move || vfs.dematerialise(&local)).await
|
||||
}
|
||||
|
||||
async fn create_dir(&self, path: &RemotePath) -> Result<(), RemoteError> {
|
||||
let local = self.resolve(path)?;
|
||||
blocking(move || {
|
||||
|
||||
Reference in New Issue
Block a user