The folder connector was pointed at a Nextcloud VFS tree and got three things wrong, the first of which loses work. **A dehydrated sidecar read as absent.** `a.drsc` does not exist when the client has dehydrated it — only `a.drsc.nextcloud` does — so `get` missed, `.ok()` swallowed the `NotFound`, and the sidecar writer took that for "there is no sidecar yet" and wrote a fresh document over the existing one. Every edit another device had put there went with it. That function's own doc comment calls this the exact loss the format's unknown-key preservation exists to prevent. **A stub was catalogued as a 1-byte image**, and ARCH §9.0 measured this machine at 121,785 placeholders against 10,267 real files — so a folder library on a synced tree was ~92% broken rows. **Identity changed on hydration**, so downloading a photograph looked like a delete and an add, orphaning its thumbnail and its face rows. Entries now carry the photograph's own name and a `materialised` flag; `get` on a stub returns the new `RemoteError::NotMaterialised`, which is distinct from `NotFound` precisely because the sidecar writer must treat them differently — it fetches the sidecar and merges, or leaves the entry queued. Hydration is a **borrow**. `BorrowPool` records what was on disk before it asked, so `release_all` dehydrates only what a pass brought and leaves what the user already had. Reference counted: the thumbnail pass and the face pass meet on the same RAW, and without counting the first to finish dehydrates the file the second is reading. A borrow against a plain folder or a server does nothing, so a pass written for VFS runs everywhere. Releasing means asking the client to dehydrate and never deleting: a deletion inside a synced tree propagates to the server and removes the photograph from every device. Not a second backend — the capability is per *connection*, not per type, since the same folder hydrates only while the client runs. The convention arrives through a detector the registry supplies, so `dr-sync-folder` still knows nothing about any client's protocol. ARCH §9.0a records this as an amendment: finding 3 rejected hydration because it costs 100× a range read, and that comparison assumed a connector was available. A folder library has none.
24 KiB
Storage backends
How DarkRoom talks to wherever a library lives, and what it takes to add somewhere new.
This document is the contract. docs/architecture.md §8 says why sync is built
on capability negotiation rather than a common denominator; this says what the
seam actually is, where each piece lives, and what a third connector has to do.
1. What "pluggable" has to mean
A trait alone does not make storage pluggable. RemoteBackend existed from the
first release and every layer above it still knew it was talking to Nextcloud:
seven files in dr-ui constructed a NextcloudBackend directly, ten functions
took one by concrete type, the account model was a server URL beside a DAV user
id, and the local cache directory was named after a hostname. The abstraction
was real and bought nothing, because everything that reached a backend was
still shaped like one product.
Pluggable means all four of these, not just the first:
- Operations — what a backend can do.
RemoteBackend. - Capabilities — what it can do cheaply, so the engine adapts instead of
assuming.
Capabilities. - Configuration — what an account is, with no server in it.
Account. - Registration — how the application discovers a connector at all, without
naming it.
BackendProvider+BackendRegistry.
Two connectors ship. Nextcloud is unchanged and keeps every one of its peculiarities — those are the point of the capability model, not an embarrassment it has to hide. The folder connector serves a plain directory and exists partly because it is genuinely useful and partly because a second implementation is the only way to find out whether the first was an abstraction.
2. Where each piece lives
core/dr-sync/ the contract, and nothing that speaks a protocol
├─ types.rs RemotePath, RemoteId, RemoteEntry, Validator, …
├─ capability.rs Capabilities, ChangeDetection, ServerPreviews
├─ error.rs RemoteError — the one error every caller handles
├─ account.rs Account, AccountStore, Secret, Connection
├─ provider.rs BackendProvider, BackendRegistry, SignIn
├─ lib.rs RemoteBackend, SyncStrategy
├─ scan.rs the walk, driven by capabilities
├─ upload.rs where an original is placed
└─ reachability.rs online/offline, inferred from observed results
core/dr-sync-nextcloud/ WebDAV, oc:fileid, chunked upload v2, Login Flow v2
core/dr-sync-folder/ a directory on a filesystem
ui/dr-ui/src/remote.rs the registry — the ONLY file above dr-sync that
names a connector
dr-sync depends on no connector. That is deliberate and load-bearing: a build
that only wants a folder library must not compile a TLS stack to get one, and
the registry therefore lives in the crate that already depends on everything —
the interface.
3. The four traits and types a connector meets
3.1 RemoteBackend — operations
#[async_trait]
pub trait RemoteBackend: Send + Sync {
fn capabilities(&self) -> &Capabilities;
fn name(&self) -> &str;
// discovery
async fn list(&self, dir: &RemotePath, since: Option<&Validator>)
-> Result<Vec<RemoteEntry>, RemoteError>;
async fn dir_validator(&self, dir: &RemotePath) -> Result<Validator, RemoteError>;
async fn delta(&self, cursor: &Cursor)
-> Result<(Vec<RemoteChange>, Cursor), RemoteError>;
// transfer
async fn get(&self, id: &RemoteId, range: Option<Range<u64>>)
-> Result<Vec<u8>, RemoteError>;
async fn put(&self, path: &RemotePath, body: Vec<u8>, precond: Option<Precondition>)
-> Result<Validator, RemoteError>;
async fn put_many(&self, items: Vec<(RemotePath, Vec<u8>)>) // defaulted
-> Result<Vec<Result<Validator, RemoteError>>, RemoteError>;
async fn delete(&self, id: &RemoteId, precond: Option<Precondition>)
-> Result<(), RemoteError>;
async fn move_to(&self, from: &RemoteId, to: &RemotePath) -> Result<(), RemoteError>;
async fn create_dir(&self, path: &RemotePath) -> Result<(), RemoteError>;
// optional
async fn thumbnail(&self, id: &RemoteId, size: u32) // defaulted to None
-> Result<Option<Vec<u8>>, RemoteError>;
}
Rules that are not obvious from the signatures:
dir_validatoranddeltaare capability-gated. ReturnRemoteError::Unsupportedunless yourChangeDetectionisPropagatingEtagsorDeltaCursorrespectively. Answeringdir_validatorwith something that does not actually propagate is worse than refusing: it lets a caller prune a subtree whose contents changed, and hides those changes for as long as the folder list holds still.gettakes an optional range, and it is a hint. A backend without cheap ranges may return the whole object; the caller slices. Correctness holds either way andCapabilities::range_readssays whether it was cheap.- Chunked upload is not in the trait. It is an implementation detail of
put, chosen by body size. Exposing it would leak one server's protocol. move_tomust preserve identity where the backend has stable ids. This is what a soft delete uses (FR-CAT-15): a move implemented as copy + delete allocates a new id, orphaning the thumbnail shard and turning a restore into a full re-download.create_dirmakes parents and succeeds if the directory exists. Callers use it to guarantee a destination, not to claim they created one.
3.2 Capabilities — what is cheap
The engine reads these once at connect time and picks a SyncStrategy. See
ARCH §8.1–8.2 for the tiers. The two that change behaviour rather than speed:
| Absent | Consequence the engine handles |
|---|---|
range_reads |
Embedded-preview extraction is impossible; browsing falls back to server previews or full download, and is refused on a metered connection |
conditional_write |
Sidecar conflict detection falls back to revision counters inside the sidecar — narrows the race, does not close it. Reported as a reduced-safety mode |
Declare what is true, not what is flattering. A backend claiming
PropagatingEtags it does not have does not merely run slowly; it silently
hides changes.
3.3 Account — configuration with no server in it
pub struct Account {
pub backend: String, // BackendProvider::id; defaults to "nextcloud" on load
pub endpoint: String, // stored as "server" — a URL, a path, a bucket
pub login: String, // empty where the connector has no notion of a user
pub user_id: String, // connector-defined sub-address; Nextcloud's DAV segment
pub root: String, // the folder chosen as the library root
pub formats: Vec<String>,
pub last_scan: Option<i64>,
}
Everything but backend is the connector's to interpret. Code above dr-sync
reads these for display and for cache keys, never for meaning.
Two properties are load-bearing:
- The on-disk form is backwards compatible.
backenddefaults to"nextcloud"andendpointis stored under its historical keyserver, so every account written before there was a choice loads unchanged. A config the app refuses to parse is an account the user has to set up again. Account::namespace()is frozen for Nextcloud. It names the directory holding the catalog, the thumbnail shards, the sidecar spool and the export outbox. Changing it does not lose that data, it abandons it — silently, as an upgrade — and costs a full rescan on top. The Nextcloud form is reproduced byte for byte from whatcatalog_pathcomputed before; every other backend is prefixed by its connector id, and long endpoints are truncated with a hash tail so two deep paths cannot collide inside one filesystem's 255-byte component limit.
3.4 Connection and Secret — the credential split
pub struct Connection { pub account: Account, pub secret: Option<Secret> }
Credentials go to platform secure storage (FR-NC-2, NFR-SEC-2). Never the
catalog, never the config file, never a log line. AccountStore writes the
account as plain JSON and the secret to the keyring, which is what lets the app
show "signed in as duncan, watching /PhotosRaw" before it has touched the
keyring at all.
Secret's inner string is reachable only through expose(), and its Debug
prints Secret(***). That closes the indirect leak — a {:?} on any struct
that happens to hold a connection — by construction rather than by review.
Connection is also what replaced a pair of arguments (credentials, user id)
threaded together through fifteen signatures in an order that could be swapped.
3.5 BackendProvider — registration
pub trait BackendProvider: Send + Sync {
fn id(&self) -> &'static str; // written to Account::backend
fn display_name(&self) -> &'static str;
fn endpoint_label(&self) -> &'static str; // "Server" / "Folder"
fn endpoint_placeholder(&self) -> &'static str;
fn sign_in(&self) -> SignIn;
fn normalise_endpoint(&self, input: &str) -> Result<String, String>;
fn account_for(&self, endpoint: &str) -> Result<Account, RemoteError>; // defaulted
fn connect(&self, conn: &Connection) -> Result<Box<dyn RemoteBackend>, RemoteError>;
}
pub enum SignIn {
/// A handshake the user completes outside the app, yielding a credential.
Browser,
/// The endpoint is the whole account. No credential, no waiting state.
EndpointOnly,
}
idis on-disk configuration. Changing it after anyone has an account orphans that account. Pick it once.normalise_endpointis where a bad endpoint is rejected, before an account is written for a library that does not exist. Its error string is shown to the user, so it says what to fix rather than naming a type. The Nextcloud provider upgradeshttp://tohttps://here (NFR-SEC-3); the folder provider canonicalises the path, so two spellings of one directory do not become two accounts indexing the same photographs.connectis synchronous and cheap. It validates configuration and builds a client; it does not talk to the remote. Workers call it per task.SignInis a shape, not a method. It would be tidier to exposeasync fn sign_in(), and wrong: Login Flow v2 is a browser handshake the user completes elsewhere while the app polls, so it is not one call, it does not finish on our schedule, and the screen has to render a URL and a waiting state in the middle of it.SignIntells the launch screen which of the two shapes to draw; the flow stays where its protocol is.
Credentials are deliberately not abstracted. An app password, an OAuth
token and a bucket key pair have no useful common shape, and inventing one
before a third backend exists would produce a wrong answer confidently. The
general form is Connection — an account plus an opaque secret — and each
connector translates that into what its protocol needs
(NextcloudProvider::credentials).
4. Adding a backend
- Implement
RemoteBackendover your protocol, in a newcore/dr-sync-<name>crate depending ondr-syncand nothing else of ours. - Declare
Capabilitieshonestly. Start fromCapabilities::minimal()and raise only what you can actually deliver. - Implement
BackendProviderbeside it. - Register it in
ui/dr-ui/src/remote.rs::registry()and add the crate toui/dr-ui/Cargo.toml.
That is the whole list. Nothing else in dr-ui changes, because nothing else in
dr-ui names a connector.
Two things to get right, because they are silent when wrong:
- Identity.
RemoteEntry::idshould beRemoteId::Stable(u64)wherever you can produce au64that names the same photograph on every device looking at the same library. The catalog keys the thumbnail shards and the face index on it (catalog.md§10.1), and an entry without one gets neither. SetCapabilities::stable_idsonly if that id also survives a rename — the two are different questions and only the second is a capability. - Path safety. A
RemotePathis built from names on the remote and from a catalog another device wrote. If you resolve one against a real filesystem, reject..before you open anything.
Register a test double the same way — BackendRegistry::register replaces an
existing id rather than shadowing it — so an integration test can stand a fake
server behind "nextcloud" without the registry knowing it happened.
5. The connectors that ship
5.1 Nextcloud (dr-sync-nextcloud, id "nextcloud")
Unchanged by the abstraction, peculiarities intact — see ARCH §8.4 for the full mapping. What matters here is that none of them had to be given up to make room for a second backend:
change_detection |
PropagatingEtags — the one-request no-op sync |
stable_ids |
yes, oc:fileid, survives server-side rename and move |
range_reads |
yes, detected by 206 vs 200, never HEAD |
chunked_upload |
v2, 5 MB – 5 GB, MKCOL → PUT chunks → MOVE .file |
bulk_upload |
yes, POST /remote.php/dav/bulk |
conditional_write |
yes, If-Match |
server_previews |
CommonFormatsOnly — stock Nextcloud renders no RAW |
| sign-in | SignIn::Browser, Login Flow v2, system browser, app password |
Also kept: the oc:permissions probe on a refused PUT, which is what
distinguishes a create-only share from a bad credential; the 423 Locked
retry classification; and the bundled ISRG Root YE certificate.
5.2 Folder (dr-sync-folder, id "folder")
A local disk, an NFS or SMB mount, an external drive, or the directory a Nextcloud desktop client already syncs. No server, no account, no credential — which makes it the route that works on a machine with no secrets daemon at all.
change_detection |
LocalEtags — see below |
stable_ids |
no — the id is a path hash and does not survive a rename |
range_reads |
yes, seek + take |
chunked_upload |
none; a write is a write |
bulk_upload |
no |
conditional_write |
yes, with a documented residual race |
server_previews |
None |
| sign-in | SignIn::EndpointOnly |
Why LocalEtags and not PropagatingEtags. A POSIX directory's mtime
changes when its own entry list changes and at no other time — not when a
child's contents are edited, and not for a grandchild. There is nothing to
propagate, so dir_validator returns Unsupported and the engine walks the
tree every scan. Which costs almost nothing, because the walk that was expensive
was expensive for a reason this backend does not have.
Measured 2026-08-28, cargo run -p dr-sync-folder --example scan: a full
uncached walk of 2,299 images across 233 directories completed in 137 ms,
and 380 images across 13 directories in 29 ms — the same engine, the same
Depth: 1-per-directory walk, with no pruning at all. The Nextcloud connector's
comparable figure is 34.1 s for 17,185 RAWs across 334 directories with
pruning available (ARCH §8.4). The capability model is what lets one engine
drive both at the speed each actually runs at, instead of forcing the fast one
down to the slow one's interface.
Identity is a hash of the path relative to the library root, FNV-1a 64
(written out, because DefaultHasher is explicitly unstable between Rust
releases and this value is written into the catalog). It gives the catalog a
u64 that names a photograph, is the same on every device looking at the same
folder, and does not change when the file is edited. It does not survive a
rename, and stable_ids: false says so: a moved photograph is seen as a delete
and an add, and its thumbnail is derived again.
The alternative — keying on the inode — is stable across a rename but differs between devices and is reused by the filesystem after a delete. Two machines would disagree about which photograph a thumbnail belonged to, and a recycled inode would silently attach an old thumbnail to a new image. Re-deriving a thumbnail is a cost; showing the wrong one is a bug.
Conditional writes. IfAbsent is genuinely atomic (O_CREAT | O_EXCL).
IfMatch is compare-then-swap: a stat, then a write to a temporary beside the
destination and a rename over it. A POSIX filesystem has no compare-and-swap,
so the race is narrowed to the microseconds between the two syscalls rather than
closed — still far tighter than the fallback the engine uses for a backend that
declares no conditional write at all, which spans a whole read-modify-write.
The capability is declared, and the residual race is documented at the call
site.
Two deliberate divergences from WebDAV semantics:
deleteis not recursive. A folder library is the user's own photographs on their own disk with no server-side trash behind it, so a caller that passed the wrong path would have no way back. Deleting a non-empty directory returnsRemoteError::Configuration. Nothing in the engine deletes a directory — the soft delete is amove_tointo the trash folder — so the guard is free.- Every filesystem call runs on the blocking pool. On a local disk that is overkill; on the NFS mount this backend is most useful over, a stalled server would otherwise wedge the async worker that made the call and every other request sharing it.
Failure classification matters as much as the operations. A vanished mount
(ESTALE, ENOTCONN, EIO) maps to RemoteError::Network, which is what puts
the app into offline mode and leaves the catalog readable — exactly as a dead
server does. A permissions problem maps to PermissionDenied and does not,
because going offline over one forbidden file would hide a fixable problem
behind a network banner. An endpoint that is not a directory at all maps to
RemoteError::Configuration: nothing was unreachable and no credential was
wrong, so neither of the other two would send the user anywhere useful.
6. Virtual filesystems
A sync client in virtual-files mode leaves a placeholder where a file is
catalogued but not downloaded. On Linux — the only mode it supports — that
means IMG.CR2 does not exist at all and IMG.CR2.nextcloud does, holding one
byte. ARCH §9.0 measured a real machine: 121,785 placeholders against 10,267
materialised files.
A folder library that ignores this is not merely degraded, it is dangerous. Before the handling below existed, the folder connector catalogued every stub as a 1-byte image, gave it an identity that changed the moment it was downloaded, and — worst — reported a dehydrated sidecar as absent, which made the sidecar writer create a fresh document over an existing one and discard every edit another device had put there.
6.1 Three questions, one trait
Everything else about a synced folder is an ordinary directory, so this is not
a second connector. dr_sync_folder::Vfs asks only what differs:
pub trait Vfs: Send + Sync {
fn name(&self) -> &'static str;
fn is_placeholder(&self, on_disk: &str) -> bool;
fn real_name<'a>(&self, on_disk: &'a str) -> &'a str;
fn placeholder_name(&self, name: &str) -> Cow<'_, str>;
fn can_materialise(&self) -> bool; // defaulted false
fn materialise(&self, local: &Path) -> Result<(), RemoteError>; // defaulted
fn dematerialise(&self, local: &Path) -> Result<(), RemoteError>; // defaulted
}
NoVfs for a plain directory; dr_sync_nextcloud::NextcloudVfs for a synced
one, wrapping the DesktopClient socket. A third convention is a third impl.
Why not a folder-vfs provider. The interesting capability is not a
property of the backend: the same directory can materialise on demand while the
client is running and cannot when it is down, so it must be computed per
connection either way. Registering two providers would ask the user to choose
between two things that differ by whether a background process is up. The
convention is detected instead, per connection, by a hook the registry supplies
(FolderProvider::with_vfs_detector) — which is what keeps dr-sync-folder
free of any client's protocol.
6.2 What the backend reports
RemoteEntry::path |
the photograph's name, never the stub's — so identity survives a download |
RemoteEntry::materialised |
false on a stub; the catalog maps it to Availability::Offline |
RemoteEntry::size |
0 on a stub, meaning unknown — see below |
get on a stub |
RemoteError::NotMaterialised, never NotFound and never the stub's one byte |
put over a stub |
refused; writing a rival file hands the client a conflict to resolve arbitrarily |
move_to a stub |
moves the stub and keeps it a stub — culling without downloading is ordinary |
delete a stub |
deletes it; a photograph is deleted whether or not its bytes are here |
capabilities().materialisation |
OnDemand with a client, Placeholders without, Always on a plain folder |
Size is genuinely unknown. A Linux suffix-mode stub is one byte and carries
no record of what it stands for. The client's ._sync_*.db has the real size,
but that is a private schema and reading it would couple us to their migrations.
FR-NC-6c wants a transfer size quoted before an operation starts; for a stub the
honest answer is that it cannot be, and the interface should say so rather than
report one byte or invent an estimate silently.
6.3 Hydration is a borrow
The rule: a file is returned to the state it was found in. What a pass
downloaded is released; what the user already had is left alone. BorrowPool
enforces it.
let pool = BorrowPool::new();
{
let held = pool.borrow(&backend, &path).await?; // downloads only if absent
// ... read it, thumbnail it, index its faces ...
} // borrow ends
let stats = pool.release_all(&backend).await; // dehydrates only what it hydrated
Three properties that are not obvious:
- Reference counted. The thumbnail pass and the face pass meet on the same RAW. Without counting, the first to finish dehydrates the file the second is reading; with it, the transfer is paid once and released when the last borrower is done.
- Prior state is read before asking. After
materialisethere is no way to tell what the pass brought from what was already there, so it is recorded first. Getting this wrong silently undoes a pin, and "my pinned trip evaporated after an indexing run" is the failure that would make people stop trusting the feature. - Being unsure is not symmetric.
borrow_known(.., Some(true))keeps a file that might have been ours — costing disk.Some(false)releases one that might have been the user's. An uncertain caller passestrueorNone, never a guess atfalse.
A borrow against a plain folder or a server backend short-circuits and does nothing, so a pass written for a VFS library runs unchanged everywhere rather than growing two code paths.
6.4 Release means dehydrate, never delete
The single most dangerous thing in this feature. A synced folder is not a
cache: deleting a materialised file inside it propagates the deletion to the
server and removes the photograph from every device the user owns. Vfs and
RemoteBackend::dematerialise both say so, and an implementation that cannot
dehydrate returns Unsupported rather than approximating it.
This is also why the originals cache (dr_catalog::cache) cannot simply be
pointed at a VFS library: Cache::release deletes bytes, which is right for a
copy under originals/ and catastrophic in place.
7. What the abstraction does not yet cover
Stated so the next person does not have to rediscover it.
- Multiple accounts at once.
AccountStoreholds a list and the launch screen uses the most recent. Nothing in the model prevents two open libraries; the interface has no place to show them. - Per-backend settings. A connector has no way to contribute a settings
page. Anything configurable is on the
Accountor is not configurable. - Capability probing at runtime.
Capabilitiesis fixed at construction. Nextcloud'sserver_previewsshould really be probed per account — a server withcamerarawpreviewsinstalled can render RAW — and today it is assumed to beCommonFormatsOnly. - A general notion of an account. Credentials stay connector-specific on
purpose (§3.5). A third connector with an OAuth flow will need a third
SignInvariant, and that is the right place for it to appear.