Files
DarkRoom/core/dr-sync/src/lib.rs
T
dtourolle 5768100816 Borrow the library to index it, and give it back
The passes that need every photograph's bytes — thumbnails, face
indexing — now borrow each one and release it at the end. On a
placeholder library that is the difference between peak disk being the
working set and being the whole library.

Including on cancellation, which was nearly missed: the face sweep
returns mid-loop when the user presses Stop, and without releasing there
the disk is spent and nothing is delivered for it.

`materialise` now answers whether *it* fetched the content. The pool used
to work that out by listing a file's parent directory — one listing per
file across a library — when the backend already had to `stat` it to
decide whether to ask. One syscall instead of a directory walk, and it
removes the bug class the tests found earlier: a file at the library root
has no `parent()`, so every one of them read as already-downloaded.

**Pinning is the retention control**, and it drives the model the catalog
already had rather than a second one. `tier_desired` is what the user
asked to keep hydrated, `pending_pins` is the resumable work list, and a
pinned collection is never dehydrated for the same reason it was never
evicted. It was in fact *broken* here before: `get` on a stub failed, and
the pin worker logged "one unreadable file must not abandon the whole
pin" and silently did nothing.

Pinned originals on such a library are recorded with `path = NULL`
(`Cache::record_in_place`) rather than copied under `originals/`. Two
reasons, and the second is the important one. A copy would hold every
pinned photograph twice, with the budget able to evict the half that was
not costing the disk. And `release` deletes the file a row names — so a
row that names none cannot delete anything, which puts the one
catastrophic operation out of reach by construction rather than by
remembering not to call it. Deleting a materialised file inside a synced
tree removes the photograph from the server and every other device.

Handing disk back is `spawn_dehydrate`, which asks the client.

Two gaps written down rather than papered over (docs/storage.md §7): a
hydrating pass cannot yet quote its cost, because a stub reports no size;
and the two sweeps hold separate pools, so a library indexed for both
fetches twice.
2026-08-29 09:57:53 +02:00

294 lines
12 KiB
Rust

//! Pluggable remote storage for DarkRoom.
//!
//! Defines the [`RemoteBackend`] trait and the capability model the sync
//! engine adapts to, plus the pieces that let the application hold a backend
//! without naming one: an [`Account`] that is configuration rather than a
//! server, and a [`BackendProvider`] registry that turns one into a live
//! connection.
//!
//! Two connectors ship: `dr-sync-nextcloud` and `dr-sync-folder`. Adding a
//! third is implementing those two traits and registering the result — see
//! [`provider`] for the whole contract.
//!
//! # Why capabilities rather than a common denominator
//!
//! Nextcloud's fast path relies on a behaviour that is *not* a WebDAV
//! guarantee: directory ETags propagate up the tree, so an unchanged root
//! ETag proves nothing anywhere in the library changed. That single property
//! turns a no-op sync over 50k images into one HTTP request.
//!
//! A trait built to what every backend can do would force full enumeration
//! every time — the exact cost the design exists to avoid. So backends
//! declare what they support and the engine picks a strategy (ARCH §8.1).
use std::ops::Range;
use async_trait::async_trait;
pub mod account;
pub mod capability;
pub mod error;
pub mod provider;
pub mod reachability;
pub mod scan;
pub mod types;
pub mod upload;
pub use account::{Account, AccountError, AccountStore, Connection, Secret, LEGACY_BACKEND};
pub use capability::{
Capabilities, ChangeDetection, ChunkConstraints, Materialisation, ServerPreviews,
};
pub use error::RemoteError;
pub use provider::{BackendProvider, BackendRegistry, SignIn};
pub use reachability::{Connectivity, Reachability};
pub use scan::{scan, ScanProgress, ScanResult};
pub use types::{
Cursor, EntryKind, Identity, Precondition, RemoteChange, RemoteEntry, RemoteId, RemotePath,
Validator,
};
pub use upload::{destination, upload_original, Placed};
/// TRACES: FR-NC-12
/// A remote storage backend.
///
/// Implementations are expected to be cheap to clone or to be used behind an
/// `Arc`; the engine may call them concurrently.
#[async_trait]
pub trait RemoteBackend: Send + Sync {
/// What this backend supports. Read once at connect time and used to pick
/// a sync strategy.
fn capabilities(&self) -> &Capabilities;
/// Human-readable backend name, for logs and the UI.
fn name(&self) -> &str;
// ---- discovery --------------------------------------------------------
/// List one directory level.
///
/// `since` carries the validator the caller last saw, so backends able to
/// skip unchanged entries may do so. Backends that cannot simply ignore
/// it.
async fn list(
&self,
dir: &RemotePath,
since: Option<&Validator>,
) -> Result<Vec<RemoteEntry>, RemoteError>;
/// Fetch a directory's validator without listing its contents.
///
/// The cheap probe that makes ETag pruning work: one request against the
/// root answers "did anything change?". Backends without
/// [`ChangeDetection::PropagatingEtags`] return
/// [`RemoteError::Unsupported`].
async fn dir_validator(&self, dir: &RemotePath) -> Result<Validator, RemoteError>;
/// Ask what changed since a cursor.
///
/// Only meaningful for [`ChangeDetection::DeltaCursor`] backends; others
/// return [`RemoteError::Unsupported`].
async fn delta(&self, cursor: &Cursor) -> Result<(Vec<RemoteChange>, Cursor), RemoteError>;
// ---- transfer ---------------------------------------------------------
/// Fetch an object, optionally a byte range.
///
/// The range is a hint, not a guarantee: backends without range support
/// may return the whole object, and the caller slices. Correctness holds
/// either way; [`Capabilities::range_reads`] says whether it was cheap.
async fn get(&self, id: &RemoteId, range: Option<Range<u64>>) -> Result<Vec<u8>, RemoteError>;
/// Upload, optionally guarded by a precondition.
///
/// Backends handle chunking internally based on body size — chunked
/// upload is an implementation detail, not part of this interface, since
/// exposing it would leak one server's protocol into the abstraction.
async fn put(
&self,
path: &RemotePath,
body: Vec<u8>,
precond: Option<Precondition>,
) -> Result<Validator, RemoteError>;
/// Upload many small objects.
///
/// Defaults to sequential [`put`](Self::put) calls; backends with a bulk
/// endpoint override it. Sidecars are the motivating case — hundreds of
/// a few KB each.
async fn put_many(
&self,
items: Vec<(RemotePath, Vec<u8>)>,
) -> Result<Vec<Result<Validator, RemoteError>>, RemoteError> {
let mut out = Vec::with_capacity(items.len());
for (path, body) in items {
out.push(self.put(&path, body, None).await);
}
Ok(out)
}
/// Delete an object.
async fn delete(&self, id: &RemoteId, precond: Option<Precondition>)
-> Result<(), RemoteError>;
/// Move an object, keeping its identity.
///
/// TRACES: FR-CAT-15
/// **The stable id must survive.** This is what a soft delete uses to put a
/// photograph in the trash folder, and what a restore uses to bring it back.
/// A move implemented as copy-then-delete would allocate a *new*
/// `oc:fileid`, which orphans the thumbnail shard entry and the sidecar
/// mapping and turns a restore into a full re-download. WebDAV `MOVE` is one
/// request and preserves the id, which is why this is its own method rather
/// than something the caller composes.
///
/// Creates missing parent directories of `to`: the trash folder does not
/// exist until the first image is trashed, and requiring the caller to
/// create it separately makes the first trash of every library a two-step
/// dance with a failure mode in the middle.
async fn move_to(&self, from: &RemoteId, to: &RemotePath) -> Result<(), RemoteError>;
/// Create a directory, and any missing parents.
///
/// Succeeds if it already exists — callers use this to guarantee a
/// destination, not to claim they created it.
async fn create_dir(&self, path: &RemotePath) -> Result<(), RemoteError>;
// ---- materialisation --------------------------------------------------
/// TRACES: FR-NC-6c
/// Ask for a placeholder's content to be brought to this device.
///
/// Only meaningful where [`Capabilities::materialisation`] is
/// [`Materialisation::OnDemand`]; others return
/// [`RemoteError::Unsupported`].
///
/// **Whole-file, and slow.** There is no partial hydration: a placeholder
/// becomes one byte or all of them, so this transfers a 27 MB RAW to
/// answer a question a 256 KB range read would have answered (ARCH §9.0
/// finding 3). It is for the originals tier — develop, export, a pin the
/// user asked for — and for passes the user has been quoted a price on and
/// agreed to. **Never for filling a grid**: doing so downloads the entire
/// library to produce thumbnails.
///
/// Returns once the content is readable, and **whether this call is what
/// brought it here** — `false` meaning it was already local.
///
/// That boolean is the whole basis of borrowing. A caller releasing what
/// it fetched must not release what the user already had, and after the
/// fact the two are indistinguishable; the backend knows because it had to
/// look before deciding whether to ask. Answering it here costs the `stat`
/// the implementation performs anyway, where a caller determining it
/// separately would pay a directory listing per file.
async fn materialise(&self, _id: &RemoteId) -> Result<bool, RemoteError> {
Err(RemoteError::Unsupported(
"this backend has no placeholders to materialise",
))
}
/// Give a placeholder's content back, freeing the disk it held.
///
/// The counterpart that makes hydration a *borrow* rather than an
/// acquisition: a pass that hydrates a library to index it can return each
/// file as it finishes, so peak disk is the working set rather than the
/// library.
///
/// **Never destructive.** On a synced folder this asks the client to
/// dehydrate; it must not delete, because a deletion in a synced tree
/// propagates to the server and removes the photograph everywhere. An
/// implementation that cannot dehydrate must return
/// [`RemoteError::Unsupported`] rather than approximating it.
async fn dematerialise(&self, _id: &RemoteId) -> Result<(), RemoteError> {
Err(RemoteError::Unsupported(
"this backend has no placeholders to release",
))
}
// ---- optional ---------------------------------------------------------
/// Server-rendered thumbnail, where available.
///
/// `Ok(None)` means the server has no preview for this object — which is
/// common for RAW, since stock Nextcloud ships no RAW preview provider
/// (ARCH §6.7). Callers must have a local fallback.
async fn thumbnail(&self, _id: &RemoteId, _size: u32) -> Result<Option<Vec<u8>>, RemoteError> {
Ok(None)
}
}
/// TRACES: FR-NC-4 | FR-NC-12
/// Which sync strategy the engine should use for a backend.
///
/// Derived from capabilities at connect time. Reported to the user so a slow
/// backend is visibly slow rather than mysteriously slow (FR-NC-12).
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum SyncStrategy {
/// Ask the server what changed. Cheapest.
Delta,
/// Probe the root ETag; recurse only where it differs. One request when
/// nothing changed.
EtagPruning,
/// Enumerate the tree, using per-entry ETags to avoid re-downloading.
FullListing,
/// Enumerate and compare modification times. Degraded — clock skew and
/// second-granularity timestamps both cause misses.
TimestampCompare,
}
impl SyncStrategy {
pub fn for_capabilities(caps: &Capabilities) -> Self {
match caps.change_detection {
ChangeDetection::DeltaCursor => SyncStrategy::Delta,
ChangeDetection::PropagatingEtags => SyncStrategy::EtagPruning,
ChangeDetection::LocalEtags => SyncStrategy::FullListing,
ChangeDetection::Timestamps => SyncStrategy::TimestampCompare,
}
}
/// A short description for the UI.
pub fn describe(self) -> &'static str {
match self {
SyncStrategy::Delta => "server change feed",
SyncStrategy::EtagPruning => "incremental (ETag pruning)",
SyncStrategy::FullListing => "full listing, cached by ETag",
SyncStrategy::TimestampCompare => "full listing by timestamp (degraded)",
}
}
/// Whether this strategy can prove "nothing changed" cheaply.
///
/// Where false, every sync costs at least one request per folder, which
/// the UI should warn about on large libraries.
pub fn has_cheap_noop(self) -> bool {
matches!(self, SyncStrategy::Delta | SyncStrategy::EtagPruning)
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn strategy_follows_capability() {
let mut caps = Capabilities::minimal();
assert_eq!(
SyncStrategy::for_capabilities(&caps),
SyncStrategy::TimestampCompare
);
caps.change_detection = ChangeDetection::PropagatingEtags;
assert_eq!(
SyncStrategy::for_capabilities(&caps),
SyncStrategy::EtagPruning
);
}
#[test]
fn only_delta_and_pruning_have_cheap_noop() {
assert!(SyncStrategy::Delta.has_cheap_noop());
assert!(SyncStrategy::EtagPruning.has_cheap_noop());
// These cost at least one request per folder, every time.
assert!(!SyncStrategy::FullListing.has_cheap_noop());
assert!(!SyncStrategy::TimestampCompare.has_cheap_noop());
}
}