The passes that need every photograph's bytes — thumbnails, face indexing — now borrow each one and release it at the end. On a placeholder library that is the difference between peak disk being the working set and being the whole library. Including on cancellation, which was nearly missed: the face sweep returns mid-loop when the user presses Stop, and without releasing there the disk is spent and nothing is delivered for it. `materialise` now answers whether *it* fetched the content. The pool used to work that out by listing a file's parent directory — one listing per file across a library — when the backend already had to `stat` it to decide whether to ask. One syscall instead of a directory walk, and it removes the bug class the tests found earlier: a file at the library root has no `parent()`, so every one of them read as already-downloaded. **Pinning is the retention control**, and it drives the model the catalog already had rather than a second one. `tier_desired` is what the user asked to keep hydrated, `pending_pins` is the resumable work list, and a pinned collection is never dehydrated for the same reason it was never evicted. It was in fact *broken* here before: `get` on a stub failed, and the pin worker logged "one unreadable file must not abandon the whole pin" and silently did nothing. Pinned originals on such a library are recorded with `path = NULL` (`Cache::record_in_place`) rather than copied under `originals/`. Two reasons, and the second is the important one. A copy would hold every pinned photograph twice, with the budget able to evict the half that was not costing the disk. And `release` deletes the file a row names — so a row that names none cannot delete anything, which puts the one catastrophic operation out of reach by construction rather than by remembering not to call it. Deleting a materialised file inside a synced tree removes the photograph from the server and every other device. Handing disk back is `spawn_dehydrate`, which asks the client. Two gaps written down rather than papered over (docs/storage.md §7): a hydrating pass cannot yet quote its cost, because a stub reports no size; and the two sweeps hold separate pools, so a library indexed for both fetches twice.
294 lines
12 KiB
Rust
294 lines
12 KiB
Rust
//! Pluggable remote storage for DarkRoom.
|
|
//!
|
|
//! Defines the [`RemoteBackend`] trait and the capability model the sync
|
|
//! engine adapts to, plus the pieces that let the application hold a backend
|
|
//! without naming one: an [`Account`] that is configuration rather than a
|
|
//! server, and a [`BackendProvider`] registry that turns one into a live
|
|
//! connection.
|
|
//!
|
|
//! Two connectors ship: `dr-sync-nextcloud` and `dr-sync-folder`. Adding a
|
|
//! third is implementing those two traits and registering the result — see
|
|
//! [`provider`] for the whole contract.
|
|
//!
|
|
//! # Why capabilities rather than a common denominator
|
|
//!
|
|
//! Nextcloud's fast path relies on a behaviour that is *not* a WebDAV
|
|
//! guarantee: directory ETags propagate up the tree, so an unchanged root
|
|
//! ETag proves nothing anywhere in the library changed. That single property
|
|
//! turns a no-op sync over 50k images into one HTTP request.
|
|
//!
|
|
//! A trait built to what every backend can do would force full enumeration
|
|
//! every time — the exact cost the design exists to avoid. So backends
|
|
//! declare what they support and the engine picks a strategy (ARCH §8.1).
|
|
|
|
use std::ops::Range;
|
|
|
|
use async_trait::async_trait;
|
|
|
|
pub mod account;
|
|
pub mod capability;
|
|
pub mod error;
|
|
pub mod provider;
|
|
pub mod reachability;
|
|
pub mod scan;
|
|
pub mod types;
|
|
pub mod upload;
|
|
|
|
pub use account::{Account, AccountError, AccountStore, Connection, Secret, LEGACY_BACKEND};
|
|
pub use capability::{
|
|
Capabilities, ChangeDetection, ChunkConstraints, Materialisation, ServerPreviews,
|
|
};
|
|
pub use error::RemoteError;
|
|
pub use provider::{BackendProvider, BackendRegistry, SignIn};
|
|
pub use reachability::{Connectivity, Reachability};
|
|
pub use scan::{scan, ScanProgress, ScanResult};
|
|
pub use types::{
|
|
Cursor, EntryKind, Identity, Precondition, RemoteChange, RemoteEntry, RemoteId, RemotePath,
|
|
Validator,
|
|
};
|
|
pub use upload::{destination, upload_original, Placed};
|
|
|
|
/// TRACES: FR-NC-12
|
|
/// A remote storage backend.
|
|
///
|
|
/// Implementations are expected to be cheap to clone or to be used behind an
|
|
/// `Arc`; the engine may call them concurrently.
|
|
#[async_trait]
|
|
pub trait RemoteBackend: Send + Sync {
|
|
/// What this backend supports. Read once at connect time and used to pick
|
|
/// a sync strategy.
|
|
fn capabilities(&self) -> &Capabilities;
|
|
|
|
/// Human-readable backend name, for logs and the UI.
|
|
fn name(&self) -> &str;
|
|
|
|
// ---- discovery --------------------------------------------------------
|
|
|
|
/// List one directory level.
|
|
///
|
|
/// `since` carries the validator the caller last saw, so backends able to
|
|
/// skip unchanged entries may do so. Backends that cannot simply ignore
|
|
/// it.
|
|
async fn list(
|
|
&self,
|
|
dir: &RemotePath,
|
|
since: Option<&Validator>,
|
|
) -> Result<Vec<RemoteEntry>, RemoteError>;
|
|
|
|
/// Fetch a directory's validator without listing its contents.
|
|
///
|
|
/// The cheap probe that makes ETag pruning work: one request against the
|
|
/// root answers "did anything change?". Backends without
|
|
/// [`ChangeDetection::PropagatingEtags`] return
|
|
/// [`RemoteError::Unsupported`].
|
|
async fn dir_validator(&self, dir: &RemotePath) -> Result<Validator, RemoteError>;
|
|
|
|
/// Ask what changed since a cursor.
|
|
///
|
|
/// Only meaningful for [`ChangeDetection::DeltaCursor`] backends; others
|
|
/// return [`RemoteError::Unsupported`].
|
|
async fn delta(&self, cursor: &Cursor) -> Result<(Vec<RemoteChange>, Cursor), RemoteError>;
|
|
|
|
// ---- transfer ---------------------------------------------------------
|
|
|
|
/// Fetch an object, optionally a byte range.
|
|
///
|
|
/// The range is a hint, not a guarantee: backends without range support
|
|
/// may return the whole object, and the caller slices. Correctness holds
|
|
/// either way; [`Capabilities::range_reads`] says whether it was cheap.
|
|
async fn get(&self, id: &RemoteId, range: Option<Range<u64>>) -> Result<Vec<u8>, RemoteError>;
|
|
|
|
/// Upload, optionally guarded by a precondition.
|
|
///
|
|
/// Backends handle chunking internally based on body size — chunked
|
|
/// upload is an implementation detail, not part of this interface, since
|
|
/// exposing it would leak one server's protocol into the abstraction.
|
|
async fn put(
|
|
&self,
|
|
path: &RemotePath,
|
|
body: Vec<u8>,
|
|
precond: Option<Precondition>,
|
|
) -> Result<Validator, RemoteError>;
|
|
|
|
/// Upload many small objects.
|
|
///
|
|
/// Defaults to sequential [`put`](Self::put) calls; backends with a bulk
|
|
/// endpoint override it. Sidecars are the motivating case — hundreds of
|
|
/// a few KB each.
|
|
async fn put_many(
|
|
&self,
|
|
items: Vec<(RemotePath, Vec<u8>)>,
|
|
) -> Result<Vec<Result<Validator, RemoteError>>, RemoteError> {
|
|
let mut out = Vec::with_capacity(items.len());
|
|
for (path, body) in items {
|
|
out.push(self.put(&path, body, None).await);
|
|
}
|
|
Ok(out)
|
|
}
|
|
|
|
/// Delete an object.
|
|
async fn delete(&self, id: &RemoteId, precond: Option<Precondition>)
|
|
-> Result<(), RemoteError>;
|
|
|
|
/// Move an object, keeping its identity.
|
|
///
|
|
/// TRACES: FR-CAT-15
|
|
/// **The stable id must survive.** This is what a soft delete uses to put a
|
|
/// photograph in the trash folder, and what a restore uses to bring it back.
|
|
/// A move implemented as copy-then-delete would allocate a *new*
|
|
/// `oc:fileid`, which orphans the thumbnail shard entry and the sidecar
|
|
/// mapping and turns a restore into a full re-download. WebDAV `MOVE` is one
|
|
/// request and preserves the id, which is why this is its own method rather
|
|
/// than something the caller composes.
|
|
///
|
|
/// Creates missing parent directories of `to`: the trash folder does not
|
|
/// exist until the first image is trashed, and requiring the caller to
|
|
/// create it separately makes the first trash of every library a two-step
|
|
/// dance with a failure mode in the middle.
|
|
async fn move_to(&self, from: &RemoteId, to: &RemotePath) -> Result<(), RemoteError>;
|
|
|
|
/// Create a directory, and any missing parents.
|
|
///
|
|
/// Succeeds if it already exists — callers use this to guarantee a
|
|
/// destination, not to claim they created it.
|
|
async fn create_dir(&self, path: &RemotePath) -> Result<(), RemoteError>;
|
|
|
|
// ---- materialisation --------------------------------------------------
|
|
|
|
/// TRACES: FR-NC-6c
|
|
/// Ask for a placeholder's content to be brought to this device.
|
|
///
|
|
/// Only meaningful where [`Capabilities::materialisation`] is
|
|
/// [`Materialisation::OnDemand`]; others return
|
|
/// [`RemoteError::Unsupported`].
|
|
///
|
|
/// **Whole-file, and slow.** There is no partial hydration: a placeholder
|
|
/// becomes one byte or all of them, so this transfers a 27 MB RAW to
|
|
/// answer a question a 256 KB range read would have answered (ARCH §9.0
|
|
/// finding 3). It is for the originals tier — develop, export, a pin the
|
|
/// user asked for — and for passes the user has been quoted a price on and
|
|
/// agreed to. **Never for filling a grid**: doing so downloads the entire
|
|
/// library to produce thumbnails.
|
|
///
|
|
/// Returns once the content is readable, and **whether this call is what
|
|
/// brought it here** — `false` meaning it was already local.
|
|
///
|
|
/// That boolean is the whole basis of borrowing. A caller releasing what
|
|
/// it fetched must not release what the user already had, and after the
|
|
/// fact the two are indistinguishable; the backend knows because it had to
|
|
/// look before deciding whether to ask. Answering it here costs the `stat`
|
|
/// the implementation performs anyway, where a caller determining it
|
|
/// separately would pay a directory listing per file.
|
|
async fn materialise(&self, _id: &RemoteId) -> Result<bool, RemoteError> {
|
|
Err(RemoteError::Unsupported(
|
|
"this backend has no placeholders to materialise",
|
|
))
|
|
}
|
|
|
|
/// Give a placeholder's content back, freeing the disk it held.
|
|
///
|
|
/// The counterpart that makes hydration a *borrow* rather than an
|
|
/// acquisition: a pass that hydrates a library to index it can return each
|
|
/// file as it finishes, so peak disk is the working set rather than the
|
|
/// library.
|
|
///
|
|
/// **Never destructive.** On a synced folder this asks the client to
|
|
/// dehydrate; it must not delete, because a deletion in a synced tree
|
|
/// propagates to the server and removes the photograph everywhere. An
|
|
/// implementation that cannot dehydrate must return
|
|
/// [`RemoteError::Unsupported`] rather than approximating it.
|
|
async fn dematerialise(&self, _id: &RemoteId) -> Result<(), RemoteError> {
|
|
Err(RemoteError::Unsupported(
|
|
"this backend has no placeholders to release",
|
|
))
|
|
}
|
|
|
|
// ---- optional ---------------------------------------------------------
|
|
|
|
/// Server-rendered thumbnail, where available.
|
|
///
|
|
/// `Ok(None)` means the server has no preview for this object — which is
|
|
/// common for RAW, since stock Nextcloud ships no RAW preview provider
|
|
/// (ARCH §6.7). Callers must have a local fallback.
|
|
async fn thumbnail(&self, _id: &RemoteId, _size: u32) -> Result<Option<Vec<u8>>, RemoteError> {
|
|
Ok(None)
|
|
}
|
|
}
|
|
|
|
/// TRACES: FR-NC-4 | FR-NC-12
|
|
/// Which sync strategy the engine should use for a backend.
|
|
///
|
|
/// Derived from capabilities at connect time. Reported to the user so a slow
|
|
/// backend is visibly slow rather than mysteriously slow (FR-NC-12).
|
|
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
|
|
pub enum SyncStrategy {
|
|
/// Ask the server what changed. Cheapest.
|
|
Delta,
|
|
/// Probe the root ETag; recurse only where it differs. One request when
|
|
/// nothing changed.
|
|
EtagPruning,
|
|
/// Enumerate the tree, using per-entry ETags to avoid re-downloading.
|
|
FullListing,
|
|
/// Enumerate and compare modification times. Degraded — clock skew and
|
|
/// second-granularity timestamps both cause misses.
|
|
TimestampCompare,
|
|
}
|
|
|
|
impl SyncStrategy {
|
|
pub fn for_capabilities(caps: &Capabilities) -> Self {
|
|
match caps.change_detection {
|
|
ChangeDetection::DeltaCursor => SyncStrategy::Delta,
|
|
ChangeDetection::PropagatingEtags => SyncStrategy::EtagPruning,
|
|
ChangeDetection::LocalEtags => SyncStrategy::FullListing,
|
|
ChangeDetection::Timestamps => SyncStrategy::TimestampCompare,
|
|
}
|
|
}
|
|
|
|
/// A short description for the UI.
|
|
pub fn describe(self) -> &'static str {
|
|
match self {
|
|
SyncStrategy::Delta => "server change feed",
|
|
SyncStrategy::EtagPruning => "incremental (ETag pruning)",
|
|
SyncStrategy::FullListing => "full listing, cached by ETag",
|
|
SyncStrategy::TimestampCompare => "full listing by timestamp (degraded)",
|
|
}
|
|
}
|
|
|
|
/// Whether this strategy can prove "nothing changed" cheaply.
|
|
///
|
|
/// Where false, every sync costs at least one request per folder, which
|
|
/// the UI should warn about on large libraries.
|
|
pub fn has_cheap_noop(self) -> bool {
|
|
matches!(self, SyncStrategy::Delta | SyncStrategy::EtagPruning)
|
|
}
|
|
}
|
|
|
|
#[cfg(test)]
|
|
mod tests {
|
|
use super::*;
|
|
|
|
#[test]
|
|
fn strategy_follows_capability() {
|
|
let mut caps = Capabilities::minimal();
|
|
assert_eq!(
|
|
SyncStrategy::for_capabilities(&caps),
|
|
SyncStrategy::TimestampCompare
|
|
);
|
|
|
|
caps.change_detection = ChangeDetection::PropagatingEtags;
|
|
assert_eq!(
|
|
SyncStrategy::for_capabilities(&caps),
|
|
SyncStrategy::EtagPruning
|
|
);
|
|
}
|
|
|
|
#[test]
|
|
fn only_delta_and_pruning_have_cheap_noop() {
|
|
assert!(SyncStrategy::Delta.has_cheap_noop());
|
|
assert!(SyncStrategy::EtagPruning.has_cheap_noop());
|
|
// These cost at least one request per folder, every time.
|
|
assert!(!SyncStrategy::FullListing.has_cheap_noop());
|
|
assert!(!SyncStrategy::TimestampCompare.has_cheap_noop());
|
|
}
|
|
}
|