Treat a placeholder as the photograph, not as a one-byte file

The folder connector was pointed at a Nextcloud VFS tree and got three
things wrong, the first of which loses work.

**A dehydrated sidecar read as absent.** `a.drsc` does not exist when the
client has dehydrated it — only `a.drsc.nextcloud` does — so `get` missed,
`.ok()` swallowed the `NotFound`, and the sidecar writer took that for
"there is no sidecar yet" and wrote a fresh document over the existing
one. Every edit another device had put there went with it. That function's
own doc comment calls this the exact loss the format's unknown-key
preservation exists to prevent.

**A stub was catalogued as a 1-byte image**, and ARCH §9.0 measured this
machine at 121,785 placeholders against 10,267 real files — so a folder
library on a synced tree was ~92% broken rows.

**Identity changed on hydration**, so downloading a photograph looked like
a delete and an add, orphaning its thumbnail and its face rows.

Entries now carry the photograph's own name and a `materialised` flag;
`get` on a stub returns the new `RemoteError::NotMaterialised`, which is
distinct from `NotFound` precisely because the sidecar writer must treat
them differently — it fetches the sidecar and merges, or leaves the entry
queued.

Hydration is a **borrow**. `BorrowPool` records what was on disk before it
asked, so `release_all` dehydrates only what a pass brought and leaves
what the user already had. Reference counted: the thumbnail pass and the
face pass meet on the same RAW, and without counting the first to finish
dehydrates the file the second is reading. A borrow against a plain folder
or a server does nothing, so a pass written for VFS runs everywhere.

Releasing means asking the client to dehydrate and never deleting: a
deletion inside a synced tree propagates to the server and removes the
photograph from every device.

Not a second backend — the capability is per *connection*, not per type,
since the same folder hydrates only while the client runs. The convention
arrives through a detector the registry supplies, so `dr-sync-folder`
still knows nothing about any client's protocol.

ARCH §9.0a records this as an amendment: finding 3 rejected hydration
because it costs 100× a range read, and that comparison assumed a
connector was available. A folder library has none.
This commit is contained in:
2026-08-29 09:57:52 +02:00
parent 6c363cee97
commit c102ba9df2
22 changed files with 1555 additions and 53 deletions
+242
View File
@@ -0,0 +1,242 @@
//! The borrow contract, against a filesystem and a fake client.
//!
//! The fake stands in for the sync client's socket, not for the filesystem:
//! it renames stubs exactly as suffix-mode VFS does, so everything under test
//! is the real path resolution and the real state tracking.
use super::*;
use crate::{FolderBackend, Vfs};
use std::path::{Path, PathBuf};
use std::sync::atomic::{AtomicUsize, Ordering};
/// A stand-in for a sync client, counting what it was asked to do.
struct FakeClient {
suffix: &'static str,
hydrations: AtomicUsize,
dehydrations: AtomicUsize,
/// When true, refuse to hydrate — the client is running but the server is
/// not reachable.
broken: bool,
}
impl FakeClient {
fn new() -> Arc<Self> {
Arc::new(Self {
suffix: ".nextcloud",
hydrations: AtomicUsize::new(0),
dehydrations: AtomicUsize::new(0),
broken: false,
})
}
fn broken() -> Arc<Self> {
Arc::new(Self {
suffix: ".nextcloud",
hydrations: AtomicUsize::new(0),
dehydrations: AtomicUsize::new(0),
broken: true,
})
}
}
impl Vfs for FakeClient {
fn name(&self) -> &'static str {
"fake"
}
fn is_placeholder(&self, on_disk: &str) -> bool {
on_disk.ends_with(self.suffix)
}
fn real_name<'a>(&self, on_disk: &'a str) -> &'a str {
on_disk.strip_suffix(self.suffix).unwrap_or(on_disk)
}
fn placeholder_name(&self, name: &str) -> std::borrow::Cow<'_, str> {
std::borrow::Cow::Owned(format!("{name}{}", self.suffix))
}
fn can_materialise(&self) -> bool {
true
}
fn materialise(&self, local: &Path) -> Result<(), RemoteError> {
self.hydrations.fetch_add(1, Ordering::SeqCst);
if self.broken {
return Err(RemoteError::Network("no server".into()));
}
// Suffix mode renames rather than filling in place, and writes the
// real content.
let real = PathBuf::from(local.to_string_lossy().strip_suffix(self.suffix).unwrap());
std::fs::write(&real, vec![9u8; 4096]).unwrap();
std::fs::remove_file(local).unwrap();
Ok(())
}
fn dematerialise(&self, local: &Path) -> Result<(), RemoteError> {
self.dehydrations.fetch_add(1, Ordering::SeqCst);
let stub = format!("{}{}", local.display(), self.suffix);
std::fs::write(&stub, [0u8]).unwrap();
std::fs::remove_file(local).unwrap();
Ok(())
}
}
struct Tmp(PathBuf);
impl Tmp {
fn new(name: &str) -> Self {
let d = std::env::temp_dir().join(format!("dr-borrow-{name}"));
let _ = std::fs::remove_dir_all(&d);
std::fs::create_dir_all(&d).unwrap();
Tmp(d)
}
/// A dehydrated photograph.
fn stub(&self, rel: &str) -> &Self {
std::fs::write(self.0.join(format!("{rel}.nextcloud")), [0u8]).unwrap();
self
}
/// One the user already has.
fn real(&self, rel: &str) -> &Self {
std::fs::write(self.0.join(rel), vec![1u8; 2048]).unwrap();
self
}
fn has(&self, rel: &str) -> bool {
self.0.join(rel).is_file()
}
fn backend(&self, vfs: Arc<dyn Vfs>) -> FolderBackend {
FolderBackend::with_vfs(&self.0, vfs).unwrap()
}
}
impl Drop for Tmp {
fn drop(&mut self) {
let _ = std::fs::remove_dir_all(&self.0);
}
}
#[tokio::test]
async fn a_borrowed_placeholder_is_downloaded_and_given_back() {
let t = Tmp::new("cycle");
t.stub("a.CR2");
let client = FakeClient::new();
let b = t.backend(client.clone());
let pool = BorrowPool::new();
let path = RemotePath::new("a.CR2");
{
let held = pool.borrow(&b, &path).await.unwrap();
assert!(held.hydrated(), "this borrow paid for it");
assert!(t.has("a.CR2"), "content is here while borrowed");
assert_eq!(
b.get(&RemoteId::Path(path.clone()), None)
.await
.unwrap()
.len(),
4096
);
}
let stats = pool.release_all(&b).await;
assert_eq!(stats.released, 1);
assert!(!t.has("a.CR2"), "given back");
assert!(t.has("a.CR2.nextcloud"), "a placeholder is left behind");
assert_eq!(client.dehydrations.load(Ordering::SeqCst), 1);
}
#[tokio::test]
async fn a_file_the_user_already_had_is_never_taken_away() {
// The rule the whole design rests on. Silently undoing a pin — or just a
// file someone opened yesterday — after an indexing run is the failure
// that would make people stop trusting this.
let t = Tmp::new("keep");
t.real("pinned.CR2");
let client = FakeClient::new();
let b = t.backend(client.clone());
let pool = BorrowPool::new();
{
let held = pool
.borrow(&b, &RemotePath::new("pinned.CR2"))
.await
.unwrap();
assert!(!held.hydrated(), "nothing was downloaded");
}
let stats = pool.release_all(&b).await;
assert_eq!(stats.released, 0);
assert!(t.has("pinned.CR2"), "still here");
assert_eq!(client.hydrations.load(Ordering::SeqCst), 0);
assert_eq!(client.dehydrations.load(Ordering::SeqCst), 0);
}
#[tokio::test]
async fn two_lanes_wanting_one_file_download_it_once() {
// The thumbnail pass and the face pass meet on the same RAW. Without
// counting, the first to finish dehydrates the file the second is reading.
let t = Tmp::new("shared");
t.stub("a.CR2");
let client = FakeClient::new();
let b = t.backend(client.clone());
let pool = BorrowPool::new();
let path = RemotePath::new("a.CR2");
let first = pool.borrow(&b, &path).await.unwrap();
let second = pool.borrow(&b, &path).await.unwrap();
assert_eq!(client.hydrations.load(Ordering::SeqCst), 1, "paid once");
drop(first);
assert!(t.has("a.CR2"), "still held by the second borrower");
drop(second);
pool.release_all(&b).await;
assert!(!t.has("a.CR2"));
}
#[tokio::test]
async fn a_failed_download_does_not_leave_a_phantom_borrow() {
// The client is up but the server is not. The pass must see the failure
// and the pool must not believe it holds anything.
let t = Tmp::new("failed");
t.stub("a.CR2");
let b = t.backend(FakeClient::broken());
let pool = BorrowPool::new();
let e = pool
.borrow(&b, &RemotePath::new("a.CR2"))
.await
.unwrap_err();
assert!(matches!(e, RemoteError::Network(_)), "{e:?}");
assert_eq!(pool.held(), 0);
assert!(t.has("a.CR2.nextcloud"), "left as it was found");
}
#[tokio::test]
async fn borrowing_against_a_plain_folder_does_nothing_at_all() {
// A caller written for a VFS library must run unchanged elsewhere, or
// every sweep grows two code paths.
let t = Tmp::new("plain");
t.real("a.CR2");
let b = FolderBackend::new(&t.0).unwrap();
let pool = BorrowPool::new();
let held = pool.borrow(&b, &RemotePath::new("a.CR2")).await.unwrap();
assert!(!held.hydrated());
drop(held);
assert_eq!(pool.release_all(&b).await.released, 0);
assert!(t.has("a.CR2"));
}
#[tokio::test]
async fn an_uncertain_caller_keeps_the_file_rather_than_releasing_it() {
// The asymmetry stated on `borrow_known`: claiming "already local" costs
// disk, claiming "not local" can release a pin.
let t = Tmp::new("uncertain");
t.real("a.CR2");
let client = FakeClient::new();
let b = t.backend(client.clone());
let pool = BorrowPool::new();
{
let _held = pool
.borrow_known(&b, &RemotePath::new("a.CR2"), Some(true))
.await
.unwrap();
}
pool.release_all(&b).await;
assert!(t.has("a.CR2"), "kept");
assert_eq!(client.dehydrations.load(Ordering::SeqCst), 0);
}