Give a shard's remote name the client that wrote it
Build and test / Desktop (Linux) (push) Failing after 50s
Build and test / Layer separation (push) Successful in 24s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
Traceability / Requirement traces (push) Failing after 59s
Build and test / Android (aarch64) (push) Failing after 9m39s

Shard ids are per store: every client fills its own numbering from 0, so
"shard 3" names different thumbnails on every device. The derived sync
published them into a flat shard-NNNN.sqlite namespace anyway, which left
two clients writing one name.

Both failures that follow were live. On upload, a client's open shard
overwrote a peer's file of the same id — content the peer still believed
was published and would never restore, because its own copy was sealed and
the name existed. On download, the loop skipped any remote id it already
held locally, which is the only safe reading of a name that says nothing
about who wrote it, so a client holding shards 0..5 never fetched the
peer's 0..5 at all. Between them, two populated clients exchanged almost
nothing: only shards numbered above the other's highest. A fresh device
worked, having no local shards to collide with, which is why this went
unnoticed — it is exactly the case the feature was written for.

The name is now shard-<client>-NNNN.sqlite. The client id is minted per
store in index.sqlite, beside the numbering it qualifies rather than in
settings: a store deleted and rebuilt restarts at shard 0 and must not
claim the remote names its predecessor wrote. Since our own ids now say
nothing about what we have taken from others, index.sqlite also keeps a
ledger of adopted remote names and the size each had when merged. A size
rather than a flag, because a peer's sealed shard never returns but its
open one grows, and re-merging the grown copy is how the thumbnails it
gained since arrive.

Flat names already on servers still parse, reporting no owner, so each
client adopts them once, and nothing is written under that form again. One
whose id and byte size match a local shard is that client's own earlier
upload by the same identity argument the upload path already makes for
sealed shards, so the rename does not cost every client a re-download of
its whole store. Older builds ignore the new names and stop receiving
shards until updated; their own uploads are still adopted, so nothing is
lost.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-08-16 21:15:12 +02:00
co-authored by Claude Opus 5
parent f7e8cc99b1
commit 9a51cc88d6
3 changed files with 233 additions and 28 deletions
+107
View File
@@ -124,6 +124,7 @@ pub struct Thumbnail {
pub struct ThumbStore {
dir: PathBuf,
index: Connection,
client: String,
}
impl ThumbStore {
@@ -139,13 +140,47 @@ impl ThumbStore {
// entries under a bare `file_id` key; bring it forward rather than
// discarding every thumbnail already fetched.
migrate_size_column(&index, "entries")?;
let client = mint_client_id(&index)?;
Ok(Self {
dir: dir.to_path_buf(),
index,
client,
})
}
/// This store's identity among the clients sharing a library.
///
/// Shard ids are per-store — every client fills its own numbering from 0 —
/// so a name that travels needs this to say *whose* shard 3 it is.
pub fn client_id(&self) -> &str {
&self.client
}
/// The size a remote shard had when it was last merged, if it ever was.
///
/// A size rather than a bare flag: a peer's sealed shard never changes
/// again, but its open one grows, and re-merging the grown copy is how the
/// thumbnails it gained since arrive.
pub fn adopted(&self, name: &str) -> Option<u64> {
self.index
.query_row("SELECT size FROM adopted WHERE name = ?1", [name], |r| {
r.get::<_, i64>(0)
})
.ok()
.map(|bytes| bytes as u64)
}
/// Record that a remote shard of this name and size has been merged.
pub fn record_adopted(&self, name: &str, size: u64) -> Result<(), ThumbError> {
self.index.execute(
"INSERT INTO adopted(name, size) VALUES (?1, ?2)
ON CONFLICT(name) DO UPDATE SET size = excluded.size",
rusqlite::params![name, size as i64],
)?;
Ok(())
}
/// Fetch a thumbnail by file id.
pub fn get(&self, file_id: u64, size: ThumbSize) -> Result<Option<Thumbnail>, ThumbError> {
let shard: Option<i64> = self
@@ -566,8 +601,43 @@ CREATE TABLE IF NOT EXISTS shards (
-- already hold it would see it change.
sealed INTEGER NOT NULL DEFAULT 0
);
-- Facts about this store rather than about any thumbnail. The index never
-- leaves the device, so what lives here is safe to be device-specific.
CREATE TABLE IF NOT EXISTS meta (
key TEXT PRIMARY KEY,
value TEXT NOT NULL
);
-- Remote shards already merged, under the name and size they had on the
-- server. Nothing in a merged thumbnail records where it came from, so without
-- this ledger a client either re-downloads every peer's shard on every sync or
-- guesses from its own numbering — and its own numbering says nothing about
-- anyone else's.
CREATE TABLE IF NOT EXISTS adopted (
name TEXT PRIMARY KEY,
size INTEGER NOT NULL
);
"#;
/// Read this store's client id, minting one on first open.
///
/// It lives in the index because that is where the shard numbering it
/// qualifies lives: a store deleted and rebuilt starts again at shard 0, and
/// must not claim the remote names its predecessor wrote. SQLite's own
/// randomness keeps it dependency-free, and six bytes separate far more
/// devices than one account ever has.
fn mint_client_id(conn: &Connection) -> Result<String, ThumbError> {
conn.execute(
"INSERT INTO meta(key, value) VALUES('client_id', lower(hex(randomblob(6))))
ON CONFLICT(key) DO NOTHING",
[],
)?;
Ok(conn.query_row("SELECT value FROM meta WHERE key = 'client_id'", [], |r| {
r.get(0)
})?)
}
const SHARD_SCHEMA: &str = r#"
CREATE TABLE IF NOT EXISTS thumbs (
file_id INTEGER NOT NULL,
@@ -985,6 +1055,43 @@ mod tests {
assert!(p.to_string_lossy().ends_with("shard-0007.sqlite"));
}
#[test]
fn every_store_gets_its_own_stable_client_id() {
let (mine, dir) = store();
let id = mine.client_id().to_string();
assert!(!id.is_empty());
drop(mine);
// Stable across reopen, or a client would orphan its own uploads and
// re-download them as if they were a peer's.
assert_eq!(ThumbStore::open(&dir).unwrap().client_id(), id);
// Distinct per store, which is the property the remote naming rests
// on: two devices both filling shard 0 must not name one file.
let (theirs, _d) = store();
assert_ne!(theirs.client_id(), id);
}
#[test]
fn the_adoption_ledger_remembers_a_merged_shard_by_size() {
let (mine, dir) = store();
assert_eq!(mine.adopted("shard-abc-0000.sqlite"), None);
mine.record_adopted("shard-abc-0000.sqlite", 1234).unwrap();
assert_eq!(mine.adopted("shard-abc-0000.sqlite"), Some(1234));
drop(mine);
// Survives reopen: the point is to not re-download across sessions.
let reopened = ThumbStore::open(&dir).unwrap();
assert_eq!(reopened.adopted("shard-abc-0000.sqlite"), Some(1234));
// A peer's open shard grows, and the new size is what says so.
reopened
.record_adopted("shard-abc-0000.sqlite", 5678)
.unwrap();
assert_eq!(reopened.adopted("shard-abc-0000.sqlite"), Some(5678));
}
#[test]
fn reopening_finds_what_was_stored() {
let (mut s, dir) = store();