Upload a snapshot of a thumbnail shard, never the live file

Every shard is in WAL mode and every put opens its own connection, so
while thumbnails are being generated on several threads — which is when
the first sync pass runs — the log is never checkpointed and the main
file holds whatever the last quiet moment left in it. For a shard created
seconds earlier that is nothing: a zero-byte file with the schema still
in the log. The sync read that file and uploaded it, and every other
device merging it failed with "no such table: thumbs" on every pass.

Copy the shard through SQLite's backup API into scratch first, which
serialises against writers and carries the log, and upload that.
This commit is contained in:
2026-09-20 00:21:14 +02:00
parent 0fa9003e54
commit 7fba28f7d8
4 changed files with 89 additions and 15 deletions
+63
View File
@@ -481,6 +481,30 @@ impl ThumbStore {
self.dir.join(format!("shard-{shard:04}.sqlite"))
}
/// Write a coherent copy of one shard to `dest`, ready to upload.
///
/// Not a file copy. Every shard is in WAL mode and every `put` opens its
/// own connection, so while thumbnails are being generated on several
/// threads at once — which is exactly when the first sync pass runs —
/// there is nearly always a connection open and the log is never
/// checkpointed. The main file then holds whatever the *last* quiet
/// moment left in it, which for a shard created seconds ago is nothing:
/// zero bytes, the schema still in the log. Reading it uploaded an empty
/// file, and every other device merging it failed with "no such table:
/// thumbs". The backup API serialises against writers and copies the
/// database as it is, log included.
pub fn snapshot_shard(&self, shard: u32, dest: &Path) -> Result<(), ThumbError> {
let source = self.open_shard(shard, false)?;
let _ = std::fs::remove_file(dest);
let mut out = Connection::open(dest)?;
let backup = rusqlite::backup::Backup::new(&source, &mut out)?;
// rusqlite asserts a positive page count where SQLite would take -1
// for "everything"; a shard is capped well under this many pages.
backup.run_to_completion(i32::MAX, std::time::Duration::ZERO, None)?;
drop(backup);
Ok(())
}
pub fn index_path(&self) -> PathBuf {
self.dir.join("index.sqlite")
}
@@ -726,6 +750,45 @@ mod tests {
}
}
#[test]
fn a_snapshot_carries_what_the_shard_file_does_not_yet() {
// A thumbnail worker holding the shard open keeps the log from being
// checkpointed; the file on disk is then not the database. The
// upload used to read that file.
let (mut s, _d) = store();
s.put(1, ThumbSize::Grid, &thumb(1024)).unwrap();
let path = s.shard_path(0);
let worker = Connection::open(&path).unwrap();
worker.pragma_update(None, "journal_mode", "WAL").unwrap();
s.put(2, ThumbSize::Grid, &thumb(2048)).unwrap();
s.put(3, ThumbSize::Grid, &thumb(2048)).unwrap();
// What a byte-for-byte reader sees is at most what the last
// checkpoint left; the log is the part a copy misses.
let copy = _d.join("copied.sqlite");
std::fs::copy(&path, &copy).unwrap();
let file_rows = Connection::open(&copy)
.ok()
.and_then(|c| {
c.query_row("SELECT count(*) FROM thumbs", [], |r| r.get::<_, i64>(0))
.ok()
})
.unwrap_or(0);
let snap = _d.join("upload.sqlite");
s.snapshot_shard(0, &snap).unwrap();
let snapshot = Connection::open(&snap).unwrap();
let rows: i64 = snapshot
.query_row("SELECT count(*) FROM thumbs", [], |r| r.get(0))
.unwrap();
assert_eq!(rows, 3, "the snapshot is the whole shard");
assert!(
file_rows < 3,
"the file alone lagged ({file_rows} rows), which is what the snapshot exists for"
);
drop(worker);
}
#[test]
fn a_forgotten_thumbnail_is_no_longer_served() {
// The point of the whole method: a purged photograph must not keep a