Write face shards the way the rest of the catalog writes

The face store was the one part of the catalog still on SQLite's default
rollback journal at `synchronous = FULL`. The catalog itself runs WAL at
`NORMAL` (`schema::configure`) and so does the thumbnail store; nothing decided
this one should differ, it was simply never set.

Measured on this project's own filesystem, that is **21.3 ms per commit against
0.05 ms** — four hundred times. And an export commits four times per
photograph: the shard's transaction, then three separate autocommitting writes
to the index. Ten thousand images is on the order of fourteen minutes spent
doing nothing but waiting for fsync, before a byte goes to the server. That is
the "checking faces…" that appeared to hang.

So: WAL and `synchronous = NORMAL`, matching the rest, and the three index
writes fold into one transaction. `NORMAL` is the same trade the catalog makes —
a shard is derived data, and losing the last commit to a power cut costs one
image re-exported.

WAL brings an obligation with it, because **a shard is uploaded by reading its
file**: the newest commits live in a `-wal` sidecar that no upload sends, so
without a checkpoint the server would receive a database missing exactly the
faces just written, and a peer would adopt it and see nothing wrong. `checkpoint`
folds the logs back in, with `TRUNCATE` rather than the default passive mode,
which gives up when a reader holds the log and would leave the same gap while
reporting success.

Two tests: that the store is in WAL like everything else, and — the one that
matters — that a checkpointed shard copied *without* its `-wal` still holds
every face. That second one fails without the checkpoint, which is how it was
confirmed to be testing something.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-28 23:08:10 +02:00
co-authored by Claude Opus 5
parent 0bba882fb1
commit 2c84aa1224
2 changed files with 104 additions and 4 deletions
+103 -3
View File
@@ -103,6 +103,7 @@ impl FaceShardStore {
pub fn open(dir: &Path) -> Result<Self, CatalogError> {
std::fs::create_dir_all(dir).map_err(|e| CatalogError::Io(e.to_string()))?;
let index = Connection::open(dir.join("index.sqlite"))?;
fast_writes(&index)?;
index.execute_batch(INDEX_SCHEMA)?;
// `INDEX_SCHEMA` is `CREATE ... IF NOT EXISTS` like the shard schema,
// so a column added to it never reaches an index already on disk. The
@@ -257,7 +258,12 @@ impl FaceShardStore {
)?;
tx.commit()?;
self.index.execute(
// One transaction, not three. Each of these was autocommitting, and an
// autocommit is a durable write — so a single photograph cost four
// commits counting the shard's own, and an export of ten thousand paid
// forty thousand of them.
let ix = self.index.unchecked_transaction()?;
ix.execute(
"INSERT INTO entries (file_id, model_id, shard, bytes, indexed_at)
VALUES (?1, ?2, ?3, ?4, ?5)
ON CONFLICT(file_id, model_id) DO UPDATE SET
@@ -272,15 +278,16 @@ impl FaceShardStore {
indexed_at
],
)?;
self.index.execute(
ix.execute(
"INSERT INTO faces_meta (file_id, model_id) VALUES (?1, ?2)
ON CONFLICT(file_id, model_id) DO NOTHING",
rusqlite::params![file_id as i64, model_id],
)?;
self.index.execute(
ix.execute(
"UPDATE shards SET bytes = bytes + ?2 WHERE id = ?1",
rusqlite::params![shard as i64, incoming as i64],
)?;
ix.commit()?;
Ok(shard)
}
@@ -341,11 +348,32 @@ impl FaceShardStore {
return Err(CatalogError::Io(format!("face shard {shard} is missing")));
}
let conn = Connection::open(&path)?;
fast_writes(&conn)?;
conn.execute_batch(SHARD_SCHEMA)?;
upgrade_shard(&conn)?;
Ok(conn)
}
/// Fold every write-ahead log back into the database files.
///
/// **A shard is uploaded by reading its file.** Under WAL the most recent
/// commits live in a `-wal` sidecar until a checkpoint moves them, so
/// reading the `.sqlite` alone would ship a database missing exactly the
/// faces just written — and a peer would adopt it and see nothing wrong.
/// The sync calls this before it reads anything.
///
/// `TRUNCATE` rather than the default passive checkpoint: passive gives up
/// when a reader holds the log, which would leave the same gap while
/// reporting success.
pub fn checkpoint(&self) -> Result<(), CatalogError> {
self.index
.execute_batch("PRAGMA wal_checkpoint(TRUNCATE)")?;
if let Some((_, conn)) = self.writer.as_ref() {
conn.execute_batch("PRAGMA wal_checkpoint(TRUNCATE)")?;
}
Ok(())
}
pub fn shard_path(&self, shard: u32) -> PathBuf {
self.dir
.join(format!("shard-{}-{shard:04}.sqlite", self.client))
@@ -469,6 +497,27 @@ impl FaceShardStore {
}
}
/// Put a connection into the mode the rest of the catalog already uses.
///
/// The face store was the one place still on SQLite's default rollback journal
/// at `synchronous = FULL`, where the catalog (`schema::configure`) and the
/// thumbnail store both run WAL at `NORMAL`. Measured on this project's own
/// filesystem the difference is **21.3 ms per commit against 0.05 ms** — four
/// hundred times — and an export commits several times per photograph, so ten
/// thousand images spent something like fourteen minutes doing nothing but
/// waiting for fsync. That was the "checking faces…" that never finished.
///
/// `NORMAL` rather than `FULL` is the same trade the rest of the catalog makes:
/// a shard is derived data, and the cost of losing the last commit to a power
/// cut is that the next pass exports that image again.
fn fast_writes(conn: &Connection) -> Result<(), CatalogError> {
// A `journal_mode` change returns the new mode as a row, so it has to be
// queried rather than executed.
conn.query_row("PRAGMA journal_mode = WAL", [], |_| Ok(()))?;
conn.execute_batch("PRAGMA synchronous = NORMAL")?;
Ok(())
}
/// Add the columns a shard written by an older build is missing.
///
/// [`SHARD_SCHEMA`] is entirely `CREATE ... IF NOT EXISTS`, which does exactly
@@ -1431,4 +1480,55 @@ mod catalog_round_trip {
assert_eq!(s.indexed_at(6, "w600k_mbf"), Some(11));
let _ = std::fs::remove_dir_all(&dir);
}
/// The face store was the one part of the catalog still on the rollback
/// journal at `synchronous = FULL`, which cost 21 ms a commit where the
/// rest pays 0.05 ms. An export commits several times per photograph.
#[test]
fn the_store_writes_the_way_the_rest_of_the_catalog_does() {
let dir = tempdir("wal");
let mut s = FaceShardStore::open(&dir).unwrap();
s.put_image(1, "w600k_mbf", 1024, &[old_face(1, 1)])
.unwrap();
for db in [dir.join("index.sqlite"), s.shard_path(0)] {
let c = Connection::open(&db).unwrap();
let mode: String = c
.query_row("PRAGMA journal_mode", [], |r| r.get(0))
.unwrap();
assert_eq!(mode, "wal", "{} is not in WAL", db.display());
}
let _ = std::fs::remove_dir_all(&dir);
}
/// **The property the upload depends on.** A shard is sent by reading its
/// file; under WAL the newest commits sit in a `-wal` sidecar that is not
/// sent with it. Without a checkpoint the server would receive a database
/// missing exactly the faces just exported, and a peer would adopt it and
/// see nothing wrong.
#[test]
fn a_checkpointed_shard_stands_alone_without_its_write_ahead_log() {
let dir = tempdir("checkpoint");
let mut s = FaceShardStore::open(&dir).unwrap();
for id in 1..=5u64 {
s.put_image(id, "w600k_mbf", 1024, &[old_face(id, id as u8)])
.unwrap();
}
s.checkpoint().unwrap();
// Copy *only* the database, exactly as the upload reads it.
let sent = dir.join("as-uploaded.sqlite");
std::fs::copy(s.shard_path(0), &sent).unwrap();
let c = Connection::open_with_flags(
&sent,
rusqlite::OpenFlags::SQLITE_OPEN_READ_ONLY | rusqlite::OpenFlags::SQLITE_OPEN_NO_MUTEX,
)
.unwrap();
let faces: i64 = c
.query_row("SELECT COUNT(*) FROM faces", [], |r| r.get(0))
.unwrap();
assert_eq!(faces, 5, "the uploaded file is missing recent writes");
let _ = std::fs::remove_dir_all(&dir);
}
}