Build the upload snapshot without the face crops instead of stripping them

Each sync pass spent 0.8-2.0 s of CPU and 1.0-5.4 s wall on the upload
snapshot of the reference catalog (24k images, 18,871 faces), ahead of the
rest of the pass. The upload itself had been crop-less since the crops
moved to the face shards. The cost was in how it got that way. The backup
API copied all 158 MB of the catalog, 96 MB of it the ~5 KB JPEG crop on
every faces row. Then `UPDATE faces SET crop = NULL` rewrote 18.9k rows
and freed their overflow chains, and VACUUM rebuilt the file again. That
wrote the catalog about three times over to upload 50 MB.

The snapshot is now built rather than copied. An empty file attaches the
catalog, creates each table from the catalog's own sqlite_master and fills
it with INSERT ... SELECT, with faces.crop selected as NULL. Indexes,
triggers and views follow, and user_version, application_id, page size and
the WAL header flag are carried over. It all runs in one transaction on the
snapshot's connection, so the catalog is read as of one moment and
concurrent writers are serialised, not raced, as the backup API did. The
build journal is in memory with synchronous off, because the file is
scratch that is rebuilt every pass and quick_check'd before upload. Foreign
keys are off on that connection. The bundled SQLite enables them, and then
a multi-row INSERT into images scans images for children of each new row
(shadowed_by is a self-reference with no index), which cost 1.2 s alone.

Measured on a .backup copy of the reference catalog with catalog_bench,
old and new binaries back to back on a loaded machine:
  before  best 1.0-5.4 s wall, 0.84-1.98 s cpu, 49.8 MB
  after   best 0.40-2.1 s wall, 0.39-0.96 s cpu, 50.4 MB
With the machine quiet the new build takes 0.31-0.43 s.

What a receiving device gets is unchanged. It is the same schema, the same
rows and a NULL crop, which is what 0.16.0 already uploads and merges. The
merge reads only a remote face's box and model (merge::match_faces) and
never writes a local crop. No device adopts a downloaded catalog as its
own, and a fresh one takes faces and crops from the shards. There is no
schema bump, so older builds still merge it. NFR-R2 backups keep using the
backup API and keep their crops.

Tests: the snapshot matches the catalog in schema, row counts, pragmas and
WAL header. A leftover file is replaced. Merging a crop-less snapshot
carries a confirmed name across by box and leaves the local crop
untouched, and does so idempotently.
This commit is contained in:
2026-09-26 14:03:27 -04:00
parent a59f14c797
commit 49b7bc2f9d
2 changed files with 401 additions and 48 deletions
+17 -3
View File
@@ -698,9 +698,23 @@ pair that had no other home.
**A WAL database is not one file.** Committed transactions can sit in `catalog.sqlite-wal` with the
main file lagging, so copying `catalog.sqlite` alone uploads a torn snapshot — internally consistent
as of some older point, silently missing everything since. Upload therefore runs a `TRUNCATE`
checkpoint and then SQLite's backup API, which serialises against concurrent writers rather than
racing them. It never copies the live file.
as of some older point, silently missing everything since. The upload therefore never copies the
live file. It builds the snapshot in an empty file. It attaches the catalog, creates each table
from the catalog's own schema, and fills it with `INSERT … SELECT`, all inside one transaction.
That transaction holds a single read snapshot of the catalog, so concurrent writers are serialised
rather than raced, as the backup API did before. The file is then checked with `quick_check` before
it goes anywhere.
**The face crops stay out of the upload.** A crop is a ~5 KB JPEG on each `faces` row. On a 19k-face
library they are 96 MB of a 158 MB catalog. The face shards carry them to other devices, once each.
The merge reads a remote face's box and model to match it to a local one, never its pixels. No
device adopts a downloaded catalog as its own: a fresh device starts empty and takes faces, crops
included, from the shards. So the snapshot's `crop` is NULL, and a merge never writes a local
crop. They were first stripped (2026-08) by copying the whole file with the backup API, setting
`crop` to NULL and `VACUUM`ing. That wrote the file about three times to upload 50 MB. Since #71
the snapshot is built without them. Its header is kept as it was (`user_version`, page size and
the WAL flag), so every earlier build merges it unchanged. NFR-R2 backups still use the backup API
and keep the crops, because a backup is a file the user may have to live on.
**Integer primary keys are not identities.** Two devices each allocate `collections.id = 1` for
different collections, so a row-level merge keyed on the integer id would collide them. Collections