Build the upload snapshot without the face crops instead of stripping them
Each sync pass spent 0.8-2.0 s of CPU and 1.0-5.4 s wall on the upload snapshot of the reference catalog (24k images, 18,871 faces), ahead of the rest of the pass. The upload itself had been crop-less since the crops moved to the face shards. The cost was in how it got that way. The backup API copied all 158 MB of the catalog, 96 MB of it the ~5 KB JPEG crop on every faces row. Then `UPDATE faces SET crop = NULL` rewrote 18.9k rows and freed their overflow chains, and VACUUM rebuilt the file again. That wrote the catalog about three times over to upload 50 MB. The snapshot is now built rather than copied. An empty file attaches the catalog, creates each table from the catalog's own sqlite_master and fills it with INSERT ... SELECT, with faces.crop selected as NULL. Indexes, triggers and views follow, and user_version, application_id, page size and the WAL header flag are carried over. It all runs in one transaction on the snapshot's connection, so the catalog is read as of one moment and concurrent writers are serialised, not raced, as the backup API did. The build journal is in memory with synchronous off, because the file is scratch that is rebuilt every pass and quick_check'd before upload. Foreign keys are off on that connection. The bundled SQLite enables them, and then a multi-row INSERT into images scans images for children of each new row (shadowed_by is a self-reference with no index), which cost 1.2 s alone. Measured on a .backup copy of the reference catalog with catalog_bench, old and new binaries back to back on a loaded machine: before best 1.0-5.4 s wall, 0.84-1.98 s cpu, 49.8 MB after best 0.40-2.1 s wall, 0.39-0.96 s cpu, 50.4 MB With the machine quiet the new build takes 0.31-0.43 s. What a receiving device gets is unchanged. It is the same schema, the same rows and a NULL crop, which is what 0.16.0 already uploads and merges. The merge reads only a remote face's box and model (merge::match_faces) and never writes a local crop. No device adopts a downloaded catalog as its own, and a fresh one takes faces and crops from the shards. There is no schema bump, so older builds still merge it. NFR-R2 backups keep using the backup API and keep their crops. Tests: the snapshot matches the catalog in schema, row counts, pragmas and WAL header. A leftover file is replaced. Merging a crop-less snapshot carries a confirmed name across by box and leaves the local crop untouched, and does so idempotently.
This commit is contained in:
+17
-3
@@ -698,9 +698,23 @@ pair that had no other home.
|
||||
|
||||
**A WAL database is not one file.** Committed transactions can sit in `catalog.sqlite-wal` with the
|
||||
main file lagging, so copying `catalog.sqlite` alone uploads a torn snapshot — internally consistent
|
||||
as of some older point, silently missing everything since. Upload therefore runs a `TRUNCATE`
|
||||
checkpoint and then SQLite's backup API, which serialises against concurrent writers rather than
|
||||
racing them. It never copies the live file.
|
||||
as of some older point, silently missing everything since. The upload therefore never copies the
|
||||
live file. It builds the snapshot in an empty file. It attaches the catalog, creates each table
|
||||
from the catalog's own schema, and fills it with `INSERT … SELECT`, all inside one transaction.
|
||||
That transaction holds a single read snapshot of the catalog, so concurrent writers are serialised
|
||||
rather than raced, as the backup API did before. The file is then checked with `quick_check` before
|
||||
it goes anywhere.
|
||||
|
||||
**The face crops stay out of the upload.** A crop is a ~5 KB JPEG on each `faces` row. On a 19k-face
|
||||
library they are 96 MB of a 158 MB catalog. The face shards carry them to other devices, once each.
|
||||
The merge reads a remote face's box and model to match it to a local one, never its pixels. No
|
||||
device adopts a downloaded catalog as its own: a fresh device starts empty and takes faces, crops
|
||||
included, from the shards. So the snapshot's `crop` is NULL, and a merge never writes a local
|
||||
crop. They were first stripped (2026-08) by copying the whole file with the backup API, setting
|
||||
`crop` to NULL and `VACUUM`ing. That wrote the file about three times to upload 50 MB. Since #71
|
||||
the snapshot is built without them. Its header is kept as it was (`user_version`, page size and
|
||||
the WAL flag), so every earlier build merges it unchanged. NFR-R2 backups still use the backup API
|
||||
and keep the crops, because a backup is a file the user may have to live on.
|
||||
|
||||
**Integer primary keys are not identities.** Two devices each allocate `collections.id = 1` for
|
||||
different collections, so a row-level merge keyed on the integer id would collide them. Collections
|
||||
|
||||
Reference in New Issue
Block a user