Files
DarkRoom/core/dr-catalog
dtourolle 49b7bc2f9d Build the upload snapshot without the face crops instead of stripping them
Each sync pass spent 0.8-2.0 s of CPU and 1.0-5.4 s wall on the upload
snapshot of the reference catalog (24k images, 18,871 faces), ahead of the
rest of the pass. The upload itself had been crop-less since the crops
moved to the face shards. The cost was in how it got that way. The backup
API copied all 158 MB of the catalog, 96 MB of it the ~5 KB JPEG crop on
every faces row. Then `UPDATE faces SET crop = NULL` rewrote 18.9k rows
and freed their overflow chains, and VACUUM rebuilt the file again. That
wrote the catalog about three times over to upload 50 MB.

The snapshot is now built rather than copied. An empty file attaches the
catalog, creates each table from the catalog's own sqlite_master and fills
it with INSERT ... SELECT, with faces.crop selected as NULL. Indexes,
triggers and views follow, and user_version, application_id, page size and
the WAL header flag are carried over. It all runs in one transaction on the
snapshot's connection, so the catalog is read as of one moment and
concurrent writers are serialised, not raced, as the backup API did. The
build journal is in memory with synchronous off, because the file is
scratch that is rebuilt every pass and quick_check'd before upload. Foreign
keys are off on that connection. The bundled SQLite enables them, and then
a multi-row INSERT into images scans images for children of each new row
(shadowed_by is a self-reference with no index), which cost 1.2 s alone.

Measured on a .backup copy of the reference catalog with catalog_bench,
old and new binaries back to back on a loaded machine:
  before  best 1.0-5.4 s wall, 0.84-1.98 s cpu, 49.8 MB
  after   best 0.40-2.1 s wall, 0.39-0.96 s cpu, 50.4 MB
With the machine quiet the new build takes 0.31-0.43 s.

What a receiving device gets is unchanged. It is the same schema, the same
rows and a NULL crop, which is what 0.16.0 already uploads and merges. The
merge reads only a remote face's box and model (merge::match_faces) and
never writes a local crop. No device adopts a downloaded catalog as its
own, and a fresh one takes faces and crops from the shards. There is no
schema bump, so older builds still merge it. NFR-R2 backups keep using the
backup API and keep their crops.

Tests: the snapshot matches the catalog in schema, row counts, pragmas and
WAL header. A leftover file is replaced. Merging a crop-less snapshot
carries a confirmed name across by box and leaves the local crop
untouched, and does so idempotently.
2026-09-26 14:03:27 -04:00
..