Every sync pass copies the whole catalog, face crops included, into an upload snapshot: 0.9–1.4 s of CPU and 1.1–2.7 s wall on the reference library, for a 49.5 MB upload.dr_catalog::sync::snapshot_for_upload (core/dr-catalog/src/sync.rs) is now the largest cost of a pass. The rest of a pass was cut in v0.16.0: the face merge dropped from ~500 to 279 ms, and the shard export and import from 119/250 to 27/87 ms.
Why
Most of those bytes are the ~5 KB JPEG crop on each of 19k faces rows (about 100 MB before compression in the file). The face shards already carry each face, so the catalog snapshot sends it a second time. The 2026-09-25 pass tried turning off the snapshot's journal and sync. That saved about 10% and changed the uploaded file's journal mode, so it was reverted. The saving has to come from not copying the crops.
Deliverable
A snapshot that leaves out what the receiving side can rebuild or already gets elsewhere, crops first. Or else a decision, recorded in docs/dev/catalog.md, that the crops belong in the catalog upload and why.
Acceptance
Snapshot time and upload size measured before and after on a .backup copy of the reference catalog (cargo run --release -p dr-catalog --example catalog_bench)
A catalog pulled from the new snapshot shows the same faces and crops on Android and desktop as before
Merge tests cover a snapshot without crops
**Every sync pass copies the whole catalog, face crops included, into an upload snapshot: 0.9–1.4 s of CPU and 1.1–2.7 s wall on the reference library, for a 49.5 MB upload.** `dr_catalog::sync::snapshot_for_upload` (`core/dr-catalog/src/sync.rs`) is now the largest cost of a pass. The rest of a pass was cut in v0.16.0: the face merge dropped from ~500 to 279 ms, and the shard export and import from 119/250 to 27/87 ms.
## Why
Most of those bytes are the ~5 KB JPEG crop on each of 19k `faces` rows (about 100 MB before compression in the file). The face shards already carry each face, so the catalog snapshot sends it a second time. The 2026-09-25 pass tried turning off the snapshot's journal and sync. That saved about 10% and changed the uploaded file's journal mode, so it was reverted. The saving has to come from not copying the crops.
## Deliverable
A snapshot that leaves out what the receiving side can rebuild or already gets elsewhere, crops first. Or else a decision, recorded in docs/dev/catalog.md, that the crops belong in the catalog upload and why.
## Acceptance
- [ ] Snapshot time and upload size measured before and after on a `.backup` copy of the reference catalog (`cargo run --release -p dr-catalog --example catalog_bench`)
- [ ] A catalog pulled from the new snapshot shows the same faces and crops on Android and desktop as before
- [ ] Merge tests cover a snapshot without crops
Done in 49b7bc2, released in v0.17.0. The premise here was wrong: the upload has left the crops out since 79c0520 (2026-08-27). The cost was how it got there: a backup-API copy of the whole 158 MB catalog, then UPDATE faces SET crop = NULL and a VACUUM. snapshot_for_upload now builds the snapshot directly, with each table created from the catalog's own schema and filled with INSERT … SELECT with the crop NULL, in one transaction. ~1.2 s → ~0.4 s CPU per pass on the reference catalog. Upload content and size are unchanged (~50 MB), and the file is identical in schema, rows, pragmas and header, so 0.16.0 peers merge it as before. Embeddings are now most of the upload; the merge will start reading them in #77.
Done in 49b7bc2, released in v0.17.0. The premise here was wrong: the upload has left the crops out since 79c0520 (2026-08-27). The cost was how it got there: a backup-API copy of the whole 158 MB catalog, then `UPDATE faces SET crop = NULL` and a VACUUM. `snapshot_for_upload` now builds the snapshot directly, with each table created from the catalog's own schema and filled with `INSERT … SELECT` with the crop NULL, in one transaction. ~1.2 s → ~0.4 s CPU per pass on the reference catalog. Upload content and size are unchanged (~50 MB), and the file is identical in schema, rows, pragmas and header, so 0.16.0 peers merge it as before. Embeddings are now most of the upload; the merge will start reading them in #77.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Every sync pass copies the whole catalog, face crops included, into an upload snapshot: 0.9–1.4 s of CPU and 1.1–2.7 s wall on the reference library, for a 49.5 MB upload.
dr_catalog::sync::snapshot_for_upload(core/dr-catalog/src/sync.rs) is now the largest cost of a pass. The rest of a pass was cut in v0.16.0: the face merge dropped from ~500 to 279 ms, and the shard export and import from 119/250 to 27/87 ms.Why
Most of those bytes are the ~5 KB JPEG crop on each of 19k
facesrows (about 100 MB before compression in the file). The face shards already carry each face, so the catalog snapshot sends it a second time. The 2026-09-25 pass tried turning off the snapshot's journal and sync. That saved about 10% and changed the uploaded file's journal mode, so it was reverted. The saving has to come from not copying the crops.Deliverable
A snapshot that leaves out what the receiving side can rebuild or already gets elsewhere, crops first. Or else a decision, recorded in docs/dev/catalog.md, that the crops belong in the catalog upload and why.
Acceptance
.backupcopy of the reference catalog (cargo run --release -p dr-catalog --example catalog_bench)Done in
49b7bc2, released in v0.17.0. The premise here was wrong: the upload has left the crops out since79c0520(2026-08-27). The cost was how it got there: a backup-API copy of the whole 158 MB catalog, thenUPDATE faces SET crop = NULLand a VACUUM.snapshot_for_uploadnow builds the snapshot directly, with each table created from the catalog's own schema and filled withINSERT … SELECTwith the crop NULL, in one transaction. ~1.2 s → ~0.4 s CPU per pass on the reference catalog. Upload content and size are unchanged (~50 MB), and the file is identical in schema, rows, pragmas and header, so 0.16.0 peers merge it as before. Embeddings are now most of the upload; the merge will start reading them in #77.