Record the catalog lessons of the 2026-09-25 performance pass

Three habits the pass found broken, in the style of the 2026-09-19 notes:
SQL text passed to `execute`/`query_row` in a loop is a prepare per row
(the merge, `persist` and the shard sync all paid it), a case-insensitive
`LIKE` cannot use the `source_ref` key where a range can, and the backfill
inside `Catalog::open` is paid by every worker thread, the develop view's
fetches included. And where the two new benches are and how to run them.
This commit is contained in:
2026-09-25 22:06:58 -04:00
parent 4583048595
commit ce72fe49a0
+33
View File
@@ -31,6 +31,30 @@ over a result set — including `deep_count` per sidebar row, which is fine
at sidebar scale and would not be at grid scale. Aggregate in one
statement and look up in memory.
**SQL text in a loop is a prepare in a loop.** rusqlite's `execute` and
`query_row` compile their statement on every call. A loop that calls them
per row pays a prepare per row even when each query is a primary-key seek:
the merge of a synced catalog prepared four statements for each of 13,000
incoming faces (450 ms of a pass that changed nothing), `persist` four per
photograph a scan listed (1.5 s for a first scan), the shard sync one per
image each way. Hoist the statement, use `prepare_cached`, or — better, when
the loop asks the same table about every row — read that table once into a
map. And do not rewrite a row with what it already holds: an upsert of
identical values still dirties the page.
**A `LIKE` is case-insensitive, and no index here serves that.**
`source_ref LIKE 'stem.%'` read every name of the root per sidecar a pull
took in. When the check that decides is exact, spell the prefix as a range
(`>= 'stem.' AND < 'stem/'`, `/` being the byte after `.`), which the
`(root_id, source_ref)` key answers with a seek.
**`Catalog::open` is not free, and every worker thread calls it.** The
backfill runs on every open, and the develop view opens a catalog to fetch
each original and again for each neighbour it prefetches. Keep each
backfill step's no-op case to a read of the small side — the unpaired
JPEGs, not every RAW; the distinct keywords, not every assignment — and
measure an open with `catalog_bench` after adding one.
**Filter and aggregate in SQL, and aggregate the small side first.**
`faces::people` read 19,000 rows, grouped, sorted them by name, and the
screen threw 17,000 away (empty unnamed groups). `people_in_use` filters in
@@ -124,6 +148,15 @@ other builds are running on the machine — the wall clock doubles under
load, the CPU figure does not. Keep the binary from before the change and
run both back to back rather than trusting numbers taken an hour apart.
`cargo run --release -p dr-catalog --example catalog_bench -- CATALOG
[FACES_DIR]` does the same for opening the catalog (the backfill step by
step), the upload snapshot, a merge, and the face shard export and import;
`persist_bench`, an ignored test in `dr-ui`'s scan module, replays a scan's
`persist` and a sidecar pull (`DR_BENCH_CATALOG=copy.sqlite cargo test
--release -p dr-ui --lib persist_bench -- --ignored --nocapture
--test-threads=1`). Both take copies; hand `catalog_bench` a copy of the
face store directory too.
Reference figures from the 2026-09-19 fixes, largest person (754 faces),
24k images, 19k faces, before → after. What one click read: `load_people`
22 ms → 12 ms, `load_faces` 316 ms → 2.4 ms, `audit` 190 ms → not run