Ask a synced keyword's tombstone once per merge, not once per assignment

`merge_remote_catalog` on the reference library (catalog_bench, a copy
merged with itself: the steady state of a sync pass) cost 228-231 ms best
of 5. Timing its phases put 91 ms in the keyword half, not in the faces the
issue named.

Both assignment unions refuse a word this device holds only as a tombstone,
with a correlated `NOT EXISTS (... deleted = 1) OR EXISTS (... deleted = 0)`
per incoming assignment. The `deleted = 1` half has no index to use --
`keyword_terms_name` is partial on `deleted = 0` -- so it scanned the whole
vocabulary for each of the 10,800 rows:

  SCAN rk
  CORRELATED SCALAR SUBQUERY 1
    SCAN t
  CORRELATED SCALAR SUBQUERY 2
    SEARCH t USING COVERING INDEX keyword_terms_name (name=?)

The refused words are one set for the whole statement, so it is asked once:
`rk.keyword NOT IN (tombstoned names EXCEPT live names)`, which is the same
condition -- refused exactly when deleted under some identity and live under
none -- and which SQLite builds as a list before the walk:

  SCAN rk
  LIST SUBQUERY 2
    MERGE (EXCEPT) ...

The file-id union alone went from 72 ms to 11 ms (sqlite3 on a copy), and
the keyword phase of the merge from 91 ms to 28-35 ms. Every table of the
catalog checksums the same after the bench as after the old build's run,
and the merge tests for tombstones and renames pass unchanged.
This commit is contained in:
2026-09-26 13:28:50 -04:00
parent 87badb6f99
commit 981022ab1d
+15 -8
View File
@@ -725,6 +725,15 @@ fn merge_keywords_within(tx: &Connection, report: &mut MergeReport) -> Result<()
/// version arrives through there, the default version is both where
/// [`crate::keywords::assign`] writes and where the panel reads — so it is the
/// one place the word can land and be seen.
///
/// A word is refused when this device holds it only as a tombstone: deleted
/// under some identity and live under none. That is a set of words, the same
/// for every row, so it is asked once -- a list SQLite builds before the walk
/// -- rather than as a correlated `NOT EXISTS ... OR EXISTS` per incoming
/// assignment. The `deleted = 1` half of that could use no index
/// (`keyword_terms_name` holds only live rows) and scanned the whole
/// vocabulary for each of 10,800 assignments: 70 ms of a sync pass that
/// changed nothing, on the reference library, against 11 ms now.
const ASSIGN_BY_FILE_ID: &str = "
INSERT OR IGNORE INTO main.keywords(version_id, keyword)
SELECT lv.id, rk.keyword
@@ -733,10 +742,9 @@ const ASSIGN_BY_FILE_ID: &str = "
JOIN remote_cat.remote rr ON rr.image_id = rv.image_id
JOIN main.remote lr ON lr.file_id = rr.file_id
JOIN main.versions lv ON lv.image_id = lr.image_id AND lv.is_default = 1
WHERE NOT EXISTS (SELECT 1 FROM main.keyword_terms t
WHERE t.name = rk.keyword AND t.deleted = 1)
OR EXISTS (SELECT 1 FROM main.keyword_terms t
WHERE t.name = rk.keyword AND t.deleted = 0)";
WHERE rk.keyword NOT IN (SELECT name FROM main.keyword_terms WHERE deleted = 1
EXCEPT
SELECT name FROM main.keyword_terms WHERE deleted = 0)";
/// The same union for a library with no server behind it.
///
@@ -753,10 +761,9 @@ const ASSIGN_BY_CONTENT_HASH: &str = "
JOIN main.images li ON li.content_hash = ri.content_hash
JOIN main.versions lv ON lv.image_id = li.id AND lv.is_default = 1
WHERE ri.content_hash IS NOT NULL
AND (NOT EXISTS (SELECT 1 FROM main.keyword_terms t
WHERE t.name = rk.keyword AND t.deleted = 1)
OR EXISTS (SELECT 1 FROM main.keyword_terms t
WHERE t.name = rk.keyword AND t.deleted = 0))";
AND rk.keyword NOT IN (SELECT name FROM main.keyword_terms WHERE deleted = 1
EXCEPT
SELECT name FROM main.keyword_terms WHERE deleted = 0)";
/// Whether an attached database holds a table of this name.
///