Serve the keyword counts from an index that carries the version

`keywords::list` is the vocabulary with a per-word photograph count, and
`keywords::for_images` calls it on every selection change to redraw the
keyword panel. On the reference library (58 words, 10,800 assignments) it
cost 3.0-3.5 ms best-of-50 by catalog_bench, `for_images` 3.1-3.6 ms (6 and
5 ms on a busy machine).

Per word, the count walks `keywords_term (keyword)` and, for each
assignment, reads the `keywords` row to learn its version before probing
`versions` for the image:

  SEARCH k USING INDEX keywords_term (keyword=?)
  SEARCH v USING INTEGER PRIMARY KEY (rowid=?)

With `keywords_term_version (keyword, version_id)` the first step is
index-only:

  SEARCH k USING COVERING INDEX keywords_term_version (keyword=?)
  SEARCH v USING INTEGER PRIMARY KEY (rowid=?)

After: `list` 1.3 ms, `for_images` 1.5 ms, with the same answers (digests of
both outputs compared on the reference library).

The index is created on first use by `list`, with CREATE INDEX IF NOT EXISTS,
not by a migration: a new schema version makes every older build refuse this
catalog's snapshot at sync (`sync::remote_is_mergeable` compares
`user_version` and nothing else), and a build that meets an extra index
ignores it. Once the index exists the statement is a schema lookup, 8 us. A
failure to create it -- a read-only or busy catalog -- is logged and the
list is read without it, as before.
This commit is contained in:
2026-09-26 13:28:50 -04:00
parent fe6e523443
commit d537e4a965
+23 -1
View File
@@ -244,7 +244,8 @@ pub fn delete(conn: &Connection, id: KeywordId) -> Result<usize, CatalogError> {
/// every assignment, and a query per keyword would be one statement per word /// every assignment, and a query per keyword would be one statement per word
/// in the library. /// in the library.
pub fn list(conn: &Connection) -> Result<Vec<Keyword>, CatalogError> { pub fn list(conn: &Connection) -> Result<Vec<Keyword>, CatalogError> {
let mut stmt = conn.prepare( ensure_term_index(conn);
let mut stmt = conn.prepare_cached(
// DISTINCT image, not row: a word on two versions of one frame is one // DISTINCT image, not row: a word on two versions of one frame is one
// photograph, and reporting two is the kind of small lie that makes a // photograph, and reporting two is the kind of small lie that makes a
// user stop trusting the counts. // user stop trusting the counts.
@@ -269,6 +270,27 @@ pub fn list(conn: &Connection) -> Result<Vec<Keyword>, CatalogError> {
Ok(rows) Ok(rows)
} }
/// The index [`list`]'s per-word count is served from: `keywords_term`
/// with the version beside the word, so the count reads no `keywords` row.
///
/// `keywords_term` alone gave the row id, and each of the 10,800 assignments
/// on the reference library cost a probe of the table for its version --
/// 4 ms of the keyword panel's redraw, on every selection change.
///
/// Created on first use rather than by a migration, for the reason
/// `duplicates::ensure_probe_table` gives: a new schema version makes every
/// older build refuse this catalog's snapshot at sync, and an older build
/// that meets an extra index ignores it. Once it exists, the statement is a
/// lookup in the schema (microseconds). A failure to create it is logged and
/// the list read without it: the index is a speed-up, never an answer.
fn ensure_term_index(conn: &Connection) {
if let Err(e) = conn.execute_batch(
"CREATE INDEX IF NOT EXISTS keywords_term_version ON keywords(keyword, version_id);",
) {
log::warn!("keywords: could not create keywords_term_version: {e}");
}
}
/// Assign a keyword to images, creating the keyword if it is new. /// Assign a keyword to images, creating the keyword if it is new.
/// ///
/// The bulk form is the *only* form, because keywording a selection is the /// The bulk form is the *only* form, because keywording a selection is the