Count the label chips from the labelled versions, not from every image

`label_histogram` runs on every label keystroke and after every batch of
judgements is saved. On the reference library it cost 7.2-8.4 ms best-of-50
by catalog_bench (13 ms on a busy machine), to report that none of 23,500
images carried a label.

The join was the rating histogram's, with one thing worse: the index does
not carry `label`, so each probe went on to read the version's row.

  SCAN i USING COVERING INDEX images_folder
  SEARCH v USING INDEX versions_judgement (image_id=?) LEFT-JOIN
  USE TEMP B-TREE FOR GROUP BY

It now takes the rating histogram's shape: only labelled default versions
are grouped, and the unlabelled slot is what is left of `judged_rows`.

  SCAN versions USING INDEX versions_judgement
  USE TEMP B-TREE FOR GROUP BY          (the labelled rows only)

That pass still reads each default version's row for `label`, but in the
index's order, which follows the table's; a partial index on the labelled
rows would make it index-only, and was not worth a new index for the
remaining 1 ms. After: 1.7-2.0 ms, the same answer on the reference library,
and a test that compares it with the old join over the awkward states the
rating test uses (a second default's label counted, unknown codes and zero
folded into unlabelled).
This commit is contained in:
2026-09-26 13:28:50 -04:00
parent 81118728f4
commit 73059f2656
2 changed files with 37 additions and 13 deletions
+33 -9
View File
@@ -449,19 +449,26 @@ pub fn toggled_label(
/// TRACES: FR-CAT-5 | FR-CAT-6
/// How the library divides by colour label, for the filter chips' counts.
///
/// Index 0 is unlabelled and index `n` the label whose code is `n`. One
/// grouped statement — the same shape as [`rating_histogram`], and for the
/// same reason it LEFT JOINs: an image without a version row is unlabelled,
/// not missing.
/// Index 0 is unlabelled and index `n` the label whose code is `n`. The
/// same shape as [`rating_histogram`], and for the same reason the
/// unlabelled slot is what is left of [`judged_rows`]: an image without a
/// version row is unlabelled, not missing.
///
/// Only labelled rows are grouped. The join this replaced (2026-09-26)
/// probed `versions_judgement` per image and then read each version's row
/// for `label`, which the index does not carry -- 10 ms on the reference
/// library, on every label keystroke, to find that none of 23,500 images
/// had one. This walks the default versions in the index's order, which
/// is close to the table's, and groups the few that are labelled.
pub fn label_histogram(conn: &Connection) -> Result<[usize; 6], CatalogError> {
let mut out = [0usize; 6];
let mut stmt = conn.prepare(
"SELECT coalesce(v.label, 0) AS l, count(*)
FROM images i
LEFT JOIN versions v ON v.image_id = i.id AND v.is_default = 1
GROUP BY l",
let mut stmt = conn.prepare_cached(
"SELECT label, count(*) FROM versions
WHERE is_default = 1 AND label IS NOT NULL
GROUP BY label",
)?;
let rows = stmt.query_map([], |r| Ok((r.get::<_, i64>(0)?, r.get::<_, i64>(1)?)))?;
let mut counted = 0usize;
for (code, count) in rows.flatten() {
// A code this build does not know counts as unlabelled, which is how
// `label_from_code` reads it everywhere else.
@@ -471,7 +478,9 @@ pub fn label_histogram(conn: &Connection) -> Result<[usize; 6], CatalogError> {
0
};
out[slot] += count as usize;
counted += count as usize;
}
out[0] += judged_rows(conn)?.saturating_sub(counted);
Ok(out)
}
@@ -1102,6 +1111,21 @@ mod tests {
assert_eq!(expected.iter().sum::<usize>(), 9);
}
#[test]
fn the_label_histogram_agrees_with_the_join_it_replaced() {
let cat = awkward();
let expected = folded(&by_join(&cat, "label"), |code| {
if label_from_code(Some(code)).is_some() {
code as usize
} else {
0
}
});
assert_eq!(label_histogram(cat.connection()).unwrap(), expected);
assert_eq!(expected[3], 1, "the second default's label is counted");
assert_eq!(expected.iter().sum::<usize>(), 9);
}
#[test]
fn flag_counts_separate_picks_from_rejects() {
let cat = with_images(5);
File diff suppressed because one or more lines are too long