Count the rating chips from the rated versions, not from every image

`rating_histogram` runs on every star keystroke. On the reference library
(24k images, 1,200 of them rated) it cost 6.3 ms best-of-50 by
catalog_bench, and up to 10-14 ms when the machine is busy.

It was `images LEFT JOIN versions ON ... AND is_default = 1 GROUP BY
rating`. The plan:

  SCAN i USING COVERING INDEX images_folder
  SEARCH v USING COVERING INDEX versions_judgement (image_id=?) LEFT-JOIN
  USE TEMP B-TREE FOR GROUP BY

A probe of the index per image, then a sort of all 23,500 rows, to put
22,000 of them in slot zero.

Now the rated rows are grouped on their own (`rating != 0`: one pass over
`versions_judgement`, a sort of 1,200 rows), and slot zero is what is left
of the join's row count. That count is three index-only aggregates -- the
library size, the default versions, and the images holding one -- so an
image with no version is still unrated, and an image with two default
versions still counts twice, exactly as the join counted it:

  SCAN versions USING COVERING INDEX versions_judgement      (x3)
  SCAN images USING COVERING INDEX images_folder

`count(DISTINCT image_id)` has its own statement because alone it reads the
distinct values off the index order; beside other aggregates SQLite builds a
temporary b-tree for it.

After: 1.2 ms. The histogram is the same on the reference library
([22364, 663, 19, 47, 115, 374]), and a new test compares it with the old
join on a catalog holding every state the schema allows: no version, only a
virtual copy, two defaults, ratings below zero and above five.
This commit is contained in:
2026-09-26 13:28:50 -04:00
parent 408f189019
commit 81118728f4
2 changed files with 127 additions and 12 deletions
+123 -8
View File
@@ -582,25 +582,61 @@ pub fn judgements(
pub fn rating_histogram(conn: &Connection) -> Result<[usize; 6], CatalogError> {
let mut out = [0usize; 6];
// LEFT JOIN, so an image whose version row is missing still counts as
// unrated rather than vanishing from the totals. The histogram has to sum
// to the library size or it is not believable.
let mut stmt = conn.prepare(
"SELECT coalesce(v.rating, 0) AS r, count(*)
FROM images i
LEFT JOIN versions v ON v.image_id = i.id AND v.is_default = 1
GROUP BY r",
// Only the rated rows are grouped; the unrated slot is what is left of
// [`judged_rows`]. So an image whose version row is missing still counts
// as unrated rather than vanishing from the totals -- the histogram has
// to sum to the library size or it is not believable.
//
// It was one `images LEFT JOIN versions ... GROUP BY` until 2026-09-26:
// a probe of `versions_judgement` per image and a sort of every row, to
// put 22,000 of 23,500 in slot zero. 9 ms on every star keystroke on the
// reference library; this is a pass over the index that sorts only the
// rated few, and [`judged_rows`] is three index-only counts.
let mut stmt = conn.prepare_cached(
"SELECT rating, count(*) FROM versions
WHERE is_default = 1 AND rating != 0
GROUP BY rating",
)?;
let rows = stmt.query_map([], |r| Ok((r.get::<_, i64>(0)?, r.get::<_, i64>(1)?)))?;
let mut counted = 0usize;
for (rating, count) in rows.flatten() {
if let Some(slot) = out.get_mut(rating.clamp(0, MAX_RATING as i64) as usize) {
*slot += count as usize;
counted += count as usize;
}
}
out[0] += judged_rows(conn)?.saturating_sub(counted);
Ok(out)
}
/// How many rows `images LEFT JOIN versions ON ... AND is_default = 1` has:
/// one per image with no default version, and one per default version for
/// the rest. The total both histograms divide up, and the unjudged slot is
/// what is left of it once the judged rows are counted.
///
/// Spelled as three counts rather than as that join because the join probes
/// `versions_judgement` once per image, where each count here is one pass
/// over an index without reading a row: the library size, the default
/// versions, and the images holding one. An image with two default versions
/// -- nothing prevents it -- is two rows of the join and one image of the
/// third count, so it adds one here exactly as it did there. A version
/// always belongs to an image; `foreign_keys` is on and deletes cascade.
///
/// `count(DISTINCT image_id)` alone in its statement: that is what lets
/// SQLite read the distinct values off the index's order instead of
/// building a temporary b-tree of them.
fn judged_rows(conn: &Connection) -> Result<usize, CatalogError> {
let n: i64 = conn
.prepare_cached(
"SELECT (SELECT count(*) FROM images)
+ (SELECT count(*) FROM versions WHERE is_default = 1)
- (SELECT count(DISTINCT image_id) FROM versions WHERE is_default = 1)",
)?
.query_row([], |r| r.get(0))?;
Ok(n.max(0) as usize)
}
/// How many images carry each flag: `(picks, rejects)`.
pub fn flag_counts(conn: &Connection) -> Result<(usize, usize), CatalogError> {
let picks: i64 = conn.query_row(
@@ -987,6 +1023,85 @@ mod tests {
assert_eq!(h.iter().sum::<usize>(), 4);
}
/// The rows of the join the histograms used to be spelled as, grouped the
/// way `rating_histogram` groups them. What the counts must still agree
/// with, in the states nothing in the schema prevents.
fn by_join(cat: &Catalog, column: &str) -> Vec<(i64, i64)> {
cat.connection()
.prepare(&format!(
"SELECT coalesce(v.{column}, 0) AS c, count(*)
FROM images i
LEFT JOIN versions v ON v.image_id = i.id AND v.is_default = 1
GROUP BY c ORDER BY c"
))
.unwrap()
.query_map([], |r| Ok((r.get(0)?, r.get(1)?)))
.unwrap()
.map(Result::unwrap)
.collect()
}
/// A library in every awkward state at once: an image with no version,
/// one with only a virtual copy, one with two default versions, and
/// values out of range on both axes.
fn awkward() -> Catalog {
let cat = with_images(8);
ensure_default_versions(cat.connection()).unwrap();
let all = ids(&cat);
let c = cat.connection();
set_rating(c, all[0], 5).unwrap();
set_rating(c, all[1], 2).unwrap();
set_label(c, all[1], Some(ColourLabel::Blue)).unwrap();
c.execute(
"DELETE FROM versions WHERE image_id = ?1",
[all[2].0 as i64],
)
.unwrap();
c.execute(
"UPDATE versions SET is_default = 0 WHERE image_id = ?1",
[all[3].0 as i64],
)
.unwrap();
c.execute(
"INSERT INTO versions(image_id, uuid, name, is_default, rating, label)
VALUES (?1, 'second-default', 'Copy', 1, 4, 3)",
[all[4].0 as i64],
)
.unwrap();
c.execute(
"UPDATE versions SET rating = -1, label = 9 WHERE image_id = ?1",
[all[5].0 as i64],
)
.unwrap();
c.execute(
"UPDATE versions SET rating = 7, label = 0 WHERE image_id = ?1",
[all[6].0 as i64],
)
.unwrap();
cat
}
/// Fold the join's rows into slots the way the old code did.
fn folded(rows: &[(i64, i64)], slot: impl Fn(i64) -> usize) -> [usize; 6] {
let mut out = [0usize; 6];
for &(code, n) in rows {
out[slot(code)] += n as usize;
}
out
}
#[test]
fn the_rating_histogram_agrees_with_the_join_it_replaced() {
let cat = awkward();
let expected = folded(&by_join(&cat, "rating"), |r| {
r.clamp(0, MAX_RATING as i64) as usize
});
assert_eq!(rating_histogram(cat.connection()).unwrap(), expected);
// Nine rows for eight images: the doubled default counts twice, as
// it always has.
assert_eq!(expected.iter().sum::<usize>(), 9);
}
#[test]
fn flag_counts_separate_picks_from_rejects() {
let cat = with_images(5);
File diff suppressed because one or more lines are too long