Count the rating chips from the rated versions, not from every image
`rating_histogram` runs on every star keystroke. On the reference library (24k images, 1,200 of them rated) it cost 6.3 ms best-of-50 by catalog_bench, and up to 10-14 ms when the machine is busy. It was `images LEFT JOIN versions ON ... AND is_default = 1 GROUP BY rating`. The plan: SCAN i USING COVERING INDEX images_folder SEARCH v USING COVERING INDEX versions_judgement (image_id=?) LEFT-JOIN USE TEMP B-TREE FOR GROUP BY A probe of the index per image, then a sort of all 23,500 rows, to put 22,000 of them in slot zero. Now the rated rows are grouped on their own (`rating != 0`: one pass over `versions_judgement`, a sort of 1,200 rows), and slot zero is what is left of the join's row count. That count is three index-only aggregates -- the library size, the default versions, and the images holding one -- so an image with no version is still unrated, and an image with two default versions still counts twice, exactly as the join counted it: SCAN versions USING COVERING INDEX versions_judgement (x3) SCAN images USING COVERING INDEX images_folder `count(DISTINCT image_id)` has its own statement because alone it reads the distinct values off the index order; beside other aggregates SQLite builds a temporary b-tree for it. After: 1.2 ms. The histogram is the same on the reference library ([22364, 663, 19, 47, 115, 374]), and a new test compares it with the old join on a catalog holding every state the schema allows: no version, only a virtual copy, two defaults, ratings below zero and above five.
This commit is contained in:
File diff suppressed because one or more lines are too long
Reference in New Issue
Block a user