Count the label chips from the labelled versions, not from every image
`label_histogram` runs on every label keystroke and after every batch of judgements is saved. On the reference library it cost 7.2-8.4 ms best-of-50 by catalog_bench (13 ms on a busy machine), to report that none of 23,500 images carried a label. The join was the rating histogram's, with one thing worse: the index does not carry `label`, so each probe went on to read the version's row. SCAN i USING COVERING INDEX images_folder SEARCH v USING INDEX versions_judgement (image_id=?) LEFT-JOIN USE TEMP B-TREE FOR GROUP BY It now takes the rating histogram's shape: only labelled default versions are grouped, and the unlabelled slot is what is left of `judged_rows`. SCAN versions USING INDEX versions_judgement USE TEMP B-TREE FOR GROUP BY (the labelled rows only) That pass still reads each default version's row for `label`, but in the index's order, which follows the table's; a partial index on the labelled rows would make it index-only, and was not worth a new index for the remaining 1 ms. After: 1.7-2.0 ms, the same answer on the reference library, and a test that compares it with the old join over the awkward states the rating test uses (a second default's label counted, unknown codes and zero folded into unlabelled).
This commit is contained in:
File diff suppressed because one or more lines are too long
Reference in New Issue
Block a user