Cover the eyes-open subquery with an index

The people filter was served from faces_image without touching a row;
reading the eye columns in the same subquery touched every one, and
ALTER TABLE had put those seven floats after the embedding and the crop
blob. One count took 24 seconds on the reference library, thirteen of
them system time. faces_eyes covers the subquery again: five
milliseconds.
This commit is contained in:
2026-09-19 14:05:52 +02:00
parent cd0ca6785f
commit d706c12d77
4 changed files with 71 additions and 24 deletions
+37 -1
View File
@@ -15,7 +15,7 @@ use rusqlite::Connection;
use crate::error::CatalogError;
/// Schema version this build writes and understands.
pub const SCHEMA_VERSION: i64 = 16;
pub const SCHEMA_VERSION: i64 = 17;
/// Apply migrations up to [`SCHEMA_VERSION`].
///
@@ -160,6 +160,13 @@ pub fn migrate(conn: &Connection) -> Result<i64, CatalogError> {
tx.commit()?;
}
if from < 17 {
let tx = conn.unchecked_transaction()?;
tx.execute_batch(V17)?;
tx.pragma_update(None, "user_version", 17)?;
tx.commit()?;
}
Ok(from)
}
@@ -744,6 +751,29 @@ CREATE TABLE IF NOT EXISTS xmp_conflicts (
);
"#;
const V17: &str = r#"
-- TRACES: FR-CULL-13 | NFR-P9
-- The eyes-open filter's index, and a lesson about where a column lands.
--
-- The people filter is a correlated EXISTS over `faces` per image, and it
-- was fast because `faces_image` *covers* it: the subquery never touched a
-- row. Reading V16's seven eye columns in the same subquery did touch the
-- row -- and `ALTER TABLE ADD COLUMN` puts a column at the end of the
-- record, after the 1 KB embedding and the ~5 KB crop, so every check
-- dragged six kilobytes off disk to reach seven floats. Measured on the
-- reference library: 24 seconds for one count, thirteen of them system
-- time. With this index the same count takes five milliseconds, because
-- the subquery is served from the index again and never reads a row.
--
-- The columns are listed in EYE_COLUMNS' order behind `image_id`, which is
-- the key the subquery searches on. Nothing else changed in V17; a catalog
-- already at V16 needs only this.
CREATE INDEX IF NOT EXISTS faces_eyes ON faces(
image_id, eye_right, eye_right_px, eye_right_sharp,
eye_left, eye_left_px, eye_left_sharp, sunglasses
);
"#;
// V16 -- TRACES: FR-CULL-13
//
// What each face's eyes are doing: for each eye P(open), the source pixels
@@ -1601,6 +1631,12 @@ mod tests {
}
c.pragma_update(None, "user_version", 15).unwrap();
assert_eq!(migrate(&c).unwrap(), 15);
let indexed: bool = c
.prepare("SELECT 1 FROM sqlite_master WHERE type = 'index' AND name = 'faces_eyes'")
.unwrap()
.exists([])
.unwrap();
assert!(indexed, "V17's covering index is there");
let v: i64 = c
.query_row("PRAGMA user_version", [], |r| r.get(0))
.unwrap();
+5
View File
@@ -767,6 +767,11 @@ CREATE TABLE faces (
detected_at INTEGER NOT NULL
);
CREATE INDEX faces_image ON faces(image_id);
-- Covers the eyes-open filter's subquery. Without it every check read the
-- whole face row -- the eye columns sit after the blobs -- and one count
-- took 24 s on the reference library (schema V17).
CREATE INDEX faces_eyes ON faces(image_id, eye_right, eye_right_px, eye_right_sharp,
eye_left, eye_left_px, eye_left_sharp, sunglasses);
CREATE TABLE face_person (
face_id INTEGER PRIMARY KEY REFERENCES faces(id) ON DELETE CASCADE,
+6
View File
@@ -1403,6 +1403,12 @@ the honest answer is "everything". A test drives the same five readings through
through `state()` and requires the two to agree, so the badge and the grid cannot say different
things.
**And an index, learned the slow way.** The people filter was always served from a covering index
on `faces(image_id)`; the moment its subquery read the eye columns it had to read the face *row*,
and `ALTER TABLE ADD COLUMN` had put those seven floats after the embedding and the crop blob —
six kilobytes to reach every one. One count took 24 seconds on the reference library, thirteen of
them system time. `faces_eyes` (schema V17) covers the subquery again: five milliseconds.
### 17.4 What the sample says about accuracy, and what it does not
The 60-proxy sample above was run at proxy resolution, where the production pass reads the native
+23 -23
View File
File diff suppressed because one or more lines are too long