Fuse every detector's faces into one population per embedder
Choosing "Thorough" made the library look empty. The detector setting writes under its own faces.model_id, and every reader of "the faces" keyed on that exact id: the clustering pass, the coverage figure, the sweep's work list, the shard export and import, and the sync merge's face matching. On the reference library that restarted coverage at 1,834 of 19,140, drew a People rail of 36 faces for a person with 520, queued a ~400 GB re-fetch on each device, and stranded the desktop's 3,583 confirmations under the old id: the tablet held the same faces under the new one and the merge refused to match them. Same photograph, same box, same embedder, two ids — that is one face, not two libraries. The embedder half of the id is now the key. embedder_of and embedder_sql give it to every query; writes keep the full id, so which detector drew a box stays on record. record_detections is unchanged and is where the generations meet: an image holds one pipeline's faces at a time, and a re-detection carries confirmations across by box overlap. The merge's match_faces applies the same rule within an embedder. The calibration is keyed on the embedder too, since the similarity space did not change. Shards travel every generation, each under its own id, and a peer adopts whichever it is sent — including a stronger detector's pass over an image it indexed itself with a weaker one, which is the re-detection its own sweep would otherwise queue, already done. Never downwards: a tablet on Fast keeps the desktop's Thorough faces. The sweep gains the same tail — images a weaker detector indexed, after the ones nothing has — driven by FaceDetector::supersedes, so choosing a stronger detector still improves the library over time without first making it disappear.
This commit is contained in:
+186
-38
@@ -26,6 +26,25 @@
|
||||
//! disposable index; the name is irreplaceable and goes to the sidecar
|
||||
//! (FR-CULL-12), which is not this module's job.
|
||||
//!
|
||||
//! # One population per embedder, not one per detector
|
||||
//!
|
||||
//! `faces.model_id` names a pipeline, `detector+embedder`, and a detector
|
||||
//! change writes new rows under a new id. What it must *not* do is split the
|
||||
//! library in two: the People screen, the clustering pass, the sync merge and
|
||||
//! the shard import all read "the faces" — and the faces are every row whose
|
||||
//! embedder half matches, because the embedder is what makes two vectors
|
||||
//! comparable and the detector only decides where the boxes are. Before this
|
||||
//! rule each of those read the exact id, and choosing a better detector
|
||||
//! emptied every screen until a 400 GB re-index had run on every device,
|
||||
//! while a name confirmed on one device could not reach the other because the
|
||||
//! two held the same face under different ids.
|
||||
//!
|
||||
//! So reads key on [`embedder_of`] the id, via [`embedder_sql`]; writes keep
|
||||
//! the exact id, so which detector drew a box stays recorded; and
|
||||
//! [`record_detections`] is where the generations meet — an image holds one
|
||||
//! pipeline's faces at a time, and a re-detection carries the user's
|
||||
//! confirmations onto the boxes that replace them.
|
||||
//!
|
||||
//! # Privacy is structural here, not policy
|
||||
//!
|
||||
//! NFR-SEC-5 puts face data under a stricter rule than the rest of the catalog:
|
||||
@@ -40,6 +59,30 @@ use dr_types::ImageId;
|
||||
|
||||
use crate::error::CatalogError;
|
||||
|
||||
/// The embedder half of a model id: what makes two faces comparable.
|
||||
///
|
||||
/// `scrfd_10g+w600k_mbf` and `w600k_mbf` are the same embedder behind two
|
||||
/// detectors, and their vectors live in one space. A bare id is its own
|
||||
/// embedder — the first pipeline was written without a detector prefix, and
|
||||
/// every library indexed before the choice existed is under that spelling.
|
||||
pub fn embedder_of(model_id: &str) -> &str {
|
||||
model_id.rsplit('+').next().unwrap_or(model_id)
|
||||
}
|
||||
|
||||
/// The SQL for [`embedder_of`] over a column, for a `WHERE` that means "the
|
||||
/// same population as this pipeline" rather than "this exact pipeline".
|
||||
///
|
||||
/// Callers pass `embedder_of(model_id)` as the bound value. A `CASE` rather
|
||||
/// than `LIKE`, so the bare and the qualified spelling compare as the same
|
||||
/// thing without a wildcard that `w600k_mbf_v2` would also match.
|
||||
pub fn embedder_sql(column: &str) -> String {
|
||||
format!(
|
||||
"CASE WHEN instr({column}, '+') > 0
|
||||
THEN substr({column}, instr({column}, '+') + 1)
|
||||
ELSE {column} END"
|
||||
)
|
||||
}
|
||||
|
||||
/// A face's row id. Local to this catalog, like [`crate::keywords::KeywordId`].
|
||||
#[derive(Debug, Clone, Copy, PartialEq, Eq, PartialOrd, Ord, Hash)]
|
||||
pub struct FaceId(pub u64);
|
||||
@@ -330,8 +373,11 @@ pub fn record_measurements(
|
||||
tx.execute("DELETE FROM faces WHERE id = ?1", [f.0 as i64])?;
|
||||
}
|
||||
let remaining: i64 = tx.query_row(
|
||||
"SELECT COUNT(*) FROM faces WHERE image_id = ?1 AND model_id = ?2",
|
||||
rusqlite::params![image_id.0 as i64, model_id],
|
||||
&format!(
|
||||
"SELECT COUNT(*) FROM faces WHERE image_id = ?1 AND {} = ?2",
|
||||
embedder_sql("model_id")
|
||||
),
|
||||
rusqlite::params![image_id.0 as i64, embedder_of(model_id)],
|
||||
|r| r.get(0),
|
||||
)?;
|
||||
tx.execute(
|
||||
@@ -365,7 +411,7 @@ pub fn unmeasured_on_image(
|
||||
) -> Result<Vec<Face>, CatalogError> {
|
||||
Ok(for_image(conn, image_id)?
|
||||
.into_iter()
|
||||
.filter(|f| f.model_id == model_id && f.quality.is_none())
|
||||
.filter(|f| embedder_of(&f.model_id) == embedder_of(model_id) && f.quality.is_none())
|
||||
.collect())
|
||||
}
|
||||
|
||||
@@ -374,8 +420,11 @@ pub fn unmeasured_on_image(
|
||||
/// What the measuring pass has left to do, for a screen that wants to say so.
|
||||
pub fn faces_unmeasured(conn: &Connection, model_id: &str) -> Result<u64, CatalogError> {
|
||||
conn.query_row(
|
||||
"SELECT COUNT(*) FROM faces WHERE model_id = ?1 AND quality IS NULL",
|
||||
[model_id],
|
||||
&format!(
|
||||
"SELECT COUNT(*) FROM faces WHERE {} = ?1 AND quality IS NULL",
|
||||
embedder_sql("model_id")
|
||||
),
|
||||
[embedder_of(model_id)],
|
||||
|r| r.get::<_, i64>(0),
|
||||
)
|
||||
.map(|n| n as u64)
|
||||
@@ -443,13 +492,19 @@ pub fn coverage(conn: &Connection, model_id: &str) -> Result<Coverage, CatalogEr
|
||||
[],
|
||||
|r| r.get(0),
|
||||
)?;
|
||||
// One row per image, whichever compatible pipeline wrote it: an image
|
||||
// holds one pipeline's faces at a time (`record_detections`), so the
|
||||
// markers of one embedder never double-count a photograph.
|
||||
let (indexed, without, faces): (i64, i64, i64) = conn.query_row(
|
||||
"SELECT COUNT(*), COALESCE(SUM(faces_found = 0), 0), COALESCE(SUM(faces_found), 0)
|
||||
FROM face_index fi
|
||||
JOIN images i ON i.id = fi.image_id
|
||||
WHERE fi.model_id = ?1
|
||||
AND i.trashed_at IS NULL AND i.shadowed_by IS NULL",
|
||||
[model_id],
|
||||
&format!(
|
||||
"SELECT COUNT(*), COALESCE(SUM(faces_found = 0), 0), COALESCE(SUM(faces_found), 0)
|
||||
FROM face_index fi
|
||||
JOIN images i ON i.id = fi.image_id
|
||||
WHERE {} = ?1
|
||||
AND i.trashed_at IS NULL AND i.shadowed_by IS NULL",
|
||||
embedder_sql("fi.model_id")
|
||||
),
|
||||
[embedder_of(model_id)],
|
||||
|r| Ok((r.get(0)?, r.get(1)?, r.get(2)?)),
|
||||
)?;
|
||||
|
||||
@@ -461,15 +516,19 @@ pub fn coverage(conn: &Connection, model_id: &str) -> Result<Coverage, CatalogEr
|
||||
})
|
||||
}
|
||||
|
||||
/// Whether one image has been through this model.
|
||||
/// Whether one image has been through this model, or one it shares an
|
||||
/// embedder with.
|
||||
pub fn is_indexed(
|
||||
conn: &Connection,
|
||||
image_id: ImageId,
|
||||
model_id: &str,
|
||||
) -> Result<bool, CatalogError> {
|
||||
conn.query_row(
|
||||
"SELECT EXISTS(SELECT 1 FROM face_index WHERE image_id = ?1 AND model_id = ?2)",
|
||||
rusqlite::params![image_id.0 as i64, model_id],
|
||||
&format!(
|
||||
"SELECT EXISTS(SELECT 1 FROM face_index WHERE image_id = ?1 AND {} = ?2)",
|
||||
embedder_sql("model_id")
|
||||
),
|
||||
rusqlite::params![image_id.0 as i64, embedder_of(model_id)],
|
||||
|r| r.get(0),
|
||||
)
|
||||
.map_err(Into::into)
|
||||
@@ -486,8 +545,11 @@ pub fn clear_index_marker(
|
||||
model_id: &str,
|
||||
) -> Result<(), CatalogError> {
|
||||
conn.execute(
|
||||
"DELETE FROM face_index WHERE image_id = ?1 AND model_id = ?2",
|
||||
rusqlite::params![image_id.0 as i64, model_id],
|
||||
&format!(
|
||||
"DELETE FROM face_index WHERE image_id = ?1 AND {} = ?2",
|
||||
embedder_sql("model_id")
|
||||
),
|
||||
rusqlite::params![image_id.0 as i64, embedder_of(model_id)],
|
||||
)?;
|
||||
Ok(())
|
||||
}
|
||||
@@ -513,12 +575,15 @@ pub fn for_image(conn: &Connection, image_id: ImageId) -> Result<Vec<Face>, Cata
|
||||
/// unassigned and still belongs in the next clustering round, just not in that
|
||||
/// person's cluster.
|
||||
pub fn unassigned(conn: &Connection, model_id: &str) -> Result<Vec<FaceId>, CatalogError> {
|
||||
let mut q = conn.prepare(
|
||||
let mut q = conn.prepare(&format!(
|
||||
"SELECT f.id FROM faces f
|
||||
LEFT JOIN face_person fp ON fp.face_id = f.id
|
||||
WHERE fp.face_id IS NULL AND f.model_id = ?1",
|
||||
)?;
|
||||
let rows = q.query_map([model_id], |r| Ok(FaceId(r.get::<_, i64>(0)? as u64)))?;
|
||||
WHERE fp.face_id IS NULL AND {} = ?1",
|
||||
embedder_sql("f.model_id")
|
||||
))?;
|
||||
let rows = q.query_map([embedder_of(model_id)], |r| {
|
||||
Ok(FaceId(r.get::<_, i64>(0)? as u64))
|
||||
})?;
|
||||
rows.collect::<Result<_, _>>().map_err(Into::into)
|
||||
}
|
||||
|
||||
@@ -539,15 +604,18 @@ pub struct StoredEmbedding {
|
||||
|
||||
/// Embeddings for clustering, oldest first so the pass is deterministic.
|
||||
///
|
||||
/// Returned as raw f16 blobs rather than decoded vectors: the caller is
|
||||
/// `dr-face`, which owns the decoding, and a catalog that widened them here
|
||||
/// would double the memory of the one operation that holds them all at once.
|
||||
/// Every face this model's embedder produced, whichever detector found it —
|
||||
/// see the module note. Returned as raw f16 blobs rather than decoded vectors:
|
||||
/// the caller is `dr-face`, which owns the decoding, and a catalog that
|
||||
/// widened them here would double the memory of the one operation that holds
|
||||
/// them all at once.
|
||||
pub fn embeddings(conn: &Connection, model_id: &str) -> Result<Vec<StoredEmbedding>, CatalogError> {
|
||||
let mut q = conn.prepare(
|
||||
let mut q = conn.prepare(&format!(
|
||||
"SELECT id, image_id, embedding, crop_px, quality FROM faces
|
||||
WHERE model_id = ?1 ORDER BY id",
|
||||
)?;
|
||||
let rows = q.query_map([model_id], |r| {
|
||||
WHERE {} = ?1 ORDER BY id",
|
||||
embedder_sql("model_id")
|
||||
))?;
|
||||
let rows = q.query_map([embedder_of(model_id)], |r| {
|
||||
Ok(StoredEmbedding {
|
||||
face: FaceId(r.get::<_, i64>(0)? as u64),
|
||||
image: ImageId(r.get::<_, i64>(1)? as u64),
|
||||
@@ -823,6 +891,11 @@ pub fn images_for_person(
|
||||
// ── calibration ───────────────────────────────────────────────────────────
|
||||
|
||||
/// Store a fitted calibration, replacing any previous fit for the model.
|
||||
///
|
||||
/// Keyed on the embedder: a calibration is a fit over pairwise similarities,
|
||||
/// and those are the same space for every detector in front of one embedder.
|
||||
/// Keying it on the full pipeline would discard the fit — and the library's
|
||||
/// sharpened confidences with it — every time the detector was changed.
|
||||
pub fn put_calibration(
|
||||
conn: &Connection,
|
||||
model_id: &str,
|
||||
@@ -842,7 +915,7 @@ pub fn put_calibration(
|
||||
face_set_hash = excluded.face_set_hash,
|
||||
fitted_at = excluded.fitted_at",
|
||||
rusqlite::params![
|
||||
model_id,
|
||||
embedder_of(model_id),
|
||||
cal.a as f64,
|
||||
cal.b as f64,
|
||||
cal.w_size as f64,
|
||||
@@ -868,7 +941,7 @@ pub fn calibration(
|
||||
conn.query_row(
|
||||
"SELECT a, b, w_size, valid, positive_pairs, negative_pairs, face_set_hash
|
||||
FROM face_calibration WHERE model_id = ?1",
|
||||
[model_id],
|
||||
[embedder_of(model_id)],
|
||||
|r| {
|
||||
Ok((
|
||||
Calibration {
|
||||
@@ -958,8 +1031,11 @@ pub fn crop(conn: &Connection, face: FaceId) -> Result<Option<Vec<u8>>, CatalogE
|
||||
/// and what a "re-index to fill these in" prompt would be counting.
|
||||
pub fn faces_without_crop(conn: &Connection, model_id: &str) -> Result<u64, CatalogError> {
|
||||
conn.query_row(
|
||||
"SELECT COUNT(*) FROM faces WHERE model_id = ?1 AND crop IS NULL",
|
||||
[model_id],
|
||||
&format!(
|
||||
"SELECT COUNT(*) FROM faces WHERE {} = ?1 AND crop IS NULL",
|
||||
embedder_sql("model_id")
|
||||
),
|
||||
[embedder_of(model_id)],
|
||||
|r| r.get::<_, i64>(0),
|
||||
)
|
||||
.map(|n| n as u64)
|
||||
@@ -1271,6 +1347,64 @@ mod tests {
|
||||
assert_eq!(for_image(&c, img).unwrap().len(), 1);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn the_embedder_is_the_half_after_the_plus() {
|
||||
assert_eq!(embedder_of("w600k_mbf"), "w600k_mbf");
|
||||
assert_eq!(embedder_of("scrfd_10g+w600k_mbf"), "w600k_mbf");
|
||||
assert_eq!(embedder_of("scrfd_2.5g+w600k_mbf"), "w600k_mbf");
|
||||
assert_eq!(embedder_of("scrfd_10g+other"), "other");
|
||||
}
|
||||
|
||||
/// A detector change must not split the library: the clustering pass,
|
||||
/// the coverage figure and the "is this indexed" question all see every
|
||||
/// generation that shares an embedder, and none of a different one.
|
||||
#[test]
|
||||
fn every_detector_in_front_of_one_embedder_is_one_population() {
|
||||
let c = db();
|
||||
let old = image(&c, 1);
|
||||
let new = image(&c, 2);
|
||||
let foreign = image(&c, 3);
|
||||
record_detections(&c, old, "w600k_mbf", 1024, &[face(1)]).unwrap();
|
||||
let mut f = face(2);
|
||||
f.model_id = "scrfd_10g+w600k_mbf".into();
|
||||
record_detections(&c, new, "scrfd_10g+w600k_mbf", 1024, &[f]).unwrap();
|
||||
let mut g = face(3);
|
||||
g.model_id = "scrfd_10g+other".into();
|
||||
record_detections(&c, foreign, "scrfd_10g+other", 1024, &[g]).unwrap();
|
||||
|
||||
for id in ["w600k_mbf", "scrfd_10g+w600k_mbf", "scrfd_2.5g+w600k_mbf"] {
|
||||
assert_eq!(embeddings(&c, id).unwrap().len(), 2, "{id}");
|
||||
assert_eq!(unassigned(&c, id).unwrap().len(), 2, "{id}");
|
||||
assert_eq!(coverage(&c, id).unwrap().indexed, 2, "{id}");
|
||||
assert!(is_indexed(&c, old, id).unwrap(), "{id}");
|
||||
assert!(is_indexed(&c, new, id).unwrap(), "{id}");
|
||||
assert!(!is_indexed(&c, foreign, id).unwrap(), "{id}");
|
||||
}
|
||||
assert_eq!(embeddings(&c, "scrfd_10g+other").unwrap().len(), 1);
|
||||
assert_eq!(coverage(&c, "scrfd_10g+other").unwrap().indexed, 1);
|
||||
}
|
||||
|
||||
/// The fit is over the embedder's similarity space, so it is the same fit
|
||||
/// whichever detector is chosen — and a detector change must not discard
|
||||
/// it.
|
||||
#[test]
|
||||
fn a_calibration_outlives_a_detector_change() {
|
||||
let c = db();
|
||||
let cal = Calibration {
|
||||
a: 1.5,
|
||||
b: -0.5,
|
||||
w_size: 0.1,
|
||||
valid: true,
|
||||
positive_pairs: 40,
|
||||
negative_pairs: 400,
|
||||
};
|
||||
put_calibration(&c, "w600k_mbf", &cal, "abc").unwrap();
|
||||
let (got, hash) = calibration(&c, "scrfd_10g+w600k_mbf").unwrap().unwrap();
|
||||
assert_eq!(got, cal);
|
||||
assert_eq!(hash, "abc");
|
||||
assert!(calibration(&c, "scrfd_10g+other").unwrap().is_none());
|
||||
}
|
||||
|
||||
/// The invariant FR-CULL-10 turns on: a re-index with a better model must
|
||||
/// not throw away the user's own labelling.
|
||||
#[test]
|
||||
@@ -1622,9 +1756,12 @@ mod tests {
|
||||
assert_eq!(edge, 2048, "the marker should record the newer proxy");
|
||||
}
|
||||
|
||||
/// A user who switches detector and back must not find the images the
|
||||
/// second pipeline visited reported as done under the first with no
|
||||
/// faces behind the marker.
|
||||
/// An image holds one pipeline's faces at a time, so the first pipeline's
|
||||
/// marker goes with its faces — but the image stays indexed for every
|
||||
/// pipeline sharing the embedder, because the faces behind the marker
|
||||
/// that remains are the same population. Switching detector and back
|
||||
/// therefore neither re-queues the image nor leaves a marker with nothing
|
||||
/// behind it.
|
||||
#[test]
|
||||
fn re_indexing_under_another_model_drops_the_first_models_marker() {
|
||||
let c = db();
|
||||
@@ -1632,13 +1769,24 @@ mod tests {
|
||||
record_detections(&c, img, "w600k_mbf", 2048, &[face(1)]).unwrap();
|
||||
record_detections(&c, img, "scrfd_2.5g+w600k_mbf", 2048, &[face(2), face(3)]).unwrap();
|
||||
|
||||
assert!(is_indexed(&c, img, "scrfd_2.5g+w600k_mbf").unwrap());
|
||||
assert!(
|
||||
!is_indexed(&c, img, "w600k_mbf").unwrap(),
|
||||
let markers: Vec<String> = {
|
||||
let mut q = c
|
||||
.prepare("SELECT model_id FROM face_index WHERE image_id = ?1")
|
||||
.unwrap();
|
||||
q.query_map([img.0 as i64], |r| r.get(0))
|
||||
.unwrap()
|
||||
.collect::<Result<_, _>>()
|
||||
.unwrap()
|
||||
};
|
||||
assert_eq!(
|
||||
markers,
|
||||
vec!["scrfd_2.5g+w600k_mbf".to_string()],
|
||||
"the first pipeline's marker outlived its faces"
|
||||
);
|
||||
assert!(is_indexed(&c, img, "scrfd_2.5g+w600k_mbf").unwrap());
|
||||
assert!(is_indexed(&c, img, "w600k_mbf").unwrap());
|
||||
assert_eq!(for_image(&c, img).unwrap().len(), 2);
|
||||
assert_eq!(coverage(&c, "w600k_mbf").unwrap().outstanding(), 1);
|
||||
assert_eq!(coverage(&c, "w600k_mbf").unwrap().outstanding(), 0);
|
||||
}
|
||||
|
||||
#[test]
|
||||
|
||||
Reference in New Issue
Block a user