The number beside a suggestion was the mean calibrated probability between the face and the rest of its group, which measures the wrong thing twice. It punishes coverage: a person with two hundred faces over fifteen years is *meant* to have members a given photograph is orthogonal to, so a correct suggestion onto a well-photographed person scored low for being well photographed. And it never asked who else the face might be — a face matching Anna at 0.95 and nobody else, and one matching Anna at 0.95 and her sister at 0.93, came out identical, when the second is the only one worth the user's attention. dr_face::assign answers both, and multiplies them: the mean of the best ten calibrated matches into the identity (the old mean, capped, which is what stops coverage counting against it), times that identity's share of the evidence against every *named* rival. Only named people compete, and per person rather than per group. Both halves of that had to be measured on a real 18,000-face library rather than reasoned about. Normalising across every group made the number useless — median suggestion 21%, four in five under half — because clustering leaves one person spread over many groups, so a face competed against itself; and keying rivals by group left Catherine competing with Catherine, median 39%. Per named person: median 99.5%. Rivals are gathered below the merge threshold, down to even odds: a named person matching at 0.6 will never be merged into but is exactly the competition to discount for. That would be a second similarity scan, the expensive half of regrouping a library, so cluster_scored scans once at the looser floor and hands the merge engine the subset at or above the threshold — pair for pair what it would have scanned for itself, held to that by a test. Leave-one-out over that library's 2,702 confirmations across 54 named people: 99.33% of faces placed on the right person against the old mean's 99.15%, and the number shown for the right person moves from a median of 90.4% to 99.3%. It errs low — 100% correct wherever it states 80% or more — which is the safe direction, and docs/faces.md §9.1 says plainly that the low bands are not calibrated. The example that measures it comes too: this is a claim about a library's numbers, and nobody should have to take it on faith. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
107 lines
4.1 KiB
Rust
107 lines
4.1 KiB
Rust
//! Faces and identity (S14, docs/faces.md).
|
|
//!
|
|
//! Two models, run over the proxy tier, producing per face a box, five
|
|
//! landmarks, a confidence and a 512-d embedding (FR-CULL-8) — and then the
|
|
//! arithmetic that turns embeddings into people (FR-CULL-9, FR-CULL-10).
|
|
//!
|
|
//! Like `dr-segment`, this crate is **device-free**: no GPU adapter, no
|
|
//! Slint, nothing that needs a display. Unlike `dr-segment`, it carries **no
|
|
//! weights at all**, and the absence is deliberate — see [`the licence
|
|
//! note`](#the-weights-are-not-in-this-repository) below.
|
|
//!
|
|
//! # The weights are not in this repository
|
|
//!
|
|
//! The models this crate is built for — SCRFD-500MF and ArcFace/MobileFaceNet
|
|
//! — are InsightFace's, and their pretrained weights carry a **non-commercial
|
|
//! research-only** grant. That is incompatible with GPL-3.0-or-later and with
|
|
//! every channel DarkRoom ships through, so the weights cannot be committed
|
|
//! here the way `dr-segment`'s can, and there is no `embedded-model` feature
|
|
//! for a packaging script to switch on. The application obtains a model at
|
|
//! runtime; this crate takes bytes and never fetches anything.
|
|
//!
|
|
//! docs/faces.md §2 is the full reading, including what would have to change
|
|
//! for that to stop being true.
|
|
//!
|
|
//! # Why the runtime is split behind a feature
|
|
//!
|
|
//! [`calibrate`], [`cluster`] and [`assign`] are where this subsystem's accuracy
|
|
//! actually lives, and all three are pure arithmetic over embeddings with no
|
|
//! model in them.
|
|
//! They build and test without `inference`, on synthetic embeddings, on a
|
|
//! machine with no weights on it — which is what lets CI cover the part most
|
|
//! likely to be subtly wrong.
|
|
|
|
pub mod align;
|
|
pub mod assign;
|
|
pub mod calibrate;
|
|
pub mod cluster;
|
|
#[cfg(feature = "inference")]
|
|
pub mod detect;
|
|
#[cfg(feature = "inference")]
|
|
pub mod embed;
|
|
pub mod embedding;
|
|
pub mod naming;
|
|
pub mod neighbours;
|
|
|
|
pub use align::{warp, Aligned112, Similarity, ALIGNED_EDGE, ARCFACE_TEMPLATE};
|
|
pub use assign::{identity_shares, RIVAL_FLOOR, TOP_MATCHES};
|
|
pub use calibrate::{Calibration, Pairs, ReliabilityBand};
|
|
pub use cluster::{
|
|
cluster, cluster_scored, split, Candidate, Cluster, Grouping, DEFAULT_MERGE_PROBABILITY,
|
|
};
|
|
#[cfg(feature = "inference")]
|
|
pub use detect::{DetectOptions, Detection, Detector};
|
|
#[cfg(feature = "inference")]
|
|
pub use embed::Embedder;
|
|
pub use embedding::{Embedding, ModelId, EMBEDDING_DIM};
|
|
pub use naming::{name_for_instance, name_instances, NamedFace};
|
|
|
|
/// What can go wrong between an image and a face.
|
|
#[derive(Debug, thiserror::Error)]
|
|
pub enum FaceError {
|
|
#[error("could not read model file: {0}")]
|
|
ModelRead(#[source] std::io::Error),
|
|
|
|
#[cfg(feature = "inference")]
|
|
#[error("inference failed: {0}")]
|
|
Inference(#[source] ort::Error),
|
|
|
|
/// The graph is not the one this decoder was written for.
|
|
///
|
|
/// Worth a distinct variant rather than a generic failure: the models in
|
|
/// this space have interchangeable *shapes* and incompatible *layouts*
|
|
/// (a YuNet export also has twelve outputs), so the failure this catches
|
|
/// is not a crash but a page of plausible numbers.
|
|
#[error("model does not look like {expected}: {detail}")]
|
|
WrongModel {
|
|
expected: &'static str,
|
|
detail: String,
|
|
},
|
|
|
|
#[error("image buffer is {got} floats, expected {expected} (RGB, three per pixel)")]
|
|
ImageShape { expected: usize, got: usize },
|
|
}
|
|
|
|
/// Install tract as `ort`'s backend.
|
|
///
|
|
/// Idempotent, and it must happen before any other `ort` call: with
|
|
/// `alternative-backend` there is no linked runtime to fall back on, so an
|
|
/// un-set API is a panic rather than a slow path. Same helper as
|
|
/// `dr-segment::semantic`, for the same reason.
|
|
#[cfg(feature = "inference")]
|
|
pub(crate) fn install_backend() {
|
|
use std::sync::Once;
|
|
static ONCE: Once = Once::new();
|
|
ONCE.call_once(|| {
|
|
let _ = ort::set_api(ort_tract::api());
|
|
});
|
|
}
|
|
|
|
/// [`install_backend`] for the M1 probe example, which drives `ort` directly
|
|
/// rather than through [`detect::Detector`] so it can report the raw error.
|
|
#[cfg(feature = "inference")]
|
|
#[doc(hidden)]
|
|
pub fn install_backend_for_probe() {
|
|
install_backend();
|
|
}
|