Two MIT classifiers from the same author as the reference pipeline's whole-body detector: OCEC answers P(open) for one 40×24 eye, SGC P(sunglasses) for a 48×48 head. Both load in tract once their batch dimension is pinned by tools/fix-face-model-shapes.sh, like the embedder. The crops come through the same fitted similarity the aligned face does, so an eye window is a constant in template units rather than a second warp, and a tilted head yields an upright eye. Measured on 60 proxies from the reference library: the eye window plateaus at 22×11, the S variant beats M and L (which overfit their own domain), and for sunglasses the aligned face beats a head framing but the higher of the two catches 11 of 12 pairs against 9 for either alone. The reading keeps both eyes and the sunglasses number apart, because a wink averages to the least informative value and a lens of dark glass draws a confident answer from the eye classifier — over a woman in sunglasses it read the right eye 0.97 open. Sunglasses take precedence, and a face behind them is neither open nor a blink.
139 lines
5.6 KiB
Rust
139 lines
5.6 KiB
Rust
//! Faces and identity (S14, docs/faces.md).
|
||
//!
|
||
//! Two models, run over the proxy tier, producing per face a box, five
|
||
//! landmarks, a confidence and a 512-d embedding (FR-CULL-8) — and then the
|
||
//! arithmetic that turns embeddings into people (FR-CULL-9, FR-CULL-10).
|
||
//!
|
||
//! Like `dr-segment`, this crate is **device-free**: no GPU adapter, no
|
||
//! Slint, nothing that needs a display. Unlike `dr-segment`, it carries **no
|
||
//! weights at all**, and the absence is deliberate — see [`the licence
|
||
//! note`](#the-weights-are-not-in-this-repository) below.
|
||
//!
|
||
//! # The weights are not in this repository
|
||
//!
|
||
//! The models this crate is built for — SCRFD-500MF and ArcFace/MobileFaceNet
|
||
//! — are InsightFace's, and their pretrained weights carry a **non-commercial
|
||
//! research-only** grant. That is incompatible with GPL-3.0-or-later and with
|
||
//! every channel DarkRoom ships through, so the weights cannot be committed
|
||
//! here the way `dr-segment`'s can, and there is no `embedded-model` feature
|
||
//! for a packaging script to switch on. The application obtains a model at
|
||
//! runtime; this crate takes bytes and never fetches anything.
|
||
//!
|
||
//! docs/faces.md §2 is the full reading, including what would have to change
|
||
//! for that to stop being true.
|
||
//!
|
||
//! # Why the runtime is split behind a feature
|
||
//!
|
||
//! [`calibrate`], [`cluster`] and [`assign`] are where this subsystem's accuracy
|
||
//! actually lives, and all three are pure arithmetic over embeddings with no
|
||
//! model in them.
|
||
//! They build and test without `inference`, on synthetic embeddings, on a
|
||
//! machine with no weights on it — which is what lets CI cover the part most
|
||
//! likely to be subtly wrong.
|
||
|
||
pub mod align;
|
||
pub mod assign;
|
||
pub mod calibrate;
|
||
#[cfg(feature = "inference")]
|
||
pub mod classify;
|
||
pub mod cluster;
|
||
#[cfg(feature = "inference")]
|
||
pub mod detect;
|
||
#[cfg(feature = "inference")]
|
||
pub mod embed;
|
||
pub mod embedding;
|
||
pub mod eyes;
|
||
pub mod naming;
|
||
pub mod neighbours;
|
||
|
||
/// Smallest long edge a face crop may be sampled from.
|
||
///
|
||
/// **A floor on the crop source, not on the detector input.** The distinction
|
||
/// is the whole of FR-CULL-8 and `docs/faces.md` §7: detection letterboxes
|
||
/// every buffer into 640×640, so its input resolution decides nothing, while
|
||
/// [`warp`] samples the 112×112 the embedder sees and so converts source
|
||
/// resolution directly into embedding quality. FR-CULL-8 requires that crop to
|
||
/// come from the native render; this is the guard that catches a caller
|
||
/// sampling from a proxy instead.
|
||
///
|
||
/// 1025 rather than 1024 because 1024 is exactly `dr_thumbs::ThumbSize::Large`,
|
||
/// the stored proxy tier a caller is most likely to reach for by mistake, so
|
||
/// the floor has to exclude it rather than admit it. Written as a minimum so
|
||
/// the test is `edge < MIN_CROP_EDGE` with no boundary to get wrong.
|
||
///
|
||
/// It is a coarse guard and deliberately so: whether any *individual* crop was
|
||
/// upsampled is answered exactly by `faces.crop_px` against [`ALIGNED_EDGE`],
|
||
/// and that is the number §7b measures. This only stops a whole pass reading
|
||
/// from the wrong tier.
|
||
pub const MIN_CROP_EDGE: u32 = 1025;
|
||
|
||
pub use align::{
|
||
eye_patches, head_views, warp, warp_pixels, Aligned112, EyePatch, EyePatches, HeadViews,
|
||
Pixels, Similarity, ALIGNED_EDGE, ARCFACE_TEMPLATE,
|
||
};
|
||
pub use assign::{identity_shares, RIVAL_FLOOR, TOP_MATCHES};
|
||
pub use calibrate::{Calibration, Pairs, ReliabilityBand};
|
||
#[cfg(feature = "inference")]
|
||
pub use classify::{EyeClassifier, EyeModels, SunglassesClassifier};
|
||
pub use cluster::{
|
||
cluster, cluster_scored, split, Candidate, Cluster, Grouping, DEFAULT_MERGE_PROBABILITY,
|
||
};
|
||
#[cfg(feature = "inference")]
|
||
pub use detect::{DetectOptions, Detection, Detector};
|
||
#[cfg(feature = "inference")]
|
||
pub use embed::{Embedded, Embedder};
|
||
pub use embedding::{
|
||
in_gallery, read_f16_bytes, Embedding, ModelId, EMBEDDING_DIM, MIN_GALLERY_QUALITY,
|
||
};
|
||
pub use eyes::{EyeReading, EyeState, EYES_OPEN_THRESHOLD, SUNGLASSES_THRESHOLD};
|
||
pub use naming::{name_for_instance, name_instances, NamedFace};
|
||
|
||
/// What can go wrong between an image and a face.
|
||
#[derive(Debug, thiserror::Error)]
|
||
pub enum FaceError {
|
||
#[error("could not read model file: {0}")]
|
||
ModelRead(#[source] std::io::Error),
|
||
|
||
#[cfg(feature = "inference")]
|
||
#[error("inference failed: {0}")]
|
||
Inference(#[source] ort::Error),
|
||
|
||
/// The graph is not the one this decoder was written for.
|
||
///
|
||
/// Worth a distinct variant rather than a generic failure: the models in
|
||
/// this space have interchangeable *shapes* and incompatible *layouts*
|
||
/// (a YuNet export also has twelve outputs), so the failure this catches
|
||
/// is not a crash but a page of plausible numbers.
|
||
#[error("model does not look like {expected}: {detail}")]
|
||
WrongModel {
|
||
expected: &'static str,
|
||
detail: String,
|
||
},
|
||
|
||
#[error("image buffer is {got} floats, expected {expected} (RGB, three per pixel)")]
|
||
ImageShape { expected: usize, got: usize },
|
||
}
|
||
|
||
/// Install tract as `ort`'s backend.
|
||
///
|
||
/// Idempotent, and it must happen before any other `ort` call: with
|
||
/// `alternative-backend` there is no linked runtime to fall back on, so an
|
||
/// un-set API is a panic rather than a slow path. Same helper as
|
||
/// `dr-segment::semantic`, for the same reason.
|
||
#[cfg(feature = "inference")]
|
||
pub(crate) fn install_backend() {
|
||
use std::sync::Once;
|
||
static ONCE: Once = Once::new();
|
||
ONCE.call_once(|| {
|
||
let _ = ort::set_api(ort_tract::api());
|
||
});
|
||
}
|
||
|
||
/// [`install_backend`] for the M1 probe example, which drives `ort` directly
|
||
/// rather than through [`detect::Detector`] so it can report the raw error.
|
||
#[cfg(feature = "inference")]
|
||
#[doc(hidden)]
|
||
pub fn install_backend_for_probe() {
|
||
install_backend();
|
||
}
|