Benchmarks / CPU and I/O (per commit) (push) Failing after 6m23s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 55s
Build and test / Layer separation (push) Successful in 27s
Traceability / Requirement traces (push) Failing after 54s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Failing after 2m21s
Build and test / Windows (x86_64, cross) (push) Failing after 3m5s
Every face the user has ruled on entered the pass as an anchor, and the scan is exhaustive by design (`dr_face::neighbours`), so a person with 750 confirmed faces cost 750 comparisons against every other face in the library — and the cost of a library grew with how well it was named. Most of those comparisons said nothing new: thirty frames from one afternoon are one point of view, not thirty, and a face that matches one of them matches the rest. Each person now enters through at most 100 of their anchored faces (`dr_face::references`). Eligible are those whose raw embedding is at least 15 long — one above the gallery floor, since a reference speaks for someone rather than merely being admitted — with an unmeasured length admitted as it is everywhere else. From those, the set spanning the greatest volume is chosen greedily: the longest vector first, then at each step the face with the largest component orthogonal to the chosen so far. That is pivoted Gram–Schmidt, and the product of the residuals it picks is the Gram determinant, so the greedy step is the exact greedy on the objective. A near-duplicate of a chosen face has no residual and is passed over; the one profile shot among two hundred frontal frames is taken early; faces inside the span of the chosen add no volume and are not taken to fill the cap. The faces not chosen keep their confirmations and are not touched by the pass — they stay in the anchor map, so it never releases them — they are simply not compared. A person none of whose faces is long enough is still stood for, by their longest, rather than losing their anchor and having their next face filed as a stranger. Under the cap nothing changes: every eligible face stands, and the short ones stay in as the probes they were. At the reference library's 3,851 confirmations the scan shrinks by about a fifth; at 15,000 it is a fifth of what it was.
146 lines
6.0 KiB
Rust
146 lines
6.0 KiB
Rust
//! Faces and identity (S14, docs/faces.md).
|
||
//!
|
||
//! Two models, run over the native render, producing per face a box, five
|
||
//! landmarks, a confidence and a 512-d embedding (FR-CULL-8) — and then the
|
||
//! arithmetic that turns embeddings into people (FR-CULL-9, FR-CULL-10). Two
|
||
//! more, optional, read each face's eyes and whether sunglasses hide them
|
||
//! (FR-CULL-8a, [`classify`] and [`eyes`]).
|
||
//!
|
||
//! Like `dr-segment`, this crate is **device-free**: no GPU adapter, no
|
||
//! Slint, nothing that needs a display. Unlike `dr-segment`, it carries **no
|
||
//! weights at all**, and the absence is deliberate — see [`the licence
|
||
//! note`](#the-weights-are-not-in-this-repository) below.
|
||
//!
|
||
//! # The weights are not in this repository
|
||
//!
|
||
//! The models this crate is built for — SCRFD-500MF and ArcFace/MobileFaceNet
|
||
//! — are InsightFace's, and their pretrained weights carry a **non-commercial
|
||
//! research-only** grant. That is incompatible with GPL-3.0-or-later and with
|
||
//! every channel DarkRoom ships through, so the weights cannot be committed
|
||
//! here the way `dr-segment`'s can, and there is no `embedded-model` feature
|
||
//! for a packaging script to switch on. The application obtains a model at
|
||
//! runtime; this crate takes bytes and never fetches anything.
|
||
//!
|
||
//! docs/faces.md §2 is the full reading, including what would have to change
|
||
//! for that to stop being true. The eye-state models are the exception: MIT,
|
||
//! weights and all, and shipped in `models/face/` (docs/faces.md §17).
|
||
//!
|
||
//! # Why the runtime is split behind a feature
|
||
//!
|
||
//! [`calibrate`], [`cluster`] and [`assign`] are where this subsystem's accuracy
|
||
//! actually lives, and all three are pure arithmetic over embeddings with no
|
||
//! model in them.
|
||
//! They build and test without `inference`, on synthetic embeddings, on a
|
||
//! machine with no weights on it — which is what lets CI cover the part most
|
||
//! likely to be subtly wrong.
|
||
|
||
pub mod align;
|
||
pub mod assign;
|
||
pub mod calibrate;
|
||
#[cfg(feature = "inference")]
|
||
pub mod classify;
|
||
pub mod cluster;
|
||
#[cfg(feature = "inference")]
|
||
pub mod detect;
|
||
#[cfg(feature = "inference")]
|
||
pub mod embed;
|
||
pub mod embedding;
|
||
pub mod eyes;
|
||
#[cfg(feature = "inference")]
|
||
pub mod landmarks;
|
||
pub mod naming;
|
||
pub mod neighbours;
|
||
pub mod references;
|
||
|
||
/// Smallest long edge a face crop may be sampled from.
|
||
///
|
||
/// **A floor on the crop source, not on the detector input.** The distinction
|
||
/// is the whole of FR-CULL-8 and `docs/faces.md` §7: detection letterboxes
|
||
/// every buffer into 640×640, so its input resolution decides nothing, while
|
||
/// [`warp`] samples the 112×112 the embedder sees and so converts source
|
||
/// resolution directly into embedding quality. FR-CULL-8 requires that crop to
|
||
/// come from the native render; this is the guard that catches a caller
|
||
/// sampling from a proxy instead.
|
||
///
|
||
/// 1025 rather than 1024 because 1024 is exactly `dr_thumbs::ThumbSize::Large`,
|
||
/// the stored proxy tier a caller is most likely to reach for by mistake, so
|
||
/// the floor has to exclude it rather than admit it. Written as a minimum so
|
||
/// the test is `edge < MIN_CROP_EDGE` with no boundary to get wrong.
|
||
///
|
||
/// It is a coarse guard and deliberately so: whether any *individual* crop was
|
||
/// upsampled is answered exactly by `faces.crop_px` against [`ALIGNED_EDGE`],
|
||
/// and that is the number §7b measures. This only stops a whole pass reading
|
||
/// from the wrong tier.
|
||
pub const MIN_CROP_EDGE: u32 = 1025;
|
||
|
||
pub use align::{
|
||
crop_box, eye_box, eye_patch, head_views, warp, warp_pixels, Aligned112, EyePatch, HeadViews,
|
||
Pixels, Similarity, ALIGNED_EDGE, ARCFACE_TEMPLATE,
|
||
};
|
||
pub use assign::{identity_shares, RIVAL_FLOOR, TOP_MATCHES};
|
||
pub use calibrate::{Calibration, Pairs, ReliabilityBand};
|
||
#[cfg(feature = "inference")]
|
||
pub use classify::{EyeClassifier, EyeModels, SunglassesClassifier};
|
||
pub use cluster::{
|
||
cluster, cluster_scored, split, Candidate, Cluster, Grouping, DEFAULT_MERGE_PROBABILITY,
|
||
};
|
||
#[cfg(feature = "inference")]
|
||
pub use detect::{DetectOptions, Detection, Detector};
|
||
#[cfg(feature = "inference")]
|
||
pub use embed::{Embedded, Embedder};
|
||
pub use embedding::{
|
||
in_gallery, read_f16_bytes, Embedding, ModelId, EMBEDDING_DIM, MIN_GALLERY_QUALITY,
|
||
};
|
||
pub use eyes::{
|
||
Eye, EyeReading, EyeState, EYES_OPEN_THRESHOLD, HIDDEN_EYE_RATIO, MIN_EYE_PX,
|
||
MIN_EYE_SHARPNESS, SUNGLASSES_THRESHOLD,
|
||
};
|
||
#[cfg(feature = "inference")]
|
||
pub use landmarks::{Landmarker, Landmarks};
|
||
pub use naming::{name_for_instance, name_instances, NamedFace};
|
||
|
||
/// What can go wrong between an image and a face.
|
||
#[derive(Debug, thiserror::Error)]
|
||
pub enum FaceError {
|
||
#[error("could not read model file: {0}")]
|
||
ModelRead(#[source] std::io::Error),
|
||
|
||
#[cfg(feature = "inference")]
|
||
#[error("inference failed: {0}")]
|
||
Inference(#[source] ort::Error),
|
||
|
||
/// The graph is not the one this decoder was written for.
|
||
///
|
||
/// Worth a distinct variant rather than a generic failure: the models in
|
||
/// this space have interchangeable *shapes* and incompatible *layouts*
|
||
/// (a YuNet export also has twelve outputs), so the failure this catches
|
||
/// is not a crash but a page of plausible numbers.
|
||
#[error("model does not look like {expected}: {detail}")]
|
||
WrongModel {
|
||
expected: &'static str,
|
||
detail: String,
|
||
},
|
||
|
||
#[error("image buffer is {got} floats, expected {expected} (RGB, three per pixel)")]
|
||
ImageShape { expected: usize, got: usize },
|
||
}
|
||
|
||
#[cfg(feature = "inference")]
|
||
impl From<dr_inference_engine::Error> for FaceError {
|
||
fn from(e: dr_inference_engine::Error) -> Self {
|
||
match e {
|
||
dr_inference_engine::Error::Inference(e) => FaceError::Inference(e),
|
||
dr_inference_engine::Error::Io(e) => FaceError::ModelRead(e),
|
||
}
|
||
}
|
||
}
|
||
|
||
/// Make sure `ort` has a backend, for the M1 probe example, which drives
|
||
/// `ort` directly rather than through [`detect::Detector`] so it can report
|
||
/// the raw error. Every other path goes through `dr-inference-engine`.
|
||
#[cfg(feature = "inference")]
|
||
#[doc(hidden)]
|
||
pub fn install_backend_for_probe() {
|
||
dr_inference_engine::ensure_runtime();
|
||
}
|