docs/ had 26 developer documents flat beside the manual, and the two audiences are very differently sized: most readers want the manual and the gesture reference, a few want the register, the designs and the measurements. The manual and gestures.md stay at the top; everything for someone changing the code moves to docs/dev/, and the two documents that name their own successors — the v0.1 milestone and the UI-refinement plan — go to docs/dev/archive/ rather than being deleted, since both are still cited. docs/README.md is the index, users first. Every reference follows: code comments, Cargo manifests, the workflows, the pre-commit hook, the bench and traceability tools (which locate the repo root by docs/dev/requirements.md now), packaging, the Docker READMEs, CLAUDE.md, CONTRIBUTING.md and the README. The matrix links one level deeper and is regenerated. Links out of the moved documents into the tree gain a level; a link checker over every Markdown file finds none broken.
146 lines
6.0 KiB
Rust
146 lines
6.0 KiB
Rust
//! Faces and identity (S14, docs/dev/faces.md).
|
||
//!
|
||
//! Two models, run over the native render, producing per face a box, five
|
||
//! landmarks, a confidence and a 512-d embedding (FR-CULL-8) — and then the
|
||
//! arithmetic that turns embeddings into people (FR-CULL-9, FR-CULL-10). Two
|
||
//! more, optional, read each face's eyes and whether sunglasses hide them
|
||
//! (FR-CULL-8a, [`classify`] and [`eyes`]).
|
||
//!
|
||
//! Like `dr-segment`, this crate is **device-free**: no GPU adapter, no
|
||
//! Slint, nothing that needs a display. Unlike `dr-segment`, it carries **no
|
||
//! weights at all**, and the absence is deliberate — see [`the licence
|
||
//! note`](#the-weights-are-not-in-this-repository) below.
|
||
//!
|
||
//! # The weights are not in this repository
|
||
//!
|
||
//! The models this crate is built for — SCRFD-500MF and ArcFace/MobileFaceNet
|
||
//! — are InsightFace's, and their pretrained weights carry a **non-commercial
|
||
//! research-only** grant. That is incompatible with GPL-3.0-or-later and with
|
||
//! every channel DarkRoom ships through, so the weights cannot be committed
|
||
//! here the way `dr-segment`'s can, and there is no `embedded-model` feature
|
||
//! for a packaging script to switch on. The application obtains a model at
|
||
//! runtime; this crate takes bytes and never fetches anything.
|
||
//!
|
||
//! docs/dev/faces.md §2 is the full reading, including what would have to change
|
||
//! for that to stop being true. The eye-state models are the exception: MIT,
|
||
//! weights and all, and shipped in `models/face/` (docs/dev/faces.md §17).
|
||
//!
|
||
//! # Why the runtime is split behind a feature
|
||
//!
|
||
//! [`calibrate`], [`cluster`] and [`assign`] are where this subsystem's accuracy
|
||
//! actually lives, and all three are pure arithmetic over embeddings with no
|
||
//! model in them.
|
||
//! They build and test without `inference`, on synthetic embeddings, on a
|
||
//! machine with no weights on it — which is what lets CI cover the part most
|
||
//! likely to be subtly wrong.
|
||
|
||
pub mod align;
|
||
pub mod assign;
|
||
pub mod calibrate;
|
||
#[cfg(feature = "inference")]
|
||
pub mod classify;
|
||
pub mod cluster;
|
||
#[cfg(feature = "inference")]
|
||
pub mod detect;
|
||
#[cfg(feature = "inference")]
|
||
pub mod embed;
|
||
pub mod embedding;
|
||
pub mod eyes;
|
||
#[cfg(feature = "inference")]
|
||
pub mod landmarks;
|
||
pub mod naming;
|
||
pub mod neighbours;
|
||
pub mod references;
|
||
|
||
/// Smallest long edge a face crop may be sampled from.
|
||
///
|
||
/// **A floor on the crop source, not on the detector input.** The distinction
|
||
/// is the whole of FR-CULL-8 and `docs/dev/faces.md` §7: detection letterboxes
|
||
/// every buffer into 640×640, so its input resolution decides nothing, while
|
||
/// [`warp`] samples the 112×112 the embedder sees and so converts source
|
||
/// resolution directly into embedding quality. FR-CULL-8 requires that crop to
|
||
/// come from the native render; this is the guard that catches a caller
|
||
/// sampling from a proxy instead.
|
||
///
|
||
/// 1025 rather than 1024 because 1024 is exactly `dr_thumbs::ThumbSize::Large`,
|
||
/// the stored proxy tier a caller is most likely to reach for by mistake, so
|
||
/// the floor has to exclude it rather than admit it. Written as a minimum so
|
||
/// the test is `edge < MIN_CROP_EDGE` with no boundary to get wrong.
|
||
///
|
||
/// It is a coarse guard and deliberately so: whether any *individual* crop was
|
||
/// upsampled is answered exactly by `faces.crop_px` against [`ALIGNED_EDGE`],
|
||
/// and that is the number §7b measures. This only stops a whole pass reading
|
||
/// from the wrong tier.
|
||
pub const MIN_CROP_EDGE: u32 = 1025;
|
||
|
||
pub use align::{
|
||
crop_box, eye_box, eye_patch, head_views, warp, warp_pixels, Aligned112, EyePatch, HeadViews,
|
||
Pixels, Similarity, ALIGNED_EDGE, ARCFACE_TEMPLATE,
|
||
};
|
||
pub use assign::{identity_shares, RIVAL_FLOOR, TOP_MATCHES};
|
||
pub use calibrate::{Calibration, Pairs, ReliabilityBand};
|
||
#[cfg(feature = "inference")]
|
||
pub use classify::{EyeClassifier, EyeModels, SunglassesClassifier};
|
||
pub use cluster::{
|
||
cluster, cluster_scored, split, Candidate, Cluster, Grouping, DEFAULT_MERGE_PROBABILITY,
|
||
};
|
||
#[cfg(feature = "inference")]
|
||
pub use detect::{DetectOptions, Detection, Detector};
|
||
#[cfg(feature = "inference")]
|
||
pub use embed::{Embedded, Embedder};
|
||
pub use embedding::{
|
||
in_gallery, read_f16_bytes, Embedding, ModelId, EMBEDDING_DIM, MIN_GALLERY_QUALITY,
|
||
};
|
||
pub use eyes::{
|
||
Eye, EyeReading, EyeState, EYES_OPEN_THRESHOLD, HIDDEN_EYE_RATIO, MIN_EYE_PX,
|
||
MIN_EYE_SHARPNESS, SUNGLASSES_THRESHOLD,
|
||
};
|
||
#[cfg(feature = "inference")]
|
||
pub use landmarks::{Landmarker, Landmarks};
|
||
pub use naming::{name_for_instance, name_instances, NamedFace};
|
||
|
||
/// What can go wrong between an image and a face.
|
||
#[derive(Debug, thiserror::Error)]
|
||
pub enum FaceError {
|
||
#[error("could not read model file: {0}")]
|
||
ModelRead(#[source] std::io::Error),
|
||
|
||
#[cfg(feature = "inference")]
|
||
#[error("inference failed: {0}")]
|
||
Inference(#[source] ort::Error),
|
||
|
||
/// The graph is not the one this decoder was written for.
|
||
///
|
||
/// Worth a distinct variant rather than a generic failure: the models in
|
||
/// this space have interchangeable *shapes* and incompatible *layouts*
|
||
/// (a YuNet export also has twelve outputs), so the failure this catches
|
||
/// is not a crash but a page of plausible numbers.
|
||
#[error("model does not look like {expected}: {detail}")]
|
||
WrongModel {
|
||
expected: &'static str,
|
||
detail: String,
|
||
},
|
||
|
||
#[error("image buffer is {got} floats, expected {expected} (RGB, three per pixel)")]
|
||
ImageShape { expected: usize, got: usize },
|
||
}
|
||
|
||
#[cfg(feature = "inference")]
|
||
impl From<dr_inference_engine::Error> for FaceError {
|
||
fn from(e: dr_inference_engine::Error) -> Self {
|
||
match e {
|
||
dr_inference_engine::Error::Inference(e) => FaceError::Inference(e),
|
||
dr_inference_engine::Error::Io(e) => FaceError::ModelRead(e),
|
||
}
|
||
}
|
||
}
|
||
|
||
/// Make sure `ort` has a backend, for the M1 probe example, which drives
|
||
/// `ort` directly rather than through [`detect::Detector`] so it can report
|
||
/// the raw error. Every other path goes through `dr-inference-engine`.
|
||
#[cfg(feature = "inference")]
|
||
#[doc(hidden)]
|
||
pub fn install_backend_for_probe() {
|
||
dr_inference_engine::ensure_runtime();
|
||
}
|