Read each face's eyes, and whether sunglasses hide them

Two MIT classifiers from the same author as the reference pipeline's
whole-body detector: OCEC answers P(open) for one 40×24 eye, SGC
P(sunglasses) for a 48×48 head. Both load in tract once their batch
dimension is pinned by tools/fix-face-model-shapes.sh, like the embedder.

The crops come through the same fitted similarity the aligned face does,
so an eye window is a constant in template units rather than a second
warp, and a tilted head yields an upright eye. Measured on 60 proxies
from the reference library: the eye window plateaus at 22×11, the S
variant beats M and L (which overfit their own domain), and for
sunglasses the aligned face beats a head framing but the higher of the
two catches 11 of 12 pairs against 9 for either alone.

The reading keeps both eyes and the sunglasses number apart, because a
wink averages to the least informative value and a lens of dark glass
draws a confident answer from the eye classifier — over a woman in
sunglasses it read the right eye 0.97 open. Sunglasses take precedence,
and a face behind them is neither open nor a blink.
This commit is contained in:
2026-09-19 14:03:31 +02:00
parent 2481904016
commit 6b51726322
7 changed files with 931 additions and 22 deletions
+205
View File
@@ -0,0 +1,205 @@
//! TRACES: FR-CULL-13
//! The two small classifiers behind a face's eye state (docs/faces.md §17).
//!
//! **OCEC** — *open closed eyes classification*, Hyodo 2025 — reads one
//! 40×24 eye and answers P(open). **SGC** — *sunglasses classification*,
//! Hyodo 2026 — reads a 48×48 head and answers P(sunglasses); it is shown
//! two framings of each face and the higher answer stands, for the reason
//! [`crate::align::SUNGLASSES_WINDOWS`] gives. Both are
//! depthwise-separable CNNs of a few hundred kilobytes, both MIT with their
//! weights, and both were exported with BatchNorm already folded, which is
//! about the friendliest graph tract can be handed.
//!
//! Neither takes a plain buffer. [`EyeClassifier::classify`] takes an
//! [`EyePatch`] and [`SunglassesClassifier::classify`] a [`HeadViews`], each
//! constructible only by the crop in [`crate::align`] that puts the right
//! pixels in it — the same defence [`crate::embed::Embedder`] makes with
//! [`crate::align::Aligned112`], for the same reason: a classifier handed the
//! wrong region returns a confident probability of nothing.
//!
//! # The graphs must have a fixed batch
//!
//! Both ship with a dynamic batch dimension, which tract will not analyse.
//! `tools/fix-face-model-shapes.sh` pins it to 1, exactly as it does for the
//! embedder; the shipped files are the pinned ones.
//!
//! # Pre-processing
//!
//! Read off the reference demos rather than assumed: RGB, `x / 255`, NCHW,
//! the crop resized to the input with bilinear interpolation and **without**
//! preserving its aspect. [`crate::align`]'s crops arrive already at the
//! input size in `0..=1`, so there is nothing left to do but lay them out.
use ndarray::Array4;
use crate::align::{
EyePatch, EyePatches, HeadViews, EYE_PATCH_HEIGHT, EYE_PATCH_WIDTH, SUNGLASSES_EDGE,
};
use crate::eyes::EyeReading;
use crate::{install_backend, FaceError};
/// A loaded OCEC graph.
pub struct EyeClassifier {
session: ort::session::Session,
}
/// A loaded SGC graph.
pub struct SunglassesClassifier {
session: ort::session::Session,
}
/// Open a single-input, single-output classifier and check it is the shape
/// the crop feeding it will be.
///
/// The check is against the *input*, because that is where these two graphs
/// differ from each other and from everything else in this crate: an SGC file
/// given to the eye classifier would otherwise be resized into by an eye
/// patch, and answer. `expected` names the model in the error.
fn open_classifier(
bytes: &[u8],
expected: &'static str,
(h, w): (usize, usize),
) -> Result<ort::session::Session, FaceError> {
install_backend();
let session = ort::session::Session::builder()
.map_err(FaceError::Inference)?
.commit_from_memory(bytes)
.map_err(FaceError::Inference)?;
let input = session.inputs().first().ok_or(FaceError::WrongModel {
expected,
detail: "model has no inputs".into(),
})?;
let shape: Option<Vec<i64>> = input.dtype().tensor_shape().map(|s| s.to_vec());
let want = [1, 3, h as i64, w as i64];
if shape.as_deref() != Some(&want[..]) {
return Err(FaceError::WrongModel {
expected,
detail: format!(
"input '{}' is {:?}, expected {:?} (batch pinned to 1)",
input.name(),
shape,
want
),
});
}
if session.outputs().len() != 1 {
return Err(FaceError::WrongModel {
expected,
detail: format!("{} outputs, expected one", session.outputs().len()),
});
}
Ok(session)
}
/// Lay a `h × w` RGB crop out as the `[1, 3, h, w]` tensor both graphs take.
fn to_nchw(pixels: &[f32], h: usize, w: usize) -> Array4<f32> {
let mut input = Array4::<f32>::zeros((1, 3, h, w));
for y in 0..h {
for x in 0..w {
for c in 0..3 {
input[[0, c, y, x]] = pixels[(y * w + x) * 3 + c];
}
}
}
input
}
/// Run a one-number classifier and read its sigmoid back, clamped.
fn run_scalar(
session: &mut ort::session::Session,
input: Array4<f32>,
expected: &'static str,
) -> Result<f32, FaceError> {
let outputs = session
.run(ort::inputs![
ort::value::Tensor::from_array(input).map_err(FaceError::Inference)?
])
.map_err(FaceError::Inference)?;
let (_, data) = outputs[0]
.try_extract_tensor::<f32>()
.map_err(FaceError::Inference)?;
let Some(&p) = data.first() else {
return Err(FaceError::WrongModel {
expected,
detail: "empty output".into(),
});
};
// The graph ends in a sigmoid, so this is a clamp against rounding and
// nothing more — the reference demo does the same.
Ok(p.clamp(0.0, 1.0))
}
impl EyeClassifier {
pub fn from_path(path: impl AsRef<std::path::Path>) -> Result<Self, FaceError> {
let bytes = std::fs::read(path).map_err(FaceError::ModelRead)?;
Self::from_bytes(&bytes)
}
pub fn from_bytes(bytes: &[u8]) -> Result<Self, FaceError> {
Ok(Self {
session: open_classifier(bytes, "OCEC", (EYE_PATCH_HEIGHT, EYE_PATCH_WIDTH))?,
})
}
/// P(open) for one eye.
pub fn classify(&mut self, eye: &EyePatch) -> Result<f32, FaceError> {
let input = to_nchw(eye.pixels(), EYE_PATCH_HEIGHT, EYE_PATCH_WIDTH);
run_scalar(&mut self.session, input, "OCEC")
}
}
impl SunglassesClassifier {
pub fn from_path(path: impl AsRef<std::path::Path>) -> Result<Self, FaceError> {
let bytes = std::fs::read(path).map_err(FaceError::ModelRead)?;
Self::from_bytes(&bytes)
}
pub fn from_bytes(bytes: &[u8]) -> Result<Self, FaceError> {
Ok(Self {
session: open_classifier(bytes, "SGC", (SUNGLASSES_EDGE, SUNGLASSES_EDGE))?,
})
}
/// P(sunglasses) for one head: the highest answer over its framings.
pub fn classify(&mut self, head: &HeadViews) -> Result<f32, FaceError> {
let mut best = 0.0_f32;
for view in head.views() {
let input = to_nchw(view, SUNGLASSES_EDGE, SUNGLASSES_EDGE);
best = best.max(run_scalar(&mut self.session, input, "SGC")?);
}
Ok(best)
}
}
/// The two classifiers together, which is how every caller holds them.
///
/// One struct rather than two optional parameters, because half a reading is
/// not a reading: an eye state with no sunglasses number behind it is exactly
/// the beach-photograph failure [`crate::eyes`] describes, so the models load
/// together or not at all.
pub struct EyeModels {
pub eyes: EyeClassifier,
pub sunglasses: SunglassesClassifier,
}
impl EyeModels {
pub fn from_paths(
eyes: impl AsRef<std::path::Path>,
sunglasses: impl AsRef<std::path::Path>,
) -> Result<Self, FaceError> {
Ok(Self {
eyes: EyeClassifier::from_path(eyes)?,
sunglasses: SunglassesClassifier::from_path(sunglasses)?,
})
}
/// Read one face's eyes.
pub fn read(&mut self, eyes: &EyePatches, head: &HeadViews) -> Result<EyeReading, FaceError> {
Ok(EyeReading {
right_open: self.eyes.classify(&eyes.right)?,
left_open: self.eyes.classify(&eyes.left)?,
sunglasses: self.sunglasses.classify(head)?,
})
}
}