Cut the eye box from a landmark contour, and refuse eyes that cannot be read

SCRFD's eye point places a face, not an eye: on turned and smiling heads
the classifier's window had the eye in a corner, and two model-free ways
of re-centring it — the darkest blob, the most contrasty window — both
lost open eyes (19 → 15 and 19 → 9 of 25). Three landmark models were
then run over the same faces; Face Mesh V2 and InsightFace's 2d106det
tied at 22 of 25 and 2d106det ships, being the cheapest by far and under
the grant the detector and embedder already carry. The eye box is the
tight bounding box of its ten lid points, cut upright from the native
render, which is what the classifier was trained on.

The larger change is that the reading now carries, per eye, the source
pixels across the box and the sharpness of the patch — because the
commonest wrong answer on the reference library was a soft eye read as
closed, and a classifier shown a smear will always say something. An eye
under either floor, or narrower than six tenths of its partner (the far
eye of a turned head, whose contour collapses), is not asked; a face with
no readable eye is a fourth state, Unreadable, that no filter drops. On
twenty native renders the one real blink is caught, the laughing faces
are closed, the profiles are judged on the near eye, and the one thing
left beyond any floor is a face with a pot held over it.
This commit is contained in:
2026-09-19 14:04:08 +02:00
parent b908d861e0
commit f5956707e7
7 changed files with 690 additions and 311 deletions
+68 -15
View File
@@ -15,7 +15,9 @@
//! constructible only by the crop in [`crate::align`] that puts the right
//! pixels in it — the same defence [`crate::embed::Embedder`] makes with
//! [`crate::align::Aligned112`], for the same reason: a classifier handed the
//! wrong region returns a confident probability of nothing.
//! wrong region returns a confident probability of nothing. Where the eye
//! box comes from is [`crate::landmarks`]; [`EyeModels::read`] is the whole
//! chain.
//!
//! # The graphs must have a fixed batch
//!
@@ -33,10 +35,12 @@
use ndarray::Array4;
use crate::align::{
EyePatch, EyePatches, HeadViews, EYE_PATCH_HEIGHT, EYE_PATCH_WIDTH, SUNGLASSES_EDGE,
eye_box, eye_patch, head_views, EyePatch, HeadViews, EYE_PATCH_HEIGHT, EYE_PATCH_WIDTH,
SUNGLASSES_EDGE,
};
use crate::eyes::EyeReading;
use crate::{install_backend, FaceError};
use crate::eyes::{Eye, EyeReading};
use crate::landmarks::Landmarker;
use crate::{install_backend, FaceError, Pixels};
/// A loaded OCEC graph.
pub struct EyeClassifier {
@@ -172,34 +176,83 @@ impl SunglassesClassifier {
}
}
/// The two classifiers together, which is how every caller holds them.
/// The three models behind a reading, which is how every caller holds them.
///
/// One struct rather than two optional parameters, because half a reading is
/// not a reading: an eye state with no sunglasses number behind it is exactly
/// the beach-photograph failure [`crate::eyes`] describes, so the models load
/// together or not at all.
/// One struct rather than three optional parameters, because a partial
/// reading is not a reading: an eye state with no sunglasses number behind
/// it is exactly the beach-photograph failure [`crate::eyes`] describes, and
/// an eye box without the landmarks is the loose one this module replaced.
/// The models load together or not at all.
pub struct EyeModels {
pub landmarks: Landmarker,
pub eyes: EyeClassifier,
pub sunglasses: SunglassesClassifier,
}
impl EyeModels {
pub fn from_paths(
landmarks: impl AsRef<std::path::Path>,
eyes: impl AsRef<std::path::Path>,
sunglasses: impl AsRef<std::path::Path>,
) -> Result<Self, FaceError> {
Ok(Self {
landmarks: Landmarker::from_path(landmarks)?,
eyes: EyeClassifier::from_path(eyes)?,
sunglasses: SunglassesClassifier::from_path(sunglasses)?,
})
}
/// Read one face's eyes.
pub fn read(&mut self, eyes: &EyePatches, head: &HeadViews) -> Result<EyeReading, FaceError> {
Ok(EyeReading {
right_open: self.eyes.classify(&eyes.right)?,
left_open: self.eyes.classify(&eyes.left)?,
sunglasses: self.sunglasses.classify(head)?,
})
///
/// `bbox` is the detector's `(x0, y0, x1, y1)` and `landmarks5` its five
/// points, both in source pixels; the buffer is the one the aligned
/// crop was taken from, so an eye is read from the same pixels the
/// embedder saw the face in. `None` where nothing could be cut — a
/// degenerate box or landmarks — which the caller stores as "not read".
pub fn read(
&mut self,
px: Pixels<'_>,
width: usize,
height: usize,
bbox: (f32, f32, f32, f32),
landmarks5: &[(f32, f32); 5],
) -> Result<Option<EyeReading>, FaceError> {
let Some(lm) = self.landmarks.landmarks(px, width, height, bbox)? else {
return Ok(None);
};
let Some(head) = head_views(px, width, height, landmarks5) else {
return Ok(None);
};
let mut eye = |contour: &[(f32, f32)]| -> Result<Eye, FaceError> {
// A hidden eye's contour can collapse to no width. Its numbers
// are then zero — no pixels, no sharpness — which is what the
// rule in `crate::eyes` reads as "not readable".
let Some(b) = eye_box(contour) else {
return Ok(Eye {
open: 0.0,
px: 0.0,
sharpness: 0.0,
});
};
let Some(patch) = eye_patch(px, width, height, b) else {
return Ok(Eye {
open: 0.0,
px: 0.0,
sharpness: 0.0,
});
};
Ok(Eye {
open: self.eyes.classify(&patch)?,
px: patch.source_px(),
sharpness: patch.sharpness(),
})
};
let right = eye(&lm.right_eye())?;
let left = eye(&lm.left_eye())?;
Ok(Some(EyeReading {
right,
left,
sunglasses: self.sunglasses.classify(&head)?,
}))
}
}