Specify eye state as a filter term, and record what was measured

FR-CULL-13, with §3.9.1's exclusion of blink detection re-read as the
exclusion of blink selection it always was: the stored fact and the chip
are built, a pass that picks the frame where everyone's eyes are open is
not. faces.md §17 has the models, the crop measurements, the four-state
rule and its floors, the native and proxy sheets read face by face, and
what remains to measure.
This commit is contained in:
2026-09-19 14:05:50 +02:00
parent 83f4253b6a
commit cd0ca6785f
3 changed files with 213 additions and 1 deletions
+14
View File
@@ -748,6 +748,19 @@ CREATE TABLE faces (
-- (faces.md §6, §9). NULL for a face stored as a unit vector before it -- (faces.md §6, §9). NULL for a face stored as a unit vector before it
-- was kept. -- was kept.
quality REAL, quality REAL,
-- What the eyes are doing (FR-CULL-13, faces.md §17): per eye P(open),
-- the source pixels across its box and the sharpness of the patch the
-- classifier saw; and P(sunglasses). All seven or none; NULL is "never
-- read", which every filter treats as unknown rather than as closed.
-- The verdict -- open, closed, sunglasses, unclear -- is a rule in
-- dr_face::eyes, not a column.
eye_right REAL,
eye_right_px REAL,
eye_right_sharp REAL,
eye_left REAL,
eye_left_px REAL,
eye_left_sharp REAL,
sunglasses REAL,
-- Which model produced this. An embedding is only comparable to others -- Which model produced this. An embedding is only comparable to others
-- from the same model; mixing them silently yields nonsense similarities. -- from the same model; mixing them silently yields nonsense similarities.
model_id TEXT NOT NULL, model_id TEXT NOT NULL,
@@ -895,4 +908,5 @@ and is not answered here.
| FR-CULL-10 | §10.1 `people` and `face_person`, merge-by-redirect, §10.2 clustering as a library pass | | FR-CULL-10 | §10.1 `people` and `face_person`, merge-by-redirect, §10.2 clustering as a library pass |
| FR-CULL-11 | §5 `Selector::Person`, confirmed-only by default | | FR-CULL-11 | §5 `Selector::Person`, confirmed-only by default |
| FR-CULL-12 | §10.1 derived data in the catalog; names to the sidecar, UUID as merge identity | | FR-CULL-12 | §10.1 derived data in the catalog; names to the sidecar, UUID as merge identity |
| FR-CULL-13 | §10.1 the seven eye columns, derived like the embedding; the filter term is `RatingFilter::eyes_open` |
| NFR-SEC-5 | §10.4 sync left undecided and off; nothing in §10 emits an embedding | | NFR-SEC-5 | §10.4 sync left undecided and off; nothing in §10 emits an embedding |
+184
View File
@@ -1266,3 +1266,187 @@ frames to family snapshots as the real unknowns.
| NFR-COMPAT-2 | §2 — why the obvious weights cannot ship, and what does instead | | NFR-COMPAT-2 | §2 — why the obvious weights cannot ship, and what does instead |
| NFR-RES-2 | §1 the cheap model pair, §12 M2/M3/M10 on a phone | | NFR-RES-2 | §1 the cheap model pair, §12 M2/M3/M10 on a phone |
| NFR-ARCH-2 | §10 background priority, preempted by visible work | | NFR-ARCH-2 | §10 background priority, preempted by visible work |
| FR-CULL-13 | §17 — eye state and sunglasses, the crops, the measurements, and the filter |
---
## 17. Eyes and sunglasses · 2026-09-19
FR-CULL-13. Three more models run over every face the pipeline already aligns, and what they
produce is a **filter term** — "eyes open", beside a person or alone — and a badge on the People
screen. Nothing acts on it. §3.9.1's exclusion of blink *detection* was an exclusion of blink
*selection*, and the requirement is written to hold that line: the chip narrows the grid the way a
star count does, and rates nothing.
### 17.1 The models, and why these
| | 2d106det | OCEC | SGC |
|---|---|---|---|
| Answers | 106 landmarks, ten round each eye's lids | P(this eye is open) | P(this head wears sunglasses) |
| Input | 192² RGB 0..255, the detector box at 1.5× | one eye, 40×24 RGB, `x/255` | one head, 48×48 RGB, `x/255` |
| Shipped | `buffalo_l`'s, 4.8 MB | S, 483 KB, F1 0.9943 on its own split | L, 6.1 MB, F1 0.9554 on its own split |
| Licence | InsightFace's research-only grant, like the pair (§2.2a) | MIT, code and weights | MIT, code and weights |
| Training data | InsightFace's | *Open and Closed Eyes* (ODC-By 1.0) + Wholebody34 crops (Apache 2.0) | **not stated** — recorded in `models/face/README.md` |
| Cost in tract | ~24 ms per face | ~6 ms per eye | ~6 ms per framing, two framings |
The two classifiers are from Katsuya Hyodo's ultra-lightweight series — the same author as the
whole-body detector §1.1's reference pipeline uses — and are the first weights in `models/face/`
that do not come out when the project publishes. The landmark model is under the grant the pair
already carries; it was chosen over two permissively licensed alternatives on a measurement
(§17.2) after the decision that this project will not be commercial, which is what §2.2a already
records for the pair.
All three load in tract as shipped, dynamic batch and all — the first graphs in this subsystem to
do so — and are pinned to a batch of 1 by `tools/fix-face-model-shapes.sh` anyway, because a graph
the engine *analyses* and a graph it has been *measured running* are different claims, and the
embedder's precedent is the safer one. Six milliseconds per classifier call against 0.4 in the
reference README is tract's per-call overhead on a graph this small; the whole reading is under
60 ms per face beside an embedding at 160 ms and a native decode in seconds.
### 17.2 Where the eye box comes from
**The eye classifier was trained on a whole-body detector's eye boxes, and this pipeline has no
eye boxes.** It has five landmarks, and SCRFD's eye point is loose: it is one of five points that
place a face, not an eye centre, and on a turned or smiling head the eye sat in a corner of a
window centred on it. Everything below was measured on 60 proxies from the reference library with
25 plainly open-eyed faces labelled by hand (`examples/eyes.rs --dump`, then a contact sheet), and
the count that matters is how many of those 25 the classifier read as open in both eyes.
**A window on the SCRFD point: 19 of 25.** Windows from 20×10 to 34×17 template units all gave
19–20; smaller lost eyes. Two model-free ways of re-centring the window were then tried and both
lost eyes: the darkest blob near the landmark is the inner corner's shadow or the lash line
(19 → 15), and the most contrasty window is the one that takes in the edge of the nose (19 → 9).
The landmark as SCRFD gives it beats either.
**A box from a landmark model's lid contour: 22 of 25.** Three models were run over the same
faces, each fed the crop its reference code feeds it, and the eye box cut as the bounding box of
the lid points grown by a margin:
| model | points | input | tract | per face | open at margin 0.1 |
|---|---|---|---|---|---|
| MediaPipe Face Mesh V2 (Apache 2.0) | 478, with z | 256² | loads | ~36 ms | 22 |
| PIPNet, PINTO's irnet18 export (WFLW, research-only) | 68 | 256² | loads | ~98 ms | 20 |
| **InsightFace 2d106det** | **106** | **192²** | **loads** | **~24 ms** | **22** |
The margin was swept on the two that tied: 22 at 0 and 0.1, 18 at 0.4, 14 at 0.6 — the training
crops were tight detector boxes, and a tight box is what the classifier wants (`EYE_BOX_MARGIN`).
2d106det ships: it tied the best, costs the least, and is under a grant the project has already
accepted. Face Mesh would be the choice if that changed; it also gives z and an iris, neither of
which this needs yet.
The box is cut **upright from the native render**, not through the face's alignment — the
training crops were detector boxes, and the contour already says where the eye is on a tilted
head (`align::eye_patch`). A shut eye's contour has no height and is given an open eye's
(`EYE_BOX_MIN_ASPECT`), so the classifier sees the same framing either way.
### 17.3 Not asking what cannot be answered
The 25 open faces were never the real problem. The real problem was the faces that were *not*
open-eyed by the classifier's account and were not blinks either, and on the reference sample they
were the commonest wrong answer of all: **a soft eye reads as closed.** A face small enough that
its eye was seven pixels wide, a motion-blurred face, a face from a 1024 proxy where the native
render should have been — each produced a confident "closed" from a classifier shown a smear. The
same failure the face's own sharpness gate exists for (§4.3), one stage down, where the face's gate
cannot see it: a face sharp enough to embed can hold an eye too soft to read, because the eye is a
fortieth of it.
So the reading is **seven numbers, not a verdict** — per eye P(open), the source pixels across its
box and the sharpness of the patch the classifier saw; and P(sunglasses) — stored as such
(`faces.eye_right`, `faces.eye_right_px`, `faces.eye_right_sharp`, likewise `eye_left`, and
`faces.sunglasses`, schema V16), and the verdict is a rule with thresholds in it,
`dr_face::eyes::EyeReading::state`, the only place the thresholds live:
```
sunglasses ≥ 0.5 → Sunglasses (whatever the eyes said)
an eye is readable when px ≥ 12
and sharpness ≥ 0.02
and px ≥ 0.6 × the other eye's px
no readable eye → Unreadable
a readable eye < 0.5 → Closed (a blink, or a wink)
otherwise → Open
```
**Sunglasses take precedence** because the eye classifier answers confidently over dark glass:
over a woman in sunglasses on the reference library it read her right eye 0.97 open. **The
pixel floor** is where the classifier's own training stopped — its reference footage averaged
15–21 pixels an eye. **The sharpness floor** is the face's measure over the patch, set where the
sample's open eyes were being called closed: the open set ran from 0.019 (a lens reflection) to
5.4, the unreadable ones under 0.02 with the pixels to match. **The width ratio** is the profile:
a landmark model's contour for the far eye of a turned head collapses towards the nose. On the
twenty native renders of §17.4, profiles put the far eye at 0.02–0.43 of the near one's width,
two three-quarter faces whose far eye was reading closed sat at 0.54, and every face looking at
the camera sat at 0.78 or more — a shut eye's box keeps its width, so a wink is not mistaken for
a turn. 0.6 splits the gap. An eye that fails any of the three is not asked, the near eye still
decides, and a face with no readable eye is *unclear* — which is not a blink, and not open, and
which no filter drops.
**The two eyes are kept apart** rather than averaged, because a wink averages to 0.5 — the one
value that says the least — and "eyes open" means every eye that could be read.
With the rule in place, the same 60 proxies read: 33 open, 14 closed, 24 sunglasses, 14 unclear.
Of the 25 labelled open faces, 22 open, 2 unclear (eye boxes of 7 and 12 pixels on a child's
face), 1 closed — a squinting smile whose contour collapsed to eleven pixels, which the classifier
is not wrong to call narrow. The 14 closed are downcast eyes, laughs, two sunglasses the head
classifier missed, and the squint. A face 141 pixels across the eye but motion-blurred to a
sharpness of 0.016 reads *unclear* where it read *closed* before, which is the change this
section is for.
**The filter drops only *Closed*.** `RatingFilter::eyes_open` compiles the rule above into a
predicate on the face row, ANDed into the chosen people's face subquery, so "Anna, eyes open" asks
about Anna's face and not about Bob blinking beside her. The chip is offered only while someone is
chosen and goes when the last person does — without a name in front of it, it would be a verdict
on everyone in the frame. The predicate still handles the empty case, as `NOT EXISTS` over every
face, for a filter arriving by another route; a landscape passes because there is no one in it to
have blinked. Sunglasses pass. Unclear passes. Never read passes — that last is what keeps an old
library from emptying its grid the moment the chip is pressed: until the measuring pass has run,
the honest answer is "everything". A test drives the same five readings through the SQL and
through `state()` and requires the two to agree, so the badge and the grid cannot say different
things.
### 17.4 What the sample says about accuracy, and what it does not
The 60-proxy sample above was run at proxy resolution, where the production pass reads the native
render; the eye box on a 200-pixel face is 40 source pixels from the proxy and 240 from the
original. So the shipped configuration was also run over **twenty native renders** from the
reference library — a wedding burst of six frames with six or seven faces each, and a dozen
singles — exported by `face_native --export` and read by `examples/eyes.rs --dump`, 62 faces in
all: 20 open, 28 closed, 12 sunglasses, 2 unclear before the width ratio was moved (below).
Read off the contact sheet, face by face: the one real blink in the set (`7884.dng`, a man with
his eyes shut) is *closed*; the laughing faces with their eyes screwed shut are *closed*, which a
photographer would call right; the downcast faces are *closed*, which is arguable; the profiles
are judged on the near eye and mostly *open*, which the SCRFD-point pass could not do. Two
faces were wrong: three-quarter views whose far eye's box came to 0.54 of the near one's and read
closed over a cheek, which is what moved the width ratio from 0.45 to 0.6. One is beyond any
floor: a face with a porcelain pot held over the eyes, whose contour is a guess and whose "eyes"
are sharp white china.
Two things no sample so far can say. None holds more than one real blink, so the precision of
*Closed* is not measured — every closed verdict on the two sheets but the occlusion and the
sunglasses misses is a narrowed or shut eye rather than a wrong one, but that is a reading of a
contact sheet, not a number. And the floors were set on a few dozen faces. Both are M11.
### 17.5 The measuring pass, and shards
A face indexed before the models existed, or on a device without them, has no reading. The sweep's
measuring pass — the one V14 built to re-embed faces stored as unit vectors — lists those faces
too, on a device that has the models, and reads their eyes from the same native render with the
box and landmarks already stored (`dr_ui::faces::measure_native`). No detector runs and no identity
moves. A device *without* the models does not list them, or it would fetch every original in the
library to do nothing to it; `dr_catalog::faces::unmeasured_sql` is the one predicate both the
count and the work list use, so the pass converges. The People screen's coverage line counts
these faces as work to measure and keeps the Index button while any remain, or an already-indexed
library could never have its readings filled in.
Shards carry the seven columns beside `quality`. A peer's faces without a reading are **adopted**,
unlike a peer's faces without a quality (§14): the measuring pass finds this work by the NULL and
not by the run marker, so adoption costs the reading nothing, and a peer with no eye models may be
the only device that has done the detection at all.
### 17.6 Still to measure
| # | Measure | Why it decides something |
|---|---|---|
| **M11** | Open-eye recall and blink precision from the **native** pass with the shipped configuration, on a labelled sample that contains real blinks — a burst with one in it is enough | §17.3's figures are proxy figures from 25 faces, with no precision beside them; this is the number FR-CULL-13's acceptance clause asks for, and where the two readability floors get set on more than 25 faces |
| **M12** | Whether the remaining open-eye failure — a lens reflection over an open eye — moves with the sharpness floor, or needs the eye classifier told about spectacles | If the latter, the fix is a classifier trained closer to this domain, and that is a different decision |
| **M13** | Sunglasses recall on more than twelve faces, and the false-positive rate on caps and clear glasses | The two sunglasses the head classifier missed on the sample became false blinks; a library of skiers would say whether that is two faces or a class |
+15 -1
View File
@@ -1304,6 +1304,20 @@ per eye, sub-millisecond on CPU, opset 17 with batch-norm already folded. Its pu
0.99 is on its own crops; ours are cut from a five-point landmark, so the number is measured on 0.99 is on its own crops; ours are cut from a five-point landmark, so the number is measured on
this library before it is believed. this library before it is believed.
**Built 2026-09-19, in part** ([faces.md §17](faces.md)). Eye state ships: OCEC over a box cut
from the lid contour of InsightFace's `2d106det`, plus a sunglasses classifier (SGC, MIT) whose
answer takes precedence over the eyes it hides, and — the part the measurement forced — per eye
the source pixels across the box and the sharpness of the patch, so an eye too small, too soft, or
hidden by the turn of the head is recorded as *unreadable* rather than read as closed. The filter
is the "Eyes open" chip on the library's people filter: with people chosen, it asks about their
faces, and it drops a frame only on a closed eye that could be read. Head pose is **not** built;
the hidden eye of a turned head is caught by its collapsed contour instead, and the selector-term
form ("everyone's eyes open", composable in a saved collection) waits on FR-CULL-11's person term,
which the grid filter also predates. The landmark model is under the InsightFace grant, accepted
on the same terms as the detector and embedder (faces.md §2.2a, decision 2026-09-19: this project
will not be commercial), so the weights test above is met by the two classifiers and not by the
third model.
*Acceptance:* on a labelled set of at least 500 faces from the reference library, eye state is *Acceptance:* on a labelled set of at least 500 faces from the reference library, eye state is
within a stated tolerance of its published F1, reported separately for glasses, profile, and within a stated tolerance of its published F1, reported separately for glasses, profile, and
faces under 60 px; facing-the-camera agrees with a hand-labelled split at a stated rate. Both are faces under 60 px; facing-the-camera agrees with a hand-labelled split at a stated rate. Both are
@@ -2240,7 +2254,7 @@ note where deferring now constrains the design later.
| Gaze and eye-contact estimation | Every open gaze model read on 2026-09-19 is trained on Gaze360, MPIIGaze or ETH-XGaze, all research-only, and Gaze360's licence restricts *models trained on it* by name — the InsightFace situation again (D13). Head pose from the five landmarks is the proxy (FR-CULL-8a). Revisit when weights with a clean data chain exist; iris offset within the eye crop is the licence-free fallback if the proxy proves too weak. | | Gaze and eye-contact estimation | Every open gaze model read on 2026-09-19 is trained on Gaze360, MPIIGaze or ETH-XGaze, all research-only, and Gaze360's licence restricts *models trained on it* by name — the InsightFace situation again (D13). Head pose from the five landmarks is the proxy (FR-CULL-8a). Revisit when weights with a clean data chain exist; iris offset within the eye crop is the licence-free fallback if the proxy proves too weak. |
| Print layout | — | | Print layout | — |
| Soft proofing | Parameterise the output colour stage by an arbitrary profile so this becomes a UI addition, not a pipeline change. FR-EXP-3's print-dimension mode already half-commits to print workflows. | | Soft proofing | Parameterise the output colour stage by an arbitrary profile so this becomes a UI addition, not a pipeline change. FR-EXP-3's print-dimension mode already half-commits to print workflows. |
| ~~Face recognition~~ | **Undeferred 2026-08-09**, in the narrower form specified in §3.9.1 (FR-CULL-8 … FR-CULL-12): people *grouping and search*, no automated selection. Reclassified as culling rather than AI — it is the same mechanical-grouping category as FR-CULL-5, not the taste operation AI masking is. Gated on spike S14 and decision D13. | | ~~Face recognition~~ | **Undeferred 2026-08-09**, in the narrower form specified in §3.9.1 (FR-CULL-8 … FR-CULL-13): people *grouping and search*, no automated selection. Reclassified as culling rather than AI — it is the same mechanical-grouping category as FR-CULL-5, not the taste operation AI masking is. Gated on spike S14 and decision D13. |
| ~~AI subject masking~~ | **Undeferred 2026-09-19** — it had been built. FR-DEV-3i states what exists: subject and category masks from local models, stored as identity, editable like a drawn mask. The shape this row asked for when it was written — point at a thing, get an editable mask — is the shape that shipped. | | ~~AI subject masking~~ | **Undeferred 2026-09-19** — it had been built. FR-DEV-3i states what exists: subject and category masks from local models, stored as identity, editable like a drawn mask. The shape this row asked for when it was written — point at a thing, get an editable mask — is the shape that shipped. |
| AI upscaling | Deferred. Lower priority than denoise, which has no manual fallback. | | AI upscaling | Deferred. Lower priority than denoise, which has no manual fallback. |
| Video | — | | Video | — |