diff --git a/docs/catalog.md b/docs/catalog.md index 30a6e66..68bc901 100644 --- a/docs/catalog.md +++ b/docs/catalog.md @@ -748,6 +748,19 @@ CREATE TABLE faces ( -- (faces.md §6, §9). NULL for a face stored as a unit vector before it -- was kept. quality REAL, + -- What the eyes are doing (FR-CULL-13, faces.md §17): per eye P(open), + -- the source pixels across its box and the sharpness of the patch the + -- classifier saw; and P(sunglasses). All seven or none; NULL is "never + -- read", which every filter treats as unknown rather than as closed. + -- The verdict -- open, closed, sunglasses, unclear -- is a rule in + -- dr_face::eyes, not a column. + eye_right REAL, + eye_right_px REAL, + eye_right_sharp REAL, + eye_left REAL, + eye_left_px REAL, + eye_left_sharp REAL, + sunglasses REAL, -- Which model produced this. An embedding is only comparable to others -- from the same model; mixing them silently yields nonsense similarities. model_id TEXT NOT NULL, @@ -895,4 +908,5 @@ and is not answered here. | FR-CULL-10 | §10.1 `people` and `face_person`, merge-by-redirect, §10.2 clustering as a library pass | | FR-CULL-11 | §5 `Selector::Person`, confirmed-only by default | | FR-CULL-12 | §10.1 derived data in the catalog; names to the sidecar, UUID as merge identity | +| FR-CULL-13 | §10.1 the seven eye columns, derived like the embedding; the filter term is `RatingFilter::eyes_open` | | NFR-SEC-5 | §10.4 sync left undecided and off; nothing in §10 emits an embedding | diff --git a/docs/faces.md b/docs/faces.md index b9cedf3..39e53f9 100644 --- a/docs/faces.md +++ b/docs/faces.md @@ -1266,3 +1266,187 @@ frames to family snapshots as the real unknowns. | NFR-COMPAT-2 | §2 — why the obvious weights cannot ship, and what does instead | | NFR-RES-2 | §1 the cheap model pair, §12 M2/M3/M10 on a phone | | NFR-ARCH-2 | §10 background priority, preempted by visible work | +| FR-CULL-13 | §17 — eye state and sunglasses, the crops, the measurements, and the filter | + +--- + +## 17. Eyes and sunglasses · 2026-09-19 + +FR-CULL-13. Three more models run over every face the pipeline already aligns, and what they +produce is a **filter term** — "eyes open", beside a person or alone — and a badge on the People +screen. Nothing acts on it. §3.9.1's exclusion of blink *detection* was an exclusion of blink +*selection*, and the requirement is written to hold that line: the chip narrows the grid the way a +star count does, and rates nothing. + +### 17.1 The models, and why these + +| | 2d106det | OCEC | SGC | +|---|---|---|---| +| Answers | 106 landmarks, ten round each eye's lids | P(this eye is open) | P(this head wears sunglasses) | +| Input | 192² RGB 0..255, the detector box at 1.5× | one eye, 40×24 RGB, `x/255` | one head, 48×48 RGB, `x/255` | +| Shipped | `buffalo_l`'s, 4.8 MB | S, 483 KB, F1 0.9943 on its own split | L, 6.1 MB, F1 0.9554 on its own split | +| Licence | InsightFace's research-only grant, like the pair (§2.2a) | MIT, code and weights | MIT, code and weights | +| Training data | InsightFace's | *Open and Closed Eyes* (ODC-By 1.0) + Wholebody34 crops (Apache 2.0) | **not stated** — recorded in `models/face/README.md` | +| Cost in tract | ~24 ms per face | ~6 ms per eye | ~6 ms per framing, two framings | + +The two classifiers are from Katsuya Hyodo's ultra-lightweight series — the same author as the +whole-body detector §1.1's reference pipeline uses — and are the first weights in `models/face/` +that do not come out when the project publishes. The landmark model is under the grant the pair +already carries; it was chosen over two permissively licensed alternatives on a measurement +(§17.2) after the decision that this project will not be commercial, which is what §2.2a already +records for the pair. + +All three load in tract as shipped, dynamic batch and all — the first graphs in this subsystem to +do so — and are pinned to a batch of 1 by `tools/fix-face-model-shapes.sh` anyway, because a graph +the engine *analyses* and a graph it has been *measured running* are different claims, and the +embedder's precedent is the safer one. Six milliseconds per classifier call against 0.4 in the +reference README is tract's per-call overhead on a graph this small; the whole reading is under +60 ms per face beside an embedding at 160 ms and a native decode in seconds. + +### 17.2 Where the eye box comes from + +**The eye classifier was trained on a whole-body detector's eye boxes, and this pipeline has no +eye boxes.** It has five landmarks, and SCRFD's eye point is loose: it is one of five points that +place a face, not an eye centre, and on a turned or smiling head the eye sat in a corner of a +window centred on it. Everything below was measured on 60 proxies from the reference library with +25 plainly open-eyed faces labelled by hand (`examples/eyes.rs --dump`, then a contact sheet), and +the count that matters is how many of those 25 the classifier read as open in both eyes. + +**A window on the SCRFD point: 19 of 25.** Windows from 20×10 to 34×17 template units all gave +19–20; smaller lost eyes. Two model-free ways of re-centring the window were then tried and both +lost eyes: the darkest blob near the landmark is the inner corner's shadow or the lash line +(19 → 15), and the most contrasty window is the one that takes in the edge of the nose (19 → 9). +The landmark as SCRFD gives it beats either. + +**A box from a landmark model's lid contour: 22 of 25.** Three models were run over the same +faces, each fed the crop its reference code feeds it, and the eye box cut as the bounding box of +the lid points grown by a margin: + +| model | points | input | tract | per face | open at margin 0.1 | +|---|---|---|---|---|---| +| MediaPipe Face Mesh V2 (Apache 2.0) | 478, with z | 256² | loads | ~36 ms | 22 | +| PIPNet, PINTO's irnet18 export (WFLW, research-only) | 68 | 256² | loads | ~98 ms | 20 | +| **InsightFace 2d106det** | **106** | **192²** | **loads** | **~24 ms** | **22** | + +The margin was swept on the two that tied: 22 at 0 and 0.1, 18 at 0.4, 14 at 0.6 — the training +crops were tight detector boxes, and a tight box is what the classifier wants (`EYE_BOX_MARGIN`). +2d106det ships: it tied the best, costs the least, and is under a grant the project has already +accepted. Face Mesh would be the choice if that changed; it also gives z and an iris, neither of +which this needs yet. + +The box is cut **upright from the native render**, not through the face's alignment — the +training crops were detector boxes, and the contour already says where the eye is on a tilted +head (`align::eye_patch`). A shut eye's contour has no height and is given an open eye's +(`EYE_BOX_MIN_ASPECT`), so the classifier sees the same framing either way. + +### 17.3 Not asking what cannot be answered + +The 25 open faces were never the real problem. The real problem was the faces that were *not* +open-eyed by the classifier's account and were not blinks either, and on the reference sample they +were the commonest wrong answer of all: **a soft eye reads as closed.** A face small enough that +its eye was seven pixels wide, a motion-blurred face, a face from a 1024 proxy where the native +render should have been — each produced a confident "closed" from a classifier shown a smear. The +same failure the face's own sharpness gate exists for (§4.3), one stage down, where the face's gate +cannot see it: a face sharp enough to embed can hold an eye too soft to read, because the eye is a +fortieth of it. + +So the reading is **seven numbers, not a verdict** — per eye P(open), the source pixels across its +box and the sharpness of the patch the classifier saw; and P(sunglasses) — stored as such +(`faces.eye_right`, `faces.eye_right_px`, `faces.eye_right_sharp`, likewise `eye_left`, and +`faces.sunglasses`, schema V16), and the verdict is a rule with thresholds in it, +`dr_face::eyes::EyeReading::state`, the only place the thresholds live: + +``` +sunglasses ≥ 0.5 → Sunglasses (whatever the eyes said) +an eye is readable when px ≥ 12 + and sharpness ≥ 0.02 + and px ≥ 0.6 × the other eye's px +no readable eye → Unreadable +a readable eye < 0.5 → Closed (a blink, or a wink) +otherwise → Open +``` + +**Sunglasses take precedence** because the eye classifier answers confidently over dark glass: +over a woman in sunglasses on the reference library it read her right eye 0.97 open. **The +pixel floor** is where the classifier's own training stopped — its reference footage averaged +15–21 pixels an eye. **The sharpness floor** is the face's measure over the patch, set where the +sample's open eyes were being called closed: the open set ran from 0.019 (a lens reflection) to +5.4, the unreadable ones under 0.02 with the pixels to match. **The width ratio** is the profile: +a landmark model's contour for the far eye of a turned head collapses towards the nose. On the +twenty native renders of §17.4, profiles put the far eye at 0.02–0.43 of the near one's width, +two three-quarter faces whose far eye was reading closed sat at 0.54, and every face looking at +the camera sat at 0.78 or more — a shut eye's box keeps its width, so a wink is not mistaken for +a turn. 0.6 splits the gap. An eye that fails any of the three is not asked, the near eye still +decides, and a face with no readable eye is *unclear* — which is not a blink, and not open, and +which no filter drops. + +**The two eyes are kept apart** rather than averaged, because a wink averages to 0.5 — the one +value that says the least — and "eyes open" means every eye that could be read. + +With the rule in place, the same 60 proxies read: 33 open, 14 closed, 24 sunglasses, 14 unclear. +Of the 25 labelled open faces, 22 open, 2 unclear (eye boxes of 7 and 12 pixels on a child's +face), 1 closed — a squinting smile whose contour collapsed to eleven pixels, which the classifier +is not wrong to call narrow. The 14 closed are downcast eyes, laughs, two sunglasses the head +classifier missed, and the squint. A face 141 pixels across the eye but motion-blurred to a +sharpness of 0.016 reads *unclear* where it read *closed* before, which is the change this +section is for. + +**The filter drops only *Closed*.** `RatingFilter::eyes_open` compiles the rule above into a +predicate on the face row, ANDed into the chosen people's face subquery, so "Anna, eyes open" asks +about Anna's face and not about Bob blinking beside her. The chip is offered only while someone is +chosen and goes when the last person does — without a name in front of it, it would be a verdict +on everyone in the frame. The predicate still handles the empty case, as `NOT EXISTS` over every +face, for a filter arriving by another route; a landscape passes because there is no one in it to +have blinked. Sunglasses pass. Unclear passes. Never read passes — that last is what keeps an old +library from emptying its grid the moment the chip is pressed: until the measuring pass has run, +the honest answer is "everything". A test drives the same five readings through the SQL and +through `state()` and requires the two to agree, so the badge and the grid cannot say different +things. + +### 17.4 What the sample says about accuracy, and what it does not + +The 60-proxy sample above was run at proxy resolution, where the production pass reads the native +render; the eye box on a 200-pixel face is 40 source pixels from the proxy and 240 from the +original. So the shipped configuration was also run over **twenty native renders** from the +reference library — a wedding burst of six frames with six or seven faces each, and a dozen +singles — exported by `face_native --export` and read by `examples/eyes.rs --dump`, 62 faces in +all: 20 open, 28 closed, 12 sunglasses, 2 unclear before the width ratio was moved (below). + +Read off the contact sheet, face by face: the one real blink in the set (`7884.dng`, a man with +his eyes shut) is *closed*; the laughing faces with their eyes screwed shut are *closed*, which a +photographer would call right; the downcast faces are *closed*, which is arguable; the profiles +are judged on the near eye and mostly *open*, which the SCRFD-point pass could not do. Two +faces were wrong: three-quarter views whose far eye's box came to 0.54 of the near one's and read +closed over a cheek, which is what moved the width ratio from 0.45 to 0.6. One is beyond any +floor: a face with a porcelain pot held over the eyes, whose contour is a guess and whose "eyes" +are sharp white china. + +Two things no sample so far can say. None holds more than one real blink, so the precision of +*Closed* is not measured — every closed verdict on the two sheets but the occlusion and the +sunglasses misses is a narrowed or shut eye rather than a wrong one, but that is a reading of a +contact sheet, not a number. And the floors were set on a few dozen faces. Both are M11. + +### 17.5 The measuring pass, and shards + +A face indexed before the models existed, or on a device without them, has no reading. The sweep's +measuring pass — the one V14 built to re-embed faces stored as unit vectors — lists those faces +too, on a device that has the models, and reads their eyes from the same native render with the +box and landmarks already stored (`dr_ui::faces::measure_native`). No detector runs and no identity +moves. A device *without* the models does not list them, or it would fetch every original in the +library to do nothing to it; `dr_catalog::faces::unmeasured_sql` is the one predicate both the +count and the work list use, so the pass converges. The People screen's coverage line counts +these faces as work to measure and keeps the Index button while any remain, or an already-indexed +library could never have its readings filled in. + +Shards carry the seven columns beside `quality`. A peer's faces without a reading are **adopted**, +unlike a peer's faces without a quality (§14): the measuring pass finds this work by the NULL and +not by the run marker, so adoption costs the reading nothing, and a peer with no eye models may be +the only device that has done the detection at all. + +### 17.6 Still to measure + +| # | Measure | Why it decides something | +|---|---|---| +| **M11** | Open-eye recall and blink precision from the **native** pass with the shipped configuration, on a labelled sample that contains real blinks — a burst with one in it is enough | §17.3's figures are proxy figures from 25 faces, with no precision beside them; this is the number FR-CULL-13's acceptance clause asks for, and where the two readability floors get set on more than 25 faces | +| **M12** | Whether the remaining open-eye failure — a lens reflection over an open eye — moves with the sharpness floor, or needs the eye classifier told about spectacles | If the latter, the fix is a classifier trained closer to this domain, and that is a different decision | +| **M13** | Sunglasses recall on more than twelve faces, and the false-positive rate on caps and clear glasses | The two sunglasses the head classifier missed on the sample became false blinks; a library of skiers would say whether that is two faces or a class | diff --git a/docs/requirements.md b/docs/requirements.md index 500fe8a..908886c 100644 --- a/docs/requirements.md +++ b/docs/requirements.md @@ -1304,6 +1304,20 @@ per eye, sub-millisecond on CPU, opset 17 with batch-norm already folded. Its pu 0.99 is on its own crops; ours are cut from a five-point landmark, so the number is measured on this library before it is believed. +**Built 2026-09-19, in part** ([faces.md §17](faces.md)). Eye state ships: OCEC over a box cut +from the lid contour of InsightFace's `2d106det`, plus a sunglasses classifier (SGC, MIT) whose +answer takes precedence over the eyes it hides, and — the part the measurement forced — per eye +the source pixels across the box and the sharpness of the patch, so an eye too small, too soft, or +hidden by the turn of the head is recorded as *unreadable* rather than read as closed. The filter +is the "Eyes open" chip on the library's people filter: with people chosen, it asks about their +faces, and it drops a frame only on a closed eye that could be read. Head pose is **not** built; +the hidden eye of a turned head is caught by its collapsed contour instead, and the selector-term +form ("everyone's eyes open", composable in a saved collection) waits on FR-CULL-11's person term, +which the grid filter also predates. The landmark model is under the InsightFace grant, accepted +on the same terms as the detector and embedder (faces.md §2.2a, decision 2026-09-19: this project +will not be commercial), so the weights test above is met by the two classifiers and not by the +third model. + *Acceptance:* on a labelled set of at least 500 faces from the reference library, eye state is within a stated tolerance of its published F1, reported separately for glasses, profile, and faces under 60 px; facing-the-camera agrees with a hand-labelled split at a stated rate. Both are @@ -2240,7 +2254,7 @@ note where deferring now constrains the design later. | Gaze and eye-contact estimation | Every open gaze model read on 2026-09-19 is trained on Gaze360, MPIIGaze or ETH-XGaze, all research-only, and Gaze360's licence restricts *models trained on it* by name — the InsightFace situation again (D13). Head pose from the five landmarks is the proxy (FR-CULL-8a). Revisit when weights with a clean data chain exist; iris offset within the eye crop is the licence-free fallback if the proxy proves too weak. | | Print layout | — | | Soft proofing | Parameterise the output colour stage by an arbitrary profile so this becomes a UI addition, not a pipeline change. FR-EXP-3's print-dimension mode already half-commits to print workflows. | -| ~~Face recognition~~ | **Undeferred 2026-08-09**, in the narrower form specified in §3.9.1 (FR-CULL-8 … FR-CULL-12): people *grouping and search*, no automated selection. Reclassified as culling rather than AI — it is the same mechanical-grouping category as FR-CULL-5, not the taste operation AI masking is. Gated on spike S14 and decision D13. | +| ~~Face recognition~~ | **Undeferred 2026-08-09**, in the narrower form specified in §3.9.1 (FR-CULL-8 … FR-CULL-13): people *grouping and search*, no automated selection. Reclassified as culling rather than AI — it is the same mechanical-grouping category as FR-CULL-5, not the taste operation AI masking is. Gated on spike S14 and decision D13. | | ~~AI subject masking~~ | **Undeferred 2026-09-19** — it had been built. FR-DEV-3i states what exists: subject and category masks from local models, stored as identity, editable like a drawn mask. The shape this row asked for when it was written — point at a thing, get an editable mask — is the shape that shipped. | | AI upscaling | Deferred. Lower priority than denoise, which has no manual fallback. | | Video | — |