764ad55eadff87ac70307c4bfdfccbc5d92a6c6e
18
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
764ad55ead |
Stand each person in the grouping pass by at most 100 references
Benchmarks / CPU and I/O (per commit) (push) Failing after 6m23s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 55s
Build and test / Layer separation (push) Successful in 27s
Traceability / Requirement traces (push) Failing after 54s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Failing after 2m21s
Build and test / Windows (x86_64, cross) (push) Failing after 3m5s
Every face the user has ruled on entered the pass as an anchor, and the scan is exhaustive by design (`dr_face::neighbours`), so a person with 750 confirmed faces cost 750 comparisons against every other face in the library — and the cost of a library grew with how well it was named. Most of those comparisons said nothing new: thirty frames from one afternoon are one point of view, not thirty, and a face that matches one of them matches the rest. Each person now enters through at most 100 of their anchored faces (`dr_face::references`). Eligible are those whose raw embedding is at least 15 long — one above the gallery floor, since a reference speaks for someone rather than merely being admitted — with an unmeasured length admitted as it is everywhere else. From those, the set spanning the greatest volume is chosen greedily: the longest vector first, then at each step the face with the largest component orthogonal to the chosen so far. That is pivoted Gram–Schmidt, and the product of the residuals it picks is the Gram determinant, so the greedy step is the exact greedy on the objective. A near-duplicate of a chosen face has no residual and is passed over; the one profile shot among two hundred frontal frames is taken early; faces inside the span of the chosen add no volume and are not taken to fill the cap. The faces not chosen keep their confirmations and are not touched by the pass — they stay in the anchor map, so it never releases them — they are simply not compared. A person none of whose faces is long enough is still stood for, by their longest, rather than losing their anchor and having their next face filed as a stranger. Under the cap nothing changes: every eligible face stands, and the short ones stay in as the probes they were. At the reference library's 3,851 confirmations the scan shrinks by about a fifth; at 15,000 it is a fifth of what it was. |
||
|
|
d15c41e699 |
Add dr-inference-engine and route every model session through it
One crate names the runtime, the providers and the devices; dr-face and dr-segment ask it for a session by role. It hands ort an API table once per process — from a libonnxruntime it dlopens when the app names a directory holding one, otherwise from tract — so the Rust build stays free of C on every target and a package can install the runtime as a file (docs/inference.md §3). Sessions live in a registry behind a Model handle that holds the bytes, not the session: every use refreshes a timestamp and a reaper unloads whatever sat idle past the decay. A scan that runs the detector on each image never lets it go idle; a click in the develop view lets the segmenter go after thirty seconds; a handle used after that reloads, and reloads on a higher rung if a compiled engine has landed meanwhile. The probe walks the platform's ladder by building strict sessions and timing them against the CPU provider, caches the choice against a fingerprint of the runtime, driver, hardware and models, and compiles engines for the selected rung in the background, smallest model first. Nothing in this commit turns the native path on: the apps still run on tract until they call init with a runtime directory. |
||
|
|
85cc2b1dcc |
Trace the eye reading to FR-CULL-8a and the chip to FR-CULL-13
The register grew both clauses the same day this was built: FR-CULL-8a is the per-face state the reading is, and FR-CULL-13 is the rule that a signal is shown and filtered and never writes a judgement. The tags, faces.md §17 and catalog.md now say which is which; FR-CULL-8a records what of it is built, and that its third model is under the InsightFace grant by the same decision as the pair. |
||
|
|
f5956707e7 |
Cut the eye box from a landmark contour, and refuse eyes that cannot be read
SCRFD's eye point places a face, not an eye: on turned and smiling heads the classifier's window had the eye in a corner, and two model-free ways of re-centring it — the darkest blob, the most contrasty window — both lost open eyes (19 → 15 and 19 → 9 of 25). Three landmark models were then run over the same faces; Face Mesh V2 and InsightFace's 2d106det tied at 22 of 25 and 2d106det ships, being the cheapest by far and under the grant the detector and embedder already carry. The eye box is the tight bounding box of its ten lid points, cut upright from the native render, which is what the classifier was trained on. The larger change is that the reading now carries, per eye, the source pixels across the box and the sharpness of the patch — because the commonest wrong answer on the reference library was a soft eye read as closed, and a classifier shown a smear will always say something. An eye under either floor, or narrower than six tenths of its partner (the far eye of a turned head, whose contour collapses), is not asked; a face with no readable eye is a fourth state, Unreadable, that no filter drops. On twenty native renders the one real blink is caught, the laughing faces are closed, the profiles are judged on the near eye, and the one thing left beyond any floor is a face with a pot held over it. |
||
|
|
6b51726322 |
Read each face's eyes, and whether sunglasses hide them
Two MIT classifiers from the same author as the reference pipeline's whole-body detector: OCEC answers P(open) for one 40×24 eye, SGC P(sunglasses) for a 48×48 head. Both load in tract once their batch dimension is pinned by tools/fix-face-model-shapes.sh, like the embedder. The crops come through the same fitted similarity the aligned face does, so an eye window is a constant in template units rather than a second warp, and a tilted head yields an upright eye. Measured on 60 proxies from the reference library: the eye window plateaus at 22×11, the S variant beats M and L (which overfit their own domain), and for sunglasses the aligned face beats a head framing but the higher of the two catches 11 of 12 pairs against 9 for either alone. The reading keeps both eyes and the sunglasses number apart, because a wink averages to the least informative value and a lens of dark glass draws a confident answer from the eye classifier — over a woman in sunglasses it read the right eye 0.97 open. Sunglasses take precedence, and a face behind them is neither open nor a blink. |
||
|
|
8b3abdb787 |
Keep each face's quality, and never compare against a poor one
The embedder's raw output has a length, and the length is a reading of how recognisable the crop was: a blur, an occlusion or a hard profile comes out short. Normalising threw it away. A short vector sits near the middle of the sphere and matches a little of everyone, which is how one bad crop bridges two people in a grouping pass. So the length is kept — the store now holds the raw vector, re-normalised on load, with the length beside it as `faces.quality` — and a face under MIN_GALLERY_QUALITY (14) is a probe: measured against the gallery and placed where it fits, but never what another face is measured against. Two probes are never paired, and a probe is nobody's evidence for a confidence. The People screen shows the number as "Quality 17.3", dimmed below the floor. Faces indexed before this stored unit vectors and have no reading; they are admitted to the gallery, and schema V14 forgets the run marker of every image holding one so the next indexing pass measures them. A peer's unmeasured shard faces are not adopted, or a sync would write that marker back. |
||
|
|
710fcbc1bd |
Format the face work with the workspace's own rustfmt
Not authored in this session. `cargo fmt --all` reformats every crate, so running it while working on `dr-segment` picked up five files from the recent face and library work that had been committed unformatted. Committed on its own rather than swept into the change that happened to produce it: the diff is pure whitespace, and mixed into a commit that alters an algorithm it would be noise in exactly the place someone is trying to read carefully. `cargo fmt --all -- --check` is a CI gate (tools/ci-local.sh), so this had to land somewhere regardless. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4af3b93dfa |
Index faces from the native render, not from a preview of it
Implements the FR-CULL-8 written two commits ago. The sweep fetched the JPEG preview embedded in each RAW and used that one buffer for both detection and the crop; it now fetches the original, renders it through the same path export uses, reduces that for the detector, and warps the crop back out of the native frame. Three pieces, and each exists for a reason worth stating. dr_face::Pixels lets the warp sample 8-bit RGBA directly. A 24 MP native frame is 96 MB as RGBA and 288 MB converted to the f32 RGB align.rs was written against, and the warp reads about forty thousand pixels out of it. Converting the whole frame to sample 0.2% of it is NFR-RES-2's budget spent on a copy, per image, for a whole library. The variant costs one branch per sample and a test asserts both layouts produce identical crops. The detector gets a box-filtered reduction to 1600px, not the native frame and not a point-sampled one. Averaging rather than sampling because the detector's job is finding small faces and decimation is precisely the operation that removes them: at 4x, fifteen of every sixteen pixels are discarded and a 40px face survives or not depending on where it falls relative to the sample grid. 1600 rather than 640 leaves the letterbox a mild 2.5x rather than a 9x, and bounds the f32 buffer at 20 MB. Landmarks come back in the reduction's coordinates and are scaled to native in one place before any crop pixel is read. This is the failure mode that would not announce itself -- unscaled landmarks put every crop near the top-left corner, which yields faces of something else, cleanly embedded and confidently clustered. The sweep fetches SWEEP_LANES-wide and renders sequentially. Not a placeholder for a parallel version: there is one GPU, so concurrent renders queue on it regardless, and each materialises a native frame. Overlapping them would multiply the one allocation that threatens the memory budget while buying parallelism that does not exist. The chunk drops from 96 to 6 for the same reason -- 96 held 8 MB previews, this holds whole RAWs. The stored edit is deliberately not applied, which is where this departs from export::render_from_library. Face geometry is normalised to the frame, so indexing a cropped render would record boxes against a frame that changes whenever the user changes their mind, and every stored box would quietly become wrong. Orientation is applied: that is a fact about the file rather than an edit. examples/face_native.rs renders one file and indexes it both ways, so the claim behind all of this can be checked against photographs rather than re-read out of the catalog it came from. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
e7b526c550 |
Specify face indexing at native resolution, and say what the proxy cost
FR-CULL-8 said detection runs against the thumbnail or proxy tier and never a full decode, and faces.md §5 said the aligned crop is sampled from that same proxy. Both are wrong in the same place: they treat detection and cropping as one resolution problem when they are two, with opposite answers. Detection does not care. §4.1 fixes the graph's input at 640x640 and letterboxes whatever arrives, so a face filling 2% of the frame reaches the model at 12px whether the buffer handed over is 1024px or 6000px. Every pixel above the detector's own input is discarded before inference. The crop cares about nothing else. §5's warp produces the fixed 112x112 ArcFace sees, so source resolution converts directly into whether those 112 pixels were photographed or interpolated. Reading crop_px across the 18,671 faces the proxy-tier implementation stored: 47.3% were upsampled to reach the embedder, 314 of them by more than 2x, the smallest from 34 source pixels. An upsampled crop does not fail loudly -- it yields a confident embedding of detail that was never there, and the damage appears three stages later as clusters that will not separate. So FR-CULL-8 now specifies four stages with the resolutions named separately: render native through FR-EXP-9's pipeline, downscale for the detector, map boxes and landmarks back to native, crop and align from the native render. The affordability the old rule bought is met instead by when the pass runs -- background, preempted, resumable -- and the requirement says plainly what it now costs on a remote library: the original rather than FR-NC-3's byte range, 412 GB across the reference library's 19,107 images, so a whole-library pass is a transfer under FR-NC-6 rather than something that may start on its own. MIN_CROP_EDGE replaces the MIN_DETECT_EDGE this branch briefly had. Same number, guarding the quantity that turned out to matter. faces.md §7b records both measurements, and marks the second as unexplained rather than dressing it as a finding. Grouped by the buffer detection ran against, faces per image was 0.078 at 1024 or below and 1.82 at 2048 or better, controlled for file type and size. That gap is real and reproducible and I cannot account for it, because the letterbox above says detector input should not matter. M4 is where it gets settled. The crop measurement does not depend on it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
144d2e4e84 |
Revert the detection floor: it guards the wrong resolution
Reverts 53f7cdf and e92d22d. The floor those added sat on Detector::detect, refusing any buffer under 1025px on the reasoning that a small buffer finds no faces. That reasoning does not survive §4.1: the detector letterboxes every input to 640x640, so a face occupying 2% of the frame presents at 12px to the model whether it is handed a 1024px buffer or a 6000px one. Detector input is precisely the quantity that does not matter. Worse than merely useless, it blocks the design FR-CULL-8 now specifies, where the detector is deliberately fed a downscale and the crop is taken from the native render. A guard on detect() rejects exactly that call. What the measurement actually supports is a floor on the *crop* source, which is where resolution converts into embedding quality, and which faces.crop_px already records: 47% of the reference library's faces were upsampled to reach 112x112. That floor is a separate change against the native-resolution path and does not belong on the detector. The 23x faces-per-image gap by source_edge that motivated the original commit is kept in faces.md §7b, restated as the unexplained observation it is rather than the causal claim it was written as. V12 stands: those runs cropped at 1024 whatever detection did, and that is reason enough to look at them again. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1d800d56b0 |
Refuse to detect faces on a proxy too small to find one
Detection was run on whatever proxy the caller happened to have. A small one does not fail -- the image is letterboxed into the detector's 640px input at any size -- so it comes back with almost nothing, and the caller then writes a face_index row saying the photograph was examined. That row is the damage. Nothing distinguishes it from "examined properly, no faces in this one", so the image is never looked at again. The measurement, on the reference library of 23,531 images. Runs against a 1024-edge proxy: 0.078 faces per image, 90% of them finding nothing at all. Runs against 2048 or better: 1.82. To rule out the obvious objection that small proxies just come from small photographs, the same comparison restricted to DNGs -- 1,592 of them averaging 21 MB against 7,724 averaging 23 MB, so the same kind of file in the same library -- gives 0.078 against 1.82 again. Twenty-three fold, on identical source material, identical weights, identical options. So the floor goes in the detector rather than in either sweep, because both of them, the example tool and any future job handler are equally entitled to get this wrong, and there is one place that sees every attempt. It is 1025, not 1024, and the odd-looking number is the point: 1024 is exactly ThumbSize::Large, the tier proxies are stored at and the tier one of the two sweeps was detecting on. A floor that admitted 1024 would admit precisely the population this exists to exclude. Written as a minimum rather than a maximum so the test at each call site is `edge < MIN_DETECT_EDGE` with no boundary left to get wrong. ProxyTooSmall is its own error variant rather than an empty result because the caller has to tell it apart from a failure: nothing is wrong with the image or the model, and the answer is to go and find better pixels, not to retry these ones. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ebb7d3cf5c |
Score a suggestion against the people the user has named
The number beside a suggestion was the mean calibrated probability between the face and the rest of its group, which measures the wrong thing twice. It punishes coverage: a person with two hundred faces over fifteen years is *meant* to have members a given photograph is orthogonal to, so a correct suggestion onto a well-photographed person scored low for being well photographed. And it never asked who else the face might be — a face matching Anna at 0.95 and nobody else, and one matching Anna at 0.95 and her sister at 0.93, came out identical, when the second is the only one worth the user's attention. dr_face::assign answers both, and multiplies them: the mean of the best ten calibrated matches into the identity (the old mean, capped, which is what stops coverage counting against it), times that identity's share of the evidence against every *named* rival. Only named people compete, and per person rather than per group. Both halves of that had to be measured on a real 18,000-face library rather than reasoned about. Normalising across every group made the number useless — median suggestion 21%, four in five under half — because clustering leaves one person spread over many groups, so a face competed against itself; and keying rivals by group left Catherine competing with Catherine, median 39%. Per named person: median 99.5%. Rivals are gathered below the merge threshold, down to even odds: a named person matching at 0.6 will never be merged into but is exactly the competition to discount for. That would be a second similarity scan, the expensive half of regrouping a library, so cluster_scored scans once at the looser floor and hands the merge engine the subset at or above the threshold — pair for pair what it would have scanned for itself, held to that by a test. Leave-one-out over that library's 2,702 confirmations across 54 named people: 99.33% of faces placed on the right person against the old mean's 99.15%, and the number shown for the right person moves from a median of 90.4% to 99.3%. It errs low — 100% correct wherever it states 80% or more — which is the safe direction, and docs/faces.md §9.1 says plainly that the low bands are not calibrated. The example that measures it comes too: this is a claim about a library's numbers, and nobody should have to take it on faith. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c10dca984f |
Regroup the library without stopping the window
Pressing Regroup on a real library did not come back. Clustering 1,813 faces is the textbook agglomeration — compute every pairwise cosine, then repeatedly scan all live group pairs, score each with average link, and merge the best — and the scan is inside the loop. Each merge rescans every surviving pair, and each score is recomputed from scratch over every cross pair. Some 1.6 million pair scores per merge, some 700 merges to do. Three changes, none of which alter the answer. Only above-threshold pairs can ever matter. An average that reaches the threshold must have at least one term at or above it, so two groups with no qualifying pair between them can never merge — not now, and not after any sequence of merges, since merging only adds terms. The new `neighbours` module produces exactly that sparse list: 7,875 pairs rather than 1.6 million on the reference library. It also means the n^2 matrix is never materialised, so memory goes from O(n^2) to O(edges) — 2.5 GB to a few hundred KB at 25,000 faces. Merges cannot cross components, so the connected components of that graph are independent problems: four hundred small agglomerations instead of one large one. Average link is additive — sum(A u B, C) = sum(A, C) + sum(B, C) — so a merged group's scores follow by addition. Kept as running (sum, count) per adjacent pair, a score costs one division instead of a nested loop, and a heap with lazy invalidation replaces the rescan. Measured on the reference library: 0.28s, release, for all 1,813 faces. An exact ANN index was tried and removed, and neighbours.rs records why so it is not rediscovered as a good idea. IVF with a triangle-inequality bound is exact and prunes beautifully on synthetic clusters; on real embeddings it prunes *nothing* — 946 of 946 cell pairs survive. Median pair angle is 88.5 degrees and the merge threshold is 66.2, so the bound needs cells of radius under ~10 degrees, but two photographs of the same person sit 36-60 degrees apart. No ball-based partition of a 512-d near-orthogonal space can be tight enough. So the scan stayed exhaustive and got an unrolled dot product and its blocks spread across cores instead. Correctness is held by keeping the old implementation as an oracle: three tests run both engines over the same population — plain, under co-occurrence and anchor constraints, and with a size-weighted calibration — and assert the clusters are identical. Determinism is asserted at a size where the threaded path is in play. 62 tests pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b846b312b8 |
Run the formatter over the face branch before it reaches CI
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 2s
Build and test / Desktop (Linux) (push) Successful in 1h21m32s
Build and test / Layer separation (push) Successful in 37s
Traceability / Requirement traces (push) Successful in 25s
Build and test / Android (aarch64) (push) Failing after 33m58s
The merge of the SCRFD/MobileFaceNet work brought 69 rustfmt diffs across
dr-catalog, dr-face and dr-ui with it, so `cargo fmt --all -- --check` fails
on master and the Desktop job stops at its Format step — before clippy, the
tests or the release build have run at all. That makes the whole desktop
half of CI blind: a real compile error behind this would look exactly the
same from the outside. There was nothing behind it, as it turns out — with
the formatting fixed, clippy, the test suite and the release build all pass.
Every .rs hunk is `cargo fmt --all` on the pinned 1.92.0 toolchain, not a
hand edit, but it is worth being precise about what that moved, because it
is more than whitespace. Besides reflowing signatures and call chains,
rustfmt reordered the `pub mod` and `pub use` items in dr-face/src/lib.rs so
the `#[cfg(feature = "inference")]` entries sort in place, added the trailing
semicolon inside `let ... else { return }` bodies in identity_ui.rs, wrapped
a bare closure body in braces in cluster.rs, adjusted trailing commas, and
dropped a stray blank line at the end of identity_ui.rs. All of it is
semantically inert; none of it changes behaviour.
docs/traceability.md rides along because it has to. The matrix records each
TRACES tag by line number, and reflowing develop.rs, lib.rs, faces.rs,
identity.rs and identity_ui.rs moved them — FR-CAT-8, FR-CAT-9, FR-CULL-10,
FR-DEV-3, FR-DEV-3a and FR-DEV-3c all shift by a line or two. The matrix was
verified up to date on
|
||
|
|
b55812812a |
Name segmented people from the faces already recognised in them
The segmenter knows it found a person; the face index knows which person. Joining them turns "person" in the mask list into "Anna", which is the difference between a vocabulary of eighty COCO classes and one that includes the user's family. Selecting a subject in a group photograph stops being a guessing game between three identical rows. Containment, not IoU. A face is a small part of the person it belongs to, so a correct pairing has an IoU near zero and anything IoU-based would reject every true match. Confirmed names only. A suggestion is the system's guess, and printing a guessed name onto a mask region would launder it into a fact. Writing the tests corrected the design once: a tight head-and-shoulders portrait, where the face fills most of the person box, is the case where naming is most certain, not least. An earlier guard rejected exactly that and has been removed, with the reasoning left as a test because it is easy to get backwards a second time. The names hang on the develop session, set when the image opens because that is the one moment the catalog and the image id are both in reach. Every segmentation run afterwards picks them up for free, and a library with no face indexing behaves exactly as it did before. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
00e78dc2ac |
Cluster faces into people, and calibrate what a similarity means
FR-CULL-9 forbids thresholding a bare cosine anywhere in the subsystem, so calibrate fits P(same person) per library and reports whether the fit is trustworthy. Two details carry most of the weight. The fit runs against a 200-bin histogram rather than a pair list: a 25,000-face library has ~3e8 pairs and no gradient descent is running over that. And a fresh library has no valid calibration, because the positives have to come from user confirmations or burst siblings -- bootstrapping them from high cosine would fit the calibration to the belief it was supposed to test. Clustering defends against the over-merging FR-CULL-10 warns about with constraints rather than a better threshold: two faces in one photograph never merge, and two groups confirmed as different people never merge. Average link rather than single link, so one strong edge cannot weld two families together. Calibration is defined once, in dr-face, and dr-catalog re-exports it. Two implementations of one probability model is exactly how a number comes to mean the wrong thing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
19981c1033 |
Detect, align and embed faces with SCRFD and MobileFaceNet
Ports the pipeline from the C++ reference in ../scene-actor-extraction (MIT, same author). End to end on real portraits it separates identities the way the reference's fitted calibration says it should: 0.596 between distinct photographs of one person, 0.05 between different people, either side of MBF's 0.267 boundary. Three things are structural rather than incidental: Aligned112 can only be built by align::warp, so Embedder::embed cannot be handed an unaligned bounding-box crop. That mistake yields 512 plausible unit-norm numbers and no error, so the type system refuses it instead. Embedding carries its ModelId and cosine() returns None across models, because a cross-model similarity is the one mistake that produces plausible garbage rather than a failure. The model-free half -- alignment, embedding arithmetic, f16 storage -- sits outside the inference feature and is covered by 11 tests that need no weights on the machine. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
72410f39c6 |
Answer M1: tract loads both face graphs once their dims are pinned
Neither InsightFace export parses as shipped -- SCRFD fails at its input node, ArcFace at the first Conv -- which is the same wall dr-segment hit on YOLO's dynamic export. Both load cleanly with the input dims frozen, so the pure-Rust runtime holds for the face pipeline too. tools/fix-face-model-shapes.sh does the freezing, and exists so the artefact is reproducible rather than a binary someone once produced. It takes two forms because the two graphs need different ones: ArcFace's batch is a named dim_param, SCRFD's H and W are dynamic but unnamed. Also notes YuNet loading with no intervention, which matters for the licence question in faces.md 2.3. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |