Commit Graph
11 Commits
Author SHA1 Message Date
dtourolleandClaude Opus 5 a1790c3e67 Stop indexing faces too small or too blurred to be anyone
The library was storing faces at 52 source pixels and embedding whatever came
back. There was a size floor, but it was 40 pixels on the *bounding box*, and
there was no blur gate at all — so a subject walking through a half-second
exposure detected confidently, aligned cleanly, and produced a perfectly
ordinary-looking 512-vector. Nothing downstream can tell that apart from a real
face, and because blurs resemble each other more than they resemble the people
they were, they cluster together and weld unrelated identities into one group.

Two floors, both measured rather than guessed. `face_index --quality` runs the
detector over real proxies with both gates disabled and prints the distribution;
over 1,503 faces in 600 images of the reference library:

   percentile   crop px   sharpness
           1%        16      0.0006
          25%        23      0.0025
          50%        38      0.0071
          75%        76      0.0284
          99%       352      0.4282

The median face in a personal library is 38 pixels. Most of what the detector
finds is background: people across a square, a face on a poster, a stranger at
the next table. They are real detections and useless identifications.

**Size, on the crop rather than the box.** "At least 64x64" has to mean the
pixels the *embedder* sees, and the box is not that — the ArcFace template
reaches past it for forehead and chin, so the aligned crop spans roughly 1.3x
the box's shorter edge. The floor is therefore `min_source_px` on the aligned
crop, applied after the warp fixes the scale, and `min_face_px` drops to 48 as
what it always really was: a cheap pre-filter set low enough that it cannot
reject a face the real floor would have kept.

**Sharpness.** Variance of the Laplacian divided by the variance of the luma it
was taken over. The division is the part that matters: raw Laplacian variance
scales with contrast, so a threshold on it would quietly discard every backlit
portrait in the library. The ratio asks how much of the crop's variation is
edges rather than broad gradients, and is invariant to exposure.

What each pair removes, cumulatively, of everything the detector finds:

    min crop   min sharp   size cut   blur cut       kept
          64       0.000        70%         0%        30%
          64       0.010        70%         3%        27%
          64       0.020        70%         7%        23%
          80       0.010        76%         2%        21%

64 and 0.020. The size floor does most of the work, and the blur floor removing
only 7% on top of it is the point rather than a disappointment: at 64 pixels
most faces are already sharp, and what it takes out is the large-but-soft one —
precisely the face that would otherwise contribute a confident, wrong embedding.

The two gates are not independent and the doc comments say so: a face under 112
pixels was upsampled to reach the embedder, and upsampling invents no edges, so
small faces score low on sharpness even when the original was crisp. That is why
`--quality` prints them together.

**This will re-index.** Around 70% of what the current settings store falls below
the new floors — faces between 20 and 40 pixels that nobody could identify. The
People screen gets shorter and every group in it gets better.

66 dr-face tests pass, including that a blurred crop scores below a sharp one,
that halving the contrast does not move the score, and that an upsampled face
scores below the same face at full size.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 22:17:38 +02:00
dtourolleandClaude Opus 5 7275c020d7 Group people at the threshold the library actually supports
0.90 left a third of the reference library ungrouped: 1,213 of 1,813 faces in a
group, and the rest sitting alone in a screen that had nothing to offer for
them.

"Is 0.90 too tight" is not answerable from the number. It is a probability, and
which cosine it lands on depends on the calibration — so the first half of this
is a way to ask the question properly. `face_index --tune` runs the real
clusterer over the real embeddings at ten thresholds and prints what each one
produces. It writes nothing; comparing thresholds by applying them would have
each one pollute the next.

On the reference library:

      P   cosine   groups  grouped  largest
   0.95    0.449      311      62%       51
   0.90    0.403      316      67%       51
   0.85    0.374      318      70%       57
   0.80    0.353      328      74%       69
   0.75    0.335      327      77%       69
   0.70    0.319      326      79%       81
   0.50    0.267      303      85%       90

The count of *groups* is the signal, not the count of grouped faces. Loosening
from 0.95 makes it climb: real people are being assembled out of fragments. It
peaks at 0.80 and then falls — and a falling group count while the grouped faces
keep rising is the shape of over-merging, separate identities being welded
together. That is the FR-CULL-10 failure, and the one the user cannot undo by
hand.

So 0.80: the loosest setting still building people rather than melting them
together. A third more of the library gets grouped than at 0.90, and the largest
group grows by eighteen faces rather than by forty.

The table is one library, and the doc comment says so — `--tune` reruns it on
any other.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 21:17:16 +02:00
dtourolleandClaude Opus 5 c10dca984f Regroup the library without stopping the window
Pressing Regroup on a real library did not come back. Clustering 1,813 faces is
the textbook agglomeration — compute every pairwise cosine, then repeatedly scan
all live group pairs, score each with average link, and merge the best — and the
scan is inside the loop. Each merge rescans every surviving pair, and each score
is recomputed from scratch over every cross pair. Some 1.6 million pair scores
per merge, some 700 merges to do.

Three changes, none of which alter the answer.

Only above-threshold pairs can ever matter. An average that reaches the
threshold must have at least one term at or above it, so two groups with no
qualifying pair between them can never merge — not now, and not after any
sequence of merges, since merging only adds terms. The new `neighbours` module
produces exactly that sparse list: 7,875 pairs rather than 1.6 million on the
reference library. It also means the n^2 matrix is never materialised, so memory
goes from O(n^2) to O(edges) — 2.5 GB to a few hundred KB at 25,000 faces.

Merges cannot cross components, so the connected components of that graph are
independent problems: four hundred small agglomerations instead of one large one.

Average link is additive — sum(A u B, C) = sum(A, C) + sum(B, C) — so a merged
group's scores follow by addition. Kept as running (sum, count) per adjacent
pair, a score costs one division instead of a nested loop, and a heap with lazy
invalidation replaces the rescan.

Measured on the reference library: 0.28s, release, for all 1,813 faces.

An exact ANN index was tried and removed, and neighbours.rs records why so it is
not rediscovered as a good idea. IVF with a triangle-inequality bound is exact
and prunes beautifully on synthetic clusters; on real embeddings it prunes
*nothing* — 946 of 946 cell pairs survive. Median pair angle is 88.5 degrees and
the merge threshold is 66.2, so the bound needs cells of radius under ~10
degrees, but two photographs of the same person sit 36-60 degrees apart. No
ball-based partition of a 512-d near-orthogonal space can be tight enough. So
the scan stayed exhaustive and got an unrolled dot product and its blocks spread
across cores instead.

Correctness is held by keeping the old implementation as an oracle: three tests
run both engines over the same population — plain, under co-occurrence and
anchor constraints, and with a size-weighted calibration — and assert the
clusters are identical. Determinism is asserted at a size where the threaded
path is in play.

62 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 21:16:55 +02:00
dtourolleandClaude Opus 5 b846b312b8 Run the formatter over the face branch before it reaches CI
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 2s
Build and test / Desktop (Linux) (push) Successful in 1h21m32s
Build and test / Layer separation (push) Successful in 37s
Traceability / Requirement traces (push) Successful in 25s
Build and test / Android (aarch64) (push) Failing after 33m58s
The merge of the SCRFD/MobileFaceNet work brought 69 rustfmt diffs across
dr-catalog, dr-face and dr-ui with it, so `cargo fmt --all -- --check` fails
on master and the Desktop job stops at its Format step — before clippy, the
tests or the release build have run at all. That makes the whole desktop
half of CI blind: a real compile error behind this would look exactly the
same from the outside. There was nothing behind it, as it turns out — with
the formatting fixed, clippy, the test suite and the release build all pass.

Every .rs hunk is `cargo fmt --all` on the pinned 1.92.0 toolchain, not a
hand edit, but it is worth being precise about what that moved, because it
is more than whitespace. Besides reflowing signatures and call chains,
rustfmt reordered the `pub mod` and `pub use` items in dr-face/src/lib.rs so
the `#[cfg(feature = "inference")]` entries sort in place, added the trailing
semicolon inside `let ... else { return }` bodies in identity_ui.rs, wrapped
a bare closure body in braces in cluster.rs, adjusted trailing commas, and
dropped a stray blank line at the end of identity_ui.rs. All of it is
semantically inert; none of it changes behaviour.

docs/traceability.md rides along because it has to. The matrix records each
TRACES tag by line number, and reflowing develop.rs, lib.rs, faces.rs,
identity.rs and identity_ui.rs moved them — FR-CAT-8, FR-CAT-9, FR-CULL-10,
FR-DEV-3, FR-DEV-3a and FR-DEV-3c all shift by a line or two. The matrix was
verified up to date on d777f7f before this commit, so this is drift these
formatting changes introduced, not pre-existing staleness being swept up.
Leaving it for a follow-up commit would hand traceability-check.yml a
failure caused entirely by a whitespace change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 13:12:49 +02:00
dtourolle d777f7f44d Merge branch 'master' into worktree-faces-scrfd-mbf
Build and test / Desktop (Linux) (push) Failing after 25s
Build and test / Layer separation (push) Successful in 22s
Traceability / Requirement traces (push) Successful in 58s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Failing after 33m26s
# Conflicts:
#	docs/traceability.md
#	ui/dr-ui/src/develop.rs
#	ui/dr-ui/src/segmentation.rs
2026-08-27 11:57:38 +02:00
dtourolleandClaude Opus 5 f00b92a0e6 Turn face boxes back into sensor space before matching regions
Faces are found on the thumbnail, which is cached the right way up --
the grid would lie on its side otherwise. Segmentation runs on a proxy
rendered through a neutral edit graph, which carries no orientation and
is therefore in sensor order. For anything shot in portrait the two
differ by a quarter turn, so a face and the person containing it were
being compared in spaces 90 degrees apart: no match, or worse, a match
against somebody else's region.

The transform goes on the face rather than on the proxy. Instance masks
are defined in the proxy's space and sampled long afterwards, so turning
that space would be a far larger change than naming a region warrants.

Also two things the first screenshot of the running app showed that no
test would have:

110 of 23,528 displayed as "0%", which reads as the feature having done
nothing. One decimal below ten percent, and a floor so real progress
never shows as none.

The rail picked some near-black covers, because the largest face in a
group is often the nearest one in a badly lit frame and a black square
beside a name identifies nobody. It now cuts the best few and takes the
first legible one, falling back to the largest when a person's every
photograph is dark -- which happens, and showing it beats showing
nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 23:28:52 +02:00
dtourolleandClaude Opus 5 b55812812a Name segmented people from the faces already recognised in them
The segmenter knows it found a person; the face index knows which person.
Joining them turns "person" in the mask list into "Anna", which is the
difference between a vocabulary of eighty COCO classes and one that
includes the user's family. Selecting a subject in a group photograph
stops being a guessing game between three identical rows.

Containment, not IoU. A face is a small part of the person it belongs to,
so a correct pairing has an IoU near zero and anything IoU-based would
reject every true match.

Confirmed names only. A suggestion is the system's guess, and printing a
guessed name onto a mask region would launder it into a fact.

Writing the tests corrected the design once: a tight head-and-shoulders
portrait, where the face fills most of the person box, is the case where
naming is most certain, not least. An earlier guard rejected exactly that
and has been removed, with the reasoning left as a test because it is
easy to get backwards a second time.

The names hang on the develop session, set when the image opens because
that is the one moment the catalog and the image id are both in reach.
Every segmentation run afterwards picks them up for free, and a library
with no face indexing behaves exactly as it did before.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 23:10:46 +02:00
dtourolleandClaude Opus 5 2ac069a6b3 Index the library's faces, and group them into people
Wires dr-face to dr-catalog: a background sweep that reads the proxy the
grid already built, detects, aligns, embeds and stores, then a clustering
pass that turns those embeddings into suggested people.

Detection runs on the Large thumbnail tier and nowhere else. That is what
makes the feature affordable -- a browsed library has already paid for
its proxies, so face indexing adds no RAW decode that was not already
happening -- and it is why an image whose proxy is missing is skipped
rather than fetched: requesting one here would put face indexing on the
network path FR-CULL-8 keeps it off.

The sweep keeps no cursor. It asks the catalog what is missing, so it
resumes after process death with no repeated work beyond the in-flight
image, and cancelling is dropping the receiver.

recluster writes only the suggested half. Confirmed faces go in as
anchors and come back untouched, and a cluster of one stays nameless --
naming every stray face would fill the People view with noise the user
then has to dismiss.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 20:49:30 +02:00
dtourolleandClaude Opus 5 00e78dc2ac Cluster faces into people, and calibrate what a similarity means
FR-CULL-9 forbids thresholding a bare cosine anywhere in the subsystem,
so calibrate fits P(same person) per library and reports whether the fit
is trustworthy. Two details carry most of the weight.

The fit runs against a 200-bin histogram rather than a pair list: a
25,000-face library has ~3e8 pairs and no gradient descent is running
over that. And a fresh library has no valid calibration, because the
positives have to come from user confirmations or burst siblings --
bootstrapping them from high cosine would fit the calibration to the
belief it was supposed to test.

Clustering defends against the over-merging FR-CULL-10 warns about with
constraints rather than a better threshold: two faces in one photograph
never merge, and two groups confirmed as different people never merge.
Average link rather than single link, so one strong edge cannot weld two
families together.

Calibration is defined once, in dr-face, and dr-catalog re-exports it.
Two implementations of one probability model is exactly how a number
comes to mean the wrong thing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 20:10:33 +02:00
dtourolleandClaude Opus 5 19981c1033 Detect, align and embed faces with SCRFD and MobileFaceNet
Ports the pipeline from the C++ reference in ../scene-actor-extraction
(MIT, same author). End to end on real portraits it separates identities
the way the reference's fitted calibration says it should: 0.596 between
distinct photographs of one person, 0.05 between different people, either
side of MBF's 0.267 boundary.

Three things are structural rather than incidental:

Aligned112 can only be built by align::warp, so Embedder::embed cannot be
handed an unaligned bounding-box crop. That mistake yields 512 plausible
unit-norm numbers and no error, so the type system refuses it instead.

Embedding carries its ModelId and cosine() returns None across models,
because a cross-model similarity is the one mistake that produces
plausible garbage rather than a failure.

The model-free half -- alignment, embedding arithmetic, f16 storage --
sits outside the inference feature and is covered by 11 tests that need
no weights on the machine.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 19:57:56 +02:00
dtourolleandClaude Opus 5 72410f39c6 Answer M1: tract loads both face graphs once their dims are pinned
Neither InsightFace export parses as shipped -- SCRFD fails at its input
node, ArcFace at the first Conv -- which is the same wall dr-segment hit
on YOLO's dynamic export. Both load cleanly with the input dims frozen,
so the pure-Rust runtime holds for the face pipeline too.

tools/fix-face-model-shapes.sh does the freezing, and exists so the
artefact is reproducible rather than a binary someone once produced. It
takes two forms because the two graphs need different ones: ArcFace's
batch is a named dim_param, SCRFD's H and W are dynamic but unnamed.

Also notes YuNet loading with no intervention, which matters for the
licence question in faces.md 2.3.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 19:51:02 +02:00