Commit Graph
5 Commits
Author SHA1 Message Date
dtourolleandClaude Opus 5 b55812812a Name segmented people from the faces already recognised in them
The segmenter knows it found a person; the face index knows which person.
Joining them turns "person" in the mask list into "Anna", which is the
difference between a vocabulary of eighty COCO classes and one that
includes the user's family. Selecting a subject in a group photograph
stops being a guessing game between three identical rows.

Containment, not IoU. A face is a small part of the person it belongs to,
so a correct pairing has an IoU near zero and anything IoU-based would
reject every true match.

Confirmed names only. A suggestion is the system's guess, and printing a
guessed name onto a mask region would launder it into a fact.

Writing the tests corrected the design once: a tight head-and-shoulders
portrait, where the face fills most of the person box, is the case where
naming is most certain, not least. An earlier guard rejected exactly that
and has been removed, with the reasoning left as a test because it is
easy to get backwards a second time.

The names hang on the develop session, set when the image opens because
that is the one moment the catalog and the image id are both in reach.
Every segmentation run afterwards picks them up for free, and a library
with no face indexing behaves exactly as it did before.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 23:10:46 +02:00
dtourolleandClaude Opus 5 2ac069a6b3 Index the library's faces, and group them into people
Wires dr-face to dr-catalog: a background sweep that reads the proxy the
grid already built, detects, aligns, embeds and stores, then a clustering
pass that turns those embeddings into suggested people.

Detection runs on the Large thumbnail tier and nowhere else. That is what
makes the feature affordable -- a browsed library has already paid for
its proxies, so face indexing adds no RAW decode that was not already
happening -- and it is why an image whose proxy is missing is skipped
rather than fetched: requesting one here would put face indexing on the
network path FR-CULL-8 keeps it off.

The sweep keeps no cursor. It asks the catalog what is missing, so it
resumes after process death with no repeated work beyond the in-flight
image, and cancelling is dropping the receiver.

recluster writes only the suggested half. Confirmed faces go in as
anchors and come back untouched, and a cluster of one stays nameless --
naming every stray face would fill the People view with noise the user
then has to dismiss.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 20:49:30 +02:00
dtourolleandClaude Opus 5 00e78dc2ac Cluster faces into people, and calibrate what a similarity means
FR-CULL-9 forbids thresholding a bare cosine anywhere in the subsystem,
so calibrate fits P(same person) per library and reports whether the fit
is trustworthy. Two details carry most of the weight.

The fit runs against a 200-bin histogram rather than a pair list: a
25,000-face library has ~3e8 pairs and no gradient descent is running
over that. And a fresh library has no valid calibration, because the
positives have to come from user confirmations or burst siblings --
bootstrapping them from high cosine would fit the calibration to the
belief it was supposed to test.

Clustering defends against the over-merging FR-CULL-10 warns about with
constraints rather than a better threshold: two faces in one photograph
never merge, and two groups confirmed as different people never merge.
Average link rather than single link, so one strong edge cannot weld two
families together.

Calibration is defined once, in dr-face, and dr-catalog re-exports it.
Two implementations of one probability model is exactly how a number
comes to mean the wrong thing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 20:10:33 +02:00
dtourolleandClaude Opus 5 19981c1033 Detect, align and embed faces with SCRFD and MobileFaceNet
Ports the pipeline from the C++ reference in ../scene-actor-extraction
(MIT, same author). End to end on real portraits it separates identities
the way the reference's fitted calibration says it should: 0.596 between
distinct photographs of one person, 0.05 between different people, either
side of MBF's 0.267 boundary.

Three things are structural rather than incidental:

Aligned112 can only be built by align::warp, so Embedder::embed cannot be
handed an unaligned bounding-box crop. That mistake yields 512 plausible
unit-norm numbers and no error, so the type system refuses it instead.

Embedding carries its ModelId and cosine() returns None across models,
because a cross-model similarity is the one mistake that produces
plausible garbage rather than a failure.

The model-free half -- alignment, embedding arithmetic, f16 storage --
sits outside the inference feature and is covered by 11 tests that need
no weights on the machine.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 19:57:56 +02:00
dtourolleandClaude Opus 5 72410f39c6 Answer M1: tract loads both face graphs once their dims are pinned
Neither InsightFace export parses as shipped -- SCRFD fails at its input
node, ArcFace at the first Conv -- which is the same wall dr-segment hit
on YOLO's dynamic export. Both load cleanly with the input dims frozen,
so the pure-Rust runtime holds for the face pipeline too.

tools/fix-face-model-shapes.sh does the freezing, and exists so the
artefact is reproducible rather than a binary someone once produced. It
takes two forms because the two graphs need different ones: ArcFace's
batch is a named dim_param, SCRFD's H and W are dynamic but unnamed.

Also notes YuNet loading with no intervention, which matters for the
licence question in faces.md 2.3.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 19:51:02 +02:00