Commit Graph
12 Commits
Author SHA1 Message Date
dtourolleandClaude Opus 5 da20d42d33 Merge master: pluggable storage, and a name that anchors
Conflicts were docs/traceability.md alone, and it is generated — so it
was regenerated rather than hand-merged. dr-face was untouched on the
other side; ui/dr-ui/src/faces.rs and identity_ui.rs auto-merged, the
first around recluster's anchoring and the second around load_faces.

Worth recording because the two branches met on the same problem from
different ends. Master's "Let a name hold a group together" is the fix
for the sixteen Catherines — fourteen of them empty — that this branch
found while measuring the library and reported without fixing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 12:30:01 +02:00
dtourolleandClaude Opus 5 ebb7d3cf5c Score a suggestion against the people the user has named
The number beside a suggestion was the mean calibrated probability
between the face and the rest of its group, which measures the wrong
thing twice. It punishes coverage: a person with two hundred faces over
fifteen years is *meant* to have members a given photograph is
orthogonal to, so a correct suggestion onto a well-photographed person
scored low for being well photographed. And it never asked who else the
face might be — a face matching Anna at 0.95 and nobody else, and one
matching Anna at 0.95 and her sister at 0.93, came out identical, when
the second is the only one worth the user's attention.

dr_face::assign answers both, and multiplies them: the mean of the best
ten calibrated matches into the identity (the old mean, capped, which is
what stops coverage counting against it), times that identity's share of
the evidence against every *named* rival.

Only named people compete, and per person rather than per group. Both
halves of that had to be measured on a real 18,000-face library rather
than reasoned about. Normalising across every group made the number
useless — median suggestion 21%, four in five under half — because
clustering leaves one person spread over many groups, so a face competed
against itself; and keying rivals by group left Catherine competing with
Catherine, median 39%. Per named person: median 99.5%.

Rivals are gathered below the merge threshold, down to even odds: a
named person matching at 0.6 will never be merged into but is exactly
the competition to discount for. That would be a second similarity scan,
the expensive half of regrouping a library, so cluster_scored scans once
at the looser floor and hands the merge engine the subset at or above
the threshold — pair for pair what it would have scanned for itself,
held to that by a test.

Leave-one-out over that library's 2,702 confirmations across 54 named
people: 99.33% of faces placed on the right person against the old
mean's 99.15%, and the number shown for the right person moves from a
median of 90.4% to 99.3%. It errs low — 100% correct wherever it states
80% or more — which is the safe direction, and docs/faces.md §9.1 says
plainly that the low bands are not calibrated.

The example that measures it comes too: this is a claim about a
library's numbers, and nobody should have to take it on faith.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 10:44:01 +02:00
dtourolleandClaude Opus 5 1b8b7998a2 Let a name hold a group together, the way a confirmation does
Build and test / Desktop (Linux) (push) Successful in 2h6m43s
Build and test / Layer separation (push) Successful in 46s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 2s
Traceability / Requirement traces (push) Successful in 29s
Build and test / Android (aarch64) (push) Successful in 21m27s
Sixteen people called Catherine, fourteen of them holding no faces at all, and
her actual photographs split across the two that did. That is not a sync fault;
it is this function, and it has been quietly doing it on every Regroup.

Naming a cluster does not confirm its faces. They stay *suggestions* — and only
confirmed faces anchored here, so the next pass cut them loose, regrouped them
into a brand new person, and left the named one holding nothing.
`prune_empty_unnamed` will not clean that up, because it has a name. Name the new
group the same thing, and it happens again. Repeat over a few sessions and you
have sixteen of her.

A name is a judgement about *this group*, of exactly the kind FR-CULL-12 says
travels and inference does not — the same argument that already anchors a group
the user set aside. So all three kinds of ruling anchor now: confirmed, ignored,
and named.

It also fixes the quieter half of the same fault. A face indexed later that
matches a named person now merges *into* them, rather than arriving as a rival
group the user has to name all over again.

Two tests, and the first fails without the change — it reports "Catherine"
finishing the pass with zero faces while a fresh unnamed person holds the two
she was named for.

This does not retro-fit an existing library: the fourteen empty Catherines stay
until they are merged by hand, and the two holding faces are separate identities
that only the user can say are one person. What it stops is making more.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 09:48:32 +02:00
dtourolleandClaude Opus 5 41daa4cca7 Keep a group set aside actually set aside
"Not interested" hid a group, and the next Regroup brought it straight back.

Reclustering anchors the faces the user has ruled on so a pass cannot move them.
It took *confirmations* as the only kind of ruling — but setting a group aside is
a ruling too, and the faces it covers are only ever suggestions. So an ignored
group's faces entered clustering loose, regrouped into a fresh person that
carried no ignore flag, and reappeared in the rail. The original group was left
behind holding nothing, hidden and empty.

Anchoring them on `ignored` as well as on `confirmed` fixes it, and does one
better: a face indexed later that matches a group which was set aside now merges
*into* it, so a stranger photographed again stays set aside instead of arriving
as somebody new. That is the case that would otherwise have made the feature
feel like it only half worked.

Three tests, and the first fails without the change — it reports the group
coming back with its two faces while the original sits ignored and empty.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 08:30:16 +02:00
dtourolleandClaude Opus 5 a1790c3e67 Stop indexing faces too small or too blurred to be anyone
The library was storing faces at 52 source pixels and embedding whatever came
back. There was a size floor, but it was 40 pixels on the *bounding box*, and
there was no blur gate at all — so a subject walking through a half-second
exposure detected confidently, aligned cleanly, and produced a perfectly
ordinary-looking 512-vector. Nothing downstream can tell that apart from a real
face, and because blurs resemble each other more than they resemble the people
they were, they cluster together and weld unrelated identities into one group.

Two floors, both measured rather than guessed. `face_index --quality` runs the
detector over real proxies with both gates disabled and prints the distribution;
over 1,503 faces in 600 images of the reference library:

   percentile   crop px   sharpness
           1%        16      0.0006
          25%        23      0.0025
          50%        38      0.0071
          75%        76      0.0284
          99%       352      0.4282

The median face in a personal library is 38 pixels. Most of what the detector
finds is background: people across a square, a face on a poster, a stranger at
the next table. They are real detections and useless identifications.

**Size, on the crop rather than the box.** "At least 64x64" has to mean the
pixels the *embedder* sees, and the box is not that — the ArcFace template
reaches past it for forehead and chin, so the aligned crop spans roughly 1.3x
the box's shorter edge. The floor is therefore `min_source_px` on the aligned
crop, applied after the warp fixes the scale, and `min_face_px` drops to 48 as
what it always really was: a cheap pre-filter set low enough that it cannot
reject a face the real floor would have kept.

**Sharpness.** Variance of the Laplacian divided by the variance of the luma it
was taken over. The division is the part that matters: raw Laplacian variance
scales with contrast, so a threshold on it would quietly discard every backlit
portrait in the library. The ratio asks how much of the crop's variation is
edges rather than broad gradients, and is invariant to exposure.

What each pair removes, cumulatively, of everything the detector finds:

    min crop   min sharp   size cut   blur cut       kept
          64       0.000        70%         0%        30%
          64       0.010        70%         3%        27%
          64       0.020        70%         7%        23%
          80       0.010        76%         2%        21%

64 and 0.020. The size floor does most of the work, and the blur floor removing
only 7% on top of it is the point rather than a disappointment: at 64 pixels
most faces are already sharp, and what it takes out is the large-but-soft one —
precisely the face that would otherwise contribute a confident, wrong embedding.

The two gates are not independent and the doc comments say so: a face under 112
pixels was upsampled to reach the embedder, and upsampling invents no edges, so
small faces score low on sharpness even when the original was crisp. That is why
`--quality` prints them together.

**This will re-index.** Around 70% of what the current settings store falls below
the new floors — faces between 20 and 40 pixels that nobody could identify. The
People screen gets shorter and every group in it gets better.

66 dr-face tests pass, including that a blurred crop scores below a sharp one,
that halving the contrast does not move the score, and that an upsampled face
scores below the same face at full size.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 22:17:38 +02:00
dtourolleandClaude Opus 5 d6380fecc8 Make the People screen a place work can be done
Five faults, all on one screen, and the Slint and Rust halves of each have to
land together.

**Regroup froze the window.** It ran inside the Slint callback, on the UI
thread. It is much faster now, but fast is not bounded — the work grows with the
library, and the one thing that must not grow with the library is how long the
window stops answering. It runs on a worker thread with an mpsc channel and a
250ms poll, like every other long pass in this module, and the button says what
it is doing instead of the window going quiet. Cancellation is dropping the
receiver. Reclustering also prunes the empty groups the previous pass left, so
pressing the button twice no longer fills the rail with "Unnamed (0 faces)".

**The faces were a single row running off the screen.** The comment on the
layout claimed to be a wrapping row; Slint has no flow layout and a
HorizontalLayout does not wrap, so a person with forty faces was a person whose
faces could not be reviewed past the fifth. It is now laid out the way the
library grid lays out thumbnails, with the same arithmetic: choose how many
columns of roughly the requested size fit, then divide the width between them so
the cells fill the row exactly and nothing overhangs.

**The header did not fit a phone.** A 240px name field beside five buttons is
wider than an Android screen — and worse than not fitting, a layout cannot be
narrower than its children's minimums, so the row reported that oversized
minimum upwards and inflated the whole screen. The faces grid is its sibling, so
it would have been measured against a width that was never on the display. The
header is now two rows, the actions sit in a Flickable that scrolls rather than
overflowing, and the rail narrows to 132px on the compact class.

**Strangers crowded out the people who matter.** Most clusters in a real library
are passers-by and other people's guests. "Not interested" sets a group aside;
the rail hides it and says how many are hidden, with one button to bring them
back. Reversible, and never a deletion — see the catalog commit for why.

**A face was a dead end.** Identifying someone and then having no way to see
their photographs is a filing cabinet with no drawer handles. "Show photos"
narrows the library grid to that person and leaves a chip on the filter bar
saying so, which is also how it is cleared. It is a term on `RatingFilter`
rather than a grid scope of its own, exactly as that struct's own doc says new
narrowing terms should be — so the count and the cells are narrowed by the same
thing, and it composes with the others for free. Suggested faces count, not only
confirmed ones, or a freshly grouped person would show an empty grid.

Crops are read from where they are now stored, falling back to cutting one out
of the proxy for faces indexed before that existed.

480 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 21:21:31 +02:00
dtourolleandClaude Opus 5 c8c6368542 Index the whole library by fetching what it has not seen
"Index faces in the whole library" could not. Its work list was intersected
with the thumbnail store at `ThumbSize::Large`, and nothing fills that class for
a whole library — `SWEEP_THUMB_SIZE` is deliberately `Grid`, because the large
class is ~860 MB of shards against ~200 MB and every syncing device pays it. So
the only images with a large proxy were the ones the user had personally zoomed
into or opened in the loupe. On this library that was 220 of 23,529.

The comment defending it misread the requirement:

    // Requesting one here would put face indexing on the network path,
    // which FR-CULL-8 explicitly keeps it off.

FR-CULL-8 keeps indexing off the **full decode**, not the network, and then says
the opposite in the same paragraph: "where no proxy exists, the job requests one
at background priority rather than decoding inline". faces.md §7 repeats it.
Neither was implemented.

So the pass fetches. Same two-stage route the thumbnail sweep uses — the header,
then the located preview's own byte range (FR-NC-3) — so no whole file is pulled
and no RAW is decoded, because an embedded preview is a JPEG. The work list is
now every visible image with no `face_index` row for the model: 23,308 here,
against nearly none before.

**It indexes at the resolution the preview actually has**, not the 1024 the old
tier would have given. `locate_preview` already picks the largest embedded
preview, and the thumbnail sweep was decoding it and throwing the detail away at
`downscale_to(256)`. A face 2% across the frame is 5 px on a grid thumbnail and
~61 px at the cap here — and 112 is what the embedder samples, so this is the
difference between an upsampled crop and a real one. `crop_px` records which,
per face, as §7 intended.

Capped at 3072 rather than truly full: `index_proxy` needs packed `f32` RGB at
12 bytes a pixel, so a 24 MP frame is ~288 MB and the fetch lanes hold one each.
The constant is named and sits next to the reason.

Orientation is applied **before** detection, not after downscaling. That costs a
permutation of a larger buffer — ~15 ms against a ~150 ms decode — and buys the
entire class of bug this codebase keeps having: detection then runs on the
photograph rather than the sensor, so every box and landmark is already in the
space the catalog stores and the overlay draws, with no second mapping to get
backwards.

One detector and one embedder serve every lane. The lanes are concurrent futures
on a single thread, not threads, and inference contains no await, so a `RefCell`
borrow never overlaps another — a pair per lane would duplicate ~16 MB of
weights for no parallelism.

Images with no face in them are recorded too. `face_index` records that
detection *ran*, and zero is its most valuable value: without the row every
landscape and document scan returns on every pass, for ever, and in a personal
library that is most of it (§7a).

The old store-only pass survives as `spawn_store_face_sweep` for
`examples/face_index.rs`, which indexes a local store with no network. The
settings copy no longer claims indexing reads "the photographs already
thumbnailed above", and the audit line says "to fetch" rather than "awaiting a
proxy", which had become a blocker that no longer blocks.

Verified against the real catalog: the new work list returns 23,308 where the
old one returned effectively nothing. 469 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 18:51:53 +02:00
dtourolleandClaude Opus 5 b846b312b8 Run the formatter over the face branch before it reaches CI
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 2s
Build and test / Desktop (Linux) (push) Successful in 1h21m32s
Build and test / Layer separation (push) Successful in 37s
Traceability / Requirement traces (push) Successful in 25s
Build and test / Android (aarch64) (push) Failing after 33m58s
The merge of the SCRFD/MobileFaceNet work brought 69 rustfmt diffs across
dr-catalog, dr-face and dr-ui with it, so `cargo fmt --all -- --check` fails
on master and the Desktop job stops at its Format step — before clippy, the
tests or the release build have run at all. That makes the whole desktop
half of CI blind: a real compile error behind this would look exactly the
same from the outside. There was nothing behind it, as it turns out — with
the formatting fixed, clippy, the test suite and the release build all pass.

Every .rs hunk is `cargo fmt --all` on the pinned 1.92.0 toolchain, not a
hand edit, but it is worth being precise about what that moved, because it
is more than whitespace. Besides reflowing signatures and call chains,
rustfmt reordered the `pub mod` and `pub use` items in dr-face/src/lib.rs so
the `#[cfg(feature = "inference")]` entries sort in place, added the trailing
semicolon inside `let ... else { return }` bodies in identity_ui.rs, wrapped
a bare closure body in braces in cluster.rs, adjusted trailing commas, and
dropped a stray blank line at the end of identity_ui.rs. All of it is
semantically inert; none of it changes behaviour.

docs/traceability.md rides along because it has to. The matrix records each
TRACES tag by line number, and reflowing develop.rs, lib.rs, faces.rs,
identity.rs and identity_ui.rs moved them — FR-CAT-8, FR-CAT-9, FR-CULL-10,
FR-DEV-3, FR-DEV-3a and FR-DEV-3c all shift by a line or two. The matrix was
verified up to date on d777f7f before this commit, so this is drift these
formatting changes introduced, not pre-existing staleness being swept up.
Leaving it for a follow-up commit would hand traceability-check.yml a
failure caused entirely by a whitespace change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 13:12:49 +02:00
dtourolleandClaude Opus 5 f00b92a0e6 Turn face boxes back into sensor space before matching regions
Faces are found on the thumbnail, which is cached the right way up --
the grid would lie on its side otherwise. Segmentation runs on a proxy
rendered through a neutral edit graph, which carries no orientation and
is therefore in sensor order. For anything shot in portrait the two
differ by a quarter turn, so a face and the person containing it were
being compared in spaces 90 degrees apart: no match, or worse, a match
against somebody else's region.

The transform goes on the face rather than on the proxy. Instance masks
are defined in the proxy's space and sampled long afterwards, so turning
that space would be a far larger change than naming a region warrants.

Also two things the first screenshot of the running app showed that no
test would have:

110 of 23,528 displayed as "0%", which reads as the feature having done
nothing. One decimal below ten percent, and a floor so real progress
never shows as none.

The rail picked some near-black covers, because the largest face in a
group is often the nearest one in a badly lit frame and a black square
beside a name identifies nobody. It now cuts the best few and takes the
first legible one, falling back to the largest when a person's every
photograph is dark -- which happens, and showing it beats showing
nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 23:28:52 +02:00
dtourolleandClaude Opus 5 26a1eb7e28 Record that face detection has run, not just what it found
An image with no faces in it was indistinguishable from one that had
never been looked at, so every landscape, still life and document scan in
the library was re-detected on every pass, for ever. In a real library
that is most of it: on the 23,527-image test library, 64 of the first 110
images indexed contain no face at all.

Schema v9 adds face_index, a run marker per (image, model) carrying the
face count and the proxy edge it read. Keyed on the model, so a model
change puts every image back in the queue by itself.

That makes a coverage figure possible, which is the thing a user actually
wants to see. The audit also splits the outstanding set by whether a
proxy exists, because 23,417 awaiting a proxy and 110 ready to index are
different problems, and telling the user to run indexing again would not
fix the first.

The Identity screen gains Index faces, Stop, and the coverage line.
examples/face_index.rs is the same check and sweep without a window,
which is the right shape for an overnight pass.

Measured on the real library in release: 3.5 images/second, 110 images
and 125 faces in 30 seconds, and a second run correctly finds nothing
left to do.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 22:13:41 +02:00
dtourolleandClaude Opus 5 10b10569e2 Add the Identity screen
A third top-level screen beside the library and develop, because naming a
cluster and pulling a stranger out of it are tasks with their own rhythm
and need the whole window.

The screen is designed around the clustering being wrong, which is
FR-CULL-10 rather than pessimism: grouping over-merges on siblings, on
parents and children, and on the same person a decade apart. So Split off
sits next to Confirm all rather than behind a menu, the confirm/reject
pair is on the face itself, and a group the system found is drawn
differently from a person the user has vouched for.

Splitting rejects before it confirms. Without that the next clustering
pass suggests the face straight back and the user's correction becomes an
argument they keep having.

Face crops come from the proxies the grid already built, one decode per
image rather than per face -- a group photograph holding six faces of one
family is one JPEG.

Where the calibration is not fitted the screen says confidence is
unavailable instead of printing a percentage that looks measured, which
is FR-CULL-9's rule at the point it becomes visible.

The verdict controls use drawn icons, not tick and cross characters:
ui/icons.slint exists because those render as tofu on Android.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 21:05:41 +02:00
dtourolleandClaude Opus 5 2ac069a6b3 Index the library's faces, and group them into people
Wires dr-face to dr-catalog: a background sweep that reads the proxy the
grid already built, detects, aligns, embeds and stores, then a clustering
pass that turns those embeddings into suggested people.

Detection runs on the Large thumbnail tier and nowhere else. That is what
makes the feature affordable -- a browsed library has already paid for
its proxies, so face indexing adds no RAW decode that was not already
happening -- and it is why an image whose proxy is missing is skipped
rather than fetched: requesting one here would put face indexing on the
network path FR-CULL-8 keeps it off.

The sweep keeps no cursor. It asks the catalog what is missing, so it
resumes after process death with no repeated work beyond the in-flight
image, and cancelling is dropping the receiver.

recluster writes only the suggested half. Confirmed faces go in as
anchors and come back untouched, and a cluster of one stays nameless --
naming every stray face would fill the People view with noise the user
then has to dismiss.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 20:49:30 +02:00