eb39599f12c54699d45da4792bb89b6e1b1f28a6
6
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1d4c3348af |
Let a 32-pixel face count, and move the blur floor with it
64 source pixels was too strict: it threw away 70% of everything the detector
finds, and plenty of what it took were faces a person could name.
Lowering it is not a one-line change, because the two floors are coupled. A face
under 112 pixels is *upsampled* to reach the embedder and upsampling invents no
edges, so a small face scores low on sharpness however crisp the original was.
Re-measured over the reference library with `face_index --quality`:
min crop min sharp size cut blur cut kept
32 0.000 44% 0% 56%
32 0.002 44% 3% 53%
32 0.005 44% 8% 48%
32 0.010 44% 16% 40%
32 0.020 44% 27% 30%
64 0.020 70% 7% 23%
Holding the blur floor at 0.020 while dropping the size floor to 32 would have
rejected a further 27% — for being small rather than for being blurred — and
kept only 30%, barely more than the 23% the strict pair kept. Most of the point
of lowering the size floor would have gone straight back out through the other
gate.
0.005 removes 8% of what the size floor leaves, which is the same job 0.020 was
doing at 64 (7%): the large-but-soft face this gate exists for. Together they
now keep 48% of what the detector finds, against 23% before.
The box pre-filter follows down to 24, staying below what the real floor accepts
so it cannot reject a face that would have cleared 32.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
a1790c3e67 |
Stop indexing faces too small or too blurred to be anyone
The library was storing faces at 52 source pixels and embedding whatever came
back. There was a size floor, but it was 40 pixels on the *bounding box*, and
there was no blur gate at all — so a subject walking through a half-second
exposure detected confidently, aligned cleanly, and produced a perfectly
ordinary-looking 512-vector. Nothing downstream can tell that apart from a real
face, and because blurs resemble each other more than they resemble the people
they were, they cluster together and weld unrelated identities into one group.
Two floors, both measured rather than guessed. `face_index --quality` runs the
detector over real proxies with both gates disabled and prints the distribution;
over 1,503 faces in 600 images of the reference library:
percentile crop px sharpness
1% 16 0.0006
25% 23 0.0025
50% 38 0.0071
75% 76 0.0284
99% 352 0.4282
The median face in a personal library is 38 pixels. Most of what the detector
finds is background: people across a square, a face on a poster, a stranger at
the next table. They are real detections and useless identifications.
**Size, on the crop rather than the box.** "At least 64x64" has to mean the
pixels the *embedder* sees, and the box is not that — the ArcFace template
reaches past it for forehead and chin, so the aligned crop spans roughly 1.3x
the box's shorter edge. The floor is therefore `min_source_px` on the aligned
crop, applied after the warp fixes the scale, and `min_face_px` drops to 48 as
what it always really was: a cheap pre-filter set low enough that it cannot
reject a face the real floor would have kept.
**Sharpness.** Variance of the Laplacian divided by the variance of the luma it
was taken over. The division is the part that matters: raw Laplacian variance
scales with contrast, so a threshold on it would quietly discard every backlit
portrait in the library. The ratio asks how much of the crop's variation is
edges rather than broad gradients, and is invariant to exposure.
What each pair removes, cumulatively, of everything the detector finds:
min crop min sharp size cut blur cut kept
64 0.000 70% 0% 30%
64 0.010 70% 3% 27%
64 0.020 70% 7% 23%
80 0.010 76% 2% 21%
64 and 0.020. The size floor does most of the work, and the blur floor removing
only 7% on top of it is the point rather than a disappointment: at 64 pixels
most faces are already sharp, and what it takes out is the large-but-soft one —
precisely the face that would otherwise contribute a confident, wrong embedding.
The two gates are not independent and the doc comments say so: a face under 112
pixels was upsampled to reach the embedder, and upsampling invents no edges, so
small faces score low on sharpness even when the original was crisp. That is why
`--quality` prints them together.
**This will re-index.** Around 70% of what the current settings store falls below
the new floors — faces between 20 and 40 pixels that nobody could identify. The
People screen gets shorter and every group in it gets better.
66 dr-face tests pass, including that a blurred crop scores below a sharp one,
that halving the contrast does not move the score, and that an upsampled face
scores below the same face at full size.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
7275c020d7 |
Group people at the threshold the library actually supports
0.90 left a third of the reference library ungrouped: 1,213 of 1,813 faces in a
group, and the rest sitting alone in a screen that had nothing to offer for
them.
"Is 0.90 too tight" is not answerable from the number. It is a probability, and
which cosine it lands on depends on the calibration — so the first half of this
is a way to ask the question properly. `face_index --tune` runs the real
clusterer over the real embeddings at ten thresholds and prints what each one
produces. It writes nothing; comparing thresholds by applying them would have
each one pollute the next.
On the reference library:
P cosine groups grouped largest
0.95 0.449 311 62% 51
0.90 0.403 316 67% 51
0.85 0.374 318 70% 57
0.80 0.353 328 74% 69
0.75 0.335 327 77% 69
0.70 0.319 326 79% 81
0.50 0.267 303 85% 90
The count of *groups* is the signal, not the count of grouped faces. Loosening
from 0.95 makes it climb: real people are being assembled out of fragments. It
peaks at 0.80 and then falls — and a falling group count while the grouped faces
keep rising is the shape of over-merging, separate identities being welded
together. That is the FR-CULL-10 failure, and the one the user cannot undo by
hand.
So 0.80: the loosest setting still building people rather than melting them
together. A third more of the library gets grouped than at 0.90, and the largest
group grows by eighteen faces rather than by forty.
The table is one library, and the doc comment says so — `--tune` reruns it on
any other.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
c8c6368542 |
Index the whole library by fetching what it has not seen
"Index faces in the whole library" could not. Its work list was intersected
with the thumbnail store at `ThumbSize::Large`, and nothing fills that class for
a whole library — `SWEEP_THUMB_SIZE` is deliberately `Grid`, because the large
class is ~860 MB of shards against ~200 MB and every syncing device pays it. So
the only images with a large proxy were the ones the user had personally zoomed
into or opened in the loupe. On this library that was 220 of 23,529.
The comment defending it misread the requirement:
// Requesting one here would put face indexing on the network path,
// which FR-CULL-8 explicitly keeps it off.
FR-CULL-8 keeps indexing off the **full decode**, not the network, and then says
the opposite in the same paragraph: "where no proxy exists, the job requests one
at background priority rather than decoding inline". faces.md §7 repeats it.
Neither was implemented.
So the pass fetches. Same two-stage route the thumbnail sweep uses — the header,
then the located preview's own byte range (FR-NC-3) — so no whole file is pulled
and no RAW is decoded, because an embedded preview is a JPEG. The work list is
now every visible image with no `face_index` row for the model: 23,308 here,
against nearly none before.
**It indexes at the resolution the preview actually has**, not the 1024 the old
tier would have given. `locate_preview` already picks the largest embedded
preview, and the thumbnail sweep was decoding it and throwing the detail away at
`downscale_to(256)`. A face 2% across the frame is 5 px on a grid thumbnail and
~61 px at the cap here — and 112 is what the embedder samples, so this is the
difference between an upsampled crop and a real one. `crop_px` records which,
per face, as §7 intended.
Capped at 3072 rather than truly full: `index_proxy` needs packed `f32` RGB at
12 bytes a pixel, so a 24 MP frame is ~288 MB and the fetch lanes hold one each.
The constant is named and sits next to the reason.
Orientation is applied **before** detection, not after downscaling. That costs a
permutation of a larger buffer — ~15 ms against a ~150 ms decode — and buys the
entire class of bug this codebase keeps having: detection then runs on the
photograph rather than the sensor, so every box and landmark is already in the
space the catalog stores and the overlay draws, with no second mapping to get
backwards.
One detector and one embedder serve every lane. The lanes are concurrent futures
on a single thread, not threads, and inference contains no await, so a `RefCell`
borrow never overlaps another — a pair per lane would duplicate ~16 MB of
weights for no parallelism.
Images with no face in them are recorded too. `face_index` records that
detection *ran*, and zero is its most valuable value: without the row every
landscape and document scan returns on every pass, for ever, and in a personal
library that is most of it (§7a).
The old store-only pass survives as `spawn_store_face_sweep` for
`examples/face_index.rs`, which indexes a local store with no network. The
settings copy no longer claims indexing reads "the photographs already
thumbnailed above", and the audit line says "to fetch" rather than "awaiting a
proxy", which had become a blocker that no longer blocks.
Verified against the real catalog: the new work list returns 23,308 where the
old one returned effectively nothing. 469 tests pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
96d07da15f |
Sync face data as sealed shards, so a second device does not re-index
Indexing 23,500 images is about two hours of CPU, and the result is byte-identical on every device: the same model over the same proxy produces the same embedding. Paying for it once per account rather than once per device is the point. Shards rather than the catalog snapshot, because the snapshot goes up whole on every sync and a fully indexed library carries roughly 30 MB of embeddings. That is exactly the cost the thumbnail store's 25 MB cap exists to bound, so face shards use the same cap -- imported from dr_thumbs rather than restated, since the number is a statement about sync cost and the two must not drift apart. The split follows the one already there: bulk immutable data in sealed shards, small mutable data in the catalog snapshot. Faces, landmarks, embeddings and run markers shard; people, names and assignments ride the catalog and merge by uuid. Keyed on oc:fileid throughout, never on image_id, because a row id means nothing on another device. The run marker travels with the faces it describes. Without it a receiving device cannot tell an image with no faces from one never examined, and would re-detect every landscape it had just adopted. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
26a1eb7e28 |
Record that face detection has run, not just what it found
An image with no faces in it was indistinguishable from one that had never been looked at, so every landscape, still life and document scan in the library was re-detected on every pass, for ever. In a real library that is most of it: on the 23,527-image test library, 64 of the first 110 images indexed contain no face at all. Schema v9 adds face_index, a run marker per (image, model) carrying the face count and the proxy edge it read. Keyed on the model, so a model change puts every image back in the queue by itself. That makes a coverage figure possible, which is the thing a user actually wants to see. The audit also splits the outstanding set by whether a proxy exists, because 23,417 awaiting a proxy and 110 ready to index are different problems, and telling the user to run indexing again would not fix the first. The Identity screen gains Index faces, Stop, and the coverage line. examples/face_index.rs is the same check and sweep without a window, which is the right shape for an overnight pass. Measured on the real library in release: 3.5 images/second, 110 images and 125 faces in 30 seconds, and a second run correctly finds nothing left to do. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |