44ea763c619dcc38bc06b76ba7d3a63301abe826
11
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
83f4253b6a |
Filter the grid to a person with their eyes open
An "Eyes open" chip beside the people chips, offered only while someone is chosen and dropped when the last person goes, so no term narrows the grid with nothing on the bar to say so. It compiles the rule in dr_face::eyes into the person's face subquery — Anna, eyes open, whoever else is blinking beside her — and drops a frame only on a closed eye that could be read: sunglasses, eyes too small or soft to read, and faces never read all pass, so an old library shows everything under the chip until the measuring pass has run. A test drives the same readings through the SQL and through the rule and requires them to agree. The People screen badges a face "Eyes closed", "Sunglasses" or "Eyes unclear" so the reason a frame is or is not in the grid can be read off the face; the sweep loads the three models when they are beside the pair and reads eyes on the indexing and measuring passes from the native render; the coverage line counts unread faces as work to measure so an already-indexed library keeps its Index button. The term travels with the place. |
||
|
|
8b3abdb787 |
Keep each face's quality, and never compare against a poor one
The embedder's raw output has a length, and the length is a reading of how recognisable the crop was: a blur, an occlusion or a hard profile comes out short. Normalising threw it away. A short vector sits near the middle of the sphere and matches a little of everyone, which is how one bad crop bridges two people in a grouping pass. So the length is kept — the store now holds the raw vector, re-normalised on load, with the length beside it as `faces.quality` — and a face under MIN_GALLERY_QUALITY (14) is a probe: measured against the gallery and placed where it fits, but never what another face is measured against. Two probes are never paired, and a probe is nobody's evidence for a confidence. The People screen shows the number as "Quality 17.3", dimmed below the floor. Faces indexed before this stored unit vectors and have no reading; they are admitted to the gallery, and schema V14 forgets the run marker of every image holding one so the next indexing pass measures them. A peer's unmeasured shard faces are not adopted, or a sync would write that marker back. |
||
|
|
9ddc1273c0 |
Make a sweep that fails everything say so
A run over 169 images failed all 169, in fourteen seconds, and reported
"0 face(s) in 0 image(s)" -- the same sentence a run that indexed
nothing because there was nothing to index produces. Three separate
places dropped the information on its way to the screen.
The progress count only moved on success. FaceSweepMessage had no
failure variant at all, so a pass where every image failed sat at 0/169
from the first tick to the last: the receiver was told the total, told
nothing, and told the pass had ended. That is indistinguishable from a
hung job, and it is what it was taken for.
Finished already carried a failed count and identity_ui matched it with
`Finished { .. }`, throwing the number away and printing the tidy
success line regardless.
And the reason each image failed was logged at debug, which is off, so
169 consecutive failures left no trace of why anywhere.
Failed { images } now carries the count back per lane batch, the
progress counter advances on it, and both the running status line and
the finishing activity row say how many could not be read. A batch
rather than one message per image because failures come back lane-sized
and the useful number is how many.
Also renames the store sweep's guard to MIN_CROP_EDGE with the rest of
that constant's move, since the two touch the same lines.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
e7b526c550 |
Specify face indexing at native resolution, and say what the proxy cost
FR-CULL-8 said detection runs against the thumbnail or proxy tier and never a full decode, and faces.md §5 said the aligned crop is sampled from that same proxy. Both are wrong in the same place: they treat detection and cropping as one resolution problem when they are two, with opposite answers. Detection does not care. §4.1 fixes the graph's input at 640x640 and letterboxes whatever arrives, so a face filling 2% of the frame reaches the model at 12px whether the buffer handed over is 1024px or 6000px. Every pixel above the detector's own input is discarded before inference. The crop cares about nothing else. §5's warp produces the fixed 112x112 ArcFace sees, so source resolution converts directly into whether those 112 pixels were photographed or interpolated. Reading crop_px across the 18,671 faces the proxy-tier implementation stored: 47.3% were upsampled to reach the embedder, 314 of them by more than 2x, the smallest from 34 source pixels. An upsampled crop does not fail loudly -- it yields a confident embedding of detail that was never there, and the damage appears three stages later as clusters that will not separate. So FR-CULL-8 now specifies four stages with the resolutions named separately: render native through FR-EXP-9's pipeline, downscale for the detector, map boxes and landmarks back to native, crop and align from the native render. The affordability the old rule bought is met instead by when the pass runs -- background, preempted, resumable -- and the requirement says plainly what it now costs on a remote library: the original rather than FR-NC-3's byte range, 412 GB across the reference library's 19,107 images, so a whole-library pass is a transfer under FR-NC-6 rather than something that may start on its own. MIN_CROP_EDGE replaces the MIN_DETECT_EDGE this branch briefly had. Same number, guarding the quantity that turned out to matter. faces.md §7b records both measurements, and marks the second as unexplained rather than dressing it as a finding. Grouped by the buffer detection ran against, faces per image was 0.078 at 1024 or below and 1.82 at 2048 or better, controlled for file type and size. That gap is real and reproducible and I cannot account for it, because the letterbox above says detector input should not matter. M4 is where it gets settled. The crop measurement does not depend on it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d0b671d4db |
Put the grouping dials where the regrouping is
The merge probability was `dr_face`'s constant and the smallest group was a bare `< 2` in the clustering pass. Both were tuned on one library — 1,813 faces of one photographer's family — and the quantity they optimise is a property of the population, not of the model. A household at close family resemblance and two thousand strangers at a wedding want different answers, and neither of them is the reference library. The doc comment already conceded the point and pointed at `face_index --tune`; a photographer does not have a terminal. So they are `FaceSettings` now, saved per device beside the cache budgets and edited from the People screen — beside the Regroup button that applies them and the rail that shows what they did, because a value changed three screens away from its effect is one nobody can tune. Moving them is safe by construction, which is why nothing asks for confirmation: a regroup writes only the suggested half, and confirmations, names and ignores enter as anchors and come back unchanged. The smallest-group rule is applied only to groups the system invented — a group the user named or set aside survives it whatever its size, because a display preference does not overrule a judgement. **Withdrawal, without which the setting does nothing visible.** Raising the smallest group stops the pass creating small groups; it does not remove the ones a previous pass made, because those still hold their suggestions, so they are not empty, so the prune leaves them. The pass now releases every unanchored face it did not place before pruning. And a dial you cannot see the effect of is not a dial. "What would this do?" runs the same population through the clusterer without opening a transaction and reports groups, faces grouped and largest group — one row of `--tune`'s table, on the user's own library, on a worker thread. The line leads with the group count because that is the number that says which side of the right setting you are on: it climbs as fragments are gathered into people and falls as separate people start being welded, while the grouped-face count rises straight through both. The preview parks its poll timer in a slot of its own. A preview and a regroup are allowed to be in flight together, and sharing the sweep's single slot would have the second to start drop the first's timer — visible as a Regroup that finished on its worker and never said so. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1d4c3348af |
Let a 32-pixel face count, and move the blur floor with it
64 source pixels was too strict: it threw away 70% of everything the detector
finds, and plenty of what it took were faces a person could name.
Lowering it is not a one-line change, because the two floors are coupled. A face
under 112 pixels is *upsampled* to reach the embedder and upsampling invents no
edges, so a small face scores low on sharpness however crisp the original was.
Re-measured over the reference library with `face_index --quality`:
min crop min sharp size cut blur cut kept
32 0.000 44% 0% 56%
32 0.002 44% 3% 53%
32 0.005 44% 8% 48%
32 0.010 44% 16% 40%
32 0.020 44% 27% 30%
64 0.020 70% 7% 23%
Holding the blur floor at 0.020 while dropping the size floor to 32 would have
rejected a further 27% — for being small rather than for being blurred — and
kept only 30%, barely more than the 23% the strict pair kept. Most of the point
of lowering the size floor would have gone straight back out through the other
gate.
0.005 removes 8% of what the size floor leaves, which is the same job 0.020 was
doing at 64 (7%): the large-but-soft face this gate exists for. Together they
now keep 48% of what the detector finds, against 23% before.
The box pre-filter follows down to 24, staying below what the real floor accepts
so it cannot reject a face that would have cleared 32.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
a1790c3e67 |
Stop indexing faces too small or too blurred to be anyone
The library was storing faces at 52 source pixels and embedding whatever came
back. There was a size floor, but it was 40 pixels on the *bounding box*, and
there was no blur gate at all — so a subject walking through a half-second
exposure detected confidently, aligned cleanly, and produced a perfectly
ordinary-looking 512-vector. Nothing downstream can tell that apart from a real
face, and because blurs resemble each other more than they resemble the people
they were, they cluster together and weld unrelated identities into one group.
Two floors, both measured rather than guessed. `face_index --quality` runs the
detector over real proxies with both gates disabled and prints the distribution;
over 1,503 faces in 600 images of the reference library:
percentile crop px sharpness
1% 16 0.0006
25% 23 0.0025
50% 38 0.0071
75% 76 0.0284
99% 352 0.4282
The median face in a personal library is 38 pixels. Most of what the detector
finds is background: people across a square, a face on a poster, a stranger at
the next table. They are real detections and useless identifications.
**Size, on the crop rather than the box.** "At least 64x64" has to mean the
pixels the *embedder* sees, and the box is not that — the ArcFace template
reaches past it for forehead and chin, so the aligned crop spans roughly 1.3x
the box's shorter edge. The floor is therefore `min_source_px` on the aligned
crop, applied after the warp fixes the scale, and `min_face_px` drops to 48 as
what it always really was: a cheap pre-filter set low enough that it cannot
reject a face the real floor would have kept.
**Sharpness.** Variance of the Laplacian divided by the variance of the luma it
was taken over. The division is the part that matters: raw Laplacian variance
scales with contrast, so a threshold on it would quietly discard every backlit
portrait in the library. The ratio asks how much of the crop's variation is
edges rather than broad gradients, and is invariant to exposure.
What each pair removes, cumulatively, of everything the detector finds:
min crop min sharp size cut blur cut kept
64 0.000 70% 0% 30%
64 0.010 70% 3% 27%
64 0.020 70% 7% 23%
80 0.010 76% 2% 21%
64 and 0.020. The size floor does most of the work, and the blur floor removing
only 7% on top of it is the point rather than a disappointment: at 64 pixels
most faces are already sharp, and what it takes out is the large-but-soft one —
precisely the face that would otherwise contribute a confident, wrong embedding.
The two gates are not independent and the doc comments say so: a face under 112
pixels was upsampled to reach the embedder, and upsampling invents no edges, so
small faces score low on sharpness even when the original was crisp. That is why
`--quality` prints them together.
**This will re-index.** Around 70% of what the current settings store falls below
the new floors — faces between 20 and 40 pixels that nobody could identify. The
People screen gets shorter and every group in it gets better.
66 dr-face tests pass, including that a blurred crop scores below a sharp one,
that halving the contrast does not move the score, and that an upsampled face
scores below the same face at full size.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
7275c020d7 |
Group people at the threshold the library actually supports
0.90 left a third of the reference library ungrouped: 1,213 of 1,813 faces in a
group, and the rest sitting alone in a screen that had nothing to offer for
them.
"Is 0.90 too tight" is not answerable from the number. It is a probability, and
which cosine it lands on depends on the calibration — so the first half of this
is a way to ask the question properly. `face_index --tune` runs the real
clusterer over the real embeddings at ten thresholds and prints what each one
produces. It writes nothing; comparing thresholds by applying them would have
each one pollute the next.
On the reference library:
P cosine groups grouped largest
0.95 0.449 311 62% 51
0.90 0.403 316 67% 51
0.85 0.374 318 70% 57
0.80 0.353 328 74% 69
0.75 0.335 327 77% 69
0.70 0.319 326 79% 81
0.50 0.267 303 85% 90
The count of *groups* is the signal, not the count of grouped faces. Loosening
from 0.95 makes it climb: real people are being assembled out of fragments. It
peaks at 0.80 and then falls — and a falling group count while the grouped faces
keep rising is the shape of over-merging, separate identities being welded
together. That is the FR-CULL-10 failure, and the one the user cannot undo by
hand.
So 0.80: the loosest setting still building people rather than melting them
together. A third more of the library gets grouped than at 0.90, and the largest
group grows by eighteen faces rather than by forty.
The table is one library, and the doc comment says so — `--tune` reruns it on
any other.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
c8c6368542 |
Index the whole library by fetching what it has not seen
"Index faces in the whole library" could not. Its work list was intersected
with the thumbnail store at `ThumbSize::Large`, and nothing fills that class for
a whole library — `SWEEP_THUMB_SIZE` is deliberately `Grid`, because the large
class is ~860 MB of shards against ~200 MB and every syncing device pays it. So
the only images with a large proxy were the ones the user had personally zoomed
into or opened in the loupe. On this library that was 220 of 23,529.
The comment defending it misread the requirement:
// Requesting one here would put face indexing on the network path,
// which FR-CULL-8 explicitly keeps it off.
FR-CULL-8 keeps indexing off the **full decode**, not the network, and then says
the opposite in the same paragraph: "where no proxy exists, the job requests one
at background priority rather than decoding inline". faces.md §7 repeats it.
Neither was implemented.
So the pass fetches. Same two-stage route the thumbnail sweep uses — the header,
then the located preview's own byte range (FR-NC-3) — so no whole file is pulled
and no RAW is decoded, because an embedded preview is a JPEG. The work list is
now every visible image with no `face_index` row for the model: 23,308 here,
against nearly none before.
**It indexes at the resolution the preview actually has**, not the 1024 the old
tier would have given. `locate_preview` already picks the largest embedded
preview, and the thumbnail sweep was decoding it and throwing the detail away at
`downscale_to(256)`. A face 2% across the frame is 5 px on a grid thumbnail and
~61 px at the cap here — and 112 is what the embedder samples, so this is the
difference between an upsampled crop and a real one. `crop_px` records which,
per face, as §7 intended.
Capped at 3072 rather than truly full: `index_proxy` needs packed `f32` RGB at
12 bytes a pixel, so a 24 MP frame is ~288 MB and the fetch lanes hold one each.
The constant is named and sits next to the reason.
Orientation is applied **before** detection, not after downscaling. That costs a
permutation of a larger buffer — ~15 ms against a ~150 ms decode — and buys the
entire class of bug this codebase keeps having: detection then runs on the
photograph rather than the sensor, so every box and landmark is already in the
space the catalog stores and the overlay draws, with no second mapping to get
backwards.
One detector and one embedder serve every lane. The lanes are concurrent futures
on a single thread, not threads, and inference contains no await, so a `RefCell`
borrow never overlaps another — a pair per lane would duplicate ~16 MB of
weights for no parallelism.
Images with no face in them are recorded too. `face_index` records that
detection *ran*, and zero is its most valuable value: without the row every
landscape and document scan returns on every pass, for ever, and in a personal
library that is most of it (§7a).
The old store-only pass survives as `spawn_store_face_sweep` for
`examples/face_index.rs`, which indexes a local store with no network. The
settings copy no longer claims indexing reads "the photographs already
thumbnailed above", and the audit line says "to fetch" rather than "awaiting a
proxy", which had become a blocker that no longer blocks.
Verified against the real catalog: the new work list returns 23,308 where the
old one returned effectively nothing. 469 tests pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
96d07da15f |
Sync face data as sealed shards, so a second device does not re-index
Indexing 23,500 images is about two hours of CPU, and the result is byte-identical on every device: the same model over the same proxy produces the same embedding. Paying for it once per account rather than once per device is the point. Shards rather than the catalog snapshot, because the snapshot goes up whole on every sync and a fully indexed library carries roughly 30 MB of embeddings. That is exactly the cost the thumbnail store's 25 MB cap exists to bound, so face shards use the same cap -- imported from dr_thumbs rather than restated, since the number is a statement about sync cost and the two must not drift apart. The split follows the one already there: bulk immutable data in sealed shards, small mutable data in the catalog snapshot. Faces, landmarks, embeddings and run markers shard; people, names and assignments ride the catalog and merge by uuid. Keyed on oc:fileid throughout, never on image_id, because a row id means nothing on another device. The run marker travels with the faces it describes. Without it a receiving device cannot tell an image with no faces from one never examined, and would re-detect every landscape it had just adopted. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
26a1eb7e28 |
Record that face detection has run, not just what it found
An image with no faces in it was indistinguishable from one that had never been looked at, so every landscape, still life and document scan in the library was re-detected on every pass, for ever. In a real library that is most of it: on the 23,527-image test library, 64 of the first 110 images indexed contain no face at all. Schema v9 adds face_index, a run marker per (image, model) carrying the face count and the proxy edge it read. Keyed on the model, so a model change puts every image back in the queue by itself. That makes a coverage figure possible, which is the thing a user actually wants to see. The audit also splits the outstanding set by whether a proxy exists, because 23,417 awaiting a proxy and 110 ready to index are different problems, and telling the user to run indexing again would not fix the first. The Identity screen gains Index faces, Stop, and the coverage line. examples/face_index.rs is the same check and sweep without a window, which is the right shape for an overnight pass. Measured on the real library in release: 3.5 images/second, 110 images and 125 faces in 30 seconds, and a second run correctly finds nothing left to do. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |