Commit Graph
5 Commits
Author SHA1 Message Date
dtourolle 6b1aac477d Put the developer docs under docs/dev and index the folder for users first
docs/ had 26 developer documents flat beside the manual, and the two
audiences are very differently sized: most readers want the manual and
the gesture reference, a few want the register, the designs and the
measurements. The manual and gestures.md stay at the top; everything for
someone changing the code moves to docs/dev/, and the two documents that
name their own successors — the v0.1 milestone and the UI-refinement plan
— go to docs/dev/archive/ rather than being deleted, since both are still
cited. docs/README.md is the index, users first.

Every reference follows: code comments, Cargo manifests, the workflows,
the pre-commit hook, the bench and traceability tools (which locate the
repo root by docs/dev/requirements.md now), packaging, the Docker READMEs,
CLAUDE.md, CONTRIBUTING.md and the README. The matrix links one level
deeper and is regenerated. Links out of the moved documents into the tree
gain a level; a link checker over every Markdown file finds none broken.
2026-09-20 16:20:15 +02:00
dtourolle 8b3abdb787 Keep each face's quality, and never compare against a poor one
The embedder's raw output has a length, and the length is a reading of
how recognisable the crop was: a blur, an occlusion or a hard profile
comes out short. Normalising threw it away. A short vector sits near
the middle of the sphere and matches a little of everyone, which is how
one bad crop bridges two people in a grouping pass.

So the length is kept — the store now holds the raw vector, re-normalised
on load, with the length beside it as `faces.quality` — and a face under
MIN_GALLERY_QUALITY (14) is a probe: measured against the gallery and
placed where it fits, but never what another face is measured against.
Two probes are never paired, and a probe is nobody's evidence for a
confidence. The People screen shows the number as "Quality 17.3", dimmed
below the floor.

Faces indexed before this stored unit vectors and have no reading; they
are admitted to the gallery, and schema V14 forgets the run marker of
every image holding one so the next indexing pass measures them. A
peer's unmeasured shard faces are not adopted, or a sync would write
that marker back.
2026-09-11 21:50:12 +02:00
dtourolleandClaude Opus 5 a67402961d Let the merge engine's dot product use the machine's kernel too
Engine::cross is the one place a dot product is computed during
agglomeration — when two groups become adjacent through a third and
their sub-threshold pairs, never summed because they were never
interesting, have to be accounted for. It was calling the portable loop
while the scan beside it had AVX2 or NEON, which on the reference
library was 1,753,514 dot products taking 0.54s.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 11:58:07 +02:00
dtourolleandClaude Opus 5 e596eb0657 Give the similarity scan the machine's SIMD, and its cache
The scan is O(n²) dot products and nothing else, so its speed is the
face subsystem's speed — and it was running at 0.7 flops per cycle.

Two separate faults, both measured over the reference 18,143-face
library on twenty cores. It walked the whole embedding array once per
row, ~336 GB of traffic, where a column tile that fits in L2 is read
once per tile of rows: 4.64s → 2.81s. And the workspace builds for
baseline x86-64 — SSE2, no FMA — into which the portable loop was not
being vectorised at all: 2.81s → 0.86s, 195 GFLOP/s.

So the dot product is now chosen per machine. AVX2 + FMA where
is_x86_feature_detected! finds it; NEON unconditionally on aarch64,
since Advanced SIMD is in that baseline and every Android device the app
builds for has it — with the explicit vfmaq, because LLVM will not fuse
a multiply and an add without being told to. The portable loop stays as
the definition the others are tested against, and
the_fastest_kernel_agrees_with_the_portable_one is the only check the
NEON path gets on a machine that is not aarch64.

Faces::embeddings is one flat buffer rather than a Vec per face: the
pointer chase defeated both the prefetcher and the tiling, and it is
also the layout a GPU pass would want.

Behaviour is unchanged and that is checked rather than asserted — the
same 1,531,969 pairs from all three kernels, and on the real library the
same 2,518 groups holding the same 16,246 faces with the same confidence
distribution. A full regroup there goes from 10.0s to 5.9s; the rest is
the agglomeration, which is a sequential heap walk and is where the next
look should go.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 11:33:15 +02:00
dtourolleandClaude Opus 5 c10dca984f Regroup the library without stopping the window
Pressing Regroup on a real library did not come back. Clustering 1,813 faces is
the textbook agglomeration — compute every pairwise cosine, then repeatedly scan
all live group pairs, score each with average link, and merge the best — and the
scan is inside the loop. Each merge rescans every surviving pair, and each score
is recomputed from scratch over every cross pair. Some 1.6 million pair scores
per merge, some 700 merges to do.

Three changes, none of which alter the answer.

Only above-threshold pairs can ever matter. An average that reaches the
threshold must have at least one term at or above it, so two groups with no
qualifying pair between them can never merge — not now, and not after any
sequence of merges, since merging only adds terms. The new `neighbours` module
produces exactly that sparse list: 7,875 pairs rather than 1.6 million on the
reference library. It also means the n^2 matrix is never materialised, so memory
goes from O(n^2) to O(edges) — 2.5 GB to a few hundred KB at 25,000 faces.

Merges cannot cross components, so the connected components of that graph are
independent problems: four hundred small agglomerations instead of one large one.

Average link is additive — sum(A u B, C) = sum(A, C) + sum(B, C) — so a merged
group's scores follow by addition. Kept as running (sum, count) per adjacent
pair, a score costs one division instead of a nested loop, and a heap with lazy
invalidation replaces the rescan.

Measured on the reference library: 0.28s, release, for all 1,813 faces.

An exact ANN index was tried and removed, and neighbours.rs records why so it is
not rediscovered as a good idea. IVF with a triangle-inequality bound is exact
and prunes beautifully on synthetic clusters; on real embeddings it prunes
*nothing* — 946 of 946 cell pairs survive. Median pair angle is 88.5 degrees and
the merge threshold is 66.2, so the bound needs cells of radius under ~10
degrees, but two photographs of the same person sit 36-60 degrees apart. No
ball-based partition of a 512-d near-orthogonal space can be tight enough. So
the scan stayed exhaustive and got an unrolled dot product and its blocks spread
across cores instead.

Correctness is held by keeping the old implementation as an oracle: three tests
run both engines over the same population — plain, under co-occurrence and
anchor constraints, and with a size-weighted calibration — and assert the
clusters are identical. Determinism is asserted at a size where the threaded
path is in play.

62 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 21:16:55 +02:00