Give the similarity scan the machine's SIMD, and its cache

The scan is O(n²) dot products and nothing else, so its speed is the
face subsystem's speed — and it was running at 0.7 flops per cycle.

Two separate faults, both measured over the reference 18,143-face
library on twenty cores. It walked the whole embedding array once per
row, ~336 GB of traffic, where a column tile that fits in L2 is read
once per tile of rows: 4.64s → 2.81s. And the workspace builds for
baseline x86-64 — SSE2, no FMA — into which the portable loop was not
being vectorised at all: 2.81s → 0.86s, 195 GFLOP/s.

So the dot product is now chosen per machine. AVX2 + FMA where
is_x86_feature_detected! finds it; NEON unconditionally on aarch64,
since Advanced SIMD is in that baseline and every Android device the app
builds for has it — with the explicit vfmaq, because LLVM will not fuse
a multiply and an add without being told to. The portable loop stays as
the definition the others are tested against, and
the_fastest_kernel_agrees_with_the_portable_one is the only check the
NEON path gets on a machine that is not aarch64.

Faces::embeddings is one flat buffer rather than a Vec per face: the
pointer chase defeated both the prefetcher and the tiling, and it is
also the layout a GPU pass would want.

Behaviour is unchanged and that is checked rather than asserted — the
same 1,531,969 pairs from all three kernels, and on the real library the
same 2,518 groups holding the same 16,246 faces with the same confidence
distribution. A full regroup there goes from 10.0s to 5.9s; the rest is
the agglomeration, which is a sequential heap walk and is where the next
look should go.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-29 11:33:15 +02:00
co-authored by Claude Opus 5
parent ebb7d3cf5c
commit e596eb0657
4 changed files with 321 additions and 43 deletions
+1 -1
View File
@@ -45,7 +45,7 @@ done | sort -rn
| Test functions | 2,042, plus 21 integration test files |
| `.unwrap()` in production code | **3** — one in `dr-gpu`, two in `dr-ingest` |
| `.unwrap()` in test code | ~1,500, which is where it belongs |
| `unsafe` blocks | 6 |
| `unsafe` blocks | 9 — three of them the face scan's SIMD kernels (faces.md §9) |
| `TRACES` tags / orphan tags | 793 / 0 |
| Resolved dependencies | 826 |
| Largest function | `dr-ui::run` — 1,855 lines |