Give the similarity scan the machine's SIMD, and its cache

The scan is O(n²) dot products and nothing else, so its speed is the
face subsystem's speed — and it was running at 0.7 flops per cycle.

Two separate faults, both measured over the reference 18,143-face
library on twenty cores. It walked the whole embedding array once per
row, ~336 GB of traffic, where a column tile that fits in L2 is read
once per tile of rows: 4.64s → 2.81s. And the workspace builds for
baseline x86-64 — SSE2, no FMA — into which the portable loop was not
being vectorised at all: 2.81s → 0.86s, 195 GFLOP/s.

So the dot product is now chosen per machine. AVX2 + FMA where
is_x86_feature_detected! finds it; NEON unconditionally on aarch64,
since Advanced SIMD is in that baseline and every Android device the app
builds for has it — with the explicit vfmaq, because LLVM will not fuse
a multiply and an add without being told to. The portable loop stays as
the definition the others are tested against, and
the_fastest_kernel_agrees_with_the_portable_one is the only check the
NEON path gets on a machine that is not aarch64.

Faces::embeddings is one flat buffer rather than a Vec per face: the
pointer chase defeated both the prefetcher and the tiling, and it is
also the layout a GPU pass would want.

Behaviour is unchanged and that is checked rather than asserted — the
same 1,531,969 pairs from all three kernels, and on the real library the
same 2,518 groups holding the same 16,246 faces with the same confidence
distribution. A full regroup there goes from 10.0s to 5.9s; the rest is
the agglomeration, which is a sequential heap walk and is where the next
look should go.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-29 11:33:15 +02:00
co-authored by Claude Opus 5
parent ebb7d3cf5c
commit e596eb0657
4 changed files with 321 additions and 43 deletions
+27
View File
@@ -653,6 +653,33 @@ its histogram contribution before the next is started, and never materialised wh
background, after an indexing sweep that took an hour. An approximate index is an optimisation to
reach for when §12 says it is needed, not before.
**The kernel is most of the cost, and it was running at a tenth of the machine.** The scan is
`O(n²)` dot products and nothing else, so its speed *is* the subsystem's speed. Two things were
wrong with the first one, both measured over a real 18,143-face library on a twenty-core desktop:
| | scan | GFLOP/s |
|---|---|---|
| a row against every other row, `&[Vec<f32>]` | 4.64 s | 36 |
| tiled on the column side too, one flat buffer | 2.81 s | 60 |
| **plus AVX2 + FMA** | **0.86 s** | **195** |
The first is memory: walking the whole embedding array once per row moves ~336 GB for that library,
where a column tile that fits in L2 is read once per *tile of rows*. The second is that the workspace
builds for baseline `x86-64` — SSE2, no FMA — and the portable loop was not being vectorised into
even that, at 0.7 flops per cycle.
So the dot product is chosen per machine: AVX2 + FMA where `is_x86_feature_detected!` finds it,
**NEON unconditionally on aarch64** — Advanced SIMD is in that baseline, so every Android device the
app builds for has it, and the explicit `vfmaq` matters because LLVM will not fuse a multiply and an
add on its own. The portable loop remains the definition the others are tested against. All three
produce the same 1,531,969 pairs.
Worth keeping in view when this is next optimised: on that library a full regroup is **scan 4.60 s ·
agglomerate 4.84 s · score 0.23 s**, so the scan was under half of it and the SIMD work moved the
whole pass from 10.0 s to 5.9 s. A GPU GEMM is the next step for the scan, and it is capped by the same
arithmetic — the agglomeration is a sequential heap walk and no amount of silicon touches
it.
**Constraints, not just thresholds:**
- **Cannot-link on co-occurrence.** Two faces in the same image are never merged. This is the same