Give the similarity scan the machine's SIMD, and its cache
The scan is O(n²) dot products and nothing else, so its speed is the face subsystem's speed — and it was running at 0.7 flops per cycle. Two separate faults, both measured over the reference 18,143-face library on twenty cores. It walked the whole embedding array once per row, ~336 GB of traffic, where a column tile that fits in L2 is read once per tile of rows: 4.64s → 2.81s. And the workspace builds for baseline x86-64 — SSE2, no FMA — into which the portable loop was not being vectorised at all: 2.81s → 0.86s, 195 GFLOP/s. So the dot product is now chosen per machine. AVX2 + FMA where is_x86_feature_detected! finds it; NEON unconditionally on aarch64, since Advanced SIMD is in that baseline and every Android device the app builds for has it — with the explicit vfmaq, because LLVM will not fuse a multiply and an add without being told to. The portable loop stays as the definition the others are tested against, and the_fastest_kernel_agrees_with_the_portable_one is the only check the NEON path gets on a machine that is not aarch64. Faces::embeddings is one flat buffer rather than a Vec per face: the pointer chase defeated both the prefetcher and the tiling, and it is also the layout a GPU pass would want. Behaviour is unchanged and that is checked rather than asserted — the same 1,531,969 pairs from all three kernels, and on the real library the same 2,518 groups holding the same 16,246 faces with the same confidence distribution. A full regroup there goes from 10.0s to 5.9s; the rest is the agglomeration, which is a sequential heap walk and is where the next look should go. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
+1
-1
@@ -45,7 +45,7 @@ done | sort -rn
|
||||
| Test functions | 2,042, plus 21 integration test files |
|
||||
| `.unwrap()` in production code | **3** — one in `dr-gpu`, two in `dr-ingest` |
|
||||
| `.unwrap()` in test code | ~1,500, which is where it belongs |
|
||||
| `unsafe` blocks | 6 |
|
||||
| `unsafe` blocks | 9 — three of them the face scan's SIMD kernels (faces.md §9) |
|
||||
| `TRACES` tags / orphan tags | 793 / 0 |
|
||||
| Resolved dependencies | 826 |
|
||||
| Largest function | `dr-ui::run` — 1,855 lines |
|
||||
|
||||
@@ -653,6 +653,33 @@ its histogram contribution before the next is started, and never materialised wh
|
||||
background, after an indexing sweep that took an hour. An approximate index is an optimisation to
|
||||
reach for when §12 says it is needed, not before.
|
||||
|
||||
**The kernel is most of the cost, and it was running at a tenth of the machine.** The scan is
|
||||
`O(n²)` dot products and nothing else, so its speed *is* the subsystem's speed. Two things were
|
||||
wrong with the first one, both measured over a real 18,143-face library on a twenty-core desktop:
|
||||
|
||||
| | scan | GFLOP/s |
|
||||
|---|---|---|
|
||||
| a row against every other row, `&[Vec<f32>]` | 4.64 s | 36 |
|
||||
| tiled on the column side too, one flat buffer | 2.81 s | 60 |
|
||||
| **plus AVX2 + FMA** | **0.86 s** | **195** |
|
||||
|
||||
The first is memory: walking the whole embedding array once per row moves ~336 GB for that library,
|
||||
where a column tile that fits in L2 is read once per *tile of rows*. The second is that the workspace
|
||||
builds for baseline `x86-64` — SSE2, no FMA — and the portable loop was not being vectorised into
|
||||
even that, at 0.7 flops per cycle.
|
||||
|
||||
So the dot product is chosen per machine: AVX2 + FMA where `is_x86_feature_detected!` finds it,
|
||||
**NEON unconditionally on aarch64** — Advanced SIMD is in that baseline, so every Android device the
|
||||
app builds for has it, and the explicit `vfmaq` matters because LLVM will not fuse a multiply and an
|
||||
add on its own. The portable loop remains the definition the others are tested against. All three
|
||||
produce the same 1,531,969 pairs.
|
||||
|
||||
Worth keeping in view when this is next optimised: on that library a full regroup is **scan 4.60 s ·
|
||||
agglomerate 4.84 s · score 0.23 s**, so the scan was under half of it and the SIMD work moved the
|
||||
whole pass from 10.0 s to 5.9 s. A GPU GEMM is the next step for the scan, and it is capped by the same
|
||||
arithmetic — the agglomeration is a sequential heap walk and no amount of silicon touches
|
||||
it.
|
||||
|
||||
**Constraints, not just thresholds:**
|
||||
|
||||
- **Cannot-link on co-occurrence.** Two faces in the same image are never merged. This is the same
|
||||
|
||||
Reference in New Issue
Block a user