Give the similarity scan the machine's SIMD, and its cache
The scan is O(n²) dot products and nothing else, so its speed is the face subsystem's speed — and it was running at 0.7 flops per cycle. Two separate faults, both measured over the reference 18,143-face library on twenty cores. It walked the whole embedding array once per row, ~336 GB of traffic, where a column tile that fits in L2 is read once per tile of rows: 4.64s → 2.81s. And the workspace builds for baseline x86-64 — SSE2, no FMA — into which the portable loop was not being vectorised at all: 2.81s → 0.86s, 195 GFLOP/s. So the dot product is now chosen per machine. AVX2 + FMA where is_x86_feature_detected! finds it; NEON unconditionally on aarch64, since Advanced SIMD is in that baseline and every Android device the app builds for has it — with the explicit vfmaq, because LLVM will not fuse a multiply and an add without being told to. The portable loop stays as the definition the others are tested against, and the_fastest_kernel_agrees_with_the_portable_one is the only check the NEON path gets on a machine that is not aarch64. Faces::embeddings is one flat buffer rather than a Vec per face: the pointer chase defeated both the prefetcher and the tiling, and it is also the layout a GPU pass would want. Behaviour is unchanged and that is checked rather than asserted — the same 1,531,969 pairs from all three kernels, and on the real library the same 2,518 groups holding the same 16,246 faces with the same confidence distribution. A full regroup there goes from 10.0s to 5.9s; the rest is the agglomeration, which is a sequential heap walk and is where the next look should go. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -220,15 +220,29 @@ pub fn cluster_scored(faces: &[Candidate], cal: &Calibration, min_probability: f
|
||||
/// columns. Both entry points do, identically, which is the only reason this is
|
||||
/// a type and not three locals.
|
||||
struct Columns {
|
||||
embeddings: Vec<Vec<f32>>,
|
||||
/// Every embedding end to end — see [`Faces::embeddings`] for why flat.
|
||||
embeddings: Vec<f32>,
|
||||
dim: usize,
|
||||
crop_px: Vec<f32>,
|
||||
images: Vec<u64>,
|
||||
}
|
||||
|
||||
impl Columns {
|
||||
fn of(faces: &[Candidate]) -> Self {
|
||||
// Ragged input would index the wrong row for every face after the odd
|
||||
// one, so the widest wins and short rows are padded with zeros: a
|
||||
// zero-padded row scores lower against everything, which is the safe
|
||||
// direction. It does not happen — one model, one dimension — and it is
|
||||
// handled rather than trusted because the failure would be silent.
|
||||
let dim = faces.iter().map(|f| f.embedding.len()).max().unwrap_or(0);
|
||||
let mut embeddings = Vec::with_capacity(faces.len() * dim);
|
||||
for f in faces {
|
||||
embeddings.extend_from_slice(&f.embedding);
|
||||
embeddings.resize(embeddings.len() + dim - f.embedding.len(), 0.0);
|
||||
}
|
||||
Self {
|
||||
embeddings: faces.iter().map(|f| f.embedding.clone()).collect(),
|
||||
embeddings,
|
||||
dim,
|
||||
crop_px: faces.iter().map(|f| f.crop_px).collect(),
|
||||
images: faces.iter().map(|f| f.image).collect(),
|
||||
}
|
||||
@@ -237,6 +251,7 @@ impl Columns {
|
||||
fn view(&self) -> Faces<'_> {
|
||||
Faces {
|
||||
embeddings: &self.embeddings,
|
||||
dim: self.dim,
|
||||
crop_px: &self.crop_px,
|
||||
images: &self.images,
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user