Measure a regroup on the tablet, not just on the desktop

The GPU question needed a number nobody had: how a regroup divides on
the hardware whose CPU is weakest. dr-face carries no weights and
touches no display, and dr-catalog's example needs only a catalog file,
so both run under adb shell against a copy of a real library.

On the same 18,143 faces — desktop against the tablet — scan 0.96s /
2.61s, agglomerate 1.69s / 2.16s, score 0.26s / 0.40s. The scan is half
the pass on the tablet and under a third on the desktop, because twenty
cores of AVX2 pull ahead of NEON much further than the merge engine's
single-threaded hashing does. So a GPU GEMM is worth roughly 2× a
regroup on the tablet and 1.5× here, and it is the tablet that should
decide whether it is built.

The two architectures agree exactly: the same 1,531,969 evidence pairs,
the same 2,518 groups holding the same 16,246 faces, the same
reliability table. That is a better check on the NEON kernel than the
unit test can be.

Two instruments, both read-only: the example now prints its phases, and
dr-face gains scan_bench, which needs no library at all and so can
answer "how fast is this machine" on a device with nothing on it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-29 12:16:48 +02:00
co-authored by Claude Opus 5
parent f4395bd17c
commit b2250cc460
4 changed files with 162 additions and 12 deletions
@@ -305,6 +305,46 @@ fn full_library(
candidates.sort_by_key(|c| c.face);
println!("\nthe whole library, at the default merge probability:");
// The three phases, separately, because "a regroup takes n seconds" does
// not tell anyone which half to optimise — and the answer differs between
// a desktop and a tablet (docs/faces.md §9).
{
let dim = candidates.first().map(|c| c.embedding.len()).unwrap_or(0);
let flat: Vec<f32> = candidates.iter().flat_map(|c| c.embedding.clone()).collect();
let crop_px: Vec<f32> = candidates.iter().map(|c| c.crop_px).collect();
let images: Vec<u64> = candidates.iter().map(|c| c.image).collect();
let view = dr_face::neighbours::Faces {
embeddings: &flat,
dim,
crop_px: &crop_px,
images: &images,
};
let t = std::time::Instant::now();
let evidence = dr_face::neighbours::above_threshold(&view, cal, dr_face::RIVAL_FLOOR);
let scan = t.elapsed().as_secs_f64();
// `cluster` runs its own scan at the merge threshold, so the
// agglomeration is what is left after taking one scan off the total.
let t = std::time::Instant::now();
let clusters = dr_face::cluster(&candidates, cal, dr_face::DEFAULT_MERGE_PROBABILITY);
let agglomerate = t.elapsed().as_secs_f64() - scan;
let t = std::time::Instant::now();
let _ = dr_face::identity_shares(
candidates.len(),
&clusters,
&evidence,
dr_face::TOP_MATCHES,
);
println!(
" scan {scan:.2}s ({} evidence pairs) · agglomerate {agglomerate:.2}s · score {:.2}s",
evidence.len(),
t.elapsed().as_secs_f64()
);
}
let start = std::time::Instant::now();
let grouping = dr_face::cluster_scored(&candidates, cal, dr_face::DEFAULT_MERGE_PROBABILITY);
let real: Vec<_> = grouping