Files
DarkRoom/core/dr-face/examples/scan_bench.rs
T
dtourolleandClaude Opus 5 b2250cc460 Measure a regroup on the tablet, not just on the desktop
The GPU question needed a number nobody had: how a regroup divides on
the hardware whose CPU is weakest. dr-face carries no weights and
touches no display, and dr-catalog's example needs only a catalog file,
so both run under adb shell against a copy of a real library.

On the same 18,143 faces — desktop against the tablet — scan 0.96s /
2.61s, agglomerate 1.69s / 2.16s, score 0.26s / 0.40s. The scan is half
the pass on the tablet and under a third on the desktop, because twenty
cores of AVX2 pull ahead of NEON much further than the merge engine's
single-threaded hashing does. So a GPU GEMM is worth roughly 2× a
regroup on the tablet and 1.5× here, and it is the tablet that should
decide whether it is built.

The two architectures agree exactly: the same 1,531,969 evidence pairs,
the same 2,518 groups holding the same 16,246 faces, the same
reliability table. That is a better check on the NEON kernel than the
unit test can be.

Two instruments, both read-only: the example now prints its phases, and
dr-face gains scan_bench, which needs no library at all and so can
answer "how fast is this machine" on a device with nothing on it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 12:16:48 +02:00

105 lines
3.7 KiB
Rust
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
//! How fast this machine can scan a library for face pairs.
//!
//! cargo run --release -p dr-face --example scan_bench [-- FACES…]
//!
//! Synthetic embeddings, because the scan's cost is `n²/2` dot products and
//! does not care what the vectors mean — which is what makes this runnable on
//! a phone with no library on it, over `adb shell`, beside
//! `tools/face-tests-on-device.sh`.
//!
//! # What it is for
//!
//! docs/faces.md §9 has the desktop numbers and the question they leave open:
//! a GPU GEMM is worth roughly 1.5× of a regroup on a twenty-core desktop,
//! because the scan is under a third of the pass there. On a tablet the CPU is
//! several times slower and the GPU is not, so the same optimisation is worth
//! something different — and nobody had measured which.
//!
//! No weights, no catalog, no display: it needs nothing but the binary.
use dr_face::neighbours::{above_threshold, Faces};
use dr_face::{Calibration, EMBEDDING_DIM, RIVAL_FLOOR};
/// Sizes to time, unless the command line names others.
const DEFAULT_SIZES: [usize; 4] = [2_000, 4_000, 8_000, 18_000];
fn main() {
let sizes: Vec<usize> = {
let given: Vec<usize> = std::env::args().skip(1).filter_map(|a| a.parse().ok()).collect();
if given.is_empty() {
DEFAULT_SIZES.to_vec()
} else {
given
}
};
// The reference curve, so the cosine floor is the one a real library with
// no fit of its own would scan at.
let cal = Calibration::default();
println!("{:>8} {:>9} {:>10} {:>9}", "faces", "scan", "pairs", "GFLOP/s");
for n in sizes {
let (embeddings, crop_px, images) = population(n);
let faces = Faces {
embeddings: &embeddings,
dim: EMBEDDING_DIM,
crop_px: &crop_px,
images: &images,
};
let start = std::time::Instant::now();
let pairs = above_threshold(&faces, &cal, RIVAL_FLOOR);
let secs = start.elapsed().as_secs_f64();
let flop = n as f64 * n as f64 / 2.0 * EMBEDDING_DIM as f64 * 2.0;
println!(
"{n:>8} {secs:>8.2}s {:>10} {:>9.1}",
pairs.len(),
flop / secs / 1e9
);
}
}
/// `n` L2-normalised embeddings in a handful of loose clusters.
///
/// Clustered rather than uniform so the scan finds a plausible number of pairs
/// to keep — a population where nothing survives the threshold would time the
/// rejection path alone, which is not the path that matters. Hashed from an
/// index rather than drawn from an RNG, so a number is reproducible from the
/// command that produced it.
fn population(n: usize) -> (Vec<f32>, Vec<f32>, Vec<u64>) {
let identities = (n / 12).max(1);
let mut embeddings = Vec::with_capacity(n * EMBEDDING_DIM);
for i in 0..n {
let mut v = unit(i % identities);
let noise = unit(i + 1_000_000);
for (x, e) in v.iter_mut().zip(&noise) {
*x = 0.75 * *x + 0.25 * e;
}
embeddings.extend(normalise(v));
}
// Every face in its own photograph: the co-occurrence rule skips pairs
// rather than scoring them, and skipped pairs are not what is being timed.
((embeddings), vec![150.0; n], (0..n as u64).collect())
}
fn unit(seed: usize) -> Vec<f32> {
let mut s = (seed as u64).wrapping_mul(0x9E37_79B9_7F4A_7C15) | 1;
let mut v = Vec::with_capacity(EMBEDDING_DIM);
for _ in 0..EMBEDDING_DIM {
s ^= s << 13;
s ^= s >> 7;
s ^= s << 17;
v.push(((s >> 11) as f64 / (1u64 << 53) as f64) as f32 - 0.5);
}
normalise(v)
}
fn normalise(mut v: Vec<f32>) -> Vec<f32> {
let len = v.iter().map(|x| x * x).sum::<f32>().sqrt();
for x in &mut v {
*x /= len;
}
v
}