Measure a regroup on the tablet, not just on the desktop

The GPU question needed a number nobody had: how a regroup divides on
the hardware whose CPU is weakest. dr-face carries no weights and
touches no display, and dr-catalog's example needs only a catalog file,
so both run under adb shell against a copy of a real library.

On the same 18,143 faces — desktop against the tablet — scan 0.96s /
2.61s, agglomerate 1.69s / 2.16s, score 0.26s / 0.40s. The scan is half
the pass on the tablet and under a third on the desktop, because twenty
cores of AVX2 pull ahead of NEON much further than the merge engine's
single-threaded hashing does. So a GPU GEMM is worth roughly 2× a
regroup on the tablet and 1.5× here, and it is the tablet that should
decide whether it is built.

The two architectures agree exactly: the same 1,531,969 evidence pairs,
the same 2,518 groups holding the same 16,246 faces, the same
reliability table. That is a better check on the NEON kernel than the
unit test can be.

Two instruments, both read-only: the example now prints its phases, and
dr-face gains scan_bench, which needs no library at all and so can
answer "how fast is this machine" on a device with nothing on it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-29 12:16:48 +02:00
co-authored by Claude Opus 5
parent f4395bd17c
commit b2250cc460
4 changed files with 162 additions and 12 deletions
+17 -11
View File
@@ -680,14 +680,15 @@ and touches no display, so its tests are a plain ARM64 binary that runs under `a
nothing installed. Worth running whenever the kernels change.
**Where a regroup's time actually goes**, on that library, because the answer moved twice while it
was being looked at:
was being looked at. Measured with `cargo run --release -p dr-catalog --example face_confidence --
CATALOG --full`, on the reference desktop and on a Honor tablet (ROD2-W09, aarch64):
| | before | after |
|---|---|---|
| scan | 4.60 s | **0.94 s** |
| agglomerate | 4.84 s | **1.75 s** |
| score | 0.23 s | 0.28 s |
| **total** | **10.0 s** | **3.0 s** |
| | desktop, before | desktop | **tablet** |
|---|---|---|---|
| scan | 4.60 s | 0.96 s | **2.61 s** |
| agglomerate | 4.84 s | 1.69 s | **2.16 s** |
| score | 0.23 s | 0.26 s | **0.40 s** |
| **total** | **10.0 s** | **3.1 s** | **5.2 s** |
The scan came down by the kernel work above. The agglomeration was not, as it looked, an irreducible
sequential heap walk: 3.06 s of it was every component scanning the *whole* pair list for the pairs
@@ -695,10 +696,15 @@ that were its own — 382 million set lookups to place 804,499 pairs — which t
one pass while it is finding the components anyway. Another 0.54 s was `Engine::cross` computing dot
products with the portable loop while the scan beside it used the machine's SIMD.
A GPU GEMM is the obvious next step for the scan, and it should be judged against the second column
rather than the first: the scan is now under a third of the pass on a desktop, so the ceiling there
is a 1.5× regroup. On a tablet, where the CPU is several times slower and the GPU is not, the split
is different and the case is stronger — which is a measurement nobody has taken yet.
**The two architectures agree exactly**, which is worth more than either column: the same 1,531,969
evidence pairs, the same 2,518 groups holding the same 16,246 faces, and the same reliability table,
from AVX2 on the desktop and NEON on the tablet. That is the cross-kernel check the unit test can
only approximate.
**The tablet is where a GPU GEMM would pay.** Its scan is *half* the pass, against under a third on
the desktop — twenty cores of AVX2 pull ahead of a tablet's NEON far more than the merge engine's
single-threaded hashing does. So a perfect GEMM is worth about 2× a regroup there and about 1.5×
here, and it is the phone and tablet story that should decide whether it gets built.
**Constraints, not just thresholds:**