Measure a regroup on the tablet, not just on the desktop
The GPU question needed a number nobody had: how a regroup divides on the hardware whose CPU is weakest. dr-face carries no weights and touches no display, and dr-catalog's example needs only a catalog file, so both run under adb shell against a copy of a real library. On the same 18,143 faces — desktop against the tablet — scan 0.96s / 2.61s, agglomerate 1.69s / 2.16s, score 0.26s / 0.40s. The scan is half the pass on the tablet and under a third on the desktop, because twenty cores of AVX2 pull ahead of NEON much further than the merge engine's single-threaded hashing does. So a GPU GEMM is worth roughly 2× a regroup on the tablet and 1.5× here, and it is the tablet that should decide whether it is built. The two architectures agree exactly: the same 1,531,969 evidence pairs, the same 2,518 groups holding the same 16,246 faces, the same reliability table. That is a better check on the NEON kernel than the unit test can be. Two instruments, both read-only: the example now prints its phases, and dr-face gains scan_bench, which needs no library at all and so can answer "how fast is this machine" on a device with nothing on it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
+17
-11
@@ -680,14 +680,15 @@ and touches no display, so its tests are a plain ARM64 binary that runs under `a
|
||||
nothing installed. Worth running whenever the kernels change.
|
||||
|
||||
**Where a regroup's time actually goes**, on that library, because the answer moved twice while it
|
||||
was being looked at:
|
||||
was being looked at. Measured with `cargo run --release -p dr-catalog --example face_confidence --
|
||||
CATALOG --full`, on the reference desktop and on a Honor tablet (ROD2-W09, aarch64):
|
||||
|
||||
| | before | after |
|
||||
|---|---|---|
|
||||
| scan | 4.60 s | **0.94 s** |
|
||||
| agglomerate | 4.84 s | **1.75 s** |
|
||||
| score | 0.23 s | 0.28 s |
|
||||
| **total** | **10.0 s** | **3.0 s** |
|
||||
| | desktop, before | desktop | **tablet** |
|
||||
|---|---|---|---|
|
||||
| scan | 4.60 s | 0.96 s | **2.61 s** |
|
||||
| agglomerate | 4.84 s | 1.69 s | **2.16 s** |
|
||||
| score | 0.23 s | 0.26 s | **0.40 s** |
|
||||
| **total** | **10.0 s** | **3.1 s** | **5.2 s** |
|
||||
|
||||
The scan came down by the kernel work above. The agglomeration was not, as it looked, an irreducible
|
||||
sequential heap walk: 3.06 s of it was every component scanning the *whole* pair list for the pairs
|
||||
@@ -695,10 +696,15 @@ that were its own — 382 million set lookups to place 804,499 pairs — which t
|
||||
one pass while it is finding the components anyway. Another 0.54 s was `Engine::cross` computing dot
|
||||
products with the portable loop while the scan beside it used the machine's SIMD.
|
||||
|
||||
A GPU GEMM is the obvious next step for the scan, and it should be judged against the second column
|
||||
rather than the first: the scan is now under a third of the pass on a desktop, so the ceiling there
|
||||
is a 1.5× regroup. On a tablet, where the CPU is several times slower and the GPU is not, the split
|
||||
is different and the case is stronger — which is a measurement nobody has taken yet.
|
||||
**The two architectures agree exactly**, which is worth more than either column: the same 1,531,969
|
||||
evidence pairs, the same 2,518 groups holding the same 16,246 faces, and the same reliability table,
|
||||
from AVX2 on the desktop and NEON on the tablet. That is the cross-kernel check the unit test can
|
||||
only approximate.
|
||||
|
||||
**The tablet is where a GPU GEMM would pay.** Its scan is *half* the pass, against under a third on
|
||||
the desktop — twenty cores of AVX2 pull ahead of a tablet's NEON far more than the merge engine's
|
||||
single-threaded hashing does. So a perfect GEMM is worth about 2× a regroup there and about 1.5×
|
||||
here, and it is the phone and tablet story that should decide whether it gets built.
|
||||
|
||||
**Constraints, not just thresholds:**
|
||||
|
||||
|
||||
Reference in New Issue
Block a user