Record where a regroup's time actually goes

The §9 note said the scan was half a regroup and implied the
agglomeration was an irreducible sequential walk. Both halves of that
are now wrong, and the numbers are the point of the section.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-29 11:59:01 +02:00
co-authored by Claude Opus 5
parent a67402961d
commit 8f596262b8
+20 -5
View File
@@ -674,11 +674,26 @@ app builds for has it, and the explicit `vfmaq` matters because LLVM will not fu
add on its own. The portable loop remains the definition the others are tested against. All three
produce the same 1,531,969 pairs.
Worth keeping in view when this is next optimised: on that library a full regroup is **scan 4.60 s ·
agglomerate 4.84 s · score 0.23 s**, so the scan was under half of it and the SIMD work moved the
whole pass from 10.0 s to 5.9 s. A GPU GEMM is the next step for the scan, and it is capped by the same
arithmetic — the agglomeration is a sequential heap walk and no amount of silicon touches
it.
**Where a regroup's time actually goes**, on that library, because the answer moved twice while it
was being looked at:
| | before | after |
|---|---|---|
| scan | 4.60 s | **0.94 s** |
| agglomerate | 4.84 s | **1.75 s** |
| score | 0.23 s | 0.28 s |
| **total** | **10.0 s** | **3.0 s** |
The scan came down by the kernel work above. The agglomeration was not, as it looked, an irreducible
sequential heap walk: 3.06 s of it was every component scanning the *whole* pair list for the pairs
that were its own — 382 million set lookups to place 804,499 pairs — which the union-find can do in
one pass while it is finding the components anyway. Another 0.54 s was `Engine::cross` computing dot
products with the portable loop while the scan beside it used the machine's SIMD.
A GPU GEMM is the obvious next step for the scan, and it should be judged against the second column
rather than the first: the scan is now under a third of the pass on a desktop, so the ceiling there
is a 1.5× regroup. On a tablet, where the CPU is several times slower and the GPU is not, the split
is different and the case is stronger — which is a measurement nobody has taken yet.
**Constraints, not just thresholds:**