Record where a regroup's time actually goes
The §9 note said the scan was half a regroup and implied the agglomeration was an irreducible sequential walk. Both halves of that are now wrong, and the numbers are the point of the section. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
+20
-5
@@ -674,11 +674,26 @@ app builds for has it, and the explicit `vfmaq` matters because LLVM will not fu
|
||||
add on its own. The portable loop remains the definition the others are tested against. All three
|
||||
produce the same 1,531,969 pairs.
|
||||
|
||||
Worth keeping in view when this is next optimised: on that library a full regroup is **scan 4.60 s ·
|
||||
agglomerate 4.84 s · score 0.23 s**, so the scan was under half of it and the SIMD work moved the
|
||||
whole pass from 10.0 s to 5.9 s. A GPU GEMM is the next step for the scan, and it is capped by the same
|
||||
arithmetic — the agglomeration is a sequential heap walk and no amount of silicon touches
|
||||
it.
|
||||
**Where a regroup's time actually goes**, on that library, because the answer moved twice while it
|
||||
was being looked at:
|
||||
|
||||
| | before | after |
|
||||
|---|---|---|
|
||||
| scan | 4.60 s | **0.94 s** |
|
||||
| agglomerate | 4.84 s | **1.75 s** |
|
||||
| score | 0.23 s | 0.28 s |
|
||||
| **total** | **10.0 s** | **3.0 s** |
|
||||
|
||||
The scan came down by the kernel work above. The agglomeration was not, as it looked, an irreducible
|
||||
sequential heap walk: 3.06 s of it was every component scanning the *whole* pair list for the pairs
|
||||
that were its own — 382 million set lookups to place 804,499 pairs — which the union-find can do in
|
||||
one pass while it is finding the components anyway. Another 0.54 s was `Engine::cross` computing dot
|
||||
products with the portable loop while the scan beside it used the machine's SIMD.
|
||||
|
||||
A GPU GEMM is the obvious next step for the scan, and it should be judged against the second column
|
||||
rather than the first: the scan is now under a third of the pass on a desktop, so the ceiling there
|
||||
is a 1.5× regroup. On a tablet, where the CPU is several times slower and the GPU is not, the split
|
||||
is different and the case is stronger — which is a measurement nobody has taken yet.
|
||||
|
||||
**Constraints, not just thresholds:**
|
||||
|
||||
|
||||
Reference in New Issue
Block a user