From 8f596262b8985643c6253c4821cc4fccb2285c80 Mon Sep 17 00:00:00 2001 From: Duncan Tourolle Date: Sat, 29 Aug 2026 11:59:01 +0200 Subject: [PATCH] Record where a regroup's time actually goes MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The §9 note said the scan was half a regroup and implied the agglomeration was an irreducible sequential walk. Both halves of that are now wrong, and the numbers are the point of the section. Co-Authored-By: Claude Opus 5 (1M context) --- docs/faces.md | 25 ++++++++++++++++++++----- 1 file changed, 20 insertions(+), 5 deletions(-) diff --git a/docs/faces.md b/docs/faces.md index c43507f..56abafb 100644 --- a/docs/faces.md +++ b/docs/faces.md @@ -674,11 +674,26 @@ app builds for has it, and the explicit `vfmaq` matters because LLVM will not fu add on its own. The portable loop remains the definition the others are tested against. All three produce the same 1,531,969 pairs. -Worth keeping in view when this is next optimised: on that library a full regroup is **scan 4.60 s · -agglomerate 4.84 s · score 0.23 s**, so the scan was under half of it and the SIMD work moved the -whole pass from 10.0 s to 5.9 s. A GPU GEMM is the next step for the scan, and it is capped by the same -arithmetic — the agglomeration is a sequential heap walk and no amount of silicon touches -it. +**Where a regroup's time actually goes**, on that library, because the answer moved twice while it +was being looked at: + +| | before | after | +|---|---|---| +| scan | 4.60 s | **0.94 s** | +| agglomerate | 4.84 s | **1.75 s** | +| score | 0.23 s | 0.28 s | +| **total** | **10.0 s** | **3.0 s** | + +The scan came down by the kernel work above. The agglomeration was not, as it looked, an irreducible +sequential heap walk: 3.06 s of it was every component scanning the *whole* pair list for the pairs +that were its own — 382 million set lookups to place 804,499 pairs — which the union-find can do in +one pass while it is finding the components anyway. Another 0.54 s was `Engine::cross` computing dot +products with the portable loop while the scan beside it used the machine's SIMD. + +A GPU GEMM is the obvious next step for the scan, and it should be judged against the second column +rather than the first: the scan is now under a third of the pass on a desktop, so the ceiling there +is a 1.5× regroup. On a tablet, where the CPU is several times slower and the GPU is not, the split +is different and the case is stronger — which is a measurement nobody has taken yet. **Constraints, not just thresholds:**