Quote the percentile the measurement can actually support

The before/after published a p99 ratio of 6.0x at 4K. It is not supported and
the correction is worth more than the number was.

The baseline run's `fit` rows spread 2.3x between median and 99th percentile
while every row of the after run spreads about 1.1x. A stage costing
`radius x pixels` has no reason to be bimodal, and `fit` is the memory-bound
configuration — it walks the whole 482 MB source on a stride where `1:1` reads
a contiguous window. Something else had the machine.

The merge brings in the cross-check that settles it: "Try every GPU, not only
the fastest one" measured the same baseline code on the same card and reports
4.67 ms p99 for M3 clarity fit at 2560x1600, against 10.40 ms here. Two
measurements of one thing differing by 2.2x mean the noisier one is wrong.

So both percentiles are now published and the p50 column is the claim: 2.8x at
4K rather than 6.0x. The p99 improvement is real and larger; this run cannot
say by how much, and says so.

What the caveat does not touch: every after figure is inside the 16 ms budget
with a p99 within 26% of its median at every size and both views, and clarity
against all-four is a within-run comparison.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-29 13:28:41 +02:00
co-authored by Claude Opus 5
parent 30162e80df
commit 480c46c41a
2 changed files with 49 additions and 26 deletions
+36 -20
View File
@@ -373,29 +373,45 @@ baseline at `0407fb8` and the result at `bff95e2`.
| Result | `bff95e2` |
| Date | 2026-08-29 |
### M3 — the neighbourhood stage alone, before and after
### M3 — clarity alone, before and after
| size | stage | view | before p99 | after p99 | |
|------------:|---------:|:-----|-----------:|----------:|-----:|
| 1920 × 1200 | clarity | fit | 3.94ms | 1.97ms | 2.0× |
| 1920 × 1200 | all four | fit | 2.57ms | 1.81ms | 1.4× |
| 1920 × 1200 | clarity | 1:1 | 3.07ms | 2.41ms | 1.3× |
| 2560 × 1600 | clarity | fit | 10.40ms | 2.37ms | 4.4× |
| 2560 × 1600 | all four | fit | 5.50ms | 2.64ms | 2.1× |
| 2560 × 1600 | clarity | 1:1 | 5.78ms | 3.85ms | 1.5× |
| 3840 × 2160 | clarity | fit | **25.05ms**|**4.17ms** | 6.0× |
| 3840 × 2160 | all four | fit | 25.48ms | 4.61ms | 5.5× |
| 3840 × 2160 | clarity | 1:1 | 27.60ms | 6.93ms | 4.0× |
| 3840 × 2160 | all four | 1:1 | 15.88ms | 7.74ms | 2.1× |
Both percentiles, because they disagree and the disagreement is the
interesting part.
The passes went from 2 to 4 and got faster, which is the point: three of the
four now run on a grid a sixteenth the area, and the fourth — the combine — is a
single bilinear read where it used to be a 105-tap convolution.
| size | view | before p50 | after p50 | | before p99 | after p99 |
|------------:|:-----|-----------:|----------:|-----:|-----------:|----------:|
| 1920 × 1200 | fit | 2.30ms | 1.56ms | 1.5× | 3.94ms | 1.97ms |
| 1920 × 1200 | 1:1 | 2.88ms | 1.99ms | 1.4× | 3.07ms | 2.41ms |
| 2560 × 1600 | fit | 4.41ms | 1.95ms | 2.3× | 10.40ms | 2.37ms |
| 2560 × 1600 | 1:1 | 5.59ms | 3.32ms | 1.7× | 5.78ms | 3.85ms |
| 3840 × 2160 | fit | 10.94ms | 3.88ms | 2.8× | 25.05ms | 4.17ms |
| 3840 × 2160 | 1:1 | 13.31ms | 6.36ms | 2.1× | 27.60ms | 6.93ms |
**Clarity is no longer the stage that dominates.** At 4K fit it is 4.17 ms
against 4.61 ms for all four neighbourhood operations together; it was 97% of
that total at every size before. Every M1 and M2 row now sits inside the 16 ms
budget on this adapter, including the six `all` rows that carried `OVER`.
**The honest headline is the p50 column: 2.8× at 4K.** An earlier draft of this
section led with the p99 ratio, which reads as 6.0× at the same size. That
number is not supported, and the reason it is not is worth recording rather
than quietly deleting.
The baseline run's `fit` rows have a p99/p50 spread of about 2.3×, while every
row of the after run sits between 1.07× and 1.26×. A stage whose cost is
`radius × pixels` has no reason to be bimodal, and the `fit` configuration is
the memory-bound one — it walks the whole 482 MB source on a stride, where
`1:1` reads a contiguous window. Something else was using the machine.
The cross-check settles it. §"Which GPU, on a machine with more than one"
above measured the *same baseline code on the same card* independently, and
reports M3 clarity, fit, 2560 × 1600 at **4.67 ms p99** — against the 10.40 ms
in the table here. Two measurements of one thing that differ by 2.2× mean the
noisier one is wrong, and it is this one.
So: the p50 ratios are the claim. The p99 improvement is real and larger, but
this run cannot say by how much, and a clean re-measurement on a quiet machine
is the way to find out.
What survives the caveat intact is the **shape** of the after column. Every
figure is inside the 16 ms budget with a p99 within 26% of its median, at every
size and both views — which is what a stage that is no longer the bottleneck
looks like, whatever the exact ratio to what it replaced.
### The one thing that is not a pure speed-up
+13 -6
View File
@@ -208,15 +208,22 @@ size on each axis. Measured before and after on the same machine, same adapter,
with only the change between them — see [frame-budget.md](frame-budget.md) §"The reduced base,
measured":
| viewport, fit | before | after | |
|---|---:|---:|---:|
| 1920 × 1200 | 3.94 ms | 1.97 ms | 2.0× |
| 2560 × 1600 | 10.40 ms | 2.37 ms | 4.4× |
| 3840 × 2160 | **25.05 ms** | **4.17 ms** | **6.0×** |
| viewport, fit | before p50 | after p50 | | before p99 | after p99 |
|---|---:|---:|---:|---:|---:|
| 1920 × 1200 | 2.30 ms | 1.56 ms | 1.5× | 3.94 ms | 1.97 ms |
| 2560 × 1600 | 4.41 ms | 1.95 ms | 2.3× | 10.40 ms | 2.37 ms |
| 3840 × 2160 | **10.94 ms** | **3.88 ms** | **2.8×** | 25.05 ms | 4.17 ms |
**Quote the p50 column.** The baseline run's p99 figures are contaminated — its `fit` rows spread
2.3× between median and 99th percentile where the after run spreads 1.1×, and
[frame-budget.md](frame-budget.md)'s own independent measurement of the same baseline on the same
card reports 4.67 ms p99 at 2560 × 1600 against the 10.40 ms here. The p99 improvement is real and
larger than 2.8×; this run cannot say by how much.
Clarity is no longer the stage that misses the budget, and no longer dominates the neighbourhood
stage: at 4K it is 4.17 ms against 4.61 ms for all four neighbourhood operations together, where it
was 97% of that total at every size.
was 97% of that total at every size. That comparison is within one run, so the contention does not
touch it.
**Two things worth recording, because neither is visible in the table.**