Quote the percentile the measurement can actually support
The before/after published a p99 ratio of 6.0x at 4K. It is not supported and the correction is worth more than the number was. The baseline run's `fit` rows spread 2.3x between median and 99th percentile while every row of the after run spreads about 1.1x. A stage costing `radius x pixels` has no reason to be bimodal, and `fit` is the memory-bound configuration — it walks the whole 482 MB source on a stride where `1:1` reads a contiguous window. Something else had the machine. The merge brings in the cross-check that settles it: "Try every GPU, not only the fastest one" measured the same baseline code on the same card and reports 4.67 ms p99 for M3 clarity fit at 2560x1600, against 10.40 ms here. Two measurements of one thing differing by 2.2x mean the noisier one is wrong. So both percentiles are now published and the p50 column is the claim: 2.8x at 4K rather than 6.0x. The p99 improvement is real and larger; this run cannot say by how much, and says so. What the caveat does not touch: every after figure is inside the 16 ms budget with a p99 within 26% of its median at every size and both views, and clarity against all-four is a within-run comparison. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
+13
-6
@@ -208,15 +208,22 @@ size on each axis. Measured before and after on the same machine, same adapter,
|
||||
with only the change between them — see [frame-budget.md](frame-budget.md) §"The reduced base,
|
||||
measured":
|
||||
|
||||
| viewport, fit | before | after | |
|
||||
|---|---:|---:|---:|
|
||||
| 1920 × 1200 | 3.94 ms | 1.97 ms | 2.0× |
|
||||
| 2560 × 1600 | 10.40 ms | 2.37 ms | 4.4× |
|
||||
| 3840 × 2160 | **25.05 ms** | **4.17 ms** | **6.0×** |
|
||||
| viewport, fit | before p50 | after p50 | | before p99 | after p99 |
|
||||
|---|---:|---:|---:|---:|---:|
|
||||
| 1920 × 1200 | 2.30 ms | 1.56 ms | 1.5× | 3.94 ms | 1.97 ms |
|
||||
| 2560 × 1600 | 4.41 ms | 1.95 ms | 2.3× | 10.40 ms | 2.37 ms |
|
||||
| 3840 × 2160 | **10.94 ms** | **3.88 ms** | **2.8×** | 25.05 ms | 4.17 ms |
|
||||
|
||||
**Quote the p50 column.** The baseline run's p99 figures are contaminated — its `fit` rows spread
|
||||
2.3× between median and 99th percentile where the after run spreads 1.1×, and
|
||||
[frame-budget.md](frame-budget.md)'s own independent measurement of the same baseline on the same
|
||||
card reports 4.67 ms p99 at 2560 × 1600 against the 10.40 ms here. The p99 improvement is real and
|
||||
larger than 2.8×; this run cannot say by how much.
|
||||
|
||||
Clarity is no longer the stage that misses the budget, and no longer dominates the neighbourhood
|
||||
stage: at 4K it is 4.17 ms against 4.61 ms for all four neighbourhood operations together, where it
|
||||
was 97% of that total at every size.
|
||||
was 97% of that total at every size. That comparison is within one run, so the contention does not
|
||||
touch it.
|
||||
|
||||
**Two things worth recording, because neither is visible in the table.**
|
||||
|
||||
|
||||
Reference in New Issue
Block a user