diff --git a/docs/frame-budget.md b/docs/frame-budget.md index 042036d..b9bf1b2 100644 --- a/docs/frame-budget.md +++ b/docs/frame-budget.md @@ -373,29 +373,45 @@ baseline at `0407fb8` and the result at `bff95e2`. | Result | `bff95e2` | | Date | 2026-08-29 | -### M3 — the neighbourhood stage alone, before and after +### M3 — clarity alone, before and after -| size | stage | view | before p99 | after p99 | | -|------------:|---------:|:-----|-----------:|----------:|-----:| -| 1920 × 1200 | clarity | fit | 3.94ms | 1.97ms | 2.0× | -| 1920 × 1200 | all four | fit | 2.57ms | 1.81ms | 1.4× | -| 1920 × 1200 | clarity | 1:1 | 3.07ms | 2.41ms | 1.3× | -| 2560 × 1600 | clarity | fit | 10.40ms | 2.37ms | 4.4× | -| 2560 × 1600 | all four | fit | 5.50ms | 2.64ms | 2.1× | -| 2560 × 1600 | clarity | 1:1 | 5.78ms | 3.85ms | 1.5× | -| 3840 × 2160 | clarity | fit | **25.05ms**|**4.17ms** | 6.0× | -| 3840 × 2160 | all four | fit | 25.48ms | 4.61ms | 5.5× | -| 3840 × 2160 | clarity | 1:1 | 27.60ms | 6.93ms | 4.0× | -| 3840 × 2160 | all four | 1:1 | 15.88ms | 7.74ms | 2.1× | +Both percentiles, because they disagree and the disagreement is the +interesting part. -The passes went from 2 to 4 and got faster, which is the point: three of the -four now run on a grid a sixteenth the area, and the fourth — the combine — is a -single bilinear read where it used to be a 105-tap convolution. +| size | view | before p50 | after p50 | | before p99 | after p99 | +|------------:|:-----|-----------:|----------:|-----:|-----------:|----------:| +| 1920 × 1200 | fit | 2.30ms | 1.56ms | 1.5× | 3.94ms | 1.97ms | +| 1920 × 1200 | 1:1 | 2.88ms | 1.99ms | 1.4× | 3.07ms | 2.41ms | +| 2560 × 1600 | fit | 4.41ms | 1.95ms | 2.3× | 10.40ms | 2.37ms | +| 2560 × 1600 | 1:1 | 5.59ms | 3.32ms | 1.7× | 5.78ms | 3.85ms | +| 3840 × 2160 | fit | 10.94ms | 3.88ms | 2.8× | 25.05ms | 4.17ms | +| 3840 × 2160 | 1:1 | 13.31ms | 6.36ms | 2.1× | 27.60ms | 6.93ms | -**Clarity is no longer the stage that dominates.** At 4K fit it is 4.17 ms -against 4.61 ms for all four neighbourhood operations together; it was 97% of -that total at every size before. Every M1 and M2 row now sits inside the 16 ms -budget on this adapter, including the six `all` rows that carried `OVER`. +**The honest headline is the p50 column: 2.8× at 4K.** An earlier draft of this +section led with the p99 ratio, which reads as 6.0× at the same size. That +number is not supported, and the reason it is not is worth recording rather +than quietly deleting. + +The baseline run's `fit` rows have a p99/p50 spread of about 2.3×, while every +row of the after run sits between 1.07× and 1.26×. A stage whose cost is +`radius × pixels` has no reason to be bimodal, and the `fit` configuration is +the memory-bound one — it walks the whole 482 MB source on a stride, where +`1:1` reads a contiguous window. Something else was using the machine. + +The cross-check settles it. §"Which GPU, on a machine with more than one" +above measured the *same baseline code on the same card* independently, and +reports M3 clarity, fit, 2560 × 1600 at **4.67 ms p99** — against the 10.40 ms +in the table here. Two measurements of one thing that differ by 2.2× mean the +noisier one is wrong, and it is this one. + +So: the p50 ratios are the claim. The p99 improvement is real and larger, but +this run cannot say by how much, and a clean re-measurement on a quiet machine +is the way to find out. + +What survives the caveat intact is the **shape** of the after column. Every +figure is inside the 16 ms budget with a p99 within 26% of its median, at every +size and both views — which is what a stage that is no longer the bottleneck +looks like, whatever the exact ratio to what it replaced. ### The one thing that is not a pure speed-up diff --git a/docs/technical-debt.md b/docs/technical-debt.md index 406aadd..057157c 100644 --- a/docs/technical-debt.md +++ b/docs/technical-debt.md @@ -208,15 +208,22 @@ size on each axis. Measured before and after on the same machine, same adapter, with only the change between them — see [frame-budget.md](frame-budget.md) §"The reduced base, measured": -| viewport, fit | before | after | | -|---|---:|---:|---:| -| 1920 × 1200 | 3.94 ms | 1.97 ms | 2.0× | -| 2560 × 1600 | 10.40 ms | 2.37 ms | 4.4× | -| 3840 × 2160 | **25.05 ms** | **4.17 ms** | **6.0×** | +| viewport, fit | before p50 | after p50 | | before p99 | after p99 | +|---|---:|---:|---:|---:|---:| +| 1920 × 1200 | 2.30 ms | 1.56 ms | 1.5× | 3.94 ms | 1.97 ms | +| 2560 × 1600 | 4.41 ms | 1.95 ms | 2.3× | 10.40 ms | 2.37 ms | +| 3840 × 2160 | **10.94 ms** | **3.88 ms** | **2.8×** | 25.05 ms | 4.17 ms | + +**Quote the p50 column.** The baseline run's p99 figures are contaminated — its `fit` rows spread +2.3× between median and 99th percentile where the after run spreads 1.1×, and +[frame-budget.md](frame-budget.md)'s own independent measurement of the same baseline on the same +card reports 4.67 ms p99 at 2560 × 1600 against the 10.40 ms here. The p99 improvement is real and +larger than 2.8×; this run cannot say by how much. Clarity is no longer the stage that misses the budget, and no longer dominates the neighbourhood stage: at 4K it is 4.17 ms against 4.61 ms for all four neighbourhood operations together, where it -was 97% of that total at every size. +was 97% of that total at every size. That comparison is within one run, so the contention does not +touch it. **Two things worth recording, because neither is visible in the table.**