The before/after published a p99 ratio of 6.0x at 4K. It is not supported and
the correction is worth more than the number was.
The baseline run's `fit` rows spread 2.3x between median and 99th percentile
while every row of the after run spreads about 1.1x. A stage costing
`radius x pixels` has no reason to be bimodal, and `fit` is the memory-bound
configuration — it walks the whole 482 MB source on a stride where `1:1` reads
a contiguous window. Something else had the machine.
The merge brings in the cross-check that settles it: "Try every GPU, not only
the fastest one" measured the same baseline code on the same card and reports
4.67 ms p99 for M3 clarity fit at 2560x1600, against 10.40 ms here. Two
measurements of one thing differing by 2.2x mean the noisier one is wrong.
So both percentiles are now published and the p50 column is the claim: 2.8x at
4K rather than 6.0x. The p99 improvement is real and larger; this run cannot
say by how much, and says so.
What the caveat does not touch: every after figure is inside the 16 ms budget
with a p99 within 26% of its median at every size and both views, and clarity
against all-four is a within-run comparison.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
TD-4 asked for the measurement as well as the change, and this is it: 25.05 ms
to 4.17 ms at 3840 x 2160, six times faster, with clarity no longer dominating
the neighbourhood stage it used to be 97% of.
Measured before and after on the same machine and the same adapter minutes
apart, baseline at the branch's merge-base, so the only variable is the change.
That adapter is not the RTX 3050 the rest of this document was measured on, so
the new table says to read it on its own rather than against the ones above —
the before/after is comparable, the absolute figures are not, and quietly
replacing the existing tables would have changed the instrument.
Also recorded: the declared halo is now quantised to multiples of the output
scale, because the 2-sigma truncation rounds on the reduced grid. 29 px becomes
28 at 1920x1200 and 38 becomes 40 at 2560x1600. It is inside what the cross-form
test holds — 0.03 stops of peak, 2% of reach — but it is a change in reach and
not only in cost, and a tile scheduler would be handed it. Better written down
now than found later as a seam.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`request_adapter` with `HighPerformance` returns one adapter and no
second chance. That is right on a healthy machine and wrong on one with
a sick GPU, which is not rare: observed 2026-08-29 on a laptop whose
discrete card had hit an NVRM assertion failure and a fullchip reset.
The driver still advertised it, wgpu dutifully picked it as the highest
performing, and the process died on it — while a working integrated GPU
and a working external card sat unused in the same enumeration. A photo
editor that will not start because the *fastest* GPU is broken, on a
machine holding two that are not, is worse than a slow one.
So: enumerate, order by preference, take the first that yields a device.
The ordering reproduces what `HighPerformance` meant, so a healthy
machine picks what it always picked and pays one enumeration for it. A
CPU adapter sorts last rather than being excluded — software rendering
is a poor experience and a working one.
Which GPU to prefer is now a policy rather than an assumption, because
the fastest is not obviously the right one. A 24 MP frame is ~96 MB of
RGBA and every upload and export readback crosses PCIe on a discrete
card, where an integrated GPU shares memory and crosses nothing — and
does not empty a battery.
Measured before choosing a default, on this machine's Iris Xe against
its RX 5700 XT. The fused colour pass is within 1.5x, which is the
shape shared memory suits. The neighbourhood stage is 5-8x slower, and
that decides it: clarity at 1920x1200 costs 20 ms on the iGPU, over the
budget on its own at the smallest size tested. So `Performance` stays
the default and `Efficiency` is offered rather than chosen
(`DARKROOM_GPU=integrated`).
docs/frame-budget.md carries the table, and says what it does *not*
show: the harness renders from a resident texture and never uploads or
reads back, so the transfer cost an iGPU avoids appears in none of it.
Import, export and the thumbnail sweeps may well go the other way.
What this cannot fix: a GPU sick enough to accept `request_device` and
segfault afterwards, which arrives as a driver crash rather than an
error. It moves the boundary from "the preferred adapter is unusable" to
"unusable and dishonest about it".
FR-DSP-3 states a latency requirement and nothing checked it, which makes it a
wish. This adds the check and the measurements it guards.
`docs/frame-budget.md` is the bench's output with the reading of §2's decision
rule attached. The short version: every point-operation chain at every viewport
size, fit and at 1:1, is inside 16 ms at the 99th percentile — the widest is
4.5 ms of GPU at 4K — so FR-DSP-2 should be rewritten rather than implemented.
The measurement did find a stage that misses the budget, and it is the one §2
predicted: clarity's 52-pixel separable kernel costs 34 ms at 4K. Tiles make
that worse rather than better, since a tiled convolution reads a halo per tile;
the fix `local_contrast` already names for itself is a base computed at reduced
resolution.
The test guards the fused path and says so, at length, rather than quietly
excluding the expensive stage and letting the tag imply otherwise (§7). What it
asserts is exactly the claim the recommendation rests on: one dispatch over a
viewport-sized target, at a full chain, is comfortably inside a frame.
Two things the numbers forced:
- The two cases are one `#[test]`. As two they ran on a thread each, contended
for the same device, and took the 1:1 case from 2.5 ms to 14.9 ms — a
measurement of the harness that would have flickered either side of the
budget forever.
- The CPU half of the frame is judged only in an optimised build. Composition
is real per-frame work on the UI thread and belongs in the budget, but the
workspace builds its own crates at `opt-level = 0` in dev and `cargo test` is
a dev build, so measuring it there measures rustc. The GPU half is asserted
either way.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>