FR-CULL-3's focus peaking. One compute dispatch measures local contrast
in WGSL and writes an overlay texture; on desktop it reaches Slint
through the same zero-copy wgpu import the canvas uses, so nothing
per-pixel touches the CPU on the frame path.
With peaking off the cost is zero and structurally so: focus_overlay
opens with `let settings = self.peaking?;` before the frame is touched,
and clearing drops both overlay textures, so no VRAM is held either.
NFR-P14 is met by construction rather than by measurement -- one
dispatch, no second render, no pipeline compile after session open, and
a test asserting allocations stay at 2 over eight frames. The budget
test asserts 50ms at 4K rather than a tight bound, deliberately: a tight
bound fails on a loaded machine and gets deleted, which is worse than a
loose one that still catches the regression that matters.
TD-1 is amended rather than joined by a TD-6: on Android the overlay
rides the readback that already exists there, roughly doubling that
transfer while peaking is on, and TD-1's own "Done when" removes both
because both are the same missing capability.
Verified: cargo fmt clean; clippy --workspace --all-targets -D warnings
green, which also compiles peaking.slint through dr-ui's build.rs; 11
focus GPU tests and 79 baseline dr-gpu tests pass; 511 dr-ui tests pass.
Not verified: the cfg(target_os = "android") arm, which the host-target
clippy never compiled.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The before/after published a p99 ratio of 6.0x at 4K. It is not supported and
the correction is worth more than the number was.
The baseline run's `fit` rows spread 2.3x between median and 99th percentile
while every row of the after run spreads about 1.1x. A stage costing
`radius x pixels` has no reason to be bimodal, and `fit` is the memory-bound
configuration — it walks the whole 482 MB source on a stride where `1:1` reads
a contiguous window. Something else had the machine.
The merge brings in the cross-check that settles it: "Try every GPU, not only
the fastest one" measured the same baseline code on the same card and reports
4.67 ms p99 for M3 clarity fit at 2560x1600, against 10.40 ms here. Two
measurements of one thing differing by 2.2x mean the noisier one is wrong.
So both percentiles are now published and the p50 column is the claim: 2.8x at
4K rather than 6.0x. The p99 improvement is real and larger; this run cannot
say by how much, and says so.
What the caveat does not touch: every after figure is inside the 16 ms budget
with a p99 within 26% of its median at every size and both views, and clarity
against all-four is a within-run comparison.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
TD-4 asked for the measurement as well as the change, and this is it: 25.05 ms
to 4.17 ms at 3840 x 2160, six times faster, with clarity no longer dominating
the neighbourhood stage it used to be 97% of.
Measured before and after on the same machine and the same adapter minutes
apart, baseline at the branch's merge-base, so the only variable is the change.
That adapter is not the RTX 3050 the rest of this document was measured on, so
the new table says to read it on its own rather than against the ones above —
the before/after is comparable, the absolute figures are not, and quietly
replacing the existing tables would have changed the instrument.
Also recorded: the declared halo is now quantised to multiples of the output
scale, because the 2-sigma truncation rounds on the reduced grid. 29 px becomes
28 at 1920x1200 and 38 becomes 40 at 2560x1600. It is inside what the cross-form
test holds — 0.03 stops of peak, 2% of reach — but it is a change in reach and
not only in cost, and a tile scheduler would be handed it. Better written down
now than found later as a seam.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three places, because the finding has three audiences.
The display spec's §1 table said FR-DSP-3 was unmeasured and FR-DSP-5 untagged.
Both are now false, and §2's decision rule has fired. The body of §2 is left as
written with the verdict quoted above it: a plan overtaken by its own evidence
reads better in order than quietly edited into agreement with the outcome.
TD-4 is the stage that misses the budget. Clarity's kernel is a fraction of the
frame, so it reaches a 52-pixel radius at 4K and costs 34 ms — seven times the
entire fused chain, for one slider. It is debt rather than a bug because the
detail stage cannot yet write a target smaller than it reads, which
`local_contrast`'s own documentation has said since it was written. The entry
says plainly that tiles are the wrong tool for it, since that is exactly the
conclusion a reader arriving from ARCH §5.3 would otherwise draw.
TD-5 is the one nobody was looking for: composing the fused shader costs
2.8–5.2 ms of CPU per frame on a full chain, on the UI thread, which at
1920x1200 is more than the dispatch it precedes. The source depends only on the
graph's structure — what `structure_hash` already identifies and what does not
move during a drag — so the fix is the cache `AdjustPass` already keeps for
compiled pipelines, one level up.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three compromises were taken deliberately over the last few days and none of
them was written anywhere a future reader would look. So `docs/technical-debt.md`,
and two corrections to the architecture document that the Android change made
untrue the moment it landed.
**TD-1, the Android readback.** §12/6.1 says GPU results never round-trip
through the CPU, and the develop view on Android now does exactly that. That
is worth recording as a breach with reasons rather than quietly leaving a
constraint the code no longer honours — the next person to read §6.1 and then
`DevelopSession::render` would otherwise conclude one of them is a mistake.
The entry carries the device measurements that forced it, why setting
`preTransform` is not a fix available to us, and the three separate things any
one of which would remove it.
**TD-2 and TD-3**, the serial thumbnail fetch and the unbounded drain, were
found while chasing the tearing and are still outstanding. Both have a known
shape for the fix; neither is a bug, and neither should be discovered again
from scratch.
The architecture document said the app renders "through wgpu to Vulkan on both
Linux and Android", which stopped being true at 6267802, and §12/6.1 claimed a
constraint with no exceptions. Both now say what the code does and point at
the debt entry for why.
The numbers are labelled with what they are. TD-1's readback cost has *not*
been measured on the device and says so, and TD-3's figures are from a debug
build and say so — a documented measurement that quietly turns out to be the
wrong build is worse than no measurement.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>