Record what the measurement found, where the next person will look for it

Three places, because the finding has three audiences.

The display spec's §1 table said FR-DSP-3 was unmeasured and FR-DSP-5 untagged.
Both are now false, and §2's decision rule has fired. The body of §2 is left as
written with the verdict quoted above it: a plan overtaken by its own evidence
reads better in order than quietly edited into agreement with the outcome.

TD-4 is the stage that misses the budget. Clarity's kernel is a fraction of the
frame, so it reaches a 52-pixel radius at 4K and costs 34 ms — seven times the
entire fused chain, for one slider. It is debt rather than a bug because the
detail stage cannot yet write a target smaller than it reads, which
`local_contrast`'s own documentation has said since it was written. The entry
says plainly that tiles are the wrong tool for it, since that is exactly the
conclusion a reader arriving from ARCH §5.3 would otherwise draw.

TD-5 is the one nobody was looking for: composing the fused shader costs
2.8–5.2 ms of CPU per frame on a full chain, on the UI thread, which at
1920x1200 is more than the dispatch it precedes. The source depends only on the
graph's structure — what `structure_hash` already identifies and what does not
move during a drag — so the fix is the cache `AdjustPass` already keeps for
compiled pipelines, one level up.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-08-27 19:06:19 +02:00
co-authored by Claude Opus 5
parent 772a69711d
commit 2e6825ded0
2 changed files with 114 additions and 4 deletions
+17 -4
View File
@@ -20,10 +20,10 @@ and the reason is that some of the work is done and untagged.
| Requirement | Reality |
|---|---|
| FR-DSP-1 proxy rendering | **Done.** The develop view renders at viewport resolution, not source. |
| FR-DSP-2 tiled computation | **Absent.** The fused pass renders the whole viewport in one dispatch. |
| FR-DSP-3 interactive latency | **Unmeasured.** No frame budget is asserted anywhere. |
| FR-DSP-4 progressive refinement | **Absent.** Every render is full quality. |
| FR-DSP-5 zoom and pan | **Substantially done, untagged.** `Framing::view` shrinks the sampled region while the render target keeps its size, so zooming *raises* the resolution the pipeline works at. That is FR-DSP-5's requirement, arrived at without tiles. |
| FR-DSP-2 tiled computation | **Absent, and §2 now says it should stay that way.** Measured: the fused pass is inside the budget everywhere. See [frame-budget.md](frame-budget.md). |
| FR-DSP-3 interactive latency | **Measured and asserted** for the fused path — `core/dr-gpu/tests/frame_budget.rs`. Missed by one operation, clarity, for the reason recorded as TD-4. |
| FR-DSP-4 progressive refinement | **Absent**, and §4's condition did not fire. Every render is full quality and can afford to be. |
| FR-DSP-5 zoom and pan | **Done and tagged**, against tests that fail if the behaviour is removed — `core/dr-gpu/tests/zoom_resolution.rs`. `Framing::view` shrinks the sampled region while the render target keeps its size, so zooming *raises* the resolution the pipeline works at. That is FR-DSP-5's requirement, arrived at without tiles. |
| FR-DSP-6 colour management | **Done.** Output space is a parameter of composition. |
| FR-DSP-7 histogram and clipping | **Done**, GPU-side, no per-frame readback. |
| FR-DSP-8 per-display colour | **Absent.** One transform, not per-display. |
@@ -32,12 +32,25 @@ and the reason is that some of the work is done and untagged.
Two of the five uncovered display requirements are therefore *measurement and tagging*, not
construction. That is worth knowing before anyone plans a quarter around them.
**Both have since been done.** [frame-budget.md](frame-budget.md) holds the measurements §2 asks
for and the reading of its decision rule; the table above is updated to match. The rest of this
document is left as it was written, because a plan that has been overtaken by its own evidence is
more useful read in order than quietly edited into agreement.
---
## 2. Measure before building tiles
**FR-DSP-2 is the one requirement in this document that may not be worth satisfying as written.**
> **Resolved.** M1–M3 were run; the numbers and the verdict are in
> [frame-budget.md](frame-budget.md). The rule below fired for *rewrite*: every point-operation
> chain is inside 16 ms at the 99th percentile at every viewport size, fit and at 1:1, the widest
> being 4.5 ms of GPU at 4K. The measurement did find a stage that misses the budget — clarity's
> 52-pixel kernel, 34 ms at 4K — and tiling makes that stage *worse*, since a tiled convolution
> reads a halo per tile. It is recorded as TD-4 with the fix its own module already names.
The requirement predates the fused-shader design. It assumes the pipeline is a chain of passes over
a large buffer, where recomputing everything on each frame would be ruinous and tiles are the way
out. What was built instead composes every active operation into **one dispatch over a