Merge branch 'worktree-agent-aa9f4356c13893373' into master
# Conflicts: # docs/traceability.md
This commit is contained in:
@@ -20,10 +20,10 @@ and the reason is that some of the work is done and untagged.
|
||||
| Requirement | Reality |
|
||||
|---|---|
|
||||
| FR-DSP-1 proxy rendering | **Done.** The develop view renders at viewport resolution, not source. |
|
||||
| FR-DSP-2 tiled computation | **Absent.** The fused pass renders the whole viewport in one dispatch. |
|
||||
| FR-DSP-3 interactive latency | **Unmeasured.** No frame budget is asserted anywhere. |
|
||||
| FR-DSP-4 progressive refinement | **Absent.** Every render is full quality. |
|
||||
| FR-DSP-5 zoom and pan | **Substantially done, untagged.** `Framing::view` shrinks the sampled region while the render target keeps its size, so zooming *raises* the resolution the pipeline works at. That is FR-DSP-5's requirement, arrived at without tiles. |
|
||||
| FR-DSP-2 tiled computation | **Absent, and §2 now says it should stay that way.** Measured: the fused pass is inside the budget everywhere. See [frame-budget.md](frame-budget.md). |
|
||||
| FR-DSP-3 interactive latency | **Measured and asserted** for the fused path — `core/dr-gpu/tests/frame_budget.rs`. Missed by one operation, clarity, for the reason recorded as TD-4. |
|
||||
| FR-DSP-4 progressive refinement | **Absent**, and §4's condition did not fire. Every render is full quality and can afford to be. |
|
||||
| FR-DSP-5 zoom and pan | **Done and tagged**, against tests that fail if the behaviour is removed — `core/dr-gpu/tests/zoom_resolution.rs`. `Framing::view` shrinks the sampled region while the render target keeps its size, so zooming *raises* the resolution the pipeline works at. That is FR-DSP-5's requirement, arrived at without tiles. |
|
||||
| FR-DSP-6 colour management | **Done.** Output space is a parameter of composition. |
|
||||
| FR-DSP-7 histogram and clipping | **Done**, GPU-side, no per-frame readback. |
|
||||
| FR-DSP-8 per-display colour | **Done**, with one caveat named in §5.4. Acquisition per display server, an sRGB fallback that is visible in About, and the canvas rendered at physical pixel size. |
|
||||
@@ -32,12 +32,25 @@ and the reason is that some of the work is done and untagged.
|
||||
Two of the five uncovered display requirements are therefore *measurement and tagging*, not
|
||||
construction. That is worth knowing before anyone plans a quarter around them.
|
||||
|
||||
**Both have since been done.** [frame-budget.md](frame-budget.md) holds the measurements §2 asks
|
||||
for and the reading of its decision rule; the table above is updated to match. The rest of this
|
||||
document is left as it was written, because a plan that has been overtaken by its own evidence is
|
||||
more useful read in order than quietly edited into agreement.
|
||||
|
||||
---
|
||||
|
||||
## 2. Measure before building tiles
|
||||
|
||||
**FR-DSP-2 is the one requirement in this document that may not be worth satisfying as written.**
|
||||
|
||||
> **Resolved.** M1–M3 were run; the numbers and the verdict are in
|
||||
> [frame-budget.md](frame-budget.md). The rule below fired for *rewrite*: every point-operation
|
||||
> chain is inside 16 ms at the 99th percentile at every viewport size, fit and at 1:1, the widest
|
||||
> being 4.5 ms of GPU at 4K. The measurement did find a stage that misses the budget — clarity's
|
||||
> 52-pixel kernel, 34 ms at 4K — and tiling makes that stage *worse*, since a tiled convolution
|
||||
> reads a halo per tile. It is recorded as TD-4 with the fix its own module already names.
|
||||
|
||||
|
||||
The requirement predates the fused-shader design. It assumes the pipeline is a chain of passes over
|
||||
a large buffer, where recomputing everything on each frame would be ruinous and tiles are the way
|
||||
out. What was built instead composes every active operation into **one dispatch over a
|
||||
|
||||
@@ -0,0 +1,297 @@
|
||||
# What a frame costs
|
||||
|
||||
**Status:** Measured · 2026-08-27
|
||||
**Companion to:** [display-and-extension.md](display-and-extension.md) §2–3 ·
|
||||
[requirements.md](requirements.md) §3.4 FR-DSP-2, FR-DSP-3, FR-DSP-4
|
||||
**Instrument:** [`core/dr-gpu/examples/frame_budget.rs`](../core/dr-gpu/examples/frame_budget.rs)
|
||||
**Guard:** [`core/dr-gpu/tests/frame_budget.rs`](../core/dr-gpu/tests/frame_budget.rs)
|
||||
|
||||
[display-and-extension.md](display-and-extension.md) §2 fixed a decision rule in
|
||||
advance and made three measurements the thing that settles it. This file is
|
||||
those measurements, and the recommendation they support.
|
||||
|
||||
Rerun with:
|
||||
|
||||
```sh
|
||||
cargo run --release -p dr-gpu --example frame_budget
|
||||
```
|
||||
|
||||
and diff this file. That is the whole point of committing numbers: a regression
|
||||
should be a diff rather than somebody's recollection of how fast it used to be.
|
||||
|
||||
---
|
||||
|
||||
## The answer, first
|
||||
|
||||
**FR-DSP-2 should be rewritten, not implemented.** M1 and M2 sit inside the
|
||||
16 ms budget at the 99th percentile for every chain of point operations at every
|
||||
viewport size measured, fit and at 1:1 — the widest case, every operation that
|
||||
contributes a fragment to the fused shader at 4K, costs **4.5 ms** on the GPU and
|
||||
**8.2 ms** including the composition that precedes it. Tiling the interactive
|
||||
path would be optimising something that is already using a quarter of its budget.
|
||||
|
||||
**But the measurement did find a budget-breaker, and it is not the one tiling
|
||||
fixes.** The neighbourhood stage — clarity in particular — costs **34 ms at 4K
|
||||
on its own**, twice the whole budget, and tiles do not help it: a tile of a
|
||||
convolution has to read its halo, so tiling raises the total tap count rather
|
||||
than lowering it. §2 predicted this exactly ("a separable blur at a large radius
|
||||
is the plausible budget-breaker, not the fused pass"), and the fix it needs is
|
||||
the one `local_contrast`'s own module documentation already names — a base
|
||||
computed at reduced resolution — which is a change to `crate::detail`, not a
|
||||
tile scheduler.
|
||||
|
||||
There is a third finding nobody was looking for: **shader composition costs
|
||||
3–5 ms of CPU per frame on a full chain**, on the UI thread, before any GPU work
|
||||
is submitted. That is a fifth to a third of the budget spent formatting strings,
|
||||
and it is invisible to any amount of tiling.
|
||||
|
||||
---
|
||||
|
||||
## Conditions
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| Adapter | NVIDIA GeForce RTX 3050 6GB Laptop GPU (Vulkan) |
|
||||
| Source | 9504 × 6336 synthetic (60.2 MP, 482 MB as `rgba16f`) |
|
||||
| Frames | 100 measured per row, 12 warm-up frames discarded |
|
||||
| Percentile | Nearest-rank, so p99 of 100 frames is the second-worst frame |
|
||||
| Build | `--release` |
|
||||
| Date | 2026-08-27 |
|
||||
|
||||
`shader` is `EditGraph::compose` alone. `cpu` adds the detail chain and the
|
||||
invalidation hash — everything `DevelopSession::render` does per frame before it
|
||||
dispatches. `gpu` is submit plus wait-for-idle, which serialises the GPU work
|
||||
into the frame that caused it and is therefore pessimistic. `TOTAL` ranks
|
||||
`cpu + gpu` summed **within each frame**, which is the column the budget is
|
||||
judged on; adding two percentiles instead would invent a stutter that no frame
|
||||
actually had.
|
||||
|
||||
Chains: `one` is exposure. `five` is exposure, contrast, highlights/shadows,
|
||||
blacks/whites, vibrance. `point` is every operation in the default chain that
|
||||
contributes a fragment to the fused shader, film stock included. `all` is `point`
|
||||
plus the four neighbourhood operations — noise reduction, capture sharpening,
|
||||
clarity and texture.
|
||||
|
||||
---
|
||||
|
||||
## M1 — the fused pass at proxy resolution
|
||||
|
||||
The develop view: the whole frame fit to the viewport.
|
||||
|
||||
| size | chain | shader | cpu p99 | gpu p50 | gpu p99 | TOTAL | |
|
||||
|------------:|------:|-------:|--------:|--------:|--------:|--------:|:-----|
|
||||
| 1920 × 1200 | one | 0.08ms | 0.10ms | 1.02ms | 1.23ms | 1.31ms | |
|
||||
| 1920 × 1200 | five | 0.18ms | 0.20ms | 1.01ms | 1.20ms | 1.36ms | |
|
||||
| 1920 × 1200 | point | 2.79ms | 2.82ms | 1.98ms | 2.18ms | 4.83ms | |
|
||||
| 1920 × 1200 | all | 3.65ms | 4.73ms | 6.86ms | 7.37ms | 12.02ms | |
|
||||
| 2560 × 1600 | one | 0.08ms | 0.10ms | 1.73ms | 2.00ms | 2.12ms | |
|
||||
| 2560 × 1600 | five | 0.24ms | 0.27ms | 1.73ms | 2.26ms | 2.46ms | |
|
||||
| 2560 × 1600 | point | 2.82ms | 2.85ms | 2.37ms | 2.65ms | 5.38ms | |
|
||||
| 2560 × 1600 | all | 3.37ms | 4.35ms | 14.31ms | 15.65ms | 18.42ms | OVER |
|
||||
| 3840 × 2160 | one | 0.10ms | 0.14ms | 3.09ms | 3.31ms | 3.42ms | |
|
||||
| 3840 × 2160 | five | 0.30ms | 0.32ms | 3.03ms | 3.40ms | 3.61ms | |
|
||||
| 3840 × 2160 | point | 3.62ms | 3.65ms | 4.12ms | 4.52ms | 8.23ms | |
|
||||
| 3840 × 2160 | all | 4.12ms | 5.07ms | 37.73ms | 40.17ms | 43.24ms | OVER |
|
||||
|
||||
Read the `point` rows: **the fused dispatch scales with pixels and almost not at
|
||||
all with chain length.** Going from one operation to the entire point chain at
|
||||
4K costs 1.2 ms of GPU. Going from 2.3 M pixels to 8.3 M costs 2.3 ms. Both are
|
||||
small, and the second is the one tiling would address.
|
||||
|
||||
The `all` rows go over, and the `point` rows in the same block are what say why:
|
||||
the difference between them is the neighbourhood stage, measured on its own in
|
||||
M3 and arriving at almost exactly the same figure.
|
||||
|
||||
## M2 — the same, zoomed to 1:1 on the 60 MP source
|
||||
|
||||
FR-DSP-5's case. `Framing::view` shrinks the sampled region while the render
|
||||
target keeps its size, so one render pixel lands on one source pixel.
|
||||
|
||||
| size | chain | shader | cpu p99 | gpu p50 | gpu p99 | TOTAL | |
|
||||
|------------:|------:|-------:|--------:|--------:|--------:|--------:|:-----|
|
||||
| 1920 × 1200 | one | 0.12ms | 0.14ms | 0.42ms | 0.66ms | 0.75ms | |
|
||||
| 1920 × 1200 | five | 0.27ms | 0.30ms | 0.49ms | 1.14ms | 1.22ms | |
|
||||
| 1920 × 1200 | point | 3.49ms | 3.52ms | 1.18ms | 1.39ms | 4.85ms | |
|
||||
| 1920 × 1200 | all | 3.64ms | 5.18ms | 8.96ms | 9.55ms | 14.30ms | |
|
||||
| 2560 × 1600 | one | 0.11ms | 0.12ms | 0.56ms | 0.99ms | 1.06ms | |
|
||||
| 2560 × 1600 | five | 0.22ms | 0.25ms | 0.77ms | 1.02ms | 1.17ms | |
|
||||
| 2560 × 1600 | point | 2.96ms | 2.99ms | 2.03ms | 2.52ms | 5.61ms | |
|
||||
| 2560 × 1600 | all | 5.24ms | 7.35ms | 18.80ms | 21.62ms | 25.81ms | OVER |
|
||||
| 3840 × 2160 | one | 0.10ms | 0.12ms | 1.31ms | 1.52ms | 1.63ms | |
|
||||
| 3840 × 2160 | five | 0.14ms | 0.27ms | 1.39ms | 1.64ms | 1.75ms | |
|
||||
| 3840 × 2160 | point | 3.14ms | 3.16ms | 4.04ms | 4.50ms | 7.21ms | |
|
||||
| 3840 × 2160 | all | 4.84ms | 6.78ms | 47.22ms | 48.79ms | 54.47ms | OVER |
|
||||
|
||||
**A 1:1 view of a 60 MP file is cheaper than the fit view of the same file**, for
|
||||
every point chain and at every size — 1.52 ms against 3.31 ms for one operation
|
||||
at 4K. That is not a rounding artefact and it is worth stating plainly, because
|
||||
it is the opposite of what "full resolution" sounds like it should cost. The
|
||||
dispatch is the same number of pixels either way; what changes is where those
|
||||
pixels read from. A fit view walks the whole 482 MB texture on a stride, and a
|
||||
1:1 view reads a contiguous window of it that fits comfortably in cache.
|
||||
|
||||
So the resolution FR-DSP-5 promises costs nothing extra on the fused path.
|
||||
Zooming is not an expensive mode to be dreaded and progressively refined into;
|
||||
it is the cheap one.
|
||||
|
||||
The `all` rows are worse at 1:1 than fit, and that is the detail stage again for
|
||||
a specific reason: noise reduction's radius is stated in *source* pixels, so
|
||||
`RenderScale::ratio` climbing to 1.0 widens its kernel. Clarity's is stated as a
|
||||
fraction of the frame and does not move. M3 separates the two.
|
||||
|
||||
## M3 — the neighbourhood stage alone
|
||||
|
||||
Timed with the fused dispatch deliberately reused: only a detail parameter moves,
|
||||
so `render_detailed` skips the colour pass (FR-DEV-3d) and what remains is the
|
||||
convolutions. `colour` counts fused dispatches over the measured frames and is
|
||||
zero on every row, which is what makes these numbers mean "detail alone" rather
|
||||
than asserting it.
|
||||
|
||||
| size | stage | view | pass | radius | colour | cpu p99 | p50 | p99 |
|
||||
|------------:|---------:|:-----|-----:|-------:|-------:|--------:|--------:|--------:|
|
||||
| 1920 × 1200 | clarity | fit | 2 | 29 | 0 | 1.60ms | 5.42ms | 5.99ms |
|
||||
| 1920 × 1200 | all four | fit | 7 | 29 | 0 | 1.71ms | 5.78ms | 6.16ms |
|
||||
| 1920 × 1200 | clarity | 1:1 | 2 | 29 | 0 | 1.12ms | 7.51ms | 8.01ms |
|
||||
| 1920 × 1200 | all four | 1:1 | 9 | 29 | 0 | 2.71ms | 8.25ms | 9.11ms |
|
||||
| 2560 × 1600 | clarity | fit | 2 | 38 | 0 | 1.03ms | 12.02ms | 12.44ms |
|
||||
| 2560 × 1600 | all four | fit | 7 | 38 | 0 | 1.87ms | 12.49ms | 13.16ms |
|
||||
| 2560 × 1600 | clarity | 1:1 | 2 | 38 | 0 | 1.87ms | 15.82ms | 16.60ms |
|
||||
| 2560 × 1600 | all four | 1:1 | 9 | 38 | 0 | 2.71ms | 17.24ms | 18.06ms |
|
||||
| 3840 × 2160 | clarity | fit | 2 | 52 | 0 | 1.75ms | 33.11ms | 33.89ms |
|
||||
| 3840 × 2160 | all four | fit | 7 | 52 | 0 | 1.76ms | 34.21ms | 35.03ms |
|
||||
| 3840 × 2160 | clarity | 1:1 | 2 | 52 | 0 | 1.08ms | 40.39ms | 41.86ms |
|
||||
| 3840 × 2160 | all four | 1:1 | 9 | 52 | 0 | 2.37ms | 43.29ms | 44.72ms |
|
||||
|
||||
`radius` is the widest halo any pass reads, in render pixels.
|
||||
|
||||
Clarity alone is 97% of the cost of all four neighbourhood operations together,
|
||||
at every size. Its σ is 1.2% of the shorter edge and it truncates at 2σ, so its
|
||||
radius is 29 px on a 1200 px viewport and **52 px at 4K** — two separable passes
|
||||
of 105 taps each, over 8.3 M pixels, which is 1.7 billion texture reads. That is
|
||||
the whole of the problem, and the numbers scale as `radius × pixels` exactly as
|
||||
that description predicts: 5.99 → 12.44 → 33.89 ms for radii of 29 → 38 → 52 over
|
||||
2.3 → 4.1 → 8.3 M pixels.
|
||||
|
||||
The extra cost at 1:1 is noise reduction and capture sharpening, whose radii are
|
||||
properties of the sensor rather than of the frame. That is the correct behaviour
|
||||
— it is why `RenderScale` has two units — and it is bounded by the kernel caps
|
||||
those operations already declare.
|
||||
|
||||
---
|
||||
|
||||
## Reading this against §2's decision rule
|
||||
|
||||
§2: *"If M1 and M2 sit inside 16 ms at the 99th percentile, FR-DSP-2 is
|
||||
rewritten rather than implemented … If they do not, the measurement tells us
|
||||
which stage to tile."*
|
||||
|
||||
Both halves of the rule fire, on different stages, and the honest reading takes
|
||||
both.
|
||||
|
||||
### FR-DSP-2 — rewrite it
|
||||
|
||||
For the fused pass the rule passes with a wide margin. Every point chain at
|
||||
every size, fit and at 1:1, is inside 16 ms — the worst `TOTAL` is 8.23 ms and
|
||||
the worst GPU figure is 4.52 ms. There is no viewport size on a desktop display
|
||||
where recomputing the entire point chain over every visible pixel is a problem.
|
||||
|
||||
Two further reasons not to build the tile scheduler as written:
|
||||
|
||||
1. **Panning, which is the case ARCH §5.3's tile cache is designed for, gets no
|
||||
benefit here.** Reusing already-valid tiles saves recomputation. Recomputing
|
||||
the whole 4K viewport costs 4.5 ms, so a perfect tile cache could save at most
|
||||
4.5 ms of a 16 ms budget, at the price of a cache keyed by
|
||||
`(VersionId, tile, zoom, graph_hash_prefix)` that has to stay correct across
|
||||
every parameter change in the graph. That is a large correctness surface
|
||||
bought with a small number.
|
||||
|
||||
2. **It would make the actual problem worse.** The stage that misses the budget
|
||||
is a convolution, and a tiled convolution reads a halo per tile. At a 52-pixel
|
||||
radius, 256-pixel tiles would read (256+104)² instead of 256² — very nearly
|
||||
*twice* the taps. Tiling is the wrong tool for the one stage that needs a
|
||||
tool.
|
||||
|
||||
So FR-DSP-2 becomes what §2 said it actually is for this architecture: a
|
||||
scheduling concern for export and thumbnailing, both of which already run off
|
||||
the frame path. The interactive path does not tile.
|
||||
|
||||
### The stage that does need work — and it is not tiling
|
||||
|
||||
The measurement's real product is naming the stage. It is `local_contrast`, and
|
||||
the fix is stated in that module's own documentation:
|
||||
|
||||
> The right optimisation is a base computed at reduced resolution, which needs a
|
||||
> detail stage that can write a smaller target than it reads; that is a change to
|
||||
> `crate::detail`, not to this file.
|
||||
|
||||
A Gaussian base at a quarter resolution is 1/16 the pixels at 1/4 the radius —
|
||||
about 1/64 of the work — and the result is visually identical because a base at
|
||||
σ = 26 px has no content above the quarter-resolution Nyquist to lose. That is a
|
||||
change to two files with a bounded blast radius, and it is what the 34 ms buys
|
||||
back. It should be tracked as its own item rather than smuggled in under a
|
||||
requirement about tiles.
|
||||
|
||||
### FR-DSP-3 — the clause that should be narrowed
|
||||
|
||||
§3.3 proposes narrowing "when a full-resolution result is needed it is computed
|
||||
asynchronously, and the proxy result remains on screen until it is ready" to
|
||||
export and 1:1 zoom, or striking it.
|
||||
|
||||
**M2 says strike it.** The clause exists to hide the latency of a
|
||||
full-resolution render behind a proxy. There is no such latency: the 1:1 view is
|
||||
*faster* than the fit view on the fused path, and there is no second
|
||||
full-resolution code path to be asynchronous about — `Framing::view` is the
|
||||
whole mechanism. Export renders its own frames on a worker already. Keeping the
|
||||
clause would mean building a progressive-swap machine to conceal a render that
|
||||
completes in 1.4 ms.
|
||||
|
||||
### FR-DSP-4 — satisfied vacuously, on the fused path
|
||||
|
||||
§4 makes progressive refinement conditional on M1 failing. On the fused path M1
|
||||
passes, so reduced-quality rendering during a drag would buy nothing and cost the
|
||||
visible softness the requirement itself warns against.
|
||||
|
||||
The neighbourhood stage is the exception, and it is worth being precise: what
|
||||
that stage needs is not *progressive* refinement — it is a permanently cheaper
|
||||
base, computed at reduced resolution and correct at any moment the user stops.
|
||||
"Render coarse while dragging, sharpen when it settles" would paper over the same
|
||||
34 ms with a visible swap. Fix the stage.
|
||||
|
||||
---
|
||||
|
||||
## What is not measured here
|
||||
|
||||
Stated because §7 of [display-and-extension.md](display-and-extension.md) asks
|
||||
for it, and because each of these could move the numbers.
|
||||
|
||||
- **Local adjustments.** The mask stack is a separate chain per layer and is not
|
||||
in any row above. `render_masked` takes them and the fused shader addresses
|
||||
them per layer, so a heavily masked edit costs more than `all`.
|
||||
- **Spot repairs.** These add detail passes, and their cost is per spot.
|
||||
- **Lens corrections.** Not part of `EditGraph::default_chain` — they are built
|
||||
from a matched profile — so the `point` row does not include the warp chain.
|
||||
- **Demosaic.** Once per photograph on a worker, not on the frame path.
|
||||
- **Presentation.** The bench waits for the device to go idle inside the frame it
|
||||
measures. A real compositor overlaps frames, so these figures are an upper
|
||||
bound rather than an estimate.
|
||||
- **One adapter.** A discrete laptop GPU. The Intel iGPU on the same machine, and
|
||||
Android, will be slower — which is an argument for the conclusion rather than
|
||||
against it: the stage with no headroom has none to lose.
|
||||
|
||||
## The CPU finding, which deserves its own item
|
||||
|
||||
`EditGraph::compose` costs 2.8–5.2 ms per frame on a full chain, at every
|
||||
resolution, because it is resolution-independent: it assembles a WGSL string and
|
||||
hashes it. On the `all` rows it is a third of what is left of the budget after
|
||||
the GPU has taken its share, and at 1920 × 1200 it is larger than the entire
|
||||
fused dispatch.
|
||||
|
||||
Nothing in this document's recommendations changes it, and it is the cheapest
|
||||
remaining win. The generated *source* depends only on the structure of the graph
|
||||
— that is what `structure_hash` already identifies, and it is precisely what does
|
||||
not change while a slider is being dragged, which is why the pipeline cache in
|
||||
`AdjustPass` does not recompile. The uniforms do change, but assembling them is a
|
||||
handful of floats per operation. So caching the source string against the
|
||||
structure hash and rebuilding only the uniforms would take these milliseconds to
|
||||
approximately nothing, on the path that needs them most. Worth its own entry in
|
||||
[technical-debt.md](technical-debt.md).
|
||||
@@ -142,6 +142,103 @@ in well under a second.
|
||||
|
||||
---
|
||||
|
||||
## TD-4 — The local-contrast base is computed at full render resolution
|
||||
|
||||
**Where:** `dr_pipeline::ops::local_contrast::LocalContrast::passes` — the `base` and `combine`
|
||||
passes, and the stage that dispatches them, `dr_pipeline::detail`.
|
||||
|
||||
Breaks **FR-DSP-3** at large viewports. Measured, and the numbers are in
|
||||
[frame-budget.md](frame-budget.md) §M3.
|
||||
|
||||
### What it does
|
||||
|
||||
Clarity's Gaussian σ is 1.2% of the frame's shorter edge, truncated at 2σ, so its kernel radius is
|
||||
a property of the *viewport*: 29 px at 1920 × 1200, 38 px at 2560 × 1600, **52 px at 4K**. The two
|
||||
separable passes therefore run 105 taps each over 8.3 M pixels at 4K, which is 1.7 billion texture
|
||||
reads for one control.
|
||||
|
||||
| viewport | radius | clarity alone, p99 |
|
||||
|---|---:|---:|
|
||||
| 1920 × 1200 | 29 | 5.99 ms |
|
||||
| 2560 × 1600 | 38 | 12.44 ms |
|
||||
| 3840 × 2160 | 52 | **33.89 ms** |
|
||||
|
||||
RTX 3050 laptop, `examples/frame_budget`, fused dispatch reused so this is the convolutions alone.
|
||||
Clarity is 97% of the cost of all four neighbourhood operations together at every size.
|
||||
|
||||
For scale: the entire fused chain — every point operation active, film stock included — costs
|
||||
4.5 ms at the same 4K viewport. **A single slider is seven times the rest of the pipeline.**
|
||||
|
||||
### Why
|
||||
|
||||
Because the stage cannot do otherwise yet. `dr_pipeline::detail` dispatches every pass at the
|
||||
render size; there is no way to express "read this target and write a smaller one". The module's own
|
||||
documentation has said so since it was written:
|
||||
|
||||
> The right optimisation is a base computed at reduced resolution, which needs a detail stage that
|
||||
> can write a smaller target than it reads; that is a change to `crate::detail`, not to this file.
|
||||
|
||||
It was the right call to ship the correct answer slowly rather than a fast approximation nobody had
|
||||
checked — the halo behaviour is the hard part of this operation and it is tested.
|
||||
|
||||
### Not a tiling problem
|
||||
|
||||
Worth saying because ARCH §5.3 offers a tile cache and this is the stage that looks like it wants
|
||||
one. It does not: a tiled convolution reads a halo per tile, so at a 52-pixel radius, 256-pixel
|
||||
tiles would read (256 + 104)² instead of 256² — very nearly **twice** the taps.
|
||||
[display-and-extension.md](display-and-extension.md) §2's decision rule was resolved on this
|
||||
evidence; see [frame-budget.md](frame-budget.md).
|
||||
|
||||
### Paying it off
|
||||
|
||||
A detail pass that declares an output scale, so the base can be computed at a quarter resolution and
|
||||
sampled back up in `combine`. A quarter-resolution base is 1/16 the pixels at 1/4 the radius —
|
||||
about **1/64 of the work** — and is visually identical, because a base at σ = 26 px holds no content
|
||||
above the quarter-resolution Nyquist to lose. Texture's σ is a decade finer and must stay at full
|
||||
resolution; the scale therefore belongs on the `DetailPass`, not on the stage.
|
||||
|
||||
**Done when:** clarity at 100% is inside the frame budget at 3840 × 2160, the halo tests in
|
||||
`tests/local_contrast.rs` still pass unchanged, and `examples/frame_budget`'s M3 table in
|
||||
[frame-budget.md](frame-budget.md) has been rerun and committed.
|
||||
|
||||
---
|
||||
|
||||
## TD-5 — The fused shader is reassembled from strings on every frame
|
||||
|
||||
**Where:** `dr_pipeline::operation::compose_full`, called from `DevelopSession::render`.
|
||||
|
||||
### What it does
|
||||
|
||||
`EditGraph::compose` walks the active operations and formats a WGSL source string, per frame, on
|
||||
the UI thread. On a full chain that is **2.8–5.2 ms** — at 1920 × 1200 it is larger than the entire
|
||||
fused dispatch it precedes, and on the chains that also carry a detail stage it is a third of what
|
||||
is left of the 16 ms budget after the GPU has taken its share. It does not vary with resolution,
|
||||
because it is not pixel work.
|
||||
|
||||
### Why
|
||||
|
||||
Because it was free until the chain got long. Composition was written when an edit was two or three
|
||||
operations, and the cost is roughly linear in generated source: the colour mixer emits twelve hue bands
|
||||
and the tone curve emits a spline evaluator, so a full chain is a large string built from scratch
|
||||
sixty times a second.
|
||||
|
||||
### Paying it off
|
||||
|
||||
The generated source depends only on the *structure* of the graph — which is precisely what
|
||||
`ComposedShader::structure_hash` already identifies, and precisely what does not change while a
|
||||
slider is being dragged. `AdjustPass` relies on that already: it caches compiled pipelines against
|
||||
that hash and does not recompile during a drag. Caching the source string against the same hash and
|
||||
rebuilding only the uniforms — a handful of floats per operation — takes this to approximately
|
||||
nothing on the path that needs it most.
|
||||
|
||||
The care needed is in what the hash covers. It deliberately excludes parameter *magnitudes*, so a
|
||||
cache keyed on it is sound for the source and would be wrong for anything else in `ComposedShader`.
|
||||
|
||||
**Done when:** the `shader` column of [frame-budget.md](frame-budget.md)'s M1 table is under a
|
||||
millisecond for the `point` and `all` chains, and the codegen tests still pass byte for byte.
|
||||
|
||||
---
|
||||
|
||||
## Related, and deliberately not here
|
||||
|
||||
The window-move rule, the grid's ordering index and the whole-library readout cache were *fixed*
|
||||
|
||||
@@ -9,17 +9,17 @@ Denominators are parsed from [`requirements.md`](requirements.md) at run time, n
|
||||
|
||||
| Metric | Value |
|
||||
|---|---|
|
||||
| Source files scanned | 265 |
|
||||
| TRACES tags found | 765 |
|
||||
| Source files scanned | 268 |
|
||||
| TRACES tags found | 770 |
|
||||
| Requirements defined | 177 |
|
||||
| Requirements covered | 102 |
|
||||
| **Coverage** | **57.6%** (102/177) |
|
||||
| Requirements covered | 104 |
|
||||
| **Coverage** | **58.8%** (104/177) |
|
||||
|
||||
### By type
|
||||
|
||||
| Type | Covered | Defined |
|
||||
|---|---|---|
|
||||
| FR | 80 | 122 |
|
||||
| FR | 82 | 122 |
|
||||
| NFR | 20 | 49 |
|
||||
| R | 2 | 6 |
|
||||
|
||||
@@ -71,6 +71,8 @@ _None._
|
||||
| FR-DEV-7 | [`core/dr-pipeline/src/history.rs:214`](../core/dr-pipeline/src/history.rs#L214), [`core/dr-pipeline/src/history.rs:499`](../core/dr-pipeline/src/history.rs#L499), [`core/dr-pipeline/src/history.rs:526`](../core/dr-pipeline/src/history.rs#L526), [`ui/dr-ui/src/develop.rs:3468`](../ui/dr-ui/src/develop.rs#L3468), [`ui/dr-ui/src/develop.rs:3500`](../ui/dr-ui/src/develop.rs#L3500), [`ui/dr-ui/src/lib.rs:1437`](../ui/dr-ui/src/lib.rs#L1437), [`ui/dr-ui/src/lib.rs:2323`](../ui/dr-ui/src/lib.rs#L2323), [`ui/dr-ui/ui/history.slint:1`](../ui/dr-ui/ui/history.slint#L1) |
|
||||
| FR-DEV-8 | [`core/dr-gpu/src/detail.rs:252`](../core/dr-gpu/src/detail.rs#L252), [`core/dr-gpu/src/detail.rs:434`](../core/dr-gpu/src/detail.rs#L434), [`core/dr-gpu/tests/detail_instances.rs:1`](../core/dr-gpu/tests/detail_instances.rs#L1), [`core/dr-gpu/tests/spot_removal.rs:1`](../core/dr-gpu/tests/spot_removal.rs#L1), [`core/dr-pipeline/src/detail.rs:363`](../core/dr-pipeline/src/detail.rs#L363), [`core/dr-pipeline/src/detail.rs:387`](../core/dr-pipeline/src/detail.rs#L387), [`core/dr-pipeline/src/detail.rs:422`](../core/dr-pipeline/src/detail.rs#L422), [`core/dr-pipeline/src/detail.rs:496`](../core/dr-pipeline/src/detail.rs#L496), [`core/dr-pipeline/src/graph.rs:113`](../core/dr-pipeline/src/graph.rs#L113), [`core/dr-pipeline/src/graph.rs:191`](../core/dr-pipeline/src/graph.rs#L191), [`core/dr-pipeline/src/graph.rs:681`](../core/dr-pipeline/src/graph.rs#L681), [`core/dr-pipeline/src/operation.rs:299`](../core/dr-pipeline/src/operation.rs#L299), [`core/dr-pipeline/src/operation.rs:517`](../core/dr-pipeline/src/operation.rs#L517), [`core/dr-pipeline/src/sidecar.rs:183`](../core/dr-pipeline/src/sidecar.rs#L183), [`core/dr-pipeline/src/sidecar.rs:352`](../core/dr-pipeline/src/sidecar.rs#L352), [`core/dr-pipeline/src/sidecar.rs:672`](../core/dr-pipeline/src/sidecar.rs#L672), [`core/dr-pipeline/src/sidecar.rs:808`](../core/dr-pipeline/src/sidecar.rs#L808), [`core/dr-pipeline/src/sidecar.rs:862`](../core/dr-pipeline/src/sidecar.rs#L862), [`core/dr-pipeline/src/sidecar.rs:892`](../core/dr-pipeline/src/sidecar.rs#L892), [`core/dr-pipeline/src/spot.rs:115`](../core/dr-pipeline/src/spot.rs#L115), [`core/dr-pipeline/src/spot.rs:151`](../core/dr-pipeline/src/spot.rs#L151), [`core/dr-pipeline/src/spot.rs:1`](../core/dr-pipeline/src/spot.rs#L1), [`core/dr-pipeline/src/spot.rs:207`](../core/dr-pipeline/src/spot.rs#L207), [`core/dr-pipeline/src/spot.rs:387`](../core/dr-pipeline/src/spot.rs#L387), [`core/dr-pipeline/src/spot.rs:472`](../core/dr-pipeline/src/spot.rs#L472), [`core/dr-pipeline/src/spot.rs:582`](../core/dr-pipeline/src/spot.rs#L582), [`core/dr-pipeline/src/spot.rs:673`](../core/dr-pipeline/src/spot.rs#L673), [`core/dr-pipeline/src/state.rs:103`](../core/dr-pipeline/src/state.rs#L103), [`core/dr-pipeline/tests/spot_sidecar.rs:1`](../core/dr-pipeline/tests/spot_sidecar.rs#L1), [`core/dr-pipeline/tests/spots.rs:1`](../core/dr-pipeline/tests/spots.rs#L1), [`ui/dr-ui/src/develop.rs:2144`](../ui/dr-ui/src/develop.rs#L2144), [`ui/dr-ui/src/develop.rs:2182`](../ui/dr-ui/src/develop.rs#L2182), [`ui/dr-ui/src/develop.rs:2247`](../ui/dr-ui/src/develop.rs#L2247), [`ui/dr-ui/src/develop.rs:2330`](../ui/dr-ui/src/develop.rs#L2330), [`ui/dr-ui/src/develop.rs:2344`](../ui/dr-ui/src/develop.rs#L2344), [`ui/dr-ui/src/develop.rs:658`](../ui/dr-ui/src/develop.rs#L658), [`ui/dr-ui/src/labels.rs:52`](../ui/dr-ui/src/labels.rs#L52), [`ui/dr-ui/src/lib.rs:1458`](../ui/dr-ui/src/lib.rs#L1458), [`ui/dr-ui/src/lib.rs:2420`](../ui/dr-ui/src/lib.rs#L2420), [`ui/dr-ui/src/lib.rs:307`](../ui/dr-ui/src/lib.rs#L307), [`ui/dr-ui/src/spots_ui.rs:19`](../ui/dr-ui/src/spots_ui.rs#L19), [`ui/dr-ui/src/spots_ui.rs:1`](../ui/dr-ui/src/spots_ui.rs#L1), [`ui/dr-ui/src/spots_ui.rs:265`](../ui/dr-ui/src/spots_ui.rs#L265), [`ui/dr-ui/ui/adjust.slint:676`](../ui/dr-ui/ui/adjust.slint#L676), [`ui/dr-ui/ui/app.slint:1853`](../ui/dr-ui/ui/app.slint#L1853), [`ui/dr-ui/ui/app.slint:2370`](../ui/dr-ui/ui/app.slint#L2370), [`ui/dr-ui/ui/app.slint:2636`](../ui/dr-ui/ui/app.slint#L2636), [`ui/dr-ui/ui/app.slint:339`](../ui/dr-ui/ui/app.slint#L339), [`ui/dr-ui/ui/spots.slint:48`](../ui/dr-ui/ui/spots.slint#L48), [`ui/dr-ui/ui/spots.slint:5`](../ui/dr-ui/ui/spots.slint#L5) |
|
||||
| FR-DSP-1 | [`core/dr-gpu/src/adjust.rs:2088`](../core/dr-gpu/src/adjust.rs#L2088), [`core/dr-gpu/src/adjust.rs:2165`](../core/dr-gpu/src/adjust.rs#L2165), [`core/dr-gpu/src/adjust.rs:2250`](../core/dr-gpu/src/adjust.rs#L2250), [`core/dr-gpu/src/adjust.rs:54`](../core/dr-gpu/src/adjust.rs#L54), [`core/dr-gpu/src/adjust.rs:770`](../core/dr-gpu/src/adjust.rs#L770), [`core/dr-gpu/src/lib.rs:54`](../core/dr-gpu/src/lib.rs#L54), [`core/dr-gpu/src/lib.rs:94`](../core/dr-gpu/src/lib.rs#L94), [`core/dr-gpu/tests/capture_sharpen.rs:200`](../core/dr-gpu/tests/capture_sharpen.rs#L200), [`core/dr-gpu/tests/detail_stage.rs:328`](../core/dr-gpu/tests/detail_stage.rs#L328), [`core/dr-gpu/tests/local_contrast.rs:263`](../core/dr-gpu/tests/local_contrast.rs#L263), [`core/dr-gpu/tests/noise_reduction.rs:378`](../core/dr-gpu/tests/noise_reduction.rs#L378), [`core/dr-pipeline/src/detail.rs:136`](../core/dr-pipeline/src/detail.rs#L136), [`core/dr-pipeline/src/detail.rs:465`](../core/dr-pipeline/src/detail.rs#L465), [`core/dr-pipeline/src/graph.rs:543`](../core/dr-pipeline/src/graph.rs#L543), [`core/dr-pipeline/src/graph.rs:573`](../core/dr-pipeline/src/graph.rs#L573), [`core/dr-pipeline/src/ops/capture_sharpen.rs:1`](../core/dr-pipeline/src/ops/capture_sharpen.rs#L1), [`core/dr-pipeline/src/ops/capture_sharpen.rs:654`](../core/dr-pipeline/src/ops/capture_sharpen.rs#L654), [`core/dr-pipeline/src/ops/local_contrast.rs:1`](../core/dr-pipeline/src/ops/local_contrast.rs#L1), [`core/dr-pipeline/src/ops/local_contrast.rs:662`](../core/dr-pipeline/src/ops/local_contrast.rs#L662), [`core/dr-pipeline/src/ops/noise_reduction.rs:695`](../core/dr-pipeline/src/ops/noise_reduction.rs#L695), [`core/dr-pipeline/src/spot.rs:673`](../core/dr-pipeline/src/spot.rs#L673), [`ui/dr-ui/src/develop.rs:2668`](../ui/dr-ui/src/develop.rs#L2668), [`ui/dr-ui/src/develop.rs:3698`](../ui/dr-ui/src/develop.rs#L3698), [`ui/dr-ui/src/develop.rs:4181`](../ui/dr-ui/src/develop.rs#L4181), [`ui/dr-ui/src/develop.rs:4215`](../ui/dr-ui/src/develop.rs#L4215), [`ui/dr-ui/src/lib.rs:71`](../ui/dr-ui/src/lib.rs#L71), [`ui/dr-ui/src/lib.rs:737`](../ui/dr-ui/src/lib.rs#L737), [`ui/dr-ui/src/lib.rs:796`](../ui/dr-ui/src/lib.rs#L796) |
|
||||
| FR-DSP-3 | [`core/dr-gpu/tests/frame_budget.rs:101`](../core/dr-gpu/tests/frame_budget.rs#L101) |
|
||||
| FR-DSP-5 | [`core/dr-gpu/tests/frame_budget.rs:101`](../core/dr-gpu/tests/frame_budget.rs#L101), [`core/dr-gpu/tests/zoom_resolution.rs:131`](../core/dr-gpu/tests/zoom_resolution.rs#L131), [`core/dr-gpu/tests/zoom_resolution.rs:162`](../core/dr-gpu/tests/zoom_resolution.rs#L162), [`core/dr-gpu/tests/zoom_resolution.rs:1`](../core/dr-gpu/tests/zoom_resolution.rs#L1), [`core/dr-gpu/tests/zoom_resolution.rs:212`](../core/dr-gpu/tests/zoom_resolution.rs#L212) |
|
||||
| FR-DSP-6 | [`core/dr-pipeline/src/operation.rs:446`](../core/dr-pipeline/src/operation.rs#L446), [`core/dr-types/src/colour.rs:1`](../core/dr-types/src/colour.rs#L1), [`ui/dr-ui/src/develop.rs:2698`](../ui/dr-ui/src/develop.rs#L2698), [`ui/dr-ui/src/develop.rs:5473`](../ui/dr-ui/src/develop.rs#L5473), [`ui/dr-ui/src/lib.rs:1414`](../ui/dr-ui/src/lib.rs#L1414), [`ui/dr-ui/src/lib.rs:2626`](../ui/dr-ui/src/lib.rs#L2626) |
|
||||
| FR-DSP-7 | [`core/dr-gpu/src/histogram.rs:147`](../core/dr-gpu/src/histogram.rs#L147), [`core/dr-gpu/src/histogram.rs:1`](../core/dr-gpu/src/histogram.rs#L1), [`core/dr-gpu/src/histogram.rs:281`](../core/dr-gpu/src/histogram.rs#L281), [`core/dr-gpu/src/histogram.rs:50`](../core/dr-gpu/src/histogram.rs#L50), [`core/dr-gpu/src/shaders/histogram.wgsl:1`](../core/dr-gpu/src/shaders/histogram.wgsl#L1), [`ui/dr-ui/src/develop.rs:2747`](../ui/dr-ui/src/develop.rs#L2747), [`ui/dr-ui/src/develop.rs:5283`](../ui/dr-ui/src/develop.rs#L5283), [`ui/dr-ui/src/develop.rs:5315`](../ui/dr-ui/src/develop.rs#L5315), [`ui/dr-ui/src/develop.rs:621`](../ui/dr-ui/src/develop.rs#L621), [`ui/dr-ui/src/histogram.rs:1`](../ui/dr-ui/src/histogram.rs#L1), [`ui/dr-ui/src/lib.rs:1514`](../ui/dr-ui/src/lib.rs#L1514), [`ui/dr-ui/src/lib.rs:297`](../ui/dr-ui/src/lib.rs#L297), [`ui/dr-ui/ui/app.slint:299`](../ui/dr-ui/ui/app.slint#L299), [`ui/dr-ui/ui/histogram.slint:122`](../ui/dr-ui/ui/histogram.slint#L122), [`ui/dr-ui/ui/histogram.slint:1`](../ui/dr-ui/ui/histogram.slint#L1) |
|
||||
| FR-DSP-8 | [`platform/dr-plat/src/display.rs:1`](../platform/dr-plat/src/display.rs#L1), [`platform/dr-plat/src/display/icc.rs:1`](../platform/dr-plat/src/display/icc.rs#L1), [`platform/dr-plat/src/display/wayland.rs:1`](../platform/dr-plat/src/display/wayland.rs#L1), [`platform/dr-plat/src/display/x11.rs:1`](../platform/dr-plat/src/display/x11.rs#L1), [`ui/dr-ui/src/develop.rs:2644`](../ui/dr-ui/src/develop.rs#L2644), [`ui/dr-ui/src/develop.rs:2698`](../ui/dr-ui/src/develop.rs#L2698), [`ui/dr-ui/src/develop.rs:5473`](../ui/dr-ui/src/develop.rs#L5473), [`ui/dr-ui/src/develop.rs:5519`](../ui/dr-ui/src/develop.rs#L5519), [`ui/dr-ui/src/develop.rs:5538`](../ui/dr-ui/src/develop.rs#L5538), [`ui/dr-ui/src/develop.rs:689`](../ui/dr-ui/src/develop.rs#L689), [`ui/dr-ui/src/display_ui.rs:192`](../ui/dr-ui/src/display_ui.rs#L192), [`ui/dr-ui/src/display_ui.rs:1`](../ui/dr-ui/src/display_ui.rs#L1), [`ui/dr-ui/src/display_ui.rs:325`](../ui/dr-ui/src/display_ui.rs#L325), [`ui/dr-ui/src/display_ui.rs:346`](../ui/dr-ui/src/display_ui.rs#L346), [`ui/dr-ui/src/display_ui.rs:379`](../ui/dr-ui/src/display_ui.rs#L379), [`ui/dr-ui/src/lib.rs:1388`](../ui/dr-ui/src/lib.rs#L1388), [`ui/dr-ui/src/lib.rs:1414`](../ui/dr-ui/src/lib.rs#L1414), [`ui/dr-ui/src/lib.rs:2603`](../ui/dr-ui/src/lib.rs#L2603), [`ui/dr-ui/src/lib.rs:2626`](../ui/dr-ui/src/lib.rs#L2626), [`ui/dr-ui/ui/app.slint:1713`](../ui/dr-ui/ui/app.slint#L1713), [`ui/dr-ui/ui/app.slint:278`](../ui/dr-ui/ui/app.slint#L278), [`ui/dr-ui/ui/settings.slint:120`](../ui/dr-ui/ui/settings.slint#L120), [`ui/dr-ui/ui/settings.slint:730`](../ui/dr-ui/ui/settings.slint#L730) |
|
||||
@@ -138,7 +140,7 @@ _None._
|
||||
|
||||
## Not yet tagged
|
||||
|
||||
75 of 177 requirements have no implementation tag. Expected while the codebase is young; each should gain one as it is built.
|
||||
73 of 177 requirements have no implementation tag. Expected while the codebase is young; each should gain one as it is built.
|
||||
|
||||
<details><summary>Show untagged requirements</summary>
|
||||
|
||||
@@ -150,9 +152,7 @@ _None._
|
||||
- FR-DEV-1
|
||||
- FR-DEV-3g
|
||||
- FR-DSP-2
|
||||
- FR-DSP-3
|
||||
- FR-DSP-4
|
||||
- FR-DSP-5
|
||||
- FR-NC-11
|
||||
- FR-PLAT-AND-2
|
||||
- FR-PLAT-AND-4
|
||||
|
||||
Reference in New Issue
Block a user