TD-4 asked for the measurement as well as the change, and this is it: 25.05 ms to 4.17 ms at 3840 x 2160, six times faster, with clarity no longer dominating the neighbourhood stage it used to be 97% of. Measured before and after on the same machine and the same adapter minutes apart, baseline at the branch's merge-base, so the only variable is the change. That adapter is not the RTX 3050 the rest of this document was measured on, so the new table says to read it on its own rather than against the ones above — the before/after is comparable, the absolute figures are not, and quietly replacing the existing tables would have changed the instrument. Also recorded: the declared halo is now quantised to multiples of the output scale, because the 2-sigma truncation rounds on the reduced grid. 29 px becomes 28 at 1920x1200 and 38 becomes 40 at 2560x1600. It is inside what the cross-form test holds — 0.03 stops of peak, 2% of reach — but it is a change in reach and not only in cost, and a tile scheduler would be handed it. Better written down now than found later as a seam. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
19 KiB
What a frame costs
Status: Measured · 2026-08-27
Companion to: display-and-extension.md §2–3 ·
requirements.md §3.4 FR-DSP-2, FR-DSP-3, FR-DSP-4
Instrument: core/dr-gpu/examples/frame_budget.rs
Guard: core/dr-gpu/tests/frame_budget.rs
display-and-extension.md §2 fixed a decision rule in advance and made three measurements the thing that settles it. This file is those measurements, and the recommendation they support.
Rerun with:
cargo run --release -p dr-gpu --example frame_budget
and diff this file. That is the whole point of committing numbers: a regression should be a diff rather than somebody's recollection of how fast it used to be.
The answer, first
FR-DSP-2 should be rewritten, not implemented. M1 and M2 sit inside the 16 ms budget at the 99th percentile for every chain of point operations at every viewport size measured, fit and at 1:1 — the widest case, every operation that contributes a fragment to the fused shader at 4K, costs 4.5 ms on the GPU and 8.2 ms including the composition that precedes it. Tiling the interactive path would be optimising something that is already using a quarter of its budget.
But the measurement did find a budget-breaker, and it is not the one tiling
fixes. The neighbourhood stage — clarity in particular — costs 34 ms at 4K
on its own, twice the whole budget, and tiles do not help it: a tile of a
convolution has to read its halo, so tiling raises the total tap count rather
than lowering it. §2 predicted this exactly ("a separable blur at a large radius
is the plausible budget-breaker, not the fused pass"), and the fix it needs is
the one local_contrast's own module documentation already names — a base
computed at reduced resolution — which is a change to crate::detail, not a
tile scheduler.
There is a third finding nobody was looking for: shader composition costs 3–5 ms of CPU per frame on a full chain, on the UI thread, before any GPU work is submitted. That is a fifth to a third of the budget spent formatting strings, and it is invisible to any amount of tiling.
Conditions
| Adapter | NVIDIA GeForce RTX 3050 6GB Laptop GPU (Vulkan) |
| Source | 9504 × 6336 synthetic (60.2 MP, 482 MB as rgba16f) |
| Frames | 100 measured per row, 12 warm-up frames discarded |
| Percentile | Nearest-rank, so p99 of 100 frames is the second-worst frame |
| Build | --release |
| Date | 2026-08-27 |
shader is EditGraph::compose alone. cpu adds the detail chain and the
invalidation hash — everything DevelopSession::render does per frame before it
dispatches. gpu is submit plus wait-for-idle, which serialises the GPU work
into the frame that caused it and is therefore pessimistic. TOTAL ranks
cpu + gpu summed within each frame, which is the column the budget is
judged on; adding two percentiles instead would invent a stutter that no frame
actually had.
Chains: one is exposure. five is exposure, contrast, highlights/shadows,
blacks/whites, vibrance. point is every operation in the default chain that
contributes a fragment to the fused shader, film stock included. all is point
plus the four neighbourhood operations — noise reduction, capture sharpening,
clarity and texture.
M1 — the fused pass at proxy resolution
The develop view: the whole frame fit to the viewport.
| size | chain | shader | cpu p99 | gpu p50 | gpu p99 | TOTAL | |
|---|---|---|---|---|---|---|---|
| 1920 × 1200 | one | 0.08ms | 0.10ms | 1.02ms | 1.23ms | 1.31ms | |
| 1920 × 1200 | five | 0.18ms | 0.20ms | 1.01ms | 1.20ms | 1.36ms | |
| 1920 × 1200 | point | 2.79ms | 2.82ms | 1.98ms | 2.18ms | 4.83ms | |
| 1920 × 1200 | all | 3.65ms | 4.73ms | 6.86ms | 7.37ms | 12.02ms | |
| 2560 × 1600 | one | 0.08ms | 0.10ms | 1.73ms | 2.00ms | 2.12ms | |
| 2560 × 1600 | five | 0.24ms | 0.27ms | 1.73ms | 2.26ms | 2.46ms | |
| 2560 × 1600 | point | 2.82ms | 2.85ms | 2.37ms | 2.65ms | 5.38ms | |
| 2560 × 1600 | all | 3.37ms | 4.35ms | 14.31ms | 15.65ms | 18.42ms | OVER |
| 3840 × 2160 | one | 0.10ms | 0.14ms | 3.09ms | 3.31ms | 3.42ms | |
| 3840 × 2160 | five | 0.30ms | 0.32ms | 3.03ms | 3.40ms | 3.61ms | |
| 3840 × 2160 | point | 3.62ms | 3.65ms | 4.12ms | 4.52ms | 8.23ms | |
| 3840 × 2160 | all | 4.12ms | 5.07ms | 37.73ms | 40.17ms | 43.24ms | OVER |
Read the point rows: the fused dispatch scales with pixels and almost not at
all with chain length. Going from one operation to the entire point chain at
4K costs 1.2 ms of GPU. Going from 2.3 M pixels to 8.3 M costs 2.3 ms. Both are
small, and the second is the one tiling would address.
The all rows go over, and the point rows in the same block are what say why:
the difference between them is the neighbourhood stage, measured on its own in
M3 and arriving at almost exactly the same figure.
M2 — the same, zoomed to 1:1 on the 60 MP source
FR-DSP-5's case. Framing::view shrinks the sampled region while the render
target keeps its size, so one render pixel lands on one source pixel.
| size | chain | shader | cpu p99 | gpu p50 | gpu p99 | TOTAL | |
|---|---|---|---|---|---|---|---|
| 1920 × 1200 | one | 0.12ms | 0.14ms | 0.42ms | 0.66ms | 0.75ms | |
| 1920 × 1200 | five | 0.27ms | 0.30ms | 0.49ms | 1.14ms | 1.22ms | |
| 1920 × 1200 | point | 3.49ms | 3.52ms | 1.18ms | 1.39ms | 4.85ms | |
| 1920 × 1200 | all | 3.64ms | 5.18ms | 8.96ms | 9.55ms | 14.30ms | |
| 2560 × 1600 | one | 0.11ms | 0.12ms | 0.56ms | 0.99ms | 1.06ms | |
| 2560 × 1600 | five | 0.22ms | 0.25ms | 0.77ms | 1.02ms | 1.17ms | |
| 2560 × 1600 | point | 2.96ms | 2.99ms | 2.03ms | 2.52ms | 5.61ms | |
| 2560 × 1600 | all | 5.24ms | 7.35ms | 18.80ms | 21.62ms | 25.81ms | OVER |
| 3840 × 2160 | one | 0.10ms | 0.12ms | 1.31ms | 1.52ms | 1.63ms | |
| 3840 × 2160 | five | 0.14ms | 0.27ms | 1.39ms | 1.64ms | 1.75ms | |
| 3840 × 2160 | point | 3.14ms | 3.16ms | 4.04ms | 4.50ms | 7.21ms | |
| 3840 × 2160 | all | 4.84ms | 6.78ms | 47.22ms | 48.79ms | 54.47ms | OVER |
A 1:1 view of a 60 MP file is cheaper than the fit view of the same file, for every point chain and at every size — 1.52 ms against 3.31 ms for one operation at 4K. That is not a rounding artefact and it is worth stating plainly, because it is the opposite of what "full resolution" sounds like it should cost. The dispatch is the same number of pixels either way; what changes is where those pixels read from. A fit view walks the whole 482 MB texture on a stride, and a 1:1 view reads a contiguous window of it that fits comfortably in cache.
So the resolution FR-DSP-5 promises costs nothing extra on the fused path. Zooming is not an expensive mode to be dreaded and progressively refined into; it is the cheap one.
The all rows are worse at 1:1 than fit, and that is the detail stage again for
a specific reason: noise reduction's radius is stated in source pixels, so
RenderScale::ratio climbing to 1.0 widens its kernel. Clarity's is stated as a
fraction of the frame and does not move. M3 separates the two.
M3 — the neighbourhood stage alone
Timed with the fused dispatch deliberately reused: only a detail parameter moves,
so render_detailed skips the colour pass (FR-DEV-3d) and what remains is the
convolutions. colour counts fused dispatches over the measured frames and is
zero on every row, which is what makes these numbers mean "detail alone" rather
than asserting it.
| size | stage | view | pass | radius | colour | cpu p99 | p50 | p99 |
|---|---|---|---|---|---|---|---|---|
| 1920 × 1200 | clarity | fit | 2 | 29 | 0 | 1.60ms | 5.42ms | 5.99ms |
| 1920 × 1200 | all four | fit | 7 | 29 | 0 | 1.71ms | 5.78ms | 6.16ms |
| 1920 × 1200 | clarity | 1:1 | 2 | 29 | 0 | 1.12ms | 7.51ms | 8.01ms |
| 1920 × 1200 | all four | 1:1 | 9 | 29 | 0 | 2.71ms | 8.25ms | 9.11ms |
| 2560 × 1600 | clarity | fit | 2 | 38 | 0 | 1.03ms | 12.02ms | 12.44ms |
| 2560 × 1600 | all four | fit | 7 | 38 | 0 | 1.87ms | 12.49ms | 13.16ms |
| 2560 × 1600 | clarity | 1:1 | 2 | 38 | 0 | 1.87ms | 15.82ms | 16.60ms |
| 2560 × 1600 | all four | 1:1 | 9 | 38 | 0 | 2.71ms | 17.24ms | 18.06ms |
| 3840 × 2160 | clarity | fit | 2 | 52 | 0 | 1.75ms | 33.11ms | 33.89ms |
| 3840 × 2160 | all four | fit | 7 | 52 | 0 | 1.76ms | 34.21ms | 35.03ms |
| 3840 × 2160 | clarity | 1:1 | 2 | 52 | 0 | 1.08ms | 40.39ms | 41.86ms |
| 3840 × 2160 | all four | 1:1 | 9 | 52 | 0 | 2.37ms | 43.29ms | 44.72ms |
radius is the widest halo any pass reads, in render pixels.
Clarity alone is 97% of the cost of all four neighbourhood operations together,
at every size. Its σ is 1.2% of the shorter edge and it truncates at 2σ, so its
radius is 29 px on a 1200 px viewport and 52 px at 4K — two separable passes
of 105 taps each, over 8.3 M pixels, which is 1.7 billion texture reads. That is
the whole of the problem, and the numbers scale as radius × pixels exactly as
that description predicts: 5.99 → 12.44 → 33.89 ms for radii of 29 → 38 → 52 over
2.3 → 4.1 → 8.3 M pixels.
The extra cost at 1:1 is noise reduction and capture sharpening, whose radii are
properties of the sensor rather than of the frame. That is the correct behaviour
— it is why RenderScale has two units — and it is bounded by the kernel caps
those operations already declare.
Reading this against §2's decision rule
§2: "If M1 and M2 sit inside 16 ms at the 99th percentile, FR-DSP-2 is rewritten rather than implemented … If they do not, the measurement tells us which stage to tile."
Both halves of the rule fire, on different stages, and the honest reading takes both.
FR-DSP-2 — rewrite it
For the fused pass the rule passes with a wide margin. Every point chain at
every size, fit and at 1:1, is inside 16 ms — the worst TOTAL is 8.23 ms and
the worst GPU figure is 4.52 ms. There is no viewport size on a desktop display
where recomputing the entire point chain over every visible pixel is a problem.
Two further reasons not to build the tile scheduler as written:
-
Panning, which is the case ARCH §5.3's tile cache is designed for, gets no benefit here. Reusing already-valid tiles saves recomputation. Recomputing the whole 4K viewport costs 4.5 ms, so a perfect tile cache could save at most 4.5 ms of a 16 ms budget, at the price of a cache keyed by
(VersionId, tile, zoom, graph_hash_prefix)that has to stay correct across every parameter change in the graph. That is a large correctness surface bought with a small number. -
It would make the actual problem worse. The stage that misses the budget is a convolution, and a tiled convolution reads a halo per tile. At a 52-pixel radius, 256-pixel tiles would read (256+104)² instead of 256² — very nearly twice the taps. Tiling is the wrong tool for the one stage that needs a tool.
So FR-DSP-2 becomes what §2 said it actually is for this architecture: a scheduling concern for export and thumbnailing, both of which already run off the frame path. The interactive path does not tile.
The stage that did need work — and it is not tiling
Resolved. The fix described below landed; the measurement is in
§The reduced base, measured at the foot of this
file, and docs/technical-debt.md TD-4 is closed. What follows is the
reasoning as it stood, kept because it is what the numbers above argue for and
because the tiling half of it is still live.
The measurement's real product is naming the stage. It is local_contrast, and
the fix is stated in that module's own documentation:
The right optimisation is a base computed at reduced resolution, which needs a detail stage that can write a smaller target than it reads; that is a change to
crate::detail, not to this file.
A Gaussian base at a quarter resolution is 1/16 the pixels at 1/4 the radius — about 1/64 of the work — and the result is visually identical because a base at σ = 26 px has no content above the quarter-resolution Nyquist to lose. That is a change to two files with a bounded blast radius, and it is what the 34 ms buys back. It should be tracked as its own item rather than smuggled in under a requirement about tiles.
FR-DSP-3 — the clause that should be narrowed
§3.3 proposes narrowing "when a full-resolution result is needed it is computed asynchronously, and the proxy result remains on screen until it is ready" to export and 1:1 zoom, or striking it.
M2 says strike it. The clause exists to hide the latency of a
full-resolution render behind a proxy. There is no such latency: the 1:1 view is
faster than the fit view on the fused path, and there is no second
full-resolution code path to be asynchronous about — Framing::view is the
whole mechanism. Export renders its own frames on a worker already. Keeping the
clause would mean building a progressive-swap machine to conceal a render that
completes in 1.4 ms.
FR-DSP-4 — satisfied vacuously, on the fused path
§4 makes progressive refinement conditional on M1 failing. On the fused path M1 passes, so reduced-quality rendering during a drag would buy nothing and cost the visible softness the requirement itself warns against.
The neighbourhood stage is the exception, and it is worth being precise: what that stage needs is not progressive refinement — it is a permanently cheaper base, computed at reduced resolution and correct at any moment the user stops. "Render coarse while dragging, sharpen when it settles" would paper over the same 34 ms with a visible swap. Fix the stage.
What is not measured here
Stated because §7 of display-and-extension.md asks for it, and because each of these could move the numbers.
- Local adjustments. The mask stack is a separate chain per layer and is not
in any row above.
render_maskedtakes them and the fused shader addresses them per layer, so a heavily masked edit costs more thanall. - Spot repairs. These add detail passes, and their cost is per spot.
- Lens corrections. Not part of
EditGraph::default_chain— they are built from a matched profile — so thepointrow does not include the warp chain. - Demosaic. Once per photograph on a worker, not on the frame path.
- Presentation. The bench waits for the device to go idle inside the frame it measures. A real compositor overlaps frames, so these figures are an upper bound rather than an estimate.
- One adapter. A discrete laptop GPU. The Intel iGPU on the same machine, and Android, will be slower — which is an argument for the conclusion rather than against it: the stage with no headroom has none to lose.
The CPU finding, which deserves its own item
EditGraph::compose costs 2.8–5.2 ms per frame on a full chain, at every
resolution, because it is resolution-independent: it assembles a WGSL string and
hashes it. On the all rows it is a third of what is left of the budget after
the GPU has taken its share, and at 1920 × 1200 it is larger than the entire
fused dispatch.
Nothing in this document's recommendations changes it, and it is the cheapest
remaining win. The generated source depends only on the structure of the graph
— that is what structure_hash already identifies, and it is precisely what does
not change while a slider is being dragged, which is why the pipeline cache in
AdjustPass does not recompile. The uniforms do change, but assembling them is a
handful of floats per operation. So caching the source string against the
structure hash and rebuilding only the uniforms would take these milliseconds to
approximately nothing, on the path that needs them most. Worth its own entry in
technical-debt.md.
The reduced base, measured
Status: Measured · 2026-08-29 · closes TD-4
DetailPass gained an output_scale, and clarity's base is now computed on a
grid a quarter the size on each axis — the change §M3 argued for above.
Read this table on its own, not against the ones above. It was taken on a
different adapter, so the absolute figures are not comparable with the RTX 3050
measurements this document is otherwise built from. What is comparable is the
before and the after, which were measured on the same machine, same card, same
release profile, minutes apart, with nothing between them but the change — the
baseline at 0407fb8 and the result at bff95e2.
| Adapter | AMD Radeon RX 5700 XT (RADV NAVI10) (Vulkan) |
| Source | 9504 × 6336 (60.2 MP, 482 MB as rgba16f) |
| Baseline | 0407fb8, the branch's merge-base |
| Result | bff95e2 |
| Date | 2026-08-29 |
M3 — the neighbourhood stage alone, before and after
| size | stage | view | before p99 | after p99 | |
|---|---|---|---|---|---|
| 1920 × 1200 | clarity | fit | 3.94ms | 1.97ms | 2.0× |
| 1920 × 1200 | all four | fit | 2.57ms | 1.81ms | 1.4× |
| 1920 × 1200 | clarity | 1:1 | 3.07ms | 2.41ms | 1.3× |
| 2560 × 1600 | clarity | fit | 10.40ms | 2.37ms | 4.4× |
| 2560 × 1600 | all four | fit | 5.50ms | 2.64ms | 2.1× |
| 2560 × 1600 | clarity | 1:1 | 5.78ms | 3.85ms | 1.5× |
| 3840 × 2160 | clarity | fit | 25.05ms | 4.17ms | 6.0× |
| 3840 × 2160 | all four | fit | 25.48ms | 4.61ms | 5.5× |
| 3840 × 2160 | clarity | 1:1 | 27.60ms | 6.93ms | 4.0× |
| 3840 × 2160 | all four | 1:1 | 15.88ms | 7.74ms | 2.1× |
The passes went from 2 to 4 and got faster, which is the point: three of the four now run on a grid a sixteenth the area, and the fourth — the combine — is a single bilinear read where it used to be a 105-tap convolution.
Clarity is no longer the stage that dominates. At 4K fit it is 4.17 ms
against 4.61 ms for all four neighbourhood operations together; it was 97% of
that total at every size before. Every M1 and M2 row now sits inside the 16 ms
budget on this adapter, including the six all rows that carried OVER.
The one thing that is not a pure speed-up
The declared halo is now quantised to multiples of output_scale. The
kernel truncates at 2σ, and that rounding now happens on the reduced grid
before being multiplied back up:
| viewport | before | after |
|---|---|---|
| 1920 × 1200 | 29 px | 28 px |
| 2560 × 1600 | 38 px | 40 px |
| 3840 × 2160 | 52 px | 52 px |
At 2σ the Gaussian is already down to e⁻² of its peak, and
crossing_the_reduction_threshold_does_not_change_the_picture holds the
difference between a quarter-scale and a half-scale base to 0.03 stops of peak
excursion and 2% of frame reach. But it is a change in reach rather than only
in cost, it is what a tile scheduler would be handed, and it is worth knowing
that the number moved rather than discovering it later as a seam.