Files
DarkRoom/docs/dev/frame-budget.md
dtourolle affdaecaee Stop describing a base curve the pipeline no longer has
D19 retired the per-body base curve, moved the matrix ahead of the
edits and the film into the view transform's place, but a dozen doc
comments still listed the curve among what a pixel passes through, or
said the film skipped it. The detail stage's module doc still drew the
matrix after the edits and the last detail pass encoding, which the
view pass took over. The film crate's README gave the base curves as
its reason for being data, and the ops README's list of hand-written
nodes had neither the view transform nor three of the five kernels.

FR-MRG-2 gave the base curve as why the merge cuts below the profile;
the view transform is why now. The decision table still said colour
defaults were a per-body curve, and FR-DEV-3j said only the default
view transform skips a JPEG, where the node skips one whatever its
sliders say. frame-budget.md records the view pass as unmeasured.
2026-09-27 19:42:41 -04:00

28 KiB
Raw Permalink Blame History

What a frame costs

Status: Measured · 2026-08-27 Companion to: display-and-extension.md §2–3 · requirements.md §3.4 FR-DSP-2, FR-DSP-3, FR-DSP-4 Instrument: core/dr-gpu/examples/frame_budget.rs Guard: core/dr-gpu/tests/frame_budget.rs

display-and-extension.md §2 fixed a decision rule in advance and made three measurements the thing that settles it. This file is those measurements, and the recommendation they support.

Rerun with:

cargo run --release -p dr-gpu --example frame_budget

and diff this file. That is the whole point of committing numbers: a regression should be a diff rather than somebody's recollection of how fast it used to be.


The answer, first

FR-DSP-2 should be rewritten, not implemented. M1 and M2 sit inside the 16 ms budget at the 99th percentile for every chain of point operations at every viewport size measured, fit and at 1:1 — the widest case, every operation that contributes a fragment to the fused shader at 4K, costs 4.5 ms on the GPU and 8.2 ms including the composition that precedes it. Tiling the interactive path would be optimising something that is already using a quarter of its budget.

But the measurement did find a budget-breaker, and it is not the one tiling fixes. The neighbourhood stage — clarity in particular — costs 34 ms at 4K on its own, twice the whole budget, and tiles do not help it: a tile of a convolution has to read its halo, so tiling raises the total tap count rather than lowering it. §2 predicted this exactly ("a separable blur at a large radius is the plausible budget-breaker, not the fused pass"), and the fix it needs is the one local_contrast's own module documentation already names — a base computed at reduced resolution — which is a change to crate::detail, not a tile scheduler.

There is a third finding nobody was looking for: shader composition costs 3–5 ms of CPU per frame on a full chain, on the UI thread, before any GPU work is submitted. That is a fifth to a third of the budget spent formatting strings, and it is invisible to any amount of tiling.


Conditions

Adapter NVIDIA GeForce RTX 3050 6GB Laptop GPU (Vulkan)
Source 9504 × 6336 synthetic (60.2 MP, 482 MB as rgba16f)
Frames 100 measured per row, 12 warm-up frames discarded
Percentile Nearest-rank, so p99 of 100 frames is the second-worst frame
Build --release
Date 2026-08-27

shader is EditGraph::compose alone. cpu adds the detail chain and the invalidation hash — everything DevelopSession::render does per frame before it dispatches. gpu is submit plus wait-for-idle, which serialises the GPU work into the frame that caused it and is therefore pessimistic. TOTAL ranks cpu + gpu summed within each frame, which is the column the budget is judged on; adding two percentiles instead would invent a stutter that no frame actually had.

Chains: one is exposure. five is exposure, contrast, highlights/shadows, blacks/whites, vibrance. point is every operation in the default chain that contributes a fragment to the fused shader, film stock included. all is point plus the four neighbourhood operations — noise reduction, capture sharpening, clarity and texture.


M1 — the fused pass at proxy resolution

The develop view: the whole frame fit to the viewport.

size chain shader cpu p99 gpu p50 gpu p99 TOTAL
1920 × 1200 one 0.08ms 0.10ms 1.02ms 1.23ms 1.31ms
1920 × 1200 five 0.18ms 0.20ms 1.01ms 1.20ms 1.36ms
1920 × 1200 point 2.79ms 2.82ms 1.98ms 2.18ms 4.83ms
1920 × 1200 all 3.65ms 4.73ms 6.86ms 7.37ms 12.02ms
2560 × 1600 one 0.08ms 0.10ms 1.73ms 2.00ms 2.12ms
2560 × 1600 five 0.24ms 0.27ms 1.73ms 2.26ms 2.46ms
2560 × 1600 point 2.82ms 2.85ms 2.37ms 2.65ms 5.38ms
2560 × 1600 all 3.37ms 4.35ms 14.31ms 15.65ms 18.42ms OVER
3840 × 2160 one 0.10ms 0.14ms 3.09ms 3.31ms 3.42ms
3840 × 2160 five 0.30ms 0.32ms 3.03ms 3.40ms 3.61ms
3840 × 2160 point 3.62ms 3.65ms 4.12ms 4.52ms 8.23ms
3840 × 2160 all 4.12ms 5.07ms 37.73ms 40.17ms 43.24ms OVER

Read the point rows: the fused dispatch scales with pixels and almost not at all with chain length. Going from one operation to the entire point chain at 4K costs 1.2 ms of GPU. Going from 2.3 M pixels to 8.3 M costs 2.3 ms. Both are small, and the second is the one tiling would address.

The all rows go over, and the point rows in the same block are what say why: the difference between them is the neighbourhood stage, measured on its own in M3 and arriving at almost exactly the same figure.

M2 — the same, zoomed to 1:1 on the 60 MP source

FR-DSP-5's case. Framing::view shrinks the sampled region while the render target keeps its size, so one render pixel lands on one source pixel.

size chain shader cpu p99 gpu p50 gpu p99 TOTAL
1920 × 1200 one 0.12ms 0.14ms 0.42ms 0.66ms 0.75ms
1920 × 1200 five 0.27ms 0.30ms 0.49ms 1.14ms 1.22ms
1920 × 1200 point 3.49ms 3.52ms 1.18ms 1.39ms 4.85ms
1920 × 1200 all 3.64ms 5.18ms 8.96ms 9.55ms 14.30ms
2560 × 1600 one 0.11ms 0.12ms 0.56ms 0.99ms 1.06ms
2560 × 1600 five 0.22ms 0.25ms 0.77ms 1.02ms 1.17ms
2560 × 1600 point 2.96ms 2.99ms 2.03ms 2.52ms 5.61ms
2560 × 1600 all 5.24ms 7.35ms 18.80ms 21.62ms 25.81ms OVER
3840 × 2160 one 0.10ms 0.12ms 1.31ms 1.52ms 1.63ms
3840 × 2160 five 0.14ms 0.27ms 1.39ms 1.64ms 1.75ms
3840 × 2160 point 3.14ms 3.16ms 4.04ms 4.50ms 7.21ms
3840 × 2160 all 4.84ms 6.78ms 47.22ms 48.79ms 54.47ms OVER

A 1:1 view of a 60 MP file is cheaper than the fit view of the same file, for every point chain and at every size — 1.52 ms against 3.31 ms for one operation at 4K. That is not a rounding artefact and it is worth stating plainly, because it is the opposite of what "full resolution" sounds like it should cost. The dispatch is the same number of pixels either way; what changes is where those pixels read from. A fit view walks the whole 482 MB texture on a stride, and a 1:1 view reads a contiguous window of it that fits comfortably in cache.

So the resolution FR-DSP-5 promises costs nothing extra on the fused path. Zooming is not an expensive mode to be dreaded and progressively refined into; it is the cheap one.

The all rows are worse at 1:1 than fit, and that is the detail stage again for a specific reason: noise reduction's radius is stated in source pixels, so RenderScale::ratio climbing to 1.0 widens its kernel. Clarity's is stated as a fraction of the frame and does not move. M3 separates the two.

M3 — the neighbourhood stage alone

Timed with the fused dispatch deliberately reused: only a detail parameter moves, so render_detailed skips the colour pass (FR-DEV-3d) and what remains is the convolutions. colour counts fused dispatches over the measured frames and is zero on every row, which is what makes these numbers mean "detail alone" rather than asserting it.

size stage view pass radius colour cpu p99 p50 p99
1920 × 1200 clarity fit 2 29 0 1.60ms 5.42ms 5.99ms
1920 × 1200 all four fit 7 29 0 1.71ms 5.78ms 6.16ms
1920 × 1200 clarity 1:1 2 29 0 1.12ms 7.51ms 8.01ms
1920 × 1200 all four 1:1 9 29 0 2.71ms 8.25ms 9.11ms
2560 × 1600 clarity fit 2 38 0 1.03ms 12.02ms 12.44ms
2560 × 1600 all four fit 7 38 0 1.87ms 12.49ms 13.16ms
2560 × 1600 clarity 1:1 2 38 0 1.87ms 15.82ms 16.60ms
2560 × 1600 all four 1:1 9 38 0 2.71ms 17.24ms 18.06ms
3840 × 2160 clarity fit 2 52 0 1.75ms 33.11ms 33.89ms
3840 × 2160 all four fit 7 52 0 1.76ms 34.21ms 35.03ms
3840 × 2160 clarity 1:1 2 52 0 1.08ms 40.39ms 41.86ms
3840 × 2160 all four 1:1 9 52 0 2.37ms 43.29ms 44.72ms

radius is the widest halo any pass reads, in render pixels.

Clarity alone is 97% of the cost of all four neighbourhood operations together, at every size. Its σ is 1.2% of the shorter edge and it truncates at 2σ, so its radius is 29 px on a 1200 px viewport and 52 px at 4K — two separable passes of 105 taps each, over 8.3 M pixels, which is 1.7 billion texture reads. That is the whole of the problem, and the numbers scale as radius × pixels exactly as that description predicts: 5.99 → 12.44 → 33.89 ms for radii of 29 → 38 → 52 over 2.3 → 4.1 → 8.3 M pixels.

The extra cost at 1:1 is noise reduction and capture sharpening, whose radii are properties of the sensor rather than of the frame. That is the correct behaviour — it is why RenderScale has two units — and it is bounded by the kernel caps those operations already declare.


Reading this against §2's decision rule

§2: "If M1 and M2 sit inside 16 ms at the 99th percentile, FR-DSP-2 is rewritten rather than implemented … If they do not, the measurement tells us which stage to tile."

Both halves of the rule fire, on different stages, and the honest reading takes both.

FR-DSP-2 — rewrite it

For the fused pass the rule passes with a wide margin. Every point chain at every size, fit and at 1:1, is inside 16 ms — the worst TOTAL is 8.23 ms and the worst GPU figure is 4.52 ms. There is no viewport size on a desktop display where recomputing the entire point chain over every visible pixel is a problem.

Two further reasons not to build the tile scheduler as written:

  1. Panning, which is the case ARCH §5.3's tile cache is designed for, gets no benefit here. Reusing already-valid tiles saves recomputation. Recomputing the whole 4K viewport costs 4.5 ms, so a perfect tile cache could save at most 4.5 ms of a 16 ms budget, at the price of a cache keyed by (VersionId, tile, zoom, graph_hash_prefix) that has to stay correct across every parameter change in the graph. That is a large correctness surface bought with a small number.

  2. It would make the actual problem worse. The stage that misses the budget is a convolution, and a tiled convolution reads a halo per tile. At a 52-pixel radius, 256-pixel tiles would read (256+104)² instead of 256² — very nearly twice the taps. Tiling is the wrong tool for the one stage that needs a tool.

So FR-DSP-2 becomes what §2 said it actually is for this architecture: a scheduling concern for export and thumbnailing, both of which already run off the frame path. The interactive path does not tile.

The stage that did need work — and it is not tiling

Resolved. The fix described below landed; the measurement is in §The reduced base, measured at the foot of this file, and docs/technical-debt.md TD-4 is closed. What follows is the reasoning as it stood, kept because it is what the numbers above argue for and because the tiling half of it is still live.

The measurement's real product is naming the stage. It is local_contrast, and the fix is stated in that module's own documentation:

The right optimisation is a base computed at reduced resolution, which needs a detail stage that can write a smaller target than it reads; that is a change to crate::detail, not to this file.

A Gaussian base at a quarter resolution is 1/16 the pixels at 1/4 the radius — about 1/64 of the work — and the result is visually identical because a base at σ = 26 px has no content above the quarter-resolution Nyquist to lose. That is a change to two files with a bounded blast radius, and it is what the 34 ms buys back. It should be tracked as its own item rather than smuggled in under a requirement about tiles.

FR-DSP-3 — the clause that should be narrowed

§3.3 proposes narrowing "when a full-resolution result is needed it is computed asynchronously, and the proxy result remains on screen until it is ready" to export and 1:1 zoom, or striking it.

M2 says strike it. The clause exists to hide the latency of a full-resolution render behind a proxy. There is no such latency: the 1:1 view is faster than the fit view on the fused path, and there is no second full-resolution code path to be asynchronous about — Framing::view is the whole mechanism. Export renders its own frames on a worker already. Keeping the clause would mean building a progressive-swap machine to conceal a render that completes in 1.4 ms.

FR-DSP-4 — satisfied vacuously, on the fused path

§4 makes progressive refinement conditional on M1 failing. On the fused path M1 passes, so reduced-quality rendering during a drag would buy nothing and cost the visible softness the requirement itself warns against.

The neighbourhood stage is the exception, and it is worth being precise: what that stage needs is not progressive refinement — it is a permanently cheaper base, computed at reduced resolution and correct at any moment the user stops. "Render coarse while dragging, sharpen when it settles" would paper over the same 34 ms with a visible swap. Fix the stage.

Since, 0.15.0. The stage was fixed (§ The reduced base, below), and the draft was built anyway, in the form this section would accept: the develop view renders at half resolution while a gesture moves and once at full resolution 120 ms after it stops (ui/dr-ui/src/refine.rs), the histogram dims while it describes an older frame, and the last draft fades out over 150 ms rather than being swapped. It is not a mask over a slow stage; it spares a drag the full-resolution frames it does not need.


Which GPU, on a machine with more than one

Measured 2026-08-29 on a laptop holding an Intel Iris Xe (RPL-P) and an AMD RX 5700 XT, same binary, adapter forced with VK_ICD_FILENAMES.

The question was whether an integrated GPU is the better choice for this application. The argument for it is good: a 24 MP frame is ~96 MB of RGBA, and on a discrete card every upload and every export readback crosses PCIe, where an iGPU shares memory with the CPU and crosses nothing. It also does not empty a battery.

The compute says otherwise, and not marginally.

2560×1600, p99 AMD RX 5700 XT Intel Iris Xe
fused pass, point 5.19 ms 7.75 ms
fused pass, all 11.70 ms 66.42 ms
M3 clarity, fit 4.67 ms 38.28 ms
M3 all four, 1:1 6.44 ms 57.67 ms
1920×1200, M3 clarity, fit 2.35 ms 19.99 ms

The fused colour pass is within a factor of 1.5 — it is one read and one write per pixel, which an iGPU does perfectly well. The neighbourhood stage is 5–8× slower, and that is what decides it: clarity at 1920×1200 costs 20 ms on the Iris Xe, so it leaves the budget on its own at the smallest size tested, before anything else in the chain runs.

So the default adapter preference stays Performance (dr_gpu::AdapterPreference).

Two things this does not show, and neither is a reason to revisit the default without measuring them:

  • It does not refute the transfer argument. This harness renders from a resident texture and never uploads or reads back, so the PCIe cost an iGPU avoids does not appear in any column above. Import, export and the thumbnail sweeps are transfer-heavy and compute-trivial, and may well go the other way — but they are not what FR-DSP-3 bounds, and one device is opened at startup and shared with the compositor, so there is currently no way to use a different adapter for a different task.
  • It says nothing about power. Efficiency remains offered (DARKROOM_GPU=integrated) because a user on battery may rationally accept a slower detail chain, and because someone whose discrete card has failed needs a way to keep working.

What is not measured here

Stated because §7 of display-and-extension.md asks for it, and because each of these could move the numbers.

  • Local adjustments. Not in any row above. Since 0.18.1 a layer is no longer a separate chain after the global one: each operation a layer touches runs a second fragment at its own place in the chain, blended by the layer's mask (architecture.md §5.2), and render_masked binds the mask array the fused shader samples. A heavily masked edit therefore costs more than all, by roughly one fragment per touched operation per layer.
  • Spot repairs. These add detail passes, and their cost is per spot.
  • Lens corrections. Not part of EditGraph::default_chain — they are built from a matched profile — so the point row does not include the warp chain.
  • Demosaic. Once per photograph on a worker, not on the frame path.
  • Presentation. The bench waits for the device to go idle inside the frame it measures. A real compositor overlaps frames, so these figures are an upper bound rather than an estimate.
  • One adapter. A discrete laptop GPU. The Intel iGPU on the same machine, and Android, will be slower — which is an argument for the conclusion rather than against it: the stage with no headroom has none to lose.

The CPU finding, which deserves its own item

EditGraph::compose costs 2.8–5.2 ms per frame on a full chain, at every resolution, because it is resolution-independent: it assembles a WGSL string and hashes it. On the all rows it is a third of what is left of the budget after the GPU has taken its share, and at 1920 × 1200 it is larger than the entire fused dispatch.

Nothing in this document's recommendations changes it, and it is the cheapest remaining win. The generated source depends only on the structure of the graph — that is what structure_hash already identifies, and it is precisely what does not change while a slider is being dragged, which is why the pipeline cache in AdjustPass does not recompile. The uniforms do change, but assembling them is a handful of floats per operation. So caching the source string against the structure hash and rebuilding only the uniforms would take these milliseconds to approximately nothing, on the path that needs them most. Worth its own entry in technical-debt.md.


The reduced base, measured

Status: Measured · 2026-08-29 · closes TD-4

DetailPass gained an output_scale, and clarity's base is now computed on a grid a quarter the size on each axis — the change §M3 argued for above.

Read this table on its own, not against the ones above. It was taken on a different adapter, so the absolute figures are not comparable with the RTX 3050 measurements this document is otherwise built from. What is comparable is the before and the after, which were measured on the same machine, same card, same release profile, minutes apart, with nothing between them but the change — the baseline at 0407fb8 and the result at bff95e2.

Adapter AMD Radeon RX 5700 XT (RADV NAVI10) (Vulkan)
Source 9504 × 6336 (60.2 MP, 482 MB as rgba16f)
Baseline 0407fb8, the branch's merge-base
Result bff95e2
Date 2026-08-29

M3 — clarity alone, before and after

Both percentiles, because they disagree and the disagreement is the interesting part.

size view before p50 after p50 before p99 after p99
1920 × 1200 fit 2.30ms 1.56ms 1.5× 3.94ms 1.97ms
1920 × 1200 1:1 2.88ms 1.99ms 1.4× 3.07ms 2.41ms
2560 × 1600 fit 4.41ms 1.95ms 2.3× 10.40ms 2.37ms
2560 × 1600 1:1 5.59ms 3.32ms 1.7× 5.78ms 3.85ms
3840 × 2160 fit 10.94ms 3.88ms 2.8× 25.05ms 4.17ms
3840 × 2160 1:1 13.31ms 6.36ms 2.1× 27.60ms 6.93ms

The honest headline is the p50 column: 2.8× at 4K. An earlier draft of this section led with the p99 ratio, which reads as 6.0× at the same size. That number is not supported, and the reason it is not is worth recording rather than quietly deleting.

The baseline run's fit rows have a p99/p50 spread of about 2.3×, while every row of the after run sits between 1.07× and 1.26×. A stage whose cost is radius × pixels has no reason to be bimodal, and the fit configuration is the memory-bound one — it walks the whole 482 MB source on a stride, where 1:1 reads a contiguous window. Something else was using the machine.

The cross-check settles it. §"Which GPU, on a machine with more than one" above measured the same baseline code on the same card independently, and reports M3 clarity, fit, 2560 × 1600 at 4.67 ms p99 — against the 10.40 ms in the table here. Two measurements of one thing that differ by 2.2× mean the noisier one is wrong, and it is this one.

So: the p50 ratios are the claim. The p99 improvement is real and larger, but this run cannot say by how much, and a clean re-measurement on a quiet machine is the way to find out.

What survives the caveat intact is the shape of the after column. Every figure is inside the 16 ms budget with a p99 within 26% of its median, at every size and both views — which is what a stage that is no longer the bottleneck looks like, whatever the exact ratio to what it replaced.

The one thing that is not a pure speed-up

The declared halo is now quantised to multiples of output_scale. The kernel truncates at 2σ, and that rounding now happens on the reduced grid before being multiplied back up:

viewport before after
1920 × 1200 29 px 28 px
2560 × 1600 38 px 40 px
3840 × 2160 52 px 52 px

At 2σ the Gaussian is already down to e⁻² of its peak, and crossing_the_reduction_threshold_does_not_change_the_picture holds the difference between a quarter-scale and a half-scale base to 0.03 stops of peak excursion and 2% of frame reach. But it is a change in reach rather than only in cost, it is what a tile scheduler would be handed, and it is worth knowing that the number moved rather than discovering it later as a seam.


The fit view, again — 2026-09-25

Status: Measured in the commits named, not re-run for this file.

Two changes to the fused path made the fit rows above cheaper again, each with its before and after in its commit message. Both were measured on the laptop RTX 3050 with its clocks held at 420/810 MHz by the power cap, on the synthetic 60 MP source of examples/frame_budget.rs, median of five alternated runs; both leave the rgba8 output bit-identical.

The source gather is read once per framing (1dc7b45). At fit every output pixel reads one texel on a stride through a source three or four times its width, and that gather was most of the fused pass. The pass now keeps a render-sized rgba16float cache of it, keyed on the framing, and reads it back while only the adjustments move:

scene before after
neutral, 2560 × 1600 fit 10.62 ms 3.88 ms
neutral, 3840 × 2160 fit 21.05 ms 7.11 ms
clarity, 3840 × 2160 fit 42.20 ms 27.88 ms
neutral, 2560 × 1600 1:1 (control) 3.83 ms 3.84 ms

An interpolated read (straightening, lens warps, CA) is not cached, and the cache is written on the second frame with a given key, so a crop or zoom drag pays nothing for it.

A detail pass that changes nothing is dropped (d430ec9). Capture sharpening at a scale too coarse to draw its radius emits an empty pass, which cost a full read and write when another neighbourhood operation followed it: sharpen with clarity at 2560 × 1600 fit went from 18.66 ms to 14.16 ms, and at 3840 × 2160 from 37.93 ms to 27.88 ms.

Dehaze in two passes — 2026-09-26

Status: Measured in the commit named, not re-run for this file.

Dehaze erodes each axis in one pass and recovers in the second (bee5c58, #74). It was five passes — a run and a span erosion along x, the same along y, and the recovery — and cost 22.9 ms of a 2560 × 1600 frame on the laptop RTX 3050, 54.1 ms at 3840 × 2160, with the memory clock held at 810 MHz by the power cap. At those clocks a detail pass costs what it reads and writes rather than what it taps: a pass with an empty body, one render-sized rgba16float read and write, measured 4.0 ms, and each dehaze pass 4.4–4.6 ms, so the taps were about 2 ms of the 22 and the four hand-offs between passes were the rest. Each axis now takes the minimum over its whole window directly, and the recovery rides in the y pass, which already holds the veil and the pixel's own colour: 36 texture reads a pixel in place of 12, nearly all cache hits, and two passes in place of five.

The same synthetic 60 MP source, only a detail parameter moving so the fused pass is reused, 30 frames a scene after six of warm-up, five runs of each binary alternated, median of the per-run p50:

scene before after
dehaze, 2560 × 1600 fit 22.88 ms 9.06 ms
dehaze, 2560 × 1600 1:1 23.41 ms 9.52 ms
dehaze, 3840 × 2160 fit 54.09 ms 28.12 ms
all five detail operations, 2560 × 1600 fit 53.11 ms 39.97 ms
all five detail operations, 2560 × 1600 1:1 67.48 ms 56.42 ms
every operation with film, 2560 × 1600 fit 57.59 ms 44.19 ms
every operation with film, 2560 × 1600 1:1 71.83 ms 57.93 ms

The five detail operations are noise reduction, sharpening, clarity, texture and dehaze; the scenes without dehaze moved within ±2%. The picture is the same bits: a minimum is exact in any order, the window is the one the split passes covered, and the rgba8 output hashed identically before and after in all 64 scene, view and size combinations measured.

The view pass after the detail stage — 2026-09-27

Status: Not measured. Every figure above predates it.

A chain with a detail stage is now one dispatch longer (c07f81e, D19). The fused pass used to end in the rendering — the base curve, then the output transform — before it stored, so every detail pass convolved display-referred values, and the last detail pass encoded. Now the fused pass stops before the view transform, every detail pass writes a scene-linear rgba16float intermediate, the last one included, and a view pass composed from the same inputs reads the result and runs the view transform (or the film stock), the output transform and the mask reveal.

What that adds, per frame with a detail stage: one render-sized read and write, and a third intermediate for a one-pass chain. By the dehaze section's own figure that is about 4 ms at 2560 × 1600 on the laptop RTX 3050 under its power cap. What it removed: a capture sharpening too fine to draw at the current scale no longer emits a pass-through, and there is no resolve pass for an active kernel with nothing to draw. A chain with no detail operation is unchanged, one fused dispatch with the view transform at its tail. The rows above that name a detail operation should be re-run before they are quoted.