The figures are bee5c58's, quoted as that commit measured them: five passes to two, 22.9 ms to 9.1 ms at 2560 x 1600 fit and 54.1 ms to 28.1 ms at 4K, the output bit-identical. Under its own heading and status line, like the 2026-09-25 section, since this file was not re-run for them.
515 lines
26 KiB
Markdown
515 lines
26 KiB
Markdown
# What a frame costs
|
||
|
||
**Status:** Measured · 2026-08-27
|
||
**Companion to:** [display-and-extension.md](display-and-extension.md) §2–3 ·
|
||
[requirements.md](requirements.md) §3.4 FR-DSP-2, FR-DSP-3, FR-DSP-4
|
||
**Instrument:** [`core/dr-gpu/examples/frame_budget.rs`](../../core/dr-gpu/examples/frame_budget.rs)
|
||
**Guard:** [`core/dr-gpu/tests/frame_budget.rs`](../../core/dr-gpu/tests/frame_budget.rs)
|
||
|
||
[display-and-extension.md](display-and-extension.md) §2 fixed a decision rule in
|
||
advance and made three measurements the thing that settles it. This file is
|
||
those measurements, and the recommendation they support.
|
||
|
||
Rerun with:
|
||
|
||
```sh
|
||
cargo run --release -p dr-gpu --example frame_budget
|
||
```
|
||
|
||
and diff this file. That is the whole point of committing numbers: a regression
|
||
should be a diff rather than somebody's recollection of how fast it used to be.
|
||
|
||
---
|
||
|
||
## The answer, first
|
||
|
||
**FR-DSP-2 should be rewritten, not implemented.** M1 and M2 sit inside the
|
||
16 ms budget at the 99th percentile for every chain of point operations at every
|
||
viewport size measured, fit and at 1:1 — the widest case, every operation that
|
||
contributes a fragment to the fused shader at 4K, costs **4.5 ms** on the GPU and
|
||
**8.2 ms** including the composition that precedes it. Tiling the interactive
|
||
path would be optimising something that is already using a quarter of its budget.
|
||
|
||
**But the measurement did find a budget-breaker, and it is not the one tiling
|
||
fixes.** The neighbourhood stage — clarity in particular — costs **34 ms at 4K
|
||
on its own**, twice the whole budget, and tiles do not help it: a tile of a
|
||
convolution has to read its halo, so tiling raises the total tap count rather
|
||
than lowering it. §2 predicted this exactly ("a separable blur at a large radius
|
||
is the plausible budget-breaker, not the fused pass"), and the fix it needs is
|
||
the one `local_contrast`'s own module documentation already names — a base
|
||
computed at reduced resolution — which is a change to `crate::detail`, not a
|
||
tile scheduler.
|
||
|
||
There is a third finding nobody was looking for: **shader composition costs
|
||
3–5 ms of CPU per frame on a full chain**, on the UI thread, before any GPU work
|
||
is submitted. That is a fifth to a third of the budget spent formatting strings,
|
||
and it is invisible to any amount of tiling.
|
||
|
||
---
|
||
|
||
## Conditions
|
||
|
||
| | |
|
||
|---|---|
|
||
| Adapter | NVIDIA GeForce RTX 3050 6GB Laptop GPU (Vulkan) |
|
||
| Source | 9504 × 6336 synthetic (60.2 MP, 482 MB as `rgba16f`) |
|
||
| Frames | 100 measured per row, 12 warm-up frames discarded |
|
||
| Percentile | Nearest-rank, so p99 of 100 frames is the second-worst frame |
|
||
| Build | `--release` |
|
||
| Date | 2026-08-27 |
|
||
|
||
`shader` is `EditGraph::compose` alone. `cpu` adds the detail chain and the
|
||
invalidation hash — everything `DevelopSession::render` does per frame before it
|
||
dispatches. `gpu` is submit plus wait-for-idle, which serialises the GPU work
|
||
into the frame that caused it and is therefore pessimistic. `TOTAL` ranks
|
||
`cpu + gpu` summed **within each frame**, which is the column the budget is
|
||
judged on; adding two percentiles instead would invent a stutter that no frame
|
||
actually had.
|
||
|
||
Chains: `one` is exposure. `five` is exposure, contrast, highlights/shadows,
|
||
blacks/whites, vibrance. `point` is every operation in the default chain that
|
||
contributes a fragment to the fused shader, film stock included. `all` is `point`
|
||
plus the four neighbourhood operations — noise reduction, capture sharpening,
|
||
clarity and texture.
|
||
|
||
---
|
||
|
||
## M1 — the fused pass at proxy resolution
|
||
|
||
The develop view: the whole frame fit to the viewport.
|
||
|
||
| size | chain | shader | cpu p99 | gpu p50 | gpu p99 | TOTAL | |
|
||
|------------:|------:|-------:|--------:|--------:|--------:|--------:|:-----|
|
||
| 1920 × 1200 | one | 0.08ms | 0.10ms | 1.02ms | 1.23ms | 1.31ms | |
|
||
| 1920 × 1200 | five | 0.18ms | 0.20ms | 1.01ms | 1.20ms | 1.36ms | |
|
||
| 1920 × 1200 | point | 2.79ms | 2.82ms | 1.98ms | 2.18ms | 4.83ms | |
|
||
| 1920 × 1200 | all | 3.65ms | 4.73ms | 6.86ms | 7.37ms | 12.02ms | |
|
||
| 2560 × 1600 | one | 0.08ms | 0.10ms | 1.73ms | 2.00ms | 2.12ms | |
|
||
| 2560 × 1600 | five | 0.24ms | 0.27ms | 1.73ms | 2.26ms | 2.46ms | |
|
||
| 2560 × 1600 | point | 2.82ms | 2.85ms | 2.37ms | 2.65ms | 5.38ms | |
|
||
| 2560 × 1600 | all | 3.37ms | 4.35ms | 14.31ms | 15.65ms | 18.42ms | OVER |
|
||
| 3840 × 2160 | one | 0.10ms | 0.14ms | 3.09ms | 3.31ms | 3.42ms | |
|
||
| 3840 × 2160 | five | 0.30ms | 0.32ms | 3.03ms | 3.40ms | 3.61ms | |
|
||
| 3840 × 2160 | point | 3.62ms | 3.65ms | 4.12ms | 4.52ms | 8.23ms | |
|
||
| 3840 × 2160 | all | 4.12ms | 5.07ms | 37.73ms | 40.17ms | 43.24ms | OVER |
|
||
|
||
Read the `point` rows: **the fused dispatch scales with pixels and almost not at
|
||
all with chain length.** Going from one operation to the entire point chain at
|
||
4K costs 1.2 ms of GPU. Going from 2.3 M pixels to 8.3 M costs 2.3 ms. Both are
|
||
small, and the second is the one tiling would address.
|
||
|
||
The `all` rows go over, and the `point` rows in the same block are what say why:
|
||
the difference between them is the neighbourhood stage, measured on its own in
|
||
M3 and arriving at almost exactly the same figure.
|
||
|
||
## M2 — the same, zoomed to 1:1 on the 60 MP source
|
||
|
||
FR-DSP-5's case. `Framing::view` shrinks the sampled region while the render
|
||
target keeps its size, so one render pixel lands on one source pixel.
|
||
|
||
| size | chain | shader | cpu p99 | gpu p50 | gpu p99 | TOTAL | |
|
||
|------------:|------:|-------:|--------:|--------:|--------:|--------:|:-----|
|
||
| 1920 × 1200 | one | 0.12ms | 0.14ms | 0.42ms | 0.66ms | 0.75ms | |
|
||
| 1920 × 1200 | five | 0.27ms | 0.30ms | 0.49ms | 1.14ms | 1.22ms | |
|
||
| 1920 × 1200 | point | 3.49ms | 3.52ms | 1.18ms | 1.39ms | 4.85ms | |
|
||
| 1920 × 1200 | all | 3.64ms | 5.18ms | 8.96ms | 9.55ms | 14.30ms | |
|
||
| 2560 × 1600 | one | 0.11ms | 0.12ms | 0.56ms | 0.99ms | 1.06ms | |
|
||
| 2560 × 1600 | five | 0.22ms | 0.25ms | 0.77ms | 1.02ms | 1.17ms | |
|
||
| 2560 × 1600 | point | 2.96ms | 2.99ms | 2.03ms | 2.52ms | 5.61ms | |
|
||
| 2560 × 1600 | all | 5.24ms | 7.35ms | 18.80ms | 21.62ms | 25.81ms | OVER |
|
||
| 3840 × 2160 | one | 0.10ms | 0.12ms | 1.31ms | 1.52ms | 1.63ms | |
|
||
| 3840 × 2160 | five | 0.14ms | 0.27ms | 1.39ms | 1.64ms | 1.75ms | |
|
||
| 3840 × 2160 | point | 3.14ms | 3.16ms | 4.04ms | 4.50ms | 7.21ms | |
|
||
| 3840 × 2160 | all | 4.84ms | 6.78ms | 47.22ms | 48.79ms | 54.47ms | OVER |
|
||
|
||
**A 1:1 view of a 60 MP file is cheaper than the fit view of the same file**, for
|
||
every point chain and at every size — 1.52 ms against 3.31 ms for one operation
|
||
at 4K. That is not a rounding artefact and it is worth stating plainly, because
|
||
it is the opposite of what "full resolution" sounds like it should cost. The
|
||
dispatch is the same number of pixels either way; what changes is where those
|
||
pixels read from. A fit view walks the whole 482 MB texture on a stride, and a
|
||
1:1 view reads a contiguous window of it that fits comfortably in cache.
|
||
|
||
So the resolution FR-DSP-5 promises costs nothing extra on the fused path.
|
||
Zooming is not an expensive mode to be dreaded and progressively refined into;
|
||
it is the cheap one.
|
||
|
||
The `all` rows are worse at 1:1 than fit, and that is the detail stage again for
|
||
a specific reason: noise reduction's radius is stated in *source* pixels, so
|
||
`RenderScale::ratio` climbing to 1.0 widens its kernel. Clarity's is stated as a
|
||
fraction of the frame and does not move. M3 separates the two.
|
||
|
||
## M3 — the neighbourhood stage alone
|
||
|
||
Timed with the fused dispatch deliberately reused: only a detail parameter moves,
|
||
so `render_detailed` skips the colour pass (FR-DEV-3d) and what remains is the
|
||
convolutions. `colour` counts fused dispatches over the measured frames and is
|
||
zero on every row, which is what makes these numbers mean "detail alone" rather
|
||
than asserting it.
|
||
|
||
| size | stage | view | pass | radius | colour | cpu p99 | p50 | p99 |
|
||
|------------:|---------:|:-----|-----:|-------:|-------:|--------:|--------:|--------:|
|
||
| 1920 × 1200 | clarity | fit | 2 | 29 | 0 | 1.60ms | 5.42ms | 5.99ms |
|
||
| 1920 × 1200 | all four | fit | 7 | 29 | 0 | 1.71ms | 5.78ms | 6.16ms |
|
||
| 1920 × 1200 | clarity | 1:1 | 2 | 29 | 0 | 1.12ms | 7.51ms | 8.01ms |
|
||
| 1920 × 1200 | all four | 1:1 | 9 | 29 | 0 | 2.71ms | 8.25ms | 9.11ms |
|
||
| 2560 × 1600 | clarity | fit | 2 | 38 | 0 | 1.03ms | 12.02ms | 12.44ms |
|
||
| 2560 × 1600 | all four | fit | 7 | 38 | 0 | 1.87ms | 12.49ms | 13.16ms |
|
||
| 2560 × 1600 | clarity | 1:1 | 2 | 38 | 0 | 1.87ms | 15.82ms | 16.60ms |
|
||
| 2560 × 1600 | all four | 1:1 | 9 | 38 | 0 | 2.71ms | 17.24ms | 18.06ms |
|
||
| 3840 × 2160 | clarity | fit | 2 | 52 | 0 | 1.75ms | 33.11ms | 33.89ms |
|
||
| 3840 × 2160 | all four | fit | 7 | 52 | 0 | 1.76ms | 34.21ms | 35.03ms |
|
||
| 3840 × 2160 | clarity | 1:1 | 2 | 52 | 0 | 1.08ms | 40.39ms | 41.86ms |
|
||
| 3840 × 2160 | all four | 1:1 | 9 | 52 | 0 | 2.37ms | 43.29ms | 44.72ms |
|
||
|
||
`radius` is the widest halo any pass reads, in render pixels.
|
||
|
||
Clarity alone is 97% of the cost of all four neighbourhood operations together,
|
||
at every size. Its σ is 1.2% of the shorter edge and it truncates at 2σ, so its
|
||
radius is 29 px on a 1200 px viewport and **52 px at 4K** — two separable passes
|
||
of 105 taps each, over 8.3 M pixels, which is 1.7 billion texture reads. That is
|
||
the whole of the problem, and the numbers scale as `radius × pixels` exactly as
|
||
that description predicts: 5.99 → 12.44 → 33.89 ms for radii of 29 → 38 → 52 over
|
||
2.3 → 4.1 → 8.3 M pixels.
|
||
|
||
The extra cost at 1:1 is noise reduction and capture sharpening, whose radii are
|
||
properties of the sensor rather than of the frame. That is the correct behaviour
|
||
— it is why `RenderScale` has two units — and it is bounded by the kernel caps
|
||
those operations already declare.
|
||
|
||
---
|
||
|
||
## Reading this against §2's decision rule
|
||
|
||
§2: *"If M1 and M2 sit inside 16 ms at the 99th percentile, FR-DSP-2 is
|
||
rewritten rather than implemented … If they do not, the measurement tells us
|
||
which stage to tile."*
|
||
|
||
Both halves of the rule fire, on different stages, and the honest reading takes
|
||
both.
|
||
|
||
### FR-DSP-2 — rewrite it
|
||
|
||
For the fused pass the rule passes with a wide margin. Every point chain at
|
||
every size, fit and at 1:1, is inside 16 ms — the worst `TOTAL` is 8.23 ms and
|
||
the worst GPU figure is 4.52 ms. There is no viewport size on a desktop display
|
||
where recomputing the entire point chain over every visible pixel is a problem.
|
||
|
||
Two further reasons not to build the tile scheduler as written:
|
||
|
||
1. **Panning, which is the case ARCH §5.3's tile cache is designed for, gets no
|
||
benefit here.** Reusing already-valid tiles saves recomputation. Recomputing
|
||
the whole 4K viewport costs 4.5 ms, so a perfect tile cache could save at most
|
||
4.5 ms of a 16 ms budget, at the price of a cache keyed by
|
||
`(VersionId, tile, zoom, graph_hash_prefix)` that has to stay correct across
|
||
every parameter change in the graph. That is a large correctness surface
|
||
bought with a small number.
|
||
|
||
2. **It would make the actual problem worse.** The stage that misses the budget
|
||
is a convolution, and a tiled convolution reads a halo per tile. At a 52-pixel
|
||
radius, 256-pixel tiles would read (256+104)² instead of 256² — very nearly
|
||
*twice* the taps. Tiling is the wrong tool for the one stage that needs a
|
||
tool.
|
||
|
||
So FR-DSP-2 becomes what §2 said it actually is for this architecture: a
|
||
scheduling concern for export and thumbnailing, both of which already run off
|
||
the frame path. The interactive path does not tile.
|
||
|
||
### The stage that did need work — and it is not tiling
|
||
|
||
**Resolved.** The fix described below landed; the measurement is in
|
||
§[The reduced base, measured](#the-reduced-base-measured) at the foot of this
|
||
file, and `docs/technical-debt.md` TD-4 is closed. What follows is the
|
||
reasoning as it stood, kept because it is what the numbers above argue for and
|
||
because the tiling half of it is still live.
|
||
|
||
|
||
The measurement's real product is naming the stage. It is `local_contrast`, and
|
||
the fix is stated in that module's own documentation:
|
||
|
||
> The right optimisation is a base computed at reduced resolution, which needs a
|
||
> detail stage that can write a smaller target than it reads; that is a change to
|
||
> `crate::detail`, not to this file.
|
||
|
||
A Gaussian base at a quarter resolution is 1/16 the pixels at 1/4 the radius —
|
||
about 1/64 of the work — and the result is visually identical because a base at
|
||
σ = 26 px has no content above the quarter-resolution Nyquist to lose. That is a
|
||
change to two files with a bounded blast radius, and it is what the 34 ms buys
|
||
back. It should be tracked as its own item rather than smuggled in under a
|
||
requirement about tiles.
|
||
|
||
### FR-DSP-3 — the clause that should be narrowed
|
||
|
||
§3.3 proposes narrowing "when a full-resolution result is needed it is computed
|
||
asynchronously, and the proxy result remains on screen until it is ready" to
|
||
export and 1:1 zoom, or striking it.
|
||
|
||
**M2 says strike it.** The clause exists to hide the latency of a
|
||
full-resolution render behind a proxy. There is no such latency: the 1:1 view is
|
||
*faster* than the fit view on the fused path, and there is no second
|
||
full-resolution code path to be asynchronous about — `Framing::view` is the
|
||
whole mechanism. Export renders its own frames on a worker already. Keeping the
|
||
clause would mean building a progressive-swap machine to conceal a render that
|
||
completes in 1.4 ms.
|
||
|
||
### FR-DSP-4 — satisfied vacuously, on the fused path
|
||
|
||
§4 makes progressive refinement conditional on M1 failing. On the fused path M1
|
||
passes, so reduced-quality rendering during a drag would buy nothing and cost the
|
||
visible softness the requirement itself warns against.
|
||
|
||
The neighbourhood stage is the exception, and it is worth being precise: what
|
||
that stage needs is not *progressive* refinement — it is a permanently cheaper
|
||
base, computed at reduced resolution and correct at any moment the user stops.
|
||
"Render coarse while dragging, sharpen when it settles" would paper over the same
|
||
34 ms with a visible swap. Fix the stage.
|
||
|
||
> **Since, 0.15.0.** The stage was fixed (§ The reduced base, below), and the
|
||
> draft was built anyway, in the form this section would accept: the develop
|
||
> view renders at half resolution while a gesture moves and once at full
|
||
> resolution 120 ms after it stops (`ui/dr-ui/src/refine.rs`), the histogram
|
||
> dims while it describes an older frame, and the last draft fades out over
|
||
> 150 ms rather than being swapped. It is not a mask over a slow stage; it
|
||
> spares a drag the full-resolution frames it does not need.
|
||
|
||
---
|
||
|
||
## Which GPU, on a machine with more than one
|
||
|
||
**Measured 2026-08-29** on a laptop holding an Intel Iris Xe (RPL-P) and an AMD
|
||
RX 5700 XT, same binary, adapter forced with `VK_ICD_FILENAMES`.
|
||
|
||
The question was whether an integrated GPU is the better choice for this
|
||
application. The argument for it is good: a 24 MP frame is ~96 MB of RGBA, and
|
||
on a discrete card every upload and every export readback crosses PCIe, where
|
||
an iGPU shares memory with the CPU and crosses nothing. It also does not empty
|
||
a battery.
|
||
|
||
The compute says otherwise, and not marginally.
|
||
|
||
| 2560×1600, p99 | AMD RX 5700 XT | Intel Iris Xe |
|
||
|---|---|---|
|
||
| fused pass, `point` | 5.19 ms | 7.75 ms |
|
||
| fused pass, `all` | 11.70 ms | **66.42 ms** |
|
||
| M3 clarity, fit | 4.67 ms | **38.28 ms** |
|
||
| M3 all four, 1:1 | 6.44 ms | **57.67 ms** |
|
||
|
||
| 1920×1200, M3 clarity, fit | 2.35 ms | **19.99 ms** |
|
||
|---|---|---|
|
||
|
||
The fused colour pass is within a factor of 1.5 — it is one read and one write
|
||
per pixel, which an iGPU does perfectly well. The **neighbourhood stage is
|
||
5–8× slower**, and that is what decides it: clarity at 1920×1200 costs 20 ms on
|
||
the Iris Xe, so it leaves the budget on its own at the smallest size tested,
|
||
before anything else in the chain runs.
|
||
|
||
**So the default adapter preference stays `Performance`** (`dr_gpu::AdapterPreference`).
|
||
|
||
Two things this does *not* show, and neither is a reason to revisit the default
|
||
without measuring them:
|
||
|
||
- **It does not refute the transfer argument.** This harness renders from a
|
||
resident texture and never uploads or reads back, so the PCIe cost an iGPU
|
||
avoids does not appear in any column above. Import, export and the thumbnail
|
||
sweeps are transfer-heavy and compute-trivial, and may well go the other way
|
||
— but they are not what FR-DSP-3 bounds, and one device is opened at startup
|
||
and shared with the compositor, so there is currently no way to use a
|
||
different adapter for a different task.
|
||
- **It says nothing about power.** `Efficiency` remains offered
|
||
(`DARKROOM_GPU=integrated`) because a user on battery may rationally accept a
|
||
slower detail chain, and because someone whose discrete card has failed needs
|
||
a way to keep working.
|
||
|
||
## What is not measured here
|
||
|
||
Stated because §7 of [display-and-extension.md](display-and-extension.md) asks
|
||
for it, and because each of these could move the numbers.
|
||
|
||
- **Local adjustments.** The mask stack is a separate chain per layer and is not
|
||
in any row above. `render_masked` takes them and the fused shader addresses
|
||
them per layer, so a heavily masked edit costs more than `all`.
|
||
- **Spot repairs.** These add detail passes, and their cost is per spot.
|
||
- **Lens corrections.** Not part of `EditGraph::default_chain` — they are built
|
||
from a matched profile — so the `point` row does not include the warp chain.
|
||
- **Demosaic.** Once per photograph on a worker, not on the frame path.
|
||
- **Presentation.** The bench waits for the device to go idle inside the frame it
|
||
measures. A real compositor overlaps frames, so these figures are an upper
|
||
bound rather than an estimate.
|
||
- **One adapter.** A discrete laptop GPU. The Intel iGPU on the same machine, and
|
||
Android, will be slower — which is an argument for the conclusion rather than
|
||
against it: the stage with no headroom has none to lose.
|
||
|
||
## The CPU finding, which deserves its own item
|
||
|
||
`EditGraph::compose` costs 2.8–5.2 ms per frame on a full chain, at every
|
||
resolution, because it is resolution-independent: it assembles a WGSL string and
|
||
hashes it. On the `all` rows it is a third of what is left of the budget after
|
||
the GPU has taken its share, and at 1920 × 1200 it is larger than the entire
|
||
fused dispatch.
|
||
|
||
Nothing in this document's recommendations changes it, and it is the cheapest
|
||
remaining win. The generated *source* depends only on the structure of the graph
|
||
— that is what `structure_hash` already identifies, and it is precisely what does
|
||
not change while a slider is being dragged, which is why the pipeline cache in
|
||
`AdjustPass` does not recompile. The uniforms do change, but assembling them is a
|
||
handful of floats per operation. So caching the source string against the
|
||
structure hash and rebuilding only the uniforms would take these milliseconds to
|
||
approximately nothing, on the path that needs them most. Worth its own entry in
|
||
[technical-debt.md](technical-debt.md).
|
||
|
||
---
|
||
|
||
## The reduced base, measured
|
||
|
||
**Status:** Measured · 2026-08-29 · closes TD-4
|
||
|
||
`DetailPass` gained an `output_scale`, and clarity's base is now computed on a
|
||
grid a quarter the size on each axis — the change §M3 argued for above.
|
||
|
||
**Read this table on its own, not against the ones above.** It was taken on a
|
||
different adapter, so the absolute figures are not comparable with the RTX 3050
|
||
measurements this document is otherwise built from. What *is* comparable is the
|
||
before and the after, which were measured on the same machine, same card, same
|
||
release profile, minutes apart, with nothing between them but the change — the
|
||
baseline at `0407fb8` and the result at `bff95e2`.
|
||
|
||
| | |
|
||
|---|---|
|
||
| Adapter | AMD Radeon RX 5700 XT (RADV NAVI10) (Vulkan) |
|
||
| Source | 9504 × 6336 (60.2 MP, 482 MB as `rgba16f`) |
|
||
| Baseline | `0407fb8`, the branch's merge-base |
|
||
| Result | `bff95e2` |
|
||
| Date | 2026-08-29 |
|
||
|
||
### M3 — clarity alone, before and after
|
||
|
||
Both percentiles, because they disagree and the disagreement is the
|
||
interesting part.
|
||
|
||
| size | view | before p50 | after p50 | | before p99 | after p99 |
|
||
|------------:|:-----|-----------:|----------:|-----:|-----------:|----------:|
|
||
| 1920 × 1200 | fit | 2.30ms | 1.56ms | 1.5× | 3.94ms | 1.97ms |
|
||
| 1920 × 1200 | 1:1 | 2.88ms | 1.99ms | 1.4× | 3.07ms | 2.41ms |
|
||
| 2560 × 1600 | fit | 4.41ms | 1.95ms | 2.3× | 10.40ms | 2.37ms |
|
||
| 2560 × 1600 | 1:1 | 5.59ms | 3.32ms | 1.7× | 5.78ms | 3.85ms |
|
||
| 3840 × 2160 | fit | 10.94ms | 3.88ms | 2.8× | 25.05ms | 4.17ms |
|
||
| 3840 × 2160 | 1:1 | 13.31ms | 6.36ms | 2.1× | 27.60ms | 6.93ms |
|
||
|
||
**The honest headline is the p50 column: 2.8× at 4K.** An earlier draft of this
|
||
section led with the p99 ratio, which reads as 6.0× at the same size. That
|
||
number is not supported, and the reason it is not is worth recording rather
|
||
than quietly deleting.
|
||
|
||
The baseline run's `fit` rows have a p99/p50 spread of about 2.3×, while every
|
||
row of the after run sits between 1.07× and 1.26×. A stage whose cost is
|
||
`radius × pixels` has no reason to be bimodal, and the `fit` configuration is
|
||
the memory-bound one — it walks the whole 482 MB source on a stride, where
|
||
`1:1` reads a contiguous window. Something else was using the machine.
|
||
|
||
The cross-check settles it. §"Which GPU, on a machine with more than one"
|
||
above measured the *same baseline code on the same card* independently, and
|
||
reports M3 clarity, fit, 2560 × 1600 at **4.67 ms p99** — against the 10.40 ms
|
||
in the table here. Two measurements of one thing that differ by 2.2× mean the
|
||
noisier one is wrong, and it is this one.
|
||
|
||
So: the p50 ratios are the claim. The p99 improvement is real and larger, but
|
||
this run cannot say by how much, and a clean re-measurement on a quiet machine
|
||
is the way to find out.
|
||
|
||
What survives the caveat intact is the **shape** of the after column. Every
|
||
figure is inside the 16 ms budget with a p99 within 26% of its median, at every
|
||
size and both views — which is what a stage that is no longer the bottleneck
|
||
looks like, whatever the exact ratio to what it replaced.
|
||
|
||
### The one thing that is not a pure speed-up
|
||
|
||
**The declared halo is now quantised to multiples of `output_scale`.** The
|
||
kernel truncates at 2σ, and that rounding now happens on the reduced grid
|
||
before being multiplied back up:
|
||
|
||
| viewport | before | after |
|
||
|---|---:|---:|
|
||
| 1920 × 1200 | 29 px | 28 px |
|
||
| 2560 × 1600 | 38 px | 40 px |
|
||
| 3840 × 2160 | 52 px | 52 px |
|
||
|
||
At 2σ the Gaussian is already down to `e⁻²` of its peak, and
|
||
`crossing_the_reduction_threshold_does_not_change_the_picture` holds the
|
||
difference between a quarter-scale and a half-scale base to 0.03 stops of peak
|
||
excursion and 2% of frame reach. But it is a change in reach rather than only
|
||
in cost, it is what a tile scheduler would be handed, and it is worth knowing
|
||
that the number moved rather than discovering it later as a seam.
|
||
|
||
---
|
||
|
||
## The fit view, again — 2026-09-25
|
||
|
||
**Status:** Measured in the commits named, not re-run for this file.
|
||
|
||
Two changes to the fused path made the `fit` rows above cheaper again,
|
||
each with its before and after in its commit message. Both were measured
|
||
on the laptop RTX 3050 with its clocks held at 420/810 MHz by the power
|
||
cap, on the synthetic 60 MP source of `examples/frame_budget.rs`, median of
|
||
five alternated runs; both leave the rgba8 output bit-identical.
|
||
|
||
**The source gather is read once per framing** (`1dc7b45`). At fit every
|
||
output pixel reads one texel on a stride through a source three or four
|
||
times its width, and that gather was most of the fused pass. The pass now
|
||
keeps a render-sized `rgba16float` cache of it, keyed on the framing, and
|
||
reads it back while only the adjustments move:
|
||
|
||
| scene | before | after |
|
||
|---|---:|---:|
|
||
| neutral, 2560 × 1600 fit | 10.62 ms | 3.88 ms |
|
||
| neutral, 3840 × 2160 fit | 21.05 ms | 7.11 ms |
|
||
| clarity, 3840 × 2160 fit | 42.20 ms | 27.88 ms |
|
||
| neutral, 2560 × 1600 1:1 (control) | 3.83 ms | 3.84 ms |
|
||
|
||
An interpolated read (straightening, lens warps, CA) is not cached, and the
|
||
cache is written on the second frame with a given key, so a crop or zoom
|
||
drag pays nothing for it.
|
||
|
||
**A detail pass that changes nothing is dropped** (`d430ec9`). Capture
|
||
sharpening at a scale too coarse to draw its radius emits an empty pass,
|
||
which cost a full read and write when another neighbourhood operation
|
||
followed it: sharpen with clarity at 2560 × 1600 fit went from 18.66 ms to
|
||
14.16 ms, and at 3840 × 2160 from 37.93 ms to 27.88 ms.
|
||
|
||
## Dehaze in two passes — 2026-09-26
|
||
|
||
**Status:** Measured in the commit named, not re-run for this file.
|
||
|
||
**Dehaze erodes each axis in one pass and recovers in the second**
|
||
(`bee5c58`, #74). It was five passes — a run and a span erosion along x,
|
||
the same along y, and the recovery — and cost 22.9 ms of a 2560 × 1600
|
||
frame on the laptop RTX 3050, 54.1 ms at 3840 × 2160, with the memory
|
||
clock held at 810 MHz by the power cap. At those clocks a detail pass costs
|
||
what it reads and writes rather than what it taps: a pass with an empty
|
||
body, one render-sized `rgba16float` read and write, measured 4.0 ms, and
|
||
each dehaze pass 4.4–4.6 ms, so the taps were about 2 ms of the 22 and the
|
||
four hand-offs between passes were the rest. Each axis now takes the
|
||
minimum over its whole window directly, and the recovery rides in the y
|
||
pass, which already holds the veil and the pixel's own colour: 36 texture
|
||
reads a pixel in place of 12, nearly all cache hits, and two passes in
|
||
place of five.
|
||
|
||
The same synthetic 60 MP source, only a detail parameter moving so the
|
||
fused pass is reused, 30 frames a scene after six of warm-up, five runs of
|
||
each binary alternated, median of the per-run p50:
|
||
|
||
| scene | before | after |
|
||
|---|---:|---:|
|
||
| dehaze, 2560 × 1600 fit | 22.88 ms | 9.06 ms |
|
||
| dehaze, 2560 × 1600 1:1 | 23.41 ms | 9.52 ms |
|
||
| dehaze, 3840 × 2160 fit | 54.09 ms | 28.12 ms |
|
||
| all five detail operations, 2560 × 1600 fit | 53.11 ms | 39.97 ms |
|
||
| all five detail operations, 2560 × 1600 1:1 | 67.48 ms | 56.42 ms |
|
||
| every operation with film, 2560 × 1600 fit | 57.59 ms | 44.19 ms |
|
||
| every operation with film, 2560 × 1600 1:1 | 71.83 ms | 57.93 ms |
|
||
|
||
The five detail operations are noise reduction, sharpening, clarity,
|
||
texture and dehaze; the scenes without dehaze moved within ±2%. The
|
||
picture is the same bits: a minimum is exact in any order, the window is
|
||
the one the split passes covered, and the rgba8 output hashed identically
|
||
before and after in all 64 scene, view and size combinations measured.
|