Record progressive refinement as built, and the fit view's two speed-ups

display-and-extension.md still called FR-DSP-4 absent, and
frame-budget.md ended its reading at "satisfied vacuously". 0.15.0 built
it (refine.rs, the provisional histogram, the fade), so the table row
says so, frame-budget.md gains a note under its FR-DSP-4 section, and the
register gains a status note like FR-RAW-2's and FR-UI-5's.

frame-budget.md also gains a short section with the figures 1dc7b45
(the fit view's gather cached per framing) and d430ec9 (an identity
detail pass dropped) measured, since its fit rows no longer describe the
fused path. They are quoted from those commits, on the machine they name,
and are marked as not re-run here.
This commit is contained in:
2026-09-26 07:45:26 -04:00
parent 79c051e8a6
commit e31990550f
3 changed files with 50 additions and 1 deletions
+43
View File
@@ -264,6 +264,14 @@ base, computed at reduced resolution and correct at any moment the user stops.
"Render coarse while dragging, sharpen when it settles" would paper over the same
34 ms with a visible swap. Fix the stage.
> **Since, 0.15.0.** The stage was fixed (§ The reduced base, below), and the
> draft was built anyway, in the form this section would accept: the develop
> view renders at half resolution while a gesture moves and once at full
> resolution 120 ms after it stops (`ui/dr-ui/src/refine.rs`), the histogram
> dims while it describes an older frame, and the last draft fades out over
> 150 ms rather than being swapped. It is not a mask over a slow stage; it
> spares a drag the full-resolution frames it does not need.
---
## Which GPU, on a machine with more than one
@@ -431,3 +439,38 @@ difference between a quarter-scale and a half-scale base to 0.03 stops of peak
excursion and 2% of frame reach. But it is a change in reach rather than only
in cost, it is what a tile scheduler would be handed, and it is worth knowing
that the number moved rather than discovering it later as a seam.
---
## The fit view, again — 2026-09-25
**Status:** Measured in the commits named, not re-run for this file.
Two changes to the fused path made the `fit` rows above cheaper again,
each with its before and after in its commit message. Both were measured
on the laptop RTX 3050 with its clocks held at 420/810 MHz by the power
cap, on the synthetic 60 MP source of `examples/frame_budget.rs`, median of
five alternated runs; both leave the rgba8 output bit-identical.
**The source gather is read once per framing** (`1dc7b45`). At fit every
output pixel reads one texel on a stride through a source three or four
times its width, and that gather was most of the fused pass. The pass now
keeps a render-sized `rgba16float` cache of it, keyed on the framing, and
reads it back while only the adjustments move:
| scene | before | after |
|---|---:|---:|
| neutral, 2560 × 1600 fit | 10.62 ms | 3.88 ms |
| neutral, 3840 × 2160 fit | 21.05 ms | 7.11 ms |
| clarity, 3840 × 2160 fit | 42.20 ms | 27.88 ms |
| neutral, 2560 × 1600 1:1 (control) | 3.83 ms | 3.84 ms |
An interpolated read (straightening, lens warps, CA) is not cached, and the
cache is written on the second frame with a given key, so a crop or zoom
drag pays nothing for it.
**A detail pass that changes nothing is dropped** (`d430ec9`). Capture
sharpening at a scale too coarse to draw its radius emits an empty pass,
which cost a full read and write when another neighbourhood operation
followed it: sharpen with clarity at 2560 × 1600 fit went from 18.66 ms to
14.16 ms, and at 3840 × 2160 from 37.93 ms to 27.88 ms.