Three places, because the finding has three audiences. The display spec's §1 table said FR-DSP-3 was unmeasured and FR-DSP-5 untagged. Both are now false, and §2's decision rule has fired. The body of §2 is left as written with the verdict quoted above it: a plan overtaken by its own evidence reads better in order than quietly edited into agreement with the outcome. TD-4 is the stage that misses the budget. Clarity's kernel is a fraction of the frame, so it reaches a 52-pixel radius at 4K and costs 34 ms — seven times the entire fused chain, for one slider. It is debt rather than a bug because the detail stage cannot yet write a target smaller than it reads, which `local_contrast`'s own documentation has said since it was written. The entry says plainly that tiles are the wrong tool for it, since that is exactly the conclusion a reader arriving from ARCH §5.3 would otherwise draw. TD-5 is the one nobody was looking for: composing the fused shader costs 2.8–5.2 ms of CPU per frame on a full chain, on the UI thread, which at 1920x1200 is more than the dispatch it precedes. The source depends only on the graph's structure — what `structure_hash` already identifies and what does not move during a drag — so the fix is the cache `AdjustPass` already keeps for compiled pipelines, one level up. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
225 lines
13 KiB
Markdown
225 lines
13 KiB
Markdown
# Finishing the display contract, and opening the pipeline
|
||
|
||
Spec for two pieces of work that turn out to be one conversation: closing **FR-DSP**, which is the
|
||
architecture's central performance claim, and reaching **FR-PLG**, which is the only requirement
|
||
family at zero.
|
||
|
||
They belong in one document because the same property decides both. The pipeline composes its work
|
||
from *declarations* — an operation says what its parameters are and contributes a WGSL fragment,
|
||
and the composer fuses the active ones into a single dispatch. That is why the display path is fast,
|
||
and it is also, already, most of a plugin format. Finishing one and opening the other are the same
|
||
seam approached from two sides.
|
||
|
||
---
|
||
|
||
## 1. What is actually true today
|
||
|
||
Stated first because both halves of this document are smaller than the requirement numbers suggest,
|
||
and the reason is that some of the work is done and untagged.
|
||
|
||
| Requirement | Reality |
|
||
|---|---|
|
||
| FR-DSP-1 proxy rendering | **Done.** The develop view renders at viewport resolution, not source. |
|
||
| FR-DSP-2 tiled computation | **Absent, and §2 now says it should stay that way.** Measured: the fused pass is inside the budget everywhere. See [frame-budget.md](frame-budget.md). |
|
||
| FR-DSP-3 interactive latency | **Measured and asserted** for the fused path — `core/dr-gpu/tests/frame_budget.rs`. Missed by one operation, clarity, for the reason recorded as TD-4. |
|
||
| FR-DSP-4 progressive refinement | **Absent**, and §4's condition did not fire. Every render is full quality and can afford to be. |
|
||
| FR-DSP-5 zoom and pan | **Done and tagged**, against tests that fail if the behaviour is removed — `core/dr-gpu/tests/zoom_resolution.rs`. `Framing::view` shrinks the sampled region while the render target keeps its size, so zooming *raises* the resolution the pipeline works at. That is FR-DSP-5's requirement, arrived at without tiles. |
|
||
| FR-DSP-6 colour management | **Done.** Output space is a parameter of composition. |
|
||
| FR-DSP-7 histogram and clipping | **Done**, GPU-side, no per-frame readback. |
|
||
| FR-DSP-8 per-display colour | **Absent.** One transform, not per-display. |
|
||
| FR-PLG-* | **Zero tagged.** But `ops/*.yaml` + `build.rs` is already the class-1 plugin format compiled at build time rather than loaded. |
|
||
|
||
Two of the five uncovered display requirements are therefore *measurement and tagging*, not
|
||
construction. That is worth knowing before anyone plans a quarter around them.
|
||
|
||
**Both have since been done.** [frame-budget.md](frame-budget.md) holds the measurements §2 asks
|
||
for and the reading of its decision rule; the table above is updated to match. The rest of this
|
||
document is left as it was written, because a plan that has been overtaken by its own evidence is
|
||
more useful read in order than quietly edited into agreement.
|
||
|
||
---
|
||
|
||
## 2. Measure before building tiles
|
||
|
||
**FR-DSP-2 is the one requirement in this document that may not be worth satisfying as written.**
|
||
|
||
> **Resolved.** M1–M3 were run; the numbers and the verdict are in
|
||
> [frame-budget.md](frame-budget.md). The rule below fired for *rewrite*: every point-operation
|
||
> chain is inside 16 ms at the 99th percentile at every viewport size, fit and at 1:1, the widest
|
||
> being 4.5 ms of GPU at 4K. The measurement did find a stage that misses the budget — clarity's
|
||
> 52-pixel kernel, 34 ms at 4K — and tiling makes that stage *worse*, since a tiled convolution
|
||
> reads a halo per tile. It is recorded as TD-4 with the fix its own module already names.
|
||
|
||
|
||
The requirement predates the fused-shader design. It assumes the pipeline is a chain of passes over
|
||
a large buffer, where recomputing everything on each frame would be ruinous and tiles are the way
|
||
out. What was built instead composes every active operation into **one dispatch over a
|
||
viewport-sized target** — at 2000×1300 that is 2.6 M pixels, once, for the whole chain.
|
||
|
||
So the question tiling was invented to answer may already be answered. Before any tile scheduler is
|
||
written:
|
||
|
||
**M1 — Frame cost at proxy resolution.** Time `render_detailed` at 1920×1200, 2560×1600 and
|
||
3840×2160, with a chain of one operation, five, and every operation active. Report the 99th
|
||
percentile, not the mean; a slider drag is judged by its worst frame.
|
||
|
||
**M2 — Frame cost at 1:1 on a large file.** The same, with `Framing::view` zoomed to 1:1 on a 60 MP
|
||
frame, which is the case FR-DSP-5 names and the one where the sampled region is smallest but the
|
||
detail chain's kernels are widest.
|
||
|
||
**M3 — Cost of the detail stage separately.** Neighbourhood operations dispatch per pass and are the
|
||
only part of the chain whose cost is not one read and one write. A separable blur at a large radius
|
||
is the plausible budget-breaker, not the fused pass.
|
||
|
||
**The decision rule, fixed in advance.** If M1 and M2 sit inside 16 ms at the 99th percentile,
|
||
**FR-DSP-2 is rewritten rather than implemented**: tiling stops being an interactive-path
|
||
requirement and becomes what it actually is for this architecture — a *scheduling* concern for
|
||
export and thumbnailing, which already run off the frame path. If they do not, the measurement tells
|
||
us which stage to tile, which is a far better starting point than tiling everything on principle.
|
||
|
||
Writing a tile scheduler that the design does not need would be the most expensive way to discover
|
||
this. ARCH §5.3's tile cache keyed by `(VersionId, tile, zoom, graph_hash_prefix)` is a good design
|
||
for a pipeline that needs it; the burden of proof is that this one does.
|
||
|
||
---
|
||
|
||
## 3. FR-DSP-3 — make the budget a test, not an aspiration
|
||
|
||
A latency requirement that nothing asserts is a wish. The work is:
|
||
|
||
**3.1** A bench in `dr-gpu` that renders a fixed chain at a fixed size and reports percentiles.
|
||
Committed with its numbers, so a regression is a diff rather than a memory.
|
||
|
||
**3.2** A test that *fails* when a frame exceeds the budget on the reference desktop, skipping where
|
||
there is no adapter — the pattern the GPU tests already use. It should assert the 99th percentile of
|
||
a hundred frames, because the failure mode being guarded against is a stutter, not an average.
|
||
|
||
**3.3** The asynchronous half of the requirement: "when a full-resolution result is needed it is
|
||
computed asynchronously, and the proxy result remains on screen until it is ready." Nothing does
|
||
this today because nothing needs a full-resolution result on the frame path — export renders its
|
||
own. This clause should be **narrowed to export and 1:1 zoom** or struck, and struck is defensible.
|
||
|
||
---
|
||
|
||
## 4. FR-DSP-4 — progressive refinement
|
||
|
||
The one genuinely new piece of interactive work, and it is small because the pipeline is already
|
||
resolution-parametric.
|
||
|
||
During a drag, render at a fraction of the viewport and let the compositor scale; when the gesture
|
||
settles, render at full viewport size. `DevelopSession` already knows when a drag is in flight —
|
||
`drag-changed` exists on every slider and is what stands the Flickable down.
|
||
|
||
Two things decide whether this is worth having, and M1 answers both. If a full-quality frame is
|
||
already inside budget, reduced-quality rendering buys nothing and costs a visible softness during
|
||
every drag — which the requirement itself warns against ("refinement is visually smooth, not a
|
||
jarring swap"). **This requirement is conditional on M1 failing.** If M1 passes, FR-DSP-4 is
|
||
satisfied vacuously: there is no rapid interaction the app cannot render at full quality, which is a
|
||
stronger outcome than refining.
|
||
|
||
---
|
||
|
||
## 5. FR-DSP-8 — per-display colour
|
||
|
||
The only display requirement needing platform work rather than pipeline work, and the only one where
|
||
being wrong is a correctness defect rather than a slow frame: a second monitor with a different
|
||
profile shows wrong colours, silently.
|
||
|
||
**5.1 Acquisition, per display server.** X11 has `_ICC_PROFILE` atoms per output. Wayland's
|
||
colour-management protocol is not universally available, and the requirement already anticipates
|
||
this by demanding "a defined fallback where Wayland provides no profile" — that fallback is sRGB,
|
||
stated in the About page beside the other diagnostics so a photographer can see which path they are
|
||
on rather than wonder.
|
||
|
||
**5.2 Reacting to a move.** The transform is selected per the display currently showing the canvas
|
||
and updates when the window moves. Slint reports window moves; the composed output space is already
|
||
a parameter of composition (`compose_with_framing(..., output)`), so a display change is a
|
||
recomposition, not a pipeline change. This is the part the existing design makes cheap.
|
||
|
||
**5.3 Fractional scaling.** "Handled without resampling artefacts in the canvas" — the canvas is a
|
||
wgpu texture handed to the compositor, so the requirement is that we render at the *physical* pixel
|
||
size rather than the logical one and let the compositor present 1:1. Worth an explicit test, since
|
||
the failure is subtle: a slightly soft canvas that looks like a bad demosaic.
|
||
|
||
---
|
||
|
||
## 6. Extensibility: the format already exists
|
||
|
||
`FR-PLG-2` says "the node declaration is the plugin format". That is already true — it is simply
|
||
resolved at build time:
|
||
|
||
```
|
||
ops/exposure.yaml ──build.rs──▶ generated Rust impl Operation ──▶ fused shader
|
||
```
|
||
|
||
A declaration names its parameters, their ranges and units, its attributes, its WGSL body and its
|
||
neutral. `build.rs` compiles that into something indistinguishable from a hand-written operation.
|
||
**Nothing about that requires the declaration to be present at compile time** — everything it
|
||
produces is data plus a WGSL string, and the composer already assembles WGSL at run time from
|
||
whatever operations are active.
|
||
|
||
So class 1 is not a new mechanism. It is the existing one, loaded later.
|
||
|
||
### 6.1 What has to change
|
||
|
||
**6.1.1 Descriptors become owned, not `&'static`.** `Operation::descriptor()` returns
|
||
`&'static OpDescriptor` today, which is what makes a build-time node free and a run-time node
|
||
impossible. This is the one invasive change in the whole plan and everything else waits behind it.
|
||
`Arc<OpDescriptor>` is the obvious shape; the cost is one refcount per descriptor read, on a path
|
||
that reads descriptors when the panel is built rather than per frame.
|
||
|
||
**6.1.2 A run-time node type.** One `DeclaredOp` implementing `Operation` from an owned
|
||
declaration, replacing *generated code per node* with *one interpreter over many declarations*. The
|
||
generated path can stay for the built-in chain — it costs nothing and keeps the built-ins
|
||
inspectable — but the two must produce identical behaviour, which is a test: parse each built-in
|
||
`ops/*.yaml` at run time and assert the composed WGSL matches the generated one byte for byte.
|
||
|
||
**6.1.3 WGSL validation at load, not at dispatch.** A plugin's fragment is a string from a stranger.
|
||
`compose` already builds a full shader and `naga` will reject bad source, but the failure currently
|
||
surfaces as a broken render. A plugin's source must be compiled and rejected at *load*, with the
|
||
error naming the plugin, because the alternative is an app that draws nothing and blames itself.
|
||
|
||
**6.1.4 Order and identity.** `order:` decides chain position and `build.rs` already refuses
|
||
duplicates — that guard becomes load-time. Plugin ids need a namespace (`author.name`) so two
|
||
plugins cannot collide, and the sidecar stores parameters by `(op_id, param_id)`, so an id collision
|
||
is a *wrong edit silently applied*, exactly the failure `MaskSource::Regions`' signature exists to
|
||
prevent.
|
||
|
||
### 6.2 What this buys immediately
|
||
|
||
The features enumerated as missing against Lightroom that are *pure point operations* become
|
||
declarations rather than code: split toning, colour zones, selective colour, creative vignette,
|
||
channel mixer variants. A photographer-author can write one without a Rust toolchain, and the
|
||
existing `ops/README.md` is already its documentation.
|
||
|
||
**It does not buy the neighbourhood operations** — dehaze, spot removal, liquify — because those are
|
||
`DetailStage` implementations with kernels and per-render scale conversion, which FR-PLG-2a
|
||
anticipates by naming "fragment nodes and pass nodes" as two templates. Pass nodes are a second
|
||
phase and should not gate the first.
|
||
|
||
### 6.3 Order of work
|
||
|
||
1. Owned descriptors (6.1.1) — invasive, unblocks everything, no user-visible change
|
||
2. `DeclaredOp` + byte-identical parity test against the generated built-ins (6.1.2)
|
||
3. Load-time WGSL validation and id namespacing (6.1.3, 6.1.4)
|
||
4. A directory that is read at startup, and one shipped example that is not a built-in
|
||
5. Pass nodes (FR-PLG-2a's second template), once 1–4 are load-bearing
|
||
|
||
Classes 2 and 3 — view plugins and computational plugins — are deliberately not in this plan.
|
||
FR-PLG-3a's "a view plugin cannot be trusted with the UI thread" and FR-PLG-4a's capability grants
|
||
are both larger design problems than class 1, and class 1 is where the requested features live.
|
||
|
||
---
|
||
|
||
## 7. What this document does not claim
|
||
|
||
Traceability counts a requirement as covered when a `TRACES` tag names it. It does not check that
|
||
the code under the tag does the thing — `FR-DEV-8` is currently tagged against instance-buffer
|
||
plumbing that a future spot-removal operation would use, and `FR-DEV-7` against a history row for a
|
||
frontend that does not exist. Both read as covered.
|
||
|
||
So the 51% figure is an overstatement of unknown size, and closing FR-DSP by tagging what already
|
||
works would make it a larger one. **Every requirement closed by this plan should be closed by a
|
||
test that would fail if the behaviour were removed**, which is the only kind of coverage worth
|
||
counting.
|