Two families were the weak points: FR-DSP at 37%, which is the architecture's central performance claim, and FR-PLG at zero. They are one document because the same property decides both — the pipeline composes its work from declarations, which is why the display path is fast and is also, already, most of a plugin format. Two findings change the size of the job. **Some of FR-DSP is done and untagged.** Zoom already meets FR-DSP-5: the sampled region shrinks while the render target keeps its size, so zooming raises the resolution the pipeline works at — arrived at without tiles. **FR-DSP-2 may not be worth satisfying as written.** It predates the fused shader and assumes a chain of passes over a large buffer, where recomputing per frame is ruinous and tiles are the way out. What exists composes every active operation into one dispatch over a viewport-sized target. So the spec fixes three measurements and a decision rule *in advance*: if a frame sits inside budget at the 99th percentile, the requirement is rewritten rather than implemented, and tiling becomes what it actually is here — a scheduling concern for export, which already runs off the frame path. Writing a tile scheduler the design does not need would be the most expensive way to find that out. **The plugin format already exists**, resolved at build time: `ops/*.yaml` through `build.rs` produces something indistinguishable from a hand-written operation, and everything it produces is data plus a WGSL string. The blocker is one line — `descriptor()` returns `&'static` — and the rest is an interpreter over declarations, load-time WGSL validation, and namespaced ids, because the sidecar stores parameters by id and a collision is a wrong edit silently applied. Recorded last and deliberately: traceability counts a tag, not a behaviour. `FR-DEV-8` is tagged against plumbing a future spot-removal op would use and `FR-DEV-7` against a history row for a frontend that does not exist, so 51% is an overstatement of unknown size. Every requirement this plan closes should be closed by a test that fails if the behaviour is removed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
212 lines
12 KiB
Markdown
212 lines
12 KiB
Markdown
# Finishing the display contract, and opening the pipeline
|
||
|
||
Spec for two pieces of work that turn out to be one conversation: closing **FR-DSP**, which is the
|
||
architecture's central performance claim, and reaching **FR-PLG**, which is the only requirement
|
||
family at zero.
|
||
|
||
They belong in one document because the same property decides both. The pipeline composes its work
|
||
from *declarations* — an operation says what its parameters are and contributes a WGSL fragment,
|
||
and the composer fuses the active ones into a single dispatch. That is why the display path is fast,
|
||
and it is also, already, most of a plugin format. Finishing one and opening the other are the same
|
||
seam approached from two sides.
|
||
|
||
---
|
||
|
||
## 1. What is actually true today
|
||
|
||
Stated first because both halves of this document are smaller than the requirement numbers suggest,
|
||
and the reason is that some of the work is done and untagged.
|
||
|
||
| Requirement | Reality |
|
||
|---|---|
|
||
| FR-DSP-1 proxy rendering | **Done.** The develop view renders at viewport resolution, not source. |
|
||
| FR-DSP-2 tiled computation | **Absent.** The fused pass renders the whole viewport in one dispatch. |
|
||
| FR-DSP-3 interactive latency | **Unmeasured.** No frame budget is asserted anywhere. |
|
||
| FR-DSP-4 progressive refinement | **Absent.** Every render is full quality. |
|
||
| FR-DSP-5 zoom and pan | **Substantially done, untagged.** `Framing::view` shrinks the sampled region while the render target keeps its size, so zooming *raises* the resolution the pipeline works at. That is FR-DSP-5's requirement, arrived at without tiles. |
|
||
| FR-DSP-6 colour management | **Done.** Output space is a parameter of composition. |
|
||
| FR-DSP-7 histogram and clipping | **Done**, GPU-side, no per-frame readback. |
|
||
| FR-DSP-8 per-display colour | **Absent.** One transform, not per-display. |
|
||
| FR-PLG-* | **Zero tagged.** But `ops/*.yaml` + `build.rs` is already the class-1 plugin format compiled at build time rather than loaded. |
|
||
|
||
Two of the five uncovered display requirements are therefore *measurement and tagging*, not
|
||
construction. That is worth knowing before anyone plans a quarter around them.
|
||
|
||
---
|
||
|
||
## 2. Measure before building tiles
|
||
|
||
**FR-DSP-2 is the one requirement in this document that may not be worth satisfying as written.**
|
||
|
||
The requirement predates the fused-shader design. It assumes the pipeline is a chain of passes over
|
||
a large buffer, where recomputing everything on each frame would be ruinous and tiles are the way
|
||
out. What was built instead composes every active operation into **one dispatch over a
|
||
viewport-sized target** — at 2000×1300 that is 2.6 M pixels, once, for the whole chain.
|
||
|
||
So the question tiling was invented to answer may already be answered. Before any tile scheduler is
|
||
written:
|
||
|
||
**M1 — Frame cost at proxy resolution.** Time `render_detailed` at 1920×1200, 2560×1600 and
|
||
3840×2160, with a chain of one operation, five, and every operation active. Report the 99th
|
||
percentile, not the mean; a slider drag is judged by its worst frame.
|
||
|
||
**M2 — Frame cost at 1:1 on a large file.** The same, with `Framing::view` zoomed to 1:1 on a 60 MP
|
||
frame, which is the case FR-DSP-5 names and the one where the sampled region is smallest but the
|
||
detail chain's kernels are widest.
|
||
|
||
**M3 — Cost of the detail stage separately.** Neighbourhood operations dispatch per pass and are the
|
||
only part of the chain whose cost is not one read and one write. A separable blur at a large radius
|
||
is the plausible budget-breaker, not the fused pass.
|
||
|
||
**The decision rule, fixed in advance.** If M1 and M2 sit inside 16 ms at the 99th percentile,
|
||
**FR-DSP-2 is rewritten rather than implemented**: tiling stops being an interactive-path
|
||
requirement and becomes what it actually is for this architecture — a *scheduling* concern for
|
||
export and thumbnailing, which already run off the frame path. If they do not, the measurement tells
|
||
us which stage to tile, which is a far better starting point than tiling everything on principle.
|
||
|
||
Writing a tile scheduler that the design does not need would be the most expensive way to discover
|
||
this. ARCH §5.3's tile cache keyed by `(VersionId, tile, zoom, graph_hash_prefix)` is a good design
|
||
for a pipeline that needs it; the burden of proof is that this one does.
|
||
|
||
---
|
||
|
||
## 3. FR-DSP-3 — make the budget a test, not an aspiration
|
||
|
||
A latency requirement that nothing asserts is a wish. The work is:
|
||
|
||
**3.1** A bench in `dr-gpu` that renders a fixed chain at a fixed size and reports percentiles.
|
||
Committed with its numbers, so a regression is a diff rather than a memory.
|
||
|
||
**3.2** A test that *fails* when a frame exceeds the budget on the reference desktop, skipping where
|
||
there is no adapter — the pattern the GPU tests already use. It should assert the 99th percentile of
|
||
a hundred frames, because the failure mode being guarded against is a stutter, not an average.
|
||
|
||
**3.3** The asynchronous half of the requirement: "when a full-resolution result is needed it is
|
||
computed asynchronously, and the proxy result remains on screen until it is ready." Nothing does
|
||
this today because nothing needs a full-resolution result on the frame path — export renders its
|
||
own. This clause should be **narrowed to export and 1:1 zoom** or struck, and struck is defensible.
|
||
|
||
---
|
||
|
||
## 4. FR-DSP-4 — progressive refinement
|
||
|
||
The one genuinely new piece of interactive work, and it is small because the pipeline is already
|
||
resolution-parametric.
|
||
|
||
During a drag, render at a fraction of the viewport and let the compositor scale; when the gesture
|
||
settles, render at full viewport size. `DevelopSession` already knows when a drag is in flight —
|
||
`drag-changed` exists on every slider and is what stands the Flickable down.
|
||
|
||
Two things decide whether this is worth having, and M1 answers both. If a full-quality frame is
|
||
already inside budget, reduced-quality rendering buys nothing and costs a visible softness during
|
||
every drag — which the requirement itself warns against ("refinement is visually smooth, not a
|
||
jarring swap"). **This requirement is conditional on M1 failing.** If M1 passes, FR-DSP-4 is
|
||
satisfied vacuously: there is no rapid interaction the app cannot render at full quality, which is a
|
||
stronger outcome than refining.
|
||
|
||
---
|
||
|
||
## 5. FR-DSP-8 — per-display colour
|
||
|
||
The only display requirement needing platform work rather than pipeline work, and the only one where
|
||
being wrong is a correctness defect rather than a slow frame: a second monitor with a different
|
||
profile shows wrong colours, silently.
|
||
|
||
**5.1 Acquisition, per display server.** X11 has `_ICC_PROFILE` atoms per output. Wayland's
|
||
colour-management protocol is not universally available, and the requirement already anticipates
|
||
this by demanding "a defined fallback where Wayland provides no profile" — that fallback is sRGB,
|
||
stated in the About page beside the other diagnostics so a photographer can see which path they are
|
||
on rather than wonder.
|
||
|
||
**5.2 Reacting to a move.** The transform is selected per the display currently showing the canvas
|
||
and updates when the window moves. Slint reports window moves; the composed output space is already
|
||
a parameter of composition (`compose_with_framing(..., output)`), so a display change is a
|
||
recomposition, not a pipeline change. This is the part the existing design makes cheap.
|
||
|
||
**5.3 Fractional scaling.** "Handled without resampling artefacts in the canvas" — the canvas is a
|
||
wgpu texture handed to the compositor, so the requirement is that we render at the *physical* pixel
|
||
size rather than the logical one and let the compositor present 1:1. Worth an explicit test, since
|
||
the failure is subtle: a slightly soft canvas that looks like a bad demosaic.
|
||
|
||
---
|
||
|
||
## 6. Extensibility: the format already exists
|
||
|
||
`FR-PLG-2` says "the node declaration is the plugin format". That is already true — it is simply
|
||
resolved at build time:
|
||
|
||
```
|
||
ops/exposure.yaml ──build.rs──▶ generated Rust impl Operation ──▶ fused shader
|
||
```
|
||
|
||
A declaration names its parameters, their ranges and units, its attributes, its WGSL body and its
|
||
neutral. `build.rs` compiles that into something indistinguishable from a hand-written operation.
|
||
**Nothing about that requires the declaration to be present at compile time** — everything it
|
||
produces is data plus a WGSL string, and the composer already assembles WGSL at run time from
|
||
whatever operations are active.
|
||
|
||
So class 1 is not a new mechanism. It is the existing one, loaded later.
|
||
|
||
### 6.1 What has to change
|
||
|
||
**6.1.1 Descriptors become owned, not `&'static`.** `Operation::descriptor()` returns
|
||
`&'static OpDescriptor` today, which is what makes a build-time node free and a run-time node
|
||
impossible. This is the one invasive change in the whole plan and everything else waits behind it.
|
||
`Arc<OpDescriptor>` is the obvious shape; the cost is one refcount per descriptor read, on a path
|
||
that reads descriptors when the panel is built rather than per frame.
|
||
|
||
**6.1.2 A run-time node type.** One `DeclaredOp` implementing `Operation` from an owned
|
||
declaration, replacing *generated code per node* with *one interpreter over many declarations*. The
|
||
generated path can stay for the built-in chain — it costs nothing and keeps the built-ins
|
||
inspectable — but the two must produce identical behaviour, which is a test: parse each built-in
|
||
`ops/*.yaml` at run time and assert the composed WGSL matches the generated one byte for byte.
|
||
|
||
**6.1.3 WGSL validation at load, not at dispatch.** A plugin's fragment is a string from a stranger.
|
||
`compose` already builds a full shader and `naga` will reject bad source, but the failure currently
|
||
surfaces as a broken render. A plugin's source must be compiled and rejected at *load*, with the
|
||
error naming the plugin, because the alternative is an app that draws nothing and blames itself.
|
||
|
||
**6.1.4 Order and identity.** `order:` decides chain position and `build.rs` already refuses
|
||
duplicates — that guard becomes load-time. Plugin ids need a namespace (`author.name`) so two
|
||
plugins cannot collide, and the sidecar stores parameters by `(op_id, param_id)`, so an id collision
|
||
is a *wrong edit silently applied*, exactly the failure `MaskSource::Regions`' signature exists to
|
||
prevent.
|
||
|
||
### 6.2 What this buys immediately
|
||
|
||
The features enumerated as missing against Lightroom that are *pure point operations* become
|
||
declarations rather than code: split toning, colour zones, selective colour, creative vignette,
|
||
channel mixer variants. A photographer-author can write one without a Rust toolchain, and the
|
||
existing `ops/README.md` is already its documentation.
|
||
|
||
**It does not buy the neighbourhood operations** — dehaze, spot removal, liquify — because those are
|
||
`DetailStage` implementations with kernels and per-render scale conversion, which FR-PLG-2a
|
||
anticipates by naming "fragment nodes and pass nodes" as two templates. Pass nodes are a second
|
||
phase and should not gate the first.
|
||
|
||
### 6.3 Order of work
|
||
|
||
1. Owned descriptors (6.1.1) — invasive, unblocks everything, no user-visible change
|
||
2. `DeclaredOp` + byte-identical parity test against the generated built-ins (6.1.2)
|
||
3. Load-time WGSL validation and id namespacing (6.1.3, 6.1.4)
|
||
4. A directory that is read at startup, and one shipped example that is not a built-in
|
||
5. Pass nodes (FR-PLG-2a's second template), once 1–4 are load-bearing
|
||
|
||
Classes 2 and 3 — view plugins and computational plugins — are deliberately not in this plan.
|
||
FR-PLG-3a's "a view plugin cannot be trusted with the UI thread" and FR-PLG-4a's capability grants
|
||
are both larger design problems than class 1, and class 1 is where the requested features live.
|
||
|
||
---
|
||
|
||
## 7. What this document does not claim
|
||
|
||
Traceability counts a requirement as covered when a `TRACES` tag names it. It does not check that
|
||
the code under the tag does the thing — `FR-DEV-8` is currently tagged against instance-buffer
|
||
plumbing that a future spot-removal operation would use, and `FR-DEV-7` against a history row for a
|
||
frontend that does not exist. Both read as covered.
|
||
|
||
So the 51% figure is an overstatement of unknown size, and closing FR-DSP by tagging what already
|
||
works would make it a larger one. **Every requirement closed by this plan should be closed by a
|
||
test that would fail if the behaviour were removed**, which is the only kind of coverage worth
|
||
counting.
|