# Finishing the display contract, and opening the pipeline Spec for two pieces of work that turn out to be one conversation: closing **FR-DSP**, which is the architecture's central performance claim, and reaching **FR-PLG**, which is the only requirement family at zero. They belong in one document because the same property decides both. The pipeline composes its work from *declarations* — an operation says what its parameters are and contributes a WGSL fragment, and the composer fuses the active ones into a single dispatch. That is why the display path is fast, and it is also, already, most of a plugin format. Finishing one and opening the other are the same seam approached from two sides. --- ## 1. What is actually true today Stated first because both halves of this document are smaller than the requirement numbers suggest, and the reason is that some of the work is done and untagged. | Requirement | Reality | |---|---| | FR-DSP-1 proxy rendering | **Done.** The develop view renders at viewport resolution, not source. | | FR-DSP-2 tiled computation | **Absent from the interactive path, and §2 now says it should stay that way.** Measured: the fused pass is inside the budget everywhere. See [frame-budget.md](frame-budget.md). *Since 0.19.0* the export of a linear DNG too large for one texture is drawn in halo-grown tiles, and the canvas renders such a file from a reduced copy and full-resolution windows (ARCH §5.3). | | FR-DSP-3 interactive latency | **Measured and asserted** for the fused path — `core/dr-gpu/tests/frame_budget.rs`. Missed by one operation, clarity, for the reason recorded as TD-4. | | FR-DSP-4 progressive refinement | **Built** (0.15.0), although §4's condition did not fire on the fused path: a half-resolution draft while a gesture moves, one sharp frame 120 ms after it stops, the histogram dimmed while it lags, and the draft faded out over 150 ms — `ui/dr-ui/src/refine.rs`. See [frame-budget.md](frame-budget.md). | | FR-DSP-5 zoom and pan | **Done and tagged**, against tests that fail if the behaviour is removed — `core/dr-gpu/tests/zoom_resolution.rs`. `Framing::view` shrinks the sampled region while the render target keeps its size, so zooming *raises* the resolution the pipeline works at. That is FR-DSP-5's requirement, arrived at without tiles. | | FR-DSP-6 colour management | **Done.** Output space is a parameter of composition. | | FR-DSP-7 histogram and clipping | **Done**, GPU-side, no per-frame readback. | | FR-DSP-8 per-display colour | **Done**, with one caveat named in §5.4. Acquisition per display server, an sRGB fallback that is visible in About, and the canvas rendered at physical pixel size. | | FR-PLG-* | **Zero tagged.** But `ops/*.yaml` + `build.rs` is already the class-1 plugin format compiled at build time rather than loaded. | Two of the five uncovered display requirements are therefore *measurement and tagging*, not construction. That is worth knowing before anyone plans a quarter around them. **Both have since been done.** [frame-budget.md](frame-budget.md) holds the measurements §2 asks for and the reading of its decision rule; the table above is updated to match. The rest of this document is left as it was written, because a plan that has been overtaken by its own evidence is more useful read in order than quietly edited into agreement. --- ## 2. Measure before building tiles **FR-DSP-2 is the one requirement in this document that may not be worth satisfying as written.** > **Resolved.** M1–M3 were run; the numbers and the verdict are in > [frame-budget.md](frame-budget.md). The rule below fired for *rewrite*: every point-operation > chain is inside 16 ms at the 99th percentile at every viewport size, fit and at 1:1, the widest > being 4.5 ms of GPU at 4K. The measurement did find a stage that misses the budget — clarity's > 52-pixel kernel, 34 ms at 4K — and tiling makes that stage *worse*, since a tiled convolution > reads a halo per tile. It is recorded as TD-4 with the fix its own module already names. The requirement predates the fused-shader design. It assumes the pipeline is a chain of passes over a large buffer, where recomputing everything on each frame would be ruinous and tiles are the way out. What was built instead composes every active operation into **one dispatch over a viewport-sized target** — at 2000×1300 that is 2.6 M pixels, once, for the whole chain. So the question tiling was invented to answer may already be answered. Before any tile scheduler is written: **M1 — Frame cost at proxy resolution.** Time `render_detailed` at 1920×1200, 2560×1600 and 3840×2160, with a chain of one operation, five, and every operation active. Report the 99th percentile, not the mean; a slider drag is judged by its worst frame. **M2 — Frame cost at 1:1 on a large file.** The same, with `Framing::view` zoomed to 1:1 on a 60 MP frame, which is the case FR-DSP-5 names and the one where the sampled region is smallest but the detail chain's kernels are widest. **M3 — Cost of the detail stage separately.** Neighbourhood operations dispatch per pass and are the only part of the chain whose cost is not one read and one write. A separable blur at a large radius is the plausible budget-breaker, not the fused pass. **The decision rule, fixed in advance.** If M1 and M2 sit inside 16 ms at the 99th percentile, **FR-DSP-2 is rewritten rather than implemented**: tiling stops being an interactive-path requirement and becomes what it actually is for this architecture — a *scheduling* concern for export and thumbnailing, which already run off the frame path. If they do not, the measurement tells us which stage to tile, which is a far better starting point than tiling everything on principle. Writing a tile scheduler that the design does not need would be the most expensive way to discover this. ARCH §5.3's tile cache keyed by `(VersionId, tile, zoom, graph_hash_prefix)` is a good design for a pipeline that needs it; the burden of proof is that this one does. --- ## 3. FR-DSP-3 — make the budget a test, not an aspiration A latency requirement that nothing asserts is a wish. The work is: **3.1** A bench in `dr-gpu` that renders a fixed chain at a fixed size and reports percentiles. Committed with its numbers, so a regression is a diff rather than a memory. **3.2** A test that *fails* when a frame exceeds the budget on the reference desktop, skipping where there is no adapter — the pattern the GPU tests already use. It should assert the 99th percentile of a hundred frames, because the failure mode being guarded against is a stutter, not an average. **3.3** The asynchronous half of the requirement: "when a full-resolution result is needed it is computed asynchronously, and the proxy result remains on screen until it is ready." Nothing does this today because nothing needs a full-resolution result on the frame path — export renders its own. This clause should be **narrowed to export and 1:1 zoom** or struck, and struck is defensible. --- ## 4. FR-DSP-4 — progressive refinement The one genuinely new piece of interactive work, and it is small because the pipeline is already resolution-parametric. During a drag, render at a fraction of the viewport and let the compositor scale; when the gesture settles, render at full viewport size. `DevelopSession` already knows when a drag is in flight — `drag-changed` exists on every slider and is what stands the Flickable down. Two things decide whether this is worth having, and M1 answers both. If a full-quality frame is already inside budget, reduced-quality rendering buys nothing and costs a visible softness during every drag — which the requirement itself warns against ("refinement is visually smooth, not a jarring swap"). **This requirement is conditional on M1 failing.** If M1 passes, FR-DSP-4 is satisfied vacuously: there is no rapid interaction the app cannot render at full quality, which is a stronger outcome than refining. --- ## 5. FR-DSP-8 — per-display colour The only display requirement needing platform work rather than pipeline work, and the only one where being wrong is a correctness defect rather than a slow frame: a second monitor with a different profile shows wrong colours, silently. **5.1 Acquisition, per display server.** X11 has `_ICC_PROFILE` atoms per output. Wayland's colour-management protocol is not universally available, and the requirement already anticipates this by demanding "a defined fallback where Wayland provides no profile" — that fallback is sRGB, stated in the About page beside the other diagnostics so a photographer can see which path they are on rather than wonder. **5.2 Reacting to a move.** The transform is selected per the display currently showing the canvas and updates when the window moves. The composed output space is already a parameter of composition (`compose_with_framing(..., output)`), so a display change is a recomposition, not a pipeline change. This is the part the existing design makes cheap. Slint turned out *not* to report window moves — there is no `on_moved` on any backend — so the window's position and scale factor are sampled twice a second and the platform re-surveyed only when they differ. See §5.4 for the display server where the position itself is unavailable. **5.3 Fractional scaling.** "Handled without resampling artefacts in the canvas" — the canvas is a wgpu texture handed to the compositor, so the requirement is that we render at the *physical* pixel size rather than the logical one and let the compositor present 1:1. Worth an explicit test, since the failure is subtle: a slightly soft canvas that looks like a bad demosaic. ### 5.4 What landed, and the one thing that did not `dr_plat::display` surveys the session's displays; `dr_ui::display_ui` decides which one is showing the canvas and keeps `DevelopSession`'s output space pointed at it. Composition was already parameterised on the space, so the pixel path changed by one argument. Two things are worth recording because they are trades rather than omissions. **A profile is matched to the nearest of four spaces, not applied.** A measured panel is none of `Srgb`, `DisplayP3`, `AdobeRgb` or `ProPhoto`, and a general ICC engine is a much larger piece of work — a CMM, rendering intents, LUT-based profiles, and a per-frame cost to argue about. The profile is reduced to its D50-adapted colorants and matched against the four; a match that is merely nearest is marked as such, and About says "nearest to Display P3" rather than "Display P3". A LUT-based profile, which is what a hardware calibrator often writes, is declined by shape and falls back to sRGB with that stated. An approximation the photographer can see beats a silent one. **On Wayland the canvas follows the first output, not the window.** A Wayland client is never told where its window is — `xdg_toplevel` carries no position, deliberately — so the "which display" question cannot be answered by geometry there. The protocol's own answer is `wp_color_management_surface_feedback_v1`, which hands a client the preferred image description for *its surface* and re-sends it on a move; it needs the application's `wl_surface`, which Slint owns and does not expose. So the profiles are read correctly for every output and the *selection* among them is right on X11 and on any single-monitor Wayland session, which is most of them. Closing the gap is a Slint surface handle, not a change to any of this. --- ## 6. Extensibility: the format already exists `FR-PLG-2` says "the node declaration is the plugin format". That is already true — it is simply resolved at build time: ``` ops/exposure.yaml ──build.rs──▶ generated Rust impl Operation ──▶ fused shader ``` A declaration names its parameters, their ranges and units, its attributes, its WGSL body and its neutral. `build.rs` compiles that into something indistinguishable from a hand-written operation. **Nothing about that requires the declaration to be present at compile time** — everything it produces is data plus a WGSL string, and the composer already assembles WGSL at run time from whatever operations are active. So class 1 is not a new mechanism. It is the existing one, loaded later. ### 6.1 What has to change **6.1.1 Descriptors become owned, not `&'static`.** `Operation::descriptor()` returns `&'static OpDescriptor` today, which is what makes a build-time node free and a run-time node impossible. This is the one invasive change in the whole plan and everything else waits behind it. `Arc` is the obvious shape; the cost is one refcount per descriptor read, on a path that reads descriptors when the panel is built rather than per frame. **6.1.2 A run-time node type.** One `DeclaredOp` implementing `Operation` from an owned declaration, replacing *generated code per node* with *one interpreter over many declarations*. The generated path can stay for the built-in chain — it costs nothing and keeps the built-ins inspectable — but the two must produce identical behaviour, which is a test: parse each built-in `ops/*.yaml` at run time and assert the composed WGSL matches the generated one byte for byte. **6.1.3 WGSL validation at load, not at dispatch.** A plugin's fragment is a string from a stranger. `compose` already builds a full shader and `naga` will reject bad source, but the failure currently surfaces as a broken render. A plugin's source must be compiled and rejected at *load*, with the error naming the plugin, because the alternative is an app that draws nothing and blames itself. **6.1.4 Order and identity.** `order:` decides chain position and `build.rs` already refuses duplicates — that guard becomes load-time. Plugin ids need a namespace (`author.name`) so two plugins cannot collide, and the sidecar stores parameters by `(op_id, param_id)`, so an id collision is a *wrong edit silently applied*, exactly the failure `MaskSource::Regions`' signature exists to prevent. ### 6.2 What this buys immediately The features enumerated as missing against Lightroom that are *pure point operations* become declarations rather than code: split toning, colour zones, selective colour, creative vignette, channel mixer variants. A photographer-author can write one without a Rust toolchain, and the existing `ops/README.md` is already its documentation. **It does not buy the neighbourhood operations** — dehaze, spot removal, liquify — because those are `DetailStage` implementations with kernels and per-render scale conversion, which FR-PLG-2a anticipates by naming "fragment nodes and pass nodes" as two templates. Pass nodes are a second phase and should not gate the first. ### 6.3 Order of work 1. Owned descriptors (6.1.1) — invasive, unblocks everything, no user-visible change 2. `DeclaredOp` + byte-identical parity test against the generated built-ins (6.1.2) 3. Load-time WGSL validation and id namespacing (6.1.3, 6.1.4) 4. A directory that is read at startup, and one shipped example that is not a built-in 5. Pass nodes (FR-PLG-2a's second template), once 1–4 are load-bearing Classes 2 and 3 — view plugins and computational plugins — are deliberately not in this plan. FR-PLG-3a's "a view plugin cannot be trusted with the UI thread" and FR-PLG-4a's capability grants are both larger design problems than class 1, and class 1 is where the requested features live. --- ## 7. What this document does not claim Traceability counts a requirement as covered when a `TRACES` tag names it. It does not check that the code under the tag does the thing — `FR-DEV-8` is currently tagged against instance-buffer plumbing that a future spot-removal operation would use, and `FR-DEV-7` against a history row for a frontend that does not exist. Both read as covered. So the 51% figure is an overstatement of unknown size, and closing FR-DSP by tagging what already works would make it a larger one. **Every requirement closed by this plan should be closed by a test that would fail if the behaviour were removed**, which is the only kind of coverage worth counting.