Files
DarkRoom/docs/display-and-extension.md
T
dtourolleandClaude Opus 5 2e6825ded0 Record what the measurement found, where the next person will look for it
Three places, because the finding has three audiences.

The display spec's §1 table said FR-DSP-3 was unmeasured and FR-DSP-5 untagged.
Both are now false, and §2's decision rule has fired. The body of §2 is left as
written with the verdict quoted above it: a plan overtaken by its own evidence
reads better in order than quietly edited into agreement with the outcome.

TD-4 is the stage that misses the budget. Clarity's kernel is a fraction of the
frame, so it reaches a 52-pixel radius at 4K and costs 34 ms — seven times the
entire fused chain, for one slider. It is debt rather than a bug because the
detail stage cannot yet write a target smaller than it reads, which
`local_contrast`'s own documentation has said since it was written. The entry
says plainly that tiles are the wrong tool for it, since that is exactly the
conclusion a reader arriving from ARCH §5.3 would otherwise draw.

TD-5 is the one nobody was looking for: composing the fused shader costs
2.8–5.2 ms of CPU per frame on a full chain, on the UI thread, which at
1920x1200 is more than the dispatch it precedes. The source depends only on the
graph's structure — what `structure_hash` already identifies and what does not
move during a drag — so the fix is the cache `AdjustPass` already keeps for
compiled pipelines, one level up.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 19:06:19 +02:00

13 KiB
Raw Blame History

Finishing the display contract, and opening the pipeline

Spec for two pieces of work that turn out to be one conversation: closing FR-DSP, which is the architecture's central performance claim, and reaching FR-PLG, which is the only requirement family at zero.

They belong in one document because the same property decides both. The pipeline composes its work from declarations — an operation says what its parameters are and contributes a WGSL fragment, and the composer fuses the active ones into a single dispatch. That is why the display path is fast, and it is also, already, most of a plugin format. Finishing one and opening the other are the same seam approached from two sides.


1. What is actually true today

Stated first because both halves of this document are smaller than the requirement numbers suggest, and the reason is that some of the work is done and untagged.

Requirement Reality
FR-DSP-1 proxy rendering Done. The develop view renders at viewport resolution, not source.
FR-DSP-2 tiled computation Absent, and §2 now says it should stay that way. Measured: the fused pass is inside the budget everywhere. See frame-budget.md.
FR-DSP-3 interactive latency Measured and asserted for the fused path — core/dr-gpu/tests/frame_budget.rs. Missed by one operation, clarity, for the reason recorded as TD-4.
FR-DSP-4 progressive refinement Absent, and §4's condition did not fire. Every render is full quality and can afford to be.
FR-DSP-5 zoom and pan Done and tagged, against tests that fail if the behaviour is removed — core/dr-gpu/tests/zoom_resolution.rs. Framing::view shrinks the sampled region while the render target keeps its size, so zooming raises the resolution the pipeline works at. That is FR-DSP-5's requirement, arrived at without tiles.
FR-DSP-6 colour management Done. Output space is a parameter of composition.
FR-DSP-7 histogram and clipping Done, GPU-side, no per-frame readback.
FR-DSP-8 per-display colour Absent. One transform, not per-display.
FR-PLG-* Zero tagged. But ops/*.yaml + build.rs is already the class-1 plugin format compiled at build time rather than loaded.

Two of the five uncovered display requirements are therefore measurement and tagging, not construction. That is worth knowing before anyone plans a quarter around them.

Both have since been done. frame-budget.md holds the measurements §2 asks for and the reading of its decision rule; the table above is updated to match. The rest of this document is left as it was written, because a plan that has been overtaken by its own evidence is more useful read in order than quietly edited into agreement.


2. Measure before building tiles

FR-DSP-2 is the one requirement in this document that may not be worth satisfying as written.

Resolved. M1–M3 were run; the numbers and the verdict are in frame-budget.md. The rule below fired for rewrite: every point-operation chain is inside 16 ms at the 99th percentile at every viewport size, fit and at 1:1, the widest being 4.5 ms of GPU at 4K. The measurement did find a stage that misses the budget — clarity's 52-pixel kernel, 34 ms at 4K — and tiling makes that stage worse, since a tiled convolution reads a halo per tile. It is recorded as TD-4 with the fix its own module already names.

The requirement predates the fused-shader design. It assumes the pipeline is a chain of passes over a large buffer, where recomputing everything on each frame would be ruinous and tiles are the way out. What was built instead composes every active operation into one dispatch over a viewport-sized target — at 2000×1300 that is 2.6 M pixels, once, for the whole chain.

So the question tiling was invented to answer may already be answered. Before any tile scheduler is written:

M1 — Frame cost at proxy resolution. Time render_detailed at 1920×1200, 2560×1600 and 3840×2160, with a chain of one operation, five, and every operation active. Report the 99th percentile, not the mean; a slider drag is judged by its worst frame.

M2 — Frame cost at 1:1 on a large file. The same, with Framing::view zoomed to 1:1 on a 60 MP frame, which is the case FR-DSP-5 names and the one where the sampled region is smallest but the detail chain's kernels are widest.

M3 — Cost of the detail stage separately. Neighbourhood operations dispatch per pass and are the only part of the chain whose cost is not one read and one write. A separable blur at a large radius is the plausible budget-breaker, not the fused pass.

The decision rule, fixed in advance. If M1 and M2 sit inside 16 ms at the 99th percentile, FR-DSP-2 is rewritten rather than implemented: tiling stops being an interactive-path requirement and becomes what it actually is for this architecture — a scheduling concern for export and thumbnailing, which already run off the frame path. If they do not, the measurement tells us which stage to tile, which is a far better starting point than tiling everything on principle.

Writing a tile scheduler that the design does not need would be the most expensive way to discover this. ARCH §5.3's tile cache keyed by (VersionId, tile, zoom, graph_hash_prefix) is a good design for a pipeline that needs it; the burden of proof is that this one does.


3. FR-DSP-3 — make the budget a test, not an aspiration

A latency requirement that nothing asserts is a wish. The work is:

3.1 A bench in dr-gpu that renders a fixed chain at a fixed size and reports percentiles. Committed with its numbers, so a regression is a diff rather than a memory.

3.2 A test that fails when a frame exceeds the budget on the reference desktop, skipping where there is no adapter — the pattern the GPU tests already use. It should assert the 99th percentile of a hundred frames, because the failure mode being guarded against is a stutter, not an average.

3.3 The asynchronous half of the requirement: "when a full-resolution result is needed it is computed asynchronously, and the proxy result remains on screen until it is ready." Nothing does this today because nothing needs a full-resolution result on the frame path — export renders its own. This clause should be narrowed to export and 1:1 zoom or struck, and struck is defensible.


4. FR-DSP-4 — progressive refinement

The one genuinely new piece of interactive work, and it is small because the pipeline is already resolution-parametric.

During a drag, render at a fraction of the viewport and let the compositor scale; when the gesture settles, render at full viewport size. DevelopSession already knows when a drag is in flight — drag-changed exists on every slider and is what stands the Flickable down.

Two things decide whether this is worth having, and M1 answers both. If a full-quality frame is already inside budget, reduced-quality rendering buys nothing and costs a visible softness during every drag — which the requirement itself warns against ("refinement is visually smooth, not a jarring swap"). This requirement is conditional on M1 failing. If M1 passes, FR-DSP-4 is satisfied vacuously: there is no rapid interaction the app cannot render at full quality, which is a stronger outcome than refining.


5. FR-DSP-8 — per-display colour

The only display requirement needing platform work rather than pipeline work, and the only one where being wrong is a correctness defect rather than a slow frame: a second monitor with a different profile shows wrong colours, silently.

5.1 Acquisition, per display server. X11 has _ICC_PROFILE atoms per output. Wayland's colour-management protocol is not universally available, and the requirement already anticipates this by demanding "a defined fallback where Wayland provides no profile" — that fallback is sRGB, stated in the About page beside the other diagnostics so a photographer can see which path they are on rather than wonder.

5.2 Reacting to a move. The transform is selected per the display currently showing the canvas and updates when the window moves. Slint reports window moves; the composed output space is already a parameter of composition (compose_with_framing(..., output)), so a display change is a recomposition, not a pipeline change. This is the part the existing design makes cheap.

5.3 Fractional scaling. "Handled without resampling artefacts in the canvas" — the canvas is a wgpu texture handed to the compositor, so the requirement is that we render at the physical pixel size rather than the logical one and let the compositor present 1:1. Worth an explicit test, since the failure is subtle: a slightly soft canvas that looks like a bad demosaic.


6. Extensibility: the format already exists

FR-PLG-2 says "the node declaration is the plugin format". That is already true — it is simply resolved at build time:

ops/exposure.yaml  ──build.rs──▶  generated Rust impl Operation  ──▶  fused shader

A declaration names its parameters, their ranges and units, its attributes, its WGSL body and its neutral. build.rs compiles that into something indistinguishable from a hand-written operation. Nothing about that requires the declaration to be present at compile time — everything it produces is data plus a WGSL string, and the composer already assembles WGSL at run time from whatever operations are active.

So class 1 is not a new mechanism. It is the existing one, loaded later.

6.1 What has to change

6.1.1 Descriptors become owned, not &'static. Operation::descriptor() returns &'static OpDescriptor today, which is what makes a build-time node free and a run-time node impossible. This is the one invasive change in the whole plan and everything else waits behind it. Arc<OpDescriptor> is the obvious shape; the cost is one refcount per descriptor read, on a path that reads descriptors when the panel is built rather than per frame.

6.1.2 A run-time node type. One DeclaredOp implementing Operation from an owned declaration, replacing generated code per node with one interpreter over many declarations. The generated path can stay for the built-in chain — it costs nothing and keeps the built-ins inspectable — but the two must produce identical behaviour, which is a test: parse each built-in ops/*.yaml at run time and assert the composed WGSL matches the generated one byte for byte.

6.1.3 WGSL validation at load, not at dispatch. A plugin's fragment is a string from a stranger. compose already builds a full shader and naga will reject bad source, but the failure currently surfaces as a broken render. A plugin's source must be compiled and rejected at load, with the error naming the plugin, because the alternative is an app that draws nothing and blames itself.

6.1.4 Order and identity. order: decides chain position and build.rs already refuses duplicates — that guard becomes load-time. Plugin ids need a namespace (author.name) so two plugins cannot collide, and the sidecar stores parameters by (op_id, param_id), so an id collision is a wrong edit silently applied, exactly the failure MaskSource::Regions' signature exists to prevent.

6.2 What this buys immediately

The features enumerated as missing against Lightroom that are pure point operations become declarations rather than code: split toning, colour zones, selective colour, creative vignette, channel mixer variants. A photographer-author can write one without a Rust toolchain, and the existing ops/README.md is already its documentation.

It does not buy the neighbourhood operations — dehaze, spot removal, liquify — because those are DetailStage implementations with kernels and per-render scale conversion, which FR-PLG-2a anticipates by naming "fragment nodes and pass nodes" as two templates. Pass nodes are a second phase and should not gate the first.

6.3 Order of work

  1. Owned descriptors (6.1.1) — invasive, unblocks everything, no user-visible change
  2. DeclaredOp + byte-identical parity test against the generated built-ins (6.1.2)
  3. Load-time WGSL validation and id namespacing (6.1.3, 6.1.4)
  4. A directory that is read at startup, and one shipped example that is not a built-in
  5. Pass nodes (FR-PLG-2a's second template), once 1–4 are load-bearing

Classes 2 and 3 — view plugins and computational plugins — are deliberately not in this plan. FR-PLG-3a's "a view plugin cannot be trusted with the UI thread" and FR-PLG-4a's capability grants are both larger design problems than class 1, and class 1 is where the requested features live.


7. What this document does not claim

Traceability counts a requirement as covered when a TRACES tag names it. It does not check that the code under the tag does the thing — FR-DEV-8 is currently tagged against instance-buffer plumbing that a future spot-removal operation would use, and FR-DEV-7 against a history row for a frontend that does not exist. Both read as covered.

So the 51% figure is an overstatement of unknown size, and closing FR-DSP by tagging what already works would make it a larger one. Every requirement closed by this plan should be closed by a test that would fail if the behaviour were removed, which is the only kind of coverage worth counting.