ARCH §5.3 still described only a tile cache nobody built; it now says
what 0.19.0 tiles and what it does not: a linear DNG past PROXY_EDGE
opens on a reduced copy, a finer render samples a full-resolution
window, and the export is cut into halo-grown tiles. display-and-
extension.md's FR-DSP-2 row said absent, outstanding.md said the halo
had nothing to read it and that such a file fell to the embedded
preview.
rawler now builds from third_party, which ARCH's stack table and §3.2,
the root Cargo.toml's comment ("two upstream crates ... for Android")
and third_party/README.md's bump procedure did not know; the README
also names each vendored crate's licence. panorama.md and the manual
say a composite this wide develops and exports.
254 lines
16 KiB
Markdown
254 lines
16 KiB
Markdown
# Finishing the display contract, and opening the pipeline
|
||
|
||
Spec for two pieces of work that turn out to be one conversation: closing **FR-DSP**, which is the
|
||
architecture's central performance claim, and reaching **FR-PLG**, which is the only requirement
|
||
family at zero.
|
||
|
||
They belong in one document because the same property decides both. The pipeline composes its work
|
||
from *declarations* — an operation says what its parameters are and contributes a WGSL fragment,
|
||
and the composer fuses the active ones into a single dispatch. That is why the display path is fast,
|
||
and it is also, already, most of a plugin format. Finishing one and opening the other are the same
|
||
seam approached from two sides.
|
||
|
||
---
|
||
|
||
## 1. What is actually true today
|
||
|
||
Stated first because both halves of this document are smaller than the requirement numbers suggest,
|
||
and the reason is that some of the work is done and untagged.
|
||
|
||
| Requirement | Reality |
|
||
|---|---|
|
||
| FR-DSP-1 proxy rendering | **Done.** The develop view renders at viewport resolution, not source. |
|
||
| FR-DSP-2 tiled computation | **Absent from the interactive path, and §2 now says it should stay that way.** Measured: the fused pass is inside the budget everywhere. See [frame-budget.md](frame-budget.md). *Since 0.19.0* the export of a linear DNG too large for one texture is drawn in halo-grown tiles, and the canvas renders such a file from a reduced copy and full-resolution windows (ARCH §5.3). |
|
||
| FR-DSP-3 interactive latency | **Measured and asserted** for the fused path — `core/dr-gpu/tests/frame_budget.rs`. Missed by one operation, clarity, for the reason recorded as TD-4. |
|
||
| FR-DSP-4 progressive refinement | **Built** (0.15.0), although §4's condition did not fire on the fused path: a half-resolution draft while a gesture moves, one sharp frame 120 ms after it stops, the histogram dimmed while it lags, and the draft faded out over 150 ms — `ui/dr-ui/src/refine.rs`. See [frame-budget.md](frame-budget.md). |
|
||
| FR-DSP-5 zoom and pan | **Done and tagged**, against tests that fail if the behaviour is removed — `core/dr-gpu/tests/zoom_resolution.rs`. `Framing::view` shrinks the sampled region while the render target keeps its size, so zooming *raises* the resolution the pipeline works at. That is FR-DSP-5's requirement, arrived at without tiles. |
|
||
| FR-DSP-6 colour management | **Done.** Output space is a parameter of composition. |
|
||
| FR-DSP-7 histogram and clipping | **Done**, GPU-side, no per-frame readback. |
|
||
| FR-DSP-8 per-display colour | **Done**, with one caveat named in §5.4. Acquisition per display server, an sRGB fallback that is visible in About, and the canvas rendered at physical pixel size. |
|
||
| FR-PLG-* | **Zero tagged.** But `ops/*.yaml` + `build.rs` is already the class-1 plugin format compiled at build time rather than loaded. |
|
||
|
||
Two of the five uncovered display requirements are therefore *measurement and tagging*, not
|
||
construction. That is worth knowing before anyone plans a quarter around them.
|
||
|
||
**Both have since been done.** [frame-budget.md](frame-budget.md) holds the measurements §2 asks
|
||
for and the reading of its decision rule; the table above is updated to match. The rest of this
|
||
document is left as it was written, because a plan that has been overtaken by its own evidence is
|
||
more useful read in order than quietly edited into agreement.
|
||
|
||
---
|
||
|
||
## 2. Measure before building tiles
|
||
|
||
**FR-DSP-2 is the one requirement in this document that may not be worth satisfying as written.**
|
||
|
||
> **Resolved.** M1–M3 were run; the numbers and the verdict are in
|
||
> [frame-budget.md](frame-budget.md). The rule below fired for *rewrite*: every point-operation
|
||
> chain is inside 16 ms at the 99th percentile at every viewport size, fit and at 1:1, the widest
|
||
> being 4.5 ms of GPU at 4K. The measurement did find a stage that misses the budget — clarity's
|
||
> 52-pixel kernel, 34 ms at 4K — and tiling makes that stage *worse*, since a tiled convolution
|
||
> reads a halo per tile. It is recorded as TD-4 with the fix its own module already names.
|
||
|
||
|
||
The requirement predates the fused-shader design. It assumes the pipeline is a chain of passes over
|
||
a large buffer, where recomputing everything on each frame would be ruinous and tiles are the way
|
||
out. What was built instead composes every active operation into **one dispatch over a
|
||
viewport-sized target** — at 2000×1300 that is 2.6 M pixels, once, for the whole chain.
|
||
|
||
So the question tiling was invented to answer may already be answered. Before any tile scheduler is
|
||
written:
|
||
|
||
**M1 — Frame cost at proxy resolution.** Time `render_detailed` at 1920×1200, 2560×1600 and
|
||
3840×2160, with a chain of one operation, five, and every operation active. Report the 99th
|
||
percentile, not the mean; a slider drag is judged by its worst frame.
|
||
|
||
**M2 — Frame cost at 1:1 on a large file.** The same, with `Framing::view` zoomed to 1:1 on a 60 MP
|
||
frame, which is the case FR-DSP-5 names and the one where the sampled region is smallest but the
|
||
detail chain's kernels are widest.
|
||
|
||
**M3 — Cost of the detail stage separately.** Neighbourhood operations dispatch per pass and are the
|
||
only part of the chain whose cost is not one read and one write. A separable blur at a large radius
|
||
is the plausible budget-breaker, not the fused pass.
|
||
|
||
**The decision rule, fixed in advance.** If M1 and M2 sit inside 16 ms at the 99th percentile,
|
||
**FR-DSP-2 is rewritten rather than implemented**: tiling stops being an interactive-path
|
||
requirement and becomes what it actually is for this architecture — a *scheduling* concern for
|
||
export and thumbnailing, which already run off the frame path. If they do not, the measurement tells
|
||
us which stage to tile, which is a far better starting point than tiling everything on principle.
|
||
|
||
Writing a tile scheduler that the design does not need would be the most expensive way to discover
|
||
this. ARCH §5.3's tile cache keyed by `(VersionId, tile, zoom, graph_hash_prefix)` is a good design
|
||
for a pipeline that needs it; the burden of proof is that this one does.
|
||
|
||
---
|
||
|
||
## 3. FR-DSP-3 — make the budget a test, not an aspiration
|
||
|
||
A latency requirement that nothing asserts is a wish. The work is:
|
||
|
||
**3.1** A bench in `dr-gpu` that renders a fixed chain at a fixed size and reports percentiles.
|
||
Committed with its numbers, so a regression is a diff rather than a memory.
|
||
|
||
**3.2** A test that *fails* when a frame exceeds the budget on the reference desktop, skipping where
|
||
there is no adapter — the pattern the GPU tests already use. It should assert the 99th percentile of
|
||
a hundred frames, because the failure mode being guarded against is a stutter, not an average.
|
||
|
||
**3.3** The asynchronous half of the requirement: "when a full-resolution result is needed it is
|
||
computed asynchronously, and the proxy result remains on screen until it is ready." Nothing does
|
||
this today because nothing needs a full-resolution result on the frame path — export renders its
|
||
own. This clause should be **narrowed to export and 1:1 zoom** or struck, and struck is defensible.
|
||
|
||
---
|
||
|
||
## 4. FR-DSP-4 — progressive refinement
|
||
|
||
The one genuinely new piece of interactive work, and it is small because the pipeline is already
|
||
resolution-parametric.
|
||
|
||
During a drag, render at a fraction of the viewport and let the compositor scale; when the gesture
|
||
settles, render at full viewport size. `DevelopSession` already knows when a drag is in flight —
|
||
`drag-changed` exists on every slider and is what stands the Flickable down.
|
||
|
||
Two things decide whether this is worth having, and M1 answers both. If a full-quality frame is
|
||
already inside budget, reduced-quality rendering buys nothing and costs a visible softness during
|
||
every drag — which the requirement itself warns against ("refinement is visually smooth, not a
|
||
jarring swap"). **This requirement is conditional on M1 failing.** If M1 passes, FR-DSP-4 is
|
||
satisfied vacuously: there is no rapid interaction the app cannot render at full quality, which is a
|
||
stronger outcome than refining.
|
||
|
||
---
|
||
|
||
## 5. FR-DSP-8 — per-display colour
|
||
|
||
The only display requirement needing platform work rather than pipeline work, and the only one where
|
||
being wrong is a correctness defect rather than a slow frame: a second monitor with a different
|
||
profile shows wrong colours, silently.
|
||
|
||
**5.1 Acquisition, per display server.** X11 has `_ICC_PROFILE` atoms per output. Wayland's
|
||
colour-management protocol is not universally available, and the requirement already anticipates
|
||
this by demanding "a defined fallback where Wayland provides no profile" — that fallback is sRGB,
|
||
stated in the About page beside the other diagnostics so a photographer can see which path they are
|
||
on rather than wonder.
|
||
|
||
**5.2 Reacting to a move.** The transform is selected per the display currently showing the canvas
|
||
and updates when the window moves. The composed output space is already a parameter of composition
|
||
(`compose_with_framing(..., output)`), so a display change is a recomposition, not a pipeline
|
||
change. This is the part the existing design makes cheap.
|
||
|
||
Slint turned out *not* to report window moves — there is no `on_moved` on any backend — so the
|
||
window's position and scale factor are sampled twice a second and the platform re-surveyed only when
|
||
they differ. See §5.4 for the display server where the position itself is unavailable.
|
||
|
||
**5.3 Fractional scaling.** "Handled without resampling artefacts in the canvas" — the canvas is a
|
||
wgpu texture handed to the compositor, so the requirement is that we render at the *physical* pixel
|
||
size rather than the logical one and let the compositor present 1:1. Worth an explicit test, since
|
||
the failure is subtle: a slightly soft canvas that looks like a bad demosaic.
|
||
|
||
### 5.4 What landed, and the one thing that did not
|
||
|
||
`dr_plat::display` surveys the session's displays; `dr_ui::display_ui` decides which one is showing
|
||
the canvas and keeps `DevelopSession`'s output space pointed at it. Composition was already
|
||
parameterised on the space, so the pixel path changed by one argument.
|
||
|
||
Two things are worth recording because they are trades rather than omissions.
|
||
|
||
**A profile is matched to the nearest of four spaces, not applied.** A measured panel is none of
|
||
`Srgb`, `DisplayP3`, `AdobeRgb` or `ProPhoto`, and a general ICC engine is a much larger piece of
|
||
work — a CMM, rendering intents, LUT-based profiles, and a per-frame cost to argue about. The
|
||
profile is reduced to its D50-adapted colorants and matched against the four; a match that is merely
|
||
nearest is marked as such, and About says "nearest to Display P3" rather than "Display P3". A
|
||
LUT-based profile, which is what a hardware calibrator often writes, is declined by shape and falls
|
||
back to sRGB with that stated. An approximation the photographer can see beats a silent one.
|
||
|
||
**On Wayland the canvas follows the first output, not the window.** A Wayland client is never told
|
||
where its window is — `xdg_toplevel` carries no position, deliberately — so the "which display"
|
||
question cannot be answered by geometry there. The protocol's own answer is
|
||
`wp_color_management_surface_feedback_v1`, which hands a client the preferred image description for
|
||
*its surface* and re-sends it on a move; it needs the application's `wl_surface`, which Slint owns
|
||
and does not expose. So the profiles are read correctly for every output and the *selection* among
|
||
them is right on X11 and on any single-monitor Wayland session, which is most of them. Closing the
|
||
gap is a Slint surface handle, not a change to any of this.
|
||
|
||
---
|
||
|
||
## 6. Extensibility: the format already exists
|
||
|
||
`FR-PLG-2` says "the node declaration is the plugin format". That is already true — it is simply
|
||
resolved at build time:
|
||
|
||
```
|
||
ops/exposure.yaml ──build.rs──▶ generated Rust impl Operation ──▶ fused shader
|
||
```
|
||
|
||
A declaration names its parameters, their ranges and units, its attributes, its WGSL body and its
|
||
neutral. `build.rs` compiles that into something indistinguishable from a hand-written operation.
|
||
**Nothing about that requires the declaration to be present at compile time** — everything it
|
||
produces is data plus a WGSL string, and the composer already assembles WGSL at run time from
|
||
whatever operations are active.
|
||
|
||
So class 1 is not a new mechanism. It is the existing one, loaded later.
|
||
|
||
### 6.1 What has to change
|
||
|
||
**6.1.1 Descriptors become owned, not `&'static`.** `Operation::descriptor()` returns
|
||
`&'static OpDescriptor` today, which is what makes a build-time node free and a run-time node
|
||
impossible. This is the one invasive change in the whole plan and everything else waits behind it.
|
||
`Arc<OpDescriptor>` is the obvious shape; the cost is one refcount per descriptor read, on a path
|
||
that reads descriptors when the panel is built rather than per frame.
|
||
|
||
**6.1.2 A run-time node type.** One `DeclaredOp` implementing `Operation` from an owned
|
||
declaration, replacing *generated code per node* with *one interpreter over many declarations*. The
|
||
generated path can stay for the built-in chain — it costs nothing and keeps the built-ins
|
||
inspectable — but the two must produce identical behaviour, which is a test: parse each built-in
|
||
`ops/*.yaml` at run time and assert the composed WGSL matches the generated one byte for byte.
|
||
|
||
**6.1.3 WGSL validation at load, not at dispatch.** A plugin's fragment is a string from a stranger.
|
||
`compose` already builds a full shader and `naga` will reject bad source, but the failure currently
|
||
surfaces as a broken render. A plugin's source must be compiled and rejected at *load*, with the
|
||
error naming the plugin, because the alternative is an app that draws nothing and blames itself.
|
||
|
||
**6.1.4 Order and identity.** `order:` decides chain position and `build.rs` already refuses
|
||
duplicates — that guard becomes load-time. Plugin ids need a namespace (`author.name`) so two
|
||
plugins cannot collide, and the sidecar stores parameters by `(op_id, param_id)`, so an id collision
|
||
is a *wrong edit silently applied*, exactly the failure `MaskSource::Regions`' signature exists to
|
||
prevent.
|
||
|
||
### 6.2 What this buys immediately
|
||
|
||
The features enumerated as missing against Lightroom that are *pure point operations* become
|
||
declarations rather than code: split toning, colour zones, selective colour, creative vignette,
|
||
channel mixer variants. A photographer-author can write one without a Rust toolchain, and the
|
||
existing `ops/README.md` is already its documentation.
|
||
|
||
**It does not buy the neighbourhood operations** — dehaze, spot removal, liquify — because those are
|
||
`DetailStage` implementations with kernels and per-render scale conversion, which FR-PLG-2a
|
||
anticipates by naming "fragment nodes and pass nodes" as two templates. Pass nodes are a second
|
||
phase and should not gate the first.
|
||
|
||
### 6.3 Order of work
|
||
|
||
1. Owned descriptors (6.1.1) — invasive, unblocks everything, no user-visible change
|
||
2. `DeclaredOp` + byte-identical parity test against the generated built-ins (6.1.2)
|
||
3. Load-time WGSL validation and id namespacing (6.1.3, 6.1.4)
|
||
4. A directory that is read at startup, and one shipped example that is not a built-in
|
||
5. Pass nodes (FR-PLG-2a's second template), once 1–4 are load-bearing
|
||
|
||
Classes 2 and 3 — view plugins and computational plugins — are deliberately not in this plan.
|
||
FR-PLG-3a's "a view plugin cannot be trusted with the UI thread" and FR-PLG-4a's capability grants
|
||
are both larger design problems than class 1, and class 1 is where the requested features live.
|
||
|
||
---
|
||
|
||
## 7. What this document does not claim
|
||
|
||
Traceability counts a requirement as covered when a `TRACES` tag names it. It does not check that
|
||
the code under the tag does the thing — `FR-DEV-8` is currently tagged against instance-buffer
|
||
plumbing that a future spot-removal operation would use, and `FR-DEV-7` against a history row for a
|
||
frontend that does not exist. Both read as covered.
|
||
|
||
So the 51% figure is an overstatement of unknown size, and closing FR-DSP by tagging what already
|
||
works would make it a larger one. **Every requirement closed by this plan should be closed by a
|
||
test that would fail if the behaviour were removed**, which is the only kind of coverage worth
|
||
counting.
|