Put the developer docs under docs/dev and index the folder for users first

docs/ had 26 developer documents flat beside the manual, and the two
audiences are very differently sized: most readers want the manual and
the gesture reference, a few want the register, the designs and the
measurements. The manual and gestures.md stay at the top; everything for
someone changing the code moves to docs/dev/, and the two documents that
name their own successors — the v0.1 milestone and the UI-refinement plan
— go to docs/dev/archive/ rather than being deleted, since both are still
cited. docs/README.md is the index, users first.

Every reference follows: code comments, Cargo manifests, the workflows,
the pre-commit hook, the bench and traceability tools (which locate the
repo root by docs/dev/requirements.md now), packaging, the Docker READMEs,
CLAUDE.md, CONTRIBUTING.md and the README. The matrix links one level
deeper and is regenerated. Links out of the moved documents into the tree
gain a level; a link checker over every Markdown file finds none broken.
This commit is contained in:
2026-09-20 16:20:15 +02:00
parent 08727cff5a
commit 6b1aac477d
137 changed files with 658 additions and 572 deletions
+86
View File
@@ -0,0 +1,86 @@
# Signing the Android build
Every build produces an APK. Which key signs it depends entirely on whether
four secrets are present:
| secret | what it is |
|---|---|
| `ANDROID_KEYSTORE_BASE64` | the keystore file, base64-encoded |
| `ANDROID_KEYSTORE_PASSWORD` | the store password |
| `ANDROID_KEY_ALIAS` | the alias of the key inside the store |
| `ANDROID_KEY_PASSWORD` | the key password |
With none of them set, `docker/android/assemble-apk.sh` generates a throwaway
debug key and signs with that. That is the right answer for a branch build or
a fork: the APK installs on a test device and nothing pretends it is a
release. With all of them set, the same script signs with the real key.
The names match JellyTau's deliberately. One convention across both Android
projects is one thing to remember instead of two.
## Making the key
Once, and then never again — keep it forever. Android identifies an app by
its signature, so an app signed with a new key is a *different* app to every
device that has the old one installed. There is no recovery from losing it
beyond telling everybody to uninstall and reinstall.
keytool -genkeypair -v \
-keystore darkroom-release.jks \
-alias darkroom \
-keyalg RSA -keysize 4096 -validity 10000 \
-dname "CN=Duncan Tourolle, O=tourolle.paris, C=FR"
`keytool` prompts for the passwords rather than taking them on the command
line, which keeps them out of shell history. Back the `.jks` up somewhere that
is not this repository and not the machine that builds it.
**The key exists, since 2026-09-11.** It was made as above, with a random
password, and the four secrets are loaded. The local copy is at
`~/.config/darkroom/signing/` on the development desktop — `darkroom-release.jks`
beside `storepass` and `keypass`, all mode 600 in a mode 700 directory. That
copy is what `package.sh` can sign with locally:
D=~/.config/darkroom/signing
KEYSTORE="$D/darkroom-release.jks" KEYSTORE_PASS="$(cat "$D/storepass")" \
KEY_ALIAS=darkroom ./docker/android/package.sh --install
Before it existed, every build — CI and local alike — was signed with a
throwaway debug key, and a debug key is exactly as durable as the cache
directory it lives in: the local one was regenerated the night the cache was
cleared, at which point no build anywhere could install over the device's copy.
Any device that received a build from before this date has to uninstall once.
## Loading the secrets
base64 -w0 darkroom-release.jks > /tmp/ks.b64
tea api --method PUT /repos/dtourolle/DarkRoom/actions/secrets/ANDROID_KEYSTORE_BASE64 \
-f data=@/tmp/ks.b64
shred -u /tmp/ks.b64
# and the three strings, read rather than typed so they miss the history
read -rs PW && tea api --method PUT \
/repos/dtourolle/DarkRoom/actions/secrets/ANDROID_KEYSTORE_PASSWORD -f data="$PW"
...and the same for `ANDROID_KEY_PASSWORD` and `ANDROID_KEY_ALIAS`. Or paste
them into Settings → Actions → Secrets in the web UI, which is less fiddly and
just as good.
## Checking which key signed a build
The packaging step prints it, and the APK carries it:
apksigner verify --print-certs darkroom.apk
A debug build says `CN=Android Debug`. Anything else is the real key.
## Signing locally
`docker/android/package.sh` takes the same environment variables, so a local
release-signed build is:
KEYSTORE=$PWD/darkroom-release.jks KEY_ALIAS=darkroom \
KEYSTORE_PASS=... KEY_PASS=... ./docker/android/package.sh
Without them it debug-signs, and keeps one debug keystore in the build cache
so repeat installs to a device do not need an uninstall first.
File diff suppressed because it is too large Load Diff
+327
View File
@@ -0,0 +1,327 @@
# DarkRoom v0.1 — Remote library viewer
**Status:** Delivered and superseded · written 2026-08-08, closed 2026-08-30
> Kept as the record of what the first milestone asked for, not as a plan.
> Everything below shipped, and the application went well past it — see
> [outstanding.md](../outstanding.md) for what is still missing at 0.9.0.
**Companion to:** [requirements.md](../requirements.md) · [architecture.md](../architecture.md)
The first buildable milestone: connect to a Nextcloud folder, index it locally, and display RAW
previews on both Linux and Android.
---
## 1. What this is
A photo *viewer*, not an editor. It connects to a Nextcloud account, lets the user pick a folder,
indexes what is there into a local catalog, and displays images by extracting their embedded JPEG
previews over byte-range requests.
**It exists to prove the architecture is sound before anything is built on it.** Four assumptions
in the design would each be expensive to discover wrong later, and this milestone tests all four
against real servers, real files, and real devices.
### 1.1 Why this shape
The alternative first milestone (D3's vertical slice: local scan → grid → open → two edits →
export) proves the pipeline but is useful to nobody and tests nothing about sync or SAF. This
milestone is comparable in size, exercises the genuinely risky parts, and produces something you
can actually point at a server and use.
### 1.2 What it deliberately excludes
No editing. No develop pipeline, no sidecars, no export, no culling mode, no ratings. Those depend
on foundations this milestone establishes; adding them before the foundations are proven risks
building on sand.
---
## 2. The four assumptions under test
Each maps to a spike in [requirements.md §9](../requirements.md).
| # | Assumption | If wrong | Spike |
|---|---|---|---|
| **A1** | Slint can composite a wgpu compute texture zero-copy, on both platforms | ARCH §6.1 fails; D1's fallbacks apply; **the framework choice is wrong** | S1, S2 |
| **A2** | Android SAF can enumerate and range-read at library scale | FR-CAT-1a and the Android performance targets fail | S10 |
| **A3** | HTTP Range extraction of embedded previews works against Nextcloud | Remote browsing on mobile data is not viable; ARCH §6.7 fails | S4 |
| **A4** | reqwest does TLS on Android without unmanageable pain | D7's escape hatch needed | S3 |
**A1 is the one that would hurt most.** It determines whether Rust + Slint was the right call at
all, so it should be proven in the first week, before catalog or sync work begins.
---
## 3. Functional scope
Requirement IDs reference [requirements.md](../requirements.md); a v0.1 suffix marks a reduced subset
of the full requirement.
### 3.1 Account and connection
**M-1 — Login.** Connect to a Nextcloud instance via Login Flow v2 (FR-NC-1): POST to
`/index.php/login/v2`, open the returned URL in the **system browser**, poll until an app password
arrives. The app never handles the user's primary password.
`User-Agent` identifies the device so the resulting app password is revocable per-device.
**M-2 — Credential storage.** Store the app password in platform secure storage (FR-NC-2) — Secret
Service on Linux, Keystore-backed on Android. Never in the catalog, never in logs.
**M-3 — Folder selection.** Browse the remote tree and choose one folder as the library root.
Recursive descent is in scope; multiple roots are not.
**M-4 — Disconnect.** Revoke via `DELETE /ocs/v2.php/core/apppassword` and clear local credentials.
### 3.2 Indexing
**M-5 — Remote listing.** `PROPFIND Depth:1` walking the chosen folder recursively, requesting
`oc:fileid`, `getetag`, `getcontentlength`, `getlastmodified`, `resourcetype`, and
`nc:has-preview`. Files whose extension matches the supported set (§3.3) are catalogued; others are
ignored.
**M-6 — Catalog.** Persist to SQLite in WAL mode. v0.1 schema is a strict subset of ARCH §6.2:
```sql
schema_version(version)
accounts(id, server_url, login_name, user_id)
folders(id, account_id, parent_id, remote_path, etag, last_listed)
images(id, account_id, file_id, remote_path, etag, size, remote_mtime,
format, captured_at, camera, lens, iso, aperture, shutter,
width, height, availability)
previews(image_id, kind, width, height, bytes, path, last_used)
```
`folders.etag` is present **from schema v1** — ARCH §6.6 requires it, and retrofitting means a
migration plus a full re-scan of every library.
`schema_version` exists from the first commit so NFR-R5's migration machinery has somewhere to
start.
**M-7 — Incremental re-listing.** On subsequent syncs use ETag pruning (FR-NC-4): `PROPFIND
Depth:0` on the root, and if its ETag is unchanged, **stop** — one request proves the whole library
is unchanged. Where changed, `Depth:1` and recurse only into folders whose ETags differ.
This is the mechanism that has to work for the library to scale, so v0.1 exercises it deliberately
rather than always re-listing.
**M-8 — Offline browsing.** The catalog is queryable with no network. Previously indexed images
appear with their metadata and any cached previews. Availability is visible per image (FR-NC-6c):
*Preview* or *Metadata only*.
### 3.3 RAW handling
**M-9 — Formats.** The FR-RAW-1 launch set: CR2, CR3, NEF, ARW, RAF, RW2, ORF, DNG. Plus JPEG, so
a mixed folder displays sensibly.
**M-10 — Range-based preview extraction.** Display without downloading whole files:
1. Range-read the first 256 KB of the file
2. Parse the container to locate the embedded JPEG preview (offset and length)
3. Range-read exactly those bytes
4. Decode the JPEG and cache it locally
Typical cost 1–3 MB against 25–100 MB for the full file. **This is what makes the app usable on
mobile data**, and it is assumption A3.
Range support is detected by issuing a `Range` request and checking for `206` versus `200` —
**never** by probing with `HEAD`, since Nextcloud does not advertise `Accept-Ranges` (ARCH §6.7).
**M-11 — Fallbacks, in order.** Where step 2 finds no usable preview:
1. Server preview via `/core/preview?fileId=…&forceIcon=false` where `nc:has-preview` is true.
**`forceIcon=false` is mandatory** — the default returns a generic mimetype icon for files the
server cannot render, which would otherwise be cached as though it were a thumbnail.
2. Full download and decode via rawler, on explicit user action only, never automatically.
3. Placeholder with an explanatory state.
**M-12 — Metadata.** Extract from the same header range already fetched in M-10: camera make and
model, lens, capture time, ISO, aperture, shutter, dimensions. No second request.
### 3.4 Display
**M-13 — Grid.** Virtualised thumbnail grid (FR-CAT-4) rendering only visible cells plus a prefetch
margin, with bounded memory independent of folder size.
**M-14 — Single image view.** Full-window display of one image, with fit and 1:1 zoom, pan, and
next/previous navigation.
**M-15 — GPU display path.** Decoded previews upload to GPU textures and composite through Slint
via `create_texture_from_hal`. **Pixels are never read back to the CPU** (ARCH §6.1).
This is assumption A1, and the reason it is in v0.1 at all: it is far cheaper to discover a
compositing problem now than after a develop pipeline is written against it.
**M-16 — Adaptive layout.** Compact and expanded layout classes (FR-UI-1) driven by window size,
not device type. Touch targets meet 44pt when touch is the active modality (FR-UI-3).
### 3.5 Platform
**M-17 — Linux.** X11 and Wayland. XDG base directories for catalog, cache, and config. Secret
Service for credentials.
**M-18 — Android.** Storage Access Framework only (FR-PLAT-AND-1) — no `MANAGE_EXTERNAL_STORAGE`,
no `READ_MEDIA_IMAGES`. v0.1 reads remote content, so SAF matters for the *cache* location and for
proving the `SourceRef` abstraction holds before local library support arrives.
Keystore-backed credential storage. Process death mid-browse restores the current folder and scroll
position (FR-PLAT-AND-3).
---
## 4. Explicitly out of scope
Listed so absence reads as a decision.
| Excluded | Arrives in |
|---|---|
| Any editing, develop pipeline, sidecars | v0.2 |
| Export | v0.2 |
| Culling mode, ratings, flags, labels | v0.3 |
| Focus peaking, raw histogram | v0.3 |
| Local (non-Nextcloud) library scanning | v0.2 |
| Upload, bi-directional sync, conflict merge | v0.4 |
| Cache rules and pinning (FR-NC-6a) | v0.4 |
| Multiple accounts or roots | later |
| Ingest from card | later |
| Search and filter beyond folder navigation | v0.3 |
**Read-only against the server.** v0.1 performs no `PUT`, `MOVE`, or `DELETE` on remote content.
This removes conflict handling and chunked upload from scope entirely, and means a bug cannot
damage the user's library.
---
## 5. Acceptance criteria
Measured against a reference Nextcloud instance holding **≥5,000 RAW files**, on the reference
desktop and **two Android devices with different GPU vendors** (Adreno and Mali).
| # | Criterion | Target |
|---|---|---|
| **AC-1** | Login through to first grid render | < 30 s for 5,000 files |
| **AC-2** | Re-sync with nothing changed | **1 HTTP request** |
| **AC-3** | Grid scroll, cached previews | 60 fps sustained, both platforms |
| **AC-4** | Bytes transferred per image, preview path | < 3 MB average |
| **AC-5** | Single-image display from cache | < 100 ms |
| **AC-6** | Offline launch → browsable grid | < 2 s desktop, < 4 s Android |
| **AC-7** | Memory, 5,000-image folder open | < 400 MB desktop, < 200 MB Android |
| **AC-8** | GPU readback of image data | **Zero occurrences** — asserted by instrumentation |
| **AC-9** | Android process death mid-browse | Folder and scroll position restored |
| **AC-10** | Catalog after forced kill during sync | Opens clean, resumes |
**AC-2 and AC-8 are the load-bearing ones.** AC-2 proves ETag pruning works, which is what lets the
library scale. AC-8 proves ARCH §6.1 holds — and it is asserted by instrumentation rather than
inspection, because a readback introduced later would otherwise pass unnoticed until it showed up
as unexplained slowness.
---
## 6. Build order
### Phase 0 — de-risk (before anything else)
Throwaway code answering the four assumptions. If A1 fails, stop and revisit D1 rather than
building on it.
1. **A1 / S1** — wgpu compute writes a texture; Slint composites UI over it; Linux. Test on
Mesa/AMD, Intel, and NVIDIA proprietary, under X11 and Wayland.
2. **A1 / S2** — the same on both Android devices.
3. **A4 / S3** — reqwest HTTPS PROPFIND from an Android device, including the
`rustls-platform-verifier` Kotlin init.
4. **A3 / S4** — range-extract an embedded JPEG from a CR3, NEF, and ARW on a real server; measure
bytes.
5. **A2 / S10** — SAF enumeration and range reads over a 10k-file tree; measure against the
NFR-P1/P3 figures.
### Phase 1 — foundations
`dr-types` (`SourceRef`, ids) · `dr-plat` traits and both implementations · `dr-catalog` (v0.1
schema, migrations from commit one) · `dr-gpu` (device, texture upload, Slint bridge).
### Phase 2 — connectivity
`dr-sync` (`RemoteBackend` trait) · `dr-sync-nextcloud` (Login Flow v2, PROPFIND, ETag pruning,
range GET) · credential storage.
### Phase 3 — imaging
`dr-decode` (container parsing, embedded-preview location, JPEG decode, metadata) · preview cache.
### Phase 4 — interface
`dr-ui` (grid, single-image view, adaptive layout, folder picker, connection flow).
### Phase 5 — hardening
Offline behaviour · process death · cancellation · error surfaces · the AC suite in CI.
**Both platforms build in CI from the first commit.** This is the point of choosing both from day
one: an Android break is caught the day it lands, not at a porting milestone.
---
## 7. Crates in play
A subset of ARCH §2. Crates not listed are not created yet.
```
core/
dr-types ✓ SourceRef, ImageId, AccountId, Validator
dr-catalog ✓ v0.1 schema subset, queries, migrations
dr-decode ✓ preview extraction + metadata; no demosaic
dr-gpu ✓ device, texture upload, Slint bridge; no pipeline
dr-sync ✓ RemoteBackend trait, ETag pruning engine
dr-sync-nextcloud ✓ the connector
dr-pipeline ✗ v0.2
dr-colour ✗ v0.2
dr-export ✗ v0.2
dr-sidecar ✗ v0.2
ui/
dr-ui ✓ grid, viewer, connection flow
dr-widgets ✗ v0.2 (no custom controls yet)
platform/
dr-plat ✓ Storage, Secrets, Lifecycle traits
dr-plat-linux ✓
dr-plat-android ✓
apps/
darkroom-desktop ✓
darkroom-android ✓
```
`dr-gpu` exists in v0.1 **only** to upload decoded JPEGs and hand textures to Slint. No compute
pipeline, no tiling, no masks. It is deliberately the thinnest thing that still proves A1.
---
## 8. Decisions this milestone informs
| Decision | What v0.1 tells us |
|---|---|
| **D12** — scope vs pace | How long a milestone of this size actually takes at the available pace. The single most useful output. |
| **D3** — first milestone | Supersedes the vertical slice, if this proves the better shape. |
| **D1** — Rust + Slint | AC-8 and A1 either confirm the framework choice or reopen it. |
| **NFR-COMPAT-1** — hardware baseline | Two real Android devices give the baseline actual numbers instead of a placeholder. |
| **D7** — network stack | Whether the reqwest Android TLS path is a half-day or a fortnight. |
---
## 9. Risks
| Risk | Likelihood | Mitigation |
|---|---|---|
| Slint `create_texture_from_hal` doesn't work as documented | Medium | Phase 0 first; D1 records fallbacks |
| SAF enumeration too slow at 10k files | Medium | S10 measures before commitment; batch and cache aggressively |
| Embedded previews too small or absent on some bodies | High | Known — Sony embeds small previews, some bodies none. M-11's fallback chain handles it; detect per camera model |
| **rawler exposes only full-resolution previews** | **Confirmed** | Measured 2026-08-09: rawler 0.7.2's CR2 decoder implements `full_image` only; `thumbnail_image`/`preview_image` are unimplemented defaults. Every rung resolves to a 5472×3648 decode at ~250 ms, 5× over NFR-P13. CR2 does carry smaller IFDs, so the fix is our own IFD walk or an upstream contribution — not a change to callers |
| **Android secret storage unimplemented** | **Confirmed** | Needs no investigation — `PlatformSecretStore` on Android is unimplemented by design, and fails loudly rather than silently no-opping (`platform/dr-plat/src/secrets.rs`). The fix is a real Keystore-over-JNI implementation (FR-PLAT-AND-1), which is `dr-plat-android` work not yet started |
| reqwest Android TLS worse than expected | Medium | D7 escape hatch: `tls_certs_only` with `webpki-roots` |
| GPU vendor divergence on Android | Medium | Two vendors in CI from the start |
| Scope creeps toward editing | **High** | §4 is explicit; v0.1 is read-only against the server |
The last one is the real risk. A viewer that works is a strong temptation to add "just one slider."
+572
View File
@@ -0,0 +1,572 @@
# UI refinement: toward a Lightroom-shaped darkroom
TRACES: FR-UI-1 | FR-UI-3 | FR-DEV-3a | FR-CAT-4
## Why
The v0.1 UI proved the architecture: capability-driven controls, a windowed
grid, a GPU canvas with no CPU round trip. What it has not yet done is *feel*
like a photo editor. The gaps are structural rather than cosmetic, and this
document names them so they can be closed independently.
Four principles guide every change below.
1. **The image is the subject.** Chrome recedes; nothing competes with the
photograph for attention or for colour.
2. **Colour is a signal, not a decoration.** The accent means *modified* or
*active*. Everywhere it currently means "heading" or "chrome", it is
spending a signal on noise.
3. **Navigation is continuous.** Moving between images should not be a change
of screen. Lightroom's filmstrip is the mechanism; modal view switching is
what it replaces.
4. **Density is earned.** A panel shows what has been touched; everything else
collapses out of the way.
## Non-goals
- No new pipeline operations, and no change to how operations reach the panel.
`AdjustPanel` must still learn its contents from the capability model and
must still name no operation (FR-DEV-3a).
- No change to the catalog schema, the sync layer, or the render path.
- No light theme. The ground stays dark (see `theme.slint` preamble).
---
## Workstream S — The style layer
Prerequisite for everything else that touches colour. Lands before B–F.
### S1 — Near-neutral palette
**Problem.** The original palette was a warm "darkroom safelight" brown —
ground `#14120F`, R twelve points above B, and the same cast through every
surface and ink. That biases the work. Simultaneous contrast pushes perception
of the image *away* from its surround, so warm chrome makes a neutral
photograph read cool; the photographer corrects toward warm to compensate and
every export drifts yellow. The file's own preamble had the right instinct —
a light UI biases judgement — and stopped one step short. Warmth biases it
too, and more quietly, because a warm cast reads as *cosy* rather than
*wrong*.
**Done.** `theme.slint` now carries near-neutral greys with a 2–3 point cool
lift (pure R=G=B reads as dead; a trace of cool reads as instrument), and the
red accent is replaced by an achromatic `active` family. `warn-ink` is the
only hue left — a caution is genuinely a different kind of thing from an
active state.
Migration shims alias `accent`/`accent-dim`/`accent-hover` onto the new
tokens, so the tree compiles while call sites migrate. **They are temporary.**
Delete them once `grep -rn 'Theme.accent' ui/` is empty.
### S2 — `style.yaml` as the token source
**Problem.** Tokens live in Slint, so tuning a palette means editing a
language file, and nothing else — docs, tooling, a future export theme — can
read them.
**Deliverable.** `ui/dr-ui/style.yaml` becomes the source of truth for
**colours and lengths**: the whole current token set, nothing more.
- `ui/dr-ui/build.rs` reads it and generates `theme.slint` at compile time.
Zero runtime cost, and a malformed file is a build error rather than a
failure in front of a photographer mid-edit.
- The generated file carries a "do not edit" banner naming its source, and
must land somewhere `.gitignore`d or clearly marked generated — a
hand-edited generated file is a bug that hides for weeks.
- **The prose survives.** The reasoning in today's `theme.slint` preamble and
its per-token comments is the most valuable thing in the file. YAML comments
carry across into the generated Slint, or the generator emits them from
structured fields. A token set with the *why* stripped out is a downgrade,
however tidy the pipeline.
- Debug-only live reload behind a feature flag: re-read the YAML at startup so
a palette can be tuned without a full rebuild. Off in release, where the
generated constants are what ship.
**Not in the YAML.** Semantic components (S3). `PanelHeading` binds colour,
size, weight and letter-spacing into one concept — that is Slint, not data.
YAML holds leaf values; components compose them.
**Dependency.** Adds a YAML parser to the UI's build-dependencies. Note that
`serde_yaml` was deprecated in 2024; prefer a maintained alternative.
### S3 — Semantic components
**Problem.** `theme.slint` says what `surface` is. It says nothing about what
a *panel heading* is — so every file re-derives one. `IMAGE`, `ADJUST`, the
six launch-screen headings, `library.slint:267` each independently spell out
colour, size, weight and letter-spacing for a single concept. That is why one
accent reached forty call sites: there was no single place to change it.
**Deliverable.** `widgets.slint` grows from three primitives into the style
layer:
- `PanelHeading` — the `IMAGE`/`ADJUST`/launch headings. One definition, one
place to decide headings are ink rather than accent.
- `Label`, `Value`, `Caption` — text roles, so `text-sm` + `ink-dim` stops
being copy-pasted. `Value` carries the modified state, since a value that
differs from its default is the one thing worth spotting at a glance.
- `Panel` — surface + rule + padding, currently rebuilt in four places.
- `Field` — the launch screen's text input with its focus border.
**The rule this establishes.** Files consume components. Raw `Theme.*` is for
*composing* a component, not for styling a call site. A new colour literal or
a bare `Theme.ink-faint` in a screen file is a signal that a component is
missing.
**Sweep.** Migrating call sites onto these components removes the `accent`
references as a side effect — the forty sites collapse into a handful of
component definitions. That is the point: fix the cause, not the symptom.
Delete the S1 shims when the grep comes back empty.
**Done when.** `grep -rn 'Theme.accent' ui/` is empty, the shims are gone, no
screen file styles a heading inline, and the UI carries no hue outside
`warn-ink`.
---
## Workstream P — Plural presentation hints
Implements ARCH §4.3a. Touches `core/dr-pipeline` and `ui/dr-ui/src/develop.rs`;
no Slint change beyond what falls out of it.
**Problem.** `Presentation.widget` is a single `WidgetKind`. An operation can
name one preferred control and nothing else, so a curve cannot say "a curve
editor is best, a parametric band control would do, and sliders are fine" and
let the frontend choose. Since layout class already drives compact-vs-expanded,
this is not hypothetical: a curve editor is comfortable in a 280px panel and
unusable in a 120px one, and today the frontend has no sanctioned way to
decide that — its only options are the named widget or nothing.
**Deliverable.**
- `Presentation.widget` becomes `widgets: &'static [WidgetKind]`, in descending
preference. The frontend takes the first it implements and can afford.
- A `WidgetDemand` describing what a widget inherently requires — candidates:
two-dimensional direct manipulation, precision pointing, a minimum count of
simultaneous values. **No pixels, no breakpoints, no DPI, no platform names.**
Those are frontend thresholds and live in `dr-ui`.
- `develop.rs` chooses per operation: walk the hint list, take the first whose
demands the current layout satisfies, else fall through to plain scalars.
The existing `ParamRow.kind` string is where exhaustive matching is currently
lost — the choice must happen in Rust against the enum, before flattening.
- The tone curve declares `[Curve]` with its real demands. Nothing else needs
to change; operations wanting sliders still say nothing at all.
**Verification.** The existing `the_curve_collapses_to_a_single_row` test in
`develop.rs` covers the happy path. Add its complement: with a layout that
cannot satisfy the curve's demands, the same capability must produce ten
addressable scalar rows and remain fully editable. That test is the contract —
it is what makes the fallback real rather than aspirational.
**Do not** let a demand grow a `min_width`. If one seems necessary, the
demand vocabulary is wrong; widen the vocabulary, not the abstraction.
---
## Workstream M — The colour mixer is unreadable
Found in use, not in review. The panel currently renders the mixer as
thirty-six anonymous sliders reading `Hue 0 / Sat 16 / Lum 0` twelve times
over, with nothing saying which band any row belongs to. Three separate
defects meet here.
### M1 — Band identity is lost (a bug, not a style issue) — **done**
`labels.rs` has no `param.mixer.*` entries, so all thirty-six keys fall
through to a default that yields the bare channel name. The core is not at
fault: it declares `param.mixer.orange.sat`, and `BANDS` carries `key:
"orange"` with `hue: 30.0`. The identity is present in the capability and
discarded at resolution.
**Deliverable.** Resolve mixer keys to their band. A row reads `Orange · Sat`,
or `Sat` under a band heading — M3 decides which.
**Landed** as a `band.*` catalogue resolving the *subject* of a faceted
parameter (M2), rather than as entries for the thirty-six `param.mixer.*`
keys. Three bands are catalogued to something other than their key: chartreuse
reads "Yellow-Green" and spring "Blue-Green", because a photographer looking
for foliage does not scan a list for "Spring".
### M2 — Bands carry their centre hue as capability data — **done**
**Deliverable.** `ParamDescriptor` gains an optional band hue in degrees, set
by `band_params!` from `BANDS`. The UI converts degrees to a swatch.
**Landed** as `descriptor::Facet` — a little wider than "a band hue", and the
width is what M3 turned out to need. A parameter may say which **aspect** it
adjusts (the channel) and which **subject** it adjusts it on (the band), with
the subject's hue attached where the subject is a colour. `ParamDescriptor` is
otherwise unchanged and `faceted()` is a const builder step, so the thirty-six
descriptors stay `static` and every other operation says nothing at all.
The hue reaches the screen as `Swatch` in `widgets.slint`, which owns the
saturation and brightness. Those two are *not* in `style.yaml`: it holds
colours and lengths, and a third section for two floats used in one component
buys less than it costs. `Theme.swatch` — the square's size — is a length and
does live there.
**Why this is not a §4.3a violation.** A band's centre hue is a *fact about
the operation* — the mixer genuinely acts on the 30° band, and that number is
what it acts on. The core says "this parameter belongs to the band centred at
30°". It does not say what colour to draw, at what saturation or lightness, or
whether to draw a swatch at all. Those conversions are presentation and live
in `dr-ui`; the swatch's saturation and lightness belong in `style.yaml`.
The line to hold: a hue in degrees is data. A hex colour in a descriptor
would be the core deciding appearance, and is forbidden.
**Swatches are the one sanctioned exception to the achromatic palette.** A
swatch is not chrome — it is data identifying which hue band a row edits,
exactly as an image is data. That is categorically different from an accent
decorating a heading, which is what the palette rule forbids. Keep them small
and let them identify, never dominate.
### M3 — Structure: three runs of twelve, not twelve of three — **done**
~~The mixer renders as twelve `Section`s — one per band, titled by band name,
carrying its swatch — each holding Hue, Sat and Lum. Collapsed by default.~~
**Superseded, twice over.** The lids came off the develop column entirely
(Workstream C's sections are now plain headings), so "collapsed by default"
had nothing left to mean. And the grouping was the wrong way round: an edit is
almost never "everything about orange", it is "the saturation of the greens",
made by comparing one channel across neighbouring bands. Twelve band sections
put the twelve rows you want to compare in twelve different places.
**Landed.** Three runs — Hue, Saturation, Luminance — of twelve rows each,
under the operation's own heading. `develop.rs` stacks the rows by aspect
(`presentation_order`) and marks the first of each run; `adjust.slint` names
the run once and draws the rest. Each row is a swatch, a track and a readout on
one line: the swatch *is* the label, which is what makes twelve rows fit where
four did, and the band name lives on as the row's accessible label so the
control is not colour-only.
The mixer's declaration order is untouched — it declares band by band, which
is the order the shader wants. Rearranging it for the panel would have been
the core laying out a screen (§4.3a); doing it in `develop.rs` is the same
frontend-side derivation that decides there are groups at all.
Nothing here is mixer-specific: any operation whose parameters carry facets
groups this way, and one that carries none is untouched.
### M4 — A single-parameter operation should not cost a heading
Vibrance and Saturation each render a section heading above one slider,
spending two lines and a visual break on one control. An operation whose
parameters number one wants to *be* a named row, not a group containing one.
**Deliverable.** The panel collapses a single-parameter operation into one
row labelled by the operation. Derived frontend-side from the parameter
count — the core says nothing about it, per §4.3a.
**Done when.** Every mixer row says which band it edits; the rows for one
channel read as a run rather than as twelve unrelated sliders; Vibrance and
Saturation are one row each; and no hex colour appears in any descriptor.
**All four are met.**
---
## Workstream V — Vertical density
The panel spends too much height on too little information. Reported from
use, and it compounds Workstream M: at thirty-six mixer rows the waste is
measured in whole screens.
**The arithmetic.** `ParamSlider` is a fixed 46px carrying an 11px label and a
3px track. The track region is `Theme.touch-target / 2` — 22px — which is a
finger-sized allowance drawn for a pointer, and the label occupies a line of
its own above it. Every section heading adds a further `Theme.gap` (12px)
spacer plus a 2px rule. Twelve mixer bands at three rows each is roughly
1650px of panel, most of it air.
**The cause is layout, not spacing tokens.** Shaving pixels uniformly would
compress the readable parts along with the waste. The row is stacked when it
could be inline: name left, value right, track beneath — which is what
Lightroom does, in about 32px.
**Deliverable.**
- Rework `ParamSlider` so label and value share one line and the track sits
under them. Target ~32px per row, down from 46px.
- The *drawn* track shrinks; the **TouchArea does not**. FR-UI-3 is about the
finger, and `Button` already establishes the pattern — draw at control
height, grow the hit target past the ink and centre it. A denser panel must
not become a less touchable one.
- Section heading spacing comes from one token, not an inline `Rectangle {
height: Theme.gap }`. A spacer rectangle written inline is how the panel
ended up with spacing nobody can adjust centrally.
- Re-check the compact layout class after the change: rows that work at 280px
may crowd at narrower widths, where the touch overhang also matters most.
**Constraint.** Density is not the goal; *legibility per pixel* is. If a row
gets shorter and harder to read, it has failed. The value readout in
particular carries the modified signal and must stay scannable.
**Done when.** A parameter row is ~32px, hit targets still meet FR-UI-3 under
touch, heading spacing is tokenised, and the panel reads as easily at the new
density as the old.
---
## Workstream A — Shared chrome primitives
**Problem.** Buttons are hand-rolled `Rectangle` + `TouchArea` pairs in
`app.slint`, `library.slint`, and `launch.slint`, at three different sizes
(64×20, 110×28, 88×28) with three near-identical hover/press treatments. Any
consistency in the chrome is currently coincidental.
**Deliverable.** A new `ui/dr-ui/ui/widgets.slint` exporting:
- `Button` — `text`, `enabled`, `primary` (bool), `clicked()`. Height meets
`Theme.touch-target` under compact layout and may be denser when expanded.
Press and hover states derive from theme tokens, not literals.
- `IconButton` — square, for toolbar affordances that carry a glyph.
- `Section` — a collapsible container: `title`, `modified` (bool),
`expanded` (in-out bool), a default child slot. Draws the disclosure
triangle and the modified dot. Workstream C consumes this.
Every existing hand-rolled button is replaced by `Button`. The visual result
should be a *narrower* range of sizes than today, not a wider one.
**Theme additions.** `theme.slint` gains what the widgets need and no more:
`radius-sm`/`radius`, a `hover` and `pressed` surface token, and a
`modified` token aliased to `accent`. Adding tokens is preferred over
literals appearing in widget bodies.
**Done when.** No `TouchArea` inside a `Rectangle` styled as a button remains
in `app.slint`, `library.slint`, or `library.slint`'s header. `cargo build`
clean, app launches, every button still fires its callback.
---
## Workstream B — Canvas presentation
**Problem.** The canvas fills its container edge to edge. The image reads as a
texture rather than a print, and there is no visual separation between the
photograph and the panel beside it.
**Deliverable.** In `app.slint`'s `canvas-area`:
- Inset the image by a margin that scales with the layout class — generous
when `expanded`, tighter when compact, never zero. The surrounding field is
`Theme.ground`.
- A subtle 1px `Theme.rule` border on the image bounds, so a dark photograph
does not bleed into the dark ground. This requires knowing the *fitted*
rectangle, not the container — if that proves awkward in Slint, a shadow or
a very slightly lighter mat behind the image is an acceptable substitute.
Pick one and say which in the summary.
- Empty and error states keep their current copy and centring.
**Constraint.** `canvas-resized` must continue to report the *drawable* pixel
size — the render target follows the image area, not the container. Getting
this wrong shows up as a soft or stretched image, so verify the reported size
changes when the margin does.
**Done when.** The image sits in a visible field with margin, the border or
mat is present, and resizing the window still produces a crisp canvas.
---
## Workstream C — Collapsible adjust sections
**Problem.** `AdjustPanel` renders every parameter of every operation, always
expanded. This is tolerable at today's operation count and unusable at
fifteen. Lightroom's right panel is a stack of collapsible modules whose
headers report whether anything inside has been touched.
**Deliverable.** Rework the `for row[i] in root.rows` body in `adjust.slint`
to group by operation and wrap each group in `Section` (Workstream A).
The hard part is that `rows` is a **flat** model with a `starts-group` flag —
Slint cannot easily nest a `for` inside a group boundary derived at runtime.
Two viable approaches; pick one and justify it briefly:
1. **Flatten the collapse.** Keep the flat `for`, add an `expanded` bool per
op-index held in the panel, and make each non-heading row `visible: false`
and zero-height when its group is collapsed. Simple, no Rust change.
2. **Nest the model.** Have Rust supply `[[ParamRow]]` — one inner model per
operation. Cleaner Slint, but changes the `ParamRow` contract and the
`develop.rs` code that builds it.
Approach 1 is likely correct for this pass; prefer it unless it proves
unworkable.
**Modified indicator.** A group is modified when any row in it has
`value != default-value`.
**This must be derived in `dr-ui`, not supplied by the core** (ARCH §4.3a).
A `group_modified` flag on a descriptor would be the core deciding the panel
has groups at all, which is a composition decision. `develop.rs` already holds
both the capabilities and the live values, so it can aggregate per operation
while flattening — that is frontend-side derivation and stays on the right
side of the line. What it must not do is ask the core for the answer.
The same reasoning condemns the existing `starts-group` flag, which is the
core telling the panel where to draw section breaks. It predates this contract;
fold it into the same pass and derive grouping from `op-index` changes instead.
**Also.** The per-group `reset` should live on the section header, alongside
the existing global `reset`.
**Invariant.** This file must still name no operation. Collapse state is keyed
by `op-index`, never by label.
**Done when.** Sections collapse and expand, collapsed state survives a slider
drag elsewhere in the panel, headers show a modified dot that appears and
disappears as values move off and back to default, and adding an operation to
the pipeline still requires no edit to `adjust.slint`.
---
## Workstream D — Grid refinement
**Problem.** Cells are boxes first and images second: `image-fit: contain` on
a square cell leaves landscape shots floating in dead space, the cell surface
contrasts with the ground so the grid reads as a rhythm of rectangles, and
there is hover state but no *selection* state.
**Deliverable.** In `library.slint`:
- Cells crop to fill (`image-fit: cover`) with `clip: true`, so the grid is a
rhythm of images. The filename caption stays.
- Cell background moves to `Theme.ground` or very near it; the frame recedes.
- A **selected** cell gets a persistent accent ring. Add
`in property <int> selected-index` to `LibraryGrid`, defaulting to -1, and
have `library_ui.rs` set it when a cell is clicked. Hover stays distinct
from selection — a dimmer treatment.
- Keyboard navigation: arrow keys move the selection, Enter opens it. This
needs a `FocusScope` over the grid and a `selection-moved(int)` callback.
**Done when.** The grid reads as images rather than boxes, the current image
is unambiguous, and arrows plus Enter navigate it without the mouse.
### Keyboard navigation — **done**
Arrows walk the grid, shift+arrow extends the selection, Home/End reach the
ends, PageUp/PageDown move by a screenful, and `Return` opens what the cursor
is on. With the judgement keys the grid already bound, a culling pass is now a
keyboard job end to end — which is the point: a cull is thousands of decisions,
and reaching for the mouse between each is the difference between an hour and
an evening.
**The cursor is a library ordinal, not a row of the loaded window** — that is
the whole design, and it is the same argument the selection already made by
keying on ids. The window is a few screenfuls around wherever the user is
looking; a cursor held as a row would stop at the window's edge or, worse, keep
counting into cells that belong to other photographs. Walking out of the window
reloads it around the new position, exactly as scrolling does.
The same fault was already live in the **anchor**, which was a window row: a
shift-click after a scroll extended from whatever image had drifted into that
row. It is now an ordinal too, and `apply_press` takes the window's offset to
map between the two. A range longer than the loaded window truncates to what is
loaded — selection is by id, and an image the catalog has not been asked for
has no id to select — which is the honest failure, not the silent one.
`select_row` is shared by the pointer and the keyboard so the two cannot drift:
"click here, shift+down twice" has to mean what "click here, shift-click there"
means. The one deliberate difference is that a plain arrow collapses the
selection onto the cursor, where a plain *click* on an already-selected cell
leaves it alone — that exception exists so a multi-image drag can start from
one of its members, and there is no drag behind a keystroke.
**Not done here:** a cursor marker distinct from the selection ring. A plain
arrow selects what it lands on, so the ring shows it; only during a shift
extension is the moving end indistinguishable from the rest of the range.
That wants a `cursor` flag on `LibraryCell` and a second ring treatment.
**Still open in D:** cells cropping to fill, the cell background receding to
the ground, and hover reading distinctly from selection.
---
## Workstream E — Chrome hierarchy
**Problem.** `StatusBar` mixes three unrelated things: navigation (`‹
Library`), identity (filename, position), and spike telemetry (fps, adapter,
backend, layout-class). The telemetry earned its place while assumption A1 was
open; it is now permanent furniture competing with the photograph.
**Deliverable.**
- The top strip carries identity and navigation only: filename, position,
and the library affordance.
- Telemetry moves behind a toggle — a keyboard shortcut (suggest `` ` ``) that
reveals a small diagnostics overlay in a corner of the canvas, carrying
backend, adapter, fps, and layout class. Default off.
- The accent stops being used for chrome. `backend` in the status bar, the
`IMAGE` and `ADJUST` panel headings, and the section headings in
`adjust.slint` all move to `Theme.ink-faint` or `ink-dim`. After this pass,
accent should appear only on: modified values, the modified dot, the curve
line, slider fill, selection, and progress.
**Done when.** A fresh launch shows no fps counter and no accent-coloured
chrome; `` ` `` toggles the diagnostics overlay; every previously-visible
diagnostic is still reachable.
---
## Workstream F — Filmstrip and unified view
**The big one.** Depends on A (for `Button`) and D (for cell treatment and
selection). Should land last.
**Problem.** Library and Develop are mutually exclusive screens
(`show-library` in `app.slint`). Every move between images is a change of
screen. Lightroom's continuity comes from the grid never fully leaving: it
collapses to a filmstrip along the bottom of the develop view, and clicking a
neighbour is navigation, not a mode change.
**Deliverable.**
- A `Filmstrip` component in a new `ui/dr-ui/ui/filmstrip.slint`, consuming
the **same** `[LibraryCell]` model and the same `selected-index` as
`LibraryGrid`. Horizontal, ~90px tall, scrolls to keep the selection
visible.
- Develop gains the filmstrip along its bottom edge, visible when a library is
open (i.e. when the current model is non-empty). Command-line file sets get
it too — they are also a list of images.
- Clicking a filmstrip cell loads that image. Arrow keys drive both the
filmstrip and the existing next/prev, which become the same action.
- `show-library` becomes a *mode* rather than a screen swap: Grid mode and
Develop mode over one shared library state, toggled by a `G`/`D` shortcut
and by the existing buttons. The `can-return-to-library` special case and
the `‹ Library` button both disappear.
**Windowing constraint.** The filmstrip and the grid must share one windowed
model, not hold two. `LibraryController` currently keys its `WINDOW` on grid
scroll position; the filmstrip's window follows the *selection* instead. This
is the genuinely hard part of the workstream — the window must move as
selection walks past its edge, and thumbnail requests must not thrash when it
does. Resolve this explicitly rather than by widening `WINDOW`.
**Done when.** Selecting an image in the grid enters develop with the
filmstrip showing neighbours; arrows walk the filmstrip and load images;
`G`/`D` toggles modes with selection preserved in both directions; a
17k-image library still holds a bounded number of live cells and does not
re-fetch thumbnails on every keystroke.
---
## Sequencing
```
A (primitives) ──┬── C (sections)
├── B (canvas) ── E (chrome)
└── D (grid) ──┐
├── F (filmstrip)
┘
```
A, B, D, and E are independent of each other once A lands; C depends on A;
F depends on A and D. B and E both touch `app.slint`, so they should not run
concurrently.
## Verification, all workstreams
- `cargo build` clean, no new Slint warnings — in particular no binding-loop
warnings, which `app.slint` already comments on at length and which can
panic at runtime.
- The app launches and reaches the grid.
- No workstream may break FR-DEV-3a: adding a pipeline operation must still
surface in the panel with no UI edit.
+113
View File
@@ -0,0 +1,113 @@
{
"_readme": [
"The committed numbers for DarkRoom's benchmark suite (docs/requirements.md §8).",
"Produced and checked by `cargo run --release -p dr-bench`; docs/benchmarks.md explains each metric.",
"",
"Two gates, and they are not the same gate. `budget` is the requirement's own threshold and never moves.",
"`recorded` is what the reference desktop last measured, and a run that drifts past `tolerance` beyond it",
"fails the build even while still inside the budget — which is how most performance rot actually arrives.",
"",
"`recorded` is null on every metric because nobody has run the suite yet. That is deliberate: writing",
"plausible-looking figures here would make every later comparison a comparison against a guess. Run",
"`cargo run --release -p dr-bench -- record --reference` on the reference desktop and commit the diff.",
"Until then the budget gate works and the regression gate says, in the report, that it cannot.",
"",
"`machine_sensitive` says whether a budget is a statement about a machine as much as about the code.",
"Those budgets are asserted only under --reference: §8 names the reference desktop, and a two-core CI",
"container cannot speak to a target written for twenty-four threads. Asserting one there would produce a",
"red gate everybody learns to ignore, which is the trap core/dr-gpu/tests/frame_budget.rs already avoids.",
"",
"Several metrics carry a requirement ID with a qualifier. Read those literally. NFR-P7's budget here is",
"checked against the encode half of an export only — no GPU render is in the figure — so it can fail the",
"requirement and cannot pass it, and no TRACES tag claims otherwise. NFR-P8 has no budget at all yet,",
"because nobody has decided how much of its 500 MB belongs to the catalog layer; this records the number",
"that decision needs."
],
"tolerance": 0.15,
"recorded_on": null,
"recorded_at_unix": null,
"fixture": null,
"metrics": {
"catalog_filtered_ms": {
"requirement": "FR-CAT-6",
"what": "Count plus first window under a rating filter, which compiles to a correlated subquery over versions.",
"unit": "ms",
"direction": "lower_is_better",
"machine_sensitive": true,
"budget": null,
"recorded": null
},
"catalog_idle_rss_mb": {
"requirement": "NFR-P8 (the catalog layer's share only — no toolkit, no adapter, no decode cache)",
"what": "Resident memory of a process that opened the 50k catalog and scrolled ten thousand rows.",
"unit": "MB",
"direction": "lower_is_better",
"machine_sensitive": false,
"budget": null,
"recorded": null
},
"catalog_open_ms": {
"requirement": "NFR-P1",
"what": "Catalog::open plus the count, first window and timeline the grid cannot paint without.",
"unit": "ms",
"direction": "lower_is_better",
"machine_sensitive": false,
"budget": 2000.0,
"recorded": null
},
"catalog_open_warm_ms": {
"requirement": "NFR-P1",
"what": "The same four calls on a second connection, with SQLite's page cache already warm.",
"unit": "ms",
"direction": "lower_is_better",
"machine_sensitive": false,
"budget": 2000.0,
"recorded": null
},
"catalog_window_p99_ms": {
"requirement": "FR-CAT-4",
"what": "One 400-row grid window at a random offset, p99 of one hundred.",
"unit": "ms",
"direction": "lower_is_better",
"machine_sensitive": true,
"budget": null,
"recorded": null
},
"export_24mp_long_edge_2048_ms": {
"requirement": "FR-EXP-3",
"what": "The web export: resample a 24 MP frame to a 2048 px long edge, sharpen, encode. p99 of five.",
"unit": "ms",
"direction": "lower_is_better",
"machine_sensitive": true,
"budget": null,
"recorded": null
},
"export_24mp_original_ms": {
"requirement": "NFR-P7 (the encode half only — the GPU render is not in this figure)",
"what": "Resample, output-sharpen and JPEG-encode a 24 MP frame at source size. p99 of five.",
"unit": "ms",
"direction": "lower_is_better",
"machine_sensitive": true,
"budget": 2000.0,
"recorded": null
},
"thumbnail_per_image_p99_ms": {
"requirement": "NFR-P3",
"what": "One thumbnail on its own lane: decode the preview, downscale, orient, encode. p99.",
"unit": "ms",
"direction": "lower_is_better",
"machine_sensitive": true,
"budget": null,
"recorded": null
},
"thumbnail_throughput_ips": {
"requirement": "NFR-P3",
"what": "Whole-sweep throughput: 1200 thumbnails through the sweep's chunk-and-lane shape, wall clock.",
"unit": "img/s",
"direction": "higher_is_better",
"machine_sensitive": true,
"budget": 100.0,
"recorded": null
}
}
}
+243
View File
@@ -0,0 +1,243 @@
# The benchmark suite
**Status:** Built, not yet recorded · 2026-08-30
**Companion to:** [requirements.md](requirements.md) §4.1 (performance targets) · §8 (verification)
**Instrument:** [`tools/bench`](../../tools/bench) — `cargo run --release -p dr-bench -- check`
**Committed numbers:** [`bench-baseline.json`](bench-baseline.json)
**GPU half:** [`core/dr-gpu/tests/frame_budget.rs`](../../core/dr-gpu/tests/frame_budget.rs) ·
[frame-budget.md](frame-budget.md)
§8 has said since it was written that performance is verified by *"an automated
benchmark suite against a synthetic 50k catalog, run per-commit … A regression
beyond stated tolerance fails the build."* Until this suite there was none. No
`benches/`, no `[[bench]]`, no criterion, no fixture — and ten performance
requirements that could therefore be neither passed nor failed, five of them
carrying a `TRACES:` tag regardless.
This file is what the suite covers, what it deliberately does not, and how to
read a failure.
---
## The state of it, first
**No numbers have been recorded yet.** Every `recorded` field in
[`bench-baseline.json`](bench-baseline.json) is `null`, on purpose: writing
plausible-looking figures into a baseline would make every later comparison a
comparison against a guess, and the first real regression would be invisible.
To record them, on the reference desktop:
```sh
cargo run --release -p dr-bench -- record --reference
```
and commit the diff. Until that happens the **budget** gate works — a catalog
that takes three seconds to open fails the build today — and the **regression**
gate reports that it has nothing to compare against, rather than pretending.
---
## What it measures
| Metric | Requirement | Gated? |
|---|---|---|
| `catalog_open_ms` | **NFR-P1**, and R2's second sentence | Yes, everywhere — budget 2000 ms |
| `catalog_open_warm_ms` | NFR-P1, page cache warm | Yes, everywhere — budget 2000 ms |
| `catalog_window_p99_ms` | FR-CAT-4 | Regression only |
| `catalog_filtered_ms` | FR-CAT-6 | Regression only |
| `thumbnail_throughput_ips` | **NFR-P3** | Budget 100 img/s, on the reference desktop |
| `thumbnail_per_image_p99_ms` | NFR-P3 | Regression only |
| `export_24mp_original_ms` | NFR-P7, **encode half only** | One-sided: can fail it, cannot pass it |
| `export_24mp_long_edge_2048_ms` | FR-EXP-3 | Regression only |
| `catalog_idle_rss_mb` | NFR-P8, **catalog layer only** | Regression only — see below |
Two of those rows carry a qualifier, and the qualifiers are the point.
### Requirements this can now pass *or* fail
**NFR-P1 — catalog open under 2 s.** The measured span is the four things the
library view cannot paint without: `Catalog::open` (which connects, migrates and
**backfills**, and the backfill is three passes over the images table on every
open), `count`, the first 400-row `window`, and the monthly `timeline`. Tagged
`TRACES: NFR-P1` in [`tools/bench/src/catalog_open.rs`](../../tools/bench/src/catalog_open.rs),
because a build that breaks it fails this gate.
**NFR-P3 — ≥ 100 images per second on the embedded preview path.** The
per-image work is exactly what `spawn_thumbnail_sweep` does — `decode_jpeg`,
`Preview::downscale_to`, `Preview::apply_orientation`, `encode_rgba`,
`ThumbStore::put` — arranged in the same shape: chunks of 96, lanes owning
disjoint slices, and the single thread that owns the store writing the finished
chunk. Tagged `TRACES: NFR-P3` in
[`tools/bench/src/thumbnails.rs`](../../tools/bench/src/thumbnails.rs).
### Requirements this can only half-answer, and is not tagged for
**NFR-P7 — 24 MP export under 2 s, full chain.** The full chain is decode,
demosaic, a full-resolution GPU render, a read-back, then resize, sharpen and
encode. Only the last three run without an adapter. So the figure here is a
**lower bound** on the requirement: exceeding 2 s in the encode alone violates
NFR-P7 no matter how fast the render is, and coming in under it proves nothing.
The budget is gated on that basis and there is no `TRACES: NFR-P7` anywhere in
`tools/bench`.
**NFR-P8 — idle memory under 500 MB.** The probe is a fresh process holding the
catalog and nothing else: no Slint, no wgpu device, no font stack, no decode
cache. Its RSS is the catalog layer's *share* of that 500 MB, not the figure the
requirement is about. It carries no budget for a reason given below.
### Requirements out of scope, listed so their absence reads as a decision
NFR-P2 (grid scroll at 60 fps), P4 (open in develop), P5 (slider to visible),
P6 (pan/zoom), P9 (UI-executor blocking), P10 (touch response), P11 (layout
transition), P12 (warm shader setup), P13 (next image in culling), P14 (focus
peaking), P15 (drawn mask stroke). Every one of them needs a frame-timing probe
inside a running Slint application, a GPU adapter, or both. None is faked here.
The GPU half of the story that *does* exist is
[frame-budget.md](frame-budget.md) and its guard test, which asserts FR-DSP-3
and skips itself where there is no adapter. `.gitea/workflows/benchmark.yml`
runs it as its own job for exactly that reason.
---
## The fixture
Fifty thousand rows over a pool of twelve real image files. Rows are cheap and
pixels are not: everything the catalog half touches is rows and is therefore
exact at full scale, and everything the pixel half touches is one file at a time
and does not care how many rows point at it. The result is ~14 MB on disk
instead of ~2 TB, and neither half is flattered by that.
| | |
|---|---|
| Rows | 50,000 images, 50,000 default versions, 400 folders, one root |
| Capture times | Twelve years from a fixed epoch, so the timeline has ~144 monthly buckets |
| Sources | 12 synthesised JPEGs at 1620 × 1080 — the size `dr-decode` records a CR2 carrying in IFD2 |
| Seed | 20260829, in [`tools/bench/src/main.rs`](../../tools/bench/src/main.rs) |
| Location | `$DR_BENCH_DIR`, else the system temporary directory |
It is reproducible from the seed, and a `stamp.json` beside it records what it
was built from — seed, row count, source count, preview size, and `dr-catalog`'s
schema version. A mismatch rebuilds rather than silently measuring a different
workload than the baseline describes.
Two honest limits on it:
- **The page cache is warm.** The fixture was written by this suite or by an
earlier run of it, so neither the catalog open nor the thumbnail sweep pays
for a cold disk. On the reference desktop's NVMe a genuinely cold read of a
14 MB catalog is tens of milliseconds; on spinning rust it is not.
- **The sources are synthetic.** A coarse gradient with a fine dither, which is
what `frame_budget.rs` synthesises for the same reason — a flat frame lets the
memory system serve every sample from one cache line and flatters a box
filter, and pure noise defeats the entropy coder in the other direction.
---
## Two gates, and how to read a failure
**Budget.** The requirement's own threshold. It does not move. Failing it means
a requirement is violated.
**Regression.** More than 15% worse than the last recorded figure *on the same
machine, against the same fixture*. Failing it means the code got slower while
still inside the requirement — which is how most performance rot actually
arrives, never over the line, always a little worse, until one day the line is
crossed by a change that was not the cause.
A metric declares whether its budget is `machine_sensitive`. Those are asserted
only under `--reference`, and reported everywhere else. §8 names *"the reference
desktop"*, not CI, and it is right to: a container with two cores cannot speak
to a throughput target written for twenty-four threads, and asserting one there
would produce exactly what `core/dr-gpu/tests/frame_budget.rs` refused to
produce — *"a red suite that everyone learns to ignore"*. Catalog open is not
machine-sensitive: 2 s against an expected figure two orders of magnitude
smaller is a threshold any machine can be held to.
Exit codes: `0` everything passed, `1` a gate failed, `2` the harness itself
could not run. Distinguished so a CI log that says "failed" does not leave
anyone guessing whether the code got slower or the fixture would not build.
**Release, always.** The workspace builds its own crates at `opt-level = 0` in
dev, and every figure here is dominated by this workspace's own code — the JPEG
decode, the box filter, the resample, the sharpen. A debug run measures rustc's
shadow. The report says which profile it was built in on its second line.
---
## NFR-P8, and the question §4.1 asks
§4.1 says NFR-P8 *"must state whether it measures RSS inclusive or exclusive of
GPU allocations, and whether it holds after SQLite's page cache warms on a 50k
catalog."* Both halves have an answer.
**On the page cache: warm.** The probe runs the count, the timeline and
twenty-five windows before it reads its counters, so SQLite's cache holds the
b-tree pages a scroll touches. That is the right side to err on — a figure taken
before the cache warms would understate a steady-state library.
**On GPU memory: RSS is exclusive of device-local allocations, and cannot be
made otherwise.** A Vulkan allocation in a device-local heap never enters the
process's address space, so nothing under `/proc/self/status` can see it. What
*does* land in RSS is the host-visible side — staging buffers, mapped upload
rings, the read-back `AdjustPass` performs on export — plus the driver's own
resident pages.
So "idle memory < 500 MB" is two questions wearing one number, and a build
holding 400 MB of RSS and 3 GB of textures would pass it.
**Recommendation: NFR-P8 should be restated as two figures** — host RSS
exclusive of device-local memory, and a separate VRAM ceiling read from the
adapter — because the second is the one that decides whether the application
survives beside a browser on an 8 GB card, and nothing in this repository
measures it today.
**And a decision is outstanding.** `catalog_idle_rss_mb` carries no budget
because nobody has decided how much of the 500 MB belongs to the catalog layer
and how much to everything above it. The suite records the number so that
decision can be taken against a measurement rather than an estimate. When it is
taken, put the figure in `budget` and the metric becomes a gate.
---
## What is not measured, and would be worth adding
- **The UI's own open.** `ui/dr-ui/src/library.rs` does not call
`Catalog::count` or `Catalog::window`; it issues its own SQL against the same
tables, with a `VISIBLE` predicate and a burst-folding clause. `dr-bench`
cannot see those without depending on `dr-ui`, which would drag Slint into a
job that has no display. **Falsifiable end:** when the grid's queries move
down into `dr-catalog` — which is where SQL over catalog tables belongs —
`catalog_open_ms` becomes the whole of the application's open and this caveat
can be deleted rather than argued about.
- **The remote sweep.** `spawn_thumbnail_sweep`'s wall clock against a real
server is latency, not CPU, and is what FR-NC-3's design is judged by. It
needs a server and belongs in a different kind of test.
- **A cold disk.** See the fixture's limits above.
- **Android.** §4.1 states a second column of targets and §8 asks for
"periodically on the named reference Android devices". Nothing here runs on a
device. Spike S10 is the piece of work that would start it.
- **Everything with a frame in it.** See the out-of-scope list above.
---
## Running it
```sh
# Measure and print. Judges nothing.
cargo run --release -p dr-bench -- run
# Measure and gate. What CI runs.
cargo run --release -p dr-bench -- check
# The same, with machine-sensitive budgets asserted too.
cargo run --release -p dr-bench -- check --reference
# Rewrite bench-baseline.json from this run, and commit the diff.
cargo run --release -p dr-bench -- record --reference
```
Useful flags: `--fixture <dir>` (or `$DR_BENCH_DIR`) to put the synthetic
catalog somewhere specific, `--lanes <n>` to pin the sweep's parallelism, and
`--thumbnails <n>` to lengthen or shorten the throughput row.
+921
View File
@@ -0,0 +1,921 @@
# DarkRoom — Catalog, library view, and background work
**Status:** Draft v0.1 · 2026-08-09
**Companion to:** [requirements.md](requirements.md), [architecture.md](architecture.md)
Specifies `dr-catalog`: the index the library view queries, how it stays current without rescanning
everything, and how thumbnails get made. [architecture.md §6.2](architecture.md) sketches the schema
in eight lines; this expands it to the point of implementability and fills the two gaps that sketch
leaves open — **incremental local scan** and **the job queue**.
Sync's remote side is already designed ([architecture.md §8](architecture.md)): ETag pruning turns a
no-op sync of 50k images into one request. Nothing equivalent existed for a local root, which is the
central problem this document solves.
---
## 1. What this must not do
Stated first because every design choice below follows from it.
| Must not | Why |
|---|---|
| Stat 50k files to open the catalog | NFR-P1: catalog open < 2 s desktop, < 4 s Android. SAF `DocumentsContract` queries are far slower than `stat` (spike S10). |
| Re-derive thumbnails for unchanged images | NFR-P3 throughput is for *new* work; redoing it on every connect makes first paint unbounded. |
| Fetch previews for remote images nobody looks at | A 50k remote library at 1–3 MB per range-extract is 50–150 GB. FR-NC-6 forbids bulk transfer by default. |
| Evaluate cache rules per grid cell | ARCH §9.5 already answers this: `tier_desired` is materialised. |
| Block the UI executor on any of it | NFR-P9, NFR-ARCH-1. |
The unifying principle: **work is proportional to what changed, or to what the user is looking at —
never to library size.**
---
## 2. Schema
Extends [architecture.md §6.2](architecture.md). Additions beyond that sketch are marked ⊕.
```sql
-- Roots -----------------------------------------------------------------
roots(
id INTEGER PRIMARY KEY,
kind TEXT, -- 'local' | 'saf' | 'remote'
grant_blob BLOB, -- SAF persisted permission; NULL on Linux
label TEXT,
last_seen INTEGER,
scan_generation INTEGER -- ⊕ bumped per completed scan; see §3.4
);
-- Folders: the unit of change detection, local and remote alike ---------
folders(
id INTEGER PRIMARY KEY,
root_id INTEGER NOT NULL REFERENCES roots(id),
parent_id INTEGER REFERENCES folders(id),
path TEXT NOT NULL,
etag TEXT, -- remote: propagating ETag (ARCH §8.4)
mtime INTEGER, -- ⊕ local: directory mtime
entry_count INTEGER, -- ⊕ local: direct children, mtime's blind spot
scanned_generation INTEGER, -- ⊕ deletion sweep; see §3.4
UNIQUE(root_id, path)
);
-- Images ----------------------------------------------------------------
images(
id INTEGER PRIMARY KEY,
root_id INTEGER NOT NULL REFERENCES roots(id),
folder_id INTEGER REFERENCES folders(id), -- ⊕ folder filter without LIKE
source_ref TEXT NOT NULL,
content_hash TEXT, -- NULL until hashed; see §3.5
format TEXT,
w INTEGER, h INTEGER,
captured_at INTEGER, -- UTC seconds; NULL if EXIF absent
captured_offset INTEGER, -- ⊕ minutes east of UTC; see §4.2
camera TEXT, lens TEXT,
iso INTEGER, aperture REAL, shutter REAL,
availability INTEGER,
file_size INTEGER, -- ⊕ cheap change signal alongside mtime
file_mtime INTEGER, -- ⊕
metadata_state INTEGER, -- ⊕ 0=none 1=stat-only 2=full EXIF; §3.5
sidecar_mtime INTEGER,
UNIQUE(root_id, source_ref)
);
-- Versions, keywords, remote, cache: per ARCH §6.2, unchanged -----------
-- Collections ⊕ ---------------------------------------------------------
collections(
id INTEGER PRIMARY KEY,
name TEXT NOT NULL,
parent_id INTEGER REFERENCES collections(id), -- collection sets
kind INTEGER NOT NULL, -- 0 = manual, 1 = smart
selector_json TEXT, -- smart only; the §5 Selector
created INTEGER
);
collection_members(
collection_id INTEGER NOT NULL REFERENCES collections(id) ON DELETE CASCADE,
image_id INTEGER NOT NULL REFERENCES images(id) ON DELETE CASCADE,
position INTEGER, -- manual ordering; NULL = by capture time
PRIMARY KEY(collection_id, image_id)
);
-- Jobs ⊕ ----------------------------------------------------------------
jobs(
id INTEGER PRIMARY KEY,
kind INTEGER NOT NULL,
subject_id INTEGER, -- image or folder, per kind
priority INTEGER NOT NULL,
state INTEGER NOT NULL, -- 0=pending 1=running 2=failed
attempts INTEGER NOT NULL DEFAULT 0,
not_before INTEGER, -- retry backoff
payload TEXT,
UNIQUE(kind, subject_id) -- coalescing; see §6.2
);
```
Indices that exist for a stated query, not speculatively:
```sql
CREATE INDEX images_captured ON images(captured_at); -- §4 timeline
CREATE INDEX images_folder ON images(folder_id);
CREATE INDEX images_hash ON images(content_hash) WHERE content_hash IS NOT NULL;
CREATE INDEX folders_parent ON folders(parent_id);
CREATE INDEX jobs_ready ON jobs(state, priority DESC, not_before);
CREATE INDEX versions_image ON versions(image_id);
CREATE INDEX members_image ON collection_members(image_id);
```
`content_hash` is indexed *partially*. It is NULL for most rows most of the time (§3.5), and a
partial index over the non-NULL subset is both smaller and what FR-CAT-9's reconnection-by-hash
and FR-CAT-11's duplicate detection actually query.
---
## 3. Incremental scan
### 3.1 The local analogue of ETag pruning
Nextcloud propagates ETags up the tree, so one request proves a whole library unchanged
([architecture.md §8.4](architecture.md)). A filesystem offers no such guarantee — a directory's
mtime changes when its *direct* entries change, and not when a grandchild does. There is no
cheap "did anything below here change" probe.
So local scan prunes at each level rather than at the root:
```
scan(folder):
(mtime, count) = stat(folder)
if (mtime, count) == stored:
# This directory's own entries are unchanged. Its files need no
# examination at all — but subdirectories may still have changed
# internally, so recurse into known children without listing.
for child in stored_children(folder):
scan(child)
else:
entries = list(folder) # the expensive call
reconcile(folder, entries) # §3.3
for child in entries.dirs: scan(child)
mark scanned(folder, current_generation)
```
Cost is **one `stat` per directory** when nothing changed, versus one per *file*. A 50k-image
library in ~2k folders costs 2k stats — a few milliseconds locally, and the difference between
meeting and missing NFR-P1 on SAF.
The recursion into unchanged directories is not redundant: it is what makes a change to one deep
file detectable at all, given no upward propagation. What it avoids is the *listing* — on SAF a
`DocumentsContract` query returning 200 rows costs far more than a metadata probe on the directory
itself.
### 3.2 Why entry-count as well as mtime
Directory mtime alone misses a real case: delete one file and create another within the same
timestamp granularity, and mtime can be unchanged while contents differ. Some filesystems and most
SAF providers report coarse timestamps, which widens the window.
Storing `(mtime, entry_count)` closes the common form of this — a paired add and remove changes
neither, but that is rarer than a bare add or remove, and both of those move the count. It is a
cheap narrowing, not a proof.
**Where correctness must not depend on it,** the user gets an explicit *Rescan folder* action
(FR-CAT-1), and reconnection matches by content hash (FR-CAT-9). Sync's remote path is unaffected —
ETags are authoritative there.
### 3.3 Reconciling a changed directory
For each entry in a listing:
| Situation | Action |
|---|---|
| Not in catalog | Insert with `metadata_state = 1`; enqueue `ExtractMetadata` |
| In catalog, `(size, mtime)` match | Nothing — the common case |
| In catalog, `(size, mtime)` differ | Re-enqueue `ExtractMetadata` and `Thumbnail`; clear `content_hash` |
| In catalog, absent from listing | Deletion candidate — §3.4 |
| Placeholder (`*.nextcloud`) | Catalogue as the image it stands for; `Availability::Offline` (ARCH §9.0) |
Sidecars are examined in the same pass: a `.drsc` whose mtime exceeds `images.sidecar_mtime` enqueues
a `ReadSidecar` job. This is how an edit made on another device — landed by the Nextcloud client,
not by us — reaches the catalog.
### 3.4 Deletion without a full sweep
A file removed outside the app appears only as an *absence*, which a pruned scan cannot see: the
folder it vanished from has a changed mtime and is listed, but a folder never visited is never
compared.
Generation counting handles this without a full pass. Each scan bumps `roots.scan_generation`, and
every folder reached — whether listed or skipped — records it. After the walk:
```sql
-- Folders never reached: their parent no longer lists them.
DELETE FROM folders
WHERE root_id = ?1 AND scanned_generation < ?2;
```
Images under a deleted folder cascade. Images missing from a *listed* folder are caught directly in
§3.3. Together these cover deletion with no additional traversal.
Deletion here means **removing the catalog row for a source proven absent**, which FR-CAT-9 sharply
distinguishes from a source merely unreachable. A root that fails to open at all — unplugged drive,
revoked SAF grant — aborts the scan and marks the root offline. It never runs the sweep, because
every folder would look unreached and the sweep would delete the entire library.
That guard is the single most dangerous line in this design, and it is stated as an invariant:
**the deletion sweep runs only after a scan that completed without a root-level access error.**
### 3.5 Metadata in two passes
Full EXIF extraction requires opening and parsing each file. At 50k images that is minutes, and it
must not stand between the user and a usable grid.
`metadata_state` records how far each image has got:
| State | Holds | Cost |
|---|---|---|
| 0 — none | Row exists, nothing read | — |
| 1 — stat-only | Name, size, mtime, format from extension | Free, from the listing |
| 2 — full | EXIF: capture time, camera, lens, exposure, dimensions | One open + parse |
The grid is usable at state 1: it can show filenames, sort by filename or file mtime, and display
placeholder cells. Promotion to state 2 runs as background jobs, prioritised by what is on screen
(§6.3), so visible images get real capture times within a frame or two of being scrolled to.
**Capture-time filtering (§4) needs state 2**, so a freshly scanned library's timeline is incomplete
until the pass finishes. The UI states this plainly — a progress affordance on the timeline, not a
silently wrong filter. Which is the FR-NC-6c principle applied to metadata rather than pixels: say
what you actually have.
`content_hash` is a *third*, still lazier tier. It requires reading the whole file, so it is computed
only when something needs it: import duplicate detection (FR-CAT-11), or reconnecting a moved source
(FR-CAT-9). Never during a routine scan.
---
## 4. The library view
### 4.1 Query model
The UI never assembles SQL. It hands the catalog a `Query` and receives a stable, windowable result:
```rust
pub struct Query {
pub filter: Selector, // §5 — same type cache rules use
pub sort: Sort,
pub descending: bool,
}
pub enum Sort {
CapturedAt,
Added,
FileName,
Rating,
/// Manual order within a collection; falls back to CapturedAt elsewhere.
CollectionPosition,
}
```
Results are fetched by window, never wholesale — FR-CAT-4 requires memory bounded independently of
catalog size:
```rust
impl Catalog {
fn count(&self, q: &Query) -> Result<usize, CatalogError>;
fn window(&self, q: &Query, range: Range<usize>) -> Result<Vec<GridRow>, CatalogError>;
}
```
`GridRow` carries exactly what a cell draws — id, thumbnail key, availability, rating, flag, capture
time — and nothing that would require a join per cell. Availability badges read `tier_desired`
directly (ARCH §9.5), so no rule evaluation happens on the render path.
A `LIMIT/OFFSET` window degrades at high offsets, since SQLite must walk the skipped rows. Scrolling
is overwhelmingly *sequential*, so the catalog keeps a keyset cursor for forward and backward paging
and falls back to OFFSET only for a scrollbar jump. Jumps are rare and single; scrolling is
continuous.
### 4.2 Time
Capture time is the spine of a photo library, and it has one persistent trap: **a photograph's
timestamp is local to where it was taken.** Store UTC alone and a shoot that ran 09:00–17:00 in
Tokyo displays as spanning two days in Paris. Store local time alone and ordering across a timezone
change is wrong.
So both: `captured_at` in UTC for ordering, `captured_offset` in minutes for display and for
day-bucketing. EXIF `OffsetTimeOriginal` supplies it where present; where absent — common on older
bodies — the offset is NULL and the catalog falls back to the library's configured display timezone,
flagged so the UI can show it as inferred.
Day, month, and year buckets are computed against **local** time. "Everything from 3 August" means
the photographer's 3 August.
The timeline affordance is a histogram of counts per bucket, which the grid uses for scrubbing:
```rust
pub enum Granularity { Year, Month, Day, Hour }
pub struct TimeBucket {
pub start: i64, // UTC seconds, bucket start
pub count: u32,
}
fn timeline(&self, q: &Query, g: Granularity) -> Result<Vec<TimeBucket>, CatalogError>;
```
This is one grouped aggregate over the `images_captured` index, not 50k rows into the UI. It is what
makes "drag across two years to find the trip" work, and it is the cheapest useful thing a library
view can offer over a flat grid.
### 4.3 Filtering interactively
FR-CAT-6 requires filter results to update interactively on 50k images. Three things make that hold:
1. **Filters compile to indexed predicates.** A `Selector` becomes a WHERE clause over indexed
columns. Keyword and collection membership become `EXISTS` subqueries against their own indices.
2. **Count and first window are one round trip.** The grid needs a row count to size its scrollbar
and the first screenful to paint; the catalog returns both together.
3. **A filter change cancels the one in flight.** Typing in a search box issues a query per
keystroke; each supersedes the last (NFR-ARCH-3). Without this the UI queues work it will discard.
---
## 5. Selectors: one type, three uses
[architecture.md §9.2](architecture.md) defines `Selector` for cache rules. The same type expresses
library filters and smart collections. This is deliberate and worth stating as a design decision,
because three near-identical predicate languages is a classic way for a catalog to rot.
| Use | Meaning |
|---|---|
| Library filter | What the grid shows now |
| Smart collection | A saved, named filter (FR-CAT-7) |
| Cache rule | What is kept locally, at which tier (FR-NC-6a) |
One consequence is directly useful: any filter the user has narrowed to can be saved as a smart
collection, and any collection can be pinned offline, with no conversion step. "Show me 5-star images
from the last 90 days" → save as a collection → pin it for the trip. Three features, one mechanism.
`Selector` moves to `dr-types` so `dr-catalog` and `dr-sync` share it without either depending on the
other. It gains variants the cache-rule sketch did not need:
```rust
pub enum Selector {
All, // ⊕ the empty filter
Collection(CollectionId),
Folder { root: RootId, path: String, recursive: bool },
DateRange(DateSelector),
Rating { min: u8 },
Label(ColourLabel),
Flag(FlagState),
Keyword(String),
Camera(String), // ⊕ FR-CAT-6 indexed field
Lens(String), // ⊕
IsoRange { min: u32, max: u32 }, // ⊕
Availability(Availability), // ⊕ "what can I edit right now"
Text(String), // ⊕ filename/keyword substring
Person { id: PersonId, include_suggested: bool }, // ⊕ §10 (FR-CULL-11)
All_(Vec<Selector>),
Any(Vec<Selector>),
Not(Box<Selector>),
}
```
`Person` carries `include_suggested` rather than defaulting silently. A saved collection built from
confirmed faces must not quietly change membership because a later indexing pass guessed at another
face; the user chose "photos of Anna", not "photos the model currently believes contain Anna". The
default is `false`, and the interactive filter offers the looser form explicitly as a way to *find*
faces to confirm.
`Availability` as a selector earns its place: on a tablet the most useful filter is often "what do I
actually have here", and it is also the natural thing to *pin* — "keep everything I've flagged that
isn't already local".
Compilation is a straightforward recursive walk producing SQL with bound parameters. **Nothing
user-supplied is ever interpolated into SQL text.** `Text` becomes a bound `LIKE` pattern with `%`,
`_`, and the escape character escaped.
---
## 6. Background work
### 6.1 Job kinds
```rust
pub enum JobKind {
ScanFolder, // §3, recursive from a folder
ExtractMetadata, // state 1 → 2
Thumbnail, // §7
ReadSidecar, // external sidecar change detected
WriteSidecar, // local edit → disk, debounced (ARCH §6.1)
ContentHash, // on demand only
FetchPreview, // remote range-extract (FR-NC-3)
FetchOriginal, // pinned or explicitly requested
DetectFaces, // §10, on the proxy tier (FR-CULL-8)
}
```
### 6.2 Coalescing is the point
`UNIQUE(kind, subject_id)` on `jobs` means enqueueing is idempotent: an image touched five times
during a scan has one thumbnail job, not five. Enqueue is
`INSERT … ON CONFLICT DO UPDATE SET priority = max(priority, excluded.priority)`, so a re-request at
higher priority promotes the existing row rather than duplicating it.
This is what makes "regenerate on update" safe to call liberally. Every code path that notices a
change can just enqueue; the table absorbs the redundancy.
### 6.3 Priority
Reuses the existing GPU scheduler classes ([architecture.md §5.3](architecture.md)) so one notion of
priority governs the whole app:
| Class | Jobs | Preempts |
|---|---|---|
| `Interactive` | Metadata and thumbnails for visible cells; preview for the open image | everything |
| `Prefetch` | The scroll margin; next image in culling | Background |
| `Background` | Bulk metadata, rule-driven fetches, hashing | — |
Visible-cell work is enqueued by the grid as it scrolls, at `Interactive`. The effect is that a
freshly scanned library fills in *where the user is looking* first, and grinds through the rest
behind them.
### 6.4 Durability and failure
Jobs live in the catalog, so they survive process death — which on Android is routine, not
exceptional (FR-PLAT-AND-3). On startup, rows in state `running` revert to `pending`: the process
that owned them is gone.
Failures increment `attempts` and set `not_before` to an exponential backoff. After a bounded retry
count the job is marked failed and attached to its image as a typed error (NFR-ARCH-4) — one
corrupt file does not stall the queue, and the user can see which files failed and why.
**A job runner never touches the UI executor**, and `Interactive` work runs on the decode pool with
the I/O pool behind it (ARCH §7.1).
---
## 7. Thumbnails
### 7.1 When
Not "on first connect" as a bulk operation. Thumbnails are generated:
- **On demand**, for cells entering the viewport plus the prefetch margin — at `Interactive`
- **On change**, when §3.3 sees a differing `(size, mtime)`
- **On rule**, for images a cache rule pins at `Preview` or above — at `Background`
- **Never** for a remote image nobody has looked at and no rule covers
For a local library this converges on "everything, eventually", because scrolling reaches everything
and the background pass has nothing else to do. For a remote library it converges on "what you
actually browsed".
**Measured on a real 17,185-RAW library, 2026-08-09:** cataloguing it by whole-file fetch would move
roughly **370 GB**; the range-extract path moves a few MB for the images actually viewed. This is
the single largest cost difference in the design, and it is why §7.1 is a list of narrow triggers
rather than "generate them all on connect".
### 7.2 How, by availability
| Availability | Source | Cost |
|---|---|---|
| `Original`, local | Embedded JPEG via `dr-decode` preview path | ~200 KB read, no demosaic |
| `Original`, no embedded preview | Full decode, downscale | Expensive — `Background` only |
| Remote | Range-extract embedded JPEG (FR-NC-3) | 1–3 MB vs 25–100 MB — **measured: 262 KB of a 21.5 MB DNG, 119 ms, 1.22% of the file** |
| Placeholder / `Offline` | None — render the offline affordance | 0 |
The remote path deliberately does **not** ask the Nextcloud client to hydrate the file. ARCH §9.0
established hydration is whole-file, so it costs ~100× what the range extract does. Hydration stays
reserved for the original tier, where the user has asked for the actual image.
Server previews (`/core/preview`) are tried only where PROPFIND reported `nc:has-preview`. ARCH §6.7
verified stock Nextcloud ships no RAW preview provider, so for RAW this is nearly always absent — it
is an opportunistic saving, never the mechanism.
### 7.3 Storage: sharded, shared, synced
Decided 2026-08-09, implemented in `dr-thumbs`. Thumbnails live in **sharded SQLite databases that
sync to Nextcloud**, so a second device gets a full grid without re-fetching a byte of RAW.
```text
thumbs/
index.sqlite fileid → shard, size accounting, client id, adoption ledger
shard-0000.sqlite ≤ 25 MB, sealed
shard-0001.sqlite ≤ 25 MB, active
```
**Why a thumbnail is worth syncing when the catalog mostly is not.** It is expensive to produce — a
range fetch plus a decode, per image — and byte-identical for every client looking at the same file.
This does not make it authoritative: losing the store costs regeneration and nothing else, so §6.12
is untouched.
**Why shards, and why small.** The 25 MB cap is about *sync granularity*, not SQLite's limits. One
growing database means every client re-downloads all of it whenever a single thumbnail is added.
With sequential fill only the newest shard is ever dirty, so an up-to-date client transfers one small
file. Sealed shards are immutable, which makes them safe to cache forever and cheap to skip.
At ~20 KB per 256px JPEG a shard holds roughly 1,200 thumbnails, so the 17,185-image reference
library lands in ~14 shards.
**Keyed on `oc:fileid`** — stable across server-side rename and move (FR-NC-5), and already in hand
from PROPFIND. Accepted consequence: shards are account-scoped, so the same photograph on two
servers is thumbnailed twice.
**Stored as JPEG, not raw pixels.** A 256×170 RGBA buffer is ~174 KB against ~15 KB encoded. Since
shards sync, that 11× is transfer cost paid by every client, not just disk.
Three invariants, each tested:
| Invariant | Why it matters |
|---|---|
| A sealed shard never reopens | Reopening one forces every client that holds it to re-download |
| Re-storing an existing id updates in place, never migrates | Migrating would rewrite a sealed shard |
| Merging another client's shard is insert-only and idempotent | Both copies derive from the same bytes by the same code, so neither is better; preferring ours avoids dirtying a shard others have synced |
**The transfer**, in `dr-ui`'s `derived_sync`, exchanges shards with `.darkroom-derived/` under the
library root. `ThumbStore::shards()` reports which are sealed, so an up-to-date client's whole pass
is one listing plus whichever shard is still open.
**Why a remote name carries a client id.** Corrected 2026-08-16. Shard ids are *per store* — every
client fills its own numbering from 0 — so the flat `shard-NNNN.sqlite` namespace the transfer first
used had two clients writing one name. Two failures followed from it, and both were live: the second
client's upload **overwrote** content the first still believed was published, and no client could
distinguish a peer's shard 3 from its own, so the only safe reading of "I already hold 3" was to skip
it. Between them, two populated clients exchanged almost nothing — only shards numbered above the
other's highest. A fresh device worked, which is why it went unnoticed: with no local shards there is
nothing to collide with.
The name is now `shard-<client>-NNNN.sqlite`, where `<client>` is minted per store in `index.sqlite`
beside the numbering it qualifies — a store deleted and rebuilt restarts at shard 0 and must not
claim its predecessor's names. Since a client's own ids no longer say anything about what it has
taken from others, `index.sqlite` also keeps an **adoption ledger** of merged remote names and the
size each had. Size, not a flag: a peer's sealed shard never returns, but its open one grows, and
re-merging the grown copy is how the thumbnails it gained arrive.
Flat names left on servers by earlier builds are still read — they report no owner, so each client
adopts them once — and nothing is written under that form again. A flat name whose id and byte size
match a local shard is that client's own earlier upload by the same identity argument used for
sealed shards, so the rename does not cost every client a re-download of its whole store. Older
builds ignore the new names, so they stop receiving shards until updated; nothing is lost, since
their own uploads are still adopted.
Two size classes remain planned — grid (256px) and filmstrip/loupe (1024px). Only the grid class is
implemented. The cache is LRU-capped per NFR-RES-4, and thumbnails evict before proxies and long
after sidecars, which never evict at all (FR-NC-6b).
---
## 7a. Editing collections
Decided 2026-08-09. The schema for collections landed with §2 and the cross-device merge rules with
§8; this is the layer between them — the operations a user actually performs, in
`dr_catalog::collections`.
### 7a.1 Hierarchy and membership are independent
Two structures, deliberately not entangled:
| | Mechanism | Meaning |
|---|---|---|
| Hierarchy | `collections.parent_id` | A collection inside a collection (Lightroom's "collection set"). A parent is an ordinary collection, not a separate kind, so a set can hold images of its own |
| Membership | `collection_members` | An image is in as many collections as the user likes. Nothing moves on disk; no collection owns an image |
**Adding an image to a child does not write a row for the parent.** A parent's contents are the union
of its own members and its descendants', computed on read. Materialising it instead would make one
add touch every ancestor, and a reparent rewrite membership — both of which §8's row-level merge
would then have to reconcile. The read path pays a bounded tree walk instead, which at sidebar scale
is nothing.
The consequence the UI depends on: dragging images onto a collection is **additive**. It does not
remove them from anywhere, which is why the gesture's default action is `copy` and not `move`.
### 7a.2 Rules that exist to prevent silent damage
| Rule | Why |
|---|---|
| Every mutation bumps `revision` | §8.4 resolves conflicts by revision. An edit that updates `modified` alone is invisible to the merge, so the *other* device silently wins and the user's work vanishes |
| A no-op add does **not** bump it | Otherwise an idle device that re-dropped the same images outranks one that did real work |
| Deleting a parent **promotes** its children | The schema's `ON DELETE CASCADE` would take the whole subtree. Losing a nested collection because its container was tidied away is not recoverable |
| Deletion leaves a tombstone | Without it, merging with a device that still holds the collection resurrects it (§8.4) |
| Cycles are refused at the write | Both kinds — parenting under a descendant, and a smart collection whose selector reaches itself. A cycle is unbounded recursion in the tree walk, so it must not be *representable*, not merely handled when drawn |
| Tree walks are depth-guarded anyway | A merge can deliver a row this device never validated. The read path must terminate, so it truncates and logs rather than hanging the UI thread |
| A drop onto a smart collection is refused | Its membership *is* its selector; member rows would be a second source of truth that nothing reads |
| Deep counts are `count(DISTINCT image_id)` | An image in both a parent and a child is one photograph. A count that disagrees with the number of cells drawn makes both untrustworthy |
`collections.uuid` is generated from the OS CSPRNG. A collision fuses two unrelated collections at
the next merge, so the fallback path (used only if `/dev/urandom` cannot be read) logs loudly rather
than degrading identity quality in silence.
### 7a.3 Drag and drop is Slint's, not ours
The first implementation hand-rolled the gesture on `TouchArea` — tracking the press, measuring
travel to distinguish a click from a drag, and deciding the drop target from the last row hovered.
**It did not work**, for a reason worth recording: an interactive `Flickable` claims any drag
beginning inside it for scrolling and *cancels* the child `TouchArea`'s press, so the gesture could
never leave the grid. It also had a correctness hole — a tree rebuilt mid-drag could redirect the
drop, since a captured pointer is invisible to every other element.
Slint 1.17's `DragArea`/`DropArea` own all of it: capture, the click-versus-drag threshold,
arbitration against the `Flickable`, the image under the cursor, and hit-testing the release. What
remains in `collections_ui` is only what Slint cannot know — the payload (which images, read from the
selection when the drag starts) and the **spring**: a dwell timer that opens a collapsed collection
so a nested child can be reached mid-drag, and closes again whatever the drag merely passed over.
One hazard survives the change and is easy to reintroduce. Every consequence of a drop — rebuilding
the tree, refreshing the badges, rereading the grid — *replaces a Slint model*, and doing that inside
the `dropped` handler destroys the elements Slint is still using to deliver that event. So the drop
records its target and `drag-finished` acts on it. This is the same hazard `sync_rows` in `lib.rs`
documents for the adjust panel, and it presents as a control that works once and then goes dead.
---
## 8. Syncing the catalog file
Decided 2026-08-09. **This qualifies [architecture.md §6.12](architecture.md)** — the catalog
remains a rebuildable index, but the file itself now travels to Nextcloud. The qualification is
worth stating precisely, because the sidecar-authoritative model is load-bearing and this is the
one place it bends.
### 8.1 Why collections forced this
Every other thing the catalog holds has authoritative backing outside it. Ratings, labels,
keywords, and edit graphs live in sidecars next to the images, so a rebuild recovers them.
**Collections do not.** A manual collection is a set of images the user assembled by hand; nothing
in the filesystem records it. Losing the catalog loses them, and no rescan brings them back.
So collections need to be durable across devices somehow. Syncing the catalog file is the chosen
mechanism.
### 8.2 What the file sync does and does not carry
Only **collections and their membership** merge. The rest of a catalog describes *local* state —
folder mtimes, cache file paths, job rows, `tier_actual` — and importing another device's version
of those would be actively wrong. The downloaded remote is read for its collections and discarded.
This is what keeps §6.12 substantially intact: nothing here makes the local database authoritative
for anything a rebuild could not recover. The catalog is still deletable. What syncs is one table
pair that had no other home.
### 8.3 Two hazards the implementation must handle
**A WAL database is not one file.** Committed transactions can sit in `catalog.sqlite-wal` with the
main file lagging, so copying `catalog.sqlite` alone uploads a torn snapshot — internally consistent
as of some older point, silently missing everything since. Upload therefore runs a `TRUNCATE`
checkpoint and then SQLite's backup API, which serialises against concurrent writers rather than
racing them. It never copies the live file.
**Integer primary keys are not identities.** Two devices each allocate `collections.id = 1` for
different collections, so a row-level merge keyed on the integer id would collide them. Collections
therefore carry a **UUID**, and membership maps across devices by **image content hash**. The
integer ids stay local and are never compared across catalogs.
### 8.4 Merge rules
| Concern | Rule | Why |
|---|---|---|
| Which collection wins | Higher `revision` — a counter bumped per local edit. `modified` only breaks an exact tie | A device with a skewed clock cannot silently overwrite real work. The same reason FR-NC-9 avoids mtime for sidecars |
| Membership | **Set union**, not last-writer-wins | Two devices adding different images to one collection keep both. The exception — a removal racing an addition — resolves toward the addition, which is recoverable by removing it again. A lost addition is not |
| Deletion | Tombstone (`deleted = 1`) carrying a revision | Without it, merging against a device that still holds the collection resurrects it. With a revision, deletion competes on equal footing with a rename |
| An image the remote has and we do not | Skip the membership row | It joins on a later merge, once a scan has catalogued the file. Not an error |
| A remote from a newer schema | Decline before attaching | Attempting it would fail mid-transaction rather than declining cleanly |
Merging is idempotent: running it twice reports no changes the second time. That property is tested,
because a merge that oscillates would upload on every sync forever.
### 8.5 What was rejected
**Replace-if-newer.** The literal reading of "sync the file and take the newer one". Rejected
because it is not a merge: whichever device syncs second loses every collection the first did not
have. Binary SQLite files do not merge, so "newer wins" means "older is destroyed".
**A `collections.drsc` sidecar at the library root.** The alternative that would have kept §6.12
untouched, merging as text the way edit sidecars do. Viable, and cheaper in machinery, but it means
a second serialisation format and a second merge implementation for the same data. Recorded here
because if the SQLite path proves troublesome, this is the fallback with a known shape.
---
## 9. What this document does not settle
- **FTS.** `Selector::Text` is a `LIKE` scan over filename and keywords. Adequate at 50k; if free
text over description and title becomes a real workflow, an FTS5 table is the answer, and it is
additive.
- **Smart collection materialisation.** Currently evaluated on read. If a smart collection's
membership needs to be *stable* — for manual ordering, or for a pinned set that must not shift
under the user — it needs materialising with an invalidation rule. Deferred until there is a
concrete need.
- **Multi-root capture-time collisions.** FR-CAT-11 detects duplicates on import; the same image
catalogued under two roots is a related but distinct case, not yet specified.
- **Timeline granularity selection.** Which bucket size the UI picks for a given zoom is a UI
concern, but the catalog should probably suggest one from the query's date span rather than have
the UI guess.
---
## 10. People and faces
Specified by FR-CULL-8 … FR-CULL-12, NFR-SEC-5, [architecture.md §6.4](architecture.md). Gated on
spike S14 and decision D13 — the runtime and the model licences are unresolved, so this is the shape
of the subsystem, not a build order.
[faces.md](faces.md) names the models this shape is filled in with, and adds one column to §10.1's
`faces` table (`crop_px`) that the calibration in its §8 depends on.
### 10.1 Schema (a v5 migration)
```sql
CREATE TABLE people (
id INTEGER PRIMARY KEY,
uuid TEXT NOT NULL UNIQUE, -- merge identity, not the name (ARCH §6.3)
name TEXT NOT NULL,
-- Tombstone-by-redirect. A merged person must outlive its merge, or a
-- device that still has it resurrects it — same hazard collections have.
merged_into INTEGER REFERENCES people(id) ON DELETE SET NULL,
created INTEGER NOT NULL,
revision INTEGER NOT NULL DEFAULT 1,
modified INTEGER NOT NULL
);
CREATE TABLE faces (
id INTEGER PRIMARY KEY,
image_id INTEGER NOT NULL REFERENCES images(id) ON DELETE CASCADE,
-- Normalised to the image's long edge, so a face survives the proxy it was
-- found on being regenerated at another resolution.
x REAL NOT NULL, y REAL NOT NULL, w REAL NOT NULL, h REAL NOT NULL,
landmarks BLOB, -- 5 × (x, y) f32, the alignment input
detector_confidence REAL NOT NULL,
embedding BLOB NOT NULL, -- 512 × f16, the raw model output; re-normalised on load
-- Length of that vector: the model's own reading of how recognisable the
-- crop was, and the gate on whether this face may be compared *against*
-- (faces.md §6, §9). NULL for a face stored as a unit vector before it
-- was kept.
quality REAL,
-- What the eyes are doing (FR-CULL-8a, faces.md §17): per eye P(open),
-- the source pixels across its box and the sharpness of the patch the
-- classifier saw; and P(sunglasses). All seven or none; NULL is "never
-- read", which every filter treats as unknown rather than as closed.
-- The verdict -- open, closed, sunglasses, unclear -- is a rule in
-- dr_face::eyes, not a column.
eye_right REAL,
eye_right_px REAL,
eye_right_sharp REAL,
eye_left REAL,
eye_left_px REAL,
eye_left_sharp REAL,
sunglasses REAL,
-- The 106 dense landmarks the eyes were read from, packed as 16-bit
-- fixed point over the frame: 424 bytes (schema V18). Kept so the next
-- per-face pass runs from the catalog rather than from the original.
landmarks_dense BLOB,
-- Which model produced this. An embedding is only comparable to others
-- from the same model; mixing them silently yields nonsense similarities.
model_id TEXT NOT NULL,
detected_at INTEGER NOT NULL
);
CREATE INDEX faces_image ON faces(image_id);
-- Covers the eyes-open filter's subquery. Without it every check read the
-- whole face row -- the eye columns sit after the blobs -- and one count
-- took 24 s on the reference library (schema V17).
CREATE INDEX faces_eyes ON faces(image_id, eye_right, eye_right_px, eye_right_sharp,
eye_left, eye_left_px, eye_left_sharp, sunglasses);
CREATE TABLE face_person (
face_id INTEGER PRIMARY KEY REFERENCES faces(id) ON DELETE CASCADE,
person_id INTEGER NOT NULL REFERENCES people(id) ON DELETE CASCADE,
-- Calibrated P(this face is this person), never a raw cosine (FR-CULL-9).
probability REAL NOT NULL,
-- The user said so. Never overwritten by a later inference pass.
confirmed INTEGER NOT NULL DEFAULT 0
);
CREATE INDEX face_person_person ON face_person(person_id, confirmed);
```
Three things in that schema are load-bearing:
**`model_id` on every face.** Embeddings from different models are not comparable — this is the one
mistake that produces plausible-looking garbage rather than an error. Storing the model with the
embedding means a model change is detectable and re-indexable, instead of quietly poisoning every
similarity in the library.
**Normalised bounding boxes.** Detection runs on whichever proxy exists (FR-CULL-8). Storing pixel
coordinates would bind a face to a resolution that the cache is entitled to evict and regenerate
differently.
**`confirmed` as a column, not a probability of 1.0.** A confirmation is a different kind of fact
from a confident guess, and collapsing them loses the ability to recompute suggestions without
touching user data.
### 10.2 Why clustering is not a job kind
Detection is per-image and parallel, so it is a job (`DetectFaces`, coalesced per image like any
other). Clustering is a *whole-library* operation over the embeddings detection produced — it has no
natural `subject_id`, and running it per-image would rebuild the world on every photograph.
It therefore runs as a debounced library-level pass, triggered when detection has been idle and the
face count has moved materially since the last clustering. The same reasoning as sidecar writes: the
work is cheap to defer, expensive to repeat, and nobody is waiting on it.
### 10.3 The calibration lives with the library
FR-CULL-9 requires similarity to be a calibrated probability, fitted from this library's own faces.
That fit is a property of the catalog and its model, so it is stored alongside — a small table
holding the fit parameters, its validity flag, and a hash of the face set it was derived from, so a
materially changed library recomputes rather than trusting a stale fit.
When the fit is not valid — a library with too few faces to have positive pairs — the UI says the
confidence is unavailable. It does not fall back to an untuned default dressed up as a measurement.
### 10.4 What this does not settle
- **Which model, and which runtime.** D13. Everything above holds regardless of the answer, which is
why it is specified in terms of "a 512-d embedding from a stated model" rather than a named one.
- **The clustering algorithm.** Density-based over the calibrated distance is the obvious starting
point, but the parameters are an S14 question, not a design-time one.
- **Whether embeddings sync.** NFR-SEC-5 permits it, opt-in. The shard mechanism in §7.3 is the
obvious carrier if they do, but nothing here depends on that decision.
- **Faces in trashed images.** FR-CAT-15's trash moves files; whether their faces stay indexed and
keep contributing to clusters is unspecified. Probably they should be excluded from suggestions but
not deleted, so a restore does not re-index.
---
## 10a. Bursts and near-duplicates
Specified by FR-CULL-5, implemented in `dr_catalog::bursts` (a v11 migration) with the pass that
feeds it in `dr_ui::bursts`.
A burst is a run of frames that are **adjacent in time and look like the frame before them**. Both
halves are load-bearing. Time alone groups a whole wedding ceremony, because a photographer working
steadily never leaves the gap that would end the run. Similarity alone groups a studio setup shot
across two days, which is a project rather than a moment. The bounds are two seconds and eight bits
of a 64-bit difference hash, and the reasoning for each figure is in the module.
**Two seconds, for a burst that fires ten frames in one.** `images.captured_at` is whole seconds:
EXIF's `DateTimeOriginal` has no sub-second field, and `SubSecTimeOriginal` is optional and widely
omitted. Ten frames of a burst therefore arrive sharing a timestamp, and any threshold finer than a
second is a threshold on information the catalog does not have. Where the pace really is faster than
two seconds, the similarity bound is what separates the frames.
**The signal is a perceptual hash of the thumbnail, not of the original.** `images.perceptual_hash`
is filled from the 256px thumbnails §7 already stores — vastly more resolution than a 9×8 reduction
uses — so a library that has been browsed has already paid for its signatures and no RAW is decoded
for this. The consequence is stated rather than hidden: an image with no thumbnail gets no
signature, and a frame with no signature never joins a burst. It is picked up by the next pass.
**It is a pass, not a job kind**, for exactly the reason §10.2 gives for face clustering: a burst is
a property of a *run* of frames and has no natural `subject_id`, so a per-image job would rebuild
the world once per photograph. It runs when the thumbnail sweep finishes, which is the first moment
the signatures can all be computed.
**A newly found burst arrives open.** The pass marks frames; it never takes them off the screen.
Collapsing on discovery would be tidier and would also mean a background pass removing photographs
from under someone part way through a cull. Folding a burst up is the user's act, it is remembered
(`burst_expanded`), and a burst that is already known keeps whatever state it is in — so the pass
that follows the next import does not spring open a morning's work.
**Nothing here ranks a frame.** The representative of a collapsed burst is its *earliest* frame,
which is a fact about the clock rather than a judgement about the photograph. FR-CULL-5 names the
failure this avoids — rejecting the only frame of an important moment because someone blinked — and
the only judgement in the subsystem is the user's own choice of representative, which lives in its
own table (`burst_pick`) so that rebuilding the grouping cannot erase it. Same argument as
`people.ignored` in §10.
**What the collapse costs the grid.** Which rows a collapsed burst hides has to be decided by the
query rather than by the cells, because the grid is a window (`LIMIT n OFFSET k`) and the frames it
hides are mostly not loaded. So the predicate joins `VISIBLE` in every query that lists or counts
cells, under the same discipline: present in four places of five, the header's count, the
scrollbar, the shift-click range and the scrub's ordinal stop describing the same list.
**What this does not settle.** Bursts are local: the tables ride along in the uploaded catalog
snapshot and nothing on the far side reads them, so a second device rebuilds its own grouping from
its own signatures. Making `burst_pick` cross-device is a merge question of the same shape as §8.4's
and is not answered here.
---
## 11. Requirements touched
| ID | How this document addresses it |
|---|---|
| FR-CAT-1 | §3 incremental scan, cancellable and resumable via §6 jobs |
| FR-CAT-3 | §7 thumbnail pyramid, two size classes, embedded-preview fast path |
| FR-CAT-4 | §4.1 windowed queries, memory independent of catalog size |
| FR-CAT-5 | §3.5 two-pass metadata |
| FR-CAT-6 | §4.3 indexed filter compilation, §5 selectors |
| FR-CAT-7 | §2 collections schema, §5 manual and smart, §7a hierarchy, membership and editing |
| FR-CAT-9 | §3.4 the offline/deleted distinction and the sweep guard |
| FR-CAT-11 | §3.5 lazy content hashing |
| FR-NC-3 | §7.2 range-extract for remote thumbnails |
| FR-NC-6a | §5 shared selector type |
| FR-NC-6c | §3.5 metadata honesty, §7.2 availability-driven sourcing |
| NFR-P1 | §3.1 one stat per directory, not per file |
| NFR-P3 | §7.1 on-demand generation |
| NFR-ARCH-2 | §6.3 priority classes shared with the GPU scheduler |
| FR-CULL-5 | §10a burst grouping: capture-time proximity and image similarity, collapse without selection |
| NFR-ARCH-3 | §4.3 query cancellation, §6 job cancellation |
| NFR-RES-4 | §7.3 LRU cap, eviction order |
| FR-CULL-8 | §10.1 `faces` schema, §6.1 `DetectFaces` job kind on the proxy tier |
| FR-CULL-9 | §10.3 per-library calibration, stored with its validity and source hash |
| FR-CULL-10 | §10.1 `people` and `face_person`, merge-by-redirect, §10.2 clustering as a library pass |
| FR-CULL-11 | §5 `Selector::Person`, confirmed-only by default |
| FR-CULL-12 | §10.1 derived data in the catalog; names to the sidecar, UUID as merge identity |
| FR-CULL-8a | §10.1 the seven eye columns, derived like the embedding; the filter term is `RatingFilter::eyes_open` |
| NFR-SEC-5 | §10.4 sync left undecided and off; nothing in §10 emits an embedding |
+366
View File
@@ -0,0 +1,366 @@
# DarkRoom — Code health and the cost of a contribution
**Status:** Audit · 2026-08-27
**Companion to:** [architecture.md](architecture.md), [technical-debt.md](technical-debt.md),
[view-composition.md](view-composition.md)
What it costs to add something to this codebase, measured rather than estimated, and the work that
would lower the price.
[technical-debt.md](technical-debt.md) records compromises that were *chosen* — each one has a
reason that outlived the person who took it. This document records the opposite: friction nobody
chose, which accumulated because no single commit was responsible for it. The distinction matters
when deciding what to touch. A TD entry is load-bearing until its "done when" is met; an entry here
is not defending anything.
It is also not a bug list. Everything below compiles, passes 2,042 tests and ships.
---
## 1. What was measured
Every figure in this document is reproducible from a clean checkout. Worktrees under `.claude/` and
build output under `target/` are excluded from all counts — including them roughly triples the line
totals and was the first thing to get wrong.
```bash
# Lines of Rust per crate
for d in core/* ui/* platform/* apps/* tools/*; do
[ -d "$d/src" ] && echo "$(find $d/src -name '*.rs' -exec cat {} + | wc -l) $d"
done | sort -rn
# unwrap() in production code only — split each file at its #[cfg(test)] marker
# (a naive grep counts ~1,500 and tells you nothing)
# Distinct window properties written per UI module
for f in ui/dr-ui/src/*.rs; do
echo "$(grep -oP '\b(w|window|win|ui)\.\Kset_[a-z0-9_]+(?=\()' "$f" | sort -u | wc -l) $f"
done | sort -rn
```
| Measure | Value |
|---|---|
| Rust across 19 crates | ~112,000 lines; ~78,000 after comments and blanks |
| Comment density | 28% overall, 20–34% per crate |
| Test functions | 2,042, plus 21 integration test files |
| `.unwrap()` in production code | **3** — one in `dr-gpu`, two in `dr-ingest` |
| `.unwrap()` in test code | ~1,500, which is where it belongs |
| `unsafe` blocks | 9 — three of them the face scan's SIMD kernels (faces.md §9) |
| `TRACES` tags / orphan tags | 793 / 0 |
| Resolved dependencies | 826 |
| Largest function | `dr-ui::run` — 1,855 lines |
| `AppWindow` members | 282 properties + 190 callbacks |
CI gates on `cargo fmt --check`, `cargo clippy --workspace --all-targets -- -D warnings`,
`cargo test --workspace`, a release build, and an Android cross-check. The traceability matrix is
regenerated and compared, with a pre-commit hook that keeps it in step.
---
## 2. The seams, graded
Six things a contributor might plausibly want to add. Each grade was checked against the tree.
| Feature | What you touch | Cost |
|---|---|---|
| A develop operation<br>*split toning, channel mixer* | One file in `core/dr-pipeline/ops/`. Nothing else. | **Trivial** |
| A RAW format | A `Format` variant and magic-byte recognition. `dr-decode` is a generic TIFF walker carrying only 3 format-specific branches. | **Easy** |
| A neighbourhood operation<br>*dehaze, a sharpener* | Rust in `dr-pipeline/src/ops/` implementing `Operation` + `DetailStage`, plus a stub YAML declaring `rust:` and `order:`. | **Moderate** |
| A parameter widget<br>*a colour wheel* | A `WidgetKind` variant, `develop::supported()`, the panel model, a Slint component. | **Moderate** |
| A second sync backend<br>*S3, WebDAV, a local folder* | `RemoteBackend` is the easy half. Seven UI files construct `NextcloudBackend` directly and ten signatures take it concretely — see [CH-2](#ch-2). | **Hard** |
| Anything with its own UI | `app.slint`'s root component, a ~1,000-line `wire()`, and `run()` at 1,855 lines — see [CH-1](#ch-1). | **Hard** |
The top half of that table is the good half, and it is very good. The bottom half is one problem
wearing two hats: **`dr-ui` has no seams, so every UI feature lands in the same three files.**
---
## 3. What is load-bearing, and must not be "tidied"
Listed before the findings deliberately. A remediation document that only enumerates problems invites
someone to fix something that was right.
**The operation declaration format.** `ops/*.yaml` + `build.rs` is a working plugin system that
happens to resolve at build time — see [display-and-extension.md §6](display-and-extension.md). Its
deliberate smallness is the point: the expression grammar is restricted so a declaration cannot
become a second, worse place to write code. Do not "improve" it by letting a node name arbitrary
Rust.
**No operation is named in `ui/`.** Verified: all fifteen built-in op ids grepped across every `.rs`
and `.slint` file in `ui/` yield exactly one hit, a localisation test in `labels.rs:353`. This is
what makes "a new operation is one file" true rather than aspirational, and it is the single
property most likely to be destroyed by a well-meaning special case in the panel. See [CH-5](#ch-5).
**Unimplemented widgets degrade rather than break.** `develop::supported()` lists every `WidgetKind`
explicitly instead of using a wildcard, so a new kind added to the core surfaces as a compile error
rather than as silence, and a widget nothing draws falls back to sliders with the edit still
working (ARCH §4.3a).
**The mpsc-plus-timer worker shape.** Slint's event loop must never block (NFR-P9). Threading is not
what any finding below proposes changing.
**`sync_rows` mutating rows in place.** Replacing the model breaks slider dragging.
---
## <a id="ch-1"></a>CH-1 — `dr-ui` has no view layer, and the cost is compounding
**Where:** `ui/dr-ui/src/lib.rs:836`, `library_ui.rs:4374`, `collections_ui.rs:1457`,
`ui/dr-ui/ui/app.slint`
### What it is
Every UI feature lands in the same three places: the Slint root component, one of the `wire()`
functions, and `run()`.
| Function | Lines |
|---|---|
| `lib.rs::run` | 1,855 |
| `library_ui::wire` | 998 |
| `collections_ui::wire` | 964 |
| `identity_ui::wire` | 478 |
| `masks_ui::wire` | 357 |
| `settings_ui::wire` | 317 |
Above them sits one `AppWindow` carrying 282 properties and 190 callbacks, with exactly one Slint
global in the whole `ui/` directory — and it is not exported. Cross-view state therefore has nowhere
to live except the root component, and every interaction is a callback registered inside a `wire`.
The core crates are healthy by the same measure: their largest functions are `compose_full` at 437
lines and `build.rs::emit_node` at 558, both of which earn their length. This is specific to `dr-ui`.
### Why it matters more than it did
**[view-composition.md](view-composition.md) already diagnosed this and specified the fix**, in
three independently landable stages, on 2026-08-09. None of the three has landed. In the eighteen
days since, the numbers that document itself used have moved:
| Its measure | 2026-08-09 | 2026-08-27 |
|---|---|---|
| `run()` | 500 lines | **1,855** |
| `library_ui` distinct window properties | 24 | **54** |
| `lib.rs` distinct window properties | 23 | **52** |
| `collections_ui` | 9 | 11 |
| `launch_ui` | 16 | 16 |
| `set_show_*` call sites | 7 | 9 |
| View-state booleans on `AppWindow` | 2 | **5** |
Four modules it did not list now write window properties too: `settings_ui` (39), `import_ui` (24),
`identity_ui` (18), `masks_ui` (16).
The prediction it made has also come true literally. It described `app.slint` compensating for two
mutually exclusive booleans with `if !root.show-launch && root.show-library` chains. There are now
five such booleans — `show-launch`, `show-library`, `show-identity`, `show-settings`, `show-import` —
and the chains at `app.slint:1405`, `:1445` and `:1654` are five-term conjunctions. Each new view
multiplies the conjunctions rather than adding to them.
None of this is difficult work. It is simply work that every feature must now do, in files every
other feature is also editing — which is why two contributors working in parallel conflict by
construction, and why a newcomer must read a 1,855-line startup sequence with real ordering
constraints before safely inserting a line into it.
### What to do
Execute [view-composition.md](view-composition.md) as written. Its analysis holds and its staging is
right; stage 1 alone removes the nullable-callback knot and the duplicated post-load sequence, and
is worth landing whether or not stages 2 and 3 follow.
One thing to add to it, because it was not in scope there: the `wire()` functions. They are already
sectioned internally by comment, so lifting each section into `fn wire_ratings(window, ctl)`,
`fn wire_keywords(...)` and so on is mechanical, checked entirely by the compiler, and can go one
section per commit. It gives a feature a *function* to own rather than a region of one, which is
what removes the conflict, and it does not wait on stage 1.
Do the extraction behind new features rather than as a big-bang refactor. The pile grows either way;
the question is only whether each new feature adds to it or subtracts.
**Done when:** no function in `dr-ui` exceeds 300 lines, a new view registers itself instead of
adding a boolean to `AppWindow`, and `active-view` has replaced the boolean set so both-true is
unrepresentable.
---
## <a id="ch-2"></a>CH-2 — `RemoteBackend` is an abstraction nothing above `dr-sync` uses
**Where:** `core/dr-sync/src/lib.rs:46`; seven files in `ui/dr-ui/src`
### What it is
The trait is carefully built. Capability negotiation decides the sync strategy; range reads are
documented as a hint rather than a guarantee so correctness holds either way; chunked upload is
deliberately kept internal so one server's protocol cannot leak into the interface. Every choice is
explained where it is made.
And then `trash.rs`, `import.rs`, `export.rs`, `derived_sync.rs`, `launch_ui.rs`, `library.rs` and
`settings_ui.rs` each construct `NextcloudBackend` directly — 34 references — and ten functions take
`&NextcloudBackend` rather than `&dyn RemoteBackend`. Exactly **two** sites in the tree take the
trait object, both inside `dr-sync` itself.
### Why it matters
The abstraction currently buys nothing it was designed for. Worse, it reads as though it does: a
contributor who wants a WebDAV or local-folder backend will find a well-documented trait, implement
it correctly, and only then discover that nothing above `dr-sync` can be handed the result.
There is no defence of this in the tree, which is what makes it an entry here rather than in
`technical-debt.md`. It is what a single-backend application looks like when the second backend has
not yet been attempted.
### What to do
Change the ten signatures to `&dyn RemoteBackend` and construct the backend once, behind something
the UI does not name — the same discipline `develop.rs` already applies to operations. The trait is
already correct, so this is a mechanical change, and it is much cheaper now than during a second
backend when it would be entangled with that backend's own problems.
Worth doing even if no second backend is ever written: it makes the sync layer testable against a
fake, which today it is not.
**Done when:** `NextcloudBackend` is named in at most one file in `ui/`, and a stub backend can be
substituted in a test without touching the UI.
**Done**, with one half deliberately left. `ui/dr-ui/src/remote.rs` is now the only file in the
interface that names a connector; the ten worker functions take `&dyn RemoteBackend` and will accept
a stub. Every method the UI ever called on a backend — `get`, `put`, `list`, `delete`, `create_dir`,
`move_to` — was already on the trait, so nothing had to be added to it.
What remains is **credentials**. `AppCredentials` is an app password obtained through Login Flow v2,
which is a Nextcloud protocol rather than a general notion of how one authenticates to a remote, and
seven files still name it. Abstracting it needs a decision about what an account *is* across
backends — an OAuth token, a bucket key pair and an app password have no useful common shape — and
making that decision before a second backend exists would produce a confident wrong answer. It is a
design problem rather than a mechanical one, and it should wait for the backend that forces it.
---
## <a id="ch-3"></a>CH-3 — There is no path in for a contributor who is not already here
**Where:** repository root
### What it is
No `CONTRIBUTING.md`. No `rust-toolchain.toml`, though CI pins 1.92.0 exactly. No issue or PR
templates.
The documentation that exists is excellent — 7,990 lines across 14 files, including 177 numbered
requirements — and all of it is written for someone who has already decided to work on this. Nothing
tells a newcomer which document to read first, that `core/dr-pipeline/ops/README.md` is the door
with the lowest bar, or that a clone without `git-lfs` needs one command before the build succeeds.
That last point is handled well in the code: `dr-segment`'s build script detects an LFS pointer file
and fails with an instruction rather than embedding 130 bytes and dying at inference time. It is
simply not written anywhere a first-time cloner would look.
### Why it matters
It is the cheapest item in this document and it gates every other contribution. A person who cannot
get a first build is not going to reach the parts that are good.
### What to do
Write `CONTRIBUTING.md` and point the first door at the operation format. "Add a develop operation"
is a genuinely one-file contribution with declared tests that run under `cargo test` — the best
first experience this codebase can offer, and it happens to teach the architecture's central idea on
the way through.
Then state the three things that are currently folklore: `git lfs` is a prerequisite, the first
build resolves 826 crates and takes a while (saying so stops it reading as a hang), and the
toolchain is 1.92.0. Add `rust-toolchain.toml` so that last one is enforced rather than documented —
a contributor on an older stable currently gets confusing type errors instead of a version message.
**Done when:** someone who has never seen the repository can clone it, build it, and land a new
`ops/*.yaml` node without asking a question.
---
## <a id="ch-4"></a>CH-4 — Coverage is counted by tagging, not by behaviour
**Where:** `docs/traceability.md`, `tools/traceability`
### What it is
Not a new finding — [display-and-extension.md §7](display-and-extension.md) states it plainly, and
`traceability.md` itself says coverage is the intersection of tagged and defined IDs. The tooling is
genuinely good: 732 tags, zero orphans, denominators parsed from `requirements.md` at run time
rather than hardcoded, regenerated in CI and guarded by a pre-commit hook.
What it cannot do is check that the code under a tag does the thing. `FR-DEV-8` is currently tagged
against instance-buffer plumbing a future spot-removal operation *would* use; `FR-DEV-7` against a
history row for a frontend that does not exist. Both read as covered.
### Why it matters
The 55.4% figure is an overstatement of unknown size, and the risk is that it is used as a planning
input. It is recorded here so that the number keeps its asterisk when read outside the document that
already qualified it.
### What to do
Nothing structural — the honest framing already exists in two places. Adopt
display-and-extension.md's rule going forward: **close a requirement with a test that would fail if
the behaviour were removed**, and let the percentage move slowly and mean something.
**Done when:** the rule is stated in `CONTRIBUTING.md` alongside the tag syntax, so it reaches
someone adding their first tag.
---
## <a id="ch-5"></a>CH-5 — The best invariant in the codebase is unprotected
**Where:** `ui/dr-ui/src`, `ui/dr-ui/ui`
### What it is
"No code in `ui/` names an operation" (FR-DEV-3a) is what makes the whole declarative pipeline pay
off, and it is currently maintained by discipline alone. Nothing fails if someone special-cases
`exposure` in the panel to fix a layout problem at five in the evening.
### Why it matters
It is one grep, it would take an hour, and it protects the property this audit rates highest. The
failure mode is silent and cumulative: the first special case is defensible, and by the fifth the
panel names half the chain and "a new operation is one file" has quietly stopped being true.
### What to do
A test that greps `ui/dr-ui/src` and `ui/dr-ui/ui` for every id in `ops/*.yaml` and fails on a hit,
with the current `labels.rs` localisation test as its one allowed exception. It belongs in CI beside
the traceability check, which is the existing precedent for a structural gate.
**Done when:** adding `window.set_exposure_slider(...)` to the panel fails CI with a message naming
FR-DEV-3a.
---
## 4. Order of work
Ordered by value per hour rather than by size.
| | Item | Effort | Why first |
|---|---|---|---|
| 1 | [CH-3](#ch-3) — `CONTRIBUTING.md` + `rust-toolchain.toml` | Half a day | Gates everything else; unblocks the trivial seam that already works |
| 2 | [CH-5](#ch-5) — CI gate on operation names in `ui/` | An hour | Protects the property everything else in the pipeline rests on |
| 3 | [CH-2](#ch-2) — make `RemoteBackend` load-bearing | 1–2 days | Mechanical now, entangled later; also makes sync testable |
| 4 | [CH-1](#ch-1) — split the `wire` functions | Incremental | No behaviour change, compiler-checked, one section per commit |
| 5 | [CH-1](#ch-1) — [view-composition.md](view-composition.md) stages 1–3 | Sustained | Largest and most invasive; every deferred month adds to the pile |
Items 1, 2 and 3 are independent of each other and of the rest. Item 4 does not wait on item 5.
---
## 5. What this document does not claim
It measures structure, discipline and coupling. It does **not** assess runtime correctness, GPU
shader behaviour, security posture, or whether any tagged requirement is actually implemented —
[CH-4](#ch-4) is precisely the observation that the last of those is unmeasured.
Line counts and function sizes are proxies. `run()` being 1,855 lines is a real problem because
every feature must edit it, not because 1,855 is a bad number; `build.rs::emit_node` at 558 lines is
not a problem at all. Where a figure appears above, the sentence around it says which of the two it
is.
The audit was first measured against `origin/master` at `d4a34ef`, then re-measured after
`android-bundled-face-models` merged in: `run()` moved from 1,810 lines to 1,855, and the whole-
library face sweep added its fetching path to `library.rs`. The seam grades are unchanged by that
merge — the work went into the seams that already existed rather than cutting new ones, which is
itself the pressure CH-1 describes.
+253
View File
@@ -0,0 +1,253 @@
# Finishing the display contract, and opening the pipeline
Spec for two pieces of work that turn out to be one conversation: closing **FR-DSP**, which is the
architecture's central performance claim, and reaching **FR-PLG**, which is the only requirement
family at zero.
They belong in one document because the same property decides both. The pipeline composes its work
from *declarations* — an operation says what its parameters are and contributes a WGSL fragment,
and the composer fuses the active ones into a single dispatch. That is why the display path is fast,
and it is also, already, most of a plugin format. Finishing one and opening the other are the same
seam approached from two sides.
---
## 1. What is actually true today
Stated first because both halves of this document are smaller than the requirement numbers suggest,
and the reason is that some of the work is done and untagged.
| Requirement | Reality |
|---|---|
| FR-DSP-1 proxy rendering | **Done.** The develop view renders at viewport resolution, not source. |
| FR-DSP-2 tiled computation | **Absent, and §2 now says it should stay that way.** Measured: the fused pass is inside the budget everywhere. See [frame-budget.md](frame-budget.md). |
| FR-DSP-3 interactive latency | **Measured and asserted** for the fused path — `core/dr-gpu/tests/frame_budget.rs`. Missed by one operation, clarity, for the reason recorded as TD-4. |
| FR-DSP-4 progressive refinement | **Absent**, and §4's condition did not fire. Every render is full quality and can afford to be. |
| FR-DSP-5 zoom and pan | **Done and tagged**, against tests that fail if the behaviour is removed — `core/dr-gpu/tests/zoom_resolution.rs`. `Framing::view` shrinks the sampled region while the render target keeps its size, so zooming *raises* the resolution the pipeline works at. That is FR-DSP-5's requirement, arrived at without tiles. |
| FR-DSP-6 colour management | **Done.** Output space is a parameter of composition. |
| FR-DSP-7 histogram and clipping | **Done**, GPU-side, no per-frame readback. |
| FR-DSP-8 per-display colour | **Done**, with one caveat named in §5.4. Acquisition per display server, an sRGB fallback that is visible in About, and the canvas rendered at physical pixel size. |
| FR-PLG-* | **Zero tagged.** But `ops/*.yaml` + `build.rs` is already the class-1 plugin format compiled at build time rather than loaded. |
Two of the five uncovered display requirements are therefore *measurement and tagging*, not
construction. That is worth knowing before anyone plans a quarter around them.
**Both have since been done.** [frame-budget.md](frame-budget.md) holds the measurements §2 asks
for and the reading of its decision rule; the table above is updated to match. The rest of this
document is left as it was written, because a plan that has been overtaken by its own evidence is
more useful read in order than quietly edited into agreement.
---
## 2. Measure before building tiles
**FR-DSP-2 is the one requirement in this document that may not be worth satisfying as written.**
> **Resolved.** M1–M3 were run; the numbers and the verdict are in
> [frame-budget.md](frame-budget.md). The rule below fired for *rewrite*: every point-operation
> chain is inside 16 ms at the 99th percentile at every viewport size, fit and at 1:1, the widest
> being 4.5 ms of GPU at 4K. The measurement did find a stage that misses the budget — clarity's
> 52-pixel kernel, 34 ms at 4K — and tiling makes that stage *worse*, since a tiled convolution
> reads a halo per tile. It is recorded as TD-4 with the fix its own module already names.
The requirement predates the fused-shader design. It assumes the pipeline is a chain of passes over
a large buffer, where recomputing everything on each frame would be ruinous and tiles are the way
out. What was built instead composes every active operation into **one dispatch over a
viewport-sized target** — at 2000×1300 that is 2.6 M pixels, once, for the whole chain.
So the question tiling was invented to answer may already be answered. Before any tile scheduler is
written:
**M1 — Frame cost at proxy resolution.** Time `render_detailed` at 1920×1200, 2560×1600 and
3840×2160, with a chain of one operation, five, and every operation active. Report the 99th
percentile, not the mean; a slider drag is judged by its worst frame.
**M2 — Frame cost at 1:1 on a large file.** The same, with `Framing::view` zoomed to 1:1 on a 60 MP
frame, which is the case FR-DSP-5 names and the one where the sampled region is smallest but the
detail chain's kernels are widest.
**M3 — Cost of the detail stage separately.** Neighbourhood operations dispatch per pass and are the
only part of the chain whose cost is not one read and one write. A separable blur at a large radius
is the plausible budget-breaker, not the fused pass.
**The decision rule, fixed in advance.** If M1 and M2 sit inside 16 ms at the 99th percentile,
**FR-DSP-2 is rewritten rather than implemented**: tiling stops being an interactive-path
requirement and becomes what it actually is for this architecture — a *scheduling* concern for
export and thumbnailing, which already run off the frame path. If they do not, the measurement tells
us which stage to tile, which is a far better starting point than tiling everything on principle.
Writing a tile scheduler that the design does not need would be the most expensive way to discover
this. ARCH §5.3's tile cache keyed by `(VersionId, tile, zoom, graph_hash_prefix)` is a good design
for a pipeline that needs it; the burden of proof is that this one does.
---
## 3. FR-DSP-3 — make the budget a test, not an aspiration
A latency requirement that nothing asserts is a wish. The work is:
**3.1** A bench in `dr-gpu` that renders a fixed chain at a fixed size and reports percentiles.
Committed with its numbers, so a regression is a diff rather than a memory.
**3.2** A test that *fails* when a frame exceeds the budget on the reference desktop, skipping where
there is no adapter — the pattern the GPU tests already use. It should assert the 99th percentile of
a hundred frames, because the failure mode being guarded against is a stutter, not an average.
**3.3** The asynchronous half of the requirement: "when a full-resolution result is needed it is
computed asynchronously, and the proxy result remains on screen until it is ready." Nothing does
this today because nothing needs a full-resolution result on the frame path — export renders its
own. This clause should be **narrowed to export and 1:1 zoom** or struck, and struck is defensible.
---
## 4. FR-DSP-4 — progressive refinement
The one genuinely new piece of interactive work, and it is small because the pipeline is already
resolution-parametric.
During a drag, render at a fraction of the viewport and let the compositor scale; when the gesture
settles, render at full viewport size. `DevelopSession` already knows when a drag is in flight —
`drag-changed` exists on every slider and is what stands the Flickable down.
Two things decide whether this is worth having, and M1 answers both. If a full-quality frame is
already inside budget, reduced-quality rendering buys nothing and costs a visible softness during
every drag — which the requirement itself warns against ("refinement is visually smooth, not a
jarring swap"). **This requirement is conditional on M1 failing.** If M1 passes, FR-DSP-4 is
satisfied vacuously: there is no rapid interaction the app cannot render at full quality, which is a
stronger outcome than refining.
---
## 5. FR-DSP-8 — per-display colour
The only display requirement needing platform work rather than pipeline work, and the only one where
being wrong is a correctness defect rather than a slow frame: a second monitor with a different
profile shows wrong colours, silently.
**5.1 Acquisition, per display server.** X11 has `_ICC_PROFILE` atoms per output. Wayland's
colour-management protocol is not universally available, and the requirement already anticipates
this by demanding "a defined fallback where Wayland provides no profile" — that fallback is sRGB,
stated in the About page beside the other diagnostics so a photographer can see which path they are
on rather than wonder.
**5.2 Reacting to a move.** The transform is selected per the display currently showing the canvas
and updates when the window moves. The composed output space is already a parameter of composition
(`compose_with_framing(..., output)`), so a display change is a recomposition, not a pipeline
change. This is the part the existing design makes cheap.
Slint turned out *not* to report window moves — there is no `on_moved` on any backend — so the
window's position and scale factor are sampled twice a second and the platform re-surveyed only when
they differ. See §5.4 for the display server where the position itself is unavailable.
**5.3 Fractional scaling.** "Handled without resampling artefacts in the canvas" — the canvas is a
wgpu texture handed to the compositor, so the requirement is that we render at the *physical* pixel
size rather than the logical one and let the compositor present 1:1. Worth an explicit test, since
the failure is subtle: a slightly soft canvas that looks like a bad demosaic.
### 5.4 What landed, and the one thing that did not
`dr_plat::display` surveys the session's displays; `dr_ui::display_ui` decides which one is showing
the canvas and keeps `DevelopSession`'s output space pointed at it. Composition was already
parameterised on the space, so the pixel path changed by one argument.
Two things are worth recording because they are trades rather than omissions.
**A profile is matched to the nearest of four spaces, not applied.** A measured panel is none of
`Srgb`, `DisplayP3`, `AdobeRgb` or `ProPhoto`, and a general ICC engine is a much larger piece of
work — a CMM, rendering intents, LUT-based profiles, and a per-frame cost to argue about. The
profile is reduced to its D50-adapted colorants and matched against the four; a match that is merely
nearest is marked as such, and About says "nearest to Display P3" rather than "Display P3". A
LUT-based profile, which is what a hardware calibrator often writes, is declined by shape and falls
back to sRGB with that stated. An approximation the photographer can see beats a silent one.
**On Wayland the canvas follows the first output, not the window.** A Wayland client is never told
where its window is — `xdg_toplevel` carries no position, deliberately — so the "which display"
question cannot be answered by geometry there. The protocol's own answer is
`wp_color_management_surface_feedback_v1`, which hands a client the preferred image description for
*its surface* and re-sends it on a move; it needs the application's `wl_surface`, which Slint owns
and does not expose. So the profiles are read correctly for every output and the *selection* among
them is right on X11 and on any single-monitor Wayland session, which is most of them. Closing the
gap is a Slint surface handle, not a change to any of this.
---
## 6. Extensibility: the format already exists
`FR-PLG-2` says "the node declaration is the plugin format". That is already true — it is simply
resolved at build time:
```
ops/exposure.yaml ──build.rs──▶ generated Rust impl Operation ──▶ fused shader
```
A declaration names its parameters, their ranges and units, its attributes, its WGSL body and its
neutral. `build.rs` compiles that into something indistinguishable from a hand-written operation.
**Nothing about that requires the declaration to be present at compile time** — everything it
produces is data plus a WGSL string, and the composer already assembles WGSL at run time from
whatever operations are active.
So class 1 is not a new mechanism. It is the existing one, loaded later.
### 6.1 What has to change
**6.1.1 Descriptors become owned, not `&'static`.** `Operation::descriptor()` returns
`&'static OpDescriptor` today, which is what makes a build-time node free and a run-time node
impossible. This is the one invasive change in the whole plan and everything else waits behind it.
`Arc<OpDescriptor>` is the obvious shape; the cost is one refcount per descriptor read, on a path
that reads descriptors when the panel is built rather than per frame.
**6.1.2 A run-time node type.** One `DeclaredOp` implementing `Operation` from an owned
declaration, replacing *generated code per node* with *one interpreter over many declarations*. The
generated path can stay for the built-in chain — it costs nothing and keeps the built-ins
inspectable — but the two must produce identical behaviour, which is a test: parse each built-in
`ops/*.yaml` at run time and assert the composed WGSL matches the generated one byte for byte.
**6.1.3 WGSL validation at load, not at dispatch.** A plugin's fragment is a string from a stranger.
`compose` already builds a full shader and `naga` will reject bad source, but the failure currently
surfaces as a broken render. A plugin's source must be compiled and rejected at *load*, with the
error naming the plugin, because the alternative is an app that draws nothing and blames itself.
**6.1.4 Order and identity.** `order:` decides chain position and `build.rs` already refuses
duplicates — that guard becomes load-time. Plugin ids need a namespace (`author.name`) so two
plugins cannot collide, and the sidecar stores parameters by `(op_id, param_id)`, so an id collision
is a *wrong edit silently applied*, exactly the failure `MaskSource::Regions`' signature exists to
prevent.
### 6.2 What this buys immediately
The features enumerated as missing against Lightroom that are *pure point operations* become
declarations rather than code: split toning, colour zones, selective colour, creative vignette,
channel mixer variants. A photographer-author can write one without a Rust toolchain, and the
existing `ops/README.md` is already its documentation.
**It does not buy the neighbourhood operations** — dehaze, spot removal, liquify — because those are
`DetailStage` implementations with kernels and per-render scale conversion, which FR-PLG-2a
anticipates by naming "fragment nodes and pass nodes" as two templates. Pass nodes are a second
phase and should not gate the first.
### 6.3 Order of work
1. Owned descriptors (6.1.1) — invasive, unblocks everything, no user-visible change
2. `DeclaredOp` + byte-identical parity test against the generated built-ins (6.1.2)
3. Load-time WGSL validation and id namespacing (6.1.3, 6.1.4)
4. A directory that is read at startup, and one shipped example that is not a built-in
5. Pass nodes (FR-PLG-2a's second template), once 1–4 are load-bearing
Classes 2 and 3 — view plugins and computational plugins — are deliberately not in this plan.
FR-PLG-3a's "a view plugin cannot be trusted with the UI thread" and FR-PLG-4a's capability grants
are both larger design problems than class 1, and class 1 is where the requested features live.
---
## 7. What this document does not claim
Traceability counts a requirement as covered when a `TRACES` tag names it. It does not check that
the code under the tag does the thing — `FR-DEV-8` is currently tagged against instance-buffer
plumbing that a future spot-removal operation would use, and `FR-DEV-7` against a history row for a
frontend that does not exist. Both read as covered.
So the 51% figure is an overstatement of unknown size, and closing FR-DSP by tagging what already
works would make it a larger one. **Every requirement closed by this plan should be closed by a
test that would fail if the behaviour were removed**, which is the only kind of coverage worth
counting.
+254
View File
@@ -0,0 +1,254 @@
# DarkRoom — Distribution
**Satisfies:** NFR-COMPAT-2 (v1 channels) · FR-PLAT-LIN-3 (sandboxed distribution)
**Companion to:** [requirements.md](requirements.md) §3.8, §4.8 · [storage.md](storage.md)
NFR-COMPAT-2 asks for the v1 channels to be *stated*, and says why in its own
second sentence: the channel decision and the storage design are coupled. A
channel is not a build target. It is a set of constraints that reach back into
the code — what the application is allowed to see, what it may ask for, and
what it must be able to do without asking. This document records which channels
v1 targets and what each one costs, and it is where to look before adding a
permission to a package rather than after.
---
## 1. The channels
| Platform | Channel | State | What it constrains |
|---|---|---|---|
| Linux | Arch source package — [`packaging/PKGBUILD`](../../packaging/PKGBUILD) | Built, in tree | Nothing. Full filesystem access, system Vulkan, system secret daemon |
| Linux | Flatpak — [`packaging/flatpak/`](../../packaging/flatpak/) | Manifest in tree, **library selection does not work** (§4) | Portals only. No `--filesystem=`, no host mount table, no typed paths |
| Linux | AppImage | v1 channel, **recipe not yet written** (§5) | Oldest supported glibc, and no sandbox at all |
| Android | F-Droid | v1 channel, not yet submitted | GPLv3-clean build, reproducible, no proprietary blobs |
| Android | Play Store | **Not v1** (§6) | Would make ARCH §6.9 binding as policy rather than as engineering |
| Windows | NSIS per-user installer, cross-built — [windows.md](windows.md) | Built by CI, **untested on Windows**. Not v1 | Known folders in place of XDG; no sandbox; unsigned until there is a certificate |
Four of these six exist as recipes and two do not. That is stated rather than
smoothed over, because the value of writing the channels down is knowing which
constraints are already being met and which are promises.
### What every channel has to get right
Independent of packaging format, and each of these has bitten a package
somewhere:
- **One identifier, four places.** `paris.tourolle.darkroom` is the AppStream
component id, the `.desktop` basename, the Flatpak application id, and the
string `dr_ui::run` sets as the Wayland `app_id` and X11 `WM_CLASS`. A rename
that misses one of them costs the icon in the shell or the association in the
software centre, and neither failure announces itself.
- **The metainfo, not just the desktop entry.**
[`packaging/paris.tourolle.darkroom.metainfo.xml`](../../packaging/paris.tourolle.darkroom.metainfo.xml)
is the single description of the application, installed by every channel that
has somewhere to put it. Its `metadata_license` is CC0-1.0 and its
`project_license` is GPL-3.0-or-later; those differ on purpose — see the
comment in the file.
- **Vulkan is a requirement, not a preference.** The develop pipeline is
compute shaders through wgpu, and NFR-R8 — how far a CPU fallback goes — is
still open, so today there is nothing behind it. A package that installs onto
a machine with no working ICD produces an application that starts and cannot
develop.
- **A Secret Service implementation, or an honest degraded mode.** FR-NC-2 is
explicit that the absence of a secrets daemon is a stated degraded mode and
never a silent fall back to plaintext. Packages express this as an optional
dependency (the PKGBUILD) or a talk hole (the Flatpak manifest), never as a
hard dependency — a headless or minimal-WM install is a supported way to run.
- **The face models are Git LFS objects.** A checkout without `git lfs pull`
has ~130-byte pointers where 11 MB models should be. Both the PKGBUILD and
the Flatpak manifest check the file size and refuse, because the alternative
is a package whose face indexing fails inside the graph loader on a user's
machine rather than on the packager's.
---
## 2. Why Flatpak is the channel that matters most
Not because it is expected to be the most used. Because it is the only one that
tests anything.
The Arch package and an AppImage both hand the application the same
unrestricted process the developer runs it in, so neither can discover that a
design assumed unrestricted access. Flatpak takes that assumption away, and
FR-PLAT-LIN-3 exists to make the discovery happen deliberately rather than in a
bug report. §4 is what it discovered.
The same argument runs the other way on Android, where SAF has been the only
option since before the first line was written (ARCH §6.9) and `SourceRef`
exists because of it. Linux got the abstraction — `LocalStorage::grant` is the
one place a `Path` enters — and never got the constraint that would have proved
it worked.
---
## 3. What already works inside the sandbox, unchanged
Worth listing, because it is the part FR-PLAT-LIN-1 quietly paid for in
advance:
- **XDG directories.** Flatpak redirects `XDG_CONFIG_HOME`, `XDG_DATA_HOME` and
`XDG_CACHE_HOME` into `~/.var/app/paris.tourolle.darkroom/`. Settings
(`settings_store.rs`), accounts (`dr_sync::account`), the catalog and the
thumbnail store all read those variables, so every one of them lands in the
application's own directory with no code change and no permission.
- **The face models.** `system_face_models_dirs()` reads `$XDG_DATA_DIRS`
rather than hard-coding `/usr/share`, which is exactly why `/app/share`
inside a Flatpak is found by the same lookup that finds the Arch package's
copy.
- **Opening a photograph from a file manager.** The `.desktop` entry declares
the RAW MIME types and `Exec=darkroom-desktop %F`; under Flatpak the file is
exported through the document portal and arrives in `argv` as a path under
`/run/user/$UID/doc/`, which is mounted in every sandbox. `main.rs` takes
paths from `argv` and `collect()` handles a file or a directory. This is
genuine portal-mediated access and it needs nothing new.
- **The Nextcloud sign-in browser.** `open_in_browser` spawns `xdg-open`; the
freedesktop runtime's `xdg-open` forwards to the OpenURI portal, and portal
calls need no `--talk-name` because Flatpak always permits them. FR-NC-1's
"system browser, never an embedded webview" therefore holds inside the
sandbox for the same reason it holds outside it.
- **Credentials.** The keyring crate speaks the Secret Service D-Bus interface,
reached through the session-bus proxy with one talk hole. The app password
stays visible to `secret-tool` and Seahorse, which is what keeps it
individually revocable by the user.
---
## 4. What does not work: choosing a library
**FR-PLAT-LIN-3 is not satisfied today, and the manifest does not pretend
otherwise.**
A folder library is chosen by typing an absolute path. `dr-sync-folder`'s
provider declares `SignIn::EndpointOnly` with the placeholder
`/home/you/Pictures`, and `normalise_endpoint` expands `~`, requires the path
to be absolute, and checks it with `std::fs`. Nothing in the tree calls the
FileChooser portal — there is no `ashpd`, no `rfd`, and no toolkit file dialog
anywhere in `ui/`, `platform/` or `core/`.
Inside a sandbox with no `--filesystem=`, `$HOME` still resolves to the real
home *path* but that directory holds only the application's own
`.var/app/…` tree. So a typed `~/Pictures` fails the `exists()` check and the
launch screen says `No folder at /home/you/Pictures.` — a truthful message
about a situation the user cannot fix from inside the application.
Import is blocked one step earlier. `dr_plat::volumes()` finds a camera card by
reading `/proc/self/mountinfo` and the `removable` flag under `/sys`. A
sandboxed process is in its own mount namespace, so the table it reads
describes the sandbox; a card mounted at `/run/media/…` on the host is not in
it. `volumes()` correctly returns an empty list, which the interface presents
as "no card found" — right for the code, wrong for the user, who is looking at
a card.
### The permission that would hide this, and why it is not in the manifest
`--filesystem=host` makes both work immediately and is the thing FR-PLAT-LIN-3
names as the alternative to portals. Granting it would mean the sandboxed build
never exercises the sandbox, which removes the entire reason for shipping one
(§2). `--filesystem=xdg-pictures` is narrower and would be tempting, but it is
still a static grant that lets a typed path resolve — it makes the same design
work by not testing it, only in a smaller directory.
So the manifest grants no filesystem access at all. The consequence is stated
plainly: **a Flatpak built from this manifest can open photographs handed to it
and cannot yet be pointed at a library.**
### What closes it
Two changes, in this order:
1. **A portal file chooser behind a platform seam.** `ashpd`'s
`OpenFileRequest` with `directory(true)` returns a URI the document portal
has exported, which the sandbox can read and which stays valid across
restarts. It resolves to a real path under `/run/user/$UID/doc/`, so
`normalise_endpoint` accepts it as it stands — `canonicalize()` on a fuse
path returns the path itself. The seam matters more than the crate: this
belongs beside `LocalStorage::grant` in `dr-plat`, which is already the one
place a `Path` enters the application, and must not become a second way for
`ui/` to learn about paths.
2. **Removable volumes through the same door.** There is no portal for "list
the mounted cards". The honest answer is that under a sandbox
`imports_supported()` should report the same `false` it reports on Android,
for the same reason it gives there — the operation cannot be performed
however hard the user tries — and the import flow should offer the folder
chooser instead of a volume list.
**Done when:** a Flatpak built from
[`packaging/flatpak/paris.tourolle.darkroom.yml`](../../packaging/flatpak/paris.tourolle.darkroom.yml),
with its `finish-args` unchanged and no `flatpak override` applied, can select a
library root, scan it, and write a sidecar back into it.
### Running a Flatpak build before then
For testing the rest of the application inside the sandbox, grant the access
per-installation rather than in the manifest, so the file that describes the
application keeps telling the truth:
```bash
flatpak override --user --filesystem=~/Pictures paris.tourolle.darkroom
```
---
## 5. AppImage
A v1 channel, and the recipe is outstanding work rather than a decision to be
made. What it will have to account for, none of which is a surprise:
- **glibc.** An AppImage links against the oldest glibc it must run on, so it
is built in a container with an old base rather than on a rolling-release
developer machine. A release binary built on a current rolling-release host carries
`GLIBC_2.44` references and would run on almost nothing else.
- **What to bundle and what not to.** The binary links fontconfig, freetype,
expat, libpng, zlib, brotli and bzip2 — bundle those. It does *not* link
Vulkan, libxkbcommon or either display-server library: wgpu `dlopen`s
`libvulkan.so.1`, and `x11rb` and `wayland-client` speak the wire protocols
in Rust. The Vulkan loader and the ICD must come from the host, and bundling
a loader is the classic way to break an AppImage on a driver it did not
expect.
- **The models.** ~15 MB of ONNX weights inside the image, or a first-run
download. In-tree is consistent with how the Lensfun database ships and with
NFR-SEC-5's local-first posture; the licence question (D13) is the same one
it is everywhere else and is not made easier or harder by this channel.
- **No sandbox.** An AppImage tests nothing about FR-PLAT-LIN-3. It is a
convenience channel for distributions the PKGBUILD does not serve, and should
never be the channel a portal problem is discovered on.
---
## 6. Android: F-Droid in v1, Play deferred
NFR-COMPAT-2 says Play distribution is what makes ARCH §6.9's constraints
binding, and that is worth reading precisely, because the constraint is already
met and would be met whatever the channel.
§6.9 is *verified*, not assumed: `MANAGE_EXTERNAL_STORAGE` is not grantable
under Play policy, and `READ_MEDIA_IMAGES` would not help because proprietary
RAW is not typed `image/*` by the platform scanner and does not appear in
`MediaStore.Images`. SAF is the only route that works, so FR-PLAT-AND-1 asks
for it unconditionally and `SourceRef` (ARCH §3.1) exists to make it possible.
A sideloaded or F-Droid build *could* ask for broader permissions; it would
gain nothing by doing so.
So the coupling runs the opposite way from how it is usually described. Play is
deferred for a reason that has nothing to do with storage: GPLv3 distribution
through Play is generally workable but has not been confirmed for this project
(ARCH §14), and F-Droid has no such question. Confirming it is a licence-reading
exercise; nothing in the storage design waits on the answer.
---
## 7. Where the recipes live
```
packaging/
PKGBUILD Arch source package
paris.tourolle.darkroom.desktop the desktop entry, installed by every channel
paris.tourolle.darkroom.metainfo.xml AppStream, installed by every channel
flatpak/
paris.tourolle.darkroom.yml the manifest, and where the permissions are argued
windows/
darkroom.nsi the installer; docker/windows/package.sh drives it
```
`packaging/` also accumulates built `.pkg.tar.zst` artefacts from local
`makepkg` runs. Those are not part of any channel and should not be committed.
+1579
View File
File diff suppressed because it is too large Load Diff
+433
View File
@@ -0,0 +1,433 @@
# What a frame costs
**Status:** Measured · 2026-08-27
**Companion to:** [display-and-extension.md](display-and-extension.md) §2–3 ·
[requirements.md](requirements.md) §3.4 FR-DSP-2, FR-DSP-3, FR-DSP-4
**Instrument:** [`core/dr-gpu/examples/frame_budget.rs`](../../core/dr-gpu/examples/frame_budget.rs)
**Guard:** [`core/dr-gpu/tests/frame_budget.rs`](../../core/dr-gpu/tests/frame_budget.rs)
[display-and-extension.md](display-and-extension.md) §2 fixed a decision rule in
advance and made three measurements the thing that settles it. This file is
those measurements, and the recommendation they support.
Rerun with:
```sh
cargo run --release -p dr-gpu --example frame_budget
```
and diff this file. That is the whole point of committing numbers: a regression
should be a diff rather than somebody's recollection of how fast it used to be.
---
## The answer, first
**FR-DSP-2 should be rewritten, not implemented.** M1 and M2 sit inside the
16 ms budget at the 99th percentile for every chain of point operations at every
viewport size measured, fit and at 1:1 — the widest case, every operation that
contributes a fragment to the fused shader at 4K, costs **4.5 ms** on the GPU and
**8.2 ms** including the composition that precedes it. Tiling the interactive
path would be optimising something that is already using a quarter of its budget.
**But the measurement did find a budget-breaker, and it is not the one tiling
fixes.** The neighbourhood stage — clarity in particular — costs **34 ms at 4K
on its own**, twice the whole budget, and tiles do not help it: a tile of a
convolution has to read its halo, so tiling raises the total tap count rather
than lowering it. §2 predicted this exactly ("a separable blur at a large radius
is the plausible budget-breaker, not the fused pass"), and the fix it needs is
the one `local_contrast`'s own module documentation already names — a base
computed at reduced resolution — which is a change to `crate::detail`, not a
tile scheduler.
There is a third finding nobody was looking for: **shader composition costs
3–5 ms of CPU per frame on a full chain**, on the UI thread, before any GPU work
is submitted. That is a fifth to a third of the budget spent formatting strings,
and it is invisible to any amount of tiling.
---
## Conditions
| | |
|---|---|
| Adapter | NVIDIA GeForce RTX 3050 6GB Laptop GPU (Vulkan) |
| Source | 9504 × 6336 synthetic (60.2 MP, 482 MB as `rgba16f`) |
| Frames | 100 measured per row, 12 warm-up frames discarded |
| Percentile | Nearest-rank, so p99 of 100 frames is the second-worst frame |
| Build | `--release` |
| Date | 2026-08-27 |
`shader` is `EditGraph::compose` alone. `cpu` adds the detail chain and the
invalidation hash — everything `DevelopSession::render` does per frame before it
dispatches. `gpu` is submit plus wait-for-idle, which serialises the GPU work
into the frame that caused it and is therefore pessimistic. `TOTAL` ranks
`cpu + gpu` summed **within each frame**, which is the column the budget is
judged on; adding two percentiles instead would invent a stutter that no frame
actually had.
Chains: `one` is exposure. `five` is exposure, contrast, highlights/shadows,
blacks/whites, vibrance. `point` is every operation in the default chain that
contributes a fragment to the fused shader, film stock included. `all` is `point`
plus the four neighbourhood operations — noise reduction, capture sharpening,
clarity and texture.
---
## M1 — the fused pass at proxy resolution
The develop view: the whole frame fit to the viewport.
| size | chain | shader | cpu p99 | gpu p50 | gpu p99 | TOTAL | |
|------------:|------:|-------:|--------:|--------:|--------:|--------:|:-----|
| 1920 × 1200 | one | 0.08ms | 0.10ms | 1.02ms | 1.23ms | 1.31ms | |
| 1920 × 1200 | five | 0.18ms | 0.20ms | 1.01ms | 1.20ms | 1.36ms | |
| 1920 × 1200 | point | 2.79ms | 2.82ms | 1.98ms | 2.18ms | 4.83ms | |
| 1920 × 1200 | all | 3.65ms | 4.73ms | 6.86ms | 7.37ms | 12.02ms | |
| 2560 × 1600 | one | 0.08ms | 0.10ms | 1.73ms | 2.00ms | 2.12ms | |
| 2560 × 1600 | five | 0.24ms | 0.27ms | 1.73ms | 2.26ms | 2.46ms | |
| 2560 × 1600 | point | 2.82ms | 2.85ms | 2.37ms | 2.65ms | 5.38ms | |
| 2560 × 1600 | all | 3.37ms | 4.35ms | 14.31ms | 15.65ms | 18.42ms | OVER |
| 3840 × 2160 | one | 0.10ms | 0.14ms | 3.09ms | 3.31ms | 3.42ms | |
| 3840 × 2160 | five | 0.30ms | 0.32ms | 3.03ms | 3.40ms | 3.61ms | |
| 3840 × 2160 | point | 3.62ms | 3.65ms | 4.12ms | 4.52ms | 8.23ms | |
| 3840 × 2160 | all | 4.12ms | 5.07ms | 37.73ms | 40.17ms | 43.24ms | OVER |
Read the `point` rows: **the fused dispatch scales with pixels and almost not at
all with chain length.** Going from one operation to the entire point chain at
4K costs 1.2 ms of GPU. Going from 2.3 M pixels to 8.3 M costs 2.3 ms. Both are
small, and the second is the one tiling would address.
The `all` rows go over, and the `point` rows in the same block are what say why:
the difference between them is the neighbourhood stage, measured on its own in
M3 and arriving at almost exactly the same figure.
## M2 — the same, zoomed to 1:1 on the 60 MP source
FR-DSP-5's case. `Framing::view` shrinks the sampled region while the render
target keeps its size, so one render pixel lands on one source pixel.
| size | chain | shader | cpu p99 | gpu p50 | gpu p99 | TOTAL | |
|------------:|------:|-------:|--------:|--------:|--------:|--------:|:-----|
| 1920 × 1200 | one | 0.12ms | 0.14ms | 0.42ms | 0.66ms | 0.75ms | |
| 1920 × 1200 | five | 0.27ms | 0.30ms | 0.49ms | 1.14ms | 1.22ms | |
| 1920 × 1200 | point | 3.49ms | 3.52ms | 1.18ms | 1.39ms | 4.85ms | |
| 1920 × 1200 | all | 3.64ms | 5.18ms | 8.96ms | 9.55ms | 14.30ms | |
| 2560 × 1600 | one | 0.11ms | 0.12ms | 0.56ms | 0.99ms | 1.06ms | |
| 2560 × 1600 | five | 0.22ms | 0.25ms | 0.77ms | 1.02ms | 1.17ms | |
| 2560 × 1600 | point | 2.96ms | 2.99ms | 2.03ms | 2.52ms | 5.61ms | |
| 2560 × 1600 | all | 5.24ms | 7.35ms | 18.80ms | 21.62ms | 25.81ms | OVER |
| 3840 × 2160 | one | 0.10ms | 0.12ms | 1.31ms | 1.52ms | 1.63ms | |
| 3840 × 2160 | five | 0.14ms | 0.27ms | 1.39ms | 1.64ms | 1.75ms | |
| 3840 × 2160 | point | 3.14ms | 3.16ms | 4.04ms | 4.50ms | 7.21ms | |
| 3840 × 2160 | all | 4.84ms | 6.78ms | 47.22ms | 48.79ms | 54.47ms | OVER |
**A 1:1 view of a 60 MP file is cheaper than the fit view of the same file**, for
every point chain and at every size — 1.52 ms against 3.31 ms for one operation
at 4K. That is not a rounding artefact and it is worth stating plainly, because
it is the opposite of what "full resolution" sounds like it should cost. The
dispatch is the same number of pixels either way; what changes is where those
pixels read from. A fit view walks the whole 482 MB texture on a stride, and a
1:1 view reads a contiguous window of it that fits comfortably in cache.
So the resolution FR-DSP-5 promises costs nothing extra on the fused path.
Zooming is not an expensive mode to be dreaded and progressively refined into;
it is the cheap one.
The `all` rows are worse at 1:1 than fit, and that is the detail stage again for
a specific reason: noise reduction's radius is stated in *source* pixels, so
`RenderScale::ratio` climbing to 1.0 widens its kernel. Clarity's is stated as a
fraction of the frame and does not move. M3 separates the two.
## M3 — the neighbourhood stage alone
Timed with the fused dispatch deliberately reused: only a detail parameter moves,
so `render_detailed` skips the colour pass (FR-DEV-3d) and what remains is the
convolutions. `colour` counts fused dispatches over the measured frames and is
zero on every row, which is what makes these numbers mean "detail alone" rather
than asserting it.
| size | stage | view | pass | radius | colour | cpu p99 | p50 | p99 |
|------------:|---------:|:-----|-----:|-------:|-------:|--------:|--------:|--------:|
| 1920 × 1200 | clarity | fit | 2 | 29 | 0 | 1.60ms | 5.42ms | 5.99ms |
| 1920 × 1200 | all four | fit | 7 | 29 | 0 | 1.71ms | 5.78ms | 6.16ms |
| 1920 × 1200 | clarity | 1:1 | 2 | 29 | 0 | 1.12ms | 7.51ms | 8.01ms |
| 1920 × 1200 | all four | 1:1 | 9 | 29 | 0 | 2.71ms | 8.25ms | 9.11ms |
| 2560 × 1600 | clarity | fit | 2 | 38 | 0 | 1.03ms | 12.02ms | 12.44ms |
| 2560 × 1600 | all four | fit | 7 | 38 | 0 | 1.87ms | 12.49ms | 13.16ms |
| 2560 × 1600 | clarity | 1:1 | 2 | 38 | 0 | 1.87ms | 15.82ms | 16.60ms |
| 2560 × 1600 | all four | 1:1 | 9 | 38 | 0 | 2.71ms | 17.24ms | 18.06ms |
| 3840 × 2160 | clarity | fit | 2 | 52 | 0 | 1.75ms | 33.11ms | 33.89ms |
| 3840 × 2160 | all four | fit | 7 | 52 | 0 | 1.76ms | 34.21ms | 35.03ms |
| 3840 × 2160 | clarity | 1:1 | 2 | 52 | 0 | 1.08ms | 40.39ms | 41.86ms |
| 3840 × 2160 | all four | 1:1 | 9 | 52 | 0 | 2.37ms | 43.29ms | 44.72ms |
`radius` is the widest halo any pass reads, in render pixels.
Clarity alone is 97% of the cost of all four neighbourhood operations together,
at every size. Its σ is 1.2% of the shorter edge and it truncates at 2σ, so its
radius is 29 px on a 1200 px viewport and **52 px at 4K** — two separable passes
of 105 taps each, over 8.3 M pixels, which is 1.7 billion texture reads. That is
the whole of the problem, and the numbers scale as `radius × pixels` exactly as
that description predicts: 5.99 → 12.44 → 33.89 ms for radii of 29 → 38 → 52 over
2.3 → 4.1 → 8.3 M pixels.
The extra cost at 1:1 is noise reduction and capture sharpening, whose radii are
properties of the sensor rather than of the frame. That is the correct behaviour
— it is why `RenderScale` has two units — and it is bounded by the kernel caps
those operations already declare.
---
## Reading this against §2's decision rule
§2: *"If M1 and M2 sit inside 16 ms at the 99th percentile, FR-DSP-2 is
rewritten rather than implemented … If they do not, the measurement tells us
which stage to tile."*
Both halves of the rule fire, on different stages, and the honest reading takes
both.
### FR-DSP-2 — rewrite it
For the fused pass the rule passes with a wide margin. Every point chain at
every size, fit and at 1:1, is inside 16 ms — the worst `TOTAL` is 8.23 ms and
the worst GPU figure is 4.52 ms. There is no viewport size on a desktop display
where recomputing the entire point chain over every visible pixel is a problem.
Two further reasons not to build the tile scheduler as written:
1. **Panning, which is the case ARCH §5.3's tile cache is designed for, gets no
benefit here.** Reusing already-valid tiles saves recomputation. Recomputing
the whole 4K viewport costs 4.5 ms, so a perfect tile cache could save at most
4.5 ms of a 16 ms budget, at the price of a cache keyed by
`(VersionId, tile, zoom, graph_hash_prefix)` that has to stay correct across
every parameter change in the graph. That is a large correctness surface
bought with a small number.
2. **It would make the actual problem worse.** The stage that misses the budget
is a convolution, and a tiled convolution reads a halo per tile. At a 52-pixel
radius, 256-pixel tiles would read (256+104)² instead of 256² — very nearly
*twice* the taps. Tiling is the wrong tool for the one stage that needs a
tool.
So FR-DSP-2 becomes what §2 said it actually is for this architecture: a
scheduling concern for export and thumbnailing, both of which already run off
the frame path. The interactive path does not tile.
### The stage that did need work — and it is not tiling
**Resolved.** The fix described below landed; the measurement is in
§[The reduced base, measured](#the-reduced-base-measured) at the foot of this
file, and `docs/technical-debt.md` TD-4 is closed. What follows is the
reasoning as it stood, kept because it is what the numbers above argue for and
because the tiling half of it is still live.
The measurement's real product is naming the stage. It is `local_contrast`, and
the fix is stated in that module's own documentation:
> The right optimisation is a base computed at reduced resolution, which needs a
> detail stage that can write a smaller target than it reads; that is a change to
> `crate::detail`, not to this file.
A Gaussian base at a quarter resolution is 1/16 the pixels at 1/4 the radius —
about 1/64 of the work — and the result is visually identical because a base at
σ = 26 px has no content above the quarter-resolution Nyquist to lose. That is a
change to two files with a bounded blast radius, and it is what the 34 ms buys
back. It should be tracked as its own item rather than smuggled in under a
requirement about tiles.
### FR-DSP-3 — the clause that should be narrowed
§3.3 proposes narrowing "when a full-resolution result is needed it is computed
asynchronously, and the proxy result remains on screen until it is ready" to
export and 1:1 zoom, or striking it.
**M2 says strike it.** The clause exists to hide the latency of a
full-resolution render behind a proxy. There is no such latency: the 1:1 view is
*faster* than the fit view on the fused path, and there is no second
full-resolution code path to be asynchronous about — `Framing::view` is the
whole mechanism. Export renders its own frames on a worker already. Keeping the
clause would mean building a progressive-swap machine to conceal a render that
completes in 1.4 ms.
### FR-DSP-4 — satisfied vacuously, on the fused path
§4 makes progressive refinement conditional on M1 failing. On the fused path M1
passes, so reduced-quality rendering during a drag would buy nothing and cost the
visible softness the requirement itself warns against.
The neighbourhood stage is the exception, and it is worth being precise: what
that stage needs is not *progressive* refinement — it is a permanently cheaper
base, computed at reduced resolution and correct at any moment the user stops.
"Render coarse while dragging, sharpen when it settles" would paper over the same
34 ms with a visible swap. Fix the stage.
---
## Which GPU, on a machine with more than one
**Measured 2026-08-29** on a laptop holding an Intel Iris Xe (RPL-P) and an AMD
RX 5700 XT, same binary, adapter forced with `VK_ICD_FILENAMES`.
The question was whether an integrated GPU is the better choice for this
application. The argument for it is good: a 24 MP frame is ~96 MB of RGBA, and
on a discrete card every upload and every export readback crosses PCIe, where
an iGPU shares memory with the CPU and crosses nothing. It also does not empty
a battery.
The compute says otherwise, and not marginally.
| 2560×1600, p99 | AMD RX 5700 XT | Intel Iris Xe |
|---|---|---|
| fused pass, `point` | 5.19 ms | 7.75 ms |
| fused pass, `all` | 11.70 ms | **66.42 ms** |
| M3 clarity, fit | 4.67 ms | **38.28 ms** |
| M3 all four, 1:1 | 6.44 ms | **57.67 ms** |
| 1920×1200, M3 clarity, fit | 2.35 ms | **19.99 ms** |
|---|---|---|
The fused colour pass is within a factor of 1.5 — it is one read and one write
per pixel, which an iGPU does perfectly well. The **neighbourhood stage is
5–8× slower**, and that is what decides it: clarity at 1920×1200 costs 20 ms on
the Iris Xe, so it leaves the budget on its own at the smallest size tested,
before anything else in the chain runs.
**So the default adapter preference stays `Performance`** (`dr_gpu::AdapterPreference`).
Two things this does *not* show, and neither is a reason to revisit the default
without measuring them:
- **It does not refute the transfer argument.** This harness renders from a
resident texture and never uploads or reads back, so the PCIe cost an iGPU
avoids does not appear in any column above. Import, export and the thumbnail
sweeps are transfer-heavy and compute-trivial, and may well go the other way
— but they are not what FR-DSP-3 bounds, and one device is opened at startup
and shared with the compositor, so there is currently no way to use a
different adapter for a different task.
- **It says nothing about power.** `Efficiency` remains offered
(`DARKROOM_GPU=integrated`) because a user on battery may rationally accept a
slower detail chain, and because someone whose discrete card has failed needs
a way to keep working.
## What is not measured here
Stated because §7 of [display-and-extension.md](display-and-extension.md) asks
for it, and because each of these could move the numbers.
- **Local adjustments.** The mask stack is a separate chain per layer and is not
in any row above. `render_masked` takes them and the fused shader addresses
them per layer, so a heavily masked edit costs more than `all`.
- **Spot repairs.** These add detail passes, and their cost is per spot.
- **Lens corrections.** Not part of `EditGraph::default_chain` — they are built
from a matched profile — so the `point` row does not include the warp chain.
- **Demosaic.** Once per photograph on a worker, not on the frame path.
- **Presentation.** The bench waits for the device to go idle inside the frame it
measures. A real compositor overlaps frames, so these figures are an upper
bound rather than an estimate.
- **One adapter.** A discrete laptop GPU. The Intel iGPU on the same machine, and
Android, will be slower — which is an argument for the conclusion rather than
against it: the stage with no headroom has none to lose.
## The CPU finding, which deserves its own item
`EditGraph::compose` costs 2.8–5.2 ms per frame on a full chain, at every
resolution, because it is resolution-independent: it assembles a WGSL string and
hashes it. On the `all` rows it is a third of what is left of the budget after
the GPU has taken its share, and at 1920 × 1200 it is larger than the entire
fused dispatch.
Nothing in this document's recommendations changes it, and it is the cheapest
remaining win. The generated *source* depends only on the structure of the graph
— that is what `structure_hash` already identifies, and it is precisely what does
not change while a slider is being dragged, which is why the pipeline cache in
`AdjustPass` does not recompile. The uniforms do change, but assembling them is a
handful of floats per operation. So caching the source string against the
structure hash and rebuilding only the uniforms would take these milliseconds to
approximately nothing, on the path that needs them most. Worth its own entry in
[technical-debt.md](technical-debt.md).
---
## The reduced base, measured
**Status:** Measured · 2026-08-29 · closes TD-4
`DetailPass` gained an `output_scale`, and clarity's base is now computed on a
grid a quarter the size on each axis — the change §M3 argued for above.
**Read this table on its own, not against the ones above.** It was taken on a
different adapter, so the absolute figures are not comparable with the RTX 3050
measurements this document is otherwise built from. What *is* comparable is the
before and the after, which were measured on the same machine, same card, same
release profile, minutes apart, with nothing between them but the change — the
baseline at `0407fb8` and the result at `bff95e2`.
| | |
|---|---|
| Adapter | AMD Radeon RX 5700 XT (RADV NAVI10) (Vulkan) |
| Source | 9504 × 6336 (60.2 MP, 482 MB as `rgba16f`) |
| Baseline | `0407fb8`, the branch's merge-base |
| Result | `bff95e2` |
| Date | 2026-08-29 |
### M3 — clarity alone, before and after
Both percentiles, because they disagree and the disagreement is the
interesting part.
| size | view | before p50 | after p50 | | before p99 | after p99 |
|------------:|:-----|-----------:|----------:|-----:|-----------:|----------:|
| 1920 × 1200 | fit | 2.30ms | 1.56ms | 1.5× | 3.94ms | 1.97ms |
| 1920 × 1200 | 1:1 | 2.88ms | 1.99ms | 1.4× | 3.07ms | 2.41ms |
| 2560 × 1600 | fit | 4.41ms | 1.95ms | 2.3× | 10.40ms | 2.37ms |
| 2560 × 1600 | 1:1 | 5.59ms | 3.32ms | 1.7× | 5.78ms | 3.85ms |
| 3840 × 2160 | fit | 10.94ms | 3.88ms | 2.8× | 25.05ms | 4.17ms |
| 3840 × 2160 | 1:1 | 13.31ms | 6.36ms | 2.1× | 27.60ms | 6.93ms |
**The honest headline is the p50 column: 2.8× at 4K.** An earlier draft of this
section led with the p99 ratio, which reads as 6.0× at the same size. That
number is not supported, and the reason it is not is worth recording rather
than quietly deleting.
The baseline run's `fit` rows have a p99/p50 spread of about 2.3×, while every
row of the after run sits between 1.07× and 1.26×. A stage whose cost is
`radius × pixels` has no reason to be bimodal, and the `fit` configuration is
the memory-bound one — it walks the whole 482 MB source on a stride, where
`1:1` reads a contiguous window. Something else was using the machine.
The cross-check settles it. §"Which GPU, on a machine with more than one"
above measured the *same baseline code on the same card* independently, and
reports M3 clarity, fit, 2560 × 1600 at **4.67 ms p99** — against the 10.40 ms
in the table here. Two measurements of one thing that differ by 2.2× mean the
noisier one is wrong, and it is this one.
So: the p50 ratios are the claim. The p99 improvement is real and larger, but
this run cannot say by how much, and a clean re-measurement on a quiet machine
is the way to find out.
What survives the caveat intact is the **shape** of the after column. Every
figure is inside the 16 ms budget with a p99 within 26% of its median, at every
size and both views — which is what a stage that is no longer the bottleneck
looks like, whatever the exact ratio to what it replaced.
### The one thing that is not a pure speed-up
**The declared halo is now quantised to multiples of `output_scale`.** The
kernel truncates at 2σ, and that rounding now happens on the reduced grid
before being multiplied back up:
| viewport | before | after |
|---|---:|---:|
| 1920 × 1200 | 29 px | 28 px |
| 2560 × 1600 | 38 px | 40 px |
| 3840 × 2160 | 52 px | 52 px |
At 2σ the Gaussian is already down to `e⁻²` of its peak, and
`crossing_the_reduction_threshold_does_not_change_the_picture` holds the
difference between a quarter-scale and a half-scale base to 0.03 stops of peak
excursion and 2% of frame reach. But it is a change in reach rather than only
in cost, it is what a tile scheduler would be handed, and it is worth knowing
that the number moved rather than discovering it later as a seam.
+484
View File
@@ -0,0 +1,484 @@
# Inference backends — the runtime and the model, chosen per device
Spec for **S16**, the build that puts the neural models on the hardware each device actually has.
Every model DarkRoom runs today — the three SCRFD detectors, the ArcFace embedder, YOLO26n-seg and
the ADE20K scene model — runs through `tract`, on one CPU core, on every platform. That was the
right first answer: D13's runtime half chose it because it costs no C dependency, and
[faces.md](faces.md) and [segmentation.md](segmentation.md) were written against it. It is also
between 20× and 300× slower than what the same devices can do, and this document is the record of
having measured that and the specification of what replaces it.
**It does not reopen D13's licensing half.** The weights are the same files under the same grant.
It does reopen the *runtime* half, and §3 is where it says how far.
---
## 1. What was measured · 2026-09-19
One benchmark, two builds of it: `ort`'s API over `tract` (exactly what the app links) and `ort`'s
API over a dynamically loaded ONNX Runtime with each execution provider in turn. Random 640×640
input, three warm-ups, the median of 15–30 timed runs, milliseconds. The same input every run, so
the numbers are compute cost and nothing else.
### 1.1 The tablet — Honor MagicPad 2, Snapdragon 8s Gen 3
SM8635: 1× Cortex-X4, 4× A720, 3× A520, Adreno 735, Hexagon V73. Android 16.
| Model | **tract** (today) | ORT CPU f32 | ORT CPU int8 | Adreno f32 ¹ | **Hexagon int8** ² |
|---|---|---|---|---|---|
| scrfd_500m (Fast) | 98 | 16 | 8 | 21 | **1.4** |
| scrfd_2.5g (Balanced) | 161 | 59 | 19 | ✗ | **1.8** |
| scrfd_10g (Thorough) | 489 | 204 | 48 | ✗ | **3.2** |
| arcface_mbf (per face) | 39 | 9 | 13 | 24 | 12 |
| yolo26n-seg | 287 | 94 | 38 | 54 | **5.1** |
| yolo26s-sem-ade20k | 408 | 154 | 46 | 47 | **3.7** |
¹ Qualcomm's own GPU backend (`libQnnGpu.so`, OpenCL). Fails on the two larger SCRFD graphs at an
`AveragePool` the layout transformer cannot place. ONNX Runtime's WebGPU provider also runs on this
GPU and was slower than the CPU on every model; it is not in the table because it is not a
candidate.
² QNN's HTP backend. The Hexagon **refuses float32 and float16 tensors** in this ORT 1.29 + QNN
2.42 pairing (error 3110 on every node, with `enable_htp_fp16_precision` set or not); int8 QDQ
graphs run with 99.6% of nodes on the NPU — 1718 of 1725 for SCRFD-500m, the remainder being the
quantise/dequantise at the graph's edges — verified from the partition log, not inferred from the
timing.
Also tried and rejected: **NNAPI** — the device registers no neural-networks HAL at all, so the
provider has nothing to talk to; Google deprecated it in Android 15 and Qualcomm stopped shipping
drivers for it. **XNNPACK** — slower than ORT's default CPU kernels on every model that loaded, and
aborts inside its partitioner on the SCRFD graphs.
### 1.2 The desktop — RTX 3050 Laptop, Raptor Lake, 20 threads
| Model | **tract** (today) | ORT CPU f32 | ORT CPU int8 | CUDA f32 | CUDA fp16 | TensorRT f32 | TensorRT fp16 | TensorRT int8 |
|---|---|---|---|---|---|---|---|---|
| scrfd_500m | 104 | 12 | 8 | 5.6 | 3.7 | 2.5 | **1.8** | ✗ ⁴ |
| scrfd_2.5g | 162 | 27 | 12 | 6.2 | 5.3 | 2.9 | **1.9** | ✗ ⁴ |
| scrfd_10g | 514 | 99 | 32 | 14.5 | 9.3 | 7.6 | **3.3** | ✗ ⁴ |
| arcface_mbf | 45 | 15 | 18 | 1.1 | 0.8 | 0.9 | 0.7 | ✗ ⁴ |
| yolo26n-seg | 307 | 60 | 38 | 9.3 | ✗ ³ | 6.8 | **5.5** | ✗ ⁴ |
| yolo26s-sem-ade20k | 395 | 61 | 31 | 10.7 | ✗ ³ | 7.9 | **3.4** | ✗ ⁴ |
³ The offline fp16 conversion (`onnxconverter-common`) left a mixed-type node the CUDA provider
rejects. TensorRT converts to fp16 itself at engine build and does not have this problem, which is
one reason it is the target and the CUDA provider is the fallback.
⁴ TensorRT refuses the QDQ form ONNX Runtime's quantiser writes for the Hexagon (uint8
activations); it wants symmetric int8. Not pursued: fp16 needs no quantisation, no calibration and
no accuracy gate, and it is already 30–60× tract.
CUDA int8 is deliberately absent: the CUDA provider has no int8 kernels and runs a QDQ graph by
dequantising it, which measured *slower* than f32 (7.0 vs 5.6 ms on scrfd_500m). Int8 on NVIDIA is
TensorRT's job.
**TensorRT's first load is 16–116 s per model in f32 and 35–290 s in fp16** (yolo26n-seg the
worst: nearly five minutes), because it is compiling an engine for this exact GPU. The engine caches to disk and the second load is milliseconds. That
number is what §6 is designed around.
### 1.3 What the numbers say
- **`tract` is single-threaded.** The tablet's one X4 core and one Raptor Lake core give the same
tract numbers. Replacing it with ONNX Runtime's CPU provider, *no accelerator involved*, is 3–6×
on the tablet and 8–10× on the desktop. That is the floor, and it is available on every platform
the app builds for.
- **The Hexagon is the standout.** A 6 W NPU running int8 beats a discrete RTX 3050 running f32 on
five of six models. A whole-library face index on the tablet goes from ~100 ms + 39 ms per face
to ~1.4 ms + 12 ms per face, and the "Thorough" detector — 3× the cost of "Fast" today — becomes
free. Its price is that the models must be **quantised to int8**, which is an accuracy question
§5 has to answer before it is believed.
- **The embedder does not gain from either accelerator.** 112×112 input, per-op overhead
dominates; it is 9 ms on the tablet's CPU and 12 ms on its NPU. It stays float, which §7 turns
from a performance footnote into a correctness rule.
- **On NVIDIA, TensorRT fp16 ≈ 3× the CUDA provider**, and the CUDA provider ≈ 2× the
multi-threaded CPU; at fp16 the detectors are 1.8–3.3 ms with no quantisation at all. Both leave the twenty cores free for decoding during a batch index, which the table does not
show and which matters more than the ratio.
---
## 2. The shape of the answer
A **ladder per platform**, walked at start-up, with the first rung that builds a real session
winning:
| Platform | 1st | 2nd | 3rd | Floor |
|---|---|---|---|---|
| Android, Qualcomm with a Hexagon the shipped QNN skel covers (V68–V81) | QNN HTP, int8 model | ORT CPU, f32 model | — | tract |
| Android, any other SoC | ORT CPU, f32 | — | — | tract |
| Linux / Windows, NVIDIA GPU | TensorRT, f32 model, fp16 engine | CUDA provider, f32 | ORT CPU, f32 | tract |
| Linux / Windows, no NVIDIA | ORT CPU, f32 | — | — | tract |
| macOS ⁵ | ORT CPU, f32 | — | — | tract |
⁵ CoreML is the obvious rung and is unmeasured; it is listed so its absence is a gap and not an
oversight.
Deliberately **not** on any ladder, with the measurement that excluded each: NNAPI (no driver),
XNNPACK (slower than CPU, aborts on SCRFD), WebGPU (slower than CPU), the Adreno through QNN (works,
but never where the Hexagon does not also), CUDA int8 (slower than CUDA f32). A rung is added to this
table by a measurement on this page, not by a provider existing.
Two things the ladder is *not*: it is not a per-model choice — one backend serves every model on a
device, because §7's identity rule needs the detector and embedder on the same runtime for the
same reason `faces.model_id` pairs them; and it is not a per-account choice — it is a property of
the hardware, like `shared_face_models_dir` is, and it lives beside it.
---
## 3. The dependency policy, and how far this reopens it
D13 chose `ort` over `tract` because `alternative-backend` made ONNX Runtime's *API* available with
none of its *C*. Every rung above the floor needs the C++ ONNX Runtime and, for the two that matter
most, vendor libraries on top: Qualcomm's QNN runtime (~60 MB for one Hexagon generation; it is
per-SoC) and NVIDIA's TensorRT plus cuDNN (~600 MB with the CUDA libraries, and cuDNN's major
version must match what ONNX Runtime was built against — the Arch package on the reference desktop
was unusable for exactly that reason).
The policy protected the **build**: no C to cross-compile under the NDK, no toolchain to keep in
step. This document keeps that intact, and the mechanism is the one thing about `ort` that makes it
possible:
**`ort::set_api` accepts any `OrtApi` table.** With `alternative-backend` on, `ort` links nothing
and asks for the table once per process. The application can `dlopen` a `libonnxruntime.so` it
finds on disk, call `OrtGetApiBase()->GetApi(version)` and hand that table over; or, if there is no
such file, hand over `ort_tract::api()`. The Rust build is identical in both cases — pure Rust,
`cargo build --target aarch64-linux-android` sees the same dependency graph it sees today. What
changes is that the runtime is a **file the package installs**, next to the models, and the app
looks for it at start-up.
Consequences that follow and are accepted:
- **The runtime is chosen once per process**, because `set_api` is once per process. The ladder
in §2 is walked at start-up and the result is what every session in that process uses. There is
no "tract for this model, ORT for that one", and there is no falling back to tract *after* ONNX
Runtime has loaded — but there does not need to be: once the library loads, its CPU provider is
always there, and every fallback the ladder needs is between providers *inside* it.
- **Feature flags stay as they are.** `dr-face`'s `inference` and `dr-segment`'s `semantic`
continue to mean "compiled against `ort`'s API"; nothing at build time knows or cares which table
will be supplied. The one addition is a `native-probe` feature on the new crate (§8) that pulls in
`libloading`, which is pure Rust and already in the tree via `wgpu`.
- **The packagers ship the runtime, not the build.** The Arch package, the Flatpak manifest, the
NSIS installer and `assemble-apk.sh` each gain the ONNX Runtime library for their platform, and
the Android and NVIDIA variants gain the vendor libraries — each under the licence the packager
reads first (§3.1). A package without them is not broken; it is the tract build, and it says so
on the about screen.
- **The NDK problem does not come back.** `libonnxruntime.so` for Android is a prebuilt from
Maven (`com.microsoft.onnxruntime:onnxruntime-android-qnn`), extracted by `assemble-apk.sh` into
`jniLibs/` the way the models are bundled as assets today. Nothing compiles it.
### 3.1 Licences the packagers read before shipping a runtime
Written down now, because [segmentation.md §7](segmentation.md) established that reading the grant
is cheaper than discovering it at packaging time.
| Component | Licence | Redistributable in a self-distributed package? |
|---|---|---|
| ONNX Runtime | MIT | Yes |
| Qualcomm QNN runtime (`com.qualcomm.qti:qnn-runtime` on Maven) | Qualcomm AI Engine Direct SDK licence — proprietary, redistribution permitted for applications using it | Yes for the APK, with the licence text shipped; not for a source distribution. **To be read in full, not summarised from memory, before the APK gains it.** |
| CUDA runtime, cuDNN, TensorRT | NVIDIA EULAs — redistributable with an application, with the licence text, not modifiable | Yes for a package that bundles them. 600 MB. The alternative is to load them from the user's system install if present and skip the rung otherwise — which is what §4's probe does anyway. |
The position this takes: the **NVIDIA libraries are not bundled**. The desktop package probes for a
system CUDA/TensorRT install and uses it if it is version-compatible; a desktop without one runs on
ORT CPU, which is still 8–10× today. Bundling 600 MB for a rung that is 2× again is not a trade
worth making unmeasured, and it can be revisited by a measurement on a batch index. The **QNN
runtime is bundled** in the APK, because the Hexagon is the difference between a tablet that
indexes a library overnight and one that does it over lunch, and the package is 60 MB larger for
it.
Both positions are D13 territory and are recorded there (§12).
---
## 4. Selection — the probe, its cache, and what it may not do
**A rung is chosen by building a real session on it, not by asking whether it exists.** Both
failure modes that are not "the provider is absent" were hit on 2026-09-19: a driver in a wedged
state where the provider registered and the session then failed, and a provider that registered,
took the graph, and rejected every node at partition time. The probe therefore:
1. Loads the runtime library (§3), or falls to tract and stops.
2. Times the **smallest detector** on the CPU provider first — the floor. Then, for each rung
in this platform's ladder, in order: builds a session for the same model on that provider
with `error_on_failure`, runs it once on a fixed input, and times three more runs. **The rung
is taken only if its median beats the floor.** That one measurement is the proof the provider
took the graph: one that silently hands the work to the CPU is the CPU rung with extra
overhead, slower than the floor, and rejected. (ONNX Runtime's
`session.disable_cpu_ep_fallback` was the first draft of this proof and refuses the Hexagon
over the ten quantise/dequantise nodes at the graph's edges that QNN declines by policy.)
3. Records the outcome — rung, runtime version, provider version, device identity (GPU name and
compute capability; SoC model and Hexagon arch), and the models' content hashes — to a small
file beside `shared_face_models_dir`. The next start-up trusts the file **unless** any of those
inputs changed, in which case it probes again. A driver update, a runtime update, a new model
file: each invalidates the cache by construction, and none needs a "reset backend" button.
What the probe may not do:
- **Block the first frame.** It runs on the same background as `install_bundled_models` and for
the same reason: a TensorRT probe can take thirty seconds cold, and a tablet that stalls that long
is an ANR. Until it reports, every model request is answered by the floor the runtime supports
(ORT CPU if the library loaded, tract otherwise), and a job that started on the floor finishes
on it — a backend does not change under a running index.
- **Retry a rung that failed within a session.** A failed probe is cached as a failure with the
same inputs; the rung is tried again when an input changes. Otherwise a wedged driver means a
thirty-second stall on every launch.
- **Choose for the user without saying so.** Settings gains one row, *Inference backend*, showing
what was chosen and why in one line ("Hexagon NPU · int8 · QNN 2.42"; "CPU · ONNX Runtime 1.30 ·
TensorRT probe failed: cuDNN 8 required"), with an override to force any lower rung. The about
screen carries the same line beside the model names NFR-SEC-5 already puts there.
---
## 5. Model variants, and who makes them
Every model exists in one **canonical** form — the f32 ONNX file the app ships or the user supplies
today — and, where a rung needs it, a **derived** form. The ladder's rungs are specified in terms
of which form they load:
| Form | Who produces it | When | Needed by |
|---|---|---|---|
| f32 ONNX, shape-fixed, **opset ≥ 13** | `tools/fix-face-model-shapes.sh`, `tools/export-seg-model.sh` | Release time, once | Every rung except Hexagon |
| int8 QDQ ONNX, per-channel, uint8 activations | `tools/quantise-models.sh` (new) | Release time, once, **calibrated on real photographs** | Hexagon |
| TensorRT engine (`.engine`, per GPU architecture and TensorRT version) | The app, from the f32 file | First run on that device, in the background | TensorRT rung |
| QNN context binary | The app, from the int8 file | First run on that device, in the background | Hexagon rung |
Two rules.
**Quantisation is a release-time step, not a device-time one.** The int8 files that produced §1's
numbers were calibrated on random noise, which is enough to time and worthless to trust. A real
int8 detector is calibrated on a few hundred real photographs and then measured against the f32
detector on the reference library by [faces.md §12.3](faces.md)'s method — faces found, per size
band, per detector — before it ships. That needs the reference library and a person reading the
result, and it happens once per model release, in `tools/`, beside the shape-fixing it already
depends on. The device never quantises anything.
The SCRFD and ArcFace files are **opset 11** as InsightFace exported them, and per-channel QDQ needs
13; `tools/fix-face-model-shapes.sh` gains an opset upgrade to 17 (`onnx.version_converter`,
`ir_version` 8), which `tract` has been verified to load and which every provider on this page
prefers. That is a change to the canonical file and so a change to the shipped models, and it
happens in the same model release as the int8 files.
**Compilation is a device-time step, and it is cached.** A TensorRT engine is specific to the GPU
it was built on and the TensorRT that built it; a QNN context binary is specific to the Hexagon
generation. Neither can ship. Both are built by the app the first time that rung is selected, in
the background (§6), and written beside the probe cache keyed by the same inputs. They are
**derived, disposable, and regenerable**: deleting the cache directory costs the next launch a
rebuild and nothing else, and the directory is excluded from anything that syncs (it is a peer of
`thumbs`, not of the catalog).
---
## 6. First run — building engines without the user waiting for them
The sequence on a device where a compiling rung (TensorRT, Hexagon) is selected:
1. **Launch.** The runtime loads; the probe (§4) starts in the background; the app serves every
model request from the floor. Face indexing, segmentation and scene grading all work, at
today's speed or better (ORT CPU).
2. **Probe reports** — say, TensorRT. The compiling rung is now *selected* but has **no engines**.
Model requests continue on the fallback rung below it (CUDA provider for TensorRT; ORT CPU for
Hexagon), which needs no compilation and is already faster than the floor.
3. **Engines build**, one model at a time, on a single low-priority background thread, smallest
model first so the detector — the one that runs per image — is ready soonest. On the reference
desktop that is ~1 minute for the first detector and ~10 minutes for all six at fp16; on the tablet the QNN
context binaries take 0.8–1.7 s each and the whole set is ready before the user has opened a
library. Each engine is written to a temporary name and renamed into place, so a request never
sees a half-written file.
4. **Requests move up as engines land.** A model whose engine exists loads it on the selected
rung; one whose engine is still building loads on the fallback. **A running job does not
switch** — an index that started on the CUDA provider finishes on it — because §7 needs one
`model_id` per job, and because a job is the wrong granularity for surprise.
5. On Android, the build runs only while the app is in the foreground and the device is not in
battery saver (NFR-RES-3): a context binary takes a second, so this costs nothing, and the rule
exists for the day a model takes longer.
Settings shows a one-line progress row while engines build ("Preparing GPU engines · 3 of 6") and
nothing when they are done. A build failure demotes the rung — it is recorded in the probe cache
as a failure with the model hash as an input, so a corrected model file retries it — and the app
carries on one rung down, saying so in the same row.
---
## 7. Identity — what changes `model_id` and what may not
`faces.model_id` exists so that two libraries indexed with different networks are never compared
as if they were one ([catalog.md §10.1](catalog.md); the trap is written up in
[faces.md §14](faces.md)). A backend that changes what a network *computes* is a different network
and must be a different `model_id`; one that changes only *where* it computes it must not be.
**The detector.** An int8 SCRFD finds a different set of faces from the f32 one — that is what
§5's acceptance measures — so **the numeric form is part of the detector's identity**:
`scrfd_500m` and `scrfd_500m_i8` are two detectors in `model_id`, and a library indexed on the
tablet's Hexagon and continued on the desktop is two populations, which a re-index on either side
reconciles the same way a switch from Fast to Thorough does today. That is acceptable because it is
already the rule for the detector and because §5 is the gate on whether the int8 form is close
enough to be *offered* at all. f32 on tract, ORT CPU, CUDA and TensorRT-f32 are one identity: the
same graph, the same arithmetic, differences at the last bit.
**The embedder** is where comparability across devices is the whole point, and it is the one
model that no accelerator helps (§1.3). So: **the embedder runs in f32 on every rung.** On TensorRT
that means the embedder's engine is built without fp16 while the detector's is built with it; on
the Hexagon it means the embedder is not on the NPU at all — it runs on the ORT CPU rung at 9 ms,
and the ladder's "one backend per device" is, precisely, one backend *per model role*, with the
embedder pinned. A `w600k_mbf` embedding from any device is comparable with one from any other,
which is the property the identity system, the calibration and the cross-device merge all rest
on, and it is not for sale for 3 ms.
If S16 wants fp16 for the embedder later, the gate is written now: over the reference library's
faces, the cosine between the f32 and fp16 embedding of the same crop exceeds 0.999 for 99.9% of
faces and the calibration's fitted threshold moves by less than its own confidence interval
([faces.md §8.3](faces.md)). Until measured, f32.
**Segmentation and the scene model** carry no identity across devices — their outputs are
recomputed per image and never stored beyond the cache — so they take whatever the rung offers,
int8 included, subject to §10's own acceptance.
---
## 8. Crate shape — `core/dr-inference-engine`
The seam is the same shape as [storage.md](storage.md)'s: a small crate below the consumers that is
the **only** place naming a provider, a library file or a vendor, with the consumers reduced to
"give me a session for these bytes in this role".
```
core/dr-inference-engine
src/lib.rs Runtime (Tract | Onnx { lib, version }), Backend (rung), Role (Detector | Embedder | Segmenter)
src/probe.rs §4 — the ladder per platform, the session-build probe, the cache file
src/engines.rs §6 — background compilation, the cache directory, progress
src/session.rs open(role, bytes) -> ort::Session, applying the rung and the role's precision rule
src/api.rs the one unsafe block: dlopen libonnxruntime, fetch OrtApi, ort::set_api — or ort_tract::api()
```
- `dr-face` and `dr-segment` **delete** their private `install_backend` and their direct
`Session::builder()` calls and take an `&dr_inference_engine::Sessions` where they take model bytes today.
Their tests keep `tract` — `dr_inference_engine::Sessions::tract()` is a constructor and the test-only path.
- `dr-inference-engine` depends on `ort` with the same workspace features as today plus `cuda`, `tensorrt`,
`qnn`: those features add option builders, not linking, under `alternative-backend`. **Verified
for the QNN, CUDA and TensorRT builders on 2026-09-19** — they go through the API table's generic
`SessionOptionsAppendExecutionProvider*`. The NNAPI builder resolves a symbol directly and would
not; it is not needed and is not enabled.
- `dr-ui` owns the settings row, the about-screen line and the progress row; it holds one
`Sessions` per process, created at launch, and passes it down. `dr_ui::library` gains
`inference_cache_dir()` beside `shared_face_models_dir()`, on the same account-independent
footing and for the same reason.
- The Android entry point's `install_bundled_models` also extracts nothing new: `jniLibs/` is
loaded by the system loader, and `dr-inference-engine` on Android looks for `libonnxruntime.so` through
`dlopen` by bare name first, which resolves to the APK's copy, before any directory.
---
## 9. Threads and memory
- ONNX Runtime's intra-op pool is sized to the physical cores minus two on desktop and to the
performance cores on Android (the X4 and the A720s; the A520s are for the compositor). tract's
single thread today is the reason a batch index leaves nineteen cores idle; ORT CPU with the
pool is the reason it will not. One session per model per process; `Session::run` is
`&mut self`-free in `ort` and internally serialised, and the index job is the only caller.
- A TensorRT session pins GPU memory for its workspace; the builder is capped at 512 MB on the
reference 6 GB card and the cap is a setting, because the develop view's tiles share the card
(NFR-RES-2). The engine cache on disk is bounded by the model set — six engines, ~80 MB — and
needs no LRU.
- The Hexagon rung sets QNN's performance mode to `Burst` for the duration of an index job and
`Default` otherwise; a 5 ms detector does not need the NPU clocked up between images.
- The probe (§4) and the engine build (§6) run on one dedicated low-priority thread. They never
share the index job's pool: a probe that competes with the job it is meant to speed up is the
frame-budget trap in a new coat.
---
## 10. What S16 measures
In order, with the gate each is:
| # | Question | Gate |
|---|---|---|
| M1 | Does one binary carry both tables? `dlopen` + `set_api` on Linux, Windows and Android; `ort_tract::api()` when the file is absent. | Go / no-go for §3. If `set_api` cannot take a dynamically fetched table on some platform, that platform ships two binaries, and the cost is stated. |
| M2 | Do the int8 SCRFD detectors, **calibrated on real photographs**, find the faces? faces.md §12.3's method over the reference library, per size band, against f32. | Ship the int8 form for a detector only if it finds ≥ 97% of the f32 detector's faces above 40 px and the difference is not concentrated in one band. Otherwise that detector's Hexagon rung is ORT CPU int8-free, and the table in §1.1 says what that costs. |
| M3 | Does the embedder on ORT CPU beside a detector on the Hexagon (§7) produce embeddings within the f32 gate? | It must — same graph, same arithmetic. This is a check that the plumbing did not quantise it by accident. |
| M4 | Is a batch index on the tablet and on the desktop faster by the ratio §1 predicts, end to end, decode included? | The face index over the reference library (18,143 faces): report wall-clock on tract, on the floor and on the selected rung, and where the time went. The prediction is that decode becomes the bottleneck on both; if it does not, say why. |
| M5 | Does the first-run sequence (§6) hold: nothing blocks the first frame, engines land, requests move up, a running job does not switch? | Observed on both devices with the app's own progress row, and with the cache directory deleted between runs. |
| M6 | What does the APK weigh with the QNN runtime, and does a non-Qualcomm Android device (any one) still launch and index on the floor? | Size reported; launch verified on one non-Qualcomm device or an emulator. |
| M7 | The segmentation and scene models at int8 on the Hexagon: does the mask boundary move? segmentation.md §6's IoU against f32 over its corpus. | Ship int8 for a model only above the IoU floor that document set for arm B. |
M1 and M2 are the ones the rest is conditional on, and M2 is the one that needs a person.
### 10.1 M2 result · 2026-09-19
Each int8 detector against its own f32 form, over 400 proxies evenly spaced through the reference
library, on ONNX Runtime's CPU provider (the int8 graph is the same file the Hexagon loads;
`ui/dr-ui/examples/face_detectors`). Calibrated on 64 proxies from the same library, disjoint
from the 400.
| Detector | f32 faces | int8 faces | both | int8 only | f32 only | found ≥ 32 px | found, all sizes |
|---|---|---|---|---|---|---|---|
| scrfd_500m | 1342 | 1319 | 1287 | 32 | 55 | 95.6% | 95.9% |
| scrfd_2.5g | 1525 | 1482 | 1478 | 4 | 47 | 96.1% | 96.9% |
| scrfd_10g | 1769 | 1760 | 1749 | 11 | 20 | 100% | 98.9% |
The 10g form clears the 97% gate; 500m and 2.5g sit one point under it. What they lose is
specific: the faces in the "f32 only" column have a **median confidence of 0.52** against a
threshold of 0.50 — detections the f32 graph itself barely made, that int8 rounding drops to the
other side of the line — and the extra faces int8 finds are the same kind (median 0.51–0.52).
Not a size-band failure: the losses are spread across bands in proportion. Shipped as they are,
with the number on record; a threshold of 0.48 for the int8 forms would recover most of the
margin, and is the first thing to try if a library's count on the tablet reads low.
Two things the calibration taught, both in `tools/quantise-models.py`: the calibration set has
to contain faces (a first attempt on landscape photographs produced a graph that found nothing —
the score head's ranges had never seen the face regime), and ONNX Runtime's own strided and
moving-average calibration modes both measurably degrade the result on these graphs, while
driving the calibrator in chunks by hand reproduces the plain min/max ranges exactly.
---
## 11. Order
1. **`dr-inference-engine` with the two tables and the floor** — `set_api` from a dlopened runtime, tract
otherwise, ORT CPU as the only rung. Consumers moved over; tests unchanged. This alone is the
3–10× and is the build most of the value sits in. M1.
2. **The probe and its cache** (§4), with the settings row and the about line. Still CPU-only;
the ladder has one rung. M5's first half.
3. **`tools/quantise-models.sh`** and the opset upgrade; the int8 detectors calibrated and
measured. M2, M7. This is the step with a person in it and it runs in parallel with 4.
4. **The Hexagon rung**, the QNN runtime in the APK, the context-binary cache. M3, M6.
5. **The TensorRT and CUDA rungs** on desktop, the engine cache, the first-run sequence. M5's
second half.
6. **M4** last, on both devices, and the number goes in this document.
The Windows installer and the Flatpak manifest are touched in steps 1 and 5 only, and only to add
a file each; the Arch package likewise.
---
## 12. Register entries
**FR-INF-1 — Runtime selection.** On launch the application shall determine, per device and
without blocking the first frame, the fastest inference backend that can build and run a session
for the shipped models, by attempting it; shall record and reuse that determination until the
runtime, driver, hardware or models change; and shall display the backend in use in Settings and
on the about screen. *Acceptance:* §10 M1 and M5.
**FR-INF-2 — Derived engines.** Backends that require device-specific compilation shall compile in
the background after selection, shall serve requests from the next lower backend until each engine
is ready, and shall not change the backend of a job in progress. *Acceptance:* M5.
**FR-INF-3 — Model forms.** Quantised model forms are produced at release time from real
calibration data and are shipped only when they meet §10's accuracy gates against the canonical
form; the application never quantises on the device. *Acceptance:* M2, M7.
**NFR-INF-1 — Embedding comparability.** Face embeddings shall be computed at a precision whose
deviation from the f32 reference is within §7's gate, on every backend, so that embeddings from any
device are comparable. *Acceptance:* M3.
**D13 — updated.** The runtime half is reopened to the extent of §3: the Rust build stays C-free
under `alternative-backend`; packages may install a dynamically loaded ONNX Runtime and, per §3.1,
the Qualcomm QNN runtime; the NVIDIA libraries are not bundled. The licensing half is unchanged.
---
## 13. Requirements touched
FR-CULL-8 (indexing time is what this exists to change), FR-CULL-9 (NFR-INF-1 is the guard on its
calibration), NFR-RES-2 (TensorRT workspace against the develop view's tiles), NFR-RES-3
(background compilation on Android), NFR-SEC-5 (the about screen's line), NFR-COMPAT-2 (each
channel gains a runtime file), FR-PLAT-AND-1 (unchanged; the runtime lives in the APK, not in
storage), ARCH §6.1 (the once-per-image budget this was sized against no longer binds; what could
run per frame is a separate question this document does not open).
+780
View File
@@ -0,0 +1,780 @@
# Editing a mask
**Status:** Draft · 2026-09-06
**Companion to:** [requirements.md](requirements.md) §3.3 FR-DEV-3, FR-DEV-10 ·
[segmentation.md](segmentation.md) · [architecture.md](architecture.md) §5.4 ·
[spot-removal.md](spot-removal.md) (the pattern a canvas tool follows here)
The model finds a subject in a second and the photographer cannot then change
it by a single pixel. This document specifies the tools that close that gap —
paint, erase, push, and combining one selection with another — and the
interface they are driven through.
---
## 1. The gap, precisely
The core is further along than the interface, and it is worth being exact
about which half is missing, because it changes the size of the work.
**Painting exists everywhere except where a finger is.**
[`MaskSource::Brush`](../../core/dr-pipeline/src/mask.rs), [`Stroke`], the
simplification and the point budgets, the sidecar's `stroke = …` line and its
parser, the GPU's per-stroke bounding-box draw with add and erase blend
states — all of it is written, tested, and reachable from no control in the
application. [`toolrail.slint:138`](../../ui/dr-ui/ui/toolrail.slint#L138) says so
in as many words: *"what is missing is the canvas interaction"*.
**A mask has exactly one source.** A layer is one `MaskSource` and a shaping
of its edge. There is no way to say *this subject **and** that one*, *the sky
**except** the branches*, or *the subject **only where** it is bright* — the
last being an intersection of a model mask with a range mask (FR-DEV-10), two
features that already exist and cannot meet.
**The edge moves all at once or not at all.** `Morphology` grows or shrinks
the *whole* boundary. When the model's coverage stops two pixels inside the
shoulder and leaks four pixels into the hair, no global number fixes both, and
that is the ordinary case rather than a corner one.
**And you cannot see the mask.** *(Built — see §6.)* The overlay on the canvas
was [`overlay_rgba`](../../ui/dr-ui/src/segmentation.rs) and nothing else — a
CPU-built false-colour picture of *what the model detected*, at proxy
resolution. It is not the layer's alpha: it knows nothing of the layer's
feather, its falloff, its morphology, its invert, or its opacity. Nobody can
refine an edge they are not being shown.
That one turned out to be load-bearing for the other three rather than the
last of four. With no way to see a mask, choosing a category produced a layer
whose extent was invisible and whose adjustment had not been touched yet — so
the correct behaviour and the broken one look identical, and "the segmentation
does not make masks" is what it reads as from the outside.
Those four are one feature, and this is its specification.
## 2. Non-goals
- **Not pixel layers.** §1.3 of the requirements excludes them. Everything
below stores geometry and parameters; no rasterised mask is ever written to
a file, and no mask ever exists in CPU memory (ARCH §5.4).
- **Not a second segmentation UI.** Choosing *which* subject or category the
model offers stays exactly as it is. This is about what happens after.
- **Not per-layer neighbourhood adjustments.** `layer_chain()` excludes detail
and optics operations for reasons that have not changed.
- **Not automatic refinement.** No "improve this mask" button that silently
redraws a boundary the photographer approved. The tools here move an edge
only where a hand is.
- **Not a mask library.** Saving a mask and applying it to another photograph
is a reasonable later feature and depends on none of this.
## 3. The five tools
Named as a photographer would say them, because these are the words the
interface will use.
| Tool | What it does | New machinery |
|------|--------------|---------------|
| **Paint** | Adds to the mask under the brush | none — `brush_add` exists |
| **Erase** | Takes away under the brush | none — `brush_erase` exists |
| **Push** | Drags the boundary itself: outward fills behind it, inward empties behind it | a warp pass |
| **Combine** | Joins another selection to this mask — add, subtract, or keep only the overlap | three blend states |
| **Show** | Draws the mask that actually results, live | two uniforms in the composed shader |
Push is the one worth defining carefully, because "tug the edge" can mean two
different operations and only one of them is right here — §5.3.
## 4. The model: a mask is a stack of parts
### 4.1 A part
The change that carries all four gaps at once is that a layer stops holding a
source and starts holding an ordered list of them.
```rust
/// One selection joined into a layer's mask.
pub struct MaskPart {
/// Stable identity, for the sidecar and for merge (FR-NC-9).
pub id: String,
/// How this part enters the mask built so far. Ignored on the first part,
/// which *is* the mask so far.
pub join: Join,
pub source: MaskSource,
/// Shaping, moved down from the layer: two parts of one mask routinely
/// want different edges — a model's soft coverage joined to a hand-painted
/// correction that must be exactly where it was painted.
pub invert: bool,
pub feather: f32,
pub falloff: Falloff,
pub morphology: Morphology,
pub morph_radius: f32,
pub refine: f32,
/// The model raster behind a Subject or Category source. Per part now,
/// for the same reason it was per layer: it materialises *this* selection.
pub coverage: Option<Arc<Coverage>>,
}
/// How a part joins the mask before it.
pub enum Join {
/// Everything either has. The default, and what "add a brush" means.
Union,
/// What the mask had, minus this. "Subtract".
Subtract,
/// Only where both agree. Where a subject meets a luminance band.
Intersect,
/// Not a set operation: this part's strokes *move* the mask under it.
/// Only ever a painted part carrying push strokes — see §5.3.
Warp,
}
```
A layer's mask is then `parts.fold(empty, join)`, and every one of §1's gaps
becomes an ordinary use of it:
- **Paint on an auto mask** — append a `Union` part whose source is painted.
- **Erase from one** — strokes inside that part carry `Erase`, or the part
itself is a `Subtract`. Both work; the panel offers the first, because a
photographer alternating add and erase over one area is drawing one
correction, not two.
- **Merge two selections** — two `Union` parts.
- **Cut one out of another** — a `Subtract` part.
- **The bright part of the sky** — a `Category` part, then an `Intersect`
luminance part.
- **Tug an edge** — a `Warp` part.
**Why a list on the layer rather than a boolean tree.** A tree expresses more
and no photographer has ever wanted the extra. A flat ordered fold is what
Lightroom, Capture One and darktable all present, it reads top to bottom in a
panel with no parentheses to draw, and — decisively here — it merges under
FR-NC-9 as a sequence of independently-keyed blocks, where a tree would merge
as a shape whose two halves can be individually won by different devices and
recombined into something neither ever had.
### 4.2 A stroke gains a mode
```rust
pub enum StrokeMode { Add, Erase, Push }
pub struct Stroke {
pub mode: StrokeMode, // was: erase: bool
pub radius: f32,
pub hardness: f32,
pub flow: f32,
/// How strongly the deposit clings to the picture's own edges: 0 paints
/// anywhere, 1 paints only what matches the colour under the point the
/// stroke began. See §5.6.
pub cling: f32,
pub points: Vec<(f32, f32)>,
}
```
`erase: bool` becomes a three-valued mode. Everything else about `Stroke` —
the grid snapping, `simplify`, `MIN_STEP_FRACTION`, `MAX_STROKE_POINTS` and
its continuation rule — is unchanged and applies to a push stroke exactly as
it does to a painted one.
### 4.3 What the layer keeps
```rust
pub struct MaskLayer {
pub id: String,
pub name: String,
pub enabled: bool,
/// Invert and opacity stay here: they are the two uniforms the generated
/// shader already reads per layer (LAYER_UNIFORM_FIELDS), they apply to
/// the finished mask, and moving them would change the composed shader.
pub invert: bool,
pub opacity: f32,
/// Never empty. `parts[0]` is the base.
pub parts: Vec<MaskPart>,
pub ops: Vec<Box<dyn Operation>>,
}
```
`MaskSource::Brush { strokes }` becomes `MaskSource::Painted { strokes }` —
the same variant under a name that no longer implies it is the only thing a
brush can touch. `MaskLayer::begin_stroke` / `extend_stroke` / `end_stroke`
keep their signatures and route to the layer's *active* part, which the panel
sets; that is the only change their callers see.
**Budgets.** `MAX_LAYER_POINTS` (4096) becomes a budget across all of a
layer's parts rather than one part's, so a layer's worst-case sidecar size and
worst-case rasterisation cost are unchanged. `MAX_PARTS = 8` per layer, on the
argument `MAX_LAYERS` makes: past that it is not a selection any more, and a
bound the panel can show is better than one a file discovers.
### 4.4 The sidecar, and why no existing file changes
Blocks are `[mask <version> <layer-id>]` today. Parts are their own blocks:
```
[mask default m1]
name = Sky
source = category
signature = 4711
category = sky
feather = 0.004
falloff = smooth
opacity = 1
[part default m1 p2]
join = subtract
source = painted
stroke = add 0.05 0.5 1 0.6 0.31,0.42 0.33,0.44 …
```
Three compatibility rules, and together they mean **every sidecar written by
every build so far loads into this one unchanged, and a layer this build
writes with one part is byte-identical to what it writes today**:
1. A `[mask …]` block with no `[part …]` blocks after it is one part. Its
source keys and its shaping keys build `parts[0]`; nothing is missing and
nothing needs a default invented for it.
2. A layer with exactly one part writes the old shape — source and shaping
inside `[mask …]`, no part block. The format grows only when the feature is
used.
3. `stroke = …` lines keep their position and their meaning. The mode token
gains `push`, and `cling` is a fourth number written only when non-zero.
An older build meeting either drops *that stroke and only that stroke*,
which is the rule [`parse_stroke`](../../core/dr-pipeline/src/sidecar.rs)
already documents and already implements.
**Merge (FR-NC-9).** A part is a block with an id, so two devices that added
different parts to the same layer merge to a layer with both, and the same
part edited on both is a conflict over that part rather than over the mask.
Part *order* is the one thing that is not per-field: it is stored as the block
order and resolved the way the layer order already is.
## 5. On the device
### 5.1 The order a layer rasterises in
Per active layer, into that layer's slice of the existing mask array:
1. **Part 0** draws as it does today — one full-screen pass through the
`switch` in `mask.wgsl`, or per-stroke bounding boxes for a painted one.
2. **Each further part** draws the same way, with the blend state its `join`
names (§5.2). No intermediate texture: the accumulator *is* the slice.
3. **A `Warp` part** copies the slice aside and redraws it displaced (§5.3).
4. `invert` and `opacity` are unchanged — two uniforms, read in the composed
shader, applied to the finished mask.
`MAX_LAYERS` and the array's memory are untouched: parts collapse into one
slice, so eight layers still cost eight channels.
### 5.2 Combining costs three blend states and no new texture
The set operations are already expressible in fixed-function blending over
`r8unorm`, which is why this is the cheap half of the feature:
| Join | `src_factor`, `dst_factor`, op | Result | Built |
|------|-------------------------------|--------|-------|
| Union | `One`, `One`, `Max` | `max(dst, src)` | M1 |
| Subtract | `Zero`, `OneMinusSrc`, `Add` | `dst · (1 − src)` | M1 |
| Intersect | `Zero`, `Src`, `Add` | `dst · src` | M2 |
The middle row is `brush_erase`, already constructed in
[`MaskPass::new`](../../core/dr-gpu/src/mask.rs). The other two are the same
three vertices with a different `BlendState`, and nothing is read back.
**One correction to the first draft of this section, found in the building.**
It claimed no second texture was needed, because a part could be blended
straight onto the layer's slice. That is wrong, and the reason is the erase
stroke: an erase inside a part means *a hole in that part*, not a hole in the
mask. Drawn straight onto the accumulator it takes away whatever the parts
before it had put there — so tidying the edge of a correction punches through
the subject underneath, and the failure reads as the model's mask having holes
in it.
So a part is drawn into one scratch texture (proxy-sized, `r8unorm`, allocated
the first time any layer has more than one part) and blended from there. A
layer of one part still takes the old path exactly — straight into its slice,
no scratch, no combine pass — which is what keeps every existing mask
rendering as it did. `an_erase_stroke_holes_its_own_part_and_not_the_mask` in
[`local_adjustments.rs`](../../core/dr-gpu/tests/local_adjustments.rs) is the test
that holds this in place.
`Max` blending on `r8unorm` is core WGPU and universally supported on the
desktop backends; **verify it on the Android adapter before M2 lands**, since
that is the platform where a blend mode is most likely to be quietly emulated
or absent. If it is missing, union is `One, OneMinusSrc, Add` — a screen blend
— which differs from `max` only where both parts are partially covered, and
never by more than the softness of their two edges.
### 5.3 Push is a warp, not a local morphology
Two mechanisms fit "drag the edge", and the choice matters.
**Local morphology** — offset the distance threshold inside the brush — is
exact, but it exists only where there is a signed distance field, which is
`Subject` and `Category` and nothing else. A push tool that works on a
model mask and does nothing on a gradient, a range or a painted part is a tool
the photographer cannot trust.
**A warp** works on any mask, because it never asks what the mask is made of.
Each dab of a push stroke contributes a displacement, and the mask is resampled
through the sum of them:
```
D(p) = Σ over dabs (b − a) · w(|p − seg(a,b)| / radius)
mask'(p) = mask(p − D(p))
```
Drag from inside the mask outward and the sample point moves back into the
interior, so the boundary follows the finger and **the area behind it fills
in**. Drag from outside inward and the vacated area samples from outside, so
**the area behind it empties**. That is exactly the pair of behaviours asked
for, from one gesture, with the direction supplied by the drag rather than by
a mode switch.
Three honest costs:
- It needs the mask it is sampling, and a pass cannot sample the target it
writes: a warp part costs one texture copy of the slice, plus one draw over
the strokes' bounding box. One scratch texture, proxy-sized, `r8unorm`,
allocated on first use and reused by every layer, since layers rasterise in
sequence.
- The displacement is a **sum**, not a sequential resample. Two crossing
push strokes therefore compose approximately rather than exactly. Bound
`|D|` at one brush radius per stroke: past that a warp tears rather than
drags, and no photographer means the difference.
- A warp moves what is under it, including detail the photographer painted by
hand. That is what it is for, and it is why push is a part in the list —
it can be removed later without disturbing the parts beneath it.
### 5.4 Painting has to be incremental
Today [`MaskPass::render`](../../core/dr-gpu/src/mask.rs) clears each slice and
redraws every stroke of the layer. That is right when a shape changes and
wrong while a finger is down: at 120 reports a second, a layer holding 4096
points redraws all of them per dab, and the cost of a stroke grows as it is
painted — the failure mode `MAX_STROKE_POINTS` already names for one stroke,
here across the layer.
**While a stroke is live, only the new segment is drawn.** The slice already
holds everything up to the previous dab, the blend states are the same ones
that would have been used in a full redraw, and the segments are drawn once
each in the same order — so the incremental result is identical to the
rebuilt one rather than an approximation of it. `MaskPass` keeps, per layer,
the `(part count, stroke count, point count)` it last drew; anything else
changing falls back to the full rebuild it does now.
This is required, not an optimisation to schedule later: it is what decides
whether painting is usable on the phone, and it is the specific failure
[`mask.rs`'s module docs](../../core/dr-pipeline/src/mask.rs) say this whole
design exists to avoid.
### 5.5 Distance fields become per part
[`SubjectMasks`](../../core/dr-gpu/src/mask.rs) uploads one signed distance field
**per active layer, in stack order**, and
[`DevelopSession`](../../ui/dr-ui/src/develop.rs) builds them on the same
indexing. With parts, a field belongs to the part that shaped it: the upload
becomes one field per *model-backed part*, flattened in `(layer, part)` order,
and the rasteriser indexes it by a running counter rather than by `slot`.
This is the most invasive change in the document — it touches the field
builder, the upload, the key that decides when to rebuild, and the shader's
`subject` binding index — and it is mechanical. It is also the reason M2 is
its own milestone rather than a rider on M1.
### 5.6 Cling: paint that stops at the picture's edge
A brush that respects the photograph's own boundaries is the difference
between refining a mask and colouring it in, and the ingredients are already
bound to this pass: the demosaiced source (for range masks) and, when there is
one, the compacted label field.
At each dab, `cling > 0` multiplies the deposit by agreement with the pixel
under the **stroke's first point**: colour distance in the same linear-sRGB
space `colour_mask` already works in, falling off over a tolerance set by
`cling`. Where a label field exists, agreement is 1 inside the same region and
falls to the colour test outside it, so the brush stops dead at a watershed
boundary and softly at a colour one.
Roughly fifteen lines of WGSL reusing `image_value` and `hue_of`, one number
in the sidecar, and it is the single control that makes a 5%-radius brush
usable along hair.
## 6. Seeing the mask
**Status: built.** `MaskStack::rendered`, `Reveal` as a list of
`(layer, colour)`, `RevealStyle`, `EditGraph::compose_revealing`,
`MaskPass::render_revealing`; on the panel, an eye and a colour per row and one
"Show masks as" strip above the stack.
The mask array is **already bound to the composed adjust shader**, so this is
almost free, and it is the first thing to build because every other tool here
is unusable without it.
### 6.1 One correction to this section, found in the building
The draft said "two uniforms — which layer to reveal (−1 for none) and which
style — always emitted and guarded by the uniform, so switching the overlay on
is a uniform write rather than a shader recompile". The second half of that is
not available, and the reason is the case the feature exists for.
A uniform can *select* a slot. It cannot conjure one. A layer with no
adjustment on it changes no pixel, so it is not `is_active`, so it occupies no
slice of the mask array and the rasteriser never draws it — and that is
precisely the layer a photographer wants to look at, for the whole of the time
between choosing a subject and deciding what to do to it. Revealing it means
*rendering* it, which changes the sequence of layers, which changes the
uniform block. The composition moves either way.
So the slot and the style are written into the source, and turning the reveal
on, off, or onto another layer recompiles the fused shader. That is a button
press rather than a frame, and the uniform would only have added a branch per
pixel on top of a recomposition that was happening anyway.
`MaskStack::rendered(reveal)` is the one sequence this rests on: `active()`
plus the layer being looked at. The rasteriser, the composer and the distance
field builder all index by position in it, so all three must be given the same
`reveal` — two of them disagreeing shows as an adjustment applied through
another layer's mask, which is why they take it as an argument rather than
reading a flag.
### 6.2 Where the block runs, and why not with the others
After the output transform, immediately before the clip and the encode — not
among the layer blocks. Everything there runs on scene-referred colour in the
working space, where a flat tint would be pushed through the base curve and
the camera matrix and arrive as some other colour, and an alpha's white on
black would arrive as neither.
### 6.3 Not on the graph
The reveal is an argument to `EditGraph::compose_revealing`, and
`compose_for` — which the exporter, the thumbnail and the neutral probe all
call — has no way to ask for one. A flag on the graph would have been fewer
parameters, would have type-checked, and would have been one forgotten reset
away from a red tint baked into an exported file.
Three styles, all read from the same alpha:
- **Tint** — the mask over the picture in its colour at ~50%. The default,
and what every editor's photographers already expect. The colour is the
mask's own, chosen from the swatches on its row — which is what answers a
red tint over a red dress.
- **Alpha** — the mask alone, white on black. For judging an edge, where a
tint over a busy picture cannot be read.
- **Edge** — the boundary outlined over the untouched picture. For checking
registration against detail the other two hide, and the same reasoning the
region overlay's white outline already carries.
**Per mask, not per selection.** The first build of this showed the *selected*
layer's mask in one global style, and it answered the wrong question. What a
photographer asks of two masks is how they meet — where the sky's edge sits
against the building's — and that needs both on screen at once, in colours that
can be told apart. So each row of the stack has an eye, and each mask a colour
from a six-entry palette (`MASK_COLOURS` in `develop.rs`); the eye is drawn in
that colour so the row says which shape on the picture is its. The style is
the one thing that stays global, because a tint beside an outline beside an
alpha would be three pictures that cannot be read against each other. Alpha
therefore draws every shown mask, each in its colour, on black.
**When it appears.** A new layer arrives with its eye open, in the first colour
nothing else is using — making a mask is asking what it selected, and for a
subject or a category that question has no other answer on screen. Arming
Paint or Erase opens the selected layer's eye if it was closed, on the same
argument `on_part_added` makes: a stroke into an invisible mask is
indistinguishable from a tool that did nothing. Pressing a swatch opens the
eye too, since colouring a mask nobody can see would change no pixel.
Only ever *this* layer's eye, and only on an explicit action. Every other eye
keeps whatever it was set to, and nothing re-arms in the background. The
existing `overlay-hidden` property is the precedent and the trap it documents
applies unchanged: an automatic reveal that re-arms a switch somebody turned
off is worse than no automatic reveal at all.
Viewing state and not edit state: eyes and colours are on the session, not on
the layer, and a photograph reopened has every eye closed.
The ~1s reveal after a shaping slider is released, from the draft, is not
built. It is a timer rather than a decision.
The region overlay stays exactly what it is — a picture of what the model
detected — and gains a name in the interface that says so, because two
overlays that look alike and mean different things is worse than either.
## 7. The interface
### 7.1 Where the tools live
**In the Local panel, not the tool rail.** The rail's entries arm a canvas
gesture for the whole photograph; a brush is meaningless without a layer to
paint into, and a rail entry that silently created one — or that lit up and
did nothing with no layer selected — is precisely the kind of surprise the
rail's own notes argue against. The tool strip sits in the Local panel's
header, enabled only when a part is selected:
```
┌ Local ─────────────────────────────┐
│ [Select] [Paint] [Erase] [Push] ◉ │ ◉ = show mask
├────────────────────────────────────┤
│ ▸ Sky ● ⌄ │
│ Category · sky │
│ ├ Subtract · Painted ✕ │
│ └ Intersect · Luminance ✕ │
│ [+ Add ⌄] [− Subtract ⌄] [∩ ⌄] │
├────────────────────────────────────┤
│ Brush size ──●─── 0.05 │
│ hardness ──●── 0.5 │
│ flow ─────● 1.0 │
│ cling ──●─── 0.4 │
└────────────────────────────────────┘
```
`toolrail.slint`'s note about `MaskSource::Brush` being the next rail entry is
superseded by this and should be replaced with the reasoning, not deleted —
the file is where somebody will next look for it.
### 7.2 The part list
Each layer row gains its parts as indented rows. A part row carries its join
(a chip that cycles add / subtract / intersect), its source name, a delete,
and selection — selecting a part is what points the brush and the shaping
controls at it. The layer row keeps invert, opacity and enable, which are the
layer's.
`[+ Add]`, `[− Subtract]` and `[∩]` each open the same menu of sources the
"new layer" buttons already offer: a gradient, a range, a subject, a category,
or painted. One code path, three joins.
**Folding two layers into one** uses the multi-selection
[`masks_ui.rs`](../../ui/dr-ui/src/masks_ui.rs) already supports: with two layers
selected, "Combine" appends the second's parts to the first and removes it.
Offered only when the second layer's adjustments are neutral, and otherwise
offered with a warning that names what will be lost — quietly discarding an
edit the user made is not a combine.
### 7.3 The brush and the cursor
Radius, hardness, flow and cling are sliders in the panel and all four are
live on the canvas as a cursor: an outer ring at the radius, an inner ring at
the hardness, and — this matters on a phone — the ring drawn at the *touch
point offset above the finger*, since the thing being painted is under the
hand that is painting it.
Radius is stored per tool, not per stroke and not per layer: a photographer
who sets a small eraser expects it to still be small the next time they erase.
### 7.4 Gestures
Each of these needs a `GESTURE:` block beside its implementation — that is the
only place [gestures.md](../gestures.md) can be written from.
| Gesture | Touch | Pointer | Keyboard |
|---------|-------|---------|----------|
| Paint into the selected part | Drag on the picture | Drag | — |
| Erase instead of paint | Hold the Erase tool | Alt-drag | — |
| Push the boundary | Drag with Push armed | Drag | — |
| Change the brush size | Drag the size slider | Scroll with Alt over the picture | `[` `]` |
| See the mask | Press the eye | Press the eye | `\` while held |
| Undo one stroke | The history list | Ctrl+Z | Ctrl+Z |
| Add a part | `[+ Add]`, pick a source | Same | — |
| Fold two layers | Select both, Combine | Same | — |
The `why` each block needs is mostly one sentence — *a stroke is a decision
and a decision is one undo step* — except for Alt-drag, which needs to say
that a modifier is the only way to alternate paint and erase without leaving
the stroke, and that touch cannot have it, which is why the tool strip is a
strip and not a single toggle.
### 7.5 The Slint hazards this walks into
Named because all of them compile:
- **The paint `TouchArea` must be declared in front of
`ScaleRotateGestureHandler`** to receive the press at all, exactly as the
region picker and the repair placer are. The consequence is that a
two-finger pinch may not reach the handler behind it while paint is armed.
Repair has the same arrangement today, so **check what repair mode actually
does with a pinch before designing around it**; if pinch is lost, the answer
is geometry — arm the paint area over the picture only — not z-order.
- **A drag must be measured in the parent frame**, never in the coordinates of
something the drag moves. `GradientHandles` in `masks.slint` is the
reference.
- **The part rows are a repeater inside the develop column**, so their chips
go through `ChipGrid` or `Segmented { columns: 3 }` — a non-wrapping chip
row sets the width of the whole sidebar.
- **None of the above is caught by the test suite.** A screenshot is the only
check; see the project's notes on capturing one under XWayland.
## 8. History, undo, labels
A stroke is one step, recorded on release: `Edit::Action`, not
`Edit::Control`. `Control` is the right variant for a dragged control and the
wrong one here precisely because it coalesces — two strokes painted a second
apart are two decisions, and coalescing them under one key would make the
second untakeable back. A push stroke follows the same rule, as does adding,
removing or re-joining a part.
The mask stack is already snapshotted per step and shared by `Arc` when a step
does not touch it, so the cost of an undoable stroke is a clone of one layer's
parts, not of the picture. New keys in
[`labels.rs`](../../ui/dr-ui/src/labels.rs):
```
history.mask_painted "Paint Mask"
history.mask_erased "Erase Mask"
history.mask_pushed "Push Mask Edge"
history.mask_part_added "Add To Mask"
history.mask_part_removed "Remove From Mask"
history.mask_joined "Change How Mask Joins"
history.masks_combined "Combine Masks"
```
## 9. Sync and merge
Nothing here stores pixels, so nothing here changes what sync carries beyond
size. A painted correction is bounded by `MAX_LAYER_POINTS` at roughly 50 kB
of text in the worst case and a few hundred bytes in the ordinary one. A model
part still carries its run-length coded `coverage` line, unchanged and still
outside `PartialEq`.
Two conflicts are new, and both resolve per block: two devices adding
different parts to one layer (both are kept, in block order), and two devices
painting the same part (a conflict over that part, arbitrated the way a layer
already is). A device that has never run a model still renders every part
correctly, because coverage travels with the part exactly as it travelled with
the layer.
## 10. Performance and budgets
The number that matters is the cost of one dab while the finger is down, and
with §5.4 it is bounded by the dab's own bounding box rather than by the
stroke's history: `(2r)²` pixels × one segment. At the default 5% radius on a
1600×1067 proxy that is about 25 000 pixels — a fraction of a millisecond, and
independent of how long the stroke has been going.
Everything else runs on a shape change, not per frame:
| Work | When | Rough cost at 1600×1067 |
|------|------|------------------------|
| Full layer rebuild | part added, source changed, resize | one draw per part |
| Warp | a push part exists and something below it changed | one copy + one bbox draw |
| Distance field | a model part's morphology changed | as today, CPU, per part |
| Overlay | never — it is two uniforms in a shader that already runs | — |
The scratch texture for warp is one proxy-sized `r8unorm` — under 2 MB — and
is allocated on first use, so a photographer who never pushes an edge never
pays for it.
Measure against [frame-budget.md](frame-budget.md), and measure it on an idle
machine: the 16 ms guard is load-sensitive enough that a busy build makes it
fail and an idle one makes it pass, so a single run proves nothing either way.
## 11. Order of work
**M1 — See it and paint it.** The overlay (§6), `parts` on the layer with
`Union` and `Subtract`, `Painted` parts, the tool strip, the brush HUD and
cursor, the canvas gestures, incremental raster (§5.4). No new shader code for
the mask pass beyond the blend variants; no change to distance fields, because
a painted part needs none. *This is the whole of the user-visible ask except
push and intersect,* and it is deliberately the milestone that stands alone.
Done: parts with `Union` and `Subtract`, painted parts, the tool strip, the
canvas gesture, and the overlay (§6). Outstanding: the brush HUD and cursor —
radius, hardness and flow are sliders with no ring drawn on the photograph,
so the size of the brush is a number rather than a thing you can see — and
the incremental raster of §5.4, without which a layer's whole stroke history
is redrawn per dab.
One lesson from the order it was actually built in, since §6 said it and the
build did not listen: the overlay is not the last quarter of M1, it is the
first. Parts and painting shipped without it and the result was a feature
nobody could tell was working — the panel listed a mask, the photograph showed
nothing, and every report of it came back as "the masks do not work".
**M2 — Combine properly.** `Intersect`, model and gradient and range parts,
per-part distance fields (§5.5), the part list with its join chips, folding
two layers.
**M3 — Push, and cling.** The warp pass (§5.3) and edge-aware deposit (§5.6).
Both are refinements of a tool that already works, which is the right place
for the two riskiest pieces.
**M4 — What the use of it asks for.** Left open on purpose. Likely candidates:
a "select the subject under the pointer" brush, pressure from a stylus, and a
per-part opacity.
## 12. Requirements to add
Proposed text for [requirements.md](requirements.md) §3.3, in the shape the
neighbouring entries take:
> **FR-DEV-19 — Mask editing.** A mask layer's coverage shall be editable by
> hand after it is created, by painting into it, erasing from it, dragging its
> boundary, and joining further selections to it. Every edit is stored as
> geometry and parameters in the edit graph; no rasterised mask is written to
> a file and none exists in CPU memory.
>
> **FR-DEV-19a — Mask composition.** A layer's mask is an ordered list of
> parts, each naming a source and how it joins the mask before it — union,
> subtraction, or intersection. A layer of one part is exactly the layer of
> today, and reads and writes the same sidecar.
>
> **FR-DEV-19b — Hand correction.** A part may be painted, with add, erase and
> push strokes, at a radius, hardness, flow and edge-clinging the photographer
> sets. Strokes are stored as normalised coordinates and rasterised on the
> device.
>
> **FR-DEV-19c — Boundary push.** A push stroke displaces the mask beneath it
> along the drag, filling behind an outward drag and emptying behind an inward
> one, on any mask source rather than only those with a distance field.
>
> **FR-DEV-19d — Mask visualisation.** The mask a layer actually produces —
> after its parts, its shaping, its inversion and its opacity — shall be
> displayable over the photograph as a tint, as an alpha, or as an outline,
> and shall appear automatically while a mask is being edited.
Each needs `TRACES:` comments at the implementation and a regenerated
[traceability.md](traceability.md) **in the same commit**, since that file
carries line numbers.
## 13. Verification
Testable without a GPU, and therefore not optional:
- `Stroke` mode round-trips through the sidecar, including a push stroke and a
non-zero cling; an unknown mode costs one stroke and not the layer.
- A one-part layer writes the byte-identical sidecar it writes today, and
every existing fixture in `mask_sidecar.rs` still loads.
- Parts merge per block: different parts added on two devices produce a layer
with both; the same part edited on both is one conflict.
- Point budgets hold across parts, and painting past them refuses rather than
dropping the oldest strokes.
- The panel model builds part rows from a hand-made stack with no device
present, the way `rows_from` does for adjustments.
- History: a stroke is one step, undo restores the parts, and a step that
leaves the masks alone still shares the `Arc`.
Needing a device, in `dr-gpu/tests`:
- A truth table for the joins: a half-covering part unioned, subtracted and
intersected with a known base gives the three expected fields.
- A warp of a known step edge moves it by the drag, and the area behind it
fills — asserted on a readback of the mask array, which is a test reading
back, not the application.
- Incremental drawing equals a full rebuild, dab for dab, on a stroke of a few
hundred points. This is the test that keeps §5.4 honest.
- `Max` blending behaves on every adapter CI runs on.
And a screenshot, because §7.5 is not reachable by any of the above. Run
`dr-ui` tests single-threaded; parallel runs segfault in this crate for
unrelated reasons.
## 14. Decisions I need
1. **Parts now, or strokes first?** M1 as written introduces `parts` and pays
the migration once. The cheaper alternative is a `strokes: Vec<Stroke>`
field on the layer, brushed over whatever the source produced, and parts
later — half the code, and it makes "subtract this from that" a second
mechanism invented afterwards. I would take the migration now.
2. **Push as a warp, or local morphology on model masks only?** §5.3 argues
the warp. It is the more expensive of the two and the only one that works
on every source.
3. **Tool strip in the Local panel, or an entry in the rail?**
`toolrail.slint` predicts a rail entry; §7.1 argues against it. The rail is
easier and worse.
4. **Do the FRs in §12 go into `requirements.md` now,** or stay proposed here
until the first milestone lands?
+495
View File
@@ -0,0 +1,495 @@
# DarkRoom — Outstanding work
**Status:** Living document · first written 2026-08-29
**Companion to:** [requirements.md §7](requirements.md), [technical-debt.md](technical-debt.md),
[traceability.md](traceability.md)
What is specified and not built, and for each cluster whether that is a decision, a dependency, or a
gap nobody has looked at.
This document exists because [traceability.md](traceability.md) cannot tell those apart. It reports
one number — the share of requirements carrying a `TRACES` tag — and a missing tag means either
"nobody has built this" or "somebody built it and did not say so". Both read the same way in the
summary table, which makes that figure pessimistic *and* uninformative at once: it understates what
works while hiding which of the remainder matters. Eight requirements gained a tag on this branch
because the code already satisfied them and nobody had said so. Everything below is the other kind.
It is also not a plan. [requirements.md §7](requirements.md) records what was deferred deliberately
and needs no argument; this records what is still nominally in scope, so that the distance between
the register and the binary is visible rather than something a reader has to reconstruct from a
percentage. Where the honest answer is "this requirement should be amended rather than met", it says
so — an unbuilt requirement that nobody intends to build is worse than a deferred one, because it
keeps costing attention.
**Several entries were struck between 2026-08-29 and 2026-08-30**, as two waves of work
landed: burst grouping (FR-CULL-5), Flatpak packaging (FR-PLAT-LIN-3), FR-CULL-3 in full,
Android memory pressure and lost-root recovery (FR-PLAT-AND-5, FR-PLAT-AND-2), image intents
(FR-PLAT-AND-6), and the job runner (FR-PLAT-AND-4's Rust half). What remains of each is
recorded where it appears rather than deleted, because a requirement that is *half* met is the
one most likely to be reported as closed.
---
## 1. Plugins — post-v1 since 2026-09-19
> **Resolved, in the register.** The contradiction below was settled on 2026-09-19 the way the
> last paragraph of this section asked: §3.10 is marked `(post-v1)` clause by clause, §7's row
> says so with a reason, D16 defers with it, and the traceability tool lists deferred clauses in
> their own table instead of counting them. Coverage went from 72.2% of 194 to 80.6% of 170 on
> that edit alone. What follows is kept as the record of what was decided and why; nothing in it
> is owed a tag.
**Untagged:** FR-PLG-1, -1a, -2a, -2b, -2c, -3, -3a, -4, -4a, -5, -5a, -5b, -5c, -6, -6a, -7, -8,
-9, -10, -11, -12.
No plugin host exists. No crate loads anything at runtime: there is no manifest reader, no WASM or
Lua engine, no registry, no signature check, no install path, no capability grant, no per-plugin
failure ledger. `declared/mod.rs` says as much in its own documentation — the operation format is
"not a plugin directory read at startup".
**Two of §3.10's requirements are met, and they are the interesting two.** FR-PLG-2 and FR-PLG-2d —
the declarative node format — are built and tagged: `core/dr-pipeline/ops/*.yaml` compiled by
`build.rs`, with the restricted expression grammar in `declared/expr.rs` and a parity test asserting
a declared operation and a hand-written one produce identical output.
[code-health.md §3](code-health.md) calls it "a working plugin system that happens to resolve at
build time", and that is exactly right. What is missing is not the format; it is everything that
would let somebody who is not in this repository use it.
**The contradiction.** [requirements.md §7](requirements.md) lists `| Plugin API | — |` among the
things deferred for v1 — a bare row, where most deferrals carry a justifying note. §3.10 then spends
roughly 280 lines and 23 requirement IDs specifying that same Plugin API in detail. Both statements
are in the register of record, and the traceability denominator counts the second one: 21 IDs, 12%
of all 179 defined requirements, worth about twelve points of coverage on their own — and nearly a
third of everything the matrix reports as uncovered. A reader looking at the coverage figure has no
way to know that, or that the subsystem behind it is one the same document says is not in this
version.
**And D16 is open.** [Decision D16](requirements.md) — plugin licensing — records that GPLv3
answers the derivative-work question differently for each of §3.10's three plugin forms, and that
this "must be answered *before* an ecosystem exists, not after", because a term introduced later
cannot be applied to plugins already written. D16 explicitly does not block FR-PLG-2; it blocks
publishing a third-party format as stable.
**What would resolve this:** an edit to `requirements.md`, not code. Either §7 drops the row, or
§3.10 is marked deferred with the two built requirements carved out. Until one of those happens the
coverage figure is measuring a decision that has already been taken, and taking it again every time
somebody reads the matrix.
---
## 2. Culling — the stated differentiator, half built
[D11](requirements.md) names culling "the core differentiator". FR-CULL-1, -2, -3, -4 and -8 through
-12 are built. Three are not.
**FR-CULL-3 — Raw-truth overlays. Built, all three bullets.** Focus peaking is
`core/dr-gpu/src/focus.rs` and `ui/dr-ui/src/peaking.rs`; the raw histogram and the raw clipping
indicators are `core/dr-gpu/src/raw_histogram.rs` and the second reading of the panel in
`histogram.slint`.
Worth recording, because it is the thing this entry previously got wrong and the next reader will
have to check again. The display histogram — `dr-gpu/src/histogram.rs`, `ui/dr-ui/src/histogram.rs`
— is **not** this requirement and never was: it is tagged FR-DSP-7, it reads `AdjustPass`'s 8-bit
output, it counts clipping as `r == 255`, and so it describes the frame the display is about to
show, after the whole develop chain. FR-CULL-3 asks for the *sensor data*, on the explicit grounds
that a rendered image "systematically lies about what is recoverable in the raw". The two now sit in
one panel behind a chip row, which is the arrangement that keeps them from being mistaken for each
other: they answer different questions and both are true.
What the raw reduction cannot answer is written down rather than left to be discovered —
[architecture.md §5.5](architecture.md) records why it reduces over the demosaiced texture instead
of the CFA samples §5.5 originally specified, and what that costs in what it can say.
**FR-CULL-5 — Burst and near-duplicate grouping. Built.** `core/dr-catalog/src/bursts.rs`:
frames join a burst when they are adjacent in time *and* look like the frame before them, compared
adjacent-pair-only in one ordered walk. Signatures are 64-bit difference hashes taken from the
thumbnails `dr-thumbs` already holds, on a background pass after the thumbnail sweep — nothing at
import, nothing at query time. Nothing scores or rejects a frame: the representative is the
earliest, a fact about the clock, and a new burst arrives *open* so the pass never takes a row off
the screen.
Two threads left hanging. `core/dr-face/src/calibrate.rs` still says "since FR-CULL-5 already
groups bursts, positives are bootstrapped from bursts" while in fact bootstrapping from confirmed
labels — that comment was a forward reference and is now simply wrong, rather than premature.
And `dr_catalog::bursts::choose_representative` is written and tested but bound to no gesture, so
today the only override is expanding the burst.
`core/dr-catalog/src/dedup.rs` remains a different thing: re-import detection under FR-CAT-11,
matching a file against one already catalogued, not two photographs against each other.
**FR-CULL-6 — Compare and survey.** Absent. No side-by-side view, no synchronised zoom or pan.
This is the one of the four with no adjacent machinery at all, and it is also the one that most
directly distinguishes culling from browsing.
**FR-CULL-7 — Culling on tablet.** Absent, and blocked by the three above rather than independent
of them: there is no separate tablet culling surface to build until there is something to put on it.
---
## 3. FR-DEV-3g — AI denoise
Promoted into v1 by [D11](requirements.md), and named there as the precondition for deferring AI
masking — the argument being that one learned stage earns the runtime that a second could then
reuse. Only classical noise reduction exists: `ops/noise_reduction.rs`, a bilateral filter in two
arrangements, exact for luminance and separable for chroma. It is good, and it is not this.
`models/` holds two face models and nothing else; `core/dr-segment/models/` holds a YOLO
segmentation model for subject masks. There is no denoise model, no learned demosaic, and no
inference path that is not face or segmentation.
The obstacle is not the pipeline. It is that [D13](requirements.md) — model licensing — is still
open for the models that already ship, and adding a third learned stage adds a third licence to
answer for. Building the runtime before that is settled means owning the same problem in one more
place.
---
## 4. The render path — FR-DSP-2, FR-DSP-4, NFR-RES-2
**FR-DSP-2 — Tiled computation. Unbuilt, and under challenge.** [architecture.md §6.2](architecture.md)
calls for tiling "from day one" on the grounds that retrofitting it is a rewrite. It was not built,
and the evidence has since moved. `core/dr-gpu/tests/frame_budget.rs` carries the argument in its
own header: one fused dispatch over a viewport-sized target is comfortably inside the frame budget,
and "if that stops being true, the recommendation to strike tiled computation from the interactive
path stops being supported, and this test is what says so."
[technical-debt.md TD-4](technical-debt.md) reaches the same place from the other direction — a
tiled convolution at clarity's radius reads nearly twice the taps that an untiled one does, so the
stage that looks most like it wants a tile cache is the stage that would be hurt most by one.
What exists is the declaration and not the mechanism: `DetailPass::radius` is documented as the halo
a tile would have to be grown by, with a test that pins it, and there is no scheduler to read it.
That is deliberate plumbing, not an oversight.
**So the open question here is not "when is tiling built" but "is FR-DSP-2 still a requirement" — and on 2026-09-19 the answer was: as written, until S6 runs.** FR-DSP-2 now carries a status note saying exactly that, and R5's note no longer claims it was rewritten.
Two measurements say it costs more than it saves on the interactive path. Neither says anything
about the export path or about a device under memory pressure, which is where the case for it
actually lives — and that is spike S6, which has not run.
**FR-DSP-4 — Progressive refinement.** Unbuilt. FR-DSP-1's proxy rendering and TD-4's
quarter-resolution base are adjacent and are not it: both are fixed choices about what resolution to
compute at, where FR-DSP-4 asks for a first frame that is deliberately cheap and a second that
replaces it. Nothing tracks a "this frame is provisional" state.
**NFR-RES-2 — Images larger than GPU memory.** Half answered. NFR-R8's "decide explicitly" was
decided on 2026-09-19: there is no CPU render pipeline, the degraded mode is the viewer on
embedded previews with develop withheld, and NFR-RES-2 no longer promises a fallback render. What
remains unbuilt is the memory half: there is no headroom budget, no allocation-failure staging,
and no spill. Spike S6 — a tiled pipeline on a
mid-range Android device with an image larger than available GPU memory — is the one that would
settle both this and FR-DSP-2, and there is no evidence it has run.
---
## 5. Android beyond running, and Flatpak
The Android app is not a stub — it builds an APK, runs the whole application, unpacks bundled face
models, and has been measured on a tablet ([faces.md §12.1](faces.md),
[technical-debt.md TD-1](technical-debt.md)). What is missing is the platform contract around it.
**FR-PLAT-AND-1 is untagged, and what it was tagged for was intent rather than code.**
The requirement demands that library access be obtained *exclusively* through the Storage Access
Framework. There is no SAF code: no `ACTION_OPEN_DOCUMENT_TREE`, no `takePersistableUriPermission`,
no `DocumentsContract`. Its two tags rested on a `SourceRef::Document` variant constructed only
inside `#[cfg(test)]` — `LocalStorage::open` refuses it, and the test that proves so is named
`a_reference_of_the_wrong_kind_is_refused_rather_than_guessed_at` — and on
`dr_plat::imports_supported`, which *returns false on Android* and whose own documentation says it
"stops being false when a SAF implementation lands". The second tag documented the absence of the
thing it was counted as evidence for. Both have been removed; this is the "plumbing a future feature
would use" case [CONTRIBUTING.md](../../CONTRIBUTING.md) and [code-health.md CH-4](code-health.md) both
warn about. Android reaches a library through a Nextcloud account or a folder, over paths, like the
desktop.
That has a consequence for the rest of the cluster: **FR-PLAT-AND-2** — detecting the loss of a
granted tree permission and marking images offline rather than deleting rows — cannot be built until
there is a permission to lose. It is listed here as unbuilt, but it is blocked, not skipped.
**FR-PLAT-AND-4 — half built.** The runner is done (`core/dr-catalog/src/runner.rs`): the
queue that `jobs.rs` always had is now claimed from, completed, failed and recovered after a
crash, which is FR-PLAT-AND-3's resumability as much as this requirement's. What is missing is
the platform half — a foreground `Service`, `FOREGROUND_SERVICE` and `POST_NOTIFICATIONS` in the
manifest, and a stated Doze behaviour. The build step that blocked it is no longer a blocker: the
APK now compiles its own Java.
Note also that **no handler is registered**, deliberately. The only enqueue site reachable in the
shipping app produces remote thumbnail jobs already served by the async grid worker, and
`walk::scan_root` — which holds the other two enqueue sites — has no caller outside an example.
Wiring the sweep to claim from the queue is the honest next step and is an async rewrite of
`library.rs`.
**FR-PLAT-AND-5 — built.** A tiered eviction registry drives GPU caches, then proxies, then
thumbnails, from `MainEvent::LowMemory` and `MainEvent::Stop`.
**FR-PLAT-AND-6 — built, with one half unwired.** VIEW, SEND and SEND_MULTIPLE filters, the launch
Intent read over JNI, and an `ExportProvider` rooted at `getFilesDir()` rather than AndroidX's
`FileProvider`. The outbound share has no caller in `ui/` yet. **None of the runtime behaviour has
been exercised on a device** — the tests read the manifest and the Java through `include_str!`,
which catches a deleted filter but not a class loader that cannot find the class.
**FR-PLAT-LIN-3 — packaged, not satisfied.** There is a Flatpak manifest now, granting no
filesystem permission of any kind, plus AppStream metainfo and `docs/distribution.md`. The
requirement is still not met, and cannot be met by packaging: a folder library is chosen by typing
an absolute path, nothing in the tree calls the FileChooser portal, and inside the sandbox `$HOME`
holds only `.var/app/...`. `dr_plat::volumes()` reads `/proc/self/mountinfo`, so a card mounted on
the host is invisible to a sandboxed process as well. The fix is an `ashpd` directory picker beside
`LocalStorage::grant`, not a change to the manifest. No Flatpak has been built here —
`flatpak-builder` is not installed — so the permission set is reasoned, not observed.
**NFR-COMPAT-2 — distribution channels. Stated, which is all this requirement asks.** The paragraph
above cites [distribution.md](distribution.md) and it is the same document that answers this: §1
names five channels and their state — Arch source package and Flatpak in tree, AppImage a v1 channel
whose recipe is not written, F-Droid a v1 channel not yet submitted, and Play explicitly **not** v1.
The requirement is to *state* the channels, and they are stated, including the deferral.
What remains is the coupling the requirement points at rather than the statement it demands. §4.8
observes that publishing on Play is what turns SAF from a preference into a constraint, and
distribution.md §6 argues the coupling runs the other way for this project — F-Droid asks nothing
that ARCH §6.9 does not already require. Spike S11, the Play permissions dry-run, has not run, and
until it does that argument is reasoned rather than confirmed. Two of the five channels also exist
as decisions rather than as recipes, and no Flatpak has been built here at all.
Related, NFR-COMPAT-1's baseline is real but scattered — API 28/36 live in the Android
Dockerfile and are checked in CI against the built ELF, which is good — while the items the
requirement singles out are missing: whether `shaderFloat16` and 16-bit storage are required (the
one it flags as jeopardising R1), minimum RAM, minimum desktop Mesa, and a named reference device
from a second GPU vendor.
**NFR-OPS-2 is met, and NFR-OPS-4 is not.** Crash reporting is `platform/dr-plat/src/crash.rs`: a
panic on either platform writes a local record with a redacted message and backtrace, ten are kept,
and there is deliberately no upload path — the requirement's "upload only on explicit opt-in" is
satisfied by there being nothing to opt into, and the module says why a transport built ahead of the
consent is the wrong order. (This paragraph said the opposite until 2026-09-12; the record had landed
on 2026-08-30 and the paragraph had not been read against it.) Update and first run are undefined;
the concrete reason NFR-OPS-4 gives — that D2 pins rawler at a non-SemVer alpha whose camera-support
fixes users will need — is unaddressed, and there is no update mechanism of any kind.
---
## 6. Accessibility and internationalisation — the hard half is done and the easy half is not
**NFR-A11Y-1 — Localisation.** `@tr(` appears **zero** times across 14,482 lines of Slint. That
number overstates the problem, because the part that is genuinely architectural was got right:
`LocalizedKey` keeps display strings out of `core/` entirely, every operation publishes a key rather
than a label, and `labels::resolve` is the single point where a key becomes text. What that single
point does, however, is a hardcoded English `match` in Rust source — so changing a translation
requires a recompile, which is the one thing the requirement explicitly forbids. There is no message
catalogue in any format, no locale-resolution rule, and no decision recorded about RTL.
The work left is therefore smaller than it looks and entirely mechanical: a catalogue format, a load
path behind `resolve`, and `@tr(` around the Slint literals. The design it needs already exists.
**NFR-A11Y-2 — Accessibility.** `accessible-*` appears five times in the whole interface, all five
on one control — the parameter slider in `adjust.slint` — and nothing is set from the Rust side at
all. Everything else in eighteen Slint files is unnamed to AT-SPI and TalkBack. The requirement's own
caveat, that Slint's Android accessibility needs verifying, is spike S13, which has not run.
**NFR-A11Y-3 — Colour-independent status.** Built where a control exists, and now tagged: the
clipping readout pairs a marker that appears or disappears with a figure in words, the rating strip
is a solid star against an outline in an achromatic palette, the pick/reject mark is a tick against
a cross, and the focus-peaking colour chips say "Red" and "Cyan" rather than showing swatches. Each
of those already carried the reasoning in a comment naming this requirement and simply had no
`TRACES` line.
Two caveats, because the tag now says more than the evidence does. **Only the clipping clause has a
test** — `a_clipping_figure_distinguishes_none_from_nearly_none`, which pins `<0.1%` apart from `0%`
so the figure cannot contradict the lit marker beside it. The three Slint components are
inspected-and-argued, not asserted, and nothing would fail if a future edit made a star differ only
in tint. **And the requirement's first named example has no interface at all**: catalog colour
labels are a nullable `label INTEGER` column on the versions table and are set and shown nowhere, so
the clause about them is untestable rather than satisfied. That clause closes when the label UI is
built, not before, and it should be built with a shape from the outset — which is the same argument
as below, for doing this alongside NFR-A11Y-2 rather than after it.
---
## 7. Catalog and sync
**FR-CAT-14 — Migration import.** Reading ratings, labels, keywords and collections out of a
Lightroom `.lrcat` or a darktable `library.db`. Unbuilt. The destination is not: keywords,
collections, ratings and the cross-device merge rules are all built and tested, and
`keywords.rs` already anticipates the arrival ("an import from Lightroom can bring in…"). What is
missing is only the two source adapters — which is a comparatively contained piece of work for a
requirement that decides whether somebody can try this software on a library they already have.
**FR-NC-11 — Initial catalog build.** Using WebDAV `SEARCH` (RFC 5323) against `/remote.php/dav/`,
filtered by mimetype and paginated, in preference to walking folders with PROPFIND. Unbuilt: no
`SEARCH` request is issued anywhere. The PROPFIND walk this exists to replace is fully built and
well optimised — ETag pruning under FR-NC-4 turns an unchanged 50k library into one request — so the
gap is narrower than it reads. It is the *first* build against a large remote library that pays, and
that is the moment a new user meets.
**FR-CAT-13 — XMP interoperability, wired on 2026-09-12.** `core/dr-xmp` reads and writes
standard XMP sidecars: `dc:subject` and `lr:hierarchicalSubject`, `xmp:Rating` and `xmp:Label`, and
the IPTC core fields, in both the attribute and the element form and whatever RDF container a file
happened to use. It states the ownership rule in one place — DarkRoom owns the properties in
`PROPERTIES` and nothing else in the document, identified by namespace URI rather than by prefix —
and enforces it by rewriting a packet event by event rather than serialising over it, so another
application's `crs:` settings, comments and processing instructions survive a write byte for byte.
**The wiring is `ui/dr-ui/src/xmp_sync.rs`.** The scan collects `.xmp` beside `.drsc` from the
listings it was already making; the pull reads each one whose ETag has moved and reconciles it
against the catalog with the catalog winning — keywords union, and a rating, label or caption taken
only where the catalog holds none. Both naming conventions resolve: `IMG_0001.CR3.xmp` names its
file, `IMG_0001.xmp` the stem, and the JPEG beside a RAW is the same photograph. A genuine
disagreement is written to `xmp_conflicts` and the settings page offers "Take the sidecars' values",
which is the reload the requirement asks for; the detection it asks for is the ETag that moved. The
write in the other direction is behind a setting that starts off (NFR-R4): a judgement or a keyword
then also rewrites the sidecar beside the original, keeping the file's own caption, copyright and
hierarchy, which the catalog has no columns for and would otherwise have deleted.
**What remains.** GPS is not carried — `exif:GPSLatitude` is a format of its own and `dr-decode`
produces no location for it to carry yet. Title, description and copyright are read and reconciled
but the catalog has nowhere to put them, so they pass through a rewrite rather than being editable.
An XMP write made offline is not queued: the catalog and the `.drsc` are authoritative, and the next
judgement online writes the file whole again.
`dr-preset-xmp` remains what it always was and is still not the counter-example it looks like: a
reader of Lightroom *presets* under FR-DEV-6, a different file for a different purpose.
---
## 8. The performance targets are half-verified, and the half that is left is the hard one
§8 and §4.1 both require the same thing in the same words: an automated benchmark suite against a
synthetic 50k catalog, run per commit, where **"a regression beyond a stated tolerance is a build
failure, not a notification."** For most of this project's life it did not exist — no `benches/`, no
criterion, no synthetic catalog, and three CI workflows that between them measured nothing.
**It exists now, for everything that does not need a frame.** [`tools/bench`](../../tools/bench) builds
a deterministic 50,000-row catalog over a pool of a dozen real files, measures against it, and fails
the build on a violated budget or a drift past tolerance;
[`.gitea/workflows/benchmark.yml`](../../.gitea/workflows/benchmark.yml) runs it on every push, and
[benchmarks.md](benchmarks.md) is the account of what it does and does not cover. **NFR-P1** and
**NFR-P3** are now genuinely gated, and R2's "catalog opens in under 2s" clause with them.
Three qualifications, all of them stated in the harness itself rather than only here:
- **The numbers have not been recorded yet.** Every `recorded` field in
[bench-baseline.json](bench-baseline.json) is `null`, deliberately: a fabricated baseline is worse
than none. Until `dr-bench record --reference` is run on the reference desktop and committed, the
budget gate works and the regression gate does not.
- **NFR-P7 and NFR-P8 are half-measured and are not tagged.** The export row covers the encode half
of the chain and no GPU render, so it can fail the requirement and cannot pass it. The memory row
covers a process holding the catalog and nothing else — no toolkit, no adapter — so it is the
catalog layer's share of the 500 MB rather than the figure NFR-P8 is about. Neither carries a
`TRACES:` tag, which is the point.
- **NFR-P8 needs a decision, not more code.** How much of its 500 MB belongs below the UI is
unstated, and until somebody says, the metric can record but not judge. [benchmarks.md](benchmarks.md)
also answers the question §4.1 raises about GPU memory — RSS cannot see device-local allocations
at all — and recommends restating the requirement as two figures.
**What is left is the frame-timing half, and it is the hard one.** NFR-P2, -P4, -P5, -P6, -P9, -P10,
-P11, -P12, -P13, -P14 and -P15 all need a probe inside a running Slint application, a GPU adapter,
or both. `dr-gpu/examples/frame_budget` is a real instrument for the GPU part and its results are
committed in [frame-budget.md](frame-budget.md) with the machine and profile named — but it is run
by hand, and the guard version in CI skips itself where there is no adapter, which is the normal
case on a runner. So the claim to take from this section is now narrower than it was, and still
true: **a scroll that dropped to 30 fps tomorrow would reach a user before it reached CI.**
---
## 9. Two core requirements that cannot be closed as written
**R1 — Cross-platform output within a bounded tolerance.** §2 states that the threshold "must be
fixed before spike S9", because S9 both validates R1 and calibrates what tolerance is achievable.
The threshold was never fixed and S9 has not run, so R1 currently has no acceptance criterion at
all — there is nothing a test could assert.
The matrix used to report R1 as *covered*, and what covered it was two string literals: fixtures
inside the traceability tool's own unit tests, which the tool scans along with everything else,
because a fixture demonstrating tag extraction was indistinguishable from a tag. The extractor now
asks where the tag sits — a tag is the first word of a comment, not a string appearing anywhere on a
line — and R1 is untagged again, which is the honest reading while it has no acceptance criterion to
tag anything against.
NFR-OPS-1 was covered by tags that were real rather than fixtures, which is the worse case of the
two: one on `compute_coverage` and one on the gesture extractor, both on the traceability tool. A
coverage calculation and a documentation generator are not diagnostics under any reading, so both
tags were removed. It is the case [CONTRIBUTING.md](../../CONTRIBUTING.md) warns about in its own words:
a tag proves a tag exists. The requirement has since been built where it says: the rotating,
size-capped log and its redaction in `platform/dr-plat/src/diagnostics.rs` (2026-08-30), and the
bundle in `diagnostics/bundle.rs` (2026-09-12) — the log, the crash records, the version, the schema
and the GPU as one text file, shown in full in Settings before a second press writes it, and sent
nowhere by either press.
**R2 — Efficient display of huge RAW libraries.** Its acceptance criterion contains "*(figure
TBD)*" — the scroll velocity below which no cell may render as a placeholder — and asks for a stated
prefetch margin and cache-hit rate. No figure is stated anywhere in the tree, neither quantity is
measured, and [TD-2](technical-debt.md) and [TD-3](technical-debt.md) both describe the thumbnail
path falling short of it in ways that were measured. R2 was deliberately left untagged on this branch
for that reason: the machinery is substantial and the criterion is unmet and partly undefined.
Both belong with §8 above. A requirement whose threshold was never chosen and a target nothing
measures fail in the same way — not by being wrong, but by being unfalsifiable.
---
## 10. Spikes
§9 defines fourteen validation spikes and says of three of them: "S1, S2 and S10 are the three that
can invalidate the architecture."
Only **S1** (Slint + wgpu zero-copy on Linux) and **S14** (the face pipeline on a real library) have
recorded results. S14's are the best evidence of any spike — a dedicated document, a measured pass
over an 18,143-face library, a named device and a reproducible command — though D13's licensing half
remains open.
**S6, S9, S10, S11 and S13 show no evidence of having run at all.** Each is referenced only from the
requirement text that asks for it:
| Spike | Would settle | Blocked on |
|---|---|---|
| S6 | FR-DSP-2, NFR-RES-2 — tiling and images larger than GPU memory | Nothing; needs a device and a large image |
| S9 | R1's tolerance threshold, and therefore R1 | Nothing; the threshold is defined *by* running it |
| S10 | Whether SAF at 10k files meets NFR-P1/P3 | §5 — there is no SAF code to measure |
| S11 | NFR-COMPAT-2, and whether Play makes SAF binding | Nothing |
| S13 | NFR-A11Y-2 on Android | §6 — there is almost nothing to test with |
S2, S3, S4, S5, S7, S8 and S12 are also unrun, several with acknowledgements in the code that say
so (`dr-sync/src/upload.rs` on S8, `dr-sync-nextcloud/src/lib.rs` on S3). S2 is one of the three
architecture-invalidating spikes and needs Adreno and Mali hardware, which the manifest notes no
emulator represents.
The pattern is worth stating rather than leaving to be inferred: the spikes that ran are the ones
whose subject was being built anyway. The ones that did not are the ones that would have said
whether something *should* be built — which is the opposite of the order §9 asks for.
---
## 11. Merging — specified 2026-09-19, nothing built
§3.11 was written on 2026-09-19 under D18, undeferring the panorama from §7 and leaving HDR merge
and focus stacking there with their data model decided. Eleven `FR-MRG` clauses and two `NFR-MRG`
figures entered the register at once with no code behind any of them, which is why the coverage
figure fell from 83.0% to 77.2% on the same day — a specification, not a regression.
[panorama.md](panorama.md) is the design, and its §10 is the order of work. Nothing starts before
**S15**: whether rawler reads back a linear DNG the application writes, whether XFeat loads under
tract at a fixed shape, whether the working-space texture can be tapped where FR-MRG-2 needs it,
and what a chunked blend of a 100 MP composite costs on the tablet. The first two are a day each
and either can change the design, which is the reason they come first.
## 12. D12, which governed all of the above
> **Decided 2026-09-19, by events.** The scope stands as calibrated, v1 has no date, and `(post-v1)`
> in §7 is the one way a clause leaves the count — used for the plugin API and nothing else. The
> argument below is kept because it is what the decision weighed; its prediction about tablet
> editing was right, and the cluster was built anyway.
[Decision D12 — scope versus pace](requirements.md) said, while it was open:
> The calibration selected an ambitious feature set — full tablet editing, full ingest, culling as a
> differentiator, complete GPU masking, AI denoise, Fuji-first colour, deep sync, sidecar durability
> — against a stated pace of evenings and weekends, indefinitely.
>
> **Those are not compatible as stated.**
Sections 1 through 10 are what that incompatibility looks like eleven versions later, and they land
almost exactly where D12 predicted: tablet editing carries SAF at unproven scale, background
execution limits and two GPU vendors to validate (§5), and every one of those is unbuilt or unrun.
The parts that *were* built — the develop pipeline, sync, faces, the catalog — are the parts that
did not need a decision first.
D12 was not resolved by choosing to work faster, and in the end not by moving clusters into §7
either, except the one: plugins. Every other cluster above stays in scope, and each still says what
it would take to build. That is the list. D3 is delivered, and
[architecture.md §11](architecture.md)'s build order is what followed.
+546
View File
@@ -0,0 +1,546 @@
# Panorama
**Status:** Draft · 2026-09-19
**Companion to:** [requirements.md](requirements.md) §3.11 FR-MRG-1 … 11, D18, S15 · [architecture.md](architecture.md) §5.2, §6.2
The first merge (§3.11): several frames, rotated about one point, become one
photograph. This document is how that lands on the pipeline that exists now —
which stages, where each runs, how the composite is produced in chunks when it
is larger than any texture or any memory, what is ported from where, and what
the keypoint model may be under D8.
---
## 1. Why it is worth the work
The audience shoots panoramas and leaves the application to stitch them. That
is the same workflow break dust was (`spot-removal.md` §1): a RAW editor that
does everything but the one thing, and the photographer's work ends up in a
JPEG produced by a tool that never saw the RAW.
It is also the merge whose alignment problem is smallest. A panorama is a
rotation — three parameters per frame plus a focal length — with no depth to
recover. HDR merge and focus stacking share its data model (D18) and most of
its machinery (FR-MRG-3, 5, 6, 7, 10, 11 are written to be general); building
the panorama first builds the shared part on the easiest geometry.
## 2. Non-goals
- **Not structure-from-motion.** No translation is solved for. A hand-held set
with parallax gets its ghosts hidden by seam placement, and a set with real
parallax is not a panorama. COLMAP's front end is the right mental model;
its back end is the wrong problem.
- **Not a multi-source Version.** D18. The composite is a file, and nothing in
the catalog, the sidecar format or sync learns about cross-references.
- **Not boundary fill.** Painting pixels that were never captured is the pixel
editing §1.3 excludes. Auto-crop is the tool.
- **Not automatic.** The tool proposes an alignment and writes nothing until
the photographer confirms. Same rule as spot removal and D17, for the same
reason: a merge that silently omits or misplaces a frame is the failure this
application must not have.
- **Not HDR-panorama in one pass.** Until HDR merge exists on its own, a
bracketed panorama is bracketed frames merged first, then stitched.
## 3. What is new, precisely
Nearly all of it, unlike spot removal. The pipeline renders one source to one
texture; nothing in the tree detects keypoints, estimates a rotation, warps
into a projection, finds a seam, or blends a pyramid. What exists and is
reused:
| Exists | Where | Reused for |
|---|---|---|
| Render a source through the fused pass, with a linear f16 output mode | `dr-gpu` demosaic → `AdjustPass`, `OutputMode::LinearWorking` | FR-MRG-2's camera-space input, as a compose entry with no operations and the profile uniforms neutral (S15.3) |
| Tiled rendering with a priority scheduler | ARCH §5.3 | Pulling source tiles on demand into an output chunk (§5 below) |
| A non-CFA source entering the pipeline | `Demosaicer::from_rgba8` | The composite's decode path, if the container is a TIFF (S15.1) |
| DNG matrices read through rawler | `dr-decode::profile` | The composite's decode path, if the container is a DNG |
| Static-shape ONNX under tract, heads decoded in Rust | `dr-segment` | The keypoint model (§6) |
| A batch worker with its own `GpuContext`, activity row, cancel | `dr-ui::export` | FR-MRG-7 verbatim |
| The 16-bit TIFF encoder with metadata sub-IFDs | `dr-export::encode` | FR-MRG-3's writer, extended to linear samples |
| Multi-select in the grid | `collections_ui::selected` | The entry point |
New: a `core/dr-pano` crate holding the geometry (keypoints, matching, the
rotation solve), a set of WGSL passes in `dr-gpu` (reprojection, gain,
seam, pyramid blend), the chunked output driver, the container writer, and
the dialog.
## 4. The stages, and where each runs
FR-MRG-10 states the rule; this is the table it was written from.
| Stage | Cost shape | Runs on | Why |
|---|---|---|---|
| Source to camera-linear | per pixel, full res | GPU, the existing pipeline | It *is* the pipeline, stopped early |
| Keypoint detection | once per frame, at 1024 px | CPU, tract (NEON on the tablet) | Bounded by frame count, not output size. Same runtime faces and masks use. Hand-written WGSL convolutions for a model that runs five times would be work with no visible gain. |
| Descriptor matching | K² × D per pair | CPU, SIMD | 2048² × 64 × 10 pairs ≈ 3 GFLOP — tens of milliseconds |
| Rotation solve, bundle adjustment | 3N + 1 parameters, Levenberg–Marquardt | CPU | Microseconds. Not parallel work. |
| Preview reprojection | per pixel, proxy res | GPU, interactive | Projection and horizon changes re-warp N proxies at frame rate |
| Full-resolution warp | per output pixel | GPU, chunked (§5) | The heaviest thing in the application |
| Gain compensation | per overlap region | GPU reduction, then N scalars | Sums, on the histogram pass's pattern (ARCH §5.5) |
| Seam finding | per overlap pixel | GPU-friendly variant | Graph cut resists the GPU; a distance-transform or per-column DP seam does not. The algorithm is chosen for the GPU, not for the paper. |
| Multi-band blend | per pixel × levels | GPU, chunked | Laplacian pyramids are separable convolutions — the detail stage's shape |
| Encode | per pixel, once | CPU, streamed per chunk row | As export does |
## 5. Chunked in output space
FR-MRG-11 forbids holding the composite as one texture, and two facts force it
before memory does:
- `max_texture_dimension_2d` is 8192 on many mobile GPUs and 16384 on desktop.
A three-row panorama is routinely 20 000 px wide.
- Five 24 MP frames at working precision are ~1 GB together. The tablet does
not have it.
**The geometry is known before any full-resolution pixel exists.** Alignment
runs on proxies; what comes out is a rotation per frame, a focal length, a
projection and an output rectangle. From those, every output pixel's source
coordinates in every frame are a closed-form function. That is what makes
chunking simple rather than clever:
```
for each output chunk C (e.g. 2048 × 2048, in output space):
frames_in(C) = frames whose projected footprint intersects C
for each frame F in frames_in(C):
source tiles T(F, C) = tiles of F that project into C, plus a margin
render T(F, C) to scene-linear through the pipeline's tile cache
warp T(F, C) into C's coordinate frame ← GPU
gain-correct, seam, blend within C ← GPU, with overlap
read C back, encode its rows ← CPU, streamed
```
The working set is one chunk, its per-frame warped copies, and the source
tiles that fed them. It does not grow with the composite.
**The blend needs a margin.** A Laplacian pyramid of L levels reads
2^L pixels beyond the chunk edge; a chunk is therefore rendered with a margin
of that width and the margin discarded after the blend. Seams cross chunk
boundaries and must agree on both sides: the seam is found once at a reduced
resolution over the whole overlap (which fits — it is a mask, not an image),
then upsampled into each chunk. The same is true of gain: the scalars are
solved once from proxy-resolution overlaps and applied everywhere.
**Source tiles are the pipeline's tiles.** ARCH §5.3's cache keys by
`(VersionId, tile, zoom, graph_hash_prefix)`; the merge asks for tiles of a
neutral graph at zoom 1 and gets the same caching every other consumer does.
A tile pulled for one chunk is usually needed by the neighbouring chunk, and
stays hot for it.
### 5.1 The tap — S15.3, answered by reading the composer
The fused shader's order, fixed by `operation.rs`'s own tests: warp → as-shot
white balance → operations → base curve → camera matrix → store. The store is
either the display encode or, in `OutputMode::LinearWorking`, an unclipped
`rgba16float` of linear sRGB. That mode exists for the detail stage and is
selected from the operations, never by a caller flag, so that a shader and
the texture bound to it cannot disagree.
The merge wants the values *before* the curve and matrix (FR-MRG-2), and the
composer already makes that a matter of uniforms rather than structure: the
white balance, the matrix and the curve's active flag are all in the reserved
uniform block, and a fused pass with no operations, `as_shot_wb = 1`,
`cam_to_srgb = I` and `base_curve_last.z = 0` stores exactly camera-linear
RGB after the warp. So the tap is:
- `EditGraph::compose_camera_linear()` — the `LinearWorking` tail with an
empty operation list and identity framing, paired by name with
- `AdjustPass::render_camera_linear()` — binds the f16 target, fills the
reserved uniforms neutral instead of from the source, returns the texture,
- and a float readback beside the existing 8-bit one.
Nothing in the chain moves. **Precision:** the tap and every chunk buffer
after it should be `rgba32float`, not f16. A 14-bit sensor has 16 384 steps
to white; f16 has 2 048 in the top octave, and a composite that is going to
be re-developed deserves the sensor's precision. The cost is 2× on buffers
FR-MRG-11 already bounds.
**What the DNG carries as a consequence:** the first source's `Make`,
`Model` and `UniqueCameraModel` — so `base_curve::for_body` finds the 6D's
curve — its `ColorMatrix1`/`2` with illuminants, and its `AsShotNeutral`. The
composite then develops through the same profile as its sources, applied
once. The spike's 64 × 48 file (§8) already carries the matrix and neutral;
the body name is a string.
## 6. The keypoint model
FR-MRG-8: works without weights, better with them. The licence read comes
first (D13's lesson, S15.2).
| Model | Licence | Fits tract? | Position |
|---|---|---|---|
| **XFeat** (CVPR 2024) | Apache-2.0 | Plain convolutions, fully convolutional, the repo ships an ONNX export | **Chosen.** Fixed 1024 px input, dense heatmap and descriptor map out, NMS and top-K in Rust — the yolo26 pattern |
| DISK | Apache-2.0 | U-Net, static | Second choice; stronger descriptors, ~3–4× the compute |
| ALIKE | BSD-3 | Plain convolutions | Fallback if XFeat's export fails F6 |
| ALIKED | BSD-3 | Deformable convolution in the descriptor head | Unlikely to load |
| SuperPoint, SuperGlue, R2D2, SiLK, MASt3R | non-commercial | — | Out on licence |
| LightGlue | Apache-2.0 | Transformer over a variable keypoint count | Not until mutual-nearest-neighbour matching fails on a real set |
**S15.2, 2026-09-19: XFeat loads under tract.** `tools/export-xfeat.sh`
exports the network alone at 768×1024 — thirteen operator types, all
standard: `Conv`, `InstanceNormalization`, `AveragePool`, `Resize`, `Slice`,
`Transpose`, `Reshape`, `Concat`, `Add`, `Relu`, `Sigmoid`, `ReduceMean`,
`Unsqueeze` — and
[`examples/onnx_probe.rs`](../../core/dr-segment/examples/onnx_probe.rs) loads
the 2.8 MB file through the app's own `ort`-over-tract backend with nothing
unsupported, in 28 ms, and runs it in **~300 ms on the reference desktop's
CPU**. The weights ship as `models/keypoints/xfeat-1024.onnx`, recorded in
`models/LICENCE.md`. **On the tablet** (S15.4's CPU half, same day):
`tools/onnx-probe-on-device.sh` cross-builds the probe, and the same file
runs in **~400 ms per frame** on the reference tablet's NEON cores (ROD2-W09,
SM8635), with output ranges identical to the desktop's — inside NFR-MRG-1's
1 s per frame with room to spare, and 1.3× the desktop rather than the 2×
faces.md §9 measured for its scan. Still to do: a keypoint-level comparison
against the PyTorch reference once the Rust decoder exists — the probe
proves the graph runs, not that the numbers match.
The outputs are three maps at 1/8 resolution, 96×128 for the export size:
64-channel descriptors, 65-channel keypoint logits (each 8×8 cell's position
plus "none"), and a reliability heatmap. The Rust decoder is: softmax over the
65, pixel-shuffle the first 64 to full resolution, 5×5 non-maximum
suppression, top-k by reliability, bilinear sampling of the descriptor at
each keypoint, L2 normalise. That is `detectAndCompute` in the reference,
minus the network.
Without weights: AKAZE (BSD, `akaze` from rust-cv), which is adequate on
well-textured overlaps and worse on sky, repeated structure and exposure
drift — which is where a learned detector earns its place.
Matching is mutual nearest neighbour with a ratio test, then RANSAC on a
rotation model. For a panorama — one lens, near-pure rotation, 20–40 %
overlap — that is what Hugin and OpenCV's stitcher use, and it is enough.
## 7. What is ported from where
Nothing is linked; everything is read.
| Source | Licence | Taken |
|---|---|---|
| OpenCV `modules/stitching` | Apache-2.0 | The stage layout — Brown & Lowe (2007) as a set of small classes with one job each — and the warpers' projection maths |
| OpenPano (ppwwyyxx) | MIT (verify on read) | The estimation and bundle-adjustment maths, function by function, with outputs diffed against it |
| enblend-enfuse | GPLv2+ | Seam-line optimisation and Burt–Adelson multi-band blending |
| Hugin `nona` | GPLv2+ | The GLSL remapper, as the reference for the WGSL warp |
The golden set (§8 of the requirements) is OpenCV's stitcher on the same
inputs: a reference output to compare against, within a tolerance calibrated
the way S9 calibrates R1.
## 8. The output file
FR-MRG-3. A linear DNG at the source's native scale: `u16` samples on the
first source's black-subtracted scale, `WhiteLevel` = its white minus its
black (13 023 for the 6D set: 15 070 − 2 047), `BlackLevel` = 0. Not rescaled
to 65 535 — the sensor had 14 bits and the file says so, and a value the
sensor could not have produced is not invented by a multiply. The first
source's `Make`, `Model`, `UniqueCameraModel`, `ColorMatrix1/2`,
`CalibrationIlluminant1/2`, `AsShotNeutral` and EXIF are carried, so the
composite develops through the same profile as its sources. Named from the
first source with a `-pano` suffix, beside it.
Three samples per pixel rather than a CFA: the warp resamples, and there is no
sensor grid to mosaic back onto. Nothing else about being a RAW is lost —
no white balance, no curve, no matrix, no clip has been applied — and the
photographer develops the panorama afterwards as one photograph.
The sources are portrait frames in the 6D set: `Orientation` is applied
before alignment (learned features are not rotation-invariant) and the
composite is written upright with `Orientation = 1`.
Two containers were candidates and S15.1 decided, on 2026-09-19:
- **Linear DNG.** `PhotometricInterpretation = LinearRaw`, three samples per
pixel, `ColorMatrix1` carried from the first source. Re-enters through
rawler as `Format::Dng` with no new decode path, *if* rawler reads it back.
What Lightroom writes.
- **Float TIFF.** `SampleFormat = IEEEFP`, 16 or 32 bits, an ICC profile for
the working space. Needs `Format::Tiff` and a decode path, but the writer is
the existing encoder with a different sample type, and nothing about it is
uncertain.
**Linear DNG.** [`examples/linear_dng.rs`](../../core/dr-decode/examples/linear_dng.rs)
hand-rolls a 64 × 48 `LinearRaw` DNG — one IFD, uncompressed 16-bit RGB,
`DNGVersion`, `ColorMatrix1`, `AsShotNeutral`, `CalibrationIlluminant1` — and
rawler 0.7 reads it back: `cpp 3`, the samples interleaved as written, the
matrix parsed into the camera definition, and `CameraProfile::extract` builds
the same profile it would for a camera file. ImageMagick's libraw reads the
same bytes. What does *not* yet work is `dr_decode::decode`, which accepts the
file as CFA and hands the pipeline three times the samples it expects: the
`cpp == 3` branch is the work, and it is the only decode work.
The composite therefore enters the pipeline as a non-CFA, *linear* source —
`from_rgba8`'s sibling with `non_linear = false` and the colour matrix carried
from the DNG — and is developed as any RAW is. The writer is the example's
IFD, grown up: tiled rather than one strip (FR-MRG-11 encodes per chunk), and
carrying the first source's EXIF in a sub-IFD as `dr-export` already does.
## 9. Interaction
- The entry is the grid's selection: two or more images, one action, "Merge
to panorama". One image, or images from different roots, and the action
says why it is unavailable.
- The dialog shows the aligned proxies in the chosen projection, with the
projection, horizon and crop controls of FR-MRG-4, and the per-frame
residuals. A frame that failed to align is named there (FR-MRG-5), and the
merge cannot be confirmed with it in the set.
- Confirm starts the FR-MRG-7 job. The composite appears in the grid when the
file is written and catalogued, beside its sources, with the merge as the
first entry in its history.
## 10. Order of work
1. **S15**, all four, before anything else. (1) and (2) are a day each and
either can change the design.
2. `dr-pano`: keypoints (AKAZE first, XFeat when S15.2 passes), matching,
RANSAC, rotation solve. Unit-tested against synthetic rotations of one
frame, where the answer is known exactly.
3. The working-space tap, and the preview reprojection pass. At this point the
dialog can show an alignment.
4. The chunked driver with a feathered blend — the whole path end to end,
writing a file, before the blend is good.
5. Gain, seams, multi-band.
6. The container, the catalog entry, provenance, the history entry.
7. Tablet: NFR-MRG-1's figure, and FR-MRG-9's ceiling.
## 11. Where it stands — 2026-09-19, end of the first day
Built, on branch `merge/panorama`, in the order §10 gave:
| Piece | Where | State |
|---|---|---|
| Geometry: keypoints, matching, homography, focal, bundle adjustment, projections | `core/dr-pano` | Done; 33 tests without a model; the fixture aligns in 4.5 s |
| XFeat at two shapes under tract | `models/keypoints`, `dr_pano::xfeat` | Done; 300 ms/frame desktop, 400 ms tablet |
| The camera-space tap | `OutputMode::CameraLinear`, `AdjustPass::render_camera_linear` | Done, `rgba32float`, tiles by view rect |
| Linear DNG writer, streamed | `dr_export::write_linear_dng` | Done; rawler reads it back |
| A three-sample `RawImage` re-entering the pipeline | `dr-decode`, `DemosaicedImage::from_linear_rgb16` | Done |
| Warp, accumulate, resolve, chunk by chunk | `dr_gpu::MergePass`, `merge.wgsl` | Done; feathered blend, scalar gain |
| The job: load, proxies, align, gains, confirm, merge, provenance | `dr_ui::merge` | Done; `examples/merge.rs` drives it headless |
| The page: table, preview, projection, Merge/Stop/Back; the grid's button | `merge.slint`, `merge_ui.rs` | Done; `DARKROOM_START_MERGE=a.CR2,b.CR2` lands on it |
| Placement beside the sources through the outbox, rescan | `merge_ui.rs` | Done, untested against a server |
**Measured on the fixture (desktop, 12 × 20 MP, Intel adapter):** proxies
and keypoints 4 s, alignment 4.5–12.6 s (load-sensitive: the matcher is
every core), the merge **26 s for a 22 993 × 5 980 composite** in twelve
bands of 2048 × 512 chunks, 45 s all told, an 825 MB DNG. NFR-MRG-1's 60 s
holds on the desktop with room; the tablet's figure is still S15.4's open
half.
**Open, in the order they matter:**
1. **Auto-crop (FR-MRG-4).** The merge returns a coverage mask per band and
the file carries the black border. The largest inscribed rectangle over
the coverage, then the DNG's `DefaultCropOrigin`/`DefaultCropSize`, so
nothing is thrown away and the develop view opens on the picture.
2. **Seams and the pyramid** (§10 step 5). The feather hides exposure and
small misalignment; parallax on the near slope will show as a soft
double edge at 1:1.
3. **Vignetting in the tap.** The lens profile's distortion is applied
before the fetch; its vignetting is an operation and is not. Frame edges
are darker than their centres by the lens's falloff, and the feather
averages them into the overlaps.
4. **The tablet:** memory (twelve 40 MB sensor buffers on the CPU, one
demosaiced frame at a time on the GPU), the figure, and FR-MRG-9's
ceiling.
5. **Horizon and drag-to-correct (FR-MRG-4, the proposed 4a).** The
alignment failed on nothing in the fixture; the interaction waits for a
set it fails on.
6. **`derived_from` names sources by file name**, not content hash: the
catalog's `content_hash` is null for most images most of the time. The
hash can join it when the catalog has one.
## 12. Filling the border instead of cropping it — MI-GAN, read and measured 2026-09-19
Raised after the first merges: the ragged border a cylinder leaves could be
*filled* rather than cropped away. FR-MRG-4 says no boundary fill, on
§1.3's "not a pixel editor"; this is the evidence for deciding whether to
revise that, not a revision.
**The candidate: MI-GAN** (Sargsyan et al., ICCV 2023, Picsart AI Research).
Image inpainting designed for mobile: ~6 M parameters, plain convolutions —
no FFT, no attention — so it quantises to int8 and runs on a phone's DSP,
with quality close to LaMa and CoModGAN.
**Licence: MIT, code and weights alike** (`LICENSE` and `LICENSE-WEIGHTS`
in the repository, read the same day). The cleanest position of any model
in the tree — GPL-compatible, store-compatible, no grant to read around.
**Export.** The HuggingFace ONNX files are the *pipeline* — uint8 image and
mask in, crop-around-mask, resize and blend inside the graph, every
dimension dynamic — and tract refuses them (F6 again). The bare generator
exports cleanly from the `migan_512_places2.pt` state dict at a fixed
`1×4×512×512` (`export_migan.py` in the spike directory; the input is
`mask − 0.5` and the masked RGB in −1..1, the output RGB in −1..1, the
caller composites). After slimming the graph is **six operator types**:
`Add, Clip, Conv, LeakyRelu, Mul, Resize`. 28 MB.
**Under tract on the reference desktop: loads in 53 ms, runs in 7.4 s per
512 × 512 tile, f32.** That is the number. The fixture's border is two
ragged bands across 22 993 px — roughly ninety 512-px tiles at full
resolution — so a CPU-f32 fill is ten minutes on the desktop and longer on
the tablet. Three ways to make it viable, none built:
1. **Fill at a quarter of the resolution and upsample.** Sky and scree
tolerate it; twenty-odd tiles, about three minutes on the desktop CPU. A
background job with the outbox's patience, not an interactive one.
2. **int8 on the tablet's Hexagon through QNN**, where the plain-conv design
is the point and the whole graph should run in milliseconds. The setup
exists from the eye-state work; MI-GAN is a candidate for the same path.
3. **A WGSL runtime for those six operators.** A project of its own, and
the only route that would make it interactive on the desktop.
Whichever, the fill is a *proposal* under FR-MRG-1's rule — shown, then
confirmed — and it would sit beside the crop, not replace it: the crop is
free and honest, the fill is invented pixels, and the photographer chooses.
## 13. The fill, built — 2026-09-19, evening
Built the same day on `merge/fill`, on the engine (S16) rather than tract,
and FR-MRG-4 revised to admit it: the border is *cropped or filled*, the
photographer's choice, the crop the default.
**What runs.** `dr_pano::fill` is the engine-independent half: an
`Inpainter` trait (a 512-px tile in, the same tile out) and `fill_border`,
which owns everything the model does not — which tiles, what context, how
to blend. `dr_pano::migan::MiGan` is the trait over the shipped generator
under `dr_inference_engine` with the new `Role::Inpainter`, so it takes
whichever rung the device has. The merge job runs the fill at **half the
composite's resolution**, in a display-ish space (white balance, camera
matrix, gamma — invertible, so the result goes back to camera-linear and
into the same linear DNG), and the full-resolution merge samples the fill
where no frame reached.
**What the spike taught, tried in order and kept or dropped.**
1. *Context across the coverage edge.* MI-GAN was trained on holes inside
pictures; given a hole at the picture's edge it invents a structure along
the open side (white streaks in the sky, on the first try). The known
content is therefore **mirrored** across the coverage edge into the hole
and into a 256-px ring, column by column for the top and bottom bands
and row by row for the sides; the model interpolates between real and
mirrored sky rather than extrapolating into nothing. *Replicated* rows
(the edge row continued flat) streaked the grass; a detrended mix (tone
replicated, texture mirrored) smeared; a low-pass extrapolation banded.
Mirror stays.
2. *Coarse to fine.* One pass at the working resolution let the boundary
leak in — each 512 tile saw only its own corner of the hole. So a
**coarse pass at a quarter** decides the structure with the whole border
in a few tiles, and **fine passes in 96-px bands** from the real edge
outward regenerate texture, each band the only unknown with the previous
band on its near side and the upsampled coarse fill on its far side.
3. *The seam.* A hard cut between real and invented showed as a sharpness
step. The known mask is eroded by a **24-px feather** (48 at half
resolution) and the fill blended in across that margin by distance to
the real edge, smoothstep.
4. *Partial pixels.* The seams were still visible until the cause was found
upstream of the fill: the camera-space tap stored **black with alpha 1**
for a pixel the lens correction pushed off the sensor, and the warp
averaged it in — a dark, poorly interpolated fringe along every frame's
edge that the fill then continued. `OutputMode::CameraLinear` now
stores alpha 0 for a pixel that is not there and the merge's warp
weights by the sampled alpha, so the fringe never enters the composite.
The mask erosion before the fill dropped from 16 px to 4.
5. *What is still wrong, and why it ships anyway.* With the seams gone the
content itself is the problem in the deep corners: the model, trained
on Places2, puts bright cloud-and-peak shapes into a sky hole and a
water-like band under grass — its prior for "top of a picture" and
"bottom of a landscape", not anything in the context (the same shapes
appear with the mirror capped, uncapped, and on the CPU as on TensorRT).
Thin borders are fine; that is most of a hand-held sweep. So the fill
ships **experimental**: opt-in, previewed, its knobs on the page and
in the sidecar, and `cargo run -p dr-ui --example fill` re-runs any
merge's dumped input (`DR_FILL_DUMP=dir`) stage by stage in seconds so
the next attempt is made from the picture, not from a seven-minute
merge. Candidates for that attempt: a context that is not a mirror at
all in deep holes (the coarse pass's own answer, iterated), a sky
detector that fills sky by extrapolating the gradient and leaves the
model to texture, or a different model.
**Measured, the fixture's twelve frames (22 991 × 5 978), 348 tiles at
half resolution.** 312 s on ONNX Runtime's CPU pool on the reference
desktop (≈ 0.8 s a tile). On TensorRT fp16: **100 s**, of which 60 ms a
tile was the engine hashing the 28 MB model on every acquire (fixed, the
hash is taken at open) and 150 ms a tile the GPU itself — throttled:
`trtexec` on the same engine read 23 ms at noon on a cool machine and
152 ms that evening after two hours of builds, nvidia-smi showing SW power
cap and thermal slowdown. Cool, the fill is ~10 s. The TensorRT engine
compiles once, in 13 minutes, cached under the inference directory.
**The runtime is a packaging matter.** Arch's `onnxruntime-opt-cuda` has
no TensorRT provider ("not enabled in this build") and its CUDA provider
does not load against cuDNN 9, so on this machine the app fell to ORT CPU
until the official `onnxruntime-linux-x64-gpu_cuda13` tarball (1.30.0,
which links the system CUDA 13.4 and TensorRT 10.16) was unpacked and
named with `DARKROOM_ORT_DIR`; `/usr/lib/darkroom` is searched too, for a
package that ships it. §12's point 2 for the tablet is unchanged.
**On the page.** A *Border* choice beside the projection — *Crop to the
picture* / *Fill the border* — with a caption saying what the fill is; the
preview re-renders filled when chosen, at preview resolution, so the
choice is seen before it is confirmed (FR-MRG-1). Greyed out with the reason
when `migan-512.onnx` is not in the model directory. Under the fill, while
it is experimental, its six knobs as sliders — working scale, edge
erosion, coarse pass, band width, mirror depth, seam feather — each
committing a redraw of the preview. A filled merge's sidecar says `border
filled` with the knobs used, and its default crop is still the inscribed
rectangle.
**Ships.** `models/inpaint/migan-512.onnx` (LFS, 28 MB, MIT,
`models/LICENCE.md`), installed by the PKGBUILD and unpacked by the APK
beside the face and scene models; `tools/export-migan.sh` regenerates it
from the upstream checkpoint.
## 14. The fill, trained — 2026-09-20
§13.5 named the remaining fault: in a deep corner the stock model puts its
Places2 prior — clouds, peaks, a water line — into a hole, because it was
trained on holes *inside* pictures and a panorama's border is a hole with
the picture on one side and nothing on the other. Every ring (mirror,
replicate, detrend) treated the symptom. The fix is a model that has seen
the real thing: **MI-GAN's 512 generator fine-tuned on border-shaped voids
cut from the user's own photographs**, in a separate repository
(`darkroom-infill`, beside this one), so the truth beyond the void is known
and the model learns one-sided extrapolation.
**What it was trained on.** Voids made the way this merge makes them:
frames with a yaw, a common pitch and per-frame roll, projected onto the
cylinder and rasterised, the canvas their union's bounding box, the void
the canvas outside the union — arcs where straight edges bent, cusps where
frames meet, the bow-tie wedge at a corner (a third of tiles are cut at a
canvas corner). Voids to 256 px deep at a 512 tile. Half the deep tiles
train the *second pass*: a no-grad first pass fills the tile, its nearest
band (64–256 px) is marked known, and the remaining void is the example —
so the model continues its own output without drift, which is how `fill`
runs deep voids. Data: ~7 400 pictures — the 1024-px proxy tier of the
library and ~2 000 raws sampled evenly across every year, developed at
half size. Loss: hole-weighted L1, VGG16 perceptual, a hinge PatchGAN.
One night on the reference desktop's RTX 3050.
**What changed here.** `FillParams::mirror_depth` **0** is now "no ring":
the void reaches the tile's edge with nothing beyond, and beyond the band
being filled the void stays *unknown* rather than presented as known coarse
fill — the two conditions the model was trained under. Defaults: mirror 0,
coarse 1 (the coarse pass seeds nothing the model is allowed to see), band
192. The ring remains on the page for the stock model's sake, at any depth
above zero. The model file is a drop-in (`models/inpaint/migan-512.onnx`,
same six operators, same tensors) and the engine loads it unchanged.
**Measured, 240 held-out tiles with projection-shaped voids (PSNR in the
hole, dB / LPIPS on the composite), stock → shipped (step 3 607):** edge
16.9 → 18.5 / 0.121 → 0.136; corner 14.6 → 16.3 / 0.183 → 0.205; interior
18.2 → 19.3 / 0.051 → 0.056. Read both columns: the fine-tune gains ~2 dB
on edges and corners because it stops inventing objects, and *loses* on
LPIPS because what it paints in a deep void is smoother than the stock
model's confident wrong texture — LPIPS rewards texture, right or wrong.
On the fixture's dump at half resolution (the merge's working size) the
sky corners are sky, with no structure and a faint tone step at worst;
the ground bands carry a fine texture at the right tone, softer than the
real scree above them. The stock model's top-left corner on the same
dump is a glowing invented structure. The training's own record — what
each loss weighting did, and the two runs abandoned (blur under L1 in
the hole; a brick pattern under a strong adversarial term against a
discriminator that had not learned) — is `runs/` in `darkroom-infill`.
**What is still wrong.** The ground fill is softer than its context —
texture, not structure, is what a night on a laptop GPU could not finish.
The levers, in order: a discriminator that learns (a pretrained one —
MI-GAN's own from the unfused checkpoint — instead of a PatchGAN from
scratch), feature matching, and more steps at 512. FR-MRG-4's
*experimental* stays.
File diff suppressed because it is too large Load Diff
+503
View File
@@ -0,0 +1,503 @@
# Region segmentation for local masking
Spec for **S15**, the spike that decides how DarkRoom finds the boundaries a local mask snaps to.
Local adjustments (FR-DEV-3, "linear gradient, radial gradient, and brush masks") need more than
placement handles to be competitive. The interactions that matter — click to select a region, drag a
contour that clings to an edge, paint without crossing a boundary — all need the same thing
underneath: **a map of where the image's regions are.**
There are two credible ways to produce that map and they are not obviously ordered. This document
specifies both, specifies the third option of combining them, and fixes the measurements that decide
between them *before* any of them is built.
---
## 1. Why this is a spike and not a build
Three properties make the choice expensive to get wrong.
**It sets the mask representation.** If regions exist, a mask is a *set of region ids* — integers,
diffable, mergeable at node level under FR-NC-9, cheap in a sidecar. If they don't, a mask is a
raster, and rasters are none of those things. This is the decision that is expensive to retrofit;
everything else in local masking sits on top of it.
**One arm collides with a settled policy.** D13 records that every dependency choice in this project
has gone the same way — rustls over aws-lc-rs, bundled SQLite, a Rust Lensfun port, zune-jpeg — to
avoid a C dependency under the Android NDK, and names `ort` as the largest exception that policy
would tolerate. Arm B needs exactly that exception. Arm A needs no new dependency at all. That
asymmetry is not a tiebreak, it is most of the cost difference, and it should be priced honestly
rather than discovered at packaging time.
**Model licensing is a distribution blocker.** D13 already establishes this for the face pipeline,
and the same reading applies here — see §7. It is a licence-reading exercise, not a research
question, and it comes first.
---
## 2. The common interface
Both arms produce the same thing. This is what makes them comparable, and what lets the choice be
deferred behind a seam rather than baked into every consumer.
```rust
/// A partition of the image into labelled regions.
pub struct RegionField {
/// Per-pixel region id at proxy resolution. R32Uint on the GPU.
labels: Texture,
/// Per-region summary: pixel count, bounding box, mean colour,
/// and (arm B only) a semantic class id.
regions: Vec<Region>,
/// Boundary strength per adjacent region pair — the edge weight
/// the merge tree is built from and the cost field reads.
adjacency: Vec<(RegionId, RegionId, f32)>,
}
```
Two consumers sit on it, and neither knows which arm produced it:
**A cost field, for contour snapping.** Live-wire — Dijkstra from the last anchor to the cursor over
a per-pixel cost that is *low* on boundaries. The cost is a **sum of terms**, which is the property
that matters: image gradient is always available, region boundary strength is added when a
`RegionField` exists, and a semantic boundary term is added when a model is present. Each source
improves the snap without changing the interface, so the arms are not exclusive here even in
principle.
**A region set, for click selection.** Click reads the label under the cursor; the mask is
`label(px) ∈ selected`. Add and subtract are set operations on ids. No flood fill, no readback, no
iteration — the whole reason the precomputed map is worth having.
Both consumers are built once, in the spike, and shared by both arms. A comparison in which each arm
gets its own consumer measures the consumers.
---
## 3. Arm A — multiscale watershed
No model, no new dependency, deterministic, works on any image.
**Gradient.** Sobel magnitude over a perceptual luma plus chroma distance, not camera-space RGB —
channel-weighted RGB gradient reads a saturated red edge as weaker than it looks. Computed after
demosaic and denoise, before the edit graph, so an exposure change does not invalidate it.
**Pre-smoothing is not optional.** Raw watershed on a noisy file makes every grain its own basin.
A guided or bilateral pre-filter, with strength tied to the file's ISO, is part of the arm rather
than a refinement of it.
**Basins.** Each pixel points downhill to its steepest neighbour; pointer-jumping resolves every
pixel to its basin root in log passes. Two compute shaders and a dispatch loop.
**The hierarchy is the cheap part.** Build the region adjacency graph, sort edges by boundary
strength, union-find over them, and *record the merge order*. That recording is the merge tree — a
click selects a leaf, and a scroll walks up through progressively coarser merges. Textbook Kruskal on
a graph of a few thousand nodes.
That node count is why the tree build is a legitimate CPU operation: it runs on the adjacency graph,
not on pixels. Pixels stay on the GPU, the graph is CPU-side — the same split ARCH §3.4 and §6.1
already draw for the edit graph, so no exception to the no-readback rule is needed.
**Known weaknesses, to be measured rather than argued about:** over-segmentation on noise and
texture, weak boundaries where contrast is low but semantics are obvious (a pale sky meeting a pale
wall), and a granularity ladder that is geometric rather than meaningful — level 7 is *a* coarser
partition, not necessarily *the* object.
---
## 4. Arm B — semantic segmentation
YOLO26-seg pretrained on ADE20K, run through `ort`.
ADE20K's 150 classes include stuff — sky, vegetation, water, wall, road — which is a far better
vocabulary for photography than COCO's 80 thing-classes. "That patch of sky" is a class here. The
nano variant is ~1.6M parameters, which is genuinely mobile-viable in a way SAM never was.
**It produces a flat partition with class ids**, so it populates `RegionField` directly: connected
components of the class map become regions, class boundaries become adjacency edges.
**What it does not produce is a hierarchy.** One partition at one semantic granularity. Click "sky"
and you get all the sky; there is no level at which you get *this part* of the sky. Adjacent
same-class regions merge whether or not you wanted them to — two different walls are one wall.
**Boundaries are semantically right and geometrically soft.** Internal stride is 4–8, upsampled to
H×W, so the class map is confident about *which* side of the boundary a pixel is on and vague about
*where* the boundary is to the pixel. Acceptable for biasing a contour. Not acceptable as a mask
edge at 100% zoom.
**Costs it brings that arm A does not:** a C dependency on the Android NDK against D13's policy, a
model to distribute and cache, an AGPL question (§7), an inference runtime per platform, and output
whose determinism across drivers is unproven (§6).
---
## 5. Arm C — semantic as a merge prior
The arms are not alternatives, and a comparison that omits their combination is a false dichotomy.
Weight each region-adjacency edge in arm A's union-find by boundary strength **and** by whether the
two regions share a semantic class. Regions that agree semantically merge earlier.
The result is a hierarchy whose coarse levels align with semantic objects and whose fine levels stay
pixel-accurate — the model doing what models are good at, which is knowing what things *are*, and
watershed doing what it is good at, which is knowing where boundaries are, exactly, at every scale.
It also repairs arm B's two weaknesses at once: the soft boundary is replaced by the watershed
boundary underneath it, and the missing granularity ladder is arm A's.
Arm C is the expected winner on quality. The question the spike actually has to answer is therefore
not "which is better" but **how much better than arm A alone, and is that increment worth D13's
cost.** §8 fixes that threshold in advance.
---
## 6. What gets measured
Per arm, over the corpus in §9, using the shared consumers from §2.
| # | Measure | Method | Why it decides anything |
|---|---|---|---|
| **M1** | **Interactions to target mask** | Clicks plus scroll steps to reach ≥95% IoU against a hand-traced mask | The real UX metric. "How many actions to get the mask I meant" is what a user experiences |
| **M2** | **Boundary accuracy** | Precision/recall of snapped-contour pixels within a 2px slack of the hand trace | Whether the edge survives 100% zoom, where masks are actually judged |
| **M3** | **Granularity coverage** | Per case, yes/no: does *any* hierarchy level produce the target region? | A hard failure mode. Arm B is expected to fail this wherever the target is not a class |
| **M4** | **Out-of-vocabulary behaviour** | M1 and M3 restricted to the OOV subset | Whether the arm degrades gracefully or produces nothing usable off-distribution |
| **M5** | **Determinism** | Same input twice on one machine; then across Mesa/AMD, NVIDIA, and Adreno | Gates whether a label field can be a cache key at all — see below |
| **M6** | **Precompute cost** | ms at proxy resolution and peak memory, on the reference desktop and one Android device | Whether it fits a background prefetch alongside the proxy |
| **M7** | **Distribution cost** | Added binary size, model size, new native dependencies, licence | D13's axis. Priced, not assumed |
**M5 deserves its own note, and it is a risk for arm A too.** ARCH §6.13 holds that cache keys are
computed over CPU-side *integer* state because GPU float results diverge across vendors. A label
field is integer state — but it is *derived from* float gradient arithmetic, so a boundary sitting
exactly between two basins could resolve differently on Adreno than on Mesa. If either arm proves
non-deterministic across vendors, its output cannot be a cache key and cannot round-trip through a
sidecar as region ids, which would push masks back toward rasters and undo most of §1's argument.
This is the measurement most likely to invalidate the whole approach, and it should be run early
rather than last.
---
## 7. Licence reading — before any code
D13's position applies unchanged: discovering at packaging time that a feature cannot ship is the
expensive failure, and it is entirely avoidable.
**Ultralytics ships YOLO under AGPL-3.0** — confirmed 2026-08-17 by reading `LICENSE` at the head of
`github.com/ultralytics/ultralytics`, which is the GNU Affero General Public License v3 verbatim.
That is deliberate on their part; the commercial licence is their business model.
GPLv3 §13 explicitly permits the combination, so this is *not* the blocker the InsightFace
non-commercial weights were: it is redistributable. But the combined work becomes effectively AGPL,
which is a change to DarkRoom's licensing posture rather than a dependency detail, and it needs to be
a decision made on purpose.
Still to verify before writing any of arm B:
- The licence on YOLO26 specifically, and on the ADE20K-pretrained weights *separately* from the
framework code — they are not necessarily the same grant.
- Whether ADE20K's own terms permit redistribution of weights derived from it.
- Whether AGPL is acceptable for DarkRoom, given Flatpak, F-Droid and Play distribution
(NFR-COMPAT-2).
**Arm A raises none of these questions**, which is worth stating plainly as part of its cost.
---
## 8. Decision criteria, fixed in advance
Stated now so the result cannot be rationalised afterwards.
- **Arm A ships alone** if it reaches within **one interaction** (M1) of arm C on the scene subset
*and* dominates arm C on the OOV subset (M4). The semantic increment does not then justify a C
dependency, an AGPL conversion, and a per-platform inference runtime.
- **Arm C ships** if it beats arm A by **two or more interactions** on the scene subset without
regressing OOV. That is a large enough difference to be felt in ordinary use, and it is what would
justify reopening D13.
- **Arm B never ships alone.** M3 is expected to fail on anything that is not an ADE20K class, and an
arm with a hard failure mode and no fallback is not a selection tool. If it surprises us and passes
M3 broadly, that is a genuine finding and this criterion is revisited on the evidence.
- **If M5 fails for an arm across vendors**, that arm cannot carry region ids into the sidecar
regardless of how it scored elsewhere.
---
## 9. Corpus
Roughly 24 images from a real library — three per category — hand-traced once and reused across all
arms. Categories chosen for the failure modes they provoke, not for coverage:
| Category | Provokes |
|---|---|
| Gradient sky | Low-contrast boundary; watershed banding |
| Foliage against sky | High-frequency boundary — arm A over-segments, arm B blurs |
| Hair against a busy background | The classic hard mask edge |
| Out-of-focus background | No edges at all; tests graceful failure in both |
| High-ISO noise | Arm A's known weakness; tests whether pre-smoothing is sufficient |
| Backlit silhouette | Strong unambiguous edge — the control case |
| Macro, abstract, still life | **OOV for ADE20K.** Arm B expected to fail M3 here |
| Architectural detail | Repeated structure; arm B merges distinct walls into one class |
Hand-tracing 24 masks is a couple of hours and it is what makes M1 and M2 mean anything. Without
ground truth this comparison is two demos and a preference.
---
## 10. Deliverables
Nothing in the UI, nothing in the graph, nothing in the sidecar.
- `core/dr-gpu/src/shaders/watershed.wgsl` — gradient and basin propagation.
- `core/dr-gpu/src/segment.rs` — the passes, producing a `RegionField`.
- RAG construction and the union-find merge tree as a pure-CPU module with unit tests and no device,
so the hierarchy is testable headless the way `dr-pipeline` is (ARCH §6.5a).
- The two shared consumers from §2 — live-wire over a summable cost field, and region-set selection.
- `core/dr-gpu/examples/segment.rs` — false-coloured PNGs at four or five hierarchy levels, plus the
M1/M2 numbers against the traced corpus.
It graduates to `core/dr-segment` if it ships; that is not a spike decision.
---
## 11. Order
1. **Licence reading (§7).** Hours, and it can eliminate arm B before anything is built.
2. **Arm A, and the shared consumers.** About a day. Look at the false-coloured PNGs — if the
granularity ladder does not feel right, nothing downstream matters and that is worth knowing
immediately.
3. **M5 across vendors, early.** It is the measurement that can invalidate the region-id
representation entirely, and it wants knowing before the corpus work is invested.
4. **The traced corpus, then M1–M4 on arm A.** Establishes the baseline every other arm is judged
against.
5. **Arms B and C**, only if §7 cleared and arm A's baseline leaves room worth closing.
Arm A is a day and needs no model, no runtime, no licence and no new dependency. It is also the
substrate every model-based arm writes into — so it is first regardless of how the comparison
eventually lands.
---
## 12. Arm A results
Built 2026-08-17. `core/dr-gpu/src/{segment.rs,hierarchy.rs}`,
`shaders/watershed.wgsl`, `examples/segment.rs`. 15 tests, 11 of them device-free.
**It works, and the hierarchy is not the expensive part.** On a 1200×800 synthetic at blur radius 2,
release build, RTX 3050 laptop: 6,730 basins and 19,223 boundaries found in **67 ms including the
readback**, and the merge tree built from them in **0.2 ms**. The tree is ~0.3% of the cost. The
estimate that priced it as a week's work was wrong by about two orders of magnitude, and the reason
is worth recording: it is Kruskal over a few thousand nodes, not a segmentation algorithm.
**The granularity ladder behaves.** At the fine end the background fragments badly — a smooth tonal
ramp bands into horizontal strips, and flat areas break into diagonal chains (see below). By
`cut_to(300)` all of that is gone: the hard-edged disc is exactly one region, the whole gradient
background is one region, and only genuine noise still fragments. The over-segmentation is absorbed
by the merge order rather than needing to be prevented, which is the property the whole design rests
on.
**Pre-smoothing is the knob it was claimed to be.** Radius 2 leaves the noisy corner fragmented at
300 regions; radius 5 largely clears it. Tying it to ISO is the right control.
Three findings that change what comes next:
**F1 — plateaux fragment into diagonal chains.** In an exactly flat region every pixel's steepest
descent is a tie, and the (value, index) tie-break sends them all up-left, so a plateau resolves into
diagonal streaks rather than one basin. Harmless here because those saddles are ~0 and the tree
merges them first — but a real sky or wall is a large plateau, and relying on the hierarchy to clean
up an artefact of the flow pass is fragile. The principled fix is a **lower-complete transform**: one
extra pass giving plateau pixels a gradient toward their nearest descending exit. Standard, cheap,
and worth doing before the corpus work.
*Attempted, and parked.* The pass exists — `plateau_init` seeds every pixel
that has a strictly lower neighbour, `plateau_step` carries a breadth-first
distance inward within a level set, and `flow` takes that distance as the
second key of a lexicographic tie-break. Bindings, ping-pong and dispatch were
all checked and are right. It is nonetheless a **measured no-op**: with a test
comparing the labelling at one iteration against sixty-four, *zero* of 9216
pixels change basin. That test is committed and ignored rather than deleted,
because it is the thing that turned "we think this works" into a fact.
Three explanations were tried and none of them was it. Exact float equality is
certainly wrong — a gradient computed from 8-bit samples is never exactly
equal across a region the eye calls flat — and a `LEVEL_EPS` tolerance now
replaces `==` and `<` in all three comparisons; it did not change the outcome.
Nor did the test image: a flat disc, a terraced disc and a constant-slope ramp
all behave identically. Worth knowing for whoever picks this up: on a
gradient-*magnitude* watershed, every flat region of the picture sits at
gradient zero, which is the global minimum, and a plateau with no descending
exit is a minimum — one basin by definition, with nothing for lower-completion
to resolve. The plateaux that do have an exit are regions of constant non-zero
gradient, which are rarer in a photograph than F1's phrasing suggests.
`plateau_iterations` therefore defaults to **0**. The pass is off, costs
nothing, and F1 stands open.
**F2 — `cut_to(N)` is a visualisation, not the interaction.** A global cut by region count spends its
budget wherever the saddles happen to be densest: at blur 5 the soft-edged disc's interior held a
cluster of near-equal saddles and ate the budget, fragmenting at a level where everything else was
clean. The real interaction walks up locally from the clicked region and has no such coupling. The
ladder in the example should not be read as what a user would experience.
**F3 — the RAG build still needs a readback.** `Segmentation::read_field` copies labels and gradient
to the CPU, gated behind the `readback` feature exactly as `read_pixels` is. Fine for a spike and
off the frame path, but a shipping build cannot take it (ARCH §6.1, AC-8), so the adjacency
accumulation has to move GPU-side with atomics. That is the largest known gap between this and
something shippable.
**M5 partially answered.** Run-to-run on one device is bit-identical, and the CPU half contributes no
nondeterminism of its own — both asserted by tests. Cross-vendor is untouched and remains the
measurement that can invalidate the region-id representation.
---
## 13. Arms B and C, and what §4 got wrong
Built 2026-08-21 on branch `local-adjustments`. `core/dr-segment`, `core/dr-gpu/src/mask.rs`,
`core/dr-pipeline/src/mask.rs`, and the develop panel.
**The dependency question dissolved rather than being decided.** §4 and D13 both priced arm B as
costing a C dependency under the Android NDK, and treated that as most of the difference between the
arms. It is not a cost that has to be paid: `ort` 2.0's `alternative-backend` feature disables the
linking entirely and lets another engine supply the `OrtApi`, and `ort-tract` — same authors,
MIT/Apache — supplies it from `tract`, which is pure Rust. So arm B runs through `ort`'s API with no
C anywhere, and D13's "largest exception the policy would tolerate" turns out not to be needed.
Measured before committing to it, because tract's operator coverage is the thing that could have
sunk it: **yolo26n-seg loads with zero unsupported operators** and runs 640×640 in ~470 ms on the
reference desktop's CPU. Correct masks on the standard `bus.jpg` — one bus and three people, outlines
following the subjects.
Three findings that contradict §4 directly, and all three change the design rather than the schedule.
**F4 — there is no ADE20K-trained YOLO.** §4's whole argument for arm B was ADE20K's 150 classes and
their *stuff* categories: "'that patch of sky' is a class here." Checked 2026-08-21: Ultralytics ships
YOLO26-seg trained on **COCO**, whose 80 classes are all *things*, and the one HuggingFace repository
claiming a YOLO/ADE20K combination is empty. ADE20K models exist as SegFormer/OneFormer/MaskFormer
transformers, not as YOLO.
So the shipped vocabulary selects **subjects**, not **stuff**. "Select the person" works; "select the
sky" does not come from the model at all and must come from the watershed. That is a narrower arm B
than §4 assumed, and it *raises* the importance of arm C rather than lowering it — the model can no
longer be the whole answer for anything.
**F5 — it is instance segmentation, not semantic segmentation.** §4 assumed a flat partition with
class ids that would "populate `RegionField` directly". YOLO-seg does not partition the image; it
finds objects, and most pixels in a landscape belong to no instance. Two consequences, one bad and
one better than expected: nothing populates a `RegionField` on its own, and two people come back as
*two* instances where a semantic model would have returned one "person" area covering both. For
selecting a subject the latter is the behaviour worth having.
**F6 — tract cannot parse a dynamic-shape export.** It fails shape inference on the neck's `Concat`.
The graph therefore ships with its input fixed at 640×640 square, and every image is letterboxed into
it. This is the constraint behind the tiling option in `semantic.rs`: with a fixed window, tiling is
the *only* route to more semantic resolution, and it costs one inference per tile (≈2.8 s for a 3×2
grid over a 1600 px proxy against 470 ms whole-frame). Defaulted off — a photographic subject is
usually large in frame, which is the case whole-frame inference handles best — and left implemented
so §9's corpus can settle it rather than an argument.
**Arm C ships, and the §8 criteria were not what decided it.** §8 asked for a two-interaction margin
over arm A on the scene subset. That comparison was never run, because F4 and F5 changed what the
arms *are*: with a model that recognises subjects and has no word for sky, arm B alone cannot be a
selection tool at all (§8's "arm B never ships alone" holds, for a stronger reason than expected),
and arm A alone cannot tell a person from the wall behind them. They are complements rather than
candidates. Arm C's implementation is `prior.rs`: instance membership re-weights the merge saddles,
so region pairs the model believes share an object merge early and pairs straddling its edge merge
late. **No boundary moves** — only the order in which boundaries dissolve — which is how the result
stays pixel-accurate at every level while its coarse levels become named things.
**M1–M4 remain unmeasured.** The 24-image corpus of §9 has not been traced, so there are no
interaction counts and no boundary-accuracy numbers. What exists is a working feature and the
evidence that each piece does what it claims in isolation. The corpus is still the thing that would
turn "this feels right" into a number, and it is the largest piece of §11 left undone.
**M5 is unchanged and still the risk it was.** Run-to-run on one device is identical, asserted by a
test. Cross-vendor is untouched. Region ids now reach the sidecar, so if the label field proves
non-deterministic across vendors a mask written on the desktop will not mean the same thing on
Android — see `MaskSource::Regions::signature`, which detects a *retuned* segmentation but not a
differently-rounded one.
**F3 still stands.** `Segmentation::read_field` still copies the label and gradient buffers to the
CPU to build the region graph. It is now behind its own `segment-readback` feature rather than
sharing `readback` — this transfer is once per image on a worker, where the one AC-8 forbids is per
frame in the render loop — but the accumulation still belongs GPU-side with atomics.
---
## 14. Register entries
**S15** — *Region segmentation for local masking* · **CLOSED 2026-08-21**. Arms A, B and C built; the
shared consumers built; the licence question resolved (§7, D14). The 24-image corpus was not traced,
so M1–M4 are unmeasured and M5 is answered only on one device. Answers: local masking snaps to arm C,
and a mask is stored as region ids. Relates to: D13, D14, FR-DEV-3, ARCH §5.4, §6.13.
**D14** — *Segmentation source for local masking* · **DECIDED 2026-08-21**: **arm C**, a watershed
hierarchy re-weighted by YOLO26n-seg instance membership, with the model optional and the watershed
sufficient without it. Weights ship in-tree under AGPL-3.0, which GPLv3 §13 permits and which makes
the combined work effectively AGPL — a deliberate change to DarkRoom's licensing posture, not a
dependency detail (`core/dr-segment/models/LICENCE.md`).
**D13** — *inference runtime* · the dependency half is **answered** for segmentation and the answer
generalises: `ort` + `ort-tract` gives ONNX inference in pure Rust, so the face pipeline of §3.9.1
needs no C dependency either. The *model licensing* half of D13 is untouched — the InsightFace
weights are still non-commercial and still unusable here.
---
## 16. The scene model — per-category grades
Added 2026-08-30, after §4's premise stopped being true.
### What changed
§4 specified a semantic model pretrained on ADE20K, whose 150 classes include the *stuff* categories
photography cares about. §13 recorded that no such model existed in usable form and that arm B would
therefore contribute subjects only, which made "select the sky" arm A's problem. Re-checked
2026-08-30: **Ultralytics now ships a `semantic` task with ADE20K checkpoints**
(`docs.ultralytics.com/tasks/semantic`). `yolo26s-sem-ade20k` is in `models/scene/`.
### It is an addition, not a correction to arm B
The instance model stays exactly where it was, and the reason is the one §13 already gave and was
right about: a semantic model merges every pixel of a class into one region, so it cannot separate
two people, and separating two people is what clicking a subject requires. Swapping arm B for this
would regress the primary interaction to fix a secondary one.
So the two divide by *what the user is doing*, not by which is better:
| | `models/segment/` (COCO instances) | `models/scene/` (ADE20K semantics) |
|---|---|---|
| Question | which pixels are *that* dog | how much of this pixel is sky |
| Granularity | per instance | per category, whole frame |
| Drives | local adjustments, subject selection | the scene tab's per-category sliders |
| Vocabulary | 80 things | 150 classes, stuff included |
### The export is truncated, and both reasons matter
Ultralytics ends the graph with `Resize → ArgMax → Cast`, returning a `[1, 640, 640]` u8 label map.
`tools/export-seg-model.sh` cuts that tail and ships the classifier's `[1, 150, 80, 80]` f32 logits.
**Cost.** The `Resize` materialises 150 × 640 × 640 × f32 — 246 MB — and the `ArgMax` then reduces
across the channel axis, striding 409,600 elements per comparison. Measured under load it was
roughly four fifths of total runtime, spent on work the application discards.
**Softness, which is the more important one.** `ArgMax` destroys the per-class scores, and the whole
design of the scene tab rests on keeping them. Softmax over the 150 channels, summed within each
category, produces per-category weights that sum to one at every pixel — a partition of unity.
Feathering that cannot double-grade a boundary. Feathering *hard labels* outward from two adjacent
categories paints both grades into the overlap, and every horizon in the frame acquires a seam.
### The resolution is 80×80, and no setting changes that
The discarded upsample was never information. `Scene` keeps the native grid and resamples on demand,
so the coarseness is visible in the type rather than hidden. A graduated grade over sky or water is
untroubled by it; a rooftop against sky at 100% zoom will show it. This is the constraint most likely
to decide whether the tab feels good, and it is not addressable by choosing a larger checkpoint —
`yolo26n-sem` and `yolo26s-sem` have the same output grid.
### Licence
Unchanged. Same AGPL-3.0 grant as the instance model, same GPLv3 §13 permission, same consequence
already accepted in D14 — so this needed no new licence decision, which is most of why it was cheap.
See `models/LICENCE.md`.
### Measurement
Timings taken while this was chosen came off a laptop compiling other things and are upper bounds
only. `cargo run -p dr-segment --example scene --release --features embedded-scene-model` reports a
median over N runs with the first excluded; a number worth quoting should come from that, on an idle
machine.
+575
View File
@@ -0,0 +1,575 @@
# Spot removal
**Status:** Draft · 2026-08-26
**Companion to:** [requirements.md](requirements.md) §3.3 FR-DEV-8 · [architecture.md](architecture.md) §5.2
The last develop feature the requirements ask for that nothing in the tree
implements. FR-DEV-8 states the shape — "non-destructive clone and heal spots
stored as parameters in the edit graph (target, radius, feather, source offset,
opacity, mode), with automatic source placement and manual override, plus a
visualise-spots mode" — and this document is how that lands on the pipeline
that exists now.
---
## 1. Why it is worth the work
Sensor dust is unavoidable with interchangeable lenses, and a dust spot is the
most common reason a photographer leaves a RAW editor for a pixel editor
mid-workflow. Every other develop operation in this application can be the best
one in its class and the workflow still breaks at the first frame with a mark on
the sky.
It is also, unusually, a feature whose cost has already been paid twice over.
The neighbourhood stage exists ([`crate::detail`](../../core/dr-pipeline/src/detail.rs)),
the convention for storing geometry in normalised source coordinates exists
([`mask.rs`](../../core/dr-pipeline/src/mask.rs)), the canvas-drag pattern exists
([`gradient.rs`](../../ui/dr-ui/src/gradient.rs)), and the merge-by-id rule exists
([`sidecar.rs`](../../core/dr-pipeline/src/sidecar.rs)). What is genuinely new is
small and is named in §3.
## 2. Non-goals
- **Not layer-based pixel editing.** §1.3 of the requirements excludes that and
this does not reopen it. A spot is a handful of numbers in the edit graph; no
pixels are stored, and the original file is never touched.
- **Not content-aware fill.** The source is a patch from the same photograph,
chosen by an offset. Synthesising texture that is not in the frame is a
different problem with a different budget.
- **Not a general clone brush.** A spot is a disc, not a stroke. A dragged
clone brush is expressible on top of this (a stroke *is* a run of discs) and
is deliberately left until the disc is finished and used.
- **Not automatic dust detection.** Finding spots without being asked is a
reasonable later feature and a bad first one: a false positive silently alters
a photograph, which is the failure this application must not have.
## 3. What is new, precisely
Four things, and it is worth being blunt about them because everything else in
this document is assembly of parts that already work:
1. **A detail pass with variable-length data.** Every [`DetailPass`] today
carries a `Vec<f32>` of uniforms fixed by its own structure. A spot list is
neither fixed nor small. §7.
2. **A neighbourhood operation whose reach is not a small kernel.** Every
existing pass declares a halo of a few pixels. A spot reads from wherever its
source is, which may be a third of the frame away. §5.3.
3. **Undo over something that is not a parameter.** [`History`] snapshots a
[`Preset`], which is a map of scalars — so mask edits are already outside
undo, and spots must not be. §10.4.
4. **A canvas mode that *creates* objects.** Crop edits one rect; local selects
a region; gradient drags an existing shape. Nothing yet makes a new thing
where the pointer went down. §10.
## 4. The model
```rust
/// TRACES: FR-DEV-8
pub struct Spot {
/// Stable across devices; see below.
pub id: String,
/// What is being covered, in normalised **source** coordinates.
pub centre: (f32, f32),
/// What covers it, as an offset from `centre` in **frame units**
/// (y spans 0..1, x spans 0..aspect — the mask convention).
pub offset: (f32, f32),
/// The radius of the disc, in frame units.
pub radius: f32,
/// Fraction of `radius` over which the edge falls away. 0 is hard.
pub feather: f32,
/// How much of the patch is laid down. 1.0 is opaque.
pub opacity: f32,
pub mode: SpotMode, // Heal | Clone
pub enabled: bool,
}
```
**Units follow the mask rule, for the mask reason.** A length stored in pixels
is a length that means something different in the preview and in the export
([`RenderScale`]'s whole documentation is this argument). `centre` is normalised
source, so a crop, a zoom, a pan and a rotation move the spot with the
photograph and no arithmetic is needed to keep it there.
Every *length* — radius, feather, offset — is in the frame's **isotropic
units**, `MaskSource::Radial`'s convention, where y spans `0..1` and x spans
`0..aspect`. Only in those units is a disc a disc: normalised coordinates would
make a spot on a 3:2 frame an ellipse half again wider than it is tall. They are
lengths against the *source* frame rather than the rendered region, so cropping
does not resize a spot already placed — a dust mark is a fact about the sensor,
not about the composition.
One unit for all three, deliberately. A radius in shorter-edge fractions beside
an offset in frame units agrees on a landscape frame and silently disagrees on a
portrait one, which is a bug that stays invisible until somebody rotates a
photograph.
**`offset` is a vector, not a second point.** Dragging the destination moves the
source with it, which is what a photographer expects when they nudge a spot half
a pixel and do not want to re-place the source. Moving the source alone is
editing `offset`.
**The id is derived, not counted.** `MaskStack::next_id` numbers layers, which
is fine for a stack a user names, and wrong here: two devices that each place a
spot offline would both produce `spot3`, and the merge in §11 would treat two
different marks as one. So the id is a short base-36 hash of the centre at
creation, and two devices that place a spot in the same place produce the same
id — which is the correct outcome, because they removed the same piece of dust.
```rust
pub struct SpotSet {
spots: Vec<Spot>, // in creation order; the order matters, see §5.2
}
```
### 4.1 Why it lives beside `ops`, not in it
The [`Operation`] trait takes a `ParamId` and returns an `f32`, and the whole
generic machinery above it — the panel, the sidecar, the presets, the history —
is built on that being true. A spot list is not scalars, and the trait says so
explicitly where it refuses a downcast for film tables.
[`EditGraph`] already holds three things that are not operations for exactly
this reason: `framing`, `masks` and `film`. `spots` is the fourth, and the
argument is the same one `masks` makes — a stack of layers is not a slider, and
folding it into the list would make every consumer that walks `ops` know that
some entries are not really operations.
Bounds, following `mask.rs`'s example of bounding what a sidecar can grow to:
| Constant | Value | Why |
|---|---|---|
| `MAX_SPOTS` | 64 | Beyond a few dozen the answer is to clean the sensor. Refuses rather than dropping, as `MaskStack::push` does. |
| `MAX_SOURCE_DISTANCE` | 0.5 | Frame units. Bounds the halo in §5.3, which is otherwise unbounded. |
| `DEFAULT_RADIUS` | 0.012 | Frame units — about 25 px on a 24 MP frame's short edge, which is a dust mark. |
| `MIN_RADIUS` / `MAX_RADIUS` | 0.001 / 0.5 | Not zero, because a spot that repairs nothing reads as a broken tool; not larger, because the halo bound has to mean something. |
| `DEFAULT_FEATHER` | 0.35 | Fraction of the radius. Soft enough that a heal on a gradient sky has no visible boundary. |
## 5. Where it runs
### 5.1 First in the detail chain
ARCH §5.2 draws spot removal *after* texture and clarity and *before* sharpen
and NR. That diagram is already out of step with the operation set — the
`order:` keys in `core/dr-pipeline/ops/` put noise reduction at 110 and capture
sharpening at 120, ahead of clarity at 130 and texture at 140 — so it needs a
correction anyway, and the correction should put spot removal **first among the
neighbourhood passes**, at a notional order of 105.
The reason is the halo. Sharpening a dust spot before removing it amplifies its
edge, and the amplified edge is wider than the spot: the sharpening kernel has
already smeared a dark ring into pixels that the spot's own disc does not cover,
so the heal leaves a faint circle of over-sharpened background around a patch
that is otherwise perfect. Removing the mark first means every later pass sees a
photograph with no mark in it, which is also the photograph the photographer
thinks they are sharpening.
Since the spot set is not in `ops`, `EditGraph::compose_detail_for` splices its
passes in front of the ops' passes rather than sorting by a declared order. That
is a two-line change and it is stated here so nobody looks for a `spots.yaml`.
### 5.2 Rounds, because sources can read destinations
Every pass reads one texture and writes another. So within a single pass, every
spot reads the *unhealed* image — and a spot whose source overlaps an earlier
spot's destination copies the mark the earlier spot was removing.
The fix is not to run one pass per spot (64 dispatches for a frame that needs
one). It is to group: walking the spots in creation order, a spot joins the
current round unless its source disc intersects the destination disc of a spot
already in that round, in which case it opens a new one. One pass per round, and
the common case — spots scattered over a sky, sources near their own
destinations — is a single round. The grouping is plain CPU code over at most 64
discs and belongs in `SpotSet`, with a test that says an overlapping pair
produces two rounds and a disjoint pair produces one.
### 5.3 The halo, honestly
[`DetailPass::radius`] is "the furthest this pass reads from the pixel it
writes", and it exists so that ARCH §5.3's tile scheduler knows how far to grow
a tile. For a spot pass that is `max(|offset| + radius)` over the pass's spots,
in render pixels — which with `MAX_SOURCE_DISTANCE` at 0.5 can approach half the
frame.
That is a real cost and it should be written down rather than discovered: a
frame with a long-armed spot is close to untileable for that one pass, so the
tiled path will compute it whole-frame. Two things keep it affordable. The pass
is cheap per pixel (§6.4), and it is only the *spot* passes that carry the halo
— the sharpening pass after it still declares its three pixels and still tiles.
The alternative, clamping the source distance to something tile-sized, would
make the tool useless exactly where it is most needed: a mark on a face is
healed from the other cheek, and that is a long way.
## 6. What a spot does to the pixels
Both modes work on the same disc. For a pixel at render coordinate `p` inside a
spot centred at `d` with radius `r`, with the source at `s = d + offset`:
```
w = falloff(|p - d| / r) // 1 at the centre, 0 at the rim
patch = bilinear(source_texture, p - d + s)
c = mix(c, patch + membrane, w * opacity)
```
`falloff` is a smoothstep over the outer `feather` fraction of the radius; a
feather of 0 is a hard disc. `bilinear` is four `tap`s and two lerps, because
the generated preamble offers `textureLoad` only and the offset is fractional in
render space — it becomes a [`Helper`], deduplicated across passes like the
existing luminance helper.
`membrane` is what separates the two modes, and it is zero for `Clone`.
### 6.1 Heal, without a Poisson solve — **implemented**
The classic heal is Poisson blending: copy the *gradients* of the source and
solve for the image whose gradients they are, subject to matching the
destination on the boundary. Solved properly that is an iterative linear system
— tens of Jacobi passes over the disc — and each iteration is a dispatch in this
architecture. Sixty dispatches to remove a dust spot is not a frame budget.
What that solve produces is a smooth membrane interpolating the boundary
difference, and a membrane can be interpolated directly instead of solved. The
shipped form samples the difference between destination and source at `K` points
around the rim and interpolates them into the interior by inverse square
distance:
```
for k in 0..K:
b_k = tap(rim_k) - tap(rim_k + offset) // boundary difference
w_k = 1 / max(|p - rim_k|², 1)
membrane = Σ w_k·b_k / Σ w_k
```
`K = 24` — `RIM_SAMPLES` in `core/dr-pipeline/src/spot.rs`, a uniform rather
than a constant in the source, so tuning it uploads a buffer instead of
recompiling. The cost is `2K` bilinear samples per pixel *inside a disc*, and
nothing at all outside one.
**What was originally specified here was mean-value seamless cloning** (Farbman
et al., 2009), whose weights are half-angle tangents over the rim rather than
inverse squares. The difference matters when the boundary difference varies
sharply around the rim; on the case that actually arises — a repair on a
smoothly varying background — both reduce to the same answer, and the inverse
square form costs two transcendentals per sample fewer. The measurement in
`core/dr-gpu/tests/spot_removal.rs` is what decides whether that trade stays
good: on a linear ramp steep enough to make a clone wrong by 38 levels out of
255, the heal is wrong by **0**. If a case turns up where it is not, the weights
are four lines and the tests are already written.
### 6.2 Why not compute the boundary statistics on the CPU
Because that means reading the rendered image back, and FR-DEV-4 forbids it in
the render path for reasons ARCH §6.1 spends a page on. A per-frame readback to
find out what colour a sky is would reintroduce exactly the stall the whole
architecture exists to avoid. The mean-value form needs no reduction at all,
which is most of why it is the right answer here.
### 6.3 Clone
`membrane = 0`. Kept because heal is wrong on a boundary: a spot straddling a
horizon healed by mean-value blending smears the horizon's contrast into the
disc, and the honest tool then is a straight copy from a matching part of the
frame. This is why FR-DEV-8 asks for both, and it costs one branch in the shader
and one segmented control in the panel.
### 6.4 Cost
Per pixel, per pass: a rejection test per spot in the pass (a squared distance
and a compare), and for the pixels actually inside a disc, `4 + 2K` taps. With
64 spots the rejection cost is the dominant term and it is about 64 × 4 ALU ops
on every pixel of the frame — call it a millisecond at 2 MP on integrated
graphics, which is affordable but not free.
If it proves not to be, the fix is the one `mask.rs` already uses for strokes: a
bounding box per spot and a dispatch sized to it. That needs the detail runner
to dispatch something other than the whole frame, which is a change to its
shape, and it is deliberately not being made until a measurement asks for it.
## 7. Getting a spot list to the GPU
[`DetailPass`] gains one field:
```rust
/// Per-instance data too large or too variable for the uniform block.
pub storage: Option<Vec<[f32; 4]>>,
```
and the generated preamble gains one binding:
```wgsl
@group(0) @binding(3) var<storage, read> instances: array<vec4<f32>>;
```
In `dr-gpu`, both bind group layouts gain a read-only storage entry at binding
3, and a pass that declares no storage binds a shared one-element dummy buffer.
wgpu permits a layout entry the shader does not use, so the two layouts stay two
rather than four, and no existing pass changes at all.
**A spot is two `vec4`s**: `(centre.x, centre.y, radius, feather)` and
`(offset.x, offset.y, opacity, flags)`, all in render pixels except `flags`,
converted on the CPU in `compose_detail_for` where the framing is in scope. This
matters: the shader never sees a normalised coordinate and never has to know
about crop, rotation or zoom — `Framing::output_at` does that map on the way in,
exactly as `gradient.rs` does it for handles. The framing is an affine
similarity, so a disc stays a disc and one radius scales by one factor:
`radius_px = radius × min(source_w, source_h) × scale.ratio()`.
**The source list does not recompile anything.** The WGSL is identical for one
spot and for sixty-four — the count is a uniform and the loop is over the
buffer — so `structure_hash` is unchanged as spots are placed, and placing the
tenth spot re-uploads a 512-byte buffer. This is the same property the fused
pass has for slider movement and it is worth a test that asserts
`cached_pipelines()` does not grow while spots are added.
**Rejected: packing spots into the uniform block.** It would touch no bind group
layout, which is genuinely attractive. It also requires the composer to emit
`vec4` uniform fields (it emits scalars), forces a fixed `MAX_SPOTS`-sized array
and its fixed upload cost into every spot pass, and gives the next operation
that wants a table — a LUT, a curve, a lens grid — nothing to build on. The
storage buffer is a few more lines once and useful again later.
**Invalidation.** The spot set folds into the `detail` key in
`EditGraph::invalidation`, alongside the detail operations. Dragging a spot
therefore re-runs the detail chain and *not* the fused colour pass or the
demosaic, which is exactly the reuse FR-DEV-3d asks for and is the difference
between a spot that follows the finger and one that stutters.
## 8. Automatic source placement
FR-DEV-8 asks for automatic placement with manual override. Two stages, because
the useful half is much cheaper than the good half.
**Stage one — a placed default.** A new spot's source is offset by `2.5 × radius`
in the direction that keeps it furthest inside the frame, biased towards the
frame centre. For dust on a sky, which is the overwhelming majority of spots,
this is right often enough to be worth having, and it is wrong in a way that is
immediately visible and one drag from fixed.
**Stage two — a scored search.** A compute dispatch per new spot scores candidate
offsets on two rings around the destination (say 32 candidates), each scored by
the sum of squared differences over the annulus just outside the destination
disc — the ring is what has to match, since the disc's interior is being
replaced anyway. Penalise candidates whose disc overlaps another spot's
destination, or the frame edge. The winner's offset is read back **once**, when
the spot is created, through `readback.rs` — a few hundred bytes, which is the
size the histogram already moves and completes in well under a frame — and
written into the spot.
Two rules about that readback, both of which are the difference between a
feature and a bug:
- It is **not** in the render loop. `readback.rs` blocks with a deadline, and
that is tolerable exactly once per placement and intolerable per frame. It
happens on the gesture, and the render that follows uses whatever the spot
currently holds.
- The result is **stored**, and the search is never re-run behind the user. A
spot whose source moved on its own when the file was reopened would be an edit
changing itself, and non-destructive editing means the sidecar decides what
the picture is.
If the readback fails, the stage-one default stands. There is no state in which
a spot has no source.
## 9. Visualising spots
FR-DEV-8's "visualise-spots mode" is two different things, and conflating them
is how one of them ends up missing:
**The overlay** — where the spots *are*. Circles for the destination, a fainter
circle for the source, a line between them for the selected spot. Drawn in Slint
over the canvas, alongside the gradient handles and by the same coordinate map,
so it costs the render path nothing and cannot leak into an export.
**The reveal** — where the spots *should be*. Lightroom's "Visualize Spots": a
high-contrast, desaturated view of the frame's high-frequency content, in which
sensor dust on a smooth sky is obvious and in a normal view is nearly invisible.
It is a detail pass appended to the chain:
```
c = abs(c - blur(c)) stretched by a threshold, greyscale, inverted
```
It is a **view**, not an edit. So it is not in the graph and not in the sidecar:
`compose_detail_for` takes a `DetailView` (`Normal` | `RevealSpots`) and the
export path passes `Normal`. A flag on the session would work until the day
somebody exports while the mode is on, and then it would produce a black-and-
white file that looks like corruption. Making the export call site name it is
what stops that from ever being possible.
## 10. Interaction
### 10.1 The mode
A `ViewMode::spots`, and a row in `ToolRail`'s table beside Crop and Local.
That table's own documentation already predicts this shape for the brush; the
spot tool is the same shape and arrives first.
(It landed as a third chip in `ModeStrip`, at the head of the develop column.
The three canvas tools have since moved out to a fixed rail down the left of
the develop view — `ui/dr-ui/ui/toolrail.slint` carries why — and what remains
of that strip is the adjustment-group filters, now `GroupStrip`. Nothing about
the mode itself changed in the move.)
### 10.2 Gestures
| Gesture | Effect |
|---|---|
| Tap / click on the photograph | Place a spot at the current radius, source auto-placed (§8), and select it |
| Drag from a spot's centre | Move the destination; the source follows |
| Drag from the source circle | Change the offset |
| Drag *out* from a fresh placement | Set the source directly, without the auto-placement |
| Scroll / pinch on a selected spot | Radius |
| Tap a spot | Select it; the panel scopes to it |
| `Delete` / `Backspace` | Remove the selected spot |
| Alt-click a spot | Remove it without selecting first |
| `Esc` | Leave the mode |
A drag is a displacement from the press, not a snap to the pointer — the rule
`gradient.rs` states and for the same reason: a finger-sized touch target snapped
to the pointer jumps by half a target the instant it is grabbed.
### 10.3 The panel
While a spot is selected, the adjust column shows radius, feather, opacity and
a Heal/Clone control for *that* spot, exactly as selecting a mask layer
re-scopes the column today. With nothing selected it shows the defaults new
spots will be created with, plus the reveal toggle.
### 10.4 Undo
[`History`] snapshots a [`Preset`], which is a parameter map — so today mask
edits are not undoable, and spot placement must not inherit that. `History`
should hold `(Preset, SpotSet)` and restore both.
That is a narrow change with a wide benefit: the same door lets the mask stack
join later, which closes a gap FR-DEV-5 has open right now. The coalescing rule
needs one addition — a drag of one spot's handle is one step, keyed by the spot
id in the same way a slider drag is keyed by its control — and placement,
deletion and mode changes each open a step of their own.
### 10.5 Touch
Every handle is a `Theme.touch-target`, per FR-UI-3. On a phone the destination
and source circles of a small spot overlap at that size, so the source handle is
drawn at a minimum arm length from the centre while the stored offset is
untouched — the trick `gradient.rs` uses with `MIN_ARM`, for the identical
reason: a handle that cannot be grabbed again is a one-way edit.
## 11. Persistence
One line per spot in the version block:
```
[version 8f04c0e2-…]
exposure.exposure = 0.75
spot.3f9k = 0.4213 0.2871 0.0120 0.35 0.0310 -0.0180 1 heal
```
Fields in order: `centre.x centre.y radius feather offset.x offset.y opacity
mode`. A line rather than a block because a spot is eight numbers and sixty-four
blocks would bury the rest of the file; a line *per spot* rather than one line
for the set because the line is the unit of merge and of a readable diff — the
same reasoning `write_strokes` gives for a line per stroke.
Coordinates are written at the same precision they are held at, as strokes are,
so a round trip is exact and two devices do not generate a diff of noise in the
sixth decimal.
**A malformed line costs that spot and not the file.** A truncated line is
dropped with a warning, exactly as `parse_stroke` drops a bad stroke: a spot
that silently lands somewhere the user never put it is worse than a spot that is
missing, because only one of the two is noticeable.
**Merge** follows `merge_masks` precisely: by id, disjoint survives, a spot both
sides edited resolves wholesale to the higher revision. Half of one device's
offset with the other's radius is a repair neither photographer made. Deletion
propagates through the base comparison exactly as a layer's does.
**Presets do not carry spots** in the first version — a preset is a look, and a
look does not include where the dust was. But dust is in the *same place on
every frame from that body*, which makes "copy spot removal to the selection"
genuinely valuable, and it is a stage of its own (§12, S7) rather than a
surprise inside the existing paste.
## 12. Stages
Each stage is shippable and each has something to look at. Test names are the
files they belong in.
**S1 — The model.** *(Done.)* `spot.rs` in `dr-pipeline`: `Spot`, `SpotMode`, `SpotSet`,
the bounds, id derivation, round grouping (§5.2). No GPU, no UI.
*Tests:* `core/dr-pipeline/tests/spots.rs` — id stability across two identical
placements, `MAX_SPOTS` refuses rather than drops, overlapping sources produce
two rounds, disjoint produce one.
**S2 — Persistence.** *(Done.)* Sidecar write, parse, round trip, merge.
*Acceptance:* a hand-written sidecar with three spots survives a load/save round
trip byte-identically, and two devices that each add a spot offline end with
both.
**S3 — The storage binding.** *(Done.)* `DetailPass::storage`, the preamble's binding 3,
the dummy buffer, the two layouts.
*Acceptance:* every existing detail test still passes untouched, and a synthetic
pass reading the buffer gets what was uploaded.
**S4 — Clone.** *(Done.)* The disc, the feather, the bilinear helper, the pass grouping,
the halo declaration, spliced first into the chain.
*Acceptance:* `core/dr-gpu/tests/spot_removal.rs` — a synthetic frame with a
black disc on a flat grey field is clean to within a tolerance after one clone
spot; the same edit at a one-quarter proxy and at full size land the disc in the
same *normalised* place; adding spots does not grow `cached_pipelines()`.
**S5 — Heal.** The membrane, the mode switch. **Done** — §6.1, and the measurement came out at 0 levels of error against a clone's 38.
*Acceptance:* a dark spot on a linear grey **gradient** — the case clone fails —
is clean to within a tolerance, and the residual at the disc boundary is below
the residual a clone leaves by an order of magnitude. This is the measurement
that decides `K`.
**S6 — The tool.** `ViewMode::spots`, the chip, placement, handles, selection,
the panel scope, deletion, history carrying the spot set, the stage-one source
default. **Done, except the reveal view** — §9's second half is the one piece
of S6 not built, and it is separable: it is a view mode over the detail chain
rather than part of the tool.
*Acceptance:* a dust mark on a real frame is gone in one click, the edit survives
a restart, and undo takes it back.
*What was verified, and how.* Everything below the interface is under test —
the model, the sidecar, the merge, the passes, both blend modes, and undo. The
interface itself was compiled, laid out and photographed: the strip renders
`Crop | Local | Repair` and the column re-scopes. It was **not** driven, because
synthetic clicks do not reach this application (the compositor refuses them),
so the gestures in §10 are as-written rather than as-felt. A first pass with a
real pointer is the outstanding work on this stage.
**S7 — The rest of FR-DEV-8.** The scored source search (§8 stage two), and
copying a spot set across a selection.
S1–S6 is the requirement met in the sense a photographer would recognise; S7 is
the sentence in FR-DEV-8 about automatic placement met in the sense the document
means it.
## 13. Documents to amend
- **requirements.md** — FR-DEV-8 has no *Acceptance:* line; every other
requirement of its weight does. Proposed: *"a dust mark on a smooth sky is
removed in one click with no visible boundary at 1:1, the spot survives a
crop, a rotation and an export at another size, and the exported file matches
the preview."*
- **architecture.md §5.2** — the stage list is out of step with the `order:`
keys in `ops/` and does not show spot removal first among the neighbourhood
passes. §5.1 above is the correction.
- **traceability.md** — regenerated, as ever, rather than edited. FR-DEV-8's row
currently points only at two comments that mention it.
## 14. Open questions
1. ~~**`K = 24`?**~~ Settled by S5's measurement: 24 samples, inverse square weights, zero error on the case the mode exists for.
2. **Does the reveal view belong to spot mode only,** or is it a view mode of
its own that a photographer can turn on while doing something else? It is
cheap to allow both; the risk is a mode nobody remembers turning on. Still
open, and now the only part of §9 unbuilt.
3. **Should a spot be clamped inside the crop?** A spot outside the current crop
costs nothing to render and is invisible, and re-cropping should bring it
back rather than find it deleted. Leaning strongly towards no clamp.
4. **One radius, or an ellipse?** Lightroom's spot tool is circular and its
users cope. An ellipse doubles the handle count for a case a second spot
already covers.
+605
View File
@@ -0,0 +1,605 @@
# Storage backends
How DarkRoom talks to wherever a library lives, and what it takes to add
somewhere new.
This document is the contract. `docs/architecture.md` §8 says why sync is built
on capability negotiation rather than a common denominator; this says what the
seam actually is, where each piece lives, and what a third connector has to do.
---
## 1. What "pluggable" has to mean
A trait alone does not make storage pluggable. `RemoteBackend` existed from the
first release and every layer above it still knew it was talking to Nextcloud:
seven files in `dr-ui` constructed a `NextcloudBackend` directly, ten functions
took one by concrete type, the account model was a server URL beside a DAV user
id, and the local cache directory was named after a hostname. The abstraction
was real and bought nothing, because everything that *reached* a backend was
still shaped like one product.
Pluggable means all four of these, not just the first:
1. **Operations** — what a backend can do. `RemoteBackend`.
2. **Capabilities** — what it can do *cheaply*, so the engine adapts instead of
assuming. `Capabilities`.
3. **Configuration** — what an account is, with no server in it. `Account`.
4. **Registration** — how the application discovers a connector at all, without
naming it. `BackendProvider` + `BackendRegistry`.
Two connectors ship. Nextcloud is unchanged and keeps every one of its
peculiarities — those are the point of the capability model, not an
embarrassment it has to hide. The folder connector serves a plain directory and
exists partly because it is genuinely useful and partly because a second
implementation is the only way to find out whether the first was an
abstraction.
---
## 2. Where each piece lives
```
core/dr-sync/ the contract, and nothing that speaks a protocol
├─ types.rs RemotePath, RemoteId, RemoteEntry, Validator, …
├─ capability.rs Capabilities, ChangeDetection, ServerPreviews
├─ error.rs RemoteError — the one error every caller handles
├─ account.rs Account, AccountStore, Secret, Connection
├─ provider.rs BackendProvider, BackendRegistry, SignIn
├─ lib.rs RemoteBackend, SyncStrategy
├─ scan.rs the walk, driven by capabilities
├─ upload.rs where an original is placed
└─ reachability.rs online/offline, inferred from observed results
core/dr-sync-nextcloud/ WebDAV, oc:fileid, chunked upload v2, Login Flow v2
core/dr-sync-folder/ a directory on a filesystem
ui/dr-ui/src/remote.rs the registry — the ONLY file above dr-sync that
names a connector
```
`dr-sync` depends on no connector. That is deliberate and load-bearing: a build
that only wants a folder library must not compile a TLS stack to get one, and
the registry therefore lives in the crate that already depends on everything —
the interface.
---
## 3. The four traits and types a connector meets
### 3.1 `RemoteBackend` — operations
```rust
#[async_trait]
pub trait RemoteBackend: Send + Sync {
fn capabilities(&self) -> &Capabilities;
fn name(&self) -> &str;
// discovery
async fn list(&self, dir: &RemotePath, since: Option<&Validator>)
-> Result<Vec<RemoteEntry>, RemoteError>;
async fn dir_validator(&self, dir: &RemotePath) -> Result<Validator, RemoteError>;
async fn delta(&self, cursor: &Cursor)
-> Result<(Vec<RemoteChange>, Cursor), RemoteError>;
// transfer
async fn get(&self, id: &RemoteId, range: Option<Range<u64>>)
-> Result<Vec<u8>, RemoteError>;
async fn put(&self, path: &RemotePath, body: Vec<u8>, precond: Option<Precondition>)
-> Result<Validator, RemoteError>;
async fn put_many(&self, items: Vec<(RemotePath, Vec<u8>)>) // defaulted
-> Result<Vec<Result<Validator, RemoteError>>, RemoteError>;
async fn delete(&self, id: &RemoteId, precond: Option<Precondition>)
-> Result<(), RemoteError>;
async fn move_to(&self, from: &RemoteId, to: &RemotePath) -> Result<(), RemoteError>;
async fn create_dir(&self, path: &RemotePath) -> Result<(), RemoteError>;
// optional
async fn thumbnail(&self, id: &RemoteId, size: u32) // defaulted to None
-> Result<Option<Vec<u8>>, RemoteError>;
}
```
Rules that are not obvious from the signatures:
- **`dir_validator` and `delta` are capability-gated.** Return
`RemoteError::Unsupported` unless your `ChangeDetection` is
`PropagatingEtags` or `DeltaCursor` respectively. Answering
`dir_validator` with something that does not actually propagate is worse than
refusing: it lets a caller prune a subtree whose contents changed, and hides
those changes for as long as the folder list holds still.
- **`get` takes an optional range, and it is a hint.** A backend without cheap
ranges may return the whole object; the caller slices. Correctness holds
either way and `Capabilities::range_reads` says whether it was cheap.
- **Chunked upload is not in the trait.** It is an implementation detail of
`put`, chosen by body size. Exposing it would leak one server's protocol.
- **`move_to` must preserve identity where the backend has stable ids.** This
is what a soft delete uses (`FR-CAT-15`): a move implemented as copy + delete
allocates a new id, orphaning the thumbnail shard and turning a restore into a
full re-download.
- **`create_dir` makes parents and succeeds if the directory exists.** Callers
use it to guarantee a destination, not to claim they created one.
### 3.2 `Capabilities` — what is cheap
The engine reads these once at connect time and picks a `SyncStrategy`. See
ARCH §8.1–8.2 for the tiers. The two that change behaviour rather than speed:
| Absent | Consequence the engine handles |
|---|---|
| `range_reads` | Embedded-preview extraction is impossible; browsing falls back to server previews or full download, and is refused on a metered connection |
| `conditional_write` | Sidecar conflict detection falls back to revision counters inside the sidecar — narrows the race, does not close it. Reported as a reduced-safety mode |
**Declare what is true, not what is flattering.** A backend claiming
`PropagatingEtags` it does not have does not merely run slowly; it silently
hides changes.
### 3.3 `Account` — configuration with no server in it
```rust
pub struct Account {
pub backend: String, // BackendProvider::id; defaults to "nextcloud" on load
pub endpoint: String, // stored as "server" — a URL, a path, a bucket
pub login: String, // empty where the connector has no notion of a user
pub user_id: String, // connector-defined sub-address; Nextcloud's DAV segment
pub root: String, // the folder chosen as the library root
pub formats: Vec<String>,
pub last_scan: Option<i64>,
}
```
Everything but `backend` is the connector's to interpret. Code above `dr-sync`
reads these for display and for cache keys, never for meaning.
Two properties are load-bearing:
- **The on-disk form is backwards compatible.** `backend` defaults to
`"nextcloud"` and `endpoint` is stored under its historical key `server`, so
every account written before there was a choice loads unchanged. A config the
app refuses to parse is an account the user has to set up again.
- **`Account::namespace()` is frozen for Nextcloud.** It names the directory
holding the catalog, the thumbnail shards, the sidecar spool and the export
outbox. Changing it does not lose that data, it *abandons* it — silently, as
an upgrade — and costs a full rescan on top. The Nextcloud form is reproduced
byte for byte from what `catalog_path` computed before; every other backend is
prefixed by its connector id, and long endpoints are truncated with a hash
tail so two deep paths cannot collide inside one filesystem's 255-byte
component limit.
### 3.4 `Connection` and `Secret` — the credential split
```rust
pub struct Connection { pub account: Account, pub secret: Option<Secret> }
```
Credentials go to platform secure storage (`FR-NC-2`, `NFR-SEC-2`). Never the
catalog, never the config file, never a log line. `AccountStore` writes the
account as plain JSON and the secret to the keyring, which is what lets the app
show "signed in as duncan, watching /PhotosRaw" before it has touched the
keyring at all.
`Secret`'s inner string is reachable only through `expose()`, and its `Debug`
prints `Secret(***)`. That closes the indirect leak — a `{:?}` on any struct
that happens to hold a connection — by construction rather than by review.
`Connection` is also what replaced a pair of arguments (credentials, user id)
threaded together through fifteen signatures in an order that could be swapped.
### 3.5 `BackendProvider` — registration
```rust
pub trait BackendProvider: Send + Sync {
fn id(&self) -> &'static str; // written to Account::backend
fn display_name(&self) -> &'static str;
fn endpoint_label(&self) -> &'static str; // "Server" / "Folder"
fn endpoint_placeholder(&self) -> &'static str;
fn sign_in(&self) -> SignIn;
fn normalise_endpoint(&self, input: &str) -> Result<String, String>;
fn account_for(&self, endpoint: &str) -> Result<Account, RemoteError>; // defaulted
fn connect(&self, conn: &Connection) -> Result<Box<dyn RemoteBackend>, RemoteError>;
}
pub enum SignIn {
/// A handshake the user completes outside the app, yielding a credential.
Browser,
/// The endpoint is the whole account. No credential, no waiting state.
EndpointOnly,
}
```
- **`id` is on-disk configuration.** Changing it after anyone has an account
orphans that account. Pick it once.
- **`normalise_endpoint` is where a bad endpoint is *rejected*,** before an
account is written for a library that does not exist. Its error string is
shown to the user, so it says what to fix rather than naming a type. The
Nextcloud provider upgrades `http://` to `https://` here (`NFR-SEC-3`); the
folder provider canonicalises the path, so two spellings of one directory do
not become two accounts indexing the same photographs.
- **`connect` is synchronous and cheap.** It validates configuration and builds
a client; it does not talk to the remote. Workers call it per task.
- **`SignIn` is a shape, not a method.** It would be tidier to expose
`async fn sign_in()`, and wrong: Login Flow v2 is a browser handshake the user
completes elsewhere while the app polls, so it is not one call, it does not
finish on our schedule, and the screen has to render a URL and a waiting state
in the middle of it. `SignIn` tells the launch screen which of the two shapes
to draw; the flow stays where its protocol is.
**Credentials are deliberately not abstracted.** An app password, an OAuth
token and a bucket key pair have no useful common shape, and inventing one
before a third backend exists would produce a wrong answer confidently. The
general form is `Connection` — an account plus an opaque secret — and each
connector translates that into what its protocol needs
(`NextcloudProvider::credentials`).
---
## 4. Adding a backend
1. **Implement `RemoteBackend`** over your protocol, in a new
`core/dr-sync-<name>` crate depending on `dr-sync` and nothing else of ours.
2. **Declare `Capabilities` honestly.** Start from `Capabilities::minimal()` and
raise only what you can actually deliver.
3. **Implement `BackendProvider`** beside it.
4. **Register it** in `ui/dr-ui/src/remote.rs::registry()` and add the crate to
`ui/dr-ui/Cargo.toml`.
That is the whole list. Nothing else in `dr-ui` changes, because nothing else in
`dr-ui` names a connector.
**Two things to get right, because they are silent when wrong:**
- **Identity.** `RemoteEntry::id` should be `RemoteId::Stable(u64)` wherever you
can produce a `u64` that names the same photograph on every device looking at
the same library. The catalog keys the thumbnail shards and the face index on
it (`catalog.md` §10.1), and an entry without one gets neither. Set
`Capabilities::stable_ids` only if that id also survives a rename — the two
are different questions and only the second is a capability.
- **Path safety.** A `RemotePath` is built from names on the remote and from a
catalog another device wrote. If you resolve one against a real filesystem,
reject `..` before you open anything.
Register a test double the same way — `BackendRegistry::register` replaces an
existing id rather than shadowing it — so an integration test can stand a fake
server behind `"nextcloud"` without the registry knowing it happened.
---
## 5. The connectors that ship
### 5.1 Nextcloud (`dr-sync-nextcloud`, id `"nextcloud"`)
Unchanged by the abstraction, peculiarities intact — see ARCH §8.4 for the full
mapping. What matters here is that none of them had to be given up to make room
for a second backend:
| | |
|---|---|
| `change_detection` | `PropagatingEtags` — the one-request no-op sync |
| `stable_ids` | yes, `oc:fileid`, survives server-side rename and move |
| `range_reads` | yes, detected by `206` vs `200`, never `HEAD` |
| `chunked_upload` | v2, 5 MB – 5 GB, `MKCOL` → `PUT` chunks → `MOVE .file` |
| `bulk_upload` | yes, `POST /remote.php/dav/bulk` |
| `conditional_write` | yes, `If-Match` |
| `server_previews` | `CommonFormatsOnly` — stock Nextcloud renders no RAW |
| sign-in | `SignIn::Browser`, Login Flow v2, system browser, app password |
Also kept: the `oc:permissions` probe on a refused `PUT`, which is what
distinguishes a create-only share from a bad credential; the `423 Locked`
retry classification; and the bundled ISRG Root YE certificate.
### 5.2 Folder (`dr-sync-folder`, id `"folder"`)
A local disk, an NFS or SMB mount, an external drive, or the directory a
Nextcloud desktop client already syncs. No server, no account, no credential —
which makes it the route that works on a machine with no secrets daemon at all.
| | |
|---|---|
| `change_detection` | `LocalEtags` — see below |
| `stable_ids` | **no** — the id is a path hash and does not survive a rename |
| `range_reads` | yes, `seek` + `take` |
| `chunked_upload` | none; a write is a write |
| `bulk_upload` | no |
| `conditional_write` | yes, with a documented residual race |
| `server_previews` | `None` |
| sign-in | `SignIn::EndpointOnly` |
**Why `LocalEtags` and not `PropagatingEtags`.** A POSIX directory's mtime
changes when its own entry list changes and at no other time — not when a
child's contents are edited, and not for a grandchild. There is nothing to
propagate, so `dir_validator` returns `Unsupported` and the engine walks the
tree every scan. Which costs almost nothing, because the walk that was expensive
was expensive for a reason this backend does not have.
**Measured 2026-08-28**, `cargo run -p dr-sync-folder --example scan`: a full
uncached walk of 2,299 images across 233 directories completed in **137 ms**,
and 380 images across 13 directories in **29 ms** — the same engine, the same
`Depth: 1`-per-directory walk, with no pruning at all. The Nextcloud connector's
comparable figure is 34.1 s for 17,185 RAWs across 334 directories *with*
pruning available (ARCH §8.4). The capability model is what lets one engine
drive both at the speed each actually runs at, instead of forcing the fast one
down to the slow one's interface.
**Identity is a hash of the path relative to the library root**, FNV-1a 64
(written out, because `DefaultHasher` is explicitly unstable between Rust
releases and this value is written into the catalog). It gives the catalog a
`u64` that names a photograph, is the same on every device looking at the same
folder, and does not change when the file is edited. It does not survive a
rename, and `stable_ids: false` says so: a moved photograph is seen as a delete
and an add, and its thumbnail is derived again.
The alternative — keying on the inode — is stable across a rename but *differs
between devices* and is reused by the filesystem after a delete. Two machines
would disagree about which photograph a thumbnail belonged to, and a recycled
inode would silently attach an old thumbnail to a new image. Re-deriving a
thumbnail is a cost; showing the wrong one is a bug.
**Conditional writes.** `IfAbsent` is genuinely atomic (`O_CREAT | O_EXCL`).
`IfMatch` is compare-then-swap: a `stat`, then a write to a temporary beside the
destination and a `rename` over it. A POSIX filesystem has no compare-and-swap,
so the race is narrowed to the microseconds between the two syscalls rather than
closed — still far tighter than the fallback the engine uses for a backend that
declares no conditional write at all, which spans a whole read-modify-write.
The capability is declared, and the residual race is documented at the call
site.
**Two deliberate divergences from WebDAV semantics:**
- **`delete` is not recursive.** A folder library is the user's own photographs
on their own disk with no server-side trash behind it, so a caller that passed
the wrong path would have no way back. Deleting a non-empty directory returns
`RemoteError::Configuration`. Nothing in the engine deletes a directory — the
soft delete is a `move_to` into the trash folder — so the guard is free.
- **Every filesystem call runs on the blocking pool.** On a local disk that is
overkill; on the NFS mount this backend is most useful over, a stalled server
would otherwise wedge the async worker that made the call and every other
request sharing it.
**Failure classification** matters as much as the operations. A vanished mount
(`ESTALE`, `ENOTCONN`, `EIO`) maps to `RemoteError::Network`, which is what puts
the app into offline mode and leaves the catalog readable — exactly as a dead
server does. A permissions problem maps to `PermissionDenied` and does *not*,
because going offline over one forbidden file would hide a fixable problem
behind a network banner. An endpoint that is not a directory at all maps to
`RemoteError::Configuration`: nothing was unreachable and no credential was
wrong, so neither of the other two would send the user anywhere useful.
---
## 6. Virtual filesystems
A sync client in virtual-files mode leaves a **placeholder** where a file is
catalogued but not downloaded. On Linux — the only mode it supports — that
means `IMG.CR2` does not exist at all and `IMG.CR2.nextcloud` does, holding one
byte. ARCH §9.0 measured a real machine: 121,785 placeholders against 10,267
materialised files.
A folder library that ignores this is not merely degraded, it is dangerous.
Before the handling below existed, the folder connector catalogued every stub
as a 1-byte image, gave it an identity that changed the moment it was
downloaded, and — worst — reported a dehydrated *sidecar* as absent, which made
the sidecar writer create a fresh document over an existing one and discard
every edit another device had put there.
### 6.1 Three questions, one trait
Everything else about a synced folder is an ordinary directory, so this is not
a second connector. `dr_sync_folder::Vfs` asks only what differs:
```rust
pub trait Vfs: Send + Sync {
fn name(&self) -> &'static str;
fn is_placeholder(&self, on_disk: &str) -> bool;
fn real_name<'a>(&self, on_disk: &'a str) -> &'a str;
fn placeholder_name(&self, name: &str) -> Cow<'_, str>;
fn can_materialise(&self) -> bool; // defaulted false
fn materialise(&self, local: &Path) -> Result<(), RemoteError>; // defaulted
fn dematerialise(&self, local: &Path) -> Result<(), RemoteError>; // defaulted
}
```
`NoVfs` for a plain directory; `dr_sync_nextcloud::NextcloudVfs` for a synced
one, wrapping the `DesktopClient` socket. A third convention is a third impl.
**Why not a `folder-vfs` provider.** The interesting capability is not a
property of the backend: the same directory can materialise on demand while the
client is running and cannot when it is down, so it must be computed per
connection either way. Registering two providers would ask the user to choose
between two things that differ by whether a background process is up. The
convention is detected instead, per connection, by a hook the registry supplies
(`FolderProvider::with_vfs_detector`) — which is what keeps `dr-sync-folder`
free of any client's protocol.
### 6.2 What the backend reports
| | |
|---|---|
| `RemoteEntry::path` | the photograph's name, never the stub's — so identity survives a download |
| `RemoteEntry::materialised` | `false` on a stub; the catalog maps it to `Availability::Offline` |
| `RemoteEntry::size` | `0` on a stub, meaning *unknown* — see below |
| `get` on a stub | `RemoteError::NotMaterialised`, **never** `NotFound` and never the stub's one byte |
| `put` over a stub, unconditional | **replaces it** — the whole file is being written, so there is nothing in the stub to keep, and the placeholder is removed after the content lands |
| `put` over a stub, `IfMatch` | `NotMaterialised` — a stub's validator describes the placeholder, so nothing here can satisfy the guard; the caller fetches and retries |
| `put` over a stub, `IfAbsent` | `PreconditionFailed` — the file *is* there, only its content is elsewhere |
| `move_to` a stub | moves the stub and keeps it a stub — culling without downloading is ordinary |
| `delete` a stub | deletes it; a photograph is deleted whether or not its bytes are here |
| `capabilities().materialisation` | `OnDemand` with a client, `Placeholders` without, `Always` on a plain folder |
**Size is genuinely unknown.** A Linux suffix-mode stub is one byte and carries
no record of what it stands for. The client's `._sync_*.db` has the real size,
but that is a private schema and reading it would couple us to their migrations.
FR-NC-6c wants a transfer size quoted before an operation starts; for a stub the
honest answer is that it cannot be, and the interface should say so rather than
report one byte or invent an estimate silently.
### 6.3 Hydration is a borrow
The rule: **a file is returned to the state it was found in.** What a pass
downloaded is released; what the user already had is left alone. `BorrowPool`
enforces it.
```rust
let pool = BorrowPool::new();
{
let held = pool.borrow(&backend, &path).await?; // downloads only if absent
// ... read it, thumbnail it, index its faces ...
} // borrow ends
let stats = pool.release_all(&backend).await; // dehydrates only what it hydrated
```
Three properties that are not obvious:
- **Reference counted.** The thumbnail pass and the face pass meet on the same
RAW. Without counting, the first to finish dehydrates the file the second is
reading; with it, the transfer is paid once and released when the last
borrower is done.
- **Prior state is read before asking.** After `materialise` there is no way to
tell what the pass brought from what was already there, so it is recorded
first. Getting this wrong silently undoes a pin, and "my pinned trip
evaporated after an indexing run" is the failure that would make people stop
trusting the feature.
- **Being unsure is not symmetric.** `borrow_known(.., Some(true))` keeps a file
that might have been ours — costing disk. `Some(false)` releases one that
might have been the user's. An uncertain caller passes `true` or `None`,
never a guess at `false`.
A borrow against a plain folder or a server backend short-circuits and does
nothing, so a pass written for a VFS library runs unchanged everywhere rather
than growing two code paths.
**Measured 2026-08-29**, `cargo run -p dr-sync-folder --example vfs_cycle`: a
library of 100 photographs at 25 MB each, 90 of them dehydrated and 10 the user
keeps. A pass over all 100, borrowing and releasing as it goes:
| | |
|---|---|
| on disk at rest | 250 MB |
| **peak during the pass** | **275 MB** — the resting set plus one photograph |
| without borrowing | 2,500 MB |
| on disk afterwards | 250 MB |
| of the 10 the user already had | 10 still there |
The peak is the working set, not the library, and the release is selective.
### 6.4 Derived state is dehydrated too
Shards, the catalog snapshot and the place record live in `.darkroom-derived/`
**inside the library folder**, so a sync client dehydrates them exactly as it
dehydrates a photograph. Unlike a photograph, none of them can be skipped: a
shard that will not open is a peer's thumbnails never merging, and a catalog
snapshot that will not open is their collections.
The folder holds three kinds of thing:
| File | What it is | How two devices reconcile it |
| --- | --- | --- |
| `shard-<client>-NNNN.sqlite` | Thumbnail and face shards | Sealed and immutable; a name match is a content match, so existence is the whole protocol |
| `catalog.sqlite` | The catalog snapshot, for its collections | Read-modify-write: **merge theirs, then push the union** |
| `place.json` | Where the photographer was (FR-UI-8) | Replace: the newer timestamp wins outright |
`place.json` is the odd one out, and deliberately. Everything else there is
*derived* — a faster way to learn something the device could have worked out for
itself from the originals and the sidecars — so losing it costs time. A place is
a fact only the other device knew, and losing it costs a scroll. That is why it
is exchanged last in the pass, why its failures are logged rather than reported,
and why it is the one file here that is replaced rather than merged: two devices
cannot both be where the photographer is, so there is nothing of theirs inside
ours to preserve.
It still refuses to upload over a copy it could not read, for a smaller version
of the reason below: a record we have not compared against might be the newer
one, and overwriting it would move the other device's photographer without ever
having seen where they were.
`derived_sync::read_derived` fetches on demand rather than giving up. More
important is what happens when it *cannot*:
The catalog sync is a read-modify-write over a file another device also writes.
It was shaped `if let Ok(bytes) = backend.get(..)`, which folded every failure
into "there is no remote catalog" and carried straight on to the upload — so a
dehydrated snapshot meant pushing ours over theirs unmerged, taking their
collections and members with it. The same shape as the sidecar bug in §6, and
the same fix: a read that fails for any reason other than `NotFound` **stops the
upload**.
That is why `NotFound` and `NotMaterialised` had to be separate errors. One
means "yours is the whole truth, write it"; the other means "do not dare".
### 6.5 Release means dehydrate, never delete
The single most dangerous thing in this feature. A synced folder is not a
cache: deleting a materialised file inside it propagates the deletion to the
server and removes the photograph from every device the user owns. `Vfs` and
`RemoteBackend::dematerialise` both say so, and an implementation that cannot
dehydrate returns `Unsupported` rather than approximating it.
This is also why the originals cache (`dr_catalog::cache`) cannot simply be
pointed at a VFS library: `Cache::release` deletes bytes, which is right for a
copy under `originals/` and catastrophic in place.
### 6.6 Which photographs stay downloaded
The user's half of the bargain: a pass borrows for a moment, but *some* of the
library should stay local — the trip you are about to take, the shoot you are
working on.
That is a **pin**, and it is the pin the originals cache already had
(`dr_catalog::cache`, FR-NC-6a). Nothing parallel was built, because the model
was already the right one:
| Cache concept | On a placeholder library |
|---|---|
| `tier_desired` | what the user asked to keep hydrated |
| `tier_actual` | what is actually materialised |
| `pending_pins()` | the work list — what to hydrate next, resumable |
| pinned rows are never evicted | a pinned collection is never dehydrated |
| passive rows, LRU under a budget | what a pass borrowed, released when it finishes |
So "keep this collection hydrated" is `Cache::pin`, and the existing pin worker
drives it — except that on a placeholder library it calls `materialise` instead
of downloading a copy.
**Why not a copy.** The original materialises *in the library folder*. Copying
it under `originals/` as well would hold every pinned photograph twice, and the
copy would be the half the budget could evict while the real disk cost stayed.
`Cache::record_in_place` records the bookkeeping with **`path = NULL`**, and
that null is load-bearing: `release` deletes the file a row names, and a row
that names none deletes nothing. The safety property is structural rather than
remembered.
Unpinning therefore frees nothing by itself — the bytes are not ours to delete.
`spawn_dehydrate` asks the client to take them back, which is what actually
returns the disk.
---
## 7. What the abstraction does not yet cover
Stated so the next person does not have to rediscover it.
- **Multiple accounts at once.** `AccountStore` holds a list and the launch
screen uses the most recent. Nothing in the model prevents two open libraries;
the interface has no place to show them.
- **Per-backend settings.** A connector has no way to contribute a settings
page. Anything configurable is on the `Account` or is not configurable.
- **Capability probing at runtime.** `Capabilities` is fixed at construction.
Nextcloud's `server_previews` should really be probed per account — a server
with `camerarawpreviews` installed can render RAW — and today it is assumed to
be `CommonFormatsOnly`.
- **A general notion of an account.** Credentials stay connector-specific on
purpose (§3.5). A third connector with an OAuth flow will need a third `SignIn`
variant, and that is the right place for it to appear.
- **A quoted cost before a hydrating pass.** FR-NC-6c wants the transfer size
stated before an operation that needs absent data. A placeholder reports no
size (§6.2), so the honest figure for "index this library" is a count and not
a byte total. The interface should say *n photographs, size unknown until
fetched* rather than estimate one silently — and it does not say anything yet.
- **Metadata-only placeholders.** Windows and macOS express these in filesystem
metadata rather than in the name, and carry the real size there. `Vfs` asks
its questions about a *name*, which is all the one convention this project has
met needs. Supporting them means widening the trait to take a `Metadata`, and
doing that before anyone has run this on those platforms would be guessing.
- **Hydration during browsing, deliberately.** It stays forbidden (ARCH §9.0
finding 3). A grid cell whose content is absent shows as not-downloaded; only
a pass the user asked for may fetch.
+451
View File
@@ -0,0 +1,451 @@
# DarkRoom — Technical debt
**Status:** Living document · first written 2026-08-26
**Companion to:** [architecture.md](architecture.md)
Deliberate compromises: things the code does knowing they are wrong, because the alternative was
worse at the time. Each entry says what the debt is, what it cost to take on, what it would take to
pay off, and how you would know it had been paid.
Not a bug list. A bug is something nobody chose. Everything here was chosen, and the point of
writing it down is that the reasoning outlives whoever chose it — so the next person can tell a
constraint from an accident, and does not "fix" something load-bearing or preserve something that
has quietly stopped being necessary.
---
## TD-1 — The Android develop view reads pixels back through the CPU
**Breaks:** [architecture.md §12 / 6.1](architecture.md) — GPU results never round-trip through the
CPU — and AC-8, on Android only. Desktop is unaffected and keeps the zero-copy path.
### What it does
`DevelopSession::render` on Android runs the compute passes on the GPU as usual, then calls
`AdjustPass::export_pixels` and hands the frame to Slint as a `SharedPixelBuffer`. That is exactly
the GPU→CPU→GPU transfer §6.1 exists to forbid, and it is on the frame path.
### Why
Zero-copy needs Slint to draw with wgpu. On Android that means wgpu's Vulkan swapchain, which
hardcodes `preTransform = VK_SURFACE_TRANSFORM_IDENTITY_BIT_KHR`
([gfx-rs/wgpu#3345](https://github.com/gfx-rs/wgpu/issues/3345)) — wgpu-hal says so in a comment
beside the line.
On a tablet whose panel is mounted landscape, a portrait window then hands Android an unrotated
buffer, every present returns `VK_SUBOPTIMAL_KHR`, and frames arrive torn. Measured on the device,
same build, only the tablet rotated:
| orientation | `bufferTransform` | composition | result |
|---|---|---|---|
| landscape | `ROT_180` | `DEVICE (2)` | clean |
| portrait | `ROT_270` | `CLIENT (1)` | torn |
Setting `preTransform` is not a fix available to us: the field is a *promise* that the content is
already rotated, so honouring it needs the renderer to rotate what it draws, which wgpu cannot do
on Skia's behalf.
So the choice was never fast-develop against slow-develop. It was a develop view that costs a
readback against a grid that tears in the orientation a tablet is mostly held in.
### What it costs
Less than §6.1's headline numbers, because `render` fits the pass to the canvas before it runs — the
readback is at viewport resolution, not sensor resolution. The 7.43 ms at 4K in §12 is the ceiling,
not the bill. **It has not been measured on the device**, which is the first thing to do if the
develop view feels heavy on the tablet; do not assume this is the cause without a number.
### And a second transfer, while focus peaking is on
Added 2026-08-29 with FR-CULL-3. The focus-peaking overlay is a compute pass writing its own
`Rgba8Unorm` texture, which on desktop reaches the compositor with no copy — but on Android there is
no more a path for *that* texture than for the frame it belongs to, and an overlay that stayed on
the device while the picture underneath it did not would simply never be seen. So
`FocusPeakPass::read_overlay` follows the frame back through memory, and the Android frame path
carries **two** full-resolution `copy_texture_to_buffer` transfers instead of one.
This is recorded under TD-1 rather than as its own entry because it is not an independent choice.
It exists only because TD-1 exists, it is bounded by the same thing — `render` fits the pass to the
canvas, so both transfers are at viewport resolution — and TD-1's "Done when" already covers it:
whichever of the three fixes above lands removes the readback for the frame and the overlay
together, because both are the same missing capability.
Two things worth saying plainly. The doubling is **reasoned, not measured on the device** — the same
gap TD-1 admits about its own cost, and the reason neither number should be quoted as a measurement.
And it is paid only while the photographer has the overlay switched on: `DevelopSession::focus_overlay`
returns on its first line when peaking is off, so with it off there is no dispatch and no transfer,
and the Android frame path is exactly what it was before this feature existed.
### Paying it off
Any one of these removes it:
- wgpu implements pre-rotation (#3345), and Android goes back on `unstable-wgpu-29`.
- Slint's Skia Vulkan surface handles `preTransform` and Android uses that instead of OpenGL.
- Skia over OpenGL grows a way to sample an external texture that wgpu can write.
**Done when:** `ui/dr-ui/Cargo.toml` no longer scopes `renderer-femtovg-wgpu` and
`unstable-wgpu-29` to non-Android, the `#[cfg(target_os = "android")]` arm of
`DevelopSession::render` is gone, and the tablet is clean in portrait.
---
## TD-2 — Thumbnails are fetched one at a time
**Where:** `library::spawn_thumbnails` — the `for req in to_fetch` loop.
### What it does
The interactive thumbnail batch fetches serially: one image at a time, and two HTTP round trips
each (a header read, then the preview's byte range). A window of a few hundred cells is that many
sequential round trips against the server.
### Why it is debt rather than a bug
It is correct, and it was fast enough when a window was one screenful. It is the *ordering* that
kept it survivable: since `fetch_rank`, on-screen cells are requested first, so the cells a person
is looking at arrive first even though the queue as a whole is slow.
Portrait makes it worse by construction — a narrow window means smaller cells, more rows, and two
to three times as many cells on screen at once, all of them ahead of the ones below in a queue that
never runs more than one request.
### Paying it off
`spawn_thumbnail_sweep` already has the pattern: `SWEEP_LANES` disjoint lanes over a chunk, joined,
with the store written on the one thread that owns it. Striping a *priority-ordered* chunk across
lanes keeps `fetch_rank`'s ordering while running several requests at once.
Not done yet because it multiplies concurrent requests against the user's Nextcloud during a
scroll, and that is a behaviour change worth deciding on deliberately rather than inheriting from a
performance fix.
**Done when:** the interactive batch runs on more than one lane, priority order is preserved
across the lanes, and a slow server still cannot stall the visible cells behind offscreen ones.
---
## TD-3 — The thumbnail drain applies an unbounded batch on the UI thread
**Where:** `library_ui::drain_thumbnails` — the `loop` inside the timer callback.
### What it does
Every message queued when the timer fires is applied in that one callback, with no ceiling. On a
library whose thumbnails are already in the store, the worker delivers a whole window at once, so a
single callback can do hundreds of `to_slint_image` calls back to back — each an allocation and a
full RGBA copy — while the grid is mid-flick.
The copy cannot move off the UI thread: `slint::SharedPixelBuffer` is not `Send`, so decoded bytes
can only become an `Image` on the thread that draws. Only the *amount done per wake* is ours to
choose, and right now it is "all of it".
### Cost
Measured with a temporary probe, **debug build**, so treat the shape rather than the size:
| class | per thumbnail | × a 280-cell window |
|---|---|---|
| grid, 256 px | 1.93 ms | 539 ms |
| large, 512 px | 7.78 ms | 2.18 s |
A release measurement was started and never completed — do not quote these as release figures.
### Paying it off
A time budget per wake and a shorter interval: apply for a few milliseconds, return without
stopping the timer, and finish on the next tick. A batch then lands in frame-sized slices rather
than one lump between two frames. Draft written and discarded during the investigation; it is a
small change.
**Done when:** one wake of the drain cannot exceed a frame, and a fully-cached window still fills
in well under a second.
---
## TD-4 — The local-contrast base is computed at full render resolution ✅ PAID OFF
**Where:** `dr_pipeline::ops::local_contrast::LocalContrast::passes` — the `base` and `combine`
passes, and the stage that dispatches them, `dr_pipeline::detail`.
Breaks **FR-DSP-3** at large viewports. Measured, and the numbers are in
[frame-budget.md](frame-budget.md) §M3.
### What it does
Clarity's Gaussian σ is 1.2% of the frame's shorter edge, truncated at 2σ, so its kernel radius is
a property of the *viewport*: 29 px at 1920 × 1200, 38 px at 2560 × 1600, **52 px at 4K**. The two
separable passes therefore run 105 taps each over 8.3 M pixels at 4K, which is 1.7 billion texture
reads for one control.
| viewport | radius | clarity alone, p99 |
|---|---:|---:|
| 1920 × 1200 | 29 | 5.99 ms |
| 2560 × 1600 | 38 | 12.44 ms |
| 3840 × 2160 | 52 | **33.89 ms** |
RTX 3050 laptop, `examples/frame_budget`, fused dispatch reused so this is the convolutions alone.
Clarity is 97% of the cost of all four neighbourhood operations together at every size.
For scale: the entire fused chain — every point operation active, film stock included — costs
4.5 ms at the same 4K viewport. **A single slider is seven times the rest of the pipeline.**
### Why
Because the stage cannot do otherwise yet. `dr_pipeline::detail` dispatches every pass at the
render size; there is no way to express "read this target and write a smaller one". The module's own
documentation has said so since it was written:
> The right optimisation is a base computed at reduced resolution, which needs a detail stage that
> can write a smaller target than it reads; that is a change to `crate::detail`, not to this file.
It was the right call to ship the correct answer slowly rather than a fast approximation nobody had
checked — the halo behaviour is the hard part of this operation and it is tested.
### Not a tiling problem
Worth saying because ARCH §5.3 offers a tile cache and this is the stage that looks like it wants
one. It does not: a tiled convolution reads a halo per tile, so at a 52-pixel radius, 256-pixel
tiles would read (256 + 104)² instead of 256² — very nearly **twice** the taps.
[display-and-extension.md](display-and-extension.md) §2's decision rule was resolved on this
evidence; see [frame-budget.md](frame-budget.md).
### Paying it off
A detail pass that declares an output scale, so the base can be computed at a quarter resolution and
sampled back up in `combine`. A quarter-resolution base is 1/16 the pixels at 1/4 the radius —
about **1/64 of the work** — and is visually identical, because a base at σ = 26 px holds no content
above the quarter-resolution Nyquist to lose. Texture's σ is a decade finer and must stay at full
resolution; the scale therefore belongs on the `DetailPass`, not on the stage.
**Done when:** clarity at 100% is inside the frame budget at 3840 × 2160, the halo tests in
`tests/local_contrast.rs` still pass unchanged, and `examples/frame_budget`'s M3 table in
[frame-budget.md](frame-budget.md) has been rerun and committed.
### Paid off
A `DetailPass` now declares `output_scale`, and clarity's base is computed on a grid a quarter the
size on each axis. Measured before and after on the same machine, same adapter, same build profile,
with only the change between them — see [frame-budget.md](frame-budget.md) §"The reduced base,
measured":
| viewport, fit | before p50 | after p50 | | before p99 | after p99 |
|---|---:|---:|---:|---:|---:|
| 1920 × 1200 | 2.30 ms | 1.56 ms | 1.5× | 3.94 ms | 1.97 ms |
| 2560 × 1600 | 4.41 ms | 1.95 ms | 2.3× | 10.40 ms | 2.37 ms |
| 3840 × 2160 | **10.94 ms** | **3.88 ms** | **2.8×** | 25.05 ms | 4.17 ms |
**Quote the p50 column.** The baseline run's p99 figures are contaminated — its `fit` rows spread
2.3× between median and 99th percentile where the after run spreads 1.1×, and
[frame-budget.md](frame-budget.md)'s own independent measurement of the same baseline on the same
card reports 4.67 ms p99 at 2560 × 1600 against the 10.40 ms here. The p99 improvement is real and
larger than 2.8×; this run cannot say by how much.
Clarity is no longer the stage that misses the budget, and no longer dominates the neighbourhood
stage: at 4K it is 4.17 ms against 4.61 ms for all four neighbourhood operations together, where it
was 97% of that total at every size. That comparison is within one run, so the contention does not
touch it.
**Two things worth recording, because neither is visible in the table.**
The declared halo is now quantised to multiples of `output_scale`. The kernel truncates at 2σ and
that rounding now happens on the reduced grid, so 1920 × 1200 reports 28 render pixels where it
reported 29, and 2560 × 1600 reports 40 where it reported 38. At 2σ the Gaussian is already down to
`e⁻²` of its peak, and the cross-form test holds the difference to 0.03 stops of peak excursion and
2% of frame reach — but it is a real change in reach, not a pure speed-up, and a tile scheduler
would see it.
The measurement was taken on an AMD RX 5700 XT, not the RTX 3050 the M1/M2/M3 tables above were
measured on, so the *absolute* figures are not comparable with those. The before/after is, because
both halves of it were measured on the same card minutes apart.
---
## TD-5 — The fused shader is reassembled from strings on every frame
**Where:** `dr_pipeline::operation::compose_full`, called from `DevelopSession::render`.
### What it does
`EditGraph::compose` walks the active operations and formats a WGSL source string, per frame, on
the UI thread. On a full chain that is **2.8–5.2 ms** — at 1920 × 1200 it is larger than the entire
fused dispatch it precedes, and on the chains that also carry a detail stage it is a third of what
is left of the 16 ms budget after the GPU has taken its share. It does not vary with resolution,
because it is not pixel work.
### Why
Because it was free until the chain got long. Composition was written when an edit was two or three
operations, and the cost is roughly linear in generated source: the colour mixer emits twelve hue bands
and the tone curve emits a spline evaluator, so a full chain is a large string built from scratch
sixty times a second.
### Paying it off
The generated source depends only on the *structure* of the graph — which is precisely what
`ComposedShader::structure_hash` already identifies, and precisely what does not change while a
slider is being dragged. `AdjustPass` relies on that already: it caches compiled pipelines against
that hash and does not recompile during a drag. Caching the source string against the same hash and
rebuilding only the uniforms — a handful of floats per operation — takes this to approximately
nothing on the path that needs it most.
The care needed is in what the hash covers. It deliberately excludes parameter *magnitudes*, so a
cache keyed on it is sound for the source and would be wrong for anything else in `ComposedShader`.
**Done when:** the `shader` column of [frame-budget.md](frame-budget.md)'s M1 table is under a
millisecond for the `point` and `all` chains, and the codegen tests still pass byte for byte.
---
## TD-6 — The quietest ink does not reach WCAG AA, and the rule does not reach 3:1
**Breaks:** [requirements.md](requirements.md) NFR-A11Y-2 — "non-canvas UI meets WCAG AA contrast".
### What it does
`style.yaml` sets three inks and four surfaces. Measured as WCAG 2 contrast ratios (sRGB relative
luminance, the standard formula), against the surfaces each ink is actually drawn on:
| ink | on `ground` | on `surface` | on `surface-raised` | on `hover` | on `selected` |
|---|---|---|---|---|---|
| `ink` #EDEEF0 | 16.02 | 14.69 | 13.03 | 11.39 | 9.68 |
| `ink-dim` #9EA1A6 | 7.18 | 6.58 | 5.84 | 5.10 | **4.34** |
| `ink-faint` #71747A | **3.97** | **3.64** | **3.23** | **2.82** | **2.40** |
| `warn-ink` #C9A05A | 7.67 | 7.03 | 6.24 | 5.45 | 4.64 |
| `rule` #323438 | **1.49** | **1.37** | **1.21** | **1.06** | **1.11** |
Every text size in the application is 11px, 13px, 17px or 24px, and WCAG's "large text" relief
begins at 18.66px bold or 24px regular — so all four of those thresholds are the normal-text one,
**4.5:1**, except the masthead. Bold entries fail it.
The inverted cases pass and are worth stating so nobody re-measures them: `ground` on `active`
(#FFFFFF) is 18.60, on `active-dim` 11.20, on `active-pressed` 5.95, on `selected-ring` 13.02. The
near-white fills that `Button.primary`, `FilterChip.active` and the held tool-rail entry use are
the *best*-contrasting text in the interface, not the worst.
So the failures are exactly two, and neither is where one would guess:
- **`ink-faint` reaches 4.5:1 nowhere at all.** It is the ink for `Caption`, `PanelHeading`,
`Disclosure`, `Value`'s placeholder state and `FilterChip`'s count — every hint, every section
name, every "3 photographs" under a title.
- **`rule` reaches 3:1 nowhere.** WCAG 1.4.11 asks 3:1 of the boundary of a control the user must
perceive, and `rule` is the border of every `Button`, `Field`, `Panel`, `ChoiceChip` and
`IconButton`. An unfilled secondary button is a 1.4:1 outline on a 1.2:1 background.
`ink-dim` on `selected` at 4.34 is a third case, marginal enough that a two-point lift fixes it.
### Why
Not an oversight — the direct consequence of the palette's own argument, which `style.yaml`'s
preamble makes at length and correctly. The chrome is deliberately quiet because a bright surround
biases how a photograph is judged, and hue is banned outright because an accent beside the image
shifts the perception of nearby colours. What is left to signal with is luminance, and the palette
spends its luminance range on the *photograph*, keeping the chrome inside a narrow band above the
ground.
A narrow band is precisely what a contrast ratio measures. `ink-faint` exists to be skipped by the
reader who did not stop to look; that is a real design intent, and "text you are meant to skip"
and "text everyone can read" are in genuine tension rather than one being a mistake.
### What it costs
The photographer who cannot read a caption cannot read *any* caption, on any screen — this is one
token, so it fails everywhere at once. The hints under the settings switches say what a setting
costs, the section names say what a panel is, and the counts say how big a filter is. None of it is
decorative.
### Paying it off
Two token changes, and the second is the awkward one.
`ink-faint` needs roughly #8A8D93 to clear 4.5:1 against `surface-raised`, the darkest surface it
is drawn on that matters — which puts it about where `ink-dim` sits today and collapses the
three-ink scale to two. So the real fix is to re-derive all three inks against the surfaces rather
than to nudge one: the scale wants to start higher and keep its steps, not compress.
`rule` needs about #4A4D52 for 3:1 against `surface`. That is a visibly stronger line, and the
preamble's "instrument rather than absence" reasoning applies to it as much as to the greys — this
is a look change, not a number change, and it should be looked at rather than computed.
Both are decisions about how the application appears next to a photograph, which is the one thing
this palette was designed around. They want a screenshot and an opinion, not a patch.
**Done when:** every `Theme` ink reaches 4.5:1 against every surface it is drawn on, `rule` reaches
3:1 against `surface` and `surface-raised`, and a test recomputes those ratios from `style.yaml` so
the next palette edit cannot quietly undo it. The table above is the baseline to compare against.
### Not in scope
The histogram's `plot-*` inks (2.36 for `plot-luma` on `ground`) are drawn *on* the canvas, and
NFR-A11Y-2 scopes contrast to non-canvas UI. NFR-A11Y-3 covers what those need instead, and is
already met — the readouts name the channel in words.
---
## TD-7 — Platform font scaling is not honoured
**Breaks:** [requirements.md](requirements.md) NFR-A11Y-2 — "platform font scaling is honoured
without clipping".
### What it does
Nothing at all, which is the entry. Every type size is a constant in `style.yaml` — 11, 13, 17, 24
— read as `Theme.text-sm` and friends at 65 call sites, and there is no multiplier anywhere between
the platform's font-size preference and those numbers. `scale_factor()` is read in `display_ui.rs`
and in `lib.rs`, but only to size the canvas in physical pixels for the render; it is display DPI,
which Slint already applies to logical lengths, and it is not the user's text-size setting. A
photographer who sets 130% text on GNOME or Android gets an application that ignores it.
### Why
Because honouring it is not a multiplier, and pretending it is would be worse than not doing it.
The layout is built on constants that are not derived from the type size: `control-height` 28,
`touch-target` 44, `row-height` 26, `rail-entry-height` 54, `panel-width` 360, and a dozen fixed
heights written at their call sites — `ParamSlider`'s 46px, the folder picker's 220px box, the
readout column's 30px. Scaling the type alone clips against every one of them, silently, because
Slint elides rather than errors. `SwatchSlider` is the sharpest case: 12 hue bands × 3 channels in
a 360px column, sized so that a track and a swatch and a three-character readout fit on one line.
And the failure is invisible to the person shipping it. The requirement's own phrase is "without
clipping", and clipping is exactly what a screenshot at 100% cannot show — the memory note on
verifying Slint changes exists because these files have a history of compiling, rendering and being
wrong.
### What it costs
The user for whom this matters most is not the screen-reader user the rest of this branch serves —
it is the one with usable but poor sight, who reads the interface and needs it larger. The
application is unusable to them at any setting, and there is no partial credit: text scaling is a
system-wide preference, so an app that ignores it is the one thing on the desktop that did.
### Paying it off
In the order the pieces depend on each other:
1. **A scale token.** `Theme.text-sm` and the rest become `base × Theme.type-scale`, with the
scale an `in-out` property Rust writes at startup from the platform. `build.rs` already emits
`in-out` tokens under `live-style`, so the codegen half of this exists and is proven — that
feature is the mechanism, one line from being general.
2. **A source for the number.** GNOME publishes `text-scaling-factor` over the settings portal;
Android has `Configuration.fontScale` through JNI, beside the calls `lib.rs` already makes for
`ACTION_VIEW`. Both want a default of 1.0 and a sane clamp — 0.8 to 2.0 — because a user who has
set 300% for a phone launcher has not asked for a 72px slider readout.
3. **The constants that are not type.** Every fixed height a *label* sits inside has to follow the
scale; every touch target must not shrink and need not grow. That is the work, and it is where
the 46px and 220px literals get read one at a time.
4. **Evidence.** Screenshots at 1.0, 1.3 and 2.0 of the develop column, the settings page and the
colour mixer — the three densest layouts — because "without clipping" is a claim about the
worst case and nothing else will show it.
**Done when:** the develop column, the settings page and the colour mixer render at a 2.0 scale
with no elided label and no touch target under 44 logical pixels.
---
## Related, and deliberately not here
The window-move rule, the grid's ordering index and the whole-library readout cache were *fixed*
rather than deferred — see the commits around `6d6ef8d`. They are mentioned only so that a reader
looking for "why was the grid slow" finds the answer in the code and its comments rather than
assuming it is still outstanding.
File diff suppressed because one or more lines are too long
+827
View File
@@ -0,0 +1,827 @@
# UI navigation: finding things once there are many
TRACES: FR-UI-1 | FR-UI-3 | FR-UI-5 | FR-DEV-3a | FR-DEV-3c
Successor to [`ui-refinement.md`](archive/ui-refinement.md), which asked how the interface should *look*.
This asks how someone finds anything in it. The two are sequenced together at
the end.
## Why now
Local adjustments landed (D14, `segmentation.md` §14) and the develop column
went from four panels to six: image, histogram, geometry, settings, local,
adjust. Adjust alone is ten operations, and the colour mixer contributes
thirty-six parameters by itself. That is already past what one scrolling
column presents well, and the operation set is meant to keep growing —
FR-DEV-3 lists texture, clarity, sharpening and noise reduction as v1, none of
which exist yet.
But the count is the lesser problem. **Local adjustments introduced a mode
without introducing a way to see it**, and that is the part that can lose
someone's work rather than merely slow them down.
---
## 1. The three problems, which are not one problem
### 1.1 Scope — invisible state
Selecting a mask layer silently re-points the adjust panel at that layer's
chain. Same thirty sliders, different meaning, and the only indication is a
caption between the two panels.
Three ways that bites:
- An exposure change lands on the whole photograph when it was meant for a
face, or the reverse. Both are silent; both are discovered later.
- **The histogram does not follow scope.** The instrument the tonal controls
are judged against reports the whole frame while the slider edits a
subject's face. FR-DSP-7 asks the histogram to describe what the
photographer is looking at; under a mask it currently does not.
- Undo interleaves global and local edits with nothing distinguishing them.
This is the classic modal fault, and the classic remedy applies: make the mode
visible, or make it not a mode. See §3.
### 1.2 Extent — six panels in 280px
The column scrolls as one (`ui-refinement.md` Workstream C explains why: a
scroller inside a scroller gives every drag a third thing to be lost to). At
six panels that scroll is long enough that the histogram — the instrument
everything tonal is judged against — is frequently off screen while the
sliders it reports on are being dragged.
### 1.3 View — grid and develop are still separate screens
Diagnosed as `ui-refinement.md` Workstream F and not yet built. Unchanged by
this document, which assumes F lands.
---
## 2. What other programs do
Worth summarising honestly, because all three problems are solved elsewhere
and the solutions have known costs.
### On a desktop, two families
**A collapsible stack.** Lightroom Classic's right panel is a vertical column
of modules — Basic, Tone Curve, HSL, Detail, Lens, Effects — each collapsible,
each reporting whether anything inside it has been touched. Everything is in
one place and in a fixed order, so muscle memory works; the cost is a long
scroll and a lot of triangles.
**Tool tabs.** Capture One puts an icon strip at the top of the tool panel and
gives each tab a curated set of tools; RawTherapee tabs the right-hand panel
the same way. Scroll is bounded and there is a clear sense of place; the cost
is that a control you cannot name is in one of eight places, and switching tabs
loses the context you were comparing against.
**Groups over a stack.** darktable combines both — a row of group icons
filtering a long list of collapsible modules, plus a search box. It is the most
powerful and the most often described as overwhelming, which is worth reading
as a warning about combining mechanisms rather than about either one.
Across all of them, **masking is its own tool**, not a panel among peers.
Lightroom opens a masking tool with its own layer list and its own canvas
overlay; Capture One makes layers a persistent selector at the top of the
adjustments tab. Nobody makes a mask a panel that silently rewires a different
panel — which is what DarkRoom currently does.
### On a phone, one family — and why it does not apply here
Lightroom Mobile, Photomator and VSCO all converge on the same shape: **a
horizontal strip of tool icons along the bottom**, and tapping one replaces a
bottom sheet with that tool's controls. Snapseed goes further — one tool fills
the screen and a vertical swipe chooses the parameter while a horizontal one
sets it.
The convergence is not fashion. Three physical facts drive it:
- **Thumb reach.** A one-handed grip reaches the bottom third. A right-hand
column is a mouse idiom.
- **The image needs the screen.** A 280px column is a fifth of a desktop
window and most of a phone in portrait.
- **There is no hover.** Disclosure triangles and hover-revealed affordances
are worth less; a control is either visible or gone.
**None of the first two apply to DarkRoom's targets**, which are a 12-inch
tablet and a desktop (§3, D-N2). Nobody thumbs a 12-inch tablet one-handed,
and its narrow dimension is not narrow. The third does apply, and is handled
already — see D-N2.
---
## 3. The decisions
### D-N1 — Local adjustment becomes a mode · **DECIDED**
Not a panel that re-points another panel. A mode, in the sense `crop-mode`
already is: it changes what the canvas does, scopes what the column shows, and
is left explicitly.
**Why this shape rather than louder signalling.** The app already has this
pattern and the user already knows it. Crop mode arms a canvas interaction,
draws an overlay, gives the column one job, and exits by the same control that
entered it. Local masking is the same animal — a canvas interaction plus a
scoped panel — and building it as a peer panel is what created §1.1. Making it
a mode removes the ambiguity by construction instead of describing it in a
caption.
It also inherits machinery that exists. `lib.rs` already resolves Escape and
the Android back gesture to "leave the innermost state first" (FR-UI-5); local
mode joins that stack and needs no new exit concept.
In local mode:
- the canvas turns the overlay on and arms click-to-select;
- the column shows the mask stack and, beneath it, the adjustments **scoped to
the selected layer**;
- the header names the scope — the layer, not "adjust";
- leaving returns to the whole photograph, by Escape, by back, or by the mode
control.
**The histogram follows the scope.** Under a mask it reduces over the masked
pixels only. This is FR-DSP-7 read literally — it asks the histogram to
describe what is being looked at — and without it the instrument and the
controls disagree about what they are measuring. Costs a mask term in the
histogram reduction, which already runs per frame over the displayed frame.
### D-N2 — One layout, because both targets are wide · **PARTLY REVERSED**
> **Reversed for navigation, 2026-09-05, by use.** The reasoning below is
> still right about *size* and still right that `cfg(target_os)` is the wrong
> axis. It is wrong in one place, and the wrong bit is the sentence "touch
> changes **hit regions, not layout**". See D-N6.
>
> **Reversed for portrait, 2026-09-06, by arithmetic.** "Both orientations of
> both targets are the expanded class" is still true and is no longer the
> point. It was worked out for a 4:3 panel; the tablet's is 25:16, and on that
> aspect a column *beside* the photograph in portrait leaves it a strip. See
> D-N7, which keeps the layout class and adds an axis D-N2 did not consider.
The question was whether desktop and Android should diverge. The answer turns
out to be that **neither the platform nor the width axis separates DarkRoom's
targets**, so there is no divergence to build.
**The targets are a 12-inch tablet and a desktop.** No phone, decided
2026-08-22. A 12-inch tablet is roughly 1024 logical pixels across in portrait
and 1400 in landscape; `EXPANDED_MIN_WIDTH` is 820. **Both orientations of both
targets are the expanded class.** The compact class now fires only when a
desktop window is dragged under 820px, which is a case to degrade gracefully
into, not a second interface to design.
**Platform would have been the wrong axis anyway**, and it is worth recording
why so it is not proposed again. A tablet in landscape wants what a desktop
wants; a desktop window dragged narrow wants what a small screen wants.
Splitting on `cfg(target_os)` gives one *physical* situation two answers
depending on which binary it happens to be. `apply_layout_class` already says
this in its own comment — "logical pixels, not a device check" — and it was
right.
**What actually differs between the two targets is input, not size**, and the
architecture has already decided that too. `WidgetDemand::precise_pointing`
exists for a frontend driving a television with a remote, and its own
documentation states the position: *touch is fine, since hit regions grow to
the modality* (FR-UI-7). Touch changes **hit regions, not layout**. A control
is drawn where it belongs and its target grows past its own bounds — which
`Check` and the mask rows already do.
So: **one develop layout, tuned for a wide viewport, with touch targets
throughout.** The consequences worth stating:
- **A guaranteed-wide viewport is an asset.** The extent problem (§1.2) can be
solved by pinning rather than by hiding — see N4.
- **No hover-only affordance may carry meaning.** Hover may *emphasise*; it may
never be the only way to discover a control. The mask rows already obey this
— the eye and the delete target are drawn, not revealed.
- **No modifier key may be required.** A tablet has no shift. Local masking
already lost its shift-click extend for this reason, and nothing should
reintroduce one as the only route to a feature.
- **Anything dragged needs a finger-sized target.** ~~This is the live one:
gradient masks have no on-canvas handles yet~~ — they have them now (N2a), and
their handles are the first control in the app designed to be dragged on a
photograph rather than in a panel. Drawn at 14px so they do not hide the edge
they sit on, with a full touch target centred on the drawing, which is the
split `Button` already establishes. Nothing about them is revealed by hover
and nothing about them is qualified by a modifier: what is drawn is all there
is.
### D-N6 — The groups move to the rail under a finger · **DECIDED**
Reported from a tablet: the tool rail is *"very useful"* there, and the same
interface with a mouse and keyboard is not ergonomic. That is D-N2's assumption
failing in the field, and it is worth being precise about which half failed.
**What D-N2 got right.** Platform is the wrong axis, and width is the wrong
axis. A tablet in landscape wants what a desktop wants; a desktop window
dragged narrow wants what a small screen wants. `apply_layout_class` still
decides the layout class from the window, and nothing here changes that.
**What it got wrong.** It identified input as the real difference between the
targets and then concluded that input changes only hit regions. Two controls
answering one question — *which group of adjustments am I looking at* —
disprove it:
- **A horizontal strip above the column.** One gesture to a target the eye has
already found; costs one row of a column with height to spare. It pans when
the operation set is rich, so a group can be off the end with nothing saying
so — which a pointer user tolerates and a finger user does not discover.
- **A vertical run down the rail.** Every entry visible at once, each a
finger-sized target, on the edge of the screen the hand is already holding.
Costs nothing extra in width, because the rail is already there and already
mandated.
Neither is better in general. The first is better with a pointer and the second
is better with a finger, which is a divergence on **input modality** — the axis
D-N2 itself named.
**The shape.** `ToolRail` grows a second section below a rule: the same
`adjust-tabs` model the strip takes, plus "All". `GroupStrip` stands down when
the rail carries them, so the two are never both on screen and there is no
state to keep in step. Mode and group stay independent axes exactly as N1
requires — one entry lit in each section, and picking a group while a tool is
held still filters without putting the tool down.
**Drawn differently, still.** N1 insisted a mode and a filter must not be told
apart by the shape of their highlight alone. The tools fill with `active-dim`
and invert their ink; the groups take a bar down the leading edge — the
underline from the horizontal strip, turned ninety degrees. The rule between
the sections is the second signal.
**The rail scrolls now.** Its own note argued against a Flickable on the
grounds that four entries were written in the file. With the groups in it the
list is generated from the operation set, which is exactly the "something the
user's data decides" the note excluded it from.
**And it is a preference, because the automatic answer is a guess.** Neither
platform can be asked what the user is actually holding — an Android tablet in
a keyboard case is being driven like a desktop, and a touchscreen laptop is
whichever its owner says. `dr_plat::is_touch_first` reports the usual case per
platform and `dr_types::GroupNavigation` lets it be overridden; Settings names
what Automatic resolves to on this device rather than leaving it to be found by
pressing.
**Still open: whether Local is a mode at all.** The rail now holds two kinds of
entry, and a third reading is available — that Compose and Repair are
categories with a canvas gesture attached, while Local edits *nothing* and
instead changes what every other category applies to. That would make it a
**scope**, not a peer of the tools, and would collapse the two sections into
one list of seven. It is the tidier model and a much larger change; deferred
until the two-section rail has been lived with. §1.1's complaint was that scope
was invisible, so this is the same argument arriving from the other end.
### D-N7 — The column docks under the photograph on a tall window · **DECIDED**
D-N2 dismissed portrait with one number: a 12-inch tablet is about 1024
logical pixels across in portrait, which clears `EXPANDED_MIN_WIDTH`, so
portrait is expanded, so there is nothing to design. The number was for a 4:3
panel. **The tablet's panel is 3000 × 1920**, which is 25:16 — closer to a
sheet of A4 than to an iPad — and the same arithmetic on that aspect comes out
the other way.
**What the photograph gets.** Logical size depends on the density Android
reports, which nothing in this repository records (N6 measures it). At a scale
of 2.0 the window is 960 × 1500 in portrait; at 1.75 it is 1097 × 1714. Take
the first, subtract the 60px rail, the 360px column and the 44px status bar,
and the canvas beside the column is **540 × 1456** — a strip two and a half
times taller than it is wide. Against a column *under* the canvas, 480px tall:
| Photograph | Beside the column | Under the column | Gain |
|---|---|---|---|
| 3:2, landscape | 540 × 360 | 900 × 600 | 2.8× the area |
| 2:3, portrait | 540 × 810 | 651 × 976 | 1.4× the area |
At 1.75 the figures move and the ratios hold (2.4× and 1.3×). So on this panel
the dock wins for **both** orientations of the photograph, not only the
landscape frame one would guess it was for. That is what makes it a decision
rather than a preference: the column is on the right because the eye's path is
tool, photograph, adjustment, and in portrait the photograph in the middle of
that path is the thing being starved.
**What it is not.** N5 said "no bottom sheet, no second layout, no tool strip
along the bottom — those solve a phone". Still true of all three. This is not
a sheet: the column does not slide over the photograph, it sits beside it on
the other axis, with the same contents, the same collapse and the same toggle.
It is not a second layout in D-N2's sense: the layout class is still decided
by width, the compact class still means what it meant, and a tall narrow
desktop window gets exactly what a portrait tablet gets, which is FR-UI-1's
rule. And the rail does not move — `toolrail.slint` argued it never should,
and a vertical list of finger-sized entries wants height, which portrait has
more of.
**The axis is aspect, not width, and it is independent of the class.** A 960
wide portrait window is expanded by width and wants the dock; a 1500 wide
landscape one is expanded by width and does not. So this is a third property
beside `layout-class` and `panel-max-width`, set from the same place for the
same reason — a size that both derives from and feeds the layout is a binding
loop in Slint, and `apply_layout_class` already measures the window. The
window-resized callback reports width alone today and grows a height. The
threshold is height above 1.2 × width, with hysteresis wide enough that a
window resized across square does not flap.
**Slint cannot turn a layout on its side**, and does not need to. The develop
view is one `HorizontalLayout` of rail, canvas and column; it becomes a
`Rectangle` whose three children take `x`, `y`, `width` and `height` from the
flag. The column's own `VerticalLayout` and `Flickable` are untouched. The
alternative — the column subtree declared twice under two `if`s — is 400 lines
of bindings copied, in a file whose own notes record conditional children in
layouts as the shape that has produced binding loops before.
**The contents are the real cost.** Everything in the column was drawn for a
360px vertical scroll, and the dock is 900 to 1040 wide by about 480 tall.
Stretched to that width the stack works — sliders get longer tracks, the
histogram divides the width, text wraps — and it is one long scroll in a short
box, with the histogram scrolling away from the sliders it serves, which is
§1.2 again. The width is two and a half to three of today's columns, so the
composition that fits it is three of them side by side (N9): the instruments,
the sliders, and the panels the mode adds. That needs the panels declared a
second time, which is cheap only once their callbacks stop being forwarded
through the window root by hand (N8). The stretched stack ships first as the
stopgap (N7), because the photograph gets its area back on day one and the
dock's contents can be got right afterwards.
**The dock's height is mandated, as the column's width is.** `panel-width`
exists because a column that sizes itself to its contents is a photograph
that changes size when a caption does; a dock has the same disease on the
other axis. `dock-height` in `style.yaml`, 480 until N6 says otherwise: room
for the pinned instruments (about 260px per N4) and a group of sliders under
them, and a 3:2 frame at 900 wide still fits above it at either scale.
**The filmstrip stays where it is.** It takes its strip off the bottom of the
photograph on demand and it keeps doing so; with the dock below it sits
between the two. The photograph loses 108px while the roll is open, which is
what it loses today, and the roll is not made part of the dock because it is a
different kind of thing — navigation, not adjustment — and D-N6 has already
been through why two kinds of entry in one control need a rule between them.
**Not remembered separately.** `PanelChoices` keeps the user's open-or-closed
override per layout class. The dock does not add a class and does not add a
remembered state: closing the column in landscape closes the dock in
portrait, because it is the same column.
### D-N3 — Collapsible panels, not tool tabs · **OPEN**
For the expanded layout, extend `ui-refinement.md` Workstream C from
sections-inside-adjust to the panels themselves: each collapses to a header
carrying a modified dot, and collapse state survives a drag elsewhere.
**Why this over tabs, on the evidence above.** Tabs need a taxonomy, and the
taxonomy is the problem. `ui-refinement.md` condemns the `starts-group` flag
for being the core telling the panel where sections go, and FR-DEV-3a requires
that adding an operation needs no UI edit. A tab strip built from a hardcoded
op-id → tab table in `dr-ui` breaks the second; one built from a `group:` field
in `ops/*.yaml` risks breaking the first.
There *is* a legitimate route to tabs, and it should be recorded rather than
discovered later: the descriptor could declare an operation's **nature** —
tone, colour, detail, optics — the same shape as `Affects` and `ParamKind`
already take. The core would be saying *what the operation is*, which is its
business, and the frontend would remain free to render that as a tab, a
section heading, or nothing at all. That stays on the right side of §4.3a.
**The recommendation is to defer it.** Ten operations do not need eight tabs,
collapse needs no taxonomy at all, and the nature field is easy to add later
and awkward to remove. Revisit when the operation count passes roughly fifteen
— which FR-DEV-3's outstanding list will reach.
**Open, because it is a taste call**: whether the expanded layout should also
gain tabs, or stay a single collapsible stack indefinitely.
---
## 4. Workstreams
Numbered N to avoid colliding with `ui-refinement.md`'s A–F.
### N1 — The mode strip — **done**
**Deliverable.** ~~One control naming the current mode, replacing the implicit
`crop-mode` boolean: **Photo · Crop · Local**. Top of the canvas in expanded,
bottom in compact.~~ The selected mode is the accent's job — it means *active*,
which is exactly this.
`crop-mode` becomes one value of a mode enum rather than its own flag, so the
two modes cannot both be on, which today they can.
**Landed as one strip, not two.** The mode control and the existing group strip
(`All · Light · Colour`, derived from operation attributes) were going to sit
beside each other above the same column, which is two controls answering one
question — *what am I working on*. They are now one:
```
Crop · Local │ All Light Colour
```
Lightroom Mobile's bottom strip mixes Crop and Masking with Light and Colour
for the same reason, and it reads naturally because from the photographer's
side they are the same kind of choice.
The two halves are **different kinds of state and are drawn differently**: a
mode is a chip that fills with the accent when it is on, a group is a word with
a rule under it. That is what lets both be read at once, and both are on at
once routinely — see below.
**Mode and group are independent axes.** Picking `Light` while a mask layer is
selected filters *that layer's* chain and does not leave local mode. The
alternative — a group press quietly dropping the scope — would be §1.1's fault
reintroduced from the other end, and it would make `Light` mean two things
depending on where it was pressed.
**Where it sits.** Pinned above the develop column, where the group strip
already was, rather than at the top of the canvas. The half that filters the
column belongs to the column, and moving it onto the photograph would put it
somewhere the four principles say chrome should not be. The canvas keeps one
button — now *"Done Cropping"* / *"Done Masking"*, naming the mode it leaves —
because the column can be closed on a narrow window and no mode may be
inescapable.
**Also landed.** `GeometryPanel`'s Crop button is gone: a second control
entering the same mode is a second thing that has to agree about which mode the
view is in.
**Done when.** ~~Entering crop from the strip does what the crop button did;
Escape and back leave the innermost mode; no two modes are ever active
together.~~ All three, checked on screen as well as in tests — `back_step` has
one `LeaveMode` step covering both modes, and the enum makes "no two at once"
unrepresentable rather than merely untested.
### N2 — Local mode — **done**
**Depends on** N1.
**Deliverable.** Entering local mode turns the overlay on and arms picking
without either being a separate toggle — they are what the mode *is*. The
column shows the mask stack, then the scoped adjustments. The adjust header
names the layer.
Leaving local mode clears the selection so the adjustments are unambiguously
global again.
**Landed.** The "Overlay" and "Select" buttons are gone; entering the mode does
both, and `region-picking` is now derived from the mode rather than toggled.
The masking panel is no longer a panel among peers in the scrolling column — it
appears only in local mode, which is what takes the column from six panels to
three there.
**The scope is the adjust panel's own heading**, not a caption in the panel
above it. `ADJUST` becomes the layer's name. That is the difference between
describing the hazard and removing it: the heading of the thing that changed
cannot be skipped on the way to a slider, and a caption in a different panel
routinely was.
**What local mode drops from the column**: the capture metadata, the framing
controls and copy/paste. None is a property of a region within the photograph,
so all three would be controls in scope of nothing. The histogram stays and
still reports the whole frame — the disagreement §1.1 names is real and is N3's
to close; removing the instrument would be a worse answer than an honest one
that is not yet scoped.
**Done when.** ~~There is no way to have a mask selected without knowing it, and
the two toggles that currently arm the overlay and picking are gone.~~ Both.
### N2a — Gradient handles — **done**
Not a numbered workstream when this was written, and it belongs beside N2: the
canvas half of local mode.
`MaskSource::Linear` and `MaskSource::Radial` could be created and then not
moved, so a radial sat at the centre of the frame at its default size for ever.
They now carry handles on the photograph — the first controls in the
application designed to be dragged there rather than in a panel, and D-N2's
"the live one".
Three faults had to be fixed before a handle was worth drawing.
**A gradient did not render at all until the model had run.** The mask
rasteriser was built on the way out of `segment`, and the array's size was read
*off* the segmentation, so a gradient added to an unsegmented photograph
produced nothing — silently, because the generated shader still emits the
layer's block and the empty placeholder multiplies it by zero. The proxy size
is a property of the photograph; both are now derived from it, deliberately at
the same size because a subject's distance field is sampled against the array.
**A gradient's geometry was measured in raw `0..1` fractions**, so a 45° ramp
was not at 45° and a radial with equal radii drew an ellipse. Angles and
distances are now in the frame's isotropic units — y spans `0..1`, x spans
`0..aspect` — converted in exactly one place, `frame_delta` in `mask.wgsl`.
Only the *meaning* of the stored numbers changed; the sidecar format did not.
**Hit-testing has to go through the framing map.** A handle is drawn in output
coordinates and stored in source ones, and the two are separated by the crop,
the zoom, the pan, the straightening and the turns. `Framing::source_at` and
`Framing::output_at` are `wgsl_prologue` evaluated on the CPU, kept in that file
beside it so the correspondence is one file's problem.
**Handles.** A linear ramp has three — centre, width, angle. A radial has three
— centre and one per semi-axis, the major one carrying the ellipse's angle as
well as its length, because where an axis is put says both. A rotation arm was
tried on the radial and taken out: standing off the shape by a fixed distance,
it began outside the photograph at the size a new radial is created at.
**A drag is a displacement applied to where the mask was when the press
landed**, not a destination the handle is snapped to. Snapping jerks the handle
by up to half a touch target on the first press, and the target is finger-sized
(FR-UI-3).
**Two faults found by looking at the screen** rather than by reading the source,
both of the kind `ui-refinement.md`'s verification section warns about. A `1px`
rule with a size and no position is *centred* by Slint, so the develop column's
seam was a hairline down the middle of the panel — twice over, once in
`app.slint` and once in `AdjustPanel`. And handing Slint a new `ModelRc` for the
handles on every pointer event made the repeater rebuild its items, taking the
`TouchArea` holding the gesture with them: the handle jumped once and then went
dead under a finger that was still down. The model is now rewritten in place.
### N3 — Scope-following histogram
**Depends on** N2, which has landed, so this is next and is the outstanding
half of §1.1: the panel now says *which* chain the sliders edit, and the
instrument beside them still measures the other one.
**Deliverable.** The histogram reduction takes an optional mask; in local mode
it reduces over the selected layer's coverage. The panel says which it is
showing, because a histogram of a face is a strange shape and the user should
know why.
**Done when.** Selecting a layer visibly changes the histogram, and leaving
local mode restores the frame's.
### N4 — Collapsible panels
**Depends on** `ui-refinement.md` A. Extends C.
**Deliverable, two halves.**
*Collapse.* Each panel collapses to its header; headers carry the modified dot
C defines; state is keyed by panel identity and survives a slider drag.
Default: histogram open, the rest collapsed until touched.
*Pin.* The histogram and the scope header do **not** scroll. They sit above the
scrolling region, always visible.
Pinning is the half a guaranteed-wide viewport buys, and it addresses §1.2
directly rather than obliquely. The complaint is not that the column is long —
it is that the instrument every tonal control is judged against scrolls away
from the controls it reports on. Collapsing panels shortens the scroll;
pinning removes the problem. Together they cost about 260px of fixed height,
which a target that is never under 820px wide and rarely under 1000 tall can
afford.
**Done when.** The histogram is visible while any tonal slider is being
dragged, at every window size the targets produce, without the user having
scrolled to arrange it.
### N5 — Compact degrades, rather than diverges
**Depends on** N4.
Not a second interface. Below `EXPANDED_MIN_WIDTH` the develop column already
overlays rather than sits beside the canvas, and `apply_layout_class` already
remembers the user's override per class. With N4's collapse in place a narrow
window is a one-panel-at-a-time column by consequence rather than by design.
**Deliverable.** Confirm the narrow case is usable and fix what is not. ~~No
bottom sheet, no second layout, no tool strip along the bottom — those solve a
phone, and there is no phone.~~ Still no sheet and still no strip; but a
*tall* window is not a narrow one, and D-N7 (2026-09-06) puts the column under
the photograph there. N6–N9 carry it. This workstream keeps the narrow case.
**Done when.** A desktop window dragged to 700px shows the photograph and a
usable column, and nothing is unreachable that was reachable at 1400px.
### N6 — Measure the tablet
**Depends on** nothing. Hours.
Every figure in D-N7 is computed at a guessed scale factor. The tablet's
panel is 3000 × 1920 physical; its logical size is that divided by whatever
density Android reports, and the two plausible answers (2.0 and 1.75) put the
dock's width at 900 or 1037 and its available height 200px apart.
**Deliverable.** Log the window's scale factor and logical size once at
startup, beside the existing `apply_layout_class` call, at a level that
reaches `adb logcat`. Run it on the tablet in both orientations. Record the
four numbers in D-N7 and, if they move the table, correct it. #30
(NFR-COMPAT-1) wants a reference device named; this is one of the numbers
that names it.
**Done when.** D-N7 cites measured logical sizes, not a scale it assumed, and
`dock-height` has been checked against the measured portrait height.
### N7 — The column docks on a tall window
**Depends on** N6 only for the value of `dock-height`; the shape does not
wait for it.
**Deliverable, three parts.**
*The flag.* `window-resized` reports height as well as width.
`apply_layout_class` derives a third property, `column-below`, from the
aspect with hysteresis (enter above 1.25, leave below 1.15, or thereabouts —
the point is that a window resized across square does not flap), and sets it
beside `layout-class` and `panel-max-width`. It is not a layout class and
`PanelChoices` does not learn about it.
*The frame.* The `HorizontalLayout` holding `ToolRail`, `canvas-area` and
`develop-column` becomes a `Rectangle` and each child takes `x`, `y`, `width`
and `height` from the flag. The rail is full height on the left in both
cases. With the flag off the geometry is what the layout produced, to the
pixel — screenshot before and after and diff them. With it on, the column is
`dock-height` tall and runs from the rail's edge to the window's; the canvas
has what is left above it. `panel-visible` collapses the dock to zero height
exactly as it collapses the column to zero width.
*The contents, as a stopgap.* The column's stack stretches to the dock's
width: the Flickable's viewport width follows the dock, the `min-width` floor
that keeps a mandated column honest is still there and is simply not binding.
Sliders take the width. Nothing is reflowed; N9 does that.
`dock-height` goes in `style.yaml` next to `panel-width`, with the reasoning
D-N7 gives, and is read as `Theme.dock-height`.
**Done when.** On the tablet in portrait, a 3:2 photograph is drawn at the
width of the canvas, not at the width the old column left it; turning the
tablet moves the column back beside it with the same scroll position and the
same panels open; a desktop window dragged taller than it is wide does the
same; and the screenshot diff in landscape is empty.
### N8 — Develop callbacks onto a global
**Depends on** nothing; can run beside N7.
The develop panels — `AdjustPanel`, `MaskPanel`, `SpotPanel`, `ComposePanel`,
`TransferPanel`, `FocusPanel`, `HistogramPanel`, `InfoPanel` — each forward
their callbacks and take their inputs through the window root, and the
instantiation in `app.slint` that wires one up is twenty to forty lines. N9
needs each of them declared a second time, and copying that wiring is the
kind of duplication that drifts.
**Deliverable.** A Slint global (or one per panel family, if a single one
reads badly) carrying the develop callbacks and the inputs the panels bind
to. A panel calls the global directly; Rust hooks the global instead of the
window. The existing instantiation in the column shrinks to the properties
that genuinely differ by placement, which should be none. `Readout` in
`adjust.slint` is the precedent for a global in this codebase.
The tests in `ui/dr-ui/src` that drive these callbacks through the window
move to the global; count them before starting, so the ticket knows its own
size.
**Done when.** No develop panel's instantiation in `app.slint` forwards a
callback by hand, every existing test passes, and a second instantiation of
any panel is under five lines.
### N9 — Three columns in the dock
**Depends on** N7 and N8.
**Deliverable.** A second composition of the same panels for the dock, in
three columns of equal width side by side, each its own scroll:
1. **Instruments** — `InfoPanel`, `HistogramPanel`, `FocusPanel`. What the
sliders are judged against, pinned by construction: it does not scroll
with them because it is not in their column.
2. **Sliders** — `GroupStrip` above `AdjustPanel`, exactly as in the column.
With a group selected this is one screen of sliders; with All it scrolls.
3. **The mode's panels** — `ComposePanel` and `TransferPanel` in photo mode,
`SpotPanel` in repair, `MaskPanel` in local. Empty otherwise, which is
a signal of its own about which mode the view is in (§1.1).
Three columns at 900 wide are 300 each, and at 1037 they are 346: inside the
280–360 band `PANEL_MIN_WIDTH`'s note says the sliders stay accurate over.
The dock takes the three-column composition only when a third of its width
is at least `PANEL_MIN_WIDTH` — 840px of dock, which both scales of the
tablet exceed. Below that it keeps N7's stretched stack. Two compositions,
not three: a dock too narrow for three columns is a desktop window in an odd
shape, and N5 says that case degrades.
The panels are declared twice, once per composition, which is what N8 made
cheap. Declaring each once and positioning it by hand in both modes was
considered and rejected: a column that scrolls is a `Flickable`, a
`Flickable`'s children are its children, and a panel cannot be in two.
**Done when.** On the tablet in portrait the histogram is visible while any
slider in any group is dragged, with no scrolling to arrange it — N4's own
criterion, met in the dock by construction; every panel reachable in
landscape is reachable in portrait; and the instantiation of each panel in
the dock is a handful of lines.
---
## 5. Sequencing
```
ui-refinement A ──┬── C ──── N4 ──┐
└── D ──── F ├── N5
N1 ── N2 ── N3 ───────────────────┘
N6 ── N7 ──┐
├── N9
N8 ──┘
```
N1–N3 are independent of `ui-refinement.md` and can start now; they touch the
canvas and the develop column's contents, not its layout. N4 needs C's
`Section`. N5 needs both and should land last, as F does — it is the one that
rearranges everything.
N6–N9 are the portrait dock (D-N7) and run beside the first row rather than
after it. N7 is the one that touches the develop view's frame, so it should
not land in the same wave as F; N8 touches only plumbing and can. N9 waits
for both, and gains from N4 if N4 is in by then — a collapsible panel in a
300px column is worth more than in a 360px one.
## 6. Invariants, for every workstream
- **FR-DEV-3a.** Adding a pipeline operation must still surface in both
layouts with no UI edit. No file under `ui/` may name an operation.
- **ARCH §4.3a.** Composition is the frontend's decision. The core may declare
what an operation *is*; it may not declare where the panel puts it.
- **FR-UI-3 / FR-UI-7.** Touch changes hit regions, not layout. Every control
keeps a finger-sized target wherever it is drawn, and no hover-only
affordance and no modifier key may be the sole route to anything — a 12-inch
tablet has neither.
- **FR-UI-5.** Escape and the Android back gesture leave the innermost state
first. Every mode added here joins that order.
---
## 7. Place — the other half of "where am I"
§1 asked how someone *finds* anything. This asks how they stop losing what they
already found. Both are navigation; the second is the one nobody notices until
it is wrong, and then notices constantly.
### 7.1 Three failures, one cause
The library grid is gated on an `if` in the markup, so **every** route away from
it destroys the subtree and rebuilds it on return. The Flickable inside passes
its viewport through zero on the way out. Three consequences, reported
separately and all the same bug:
- *"Opening Settings and coming back puts me at the top."* The guard on the
scroll handler tested `show-library`, which stays true while Settings, Import,
People or the launch screen covers the grid. The teardown's scroll-to-zero
passed it, and the remembered position was overwritten with 0.
- *"Coming out of develop I lose the photograph I was editing."* The position
restored was where the *grid* was, not what was open — and walking the photo
roll moves the second a long way from the first.
- *"Launching puts me at the beginning."* Nothing was written down at all.
`app.slint` now computes `library-visible` once — the same expression the `if`
is spelled from — and Rust reads that rather than `show-library`. The two cannot
drift apart, which is what let them drift in the first place.
### 7.2 Two positions, and which one wins
Returning from develop has two candidates: `resume_at`, where the grid was, and
the open photograph. They agree in the ordinary case and disagree after a walk
along the roll.
The rule is not "pick one". The grid seeks to the remembered position, then
`reveal()`s the keyboard cursor, which is on the open photograph and which moves
the viewport as little as will bring it into view. A frame inside the remembered
screenful moves nothing; one outside it scrolls exactly far enough. One rule,
both behaviours — and the cursor rather than the selection, so a set of forty
photographs assembled in the grid survives having one of them opened.
### 7.3 A place is not a scroll position
What gets written down is the view, the scope, the filter and the photograph —
because a position without the filter that produced it names a row of a list
that no longer exists. Restoring them has an order for the same reason: scope,
then filter, then position, then the view. Each step changes what an ordinal
*means*.
Addressed by remote path and collection UUID, never by an ordinal or a row id.
See FR-UI-8 and `dr_types::place` for why, and `library::ordinal_of_path` for the
one place the ordering is inverted — through the grid's own `ORDER BY`, taken
verbatim, rather than spelled a second time.
### 7.4 The handover, and when to refuse it
The record travels through `.darkroom-derived/place.json`, so a session begun on
the desktop continues on the tablet. Newest wins; there is nothing to merge.
The interesting decision is the refusal. A place arriving from another device is
welcome on the way in and unwelcome the moment the photographer has started
working — a grid that jumped somewhere else mid-scroll because a round trip
finally landed would have lost their place to the feature meant to keep it. So
any scroll, scrub, scope change, filter or opened photograph closes the latch,
and a record that arrives after that is still written to disk and simply takes
effect at the next launch.
### 7.5 Two smaller instruments that were saying nothing
Both were "correct" in the sense of not being wrong, and both were useless.
- **The capture-time marker** rested greyed at mid-track until the first scroll,
on the reasoning that anchoring it would imply a choice the user had not made.
But the sidebar's claim is to say *when* you are, and that is known from the
first frame. It is now seeded from wherever the view sits.
- **The photo roll** brought the open frame into view by the shortest move,
which put it hard against one edge with nothing on that side. It now centres
on the first reveal of a develop session and steps minimally thereafter —
a one-shot request the strip consumes, so an overlay screen rebuilding the
view does not undo a roll the user has scrolled by hand.
+296
View File
@@ -0,0 +1,296 @@
# View composition: a controller for the display layer
TRACES: FR-UI-1 | FR-UI-6 | FR-UI-8 | FR-DEV-3a | NFR-P9
**Status:** Draft · 2026-08-09
**Companion to:** [architecture.md](architecture.md) §4.3a, [ui-refinement.md](archive/ui-refinement.md)
## Why
`dr-ui` has four views — launch, library, develop, collections — and no view
layer. `run()` in `ui/dr-ui/src/lib.rs` is 500 lines that construct every
controller, wire every cross-view callback, own the develop session, and hold
the only complete picture of what is on screen. It has become the controller
by accretion rather than by design, and it shows in three specific ways.
**The back-patched knot.** `run()` declares
```rust
let open_from_library: Rc<RefCell<Option<Rc<dyn Fn(String)>>>> = ...
```
wired empty at line 328 and filled at line 604, because the library grid needs
to open an image in develop and the develop closure needs the GPU context that
is built after the library is wired. The comment in the source is candid about
it: *"This cell is the knot between them: wired empty here, filled once `show`
exists."* A nullable function slot resolved at runtime is what a dependency
cycle looks like when the language will not let you write one directly.
**Two copies of load-and-display.** `show` (lib.rs:524) and the remote-fetch
timer body (lib.rs:657) run the same sequence after their inputs diverge: set
`load_error` empty, set camera / exposure / dimensions from metadata, branch on
whether sensor data was recovered, then either `set_adjust_enabled(true)` +
`sync_rows` + `redraw`, or clear the session, empty the rows, disable adjust
and show the fallback image — and on failure, clear the session and report the
error. One takes a path and one takes bytes; everything downstream is written
twice.
The shared preamble has *already* been factored out into `reset_view_state`,
which both branches call, so the direction is established — the post-load half
is simply the part that has not been done yet. The two halves have not yet
drifted in the fields they set; the argument for consolidating is to keep it
that way, since every future display field must currently be added in two
places.
**Diffuse ownership of window state.** Distinct `window.set_*` calls by module:
| Module | Distinct window properties written |
|---|---|
| `library_ui` | 24 |
| `lib.rs` (`run`) | 23 |
| `launch_ui` | 16 |
| `collections_ui` | 9 |
No one owns "what is on screen". The clearest symptom is view switching
itself: `show-launch` and `show-library` are two booleans encoding one piece
of state, written from seven call sites across three modules
(`lib.rs:353`, `launch_ui.rs:171`, `library_ui.rs:246/254/1486/1645/1655`).
Nothing prevents both being true, and `app.slint` compensates with
`if !root.show-launch && root.show-library` chains at lines 349, 386 and 491.
`develop.rs` is the exception that proves the point: it writes no window
properties at all, taking capabilities in and returning rows and images out.
It is the one module already shaped the way this document argues for.
None of this is broken. It works, and several of the surrounding patterns are
load-bearing and correct — the mpsc-plus-timer worker shape exists because
Slint's event loop must never block (NFR-P9), and `sync_rows` mutates rows in
place because replacing the model breaks slider dragging. This document
changes ownership, not threading and not Slint model handling.
## Non-goals
- No change to the threading model. Workers stay on threads; results still
return through channels drained by Slint timers (NFR-P9).
- No change to the pipeline, catalog, sync layer, or render path.
- No change to how operations reach the adjust panel. `develop.rs` already
composes from capabilities and names no operation (FR-DEV-3a); that stays.
- **No dynamic view instantiation.** See the constraint below.
## The constraint that shapes all of this
Slint has no runtime component instantiation. Views are selected by statically
compiled conditionals over properties — `if root.show-launch: LaunchScreen`
and friends — so every view must exist in the `.slint` source at build time.
Descriptors therefore drive **composition and chrome**: which views exist,
their labels, their order, whether each is currently available, and which one
is active. They cannot conjure a view body. This is a smaller claim than
"self-describing modules" might suggest, and stating it here is deliberate:
a design that assumed otherwise would hit the wall at codegen.
`build.rs` already generates `theme.slint` from `style.yaml`, so generating a
tab bar from the descriptor set is *possible* later. It is not proposed now —
generated `.slint` is markedly harder to debug than generated tokens, and the
win does not yet justify it.
---
## Stage 1 — `ViewController`
Independently landable, and worth landing whether or not stages 2 and 3
follow.
A `ViewController` owns the window handle and the state that currently floats
in `run()`'s closure captures: `session`, `entries`, `index`, `viewport`,
`rows`, and `redraw`. Sub-controllers are held by it rather than wired to each
other through `run()`.
```rust
pub struct ViewController {
window: slint::Weak<AppWindow>,
gpu: Option<GpuContext>,
session: RefCell<Option<DevelopSession>>,
entries: RefCell<Vec<PathBuf>>,
index: Cell<usize>,
viewport: Cell<(u32, u32)>,
rows: Rc<slint::VecModel<ParamRow>>,
library: Rc<LibraryController>,
collections: Rc<CollectionsController>,
launch: Rc<LaunchController>,
}
```
Three changes follow, in order:
**1a. One `present()`.** Both load routes converge on a single method:
```rust
impl ViewController {
fn present(&self, loaded: Result<Loaded, String>, name: &str) { ... }
}
```
The local path calls `self.present(load(gpu, path), &name)`; the remote timer
calls `self.present(load_bytes(gpu, &bytes), &name)`. The duplicated post-load
sequence exists once, so a new display field can only be added in one place.
**1b. `Rc<ViewController>` replaces the nullable callback cell.** The grid's
click handler captures the controller and calls
`controller.open_remote(path)`. Both sides now depend on the controller rather
than on each other, so the cycle disappears and the `Option` slot with it.
**1c. Collapse the four adjust callbacks.** `on_param_changed` (lib.rs:702),
`on_param_reset` (716), `on_reset_all` (730) and `on_curve_reset` (747) are
four blocks differing only in which `DevelopSession` method they call; each
then runs the identical `sync_rows` + `redraw` pair. A single
`self.mutate_session(|s| ...)` helper that performs the mutation and then
re-syncs makes the shared tail impossible to forget.
Note that not every handler wants that tail — the view/pan handlers below them
deliberately `redraw` without `sync_rows`, because panning changes no
parameter. The helper must therefore be the *opinionated* path for parameter
mutation, not a mandatory funnel for everything that touches the session.
**Verification.** These are refactors with no behavioural change: existing
tests in `lib.rs` (`describe_camera`, `describe_exposure`, `collect`) must pass
untouched, and manual checks cover launch → library → develop, next/previous,
slider drag, reset, and canvas resize.
---
## Stage 2 — The `View` trait
The same discipline `descriptor.rs` already applies to operations, applied one
level up. That precedent matters: the core publishes `OpDescriptor` /
`ParamKind` / `Presentation`, and `develop.rs` builds controls from it without
naming a single operation. Views are the same shape of problem.
```rust
pub struct ViewDescriptor {
pub id: ViewId,
pub label: LocalizedKey,
pub kind: ViewKind,
}
/// Closed, not a string — a controller must be able to match exhaustively
/// and know it has covered everything. Same reasoning as `WidgetKind`.
pub enum ViewKind {
/// Occupies the window alone; no chrome, not tabbable. Launch.
Modal,
/// Participates in the tab set. Library, develop.
Primary,
/// Renders beside a primary view. Collections.
Adjunct,
}
pub enum Availability {
Available,
/// Greyed out, with a reason the UI can show. Not hidden — a missing
/// tab is indistinguishable from a bug.
Unavailable(LocalizedKey),
}
pub trait View {
fn descriptor(&self) -> ViewDescriptor;
fn availability(&self, ctx: &AppContext) -> Availability;
fn activate(&self, ctx: &AppContext) {}
fn deactivate(&self, ctx: &AppContext) {}
}
```
Two properties are carried up from `descriptor.rs` deliberately, because they
are why that design works:
- **`ViewKind` is a closed enum.** Its `WidgetKind` counterpart says so
explicitly: *"a UI must be able to match exhaustively and know it has
covered everything the core can ask for."*
- **Descriptors are hints with a working fallback.** A curve degrades to
sliders. A view whose `kind` a shell does not implement still renders
standalone; nothing about tabbing is required for a view to function.
Labels are `LocalizedKey`, reusing the type `dr-pipeline` already exports —
`dr-ui` depends on `dr-pipeline`, so this adds no coupling, and it keeps view
labels on the same footing as operation labels. `labels::resolve` already takes
a `&str` and derives a readable fallback for uncatalogued keys, so a new view
appears with a sensible label before anyone writes a translation, exactly as a
new operation does today.
`availability` is the part that earns the trait rather than merely tidying.
Develop-without-a-GPU and collections-without-a-catalog are handled
inconsistently today — `run()` logs a warning and sets `backend` to "NO GPU",
`library_ui` guards each catalog access separately — and the differences are
not intentional. One method, asked before a view is offered, makes those cases
uniform and gives the shell something honest to display.
`activate`/`deactivate` exist so switching away can stop timers and release
GPU resources instead of leaking them, which nothing does today.
**Scope check.** Four views, all known, all in-tree, no third-party authors.
This abstraction has to justify itself against an `if` chain, and on tab
composition alone it would be close. The `availability` consolidation is what
tips it, because that is a correctness fix rather than a tidiness one.
**Verification.** Stage 2 retrofits the trait to the four existing views and
changes no Slint. The descriptors must reproduce current behaviour exactly
before stage 3 consumes them.
---
## Stage 3 — Data-driven view switching
With descriptors in place, the boolean pair collapses:
```
- in property <bool> show-launch;
- in property <bool> show-library;
+ in property <string> active-view;
+ in-out property <[TabEntry]> view-tabs;
```
`app.slint`'s conditionals key off `active-view == "library"` rather than a
two-boolean conjunction, and the illegal both-true state stops being
representable. `view-tabs` is a model the controller publishes from the
descriptor set, carrying label, id, and availability — so the tab bar is built
from data even though the view bodies are static.
Only `ViewController` writes `active-view`. The seven scattered `set_show_*`
call sites become `controller.activate(ViewId::Library)`, which is also the
hook where `deactivate` on the outgoing view runs.
**Verification.** Behavioural parity on every transition currently reachable:
launch → library on sign-in, library → develop on cell click, develop →
library on back, and the startup paths in `launch::Startup` (all three arms).
---
## Sequencing
| Stage | Depends on | Independently valuable |
|---|---|---|
| 1 — `ViewController` + `present()` | — | Yes: removes the cycle and the duplication |
| 2 — `View` trait | 1 | Yes: uniform availability handling |
| 3 — `active-view` | 2 | Yes: illegal states unrepresentable |
Stage 1 first regardless. The descriptor layer needs a controller to live in,
and the `present()` duplication is a live bug source independent of how views
are composed.
## Requirements
Stages 1 and 2 are traced by existing IDs — FR-UI-6 (shared components),
FR-DEV-3a (self-describing), NFR-P9 (no UI-thread blocking). Stage 3's claim,
that view composition is data-driven and view identity single-valued, has no
requirement covering it. Proposed for `requirements.md` §3.5, after FR-UI-7:
> **FR-UI-8 — Self-describing views.** Each view publishes a descriptor —
> identity, label key, kind, and current availability — and a single
> controller composes the interface from the descriptor set. View identity is
> single-valued: exactly one primary view is active at a time. A view that
> cannot currently function reports why, and the shell presents it as
> unavailable rather than omitting it. Adding a view requires no change to the
> shell beyond registering it and declaring its body.
Not added to the register by this document; adding it is a separate edit,
since `requirements.md` is the register of record and renumbering there ripples
into `traceability.md`.
+487
View File
@@ -0,0 +1,487 @@
# DarkRoom — A Windows installer from the Linux CI
**Satisfies:** FR-PLAT-WIN-1 · FR-PLAT-WIN-2 · FR-PLAT-WIN-3 · NFR-COMPAT-2 (a stated channel)
**Companion to:** [distribution.md](distribution.md) · [requirements.md](requirements.md) §3.8, §4.4 ·
[android-signing.md](android-signing.md)
Spec for producing `DarkRoom-<version>-x86_64-setup.exe` from the same Gitea runner that builds the
Arch package and the APK, with no Windows machine in the loop. It names the toolchain, what the
tree has to change to compile for the target, what the installer does, how the CI job is shaped,
and — because there is no Windows hardware on the runner — exactly how much of the result can be
verified before a person double-clicks it.
**Written as a spec; §10 is the report.** Every step of §9 has since been run —
[`docker/windows/`](../../docker/windows/) is the container, [`packaging/windows/darkroom.nsi`](../../packaging/windows/darkroom.nsi)
the installer, [`.gitea/workflows/windows-image.yml`](../../.gitea/workflows/windows-image.yml) and
the `windows` job in `build-and-test.yml` the CI leg, and §6's gate passes through row 4 under
Wine. Four claims in the first draft were wrong and are corrected in place with a note; §10 lists
them. Where a claim still rests on reading rather than running, it says so.
---
## 1. Why this is nearly free, and where it is not
The reason to write this at all is that the tree is closer to Windows than a Linux-only project
usually is. The three things that ordinarily make a cross-build to Windows a week of work are all
absent:
| Usual obstacle | Here |
|---|---|
| A C image library (libraw, libjpeg-turbo, lcms) | `rawler`, `zune-jpeg`, `jpeg-encoder`, all pure Rust |
| OpenSSL, or a TLS stack with a system dependency | `reqwest` on `rustls` + `webpki-roots`; `ring` cross-compiles to the GNU target |
| A GUI toolkit with a platform-specific build | Slint on `winit` + wgpu, which already runs the same code on Linux and Android |
`rusqlite` is `bundled`, so SQLite compiles with whatever C compiler the target has; that is the one
place a cross C compiler is required, and it is a package install rather than a port. The inference
engine is tract, in Rust, which is what [faces.md §3](faces.md) chose it for — and this is the
second time that choice pays: the C++ ONNX Runtime would have needed a prebuilt Windows binary
fetched at build time.
**Where it is not free** is `dr-plat` and the handful of paths above it, which is exactly where
NFR-PORT-1 says platform code should be and where §3 finds it. That the list in §3 is short and
every item on it is already behind a `cfg` is the measure of whether NFR-PORT-3 ("a third platform
requires implementing the platform interfaces only") was met. It was, nearly: the gaps are in
things that grew *above* `dr-plat` — a settings file path, an `xdg-open` — rather than in the
interfaces themselves.
---
## 2. Toolchain: the GNU target, from a container
Two Rust targets can produce a Windows binary from Linux.
| Target | Linker | What it needs on the runner | What it costs |
|---|---|---|---|
| **`x86_64-pc-windows-gnu`** | MinGW-w64 `gcc` | `gcc-mingw-w64-x86-64` (Debian/Ubuntu), `mingw-w64-gcc` (Arch) — one apt/pacman install | Binaries link `libgcc_s` and `libwinpthread` unless told not to; the SEH unwinder is MinGW's rather than MSVC's; DirectX bindings are less exercised (not used — §2.1) |
| `x86_64-pc-windows-msvc` | `lld-link` via [`cargo-xwin`](https://github.com/rust-cross/cargo-xwin) | The MSVC CRT and Windows SDK headers, fetched from Microsoft's servers by `xwin` on first use (~1.5 GB, licence-accepted by flag) | A download step in CI that depends on Microsoft keeping those URLs stable, and a licence the runner accepts on the project's behalf |
**Decision: GNU.** It is a package install, it is what `rustup target add` supports out of the
box, and every crate in the dependency graph that carries a C component (`ring`, `libsqlite3-sys`,
`zstd-sys` if present) builds against MinGW today. The MSVC route produces a marginally more
conventional binary — the same CRT every other Windows application links — and costs a
1.5 GB fetch of Microsoft-licensed headers on every cold CI run. That is the wrong trade for a
channel whose users are, for now, the author.
Static-link the MinGW runtime so the installer carries one file rather than three. The
configuration lives in the container as environment variables rather than in a `.cargo/config.toml`
— that file is untracked here on purpose, and the Android image sets its linkers the same way:
```sh
CARGO_TARGET_X86_64_PC_WINDOWS_GNU_LINKER=x86_64-w64-mingw32-gcc-posix
CARGO_TARGET_X86_64_PC_WINDOWS_GNU_RUSTFLAGS="-C link-args=-static-libgcc -C link-args=-static-libstdc++"
```
**Corrected.** The first draft added `-Wl,--whole-archive -lwinpthread` "so nothing imports
`libwinpthread-1.dll`". That flag breaks the link: forcing the whole archive drags in unused
winpthread objects whose kernel32 and msvcrt references land after those libraries on the link
line, and the build dies on a hundred undefined `__imp_` symbols. It was also unnecessary —
rustc's windows-gnu target links its own winpthread in self-contained mode, and the built
executable imports no MinGW library at all (§10). `-posix` is stated because Debian's bare
`x86_64-w64-mingw32-gcc` is an alternatives symlink to either thread model.
Two more things the container needs that the draft did not name: a **host** `gcc`, because build
scripts and proc-macros compile for Linux whatever the target and the very first one fails with
"linker `cc` not found" without it; and **Wine 10**, because rustc's std imports
`bcryptprimitives.dll` for its random source and Debian bookworm's Wine 8.0 does not have it —
the smoke test dies at load with `c0000135` before the first instruction. So the image is
`debian:trixie-slim`, which also ships Node 20 natively.
### 2.1 What the binary reaches at runtime
Nothing the installer has to carry. wgpu opens **Vulkan** (`dr_gpu::new_shared`, D1 — Vulkan on
both targets, and the shared-device path permits nothing else), and on Windows the Vulkan loader
`vulkan-1.dll` is installed by every GPU vendor's driver. Slint's femtovg renderer finds system
fonts through `fontdb`, so the `fontconfig` the Linux CI job installs has no Windows counterpart.
There is no `libxkbcommon`, no display-server library: `winit` speaks Win32 directly.
**DirectX 12 is deliberately not enabled.** wgpu supports it and on Windows would be the
conventional choice, but the develop pipeline's compute shaders are written once for Vulkan
(NFR-PORT-2) and validated on two Vulkan drivers already; a third backend is a third set of
driver behaviours to characterise (NFR-R1's tolerance argument), and no Windows machine that can
run this application lacks a Vulkan ICD. The same reasoning that keeps GL out of `new_shared`
keeps DX12 out here. It is one flag away if that turns out to be wrong.
### 2.2 Where it builds
The same shape as the Android leg: a job container built from a Dockerfile in the tree and pushed
to the Gitea registry, tagged by the tree id of its directory so an unrelated push reuses it
([`android-image.yml`](../../.gitea/workflows/android-image.yml) already does this and the comment
there explains why).
```
docker/windows/
Dockerfile debian:trixie + rustup (1.92.0, target x86_64-pc-windows-gnu) + gcc + gcc-mingw-w64-x86-64 + nsis + wine
build.sh run a command in the container; caches registry, target and the Wine prefix
package.sh the LFS guard, staging, then makensis
```
`makensis` is a Linux binary; NSIS has always been buildable and runnable on POSIX hosts, and
Debian ships it as `nsis`. Wine is in the image for §6, not for the build. `osslsigncode` is not
in it until there is a certificate to give it (§5.4).
---
## 3. What the tree has to change
Read from the source, not run. Everything here is `dr-plat` or the thin layer above it — nothing
in `core/` is touched, which is NFR-PORT-1 holding.
### 3.1 Already handled
- **`volumes.rs`** — card detection reads `/proc/mounts` and `/sys/block` under
`cfg(target_os = "linux")` and returns an empty list elsewhere. Windows gets no card detection
in this pass; the import flow's path picker still works. (A `GetDriveType`/`DRIVE_REMOVABLE`
implementation is a screen of code and a follow-up.)
- **`display.rs`** — the X11 and Wayland colour-profile readers are `cfg(all(unix, not(android)))`;
the fallback is FR-DSP-8's stated one. Windows ICC profiles via `GetICMProfile` are a follow-up
for the same reason.
- **`desktop_client.rs`** — the Nextcloud desktop client's Unix socket is `cfg(unix)`. On Windows
the client listens on a named pipe (`\\.\pipe\...`); until that is implemented FR-NC-6c's
integration is absent and the app behaves as it does on a Linux machine with no client running.
- **`secrets.rs`** — has a `PlatformSecretStore` for "any platform without an implementation"
that returns `SecretError::Unavailable` on every call. It is loud on purpose, so a Windows build
made with no further change *compiles*, starts, and fails at sign-in with a clear message. §3.2
is what turns that into a working store.
- **`keyring`, `x11rb`, `wayland-*`** are target-scoped dependencies already, so the Linux-only
crates are not even compiled.
### 3.2 Required before the installer is worth shipping
Ordered by what blocks a first sign-in. **All four are done**; each item says how.
1. **Secret store.** *Done.* `keyring` 4's `v1` feature set — the one the workspace already
asks for — includes `windows-native-keyring-store`, so the Credential Manager backend needed
no new feature name, only the crate as a `cfg(windows)` target dependency and the existing
Secret Service implementation's `cfg` widened to include Windows. One implementation over
both, because `keyring::Entry` is the same API over either; the only difference is that
`is_available`'s probe always succeeds on Windows, which is correct — Credential Manager is
always present, so FR-NC-2's degraded mode does not arise.
2. **Paths.** *Done* — [`platform/dr-plat/src/dirs.rs`](../../platform/dr-plat/src/dirs.rs). FR-PLAT-LIN-1
says XDG, and the code said it in five places by reading `XDG_*_HOME` and falling back to
`$HOME/.local/...`. On Windows `HOME` is normally unset, so every one of these degraded to a
relative path from the working directory — which for a Start Menu launch is
`C:\Windows\System32`. Now one function per kind of directory in `dr-plat`, with the Windows
branch reading `%APPDATA%` (config; roams) and `%LOCALAPPDATA%` (data, cache, state; does
not), and the five call sites using it. The Android overrides (`set_state_dir`,
`set_data_dir`) stay where they were; only the fallback behind them moved. Both platforms'
rules are unit-tested on either host, and the Windows one was confirmed by running the
application under Wine: the log landed in `AppData\Local\darkroom\state` and nothing was
written anywhere else.
| Kind | Linux today | Windows |
|---|---|---|
| config (`settings.json`, accounts) | `$XDG_CONFIG_HOME/darkroom` — `dr_sync::config_dir`, `settings_store.rs` | `%APPDATA%\darkroom` |
| data (catalog, thumbnails, faces) | `$XDG_DATA_HOME/darkroom` — `library::data_root` | `%LOCALAPPDATA%\darkroom` |
| state (crash reports, diagnostics) | `$XDG_STATE_HOME/darkroom` — `state.rs`, `crash.rs` | `%LOCALAPPDATA%\darkroom\state` |
FR-PLAT-WIN-1 states this as the requirement. The catalog and thumbnail *formats* do not change,
so a library directory copied from a Linux machine opens.
3. **Face models.** *Done.* `library::system_face_models_dirs` walked `$XDG_DATA_DIRS`, which
does not exist on Windows. The rule moved to `dr_plat::system_data_dirs`: the installer puts
the models beside the executable (§5), so the Windows branch returns the executable's own
directory. The user-directory lookups above it are unchanged, so a hand-placed pair still
outranks the installed one, exactly as on Linux.
4. **Opening the sign-in URL.** *Done.* `launch_ui.rs` shelled out to `xdg-open`. The Windows
branch runs `rundll32 url.dll,FileProtocolHandler <url>`, which is `ShellExecute` on the URL
and needs no crate — chosen over `cmd /C start`, whose quoting of `&` in a query string is a
known trap, and over the `open` crate, which would be a dependency for one line. Android has
its own Intent path already, so this is the third branch of a function that already had two.
*Not verified*: Wine has no browser to open.
5. **`std::os::unix` uses outside a `cfg`.** `diagnostics.rs` and `presets.rs` use
`PermissionsExt` for mode bits on written files. Most are inside `#[cfg(unix)]` blocks already;
the first cross-compile will name any that are not, and the fix is a `cfg` rather than a
Windows ACL equivalent — the files in question are the user's own.
6. **The executable's identity.** *Done.* Windows takes the icon and the version block from a
resource compiled into the `.exe`, not from a `.desktop` file.
[`apps/darkroom-desktop/build.rs`](../../apps/darkroom-desktop/build.rs) uses `winresource`
(which invokes MinGW's `windres` when cross-compiling) to embed
[`ui/dr-ui/ui/app-icon.png`](../../ui/dr-ui/ui/app-icon.png) — wrapped into an `.ico` in
`OUT_DIR` at build time, since an ICO entry may be a PNG, so no generated binary is committed —
plus the version from `CARGO_PKG_VERSION` and the product name. The script returns before
touching the crate on every other target, and `winresource` is an unconditional
build-dependency because **a `cfg(windows)` on a build-dependency is evaluated against the
host**, which is Linux. This is the fifth place the identifier lives, and
[`tools/set-version.sh`](../../tools/set-version.sh) does not need to learn it: the resource
reads the version cargo already knows. The same commit made the release binary a GUI-subsystem
executable (`windows_subsystem = "windows"`), or Windows keeps a console window open behind
the application.
Everything in this list is `cfg(windows)` code in `dr-plat` or a call-site switch in `dr-ui`, and
none of it touches the image core, the catalog schema, or the edit pipeline. That is the NFR-PORT-3
test, and it should be stated in the commit that closes the list whether it passed.
**What the first cross-compile actually found** (§9 step 2): nothing in this list blocked the
link. The whole graph compiled; the only warnings were two constants — `SERVICE` in `secrets.rs`
and `TIMEOUT` in `desktop_client.rs` — left unused by the `cfg`s that already shadow their users,
now guarded the same way. Items 1–4 are still open, and the binary starts without them; it just
cannot sign in.
### 3.3 Explicitly not in this pass
- **MIME/file-type registration** — FR-PLAT-LIN-1's `.desktop` MIME entries have a registry
equivalent (`HKCU\Software\Classes\.cr2` etc.). Not until the application opens a file from the
command line usefully, which `main.rs` accepts but the launch flow does not yet act on.
- **High-DPI declaration** — winit sets per-monitor-v2 awareness through its manifest by default.
Verified in winit's source, not on a monitor; if text is blurry on a 150% display this is the
first suspect.
- **Card detection, ICC profiles, the desktop-client pipe** — §3.1's three follow-ups.
- **A GL or DX12 fallback** — §2.1. A machine without Vulkan gets the library and no develop
path, which is what it gets on Linux too.
---
## 4. The build script
`docker/windows/build.sh`, in the shape of the Android one and with the same rules:
```sh
cargo build --release --target x86_64-pc-windows-gnu -p darkroom-desktop
```
Release only, with `CARGO_TARGET_DIR` inside the workspace so the CI cache key
(`windows-${{ hashFiles('**/Cargo.lock') }}`) covers it. The whole workspace is *not* built for
the target: `darkroom-android` cannot be, and the examples that need a display or a catalog on
disk have nothing to run against. `cargo clippy --target x86_64-pc-windows-gnu -p darkroom-desktop`
is worth running in the same job, because the `cfg(windows)` branches from §3 are otherwise
never linted — the Linux job cannot see them.
Not `cargo test --target x86_64-pc-windows-gnu`: the test binaries would be Windows executables,
and running them means Wine. §6 does that for exactly one binary, deliberately.
---
## 5. The installer
[`packaging/windows/darkroom.nsi`](../../packaging/windows/darkroom.nsi), compiled by `makensis` on
the runner into `DarkRoom-<version>-x86_64-setup.exe`. `package.sh` passes the version in
(`/DVERSION=…`, from `tools/set-version.sh`'s single source, the workspace `Cargo.toml`) and refuses
to run if any `models/face/*.onnx` is smaller than 100 KB — the LFS-pointer guard every other
packager carries, for the reason [distribution.md §1](distribution.md) gives.
### 5.1 Per-user, not per-machine
Install to `$LOCALAPPDATA\Programs\DarkRoom`, register the uninstaller under
`HKCU\Software\Microsoft\Windows\CurrentVersion\Uninstall\DarkRoom`, `RequestExecutionLevel user`.
No UAC prompt, no `Program Files`, no writes outside the user's profile. This is the shape VS Code's
"User Installer" and most Electron applications use, and it is right for this project for two
reasons: an unsigned installer that also asks for administrator rights is the most alarming thing
Windows can show a user (§5.4), and a per-user install means the application's own data directories
(§3.2) and its binaries are governed by the same account, which is what NFR-SEC-5's
"the user's own hardware" means on a shared machine.
### 5.2 What it puts on disk
```
$LOCALAPPDATA\Programs\DarkRoom\
darkroom.exe
models\
scrfd_500m_640.onnx scrfd_2.5g_640.onnx scrfd_10g_640.onnx arcface_mbf_b1.onnx
2d106det_b1.onnx ocec_s_b1.onnx sgc_l_48_b1.onnx
yolo26s-sem-ade20k.onnx yolo26s-sem-ade20k.classes.json categories.txt
LICENSE
uninstall.exe
```
Plus a Start Menu shortcut, and nothing on the desktop unless the user ticks it. The models are
the same ten files the APK bundles and the PKGBUILD installs; `models\` beside the executable is
where §3.2's lookup finds them. **No `LICENSE` yet**: the repository has no licence file at its
root (the Arch package points at the system's shared GPL text), so the installer has no licence
page until one is added — a one-file change, and the `.nsi` says where the page then goes. The face weights carry the research-only grant that
[faces.md §2](faces.md) records, and this channel changes nothing about that: the installer is
for the author's own machines until §2.2a's caveat is resolved, exactly as the APK is.
### 5.3 Uninstall
Removes the install directory, the shortcut and the registry key. **Does not touch
`%LOCALAPPDATA%\darkroom` or `%APPDATA%\darkroom`** — the catalog, the thumbnails, the face
index, the settings. An uninstaller that deletes a library index the user spent two hours building
is the kind of destructive default FR-CULL-12 and NFR-SEC-5's "disabling deletes nothing" both
argue against. The uninstaller says so on its one page, and names the two directories so a user who
does want them gone knows where they are.
### 5.4 Signing, and the warning that results from not doing it
An unsigned installer triggers SmartScreen's "Windows protected your PC" interstitial, dismissable
through "More info → Run anyway". Signing needs an Authenticode certificate, which is paid and
identity-verified; from Linux the signing itself is `osslsigncode`, which is why it is in the
container image, but there is no certificate to give it. **This spec ships unsigned** and the
release notes say what the interstitial looks like. An OV certificate is a cost decision to make
if this channel ever has a user who is not the author; an EV one buys instant reputation and costs
a hardware token. Neither is a build problem.
The APK went through the same sequence — [android-signing.md](android-signing.md) records a
debug-signed build becoming a release-signed one when it mattered — and this channel should be
allowed to do the same.
### 5.5 What NSIS is chosen over
WiX produces an MSI, which is what enterprise deployment tooling wants and what nobody deploying
a photo editor to their own laptop cares about; its Linux story is `wixl` from msitools, which is
real but thinly used. Inno Setup runs only under Wine. NSIS is scriptable in plain text, builds
natively on Linux, produces a single self-contained `.exe`, and the script for §5.2 is under a
hundred lines. It is the conventional answer for exactly this situation.
One choice inside NSIS: `Target amd64-unicode`, a 64-bit installer rather than the default 32-bit
stub. The application is x86_64 only so nothing is lost, and it is what lets §6's install test run
under a 64-bit-only Wine — the 32-bit stub needs an i386 multiarch Wine and dies loading the WoW64
`ntdll` without one.
---
## 6. Verifying without Windows
This is the part to be honest about. The runner has no Windows, no GPU it can hand to a Windows
process, and no display. What *can* be checked, in increasing cost and decreasing certainty:
| Check | How | What it proves |
|---|---|---|
| **It links** | the build succeeds | Every `cfg(windows)` branch compiles; no `unix`-only symbol leaked past a `cfg` |
| **It is a Windows executable** | `file darkroom.exe` reports PE32+; `x86_64-w64-mingw32-objdump -p` lists the DLLs it imports and none are MinGW's | The static-runtime flags in §2 held |
| **It starts** | `wine64 darkroom.exe --version` exits 0 and prints the version | The CRT, the resource block and `main` are sound; paths in §3.2 resolve (Wine sets `LOCALAPPDATA`) |
| **The installer runs** | `wine64 DarkRoom-setup.exe /S` then the install directory exists with the eight files, and `wine64 uninstall.exe /S` removes it | The NSIS script's file list, sections and uninstaller are right |
| **It draws a window** | `xvfb-run wine64 darkroom.exe` with `SLINT_WGPU_CPU` and a lavapipe ICD exposed through `winevulkan` | That Slint's winit backend initialises on Win32 — and this is where the chain gets long enough that a failure says more about Wine than about DarkRoom |
The first four are the CI gate. The fifth is worth trying once by hand and not putting in CI:
it needs `winevulkan` to find a host ICD, `xvfb`, and a Wine prefix warmed up in the container,
and every one of those is a moving part that has nothing to do with whether the application works
on Windows.
`--version` exists for this — a smoke test needs an exit that opens no window and touches no
directory — and it is answered before the logger and the crash hook install, so it proves the CRT
and the resource block and nothing above them.
**One more row the table missed:** the Start Menu shortcut. `CreateShortcut` is `IShellLink`,
which does nothing under a headless Wine while the `CreateDirectory` beside it succeeds, so an
installer that installs and uninstalls cleanly here can still have a broken shortcut. That row is
on Windows only.
**What none of this proves:** that wgpu opens a Vulkan device on a real driver, that a 6000-px
render completes, that fonts are found, that the secret store round-trips. Those are a person with
a Windows machine, once per release, until there is a Windows runner — and a self-hosted Windows
act_runner is how that would be done, not a cloud service. The release notes for the first build
say which of these were checked and on what.
---
## 7. The CI job
A fourth leg of [`build-and-test.yml`](../../.gitea/workflows/build-and-test.yml), beside desktop,
Android and traceability:
```yaml
windows-image:
uses: ./.gitea/workflows/windows-image.yml # same shape as android-image.yml
windows:
runs-on: linux/amd64
name: Windows (x86_64, cross)
needs: windows-image
container:
image: gitea.tourolle.paris/dtourolle/darkroom-windows:latest
steps:
- checkout, LFS pull # copied from the desktop leg
- cache: ~/.cargo, target # key: windows-${{ hashFiles('**/Cargo.lock') }}
- docker/windows/build.sh # cargo build + clippy, --target x86_64-pc-windows-gnu
- smoke: file, objdump, wine64 --version # §6 rows 1–3
- docker/windows/package.sh # LFS guard, makensis
- smoke: wine64 setup.exe /S; ls; uninstall # §6 row 4
- upload artefact: DarkRoom-*-setup.exe # on tags only, like the APK
```
Same gotchas as the Android leg, which its comments already record: the host has no Node, so the
checkout is plain `git`; workflow inputs arrive as strings; the image build needs the host Docker
daemon and runs outside a container. None of that is new.
**Cost.** A cold build of the whole graph for a second target is roughly the desktop leg again —
tract, Slint's compiler, wgpu — so with the cargo cache warm it is minutes and cold it is the
better part of half an hour. Worth noting because the runner is one machine and the legs run in
parallel on it; if it starts starving the desktop leg, `needs: desktop` serialises them.
---
## 8. Requirements
Three, added to [requirements.md §3.8](requirements.md) under a `#### Windows` heading beside the
Linux ones. Phrased to be testable, and each one is something §3 or §5 would otherwise leave as a
convention.
**FR-PLAT-WIN-1 — Known folders.** Configuration under `%APPDATA%\darkroom`; data, cache and
state under `%LOCALAPPDATA%\darkroom`. No file under the user's profile root and nothing relative
to the working directory. The directory *layout* beneath those roots is the same as under XDG, so
a library directory moves between platforms unchanged.
**FR-PLAT-WIN-2 — Installer.** A per-user installer that needs no elevation, registers an
uninstaller, and whose uninstaller removes what the installer wrote and nothing the application
wrote. Models are installed beside the executable and found there last, after the user's own
directories.
**FR-PLAT-WIN-3 — Built from Linux.** The Windows binary and its installer are produced by the
Linux CI from the same commit as every other channel, with no Windows machine in the build.
Verification on Windows is a release step, recorded per release, not a build step.
NFR-COMPAT-2's channel table in [distribution.md §1](distribution.md) gains a row. NFR-PORT-3 gets
its first real test, and the commit that closes §3.2 records the answer.
---
## 9. Order
1. `--version` in `main.rs`, and the `.cargo/config.toml` target block. Trivial, and the smoke
test in §6 needs both before anything else can be measured.
2. `rustup target add x86_64-pc-windows-gnu`, `pacman -S mingw-w64-gcc`, and a first
`cargo build --target …` on the developer machine — **before the container exists**, because
the list in §3.2 is a reading of the source and the compiler's list will be longer. Fix the
`cfg` fallout as it appears. This is the afternoon that decides whether §1's optimism holds.
3. §3.2 items 1–4, each its own commit, each stating which NFR-PORT interface it implemented.
4. §3.2 item 6 — the resource block — and the NSIS script; `makensis` by hand; `wine64 setup.exe /S`
by hand. Now there is an artefact.
5. The container, the image workflow, the CI leg. Only after 4 works locally, for the same reason
the Android image was reproduced from the tree after it had lived on one laptop.
6. A build on a real Windows machine, and a note in the release saying what was checked.
Steps 1–2 are cheap and either confirm this document or replace §3.2 with the true list. Nothing
past step 2 should be started on the strength of this document alone.
---
## 10. Report · 2026-09-12
Steps 1, 2 and 4 run, in the container rather than on the developer machine, because the
container was the cheaper way to get a pinned MinGW and a Wine that could be thrown away.
| §6 row | Result |
|---|---|
| It links | Yes, first attempt once the link flags were right. 115 MB, `PE32+ … (GUI)`. Two dead-code warnings, both `cfg`-shadowed constants, fixed. |
| It is a Windows executable | 26 imports, all Windows system DLLs. No MinGW runtime. `.rsrc` carries `PRODUCTVERSION 0,12,0,0`, `ProductName DarkRoom`, the icon. |
| It starts | `wine darkroom-desktop.exe --version` → `darkroom-desktop 0.12.0`, exit 0, 0.1 s. |
| The installer runs | `makensis` → 105 MB. `/S` installs the exe and ten models to `AppData\Local\Programs\DarkRoom`, writes the `HKCU` uninstall key; the installed exe runs; `uninstall.exe /S` removes directory and key. Shortcut unverifiable (§6). |
| It draws a window | Not attempted. |
**What the first draft got wrong**, kept in place above with a note rather than rewritten, because
the reasoning that produced each mistake is the thing a reader will otherwise repeat:
1. The `--whole-archive -lwinpthread` link flag (§2) — breaks the link and was never needed.
2. No host C compiler in the image (§2) — build scripts are host binaries.
3. Debian bookworm's Wine (§2) — lacks `bcryptprimitives.dll`, which rustc's std imports.
4. A `cfg(windows)` on the `winresource` build-dependency (§3.2 item 6) — evaluated against the
host, so the crate was silently absent from the cross-build.
And two things it did not know to say: NSIS's default stub is 32-bit (§5.5), and `CreateShortcut`
cannot be verified headless (§6).
**Closed since**, same day: all four §3.2 items (each says how), a `LICENSE` at the repository
root so the installer has its licence page, and §7's CI leg — `windows-image.yml` and the
`windows` job, every step of which was run by hand in the same container first. The Windows
target is also linted now, with `cargo clippy --target x86_64-pc-windows-gnu -- -D warnings` in
that job, which is the only place the `cfg(windows)` branches are ever compiled by CI.
**What the first real Windows run has to check**, in order, because Wine cannot: that a Vulkan
device opens on a real driver and a render completes; that fonts are found; that Credential
Manager round-trips a sign-in and the browser opens for Login Flow v2; that the Start Menu
shortcut exists; and that text is sharp on a scaled display (§3.3). The release notes for the
first build should say which of these were checked and on what machine.