Put the developer docs under docs/dev and index the folder for users first
docs/ had 26 developer documents flat beside the manual, and the two audiences are very differently sized: most readers want the manual and the gesture reference, a few want the register, the designs and the measurements. The manual and gestures.md stay at the top; everything for someone changing the code moves to docs/dev/, and the two documents that name their own successors — the v0.1 milestone and the UI-refinement plan — go to docs/dev/archive/ rather than being deleted, since both are still cited. docs/README.md is the index, users first. Every reference follows: code comments, Cargo manifests, the workflows, the pre-commit hook, the bench and traceability tools (which locate the repo root by docs/dev/requirements.md now), packaging, the Docker READMEs, CLAUDE.md, CONTRIBUTING.md and the README. The matrix links one level deeper and is regenerated. Links out of the moved documents into the tree gain a level; a link checker over every Markdown file finds none broken.
This commit is contained in:
@@ -0,0 +1,86 @@
|
||||
# Signing the Android build
|
||||
|
||||
Every build produces an APK. Which key signs it depends entirely on whether
|
||||
four secrets are present:
|
||||
|
||||
| secret | what it is |
|
||||
|---|---|
|
||||
| `ANDROID_KEYSTORE_BASE64` | the keystore file, base64-encoded |
|
||||
| `ANDROID_KEYSTORE_PASSWORD` | the store password |
|
||||
| `ANDROID_KEY_ALIAS` | the alias of the key inside the store |
|
||||
| `ANDROID_KEY_PASSWORD` | the key password |
|
||||
|
||||
With none of them set, `docker/android/assemble-apk.sh` generates a throwaway
|
||||
debug key and signs with that. That is the right answer for a branch build or
|
||||
a fork: the APK installs on a test device and nothing pretends it is a
|
||||
release. With all of them set, the same script signs with the real key.
|
||||
|
||||
The names match JellyTau's deliberately. One convention across both Android
|
||||
projects is one thing to remember instead of two.
|
||||
|
||||
## Making the key
|
||||
|
||||
Once, and then never again — keep it forever. Android identifies an app by
|
||||
its signature, so an app signed with a new key is a *different* app to every
|
||||
device that has the old one installed. There is no recovery from losing it
|
||||
beyond telling everybody to uninstall and reinstall.
|
||||
|
||||
keytool -genkeypair -v \
|
||||
-keystore darkroom-release.jks \
|
||||
-alias darkroom \
|
||||
-keyalg RSA -keysize 4096 -validity 10000 \
|
||||
-dname "CN=Duncan Tourolle, O=tourolle.paris, C=FR"
|
||||
|
||||
`keytool` prompts for the passwords rather than taking them on the command
|
||||
line, which keeps them out of shell history. Back the `.jks` up somewhere that
|
||||
is not this repository and not the machine that builds it.
|
||||
|
||||
**The key exists, since 2026-09-11.** It was made as above, with a random
|
||||
password, and the four secrets are loaded. The local copy is at
|
||||
`~/.config/darkroom/signing/` on the development desktop — `darkroom-release.jks`
|
||||
beside `storepass` and `keypass`, all mode 600 in a mode 700 directory. That
|
||||
copy is what `package.sh` can sign with locally:
|
||||
|
||||
D=~/.config/darkroom/signing
|
||||
KEYSTORE="$D/darkroom-release.jks" KEYSTORE_PASS="$(cat "$D/storepass")" \
|
||||
KEY_ALIAS=darkroom ./docker/android/package.sh --install
|
||||
|
||||
Before it existed, every build — CI and local alike — was signed with a
|
||||
throwaway debug key, and a debug key is exactly as durable as the cache
|
||||
directory it lives in: the local one was regenerated the night the cache was
|
||||
cleared, at which point no build anywhere could install over the device's copy.
|
||||
Any device that received a build from before this date has to uninstall once.
|
||||
|
||||
## Loading the secrets
|
||||
|
||||
base64 -w0 darkroom-release.jks > /tmp/ks.b64
|
||||
tea api --method PUT /repos/dtourolle/DarkRoom/actions/secrets/ANDROID_KEYSTORE_BASE64 \
|
||||
-f data=@/tmp/ks.b64
|
||||
shred -u /tmp/ks.b64
|
||||
|
||||
# and the three strings, read rather than typed so they miss the history
|
||||
read -rs PW && tea api --method PUT \
|
||||
/repos/dtourolle/DarkRoom/actions/secrets/ANDROID_KEYSTORE_PASSWORD -f data="$PW"
|
||||
|
||||
...and the same for `ANDROID_KEY_PASSWORD` and `ANDROID_KEY_ALIAS`. Or paste
|
||||
them into Settings → Actions → Secrets in the web UI, which is less fiddly and
|
||||
just as good.
|
||||
|
||||
## Checking which key signed a build
|
||||
|
||||
The packaging step prints it, and the APK carries it:
|
||||
|
||||
apksigner verify --print-certs darkroom.apk
|
||||
|
||||
A debug build says `CN=Android Debug`. Anything else is the real key.
|
||||
|
||||
## Signing locally
|
||||
|
||||
`docker/android/package.sh` takes the same environment variables, so a local
|
||||
release-signed build is:
|
||||
|
||||
KEYSTORE=$PWD/darkroom-release.jks KEY_ALIAS=darkroom \
|
||||
KEYSTORE_PASS=... KEY_PASS=... ./docker/android/package.sh
|
||||
|
||||
Without them it debug-signs, and keeps one debug keystore in the build cache
|
||||
so repeat installs to a device do not need an uninstall first.
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,327 @@
|
||||
# DarkRoom v0.1 — Remote library viewer
|
||||
|
||||
**Status:** Delivered and superseded · written 2026-08-08, closed 2026-08-30
|
||||
|
||||
> Kept as the record of what the first milestone asked for, not as a plan.
|
||||
> Everything below shipped, and the application went well past it — see
|
||||
> [outstanding.md](../outstanding.md) for what is still missing at 0.9.0.
|
||||
**Companion to:** [requirements.md](../requirements.md) · [architecture.md](../architecture.md)
|
||||
|
||||
The first buildable milestone: connect to a Nextcloud folder, index it locally, and display RAW
|
||||
previews on both Linux and Android.
|
||||
|
||||
---
|
||||
|
||||
## 1. What this is
|
||||
|
||||
A photo *viewer*, not an editor. It connects to a Nextcloud account, lets the user pick a folder,
|
||||
indexes what is there into a local catalog, and displays images by extracting their embedded JPEG
|
||||
previews over byte-range requests.
|
||||
|
||||
**It exists to prove the architecture is sound before anything is built on it.** Four assumptions
|
||||
in the design would each be expensive to discover wrong later, and this milestone tests all four
|
||||
against real servers, real files, and real devices.
|
||||
|
||||
### 1.1 Why this shape
|
||||
|
||||
The alternative first milestone (D3's vertical slice: local scan → grid → open → two edits →
|
||||
export) proves the pipeline but is useful to nobody and tests nothing about sync or SAF. This
|
||||
milestone is comparable in size, exercises the genuinely risky parts, and produces something you
|
||||
can actually point at a server and use.
|
||||
|
||||
### 1.2 What it deliberately excludes
|
||||
|
||||
No editing. No develop pipeline, no sidecars, no export, no culling mode, no ratings. Those depend
|
||||
on foundations this milestone establishes; adding them before the foundations are proven risks
|
||||
building on sand.
|
||||
|
||||
---
|
||||
|
||||
## 2. The four assumptions under test
|
||||
|
||||
Each maps to a spike in [requirements.md §9](../requirements.md).
|
||||
|
||||
| # | Assumption | If wrong | Spike |
|
||||
|---|---|---|---|
|
||||
| **A1** | Slint can composite a wgpu compute texture zero-copy, on both platforms | ARCH §6.1 fails; D1's fallbacks apply; **the framework choice is wrong** | S1, S2 |
|
||||
| **A2** | Android SAF can enumerate and range-read at library scale | FR-CAT-1a and the Android performance targets fail | S10 |
|
||||
| **A3** | HTTP Range extraction of embedded previews works against Nextcloud | Remote browsing on mobile data is not viable; ARCH §6.7 fails | S4 |
|
||||
| **A4** | reqwest does TLS on Android without unmanageable pain | D7's escape hatch needed | S3 |
|
||||
|
||||
**A1 is the one that would hurt most.** It determines whether Rust + Slint was the right call at
|
||||
all, so it should be proven in the first week, before catalog or sync work begins.
|
||||
|
||||
---
|
||||
|
||||
## 3. Functional scope
|
||||
|
||||
Requirement IDs reference [requirements.md](../requirements.md); a v0.1 suffix marks a reduced subset
|
||||
of the full requirement.
|
||||
|
||||
### 3.1 Account and connection
|
||||
|
||||
**M-1 — Login.** Connect to a Nextcloud instance via Login Flow v2 (FR-NC-1): POST to
|
||||
`/index.php/login/v2`, open the returned URL in the **system browser**, poll until an app password
|
||||
arrives. The app never handles the user's primary password.
|
||||
|
||||
`User-Agent` identifies the device so the resulting app password is revocable per-device.
|
||||
|
||||
**M-2 — Credential storage.** Store the app password in platform secure storage (FR-NC-2) — Secret
|
||||
Service on Linux, Keystore-backed on Android. Never in the catalog, never in logs.
|
||||
|
||||
**M-3 — Folder selection.** Browse the remote tree and choose one folder as the library root.
|
||||
Recursive descent is in scope; multiple roots are not.
|
||||
|
||||
**M-4 — Disconnect.** Revoke via `DELETE /ocs/v2.php/core/apppassword` and clear local credentials.
|
||||
|
||||
### 3.2 Indexing
|
||||
|
||||
**M-5 — Remote listing.** `PROPFIND Depth:1` walking the chosen folder recursively, requesting
|
||||
`oc:fileid`, `getetag`, `getcontentlength`, `getlastmodified`, `resourcetype`, and
|
||||
`nc:has-preview`. Files whose extension matches the supported set (§3.3) are catalogued; others are
|
||||
ignored.
|
||||
|
||||
**M-6 — Catalog.** Persist to SQLite in WAL mode. v0.1 schema is a strict subset of ARCH §6.2:
|
||||
|
||||
```sql
|
||||
schema_version(version)
|
||||
|
||||
accounts(id, server_url, login_name, user_id)
|
||||
|
||||
folders(id, account_id, parent_id, remote_path, etag, last_listed)
|
||||
|
||||
images(id, account_id, file_id, remote_path, etag, size, remote_mtime,
|
||||
format, captured_at, camera, lens, iso, aperture, shutter,
|
||||
width, height, availability)
|
||||
|
||||
previews(image_id, kind, width, height, bytes, path, last_used)
|
||||
```
|
||||
|
||||
`folders.etag` is present **from schema v1** — ARCH §6.6 requires it, and retrofitting means a
|
||||
migration plus a full re-scan of every library.
|
||||
|
||||
`schema_version` exists from the first commit so NFR-R5's migration machinery has somewhere to
|
||||
start.
|
||||
|
||||
**M-7 — Incremental re-listing.** On subsequent syncs use ETag pruning (FR-NC-4): `PROPFIND
|
||||
Depth:0` on the root, and if its ETag is unchanged, **stop** — one request proves the whole library
|
||||
is unchanged. Where changed, `Depth:1` and recurse only into folders whose ETags differ.
|
||||
|
||||
This is the mechanism that has to work for the library to scale, so v0.1 exercises it deliberately
|
||||
rather than always re-listing.
|
||||
|
||||
**M-8 — Offline browsing.** The catalog is queryable with no network. Previously indexed images
|
||||
appear with their metadata and any cached previews. Availability is visible per image (FR-NC-6c):
|
||||
*Preview* or *Metadata only*.
|
||||
|
||||
### 3.3 RAW handling
|
||||
|
||||
**M-9 — Formats.** The FR-RAW-1 launch set: CR2, CR3, NEF, ARW, RAF, RW2, ORF, DNG. Plus JPEG, so
|
||||
a mixed folder displays sensibly.
|
||||
|
||||
**M-10 — Range-based preview extraction.** Display without downloading whole files:
|
||||
|
||||
1. Range-read the first 256 KB of the file
|
||||
2. Parse the container to locate the embedded JPEG preview (offset and length)
|
||||
3. Range-read exactly those bytes
|
||||
4. Decode the JPEG and cache it locally
|
||||
|
||||
Typical cost 1–3 MB against 25–100 MB for the full file. **This is what makes the app usable on
|
||||
mobile data**, and it is assumption A3.
|
||||
|
||||
Range support is detected by issuing a `Range` request and checking for `206` versus `200` —
|
||||
**never** by probing with `HEAD`, since Nextcloud does not advertise `Accept-Ranges` (ARCH §6.7).
|
||||
|
||||
**M-11 — Fallbacks, in order.** Where step 2 finds no usable preview:
|
||||
|
||||
1. Server preview via `/core/preview?fileId=…&forceIcon=false` where `nc:has-preview` is true.
|
||||
**`forceIcon=false` is mandatory** — the default returns a generic mimetype icon for files the
|
||||
server cannot render, which would otherwise be cached as though it were a thumbnail.
|
||||
2. Full download and decode via rawler, on explicit user action only, never automatically.
|
||||
3. Placeholder with an explanatory state.
|
||||
|
||||
**M-12 — Metadata.** Extract from the same header range already fetched in M-10: camera make and
|
||||
model, lens, capture time, ISO, aperture, shutter, dimensions. No second request.
|
||||
|
||||
### 3.4 Display
|
||||
|
||||
**M-13 — Grid.** Virtualised thumbnail grid (FR-CAT-4) rendering only visible cells plus a prefetch
|
||||
margin, with bounded memory independent of folder size.
|
||||
|
||||
**M-14 — Single image view.** Full-window display of one image, with fit and 1:1 zoom, pan, and
|
||||
next/previous navigation.
|
||||
|
||||
**M-15 — GPU display path.** Decoded previews upload to GPU textures and composite through Slint
|
||||
via `create_texture_from_hal`. **Pixels are never read back to the CPU** (ARCH §6.1).
|
||||
|
||||
This is assumption A1, and the reason it is in v0.1 at all: it is far cheaper to discover a
|
||||
compositing problem now than after a develop pipeline is written against it.
|
||||
|
||||
**M-16 — Adaptive layout.** Compact and expanded layout classes (FR-UI-1) driven by window size,
|
||||
not device type. Touch targets meet 44pt when touch is the active modality (FR-UI-3).
|
||||
|
||||
### 3.5 Platform
|
||||
|
||||
**M-17 — Linux.** X11 and Wayland. XDG base directories for catalog, cache, and config. Secret
|
||||
Service for credentials.
|
||||
|
||||
**M-18 — Android.** Storage Access Framework only (FR-PLAT-AND-1) — no `MANAGE_EXTERNAL_STORAGE`,
|
||||
no `READ_MEDIA_IMAGES`. v0.1 reads remote content, so SAF matters for the *cache* location and for
|
||||
proving the `SourceRef` abstraction holds before local library support arrives.
|
||||
|
||||
Keystore-backed credential storage. Process death mid-browse restores the current folder and scroll
|
||||
position (FR-PLAT-AND-3).
|
||||
|
||||
---
|
||||
|
||||
## 4. Explicitly out of scope
|
||||
|
||||
Listed so absence reads as a decision.
|
||||
|
||||
| Excluded | Arrives in |
|
||||
|---|---|
|
||||
| Any editing, develop pipeline, sidecars | v0.2 |
|
||||
| Export | v0.2 |
|
||||
| Culling mode, ratings, flags, labels | v0.3 |
|
||||
| Focus peaking, raw histogram | v0.3 |
|
||||
| Local (non-Nextcloud) library scanning | v0.2 |
|
||||
| Upload, bi-directional sync, conflict merge | v0.4 |
|
||||
| Cache rules and pinning (FR-NC-6a) | v0.4 |
|
||||
| Multiple accounts or roots | later |
|
||||
| Ingest from card | later |
|
||||
| Search and filter beyond folder navigation | v0.3 |
|
||||
|
||||
**Read-only against the server.** v0.1 performs no `PUT`, `MOVE`, or `DELETE` on remote content.
|
||||
This removes conflict handling and chunked upload from scope entirely, and means a bug cannot
|
||||
damage the user's library.
|
||||
|
||||
---
|
||||
|
||||
## 5. Acceptance criteria
|
||||
|
||||
Measured against a reference Nextcloud instance holding **≥5,000 RAW files**, on the reference
|
||||
desktop and **two Android devices with different GPU vendors** (Adreno and Mali).
|
||||
|
||||
| # | Criterion | Target |
|
||||
|---|---|---|
|
||||
| **AC-1** | Login through to first grid render | < 30 s for 5,000 files |
|
||||
| **AC-2** | Re-sync with nothing changed | **1 HTTP request** |
|
||||
| **AC-3** | Grid scroll, cached previews | 60 fps sustained, both platforms |
|
||||
| **AC-4** | Bytes transferred per image, preview path | < 3 MB average |
|
||||
| **AC-5** | Single-image display from cache | < 100 ms |
|
||||
| **AC-6** | Offline launch → browsable grid | < 2 s desktop, < 4 s Android |
|
||||
| **AC-7** | Memory, 5,000-image folder open | < 400 MB desktop, < 200 MB Android |
|
||||
| **AC-8** | GPU readback of image data | **Zero occurrences** — asserted by instrumentation |
|
||||
| **AC-9** | Android process death mid-browse | Folder and scroll position restored |
|
||||
| **AC-10** | Catalog after forced kill during sync | Opens clean, resumes |
|
||||
|
||||
**AC-2 and AC-8 are the load-bearing ones.** AC-2 proves ETag pruning works, which is what lets the
|
||||
library scale. AC-8 proves ARCH §6.1 holds — and it is asserted by instrumentation rather than
|
||||
inspection, because a readback introduced later would otherwise pass unnoticed until it showed up
|
||||
as unexplained slowness.
|
||||
|
||||
---
|
||||
|
||||
## 6. Build order
|
||||
|
||||
### Phase 0 — de-risk (before anything else)
|
||||
|
||||
Throwaway code answering the four assumptions. If A1 fails, stop and revisit D1 rather than
|
||||
building on it.
|
||||
|
||||
1. **A1 / S1** — wgpu compute writes a texture; Slint composites UI over it; Linux. Test on
|
||||
Mesa/AMD, Intel, and NVIDIA proprietary, under X11 and Wayland.
|
||||
2. **A1 / S2** — the same on both Android devices.
|
||||
3. **A4 / S3** — reqwest HTTPS PROPFIND from an Android device, including the
|
||||
`rustls-platform-verifier` Kotlin init.
|
||||
4. **A3 / S4** — range-extract an embedded JPEG from a CR3, NEF, and ARW on a real server; measure
|
||||
bytes.
|
||||
5. **A2 / S10** — SAF enumeration and range reads over a 10k-file tree; measure against the
|
||||
NFR-P1/P3 figures.
|
||||
|
||||
### Phase 1 — foundations
|
||||
|
||||
`dr-types` (`SourceRef`, ids) · `dr-plat` traits and both implementations · `dr-catalog` (v0.1
|
||||
schema, migrations from commit one) · `dr-gpu` (device, texture upload, Slint bridge).
|
||||
|
||||
### Phase 2 — connectivity
|
||||
|
||||
`dr-sync` (`RemoteBackend` trait) · `dr-sync-nextcloud` (Login Flow v2, PROPFIND, ETag pruning,
|
||||
range GET) · credential storage.
|
||||
|
||||
### Phase 3 — imaging
|
||||
|
||||
`dr-decode` (container parsing, embedded-preview location, JPEG decode, metadata) · preview cache.
|
||||
|
||||
### Phase 4 — interface
|
||||
|
||||
`dr-ui` (grid, single-image view, adaptive layout, folder picker, connection flow).
|
||||
|
||||
### Phase 5 — hardening
|
||||
|
||||
Offline behaviour · process death · cancellation · error surfaces · the AC suite in CI.
|
||||
|
||||
**Both platforms build in CI from the first commit.** This is the point of choosing both from day
|
||||
one: an Android break is caught the day it lands, not at a porting milestone.
|
||||
|
||||
---
|
||||
|
||||
## 7. Crates in play
|
||||
|
||||
A subset of ARCH §2. Crates not listed are not created yet.
|
||||
|
||||
```
|
||||
core/
|
||||
dr-types ✓ SourceRef, ImageId, AccountId, Validator
|
||||
dr-catalog ✓ v0.1 schema subset, queries, migrations
|
||||
dr-decode ✓ preview extraction + metadata; no demosaic
|
||||
dr-gpu ✓ device, texture upload, Slint bridge; no pipeline
|
||||
dr-sync ✓ RemoteBackend trait, ETag pruning engine
|
||||
dr-sync-nextcloud ✓ the connector
|
||||
dr-pipeline ✗ v0.2
|
||||
dr-colour ✗ v0.2
|
||||
dr-export ✗ v0.2
|
||||
dr-sidecar ✗ v0.2
|
||||
ui/
|
||||
dr-ui ✓ grid, viewer, connection flow
|
||||
dr-widgets ✗ v0.2 (no custom controls yet)
|
||||
platform/
|
||||
dr-plat ✓ Storage, Secrets, Lifecycle traits
|
||||
dr-plat-linux ✓
|
||||
dr-plat-android ✓
|
||||
apps/
|
||||
darkroom-desktop ✓
|
||||
darkroom-android ✓
|
||||
```
|
||||
|
||||
`dr-gpu` exists in v0.1 **only** to upload decoded JPEGs and hand textures to Slint. No compute
|
||||
pipeline, no tiling, no masks. It is deliberately the thinnest thing that still proves A1.
|
||||
|
||||
---
|
||||
|
||||
## 8. Decisions this milestone informs
|
||||
|
||||
| Decision | What v0.1 tells us |
|
||||
|---|---|
|
||||
| **D12** — scope vs pace | How long a milestone of this size actually takes at the available pace. The single most useful output. |
|
||||
| **D3** — first milestone | Supersedes the vertical slice, if this proves the better shape. |
|
||||
| **D1** — Rust + Slint | AC-8 and A1 either confirm the framework choice or reopen it. |
|
||||
| **NFR-COMPAT-1** — hardware baseline | Two real Android devices give the baseline actual numbers instead of a placeholder. |
|
||||
| **D7** — network stack | Whether the reqwest Android TLS path is a half-day or a fortnight. |
|
||||
|
||||
---
|
||||
|
||||
## 9. Risks
|
||||
|
||||
| Risk | Likelihood | Mitigation |
|
||||
|---|---|---|
|
||||
| Slint `create_texture_from_hal` doesn't work as documented | Medium | Phase 0 first; D1 records fallbacks |
|
||||
| SAF enumeration too slow at 10k files | Medium | S10 measures before commitment; batch and cache aggressively |
|
||||
| Embedded previews too small or absent on some bodies | High | Known — Sony embeds small previews, some bodies none. M-11's fallback chain handles it; detect per camera model |
|
||||
| **rawler exposes only full-resolution previews** | **Confirmed** | Measured 2026-08-09: rawler 0.7.2's CR2 decoder implements `full_image` only; `thumbnail_image`/`preview_image` are unimplemented defaults. Every rung resolves to a 5472×3648 decode at ~250 ms, 5× over NFR-P13. CR2 does carry smaller IFDs, so the fix is our own IFD walk or an upstream contribution — not a change to callers |
|
||||
| **Android secret storage unimplemented** | **Confirmed** | Needs no investigation — `PlatformSecretStore` on Android is unimplemented by design, and fails loudly rather than silently no-opping (`platform/dr-plat/src/secrets.rs`). The fix is a real Keystore-over-JNI implementation (FR-PLAT-AND-1), which is `dr-plat-android` work not yet started |
|
||||
| reqwest Android TLS worse than expected | Medium | D7 escape hatch: `tls_certs_only` with `webpki-roots` |
|
||||
| GPU vendor divergence on Android | Medium | Two vendors in CI from the start |
|
||||
| Scope creeps toward editing | **High** | §4 is explicit; v0.1 is read-only against the server |
|
||||
|
||||
The last one is the real risk. A viewer that works is a strong temptation to add "just one slider."
|
||||
@@ -0,0 +1,572 @@
|
||||
# UI refinement: toward a Lightroom-shaped darkroom
|
||||
|
||||
TRACES: FR-UI-1 | FR-UI-3 | FR-DEV-3a | FR-CAT-4
|
||||
|
||||
## Why
|
||||
|
||||
The v0.1 UI proved the architecture: capability-driven controls, a windowed
|
||||
grid, a GPU canvas with no CPU round trip. What it has not yet done is *feel*
|
||||
like a photo editor. The gaps are structural rather than cosmetic, and this
|
||||
document names them so they can be closed independently.
|
||||
|
||||
Four principles guide every change below.
|
||||
|
||||
1. **The image is the subject.** Chrome recedes; nothing competes with the
|
||||
photograph for attention or for colour.
|
||||
2. **Colour is a signal, not a decoration.** The accent means *modified* or
|
||||
*active*. Everywhere it currently means "heading" or "chrome", it is
|
||||
spending a signal on noise.
|
||||
3. **Navigation is continuous.** Moving between images should not be a change
|
||||
of screen. Lightroom's filmstrip is the mechanism; modal view switching is
|
||||
what it replaces.
|
||||
4. **Density is earned.** A panel shows what has been touched; everything else
|
||||
collapses out of the way.
|
||||
|
||||
## Non-goals
|
||||
|
||||
- No new pipeline operations, and no change to how operations reach the panel.
|
||||
`AdjustPanel` must still learn its contents from the capability model and
|
||||
must still name no operation (FR-DEV-3a).
|
||||
- No change to the catalog schema, the sync layer, or the render path.
|
||||
- No light theme. The ground stays dark (see `theme.slint` preamble).
|
||||
|
||||
---
|
||||
|
||||
## Workstream S — The style layer
|
||||
|
||||
Prerequisite for everything else that touches colour. Lands before B–F.
|
||||
|
||||
### S1 — Near-neutral palette
|
||||
|
||||
**Problem.** The original palette was a warm "darkroom safelight" brown —
|
||||
ground `#14120F`, R twelve points above B, and the same cast through every
|
||||
surface and ink. That biases the work. Simultaneous contrast pushes perception
|
||||
of the image *away* from its surround, so warm chrome makes a neutral
|
||||
photograph read cool; the photographer corrects toward warm to compensate and
|
||||
every export drifts yellow. The file's own preamble had the right instinct —
|
||||
a light UI biases judgement — and stopped one step short. Warmth biases it
|
||||
too, and more quietly, because a warm cast reads as *cosy* rather than
|
||||
*wrong*.
|
||||
|
||||
**Done.** `theme.slint` now carries near-neutral greys with a 2–3 point cool
|
||||
lift (pure R=G=B reads as dead; a trace of cool reads as instrument), and the
|
||||
red accent is replaced by an achromatic `active` family. `warn-ink` is the
|
||||
only hue left — a caution is genuinely a different kind of thing from an
|
||||
active state.
|
||||
|
||||
Migration shims alias `accent`/`accent-dim`/`accent-hover` onto the new
|
||||
tokens, so the tree compiles while call sites migrate. **They are temporary.**
|
||||
Delete them once `grep -rn 'Theme.accent' ui/` is empty.
|
||||
|
||||
### S2 — `style.yaml` as the token source
|
||||
|
||||
**Problem.** Tokens live in Slint, so tuning a palette means editing a
|
||||
language file, and nothing else — docs, tooling, a future export theme — can
|
||||
read them.
|
||||
|
||||
**Deliverable.** `ui/dr-ui/style.yaml` becomes the source of truth for
|
||||
**colours and lengths**: the whole current token set, nothing more.
|
||||
|
||||
- `ui/dr-ui/build.rs` reads it and generates `theme.slint` at compile time.
|
||||
Zero runtime cost, and a malformed file is a build error rather than a
|
||||
failure in front of a photographer mid-edit.
|
||||
- The generated file carries a "do not edit" banner naming its source, and
|
||||
must land somewhere `.gitignore`d or clearly marked generated — a
|
||||
hand-edited generated file is a bug that hides for weeks.
|
||||
- **The prose survives.** The reasoning in today's `theme.slint` preamble and
|
||||
its per-token comments is the most valuable thing in the file. YAML comments
|
||||
carry across into the generated Slint, or the generator emits them from
|
||||
structured fields. A token set with the *why* stripped out is a downgrade,
|
||||
however tidy the pipeline.
|
||||
- Debug-only live reload behind a feature flag: re-read the YAML at startup so
|
||||
a palette can be tuned without a full rebuild. Off in release, where the
|
||||
generated constants are what ship.
|
||||
|
||||
**Not in the YAML.** Semantic components (S3). `PanelHeading` binds colour,
|
||||
size, weight and letter-spacing into one concept — that is Slint, not data.
|
||||
YAML holds leaf values; components compose them.
|
||||
|
||||
**Dependency.** Adds a YAML parser to the UI's build-dependencies. Note that
|
||||
`serde_yaml` was deprecated in 2024; prefer a maintained alternative.
|
||||
|
||||
### S3 — Semantic components
|
||||
|
||||
**Problem.** `theme.slint` says what `surface` is. It says nothing about what
|
||||
a *panel heading* is — so every file re-derives one. `IMAGE`, `ADJUST`, the
|
||||
six launch-screen headings, `library.slint:267` each independently spell out
|
||||
colour, size, weight and letter-spacing for a single concept. That is why one
|
||||
accent reached forty call sites: there was no single place to change it.
|
||||
|
||||
**Deliverable.** `widgets.slint` grows from three primitives into the style
|
||||
layer:
|
||||
|
||||
- `PanelHeading` — the `IMAGE`/`ADJUST`/launch headings. One definition, one
|
||||
place to decide headings are ink rather than accent.
|
||||
- `Label`, `Value`, `Caption` — text roles, so `text-sm` + `ink-dim` stops
|
||||
being copy-pasted. `Value` carries the modified state, since a value that
|
||||
differs from its default is the one thing worth spotting at a glance.
|
||||
- `Panel` — surface + rule + padding, currently rebuilt in four places.
|
||||
- `Field` — the launch screen's text input with its focus border.
|
||||
|
||||
**The rule this establishes.** Files consume components. Raw `Theme.*` is for
|
||||
*composing* a component, not for styling a call site. A new colour literal or
|
||||
a bare `Theme.ink-faint` in a screen file is a signal that a component is
|
||||
missing.
|
||||
|
||||
**Sweep.** Migrating call sites onto these components removes the `accent`
|
||||
references as a side effect — the forty sites collapse into a handful of
|
||||
component definitions. That is the point: fix the cause, not the symptom.
|
||||
Delete the S1 shims when the grep comes back empty.
|
||||
|
||||
**Done when.** `grep -rn 'Theme.accent' ui/` is empty, the shims are gone, no
|
||||
screen file styles a heading inline, and the UI carries no hue outside
|
||||
`warn-ink`.
|
||||
|
||||
---
|
||||
|
||||
## Workstream P — Plural presentation hints
|
||||
|
||||
Implements ARCH §4.3a. Touches `core/dr-pipeline` and `ui/dr-ui/src/develop.rs`;
|
||||
no Slint change beyond what falls out of it.
|
||||
|
||||
**Problem.** `Presentation.widget` is a single `WidgetKind`. An operation can
|
||||
name one preferred control and nothing else, so a curve cannot say "a curve
|
||||
editor is best, a parametric band control would do, and sliders are fine" and
|
||||
let the frontend choose. Since layout class already drives compact-vs-expanded,
|
||||
this is not hypothetical: a curve editor is comfortable in a 280px panel and
|
||||
unusable in a 120px one, and today the frontend has no sanctioned way to
|
||||
decide that — its only options are the named widget or nothing.
|
||||
|
||||
**Deliverable.**
|
||||
|
||||
- `Presentation.widget` becomes `widgets: &'static [WidgetKind]`, in descending
|
||||
preference. The frontend takes the first it implements and can afford.
|
||||
- A `WidgetDemand` describing what a widget inherently requires — candidates:
|
||||
two-dimensional direct manipulation, precision pointing, a minimum count of
|
||||
simultaneous values. **No pixels, no breakpoints, no DPI, no platform names.**
|
||||
Those are frontend thresholds and live in `dr-ui`.
|
||||
- `develop.rs` chooses per operation: walk the hint list, take the first whose
|
||||
demands the current layout satisfies, else fall through to plain scalars.
|
||||
The existing `ParamRow.kind` string is where exhaustive matching is currently
|
||||
lost — the choice must happen in Rust against the enum, before flattening.
|
||||
- The tone curve declares `[Curve]` with its real demands. Nothing else needs
|
||||
to change; operations wanting sliders still say nothing at all.
|
||||
|
||||
**Verification.** The existing `the_curve_collapses_to_a_single_row` test in
|
||||
`develop.rs` covers the happy path. Add its complement: with a layout that
|
||||
cannot satisfy the curve's demands, the same capability must produce ten
|
||||
addressable scalar rows and remain fully editable. That test is the contract —
|
||||
it is what makes the fallback real rather than aspirational.
|
||||
|
||||
**Do not** let a demand grow a `min_width`. If one seems necessary, the
|
||||
demand vocabulary is wrong; widen the vocabulary, not the abstraction.
|
||||
|
||||
---
|
||||
|
||||
## Workstream M — The colour mixer is unreadable
|
||||
|
||||
Found in use, not in review. The panel currently renders the mixer as
|
||||
thirty-six anonymous sliders reading `Hue 0 / Sat 16 / Lum 0` twelve times
|
||||
over, with nothing saying which band any row belongs to. Three separate
|
||||
defects meet here.
|
||||
|
||||
### M1 — Band identity is lost (a bug, not a style issue) — **done**
|
||||
|
||||
`labels.rs` has no `param.mixer.*` entries, so all thirty-six keys fall
|
||||
through to a default that yields the bare channel name. The core is not at
|
||||
fault: it declares `param.mixer.orange.sat`, and `BANDS` carries `key:
|
||||
"orange"` with `hue: 30.0`. The identity is present in the capability and
|
||||
discarded at resolution.
|
||||
|
||||
**Deliverable.** Resolve mixer keys to their band. A row reads `Orange · Sat`,
|
||||
or `Sat` under a band heading — M3 decides which.
|
||||
|
||||
**Landed** as a `band.*` catalogue resolving the *subject* of a faceted
|
||||
parameter (M2), rather than as entries for the thirty-six `param.mixer.*`
|
||||
keys. Three bands are catalogued to something other than their key: chartreuse
|
||||
reads "Yellow-Green" and spring "Blue-Green", because a photographer looking
|
||||
for foliage does not scan a list for "Spring".
|
||||
|
||||
### M2 — Bands carry their centre hue as capability data — **done**
|
||||
|
||||
**Deliverable.** `ParamDescriptor` gains an optional band hue in degrees, set
|
||||
by `band_params!` from `BANDS`. The UI converts degrees to a swatch.
|
||||
|
||||
**Landed** as `descriptor::Facet` — a little wider than "a band hue", and the
|
||||
width is what M3 turned out to need. A parameter may say which **aspect** it
|
||||
adjusts (the channel) and which **subject** it adjusts it on (the band), with
|
||||
the subject's hue attached where the subject is a colour. `ParamDescriptor` is
|
||||
otherwise unchanged and `faceted()` is a const builder step, so the thirty-six
|
||||
descriptors stay `static` and every other operation says nothing at all.
|
||||
|
||||
The hue reaches the screen as `Swatch` in `widgets.slint`, which owns the
|
||||
saturation and brightness. Those two are *not* in `style.yaml`: it holds
|
||||
colours and lengths, and a third section for two floats used in one component
|
||||
buys less than it costs. `Theme.swatch` — the square's size — is a length and
|
||||
does live there.
|
||||
|
||||
**Why this is not a §4.3a violation.** A band's centre hue is a *fact about
|
||||
the operation* — the mixer genuinely acts on the 30° band, and that number is
|
||||
what it acts on. The core says "this parameter belongs to the band centred at
|
||||
30°". It does not say what colour to draw, at what saturation or lightness, or
|
||||
whether to draw a swatch at all. Those conversions are presentation and live
|
||||
in `dr-ui`; the swatch's saturation and lightness belong in `style.yaml`.
|
||||
|
||||
The line to hold: a hue in degrees is data. A hex colour in a descriptor
|
||||
would be the core deciding appearance, and is forbidden.
|
||||
|
||||
**Swatches are the one sanctioned exception to the achromatic palette.** A
|
||||
swatch is not chrome — it is data identifying which hue band a row edits,
|
||||
exactly as an image is data. That is categorically different from an accent
|
||||
decorating a heading, which is what the palette rule forbids. Keep them small
|
||||
and let them identify, never dominate.
|
||||
|
||||
### M3 — Structure: three runs of twelve, not twelve of three — **done**
|
||||
|
||||
~~The mixer renders as twelve `Section`s — one per band, titled by band name,
|
||||
carrying its swatch — each holding Hue, Sat and Lum. Collapsed by default.~~
|
||||
|
||||
**Superseded, twice over.** The lids came off the develop column entirely
|
||||
(Workstream C's sections are now plain headings), so "collapsed by default"
|
||||
had nothing left to mean. And the grouping was the wrong way round: an edit is
|
||||
almost never "everything about orange", it is "the saturation of the greens",
|
||||
made by comparing one channel across neighbouring bands. Twelve band sections
|
||||
put the twelve rows you want to compare in twelve different places.
|
||||
|
||||
**Landed.** Three runs — Hue, Saturation, Luminance — of twelve rows each,
|
||||
under the operation's own heading. `develop.rs` stacks the rows by aspect
|
||||
(`presentation_order`) and marks the first of each run; `adjust.slint` names
|
||||
the run once and draws the rest. Each row is a swatch, a track and a readout on
|
||||
one line: the swatch *is* the label, which is what makes twelve rows fit where
|
||||
four did, and the band name lives on as the row's accessible label so the
|
||||
control is not colour-only.
|
||||
|
||||
The mixer's declaration order is untouched — it declares band by band, which
|
||||
is the order the shader wants. Rearranging it for the panel would have been
|
||||
the core laying out a screen (§4.3a); doing it in `develop.rs` is the same
|
||||
frontend-side derivation that decides there are groups at all.
|
||||
|
||||
Nothing here is mixer-specific: any operation whose parameters carry facets
|
||||
groups this way, and one that carries none is untouched.
|
||||
|
||||
### M4 — A single-parameter operation should not cost a heading
|
||||
|
||||
Vibrance and Saturation each render a section heading above one slider,
|
||||
spending two lines and a visual break on one control. An operation whose
|
||||
parameters number one wants to *be* a named row, not a group containing one.
|
||||
|
||||
**Deliverable.** The panel collapses a single-parameter operation into one
|
||||
row labelled by the operation. Derived frontend-side from the parameter
|
||||
count — the core says nothing about it, per §4.3a.
|
||||
|
||||
**Done when.** Every mixer row says which band it edits; the rows for one
|
||||
channel read as a run rather than as twelve unrelated sliders; Vibrance and
|
||||
Saturation are one row each; and no hex colour appears in any descriptor.
|
||||
**All four are met.**
|
||||
|
||||
---
|
||||
|
||||
## Workstream V — Vertical density
|
||||
|
||||
The panel spends too much height on too little information. Reported from
|
||||
use, and it compounds Workstream M: at thirty-six mixer rows the waste is
|
||||
measured in whole screens.
|
||||
|
||||
**The arithmetic.** `ParamSlider` is a fixed 46px carrying an 11px label and a
|
||||
3px track. The track region is `Theme.touch-target / 2` — 22px — which is a
|
||||
finger-sized allowance drawn for a pointer, and the label occupies a line of
|
||||
its own above it. Every section heading adds a further `Theme.gap` (12px)
|
||||
spacer plus a 2px rule. Twelve mixer bands at three rows each is roughly
|
||||
1650px of panel, most of it air.
|
||||
|
||||
**The cause is layout, not spacing tokens.** Shaving pixels uniformly would
|
||||
compress the readable parts along with the waste. The row is stacked when it
|
||||
could be inline: name left, value right, track beneath — which is what
|
||||
Lightroom does, in about 32px.
|
||||
|
||||
**Deliverable.**
|
||||
|
||||
- Rework `ParamSlider` so label and value share one line and the track sits
|
||||
under them. Target ~32px per row, down from 46px.
|
||||
- The *drawn* track shrinks; the **TouchArea does not**. FR-UI-3 is about the
|
||||
finger, and `Button` already establishes the pattern — draw at control
|
||||
height, grow the hit target past the ink and centre it. A denser panel must
|
||||
not become a less touchable one.
|
||||
- Section heading spacing comes from one token, not an inline `Rectangle {
|
||||
height: Theme.gap }`. A spacer rectangle written inline is how the panel
|
||||
ended up with spacing nobody can adjust centrally.
|
||||
- Re-check the compact layout class after the change: rows that work at 280px
|
||||
may crowd at narrower widths, where the touch overhang also matters most.
|
||||
|
||||
**Constraint.** Density is not the goal; *legibility per pixel* is. If a row
|
||||
gets shorter and harder to read, it has failed. The value readout in
|
||||
particular carries the modified signal and must stay scannable.
|
||||
|
||||
**Done when.** A parameter row is ~32px, hit targets still meet FR-UI-3 under
|
||||
touch, heading spacing is tokenised, and the panel reads as easily at the new
|
||||
density as the old.
|
||||
|
||||
---
|
||||
|
||||
## Workstream A — Shared chrome primitives
|
||||
|
||||
**Problem.** Buttons are hand-rolled `Rectangle` + `TouchArea` pairs in
|
||||
`app.slint`, `library.slint`, and `launch.slint`, at three different sizes
|
||||
(64×20, 110×28, 88×28) with three near-identical hover/press treatments. Any
|
||||
consistency in the chrome is currently coincidental.
|
||||
|
||||
**Deliverable.** A new `ui/dr-ui/ui/widgets.slint` exporting:
|
||||
|
||||
- `Button` — `text`, `enabled`, `primary` (bool), `clicked()`. Height meets
|
||||
`Theme.touch-target` under compact layout and may be denser when expanded.
|
||||
Press and hover states derive from theme tokens, not literals.
|
||||
- `IconButton` — square, for toolbar affordances that carry a glyph.
|
||||
- `Section` — a collapsible container: `title`, `modified` (bool),
|
||||
`expanded` (in-out bool), a default child slot. Draws the disclosure
|
||||
triangle and the modified dot. Workstream C consumes this.
|
||||
|
||||
Every existing hand-rolled button is replaced by `Button`. The visual result
|
||||
should be a *narrower* range of sizes than today, not a wider one.
|
||||
|
||||
**Theme additions.** `theme.slint` gains what the widgets need and no more:
|
||||
`radius-sm`/`radius`, a `hover` and `pressed` surface token, and a
|
||||
`modified` token aliased to `accent`. Adding tokens is preferred over
|
||||
literals appearing in widget bodies.
|
||||
|
||||
**Done when.** No `TouchArea` inside a `Rectangle` styled as a button remains
|
||||
in `app.slint`, `library.slint`, or `library.slint`'s header. `cargo build`
|
||||
clean, app launches, every button still fires its callback.
|
||||
|
||||
---
|
||||
|
||||
## Workstream B — Canvas presentation
|
||||
|
||||
**Problem.** The canvas fills its container edge to edge. The image reads as a
|
||||
texture rather than a print, and there is no visual separation between the
|
||||
photograph and the panel beside it.
|
||||
|
||||
**Deliverable.** In `app.slint`'s `canvas-area`:
|
||||
|
||||
- Inset the image by a margin that scales with the layout class — generous
|
||||
when `expanded`, tighter when compact, never zero. The surrounding field is
|
||||
`Theme.ground`.
|
||||
- A subtle 1px `Theme.rule` border on the image bounds, so a dark photograph
|
||||
does not bleed into the dark ground. This requires knowing the *fitted*
|
||||
rectangle, not the container — if that proves awkward in Slint, a shadow or
|
||||
a very slightly lighter mat behind the image is an acceptable substitute.
|
||||
Pick one and say which in the summary.
|
||||
- Empty and error states keep their current copy and centring.
|
||||
|
||||
**Constraint.** `canvas-resized` must continue to report the *drawable* pixel
|
||||
size — the render target follows the image area, not the container. Getting
|
||||
this wrong shows up as a soft or stretched image, so verify the reported size
|
||||
changes when the margin does.
|
||||
|
||||
**Done when.** The image sits in a visible field with margin, the border or
|
||||
mat is present, and resizing the window still produces a crisp canvas.
|
||||
|
||||
---
|
||||
|
||||
## Workstream C — Collapsible adjust sections
|
||||
|
||||
**Problem.** `AdjustPanel` renders every parameter of every operation, always
|
||||
expanded. This is tolerable at today's operation count and unusable at
|
||||
fifteen. Lightroom's right panel is a stack of collapsible modules whose
|
||||
headers report whether anything inside has been touched.
|
||||
|
||||
**Deliverable.** Rework the `for row[i] in root.rows` body in `adjust.slint`
|
||||
to group by operation and wrap each group in `Section` (Workstream A).
|
||||
|
||||
The hard part is that `rows` is a **flat** model with a `starts-group` flag —
|
||||
Slint cannot easily nest a `for` inside a group boundary derived at runtime.
|
||||
Two viable approaches; pick one and justify it briefly:
|
||||
|
||||
1. **Flatten the collapse.** Keep the flat `for`, add an `expanded` bool per
|
||||
op-index held in the panel, and make each non-heading row `visible: false`
|
||||
and zero-height when its group is collapsed. Simple, no Rust change.
|
||||
2. **Nest the model.** Have Rust supply `[[ParamRow]]` — one inner model per
|
||||
operation. Cleaner Slint, but changes the `ParamRow` contract and the
|
||||
`develop.rs` code that builds it.
|
||||
|
||||
Approach 1 is likely correct for this pass; prefer it unless it proves
|
||||
unworkable.
|
||||
|
||||
**Modified indicator.** A group is modified when any row in it has
|
||||
`value != default-value`.
|
||||
|
||||
**This must be derived in `dr-ui`, not supplied by the core** (ARCH §4.3a).
|
||||
A `group_modified` flag on a descriptor would be the core deciding the panel
|
||||
has groups at all, which is a composition decision. `develop.rs` already holds
|
||||
both the capabilities and the live values, so it can aggregate per operation
|
||||
while flattening — that is frontend-side derivation and stays on the right
|
||||
side of the line. What it must not do is ask the core for the answer.
|
||||
|
||||
The same reasoning condemns the existing `starts-group` flag, which is the
|
||||
core telling the panel where to draw section breaks. It predates this contract;
|
||||
fold it into the same pass and derive grouping from `op-index` changes instead.
|
||||
|
||||
**Also.** The per-group `reset` should live on the section header, alongside
|
||||
the existing global `reset`.
|
||||
|
||||
**Invariant.** This file must still name no operation. Collapse state is keyed
|
||||
by `op-index`, never by label.
|
||||
|
||||
**Done when.** Sections collapse and expand, collapsed state survives a slider
|
||||
drag elsewhere in the panel, headers show a modified dot that appears and
|
||||
disappears as values move off and back to default, and adding an operation to
|
||||
the pipeline still requires no edit to `adjust.slint`.
|
||||
|
||||
---
|
||||
|
||||
## Workstream D — Grid refinement
|
||||
|
||||
**Problem.** Cells are boxes first and images second: `image-fit: contain` on
|
||||
a square cell leaves landscape shots floating in dead space, the cell surface
|
||||
contrasts with the ground so the grid reads as a rhythm of rectangles, and
|
||||
there is hover state but no *selection* state.
|
||||
|
||||
**Deliverable.** In `library.slint`:
|
||||
|
||||
- Cells crop to fill (`image-fit: cover`) with `clip: true`, so the grid is a
|
||||
rhythm of images. The filename caption stays.
|
||||
- Cell background moves to `Theme.ground` or very near it; the frame recedes.
|
||||
- A **selected** cell gets a persistent accent ring. Add
|
||||
`in property <int> selected-index` to `LibraryGrid`, defaulting to -1, and
|
||||
have `library_ui.rs` set it when a cell is clicked. Hover stays distinct
|
||||
from selection — a dimmer treatment.
|
||||
- Keyboard navigation: arrow keys move the selection, Enter opens it. This
|
||||
needs a `FocusScope` over the grid and a `selection-moved(int)` callback.
|
||||
|
||||
**Done when.** The grid reads as images rather than boxes, the current image
|
||||
is unambiguous, and arrows plus Enter navigate it without the mouse.
|
||||
|
||||
### Keyboard navigation — **done**
|
||||
|
||||
Arrows walk the grid, shift+arrow extends the selection, Home/End reach the
|
||||
ends, PageUp/PageDown move by a screenful, and `Return` opens what the cursor
|
||||
is on. With the judgement keys the grid already bound, a culling pass is now a
|
||||
keyboard job end to end — which is the point: a cull is thousands of decisions,
|
||||
and reaching for the mouse between each is the difference between an hour and
|
||||
an evening.
|
||||
|
||||
**The cursor is a library ordinal, not a row of the loaded window** — that is
|
||||
the whole design, and it is the same argument the selection already made by
|
||||
keying on ids. The window is a few screenfuls around wherever the user is
|
||||
looking; a cursor held as a row would stop at the window's edge or, worse, keep
|
||||
counting into cells that belong to other photographs. Walking out of the window
|
||||
reloads it around the new position, exactly as scrolling does.
|
||||
|
||||
The same fault was already live in the **anchor**, which was a window row: a
|
||||
shift-click after a scroll extended from whatever image had drifted into that
|
||||
row. It is now an ordinal too, and `apply_press` takes the window's offset to
|
||||
map between the two. A range longer than the loaded window truncates to what is
|
||||
loaded — selection is by id, and an image the catalog has not been asked for
|
||||
has no id to select — which is the honest failure, not the silent one.
|
||||
|
||||
`select_row` is shared by the pointer and the keyboard so the two cannot drift:
|
||||
"click here, shift+down twice" has to mean what "click here, shift-click there"
|
||||
means. The one deliberate difference is that a plain arrow collapses the
|
||||
selection onto the cursor, where a plain *click* on an already-selected cell
|
||||
leaves it alone — that exception exists so a multi-image drag can start from
|
||||
one of its members, and there is no drag behind a keystroke.
|
||||
|
||||
**Not done here:** a cursor marker distinct from the selection ring. A plain
|
||||
arrow selects what it lands on, so the ring shows it; only during a shift
|
||||
extension is the moving end indistinguishable from the rest of the range.
|
||||
That wants a `cursor` flag on `LibraryCell` and a second ring treatment.
|
||||
|
||||
**Still open in D:** cells cropping to fill, the cell background receding to
|
||||
the ground, and hover reading distinctly from selection.
|
||||
|
||||
---
|
||||
|
||||
## Workstream E — Chrome hierarchy
|
||||
|
||||
**Problem.** `StatusBar` mixes three unrelated things: navigation (`‹
|
||||
Library`), identity (filename, position), and spike telemetry (fps, adapter,
|
||||
backend, layout-class). The telemetry earned its place while assumption A1 was
|
||||
open; it is now permanent furniture competing with the photograph.
|
||||
|
||||
**Deliverable.**
|
||||
|
||||
- The top strip carries identity and navigation only: filename, position,
|
||||
and the library affordance.
|
||||
- Telemetry moves behind a toggle — a keyboard shortcut (suggest `` ` ``) that
|
||||
reveals a small diagnostics overlay in a corner of the canvas, carrying
|
||||
backend, adapter, fps, and layout class. Default off.
|
||||
- The accent stops being used for chrome. `backend` in the status bar, the
|
||||
`IMAGE` and `ADJUST` panel headings, and the section headings in
|
||||
`adjust.slint` all move to `Theme.ink-faint` or `ink-dim`. After this pass,
|
||||
accent should appear only on: modified values, the modified dot, the curve
|
||||
line, slider fill, selection, and progress.
|
||||
|
||||
**Done when.** A fresh launch shows no fps counter and no accent-coloured
|
||||
chrome; `` ` `` toggles the diagnostics overlay; every previously-visible
|
||||
diagnostic is still reachable.
|
||||
|
||||
---
|
||||
|
||||
## Workstream F — Filmstrip and unified view
|
||||
|
||||
**The big one.** Depends on A (for `Button`) and D (for cell treatment and
|
||||
selection). Should land last.
|
||||
|
||||
**Problem.** Library and Develop are mutually exclusive screens
|
||||
(`show-library` in `app.slint`). Every move between images is a change of
|
||||
screen. Lightroom's continuity comes from the grid never fully leaving: it
|
||||
collapses to a filmstrip along the bottom of the develop view, and clicking a
|
||||
neighbour is navigation, not a mode change.
|
||||
|
||||
**Deliverable.**
|
||||
|
||||
- A `Filmstrip` component in a new `ui/dr-ui/ui/filmstrip.slint`, consuming
|
||||
the **same** `[LibraryCell]` model and the same `selected-index` as
|
||||
`LibraryGrid`. Horizontal, ~90px tall, scrolls to keep the selection
|
||||
visible.
|
||||
- Develop gains the filmstrip along its bottom edge, visible when a library is
|
||||
open (i.e. when the current model is non-empty). Command-line file sets get
|
||||
it too — they are also a list of images.
|
||||
- Clicking a filmstrip cell loads that image. Arrow keys drive both the
|
||||
filmstrip and the existing next/prev, which become the same action.
|
||||
- `show-library` becomes a *mode* rather than a screen swap: Grid mode and
|
||||
Develop mode over one shared library state, toggled by a `G`/`D` shortcut
|
||||
and by the existing buttons. The `can-return-to-library` special case and
|
||||
the `‹ Library` button both disappear.
|
||||
|
||||
**Windowing constraint.** The filmstrip and the grid must share one windowed
|
||||
model, not hold two. `LibraryController` currently keys its `WINDOW` on grid
|
||||
scroll position; the filmstrip's window follows the *selection* instead. This
|
||||
is the genuinely hard part of the workstream — the window must move as
|
||||
selection walks past its edge, and thumbnail requests must not thrash when it
|
||||
does. Resolve this explicitly rather than by widening `WINDOW`.
|
||||
|
||||
**Done when.** Selecting an image in the grid enters develop with the
|
||||
filmstrip showing neighbours; arrows walk the filmstrip and load images;
|
||||
`G`/`D` toggles modes with selection preserved in both directions; a
|
||||
17k-image library still holds a bounded number of live cells and does not
|
||||
re-fetch thumbnails on every keystroke.
|
||||
|
||||
---
|
||||
|
||||
## Sequencing
|
||||
|
||||
```
|
||||
A (primitives) ──┬── C (sections)
|
||||
├── B (canvas) ── E (chrome)
|
||||
└── D (grid) ──┐
|
||||
├── F (filmstrip)
|
||||
┘
|
||||
```
|
||||
|
||||
A, B, D, and E are independent of each other once A lands; C depends on A;
|
||||
F depends on A and D. B and E both touch `app.slint`, so they should not run
|
||||
concurrently.
|
||||
|
||||
## Verification, all workstreams
|
||||
|
||||
- `cargo build` clean, no new Slint warnings — in particular no binding-loop
|
||||
warnings, which `app.slint` already comments on at length and which can
|
||||
panic at runtime.
|
||||
- The app launches and reaches the grid.
|
||||
- No workstream may break FR-DEV-3a: adding a pipeline operation must still
|
||||
surface in the panel with no UI edit.
|
||||
@@ -0,0 +1,113 @@
|
||||
{
|
||||
"_readme": [
|
||||
"The committed numbers for DarkRoom's benchmark suite (docs/requirements.md §8).",
|
||||
"Produced and checked by `cargo run --release -p dr-bench`; docs/benchmarks.md explains each metric.",
|
||||
"",
|
||||
"Two gates, and they are not the same gate. `budget` is the requirement's own threshold and never moves.",
|
||||
"`recorded` is what the reference desktop last measured, and a run that drifts past `tolerance` beyond it",
|
||||
"fails the build even while still inside the budget — which is how most performance rot actually arrives.",
|
||||
"",
|
||||
"`recorded` is null on every metric because nobody has run the suite yet. That is deliberate: writing",
|
||||
"plausible-looking figures here would make every later comparison a comparison against a guess. Run",
|
||||
"`cargo run --release -p dr-bench -- record --reference` on the reference desktop and commit the diff.",
|
||||
"Until then the budget gate works and the regression gate says, in the report, that it cannot.",
|
||||
"",
|
||||
"`machine_sensitive` says whether a budget is a statement about a machine as much as about the code.",
|
||||
"Those budgets are asserted only under --reference: §8 names the reference desktop, and a two-core CI",
|
||||
"container cannot speak to a target written for twenty-four threads. Asserting one there would produce a",
|
||||
"red gate everybody learns to ignore, which is the trap core/dr-gpu/tests/frame_budget.rs already avoids.",
|
||||
"",
|
||||
"Several metrics carry a requirement ID with a qualifier. Read those literally. NFR-P7's budget here is",
|
||||
"checked against the encode half of an export only — no GPU render is in the figure — so it can fail the",
|
||||
"requirement and cannot pass it, and no TRACES tag claims otherwise. NFR-P8 has no budget at all yet,",
|
||||
"because nobody has decided how much of its 500 MB belongs to the catalog layer; this records the number",
|
||||
"that decision needs."
|
||||
],
|
||||
"tolerance": 0.15,
|
||||
"recorded_on": null,
|
||||
"recorded_at_unix": null,
|
||||
"fixture": null,
|
||||
"metrics": {
|
||||
"catalog_filtered_ms": {
|
||||
"requirement": "FR-CAT-6",
|
||||
"what": "Count plus first window under a rating filter, which compiles to a correlated subquery over versions.",
|
||||
"unit": "ms",
|
||||
"direction": "lower_is_better",
|
||||
"machine_sensitive": true,
|
||||
"budget": null,
|
||||
"recorded": null
|
||||
},
|
||||
"catalog_idle_rss_mb": {
|
||||
"requirement": "NFR-P8 (the catalog layer's share only — no toolkit, no adapter, no decode cache)",
|
||||
"what": "Resident memory of a process that opened the 50k catalog and scrolled ten thousand rows.",
|
||||
"unit": "MB",
|
||||
"direction": "lower_is_better",
|
||||
"machine_sensitive": false,
|
||||
"budget": null,
|
||||
"recorded": null
|
||||
},
|
||||
"catalog_open_ms": {
|
||||
"requirement": "NFR-P1",
|
||||
"what": "Catalog::open plus the count, first window and timeline the grid cannot paint without.",
|
||||
"unit": "ms",
|
||||
"direction": "lower_is_better",
|
||||
"machine_sensitive": false,
|
||||
"budget": 2000.0,
|
||||
"recorded": null
|
||||
},
|
||||
"catalog_open_warm_ms": {
|
||||
"requirement": "NFR-P1",
|
||||
"what": "The same four calls on a second connection, with SQLite's page cache already warm.",
|
||||
"unit": "ms",
|
||||
"direction": "lower_is_better",
|
||||
"machine_sensitive": false,
|
||||
"budget": 2000.0,
|
||||
"recorded": null
|
||||
},
|
||||
"catalog_window_p99_ms": {
|
||||
"requirement": "FR-CAT-4",
|
||||
"what": "One 400-row grid window at a random offset, p99 of one hundred.",
|
||||
"unit": "ms",
|
||||
"direction": "lower_is_better",
|
||||
"machine_sensitive": true,
|
||||
"budget": null,
|
||||
"recorded": null
|
||||
},
|
||||
"export_24mp_long_edge_2048_ms": {
|
||||
"requirement": "FR-EXP-3",
|
||||
"what": "The web export: resample a 24 MP frame to a 2048 px long edge, sharpen, encode. p99 of five.",
|
||||
"unit": "ms",
|
||||
"direction": "lower_is_better",
|
||||
"machine_sensitive": true,
|
||||
"budget": null,
|
||||
"recorded": null
|
||||
},
|
||||
"export_24mp_original_ms": {
|
||||
"requirement": "NFR-P7 (the encode half only — the GPU render is not in this figure)",
|
||||
"what": "Resample, output-sharpen and JPEG-encode a 24 MP frame at source size. p99 of five.",
|
||||
"unit": "ms",
|
||||
"direction": "lower_is_better",
|
||||
"machine_sensitive": true,
|
||||
"budget": 2000.0,
|
||||
"recorded": null
|
||||
},
|
||||
"thumbnail_per_image_p99_ms": {
|
||||
"requirement": "NFR-P3",
|
||||
"what": "One thumbnail on its own lane: decode the preview, downscale, orient, encode. p99.",
|
||||
"unit": "ms",
|
||||
"direction": "lower_is_better",
|
||||
"machine_sensitive": true,
|
||||
"budget": null,
|
||||
"recorded": null
|
||||
},
|
||||
"thumbnail_throughput_ips": {
|
||||
"requirement": "NFR-P3",
|
||||
"what": "Whole-sweep throughput: 1200 thumbnails through the sweep's chunk-and-lane shape, wall clock.",
|
||||
"unit": "img/s",
|
||||
"direction": "higher_is_better",
|
||||
"machine_sensitive": true,
|
||||
"budget": 100.0,
|
||||
"recorded": null
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,243 @@
|
||||
# The benchmark suite
|
||||
|
||||
**Status:** Built, not yet recorded · 2026-08-30
|
||||
**Companion to:** [requirements.md](requirements.md) §4.1 (performance targets) · §8 (verification)
|
||||
**Instrument:** [`tools/bench`](../../tools/bench) — `cargo run --release -p dr-bench -- check`
|
||||
**Committed numbers:** [`bench-baseline.json`](bench-baseline.json)
|
||||
**GPU half:** [`core/dr-gpu/tests/frame_budget.rs`](../../core/dr-gpu/tests/frame_budget.rs) ·
|
||||
[frame-budget.md](frame-budget.md)
|
||||
|
||||
§8 has said since it was written that performance is verified by *"an automated
|
||||
benchmark suite against a synthetic 50k catalog, run per-commit … A regression
|
||||
beyond stated tolerance fails the build."* Until this suite there was none. No
|
||||
`benches/`, no `[[bench]]`, no criterion, no fixture — and ten performance
|
||||
requirements that could therefore be neither passed nor failed, five of them
|
||||
carrying a `TRACES:` tag regardless.
|
||||
|
||||
This file is what the suite covers, what it deliberately does not, and how to
|
||||
read a failure.
|
||||
|
||||
---
|
||||
|
||||
## The state of it, first
|
||||
|
||||
**No numbers have been recorded yet.** Every `recorded` field in
|
||||
[`bench-baseline.json`](bench-baseline.json) is `null`, on purpose: writing
|
||||
plausible-looking figures into a baseline would make every later comparison a
|
||||
comparison against a guess, and the first real regression would be invisible.
|
||||
|
||||
To record them, on the reference desktop:
|
||||
|
||||
```sh
|
||||
cargo run --release -p dr-bench -- record --reference
|
||||
```
|
||||
|
||||
and commit the diff. Until that happens the **budget** gate works — a catalog
|
||||
that takes three seconds to open fails the build today — and the **regression**
|
||||
gate reports that it has nothing to compare against, rather than pretending.
|
||||
|
||||
---
|
||||
|
||||
## What it measures
|
||||
|
||||
| Metric | Requirement | Gated? |
|
||||
|---|---|---|
|
||||
| `catalog_open_ms` | **NFR-P1**, and R2's second sentence | Yes, everywhere — budget 2000 ms |
|
||||
| `catalog_open_warm_ms` | NFR-P1, page cache warm | Yes, everywhere — budget 2000 ms |
|
||||
| `catalog_window_p99_ms` | FR-CAT-4 | Regression only |
|
||||
| `catalog_filtered_ms` | FR-CAT-6 | Regression only |
|
||||
| `thumbnail_throughput_ips` | **NFR-P3** | Budget 100 img/s, on the reference desktop |
|
||||
| `thumbnail_per_image_p99_ms` | NFR-P3 | Regression only |
|
||||
| `export_24mp_original_ms` | NFR-P7, **encode half only** | One-sided: can fail it, cannot pass it |
|
||||
| `export_24mp_long_edge_2048_ms` | FR-EXP-3 | Regression only |
|
||||
| `catalog_idle_rss_mb` | NFR-P8, **catalog layer only** | Regression only — see below |
|
||||
|
||||
Two of those rows carry a qualifier, and the qualifiers are the point.
|
||||
|
||||
### Requirements this can now pass *or* fail
|
||||
|
||||
**NFR-P1 — catalog open under 2 s.** The measured span is the four things the
|
||||
library view cannot paint without: `Catalog::open` (which connects, migrates and
|
||||
**backfills**, and the backfill is three passes over the images table on every
|
||||
open), `count`, the first 400-row `window`, and the monthly `timeline`. Tagged
|
||||
`TRACES: NFR-P1` in [`tools/bench/src/catalog_open.rs`](../../tools/bench/src/catalog_open.rs),
|
||||
because a build that breaks it fails this gate.
|
||||
|
||||
**NFR-P3 — ≥ 100 images per second on the embedded preview path.** The
|
||||
per-image work is exactly what `spawn_thumbnail_sweep` does — `decode_jpeg`,
|
||||
`Preview::downscale_to`, `Preview::apply_orientation`, `encode_rgba`,
|
||||
`ThumbStore::put` — arranged in the same shape: chunks of 96, lanes owning
|
||||
disjoint slices, and the single thread that owns the store writing the finished
|
||||
chunk. Tagged `TRACES: NFR-P3` in
|
||||
[`tools/bench/src/thumbnails.rs`](../../tools/bench/src/thumbnails.rs).
|
||||
|
||||
### Requirements this can only half-answer, and is not tagged for
|
||||
|
||||
**NFR-P7 — 24 MP export under 2 s, full chain.** The full chain is decode,
|
||||
demosaic, a full-resolution GPU render, a read-back, then resize, sharpen and
|
||||
encode. Only the last three run without an adapter. So the figure here is a
|
||||
**lower bound** on the requirement: exceeding 2 s in the encode alone violates
|
||||
NFR-P7 no matter how fast the render is, and coming in under it proves nothing.
|
||||
The budget is gated on that basis and there is no `TRACES: NFR-P7` anywhere in
|
||||
`tools/bench`.
|
||||
|
||||
**NFR-P8 — idle memory under 500 MB.** The probe is a fresh process holding the
|
||||
catalog and nothing else: no Slint, no wgpu device, no font stack, no decode
|
||||
cache. Its RSS is the catalog layer's *share* of that 500 MB, not the figure the
|
||||
requirement is about. It carries no budget for a reason given below.
|
||||
|
||||
### Requirements out of scope, listed so their absence reads as a decision
|
||||
|
||||
NFR-P2 (grid scroll at 60 fps), P4 (open in develop), P5 (slider to visible),
|
||||
P6 (pan/zoom), P9 (UI-executor blocking), P10 (touch response), P11 (layout
|
||||
transition), P12 (warm shader setup), P13 (next image in culling), P14 (focus
|
||||
peaking), P15 (drawn mask stroke). Every one of them needs a frame-timing probe
|
||||
inside a running Slint application, a GPU adapter, or both. None is faked here.
|
||||
|
||||
The GPU half of the story that *does* exist is
|
||||
[frame-budget.md](frame-budget.md) and its guard test, which asserts FR-DSP-3
|
||||
and skips itself where there is no adapter. `.gitea/workflows/benchmark.yml`
|
||||
runs it as its own job for exactly that reason.
|
||||
|
||||
---
|
||||
|
||||
## The fixture
|
||||
|
||||
Fifty thousand rows over a pool of twelve real image files. Rows are cheap and
|
||||
pixels are not: everything the catalog half touches is rows and is therefore
|
||||
exact at full scale, and everything the pixel half touches is one file at a time
|
||||
and does not care how many rows point at it. The result is ~14 MB on disk
|
||||
instead of ~2 TB, and neither half is flattered by that.
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| Rows | 50,000 images, 50,000 default versions, 400 folders, one root |
|
||||
| Capture times | Twelve years from a fixed epoch, so the timeline has ~144 monthly buckets |
|
||||
| Sources | 12 synthesised JPEGs at 1620 × 1080 — the size `dr-decode` records a CR2 carrying in IFD2 |
|
||||
| Seed | 20260829, in [`tools/bench/src/main.rs`](../../tools/bench/src/main.rs) |
|
||||
| Location | `$DR_BENCH_DIR`, else the system temporary directory |
|
||||
|
||||
It is reproducible from the seed, and a `stamp.json` beside it records what it
|
||||
was built from — seed, row count, source count, preview size, and `dr-catalog`'s
|
||||
schema version. A mismatch rebuilds rather than silently measuring a different
|
||||
workload than the baseline describes.
|
||||
|
||||
Two honest limits on it:
|
||||
|
||||
- **The page cache is warm.** The fixture was written by this suite or by an
|
||||
earlier run of it, so neither the catalog open nor the thumbnail sweep pays
|
||||
for a cold disk. On the reference desktop's NVMe a genuinely cold read of a
|
||||
14 MB catalog is tens of milliseconds; on spinning rust it is not.
|
||||
- **The sources are synthetic.** A coarse gradient with a fine dither, which is
|
||||
what `frame_budget.rs` synthesises for the same reason — a flat frame lets the
|
||||
memory system serve every sample from one cache line and flatters a box
|
||||
filter, and pure noise defeats the entropy coder in the other direction.
|
||||
|
||||
---
|
||||
|
||||
## Two gates, and how to read a failure
|
||||
|
||||
**Budget.** The requirement's own threshold. It does not move. Failing it means
|
||||
a requirement is violated.
|
||||
|
||||
**Regression.** More than 15% worse than the last recorded figure *on the same
|
||||
machine, against the same fixture*. Failing it means the code got slower while
|
||||
still inside the requirement — which is how most performance rot actually
|
||||
arrives, never over the line, always a little worse, until one day the line is
|
||||
crossed by a change that was not the cause.
|
||||
|
||||
A metric declares whether its budget is `machine_sensitive`. Those are asserted
|
||||
only under `--reference`, and reported everywhere else. §8 names *"the reference
|
||||
desktop"*, not CI, and it is right to: a container with two cores cannot speak
|
||||
to a throughput target written for twenty-four threads, and asserting one there
|
||||
would produce exactly what `core/dr-gpu/tests/frame_budget.rs` refused to
|
||||
produce — *"a red suite that everyone learns to ignore"*. Catalog open is not
|
||||
machine-sensitive: 2 s against an expected figure two orders of magnitude
|
||||
smaller is a threshold any machine can be held to.
|
||||
|
||||
Exit codes: `0` everything passed, `1` a gate failed, `2` the harness itself
|
||||
could not run. Distinguished so a CI log that says "failed" does not leave
|
||||
anyone guessing whether the code got slower or the fixture would not build.
|
||||
|
||||
**Release, always.** The workspace builds its own crates at `opt-level = 0` in
|
||||
dev, and every figure here is dominated by this workspace's own code — the JPEG
|
||||
decode, the box filter, the resample, the sharpen. A debug run measures rustc's
|
||||
shadow. The report says which profile it was built in on its second line.
|
||||
|
||||
---
|
||||
|
||||
## NFR-P8, and the question §4.1 asks
|
||||
|
||||
§4.1 says NFR-P8 *"must state whether it measures RSS inclusive or exclusive of
|
||||
GPU allocations, and whether it holds after SQLite's page cache warms on a 50k
|
||||
catalog."* Both halves have an answer.
|
||||
|
||||
**On the page cache: warm.** The probe runs the count, the timeline and
|
||||
twenty-five windows before it reads its counters, so SQLite's cache holds the
|
||||
b-tree pages a scroll touches. That is the right side to err on — a figure taken
|
||||
before the cache warms would understate a steady-state library.
|
||||
|
||||
**On GPU memory: RSS is exclusive of device-local allocations, and cannot be
|
||||
made otherwise.** A Vulkan allocation in a device-local heap never enters the
|
||||
process's address space, so nothing under `/proc/self/status` can see it. What
|
||||
*does* land in RSS is the host-visible side — staging buffers, mapped upload
|
||||
rings, the read-back `AdjustPass` performs on export — plus the driver's own
|
||||
resident pages.
|
||||
|
||||
So "idle memory < 500 MB" is two questions wearing one number, and a build
|
||||
holding 400 MB of RSS and 3 GB of textures would pass it.
|
||||
|
||||
**Recommendation: NFR-P8 should be restated as two figures** — host RSS
|
||||
exclusive of device-local memory, and a separate VRAM ceiling read from the
|
||||
adapter — because the second is the one that decides whether the application
|
||||
survives beside a browser on an 8 GB card, and nothing in this repository
|
||||
measures it today.
|
||||
|
||||
**And a decision is outstanding.** `catalog_idle_rss_mb` carries no budget
|
||||
because nobody has decided how much of the 500 MB belongs to the catalog layer
|
||||
and how much to everything above it. The suite records the number so that
|
||||
decision can be taken against a measurement rather than an estimate. When it is
|
||||
taken, put the figure in `budget` and the metric becomes a gate.
|
||||
|
||||
---
|
||||
|
||||
## What is not measured, and would be worth adding
|
||||
|
||||
- **The UI's own open.** `ui/dr-ui/src/library.rs` does not call
|
||||
`Catalog::count` or `Catalog::window`; it issues its own SQL against the same
|
||||
tables, with a `VISIBLE` predicate and a burst-folding clause. `dr-bench`
|
||||
cannot see those without depending on `dr-ui`, which would drag Slint into a
|
||||
job that has no display. **Falsifiable end:** when the grid's queries move
|
||||
down into `dr-catalog` — which is where SQL over catalog tables belongs —
|
||||
`catalog_open_ms` becomes the whole of the application's open and this caveat
|
||||
can be deleted rather than argued about.
|
||||
- **The remote sweep.** `spawn_thumbnail_sweep`'s wall clock against a real
|
||||
server is latency, not CPU, and is what FR-NC-3's design is judged by. It
|
||||
needs a server and belongs in a different kind of test.
|
||||
- **A cold disk.** See the fixture's limits above.
|
||||
- **Android.** §4.1 states a second column of targets and §8 asks for
|
||||
"periodically on the named reference Android devices". Nothing here runs on a
|
||||
device. Spike S10 is the piece of work that would start it.
|
||||
- **Everything with a frame in it.** See the out-of-scope list above.
|
||||
|
||||
---
|
||||
|
||||
## Running it
|
||||
|
||||
```sh
|
||||
# Measure and print. Judges nothing.
|
||||
cargo run --release -p dr-bench -- run
|
||||
|
||||
# Measure and gate. What CI runs.
|
||||
cargo run --release -p dr-bench -- check
|
||||
|
||||
# The same, with machine-sensitive budgets asserted too.
|
||||
cargo run --release -p dr-bench -- check --reference
|
||||
|
||||
# Rewrite bench-baseline.json from this run, and commit the diff.
|
||||
cargo run --release -p dr-bench -- record --reference
|
||||
```
|
||||
|
||||
Useful flags: `--fixture <dir>` (or `$DR_BENCH_DIR`) to put the synthetic
|
||||
catalog somewhere specific, `--lanes <n>` to pin the sweep's parallelism, and
|
||||
`--thumbnails <n>` to lengthen or shorten the throughput row.
|
||||
@@ -0,0 +1,921 @@
|
||||
# DarkRoom — Catalog, library view, and background work
|
||||
|
||||
**Status:** Draft v0.1 · 2026-08-09
|
||||
**Companion to:** [requirements.md](requirements.md), [architecture.md](architecture.md)
|
||||
|
||||
Specifies `dr-catalog`: the index the library view queries, how it stays current without rescanning
|
||||
everything, and how thumbnails get made. [architecture.md §6.2](architecture.md) sketches the schema
|
||||
in eight lines; this expands it to the point of implementability and fills the two gaps that sketch
|
||||
leaves open — **incremental local scan** and **the job queue**.
|
||||
|
||||
Sync's remote side is already designed ([architecture.md §8](architecture.md)): ETag pruning turns a
|
||||
no-op sync of 50k images into one request. Nothing equivalent existed for a local root, which is the
|
||||
central problem this document solves.
|
||||
|
||||
---
|
||||
|
||||
## 1. What this must not do
|
||||
|
||||
Stated first because every design choice below follows from it.
|
||||
|
||||
| Must not | Why |
|
||||
|---|---|
|
||||
| Stat 50k files to open the catalog | NFR-P1: catalog open < 2 s desktop, < 4 s Android. SAF `DocumentsContract` queries are far slower than `stat` (spike S10). |
|
||||
| Re-derive thumbnails for unchanged images | NFR-P3 throughput is for *new* work; redoing it on every connect makes first paint unbounded. |
|
||||
| Fetch previews for remote images nobody looks at | A 50k remote library at 1–3 MB per range-extract is 50–150 GB. FR-NC-6 forbids bulk transfer by default. |
|
||||
| Evaluate cache rules per grid cell | ARCH §9.5 already answers this: `tier_desired` is materialised. |
|
||||
| Block the UI executor on any of it | NFR-P9, NFR-ARCH-1. |
|
||||
|
||||
The unifying principle: **work is proportional to what changed, or to what the user is looking at —
|
||||
never to library size.**
|
||||
|
||||
---
|
||||
|
||||
## 2. Schema
|
||||
|
||||
Extends [architecture.md §6.2](architecture.md). Additions beyond that sketch are marked ⊕.
|
||||
|
||||
```sql
|
||||
-- Roots -----------------------------------------------------------------
|
||||
roots(
|
||||
id INTEGER PRIMARY KEY,
|
||||
kind TEXT, -- 'local' | 'saf' | 'remote'
|
||||
grant_blob BLOB, -- SAF persisted permission; NULL on Linux
|
||||
label TEXT,
|
||||
last_seen INTEGER,
|
||||
scan_generation INTEGER -- ⊕ bumped per completed scan; see §3.4
|
||||
);
|
||||
|
||||
-- Folders: the unit of change detection, local and remote alike ---------
|
||||
folders(
|
||||
id INTEGER PRIMARY KEY,
|
||||
root_id INTEGER NOT NULL REFERENCES roots(id),
|
||||
parent_id INTEGER REFERENCES folders(id),
|
||||
path TEXT NOT NULL,
|
||||
etag TEXT, -- remote: propagating ETag (ARCH §8.4)
|
||||
mtime INTEGER, -- ⊕ local: directory mtime
|
||||
entry_count INTEGER, -- ⊕ local: direct children, mtime's blind spot
|
||||
scanned_generation INTEGER, -- ⊕ deletion sweep; see §3.4
|
||||
UNIQUE(root_id, path)
|
||||
);
|
||||
|
||||
-- Images ----------------------------------------------------------------
|
||||
images(
|
||||
id INTEGER PRIMARY KEY,
|
||||
root_id INTEGER NOT NULL REFERENCES roots(id),
|
||||
folder_id INTEGER REFERENCES folders(id), -- ⊕ folder filter without LIKE
|
||||
source_ref TEXT NOT NULL,
|
||||
content_hash TEXT, -- NULL until hashed; see §3.5
|
||||
format TEXT,
|
||||
w INTEGER, h INTEGER,
|
||||
captured_at INTEGER, -- UTC seconds; NULL if EXIF absent
|
||||
captured_offset INTEGER, -- ⊕ minutes east of UTC; see §4.2
|
||||
camera TEXT, lens TEXT,
|
||||
iso INTEGER, aperture REAL, shutter REAL,
|
||||
availability INTEGER,
|
||||
file_size INTEGER, -- ⊕ cheap change signal alongside mtime
|
||||
file_mtime INTEGER, -- ⊕
|
||||
metadata_state INTEGER, -- ⊕ 0=none 1=stat-only 2=full EXIF; §3.5
|
||||
sidecar_mtime INTEGER,
|
||||
UNIQUE(root_id, source_ref)
|
||||
);
|
||||
|
||||
-- Versions, keywords, remote, cache: per ARCH §6.2, unchanged -----------
|
||||
|
||||
-- Collections ⊕ ---------------------------------------------------------
|
||||
collections(
|
||||
id INTEGER PRIMARY KEY,
|
||||
name TEXT NOT NULL,
|
||||
parent_id INTEGER REFERENCES collections(id), -- collection sets
|
||||
kind INTEGER NOT NULL, -- 0 = manual, 1 = smart
|
||||
selector_json TEXT, -- smart only; the §5 Selector
|
||||
created INTEGER
|
||||
);
|
||||
|
||||
collection_members(
|
||||
collection_id INTEGER NOT NULL REFERENCES collections(id) ON DELETE CASCADE,
|
||||
image_id INTEGER NOT NULL REFERENCES images(id) ON DELETE CASCADE,
|
||||
position INTEGER, -- manual ordering; NULL = by capture time
|
||||
PRIMARY KEY(collection_id, image_id)
|
||||
);
|
||||
|
||||
-- Jobs ⊕ ----------------------------------------------------------------
|
||||
jobs(
|
||||
id INTEGER PRIMARY KEY,
|
||||
kind INTEGER NOT NULL,
|
||||
subject_id INTEGER, -- image or folder, per kind
|
||||
priority INTEGER NOT NULL,
|
||||
state INTEGER NOT NULL, -- 0=pending 1=running 2=failed
|
||||
attempts INTEGER NOT NULL DEFAULT 0,
|
||||
not_before INTEGER, -- retry backoff
|
||||
payload TEXT,
|
||||
UNIQUE(kind, subject_id) -- coalescing; see §6.2
|
||||
);
|
||||
```
|
||||
|
||||
Indices that exist for a stated query, not speculatively:
|
||||
|
||||
```sql
|
||||
CREATE INDEX images_captured ON images(captured_at); -- §4 timeline
|
||||
CREATE INDEX images_folder ON images(folder_id);
|
||||
CREATE INDEX images_hash ON images(content_hash) WHERE content_hash IS NOT NULL;
|
||||
CREATE INDEX folders_parent ON folders(parent_id);
|
||||
CREATE INDEX jobs_ready ON jobs(state, priority DESC, not_before);
|
||||
CREATE INDEX versions_image ON versions(image_id);
|
||||
CREATE INDEX members_image ON collection_members(image_id);
|
||||
```
|
||||
|
||||
`content_hash` is indexed *partially*. It is NULL for most rows most of the time (§3.5), and a
|
||||
partial index over the non-NULL subset is both smaller and what FR-CAT-9's reconnection-by-hash
|
||||
and FR-CAT-11's duplicate detection actually query.
|
||||
|
||||
---
|
||||
|
||||
## 3. Incremental scan
|
||||
|
||||
### 3.1 The local analogue of ETag pruning
|
||||
|
||||
Nextcloud propagates ETags up the tree, so one request proves a whole library unchanged
|
||||
([architecture.md §8.4](architecture.md)). A filesystem offers no such guarantee — a directory's
|
||||
mtime changes when its *direct* entries change, and not when a grandchild does. There is no
|
||||
cheap "did anything below here change" probe.
|
||||
|
||||
So local scan prunes at each level rather than at the root:
|
||||
|
||||
```
|
||||
scan(folder):
|
||||
(mtime, count) = stat(folder)
|
||||
if (mtime, count) == stored:
|
||||
# This directory's own entries are unchanged. Its files need no
|
||||
# examination at all — but subdirectories may still have changed
|
||||
# internally, so recurse into known children without listing.
|
||||
for child in stored_children(folder):
|
||||
scan(child)
|
||||
else:
|
||||
entries = list(folder) # the expensive call
|
||||
reconcile(folder, entries) # §3.3
|
||||
for child in entries.dirs: scan(child)
|
||||
mark scanned(folder, current_generation)
|
||||
```
|
||||
|
||||
Cost is **one `stat` per directory** when nothing changed, versus one per *file*. A 50k-image
|
||||
library in ~2k folders costs 2k stats — a few milliseconds locally, and the difference between
|
||||
meeting and missing NFR-P1 on SAF.
|
||||
|
||||
The recursion into unchanged directories is not redundant: it is what makes a change to one deep
|
||||
file detectable at all, given no upward propagation. What it avoids is the *listing* — on SAF a
|
||||
`DocumentsContract` query returning 200 rows costs far more than a metadata probe on the directory
|
||||
itself.
|
||||
|
||||
### 3.2 Why entry-count as well as mtime
|
||||
|
||||
Directory mtime alone misses a real case: delete one file and create another within the same
|
||||
timestamp granularity, and mtime can be unchanged while contents differ. Some filesystems and most
|
||||
SAF providers report coarse timestamps, which widens the window.
|
||||
|
||||
Storing `(mtime, entry_count)` closes the common form of this — a paired add and remove changes
|
||||
neither, but that is rarer than a bare add or remove, and both of those move the count. It is a
|
||||
cheap narrowing, not a proof.
|
||||
|
||||
**Where correctness must not depend on it,** the user gets an explicit *Rescan folder* action
|
||||
(FR-CAT-1), and reconnection matches by content hash (FR-CAT-9). Sync's remote path is unaffected —
|
||||
ETags are authoritative there.
|
||||
|
||||
### 3.3 Reconciling a changed directory
|
||||
|
||||
For each entry in a listing:
|
||||
|
||||
| Situation | Action |
|
||||
|---|---|
|
||||
| Not in catalog | Insert with `metadata_state = 1`; enqueue `ExtractMetadata` |
|
||||
| In catalog, `(size, mtime)` match | Nothing — the common case |
|
||||
| In catalog, `(size, mtime)` differ | Re-enqueue `ExtractMetadata` and `Thumbnail`; clear `content_hash` |
|
||||
| In catalog, absent from listing | Deletion candidate — §3.4 |
|
||||
| Placeholder (`*.nextcloud`) | Catalogue as the image it stands for; `Availability::Offline` (ARCH §9.0) |
|
||||
|
||||
Sidecars are examined in the same pass: a `.drsc` whose mtime exceeds `images.sidecar_mtime` enqueues
|
||||
a `ReadSidecar` job. This is how an edit made on another device — landed by the Nextcloud client,
|
||||
not by us — reaches the catalog.
|
||||
|
||||
### 3.4 Deletion without a full sweep
|
||||
|
||||
A file removed outside the app appears only as an *absence*, which a pruned scan cannot see: the
|
||||
folder it vanished from has a changed mtime and is listed, but a folder never visited is never
|
||||
compared.
|
||||
|
||||
Generation counting handles this without a full pass. Each scan bumps `roots.scan_generation`, and
|
||||
every folder reached — whether listed or skipped — records it. After the walk:
|
||||
|
||||
```sql
|
||||
-- Folders never reached: their parent no longer lists them.
|
||||
DELETE FROM folders
|
||||
WHERE root_id = ?1 AND scanned_generation < ?2;
|
||||
```
|
||||
|
||||
Images under a deleted folder cascade. Images missing from a *listed* folder are caught directly in
|
||||
§3.3. Together these cover deletion with no additional traversal.
|
||||
|
||||
Deletion here means **removing the catalog row for a source proven absent**, which FR-CAT-9 sharply
|
||||
distinguishes from a source merely unreachable. A root that fails to open at all — unplugged drive,
|
||||
revoked SAF grant — aborts the scan and marks the root offline. It never runs the sweep, because
|
||||
every folder would look unreached and the sweep would delete the entire library.
|
||||
|
||||
That guard is the single most dangerous line in this design, and it is stated as an invariant:
|
||||
**the deletion sweep runs only after a scan that completed without a root-level access error.**
|
||||
|
||||
### 3.5 Metadata in two passes
|
||||
|
||||
Full EXIF extraction requires opening and parsing each file. At 50k images that is minutes, and it
|
||||
must not stand between the user and a usable grid.
|
||||
|
||||
`metadata_state` records how far each image has got:
|
||||
|
||||
| State | Holds | Cost |
|
||||
|---|---|---|
|
||||
| 0 — none | Row exists, nothing read | — |
|
||||
| 1 — stat-only | Name, size, mtime, format from extension | Free, from the listing |
|
||||
| 2 — full | EXIF: capture time, camera, lens, exposure, dimensions | One open + parse |
|
||||
|
||||
The grid is usable at state 1: it can show filenames, sort by filename or file mtime, and display
|
||||
placeholder cells. Promotion to state 2 runs as background jobs, prioritised by what is on screen
|
||||
(§6.3), so visible images get real capture times within a frame or two of being scrolled to.
|
||||
|
||||
**Capture-time filtering (§4) needs state 2**, so a freshly scanned library's timeline is incomplete
|
||||
until the pass finishes. The UI states this plainly — a progress affordance on the timeline, not a
|
||||
silently wrong filter. Which is the FR-NC-6c principle applied to metadata rather than pixels: say
|
||||
what you actually have.
|
||||
|
||||
`content_hash` is a *third*, still lazier tier. It requires reading the whole file, so it is computed
|
||||
only when something needs it: import duplicate detection (FR-CAT-11), or reconnecting a moved source
|
||||
(FR-CAT-9). Never during a routine scan.
|
||||
|
||||
---
|
||||
|
||||
## 4. The library view
|
||||
|
||||
### 4.1 Query model
|
||||
|
||||
The UI never assembles SQL. It hands the catalog a `Query` and receives a stable, windowable result:
|
||||
|
||||
```rust
|
||||
pub struct Query {
|
||||
pub filter: Selector, // §5 — same type cache rules use
|
||||
pub sort: Sort,
|
||||
pub descending: bool,
|
||||
}
|
||||
|
||||
pub enum Sort {
|
||||
CapturedAt,
|
||||
Added,
|
||||
FileName,
|
||||
Rating,
|
||||
/// Manual order within a collection; falls back to CapturedAt elsewhere.
|
||||
CollectionPosition,
|
||||
}
|
||||
```
|
||||
|
||||
Results are fetched by window, never wholesale — FR-CAT-4 requires memory bounded independently of
|
||||
catalog size:
|
||||
|
||||
```rust
|
||||
impl Catalog {
|
||||
fn count(&self, q: &Query) -> Result<usize, CatalogError>;
|
||||
fn window(&self, q: &Query, range: Range<usize>) -> Result<Vec<GridRow>, CatalogError>;
|
||||
}
|
||||
```
|
||||
|
||||
`GridRow` carries exactly what a cell draws — id, thumbnail key, availability, rating, flag, capture
|
||||
time — and nothing that would require a join per cell. Availability badges read `tier_desired`
|
||||
directly (ARCH §9.5), so no rule evaluation happens on the render path.
|
||||
|
||||
A `LIMIT/OFFSET` window degrades at high offsets, since SQLite must walk the skipped rows. Scrolling
|
||||
is overwhelmingly *sequential*, so the catalog keeps a keyset cursor for forward and backward paging
|
||||
and falls back to OFFSET only for a scrollbar jump. Jumps are rare and single; scrolling is
|
||||
continuous.
|
||||
|
||||
### 4.2 Time
|
||||
|
||||
Capture time is the spine of a photo library, and it has one persistent trap: **a photograph's
|
||||
timestamp is local to where it was taken.** Store UTC alone and a shoot that ran 09:00–17:00 in
|
||||
Tokyo displays as spanning two days in Paris. Store local time alone and ordering across a timezone
|
||||
change is wrong.
|
||||
|
||||
So both: `captured_at` in UTC for ordering, `captured_offset` in minutes for display and for
|
||||
day-bucketing. EXIF `OffsetTimeOriginal` supplies it where present; where absent — common on older
|
||||
bodies — the offset is NULL and the catalog falls back to the library's configured display timezone,
|
||||
flagged so the UI can show it as inferred.
|
||||
|
||||
Day, month, and year buckets are computed against **local** time. "Everything from 3 August" means
|
||||
the photographer's 3 August.
|
||||
|
||||
The timeline affordance is a histogram of counts per bucket, which the grid uses for scrubbing:
|
||||
|
||||
```rust
|
||||
pub enum Granularity { Year, Month, Day, Hour }
|
||||
|
||||
pub struct TimeBucket {
|
||||
pub start: i64, // UTC seconds, bucket start
|
||||
pub count: u32,
|
||||
}
|
||||
|
||||
fn timeline(&self, q: &Query, g: Granularity) -> Result<Vec<TimeBucket>, CatalogError>;
|
||||
```
|
||||
|
||||
This is one grouped aggregate over the `images_captured` index, not 50k rows into the UI. It is what
|
||||
makes "drag across two years to find the trip" work, and it is the cheapest useful thing a library
|
||||
view can offer over a flat grid.
|
||||
|
||||
### 4.3 Filtering interactively
|
||||
|
||||
FR-CAT-6 requires filter results to update interactively on 50k images. Three things make that hold:
|
||||
|
||||
1. **Filters compile to indexed predicates.** A `Selector` becomes a WHERE clause over indexed
|
||||
columns. Keyword and collection membership become `EXISTS` subqueries against their own indices.
|
||||
2. **Count and first window are one round trip.** The grid needs a row count to size its scrollbar
|
||||
and the first screenful to paint; the catalog returns both together.
|
||||
3. **A filter change cancels the one in flight.** Typing in a search box issues a query per
|
||||
keystroke; each supersedes the last (NFR-ARCH-3). Without this the UI queues work it will discard.
|
||||
|
||||
---
|
||||
|
||||
## 5. Selectors: one type, three uses
|
||||
|
||||
[architecture.md §9.2](architecture.md) defines `Selector` for cache rules. The same type expresses
|
||||
library filters and smart collections. This is deliberate and worth stating as a design decision,
|
||||
because three near-identical predicate languages is a classic way for a catalog to rot.
|
||||
|
||||
| Use | Meaning |
|
||||
|---|---|
|
||||
| Library filter | What the grid shows now |
|
||||
| Smart collection | A saved, named filter (FR-CAT-7) |
|
||||
| Cache rule | What is kept locally, at which tier (FR-NC-6a) |
|
||||
|
||||
One consequence is directly useful: any filter the user has narrowed to can be saved as a smart
|
||||
collection, and any collection can be pinned offline, with no conversion step. "Show me 5-star images
|
||||
from the last 90 days" → save as a collection → pin it for the trip. Three features, one mechanism.
|
||||
|
||||
`Selector` moves to `dr-types` so `dr-catalog` and `dr-sync` share it without either depending on the
|
||||
other. It gains variants the cache-rule sketch did not need:
|
||||
|
||||
```rust
|
||||
pub enum Selector {
|
||||
All, // ⊕ the empty filter
|
||||
Collection(CollectionId),
|
||||
Folder { root: RootId, path: String, recursive: bool },
|
||||
DateRange(DateSelector),
|
||||
Rating { min: u8 },
|
||||
Label(ColourLabel),
|
||||
Flag(FlagState),
|
||||
Keyword(String),
|
||||
Camera(String), // ⊕ FR-CAT-6 indexed field
|
||||
Lens(String), // ⊕
|
||||
IsoRange { min: u32, max: u32 }, // ⊕
|
||||
Availability(Availability), // ⊕ "what can I edit right now"
|
||||
Text(String), // ⊕ filename/keyword substring
|
||||
Person { id: PersonId, include_suggested: bool }, // ⊕ §10 (FR-CULL-11)
|
||||
All_(Vec<Selector>),
|
||||
Any(Vec<Selector>),
|
||||
Not(Box<Selector>),
|
||||
}
|
||||
```
|
||||
|
||||
`Person` carries `include_suggested` rather than defaulting silently. A saved collection built from
|
||||
confirmed faces must not quietly change membership because a later indexing pass guessed at another
|
||||
face; the user chose "photos of Anna", not "photos the model currently believes contain Anna". The
|
||||
default is `false`, and the interactive filter offers the looser form explicitly as a way to *find*
|
||||
faces to confirm.
|
||||
|
||||
`Availability` as a selector earns its place: on a tablet the most useful filter is often "what do I
|
||||
actually have here", and it is also the natural thing to *pin* — "keep everything I've flagged that
|
||||
isn't already local".
|
||||
|
||||
Compilation is a straightforward recursive walk producing SQL with bound parameters. **Nothing
|
||||
user-supplied is ever interpolated into SQL text.** `Text` becomes a bound `LIKE` pattern with `%`,
|
||||
`_`, and the escape character escaped.
|
||||
|
||||
---
|
||||
|
||||
## 6. Background work
|
||||
|
||||
### 6.1 Job kinds
|
||||
|
||||
```rust
|
||||
pub enum JobKind {
|
||||
ScanFolder, // §3, recursive from a folder
|
||||
ExtractMetadata, // state 1 → 2
|
||||
Thumbnail, // §7
|
||||
ReadSidecar, // external sidecar change detected
|
||||
WriteSidecar, // local edit → disk, debounced (ARCH §6.1)
|
||||
ContentHash, // on demand only
|
||||
FetchPreview, // remote range-extract (FR-NC-3)
|
||||
FetchOriginal, // pinned or explicitly requested
|
||||
DetectFaces, // §10, on the proxy tier (FR-CULL-8)
|
||||
}
|
||||
```
|
||||
|
||||
### 6.2 Coalescing is the point
|
||||
|
||||
`UNIQUE(kind, subject_id)` on `jobs` means enqueueing is idempotent: an image touched five times
|
||||
during a scan has one thumbnail job, not five. Enqueue is
|
||||
`INSERT … ON CONFLICT DO UPDATE SET priority = max(priority, excluded.priority)`, so a re-request at
|
||||
higher priority promotes the existing row rather than duplicating it.
|
||||
|
||||
This is what makes "regenerate on update" safe to call liberally. Every code path that notices a
|
||||
change can just enqueue; the table absorbs the redundancy.
|
||||
|
||||
### 6.3 Priority
|
||||
|
||||
Reuses the existing GPU scheduler classes ([architecture.md §5.3](architecture.md)) so one notion of
|
||||
priority governs the whole app:
|
||||
|
||||
| Class | Jobs | Preempts |
|
||||
|---|---|---|
|
||||
| `Interactive` | Metadata and thumbnails for visible cells; preview for the open image | everything |
|
||||
| `Prefetch` | The scroll margin; next image in culling | Background |
|
||||
| `Background` | Bulk metadata, rule-driven fetches, hashing | — |
|
||||
|
||||
Visible-cell work is enqueued by the grid as it scrolls, at `Interactive`. The effect is that a
|
||||
freshly scanned library fills in *where the user is looking* first, and grinds through the rest
|
||||
behind them.
|
||||
|
||||
### 6.4 Durability and failure
|
||||
|
||||
Jobs live in the catalog, so they survive process death — which on Android is routine, not
|
||||
exceptional (FR-PLAT-AND-3). On startup, rows in state `running` revert to `pending`: the process
|
||||
that owned them is gone.
|
||||
|
||||
Failures increment `attempts` and set `not_before` to an exponential backoff. After a bounded retry
|
||||
count the job is marked failed and attached to its image as a typed error (NFR-ARCH-4) — one
|
||||
corrupt file does not stall the queue, and the user can see which files failed and why.
|
||||
|
||||
**A job runner never touches the UI executor**, and `Interactive` work runs on the decode pool with
|
||||
the I/O pool behind it (ARCH §7.1).
|
||||
|
||||
---
|
||||
|
||||
## 7. Thumbnails
|
||||
|
||||
### 7.1 When
|
||||
|
||||
Not "on first connect" as a bulk operation. Thumbnails are generated:
|
||||
|
||||
- **On demand**, for cells entering the viewport plus the prefetch margin — at `Interactive`
|
||||
- **On change**, when §3.3 sees a differing `(size, mtime)`
|
||||
- **On rule**, for images a cache rule pins at `Preview` or above — at `Background`
|
||||
- **Never** for a remote image nobody has looked at and no rule covers
|
||||
|
||||
For a local library this converges on "everything, eventually", because scrolling reaches everything
|
||||
and the background pass has nothing else to do. For a remote library it converges on "what you
|
||||
actually browsed".
|
||||
|
||||
**Measured on a real 17,185-RAW library, 2026-08-09:** cataloguing it by whole-file fetch would move
|
||||
roughly **370 GB**; the range-extract path moves a few MB for the images actually viewed. This is
|
||||
the single largest cost difference in the design, and it is why §7.1 is a list of narrow triggers
|
||||
rather than "generate them all on connect".
|
||||
|
||||
### 7.2 How, by availability
|
||||
|
||||
| Availability | Source | Cost |
|
||||
|---|---|---|
|
||||
| `Original`, local | Embedded JPEG via `dr-decode` preview path | ~200 KB read, no demosaic |
|
||||
| `Original`, no embedded preview | Full decode, downscale | Expensive — `Background` only |
|
||||
| Remote | Range-extract embedded JPEG (FR-NC-3) | 1–3 MB vs 25–100 MB — **measured: 262 KB of a 21.5 MB DNG, 119 ms, 1.22% of the file** |
|
||||
| Placeholder / `Offline` | None — render the offline affordance | 0 |
|
||||
|
||||
The remote path deliberately does **not** ask the Nextcloud client to hydrate the file. ARCH §9.0
|
||||
established hydration is whole-file, so it costs ~100× what the range extract does. Hydration stays
|
||||
reserved for the original tier, where the user has asked for the actual image.
|
||||
|
||||
Server previews (`/core/preview`) are tried only where PROPFIND reported `nc:has-preview`. ARCH §6.7
|
||||
verified stock Nextcloud ships no RAW preview provider, so for RAW this is nearly always absent — it
|
||||
is an opportunistic saving, never the mechanism.
|
||||
|
||||
### 7.3 Storage: sharded, shared, synced
|
||||
|
||||
Decided 2026-08-09, implemented in `dr-thumbs`. Thumbnails live in **sharded SQLite databases that
|
||||
sync to Nextcloud**, so a second device gets a full grid without re-fetching a byte of RAW.
|
||||
|
||||
```text
|
||||
thumbs/
|
||||
index.sqlite fileid → shard, size accounting, client id, adoption ledger
|
||||
shard-0000.sqlite ≤ 25 MB, sealed
|
||||
shard-0001.sqlite ≤ 25 MB, active
|
||||
```
|
||||
|
||||
**Why a thumbnail is worth syncing when the catalog mostly is not.** It is expensive to produce — a
|
||||
range fetch plus a decode, per image — and byte-identical for every client looking at the same file.
|
||||
This does not make it authoritative: losing the store costs regeneration and nothing else, so §6.12
|
||||
is untouched.
|
||||
|
||||
**Why shards, and why small.** The 25 MB cap is about *sync granularity*, not SQLite's limits. One
|
||||
growing database means every client re-downloads all of it whenever a single thumbnail is added.
|
||||
With sequential fill only the newest shard is ever dirty, so an up-to-date client transfers one small
|
||||
file. Sealed shards are immutable, which makes them safe to cache forever and cheap to skip.
|
||||
|
||||
At ~20 KB per 256px JPEG a shard holds roughly 1,200 thumbnails, so the 17,185-image reference
|
||||
library lands in ~14 shards.
|
||||
|
||||
**Keyed on `oc:fileid`** — stable across server-side rename and move (FR-NC-5), and already in hand
|
||||
from PROPFIND. Accepted consequence: shards are account-scoped, so the same photograph on two
|
||||
servers is thumbnailed twice.
|
||||
|
||||
**Stored as JPEG, not raw pixels.** A 256×170 RGBA buffer is ~174 KB against ~15 KB encoded. Since
|
||||
shards sync, that 11× is transfer cost paid by every client, not just disk.
|
||||
|
||||
Three invariants, each tested:
|
||||
|
||||
| Invariant | Why it matters |
|
||||
|---|---|
|
||||
| A sealed shard never reopens | Reopening one forces every client that holds it to re-download |
|
||||
| Re-storing an existing id updates in place, never migrates | Migrating would rewrite a sealed shard |
|
||||
| Merging another client's shard is insert-only and idempotent | Both copies derive from the same bytes by the same code, so neither is better; preferring ours avoids dirtying a shard others have synced |
|
||||
|
||||
**The transfer**, in `dr-ui`'s `derived_sync`, exchanges shards with `.darkroom-derived/` under the
|
||||
library root. `ThumbStore::shards()` reports which are sealed, so an up-to-date client's whole pass
|
||||
is one listing plus whichever shard is still open.
|
||||
|
||||
**Why a remote name carries a client id.** Corrected 2026-08-16. Shard ids are *per store* — every
|
||||
client fills its own numbering from 0 — so the flat `shard-NNNN.sqlite` namespace the transfer first
|
||||
used had two clients writing one name. Two failures followed from it, and both were live: the second
|
||||
client's upload **overwrote** content the first still believed was published, and no client could
|
||||
distinguish a peer's shard 3 from its own, so the only safe reading of "I already hold 3" was to skip
|
||||
it. Between them, two populated clients exchanged almost nothing — only shards numbered above the
|
||||
other's highest. A fresh device worked, which is why it went unnoticed: with no local shards there is
|
||||
nothing to collide with.
|
||||
|
||||
The name is now `shard-<client>-NNNN.sqlite`, where `<client>` is minted per store in `index.sqlite`
|
||||
beside the numbering it qualifies — a store deleted and rebuilt restarts at shard 0 and must not
|
||||
claim its predecessor's names. Since a client's own ids no longer say anything about what it has
|
||||
taken from others, `index.sqlite` also keeps an **adoption ledger** of merged remote names and the
|
||||
size each had. Size, not a flag: a peer's sealed shard never returns, but its open one grows, and
|
||||
re-merging the grown copy is how the thumbnails it gained arrive.
|
||||
|
||||
Flat names left on servers by earlier builds are still read — they report no owner, so each client
|
||||
adopts them once — and nothing is written under that form again. A flat name whose id and byte size
|
||||
match a local shard is that client's own earlier upload by the same identity argument used for
|
||||
sealed shards, so the rename does not cost every client a re-download of its whole store. Older
|
||||
builds ignore the new names, so they stop receiving shards until updated; nothing is lost, since
|
||||
their own uploads are still adopted.
|
||||
|
||||
Two size classes remain planned — grid (256px) and filmstrip/loupe (1024px). Only the grid class is
|
||||
implemented. The cache is LRU-capped per NFR-RES-4, and thumbnails evict before proxies and long
|
||||
after sidecars, which never evict at all (FR-NC-6b).
|
||||
|
||||
---
|
||||
|
||||
## 7a. Editing collections
|
||||
|
||||
Decided 2026-08-09. The schema for collections landed with §2 and the cross-device merge rules with
|
||||
§8; this is the layer between them — the operations a user actually performs, in
|
||||
`dr_catalog::collections`.
|
||||
|
||||
### 7a.1 Hierarchy and membership are independent
|
||||
|
||||
Two structures, deliberately not entangled:
|
||||
|
||||
| | Mechanism | Meaning |
|
||||
|---|---|---|
|
||||
| Hierarchy | `collections.parent_id` | A collection inside a collection (Lightroom's "collection set"). A parent is an ordinary collection, not a separate kind, so a set can hold images of its own |
|
||||
| Membership | `collection_members` | An image is in as many collections as the user likes. Nothing moves on disk; no collection owns an image |
|
||||
|
||||
**Adding an image to a child does not write a row for the parent.** A parent's contents are the union
|
||||
of its own members and its descendants', computed on read. Materialising it instead would make one
|
||||
add touch every ancestor, and a reparent rewrite membership — both of which §8's row-level merge
|
||||
would then have to reconcile. The read path pays a bounded tree walk instead, which at sidebar scale
|
||||
is nothing.
|
||||
|
||||
The consequence the UI depends on: dragging images onto a collection is **additive**. It does not
|
||||
remove them from anywhere, which is why the gesture's default action is `copy` and not `move`.
|
||||
|
||||
### 7a.2 Rules that exist to prevent silent damage
|
||||
|
||||
| Rule | Why |
|
||||
|---|---|
|
||||
| Every mutation bumps `revision` | §8.4 resolves conflicts by revision. An edit that updates `modified` alone is invisible to the merge, so the *other* device silently wins and the user's work vanishes |
|
||||
| A no-op add does **not** bump it | Otherwise an idle device that re-dropped the same images outranks one that did real work |
|
||||
| Deleting a parent **promotes** its children | The schema's `ON DELETE CASCADE` would take the whole subtree. Losing a nested collection because its container was tidied away is not recoverable |
|
||||
| Deletion leaves a tombstone | Without it, merging with a device that still holds the collection resurrects it (§8.4) |
|
||||
| Cycles are refused at the write | Both kinds — parenting under a descendant, and a smart collection whose selector reaches itself. A cycle is unbounded recursion in the tree walk, so it must not be *representable*, not merely handled when drawn |
|
||||
| Tree walks are depth-guarded anyway | A merge can deliver a row this device never validated. The read path must terminate, so it truncates and logs rather than hanging the UI thread |
|
||||
| A drop onto a smart collection is refused | Its membership *is* its selector; member rows would be a second source of truth that nothing reads |
|
||||
| Deep counts are `count(DISTINCT image_id)` | An image in both a parent and a child is one photograph. A count that disagrees with the number of cells drawn makes both untrustworthy |
|
||||
|
||||
`collections.uuid` is generated from the OS CSPRNG. A collision fuses two unrelated collections at
|
||||
the next merge, so the fallback path (used only if `/dev/urandom` cannot be read) logs loudly rather
|
||||
than degrading identity quality in silence.
|
||||
|
||||
### 7a.3 Drag and drop is Slint's, not ours
|
||||
|
||||
The first implementation hand-rolled the gesture on `TouchArea` — tracking the press, measuring
|
||||
travel to distinguish a click from a drag, and deciding the drop target from the last row hovered.
|
||||
**It did not work**, for a reason worth recording: an interactive `Flickable` claims any drag
|
||||
beginning inside it for scrolling and *cancels* the child `TouchArea`'s press, so the gesture could
|
||||
never leave the grid. It also had a correctness hole — a tree rebuilt mid-drag could redirect the
|
||||
drop, since a captured pointer is invisible to every other element.
|
||||
|
||||
Slint 1.17's `DragArea`/`DropArea` own all of it: capture, the click-versus-drag threshold,
|
||||
arbitration against the `Flickable`, the image under the cursor, and hit-testing the release. What
|
||||
remains in `collections_ui` is only what Slint cannot know — the payload (which images, read from the
|
||||
selection when the drag starts) and the **spring**: a dwell timer that opens a collapsed collection
|
||||
so a nested child can be reached mid-drag, and closes again whatever the drag merely passed over.
|
||||
|
||||
One hazard survives the change and is easy to reintroduce. Every consequence of a drop — rebuilding
|
||||
the tree, refreshing the badges, rereading the grid — *replaces a Slint model*, and doing that inside
|
||||
the `dropped` handler destroys the elements Slint is still using to deliver that event. So the drop
|
||||
records its target and `drag-finished` acts on it. This is the same hazard `sync_rows` in `lib.rs`
|
||||
documents for the adjust panel, and it presents as a control that works once and then goes dead.
|
||||
|
||||
---
|
||||
|
||||
## 8. Syncing the catalog file
|
||||
|
||||
Decided 2026-08-09. **This qualifies [architecture.md §6.12](architecture.md)** — the catalog
|
||||
remains a rebuildable index, but the file itself now travels to Nextcloud. The qualification is
|
||||
worth stating precisely, because the sidecar-authoritative model is load-bearing and this is the
|
||||
one place it bends.
|
||||
|
||||
### 8.1 Why collections forced this
|
||||
|
||||
Every other thing the catalog holds has authoritative backing outside it. Ratings, labels,
|
||||
keywords, and edit graphs live in sidecars next to the images, so a rebuild recovers them.
|
||||
**Collections do not.** A manual collection is a set of images the user assembled by hand; nothing
|
||||
in the filesystem records it. Losing the catalog loses them, and no rescan brings them back.
|
||||
|
||||
So collections need to be durable across devices somehow. Syncing the catalog file is the chosen
|
||||
mechanism.
|
||||
|
||||
### 8.2 What the file sync does and does not carry
|
||||
|
||||
Only **collections and their membership** merge. The rest of a catalog describes *local* state —
|
||||
folder mtimes, cache file paths, job rows, `tier_actual` — and importing another device's version
|
||||
of those would be actively wrong. The downloaded remote is read for its collections and discarded.
|
||||
|
||||
This is what keeps §6.12 substantially intact: nothing here makes the local database authoritative
|
||||
for anything a rebuild could not recover. The catalog is still deletable. What syncs is one table
|
||||
pair that had no other home.
|
||||
|
||||
### 8.3 Two hazards the implementation must handle
|
||||
|
||||
**A WAL database is not one file.** Committed transactions can sit in `catalog.sqlite-wal` with the
|
||||
main file lagging, so copying `catalog.sqlite` alone uploads a torn snapshot — internally consistent
|
||||
as of some older point, silently missing everything since. Upload therefore runs a `TRUNCATE`
|
||||
checkpoint and then SQLite's backup API, which serialises against concurrent writers rather than
|
||||
racing them. It never copies the live file.
|
||||
|
||||
**Integer primary keys are not identities.** Two devices each allocate `collections.id = 1` for
|
||||
different collections, so a row-level merge keyed on the integer id would collide them. Collections
|
||||
therefore carry a **UUID**, and membership maps across devices by **image content hash**. The
|
||||
integer ids stay local and are never compared across catalogs.
|
||||
|
||||
### 8.4 Merge rules
|
||||
|
||||
| Concern | Rule | Why |
|
||||
|---|---|---|
|
||||
| Which collection wins | Higher `revision` — a counter bumped per local edit. `modified` only breaks an exact tie | A device with a skewed clock cannot silently overwrite real work. The same reason FR-NC-9 avoids mtime for sidecars |
|
||||
| Membership | **Set union**, not last-writer-wins | Two devices adding different images to one collection keep both. The exception — a removal racing an addition — resolves toward the addition, which is recoverable by removing it again. A lost addition is not |
|
||||
| Deletion | Tombstone (`deleted = 1`) carrying a revision | Without it, merging against a device that still holds the collection resurrects it. With a revision, deletion competes on equal footing with a rename |
|
||||
| An image the remote has and we do not | Skip the membership row | It joins on a later merge, once a scan has catalogued the file. Not an error |
|
||||
| A remote from a newer schema | Decline before attaching | Attempting it would fail mid-transaction rather than declining cleanly |
|
||||
|
||||
Merging is idempotent: running it twice reports no changes the second time. That property is tested,
|
||||
because a merge that oscillates would upload on every sync forever.
|
||||
|
||||
### 8.5 What was rejected
|
||||
|
||||
**Replace-if-newer.** The literal reading of "sync the file and take the newer one". Rejected
|
||||
because it is not a merge: whichever device syncs second loses every collection the first did not
|
||||
have. Binary SQLite files do not merge, so "newer wins" means "older is destroyed".
|
||||
|
||||
**A `collections.drsc` sidecar at the library root.** The alternative that would have kept §6.12
|
||||
untouched, merging as text the way edit sidecars do. Viable, and cheaper in machinery, but it means
|
||||
a second serialisation format and a second merge implementation for the same data. Recorded here
|
||||
because if the SQLite path proves troublesome, this is the fallback with a known shape.
|
||||
|
||||
---
|
||||
|
||||
## 9. What this document does not settle
|
||||
|
||||
- **FTS.** `Selector::Text` is a `LIKE` scan over filename and keywords. Adequate at 50k; if free
|
||||
text over description and title becomes a real workflow, an FTS5 table is the answer, and it is
|
||||
additive.
|
||||
- **Smart collection materialisation.** Currently evaluated on read. If a smart collection's
|
||||
membership needs to be *stable* — for manual ordering, or for a pinned set that must not shift
|
||||
under the user — it needs materialising with an invalidation rule. Deferred until there is a
|
||||
concrete need.
|
||||
- **Multi-root capture-time collisions.** FR-CAT-11 detects duplicates on import; the same image
|
||||
catalogued under two roots is a related but distinct case, not yet specified.
|
||||
- **Timeline granularity selection.** Which bucket size the UI picks for a given zoom is a UI
|
||||
concern, but the catalog should probably suggest one from the query's date span rather than have
|
||||
the UI guess.
|
||||
|
||||
---
|
||||
|
||||
## 10. People and faces
|
||||
|
||||
Specified by FR-CULL-8 … FR-CULL-12, NFR-SEC-5, [architecture.md §6.4](architecture.md). Gated on
|
||||
spike S14 and decision D13 — the runtime and the model licences are unresolved, so this is the shape
|
||||
of the subsystem, not a build order.
|
||||
|
||||
[faces.md](faces.md) names the models this shape is filled in with, and adds one column to §10.1's
|
||||
`faces` table (`crop_px`) that the calibration in its §8 depends on.
|
||||
|
||||
### 10.1 Schema (a v5 migration)
|
||||
|
||||
```sql
|
||||
CREATE TABLE people (
|
||||
id INTEGER PRIMARY KEY,
|
||||
uuid TEXT NOT NULL UNIQUE, -- merge identity, not the name (ARCH §6.3)
|
||||
name TEXT NOT NULL,
|
||||
-- Tombstone-by-redirect. A merged person must outlive its merge, or a
|
||||
-- device that still has it resurrects it — same hazard collections have.
|
||||
merged_into INTEGER REFERENCES people(id) ON DELETE SET NULL,
|
||||
created INTEGER NOT NULL,
|
||||
revision INTEGER NOT NULL DEFAULT 1,
|
||||
modified INTEGER NOT NULL
|
||||
);
|
||||
|
||||
CREATE TABLE faces (
|
||||
id INTEGER PRIMARY KEY,
|
||||
image_id INTEGER NOT NULL REFERENCES images(id) ON DELETE CASCADE,
|
||||
-- Normalised to the image's long edge, so a face survives the proxy it was
|
||||
-- found on being regenerated at another resolution.
|
||||
x REAL NOT NULL, y REAL NOT NULL, w REAL NOT NULL, h REAL NOT NULL,
|
||||
landmarks BLOB, -- 5 × (x, y) f32, the alignment input
|
||||
detector_confidence REAL NOT NULL,
|
||||
embedding BLOB NOT NULL, -- 512 × f16, the raw model output; re-normalised on load
|
||||
-- Length of that vector: the model's own reading of how recognisable the
|
||||
-- crop was, and the gate on whether this face may be compared *against*
|
||||
-- (faces.md §6, §9). NULL for a face stored as a unit vector before it
|
||||
-- was kept.
|
||||
quality REAL,
|
||||
-- What the eyes are doing (FR-CULL-8a, faces.md §17): per eye P(open),
|
||||
-- the source pixels across its box and the sharpness of the patch the
|
||||
-- classifier saw; and P(sunglasses). All seven or none; NULL is "never
|
||||
-- read", which every filter treats as unknown rather than as closed.
|
||||
-- The verdict -- open, closed, sunglasses, unclear -- is a rule in
|
||||
-- dr_face::eyes, not a column.
|
||||
eye_right REAL,
|
||||
eye_right_px REAL,
|
||||
eye_right_sharp REAL,
|
||||
eye_left REAL,
|
||||
eye_left_px REAL,
|
||||
eye_left_sharp REAL,
|
||||
sunglasses REAL,
|
||||
-- The 106 dense landmarks the eyes were read from, packed as 16-bit
|
||||
-- fixed point over the frame: 424 bytes (schema V18). Kept so the next
|
||||
-- per-face pass runs from the catalog rather than from the original.
|
||||
landmarks_dense BLOB,
|
||||
-- Which model produced this. An embedding is only comparable to others
|
||||
-- from the same model; mixing them silently yields nonsense similarities.
|
||||
model_id TEXT NOT NULL,
|
||||
detected_at INTEGER NOT NULL
|
||||
);
|
||||
CREATE INDEX faces_image ON faces(image_id);
|
||||
-- Covers the eyes-open filter's subquery. Without it every check read the
|
||||
-- whole face row -- the eye columns sit after the blobs -- and one count
|
||||
-- took 24 s on the reference library (schema V17).
|
||||
CREATE INDEX faces_eyes ON faces(image_id, eye_right, eye_right_px, eye_right_sharp,
|
||||
eye_left, eye_left_px, eye_left_sharp, sunglasses);
|
||||
|
||||
CREATE TABLE face_person (
|
||||
face_id INTEGER PRIMARY KEY REFERENCES faces(id) ON DELETE CASCADE,
|
||||
person_id INTEGER NOT NULL REFERENCES people(id) ON DELETE CASCADE,
|
||||
-- Calibrated P(this face is this person), never a raw cosine (FR-CULL-9).
|
||||
probability REAL NOT NULL,
|
||||
-- The user said so. Never overwritten by a later inference pass.
|
||||
confirmed INTEGER NOT NULL DEFAULT 0
|
||||
);
|
||||
CREATE INDEX face_person_person ON face_person(person_id, confirmed);
|
||||
```
|
||||
|
||||
Three things in that schema are load-bearing:
|
||||
|
||||
**`model_id` on every face.** Embeddings from different models are not comparable — this is the one
|
||||
mistake that produces plausible-looking garbage rather than an error. Storing the model with the
|
||||
embedding means a model change is detectable and re-indexable, instead of quietly poisoning every
|
||||
similarity in the library.
|
||||
|
||||
**Normalised bounding boxes.** Detection runs on whichever proxy exists (FR-CULL-8). Storing pixel
|
||||
coordinates would bind a face to a resolution that the cache is entitled to evict and regenerate
|
||||
differently.
|
||||
|
||||
**`confirmed` as a column, not a probability of 1.0.** A confirmation is a different kind of fact
|
||||
from a confident guess, and collapsing them loses the ability to recompute suggestions without
|
||||
touching user data.
|
||||
|
||||
### 10.2 Why clustering is not a job kind
|
||||
|
||||
Detection is per-image and parallel, so it is a job (`DetectFaces`, coalesced per image like any
|
||||
other). Clustering is a *whole-library* operation over the embeddings detection produced — it has no
|
||||
natural `subject_id`, and running it per-image would rebuild the world on every photograph.
|
||||
|
||||
It therefore runs as a debounced library-level pass, triggered when detection has been idle and the
|
||||
face count has moved materially since the last clustering. The same reasoning as sidecar writes: the
|
||||
work is cheap to defer, expensive to repeat, and nobody is waiting on it.
|
||||
|
||||
### 10.3 The calibration lives with the library
|
||||
|
||||
FR-CULL-9 requires similarity to be a calibrated probability, fitted from this library's own faces.
|
||||
That fit is a property of the catalog and its model, so it is stored alongside — a small table
|
||||
holding the fit parameters, its validity flag, and a hash of the face set it was derived from, so a
|
||||
materially changed library recomputes rather than trusting a stale fit.
|
||||
|
||||
When the fit is not valid — a library with too few faces to have positive pairs — the UI says the
|
||||
confidence is unavailable. It does not fall back to an untuned default dressed up as a measurement.
|
||||
|
||||
### 10.4 What this does not settle
|
||||
|
||||
- **Which model, and which runtime.** D13. Everything above holds regardless of the answer, which is
|
||||
why it is specified in terms of "a 512-d embedding from a stated model" rather than a named one.
|
||||
- **The clustering algorithm.** Density-based over the calibrated distance is the obvious starting
|
||||
point, but the parameters are an S14 question, not a design-time one.
|
||||
- **Whether embeddings sync.** NFR-SEC-5 permits it, opt-in. The shard mechanism in §7.3 is the
|
||||
obvious carrier if they do, but nothing here depends on that decision.
|
||||
- **Faces in trashed images.** FR-CAT-15's trash moves files; whether their faces stay indexed and
|
||||
keep contributing to clusters is unspecified. Probably they should be excluded from suggestions but
|
||||
not deleted, so a restore does not re-index.
|
||||
|
||||
---
|
||||
|
||||
## 10a. Bursts and near-duplicates
|
||||
|
||||
Specified by FR-CULL-5, implemented in `dr_catalog::bursts` (a v11 migration) with the pass that
|
||||
feeds it in `dr_ui::bursts`.
|
||||
|
||||
A burst is a run of frames that are **adjacent in time and look like the frame before them**. Both
|
||||
halves are load-bearing. Time alone groups a whole wedding ceremony, because a photographer working
|
||||
steadily never leaves the gap that would end the run. Similarity alone groups a studio setup shot
|
||||
across two days, which is a project rather than a moment. The bounds are two seconds and eight bits
|
||||
of a 64-bit difference hash, and the reasoning for each figure is in the module.
|
||||
|
||||
**Two seconds, for a burst that fires ten frames in one.** `images.captured_at` is whole seconds:
|
||||
EXIF's `DateTimeOriginal` has no sub-second field, and `SubSecTimeOriginal` is optional and widely
|
||||
omitted. Ten frames of a burst therefore arrive sharing a timestamp, and any threshold finer than a
|
||||
second is a threshold on information the catalog does not have. Where the pace really is faster than
|
||||
two seconds, the similarity bound is what separates the frames.
|
||||
|
||||
**The signal is a perceptual hash of the thumbnail, not of the original.** `images.perceptual_hash`
|
||||
is filled from the 256px thumbnails §7 already stores — vastly more resolution than a 9×8 reduction
|
||||
uses — so a library that has been browsed has already paid for its signatures and no RAW is decoded
|
||||
for this. The consequence is stated rather than hidden: an image with no thumbnail gets no
|
||||
signature, and a frame with no signature never joins a burst. It is picked up by the next pass.
|
||||
|
||||
**It is a pass, not a job kind**, for exactly the reason §10.2 gives for face clustering: a burst is
|
||||
a property of a *run* of frames and has no natural `subject_id`, so a per-image job would rebuild
|
||||
the world once per photograph. It runs when the thumbnail sweep finishes, which is the first moment
|
||||
the signatures can all be computed.
|
||||
|
||||
**A newly found burst arrives open.** The pass marks frames; it never takes them off the screen.
|
||||
Collapsing on discovery would be tidier and would also mean a background pass removing photographs
|
||||
from under someone part way through a cull. Folding a burst up is the user's act, it is remembered
|
||||
(`burst_expanded`), and a burst that is already known keeps whatever state it is in — so the pass
|
||||
that follows the next import does not spring open a morning's work.
|
||||
|
||||
**Nothing here ranks a frame.** The representative of a collapsed burst is its *earliest* frame,
|
||||
which is a fact about the clock rather than a judgement about the photograph. FR-CULL-5 names the
|
||||
failure this avoids — rejecting the only frame of an important moment because someone blinked — and
|
||||
the only judgement in the subsystem is the user's own choice of representative, which lives in its
|
||||
own table (`burst_pick`) so that rebuilding the grouping cannot erase it. Same argument as
|
||||
`people.ignored` in §10.
|
||||
|
||||
**What the collapse costs the grid.** Which rows a collapsed burst hides has to be decided by the
|
||||
query rather than by the cells, because the grid is a window (`LIMIT n OFFSET k`) and the frames it
|
||||
hides are mostly not loaded. So the predicate joins `VISIBLE` in every query that lists or counts
|
||||
cells, under the same discipline: present in four places of five, the header's count, the
|
||||
scrollbar, the shift-click range and the scrub's ordinal stop describing the same list.
|
||||
|
||||
**What this does not settle.** Bursts are local: the tables ride along in the uploaded catalog
|
||||
snapshot and nothing on the far side reads them, so a second device rebuilds its own grouping from
|
||||
its own signatures. Making `burst_pick` cross-device is a merge question of the same shape as §8.4's
|
||||
and is not answered here.
|
||||
|
||||
---
|
||||
|
||||
## 11. Requirements touched
|
||||
|
||||
| ID | How this document addresses it |
|
||||
|---|---|
|
||||
| FR-CAT-1 | §3 incremental scan, cancellable and resumable via §6 jobs |
|
||||
| FR-CAT-3 | §7 thumbnail pyramid, two size classes, embedded-preview fast path |
|
||||
| FR-CAT-4 | §4.1 windowed queries, memory independent of catalog size |
|
||||
| FR-CAT-5 | §3.5 two-pass metadata |
|
||||
| FR-CAT-6 | §4.3 indexed filter compilation, §5 selectors |
|
||||
| FR-CAT-7 | §2 collections schema, §5 manual and smart, §7a hierarchy, membership and editing |
|
||||
| FR-CAT-9 | §3.4 the offline/deleted distinction and the sweep guard |
|
||||
| FR-CAT-11 | §3.5 lazy content hashing |
|
||||
| FR-NC-3 | §7.2 range-extract for remote thumbnails |
|
||||
| FR-NC-6a | §5 shared selector type |
|
||||
| FR-NC-6c | §3.5 metadata honesty, §7.2 availability-driven sourcing |
|
||||
| NFR-P1 | §3.1 one stat per directory, not per file |
|
||||
| NFR-P3 | §7.1 on-demand generation |
|
||||
| NFR-ARCH-2 | §6.3 priority classes shared with the GPU scheduler |
|
||||
| FR-CULL-5 | §10a burst grouping: capture-time proximity and image similarity, collapse without selection |
|
||||
| NFR-ARCH-3 | §4.3 query cancellation, §6 job cancellation |
|
||||
| NFR-RES-4 | §7.3 LRU cap, eviction order |
|
||||
| FR-CULL-8 | §10.1 `faces` schema, §6.1 `DetectFaces` job kind on the proxy tier |
|
||||
| FR-CULL-9 | §10.3 per-library calibration, stored with its validity and source hash |
|
||||
| FR-CULL-10 | §10.1 `people` and `face_person`, merge-by-redirect, §10.2 clustering as a library pass |
|
||||
| FR-CULL-11 | §5 `Selector::Person`, confirmed-only by default |
|
||||
| FR-CULL-12 | §10.1 derived data in the catalog; names to the sidecar, UUID as merge identity |
|
||||
| FR-CULL-8a | §10.1 the seven eye columns, derived like the embedding; the filter term is `RatingFilter::eyes_open` |
|
||||
| NFR-SEC-5 | §10.4 sync left undecided and off; nothing in §10 emits an embedding |
|
||||
@@ -0,0 +1,366 @@
|
||||
# DarkRoom — Code health and the cost of a contribution
|
||||
|
||||
**Status:** Audit · 2026-08-27
|
||||
**Companion to:** [architecture.md](architecture.md), [technical-debt.md](technical-debt.md),
|
||||
[view-composition.md](view-composition.md)
|
||||
|
||||
What it costs to add something to this codebase, measured rather than estimated, and the work that
|
||||
would lower the price.
|
||||
|
||||
[technical-debt.md](technical-debt.md) records compromises that were *chosen* — each one has a
|
||||
reason that outlived the person who took it. This document records the opposite: friction nobody
|
||||
chose, which accumulated because no single commit was responsible for it. The distinction matters
|
||||
when deciding what to touch. A TD entry is load-bearing until its "done when" is met; an entry here
|
||||
is not defending anything.
|
||||
|
||||
It is also not a bug list. Everything below compiles, passes 2,042 tests and ships.
|
||||
|
||||
---
|
||||
|
||||
## 1. What was measured
|
||||
|
||||
Every figure in this document is reproducible from a clean checkout. Worktrees under `.claude/` and
|
||||
build output under `target/` are excluded from all counts — including them roughly triples the line
|
||||
totals and was the first thing to get wrong.
|
||||
|
||||
```bash
|
||||
# Lines of Rust per crate
|
||||
for d in core/* ui/* platform/* apps/* tools/*; do
|
||||
[ -d "$d/src" ] && echo "$(find $d/src -name '*.rs' -exec cat {} + | wc -l) $d"
|
||||
done | sort -rn
|
||||
|
||||
# unwrap() in production code only — split each file at its #[cfg(test)] marker
|
||||
# (a naive grep counts ~1,500 and tells you nothing)
|
||||
|
||||
# Distinct window properties written per UI module
|
||||
for f in ui/dr-ui/src/*.rs; do
|
||||
echo "$(grep -oP '\b(w|window|win|ui)\.\Kset_[a-z0-9_]+(?=\()' "$f" | sort -u | wc -l) $f"
|
||||
done | sort -rn
|
||||
```
|
||||
|
||||
| Measure | Value |
|
||||
|---|---|
|
||||
| Rust across 19 crates | ~112,000 lines; ~78,000 after comments and blanks |
|
||||
| Comment density | 28% overall, 20–34% per crate |
|
||||
| Test functions | 2,042, plus 21 integration test files |
|
||||
| `.unwrap()` in production code | **3** — one in `dr-gpu`, two in `dr-ingest` |
|
||||
| `.unwrap()` in test code | ~1,500, which is where it belongs |
|
||||
| `unsafe` blocks | 9 — three of them the face scan's SIMD kernels (faces.md §9) |
|
||||
| `TRACES` tags / orphan tags | 793 / 0 |
|
||||
| Resolved dependencies | 826 |
|
||||
| Largest function | `dr-ui::run` — 1,855 lines |
|
||||
| `AppWindow` members | 282 properties + 190 callbacks |
|
||||
|
||||
CI gates on `cargo fmt --check`, `cargo clippy --workspace --all-targets -- -D warnings`,
|
||||
`cargo test --workspace`, a release build, and an Android cross-check. The traceability matrix is
|
||||
regenerated and compared, with a pre-commit hook that keeps it in step.
|
||||
|
||||
---
|
||||
|
||||
## 2. The seams, graded
|
||||
|
||||
Six things a contributor might plausibly want to add. Each grade was checked against the tree.
|
||||
|
||||
| Feature | What you touch | Cost |
|
||||
|---|---|---|
|
||||
| A develop operation<br>*split toning, channel mixer* | One file in `core/dr-pipeline/ops/`. Nothing else. | **Trivial** |
|
||||
| A RAW format | A `Format` variant and magic-byte recognition. `dr-decode` is a generic TIFF walker carrying only 3 format-specific branches. | **Easy** |
|
||||
| A neighbourhood operation<br>*dehaze, a sharpener* | Rust in `dr-pipeline/src/ops/` implementing `Operation` + `DetailStage`, plus a stub YAML declaring `rust:` and `order:`. | **Moderate** |
|
||||
| A parameter widget<br>*a colour wheel* | A `WidgetKind` variant, `develop::supported()`, the panel model, a Slint component. | **Moderate** |
|
||||
| A second sync backend<br>*S3, WebDAV, a local folder* | `RemoteBackend` is the easy half. Seven UI files construct `NextcloudBackend` directly and ten signatures take it concretely — see [CH-2](#ch-2). | **Hard** |
|
||||
| Anything with its own UI | `app.slint`'s root component, a ~1,000-line `wire()`, and `run()` at 1,855 lines — see [CH-1](#ch-1). | **Hard** |
|
||||
|
||||
The top half of that table is the good half, and it is very good. The bottom half is one problem
|
||||
wearing two hats: **`dr-ui` has no seams, so every UI feature lands in the same three files.**
|
||||
|
||||
---
|
||||
|
||||
## 3. What is load-bearing, and must not be "tidied"
|
||||
|
||||
Listed before the findings deliberately. A remediation document that only enumerates problems invites
|
||||
someone to fix something that was right.
|
||||
|
||||
**The operation declaration format.** `ops/*.yaml` + `build.rs` is a working plugin system that
|
||||
happens to resolve at build time — see [display-and-extension.md §6](display-and-extension.md). Its
|
||||
deliberate smallness is the point: the expression grammar is restricted so a declaration cannot
|
||||
become a second, worse place to write code. Do not "improve" it by letting a node name arbitrary
|
||||
Rust.
|
||||
|
||||
**No operation is named in `ui/`.** Verified: all fifteen built-in op ids grepped across every `.rs`
|
||||
and `.slint` file in `ui/` yield exactly one hit, a localisation test in `labels.rs:353`. This is
|
||||
what makes "a new operation is one file" true rather than aspirational, and it is the single
|
||||
property most likely to be destroyed by a well-meaning special case in the panel. See [CH-5](#ch-5).
|
||||
|
||||
**Unimplemented widgets degrade rather than break.** `develop::supported()` lists every `WidgetKind`
|
||||
explicitly instead of using a wildcard, so a new kind added to the core surfaces as a compile error
|
||||
rather than as silence, and a widget nothing draws falls back to sliders with the edit still
|
||||
working (ARCH §4.3a).
|
||||
|
||||
**The mpsc-plus-timer worker shape.** Slint's event loop must never block (NFR-P9). Threading is not
|
||||
what any finding below proposes changing.
|
||||
|
||||
**`sync_rows` mutating rows in place.** Replacing the model breaks slider dragging.
|
||||
|
||||
---
|
||||
|
||||
## <a id="ch-1"></a>CH-1 — `dr-ui` has no view layer, and the cost is compounding
|
||||
|
||||
**Where:** `ui/dr-ui/src/lib.rs:836`, `library_ui.rs:4374`, `collections_ui.rs:1457`,
|
||||
`ui/dr-ui/ui/app.slint`
|
||||
|
||||
### What it is
|
||||
|
||||
Every UI feature lands in the same three places: the Slint root component, one of the `wire()`
|
||||
functions, and `run()`.
|
||||
|
||||
| Function | Lines |
|
||||
|---|---|
|
||||
| `lib.rs::run` | 1,855 |
|
||||
| `library_ui::wire` | 998 |
|
||||
| `collections_ui::wire` | 964 |
|
||||
| `identity_ui::wire` | 478 |
|
||||
| `masks_ui::wire` | 357 |
|
||||
| `settings_ui::wire` | 317 |
|
||||
|
||||
Above them sits one `AppWindow` carrying 282 properties and 190 callbacks, with exactly one Slint
|
||||
global in the whole `ui/` directory — and it is not exported. Cross-view state therefore has nowhere
|
||||
to live except the root component, and every interaction is a callback registered inside a `wire`.
|
||||
|
||||
The core crates are healthy by the same measure: their largest functions are `compose_full` at 437
|
||||
lines and `build.rs::emit_node` at 558, both of which earn their length. This is specific to `dr-ui`.
|
||||
|
||||
### Why it matters more than it did
|
||||
|
||||
**[view-composition.md](view-composition.md) already diagnosed this and specified the fix**, in
|
||||
three independently landable stages, on 2026-08-09. None of the three has landed. In the eighteen
|
||||
days since, the numbers that document itself used have moved:
|
||||
|
||||
| Its measure | 2026-08-09 | 2026-08-27 |
|
||||
|---|---|---|
|
||||
| `run()` | 500 lines | **1,855** |
|
||||
| `library_ui` distinct window properties | 24 | **54** |
|
||||
| `lib.rs` distinct window properties | 23 | **52** |
|
||||
| `collections_ui` | 9 | 11 |
|
||||
| `launch_ui` | 16 | 16 |
|
||||
| `set_show_*` call sites | 7 | 9 |
|
||||
| View-state booleans on `AppWindow` | 2 | **5** |
|
||||
|
||||
Four modules it did not list now write window properties too: `settings_ui` (39), `import_ui` (24),
|
||||
`identity_ui` (18), `masks_ui` (16).
|
||||
|
||||
The prediction it made has also come true literally. It described `app.slint` compensating for two
|
||||
mutually exclusive booleans with `if !root.show-launch && root.show-library` chains. There are now
|
||||
five such booleans — `show-launch`, `show-library`, `show-identity`, `show-settings`, `show-import` —
|
||||
and the chains at `app.slint:1405`, `:1445` and `:1654` are five-term conjunctions. Each new view
|
||||
multiplies the conjunctions rather than adding to them.
|
||||
|
||||
None of this is difficult work. It is simply work that every feature must now do, in files every
|
||||
other feature is also editing — which is why two contributors working in parallel conflict by
|
||||
construction, and why a newcomer must read a 1,855-line startup sequence with real ordering
|
||||
constraints before safely inserting a line into it.
|
||||
|
||||
### What to do
|
||||
|
||||
Execute [view-composition.md](view-composition.md) as written. Its analysis holds and its staging is
|
||||
right; stage 1 alone removes the nullable-callback knot and the duplicated post-load sequence, and
|
||||
is worth landing whether or not stages 2 and 3 follow.
|
||||
|
||||
One thing to add to it, because it was not in scope there: the `wire()` functions. They are already
|
||||
sectioned internally by comment, so lifting each section into `fn wire_ratings(window, ctl)`,
|
||||
`fn wire_keywords(...)` and so on is mechanical, checked entirely by the compiler, and can go one
|
||||
section per commit. It gives a feature a *function* to own rather than a region of one, which is
|
||||
what removes the conflict, and it does not wait on stage 1.
|
||||
|
||||
Do the extraction behind new features rather than as a big-bang refactor. The pile grows either way;
|
||||
the question is only whether each new feature adds to it or subtracts.
|
||||
|
||||
**Done when:** no function in `dr-ui` exceeds 300 lines, a new view registers itself instead of
|
||||
adding a boolean to `AppWindow`, and `active-view` has replaced the boolean set so both-true is
|
||||
unrepresentable.
|
||||
|
||||
---
|
||||
|
||||
## <a id="ch-2"></a>CH-2 — `RemoteBackend` is an abstraction nothing above `dr-sync` uses
|
||||
|
||||
**Where:** `core/dr-sync/src/lib.rs:46`; seven files in `ui/dr-ui/src`
|
||||
|
||||
### What it is
|
||||
|
||||
The trait is carefully built. Capability negotiation decides the sync strategy; range reads are
|
||||
documented as a hint rather than a guarantee so correctness holds either way; chunked upload is
|
||||
deliberately kept internal so one server's protocol cannot leak into the interface. Every choice is
|
||||
explained where it is made.
|
||||
|
||||
And then `trash.rs`, `import.rs`, `export.rs`, `derived_sync.rs`, `launch_ui.rs`, `library.rs` and
|
||||
`settings_ui.rs` each construct `NextcloudBackend` directly — 34 references — and ten functions take
|
||||
`&NextcloudBackend` rather than `&dyn RemoteBackend`. Exactly **two** sites in the tree take the
|
||||
trait object, both inside `dr-sync` itself.
|
||||
|
||||
### Why it matters
|
||||
|
||||
The abstraction currently buys nothing it was designed for. Worse, it reads as though it does: a
|
||||
contributor who wants a WebDAV or local-folder backend will find a well-documented trait, implement
|
||||
it correctly, and only then discover that nothing above `dr-sync` can be handed the result.
|
||||
|
||||
There is no defence of this in the tree, which is what makes it an entry here rather than in
|
||||
`technical-debt.md`. It is what a single-backend application looks like when the second backend has
|
||||
not yet been attempted.
|
||||
|
||||
### What to do
|
||||
|
||||
Change the ten signatures to `&dyn RemoteBackend` and construct the backend once, behind something
|
||||
the UI does not name — the same discipline `develop.rs` already applies to operations. The trait is
|
||||
already correct, so this is a mechanical change, and it is much cheaper now than during a second
|
||||
backend when it would be entangled with that backend's own problems.
|
||||
|
||||
Worth doing even if no second backend is ever written: it makes the sync layer testable against a
|
||||
fake, which today it is not.
|
||||
|
||||
**Done when:** `NextcloudBackend` is named in at most one file in `ui/`, and a stub backend can be
|
||||
substituted in a test without touching the UI.
|
||||
|
||||
**Done**, with one half deliberately left. `ui/dr-ui/src/remote.rs` is now the only file in the
|
||||
interface that names a connector; the ten worker functions take `&dyn RemoteBackend` and will accept
|
||||
a stub. Every method the UI ever called on a backend — `get`, `put`, `list`, `delete`, `create_dir`,
|
||||
`move_to` — was already on the trait, so nothing had to be added to it.
|
||||
|
||||
What remains is **credentials**. `AppCredentials` is an app password obtained through Login Flow v2,
|
||||
which is a Nextcloud protocol rather than a general notion of how one authenticates to a remote, and
|
||||
seven files still name it. Abstracting it needs a decision about what an account *is* across
|
||||
backends — an OAuth token, a bucket key pair and an app password have no useful common shape — and
|
||||
making that decision before a second backend exists would produce a confident wrong answer. It is a
|
||||
design problem rather than a mechanical one, and it should wait for the backend that forces it.
|
||||
|
||||
---
|
||||
|
||||
## <a id="ch-3"></a>CH-3 — There is no path in for a contributor who is not already here
|
||||
|
||||
**Where:** repository root
|
||||
|
||||
### What it is
|
||||
|
||||
No `CONTRIBUTING.md`. No `rust-toolchain.toml`, though CI pins 1.92.0 exactly. No issue or PR
|
||||
templates.
|
||||
|
||||
The documentation that exists is excellent — 7,990 lines across 14 files, including 177 numbered
|
||||
requirements — and all of it is written for someone who has already decided to work on this. Nothing
|
||||
tells a newcomer which document to read first, that `core/dr-pipeline/ops/README.md` is the door
|
||||
with the lowest bar, or that a clone without `git-lfs` needs one command before the build succeeds.
|
||||
|
||||
That last point is handled well in the code: `dr-segment`'s build script detects an LFS pointer file
|
||||
and fails with an instruction rather than embedding 130 bytes and dying at inference time. It is
|
||||
simply not written anywhere a first-time cloner would look.
|
||||
|
||||
### Why it matters
|
||||
|
||||
It is the cheapest item in this document and it gates every other contribution. A person who cannot
|
||||
get a first build is not going to reach the parts that are good.
|
||||
|
||||
### What to do
|
||||
|
||||
Write `CONTRIBUTING.md` and point the first door at the operation format. "Add a develop operation"
|
||||
is a genuinely one-file contribution with declared tests that run under `cargo test` — the best
|
||||
first experience this codebase can offer, and it happens to teach the architecture's central idea on
|
||||
the way through.
|
||||
|
||||
Then state the three things that are currently folklore: `git lfs` is a prerequisite, the first
|
||||
build resolves 826 crates and takes a while (saying so stops it reading as a hang), and the
|
||||
toolchain is 1.92.0. Add `rust-toolchain.toml` so that last one is enforced rather than documented —
|
||||
a contributor on an older stable currently gets confusing type errors instead of a version message.
|
||||
|
||||
**Done when:** someone who has never seen the repository can clone it, build it, and land a new
|
||||
`ops/*.yaml` node without asking a question.
|
||||
|
||||
---
|
||||
|
||||
## <a id="ch-4"></a>CH-4 — Coverage is counted by tagging, not by behaviour
|
||||
|
||||
**Where:** `docs/traceability.md`, `tools/traceability`
|
||||
|
||||
### What it is
|
||||
|
||||
Not a new finding — [display-and-extension.md §7](display-and-extension.md) states it plainly, and
|
||||
`traceability.md` itself says coverage is the intersection of tagged and defined IDs. The tooling is
|
||||
genuinely good: 732 tags, zero orphans, denominators parsed from `requirements.md` at run time
|
||||
rather than hardcoded, regenerated in CI and guarded by a pre-commit hook.
|
||||
|
||||
What it cannot do is check that the code under a tag does the thing. `FR-DEV-8` is currently tagged
|
||||
against instance-buffer plumbing a future spot-removal operation *would* use; `FR-DEV-7` against a
|
||||
history row for a frontend that does not exist. Both read as covered.
|
||||
|
||||
### Why it matters
|
||||
|
||||
The 55.4% figure is an overstatement of unknown size, and the risk is that it is used as a planning
|
||||
input. It is recorded here so that the number keeps its asterisk when read outside the document that
|
||||
already qualified it.
|
||||
|
||||
### What to do
|
||||
|
||||
Nothing structural — the honest framing already exists in two places. Adopt
|
||||
display-and-extension.md's rule going forward: **close a requirement with a test that would fail if
|
||||
the behaviour were removed**, and let the percentage move slowly and mean something.
|
||||
|
||||
**Done when:** the rule is stated in `CONTRIBUTING.md` alongside the tag syntax, so it reaches
|
||||
someone adding their first tag.
|
||||
|
||||
---
|
||||
|
||||
## <a id="ch-5"></a>CH-5 — The best invariant in the codebase is unprotected
|
||||
|
||||
**Where:** `ui/dr-ui/src`, `ui/dr-ui/ui`
|
||||
|
||||
### What it is
|
||||
|
||||
"No code in `ui/` names an operation" (FR-DEV-3a) is what makes the whole declarative pipeline pay
|
||||
off, and it is currently maintained by discipline alone. Nothing fails if someone special-cases
|
||||
`exposure` in the panel to fix a layout problem at five in the evening.
|
||||
|
||||
### Why it matters
|
||||
|
||||
It is one grep, it would take an hour, and it protects the property this audit rates highest. The
|
||||
failure mode is silent and cumulative: the first special case is defensible, and by the fifth the
|
||||
panel names half the chain and "a new operation is one file" has quietly stopped being true.
|
||||
|
||||
### What to do
|
||||
|
||||
A test that greps `ui/dr-ui/src` and `ui/dr-ui/ui` for every id in `ops/*.yaml` and fails on a hit,
|
||||
with the current `labels.rs` localisation test as its one allowed exception. It belongs in CI beside
|
||||
the traceability check, which is the existing precedent for a structural gate.
|
||||
|
||||
**Done when:** adding `window.set_exposure_slider(...)` to the panel fails CI with a message naming
|
||||
FR-DEV-3a.
|
||||
|
||||
---
|
||||
|
||||
## 4. Order of work
|
||||
|
||||
Ordered by value per hour rather than by size.
|
||||
|
||||
| | Item | Effort | Why first |
|
||||
|---|---|---|---|
|
||||
| 1 | [CH-3](#ch-3) — `CONTRIBUTING.md` + `rust-toolchain.toml` | Half a day | Gates everything else; unblocks the trivial seam that already works |
|
||||
| 2 | [CH-5](#ch-5) — CI gate on operation names in `ui/` | An hour | Protects the property everything else in the pipeline rests on |
|
||||
| 3 | [CH-2](#ch-2) — make `RemoteBackend` load-bearing | 1–2 days | Mechanical now, entangled later; also makes sync testable |
|
||||
| 4 | [CH-1](#ch-1) — split the `wire` functions | Incremental | No behaviour change, compiler-checked, one section per commit |
|
||||
| 5 | [CH-1](#ch-1) — [view-composition.md](view-composition.md) stages 1–3 | Sustained | Largest and most invasive; every deferred month adds to the pile |
|
||||
|
||||
Items 1, 2 and 3 are independent of each other and of the rest. Item 4 does not wait on item 5.
|
||||
|
||||
---
|
||||
|
||||
## 5. What this document does not claim
|
||||
|
||||
It measures structure, discipline and coupling. It does **not** assess runtime correctness, GPU
|
||||
shader behaviour, security posture, or whether any tagged requirement is actually implemented —
|
||||
[CH-4](#ch-4) is precisely the observation that the last of those is unmeasured.
|
||||
|
||||
Line counts and function sizes are proxies. `run()` being 1,855 lines is a real problem because
|
||||
every feature must edit it, not because 1,855 is a bad number; `build.rs::emit_node` at 558 lines is
|
||||
not a problem at all. Where a figure appears above, the sentence around it says which of the two it
|
||||
is.
|
||||
|
||||
The audit was first measured against `origin/master` at `d4a34ef`, then re-measured after
|
||||
`android-bundled-face-models` merged in: `run()` moved from 1,810 lines to 1,855, and the whole-
|
||||
library face sweep added its fetching path to `library.rs`. The seam grades are unchanged by that
|
||||
merge — the work went into the seams that already existed rather than cutting new ones, which is
|
||||
itself the pressure CH-1 describes.
|
||||
@@ -0,0 +1,253 @@
|
||||
# Finishing the display contract, and opening the pipeline
|
||||
|
||||
Spec for two pieces of work that turn out to be one conversation: closing **FR-DSP**, which is the
|
||||
architecture's central performance claim, and reaching **FR-PLG**, which is the only requirement
|
||||
family at zero.
|
||||
|
||||
They belong in one document because the same property decides both. The pipeline composes its work
|
||||
from *declarations* — an operation says what its parameters are and contributes a WGSL fragment,
|
||||
and the composer fuses the active ones into a single dispatch. That is why the display path is fast,
|
||||
and it is also, already, most of a plugin format. Finishing one and opening the other are the same
|
||||
seam approached from two sides.
|
||||
|
||||
---
|
||||
|
||||
## 1. What is actually true today
|
||||
|
||||
Stated first because both halves of this document are smaller than the requirement numbers suggest,
|
||||
and the reason is that some of the work is done and untagged.
|
||||
|
||||
| Requirement | Reality |
|
||||
|---|---|
|
||||
| FR-DSP-1 proxy rendering | **Done.** The develop view renders at viewport resolution, not source. |
|
||||
| FR-DSP-2 tiled computation | **Absent, and §2 now says it should stay that way.** Measured: the fused pass is inside the budget everywhere. See [frame-budget.md](frame-budget.md). |
|
||||
| FR-DSP-3 interactive latency | **Measured and asserted** for the fused path — `core/dr-gpu/tests/frame_budget.rs`. Missed by one operation, clarity, for the reason recorded as TD-4. |
|
||||
| FR-DSP-4 progressive refinement | **Absent**, and §4's condition did not fire. Every render is full quality and can afford to be. |
|
||||
| FR-DSP-5 zoom and pan | **Done and tagged**, against tests that fail if the behaviour is removed — `core/dr-gpu/tests/zoom_resolution.rs`. `Framing::view` shrinks the sampled region while the render target keeps its size, so zooming *raises* the resolution the pipeline works at. That is FR-DSP-5's requirement, arrived at without tiles. |
|
||||
| FR-DSP-6 colour management | **Done.** Output space is a parameter of composition. |
|
||||
| FR-DSP-7 histogram and clipping | **Done**, GPU-side, no per-frame readback. |
|
||||
| FR-DSP-8 per-display colour | **Done**, with one caveat named in §5.4. Acquisition per display server, an sRGB fallback that is visible in About, and the canvas rendered at physical pixel size. |
|
||||
| FR-PLG-* | **Zero tagged.** But `ops/*.yaml` + `build.rs` is already the class-1 plugin format compiled at build time rather than loaded. |
|
||||
|
||||
Two of the five uncovered display requirements are therefore *measurement and tagging*, not
|
||||
construction. That is worth knowing before anyone plans a quarter around them.
|
||||
|
||||
**Both have since been done.** [frame-budget.md](frame-budget.md) holds the measurements §2 asks
|
||||
for and the reading of its decision rule; the table above is updated to match. The rest of this
|
||||
document is left as it was written, because a plan that has been overtaken by its own evidence is
|
||||
more useful read in order than quietly edited into agreement.
|
||||
|
||||
---
|
||||
|
||||
## 2. Measure before building tiles
|
||||
|
||||
**FR-DSP-2 is the one requirement in this document that may not be worth satisfying as written.**
|
||||
|
||||
> **Resolved.** M1–M3 were run; the numbers and the verdict are in
|
||||
> [frame-budget.md](frame-budget.md). The rule below fired for *rewrite*: every point-operation
|
||||
> chain is inside 16 ms at the 99th percentile at every viewport size, fit and at 1:1, the widest
|
||||
> being 4.5 ms of GPU at 4K. The measurement did find a stage that misses the budget — clarity's
|
||||
> 52-pixel kernel, 34 ms at 4K — and tiling makes that stage *worse*, since a tiled convolution
|
||||
> reads a halo per tile. It is recorded as TD-4 with the fix its own module already names.
|
||||
|
||||
|
||||
The requirement predates the fused-shader design. It assumes the pipeline is a chain of passes over
|
||||
a large buffer, where recomputing everything on each frame would be ruinous and tiles are the way
|
||||
out. What was built instead composes every active operation into **one dispatch over a
|
||||
viewport-sized target** — at 2000×1300 that is 2.6 M pixels, once, for the whole chain.
|
||||
|
||||
So the question tiling was invented to answer may already be answered. Before any tile scheduler is
|
||||
written:
|
||||
|
||||
**M1 — Frame cost at proxy resolution.** Time `render_detailed` at 1920×1200, 2560×1600 and
|
||||
3840×2160, with a chain of one operation, five, and every operation active. Report the 99th
|
||||
percentile, not the mean; a slider drag is judged by its worst frame.
|
||||
|
||||
**M2 — Frame cost at 1:1 on a large file.** The same, with `Framing::view` zoomed to 1:1 on a 60 MP
|
||||
frame, which is the case FR-DSP-5 names and the one where the sampled region is smallest but the
|
||||
detail chain's kernels are widest.
|
||||
|
||||
**M3 — Cost of the detail stage separately.** Neighbourhood operations dispatch per pass and are the
|
||||
only part of the chain whose cost is not one read and one write. A separable blur at a large radius
|
||||
is the plausible budget-breaker, not the fused pass.
|
||||
|
||||
**The decision rule, fixed in advance.** If M1 and M2 sit inside 16 ms at the 99th percentile,
|
||||
**FR-DSP-2 is rewritten rather than implemented**: tiling stops being an interactive-path
|
||||
requirement and becomes what it actually is for this architecture — a *scheduling* concern for
|
||||
export and thumbnailing, which already run off the frame path. If they do not, the measurement tells
|
||||
us which stage to tile, which is a far better starting point than tiling everything on principle.
|
||||
|
||||
Writing a tile scheduler that the design does not need would be the most expensive way to discover
|
||||
this. ARCH §5.3's tile cache keyed by `(VersionId, tile, zoom, graph_hash_prefix)` is a good design
|
||||
for a pipeline that needs it; the burden of proof is that this one does.
|
||||
|
||||
---
|
||||
|
||||
## 3. FR-DSP-3 — make the budget a test, not an aspiration
|
||||
|
||||
A latency requirement that nothing asserts is a wish. The work is:
|
||||
|
||||
**3.1** A bench in `dr-gpu` that renders a fixed chain at a fixed size and reports percentiles.
|
||||
Committed with its numbers, so a regression is a diff rather than a memory.
|
||||
|
||||
**3.2** A test that *fails* when a frame exceeds the budget on the reference desktop, skipping where
|
||||
there is no adapter — the pattern the GPU tests already use. It should assert the 99th percentile of
|
||||
a hundred frames, because the failure mode being guarded against is a stutter, not an average.
|
||||
|
||||
**3.3** The asynchronous half of the requirement: "when a full-resolution result is needed it is
|
||||
computed asynchronously, and the proxy result remains on screen until it is ready." Nothing does
|
||||
this today because nothing needs a full-resolution result on the frame path — export renders its
|
||||
own. This clause should be **narrowed to export and 1:1 zoom** or struck, and struck is defensible.
|
||||
|
||||
---
|
||||
|
||||
## 4. FR-DSP-4 — progressive refinement
|
||||
|
||||
The one genuinely new piece of interactive work, and it is small because the pipeline is already
|
||||
resolution-parametric.
|
||||
|
||||
During a drag, render at a fraction of the viewport and let the compositor scale; when the gesture
|
||||
settles, render at full viewport size. `DevelopSession` already knows when a drag is in flight —
|
||||
`drag-changed` exists on every slider and is what stands the Flickable down.
|
||||
|
||||
Two things decide whether this is worth having, and M1 answers both. If a full-quality frame is
|
||||
already inside budget, reduced-quality rendering buys nothing and costs a visible softness during
|
||||
every drag — which the requirement itself warns against ("refinement is visually smooth, not a
|
||||
jarring swap"). **This requirement is conditional on M1 failing.** If M1 passes, FR-DSP-4 is
|
||||
satisfied vacuously: there is no rapid interaction the app cannot render at full quality, which is a
|
||||
stronger outcome than refining.
|
||||
|
||||
---
|
||||
|
||||
## 5. FR-DSP-8 — per-display colour
|
||||
|
||||
The only display requirement needing platform work rather than pipeline work, and the only one where
|
||||
being wrong is a correctness defect rather than a slow frame: a second monitor with a different
|
||||
profile shows wrong colours, silently.
|
||||
|
||||
**5.1 Acquisition, per display server.** X11 has `_ICC_PROFILE` atoms per output. Wayland's
|
||||
colour-management protocol is not universally available, and the requirement already anticipates
|
||||
this by demanding "a defined fallback where Wayland provides no profile" — that fallback is sRGB,
|
||||
stated in the About page beside the other diagnostics so a photographer can see which path they are
|
||||
on rather than wonder.
|
||||
|
||||
**5.2 Reacting to a move.** The transform is selected per the display currently showing the canvas
|
||||
and updates when the window moves. The composed output space is already a parameter of composition
|
||||
(`compose_with_framing(..., output)`), so a display change is a recomposition, not a pipeline
|
||||
change. This is the part the existing design makes cheap.
|
||||
|
||||
Slint turned out *not* to report window moves — there is no `on_moved` on any backend — so the
|
||||
window's position and scale factor are sampled twice a second and the platform re-surveyed only when
|
||||
they differ. See §5.4 for the display server where the position itself is unavailable.
|
||||
|
||||
**5.3 Fractional scaling.** "Handled without resampling artefacts in the canvas" — the canvas is a
|
||||
wgpu texture handed to the compositor, so the requirement is that we render at the *physical* pixel
|
||||
size rather than the logical one and let the compositor present 1:1. Worth an explicit test, since
|
||||
the failure is subtle: a slightly soft canvas that looks like a bad demosaic.
|
||||
|
||||
### 5.4 What landed, and the one thing that did not
|
||||
|
||||
`dr_plat::display` surveys the session's displays; `dr_ui::display_ui` decides which one is showing
|
||||
the canvas and keeps `DevelopSession`'s output space pointed at it. Composition was already
|
||||
parameterised on the space, so the pixel path changed by one argument.
|
||||
|
||||
Two things are worth recording because they are trades rather than omissions.
|
||||
|
||||
**A profile is matched to the nearest of four spaces, not applied.** A measured panel is none of
|
||||
`Srgb`, `DisplayP3`, `AdobeRgb` or `ProPhoto`, and a general ICC engine is a much larger piece of
|
||||
work — a CMM, rendering intents, LUT-based profiles, and a per-frame cost to argue about. The
|
||||
profile is reduced to its D50-adapted colorants and matched against the four; a match that is merely
|
||||
nearest is marked as such, and About says "nearest to Display P3" rather than "Display P3". A
|
||||
LUT-based profile, which is what a hardware calibrator often writes, is declined by shape and falls
|
||||
back to sRGB with that stated. An approximation the photographer can see beats a silent one.
|
||||
|
||||
**On Wayland the canvas follows the first output, not the window.** A Wayland client is never told
|
||||
where its window is — `xdg_toplevel` carries no position, deliberately — so the "which display"
|
||||
question cannot be answered by geometry there. The protocol's own answer is
|
||||
`wp_color_management_surface_feedback_v1`, which hands a client the preferred image description for
|
||||
*its surface* and re-sends it on a move; it needs the application's `wl_surface`, which Slint owns
|
||||
and does not expose. So the profiles are read correctly for every output and the *selection* among
|
||||
them is right on X11 and on any single-monitor Wayland session, which is most of them. Closing the
|
||||
gap is a Slint surface handle, not a change to any of this.
|
||||
|
||||
---
|
||||
|
||||
## 6. Extensibility: the format already exists
|
||||
|
||||
`FR-PLG-2` says "the node declaration is the plugin format". That is already true — it is simply
|
||||
resolved at build time:
|
||||
|
||||
```
|
||||
ops/exposure.yaml ──build.rs──▶ generated Rust impl Operation ──▶ fused shader
|
||||
```
|
||||
|
||||
A declaration names its parameters, their ranges and units, its attributes, its WGSL body and its
|
||||
neutral. `build.rs` compiles that into something indistinguishable from a hand-written operation.
|
||||
**Nothing about that requires the declaration to be present at compile time** — everything it
|
||||
produces is data plus a WGSL string, and the composer already assembles WGSL at run time from
|
||||
whatever operations are active.
|
||||
|
||||
So class 1 is not a new mechanism. It is the existing one, loaded later.
|
||||
|
||||
### 6.1 What has to change
|
||||
|
||||
**6.1.1 Descriptors become owned, not `&'static`.** `Operation::descriptor()` returns
|
||||
`&'static OpDescriptor` today, which is what makes a build-time node free and a run-time node
|
||||
impossible. This is the one invasive change in the whole plan and everything else waits behind it.
|
||||
`Arc<OpDescriptor>` is the obvious shape; the cost is one refcount per descriptor read, on a path
|
||||
that reads descriptors when the panel is built rather than per frame.
|
||||
|
||||
**6.1.2 A run-time node type.** One `DeclaredOp` implementing `Operation` from an owned
|
||||
declaration, replacing *generated code per node* with *one interpreter over many declarations*. The
|
||||
generated path can stay for the built-in chain — it costs nothing and keeps the built-ins
|
||||
inspectable — but the two must produce identical behaviour, which is a test: parse each built-in
|
||||
`ops/*.yaml` at run time and assert the composed WGSL matches the generated one byte for byte.
|
||||
|
||||
**6.1.3 WGSL validation at load, not at dispatch.** A plugin's fragment is a string from a stranger.
|
||||
`compose` already builds a full shader and `naga` will reject bad source, but the failure currently
|
||||
surfaces as a broken render. A plugin's source must be compiled and rejected at *load*, with the
|
||||
error naming the plugin, because the alternative is an app that draws nothing and blames itself.
|
||||
|
||||
**6.1.4 Order and identity.** `order:` decides chain position and `build.rs` already refuses
|
||||
duplicates — that guard becomes load-time. Plugin ids need a namespace (`author.name`) so two
|
||||
plugins cannot collide, and the sidecar stores parameters by `(op_id, param_id)`, so an id collision
|
||||
is a *wrong edit silently applied*, exactly the failure `MaskSource::Regions`' signature exists to
|
||||
prevent.
|
||||
|
||||
### 6.2 What this buys immediately
|
||||
|
||||
The features enumerated as missing against Lightroom that are *pure point operations* become
|
||||
declarations rather than code: split toning, colour zones, selective colour, creative vignette,
|
||||
channel mixer variants. A photographer-author can write one without a Rust toolchain, and the
|
||||
existing `ops/README.md` is already its documentation.
|
||||
|
||||
**It does not buy the neighbourhood operations** — dehaze, spot removal, liquify — because those are
|
||||
`DetailStage` implementations with kernels and per-render scale conversion, which FR-PLG-2a
|
||||
anticipates by naming "fragment nodes and pass nodes" as two templates. Pass nodes are a second
|
||||
phase and should not gate the first.
|
||||
|
||||
### 6.3 Order of work
|
||||
|
||||
1. Owned descriptors (6.1.1) — invasive, unblocks everything, no user-visible change
|
||||
2. `DeclaredOp` + byte-identical parity test against the generated built-ins (6.1.2)
|
||||
3. Load-time WGSL validation and id namespacing (6.1.3, 6.1.4)
|
||||
4. A directory that is read at startup, and one shipped example that is not a built-in
|
||||
5. Pass nodes (FR-PLG-2a's second template), once 1–4 are load-bearing
|
||||
|
||||
Classes 2 and 3 — view plugins and computational plugins — are deliberately not in this plan.
|
||||
FR-PLG-3a's "a view plugin cannot be trusted with the UI thread" and FR-PLG-4a's capability grants
|
||||
are both larger design problems than class 1, and class 1 is where the requested features live.
|
||||
|
||||
---
|
||||
|
||||
## 7. What this document does not claim
|
||||
|
||||
Traceability counts a requirement as covered when a `TRACES` tag names it. It does not check that
|
||||
the code under the tag does the thing — `FR-DEV-8` is currently tagged against instance-buffer
|
||||
plumbing that a future spot-removal operation would use, and `FR-DEV-7` against a history row for a
|
||||
frontend that does not exist. Both read as covered.
|
||||
|
||||
So the 51% figure is an overstatement of unknown size, and closing FR-DSP by tagging what already
|
||||
works would make it a larger one. **Every requirement closed by this plan should be closed by a
|
||||
test that would fail if the behaviour were removed**, which is the only kind of coverage worth
|
||||
counting.
|
||||
@@ -0,0 +1,254 @@
|
||||
# DarkRoom — Distribution
|
||||
|
||||
**Satisfies:** NFR-COMPAT-2 (v1 channels) · FR-PLAT-LIN-3 (sandboxed distribution)
|
||||
**Companion to:** [requirements.md](requirements.md) §3.8, §4.8 · [storage.md](storage.md)
|
||||
|
||||
NFR-COMPAT-2 asks for the v1 channels to be *stated*, and says why in its own
|
||||
second sentence: the channel decision and the storage design are coupled. A
|
||||
channel is not a build target. It is a set of constraints that reach back into
|
||||
the code — what the application is allowed to see, what it may ask for, and
|
||||
what it must be able to do without asking. This document records which channels
|
||||
v1 targets and what each one costs, and it is where to look before adding a
|
||||
permission to a package rather than after.
|
||||
|
||||
---
|
||||
|
||||
## 1. The channels
|
||||
|
||||
| Platform | Channel | State | What it constrains |
|
||||
|---|---|---|---|
|
||||
| Linux | Arch source package — [`packaging/PKGBUILD`](../../packaging/PKGBUILD) | Built, in tree | Nothing. Full filesystem access, system Vulkan, system secret daemon |
|
||||
| Linux | Flatpak — [`packaging/flatpak/`](../../packaging/flatpak/) | Manifest in tree, **library selection does not work** (§4) | Portals only. No `--filesystem=`, no host mount table, no typed paths |
|
||||
| Linux | AppImage | v1 channel, **recipe not yet written** (§5) | Oldest supported glibc, and no sandbox at all |
|
||||
| Android | F-Droid | v1 channel, not yet submitted | GPLv3-clean build, reproducible, no proprietary blobs |
|
||||
| Android | Play Store | **Not v1** (§6) | Would make ARCH §6.9 binding as policy rather than as engineering |
|
||||
| Windows | NSIS per-user installer, cross-built — [windows.md](windows.md) | Built by CI, **untested on Windows**. Not v1 | Known folders in place of XDG; no sandbox; unsigned until there is a certificate |
|
||||
|
||||
Four of these six exist as recipes and two do not. That is stated rather than
|
||||
smoothed over, because the value of writing the channels down is knowing which
|
||||
constraints are already being met and which are promises.
|
||||
|
||||
### What every channel has to get right
|
||||
|
||||
Independent of packaging format, and each of these has bitten a package
|
||||
somewhere:
|
||||
|
||||
- **One identifier, four places.** `paris.tourolle.darkroom` is the AppStream
|
||||
component id, the `.desktop` basename, the Flatpak application id, and the
|
||||
string `dr_ui::run` sets as the Wayland `app_id` and X11 `WM_CLASS`. A rename
|
||||
that misses one of them costs the icon in the shell or the association in the
|
||||
software centre, and neither failure announces itself.
|
||||
- **The metainfo, not just the desktop entry.**
|
||||
[`packaging/paris.tourolle.darkroom.metainfo.xml`](../../packaging/paris.tourolle.darkroom.metainfo.xml)
|
||||
is the single description of the application, installed by every channel that
|
||||
has somewhere to put it. Its `metadata_license` is CC0-1.0 and its
|
||||
`project_license` is GPL-3.0-or-later; those differ on purpose — see the
|
||||
comment in the file.
|
||||
- **Vulkan is a requirement, not a preference.** The develop pipeline is
|
||||
compute shaders through wgpu, and NFR-R8 — how far a CPU fallback goes — is
|
||||
still open, so today there is nothing behind it. A package that installs onto
|
||||
a machine with no working ICD produces an application that starts and cannot
|
||||
develop.
|
||||
- **A Secret Service implementation, or an honest degraded mode.** FR-NC-2 is
|
||||
explicit that the absence of a secrets daemon is a stated degraded mode and
|
||||
never a silent fall back to plaintext. Packages express this as an optional
|
||||
dependency (the PKGBUILD) or a talk hole (the Flatpak manifest), never as a
|
||||
hard dependency — a headless or minimal-WM install is a supported way to run.
|
||||
- **The face models are Git LFS objects.** A checkout without `git lfs pull`
|
||||
has ~130-byte pointers where 11 MB models should be. Both the PKGBUILD and
|
||||
the Flatpak manifest check the file size and refuse, because the alternative
|
||||
is a package whose face indexing fails inside the graph loader on a user's
|
||||
machine rather than on the packager's.
|
||||
|
||||
---
|
||||
|
||||
## 2. Why Flatpak is the channel that matters most
|
||||
|
||||
Not because it is expected to be the most used. Because it is the only one that
|
||||
tests anything.
|
||||
|
||||
The Arch package and an AppImage both hand the application the same
|
||||
unrestricted process the developer runs it in, so neither can discover that a
|
||||
design assumed unrestricted access. Flatpak takes that assumption away, and
|
||||
FR-PLAT-LIN-3 exists to make the discovery happen deliberately rather than in a
|
||||
bug report. §4 is what it discovered.
|
||||
|
||||
The same argument runs the other way on Android, where SAF has been the only
|
||||
option since before the first line was written (ARCH §6.9) and `SourceRef`
|
||||
exists because of it. Linux got the abstraction — `LocalStorage::grant` is the
|
||||
one place a `Path` enters — and never got the constraint that would have proved
|
||||
it worked.
|
||||
|
||||
---
|
||||
|
||||
## 3. What already works inside the sandbox, unchanged
|
||||
|
||||
Worth listing, because it is the part FR-PLAT-LIN-1 quietly paid for in
|
||||
advance:
|
||||
|
||||
- **XDG directories.** Flatpak redirects `XDG_CONFIG_HOME`, `XDG_DATA_HOME` and
|
||||
`XDG_CACHE_HOME` into `~/.var/app/paris.tourolle.darkroom/`. Settings
|
||||
(`settings_store.rs`), accounts (`dr_sync::account`), the catalog and the
|
||||
thumbnail store all read those variables, so every one of them lands in the
|
||||
application's own directory with no code change and no permission.
|
||||
- **The face models.** `system_face_models_dirs()` reads `$XDG_DATA_DIRS`
|
||||
rather than hard-coding `/usr/share`, which is exactly why `/app/share`
|
||||
inside a Flatpak is found by the same lookup that finds the Arch package's
|
||||
copy.
|
||||
- **Opening a photograph from a file manager.** The `.desktop` entry declares
|
||||
the RAW MIME types and `Exec=darkroom-desktop %F`; under Flatpak the file is
|
||||
exported through the document portal and arrives in `argv` as a path under
|
||||
`/run/user/$UID/doc/`, which is mounted in every sandbox. `main.rs` takes
|
||||
paths from `argv` and `collect()` handles a file or a directory. This is
|
||||
genuine portal-mediated access and it needs nothing new.
|
||||
- **The Nextcloud sign-in browser.** `open_in_browser` spawns `xdg-open`; the
|
||||
freedesktop runtime's `xdg-open` forwards to the OpenURI portal, and portal
|
||||
calls need no `--talk-name` because Flatpak always permits them. FR-NC-1's
|
||||
"system browser, never an embedded webview" therefore holds inside the
|
||||
sandbox for the same reason it holds outside it.
|
||||
- **Credentials.** The keyring crate speaks the Secret Service D-Bus interface,
|
||||
reached through the session-bus proxy with one talk hole. The app password
|
||||
stays visible to `secret-tool` and Seahorse, which is what keeps it
|
||||
individually revocable by the user.
|
||||
|
||||
---
|
||||
|
||||
## 4. What does not work: choosing a library
|
||||
|
||||
**FR-PLAT-LIN-3 is not satisfied today, and the manifest does not pretend
|
||||
otherwise.**
|
||||
|
||||
A folder library is chosen by typing an absolute path. `dr-sync-folder`'s
|
||||
provider declares `SignIn::EndpointOnly` with the placeholder
|
||||
`/home/you/Pictures`, and `normalise_endpoint` expands `~`, requires the path
|
||||
to be absolute, and checks it with `std::fs`. Nothing in the tree calls the
|
||||
FileChooser portal — there is no `ashpd`, no `rfd`, and no toolkit file dialog
|
||||
anywhere in `ui/`, `platform/` or `core/`.
|
||||
|
||||
Inside a sandbox with no `--filesystem=`, `$HOME` still resolves to the real
|
||||
home *path* but that directory holds only the application's own
|
||||
`.var/app/…` tree. So a typed `~/Pictures` fails the `exists()` check and the
|
||||
launch screen says `No folder at /home/you/Pictures.` — a truthful message
|
||||
about a situation the user cannot fix from inside the application.
|
||||
|
||||
Import is blocked one step earlier. `dr_plat::volumes()` finds a camera card by
|
||||
reading `/proc/self/mountinfo` and the `removable` flag under `/sys`. A
|
||||
sandboxed process is in its own mount namespace, so the table it reads
|
||||
describes the sandbox; a card mounted at `/run/media/…` on the host is not in
|
||||
it. `volumes()` correctly returns an empty list, which the interface presents
|
||||
as "no card found" — right for the code, wrong for the user, who is looking at
|
||||
a card.
|
||||
|
||||
### The permission that would hide this, and why it is not in the manifest
|
||||
|
||||
`--filesystem=host` makes both work immediately and is the thing FR-PLAT-LIN-3
|
||||
names as the alternative to portals. Granting it would mean the sandboxed build
|
||||
never exercises the sandbox, which removes the entire reason for shipping one
|
||||
(§2). `--filesystem=xdg-pictures` is narrower and would be tempting, but it is
|
||||
still a static grant that lets a typed path resolve — it makes the same design
|
||||
work by not testing it, only in a smaller directory.
|
||||
|
||||
So the manifest grants no filesystem access at all. The consequence is stated
|
||||
plainly: **a Flatpak built from this manifest can open photographs handed to it
|
||||
and cannot yet be pointed at a library.**
|
||||
|
||||
### What closes it
|
||||
|
||||
Two changes, in this order:
|
||||
|
||||
1. **A portal file chooser behind a platform seam.** `ashpd`'s
|
||||
`OpenFileRequest` with `directory(true)` returns a URI the document portal
|
||||
has exported, which the sandbox can read and which stays valid across
|
||||
restarts. It resolves to a real path under `/run/user/$UID/doc/`, so
|
||||
`normalise_endpoint` accepts it as it stands — `canonicalize()` on a fuse
|
||||
path returns the path itself. The seam matters more than the crate: this
|
||||
belongs beside `LocalStorage::grant` in `dr-plat`, which is already the one
|
||||
place a `Path` enters the application, and must not become a second way for
|
||||
`ui/` to learn about paths.
|
||||
2. **Removable volumes through the same door.** There is no portal for "list
|
||||
the mounted cards". The honest answer is that under a sandbox
|
||||
`imports_supported()` should report the same `false` it reports on Android,
|
||||
for the same reason it gives there — the operation cannot be performed
|
||||
however hard the user tries — and the import flow should offer the folder
|
||||
chooser instead of a volume list.
|
||||
|
||||
**Done when:** a Flatpak built from
|
||||
[`packaging/flatpak/paris.tourolle.darkroom.yml`](../../packaging/flatpak/paris.tourolle.darkroom.yml),
|
||||
with its `finish-args` unchanged and no `flatpak override` applied, can select a
|
||||
library root, scan it, and write a sidecar back into it.
|
||||
|
||||
### Running a Flatpak build before then
|
||||
|
||||
For testing the rest of the application inside the sandbox, grant the access
|
||||
per-installation rather than in the manifest, so the file that describes the
|
||||
application keeps telling the truth:
|
||||
|
||||
```bash
|
||||
flatpak override --user --filesystem=~/Pictures paris.tourolle.darkroom
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 5. AppImage
|
||||
|
||||
A v1 channel, and the recipe is outstanding work rather than a decision to be
|
||||
made. What it will have to account for, none of which is a surprise:
|
||||
|
||||
- **glibc.** An AppImage links against the oldest glibc it must run on, so it
|
||||
is built in a container with an old base rather than on a rolling-release
|
||||
developer machine. A release binary built on a current rolling-release host carries
|
||||
`GLIBC_2.44` references and would run on almost nothing else.
|
||||
- **What to bundle and what not to.** The binary links fontconfig, freetype,
|
||||
expat, libpng, zlib, brotli and bzip2 — bundle those. It does *not* link
|
||||
Vulkan, libxkbcommon or either display-server library: wgpu `dlopen`s
|
||||
`libvulkan.so.1`, and `x11rb` and `wayland-client` speak the wire protocols
|
||||
in Rust. The Vulkan loader and the ICD must come from the host, and bundling
|
||||
a loader is the classic way to break an AppImage on a driver it did not
|
||||
expect.
|
||||
- **The models.** ~15 MB of ONNX weights inside the image, or a first-run
|
||||
download. In-tree is consistent with how the Lensfun database ships and with
|
||||
NFR-SEC-5's local-first posture; the licence question (D13) is the same one
|
||||
it is everywhere else and is not made easier or harder by this channel.
|
||||
- **No sandbox.** An AppImage tests nothing about FR-PLAT-LIN-3. It is a
|
||||
convenience channel for distributions the PKGBUILD does not serve, and should
|
||||
never be the channel a portal problem is discovered on.
|
||||
|
||||
---
|
||||
|
||||
## 6. Android: F-Droid in v1, Play deferred
|
||||
|
||||
NFR-COMPAT-2 says Play distribution is what makes ARCH §6.9's constraints
|
||||
binding, and that is worth reading precisely, because the constraint is already
|
||||
met and would be met whatever the channel.
|
||||
|
||||
§6.9 is *verified*, not assumed: `MANAGE_EXTERNAL_STORAGE` is not grantable
|
||||
under Play policy, and `READ_MEDIA_IMAGES` would not help because proprietary
|
||||
RAW is not typed `image/*` by the platform scanner and does not appear in
|
||||
`MediaStore.Images`. SAF is the only route that works, so FR-PLAT-AND-1 asks
|
||||
for it unconditionally and `SourceRef` (ARCH §3.1) exists to make it possible.
|
||||
A sideloaded or F-Droid build *could* ask for broader permissions; it would
|
||||
gain nothing by doing so.
|
||||
|
||||
So the coupling runs the opposite way from how it is usually described. Play is
|
||||
deferred for a reason that has nothing to do with storage: GPLv3 distribution
|
||||
through Play is generally workable but has not been confirmed for this project
|
||||
(ARCH §14), and F-Droid has no such question. Confirming it is a licence-reading
|
||||
exercise; nothing in the storage design waits on the answer.
|
||||
|
||||
---
|
||||
|
||||
## 7. Where the recipes live
|
||||
|
||||
```
|
||||
packaging/
|
||||
PKGBUILD Arch source package
|
||||
paris.tourolle.darkroom.desktop the desktop entry, installed by every channel
|
||||
paris.tourolle.darkroom.metainfo.xml AppStream, installed by every channel
|
||||
flatpak/
|
||||
paris.tourolle.darkroom.yml the manifest, and where the permissions are argued
|
||||
windows/
|
||||
darkroom.nsi the installer; docker/windows/package.sh drives it
|
||||
```
|
||||
|
||||
`packaging/` also accumulates built `.pkg.tar.zst` artefacts from local
|
||||
`makepkg` runs. Those are not part of any channel and should not be committed.
|
||||
+1579
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,433 @@
|
||||
# What a frame costs
|
||||
|
||||
**Status:** Measured · 2026-08-27
|
||||
**Companion to:** [display-and-extension.md](display-and-extension.md) §2–3 ·
|
||||
[requirements.md](requirements.md) §3.4 FR-DSP-2, FR-DSP-3, FR-DSP-4
|
||||
**Instrument:** [`core/dr-gpu/examples/frame_budget.rs`](../../core/dr-gpu/examples/frame_budget.rs)
|
||||
**Guard:** [`core/dr-gpu/tests/frame_budget.rs`](../../core/dr-gpu/tests/frame_budget.rs)
|
||||
|
||||
[display-and-extension.md](display-and-extension.md) §2 fixed a decision rule in
|
||||
advance and made three measurements the thing that settles it. This file is
|
||||
those measurements, and the recommendation they support.
|
||||
|
||||
Rerun with:
|
||||
|
||||
```sh
|
||||
cargo run --release -p dr-gpu --example frame_budget
|
||||
```
|
||||
|
||||
and diff this file. That is the whole point of committing numbers: a regression
|
||||
should be a diff rather than somebody's recollection of how fast it used to be.
|
||||
|
||||
---
|
||||
|
||||
## The answer, first
|
||||
|
||||
**FR-DSP-2 should be rewritten, not implemented.** M1 and M2 sit inside the
|
||||
16 ms budget at the 99th percentile for every chain of point operations at every
|
||||
viewport size measured, fit and at 1:1 — the widest case, every operation that
|
||||
contributes a fragment to the fused shader at 4K, costs **4.5 ms** on the GPU and
|
||||
**8.2 ms** including the composition that precedes it. Tiling the interactive
|
||||
path would be optimising something that is already using a quarter of its budget.
|
||||
|
||||
**But the measurement did find a budget-breaker, and it is not the one tiling
|
||||
fixes.** The neighbourhood stage — clarity in particular — costs **34 ms at 4K
|
||||
on its own**, twice the whole budget, and tiles do not help it: a tile of a
|
||||
convolution has to read its halo, so tiling raises the total tap count rather
|
||||
than lowering it. §2 predicted this exactly ("a separable blur at a large radius
|
||||
is the plausible budget-breaker, not the fused pass"), and the fix it needs is
|
||||
the one `local_contrast`'s own module documentation already names — a base
|
||||
computed at reduced resolution — which is a change to `crate::detail`, not a
|
||||
tile scheduler.
|
||||
|
||||
There is a third finding nobody was looking for: **shader composition costs
|
||||
3–5 ms of CPU per frame on a full chain**, on the UI thread, before any GPU work
|
||||
is submitted. That is a fifth to a third of the budget spent formatting strings,
|
||||
and it is invisible to any amount of tiling.
|
||||
|
||||
---
|
||||
|
||||
## Conditions
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| Adapter | NVIDIA GeForce RTX 3050 6GB Laptop GPU (Vulkan) |
|
||||
| Source | 9504 × 6336 synthetic (60.2 MP, 482 MB as `rgba16f`) |
|
||||
| Frames | 100 measured per row, 12 warm-up frames discarded |
|
||||
| Percentile | Nearest-rank, so p99 of 100 frames is the second-worst frame |
|
||||
| Build | `--release` |
|
||||
| Date | 2026-08-27 |
|
||||
|
||||
`shader` is `EditGraph::compose` alone. `cpu` adds the detail chain and the
|
||||
invalidation hash — everything `DevelopSession::render` does per frame before it
|
||||
dispatches. `gpu` is submit plus wait-for-idle, which serialises the GPU work
|
||||
into the frame that caused it and is therefore pessimistic. `TOTAL` ranks
|
||||
`cpu + gpu` summed **within each frame**, which is the column the budget is
|
||||
judged on; adding two percentiles instead would invent a stutter that no frame
|
||||
actually had.
|
||||
|
||||
Chains: `one` is exposure. `five` is exposure, contrast, highlights/shadows,
|
||||
blacks/whites, vibrance. `point` is every operation in the default chain that
|
||||
contributes a fragment to the fused shader, film stock included. `all` is `point`
|
||||
plus the four neighbourhood operations — noise reduction, capture sharpening,
|
||||
clarity and texture.
|
||||
|
||||
---
|
||||
|
||||
## M1 — the fused pass at proxy resolution
|
||||
|
||||
The develop view: the whole frame fit to the viewport.
|
||||
|
||||
| size | chain | shader | cpu p99 | gpu p50 | gpu p99 | TOTAL | |
|
||||
|------------:|------:|-------:|--------:|--------:|--------:|--------:|:-----|
|
||||
| 1920 × 1200 | one | 0.08ms | 0.10ms | 1.02ms | 1.23ms | 1.31ms | |
|
||||
| 1920 × 1200 | five | 0.18ms | 0.20ms | 1.01ms | 1.20ms | 1.36ms | |
|
||||
| 1920 × 1200 | point | 2.79ms | 2.82ms | 1.98ms | 2.18ms | 4.83ms | |
|
||||
| 1920 × 1200 | all | 3.65ms | 4.73ms | 6.86ms | 7.37ms | 12.02ms | |
|
||||
| 2560 × 1600 | one | 0.08ms | 0.10ms | 1.73ms | 2.00ms | 2.12ms | |
|
||||
| 2560 × 1600 | five | 0.24ms | 0.27ms | 1.73ms | 2.26ms | 2.46ms | |
|
||||
| 2560 × 1600 | point | 2.82ms | 2.85ms | 2.37ms | 2.65ms | 5.38ms | |
|
||||
| 2560 × 1600 | all | 3.37ms | 4.35ms | 14.31ms | 15.65ms | 18.42ms | OVER |
|
||||
| 3840 × 2160 | one | 0.10ms | 0.14ms | 3.09ms | 3.31ms | 3.42ms | |
|
||||
| 3840 × 2160 | five | 0.30ms | 0.32ms | 3.03ms | 3.40ms | 3.61ms | |
|
||||
| 3840 × 2160 | point | 3.62ms | 3.65ms | 4.12ms | 4.52ms | 8.23ms | |
|
||||
| 3840 × 2160 | all | 4.12ms | 5.07ms | 37.73ms | 40.17ms | 43.24ms | OVER |
|
||||
|
||||
Read the `point` rows: **the fused dispatch scales with pixels and almost not at
|
||||
all with chain length.** Going from one operation to the entire point chain at
|
||||
4K costs 1.2 ms of GPU. Going from 2.3 M pixels to 8.3 M costs 2.3 ms. Both are
|
||||
small, and the second is the one tiling would address.
|
||||
|
||||
The `all` rows go over, and the `point` rows in the same block are what say why:
|
||||
the difference between them is the neighbourhood stage, measured on its own in
|
||||
M3 and arriving at almost exactly the same figure.
|
||||
|
||||
## M2 — the same, zoomed to 1:1 on the 60 MP source
|
||||
|
||||
FR-DSP-5's case. `Framing::view` shrinks the sampled region while the render
|
||||
target keeps its size, so one render pixel lands on one source pixel.
|
||||
|
||||
| size | chain | shader | cpu p99 | gpu p50 | gpu p99 | TOTAL | |
|
||||
|------------:|------:|-------:|--------:|--------:|--------:|--------:|:-----|
|
||||
| 1920 × 1200 | one | 0.12ms | 0.14ms | 0.42ms | 0.66ms | 0.75ms | |
|
||||
| 1920 × 1200 | five | 0.27ms | 0.30ms | 0.49ms | 1.14ms | 1.22ms | |
|
||||
| 1920 × 1200 | point | 3.49ms | 3.52ms | 1.18ms | 1.39ms | 4.85ms | |
|
||||
| 1920 × 1200 | all | 3.64ms | 5.18ms | 8.96ms | 9.55ms | 14.30ms | |
|
||||
| 2560 × 1600 | one | 0.11ms | 0.12ms | 0.56ms | 0.99ms | 1.06ms | |
|
||||
| 2560 × 1600 | five | 0.22ms | 0.25ms | 0.77ms | 1.02ms | 1.17ms | |
|
||||
| 2560 × 1600 | point | 2.96ms | 2.99ms | 2.03ms | 2.52ms | 5.61ms | |
|
||||
| 2560 × 1600 | all | 5.24ms | 7.35ms | 18.80ms | 21.62ms | 25.81ms | OVER |
|
||||
| 3840 × 2160 | one | 0.10ms | 0.12ms | 1.31ms | 1.52ms | 1.63ms | |
|
||||
| 3840 × 2160 | five | 0.14ms | 0.27ms | 1.39ms | 1.64ms | 1.75ms | |
|
||||
| 3840 × 2160 | point | 3.14ms | 3.16ms | 4.04ms | 4.50ms | 7.21ms | |
|
||||
| 3840 × 2160 | all | 4.84ms | 6.78ms | 47.22ms | 48.79ms | 54.47ms | OVER |
|
||||
|
||||
**A 1:1 view of a 60 MP file is cheaper than the fit view of the same file**, for
|
||||
every point chain and at every size — 1.52 ms against 3.31 ms for one operation
|
||||
at 4K. That is not a rounding artefact and it is worth stating plainly, because
|
||||
it is the opposite of what "full resolution" sounds like it should cost. The
|
||||
dispatch is the same number of pixels either way; what changes is where those
|
||||
pixels read from. A fit view walks the whole 482 MB texture on a stride, and a
|
||||
1:1 view reads a contiguous window of it that fits comfortably in cache.
|
||||
|
||||
So the resolution FR-DSP-5 promises costs nothing extra on the fused path.
|
||||
Zooming is not an expensive mode to be dreaded and progressively refined into;
|
||||
it is the cheap one.
|
||||
|
||||
The `all` rows are worse at 1:1 than fit, and that is the detail stage again for
|
||||
a specific reason: noise reduction's radius is stated in *source* pixels, so
|
||||
`RenderScale::ratio` climbing to 1.0 widens its kernel. Clarity's is stated as a
|
||||
fraction of the frame and does not move. M3 separates the two.
|
||||
|
||||
## M3 — the neighbourhood stage alone
|
||||
|
||||
Timed with the fused dispatch deliberately reused: only a detail parameter moves,
|
||||
so `render_detailed` skips the colour pass (FR-DEV-3d) and what remains is the
|
||||
convolutions. `colour` counts fused dispatches over the measured frames and is
|
||||
zero on every row, which is what makes these numbers mean "detail alone" rather
|
||||
than asserting it.
|
||||
|
||||
| size | stage | view | pass | radius | colour | cpu p99 | p50 | p99 |
|
||||
|------------:|---------:|:-----|-----:|-------:|-------:|--------:|--------:|--------:|
|
||||
| 1920 × 1200 | clarity | fit | 2 | 29 | 0 | 1.60ms | 5.42ms | 5.99ms |
|
||||
| 1920 × 1200 | all four | fit | 7 | 29 | 0 | 1.71ms | 5.78ms | 6.16ms |
|
||||
| 1920 × 1200 | clarity | 1:1 | 2 | 29 | 0 | 1.12ms | 7.51ms | 8.01ms |
|
||||
| 1920 × 1200 | all four | 1:1 | 9 | 29 | 0 | 2.71ms | 8.25ms | 9.11ms |
|
||||
| 2560 × 1600 | clarity | fit | 2 | 38 | 0 | 1.03ms | 12.02ms | 12.44ms |
|
||||
| 2560 × 1600 | all four | fit | 7 | 38 | 0 | 1.87ms | 12.49ms | 13.16ms |
|
||||
| 2560 × 1600 | clarity | 1:1 | 2 | 38 | 0 | 1.87ms | 15.82ms | 16.60ms |
|
||||
| 2560 × 1600 | all four | 1:1 | 9 | 38 | 0 | 2.71ms | 17.24ms | 18.06ms |
|
||||
| 3840 × 2160 | clarity | fit | 2 | 52 | 0 | 1.75ms | 33.11ms | 33.89ms |
|
||||
| 3840 × 2160 | all four | fit | 7 | 52 | 0 | 1.76ms | 34.21ms | 35.03ms |
|
||||
| 3840 × 2160 | clarity | 1:1 | 2 | 52 | 0 | 1.08ms | 40.39ms | 41.86ms |
|
||||
| 3840 × 2160 | all four | 1:1 | 9 | 52 | 0 | 2.37ms | 43.29ms | 44.72ms |
|
||||
|
||||
`radius` is the widest halo any pass reads, in render pixels.
|
||||
|
||||
Clarity alone is 97% of the cost of all four neighbourhood operations together,
|
||||
at every size. Its σ is 1.2% of the shorter edge and it truncates at 2σ, so its
|
||||
radius is 29 px on a 1200 px viewport and **52 px at 4K** — two separable passes
|
||||
of 105 taps each, over 8.3 M pixels, which is 1.7 billion texture reads. That is
|
||||
the whole of the problem, and the numbers scale as `radius × pixels` exactly as
|
||||
that description predicts: 5.99 → 12.44 → 33.89 ms for radii of 29 → 38 → 52 over
|
||||
2.3 → 4.1 → 8.3 M pixels.
|
||||
|
||||
The extra cost at 1:1 is noise reduction and capture sharpening, whose radii are
|
||||
properties of the sensor rather than of the frame. That is the correct behaviour
|
||||
— it is why `RenderScale` has two units — and it is bounded by the kernel caps
|
||||
those operations already declare.
|
||||
|
||||
---
|
||||
|
||||
## Reading this against §2's decision rule
|
||||
|
||||
§2: *"If M1 and M2 sit inside 16 ms at the 99th percentile, FR-DSP-2 is
|
||||
rewritten rather than implemented … If they do not, the measurement tells us
|
||||
which stage to tile."*
|
||||
|
||||
Both halves of the rule fire, on different stages, and the honest reading takes
|
||||
both.
|
||||
|
||||
### FR-DSP-2 — rewrite it
|
||||
|
||||
For the fused pass the rule passes with a wide margin. Every point chain at
|
||||
every size, fit and at 1:1, is inside 16 ms — the worst `TOTAL` is 8.23 ms and
|
||||
the worst GPU figure is 4.52 ms. There is no viewport size on a desktop display
|
||||
where recomputing the entire point chain over every visible pixel is a problem.
|
||||
|
||||
Two further reasons not to build the tile scheduler as written:
|
||||
|
||||
1. **Panning, which is the case ARCH §5.3's tile cache is designed for, gets no
|
||||
benefit here.** Reusing already-valid tiles saves recomputation. Recomputing
|
||||
the whole 4K viewport costs 4.5 ms, so a perfect tile cache could save at most
|
||||
4.5 ms of a 16 ms budget, at the price of a cache keyed by
|
||||
`(VersionId, tile, zoom, graph_hash_prefix)` that has to stay correct across
|
||||
every parameter change in the graph. That is a large correctness surface
|
||||
bought with a small number.
|
||||
|
||||
2. **It would make the actual problem worse.** The stage that misses the budget
|
||||
is a convolution, and a tiled convolution reads a halo per tile. At a 52-pixel
|
||||
radius, 256-pixel tiles would read (256+104)² instead of 256² — very nearly
|
||||
*twice* the taps. Tiling is the wrong tool for the one stage that needs a
|
||||
tool.
|
||||
|
||||
So FR-DSP-2 becomes what §2 said it actually is for this architecture: a
|
||||
scheduling concern for export and thumbnailing, both of which already run off
|
||||
the frame path. The interactive path does not tile.
|
||||
|
||||
### The stage that did need work — and it is not tiling
|
||||
|
||||
**Resolved.** The fix described below landed; the measurement is in
|
||||
§[The reduced base, measured](#the-reduced-base-measured) at the foot of this
|
||||
file, and `docs/technical-debt.md` TD-4 is closed. What follows is the
|
||||
reasoning as it stood, kept because it is what the numbers above argue for and
|
||||
because the tiling half of it is still live.
|
||||
|
||||
|
||||
The measurement's real product is naming the stage. It is `local_contrast`, and
|
||||
the fix is stated in that module's own documentation:
|
||||
|
||||
> The right optimisation is a base computed at reduced resolution, which needs a
|
||||
> detail stage that can write a smaller target than it reads; that is a change to
|
||||
> `crate::detail`, not to this file.
|
||||
|
||||
A Gaussian base at a quarter resolution is 1/16 the pixels at 1/4 the radius —
|
||||
about 1/64 of the work — and the result is visually identical because a base at
|
||||
σ = 26 px has no content above the quarter-resolution Nyquist to lose. That is a
|
||||
change to two files with a bounded blast radius, and it is what the 34 ms buys
|
||||
back. It should be tracked as its own item rather than smuggled in under a
|
||||
requirement about tiles.
|
||||
|
||||
### FR-DSP-3 — the clause that should be narrowed
|
||||
|
||||
§3.3 proposes narrowing "when a full-resolution result is needed it is computed
|
||||
asynchronously, and the proxy result remains on screen until it is ready" to
|
||||
export and 1:1 zoom, or striking it.
|
||||
|
||||
**M2 says strike it.** The clause exists to hide the latency of a
|
||||
full-resolution render behind a proxy. There is no such latency: the 1:1 view is
|
||||
*faster* than the fit view on the fused path, and there is no second
|
||||
full-resolution code path to be asynchronous about — `Framing::view` is the
|
||||
whole mechanism. Export renders its own frames on a worker already. Keeping the
|
||||
clause would mean building a progressive-swap machine to conceal a render that
|
||||
completes in 1.4 ms.
|
||||
|
||||
### FR-DSP-4 — satisfied vacuously, on the fused path
|
||||
|
||||
§4 makes progressive refinement conditional on M1 failing. On the fused path M1
|
||||
passes, so reduced-quality rendering during a drag would buy nothing and cost the
|
||||
visible softness the requirement itself warns against.
|
||||
|
||||
The neighbourhood stage is the exception, and it is worth being precise: what
|
||||
that stage needs is not *progressive* refinement — it is a permanently cheaper
|
||||
base, computed at reduced resolution and correct at any moment the user stops.
|
||||
"Render coarse while dragging, sharpen when it settles" would paper over the same
|
||||
34 ms with a visible swap. Fix the stage.
|
||||
|
||||
---
|
||||
|
||||
## Which GPU, on a machine with more than one
|
||||
|
||||
**Measured 2026-08-29** on a laptop holding an Intel Iris Xe (RPL-P) and an AMD
|
||||
RX 5700 XT, same binary, adapter forced with `VK_ICD_FILENAMES`.
|
||||
|
||||
The question was whether an integrated GPU is the better choice for this
|
||||
application. The argument for it is good: a 24 MP frame is ~96 MB of RGBA, and
|
||||
on a discrete card every upload and every export readback crosses PCIe, where
|
||||
an iGPU shares memory with the CPU and crosses nothing. It also does not empty
|
||||
a battery.
|
||||
|
||||
The compute says otherwise, and not marginally.
|
||||
|
||||
| 2560×1600, p99 | AMD RX 5700 XT | Intel Iris Xe |
|
||||
|---|---|---|
|
||||
| fused pass, `point` | 5.19 ms | 7.75 ms |
|
||||
| fused pass, `all` | 11.70 ms | **66.42 ms** |
|
||||
| M3 clarity, fit | 4.67 ms | **38.28 ms** |
|
||||
| M3 all four, 1:1 | 6.44 ms | **57.67 ms** |
|
||||
|
||||
| 1920×1200, M3 clarity, fit | 2.35 ms | **19.99 ms** |
|
||||
|---|---|---|
|
||||
|
||||
The fused colour pass is within a factor of 1.5 — it is one read and one write
|
||||
per pixel, which an iGPU does perfectly well. The **neighbourhood stage is
|
||||
5–8× slower**, and that is what decides it: clarity at 1920×1200 costs 20 ms on
|
||||
the Iris Xe, so it leaves the budget on its own at the smallest size tested,
|
||||
before anything else in the chain runs.
|
||||
|
||||
**So the default adapter preference stays `Performance`** (`dr_gpu::AdapterPreference`).
|
||||
|
||||
Two things this does *not* show, and neither is a reason to revisit the default
|
||||
without measuring them:
|
||||
|
||||
- **It does not refute the transfer argument.** This harness renders from a
|
||||
resident texture and never uploads or reads back, so the PCIe cost an iGPU
|
||||
avoids does not appear in any column above. Import, export and the thumbnail
|
||||
sweeps are transfer-heavy and compute-trivial, and may well go the other way
|
||||
— but they are not what FR-DSP-3 bounds, and one device is opened at startup
|
||||
and shared with the compositor, so there is currently no way to use a
|
||||
different adapter for a different task.
|
||||
- **It says nothing about power.** `Efficiency` remains offered
|
||||
(`DARKROOM_GPU=integrated`) because a user on battery may rationally accept a
|
||||
slower detail chain, and because someone whose discrete card has failed needs
|
||||
a way to keep working.
|
||||
|
||||
## What is not measured here
|
||||
|
||||
Stated because §7 of [display-and-extension.md](display-and-extension.md) asks
|
||||
for it, and because each of these could move the numbers.
|
||||
|
||||
- **Local adjustments.** The mask stack is a separate chain per layer and is not
|
||||
in any row above. `render_masked` takes them and the fused shader addresses
|
||||
them per layer, so a heavily masked edit costs more than `all`.
|
||||
- **Spot repairs.** These add detail passes, and their cost is per spot.
|
||||
- **Lens corrections.** Not part of `EditGraph::default_chain` — they are built
|
||||
from a matched profile — so the `point` row does not include the warp chain.
|
||||
- **Demosaic.** Once per photograph on a worker, not on the frame path.
|
||||
- **Presentation.** The bench waits for the device to go idle inside the frame it
|
||||
measures. A real compositor overlaps frames, so these figures are an upper
|
||||
bound rather than an estimate.
|
||||
- **One adapter.** A discrete laptop GPU. The Intel iGPU on the same machine, and
|
||||
Android, will be slower — which is an argument for the conclusion rather than
|
||||
against it: the stage with no headroom has none to lose.
|
||||
|
||||
## The CPU finding, which deserves its own item
|
||||
|
||||
`EditGraph::compose` costs 2.8–5.2 ms per frame on a full chain, at every
|
||||
resolution, because it is resolution-independent: it assembles a WGSL string and
|
||||
hashes it. On the `all` rows it is a third of what is left of the budget after
|
||||
the GPU has taken its share, and at 1920 × 1200 it is larger than the entire
|
||||
fused dispatch.
|
||||
|
||||
Nothing in this document's recommendations changes it, and it is the cheapest
|
||||
remaining win. The generated *source* depends only on the structure of the graph
|
||||
— that is what `structure_hash` already identifies, and it is precisely what does
|
||||
not change while a slider is being dragged, which is why the pipeline cache in
|
||||
`AdjustPass` does not recompile. The uniforms do change, but assembling them is a
|
||||
handful of floats per operation. So caching the source string against the
|
||||
structure hash and rebuilding only the uniforms would take these milliseconds to
|
||||
approximately nothing, on the path that needs them most. Worth its own entry in
|
||||
[technical-debt.md](technical-debt.md).
|
||||
|
||||
---
|
||||
|
||||
## The reduced base, measured
|
||||
|
||||
**Status:** Measured · 2026-08-29 · closes TD-4
|
||||
|
||||
`DetailPass` gained an `output_scale`, and clarity's base is now computed on a
|
||||
grid a quarter the size on each axis — the change §M3 argued for above.
|
||||
|
||||
**Read this table on its own, not against the ones above.** It was taken on a
|
||||
different adapter, so the absolute figures are not comparable with the RTX 3050
|
||||
measurements this document is otherwise built from. What *is* comparable is the
|
||||
before and the after, which were measured on the same machine, same card, same
|
||||
release profile, minutes apart, with nothing between them but the change — the
|
||||
baseline at `0407fb8` and the result at `bff95e2`.
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| Adapter | AMD Radeon RX 5700 XT (RADV NAVI10) (Vulkan) |
|
||||
| Source | 9504 × 6336 (60.2 MP, 482 MB as `rgba16f`) |
|
||||
| Baseline | `0407fb8`, the branch's merge-base |
|
||||
| Result | `bff95e2` |
|
||||
| Date | 2026-08-29 |
|
||||
|
||||
### M3 — clarity alone, before and after
|
||||
|
||||
Both percentiles, because they disagree and the disagreement is the
|
||||
interesting part.
|
||||
|
||||
| size | view | before p50 | after p50 | | before p99 | after p99 |
|
||||
|------------:|:-----|-----------:|----------:|-----:|-----------:|----------:|
|
||||
| 1920 × 1200 | fit | 2.30ms | 1.56ms | 1.5× | 3.94ms | 1.97ms |
|
||||
| 1920 × 1200 | 1:1 | 2.88ms | 1.99ms | 1.4× | 3.07ms | 2.41ms |
|
||||
| 2560 × 1600 | fit | 4.41ms | 1.95ms | 2.3× | 10.40ms | 2.37ms |
|
||||
| 2560 × 1600 | 1:1 | 5.59ms | 3.32ms | 1.7× | 5.78ms | 3.85ms |
|
||||
| 3840 × 2160 | fit | 10.94ms | 3.88ms | 2.8× | 25.05ms | 4.17ms |
|
||||
| 3840 × 2160 | 1:1 | 13.31ms | 6.36ms | 2.1× | 27.60ms | 6.93ms |
|
||||
|
||||
**The honest headline is the p50 column: 2.8× at 4K.** An earlier draft of this
|
||||
section led with the p99 ratio, which reads as 6.0× at the same size. That
|
||||
number is not supported, and the reason it is not is worth recording rather
|
||||
than quietly deleting.
|
||||
|
||||
The baseline run's `fit` rows have a p99/p50 spread of about 2.3×, while every
|
||||
row of the after run sits between 1.07× and 1.26×. A stage whose cost is
|
||||
`radius × pixels` has no reason to be bimodal, and the `fit` configuration is
|
||||
the memory-bound one — it walks the whole 482 MB source on a stride, where
|
||||
`1:1` reads a contiguous window. Something else was using the machine.
|
||||
|
||||
The cross-check settles it. §"Which GPU, on a machine with more than one"
|
||||
above measured the *same baseline code on the same card* independently, and
|
||||
reports M3 clarity, fit, 2560 × 1600 at **4.67 ms p99** — against the 10.40 ms
|
||||
in the table here. Two measurements of one thing that differ by 2.2× mean the
|
||||
noisier one is wrong, and it is this one.
|
||||
|
||||
So: the p50 ratios are the claim. The p99 improvement is real and larger, but
|
||||
this run cannot say by how much, and a clean re-measurement on a quiet machine
|
||||
is the way to find out.
|
||||
|
||||
What survives the caveat intact is the **shape** of the after column. Every
|
||||
figure is inside the 16 ms budget with a p99 within 26% of its median, at every
|
||||
size and both views — which is what a stage that is no longer the bottleneck
|
||||
looks like, whatever the exact ratio to what it replaced.
|
||||
|
||||
### The one thing that is not a pure speed-up
|
||||
|
||||
**The declared halo is now quantised to multiples of `output_scale`.** The
|
||||
kernel truncates at 2σ, and that rounding now happens on the reduced grid
|
||||
before being multiplied back up:
|
||||
|
||||
| viewport | before | after |
|
||||
|---|---:|---:|
|
||||
| 1920 × 1200 | 29 px | 28 px |
|
||||
| 2560 × 1600 | 38 px | 40 px |
|
||||
| 3840 × 2160 | 52 px | 52 px |
|
||||
|
||||
At 2σ the Gaussian is already down to `e⁻²` of its peak, and
|
||||
`crossing_the_reduction_threshold_does_not_change_the_picture` holds the
|
||||
difference between a quarter-scale and a half-scale base to 0.03 stops of peak
|
||||
excursion and 2% of frame reach. But it is a change in reach rather than only
|
||||
in cost, it is what a tile scheduler would be handed, and it is worth knowing
|
||||
that the number moved rather than discovering it later as a seam.
|
||||
@@ -0,0 +1,484 @@
|
||||
# Inference backends — the runtime and the model, chosen per device
|
||||
|
||||
Spec for **S16**, the build that puts the neural models on the hardware each device actually has.
|
||||
|
||||
Every model DarkRoom runs today — the three SCRFD detectors, the ArcFace embedder, YOLO26n-seg and
|
||||
the ADE20K scene model — runs through `tract`, on one CPU core, on every platform. That was the
|
||||
right first answer: D13's runtime half chose it because it costs no C dependency, and
|
||||
[faces.md](faces.md) and [segmentation.md](segmentation.md) were written against it. It is also
|
||||
between 20× and 300× slower than what the same devices can do, and this document is the record of
|
||||
having measured that and the specification of what replaces it.
|
||||
|
||||
**It does not reopen D13's licensing half.** The weights are the same files under the same grant.
|
||||
It does reopen the *runtime* half, and §3 is where it says how far.
|
||||
|
||||
---
|
||||
|
||||
## 1. What was measured · 2026-09-19
|
||||
|
||||
One benchmark, two builds of it: `ort`'s API over `tract` (exactly what the app links) and `ort`'s
|
||||
API over a dynamically loaded ONNX Runtime with each execution provider in turn. Random 640×640
|
||||
input, three warm-ups, the median of 15–30 timed runs, milliseconds. The same input every run, so
|
||||
the numbers are compute cost and nothing else.
|
||||
|
||||
### 1.1 The tablet — Honor MagicPad 2, Snapdragon 8s Gen 3
|
||||
|
||||
SM8635: 1× Cortex-X4, 4× A720, 3× A520, Adreno 735, Hexagon V73. Android 16.
|
||||
|
||||
| Model | **tract** (today) | ORT CPU f32 | ORT CPU int8 | Adreno f32 ¹ | **Hexagon int8** ² |
|
||||
|---|---|---|---|---|---|
|
||||
| scrfd_500m (Fast) | 98 | 16 | 8 | 21 | **1.4** |
|
||||
| scrfd_2.5g (Balanced) | 161 | 59 | 19 | ✗ | **1.8** |
|
||||
| scrfd_10g (Thorough) | 489 | 204 | 48 | ✗ | **3.2** |
|
||||
| arcface_mbf (per face) | 39 | 9 | 13 | 24 | 12 |
|
||||
| yolo26n-seg | 287 | 94 | 38 | 54 | **5.1** |
|
||||
| yolo26s-sem-ade20k | 408 | 154 | 46 | 47 | **3.7** |
|
||||
|
||||
¹ Qualcomm's own GPU backend (`libQnnGpu.so`, OpenCL). Fails on the two larger SCRFD graphs at an
|
||||
`AveragePool` the layout transformer cannot place. ONNX Runtime's WebGPU provider also runs on this
|
||||
GPU and was slower than the CPU on every model; it is not in the table because it is not a
|
||||
candidate.
|
||||
² QNN's HTP backend. The Hexagon **refuses float32 and float16 tensors** in this ORT 1.29 + QNN
|
||||
2.42 pairing (error 3110 on every node, with `enable_htp_fp16_precision` set or not); int8 QDQ
|
||||
graphs run with 99.6% of nodes on the NPU — 1718 of 1725 for SCRFD-500m, the remainder being the
|
||||
quantise/dequantise at the graph's edges — verified from the partition log, not inferred from the
|
||||
timing.
|
||||
|
||||
Also tried and rejected: **NNAPI** — the device registers no neural-networks HAL at all, so the
|
||||
provider has nothing to talk to; Google deprecated it in Android 15 and Qualcomm stopped shipping
|
||||
drivers for it. **XNNPACK** — slower than ORT's default CPU kernels on every model that loaded, and
|
||||
aborts inside its partitioner on the SCRFD graphs.
|
||||
|
||||
### 1.2 The desktop — RTX 3050 Laptop, Raptor Lake, 20 threads
|
||||
|
||||
| Model | **tract** (today) | ORT CPU f32 | ORT CPU int8 | CUDA f32 | CUDA fp16 | TensorRT f32 | TensorRT fp16 | TensorRT int8 |
|
||||
|---|---|---|---|---|---|---|---|---|
|
||||
| scrfd_500m | 104 | 12 | 8 | 5.6 | 3.7 | 2.5 | **1.8** | ✗ ⁴ |
|
||||
| scrfd_2.5g | 162 | 27 | 12 | 6.2 | 5.3 | 2.9 | **1.9** | ✗ ⁴ |
|
||||
| scrfd_10g | 514 | 99 | 32 | 14.5 | 9.3 | 7.6 | **3.3** | ✗ ⁴ |
|
||||
| arcface_mbf | 45 | 15 | 18 | 1.1 | 0.8 | 0.9 | 0.7 | ✗ ⁴ |
|
||||
| yolo26n-seg | 307 | 60 | 38 | 9.3 | ✗ ³ | 6.8 | **5.5** | ✗ ⁴ |
|
||||
| yolo26s-sem-ade20k | 395 | 61 | 31 | 10.7 | ✗ ³ | 7.9 | **3.4** | ✗ ⁴ |
|
||||
|
||||
³ The offline fp16 conversion (`onnxconverter-common`) left a mixed-type node the CUDA provider
|
||||
rejects. TensorRT converts to fp16 itself at engine build and does not have this problem, which is
|
||||
one reason it is the target and the CUDA provider is the fallback.
|
||||
⁴ TensorRT refuses the QDQ form ONNX Runtime's quantiser writes for the Hexagon (uint8
|
||||
activations); it wants symmetric int8. Not pursued: fp16 needs no quantisation, no calibration and
|
||||
no accuracy gate, and it is already 30–60× tract.
|
||||
|
||||
CUDA int8 is deliberately absent: the CUDA provider has no int8 kernels and runs a QDQ graph by
|
||||
dequantising it, which measured *slower* than f32 (7.0 vs 5.6 ms on scrfd_500m). Int8 on NVIDIA is
|
||||
TensorRT's job.
|
||||
|
||||
**TensorRT's first load is 16–116 s per model in f32 and 35–290 s in fp16** (yolo26n-seg the
|
||||
worst: nearly five minutes), because it is compiling an engine for this exact GPU. The engine caches to disk and the second load is milliseconds. That
|
||||
number is what §6 is designed around.
|
||||
|
||||
### 1.3 What the numbers say
|
||||
|
||||
- **`tract` is single-threaded.** The tablet's one X4 core and one Raptor Lake core give the same
|
||||
tract numbers. Replacing it with ONNX Runtime's CPU provider, *no accelerator involved*, is 3–6×
|
||||
on the tablet and 8–10× on the desktop. That is the floor, and it is available on every platform
|
||||
the app builds for.
|
||||
- **The Hexagon is the standout.** A 6 W NPU running int8 beats a discrete RTX 3050 running f32 on
|
||||
five of six models. A whole-library face index on the tablet goes from ~100 ms + 39 ms per face
|
||||
to ~1.4 ms + 12 ms per face, and the "Thorough" detector — 3× the cost of "Fast" today — becomes
|
||||
free. Its price is that the models must be **quantised to int8**, which is an accuracy question
|
||||
§5 has to answer before it is believed.
|
||||
- **The embedder does not gain from either accelerator.** 112×112 input, per-op overhead
|
||||
dominates; it is 9 ms on the tablet's CPU and 12 ms on its NPU. It stays float, which §7 turns
|
||||
from a performance footnote into a correctness rule.
|
||||
- **On NVIDIA, TensorRT fp16 ≈ 3× the CUDA provider**, and the CUDA provider ≈ 2× the
|
||||
multi-threaded CPU; at fp16 the detectors are 1.8–3.3 ms with no quantisation at all. Both leave the twenty cores free for decoding during a batch index, which the table does not
|
||||
show and which matters more than the ratio.
|
||||
|
||||
---
|
||||
|
||||
## 2. The shape of the answer
|
||||
|
||||
A **ladder per platform**, walked at start-up, with the first rung that builds a real session
|
||||
winning:
|
||||
|
||||
| Platform | 1st | 2nd | 3rd | Floor |
|
||||
|---|---|---|---|---|
|
||||
| Android, Qualcomm with a Hexagon the shipped QNN skel covers (V68–V81) | QNN HTP, int8 model | ORT CPU, f32 model | — | tract |
|
||||
| Android, any other SoC | ORT CPU, f32 | — | — | tract |
|
||||
| Linux / Windows, NVIDIA GPU | TensorRT, f32 model, fp16 engine | CUDA provider, f32 | ORT CPU, f32 | tract |
|
||||
| Linux / Windows, no NVIDIA | ORT CPU, f32 | — | — | tract |
|
||||
| macOS ⁵ | ORT CPU, f32 | — | — | tract |
|
||||
|
||||
⁵ CoreML is the obvious rung and is unmeasured; it is listed so its absence is a gap and not an
|
||||
oversight.
|
||||
|
||||
Deliberately **not** on any ladder, with the measurement that excluded each: NNAPI (no driver),
|
||||
XNNPACK (slower than CPU, aborts on SCRFD), WebGPU (slower than CPU), the Adreno through QNN (works,
|
||||
but never where the Hexagon does not also), CUDA int8 (slower than CUDA f32). A rung is added to this
|
||||
table by a measurement on this page, not by a provider existing.
|
||||
|
||||
Two things the ladder is *not*: it is not a per-model choice — one backend serves every model on a
|
||||
device, because §7's identity rule needs the detector and embedder on the same runtime for the
|
||||
same reason `faces.model_id` pairs them; and it is not a per-account choice — it is a property of
|
||||
the hardware, like `shared_face_models_dir` is, and it lives beside it.
|
||||
|
||||
---
|
||||
|
||||
## 3. The dependency policy, and how far this reopens it
|
||||
|
||||
D13 chose `ort` over `tract` because `alternative-backend` made ONNX Runtime's *API* available with
|
||||
none of its *C*. Every rung above the floor needs the C++ ONNX Runtime and, for the two that matter
|
||||
most, vendor libraries on top: Qualcomm's QNN runtime (~60 MB for one Hexagon generation; it is
|
||||
per-SoC) and NVIDIA's TensorRT plus cuDNN (~600 MB with the CUDA libraries, and cuDNN's major
|
||||
version must match what ONNX Runtime was built against — the Arch package on the reference desktop
|
||||
was unusable for exactly that reason).
|
||||
|
||||
The policy protected the **build**: no C to cross-compile under the NDK, no toolchain to keep in
|
||||
step. This document keeps that intact, and the mechanism is the one thing about `ort` that makes it
|
||||
possible:
|
||||
|
||||
**`ort::set_api` accepts any `OrtApi` table.** With `alternative-backend` on, `ort` links nothing
|
||||
and asks for the table once per process. The application can `dlopen` a `libonnxruntime.so` it
|
||||
finds on disk, call `OrtGetApiBase()->GetApi(version)` and hand that table over; or, if there is no
|
||||
such file, hand over `ort_tract::api()`. The Rust build is identical in both cases — pure Rust,
|
||||
`cargo build --target aarch64-linux-android` sees the same dependency graph it sees today. What
|
||||
changes is that the runtime is a **file the package installs**, next to the models, and the app
|
||||
looks for it at start-up.
|
||||
|
||||
Consequences that follow and are accepted:
|
||||
|
||||
- **The runtime is chosen once per process**, because `set_api` is once per process. The ladder
|
||||
in §2 is walked at start-up and the result is what every session in that process uses. There is
|
||||
no "tract for this model, ORT for that one", and there is no falling back to tract *after* ONNX
|
||||
Runtime has loaded — but there does not need to be: once the library loads, its CPU provider is
|
||||
always there, and every fallback the ladder needs is between providers *inside* it.
|
||||
- **Feature flags stay as they are.** `dr-face`'s `inference` and `dr-segment`'s `semantic`
|
||||
continue to mean "compiled against `ort`'s API"; nothing at build time knows or cares which table
|
||||
will be supplied. The one addition is a `native-probe` feature on the new crate (§8) that pulls in
|
||||
`libloading`, which is pure Rust and already in the tree via `wgpu`.
|
||||
- **The packagers ship the runtime, not the build.** The Arch package, the Flatpak manifest, the
|
||||
NSIS installer and `assemble-apk.sh` each gain the ONNX Runtime library for their platform, and
|
||||
the Android and NVIDIA variants gain the vendor libraries — each under the licence the packager
|
||||
reads first (§3.1). A package without them is not broken; it is the tract build, and it says so
|
||||
on the about screen.
|
||||
- **The NDK problem does not come back.** `libonnxruntime.so` for Android is a prebuilt from
|
||||
Maven (`com.microsoft.onnxruntime:onnxruntime-android-qnn`), extracted by `assemble-apk.sh` into
|
||||
`jniLibs/` the way the models are bundled as assets today. Nothing compiles it.
|
||||
|
||||
### 3.1 Licences the packagers read before shipping a runtime
|
||||
|
||||
Written down now, because [segmentation.md §7](segmentation.md) established that reading the grant
|
||||
is cheaper than discovering it at packaging time.
|
||||
|
||||
| Component | Licence | Redistributable in a self-distributed package? |
|
||||
|---|---|---|
|
||||
| ONNX Runtime | MIT | Yes |
|
||||
| Qualcomm QNN runtime (`com.qualcomm.qti:qnn-runtime` on Maven) | Qualcomm AI Engine Direct SDK licence — proprietary, redistribution permitted for applications using it | Yes for the APK, with the licence text shipped; not for a source distribution. **To be read in full, not summarised from memory, before the APK gains it.** |
|
||||
| CUDA runtime, cuDNN, TensorRT | NVIDIA EULAs — redistributable with an application, with the licence text, not modifiable | Yes for a package that bundles them. 600 MB. The alternative is to load them from the user's system install if present and skip the rung otherwise — which is what §4's probe does anyway. |
|
||||
|
||||
The position this takes: the **NVIDIA libraries are not bundled**. The desktop package probes for a
|
||||
system CUDA/TensorRT install and uses it if it is version-compatible; a desktop without one runs on
|
||||
ORT CPU, which is still 8–10× today. Bundling 600 MB for a rung that is 2× again is not a trade
|
||||
worth making unmeasured, and it can be revisited by a measurement on a batch index. The **QNN
|
||||
runtime is bundled** in the APK, because the Hexagon is the difference between a tablet that
|
||||
indexes a library overnight and one that does it over lunch, and the package is 60 MB larger for
|
||||
it.
|
||||
|
||||
Both positions are D13 territory and are recorded there (§12).
|
||||
|
||||
---
|
||||
|
||||
## 4. Selection — the probe, its cache, and what it may not do
|
||||
|
||||
**A rung is chosen by building a real session on it, not by asking whether it exists.** Both
|
||||
failure modes that are not "the provider is absent" were hit on 2026-09-19: a driver in a wedged
|
||||
state where the provider registered and the session then failed, and a provider that registered,
|
||||
took the graph, and rejected every node at partition time. The probe therefore:
|
||||
|
||||
1. Loads the runtime library (§3), or falls to tract and stops.
|
||||
2. Times the **smallest detector** on the CPU provider first — the floor. Then, for each rung
|
||||
in this platform's ladder, in order: builds a session for the same model on that provider
|
||||
with `error_on_failure`, runs it once on a fixed input, and times three more runs. **The rung
|
||||
is taken only if its median beats the floor.** That one measurement is the proof the provider
|
||||
took the graph: one that silently hands the work to the CPU is the CPU rung with extra
|
||||
overhead, slower than the floor, and rejected. (ONNX Runtime's
|
||||
`session.disable_cpu_ep_fallback` was the first draft of this proof and refuses the Hexagon
|
||||
over the ten quantise/dequantise nodes at the graph's edges that QNN declines by policy.)
|
||||
3. Records the outcome — rung, runtime version, provider version, device identity (GPU name and
|
||||
compute capability; SoC model and Hexagon arch), and the models' content hashes — to a small
|
||||
file beside `shared_face_models_dir`. The next start-up trusts the file **unless** any of those
|
||||
inputs changed, in which case it probes again. A driver update, a runtime update, a new model
|
||||
file: each invalidates the cache by construction, and none needs a "reset backend" button.
|
||||
|
||||
What the probe may not do:
|
||||
|
||||
- **Block the first frame.** It runs on the same background as `install_bundled_models` and for
|
||||
the same reason: a TensorRT probe can take thirty seconds cold, and a tablet that stalls that long
|
||||
is an ANR. Until it reports, every model request is answered by the floor the runtime supports
|
||||
(ORT CPU if the library loaded, tract otherwise), and a job that started on the floor finishes
|
||||
on it — a backend does not change under a running index.
|
||||
- **Retry a rung that failed within a session.** A failed probe is cached as a failure with the
|
||||
same inputs; the rung is tried again when an input changes. Otherwise a wedged driver means a
|
||||
thirty-second stall on every launch.
|
||||
- **Choose for the user without saying so.** Settings gains one row, *Inference backend*, showing
|
||||
what was chosen and why in one line ("Hexagon NPU · int8 · QNN 2.42"; "CPU · ONNX Runtime 1.30 ·
|
||||
TensorRT probe failed: cuDNN 8 required"), with an override to force any lower rung. The about
|
||||
screen carries the same line beside the model names NFR-SEC-5 already puts there.
|
||||
|
||||
---
|
||||
|
||||
## 5. Model variants, and who makes them
|
||||
|
||||
Every model exists in one **canonical** form — the f32 ONNX file the app ships or the user supplies
|
||||
today — and, where a rung needs it, a **derived** form. The ladder's rungs are specified in terms
|
||||
of which form they load:
|
||||
|
||||
| Form | Who produces it | When | Needed by |
|
||||
|---|---|---|---|
|
||||
| f32 ONNX, shape-fixed, **opset ≥ 13** | `tools/fix-face-model-shapes.sh`, `tools/export-seg-model.sh` | Release time, once | Every rung except Hexagon |
|
||||
| int8 QDQ ONNX, per-channel, uint8 activations | `tools/quantise-models.sh` (new) | Release time, once, **calibrated on real photographs** | Hexagon |
|
||||
| TensorRT engine (`.engine`, per GPU architecture and TensorRT version) | The app, from the f32 file | First run on that device, in the background | TensorRT rung |
|
||||
| QNN context binary | The app, from the int8 file | First run on that device, in the background | Hexagon rung |
|
||||
|
||||
Two rules.
|
||||
|
||||
**Quantisation is a release-time step, not a device-time one.** The int8 files that produced §1's
|
||||
numbers were calibrated on random noise, which is enough to time and worthless to trust. A real
|
||||
int8 detector is calibrated on a few hundred real photographs and then measured against the f32
|
||||
detector on the reference library by [faces.md §12.3](faces.md)'s method — faces found, per size
|
||||
band, per detector — before it ships. That needs the reference library and a person reading the
|
||||
result, and it happens once per model release, in `tools/`, beside the shape-fixing it already
|
||||
depends on. The device never quantises anything.
|
||||
|
||||
The SCRFD and ArcFace files are **opset 11** as InsightFace exported them, and per-channel QDQ needs
|
||||
13; `tools/fix-face-model-shapes.sh` gains an opset upgrade to 17 (`onnx.version_converter`,
|
||||
`ir_version` 8), which `tract` has been verified to load and which every provider on this page
|
||||
prefers. That is a change to the canonical file and so a change to the shipped models, and it
|
||||
happens in the same model release as the int8 files.
|
||||
|
||||
**Compilation is a device-time step, and it is cached.** A TensorRT engine is specific to the GPU
|
||||
it was built on and the TensorRT that built it; a QNN context binary is specific to the Hexagon
|
||||
generation. Neither can ship. Both are built by the app the first time that rung is selected, in
|
||||
the background (§6), and written beside the probe cache keyed by the same inputs. They are
|
||||
**derived, disposable, and regenerable**: deleting the cache directory costs the next launch a
|
||||
rebuild and nothing else, and the directory is excluded from anything that syncs (it is a peer of
|
||||
`thumbs`, not of the catalog).
|
||||
|
||||
---
|
||||
|
||||
## 6. First run — building engines without the user waiting for them
|
||||
|
||||
The sequence on a device where a compiling rung (TensorRT, Hexagon) is selected:
|
||||
|
||||
1. **Launch.** The runtime loads; the probe (§4) starts in the background; the app serves every
|
||||
model request from the floor. Face indexing, segmentation and scene grading all work, at
|
||||
today's speed or better (ORT CPU).
|
||||
2. **Probe reports** — say, TensorRT. The compiling rung is now *selected* but has **no engines**.
|
||||
Model requests continue on the fallback rung below it (CUDA provider for TensorRT; ORT CPU for
|
||||
Hexagon), which needs no compilation and is already faster than the floor.
|
||||
3. **Engines build**, one model at a time, on a single low-priority background thread, smallest
|
||||
model first so the detector — the one that runs per image — is ready soonest. On the reference
|
||||
desktop that is ~1 minute for the first detector and ~10 minutes for all six at fp16; on the tablet the QNN
|
||||
context binaries take 0.8–1.7 s each and the whole set is ready before the user has opened a
|
||||
library. Each engine is written to a temporary name and renamed into place, so a request never
|
||||
sees a half-written file.
|
||||
4. **Requests move up as engines land.** A model whose engine exists loads it on the selected
|
||||
rung; one whose engine is still building loads on the fallback. **A running job does not
|
||||
switch** — an index that started on the CUDA provider finishes on it — because §7 needs one
|
||||
`model_id` per job, and because a job is the wrong granularity for surprise.
|
||||
5. On Android, the build runs only while the app is in the foreground and the device is not in
|
||||
battery saver (NFR-RES-3): a context binary takes a second, so this costs nothing, and the rule
|
||||
exists for the day a model takes longer.
|
||||
|
||||
Settings shows a one-line progress row while engines build ("Preparing GPU engines · 3 of 6") and
|
||||
nothing when they are done. A build failure demotes the rung — it is recorded in the probe cache
|
||||
as a failure with the model hash as an input, so a corrected model file retries it — and the app
|
||||
carries on one rung down, saying so in the same row.
|
||||
|
||||
---
|
||||
|
||||
## 7. Identity — what changes `model_id` and what may not
|
||||
|
||||
`faces.model_id` exists so that two libraries indexed with different networks are never compared
|
||||
as if they were one ([catalog.md §10.1](catalog.md); the trap is written up in
|
||||
[faces.md §14](faces.md)). A backend that changes what a network *computes* is a different network
|
||||
and must be a different `model_id`; one that changes only *where* it computes it must not be.
|
||||
|
||||
**The detector.** An int8 SCRFD finds a different set of faces from the f32 one — that is what
|
||||
§5's acceptance measures — so **the numeric form is part of the detector's identity**:
|
||||
`scrfd_500m` and `scrfd_500m_i8` are two detectors in `model_id`, and a library indexed on the
|
||||
tablet's Hexagon and continued on the desktop is two populations, which a re-index on either side
|
||||
reconciles the same way a switch from Fast to Thorough does today. That is acceptable because it is
|
||||
already the rule for the detector and because §5 is the gate on whether the int8 form is close
|
||||
enough to be *offered* at all. f32 on tract, ORT CPU, CUDA and TensorRT-f32 are one identity: the
|
||||
same graph, the same arithmetic, differences at the last bit.
|
||||
|
||||
**The embedder** is where comparability across devices is the whole point, and it is the one
|
||||
model that no accelerator helps (§1.3). So: **the embedder runs in f32 on every rung.** On TensorRT
|
||||
that means the embedder's engine is built without fp16 while the detector's is built with it; on
|
||||
the Hexagon it means the embedder is not on the NPU at all — it runs on the ORT CPU rung at 9 ms,
|
||||
and the ladder's "one backend per device" is, precisely, one backend *per model role*, with the
|
||||
embedder pinned. A `w600k_mbf` embedding from any device is comparable with one from any other,
|
||||
which is the property the identity system, the calibration and the cross-device merge all rest
|
||||
on, and it is not for sale for 3 ms.
|
||||
|
||||
If S16 wants fp16 for the embedder later, the gate is written now: over the reference library's
|
||||
faces, the cosine between the f32 and fp16 embedding of the same crop exceeds 0.999 for 99.9% of
|
||||
faces and the calibration's fitted threshold moves by less than its own confidence interval
|
||||
([faces.md §8.3](faces.md)). Until measured, f32.
|
||||
|
||||
**Segmentation and the scene model** carry no identity across devices — their outputs are
|
||||
recomputed per image and never stored beyond the cache — so they take whatever the rung offers,
|
||||
int8 included, subject to §10's own acceptance.
|
||||
|
||||
---
|
||||
|
||||
## 8. Crate shape — `core/dr-inference-engine`
|
||||
|
||||
The seam is the same shape as [storage.md](storage.md)'s: a small crate below the consumers that is
|
||||
the **only** place naming a provider, a library file or a vendor, with the consumers reduced to
|
||||
"give me a session for these bytes in this role".
|
||||
|
||||
```
|
||||
core/dr-inference-engine
|
||||
src/lib.rs Runtime (Tract | Onnx { lib, version }), Backend (rung), Role (Detector | Embedder | Segmenter)
|
||||
src/probe.rs §4 — the ladder per platform, the session-build probe, the cache file
|
||||
src/engines.rs §6 — background compilation, the cache directory, progress
|
||||
src/session.rs open(role, bytes) -> ort::Session, applying the rung and the role's precision rule
|
||||
src/api.rs the one unsafe block: dlopen libonnxruntime, fetch OrtApi, ort::set_api — or ort_tract::api()
|
||||
```
|
||||
|
||||
- `dr-face` and `dr-segment` **delete** their private `install_backend` and their direct
|
||||
`Session::builder()` calls and take an `&dr_inference_engine::Sessions` where they take model bytes today.
|
||||
Their tests keep `tract` — `dr_inference_engine::Sessions::tract()` is a constructor and the test-only path.
|
||||
- `dr-inference-engine` depends on `ort` with the same workspace features as today plus `cuda`, `tensorrt`,
|
||||
`qnn`: those features add option builders, not linking, under `alternative-backend`. **Verified
|
||||
for the QNN, CUDA and TensorRT builders on 2026-09-19** — they go through the API table's generic
|
||||
`SessionOptionsAppendExecutionProvider*`. The NNAPI builder resolves a symbol directly and would
|
||||
not; it is not needed and is not enabled.
|
||||
- `dr-ui` owns the settings row, the about-screen line and the progress row; it holds one
|
||||
`Sessions` per process, created at launch, and passes it down. `dr_ui::library` gains
|
||||
`inference_cache_dir()` beside `shared_face_models_dir()`, on the same account-independent
|
||||
footing and for the same reason.
|
||||
- The Android entry point's `install_bundled_models` also extracts nothing new: `jniLibs/` is
|
||||
loaded by the system loader, and `dr-inference-engine` on Android looks for `libonnxruntime.so` through
|
||||
`dlopen` by bare name first, which resolves to the APK's copy, before any directory.
|
||||
|
||||
---
|
||||
|
||||
## 9. Threads and memory
|
||||
|
||||
- ONNX Runtime's intra-op pool is sized to the physical cores minus two on desktop and to the
|
||||
performance cores on Android (the X4 and the A720s; the A520s are for the compositor). tract's
|
||||
single thread today is the reason a batch index leaves nineteen cores idle; ORT CPU with the
|
||||
pool is the reason it will not. One session per model per process; `Session::run` is
|
||||
`&mut self`-free in `ort` and internally serialised, and the index job is the only caller.
|
||||
- A TensorRT session pins GPU memory for its workspace; the builder is capped at 512 MB on the
|
||||
reference 6 GB card and the cap is a setting, because the develop view's tiles share the card
|
||||
(NFR-RES-2). The engine cache on disk is bounded by the model set — six engines, ~80 MB — and
|
||||
needs no LRU.
|
||||
- The Hexagon rung sets QNN's performance mode to `Burst` for the duration of an index job and
|
||||
`Default` otherwise; a 5 ms detector does not need the NPU clocked up between images.
|
||||
- The probe (§4) and the engine build (§6) run on one dedicated low-priority thread. They never
|
||||
share the index job's pool: a probe that competes with the job it is meant to speed up is the
|
||||
frame-budget trap in a new coat.
|
||||
|
||||
---
|
||||
|
||||
## 10. What S16 measures
|
||||
|
||||
In order, with the gate each is:
|
||||
|
||||
| # | Question | Gate |
|
||||
|---|---|---|
|
||||
| M1 | Does one binary carry both tables? `dlopen` + `set_api` on Linux, Windows and Android; `ort_tract::api()` when the file is absent. | Go / no-go for §3. If `set_api` cannot take a dynamically fetched table on some platform, that platform ships two binaries, and the cost is stated. |
|
||||
| M2 | Do the int8 SCRFD detectors, **calibrated on real photographs**, find the faces? faces.md §12.3's method over the reference library, per size band, against f32. | Ship the int8 form for a detector only if it finds ≥ 97% of the f32 detector's faces above 40 px and the difference is not concentrated in one band. Otherwise that detector's Hexagon rung is ORT CPU int8-free, and the table in §1.1 says what that costs. |
|
||||
| M3 | Does the embedder on ORT CPU beside a detector on the Hexagon (§7) produce embeddings within the f32 gate? | It must — same graph, same arithmetic. This is a check that the plumbing did not quantise it by accident. |
|
||||
| M4 | Is a batch index on the tablet and on the desktop faster by the ratio §1 predicts, end to end, decode included? | The face index over the reference library (18,143 faces): report wall-clock on tract, on the floor and on the selected rung, and where the time went. The prediction is that decode becomes the bottleneck on both; if it does not, say why. |
|
||||
| M5 | Does the first-run sequence (§6) hold: nothing blocks the first frame, engines land, requests move up, a running job does not switch? | Observed on both devices with the app's own progress row, and with the cache directory deleted between runs. |
|
||||
| M6 | What does the APK weigh with the QNN runtime, and does a non-Qualcomm Android device (any one) still launch and index on the floor? | Size reported; launch verified on one non-Qualcomm device or an emulator. |
|
||||
| M7 | The segmentation and scene models at int8 on the Hexagon: does the mask boundary move? segmentation.md §6's IoU against f32 over its corpus. | Ship int8 for a model only above the IoU floor that document set for arm B. |
|
||||
|
||||
M1 and M2 are the ones the rest is conditional on, and M2 is the one that needs a person.
|
||||
|
||||
### 10.1 M2 result · 2026-09-19
|
||||
|
||||
Each int8 detector against its own f32 form, over 400 proxies evenly spaced through the reference
|
||||
library, on ONNX Runtime's CPU provider (the int8 graph is the same file the Hexagon loads;
|
||||
`ui/dr-ui/examples/face_detectors`). Calibrated on 64 proxies from the same library, disjoint
|
||||
from the 400.
|
||||
|
||||
| Detector | f32 faces | int8 faces | both | int8 only | f32 only | found ≥ 32 px | found, all sizes |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| scrfd_500m | 1342 | 1319 | 1287 | 32 | 55 | 95.6% | 95.9% |
|
||||
| scrfd_2.5g | 1525 | 1482 | 1478 | 4 | 47 | 96.1% | 96.9% |
|
||||
| scrfd_10g | 1769 | 1760 | 1749 | 11 | 20 | 100% | 98.9% |
|
||||
|
||||
The 10g form clears the 97% gate; 500m and 2.5g sit one point under it. What they lose is
|
||||
specific: the faces in the "f32 only" column have a **median confidence of 0.52** against a
|
||||
threshold of 0.50 — detections the f32 graph itself barely made, that int8 rounding drops to the
|
||||
other side of the line — and the extra faces int8 finds are the same kind (median 0.51–0.52).
|
||||
Not a size-band failure: the losses are spread across bands in proportion. Shipped as they are,
|
||||
with the number on record; a threshold of 0.48 for the int8 forms would recover most of the
|
||||
margin, and is the first thing to try if a library's count on the tablet reads low.
|
||||
|
||||
Two things the calibration taught, both in `tools/quantise-models.py`: the calibration set has
|
||||
to contain faces (a first attempt on landscape photographs produced a graph that found nothing —
|
||||
the score head's ranges had never seen the face regime), and ONNX Runtime's own strided and
|
||||
moving-average calibration modes both measurably degrade the result on these graphs, while
|
||||
driving the calibrator in chunks by hand reproduces the plain min/max ranges exactly.
|
||||
|
||||
---
|
||||
|
||||
## 11. Order
|
||||
|
||||
1. **`dr-inference-engine` with the two tables and the floor** — `set_api` from a dlopened runtime, tract
|
||||
otherwise, ORT CPU as the only rung. Consumers moved over; tests unchanged. This alone is the
|
||||
3–10× and is the build most of the value sits in. M1.
|
||||
2. **The probe and its cache** (§4), with the settings row and the about line. Still CPU-only;
|
||||
the ladder has one rung. M5's first half.
|
||||
3. **`tools/quantise-models.sh`** and the opset upgrade; the int8 detectors calibrated and
|
||||
measured. M2, M7. This is the step with a person in it and it runs in parallel with 4.
|
||||
4. **The Hexagon rung**, the QNN runtime in the APK, the context-binary cache. M3, M6.
|
||||
5. **The TensorRT and CUDA rungs** on desktop, the engine cache, the first-run sequence. M5's
|
||||
second half.
|
||||
6. **M4** last, on both devices, and the number goes in this document.
|
||||
|
||||
The Windows installer and the Flatpak manifest are touched in steps 1 and 5 only, and only to add
|
||||
a file each; the Arch package likewise.
|
||||
|
||||
---
|
||||
|
||||
## 12. Register entries
|
||||
|
||||
**FR-INF-1 — Runtime selection.** On launch the application shall determine, per device and
|
||||
without blocking the first frame, the fastest inference backend that can build and run a session
|
||||
for the shipped models, by attempting it; shall record and reuse that determination until the
|
||||
runtime, driver, hardware or models change; and shall display the backend in use in Settings and
|
||||
on the about screen. *Acceptance:* §10 M1 and M5.
|
||||
|
||||
**FR-INF-2 — Derived engines.** Backends that require device-specific compilation shall compile in
|
||||
the background after selection, shall serve requests from the next lower backend until each engine
|
||||
is ready, and shall not change the backend of a job in progress. *Acceptance:* M5.
|
||||
|
||||
**FR-INF-3 — Model forms.** Quantised model forms are produced at release time from real
|
||||
calibration data and are shipped only when they meet §10's accuracy gates against the canonical
|
||||
form; the application never quantises on the device. *Acceptance:* M2, M7.
|
||||
|
||||
**NFR-INF-1 — Embedding comparability.** Face embeddings shall be computed at a precision whose
|
||||
deviation from the f32 reference is within §7's gate, on every backend, so that embeddings from any
|
||||
device are comparable. *Acceptance:* M3.
|
||||
|
||||
**D13 — updated.** The runtime half is reopened to the extent of §3: the Rust build stays C-free
|
||||
under `alternative-backend`; packages may install a dynamically loaded ONNX Runtime and, per §3.1,
|
||||
the Qualcomm QNN runtime; the NVIDIA libraries are not bundled. The licensing half is unchanged.
|
||||
|
||||
---
|
||||
|
||||
## 13. Requirements touched
|
||||
|
||||
FR-CULL-8 (indexing time is what this exists to change), FR-CULL-9 (NFR-INF-1 is the guard on its
|
||||
calibration), NFR-RES-2 (TensorRT workspace against the develop view's tiles), NFR-RES-3
|
||||
(background compilation on Android), NFR-SEC-5 (the about screen's line), NFR-COMPAT-2 (each
|
||||
channel gains a runtime file), FR-PLAT-AND-1 (unchanged; the runtime lives in the APK, not in
|
||||
storage), ARCH §6.1 (the once-per-image budget this was sized against no longer binds; what could
|
||||
run per frame is a separate question this document does not open).
|
||||
@@ -0,0 +1,780 @@
|
||||
# Editing a mask
|
||||
|
||||
**Status:** Draft · 2026-09-06
|
||||
**Companion to:** [requirements.md](requirements.md) §3.3 FR-DEV-3, FR-DEV-10 ·
|
||||
[segmentation.md](segmentation.md) · [architecture.md](architecture.md) §5.4 ·
|
||||
[spot-removal.md](spot-removal.md) (the pattern a canvas tool follows here)
|
||||
|
||||
The model finds a subject in a second and the photographer cannot then change
|
||||
it by a single pixel. This document specifies the tools that close that gap —
|
||||
paint, erase, push, and combining one selection with another — and the
|
||||
interface they are driven through.
|
||||
|
||||
---
|
||||
|
||||
## 1. The gap, precisely
|
||||
|
||||
The core is further along than the interface, and it is worth being exact
|
||||
about which half is missing, because it changes the size of the work.
|
||||
|
||||
**Painting exists everywhere except where a finger is.**
|
||||
[`MaskSource::Brush`](../../core/dr-pipeline/src/mask.rs), [`Stroke`], the
|
||||
simplification and the point budgets, the sidecar's `stroke = …` line and its
|
||||
parser, the GPU's per-stroke bounding-box draw with add and erase blend
|
||||
states — all of it is written, tested, and reachable from no control in the
|
||||
application. [`toolrail.slint:138`](../../ui/dr-ui/ui/toolrail.slint#L138) says so
|
||||
in as many words: *"what is missing is the canvas interaction"*.
|
||||
|
||||
**A mask has exactly one source.** A layer is one `MaskSource` and a shaping
|
||||
of its edge. There is no way to say *this subject **and** that one*, *the sky
|
||||
**except** the branches*, or *the subject **only where** it is bright* — the
|
||||
last being an intersection of a model mask with a range mask (FR-DEV-10), two
|
||||
features that already exist and cannot meet.
|
||||
|
||||
**The edge moves all at once or not at all.** `Morphology` grows or shrinks
|
||||
the *whole* boundary. When the model's coverage stops two pixels inside the
|
||||
shoulder and leaks four pixels into the hair, no global number fixes both, and
|
||||
that is the ordinary case rather than a corner one.
|
||||
|
||||
**And you cannot see the mask.** *(Built — see §6.)* The overlay on the canvas
|
||||
was [`overlay_rgba`](../../ui/dr-ui/src/segmentation.rs) and nothing else — a
|
||||
CPU-built false-colour picture of *what the model detected*, at proxy
|
||||
resolution. It is not the layer's alpha: it knows nothing of the layer's
|
||||
feather, its falloff, its morphology, its invert, or its opacity. Nobody can
|
||||
refine an edge they are not being shown.
|
||||
|
||||
That one turned out to be load-bearing for the other three rather than the
|
||||
last of four. With no way to see a mask, choosing a category produced a layer
|
||||
whose extent was invisible and whose adjustment had not been touched yet — so
|
||||
the correct behaviour and the broken one look identical, and "the segmentation
|
||||
does not make masks" is what it reads as from the outside.
|
||||
|
||||
Those four are one feature, and this is its specification.
|
||||
|
||||
## 2. Non-goals
|
||||
|
||||
- **Not pixel layers.** §1.3 of the requirements excludes them. Everything
|
||||
below stores geometry and parameters; no rasterised mask is ever written to
|
||||
a file, and no mask ever exists in CPU memory (ARCH §5.4).
|
||||
- **Not a second segmentation UI.** Choosing *which* subject or category the
|
||||
model offers stays exactly as it is. This is about what happens after.
|
||||
- **Not per-layer neighbourhood adjustments.** `layer_chain()` excludes detail
|
||||
and optics operations for reasons that have not changed.
|
||||
- **Not automatic refinement.** No "improve this mask" button that silently
|
||||
redraws a boundary the photographer approved. The tools here move an edge
|
||||
only where a hand is.
|
||||
- **Not a mask library.** Saving a mask and applying it to another photograph
|
||||
is a reasonable later feature and depends on none of this.
|
||||
|
||||
## 3. The five tools
|
||||
|
||||
Named as a photographer would say them, because these are the words the
|
||||
interface will use.
|
||||
|
||||
| Tool | What it does | New machinery |
|
||||
|------|--------------|---------------|
|
||||
| **Paint** | Adds to the mask under the brush | none — `brush_add` exists |
|
||||
| **Erase** | Takes away under the brush | none — `brush_erase` exists |
|
||||
| **Push** | Drags the boundary itself: outward fills behind it, inward empties behind it | a warp pass |
|
||||
| **Combine** | Joins another selection to this mask — add, subtract, or keep only the overlap | three blend states |
|
||||
| **Show** | Draws the mask that actually results, live | two uniforms in the composed shader |
|
||||
|
||||
Push is the one worth defining carefully, because "tug the edge" can mean two
|
||||
different operations and only one of them is right here — §5.3.
|
||||
|
||||
## 4. The model: a mask is a stack of parts
|
||||
|
||||
### 4.1 A part
|
||||
|
||||
The change that carries all four gaps at once is that a layer stops holding a
|
||||
source and starts holding an ordered list of them.
|
||||
|
||||
```rust
|
||||
/// One selection joined into a layer's mask.
|
||||
pub struct MaskPart {
|
||||
/// Stable identity, for the sidecar and for merge (FR-NC-9).
|
||||
pub id: String,
|
||||
/// How this part enters the mask built so far. Ignored on the first part,
|
||||
/// which *is* the mask so far.
|
||||
pub join: Join,
|
||||
pub source: MaskSource,
|
||||
/// Shaping, moved down from the layer: two parts of one mask routinely
|
||||
/// want different edges — a model's soft coverage joined to a hand-painted
|
||||
/// correction that must be exactly where it was painted.
|
||||
pub invert: bool,
|
||||
pub feather: f32,
|
||||
pub falloff: Falloff,
|
||||
pub morphology: Morphology,
|
||||
pub morph_radius: f32,
|
||||
pub refine: f32,
|
||||
/// The model raster behind a Subject or Category source. Per part now,
|
||||
/// for the same reason it was per layer: it materialises *this* selection.
|
||||
pub coverage: Option<Arc<Coverage>>,
|
||||
}
|
||||
|
||||
/// How a part joins the mask before it.
|
||||
pub enum Join {
|
||||
/// Everything either has. The default, and what "add a brush" means.
|
||||
Union,
|
||||
/// What the mask had, minus this. "Subtract".
|
||||
Subtract,
|
||||
/// Only where both agree. Where a subject meets a luminance band.
|
||||
Intersect,
|
||||
/// Not a set operation: this part's strokes *move* the mask under it.
|
||||
/// Only ever a painted part carrying push strokes — see §5.3.
|
||||
Warp,
|
||||
}
|
||||
```
|
||||
|
||||
A layer's mask is then `parts.fold(empty, join)`, and every one of §1's gaps
|
||||
becomes an ordinary use of it:
|
||||
|
||||
- **Paint on an auto mask** — append a `Union` part whose source is painted.
|
||||
- **Erase from one** — strokes inside that part carry `Erase`, or the part
|
||||
itself is a `Subtract`. Both work; the panel offers the first, because a
|
||||
photographer alternating add and erase over one area is drawing one
|
||||
correction, not two.
|
||||
- **Merge two selections** — two `Union` parts.
|
||||
- **Cut one out of another** — a `Subtract` part.
|
||||
- **The bright part of the sky** — a `Category` part, then an `Intersect`
|
||||
luminance part.
|
||||
- **Tug an edge** — a `Warp` part.
|
||||
|
||||
**Why a list on the layer rather than a boolean tree.** A tree expresses more
|
||||
and no photographer has ever wanted the extra. A flat ordered fold is what
|
||||
Lightroom, Capture One and darktable all present, it reads top to bottom in a
|
||||
panel with no parentheses to draw, and — decisively here — it merges under
|
||||
FR-NC-9 as a sequence of independently-keyed blocks, where a tree would merge
|
||||
as a shape whose two halves can be individually won by different devices and
|
||||
recombined into something neither ever had.
|
||||
|
||||
### 4.2 A stroke gains a mode
|
||||
|
||||
```rust
|
||||
pub enum StrokeMode { Add, Erase, Push }
|
||||
|
||||
pub struct Stroke {
|
||||
pub mode: StrokeMode, // was: erase: bool
|
||||
pub radius: f32,
|
||||
pub hardness: f32,
|
||||
pub flow: f32,
|
||||
/// How strongly the deposit clings to the picture's own edges: 0 paints
|
||||
/// anywhere, 1 paints only what matches the colour under the point the
|
||||
/// stroke began. See §5.6.
|
||||
pub cling: f32,
|
||||
pub points: Vec<(f32, f32)>,
|
||||
}
|
||||
```
|
||||
|
||||
`erase: bool` becomes a three-valued mode. Everything else about `Stroke` —
|
||||
the grid snapping, `simplify`, `MIN_STEP_FRACTION`, `MAX_STROKE_POINTS` and
|
||||
its continuation rule — is unchanged and applies to a push stroke exactly as
|
||||
it does to a painted one.
|
||||
|
||||
### 4.3 What the layer keeps
|
||||
|
||||
```rust
|
||||
pub struct MaskLayer {
|
||||
pub id: String,
|
||||
pub name: String,
|
||||
pub enabled: bool,
|
||||
/// Invert and opacity stay here: they are the two uniforms the generated
|
||||
/// shader already reads per layer (LAYER_UNIFORM_FIELDS), they apply to
|
||||
/// the finished mask, and moving them would change the composed shader.
|
||||
pub invert: bool,
|
||||
pub opacity: f32,
|
||||
/// Never empty. `parts[0]` is the base.
|
||||
pub parts: Vec<MaskPart>,
|
||||
pub ops: Vec<Box<dyn Operation>>,
|
||||
}
|
||||
```
|
||||
|
||||
`MaskSource::Brush { strokes }` becomes `MaskSource::Painted { strokes }` —
|
||||
the same variant under a name that no longer implies it is the only thing a
|
||||
brush can touch. `MaskLayer::begin_stroke` / `extend_stroke` / `end_stroke`
|
||||
keep their signatures and route to the layer's *active* part, which the panel
|
||||
sets; that is the only change their callers see.
|
||||
|
||||
**Budgets.** `MAX_LAYER_POINTS` (4096) becomes a budget across all of a
|
||||
layer's parts rather than one part's, so a layer's worst-case sidecar size and
|
||||
worst-case rasterisation cost are unchanged. `MAX_PARTS = 8` per layer, on the
|
||||
argument `MAX_LAYERS` makes: past that it is not a selection any more, and a
|
||||
bound the panel can show is better than one a file discovers.
|
||||
|
||||
### 4.4 The sidecar, and why no existing file changes
|
||||
|
||||
Blocks are `[mask <version> <layer-id>]` today. Parts are their own blocks:
|
||||
|
||||
```
|
||||
[mask default m1]
|
||||
name = Sky
|
||||
source = category
|
||||
signature = 4711
|
||||
category = sky
|
||||
feather = 0.004
|
||||
falloff = smooth
|
||||
opacity = 1
|
||||
|
||||
[part default m1 p2]
|
||||
join = subtract
|
||||
source = painted
|
||||
stroke = add 0.05 0.5 1 0.6 0.31,0.42 0.33,0.44 …
|
||||
```
|
||||
|
||||
Three compatibility rules, and together they mean **every sidecar written by
|
||||
every build so far loads into this one unchanged, and a layer this build
|
||||
writes with one part is byte-identical to what it writes today**:
|
||||
|
||||
1. A `[mask …]` block with no `[part …]` blocks after it is one part. Its
|
||||
source keys and its shaping keys build `parts[0]`; nothing is missing and
|
||||
nothing needs a default invented for it.
|
||||
2. A layer with exactly one part writes the old shape — source and shaping
|
||||
inside `[mask …]`, no part block. The format grows only when the feature is
|
||||
used.
|
||||
3. `stroke = …` lines keep their position and their meaning. The mode token
|
||||
gains `push`, and `cling` is a fourth number written only when non-zero.
|
||||
An older build meeting either drops *that stroke and only that stroke*,
|
||||
which is the rule [`parse_stroke`](../../core/dr-pipeline/src/sidecar.rs)
|
||||
already documents and already implements.
|
||||
|
||||
**Merge (FR-NC-9).** A part is a block with an id, so two devices that added
|
||||
different parts to the same layer merge to a layer with both, and the same
|
||||
part edited on both is a conflict over that part rather than over the mask.
|
||||
Part *order* is the one thing that is not per-field: it is stored as the block
|
||||
order and resolved the way the layer order already is.
|
||||
|
||||
## 5. On the device
|
||||
|
||||
### 5.1 The order a layer rasterises in
|
||||
|
||||
Per active layer, into that layer's slice of the existing mask array:
|
||||
|
||||
1. **Part 0** draws as it does today — one full-screen pass through the
|
||||
`switch` in `mask.wgsl`, or per-stroke bounding boxes for a painted one.
|
||||
2. **Each further part** draws the same way, with the blend state its `join`
|
||||
names (§5.2). No intermediate texture: the accumulator *is* the slice.
|
||||
3. **A `Warp` part** copies the slice aside and redraws it displaced (§5.3).
|
||||
4. `invert` and `opacity` are unchanged — two uniforms, read in the composed
|
||||
shader, applied to the finished mask.
|
||||
|
||||
`MAX_LAYERS` and the array's memory are untouched: parts collapse into one
|
||||
slice, so eight layers still cost eight channels.
|
||||
|
||||
### 5.2 Combining costs three blend states and no new texture
|
||||
|
||||
The set operations are already expressible in fixed-function blending over
|
||||
`r8unorm`, which is why this is the cheap half of the feature:
|
||||
|
||||
| Join | `src_factor`, `dst_factor`, op | Result | Built |
|
||||
|------|-------------------------------|--------|-------|
|
||||
| Union | `One`, `One`, `Max` | `max(dst, src)` | M1 |
|
||||
| Subtract | `Zero`, `OneMinusSrc`, `Add` | `dst · (1 − src)` | M1 |
|
||||
| Intersect | `Zero`, `Src`, `Add` | `dst · src` | M2 |
|
||||
|
||||
The middle row is `brush_erase`, already constructed in
|
||||
[`MaskPass::new`](../../core/dr-gpu/src/mask.rs). The other two are the same
|
||||
three vertices with a different `BlendState`, and nothing is read back.
|
||||
|
||||
**One correction to the first draft of this section, found in the building.**
|
||||
It claimed no second texture was needed, because a part could be blended
|
||||
straight onto the layer's slice. That is wrong, and the reason is the erase
|
||||
stroke: an erase inside a part means *a hole in that part*, not a hole in the
|
||||
mask. Drawn straight onto the accumulator it takes away whatever the parts
|
||||
before it had put there — so tidying the edge of a correction punches through
|
||||
the subject underneath, and the failure reads as the model's mask having holes
|
||||
in it.
|
||||
|
||||
So a part is drawn into one scratch texture (proxy-sized, `r8unorm`, allocated
|
||||
the first time any layer has more than one part) and blended from there. A
|
||||
layer of one part still takes the old path exactly — straight into its slice,
|
||||
no scratch, no combine pass — which is what keeps every existing mask
|
||||
rendering as it did. `an_erase_stroke_holes_its_own_part_and_not_the_mask` in
|
||||
[`local_adjustments.rs`](../../core/dr-gpu/tests/local_adjustments.rs) is the test
|
||||
that holds this in place.
|
||||
|
||||
`Max` blending on `r8unorm` is core WGPU and universally supported on the
|
||||
desktop backends; **verify it on the Android adapter before M2 lands**, since
|
||||
that is the platform where a blend mode is most likely to be quietly emulated
|
||||
or absent. If it is missing, union is `One, OneMinusSrc, Add` — a screen blend
|
||||
— which differs from `max` only where both parts are partially covered, and
|
||||
never by more than the softness of their two edges.
|
||||
|
||||
### 5.3 Push is a warp, not a local morphology
|
||||
|
||||
Two mechanisms fit "drag the edge", and the choice matters.
|
||||
|
||||
**Local morphology** — offset the distance threshold inside the brush — is
|
||||
exact, but it exists only where there is a signed distance field, which is
|
||||
`Subject` and `Category` and nothing else. A push tool that works on a
|
||||
model mask and does nothing on a gradient, a range or a painted part is a tool
|
||||
the photographer cannot trust.
|
||||
|
||||
**A warp** works on any mask, because it never asks what the mask is made of.
|
||||
Each dab of a push stroke contributes a displacement, and the mask is resampled
|
||||
through the sum of them:
|
||||
|
||||
```
|
||||
D(p) = Σ over dabs (b − a) · w(|p − seg(a,b)| / radius)
|
||||
mask'(p) = mask(p − D(p))
|
||||
```
|
||||
|
||||
Drag from inside the mask outward and the sample point moves back into the
|
||||
interior, so the boundary follows the finger and **the area behind it fills
|
||||
in**. Drag from outside inward and the vacated area samples from outside, so
|
||||
**the area behind it empties**. That is exactly the pair of behaviours asked
|
||||
for, from one gesture, with the direction supplied by the drag rather than by
|
||||
a mode switch.
|
||||
|
||||
Three honest costs:
|
||||
|
||||
- It needs the mask it is sampling, and a pass cannot sample the target it
|
||||
writes: a warp part costs one texture copy of the slice, plus one draw over
|
||||
the strokes' bounding box. One scratch texture, proxy-sized, `r8unorm`,
|
||||
allocated on first use and reused by every layer, since layers rasterise in
|
||||
sequence.
|
||||
- The displacement is a **sum**, not a sequential resample. Two crossing
|
||||
push strokes therefore compose approximately rather than exactly. Bound
|
||||
`|D|` at one brush radius per stroke: past that a warp tears rather than
|
||||
drags, and no photographer means the difference.
|
||||
- A warp moves what is under it, including detail the photographer painted by
|
||||
hand. That is what it is for, and it is why push is a part in the list —
|
||||
it can be removed later without disturbing the parts beneath it.
|
||||
|
||||
### 5.4 Painting has to be incremental
|
||||
|
||||
Today [`MaskPass::render`](../../core/dr-gpu/src/mask.rs) clears each slice and
|
||||
redraws every stroke of the layer. That is right when a shape changes and
|
||||
wrong while a finger is down: at 120 reports a second, a layer holding 4096
|
||||
points redraws all of them per dab, and the cost of a stroke grows as it is
|
||||
painted — the failure mode `MAX_STROKE_POINTS` already names for one stroke,
|
||||
here across the layer.
|
||||
|
||||
**While a stroke is live, only the new segment is drawn.** The slice already
|
||||
holds everything up to the previous dab, the blend states are the same ones
|
||||
that would have been used in a full redraw, and the segments are drawn once
|
||||
each in the same order — so the incremental result is identical to the
|
||||
rebuilt one rather than an approximation of it. `MaskPass` keeps, per layer,
|
||||
the `(part count, stroke count, point count)` it last drew; anything else
|
||||
changing falls back to the full rebuild it does now.
|
||||
|
||||
This is required, not an optimisation to schedule later: it is what decides
|
||||
whether painting is usable on the phone, and it is the specific failure
|
||||
[`mask.rs`'s module docs](../../core/dr-pipeline/src/mask.rs) say this whole
|
||||
design exists to avoid.
|
||||
|
||||
### 5.5 Distance fields become per part
|
||||
|
||||
[`SubjectMasks`](../../core/dr-gpu/src/mask.rs) uploads one signed distance field
|
||||
**per active layer, in stack order**, and
|
||||
[`DevelopSession`](../../ui/dr-ui/src/develop.rs) builds them on the same
|
||||
indexing. With parts, a field belongs to the part that shaped it: the upload
|
||||
becomes one field per *model-backed part*, flattened in `(layer, part)` order,
|
||||
and the rasteriser indexes it by a running counter rather than by `slot`.
|
||||
|
||||
This is the most invasive change in the document — it touches the field
|
||||
builder, the upload, the key that decides when to rebuild, and the shader's
|
||||
`subject` binding index — and it is mechanical. It is also the reason M2 is
|
||||
its own milestone rather than a rider on M1.
|
||||
|
||||
### 5.6 Cling: paint that stops at the picture's edge
|
||||
|
||||
A brush that respects the photograph's own boundaries is the difference
|
||||
between refining a mask and colouring it in, and the ingredients are already
|
||||
bound to this pass: the demosaiced source (for range masks) and, when there is
|
||||
one, the compacted label field.
|
||||
|
||||
At each dab, `cling > 0` multiplies the deposit by agreement with the pixel
|
||||
under the **stroke's first point**: colour distance in the same linear-sRGB
|
||||
space `colour_mask` already works in, falling off over a tolerance set by
|
||||
`cling`. Where a label field exists, agreement is 1 inside the same region and
|
||||
falls to the colour test outside it, so the brush stops dead at a watershed
|
||||
boundary and softly at a colour one.
|
||||
|
||||
Roughly fifteen lines of WGSL reusing `image_value` and `hue_of`, one number
|
||||
in the sidecar, and it is the single control that makes a 5%-radius brush
|
||||
usable along hair.
|
||||
|
||||
## 6. Seeing the mask
|
||||
|
||||
**Status: built.** `MaskStack::rendered`, `Reveal` as a list of
|
||||
`(layer, colour)`, `RevealStyle`, `EditGraph::compose_revealing`,
|
||||
`MaskPass::render_revealing`; on the panel, an eye and a colour per row and one
|
||||
"Show masks as" strip above the stack.
|
||||
|
||||
The mask array is **already bound to the composed adjust shader**, so this is
|
||||
almost free, and it is the first thing to build because every other tool here
|
||||
is unusable without it.
|
||||
|
||||
### 6.1 One correction to this section, found in the building
|
||||
|
||||
The draft said "two uniforms — which layer to reveal (−1 for none) and which
|
||||
style — always emitted and guarded by the uniform, so switching the overlay on
|
||||
is a uniform write rather than a shader recompile". The second half of that is
|
||||
not available, and the reason is the case the feature exists for.
|
||||
|
||||
A uniform can *select* a slot. It cannot conjure one. A layer with no
|
||||
adjustment on it changes no pixel, so it is not `is_active`, so it occupies no
|
||||
slice of the mask array and the rasteriser never draws it — and that is
|
||||
precisely the layer a photographer wants to look at, for the whole of the time
|
||||
between choosing a subject and deciding what to do to it. Revealing it means
|
||||
*rendering* it, which changes the sequence of layers, which changes the
|
||||
uniform block. The composition moves either way.
|
||||
|
||||
So the slot and the style are written into the source, and turning the reveal
|
||||
on, off, or onto another layer recompiles the fused shader. That is a button
|
||||
press rather than a frame, and the uniform would only have added a branch per
|
||||
pixel on top of a recomposition that was happening anyway.
|
||||
|
||||
`MaskStack::rendered(reveal)` is the one sequence this rests on: `active()`
|
||||
plus the layer being looked at. The rasteriser, the composer and the distance
|
||||
field builder all index by position in it, so all three must be given the same
|
||||
`reveal` — two of them disagreeing shows as an adjustment applied through
|
||||
another layer's mask, which is why they take it as an argument rather than
|
||||
reading a flag.
|
||||
|
||||
### 6.2 Where the block runs, and why not with the others
|
||||
|
||||
After the output transform, immediately before the clip and the encode — not
|
||||
among the layer blocks. Everything there runs on scene-referred colour in the
|
||||
working space, where a flat tint would be pushed through the base curve and
|
||||
the camera matrix and arrive as some other colour, and an alpha's white on
|
||||
black would arrive as neither.
|
||||
|
||||
### 6.3 Not on the graph
|
||||
|
||||
The reveal is an argument to `EditGraph::compose_revealing`, and
|
||||
`compose_for` — which the exporter, the thumbnail and the neutral probe all
|
||||
call — has no way to ask for one. A flag on the graph would have been fewer
|
||||
parameters, would have type-checked, and would have been one forgotten reset
|
||||
away from a red tint baked into an exported file.
|
||||
|
||||
Three styles, all read from the same alpha:
|
||||
|
||||
- **Tint** — the mask over the picture in its colour at ~50%. The default,
|
||||
and what every editor's photographers already expect. The colour is the
|
||||
mask's own, chosen from the swatches on its row — which is what answers a
|
||||
red tint over a red dress.
|
||||
- **Alpha** — the mask alone, white on black. For judging an edge, where a
|
||||
tint over a busy picture cannot be read.
|
||||
- **Edge** — the boundary outlined over the untouched picture. For checking
|
||||
registration against detail the other two hide, and the same reasoning the
|
||||
region overlay's white outline already carries.
|
||||
|
||||
**Per mask, not per selection.** The first build of this showed the *selected*
|
||||
layer's mask in one global style, and it answered the wrong question. What a
|
||||
photographer asks of two masks is how they meet — where the sky's edge sits
|
||||
against the building's — and that needs both on screen at once, in colours that
|
||||
can be told apart. So each row of the stack has an eye, and each mask a colour
|
||||
from a six-entry palette (`MASK_COLOURS` in `develop.rs`); the eye is drawn in
|
||||
that colour so the row says which shape on the picture is its. The style is
|
||||
the one thing that stays global, because a tint beside an outline beside an
|
||||
alpha would be three pictures that cannot be read against each other. Alpha
|
||||
therefore draws every shown mask, each in its colour, on black.
|
||||
|
||||
**When it appears.** A new layer arrives with its eye open, in the first colour
|
||||
nothing else is using — making a mask is asking what it selected, and for a
|
||||
subject or a category that question has no other answer on screen. Arming
|
||||
Paint or Erase opens the selected layer's eye if it was closed, on the same
|
||||
argument `on_part_added` makes: a stroke into an invisible mask is
|
||||
indistinguishable from a tool that did nothing. Pressing a swatch opens the
|
||||
eye too, since colouring a mask nobody can see would change no pixel.
|
||||
|
||||
Only ever *this* layer's eye, and only on an explicit action. Every other eye
|
||||
keeps whatever it was set to, and nothing re-arms in the background. The
|
||||
existing `overlay-hidden` property is the precedent and the trap it documents
|
||||
applies unchanged: an automatic reveal that re-arms a switch somebody turned
|
||||
off is worse than no automatic reveal at all.
|
||||
|
||||
Viewing state and not edit state: eyes and colours are on the session, not on
|
||||
the layer, and a photograph reopened has every eye closed.
|
||||
|
||||
The ~1s reveal after a shaping slider is released, from the draft, is not
|
||||
built. It is a timer rather than a decision.
|
||||
|
||||
The region overlay stays exactly what it is — a picture of what the model
|
||||
detected — and gains a name in the interface that says so, because two
|
||||
overlays that look alike and mean different things is worse than either.
|
||||
|
||||
## 7. The interface
|
||||
|
||||
### 7.1 Where the tools live
|
||||
|
||||
**In the Local panel, not the tool rail.** The rail's entries arm a canvas
|
||||
gesture for the whole photograph; a brush is meaningless without a layer to
|
||||
paint into, and a rail entry that silently created one — or that lit up and
|
||||
did nothing with no layer selected — is precisely the kind of surprise the
|
||||
rail's own notes argue against. The tool strip sits in the Local panel's
|
||||
header, enabled only when a part is selected:
|
||||
|
||||
```
|
||||
┌ Local ─────────────────────────────┐
|
||||
│ [Select] [Paint] [Erase] [Push] ◉ │ ◉ = show mask
|
||||
├────────────────────────────────────┤
|
||||
│ ▸ Sky ● ⌄ │
|
||||
│ Category · sky │
|
||||
│ ├ Subtract · Painted ✕ │
|
||||
│ └ Intersect · Luminance ✕ │
|
||||
│ [+ Add ⌄] [− Subtract ⌄] [∩ ⌄] │
|
||||
├────────────────────────────────────┤
|
||||
│ Brush size ──●─── 0.05 │
|
||||
│ hardness ──●── 0.5 │
|
||||
│ flow ─────● 1.0 │
|
||||
│ cling ──●─── 0.4 │
|
||||
└────────────────────────────────────┘
|
||||
```
|
||||
|
||||
`toolrail.slint`'s note about `MaskSource::Brush` being the next rail entry is
|
||||
superseded by this and should be replaced with the reasoning, not deleted —
|
||||
the file is where somebody will next look for it.
|
||||
|
||||
### 7.2 The part list
|
||||
|
||||
Each layer row gains its parts as indented rows. A part row carries its join
|
||||
(a chip that cycles add / subtract / intersect), its source name, a delete,
|
||||
and selection — selecting a part is what points the brush and the shaping
|
||||
controls at it. The layer row keeps invert, opacity and enable, which are the
|
||||
layer's.
|
||||
|
||||
`[+ Add]`, `[− Subtract]` and `[∩]` each open the same menu of sources the
|
||||
"new layer" buttons already offer: a gradient, a range, a subject, a category,
|
||||
or painted. One code path, three joins.
|
||||
|
||||
**Folding two layers into one** uses the multi-selection
|
||||
[`masks_ui.rs`](../../ui/dr-ui/src/masks_ui.rs) already supports: with two layers
|
||||
selected, "Combine" appends the second's parts to the first and removes it.
|
||||
Offered only when the second layer's adjustments are neutral, and otherwise
|
||||
offered with a warning that names what will be lost — quietly discarding an
|
||||
edit the user made is not a combine.
|
||||
|
||||
### 7.3 The brush and the cursor
|
||||
|
||||
Radius, hardness, flow and cling are sliders in the panel and all four are
|
||||
live on the canvas as a cursor: an outer ring at the radius, an inner ring at
|
||||
the hardness, and — this matters on a phone — the ring drawn at the *touch
|
||||
point offset above the finger*, since the thing being painted is under the
|
||||
hand that is painting it.
|
||||
|
||||
Radius is stored per tool, not per stroke and not per layer: a photographer
|
||||
who sets a small eraser expects it to still be small the next time they erase.
|
||||
|
||||
### 7.4 Gestures
|
||||
|
||||
Each of these needs a `GESTURE:` block beside its implementation — that is the
|
||||
only place [gestures.md](../gestures.md) can be written from.
|
||||
|
||||
| Gesture | Touch | Pointer | Keyboard |
|
||||
|---------|-------|---------|----------|
|
||||
| Paint into the selected part | Drag on the picture | Drag | — |
|
||||
| Erase instead of paint | Hold the Erase tool | Alt-drag | — |
|
||||
| Push the boundary | Drag with Push armed | Drag | — |
|
||||
| Change the brush size | Drag the size slider | Scroll with Alt over the picture | `[` `]` |
|
||||
| See the mask | Press the eye | Press the eye | `\` while held |
|
||||
| Undo one stroke | The history list | Ctrl+Z | Ctrl+Z |
|
||||
| Add a part | `[+ Add]`, pick a source | Same | — |
|
||||
| Fold two layers | Select both, Combine | Same | — |
|
||||
|
||||
The `why` each block needs is mostly one sentence — *a stroke is a decision
|
||||
and a decision is one undo step* — except for Alt-drag, which needs to say
|
||||
that a modifier is the only way to alternate paint and erase without leaving
|
||||
the stroke, and that touch cannot have it, which is why the tool strip is a
|
||||
strip and not a single toggle.
|
||||
|
||||
### 7.5 The Slint hazards this walks into
|
||||
|
||||
Named because all of them compile:
|
||||
|
||||
- **The paint `TouchArea` must be declared in front of
|
||||
`ScaleRotateGestureHandler`** to receive the press at all, exactly as the
|
||||
region picker and the repair placer are. The consequence is that a
|
||||
two-finger pinch may not reach the handler behind it while paint is armed.
|
||||
Repair has the same arrangement today, so **check what repair mode actually
|
||||
does with a pinch before designing around it**; if pinch is lost, the answer
|
||||
is geometry — arm the paint area over the picture only — not z-order.
|
||||
- **A drag must be measured in the parent frame**, never in the coordinates of
|
||||
something the drag moves. `GradientHandles` in `masks.slint` is the
|
||||
reference.
|
||||
- **The part rows are a repeater inside the develop column**, so their chips
|
||||
go through `ChipGrid` or `Segmented { columns: 3 }` — a non-wrapping chip
|
||||
row sets the width of the whole sidebar.
|
||||
- **None of the above is caught by the test suite.** A screenshot is the only
|
||||
check; see the project's notes on capturing one under XWayland.
|
||||
|
||||
## 8. History, undo, labels
|
||||
|
||||
A stroke is one step, recorded on release: `Edit::Action`, not
|
||||
`Edit::Control`. `Control` is the right variant for a dragged control and the
|
||||
wrong one here precisely because it coalesces — two strokes painted a second
|
||||
apart are two decisions, and coalescing them under one key would make the
|
||||
second untakeable back. A push stroke follows the same rule, as does adding,
|
||||
removing or re-joining a part.
|
||||
|
||||
The mask stack is already snapshotted per step and shared by `Arc` when a step
|
||||
does not touch it, so the cost of an undoable stroke is a clone of one layer's
|
||||
parts, not of the picture. New keys in
|
||||
[`labels.rs`](../../ui/dr-ui/src/labels.rs):
|
||||
|
||||
```
|
||||
history.mask_painted "Paint Mask"
|
||||
history.mask_erased "Erase Mask"
|
||||
history.mask_pushed "Push Mask Edge"
|
||||
history.mask_part_added "Add To Mask"
|
||||
history.mask_part_removed "Remove From Mask"
|
||||
history.mask_joined "Change How Mask Joins"
|
||||
history.masks_combined "Combine Masks"
|
||||
```
|
||||
|
||||
## 9. Sync and merge
|
||||
|
||||
Nothing here stores pixels, so nothing here changes what sync carries beyond
|
||||
size. A painted correction is bounded by `MAX_LAYER_POINTS` at roughly 50 kB
|
||||
of text in the worst case and a few hundred bytes in the ordinary one. A model
|
||||
part still carries its run-length coded `coverage` line, unchanged and still
|
||||
outside `PartialEq`.
|
||||
|
||||
Two conflicts are new, and both resolve per block: two devices adding
|
||||
different parts to one layer (both are kept, in block order), and two devices
|
||||
painting the same part (a conflict over that part, arbitrated the way a layer
|
||||
already is). A device that has never run a model still renders every part
|
||||
correctly, because coverage travels with the part exactly as it travelled with
|
||||
the layer.
|
||||
|
||||
## 10. Performance and budgets
|
||||
|
||||
The number that matters is the cost of one dab while the finger is down, and
|
||||
with §5.4 it is bounded by the dab's own bounding box rather than by the
|
||||
stroke's history: `(2r)²` pixels × one segment. At the default 5% radius on a
|
||||
1600×1067 proxy that is about 25 000 pixels — a fraction of a millisecond, and
|
||||
independent of how long the stroke has been going.
|
||||
|
||||
Everything else runs on a shape change, not per frame:
|
||||
|
||||
| Work | When | Rough cost at 1600×1067 |
|
||||
|------|------|------------------------|
|
||||
| Full layer rebuild | part added, source changed, resize | one draw per part |
|
||||
| Warp | a push part exists and something below it changed | one copy + one bbox draw |
|
||||
| Distance field | a model part's morphology changed | as today, CPU, per part |
|
||||
| Overlay | never — it is two uniforms in a shader that already runs | — |
|
||||
|
||||
The scratch texture for warp is one proxy-sized `r8unorm` — under 2 MB — and
|
||||
is allocated on first use, so a photographer who never pushes an edge never
|
||||
pays for it.
|
||||
|
||||
Measure against [frame-budget.md](frame-budget.md), and measure it on an idle
|
||||
machine: the 16 ms guard is load-sensitive enough that a busy build makes it
|
||||
fail and an idle one makes it pass, so a single run proves nothing either way.
|
||||
|
||||
## 11. Order of work
|
||||
|
||||
**M1 — See it and paint it.** The overlay (§6), `parts` on the layer with
|
||||
`Union` and `Subtract`, `Painted` parts, the tool strip, the brush HUD and
|
||||
cursor, the canvas gestures, incremental raster (§5.4). No new shader code for
|
||||
the mask pass beyond the blend variants; no change to distance fields, because
|
||||
a painted part needs none. *This is the whole of the user-visible ask except
|
||||
push and intersect,* and it is deliberately the milestone that stands alone.
|
||||
|
||||
Done: parts with `Union` and `Subtract`, painted parts, the tool strip, the
|
||||
canvas gesture, and the overlay (§6). Outstanding: the brush HUD and cursor —
|
||||
radius, hardness and flow are sliders with no ring drawn on the photograph,
|
||||
so the size of the brush is a number rather than a thing you can see — and
|
||||
the incremental raster of §5.4, without which a layer's whole stroke history
|
||||
is redrawn per dab.
|
||||
|
||||
One lesson from the order it was actually built in, since §6 said it and the
|
||||
build did not listen: the overlay is not the last quarter of M1, it is the
|
||||
first. Parts and painting shipped without it and the result was a feature
|
||||
nobody could tell was working — the panel listed a mask, the photograph showed
|
||||
nothing, and every report of it came back as "the masks do not work".
|
||||
|
||||
**M2 — Combine properly.** `Intersect`, model and gradient and range parts,
|
||||
per-part distance fields (§5.5), the part list with its join chips, folding
|
||||
two layers.
|
||||
|
||||
**M3 — Push, and cling.** The warp pass (§5.3) and edge-aware deposit (§5.6).
|
||||
Both are refinements of a tool that already works, which is the right place
|
||||
for the two riskiest pieces.
|
||||
|
||||
**M4 — What the use of it asks for.** Left open on purpose. Likely candidates:
|
||||
a "select the subject under the pointer" brush, pressure from a stylus, and a
|
||||
per-part opacity.
|
||||
|
||||
## 12. Requirements to add
|
||||
|
||||
Proposed text for [requirements.md](requirements.md) §3.3, in the shape the
|
||||
neighbouring entries take:
|
||||
|
||||
> **FR-DEV-19 — Mask editing.** A mask layer's coverage shall be editable by
|
||||
> hand after it is created, by painting into it, erasing from it, dragging its
|
||||
> boundary, and joining further selections to it. Every edit is stored as
|
||||
> geometry and parameters in the edit graph; no rasterised mask is written to
|
||||
> a file and none exists in CPU memory.
|
||||
>
|
||||
> **FR-DEV-19a — Mask composition.** A layer's mask is an ordered list of
|
||||
> parts, each naming a source and how it joins the mask before it — union,
|
||||
> subtraction, or intersection. A layer of one part is exactly the layer of
|
||||
> today, and reads and writes the same sidecar.
|
||||
>
|
||||
> **FR-DEV-19b — Hand correction.** A part may be painted, with add, erase and
|
||||
> push strokes, at a radius, hardness, flow and edge-clinging the photographer
|
||||
> sets. Strokes are stored as normalised coordinates and rasterised on the
|
||||
> device.
|
||||
>
|
||||
> **FR-DEV-19c — Boundary push.** A push stroke displaces the mask beneath it
|
||||
> along the drag, filling behind an outward drag and emptying behind an inward
|
||||
> one, on any mask source rather than only those with a distance field.
|
||||
>
|
||||
> **FR-DEV-19d — Mask visualisation.** The mask a layer actually produces —
|
||||
> after its parts, its shaping, its inversion and its opacity — shall be
|
||||
> displayable over the photograph as a tint, as an alpha, or as an outline,
|
||||
> and shall appear automatically while a mask is being edited.
|
||||
|
||||
Each needs `TRACES:` comments at the implementation and a regenerated
|
||||
[traceability.md](traceability.md) **in the same commit**, since that file
|
||||
carries line numbers.
|
||||
|
||||
## 13. Verification
|
||||
|
||||
Testable without a GPU, and therefore not optional:
|
||||
|
||||
- `Stroke` mode round-trips through the sidecar, including a push stroke and a
|
||||
non-zero cling; an unknown mode costs one stroke and not the layer.
|
||||
- A one-part layer writes the byte-identical sidecar it writes today, and
|
||||
every existing fixture in `mask_sidecar.rs` still loads.
|
||||
- Parts merge per block: different parts added on two devices produce a layer
|
||||
with both; the same part edited on both is one conflict.
|
||||
- Point budgets hold across parts, and painting past them refuses rather than
|
||||
dropping the oldest strokes.
|
||||
- The panel model builds part rows from a hand-made stack with no device
|
||||
present, the way `rows_from` does for adjustments.
|
||||
- History: a stroke is one step, undo restores the parts, and a step that
|
||||
leaves the masks alone still shares the `Arc`.
|
||||
|
||||
Needing a device, in `dr-gpu/tests`:
|
||||
|
||||
- A truth table for the joins: a half-covering part unioned, subtracted and
|
||||
intersected with a known base gives the three expected fields.
|
||||
- A warp of a known step edge moves it by the drag, and the area behind it
|
||||
fills — asserted on a readback of the mask array, which is a test reading
|
||||
back, not the application.
|
||||
- Incremental drawing equals a full rebuild, dab for dab, on a stroke of a few
|
||||
hundred points. This is the test that keeps §5.4 honest.
|
||||
- `Max` blending behaves on every adapter CI runs on.
|
||||
|
||||
And a screenshot, because §7.5 is not reachable by any of the above. Run
|
||||
`dr-ui` tests single-threaded; parallel runs segfault in this crate for
|
||||
unrelated reasons.
|
||||
|
||||
## 14. Decisions I need
|
||||
|
||||
1. **Parts now, or strokes first?** M1 as written introduces `parts` and pays
|
||||
the migration once. The cheaper alternative is a `strokes: Vec<Stroke>`
|
||||
field on the layer, brushed over whatever the source produced, and parts
|
||||
later — half the code, and it makes "subtract this from that" a second
|
||||
mechanism invented afterwards. I would take the migration now.
|
||||
2. **Push as a warp, or local morphology on model masks only?** §5.3 argues
|
||||
the warp. It is the more expensive of the two and the only one that works
|
||||
on every source.
|
||||
3. **Tool strip in the Local panel, or an entry in the rail?**
|
||||
`toolrail.slint` predicts a rail entry; §7.1 argues against it. The rail is
|
||||
easier and worse.
|
||||
4. **Do the FRs in §12 go into `requirements.md` now,** or stay proposed here
|
||||
until the first milestone lands?
|
||||
@@ -0,0 +1,495 @@
|
||||
# DarkRoom — Outstanding work
|
||||
|
||||
**Status:** Living document · first written 2026-08-29
|
||||
**Companion to:** [requirements.md §7](requirements.md), [technical-debt.md](technical-debt.md),
|
||||
[traceability.md](traceability.md)
|
||||
|
||||
What is specified and not built, and for each cluster whether that is a decision, a dependency, or a
|
||||
gap nobody has looked at.
|
||||
|
||||
This document exists because [traceability.md](traceability.md) cannot tell those apart. It reports
|
||||
one number — the share of requirements carrying a `TRACES` tag — and a missing tag means either
|
||||
"nobody has built this" or "somebody built it and did not say so". Both read the same way in the
|
||||
summary table, which makes that figure pessimistic *and* uninformative at once: it understates what
|
||||
works while hiding which of the remainder matters. Eight requirements gained a tag on this branch
|
||||
because the code already satisfied them and nobody had said so. Everything below is the other kind.
|
||||
|
||||
It is also not a plan. [requirements.md §7](requirements.md) records what was deferred deliberately
|
||||
and needs no argument; this records what is still nominally in scope, so that the distance between
|
||||
the register and the binary is visible rather than something a reader has to reconstruct from a
|
||||
percentage. Where the honest answer is "this requirement should be amended rather than met", it says
|
||||
so — an unbuilt requirement that nobody intends to build is worse than a deferred one, because it
|
||||
keeps costing attention.
|
||||
|
||||
**Several entries were struck between 2026-08-29 and 2026-08-30**, as two waves of work
|
||||
landed: burst grouping (FR-CULL-5), Flatpak packaging (FR-PLAT-LIN-3), FR-CULL-3 in full,
|
||||
Android memory pressure and lost-root recovery (FR-PLAT-AND-5, FR-PLAT-AND-2), image intents
|
||||
(FR-PLAT-AND-6), and the job runner (FR-PLAT-AND-4's Rust half). What remains of each is
|
||||
recorded where it appears rather than deleted, because a requirement that is *half* met is the
|
||||
one most likely to be reported as closed.
|
||||
|
||||
---
|
||||
|
||||
## 1. Plugins — post-v1 since 2026-09-19
|
||||
|
||||
> **Resolved, in the register.** The contradiction below was settled on 2026-09-19 the way the
|
||||
> last paragraph of this section asked: §3.10 is marked `(post-v1)` clause by clause, §7's row
|
||||
> says so with a reason, D16 defers with it, and the traceability tool lists deferred clauses in
|
||||
> their own table instead of counting them. Coverage went from 72.2% of 194 to 80.6% of 170 on
|
||||
> that edit alone. What follows is kept as the record of what was decided and why; nothing in it
|
||||
> is owed a tag.
|
||||
|
||||
|
||||
**Untagged:** FR-PLG-1, -1a, -2a, -2b, -2c, -3, -3a, -4, -4a, -5, -5a, -5b, -5c, -6, -6a, -7, -8,
|
||||
-9, -10, -11, -12.
|
||||
|
||||
No plugin host exists. No crate loads anything at runtime: there is no manifest reader, no WASM or
|
||||
Lua engine, no registry, no signature check, no install path, no capability grant, no per-plugin
|
||||
failure ledger. `declared/mod.rs` says as much in its own documentation — the operation format is
|
||||
"not a plugin directory read at startup".
|
||||
|
||||
**Two of §3.10's requirements are met, and they are the interesting two.** FR-PLG-2 and FR-PLG-2d —
|
||||
the declarative node format — are built and tagged: `core/dr-pipeline/ops/*.yaml` compiled by
|
||||
`build.rs`, with the restricted expression grammar in `declared/expr.rs` and a parity test asserting
|
||||
a declared operation and a hand-written one produce identical output.
|
||||
[code-health.md §3](code-health.md) calls it "a working plugin system that happens to resolve at
|
||||
build time", and that is exactly right. What is missing is not the format; it is everything that
|
||||
would let somebody who is not in this repository use it.
|
||||
|
||||
**The contradiction.** [requirements.md §7](requirements.md) lists `| Plugin API | — |` among the
|
||||
things deferred for v1 — a bare row, where most deferrals carry a justifying note. §3.10 then spends
|
||||
roughly 280 lines and 23 requirement IDs specifying that same Plugin API in detail. Both statements
|
||||
are in the register of record, and the traceability denominator counts the second one: 21 IDs, 12%
|
||||
of all 179 defined requirements, worth about twelve points of coverage on their own — and nearly a
|
||||
third of everything the matrix reports as uncovered. A reader looking at the coverage figure has no
|
||||
way to know that, or that the subsystem behind it is one the same document says is not in this
|
||||
version.
|
||||
|
||||
**And D16 is open.** [Decision D16](requirements.md) — plugin licensing — records that GPLv3
|
||||
answers the derivative-work question differently for each of §3.10's three plugin forms, and that
|
||||
this "must be answered *before* an ecosystem exists, not after", because a term introduced later
|
||||
cannot be applied to plugins already written. D16 explicitly does not block FR-PLG-2; it blocks
|
||||
publishing a third-party format as stable.
|
||||
|
||||
**What would resolve this:** an edit to `requirements.md`, not code. Either §7 drops the row, or
|
||||
§3.10 is marked deferred with the two built requirements carved out. Until one of those happens the
|
||||
coverage figure is measuring a decision that has already been taken, and taking it again every time
|
||||
somebody reads the matrix.
|
||||
|
||||
---
|
||||
|
||||
## 2. Culling — the stated differentiator, half built
|
||||
|
||||
[D11](requirements.md) names culling "the core differentiator". FR-CULL-1, -2, -3, -4 and -8 through
|
||||
-12 are built. Three are not.
|
||||
|
||||
**FR-CULL-3 — Raw-truth overlays. Built, all three bullets.** Focus peaking is
|
||||
`core/dr-gpu/src/focus.rs` and `ui/dr-ui/src/peaking.rs`; the raw histogram and the raw clipping
|
||||
indicators are `core/dr-gpu/src/raw_histogram.rs` and the second reading of the panel in
|
||||
`histogram.slint`.
|
||||
|
||||
Worth recording, because it is the thing this entry previously got wrong and the next reader will
|
||||
have to check again. The display histogram — `dr-gpu/src/histogram.rs`, `ui/dr-ui/src/histogram.rs`
|
||||
— is **not** this requirement and never was: it is tagged FR-DSP-7, it reads `AdjustPass`'s 8-bit
|
||||
output, it counts clipping as `r == 255`, and so it describes the frame the display is about to
|
||||
show, after the whole develop chain. FR-CULL-3 asks for the *sensor data*, on the explicit grounds
|
||||
that a rendered image "systematically lies about what is recoverable in the raw". The two now sit in
|
||||
one panel behind a chip row, which is the arrangement that keeps them from being mistaken for each
|
||||
other: they answer different questions and both are true.
|
||||
|
||||
What the raw reduction cannot answer is written down rather than left to be discovered —
|
||||
[architecture.md §5.5](architecture.md) records why it reduces over the demosaiced texture instead
|
||||
of the CFA samples §5.5 originally specified, and what that costs in what it can say.
|
||||
|
||||
**FR-CULL-5 — Burst and near-duplicate grouping. Built.** `core/dr-catalog/src/bursts.rs`:
|
||||
frames join a burst when they are adjacent in time *and* look like the frame before them, compared
|
||||
adjacent-pair-only in one ordered walk. Signatures are 64-bit difference hashes taken from the
|
||||
thumbnails `dr-thumbs` already holds, on a background pass after the thumbnail sweep — nothing at
|
||||
import, nothing at query time. Nothing scores or rejects a frame: the representative is the
|
||||
earliest, a fact about the clock, and a new burst arrives *open* so the pass never takes a row off
|
||||
the screen.
|
||||
|
||||
Two threads left hanging. `core/dr-face/src/calibrate.rs` still says "since FR-CULL-5 already
|
||||
groups bursts, positives are bootstrapped from bursts" while in fact bootstrapping from confirmed
|
||||
labels — that comment was a forward reference and is now simply wrong, rather than premature.
|
||||
And `dr_catalog::bursts::choose_representative` is written and tested but bound to no gesture, so
|
||||
today the only override is expanding the burst.
|
||||
|
||||
`core/dr-catalog/src/dedup.rs` remains a different thing: re-import detection under FR-CAT-11,
|
||||
matching a file against one already catalogued, not two photographs against each other.
|
||||
|
||||
**FR-CULL-6 — Compare and survey.** Absent. No side-by-side view, no synchronised zoom or pan.
|
||||
This is the one of the four with no adjacent machinery at all, and it is also the one that most
|
||||
directly distinguishes culling from browsing.
|
||||
|
||||
**FR-CULL-7 — Culling on tablet.** Absent, and blocked by the three above rather than independent
|
||||
of them: there is no separate tablet culling surface to build until there is something to put on it.
|
||||
|
||||
---
|
||||
|
||||
## 3. FR-DEV-3g — AI denoise
|
||||
|
||||
Promoted into v1 by [D11](requirements.md), and named there as the precondition for deferring AI
|
||||
masking — the argument being that one learned stage earns the runtime that a second could then
|
||||
reuse. Only classical noise reduction exists: `ops/noise_reduction.rs`, a bilateral filter in two
|
||||
arrangements, exact for luminance and separable for chroma. It is good, and it is not this.
|
||||
|
||||
`models/` holds two face models and nothing else; `core/dr-segment/models/` holds a YOLO
|
||||
segmentation model for subject masks. There is no denoise model, no learned demosaic, and no
|
||||
inference path that is not face or segmentation.
|
||||
|
||||
The obstacle is not the pipeline. It is that [D13](requirements.md) — model licensing — is still
|
||||
open for the models that already ship, and adding a third learned stage adds a third licence to
|
||||
answer for. Building the runtime before that is settled means owning the same problem in one more
|
||||
place.
|
||||
|
||||
---
|
||||
|
||||
## 4. The render path — FR-DSP-2, FR-DSP-4, NFR-RES-2
|
||||
|
||||
**FR-DSP-2 — Tiled computation. Unbuilt, and under challenge.** [architecture.md §6.2](architecture.md)
|
||||
calls for tiling "from day one" on the grounds that retrofitting it is a rewrite. It was not built,
|
||||
and the evidence has since moved. `core/dr-gpu/tests/frame_budget.rs` carries the argument in its
|
||||
own header: one fused dispatch over a viewport-sized target is comfortably inside the frame budget,
|
||||
and "if that stops being true, the recommendation to strike tiled computation from the interactive
|
||||
path stops being supported, and this test is what says so."
|
||||
[technical-debt.md TD-4](technical-debt.md) reaches the same place from the other direction — a
|
||||
tiled convolution at clarity's radius reads nearly twice the taps that an untiled one does, so the
|
||||
stage that looks most like it wants a tile cache is the stage that would be hurt most by one.
|
||||
|
||||
What exists is the declaration and not the mechanism: `DetailPass::radius` is documented as the halo
|
||||
a tile would have to be grown by, with a test that pins it, and there is no scheduler to read it.
|
||||
That is deliberate plumbing, not an oversight.
|
||||
|
||||
**So the open question here is not "when is tiling built" but "is FR-DSP-2 still a requirement" — and on 2026-09-19 the answer was: as written, until S6 runs.** FR-DSP-2 now carries a status note saying exactly that, and R5's note no longer claims it was rewritten.
|
||||
Two measurements say it costs more than it saves on the interactive path. Neither says anything
|
||||
about the export path or about a device under memory pressure, which is where the case for it
|
||||
actually lives — and that is spike S6, which has not run.
|
||||
|
||||
**FR-DSP-4 — Progressive refinement.** Unbuilt. FR-DSP-1's proxy rendering and TD-4's
|
||||
quarter-resolution base are adjacent and are not it: both are fixed choices about what resolution to
|
||||
compute at, where FR-DSP-4 asks for a first frame that is deliberately cheap and a second that
|
||||
replaces it. Nothing tracks a "this frame is provisional" state.
|
||||
|
||||
**NFR-RES-2 — Images larger than GPU memory.** Half answered. NFR-R8's "decide explicitly" was
|
||||
decided on 2026-09-19: there is no CPU render pipeline, the degraded mode is the viewer on
|
||||
embedded previews with develop withheld, and NFR-RES-2 no longer promises a fallback render. What
|
||||
remains unbuilt is the memory half: there is no headroom budget, no allocation-failure staging,
|
||||
and no spill. Spike S6 — a tiled pipeline on a
|
||||
mid-range Android device with an image larger than available GPU memory — is the one that would
|
||||
settle both this and FR-DSP-2, and there is no evidence it has run.
|
||||
|
||||
---
|
||||
|
||||
## 5. Android beyond running, and Flatpak
|
||||
|
||||
The Android app is not a stub — it builds an APK, runs the whole application, unpacks bundled face
|
||||
models, and has been measured on a tablet ([faces.md §12.1](faces.md),
|
||||
[technical-debt.md TD-1](technical-debt.md)). What is missing is the platform contract around it.
|
||||
|
||||
**FR-PLAT-AND-1 is untagged, and what it was tagged for was intent rather than code.**
|
||||
The requirement demands that library access be obtained *exclusively* through the Storage Access
|
||||
Framework. There is no SAF code: no `ACTION_OPEN_DOCUMENT_TREE`, no `takePersistableUriPermission`,
|
||||
no `DocumentsContract`. Its two tags rested on a `SourceRef::Document` variant constructed only
|
||||
inside `#[cfg(test)]` — `LocalStorage::open` refuses it, and the test that proves so is named
|
||||
`a_reference_of_the_wrong_kind_is_refused_rather_than_guessed_at` — and on
|
||||
`dr_plat::imports_supported`, which *returns false on Android* and whose own documentation says it
|
||||
"stops being false when a SAF implementation lands". The second tag documented the absence of the
|
||||
thing it was counted as evidence for. Both have been removed; this is the "plumbing a future feature
|
||||
would use" case [CONTRIBUTING.md](../../CONTRIBUTING.md) and [code-health.md CH-4](code-health.md) both
|
||||
warn about. Android reaches a library through a Nextcloud account or a folder, over paths, like the
|
||||
desktop.
|
||||
|
||||
That has a consequence for the rest of the cluster: **FR-PLAT-AND-2** — detecting the loss of a
|
||||
granted tree permission and marking images offline rather than deleting rows — cannot be built until
|
||||
there is a permission to lose. It is listed here as unbuilt, but it is blocked, not skipped.
|
||||
|
||||
**FR-PLAT-AND-4 — half built.** The runner is done (`core/dr-catalog/src/runner.rs`): the
|
||||
queue that `jobs.rs` always had is now claimed from, completed, failed and recovered after a
|
||||
crash, which is FR-PLAT-AND-3's resumability as much as this requirement's. What is missing is
|
||||
the platform half — a foreground `Service`, `FOREGROUND_SERVICE` and `POST_NOTIFICATIONS` in the
|
||||
manifest, and a stated Doze behaviour. The build step that blocked it is no longer a blocker: the
|
||||
APK now compiles its own Java.
|
||||
|
||||
Note also that **no handler is registered**, deliberately. The only enqueue site reachable in the
|
||||
shipping app produces remote thumbnail jobs already served by the async grid worker, and
|
||||
`walk::scan_root` — which holds the other two enqueue sites — has no caller outside an example.
|
||||
Wiring the sweep to claim from the queue is the honest next step and is an async rewrite of
|
||||
`library.rs`.
|
||||
|
||||
**FR-PLAT-AND-5 — built.** A tiered eviction registry drives GPU caches, then proxies, then
|
||||
thumbnails, from `MainEvent::LowMemory` and `MainEvent::Stop`.
|
||||
|
||||
**FR-PLAT-AND-6 — built, with one half unwired.** VIEW, SEND and SEND_MULTIPLE filters, the launch
|
||||
Intent read over JNI, and an `ExportProvider` rooted at `getFilesDir()` rather than AndroidX's
|
||||
`FileProvider`. The outbound share has no caller in `ui/` yet. **None of the runtime behaviour has
|
||||
been exercised on a device** — the tests read the manifest and the Java through `include_str!`,
|
||||
which catches a deleted filter but not a class loader that cannot find the class.
|
||||
|
||||
**FR-PLAT-LIN-3 — packaged, not satisfied.** There is a Flatpak manifest now, granting no
|
||||
filesystem permission of any kind, plus AppStream metainfo and `docs/distribution.md`. The
|
||||
requirement is still not met, and cannot be met by packaging: a folder library is chosen by typing
|
||||
an absolute path, nothing in the tree calls the FileChooser portal, and inside the sandbox `$HOME`
|
||||
holds only `.var/app/...`. `dr_plat::volumes()` reads `/proc/self/mountinfo`, so a card mounted on
|
||||
the host is invisible to a sandboxed process as well. The fix is an `ashpd` directory picker beside
|
||||
`LocalStorage::grant`, not a change to the manifest. No Flatpak has been built here —
|
||||
`flatpak-builder` is not installed — so the permission set is reasoned, not observed.
|
||||
|
||||
**NFR-COMPAT-2 — distribution channels. Stated, which is all this requirement asks.** The paragraph
|
||||
above cites [distribution.md](distribution.md) and it is the same document that answers this: §1
|
||||
names five channels and their state — Arch source package and Flatpak in tree, AppImage a v1 channel
|
||||
whose recipe is not written, F-Droid a v1 channel not yet submitted, and Play explicitly **not** v1.
|
||||
The requirement is to *state* the channels, and they are stated, including the deferral.
|
||||
|
||||
What remains is the coupling the requirement points at rather than the statement it demands. §4.8
|
||||
observes that publishing on Play is what turns SAF from a preference into a constraint, and
|
||||
distribution.md §6 argues the coupling runs the other way for this project — F-Droid asks nothing
|
||||
that ARCH §6.9 does not already require. Spike S11, the Play permissions dry-run, has not run, and
|
||||
until it does that argument is reasoned rather than confirmed. Two of the five channels also exist
|
||||
as decisions rather than as recipes, and no Flatpak has been built here at all.
|
||||
|
||||
Related, NFR-COMPAT-1's baseline is real but scattered — API 28/36 live in the Android
|
||||
Dockerfile and are checked in CI against the built ELF, which is good — while the items the
|
||||
requirement singles out are missing: whether `shaderFloat16` and 16-bit storage are required (the
|
||||
one it flags as jeopardising R1), minimum RAM, minimum desktop Mesa, and a named reference device
|
||||
from a second GPU vendor.
|
||||
|
||||
**NFR-OPS-2 is met, and NFR-OPS-4 is not.** Crash reporting is `platform/dr-plat/src/crash.rs`: a
|
||||
panic on either platform writes a local record with a redacted message and backtrace, ten are kept,
|
||||
and there is deliberately no upload path — the requirement's "upload only on explicit opt-in" is
|
||||
satisfied by there being nothing to opt into, and the module says why a transport built ahead of the
|
||||
consent is the wrong order. (This paragraph said the opposite until 2026-09-12; the record had landed
|
||||
on 2026-08-30 and the paragraph had not been read against it.) Update and first run are undefined;
|
||||
the concrete reason NFR-OPS-4 gives — that D2 pins rawler at a non-SemVer alpha whose camera-support
|
||||
fixes users will need — is unaddressed, and there is no update mechanism of any kind.
|
||||
|
||||
---
|
||||
|
||||
## 6. Accessibility and internationalisation — the hard half is done and the easy half is not
|
||||
|
||||
**NFR-A11Y-1 — Localisation.** `@tr(` appears **zero** times across 14,482 lines of Slint. That
|
||||
number overstates the problem, because the part that is genuinely architectural was got right:
|
||||
`LocalizedKey` keeps display strings out of `core/` entirely, every operation publishes a key rather
|
||||
than a label, and `labels::resolve` is the single point where a key becomes text. What that single
|
||||
point does, however, is a hardcoded English `match` in Rust source — so changing a translation
|
||||
requires a recompile, which is the one thing the requirement explicitly forbids. There is no message
|
||||
catalogue in any format, no locale-resolution rule, and no decision recorded about RTL.
|
||||
|
||||
The work left is therefore smaller than it looks and entirely mechanical: a catalogue format, a load
|
||||
path behind `resolve`, and `@tr(` around the Slint literals. The design it needs already exists.
|
||||
|
||||
**NFR-A11Y-2 — Accessibility.** `accessible-*` appears five times in the whole interface, all five
|
||||
on one control — the parameter slider in `adjust.slint` — and nothing is set from the Rust side at
|
||||
all. Everything else in eighteen Slint files is unnamed to AT-SPI and TalkBack. The requirement's own
|
||||
caveat, that Slint's Android accessibility needs verifying, is spike S13, which has not run.
|
||||
|
||||
**NFR-A11Y-3 — Colour-independent status.** Built where a control exists, and now tagged: the
|
||||
clipping readout pairs a marker that appears or disappears with a figure in words, the rating strip
|
||||
is a solid star against an outline in an achromatic palette, the pick/reject mark is a tick against
|
||||
a cross, and the focus-peaking colour chips say "Red" and "Cyan" rather than showing swatches. Each
|
||||
of those already carried the reasoning in a comment naming this requirement and simply had no
|
||||
`TRACES` line.
|
||||
|
||||
Two caveats, because the tag now says more than the evidence does. **Only the clipping clause has a
|
||||
test** — `a_clipping_figure_distinguishes_none_from_nearly_none`, which pins `<0.1%` apart from `0%`
|
||||
so the figure cannot contradict the lit marker beside it. The three Slint components are
|
||||
inspected-and-argued, not asserted, and nothing would fail if a future edit made a star differ only
|
||||
in tint. **And the requirement's first named example has no interface at all**: catalog colour
|
||||
labels are a nullable `label INTEGER` column on the versions table and are set and shown nowhere, so
|
||||
the clause about them is untestable rather than satisfied. That clause closes when the label UI is
|
||||
built, not before, and it should be built with a shape from the outset — which is the same argument
|
||||
as below, for doing this alongside NFR-A11Y-2 rather than after it.
|
||||
|
||||
---
|
||||
|
||||
## 7. Catalog and sync
|
||||
|
||||
**FR-CAT-14 — Migration import.** Reading ratings, labels, keywords and collections out of a
|
||||
Lightroom `.lrcat` or a darktable `library.db`. Unbuilt. The destination is not: keywords,
|
||||
collections, ratings and the cross-device merge rules are all built and tested, and
|
||||
`keywords.rs` already anticipates the arrival ("an import from Lightroom can bring in…"). What is
|
||||
missing is only the two source adapters — which is a comparatively contained piece of work for a
|
||||
requirement that decides whether somebody can try this software on a library they already have.
|
||||
|
||||
**FR-NC-11 — Initial catalog build.** Using WebDAV `SEARCH` (RFC 5323) against `/remote.php/dav/`,
|
||||
filtered by mimetype and paginated, in preference to walking folders with PROPFIND. Unbuilt: no
|
||||
`SEARCH` request is issued anywhere. The PROPFIND walk this exists to replace is fully built and
|
||||
well optimised — ETag pruning under FR-NC-4 turns an unchanged 50k library into one request — so the
|
||||
gap is narrower than it reads. It is the *first* build against a large remote library that pays, and
|
||||
that is the moment a new user meets.
|
||||
|
||||
**FR-CAT-13 — XMP interoperability, wired on 2026-09-12.** `core/dr-xmp` reads and writes
|
||||
standard XMP sidecars: `dc:subject` and `lr:hierarchicalSubject`, `xmp:Rating` and `xmp:Label`, and
|
||||
the IPTC core fields, in both the attribute and the element form and whatever RDF container a file
|
||||
happened to use. It states the ownership rule in one place — DarkRoom owns the properties in
|
||||
`PROPERTIES` and nothing else in the document, identified by namespace URI rather than by prefix —
|
||||
and enforces it by rewriting a packet event by event rather than serialising over it, so another
|
||||
application's `crs:` settings, comments and processing instructions survive a write byte for byte.
|
||||
|
||||
**The wiring is `ui/dr-ui/src/xmp_sync.rs`.** The scan collects `.xmp` beside `.drsc` from the
|
||||
listings it was already making; the pull reads each one whose ETag has moved and reconciles it
|
||||
against the catalog with the catalog winning — keywords union, and a rating, label or caption taken
|
||||
only where the catalog holds none. Both naming conventions resolve: `IMG_0001.CR3.xmp` names its
|
||||
file, `IMG_0001.xmp` the stem, and the JPEG beside a RAW is the same photograph. A genuine
|
||||
disagreement is written to `xmp_conflicts` and the settings page offers "Take the sidecars' values",
|
||||
which is the reload the requirement asks for; the detection it asks for is the ETag that moved. The
|
||||
write in the other direction is behind a setting that starts off (NFR-R4): a judgement or a keyword
|
||||
then also rewrites the sidecar beside the original, keeping the file's own caption, copyright and
|
||||
hierarchy, which the catalog has no columns for and would otherwise have deleted.
|
||||
|
||||
**What remains.** GPS is not carried — `exif:GPSLatitude` is a format of its own and `dr-decode`
|
||||
produces no location for it to carry yet. Title, description and copyright are read and reconciled
|
||||
but the catalog has nowhere to put them, so they pass through a rewrite rather than being editable.
|
||||
An XMP write made offline is not queued: the catalog and the `.drsc` are authoritative, and the next
|
||||
judgement online writes the file whole again.
|
||||
|
||||
`dr-preset-xmp` remains what it always was and is still not the counter-example it looks like: a
|
||||
reader of Lightroom *presets* under FR-DEV-6, a different file for a different purpose.
|
||||
|
||||
---
|
||||
|
||||
## 8. The performance targets are half-verified, and the half that is left is the hard one
|
||||
|
||||
§8 and §4.1 both require the same thing in the same words: an automated benchmark suite against a
|
||||
synthetic 50k catalog, run per commit, where **"a regression beyond a stated tolerance is a build
|
||||
failure, not a notification."** For most of this project's life it did not exist — no `benches/`, no
|
||||
criterion, no synthetic catalog, and three CI workflows that between them measured nothing.
|
||||
|
||||
**It exists now, for everything that does not need a frame.** [`tools/bench`](../../tools/bench) builds
|
||||
a deterministic 50,000-row catalog over a pool of a dozen real files, measures against it, and fails
|
||||
the build on a violated budget or a drift past tolerance;
|
||||
[`.gitea/workflows/benchmark.yml`](../../.gitea/workflows/benchmark.yml) runs it on every push, and
|
||||
[benchmarks.md](benchmarks.md) is the account of what it does and does not cover. **NFR-P1** and
|
||||
**NFR-P3** are now genuinely gated, and R2's "catalog opens in under 2s" clause with them.
|
||||
|
||||
Three qualifications, all of them stated in the harness itself rather than only here:
|
||||
|
||||
- **The numbers have not been recorded yet.** Every `recorded` field in
|
||||
[bench-baseline.json](bench-baseline.json) is `null`, deliberately: a fabricated baseline is worse
|
||||
than none. Until `dr-bench record --reference` is run on the reference desktop and committed, the
|
||||
budget gate works and the regression gate does not.
|
||||
- **NFR-P7 and NFR-P8 are half-measured and are not tagged.** The export row covers the encode half
|
||||
of the chain and no GPU render, so it can fail the requirement and cannot pass it. The memory row
|
||||
covers a process holding the catalog and nothing else — no toolkit, no adapter — so it is the
|
||||
catalog layer's share of the 500 MB rather than the figure NFR-P8 is about. Neither carries a
|
||||
`TRACES:` tag, which is the point.
|
||||
- **NFR-P8 needs a decision, not more code.** How much of its 500 MB belongs below the UI is
|
||||
unstated, and until somebody says, the metric can record but not judge. [benchmarks.md](benchmarks.md)
|
||||
also answers the question §4.1 raises about GPU memory — RSS cannot see device-local allocations
|
||||
at all — and recommends restating the requirement as two figures.
|
||||
|
||||
**What is left is the frame-timing half, and it is the hard one.** NFR-P2, -P4, -P5, -P6, -P9, -P10,
|
||||
-P11, -P12, -P13, -P14 and -P15 all need a probe inside a running Slint application, a GPU adapter,
|
||||
or both. `dr-gpu/examples/frame_budget` is a real instrument for the GPU part and its results are
|
||||
committed in [frame-budget.md](frame-budget.md) with the machine and profile named — but it is run
|
||||
by hand, and the guard version in CI skips itself where there is no adapter, which is the normal
|
||||
case on a runner. So the claim to take from this section is now narrower than it was, and still
|
||||
true: **a scroll that dropped to 30 fps tomorrow would reach a user before it reached CI.**
|
||||
|
||||
---
|
||||
|
||||
## 9. Two core requirements that cannot be closed as written
|
||||
|
||||
**R1 — Cross-platform output within a bounded tolerance.** §2 states that the threshold "must be
|
||||
fixed before spike S9", because S9 both validates R1 and calibrates what tolerance is achievable.
|
||||
The threshold was never fixed and S9 has not run, so R1 currently has no acceptance criterion at
|
||||
all — there is nothing a test could assert.
|
||||
|
||||
The matrix used to report R1 as *covered*, and what covered it was two string literals: fixtures
|
||||
inside the traceability tool's own unit tests, which the tool scans along with everything else,
|
||||
because a fixture demonstrating tag extraction was indistinguishable from a tag. The extractor now
|
||||
asks where the tag sits — a tag is the first word of a comment, not a string appearing anywhere on a
|
||||
line — and R1 is untagged again, which is the honest reading while it has no acceptance criterion to
|
||||
tag anything against.
|
||||
NFR-OPS-1 was covered by tags that were real rather than fixtures, which is the worse case of the
|
||||
two: one on `compute_coverage` and one on the gesture extractor, both on the traceability tool. A
|
||||
coverage calculation and a documentation generator are not diagnostics under any reading, so both
|
||||
tags were removed. It is the case [CONTRIBUTING.md](../../CONTRIBUTING.md) warns about in its own words:
|
||||
a tag proves a tag exists. The requirement has since been built where it says: the rotating,
|
||||
size-capped log and its redaction in `platform/dr-plat/src/diagnostics.rs` (2026-08-30), and the
|
||||
bundle in `diagnostics/bundle.rs` (2026-09-12) — the log, the crash records, the version, the schema
|
||||
and the GPU as one text file, shown in full in Settings before a second press writes it, and sent
|
||||
nowhere by either press.
|
||||
|
||||
**R2 — Efficient display of huge RAW libraries.** Its acceptance criterion contains "*(figure
|
||||
TBD)*" — the scroll velocity below which no cell may render as a placeholder — and asks for a stated
|
||||
prefetch margin and cache-hit rate. No figure is stated anywhere in the tree, neither quantity is
|
||||
measured, and [TD-2](technical-debt.md) and [TD-3](technical-debt.md) both describe the thumbnail
|
||||
path falling short of it in ways that were measured. R2 was deliberately left untagged on this branch
|
||||
for that reason: the machinery is substantial and the criterion is unmet and partly undefined.
|
||||
|
||||
Both belong with §8 above. A requirement whose threshold was never chosen and a target nothing
|
||||
measures fail in the same way — not by being wrong, but by being unfalsifiable.
|
||||
|
||||
---
|
||||
|
||||
## 10. Spikes
|
||||
|
||||
§9 defines fourteen validation spikes and says of three of them: "S1, S2 and S10 are the three that
|
||||
can invalidate the architecture."
|
||||
|
||||
Only **S1** (Slint + wgpu zero-copy on Linux) and **S14** (the face pipeline on a real library) have
|
||||
recorded results. S14's are the best evidence of any spike — a dedicated document, a measured pass
|
||||
over an 18,143-face library, a named device and a reproducible command — though D13's licensing half
|
||||
remains open.
|
||||
|
||||
**S6, S9, S10, S11 and S13 show no evidence of having run at all.** Each is referenced only from the
|
||||
requirement text that asks for it:
|
||||
|
||||
| Spike | Would settle | Blocked on |
|
||||
|---|---|---|
|
||||
| S6 | FR-DSP-2, NFR-RES-2 — tiling and images larger than GPU memory | Nothing; needs a device and a large image |
|
||||
| S9 | R1's tolerance threshold, and therefore R1 | Nothing; the threshold is defined *by* running it |
|
||||
| S10 | Whether SAF at 10k files meets NFR-P1/P3 | §5 — there is no SAF code to measure |
|
||||
| S11 | NFR-COMPAT-2, and whether Play makes SAF binding | Nothing |
|
||||
| S13 | NFR-A11Y-2 on Android | §6 — there is almost nothing to test with |
|
||||
|
||||
S2, S3, S4, S5, S7, S8 and S12 are also unrun, several with acknowledgements in the code that say
|
||||
so (`dr-sync/src/upload.rs` on S8, `dr-sync-nextcloud/src/lib.rs` on S3). S2 is one of the three
|
||||
architecture-invalidating spikes and needs Adreno and Mali hardware, which the manifest notes no
|
||||
emulator represents.
|
||||
|
||||
The pattern is worth stating rather than leaving to be inferred: the spikes that ran are the ones
|
||||
whose subject was being built anyway. The ones that did not are the ones that would have said
|
||||
whether something *should* be built — which is the opposite of the order §9 asks for.
|
||||
|
||||
---
|
||||
|
||||
## 11. Merging — specified 2026-09-19, nothing built
|
||||
|
||||
§3.11 was written on 2026-09-19 under D18, undeferring the panorama from §7 and leaving HDR merge
|
||||
and focus stacking there with their data model decided. Eleven `FR-MRG` clauses and two `NFR-MRG`
|
||||
figures entered the register at once with no code behind any of them, which is why the coverage
|
||||
figure fell from 83.0% to 77.2% on the same day — a specification, not a regression.
|
||||
|
||||
[panorama.md](panorama.md) is the design, and its §10 is the order of work. Nothing starts before
|
||||
**S15**: whether rawler reads back a linear DNG the application writes, whether XFeat loads under
|
||||
tract at a fixed shape, whether the working-space texture can be tapped where FR-MRG-2 needs it,
|
||||
and what a chunked blend of a 100 MP composite costs on the tablet. The first two are a day each
|
||||
and either can change the design, which is the reason they come first.
|
||||
|
||||
## 12. D12, which governed all of the above
|
||||
|
||||
> **Decided 2026-09-19, by events.** The scope stands as calibrated, v1 has no date, and `(post-v1)`
|
||||
> in §7 is the one way a clause leaves the count — used for the plugin API and nothing else. The
|
||||
> argument below is kept because it is what the decision weighed; its prediction about tablet
|
||||
> editing was right, and the cluster was built anyway.
|
||||
|
||||
[Decision D12 — scope versus pace](requirements.md) said, while it was open:
|
||||
|
||||
> The calibration selected an ambitious feature set — full tablet editing, full ingest, culling as a
|
||||
> differentiator, complete GPU masking, AI denoise, Fuji-first colour, deep sync, sidecar durability
|
||||
> — against a stated pace of evenings and weekends, indefinitely.
|
||||
>
|
||||
> **Those are not compatible as stated.**
|
||||
|
||||
Sections 1 through 10 are what that incompatibility looks like eleven versions later, and they land
|
||||
almost exactly where D12 predicted: tablet editing carries SAF at unproven scale, background
|
||||
execution limits and two GPU vendors to validate (§5), and every one of those is unbuilt or unrun.
|
||||
The parts that *were* built — the develop pipeline, sync, faces, the catalog — are the parts that
|
||||
did not need a decision first.
|
||||
|
||||
D12 was not resolved by choosing to work faster, and in the end not by moving clusters into §7
|
||||
either, except the one: plugins. Every other cluster above stays in scope, and each still says what
|
||||
it would take to build. That is the list. D3 is delivered, and
|
||||
[architecture.md §11](architecture.md)'s build order is what followed.
|
||||
@@ -0,0 +1,546 @@
|
||||
# Panorama
|
||||
|
||||
**Status:** Draft · 2026-09-19
|
||||
**Companion to:** [requirements.md](requirements.md) §3.11 FR-MRG-1 … 11, D18, S15 · [architecture.md](architecture.md) §5.2, §6.2
|
||||
|
||||
The first merge (§3.11): several frames, rotated about one point, become one
|
||||
photograph. This document is how that lands on the pipeline that exists now —
|
||||
which stages, where each runs, how the composite is produced in chunks when it
|
||||
is larger than any texture or any memory, what is ported from where, and what
|
||||
the keypoint model may be under D8.
|
||||
|
||||
---
|
||||
|
||||
## 1. Why it is worth the work
|
||||
|
||||
The audience shoots panoramas and leaves the application to stitch them. That
|
||||
is the same workflow break dust was (`spot-removal.md` §1): a RAW editor that
|
||||
does everything but the one thing, and the photographer's work ends up in a
|
||||
JPEG produced by a tool that never saw the RAW.
|
||||
|
||||
It is also the merge whose alignment problem is smallest. A panorama is a
|
||||
rotation — three parameters per frame plus a focal length — with no depth to
|
||||
recover. HDR merge and focus stacking share its data model (D18) and most of
|
||||
its machinery (FR-MRG-3, 5, 6, 7, 10, 11 are written to be general); building
|
||||
the panorama first builds the shared part on the easiest geometry.
|
||||
|
||||
## 2. Non-goals
|
||||
|
||||
- **Not structure-from-motion.** No translation is solved for. A hand-held set
|
||||
with parallax gets its ghosts hidden by seam placement, and a set with real
|
||||
parallax is not a panorama. COLMAP's front end is the right mental model;
|
||||
its back end is the wrong problem.
|
||||
- **Not a multi-source Version.** D18. The composite is a file, and nothing in
|
||||
the catalog, the sidecar format or sync learns about cross-references.
|
||||
- **Not boundary fill.** Painting pixels that were never captured is the pixel
|
||||
editing §1.3 excludes. Auto-crop is the tool.
|
||||
- **Not automatic.** The tool proposes an alignment and writes nothing until
|
||||
the photographer confirms. Same rule as spot removal and D17, for the same
|
||||
reason: a merge that silently omits or misplaces a frame is the failure this
|
||||
application must not have.
|
||||
- **Not HDR-panorama in one pass.** Until HDR merge exists on its own, a
|
||||
bracketed panorama is bracketed frames merged first, then stitched.
|
||||
|
||||
## 3. What is new, precisely
|
||||
|
||||
Nearly all of it, unlike spot removal. The pipeline renders one source to one
|
||||
texture; nothing in the tree detects keypoints, estimates a rotation, warps
|
||||
into a projection, finds a seam, or blends a pyramid. What exists and is
|
||||
reused:
|
||||
|
||||
| Exists | Where | Reused for |
|
||||
|---|---|---|
|
||||
| Render a source through the fused pass, with a linear f16 output mode | `dr-gpu` demosaic → `AdjustPass`, `OutputMode::LinearWorking` | FR-MRG-2's camera-space input, as a compose entry with no operations and the profile uniforms neutral (S15.3) |
|
||||
| Tiled rendering with a priority scheduler | ARCH §5.3 | Pulling source tiles on demand into an output chunk (§5 below) |
|
||||
| A non-CFA source entering the pipeline | `Demosaicer::from_rgba8` | The composite's decode path, if the container is a TIFF (S15.1) |
|
||||
| DNG matrices read through rawler | `dr-decode::profile` | The composite's decode path, if the container is a DNG |
|
||||
| Static-shape ONNX under tract, heads decoded in Rust | `dr-segment` | The keypoint model (§6) |
|
||||
| A batch worker with its own `GpuContext`, activity row, cancel | `dr-ui::export` | FR-MRG-7 verbatim |
|
||||
| The 16-bit TIFF encoder with metadata sub-IFDs | `dr-export::encode` | FR-MRG-3's writer, extended to linear samples |
|
||||
| Multi-select in the grid | `collections_ui::selected` | The entry point |
|
||||
|
||||
New: a `core/dr-pano` crate holding the geometry (keypoints, matching, the
|
||||
rotation solve), a set of WGSL passes in `dr-gpu` (reprojection, gain,
|
||||
seam, pyramid blend), the chunked output driver, the container writer, and
|
||||
the dialog.
|
||||
|
||||
## 4. The stages, and where each runs
|
||||
|
||||
FR-MRG-10 states the rule; this is the table it was written from.
|
||||
|
||||
| Stage | Cost shape | Runs on | Why |
|
||||
|---|---|---|---|
|
||||
| Source to camera-linear | per pixel, full res | GPU, the existing pipeline | It *is* the pipeline, stopped early |
|
||||
| Keypoint detection | once per frame, at 1024 px | CPU, tract (NEON on the tablet) | Bounded by frame count, not output size. Same runtime faces and masks use. Hand-written WGSL convolutions for a model that runs five times would be work with no visible gain. |
|
||||
| Descriptor matching | K² × D per pair | CPU, SIMD | 2048² × 64 × 10 pairs ≈ 3 GFLOP — tens of milliseconds |
|
||||
| Rotation solve, bundle adjustment | 3N + 1 parameters, Levenberg–Marquardt | CPU | Microseconds. Not parallel work. |
|
||||
| Preview reprojection | per pixel, proxy res | GPU, interactive | Projection and horizon changes re-warp N proxies at frame rate |
|
||||
| Full-resolution warp | per output pixel | GPU, chunked (§5) | The heaviest thing in the application |
|
||||
| Gain compensation | per overlap region | GPU reduction, then N scalars | Sums, on the histogram pass's pattern (ARCH §5.5) |
|
||||
| Seam finding | per overlap pixel | GPU-friendly variant | Graph cut resists the GPU; a distance-transform or per-column DP seam does not. The algorithm is chosen for the GPU, not for the paper. |
|
||||
| Multi-band blend | per pixel × levels | GPU, chunked | Laplacian pyramids are separable convolutions — the detail stage's shape |
|
||||
| Encode | per pixel, once | CPU, streamed per chunk row | As export does |
|
||||
|
||||
## 5. Chunked in output space
|
||||
|
||||
FR-MRG-11 forbids holding the composite as one texture, and two facts force it
|
||||
before memory does:
|
||||
|
||||
- `max_texture_dimension_2d` is 8192 on many mobile GPUs and 16384 on desktop.
|
||||
A three-row panorama is routinely 20 000 px wide.
|
||||
- Five 24 MP frames at working precision are ~1 GB together. The tablet does
|
||||
not have it.
|
||||
|
||||
**The geometry is known before any full-resolution pixel exists.** Alignment
|
||||
runs on proxies; what comes out is a rotation per frame, a focal length, a
|
||||
projection and an output rectangle. From those, every output pixel's source
|
||||
coordinates in every frame are a closed-form function. That is what makes
|
||||
chunking simple rather than clever:
|
||||
|
||||
```
|
||||
for each output chunk C (e.g. 2048 × 2048, in output space):
|
||||
frames_in(C) = frames whose projected footprint intersects C
|
||||
for each frame F in frames_in(C):
|
||||
source tiles T(F, C) = tiles of F that project into C, plus a margin
|
||||
render T(F, C) to scene-linear through the pipeline's tile cache
|
||||
warp T(F, C) into C's coordinate frame ← GPU
|
||||
gain-correct, seam, blend within C ← GPU, with overlap
|
||||
read C back, encode its rows ← CPU, streamed
|
||||
```
|
||||
|
||||
The working set is one chunk, its per-frame warped copies, and the source
|
||||
tiles that fed them. It does not grow with the composite.
|
||||
|
||||
**The blend needs a margin.** A Laplacian pyramid of L levels reads
|
||||
2^L pixels beyond the chunk edge; a chunk is therefore rendered with a margin
|
||||
of that width and the margin discarded after the blend. Seams cross chunk
|
||||
boundaries and must agree on both sides: the seam is found once at a reduced
|
||||
resolution over the whole overlap (which fits — it is a mask, not an image),
|
||||
then upsampled into each chunk. The same is true of gain: the scalars are
|
||||
solved once from proxy-resolution overlaps and applied everywhere.
|
||||
|
||||
**Source tiles are the pipeline's tiles.** ARCH §5.3's cache keys by
|
||||
`(VersionId, tile, zoom, graph_hash_prefix)`; the merge asks for tiles of a
|
||||
neutral graph at zoom 1 and gets the same caching every other consumer does.
|
||||
A tile pulled for one chunk is usually needed by the neighbouring chunk, and
|
||||
stays hot for it.
|
||||
|
||||
### 5.1 The tap — S15.3, answered by reading the composer
|
||||
|
||||
The fused shader's order, fixed by `operation.rs`'s own tests: warp → as-shot
|
||||
white balance → operations → base curve → camera matrix → store. The store is
|
||||
either the display encode or, in `OutputMode::LinearWorking`, an unclipped
|
||||
`rgba16float` of linear sRGB. That mode exists for the detail stage and is
|
||||
selected from the operations, never by a caller flag, so that a shader and
|
||||
the texture bound to it cannot disagree.
|
||||
|
||||
The merge wants the values *before* the curve and matrix (FR-MRG-2), and the
|
||||
composer already makes that a matter of uniforms rather than structure: the
|
||||
white balance, the matrix and the curve's active flag are all in the reserved
|
||||
uniform block, and a fused pass with no operations, `as_shot_wb = 1`,
|
||||
`cam_to_srgb = I` and `base_curve_last.z = 0` stores exactly camera-linear
|
||||
RGB after the warp. So the tap is:
|
||||
|
||||
- `EditGraph::compose_camera_linear()` — the `LinearWorking` tail with an
|
||||
empty operation list and identity framing, paired by name with
|
||||
- `AdjustPass::render_camera_linear()` — binds the f16 target, fills the
|
||||
reserved uniforms neutral instead of from the source, returns the texture,
|
||||
- and a float readback beside the existing 8-bit one.
|
||||
|
||||
Nothing in the chain moves. **Precision:** the tap and every chunk buffer
|
||||
after it should be `rgba32float`, not f16. A 14-bit sensor has 16 384 steps
|
||||
to white; f16 has 2 048 in the top octave, and a composite that is going to
|
||||
be re-developed deserves the sensor's precision. The cost is 2× on buffers
|
||||
FR-MRG-11 already bounds.
|
||||
|
||||
**What the DNG carries as a consequence:** the first source's `Make`,
|
||||
`Model` and `UniqueCameraModel` — so `base_curve::for_body` finds the 6D's
|
||||
curve — its `ColorMatrix1`/`2` with illuminants, and its `AsShotNeutral`. The
|
||||
composite then develops through the same profile as its sources, applied
|
||||
once. The spike's 64 × 48 file (§8) already carries the matrix and neutral;
|
||||
the body name is a string.
|
||||
|
||||
## 6. The keypoint model
|
||||
|
||||
FR-MRG-8: works without weights, better with them. The licence read comes
|
||||
first (D13's lesson, S15.2).
|
||||
|
||||
| Model | Licence | Fits tract? | Position |
|
||||
|---|---|---|---|
|
||||
| **XFeat** (CVPR 2024) | Apache-2.0 | Plain convolutions, fully convolutional, the repo ships an ONNX export | **Chosen.** Fixed 1024 px input, dense heatmap and descriptor map out, NMS and top-K in Rust — the yolo26 pattern |
|
||||
| DISK | Apache-2.0 | U-Net, static | Second choice; stronger descriptors, ~3–4× the compute |
|
||||
| ALIKE | BSD-3 | Plain convolutions | Fallback if XFeat's export fails F6 |
|
||||
| ALIKED | BSD-3 | Deformable convolution in the descriptor head | Unlikely to load |
|
||||
| SuperPoint, SuperGlue, R2D2, SiLK, MASt3R | non-commercial | — | Out on licence |
|
||||
| LightGlue | Apache-2.0 | Transformer over a variable keypoint count | Not until mutual-nearest-neighbour matching fails on a real set |
|
||||
|
||||
**S15.2, 2026-09-19: XFeat loads under tract.** `tools/export-xfeat.sh`
|
||||
exports the network alone at 768×1024 — thirteen operator types, all
|
||||
standard: `Conv`, `InstanceNormalization`, `AveragePool`, `Resize`, `Slice`,
|
||||
`Transpose`, `Reshape`, `Concat`, `Add`, `Relu`, `Sigmoid`, `ReduceMean`,
|
||||
`Unsqueeze` — and
|
||||
[`examples/onnx_probe.rs`](../../core/dr-segment/examples/onnx_probe.rs) loads
|
||||
the 2.8 MB file through the app's own `ort`-over-tract backend with nothing
|
||||
unsupported, in 28 ms, and runs it in **~300 ms on the reference desktop's
|
||||
CPU**. The weights ship as `models/keypoints/xfeat-1024.onnx`, recorded in
|
||||
`models/LICENCE.md`. **On the tablet** (S15.4's CPU half, same day):
|
||||
`tools/onnx-probe-on-device.sh` cross-builds the probe, and the same file
|
||||
runs in **~400 ms per frame** on the reference tablet's NEON cores (ROD2-W09,
|
||||
SM8635), with output ranges identical to the desktop's — inside NFR-MRG-1's
|
||||
1 s per frame with room to spare, and 1.3× the desktop rather than the 2×
|
||||
faces.md §9 measured for its scan. Still to do: a keypoint-level comparison
|
||||
against the PyTorch reference once the Rust decoder exists — the probe
|
||||
proves the graph runs, not that the numbers match.
|
||||
|
||||
The outputs are three maps at 1/8 resolution, 96×128 for the export size:
|
||||
64-channel descriptors, 65-channel keypoint logits (each 8×8 cell's position
|
||||
plus "none"), and a reliability heatmap. The Rust decoder is: softmax over the
|
||||
65, pixel-shuffle the first 64 to full resolution, 5×5 non-maximum
|
||||
suppression, top-k by reliability, bilinear sampling of the descriptor at
|
||||
each keypoint, L2 normalise. That is `detectAndCompute` in the reference,
|
||||
minus the network.
|
||||
|
||||
Without weights: AKAZE (BSD, `akaze` from rust-cv), which is adequate on
|
||||
well-textured overlaps and worse on sky, repeated structure and exposure
|
||||
drift — which is where a learned detector earns its place.
|
||||
|
||||
Matching is mutual nearest neighbour with a ratio test, then RANSAC on a
|
||||
rotation model. For a panorama — one lens, near-pure rotation, 20–40 %
|
||||
overlap — that is what Hugin and OpenCV's stitcher use, and it is enough.
|
||||
|
||||
## 7. What is ported from where
|
||||
|
||||
Nothing is linked; everything is read.
|
||||
|
||||
| Source | Licence | Taken |
|
||||
|---|---|---|
|
||||
| OpenCV `modules/stitching` | Apache-2.0 | The stage layout — Brown & Lowe (2007) as a set of small classes with one job each — and the warpers' projection maths |
|
||||
| OpenPano (ppwwyyxx) | MIT (verify on read) | The estimation and bundle-adjustment maths, function by function, with outputs diffed against it |
|
||||
| enblend-enfuse | GPLv2+ | Seam-line optimisation and Burt–Adelson multi-band blending |
|
||||
| Hugin `nona` | GPLv2+ | The GLSL remapper, as the reference for the WGSL warp |
|
||||
|
||||
The golden set (§8 of the requirements) is OpenCV's stitcher on the same
|
||||
inputs: a reference output to compare against, within a tolerance calibrated
|
||||
the way S9 calibrates R1.
|
||||
|
||||
## 8. The output file
|
||||
|
||||
FR-MRG-3. A linear DNG at the source's native scale: `u16` samples on the
|
||||
first source's black-subtracted scale, `WhiteLevel` = its white minus its
|
||||
black (13 023 for the 6D set: 15 070 − 2 047), `BlackLevel` = 0. Not rescaled
|
||||
to 65 535 — the sensor had 14 bits and the file says so, and a value the
|
||||
sensor could not have produced is not invented by a multiply. The first
|
||||
source's `Make`, `Model`, `UniqueCameraModel`, `ColorMatrix1/2`,
|
||||
`CalibrationIlluminant1/2`, `AsShotNeutral` and EXIF are carried, so the
|
||||
composite develops through the same profile as its sources. Named from the
|
||||
first source with a `-pano` suffix, beside it.
|
||||
|
||||
Three samples per pixel rather than a CFA: the warp resamples, and there is no
|
||||
sensor grid to mosaic back onto. Nothing else about being a RAW is lost —
|
||||
no white balance, no curve, no matrix, no clip has been applied — and the
|
||||
photographer develops the panorama afterwards as one photograph.
|
||||
|
||||
The sources are portrait frames in the 6D set: `Orientation` is applied
|
||||
before alignment (learned features are not rotation-invariant) and the
|
||||
composite is written upright with `Orientation = 1`.
|
||||
|
||||
Two containers were candidates and S15.1 decided, on 2026-09-19:
|
||||
|
||||
- **Linear DNG.** `PhotometricInterpretation = LinearRaw`, three samples per
|
||||
pixel, `ColorMatrix1` carried from the first source. Re-enters through
|
||||
rawler as `Format::Dng` with no new decode path, *if* rawler reads it back.
|
||||
What Lightroom writes.
|
||||
- **Float TIFF.** `SampleFormat = IEEEFP`, 16 or 32 bits, an ICC profile for
|
||||
the working space. Needs `Format::Tiff` and a decode path, but the writer is
|
||||
the existing encoder with a different sample type, and nothing about it is
|
||||
uncertain.
|
||||
|
||||
**Linear DNG.** [`examples/linear_dng.rs`](../../core/dr-decode/examples/linear_dng.rs)
|
||||
hand-rolls a 64 × 48 `LinearRaw` DNG — one IFD, uncompressed 16-bit RGB,
|
||||
`DNGVersion`, `ColorMatrix1`, `AsShotNeutral`, `CalibrationIlluminant1` — and
|
||||
rawler 0.7 reads it back: `cpp 3`, the samples interleaved as written, the
|
||||
matrix parsed into the camera definition, and `CameraProfile::extract` builds
|
||||
the same profile it would for a camera file. ImageMagick's libraw reads the
|
||||
same bytes. What does *not* yet work is `dr_decode::decode`, which accepts the
|
||||
file as CFA and hands the pipeline three times the samples it expects: the
|
||||
`cpp == 3` branch is the work, and it is the only decode work.
|
||||
|
||||
The composite therefore enters the pipeline as a non-CFA, *linear* source —
|
||||
`from_rgba8`'s sibling with `non_linear = false` and the colour matrix carried
|
||||
from the DNG — and is developed as any RAW is. The writer is the example's
|
||||
IFD, grown up: tiled rather than one strip (FR-MRG-11 encodes per chunk), and
|
||||
carrying the first source's EXIF in a sub-IFD as `dr-export` already does.
|
||||
|
||||
## 9. Interaction
|
||||
|
||||
- The entry is the grid's selection: two or more images, one action, "Merge
|
||||
to panorama". One image, or images from different roots, and the action
|
||||
says why it is unavailable.
|
||||
- The dialog shows the aligned proxies in the chosen projection, with the
|
||||
projection, horizon and crop controls of FR-MRG-4, and the per-frame
|
||||
residuals. A frame that failed to align is named there (FR-MRG-5), and the
|
||||
merge cannot be confirmed with it in the set.
|
||||
- Confirm starts the FR-MRG-7 job. The composite appears in the grid when the
|
||||
file is written and catalogued, beside its sources, with the merge as the
|
||||
first entry in its history.
|
||||
|
||||
## 10. Order of work
|
||||
|
||||
1. **S15**, all four, before anything else. (1) and (2) are a day each and
|
||||
either can change the design.
|
||||
2. `dr-pano`: keypoints (AKAZE first, XFeat when S15.2 passes), matching,
|
||||
RANSAC, rotation solve. Unit-tested against synthetic rotations of one
|
||||
frame, where the answer is known exactly.
|
||||
3. The working-space tap, and the preview reprojection pass. At this point the
|
||||
dialog can show an alignment.
|
||||
4. The chunked driver with a feathered blend — the whole path end to end,
|
||||
writing a file, before the blend is good.
|
||||
5. Gain, seams, multi-band.
|
||||
6. The container, the catalog entry, provenance, the history entry.
|
||||
7. Tablet: NFR-MRG-1's figure, and FR-MRG-9's ceiling.
|
||||
|
||||
## 11. Where it stands — 2026-09-19, end of the first day
|
||||
|
||||
Built, on branch `merge/panorama`, in the order §10 gave:
|
||||
|
||||
| Piece | Where | State |
|
||||
|---|---|---|
|
||||
| Geometry: keypoints, matching, homography, focal, bundle adjustment, projections | `core/dr-pano` | Done; 33 tests without a model; the fixture aligns in 4.5 s |
|
||||
| XFeat at two shapes under tract | `models/keypoints`, `dr_pano::xfeat` | Done; 300 ms/frame desktop, 400 ms tablet |
|
||||
| The camera-space tap | `OutputMode::CameraLinear`, `AdjustPass::render_camera_linear` | Done, `rgba32float`, tiles by view rect |
|
||||
| Linear DNG writer, streamed | `dr_export::write_linear_dng` | Done; rawler reads it back |
|
||||
| A three-sample `RawImage` re-entering the pipeline | `dr-decode`, `DemosaicedImage::from_linear_rgb16` | Done |
|
||||
| Warp, accumulate, resolve, chunk by chunk | `dr_gpu::MergePass`, `merge.wgsl` | Done; feathered blend, scalar gain |
|
||||
| The job: load, proxies, align, gains, confirm, merge, provenance | `dr_ui::merge` | Done; `examples/merge.rs` drives it headless |
|
||||
| The page: table, preview, projection, Merge/Stop/Back; the grid's button | `merge.slint`, `merge_ui.rs` | Done; `DARKROOM_START_MERGE=a.CR2,b.CR2` lands on it |
|
||||
| Placement beside the sources through the outbox, rescan | `merge_ui.rs` | Done, untested against a server |
|
||||
|
||||
**Measured on the fixture (desktop, 12 × 20 MP, Intel adapter):** proxies
|
||||
and keypoints 4 s, alignment 4.5–12.6 s (load-sensitive: the matcher is
|
||||
every core), the merge **26 s for a 22 993 × 5 980 composite** in twelve
|
||||
bands of 2048 × 512 chunks, 45 s all told, an 825 MB DNG. NFR-MRG-1's 60 s
|
||||
holds on the desktop with room; the tablet's figure is still S15.4's open
|
||||
half.
|
||||
|
||||
**Open, in the order they matter:**
|
||||
|
||||
1. **Auto-crop (FR-MRG-4).** The merge returns a coverage mask per band and
|
||||
the file carries the black border. The largest inscribed rectangle over
|
||||
the coverage, then the DNG's `DefaultCropOrigin`/`DefaultCropSize`, so
|
||||
nothing is thrown away and the develop view opens on the picture.
|
||||
2. **Seams and the pyramid** (§10 step 5). The feather hides exposure and
|
||||
small misalignment; parallax on the near slope will show as a soft
|
||||
double edge at 1:1.
|
||||
3. **Vignetting in the tap.** The lens profile's distortion is applied
|
||||
before the fetch; its vignetting is an operation and is not. Frame edges
|
||||
are darker than their centres by the lens's falloff, and the feather
|
||||
averages them into the overlaps.
|
||||
4. **The tablet:** memory (twelve 40 MB sensor buffers on the CPU, one
|
||||
demosaiced frame at a time on the GPU), the figure, and FR-MRG-9's
|
||||
ceiling.
|
||||
5. **Horizon and drag-to-correct (FR-MRG-4, the proposed 4a).** The
|
||||
alignment failed on nothing in the fixture; the interaction waits for a
|
||||
set it fails on.
|
||||
6. **`derived_from` names sources by file name**, not content hash: the
|
||||
catalog's `content_hash` is null for most images most of the time. The
|
||||
hash can join it when the catalog has one.
|
||||
|
||||
## 12. Filling the border instead of cropping it — MI-GAN, read and measured 2026-09-19
|
||||
|
||||
Raised after the first merges: the ragged border a cylinder leaves could be
|
||||
*filled* rather than cropped away. FR-MRG-4 says no boundary fill, on
|
||||
§1.3's "not a pixel editor"; this is the evidence for deciding whether to
|
||||
revise that, not a revision.
|
||||
|
||||
**The candidate: MI-GAN** (Sargsyan et al., ICCV 2023, Picsart AI Research).
|
||||
Image inpainting designed for mobile: ~6 M parameters, plain convolutions —
|
||||
no FFT, no attention — so it quantises to int8 and runs on a phone's DSP,
|
||||
with quality close to LaMa and CoModGAN.
|
||||
|
||||
**Licence: MIT, code and weights alike** (`LICENSE` and `LICENSE-WEIGHTS`
|
||||
in the repository, read the same day). The cleanest position of any model
|
||||
in the tree — GPL-compatible, store-compatible, no grant to read around.
|
||||
|
||||
**Export.** The HuggingFace ONNX files are the *pipeline* — uint8 image and
|
||||
mask in, crop-around-mask, resize and blend inside the graph, every
|
||||
dimension dynamic — and tract refuses them (F6 again). The bare generator
|
||||
exports cleanly from the `migan_512_places2.pt` state dict at a fixed
|
||||
`1×4×512×512` (`export_migan.py` in the spike directory; the input is
|
||||
`mask − 0.5` and the masked RGB in −1..1, the output RGB in −1..1, the
|
||||
caller composites). After slimming the graph is **six operator types**:
|
||||
`Add, Clip, Conv, LeakyRelu, Mul, Resize`. 28 MB.
|
||||
|
||||
**Under tract on the reference desktop: loads in 53 ms, runs in 7.4 s per
|
||||
512 × 512 tile, f32.** That is the number. The fixture's border is two
|
||||
ragged bands across 22 993 px — roughly ninety 512-px tiles at full
|
||||
resolution — so a CPU-f32 fill is ten minutes on the desktop and longer on
|
||||
the tablet. Three ways to make it viable, none built:
|
||||
|
||||
1. **Fill at a quarter of the resolution and upsample.** Sky and scree
|
||||
tolerate it; twenty-odd tiles, about three minutes on the desktop CPU. A
|
||||
background job with the outbox's patience, not an interactive one.
|
||||
2. **int8 on the tablet's Hexagon through QNN**, where the plain-conv design
|
||||
is the point and the whole graph should run in milliseconds. The setup
|
||||
exists from the eye-state work; MI-GAN is a candidate for the same path.
|
||||
3. **A WGSL runtime for those six operators.** A project of its own, and
|
||||
the only route that would make it interactive on the desktop.
|
||||
|
||||
Whichever, the fill is a *proposal* under FR-MRG-1's rule — shown, then
|
||||
confirmed — and it would sit beside the crop, not replace it: the crop is
|
||||
free and honest, the fill is invented pixels, and the photographer chooses.
|
||||
|
||||
## 13. The fill, built — 2026-09-19, evening
|
||||
|
||||
Built the same day on `merge/fill`, on the engine (S16) rather than tract,
|
||||
and FR-MRG-4 revised to admit it: the border is *cropped or filled*, the
|
||||
photographer's choice, the crop the default.
|
||||
|
||||
**What runs.** `dr_pano::fill` is the engine-independent half: an
|
||||
`Inpainter` trait (a 512-px tile in, the same tile out) and `fill_border`,
|
||||
which owns everything the model does not — which tiles, what context, how
|
||||
to blend. `dr_pano::migan::MiGan` is the trait over the shipped generator
|
||||
under `dr_inference_engine` with the new `Role::Inpainter`, so it takes
|
||||
whichever rung the device has. The merge job runs the fill at **half the
|
||||
composite's resolution**, in a display-ish space (white balance, camera
|
||||
matrix, gamma — invertible, so the result goes back to camera-linear and
|
||||
into the same linear DNG), and the full-resolution merge samples the fill
|
||||
where no frame reached.
|
||||
|
||||
**What the spike taught, tried in order and kept or dropped.**
|
||||
|
||||
1. *Context across the coverage edge.* MI-GAN was trained on holes inside
|
||||
pictures; given a hole at the picture's edge it invents a structure along
|
||||
the open side (white streaks in the sky, on the first try). The known
|
||||
content is therefore **mirrored** across the coverage edge into the hole
|
||||
and into a 256-px ring, column by column for the top and bottom bands
|
||||
and row by row for the sides; the model interpolates between real and
|
||||
mirrored sky rather than extrapolating into nothing. *Replicated* rows
|
||||
(the edge row continued flat) streaked the grass; a detrended mix (tone
|
||||
replicated, texture mirrored) smeared; a low-pass extrapolation banded.
|
||||
Mirror stays.
|
||||
2. *Coarse to fine.* One pass at the working resolution let the boundary
|
||||
leak in — each 512 tile saw only its own corner of the hole. So a
|
||||
**coarse pass at a quarter** decides the structure with the whole border
|
||||
in a few tiles, and **fine passes in 96-px bands** from the real edge
|
||||
outward regenerate texture, each band the only unknown with the previous
|
||||
band on its near side and the upsampled coarse fill on its far side.
|
||||
3. *The seam.* A hard cut between real and invented showed as a sharpness
|
||||
step. The known mask is eroded by a **24-px feather** (48 at half
|
||||
resolution) and the fill blended in across that margin by distance to
|
||||
the real edge, smoothstep.
|
||||
4. *Partial pixels.* The seams were still visible until the cause was found
|
||||
upstream of the fill: the camera-space tap stored **black with alpha 1**
|
||||
for a pixel the lens correction pushed off the sensor, and the warp
|
||||
averaged it in — a dark, poorly interpolated fringe along every frame's
|
||||
edge that the fill then continued. `OutputMode::CameraLinear` now
|
||||
stores alpha 0 for a pixel that is not there and the merge's warp
|
||||
weights by the sampled alpha, so the fringe never enters the composite.
|
||||
The mask erosion before the fill dropped from 16 px to 4.
|
||||
|
||||
5. *What is still wrong, and why it ships anyway.* With the seams gone the
|
||||
content itself is the problem in the deep corners: the model, trained
|
||||
on Places2, puts bright cloud-and-peak shapes into a sky hole and a
|
||||
water-like band under grass — its prior for "top of a picture" and
|
||||
"bottom of a landscape", not anything in the context (the same shapes
|
||||
appear with the mirror capped, uncapped, and on the CPU as on TensorRT).
|
||||
Thin borders are fine; that is most of a hand-held sweep. So the fill
|
||||
ships **experimental**: opt-in, previewed, its knobs on the page and
|
||||
in the sidecar, and `cargo run -p dr-ui --example fill` re-runs any
|
||||
merge's dumped input (`DR_FILL_DUMP=dir`) stage by stage in seconds so
|
||||
the next attempt is made from the picture, not from a seven-minute
|
||||
merge. Candidates for that attempt: a context that is not a mirror at
|
||||
all in deep holes (the coarse pass's own answer, iterated), a sky
|
||||
detector that fills sky by extrapolating the gradient and leaves the
|
||||
model to texture, or a different model.
|
||||
|
||||
**Measured, the fixture's twelve frames (22 991 × 5 978), 348 tiles at
|
||||
half resolution.** 312 s on ONNX Runtime's CPU pool on the reference
|
||||
desktop (≈ 0.8 s a tile). On TensorRT fp16: **100 s**, of which 60 ms a
|
||||
tile was the engine hashing the 28 MB model on every acquire (fixed, the
|
||||
hash is taken at open) and 150 ms a tile the GPU itself — throttled:
|
||||
`trtexec` on the same engine read 23 ms at noon on a cool machine and
|
||||
152 ms that evening after two hours of builds, nvidia-smi showing SW power
|
||||
cap and thermal slowdown. Cool, the fill is ~10 s. The TensorRT engine
|
||||
compiles once, in 13 minutes, cached under the inference directory.
|
||||
|
||||
**The runtime is a packaging matter.** Arch's `onnxruntime-opt-cuda` has
|
||||
no TensorRT provider ("not enabled in this build") and its CUDA provider
|
||||
does not load against cuDNN 9, so on this machine the app fell to ORT CPU
|
||||
until the official `onnxruntime-linux-x64-gpu_cuda13` tarball (1.30.0,
|
||||
which links the system CUDA 13.4 and TensorRT 10.16) was unpacked and
|
||||
named with `DARKROOM_ORT_DIR`; `/usr/lib/darkroom` is searched too, for a
|
||||
package that ships it. §12's point 2 for the tablet is unchanged.
|
||||
|
||||
**On the page.** A *Border* choice beside the projection — *Crop to the
|
||||
picture* / *Fill the border* — with a caption saying what the fill is; the
|
||||
preview re-renders filled when chosen, at preview resolution, so the
|
||||
choice is seen before it is confirmed (FR-MRG-1). Greyed out with the reason
|
||||
when `migan-512.onnx` is not in the model directory. Under the fill, while
|
||||
it is experimental, its six knobs as sliders — working scale, edge
|
||||
erosion, coarse pass, band width, mirror depth, seam feather — each
|
||||
committing a redraw of the preview. A filled merge's sidecar says `border
|
||||
filled` with the knobs used, and its default crop is still the inscribed
|
||||
rectangle.
|
||||
|
||||
**Ships.** `models/inpaint/migan-512.onnx` (LFS, 28 MB, MIT,
|
||||
`models/LICENCE.md`), installed by the PKGBUILD and unpacked by the APK
|
||||
beside the face and scene models; `tools/export-migan.sh` regenerates it
|
||||
from the upstream checkpoint.
|
||||
|
||||
## 14. The fill, trained — 2026-09-20
|
||||
|
||||
§13.5 named the remaining fault: in a deep corner the stock model puts its
|
||||
Places2 prior — clouds, peaks, a water line — into a hole, because it was
|
||||
trained on holes *inside* pictures and a panorama's border is a hole with
|
||||
the picture on one side and nothing on the other. Every ring (mirror,
|
||||
replicate, detrend) treated the symptom. The fix is a model that has seen
|
||||
the real thing: **MI-GAN's 512 generator fine-tuned on border-shaped voids
|
||||
cut from the user's own photographs**, in a separate repository
|
||||
(`darkroom-infill`, beside this one), so the truth beyond the void is known
|
||||
and the model learns one-sided extrapolation.
|
||||
|
||||
**What it was trained on.** Voids made the way this merge makes them:
|
||||
frames with a yaw, a common pitch and per-frame roll, projected onto the
|
||||
cylinder and rasterised, the canvas their union's bounding box, the void
|
||||
the canvas outside the union — arcs where straight edges bent, cusps where
|
||||
frames meet, the bow-tie wedge at a corner (a third of tiles are cut at a
|
||||
canvas corner). Voids to 256 px deep at a 512 tile. Half the deep tiles
|
||||
train the *second pass*: a no-grad first pass fills the tile, its nearest
|
||||
band (64–256 px) is marked known, and the remaining void is the example —
|
||||
so the model continues its own output without drift, which is how `fill`
|
||||
runs deep voids. Data: ~7 400 pictures — the 1024-px proxy tier of the
|
||||
library and ~2 000 raws sampled evenly across every year, developed at
|
||||
half size. Loss: hole-weighted L1, VGG16 perceptual, a hinge PatchGAN.
|
||||
One night on the reference desktop's RTX 3050.
|
||||
|
||||
**What changed here.** `FillParams::mirror_depth` **0** is now "no ring":
|
||||
the void reaches the tile's edge with nothing beyond, and beyond the band
|
||||
being filled the void stays *unknown* rather than presented as known coarse
|
||||
fill — the two conditions the model was trained under. Defaults: mirror 0,
|
||||
coarse 1 (the coarse pass seeds nothing the model is allowed to see), band
|
||||
192. The ring remains on the page for the stock model's sake, at any depth
|
||||
above zero. The model file is a drop-in (`models/inpaint/migan-512.onnx`,
|
||||
same six operators, same tensors) and the engine loads it unchanged.
|
||||
|
||||
**Measured, 240 held-out tiles with projection-shaped voids (PSNR in the
|
||||
hole, dB / LPIPS on the composite), stock → shipped (step 3 607):** edge
|
||||
16.9 → 18.5 / 0.121 → 0.136; corner 14.6 → 16.3 / 0.183 → 0.205; interior
|
||||
18.2 → 19.3 / 0.051 → 0.056. Read both columns: the fine-tune gains ~2 dB
|
||||
on edges and corners because it stops inventing objects, and *loses* on
|
||||
LPIPS because what it paints in a deep void is smoother than the stock
|
||||
model's confident wrong texture — LPIPS rewards texture, right or wrong.
|
||||
On the fixture's dump at half resolution (the merge's working size) the
|
||||
sky corners are sky, with no structure and a faint tone step at worst;
|
||||
the ground bands carry a fine texture at the right tone, softer than the
|
||||
real scree above them. The stock model's top-left corner on the same
|
||||
dump is a glowing invented structure. The training's own record — what
|
||||
each loss weighting did, and the two runs abandoned (blur under L1 in
|
||||
the hole; a brick pattern under a strong adversarial term against a
|
||||
discriminator that had not learned) — is `runs/` in `darkroom-infill`.
|
||||
|
||||
**What is still wrong.** The ground fill is softer than its context —
|
||||
texture, not structure, is what a night on a laptop GPU could not finish.
|
||||
The levers, in order: a discriminator that learns (a pretrained one —
|
||||
MI-GAN's own from the unfused checkpoint — instead of a PatchGAN from
|
||||
scratch), feature matching, and more steps at 512. FR-MRG-4's
|
||||
*experimental* stays.
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,503 @@
|
||||
# Region segmentation for local masking
|
||||
|
||||
Spec for **S15**, the spike that decides how DarkRoom finds the boundaries a local mask snaps to.
|
||||
|
||||
Local adjustments (FR-DEV-3, "linear gradient, radial gradient, and brush masks") need more than
|
||||
placement handles to be competitive. The interactions that matter — click to select a region, drag a
|
||||
contour that clings to an edge, paint without crossing a boundary — all need the same thing
|
||||
underneath: **a map of where the image's regions are.**
|
||||
|
||||
There are two credible ways to produce that map and they are not obviously ordered. This document
|
||||
specifies both, specifies the third option of combining them, and fixes the measurements that decide
|
||||
between them *before* any of them is built.
|
||||
|
||||
---
|
||||
|
||||
## 1. Why this is a spike and not a build
|
||||
|
||||
Three properties make the choice expensive to get wrong.
|
||||
|
||||
**It sets the mask representation.** If regions exist, a mask is a *set of region ids* — integers,
|
||||
diffable, mergeable at node level under FR-NC-9, cheap in a sidecar. If they don't, a mask is a
|
||||
raster, and rasters are none of those things. This is the decision that is expensive to retrofit;
|
||||
everything else in local masking sits on top of it.
|
||||
|
||||
**One arm collides with a settled policy.** D13 records that every dependency choice in this project
|
||||
has gone the same way — rustls over aws-lc-rs, bundled SQLite, a Rust Lensfun port, zune-jpeg — to
|
||||
avoid a C dependency under the Android NDK, and names `ort` as the largest exception that policy
|
||||
would tolerate. Arm B needs exactly that exception. Arm A needs no new dependency at all. That
|
||||
asymmetry is not a tiebreak, it is most of the cost difference, and it should be priced honestly
|
||||
rather than discovered at packaging time.
|
||||
|
||||
**Model licensing is a distribution blocker.** D13 already establishes this for the face pipeline,
|
||||
and the same reading applies here — see §7. It is a licence-reading exercise, not a research
|
||||
question, and it comes first.
|
||||
|
||||
---
|
||||
|
||||
## 2. The common interface
|
||||
|
||||
Both arms produce the same thing. This is what makes them comparable, and what lets the choice be
|
||||
deferred behind a seam rather than baked into every consumer.
|
||||
|
||||
```rust
|
||||
/// A partition of the image into labelled regions.
|
||||
pub struct RegionField {
|
||||
/// Per-pixel region id at proxy resolution. R32Uint on the GPU.
|
||||
labels: Texture,
|
||||
/// Per-region summary: pixel count, bounding box, mean colour,
|
||||
/// and (arm B only) a semantic class id.
|
||||
regions: Vec<Region>,
|
||||
/// Boundary strength per adjacent region pair — the edge weight
|
||||
/// the merge tree is built from and the cost field reads.
|
||||
adjacency: Vec<(RegionId, RegionId, f32)>,
|
||||
}
|
||||
```
|
||||
|
||||
Two consumers sit on it, and neither knows which arm produced it:
|
||||
|
||||
**A cost field, for contour snapping.** Live-wire — Dijkstra from the last anchor to the cursor over
|
||||
a per-pixel cost that is *low* on boundaries. The cost is a **sum of terms**, which is the property
|
||||
that matters: image gradient is always available, region boundary strength is added when a
|
||||
`RegionField` exists, and a semantic boundary term is added when a model is present. Each source
|
||||
improves the snap without changing the interface, so the arms are not exclusive here even in
|
||||
principle.
|
||||
|
||||
**A region set, for click selection.** Click reads the label under the cursor; the mask is
|
||||
`label(px) ∈ selected`. Add and subtract are set operations on ids. No flood fill, no readback, no
|
||||
iteration — the whole reason the precomputed map is worth having.
|
||||
|
||||
Both consumers are built once, in the spike, and shared by both arms. A comparison in which each arm
|
||||
gets its own consumer measures the consumers.
|
||||
|
||||
---
|
||||
|
||||
## 3. Arm A — multiscale watershed
|
||||
|
||||
No model, no new dependency, deterministic, works on any image.
|
||||
|
||||
**Gradient.** Sobel magnitude over a perceptual luma plus chroma distance, not camera-space RGB —
|
||||
channel-weighted RGB gradient reads a saturated red edge as weaker than it looks. Computed after
|
||||
demosaic and denoise, before the edit graph, so an exposure change does not invalidate it.
|
||||
|
||||
**Pre-smoothing is not optional.** Raw watershed on a noisy file makes every grain its own basin.
|
||||
A guided or bilateral pre-filter, with strength tied to the file's ISO, is part of the arm rather
|
||||
than a refinement of it.
|
||||
|
||||
**Basins.** Each pixel points downhill to its steepest neighbour; pointer-jumping resolves every
|
||||
pixel to its basin root in log passes. Two compute shaders and a dispatch loop.
|
||||
|
||||
**The hierarchy is the cheap part.** Build the region adjacency graph, sort edges by boundary
|
||||
strength, union-find over them, and *record the merge order*. That recording is the merge tree — a
|
||||
click selects a leaf, and a scroll walks up through progressively coarser merges. Textbook Kruskal on
|
||||
a graph of a few thousand nodes.
|
||||
|
||||
That node count is why the tree build is a legitimate CPU operation: it runs on the adjacency graph,
|
||||
not on pixels. Pixels stay on the GPU, the graph is CPU-side — the same split ARCH §3.4 and §6.1
|
||||
already draw for the edit graph, so no exception to the no-readback rule is needed.
|
||||
|
||||
**Known weaknesses, to be measured rather than argued about:** over-segmentation on noise and
|
||||
texture, weak boundaries where contrast is low but semantics are obvious (a pale sky meeting a pale
|
||||
wall), and a granularity ladder that is geometric rather than meaningful — level 7 is *a* coarser
|
||||
partition, not necessarily *the* object.
|
||||
|
||||
---
|
||||
|
||||
## 4. Arm B — semantic segmentation
|
||||
|
||||
YOLO26-seg pretrained on ADE20K, run through `ort`.
|
||||
|
||||
ADE20K's 150 classes include stuff — sky, vegetation, water, wall, road — which is a far better
|
||||
vocabulary for photography than COCO's 80 thing-classes. "That patch of sky" is a class here. The
|
||||
nano variant is ~1.6M parameters, which is genuinely mobile-viable in a way SAM never was.
|
||||
|
||||
**It produces a flat partition with class ids**, so it populates `RegionField` directly: connected
|
||||
components of the class map become regions, class boundaries become adjacency edges.
|
||||
|
||||
**What it does not produce is a hierarchy.** One partition at one semantic granularity. Click "sky"
|
||||
and you get all the sky; there is no level at which you get *this part* of the sky. Adjacent
|
||||
same-class regions merge whether or not you wanted them to — two different walls are one wall.
|
||||
|
||||
**Boundaries are semantically right and geometrically soft.** Internal stride is 4–8, upsampled to
|
||||
H×W, so the class map is confident about *which* side of the boundary a pixel is on and vague about
|
||||
*where* the boundary is to the pixel. Acceptable for biasing a contour. Not acceptable as a mask
|
||||
edge at 100% zoom.
|
||||
|
||||
**Costs it brings that arm A does not:** a C dependency on the Android NDK against D13's policy, a
|
||||
model to distribute and cache, an AGPL question (§7), an inference runtime per platform, and output
|
||||
whose determinism across drivers is unproven (§6).
|
||||
|
||||
---
|
||||
|
||||
## 5. Arm C — semantic as a merge prior
|
||||
|
||||
The arms are not alternatives, and a comparison that omits their combination is a false dichotomy.
|
||||
|
||||
Weight each region-adjacency edge in arm A's union-find by boundary strength **and** by whether the
|
||||
two regions share a semantic class. Regions that agree semantically merge earlier.
|
||||
|
||||
The result is a hierarchy whose coarse levels align with semantic objects and whose fine levels stay
|
||||
pixel-accurate — the model doing what models are good at, which is knowing what things *are*, and
|
||||
watershed doing what it is good at, which is knowing where boundaries are, exactly, at every scale.
|
||||
It also repairs arm B's two weaknesses at once: the soft boundary is replaced by the watershed
|
||||
boundary underneath it, and the missing granularity ladder is arm A's.
|
||||
|
||||
Arm C is the expected winner on quality. The question the spike actually has to answer is therefore
|
||||
not "which is better" but **how much better than arm A alone, and is that increment worth D13's
|
||||
cost.** §8 fixes that threshold in advance.
|
||||
|
||||
---
|
||||
|
||||
## 6. What gets measured
|
||||
|
||||
Per arm, over the corpus in §9, using the shared consumers from §2.
|
||||
|
||||
| # | Measure | Method | Why it decides anything |
|
||||
|---|---|---|---|
|
||||
| **M1** | **Interactions to target mask** | Clicks plus scroll steps to reach ≥95% IoU against a hand-traced mask | The real UX metric. "How many actions to get the mask I meant" is what a user experiences |
|
||||
| **M2** | **Boundary accuracy** | Precision/recall of snapped-contour pixels within a 2px slack of the hand trace | Whether the edge survives 100% zoom, where masks are actually judged |
|
||||
| **M3** | **Granularity coverage** | Per case, yes/no: does *any* hierarchy level produce the target region? | A hard failure mode. Arm B is expected to fail this wherever the target is not a class |
|
||||
| **M4** | **Out-of-vocabulary behaviour** | M1 and M3 restricted to the OOV subset | Whether the arm degrades gracefully or produces nothing usable off-distribution |
|
||||
| **M5** | **Determinism** | Same input twice on one machine; then across Mesa/AMD, NVIDIA, and Adreno | Gates whether a label field can be a cache key at all — see below |
|
||||
| **M6** | **Precompute cost** | ms at proxy resolution and peak memory, on the reference desktop and one Android device | Whether it fits a background prefetch alongside the proxy |
|
||||
| **M7** | **Distribution cost** | Added binary size, model size, new native dependencies, licence | D13's axis. Priced, not assumed |
|
||||
|
||||
**M5 deserves its own note, and it is a risk for arm A too.** ARCH §6.13 holds that cache keys are
|
||||
computed over CPU-side *integer* state because GPU float results diverge across vendors. A label
|
||||
field is integer state — but it is *derived from* float gradient arithmetic, so a boundary sitting
|
||||
exactly between two basins could resolve differently on Adreno than on Mesa. If either arm proves
|
||||
non-deterministic across vendors, its output cannot be a cache key and cannot round-trip through a
|
||||
sidecar as region ids, which would push masks back toward rasters and undo most of §1's argument.
|
||||
This is the measurement most likely to invalidate the whole approach, and it should be run early
|
||||
rather than last.
|
||||
|
||||
---
|
||||
|
||||
## 7. Licence reading — before any code
|
||||
|
||||
D13's position applies unchanged: discovering at packaging time that a feature cannot ship is the
|
||||
expensive failure, and it is entirely avoidable.
|
||||
|
||||
**Ultralytics ships YOLO under AGPL-3.0** — confirmed 2026-08-17 by reading `LICENSE` at the head of
|
||||
`github.com/ultralytics/ultralytics`, which is the GNU Affero General Public License v3 verbatim.
|
||||
That is deliberate on their part; the commercial licence is their business model.
|
||||
|
||||
GPLv3 §13 explicitly permits the combination, so this is *not* the blocker the InsightFace
|
||||
non-commercial weights were: it is redistributable. But the combined work becomes effectively AGPL,
|
||||
which is a change to DarkRoom's licensing posture rather than a dependency detail, and it needs to be
|
||||
a decision made on purpose.
|
||||
|
||||
Still to verify before writing any of arm B:
|
||||
|
||||
- The licence on YOLO26 specifically, and on the ADE20K-pretrained weights *separately* from the
|
||||
framework code — they are not necessarily the same grant.
|
||||
- Whether ADE20K's own terms permit redistribution of weights derived from it.
|
||||
- Whether AGPL is acceptable for DarkRoom, given Flatpak, F-Droid and Play distribution
|
||||
(NFR-COMPAT-2).
|
||||
|
||||
**Arm A raises none of these questions**, which is worth stating plainly as part of its cost.
|
||||
|
||||
---
|
||||
|
||||
## 8. Decision criteria, fixed in advance
|
||||
|
||||
Stated now so the result cannot be rationalised afterwards.
|
||||
|
||||
- **Arm A ships alone** if it reaches within **one interaction** (M1) of arm C on the scene subset
|
||||
*and* dominates arm C on the OOV subset (M4). The semantic increment does not then justify a C
|
||||
dependency, an AGPL conversion, and a per-platform inference runtime.
|
||||
- **Arm C ships** if it beats arm A by **two or more interactions** on the scene subset without
|
||||
regressing OOV. That is a large enough difference to be felt in ordinary use, and it is what would
|
||||
justify reopening D13.
|
||||
- **Arm B never ships alone.** M3 is expected to fail on anything that is not an ADE20K class, and an
|
||||
arm with a hard failure mode and no fallback is not a selection tool. If it surprises us and passes
|
||||
M3 broadly, that is a genuine finding and this criterion is revisited on the evidence.
|
||||
- **If M5 fails for an arm across vendors**, that arm cannot carry region ids into the sidecar
|
||||
regardless of how it scored elsewhere.
|
||||
|
||||
---
|
||||
|
||||
## 9. Corpus
|
||||
|
||||
Roughly 24 images from a real library — three per category — hand-traced once and reused across all
|
||||
arms. Categories chosen for the failure modes they provoke, not for coverage:
|
||||
|
||||
| Category | Provokes |
|
||||
|---|---|
|
||||
| Gradient sky | Low-contrast boundary; watershed banding |
|
||||
| Foliage against sky | High-frequency boundary — arm A over-segments, arm B blurs |
|
||||
| Hair against a busy background | The classic hard mask edge |
|
||||
| Out-of-focus background | No edges at all; tests graceful failure in both |
|
||||
| High-ISO noise | Arm A's known weakness; tests whether pre-smoothing is sufficient |
|
||||
| Backlit silhouette | Strong unambiguous edge — the control case |
|
||||
| Macro, abstract, still life | **OOV for ADE20K.** Arm B expected to fail M3 here |
|
||||
| Architectural detail | Repeated structure; arm B merges distinct walls into one class |
|
||||
|
||||
Hand-tracing 24 masks is a couple of hours and it is what makes M1 and M2 mean anything. Without
|
||||
ground truth this comparison is two demos and a preference.
|
||||
|
||||
---
|
||||
|
||||
## 10. Deliverables
|
||||
|
||||
Nothing in the UI, nothing in the graph, nothing in the sidecar.
|
||||
|
||||
- `core/dr-gpu/src/shaders/watershed.wgsl` — gradient and basin propagation.
|
||||
- `core/dr-gpu/src/segment.rs` — the passes, producing a `RegionField`.
|
||||
- RAG construction and the union-find merge tree as a pure-CPU module with unit tests and no device,
|
||||
so the hierarchy is testable headless the way `dr-pipeline` is (ARCH §6.5a).
|
||||
- The two shared consumers from §2 — live-wire over a summable cost field, and region-set selection.
|
||||
- `core/dr-gpu/examples/segment.rs` — false-coloured PNGs at four or five hierarchy levels, plus the
|
||||
M1/M2 numbers against the traced corpus.
|
||||
|
||||
It graduates to `core/dr-segment` if it ships; that is not a spike decision.
|
||||
|
||||
---
|
||||
|
||||
## 11. Order
|
||||
|
||||
1. **Licence reading (§7).** Hours, and it can eliminate arm B before anything is built.
|
||||
2. **Arm A, and the shared consumers.** About a day. Look at the false-coloured PNGs — if the
|
||||
granularity ladder does not feel right, nothing downstream matters and that is worth knowing
|
||||
immediately.
|
||||
3. **M5 across vendors, early.** It is the measurement that can invalidate the region-id
|
||||
representation entirely, and it wants knowing before the corpus work is invested.
|
||||
4. **The traced corpus, then M1–M4 on arm A.** Establishes the baseline every other arm is judged
|
||||
against.
|
||||
5. **Arms B and C**, only if §7 cleared and arm A's baseline leaves room worth closing.
|
||||
|
||||
Arm A is a day and needs no model, no runtime, no licence and no new dependency. It is also the
|
||||
substrate every model-based arm writes into — so it is first regardless of how the comparison
|
||||
eventually lands.
|
||||
|
||||
---
|
||||
|
||||
## 12. Arm A results
|
||||
|
||||
Built 2026-08-17. `core/dr-gpu/src/{segment.rs,hierarchy.rs}`,
|
||||
`shaders/watershed.wgsl`, `examples/segment.rs`. 15 tests, 11 of them device-free.
|
||||
|
||||
**It works, and the hierarchy is not the expensive part.** On a 1200×800 synthetic at blur radius 2,
|
||||
release build, RTX 3050 laptop: 6,730 basins and 19,223 boundaries found in **67 ms including the
|
||||
readback**, and the merge tree built from them in **0.2 ms**. The tree is ~0.3% of the cost. The
|
||||
estimate that priced it as a week's work was wrong by about two orders of magnitude, and the reason
|
||||
is worth recording: it is Kruskal over a few thousand nodes, not a segmentation algorithm.
|
||||
|
||||
**The granularity ladder behaves.** At the fine end the background fragments badly — a smooth tonal
|
||||
ramp bands into horizontal strips, and flat areas break into diagonal chains (see below). By
|
||||
`cut_to(300)` all of that is gone: the hard-edged disc is exactly one region, the whole gradient
|
||||
background is one region, and only genuine noise still fragments. The over-segmentation is absorbed
|
||||
by the merge order rather than needing to be prevented, which is the property the whole design rests
|
||||
on.
|
||||
|
||||
**Pre-smoothing is the knob it was claimed to be.** Radius 2 leaves the noisy corner fragmented at
|
||||
300 regions; radius 5 largely clears it. Tying it to ISO is the right control.
|
||||
|
||||
Three findings that change what comes next:
|
||||
|
||||
**F1 — plateaux fragment into diagonal chains.** In an exactly flat region every pixel's steepest
|
||||
descent is a tie, and the (value, index) tie-break sends them all up-left, so a plateau resolves into
|
||||
diagonal streaks rather than one basin. Harmless here because those saddles are ~0 and the tree
|
||||
merges them first — but a real sky or wall is a large plateau, and relying on the hierarchy to clean
|
||||
up an artefact of the flow pass is fragile. The principled fix is a **lower-complete transform**: one
|
||||
extra pass giving plateau pixels a gradient toward their nearest descending exit. Standard, cheap,
|
||||
and worth doing before the corpus work.
|
||||
|
||||
*Attempted, and parked.* The pass exists — `plateau_init` seeds every pixel
|
||||
that has a strictly lower neighbour, `plateau_step` carries a breadth-first
|
||||
distance inward within a level set, and `flow` takes that distance as the
|
||||
second key of a lexicographic tie-break. Bindings, ping-pong and dispatch were
|
||||
all checked and are right. It is nonetheless a **measured no-op**: with a test
|
||||
comparing the labelling at one iteration against sixty-four, *zero* of 9216
|
||||
pixels change basin. That test is committed and ignored rather than deleted,
|
||||
because it is the thing that turned "we think this works" into a fact.
|
||||
|
||||
Three explanations were tried and none of them was it. Exact float equality is
|
||||
certainly wrong — a gradient computed from 8-bit samples is never exactly
|
||||
equal across a region the eye calls flat — and a `LEVEL_EPS` tolerance now
|
||||
replaces `==` and `<` in all three comparisons; it did not change the outcome.
|
||||
Nor did the test image: a flat disc, a terraced disc and a constant-slope ramp
|
||||
all behave identically. Worth knowing for whoever picks this up: on a
|
||||
gradient-*magnitude* watershed, every flat region of the picture sits at
|
||||
gradient zero, which is the global minimum, and a plateau with no descending
|
||||
exit is a minimum — one basin by definition, with nothing for lower-completion
|
||||
to resolve. The plateaux that do have an exit are regions of constant non-zero
|
||||
gradient, which are rarer in a photograph than F1's phrasing suggests.
|
||||
|
||||
`plateau_iterations` therefore defaults to **0**. The pass is off, costs
|
||||
nothing, and F1 stands open.
|
||||
|
||||
**F2 — `cut_to(N)` is a visualisation, not the interaction.** A global cut by region count spends its
|
||||
budget wherever the saddles happen to be densest: at blur 5 the soft-edged disc's interior held a
|
||||
cluster of near-equal saddles and ate the budget, fragmenting at a level where everything else was
|
||||
clean. The real interaction walks up locally from the clicked region and has no such coupling. The
|
||||
ladder in the example should not be read as what a user would experience.
|
||||
|
||||
**F3 — the RAG build still needs a readback.** `Segmentation::read_field` copies labels and gradient
|
||||
to the CPU, gated behind the `readback` feature exactly as `read_pixels` is. Fine for a spike and
|
||||
off the frame path, but a shipping build cannot take it (ARCH §6.1, AC-8), so the adjacency
|
||||
accumulation has to move GPU-side with atomics. That is the largest known gap between this and
|
||||
something shippable.
|
||||
|
||||
**M5 partially answered.** Run-to-run on one device is bit-identical, and the CPU half contributes no
|
||||
nondeterminism of its own — both asserted by tests. Cross-vendor is untouched and remains the
|
||||
measurement that can invalidate the region-id representation.
|
||||
|
||||
---
|
||||
|
||||
## 13. Arms B and C, and what §4 got wrong
|
||||
|
||||
Built 2026-08-21 on branch `local-adjustments`. `core/dr-segment`, `core/dr-gpu/src/mask.rs`,
|
||||
`core/dr-pipeline/src/mask.rs`, and the develop panel.
|
||||
|
||||
**The dependency question dissolved rather than being decided.** §4 and D13 both priced arm B as
|
||||
costing a C dependency under the Android NDK, and treated that as most of the difference between the
|
||||
arms. It is not a cost that has to be paid: `ort` 2.0's `alternative-backend` feature disables the
|
||||
linking entirely and lets another engine supply the `OrtApi`, and `ort-tract` — same authors,
|
||||
MIT/Apache — supplies it from `tract`, which is pure Rust. So arm B runs through `ort`'s API with no
|
||||
C anywhere, and D13's "largest exception the policy would tolerate" turns out not to be needed.
|
||||
|
||||
Measured before committing to it, because tract's operator coverage is the thing that could have
|
||||
sunk it: **yolo26n-seg loads with zero unsupported operators** and runs 640×640 in ~470 ms on the
|
||||
reference desktop's CPU. Correct masks on the standard `bus.jpg` — one bus and three people, outlines
|
||||
following the subjects.
|
||||
|
||||
Three findings that contradict §4 directly, and all three change the design rather than the schedule.
|
||||
|
||||
**F4 — there is no ADE20K-trained YOLO.** §4's whole argument for arm B was ADE20K's 150 classes and
|
||||
their *stuff* categories: "'that patch of sky' is a class here." Checked 2026-08-21: Ultralytics ships
|
||||
YOLO26-seg trained on **COCO**, whose 80 classes are all *things*, and the one HuggingFace repository
|
||||
claiming a YOLO/ADE20K combination is empty. ADE20K models exist as SegFormer/OneFormer/MaskFormer
|
||||
transformers, not as YOLO.
|
||||
|
||||
So the shipped vocabulary selects **subjects**, not **stuff**. "Select the person" works; "select the
|
||||
sky" does not come from the model at all and must come from the watershed. That is a narrower arm B
|
||||
than §4 assumed, and it *raises* the importance of arm C rather than lowering it — the model can no
|
||||
longer be the whole answer for anything.
|
||||
|
||||
**F5 — it is instance segmentation, not semantic segmentation.** §4 assumed a flat partition with
|
||||
class ids that would "populate `RegionField` directly". YOLO-seg does not partition the image; it
|
||||
finds objects, and most pixels in a landscape belong to no instance. Two consequences, one bad and
|
||||
one better than expected: nothing populates a `RegionField` on its own, and two people come back as
|
||||
*two* instances where a semantic model would have returned one "person" area covering both. For
|
||||
selecting a subject the latter is the behaviour worth having.
|
||||
|
||||
**F6 — tract cannot parse a dynamic-shape export.** It fails shape inference on the neck's `Concat`.
|
||||
The graph therefore ships with its input fixed at 640×640 square, and every image is letterboxed into
|
||||
it. This is the constraint behind the tiling option in `semantic.rs`: with a fixed window, tiling is
|
||||
the *only* route to more semantic resolution, and it costs one inference per tile (≈2.8 s for a 3×2
|
||||
grid over a 1600 px proxy against 470 ms whole-frame). Defaulted off — a photographic subject is
|
||||
usually large in frame, which is the case whole-frame inference handles best — and left implemented
|
||||
so §9's corpus can settle it rather than an argument.
|
||||
|
||||
**Arm C ships, and the §8 criteria were not what decided it.** §8 asked for a two-interaction margin
|
||||
over arm A on the scene subset. That comparison was never run, because F4 and F5 changed what the
|
||||
arms *are*: with a model that recognises subjects and has no word for sky, arm B alone cannot be a
|
||||
selection tool at all (§8's "arm B never ships alone" holds, for a stronger reason than expected),
|
||||
and arm A alone cannot tell a person from the wall behind them. They are complements rather than
|
||||
candidates. Arm C's implementation is `prior.rs`: instance membership re-weights the merge saddles,
|
||||
so region pairs the model believes share an object merge early and pairs straddling its edge merge
|
||||
late. **No boundary moves** — only the order in which boundaries dissolve — which is how the result
|
||||
stays pixel-accurate at every level while its coarse levels become named things.
|
||||
|
||||
**M1–M4 remain unmeasured.** The 24-image corpus of §9 has not been traced, so there are no
|
||||
interaction counts and no boundary-accuracy numbers. What exists is a working feature and the
|
||||
evidence that each piece does what it claims in isolation. The corpus is still the thing that would
|
||||
turn "this feels right" into a number, and it is the largest piece of §11 left undone.
|
||||
|
||||
**M5 is unchanged and still the risk it was.** Run-to-run on one device is identical, asserted by a
|
||||
test. Cross-vendor is untouched. Region ids now reach the sidecar, so if the label field proves
|
||||
non-deterministic across vendors a mask written on the desktop will not mean the same thing on
|
||||
Android — see `MaskSource::Regions::signature`, which detects a *retuned* segmentation but not a
|
||||
differently-rounded one.
|
||||
|
||||
**F3 still stands.** `Segmentation::read_field` still copies the label and gradient buffers to the
|
||||
CPU to build the region graph. It is now behind its own `segment-readback` feature rather than
|
||||
sharing `readback` — this transfer is once per image on a worker, where the one AC-8 forbids is per
|
||||
frame in the render loop — but the accumulation still belongs GPU-side with atomics.
|
||||
|
||||
---
|
||||
|
||||
## 14. Register entries
|
||||
|
||||
**S15** — *Region segmentation for local masking* · **CLOSED 2026-08-21**. Arms A, B and C built; the
|
||||
shared consumers built; the licence question resolved (§7, D14). The 24-image corpus was not traced,
|
||||
so M1–M4 are unmeasured and M5 is answered only on one device. Answers: local masking snaps to arm C,
|
||||
and a mask is stored as region ids. Relates to: D13, D14, FR-DEV-3, ARCH §5.4, §6.13.
|
||||
|
||||
**D14** — *Segmentation source for local masking* · **DECIDED 2026-08-21**: **arm C**, a watershed
|
||||
hierarchy re-weighted by YOLO26n-seg instance membership, with the model optional and the watershed
|
||||
sufficient without it. Weights ship in-tree under AGPL-3.0, which GPLv3 §13 permits and which makes
|
||||
the combined work effectively AGPL — a deliberate change to DarkRoom's licensing posture, not a
|
||||
dependency detail (`core/dr-segment/models/LICENCE.md`).
|
||||
|
||||
**D13** — *inference runtime* · the dependency half is **answered** for segmentation and the answer
|
||||
generalises: `ort` + `ort-tract` gives ONNX inference in pure Rust, so the face pipeline of §3.9.1
|
||||
needs no C dependency either. The *model licensing* half of D13 is untouched — the InsightFace
|
||||
weights are still non-commercial and still unusable here.
|
||||
|
||||
---
|
||||
|
||||
## 16. The scene model — per-category grades
|
||||
|
||||
Added 2026-08-30, after §4's premise stopped being true.
|
||||
|
||||
### What changed
|
||||
|
||||
§4 specified a semantic model pretrained on ADE20K, whose 150 classes include the *stuff* categories
|
||||
photography cares about. §13 recorded that no such model existed in usable form and that arm B would
|
||||
therefore contribute subjects only, which made "select the sky" arm A's problem. Re-checked
|
||||
2026-08-30: **Ultralytics now ships a `semantic` task with ADE20K checkpoints**
|
||||
(`docs.ultralytics.com/tasks/semantic`). `yolo26s-sem-ade20k` is in `models/scene/`.
|
||||
|
||||
### It is an addition, not a correction to arm B
|
||||
|
||||
The instance model stays exactly where it was, and the reason is the one §13 already gave and was
|
||||
right about: a semantic model merges every pixel of a class into one region, so it cannot separate
|
||||
two people, and separating two people is what clicking a subject requires. Swapping arm B for this
|
||||
would regress the primary interaction to fix a secondary one.
|
||||
|
||||
So the two divide by *what the user is doing*, not by which is better:
|
||||
|
||||
| | `models/segment/` (COCO instances) | `models/scene/` (ADE20K semantics) |
|
||||
|---|---|---|
|
||||
| Question | which pixels are *that* dog | how much of this pixel is sky |
|
||||
| Granularity | per instance | per category, whole frame |
|
||||
| Drives | local adjustments, subject selection | the scene tab's per-category sliders |
|
||||
| Vocabulary | 80 things | 150 classes, stuff included |
|
||||
|
||||
### The export is truncated, and both reasons matter
|
||||
|
||||
Ultralytics ends the graph with `Resize → ArgMax → Cast`, returning a `[1, 640, 640]` u8 label map.
|
||||
`tools/export-seg-model.sh` cuts that tail and ships the classifier's `[1, 150, 80, 80]` f32 logits.
|
||||
|
||||
**Cost.** The `Resize` materialises 150 × 640 × 640 × f32 — 246 MB — and the `ArgMax` then reduces
|
||||
across the channel axis, striding 409,600 elements per comparison. Measured under load it was
|
||||
roughly four fifths of total runtime, spent on work the application discards.
|
||||
|
||||
**Softness, which is the more important one.** `ArgMax` destroys the per-class scores, and the whole
|
||||
design of the scene tab rests on keeping them. Softmax over the 150 channels, summed within each
|
||||
category, produces per-category weights that sum to one at every pixel — a partition of unity.
|
||||
Feathering that cannot double-grade a boundary. Feathering *hard labels* outward from two adjacent
|
||||
categories paints both grades into the overlap, and every horizon in the frame acquires a seam.
|
||||
|
||||
### The resolution is 80×80, and no setting changes that
|
||||
|
||||
The discarded upsample was never information. `Scene` keeps the native grid and resamples on demand,
|
||||
so the coarseness is visible in the type rather than hidden. A graduated grade over sky or water is
|
||||
untroubled by it; a rooftop against sky at 100% zoom will show it. This is the constraint most likely
|
||||
to decide whether the tab feels good, and it is not addressable by choosing a larger checkpoint —
|
||||
`yolo26n-sem` and `yolo26s-sem` have the same output grid.
|
||||
|
||||
### Licence
|
||||
|
||||
Unchanged. Same AGPL-3.0 grant as the instance model, same GPLv3 §13 permission, same consequence
|
||||
already accepted in D14 — so this needed no new licence decision, which is most of why it was cheap.
|
||||
See `models/LICENCE.md`.
|
||||
|
||||
### Measurement
|
||||
|
||||
Timings taken while this was chosen came off a laptop compiling other things and are upper bounds
|
||||
only. `cargo run -p dr-segment --example scene --release --features embedded-scene-model` reports a
|
||||
median over N runs with the first excluded; a number worth quoting should come from that, on an idle
|
||||
machine.
|
||||
@@ -0,0 +1,575 @@
|
||||
# Spot removal
|
||||
|
||||
**Status:** Draft · 2026-08-26
|
||||
**Companion to:** [requirements.md](requirements.md) §3.3 FR-DEV-8 · [architecture.md](architecture.md) §5.2
|
||||
|
||||
The last develop feature the requirements ask for that nothing in the tree
|
||||
implements. FR-DEV-8 states the shape — "non-destructive clone and heal spots
|
||||
stored as parameters in the edit graph (target, radius, feather, source offset,
|
||||
opacity, mode), with automatic source placement and manual override, plus a
|
||||
visualise-spots mode" — and this document is how that lands on the pipeline
|
||||
that exists now.
|
||||
|
||||
---
|
||||
|
||||
## 1. Why it is worth the work
|
||||
|
||||
Sensor dust is unavoidable with interchangeable lenses, and a dust spot is the
|
||||
most common reason a photographer leaves a RAW editor for a pixel editor
|
||||
mid-workflow. Every other develop operation in this application can be the best
|
||||
one in its class and the workflow still breaks at the first frame with a mark on
|
||||
the sky.
|
||||
|
||||
It is also, unusually, a feature whose cost has already been paid twice over.
|
||||
The neighbourhood stage exists ([`crate::detail`](../../core/dr-pipeline/src/detail.rs)),
|
||||
the convention for storing geometry in normalised source coordinates exists
|
||||
([`mask.rs`](../../core/dr-pipeline/src/mask.rs)), the canvas-drag pattern exists
|
||||
([`gradient.rs`](../../ui/dr-ui/src/gradient.rs)), and the merge-by-id rule exists
|
||||
([`sidecar.rs`](../../core/dr-pipeline/src/sidecar.rs)). What is genuinely new is
|
||||
small and is named in §3.
|
||||
|
||||
## 2. Non-goals
|
||||
|
||||
- **Not layer-based pixel editing.** §1.3 of the requirements excludes that and
|
||||
this does not reopen it. A spot is a handful of numbers in the edit graph; no
|
||||
pixels are stored, and the original file is never touched.
|
||||
- **Not content-aware fill.** The source is a patch from the same photograph,
|
||||
chosen by an offset. Synthesising texture that is not in the frame is a
|
||||
different problem with a different budget.
|
||||
- **Not a general clone brush.** A spot is a disc, not a stroke. A dragged
|
||||
clone brush is expressible on top of this (a stroke *is* a run of discs) and
|
||||
is deliberately left until the disc is finished and used.
|
||||
- **Not automatic dust detection.** Finding spots without being asked is a
|
||||
reasonable later feature and a bad first one: a false positive silently alters
|
||||
a photograph, which is the failure this application must not have.
|
||||
|
||||
## 3. What is new, precisely
|
||||
|
||||
Four things, and it is worth being blunt about them because everything else in
|
||||
this document is assembly of parts that already work:
|
||||
|
||||
1. **A detail pass with variable-length data.** Every [`DetailPass`] today
|
||||
carries a `Vec<f32>` of uniforms fixed by its own structure. A spot list is
|
||||
neither fixed nor small. §7.
|
||||
2. **A neighbourhood operation whose reach is not a small kernel.** Every
|
||||
existing pass declares a halo of a few pixels. A spot reads from wherever its
|
||||
source is, which may be a third of the frame away. §5.3.
|
||||
3. **Undo over something that is not a parameter.** [`History`] snapshots a
|
||||
[`Preset`], which is a map of scalars — so mask edits are already outside
|
||||
undo, and spots must not be. §10.4.
|
||||
4. **A canvas mode that *creates* objects.** Crop edits one rect; local selects
|
||||
a region; gradient drags an existing shape. Nothing yet makes a new thing
|
||||
where the pointer went down. §10.
|
||||
|
||||
## 4. The model
|
||||
|
||||
```rust
|
||||
/// TRACES: FR-DEV-8
|
||||
pub struct Spot {
|
||||
/// Stable across devices; see below.
|
||||
pub id: String,
|
||||
/// What is being covered, in normalised **source** coordinates.
|
||||
pub centre: (f32, f32),
|
||||
/// What covers it, as an offset from `centre` in **frame units**
|
||||
/// (y spans 0..1, x spans 0..aspect — the mask convention).
|
||||
pub offset: (f32, f32),
|
||||
/// The radius of the disc, in frame units.
|
||||
pub radius: f32,
|
||||
/// Fraction of `radius` over which the edge falls away. 0 is hard.
|
||||
pub feather: f32,
|
||||
/// How much of the patch is laid down. 1.0 is opaque.
|
||||
pub opacity: f32,
|
||||
pub mode: SpotMode, // Heal | Clone
|
||||
pub enabled: bool,
|
||||
}
|
||||
```
|
||||
|
||||
**Units follow the mask rule, for the mask reason.** A length stored in pixels
|
||||
is a length that means something different in the preview and in the export
|
||||
([`RenderScale`]'s whole documentation is this argument). `centre` is normalised
|
||||
source, so a crop, a zoom, a pan and a rotation move the spot with the
|
||||
photograph and no arithmetic is needed to keep it there.
|
||||
|
||||
Every *length* — radius, feather, offset — is in the frame's **isotropic
|
||||
units**, `MaskSource::Radial`'s convention, where y spans `0..1` and x spans
|
||||
`0..aspect`. Only in those units is a disc a disc: normalised coordinates would
|
||||
make a spot on a 3:2 frame an ellipse half again wider than it is tall. They are
|
||||
lengths against the *source* frame rather than the rendered region, so cropping
|
||||
does not resize a spot already placed — a dust mark is a fact about the sensor,
|
||||
not about the composition.
|
||||
|
||||
One unit for all three, deliberately. A radius in shorter-edge fractions beside
|
||||
an offset in frame units agrees on a landscape frame and silently disagrees on a
|
||||
portrait one, which is a bug that stays invisible until somebody rotates a
|
||||
photograph.
|
||||
|
||||
**`offset` is a vector, not a second point.** Dragging the destination moves the
|
||||
source with it, which is what a photographer expects when they nudge a spot half
|
||||
a pixel and do not want to re-place the source. Moving the source alone is
|
||||
editing `offset`.
|
||||
|
||||
**The id is derived, not counted.** `MaskStack::next_id` numbers layers, which
|
||||
is fine for a stack a user names, and wrong here: two devices that each place a
|
||||
spot offline would both produce `spot3`, and the merge in §11 would treat two
|
||||
different marks as one. So the id is a short base-36 hash of the centre at
|
||||
creation, and two devices that place a spot in the same place produce the same
|
||||
id — which is the correct outcome, because they removed the same piece of dust.
|
||||
|
||||
```rust
|
||||
pub struct SpotSet {
|
||||
spots: Vec<Spot>, // in creation order; the order matters, see §5.2
|
||||
}
|
||||
```
|
||||
|
||||
### 4.1 Why it lives beside `ops`, not in it
|
||||
|
||||
The [`Operation`] trait takes a `ParamId` and returns an `f32`, and the whole
|
||||
generic machinery above it — the panel, the sidecar, the presets, the history —
|
||||
is built on that being true. A spot list is not scalars, and the trait says so
|
||||
explicitly where it refuses a downcast for film tables.
|
||||
|
||||
[`EditGraph`] already holds three things that are not operations for exactly
|
||||
this reason: `framing`, `masks` and `film`. `spots` is the fourth, and the
|
||||
argument is the same one `masks` makes — a stack of layers is not a slider, and
|
||||
folding it into the list would make every consumer that walks `ops` know that
|
||||
some entries are not really operations.
|
||||
|
||||
Bounds, following `mask.rs`'s example of bounding what a sidecar can grow to:
|
||||
|
||||
| Constant | Value | Why |
|
||||
|---|---|---|
|
||||
| `MAX_SPOTS` | 64 | Beyond a few dozen the answer is to clean the sensor. Refuses rather than dropping, as `MaskStack::push` does. |
|
||||
| `MAX_SOURCE_DISTANCE` | 0.5 | Frame units. Bounds the halo in §5.3, which is otherwise unbounded. |
|
||||
| `DEFAULT_RADIUS` | 0.012 | Frame units — about 25 px on a 24 MP frame's short edge, which is a dust mark. |
|
||||
| `MIN_RADIUS` / `MAX_RADIUS` | 0.001 / 0.5 | Not zero, because a spot that repairs nothing reads as a broken tool; not larger, because the halo bound has to mean something. |
|
||||
| `DEFAULT_FEATHER` | 0.35 | Fraction of the radius. Soft enough that a heal on a gradient sky has no visible boundary. |
|
||||
|
||||
## 5. Where it runs
|
||||
|
||||
### 5.1 First in the detail chain
|
||||
|
||||
ARCH §5.2 draws spot removal *after* texture and clarity and *before* sharpen
|
||||
and NR. That diagram is already out of step with the operation set — the
|
||||
`order:` keys in `core/dr-pipeline/ops/` put noise reduction at 110 and capture
|
||||
sharpening at 120, ahead of clarity at 130 and texture at 140 — so it needs a
|
||||
correction anyway, and the correction should put spot removal **first among the
|
||||
neighbourhood passes**, at a notional order of 105.
|
||||
|
||||
The reason is the halo. Sharpening a dust spot before removing it amplifies its
|
||||
edge, and the amplified edge is wider than the spot: the sharpening kernel has
|
||||
already smeared a dark ring into pixels that the spot's own disc does not cover,
|
||||
so the heal leaves a faint circle of over-sharpened background around a patch
|
||||
that is otherwise perfect. Removing the mark first means every later pass sees a
|
||||
photograph with no mark in it, which is also the photograph the photographer
|
||||
thinks they are sharpening.
|
||||
|
||||
Since the spot set is not in `ops`, `EditGraph::compose_detail_for` splices its
|
||||
passes in front of the ops' passes rather than sorting by a declared order. That
|
||||
is a two-line change and it is stated here so nobody looks for a `spots.yaml`.
|
||||
|
||||
### 5.2 Rounds, because sources can read destinations
|
||||
|
||||
Every pass reads one texture and writes another. So within a single pass, every
|
||||
spot reads the *unhealed* image — and a spot whose source overlaps an earlier
|
||||
spot's destination copies the mark the earlier spot was removing.
|
||||
|
||||
The fix is not to run one pass per spot (64 dispatches for a frame that needs
|
||||
one). It is to group: walking the spots in creation order, a spot joins the
|
||||
current round unless its source disc intersects the destination disc of a spot
|
||||
already in that round, in which case it opens a new one. One pass per round, and
|
||||
the common case — spots scattered over a sky, sources near their own
|
||||
destinations — is a single round. The grouping is plain CPU code over at most 64
|
||||
discs and belongs in `SpotSet`, with a test that says an overlapping pair
|
||||
produces two rounds and a disjoint pair produces one.
|
||||
|
||||
### 5.3 The halo, honestly
|
||||
|
||||
[`DetailPass::radius`] is "the furthest this pass reads from the pixel it
|
||||
writes", and it exists so that ARCH §5.3's tile scheduler knows how far to grow
|
||||
a tile. For a spot pass that is `max(|offset| + radius)` over the pass's spots,
|
||||
in render pixels — which with `MAX_SOURCE_DISTANCE` at 0.5 can approach half the
|
||||
frame.
|
||||
|
||||
That is a real cost and it should be written down rather than discovered: a
|
||||
frame with a long-armed spot is close to untileable for that one pass, so the
|
||||
tiled path will compute it whole-frame. Two things keep it affordable. The pass
|
||||
is cheap per pixel (§6.4), and it is only the *spot* passes that carry the halo
|
||||
— the sharpening pass after it still declares its three pixels and still tiles.
|
||||
The alternative, clamping the source distance to something tile-sized, would
|
||||
make the tool useless exactly where it is most needed: a mark on a face is
|
||||
healed from the other cheek, and that is a long way.
|
||||
|
||||
## 6. What a spot does to the pixels
|
||||
|
||||
Both modes work on the same disc. For a pixel at render coordinate `p` inside a
|
||||
spot centred at `d` with radius `r`, with the source at `s = d + offset`:
|
||||
|
||||
```
|
||||
w = falloff(|p - d| / r) // 1 at the centre, 0 at the rim
|
||||
patch = bilinear(source_texture, p - d + s)
|
||||
c = mix(c, patch + membrane, w * opacity)
|
||||
```
|
||||
|
||||
`falloff` is a smoothstep over the outer `feather` fraction of the radius; a
|
||||
feather of 0 is a hard disc. `bilinear` is four `tap`s and two lerps, because
|
||||
the generated preamble offers `textureLoad` only and the offset is fractional in
|
||||
render space — it becomes a [`Helper`], deduplicated across passes like the
|
||||
existing luminance helper.
|
||||
|
||||
`membrane` is what separates the two modes, and it is zero for `Clone`.
|
||||
|
||||
### 6.1 Heal, without a Poisson solve — **implemented**
|
||||
|
||||
The classic heal is Poisson blending: copy the *gradients* of the source and
|
||||
solve for the image whose gradients they are, subject to matching the
|
||||
destination on the boundary. Solved properly that is an iterative linear system
|
||||
— tens of Jacobi passes over the disc — and each iteration is a dispatch in this
|
||||
architecture. Sixty dispatches to remove a dust spot is not a frame budget.
|
||||
|
||||
What that solve produces is a smooth membrane interpolating the boundary
|
||||
difference, and a membrane can be interpolated directly instead of solved. The
|
||||
shipped form samples the difference between destination and source at `K` points
|
||||
around the rim and interpolates them into the interior by inverse square
|
||||
distance:
|
||||
|
||||
```
|
||||
for k in 0..K:
|
||||
b_k = tap(rim_k) - tap(rim_k + offset) // boundary difference
|
||||
w_k = 1 / max(|p - rim_k|², 1)
|
||||
membrane = Σ w_k·b_k / Σ w_k
|
||||
```
|
||||
|
||||
`K = 24` — `RIM_SAMPLES` in `core/dr-pipeline/src/spot.rs`, a uniform rather
|
||||
than a constant in the source, so tuning it uploads a buffer instead of
|
||||
recompiling. The cost is `2K` bilinear samples per pixel *inside a disc*, and
|
||||
nothing at all outside one.
|
||||
|
||||
**What was originally specified here was mean-value seamless cloning** (Farbman
|
||||
et al., 2009), whose weights are half-angle tangents over the rim rather than
|
||||
inverse squares. The difference matters when the boundary difference varies
|
||||
sharply around the rim; on the case that actually arises — a repair on a
|
||||
smoothly varying background — both reduce to the same answer, and the inverse
|
||||
square form costs two transcendentals per sample fewer. The measurement in
|
||||
`core/dr-gpu/tests/spot_removal.rs` is what decides whether that trade stays
|
||||
good: on a linear ramp steep enough to make a clone wrong by 38 levels out of
|
||||
255, the heal is wrong by **0**. If a case turns up where it is not, the weights
|
||||
are four lines and the tests are already written.
|
||||
|
||||
### 6.2 Why not compute the boundary statistics on the CPU
|
||||
|
||||
Because that means reading the rendered image back, and FR-DEV-4 forbids it in
|
||||
the render path for reasons ARCH §6.1 spends a page on. A per-frame readback to
|
||||
find out what colour a sky is would reintroduce exactly the stall the whole
|
||||
architecture exists to avoid. The mean-value form needs no reduction at all,
|
||||
which is most of why it is the right answer here.
|
||||
|
||||
### 6.3 Clone
|
||||
|
||||
`membrane = 0`. Kept because heal is wrong on a boundary: a spot straddling a
|
||||
horizon healed by mean-value blending smears the horizon's contrast into the
|
||||
disc, and the honest tool then is a straight copy from a matching part of the
|
||||
frame. This is why FR-DEV-8 asks for both, and it costs one branch in the shader
|
||||
and one segmented control in the panel.
|
||||
|
||||
### 6.4 Cost
|
||||
|
||||
Per pixel, per pass: a rejection test per spot in the pass (a squared distance
|
||||
and a compare), and for the pixels actually inside a disc, `4 + 2K` taps. With
|
||||
64 spots the rejection cost is the dominant term and it is about 64 × 4 ALU ops
|
||||
on every pixel of the frame — call it a millisecond at 2 MP on integrated
|
||||
graphics, which is affordable but not free.
|
||||
|
||||
If it proves not to be, the fix is the one `mask.rs` already uses for strokes: a
|
||||
bounding box per spot and a dispatch sized to it. That needs the detail runner
|
||||
to dispatch something other than the whole frame, which is a change to its
|
||||
shape, and it is deliberately not being made until a measurement asks for it.
|
||||
|
||||
## 7. Getting a spot list to the GPU
|
||||
|
||||
[`DetailPass`] gains one field:
|
||||
|
||||
```rust
|
||||
/// Per-instance data too large or too variable for the uniform block.
|
||||
pub storage: Option<Vec<[f32; 4]>>,
|
||||
```
|
||||
|
||||
and the generated preamble gains one binding:
|
||||
|
||||
```wgsl
|
||||
@group(0) @binding(3) var<storage, read> instances: array<vec4<f32>>;
|
||||
```
|
||||
|
||||
In `dr-gpu`, both bind group layouts gain a read-only storage entry at binding
|
||||
3, and a pass that declares no storage binds a shared one-element dummy buffer.
|
||||
wgpu permits a layout entry the shader does not use, so the two layouts stay two
|
||||
rather than four, and no existing pass changes at all.
|
||||
|
||||
**A spot is two `vec4`s**: `(centre.x, centre.y, radius, feather)` and
|
||||
`(offset.x, offset.y, opacity, flags)`, all in render pixels except `flags`,
|
||||
converted on the CPU in `compose_detail_for` where the framing is in scope. This
|
||||
matters: the shader never sees a normalised coordinate and never has to know
|
||||
about crop, rotation or zoom — `Framing::output_at` does that map on the way in,
|
||||
exactly as `gradient.rs` does it for handles. The framing is an affine
|
||||
similarity, so a disc stays a disc and one radius scales by one factor:
|
||||
`radius_px = radius × min(source_w, source_h) × scale.ratio()`.
|
||||
|
||||
**The source list does not recompile anything.** The WGSL is identical for one
|
||||
spot and for sixty-four — the count is a uniform and the loop is over the
|
||||
buffer — so `structure_hash` is unchanged as spots are placed, and placing the
|
||||
tenth spot re-uploads a 512-byte buffer. This is the same property the fused
|
||||
pass has for slider movement and it is worth a test that asserts
|
||||
`cached_pipelines()` does not grow while spots are added.
|
||||
|
||||
**Rejected: packing spots into the uniform block.** It would touch no bind group
|
||||
layout, which is genuinely attractive. It also requires the composer to emit
|
||||
`vec4` uniform fields (it emits scalars), forces a fixed `MAX_SPOTS`-sized array
|
||||
and its fixed upload cost into every spot pass, and gives the next operation
|
||||
that wants a table — a LUT, a curve, a lens grid — nothing to build on. The
|
||||
storage buffer is a few more lines once and useful again later.
|
||||
|
||||
**Invalidation.** The spot set folds into the `detail` key in
|
||||
`EditGraph::invalidation`, alongside the detail operations. Dragging a spot
|
||||
therefore re-runs the detail chain and *not* the fused colour pass or the
|
||||
demosaic, which is exactly the reuse FR-DEV-3d asks for and is the difference
|
||||
between a spot that follows the finger and one that stutters.
|
||||
|
||||
## 8. Automatic source placement
|
||||
|
||||
FR-DEV-8 asks for automatic placement with manual override. Two stages, because
|
||||
the useful half is much cheaper than the good half.
|
||||
|
||||
**Stage one — a placed default.** A new spot's source is offset by `2.5 × radius`
|
||||
in the direction that keeps it furthest inside the frame, biased towards the
|
||||
frame centre. For dust on a sky, which is the overwhelming majority of spots,
|
||||
this is right often enough to be worth having, and it is wrong in a way that is
|
||||
immediately visible and one drag from fixed.
|
||||
|
||||
**Stage two — a scored search.** A compute dispatch per new spot scores candidate
|
||||
offsets on two rings around the destination (say 32 candidates), each scored by
|
||||
the sum of squared differences over the annulus just outside the destination
|
||||
disc — the ring is what has to match, since the disc's interior is being
|
||||
replaced anyway. Penalise candidates whose disc overlaps another spot's
|
||||
destination, or the frame edge. The winner's offset is read back **once**, when
|
||||
the spot is created, through `readback.rs` — a few hundred bytes, which is the
|
||||
size the histogram already moves and completes in well under a frame — and
|
||||
written into the spot.
|
||||
|
||||
Two rules about that readback, both of which are the difference between a
|
||||
feature and a bug:
|
||||
|
||||
- It is **not** in the render loop. `readback.rs` blocks with a deadline, and
|
||||
that is tolerable exactly once per placement and intolerable per frame. It
|
||||
happens on the gesture, and the render that follows uses whatever the spot
|
||||
currently holds.
|
||||
- The result is **stored**, and the search is never re-run behind the user. A
|
||||
spot whose source moved on its own when the file was reopened would be an edit
|
||||
changing itself, and non-destructive editing means the sidecar decides what
|
||||
the picture is.
|
||||
|
||||
If the readback fails, the stage-one default stands. There is no state in which
|
||||
a spot has no source.
|
||||
|
||||
## 9. Visualising spots
|
||||
|
||||
FR-DEV-8's "visualise-spots mode" is two different things, and conflating them
|
||||
is how one of them ends up missing:
|
||||
|
||||
**The overlay** — where the spots *are*. Circles for the destination, a fainter
|
||||
circle for the source, a line between them for the selected spot. Drawn in Slint
|
||||
over the canvas, alongside the gradient handles and by the same coordinate map,
|
||||
so it costs the render path nothing and cannot leak into an export.
|
||||
|
||||
**The reveal** — where the spots *should be*. Lightroom's "Visualize Spots": a
|
||||
high-contrast, desaturated view of the frame's high-frequency content, in which
|
||||
sensor dust on a smooth sky is obvious and in a normal view is nearly invisible.
|
||||
It is a detail pass appended to the chain:
|
||||
|
||||
```
|
||||
c = abs(c - blur(c)) stretched by a threshold, greyscale, inverted
|
||||
```
|
||||
|
||||
It is a **view**, not an edit. So it is not in the graph and not in the sidecar:
|
||||
`compose_detail_for` takes a `DetailView` (`Normal` | `RevealSpots`) and the
|
||||
export path passes `Normal`. A flag on the session would work until the day
|
||||
somebody exports while the mode is on, and then it would produce a black-and-
|
||||
white file that looks like corruption. Making the export call site name it is
|
||||
what stops that from ever being possible.
|
||||
|
||||
## 10. Interaction
|
||||
|
||||
### 10.1 The mode
|
||||
|
||||
A `ViewMode::spots`, and a row in `ToolRail`'s table beside Crop and Local.
|
||||
That table's own documentation already predicts this shape for the brush; the
|
||||
spot tool is the same shape and arrives first.
|
||||
|
||||
(It landed as a third chip in `ModeStrip`, at the head of the develop column.
|
||||
The three canvas tools have since moved out to a fixed rail down the left of
|
||||
the develop view — `ui/dr-ui/ui/toolrail.slint` carries why — and what remains
|
||||
of that strip is the adjustment-group filters, now `GroupStrip`. Nothing about
|
||||
the mode itself changed in the move.)
|
||||
|
||||
### 10.2 Gestures
|
||||
|
||||
| Gesture | Effect |
|
||||
|---|---|
|
||||
| Tap / click on the photograph | Place a spot at the current radius, source auto-placed (§8), and select it |
|
||||
| Drag from a spot's centre | Move the destination; the source follows |
|
||||
| Drag from the source circle | Change the offset |
|
||||
| Drag *out* from a fresh placement | Set the source directly, without the auto-placement |
|
||||
| Scroll / pinch on a selected spot | Radius |
|
||||
| Tap a spot | Select it; the panel scopes to it |
|
||||
| `Delete` / `Backspace` | Remove the selected spot |
|
||||
| Alt-click a spot | Remove it without selecting first |
|
||||
| `Esc` | Leave the mode |
|
||||
|
||||
A drag is a displacement from the press, not a snap to the pointer — the rule
|
||||
`gradient.rs` states and for the same reason: a finger-sized touch target snapped
|
||||
to the pointer jumps by half a target the instant it is grabbed.
|
||||
|
||||
### 10.3 The panel
|
||||
|
||||
While a spot is selected, the adjust column shows radius, feather, opacity and
|
||||
a Heal/Clone control for *that* spot, exactly as selecting a mask layer
|
||||
re-scopes the column today. With nothing selected it shows the defaults new
|
||||
spots will be created with, plus the reveal toggle.
|
||||
|
||||
### 10.4 Undo
|
||||
|
||||
[`History`] snapshots a [`Preset`], which is a parameter map — so today mask
|
||||
edits are not undoable, and spot placement must not inherit that. `History`
|
||||
should hold `(Preset, SpotSet)` and restore both.
|
||||
|
||||
That is a narrow change with a wide benefit: the same door lets the mask stack
|
||||
join later, which closes a gap FR-DEV-5 has open right now. The coalescing rule
|
||||
needs one addition — a drag of one spot's handle is one step, keyed by the spot
|
||||
id in the same way a slider drag is keyed by its control — and placement,
|
||||
deletion and mode changes each open a step of their own.
|
||||
|
||||
### 10.5 Touch
|
||||
|
||||
Every handle is a `Theme.touch-target`, per FR-UI-3. On a phone the destination
|
||||
and source circles of a small spot overlap at that size, so the source handle is
|
||||
drawn at a minimum arm length from the centre while the stored offset is
|
||||
untouched — the trick `gradient.rs` uses with `MIN_ARM`, for the identical
|
||||
reason: a handle that cannot be grabbed again is a one-way edit.
|
||||
|
||||
## 11. Persistence
|
||||
|
||||
One line per spot in the version block:
|
||||
|
||||
```
|
||||
[version 8f04c0e2-…]
|
||||
exposure.exposure = 0.75
|
||||
spot.3f9k = 0.4213 0.2871 0.0120 0.35 0.0310 -0.0180 1 heal
|
||||
```
|
||||
|
||||
Fields in order: `centre.x centre.y radius feather offset.x offset.y opacity
|
||||
mode`. A line rather than a block because a spot is eight numbers and sixty-four
|
||||
blocks would bury the rest of the file; a line *per spot* rather than one line
|
||||
for the set because the line is the unit of merge and of a readable diff — the
|
||||
same reasoning `write_strokes` gives for a line per stroke.
|
||||
|
||||
Coordinates are written at the same precision they are held at, as strokes are,
|
||||
so a round trip is exact and two devices do not generate a diff of noise in the
|
||||
sixth decimal.
|
||||
|
||||
**A malformed line costs that spot and not the file.** A truncated line is
|
||||
dropped with a warning, exactly as `parse_stroke` drops a bad stroke: a spot
|
||||
that silently lands somewhere the user never put it is worse than a spot that is
|
||||
missing, because only one of the two is noticeable.
|
||||
|
||||
**Merge** follows `merge_masks` precisely: by id, disjoint survives, a spot both
|
||||
sides edited resolves wholesale to the higher revision. Half of one device's
|
||||
offset with the other's radius is a repair neither photographer made. Deletion
|
||||
propagates through the base comparison exactly as a layer's does.
|
||||
|
||||
**Presets do not carry spots** in the first version — a preset is a look, and a
|
||||
look does not include where the dust was. But dust is in the *same place on
|
||||
every frame from that body*, which makes "copy spot removal to the selection"
|
||||
genuinely valuable, and it is a stage of its own (§12, S7) rather than a
|
||||
surprise inside the existing paste.
|
||||
|
||||
## 12. Stages
|
||||
|
||||
Each stage is shippable and each has something to look at. Test names are the
|
||||
files they belong in.
|
||||
|
||||
**S1 — The model.** *(Done.)* `spot.rs` in `dr-pipeline`: `Spot`, `SpotMode`, `SpotSet`,
|
||||
the bounds, id derivation, round grouping (§5.2). No GPU, no UI.
|
||||
*Tests:* `core/dr-pipeline/tests/spots.rs` — id stability across two identical
|
||||
placements, `MAX_SPOTS` refuses rather than drops, overlapping sources produce
|
||||
two rounds, disjoint produce one.
|
||||
|
||||
**S2 — Persistence.** *(Done.)* Sidecar write, parse, round trip, merge.
|
||||
*Acceptance:* a hand-written sidecar with three spots survives a load/save round
|
||||
trip byte-identically, and two devices that each add a spot offline end with
|
||||
both.
|
||||
|
||||
**S3 — The storage binding.** *(Done.)* `DetailPass::storage`, the preamble's binding 3,
|
||||
the dummy buffer, the two layouts.
|
||||
*Acceptance:* every existing detail test still passes untouched, and a synthetic
|
||||
pass reading the buffer gets what was uploaded.
|
||||
|
||||
**S4 — Clone.** *(Done.)* The disc, the feather, the bilinear helper, the pass grouping,
|
||||
the halo declaration, spliced first into the chain.
|
||||
*Acceptance:* `core/dr-gpu/tests/spot_removal.rs` — a synthetic frame with a
|
||||
black disc on a flat grey field is clean to within a tolerance after one clone
|
||||
spot; the same edit at a one-quarter proxy and at full size land the disc in the
|
||||
same *normalised* place; adding spots does not grow `cached_pipelines()`.
|
||||
|
||||
**S5 — Heal.** The membrane, the mode switch. **Done** — §6.1, and the measurement came out at 0 levels of error against a clone's 38.
|
||||
*Acceptance:* a dark spot on a linear grey **gradient** — the case clone fails —
|
||||
is clean to within a tolerance, and the residual at the disc boundary is below
|
||||
the residual a clone leaves by an order of magnitude. This is the measurement
|
||||
that decides `K`.
|
||||
|
||||
**S6 — The tool.** `ViewMode::spots`, the chip, placement, handles, selection,
|
||||
the panel scope, deletion, history carrying the spot set, the stage-one source
|
||||
default. **Done, except the reveal view** — §9's second half is the one piece
|
||||
of S6 not built, and it is separable: it is a view mode over the detail chain
|
||||
rather than part of the tool.
|
||||
*Acceptance:* a dust mark on a real frame is gone in one click, the edit survives
|
||||
a restart, and undo takes it back.
|
||||
|
||||
*What was verified, and how.* Everything below the interface is under test —
|
||||
the model, the sidecar, the merge, the passes, both blend modes, and undo. The
|
||||
interface itself was compiled, laid out and photographed: the strip renders
|
||||
`Crop | Local | Repair` and the column re-scopes. It was **not** driven, because
|
||||
synthetic clicks do not reach this application (the compositor refuses them),
|
||||
so the gestures in §10 are as-written rather than as-felt. A first pass with a
|
||||
real pointer is the outstanding work on this stage.
|
||||
|
||||
**S7 — The rest of FR-DEV-8.** The scored source search (§8 stage two), and
|
||||
copying a spot set across a selection.
|
||||
|
||||
S1–S6 is the requirement met in the sense a photographer would recognise; S7 is
|
||||
the sentence in FR-DEV-8 about automatic placement met in the sense the document
|
||||
means it.
|
||||
|
||||
## 13. Documents to amend
|
||||
|
||||
- **requirements.md** — FR-DEV-8 has no *Acceptance:* line; every other
|
||||
requirement of its weight does. Proposed: *"a dust mark on a smooth sky is
|
||||
removed in one click with no visible boundary at 1:1, the spot survives a
|
||||
crop, a rotation and an export at another size, and the exported file matches
|
||||
the preview."*
|
||||
- **architecture.md §5.2** — the stage list is out of step with the `order:`
|
||||
keys in `ops/` and does not show spot removal first among the neighbourhood
|
||||
passes. §5.1 above is the correction.
|
||||
- **traceability.md** — regenerated, as ever, rather than edited. FR-DEV-8's row
|
||||
currently points only at two comments that mention it.
|
||||
|
||||
## 14. Open questions
|
||||
|
||||
1. ~~**`K = 24`?**~~ Settled by S5's measurement: 24 samples, inverse square weights, zero error on the case the mode exists for.
|
||||
2. **Does the reveal view belong to spot mode only,** or is it a view mode of
|
||||
its own that a photographer can turn on while doing something else? It is
|
||||
cheap to allow both; the risk is a mode nobody remembers turning on. Still
|
||||
open, and now the only part of §9 unbuilt.
|
||||
3. **Should a spot be clamped inside the crop?** A spot outside the current crop
|
||||
costs nothing to render and is invisible, and re-cropping should bring it
|
||||
back rather than find it deleted. Leaning strongly towards no clamp.
|
||||
4. **One radius, or an ellipse?** Lightroom's spot tool is circular and its
|
||||
users cope. An ellipse doubles the handle count for a case a second spot
|
||||
already covers.
|
||||
@@ -0,0 +1,605 @@
|
||||
# Storage backends
|
||||
|
||||
How DarkRoom talks to wherever a library lives, and what it takes to add
|
||||
somewhere new.
|
||||
|
||||
This document is the contract. `docs/architecture.md` §8 says why sync is built
|
||||
on capability negotiation rather than a common denominator; this says what the
|
||||
seam actually is, where each piece lives, and what a third connector has to do.
|
||||
|
||||
---
|
||||
|
||||
## 1. What "pluggable" has to mean
|
||||
|
||||
A trait alone does not make storage pluggable. `RemoteBackend` existed from the
|
||||
first release and every layer above it still knew it was talking to Nextcloud:
|
||||
seven files in `dr-ui` constructed a `NextcloudBackend` directly, ten functions
|
||||
took one by concrete type, the account model was a server URL beside a DAV user
|
||||
id, and the local cache directory was named after a hostname. The abstraction
|
||||
was real and bought nothing, because everything that *reached* a backend was
|
||||
still shaped like one product.
|
||||
|
||||
Pluggable means all four of these, not just the first:
|
||||
|
||||
1. **Operations** — what a backend can do. `RemoteBackend`.
|
||||
2. **Capabilities** — what it can do *cheaply*, so the engine adapts instead of
|
||||
assuming. `Capabilities`.
|
||||
3. **Configuration** — what an account is, with no server in it. `Account`.
|
||||
4. **Registration** — how the application discovers a connector at all, without
|
||||
naming it. `BackendProvider` + `BackendRegistry`.
|
||||
|
||||
Two connectors ship. Nextcloud is unchanged and keeps every one of its
|
||||
peculiarities — those are the point of the capability model, not an
|
||||
embarrassment it has to hide. The folder connector serves a plain directory and
|
||||
exists partly because it is genuinely useful and partly because a second
|
||||
implementation is the only way to find out whether the first was an
|
||||
abstraction.
|
||||
|
||||
---
|
||||
|
||||
## 2. Where each piece lives
|
||||
|
||||
```
|
||||
core/dr-sync/ the contract, and nothing that speaks a protocol
|
||||
├─ types.rs RemotePath, RemoteId, RemoteEntry, Validator, …
|
||||
├─ capability.rs Capabilities, ChangeDetection, ServerPreviews
|
||||
├─ error.rs RemoteError — the one error every caller handles
|
||||
├─ account.rs Account, AccountStore, Secret, Connection
|
||||
├─ provider.rs BackendProvider, BackendRegistry, SignIn
|
||||
├─ lib.rs RemoteBackend, SyncStrategy
|
||||
├─ scan.rs the walk, driven by capabilities
|
||||
├─ upload.rs where an original is placed
|
||||
└─ reachability.rs online/offline, inferred from observed results
|
||||
|
||||
core/dr-sync-nextcloud/ WebDAV, oc:fileid, chunked upload v2, Login Flow v2
|
||||
core/dr-sync-folder/ a directory on a filesystem
|
||||
|
||||
ui/dr-ui/src/remote.rs the registry — the ONLY file above dr-sync that
|
||||
names a connector
|
||||
```
|
||||
|
||||
`dr-sync` depends on no connector. That is deliberate and load-bearing: a build
|
||||
that only wants a folder library must not compile a TLS stack to get one, and
|
||||
the registry therefore lives in the crate that already depends on everything —
|
||||
the interface.
|
||||
|
||||
---
|
||||
|
||||
## 3. The four traits and types a connector meets
|
||||
|
||||
### 3.1 `RemoteBackend` — operations
|
||||
|
||||
```rust
|
||||
#[async_trait]
|
||||
pub trait RemoteBackend: Send + Sync {
|
||||
fn capabilities(&self) -> &Capabilities;
|
||||
fn name(&self) -> &str;
|
||||
|
||||
// discovery
|
||||
async fn list(&self, dir: &RemotePath, since: Option<&Validator>)
|
||||
-> Result<Vec<RemoteEntry>, RemoteError>;
|
||||
async fn dir_validator(&self, dir: &RemotePath) -> Result<Validator, RemoteError>;
|
||||
async fn delta(&self, cursor: &Cursor)
|
||||
-> Result<(Vec<RemoteChange>, Cursor), RemoteError>;
|
||||
|
||||
// transfer
|
||||
async fn get(&self, id: &RemoteId, range: Option<Range<u64>>)
|
||||
-> Result<Vec<u8>, RemoteError>;
|
||||
async fn put(&self, path: &RemotePath, body: Vec<u8>, precond: Option<Precondition>)
|
||||
-> Result<Validator, RemoteError>;
|
||||
async fn put_many(&self, items: Vec<(RemotePath, Vec<u8>)>) // defaulted
|
||||
-> Result<Vec<Result<Validator, RemoteError>>, RemoteError>;
|
||||
async fn delete(&self, id: &RemoteId, precond: Option<Precondition>)
|
||||
-> Result<(), RemoteError>;
|
||||
async fn move_to(&self, from: &RemoteId, to: &RemotePath) -> Result<(), RemoteError>;
|
||||
async fn create_dir(&self, path: &RemotePath) -> Result<(), RemoteError>;
|
||||
|
||||
// optional
|
||||
async fn thumbnail(&self, id: &RemoteId, size: u32) // defaulted to None
|
||||
-> Result<Option<Vec<u8>>, RemoteError>;
|
||||
}
|
||||
```
|
||||
|
||||
Rules that are not obvious from the signatures:
|
||||
|
||||
- **`dir_validator` and `delta` are capability-gated.** Return
|
||||
`RemoteError::Unsupported` unless your `ChangeDetection` is
|
||||
`PropagatingEtags` or `DeltaCursor` respectively. Answering
|
||||
`dir_validator` with something that does not actually propagate is worse than
|
||||
refusing: it lets a caller prune a subtree whose contents changed, and hides
|
||||
those changes for as long as the folder list holds still.
|
||||
- **`get` takes an optional range, and it is a hint.** A backend without cheap
|
||||
ranges may return the whole object; the caller slices. Correctness holds
|
||||
either way and `Capabilities::range_reads` says whether it was cheap.
|
||||
- **Chunked upload is not in the trait.** It is an implementation detail of
|
||||
`put`, chosen by body size. Exposing it would leak one server's protocol.
|
||||
- **`move_to` must preserve identity where the backend has stable ids.** This
|
||||
is what a soft delete uses (`FR-CAT-15`): a move implemented as copy + delete
|
||||
allocates a new id, orphaning the thumbnail shard and turning a restore into a
|
||||
full re-download.
|
||||
- **`create_dir` makes parents and succeeds if the directory exists.** Callers
|
||||
use it to guarantee a destination, not to claim they created one.
|
||||
|
||||
### 3.2 `Capabilities` — what is cheap
|
||||
|
||||
The engine reads these once at connect time and picks a `SyncStrategy`. See
|
||||
ARCH §8.1–8.2 for the tiers. The two that change behaviour rather than speed:
|
||||
|
||||
| Absent | Consequence the engine handles |
|
||||
|---|---|
|
||||
| `range_reads` | Embedded-preview extraction is impossible; browsing falls back to server previews or full download, and is refused on a metered connection |
|
||||
| `conditional_write` | Sidecar conflict detection falls back to revision counters inside the sidecar — narrows the race, does not close it. Reported as a reduced-safety mode |
|
||||
|
||||
**Declare what is true, not what is flattering.** A backend claiming
|
||||
`PropagatingEtags` it does not have does not merely run slowly; it silently
|
||||
hides changes.
|
||||
|
||||
### 3.3 `Account` — configuration with no server in it
|
||||
|
||||
```rust
|
||||
pub struct Account {
|
||||
pub backend: String, // BackendProvider::id; defaults to "nextcloud" on load
|
||||
pub endpoint: String, // stored as "server" — a URL, a path, a bucket
|
||||
pub login: String, // empty where the connector has no notion of a user
|
||||
pub user_id: String, // connector-defined sub-address; Nextcloud's DAV segment
|
||||
pub root: String, // the folder chosen as the library root
|
||||
pub formats: Vec<String>,
|
||||
pub last_scan: Option<i64>,
|
||||
}
|
||||
```
|
||||
|
||||
Everything but `backend` is the connector's to interpret. Code above `dr-sync`
|
||||
reads these for display and for cache keys, never for meaning.
|
||||
|
||||
Two properties are load-bearing:
|
||||
|
||||
- **The on-disk form is backwards compatible.** `backend` defaults to
|
||||
`"nextcloud"` and `endpoint` is stored under its historical key `server`, so
|
||||
every account written before there was a choice loads unchanged. A config the
|
||||
app refuses to parse is an account the user has to set up again.
|
||||
- **`Account::namespace()` is frozen for Nextcloud.** It names the directory
|
||||
holding the catalog, the thumbnail shards, the sidecar spool and the export
|
||||
outbox. Changing it does not lose that data, it *abandons* it — silently, as
|
||||
an upgrade — and costs a full rescan on top. The Nextcloud form is reproduced
|
||||
byte for byte from what `catalog_path` computed before; every other backend is
|
||||
prefixed by its connector id, and long endpoints are truncated with a hash
|
||||
tail so two deep paths cannot collide inside one filesystem's 255-byte
|
||||
component limit.
|
||||
|
||||
### 3.4 `Connection` and `Secret` — the credential split
|
||||
|
||||
```rust
|
||||
pub struct Connection { pub account: Account, pub secret: Option<Secret> }
|
||||
```
|
||||
|
||||
Credentials go to platform secure storage (`FR-NC-2`, `NFR-SEC-2`). Never the
|
||||
catalog, never the config file, never a log line. `AccountStore` writes the
|
||||
account as plain JSON and the secret to the keyring, which is what lets the app
|
||||
show "signed in as duncan, watching /PhotosRaw" before it has touched the
|
||||
keyring at all.
|
||||
|
||||
`Secret`'s inner string is reachable only through `expose()`, and its `Debug`
|
||||
prints `Secret(***)`. That closes the indirect leak — a `{:?}` on any struct
|
||||
that happens to hold a connection — by construction rather than by review.
|
||||
|
||||
`Connection` is also what replaced a pair of arguments (credentials, user id)
|
||||
threaded together through fifteen signatures in an order that could be swapped.
|
||||
|
||||
### 3.5 `BackendProvider` — registration
|
||||
|
||||
```rust
|
||||
pub trait BackendProvider: Send + Sync {
|
||||
fn id(&self) -> &'static str; // written to Account::backend
|
||||
fn display_name(&self) -> &'static str;
|
||||
fn endpoint_label(&self) -> &'static str; // "Server" / "Folder"
|
||||
fn endpoint_placeholder(&self) -> &'static str;
|
||||
fn sign_in(&self) -> SignIn;
|
||||
fn normalise_endpoint(&self, input: &str) -> Result<String, String>;
|
||||
fn account_for(&self, endpoint: &str) -> Result<Account, RemoteError>; // defaulted
|
||||
fn connect(&self, conn: &Connection) -> Result<Box<dyn RemoteBackend>, RemoteError>;
|
||||
}
|
||||
|
||||
pub enum SignIn {
|
||||
/// A handshake the user completes outside the app, yielding a credential.
|
||||
Browser,
|
||||
/// The endpoint is the whole account. No credential, no waiting state.
|
||||
EndpointOnly,
|
||||
}
|
||||
```
|
||||
|
||||
- **`id` is on-disk configuration.** Changing it after anyone has an account
|
||||
orphans that account. Pick it once.
|
||||
- **`normalise_endpoint` is where a bad endpoint is *rejected*,** before an
|
||||
account is written for a library that does not exist. Its error string is
|
||||
shown to the user, so it says what to fix rather than naming a type. The
|
||||
Nextcloud provider upgrades `http://` to `https://` here (`NFR-SEC-3`); the
|
||||
folder provider canonicalises the path, so two spellings of one directory do
|
||||
not become two accounts indexing the same photographs.
|
||||
- **`connect` is synchronous and cheap.** It validates configuration and builds
|
||||
a client; it does not talk to the remote. Workers call it per task.
|
||||
- **`SignIn` is a shape, not a method.** It would be tidier to expose
|
||||
`async fn sign_in()`, and wrong: Login Flow v2 is a browser handshake the user
|
||||
completes elsewhere while the app polls, so it is not one call, it does not
|
||||
finish on our schedule, and the screen has to render a URL and a waiting state
|
||||
in the middle of it. `SignIn` tells the launch screen which of the two shapes
|
||||
to draw; the flow stays where its protocol is.
|
||||
|
||||
**Credentials are deliberately not abstracted.** An app password, an OAuth
|
||||
token and a bucket key pair have no useful common shape, and inventing one
|
||||
before a third backend exists would produce a wrong answer confidently. The
|
||||
general form is `Connection` — an account plus an opaque secret — and each
|
||||
connector translates that into what its protocol needs
|
||||
(`NextcloudProvider::credentials`).
|
||||
|
||||
---
|
||||
|
||||
## 4. Adding a backend
|
||||
|
||||
1. **Implement `RemoteBackend`** over your protocol, in a new
|
||||
`core/dr-sync-<name>` crate depending on `dr-sync` and nothing else of ours.
|
||||
2. **Declare `Capabilities` honestly.** Start from `Capabilities::minimal()` and
|
||||
raise only what you can actually deliver.
|
||||
3. **Implement `BackendProvider`** beside it.
|
||||
4. **Register it** in `ui/dr-ui/src/remote.rs::registry()` and add the crate to
|
||||
`ui/dr-ui/Cargo.toml`.
|
||||
|
||||
That is the whole list. Nothing else in `dr-ui` changes, because nothing else in
|
||||
`dr-ui` names a connector.
|
||||
|
||||
**Two things to get right, because they are silent when wrong:**
|
||||
|
||||
- **Identity.** `RemoteEntry::id` should be `RemoteId::Stable(u64)` wherever you
|
||||
can produce a `u64` that names the same photograph on every device looking at
|
||||
the same library. The catalog keys the thumbnail shards and the face index on
|
||||
it (`catalog.md` §10.1), and an entry without one gets neither. Set
|
||||
`Capabilities::stable_ids` only if that id also survives a rename — the two
|
||||
are different questions and only the second is a capability.
|
||||
- **Path safety.** A `RemotePath` is built from names on the remote and from a
|
||||
catalog another device wrote. If you resolve one against a real filesystem,
|
||||
reject `..` before you open anything.
|
||||
|
||||
Register a test double the same way — `BackendRegistry::register` replaces an
|
||||
existing id rather than shadowing it — so an integration test can stand a fake
|
||||
server behind `"nextcloud"` without the registry knowing it happened.
|
||||
|
||||
---
|
||||
|
||||
## 5. The connectors that ship
|
||||
|
||||
### 5.1 Nextcloud (`dr-sync-nextcloud`, id `"nextcloud"`)
|
||||
|
||||
Unchanged by the abstraction, peculiarities intact — see ARCH §8.4 for the full
|
||||
mapping. What matters here is that none of them had to be given up to make room
|
||||
for a second backend:
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| `change_detection` | `PropagatingEtags` — the one-request no-op sync |
|
||||
| `stable_ids` | yes, `oc:fileid`, survives server-side rename and move |
|
||||
| `range_reads` | yes, detected by `206` vs `200`, never `HEAD` |
|
||||
| `chunked_upload` | v2, 5 MB – 5 GB, `MKCOL` → `PUT` chunks → `MOVE .file` |
|
||||
| `bulk_upload` | yes, `POST /remote.php/dav/bulk` |
|
||||
| `conditional_write` | yes, `If-Match` |
|
||||
| `server_previews` | `CommonFormatsOnly` — stock Nextcloud renders no RAW |
|
||||
| sign-in | `SignIn::Browser`, Login Flow v2, system browser, app password |
|
||||
|
||||
Also kept: the `oc:permissions` probe on a refused `PUT`, which is what
|
||||
distinguishes a create-only share from a bad credential; the `423 Locked`
|
||||
retry classification; and the bundled ISRG Root YE certificate.
|
||||
|
||||
### 5.2 Folder (`dr-sync-folder`, id `"folder"`)
|
||||
|
||||
A local disk, an NFS or SMB mount, an external drive, or the directory a
|
||||
Nextcloud desktop client already syncs. No server, no account, no credential —
|
||||
which makes it the route that works on a machine with no secrets daemon at all.
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| `change_detection` | `LocalEtags` — see below |
|
||||
| `stable_ids` | **no** — the id is a path hash and does not survive a rename |
|
||||
| `range_reads` | yes, `seek` + `take` |
|
||||
| `chunked_upload` | none; a write is a write |
|
||||
| `bulk_upload` | no |
|
||||
| `conditional_write` | yes, with a documented residual race |
|
||||
| `server_previews` | `None` |
|
||||
| sign-in | `SignIn::EndpointOnly` |
|
||||
|
||||
**Why `LocalEtags` and not `PropagatingEtags`.** A POSIX directory's mtime
|
||||
changes when its own entry list changes and at no other time — not when a
|
||||
child's contents are edited, and not for a grandchild. There is nothing to
|
||||
propagate, so `dir_validator` returns `Unsupported` and the engine walks the
|
||||
tree every scan. Which costs almost nothing, because the walk that was expensive
|
||||
was expensive for a reason this backend does not have.
|
||||
|
||||
**Measured 2026-08-28**, `cargo run -p dr-sync-folder --example scan`: a full
|
||||
uncached walk of 2,299 images across 233 directories completed in **137 ms**,
|
||||
and 380 images across 13 directories in **29 ms** — the same engine, the same
|
||||
`Depth: 1`-per-directory walk, with no pruning at all. The Nextcloud connector's
|
||||
comparable figure is 34.1 s for 17,185 RAWs across 334 directories *with*
|
||||
pruning available (ARCH §8.4). The capability model is what lets one engine
|
||||
drive both at the speed each actually runs at, instead of forcing the fast one
|
||||
down to the slow one's interface.
|
||||
|
||||
**Identity is a hash of the path relative to the library root**, FNV-1a 64
|
||||
(written out, because `DefaultHasher` is explicitly unstable between Rust
|
||||
releases and this value is written into the catalog). It gives the catalog a
|
||||
`u64` that names a photograph, is the same on every device looking at the same
|
||||
folder, and does not change when the file is edited. It does not survive a
|
||||
rename, and `stable_ids: false` says so: a moved photograph is seen as a delete
|
||||
and an add, and its thumbnail is derived again.
|
||||
|
||||
The alternative — keying on the inode — is stable across a rename but *differs
|
||||
between devices* and is reused by the filesystem after a delete. Two machines
|
||||
would disagree about which photograph a thumbnail belonged to, and a recycled
|
||||
inode would silently attach an old thumbnail to a new image. Re-deriving a
|
||||
thumbnail is a cost; showing the wrong one is a bug.
|
||||
|
||||
**Conditional writes.** `IfAbsent` is genuinely atomic (`O_CREAT | O_EXCL`).
|
||||
`IfMatch` is compare-then-swap: a `stat`, then a write to a temporary beside the
|
||||
destination and a `rename` over it. A POSIX filesystem has no compare-and-swap,
|
||||
so the race is narrowed to the microseconds between the two syscalls rather than
|
||||
closed — still far tighter than the fallback the engine uses for a backend that
|
||||
declares no conditional write at all, which spans a whole read-modify-write.
|
||||
The capability is declared, and the residual race is documented at the call
|
||||
site.
|
||||
|
||||
**Two deliberate divergences from WebDAV semantics:**
|
||||
|
||||
- **`delete` is not recursive.** A folder library is the user's own photographs
|
||||
on their own disk with no server-side trash behind it, so a caller that passed
|
||||
the wrong path would have no way back. Deleting a non-empty directory returns
|
||||
`RemoteError::Configuration`. Nothing in the engine deletes a directory — the
|
||||
soft delete is a `move_to` into the trash folder — so the guard is free.
|
||||
- **Every filesystem call runs on the blocking pool.** On a local disk that is
|
||||
overkill; on the NFS mount this backend is most useful over, a stalled server
|
||||
would otherwise wedge the async worker that made the call and every other
|
||||
request sharing it.
|
||||
|
||||
**Failure classification** matters as much as the operations. A vanished mount
|
||||
(`ESTALE`, `ENOTCONN`, `EIO`) maps to `RemoteError::Network`, which is what puts
|
||||
the app into offline mode and leaves the catalog readable — exactly as a dead
|
||||
server does. A permissions problem maps to `PermissionDenied` and does *not*,
|
||||
because going offline over one forbidden file would hide a fixable problem
|
||||
behind a network banner. An endpoint that is not a directory at all maps to
|
||||
`RemoteError::Configuration`: nothing was unreachable and no credential was
|
||||
wrong, so neither of the other two would send the user anywhere useful.
|
||||
|
||||
---
|
||||
|
||||
## 6. Virtual filesystems
|
||||
|
||||
A sync client in virtual-files mode leaves a **placeholder** where a file is
|
||||
catalogued but not downloaded. On Linux — the only mode it supports — that
|
||||
means `IMG.CR2` does not exist at all and `IMG.CR2.nextcloud` does, holding one
|
||||
byte. ARCH §9.0 measured a real machine: 121,785 placeholders against 10,267
|
||||
materialised files.
|
||||
|
||||
A folder library that ignores this is not merely degraded, it is dangerous.
|
||||
Before the handling below existed, the folder connector catalogued every stub
|
||||
as a 1-byte image, gave it an identity that changed the moment it was
|
||||
downloaded, and — worst — reported a dehydrated *sidecar* as absent, which made
|
||||
the sidecar writer create a fresh document over an existing one and discard
|
||||
every edit another device had put there.
|
||||
|
||||
### 6.1 Three questions, one trait
|
||||
|
||||
Everything else about a synced folder is an ordinary directory, so this is not
|
||||
a second connector. `dr_sync_folder::Vfs` asks only what differs:
|
||||
|
||||
```rust
|
||||
pub trait Vfs: Send + Sync {
|
||||
fn name(&self) -> &'static str;
|
||||
fn is_placeholder(&self, on_disk: &str) -> bool;
|
||||
fn real_name<'a>(&self, on_disk: &'a str) -> &'a str;
|
||||
fn placeholder_name(&self, name: &str) -> Cow<'_, str>;
|
||||
fn can_materialise(&self) -> bool; // defaulted false
|
||||
fn materialise(&self, local: &Path) -> Result<(), RemoteError>; // defaulted
|
||||
fn dematerialise(&self, local: &Path) -> Result<(), RemoteError>; // defaulted
|
||||
}
|
||||
```
|
||||
|
||||
`NoVfs` for a plain directory; `dr_sync_nextcloud::NextcloudVfs` for a synced
|
||||
one, wrapping the `DesktopClient` socket. A third convention is a third impl.
|
||||
|
||||
**Why not a `folder-vfs` provider.** The interesting capability is not a
|
||||
property of the backend: the same directory can materialise on demand while the
|
||||
client is running and cannot when it is down, so it must be computed per
|
||||
connection either way. Registering two providers would ask the user to choose
|
||||
between two things that differ by whether a background process is up. The
|
||||
convention is detected instead, per connection, by a hook the registry supplies
|
||||
(`FolderProvider::with_vfs_detector`) — which is what keeps `dr-sync-folder`
|
||||
free of any client's protocol.
|
||||
|
||||
### 6.2 What the backend reports
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| `RemoteEntry::path` | the photograph's name, never the stub's — so identity survives a download |
|
||||
| `RemoteEntry::materialised` | `false` on a stub; the catalog maps it to `Availability::Offline` |
|
||||
| `RemoteEntry::size` | `0` on a stub, meaning *unknown* — see below |
|
||||
| `get` on a stub | `RemoteError::NotMaterialised`, **never** `NotFound` and never the stub's one byte |
|
||||
| `put` over a stub, unconditional | **replaces it** — the whole file is being written, so there is nothing in the stub to keep, and the placeholder is removed after the content lands |
|
||||
| `put` over a stub, `IfMatch` | `NotMaterialised` — a stub's validator describes the placeholder, so nothing here can satisfy the guard; the caller fetches and retries |
|
||||
| `put` over a stub, `IfAbsent` | `PreconditionFailed` — the file *is* there, only its content is elsewhere |
|
||||
| `move_to` a stub | moves the stub and keeps it a stub — culling without downloading is ordinary |
|
||||
| `delete` a stub | deletes it; a photograph is deleted whether or not its bytes are here |
|
||||
| `capabilities().materialisation` | `OnDemand` with a client, `Placeholders` without, `Always` on a plain folder |
|
||||
|
||||
**Size is genuinely unknown.** A Linux suffix-mode stub is one byte and carries
|
||||
no record of what it stands for. The client's `._sync_*.db` has the real size,
|
||||
but that is a private schema and reading it would couple us to their migrations.
|
||||
FR-NC-6c wants a transfer size quoted before an operation starts; for a stub the
|
||||
honest answer is that it cannot be, and the interface should say so rather than
|
||||
report one byte or invent an estimate silently.
|
||||
|
||||
### 6.3 Hydration is a borrow
|
||||
|
||||
The rule: **a file is returned to the state it was found in.** What a pass
|
||||
downloaded is released; what the user already had is left alone. `BorrowPool`
|
||||
enforces it.
|
||||
|
||||
```rust
|
||||
let pool = BorrowPool::new();
|
||||
{
|
||||
let held = pool.borrow(&backend, &path).await?; // downloads only if absent
|
||||
// ... read it, thumbnail it, index its faces ...
|
||||
} // borrow ends
|
||||
let stats = pool.release_all(&backend).await; // dehydrates only what it hydrated
|
||||
```
|
||||
|
||||
Three properties that are not obvious:
|
||||
|
||||
- **Reference counted.** The thumbnail pass and the face pass meet on the same
|
||||
RAW. Without counting, the first to finish dehydrates the file the second is
|
||||
reading; with it, the transfer is paid once and released when the last
|
||||
borrower is done.
|
||||
- **Prior state is read before asking.** After `materialise` there is no way to
|
||||
tell what the pass brought from what was already there, so it is recorded
|
||||
first. Getting this wrong silently undoes a pin, and "my pinned trip
|
||||
evaporated after an indexing run" is the failure that would make people stop
|
||||
trusting the feature.
|
||||
- **Being unsure is not symmetric.** `borrow_known(.., Some(true))` keeps a file
|
||||
that might have been ours — costing disk. `Some(false)` releases one that
|
||||
might have been the user's. An uncertain caller passes `true` or `None`,
|
||||
never a guess at `false`.
|
||||
|
||||
A borrow against a plain folder or a server backend short-circuits and does
|
||||
nothing, so a pass written for a VFS library runs unchanged everywhere rather
|
||||
than growing two code paths.
|
||||
|
||||
**Measured 2026-08-29**, `cargo run -p dr-sync-folder --example vfs_cycle`: a
|
||||
library of 100 photographs at 25 MB each, 90 of them dehydrated and 10 the user
|
||||
keeps. A pass over all 100, borrowing and releasing as it goes:
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| on disk at rest | 250 MB |
|
||||
| **peak during the pass** | **275 MB** — the resting set plus one photograph |
|
||||
| without borrowing | 2,500 MB |
|
||||
| on disk afterwards | 250 MB |
|
||||
| of the 10 the user already had | 10 still there |
|
||||
|
||||
The peak is the working set, not the library, and the release is selective.
|
||||
|
||||
|
||||
### 6.4 Derived state is dehydrated too
|
||||
|
||||
Shards, the catalog snapshot and the place record live in `.darkroom-derived/`
|
||||
**inside the library folder**, so a sync client dehydrates them exactly as it
|
||||
dehydrates a photograph. Unlike a photograph, none of them can be skipped: a
|
||||
shard that will not open is a peer's thumbnails never merging, and a catalog
|
||||
snapshot that will not open is their collections.
|
||||
|
||||
The folder holds three kinds of thing:
|
||||
|
||||
| File | What it is | How two devices reconcile it |
|
||||
| --- | --- | --- |
|
||||
| `shard-<client>-NNNN.sqlite` | Thumbnail and face shards | Sealed and immutable; a name match is a content match, so existence is the whole protocol |
|
||||
| `catalog.sqlite` | The catalog snapshot, for its collections | Read-modify-write: **merge theirs, then push the union** |
|
||||
| `place.json` | Where the photographer was (FR-UI-8) | Replace: the newer timestamp wins outright |
|
||||
|
||||
`place.json` is the odd one out, and deliberately. Everything else there is
|
||||
*derived* — a faster way to learn something the device could have worked out for
|
||||
itself from the originals and the sidecars — so losing it costs time. A place is
|
||||
a fact only the other device knew, and losing it costs a scroll. That is why it
|
||||
is exchanged last in the pass, why its failures are logged rather than reported,
|
||||
and why it is the one file here that is replaced rather than merged: two devices
|
||||
cannot both be where the photographer is, so there is nothing of theirs inside
|
||||
ours to preserve.
|
||||
|
||||
It still refuses to upload over a copy it could not read, for a smaller version
|
||||
of the reason below: a record we have not compared against might be the newer
|
||||
one, and overwriting it would move the other device's photographer without ever
|
||||
having seen where they were.
|
||||
|
||||
`derived_sync::read_derived` fetches on demand rather than giving up. More
|
||||
important is what happens when it *cannot*:
|
||||
|
||||
The catalog sync is a read-modify-write over a file another device also writes.
|
||||
It was shaped `if let Ok(bytes) = backend.get(..)`, which folded every failure
|
||||
into "there is no remote catalog" and carried straight on to the upload — so a
|
||||
dehydrated snapshot meant pushing ours over theirs unmerged, taking their
|
||||
collections and members with it. The same shape as the sidecar bug in §6, and
|
||||
the same fix: a read that fails for any reason other than `NotFound` **stops the
|
||||
upload**.
|
||||
|
||||
That is why `NotFound` and `NotMaterialised` had to be separate errors. One
|
||||
means "yours is the whole truth, write it"; the other means "do not dare".
|
||||
|
||||
### 6.5 Release means dehydrate, never delete
|
||||
|
||||
The single most dangerous thing in this feature. A synced folder is not a
|
||||
cache: deleting a materialised file inside it propagates the deletion to the
|
||||
server and removes the photograph from every device the user owns. `Vfs` and
|
||||
`RemoteBackend::dematerialise` both say so, and an implementation that cannot
|
||||
dehydrate returns `Unsupported` rather than approximating it.
|
||||
|
||||
This is also why the originals cache (`dr_catalog::cache`) cannot simply be
|
||||
pointed at a VFS library: `Cache::release` deletes bytes, which is right for a
|
||||
copy under `originals/` and catastrophic in place.
|
||||
|
||||
### 6.6 Which photographs stay downloaded
|
||||
|
||||
The user's half of the bargain: a pass borrows for a moment, but *some* of the
|
||||
library should stay local — the trip you are about to take, the shoot you are
|
||||
working on.
|
||||
|
||||
That is a **pin**, and it is the pin the originals cache already had
|
||||
(`dr_catalog::cache`, FR-NC-6a). Nothing parallel was built, because the model
|
||||
was already the right one:
|
||||
|
||||
| Cache concept | On a placeholder library |
|
||||
|---|---|
|
||||
| `tier_desired` | what the user asked to keep hydrated |
|
||||
| `tier_actual` | what is actually materialised |
|
||||
| `pending_pins()` | the work list — what to hydrate next, resumable |
|
||||
| pinned rows are never evicted | a pinned collection is never dehydrated |
|
||||
| passive rows, LRU under a budget | what a pass borrowed, released when it finishes |
|
||||
|
||||
So "keep this collection hydrated" is `Cache::pin`, and the existing pin worker
|
||||
drives it — except that on a placeholder library it calls `materialise` instead
|
||||
of downloading a copy.
|
||||
|
||||
**Why not a copy.** The original materialises *in the library folder*. Copying
|
||||
it under `originals/` as well would hold every pinned photograph twice, and the
|
||||
copy would be the half the budget could evict while the real disk cost stayed.
|
||||
`Cache::record_in_place` records the bookkeeping with **`path = NULL`**, and
|
||||
that null is load-bearing: `release` deletes the file a row names, and a row
|
||||
that names none deletes nothing. The safety property is structural rather than
|
||||
remembered.
|
||||
|
||||
Unpinning therefore frees nothing by itself — the bytes are not ours to delete.
|
||||
`spawn_dehydrate` asks the client to take them back, which is what actually
|
||||
returns the disk.
|
||||
|
||||
---
|
||||
|
||||
## 7. What the abstraction does not yet cover
|
||||
|
||||
Stated so the next person does not have to rediscover it.
|
||||
|
||||
- **Multiple accounts at once.** `AccountStore` holds a list and the launch
|
||||
screen uses the most recent. Nothing in the model prevents two open libraries;
|
||||
the interface has no place to show them.
|
||||
- **Per-backend settings.** A connector has no way to contribute a settings
|
||||
page. Anything configurable is on the `Account` or is not configurable.
|
||||
- **Capability probing at runtime.** `Capabilities` is fixed at construction.
|
||||
Nextcloud's `server_previews` should really be probed per account — a server
|
||||
with `camerarawpreviews` installed can render RAW — and today it is assumed to
|
||||
be `CommonFormatsOnly`.
|
||||
- **A general notion of an account.** Credentials stay connector-specific on
|
||||
purpose (§3.5). A third connector with an OAuth flow will need a third `SignIn`
|
||||
variant, and that is the right place for it to appear.
|
||||
- **A quoted cost before a hydrating pass.** FR-NC-6c wants the transfer size
|
||||
stated before an operation that needs absent data. A placeholder reports no
|
||||
size (§6.2), so the honest figure for "index this library" is a count and not
|
||||
a byte total. The interface should say *n photographs, size unknown until
|
||||
fetched* rather than estimate one silently — and it does not say anything yet.
|
||||
- **Metadata-only placeholders.** Windows and macOS express these in filesystem
|
||||
metadata rather than in the name, and carry the real size there. `Vfs` asks
|
||||
its questions about a *name*, which is all the one convention this project has
|
||||
met needs. Supporting them means widening the trait to take a `Metadata`, and
|
||||
doing that before anyone has run this on those platforms would be guessing.
|
||||
- **Hydration during browsing, deliberately.** It stays forbidden (ARCH §9.0
|
||||
finding 3). A grid cell whose content is absent shows as not-downloaded; only
|
||||
a pass the user asked for may fetch.
|
||||
@@ -0,0 +1,451 @@
|
||||
# DarkRoom — Technical debt
|
||||
|
||||
**Status:** Living document · first written 2026-08-26
|
||||
**Companion to:** [architecture.md](architecture.md)
|
||||
|
||||
Deliberate compromises: things the code does knowing they are wrong, because the alternative was
|
||||
worse at the time. Each entry says what the debt is, what it cost to take on, what it would take to
|
||||
pay off, and how you would know it had been paid.
|
||||
|
||||
Not a bug list. A bug is something nobody chose. Everything here was chosen, and the point of
|
||||
writing it down is that the reasoning outlives whoever chose it — so the next person can tell a
|
||||
constraint from an accident, and does not "fix" something load-bearing or preserve something that
|
||||
has quietly stopped being necessary.
|
||||
|
||||
---
|
||||
|
||||
## TD-1 — The Android develop view reads pixels back through the CPU
|
||||
|
||||
**Breaks:** [architecture.md §12 / 6.1](architecture.md) — GPU results never round-trip through the
|
||||
CPU — and AC-8, on Android only. Desktop is unaffected and keeps the zero-copy path.
|
||||
|
||||
### What it does
|
||||
|
||||
`DevelopSession::render` on Android runs the compute passes on the GPU as usual, then calls
|
||||
`AdjustPass::export_pixels` and hands the frame to Slint as a `SharedPixelBuffer`. That is exactly
|
||||
the GPU→CPU→GPU transfer §6.1 exists to forbid, and it is on the frame path.
|
||||
|
||||
### Why
|
||||
|
||||
Zero-copy needs Slint to draw with wgpu. On Android that means wgpu's Vulkan swapchain, which
|
||||
hardcodes `preTransform = VK_SURFACE_TRANSFORM_IDENTITY_BIT_KHR`
|
||||
([gfx-rs/wgpu#3345](https://github.com/gfx-rs/wgpu/issues/3345)) — wgpu-hal says so in a comment
|
||||
beside the line.
|
||||
|
||||
On a tablet whose panel is mounted landscape, a portrait window then hands Android an unrotated
|
||||
buffer, every present returns `VK_SUBOPTIMAL_KHR`, and frames arrive torn. Measured on the device,
|
||||
same build, only the tablet rotated:
|
||||
|
||||
| orientation | `bufferTransform` | composition | result |
|
||||
|---|---|---|---|
|
||||
| landscape | `ROT_180` | `DEVICE (2)` | clean |
|
||||
| portrait | `ROT_270` | `CLIENT (1)` | torn |
|
||||
|
||||
Setting `preTransform` is not a fix available to us: the field is a *promise* that the content is
|
||||
already rotated, so honouring it needs the renderer to rotate what it draws, which wgpu cannot do
|
||||
on Skia's behalf.
|
||||
|
||||
So the choice was never fast-develop against slow-develop. It was a develop view that costs a
|
||||
readback against a grid that tears in the orientation a tablet is mostly held in.
|
||||
|
||||
### What it costs
|
||||
|
||||
Less than §6.1's headline numbers, because `render` fits the pass to the canvas before it runs — the
|
||||
readback is at viewport resolution, not sensor resolution. The 7.43 ms at 4K in §12 is the ceiling,
|
||||
not the bill. **It has not been measured on the device**, which is the first thing to do if the
|
||||
develop view feels heavy on the tablet; do not assume this is the cause without a number.
|
||||
|
||||
### And a second transfer, while focus peaking is on
|
||||
|
||||
Added 2026-08-29 with FR-CULL-3. The focus-peaking overlay is a compute pass writing its own
|
||||
`Rgba8Unorm` texture, which on desktop reaches the compositor with no copy — but on Android there is
|
||||
no more a path for *that* texture than for the frame it belongs to, and an overlay that stayed on
|
||||
the device while the picture underneath it did not would simply never be seen. So
|
||||
`FocusPeakPass::read_overlay` follows the frame back through memory, and the Android frame path
|
||||
carries **two** full-resolution `copy_texture_to_buffer` transfers instead of one.
|
||||
|
||||
This is recorded under TD-1 rather than as its own entry because it is not an independent choice.
|
||||
It exists only because TD-1 exists, it is bounded by the same thing — `render` fits the pass to the
|
||||
canvas, so both transfers are at viewport resolution — and TD-1's "Done when" already covers it:
|
||||
whichever of the three fixes above lands removes the readback for the frame and the overlay
|
||||
together, because both are the same missing capability.
|
||||
|
||||
Two things worth saying plainly. The doubling is **reasoned, not measured on the device** — the same
|
||||
gap TD-1 admits about its own cost, and the reason neither number should be quoted as a measurement.
|
||||
And it is paid only while the photographer has the overlay switched on: `DevelopSession::focus_overlay`
|
||||
returns on its first line when peaking is off, so with it off there is no dispatch and no transfer,
|
||||
and the Android frame path is exactly what it was before this feature existed.
|
||||
|
||||
### Paying it off
|
||||
|
||||
Any one of these removes it:
|
||||
|
||||
- wgpu implements pre-rotation (#3345), and Android goes back on `unstable-wgpu-29`.
|
||||
- Slint's Skia Vulkan surface handles `preTransform` and Android uses that instead of OpenGL.
|
||||
- Skia over OpenGL grows a way to sample an external texture that wgpu can write.
|
||||
|
||||
**Done when:** `ui/dr-ui/Cargo.toml` no longer scopes `renderer-femtovg-wgpu` and
|
||||
`unstable-wgpu-29` to non-Android, the `#[cfg(target_os = "android")]` arm of
|
||||
`DevelopSession::render` is gone, and the tablet is clean in portrait.
|
||||
|
||||
---
|
||||
|
||||
## TD-2 — Thumbnails are fetched one at a time
|
||||
|
||||
**Where:** `library::spawn_thumbnails` — the `for req in to_fetch` loop.
|
||||
|
||||
### What it does
|
||||
|
||||
The interactive thumbnail batch fetches serially: one image at a time, and two HTTP round trips
|
||||
each (a header read, then the preview's byte range). A window of a few hundred cells is that many
|
||||
sequential round trips against the server.
|
||||
|
||||
### Why it is debt rather than a bug
|
||||
|
||||
It is correct, and it was fast enough when a window was one screenful. It is the *ordering* that
|
||||
kept it survivable: since `fetch_rank`, on-screen cells are requested first, so the cells a person
|
||||
is looking at arrive first even though the queue as a whole is slow.
|
||||
|
||||
Portrait makes it worse by construction — a narrow window means smaller cells, more rows, and two
|
||||
to three times as many cells on screen at once, all of them ahead of the ones below in a queue that
|
||||
never runs more than one request.
|
||||
|
||||
### Paying it off
|
||||
|
||||
`spawn_thumbnail_sweep` already has the pattern: `SWEEP_LANES` disjoint lanes over a chunk, joined,
|
||||
with the store written on the one thread that owns it. Striping a *priority-ordered* chunk across
|
||||
lanes keeps `fetch_rank`'s ordering while running several requests at once.
|
||||
|
||||
Not done yet because it multiplies concurrent requests against the user's Nextcloud during a
|
||||
scroll, and that is a behaviour change worth deciding on deliberately rather than inheriting from a
|
||||
performance fix.
|
||||
|
||||
**Done when:** the interactive batch runs on more than one lane, priority order is preserved
|
||||
across the lanes, and a slow server still cannot stall the visible cells behind offscreen ones.
|
||||
|
||||
---
|
||||
|
||||
## TD-3 — The thumbnail drain applies an unbounded batch on the UI thread
|
||||
|
||||
**Where:** `library_ui::drain_thumbnails` — the `loop` inside the timer callback.
|
||||
|
||||
### What it does
|
||||
|
||||
Every message queued when the timer fires is applied in that one callback, with no ceiling. On a
|
||||
library whose thumbnails are already in the store, the worker delivers a whole window at once, so a
|
||||
single callback can do hundreds of `to_slint_image` calls back to back — each an allocation and a
|
||||
full RGBA copy — while the grid is mid-flick.
|
||||
|
||||
The copy cannot move off the UI thread: `slint::SharedPixelBuffer` is not `Send`, so decoded bytes
|
||||
can only become an `Image` on the thread that draws. Only the *amount done per wake* is ours to
|
||||
choose, and right now it is "all of it".
|
||||
|
||||
### Cost
|
||||
|
||||
Measured with a temporary probe, **debug build**, so treat the shape rather than the size:
|
||||
|
||||
| class | per thumbnail | × a 280-cell window |
|
||||
|---|---|---|
|
||||
| grid, 256 px | 1.93 ms | 539 ms |
|
||||
| large, 512 px | 7.78 ms | 2.18 s |
|
||||
|
||||
A release measurement was started and never completed — do not quote these as release figures.
|
||||
|
||||
### Paying it off
|
||||
|
||||
A time budget per wake and a shorter interval: apply for a few milliseconds, return without
|
||||
stopping the timer, and finish on the next tick. A batch then lands in frame-sized slices rather
|
||||
than one lump between two frames. Draft written and discarded during the investigation; it is a
|
||||
small change.
|
||||
|
||||
**Done when:** one wake of the drain cannot exceed a frame, and a fully-cached window still fills
|
||||
in well under a second.
|
||||
|
||||
---
|
||||
|
||||
## TD-4 — The local-contrast base is computed at full render resolution ✅ PAID OFF
|
||||
|
||||
**Where:** `dr_pipeline::ops::local_contrast::LocalContrast::passes` — the `base` and `combine`
|
||||
passes, and the stage that dispatches them, `dr_pipeline::detail`.
|
||||
|
||||
Breaks **FR-DSP-3** at large viewports. Measured, and the numbers are in
|
||||
[frame-budget.md](frame-budget.md) §M3.
|
||||
|
||||
### What it does
|
||||
|
||||
Clarity's Gaussian σ is 1.2% of the frame's shorter edge, truncated at 2σ, so its kernel radius is
|
||||
a property of the *viewport*: 29 px at 1920 × 1200, 38 px at 2560 × 1600, **52 px at 4K**. The two
|
||||
separable passes therefore run 105 taps each over 8.3 M pixels at 4K, which is 1.7 billion texture
|
||||
reads for one control.
|
||||
|
||||
| viewport | radius | clarity alone, p99 |
|
||||
|---|---:|---:|
|
||||
| 1920 × 1200 | 29 | 5.99 ms |
|
||||
| 2560 × 1600 | 38 | 12.44 ms |
|
||||
| 3840 × 2160 | 52 | **33.89 ms** |
|
||||
|
||||
RTX 3050 laptop, `examples/frame_budget`, fused dispatch reused so this is the convolutions alone.
|
||||
Clarity is 97% of the cost of all four neighbourhood operations together at every size.
|
||||
|
||||
For scale: the entire fused chain — every point operation active, film stock included — costs
|
||||
4.5 ms at the same 4K viewport. **A single slider is seven times the rest of the pipeline.**
|
||||
|
||||
### Why
|
||||
|
||||
Because the stage cannot do otherwise yet. `dr_pipeline::detail` dispatches every pass at the
|
||||
render size; there is no way to express "read this target and write a smaller one". The module's own
|
||||
documentation has said so since it was written:
|
||||
|
||||
> The right optimisation is a base computed at reduced resolution, which needs a detail stage that
|
||||
> can write a smaller target than it reads; that is a change to `crate::detail`, not to this file.
|
||||
|
||||
It was the right call to ship the correct answer slowly rather than a fast approximation nobody had
|
||||
checked — the halo behaviour is the hard part of this operation and it is tested.
|
||||
|
||||
### Not a tiling problem
|
||||
|
||||
Worth saying because ARCH §5.3 offers a tile cache and this is the stage that looks like it wants
|
||||
one. It does not: a tiled convolution reads a halo per tile, so at a 52-pixel radius, 256-pixel
|
||||
tiles would read (256 + 104)² instead of 256² — very nearly **twice** the taps.
|
||||
[display-and-extension.md](display-and-extension.md) §2's decision rule was resolved on this
|
||||
evidence; see [frame-budget.md](frame-budget.md).
|
||||
|
||||
### Paying it off
|
||||
|
||||
A detail pass that declares an output scale, so the base can be computed at a quarter resolution and
|
||||
sampled back up in `combine`. A quarter-resolution base is 1/16 the pixels at 1/4 the radius —
|
||||
about **1/64 of the work** — and is visually identical, because a base at σ = 26 px holds no content
|
||||
above the quarter-resolution Nyquist to lose. Texture's σ is a decade finer and must stay at full
|
||||
resolution; the scale therefore belongs on the `DetailPass`, not on the stage.
|
||||
|
||||
**Done when:** clarity at 100% is inside the frame budget at 3840 × 2160, the halo tests in
|
||||
`tests/local_contrast.rs` still pass unchanged, and `examples/frame_budget`'s M3 table in
|
||||
[frame-budget.md](frame-budget.md) has been rerun and committed.
|
||||
|
||||
### Paid off
|
||||
|
||||
A `DetailPass` now declares `output_scale`, and clarity's base is computed on a grid a quarter the
|
||||
size on each axis. Measured before and after on the same machine, same adapter, same build profile,
|
||||
with only the change between them — see [frame-budget.md](frame-budget.md) §"The reduced base,
|
||||
measured":
|
||||
|
||||
| viewport, fit | before p50 | after p50 | | before p99 | after p99 |
|
||||
|---|---:|---:|---:|---:|---:|
|
||||
| 1920 × 1200 | 2.30 ms | 1.56 ms | 1.5× | 3.94 ms | 1.97 ms |
|
||||
| 2560 × 1600 | 4.41 ms | 1.95 ms | 2.3× | 10.40 ms | 2.37 ms |
|
||||
| 3840 × 2160 | **10.94 ms** | **3.88 ms** | **2.8×** | 25.05 ms | 4.17 ms |
|
||||
|
||||
**Quote the p50 column.** The baseline run's p99 figures are contaminated — its `fit` rows spread
|
||||
2.3× between median and 99th percentile where the after run spreads 1.1×, and
|
||||
[frame-budget.md](frame-budget.md)'s own independent measurement of the same baseline on the same
|
||||
card reports 4.67 ms p99 at 2560 × 1600 against the 10.40 ms here. The p99 improvement is real and
|
||||
larger than 2.8×; this run cannot say by how much.
|
||||
|
||||
Clarity is no longer the stage that misses the budget, and no longer dominates the neighbourhood
|
||||
stage: at 4K it is 4.17 ms against 4.61 ms for all four neighbourhood operations together, where it
|
||||
was 97% of that total at every size. That comparison is within one run, so the contention does not
|
||||
touch it.
|
||||
|
||||
**Two things worth recording, because neither is visible in the table.**
|
||||
|
||||
The declared halo is now quantised to multiples of `output_scale`. The kernel truncates at 2σ and
|
||||
that rounding now happens on the reduced grid, so 1920 × 1200 reports 28 render pixels where it
|
||||
reported 29, and 2560 × 1600 reports 40 where it reported 38. At 2σ the Gaussian is already down to
|
||||
`e⁻²` of its peak, and the cross-form test holds the difference to 0.03 stops of peak excursion and
|
||||
2% of frame reach — but it is a real change in reach, not a pure speed-up, and a tile scheduler
|
||||
would see it.
|
||||
|
||||
The measurement was taken on an AMD RX 5700 XT, not the RTX 3050 the M1/M2/M3 tables above were
|
||||
measured on, so the *absolute* figures are not comparable with those. The before/after is, because
|
||||
both halves of it were measured on the same card minutes apart.
|
||||
|
||||
---
|
||||
|
||||
## TD-5 — The fused shader is reassembled from strings on every frame
|
||||
|
||||
**Where:** `dr_pipeline::operation::compose_full`, called from `DevelopSession::render`.
|
||||
|
||||
### What it does
|
||||
|
||||
`EditGraph::compose` walks the active operations and formats a WGSL source string, per frame, on
|
||||
the UI thread. On a full chain that is **2.8–5.2 ms** — at 1920 × 1200 it is larger than the entire
|
||||
fused dispatch it precedes, and on the chains that also carry a detail stage it is a third of what
|
||||
is left of the 16 ms budget after the GPU has taken its share. It does not vary with resolution,
|
||||
because it is not pixel work.
|
||||
|
||||
### Why
|
||||
|
||||
Because it was free until the chain got long. Composition was written when an edit was two or three
|
||||
operations, and the cost is roughly linear in generated source: the colour mixer emits twelve hue bands
|
||||
and the tone curve emits a spline evaluator, so a full chain is a large string built from scratch
|
||||
sixty times a second.
|
||||
|
||||
### Paying it off
|
||||
|
||||
The generated source depends only on the *structure* of the graph — which is precisely what
|
||||
`ComposedShader::structure_hash` already identifies, and precisely what does not change while a
|
||||
slider is being dragged. `AdjustPass` relies on that already: it caches compiled pipelines against
|
||||
that hash and does not recompile during a drag. Caching the source string against the same hash and
|
||||
rebuilding only the uniforms — a handful of floats per operation — takes this to approximately
|
||||
nothing on the path that needs it most.
|
||||
|
||||
The care needed is in what the hash covers. It deliberately excludes parameter *magnitudes*, so a
|
||||
cache keyed on it is sound for the source and would be wrong for anything else in `ComposedShader`.
|
||||
|
||||
**Done when:** the `shader` column of [frame-budget.md](frame-budget.md)'s M1 table is under a
|
||||
millisecond for the `point` and `all` chains, and the codegen tests still pass byte for byte.
|
||||
|
||||
---
|
||||
|
||||
## TD-6 — The quietest ink does not reach WCAG AA, and the rule does not reach 3:1
|
||||
|
||||
**Breaks:** [requirements.md](requirements.md) NFR-A11Y-2 — "non-canvas UI meets WCAG AA contrast".
|
||||
|
||||
### What it does
|
||||
|
||||
`style.yaml` sets three inks and four surfaces. Measured as WCAG 2 contrast ratios (sRGB relative
|
||||
luminance, the standard formula), against the surfaces each ink is actually drawn on:
|
||||
|
||||
| ink | on `ground` | on `surface` | on `surface-raised` | on `hover` | on `selected` |
|
||||
|---|---|---|---|---|---|
|
||||
| `ink` #EDEEF0 | 16.02 | 14.69 | 13.03 | 11.39 | 9.68 |
|
||||
| `ink-dim` #9EA1A6 | 7.18 | 6.58 | 5.84 | 5.10 | **4.34** |
|
||||
| `ink-faint` #71747A | **3.97** | **3.64** | **3.23** | **2.82** | **2.40** |
|
||||
| `warn-ink` #C9A05A | 7.67 | 7.03 | 6.24 | 5.45 | 4.64 |
|
||||
| `rule` #323438 | **1.49** | **1.37** | **1.21** | **1.06** | **1.11** |
|
||||
|
||||
Every text size in the application is 11px, 13px, 17px or 24px, and WCAG's "large text" relief
|
||||
begins at 18.66px bold or 24px regular — so all four of those thresholds are the normal-text one,
|
||||
**4.5:1**, except the masthead. Bold entries fail it.
|
||||
|
||||
The inverted cases pass and are worth stating so nobody re-measures them: `ground` on `active`
|
||||
(#FFFFFF) is 18.60, on `active-dim` 11.20, on `active-pressed` 5.95, on `selected-ring` 13.02. The
|
||||
near-white fills that `Button.primary`, `FilterChip.active` and the held tool-rail entry use are
|
||||
the *best*-contrasting text in the interface, not the worst.
|
||||
|
||||
So the failures are exactly two, and neither is where one would guess:
|
||||
|
||||
- **`ink-faint` reaches 4.5:1 nowhere at all.** It is the ink for `Caption`, `PanelHeading`,
|
||||
`Disclosure`, `Value`'s placeholder state and `FilterChip`'s count — every hint, every section
|
||||
name, every "3 photographs" under a title.
|
||||
- **`rule` reaches 3:1 nowhere.** WCAG 1.4.11 asks 3:1 of the boundary of a control the user must
|
||||
perceive, and `rule` is the border of every `Button`, `Field`, `Panel`, `ChoiceChip` and
|
||||
`IconButton`. An unfilled secondary button is a 1.4:1 outline on a 1.2:1 background.
|
||||
|
||||
`ink-dim` on `selected` at 4.34 is a third case, marginal enough that a two-point lift fixes it.
|
||||
|
||||
### Why
|
||||
|
||||
Not an oversight — the direct consequence of the palette's own argument, which `style.yaml`'s
|
||||
preamble makes at length and correctly. The chrome is deliberately quiet because a bright surround
|
||||
biases how a photograph is judged, and hue is banned outright because an accent beside the image
|
||||
shifts the perception of nearby colours. What is left to signal with is luminance, and the palette
|
||||
spends its luminance range on the *photograph*, keeping the chrome inside a narrow band above the
|
||||
ground.
|
||||
|
||||
A narrow band is precisely what a contrast ratio measures. `ink-faint` exists to be skipped by the
|
||||
reader who did not stop to look; that is a real design intent, and "text you are meant to skip"
|
||||
and "text everyone can read" are in genuine tension rather than one being a mistake.
|
||||
|
||||
### What it costs
|
||||
|
||||
The photographer who cannot read a caption cannot read *any* caption, on any screen — this is one
|
||||
token, so it fails everywhere at once. The hints under the settings switches say what a setting
|
||||
costs, the section names say what a panel is, and the counts say how big a filter is. None of it is
|
||||
decorative.
|
||||
|
||||
### Paying it off
|
||||
|
||||
Two token changes, and the second is the awkward one.
|
||||
|
||||
`ink-faint` needs roughly #8A8D93 to clear 4.5:1 against `surface-raised`, the darkest surface it
|
||||
is drawn on that matters — which puts it about where `ink-dim` sits today and collapses the
|
||||
three-ink scale to two. So the real fix is to re-derive all three inks against the surfaces rather
|
||||
than to nudge one: the scale wants to start higher and keep its steps, not compress.
|
||||
|
||||
`rule` needs about #4A4D52 for 3:1 against `surface`. That is a visibly stronger line, and the
|
||||
preamble's "instrument rather than absence" reasoning applies to it as much as to the greys — this
|
||||
is a look change, not a number change, and it should be looked at rather than computed.
|
||||
|
||||
Both are decisions about how the application appears next to a photograph, which is the one thing
|
||||
this palette was designed around. They want a screenshot and an opinion, not a patch.
|
||||
|
||||
**Done when:** every `Theme` ink reaches 4.5:1 against every surface it is drawn on, `rule` reaches
|
||||
3:1 against `surface` and `surface-raised`, and a test recomputes those ratios from `style.yaml` so
|
||||
the next palette edit cannot quietly undo it. The table above is the baseline to compare against.
|
||||
|
||||
### Not in scope
|
||||
|
||||
The histogram's `plot-*` inks (2.36 for `plot-luma` on `ground`) are drawn *on* the canvas, and
|
||||
NFR-A11Y-2 scopes contrast to non-canvas UI. NFR-A11Y-3 covers what those need instead, and is
|
||||
already met — the readouts name the channel in words.
|
||||
|
||||
---
|
||||
|
||||
## TD-7 — Platform font scaling is not honoured
|
||||
|
||||
**Breaks:** [requirements.md](requirements.md) NFR-A11Y-2 — "platform font scaling is honoured
|
||||
without clipping".
|
||||
|
||||
### What it does
|
||||
|
||||
Nothing at all, which is the entry. Every type size is a constant in `style.yaml` — 11, 13, 17, 24
|
||||
— read as `Theme.text-sm` and friends at 65 call sites, and there is no multiplier anywhere between
|
||||
the platform's font-size preference and those numbers. `scale_factor()` is read in `display_ui.rs`
|
||||
and in `lib.rs`, but only to size the canvas in physical pixels for the render; it is display DPI,
|
||||
which Slint already applies to logical lengths, and it is not the user's text-size setting. A
|
||||
photographer who sets 130% text on GNOME or Android gets an application that ignores it.
|
||||
|
||||
### Why
|
||||
|
||||
Because honouring it is not a multiplier, and pretending it is would be worse than not doing it.
|
||||
|
||||
The layout is built on constants that are not derived from the type size: `control-height` 28,
|
||||
`touch-target` 44, `row-height` 26, `rail-entry-height` 54, `panel-width` 360, and a dozen fixed
|
||||
heights written at their call sites — `ParamSlider`'s 46px, the folder picker's 220px box, the
|
||||
readout column's 30px. Scaling the type alone clips against every one of them, silently, because
|
||||
Slint elides rather than errors. `SwatchSlider` is the sharpest case: 12 hue bands × 3 channels in
|
||||
a 360px column, sized so that a track and a swatch and a three-character readout fit on one line.
|
||||
|
||||
And the failure is invisible to the person shipping it. The requirement's own phrase is "without
|
||||
clipping", and clipping is exactly what a screenshot at 100% cannot show — the memory note on
|
||||
verifying Slint changes exists because these files have a history of compiling, rendering and being
|
||||
wrong.
|
||||
|
||||
### What it costs
|
||||
|
||||
The user for whom this matters most is not the screen-reader user the rest of this branch serves —
|
||||
it is the one with usable but poor sight, who reads the interface and needs it larger. The
|
||||
application is unusable to them at any setting, and there is no partial credit: text scaling is a
|
||||
system-wide preference, so an app that ignores it is the one thing on the desktop that did.
|
||||
|
||||
### Paying it off
|
||||
|
||||
In the order the pieces depend on each other:
|
||||
|
||||
1. **A scale token.** `Theme.text-sm` and the rest become `base × Theme.type-scale`, with the
|
||||
scale an `in-out` property Rust writes at startup from the platform. `build.rs` already emits
|
||||
`in-out` tokens under `live-style`, so the codegen half of this exists and is proven — that
|
||||
feature is the mechanism, one line from being general.
|
||||
2. **A source for the number.** GNOME publishes `text-scaling-factor` over the settings portal;
|
||||
Android has `Configuration.fontScale` through JNI, beside the calls `lib.rs` already makes for
|
||||
`ACTION_VIEW`. Both want a default of 1.0 and a sane clamp — 0.8 to 2.0 — because a user who has
|
||||
set 300% for a phone launcher has not asked for a 72px slider readout.
|
||||
3. **The constants that are not type.** Every fixed height a *label* sits inside has to follow the
|
||||
scale; every touch target must not shrink and need not grow. That is the work, and it is where
|
||||
the 46px and 220px literals get read one at a time.
|
||||
4. **Evidence.** Screenshots at 1.0, 1.3 and 2.0 of the develop column, the settings page and the
|
||||
colour mixer — the three densest layouts — because "without clipping" is a claim about the
|
||||
worst case and nothing else will show it.
|
||||
|
||||
**Done when:** the develop column, the settings page and the colour mixer render at a 2.0 scale
|
||||
with no elided label and no touch target under 44 logical pixels.
|
||||
|
||||
---
|
||||
|
||||
## Related, and deliberately not here
|
||||
|
||||
The window-move rule, the grid's ordering index and the whole-library readout cache were *fixed*
|
||||
rather than deferred — see the commits around `6d6ef8d`. They are mentioned only so that a reader
|
||||
looking for "why was the grid slow" finds the answer in the code and its comments rather than
|
||||
assuming it is still outstanding.
|
||||
File diff suppressed because one or more lines are too long
@@ -0,0 +1,827 @@
|
||||
# UI navigation: finding things once there are many
|
||||
|
||||
TRACES: FR-UI-1 | FR-UI-3 | FR-UI-5 | FR-DEV-3a | FR-DEV-3c
|
||||
|
||||
Successor to [`ui-refinement.md`](archive/ui-refinement.md), which asked how the interface should *look*.
|
||||
This asks how someone finds anything in it. The two are sequenced together at
|
||||
the end.
|
||||
|
||||
## Why now
|
||||
|
||||
Local adjustments landed (D14, `segmentation.md` §14) and the develop column
|
||||
went from four panels to six: image, histogram, geometry, settings, local,
|
||||
adjust. Adjust alone is ten operations, and the colour mixer contributes
|
||||
thirty-six parameters by itself. That is already past what one scrolling
|
||||
column presents well, and the operation set is meant to keep growing —
|
||||
FR-DEV-3 lists texture, clarity, sharpening and noise reduction as v1, none of
|
||||
which exist yet.
|
||||
|
||||
But the count is the lesser problem. **Local adjustments introduced a mode
|
||||
without introducing a way to see it**, and that is the part that can lose
|
||||
someone's work rather than merely slow them down.
|
||||
|
||||
---
|
||||
|
||||
## 1. The three problems, which are not one problem
|
||||
|
||||
### 1.1 Scope — invisible state
|
||||
|
||||
Selecting a mask layer silently re-points the adjust panel at that layer's
|
||||
chain. Same thirty sliders, different meaning, and the only indication is a
|
||||
caption between the two panels.
|
||||
|
||||
Three ways that bites:
|
||||
|
||||
- An exposure change lands on the whole photograph when it was meant for a
|
||||
face, or the reverse. Both are silent; both are discovered later.
|
||||
- **The histogram does not follow scope.** The instrument the tonal controls
|
||||
are judged against reports the whole frame while the slider edits a
|
||||
subject's face. FR-DSP-7 asks the histogram to describe what the
|
||||
photographer is looking at; under a mask it currently does not.
|
||||
- Undo interleaves global and local edits with nothing distinguishing them.
|
||||
|
||||
This is the classic modal fault, and the classic remedy applies: make the mode
|
||||
visible, or make it not a mode. See §3.
|
||||
|
||||
### 1.2 Extent — six panels in 280px
|
||||
|
||||
The column scrolls as one (`ui-refinement.md` Workstream C explains why: a
|
||||
scroller inside a scroller gives every drag a third thing to be lost to). At
|
||||
six panels that scroll is long enough that the histogram — the instrument
|
||||
everything tonal is judged against — is frequently off screen while the
|
||||
sliders it reports on are being dragged.
|
||||
|
||||
### 1.3 View — grid and develop are still separate screens
|
||||
|
||||
Diagnosed as `ui-refinement.md` Workstream F and not yet built. Unchanged by
|
||||
this document, which assumes F lands.
|
||||
|
||||
---
|
||||
|
||||
## 2. What other programs do
|
||||
|
||||
Worth summarising honestly, because all three problems are solved elsewhere
|
||||
and the solutions have known costs.
|
||||
|
||||
### On a desktop, two families
|
||||
|
||||
**A collapsible stack.** Lightroom Classic's right panel is a vertical column
|
||||
of modules — Basic, Tone Curve, HSL, Detail, Lens, Effects — each collapsible,
|
||||
each reporting whether anything inside it has been touched. Everything is in
|
||||
one place and in a fixed order, so muscle memory works; the cost is a long
|
||||
scroll and a lot of triangles.
|
||||
|
||||
**Tool tabs.** Capture One puts an icon strip at the top of the tool panel and
|
||||
gives each tab a curated set of tools; RawTherapee tabs the right-hand panel
|
||||
the same way. Scroll is bounded and there is a clear sense of place; the cost
|
||||
is that a control you cannot name is in one of eight places, and switching tabs
|
||||
loses the context you were comparing against.
|
||||
|
||||
**Groups over a stack.** darktable combines both — a row of group icons
|
||||
filtering a long list of collapsible modules, plus a search box. It is the most
|
||||
powerful and the most often described as overwhelming, which is worth reading
|
||||
as a warning about combining mechanisms rather than about either one.
|
||||
|
||||
Across all of them, **masking is its own tool**, not a panel among peers.
|
||||
Lightroom opens a masking tool with its own layer list and its own canvas
|
||||
overlay; Capture One makes layers a persistent selector at the top of the
|
||||
adjustments tab. Nobody makes a mask a panel that silently rewires a different
|
||||
panel — which is what DarkRoom currently does.
|
||||
|
||||
### On a phone, one family — and why it does not apply here
|
||||
|
||||
Lightroom Mobile, Photomator and VSCO all converge on the same shape: **a
|
||||
horizontal strip of tool icons along the bottom**, and tapping one replaces a
|
||||
bottom sheet with that tool's controls. Snapseed goes further — one tool fills
|
||||
the screen and a vertical swipe chooses the parameter while a horizontal one
|
||||
sets it.
|
||||
|
||||
The convergence is not fashion. Three physical facts drive it:
|
||||
|
||||
- **Thumb reach.** A one-handed grip reaches the bottom third. A right-hand
|
||||
column is a mouse idiom.
|
||||
- **The image needs the screen.** A 280px column is a fifth of a desktop
|
||||
window and most of a phone in portrait.
|
||||
- **There is no hover.** Disclosure triangles and hover-revealed affordances
|
||||
are worth less; a control is either visible or gone.
|
||||
|
||||
**None of the first two apply to DarkRoom's targets**, which are a 12-inch
|
||||
tablet and a desktop (§3, D-N2). Nobody thumbs a 12-inch tablet one-handed,
|
||||
and its narrow dimension is not narrow. The third does apply, and is handled
|
||||
already — see D-N2.
|
||||
|
||||
---
|
||||
|
||||
## 3. The decisions
|
||||
|
||||
### D-N1 — Local adjustment becomes a mode · **DECIDED**
|
||||
|
||||
Not a panel that re-points another panel. A mode, in the sense `crop-mode`
|
||||
already is: it changes what the canvas does, scopes what the column shows, and
|
||||
is left explicitly.
|
||||
|
||||
**Why this shape rather than louder signalling.** The app already has this
|
||||
pattern and the user already knows it. Crop mode arms a canvas interaction,
|
||||
draws an overlay, gives the column one job, and exits by the same control that
|
||||
entered it. Local masking is the same animal — a canvas interaction plus a
|
||||
scoped panel — and building it as a peer panel is what created §1.1. Making it
|
||||
a mode removes the ambiguity by construction instead of describing it in a
|
||||
caption.
|
||||
|
||||
It also inherits machinery that exists. `lib.rs` already resolves Escape and
|
||||
the Android back gesture to "leave the innermost state first" (FR-UI-5); local
|
||||
mode joins that stack and needs no new exit concept.
|
||||
|
||||
In local mode:
|
||||
|
||||
- the canvas turns the overlay on and arms click-to-select;
|
||||
- the column shows the mask stack and, beneath it, the adjustments **scoped to
|
||||
the selected layer**;
|
||||
- the header names the scope — the layer, not "adjust";
|
||||
- leaving returns to the whole photograph, by Escape, by back, or by the mode
|
||||
control.
|
||||
|
||||
**The histogram follows the scope.** Under a mask it reduces over the masked
|
||||
pixels only. This is FR-DSP-7 read literally — it asks the histogram to
|
||||
describe what is being looked at — and without it the instrument and the
|
||||
controls disagree about what they are measuring. Costs a mask term in the
|
||||
histogram reduction, which already runs per frame over the displayed frame.
|
||||
|
||||
### D-N2 — One layout, because both targets are wide · **PARTLY REVERSED**
|
||||
|
||||
> **Reversed for navigation, 2026-09-05, by use.** The reasoning below is
|
||||
> still right about *size* and still right that `cfg(target_os)` is the wrong
|
||||
> axis. It is wrong in one place, and the wrong bit is the sentence "touch
|
||||
> changes **hit regions, not layout**". See D-N6.
|
||||
>
|
||||
> **Reversed for portrait, 2026-09-06, by arithmetic.** "Both orientations of
|
||||
> both targets are the expanded class" is still true and is no longer the
|
||||
> point. It was worked out for a 4:3 panel; the tablet's is 25:16, and on that
|
||||
> aspect a column *beside* the photograph in portrait leaves it a strip. See
|
||||
> D-N7, which keeps the layout class and adds an axis D-N2 did not consider.
|
||||
|
||||
The question was whether desktop and Android should diverge. The answer turns
|
||||
out to be that **neither the platform nor the width axis separates DarkRoom's
|
||||
targets**, so there is no divergence to build.
|
||||
|
||||
**The targets are a 12-inch tablet and a desktop.** No phone, decided
|
||||
2026-08-22. A 12-inch tablet is roughly 1024 logical pixels across in portrait
|
||||
and 1400 in landscape; `EXPANDED_MIN_WIDTH` is 820. **Both orientations of both
|
||||
targets are the expanded class.** The compact class now fires only when a
|
||||
desktop window is dragged under 820px, which is a case to degrade gracefully
|
||||
into, not a second interface to design.
|
||||
|
||||
**Platform would have been the wrong axis anyway**, and it is worth recording
|
||||
why so it is not proposed again. A tablet in landscape wants what a desktop
|
||||
wants; a desktop window dragged narrow wants what a small screen wants.
|
||||
Splitting on `cfg(target_os)` gives one *physical* situation two answers
|
||||
depending on which binary it happens to be. `apply_layout_class` already says
|
||||
this in its own comment — "logical pixels, not a device check" — and it was
|
||||
right.
|
||||
|
||||
**What actually differs between the two targets is input, not size**, and the
|
||||
architecture has already decided that too. `WidgetDemand::precise_pointing`
|
||||
exists for a frontend driving a television with a remote, and its own
|
||||
documentation states the position: *touch is fine, since hit regions grow to
|
||||
the modality* (FR-UI-7). Touch changes **hit regions, not layout**. A control
|
||||
is drawn where it belongs and its target grows past its own bounds — which
|
||||
`Check` and the mask rows already do.
|
||||
|
||||
So: **one develop layout, tuned for a wide viewport, with touch targets
|
||||
throughout.** The consequences worth stating:
|
||||
|
||||
- **A guaranteed-wide viewport is an asset.** The extent problem (§1.2) can be
|
||||
solved by pinning rather than by hiding — see N4.
|
||||
- **No hover-only affordance may carry meaning.** Hover may *emphasise*; it may
|
||||
never be the only way to discover a control. The mask rows already obey this
|
||||
— the eye and the delete target are drawn, not revealed.
|
||||
- **No modifier key may be required.** A tablet has no shift. Local masking
|
||||
already lost its shift-click extend for this reason, and nothing should
|
||||
reintroduce one as the only route to a feature.
|
||||
- **Anything dragged needs a finger-sized target.** ~~This is the live one:
|
||||
gradient masks have no on-canvas handles yet~~ — they have them now (N2a), and
|
||||
their handles are the first control in the app designed to be dragged on a
|
||||
photograph rather than in a panel. Drawn at 14px so they do not hide the edge
|
||||
they sit on, with a full touch target centred on the drawing, which is the
|
||||
split `Button` already establishes. Nothing about them is revealed by hover
|
||||
and nothing about them is qualified by a modifier: what is drawn is all there
|
||||
is.
|
||||
|
||||
### D-N6 — The groups move to the rail under a finger · **DECIDED**
|
||||
|
||||
Reported from a tablet: the tool rail is *"very useful"* there, and the same
|
||||
interface with a mouse and keyboard is not ergonomic. That is D-N2's assumption
|
||||
failing in the field, and it is worth being precise about which half failed.
|
||||
|
||||
**What D-N2 got right.** Platform is the wrong axis, and width is the wrong
|
||||
axis. A tablet in landscape wants what a desktop wants; a desktop window
|
||||
dragged narrow wants what a small screen wants. `apply_layout_class` still
|
||||
decides the layout class from the window, and nothing here changes that.
|
||||
|
||||
**What it got wrong.** It identified input as the real difference between the
|
||||
targets and then concluded that input changes only hit regions. Two controls
|
||||
answering one question — *which group of adjustments am I looking at* —
|
||||
disprove it:
|
||||
|
||||
- **A horizontal strip above the column.** One gesture to a target the eye has
|
||||
already found; costs one row of a column with height to spare. It pans when
|
||||
the operation set is rich, so a group can be off the end with nothing saying
|
||||
so — which a pointer user tolerates and a finger user does not discover.
|
||||
- **A vertical run down the rail.** Every entry visible at once, each a
|
||||
finger-sized target, on the edge of the screen the hand is already holding.
|
||||
Costs nothing extra in width, because the rail is already there and already
|
||||
mandated.
|
||||
|
||||
Neither is better in general. The first is better with a pointer and the second
|
||||
is better with a finger, which is a divergence on **input modality** — the axis
|
||||
D-N2 itself named.
|
||||
|
||||
**The shape.** `ToolRail` grows a second section below a rule: the same
|
||||
`adjust-tabs` model the strip takes, plus "All". `GroupStrip` stands down when
|
||||
the rail carries them, so the two are never both on screen and there is no
|
||||
state to keep in step. Mode and group stay independent axes exactly as N1
|
||||
requires — one entry lit in each section, and picking a group while a tool is
|
||||
held still filters without putting the tool down.
|
||||
|
||||
**Drawn differently, still.** N1 insisted a mode and a filter must not be told
|
||||
apart by the shape of their highlight alone. The tools fill with `active-dim`
|
||||
and invert their ink; the groups take a bar down the leading edge — the
|
||||
underline from the horizontal strip, turned ninety degrees. The rule between
|
||||
the sections is the second signal.
|
||||
|
||||
**The rail scrolls now.** Its own note argued against a Flickable on the
|
||||
grounds that four entries were written in the file. With the groups in it the
|
||||
list is generated from the operation set, which is exactly the "something the
|
||||
user's data decides" the note excluded it from.
|
||||
|
||||
**And it is a preference, because the automatic answer is a guess.** Neither
|
||||
platform can be asked what the user is actually holding — an Android tablet in
|
||||
a keyboard case is being driven like a desktop, and a touchscreen laptop is
|
||||
whichever its owner says. `dr_plat::is_touch_first` reports the usual case per
|
||||
platform and `dr_types::GroupNavigation` lets it be overridden; Settings names
|
||||
what Automatic resolves to on this device rather than leaving it to be found by
|
||||
pressing.
|
||||
|
||||
**Still open: whether Local is a mode at all.** The rail now holds two kinds of
|
||||
entry, and a third reading is available — that Compose and Repair are
|
||||
categories with a canvas gesture attached, while Local edits *nothing* and
|
||||
instead changes what every other category applies to. That would make it a
|
||||
**scope**, not a peer of the tools, and would collapse the two sections into
|
||||
one list of seven. It is the tidier model and a much larger change; deferred
|
||||
until the two-section rail has been lived with. §1.1's complaint was that scope
|
||||
was invisible, so this is the same argument arriving from the other end.
|
||||
|
||||
### D-N7 — The column docks under the photograph on a tall window · **DECIDED**
|
||||
|
||||
D-N2 dismissed portrait with one number: a 12-inch tablet is about 1024
|
||||
logical pixels across in portrait, which clears `EXPANDED_MIN_WIDTH`, so
|
||||
portrait is expanded, so there is nothing to design. The number was for a 4:3
|
||||
panel. **The tablet's panel is 3000 × 1920**, which is 25:16 — closer to a
|
||||
sheet of A4 than to an iPad — and the same arithmetic on that aspect comes out
|
||||
the other way.
|
||||
|
||||
**What the photograph gets.** Logical size depends on the density Android
|
||||
reports, which nothing in this repository records (N6 measures it). At a scale
|
||||
of 2.0 the window is 960 × 1500 in portrait; at 1.75 it is 1097 × 1714. Take
|
||||
the first, subtract the 60px rail, the 360px column and the 44px status bar,
|
||||
and the canvas beside the column is **540 × 1456** — a strip two and a half
|
||||
times taller than it is wide. Against a column *under* the canvas, 480px tall:
|
||||
|
||||
| Photograph | Beside the column | Under the column | Gain |
|
||||
|---|---|---|---|
|
||||
| 3:2, landscape | 540 × 360 | 900 × 600 | 2.8× the area |
|
||||
| 2:3, portrait | 540 × 810 | 651 × 976 | 1.4× the area |
|
||||
|
||||
At 1.75 the figures move and the ratios hold (2.4× and 1.3×). So on this panel
|
||||
the dock wins for **both** orientations of the photograph, not only the
|
||||
landscape frame one would guess it was for. That is what makes it a decision
|
||||
rather than a preference: the column is on the right because the eye's path is
|
||||
tool, photograph, adjustment, and in portrait the photograph in the middle of
|
||||
that path is the thing being starved.
|
||||
|
||||
**What it is not.** N5 said "no bottom sheet, no second layout, no tool strip
|
||||
along the bottom — those solve a phone". Still true of all three. This is not
|
||||
a sheet: the column does not slide over the photograph, it sits beside it on
|
||||
the other axis, with the same contents, the same collapse and the same toggle.
|
||||
It is not a second layout in D-N2's sense: the layout class is still decided
|
||||
by width, the compact class still means what it meant, and a tall narrow
|
||||
desktop window gets exactly what a portrait tablet gets, which is FR-UI-1's
|
||||
rule. And the rail does not move — `toolrail.slint` argued it never should,
|
||||
and a vertical list of finger-sized entries wants height, which portrait has
|
||||
more of.
|
||||
|
||||
**The axis is aspect, not width, and it is independent of the class.** A 960
|
||||
wide portrait window is expanded by width and wants the dock; a 1500 wide
|
||||
landscape one is expanded by width and does not. So this is a third property
|
||||
beside `layout-class` and `panel-max-width`, set from the same place for the
|
||||
same reason — a size that both derives from and feeds the layout is a binding
|
||||
loop in Slint, and `apply_layout_class` already measures the window. The
|
||||
window-resized callback reports width alone today and grows a height. The
|
||||
threshold is height above 1.2 × width, with hysteresis wide enough that a
|
||||
window resized across square does not flap.
|
||||
|
||||
**Slint cannot turn a layout on its side**, and does not need to. The develop
|
||||
view is one `HorizontalLayout` of rail, canvas and column; it becomes a
|
||||
`Rectangle` whose three children take `x`, `y`, `width` and `height` from the
|
||||
flag. The column's own `VerticalLayout` and `Flickable` are untouched. The
|
||||
alternative — the column subtree declared twice under two `if`s — is 400 lines
|
||||
of bindings copied, in a file whose own notes record conditional children in
|
||||
layouts as the shape that has produced binding loops before.
|
||||
|
||||
**The contents are the real cost.** Everything in the column was drawn for a
|
||||
360px vertical scroll, and the dock is 900 to 1040 wide by about 480 tall.
|
||||
Stretched to that width the stack works — sliders get longer tracks, the
|
||||
histogram divides the width, text wraps — and it is one long scroll in a short
|
||||
box, with the histogram scrolling away from the sliders it serves, which is
|
||||
§1.2 again. The width is two and a half to three of today's columns, so the
|
||||
composition that fits it is three of them side by side (N9): the instruments,
|
||||
the sliders, and the panels the mode adds. That needs the panels declared a
|
||||
second time, which is cheap only once their callbacks stop being forwarded
|
||||
through the window root by hand (N8). The stretched stack ships first as the
|
||||
stopgap (N7), because the photograph gets its area back on day one and the
|
||||
dock's contents can be got right afterwards.
|
||||
|
||||
**The dock's height is mandated, as the column's width is.** `panel-width`
|
||||
exists because a column that sizes itself to its contents is a photograph
|
||||
that changes size when a caption does; a dock has the same disease on the
|
||||
other axis. `dock-height` in `style.yaml`, 480 until N6 says otherwise: room
|
||||
for the pinned instruments (about 260px per N4) and a group of sliders under
|
||||
them, and a 3:2 frame at 900 wide still fits above it at either scale.
|
||||
|
||||
**The filmstrip stays where it is.** It takes its strip off the bottom of the
|
||||
photograph on demand and it keeps doing so; with the dock below it sits
|
||||
between the two. The photograph loses 108px while the roll is open, which is
|
||||
what it loses today, and the roll is not made part of the dock because it is a
|
||||
different kind of thing — navigation, not adjustment — and D-N6 has already
|
||||
been through why two kinds of entry in one control need a rule between them.
|
||||
|
||||
**Not remembered separately.** `PanelChoices` keeps the user's open-or-closed
|
||||
override per layout class. The dock does not add a class and does not add a
|
||||
remembered state: closing the column in landscape closes the dock in
|
||||
portrait, because it is the same column.
|
||||
|
||||
### D-N3 — Collapsible panels, not tool tabs · **OPEN**
|
||||
|
||||
For the expanded layout, extend `ui-refinement.md` Workstream C from
|
||||
sections-inside-adjust to the panels themselves: each collapses to a header
|
||||
carrying a modified dot, and collapse state survives a drag elsewhere.
|
||||
|
||||
**Why this over tabs, on the evidence above.** Tabs need a taxonomy, and the
|
||||
taxonomy is the problem. `ui-refinement.md` condemns the `starts-group` flag
|
||||
for being the core telling the panel where sections go, and FR-DEV-3a requires
|
||||
that adding an operation needs no UI edit. A tab strip built from a hardcoded
|
||||
op-id → tab table in `dr-ui` breaks the second; one built from a `group:` field
|
||||
in `ops/*.yaml` risks breaking the first.
|
||||
|
||||
There *is* a legitimate route to tabs, and it should be recorded rather than
|
||||
discovered later: the descriptor could declare an operation's **nature** —
|
||||
tone, colour, detail, optics — the same shape as `Affects` and `ParamKind`
|
||||
already take. The core would be saying *what the operation is*, which is its
|
||||
business, and the frontend would remain free to render that as a tab, a
|
||||
section heading, or nothing at all. That stays on the right side of §4.3a.
|
||||
|
||||
**The recommendation is to defer it.** Ten operations do not need eight tabs,
|
||||
collapse needs no taxonomy at all, and the nature field is easy to add later
|
||||
and awkward to remove. Revisit when the operation count passes roughly fifteen
|
||||
— which FR-DEV-3's outstanding list will reach.
|
||||
|
||||
**Open, because it is a taste call**: whether the expanded layout should also
|
||||
gain tabs, or stay a single collapsible stack indefinitely.
|
||||
|
||||
---
|
||||
|
||||
## 4. Workstreams
|
||||
|
||||
Numbered N to avoid colliding with `ui-refinement.md`'s A–F.
|
||||
|
||||
### N1 — The mode strip — **done**
|
||||
|
||||
**Deliverable.** ~~One control naming the current mode, replacing the implicit
|
||||
`crop-mode` boolean: **Photo · Crop · Local**. Top of the canvas in expanded,
|
||||
bottom in compact.~~ The selected mode is the accent's job — it means *active*,
|
||||
which is exactly this.
|
||||
|
||||
`crop-mode` becomes one value of a mode enum rather than its own flag, so the
|
||||
two modes cannot both be on, which today they can.
|
||||
|
||||
**Landed as one strip, not two.** The mode control and the existing group strip
|
||||
(`All · Light · Colour`, derived from operation attributes) were going to sit
|
||||
beside each other above the same column, which is two controls answering one
|
||||
question — *what am I working on*. They are now one:
|
||||
|
||||
```
|
||||
Crop · Local │ All Light Colour
|
||||
```
|
||||
|
||||
Lightroom Mobile's bottom strip mixes Crop and Masking with Light and Colour
|
||||
for the same reason, and it reads naturally because from the photographer's
|
||||
side they are the same kind of choice.
|
||||
|
||||
The two halves are **different kinds of state and are drawn differently**: a
|
||||
mode is a chip that fills with the accent when it is on, a group is a word with
|
||||
a rule under it. That is what lets both be read at once, and both are on at
|
||||
once routinely — see below.
|
||||
|
||||
**Mode and group are independent axes.** Picking `Light` while a mask layer is
|
||||
selected filters *that layer's* chain and does not leave local mode. The
|
||||
alternative — a group press quietly dropping the scope — would be §1.1's fault
|
||||
reintroduced from the other end, and it would make `Light` mean two things
|
||||
depending on where it was pressed.
|
||||
|
||||
**Where it sits.** Pinned above the develop column, where the group strip
|
||||
already was, rather than at the top of the canvas. The half that filters the
|
||||
column belongs to the column, and moving it onto the photograph would put it
|
||||
somewhere the four principles say chrome should not be. The canvas keeps one
|
||||
button — now *"Done Cropping"* / *"Done Masking"*, naming the mode it leaves —
|
||||
because the column can be closed on a narrow window and no mode may be
|
||||
inescapable.
|
||||
|
||||
**Also landed.** `GeometryPanel`'s Crop button is gone: a second control
|
||||
entering the same mode is a second thing that has to agree about which mode the
|
||||
view is in.
|
||||
|
||||
**Done when.** ~~Entering crop from the strip does what the crop button did;
|
||||
Escape and back leave the innermost mode; no two modes are ever active
|
||||
together.~~ All three, checked on screen as well as in tests — `back_step` has
|
||||
one `LeaveMode` step covering both modes, and the enum makes "no two at once"
|
||||
unrepresentable rather than merely untested.
|
||||
|
||||
### N2 — Local mode — **done**
|
||||
|
||||
**Depends on** N1.
|
||||
|
||||
**Deliverable.** Entering local mode turns the overlay on and arms picking
|
||||
without either being a separate toggle — they are what the mode *is*. The
|
||||
column shows the mask stack, then the scoped adjustments. The adjust header
|
||||
names the layer.
|
||||
|
||||
Leaving local mode clears the selection so the adjustments are unambiguously
|
||||
global again.
|
||||
|
||||
**Landed.** The "Overlay" and "Select" buttons are gone; entering the mode does
|
||||
both, and `region-picking` is now derived from the mode rather than toggled.
|
||||
The masking panel is no longer a panel among peers in the scrolling column — it
|
||||
appears only in local mode, which is what takes the column from six panels to
|
||||
three there.
|
||||
|
||||
**The scope is the adjust panel's own heading**, not a caption in the panel
|
||||
above it. `ADJUST` becomes the layer's name. That is the difference between
|
||||
describing the hazard and removing it: the heading of the thing that changed
|
||||
cannot be skipped on the way to a slider, and a caption in a different panel
|
||||
routinely was.
|
||||
|
||||
**What local mode drops from the column**: the capture metadata, the framing
|
||||
controls and copy/paste. None is a property of a region within the photograph,
|
||||
so all three would be controls in scope of nothing. The histogram stays and
|
||||
still reports the whole frame — the disagreement §1.1 names is real and is N3's
|
||||
to close; removing the instrument would be a worse answer than an honest one
|
||||
that is not yet scoped.
|
||||
|
||||
**Done when.** ~~There is no way to have a mask selected without knowing it, and
|
||||
the two toggles that currently arm the overlay and picking are gone.~~ Both.
|
||||
|
||||
### N2a — Gradient handles — **done**
|
||||
|
||||
Not a numbered workstream when this was written, and it belongs beside N2: the
|
||||
canvas half of local mode.
|
||||
|
||||
`MaskSource::Linear` and `MaskSource::Radial` could be created and then not
|
||||
moved, so a radial sat at the centre of the frame at its default size for ever.
|
||||
They now carry handles on the photograph — the first controls in the
|
||||
application designed to be dragged there rather than in a panel, and D-N2's
|
||||
"the live one".
|
||||
|
||||
Three faults had to be fixed before a handle was worth drawing.
|
||||
|
||||
**A gradient did not render at all until the model had run.** The mask
|
||||
rasteriser was built on the way out of `segment`, and the array's size was read
|
||||
*off* the segmentation, so a gradient added to an unsegmented photograph
|
||||
produced nothing — silently, because the generated shader still emits the
|
||||
layer's block and the empty placeholder multiplies it by zero. The proxy size
|
||||
is a property of the photograph; both are now derived from it, deliberately at
|
||||
the same size because a subject's distance field is sampled against the array.
|
||||
|
||||
**A gradient's geometry was measured in raw `0..1` fractions**, so a 45° ramp
|
||||
was not at 45° and a radial with equal radii drew an ellipse. Angles and
|
||||
distances are now in the frame's isotropic units — y spans `0..1`, x spans
|
||||
`0..aspect` — converted in exactly one place, `frame_delta` in `mask.wgsl`.
|
||||
Only the *meaning* of the stored numbers changed; the sidecar format did not.
|
||||
|
||||
**Hit-testing has to go through the framing map.** A handle is drawn in output
|
||||
coordinates and stored in source ones, and the two are separated by the crop,
|
||||
the zoom, the pan, the straightening and the turns. `Framing::source_at` and
|
||||
`Framing::output_at` are `wgsl_prologue` evaluated on the CPU, kept in that file
|
||||
beside it so the correspondence is one file's problem.
|
||||
|
||||
**Handles.** A linear ramp has three — centre, width, angle. A radial has three
|
||||
— centre and one per semi-axis, the major one carrying the ellipse's angle as
|
||||
well as its length, because where an axis is put says both. A rotation arm was
|
||||
tried on the radial and taken out: standing off the shape by a fixed distance,
|
||||
it began outside the photograph at the size a new radial is created at.
|
||||
|
||||
**A drag is a displacement applied to where the mask was when the press
|
||||
landed**, not a destination the handle is snapped to. Snapping jerks the handle
|
||||
by up to half a touch target on the first press, and the target is finger-sized
|
||||
(FR-UI-3).
|
||||
|
||||
**Two faults found by looking at the screen** rather than by reading the source,
|
||||
both of the kind `ui-refinement.md`'s verification section warns about. A `1px`
|
||||
rule with a size and no position is *centred* by Slint, so the develop column's
|
||||
seam was a hairline down the middle of the panel — twice over, once in
|
||||
`app.slint` and once in `AdjustPanel`. And handing Slint a new `ModelRc` for the
|
||||
handles on every pointer event made the repeater rebuild its items, taking the
|
||||
`TouchArea` holding the gesture with them: the handle jumped once and then went
|
||||
dead under a finger that was still down. The model is now rewritten in place.
|
||||
|
||||
### N3 — Scope-following histogram
|
||||
|
||||
**Depends on** N2, which has landed, so this is next and is the outstanding
|
||||
half of §1.1: the panel now says *which* chain the sliders edit, and the
|
||||
instrument beside them still measures the other one.
|
||||
|
||||
**Deliverable.** The histogram reduction takes an optional mask; in local mode
|
||||
it reduces over the selected layer's coverage. The panel says which it is
|
||||
showing, because a histogram of a face is a strange shape and the user should
|
||||
know why.
|
||||
|
||||
**Done when.** Selecting a layer visibly changes the histogram, and leaving
|
||||
local mode restores the frame's.
|
||||
|
||||
### N4 — Collapsible panels
|
||||
|
||||
**Depends on** `ui-refinement.md` A. Extends C.
|
||||
|
||||
**Deliverable, two halves.**
|
||||
|
||||
*Collapse.* Each panel collapses to its header; headers carry the modified dot
|
||||
C defines; state is keyed by panel identity and survives a slider drag.
|
||||
Default: histogram open, the rest collapsed until touched.
|
||||
|
||||
*Pin.* The histogram and the scope header do **not** scroll. They sit above the
|
||||
scrolling region, always visible.
|
||||
|
||||
Pinning is the half a guaranteed-wide viewport buys, and it addresses §1.2
|
||||
directly rather than obliquely. The complaint is not that the column is long —
|
||||
it is that the instrument every tonal control is judged against scrolls away
|
||||
from the controls it reports on. Collapsing panels shortens the scroll;
|
||||
pinning removes the problem. Together they cost about 260px of fixed height,
|
||||
which a target that is never under 820px wide and rarely under 1000 tall can
|
||||
afford.
|
||||
|
||||
**Done when.** The histogram is visible while any tonal slider is being
|
||||
dragged, at every window size the targets produce, without the user having
|
||||
scrolled to arrange it.
|
||||
|
||||
### N5 — Compact degrades, rather than diverges
|
||||
|
||||
**Depends on** N4.
|
||||
|
||||
Not a second interface. Below `EXPANDED_MIN_WIDTH` the develop column already
|
||||
overlays rather than sits beside the canvas, and `apply_layout_class` already
|
||||
remembers the user's override per class. With N4's collapse in place a narrow
|
||||
window is a one-panel-at-a-time column by consequence rather than by design.
|
||||
|
||||
**Deliverable.** Confirm the narrow case is usable and fix what is not. ~~No
|
||||
bottom sheet, no second layout, no tool strip along the bottom — those solve a
|
||||
phone, and there is no phone.~~ Still no sheet and still no strip; but a
|
||||
*tall* window is not a narrow one, and D-N7 (2026-09-06) puts the column under
|
||||
the photograph there. N6–N9 carry it. This workstream keeps the narrow case.
|
||||
|
||||
**Done when.** A desktop window dragged to 700px shows the photograph and a
|
||||
usable column, and nothing is unreachable that was reachable at 1400px.
|
||||
|
||||
### N6 — Measure the tablet
|
||||
|
||||
**Depends on** nothing. Hours.
|
||||
|
||||
Every figure in D-N7 is computed at a guessed scale factor. The tablet's
|
||||
panel is 3000 × 1920 physical; its logical size is that divided by whatever
|
||||
density Android reports, and the two plausible answers (2.0 and 1.75) put the
|
||||
dock's width at 900 or 1037 and its available height 200px apart.
|
||||
|
||||
**Deliverable.** Log the window's scale factor and logical size once at
|
||||
startup, beside the existing `apply_layout_class` call, at a level that
|
||||
reaches `adb logcat`. Run it on the tablet in both orientations. Record the
|
||||
four numbers in D-N7 and, if they move the table, correct it. #30
|
||||
(NFR-COMPAT-1) wants a reference device named; this is one of the numbers
|
||||
that names it.
|
||||
|
||||
**Done when.** D-N7 cites measured logical sizes, not a scale it assumed, and
|
||||
`dock-height` has been checked against the measured portrait height.
|
||||
|
||||
### N7 — The column docks on a tall window
|
||||
|
||||
**Depends on** N6 only for the value of `dock-height`; the shape does not
|
||||
wait for it.
|
||||
|
||||
**Deliverable, three parts.**
|
||||
|
||||
*The flag.* `window-resized` reports height as well as width.
|
||||
`apply_layout_class` derives a third property, `column-below`, from the
|
||||
aspect with hysteresis (enter above 1.25, leave below 1.15, or thereabouts —
|
||||
the point is that a window resized across square does not flap), and sets it
|
||||
beside `layout-class` and `panel-max-width`. It is not a layout class and
|
||||
`PanelChoices` does not learn about it.
|
||||
|
||||
*The frame.* The `HorizontalLayout` holding `ToolRail`, `canvas-area` and
|
||||
`develop-column` becomes a `Rectangle` and each child takes `x`, `y`, `width`
|
||||
and `height` from the flag. The rail is full height on the left in both
|
||||
cases. With the flag off the geometry is what the layout produced, to the
|
||||
pixel — screenshot before and after and diff them. With it on, the column is
|
||||
`dock-height` tall and runs from the rail's edge to the window's; the canvas
|
||||
has what is left above it. `panel-visible` collapses the dock to zero height
|
||||
exactly as it collapses the column to zero width.
|
||||
|
||||
*The contents, as a stopgap.* The column's stack stretches to the dock's
|
||||
width: the Flickable's viewport width follows the dock, the `min-width` floor
|
||||
that keeps a mandated column honest is still there and is simply not binding.
|
||||
Sliders take the width. Nothing is reflowed; N9 does that.
|
||||
|
||||
`dock-height` goes in `style.yaml` next to `panel-width`, with the reasoning
|
||||
D-N7 gives, and is read as `Theme.dock-height`.
|
||||
|
||||
**Done when.** On the tablet in portrait, a 3:2 photograph is drawn at the
|
||||
width of the canvas, not at the width the old column left it; turning the
|
||||
tablet moves the column back beside it with the same scroll position and the
|
||||
same panels open; a desktop window dragged taller than it is wide does the
|
||||
same; and the screenshot diff in landscape is empty.
|
||||
|
||||
### N8 — Develop callbacks onto a global
|
||||
|
||||
**Depends on** nothing; can run beside N7.
|
||||
|
||||
The develop panels — `AdjustPanel`, `MaskPanel`, `SpotPanel`, `ComposePanel`,
|
||||
`TransferPanel`, `FocusPanel`, `HistogramPanel`, `InfoPanel` — each forward
|
||||
their callbacks and take their inputs through the window root, and the
|
||||
instantiation in `app.slint` that wires one up is twenty to forty lines. N9
|
||||
needs each of them declared a second time, and copying that wiring is the
|
||||
kind of duplication that drifts.
|
||||
|
||||
**Deliverable.** A Slint global (or one per panel family, if a single one
|
||||
reads badly) carrying the develop callbacks and the inputs the panels bind
|
||||
to. A panel calls the global directly; Rust hooks the global instead of the
|
||||
window. The existing instantiation in the column shrinks to the properties
|
||||
that genuinely differ by placement, which should be none. `Readout` in
|
||||
`adjust.slint` is the precedent for a global in this codebase.
|
||||
|
||||
The tests in `ui/dr-ui/src` that drive these callbacks through the window
|
||||
move to the global; count them before starting, so the ticket knows its own
|
||||
size.
|
||||
|
||||
**Done when.** No develop panel's instantiation in `app.slint` forwards a
|
||||
callback by hand, every existing test passes, and a second instantiation of
|
||||
any panel is under five lines.
|
||||
|
||||
### N9 — Three columns in the dock
|
||||
|
||||
**Depends on** N7 and N8.
|
||||
|
||||
**Deliverable.** A second composition of the same panels for the dock, in
|
||||
three columns of equal width side by side, each its own scroll:
|
||||
|
||||
1. **Instruments** — `InfoPanel`, `HistogramPanel`, `FocusPanel`. What the
|
||||
sliders are judged against, pinned by construction: it does not scroll
|
||||
with them because it is not in their column.
|
||||
2. **Sliders** — `GroupStrip` above `AdjustPanel`, exactly as in the column.
|
||||
With a group selected this is one screen of sliders; with All it scrolls.
|
||||
3. **The mode's panels** — `ComposePanel` and `TransferPanel` in photo mode,
|
||||
`SpotPanel` in repair, `MaskPanel` in local. Empty otherwise, which is
|
||||
a signal of its own about which mode the view is in (§1.1).
|
||||
|
||||
Three columns at 900 wide are 300 each, and at 1037 they are 346: inside the
|
||||
280–360 band `PANEL_MIN_WIDTH`'s note says the sliders stay accurate over.
|
||||
The dock takes the three-column composition only when a third of its width
|
||||
is at least `PANEL_MIN_WIDTH` — 840px of dock, which both scales of the
|
||||
tablet exceed. Below that it keeps N7's stretched stack. Two compositions,
|
||||
not three: a dock too narrow for three columns is a desktop window in an odd
|
||||
shape, and N5 says that case degrades.
|
||||
|
||||
The panels are declared twice, once per composition, which is what N8 made
|
||||
cheap. Declaring each once and positioning it by hand in both modes was
|
||||
considered and rejected: a column that scrolls is a `Flickable`, a
|
||||
`Flickable`'s children are its children, and a panel cannot be in two.
|
||||
|
||||
**Done when.** On the tablet in portrait the histogram is visible while any
|
||||
slider in any group is dragged, with no scrolling to arrange it — N4's own
|
||||
criterion, met in the dock by construction; every panel reachable in
|
||||
landscape is reachable in portrait; and the instantiation of each panel in
|
||||
the dock is a handful of lines.
|
||||
|
||||
---
|
||||
|
||||
## 5. Sequencing
|
||||
|
||||
```
|
||||
ui-refinement A ──┬── C ──── N4 ──┐
|
||||
└── D ──── F ├── N5
|
||||
N1 ── N2 ── N3 ───────────────────┘
|
||||
|
||||
N6 ── N7 ──┐
|
||||
├── N9
|
||||
N8 ──┘
|
||||
```
|
||||
|
||||
N1–N3 are independent of `ui-refinement.md` and can start now; they touch the
|
||||
canvas and the develop column's contents, not its layout. N4 needs C's
|
||||
`Section`. N5 needs both and should land last, as F does — it is the one that
|
||||
rearranges everything.
|
||||
|
||||
N6–N9 are the portrait dock (D-N7) and run beside the first row rather than
|
||||
after it. N7 is the one that touches the develop view's frame, so it should
|
||||
not land in the same wave as F; N8 touches only plumbing and can. N9 waits
|
||||
for both, and gains from N4 if N4 is in by then — a collapsible panel in a
|
||||
300px column is worth more than in a 360px one.
|
||||
|
||||
## 6. Invariants, for every workstream
|
||||
|
||||
- **FR-DEV-3a.** Adding a pipeline operation must still surface in both
|
||||
layouts with no UI edit. No file under `ui/` may name an operation.
|
||||
- **ARCH §4.3a.** Composition is the frontend's decision. The core may declare
|
||||
what an operation *is*; it may not declare where the panel puts it.
|
||||
- **FR-UI-3 / FR-UI-7.** Touch changes hit regions, not layout. Every control
|
||||
keeps a finger-sized target wherever it is drawn, and no hover-only
|
||||
affordance and no modifier key may be the sole route to anything — a 12-inch
|
||||
tablet has neither.
|
||||
- **FR-UI-5.** Escape and the Android back gesture leave the innermost state
|
||||
first. Every mode added here joins that order.
|
||||
|
||||
---
|
||||
|
||||
## 7. Place — the other half of "where am I"
|
||||
|
||||
§1 asked how someone *finds* anything. This asks how they stop losing what they
|
||||
already found. Both are navigation; the second is the one nobody notices until
|
||||
it is wrong, and then notices constantly.
|
||||
|
||||
### 7.1 Three failures, one cause
|
||||
|
||||
The library grid is gated on an `if` in the markup, so **every** route away from
|
||||
it destroys the subtree and rebuilds it on return. The Flickable inside passes
|
||||
its viewport through zero on the way out. Three consequences, reported
|
||||
separately and all the same bug:
|
||||
|
||||
- *"Opening Settings and coming back puts me at the top."* The guard on the
|
||||
scroll handler tested `show-library`, which stays true while Settings, Import,
|
||||
People or the launch screen covers the grid. The teardown's scroll-to-zero
|
||||
passed it, and the remembered position was overwritten with 0.
|
||||
- *"Coming out of develop I lose the photograph I was editing."* The position
|
||||
restored was where the *grid* was, not what was open — and walking the photo
|
||||
roll moves the second a long way from the first.
|
||||
- *"Launching puts me at the beginning."* Nothing was written down at all.
|
||||
|
||||
`app.slint` now computes `library-visible` once — the same expression the `if`
|
||||
is spelled from — and Rust reads that rather than `show-library`. The two cannot
|
||||
drift apart, which is what let them drift in the first place.
|
||||
|
||||
### 7.2 Two positions, and which one wins
|
||||
|
||||
Returning from develop has two candidates: `resume_at`, where the grid was, and
|
||||
the open photograph. They agree in the ordinary case and disagree after a walk
|
||||
along the roll.
|
||||
|
||||
The rule is not "pick one". The grid seeks to the remembered position, then
|
||||
`reveal()`s the keyboard cursor, which is on the open photograph and which moves
|
||||
the viewport as little as will bring it into view. A frame inside the remembered
|
||||
screenful moves nothing; one outside it scrolls exactly far enough. One rule,
|
||||
both behaviours — and the cursor rather than the selection, so a set of forty
|
||||
photographs assembled in the grid survives having one of them opened.
|
||||
|
||||
### 7.3 A place is not a scroll position
|
||||
|
||||
What gets written down is the view, the scope, the filter and the photograph —
|
||||
because a position without the filter that produced it names a row of a list
|
||||
that no longer exists. Restoring them has an order for the same reason: scope,
|
||||
then filter, then position, then the view. Each step changes what an ordinal
|
||||
*means*.
|
||||
|
||||
Addressed by remote path and collection UUID, never by an ordinal or a row id.
|
||||
See FR-UI-8 and `dr_types::place` for why, and `library::ordinal_of_path` for the
|
||||
one place the ordering is inverted — through the grid's own `ORDER BY`, taken
|
||||
verbatim, rather than spelled a second time.
|
||||
|
||||
### 7.4 The handover, and when to refuse it
|
||||
|
||||
The record travels through `.darkroom-derived/place.json`, so a session begun on
|
||||
the desktop continues on the tablet. Newest wins; there is nothing to merge.
|
||||
|
||||
The interesting decision is the refusal. A place arriving from another device is
|
||||
welcome on the way in and unwelcome the moment the photographer has started
|
||||
working — a grid that jumped somewhere else mid-scroll because a round trip
|
||||
finally landed would have lost their place to the feature meant to keep it. So
|
||||
any scroll, scrub, scope change, filter or opened photograph closes the latch,
|
||||
and a record that arrives after that is still written to disk and simply takes
|
||||
effect at the next launch.
|
||||
|
||||
### 7.5 Two smaller instruments that were saying nothing
|
||||
|
||||
Both were "correct" in the sense of not being wrong, and both were useless.
|
||||
|
||||
- **The capture-time marker** rested greyed at mid-track until the first scroll,
|
||||
on the reasoning that anchoring it would imply a choice the user had not made.
|
||||
But the sidebar's claim is to say *when* you are, and that is known from the
|
||||
first frame. It is now seeded from wherever the view sits.
|
||||
- **The photo roll** brought the open frame into view by the shortest move,
|
||||
which put it hard against one edge with nothing on that side. It now centres
|
||||
on the first reveal of a develop session and steps minimally thereafter —
|
||||
a one-shot request the strip consumes, so an overlay screen rebuilding the
|
||||
view does not undo a roll the user has scrolled by hand.
|
||||
@@ -0,0 +1,296 @@
|
||||
# View composition: a controller for the display layer
|
||||
|
||||
TRACES: FR-UI-1 | FR-UI-6 | FR-UI-8 | FR-DEV-3a | NFR-P9
|
||||
|
||||
**Status:** Draft · 2026-08-09
|
||||
**Companion to:** [architecture.md](architecture.md) §4.3a, [ui-refinement.md](archive/ui-refinement.md)
|
||||
|
||||
## Why
|
||||
|
||||
`dr-ui` has four views — launch, library, develop, collections — and no view
|
||||
layer. `run()` in `ui/dr-ui/src/lib.rs` is 500 lines that construct every
|
||||
controller, wire every cross-view callback, own the develop session, and hold
|
||||
the only complete picture of what is on screen. It has become the controller
|
||||
by accretion rather than by design, and it shows in three specific ways.
|
||||
|
||||
**The back-patched knot.** `run()` declares
|
||||
|
||||
```rust
|
||||
let open_from_library: Rc<RefCell<Option<Rc<dyn Fn(String)>>>> = ...
|
||||
```
|
||||
|
||||
wired empty at line 328 and filled at line 604, because the library grid needs
|
||||
to open an image in develop and the develop closure needs the GPU context that
|
||||
is built after the library is wired. The comment in the source is candid about
|
||||
it: *"This cell is the knot between them: wired empty here, filled once `show`
|
||||
exists."* A nullable function slot resolved at runtime is what a dependency
|
||||
cycle looks like when the language will not let you write one directly.
|
||||
|
||||
**Two copies of load-and-display.** `show` (lib.rs:524) and the remote-fetch
|
||||
timer body (lib.rs:657) run the same sequence after their inputs diverge: set
|
||||
`load_error` empty, set camera / exposure / dimensions from metadata, branch on
|
||||
whether sensor data was recovered, then either `set_adjust_enabled(true)` +
|
||||
`sync_rows` + `redraw`, or clear the session, empty the rows, disable adjust
|
||||
and show the fallback image — and on failure, clear the session and report the
|
||||
error. One takes a path and one takes bytes; everything downstream is written
|
||||
twice.
|
||||
|
||||
The shared preamble has *already* been factored out into `reset_view_state`,
|
||||
which both branches call, so the direction is established — the post-load half
|
||||
is simply the part that has not been done yet. The two halves have not yet
|
||||
drifted in the fields they set; the argument for consolidating is to keep it
|
||||
that way, since every future display field must currently be added in two
|
||||
places.
|
||||
|
||||
**Diffuse ownership of window state.** Distinct `window.set_*` calls by module:
|
||||
|
||||
| Module | Distinct window properties written |
|
||||
|---|---|
|
||||
| `library_ui` | 24 |
|
||||
| `lib.rs` (`run`) | 23 |
|
||||
| `launch_ui` | 16 |
|
||||
| `collections_ui` | 9 |
|
||||
|
||||
No one owns "what is on screen". The clearest symptom is view switching
|
||||
itself: `show-launch` and `show-library` are two booleans encoding one piece
|
||||
of state, written from seven call sites across three modules
|
||||
(`lib.rs:353`, `launch_ui.rs:171`, `library_ui.rs:246/254/1486/1645/1655`).
|
||||
Nothing prevents both being true, and `app.slint` compensates with
|
||||
`if !root.show-launch && root.show-library` chains at lines 349, 386 and 491.
|
||||
|
||||
`develop.rs` is the exception that proves the point: it writes no window
|
||||
properties at all, taking capabilities in and returning rows and images out.
|
||||
It is the one module already shaped the way this document argues for.
|
||||
|
||||
None of this is broken. It works, and several of the surrounding patterns are
|
||||
load-bearing and correct — the mpsc-plus-timer worker shape exists because
|
||||
Slint's event loop must never block (NFR-P9), and `sync_rows` mutates rows in
|
||||
place because replacing the model breaks slider dragging. This document
|
||||
changes ownership, not threading and not Slint model handling.
|
||||
|
||||
## Non-goals
|
||||
|
||||
- No change to the threading model. Workers stay on threads; results still
|
||||
return through channels drained by Slint timers (NFR-P9).
|
||||
- No change to the pipeline, catalog, sync layer, or render path.
|
||||
- No change to how operations reach the adjust panel. `develop.rs` already
|
||||
composes from capabilities and names no operation (FR-DEV-3a); that stays.
|
||||
- **No dynamic view instantiation.** See the constraint below.
|
||||
|
||||
## The constraint that shapes all of this
|
||||
|
||||
Slint has no runtime component instantiation. Views are selected by statically
|
||||
compiled conditionals over properties — `if root.show-launch: LaunchScreen`
|
||||
and friends — so every view must exist in the `.slint` source at build time.
|
||||
|
||||
Descriptors therefore drive **composition and chrome**: which views exist,
|
||||
their labels, their order, whether each is currently available, and which one
|
||||
is active. They cannot conjure a view body. This is a smaller claim than
|
||||
"self-describing modules" might suggest, and stating it here is deliberate:
|
||||
a design that assumed otherwise would hit the wall at codegen.
|
||||
|
||||
`build.rs` already generates `theme.slint` from `style.yaml`, so generating a
|
||||
tab bar from the descriptor set is *possible* later. It is not proposed now —
|
||||
generated `.slint` is markedly harder to debug than generated tokens, and the
|
||||
win does not yet justify it.
|
||||
|
||||
---
|
||||
|
||||
## Stage 1 — `ViewController`
|
||||
|
||||
Independently landable, and worth landing whether or not stages 2 and 3
|
||||
follow.
|
||||
|
||||
A `ViewController` owns the window handle and the state that currently floats
|
||||
in `run()`'s closure captures: `session`, `entries`, `index`, `viewport`,
|
||||
`rows`, and `redraw`. Sub-controllers are held by it rather than wired to each
|
||||
other through `run()`.
|
||||
|
||||
```rust
|
||||
pub struct ViewController {
|
||||
window: slint::Weak<AppWindow>,
|
||||
gpu: Option<GpuContext>,
|
||||
session: RefCell<Option<DevelopSession>>,
|
||||
entries: RefCell<Vec<PathBuf>>,
|
||||
index: Cell<usize>,
|
||||
viewport: Cell<(u32, u32)>,
|
||||
rows: Rc<slint::VecModel<ParamRow>>,
|
||||
library: Rc<LibraryController>,
|
||||
collections: Rc<CollectionsController>,
|
||||
launch: Rc<LaunchController>,
|
||||
}
|
||||
```
|
||||
|
||||
Three changes follow, in order:
|
||||
|
||||
**1a. One `present()`.** Both load routes converge on a single method:
|
||||
|
||||
```rust
|
||||
impl ViewController {
|
||||
fn present(&self, loaded: Result<Loaded, String>, name: &str) { ... }
|
||||
}
|
||||
```
|
||||
|
||||
The local path calls `self.present(load(gpu, path), &name)`; the remote timer
|
||||
calls `self.present(load_bytes(gpu, &bytes), &name)`. The duplicated post-load
|
||||
sequence exists once, so a new display field can only be added in one place.
|
||||
|
||||
**1b. `Rc<ViewController>` replaces the nullable callback cell.** The grid's
|
||||
click handler captures the controller and calls
|
||||
`controller.open_remote(path)`. Both sides now depend on the controller rather
|
||||
than on each other, so the cycle disappears and the `Option` slot with it.
|
||||
|
||||
**1c. Collapse the four adjust callbacks.** `on_param_changed` (lib.rs:702),
|
||||
`on_param_reset` (716), `on_reset_all` (730) and `on_curve_reset` (747) are
|
||||
four blocks differing only in which `DevelopSession` method they call; each
|
||||
then runs the identical `sync_rows` + `redraw` pair. A single
|
||||
`self.mutate_session(|s| ...)` helper that performs the mutation and then
|
||||
re-syncs makes the shared tail impossible to forget.
|
||||
|
||||
Note that not every handler wants that tail — the view/pan handlers below them
|
||||
deliberately `redraw` without `sync_rows`, because panning changes no
|
||||
parameter. The helper must therefore be the *opinionated* path for parameter
|
||||
mutation, not a mandatory funnel for everything that touches the session.
|
||||
|
||||
**Verification.** These are refactors with no behavioural change: existing
|
||||
tests in `lib.rs` (`describe_camera`, `describe_exposure`, `collect`) must pass
|
||||
untouched, and manual checks cover launch → library → develop, next/previous,
|
||||
slider drag, reset, and canvas resize.
|
||||
|
||||
---
|
||||
|
||||
## Stage 2 — The `View` trait
|
||||
|
||||
The same discipline `descriptor.rs` already applies to operations, applied one
|
||||
level up. That precedent matters: the core publishes `OpDescriptor` /
|
||||
`ParamKind` / `Presentation`, and `develop.rs` builds controls from it without
|
||||
naming a single operation. Views are the same shape of problem.
|
||||
|
||||
```rust
|
||||
pub struct ViewDescriptor {
|
||||
pub id: ViewId,
|
||||
pub label: LocalizedKey,
|
||||
pub kind: ViewKind,
|
||||
}
|
||||
|
||||
/// Closed, not a string — a controller must be able to match exhaustively
|
||||
/// and know it has covered everything. Same reasoning as `WidgetKind`.
|
||||
pub enum ViewKind {
|
||||
/// Occupies the window alone; no chrome, not tabbable. Launch.
|
||||
Modal,
|
||||
/// Participates in the tab set. Library, develop.
|
||||
Primary,
|
||||
/// Renders beside a primary view. Collections.
|
||||
Adjunct,
|
||||
}
|
||||
|
||||
pub enum Availability {
|
||||
Available,
|
||||
/// Greyed out, with a reason the UI can show. Not hidden — a missing
|
||||
/// tab is indistinguishable from a bug.
|
||||
Unavailable(LocalizedKey),
|
||||
}
|
||||
|
||||
pub trait View {
|
||||
fn descriptor(&self) -> ViewDescriptor;
|
||||
fn availability(&self, ctx: &AppContext) -> Availability;
|
||||
fn activate(&self, ctx: &AppContext) {}
|
||||
fn deactivate(&self, ctx: &AppContext) {}
|
||||
}
|
||||
```
|
||||
|
||||
Two properties are carried up from `descriptor.rs` deliberately, because they
|
||||
are why that design works:
|
||||
|
||||
- **`ViewKind` is a closed enum.** Its `WidgetKind` counterpart says so
|
||||
explicitly: *"a UI must be able to match exhaustively and know it has
|
||||
covered everything the core can ask for."*
|
||||
- **Descriptors are hints with a working fallback.** A curve degrades to
|
||||
sliders. A view whose `kind` a shell does not implement still renders
|
||||
standalone; nothing about tabbing is required for a view to function.
|
||||
|
||||
Labels are `LocalizedKey`, reusing the type `dr-pipeline` already exports —
|
||||
`dr-ui` depends on `dr-pipeline`, so this adds no coupling, and it keeps view
|
||||
labels on the same footing as operation labels. `labels::resolve` already takes
|
||||
a `&str` and derives a readable fallback for uncatalogued keys, so a new view
|
||||
appears with a sensible label before anyone writes a translation, exactly as a
|
||||
new operation does today.
|
||||
|
||||
`availability` is the part that earns the trait rather than merely tidying.
|
||||
Develop-without-a-GPU and collections-without-a-catalog are handled
|
||||
inconsistently today — `run()` logs a warning and sets `backend` to "NO GPU",
|
||||
`library_ui` guards each catalog access separately — and the differences are
|
||||
not intentional. One method, asked before a view is offered, makes those cases
|
||||
uniform and gives the shell something honest to display.
|
||||
|
||||
`activate`/`deactivate` exist so switching away can stop timers and release
|
||||
GPU resources instead of leaking them, which nothing does today.
|
||||
|
||||
**Scope check.** Four views, all known, all in-tree, no third-party authors.
|
||||
This abstraction has to justify itself against an `if` chain, and on tab
|
||||
composition alone it would be close. The `availability` consolidation is what
|
||||
tips it, because that is a correctness fix rather than a tidiness one.
|
||||
|
||||
**Verification.** Stage 2 retrofits the trait to the four existing views and
|
||||
changes no Slint. The descriptors must reproduce current behaviour exactly
|
||||
before stage 3 consumes them.
|
||||
|
||||
---
|
||||
|
||||
## Stage 3 — Data-driven view switching
|
||||
|
||||
With descriptors in place, the boolean pair collapses:
|
||||
|
||||
```
|
||||
- in property <bool> show-launch;
|
||||
- in property <bool> show-library;
|
||||
+ in property <string> active-view;
|
||||
+ in-out property <[TabEntry]> view-tabs;
|
||||
```
|
||||
|
||||
`app.slint`'s conditionals key off `active-view == "library"` rather than a
|
||||
two-boolean conjunction, and the illegal both-true state stops being
|
||||
representable. `view-tabs` is a model the controller publishes from the
|
||||
descriptor set, carrying label, id, and availability — so the tab bar is built
|
||||
from data even though the view bodies are static.
|
||||
|
||||
Only `ViewController` writes `active-view`. The seven scattered `set_show_*`
|
||||
call sites become `controller.activate(ViewId::Library)`, which is also the
|
||||
hook where `deactivate` on the outgoing view runs.
|
||||
|
||||
**Verification.** Behavioural parity on every transition currently reachable:
|
||||
launch → library on sign-in, library → develop on cell click, develop →
|
||||
library on back, and the startup paths in `launch::Startup` (all three arms).
|
||||
|
||||
---
|
||||
|
||||
## Sequencing
|
||||
|
||||
| Stage | Depends on | Independently valuable |
|
||||
|---|---|---|
|
||||
| 1 — `ViewController` + `present()` | — | Yes: removes the cycle and the duplication |
|
||||
| 2 — `View` trait | 1 | Yes: uniform availability handling |
|
||||
| 3 — `active-view` | 2 | Yes: illegal states unrepresentable |
|
||||
|
||||
Stage 1 first regardless. The descriptor layer needs a controller to live in,
|
||||
and the `present()` duplication is a live bug source independent of how views
|
||||
are composed.
|
||||
|
||||
## Requirements
|
||||
|
||||
Stages 1 and 2 are traced by existing IDs — FR-UI-6 (shared components),
|
||||
FR-DEV-3a (self-describing), NFR-P9 (no UI-thread blocking). Stage 3's claim,
|
||||
that view composition is data-driven and view identity single-valued, has no
|
||||
requirement covering it. Proposed for `requirements.md` §3.5, after FR-UI-7:
|
||||
|
||||
> **FR-UI-8 — Self-describing views.** Each view publishes a descriptor —
|
||||
> identity, label key, kind, and current availability — and a single
|
||||
> controller composes the interface from the descriptor set. View identity is
|
||||
> single-valued: exactly one primary view is active at a time. A view that
|
||||
> cannot currently function reports why, and the shell presents it as
|
||||
> unavailable rather than omitting it. Adding a view requires no change to the
|
||||
> shell beyond registering it and declaring its body.
|
||||
|
||||
Not added to the register by this document; adding it is a separate edit,
|
||||
since `requirements.md` is the register of record and renumbering there ripples
|
||||
into `traceability.md`.
|
||||
@@ -0,0 +1,487 @@
|
||||
# DarkRoom — A Windows installer from the Linux CI
|
||||
|
||||
**Satisfies:** FR-PLAT-WIN-1 · FR-PLAT-WIN-2 · FR-PLAT-WIN-3 · NFR-COMPAT-2 (a stated channel)
|
||||
**Companion to:** [distribution.md](distribution.md) · [requirements.md](requirements.md) §3.8, §4.4 ·
|
||||
[android-signing.md](android-signing.md)
|
||||
|
||||
Spec for producing `DarkRoom-<version>-x86_64-setup.exe` from the same Gitea runner that builds the
|
||||
Arch package and the APK, with no Windows machine in the loop. It names the toolchain, what the
|
||||
tree has to change to compile for the target, what the installer does, how the CI job is shaped,
|
||||
and — because there is no Windows hardware on the runner — exactly how much of the result can be
|
||||
verified before a person double-clicks it.
|
||||
|
||||
**Written as a spec; §10 is the report.** Every step of §9 has since been run —
|
||||
[`docker/windows/`](../../docker/windows/) is the container, [`packaging/windows/darkroom.nsi`](../../packaging/windows/darkroom.nsi)
|
||||
the installer, [`.gitea/workflows/windows-image.yml`](../../.gitea/workflows/windows-image.yml) and
|
||||
the `windows` job in `build-and-test.yml` the CI leg, and §6's gate passes through row 4 under
|
||||
Wine. Four claims in the first draft were wrong and are corrected in place with a note; §10 lists
|
||||
them. Where a claim still rests on reading rather than running, it says so.
|
||||
|
||||
---
|
||||
|
||||
## 1. Why this is nearly free, and where it is not
|
||||
|
||||
The reason to write this at all is that the tree is closer to Windows than a Linux-only project
|
||||
usually is. The three things that ordinarily make a cross-build to Windows a week of work are all
|
||||
absent:
|
||||
|
||||
| Usual obstacle | Here |
|
||||
|---|---|
|
||||
| A C image library (libraw, libjpeg-turbo, lcms) | `rawler`, `zune-jpeg`, `jpeg-encoder`, all pure Rust |
|
||||
| OpenSSL, or a TLS stack with a system dependency | `reqwest` on `rustls` + `webpki-roots`; `ring` cross-compiles to the GNU target |
|
||||
| A GUI toolkit with a platform-specific build | Slint on `winit` + wgpu, which already runs the same code on Linux and Android |
|
||||
|
||||
`rusqlite` is `bundled`, so SQLite compiles with whatever C compiler the target has; that is the one
|
||||
place a cross C compiler is required, and it is a package install rather than a port. The inference
|
||||
engine is tract, in Rust, which is what [faces.md §3](faces.md) chose it for — and this is the
|
||||
second time that choice pays: the C++ ONNX Runtime would have needed a prebuilt Windows binary
|
||||
fetched at build time.
|
||||
|
||||
**Where it is not free** is `dr-plat` and the handful of paths above it, which is exactly where
|
||||
NFR-PORT-1 says platform code should be and where §3 finds it. That the list in §3 is short and
|
||||
every item on it is already behind a `cfg` is the measure of whether NFR-PORT-3 ("a third platform
|
||||
requires implementing the platform interfaces only") was met. It was, nearly: the gaps are in
|
||||
things that grew *above* `dr-plat` — a settings file path, an `xdg-open` — rather than in the
|
||||
interfaces themselves.
|
||||
|
||||
---
|
||||
|
||||
## 2. Toolchain: the GNU target, from a container
|
||||
|
||||
Two Rust targets can produce a Windows binary from Linux.
|
||||
|
||||
| Target | Linker | What it needs on the runner | What it costs |
|
||||
|---|---|---|---|
|
||||
| **`x86_64-pc-windows-gnu`** | MinGW-w64 `gcc` | `gcc-mingw-w64-x86-64` (Debian/Ubuntu), `mingw-w64-gcc` (Arch) — one apt/pacman install | Binaries link `libgcc_s` and `libwinpthread` unless told not to; the SEH unwinder is MinGW's rather than MSVC's; DirectX bindings are less exercised (not used — §2.1) |
|
||||
| `x86_64-pc-windows-msvc` | `lld-link` via [`cargo-xwin`](https://github.com/rust-cross/cargo-xwin) | The MSVC CRT and Windows SDK headers, fetched from Microsoft's servers by `xwin` on first use (~1.5 GB, licence-accepted by flag) | A download step in CI that depends on Microsoft keeping those URLs stable, and a licence the runner accepts on the project's behalf |
|
||||
|
||||
**Decision: GNU.** It is a package install, it is what `rustup target add` supports out of the
|
||||
box, and every crate in the dependency graph that carries a C component (`ring`, `libsqlite3-sys`,
|
||||
`zstd-sys` if present) builds against MinGW today. The MSVC route produces a marginally more
|
||||
conventional binary — the same CRT every other Windows application links — and costs a
|
||||
1.5 GB fetch of Microsoft-licensed headers on every cold CI run. That is the wrong trade for a
|
||||
channel whose users are, for now, the author.
|
||||
|
||||
Static-link the MinGW runtime so the installer carries one file rather than three. The
|
||||
configuration lives in the container as environment variables rather than in a `.cargo/config.toml`
|
||||
— that file is untracked here on purpose, and the Android image sets its linkers the same way:
|
||||
|
||||
```sh
|
||||
CARGO_TARGET_X86_64_PC_WINDOWS_GNU_LINKER=x86_64-w64-mingw32-gcc-posix
|
||||
CARGO_TARGET_X86_64_PC_WINDOWS_GNU_RUSTFLAGS="-C link-args=-static-libgcc -C link-args=-static-libstdc++"
|
||||
```
|
||||
|
||||
**Corrected.** The first draft added `-Wl,--whole-archive -lwinpthread` "so nothing imports
|
||||
`libwinpthread-1.dll`". That flag breaks the link: forcing the whole archive drags in unused
|
||||
winpthread objects whose kernel32 and msvcrt references land after those libraries on the link
|
||||
line, and the build dies on a hundred undefined `__imp_` symbols. It was also unnecessary —
|
||||
rustc's windows-gnu target links its own winpthread in self-contained mode, and the built
|
||||
executable imports no MinGW library at all (§10). `-posix` is stated because Debian's bare
|
||||
`x86_64-w64-mingw32-gcc` is an alternatives symlink to either thread model.
|
||||
|
||||
Two more things the container needs that the draft did not name: a **host** `gcc`, because build
|
||||
scripts and proc-macros compile for Linux whatever the target and the very first one fails with
|
||||
"linker `cc` not found" without it; and **Wine 10**, because rustc's std imports
|
||||
`bcryptprimitives.dll` for its random source and Debian bookworm's Wine 8.0 does not have it —
|
||||
the smoke test dies at load with `c0000135` before the first instruction. So the image is
|
||||
`debian:trixie-slim`, which also ships Node 20 natively.
|
||||
|
||||
### 2.1 What the binary reaches at runtime
|
||||
|
||||
Nothing the installer has to carry. wgpu opens **Vulkan** (`dr_gpu::new_shared`, D1 — Vulkan on
|
||||
both targets, and the shared-device path permits nothing else), and on Windows the Vulkan loader
|
||||
`vulkan-1.dll` is installed by every GPU vendor's driver. Slint's femtovg renderer finds system
|
||||
fonts through `fontdb`, so the `fontconfig` the Linux CI job installs has no Windows counterpart.
|
||||
There is no `libxkbcommon`, no display-server library: `winit` speaks Win32 directly.
|
||||
|
||||
**DirectX 12 is deliberately not enabled.** wgpu supports it and on Windows would be the
|
||||
conventional choice, but the develop pipeline's compute shaders are written once for Vulkan
|
||||
(NFR-PORT-2) and validated on two Vulkan drivers already; a third backend is a third set of
|
||||
driver behaviours to characterise (NFR-R1's tolerance argument), and no Windows machine that can
|
||||
run this application lacks a Vulkan ICD. The same reasoning that keeps GL out of `new_shared`
|
||||
keeps DX12 out here. It is one flag away if that turns out to be wrong.
|
||||
|
||||
### 2.2 Where it builds
|
||||
|
||||
The same shape as the Android leg: a job container built from a Dockerfile in the tree and pushed
|
||||
to the Gitea registry, tagged by the tree id of its directory so an unrelated push reuses it
|
||||
([`android-image.yml`](../../.gitea/workflows/android-image.yml) already does this and the comment
|
||||
there explains why).
|
||||
|
||||
```
|
||||
docker/windows/
|
||||
Dockerfile debian:trixie + rustup (1.92.0, target x86_64-pc-windows-gnu) + gcc + gcc-mingw-w64-x86-64 + nsis + wine
|
||||
build.sh run a command in the container; caches registry, target and the Wine prefix
|
||||
package.sh the LFS guard, staging, then makensis
|
||||
```
|
||||
|
||||
`makensis` is a Linux binary; NSIS has always been buildable and runnable on POSIX hosts, and
|
||||
Debian ships it as `nsis`. Wine is in the image for §6, not for the build. `osslsigncode` is not
|
||||
in it until there is a certificate to give it (§5.4).
|
||||
|
||||
---
|
||||
|
||||
## 3. What the tree has to change
|
||||
|
||||
Read from the source, not run. Everything here is `dr-plat` or the thin layer above it — nothing
|
||||
in `core/` is touched, which is NFR-PORT-1 holding.
|
||||
|
||||
### 3.1 Already handled
|
||||
|
||||
- **`volumes.rs`** — card detection reads `/proc/mounts` and `/sys/block` under
|
||||
`cfg(target_os = "linux")` and returns an empty list elsewhere. Windows gets no card detection
|
||||
in this pass; the import flow's path picker still works. (A `GetDriveType`/`DRIVE_REMOVABLE`
|
||||
implementation is a screen of code and a follow-up.)
|
||||
- **`display.rs`** — the X11 and Wayland colour-profile readers are `cfg(all(unix, not(android)))`;
|
||||
the fallback is FR-DSP-8's stated one. Windows ICC profiles via `GetICMProfile` are a follow-up
|
||||
for the same reason.
|
||||
- **`desktop_client.rs`** — the Nextcloud desktop client's Unix socket is `cfg(unix)`. On Windows
|
||||
the client listens on a named pipe (`\\.\pipe\...`); until that is implemented FR-NC-6c's
|
||||
integration is absent and the app behaves as it does on a Linux machine with no client running.
|
||||
- **`secrets.rs`** — has a `PlatformSecretStore` for "any platform without an implementation"
|
||||
that returns `SecretError::Unavailable` on every call. It is loud on purpose, so a Windows build
|
||||
made with no further change *compiles*, starts, and fails at sign-in with a clear message. §3.2
|
||||
is what turns that into a working store.
|
||||
- **`keyring`, `x11rb`, `wayland-*`** are target-scoped dependencies already, so the Linux-only
|
||||
crates are not even compiled.
|
||||
|
||||
### 3.2 Required before the installer is worth shipping
|
||||
|
||||
Ordered by what blocks a first sign-in. **All four are done**; each item says how.
|
||||
|
||||
1. **Secret store.** *Done.* `keyring` 4's `v1` feature set — the one the workspace already
|
||||
asks for — includes `windows-native-keyring-store`, so the Credential Manager backend needed
|
||||
no new feature name, only the crate as a `cfg(windows)` target dependency and the existing
|
||||
Secret Service implementation's `cfg` widened to include Windows. One implementation over
|
||||
both, because `keyring::Entry` is the same API over either; the only difference is that
|
||||
`is_available`'s probe always succeeds on Windows, which is correct — Credential Manager is
|
||||
always present, so FR-NC-2's degraded mode does not arise.
|
||||
|
||||
2. **Paths.** *Done* — [`platform/dr-plat/src/dirs.rs`](../../platform/dr-plat/src/dirs.rs). FR-PLAT-LIN-1
|
||||
says XDG, and the code said it in five places by reading `XDG_*_HOME` and falling back to
|
||||
`$HOME/.local/...`. On Windows `HOME` is normally unset, so every one of these degraded to a
|
||||
relative path from the working directory — which for a Start Menu launch is
|
||||
`C:\Windows\System32`. Now one function per kind of directory in `dr-plat`, with the Windows
|
||||
branch reading `%APPDATA%` (config; roams) and `%LOCALAPPDATA%` (data, cache, state; does
|
||||
not), and the five call sites using it. The Android overrides (`set_state_dir`,
|
||||
`set_data_dir`) stay where they were; only the fallback behind them moved. Both platforms'
|
||||
rules are unit-tested on either host, and the Windows one was confirmed by running the
|
||||
application under Wine: the log landed in `AppData\Local\darkroom\state` and nothing was
|
||||
written anywhere else.
|
||||
|
||||
| Kind | Linux today | Windows |
|
||||
|---|---|---|
|
||||
| config (`settings.json`, accounts) | `$XDG_CONFIG_HOME/darkroom` — `dr_sync::config_dir`, `settings_store.rs` | `%APPDATA%\darkroom` |
|
||||
| data (catalog, thumbnails, faces) | `$XDG_DATA_HOME/darkroom` — `library::data_root` | `%LOCALAPPDATA%\darkroom` |
|
||||
| state (crash reports, diagnostics) | `$XDG_STATE_HOME/darkroom` — `state.rs`, `crash.rs` | `%LOCALAPPDATA%\darkroom\state` |
|
||||
|
||||
FR-PLAT-WIN-1 states this as the requirement. The catalog and thumbnail *formats* do not change,
|
||||
so a library directory copied from a Linux machine opens.
|
||||
|
||||
3. **Face models.** *Done.* `library::system_face_models_dirs` walked `$XDG_DATA_DIRS`, which
|
||||
does not exist on Windows. The rule moved to `dr_plat::system_data_dirs`: the installer puts
|
||||
the models beside the executable (§5), so the Windows branch returns the executable's own
|
||||
directory. The user-directory lookups above it are unchanged, so a hand-placed pair still
|
||||
outranks the installed one, exactly as on Linux.
|
||||
|
||||
4. **Opening the sign-in URL.** *Done.* `launch_ui.rs` shelled out to `xdg-open`. The Windows
|
||||
branch runs `rundll32 url.dll,FileProtocolHandler <url>`, which is `ShellExecute` on the URL
|
||||
and needs no crate — chosen over `cmd /C start`, whose quoting of `&` in a query string is a
|
||||
known trap, and over the `open` crate, which would be a dependency for one line. Android has
|
||||
its own Intent path already, so this is the third branch of a function that already had two.
|
||||
*Not verified*: Wine has no browser to open.
|
||||
|
||||
5. **`std::os::unix` uses outside a `cfg`.** `diagnostics.rs` and `presets.rs` use
|
||||
`PermissionsExt` for mode bits on written files. Most are inside `#[cfg(unix)]` blocks already;
|
||||
the first cross-compile will name any that are not, and the fix is a `cfg` rather than a
|
||||
Windows ACL equivalent — the files in question are the user's own.
|
||||
|
||||
6. **The executable's identity.** *Done.* Windows takes the icon and the version block from a
|
||||
resource compiled into the `.exe`, not from a `.desktop` file.
|
||||
[`apps/darkroom-desktop/build.rs`](../../apps/darkroom-desktop/build.rs) uses `winresource`
|
||||
(which invokes MinGW's `windres` when cross-compiling) to embed
|
||||
[`ui/dr-ui/ui/app-icon.png`](../../ui/dr-ui/ui/app-icon.png) — wrapped into an `.ico` in
|
||||
`OUT_DIR` at build time, since an ICO entry may be a PNG, so no generated binary is committed —
|
||||
plus the version from `CARGO_PKG_VERSION` and the product name. The script returns before
|
||||
touching the crate on every other target, and `winresource` is an unconditional
|
||||
build-dependency because **a `cfg(windows)` on a build-dependency is evaluated against the
|
||||
host**, which is Linux. This is the fifth place the identifier lives, and
|
||||
[`tools/set-version.sh`](../../tools/set-version.sh) does not need to learn it: the resource
|
||||
reads the version cargo already knows. The same commit made the release binary a GUI-subsystem
|
||||
executable (`windows_subsystem = "windows"`), or Windows keeps a console window open behind
|
||||
the application.
|
||||
|
||||
Everything in this list is `cfg(windows)` code in `dr-plat` or a call-site switch in `dr-ui`, and
|
||||
none of it touches the image core, the catalog schema, or the edit pipeline. That is the NFR-PORT-3
|
||||
test, and it should be stated in the commit that closes the list whether it passed.
|
||||
|
||||
**What the first cross-compile actually found** (§9 step 2): nothing in this list blocked the
|
||||
link. The whole graph compiled; the only warnings were two constants — `SERVICE` in `secrets.rs`
|
||||
and `TIMEOUT` in `desktop_client.rs` — left unused by the `cfg`s that already shadow their users,
|
||||
now guarded the same way. Items 1–4 are still open, and the binary starts without them; it just
|
||||
cannot sign in.
|
||||
|
||||
### 3.3 Explicitly not in this pass
|
||||
|
||||
- **MIME/file-type registration** — FR-PLAT-LIN-1's `.desktop` MIME entries have a registry
|
||||
equivalent (`HKCU\Software\Classes\.cr2` etc.). Not until the application opens a file from the
|
||||
command line usefully, which `main.rs` accepts but the launch flow does not yet act on.
|
||||
- **High-DPI declaration** — winit sets per-monitor-v2 awareness through its manifest by default.
|
||||
Verified in winit's source, not on a monitor; if text is blurry on a 150% display this is the
|
||||
first suspect.
|
||||
- **Card detection, ICC profiles, the desktop-client pipe** — §3.1's three follow-ups.
|
||||
- **A GL or DX12 fallback** — §2.1. A machine without Vulkan gets the library and no develop
|
||||
path, which is what it gets on Linux too.
|
||||
|
||||
---
|
||||
|
||||
## 4. The build script
|
||||
|
||||
`docker/windows/build.sh`, in the shape of the Android one and with the same rules:
|
||||
|
||||
```sh
|
||||
cargo build --release --target x86_64-pc-windows-gnu -p darkroom-desktop
|
||||
```
|
||||
|
||||
Release only, with `CARGO_TARGET_DIR` inside the workspace so the CI cache key
|
||||
(`windows-${{ hashFiles('**/Cargo.lock') }}`) covers it. The whole workspace is *not* built for
|
||||
the target: `darkroom-android` cannot be, and the examples that need a display or a catalog on
|
||||
disk have nothing to run against. `cargo clippy --target x86_64-pc-windows-gnu -p darkroom-desktop`
|
||||
is worth running in the same job, because the `cfg(windows)` branches from §3 are otherwise
|
||||
never linted — the Linux job cannot see them.
|
||||
|
||||
Not `cargo test --target x86_64-pc-windows-gnu`: the test binaries would be Windows executables,
|
||||
and running them means Wine. §6 does that for exactly one binary, deliberately.
|
||||
|
||||
---
|
||||
|
||||
## 5. The installer
|
||||
|
||||
[`packaging/windows/darkroom.nsi`](../../packaging/windows/darkroom.nsi), compiled by `makensis` on
|
||||
the runner into `DarkRoom-<version>-x86_64-setup.exe`. `package.sh` passes the version in
|
||||
(`/DVERSION=…`, from `tools/set-version.sh`'s single source, the workspace `Cargo.toml`) and refuses
|
||||
to run if any `models/face/*.onnx` is smaller than 100 KB — the LFS-pointer guard every other
|
||||
packager carries, for the reason [distribution.md §1](distribution.md) gives.
|
||||
|
||||
### 5.1 Per-user, not per-machine
|
||||
|
||||
Install to `$LOCALAPPDATA\Programs\DarkRoom`, register the uninstaller under
|
||||
`HKCU\Software\Microsoft\Windows\CurrentVersion\Uninstall\DarkRoom`, `RequestExecutionLevel user`.
|
||||
No UAC prompt, no `Program Files`, no writes outside the user's profile. This is the shape VS Code's
|
||||
"User Installer" and most Electron applications use, and it is right for this project for two
|
||||
reasons: an unsigned installer that also asks for administrator rights is the most alarming thing
|
||||
Windows can show a user (§5.4), and a per-user install means the application's own data directories
|
||||
(§3.2) and its binaries are governed by the same account, which is what NFR-SEC-5's
|
||||
"the user's own hardware" means on a shared machine.
|
||||
|
||||
### 5.2 What it puts on disk
|
||||
|
||||
```
|
||||
$LOCALAPPDATA\Programs\DarkRoom\
|
||||
darkroom.exe
|
||||
models\
|
||||
scrfd_500m_640.onnx scrfd_2.5g_640.onnx scrfd_10g_640.onnx arcface_mbf_b1.onnx
|
||||
2d106det_b1.onnx ocec_s_b1.onnx sgc_l_48_b1.onnx
|
||||
yolo26s-sem-ade20k.onnx yolo26s-sem-ade20k.classes.json categories.txt
|
||||
LICENSE
|
||||
uninstall.exe
|
||||
```
|
||||
|
||||
Plus a Start Menu shortcut, and nothing on the desktop unless the user ticks it. The models are
|
||||
the same ten files the APK bundles and the PKGBUILD installs; `models\` beside the executable is
|
||||
where §3.2's lookup finds them. **No `LICENSE` yet**: the repository has no licence file at its
|
||||
root (the Arch package points at the system's shared GPL text), so the installer has no licence
|
||||
page until one is added — a one-file change, and the `.nsi` says where the page then goes. The face weights carry the research-only grant that
|
||||
[faces.md §2](faces.md) records, and this channel changes nothing about that: the installer is
|
||||
for the author's own machines until §2.2a's caveat is resolved, exactly as the APK is.
|
||||
|
||||
### 5.3 Uninstall
|
||||
|
||||
Removes the install directory, the shortcut and the registry key. **Does not touch
|
||||
`%LOCALAPPDATA%\darkroom` or `%APPDATA%\darkroom`** — the catalog, the thumbnails, the face
|
||||
index, the settings. An uninstaller that deletes a library index the user spent two hours building
|
||||
is the kind of destructive default FR-CULL-12 and NFR-SEC-5's "disabling deletes nothing" both
|
||||
argue against. The uninstaller says so on its one page, and names the two directories so a user who
|
||||
does want them gone knows where they are.
|
||||
|
||||
### 5.4 Signing, and the warning that results from not doing it
|
||||
|
||||
An unsigned installer triggers SmartScreen's "Windows protected your PC" interstitial, dismissable
|
||||
through "More info → Run anyway". Signing needs an Authenticode certificate, which is paid and
|
||||
identity-verified; from Linux the signing itself is `osslsigncode`, which is why it is in the
|
||||
container image, but there is no certificate to give it. **This spec ships unsigned** and the
|
||||
release notes say what the interstitial looks like. An OV certificate is a cost decision to make
|
||||
if this channel ever has a user who is not the author; an EV one buys instant reputation and costs
|
||||
a hardware token. Neither is a build problem.
|
||||
|
||||
The APK went through the same sequence — [android-signing.md](android-signing.md) records a
|
||||
debug-signed build becoming a release-signed one when it mattered — and this channel should be
|
||||
allowed to do the same.
|
||||
|
||||
### 5.5 What NSIS is chosen over
|
||||
|
||||
WiX produces an MSI, which is what enterprise deployment tooling wants and what nobody deploying
|
||||
a photo editor to their own laptop cares about; its Linux story is `wixl` from msitools, which is
|
||||
real but thinly used. Inno Setup runs only under Wine. NSIS is scriptable in plain text, builds
|
||||
natively on Linux, produces a single self-contained `.exe`, and the script for §5.2 is under a
|
||||
hundred lines. It is the conventional answer for exactly this situation.
|
||||
|
||||
One choice inside NSIS: `Target amd64-unicode`, a 64-bit installer rather than the default 32-bit
|
||||
stub. The application is x86_64 only so nothing is lost, and it is what lets §6's install test run
|
||||
under a 64-bit-only Wine — the 32-bit stub needs an i386 multiarch Wine and dies loading the WoW64
|
||||
`ntdll` without one.
|
||||
|
||||
---
|
||||
|
||||
## 6. Verifying without Windows
|
||||
|
||||
This is the part to be honest about. The runner has no Windows, no GPU it can hand to a Windows
|
||||
process, and no display. What *can* be checked, in increasing cost and decreasing certainty:
|
||||
|
||||
| Check | How | What it proves |
|
||||
|---|---|---|
|
||||
| **It links** | the build succeeds | Every `cfg(windows)` branch compiles; no `unix`-only symbol leaked past a `cfg` |
|
||||
| **It is a Windows executable** | `file darkroom.exe` reports PE32+; `x86_64-w64-mingw32-objdump -p` lists the DLLs it imports and none are MinGW's | The static-runtime flags in §2 held |
|
||||
| **It starts** | `wine64 darkroom.exe --version` exits 0 and prints the version | The CRT, the resource block and `main` are sound; paths in §3.2 resolve (Wine sets `LOCALAPPDATA`) |
|
||||
| **The installer runs** | `wine64 DarkRoom-setup.exe /S` then the install directory exists with the eight files, and `wine64 uninstall.exe /S` removes it | The NSIS script's file list, sections and uninstaller are right |
|
||||
| **It draws a window** | `xvfb-run wine64 darkroom.exe` with `SLINT_WGPU_CPU` and a lavapipe ICD exposed through `winevulkan` | That Slint's winit backend initialises on Win32 — and this is where the chain gets long enough that a failure says more about Wine than about DarkRoom |
|
||||
|
||||
The first four are the CI gate. The fifth is worth trying once by hand and not putting in CI:
|
||||
it needs `winevulkan` to find a host ICD, `xvfb`, and a Wine prefix warmed up in the container,
|
||||
and every one of those is a moving part that has nothing to do with whether the application works
|
||||
on Windows.
|
||||
|
||||
`--version` exists for this — a smoke test needs an exit that opens no window and touches no
|
||||
directory — and it is answered before the logger and the crash hook install, so it proves the CRT
|
||||
and the resource block and nothing above them.
|
||||
|
||||
**One more row the table missed:** the Start Menu shortcut. `CreateShortcut` is `IShellLink`,
|
||||
which does nothing under a headless Wine while the `CreateDirectory` beside it succeeds, so an
|
||||
installer that installs and uninstalls cleanly here can still have a broken shortcut. That row is
|
||||
on Windows only.
|
||||
|
||||
**What none of this proves:** that wgpu opens a Vulkan device on a real driver, that a 6000-px
|
||||
render completes, that fonts are found, that the secret store round-trips. Those are a person with
|
||||
a Windows machine, once per release, until there is a Windows runner — and a self-hosted Windows
|
||||
act_runner is how that would be done, not a cloud service. The release notes for the first build
|
||||
say which of these were checked and on what.
|
||||
|
||||
---
|
||||
|
||||
## 7. The CI job
|
||||
|
||||
A fourth leg of [`build-and-test.yml`](../../.gitea/workflows/build-and-test.yml), beside desktop,
|
||||
Android and traceability:
|
||||
|
||||
```yaml
|
||||
windows-image:
|
||||
uses: ./.gitea/workflows/windows-image.yml # same shape as android-image.yml
|
||||
|
||||
windows:
|
||||
runs-on: linux/amd64
|
||||
name: Windows (x86_64, cross)
|
||||
needs: windows-image
|
||||
container:
|
||||
image: gitea.tourolle.paris/dtourolle/darkroom-windows:latest
|
||||
steps:
|
||||
- checkout, LFS pull # copied from the desktop leg
|
||||
- cache: ~/.cargo, target # key: windows-${{ hashFiles('**/Cargo.lock') }}
|
||||
- docker/windows/build.sh # cargo build + clippy, --target x86_64-pc-windows-gnu
|
||||
- smoke: file, objdump, wine64 --version # §6 rows 1–3
|
||||
- docker/windows/package.sh # LFS guard, makensis
|
||||
- smoke: wine64 setup.exe /S; ls; uninstall # §6 row 4
|
||||
- upload artefact: DarkRoom-*-setup.exe # on tags only, like the APK
|
||||
```
|
||||
|
||||
Same gotchas as the Android leg, which its comments already record: the host has no Node, so the
|
||||
checkout is plain `git`; workflow inputs arrive as strings; the image build needs the host Docker
|
||||
daemon and runs outside a container. None of that is new.
|
||||
|
||||
**Cost.** A cold build of the whole graph for a second target is roughly the desktop leg again —
|
||||
tract, Slint's compiler, wgpu — so with the cargo cache warm it is minutes and cold it is the
|
||||
better part of half an hour. Worth noting because the runner is one machine and the legs run in
|
||||
parallel on it; if it starts starving the desktop leg, `needs: desktop` serialises them.
|
||||
|
||||
---
|
||||
|
||||
## 8. Requirements
|
||||
|
||||
Three, added to [requirements.md §3.8](requirements.md) under a `#### Windows` heading beside the
|
||||
Linux ones. Phrased to be testable, and each one is something §3 or §5 would otherwise leave as a
|
||||
convention.
|
||||
|
||||
**FR-PLAT-WIN-1 — Known folders.** Configuration under `%APPDATA%\darkroom`; data, cache and
|
||||
state under `%LOCALAPPDATA%\darkroom`. No file under the user's profile root and nothing relative
|
||||
to the working directory. The directory *layout* beneath those roots is the same as under XDG, so
|
||||
a library directory moves between platforms unchanged.
|
||||
|
||||
**FR-PLAT-WIN-2 — Installer.** A per-user installer that needs no elevation, registers an
|
||||
uninstaller, and whose uninstaller removes what the installer wrote and nothing the application
|
||||
wrote. Models are installed beside the executable and found there last, after the user's own
|
||||
directories.
|
||||
|
||||
**FR-PLAT-WIN-3 — Built from Linux.** The Windows binary and its installer are produced by the
|
||||
Linux CI from the same commit as every other channel, with no Windows machine in the build.
|
||||
Verification on Windows is a release step, recorded per release, not a build step.
|
||||
|
||||
NFR-COMPAT-2's channel table in [distribution.md §1](distribution.md) gains a row. NFR-PORT-3 gets
|
||||
its first real test, and the commit that closes §3.2 records the answer.
|
||||
|
||||
---
|
||||
|
||||
## 9. Order
|
||||
|
||||
1. `--version` in `main.rs`, and the `.cargo/config.toml` target block. Trivial, and the smoke
|
||||
test in §6 needs both before anything else can be measured.
|
||||
2. `rustup target add x86_64-pc-windows-gnu`, `pacman -S mingw-w64-gcc`, and a first
|
||||
`cargo build --target …` on the developer machine — **before the container exists**, because
|
||||
the list in §3.2 is a reading of the source and the compiler's list will be longer. Fix the
|
||||
`cfg` fallout as it appears. This is the afternoon that decides whether §1's optimism holds.
|
||||
3. §3.2 items 1–4, each its own commit, each stating which NFR-PORT interface it implemented.
|
||||
4. §3.2 item 6 — the resource block — and the NSIS script; `makensis` by hand; `wine64 setup.exe /S`
|
||||
by hand. Now there is an artefact.
|
||||
5. The container, the image workflow, the CI leg. Only after 4 works locally, for the same reason
|
||||
the Android image was reproduced from the tree after it had lived on one laptop.
|
||||
6. A build on a real Windows machine, and a note in the release saying what was checked.
|
||||
|
||||
Steps 1–2 are cheap and either confirm this document or replace §3.2 with the true list. Nothing
|
||||
past step 2 should be started on the strength of this document alone.
|
||||
|
||||
---
|
||||
|
||||
## 10. Report · 2026-09-12
|
||||
|
||||
Steps 1, 2 and 4 run, in the container rather than on the developer machine, because the
|
||||
container was the cheaper way to get a pinned MinGW and a Wine that could be thrown away.
|
||||
|
||||
| §6 row | Result |
|
||||
|---|---|
|
||||
| It links | Yes, first attempt once the link flags were right. 115 MB, `PE32+ … (GUI)`. Two dead-code warnings, both `cfg`-shadowed constants, fixed. |
|
||||
| It is a Windows executable | 26 imports, all Windows system DLLs. No MinGW runtime. `.rsrc` carries `PRODUCTVERSION 0,12,0,0`, `ProductName DarkRoom`, the icon. |
|
||||
| It starts | `wine darkroom-desktop.exe --version` → `darkroom-desktop 0.12.0`, exit 0, 0.1 s. |
|
||||
| The installer runs | `makensis` → 105 MB. `/S` installs the exe and ten models to `AppData\Local\Programs\DarkRoom`, writes the `HKCU` uninstall key; the installed exe runs; `uninstall.exe /S` removes directory and key. Shortcut unverifiable (§6). |
|
||||
| It draws a window | Not attempted. |
|
||||
|
||||
**What the first draft got wrong**, kept in place above with a note rather than rewritten, because
|
||||
the reasoning that produced each mistake is the thing a reader will otherwise repeat:
|
||||
|
||||
1. The `--whole-archive -lwinpthread` link flag (§2) — breaks the link and was never needed.
|
||||
2. No host C compiler in the image (§2) — build scripts are host binaries.
|
||||
3. Debian bookworm's Wine (§2) — lacks `bcryptprimitives.dll`, which rustc's std imports.
|
||||
4. A `cfg(windows)` on the `winresource` build-dependency (§3.2 item 6) — evaluated against the
|
||||
host, so the crate was silently absent from the cross-build.
|
||||
|
||||
And two things it did not know to say: NSIS's default stub is 32-bit (§5.5), and `CreateShortcut`
|
||||
cannot be verified headless (§6).
|
||||
|
||||
**Closed since**, same day: all four §3.2 items (each says how), a `LICENSE` at the repository
|
||||
root so the installer has its licence page, and §7's CI leg — `windows-image.yml` and the
|
||||
`windows` job, every step of which was run by hand in the same container first. The Windows
|
||||
target is also linted now, with `cargo clippy --target x86_64-pc-windows-gnu -- -D warnings` in
|
||||
that job, which is the only place the `cfg(windows)` branches are ever compiled by CI.
|
||||
|
||||
**What the first real Windows run has to check**, in order, because Wine cannot: that a Vulkan
|
||||
device opens on a real driver and a render completes; that fonts are found; that Credential
|
||||
Manager round-trips a sign-in and the browser opens for Login Flow v2; that the Start Menu
|
||||
shortcut exists; and that text is sharp on a scaled display (§3.3). The release notes for the
|
||||
first build should say which of these were checked and on what machine.
|
||||
Reference in New Issue
Block a user