Files
DarkRoom/docs/requirements.md
T
dtourolle 696bafa9d5 Undefer AI subject masking, which shipped, and give it a clause
§7 still listed "AI subject masking — deferred per D11" while
MaskSource::Subject and MaskSource::Category, backed by dr-segment's
instance and semantic models, had been the primary way a local
adjustment is made for weeks. The code was tagged FR-DEV-3, which
names gradients and brushes and says nothing about a model.

FR-DEV-3i now states what exists: a subject or a category found by a
local model, stored as identity with the run's signature so that it
merges per field and reads as stale rather than wrong, then treated as
any other layer by the edge, stroke, composition and reveal clauses.
The one place it departs from FR-DEV-19 — coverage written run-length
coded beside the layer, so a stored subject renders without a model —
is recorded in the clause instead of left for the next audit to find.
The segmentation crate and the UI's selection module are tagged to it.
2026-09-19 12:25:03 +02:00

2321 lines
149 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# DarkRoom — Requirements Specification
**Status:** Living document · first written 2026-08-08 · audited 2026-09-19
**Owner:** Duncan Tourolle
A cross-platform, non-destructive RAW photo editor for Linux desktop and Android, in the
Lightroom idiom: a catalog of many thousands of images, a develop module with GPU-accelerated
adjustments, and export to standard 8-bit (or higher) deliverables.
---
## 1. Scope and intent
### 1.1 What this is
DarkRoom is a photo *library* and *develop* application. It manages large collections of camera
RAW files, renders them to screen with GPU acceleration, applies non-destructive edits stored as
metadata, and exports finished images.
### 1.2 Primary platforms
| Platform | Priority | Notes |
|---|---|---|
| Linux desktop | Primary | X11 and Wayland. Development and reference platform. |
| Android | Primary | A 12-inch tablet (D15). Phones are not a target: the build runs on one, and nothing is designed for one. Shares the image core. |
Other platforms (Windows, macOS, iOS) are explicitly out of scope for v1, but the architecture
must not foreclose them. In practice this means the GPU abstraction and the image core must not
hard-code Vulkan-only or Linux-only assumptions at their public interfaces.
### 1.3 What this is not
- Not a DAM with server-side multi-user collaboration
- Not a pixel editor (no layers, no brushes in v1 beyond local-adjustment masks)
- Not a printing/soft-proofing suite in v1
### 1.4 Architecture
The technology stack, crate layout, and design are specified separately in
[architecture.md](architecture.md). This document states *what* the software must do;
the architecture document states *how*.
Requirements here reference architectural constraints as `ARCH §n` where the constraint
materially shapes what is testable.
---
## 2. Core requirements (from the brief)
These are the user's stated requirements, restated as testable criteria.
| ID | Requirement | Acceptance criterion |
|---|---|---|
| **R1** | Cross-platform: Linux + Android | Same image core compiles and runs on both. For the same input and edit graph, output is **perceptually identical within a bounded tolerance** — see below. |
| **R2** | Efficient display of huge RAW libraries | A 50,000-image catalog scrolls at 60fps sustained, with a stated prefetch margin and cache-hit rate sufficient that no cell renders as a placeholder at a scroll velocity of *(figure TBD)* rows/second. Catalog opens in under 2s. |
| **R3** | Make a RAW beautiful at 8-bit output | Full non-destructive develop chain at high internal precision, with camera input profiles (FR-DEV-3e) and a colour-managed path to 8/16-bit export. |
| **R4** | HW acceleration and parallelism | All per-pixel work runs on GPU compute. CPU work (decode, I/O) is parallelised across cores. The UI executor never blocks on image work (NFR-ARCH-1). |
| **R5** | Work on downscaled proxies for display | Display pipeline operates at viewport resolution, not source resolution. The tiling clauses this criterion used to carry have been struck — see below. |
| **R6** | Nextcloud integration | Browse, download, and upload images and edit metadata against a Nextcloud instance, offline-capable. |
| **R7** | Judge anywhere, on evidence | Rating and flag are reachable from every view that shows a photograph, and apply to the one on screen. Everything the app computes about a frame in aid of culling — clipping, focus, burst membership, per-face state — is shown as evidence the photographer reads, and **no code path writes a rating or flag without a user action** (FR-CULL-13). |
**On R7, added 2026-09-19.** Stated from use rather than from the brief. Two things prompted it.
Judgement keys had been built into the grid alone, so a photograph opened in develop — the view a
photographer is most sure about — could not be rated without leaving it; the rule is that judging
follows the photograph, not the view. And specifying per-face signals (FR-CULL-8a) forced the
question of what a signal is *for*, which sharpened what this document already said in FR-CULL-5:
the app may know a great deal about a frame and may say all of it, and it never holds the pen.
R7 is the user-level statement; FR-CULL-13 is the testable one.
**On R1's tolerance.** An earlier draft required output to be *bit-identical* across platforms.
That is not achievable and the requirement has been corrected. Floating-point compute results
differ between GPU vendors: transcendental function implementations vary, drivers apply different
optimisations, and f16 rounding diverges. A checksum comparison across Adreno and Mesa would fail
for reasons that have nothing to do with correctness.
R1 is therefore stated as a bounded tolerance — a defined maximum per-pixel deviation, expressed in
ΔE2000 for colour or ULPs at the working precision. **The threshold must be fixed before spike S9**,
because S9 both validates R1 and calibrates what the achievable tolerance actually is.
Where genuine bit-identity is required — cache keys, edit-graph hashing (§5.2 invariant 3) — it
applies to *integer* operations on CPU-side state, which are deterministic, never to GPU float
results.
**On R5's tiling.** An earlier draft added two clauses to R5's criterion: *"only visible tiles are
computed; panning recomputes only newly exposed tiles"*. They have been struck, and the reason is
the one [frame-budget.md](frame-budget.md) measured — the same measurement that argues FR-DSP-2
should be rewritten rather than implemented. FR-DSP-2 itself has **not** been rewritten: it
stands as written until spike S6 has run on constrained Android hardware, because the desktop
measurement cannot speak for a device whose GPU memory the image exceeds (decided 2026-09-19;
see the note under FR-DSP-2).
R5's actual demand is met and tested. The display pipeline works at viewport resolution:
`Framing::view` shrinks the sampled region while the render target keeps its size, so zooming raises
the resolution the pipeline works at rather than magnifying pixels already drawn, and
`core/dr-gpu/tests/zoom_resolution.rs` establishes it as a pixel equality rather than an impression
of sharpness.
Tiling is a different claim, and it was written in as though it were the mechanism by which the
first one is achieved. It is not. Recomputing the *entire* 4K viewport costs 4.5 ms of a 16 ms
budget, so a perfect tile cache saves at most that, in exchange for a cache keyed by
`(VersionId, tile, zoom, graph_hash_prefix)` that has to stay correct across every parameter change
in the graph — a large correctness surface bought with a small number. And for the one stage that
does miss the budget, tiling makes it worse: that stage is a convolution, and a tiled convolution
reads a halo per tile, so at the 52 px radius measured at 4K a 256 px tile would read (256+104)²
taps instead of 256², very nearly twice the work.
The intent behind the struck clauses — that the display path must not do work proportional to the
source image — survives in the clause that remains, which is the honest statement of it. Tiling
stays where FR-DSP-2 puts it: a scheduling concern for export and thumbnailing, both of which
already run off the frame path.
---
## 3. Functional requirements
### 3.1 Catalog and library management
**FR-CAT-1 — Scan.** The app shall scan one or more user-granted library roots for supported image
files, recursively, without blocking the UI. Progress is reported and the scan is cancellable and
resumable. A "root" is a platform-specific grant (a directory on Linux, a persisted document tree
on Android — see FR-PLAT-AND-1), not necessarily a filesystem path.
**FR-CAT-1a — Source addressing.** The catalog and decode layers shall address source data through
an opaque `SourceRef` that resolves to a seekable byte stream, **never through a filesystem path**.
Android's Storage Access Framework provides no usable path (ARCH §6.9), so a path-based API would not be
portable. `SourceRef` carries enough information to re-resolve after an app restart or a permission
re-grant.
**FR-CAT-2 — Catalog store.** All catalog metadata (source references, EXIF, ratings, labels, edit
graphs, sync state) is stored in a local embedded database. The database is the source of truth for
the UI; sources are scanned into it, never queried directly on the UI path.
**FR-CAT-3 — Thumbnail pyramid.** For each image the app maintains a cached multi-resolution
thumbnail set. Initial thumbnails are extracted from the RAW's embedded JPEG preview where
present (fast path, no demosaic). Higher-quality proxies are generated lazily from the full
decode when the image is first opened in develop.
**FR-CAT-4 — Virtualised grid.** The library grid shall render only visible cells plus a small
prefetch margin. Memory use is bounded and independent of catalog size.
**FR-CAT-5 — Metadata.** Read EXIF, camera make/model, lens, capture time, ISO/aperture/shutter,
GPS. Support user-assigned star ratings, colour labels, flags, and keywords.
**FR-CAT-6 — Search and filter.** Filter the catalog by any indexed metadata field, rating,
label, folder, and keyword, with results updating interactively on a 50k catalog.
**FR-CAT-7 — Collections.** User-defined collections that reference images without moving files.
**FR-CAT-8 — Sidecar persistence, independent of sync.** Edit graphs shall be written to per-image
sidecars for **all** catalogued images, whether or not a Nextcloud account exists. Invariant 5.2.4
(catalog rebuildable from sources plus sidecars) otherwise fails for local-only users, leaving every
edit in a single SQLite file with no recovery path.
Where the source location is not writable — read-only mounts, and commonly Android SAF trees —
sidecars are written to an app-managed store keyed by `SourceRef`, and the app shall state which
location is in use.
**FR-CAT-9 — Offline and relocated sources.** An image whose source is unreachable shall be marked
*offline*, never silently removed. Cached previews, metadata, ratings, and edits remain browsable
and editable while offline; edits queue and apply when the source returns.
The app shall support folder-level and image-level reconnection, matching candidates by content
hash and filename, and shall auto-reconnect a volume or tree when it reappears. A source deleted
outside the app shall be distinguished from one merely unreachable before any destructive catalog
action is offered.
This matters more than it appears: external drives, SD cards, and network mounts disappear
routinely, and on Android a tree permission can be revoked or lost on reinstall.
**FR-CAT-10 — Import and ingest.** Copy or move files from a source volume into a destination
structured by a date/metadata template, with rename-on-import, an optional simultaneous
second-destination backup copy, and per-file verification against a checksum. Removable-volume
insertion is detected where the platform permits.
Distinct from FR-CAT-1: scanning catalogues files where they already are; import moves them from a
card into the library. Both are needed.
**FR-CAT-11 — Duplicate detection.** Detect duplicates on import by (capture time + camera serial +
original filename) and by content hash, offering skip or import-as-new. Camera filenames wrap at
`IMG_9999`, so filename alone is insufficient. Existing catalog duplicates are detectable on demand.
**FR-CAT-12 — Versions (virtual copies).** An image may carry multiple named `Version`s, each with
an independent edit graph, without duplicating source data. Versions are creatable, nameable,
deletable, and independently exportable; one is the default.
**FR-CAT-13 — XMP interoperability.** Read and write standard XMP sidecars for ratings, colour
labels, keywords and hierarchical subjects, title, description, copyright, and GPS, using standard
`xmp:`/`dc:`/`lr:` schemas so other tools interoperate.
DarkRoom's edit graph lives in a private namespace and shall neither be interpreted by, nor
corrupt, other tools' XMP. Writing to source-adjacent XMP is off by default (NFR-R4). External
modification of an XMP sidecar shall be detected and a metadata reload offered.
**FR-CAT-14 — Migration import.** Import ratings, labels, keywords, and collections from a
Lightroom `.lrcat` and a darktable `library.db`. Edit graphs are explicitly **not** migrated —
develop parameters do not translate meaningfully between pipelines, and a partial translation is
worse than none. This is the path in for users with existing libraries.
**FR-CAT-15 — Trash and permanent delete.** Deleting an image shall be reversible by default. A
soft delete **moves the file** into a `.darkroom-trash/` folder under the library root and records
in the catalog when it was trashed and the path it came from; restore moves it back to that path.
Permanent delete removes the file first and the catalog row second, and a delete of something
already gone counts as success.
A flag alone would not survive invariant 5.2.4: the catalog is rebuildable from sources, so a
rescan would find every "deleted" file still in the library and re-index it. The folder is the
durable fact and the row is the convenience — which also means the scanner shall exclude the trash
folder, and that a user can recover by hand without DarkRoom. Derived data keyed on the file
(thumbnails, cached previews) is dropped when the image is permanently deleted, not when it is
trashed.
The trash shall be listable newest-first, with the count and total bytes it holds shown before any
destructive action, since that figure is what tells the user whether they meant it.
### 3.2 RAW decoding
**FR-RAW-1 — Format support.** Decode mainstream RAW formats. Minimum launch set: Canon (CR2,
CR3), Nikon (NEF), Sony (ARW), Fujifilm (RAF, including X-Trans), Panasonic (RW2), Olympus (ORF),
Adobe DNG. Additional formats are a coverage goal, not a launch blocker.
**FR-RAW-2 — Decoder abstraction.** RAW decoding takes **bytes**, never a filesystem path and never
a reference it would have to resolve. Resolving a `SourceRef` (FR-CAT-1a) to bytes is
`Storage::open`'s job and happens at the caller, so the same decoder works over a local file, an
Android SAF document, or a byte range fetched from Nextcloud. A second implementation may be added
for broader camera coverage without changing callers (D2).
*On the change of mechanism.* This clause used to require "a trait taking a `SourceRef`". The
purpose — that no decoder API takes a path, so nothing in the decode path assumes a filesystem —
is met and is not in question: `dr_decode::decode` takes `&[u8]`, and there is no path-based entry
point in the crate. The mechanism was wrong, and stating it that way would have made the design
worse.
A `SourceRef` is opaque by construction; the only thing that turns one into readable bytes is
`Storage`, in `platform/dr-plat`. A decoder taking a `SourceRef` would therefore have to take a
`Storage` alongside it, which moves retry, permission loss and remote fetching inside the decoder
and leaves it constructible only where a `Storage` exists. Bytes in, image out, is both narrower and
more portable: the decoder has no idea where its input came from, which is the property this
requirement is actually asking for.
It also serves the Nextcloud case better rather than worse, which is the one a byte-oriented API
looks like it would lose. `dr_decode::HEADER_BYTES` declares how much of a file the decoder needs to
read metadata, and `import.rs` fetches exactly that range through `Storage::read_range` before
calling `dr_decode::metadata`. The decoder states its requirement and the storage layer satisfies
it; a decoder holding its own `SourceRef` would have had to implement the range policy itself.
What is genuinely not built is the trait. There is one decoder, reached through free functions, so
"without changing callers" is a claim nothing yet tests. The clause stands as written and is
outstanding work, not a satisfied one.
**FR-RAW-3 — Sensor data handling.** Correctly apply per-camera black/white levels, CFA pattern
identification, and camera-native colour matrices. Demosaic quality shall be selectable, with at
least a fast method for preview and a high-quality method for export (FR-EXP-9 requires export to
use the latter).
**FR-RAW-4 — Robustness.** A malformed or hostile RAW file shall not crash the application or
compromise the process. Decode failures are reported per-file and do not abort a batch.
**FR-RAW-5 — X-Trans as a first-class path.** Per D11, Fujifilm is explicitly targeted:
- **Markesteijn-class demosaic as the default** for X-Trans sensors, not an opt-in advanced setting
- **X-Trans-aware sharpening**, since the non-Bayer CFA responds differently
- In-RAF film simulation tag read and matched (FR-DEV-3f)
This targets the market's best-documented colour grievance. Adobe's X-Trans "worms" artefact is a
decade-old unresolved complaint; darktable and RawTherapee have the better algorithm but poor
defaults; Capture One has the best film-simulation support but drops X-Trans I and II.
*Cost to note:* X-Trans demosaic is documented at **at least 2× the processing cost of Bayer**,
which affects the NFR-P4 and NFR-P7 budgets for Fuji files specifically.
### 3.3 Develop pipeline
**FR-DEV-1 — Non-destructive edit graph.** All edits are stored as parameters in an ordered edit
graph attached to the image. Source files are never modified. Any rendered output is reproducible
from source + graph.
**FR-DEV-2 — Internal precision.** The pipeline operates internally at a minimum of 16-bit float
per channel in a wide-gamut linear working space. Quantisation to the output bit depth happens
once, at the final export or display stage.
**FR-DEV-3 — Adjustment set (v1).**
- White balance (temperature/tint, and picker)
- Exposure, contrast
- Highlights / shadows / whites / blacks recovery
- Tone curve (RGB and per-channel)
- HSL / colour mixer per colour band
- Vibrance and saturation
- Texture / clarity
- Sharpening and noise reduction (luminance and chroma)
- Lens corrections: distortion, chromatic aberration, vignetting — driven by the lens profile the
file's own EXIF matches in the bundled Lensfun database where there is one, and by hand where
there is not. The profile is a control rather than a silent step: a photograph whose lens is
recognised carries one switch that accepts or declines the measurement, and the manual sliders
trim whatever it leaves. A photograph whose lens is unrecognised is told so in words and offered
no switch, because a correction that looks available and does nothing is worse than one that is
visibly unavailable
- Crop, straighten, rotate, flip
- Local adjustments: linear gradient, radial gradient, and brush masks
**FR-DEV-3a — Self-describing operations.** Every processing operation shall declare its own
parameters through a descriptor, so that adding an operation requires no changes to frontend code.
An operation declares *what* its parameters are; the frontend decides *how* to present them.
```rust
pub trait Operation: Send + Sync {
/// Static description of this op's parameters. Drives UI generation.
fn descriptor() -> OpDescriptor where Self: Sized;
/// Parameter values → GPU work. No UI types cross this boundary.
fn encode(&self, enc: &mut ComputeEncoder, ctx: &TileContext);
/// Identity for cache invalidation (see §5.2 invariant 3).
fn params_hash(&self) -> u64;
}
pub struct ParamDescriptor {
pub id: ParamId,
pub label: LocalizedString,
pub kind: ParamKind,
pub default: ParamValue,
pub affects: Affects, // Geometry | Colour | Detail — drives invalidation scope
}
pub enum ParamKind {
/// Ordinary numeric parameter. Frontend picks slider / drag-strip / dial by modality.
Scalar { min: f32, max: f32, scale: Scale, unit: Unit, precision: u8 },
Bool,
Enum { variants: Vec<(EnumId, LocalizedString)> },
Colour { has_alpha: bool },
/// Escape hatch: a control that does not reduce to a primitive.
/// The frontend owns the implementation; the op only names the kind
/// and defines the data it exchanges.
Custom { widget: WidgetKind, data: CustomParamSchema },
}
pub enum WidgetKind {
ToneCurve, // per-channel curve editor
ColourWheel, // colour grading wheels
CropOverlay, // on-canvas crop and straighten handles
GradientHandle, // on-canvas linear/radial mask placement
BrushMask, // on-canvas brush strokes
WhiteBalancePick, // eyedropper bound to canvas
}
```
**The pipeline crate shall not depend on the UI toolkit.** Descriptors carry data, never widgets.
This keeps the edit chain testable headless (see §9's golden-image tests, which must link no UI)
and is what allows one operation to render differently on touch and desktop.
**FR-DEV-3b — Frontend presentation mapping.** The frontend maps `ParamKind` to a concrete control
based on input modality and available space. The same descriptor yields different presentations:
| `ParamKind` | Desktop | Touch (tablet) |
|---|---|---|
| `Scalar` | Slider with numeric entry, scroll-wheel fine adjust | Large drag-strip, double-tap to reset, no keyboard entry |
| `Bool` | Checkbox | Switch, minimum 44pt target |
| `Enum` | Dropdown | Segmented control or sheet |
| `Colour` | Swatch opening a picker popover | Swatch opening a full-width sheet |
| `Custom` | Frontend-supplied control for that `WidgetKind` | Same control, touch-tuned hit targets |
**FR-DEV-3c — Operation registry.** Operations register themselves at startup. The develop panel is
generated by walking the registry, so a new operation appears in the UI without any frontend
change. Registration order defines default pipeline order; the ordering itself is data, not code.
**FR-DEV-3d — Invalidation scope.** Each parameter declares what it `affects`, so a change
invalidates only the necessary part of the pipeline. Adjusting exposure shall not re-run lens
correction or re-tile geometry. This is what makes FR-DSP-3's one-frame slider response achievable.
**FR-DEV-3e — Camera input profiles.** The pipeline shall include a camera-profile stage between
demosaic and the working-space conversion.
**v1 scope** (per D11 — good defaults rather than exhaustive colour science):
1. Embedded DNG `ColorMatrix1/2` and `ForwardMatrix1/2` tags
2. A hand-tuned base curve per launch camera body, shipped with the app
3. HaldCLUT import (FR-DEV-3f)
**Deferred but not foreclosed:** full `.dcp` support with `HueSatDeltas`, `ProfileLookTable`, and
dual-illuminant interpolation. The stage shall be structured so these are additions rather than a
pipeline reordering.
Rationale for the reduced scope: a bare 3×3 matrix produces the flat, poor-skin-tone rendering
characteristic of dcraw defaults, which is the documented reason people abandon darktable in the
first hour. A per-body base curve fixes most of that at a fraction of the cost of a full DCP
implementation. The profile database ships **versioned independently of the app binary** so bodies
and curves can be added without a release — and, under D8's GPLv3, contributed by users.
*Acceptance:* for each launch body, the default render is subjectively comparable to the camera's
own JPEG. ΔE2000 validation against ColorChecker references applies once DCP support lands.
**FR-DEV-3f — Look emulation.** Support HaldCLUT import, which inherits the existing free film
simulation ecosystem at near-zero implementation cost, plus reading the in-RAF film simulation tag
to auto-apply a matching render for Fujifilm files.
**Spectral film simulation, in addition rather than instead** (`dr-film`). Where a stock's
measurements exist, simulate the physics instead of replaying a grade: spectral sensitivity exposes
three emulsion layers, characteristic curves develop them to densities, dye densities absorb, and a
paper profile prints the negative with the enlarger's filtration solved rather than dialled. A
scanned negative is therefore orange and inverted, because that is what a negative is.
Two things this buys that a LUT cannot. The parameters stay **physical** — opening up a stop moves
the picture along the film's own curve, shoulder and all, rather than scaling a number baked at one
exposure. And the **data cost inverts**: a stock is ~17 kB of published measurements where one
HaldCLUT is ~800 kB of one person's grade.
A film simulation is a *rendering*, not an adjustment, so it replaces the camera profile's base
curve and the conversion out of camera space (`Operation::renders`) — applying both would render
the scene twice.
*Acceptance:* a neutral scene printed through a colour negative's own paper renders neutral to
within 0.06 in linear sRGB; the baked lookup's interpolation error stays under one 8-bit code
value; and the shader agrees with the CPU model, which agrees in turn with an independent
reference implementation.
**Open:** how the chosen stock persists. Sidecar parameters are `f32` and the stock list is
data-driven, so neither an index nor a name fits the existing shape.
**FR-DEV-3g — AI denoise.** Learned denoising operating in the raw domain, ideally jointly with
demosaic.
Promoted into v1 scope per D11. The reasoning: unlike AI masking, denoise has **no manual fallback**
— it reaches a quality ceiling no conventional method matches, which is why photographers run a
second application for it. Raw-domain joint demosaic-and-denoise is also markedly easier to build
into a new pipeline than to retrofit, and the same component attacks the X-Trans artefact problem
(FR-RAW-5).
Inference is **local only** — no cloud, no telemetry (NFR-SEC-4). The stage is optional at runtime
and its absence degrades gracefully.
**FR-DEV-3h — Stored orientation is honoured, not edited.** An image shall be shown the way the
photograph was taken, from its EXIF orientation tag (`0x0112`), everywhere it appears: the grid's
thumbnails, the develop canvas, and the read-only preview shown when no decoder can open the file.
The tag shall be applied as a property of **reading the file**, at the same standing as a RAW's
masked-photosite crop (FR-RAW-3) — never as an edit. Concretely:
- Opening a frame the camera stored sideways shall not mark it modified, shall not enable the
framing reset, and shall write nothing to its sidecar.
- "Reset framing" shall return the image to *upright*, not to the sensor's scan order.
- A sidecar shall never carry the orientation. Edits are shared between devices and bodies
(FR-NC-9); one camera's sensor scan must not be applied to another's file.
- A user's own quarter turns compose *on top* of it, so one press of the rotate button moves the
image by 90° whatever the file's baseline.
A file carrying no tag, or a value outside 1..=8, is displayed as stored. Guessing would turn a
missing tag into a visibly wrong image, and most files have no tag.
*Acceptance:* a portrait frame from a phone or a body held sideways appears upright in the grid and
in develop with no user action, and its sidecar is byte-identical to that of the same frame shot in
landscape.
**FR-DEV-3i — Model-found masks.** A mask layer's part may be a thing a model found in the
photograph, selected by pointing at it: one **subject** — this dog, not that one — from an
instance model, or one **category** — all the sky, all the foliage — from a semantic model. The
selection is stored as *identity* (which run, which instance or which category name, with the
run's signature) so that two devices selecting the same subject hold the same value and merge per
field under FR-NC-9, and a layer whose run no longer matches reads as **stale** rather than
silently masking something else. The mask arrives approximately right and soft, and it is then an
ordinary layer: FR-DEV-3's edge treatment shapes it, FR-DEV-19b's strokes correct it, FR-DEV-19a
composes it with a gradient or a range, and FR-DEV-19c shows it. Inference is local
(NFR-SEC-4); the models and their grants are named in `models/LICENCE.md` and D14.
*Added 2026-09-19, because it was built and §7 still said it was deferred.* The deferral treated
subject masking as an AI feature to be copied from Adobe or darktable; what was built is the
shape §7's own note asked for — a mask that behaves like a hand-drawn one — and it has been the
primary way a local adjustment is made since the watershed hierarchy failed on real photographs
(`docs/segmentation.md` §15).
One departure from FR-DEV-19 is recorded rather than hidden. A model-found part's coverage is
also written to the sidecar, run-length coded beside the layer, because a stored subject that is
"reproducible by running the model again" renders as nothing on a device or in a batch export
that never runs a model. It is a materialisation of the identity, not the edit: it takes no part
in equality or merge, and the identity remains what the part means.
**FR-DEV-4 — Ordered, GPU-resident execution.** The pipeline executes as a sequence of GPU
compute stages. Intermediate results remain in GPU memory between stages. **Processed pixels
shall reach the display without a CPU round-trip.** *(This is a hard architectural constraint —
see ARCH §6.1.)*
**FR-DEV-5 — Edit history.** Per-image undo/redo of edit operations, persisted with the catalog
so history survives a restart. Named snapshots of an edit state.
**FR-DEV-6 — Presets.** Save, apply, and manage named presets covering a subset of the edit
graph. Copy/paste settings between images. Batch-apply to a selection.
**FR-DEV-7 — Before/after.** Compare current edit state against the unedited original or against
a chosen history state.
**FR-DEV-8 — Spot removal.** Non-destructive clone and heal spots stored as parameters in the edit
graph (target, radius, feather, source offset, opacity, mode), with automatic source placement and
manual override, plus a visualise-spots mode.
Sensor dust is unavoidable with interchangeable lenses, and dust spots are the most common reason a
photographer leaves a RAW editor for a pixel editor mid-workflow. This is not the layer-based
pixel editing excluded by §1.3 — it is a standard parameterised develop operation, and the brush
infrastructure required by FR-DEV-3's masks already covers most of the cost.
**FR-DEV-12 — Colour grading by tonal range.** Hue and strength applied independently to the
shadows, the midtones and the highlights, plus a global cast over the whole frame. Ordinary
parameters in the edit graph like any other adjustment, presented as colour wheels where the
frontend implements them and as sliders where it does not.
FR-DEV-3's colour mixer already adjusts hues, and this is not a second copy of it. The mixer acts
on the colours that are *in* the frame and can only turn what it finds; grading acts on a *tonal
range* and puts colour where there was none. That is the difference between correcting a colour and
choosing one: the mixer cannot warm an already-neutral highlight, and it cannot tone a monochrome
conversion at all. Split toning — cool shadows against warm highlights — is the look this makes
possible, it is the one every competing developer ships, and there is no way to reach it from the
adjustments above.
**FR-DEV-18 — Dehaze.** Remove, or add, the atmospheric veil that distance puts between the camera
and the subject, as a develop operation with a single symmetric amount. The transmission is
estimated from the photograph rather than supplied by the photographer, and every length the
estimate depends on is stored as a fraction of the frame, so that what is judged on screen is what
lands in the exported file (FR-DSP-1).
Haze is the one degradation the tone and colour controls cannot reach, because it is spatially
varying: a black point that clears the mountains crushes the foreground, and a contrast curve that
clears the mountains does the same. It belongs with texture and clarity in FR-DEV-3's adjustment
set as a member of the compositional detail family — the operations whose radius is a property of
the picture rather than of the sensor — and it runs in the same neighbourhood stage, for the same
reason: it is defined by what the pixels around a pixel are doing.
**FR-DEV-10 — Range masks.** A mask layer may select by a band over the photograph's own values —
luminance, or colour — with soft edges, stored as bounds and softness rather than as pixels and
rasterised on the device in the existing mask pass.
Every mask in FR-DEV-3 begins from a shape: painted, drawn, or found by a model. A range begins
from a property, which is what a luminosity mask has always been and what makes a local adjustment
blend without a halo — the boundary is the picture's own, at whatever detail the picture has. It is
also the only route to selecting by skin tone, and the prerequisite for composing a selection from
several criteria at once.
**FR-DEV-19 — Mask editing.** A mask layer's coverage shall be editable by hand after it is
created: painted into, erased from, and built from more than one selection. Every edit is stored as
geometry and parameters in the edit graph; no rasterised mask is written to a file and none exists
in CPU memory (ARCH §5.4).
A model finds a subject in a second and, before this, the photographer could not then change it by a
single pixel. The coverage arrives approximately right — stopping inside a shoulder, leaking into
the hair — and FR-DEV-3's edge controls move the *whole* boundary, so no value of either fixes two
errors that go opposite ways. Every competing developer has had a brush over its selections for a
decade; a mask that cannot be corrected is the reason an edit leaves this application for a pixel
editor.
**FR-DEV-19a — Mask composition.** A layer's mask is an ordered list of parts, each naming a source
and how it joins the mask before it — added to it, or taken out of it. A layer of one part is
exactly the layer that existed before this, and reads and writes the same sidecar. The parts of a
layer merge under FR-NC-9 as the layer does, and each carries its own edge treatment.
**FR-DEV-19b — Hand correction.** A part may be painted, with add and erase strokes, at a radius,
hardness and flow the photographer sets. Strokes are stored as normalised source coordinates and
rasterised on the device, so a correction stays on what it was painted on through a crop, a zoom and
an export at any size. A whole stroke is one step in the history.
**FR-DEV-19c — Seeing the mask.** Each layer's mask shall be drawable over the photograph, switched
per layer by an eye on its row and drawn in a colour of that layer's own, so that several can be
shown at once and told apart; the style — a tint, an alpha, or an outline — is one setting for all
of them. What is shown is the finished mask — every part folded, with the layer's feather, falloff,
morphology, invert and opacity applied — and it is shown for a layer that carries no adjustment
yet, which is the state every mask is in for its first few seconds. No rendered output ever carries
it: the reveal is a property of looking at an edit, not of the edit, so it reaches the pipeline
through the composition that draws the canvas and through no other, and it does not travel in a
sidecar.
Nobody can refine an edge they are not being shown. Before this the only thing drawn on the
photograph was FR-DEV-3's region overlay — a false-coloured picture of what the *model detected*,
which knows nothing of a layer's shaping and nothing at all about a gradient, a range or a stroke —
so choosing a subject or a category produced a layer whose extent was invisible, and every control
in FR-DEV-19a and FR-DEV-19b acted on something the photographer could not see.
**FR-DEV-16 — A keyboard vocabulary for develop.** Every develop gesture that can be reached from
the keyboard is bound, tagged beside its implementation, and generated into the gesture book
(FR-UI-4): stepping through the folder, fit and 1:1, holding the original, undo and redo, copying
and pasting settings, cycling the adjustment groups, resetting the control last moved, and showing
or hiding the selected mask layer. None is keyboard-only — each has a pointer and a touch route
(FR-DEV-3b) — and the book is regenerated from the tags, so it cannot describe a binding the
application does not have.
Editing rhythm depends on the hands staying put: reaching for a menu breaks the concentration of an
edit the same way it breaks the pace of a cull. A binding nobody can discover is the same as no
binding, which is why the generated book is part of the requirement rather than documentation of it.
### 3.4 Display and interaction
**FR-DSP-1 — Proxy-resolution rendering.** The develop view renders at the resolution actually
required by the viewport, not the source resolution. A 60MP image displayed in a 2000px viewport
processes approximately 2000px of data, not 60MP.
**FR-DSP-2 — Tiled computation.** The visible region is divided into tiles. Only tiles
intersecting the viewport are computed. Panning computes only newly exposed tiles; already-valid
tiles are reused.
> **Status, 2026-09-19: as written, unbuilt, and waiting on S6.** The interactive path does not
> tile, and [frame-budget.md](frame-budget.md) argues it should not on the reference desktop: a
> fused pass over a 4K viewport costs 4.5 ms of a 16 ms budget, a tile cache could save at most
> that, and the one stage over budget is a convolution that tiling makes worse. That argument is
> a desktop measurement. The case this clause was written for — an image larger than the GPU
> memory of a mid-range Android device (NFR-RES-2, ARCH §6.2) — has not been measured, and S6 is
> the spike that measures it. Until it runs the clause stands, so that a decision is taken on a
> number from the device the clause is about rather than from the one it is not. If S6 finds the
> fused pass inside budget there too, FR-DSP-2 becomes a scheduling concern for export and
> thumbnailing as frame-budget.md proposes; if not, S6 names the stage to tile.
**FR-DSP-3 — Interactive latency.** Moving a slider updates the visible region within one frame
budget at proxy resolution. When a full-resolution result is needed it is computed
asynchronously, and the proxy result remains on screen until it is ready.
**FR-DSP-4 — Progressive refinement.** During rapid interaction the app may render at reduced
quality or resolution, refining to full quality when interaction settles. Refinement is visually
smooth, not a jarring swap.
**FR-DSP-5 — Zoom and pan.** Fit, 1:1, and arbitrary zoom levels. At 1:1 and above, the pipeline
operates on the visible crop at full source resolution.
**FR-DSP-6 — Colour management.** The display path is colour-managed via the output device
profile. Where the platform and display support it, output at greater than 8 bits per channel
and in a wide gamut. On Android this means using the wide-gamut display path where available.
**FR-DSP-7 — Image evaluation.** Provide a live histogram (luminance and per-channel, in the output
colour space), highlight and shadow clipping indicators, and a pixel colour readout under the
cursor or touch point.
**These derive from a GPU-side reduction into a small buffer. Per-frame CPU readback of image data
is prohibited** — it would violate ARCH §6.1 on every frame, which is precisely the bottleneck darktable
documents. Histogram computation shall not extend the FR-DSP-3 frame budget.
Without this a photographer cannot see what highlight recovery is actually doing, which makes the
FR-DEV-3 adjustment set substantially less usable.
**FR-DSP-8 — Per-display colour and scaling.** The display transform is selected per the display
currently showing the canvas, and updates when the window moves between displays. Fractional and
mixed DPI scaling are handled without resampling artefacts in the canvas.
The profile-acquisition mechanism is stated per display server, with a defined fallback where
Wayland provides no profile. On a multi-monitor desktop with differing profiles, showing wrong
colours on the second display is a correctness defect, not a polish item.
### 3.5 Adaptive interface
DarkRoom ships **one adaptive interface**, not separate touch and desktop applications. A single
Slint codebase reflows by available space and input modality, guaranteeing feature parity by
construction. Phones are not a target (D15); the layout family spans tablet and desktop.
**FR-UI-1 — Layout breakpoints.** The interface adapts across at least two layout classes:
| Class | Typical | Develop layout |
|---|---|---|
| Compact | ~~Tablet portrait,~~ Narrow desktop window | Canvas full-width; one collapsible panel at a time; filmstrip on demand |
| Expanded | Tablet portrait and landscape, desktop | Filmstrip, canvas, and adjustment panel simultaneously |
Layout class is a function of window size, not device type — a narrow window on desktop uses the
compact layout, and the transition is continuous rather than a mode switch.
The develop column's side — beside the canvas, or docked under it on a tall window — follows the
window's aspect, independent of the layout class (`docs/ui-navigation.md` D-N7). A placement, not a
mode.
> **Amended 2026-09-06.** Tablet portrait is the expanded class, not the compact one: both
> orientations of the 12-inch tablet clear `EXPANDED_MIN_WIDTH`, 820 logical pixels (D-N2). Compact
> fires only on a desktop window dragged narrow. What portrait needs is the dock, which is the
> aspect axis above and not a second class (D-N7).
> **Amended 2026-09-19.** The expanded row has said "filmstrip" since the table was written, and
> the build had the photo roll on demand in both classes — the compact row applied everywhere. In
> the expanded class the roll is **open by default** and closable, and it is shown for a set of
> files named on the command line as much as for a library, because a set of photographs is a set.
> It carries the place FR-UI-8 remembers — how many, which one, its name, and what the filter is
> narrowing to — so that develop and the grid read as one interface with two views of the same
> set, not two screens joined by a button.
**FR-UI-2 — Input modality.** The interface detects and adapts to the active input method, which
is independent of layout class: a tablet may have a keyboard and pointer attached, and a desktop
may have a touchscreen. Modality affects control sizing and affordances (FR-DEV-3b), not layout —
save for the group selector, which moves from above the develop column into the tool rail under
touch (`docs/ui-navigation.md` D-N6). Switching input mid-session shall be handled without restart.
**FR-UI-3 — Touch targets.** Interactive controls present a minimum 44pt hit target when touch is
the active modality. Hit targets may exceed the drawn control bounds.
**FR-UI-4 — Gestures.** The canvas supports pinch-zoom, two-finger pan, and double-tap to toggle
fit/1:1. Gestures are additive: every gesture-driven action has a non-gesture equivalent, so no
functionality is touch-only.
**FR-UI-5 — Pointer and keyboard.** Where a pointer is present: hover states, right-click context
menus, and scroll-wheel adjustment on numeric controls. Keyboard shortcuts cover navigation,
rating, and common adjustments. Neither is required for any operation to be reachable.
> **Amended 2026-09-19.** "Rating" above is not qualified by view and was built as though it were:
> the judgement keys lived in the grid alone. They shall work in every view that shows a
> photograph, applying to the **one on screen** — in develop, the open photograph, never a
> selection left behind in the grid — with a pointer and touch equivalent in the same view, since
> a tablet has no number row. Judging in develop does **not** advance to the next frame:
> auto-advance belongs to FR-CULL-4's mode, where moving on is the point, and in develop the
> photographer is working on the frame in front of them. The current rating and flag are shown
> wherever they can be set, and on the roll's cells, so stepping along a set shows what has been
> judged.
**FR-UI-6 — Shared component library.** Touch and desktop presentations are variants of shared
components, not parallel implementations. A new operation (FR-DEV-3c) becomes usable on both
without frontend work.
**FR-UI-7 — On-canvas controls.** Custom controls that operate on the canvas — crop handles,
gradient placement, brush strokes (`WidgetKind` in FR-DEV-3a) — size their interaction regions to
the active modality while rendering identically. A crop handle drawn at 8px may carry a 44pt touch
region.
**FR-UI-8 — Resumed place.** The application shall remember where the photographer was and return
them to it. "Where" is the view (grid or develop), the scope (whole library, a collection, or the
trash), the rating filter narrowing it, and the photograph on screen — the open one in develop, the
first visible one in the grid. It shall be restored on launch, and preserved across every transition
between screens within a session: leaving the grid for Settings, Import, People or the develop view
and returning shall land where it was left, never at the top of the library.
*A place is addressed by what every device agrees on.* The photograph by its remote path, the
collection by its UUID; never by a grid ordinal, `images.id` or `collections.id`, all of which are
local to one catalog and to one ordering. Where the ordinal is needed it shall be computed through
the same ordering the grid draws with, so that a restored position names the photograph the
photographer actually left. Capture time is the permitted fallback when the photograph is gone.
*A place travels.* One record per library is exchanged through the same derived folder as the
thumbnail shards and the catalog snapshot (FR-NC-3), so that a session begun on one device can be
continued on another. The newer of two records wins; there is nothing to merge, because two devices
cannot both be where the photographer is.
*And it is always advisory.* A place that cannot be read, names a collection this device has not
merged, or points at a photograph that has since been deleted shall degrade to the nearest sensible
position and never to an error, an empty view, or a refusal to start. A record arriving from another
device shall not be applied once the photographer has begun working in this session — it is a
handover, not an interruption.
### 3.6 Export
**FR-EXP-1 — Formats.** Export to JPEG, PNG, and TIFF (8 and 16-bit). Quality, chroma subsampling,
and bit depth are configurable.
**AVIF and JPEG XL are post-v1** and are not part of this requirement's acceptance. Either may still
be listed in the settings page before its encoder exists, on one condition: choosing it shall fail
with a typed error naming the format, never with a file. The test
`every_offered_format_either_encodes_or_explains_itself` walks every format the page offers and
enforces exactly that, so a format cannot be added to the picker and quietly reach an encoder that
does not handle it.
*On splitting this requirement.* It read as one undifferentiated list — "JPEG, PNG, TIFF (8 and
16-bit), and AVIF or JPEG XL" — which left it neither met nor unmet. Three formats, both TIFF
depths, and the configurability clause are built, encode, embed their profile, and are tested; the
fourth item is a deliberate deferral, and the encoders for it are the two with the least settled
library support. Fused into one sentence, the only choices were a tag asserting something untrue or
no tag at all, and the second is the worse of the two: it would have removed the register's record
of four-fifths of a requirement that is finished. The deferral is now stated where it can be read as
a decision rather than inferred from an error variant.
**FR-EXP-2 — Colour space.** Export in a selectable output colour space (sRGB, Display P3,
Adobe RGB, ProPhoto), with the correct ICC profile embedded.
**FR-EXP-3 — Output sizing.** Export size shall be specifiable by any of the following modes:
| Mode | Behaviour |
|---|---|
| Original | Full source resolution after crop |
| Long edge | Specified pixels on the longer dimension; aspect preserved |
| Short edge | Specified pixels on the shorter dimension; aspect preserved |
| Width × Height (fit) | Scaled to fit within the box; aspect preserved; result may be smaller in one dimension |
| Width × Height (fill) | Scaled to cover the box and centre-cropped to exactly those dimensions |
| Percentage | Scaled by a factor of the source |
| Megapixels | Scaled so the result approximates a target pixel count |
| Print dimensions | Physical size (mm or inches) at a specified DPI, resolved to pixels |
Additional constraints:
- **Upscaling** is permitted but shall be off by default, with an explicit opt-in. Where disabled,
a request larger than the source exports at source size rather than failing.
- **DPI metadata** is settable independently of pixel dimensions, for print workflows.
- **File-size ceiling** (JPEG/AVIF/JPEG XL): optionally target a maximum output size in KB/MB, with
the encoder iterating quality to meet it. Useful for upload limits.
- Sizing operates on the **cropped** result, so the crop rectangle defines the aspect ratio unless
a fill mode overrides it.
**FR-EXP-4 — Resampling and output sharpening.** Resizing uses a quality resampler (Lanczos or
equivalent) operating on linear-light data at pipeline precision, before quantisation to the output
bit depth. Output sharpening is selectable (none / screen / matte paper / glossy paper) and its
strength scales with the resize factor, since downscaling softens.
**FR-EXP-5 — Export presets.** Named presets capture format, quality, colour space, sizing mode,
sharpening, metadata policy, and destination. A preset is applicable to a single image or a batch.
Multiple presets may be applied in one operation, producing several outputs per image — e.g. a
full-size TIFF alongside a 2048px sRGB JPEG.
**FR-EXP-6 — Naming and destination.** Output filenames are generated from a template supporting at
minimum: original filename, sequence number, capture date, export dimensions, and preset name.
Collision policy (overwrite / skip / auto-increment) is configurable. Destinations include a local
path and a Nextcloud remote path (FR-NC-7).
**FR-EXP-7 — Batch export.** Export a selection with one or more presets, running in the background
with progress and cancellation. Uses all available cores and the GPU. A failure on one image is
reported and does not abort the batch.
**FR-EXP-8 — Metadata on export.** Configurable EXIF/IPTC/XMP retention, including an option to
strip GPS and personal metadata. Copyright and contact fields are settable per-preset.
**FR-EXP-9 — Full-quality path.** Export always uses the full-resolution, highest-quality pipeline
regardless of what the display was showing — including the high-quality demosaic (FR-RAW-3), never
the fast preview method.
### 3.7 Nextcloud integration
Mechanics below are verified against Nextcloud 34 documentation and server/desktop-client source.
Three findings shape this section and are recorded as constraints in ARCH §6.6–ARCH §6.8.
**FR-NC-1 — Account setup.** Connect via **Login Flow v2**: `POST /index.php/login/v2` returns a
browser URL and a poll token; the app opens the URL in the *system browser* (never an embedded
webview) and polls `POST /login/v2/poll` until it returns an app password. The token is valid 20
minutes and the success response is returned exactly once. The app never sees the user's primary
password.
The `User-Agent` sent during the flow names the resulting app password in the user's security
settings, so it shall identify the device (e.g. `DarkRoom (Linux desktop)`), allowing per-device
revocation. Logout shall call `DELETE /ocs/v2.php/core/apppassword` to revoke cleanly.
Manual app passwords are supported as a fallback for unusual server configurations.
**FR-NC-2 — Credential storage.** Linux: Secret Service via libsecret. KDE exposes the same
interface through `ksecretd` since KF5.97, so one code path covers GNOME and KDE. Where no
secrets daemon is running, the app shall enter an explicit degraded mode rather than silently
storing credentials in plaintext.
Android: Keystore-backed encryption. Note `EncryptedSharedPreferences` is deprecated; the current
approach is DataStore for persistence with Tink for encryption and Keystore for key protection.
Keys must not require user authentication, or background sync will fail.
**FR-NC-3 — Remote browsing without full download.** The app shall display a remote library's
thumbnails without transferring full RAW files. Two mechanisms, selected per-account by capability
probe at setup:
1. *Server previews* — `GET /core/preview?fileId=…` where available. **`forceIcon=false` is
mandatory**: the default returns a generic mimetype icon when the server cannot render the
file, which would otherwise be cached as though it were a thumbnail. The `nc:has-preview`
property in PROPFIND indicates per-file availability.
2. *Range-based embedded preview extraction* — the required fallback (see ARCH §6.7). Fetch the first
64–256KB via HTTP `Range`, parse the container to locate the embedded JPEG preview, then fetch
exactly that byte range. Typical cost 1–3MB versus 25–100MB for the full file.
**FR-NC-4 — Change detection.** Sync shall use recursive ETag pruning, matching the official
desktop client's discovery algorithm:
1. `PROPFIND Depth: 0` on the sync root requesting `getetag`. If unchanged from the stored value,
nothing anywhere in the library has changed — sync completes in one request.
2. Where changed, `PROPFIND Depth: 1` and recurse only into child folders whose ETag differs.
Cost is proportional to the changed subtree, not to library size. `Depth: infinity` shall not be
relied upon (frequently disabled or prohibitively expensive). ETags shall be normalised for quote
inconsistencies before comparison, or spurious full rescans result.
**FR-NC-5 — Identity.** The catalog shall key remote files on Nextcloud's `oc:fileid`, which is
stable across renames and moves, so that a server-side move is detected as a move rather than as a
delete plus a re-download of a 100MB file.
**FR-NC-6 — Selective download.** Downloads are on-demand and resumable, running in the
background. RAW files are never bulk-synced by default. On Android, transfers respect
unmetered-network and charging constraints.
Three storage tiers:
| Tier | Content | Policy |
|---|---|---|
| Metadata | Catalog rows, edit-graph sidecars | Always synced; kilobytes; sync even on metered connections |
| Previews | Embedded JPEGs or server previews | LRU-evicted, size-capped; what the grid browses |
| Full RAW | Source files | Explicit pin or on-demand open only |
**FR-NC-6a — Cache rules.** The user shall be able to pin a *set* of images at a chosen tier, with
the set defined by a rule that the app re-evaluates as the catalog changes. Selectors shall include
at minimum:
- **Collection** — "this trip is available offline"
- **Folder**, optionally recursive
- **Date range**, absolute or **rolling** ("the last 90 days", which moves with the clock)
- **Rating**, **colour label**, **flag**, or **keyword** — "every 5-star image, always"
- **Boolean composition** of the above
A rolling window shall stay current without user intervention. Where rules disagree about an image,
the most generous tier wins.
*Rationale:* selective sync is only usable if the selection can be expressed as intent rather than
enumerated by hand. "Keep this shoot and everything from the last three months" is a sentence a
photographer will say; selecting four thousand files individually is not.
**FR-NC-6b — Lazy eviction.** An image that stops matching a rule is **not** deleted immediately; it
becomes the first candidate for eviction when the cache cap (NFR-RES-4) or platform memory pressure
(FR-PLAT-AND-5) actually requires space. Eviction order is unpinned originals by last use, then
proxies, then thumbnails. **Metadata and sidecars are never evicted** — they are authoritative
(ARCH §6.12) and small.
**FR-NC-6c — Availability is visible.** Every image shall carry a visible availability state:
*Original*, *Preview*, *Metadata only*, or *Offline*.
- An operation requiring absent data shall say so, with the transfer size, *before* starting
- Export from a preview-only image is **refused**, not silently degraded
- A pinned set reports its true byte cost before the user commits
*Rationale:* Lightroom Classic syncs 2560px proxies while displaying the original's filename,
extension, and size, so users do not know what they actually have. Sync failures of legibility are
more damaging than failures of transport.
**FR-NC-6d — Placeholder libraries.** Where a library is a folder kept by a sync client in
virtual-files mode, the app shall treat a placeholder as *the photograph, not downloaded* — never as
a one-byte file and never as a missing one.
- A placeholder is catalogued under the photograph's own name, with an identity that does not change
when it is downloaded
- Reading one yields a distinct, actionable error; it shall **not** be reported as absent, because
the sidecar writer creates a new document when a sidecar is absent and would discard the existing
one (FR-CAT-8)
- Its size is reported as unknown rather than as the stub's byte count
Where the client offers hydration, content may be fetched **as a borrow**: a file is returned to the
state it was found in, so a pass releases what it downloaded and leaves alone what the user already
had. Releasing means asking the client to dehydrate — **never deleting**, which inside a synced tree
would propagate to the server and remove the photograph everywhere.
Hydration is whole-file and shall never serve browsing (ARCH §9.0 finding 3, §9.0a). It is for the
originals tier and for passes the user has been quoted a cost on and has agreed to.
**FR-NC-7 — Upload.** Files above 5MB use **chunked upload v2** against
`/remote.php/dav/uploads/<userid>/`: `MKCOL` to create the upload folder, `PUT` each chunk, then
`MOVE` the `.file` pseudo-entry to the destination. Chunks are 5MB–5GB and named 1–10000.
`OC-Total-Length` shall always be sent so quota is checked up front rather than at assembly time.
Upload folders expire after 24h of inactivity; the app shall persist upload state and either
resume or `DELETE` stranded uploads on startup.
Small files (sidecars) use **bulk upload** via `POST /remote.php/dav/bulk` with a
`multipart/related` body, allowing hundreds of edit-graph sidecars in a single request.
**FR-NC-7a — Remote layout.** FR-NC-7 says how bytes travel; this says where they land. An
uploaded original shall be placed under the account's library root in a directory expanded from a
date template, defaulting to `{yyyy}/{yyyy}-{mm}-{dd}` — one directory per year, one per capture
day beneath it. The template is configurable per account and uses the same token vocabulary as
export naming (FR-EXP-6), so `{date}` means the image's **capture** date and never today's; an
image with no readable capture time falls back to file mtime, and the fallback is visible in the
import report rather than silent.
Two consequences the mechanics do not give for free:
- **The path is derived, not remembered.** The same image uploaded twice from two devices must
compute the same destination, so expansion is a pure function of capture metadata and the
template — never of local library layout, which differs per device.
- **Existing trees are not restructured.** An image already present remotely stays where it is.
The template governs placement on upload only; DarkRoom shall not move server-side files to
conform, because the remote library is also reachable by other clients (§1.3).
Sidecars follow their image, not the template: they are named from `oc:fileid` (FR-NC-8) and live
beside the file they describe.
**FR-NC-7b — Ingest to remote.** Import (FR-CAT-10) and upload (FR-NC-7) compose, and the library
an import targets is **always** the remote one. There is no local library for a card to land in:
FR-NC-6 puts the library on the server, so an import has exactly one destination and the interface
shall not offer a choice of another.
What lands on the device is a **staging copy**, in a location the app owns, and it is not a
library: no view lists it, and it is removed once the server confirms the file. A staged file the
server did not take remains queued, and a later import drains the queue. An import with no
reachable server therefore succeeds and defers, rather than failing (FR-NC-10); an import with no
*account* has nowhere to go at all and shall be refused.
The bytes shall reach the device before they reach the network. Streaming a card straight to the
server would make a move-import erase a card against an in-flight upload, and would make importing
impossible offline.
**On erasing the card.** A move-import shall delete from the card only those photographs the
**server has confirmed** — not those merely written to the staging copy, since that copy is removed
as soon as the upload succeeds and would otherwise be the only remaining copy. Anything that does
not upload keeps its card copy. A card erased against an unfinished upload is unrecoverable, which
is the one failure in this app with no undo.
Duplicate detection (FR-CAT-11) runs against the catalog before upload, so re-inserting a card
that was already imported transfers nothing. Where the catalog cannot answer — a file already in
the destination folder, locally or on the server — the folder's own listing shall answer instead,
so a re-import neither duplicates nor renames what is already held.
**FR-NC-8 — Edit metadata sync.** Edit graphs sync bidirectionally as sidecars, one per image,
named deterministically from `oc:fileid`. Each sidecar carries a monotonic revision counter, a
per-device UUID, and a last-edit timestamp.
A sidecar holds a **keyed set of Versions** (FR-CAT-12), not a single edit graph — one image may
carry several virtual copies, and a single-graph format could not represent them. Version identity
is part of the sidecar schema, so conflict merge (FR-NC-9) operates per-version.
**FR-NC-9 — Conflict handling.** Sidecar updates use `If-Match` with the known ETag for optimistic
concurrency (`If-None-Match: *` for creates). On `412 Precondition Failed` the app shall fetch the
remote sidecar and **merge at the edit-graph node level** — disjoint edits (e.g. a crop on one
device, an exposure change on the other) both survive; genuinely conflicting nodes resolve by
timestamp — then retry with the new ETag under a bounded retry count.
The app shall **not** replicate the desktop client's `(conflicted copy)` file behaviour. Sidecars
are structured data of a few KB; a read-merge-rewrite cycle is cheap and preserves user intent.
Only genuinely ambiguous merges surface to the UI. Source RAW files are write-once and shall never
generate a conflict.
**FR-NC-10 — Offline-first.** The app is fully functional offline against cached content. Local
sidecar writes are atomic (temp file plus rename) and committed locally *before* any network
round-trip, so editing never blocks on connectivity. Sync resumes automatically when connectivity
returns.
**FR-NC-11 — Initial catalog build.** For first sync of a large remote library, the app may use
WebDAV `SEARCH` (RFC 5323) against `/remote.php/dav/` filtered by mimetype and paginated via
`d:limit`/`d:nresults`, in preference to walking thousands of folders with PROPFIND.
**FR-NC-12 — Backend independence.** Sync shall be implemented against a backend interface. No
protocol detail specific to any one backend may appear outside its connector, and no layer above
the interface may name a connector — with the single exception of the registry that constructs them
(`dr_ui::remote`).
A trait over operations is not sufficient on its own, and the first release proved it: `dr-ui`
constructed the Nextcloud backend directly in seven files, an account *was* a server URL beside a
DAV user id, and the local cache directory was named after a hostname. Independence requires four
things — operations, declared capabilities, an account model with no server in it, and a
registration mechanism (ARCH §8.0, `docs/storage.md`).
Backends **declare capabilities** rather than conforming to a lowest common denominator, because
the property that makes Nextcloud sync fast — directory ETags propagating up the tree, so an
unchanged root proves an unchanged library — is not a general guarantee. An interface built to the
common subset would force full enumeration on every sync (ARCH §8.1).
Where a capability is absent the app shall **degrade visibly, not silently**:
- Sync strategy in use is reportable to the user, so a slow backend is visibly slow
- Without byte-range reads, remote browsing cannot extract embedded previews; the app shall refuse
full downloads for browsing on a metered connection and explain why
- Without conditional writes, sidecar conflict detection falls back to revision comparison, which
narrows but does not close the race; this is surfaced as a reduced-safety mode
**FR-NC-13 — Folder libraries.** A library shall be openable as a **plain directory** — a local
disk, a network mount, an external drive, or a folder another client already syncs — with no
account, no server and no credential.
This is a requirement rather than a convenience for three reasons. It is what a photographer with
an archive drive and no server actually has. It is the only route that works where no secrets
daemon exists, which FR-NC-2 otherwise treats as a degraded mode. And a second connector is the
only way to keep FR-NC-12 honest: an interface with one implementation cannot be shown to be an
interface.
The folder connector shall declare its capabilities truthfully rather than flatteringly — in
particular it shall **not** claim propagating directory ETags, because a POSIX directory's mtime
describes its own entry list and nothing beneath it, and a backend that claimed otherwise would
hide edits rather than merely run slowly (ARCH §8.4a).
### 3.8 Platform integration
#### Android
**FR-PLAT-AND-1 — Storage access.** Library access is obtained **exclusively via the Storage Access
Framework**: the user grants one or more document trees through `ACTION_OPEN_DOCUMENT_TREE`,
persisted with `takePersistableUriPermission` and enumerated via `DocumentsContract`.
The app shall **not** request `MANAGE_EXTERNAL_STORAGE` and shall **not** depend on
`READ_MEDIA_IMAGES` for RAW discovery. See ARCH §6.9 for why neither is viable.
**FR-PLAT-AND-2 — Permission loss.** Loss of a previously granted tree permission — revocation,
reinstall, removed SD card — shall be detected and surfaced, marking affected images offline per
FR-CAT-9 rather than deleting catalog rows.
**FR-PLAT-AND-3 — Process lifecycle.** An Android process may be killed at any moment. Edit state
shall be durable such that process death loses at most the last uncommitted parameter change. On
resume the app restores the develop session, its image, and its viewport.
**FR-PLAT-AND-4 — Background execution.** Long-running sync and export use the platform's managed
background execution with the constraints in FR-NC-6, and a foreground service with notification
for user-initiated exports. Behaviour under Doze and battery-saver is specified and tested.
**FR-PLAT-AND-5 — Memory pressure.** The app shall respond to `onTrimMemory` /
`ComponentCallbacks2` by evicting caches per NFR-RES-1, in a stated eviction order (GPU tiles
first, then proxies, then thumbnails).
**FR-PLAT-AND-6 — Intents.** Register as a receiver for image view and share intents, and provide
share-out of exported results via `FileProvider`.
#### Linux
**FR-PLAT-LIN-1 — Desktop integration.** Follow the XDG Base Directory specification for config,
data, cache, and state. Register MIME associations for supported RAW types and ship a `.desktop`
entry.
**FR-PLAT-LIN-2 — Display server.** Support X11 and Wayland. Where Wayland's colour-management
protocol is unavailable, FR-DSP-8's stated fallback applies.
**FR-PLAT-LIN-3 — Sandboxed distribution.** Where distributed as Flatpak, filesystem access uses
portals and credential storage uses the Secret Service portal, both verified to satisfy FR-NC-2 and
FR-CAT-1 within the sandbox.
#### Windows
Specified in [windows.md](windows.md); a stated channel under NFR-COMPAT-2, not a v1 one.
**FR-PLAT-WIN-1 — Known folders.** Configuration under `%APPDATA%\darkroom`; data, cache and
state under `%LOCALAPPDATA%\darkroom`. Nothing under the profile root and nothing relative to the
working directory. The layout beneath those roots is the same as under XDG, so a library directory
moves between platforms unchanged.
**FR-PLAT-WIN-2 — Installer.** A per-user installer that needs no elevation, registers an
uninstaller, and whose uninstaller removes what the installer wrote and nothing the application
wrote. Models are installed beside the executable and found there last, after the user's own
directories.
**FR-PLAT-WIN-3 — Built from Linux.** The Windows binary and its installer are produced by the
Linux CI from the same commit as every other channel, with no Windows machine in the build.
Verification on Windows is a release step, recorded per release, not a build step.
### 3.9 Culling
Per D11 this is the product's primary differentiator, not an incidental capability.
**The opportunity, stated plainly.** Photo Mechanic is fast because it displays the camera's
embedded JPEG — no demosaic, no database, no import step. FastRawViewer is *truthful* because it
shows a genuine raw histogram, raw-derived clipping, and focus peaking. **No shipping tool combines
both.** The culler/editor split exists only because Lightroom's culling is slow — it is a workaround
photographers tolerate, not a workflow they want. A tool that is genuinely Photo Mechanic-fast
eliminates the handoff rather than improving it.
**FR-CULL-1 — Instant display.** Displaying the next image shall not wait on demosaic, catalog
import, or full decode. The embedded JPEG preview is shown immediately; higher-quality renders
replace it progressively.
*Acceptance:* next-image display within **50 ms** of the input event, sustained across a
3,000-image folder, on both platforms. This is the single most important performance figure in the
document — Lightroom's ~2 s stall is the entire reason a competing product category exists.
**FR-CULL-2 — Preview ladder.** Previews resolve through tiers, each falling through to the next:
1. Embedded JPEG preview from the RAW container (instant)
2. Cached proxy from a previous visit
3. Background full decode, promoted when ready
Some cameras embed previews below sensor resolution, and some embed none. The app shall **detect
this per camera model** and pre-emptively background-render where the embedded preview is
insufficient, rather than showing the user a soft image and letting them discover it at zoom.
**FR-CULL-3 — Raw-truth overlays.** Culling decisions are made against raw data, not the embedded
JPEG:
- **Raw histogram** — computed from sensor data, not the preview. The embedded JPEG's histogram
misrepresents available highlight headroom.
- **Raw-derived clipping indicators** — a JPEG's clipping warnings systematically lie about what is
recoverable in the raw.
- **Focus peaking** — overlays in-focus regions on the preview, so focus is verifiable **without
zooming to 100%**. This removes the largest single source of culling latency from the critical
path.
Per D11, zoom-to-100% remains available for certainty; peaking makes it optional rather than
mandatory.
**FR-CULL-4 — Culling mode.** A dedicated full-screen mode with:
- **Auto-advance** on judgement — users independently reinvent this in darktable, Lightroom desktop,
and Lightroom mobile, which is strong evidence it should be the default rather than an option
- **One-key reject**, plus the full rating, flag, and colour-label axes
- **Filter to unjudged**, so a session resumes where it stopped
- Keyboard-driven on desktop; single-thumb reachable on tablet
- **Evidence beside the frame** (added 2026-09-19) — FR-CULL-13's signals, readable at a glance
without leaving the mode or opening the frame
- **The last import as a scope** (added 2026-09-19) — the set a culling session most often starts
from, reachable in one step and counted, expressed as a term of the §5 selector language rather
than as a special collection
**FR-CULL-5 — Burst and near-duplicate grouping.** Group frames by capture-time proximity and image
similarity, allowing a burst to collapse to one representative and be judged as a unit.
This is the one automated capability photographers consistently praise, precisely because it is a
mechanical grouping problem rather than a taste judgement. Automated *selection* is distrusted — the
documented failure is rejecting the only frame of an important moment because someone blinked.
**FR-CULL-6 — Comparison.** Side-by-side and survey comparison of a selection, with synchronised
zoom and pan, for choosing among near-identical frames.
**FR-CULL-7 — Tablet culling.** Culling shall be fully usable on tablet, as the validated
multi-device workflow (§3.5, D11). Requires only ratings and small proxies to sync, not full
originals — a substantially smaller sync problem than develop parity.
*Design note:* pinch-zoom accidentally triggering ratings is a documented defect in Lightroom
mobile. Gesture and rating targets must not overlap.
**FR-CULL-13 — Evidence, never verdicts.** Added 2026-09-19; numbered after the people clauses
because it governs them too. Everything the app computes about a frame in aid of culling is
**evidence**: shown beside the frame, filterable and sortable through the §5 selector language,
and — within a burst — permitted to *propose* which frame represents the group. The signals are
the raw histogram, clipping and focus (FR-CULL-3), burst membership (FR-CULL-5), and per-face eye
state and head pose (FR-CULL-8a). The list grows by adding to it, never by adding a second kind of
thing.
What evidence may not do: change a rating, a flag, a colour label, or trash membership. There shall
be **no code path from a signal to a judgement write** without a user action between them, and a
proposal — a representative, "three of these have eyes closed" — is accepted by a press, never by a
timeout or a default. This restates FR-CULL-5's ground rather than extending it: automated
selection is distrusted because its documented failure is rejecting the only frame of a moment
because someone blinked, and the remedy is not a better classifier, it is that the classifier does
not hold the pen.
Evidence says what it measured. A focus figure names its region; a per-face count names the
faces; a signal that could not be computed — no faces found, the original not on this device — is
shown as absent, never as zero.
*Acceptance:* a test enumerates every write to the rating and flag axes and shows each reachable
only from an input event. The evidence for a frame is visible in FR-CULL-4's mode and on its grid
cell without opening it, and a filter on any one signal returns exactly the set whose chips show
it.
### 3.9.1 People
Face recognition was deferred in §7 through the 2026-08-08 calibration. It is undeferred here in a
narrower form, and the narrowing is the point.
**What changed.** The deferral treated "face recognition" as an AI feature adjacent to subject
masking. It is not the same kind of thing. Masking is a *taste* operation applied to one image;
grouping photographs by who is in them is a **mechanical grouping problem over the whole library**,
which is the category FR-CULL-5 already commits to and already justifies: grouping is the automated
capability photographers consistently praise, because it organises without deciding. Every argument
FR-CULL-5 makes for burst grouping applies unchanged to people grouping. Answering "where are the
frames with the bride in them" across a 4,000-image wedding is a culling operation, and culling is
the differentiator.
**What is deliberately not in scope**, because it is the failure FR-CULL-5 names: no automated
*selection*. Nothing here rejects a frame, ranks a face, scores a smile, ~~or detects a blink~~. The
feature produces a **filter**, never a judgement. The user's rating axes remain the only thing that
rejects a photograph.
> **Amended 2026-09-19.** "Detects a blink" is struck. The exclusion was always of *judgement*,
> and a blink is a fact about a frame of the same kind as a clipped highlight: reporting it is
> evidence, acting on it is the failure. FR-CULL-8a detects eye state and head pose; FR-CULL-13
> says what may be done with them — shown, filtered, sorted, proposed — and what may not. The
> sentence that survives is the one that matters: the user's rating axes remain the only thing
> that rejects a photograph.
**FR-CULL-8 — Face detection.** The app shall detect faces in library images as a background job,
producing per-face a bounding box, five-point landmarks, a detector confidence, and a 512-dimension
embedding.
**The two resolutions are separate, and conflating them is the failure this clause exists to
prevent.** Detection and cropping have opposite resolution needs, and a single buffer cannot serve
both well:
1. **Source.** The image is rendered at **native resolution** through the full-quality path
(FR-EXP-9's pipeline, including the high-quality demosaic of FR-RAW-3). This is the same render
export uses and is deliberately not the FR-CULL-2 preview ladder.
2. **Detector input.** That render is downscaled for the detector, which fixes its input at 640×640
regardless (faces.md §4.1). Detection gains nothing from more pixels than its own input, so the
downscale is free accuracy-wise and is what keeps the pass affordable in CPU.
3. **Crop.** Boxes and landmarks are mapped **back to native coordinates**, and the aligned crop is
sampled from the native render — never from the downscale the detector saw.
4. **Embedding.** The aligned crop is warped to 112×112 in one bilinear step (faces.md §5).
The crop is the reason. ArcFace receives a fixed 112×112 whatever it is given, so the only question
that matters is whether those 112 pixels are real pixels or interpolated ones. Sampling the crop
from a preview means a face occupying a small part of the frame is *upsampled* to reach the
embedder, and an upsampled crop yields a confident embedding of detail that was never there —
which does not fail loudly, it degrades clustering three stages later. Measured on the reference
library under the previous preview-tier implementation: **47% of all stored faces had been
upsampled to reach 112×112**, with `crop_px` as low as 34.
This supersedes the previous rule that detection ran against the thumbnail or proxy tier and never
a full decode. That rule was adopted for affordability and it bought exactly that, at a cost to
crop quality that was not measured until the library was large. Affordability is now met by *when*
the pass runs rather than by *what* it reads: it is background work, preempted by everything
visible, and resumable per image.
**On a remote library this needs the original**, not FR-NC-3's byte-ranged preview — on the
reference library, 412 GB across 19,107 images rather than a range request each. So a whole-library
pass is a **transfer under FR-NC-6**: never automatic, subject to the unmetered-network and
charging constraints, and reported as the download it is before it starts rather than presented as
a local operation. An original already on the device is indexed from what is there. Nothing here
requires the original to be *kept*: it is rendered, cropped, and given back under the same rules as
any other borrowed file (ARCH §9.0a), so the pass costs transfer and time rather than permanent
disk.
Detection is a job in the FR-CAT-3 queue and inherits its properties without exception: coalesced
per image, interruptible, resumable across process death (FR-PLAT-AND-3), and strictly preempted by
visible work (NFR-ARCH-2). A library indexes while idle or it does not index; it never competes with
the grid. `face_index.source_edge` records the native edge each run was made at, and `faces.crop_px`
the pixels behind each individual crop, so raising the standard
later re-indexes only the images that stand to gain rather than all of them.
*Acceptance:* indexing a 10k-image library completes without the grid dropping below NFR-P9's
interaction target at any point, and survives being killed and restarted with no repeated work
beyond the in-flight image. No face is stored whose aligned crop was upsampled beyond a stated
factor; the crop source resolution is recorded per face (`crop_px`) and is auditable.
**FR-CULL-8a — Per-face state.** Added 2026-09-19. For every detected face the app shall record,
as derived data under FR-CULL-12:
- **Eye state** — one open-probability per eye, from a classifier over a crop around each eye
landmark. Per face this reads as both open, one closed, or both closed; per frame, as a count.
- **Head pose** — yaw, pitch and roll, solved from the five landmarks against a generic face
template. No model: this is geometry the detector has already paid for. *Facing the camera* is
the pose within a stated band, and is the proxy for eye contact this document adopts — because
gaze estimation has no redistributable weights (D13), and because in the photographs where it
matters, groups and children and events, a turned head is the decision and averted eyes are
often the better frame.
Both are terms in the §5 selector language (FR-CULL-11) — "everyone's eyes open", "facing the
camera" — composable with a person and with every other term. Both feed FR-CULL-13's evidence and
may propose a burst representative (FR-CULL-5); neither may set a rating or a flag.
**The weights ship under the same test as every other model**: redistributable under a licence
compatible with GPLv3 and with Flatpak, F-Droid and Play, with the training data's terms read as
well as the weights' (D13). The candidate identified 2026-09-19 is OCEC — MIT for code and
weights, data under ODC-By 1.0 and Apache 2.0, six variants from 112 KB to 6.4 MB, a 24×40 crop
per eye, sub-millisecond on CPU, opset 17 with batch-norm already folded. Its published F1 of
0.99 is on its own crops; ours are cut from a five-point landmark, so the number is measured on
this library before it is believed.
*Acceptance:* on a labelled set of at least 500 faces from the reference library, eye state is
within a stated tolerance of its published F1, reported separately for glasses, profile, and
faces under 60 px; facing-the-camera agrees with a hand-labelled split at a stated rate. Both are
recomputed by re-indexing, neither is written to a sidecar, and the indexing pass stays inside
FR-CULL-8's acceptance.
**FR-CULL-9 — Calibrated identity.** Face similarity shall be expressed as a **calibrated
probability that two faces are the same person**, not as a raw embedding distance. Every threshold
in the subsystem — clustering, suggestion, auto-confirmation — shall be stated in that probability
space, and no code path may threshold a bare cosine similarity.
This is a hard requirement rather than an implementation detail because the failure mode is
invisible. A raw cosine means something different for every model, every population, and every face
size; a threshold tuned on one library silently misbehaves on another, and an uncalibrated
similarity still *looks* like a plausible number all the way to the user interface. A displayed
confidence that does not mean what it says is worse than no confidence, because it is trusted.
The calibration shall be fitted per library from that library's own faces, and shall report whether
it is valid. Where it is not — too few examples to fit — the app shall fall back to a **documented,
published operating point** (the reference implementation's fitted curve) and shall say, at the
screen level, that the confidences come from it. What is forbidden is presenting an untuned default
*as though it were measured on this library*; withholding the number entirely is not required and
shall not be done, because a screen of unranked suggestions is the state most libraries would
permanently sit in — the fit needs confirmations, and confirmations need a ranked screen to be made
on.
*Acceptance:* on a labelled corpus, the stated probability is within a documented tolerance of the
observed match rate across the probability range (a reliability-diagram check, not a single
accuracy figure).
**FR-CULL-10 — Clustering and naming.** Detected faces shall be clustered into unnamed groups. The
user names a group, and that name applies to its members. A person is thereafter a first-class
catalog entity with a stable UUID, independent of any name given to them.
The user shall be able to **merge** two groups that are the same person, **split** a group that is
not, **remove** a face from a person, and **rename** a person, at any time and without re-indexing.
Splitting must be as easy as merging: clustering will over-merge on siblings, on parents and
children, and on the same person a decade apart, and a tool that can only merge makes its own errors
permanent.
Confirmation is explicit. A face is either **suggested** (the system's inference) or **confirmed**
(the user's judgement), and the two are never conflated in storage or in display. Suggestions may be
recomputed freely; confirmations are user data and are never overwritten by a later inference pass.
**FR-CULL-11 — People as a selector term.** A person shall be a term in the §5 selector language,
composable with every other term.
This is the requirement that pays for the subsystem, and it is nearly free once FR-CULL-10 exists:
because one predicate language serves the library filter, smart collections, and cache rules, a
person term yields all three at once — filter the grid to a person, save "every photo of Anna rated
three or higher" as a smart collection, and pin "every photo of my children" to stay local on the
tablet. The last is a genuinely new capability, not a restatement of the first two.
Selectors shall distinguish confirmed from suggested membership, defaulting to confirmed-only, so a
saved collection does not silently change membership when a later indexing pass revises a guess.
**FR-CULL-12 — Names are user data; embeddings are not.** A confirmed person name is a user
judgement of the same class as a rating or a keyword, and shall be written to the sidecar
(FR-CAT-8), so it survives catalog deletion and travels with the photograph.
Embeddings, detections, cluster assignments, and unconfirmed suggestions are **derived data**. They
live in the catalog only, are rebuildable by re-indexing, and are never written to a sidecar. This
follows ARCH §6.12 exactly: the expensive-but-reproducible artefact stays in the disposable index,
and only the irreplaceable human judgement enters the trust path.
The person UUID is what a cross-device merge keys on, in the same way collections merge (FR-CAT-7).
Two devices that independently name the same cluster produce two people; merging them is the
ordinary FR-CULL-10 merge, not a special case.
### 3.10 Extensibility and plugins
> **Post-v1, decided 2026-09-19.** Every clause in this section, and NFR-SEC-6 which exists for
> it, is marked `(post-v1)` on its defining line and is outside the v1 count. The section stays as
> the design of record — an extension point that is built as though a plugin might one day reach it
> is cheaper than one retrofitted — but nothing here is owed a tag, and §7's row is the statement
> of scope. The audit that prompted this found the register saying both things at once: §7 had
> deferred "Plugin API" in a bare row since the first draft while these 23 clauses counted against
> coverage, which measured the contradiction rather than the software. D16 defers with the section.
FR-DEV-3c already buys one form of this: an operation is added by writing a declaration, and the
develop panel grows its controls without a frontend change. This section extends that property past
the compiler. **A plugin is a file dropped into a directory; the application uses it without being
rebuilt.**
The reason FR-DEV-3c was cheap is worth stating, because it decides everything below. An
`Operation` never touches a pixel. It publishes a descriptor and returns WGSL text and a handful of
floats, and the fused shader does the work (ARCH §3.3, ARCH §5.2). Any extension point that can be
given that shape — data in, data out, no execution — gets plugins for almost nothing. Any extension
point that cannot needs a sandbox to run code in, and that is where the entire cost of this section
lives.
So the classes are ordered by expense, and the rule is: **an extension point is declarative unless
it is demonstrably impossible to make it so.**
**FR-PLG-1 — Three plugin classes.** *(post-v1)* The app shall support exactly three, and no fourth shall be
introduced without a recorded decision.
| Class | Mechanism | Covers | Executes code |
|---|---|---|---|
| **Declarative** | YAML declaration plus WGSL | develop operations, mask generators, scope computation, camera profiles, LUTs, presets, localisations | No — shader source only, validated before use |
| **View** | Slint compiled at runtime into a declared slot | toolbar and panel widgets, the drawing half of a scope | On the UI thread; not sandboxed |
| **Computational** | WebAssembly component against a versioned interface | segmentation strategies, decoders, sync backends, metadata extractors | Yes, sandboxed |
The expected distribution is heavily weighted to the first. A plugin author reaches for class 2 only
to *draw* something new and for class 3 only when an algorithm cannot be expressed as GPU passes.
**Class 3 shall be WebAssembly and shall not be native shared objects.** Three reasons, each
independently sufficient: a native object cannot be loaded on Android (§1.2 names it a primary
platform) or on any future iOS build; Rust has no stable ABI, so a native plugin would be locked to
one compiler version and one build of every crate it touches; and an in-process native object holds
the whole application's privileges, which would make NFR-SEC-4 a promise about third-party code
rather than a property of the system.
**FR-PLG-1a — Plugins orchestrate; the GPU does pixels.** *(post-v1)* No plugin interface shall pass
full-resolution pixel data across a sandbox boundary. A computational plugin receives handles and
small buffers, and expresses per-pixel work as class-1 shader passes it declares. This is what keeps
WebAssembly's arithmetic penalty irrelevant, and it is also ARCH §6.1 applied to plugins: a plugin
must not be the reason a result round-trips through the CPU.
#### Declarative plugins
**FR-PLG-2 — The node declaration is the plugin format.** *(post-v1)* The schema documented in
`core/dr-pipeline/ops/README.md` — parameters, uniform expressions, WGSL, helpers, activity,
presentation, attributes, order, tests — shall be readable at **load time** as well as build time,
from a plugin directory, with no change to what a declaration means.
One consequence is deliberate: a bundled operation and a third-party plugin are the same kind of
thing, differing only in where the file was found. There is no second, weaker format for outsiders,
and no path by which the bundled operations acquire capabilities plugins cannot reach.
The build-time path is not removed. Operations that ship with the app remain compiled, because a
generated `match` is faster than an interpreted one and because their tests must run under
`cargo test`. The two paths shall produce descriptors indistinguishable to everything downstream, in
the same way and for the same reason that a declared node is today indistinguishable from a
hand-written one.
**FR-PLG-2a — Fragment nodes and pass nodes.** *(post-v1)* Two templates, and a plugin author chooses by
answering one question: does this operation need to read a pixel other than its own?
| | Fragment node | Pass node |
|---|---|---|
| Declares | a WGSL body over `c` | a complete compute shader with input and output textures |
| Cost | none — fuses into the existing dispatch | one full-resolution texture round-trip |
| Reaches | per-pixel colour transforms | neighbourhood operations — blur, clarity, halation, structured grain |
Fragment nodes shall continue to compose into a single dispatch (ARCH §5.2). A pass node breaks that
fusion **for itself only**: the fragments before it and after it still fuse into one dispatch each.
A declaration shall state which it is, and the cost shall be visible to the user in the plugin
listing, because a chain of pass nodes is how a fast application becomes a slow one without any
single decision having been wrong.
**FR-PLG-2b — Mask generators are declarative.** *(post-v1)* `Linear` and `Radial` mask sources are already
geometry in normalised coordinates rasterised by a shader (ARCH §5.4). A plugin shall be able to
contribute a mask generator on the same terms — declared parameters plus a WGSL function from
normalised coordinates to coverage — reaching luminosity-range, colour-range, and further gradient
forms with no code. Mask sources that are *identity into a segmentation* (`Regions`, `Subject`)
are not declarative and belong to class 3.
**FR-PLG-2c — Scopes split at the existing seam.** *(post-v1)* A scope plugin is a class-1 compute shader
producing a small bin buffer plus a class-2 view drawing it. This is the split already in place for
the histogram — `dr-gpu` counts, `dr-ui` shapes, Slint draws — and it holds for waveform,
vectorscope and RGB parade without change. The counting half shall not read back full-resolution
pixels (ARCH §5.5).
**FR-PLG-2d — The vocabularies stay closed.** *(post-v1)* `WidgetKind`, `attributes`, and the parameter `kind`
list remain closed enumerations, and a plugin may use them but shall not extend them. The reason
given in the node README strengthens rather than weakens here: a typo that creates a new category
containing exactly one control is indistinguishable from a deliberate new category until somebody
notices. An unrecognised attribute shall place the operation in a clearly-labelled fallback group
and warn — never silently omit it, which would make a control that does not exist look like a
control that was never written.
#### View plugins
**FR-PLG-3 — Declared slots, typed contracts.** *(post-v1)* A view plugin shall be a Slint component compiled at
runtime and instantiated into a **named slot** the application declares — not a licence to draw
anywhere in the window. Each slot states the data it provides and the callbacks it accepts, and a
component that does not match its slot's contract shall be rejected at load with a message naming
the mismatch.
Slots are a closed list under the same reasoning as FR-PLG-2d, and the composition rules of
ARCH §4.3a continue to apply: a slot describes what a view *is for*, never how much room it has.
**FR-PLG-3a — A view plugin cannot be trusted with the UI thread.** *(post-v1)* Class 2 is the one class with no
sandbox — an interpreted component runs on the UI executor and can violate NFR-ARCH-1 by looping.
The application shall therefore watchdog slot rendering, disable a component that exceeds a stated
budget, and report which plugin was disabled. A view plugin that fails shall leave the slot empty
and the application usable; it shall never take down the window.
#### Computational plugins
**FR-PLG-4 — One interface per extension point, versioned.** *(post-v1)* Each class-3 extension point shall
define an explicit interface, versioned independently, and a plugin shall declare which version it
implements. Interfaces are the only surface a computational plugin can reach: there is no ambient
filesystem, no network, and no access to the catalog.
**FR-PLG-4a — Capabilities are granted, never assumed.** *(post-v1)* A plugin that needs to read a file or
reach the network shall declare the capability, and the user shall grant it explicitly with the
reason shown. A plugin's declared capabilities shall be visible before installation, and a plugin
that requests none — which is the expected case for a segmentation strategy — shall be installable
without a security decision.
This is what makes NFR-SEC-4 hold under a plugin ecosystem. A segmentation plugin that cannot open a
socket cannot send a photograph anywhere, and that is a structural property rather than a promise.
#### Versioning
**FR-PLG-5 — Three independent version numbers.** *(post-v1)* Conflating any two of these produces a wrong
answer in both directions, so the format shall carry all three.
| Number | Owned by | Governs | Moves when |
|---|---|---|---|
| **Interface version** | the application | whether the plugin loads at all | the host changes what it offers |
| **Plugin version** | the author | updates, provenance, and the trust record | any release |
| **Parameter schema version** | the author | whether an existing sidecar still reads | parameters change incompatibly |
The common case is a plugin renaming a parameter: sidecars break while the interface version never
moves. The converse also occurs. Binding sidecar compatibility to the interface version would be
wrong in both.
**FR-PLG-5a — Declared, never inferred.** *(post-v1)* A plugin shall state its interface version explicitly.
Deducing it from which keys are present produces files that are ambiguous between two versions, and
the ambiguity surfaces years later as a wrong render rather than as an error.
**FR-PLG-5b — A supported window, and a written policy.** *(post-v1)* The application shall support the current
interface version and at least one predecessor, and the deprecation policy shall be stated in the
plugin authoring documentation rather than decided per release under pressure.
**Additive changes shall not bump the version.** A new optional key is compatible because absent
means default — the rule `active:`, `presentation:` and `define:` already follow. Holding that
discipline is what keeps the version number nearly stationary.
**FR-PLG-5c — Adapt at the boundary, normalise inward.** *(post-v1)* A plugin declaring an older interface
version shall be adapted at load into the current internal representation, and nothing downstream
shall be able to tell. Version branches threaded through the pipeline are how this becomes
unmaintainable; the single adaptation point is the same discipline that lets a declared node and a
hand-written one be one thing by the time anything reads them.
**FR-PLG-6 — Migrations are data.** *(post-v1)* Because every parameter is an `f32` addressed by a flat
`op_id.param_id` key, a schema migration is a rewrite table rather than code. A plugin shall be able
to declare migrations between consecutive parameter schema versions, supporting at minimum rename,
rescale, and default-for-a-new-parameter.
The **application** applies them, chained, at load, so that everything downstream sees only
current-schema parameters. Migrations shall be testable through the same declared-test mechanism as
the node itself.
**FR-PLG-6a — `params_version` is a promise, not a hint.** *(post-v1)* Bumping it locks every older build out of
the edits that use it (FR-PLG-9). It shall be bumped only when an older build would genuinely
*misread* the file, and never merely because a parameter was added — absence already means default,
which already means neutral. This obligation belongs in the authoring documentation in as many words.
#### Sidecars
**FR-PLG-7 — The sidecar records identity, not location.** *(post-v1)* Each version block shall record, for
every plugin it depends on: the plugin id, its version, its parameter schema version, and a content
hash of the artefact. These merge key-wise like every other line in the format (FR-NC-9).
**It shall not record an install URL.** Sidecars arrive from elsewhere — they sync (FR-NC-9), they
travel with shared photographs, they come from other people's catalogs. A sidecar that names where
to fetch code lets whoever wrote it choose what the user is prompted to install, which is a
confused deputy with a friendly dialog in front of it. Resolution from identity to location is a
decision the user made when they configured a registry, not one an incoming file makes for them.
The content hash carries a second benefit: "install filmic 2.1" resolves to exactly the bytes the
original edit was rendered with, which makes substitution and silent render drift detectable rather
than merely regrettable.
**FR-PLG-8 — A missing plugin shall never cost an edit.** *(post-v1)* The sidecar already preserves lines it
does not understand verbatim and writes them back untouched, so a machine lacking a plugin cannot
destroy an edit that uses it. That property is now load-bearing and shall be treated as such.
Beyond preservation:
- **Alert, aggregated.** Missing plugins shall be reported once per import or session and listed in
one catalog-wide view — never once per image. A folder of five hundred synced photographs sharing
one missing plugin is one notice.
- **Non-blocking.** Nothing here is urgent, because the edit is safe either way. The image shows a
clear mark that an operation is unavailable; the install dialog appears when the user asks for it.
- **Never during unattended work.** No such prompt shall interrupt a background sync or a batch
export, where a dialog becomes either a stalled job or a reflex click.
- **Render without, never render a guess.** The image shall be rendered omitting the unavailable
operation. Interpreting its parameters under different semantics is forbidden: a plausible wrong
render and a correct one look equally plausible, which makes the wrong one the more dangerous
output.
- **Export is gated harder than display.** Exporting an edit with an unavailable operation shall
require explicit acknowledgement. A slightly wrong screen is recoverable; a delivered file that
silently omits an adjustment is not.
*Acceptance:* a sidecar written with a plugin installed, opened and saved on a machine without it,
is byte-identical to the original.
**FR-PLG-9 — Forward skew is quarantined, not guessed.** *(post-v1)* Where a version block names a parameter
schema version **newer** than the installed plugin declares, the application shall divert that
operation's parameters into the preserved-verbatim path *before applying any of them*, and mark the
version quarantined.
This does not fall out of FR-PLG-8. A forward-skewed operation is a **recognised** id, so the
existing unknown-key path never sees it: its parameters parse, apply under the older meaning, render
plausibly, and are written back — silently downgrading the edit, on every device it syncs to. The
hazard is the write-back, not the display.
Quarantine means, precisely:
- **No apply, no edit, no save, no export** for that version. These are the operations that lose data
or ship a wrong file.
- **The photograph still opens.** Decoding needs no plugin, and refusing to display a file because
one adjustment is from the future holds the photograph hostage over an edit.
- **Per version, not per image.** A sidecar holds several versions (FR-CAT-12); one may be
quarantined while the others open normally.
- **An upgrade is offered**, resolved through FR-PLG-10 like any other install — and it shall be
allowed to fail. Offline, declined, or delisted all fall back to quarantine, never to deletion and
never to opening anyway.
- **Bundled plugins say so.** Where the plugin ships with the application, the remedy is an
application update, and the message shall say that rather than offering an install that cannot help.
*Acceptance:* a sidecar written by a newer plugin version, opened, browsed, and closed on a build
with an older one, is byte-identical afterwards.
#### Distribution
**FR-PLG-10 — Registry resolution and verified artefacts.** *(post-v1)* Installation shall resolve a plugin id
through a registry the user has configured, with a default registry shipped. The application shall
verify the artefact against the hash or signature the registry states before loading it.
- **An unrecognised id is the loud case.** Where no configured registry knows the plugin, the
application shall say so and show what the sidecar claims, as text, for the user to act on
deliberately. There shall be no one-click install of a location supplied by a file.
- **Provenance appears in the prompt.** Plugin name, version, registry, and publisher, so "from the
registry you trust" and "from somewhere you have never heard of" do not look alike.
- **Offline degrades, it does not block.** With no network: say what is missing, keep the edit
intact, open the photograph. An application whose privacy claim is local-only shall not need the
internet to show a picture.
A release page on a code-hosting service is a distribution mechanism, not an identity. Binding
identity to one would give dead links on a rename, no mirroring, no offline install, and a
dependency on one company's availability.
#### Authoring and operations
**FR-PLG-11 — Plugins are validated, and validation is the author's tool.** *(post-v1)* A declared plugin's
`tests:` shall be runnable outside the application, against the same interpreter that loads it, via
a command-line validator. The application shall ship a scaffold command producing a minimal working
plugin of each class.
Load-time failures shall be collected per file and reported with the key that was wrong, in the
manner `build.rs` already reports them. **A malformed plugin shall be skipped, never fatal.** The
build script's exit-on-error discipline is right for an author with a compiler open and wrong for a
user opening their library.
**FR-PLG-12 — Failure is attributable and revocable.** *(post-v1)* The application shall record per-plugin
timing and error counts, surface them in the plugin listing, and allow any plugin to be disabled
without uninstalling it — including on the next launch after a crash, so a plugin that prevents
startup can be disabled by someone who cannot start the application.
*Acceptance:* with a plugin deliberately made to fail at load, at render, and at UI paint, the
application starts, opens a photograph, names the responsible plugin, and continues.
---
## 4. Non-functional requirements
### 4.1 Performance targets
These are targets to design against and measure, on the reference desktop
(AMD Threadripper 2920X, 24 threads, discrete GPU) and a mid-range Android device.
| ID | Metric | Desktop target | Android target |
|---|---|---|---|
| **NFR-P1** | Catalog open (50k images) | < 2 s | < 4 s |
| **NFR-P2** | Grid scroll | Sustained 60 fps | Sustained 60 fps |
| **NFR-P3** | Thumbnail generation throughput | ≥ 100 img/s (embedded preview path) | ≥ 25 img/s |
| **NFR-P4** | Open image in develop (to first proxy on screen) | < 400 ms | < 1000 ms |
| **NFR-P5** | Slider adjustment → visible update | < 16 ms (one frame) | < 33 ms |
| **NFR-P6** | Pan/zoom responsiveness | No dropped frames at 60 fps | No dropped frames |
| **NFR-P7** | Full-resolution export (24MP, full chain) | < 2 s | < 8 s |
| **NFR-P8** | Idle memory (50k catalog, nothing open) | < 500 MB | < 250 MB |
| **NFR-P9** | **UI-executor** blocking (not "any operation") | Never > 16 ms | Never > 16 ms |
| **NFR-P13** | **Next image in culling mode** (FR-CULL-1) | **< 50 ms** | **< 50 ms** |
| **NFR-P14** | Focus peaking overlay ready | < 100 ms after preview | < 150 ms |
| **NFR-P15** | Drawn mask stroke → visible (ARCH §6.11) | < 16 ms, no cursor lag | < 16 ms |
| **NFR-P10** | Touch gesture → visual response | < 16 ms | < 16 ms |
| **NFR-P11** | Layout class transition (window resize) | No dropped frames. No loss of *photographic* state: the open image and version, the selection, scroll position, the in-progress edit and its undo history, and the current mode. Panel disclosure is explicitly exempt — see below | n/a |
| **NFR-P12** | Warm-start shader pipeline setup (cached) | < 100 ms | < 100 ms |
Every target above requires a stated measurement method, workload, and pass threshold before it is
testable. NFR-P8 in particular must state whether it measures RSS inclusive or exclusive of GPU
allocations, and whether it holds after SQLite's page cache warms on a 50k catalog.
**On what NFR-P11 means by state.** It said "no state loss", which the implementation contradicts on
purpose, so the requirement has been made specific rather than left to be read as forbidding
something it should not. `apply_layout_class` discards the user's panel open/closed choices when the
class changes, and the argument for that is sound: a choice made in landscape answers a different
question from the one portrait asks, and carrying it across is how a photographer ends up with a
232 px sidebar on a screen with no room for it and no memory of having asked for it. Panel
disclosure is a *default*, re-derived per class, with the user's disagreement remembered only within
the class where it was expressed.
Everything the photographer produced or navigated to is a different matter, and none of it may be
touched by a resize. That is the list in the criterion, and it is the testable half.
**Performance regressions fail the build.** §9's benchmark suite runs per-commit; a regression
beyond a stated tolerance is a build failure, not a notification. Performance work rots otherwise.
### 4.2 Reliability
**NFR-R1** — No operation shall lose user edit data. The catalog database uses write-ahead
logging and survives power loss without corruption.
**NFR-R2** — The catalog is backed up automatically on a schedule and before schema migrations.
**NFR-R3** — A crash in decode or GPU work shall not take down the application where it can be
isolated; the affected image is marked as failed and the app continues.
**NFR-R4** — Source image files are strictly read-only to the application, except where the user
explicitly requests a destructive operation (e.g. delete, or writing XMP sidecars).
**NFR-R5 — Schema versioning.** The catalog carries a monotonic schema version. Migrations are
forward-only, transactional, and idempotent on retry. Each migration ships with a test that migrates
a fixture catalog from **every** prior released version. The app refuses to open a catalog with a
newer schema version rather than corrupting it, and says so.
ARCH §6.6 already anticipates one migration (folder ETags); there will be others, and the machinery must
exist before the first one.
**NFR-R6 — Corruption recovery.** On failing an integrity check at startup, the app offers restore
from the NFR-R2 backup, and failing that, rebuild from sources plus sidecars per invariant 5.2.4.
FR-CAT-8 is what makes that second path real for local-only users.
**NFR-R7 — GPU device loss.** The GPU layer treats device loss as an expected event (ARCH §6.10): detect
it, tear down and recreate the device and all derived resources, and re-drive the current render
from the edit graph. **No user edit is lost.** Recovery is exercised by a test that induces device
loss mid-render.
**NFR-R8 — No suitable GPU.** Where no Vulkan device meeting the NFR-COMPAT-1 baseline is
available, the app starts in a stated degraded mode with defined capability limits rather than
failing to launch.
**Decided 2026-09-19: there is no CPU render pipeline.** The degraded mode is the one the
viewer already has (`dr_ui::shared_gpu` returning `None`): the library opens, the grid and the
culling views run on embedded previews and cached proxies, metadata, ratings and collections are
fully editable, and develop and export are unavailable and say so. "CPU fallback" in this
document means nothing more than staging through host memory when GPU memory is short; it never
means a second implementation of the operations. ARCH §6.4 stands as written, and NFR-RES-2's
fallback clause has been reworded to match. A second full pipeline was the alternative, and it was
declined for the reason ARCH §6.4 gives: the GPU path is the product, not an optimisation of it.
### 4.3 Resource behaviour
**NFR-RES-1 — Bounded memory.** Memory use is bounded and configurable, independent of catalog
size and image count. Caches are evictable under pressure.
**NFR-RES-2 — GPU memory.** The pipeline shall handle images larger than available GPU memory by
tiling. GPU memory headroom is configurable. Where an allocation fails, work is staged through host
memory or refused with a typed error (`GpuError::TooLarge`) — never rendered by a CPU pipeline,
which does not exist (NFR-R8).
**NFR-RES-3 — Mobile power.** On Android the app shall not render continuously when idle. Battery
and thermal behaviour are first-class concerns; background sync respects metered-connection and
battery-saver settings.
**NFR-RES-4 — Disk cache.** Thumbnail and proxy caches have a configurable size cap with LRU
eviction.
### 4.4 Portability
**NFR-PORT-1** — Platform-specific code is isolated behind interfaces. The image core, catalog,
and edit pipeline contain no platform conditionals.
**NFR-PORT-2** — GPU shaders are authored once and used on both platforms.
**NFR-PORT-3** — Adding a third platform requires implementing the platform interfaces only, not
changes to the core.
### 4.5 Security and privacy
**NFR-SEC-1** — RAW parsing treats input as untrusted. Parser hardening, fuzzing, and where
practical process or memory isolation for decode.
**NFR-SEC-2** — Credentials are never written to the catalog, logs, or plain files. Platform
secure storage only.
**NFR-SEC-3** — All network traffic uses TLS with certificate validation. No option to disable
validation in release builds.
**NFR-SEC-4** — No telemetry without explicit opt-in.
**NFR-SEC-5 — Face data stays on the user's own hardware.** Face embeddings (FR-CULL-8) are handled
under a stricter rule than the rest of the catalog.
This is a personal tool for personal libraries (§1.1, D11) — the people in these photographs are the
user's family and friends. That is the reason for the rule, not a reason to relax it: the data is
sensitive precisely because it is personal, and the user is the only party with any claim on it.
- **Never leave the device by default.** Embeddings, face crops, and cluster assignments shall not be
transmitted, uploaded, or included in any diagnostics bundle (NFR-OPS-1) or crash report
(NFR-OPS-2), under any configuration. The diagnostics path has no opt-in for this; it is excluded
outright.
- **Sync is opt-in and separately consented.** Syncing embeddings to the user's own Nextcloud is
permitted — it is their server and their photographs, and it saves re-indexing a library per device
— but it is off by default, is not implied by enabling photo sync, and the consent states in plain
language what is being uploaded and why. Person *names*, being sidecar data (FR-CULL-12), sync with
the sidecar as ordinary metadata.
- **No third-party inference.** Face detection and embedding run locally. No image, crop, or
embedding is sent to a remote inference service, and the app ships no capability to do so.
- **Deletable, in one action.** The user shall be able to delete all face data — embeddings,
detections, clusters, and people — from a single control, without deleting the catalog or any
photograph, and to disable face indexing entirely so that no such data is produced.
- **Model weights are inspectable.** The models used shall be named and versioned in the about
screen, with their licences, so a user can determine what is running on their photographs.
*Rationale:* the rest of this document treats privacy as a property of the network boundary — TLS,
credentials in secure storage, opt-in telemetry. Face data needs more than a well-defended boundary,
because it is not revocable once it has crossed one, and because it describes people who are not the
user. The prohibition is therefore structural rather than configurable: the code paths that would
upload an embedding to anyone but the user's own server do not exist. A setting can be changed by
accident, or by a future maintainer who has forgotten why it was there; an absent code path cannot.
**NFR-SEC-6 — Third-party code runs inside a boundary, not beside the application.** *(post-v1)* FR-PLG-1
admits code the user did not write and the project did not review. The privacy properties asserted
elsewhere in §4.5 are properties of *this* codebase, and none of them survives a plugin that can
open a socket.
- **Sandboxed by default, with granted exceptions.** Computational plugins execute with no
filesystem, no network, and no catalog access. Anything more is a declared capability, shown
before installation and granted explicitly by the user (FR-PLG-4a).
- **No plugin reaches face data.** Embeddings, detections, crops, and cluster assignments are
outside every plugin interface, under NFR-SEC-5's structural rule: the code path does not exist,
so no grant can create one.
- **Artefacts are verified before they are loaded**, against the hash or signature a configured
registry states (FR-PLG-10). An unverifiable artefact is not loaded.
- **Nothing is fetched on a file's say-so.** Installation is resolved through user-configured
registries; a sidecar carries identity only (FR-PLG-7).
- **Native shared objects are not a plugin format** (FR-PLG-1). An in-process native object holds
the application's full privileges, which would make every rule above unenforceable.
*Rationale:* an extensible application inherits the trust properties of its weakest plugin unless
the boundary is structural. The same reasoning as NFR-SEC-5 applies for the same reason — a setting
can be changed by accident or by a maintainer who has forgotten why it was there, and an absent
capability cannot.
### 4.6 Execution model
**NFR-ARCH-1 — Named executors.** The app defines distinct executors — UI, GPU submission, decode
pool, I/O pool, network — with stated thread counts and the invariant that **no blocking call
occurs on the UI executor**. This is the mechanism behind R4 and NFR-P9, which currently assert an
outcome with no stated means.
**NFR-ARCH-2 — Scheduler priority.** The tiling scheduler assigns priority classes, with
visible-tile work **strictly preempting** background export and thumbnail work. Without this,
NFR-P5's slider latency fails during a batch export — the common case, not an edge case.
**NFR-ARCH-3 — Cancellation.** Cancellation is cooperative with a bounded worst-case latency
(target: observed within 100 ms), and covers in-flight GPU submissions. Every long-running
operation named in FR-CAT-1, FR-EXP-7, and FR-NC-6 is cancellable under this model.
**NFR-ARCH-4 — Error propagation.** No worker error may panic the process. Errors surface as typed
results attached to the affected image or job, consistent with NFR-R3 and FR-RAW-4.
### 4.7 Operations
**NFR-OPS-1 — Diagnostics.** Structured levelled logging to a rotating, size-capped on-disk log in
the XDG state directory or Android app directory, with **automatic redaction of credentials and
tokens** (required by NFR-SEC-2, which currently forbids credentials in logs that are never
otherwise specified). A one-click diagnostics bundle includes log, schema version, GPU and driver
identification, and app version — with an explicit preview-and-consent step before anything leaves
the device.
**NFR-OPS-2 — Crash reporting.** Local crash capture always; upload only on explicit opt-in
(NFR-SEC-4).
**NFR-OPS-3 — Preferences.** A single versioned preferences store, **separate from the catalog**, so
preferences survive catalog rebuild and multiple catalogs. At least eight requirements refer to
configurable settings with no store defined. Device-specific settings (GPU headroom, cache caps) do
not sync between devices.
**NFR-OPS-4 — Update and first run.** State delivery channels and their update mechanisms. This
matters concretely because D2 pins rawler at an alpha, non-SemVer version whose camera-support fixes
users will need. First-run flow is defined, including platform permission acquisition (FR-PLAT-AND-1
makes first run a permission negotiation on Android, not a welcome screen) and initial root
selection.
### 4.8 Compatibility baseline
**NFR-COMPAT-1 — Supported hardware.** Every §4.1 Android figure is meaningless without this. State:
- Minimum and target Android API level (targetSdk 36 is currently required for Play distribution)
- Minimum Vulkan version and the required feature and limit set — including whether `shaderFloat16`
and 16-bit storage are required, since **FR-DEV-2's f16 pipeline depends on them and their
absence would jeopardise R1**
- Minimum device RAM, and minimum desktop Vulkan/Mesa versions
- The **specific** reference Android device the §4.1 column is measured on, plus a secondary device
from a different GPU vendor
Adreno, Mali, and PowerVR diverge significantly in compute behaviour and in external-memory interop
— exactly what spike S1 tests. The spec already applies this reasoning to desktop drivers; it
applies at least as strongly on Android.
**NFR-COMPAT-2 — Distribution channels.** State the v1 channels (e.g. Flatpak and AppImage on Linux;
Play Store and/or F-Droid on Android). Play distribution is what makes ARCH §6.9's constraints binding —
a sideloaded or F-Droid build could use different permissions, so the channel decision and the
storage design are coupled.
### 4.9 Accessibility and internationalisation
**NFR-A11Y-1 — Localisation.** All user-facing strings, including operation and parameter labels
resolved from `LocalizedString` (FR-DEV-3a), are externalised and translatable without
recompilation. Note the constraint this creates: those labels live in core crates that **cannot
depend on the UI** (ARCH §6.5a), so the localisation mechanism must itself be UI-independent. State the
format, the locale-resolution rule, and whether RTL layout is in v1 scope.
**NFR-A11Y-2 — Accessibility.** Controls expose accessible names, roles, and values to the platform
accessibility layer (AT-SPI on Linux, TalkBack on Android). Platform font scaling is honoured
without clipping. Non-canvas UI meets WCAG AA contrast.
**Slint's accessibility support on Android requires verification** — this may be a toolkit gap, and
it is far cheaper to discover now than after the UI is built.
**NFR-A11Y-3 — Colour-independent status.** No status is conveyed by hue alone. Colour labels,
clipping indicators (FR-DSP-7), and the HSL mixer carry a shape or text affordance. This matters
more in a colour-grading application than in most software.
---
## 5. Data model and architecture
Entity definitions, invariants, and all architectural constraints are specified in
[architecture.md](architecture.md) — §3 (core abstractions), §6 (data architecture), and
§11 (architectural constraints).
Requirements in this document that depend on an architectural guarantee cite it inline. The
constraints most load-bearing for testability are:
| Constraint | Why a requirement depends on it |
|---|---|
| ARCH §6.1 — no CPU round-trip | FR-DSP-3, FR-DSP-7, NFR-P5 are unachievable without it |
| ARCH §6.11 — GPU-rasterised masks | NFR-P15 (no brush lag) |
| ARCH §6.12 — sidecars authoritative | NFR-R6, invariant behind FR-CAT-8 |
| ARCH §6.9 — Android SAF only | FR-CAT-1a, FR-PLAT-AND-1, and the NFR-P1/P3 Android figures |
| ARCH §6.6 — no sync tokens | FR-NC-4 |
| ARCH §6.13 — integer-only bit-identity | R1's tolerance, §9 golden images |
## 6. Decisions
Rationale, evidence, and the eliminated alternatives are recorded in
[architecture.md §12](architecture.md). Outcomes only:
| # | Decision | Outcome |
|---|---|---|
| D1 | Language and UI framework | Rust + Slint, rendering through wgpu |
| D2 | RAW decoder | rawler; LibRaw fallback behind a trait |
| D3 | First milestone | Delivered — [milestone-v0.1.md](milestone-v0.1.md), closed 2026-08-30 |
| D4 | Nextcloud sync mechanism | ETag pruning, chunked upload v2, Login Flow v2 |
| D5 | Colour management | lcms2 + GPU-side matrix/LUT transforms |
| D6 | Shader authoring | Hand-written WGSL |
| D7 | Network stack | reqwest + quick-xml |
| D8 | Licence | **GPLv3** |
| D9 | Operation UI model | Declarative parameter descriptors |
| D10 | Interface strategy | One adaptive UI, tablet + desktop |
| D11 | Product positioning | Culling-first differentiator; see below |
| D12 | Scope versus pace | **DECIDED 2026-09-19** — settled by events; full scope stands, no v1 date |
### D11 — product positioning
Settled by requirements calibration, 2026-08-08.
| Dimension | Decision |
|---|---|
| Audience | RAW-literate photographers, Linux-comfortable. Docs matter; hand-holding does not. |
| Library scale | 10k–50k images |
| Culling | **The core differentiator** (§3.9) |
| Focus checking | Peaking *and* zoom |
| Ingest | Full workflow — template rename, checksum verify, dual-destination |
| Colour defaults | Good, not obsessive — matrices plus per-body base curve |
| Film simulation | Fujifilm explicitly targeted |
| AI | Denoise in v1; masking deferred. Per-face eye state and head pose are in v1 **as culling evidence, not AI** (FR-CULL-8a, FR-CULL-13); gaze deferred (§7) |
| Local adjustments | Full masking, GPU-rasterised |
| Sync | The reason the project exists |
| Durability | Sidecar-first |
| Licence | GPLv3 |
| Pace | Evenings and weekends, indefinite |
### D12 — scope versus pace · **DECIDED 2026-09-19**
**Settled by events.** The reconciliation this decision asked for never happened as a decision;
it happened as a build. Between the calibration and this audit the milestone D3 pointed at was
delivered and closed (2026-08-30), and the application went through eleven further releases to
0.12.2 carrying culling, develop, masks, faces, sync, Flatpak and a Windows channel. The scope as
calibrated stands as written, and the pace is the pace. What was reconciled is the *date*: v1 has
none. A requirement in this document is in scope until §7 says otherwise, and "post-v1" in §7 is
the only mechanism by which something leaves the count — used on 2026-09-19 for the plugin API,
and for nothing else.
The two tensions below are kept as the record of what the decision weighed. Both resolved
themselves the same way: the expensive selection was built anyway, and the differentiator was
built alongside the develop chain rather than instead of it.
The calibration selected an ambitious feature set — full tablet editing, full ingest, culling as a
differentiator, complete GPU masking, AI denoise, Fuji-first colour, deep sync, sidecar durability
— against a stated pace of evenings and weekends, indefinitely.
Those are not compatible as stated. This is not an argument against any individual choice; it is
that the two must be reconciled before a build order can be set.
Two specific tensions:
**1. Full tablet editing is the most expensive selection**, chosen against the research finding
that volume photographers do not edit on tablets. It carries SAF storage at unproven scale (S10),
process-death durability, background execution limits, and two GPU vendors to validate. Tablet
*culling* is the validated workflow, is what the differentiator points at, and costs a fraction as
much.
**2. The v1 milestone and the differentiator disagree.** D3's vertical slice proves the
architecture but is useful to nobody. If culling is what no existing tool does well, a culler is
both a smaller build and a usable one — it needs no develop chain.
D12 was what set D3, and [architecture.md §11](architecture.md)'s build order followed it.
### D15 — target devices · **DECIDED 2026-08-22**
**A 12-inch tablet and a desktop. No phone.**
Recorded because it is load-bearing for the interface and invisible in the
code. Every phone-shaped answer — a bottom tool strip, a one-tool-at-a-time
sheet, thumb-reach zones — is designing for hardware this project does not
target, and each would have cost a second layout to keep in step with the
first.
What survives the decision is the *input* difference rather than the size one:
a 12-inch tablet is touched, and NFR/FR-UI-7's position already covers it —
hit regions grow to the modality, layout does not move. The practical rules
that fall out are in `docs/ui-navigation.md` D-N2: no hover-only affordance
and no modifier key may be the sole route to anything, because a tablet has
neither.
`EXPANDED_MIN_WIDTH` is 820 logical pixels and a 12-inch tablet is ~1024
across in portrait, so **both orientations of both targets are the expanded
layout**. The compact class remains as graceful degradation for a narrowed
desktop window, not as a second interface.
---
### D13 — face inference runtime and model licensing · **RUNTIME ANSWERED · LICENSING POSITION RECORDED 2026-09-19**
> **Position, 2026-09-19.** DarkRoom is non-commercial software, built and installed by its
> author for personal libraries, and it uses the InsightFace SCRFD detectors and ArcFace embedder
> under their **non-commercial research grant** as such. That is the position, and it is taken
> with the risks written down rather than around them:
>
> - The grant is a **use** restriction and binds every user of the app, not only the project. It
> is not GPL-compatible and cannot become so; D8's licence covers this codebase and not those
> weights, which `models/face/README.md` says in as many words.
> - It is incompatible with **every public channel** — Flathub, F-Droid, Play. NFR-COMPAT-2's
> channels are therefore all self-distribution (a CI-built APK, a Flatpak and an Arch package
> installed by hand, an NSIS installer), and nothing is published to a store or a catalogue
> while these weights are in the tree. Publishing is the event that reopens this decision, and
> S14's licence search is what would close it: the OCEC eye-state weights showed a clean chain is
> possible, and a clean detector and embedder are what is missing.
> - The about screen names the models and their grant (NFR-SEC-5), so the person running the app
> can read the restriction they are under.
>
> The runtime half is unchanged below.
> **Updated 2026-08-21.** The runtime half of this decision is settled, and by a route the table
> below does not contain. `ort` 2.0's `alternative-backend` feature *disables its linking entirely*
> and lets another engine supply the `OrtApi`; `ort-tract` supplies it from `tract`, which is pure
> Rust. So the third option's operator coverage comes with the first option's dependency profile —
> no C, no NDK problem, no exception to the policy. Measured on a real graph before being relied on:
> YOLO26n-seg loads with zero unsupported operators and runs 640×640 in ~470 ms of CPU
> (docs/segmentation.md §13). §3.9.1's detector and embedder are different graphs and their coverage
> has not been checked, but the *approach* no longer needs a decision.
>
> **The licensing half is untouched.** The InsightFace weights are still non-commercial and still
> unusable here. That remains what S14 has to resolve first.
>
> **Updated 2026-09-19.** Two further models were read for FR-CULL-8a. **Eye state:** OCEC
> (PINTO0309) is MIT for code and weights and trained on ODC-By 1.0 and Apache 2.0 data — the
> first face-adjacent weights found with a clean chain end to end, and the smallest by two orders
> of magnitude. **Gaze:** none. MobileGaze, L2CS-Net and their descendants carry MIT on the
> repository and Gaze360 in the weights, and Gaze360's research licence restricts *"models trained
> on dataset"* by name; MPIIGaze and ETH-XGaze are no better. Head pose needs no weights at all.
> None of this moves the detector and embedder, which are still the buffalo grant and still what
> S14 resolves first: clean eye weights behind an unshippable detector ship nothing.
§3.9.1 needs to run two neural networks locally. That collides with two settled positions, and
neither collision is small enough to leave implicit.
**1. The pure-Rust dependency policy.** Every dependency choice in this project has gone the same
way, for the same stated reason: rustls over aws-lc-rs, bundled SQLite over the system library, a
Rust Lensfun port over liblensfun, zune-jpeg over libjpeg — no C dependency to satisfy under the
Android NDK (D1's whole premise). The obvious way to run ONNX models is the ONNX Runtime C++ library,
which would be the largest exception to that policy in the codebase, and it would land on the
platform the policy exists to protect.
The options, in the order I would try them:
| Option | Cost |
|---|---|
| **wgpu compute**, models hand-ported to WGSL | No new dependency at all — the GPU device and shader infrastructure already exist (ARCH §5). Highest implementation effort, and a ViT is a lot of shader. |
| **`burn`** with the wgpu backend | Pure Rust, uses the existing GPU. Young, and ONNX import maturity needs checking against these two specific graphs. |
| **`ort`** (ONNX Runtime bindings) | Fastest to working code, best operator coverage. Reintroduces the C dependency and the NDK cross-compilation problem the policy avoids. |
The tension is real: the cheapest path is the one that breaks the rule. This is worth an explicit
decision rather than a default, and S14 is what informs it.
**2. Model licensing is a distribution blocker, not a detail.** The obvious pretrained weights are
not redistributable under GPLv3. The InsightFace "buffalo" family — ArcFace and the SCRFD detector,
the standard choices — are **licensed for non-commercial research use only**, which is incompatible
with this project's licence and with Flatpak, F-Droid, and Play distribution (NFR-COMPAT-2). Other
candidate weights need their licences read individually rather than assumed.
Two ways out, both with costs:
- **Find permissively-licensed weights** and ship them in-tree. Clean, offline-first, consistent with
how the Lensfun database ships. Requires that suitable weights exist at acceptable accuracy.
- **Download models on first use**, with the user accepting the upstream licence. Sidesteps
redistribution but adds a network dependency to a feature that is otherwise entirely local, needs a
hosting story, and sits badly with the local-first posture of NFR-SEC-5.
**This must be resolved before implementation, not during it.** Discovering at packaging time that
the feature cannot ship is the expensive failure, and it is entirely avoidable — it is a licence-
reading exercise, not a research question. S14 therefore puts it first.
*Prior art available:* `../scene-actor-extraction` is a working implementation of this pipeline
(SCRFD detect → 5-point align → 512-d embedding → Platt-calibrated similarity), benchmarked at 67.4%
macro-F1 on held-out films. Its C++ does not port — different language, OpenCV and TensorRT
dependencies — but its **design decisions do**, and they are the expensive part: the calibrated
probability space that FR-CULL-9 requires, the discipline of never thresholding a bare cosine, and
the practice of leaving an uncertain face honestly unnamed. Personal libraries should also score
better than its film benchmark: cooperative subjects, better lighting, and a closed gallery of dozens
rather than thousands.
---
### D17 — cross-frame face repair · **OPEN**
Raised 2026-09-19: pair FR-CULL-8a's eye state with the burst, and let a closed-eyed face in the
chosen frame be replaced by the same person's open-eyed face from a neighbouring frame. It inverts
FR-CULL-5's fear — the tool rescues the only frame of the moment instead of rejecting it — and it
collides with four things this document says.
1. **§1.3, "not a pixel editor."** Survivable by the FR-DEV-8 precedent: a repair is numbers in
the graph and no pixels are stored. A face repair is a spot whose source is another frame,
aligned by the similarity the face subsystem already fits between two landmark sets and blended
by the membrane heal already shipped. It reopens `spot-removal.md`'s non-goal that the source is
"a patch from the same photograph", which would be revised deliberately, not quietly.
2. **§5.1, single-source `Image`.** The real one. §7 defers panorama, HDR and focus stacking on the
grounds that the schema cannot express an image derived from several sources. This is that case
in a milder form — the output is still frame A, but A's sidecar now names B — and it brings the
rules the tiers already have: B's original must be present to render A (FR-NC-6c: visible and
priced before starting), trashing B must know A depends on it (FR-CAT-15), and `Version::merge`
has never seen a cross-reference. Reopening §5.1 for this reopens it for the three deferred
rows at once, which is the argument for doing it once and properly rather than for this alone.
3. **Tone.** The source patch goes through A's chain, not B's — B demosaiced and run through A's
parameters to the head of the detail chain, then sampled. A second small pipeline at proxy
resolution; a second full demosaic at export. The heal hides lighting drift between frames; it
does not hide a turned head, and no blend does.
4. **Never automatic.** `spot-removal.md`'s own rule: a false positive silently alters a
photograph, which is the failure this application must not have. The tool may propose — same
person, closed here, open two frames on, small alignment residual — and applies on a press.
And one thing the document does not say: **provenance.** The audience is RAW-literate (D11). A
composite declares itself — the source frame in the sidecar and the history, and in export
metadata. Nothing here specifies C2PA; that is a decision this one depends on.
Deferred under D12 until FR-CULL-8a exists and spot removal's disc has, in its own words, been
finished and used. The first cut, when it comes, is "clone from a neighbouring frame" as a spot
source; the face-aware proposal is a layer over that.
### D16 — plugin licensing · **OPEN, post-v1**
> Deferred with §3.10 on 2026-09-19. Still to be answered before the format is published as
> stable, which is now a post-v1 event; nothing in v1 waits on it.
D8 puts the application under GPLv3. §3.10 admits third-party plugins in three forms, and the
derivative-work question is answered differently for each — a YAML-and-WGSL declaration is data of
the kind the GPL has never claimed, a WebAssembly component communicating over a defined interface
is arguably at arm's length, and an interpreted Slint component compiled into the application's own
widget tree is not.
This must be answered before an ecosystem exists, not after. Contributors will not adopt a plugin
format whose licence terms are unstated, and a term introduced later cannot be applied to plugins
already written.
Three questions, in order of how much they constrain the design:
1. **May a plugin be non-free?** If yes, the interfaces are a deliberate licence boundary and must
be documented as one. If no, the registry (FR-PLG-10) enforces it and the default registry lists
only GPL-compatible plugins.
2. **Does the answer differ by class?** Declaring class 1 unambiguously data, whatever is decided
for class 3, is defensible and costs nothing.
3. **What does the default registry require?** Licence metadata is a field in the registry index
either way, so the field should exist from the first release regardless of what policy is
attached to it.
D16 does not block FR-PLG-2, which concerns operations shipped in this repository under D8 already.
It blocks publishing a third-party plugin format as stable.
## 7. Out of scope for v1
Deferred deliberately. Listed so their absence reads as a decision rather than an oversight, with a
note where deferring now constrains the design later.
| Deferred | Note |
|---|---|
| Tethered shooting | — |
| Panorama and HDR merge | **Keep the schema open** — these produce images derived from multiple sources, which §5.1's single-source `Image` cannot express. |
| Focus stacking | Same provenance consideration. |
| Cross-frame face repair ("best take") | A face from a neighbouring frame of the same burst, aligned by its landmarks and blended by FR-DEV-8's heal. Mechanically a spot whose source is another photograph; **the same multi-source schema question as the two rows above, arriving early** — D17. Deferred rather than refused, with its non-goals fixed now: never automatic, geometry not corrected, source frame declared in sidecar, history and export. |
| Gaze and eye-contact estimation | Every open gaze model read on 2026-09-19 is trained on Gaze360, MPIIGaze or ETH-XGaze, all research-only, and Gaze360's licence restricts *models trained on it* by name — the InsightFace situation again (D13). Head pose from the five landmarks is the proxy (FR-CULL-8a). Revisit when weights with a clean data chain exist; iris offset within the eye crop is the licence-free fallback if the proxy proves too weak. |
| Print layout | — |
| Soft proofing | Parameterise the output colour stage by an arbitrary profile so this becomes a UI addition, not a pipeline change. FR-EXP-3's print-dimension mode already half-commits to print workflows. |
| ~~Face recognition~~ | **Undeferred 2026-08-09**, in the narrower form specified in §3.9.1 (FR-CULL-8 … FR-CULL-12): people *grouping and search*, no automated selection. Reclassified as culling rather than AI — it is the same mechanical-grouping category as FR-CULL-5, not the taste operation AI masking is. Gated on spike S14 and decision D13. |
| ~~AI subject masking~~ | **Undeferred 2026-09-19** — it had been built. FR-DEV-3i states what exists: subject and category masks from local models, stored as identity, editable like a drawn mask. The shape this row asked for when it was written — point at a thing, get an editable mask — is the shape that shipped. |
| AI upscaling | Deferred. Lower priority than denoise, which has no manual fallback. |
| Video | — |
| Plugin API | **Post-v1**, decided 2026-09-19. Specified in full in §3.10 as the design of record; every clause there is marked `(post-v1)` and sits outside the coverage denominator. D16 (plugin licensing) defers with it. FR-DEV-3c's compile-time declarations are not plugins and remain in v1. |
| Multi-user / server-side catalog | — |
| Watermarking | Cheap *if* the export pipeline anticipates a compositing stage; expensive to retrofit otherwise. Consider reserving the stage now. |
| Multiple catalogs, catalog merge | Interacts with NFR-OPS-3: preferences must not live in the catalog. |
| Geotagging and map view | — |
| Web gallery, slideshow | — |
| DNG conversion | — |
---
## 8. Verification approach
| Requirement class | How verified |
|---|---|
| Performance (§4.1) | Automated benchmark suite against a synthetic 50k catalog, run per-commit on the reference desktop and periodically on the named reference Android devices. **A regression beyond stated tolerance fails the build.** |
| Rendering correctness | Golden-image tests: fixed source + fixed edit graph → comparison **within R1's stated tolerance**, not checksum equality. Run on both platforms and both Android GPU vendors. |
| Colour accuracy (FR-DEV-3e) | ColorChecker exposures per launch body, asserting ΔE2000 within threshold against reference values. |
| RAW decode coverage | Corpus of sample files per supported camera body; decode-and-checksum regression suite. |
| Robustness (NFR-SEC-1) | Continuous fuzzing of the decode path. |
| Sync correctness | Simulated two-device scenarios including conflict, offline edit, and interrupted transfer. |
| Memory bounds | Long-running soak test scrolling a large catalog, asserting bounded RSS and GPU memory. |
| Schema migration (NFR-R5) | Fixture catalogs from every prior released version, migrated forward and verified. |
| Device loss (NFR-R7) | Induced `VK_ERROR_DEVICE_LOST` mid-render; assert recovery with no lost edits. |
| Process death (FR-PLAT-AND-3) | Kill the Android process mid-edit; assert session and viewport restore with at most the last uncommitted change lost. |
| Source relocation (FR-CAT-9) | Move, rename, and disconnect sources; assert offline marking, reconnection by hash, and no catalog row loss. |
| Cancellation (NFR-ARCH-3) | Assert every long-running operation observes cancellation within the stated bound, including in-flight GPU work. |
| Layer separation (ARCH §6.5a) | CI dependency-tree assertion: no `core/*` crate may transitively depend on a UI toolkit. |
| Operation self-description (FR-DEV-3c) | A test operation added to the registry appears in a generated panel with no frontend change. |
| Adaptive layout (§3.5) | Snapshot tests at each breakpoint, and a resize test asserting no state loss across a layout-class transition. |
| Touch targets (FR-UI-3) | Automated check that interactive elements meet the 44pt minimum in touch modality. |
| Export sizing (FR-EXP-3) | Per-mode dimension assertions, including aspect preservation, fill-crop centring, and the upscale-disabled fallback. |
| Identity calibration (FR-CULL-9) | Reliability diagram over a hand-labelled corpus: stated probability against observed match rate, asserted within tolerance across the range — not a single accuracy figure, which would hide exactly the miscalibration this tests for. Plus a static assertion that no comparison thresholds a raw similarity. |
| Face data confinement (NFR-SEC-5) | Assert that a generated diagnostics bundle contains no embedding or face crop, and that with sync disabled no face data appears in any outbound request. Verified by inspecting what the code *can* emit, since the requirement is the absence of a path. |
---
## 9. Validation spikes
Small experiments that de-risk the highest-uncertainty assumptions before substantial build work.
Ordered by risk. With D1 settled these validate the chosen stack rather than choosing between
stacks.
### Tier 1 — before any substantial build work
| # | Spike | Answers | Relates to |
|---|---|---|---|
| **S1** | **Slint + wgpu zero-copy on Linux:** a compute shader writes a texture, `create_texture_from_hal` imports it, Slint composites UI over it. Drag a slider for 10 minutes watching for tearing, leaks, and sync bugs | Whether ARCH §6.1 holds in the chosen stack | D1, ARCH §6.1 |
| **S2** | **Slint + wgpu on Android**, on two devices from **different GPU vendors** (Adreno and Mali) | Whether the Android GPU path holds across vendor divergence | D1 |
| **S10** | **Android SAF at scale:** enumerate a 10k-file document tree and perform random-access range reads over a document fd. Measure against NFR-P1 and NFR-P3 | Whether ARCH §6.9's forced storage model meets the stated Android performance targets | ARCH §6.9, FR-PLAT-AND-1 |
| **S11** | **Play Console permissions dry-run:** submit an actual declaration for this app category before committing to the storage design | Whether Google approves anything beyond SAF | ARCH §6.9 |
| **S9** | **Golden-image comparison** of one edit graph rendered on desktop and Android; **calibrate the achievable tolerance** | What R1's tolerance threshold should actually be | R1 |
### Tier 2 — before the corresponding subsystem is built
| # | Spike | Answers | Relates to |
|---|---|---|---|
| **S3** | **reqwest HTTPS PROPFIND on a real Android device**, including the `rustls-platform-verifier` Kotlin init | The largest known Rust-on-Android networking risk | D7 |
| **S4** | **Range-extract an embedded JPEG** from CR3/NEF/ARW over WebDAV; measure bytes transferred | Whether remote browsing on mobile data is viable | ARCH §6.7, FR-NC-3 |
| **S5** | **ETag pruning against a 10k-file library:** confirm one-request no-op sync, and correct propagation on a single deep-file change | Whether FR-NC-4 scales as designed | ARCH §6.6 |
| **S6** | **Tiled GPU pipeline on a mid-range Android device**, with an image larger than available GPU memory | Whether ARCH §6.2 holds on constrained hardware | NFR-RES-2 |
| **S7** | **rawler decode coverage** across the FR-RAW-1 launch set, on real files from each body | Whether the LibRaw fallback is needed at launch or later | D2, FR-RAW-1 |
| **S8** | **Chunked upload v2** round-trip of a 100MB RAW, including resume after process kill | FR-NC-7 correctness | FR-NC-7 |
| **S12** | **GPU device loss recovery:** induce `VK_ERROR_DEVICE_LOST` mid-render, verify recreation from the edit graph with no lost edits | Whether ARCH §6.10 and NFR-R7 hold | ARCH §6.10 |
| **S13** | **Slint accessibility on Android:** verify TalkBack exposure of names, roles, and values | Whether NFR-A11Y-2 is achievable in the chosen toolkit | NFR-A11Y-2 |
| **S14** | **Face pipeline in Rust, on a real personal library:** run a detector plus an embedder over ~2,000 images through a Rust ONNX runtime, at proxy resolution, on the reference desktop. Measure per-image latency, cluster purity against hand-labelled truth, and fit the FR-CULL-9 calibration to see whether it converges on a library-sized sample. **Resolve the model licence question before writing any of it** | Whether §3.9.1 is buildable without breaking the pure-Rust dependency policy, and whether the accuracy is worth the subsystem | D13, FR-CULL-8, FR-CULL-9 |
### Why this order
**S1, S2, and S10 are the three that can invalidate the architecture.** S1 and S2 test ARCH §6.1 — the
constraint the whole design is built around, and the one darktable's documentation identifies as
their biggest bottleneck. S10 tests whether the Android storage model forced by ARCH §6.9 can actually
meet the performance targets; it is the highest-uncertainty assumption in the document because until
this revision it was unstated.
**S11 costs almost nothing and de-risks S10 definitively.** Confirming what Google will approve for
this app category before designing around it is far cheaper than discovering it at submission.
**S9 moved to Tier 1** because it does not merely test R1 — it *calibrates* it. R1's tolerance
threshold cannot be fixed sensibly without knowing the real cross-vendor deviation, and the §9
golden-image strategy depends on that number.
Test S1 on Mesa/AMD, Intel, and NVIDIA proprietary drivers, under both X11 and Wayland. FD-based
external memory has well-documented driver divergence, and the reference machine's discrete GPU will
not surface Intel or Mesa-specific issues on its own. The same reasoning is why S2 requires two
Android GPU vendors.
---
## 10. Glossary
- **Edit graph** — the ordered set of parameterised operations defining how an image is rendered.
- **Proxy** — a reduced-resolution render used for display.
- **Tile** — a sub-rectangle of an image processed independently.
- **Demosaic** — reconstructing full RGB from a colour-filter-array sensor capture.
- **CFA** — colour filter array (Bayer, X-Trans).
- **Sidecar** — a small file alongside the source holding edit metadata.
- **Pixel pipeline** — the ordered chain of processing stages from sensor data to output.