Files
DarkRoom/docs/requirements.md
T
dtourolle acd694c2cf No phone, so stop designing for one
Targets are a 12-inch tablet and a desktop (D15). That removes most of the
navigation question rather than answering it.

`EXPANDED_MIN_WIDTH` is 820 logical pixels and a 12-inch tablet is ~1024
across in portrait, so both orientations of both targets are the expanded
class. The compact class now fires only when a desktop window is dragged
narrow — graceful degradation, not a second interface. The bottom tool strip
and the one-tool-at-a-time sheet were solving a phone, and there is no phone.

What survives is input, not size, and the architecture had already decided
it: `WidgetDemand::precise_pointing` exists for a television remote, and its
own documentation says touch is fine because hit regions grow to the
modality. Touch changes hit regions, not layout. The rules that fall out are
worth stating because they are easy to violate by accident — no hover-only
affordance and no modifier key may be the sole route to anything, since a
tablet has neither. Local masking already lost its shift-click extend for
exactly this reason.

The guaranteed-wide viewport also pays for a better answer to the extent
problem than hiding things. The complaint was never that the column is long;
it is that the histogram scrolls away from the sliders it reports on.
Collapsing shortens the scroll, pinning removes the problem, and ~260px of
fixed height is affordable on a viewport that is never under 820 wide.
2026-08-22 09:55:09 +02:00

1396 lines
83 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# DarkRoom — Requirements Specification
**Status:** Draft v0.1 · 2026-08-08
**Owner:** Duncan Tourolle
A cross-platform, non-destructive RAW photo editor for Linux desktop and Android, in the
Lightroom idiom: a catalog of many thousands of images, a develop module with GPU-accelerated
adjustments, and export to standard 8-bit (or higher) deliverables.
---
## 1. Scope and intent
### 1.1 What this is
DarkRoom is a photo *library* and *develop* application. It manages large collections of camera
RAW files, renders them to screen with GPU acceleration, applies non-destructive edits stored as
metadata, and exports finished images.
### 1.2 Primary platforms
| Platform | Priority | Notes |
|---|---|---|
| Linux desktop | Primary | X11 and Wayland. Development and reference platform. |
| Android | Primary | Tablet-first; phone supported. Shares the image core. |
Other platforms (Windows, macOS, iOS) are explicitly out of scope for v1, but the architecture
must not foreclose them. In practice this means the GPU abstraction and the image core must not
hard-code Vulkan-only or Linux-only assumptions at their public interfaces.
### 1.3 What this is not
- Not a DAM with server-side multi-user collaboration
- Not a pixel editor (no layers, no brushes in v1 beyond local-adjustment masks)
- Not a printing/soft-proofing suite in v1
### 1.4 Architecture
The technology stack, crate layout, and design are specified separately in
[architecture.md](architecture.md). This document states *what* the software must do;
the architecture document states *how*.
Requirements here reference architectural constraints as `ARCH §n` where the constraint
materially shapes what is testable.
---
## 2. Core requirements (from the brief)
These are the user's stated requirements, restated as testable criteria.
| ID | Requirement | Acceptance criterion |
|---|---|---|
| **R1** | Cross-platform: Linux + Android | Same image core compiles and runs on both. For the same input and edit graph, output is **perceptually identical within a bounded tolerance** — see below. |
| **R2** | Efficient display of huge RAW libraries | A 50,000-image catalog scrolls at 60fps sustained, with a stated prefetch margin and cache-hit rate sufficient that no cell renders as a placeholder at a scroll velocity of *(figure TBD)* rows/second. Catalog opens in under 2s. |
| **R3** | Make a RAW beautiful at 8-bit output | Full non-destructive develop chain at high internal precision, with camera input profiles (FR-DEV-3e) and a colour-managed path to 8/16-bit export. |
| **R4** | HW acceleration and parallelism | All per-pixel work runs on GPU compute. CPU work (decode, I/O) is parallelised across cores. The UI executor never blocks on image work (NFR-ARCH-1). |
| **R5** | Work on downscaled proxies for display | Display pipeline operates at viewport resolution, not source resolution. Only visible tiles are computed; panning recomputes only newly exposed tiles. |
| **R6** | Nextcloud integration | Browse, download, and upload images and edit metadata against a Nextcloud instance, offline-capable. |
**On R1's tolerance.** An earlier draft required output to be *bit-identical* across platforms.
That is not achievable and the requirement has been corrected. Floating-point compute results
differ between GPU vendors: transcendental function implementations vary, drivers apply different
optimisations, and f16 rounding diverges. A checksum comparison across Adreno and Mesa would fail
for reasons that have nothing to do with correctness.
R1 is therefore stated as a bounded tolerance — a defined maximum per-pixel deviation, expressed in
ΔE2000 for colour or ULPs at the working precision. **The threshold must be fixed before spike S9**,
because S9 both validates R1 and calibrates what the achievable tolerance actually is.
Where genuine bit-identity is required — cache keys, edit-graph hashing (§5.2 invariant 3) — it
applies to *integer* operations on CPU-side state, which are deterministic, never to GPU float
results.
---
## 3. Functional requirements
### 3.1 Catalog and library management
**FR-CAT-1 — Scan.** The app shall scan one or more user-granted library roots for supported image
files, recursively, without blocking the UI. Progress is reported and the scan is cancellable and
resumable. A "root" is a platform-specific grant (a directory on Linux, a persisted document tree
on Android — see FR-PLAT-AND-1), not necessarily a filesystem path.
**FR-CAT-1a — Source addressing.** The catalog and decode layers shall address source data through
an opaque `SourceRef` that resolves to a seekable byte stream, **never through a filesystem path**.
Android's Storage Access Framework provides no usable path (ARCH §6.9), so a path-based API would not be
portable. `SourceRef` carries enough information to re-resolve after an app restart or a permission
re-grant.
**FR-CAT-2 — Catalog store.** All catalog metadata (source references, EXIF, ratings, labels, edit
graphs, sync state) is stored in a local embedded database. The database is the source of truth for
the UI; sources are scanned into it, never queried directly on the UI path.
**FR-CAT-3 — Thumbnail pyramid.** For each image the app maintains a cached multi-resolution
thumbnail set. Initial thumbnails are extracted from the RAW's embedded JPEG preview where
present (fast path, no demosaic). Higher-quality proxies are generated lazily from the full
decode when the image is first opened in develop.
**FR-CAT-4 — Virtualised grid.** The library grid shall render only visible cells plus a small
prefetch margin. Memory use is bounded and independent of catalog size.
**FR-CAT-5 — Metadata.** Read EXIF, camera make/model, lens, capture time, ISO/aperture/shutter,
GPS. Support user-assigned star ratings, colour labels, flags, and keywords.
**FR-CAT-6 — Search and filter.** Filter the catalog by any indexed metadata field, rating,
label, folder, and keyword, with results updating interactively on a 50k catalog.
**FR-CAT-7 — Collections.** User-defined collections that reference images without moving files.
**FR-CAT-8 — Sidecar persistence, independent of sync.** Edit graphs shall be written to per-image
sidecars for **all** catalogued images, whether or not a Nextcloud account exists. Invariant 5.2.4
(catalog rebuildable from sources plus sidecars) otherwise fails for local-only users, leaving every
edit in a single SQLite file with no recovery path.
Where the source location is not writable — read-only mounts, and commonly Android SAF trees —
sidecars are written to an app-managed store keyed by `SourceRef`, and the app shall state which
location is in use.
**FR-CAT-9 — Offline and relocated sources.** An image whose source is unreachable shall be marked
*offline*, never silently removed. Cached previews, metadata, ratings, and edits remain browsable
and editable while offline; edits queue and apply when the source returns.
The app shall support folder-level and image-level reconnection, matching candidates by content
hash and filename, and shall auto-reconnect a volume or tree when it reappears. A source deleted
outside the app shall be distinguished from one merely unreachable before any destructive catalog
action is offered.
This matters more than it appears: external drives, SD cards, and network mounts disappear
routinely, and on Android a tree permission can be revoked or lost on reinstall.
**FR-CAT-10 — Import and ingest.** Copy or move files from a source volume into a destination
structured by a date/metadata template, with rename-on-import, an optional simultaneous
second-destination backup copy, and per-file verification against a checksum. Removable-volume
insertion is detected where the platform permits.
Distinct from FR-CAT-1: scanning catalogues files where they already are; import moves them from a
card into the library. Both are needed.
**FR-CAT-11 — Duplicate detection.** Detect duplicates on import by (capture time + camera serial +
original filename) and by content hash, offering skip or import-as-new. Camera filenames wrap at
`IMG_9999`, so filename alone is insufficient. Existing catalog duplicates are detectable on demand.
**FR-CAT-12 — Versions (virtual copies).** An image may carry multiple named `Version`s, each with
an independent edit graph, without duplicating source data. Versions are creatable, nameable,
deletable, and independently exportable; one is the default.
**FR-CAT-13 — XMP interoperability.** Read and write standard XMP sidecars for ratings, colour
labels, keywords and hierarchical subjects, title, description, copyright, and GPS, using standard
`xmp:`/`dc:`/`lr:` schemas so other tools interoperate.
DarkRoom's edit graph lives in a private namespace and shall neither be interpreted by, nor
corrupt, other tools' XMP. Writing to source-adjacent XMP is off by default (NFR-R4). External
modification of an XMP sidecar shall be detected and a metadata reload offered.
**FR-CAT-14 — Migration import.** Import ratings, labels, keywords, and collections from a
Lightroom `.lrcat` and a darktable `library.db`. Edit graphs are explicitly **not** migrated —
develop parameters do not translate meaningfully between pipelines, and a partial translation is
worse than none. This is the path in for users with existing libraries.
**FR-CAT-15 — Trash and permanent delete.** Deleting an image shall be reversible by default. A
soft delete **moves the file** into a `.darkroom-trash/` folder under the library root and records
in the catalog when it was trashed and the path it came from; restore moves it back to that path.
Permanent delete removes the file first and the catalog row second, and a delete of something
already gone counts as success.
A flag alone would not survive invariant 5.2.4: the catalog is rebuildable from sources, so a
rescan would find every "deleted" file still in the library and re-index it. The folder is the
durable fact and the row is the convenience — which also means the scanner shall exclude the trash
folder, and that a user can recover by hand without DarkRoom. Derived data keyed on the file
(thumbnails, cached previews) is dropped when the image is permanently deleted, not when it is
trashed.
The trash shall be listable newest-first, with the count and total bytes it holds shown before any
destructive action, since that figure is what tells the user whether they meant it.
### 3.2 RAW decoding
**FR-RAW-1 — Format support.** Decode mainstream RAW formats. Minimum launch set: Canon (CR2,
CR3), Nikon (NEF), Sony (ARW), Fujifilm (RAF, including X-Trans), Panasonic (RW2), Olympus (ORF),
Adobe DNG. Additional formats are a coverage goal, not a launch blocker.
**FR-RAW-2 — Decoder abstraction.** RAW decoding sits behind a trait taking a `SourceRef`
(FR-CAT-1a), not a filesystem path, so the same decoder works over a local file, an Android SAF
document, or a byte range fetched from Nextcloud. A second implementation may be added for broader
camera coverage without changing callers (D2).
**FR-RAW-3 — Sensor data handling.** Correctly apply per-camera black/white levels, CFA pattern
identification, and camera-native colour matrices. Demosaic quality shall be selectable, with at
least a fast method for preview and a high-quality method for export (FR-EXP-9 requires export to
use the latter).
**FR-RAW-4 — Robustness.** A malformed or hostile RAW file shall not crash the application or
compromise the process. Decode failures are reported per-file and do not abort a batch.
**FR-RAW-5 — X-Trans as a first-class path.** Per D11, Fujifilm is explicitly targeted:
- **Markesteijn-class demosaic as the default** for X-Trans sensors, not an opt-in advanced setting
- **X-Trans-aware sharpening**, since the non-Bayer CFA responds differently
- In-RAF film simulation tag read and matched (FR-DEV-3f)
This targets the market's best-documented colour grievance. Adobe's X-Trans "worms" artefact is a
decade-old unresolved complaint; darktable and RawTherapee have the better algorithm but poor
defaults; Capture One has the best film-simulation support but drops X-Trans I and II.
*Cost to note:* X-Trans demosaic is documented at **at least 2× the processing cost of Bayer**,
which affects the NFR-P4 and NFR-P7 budgets for Fuji files specifically.
### 3.3 Develop pipeline
**FR-DEV-1 — Non-destructive edit graph.** All edits are stored as parameters in an ordered edit
graph attached to the image. Source files are never modified. Any rendered output is reproducible
from source + graph.
**FR-DEV-2 — Internal precision.** The pipeline operates internally at a minimum of 16-bit float
per channel in a wide-gamut linear working space. Quantisation to the output bit depth happens
once, at the final export or display stage.
**FR-DEV-3 — Adjustment set (v1).**
- White balance (temperature/tint, and picker)
- Exposure, contrast
- Highlights / shadows / whites / blacks recovery
- Tone curve (RGB and per-channel)
- HSL / colour mixer per colour band
- Vibrance and saturation
- Texture / clarity
- Sharpening and noise reduction (luminance and chroma)
- Lens corrections: distortion, chromatic aberration, vignetting
- Crop, straighten, rotate, flip
- Local adjustments: linear gradient, radial gradient, and brush masks
**FR-DEV-3a — Self-describing operations.** Every processing operation shall declare its own
parameters through a descriptor, so that adding an operation requires no changes to frontend code.
An operation declares *what* its parameters are; the frontend decides *how* to present them.
```rust
pub trait Operation: Send + Sync {
/// Static description of this op's parameters. Drives UI generation.
fn descriptor() -> OpDescriptor where Self: Sized;
/// Parameter values → GPU work. No UI types cross this boundary.
fn encode(&self, enc: &mut ComputeEncoder, ctx: &TileContext);
/// Identity for cache invalidation (see §5.2 invariant 3).
fn params_hash(&self) -> u64;
}
pub struct ParamDescriptor {
pub id: ParamId,
pub label: LocalizedString,
pub kind: ParamKind,
pub default: ParamValue,
pub affects: Affects, // Geometry | Colour | Detail — drives invalidation scope
}
pub enum ParamKind {
/// Ordinary numeric parameter. Frontend picks slider / drag-strip / dial by modality.
Scalar { min: f32, max: f32, scale: Scale, unit: Unit, precision: u8 },
Bool,
Enum { variants: Vec<(EnumId, LocalizedString)> },
Colour { has_alpha: bool },
/// Escape hatch: a control that does not reduce to a primitive.
/// The frontend owns the implementation; the op only names the kind
/// and defines the data it exchanges.
Custom { widget: WidgetKind, data: CustomParamSchema },
}
pub enum WidgetKind {
ToneCurve, // per-channel curve editor
ColourWheel, // colour grading wheels
CropOverlay, // on-canvas crop and straighten handles
GradientHandle, // on-canvas linear/radial mask placement
BrushMask, // on-canvas brush strokes
WhiteBalancePick, // eyedropper bound to canvas
}
```
**The pipeline crate shall not depend on the UI toolkit.** Descriptors carry data, never widgets.
This keeps the edit chain testable headless (see §9's golden-image tests, which must link no UI)
and is what allows one operation to render differently on touch and desktop.
**FR-DEV-3b — Frontend presentation mapping.** The frontend maps `ParamKind` to a concrete control
based on input modality and available space. The same descriptor yields different presentations:
| `ParamKind` | Desktop | Touch (tablet) |
|---|---|---|
| `Scalar` | Slider with numeric entry, scroll-wheel fine adjust | Large drag-strip, double-tap to reset, no keyboard entry |
| `Bool` | Checkbox | Switch, minimum 44pt target |
| `Enum` | Dropdown | Segmented control or sheet |
| `Colour` | Swatch opening a picker popover | Swatch opening a full-width sheet |
| `Custom` | Frontend-supplied control for that `WidgetKind` | Same control, touch-tuned hit targets |
**FR-DEV-3c — Operation registry.** Operations register themselves at startup. The develop panel is
generated by walking the registry, so a new operation appears in the UI without any frontend
change. Registration order defines default pipeline order; the ordering itself is data, not code.
**FR-DEV-3d — Invalidation scope.** Each parameter declares what it `affects`, so a change
invalidates only the necessary part of the pipeline. Adjusting exposure shall not re-run lens
correction or re-tile geometry. This is what makes FR-DSP-3's one-frame slider response achievable.
**FR-DEV-3e — Camera input profiles.** The pipeline shall include a camera-profile stage between
demosaic and the working-space conversion.
**v1 scope** (per D11 — good defaults rather than exhaustive colour science):
1. Embedded DNG `ColorMatrix1/2` and `ForwardMatrix1/2` tags
2. A hand-tuned base curve per launch camera body, shipped with the app
3. HaldCLUT import (FR-DEV-3f)
**Deferred but not foreclosed:** full `.dcp` support with `HueSatDeltas`, `ProfileLookTable`, and
dual-illuminant interpolation. The stage shall be structured so these are additions rather than a
pipeline reordering.
Rationale for the reduced scope: a bare 3×3 matrix produces the flat, poor-skin-tone rendering
characteristic of dcraw defaults, which is the documented reason people abandon darktable in the
first hour. A per-body base curve fixes most of that at a fraction of the cost of a full DCP
implementation. The profile database ships **versioned independently of the app binary** so bodies
and curves can be added without a release — and, under D8's GPLv3, contributed by users.
*Acceptance:* for each launch body, the default render is subjectively comparable to the camera's
own JPEG. ΔE2000 validation against ColorChecker references applies once DCP support lands.
**FR-DEV-3f — Look emulation.** Support HaldCLUT import, which inherits the existing free film
simulation ecosystem at near-zero implementation cost, plus reading the in-RAF film simulation tag
to auto-apply a matching render for Fujifilm files.
**FR-DEV-3g — AI denoise.** Learned denoising operating in the raw domain, ideally jointly with
demosaic.
Promoted into v1 scope per D11. The reasoning: unlike AI masking, denoise has **no manual fallback**
— it reaches a quality ceiling no conventional method matches, which is why photographers run a
second application for it. Raw-domain joint demosaic-and-denoise is also markedly easier to build
into a new pipeline than to retrofit, and the same component attacks the X-Trans artefact problem
(FR-RAW-5).
Inference is **local only** — no cloud, no telemetry (NFR-SEC-4). The stage is optional at runtime
and its absence degrades gracefully.
**FR-DEV-3h — Stored orientation is honoured, not edited.** An image shall be shown the way the
photograph was taken, from its EXIF orientation tag (`0x0112`), everywhere it appears: the grid's
thumbnails, the develop canvas, and the read-only preview shown when no decoder can open the file.
The tag shall be applied as a property of **reading the file**, at the same standing as a RAW's
masked-photosite crop (FR-RAW-3) — never as an edit. Concretely:
- Opening a frame the camera stored sideways shall not mark it modified, shall not enable the
framing reset, and shall write nothing to its sidecar.
- "Reset framing" shall return the image to *upright*, not to the sensor's scan order.
- A sidecar shall never carry the orientation. Edits are shared between devices and bodies
(FR-NC-9); one camera's sensor scan must not be applied to another's file.
- A user's own quarter turns compose *on top* of it, so one press of the rotate button moves the
image by 90° whatever the file's baseline.
A file carrying no tag, or a value outside 1..=8, is displayed as stored. Guessing would turn a
missing tag into a visibly wrong image, and most files have no tag.
*Acceptance:* a portrait frame from a phone or a body held sideways appears upright in the grid and
in develop with no user action, and its sidecar is byte-identical to that of the same frame shot in
landscape.
**FR-DEV-4 — Ordered, GPU-resident execution.** The pipeline executes as a sequence of GPU
compute stages. Intermediate results remain in GPU memory between stages. **Processed pixels
shall reach the display without a CPU round-trip.** *(This is a hard architectural constraint —
see ARCH §6.1.)*
**FR-DEV-5 — Edit history.** Per-image undo/redo of edit operations, persisted with the catalog
so history survives a restart. Named snapshots of an edit state.
**FR-DEV-6 — Presets.** Save, apply, and manage named presets covering a subset of the edit
graph. Copy/paste settings between images. Batch-apply to a selection.
**FR-DEV-7 — Before/after.** Compare current edit state against the unedited original or against
a chosen history state.
**FR-DEV-8 — Spot removal.** Non-destructive clone and heal spots stored as parameters in the edit
graph (target, radius, feather, source offset, opacity, mode), with automatic source placement and
manual override, plus a visualise-spots mode.
Sensor dust is unavoidable with interchangeable lenses, and dust spots are the most common reason a
photographer leaves a RAW editor for a pixel editor mid-workflow. This is not the layer-based
pixel editing excluded by §1.3 — it is a standard parameterised develop operation, and the brush
infrastructure required by FR-DEV-3's masks already covers most of the cost.
### 3.4 Display and interaction
**FR-DSP-1 — Proxy-resolution rendering.** The develop view renders at the resolution actually
required by the viewport, not the source resolution. A 60MP image displayed in a 2000px viewport
processes approximately 2000px of data, not 60MP.
**FR-DSP-2 — Tiled computation.** The visible region is divided into tiles. Only tiles
intersecting the viewport are computed. Panning computes only newly exposed tiles; already-valid
tiles are reused.
**FR-DSP-3 — Interactive latency.** Moving a slider updates the visible region within one frame
budget at proxy resolution. When a full-resolution result is needed it is computed
asynchronously, and the proxy result remains on screen until it is ready.
**FR-DSP-4 — Progressive refinement.** During rapid interaction the app may render at reduced
quality or resolution, refining to full quality when interaction settles. Refinement is visually
smooth, not a jarring swap.
**FR-DSP-5 — Zoom and pan.** Fit, 1:1, and arbitrary zoom levels. At 1:1 and above, the pipeline
operates on the visible crop at full source resolution.
**FR-DSP-6 — Colour management.** The display path is colour-managed via the output device
profile. Where the platform and display support it, output at greater than 8 bits per channel
and in a wide gamut. On Android this means using the wide-gamut display path where available.
**FR-DSP-7 — Image evaluation.** Provide a live histogram (luminance and per-channel, in the output
colour space), highlight and shadow clipping indicators, and a pixel colour readout under the
cursor or touch point.
**These derive from a GPU-side reduction into a small buffer. Per-frame CPU readback of image data
is prohibited** — it would violate ARCH §6.1 on every frame, which is precisely the bottleneck darktable
documents. Histogram computation shall not extend the FR-DSP-3 frame budget.
Without this a photographer cannot see what highlight recovery is actually doing, which makes the
FR-DEV-3 adjustment set substantially less usable.
**FR-DSP-8 — Per-display colour and scaling.** The display transform is selected per the display
currently showing the canvas, and updates when the window moves between displays. Fractional and
mixed DPI scaling are handled without resampling artefacts in the canvas.
The profile-acquisition mechanism is stated per display server, with a defined fallback where
Wayland provides no profile. On a multi-monitor desktop with differing profiles, showing wrong
colours on the second display is a correctness defect, not a polish item.
### 3.5 Adaptive interface
DarkRoom ships **one adaptive interface**, not separate touch and desktop applications. A single
Slint codebase reflows by available space and input modality, guaranteeing feature parity by
construction. Phones are out of scope for v1 (§1.3); the layout family spans tablet and desktop.
**FR-UI-1 — Layout breakpoints.** The interface adapts across at least two layout classes:
| Class | Typical | Develop layout |
|---|---|---|
| Compact | Tablet portrait, narrow desktop window | Canvas full-width; one collapsible panel at a time; filmstrip on demand |
| Expanded | Tablet landscape, desktop | Filmstrip, canvas, and adjustment panel simultaneously |
Layout class is a function of window size, not device type — a narrow window on desktop uses the
compact layout, and the transition is continuous rather than a mode switch.
**FR-UI-2 — Input modality.** The interface detects and adapts to the active input method, which
is independent of layout class: a tablet may have a keyboard and pointer attached, and a desktop
may have a touchscreen. Modality affects control sizing and affordances (FR-DEV-3b), not layout.
Switching input mid-session shall be handled without restart.
**FR-UI-3 — Touch targets.** Interactive controls present a minimum 44pt hit target when touch is
the active modality. Hit targets may exceed the drawn control bounds.
**FR-UI-4 — Gestures.** The canvas supports pinch-zoom, two-finger pan, and double-tap to toggle
fit/1:1. Gestures are additive: every gesture-driven action has a non-gesture equivalent, so no
functionality is touch-only.
**FR-UI-5 — Pointer and keyboard.** Where a pointer is present: hover states, right-click context
menus, and scroll-wheel adjustment on numeric controls. Keyboard shortcuts cover navigation,
rating, and common adjustments. Neither is required for any operation to be reachable.
**FR-UI-6 — Shared component library.** Touch and desktop presentations are variants of shared
components, not parallel implementations. A new operation (FR-DEV-3c) becomes usable on both
without frontend work.
**FR-UI-7 — On-canvas controls.** Custom controls that operate on the canvas — crop handles,
gradient placement, brush strokes (`WidgetKind` in FR-DEV-3a) — size their interaction regions to
the active modality while rendering identically. A crop handle drawn at 8px may carry a 44pt touch
region.
### 3.6 Export
**FR-EXP-1 — Formats.** Export to JPEG, PNG, TIFF (8 and 16-bit), and AVIF or JPEG XL. Quality,
chroma subsampling, and bit depth are configurable.
**FR-EXP-2 — Colour space.** Export in a selectable output colour space (sRGB, Display P3,
Adobe RGB, ProPhoto), with the correct ICC profile embedded.
**FR-EXP-3 — Output sizing.** Export size shall be specifiable by any of the following modes:
| Mode | Behaviour |
|---|---|
| Original | Full source resolution after crop |
| Long edge | Specified pixels on the longer dimension; aspect preserved |
| Short edge | Specified pixels on the shorter dimension; aspect preserved |
| Width × Height (fit) | Scaled to fit within the box; aspect preserved; result may be smaller in one dimension |
| Width × Height (fill) | Scaled to cover the box and centre-cropped to exactly those dimensions |
| Percentage | Scaled by a factor of the source |
| Megapixels | Scaled so the result approximates a target pixel count |
| Print dimensions | Physical size (mm or inches) at a specified DPI, resolved to pixels |
Additional constraints:
- **Upscaling** is permitted but shall be off by default, with an explicit opt-in. Where disabled,
a request larger than the source exports at source size rather than failing.
- **DPI metadata** is settable independently of pixel dimensions, for print workflows.
- **File-size ceiling** (JPEG/AVIF/JPEG XL): optionally target a maximum output size in KB/MB, with
the encoder iterating quality to meet it. Useful for upload limits.
- Sizing operates on the **cropped** result, so the crop rectangle defines the aspect ratio unless
a fill mode overrides it.
**FR-EXP-4 — Resampling and output sharpening.** Resizing uses a quality resampler (Lanczos or
equivalent) operating on linear-light data at pipeline precision, before quantisation to the output
bit depth. Output sharpening is selectable (none / screen / matte paper / glossy paper) and its
strength scales with the resize factor, since downscaling softens.
**FR-EXP-5 — Export presets.** Named presets capture format, quality, colour space, sizing mode,
sharpening, metadata policy, and destination. A preset is applicable to a single image or a batch.
Multiple presets may be applied in one operation, producing several outputs per image — e.g. a
full-size TIFF alongside a 2048px sRGB JPEG.
**FR-EXP-6 — Naming and destination.** Output filenames are generated from a template supporting at
minimum: original filename, sequence number, capture date, export dimensions, and preset name.
Collision policy (overwrite / skip / auto-increment) is configurable. Destinations include a local
path and a Nextcloud remote path (FR-NC-7).
**FR-EXP-7 — Batch export.** Export a selection with one or more presets, running in the background
with progress and cancellation. Uses all available cores and the GPU. A failure on one image is
reported and does not abort the batch.
**FR-EXP-8 — Metadata on export.** Configurable EXIF/IPTC/XMP retention, including an option to
strip GPS and personal metadata. Copyright and contact fields are settable per-preset.
**FR-EXP-9 — Full-quality path.** Export always uses the full-resolution, highest-quality pipeline
regardless of what the display was showing — including the high-quality demosaic (FR-RAW-3), never
the fast preview method.
### 3.7 Nextcloud integration
Mechanics below are verified against Nextcloud 34 documentation and server/desktop-client source.
Three findings shape this section and are recorded as constraints in ARCH §6.6–ARCH §6.8.
**FR-NC-1 — Account setup.** Connect via **Login Flow v2**: `POST /index.php/login/v2` returns a
browser URL and a poll token; the app opens the URL in the *system browser* (never an embedded
webview) and polls `POST /login/v2/poll` until it returns an app password. The token is valid 20
minutes and the success response is returned exactly once. The app never sees the user's primary
password.
The `User-Agent` sent during the flow names the resulting app password in the user's security
settings, so it shall identify the device (e.g. `DarkRoom (Linux desktop)`), allowing per-device
revocation. Logout shall call `DELETE /ocs/v2.php/core/apppassword` to revoke cleanly.
Manual app passwords are supported as a fallback for unusual server configurations.
**FR-NC-2 — Credential storage.** Linux: Secret Service via libsecret. KDE exposes the same
interface through `ksecretd` since KF5.97, so one code path covers GNOME and KDE. Where no
secrets daemon is running, the app shall enter an explicit degraded mode rather than silently
storing credentials in plaintext.
Android: Keystore-backed encryption. Note `EncryptedSharedPreferences` is deprecated; the current
approach is DataStore for persistence with Tink for encryption and Keystore for key protection.
Keys must not require user authentication, or background sync will fail.
**FR-NC-3 — Remote browsing without full download.** The app shall display a remote library's
thumbnails without transferring full RAW files. Two mechanisms, selected per-account by capability
probe at setup:
1. *Server previews* — `GET /core/preview?fileId=…` where available. **`forceIcon=false` is
mandatory**: the default returns a generic mimetype icon when the server cannot render the
file, which would otherwise be cached as though it were a thumbnail. The `nc:has-preview`
property in PROPFIND indicates per-file availability.
2. *Range-based embedded preview extraction* — the required fallback (see ARCH §6.7). Fetch the first
64–256KB via HTTP `Range`, parse the container to locate the embedded JPEG preview, then fetch
exactly that byte range. Typical cost 1–3MB versus 25–100MB for the full file.
**FR-NC-4 — Change detection.** Sync shall use recursive ETag pruning, matching the official
desktop client's discovery algorithm:
1. `PROPFIND Depth: 0` on the sync root requesting `getetag`. If unchanged from the stored value,
nothing anywhere in the library has changed — sync completes in one request.
2. Where changed, `PROPFIND Depth: 1` and recurse only into child folders whose ETag differs.
Cost is proportional to the changed subtree, not to library size. `Depth: infinity` shall not be
relied upon (frequently disabled or prohibitively expensive). ETags shall be normalised for quote
inconsistencies before comparison, or spurious full rescans result.
**FR-NC-5 — Identity.** The catalog shall key remote files on Nextcloud's `oc:fileid`, which is
stable across renames and moves, so that a server-side move is detected as a move rather than as a
delete plus a re-download of a 100MB file.
**FR-NC-6 — Selective download.** Downloads are on-demand and resumable, running in the
background. RAW files are never bulk-synced by default. On Android, transfers respect
unmetered-network and charging constraints.
Three storage tiers:
| Tier | Content | Policy |
|---|---|---|
| Metadata | Catalog rows, edit-graph sidecars | Always synced; kilobytes; sync even on metered connections |
| Previews | Embedded JPEGs or server previews | LRU-evicted, size-capped; what the grid browses |
| Full RAW | Source files | Explicit pin or on-demand open only |
**FR-NC-6a — Cache rules.** The user shall be able to pin a *set* of images at a chosen tier, with
the set defined by a rule that the app re-evaluates as the catalog changes. Selectors shall include
at minimum:
- **Collection** — "this trip is available offline"
- **Folder**, optionally recursive
- **Date range**, absolute or **rolling** ("the last 90 days", which moves with the clock)
- **Rating**, **colour label**, **flag**, or **keyword** — "every 5-star image, always"
- **Boolean composition** of the above
A rolling window shall stay current without user intervention. Where rules disagree about an image,
the most generous tier wins.
*Rationale:* selective sync is only usable if the selection can be expressed as intent rather than
enumerated by hand. "Keep this shoot and everything from the last three months" is a sentence a
photographer will say; selecting four thousand files individually is not.
**FR-NC-6b — Lazy eviction.** An image that stops matching a rule is **not** deleted immediately; it
becomes the first candidate for eviction when the cache cap (NFR-RES-4) or platform memory pressure
(FR-PLAT-AND-5) actually requires space. Eviction order is unpinned originals by last use, then
proxies, then thumbnails. **Metadata and sidecars are never evicted** — they are authoritative
(ARCH §6.12) and small.
**FR-NC-6c — Availability is visible.** Every image shall carry a visible availability state:
*Original*, *Preview*, *Metadata only*, or *Offline*.
- An operation requiring absent data shall say so, with the transfer size, *before* starting
- Export from a preview-only image is **refused**, not silently degraded
- A pinned set reports its true byte cost before the user commits
*Rationale:* Lightroom Classic syncs 2560px proxies while displaying the original's filename,
extension, and size, so users do not know what they actually have. Sync failures of legibility are
more damaging than failures of transport.
**FR-NC-7 — Upload.** Files above 5MB use **chunked upload v2** against
`/remote.php/dav/uploads/<userid>/`: `MKCOL` to create the upload folder, `PUT` each chunk, then
`MOVE` the `.file` pseudo-entry to the destination. Chunks are 5MB–5GB and named 1–10000.
`OC-Total-Length` shall always be sent so quota is checked up front rather than at assembly time.
Upload folders expire after 24h of inactivity; the app shall persist upload state and either
resume or `DELETE` stranded uploads on startup.
Small files (sidecars) use **bulk upload** via `POST /remote.php/dav/bulk` with a
`multipart/related` body, allowing hundreds of edit-graph sidecars in a single request.
**FR-NC-8 — Edit metadata sync.** Edit graphs sync bidirectionally as sidecars, one per image,
named deterministically from `oc:fileid`. Each sidecar carries a monotonic revision counter, a
per-device UUID, and a last-edit timestamp.
A sidecar holds a **keyed set of Versions** (FR-CAT-12), not a single edit graph — one image may
carry several virtual copies, and a single-graph format could not represent them. Version identity
is part of the sidecar schema, so conflict merge (FR-NC-9) operates per-version.
**FR-NC-9 — Conflict handling.** Sidecar updates use `If-Match` with the known ETag for optimistic
concurrency (`If-None-Match: *` for creates). On `412 Precondition Failed` the app shall fetch the
remote sidecar and **merge at the edit-graph node level** — disjoint edits (e.g. a crop on one
device, an exposure change on the other) both survive; genuinely conflicting nodes resolve by
timestamp — then retry with the new ETag under a bounded retry count.
The app shall **not** replicate the desktop client's `(conflicted copy)` file behaviour. Sidecars
are structured data of a few KB; a read-merge-rewrite cycle is cheap and preserves user intent.
Only genuinely ambiguous merges surface to the UI. Source RAW files are write-once and shall never
generate a conflict.
**FR-NC-10 — Offline-first.** The app is fully functional offline against cached content. Local
sidecar writes are atomic (temp file plus rename) and committed locally *before* any network
round-trip, so editing never blocks on connectivity. Sync resumes automatically when connectivity
returns.
**FR-NC-11 — Initial catalog build.** For first sync of a large remote library, the app may use
WebDAV `SEARCH` (RFC 5323) against `/remote.php/dav/` filtered by mimetype and paginated via
`d:limit`/`d:nresults`, in preference to walking thousands of folders with PROPFIND.
**FR-NC-12 — Backend independence.** Sync shall be implemented against a backend interface, with
Nextcloud as the only implementation in v1. No protocol detail specific to Nextcloud may appear
outside its connector.
Backends **declare capabilities** rather than conforming to a lowest common denominator, because
the property that makes Nextcloud sync fast — directory ETags propagating up the tree, so an
unchanged root proves an unchanged library — is not a general guarantee. An interface built to the
common subset would force full enumeration on every sync (ARCH §8.1).
Where a capability is absent the app shall **degrade visibly, not silently**:
- Sync strategy in use is reportable to the user, so a slow backend is visibly slow
- Without byte-range reads, remote browsing cannot extract embedded previews; the app shall refuse
full downloads for browsing on a metered connection and explain why
- Without conditional writes, sidecar conflict detection falls back to revision comparison, which
narrows but does not close the race; this is surfaced as a reduced-safety mode
### 3.8 Platform integration
#### Android
**FR-PLAT-AND-1 — Storage access.** Library access is obtained **exclusively via the Storage Access
Framework**: the user grants one or more document trees through `ACTION_OPEN_DOCUMENT_TREE`,
persisted with `takePersistableUriPermission` and enumerated via `DocumentsContract`.
The app shall **not** request `MANAGE_EXTERNAL_STORAGE` and shall **not** depend on
`READ_MEDIA_IMAGES` for RAW discovery. See ARCH §6.9 for why neither is viable.
**FR-PLAT-AND-2 — Permission loss.** Loss of a previously granted tree permission — revocation,
reinstall, removed SD card — shall be detected and surfaced, marking affected images offline per
FR-CAT-9 rather than deleting catalog rows.
**FR-PLAT-AND-3 — Process lifecycle.** An Android process may be killed at any moment. Edit state
shall be durable such that process death loses at most the last uncommitted parameter change. On
resume the app restores the develop session, its image, and its viewport.
**FR-PLAT-AND-4 — Background execution.** Long-running sync and export use the platform's managed
background execution with the constraints in FR-NC-6, and a foreground service with notification
for user-initiated exports. Behaviour under Doze and battery-saver is specified and tested.
**FR-PLAT-AND-5 — Memory pressure.** The app shall respond to `onTrimMemory` /
`ComponentCallbacks2` by evicting caches per NFR-RES-1, in a stated eviction order (GPU tiles
first, then proxies, then thumbnails).
**FR-PLAT-AND-6 — Intents.** Register as a receiver for image view and share intents, and provide
share-out of exported results via `FileProvider`.
#### Linux
**FR-PLAT-LIN-1 — Desktop integration.** Follow the XDG Base Directory specification for config,
data, cache, and state. Register MIME associations for supported RAW types and ship a `.desktop`
entry.
**FR-PLAT-LIN-2 — Display server.** Support X11 and Wayland. Where Wayland's colour-management
protocol is unavailable, FR-DSP-8's stated fallback applies.
**FR-PLAT-LIN-3 — Sandboxed distribution.** Where distributed as Flatpak, filesystem access uses
portals and credential storage uses the Secret Service portal, both verified to satisfy FR-NC-2 and
FR-CAT-1 within the sandbox.
### 3.9 Culling
Per D11 this is the product's primary differentiator, not an incidental capability.
**The opportunity, stated plainly.** Photo Mechanic is fast because it displays the camera's
embedded JPEG — no demosaic, no database, no import step. FastRawViewer is *truthful* because it
shows a genuine raw histogram, raw-derived clipping, and focus peaking. **No shipping tool combines
both.** The culler/editor split exists only because Lightroom's culling is slow — it is a workaround
photographers tolerate, not a workflow they want. A tool that is genuinely Photo Mechanic-fast
eliminates the handoff rather than improving it.
**FR-CULL-1 — Instant display.** Displaying the next image shall not wait on demosaic, catalog
import, or full decode. The embedded JPEG preview is shown immediately; higher-quality renders
replace it progressively.
*Acceptance:* next-image display within **50 ms** of the input event, sustained across a
3,000-image folder, on both platforms. This is the single most important performance figure in the
document — Lightroom's ~2 s stall is the entire reason a competing product category exists.
**FR-CULL-2 — Preview ladder.** Previews resolve through tiers, each falling through to the next:
1. Embedded JPEG preview from the RAW container (instant)
2. Cached proxy from a previous visit
3. Background full decode, promoted when ready
Some cameras embed previews below sensor resolution, and some embed none. The app shall **detect
this per camera model** and pre-emptively background-render where the embedded preview is
insufficient, rather than showing the user a soft image and letting them discover it at zoom.
**FR-CULL-3 — Raw-truth overlays.** Culling decisions are made against raw data, not the embedded
JPEG:
- **Raw histogram** — computed from sensor data, not the preview. The embedded JPEG's histogram
misrepresents available highlight headroom.
- **Raw-derived clipping indicators** — a JPEG's clipping warnings systematically lie about what is
recoverable in the raw.
- **Focus peaking** — overlays in-focus regions on the preview, so focus is verifiable **without
zooming to 100%**. This removes the largest single source of culling latency from the critical
path.
Per D11, zoom-to-100% remains available for certainty; peaking makes it optional rather than
mandatory.
**FR-CULL-4 — Culling mode.** A dedicated full-screen mode with:
- **Auto-advance** on judgement — users independently reinvent this in darktable, Lightroom desktop,
and Lightroom mobile, which is strong evidence it should be the default rather than an option
- **One-key reject**, plus the full rating, flag, and colour-label axes
- **Filter to unjudged**, so a session resumes where it stopped
- Keyboard-driven on desktop; single-thumb reachable on tablet
**FR-CULL-5 — Burst and near-duplicate grouping.** Group frames by capture-time proximity and image
similarity, allowing a burst to collapse to one representative and be judged as a unit.
This is the one automated capability photographers consistently praise, precisely because it is a
mechanical grouping problem rather than a taste judgement. Automated *selection* is distrusted — the
documented failure is rejecting the only frame of an important moment because someone blinked.
**FR-CULL-6 — Comparison.** Side-by-side and survey comparison of a selection, with synchronised
zoom and pan, for choosing among near-identical frames.
**FR-CULL-7 — Tablet culling.** Culling shall be fully usable on tablet, as the validated
multi-device workflow (§3.5, D11). Requires only ratings and small proxies to sync, not full
originals — a substantially smaller sync problem than develop parity.
*Design note:* pinch-zoom accidentally triggering ratings is a documented defect in Lightroom
mobile. Gesture and rating targets must not overlap.
### 3.9.1 People
Face recognition was deferred in §7 through the 2026-08-08 calibration. It is undeferred here in a
narrower form, and the narrowing is the point.
**What changed.** The deferral treated "face recognition" as an AI feature adjacent to subject
masking. It is not the same kind of thing. Masking is a *taste* operation applied to one image;
grouping photographs by who is in them is a **mechanical grouping problem over the whole library**,
which is the category FR-CULL-5 already commits to and already justifies: grouping is the automated
capability photographers consistently praise, because it organises without deciding. Every argument
FR-CULL-5 makes for burst grouping applies unchanged to people grouping. Answering "where are the
frames with the bride in them" across a 4,000-image wedding is a culling operation, and culling is
the differentiator.
**What is deliberately not in scope**, because it is the failure FR-CULL-5 names: no automated
*selection*. Nothing here rejects a frame, ranks a face, scores a smile, or detects a blink. The
feature produces a **filter**, never a judgement. The user's rating axes remain the only thing that
rejects a photograph.
**FR-CULL-8 — Face detection.** The app shall detect faces in library images as a background job,
producing per-face a bounding box, five-point landmarks, a detector confidence, and a 512-dimension
embedding.
Detection runs against the **thumbnail or proxy tier, never a full decode** (FR-CULL-2's ladder).
This is what makes indexing affordable: a library that has been browsed has already paid for its
proxies, so face indexing adds no RAW decodes that were not already happening. Where no proxy
exists, the job requests one at background priority rather than decoding inline.
Detection is a job in the FR-CAT-3 queue and inherits its properties without exception: coalesced
per image, interruptible, resumable across process death (FR-PLAT-AND-3), and strictly preempted by
visible work (NFR-ARCH-2). A library indexes while idle or it does not index; it never competes with
the grid.
*Acceptance:* indexing a 10k-image library completes without the grid dropping below NFR-P9's
interaction target at any point, and survives being killed and restarted with no repeated work
beyond the in-flight image.
**FR-CULL-9 — Calibrated identity.** Face similarity shall be expressed as a **calibrated
probability that two faces are the same person**, not as a raw embedding distance. Every threshold
in the subsystem — clustering, suggestion, auto-confirmation — shall be stated in that probability
space, and no code path may threshold a bare cosine similarity.
This is a hard requirement rather than an implementation detail because the failure mode is
invisible. A raw cosine means something different for every model, every population, and every face
size; a threshold tuned on one library silently misbehaves on another, and an uncalibrated
similarity still *looks* like a plausible number all the way to the user interface. A displayed
confidence that does not mean what it says is worse than no confidence, because it is trusted.
The calibration shall be fitted per library from that library's own faces, and shall report whether
it is valid. Where it is not — too few examples to fit — the app shall say the confidence is
unavailable rather than present an untuned default as though it were measured.
*Acceptance:* on a labelled corpus, the stated probability is within a documented tolerance of the
observed match rate across the probability range (a reliability-diagram check, not a single
accuracy figure).
**FR-CULL-10 — Clustering and naming.** Detected faces shall be clustered into unnamed groups. The
user names a group, and that name applies to its members. A person is thereafter a first-class
catalog entity with a stable UUID, independent of any name given to them.
The user shall be able to **merge** two groups that are the same person, **split** a group that is
not, **remove** a face from a person, and **rename** a person, at any time and without re-indexing.
Splitting must be as easy as merging: clustering will over-merge on siblings, on parents and
children, and on the same person a decade apart, and a tool that can only merge makes its own errors
permanent.
Confirmation is explicit. A face is either **suggested** (the system's inference) or **confirmed**
(the user's judgement), and the two are never conflated in storage or in display. Suggestions may be
recomputed freely; confirmations are user data and are never overwritten by a later inference pass.
**FR-CULL-11 — People as a selector term.** A person shall be a term in the §5 selector language,
composable with every other term.
This is the requirement that pays for the subsystem, and it is nearly free once FR-CULL-10 exists:
because one predicate language serves the library filter, smart collections, and cache rules, a
person term yields all three at once — filter the grid to a person, save "every photo of Anna rated
three or higher" as a smart collection, and pin "every photo of my children" to stay local on the
tablet. The last is a genuinely new capability, not a restatement of the first two.
Selectors shall distinguish confirmed from suggested membership, defaulting to confirmed-only, so a
saved collection does not silently change membership when a later indexing pass revises a guess.
**FR-CULL-12 — Names are user data; embeddings are not.** A confirmed person name is a user
judgement of the same class as a rating or a keyword, and shall be written to the sidecar
(FR-CAT-8), so it survives catalog deletion and travels with the photograph.
Embeddings, detections, cluster assignments, and unconfirmed suggestions are **derived data**. They
live in the catalog only, are rebuildable by re-indexing, and are never written to a sidecar. This
follows ARCH §6.12 exactly: the expensive-but-reproducible artefact stays in the disposable index,
and only the irreplaceable human judgement enters the trust path.
The person UUID is what a cross-device merge keys on, in the same way collections merge (FR-CAT-7).
Two devices that independently name the same cluster produce two people; merging them is the
ordinary FR-CULL-10 merge, not a special case.
---
## 4. Non-functional requirements
### 4.1 Performance targets
These are targets to design against and measure, on the reference desktop
(AMD Threadripper 2920X, 24 threads, discrete GPU) and a mid-range Android device.
| ID | Metric | Desktop target | Android target |
|---|---|---|---|
| **NFR-P1** | Catalog open (50k images) | < 2 s | < 4 s |
| **NFR-P2** | Grid scroll | Sustained 60 fps | Sustained 60 fps |
| **NFR-P3** | Thumbnail generation throughput | ≥ 100 img/s (embedded preview path) | ≥ 25 img/s |
| **NFR-P4** | Open image in develop (to first proxy on screen) | < 400 ms | < 1000 ms |
| **NFR-P5** | Slider adjustment → visible update | < 16 ms (one frame) | < 33 ms |
| **NFR-P6** | Pan/zoom responsiveness | No dropped frames at 60 fps | No dropped frames |
| **NFR-P7** | Full-resolution export (24MP, full chain) | < 2 s | < 8 s |
| **NFR-P8** | Idle memory (50k catalog, nothing open) | < 500 MB | < 250 MB |
| **NFR-P9** | **UI-executor** blocking (not "any operation") | Never > 16 ms | Never > 16 ms |
| **NFR-P13** | **Next image in culling mode** (FR-CULL-1) | **< 50 ms** | **< 50 ms** |
| **NFR-P14** | Focus peaking overlay ready | < 100 ms after preview | < 150 ms |
| **NFR-P15** | Drawn mask stroke → visible (ARCH §6.11) | < 16 ms, no cursor lag | < 16 ms |
| **NFR-P10** | Touch gesture → visual response | < 16 ms | < 16 ms |
| **NFR-P11** | Layout class transition (window resize) | No dropped frames, no state loss | n/a |
| **NFR-P12** | Warm-start shader pipeline setup (cached) | < 100 ms | < 100 ms |
Every target above requires a stated measurement method, workload, and pass threshold before it is
testable. NFR-P8 in particular must state whether it measures RSS inclusive or exclusive of GPU
allocations, and whether it holds after SQLite's page cache warms on a 50k catalog.
**Performance regressions fail the build.** §9's benchmark suite runs per-commit; a regression
beyond a stated tolerance is a build failure, not a notification. Performance work rots otherwise.
### 4.2 Reliability
**NFR-R1** — No operation shall lose user edit data. The catalog database uses write-ahead
logging and survives power loss without corruption.
**NFR-R2** — The catalog is backed up automatically on a schedule and before schema migrations.
**NFR-R3** — A crash in decode or GPU work shall not take down the application where it can be
isolated; the affected image is marked as failed and the app continues.
**NFR-R4** — Source image files are strictly read-only to the application, except where the user
explicitly requests a destructive operation (e.g. delete, or writing XMP sidecars).
**NFR-R5 — Schema versioning.** The catalog carries a monotonic schema version. Migrations are
forward-only, transactional, and idempotent on retry. Each migration ships with a test that migrates
a fixture catalog from **every** prior released version. The app refuses to open a catalog with a
newer schema version rather than corrupting it, and says so.
ARCH §6.6 already anticipates one migration (folder ETags); there will be others, and the machinery must
exist before the first one.
**NFR-R6 — Corruption recovery.** On failing an integrity check at startup, the app offers restore
from the NFR-R2 backup, and failing that, rebuild from sources plus sidecars per invariant 5.2.4.
FR-CAT-8 is what makes that second path real for local-only users.
**NFR-R7 — GPU device loss.** The GPU layer treats device loss as an expected event (ARCH §6.10): detect
it, tear down and recreate the device and all derived resources, and re-drive the current render
from the edit graph. **No user edit is lost.** Recovery is exercised by a test that induces device
loss mid-render.
**NFR-R8 — No suitable GPU.** Where no Vulkan device meeting the NFR-COMPAT-1 baseline is
available, the app starts in a stated degraded mode with defined capability limits rather than
failing to launch.
*This requires reconciling ARCH §6.4 and NFR-RES-2:* ARCH §6.4 says the GPU path is primary rather than an
optimisation, while NFR-RES-2 assumes a CPU fallback on allocation failure. **Decide explicitly**
whether v1 includes a full CPU pipeline, or whether "CPU fallback" means only tile-spill staging
with no independent CPU render path. The latter is recommended; the former is a second full
implementation.
### 4.3 Resource behaviour
**NFR-RES-1 — Bounded memory.** Memory use is bounded and configurable, independent of catalog
size and image count. Caches are evictable under pressure.
**NFR-RES-2 — GPU memory.** The pipeline shall handle images larger than available GPU memory by
tiling. GPU memory headroom is configurable, with a CPU fallback path if allocation fails.
**NFR-RES-3 — Mobile power.** On Android the app shall not render continuously when idle. Battery
and thermal behaviour are first-class concerns; background sync respects metered-connection and
battery-saver settings.
**NFR-RES-4 — Disk cache.** Thumbnail and proxy caches have a configurable size cap with LRU
eviction.
### 4.4 Portability
**NFR-PORT-1** — Platform-specific code is isolated behind interfaces. The image core, catalog,
and edit pipeline contain no platform conditionals.
**NFR-PORT-2** — GPU shaders are authored once and used on both platforms.
**NFR-PORT-3** — Adding a third platform requires implementing the platform interfaces only, not
changes to the core.
### 4.5 Security and privacy
**NFR-SEC-1** — RAW parsing treats input as untrusted. Parser hardening, fuzzing, and where
practical process or memory isolation for decode.
**NFR-SEC-2** — Credentials are never written to the catalog, logs, or plain files. Platform
secure storage only.
**NFR-SEC-3** — All network traffic uses TLS with certificate validation. No option to disable
validation in release builds.
**NFR-SEC-4** — No telemetry without explicit opt-in.
**NFR-SEC-5 — Face data stays on the user's own hardware.** Face embeddings (FR-CULL-8) are handled
under a stricter rule than the rest of the catalog.
This is a personal tool for personal libraries (§1.1, D11) — the people in these photographs are the
user's family and friends. That is the reason for the rule, not a reason to relax it: the data is
sensitive precisely because it is personal, and the user is the only party with any claim on it.
- **Never leave the device by default.** Embeddings, face crops, and cluster assignments shall not be
transmitted, uploaded, or included in any diagnostics bundle (NFR-OPS-1) or crash report
(NFR-OPS-2), under any configuration. The diagnostics path has no opt-in for this; it is excluded
outright.
- **Sync is opt-in and separately consented.** Syncing embeddings to the user's own Nextcloud is
permitted — it is their server and their photographs, and it saves re-indexing a library per device
— but it is off by default, is not implied by enabling photo sync, and the consent states in plain
language what is being uploaded and why. Person *names*, being sidecar data (FR-CULL-12), sync with
the sidecar as ordinary metadata.
- **No third-party inference.** Face detection and embedding run locally. No image, crop, or
embedding is sent to a remote inference service, and the app ships no capability to do so.
- **Deletable, in one action.** The user shall be able to delete all face data — embeddings,
detections, clusters, and people — from a single control, without deleting the catalog or any
photograph, and to disable face indexing entirely so that no such data is produced.
- **Model weights are inspectable.** The models used shall be named and versioned in the about
screen, with their licences, so a user can determine what is running on their photographs.
*Rationale:* the rest of this document treats privacy as a property of the network boundary — TLS,
credentials in secure storage, opt-in telemetry. Face data needs more than a well-defended boundary,
because it is not revocable once it has crossed one, and because it describes people who are not the
user. The prohibition is therefore structural rather than configurable: the code paths that would
upload an embedding to anyone but the user's own server do not exist. A setting can be changed by
accident, or by a future maintainer who has forgotten why it was there; an absent code path cannot.
### 4.6 Execution model
**NFR-ARCH-1 — Named executors.** The app defines distinct executors — UI, GPU submission, decode
pool, I/O pool, network — with stated thread counts and the invariant that **no blocking call
occurs on the UI executor**. This is the mechanism behind R4 and NFR-P9, which currently assert an
outcome with no stated means.
**NFR-ARCH-2 — Scheduler priority.** The tiling scheduler assigns priority classes, with
visible-tile work **strictly preempting** background export and thumbnail work. Without this,
NFR-P5's slider latency fails during a batch export — the common case, not an edge case.
**NFR-ARCH-3 — Cancellation.** Cancellation is cooperative with a bounded worst-case latency
(target: observed within 100 ms), and covers in-flight GPU submissions. Every long-running
operation named in FR-CAT-1, FR-EXP-7, and FR-NC-6 is cancellable under this model.
**NFR-ARCH-4 — Error propagation.** No worker error may panic the process. Errors surface as typed
results attached to the affected image or job, consistent with NFR-R3 and FR-RAW-4.
### 4.7 Operations
**NFR-OPS-1 — Diagnostics.** Structured levelled logging to a rotating, size-capped on-disk log in
the XDG state directory or Android app directory, with **automatic redaction of credentials and
tokens** (required by NFR-SEC-2, which currently forbids credentials in logs that are never
otherwise specified). A one-click diagnostics bundle includes log, schema version, GPU and driver
identification, and app version — with an explicit preview-and-consent step before anything leaves
the device.
**NFR-OPS-2 — Crash reporting.** Local crash capture always; upload only on explicit opt-in
(NFR-SEC-4).
**NFR-OPS-3 — Preferences.** A single versioned preferences store, **separate from the catalog**, so
preferences survive catalog rebuild and multiple catalogs. At least eight requirements refer to
configurable settings with no store defined. Device-specific settings (GPU headroom, cache caps) do
not sync between devices.
**NFR-OPS-4 — Update and first run.** State delivery channels and their update mechanisms. This
matters concretely because D2 pins rawler at an alpha, non-SemVer version whose camera-support fixes
users will need. First-run flow is defined, including platform permission acquisition (FR-PLAT-AND-1
makes first run a permission negotiation on Android, not a welcome screen) and initial root
selection.
### 4.8 Compatibility baseline
**NFR-COMPAT-1 — Supported hardware.** Every §4.1 Android figure is meaningless without this. State:
- Minimum and target Android API level (targetSdk 36 is currently required for Play distribution)
- Minimum Vulkan version and the required feature and limit set — including whether `shaderFloat16`
and 16-bit storage are required, since **FR-DEV-2's f16 pipeline depends on them and their
absence would jeopardise R1**
- Minimum device RAM, and minimum desktop Vulkan/Mesa versions
- The **specific** reference Android device the §4.1 column is measured on, plus a secondary device
from a different GPU vendor
Adreno, Mali, and PowerVR diverge significantly in compute behaviour and in external-memory interop
— exactly what spike S1 tests. The spec already applies this reasoning to desktop drivers; it
applies at least as strongly on Android.
**NFR-COMPAT-2 — Distribution channels.** State the v1 channels (e.g. Flatpak and AppImage on Linux;
Play Store and/or F-Droid on Android). Play distribution is what makes ARCH §6.9's constraints binding —
a sideloaded or F-Droid build could use different permissions, so the channel decision and the
storage design are coupled.
### 4.9 Accessibility and internationalisation
**NFR-A11Y-1 — Localisation.** All user-facing strings, including operation and parameter labels
resolved from `LocalizedString` (FR-DEV-3a), are externalised and translatable without
recompilation. Note the constraint this creates: those labels live in core crates that **cannot
depend on the UI** (ARCH §6.5a), so the localisation mechanism must itself be UI-independent. State the
format, the locale-resolution rule, and whether RTL layout is in v1 scope.
**NFR-A11Y-2 — Accessibility.** Controls expose accessible names, roles, and values to the platform
accessibility layer (AT-SPI on Linux, TalkBack on Android). Platform font scaling is honoured
without clipping. Non-canvas UI meets WCAG AA contrast.
**Slint's accessibility support on Android requires verification** — this may be a toolkit gap, and
it is far cheaper to discover now than after the UI is built.
**NFR-A11Y-3 — Colour-independent status.** No status is conveyed by hue alone. Colour labels,
clipping indicators (FR-DSP-7), and the HSL mixer carry a shape or text affordance. This matters
more in a colour-grading application than in most software.
---
## 5. Data model and architecture
Entity definitions, invariants, and all architectural constraints are specified in
[architecture.md](architecture.md) — §3 (core abstractions), §6 (data architecture), and
§11 (architectural constraints).
Requirements in this document that depend on an architectural guarantee cite it inline. The
constraints most load-bearing for testability are:
| Constraint | Why a requirement depends on it |
|---|---|
| ARCH §6.1 — no CPU round-trip | FR-DSP-3, FR-DSP-7, NFR-P5 are unachievable without it |
| ARCH §6.11 — GPU-rasterised masks | NFR-P15 (no brush lag) |
| ARCH §6.12 — sidecars authoritative | NFR-R6, invariant behind FR-CAT-8 |
| ARCH §6.9 — Android SAF only | FR-CAT-1a, FR-PLAT-AND-1, and the NFR-P1/P3 Android figures |
| ARCH §6.6 — no sync tokens | FR-NC-4 |
| ARCH §6.13 — integer-only bit-identity | R1's tolerance, §9 golden images |
## 6. Decisions
Rationale, evidence, and the eliminated alternatives are recorded in
[architecture.md §12](architecture.md). Outcomes only:
| # | Decision | Outcome |
|---|---|---|
| D1 | Language and UI framework | Rust + Slint, rendering through wgpu |
| D2 | RAW decoder | rawler; LibRaw fallback behind a trait |
| D3 | First milestone | **Provisional — depends on D12** |
| D4 | Nextcloud sync mechanism | ETag pruning, chunked upload v2, Login Flow v2 |
| D5 | Colour management | lcms2 + GPU-side matrix/LUT transforms |
| D6 | Shader authoring | Hand-written WGSL |
| D7 | Network stack | reqwest + quick-xml |
| D8 | Licence | **GPLv3** |
| D9 | Operation UI model | Declarative parameter descriptors |
| D10 | Interface strategy | One adaptive UI, tablet + desktop |
| D11 | Product positioning | Culling-first differentiator; see below |
| D12 | Scope versus pace | **OPEN** |
### D11 — product positioning
Settled by requirements calibration, 2026-08-08.
| Dimension | Decision |
|---|---|
| Audience | RAW-literate photographers, Linux-comfortable. Docs matter; hand-holding does not. |
| Library scale | 10k–50k images |
| Culling | **The core differentiator** (§3.9) |
| Focus checking | Peaking *and* zoom |
| Ingest | Full workflow — template rename, checksum verify, dual-destination |
| Colour defaults | Good, not obsessive — matrices plus per-body base curve |
| Film simulation | Fujifilm explicitly targeted |
| AI | Denoise in v1; masking deferred |
| Local adjustments | Full masking, GPU-rasterised |
| Sync | The reason the project exists |
| Durability | Sidecar-first |
| Licence | GPLv3 |
| Pace | Evenings and weekends, indefinite |
### D12 — scope versus pace · **OPEN**
The calibration selected an ambitious feature set — full tablet editing, full ingest, culling as a
differentiator, complete GPU masking, AI denoise, Fuji-first colour, deep sync, sidecar durability
— against a stated pace of evenings and weekends, indefinitely.
Those are not compatible as stated. This is not an argument against any individual choice; it is
that the two must be reconciled before a build order can be set.
Two specific tensions:
**1. Full tablet editing is the most expensive selection**, chosen against the research finding
that volume photographers do not edit on tablets. It carries SAF storage at unproven scale (S10),
process-death durability, background execution limits, and two GPU vendors to validate. Tablet
*culling* is the validated workflow, is what the differentiator points at, and costs a fraction as
much.
**2. The v1 milestone and the differentiator disagree.** D3's vertical slice proves the
architecture but is useful to nobody. If culling is what no existing tool does well, a culler is
both a smaller build and a usable one — it needs no develop chain.
Resolving D12 sets D3 and [architecture.md §10](architecture.md)'s Phase 2.
### D15 — target devices · **DECIDED 2026-08-22**
**A 12-inch tablet and a desktop. No phone.**
Recorded because it is load-bearing for the interface and invisible in the
code. Every phone-shaped answer — a bottom tool strip, a one-tool-at-a-time
sheet, thumb-reach zones — is designing for hardware this project does not
target, and each would have cost a second layout to keep in step with the
first.
What survives the decision is the *input* difference rather than the size one:
a 12-inch tablet is touched, and NFR/FR-UI-7's position already covers it —
hit regions grow to the modality, layout does not move. The practical rules
that fall out are in `docs/ui-navigation.md` D-N2: no hover-only affordance
and no modifier key may be the sole route to anything, because a tablet has
neither.
`EXPANDED_MIN_WIDTH` is 820 logical pixels and a 12-inch tablet is ~1024
across in portrait, so **both orientations of both targets are the expanded
layout**. The compact class remains as graceful degradation for a narrowed
desktop window, not as a second interface.
---
### D13 — face inference runtime and model licensing · **RUNTIME ANSWERED, LICENSING OPEN**
> **Updated 2026-08-21.** The runtime half of this decision is settled, and by a route the table
> below does not contain. `ort` 2.0's `alternative-backend` feature *disables its linking entirely*
> and lets another engine supply the `OrtApi`; `ort-tract` supplies it from `tract`, which is pure
> Rust. So the third option's operator coverage comes with the first option's dependency profile —
> no C, no NDK problem, no exception to the policy. Measured on a real graph before being relied on:
> YOLO26n-seg loads with zero unsupported operators and runs 640×640 in ~470 ms of CPU
> (docs/segmentation.md §13). §3.9.1's detector and embedder are different graphs and their coverage
> has not been checked, but the *approach* no longer needs a decision.
>
> **The licensing half is untouched.** The InsightFace weights are still non-commercial and still
> unusable here. That remains what S14 has to resolve first.
§3.9.1 needs to run two neural networks locally. That collides with two settled positions, and
neither collision is small enough to leave implicit.
**1. The pure-Rust dependency policy.** Every dependency choice in this project has gone the same
way, for the same stated reason: rustls over aws-lc-rs, bundled SQLite over the system library, a
Rust Lensfun port over liblensfun, zune-jpeg over libjpeg — no C dependency to satisfy under the
Android NDK (D1's whole premise). The obvious way to run ONNX models is the ONNX Runtime C++ library,
which would be the largest exception to that policy in the codebase, and it would land on the
platform the policy exists to protect.
The options, in the order I would try them:
| Option | Cost |
|---|---|
| **wgpu compute**, models hand-ported to WGSL | No new dependency at all — the GPU device and shader infrastructure already exist (ARCH §5). Highest implementation effort, and a ViT is a lot of shader. |
| **`burn`** with the wgpu backend | Pure Rust, uses the existing GPU. Young, and ONNX import maturity needs checking against these two specific graphs. |
| **`ort`** (ONNX Runtime bindings) | Fastest to working code, best operator coverage. Reintroduces the C dependency and the NDK cross-compilation problem the policy avoids. |
The tension is real: the cheapest path is the one that breaks the rule. This is worth an explicit
decision rather than a default, and S14 is what informs it.
**2. Model licensing is a distribution blocker, not a detail.** The obvious pretrained weights are
not redistributable under GPLv3. The InsightFace "buffalo" family — ArcFace and the SCRFD detector,
the standard choices — are **licensed for non-commercial research use only**, which is incompatible
with this project's licence and with Flatpak, F-Droid, and Play distribution (NFR-COMPAT-2). Other
candidate weights need their licences read individually rather than assumed.
Two ways out, both with costs:
- **Find permissively-licensed weights** and ship them in-tree. Clean, offline-first, consistent with
how the Lensfun database ships. Requires that suitable weights exist at acceptable accuracy.
- **Download models on first use**, with the user accepting the upstream licence. Sidesteps
redistribution but adds a network dependency to a feature that is otherwise entirely local, needs a
hosting story, and sits badly with the local-first posture of NFR-SEC-5.
**This must be resolved before implementation, not during it.** Discovering at packaging time that
the feature cannot ship is the expensive failure, and it is entirely avoidable — it is a licence-
reading exercise, not a research question. S14 therefore puts it first.
*Prior art available:* `../scene-actor-extraction` is a working implementation of this pipeline
(SCRFD detect → 5-point align → 512-d embedding → Platt-calibrated similarity), benchmarked at 67.4%
macro-F1 on held-out films. Its C++ does not port — different language, OpenCV and TensorRT
dependencies — but its **design decisions do**, and they are the expensive part: the calibrated
probability space that FR-CULL-9 requires, the discipline of never thresholding a bare cosine, and
the practice of leaving an uncertain face honestly unnamed. Personal libraries should also score
better than its film benchmark: cooperative subjects, better lighting, and a closed gallery of dozens
rather than thousands.
---
## 7. Out of scope for v1
Deferred deliberately. Listed so their absence reads as a decision rather than an oversight, with a
note where deferring now constrains the design later.
| Deferred | Note |
|---|---|
| Tethered shooting | — |
| Panorama and HDR merge | **Keep the schema open** — these produce images derived from multiple sources, which §5.1's single-source `Image` cannot express. |
| Focus stacking | Same provenance consideration. |
| Print layout | — |
| Soft proofing | Parameterise the output colour stage by an arbitrary profile so this becomes a UI addition, not a pipeline change. FR-EXP-3's print-dimension mode already half-commits to print workflows. |
| ~~Face recognition~~ | **Undeferred 2026-08-09**, in the narrower form specified in §3.9.1 (FR-CULL-8 … FR-CULL-12): people *grouping and search*, no automated selection. Reclassified as culling rather than AI — it is the same mechanical-grouping category as FR-CULL-5, not the taste operation AI masking is. Gated on spike S14 and decision D13. |
| AI subject masking | Deferred per D11. Note darktable shipped this in 5.6 (June 2026), so the gap is now visible. **Conditions for deferring safely:** AI denoise ships in v1 (FR-DEV-3g ✓), manual masking is excellent including GPU-rasterised drawn masks (ARCH §6.11 ✓), and the product has a clear differentiator (culling, §3.9 ✓). When it does land, copy darktable's shape — prompt-point segmentation producing an *editable* mask that behaves like a hand-drawn one — not Adobe's opaque version. |
| AI upscaling | Deferred. Lower priority than denoise, which has no manual fallback. |
| Video | — |
| Plugin API | — |
| Multi-user / server-side catalog | — |
| Watermarking | Cheap *if* the export pipeline anticipates a compositing stage; expensive to retrofit otherwise. Consider reserving the stage now. |
| Multiple catalogs, catalog merge | Interacts with NFR-OPS-3: preferences must not live in the catalog. |
| Geotagging and map view | — |
| Web gallery, slideshow | — |
| DNG conversion | — |
---
## 8. Verification approach
| Requirement class | How verified |
|---|---|
| Performance (§4.1) | Automated benchmark suite against a synthetic 50k catalog, run per-commit on the reference desktop and periodically on the named reference Android devices. **A regression beyond stated tolerance fails the build.** |
| Rendering correctness | Golden-image tests: fixed source + fixed edit graph → comparison **within R1's stated tolerance**, not checksum equality. Run on both platforms and both Android GPU vendors. |
| Colour accuracy (FR-DEV-3e) | ColorChecker exposures per launch body, asserting ΔE2000 within threshold against reference values. |
| RAW decode coverage | Corpus of sample files per supported camera body; decode-and-checksum regression suite. |
| Robustness (NFR-SEC-1) | Continuous fuzzing of the decode path. |
| Sync correctness | Simulated two-device scenarios including conflict, offline edit, and interrupted transfer. |
| Memory bounds | Long-running soak test scrolling a large catalog, asserting bounded RSS and GPU memory. |
| Schema migration (NFR-R5) | Fixture catalogs from every prior released version, migrated forward and verified. |
| Device loss (NFR-R7) | Induced `VK_ERROR_DEVICE_LOST` mid-render; assert recovery with no lost edits. |
| Process death (FR-PLAT-AND-3) | Kill the Android process mid-edit; assert session and viewport restore with at most the last uncommitted change lost. |
| Source relocation (FR-CAT-9) | Move, rename, and disconnect sources; assert offline marking, reconnection by hash, and no catalog row loss. |
| Cancellation (NFR-ARCH-3) | Assert every long-running operation observes cancellation within the stated bound, including in-flight GPU work. |
| Layer separation (ARCH §6.5a) | CI dependency-tree assertion: no `core/*` crate may transitively depend on a UI toolkit. |
| Operation self-description (FR-DEV-3c) | A test operation added to the registry appears in a generated panel with no frontend change. |
| Adaptive layout (§3.5) | Snapshot tests at each breakpoint, and a resize test asserting no state loss across a layout-class transition. |
| Touch targets (FR-UI-3) | Automated check that interactive elements meet the 44pt minimum in touch modality. |
| Export sizing (FR-EXP-3) | Per-mode dimension assertions, including aspect preservation, fill-crop centring, and the upscale-disabled fallback. |
| Identity calibration (FR-CULL-9) | Reliability diagram over a hand-labelled corpus: stated probability against observed match rate, asserted within tolerance across the range — not a single accuracy figure, which would hide exactly the miscalibration this tests for. Plus a static assertion that no comparison thresholds a raw similarity. |
| Face data confinement (NFR-SEC-5) | Assert that a generated diagnostics bundle contains no embedding or face crop, and that with sync disabled no face data appears in any outbound request. Verified by inspecting what the code *can* emit, since the requirement is the absence of a path. |
---
## 9. Validation spikes
Small experiments that de-risk the highest-uncertainty assumptions before substantial build work.
Ordered by risk. With D1 settled these validate the chosen stack rather than choosing between
stacks.
### Tier 1 — before any substantial build work
| # | Spike | Answers | Relates to |
|---|---|---|---|
| **S1** | **Slint + wgpu zero-copy on Linux:** a compute shader writes a texture, `create_texture_from_hal` imports it, Slint composites UI over it. Drag a slider for 10 minutes watching for tearing, leaks, and sync bugs | Whether ARCH §6.1 holds in the chosen stack | D1, ARCH §6.1 |
| **S2** | **Slint + wgpu on Android**, on two devices from **different GPU vendors** (Adreno and Mali) | Whether the Android GPU path holds across vendor divergence | D1 |
| **S10** | **Android SAF at scale:** enumerate a 10k-file document tree and perform random-access range reads over a document fd. Measure against NFR-P1 and NFR-P3 | Whether ARCH §6.9's forced storage model meets the stated Android performance targets | ARCH §6.9, FR-PLAT-AND-1 |
| **S11** | **Play Console permissions dry-run:** submit an actual declaration for this app category before committing to the storage design | Whether Google approves anything beyond SAF | ARCH §6.9 |
| **S9** | **Golden-image comparison** of one edit graph rendered on desktop and Android; **calibrate the achievable tolerance** | What R1's tolerance threshold should actually be | R1 |
### Tier 2 — before the corresponding subsystem is built
| # | Spike | Answers | Relates to |
|---|---|---|---|
| **S3** | **reqwest HTTPS PROPFIND on a real Android device**, including the `rustls-platform-verifier` Kotlin init | The largest known Rust-on-Android networking risk | D7 |
| **S4** | **Range-extract an embedded JPEG** from CR3/NEF/ARW over WebDAV; measure bytes transferred | Whether remote browsing on mobile data is viable | ARCH §6.7, FR-NC-3 |
| **S5** | **ETag pruning against a 10k-file library:** confirm one-request no-op sync, and correct propagation on a single deep-file change | Whether FR-NC-4 scales as designed | ARCH §6.6 |
| **S6** | **Tiled GPU pipeline on a mid-range Android device**, with an image larger than available GPU memory | Whether ARCH §6.2 holds on constrained hardware | NFR-RES-2 |
| **S7** | **rawler decode coverage** across the FR-RAW-1 launch set, on real files from each body | Whether the LibRaw fallback is needed at launch or later | D2, FR-RAW-1 |
| **S8** | **Chunked upload v2** round-trip of a 100MB RAW, including resume after process kill | FR-NC-7 correctness | FR-NC-7 |
| **S12** | **GPU device loss recovery:** induce `VK_ERROR_DEVICE_LOST` mid-render, verify recreation from the edit graph with no lost edits | Whether ARCH §6.10 and NFR-R7 hold | ARCH §6.10 |
| **S13** | **Slint accessibility on Android:** verify TalkBack exposure of names, roles, and values | Whether NFR-A11Y-2 is achievable in the chosen toolkit | NFR-A11Y-2 |
| **S14** | **Face pipeline in Rust, on a real personal library:** run a detector plus an embedder over ~2,000 images through a Rust ONNX runtime, at proxy resolution, on the reference desktop. Measure per-image latency, cluster purity against hand-labelled truth, and fit the FR-CULL-9 calibration to see whether it converges on a library-sized sample. **Resolve the model licence question before writing any of it** | Whether §3.9.1 is buildable without breaking the pure-Rust dependency policy, and whether the accuracy is worth the subsystem | D13, FR-CULL-8, FR-CULL-9 |
### Why this order
**S1, S2, and S10 are the three that can invalidate the architecture.** S1 and S2 test ARCH §6.1 — the
constraint the whole design is built around, and the one darktable's documentation identifies as
their biggest bottleneck. S10 tests whether the Android storage model forced by ARCH §6.9 can actually
meet the performance targets; it is the highest-uncertainty assumption in the document because until
this revision it was unstated.
**S11 costs almost nothing and de-risks S10 definitively.** Confirming what Google will approve for
this app category before designing around it is far cheaper than discovering it at submission.
**S9 moved to Tier 1** because it does not merely test R1 — it *calibrates* it. R1's tolerance
threshold cannot be fixed sensibly without knowing the real cross-vendor deviation, and the §9
golden-image strategy depends on that number.
Test S1 on Mesa/AMD, Intel, and NVIDIA proprietary drivers, under both X11 and Wayland. FD-based
external memory has well-documented driver divergence, and the reference machine's discrete GPU will
not surface Intel or Mesa-specific issues on its own. The same reasoning is why S2 requires two
Android GPU vendors.
---
## 10. Glossary
- **Edit graph** — the ordered set of parameterised operations defining how an image is rendered.
- **Proxy** — a reduced-resolution render used for display.
- **Tile** — a sub-rectangle of an image processed independently.
- **Demosaic** — reconstructing full RGB from a colour-filter-array sensor capture.
- **CFA** — colour filter array (Bayer, X-Trans).
- **Sidecar** — a small file alongside the source holding edit metadata.
- **Pixel pipeline** — the ordered chain of processing stages from sensor data to output.