Files
DarkRoom/docs/requirements.md
T
dtourolle 12d320cf33 Record what the spec got wrong about the model that exists
docs/segmentation.md §4 priced arm B as costing a C dependency under the
NDK and treated that as most of the difference between the arms. It is not
a cost that has to be paid: `ort`'s `alternative-backend` disables its
linking entirely and `ort-tract` supplies the API from tract, which is pure
Rust. D13's "largest exception the policy would tolerate" turns out not to
be needed, and the answer generalises to the face pipeline — so D13's
runtime half is now answered and only its licensing half is open.

Three findings contradict §4 outright and are recorded as F4-F6 rather than
quietly designed around. There is no ADE20K-trained YOLO, so the shipped
vocabulary selects subjects and not stuff — "select the sky" comes from the
watershed or from nowhere. It is instance segmentation, so it partitions
nothing and two people come back as two instances. And tract cannot parse a
dynamic-shape export, which fixes the input at 640 square and makes tiling
the only route to more semantic resolution.

Arm C ships, but §8's criteria are not what decided it, and saying so
matters more than claiming the process worked. §8 asked for a two-
interaction margin over arm A on a traced corpus. That comparison was never
run: F4 and F5 changed what the arms are, and a model that recognises
subjects but has no word for sky cannot be a selection tool alone, while a
watershed cannot tell a person from the wall behind them. They stopped
being candidates and became complements.

What is *not* done is written down as plainly: the 24-image corpus is
untraced, so M1-M4 have no numbers and "this feels right" has not become
one. M5 is answered on one device only, and region ids now reach the
sidecar — so a cross-vendor divergence would mean a mask written on the
desktop meaning something else on Android. F3 stands.
2026-08-22 08:39:17 +02:00

82 KiB
Raw Blame History

DarkRoom — Requirements Specification

Status: Draft v0.1 · 2026-08-08 Owner: Duncan Tourolle

A cross-platform, non-destructive RAW photo editor for Linux desktop and Android, in the Lightroom idiom: a catalog of many thousands of images, a develop module with GPU-accelerated adjustments, and export to standard 8-bit (or higher) deliverables.


1. Scope and intent

1.1 What this is

DarkRoom is a photo library and develop application. It manages large collections of camera RAW files, renders them to screen with GPU acceleration, applies non-destructive edits stored as metadata, and exports finished images.

1.2 Primary platforms

Platform Priority Notes
Linux desktop Primary X11 and Wayland. Development and reference platform.
Android Primary Tablet-first; phone supported. Shares the image core.

Other platforms (Windows, macOS, iOS) are explicitly out of scope for v1, but the architecture must not foreclose them. In practice this means the GPU abstraction and the image core must not hard-code Vulkan-only or Linux-only assumptions at their public interfaces.

1.3 What this is not

  • Not a DAM with server-side multi-user collaboration
  • Not a pixel editor (no layers, no brushes in v1 beyond local-adjustment masks)
  • Not a printing/soft-proofing suite in v1

1.4 Architecture

The technology stack, crate layout, and design are specified separately in architecture.md. This document states what the software must do; the architecture document states how.

Requirements here reference architectural constraints as ARCH §n where the constraint materially shapes what is testable.


2. Core requirements (from the brief)

These are the user's stated requirements, restated as testable criteria.

ID Requirement Acceptance criterion
R1 Cross-platform: Linux + Android Same image core compiles and runs on both. For the same input and edit graph, output is perceptually identical within a bounded tolerance — see below.
R2 Efficient display of huge RAW libraries A 50,000-image catalog scrolls at 60fps sustained, with a stated prefetch margin and cache-hit rate sufficient that no cell renders as a placeholder at a scroll velocity of (figure TBD) rows/second. Catalog opens in under 2s.
R3 Make a RAW beautiful at 8-bit output Full non-destructive develop chain at high internal precision, with camera input profiles (FR-DEV-3e) and a colour-managed path to 8/16-bit export.
R4 HW acceleration and parallelism All per-pixel work runs on GPU compute. CPU work (decode, I/O) is parallelised across cores. The UI executor never blocks on image work (NFR-ARCH-1).
R5 Work on downscaled proxies for display Display pipeline operates at viewport resolution, not source resolution. Only visible tiles are computed; panning recomputes only newly exposed tiles.
R6 Nextcloud integration Browse, download, and upload images and edit metadata against a Nextcloud instance, offline-capable.

On R1's tolerance. An earlier draft required output to be bit-identical across platforms. That is not achievable and the requirement has been corrected. Floating-point compute results differ between GPU vendors: transcendental function implementations vary, drivers apply different optimisations, and f16 rounding diverges. A checksum comparison across Adreno and Mesa would fail for reasons that have nothing to do with correctness.

R1 is therefore stated as a bounded tolerance — a defined maximum per-pixel deviation, expressed in ΔE2000 for colour or ULPs at the working precision. The threshold must be fixed before spike S9, because S9 both validates R1 and calibrates what the achievable tolerance actually is.

Where genuine bit-identity is required — cache keys, edit-graph hashing (§5.2 invariant 3) — it applies to integer operations on CPU-side state, which are deterministic, never to GPU float results.


3. Functional requirements

3.1 Catalog and library management

FR-CAT-1 — Scan. The app shall scan one or more user-granted library roots for supported image files, recursively, without blocking the UI. Progress is reported and the scan is cancellable and resumable. A "root" is a platform-specific grant (a directory on Linux, a persisted document tree on Android — see FR-PLAT-AND-1), not necessarily a filesystem path.

FR-CAT-1a — Source addressing. The catalog and decode layers shall address source data through an opaque SourceRef that resolves to a seekable byte stream, never through a filesystem path. Android's Storage Access Framework provides no usable path (ARCH §6.9), so a path-based API would not be portable. SourceRef carries enough information to re-resolve after an app restart or a permission re-grant.

FR-CAT-2 — Catalog store. All catalog metadata (source references, EXIF, ratings, labels, edit graphs, sync state) is stored in a local embedded database. The database is the source of truth for the UI; sources are scanned into it, never queried directly on the UI path.

FR-CAT-3 — Thumbnail pyramid. For each image the app maintains a cached multi-resolution thumbnail set. Initial thumbnails are extracted from the RAW's embedded JPEG preview where present (fast path, no demosaic). Higher-quality proxies are generated lazily from the full decode when the image is first opened in develop.

FR-CAT-4 — Virtualised grid. The library grid shall render only visible cells plus a small prefetch margin. Memory use is bounded and independent of catalog size.

FR-CAT-5 — Metadata. Read EXIF, camera make/model, lens, capture time, ISO/aperture/shutter, GPS. Support user-assigned star ratings, colour labels, flags, and keywords.

FR-CAT-6 — Search and filter. Filter the catalog by any indexed metadata field, rating, label, folder, and keyword, with results updating interactively on a 50k catalog.

FR-CAT-7 — Collections. User-defined collections that reference images without moving files.

FR-CAT-8 — Sidecar persistence, independent of sync. Edit graphs shall be written to per-image sidecars for all catalogued images, whether or not a Nextcloud account exists. Invariant 5.2.4 (catalog rebuildable from sources plus sidecars) otherwise fails for local-only users, leaving every edit in a single SQLite file with no recovery path.

Where the source location is not writable — read-only mounts, and commonly Android SAF trees — sidecars are written to an app-managed store keyed by SourceRef, and the app shall state which location is in use.

FR-CAT-9 — Offline and relocated sources. An image whose source is unreachable shall be marked offline, never silently removed. Cached previews, metadata, ratings, and edits remain browsable and editable while offline; edits queue and apply when the source returns.

The app shall support folder-level and image-level reconnection, matching candidates by content hash and filename, and shall auto-reconnect a volume or tree when it reappears. A source deleted outside the app shall be distinguished from one merely unreachable before any destructive catalog action is offered.

This matters more than it appears: external drives, SD cards, and network mounts disappear routinely, and on Android a tree permission can be revoked or lost on reinstall.

FR-CAT-10 — Import and ingest. Copy or move files from a source volume into a destination structured by a date/metadata template, with rename-on-import, an optional simultaneous second-destination backup copy, and per-file verification against a checksum. Removable-volume insertion is detected where the platform permits.

Distinct from FR-CAT-1: scanning catalogues files where they already are; import moves them from a card into the library. Both are needed.

FR-CAT-11 — Duplicate detection. Detect duplicates on import by (capture time + camera serial + original filename) and by content hash, offering skip or import-as-new. Camera filenames wrap at IMG_9999, so filename alone is insufficient. Existing catalog duplicates are detectable on demand.

FR-CAT-12 — Versions (virtual copies). An image may carry multiple named Versions, each with an independent edit graph, without duplicating source data. Versions are creatable, nameable, deletable, and independently exportable; one is the default.

FR-CAT-13 — XMP interoperability. Read and write standard XMP sidecars for ratings, colour labels, keywords and hierarchical subjects, title, description, copyright, and GPS, using standard xmp:/dc:/lr: schemas so other tools interoperate.

DarkRoom's edit graph lives in a private namespace and shall neither be interpreted by, nor corrupt, other tools' XMP. Writing to source-adjacent XMP is off by default (NFR-R4). External modification of an XMP sidecar shall be detected and a metadata reload offered.

FR-CAT-14 — Migration import. Import ratings, labels, keywords, and collections from a Lightroom .lrcat and a darktable library.db. Edit graphs are explicitly not migrated — develop parameters do not translate meaningfully between pipelines, and a partial translation is worse than none. This is the path in for users with existing libraries.

FR-CAT-15 — Trash and permanent delete. Deleting an image shall be reversible by default. A soft delete moves the file into a .darkroom-trash/ folder under the library root and records in the catalog when it was trashed and the path it came from; restore moves it back to that path. Permanent delete removes the file first and the catalog row second, and a delete of something already gone counts as success.

A flag alone would not survive invariant 5.2.4: the catalog is rebuildable from sources, so a rescan would find every "deleted" file still in the library and re-index it. The folder is the durable fact and the row is the convenience — which also means the scanner shall exclude the trash folder, and that a user can recover by hand without DarkRoom. Derived data keyed on the file (thumbnails, cached previews) is dropped when the image is permanently deleted, not when it is trashed.

The trash shall be listable newest-first, with the count and total bytes it holds shown before any destructive action, since that figure is what tells the user whether they meant it.

3.2 RAW decoding

FR-RAW-1 — Format support. Decode mainstream RAW formats. Minimum launch set: Canon (CR2, CR3), Nikon (NEF), Sony (ARW), Fujifilm (RAF, including X-Trans), Panasonic (RW2), Olympus (ORF), Adobe DNG. Additional formats are a coverage goal, not a launch blocker.

FR-RAW-2 — Decoder abstraction. RAW decoding sits behind a trait taking a SourceRef (FR-CAT-1a), not a filesystem path, so the same decoder works over a local file, an Android SAF document, or a byte range fetched from Nextcloud. A second implementation may be added for broader camera coverage without changing callers (D2).

FR-RAW-3 — Sensor data handling. Correctly apply per-camera black/white levels, CFA pattern identification, and camera-native colour matrices. Demosaic quality shall be selectable, with at least a fast method for preview and a high-quality method for export (FR-EXP-9 requires export to use the latter).

FR-RAW-4 — Robustness. A malformed or hostile RAW file shall not crash the application or compromise the process. Decode failures are reported per-file and do not abort a batch.

FR-RAW-5 — X-Trans as a first-class path. Per D11, Fujifilm is explicitly targeted:

  • Markesteijn-class demosaic as the default for X-Trans sensors, not an opt-in advanced setting
  • X-Trans-aware sharpening, since the non-Bayer CFA responds differently
  • In-RAF film simulation tag read and matched (FR-DEV-3f)

This targets the market's best-documented colour grievance. Adobe's X-Trans "worms" artefact is a decade-old unresolved complaint; darktable and RawTherapee have the better algorithm but poor defaults; Capture One has the best film-simulation support but drops X-Trans I and II.

Cost to note: X-Trans demosaic is documented at at least 2× the processing cost of Bayer, which affects the NFR-P4 and NFR-P7 budgets for Fuji files specifically.

3.3 Develop pipeline

FR-DEV-1 — Non-destructive edit graph. All edits are stored as parameters in an ordered edit graph attached to the image. Source files are never modified. Any rendered output is reproducible from source + graph.

FR-DEV-2 — Internal precision. The pipeline operates internally at a minimum of 16-bit float per channel in a wide-gamut linear working space. Quantisation to the output bit depth happens once, at the final export or display stage.

FR-DEV-3 — Adjustment set (v1).

  • White balance (temperature/tint, and picker)
  • Exposure, contrast
  • Highlights / shadows / whites / blacks recovery
  • Tone curve (RGB and per-channel)
  • HSL / colour mixer per colour band
  • Vibrance and saturation
  • Texture / clarity
  • Sharpening and noise reduction (luminance and chroma)
  • Lens corrections: distortion, chromatic aberration, vignetting
  • Crop, straighten, rotate, flip
  • Local adjustments: linear gradient, radial gradient, and brush masks

FR-DEV-3a — Self-describing operations. Every processing operation shall declare its own parameters through a descriptor, so that adding an operation requires no changes to frontend code. An operation declares what its parameters are; the frontend decides how to present them.

pub trait Operation: Send + Sync {
    /// Static description of this op's parameters. Drives UI generation.
    fn descriptor() -> OpDescriptor where Self: Sized;

    /// Parameter values → GPU work. No UI types cross this boundary.
    fn encode(&self, enc: &mut ComputeEncoder, ctx: &TileContext);

    /// Identity for cache invalidation (see §5.2 invariant 3).
    fn params_hash(&self) -> u64;
}

pub struct ParamDescriptor {
    pub id: ParamId,
    pub label: LocalizedString,
    pub kind: ParamKind,
    pub default: ParamValue,
    pub affects: Affects,      // Geometry | Colour | Detail — drives invalidation scope
}

pub enum ParamKind {
    /// Ordinary numeric parameter. Frontend picks slider / drag-strip / dial by modality.
    Scalar { min: f32, max: f32, scale: Scale, unit: Unit, precision: u8 },
    Bool,
    Enum { variants: Vec<(EnumId, LocalizedString)> },
    Colour { has_alpha: bool },
    /// Escape hatch: a control that does not reduce to a primitive.
    /// The frontend owns the implementation; the op only names the kind
    /// and defines the data it exchanges.
    Custom { widget: WidgetKind, data: CustomParamSchema },
}

pub enum WidgetKind {
    ToneCurve,        // per-channel curve editor
    ColourWheel,      // colour grading wheels
    CropOverlay,      // on-canvas crop and straighten handles
    GradientHandle,   // on-canvas linear/radial mask placement
    BrushMask,        // on-canvas brush strokes
    WhiteBalancePick, // eyedropper bound to canvas
}

The pipeline crate shall not depend on the UI toolkit. Descriptors carry data, never widgets. This keeps the edit chain testable headless (see §9's golden-image tests, which must link no UI) and is what allows one operation to render differently on touch and desktop.

FR-DEV-3b — Frontend presentation mapping. The frontend maps ParamKind to a concrete control based on input modality and available space. The same descriptor yields different presentations:

ParamKind Desktop Touch (tablet)
Scalar Slider with numeric entry, scroll-wheel fine adjust Large drag-strip, double-tap to reset, no keyboard entry
Bool Checkbox Switch, minimum 44pt target
Enum Dropdown Segmented control or sheet
Colour Swatch opening a picker popover Swatch opening a full-width sheet
Custom Frontend-supplied control for that WidgetKind Same control, touch-tuned hit targets

FR-DEV-3c — Operation registry. Operations register themselves at startup. The develop panel is generated by walking the registry, so a new operation appears in the UI without any frontend change. Registration order defines default pipeline order; the ordering itself is data, not code.

FR-DEV-3d — Invalidation scope. Each parameter declares what it affects, so a change invalidates only the necessary part of the pipeline. Adjusting exposure shall not re-run lens correction or re-tile geometry. This is what makes FR-DSP-3's one-frame slider response achievable.

FR-DEV-3e — Camera input profiles. The pipeline shall include a camera-profile stage between demosaic and the working-space conversion.

v1 scope (per D11 — good defaults rather than exhaustive colour science):

  1. Embedded DNG ColorMatrix1/2 and ForwardMatrix1/2 tags
  2. A hand-tuned base curve per launch camera body, shipped with the app
  3. HaldCLUT import (FR-DEV-3f)

Deferred but not foreclosed: full .dcp support with HueSatDeltas, ProfileLookTable, and dual-illuminant interpolation. The stage shall be structured so these are additions rather than a pipeline reordering.

Rationale for the reduced scope: a bare 3×3 matrix produces the flat, poor-skin-tone rendering characteristic of dcraw defaults, which is the documented reason people abandon darktable in the first hour. A per-body base curve fixes most of that at a fraction of the cost of a full DCP implementation. The profile database ships versioned independently of the app binary so bodies and curves can be added without a release — and, under D8's GPLv3, contributed by users.

Acceptance: for each launch body, the default render is subjectively comparable to the camera's own JPEG. ΔE2000 validation against ColorChecker references applies once DCP support lands.

FR-DEV-3f — Look emulation. Support HaldCLUT import, which inherits the existing free film simulation ecosystem at near-zero implementation cost, plus reading the in-RAF film simulation tag to auto-apply a matching render for Fujifilm files.

FR-DEV-3g — AI denoise. Learned denoising operating in the raw domain, ideally jointly with demosaic.

Promoted into v1 scope per D11. The reasoning: unlike AI masking, denoise has no manual fallback — it reaches a quality ceiling no conventional method matches, which is why photographers run a second application for it. Raw-domain joint demosaic-and-denoise is also markedly easier to build into a new pipeline than to retrofit, and the same component attacks the X-Trans artefact problem (FR-RAW-5).

Inference is local only — no cloud, no telemetry (NFR-SEC-4). The stage is optional at runtime and its absence degrades gracefully.

FR-DEV-3h — Stored orientation is honoured, not edited. An image shall be shown the way the photograph was taken, from its EXIF orientation tag (0x0112), everywhere it appears: the grid's thumbnails, the develop canvas, and the read-only preview shown when no decoder can open the file.

The tag shall be applied as a property of reading the file, at the same standing as a RAW's masked-photosite crop (FR-RAW-3) — never as an edit. Concretely:

  • Opening a frame the camera stored sideways shall not mark it modified, shall not enable the framing reset, and shall write nothing to its sidecar.
  • "Reset framing" shall return the image to upright, not to the sensor's scan order.
  • A sidecar shall never carry the orientation. Edits are shared between devices and bodies (FR-NC-9); one camera's sensor scan must not be applied to another's file.
  • A user's own quarter turns compose on top of it, so one press of the rotate button moves the image by 90° whatever the file's baseline.

A file carrying no tag, or a value outside 1..=8, is displayed as stored. Guessing would turn a missing tag into a visibly wrong image, and most files have no tag.

Acceptance: a portrait frame from a phone or a body held sideways appears upright in the grid and in develop with no user action, and its sidecar is byte-identical to that of the same frame shot in landscape.

FR-DEV-4 — Ordered, GPU-resident execution. The pipeline executes as a sequence of GPU compute stages. Intermediate results remain in GPU memory between stages. Processed pixels shall reach the display without a CPU round-trip. (This is a hard architectural constraint — see ARCH §6.1.)

FR-DEV-5 — Edit history. Per-image undo/redo of edit operations, persisted with the catalog so history survives a restart. Named snapshots of an edit state.

FR-DEV-6 — Presets. Save, apply, and manage named presets covering a subset of the edit graph. Copy/paste settings between images. Batch-apply to a selection.

FR-DEV-7 — Before/after. Compare current edit state against the unedited original or against a chosen history state.

FR-DEV-8 — Spot removal. Non-destructive clone and heal spots stored as parameters in the edit graph (target, radius, feather, source offset, opacity, mode), with automatic source placement and manual override, plus a visualise-spots mode.

Sensor dust is unavoidable with interchangeable lenses, and dust spots are the most common reason a photographer leaves a RAW editor for a pixel editor mid-workflow. This is not the layer-based pixel editing excluded by §1.3 — it is a standard parameterised develop operation, and the brush infrastructure required by FR-DEV-3's masks already covers most of the cost.

3.4 Display and interaction

FR-DSP-1 — Proxy-resolution rendering. The develop view renders at the resolution actually required by the viewport, not the source resolution. A 60MP image displayed in a 2000px viewport processes approximately 2000px of data, not 60MP.

FR-DSP-2 — Tiled computation. The visible region is divided into tiles. Only tiles intersecting the viewport are computed. Panning computes only newly exposed tiles; already-valid tiles are reused.

FR-DSP-3 — Interactive latency. Moving a slider updates the visible region within one frame budget at proxy resolution. When a full-resolution result is needed it is computed asynchronously, and the proxy result remains on screen until it is ready.

FR-DSP-4 — Progressive refinement. During rapid interaction the app may render at reduced quality or resolution, refining to full quality when interaction settles. Refinement is visually smooth, not a jarring swap.

FR-DSP-5 — Zoom and pan. Fit, 1:1, and arbitrary zoom levels. At 1:1 and above, the pipeline operates on the visible crop at full source resolution.

FR-DSP-6 — Colour management. The display path is colour-managed via the output device profile. Where the platform and display support it, output at greater than 8 bits per channel and in a wide gamut. On Android this means using the wide-gamut display path where available.

FR-DSP-7 — Image evaluation. Provide a live histogram (luminance and per-channel, in the output colour space), highlight and shadow clipping indicators, and a pixel colour readout under the cursor or touch point.

These derive from a GPU-side reduction into a small buffer. Per-frame CPU readback of image data is prohibited — it would violate ARCH §6.1 on every frame, which is precisely the bottleneck darktable documents. Histogram computation shall not extend the FR-DSP-3 frame budget.

Without this a photographer cannot see what highlight recovery is actually doing, which makes the FR-DEV-3 adjustment set substantially less usable.

FR-DSP-8 — Per-display colour and scaling. The display transform is selected per the display currently showing the canvas, and updates when the window moves between displays. Fractional and mixed DPI scaling are handled without resampling artefacts in the canvas.

The profile-acquisition mechanism is stated per display server, with a defined fallback where Wayland provides no profile. On a multi-monitor desktop with differing profiles, showing wrong colours on the second display is a correctness defect, not a polish item.

3.5 Adaptive interface

DarkRoom ships one adaptive interface, not separate touch and desktop applications. A single Slint codebase reflows by available space and input modality, guaranteeing feature parity by construction. Phones are out of scope for v1 (§1.3); the layout family spans tablet and desktop.

FR-UI-1 — Layout breakpoints. The interface adapts across at least two layout classes:

Class Typical Develop layout
Compact Tablet portrait, narrow desktop window Canvas full-width; one collapsible panel at a time; filmstrip on demand
Expanded Tablet landscape, desktop Filmstrip, canvas, and adjustment panel simultaneously

Layout class is a function of window size, not device type — a narrow window on desktop uses the compact layout, and the transition is continuous rather than a mode switch.

FR-UI-2 — Input modality. The interface detects and adapts to the active input method, which is independent of layout class: a tablet may have a keyboard and pointer attached, and a desktop may have a touchscreen. Modality affects control sizing and affordances (FR-DEV-3b), not layout. Switching input mid-session shall be handled without restart.

FR-UI-3 — Touch targets. Interactive controls present a minimum 44pt hit target when touch is the active modality. Hit targets may exceed the drawn control bounds.

FR-UI-4 — Gestures. The canvas supports pinch-zoom, two-finger pan, and double-tap to toggle fit/1:1. Gestures are additive: every gesture-driven action has a non-gesture equivalent, so no functionality is touch-only.

FR-UI-5 — Pointer and keyboard. Where a pointer is present: hover states, right-click context menus, and scroll-wheel adjustment on numeric controls. Keyboard shortcuts cover navigation, rating, and common adjustments. Neither is required for any operation to be reachable.

FR-UI-6 — Shared component library. Touch and desktop presentations are variants of shared components, not parallel implementations. A new operation (FR-DEV-3c) becomes usable on both without frontend work.

FR-UI-7 — On-canvas controls. Custom controls that operate on the canvas — crop handles, gradient placement, brush strokes (WidgetKind in FR-DEV-3a) — size their interaction regions to the active modality while rendering identically. A crop handle drawn at 8px may carry a 44pt touch region.

3.6 Export

FR-EXP-1 — Formats. Export to JPEG, PNG, TIFF (8 and 16-bit), and AVIF or JPEG XL. Quality, chroma subsampling, and bit depth are configurable.

FR-EXP-2 — Colour space. Export in a selectable output colour space (sRGB, Display P3, Adobe RGB, ProPhoto), with the correct ICC profile embedded.

FR-EXP-3 — Output sizing. Export size shall be specifiable by any of the following modes:

Mode Behaviour
Original Full source resolution after crop
Long edge Specified pixels on the longer dimension; aspect preserved
Short edge Specified pixels on the shorter dimension; aspect preserved
Width × Height (fit) Scaled to fit within the box; aspect preserved; result may be smaller in one dimension
Width × Height (fill) Scaled to cover the box and centre-cropped to exactly those dimensions
Percentage Scaled by a factor of the source
Megapixels Scaled so the result approximates a target pixel count
Print dimensions Physical size (mm or inches) at a specified DPI, resolved to pixels

Additional constraints:

  • Upscaling is permitted but shall be off by default, with an explicit opt-in. Where disabled, a request larger than the source exports at source size rather than failing.
  • DPI metadata is settable independently of pixel dimensions, for print workflows.
  • File-size ceiling (JPEG/AVIF/JPEG XL): optionally target a maximum output size in KB/MB, with the encoder iterating quality to meet it. Useful for upload limits.
  • Sizing operates on the cropped result, so the crop rectangle defines the aspect ratio unless a fill mode overrides it.

FR-EXP-4 — Resampling and output sharpening. Resizing uses a quality resampler (Lanczos or equivalent) operating on linear-light data at pipeline precision, before quantisation to the output bit depth. Output sharpening is selectable (none / screen / matte paper / glossy paper) and its strength scales with the resize factor, since downscaling softens.

FR-EXP-5 — Export presets. Named presets capture format, quality, colour space, sizing mode, sharpening, metadata policy, and destination. A preset is applicable to a single image or a batch. Multiple presets may be applied in one operation, producing several outputs per image — e.g. a full-size TIFF alongside a 2048px sRGB JPEG.

FR-EXP-6 — Naming and destination. Output filenames are generated from a template supporting at minimum: original filename, sequence number, capture date, export dimensions, and preset name. Collision policy (overwrite / skip / auto-increment) is configurable. Destinations include a local path and a Nextcloud remote path (FR-NC-7).

FR-EXP-7 — Batch export. Export a selection with one or more presets, running in the background with progress and cancellation. Uses all available cores and the GPU. A failure on one image is reported and does not abort the batch.

FR-EXP-8 — Metadata on export. Configurable EXIF/IPTC/XMP retention, including an option to strip GPS and personal metadata. Copyright and contact fields are settable per-preset.

FR-EXP-9 — Full-quality path. Export always uses the full-resolution, highest-quality pipeline regardless of what the display was showing — including the high-quality demosaic (FR-RAW-3), never the fast preview method.

3.7 Nextcloud integration

Mechanics below are verified against Nextcloud 34 documentation and server/desktop-client source. Three findings shape this section and are recorded as constraints in ARCH §6.6–ARCH §6.8.

FR-NC-1 — Account setup. Connect via Login Flow v2: POST /index.php/login/v2 returns a browser URL and a poll token; the app opens the URL in the system browser (never an embedded webview) and polls POST /login/v2/poll until it returns an app password. The token is valid 20 minutes and the success response is returned exactly once. The app never sees the user's primary password.

The User-Agent sent during the flow names the resulting app password in the user's security settings, so it shall identify the device (e.g. DarkRoom (Linux desktop)), allowing per-device revocation. Logout shall call DELETE /ocs/v2.php/core/apppassword to revoke cleanly.

Manual app passwords are supported as a fallback for unusual server configurations.

FR-NC-2 — Credential storage. Linux: Secret Service via libsecret. KDE exposes the same interface through ksecretd since KF5.97, so one code path covers GNOME and KDE. Where no secrets daemon is running, the app shall enter an explicit degraded mode rather than silently storing credentials in plaintext.

Android: Keystore-backed encryption. Note EncryptedSharedPreferences is deprecated; the current approach is DataStore for persistence with Tink for encryption and Keystore for key protection. Keys must not require user authentication, or background sync will fail.

FR-NC-3 — Remote browsing without full download. The app shall display a remote library's thumbnails without transferring full RAW files. Two mechanisms, selected per-account by capability probe at setup:

  1. Server previews — GET /core/preview?fileId=… where available. forceIcon=false is mandatory: the default returns a generic mimetype icon when the server cannot render the file, which would otherwise be cached as though it were a thumbnail. The nc:has-preview property in PROPFIND indicates per-file availability.
  2. Range-based embedded preview extraction — the required fallback (see ARCH §6.7). Fetch the first 64–256KB via HTTP Range, parse the container to locate the embedded JPEG preview, then fetch exactly that byte range. Typical cost 1–3MB versus 25–100MB for the full file.

FR-NC-4 — Change detection. Sync shall use recursive ETag pruning, matching the official desktop client's discovery algorithm:

  1. PROPFIND Depth: 0 on the sync root requesting getetag. If unchanged from the stored value, nothing anywhere in the library has changed — sync completes in one request.
  2. Where changed, PROPFIND Depth: 1 and recurse only into child folders whose ETag differs.

Cost is proportional to the changed subtree, not to library size. Depth: infinity shall not be relied upon (frequently disabled or prohibitively expensive). ETags shall be normalised for quote inconsistencies before comparison, or spurious full rescans result.

FR-NC-5 — Identity. The catalog shall key remote files on Nextcloud's oc:fileid, which is stable across renames and moves, so that a server-side move is detected as a move rather than as a delete plus a re-download of a 100MB file.

FR-NC-6 — Selective download. Downloads are on-demand and resumable, running in the background. RAW files are never bulk-synced by default. On Android, transfers respect unmetered-network and charging constraints.

Three storage tiers:

Tier Content Policy
Metadata Catalog rows, edit-graph sidecars Always synced; kilobytes; sync even on metered connections
Previews Embedded JPEGs or server previews LRU-evicted, size-capped; what the grid browses
Full RAW Source files Explicit pin or on-demand open only

FR-NC-6a — Cache rules. The user shall be able to pin a set of images at a chosen tier, with the set defined by a rule that the app re-evaluates as the catalog changes. Selectors shall include at minimum:

  • Collection — "this trip is available offline"
  • Folder, optionally recursive
  • Date range, absolute or rolling ("the last 90 days", which moves with the clock)
  • Rating, colour label, flag, or keyword — "every 5-star image, always"
  • Boolean composition of the above

A rolling window shall stay current without user intervention. Where rules disagree about an image, the most generous tier wins.

Rationale: selective sync is only usable if the selection can be expressed as intent rather than enumerated by hand. "Keep this shoot and everything from the last three months" is a sentence a photographer will say; selecting four thousand files individually is not.

FR-NC-6b — Lazy eviction. An image that stops matching a rule is not deleted immediately; it becomes the first candidate for eviction when the cache cap (NFR-RES-4) or platform memory pressure (FR-PLAT-AND-5) actually requires space. Eviction order is unpinned originals by last use, then proxies, then thumbnails. Metadata and sidecars are never evicted — they are authoritative (ARCH §6.12) and small.

FR-NC-6c — Availability is visible. Every image shall carry a visible availability state: Original, Preview, Metadata only, or Offline.

  • An operation requiring absent data shall say so, with the transfer size, before starting
  • Export from a preview-only image is refused, not silently degraded
  • A pinned set reports its true byte cost before the user commits

Rationale: Lightroom Classic syncs 2560px proxies while displaying the original's filename, extension, and size, so users do not know what they actually have. Sync failures of legibility are more damaging than failures of transport.

FR-NC-7 — Upload. Files above 5MB use chunked upload v2 against /remote.php/dav/uploads/<userid>/: MKCOL to create the upload folder, PUT each chunk, then MOVE the .file pseudo-entry to the destination. Chunks are 5MB–5GB and named 1–10000. OC-Total-Length shall always be sent so quota is checked up front rather than at assembly time. Upload folders expire after 24h of inactivity; the app shall persist upload state and either resume or DELETE stranded uploads on startup.

Small files (sidecars) use bulk upload via POST /remote.php/dav/bulk with a multipart/related body, allowing hundreds of edit-graph sidecars in a single request.

FR-NC-8 — Edit metadata sync. Edit graphs sync bidirectionally as sidecars, one per image, named deterministically from oc:fileid. Each sidecar carries a monotonic revision counter, a per-device UUID, and a last-edit timestamp.

A sidecar holds a keyed set of Versions (FR-CAT-12), not a single edit graph — one image may carry several virtual copies, and a single-graph format could not represent them. Version identity is part of the sidecar schema, so conflict merge (FR-NC-9) operates per-version.

FR-NC-9 — Conflict handling. Sidecar updates use If-Match with the known ETag for optimistic concurrency (If-None-Match: * for creates). On 412 Precondition Failed the app shall fetch the remote sidecar and merge at the edit-graph node level — disjoint edits (e.g. a crop on one device, an exposure change on the other) both survive; genuinely conflicting nodes resolve by timestamp — then retry with the new ETag under a bounded retry count.

The app shall not replicate the desktop client's (conflicted copy) file behaviour. Sidecars are structured data of a few KB; a read-merge-rewrite cycle is cheap and preserves user intent. Only genuinely ambiguous merges surface to the UI. Source RAW files are write-once and shall never generate a conflict.

FR-NC-10 — Offline-first. The app is fully functional offline against cached content. Local sidecar writes are atomic (temp file plus rename) and committed locally before any network round-trip, so editing never blocks on connectivity. Sync resumes automatically when connectivity returns.

FR-NC-11 — Initial catalog build. For first sync of a large remote library, the app may use WebDAV SEARCH (RFC 5323) against /remote.php/dav/ filtered by mimetype and paginated via d:limit/d:nresults, in preference to walking thousands of folders with PROPFIND.

FR-NC-12 — Backend independence. Sync shall be implemented against a backend interface, with Nextcloud as the only implementation in v1. No protocol detail specific to Nextcloud may appear outside its connector.

Backends declare capabilities rather than conforming to a lowest common denominator, because the property that makes Nextcloud sync fast — directory ETags propagating up the tree, so an unchanged root proves an unchanged library — is not a general guarantee. An interface built to the common subset would force full enumeration on every sync (ARCH §8.1).

Where a capability is absent the app shall degrade visibly, not silently:

  • Sync strategy in use is reportable to the user, so a slow backend is visibly slow
  • Without byte-range reads, remote browsing cannot extract embedded previews; the app shall refuse full downloads for browsing on a metered connection and explain why
  • Without conditional writes, sidecar conflict detection falls back to revision comparison, which narrows but does not close the race; this is surfaced as a reduced-safety mode

3.8 Platform integration

Android

FR-PLAT-AND-1 — Storage access. Library access is obtained exclusively via the Storage Access Framework: the user grants one or more document trees through ACTION_OPEN_DOCUMENT_TREE, persisted with takePersistableUriPermission and enumerated via DocumentsContract.

The app shall not request MANAGE_EXTERNAL_STORAGE and shall not depend on READ_MEDIA_IMAGES for RAW discovery. See ARCH §6.9 for why neither is viable.

FR-PLAT-AND-2 — Permission loss. Loss of a previously granted tree permission — revocation, reinstall, removed SD card — shall be detected and surfaced, marking affected images offline per FR-CAT-9 rather than deleting catalog rows.

FR-PLAT-AND-3 — Process lifecycle. An Android process may be killed at any moment. Edit state shall be durable such that process death loses at most the last uncommitted parameter change. On resume the app restores the develop session, its image, and its viewport.

FR-PLAT-AND-4 — Background execution. Long-running sync and export use the platform's managed background execution with the constraints in FR-NC-6, and a foreground service with notification for user-initiated exports. Behaviour under Doze and battery-saver is specified and tested.

FR-PLAT-AND-5 — Memory pressure. The app shall respond to onTrimMemory / ComponentCallbacks2 by evicting caches per NFR-RES-1, in a stated eviction order (GPU tiles first, then proxies, then thumbnails).

FR-PLAT-AND-6 — Intents. Register as a receiver for image view and share intents, and provide share-out of exported results via FileProvider.

Linux

FR-PLAT-LIN-1 — Desktop integration. Follow the XDG Base Directory specification for config, data, cache, and state. Register MIME associations for supported RAW types and ship a .desktop entry.

FR-PLAT-LIN-2 — Display server. Support X11 and Wayland. Where Wayland's colour-management protocol is unavailable, FR-DSP-8's stated fallback applies.

FR-PLAT-LIN-3 — Sandboxed distribution. Where distributed as Flatpak, filesystem access uses portals and credential storage uses the Secret Service portal, both verified to satisfy FR-NC-2 and FR-CAT-1 within the sandbox.

3.9 Culling

Per D11 this is the product's primary differentiator, not an incidental capability.

The opportunity, stated plainly. Photo Mechanic is fast because it displays the camera's embedded JPEG — no demosaic, no database, no import step. FastRawViewer is truthful because it shows a genuine raw histogram, raw-derived clipping, and focus peaking. No shipping tool combines both. The culler/editor split exists only because Lightroom's culling is slow — it is a workaround photographers tolerate, not a workflow they want. A tool that is genuinely Photo Mechanic-fast eliminates the handoff rather than improving it.

FR-CULL-1 — Instant display. Displaying the next image shall not wait on demosaic, catalog import, or full decode. The embedded JPEG preview is shown immediately; higher-quality renders replace it progressively.

Acceptance: next-image display within 50 ms of the input event, sustained across a 3,000-image folder, on both platforms. This is the single most important performance figure in the document — Lightroom's ~2 s stall is the entire reason a competing product category exists.

FR-CULL-2 — Preview ladder. Previews resolve through tiers, each falling through to the next:

  1. Embedded JPEG preview from the RAW container (instant)
  2. Cached proxy from a previous visit
  3. Background full decode, promoted when ready

Some cameras embed previews below sensor resolution, and some embed none. The app shall detect this per camera model and pre-emptively background-render where the embedded preview is insufficient, rather than showing the user a soft image and letting them discover it at zoom.

FR-CULL-3 — Raw-truth overlays. Culling decisions are made against raw data, not the embedded JPEG:

  • Raw histogram — computed from sensor data, not the preview. The embedded JPEG's histogram misrepresents available highlight headroom.
  • Raw-derived clipping indicators — a JPEG's clipping warnings systematically lie about what is recoverable in the raw.
  • Focus peaking — overlays in-focus regions on the preview, so focus is verifiable without zooming to 100%. This removes the largest single source of culling latency from the critical path.

Per D11, zoom-to-100% remains available for certainty; peaking makes it optional rather than mandatory.

FR-CULL-4 — Culling mode. A dedicated full-screen mode with:

  • Auto-advance on judgement — users independently reinvent this in darktable, Lightroom desktop, and Lightroom mobile, which is strong evidence it should be the default rather than an option
  • One-key reject, plus the full rating, flag, and colour-label axes
  • Filter to unjudged, so a session resumes where it stopped
  • Keyboard-driven on desktop; single-thumb reachable on tablet

FR-CULL-5 — Burst and near-duplicate grouping. Group frames by capture-time proximity and image similarity, allowing a burst to collapse to one representative and be judged as a unit.

This is the one automated capability photographers consistently praise, precisely because it is a mechanical grouping problem rather than a taste judgement. Automated selection is distrusted — the documented failure is rejecting the only frame of an important moment because someone blinked.

FR-CULL-6 — Comparison. Side-by-side and survey comparison of a selection, with synchronised zoom and pan, for choosing among near-identical frames.

FR-CULL-7 — Tablet culling. Culling shall be fully usable on tablet, as the validated multi-device workflow (§3.5, D11). Requires only ratings and small proxies to sync, not full originals — a substantially smaller sync problem than develop parity.

Design note: pinch-zoom accidentally triggering ratings is a documented defect in Lightroom mobile. Gesture and rating targets must not overlap.

3.9.1 People

Face recognition was deferred in §7 through the 2026-08-08 calibration. It is undeferred here in a narrower form, and the narrowing is the point.

What changed. The deferral treated "face recognition" as an AI feature adjacent to subject masking. It is not the same kind of thing. Masking is a taste operation applied to one image; grouping photographs by who is in them is a mechanical grouping problem over the whole library, which is the category FR-CULL-5 already commits to and already justifies: grouping is the automated capability photographers consistently praise, because it organises without deciding. Every argument FR-CULL-5 makes for burst grouping applies unchanged to people grouping. Answering "where are the frames with the bride in them" across a 4,000-image wedding is a culling operation, and culling is the differentiator.

What is deliberately not in scope, because it is the failure FR-CULL-5 names: no automated selection. Nothing here rejects a frame, ranks a face, scores a smile, or detects a blink. The feature produces a filter, never a judgement. The user's rating axes remain the only thing that rejects a photograph.

FR-CULL-8 — Face detection. The app shall detect faces in library images as a background job, producing per-face a bounding box, five-point landmarks, a detector confidence, and a 512-dimension embedding.

Detection runs against the thumbnail or proxy tier, never a full decode (FR-CULL-2's ladder). This is what makes indexing affordable: a library that has been browsed has already paid for its proxies, so face indexing adds no RAW decodes that were not already happening. Where no proxy exists, the job requests one at background priority rather than decoding inline.

Detection is a job in the FR-CAT-3 queue and inherits its properties without exception: coalesced per image, interruptible, resumable across process death (FR-PLAT-AND-3), and strictly preempted by visible work (NFR-ARCH-2). A library indexes while idle or it does not index; it never competes with the grid.

Acceptance: indexing a 10k-image library completes without the grid dropping below NFR-P9's interaction target at any point, and survives being killed and restarted with no repeated work beyond the in-flight image.

FR-CULL-9 — Calibrated identity. Face similarity shall be expressed as a calibrated probability that two faces are the same person, not as a raw embedding distance. Every threshold in the subsystem — clustering, suggestion, auto-confirmation — shall be stated in that probability space, and no code path may threshold a bare cosine similarity.

This is a hard requirement rather than an implementation detail because the failure mode is invisible. A raw cosine means something different for every model, every population, and every face size; a threshold tuned on one library silently misbehaves on another, and an uncalibrated similarity still looks like a plausible number all the way to the user interface. A displayed confidence that does not mean what it says is worse than no confidence, because it is trusted.

The calibration shall be fitted per library from that library's own faces, and shall report whether it is valid. Where it is not — too few examples to fit — the app shall say the confidence is unavailable rather than present an untuned default as though it were measured.

Acceptance: on a labelled corpus, the stated probability is within a documented tolerance of the observed match rate across the probability range (a reliability-diagram check, not a single accuracy figure).

FR-CULL-10 — Clustering and naming. Detected faces shall be clustered into unnamed groups. The user names a group, and that name applies to its members. A person is thereafter a first-class catalog entity with a stable UUID, independent of any name given to them.

The user shall be able to merge two groups that are the same person, split a group that is not, remove a face from a person, and rename a person, at any time and without re-indexing. Splitting must be as easy as merging: clustering will over-merge on siblings, on parents and children, and on the same person a decade apart, and a tool that can only merge makes its own errors permanent.

Confirmation is explicit. A face is either suggested (the system's inference) or confirmed (the user's judgement), and the two are never conflated in storage or in display. Suggestions may be recomputed freely; confirmations are user data and are never overwritten by a later inference pass.

FR-CULL-11 — People as a selector term. A person shall be a term in the §5 selector language, composable with every other term.

This is the requirement that pays for the subsystem, and it is nearly free once FR-CULL-10 exists: because one predicate language serves the library filter, smart collections, and cache rules, a person term yields all three at once — filter the grid to a person, save "every photo of Anna rated three or higher" as a smart collection, and pin "every photo of my children" to stay local on the tablet. The last is a genuinely new capability, not a restatement of the first two.

Selectors shall distinguish confirmed from suggested membership, defaulting to confirmed-only, so a saved collection does not silently change membership when a later indexing pass revises a guess.

FR-CULL-12 — Names are user data; embeddings are not. A confirmed person name is a user judgement of the same class as a rating or a keyword, and shall be written to the sidecar (FR-CAT-8), so it survives catalog deletion and travels with the photograph.

Embeddings, detections, cluster assignments, and unconfirmed suggestions are derived data. They live in the catalog only, are rebuildable by re-indexing, and are never written to a sidecar. This follows ARCH §6.12 exactly: the expensive-but-reproducible artefact stays in the disposable index, and only the irreplaceable human judgement enters the trust path.

The person UUID is what a cross-device merge keys on, in the same way collections merge (FR-CAT-7). Two devices that independently name the same cluster produce two people; merging them is the ordinary FR-CULL-10 merge, not a special case.


4. Non-functional requirements

4.1 Performance targets

These are targets to design against and measure, on the reference desktop (AMD Threadripper 2920X, 24 threads, discrete GPU) and a mid-range Android device.

ID Metric Desktop target Android target
NFR-P1 Catalog open (50k images) < 2 s < 4 s
NFR-P2 Grid scroll Sustained 60 fps Sustained 60 fps
NFR-P3 Thumbnail generation throughput ≥ 100 img/s (embedded preview path) ≥ 25 img/s
NFR-P4 Open image in develop (to first proxy on screen) < 400 ms < 1000 ms
NFR-P5 Slider adjustment → visible update < 16 ms (one frame) < 33 ms
NFR-P6 Pan/zoom responsiveness No dropped frames at 60 fps No dropped frames
NFR-P7 Full-resolution export (24MP, full chain) < 2 s < 8 s
NFR-P8 Idle memory (50k catalog, nothing open) < 500 MB < 250 MB
NFR-P9 UI-executor blocking (not "any operation") Never > 16 ms Never > 16 ms
NFR-P13 Next image in culling mode (FR-CULL-1) < 50 ms < 50 ms
NFR-P14 Focus peaking overlay ready < 100 ms after preview < 150 ms
NFR-P15 Drawn mask stroke → visible (ARCH §6.11) < 16 ms, no cursor lag < 16 ms
NFR-P10 Touch gesture → visual response < 16 ms < 16 ms
NFR-P11 Layout class transition (window resize) No dropped frames, no state loss n/a
NFR-P12 Warm-start shader pipeline setup (cached) < 100 ms < 100 ms

Every target above requires a stated measurement method, workload, and pass threshold before it is testable. NFR-P8 in particular must state whether it measures RSS inclusive or exclusive of GPU allocations, and whether it holds after SQLite's page cache warms on a 50k catalog.

Performance regressions fail the build. §9's benchmark suite runs per-commit; a regression beyond a stated tolerance is a build failure, not a notification. Performance work rots otherwise.

4.2 Reliability

NFR-R1 — No operation shall lose user edit data. The catalog database uses write-ahead logging and survives power loss without corruption.

NFR-R2 — The catalog is backed up automatically on a schedule and before schema migrations.

NFR-R3 — A crash in decode or GPU work shall not take down the application where it can be isolated; the affected image is marked as failed and the app continues.

NFR-R4 — Source image files are strictly read-only to the application, except where the user explicitly requests a destructive operation (e.g. delete, or writing XMP sidecars).

NFR-R5 — Schema versioning. The catalog carries a monotonic schema version. Migrations are forward-only, transactional, and idempotent on retry. Each migration ships with a test that migrates a fixture catalog from every prior released version. The app refuses to open a catalog with a newer schema version rather than corrupting it, and says so.

ARCH §6.6 already anticipates one migration (folder ETags); there will be others, and the machinery must exist before the first one.

NFR-R6 — Corruption recovery. On failing an integrity check at startup, the app offers restore from the NFR-R2 backup, and failing that, rebuild from sources plus sidecars per invariant 5.2.4. FR-CAT-8 is what makes that second path real for local-only users.

NFR-R7 — GPU device loss. The GPU layer treats device loss as an expected event (ARCH §6.10): detect it, tear down and recreate the device and all derived resources, and re-drive the current render from the edit graph. No user edit is lost. Recovery is exercised by a test that induces device loss mid-render.

NFR-R8 — No suitable GPU. Where no Vulkan device meeting the NFR-COMPAT-1 baseline is available, the app starts in a stated degraded mode with defined capability limits rather than failing to launch.

This requires reconciling ARCH §6.4 and NFR-RES-2: ARCH §6.4 says the GPU path is primary rather than an optimisation, while NFR-RES-2 assumes a CPU fallback on allocation failure. Decide explicitly whether v1 includes a full CPU pipeline, or whether "CPU fallback" means only tile-spill staging with no independent CPU render path. The latter is recommended; the former is a second full implementation.

4.3 Resource behaviour

NFR-RES-1 — Bounded memory. Memory use is bounded and configurable, independent of catalog size and image count. Caches are evictable under pressure.

NFR-RES-2 — GPU memory. The pipeline shall handle images larger than available GPU memory by tiling. GPU memory headroom is configurable, with a CPU fallback path if allocation fails.

NFR-RES-3 — Mobile power. On Android the app shall not render continuously when idle. Battery and thermal behaviour are first-class concerns; background sync respects metered-connection and battery-saver settings.

NFR-RES-4 — Disk cache. Thumbnail and proxy caches have a configurable size cap with LRU eviction.

4.4 Portability

NFR-PORT-1 — Platform-specific code is isolated behind interfaces. The image core, catalog, and edit pipeline contain no platform conditionals.

NFR-PORT-2 — GPU shaders are authored once and used on both platforms.

NFR-PORT-3 — Adding a third platform requires implementing the platform interfaces only, not changes to the core.

4.5 Security and privacy

NFR-SEC-1 — RAW parsing treats input as untrusted. Parser hardening, fuzzing, and where practical process or memory isolation for decode.

NFR-SEC-2 — Credentials are never written to the catalog, logs, or plain files. Platform secure storage only.

NFR-SEC-3 — All network traffic uses TLS with certificate validation. No option to disable validation in release builds.

NFR-SEC-4 — No telemetry without explicit opt-in.

NFR-SEC-5 — Face data stays on the user's own hardware. Face embeddings (FR-CULL-8) are handled under a stricter rule than the rest of the catalog.

This is a personal tool for personal libraries (§1.1, D11) — the people in these photographs are the user's family and friends. That is the reason for the rule, not a reason to relax it: the data is sensitive precisely because it is personal, and the user is the only party with any claim on it.

  • Never leave the device by default. Embeddings, face crops, and cluster assignments shall not be transmitted, uploaded, or included in any diagnostics bundle (NFR-OPS-1) or crash report (NFR-OPS-2), under any configuration. The diagnostics path has no opt-in for this; it is excluded outright.
  • Sync is opt-in and separately consented. Syncing embeddings to the user's own Nextcloud is permitted — it is their server and their photographs, and it saves re-indexing a library per device — but it is off by default, is not implied by enabling photo sync, and the consent states in plain language what is being uploaded and why. Person names, being sidecar data (FR-CULL-12), sync with the sidecar as ordinary metadata.
  • No third-party inference. Face detection and embedding run locally. No image, crop, or embedding is sent to a remote inference service, and the app ships no capability to do so.
  • Deletable, in one action. The user shall be able to delete all face data — embeddings, detections, clusters, and people — from a single control, without deleting the catalog or any photograph, and to disable face indexing entirely so that no such data is produced.
  • Model weights are inspectable. The models used shall be named and versioned in the about screen, with their licences, so a user can determine what is running on their photographs.

Rationale: the rest of this document treats privacy as a property of the network boundary — TLS, credentials in secure storage, opt-in telemetry. Face data needs more than a well-defended boundary, because it is not revocable once it has crossed one, and because it describes people who are not the user. The prohibition is therefore structural rather than configurable: the code paths that would upload an embedding to anyone but the user's own server do not exist. A setting can be changed by accident, or by a future maintainer who has forgotten why it was there; an absent code path cannot.

4.6 Execution model

NFR-ARCH-1 — Named executors. The app defines distinct executors — UI, GPU submission, decode pool, I/O pool, network — with stated thread counts and the invariant that no blocking call occurs on the UI executor. This is the mechanism behind R4 and NFR-P9, which currently assert an outcome with no stated means.

NFR-ARCH-2 — Scheduler priority. The tiling scheduler assigns priority classes, with visible-tile work strictly preempting background export and thumbnail work. Without this, NFR-P5's slider latency fails during a batch export — the common case, not an edge case.

NFR-ARCH-3 — Cancellation. Cancellation is cooperative with a bounded worst-case latency (target: observed within 100 ms), and covers in-flight GPU submissions. Every long-running operation named in FR-CAT-1, FR-EXP-7, and FR-NC-6 is cancellable under this model.

NFR-ARCH-4 — Error propagation. No worker error may panic the process. Errors surface as typed results attached to the affected image or job, consistent with NFR-R3 and FR-RAW-4.

4.7 Operations

NFR-OPS-1 — Diagnostics. Structured levelled logging to a rotating, size-capped on-disk log in the XDG state directory or Android app directory, with automatic redaction of credentials and tokens (required by NFR-SEC-2, which currently forbids credentials in logs that are never otherwise specified). A one-click diagnostics bundle includes log, schema version, GPU and driver identification, and app version — with an explicit preview-and-consent step before anything leaves the device.

NFR-OPS-2 — Crash reporting. Local crash capture always; upload only on explicit opt-in (NFR-SEC-4).

NFR-OPS-3 — Preferences. A single versioned preferences store, separate from the catalog, so preferences survive catalog rebuild and multiple catalogs. At least eight requirements refer to configurable settings with no store defined. Device-specific settings (GPU headroom, cache caps) do not sync between devices.

NFR-OPS-4 — Update and first run. State delivery channels and their update mechanisms. This matters concretely because D2 pins rawler at an alpha, non-SemVer version whose camera-support fixes users will need. First-run flow is defined, including platform permission acquisition (FR-PLAT-AND-1 makes first run a permission negotiation on Android, not a welcome screen) and initial root selection.

4.8 Compatibility baseline

NFR-COMPAT-1 — Supported hardware. Every §4.1 Android figure is meaningless without this. State:

  • Minimum and target Android API level (targetSdk 36 is currently required for Play distribution)
  • Minimum Vulkan version and the required feature and limit set — including whether shaderFloat16 and 16-bit storage are required, since FR-DEV-2's f16 pipeline depends on them and their absence would jeopardise R1
  • Minimum device RAM, and minimum desktop Vulkan/Mesa versions
  • The specific reference Android device the §4.1 column is measured on, plus a secondary device from a different GPU vendor

Adreno, Mali, and PowerVR diverge significantly in compute behaviour and in external-memory interop — exactly what spike S1 tests. The spec already applies this reasoning to desktop drivers; it applies at least as strongly on Android.

NFR-COMPAT-2 — Distribution channels. State the v1 channels (e.g. Flatpak and AppImage on Linux; Play Store and/or F-Droid on Android). Play distribution is what makes ARCH §6.9's constraints binding — a sideloaded or F-Droid build could use different permissions, so the channel decision and the storage design are coupled.

4.9 Accessibility and internationalisation

NFR-A11Y-1 — Localisation. All user-facing strings, including operation and parameter labels resolved from LocalizedString (FR-DEV-3a), are externalised and translatable without recompilation. Note the constraint this creates: those labels live in core crates that cannot depend on the UI (ARCH §6.5a), so the localisation mechanism must itself be UI-independent. State the format, the locale-resolution rule, and whether RTL layout is in v1 scope.

NFR-A11Y-2 — Accessibility. Controls expose accessible names, roles, and values to the platform accessibility layer (AT-SPI on Linux, TalkBack on Android). Platform font scaling is honoured without clipping. Non-canvas UI meets WCAG AA contrast.

Slint's accessibility support on Android requires verification — this may be a toolkit gap, and it is far cheaper to discover now than after the UI is built.

NFR-A11Y-3 — Colour-independent status. No status is conveyed by hue alone. Colour labels, clipping indicators (FR-DSP-7), and the HSL mixer carry a shape or text affordance. This matters more in a colour-grading application than in most software.


5. Data model and architecture

Entity definitions, invariants, and all architectural constraints are specified in architecture.md — §3 (core abstractions), §6 (data architecture), and §11 (architectural constraints).

Requirements in this document that depend on an architectural guarantee cite it inline. The constraints most load-bearing for testability are:

Constraint Why a requirement depends on it
ARCH §6.1 — no CPU round-trip FR-DSP-3, FR-DSP-7, NFR-P5 are unachievable without it
ARCH §6.11 — GPU-rasterised masks NFR-P15 (no brush lag)
ARCH §6.12 — sidecars authoritative NFR-R6, invariant behind FR-CAT-8
ARCH §6.9 — Android SAF only FR-CAT-1a, FR-PLAT-AND-1, and the NFR-P1/P3 Android figures
ARCH §6.6 — no sync tokens FR-NC-4
ARCH §6.13 — integer-only bit-identity R1's tolerance, §9 golden images

6. Decisions

Rationale, evidence, and the eliminated alternatives are recorded in architecture.md §12. Outcomes only:

# Decision Outcome
D1 Language and UI framework Rust + Slint, rendering through wgpu
D2 RAW decoder rawler; LibRaw fallback behind a trait
D3 First milestone Provisional — depends on D12
D4 Nextcloud sync mechanism ETag pruning, chunked upload v2, Login Flow v2
D5 Colour management lcms2 + GPU-side matrix/LUT transforms
D6 Shader authoring Hand-written WGSL
D7 Network stack reqwest + quick-xml
D8 Licence GPLv3
D9 Operation UI model Declarative parameter descriptors
D10 Interface strategy One adaptive UI, tablet + desktop
D11 Product positioning Culling-first differentiator; see below
D12 Scope versus pace OPEN

D11 — product positioning

Settled by requirements calibration, 2026-08-08.

Dimension Decision
Audience RAW-literate photographers, Linux-comfortable. Docs matter; hand-holding does not.
Library scale 10k–50k images
Culling The core differentiator (§3.9)
Focus checking Peaking and zoom
Ingest Full workflow — template rename, checksum verify, dual-destination
Colour defaults Good, not obsessive — matrices plus per-body base curve
Film simulation Fujifilm explicitly targeted
AI Denoise in v1; masking deferred
Local adjustments Full masking, GPU-rasterised
Sync The reason the project exists
Durability Sidecar-first
Licence GPLv3
Pace Evenings and weekends, indefinite

D12 — scope versus pace · OPEN

The calibration selected an ambitious feature set — full tablet editing, full ingest, culling as a differentiator, complete GPU masking, AI denoise, Fuji-first colour, deep sync, sidecar durability — against a stated pace of evenings and weekends, indefinitely.

Those are not compatible as stated. This is not an argument against any individual choice; it is that the two must be reconciled before a build order can be set.

Two specific tensions:

1. Full tablet editing is the most expensive selection, chosen against the research finding that volume photographers do not edit on tablets. It carries SAF storage at unproven scale (S10), process-death durability, background execution limits, and two GPU vendors to validate. Tablet culling is the validated workflow, is what the differentiator points at, and costs a fraction as much.

2. The v1 milestone and the differentiator disagree. D3's vertical slice proves the architecture but is useful to nobody. If culling is what no existing tool does well, a culler is both a smaller build and a usable one — it needs no develop chain.

Resolving D12 sets D3 and architecture.md §10's Phase 2.

D13 — face inference runtime and model licensing · RUNTIME ANSWERED, LICENSING OPEN

Updated 2026-08-21. The runtime half of this decision is settled, and by a route the table below does not contain. ort 2.0's alternative-backend feature disables its linking entirely and lets another engine supply the OrtApi; ort-tract supplies it from tract, which is pure Rust. So the third option's operator coverage comes with the first option's dependency profile — no C, no NDK problem, no exception to the policy. Measured on a real graph before being relied on: YOLO26n-seg loads with zero unsupported operators and runs 640×640 in ~470 ms of CPU (docs/segmentation.md §13). §3.9.1's detector and embedder are different graphs and their coverage has not been checked, but the approach no longer needs a decision.

The licensing half is untouched. The InsightFace weights are still non-commercial and still unusable here. That remains what S14 has to resolve first.

§3.9.1 needs to run two neural networks locally. That collides with two settled positions, and neither collision is small enough to leave implicit.

1. The pure-Rust dependency policy. Every dependency choice in this project has gone the same way, for the same stated reason: rustls over aws-lc-rs, bundled SQLite over the system library, a Rust Lensfun port over liblensfun, zune-jpeg over libjpeg — no C dependency to satisfy under the Android NDK (D1's whole premise). The obvious way to run ONNX models is the ONNX Runtime C++ library, which would be the largest exception to that policy in the codebase, and it would land on the platform the policy exists to protect.

The options, in the order I would try them:

Option Cost
wgpu compute, models hand-ported to WGSL No new dependency at all — the GPU device and shader infrastructure already exist (ARCH §5). Highest implementation effort, and a ViT is a lot of shader.
burn with the wgpu backend Pure Rust, uses the existing GPU. Young, and ONNX import maturity needs checking against these two specific graphs.
ort (ONNX Runtime bindings) Fastest to working code, best operator coverage. Reintroduces the C dependency and the NDK cross-compilation problem the policy avoids.

The tension is real: the cheapest path is the one that breaks the rule. This is worth an explicit decision rather than a default, and S14 is what informs it.

2. Model licensing is a distribution blocker, not a detail. The obvious pretrained weights are not redistributable under GPLv3. The InsightFace "buffalo" family — ArcFace and the SCRFD detector, the standard choices — are licensed for non-commercial research use only, which is incompatible with this project's licence and with Flatpak, F-Droid, and Play distribution (NFR-COMPAT-2). Other candidate weights need their licences read individually rather than assumed.

Two ways out, both with costs:

  • Find permissively-licensed weights and ship them in-tree. Clean, offline-first, consistent with how the Lensfun database ships. Requires that suitable weights exist at acceptable accuracy.
  • Download models on first use, with the user accepting the upstream licence. Sidesteps redistribution but adds a network dependency to a feature that is otherwise entirely local, needs a hosting story, and sits badly with the local-first posture of NFR-SEC-5.

This must be resolved before implementation, not during it. Discovering at packaging time that the feature cannot ship is the expensive failure, and it is entirely avoidable — it is a licence- reading exercise, not a research question. S14 therefore puts it first.

Prior art available: ../scene-actor-extraction is a working implementation of this pipeline (SCRFD detect → 5-point align → 512-d embedding → Platt-calibrated similarity), benchmarked at 67.4% macro-F1 on held-out films. Its C++ does not port — different language, OpenCV and TensorRT dependencies — but its design decisions do, and they are the expensive part: the calibrated probability space that FR-CULL-9 requires, the discipline of never thresholding a bare cosine, and the practice of leaving an uncertain face honestly unnamed. Personal libraries should also score better than its film benchmark: cooperative subjects, better lighting, and a closed gallery of dozens rather than thousands.


7. Out of scope for v1

Deferred deliberately. Listed so their absence reads as a decision rather than an oversight, with a note where deferring now constrains the design later.

Deferred Note
Tethered shooting —
Panorama and HDR merge Keep the schema open — these produce images derived from multiple sources, which §5.1's single-source Image cannot express.
Focus stacking Same provenance consideration.
Print layout —
Soft proofing Parameterise the output colour stage by an arbitrary profile so this becomes a UI addition, not a pipeline change. FR-EXP-3's print-dimension mode already half-commits to print workflows.
Face recognition Undeferred 2026-08-09, in the narrower form specified in §3.9.1 (FR-CULL-8 … FR-CULL-12): people grouping and search, no automated selection. Reclassified as culling rather than AI — it is the same mechanical-grouping category as FR-CULL-5, not the taste operation AI masking is. Gated on spike S14 and decision D13.
AI subject masking Deferred per D11. Note darktable shipped this in 5.6 (June 2026), so the gap is now visible. Conditions for deferring safely: AI denoise ships in v1 (FR-DEV-3g ✓), manual masking is excellent including GPU-rasterised drawn masks (ARCH §6.11 ✓), and the product has a clear differentiator (culling, §3.9 ✓). When it does land, copy darktable's shape — prompt-point segmentation producing an editable mask that behaves like a hand-drawn one — not Adobe's opaque version.
AI upscaling Deferred. Lower priority than denoise, which has no manual fallback.
Video —
Plugin API —
Multi-user / server-side catalog —
Watermarking Cheap if the export pipeline anticipates a compositing stage; expensive to retrofit otherwise. Consider reserving the stage now.
Multiple catalogs, catalog merge Interacts with NFR-OPS-3: preferences must not live in the catalog.
Geotagging and map view —
Web gallery, slideshow —
DNG conversion —

8. Verification approach

Requirement class How verified
Performance (§4.1) Automated benchmark suite against a synthetic 50k catalog, run per-commit on the reference desktop and periodically on the named reference Android devices. A regression beyond stated tolerance fails the build.
Rendering correctness Golden-image tests: fixed source + fixed edit graph → comparison within R1's stated tolerance, not checksum equality. Run on both platforms and both Android GPU vendors.
Colour accuracy (FR-DEV-3e) ColorChecker exposures per launch body, asserting ΔE2000 within threshold against reference values.
RAW decode coverage Corpus of sample files per supported camera body; decode-and-checksum regression suite.
Robustness (NFR-SEC-1) Continuous fuzzing of the decode path.
Sync correctness Simulated two-device scenarios including conflict, offline edit, and interrupted transfer.
Memory bounds Long-running soak test scrolling a large catalog, asserting bounded RSS and GPU memory.
Schema migration (NFR-R5) Fixture catalogs from every prior released version, migrated forward and verified.
Device loss (NFR-R7) Induced VK_ERROR_DEVICE_LOST mid-render; assert recovery with no lost edits.
Process death (FR-PLAT-AND-3) Kill the Android process mid-edit; assert session and viewport restore with at most the last uncommitted change lost.
Source relocation (FR-CAT-9) Move, rename, and disconnect sources; assert offline marking, reconnection by hash, and no catalog row loss.
Cancellation (NFR-ARCH-3) Assert every long-running operation observes cancellation within the stated bound, including in-flight GPU work.
Layer separation (ARCH §6.5a) CI dependency-tree assertion: no core/* crate may transitively depend on a UI toolkit.
Operation self-description (FR-DEV-3c) A test operation added to the registry appears in a generated panel with no frontend change.
Adaptive layout (§3.5) Snapshot tests at each breakpoint, and a resize test asserting no state loss across a layout-class transition.
Touch targets (FR-UI-3) Automated check that interactive elements meet the 44pt minimum in touch modality.
Export sizing (FR-EXP-3) Per-mode dimension assertions, including aspect preservation, fill-crop centring, and the upscale-disabled fallback.
Identity calibration (FR-CULL-9) Reliability diagram over a hand-labelled corpus: stated probability against observed match rate, asserted within tolerance across the range — not a single accuracy figure, which would hide exactly the miscalibration this tests for. Plus a static assertion that no comparison thresholds a raw similarity.
Face data confinement (NFR-SEC-5) Assert that a generated diagnostics bundle contains no embedding or face crop, and that with sync disabled no face data appears in any outbound request. Verified by inspecting what the code can emit, since the requirement is the absence of a path.

9. Validation spikes

Small experiments that de-risk the highest-uncertainty assumptions before substantial build work. Ordered by risk. With D1 settled these validate the chosen stack rather than choosing between stacks.

Tier 1 — before any substantial build work

# Spike Answers Relates to
S1 Slint + wgpu zero-copy on Linux: a compute shader writes a texture, create_texture_from_hal imports it, Slint composites UI over it. Drag a slider for 10 minutes watching for tearing, leaks, and sync bugs Whether ARCH §6.1 holds in the chosen stack D1, ARCH §6.1
S2 Slint + wgpu on Android, on two devices from different GPU vendors (Adreno and Mali) Whether the Android GPU path holds across vendor divergence D1
S10 Android SAF at scale: enumerate a 10k-file document tree and perform random-access range reads over a document fd. Measure against NFR-P1 and NFR-P3 Whether ARCH §6.9's forced storage model meets the stated Android performance targets ARCH §6.9, FR-PLAT-AND-1
S11 Play Console permissions dry-run: submit an actual declaration for this app category before committing to the storage design Whether Google approves anything beyond SAF ARCH §6.9
S9 Golden-image comparison of one edit graph rendered on desktop and Android; calibrate the achievable tolerance What R1's tolerance threshold should actually be R1

Tier 2 — before the corresponding subsystem is built

# Spike Answers Relates to
S3 reqwest HTTPS PROPFIND on a real Android device, including the rustls-platform-verifier Kotlin init The largest known Rust-on-Android networking risk D7
S4 Range-extract an embedded JPEG from CR3/NEF/ARW over WebDAV; measure bytes transferred Whether remote browsing on mobile data is viable ARCH §6.7, FR-NC-3
S5 ETag pruning against a 10k-file library: confirm one-request no-op sync, and correct propagation on a single deep-file change Whether FR-NC-4 scales as designed ARCH §6.6
S6 Tiled GPU pipeline on a mid-range Android device, with an image larger than available GPU memory Whether ARCH §6.2 holds on constrained hardware NFR-RES-2
S7 rawler decode coverage across the FR-RAW-1 launch set, on real files from each body Whether the LibRaw fallback is needed at launch or later D2, FR-RAW-1
S8 Chunked upload v2 round-trip of a 100MB RAW, including resume after process kill FR-NC-7 correctness FR-NC-7
S12 GPU device loss recovery: induce VK_ERROR_DEVICE_LOST mid-render, verify recreation from the edit graph with no lost edits Whether ARCH §6.10 and NFR-R7 hold ARCH §6.10
S13 Slint accessibility on Android: verify TalkBack exposure of names, roles, and values Whether NFR-A11Y-2 is achievable in the chosen toolkit NFR-A11Y-2
S14 Face pipeline in Rust, on a real personal library: run a detector plus an embedder over ~2,000 images through a Rust ONNX runtime, at proxy resolution, on the reference desktop. Measure per-image latency, cluster purity against hand-labelled truth, and fit the FR-CULL-9 calibration to see whether it converges on a library-sized sample. Resolve the model licence question before writing any of it Whether §3.9.1 is buildable without breaking the pure-Rust dependency policy, and whether the accuracy is worth the subsystem D13, FR-CULL-8, FR-CULL-9

Why this order

S1, S2, and S10 are the three that can invalidate the architecture. S1 and S2 test ARCH §6.1 — the constraint the whole design is built around, and the one darktable's documentation identifies as their biggest bottleneck. S10 tests whether the Android storage model forced by ARCH §6.9 can actually meet the performance targets; it is the highest-uncertainty assumption in the document because until this revision it was unstated.

S11 costs almost nothing and de-risks S10 definitively. Confirming what Google will approve for this app category before designing around it is far cheaper than discovering it at submission.

S9 moved to Tier 1 because it does not merely test R1 — it calibrates it. R1's tolerance threshold cannot be fixed sensibly without knowing the real cross-vendor deviation, and the §9 golden-image strategy depends on that number.

Test S1 on Mesa/AMD, Intel, and NVIDIA proprietary drivers, under both X11 and Wayland. FD-based external memory has well-documented driver divergence, and the reference machine's discrete GPU will not surface Intel or Mesa-specific issues on its own. The same reasoning is why S2 requires two Android GPU vendors.


10. Glossary

  • Edit graph — the ordered set of parameterised operations defining how an image is rendered.
  • Proxy — a reduced-resolution render used for display.
  • Tile — a sub-rectangle of an image processed independently.
  • Demosaic — reconstructing full RGB from a colour-filter-array sensor capture.
  • CFA — colour filter array (Bayer, X-Trans).
  • Sidecar — a small file alongside the source holding edit metadata.
  • Pixel pipeline — the ordered chain of processing stages from sensor data to output.