Every raw rendered through a camera profile — the library's DNGs with an embedded profile, and CR2s given one — now renders differently: more colourful in near-neutral tones. The profile's look table is no longer applied unless its slider is raised; PROFILE_LOOK names the strength the profile states. Against the photographer's earlier exports with no look applied, the default rendering scores the same with the look table at 100, 50 or 0 (held-out MSE 140, 140, 143), and is 9 % more colourful at 0: the table lowers the saturation of near-neutral tones, which is exactly where the default rendering was short of those exports. The user chose more colour.
199 KiB
DarkRoom — Requirements Specification
Status: Living document · first written 2026-08-08 · audited 2026-09-19 Owner: Duncan Tourolle
A cross-platform, non-destructive RAW photo editor for Linux desktop and Android, in the Lightroom idiom: a catalog of many thousands of images, a develop module with GPU-accelerated adjustments, and export to standard 8-bit (or higher) deliverables.
1. Scope and intent
1.1 What this is
DarkRoom is a photo library and develop application. It manages large collections of camera RAW files, renders them to screen with GPU acceleration, applies non-destructive edits stored as metadata, and exports finished images.
1.2 Primary platforms
| Platform | Priority | Notes |
|---|---|---|
| Linux desktop | Primary | X11 and Wayland. Development and reference platform. |
| Android | Primary | A 12-inch tablet (D15). Phones are not a target: the build runs on one, and nothing is designed for one. Shares the image core. |
Other platforms (Windows, macOS, iOS) are explicitly out of scope for v1, but the architecture must not foreclose them. In practice this means the GPU abstraction and the image core must not hard-code Vulkan-only or Linux-only assumptions at their public interfaces.
1.3 What this is not
- Not a DAM with server-side multi-user collaboration
- Not a pixel editor (no layers, no brushes in v1 beyond local-adjustment masks)
- Not a printing/soft-proofing suite in v1
1.4 Architecture
The technology stack, crate layout, and design are specified separately in architecture.md. This document states what the software must do; the architecture document states how.
Requirements here reference architectural constraints as ARCH §n where the constraint
materially shapes what is testable.
2. Core requirements (from the brief)
These are the user's stated requirements, restated as testable criteria.
| ID | Requirement | Acceptance criterion |
|---|---|---|
| R1 | Cross-platform: Linux + Android | Same image core compiles and runs on both. For the same input and edit graph, output is perceptually identical within a bounded tolerance — see below. |
| R2 | Efficient display of huge RAW libraries | A 50,000-image catalog scrolls at 60fps sustained, with a stated prefetch margin and cache-hit rate sufficient that no cell renders as a placeholder at a scroll velocity of (figure TBD) rows/second. Catalog opens in under 2s. |
| R3 | Make a RAW beautiful at 8-bit output | Full non-destructive develop chain at high internal precision, with camera input profiles (FR-DEV-3e) and a colour-managed path to 8/16-bit export. |
| R4 | HW acceleration and parallelism | All per-pixel work runs on GPU compute. CPU work (decode, I/O) is parallelised across cores. The UI executor never blocks on image work (NFR-ARCH-1). |
| R5 | Work on downscaled proxies for display | Display pipeline operates at viewport resolution, not source resolution. The tiling clauses this criterion used to carry have been struck — see below. |
| R6 | Nextcloud integration | Browse, download, and upload images and edit metadata against a Nextcloud instance, offline-capable. |
| R7 | Judge anywhere, on evidence | Rating and flag are reachable from every view that shows a photograph, and apply to the one on screen. Everything the app computes about a frame in aid of culling — clipping, focus, burst membership, per-face state — is shown as evidence the photographer reads, and no code path writes a rating or flag without a user action (FR-CULL-13). |
On R7, added 2026-09-19. Stated from use rather than from the brief. Two things prompted it. Judgement keys had been built into the grid alone, so a photograph opened in develop — the view a photographer is most sure about — could not be rated without leaving it; the rule is that judging follows the photograph, not the view. And specifying per-face signals (FR-CULL-8a) forced the question of what a signal is for, which sharpened what this document already said in FR-CULL-5: the app may know a great deal about a frame and may say all of it, and it never holds the pen. R7 is the user-level statement; FR-CULL-13 is the testable one.
On R1's tolerance. An earlier draft required output to be bit-identical across platforms. That is not achievable and the requirement has been corrected. Floating-point compute results differ between GPU vendors: transcendental function implementations vary, drivers apply different optimisations, and f16 rounding diverges. A checksum comparison across Adreno and Mesa would fail for reasons that have nothing to do with correctness.
R1 is therefore stated as a bounded tolerance — a defined maximum per-pixel deviation, expressed in ΔE2000 for colour or ULPs at the working precision. The threshold must be fixed before spike S9, because S9 both validates R1 and calibrates what the achievable tolerance actually is.
Where genuine bit-identity is required — cache keys, edit-graph hashing (ARCH §3.4, ARCH §6.13) — it applies to integer operations on CPU-side state, which are deterministic, never to GPU float results.
On R5's tiling. An earlier draft added two clauses to R5's criterion: "only visible tiles are computed; panning recomputes only newly exposed tiles". They have been struck, and the reason is the one frame-budget.md measured — the same measurement that argues FR-DSP-2 should be rewritten rather than implemented. FR-DSP-2 itself has not been rewritten: it stands as written until spike S6 has run on constrained Android hardware, because the desktop measurement cannot speak for a device whose GPU memory the image exceeds (decided 2026-09-19; see the note under FR-DSP-2).
R5's actual demand is met and tested. The display pipeline works at viewport resolution:
Framing::view shrinks the sampled region while the render target keeps its size, so zooming raises
the resolution the pipeline works at rather than magnifying pixels already drawn, and
core/dr-gpu/tests/zoom_resolution.rs establishes it as a pixel equality rather than an impression
of sharpness.
Tiling is a different claim, and it was written in as though it were the mechanism by which the
first one is achieved. It is not. Recomputing the entire 4K viewport costs 4.5 ms of a 16 ms
budget, so a perfect tile cache saves at most that, in exchange for a cache keyed by
(VersionId, tile, zoom, graph_hash_prefix) that has to stay correct across every parameter change
in the graph — a large correctness surface bought with a small number. And for the one stage that
does miss the budget, tiling makes it worse: that stage is a convolution, and a tiled convolution
reads a halo per tile, so at the 52 px radius measured at 4K a 256 px tile would read (256+104)²
taps instead of 256², very nearly twice the work.
The intent behind the struck clauses — that the display path must not do work proportional to the source image — survives in the clause that remains, which is the honest statement of it. Tiling stays where FR-DSP-2 puts it: a scheduling concern for export and thumbnailing, both of which already run off the frame path.
3. Functional requirements
3.1 Catalog and library management
FR-CAT-1 — Scan. The app shall scan one or more user-granted library roots for supported image files, recursively, without blocking the UI. Progress is reported and the scan is cancellable and resumable. A "root" is a platform-specific grant (a directory on Linux, a persisted document tree on Android — see FR-PLAT-AND-1), not necessarily a filesystem path.
FR-CAT-1a — Source addressing. The catalog and decode layers shall address source data through
an opaque SourceRef that resolves to a seekable byte stream, never through a filesystem path.
Android's Storage Access Framework provides no usable path (ARCH §6.9), so a path-based API would not be
portable. SourceRef carries enough information to re-resolve after an app restart or a permission
re-grant.
FR-CAT-2 — Catalog store. All catalog metadata (source references, EXIF, ratings, labels, edit graphs, sync state) is stored in a local embedded database. The database is the source of truth for the UI; sources are scanned into it, never queried directly on the UI path.
FR-CAT-3 — Thumbnail pyramid. For each image the app maintains a cached multi-resolution thumbnail set. Initial thumbnails are extracted from the RAW's embedded JPEG preview where present (fast path, no demosaic). Higher-quality proxies are generated lazily from the full decode when the image is first opened in develop.
FR-CAT-4 — Virtualised grid. The library grid shall render only visible cells plus a small prefetch margin. Memory use is bounded and independent of catalog size.
FR-CAT-5 — Metadata. Read EXIF, camera make/model, lens, capture time, ISO/aperture/shutter, GPS. Support user-assigned star ratings, colour labels, flags, and keywords.
FR-CAT-6 — Search and filter. Filter the catalog by any indexed metadata field, rating, label, folder, and keyword, with results updating interactively on a 50k catalog.
FR-CAT-7 — Collections. User-defined collections that reference images without moving files.
FR-CAT-8 — Sidecar persistence, independent of sync. Edit graphs shall be written to per-image sidecars for all catalogued images, whether or not a Nextcloud account exists. Invariant 5.2.4 (catalog rebuildable from sources plus sidecars) otherwise fails for local-only users, leaving every edit in a single SQLite file with no recovery path.
Where the source location is not writable — read-only mounts, and commonly Android SAF trees —
sidecars are written to an app-managed store keyed by SourceRef, and the app shall state which
location is in use.
FR-CAT-9 — Offline and relocated sources. An image whose source is unreachable shall be marked offline, never silently removed. Cached previews, metadata, ratings, and edits remain browsable and editable while offline; edits queue and apply when the source returns.
The app shall support folder-level and image-level reconnection, matching candidates by content hash and filename, and shall auto-reconnect a volume or tree when it reappears. A source deleted outside the app shall be distinguished from one merely unreachable before any destructive catalog action is offered.
This matters more than it appears: external drives, SD cards, and network mounts disappear routinely, and on Android a tree permission can be revoked or lost on reinstall.
FR-CAT-10 — Import and ingest. Copy or move files from a source volume into a destination structured by a date/metadata template, with rename-on-import, an optional simultaneous second-destination backup copy, and per-file verification against a checksum. Removable-volume insertion is detected where the platform permits.
Distinct from FR-CAT-1: scanning catalogues files where they already are; import moves them from a card into the library. Both are needed.
FR-CAT-11 — Duplicate detection. Detect duplicates on import by (capture time + camera serial +
original filename) and by content hash, offering skip or import-as-new. Camera filenames wrap at
IMG_9999, so filename alone is insufficient. Existing catalog duplicates are detectable on demand.
FR-CAT-11a — Consolidating catalog duplicates. Files already in the library more than once (same root, camera, capture time and size) shall be listed on demand as groups, from the library and from Settings. Before anything moves, each group shall be proved the same file by a stored full digest or by a digest of the first and last megabyte of every copy, and its copies' develop edits compared; a group that differs in either is left out and the review says why. One copy survives — not under a backup-looking folder, then camera-named, then oldest, overridable per group — and the others' collections, keywords, highest rating, agreed flag and label, and faces are merged onto it, with disagreements reported rather than decided silently. Each group is merged and its other copies moved to the trash (FR-CAT-15) in one transaction: wholly consolidated or untouched. Nothing is deleted.
FR-CAT-12 — Versions (virtual copies). An image may carry multiple named Versions, each with
an independent edit graph, without duplicating source data. Versions are creatable, nameable,
deletable, and independently exportable; one is the default.
FR-CAT-13 — XMP interoperability. Read and write standard XMP sidecars for ratings, colour
labels, keywords and hierarchical subjects, title, description, copyright, and GPS, using standard
xmp:/dc:/lr: schemas so other tools interoperate.
DarkRoom's edit graph lives in a private namespace and shall neither be interpreted by, nor corrupt, other tools' XMP. Writing to source-adjacent XMP is off by default (NFR-R4). External modification of an XMP sidecar shall be detected and a metadata reload offered.
FR-CAT-14 — Migration import. Import ratings, labels, keywords, and collections from a
Lightroom .lrcat and a darktable library.db. Edit graphs are explicitly not migrated —
develop parameters do not translate meaningfully between pipelines, and a partial translation is
worse than none. This is the path in for users with existing libraries.
FR-CAT-15 — Trash and permanent delete. Deleting an image shall be reversible by default. A
soft delete moves the file into a .darkroom-trash/ folder under the library root and records
in the catalog when it was trashed and the path it came from; restore moves it back to that path.
Permanent delete removes the file first and the catalog row second, and a delete of something
already gone counts as success.
A flag alone would not survive ARCH §6.12 (the catalog is a rebuildable index): the catalog is rebuildable from sources, so a rescan would find every "deleted" file still in the library and re-index it. The folder is the durable fact and the row is the convenience — which also means the scanner shall exclude the trash folder, and that a user can recover by hand without DarkRoom. Derived data keyed on the file (thumbnails, cached previews) is dropped when the image is permanently deleted, not when it is trashed.
The trash shall be listable newest-first, with the count and total bytes it holds shown before any destructive action, since that figure is what tells the user whether they meant it.
3.2 RAW decoding
FR-RAW-1 — Format support. Decode mainstream RAW formats. Minimum launch set: Canon (CR2, CR3), Nikon (NEF), Sony (ARW), Fujifilm (RAF, including X-Trans), Panasonic (RW2), Olympus (ORF), Adobe DNG. Additional formats are a coverage goal, not a launch blocker.
FR-RAW-2 — Decoder abstraction. RAW decoding takes bytes, never a filesystem path and never
a reference it would have to resolve. Resolving a SourceRef (FR-CAT-1a) to bytes is
Storage::open's job and happens at the caller, so the same decoder works over a local file, an
Android SAF document, or a byte range fetched from Nextcloud. A second implementation may be added
for broader camera coverage without changing callers (D2).
On the change of mechanism. This clause used to require "a trait taking a SourceRef". The
purpose — that no decoder API takes a path, so nothing in the decode path assumes a filesystem —
is met and is not in question: dr_decode::decode takes &[u8], and there is no path-based entry
point in the crate. The mechanism was wrong, and stating it that way would have made the design
worse.
A SourceRef is opaque by construction; the only thing that turns one into readable bytes is
Storage, in platform/dr-plat. A decoder taking a SourceRef would therefore have to take a
Storage alongside it, which moves retry, permission loss and remote fetching inside the decoder
and leaves it constructible only where a Storage exists. Bytes in, image out, is both narrower and
more portable: the decoder has no idea where its input came from, which is the property this
requirement is actually asking for.
It also serves the Nextcloud case better rather than worse, which is the one a byte-oriented API
looks like it would lose. dr_decode::HEADER_BYTES declares how much of a file the decoder needs to
read metadata, and import.rs fetches exactly that range through Storage::read_range before
calling dr_decode::metadata. The decoder states its requirement and the storage layer satisfies
it; a decoder holding its own SourceRef would have had to implement the range policy itself.
Status (2026-09-24). The trait is built: dr_decode::Decoder, over bytes — header_bytes,
metadata, orientation, locate_preview, preview and decode — with dr_decode::Rawler as
its one implementation, delegating to the free functions that were there before. The catalog scan,
the thumbnail ladder, import, the viewer, export, merge and repairs take a &dyn Decoder; only the
places that start a job name dr_decode::default(). "Without changing callers" is tested by
dr-ui's decoder_seam tests, which hand a stub decoder for a container no real decoder reads to
the scan, the ladder and export, and fail if any of them reaches past the trait. Nothing in the
trait takes a path or a SourceRef. The second decoder itself (LibRaw, D2) is not built; S7 (#48)
is what would say when it is needed.
FR-RAW-3 — Sensor data handling. Correctly apply per-camera black/white levels, CFA pattern identification, and camera-native colour matrices. Demosaic quality shall be selectable, with at least a fast method for preview and a high-quality method for export (FR-EXP-9 requires export to use the latter).
Status (2026-09-27), defective photosites. Hot and dead photosites are repaired on the mosaic,
before the demosaic, where each is still one wrong value rather than a coloured cross three pixels
wide (core/dr-gpu/src/shaders/hot_pixels.wgsl). A photosite is repaired only where it stands apart
from every same-colour photosite in its 5×5 window and from each of its eight immediate neighbours,
which leaves stars and glints alone, and it takes the value of its brightest (or, when dead,
darkest) same-colour neighbour, so nothing is invented. A 6×6 sensor-anchored colour tile serves
Bayer and X-Trans alike; export and every other path that demosaics get the repair, and there is no
setting. core/dr-gpu/tests/hot_pixels.rs renders a frame with and without a defect and compares
the finished pixels. The DNG defect map dr_decode::defects reads is not used: few files carry
one, and a CR2 none.
FR-RAW-4 — Robustness. A malformed or hostile RAW file shall not crash the application or compromise the process. Decode failures are reported per-file and do not abort a batch.
FR-RAW-5 — X-Trans as a first-class path. Per D11, Fujifilm is explicitly targeted:
- Markesteijn-class demosaic as the default for X-Trans sensors, not an opt-in advanced setting
- X-Trans-aware sharpening, since the non-Bayer CFA responds differently
- In-RAF film simulation tag read and matched (FR-DEV-3f)
This targets the market's best-documented colour grievance. Adobe's X-Trans "worms" artefact is a decade-old unresolved complaint; darktable and RawTherapee have the better algorithm but poor defaults; Capture One has the best film-simulation support but drops X-Trans I and II.
Cost to note: X-Trans demosaic is documented at at least 2× the processing cost of Bayer, which affects the NFR-P4 and NFR-P7 budgets for Fuji files specifically.
3.3 Develop pipeline
FR-DEV-1 — Non-destructive edit graph. All edits are stored as parameters in an ordered edit graph attached to the image. Source files are never modified. Any rendered output is reproducible from source + graph.
FR-DEV-2 — Internal precision. The pipeline operates internally at a minimum of 16-bit float per channel in a wide-gamut linear working space. Quantisation to the output bit depth happens once, at the final export or display stage.
Amended 2026-09-27 (D19): quantisation is one of three things deferred to the end, not the only one. Until the view transform (FR-DEV-3j), values are scene-linear and unbounded: nothing clamps above 1.0, nothing applies a transfer function, and nothing maps to a display gamut. Every operation between the camera matrix and the view transform receives and returns that. The working space's primaries are linear Rec.709, carried unbounded, so a colour outside sRGB is a negative component rather than a clipped one. That is wide-gamut in range, not in the primaries the operations measure hue against; moving the primaries to Rec.2020 is deferred (D19).
Acceptance: every point operation at non-neutral settings, handed a ramp to 16.0, returns values
that are still monotone in the ramp and still above 1.0 where the ramp is, before the view
transform (scene_referred_until_the_view, core/dr-gpu/tests/scene_referred.rs, rendered on a
device: dr-pipeline has none). The view transform itself and film simulation are excluded,
because clipping into a display range is their job, and so is the detail stage, which a flat frame
cannot exercise.
FR-DEV-3 — Adjustment set (v1).
- White balance (temperature/tint, and picker)
- Exposure, contrast
- Highlights / shadows / whites / blacks recovery
- Tone curve (RGB and per-channel)
- HSL / colour mixer per colour band
- Vibrance and saturation
- Texture / clarity
- Sharpening and noise reduction (luminance and chroma)
- Lens corrections: distortion, chromatic aberration, vignetting — driven by the lens profile the file's own EXIF matches in the bundled Lensfun database where there is one, and by hand where there is not. The profile is a control rather than a silent step: a photograph whose lens is recognised carries one switch that accepts or declines the measurement, and the manual sliders trim whatever it leaves. A photograph whose lens is unrecognised is told so in words and offered no switch, because a correction that looks available and does nothing is worse than one that is visibly unavailable
- Crop, straighten, rotate, flip
- Local adjustments: linear gradient, radial gradient, and brush masks
Resolved 2026-09-26: a mask layer's settings are offsets to the photograph's, applied at each
operation's own place in the chain — global contrast −30 under a layer at −20 is −50 inside the
mask, applied once. A moved switch or choice replaces the global one, and an offset that brings an
operation back to neutral undoes the global setting inside the mask. Layers used to run as a second
chain after every global operation, which compounded the two edits in ways neither slider showed
(architecture.md §5.2; core/dr-gpu/tests/local_adjustments.rs).
FR-DEV-3a — Self-describing operations. Every processing operation shall declare its own parameters through a descriptor, so that adding an operation requires no changes to frontend code. An operation declares what its parameters are; the frontend decides how to present them.
pub trait Operation: Send + Sync {
/// Static description of this op's parameters. Drives UI generation.
fn descriptor() -> OpDescriptor where Self: Sized;
/// Parameter values → GPU work. No UI types cross this boundary.
fn encode(&self, enc: &mut ComputeEncoder, ctx: &TileContext);
/// Identity for cache invalidation (see ARCH §3.4, ARCH §6.13).
fn params_hash(&self) -> u64;
}
pub struct ParamDescriptor {
pub id: ParamId,
pub label: LocalizedString,
pub kind: ParamKind,
pub default: ParamValue,
pub affects: Affects, // Geometry | Colour | Detail — drives invalidation scope
}
pub enum ParamKind {
/// Ordinary numeric parameter. Frontend picks slider / drag-strip / dial by modality.
Scalar { min: f32, max: f32, scale: Scale, unit: Unit, precision: u8 },
Bool,
Enum { variants: Vec<(EnumId, LocalizedString)> },
Colour { has_alpha: bool },
/// Escape hatch: a control that does not reduce to a primitive.
/// The frontend owns the implementation; the op only names the kind
/// and defines the data it exchanges.
Custom { widget: WidgetKind, data: CustomParamSchema },
}
pub enum WidgetKind {
ToneCurve, // per-channel curve editor
ColourWheel, // colour grading wheels
CropOverlay, // on-canvas crop and straighten handles
GradientHandle, // on-canvas linear/radial mask placement
BrushMask, // on-canvas brush strokes
WhiteBalancePick, // eyedropper bound to canvas
}
The pipeline crate shall not depend on the UI toolkit. Descriptors carry data, never widgets. This keeps the edit chain testable headless (see §8's golden-image tests, which must link no UI) and is what allows one operation to render differently on touch and desktop.
FR-DEV-3b — Frontend presentation mapping. The frontend maps ParamKind to a concrete control
based on input modality and available space. The same descriptor yields different presentations:
ParamKind |
Desktop | Touch (tablet) |
|---|---|---|
Scalar |
Slider with numeric entry, scroll-wheel fine adjust | Large drag-strip, double-tap to reset, no keyboard entry |
Bool |
Checkbox | Switch, minimum 44pt target |
Enum |
Dropdown | Segmented control or sheet |
Colour |
Swatch opening a picker popover | Swatch opening a full-width sheet |
Custom |
Frontend-supplied control for that WidgetKind |
Same control, touch-tuned hit targets |
FR-DEV-3c — Operation registry. Operations register themselves at startup. The develop panel is generated by walking the registry, so a new operation appears in the UI without any frontend change. Registration order defines default pipeline order; the ordering itself is data, not code.
FR-DEV-3d — Invalidation scope. Each parameter declares what it affects, so a change
invalidates only the necessary part of the pipeline. Adjusting exposure shall not re-run lens
correction or re-tile geometry. This is what makes FR-DSP-3's one-frame slider response achievable.
FR-DEV-3e — Camera input profiles. The pipeline shall include a camera-profile stage between demosaic and the working-space conversion.
v1 scope (per D11 — good defaults rather than exhaustive colour science):
- Embedded DNG
ColorMatrix1/2andForwardMatrix1/2tags A hand-tuned base curve per launch camera body, shipped with the app— retired 2026-09-27 (D19). The tone half of "the camera's look" is the view transform's (FR-DEV-3j), one for every body and adjustable. The colour half stays here, in the matrix and later the DCP.- HaldCLUT import (FR-DEV-3f)
The camera profile ends at the matrix, and the matrix runs first: white balance is applied in
camera RGB, where its multipliers are defined, and every other operation receives working-space
colour. Before D19 the edits ran in camera RGB and the matrix came after them, so a hue in the
colour mixer and the weights in luminance() meant something different on every body.
Deferred but not foreclosed: full Amended 2026-10-02 (D20): dual-illuminant interpolation of the matrices
was built with item 1. The tables follow, designed in camera-profiles.md:.dcp support with HueSatDeltas, ProfileLookTable, and
dual-illuminant interpolation. The stage shall be structured so these are additions rather than a
pipeline reordering.
- DCP tables.
ProfileHueSatMap(both illuminants, blended as the matrices are) andProfileLookTable, read from the profile embedded in a DNG or from a.dcpfile in the profiles directory matched byUniqueCameraModel, the embedded one first. They are applied by acamera_profilescene operation after exposure, with a switch and a look strength (0–200 %), on by default where a profile exists.ProfileToneCurveis read and not applied: tone is the view transform's (D19). An embedded profile whoseProfileEmbedPolicyallows copying can be saved as a.dcpfor other files from the same body. The application ships no profile.
Rationale for the reduced scope: a bare 3×3 matrix produces the flat, poor-skin-tone rendering
characteristic of dcraw defaults, which is the documented reason people abandon darktable in the
first hour. A per-body base curve fixes most of that at a fraction of the cost of a full DCP
implementation. The flat render is a missing view transform, not a missing per-body curve:
darktable's own answer to the first-hour complaint was a scene-referred default, and Ansel's is
the same. The per-body curves this clause shipped described themselves as hand-tuned shapes, not
measurements, and their provenance was not known well enough to keep them as defaults (D19).
Acceptance: the default render is subjectively comparable to the camera's own JPEG — through
FR-DEV-3j's default, for every body. ΔE2000 validation against ColorChecker references applies
once DCP support lands.
For item 4: the lookup follows the DNG SDK's on [0, 1] and leaves values above 1.0 above it;
grey and an identity table pass through unchanged; the shader agrees with the CPU reference; with
the switch off the render is to the bit the one with no profile (camera-profiles.md §8).
FR-DEV-3f — Look emulation. Support HaldCLUT import, which inherits the existing free film simulation ecosystem at near-zero implementation cost, plus reading the in-RAF film simulation tag to auto-apply a matching render for Fujifilm files.
Spectral film simulation, in addition rather than instead (dr-film). Where a stock's
measurements exist, simulate the physics instead of replaying a grade: spectral sensitivity exposes
three emulsion layers, characteristic curves develop them to densities, dye densities absorb, and a
paper profile prints the negative with the enlarger's filtration solved rather than dialled. A
scanned negative is therefore orange and inverted, because that is what a negative is.
Two things this buys that a LUT cannot. The parameters stay physical — opening up a stop moves the picture along the film's own curve, shoulder and all, rather than scaling a number baked at one exposure. And the data cost inverts: a stock is ~17 kB of published measurements where one HaldCLUT is ~800 kB of one person's grade.
A film simulation is a rendering, not an adjustment, so it is the view transform when a stock is chosen (FR-DEV-3j): it runs last, after every adjustment and after the detail stage, in place of the default sigmoid, and never in addition to it. Amended 2026-09-27 (D19): it ran at order 25 before this, after exposure and before everything else, so the edits below it acted on the film's output. They now act on the scene the film is shown: an edit is a decision about the exposure the negative receives, and the film is the last thing that happens to the picture. Existing edits that combine a stock with tone or colour operations render differently.
Acceptance: a neutral scene printed through a colour negative's own paper renders neutral to within 0.06 in linear sRGB; the baked lookup's interpolation error stays under one 8-bit code value; and the shader agrees with the CPU model, which agrees in turn with an independent reference implementation.
Resolved 2026-09-19: the chosen stock persists by id in its own sidecar field, not as a
parameter — core/dr-pipeline/src/sidecar.rs records why an index was rejected (installing a
profile would silently change which film every existing photograph was developed on).
Resolved 2026-09-26: the film's settings — exposure, push, print exposure, format — are
per-pixel, evaluated by the shader against tables that hold none of them, so a mask layer can
hold its own. A layer's settings are offsets to the photograph's, and where layers overlap a pixel
takes the weighted average of what each asks for, the photograph's setting taking whatever weight
the layers leave (operation::local_settings_block). The stock and its paper stay
photograph-wide: a layer has no picker. The print is split at the paper's log exposure, so print
exposure is an addition between two lookups and exact at any setting; push interpolates the
stock's measured processes. Before this a layer offered the film's sliders and they moved nothing.
FR-DEV-3g — AI denoise. Learned denoising operating in the raw domain, ideally jointly with demosaic.
Promoted into v1 scope per D11. The reasoning: unlike AI masking, denoise has no manual fallback — it reaches a quality ceiling no conventional method matches, which is why photographers run a second application for it. Raw-domain joint demosaic-and-denoise is also markedly easier to build into a new pipeline than to retrofit, and the same component attacks the X-Trans artefact problem (FR-RAW-5).
Inference is local only — no cloud, no telemetry (NFR-SEC-4). The stage is optional at runtime and its absence degrades gracefully.
FR-DEV-3h — Stored orientation is honoured, not edited. An image shall be shown the way the
photograph was taken, from its EXIF orientation tag (0x0112), everywhere it appears: the grid's
thumbnails, the develop canvas, and the read-only preview shown when no decoder can open the file.
The tag shall be applied as a property of reading the file, at the same standing as a RAW's masked-photosite crop (FR-RAW-3) — never as an edit. Concretely:
- Opening a frame the camera stored sideways shall not mark it modified, shall not enable the framing reset, and shall write nothing to its sidecar.
- "Reset framing" shall return the image to upright, not to the sensor's scan order.
- A sidecar shall never carry the orientation. Edits are shared between devices and bodies (FR-NC-9); one camera's sensor scan must not be applied to another's file.
- A user's own quarter turns compose on top of it, so one press of the rotate button moves the image by 90° whatever the file's baseline.
A file carrying no tag, or a value outside 1..=8, is displayed as stored. Guessing would turn a missing tag into a visibly wrong image, and most files have no tag.
Acceptance: a portrait frame from a phone or a body held sideways appears upright in the grid and in develop with no user action, and its sidecar is byte-identical to that of the same frame shot in landscape.
FR-DEV-3i — Model-found masks. A mask layer's part may be a thing a model found in the
photograph, selected by pointing at it: one subject — this dog, not that one — from an
instance model, or one category — all the sky, all the foliage — from a semantic model. The
selection is stored as identity (which run, which instance or which category name, with the
run's signature) so that two devices selecting the same subject hold the same value and merge per
field under FR-NC-9, and a layer whose run no longer matches reads as stale rather than
silently masking something else. The mask arrives approximately right and soft, and it is then an
ordinary layer: FR-DEV-3's edge treatment shapes it, FR-DEV-19b's strokes correct it, FR-DEV-19a
composes it with a gradient or a range, and FR-DEV-19c shows it. Inference is local
(NFR-SEC-4); the models and their grants are named in models/LICENCE.md and D14.
Added 2026-09-19, because it was built and §7 still said it was deferred. The deferral treated
subject masking as an AI feature to be copied from Adobe or darktable; what was built is the
shape §7's own note asked for — a mask that behaves like a hand-drawn one — and it has been the
primary way a local adjustment is made since the watershed hierarchy failed on real photographs
(docs/segmentation.md §15).
One departure from FR-DEV-19 is recorded rather than hidden. A model-found part's coverage is also written to the sidecar, run-length coded beside the layer, because a stored subject that is "reproducible by running the model again" renders as nothing on a device or in a batch export that never runs a model. It is a materialisation of the identity, not the edit: it takes no part in equality or merge, and the identity remains what the part means.
FR-DEV-3j — View transform. The last stage of the develop pipeline maps scene-linear colour to a display range, and it is the only stage that may. By default it is a log-logistic sigmoid applied per channel, with the middle channel's position between the other two restored afterwards so a hue survives the shoulder, and the result clipped only by the output transform. A stock chosen under FR-DEV-3f replaces it.
It is an operation with two parameters, persisted in the sidecar, adjustable in the develop panel, and held per mask layer like any other:
- Contrast — the sigmoid's slope. Default 1.4.
- White — how far above middle grey, in stops, the scene reaches display white. Default 4.0, so a highlight a stop past sensor saturation still rolls into white rather than clipping at it.
Scene middle grey is 0.13, where the retired default curve placed it (FR-DEV-3e), and it maps to display 0.18. A photograph with the view transform at its defaults is unedited: the operation is always composed, and "active" keeps meaning "moved from the defaults", so an untouched image writes no parameters and every other operation's neutral is still the image.
An already-rendered source — a JPEG — is not rendered again: the view transform is skipped for it, as the base curve was, so its two sliders do not move a JPEG. A film stock is not skipped, because choosing one is an edit.
Acceptance: monotone in each channel; a neutral stays neutral; middle grey lands within 0.01 of
0.18; between scene 0.03 and 1.0, the default is within 0.3 EV of the retired default curve; the
scene value 0.13 · 2^white reaches 1.0; and the shader agrees with the CPU reference.
Added 2026-09-27 (D19).
FR-DEV-4 — Ordered, GPU-resident execution. The pipeline executes as a sequence of GPU compute stages. Intermediate results remain in GPU memory between stages. Processed pixels shall reach the display without a CPU round-trip. (This is a hard architectural constraint — see ARCH §6.1.)
FR-DEV-5 — Edit history. Per-image undo/redo of edit operations, persisted with the catalog so history survives a restart. Named snapshots of an edit state.
FR-DEV-6 — Presets. Save, apply, and manage named presets covering a subset of the edit graph. Copy/paste settings between images. Batch-apply to a selection.
A preset may name a film stock, which travels with the film node's own parameters. The application ships a read-only collection of presets, updated with each release and never written into the user's library; a user preset saved under a shipped name overrides it until deleted or renamed. Shipped and imported presets change only the operations they name, so a look applied to a corrected photograph keeps the correction; a copy or a saved edit replaces everything in scope.
The user's presets sync between devices through each library they open, as one file beside the camera profiles. Each name merges on its own against what the last exchange left both sides holding, so presets added on two devices both survive, a deletion on one reaches the other rather than being restored by it, and an edit outlives a deletion made elsewhere. The write is conditional on the server's copy, so two devices exchanging at once cannot save over each other.
FR-DEV-7 — Before/after. Compare current edit state against the unedited original or against a chosen history state.
FR-DEV-8 — Spot removal. Non-destructive clone and heal spots stored as parameters in the edit graph (target, radius, feather, source offset, opacity, mode), with automatic source placement and manual override, plus a visualise-spots mode.
Sensor dust is unavoidable with interchangeable lenses, and dust spots are the most common reason a photographer leaves a RAW editor for a pixel editor mid-workflow. This is not the layer-based pixel editing excluded by §1.3 — it is a standard parameterised develop operation, and the brush infrastructure required by FR-DEV-3's masks already covers most of the cost.
FR-DEV-12 — Colour grading by tonal range. Hue and strength applied independently to the shadows, the midtones and the highlights, plus a global cast over the whole frame. Ordinary parameters in the edit graph like any other adjustment, presented as colour wheels where the frontend implements them and as sliders where it does not.
FR-DEV-3's colour mixer already adjusts hues, and this is not a second copy of it. The mixer acts on the colours that are in the frame and can only turn what it finds; grading acts on a tonal range and puts colour where there was none. That is the difference between correcting a colour and choosing one: the mixer cannot warm an already-neutral highlight, and it cannot tone a monochrome conversion at all. Split toning — cool shadows against warm highlights — is the look this makes possible, it is the one every competing developer ships, and there is no way to reach it from the adjustments above.
FR-DEV-18 — Dehaze. Remove, or add, the atmospheric veil that distance puts between the camera and the subject, as a develop operation with a single symmetric amount. The transmission is estimated from the photograph rather than supplied by the photographer, and every length the estimate depends on is stored as a fraction of the frame, so that what is judged on screen is what lands in the exported file (FR-DSP-1). Haze is the one degradation the tone and colour controls cannot reach, because it is spatially varying: a black point that clears the mountains crushes the foreground, and a contrast curve that clears the mountains does the same. It belongs with texture and clarity in FR-DEV-3's adjustment set as a member of the compositional detail family — the operations whose radius is a property of the picture rather than of the sensor — and it runs in the same neighbourhood stage, for the same reason: it is defined by what the pixels around a pixel are doing.
FR-DEV-10 — Range masks. A mask layer may select by a band over the photograph's own values — luminance, or colour — with soft edges, stored as bounds and softness rather than as pixels and rasterised on the device in the existing mask pass. Every mask in FR-DEV-3 begins from a shape: painted, drawn, or found by a model. A range begins from a property, which is what a luminosity mask has always been and what makes a local adjustment blend without a halo — the boundary is the picture's own, at whatever detail the picture has. It is also the only route to selecting by skin tone, and the prerequisite for composing a selection from several criteria at once.
FR-DEV-19 — Mask editing. A mask layer's coverage shall be editable by hand after it is created: painted into, erased from, and built from more than one selection. Every edit is stored as geometry and parameters in the edit graph; no rasterised mask is written to a file and none exists in CPU memory (ARCH §5.4).
A model finds a subject in a second and, before this, the photographer could not then change it by a single pixel. The coverage arrives approximately right — stopping inside a shoulder, leaking into the hair — and FR-DEV-3's edge controls move the whole boundary, so no value of either fixes two errors that go opposite ways. Every competing developer has had a brush over its selections for a decade; a mask that cannot be corrected is the reason an edit leaves this application for a pixel editor.
FR-DEV-19a — Mask composition. A layer's mask is an ordered list of parts, each naming a source and how it joins the mask before it — added to it, taken out of it, or intersected with it, keeping only where both agree. A layer of one part is exactly the layer that existed before this, and reads and writes the same sidecar. The parts of a layer merge under FR-NC-9 as the layer does, and each carries its own edge treatment.
Intersection added 2026-09-24 (issue #9, mask-editing.md M2). Union and subtraction were built
first; the selections that most need composing — the sky that is also bright, the subject that is
also skin — are intersections, and spelling one as a subtraction needs a part that selects the
complement. Intersection is the product of the two coverages, which is the minimum wherever either
side is fully in or out. A sidecar written before it never names it; a build from before it reads
the word as a union, which keeps the part visible rather than dropping it.
FR-DEV-19b — Hand correction. A part may be painted, with add and erase strokes, at a radius, hardness and flow the photographer sets. Strokes are stored as normalised source coordinates and rasterised on the device, so a correction stays on what it was painted on through a crop, a zoom and an export at any size. A whole stroke is one step in the history.
FR-DEV-19c — Seeing the mask. Each layer's mask shall be drawable over the photograph, switched per layer by an eye on its row and drawn in a colour of that layer's own, so that several can be shown at once and told apart; the style — a tint, an alpha, or an outline — is one setting for all of them. What is shown is the finished mask — every part folded, with the layer's feather, falloff, morphology, invert and opacity applied — and it is shown for a layer that carries no adjustment yet, which is the state every mask is in for its first few seconds. No rendered output ever carries it: the reveal is a property of looking at an edit, not of the edit, so it reaches the pipeline through the composition that draws the canvas and through no other, and it does not travel in a sidecar.
Nobody can refine an edge they are not being shown. Before this the only thing drawn on the photograph was FR-DEV-3's region overlay — a false-coloured picture of what the model detected, which knows nothing of a layer's shaping and nothing at all about a gradient, a range or a stroke — so choosing a subject or a category produced a layer whose extent was invisible, and every control in FR-DEV-19a and FR-DEV-19b acted on something the photographer could not see.
FR-DEV-17 — A crop that orphans a mask says so. When a crop is let go — the end of the drag, not any frame of it — and it has left a mask layer's coverage entirely or mostly outside the frame, the photographer is told how many layers, and which, and offered to keep the crop or take it back. The crop is applied either way and is never refused; taking it back is the ordinary undo, so the crop and its warning are one step in the history. A crop that strands nothing, and a layer already outside the frame before the crop moved, produce no notice at all. Mask geometry is stored in source coordinates, so a tighter crop does not destroy a layer: it makes it invisible, and nothing announces it. That silent loss is the whole case for cropping last; a notice at the moment it happens turns it into a decision, and lets the crop stay where the histogram argues it belongs — early.
FR-DEV-16 — A keyboard vocabulary for develop. Every develop gesture that can be reached from the keyboard is bound, tagged beside its implementation, and generated into the gesture book (FR-UI-4): stepping through the folder, fit and 1:1, holding the original, undo and redo, copying and pasting settings, cycling the adjustment groups, resetting the control last moved, and showing or hiding the selected mask layer. None is keyboard-only — each has a pointer and a touch route (FR-DEV-3b) — and the book is regenerated from the tags, so it cannot describe a binding the application does not have. Editing rhythm depends on the hands staying put: reaching for a menu breaks the concentration of an edit the same way it breaks the pace of a cull. A binding nobody can discover is the same as no binding, which is why the generated book is part of the requirement rather than documentation of it.
Status (2026-09-24). Met. Every binding named above is bound in develop's key handler and tagged
beside it, and develop also answers zoom in, out, fit and 1:1, panning a magnified view, rating,
flagging and labelling the open photograph, nudging the control last moved (the framing sliders
included), turning a mask part's join, keeping a crop that hid a mask, and going back to the grid.
"Cannot describe a binding the application does not have" is now checked rather than argued:
traces gestures-check reads the chord every handler compares against (Keys.chord, in
ui/dr-ui/ui/keys.slint) and fails when a bound key has no tag in its place or a tagged key has no
handler (tools/traceability/src/keymap.rs).
FR-DEV-20 — Perspective correction. Framing shall include perspective correction — a vertical
and a horizontal keystone — applied before the crop. It runs in the coordinate chain after the
crop and the straightening and before the stored orientation and the lens warp, so that "vertical"
means vertical in the photograph as it is shown and the lens is still corrected on the full frame
it projected. It is part of framing and carries its Compose attribute, so a settings paste
withholds it by default for the same reason it withholds the crop. The correction never exposes
undefined area on its own; where it is combined with a straightening angle, the crop that avoids
the empty corners (max_inscribed_crop) accounts for it, and the crop is refitted when the gesture
ends.
Converging verticals are the commonest geometric fault in architectural and interior work, and
correcting them reframes the image — which is why the order relative to the crop is not a
preference, and why the correction belongs to composition and not to the colour operations.
3.4 Display and interaction
FR-DSP-1 — Proxy-resolution rendering. The develop view renders at the resolution actually required by the viewport, not the source resolution. A 60MP image displayed in a 2000px viewport processes approximately 2000px of data, not 60MP.
FR-DSP-2 — Tiled computation. The visible region is divided into tiles. Only tiles intersecting the viewport are computed. Panning computes only newly exposed tiles; already-valid tiles are reused.
Status, 2026-09-19: as written, unbuilt, and waiting on S6. The interactive path does not tile, and frame-budget.md argues it should not on the reference desktop: a fused pass over a 4K viewport costs 4.5 ms of a 16 ms budget, a tile cache could save at most that, and the one stage over budget is a convolution that tiling makes worse. That argument is a desktop measurement. The case this clause was written for — an image larger than the GPU memory of a mid-range Android device (NFR-RES-2, ARCH §6.2) — has not been measured, and S6 is the spike that measures it. Until it runs the clause stands, so that a decision is taken on a number from the device the clause is about rather than from the one it is not. If S6 finds the fused pass inside budget there too, FR-DSP-2 becomes a scheduling concern for export and thumbnailing as frame-budget.md proposes; if not, S6 names the stage to tile.
2026-09-27: the export half is built, for sources larger than one texture. A linear DNG past 8192 pixels (a stitched panorama, 22927×8966 in the case that prompted it) opens on a reduced copy, and a render finer than the copy samples a window of the full resolution through the fused shader's source-window uniforms. The export is drawn in 4096-pixel tiles grown by the detail chain's reach (
dr_pipeline::tiles,ComposedDetail::reach) and matches the untiled render to within one code value (core/dr-gpu/tests/source_window.rs). The interactive path is still one dispatch over the viewport: a zoomed canvas cuts one window and keeps it while the view stays inside, which is not the tile cache this clause describes.
FR-DSP-3 — Interactive latency. Moving a slider updates the visible region within one frame budget at proxy resolution. When a full-resolution result is needed it is computed asynchronously, and the proxy result remains on screen until it is ready.
FR-DSP-4 — Progressive refinement. During rapid interaction the app may render at reduced quality or resolution, refining to full quality when interaction settles. Refinement is visually smooth, not a jarring swap.
Status (2026-09-26). Met in 0.15.0. A gesture renders half-resolution drafts, and the sharp frame
lands once, 120 ms after the last movement rather than on a timer counted from the first
(ui/dr-ui/src/refine.rs). While a draft is up the histogram's display reading is dimmed, since it
describes the last settled frame; the last draft is kept over the sharp one and faded out in
150 ms, so the refinement is not a swap.
FR-DSP-5 — Zoom and pan. Fit, 1:1, and arbitrary zoom levels. At 1:1 and above, the pipeline operates on the visible crop at full source resolution.
FR-DSP-6 — Colour management. The display path is colour-managed via the output device profile. Where the platform and display support it, output at greater than 8 bits per channel and in a wide gamut. On Android this means using the wide-gamut display path where available.
FR-DSP-7 — Image evaluation. Provide a live histogram (luminance and per-channel, in the output colour space), highlight and shadow clipping indicators, and a pixel colour readout under the cursor or touch point.
These derive from a GPU-side reduction into a small buffer. Per-frame CPU readback of image data is prohibited — it would violate ARCH §6.1 on every frame, which is precisely the bottleneck darktable documents. Histogram computation shall not extend the FR-DSP-3 frame budget.
Without this a photographer cannot see what highlight recovery is actually doing, which makes the FR-DEV-3 adjustment set substantially less usable.
FR-DSP-8 — Per-display colour and scaling. The display transform is selected per the display currently showing the canvas, and updates when the window moves between displays. Fractional and mixed DPI scaling are handled without resampling artefacts in the canvas.
The profile-acquisition mechanism is stated per display server, with a defined fallback where Wayland provides no profile. On a multi-monitor desktop with differing profiles, showing wrong colours on the second display is a correctness defect, not a polish item.
3.5 Adaptive interface
DarkRoom ships one adaptive interface, not separate touch and desktop applications. A single Slint codebase reflows by available space and input modality, guaranteeing feature parity by construction. Phones are not a target (D15); the layout family spans tablet and desktop.
FR-UI-1 — Layout breakpoints. The interface adapts across at least two layout classes:
| Class | Typical | Develop layout |
|---|---|---|
| Compact | Canvas full-width; one collapsible panel at a time; filmstrip on demand | |
| Expanded | Tablet portrait and landscape, desktop | Filmstrip, canvas, and adjustment panel simultaneously |
Layout class is a function of window size, not device type — a narrow window on desktop uses the compact layout, and the transition is continuous rather than a mode switch.
The develop column's side — beside the canvas, or docked under it on a tall window — follows the
window's aspect, independent of the layout class (docs/ui-navigation.md D-N7). A placement, not a
mode.
Amended 2026-09-06. Tablet portrait is the expanded class, not the compact one: both orientations of the 12-inch tablet clear
EXPANDED_MIN_WIDTH, 820 logical pixels (D-N2). Compact fires only on a desktop window dragged narrow. What portrait needs is the dock, which is the aspect axis above and not a second class (D-N7).
Amended 2026-09-19. The expanded row has said "filmstrip" since the table was written, and the build had the photo roll on demand in both classes — the compact row applied everywhere. In the expanded class the roll is open by default and closable, and it is shown for a set of files named on the command line as much as for a library, because a set of photographs is a set. It carries the place FR-UI-8 remembers — how many, which one, its name, and what the filter is narrowing to — so that develop and the grid read as one interface with two views of the same set, not two screens joined by a button.
FR-UI-2 — Input modality. The interface detects and adapts to the active input method, which
is independent of layout class: a tablet may have a keyboard and pointer attached, and a desktop
may have a touchscreen. Modality affects control sizing and affordances (FR-DEV-3b), not layout —
save for the group selector, which moves from above the develop column into the tool rail under
touch (docs/ui-navigation.md D-N6). Switching input mid-session shall be handled without restart.
FR-UI-3 — Touch targets. Interactive controls present a minimum 44pt hit target when touch is the active modality. Hit targets may exceed the drawn control bounds.
FR-UI-4 — Gestures. The canvas supports pinch-zoom, two-finger pan, and double-tap to toggle fit/1:1. Gestures are additive: every gesture-driven action has a non-gesture equivalent, so no functionality is touch-only.
FR-UI-5 — Pointer and keyboard. Where a pointer is present: hover states, right-click context menus, and scroll-wheel adjustment on numeric controls. Keyboard shortcuts cover navigation, rating, and common adjustments. Neither is required for any operation to be reachable.
Amended 2026-09-19. "Rating" above is not qualified by view and was built as though it were: the judgement keys lived in the grid alone. They shall work in every view that shows a photograph, applying to the one on screen — in develop, the open photograph, never a selection left behind in the grid — with a pointer and touch equivalent in the same view, since a tablet has no number row. Judging in develop does not advance to the next frame: auto-advance belongs to FR-CULL-4's mode, where moving on is the point, and in develop the photographer is working on the frame in front of them. The current rating and flag are shown wherever they can be set, and on the roll's cells, so stepping along a set shows what has been judged.
Status (2026-09-24). The keyboard half and the amendment are met; the pointer half is not yet. Keys cover navigation, rating and common adjustments in the grid, develop, the collections sidebar, People and the export and copy sheets, each one in the gesture book and held to its handler by the gestures gate. Rating and flagging work in develop on the open photograph without advancing, with stars and Pick/Reject in the top bar, and the roll's cells carry each frame's flag and stars; pick and reject also gained a pointer and touch route in the grid (Flag in the selection bar), where they had been keyboard-only. Outstanding: scroll-wheel adjustment on numeric controls — the develop sliders do not take the wheel.
FR-UI-6 — Shared component library. Touch and desktop presentations are variants of shared components, not parallel implementations. A new operation (FR-DEV-3c) becomes usable on both without frontend work.
FR-UI-7 — On-canvas controls. Custom controls that operate on the canvas — crop handles,
gradient placement, brush strokes (WidgetKind in FR-DEV-3a) — size their interaction regions to
the active modality while rendering identically. A crop handle drawn at 8px may carry a 44pt touch
region.
FR-UI-8 — Resumed place. The application shall remember where the photographer was and return them to it. "Where" is the view (grid or develop), the scope (whole library, a collection, or the trash), the rating filter narrowing it, and the photograph on screen — the open one in develop, the first visible one in the grid. It shall be restored on launch, and preserved across every transition between screens within a session: leaving the grid for Settings, Import, People or the develop view and returning shall land where it was left, never at the top of the library.
A place is addressed by what every device agrees on. The photograph by its remote path, the
collection by its UUID; never by a grid ordinal, images.id or collections.id, all of which are
local to one catalog and to one ordering. Where the ordinal is needed it shall be computed through
the same ordering the grid draws with, so that a restored position names the photograph the
photographer actually left. Capture time is the permitted fallback when the photograph is gone.
A place travels. One record per library is exchanged through the same derived folder as the thumbnail shards and the catalog snapshot (FR-NC-3), so that a session begun on one device can be continued on another. The newer of two records wins; there is nothing to merge, because two devices cannot both be where the photographer is.
And it is always advisory. A place that cannot be read, names a collection this device has not merged, or points at a photograph that has since been deleted shall degrade to the nearest sensible position and never to an error, an empty view, or a refusal to start. A record arriving from another device shall not be applied once the photographer has begun working in this session — it is a handover, not an interruption.
3.6 Export
FR-EXP-1 — Formats. Export to JPEG, PNG, and TIFF (8 and 16-bit). Quality, chroma subsampling, and bit depth are configurable.
AVIF and JPEG XL are post-v1 and are not part of this requirement's acceptance. Either may still
be listed in the settings page before its encoder exists, on one condition: choosing it shall fail
with a typed error naming the format, never with a file. The test
every_offered_format_either_encodes_or_explains_itself walks every format the page offers and
enforces exactly that, so a format cannot be added to the picker and quietly reach an encoder that
does not handle it.
On splitting this requirement. It read as one undifferentiated list — "JPEG, PNG, TIFF (8 and 16-bit), and AVIF or JPEG XL" — which left it neither met nor unmet. Three formats, both TIFF depths, and the configurability clause are built, encode, embed their profile, and are tested; the fourth item is a deliberate deferral, and the encoders for it are the two with the least settled library support. Fused into one sentence, the only choices were a tag asserting something untrue or no tag at all, and the second is the worse of the two: it would have removed the register's record of four-fifths of a requirement that is finished. The deferral is now stated where it can be read as a decision rather than inferred from an error variant.
FR-EXP-2 — Colour space. Export in a selectable output colour space (sRGB, Display P3, Adobe RGB, ProPhoto), with the correct ICC profile embedded.
FR-EXP-3 — Output sizing. Export size shall be specifiable by any of the following modes:
| Mode | Behaviour |
|---|---|
| Original | Full source resolution after crop |
| Long edge | Specified pixels on the longer dimension; aspect preserved |
| Short edge | Specified pixels on the shorter dimension; aspect preserved |
| Width × Height (fit) | Scaled to fit within the box; aspect preserved; result may be smaller in one dimension |
| Width × Height (fill) | Scaled to cover the box and centre-cropped to exactly those dimensions |
| Percentage | Scaled by a factor of the source |
| Megapixels | Scaled so the result approximates a target pixel count |
| Print dimensions | Physical size (mm or inches) at a specified DPI, resolved to pixels |
Additional constraints:
- Upscaling is permitted but shall be off by default, with an explicit opt-in. Where disabled, a request larger than the source exports at source size rather than failing.
- DPI metadata is settable independently of pixel dimensions, for print workflows.
- File-size ceiling (JPEG/AVIF/JPEG XL): optionally target a maximum output size in KB/MB, with the encoder iterating quality to meet it. Useful for upload limits.
- Sizing operates on the cropped result, so the crop rectangle defines the aspect ratio unless a fill mode overrides it.
FR-EXP-4 — Resampling and output sharpening. Resizing uses a quality resampler (Lanczos or equivalent) operating on linear-light data at pipeline precision, before quantisation to the output bit depth. Output sharpening is selectable (none / screen / matte paper / glossy paper) and its strength scales with the resize factor, since downscaling softens.
FR-EXP-5 — Export presets. Named presets capture format, quality, colour space, sizing mode, sharpening, metadata policy, and destination. A preset is applicable to a single image or a batch. Multiple presets may be applied in one operation, producing several outputs per image — e.g. a full-size TIFF alongside a 2048px sRGB JPEG.
FR-EXP-6 — Naming and destination. Output filenames are generated from a template supporting at minimum: original filename, sequence number, capture date, export dimensions, and preset name. Collision policy (overwrite / skip / auto-increment) is configurable. The destination is an album (FR-EXP-10), whose folder is on this device or on the library's server (FR-NC-7) — never inside the library itself, where a scan would catalogue the exported files as photographs.
A folder is chosen by pointing, never by typing a path: the platform's own dialogue on the desktop (the XDG desktop portal on Linux, which reaches the user's files from inside the Flatpak; the common item dialogue on Windows), the Storage Access Framework's tree picker on Android, and an in-app browser for a folder on the server. Every one of them can make a new folder as well as open one. The same applies to the other folders the application asks for on the desktop — the library folder, an import's source and second copy, and presets brought across from Lightroom.
FR-EXP-7 — Batch export. Export a selection with one or more presets, running in the background with progress and cancellation. Uses all available cores and the GPU. A failure on one image is reported and does not abort the batch.
FR-EXP-8 — Metadata on export. Configurable EXIF/IPTC/XMP retention, including an option to strip GPS and personal metadata. Copyright and contact fields are settable per-preset.
FR-EXP-9 — Full-quality path. Export always uses the full-resolution, highest-quality pipeline regardless of what the display was showing — including the high-quality demosaic (FR-RAW-3), never the fast preview method.
FR-EXP-10 — Albums. An album is a named export destination listed in the library sidebar beneath the collections. Its folder holds only the exported files; the catalog records, for each file written there, the image it was rendered from, so that selecting the album shows the originals behind its files and re-exporting after an edit is one selection away. The export sheet chooses an album by name, and an export with no album chosen is refused and says so.
Albums sync between devices as collections do (FR-CAT-7): by uuid and revision, deletions as tombstones, exports as a set union keyed on the server's file id. A folder on the server syncs with the album; a folder on the device does not, because a path or a SAF grant on one device means nothing on another — an album made elsewhere arrives with no folder here until one is chosen. Deleting an album never deletes the files in its folder. The album tables are created on first use rather than by a schema migration, so a device on an older build keeps merging the rest of the catalog.
3.7 Nextcloud integration
Mechanics below are verified against Nextcloud 34 documentation and server/desktop-client source. Three findings shape this section and are recorded as constraints in ARCH §6.6–ARCH §6.8.
FR-NC-1 — Account setup. Connect via Login Flow v2: POST /index.php/login/v2 returns a
browser URL and a poll token; the app opens the URL in the system browser (never an embedded
webview) and polls POST /login/v2/poll until it returns an app password. The token is valid 20
minutes and the success response is returned exactly once. The app never sees the user's primary
password.
The User-Agent sent during the flow names the resulting app password in the user's security
settings, so it shall identify the device (e.g. DarkRoom (Linux desktop)), allowing per-device
revocation. Logout shall call DELETE /ocs/v2.php/core/apppassword to revoke cleanly.
Manual app passwords are supported as a fallback for unusual server configurations.
FR-NC-2 — Credential storage. Linux: Secret Service via libsecret. KDE exposes the same
interface through ksecretd since KF5.97, so one code path covers GNOME and KDE. Where no
secrets daemon is running, the app shall enter an explicit degraded mode rather than silently
storing credentials in plaintext.
Android: Keystore-backed encryption. Note EncryptedSharedPreferences is deprecated; the current
approach is DataStore for persistence with Tink for encryption and Keystore for key protection.
Keys must not require user authentication, or background sync will fail.
FR-NC-3 — Remote browsing without full download. The app shall display a remote library's thumbnails without transferring full RAW files. Two mechanisms, selected per-account by capability probe at setup:
- Server previews —
GET /core/preview?fileId=…where available.forceIcon=falseis mandatory: the default returns a generic mimetype icon when the server cannot render the file, which would otherwise be cached as though it were a thumbnail. Thenc:has-previewproperty in PROPFIND indicates per-file availability. - Range-based embedded preview extraction — the required fallback (see ARCH §6.7). Fetch the first
64–256KB via HTTP
Range, parse the container to locate the embedded JPEG preview, then fetch exactly that byte range. Typical cost 1–3MB versus 25–100MB for the full file.
FR-NC-4 — Change detection. Sync shall use recursive ETag pruning, matching the official desktop client's discovery algorithm:
PROPFIND Depth: 0on the sync root requestinggetetag. If unchanged from the stored value, nothing anywhere in the library has changed — sync completes in one request.- Where changed,
PROPFIND Depth: 1and recurse only into child folders whose ETag differs.
Cost is proportional to the changed subtree, not to library size. Depth: infinity shall not be
relied upon (frequently disabled or prohibitively expensive). ETags shall be normalised for quote
inconsistencies before comparison, or spurious full rescans result.
FR-NC-5 — Identity. The catalog shall key remote files on Nextcloud's oc:fileid, which is
stable across renames and moves, so that a server-side move is detected as a move rather than as a
delete plus a re-download of a 100MB file.
FR-NC-6 — Selective download. Downloads are on-demand and resumable, running in the background. RAW files are never bulk-synced by default. On Android, transfers respect unmetered-network and charging constraints.
Three storage tiers:
| Tier | Content | Policy |
|---|---|---|
| Metadata | Catalog rows, edit-graph sidecars | Always synced; kilobytes; sync even on metered connections |
| Previews | Embedded JPEGs or server previews | LRU-evicted, size-capped; what the grid browses |
| Full RAW | Source files | Explicit pin or on-demand open only |
FR-NC-6a — Cache rules. The user shall be able to pin a set of images at a chosen tier, with the set defined by a rule that the app re-evaluates as the catalog changes. Selectors shall include at minimum:
- Collection — "this trip is available offline"
- Folder, optionally recursive
- Date range, absolute or rolling ("the last 90 days", which moves with the clock)
- Rating, colour label, flag, or keyword — "every 5-star image, always"
- Boolean composition of the above
A rolling window shall stay current without user intervention. Where rules disagree about an image, the most generous tier wins.
Rationale: selective sync is only usable if the selection can be expressed as intent rather than enumerated by hand. "Keep this shoot and everything from the last three months" is a sentence a photographer will say; selecting four thousand files individually is not.
FR-NC-6b — Lazy eviction. An image that stops matching a rule is not deleted immediately; it becomes the first candidate for eviction when the cache cap (NFR-RES-4) or platform memory pressure (FR-PLAT-AND-5) actually requires space. Eviction order is unpinned originals by last use, then proxies, then thumbnails. Metadata and sidecars are never evicted — they are authoritative (ARCH §6.12) and small.
FR-NC-6c — Availability is visible. Every image shall carry a visible availability state: Original, Preview, Metadata only, or Offline.
- An operation requiring absent data shall say so, with the transfer size, before starting
- Export from a preview-only image is refused, not silently degraded
- A pinned set reports its true byte cost before the user commits
Rationale: Lightroom Classic syncs 2560px proxies while displaying the original's filename, extension, and size, so users do not know what they actually have. Sync failures of legibility are more damaging than failures of transport.
FR-NC-6d — Placeholder libraries. Where a library is a folder kept by a sync client in virtual-files mode, the app shall treat a placeholder as the photograph, not downloaded — never as a one-byte file and never as a missing one.
- A placeholder is catalogued under the photograph's own name, with an identity that does not change when it is downloaded
- Reading one yields a distinct, actionable error; it shall not be reported as absent, because the sidecar writer creates a new document when a sidecar is absent and would discard the existing one (FR-CAT-8)
- Its size is reported as unknown rather than as the stub's byte count
Where the client offers hydration, content may be fetched as a borrow: a file is returned to the state it was found in, so a pass releases what it downloaded and leaves alone what the user already had. Releasing means asking the client to dehydrate — never deleting, which inside a synced tree would propagate to the server and remove the photograph everywhere.
Hydration is whole-file and shall never serve browsing (ARCH §9.0 finding 3, §9.0a). It is for the originals tier and for passes the user has been quoted a cost on and has agreed to.
FR-NC-7 — Upload. Files above 5MB use chunked upload v2 against
/remote.php/dav/uploads/<userid>/: MKCOL to create the upload folder, PUT each chunk, then
MOVE the .file pseudo-entry to the destination. Chunks are 5MB–5GB and named 1–10000.
OC-Total-Length shall always be sent so quota is checked up front rather than at assembly time.
Upload folders expire after 24h of inactivity; the app shall persist upload state and either
resume or DELETE stranded uploads on startup.
Small files (sidecars) use bulk upload via POST /remote.php/dav/bulk with a
multipart/related body, allowing hundreds of edit-graph sidecars in a single request.
FR-NC-7a — Remote layout. FR-NC-7 says how bytes travel; this says where they land. An
uploaded original shall be placed under the account's library root in a directory expanded from a
date template, defaulting to {yyyy}/{yyyy}-{mm}-{dd} — one directory per year, one per capture
day beneath it. The template is configurable per account and uses the same token vocabulary as
export naming (FR-EXP-6), so {date} means the image's capture date and never today's; an
image with no readable capture time falls back to file mtime, and the fallback is visible in the
import report rather than silent.
Two consequences the mechanics do not give for free:
- The path is derived, not remembered. The same image uploaded twice from two devices must compute the same destination, so expansion is a pure function of capture metadata and the template — never of local library layout, which differs per device.
- Existing trees are not restructured. An image already present remotely stays where it is. The template governs placement on upload only; DarkRoom shall not move server-side files to conform, because the remote library is also reachable by other clients (§1.3).
Sidecars follow their image, not the template: they are named from oc:fileid (FR-NC-8) and live
beside the file they describe.
FR-NC-7b — Ingest to remote. Import (FR-CAT-10) and upload (FR-NC-7) compose, and the library an import targets is always the remote one. There is no local library for a card to land in: FR-NC-6 puts the library on the server, so an import has exactly one destination and the interface shall not offer a choice of another.
What lands on the device is a staging copy, in a location the app owns, and it is not a library: no view lists it, and it is removed once the server confirms the file. A staged file the server did not take remains queued, and a later import drains the queue. An import with no reachable server therefore succeeds and defers, rather than failing (FR-NC-10); an import with no account has nowhere to go at all and shall be refused.
The bytes shall reach the device before they reach the network. Streaming a card straight to the server would make a move-import erase a card against an in-flight upload, and would make importing impossible offline.
On erasing the card. A move-import shall delete from the card only those photographs the server has confirmed — not those merely written to the staging copy, since that copy is removed as soon as the upload succeeds and would otherwise be the only remaining copy. Anything that does not upload keeps its card copy. A card erased against an unfinished upload is unrecoverable, which is the one failure in this app with no undo.
Duplicate detection (FR-CAT-11) runs against the catalog before upload, so re-inserting a card that was already imported transfers nothing. Where the catalog cannot answer — a file already in the destination folder, locally or on the server — the folder's own listing shall answer instead, so a re-import neither duplicates nor renames what is already held.
FR-NC-8 — Edit metadata sync. Edit graphs sync bidirectionally as sidecars, one per image,
named deterministically from oc:fileid. Each sidecar carries a monotonic revision counter, a
per-device UUID, and a last-edit timestamp.
A sidecar holds a keyed set of Versions (FR-CAT-12), not a single edit graph — one image may carry several virtual copies, and a single-graph format could not represent them. Version identity is part of the sidecar schema, so conflict merge (FR-NC-9) operates per-version.
FR-NC-9 — Conflict handling. Sidecar updates use If-Match with the known ETag for optimistic
concurrency (If-None-Match: * for creates). On 412 Precondition Failed the app shall fetch the
remote sidecar and merge at the edit-graph node level — disjoint edits (e.g. a crop on one
device, an exposure change on the other) both survive; genuinely conflicting nodes resolve by
timestamp — then retry with the new ETag under a bounded retry count.
The app shall not replicate the desktop client's (conflicted copy) file behaviour. Sidecars
are structured data of a few KB; a read-merge-rewrite cycle is cheap and preserves user intent.
Only genuinely ambiguous merges surface to the UI. Source RAW files are write-once and shall never
generate a conflict.
FR-NC-10 — Offline-first. The app is fully functional offline against cached content. Local sidecar writes are atomic (temp file plus rename) and committed locally before any network round-trip, so editing never blocks on connectivity. Sync resumes automatically when connectivity returns.
FR-NC-11 — Initial catalog build. For first sync of a large remote library, the app may use
WebDAV SEARCH (RFC 5323) against /remote.php/dav/ filtered by mimetype and paginated via
d:limit/d:nresults, in preference to walking thousands of folders with PROPFIND.
FR-NC-12 — Backend independence. Sync shall be implemented against a backend interface. No
protocol detail specific to any one backend may appear outside its connector, and no layer above
the interface may name a connector — with the single exception of the registry that constructs them
(dr_ui::remote).
A trait over operations is not sufficient on its own, and the first release proved it: dr-ui
constructed the Nextcloud backend directly in seven files, an account was a server URL beside a
DAV user id, and the local cache directory was named after a hostname. Independence requires four
things — operations, declared capabilities, an account model with no server in it, and a
registration mechanism (ARCH §8.0, docs/storage.md).
Backends declare capabilities rather than conforming to a lowest common denominator, because the property that makes Nextcloud sync fast — directory ETags propagating up the tree, so an unchanged root proves an unchanged library — is not a general guarantee. An interface built to the common subset would force full enumeration on every sync (ARCH §8.1).
Where a capability is absent the app shall degrade visibly, not silently:
- Sync strategy in use is reportable to the user, so a slow backend is visibly slow
- Without byte-range reads, remote browsing cannot extract embedded previews; the app shall refuse full downloads for browsing on a metered connection and explain why
- Without conditional writes, sidecar conflict detection falls back to revision comparison, which narrows but does not close the race; this is surfaced as a reduced-safety mode
FR-NC-13 — Folder libraries. A library shall be openable as a plain directory — a local disk, a network mount, an external drive, or a folder another client already syncs — with no account, no server and no credential.
This is a requirement rather than a convenience for three reasons. It is what a photographer with an archive drive and no server actually has. It is the only route that works where no secrets daemon exists, which FR-NC-2 otherwise treats as a degraded mode. And a second connector is the only way to keep FR-NC-12 honest: an interface with one implementation cannot be shown to be an interface.
The folder connector shall declare its capabilities truthfully rather than flatteringly — in particular it shall not claim propagating directory ETags, because a POSIX directory's mtime describes its own entry list and nothing beneath it, and a backend that claimed otherwise would hide edits rather than merely run slowly (ARCH §8.4a).
3.8 Platform integration
Android
FR-PLAT-AND-1 — Storage access. Library access is obtained exclusively via the Storage Access
Framework: the user grants one or more document trees through ACTION_OPEN_DOCUMENT_TREE,
persisted with takePersistableUriPermission and enumerated via DocumentsContract.
The app shall not request MANAGE_EXTERNAL_STORAGE and shall not depend on
READ_MEDIA_IMAGES for RAW discovery. See ARCH §6.9 for why neither is viable.
FR-PLAT-AND-2 — Permission loss. Loss of a previously granted tree permission — revocation, reinstall, removed SD card — shall be detected and surfaced, marking affected images offline per FR-CAT-9 rather than deleting catalog rows.
FR-PLAT-AND-3 — Process lifecycle. An Android process may be killed at any moment. Edit state shall be durable such that process death loses at most the last uncommitted parameter change. On resume the app restores the develop session, its image, and its viewport.
FR-PLAT-AND-4 — Background execution. Long-running sync and export use the platform's managed background execution with the constraints in FR-NC-6, and a foreground service with notification for user-initiated exports. Behaviour under Doze and battery-saver is specified and tested.
FR-PLAT-AND-5 — Memory pressure. The app shall respond to onTrimMemory /
ComponentCallbacks2 by evicting caches per NFR-RES-1, in a stated eviction order (GPU tiles
first, then proxies, then thumbnails).
FR-PLAT-AND-6 — Intents. Register as a receiver for image view and share intents, and provide
share-out of exported results via FileProvider.
Linux
FR-PLAT-LIN-1 — Desktop integration. Follow the XDG Base Directory specification for config,
data, cache, and state. Register MIME associations for supported RAW types and ship a .desktop
entry.
FR-PLAT-LIN-2 — Display server. Support X11 and Wayland. Where Wayland's colour-management protocol is unavailable, FR-DSP-8's stated fallback applies.
FR-PLAT-LIN-3 — Sandboxed distribution. Where distributed as Flatpak, filesystem access uses portals and credential storage uses the Secret Service portal, both verified to satisfy FR-NC-2 and FR-CAT-1 within the sandbox.
Status (2026-09-26). Not verified, because no Flatpak has been built. Since 0.17.0 every folder the desktop asks for is chosen through the FileChooser portal (FR-EXP-6), so the filesystem half has the code it needs; whether the path the portal returns opens, persists and takes a sidecar inside the sandbox is unobserved (distribution.md §4). Credentials reach the session's secret daemon through a talk hole rather than the Secret Service portal, for the reason the manifest gives.
Windows
Specified in windows.md; a stated channel under NFR-COMPAT-2, not a v1 one.
FR-PLAT-WIN-1 — Known folders. Configuration under %APPDATA%\darkroom; data, cache and
state under %LOCALAPPDATA%\darkroom. Nothing under the profile root and nothing relative to the
working directory. The layout beneath those roots is the same as under XDG, so a library directory
moves between platforms unchanged.
FR-PLAT-WIN-2 — Installer. A per-user installer that needs no elevation, registers an uninstaller, and whose uninstaller removes what the installer wrote and nothing the application wrote. Models are installed beside the executable and found there last, after the user's own directories.
FR-PLAT-WIN-3 — Built from Linux. The Windows binary and its installer are produced by the Linux CI from the same commit as every other channel, with no Windows machine in the build. Verification on Windows is a release step, recorded per release, not a build step.
3.9 Culling
Per D11 this is the product's primary differentiator, not an incidental capability.
The opportunity, stated plainly. Photo Mechanic is fast because it displays the camera's embedded JPEG — no demosaic, no database, no import step. FastRawViewer is truthful because it shows a genuine raw histogram, raw-derived clipping, and focus peaking. No shipping tool combines both. The culler/editor split exists only because Lightroom's culling is slow — it is a workaround photographers tolerate, not a workflow they want. A tool that is genuinely Photo Mechanic-fast eliminates the handoff rather than improving it.
FR-CULL-1 — Instant display. Displaying the next image shall not wait on demosaic, catalog import, or full decode. The embedded JPEG preview is shown immediately; higher-quality renders replace it progressively.
Acceptance: next-image display within 50 ms of the input event, sustained across a 3,000-image folder, on both platforms. This is the single most important performance figure in the document — Lightroom's ~2 s stall is the entire reason a competing product category exists.
FR-CULL-2 — Preview ladder. Previews resolve through tiers, each falling through to the next:
- Embedded JPEG preview from the RAW container (instant)
- Cached proxy from a previous visit
- Background full decode, promoted when ready
Some cameras embed previews below sensor resolution, and some embed none. The app shall detect this per camera model and pre-emptively background-render where the embedded preview is insufficient, rather than showing the user a soft image and letting them discover it at zoom.
FR-CULL-3 — Raw-truth overlays. Culling decisions are made against raw data, not the embedded JPEG:
- Raw histogram — computed from sensor data, not the preview. The embedded JPEG's histogram misrepresents available highlight headroom.
- Raw-derived clipping indicators — a JPEG's clipping warnings systematically lie about what is recoverable in the raw.
- Focus peaking — overlays in-focus regions on the preview, so focus is verifiable without zooming to 100%. This removes the largest single source of culling latency from the critical path.
Per D11, zoom-to-100% remains available for certainty; peaking makes it optional rather than mandatory.
FR-CULL-4 — Culling mode. A dedicated full-screen mode with:
- Auto-advance on judgement — users independently reinvent this in darktable, Lightroom desktop, and Lightroom mobile, which is strong evidence it should be the default rather than an option
- One-key reject, plus the full rating, flag, and colour-label axes
- Filter to unjudged, so a session resumes where it stopped
- Keyboard-driven on desktop; single-thumb reachable on tablet
- Evidence beside the frame (added 2026-09-19) — FR-CULL-13's signals, readable at a glance without leaving the mode or opening the frame
- The last import as a scope (added 2026-09-19) — the set a culling session most often starts from, reachable in one step and counted, expressed as a term of the selector language (ARCH §9.2) rather than as a special collection
FR-CULL-5 — Burst and near-duplicate grouping. Group frames by capture-time proximity and image similarity, allowing a burst to collapse to one representative and be judged as a unit.
This is the one automated capability photographers consistently praise, precisely because it is a mechanical grouping problem rather than a taste judgement. Automated selection is distrusted — the documented failure is rejecting the only frame of an important moment because someone blinked.
FR-CULL-6 — Comparison. Side-by-side and survey comparison of a selection, with synchronised zoom and pan, for choosing among near-identical frames.
FR-CULL-7 — Tablet culling. Culling shall be fully usable on tablet, as the validated multi-device workflow (§3.5, D11). Requires only ratings and small proxies to sync, not full originals — a substantially smaller sync problem than develop parity.
Design note: pinch-zoom accidentally triggering ratings is a documented defect in Lightroom mobile. Gesture and rating targets must not overlap.
FR-CULL-13 — Evidence, never verdicts. Added 2026-09-19; numbered after the people clauses because it governs them too. Everything the app computes about a frame in aid of culling is evidence: shown beside the frame, filterable and sortable through the selector language (ARCH §9.2), and — within a burst — permitted to propose which frame represents the group. The signals are the raw histogram, clipping and focus (FR-CULL-3), burst membership (FR-CULL-5), and per-face eye state and head pose (FR-CULL-8a). The list grows by adding to it, never by adding a second kind of thing.
What evidence may not do: change a rating, a flag, a colour label, or trash membership. There shall be no code path from a signal to a judgement write without a user action between them, and a proposal — a representative, "three of these have eyes closed" — is accepted by a press, never by a timeout or a default. This restates FR-CULL-5's ground rather than extending it: automated selection is distrusted because its documented failure is rejecting the only frame of a moment because someone blinked, and the remedy is not a better classifier, it is that the classifier does not hold the pen.
Evidence says what it measured. A focus figure names its region; a per-face count names the faces; a signal that could not be computed — no faces found, the original not on this device — is shown as absent, never as zero.
Acceptance: a test enumerates every write to the rating and flag axes and shows each reachable only from an input event. The evidence for a frame is visible in FR-CULL-4's mode and on its grid cell without opening it, and a filter on any one signal returns exactly the set whose chips show it.
Status (2026-09-26). The write-path half is met; the evidence half is not. cargo test -p traceability enumerates every write of a rating, flag, colour label or trash membership in the
shipped code — the catalog setters, any SQL that assigns those columns, the sidecar's judgement
amendment and the fields that carry one — and holds each to a hand-written list
(tools/traceability/src/verdicts.rs). Each listed write is a key, click or tap (a Slint on_*
callback, checked structurally), a function writing for its caller, whose callers are then checked
in turn, or a verdict carried from elsewhere: a sidecar or .xmp pull, the sync merge of two
devices' sidecars, the catalog mirrored out to a file, and duplicates consolidation, which moves
the copies' own verdicts onto the survivor and invents none. A new writer, including one in an
evidence producer, fails the test until it is listed with a reason; traces verdicts prints the
list. Choosing a burst's representative writes the grouping, not a verdict, and takes a press.
Outstanding: evidence chips (clipping, focus, burst membership, face counts) on the grid cell and in
FR-CULL-4's mode, shown as absent rather than zero, and a filter per signal. Eye state alone is
shown today, on People's face cells and as a library filter.
3.9.1 People
Face recognition was deferred in §7 through the 2026-08-08 calibration. It is undeferred here in a narrower form, and the narrowing is the point.
What changed. The deferral treated "face recognition" as an AI feature adjacent to subject masking. It is not the same kind of thing. Masking is a taste operation applied to one image; grouping photographs by who is in them is a mechanical grouping problem over the whole library, which is the category FR-CULL-5 already commits to and already justifies: grouping is the automated capability photographers consistently praise, because it organises without deciding. Every argument FR-CULL-5 makes for burst grouping applies unchanged to people grouping. Answering "where are the frames with the bride in them" across a 4,000-image wedding is a culling operation, and culling is the differentiator.
What is deliberately not in scope, because it is the failure FR-CULL-5 names: no automated
selection. Nothing here rejects a frame, ranks a face, scores a smile, or detects a blink. The
feature produces a filter, never a judgement. The user's rating axes remain the only thing that
rejects a photograph.
Amended 2026-09-19. "Detects a blink" is struck. The exclusion was always of judgement, and a blink is a fact about a frame of the same kind as a clipped highlight: reporting it is evidence, acting on it is the failure. FR-CULL-8a detects eye state and head pose; FR-CULL-13 says what may be done with them — shown, filtered, sorted, proposed — and what may not. The sentence that survives is the one that matters: the user's rating axes remain the only thing that rejects a photograph.
FR-CULL-8 — Face detection. The app shall detect faces in library images as a background job, producing per-face a bounding box, five-point landmarks, a detector confidence, and a 512-dimension embedding.
The two resolutions are separate, and conflating them is the failure this clause exists to prevent. Detection and cropping have opposite resolution needs, and a single buffer cannot serve both well:
- Source. The image is rendered at native resolution through the full-quality path (FR-EXP-9's pipeline, including the high-quality demosaic of FR-RAW-3). This is the same render export uses and is deliberately not the FR-CULL-2 preview ladder.
- Detector input. That render is downscaled for the detector, which fixes its input at 640×640 regardless (faces.md §4.1). Detection gains nothing from more pixels than its own input, so the downscale is free accuracy-wise and is what keeps the pass affordable in CPU.
- Crop. Boxes and landmarks are mapped back to native coordinates, and the aligned crop is sampled from the native render — never from the downscale the detector saw.
- Embedding. The aligned crop is warped to 112×112 in one bilinear step (faces.md §5).
The crop is the reason. ArcFace receives a fixed 112×112 whatever it is given, so the only question
that matters is whether those 112 pixels are real pixels or interpolated ones. Sampling the crop
from a preview means a face occupying a small part of the frame is upsampled to reach the
embedder, and an upsampled crop yields a confident embedding of detail that was never there —
which does not fail loudly, it degrades clustering three stages later. Measured on the reference
library under the previous preview-tier implementation: 47% of all stored faces had been
upsampled to reach 112×112, with crop_px as low as 34.
This supersedes the previous rule that detection ran against the thumbnail or proxy tier and never a full decode. That rule was adopted for affordability and it bought exactly that, at a cost to crop quality that was not measured until the library was large. Affordability is now met by when the pass runs rather than by what it reads: it is background work, preempted by everything visible, and resumable per image.
On a remote library this needs the original, not FR-NC-3's byte-ranged preview — on the reference library, 412 GB across 19,107 images rather than a range request each. So a whole-library pass is a transfer under FR-NC-6: never automatic, subject to the unmetered-network and charging constraints, and reported as the download it is before it starts rather than presented as a local operation. An original already on the device is indexed from what is there. Nothing here requires the original to be kept: it is rendered, cropped, and given back under the same rules as any other borrowed file (ARCH §9.0a), so the pass costs transfer and time rather than permanent disk.
Detection is a job in the FR-CAT-3 queue and inherits its properties without exception: coalesced
per image, interruptible, resumable across process death (FR-PLAT-AND-3), and strictly preempted by
visible work (NFR-ARCH-2). A library indexes while idle or it does not index; it never competes with
the grid. face_index.source_edge records the native edge each run was made at, and faces.crop_px
the pixels behind each individual crop, so raising the standard
later re-indexes only the images that stand to gain rather than all of them.
Acceptance: indexing a 10k-image library completes without the grid dropping below NFR-P9's
interaction target at any point, and survives being killed and restarted with no repeated work
beyond the in-flight image. No face is stored whose aligned crop was upsampled beyond a stated
factor; the crop source resolution is recorded per face (crop_px) and is auditable.
FR-CULL-8a — Per-face state. Added 2026-09-19. For every detected face the app shall record, as derived data under FR-CULL-12:
- Eye state — one open-probability per eye, from a classifier over a crop around each eye landmark. Per face this reads as both open, one closed, or both closed; per frame, as a count.
- Head pose — yaw, pitch and roll, solved from the five landmarks against a generic face template. No model: this is geometry the detector has already paid for. Facing the camera is the pose within a stated band, and is the proxy for eye contact this document adopts — because gaze estimation has no redistributable weights (D13), and because in the photographs where it matters, groups and children and events, a turned head is the decision and averted eyes are often the better frame.
Both are terms in the selector language (ARCH §9.2, FR-CULL-11) — "everyone's eyes open", "facing the camera" — composable with a person and with every other term. Both feed FR-CULL-13's evidence and may propose a burst representative (FR-CULL-5); neither may set a rating or a flag.
The weights ship under the same test as every other model: redistributable under a licence compatible with GPLv3 and with Flatpak, F-Droid and Play, with the training data's terms read as well as the weights' (D13). The candidate identified 2026-09-19 is OCEC — MIT for code and weights, data under ODC-By 1.0 and Apache 2.0, six variants from 112 KB to 6.4 MB, a 24×40 crop per eye, sub-millisecond on CPU, opset 17 with batch-norm already folded. Its published F1 of 0.99 is on its own crops; ours are cut from a five-point landmark, so the number is measured on this library before it is believed.
Built 2026-09-19, in part (faces.md §17). Eye state ships: OCEC over a box cut
from the lid contour of InsightFace's 2d106det, plus a sunglasses classifier (SGC, MIT) whose
answer takes precedence over the eyes it hides, and — the part the measurement forced — per eye
the source pixels across the box and the sharpness of the patch, so an eye too small, too soft, or
hidden by the turn of the head is recorded as unreadable rather than read as closed. The filter
is the "Eyes open" chip on the library's people filter: with people chosen, it asks about their
faces, and it drops a frame only on a closed eye that could be read. Head pose is not built;
the hidden eye of a turned head is caught by its collapsed contour instead, and the selector-term
form ("everyone's eyes open", composable in a saved collection) waits on FR-CULL-11's person term,
which the grid filter also predates. The landmark model is under the InsightFace grant, accepted
on the same terms as the detector and embedder (faces.md §2.2a, decision 2026-09-19: this project
will not be commercial), so the weights test above is met by the two classifiers and not by the
third model.
Acceptance: on a labelled set of at least 500 faces from the reference library, eye state is within a stated tolerance of its published F1, reported separately for glasses, profile, and faces under 60 px; facing-the-camera agrees with a hand-labelled split at a stated rate. Both are recomputed by re-indexing, neither is written to a sidecar, and the indexing pass stays inside FR-CULL-8's acceptance.
FR-CULL-9 — Calibrated identity. Face similarity shall be expressed as a calibrated probability that two faces are the same person, not as a raw embedding distance. Every threshold in the subsystem — clustering, suggestion, auto-confirmation — shall be stated in that probability space, and no code path may threshold a bare cosine similarity.
This is a hard requirement rather than an implementation detail because the failure mode is invisible. A raw cosine means something different for every model, every population, and every face size; a threshold tuned on one library silently misbehaves on another, and an uncalibrated similarity still looks like a plausible number all the way to the user interface. A displayed confidence that does not mean what it says is worse than no confidence, because it is trusted.
The calibration shall be fitted per library from that library's own faces, and shall report whether it is valid. Where it is not — too few examples to fit — the app shall fall back to a documented, published operating point (the reference implementation's fitted curve) and shall say, at the screen level, that the confidences come from it. What is forbidden is presenting an untuned default as though it were measured on this library; withholding the number entirely is not required and shall not be done, because a screen of unranked suggestions is the state most libraries would permanently sit in — the fit needs confirmations, and confirmations need a ranked screen to be made on.
Acceptance: on a labelled corpus, the stated probability is within a documented tolerance of the observed match rate across the probability range (a reliability-diagram check, not a single accuracy figure).
FR-CULL-10 — Clustering and naming. Detected faces shall be clustered into unnamed groups. The user names a group, and that name applies to its members. A person is thereafter a first-class catalog entity with a stable UUID, independent of any name given to them.
The user shall be able to merge two groups that are the same person, split a group that is not, remove a face from a person, and rename a person, at any time and without re-indexing. Splitting must be as easy as merging: clustering will over-merge on siblings, on parents and children, and on the same person a decade apart, and a tool that can only merge makes its own errors permanent.
Confirmation is explicit. A face is either suggested (the system's inference) or confirmed (the user's judgement), and the two are never conflated in storage or in display. Suggestions may be recomputed freely; confirmations are user data and are never overwritten by a later inference pass.
FR-CULL-11 — People as a selector term. A person shall be a term in the selector language (ARCH §9.2), composable with every other term.
This is the requirement that pays for the subsystem, and it is nearly free once FR-CULL-10 exists: because one predicate language serves the library filter, smart collections, and cache rules, a person term yields all three at once — filter the grid to a person, save "every photo of Anna rated three or higher" as a smart collection, and pin "every photo of my children" to stay local on the tablet. The last is a genuinely new capability, not a restatement of the first two.
Selectors shall distinguish confirmed from suggested membership, defaulting to confirmed-only, so a saved collection does not silently change membership when a later indexing pass revises a guess.
FR-CULL-12 — Names are user data; embeddings are not. A confirmed person name is a user judgement of the same class as a rating or a keyword, and shall be written to the sidecar (FR-CAT-8), so it survives catalog deletion and travels with the photograph.
Embeddings, detections, cluster assignments, and unconfirmed suggestions are derived data. They live in the catalog only, are rebuildable by re-indexing, and are never written to a sidecar. This follows ARCH §6.12 exactly: the expensive-but-reproducible artefact stays in the disposable index, and only the irreplaceable human judgement enters the trust path.
The person UUID is what a cross-device merge keys on, in the same way collections merge (FR-CAT-7). Two devices that independently name the same cluster produce two people; merging them is the ordinary FR-CULL-10 merge, not a special case.
3.10 Extensibility and plugins
Post-v1, decided 2026-09-19. Every clause in this section, and NFR-SEC-6 which exists for it, is marked
(post-v1)on its defining line and is outside the v1 count. The section stays as the design of record — an extension point that is built as though a plugin might one day reach it is cheaper than one retrofitted — but nothing here is owed a tag, and §7's row is the statement of scope. The audit that prompted this found the register saying both things at once: §7 had deferred "Plugin API" in a bare row since the first draft while these 23 clauses counted against coverage, which measured the contradiction rather than the software. D16 defers with the section.
FR-DEV-3c already buys one form of this: an operation is added by writing a declaration, and the develop panel grows its controls without a frontend change. This section extends that property past the compiler. A plugin is a file dropped into a directory; the application uses it without being rebuilt.
The reason FR-DEV-3c was cheap is worth stating, because it decides everything below. An
Operation never touches a pixel. It publishes a descriptor and returns WGSL text and a handful of
floats, and the fused shader does the work (ARCH §3.3, ARCH §5.2). Any extension point that can be
given that shape — data in, data out, no execution — gets plugins for almost nothing. Any extension
point that cannot needs a sandbox to run code in, and that is where the entire cost of this section
lives.
So the classes are ordered by expense, and the rule is: an extension point is declarative unless it is demonstrably impossible to make it so.
FR-PLG-1 — Three plugin classes. (post-v1) The app shall support exactly three, and no fourth shall be introduced without a recorded decision.
| Class | Mechanism | Covers | Executes code |
|---|---|---|---|
| Declarative | YAML declaration plus WGSL | develop operations, mask generators, scope computation, camera profiles, LUTs, presets, localisations | No — shader source only, validated before use |
| View | Slint compiled at runtime into a declared slot | toolbar and panel widgets, the drawing half of a scope | On the UI thread; not sandboxed |
| Computational | WebAssembly component against a versioned interface | segmentation strategies, decoders, sync backends, metadata extractors | Yes, sandboxed |
The expected distribution is heavily weighted to the first. A plugin author reaches for class 2 only to draw something new and for class 3 only when an algorithm cannot be expressed as GPU passes.
Class 3 shall be WebAssembly and shall not be native shared objects. Three reasons, each independently sufficient: a native object cannot be loaded on Android (§1.2 names it a primary platform) or on any future iOS build; Rust has no stable ABI, so a native plugin would be locked to one compiler version and one build of every crate it touches; and an in-process native object holds the whole application's privileges, which would make NFR-SEC-4 a promise about third-party code rather than a property of the system.
FR-PLG-1a — Plugins orchestrate; the GPU does pixels. (post-v1) No plugin interface shall pass full-resolution pixel data across a sandbox boundary. A computational plugin receives handles and small buffers, and expresses per-pixel work as class-1 shader passes it declares. This is what keeps WebAssembly's arithmetic penalty irrelevant, and it is also ARCH §6.1 applied to plugins: a plugin must not be the reason a result round-trips through the CPU.
Declarative plugins
FR-PLG-2 — The node declaration is the plugin format. (post-v1) The schema documented in
core/dr-pipeline/ops/README.md — parameters, uniform expressions, WGSL, helpers, activity,
presentation, attributes, order, tests — shall be readable at load time as well as build time,
from a plugin directory, with no change to what a declaration means.
One consequence is deliberate: a bundled operation and a third-party plugin are the same kind of thing, differing only in where the file was found. There is no second, weaker format for outsiders, and no path by which the bundled operations acquire capabilities plugins cannot reach.
The build-time path is not removed. Operations that ship with the app remain compiled, because a
generated match is faster than an interpreted one and because their tests must run under
cargo test. The two paths shall produce descriptors indistinguishable to everything downstream, in
the same way and for the same reason that a declared node is today indistinguishable from a
hand-written one.
FR-PLG-2a — Fragment nodes and pass nodes. (post-v1) Two templates, and a plugin author chooses by answering one question: does this operation need to read a pixel other than its own?
| Fragment node | Pass node | |
|---|---|---|
| Declares | a WGSL body over c |
a complete compute shader with input and output textures |
| Cost | none — fuses into the existing dispatch | one full-resolution texture round-trip |
| Reaches | per-pixel colour transforms | neighbourhood operations — blur, clarity, halation, structured grain |
Fragment nodes shall continue to compose into a single dispatch (ARCH §5.2). A pass node breaks that fusion for itself only: the fragments before it and after it still fuse into one dispatch each. A declaration shall state which it is, and the cost shall be visible to the user in the plugin listing, because a chain of pass nodes is how a fast application becomes a slow one without any single decision having been wrong.
FR-PLG-2b — Mask generators are declarative. (post-v1) Linear and Radial mask sources are already
geometry in normalised coordinates rasterised by a shader (ARCH §5.4). A plugin shall be able to
contribute a mask generator on the same terms — declared parameters plus a WGSL function from
normalised coordinates to coverage — reaching luminosity-range, colour-range, and further gradient
forms with no code. Mask sources that are identity into a segmentation (Regions, Subject)
are not declarative and belong to class 3.
FR-PLG-2c — Scopes split at the existing seam. (post-v1) A scope plugin is a class-1 compute shader
producing a small bin buffer plus a class-2 view drawing it. This is the split already in place for
the histogram — dr-gpu counts, dr-ui shapes, Slint draws — and it holds for waveform,
vectorscope and RGB parade without change. The counting half shall not read back full-resolution
pixels (ARCH §5.5).
FR-PLG-2d — The vocabularies stay closed. (post-v1) WidgetKind, attributes, and the parameter kind
list remain closed enumerations, and a plugin may use them but shall not extend them. The reason
given in the node README strengthens rather than weakens here: a typo that creates a new category
containing exactly one control is indistinguishable from a deliberate new category until somebody
notices. An unrecognised attribute shall place the operation in a clearly-labelled fallback group
and warn — never silently omit it, which would make a control that does not exist look like a
control that was never written.
View plugins
FR-PLG-3 — Declared slots, typed contracts. (post-v1) A view plugin shall be a Slint component compiled at runtime and instantiated into a named slot the application declares — not a licence to draw anywhere in the window. Each slot states the data it provides and the callbacks it accepts, and a component that does not match its slot's contract shall be rejected at load with a message naming the mismatch.
Slots are a closed list under the same reasoning as FR-PLG-2d, and the composition rules of ARCH §4.3a continue to apply: a slot describes what a view is for, never how much room it has.
FR-PLG-3a — A view plugin cannot be trusted with the UI thread. (post-v1) Class 2 is the one class with no sandbox — an interpreted component runs on the UI executor and can violate NFR-ARCH-1 by looping. The application shall therefore watchdog slot rendering, disable a component that exceeds a stated budget, and report which plugin was disabled. A view plugin that fails shall leave the slot empty and the application usable; it shall never take down the window.
Computational plugins
FR-PLG-4 — One interface per extension point, versioned. (post-v1) Each class-3 extension point shall define an explicit interface, versioned independently, and a plugin shall declare which version it implements. Interfaces are the only surface a computational plugin can reach: there is no ambient filesystem, no network, and no access to the catalog.
FR-PLG-4a — Capabilities are granted, never assumed. (post-v1) A plugin that needs to read a file or reach the network shall declare the capability, and the user shall grant it explicitly with the reason shown. A plugin's declared capabilities shall be visible before installation, and a plugin that requests none — which is the expected case for a segmentation strategy — shall be installable without a security decision.
This is what makes NFR-SEC-4 hold under a plugin ecosystem. A segmentation plugin that cannot open a socket cannot send a photograph anywhere, and that is a structural property rather than a promise.
Versioning
FR-PLG-5 — Three independent version numbers. (post-v1) Conflating any two of these produces a wrong answer in both directions, so the format shall carry all three.
| Number | Owned by | Governs | Moves when |
|---|---|---|---|
| Interface version | the application | whether the plugin loads at all | the host changes what it offers |
| Plugin version | the author | updates, provenance, and the trust record | any release |
| Parameter schema version | the author | whether an existing sidecar still reads | parameters change incompatibly |
The common case is a plugin renaming a parameter: sidecars break while the interface version never moves. The converse also occurs. Binding sidecar compatibility to the interface version would be wrong in both.
FR-PLG-5a — Declared, never inferred. (post-v1) A plugin shall state its interface version explicitly. Deducing it from which keys are present produces files that are ambiguous between two versions, and the ambiguity surfaces years later as a wrong render rather than as an error.
FR-PLG-5b — A supported window, and a written policy. (post-v1) The application shall support the current interface version and at least one predecessor, and the deprecation policy shall be stated in the plugin authoring documentation rather than decided per release under pressure.
Additive changes shall not bump the version. A new optional key is compatible because absent
means default — the rule active:, presentation: and define: already follow. Holding that
discipline is what keeps the version number nearly stationary.
FR-PLG-5c — Adapt at the boundary, normalise inward. (post-v1) A plugin declaring an older interface version shall be adapted at load into the current internal representation, and nothing downstream shall be able to tell. Version branches threaded through the pipeline are how this becomes unmaintainable; the single adaptation point is the same discipline that lets a declared node and a hand-written one be one thing by the time anything reads them.
FR-PLG-6 — Migrations are data. (post-v1) Because every parameter is an f32 addressed by a flat
op_id.param_id key, a schema migration is a rewrite table rather than code. A plugin shall be able
to declare migrations between consecutive parameter schema versions, supporting at minimum rename,
rescale, and default-for-a-new-parameter.
The application applies them, chained, at load, so that everything downstream sees only current-schema parameters. Migrations shall be testable through the same declared-test mechanism as the node itself.
FR-PLG-6a — params_version is a promise, not a hint. (post-v1) Bumping it locks every older build out of
the edits that use it (FR-PLG-9). It shall be bumped only when an older build would genuinely
misread the file, and never merely because a parameter was added — absence already means default,
which already means neutral. This obligation belongs in the authoring documentation in as many words.
Sidecars
FR-PLG-7 — The sidecar records identity, not location. (post-v1) Each version block shall record, for every plugin it depends on: the plugin id, its version, its parameter schema version, and a content hash of the artefact. These merge key-wise like every other line in the format (FR-NC-9).
It shall not record an install URL. Sidecars arrive from elsewhere — they sync (FR-NC-9), they travel with shared photographs, they come from other people's catalogs. A sidecar that names where to fetch code lets whoever wrote it choose what the user is prompted to install, which is a confused deputy with a friendly dialog in front of it. Resolution from identity to location is a decision the user made when they configured a registry, not one an incoming file makes for them.
The content hash carries a second benefit: "install filmic 2.1" resolves to exactly the bytes the original edit was rendered with, which makes substitution and silent render drift detectable rather than merely regrettable.
FR-PLG-8 — A missing plugin shall never cost an edit. (post-v1) The sidecar already preserves lines it does not understand verbatim and writes them back untouched, so a machine lacking a plugin cannot destroy an edit that uses it. That property is now load-bearing and shall be treated as such.
Beyond preservation:
- Alert, aggregated. Missing plugins shall be reported once per import or session and listed in one catalog-wide view — never once per image. A folder of five hundred synced photographs sharing one missing plugin is one notice.
- Non-blocking. Nothing here is urgent, because the edit is safe either way. The image shows a clear mark that an operation is unavailable; the install dialog appears when the user asks for it.
- Never during unattended work. No such prompt shall interrupt a background sync or a batch export, where a dialog becomes either a stalled job or a reflex click.
- Render without, never render a guess. The image shall be rendered omitting the unavailable operation. Interpreting its parameters under different semantics is forbidden: a plausible wrong render and a correct one look equally plausible, which makes the wrong one the more dangerous output.
- Export is gated harder than display. Exporting an edit with an unavailable operation shall require explicit acknowledgement. A slightly wrong screen is recoverable; a delivered file that silently omits an adjustment is not.
Acceptance: a sidecar written with a plugin installed, opened and saved on a machine without it, is byte-identical to the original.
FR-PLG-9 — Forward skew is quarantined, not guessed. (post-v1) Where a version block names a parameter schema version newer than the installed plugin declares, the application shall divert that operation's parameters into the preserved-verbatim path before applying any of them, and mark the version quarantined.
This does not fall out of FR-PLG-8. A forward-skewed operation is a recognised id, so the existing unknown-key path never sees it: its parameters parse, apply under the older meaning, render plausibly, and are written back — silently downgrading the edit, on every device it syncs to. The hazard is the write-back, not the display.
Quarantine means, precisely:
- No apply, no edit, no save, no export for that version. These are the operations that lose data or ship a wrong file.
- The photograph still opens. Decoding needs no plugin, and refusing to display a file because one adjustment is from the future holds the photograph hostage over an edit.
- Per version, not per image. A sidecar holds several versions (FR-CAT-12); one may be quarantined while the others open normally.
- An upgrade is offered, resolved through FR-PLG-10 like any other install — and it shall be allowed to fail. Offline, declined, or delisted all fall back to quarantine, never to deletion and never to opening anyway.
- Bundled plugins say so. Where the plugin ships with the application, the remedy is an application update, and the message shall say that rather than offering an install that cannot help.
Acceptance: a sidecar written by a newer plugin version, opened, browsed, and closed on a build with an older one, is byte-identical afterwards.
Distribution
FR-PLG-10 — Registry resolution and verified artefacts. (post-v1) Installation shall resolve a plugin id through a registry the user has configured, with a default registry shipped. The application shall verify the artefact against the hash or signature the registry states before loading it.
- An unrecognised id is the loud case. Where no configured registry knows the plugin, the application shall say so and show what the sidecar claims, as text, for the user to act on deliberately. There shall be no one-click install of a location supplied by a file.
- Provenance appears in the prompt. Plugin name, version, registry, and publisher, so "from the registry you trust" and "from somewhere you have never heard of" do not look alike.
- Offline degrades, it does not block. With no network: say what is missing, keep the edit intact, open the photograph. An application whose privacy claim is local-only shall not need the internet to show a picture.
A release page on a code-hosting service is a distribution mechanism, not an identity. Binding identity to one would give dead links on a rename, no mirroring, no offline install, and a dependency on one company's availability.
Authoring and operations
FR-PLG-11 — Plugins are validated, and validation is the author's tool. (post-v1) A declared plugin's
tests: shall be runnable outside the application, against the same interpreter that loads it, via
a command-line validator. The application shall ship a scaffold command producing a minimal working
plugin of each class.
Load-time failures shall be collected per file and reported with the key that was wrong, in the
manner build.rs already reports them. A malformed plugin shall be skipped, never fatal. The
build script's exit-on-error discipline is right for an author with a compiler open and wrong for a
user opening their library.
FR-PLG-12 — Failure is attributable and revocable. (post-v1) The application shall record per-plugin timing and error counts, surface them in the plugin listing, and allow any plugin to be disabled without uninstalling it — including on the next launch after a crash, so a plugin that prevents startup can be disabled by someone who cannot start the application.
Acceptance: with a plugin deliberately made to fail at load, at render, and at UI paint, the application starts, opens a photograph, names the responsible plugin, and continues.
3.11 Merging images
Several photographs become one. Panorama is the first merge and the only one specified; HDR merge and focus stacking share its data model (D18) and are still deferred in §7. The clauses below are written for the panorama and, where a clause is general to any merge, say so.
FR-MRG-1 — Panorama from a selection. Two or more selected images are aligned and blended into one composite, which is written as a new source file per D18. The tool is never automatic: it proposes an alignment, the photographer sees it and confirms, and nothing is written before that press.
The stated audience (D11) shoots panoramas and currently leaves the application to stitch them, which is the workflow break FR-DEV-8 was added to close for dust. It is also the first of the three §7 merges, and the one whose alignment problem is smallest — a rotation about one point, with no depth to recover — so it is where the shared machinery is built.
FR-MRG-2 — What is stitched. Each source enters the merge in camera space: after black and white levels, demosaic and lens distortion correction, and before everything else — no white balance, no camera matrix, no edit, no view transform. The composite carries the first source's body, colour matrix and as-shot neutral, so that it is developed afterwards exactly as one of its sources would be: the camera profile, the white balance and every operation in §3.3 are applied once, to the composite, in its own develop.
This is the clause that decides what the output is. Stitching the rendered edits is what a JPEG stitcher does; the result cannot be re-developed, and any difference between the frames' edits becomes a seam. Stitching camera-space pixels produces a photograph the camera could have taken, and nothing is applied twice. The cut sits below the profile, not above it, for a reason S15.3 found in the pipeline: the profile's rendering was applied to every frame of a known body, so a composite that baked it in and then developed as one would render it twice. (Amended 2026-09-27, D19: that rendering was the per-body base curve, retired; the view transform that replaces it is applied to every develop, so the reason stands.) Lens correction alone sits above the cut, because a distorted frame does not align. White balance sits below it because the sensor saw the same light in every frame: un-balanced camera RGB agrees across the overlaps whether or not the camera's auto white balance drifted, and the balanced values would not.
FR-MRG-3 — The output file. (general to any merge) A RAW, in every sense that survives a warp: camera-linear samples at the source's native scale — integers on the first source's black-subtracted scale with its white level, never rescaled to fill 16 bits — with the first source's body, colour matrices, illuminants and as-shot neutral carried, and its capture metadata. What a merge cannot preserve is the colour filter array: a warped image has no sensor grid to re-mosaic onto, so the composite is three samples per pixel. Named from the first source with a stated suffix and placed beside it. Where the sources' folder is not writable — a remote-only tier, a read-only mount — it goes where an export goes (FR-EXP-6) and the interface says so before the merge starts.
The container is a linear DNG (S15.1, decided 2026-09-19): rawler reads back what the application
writes, and the composite re-enters as Format::Dng through the decoder every camera DNG uses.
The point of the whole clause is that the panorama is developed afterwards — white balance,
profile, exposure, everything in §3.3 — as one photograph, from the sensor's own numbers.
FR-MRG-4 — Projection and framing. Cylindrical, spherical or perspective, chosen from the field of view and overridable; the horizon levelled from the estimated rotations, overridable by a drag; the border either cropped to the largest inscribed rectangle or filled, the photographer's choice, the crop the default.
The auto-crop is non-destructive (built 2026-09-19): it is the composite's default crop, not a cut — the whole merge including its border is in the file, and resetting the crop shows it.
The fill is generative and opt-in (revised the same day, having first said "no boundary
fill"): MI-GAN (Sargsyan et al., ICCV 2023; MIT code and weights, models/LICENCE.md) paints
the uncovered border from the picture's own edge, under the inference engine (S16, inference.md). It is a
proposal under FR-MRG-1 — shown on the page, chosen against the crop, confirmed before the merge
— never the default, and a merge that used it says so in its sidecar (border filled) so the
invented pixels are declared, not passed off as captured. §1.3 still holds: the fill touches only
pixels no frame reached, never the photograph; a filled merge keeps the inscribed crop as its
default so the honest picture is one reset away. The filler is loaded from the model directory
when present and the choice is greyed out, with the reason, when it is not.
Experimental, as shipped 2026-09-19. The fill is right in thin borders and wrong in deep
corners, where the model invents cloud and water where there is sky and grass
(panorama.md §13); it ships opt-in with every knob on the page — working scale, edge
erosion, coarse pass, band width, mirror depth, seam feather — each redrawing the preview, and
the knobs used are written into the sidecar's merge line beside border filled. The knobs
leave the page when the defaults are right; the sidecar record stays.
FR-MRG-5 — Honesty of failure. (general to any merge) A frame that cannot be aligned is named, with why — too few matches, no overlap with any other frame, a residual above the stated bound — and the merge stops. Never a silent drop, never a best-effort composite with a frame missing.
The same rule as spot-removal.md's and D17's: a tool that quietly alters or omits part of a
photograph is the failure this application must not have, and here the omission would be an
entire frame.
Amended 2026-09-27: the job no longer stops. The frame is named on its row, with why, and the merge cannot be confirmed until the photographer unticks it; the rest are then solved again from the pairs already measured (panorama.md §15). The omission is the photographer's, made in view, which is what this clause asks — never a silent drop.
FR-MRG-6 — Provenance. (general to any merge) The composite's sidecar carries
derived_from: the content hashes of its sources in order, and the merge parameters. The history
records the merge as the first entry, and export metadata declares the composite as one. Sources
trashed later leave the list dangling; the panel says so and nothing is blocked.
Provenance, not dependency. The composite renders from itself alone; derived_from exists so the
photographer, and anyone they hand the file to, can see what it is made of. The audience is
RAW-literate and a composite declares itself (D17). Nothing here specifies C2PA.
FR-MRG-7 — Execution. (general to any merge) A background job on the pattern FR-EXP-7
established: its own thread, its own GpuContext, a row in the activity panel, cancellable with
NFR-ARCH-3's bound. Alignment runs at proxy resolution and drives the preview; the full-resolution
warp and blend run only on confirm.
FR-MRG-8 — Model-optional. Keypoint detection works without any model weights and better
with them, the convention dr-segment set. Weights that ship are recorded in models/LICENCE.md
before they land, under D8's compatibility test, and their absence degrades quality rather than
disabling the feature.
FR-MRG-9 — Platform. Desktop first. Android runs the same code within NFR-RES-2 and FR-MRG-11, with a stated ceiling on frame count and source resolution, refused with a message, rather than an out-of-memory kill.
FR-MRG-10 — Where the work runs. (general to any merge) Every per-pixel stage of a merge — rendering the sources, the preview reprojection, the full-resolution warp, gain compensation, the seam and the blend — runs on the GPU as WGSL, under ARCH §6.4. The stages that are not per-pixel — keypoint detection at proxy resolution, descriptor matching, and the rotation solve over a few parameters per frame — run on the CPU, and the specification says so rather than leaving it to be "moved later".
The per-pixel stages are the whole cost, and the tablet is where the cost is paid: the output is larger than any single photograph the pipeline has rendered, and a CPU blend of it would take minutes there. The CPU stages are bounded by frame count, not by output size — detection is once per frame at 1024 px, on the same runtime faces and masks use — and moving a small model's convolutions to hand-written WGSL is real work for no visible gain. Seam finding is the one classic stage that resists the GPU; the seam algorithm is chosen for the GPU, not for the paper.
FR-MRG-11 — Tiled in output space. (general to any merge) No stage may hold the composite as one texture, on any platform. The warp, seam, blend and encode proceed in output-space chunks, each pulling only the source tiles that project into it, so the working set is one chunk plus its sources' tiles regardless of how large the composite is.
Two facts force this before memory does. max_texture_dimension_2d is 8192 on many mobile GPUs
and 16384 on desktop, and a three-row panorama is routinely 20 000 px wide — the composite would
not fit a texture even with the memory to spare. And five 24 MP frames at the working precision
are ~1 GB together, which the tablet does not have. ARCH §6.2 applies to the composite as it
applies to a source, and retrofitting it would be the rewrite it warns about.
Non-goals, fixed now. No translation solve or parallax correction — seam placement is the tool for a hand-held set, and a photograph with real parallax is not a panorama. No HDR panorama in one operation until HDR merge exists on its own. No live re-stitch: a different projection or crop after the fact is a new file, not an edit. No video.
3.12 Inference runtime
The face, segmentation and border-fill models run under an inference engine the application selects per device. inference.md is the design: its §3 sets the dependency policy the runtime may reopen, its §10 the milestones (M1–M7) the clauses below cite as acceptance. The clauses are the register entries inference.md §12 promised; they are stated here so the traceability gate can count them.
FR-INF-1 — Runtime selection. On launch the application shall determine, per device and without blocking the first frame, the fastest inference backend that can build and run a session for the shipped models, by attempting it; shall record and reuse that determination until the runtime, driver, hardware or models change; and shall display the backend in use in Settings and on the about screen. Acceptance: inference.md §10 M1 and M5.
FR-INF-2 — Derived engines. Backends that require device-specific compilation shall compile in the background after selection, shall serve requests from the next lower backend until each engine is ready, and shall not change the backend of a job in progress. Acceptance: M5.
FR-INF-3 — Model forms. Quantised model forms are produced at release time from real calibration data and are shipped only when they meet inference.md §10's accuracy gates against the canonical form; the application never quantises on the device. Acceptance: M2, M7.
4. Non-functional requirements
4.1 Performance targets
These are targets to design against and measure, on the reference desktop (AMD Threadripper 2920X, 24 threads, discrete GPU) and a mid-range Android device.
| ID | Metric | Desktop target | Android target |
|---|---|---|---|
| NFR-P1 | Catalog open (50k images) | < 2 s | < 4 s |
| NFR-P2 | Grid scroll | Sustained 60 fps | Sustained 60 fps |
| NFR-P3 | Thumbnail generation throughput | ≥ 100 img/s (embedded preview path) | ≥ 25 img/s |
| NFR-P4 | Open image in develop (to first proxy on screen) | < 400 ms | < 1000 ms |
| NFR-P5 | Slider adjustment → visible update | < 16 ms (one frame) | < 33 ms |
| NFR-P6 | Pan/zoom responsiveness | No dropped frames at 60 fps | No dropped frames |
| NFR-P7 | Full-resolution export (24MP, full chain) | < 2 s | < 8 s |
| NFR-P8 | Idle memory (50k catalog, nothing open) | < 500 MB | < 250 MB |
| NFR-P9 | UI-executor blocking (not "any operation") | Never > 16 ms | Never > 16 ms |
| NFR-P13 | Next image in culling mode (FR-CULL-1) | < 50 ms | < 50 ms |
| NFR-P14 | Focus peaking overlay ready | < 100 ms after preview | < 150 ms |
| NFR-P15 | Drawn mask stroke → visible (ARCH §6.11) | < 16 ms, no cursor lag | < 16 ms |
| NFR-P10 | Touch gesture → visual response | < 16 ms | < 16 ms |
| NFR-P11 | Layout class transition (window resize) | No dropped frames. No loss of photographic state: the open image and version, the selection, scroll position, the in-progress edit and its undo history, and the current mode. Panel disclosure is explicitly exempt — see below | n/a |
| NFR-P12 | Warm-start shader pipeline setup (cached) | < 100 ms | < 100 ms |
Every target above requires a stated measurement method, workload, and pass threshold before it is testable. NFR-P8 in particular must state whether it measures RSS inclusive or exclusive of GPU allocations, and whether it holds after SQLite's page cache warms on a 50k catalog.
On what NFR-P11 means by state. It said "no state loss", which the implementation contradicts on
purpose, so the requirement has been made specific rather than left to be read as forbidding
something it should not. apply_layout_class discards the user's panel open/closed choices when the
class changes, and the argument for that is sound: a choice made in landscape answers a different
question from the one portrait asks, and carrying it across is how a photographer ends up with a
232 px sidebar on a screen with no room for it and no memory of having asked for it. Panel
disclosure is a default, re-derived per class, with the user's disagreement remembered only within
the class where it was expressed.
Everything the photographer produced or navigated to is a different matter, and none of it may be touched by a resize. That is the list in the criterion, and it is the testable half.
NFR-MRG-1 — Merge latency. For five 24 MP frames: the alignment preview (FR-MRG-7) within 5 s on the reference desktop and 15 s on the reference tablet, of which keypoint detection is at most 1 s per frame on the tablet's CPU (measured 2026-09-19 at ~0.4 s, S15.4); the full merge written to disk within 60 s on the desktop. The tablet's full-merge figure is set by S15's blend measurement, not guessed here.
Performance regressions fail the build. §8's benchmark suite runs per-commit; a regression beyond a stated tolerance is a build failure, not a notification. Performance work rots otherwise.
4.2 Reliability
NFR-R1 — No operation shall lose user edit data. The catalog database uses write-ahead logging and survives power loss without corruption.
NFR-R2 — The catalog is backed up automatically on a schedule and before schema migrations.
NFR-R3 — A crash in decode or GPU work shall not take down the application where it can be isolated; the affected image is marked as failed and the app continues.
NFR-R4 — Source image files are strictly read-only to the application, except where the user explicitly requests a destructive operation (e.g. delete, or writing XMP sidecars).
NFR-R5 — Schema versioning. The catalog carries a monotonic schema version. Migrations are forward-only, transactional, and idempotent on retry. Each migration ships with a test that migrates a fixture catalog from every prior released version. The app refuses to open a catalog with a newer schema version rather than corrupting it, and says so.
ARCH §6.6 already anticipates one migration (folder ETags); there will be others, and the machinery must exist before the first one.
NFR-R6 — Corruption recovery. On failing an integrity check at startup, the app offers restore from the NFR-R2 backup, and failing that, rebuild from sources plus sidecars per ARCH §6.12. FR-CAT-8 is what makes that second path real for local-only users.
NFR-R7 — GPU device loss. The GPU layer treats device loss as an expected event (ARCH §6.10): detect it, tear down and recreate the device and all derived resources, and re-drive the current render from the edit graph. No user edit is lost. Recovery is exercised by a test that induces device loss mid-render.
NFR-R8 — No suitable GPU. Where no Vulkan device meeting the NFR-COMPAT-1 baseline is available, the app starts in a stated degraded mode with defined capability limits rather than failing to launch.
Decided 2026-09-19: there is no CPU render pipeline. The degraded mode is the one the
viewer already has (dr_ui::shared_gpu returning None): the library opens, the grid and the
culling views run on embedded previews and cached proxies, metadata, ratings and collections are
fully editable, and develop and export are unavailable and say so. "CPU fallback" in this
document means nothing more than staging through host memory when GPU memory is short; it never
means a second implementation of the operations. ARCH §6.4 stands as written, and NFR-RES-2's
fallback clause has been reworded to match. A second full pipeline was the alternative, and it was
declined for the reason ARCH §6.4 gives: the GPU path is the product, not an optimisation of it.
NFR-MRG-2 — Reproducible merges. The same sources, the same settings and the same device produce a byte-identical composite. Across devices the comparison is tolerance-based, calibrated as S9 calibrates R1: the merge is float work end to end, and ARCH §6.13's bit-identity applies to integer state only.
4.3 Resource behaviour
NFR-RES-1 — Bounded memory. Memory use is bounded and configurable, independent of catalog size and image count. Caches are evictable under pressure.
NFR-RES-2 — GPU memory. The pipeline shall handle images larger than available GPU memory by
tiling. GPU memory headroom is configurable. Where an allocation fails, work is staged through host
memory or refused with a typed error (GpuError::TooLarge) — never rendered by a CPU pipeline,
which does not exist (NFR-R8).
2026-09-27: a linear DNG larger than one texture is no longer refused. It is developed from a reduced copy and full-resolution windows, and exported in tiles (FR-DSP-2's note). A CFA file too large for one texture, or whose photosites overrun the device's storage-buffer limits on upload, is still refused, and there is still no headroom budget.
NFR-RES-3 — Mobile power. On Android the app shall not render continuously when idle. Battery and thermal behaviour are first-class concerns; background sync respects metered-connection and battery-saver settings.
NFR-RES-4 — Disk cache. Thumbnail and proxy caches have a configurable size cap with LRU eviction.
4.4 Portability
NFR-PORT-1 — Platform-specific code is isolated behind interfaces. The image core, catalog, and edit pipeline contain no platform conditionals.
NFR-PORT-2 — GPU shaders are authored once and used on both platforms.
NFR-PORT-3 — Adding a third platform requires implementing the platform interfaces only, not changes to the core.
4.5 Security and privacy
NFR-SEC-1 — RAW parsing treats input as untrusted. Parser hardening, fuzzing, and where practical process or memory isolation for decode.
NFR-SEC-2 — Credentials are never written to the catalog, logs, or plain files. Platform secure storage only.
NFR-SEC-3 — All network traffic uses TLS with certificate validation. No option to disable validation in release builds.
NFR-SEC-4 — No telemetry without explicit opt-in.
NFR-SEC-5 — Face data stays on the user's own hardware. Face embeddings (FR-CULL-8) are handled under a stricter rule than the rest of the catalog.
This is a personal tool for personal libraries (§1.1, D11) — the people in these photographs are the user's family and friends. That is the reason for the rule, not a reason to relax it: the data is sensitive precisely because it is personal, and the user is the only party with any claim on it.
- Never leave the device by default. Embeddings, face crops, and cluster assignments shall not be transmitted, uploaded, or included in any diagnostics bundle (NFR-OPS-1) or crash report (NFR-OPS-2), under any configuration. The diagnostics path has no opt-in for this; it is excluded outright.
- Sync is opt-in and separately consented. Syncing embeddings to the user's own Nextcloud is permitted — it is their server and their photographs, and it saves re-indexing a library per device — but it is off by default, is not implied by enabling photo sync, and the consent states in plain language what is being uploaded and why. Person names, being sidecar data (FR-CULL-12), sync with the sidecar as ordinary metadata.
- No third-party inference. Face detection and embedding run locally. No image, crop, or embedding is sent to a remote inference service, and the app ships no capability to do so.
- Deletable, in one action. The user shall be able to delete all face data — embeddings, detections, clusters, and people — from a single control, without deleting the catalog or any photograph, and to disable face indexing entirely so that no such data is produced.
- Model weights are inspectable. The models used shall be named and versioned in the about screen, with their licences, so a user can determine what is running on their photographs.
Rationale: the rest of this document treats privacy as a property of the network boundary — TLS, credentials in secure storage, opt-in telemetry. Face data needs more than a well-defended boundary, because it is not revocable once it has crossed one, and because it describes people who are not the user. The prohibition is therefore structural rather than configurable: the code paths that would upload an embedding to anyone but the user's own server do not exist. A setting can be changed by accident, or by a future maintainer who has forgotten why it was there; an absent code path cannot.
NFR-SEC-6 — Third-party code runs inside a boundary, not beside the application. (post-v1) FR-PLG-1 admits code the user did not write and the project did not review. The privacy properties asserted elsewhere in §4.5 are properties of this codebase, and none of them survives a plugin that can open a socket.
- Sandboxed by default, with granted exceptions. Computational plugins execute with no filesystem, no network, and no catalog access. Anything more is a declared capability, shown before installation and granted explicitly by the user (FR-PLG-4a).
- No plugin reaches face data. Embeddings, detections, crops, and cluster assignments are outside every plugin interface, under NFR-SEC-5's structural rule: the code path does not exist, so no grant can create one.
- Artefacts are verified before they are loaded, against the hash or signature a configured registry states (FR-PLG-10). An unverifiable artefact is not loaded.
- Nothing is fetched on a file's say-so. Installation is resolved through user-configured registries; a sidecar carries identity only (FR-PLG-7).
- Native shared objects are not a plugin format (FR-PLG-1). An in-process native object holds the application's full privileges, which would make every rule above unenforceable.
Rationale: an extensible application inherits the trust properties of its weakest plugin unless the boundary is structural. The same reasoning as NFR-SEC-5 applies for the same reason — a setting can be changed by accident or by a maintainer who has forgotten why it was there, and an absent capability cannot.
4.6 Execution model
NFR-ARCH-1 — Named executors. The app defines distinct executors — UI, GPU submission, decode pool, I/O pool, network — with stated thread counts and the invariant that no blocking call occurs on the UI executor. This is the mechanism behind R4 and NFR-P9, which currently assert an outcome with no stated means.
Status (2026-09-26). Named and guarded; not yet bounded. dr_ui::executors defines the five
executors with the thread counts of architecture.md §7.1 and the reason for each, and every
long-lived worker in dr-ui and the Android entry point is started through its spawn, which
names the thread <executor>:<role> (net:sync, decode:thumbs). run marks its thread as the
UI executor, and net_runtime's block_on asserts in debug and test builds that it is not called
there; executors' tests show the panic on a thread marked as the UI one and the same call passing
on a worker. Outstanding: the counts are a stated budget, not a limit — each job still gets a
thread of its own, and a pool sized from Executor::threads is the change NFR-ARCH-2's priorities
need; the guard covers block_on only, not a synchronous file read or a catalog query on the UI
thread, several of which the library and People screens make on a click by design (catalog.md
§1); the two mask workers in masks_ui.rs still use std::thread::spawn; and threads the core
crates start (the inference engine's reaper and probe, already named) are outside the module.
NFR-ARCH-2 — Scheduler priority. The tiling scheduler assigns priority classes, with visible-tile work strictly preempting background export and thumbnail work. Without this, NFR-P5's slider latency fails during a batch export — the common case, not an edge case.
NFR-ARCH-3 — Cancellation. Cancellation is cooperative with a bounded worst-case latency (target: observed within 100 ms), and covers in-flight GPU submissions. Every long-running operation named in FR-CAT-1, FR-EXP-7, and FR-NC-6 is cancellable under this model.
NFR-ARCH-4 — Error propagation. No worker error may panic the process. Errors surface as typed results attached to the affected image or job, consistent with NFR-R3 and FR-RAW-4.
4.7 Operations
NFR-OPS-1 — Diagnostics. Structured levelled logging to a rotating, size-capped on-disk log in the XDG state directory or Android app directory, with automatic redaction of credentials and tokens (required by NFR-SEC-2, which currently forbids credentials in logs that are never otherwise specified). A one-click diagnostics bundle includes log, schema version, GPU and driver identification, and app version — with an explicit preview-and-consent step before anything leaves the device.
NFR-OPS-2 — Crash reporting. Local crash capture always; upload only on explicit opt-in (NFR-SEC-4).
NFR-OPS-3 — Preferences. A single versioned preferences store, separate from the catalog, so preferences survive catalog rebuild and multiple catalogs. At least eight requirements refer to configurable settings with no store defined. Device-specific settings (GPU headroom, cache caps) do not sync between devices.
NFR-OPS-4 — Update and first run. State delivery channels and their update mechanisms. This matters concretely because D2 pins rawler at an alpha, non-SemVer version whose camera-support fixes users will need. First-run flow is defined, including platform permission acquisition (FR-PLAT-AND-1 makes first run a permission negotiation on Android, not a welcome screen) and initial root selection.
4.8 Compatibility baseline
NFR-COMPAT-1 — Supported hardware. Stated 2026-09-19 from what the build enforces and the hardware the figures are taken on. Every §4.1 Android figure is measured against this.
| Baseline | |
|---|---|
| Android API | minSdk 28, targetSdk 36 (docker/android/Dockerfile MIN_API/ANDROID_API; the linker targets 28, so the binary runs on the oldest level it claims) |
| GPU, both platforms | A Vulkan adapter accepted by wgpu at its default limits — not the downlevel tier — because storage textures in compute shaders are required; dr_gpu::GpuContext::device_from states this as the floor. No optional feature is required: required_features is empty. The pipeline stores intermediates in Rgba16Float textures, which are core Vulkan; shaderFloat16 arithmetic is not required and FR-DEV-2's "16-bit float" is a storage precision, not a shader-arithmetic one. The GL backend is accepted on desktop as a last resort |
| Desktop | Linux with a Mesa or vendor Vulkan driver the reference machine's era or later; Windows under NFR-COMPAT-2's installer. No minimum Mesa version is asserted beyond "wgpu's default limits are met" |
| RAM | No figure asserted; NFR-P8's budgets are the constraint, not the device's total |
| Reference Android device | HONOR ROD2-W09, Qualcomm SM8635 (Adreno), 12 GB, Android 16 / API 36, 1920×3000 at 400 dpi. Every Android column in §4.1 is measured on it |
| Second-vendor device | None. The clause asking for a Mali or PowerVR device is unmet and known unmet: no such hardware exists in the project, and S2's two-vendor run is blocked on it (#47). Until one exists, an Android figure in this document is an Adreno figure |
Adreno, Mali, and PowerVR diverge significantly in compute behaviour and in external-memory interop — exactly what spike S1 tests. The spec already applies this reasoning to desktop drivers; it applies at least as strongly on Android, which is why the missing second vendor is recorded as a gap rather than dropped.
A figure worth noting. At 400 dpi the reference tablet's portrait width is roughly 768 logical
pixels, which is under D15's assumed ~1024 and under EXPANDED_MIN_WIDTH. Whether Slint's scale
factor agrees with the platform density is open issue N6; this row records the measurement and
does not resolve it.
NFR-COMPAT-2 — Distribution channels. Decided 2026-09-19. Every channel is self-distribution: built by the Gitea CI from one commit, installed by hand, and published to no store or catalogue. That is a consequence of D13 — the face weights' grant rules out every public channel — and it is the position until those weights are replaced.
| Platform | Channel | Built by |
|---|---|---|
| Linux | Arch package (packaging/PKGBUILD) |
CI, packaging/pkg/ |
| Linux | Flatpak from the manifest in packaging/flatpak/, installed locally — not submitted to Flathub |
CI |
| Android | Signed release APK, sideloaded (docker/android/assemble-apk.sh, the release key of 2026-09-11) — not Play, not F-Droid |
CI |
| Windows | NSIS per-user installer cross-built from Linux (docker/windows/, FR-PLAT-WIN-3) |
CI |
Status (2026-09-26). The Flatpak row is the decision, not yet the state: no workflow under
.gitea/workflows/ builds it, and none has been built by hand (FR-PLAT-LIN-3).
Two consequences the channel decision has on the requirements above it. ARCH §6.9's SAF-only storage model was written because Play would reject anything else; a sideloaded build could request broader permissions, and FR-PLAT-AND-1 keeps SAF anyway, because the design is right on its own terms and because the day a public channel opens is not the day to redesign storage. And S11, the Play permissions dry-run, becomes a pre-publication step rather than a Tier-1 spike: there is nothing to submit until D13 is closed.
4.9 Accessibility and internationalisation
NFR-A11Y-1 — Localisation. All user-facing strings, including operation and parameter labels
resolved from LocalizedString (FR-DEV-3a), are externalised and translatable without
recompilation. Note the constraint this creates: those labels live in core crates that cannot
depend on the UI (ARCH §6.5a), so the localisation mechanism must itself be UI-independent. State the
format, the locale-resolution rule, and whether RTL layout is in v1 scope.
NFR-A11Y-2 — Accessibility. Controls expose accessible names, roles, and values to the platform accessibility layer (AT-SPI on Linux, TalkBack on Android). Platform font scaling is honoured without clipping. Non-canvas UI meets WCAG AA contrast.
Slint's accessibility support on Android requires verification — this may be a toolkit gap, and it is far cheaper to discover now than after the UI is built.
NFR-A11Y-3 — Colour-independent status. No status is conveyed by hue alone. Colour labels, clipping indicators (FR-DSP-7), and the HSL mixer carry a shape or text affordance. This matters more in a colour-grading application than in most software.
4.10 Inference
NFR-INF-1 — Embedding comparability. Face embeddings shall be computed at a precision whose deviation from the f32 reference is within inference.md §7's gate, on every backend, so that embeddings from any device are comparable — FR-CULL-9's calibration depends on it. Acceptance: inference.md §10 M3.
5. Data model and architecture
Entity definitions, invariants, and all architectural constraints are specified in
architecture.md — §3 (core abstractions), §6 (data architecture), and
§12 (architectural constraints — whose subsections keep their original 6.x numbers, so ARCH §6.1 in this document means §12's first constraint, not §6.1 of the data architecture).
Requirements in this document that depend on an architectural guarantee cite it inline. The constraints most load-bearing for testability are:
| Constraint | Why a requirement depends on it |
|---|---|
| ARCH §6.1 — no CPU round-trip | FR-DSP-3, FR-DSP-7, NFR-P5 are unachievable without it |
| ARCH §6.11 — GPU-rasterised masks | NFR-P15 (no brush lag) |
| ARCH §6.12 — sidecars authoritative | NFR-R6, invariant behind FR-CAT-8 |
| ARCH §6.9 — Android SAF only | FR-CAT-1a, FR-PLAT-AND-1, and the NFR-P1/P3 Android figures |
| ARCH §6.6 — no sync tokens | FR-NC-4 |
| ARCH §6.13 — integer-only bit-identity | R1's tolerance, §9 golden images |
6. Decisions
Rationale, evidence, and the eliminated alternatives are recorded in architecture.md §13. Outcomes only:
| # | Decision | Outcome |
|---|---|---|
| D1 | Language and UI framework | Rust + Slint, rendering through wgpu |
| D2 | RAW decoder | rawler; LibRaw fallback behind a trait |
| D3 | First milestone | Delivered — milestone-v0.1.md, closed 2026-08-30 |
| D4 | Nextcloud sync mechanism | ETag pruning, chunked upload v2, Login Flow v2 |
| D5 | Colour management | lcms2 + GPU-side matrix/LUT transforms |
| D6 | Shader authoring | Hand-written WGSL |
| D7 | Network stack | reqwest + quick-xml |
| D8 | Licence | GPLv3 |
| D9 | Operation UI model | Declarative parameter descriptors |
| D10 | Interface strategy | One adaptive UI, tablet + desktop |
| D11 | Product positioning | Culling-first differentiator; see below |
| D12 | Scope versus pace | DECIDED 2026-09-19 — settled by events; full scope stands, no v1 date |
| D18 | Derived images | DECIDED 2026-09-19 — a merge writes a new source file; no multi-source Version |
| D19 | Scene-referred pipeline | DECIDED 2026-09-27 — edits on unbounded scene-linear colour; one view transform, last; per-body base curves retired |
| D21 | DNG reference tone for raws | DECIDED 2026-10-03 — the view transform's DNG reference curve (profile's, else ACR3 default, via RGBTone in ProPhoto) is the default for every raw, at contrast 1.5 (a ×1.07 power about grey); measured against Lightroom exports of photographs with neutral look settings; the sigmoid stays a choice |
| D20 | DCP camera profiles | DECIDED 2026-10-02 — HueSatMap and LookTable as a scene operation after exposure; embedded profile first, then a matched .dcp; tone curve not applied; none shipped |
D11 — product positioning
Settled by requirements calibration, 2026-08-08.
| Dimension | Decision |
|---|---|
| Audience | RAW-literate photographers, Linux-comfortable. Docs matter; hand-holding does not. |
| Library scale | 10k–50k images |
| Culling | The core differentiator (§3.9) |
| Focus checking | Peaking and zoom |
| Ingest | Full workflow — template rename, checksum verify, dual-destination |
| Colour defaults | Good, not obsessive — matrices plus one scene-referred view transform for every body (D19; the per-body base curves are retired) |
| Film simulation | Fujifilm explicitly targeted |
| AI | Denoise in v1; masking deferred. Per-face eye state and head pose are in v1 as culling evidence, not AI (FR-CULL-8a, FR-CULL-13); gaze deferred (§7) |
| Local adjustments | Full masking, GPU-rasterised |
| Sync | The reason the project exists |
| Durability | Sidecar-first |
| Licence | GPLv3 |
| Pace | Evenings and weekends, indefinite |
D12 — scope versus pace · DECIDED 2026-09-19
Settled by events. The reconciliation this decision asked for never happened as a decision; it happened as a build. Between the calibration and this audit the milestone D3 pointed at was delivered and closed (2026-08-30), and the application went through eleven further releases to 0.12.2 carrying culling, develop, masks, faces, sync, Flatpak and a Windows channel. The scope as calibrated stands as written, and the pace is the pace. What was reconciled is the date: v1 has none. A requirement in this document is in scope until §7 says otherwise, and "post-v1" in §7 is the only mechanism by which something leaves the count — used on 2026-09-19 for the plugin API, and for nothing else.
The two tensions below are kept as the record of what the decision weighed. Both resolved themselves the same way: the expensive selection was built anyway, and the differentiator was built alongside the develop chain rather than instead of it.
The calibration selected an ambitious feature set — full tablet editing, full ingest, culling as a differentiator, complete GPU masking, AI denoise, Fuji-first colour, deep sync, sidecar durability — against a stated pace of evenings and weekends, indefinitely.
Those are not compatible as stated. This is not an argument against any individual choice; it is that the two must be reconciled before a build order can be set.
Two specific tensions:
1. Full tablet editing is the most expensive selection, chosen against the research finding that volume photographers do not edit on tablets. It carries SAF storage at unproven scale (S10), process-death durability, background execution limits, and two GPU vendors to validate. Tablet culling is the validated workflow, is what the differentiator points at, and costs a fraction as much.
2. The v1 milestone and the differentiator disagree. D3's vertical slice proves the architecture but is useful to nobody. If culling is what no existing tool does well, a culler is both a smaller build and a usable one — it needs no develop chain.
D12 was what set D3, and architecture.md §11's build order followed it.
D15 — target devices · DECIDED 2026-08-22
A 12-inch tablet and a desktop. No phone.
Recorded because it is load-bearing for the interface and invisible in the code. Every phone-shaped answer — a bottom tool strip, a one-tool-at-a-time sheet, thumb-reach zones — is designing for hardware this project does not target, and each would have cost a second layout to keep in step with the first.
What survives the decision is the input difference rather than the size one:
a 12-inch tablet is touched, and NFR/FR-UI-7's position already covers it —
hit regions grow to the modality, layout does not move. The practical rules
that fall out are in docs/ui-navigation.md D-N2: no hover-only affordance
and no modifier key may be the sole route to anything, because a tablet has
neither.
EXPANDED_MIN_WIDTH is 820 logical pixels and a 12-inch tablet is ~1024
across in portrait, so both orientations of both targets are the expanded
layout. The compact class remains as graceful degradation for a narrowed
desktop window, not as a second interface.
D13 — face inference runtime and model licensing · RUNTIME ANSWERED · LICENSING POSITION RECORDED 2026-09-19
Runtime, reopened 2026-09-19 — to the extent of inference.md §3. The pure-Rust build stands:
ortstill links nothing. What changed is thatort::set_apican be handed the table of alibonnxruntimethe package installs, and the app now looks for one at launch and runs on tract only when there is none. Measured before it was built: tract runs every model on one core at the same speed on a tablet and a twenty-core desktop; ONNX Runtime's CPU provider alone is 3–10× that, the Hexagon at int8 runs the detectors in 1–3 ms, TensorRT at fp16 in 2–3 ms. The Android APK bundles ONNX Runtime and Qualcomm's HTP libraries (§3.1 — the QNN licence is read, not summarised, before a release carries them); the desktop packages bundle nothing NVIDIA and use a system CUDA/TensorRT if the probe finds one that works. The licensing half is unchanged.
Position, 2026-09-19. DarkRoom is non-commercial software, built and installed by its author for personal libraries, and it uses the InsightFace SCRFD detectors and ArcFace embedder under their non-commercial research grant as such. That is the position, and it is taken with the risks written down rather than around them:
- The grant is a use restriction and binds every user of the app, not only the project. It is not GPL-compatible and cannot become so; D8's licence covers this codebase and not those weights, which
models/face/README.mdsays in as many words.- It is incompatible with every public channel — Flathub, F-Droid, Play. NFR-COMPAT-2's channels are therefore all self-distribution (a CI-built APK, a Flatpak and an Arch package installed by hand, an NSIS installer), and nothing is published to a store or a catalogue while these weights are in the tree. Publishing is the event that reopens this decision, and S14's licence search is what would close it: the OCEC eye-state weights showed a clean chain is possible, and a clean detector and embedder are what is missing.
- The about screen names the models and their grant (NFR-SEC-5), so the person running the app can read the restriction they are under.
The runtime half is unchanged below.
Updated 2026-08-21. The runtime half of this decision is settled, and by a route the table below does not contain.
ort2.0'salternative-backendfeature disables its linking entirely and lets another engine supply theOrtApi;ort-tractsupplies it fromtract, which is pure Rust. So the third option's operator coverage comes with the first option's dependency profile — no C, no NDK problem, no exception to the policy. Measured on a real graph before being relied on: YOLO26n-seg loads with zero unsupported operators and runs 640×640 in ~470 ms of CPU (docs/segmentation.md §13). §3.9.1's detector and embedder are different graphs and their coverage has not been checked, but the approach no longer needs a decision.The licensing half is untouched. The InsightFace weights are still non-commercial and still unusable here. That remains what S14 has to resolve first.
Updated 2026-09-19. Two further models were read for FR-CULL-8a. Eye state: OCEC (PINTO0309) is MIT for code and weights and trained on ODC-By 1.0 and Apache 2.0 data — the first face-adjacent weights found with a clean chain end to end, and the smallest by two orders of magnitude. Gaze: none. MobileGaze, L2CS-Net and their descendants carry MIT on the repository and Gaze360 in the weights, and Gaze360's research licence restricts "models trained on dataset" by name; MPIIGaze and ETH-XGaze are no better. Head pose needs no weights at all. None of this moves the detector and embedder, which are still the buffalo grant and still what S14 resolves first: clean eye weights behind an unshippable detector ship nothing.
§3.9.1 needs to run two neural networks locally. That collides with two settled positions, and neither collision is small enough to leave implicit.
1. The pure-Rust dependency policy. Every dependency choice in this project has gone the same way, for the same stated reason: rustls over aws-lc-rs, bundled SQLite over the system library, a Rust Lensfun port over liblensfun, zune-jpeg over libjpeg — no C dependency to satisfy under the Android NDK (D1's whole premise). The obvious way to run ONNX models is the ONNX Runtime C++ library, which would be the largest exception to that policy in the codebase, and it would land on the platform the policy exists to protect.
The options, in the order I would try them:
| Option | Cost |
|---|---|
| wgpu compute, models hand-ported to WGSL | No new dependency at all — the GPU device and shader infrastructure already exist (ARCH §5). Highest implementation effort, and a ViT is a lot of shader. |
burn with the wgpu backend |
Pure Rust, uses the existing GPU. Young, and ONNX import maturity needs checking against these two specific graphs. |
ort (ONNX Runtime bindings) |
Fastest to working code, best operator coverage. Reintroduces the C dependency and the NDK cross-compilation problem the policy avoids. |
The tension is real: the cheapest path is the one that breaks the rule. This is worth an explicit decision rather than a default, and S14 is what informs it.
2. Model licensing is a distribution blocker, not a detail. The obvious pretrained weights are not redistributable under GPLv3. The InsightFace "buffalo" family — ArcFace and the SCRFD detector, the standard choices — are licensed for non-commercial research use only, which is incompatible with this project's licence and with Flatpak, F-Droid, and Play distribution (NFR-COMPAT-2). Other candidate weights need their licences read individually rather than assumed.
Two ways out, both with costs:
- Find permissively-licensed weights and ship them in-tree. Clean, offline-first, consistent with how the Lensfun database ships. Requires that suitable weights exist at acceptable accuracy.
- Download models on first use, with the user accepting the upstream licence. Sidesteps redistribution but adds a network dependency to a feature that is otherwise entirely local, needs a hosting story, and sits badly with the local-first posture of NFR-SEC-5.
This must be resolved before implementation, not during it. Discovering at packaging time that the feature cannot ship is the expensive failure, and it is entirely avoidable — it is a licence- reading exercise, not a research question. S14 therefore puts it first.
Prior art available: ../scene-actor-extraction is a working implementation of this pipeline
(SCRFD detect → 5-point align → 512-d embedding → Platt-calibrated similarity), benchmarked at 67.4%
macro-F1 on held-out films. Its C++ does not port — different language, OpenCV and TensorRT
dependencies — but its design decisions do, and they are the expensive part: the calibrated
probability space that FR-CULL-9 requires, the discipline of never thresholding a bare cosine, and
the practice of leaving an uncertain face honestly unnamed. Personal libraries should also score
better than its film benchmark: cooperative subjects, better lighting, and a closed gallery of dozens
rather than thousands.
D17 — cross-frame face repair · OPEN
Raised 2026-09-19: pair FR-CULL-8a's eye state with the burst, and let a closed-eyed face in the chosen frame be replaced by the same person's open-eyed face from a neighbouring frame. It inverts FR-CULL-5's fear — the tool rescues the only frame of the moment instead of rejecting it — and it collides with four things this document says.
- §1.3, "not a pixel editor." Survivable by the FR-DEV-8 precedent: a repair is numbers in
the graph and no pixels are stored. A face repair is a spot whose source is another frame,
aligned by the similarity the face subsystem already fits between two landmark sets and blended
by the membrane heal already shipped. It reopens
spot-removal.md's non-goal that the source is "a patch from the same photograph", which would be revised deliberately, not quietly. - ARCH §6.3, single-source
Image. The real one. §7 defers panorama, HDR and focus stacking on the grounds that the schema cannot express an image derived from several sources. This is that case in a milder form — the output is still frame A, but A's sidecar now names B — and it brings the rules the tiers already have: B's original must be present to render A (FR-NC-6c: visible and priced before starting), trashing B must know A depends on it (FR-CAT-15), andVersion::mergehas never seen a cross-reference. Reopening ARCH §6.3 for this reopens it for the three deferred rows at once, which is the argument for doing it once and properly rather than for this alone. D18 has since done it once, the other way: a merge writes a new file and ARCH §6.3 stands. That leaves this case as the only one that would still need a cross-reference — the output is frame A, not a new file — so the cost above is now this feature's alone to justify. - Tone. The source patch goes through A's chain, not B's — B demosaiced and run through A's parameters to the head of the detail chain, then sampled. A second small pipeline at proxy resolution; a second full demosaic at export. The heal hides lighting drift between frames; it does not hide a turned head, and no blend does.
- Never automatic.
spot-removal.md's own rule: a false positive silently alters a photograph, which is the failure this application must not have. The tool may propose — same person, closed here, open two frames on, small alignment residual — and applies on a press.
And one thing the document does not say: provenance. The audience is RAW-literate (D11). A composite declares itself — the source frame in the sidecar and the history, and in export metadata. Nothing here specifies C2PA; that is a decision this one depends on.
Deferred under D12 until FR-CULL-8a exists and spot removal's disc has, in its own words, been finished and used. The first cut, when it comes, is "clone from a neighbouring frame" as a spot source; the face-aware proposal is a layer over that.
D18 — derived images · DECIDED 2026-09-19
A merge produces a new source file, not a multi-source Version. The composite is written
beside its sources (FR-MRG-3), gets its own sidecar and content-hash identity, and is from then on
an ordinary Image: developed, synced, exported and trashed like any other. Its sidecar carries
derived_from (FR-MRG-6) as provenance, not as a dependency — trashing a source does not break
the composite, and rendering it needs nothing but itself.
This is the ARCH §6.3 question §7 has been keeping open for panorama, HDR and focus stacking,
answered once for all three. The alternative — a Version whose inputs are several other images,
rendered live — was priced in D17: sidecar cross-references, Version::merge seeing a reference
for the first time, FR-CAT-15 knowing that trashing B breaks A, FR-NC-6c requiring every source
present and priced before A can render, and an export that decodes N files. Every one of those is
a change to a subsystem that works today, and none of them buys the photographer anything a file
does not. It is also what Lightroom does, and the audience (D11) knows it.
What it forecloses, so that it reads as a decision: re-merging with different settings is a new file, not an edit to the old one; and the composite occupies disk — a five-frame panorama is a 100–200 MB file, which the photographer chose to make. D17 narrows accordingly to the one case where the output is still frame A, and inherits nothing from this decision but the provenance rule.
D19 — scene-referred pipeline · DECIDED 2026-09-27
Edits operate on scene-linear, unbounded colour in the working space, and one view transform, last, maps it to a display range. Range, encoding and gamut are all deferred to that point, as quantisation already was (FR-DEV-2).
Why now. The spec missed Ansel, Aurélien Pierre's fork of darktable 4.0, and with it the argument he spent years making in darktable: a display-referred curve early in the pipeline throws away what every later stage needs. Reading the code against that argument found four places it applied:
- The base curve clipped. It was a five-point spline on the unit square, flat past its last point, so every value above 1.0 — every recovered highlight — left it at the same number, per channel.
- The detail stage was handed non-linear data. The fused pass stops at "linear working values" when a sharpener or a blur follows, but it stopped after the base curve, so the neighbourhood operations convolved curved, clipped values while their comments promised the opposite.
- The edits ran in camera RGB. The matrix came after them, so
luminance()'s Rec.709 weights were applied to camera primaries and a hue in the colour mixer was a different hue on each body. ARCH §5.2 had always drawn the matrix first; the code had drifted. - The tone curve clamped to [0, 1] and applied a 2.2 gamma around its spline, mid-chain.
What changes. The order becomes: demosaic → as-shot white balance and the white balance operation, in camera RGB → the camera matrix → every other point operation and every mask layer → the detail stage → the view transform (FR-DEV-3j), or the film stock (FR-DEV-3f) when one is chosen → the output transform. With a detail stage the view transform is a dispatch of its own after it, composed by the same generator as the fused pass. Nothing before the view transform clamps above 1.0 or display-encodes, and a test says so (FR-DEV-2).
What is retired. The per-body base curves and their database (FR-DEV-3e). Their own file called them hand-tuned shapes rather than measurements, and not enough was known about where the shapes came from to keep them as defaults behind sliders. Body character is the matrix's, and the DCP's when it lands.
What it costs.
- Every photograph renders differently. The default view transform was fitted so middle grey lands where the retired default curve put it and midtones stay within 0.3 EV of it, but the upper midtones are darker and the highlights roll off over two more stops. Previews rendered before the change keep the old look until they are rendered again.
- Film edits change meaning. A tone or colour operation beside a stock used to act on the film's output; it now acts on the scene the film receives.
- Tablet and desktop must be released together. No schema changes and the sidecar gains only ordinary parameters, but two peers on different builds render the same edit differently.
- One more dispatch with a detail stage, for the view transform after it.
Rejected. Keeping the per-body curves as the view transform's per-body defaults, for the provenance reason above. Leaving the film at order 25 and having it suppress the view transform: simpler, and it kept existing film edits' meaning, but it left a display-referred rendering in the middle of the chain, which is the thing this decision removes. A fixed view transform with no controls: it would have been smaller, but a scene-referred pipeline whose white point cannot be moved hands the photographer a shoulder they cannot place.
Deferred. Working-space primaries of Rec.2020 rather than Rec.709. The range is already unbounded, but several fragments floor at zero, which clips a colour outside sRGB, and the colour mixer's bands and the colour grading wheel would need their hues re-measured. Gamut compression beyond the output transform's clip goes with it.
D20 — DCP camera profiles · DECIDED 2026-10-02
A camera profile's HueSatMap and LookTable are applied by a camera_profile scene
operation at order 25, after exposure, converting into linear ProPhoto and back inside its own
fragment. Design and the full argument: camera-profiles.md.
Why now. The library's 9,348 Canon 6D DNGs carry Adobe Standard's tables, which Lightroom rendered them through, and DarkRoom ignored them, so every hue on those files sat somewhere other than where Lightroom put it. Measured after building it: the tables are not why Lightroom's rendering looks richer — at defaults they lower mean saturation by 3–9 %, because Adobe Standard's look desaturates dark tones to sit under Camera Raw's tone curve, which DarkRoom does not apply. The richer colour is tone, and the Vivid presets (FR-DEV-6) are what answers it today (camera-profiles.md §1).
Why there. The matrix snippet stays what D19 made it, and every copy of it (masks, picker, camera-space tap) stays correct without changing. Hue and saturation are invariant under the uniform gains that precede order 25, so a 2.5-D HueSatMap gives the same answer there as straight after the matrix, and the LookTable sees the photographer's exposure, as it does in the SDK.
Rejected. Extending the matrix snippet: every duplicate of it would have had to follow. Two
operations, one per table: the HueSatMap has no control of its own and commutes to the same place.
Applying ProfileToneCurve: a per-body tone curve is what D19 retired. Shipping Adobe's profiles:
they are not ours to ship. Handing the tables to the graph through a setter, as lens profiles are:
every render path would have to remember to call it. They travel with the decoded image, as the
matrix does.
What it costs. Every DNG with an embedded profile renders differently; previews refresh only when rendered again; tablet and desktop release together. The profiles directory syncs through the library's derived folder (camera-profiles.md §13).
D21 — DNG reference tone for raws · DECIDED 2026-10-03
Decided by measurement, later the same day. The library's photo gallery holds Lightroom 6 exports of raws that are in the library, each carrying its Camera Raw settings. Clustered by those settings, 663 exports had none of the house look (Linear curve, no HSL, no parametric curve, no split toning); 60 of them with their raws, two thirds fitted and one third held out, rendered by DarkRoom against Lightroom's JPEG (MSE, sRGB 8-bit):
| Rendering | Held-out MSE |
|---|---|
| 0.20.0's sigmoid at its defaults | ~1200 — about 0.7 EV darker, and flatter |
| Sigmoid, exposure, contrast and white fitted | ~150 (contrast 1.73, +0.73 EV) |
| DNG reference curve after baseline exposure, at contrast 1.4 | 224 |
| DNG reference curve, contrast 1.5 | ~150 |
| DNG reference curve, exposure, contrast and white fitted | 143 |
So the DNG reference curve is the default for every raw, and the default contrast is 1.5 — under that
curve a power of 1.5/1.4 about grey (REFERENCE_CONTRAST is where the curve is untouched). The
brightness needs nothing: baseline exposure and the curve together land where the earlier exports do. The
profile's look strength, vibrance and saturation bought nothing measurable on those exports. The
user chose to change every photograph rather than keep edited ones on the old rendering. The
fitting tools live outside the repository (darkroom-lrfit). Amended 2026-10-04: the look strength now defaults to 0. It scored the same at 100, 50 and 0 (held-out MSE 140, 140, 143) and the rendering is 9 % more colourful without it — the table desaturates near-neutral tones, where the default was short of those exports; the user chose more colour.
Amended earlier the same day: the default was not decided. The measurement below was against
Lightroom previews of photographs carrying the user's Lightroom edits — a house look of HSL
saturation (Blue +58, Aqua +50, …), Highlights −40 and Blacks −20 in every DNG's XMP — so it said
nothing about Camera Raw's base rendering. Under the DNG reference curve _MG_9080 renders brighter
than its Lightroom preview (mean 0.39 against 0.31). The sigmoid stays the default; the curve
below is a choice; the default is decided by measurement against Lightroom exports of unedited
photographs (the lr-fit work). What follows is the original text.
The view transform has two curves, and Camera Raw's is the default for every raw. It is the
profile's ProfileToneCurve, or the ACR3 default curve where the profile has none or there is no
profile, applied Camera Raw's way — on the largest and smallest channel in linear ProPhoto, the
middle placed proportionally — after BaselineExposure. D19's sigmoid stays as the other choice.
Design: camera-profiles.md §11–§13.
Why. Measured on the library's 6D DNGs after D20 (camera-profiles.md §1): the profile tables lowered saturation, because Adobe's look tables were tuned to sit under this curve. The user's complaint was that Lightroom's rendering is more colourful, and this curve is most of the reason. Chosen by the user over limiting it to raws with a profile, or making it opt-in.
What it reverses in D19. D19 rejected per-body curves as defaults because their provenance was unknown. A DCP's curve and the ACR3 table have known provenance — Adobe's, published — and so the objection that retired the base curves does not apply. D19's other half stands: nothing before the view transform clamps, and the curve is the view transform, last.
What it costs. Every raw renders differently again, and highlights above display white clip where the sigmoid rolled them off; Sigmoid is one click away. Previews refresh only when rendered again; tablet and desktop release together.
D16 — plugin licensing · OPEN, post-v1
Deferred with §3.10 on 2026-09-19. Still to be answered before the format is published as stable, which is now a post-v1 event; nothing in v1 waits on it.
D8 puts the application under GPLv3. §3.10 admits third-party plugins in three forms, and the derivative-work question is answered differently for each — a YAML-and-WGSL declaration is data of the kind the GPL has never claimed, a WebAssembly component communicating over a defined interface is arguably at arm's length, and an interpreted Slint component compiled into the application's own widget tree is not.
This must be answered before an ecosystem exists, not after. Contributors will not adopt a plugin format whose licence terms are unstated, and a term introduced later cannot be applied to plugins already written.
Three questions, in order of how much they constrain the design:
- May a plugin be non-free? If yes, the interfaces are a deliberate licence boundary and must be documented as one. If no, the registry (FR-PLG-10) enforces it and the default registry lists only GPL-compatible plugins.
- Does the answer differ by class? Declaring class 1 unambiguously data, whatever is decided for class 3, is defensible and costs nothing.
- What does the default registry require? Licence metadata is a field in the registry index either way, so the field should exist from the first release regardless of what policy is attached to it.
D16 does not block FR-PLG-2, which concerns operations shipped in this repository under D8 already. It blocks publishing a third-party plugin format as stable.
7. Out of scope for v1
Deferred deliberately. Listed so their absence reads as a decision rather than an oversight, with a note where deferring now constrains the design later.
| Deferred | Note |
|---|---|
| Tethered shooting | — |
Undeferred 2026-09-19 — §3.11 (FR-MRG-1 … FR-MRG-11), under D18, which answers the schema question this row was holding open: a merge is a new source file, and ARCH §6.3's single-source Image does not change. |
|
| HDR merge | Deferred. D18 answers the data model; the merge itself — exposure alignment, ghost handling, the tone of the result — is not specified. FR-MRG-3, 5, 6, 7, 10 and 11 are written to be general to it. |
| Focus stacking | Deferred, on the same terms as HDR merge. |
| Cross-frame face repair ("best take") | A face from a neighbouring frame of the same burst, aligned by its landmarks and blended by FR-DEV-8's heal. Mechanically a spot whose source is another photograph; the same multi-source schema question as the two rows above, arriving early — D17. Deferred rather than refused, with its non-goals fixed now: never automatic, geometry not corrected, source frame declared in sidecar, history and export. |
| Gaze and eye-contact estimation | Every open gaze model read on 2026-09-19 is trained on Gaze360, MPIIGaze or ETH-XGaze, all research-only, and Gaze360's licence restricts models trained on it by name — the InsightFace situation again (D13). Head pose from the five landmarks is the proxy (FR-CULL-8a). Revisit when weights with a clean data chain exist; iris offset within the eye crop is the licence-free fallback if the proxy proves too weak. |
| Print layout | — |
| Soft proofing | Parameterise the output colour stage by an arbitrary profile so this becomes a UI addition, not a pipeline change. FR-EXP-3's print-dimension mode already half-commits to print workflows. |
| Undeferred 2026-08-09, in the narrower form specified in §3.9.1 (FR-CULL-8 … FR-CULL-13): people grouping and search, no automated selection. Reclassified as culling rather than AI — it is the same mechanical-grouping category as FR-CULL-5, not the taste operation AI masking is. Gated on spike S14 and decision D13. | |
| Undeferred 2026-09-19 — it had been built. FR-DEV-3i states what exists: subject and category masks from local models, stored as identity, editable like a drawn mask. The shape this row asked for when it was written — point at a thing, get an editable mask — is the shape that shipped. | |
| AI upscaling | Deferred. Lower priority than denoise, which has no manual fallback. |
| Video | — |
| Plugin API | Post-v1, decided 2026-09-19. Specified in full in §3.10 as the design of record; every clause there is marked (post-v1) and sits outside the coverage denominator. D16 (plugin licensing) defers with it. FR-DEV-3c's compile-time declarations are not plugins and remain in v1. |
| Multi-user / server-side catalog | — |
| Watermarking | Cheap if the export pipeline anticipates a compositing stage; expensive to retrofit otherwise. Consider reserving the stage now. |
| Multiple catalogs, catalog merge | Interacts with NFR-OPS-3: preferences must not live in the catalog. |
| Geotagging and map view | — |
| Web gallery, slideshow | — |
| DNG conversion | — |
8. Verification approach
| Requirement class | How verified |
|---|---|
| Performance (§4.1) | Automated benchmark suite against a synthetic 50k catalog, run per-commit on the reference desktop and periodically on the named reference Android devices. A regression beyond stated tolerance fails the build. |
| Rendering correctness | Golden-image tests: fixed source + fixed edit graph → comparison within R1's stated tolerance, not checksum equality. Run on both platforms and both Android GPU vendors. |
| Colour accuracy (FR-DEV-3e) | ColorChecker exposures per launch body, asserting ΔE2000 within threshold against reference values. |
| RAW decode coverage | Corpus of sample files per supported camera body; decode-and-checksum regression suite. |
| Robustness (NFR-SEC-1) | Continuous fuzzing of the decode path. |
| Sync correctness | Simulated two-device scenarios including conflict, offline edit, and interrupted transfer. |
| Memory bounds | Long-running soak test scrolling a large catalog, asserting bounded RSS and GPU memory. |
| Schema migration (NFR-R5) | Fixture catalogs from every prior released version, migrated forward and verified. |
| Device loss (NFR-R7) | Induced VK_ERROR_DEVICE_LOST mid-render; assert recovery with no lost edits. |
| Process death (FR-PLAT-AND-3) | Kill the Android process mid-edit; assert session and viewport restore with at most the last uncommitted change lost. |
| Source relocation (FR-CAT-9) | Move, rename, and disconnect sources; assert offline marking, reconnection by hash, and no catalog row loss. |
| Cancellation (NFR-ARCH-3) | Assert every long-running operation observes cancellation within the stated bound, including in-flight GPU work. |
| Layer separation (ARCH §6.5a) | CI dependency-tree assertion: no core/* crate may transitively depend on a UI toolkit. |
| Operation self-description (FR-DEV-3c) | A test operation added to the registry appears in a generated panel with no frontend change. |
| Adaptive layout (§3.5) | Snapshot tests at each breakpoint, and a resize test asserting no loss of photographic state across a layout-class transition — NFR-P11's list, with panel disclosure exempt as that clause says. |
| Touch targets (FR-UI-3) | Automated check that interactive elements meet the 44pt minimum in touch modality. |
| Export sizing (FR-EXP-3) | Per-mode dimension assertions, including aspect preservation, fill-crop centring, and the upscale-disabled fallback. |
| Identity calibration (FR-CULL-9) | Reliability diagram over a hand-labelled corpus: stated probability against observed match rate, asserted within tolerance across the range — not a single accuracy figure, which would hide exactly the miscalibration this tests for. Plus a static assertion that no comparison thresholds a raw similarity. |
| Face data confinement (NFR-SEC-5) | Assert that a generated diagnostics bundle contains no embedding or face crop, and that with sync disabled no face data appears in any outbound request. Verified by inspecting what the code can emit, since the requirement is the absence of a path. |
9. Validation spikes
Small experiments that de-risk the highest-uncertainty assumptions before substantial build work. Ordered by risk. With D1 settled these validate the chosen stack rather than choosing between stacks.
Tier 1 — before any substantial build work
| # | Spike | Answers | Relates to |
|---|---|---|---|
| S1 | Slint + wgpu zero-copy on Linux: a compute shader writes a texture, create_texture_from_hal imports it, Slint composites UI over it. Drag a slider for 10 minutes watching for tearing, leaks, and sync bugs |
Whether ARCH §6.1 holds in the chosen stack | D1, ARCH §6.1 |
| S2 | Slint + wgpu on Android, on two devices from different GPU vendors (Adreno and Mali) | Whether the Android GPU path holds across vendor divergence | D1 |
| S10 | Android SAF at scale: enumerate a 10k-file document tree and perform random-access range reads over a document fd. Measure against NFR-P1 and NFR-P3 | Whether ARCH §6.9's forced storage model meets the stated Android performance targets | ARCH §6.9, FR-PLAT-AND-1 |
| S11 | Play Console permissions dry-run: submit an actual declaration for this app category before committing to the storage design | Whether Google approves anything beyond SAF | ARCH §6.9 |
| S9 | Golden-image comparison of one edit graph rendered on desktop and Android; calibrate the achievable tolerance | What R1's tolerance threshold should actually be | R1 |
Tier 2 — before the corresponding subsystem is built
| # | Spike | Answers | Relates to |
|---|---|---|---|
| S3 | reqwest HTTPS PROPFIND on a real Android device, including the rustls-platform-verifier Kotlin init |
The largest known Rust-on-Android networking risk | D7 |
| S4 | Range-extract an embedded JPEG from CR3/NEF/ARW over WebDAV; measure bytes transferred | Whether remote browsing on mobile data is viable | ARCH §6.7, FR-NC-3 |
| S5 | ETag pruning against a 10k-file library: confirm one-request no-op sync, and correct propagation on a single deep-file change | Whether FR-NC-4 scales as designed | ARCH §6.6 |
| S6 | Tiled GPU pipeline on a mid-range Android device, with an image larger than available GPU memory | Whether ARCH §6.2 holds on constrained hardware | NFR-RES-2 |
| S7 | rawler decode coverage across the FR-RAW-1 launch set, on real files from each body | Whether the LibRaw fallback is needed at launch or later | D2, FR-RAW-1 |
| S8 | Chunked upload v2 round-trip of a 100MB RAW, including resume after process kill | FR-NC-7 correctness | FR-NC-7 |
| S12 | GPU device loss recovery: induce VK_ERROR_DEVICE_LOST mid-render, verify recreation from the edit graph with no lost edits |
Whether ARCH §6.10 and NFR-R7 hold | ARCH §6.10 |
| S13 | Slint accessibility on Android: verify TalkBack exposure of names, roles, and values | Whether NFR-A11Y-2 is achievable in the chosen toolkit | NFR-A11Y-2 |
| S14 | Face pipeline in Rust, on a real personal library: run a detector plus an embedder over ~2,000 images through a Rust ONNX runtime, at proxy resolution, on the reference desktop. Measure per-image latency, cluster purity against hand-labelled truth, and fit the FR-CULL-9 calibration to see whether it converges on a library-sized sample. Resolve the model licence question before writing any of it | Whether §3.9.1 is buildable without breaking the pure-Rust dependency policy, and whether the accuracy is worth the subsystem | D13, FR-CULL-8, FR-CULL-9 |
| S15 | Panorama pre-conditions, in order of what can kill it: (1) write a linear DNG with the tiff crate and read it back through rawler — decides FR-MRG-3's container; (2) export XFeat to ONNX at a fixed 1024 px input and load it under tract with zero unsupported operators — the F6 check segmentation.md records, and the licence read first; (3) find where in the fused chain FR-MRG-2's camera-space tap sits and what it costs to expose; (4) a tiled multi-band blend of a 100 MP output on the reference tablet, and XFeat's per-frame time on its CPU — the two halves of NFR-MRG-1 | Whether §3.11 is buildable on the pipeline as it stands, and what the tablet figure is | D18, FR-MRG-2, FR-MRG-3, FR-MRG-8, FR-MRG-11, NFR-MRG-1 |
Why this order
S1, S2, and S10 are the three that can invalidate the architecture. S1 and S2 test ARCH §6.1 — the constraint the whole design is built around, and the one darktable's documentation identifies as their biggest bottleneck. S10 tests whether the Android storage model forced by ARCH §6.9 can actually meet the performance targets; it is the highest-uncertainty assumption in the document because until this revision it was unstated.
S11 costs almost nothing and de-risks S10 definitively. Confirming what Google will approve for this app category before designing around it is far cheaper than discovering it at submission.
S9 moved to Tier 1 because it does not merely test R1 — it calibrates it. R1's tolerance threshold cannot be fixed sensibly without knowing the real cross-vendor deviation, and the §8 golden-image strategy depends on that number.
Test S1 on Mesa/AMD, Intel, and NVIDIA proprietary drivers, under both X11 and Wayland. FD-based external memory has well-documented driver divergence, and the reference machine's discrete GPU will not surface Intel or Mesa-specific issues on its own. The same reasoning is why S2 requires two Android GPU vendors.
10. Glossary
- Edit graph — the ordered set of parameterised operations defining how an image is rendered.
- Proxy — a reduced-resolution render used for display.
- Tile — a sub-rectangle of an image processed independently.
- Demosaic — reconstructing full RGB from a colour-filter-array sensor capture.
- CFA — colour filter array (Bayer, X-Trans).
- Sidecar — a small file alongside the source holding edit metadata.
- Composite — an image produced by a merge (§3.11) from several sources; a source file in its own right under D18.
- Merge — an operation that produces a composite: panorama, HDR merge, focus stacking.
- Pixel pipeline — the ordered chain of processing stages from sensor data to output.