Files
DarkRoom/docs/dev/architecture.md
dtourolle 4bd8d86c00 Record where a linear DNG too large for one texture goes
ARCH §5.3 still described only a tile cache nobody built; it now says
what 0.19.0 tiles and what it does not: a linear DNG past PROXY_EDGE
opens on a reduced copy, a finer render samples a full-resolution
window, and the export is cut into halo-grown tiles. display-and-
extension.md's FR-DSP-2 row said absent, outstanding.md said the halo
had nothing to read it and that such a file fell to the embedded
preview.

rawler now builds from third_party, which ARCH's stack table and §3.2,
the root Cargo.toml's comment ("two upstream crates ... for Android")
and third_party/README.md's bump procedure did not know; the README
also names each vendored crate's licence. panorama.md and the manual
say a composite this wide develops and exports.
2026-09-27 19:42:42 -04:00

65 KiB
Raw Permalink Blame History

DarkRoom — Architecture

Status: Draft v0.1 · 2026-08-08 Companion to: requirements.md

How DarkRoom is built. The requirements document says what the software must do; this says how, and records the decisions and constraints behind the design.


1. Overview

DarkRoom is a Rust application with a Slint interface. On Linux it renders through wgpu to Vulkan. On Android it renders through Skia, also on wgpu's Vulkan swapchain, which a patched wgpu-hal and Skia renderer pre-rotate for a landscape-mounted panel (technical-debt.md TD-1, third_party/). The compute passes are wgpu on both, on the same device the compositor draws with. The design is organised around four ideas, each of which the rest of this document elaborates:

  1. Pixels stay on the GPU. From decode to display, image data never round-trips through the CPU. This is the constraint that most shapes the codebase (§6.1).
  2. Sidecars are authoritative; the catalog is a rebuildable index. Durability comes from plain text files next to the images, not from a database (§6.12).
  3. Operations describe themselves. The develop pipeline emits parameter descriptors as data; the UI generates controls from them. The core never depends on the UI toolkit (§4.3, §6.5a).
  4. Everything is tiled and cancellable. Work is decomposed into tiles scheduled by visibility, so interaction always preempts background work (§5.3).

1.1 Stack

Layer Choice Decision
Language Rust D1
UI Slint D1, D8
GPU wgpu → Vulkan (Linux + Android) D1
Shaders Hand-written WGSL D6
RAW decode rawler (0.7.2, carried patched in third_party/); LibRaw fallback behind a trait D2
Catalog SQLite (WAL) — a rebuildable index D5, §6.12
Colour lcms2 + GPU-side matrix/LUT transforms D5
Network reqwest + quick-xml D7
Licence GPLv3 D8

2. Crate layout

Dependency direction encodes the layering rule: nothing in core/ may depend on anything in ui/.

darkroom/
├── core/
│   ├── dr-types        SourceRef, ImageId, VersionId, ParamValue — shared vocabulary
│   ├── dr-catalog      SQLite index, scan, query, metadata
│   ├── dr-sidecar      the authoritative edit store (§6.12)
│   ├── dr-decode       Decoder trait (§3.2), rawler impl, embedded-preview extraction
│   ├── dr-pipeline     Operation trait, descriptors, edit graph, registry
│   ├── dr-gpu          wgpu device, tile scheduler, WGSL shaders, mask rasteriser
│   ├── dr-colour       lcms2 bindings, camera profiles, working-space transforms
│   ├── dr-export       encoders, resampling, output sizing
│   ├── dr-sync         RemoteBackend + BackendProvider, Account, sync engine, merge
│   ├── dr-sync-nextcloud   WebDAV, oc:fileid, chunked v2, Login Flow v2 (§8.4)
│   └── dr-sync-folder      a plain directory: disk, mount, synced folder (§8.4a)
├── ui/
│   ├── dr-ui           Slint components, adaptive layout, descriptor→control mapping
│   └── dr-widgets      custom controls per WidgetKind (curve, wheel, crop, brush)
├── platform/
│   ├── dr-plat         trait definitions: storage, secrets, lifecycle, power
│   ├── dr-plat-linux   XDG dirs, Secret Service, filesystem
│   └── dr-plat-android SAF, Keystore, WorkManager, lifecycle
└── apps/
    ├── darkroom-desktop
    └── darkroom-android

CI asserts the dependency rule. A UI dependency appearing in any core/ crate's tree fails the build. This is easy to violate accidentally — one use slint:: undoes headless testability and the one-operation-two-presentations property together.

platform/ exists so core/ contains no #[cfg(target_os)]. Platform traits are defined in dr-plat and injected at construction by the app crates.


3. Core abstractions

3.1 SourceRef — addressing image data

Android's Storage Access Framework provides no filesystem path (§6.9), so no core API takes a Path.

/// An opaque, re-resolvable reference to source image data.
pub struct SourceRef(SourceRefInner);

enum SourceRefInner {
    /// Desktop: an absolute path under a granted root.
    Path { root: RootId, relative: PathBuf },
    /// Android: a SAF document URI under a persisted tree grant.
    Document { tree: RootId, document_id: String },
    /// Remote: resolved through the sync layer, possibly byte-range only.
    Remote { file_id: u64, path: String },
}

impl SourceRef {
    /// Open a seekable stream. May fail if the source is offline (FR-CAT-9).
    pub fn open(&self, plat: &dyn Storage) -> Result<Box<dyn SeekableRead>, SourceError>;
    /// Read a byte range without opening the whole source — the fast path for
    /// embedded preview extraction (FR-CULL-2) and remote browsing (FR-NC-3).
    pub fn read_range(&self, plat: &dyn Storage, r: Range<u64>) -> Result<Vec<u8>, SourceError>;
}

read_range is deliberately first-class rather than a convenience over open. Extracting an embedded JPEG preview from a remote 80 MB RAW must transfer ~1–3 MB, and the culling preview ladder depends on the same primitive locally.

3.2 RawDecoder

pub trait RawDecoder: Send + Sync {
    fn probe(&self, header: &[u8]) -> Option<Format>;
    fn embedded_preview(&self, src: &mut dyn SeekableRead) -> Result<Option<Preview>, DecodeError>;
    fn decode(&self, src: &mut dyn SeekableRead) -> Result<RawImage, DecodeError>;
    fn metadata(&self, src: &mut dyn SeekableRead) -> Result<Metadata, DecodeError>;
}

Four separate entry points because the caller's needs differ sharply by phase. Culling wants embedded_preview and nothing else; the grid wants metadata; only develop and export need decode. Fusing them would force full decode where a 200 KB range read suffices.

RawImage carries CFA-pattern sensor data plus black/white levels and camera colour matrices — it is not demosaiced. Demosaic is a GPU pipeline stage (§5.2).

As built (FR-RAW-2, 0.15.0). The trait is dr_decode::Decoder, and it takes bytes rather than a reader: header_bytes says how much of a file metadata needs, metadata and orientation read it, locate_preview says where the embedded preview sits so the caller's storage can fetch that range, and preview and decode take the bytes fetched. The split this section argues for survives; what moved is who reads the file, which is the storage layer (§3.1's read_range), not the decoder. dr_decode::Rawler is the one implementation, and only the places that start a job name dr_decode::default(); everything below them takes a &dyn Decoder.

rawler itself is built from third_party/rawler-0.7.2 since 0.19.0: the crate as published, with its allocation guard raised so that a linear DNG wider than about 16 700 pixels (a stitched panorama) decodes rather than being refused (third_party/README.md).

3.3 Operation and descriptors

The develop pipeline is a sequence of operations with a uniform interface. Polymorphism is by trait, never inheritance — there is no shared implementation to inherit, only a shared shape.

pub trait Operation: Send + Sync {
    /// Static parameter description. Drives UI generation (FR-DEV-3a).
    fn descriptor() -> OpDescriptor where Self: Sized;

    /// Encode GPU work for one tile. No UI types cross this boundary.
    fn encode(&self, enc: &mut ComputeEncoder, ctx: &TileContext);

    /// Identity for cache invalidation. Integer state only — exactly
    /// deterministic, unlike GPU float output (§6.13).
    fn params_hash(&self) -> u64;

    /// What this operation's parameters affect, for invalidation scoping.
    fn affects(&self) -> Affects;
}

Descriptors are pure data:

pub struct ParamDescriptor {
    pub id: ParamId,
    pub label: LocalizedKey,      // a key, not a string — core has no localiser
    pub kind: ParamKind,
    pub default: ParamValue,
    pub affects: Affects,
}

pub enum ParamKind {
    Scalar { min: f32, max: f32, scale: Scale, unit: Unit, precision: u8 },
    Bool,
    Enum { variants: Vec<(EnumId, LocalizedKey)> },
    Colour { has_alpha: bool },
    /// Controls that don't reduce to primitives. Core names the *kind*;
    /// dr-widgets owns the implementation.
    Custom { widget: WidgetKind, schema: CustomParamSchema },
}

LocalizedKey rather than a resolved string is what lets NFR-A11Y-1 work without core/ depending on a localisation library that pulls in UI concerns. The UI resolves keys against its catalogue.

3.4 Edit graph

pub struct EditGraph {
    version: VersionId,
    ops: Vec<OpInstance>,     // ordered; order is data, not code
    masks: Vec<MaskDef>,
}

pub struct OpInstance {
    kind: OpKind,
    enabled: bool,
    params: BTreeMap<ParamId, ParamValue>,
    mask: Option<MaskId>,     // None = global
}

BTreeMap rather than HashMap so serialisation is deterministic — the sidecar format is diffable and the graph hash is stable across runs.

The graph is CPU-side state, which is what makes GPU device-loss recovery tractable (§6.10): the device can be destroyed and rebuilt, and the render re-driven from the graph with nothing lost.


4. Layering

4.1 The four layers

┌──────────────────────────────────────────────────────┐
│  apps/          wiring, platform injection           │
├──────────────────────────────────────────────────────┤
│  ui/            Slint; descriptor → control mapping  │
│                 owns: presentation, gestures, layout │
├──────────────────────────────────────────────────────┤
│  core/          catalog, pipeline, sync, export      │
│                 owns: all behaviour and state        │
├──────────────────────────────────────────────────────┤
│  platform/      storage, secrets, lifecycle, power   │
└──────────────────────────────────────────────────────┘

Calls go downward. core/ reaches platform/ through traits; ui/ reaches core/ through a session API. Nothing calls upward — the UI observes change through subscriptions, not callbacks registered into the core.

4.2 The session API

ui/ sees a small surface, not the internals:

pub struct App { /* … */ }

impl App {
    pub fn catalog(&self) -> &Catalog;
    pub fn open_develop(&self, v: VersionId) -> DevelopSession;
    pub fn open_cull(&self, filter: Filter) -> CullSession;
    pub fn export(&self, sel: &[VersionId], preset: &ExportPreset) -> JobHandle;
    pub fn sync(&self) -> &SyncEngine;
    /// Change notifications. The UI subscribes; the core never holds UI callbacks.
    pub fn subscribe(&self) -> Receiver<CoreEvent>;
}

pub struct DevelopSession { /* … */ }

impl DevelopSession {
    pub fn descriptors(&self) -> &[OpDescriptor];   // drives panel generation
    pub fn set_param(&mut self, op: OpId, p: ParamId, v: ParamValue);
    pub fn set_viewport(&mut self, vp: Viewport);
    /// The rendered result as a GPU texture handle. Never pixels.
    pub fn texture(&self) -> TextureHandle;
    pub fn histogram(&self) -> &HistogramBuffer;    // GPU reduction, not readback
    pub fn undo(&mut self); pub fn redo(&mut self);
}

texture() returning a handle rather than pixel data is the API-level expression of §6.1.

4.3 Descriptor → control mapping

dr-ui maps ParamKind to a control by input modality (FR-DEV-3b). The mapping table lives in the UI; the core is unaware presentation varies.

ParamKind Pointer Touch
Scalar Slider + numeric entry, wheel fine-adjust Drag-strip, double-tap reset
Bool Checkbox Switch, ≥44pt
Enum Dropdown Segmented control or sheet
Colour Swatch → popover Swatch → sheet
Custom dr-widgets control Same control, touch hit-targets

Adding an operation therefore requires no UI change (FR-DEV-3c) unless it needs a new WidgetKind.

4.3a The presentation contract

The division of labour, stated once so neither side drifts:

The core declares capabilities and hints. The frontend composes.

The core says what a parameter is (ParamKind), what an operation would like (Presentation), and what a widget inherently demands (WidgetDemand). It never says what is drawn, where it sits, how wide it is, or whether it is currently visible. Those are compositional decisions and they belong to whoever knows the window, the input modality and the platform — which is never the core.

Hints are plural and ordered. Presentation.widgets is a list of WidgetKind in descending preference. The frontend walks it and takes the first it both implements and can afford. Falling off the end is not an error: every parameter remains an individually addressable scalar, so plain sliders are always the final fallback and the edit still works — merely more tediously. That fallback is why curve points are scalars rather than an opaque blob.

Demands describe the control, not the screen. A hint may carry what the widget inherently needs — two-dimensional direct manipulation, precision pointing, a minimum useful number of simultaneous values. It must never carry pixels, breakpoints, DPI, or a platform name. min_width: 240px in a descriptor is the core making a layout decision, and a core that reasons about pixels will eventually be wrong about a display it never saw. The frontend maps demands onto its own thresholds; those thresholds live in dr-ui and may differ per platform without the core knowing.

Presentation state is derived, never transmitted. Whether a section is collapsed, whether a group contains a modified value, where a heading falls — all of it is computed frontend-side from capabilities plus current values. The core exposing a group_modified flag or a starts_group marker would be the core deciding the panel has groups at all, which is a composition decision. The frontend has the descriptors and the values; that is sufficient to derive any of it.

The test. A second frontend — a CLI, a test harness, a differently-shaped mobile UI — must be able to consume the same capability output and compose something entirely different, without the core changing. If a core change would be needed to lay something out differently, the boundary has been crossed.


5. GPU architecture

5.1 Device and surface

One wgpu::Device shared by the pipeline and Slint, so compute output composites without an interop layer. Slint's create_texture_from_hal imports the texture directly.

Spike S1 validates this, and it is the project's highest-risk assumption. If it fails, D1's recorded fallbacks apply.

5.2 Pipeline stages

Ordered; each consumes and produces GPU textures. Order is data (§3.4), so operations can be reordered without code changes.

RawImage (sensor data, CPU)
   │  upload
   ▼
┌─────────────────────┐
│ hot/dead photosites │  repaired on the mosaic (FR-RAW-3)
├─────────────────────┤
│ black/white levels  │  integer normalise
├─────────────────────┤
│ demosaic            │  Bayer or Markesteijn (X-Trans, FR-RAW-5)
├─────────────────────┤
│ AI denoise          │  optional; raw-domain, joint with demosaic where possible
├─────────────────────┤
│ white balance       │  camera RGB: as-shot, then the operation
├─────────────────────┤
│ camera profile      │  the matrix (FR-DEV-3e) — no curve (D19)
├─────────────────────┤
│ → working space     │  linear, unbounded, f16
├─────────────────────┤
│ exposure/contrast   │
│ highlights/shadows  │  ← masks apply per-op from here down
│ tone curve          │
│ HSL / colour mixer  │
│ texture / clarity   │
│ spot removal        │
│ sharpen / NR        │
│ lens corrections    │
├─────────────────────┤
│ geometry            │  crop, straighten, rotate
├─────────────────────┤
│ view transform      │  sigmoid, or the film stock (FR-DEV-3j)
├─────────────────────┤
│ output transform    │  → display or export profile
└─────────────────────┘
   │
   ├─→ display texture (composited by Slint — never read back)
   └─→ histogram reduction (compute → small buffer, §5.5)

Working precision is f16 in a linear wide-gamut space, quantising once at the output transform.

Scene-referred until the view transform (D19, §6.14). Everything between the matrix and the view transform is linear and unbounded. The view transform is the one stage allowed to compress the scene into a display range. With a detail stage it runs as a dispatch of its own after the detail passes, generated by the same composer as the fused pass so that it gets the mask layers, the film tables and the grain's source position. Without one it is the fused pass's tail. Both the view transform and the output transform (primaries, gamut clip, encode) come after everything that reads a neighbourhood.

A hot or dead photosite is repaired before the demosaic, not after. Past it, one photosite of nonsense is a coloured cross three pixels wide that no later stage can tell from detail. The pass (shaders/hot_pixels.wgsl, run by Demosaicer::run into a second buffer) replaces a photosite that stands apart from every same-colour photosite in its 5×5 window and from each of its eight immediate neighbours with the nearest value that neighbourhood vouches for. The second test is what keeps a star: real light arrives through a lens and lights a patch, so its neighbours are lit too. A 6×6 sensor-anchored colour tile serves Bayer and X-Trans alike, every path that demosaics gets it, and there is no setting.

A mask layer runs inside this chain, not after it. Its settings are offsets to the global ones (mask::offset_onto), and each operation a layer touches is composed at the operation's own place: the global fragment and each layer's combined fragment read the same input, and the pixel moves by each layer's weighted difference, c_g + Σ wᵢ(cᵢ − c_g). At full weight that is the combined setting exactly and at zero the global result exactly, so global contrast −30 under a layer at −20 is contrast −50 where contrast runs, never −30 now and −20 again later — which is what the layer chain did before 0.18.1, and how the shadows of a night shot went magenta. A photograph with no layers composes to the same shader byte for byte. The film is the one exception, because it is a rendering rather than an adjustment and cross-fading two developments is not what a region of a pushed negative looks like: an operation that blends_settings has its uniforms averaged by mask weight instead, and runs once (FR-DEV-3f).

There is a second reduction that does not hang off the bottom of this chain. The raw histogram (FR-CULL-3) taps the demosaiced scene-linear texture directly — the demosaic box's output — because what it measures is the file rather than the render. See §5.5.

5.3 Tiling and scheduling

Work decomposes into tiles (default 256×256) scheduled by priority class:

Class Work Preempts
Interactive Visible tiles at current viewport everything
Prefetch Tiles just outside viewport; next cull image Background
Background Export, thumbnail generation, proxy building —

Interactive strictly preempts Background. Without this, moving a slider during a batch export misses its frame budget — the common case, not an edge case.

Tile results cache keyed by (VersionId, tile, zoom, graph_hash_prefix), where the prefix covers operations up to the first Affects change. Adjusting exposure reuses cached demosaic and camera profile output for every tile.

As built (0.19.0). The interactive path does not tile: one fused dispatch over the viewport is inside the frame budget (frame-budget.md), and there is no scheduler or tile cache. What tiles is a source too large for one texture. A linear DNG whose long edge passes PROXY_EDGE (8192) — a stitched panorama — is held at full resolution on the CPU and opened on a box-reduced copy, from which the canvas at fit, the thumbnail, the histograms and the masks work. A render finer than the copy — the canvas zoomed in, a tile of the export — samples a window cut from the full resolution (DemosaicedImage::linear_rgb16_window), which the fused shader addresses through two source-window uniforms so that crops, warps and grain seeds stay where they are in the frame. The canvas keeps one window while the view stays inside it. The export is cut by dr_pipeline::tiles::plan into 4096-pixel tiles on a 16-pixel grid, each grown by the detail chain's reach (ComposedDetail::reach, the sum of its passes' radii), and reassembled; core/dr-gpu/tests/source_window.rs holds it to the untiled render within one code value. A CFA file too large for one texture is still refused.

5.4 Mask rasterisation

All masks rasterise on the GPU, including drawn brush strokes (§6.11). Strokes arrive as parameters — points, radius, hardness, flow — and a compute shader rasterises them. The mask never exists in CPU memory.

This is a direct response to darktable, where drawn masks are CPU-rasterised and users consistently describe brush lag as making them "unworkable". The problem is architectural, not performance tuning, and not fixable with a faster CPU.

5.5 Histogram without readback

The histogram is a compute-shader reduction into a small storage buffer, read once per frame at most, and only the bins — never image data. A per-frame CPU readback of pixels would reintroduce exactly the stall §6.1 exists to prevent.

Two reductions, not one. The display histogram (FR-DSP-7) counts the frame the output transform produced: its axis is the output code value, and a clipped bin means a highlight that is gone as the image currently stands. The raw histogram for culling (FR-CULL-3) counts the demosaiced scene-linear texture — before white balance, the camera matrix, the tone chain and the view transform — on an axis of stops below sensor saturation, which is how it reports headroom the embedded JPEG's histogram cannot. A culling decision needs the second, an export decision needs the first, and neither answers for the other. Both are drawn by the same panel and chosen between.

The raw reduction runs after the demosaic, not before it. This section previously specified the pre-demosaic CFA samples, and that is the more complete instrument: it counts photosites rather than pixels, so no interpolation smears a clipped site across its neighbours, and it can name which channel of the mosaic saturated first. It was not worth its cost. Demosaicer::run uploads the packed sample buffer and drops it the moment the dispatch is encoded; retaining it is 48 MB at 24 MP and 120 MB at 60 MP, resident per open photograph whether or not anyone looks at the histogram, on a platform §6.2 exists because memory is scarce on. The record here is of the choice, not of the intent — a specification the code contradicts is worse than either of the two things it could say.

The texture that is reduced over instead is camera-native, unbalanced, unmatrixed and uncurved, and is normalised by the sensor's own black and white levels: 1.0 is saturation by construction, so the distribution below it is the headroom question with no calibration to carry and no origin to choose. What it cannot answer, and the CFA reduction could, is which photosite clipped rather than which pixel, and how far above the white level a sample reached — the demosaic clamps there, for its own good reasons, so "at saturation" and "a stop past it" share a bin. Both limits are stated again where the code is, in core/dr-gpu/src/raw_histogram.rs.

It is also the one reduction on this path that is not per frame. Nothing downstream of the demosaic can move a count in it, so it is computed once per photograph and cached — which is what makes it affordable during a cull, where the display histogram's per-frame cost would be paid three thousand times.

5.6 Device loss

Handled as expected, not exceptional (§6.10):

  1. Detect wgpu::SurfaceError::Lost or device-lost callback
  2. Discard all GPU-side state — textures, buffers, pipelines
  3. Recreate device, recompile from the shader cache
  4. Re-drive the current render from the edit graph

No user edit is lost, because the graph is CPU-side. Shader pipelines cache on disk keyed by device, driver version, and shader hash (NFR-P12).


6. Data architecture

6.1 Sidecars are authoritative

The durability model inverts the usual arrangement (§6.12):

image.CR3          ← never written to
image.CR3.drsc     ← authoritative edit state, plain text
                     (or app-managed store where the location is read-only)
catalog.sqlite     ← index; deletable and rebuildable at any time

Sidecar format is versioned plain text holding all Versions for an image:

schema = 1
image_hash = "blake3:..."

[versions.default]
name = "Default"
created = 2026-08-08T14:02:00Z
revision = 7
device = "…"

[versions.default.ops.exposure]
enabled = true
exposure = 0.35
contrast = 12.0

Writing discipline, each point addressing a documented failure elsewhere:

  • Atomic — temp plus rename, never partial
  • Debounced — not per slider tick
  • Only on real content change — darktable rewrites sidecars without edits, breaking backup deduplication and mtime-based sync
  • Never embedded in the RAW — would force re-backup of the whole file

6.2 The catalog as index

SQLite in WAL mode, holding what is needed for interactive query over 50k images and nothing authoritative:

roots(id, kind, grant_blob, label, last_seen)
images(id, root_id, source_ref, content_hash, format, w, h,
       captured_at, camera, lens, iso, aperture, shutter,
       availability, sidecar_mtime)
versions(id, image_id, name, is_default, graph_hash, rating, label, flag)
keywords(version_id, keyword)
folders(id, root_id, path, etag, parent_id)      -- ETag pruning (§6.6)
remote(image_id, file_id, etag, sync_state, remote_path)
cache(version_id, kind, resolution, graph_hash, path, bytes, last_used)
jobs(id, kind, state, payload, attempts)         -- resumable across process death

folders.etag must exist from schema v1 — sync pruning cannot work without persisted folder ETags, and adding it later means a migration plus a full re-scan of every library (§6.6).

Rebuild reads sidecars and re-derives everything else. This makes NFR-R6's recovery path the normal mechanism rather than a last resort.

6.3 Identity

Entity Key Stable across
Image content hash rename, move, re-import
Version UUID everything
Remote file oc:fileid server-side rename/move
Cache entry (version, kind, resolution, graph_hash) —
Person UUID rename, merge, re-index

Content-hash identity is what makes FR-CAT-9's reconnection work: a moved file is recognised rather than re-imported, and a server-side move is not a re-download of 80 MB.

A person's identity is their UUID, never their name: renaming "Mum" to "Sarah" must not create a second person, and two devices that name the same face independently must be mergeable rather than duplicated (FR-CULL-12). Same reasoning as collections, same mechanism.

6.4 People and faces

Specified by FR-CULL-8 … FR-CULL-12 and NFR-SEC-5. Sketched here because the split across the trust boundary is an architectural decision, not a schema detail.

Three entities. A face is a detection: an image, a box, landmarks, a detector confidence, and an embedding. A person is a UUID and a name. face_person links them, carrying a calibrated probability and — critically — a confirmed flag separating what the user asserted from what the system guessed.

That flag is the whole design. Suggestions are derived data and may be recomputed at will; a confirmation is a user judgement and is never overwritten by a later inference pass. Conflating them would mean a model upgrade silently rewriting the user's own labelling, which is the kind of loss this architecture exists to prevent.

Where each part lives, and why they differ:

Data Home Rebuildable Rationale
Embeddings, boxes, landmarks Catalog only Yes — re-index Expensive but reproducible. ARCH §6.12: derived data belongs in the disposable index.
Cluster assignments, suggestions Catalog only Yes Inference output; changes whenever the model or calibration does.
Confirmed person name on an image Sidecar No A human judgement, same class as a rating or keyword (FR-CAT-8). Must survive catalog deletion.
Person entity (UUID, name) Catalog, synced No The merge identity; travels with collections under FR-CAT-7's rules.

The asymmetry is deliberate: deleting the catalog costs an afternoon of re-indexing and loses nothing the user typed. That is the same bargain §6.12 already makes everywhere else.

Detection runs on proxies, not originals. FR-CULL-8 pins this to the FR-CULL-2 preview ladder, so face indexing consumes the same artefacts the grid already built rather than forcing RAW decodes. The consequence for §5's GPU budget is that face inference competes with thumbnailing, not with rendering, and NFR-ARCH-2's priority classes already express that.

Embeddings never enter the diagnostics or crash paths (NFR-SEC-5). This is a structural exclusion, not a redaction rule: NFR-OPS-1's bundle is assembled from an allowlist, so a new table does not silently become uploadable by existing.


7. Concurrency

7.1 Executors

Executor Threads Work
UI 1 Slint event loop. Never blocks.
GPU submit 1 Command encoding and queue submission
Decode cores − 2 RAW decode, preview extraction
I/O 4 Catalog, sidecar, cache, filesystem
Network 2 Nextcloud transfer

The UI executor never blocks — this is the mechanism behind R4 and NFR-P9, which the requirements state as outcomes without saying how.

In code: ui/dr-ui/src/executors.rs. Executor names the five and states each count with its reason (Executor::threads); executors::spawn(executor, role, f) starts a worker named <executor>:<role> and marks it with its executor. run marks its own thread as the UI executor before the window exists, and executors::assert_not_ui — called by net_runtime's block_on — fails a debug or test build that blocks there. The counts are not yet enforced: a job still gets a thread of its own, and pooling by executor is where NFR-ARCH-2's priority classes will live.

7.2 Cancellation

Cooperative, with tokens threaded through every long operation. Observed within 100 ms (NFR-ARCH-3), including in-flight GPU submissions. Cancelling a scan, export, or sync leaves no partial state beyond what is separately resumable.

7.3 Errors

No worker panics the process. Errors are typed and attach to the affected image or job:

pub enum ImageError { Decode(DecodeError), SourceOffline, Gpu(GpuError), … }

A decode failure marks one image and continues the batch (FR-RAW-4). A GPU error triggers §5.6 recovery. Panics in decode are caught at the boundary, since RAW parsing handles untrusted input.


8. Sync architecture

Sync is pluggable. dr-sync defines a RemoteBackend trait and the capability model the engine adapts to; two connectors implement it — dr-sync-nextcloud and dr-sync-folder — and S3, generic WebDAV, or a self-hosted photo server can be added without touching the engine.

docs/storage.md is the contract: the four traits a connector meets, the four steps to add one, and what each shipped connector actually declares. This section says why the seam is shaped the way it is; that document says how to use it.

8.0 A trait is not a seam

Worth stating because this was got wrong for a release. RemoteBackend existed from the start and the code above it still knew it was talking to Nextcloud: seven files in dr-ui constructed a NextcloudBackend directly, ten functions took one by concrete type, an account was a server URL beside a DAV user id, and the local cache directory was named after a hostname. The abstraction was real and bought nothing.

Pluggable storage needs four things, and only the first is a trait over operations:

What Where
1 Operations RemoteBackend (§8.3)
2 Capabilities — what is cheap, so the engine adapts rather than assumes Capabilities (§8.1)
3 Configuration — what an account is, with no server in it Account, Connection
4 Registration — how a connector is discovered without being named BackendProvider, BackendRegistry

ui/dr-ui/src/remote.rs is the only file above dr-sync that names a connector. dr-sync itself depends on none of them, so a build that only wants a folder library does not compile a TLS stack to get one.

8.1 Why capability negotiation, not a common denominator

The obvious abstraction — a trait exposing only what every backend can do — would be a mistake here, and it is worth being explicit about why.

Nextcloud's fast path depends on a behaviour that is not a WebDAV guarantee: it propagates ETag changes up the folder tree, so an unchanged root ETag proves nothing anywhere in the library changed. That single property is what turns a 50k-image no-op sync into one HTTP request. A plain WebDAV server (Apache mod_dav, an SFTP mount, a raw S3 bucket) offers no such guarantee, and a trait built to their common subset would force full enumeration on every sync — the same 50k requests the design exists to avoid.

So backends declare capabilities, and the sync engine picks the best available strategy:

pub struct Capabilities {
    /// How the backend reports what changed. Determines sync cost.
    pub change_detection: ChangeDetection,
    /// Stable identity across server-side rename/move.
    pub stable_ids: bool,
    /// Byte-range reads — required for embedded-preview extraction.
    pub range_reads: bool,
    /// Resumable upload for large files.
    pub chunked_upload: Option<ChunkConstraints>,
    /// Many-small-files upload in one request (sidecars).
    pub bulk_upload: bool,
    /// Conditional write for optimistic concurrency.
    pub conditional_write: bool,
    /// Server-rendered thumbnails, if any.
    pub server_previews: ServerPreviews,
}

pub enum ChangeDetection {
    /// Backend hands us a cursor; we ask "what changed since?".
    /// Cheapest possible. (No current backend, but the shape most
    /// object stores and delta APIs fit.)
    DeltaCursor,
    /// Directory ETags propagate upward — unchanged parent proves
    /// unchanged subtree. Nextcloud. One request for a no-op sync.
    PropagatingEtags,
    /// ETags exist but only on the entry itself. Must enumerate the
    /// tree; ETags then avoid re-downloading unchanged content.
    LocalEtags,
    /// Nothing but modification times. Enumerate and compare.
    Timestamps,
}

8.2 Strategy per capability tier

The engine has one strategy per ChangeDetection variant, chosen at connect time:

Tier Detection No-op sync cost, 50k images Notes
1 DeltaCursor 1 request Cursor persisted in remote_state
2 PropagatingEtags 1 request Nextcloud. Root ETag unchanged → done
3 LocalEtags ~1 per folder Enumerate tree; ETags prevent re-download
4 Timestamps ~1 per folder + clock-skew risk Degraded; warn the user

Tiers 3 and 4 are honest degradations, not silent ones — the UI reports the sync strategy in use so a slow backend is visibly slow rather than mysteriously slow.

Capability absence never breaks correctness, only speed — with two exceptions the engine must handle explicitly:

  • No range_reads → embedded-preview extraction is impossible, so remote browsing must fall back to server previews or full download. On a metered connection the engine refuses full downloads for browsing and reports why.
  • No conditional_write → sidecar conflict detection loses its guarantee. The engine falls back to revision-counter comparison inside the sidecar body, which narrows but does not close the race. This is reported as a reduced-safety mode.

8.3 The trait

#[async_trait]
pub trait RemoteBackend: Send + Sync {
    fn capabilities(&self) -> &Capabilities;
    fn connect(&mut self, auth: AuthHandle) -> Result<Identity, RemoteError>;

    // ---- discovery -------------------------------------------------
    /// One level. `since` carries a per-entry validator (ETag/mtime)
    /// so unchanged entries can be skipped by the backend where possible.
    async fn list(&self, dir: &RemotePath, since: Option<&Validator>)
        -> Result<Vec<RemoteEntry>, RemoteError>;

    /// Tier 1 only — capability-gated, returns Unsupported otherwise.
    async fn delta(&self, cursor: &Cursor)
        -> Result<(Vec<RemoteChange>, Cursor), RemoteError>;

    /// Tier 2 only — cheap validator for a directory, without listing it.
    async fn dir_validator(&self, dir: &RemotePath)
        -> Result<Validator, RemoteError>;

    // ---- transfer --------------------------------------------------
    async fn get(&self, id: &RemoteId, range: Option<Range<u64>>)
        -> Result<Bytes, RemoteError>;
    async fn put(&self, path: &RemotePath, body: Body, precond: Option<Precondition>)
        -> Result<Validator, RemoteError>;
    async fn put_many(&self, items: Vec<(RemotePath, Bytes)>)
        -> Result<Vec<Result<Validator, RemoteError>>, RemoteError>;
    async fn delete(&self, id: &RemoteId, precond: Option<Precondition>)
        -> Result<(), RemoteError>;

    // ---- optional --------------------------------------------------
    async fn thumbnail(&self, id: &RemoteId, size: u32)
        -> Result<Option<Bytes>, RemoteError>;
}

Design notes worth keeping:

  • get takes an optional range rather than having a separate get_range. Callers always express what they need; backends without range support return the whole object and the engine slices, so correctness holds while the capability flag tells the caller whether it was cheap.
  • put takes a Precondition, not a bare ETag — so IfMatch, IfNoneMatch, and None are all expressible, and a backend without conditional writes can reject at compile-time-visible runtime rather than silently ignoring.
  • put_many is a required method with a default implementation looping over put. Nextcloud overrides it with bulk upload; other backends get correct behaviour for free.
  • Chunked upload is not in the trait. It is an implementation detail of put — the backend decides based on body size and its own constraints. Exposing it would leak Nextcloud's protocol into the interface.

8.4 The Nextcloud connector

The reference implementation, and the one whose peculiarities the capability model exists to keep. Mapping to the trait:

Trait method Nextcloud
capabilities PropagatingEtags, stable ids, ranges, chunked (5MB–5GB), bulk, conditional
list PROPFIND Depth:1 with oc:fileid, getetag, nc:has-preview
dir_validator PROPFIND Depth:0 requesting getetag only
delta Unsupported — no RFC 6578 for files (ARCH §6.6)
get + range GET with Range:; detect support by 206 vs 200, never HEAD
put large Chunked v2: MKCOL upload dir → PUT chunks → MOVE .file
put_many POST /remote.php/dav/bulk, multipart/related
thumbnail /core/preview?fileId=…&forceIcon=false — false is mandatory, or a server that cannot render RAW returns a generic mimetype icon that we would cache as a thumbnail
auth Login Flow v2, system browser, app password

Pruning walk, the Tier 2 strategy in full:

dir_validator(root)
  ├─ unchanged vs stored → done. One request, whole library.
  └─ changed → list(dir, since)
       ├─ child dirs whose validator differs → recurse
       └─ files whose validator differs → queue transfer

folders.etag must exist from schema v1. Adding it later means a migration plus a full re-scan of every user's library (ARCH §6.6).

Validated against a real library, 2026-08-09. A cold recursive scan of a 17,185-RAW library (7,836 CR2 + 9,349 DNG) across 334 directories completed in 34.1 s — Depth: 1 per directory, never Depth: infinity. Two implementation details worth keeping:

  • The format filter sees through VFS placeholder suffixes, so a dehydrated IMG.CR2.nextcloud matches as the CR2 it stands for rather than being skipped as an unknown type (§9.0).
  • Pruning is capability-gated, not assumed. With LocalEtags a directory probe costs a request and proves nothing about children, so it is pure overhead; a test asserts zero probes in that case. Only PropagatingEtags makes an unchanged parent prove an unchanged subtree.

8.4a The folder connector

A plain directory: a local disk, a network mount, an external drive, or the folder a Nextcloud desktop client already syncs. No server, no account, no credential — which makes it the route that works on a machine with no secrets daemon.

It exists for two reasons. It is genuinely useful, and a second implementation is the only way to find out whether the first was an abstraction or a description. Adding it is what turned §8.0's four items from a claim into a fact.

Trait method Folder
capabilities LocalEtags, no stable ids, ranges, no chunking, conditional
list read_dir + metadata; validator is size-mtime
dir_validator Unsupported — a directory's mtime does not propagate
delta Unsupported — a folder keeps no change feed
get + range seek + take; a short read past the end is not an error
put write to a temporary beside the destination, rename over it
put IfAbsent `O_CREAT
put IfMatch stat, compare, then the same rename — narrows the race, does not close it
delete files, and empty directories only — see below
move_to rename, falling back to copy + unlink across a mount boundary
auth none

Three things worth carrying forward:

  • LocalEtags is the honest answer, and it costs nothing. A POSIX directory's mtime describes its own entry list and nothing below it, so there is no propagation to exploit and the engine walks the tree every scan. Measured 2026-08-28: a full uncached walk of 2,299 images across 233 directories took 137 ms, with no pruning at all — against 34.1 s for 17,185 RAWs over WebDAV with pruning (§8.4). The walk that was expensive was expensive because it was thousands of PROPFINDs. This is the capability model paying for itself — one engine, two backends, each running at the speed it actually runs at.
  • Identity is a path hash, not an inode. An inode is stable across a rename but differs between devices and is reused after a delete, so two machines would disagree about which photograph a thumbnail belonged to and a recycled inode would attach an old thumbnail to a new image. Re-deriving a thumbnail is a cost; showing the wrong one is a bug. stable_ids: false reports the consequence.
  • delete is deliberately not recursive, unlike WebDAV's DELETE on a collection. There is no server-side trash behind a local folder, so a caller with a wrong path would have no way back. Nothing in the engine deletes a directory — the soft delete is a move_to (§FR-CAT-15) — so the guard is free.

8.5 Sidecar conflict resolution

put with Precondition::IfMatch(validator). On precondition failure:

  1. get the remote sidecar
  2. Merge per Version, per operation — a crop on one device and an exposure change on another both survive, because they touch disjoint operations
  3. Genuinely conflicting operations resolve by device timestamp
  4. Retry with the new validator, bounded retry count

No (conflicted copy) files. Sidecars are a few KB of structured data, so read-merge-rewrite is cheap and preserves intent — a meaningful improvement over file-level conflict copies, which is what the sidecar-authoritative model (ARCH §6.12) buys.


9. Caching and availability policy

Sync decides what changed. Caching policy decides what is kept locally — a separate concern, and the one that determines whether the app is usable on a tablet with a 2 TB library behind it.

9.0 Why not the Nextcloud client's Virtual Files

An appealing shortcut: sync the library with the official desktop client in VFS mode and let DarkRoom read ordinary paths, inheriting their sync engine, upload, and conflict handling for free.

Measured on this machine (client 4.0.7, 2026-08-09) and rejected. The configured folder uses virtualFilesMode=suffix, the only mode Linux supports, holding 121,785 placeholders against 10,267 materialised files — including 7,037 CR2 and 9,411 DNG placeholders.

Three findings, each independently disqualifying:

  1. A dehydrated file exists only under a different name. IMG.CR2 is absent; only IMG.CR2.nextcloud exists, containing exactly one byte. Any extension-based scan sees .nextcloud, so the app needs placeholder-aware code regardless — VFS is not transparent.

  2. Reads do not hydrate, but hydration can be requested. Reading a stub returns its one byte and triggers nothing — there is no FUSE layer intercepting reads. However the client exposes a local socket at $XDG_RUNTIME_DIR/Nextcloud/socket speaking a newline-delimited COMMAND:argument protocol, and MAKE_AVAILABLE_LOCALLY:<path> does fetch the file. Verified 2026-08-09: a 1-byte .nextcloud stub was replaced by the real 2.7 MB file within seconds. MAKE_ONLINE_ONLY dehydrates again.

  3. Even with hydration, granularity is wrong. VFS has two states, 1 byte or all bytes. The preview tier — the one that makes remote browsing viable on mobile data — needs a ~256 KB prefix of a 27 MB file. A hydrating VFS would transfer ~100× what FR-NC-3 requires, which is precisely the cost range extraction exists to avoid.

What this changes, and what it does not. Hydration-on-request makes VFS a usable original tier: for an image the user opens in develop or exports, asking the client to fetch it is a legitimate alternative to fetching it ourselves, and it inherits their transfer, resume and conflict handling for free.

It does not rescue the preview tier, which is the one that matters for browsing. Finding 3 stands: hydration is whole-file, so filling a grid still costs the entire library. Range extraction remains the only mechanism that satisfies FR-NC-3.

Design consequence. VFS is supported as an optional source, not as the transfer layer:

  • Recognise *.nextcloud stubs and report them as Availability::Offline (FR-NC-6c) rather than as corrupt files.
  • Where a library lives under a VFS-synced folder, offer "download" on a stub by writing MAKE_AVAILABLE_LOCALLY:<path> to the socket, rather than fetching a second copy over WebDAV and leaving the client's own state inconsistent.
  • Never depend on it: the socket is Linux-only, absent on Android, and absent when the client is not running. The direct connector remains the primary path.

9.0a Amendment, 2026-08-29 — VFS as a managed tier for folder libraries

The rejection above was written when the only backend was the direct Nextcloud connector, and it assumed the alternative to hydration was a range read. Since dr-sync-folder exists (§8.4a) a library can be opened as a directory with no server connection at all, and that changes what finding 3 is comparing against.

What finding 3 actually says. Hydration transfers ~100× what a preview needs when a range read is available. On a folder library there is no connector, so there are no range reads: the choice is not hydrate-versus-range-read, it is hydrate-once or never have a thumbnail. The finding stands unamended for the direct connector, which must still never hydrate to browse.

What makes the cost acceptable on a folder library, and all three are required:

  1. Hydration is a borrow, not an acquisition. A file is returned to the state it was found in — what the pass downloaded is released, what the user already had is left alone. Peak disk is the working set, not the library. Measured 2026-08-29 over 100 photographs of 25 MB, 90 of them dehydrated: peak 275 MB against 2,500 MB unborrowed, back to 250 MB afterwards, and all ten the user already held still there.
  2. It is paid once. Thumbnails are kept, and derived_sync pushes the shards to the server, so a second device downloads 200 MB of shards instead of hydrating 340 GB of RAWs.
  3. It is quoted and consented to. Never automatic, never on the browsing path, always resumable and cancellable (FR-NC-6c).

What this does not change. Findings 1 and 2 stand and are now implemented rather than merely noted: a stub is not transparent, so RemoteEntry carries materialised and the backend reports RemoteError::NotMaterialised rather than a miss; and the socket remains the only way to hydrate, so it stays Linux-only, optional, and absent on Android.

The one thing that must never be got wrong. Releasing content means asking the client to dehydrate — never deleting the file. A deletion inside a synced tree propagates to the server and removes the photograph from every device the user owns. RemoteBackend::dematerialise says so, the Vfs trait says so, and an implementation that cannot dehydrate returns Unsupported rather than approximating it with remove_file.

9.1 The three tiers, restated as policy

Tier Content Default
Metadata Catalog rows, sidecars, ratings Always synced; kilobytes; even on metered
Preview Embedded JPEGs, display proxies LRU within a size cap
Original Full RAW Never by default — explicit pin or on-demand

Never bulk-sync originals. A 2 TB library against a 128 GB tablet makes mirror-style sync unusable (ARCH §6.8) — which is precisely where the official Nextcloud desktop client's model fails.

9.2 Cache rules

A cache rule pins a set of images at a tier. Rules are user-authored, evaluated against the catalog, and re-evaluated as it changes — so "the last three months" stays current without the user touching it.

pub struct CacheRule {
    pub id: RuleId,
    pub selector: Selector,
    pub tier: Tier,            // Preview | Original
    pub priority: u8,          // eviction order; higher survives longer
    pub enabled: bool,
}

pub enum Selector {
    Collection(CollectionId),
    Folder { root: RootId, path: PathBuf, recursive: bool },
    /// Absolute or rolling. `Rolling` re-evaluates daily.
    DateRange(DateSelector),
    Rating { min: u8 },
    Label(ColourLabel),
    Flag(FlagState),
    Keyword(String),
    /// Boolean composition, so "5-star OR flagged, in the last year" works.
    All(Vec<Selector>),
    Any(Vec<Selector>),
    Not(Box<Selector>),
}

pub enum DateSelector {
    Between { from: Date, to: Date },
    /// "Last 90 days" — window moves with the clock.
    Rolling { days: u32 },
    /// "This trip" — bounded by a collection's own capture range.
    CollectionSpan(CollectionId),
}

Worked examples of what this expresses:

Intent Rule
Trip on the tablet Collection(trip) → Original
Recent work at full quality Rolling{90} → Original
Portfolio always available Rating{min:5} → Original, high priority
Everything browsable All → Preview (the implicit default rule)
Client edit on the road All([Collection(job), Flag(Pick)]) → Original

9.3 Evaluation and eviction

Rules produce a desired state per image; the sync engine reconciles actual against desired:

for each image:
    desired = max(tier of every matching enabled rule)   // most generous wins
    if desired > actual → enqueue fetch  (priority = rule priority)
    if desired < actual → mark evictable (not evicted immediately)

Eviction is lazy and only under pressure — cache cap reached (NFR-RES-4) or platform memory pressure (FR-PLAT-AND-5). Order: unpinned originals by last-used, then proxies, then thumbnails. Never metadata or sidecars — those are authoritative (ARCH §6.12) and tiny.

An image that leaves a rolling window is not deleted at midnight; it becomes the first candidate when space is actually needed. Fetching is likewise opportunistic: queued at the rule's priority behind anything interactive, and on Android gated by the unmetered-and-charging constraints in FR-NC-6.

9.4 Honesty about what is local

Adobe's sync fails on legibility more than transport: Lightroom Classic syncs only 2560px proxies while displaying the original's filename, extension, and size, so users genuinely do not know what they have. DarkRoom shall not repeat this.

  • Every image carries a visible availability state: Original · Preview · Metadata only · Offline
  • Operations needing data that is absent say so before starting, with the transfer size
  • Export from a preview-only image is refused, not silently degraded
  • A pinned set reports its true byte cost before the user commits to it

9.5 Storage

cache_rules(id, selector_json, tier, priority, enabled)
image_cache(image_id, tier_actual, tier_desired, bytes, last_used, pinned_by_rule)

tier_desired is materialised rather than recomputed per query, so the grid can render availability badges without evaluating every rule for every visible cell.


10. Platform abstraction

pub trait Storage: Send + Sync {
    fn roots(&self) -> Vec<RootId>;
    fn enumerate(&self, root: RootId, cb: &mut dyn FnMut(SourceRef)) -> Result<(), StorageError>;
    fn open(&self, src: &SourceRef) -> Result<Box<dyn SeekableRead>, StorageError>;
    fn read_range(&self, src: &SourceRef, r: Range<u64>) -> Result<Vec<u8>, StorageError>;
    fn writable_sidecar_dir(&self, src: &SourceRef) -> Option<PathBuf>;
}

pub trait Secrets: Send + Sync { /* Secret Service | Android Keystore */ }
pub trait Lifecycle: Send + Sync { /* process death, memory pressure, background */ }
Concern Linux Android
Storage Filesystem under granted roots SAF tree URIs, DocumentsContract
Secrets Secret Service (libsecret) Keystore + DataStore + Tink
Background Threads WorkManager, foreground service
Lifecycle Process lives Death at any moment; onTrimMemory

writable_sidecar_dir returns None on read-only mounts and most SAF trees, which is when the app-managed sidecar store is used.


11. Build order

Independent of D12's resolution, the foundations are shared. What differs is what gets built on top first.

Phase 0 — spikes. S1 (Slint+wgpu zero-copy), S2 (Android, two GPU vendors), S10 (SAF at 10k files), S11 (Play permissions), S9 (cross-platform tolerance). These can invalidate the architecture; everything else assumes they pass.

Phase 1 — foundations, needed either way. dr-types, dr-plat + both implementations, dr-catalog, dr-sidecar, dr-decode (preview path first), dr-gpu device and tile scheduler.

Phase 2 — depends on D12:

  • Culler-first: preview ladder, culling mode, raw histogram, focus peaking, ratings, burst grouping. Ships something usable without any develop chain.
  • Vertical-slice-first: minimal develop chain (exposure, curve), export, proving the full architecture end to end but not yet useful.

Phase 3+: the remaining develop operations, masking, sync, ingest, AI denoise.


12. Architectural constraints

Decisions treated as fixed because reversing them is expensive. Each is evidence-backed.

On the numbering. The subsections below are numbered 6.1–6.13, not 12.x. They were §6 when this document was first written, and every ARCH §6.n citation in requirements.md, in docs/*.md and in source comments names them by those numbers. Renumbering would break several hundred citations for no gain, so the numbers stay: ARCH §6.1 means the first constraint here, and never §6.1 of the data architecture above, which is cited as §6 Data architecture or by its title.

6.1 GPU results never round-trip through the CPU

darktable identifies this as their single biggest bottleneck: OpenCL output returns to the GTK thread, which paints it on one CPU core. It constrains the UI framework choice more than any other requirement — whatever renders the UI must composite our GPU texture directly.

Measured on this project (Radeon RX 7900 XTX, cargo run -p dr-gpu --example bench --features readback), comparing the compute pass alone against compute plus a CPU readback:

Canvas Compute With readback Readback share
840×692 0.06 ms 0.63 ms 90%
1920×1080 0.11 ms 2.35 ms 95%
3840×2160 0.28 ms 7.43 ms 96%

At 4K the shader finishes in 0.28 ms and then 7.15 ms is spent moving pixels through the CPU — a 26× overhead that scales with area, which is why an uncapped window resize falls off a cliff. The constraint is not a stylistic preference; it is the dominant cost in the frame.

No exceptions since 0.15.0. Android read the develop frame back until then, because wgpu's Vulkan swapchain could not pre-rotate and a portrait window tore on the tablet. Two local patches removed the need (technical-debt.md TD-1), so the frame reaches the compositor as a texture on both platforms. The constraint governs every display path. The export pipeline reads pixels back by design, because a file is its output.

6.2 Tiling from day one

Mobile GPUs have far less memory. Retrofitting tiling into a whole-image pipeline is a rewrite.

6.3 Non-destructive edit graph, GPU-resident

Recompute from the graph rather than caching CPU bitmaps between stages.

6.4 GPU-first, not GPU-accelerated

RawTherapee shows excellent quality is achievable on CPU — and that retrofitting GPU into a mature CPU pipeline effectively never happens.

6.5 The catalog is the UI's source of truth for queries

The UI never touches storage or network synchronously.

6.5a The core has no UI dependency

Operations describe parameters as data. This buys headless golden-image testing, one-operation-two- presentations, and UI replaceability. CI-enforced.

6.6 Sync cannot rely on server-side change tokens

Verified: Nextcloud's Directory.php does not implement ISyncCollection; sync tokens exist only for CalDAV/CardDAV (nextcloud/server#22584, open since 2020). Folder ETags must be persisted from schema v1.

6.7 Server-side RAW previews are not assumable

Verified: no RAW preview provider ships with Nextcloud. Range-based embedded extraction is the primary path; server previews are opportunistic. preview_max_filesize_image defaults to 50 MB, excluding many RAWs even where a provider exists.

Measured against a real instance, 2026-08-09, and the case is stronger than assumed. /core/preview returned HTTP 400 for every parameter combination attempted — including bare ?fileId=N, and including a JPEG the same server reported as nc:has-preview=true. The cause is unexplained: it is a server-side preview configuration issue, not a request-shape error on our side.

Recorded as unexplained rather than understood, because the distinction matters if someone later tries to depend on this endpoint. Nothing does today — the connector treats a preview failure as a miss and falls through to range extraction, which is the designed behaviour rather than a workaround.

Range extraction validated on the same library. 262 KB read from a 21.5 MB DNG in 119 ms — 1.22% of the file — yielded camera model and ISO. Extrapolated across the 17,185-file test library, cataloguing by whole-file fetch would move roughly 370 GB; the range path moves a few MB. That ratio is the difference between a viable mobile experience and an unusable one, and it is why §6.7 treats range reads as the mechanism and server previews as a bonus.

6.8 Sync is selective, not mirror-style

6.9 Android forbids the filesystem-scan model

Verified: MANAGE_EXTERNAL_STORAGE is not grantable — Play policy explicitly disallows "media files access" as a permitted use. READ_MEDIA_IMAGES would not help: proprietary RAW is not typed image/* by the platform scanner, so it does not appear in MediaStore.Images. SAF is the only path, and it provides no filesystem path — hence SourceRef (§3.1).

6.10 GPU device loss is expected

Routine on Android: backgrounding, driver resets, thermal events. Recovery is tractable because the edit graph is CPU-side.

6.11 Drawn masks are GPU-rasterised

darktable rasterises parametric masks on GPU and drawn masks on CPU; users describe drawn masking as "unworkable" and the lag is not fixable with more cores. Highest-value architectural win available relative to the incumbents, and free if designed in now.

6.12 Sidecars authoritative, catalog disposable

darktable maintains both a database and sidecars while achieving the reliability of neither — documented lockouts, corruption where all three recovery paths failed, sidecars silently not written in 5.0.0. RawTherapee's .pp3 is the model: plain text, next to the file, no database in the trust path.

6.13 Bit-identity applies to integer state only

GPU float results diverge across vendors — transcendental implementations differ, drivers optimise differently, f16 rounding varies. Cache keys and graph hashes are computed over CPU-side integer state, which is exactly deterministic. Cross-platform rendering equality is a bounded tolerance (R1), not a checksum.

6.14 Scene-referred until the view transform

Added 2026-09-27 (D19). Between the camera matrix and the view transform, values are scene-linear and unbounded, and no operation clamps above 1.0, applies a transfer function or maps to a display gamut. The view transform (FR-DEV-3j) is the one stage that does, and the output transform after it clips and encodes. This is the constraint the base curve broke and ARCH §5.2 had drawn all along: a display-referred curve in the middle of the chain throws away what every later stage, the neighbourhood ones above all, needs. A test (scene_referred_until_the_view, in dr-gpu) runs every point operation over a ramp to 16.0 so that a fragment that clips fails the build rather than the photograph.


13. Decisions

Full rationale in requirements.md §8. Summary:

# Decision Status
D1 Rust + Slint + wgpu Decided
D2 rawler, LibRaw fallback behind a trait Decided
D3 First milestone Provisional — depends on D12
D4 Nextcloud: ETag pruning, chunked upload, Login Flow v2 Resolved
D5 lcms2 + GPU-side transforms Decided
D6 Hand-written WGSL Decided
D7 reqwest + quick-xml Decided
D8 GPLv3 Decided
D9 Declarative operation descriptors Decided
D10 Single adaptive interface Decided
D11 Product positioning Decided
D12 Scope versus pace Open
D13 Face inference runtime and model licensing Runtime answered, reopened for per-device backends (docs/inference.md); licensing open
D14 Segmentation source for local masking Decided — arm C (docs/segmentation.md §14)
D15 Target devices — 12-inch tablet and desktop, no phone Decided (requirements D15)
D19 Scene-referred pipeline, one view transform last Decided (requirements D19)

14. Open questions

D12 — scope versus pace. The selected feature set implies years of full-time work against a stated evenings-and-weekends budget. Resolving it sets D3's milestone and §11's Phase 2.

NFR-R8 — CPU fallback extent. §6.4 says GPU-first; NFR-RES-2 assumes a CPU fallback on allocation failure. Decide whether that means a full CPU pipeline (a second implementation) or only tile-spill staging. Recommend the latter.

Slint accessibility on Android. Unverified; may be a toolkit gap (NFR-A11Y-2). Cheaper to discover before the UI is built.

Play Store versus GPLv3. Generally workable but should be confirmed, with F-Droid as fallback.