docs/segmentation.md §4 priced arm B as costing a C dependency under the NDK and treated that as most of the difference between the arms. It is not a cost that has to be paid: `ort`'s `alternative-backend` disables its linking entirely and `ort-tract` supplies the API from tract, which is pure Rust. D13's "largest exception the policy would tolerate" turns out not to be needed, and the answer generalises to the face pipeline — so D13's runtime half is now answered and only its licensing half is open. Three findings contradict §4 outright and are recorded as F4-F6 rather than quietly designed around. There is no ADE20K-trained YOLO, so the shipped vocabulary selects subjects and not stuff — "select the sky" comes from the watershed or from nowhere. It is instance segmentation, so it partitions nothing and two people come back as two instances. And tract cannot parse a dynamic-shape export, which fixes the input at 640 square and makes tiling the only route to more semantic resolution. Arm C ships, but §8's criteria are not what decided it, and saying so matters more than claiming the process worked. §8 asked for a two- interaction margin over arm A on a traced corpus. That comparison was never run: F4 and F5 changed what the arms are, and a model that recognises subjects but has no word for sky cannot be a selection tool alone, while a watershed cannot tell a person from the wall behind them. They stopped being candidates and became complements. What is *not* done is written down as plainly: the 24-image corpus is untraced, so M1-M4 have no numbers and "this feels right" has not become one. M5 is answered on one device only, and region ids now reach the sidecar — so a cross-vendor divergence would mean a mask written on the desktop meaning something else on Android. F3 stands.
48 KiB
DarkRoom — Architecture
Status: Draft v0.1 · 2026-08-08 Companion to: requirements.md
How DarkRoom is built. The requirements document says what the software must do; this says how, and records the decisions and constraints behind the design.
1. Overview
DarkRoom is a Rust application with a Slint interface, rendering through wgpu to Vulkan on both Linux and Android. The design is organised around four ideas, each of which the rest of this document elaborates:
- Pixels stay on the GPU. From decode to display, image data never round-trips through the CPU. This is the constraint that most shapes the codebase (§6.1).
- Sidecars are authoritative; the catalog is a rebuildable index. Durability comes from plain text files next to the images, not from a database (§6.12).
- Operations describe themselves. The develop pipeline emits parameter descriptors as data; the UI generates controls from them. The core never depends on the UI toolkit (§4.3, §6.5a).
- Everything is tiled and cancellable. Work is decomposed into tiles scheduled by visibility, so interaction always preempts background work (§5.3).
1.1 Stack
| Layer | Choice | Decision |
|---|---|---|
| Language | Rust | D1 |
| UI | Slint | D1, D8 |
| GPU | wgpu → Vulkan (Linux + Android) | D1 |
| Shaders | Hand-written WGSL | D6 |
| RAW decode | rawler; LibRaw fallback behind a trait | D2 |
| Catalog | SQLite (WAL) — a rebuildable index | D5, §6.12 |
| Colour | lcms2 + GPU-side matrix/LUT transforms | D5 |
| Network | reqwest + quick-xml | D7 |
| Licence | GPLv3 | D8 |
2. Crate layout
Dependency direction encodes the layering rule: nothing in core/ may depend on anything in ui/.
darkroom/
├── core/
│ ├── dr-types SourceRef, ImageId, VersionId, ParamValue — shared vocabulary
│ ├── dr-catalog SQLite index, scan, query, metadata
│ ├── dr-sidecar the authoritative edit store (§6.12)
│ ├── dr-decode RawDecoder trait, rawler impl, embedded-preview extraction
│ ├── dr-pipeline Operation trait, descriptors, edit graph, registry
│ ├── dr-gpu wgpu device, tile scheduler, WGSL shaders, mask rasteriser
│ ├── dr-colour lcms2 bindings, camera profiles, working-space transforms
│ ├── dr-export encoders, resampling, output sizing
│ ├── dr-sync RemoteBackend trait, sync engine, cache rules, merge
│ └── dr-sync-nextcloud the only backend implementation (§8.4)
├── ui/
│ ├── dr-ui Slint components, adaptive layout, descriptor→control mapping
│ └── dr-widgets custom controls per WidgetKind (curve, wheel, crop, brush)
├── platform/
│ ├── dr-plat trait definitions: storage, secrets, lifecycle, power
│ ├── dr-plat-linux XDG dirs, Secret Service, filesystem
│ └── dr-plat-android SAF, Keystore, WorkManager, lifecycle
└── apps/
├── darkroom-desktop
└── darkroom-android
CI asserts the dependency rule. A UI dependency appearing in any core/ crate's tree fails the
build. This is easy to violate accidentally — one use slint:: undoes headless testability and the
one-operation-two-presentations property together.
platform/ exists so core/ contains no #[cfg(target_os)]. Platform traits are defined in
dr-plat and injected at construction by the app crates.
3. Core abstractions
3.1 SourceRef — addressing image data
Android's Storage Access Framework provides no filesystem path (§6.9), so no core API takes a
Path.
/// An opaque, re-resolvable reference to source image data.
pub struct SourceRef(SourceRefInner);
enum SourceRefInner {
/// Desktop: an absolute path under a granted root.
Path { root: RootId, relative: PathBuf },
/// Android: a SAF document URI under a persisted tree grant.
Document { tree: RootId, document_id: String },
/// Remote: resolved through the sync layer, possibly byte-range only.
Remote { file_id: u64, path: String },
}
impl SourceRef {
/// Open a seekable stream. May fail if the source is offline (FR-CAT-9).
pub fn open(&self, plat: &dyn Storage) -> Result<Box<dyn SeekableRead>, SourceError>;
/// Read a byte range without opening the whole source — the fast path for
/// embedded preview extraction (FR-CULL-2) and remote browsing (FR-NC-3).
pub fn read_range(&self, plat: &dyn Storage, r: Range<u64>) -> Result<Vec<u8>, SourceError>;
}
read_range is deliberately first-class rather than a convenience over open. Extracting an
embedded JPEG preview from a remote 80 MB RAW must transfer ~1–3 MB, and the culling preview ladder
depends on the same primitive locally.
3.2 RawDecoder
pub trait RawDecoder: Send + Sync {
fn probe(&self, header: &[u8]) -> Option<Format>;
fn embedded_preview(&self, src: &mut dyn SeekableRead) -> Result<Option<Preview>, DecodeError>;
fn decode(&self, src: &mut dyn SeekableRead) -> Result<RawImage, DecodeError>;
fn metadata(&self, src: &mut dyn SeekableRead) -> Result<Metadata, DecodeError>;
}
Four separate entry points because the caller's needs differ sharply by phase. Culling wants
embedded_preview and nothing else; the grid wants metadata; only develop and export need
decode. Fusing them would force full decode where a 200 KB range read suffices.
RawImage carries CFA-pattern sensor data plus black/white levels and camera colour matrices — it
is not demosaiced. Demosaic is a GPU pipeline stage (§5.2).
3.3 Operation and descriptors
The develop pipeline is a sequence of operations with a uniform interface. Polymorphism is by trait, never inheritance — there is no shared implementation to inherit, only a shared shape.
pub trait Operation: Send + Sync {
/// Static parameter description. Drives UI generation (FR-DEV-3a).
fn descriptor() -> OpDescriptor where Self: Sized;
/// Encode GPU work for one tile. No UI types cross this boundary.
fn encode(&self, enc: &mut ComputeEncoder, ctx: &TileContext);
/// Identity for cache invalidation. Integer state only — exactly
/// deterministic, unlike GPU float output (§6.13).
fn params_hash(&self) -> u64;
/// What this operation's parameters affect, for invalidation scoping.
fn affects(&self) -> Affects;
}
Descriptors are pure data:
pub struct ParamDescriptor {
pub id: ParamId,
pub label: LocalizedKey, // a key, not a string — core has no localiser
pub kind: ParamKind,
pub default: ParamValue,
pub affects: Affects,
}
pub enum ParamKind {
Scalar { min: f32, max: f32, scale: Scale, unit: Unit, precision: u8 },
Bool,
Enum { variants: Vec<(EnumId, LocalizedKey)> },
Colour { has_alpha: bool },
/// Controls that don't reduce to primitives. Core names the *kind*;
/// dr-widgets owns the implementation.
Custom { widget: WidgetKind, schema: CustomParamSchema },
}
LocalizedKey rather than a resolved string is what lets NFR-A11Y-1 work without core/ depending
on a localisation library that pulls in UI concerns. The UI resolves keys against its catalogue.
3.4 Edit graph
pub struct EditGraph {
version: VersionId,
ops: Vec<OpInstance>, // ordered; order is data, not code
masks: Vec<MaskDef>,
}
pub struct OpInstance {
kind: OpKind,
enabled: bool,
params: BTreeMap<ParamId, ParamValue>,
mask: Option<MaskId>, // None = global
}
BTreeMap rather than HashMap so serialisation is deterministic — the sidecar format is diffable
and the graph hash is stable across runs.
The graph is CPU-side state, which is what makes GPU device-loss recovery tractable (§6.10): the device can be destroyed and rebuilt, and the render re-driven from the graph with nothing lost.
4. Layering
4.1 The four layers
┌──────────────────────────────────────────────────────┐
│ apps/ wiring, platform injection │
├──────────────────────────────────────────────────────┤
│ ui/ Slint; descriptor → control mapping │
│ owns: presentation, gestures, layout │
├──────────────────────────────────────────────────────┤
│ core/ catalog, pipeline, sync, export │
│ owns: all behaviour and state │
├──────────────────────────────────────────────────────┤
│ platform/ storage, secrets, lifecycle, power │
└──────────────────────────────────────────────────────┘
Calls go downward. core/ reaches platform/ through traits; ui/ reaches core/ through a
session API. Nothing calls upward — the UI observes change through subscriptions, not callbacks
registered into the core.
4.2 The session API
ui/ sees a small surface, not the internals:
pub struct App { /* … */ }
impl App {
pub fn catalog(&self) -> &Catalog;
pub fn open_develop(&self, v: VersionId) -> DevelopSession;
pub fn open_cull(&self, filter: Filter) -> CullSession;
pub fn export(&self, sel: &[VersionId], preset: &ExportPreset) -> JobHandle;
pub fn sync(&self) -> &SyncEngine;
/// Change notifications. The UI subscribes; the core never holds UI callbacks.
pub fn subscribe(&self) -> Receiver<CoreEvent>;
}
pub struct DevelopSession { /* … */ }
impl DevelopSession {
pub fn descriptors(&self) -> &[OpDescriptor]; // drives panel generation
pub fn set_param(&mut self, op: OpId, p: ParamId, v: ParamValue);
pub fn set_viewport(&mut self, vp: Viewport);
/// The rendered result as a GPU texture handle. Never pixels.
pub fn texture(&self) -> TextureHandle;
pub fn histogram(&self) -> &HistogramBuffer; // GPU reduction, not readback
pub fn undo(&mut self); pub fn redo(&mut self);
}
texture() returning a handle rather than pixel data is the API-level expression of §6.1.
4.3 Descriptor → control mapping
dr-ui maps ParamKind to a control by input modality (FR-DEV-3b). The mapping table lives in the
UI; the core is unaware presentation varies.
ParamKind |
Pointer | Touch |
|---|---|---|
Scalar |
Slider + numeric entry, wheel fine-adjust | Drag-strip, double-tap reset |
Bool |
Checkbox | Switch, ≥44pt |
Enum |
Dropdown | Segmented control or sheet |
Colour |
Swatch → popover | Swatch → sheet |
Custom |
dr-widgets control |
Same control, touch hit-targets |
Adding an operation therefore requires no UI change (FR-DEV-3c) unless it needs a new WidgetKind.
4.3a The presentation contract
The division of labour, stated once so neither side drifts:
The core declares capabilities and hints. The frontend composes.
The core says what a parameter is (ParamKind), what an operation would
like (Presentation), and what a widget inherently demands
(WidgetDemand). It never says what is drawn, where it sits, how wide it is,
or whether it is currently visible. Those are compositional decisions and they
belong to whoever knows the window, the input modality and the platform —
which is never the core.
Hints are plural and ordered. Presentation.widgets is a list of
WidgetKind in descending preference. The frontend walks it and takes the
first it both implements and can afford. Falling off the end is not an error:
every parameter remains an individually addressable scalar, so plain sliders
are always the final fallback and the edit still works — merely more tediously.
That fallback is why curve points are scalars rather than an opaque blob.
Demands describe the control, not the screen. A hint may carry what the
widget inherently needs — two-dimensional direct manipulation, precision
pointing, a minimum useful number of simultaneous values. It must never carry
pixels, breakpoints, DPI, or a platform name. min_width: 240px in a
descriptor is the core making a layout decision, and a core that reasons about
pixels will eventually be wrong about a display it never saw. The frontend maps
demands onto its own thresholds; those thresholds live in dr-ui and may
differ per platform without the core knowing.
Presentation state is derived, never transmitted. Whether a section is
collapsed, whether a group contains a modified value, where a heading falls —
all of it is computed frontend-side from capabilities plus current values. The
core exposing a group_modified flag or a starts_group marker would be the
core deciding the panel has groups at all, which is a composition decision. The
frontend has the descriptors and the values; that is sufficient to derive any
of it.
The test. A second frontend — a CLI, a test harness, a differently-shaped mobile UI — must be able to consume the same capability output and compose something entirely different, without the core changing. If a core change would be needed to lay something out differently, the boundary has been crossed.
5. GPU architecture
5.1 Device and surface
One wgpu::Device shared by the pipeline and Slint, so compute output composites without an
interop layer. Slint's create_texture_from_hal imports the texture directly.
Spike S1 validates this, and it is the project's highest-risk assumption. If it fails, D1's recorded fallbacks apply.
5.2 Pipeline stages
Ordered; each consumes and produces GPU textures. Order is data (§3.4), so operations can be reordered without code changes.
RawImage (sensor data, CPU)
│ upload
▼
┌─────────────────────┐
│ black/white levels │ integer normalise
├─────────────────────┤
│ demosaic │ Bayer or Markesteijn (X-Trans, FR-RAW-5)
├─────────────────────┤
│ AI denoise │ optional; raw-domain, joint with demosaic where possible
├─────────────────────┤
│ camera profile │ matrices + per-body base curve (FR-DEV-3e)
├─────────────────────┤
│ → working space │ linear, wide-gamut, f16
├─────────────────────┤
│ white balance │
│ exposure/contrast │
│ highlights/shadows │ ← masks apply per-op from here down
│ tone curve │
│ HSL / colour mixer │
│ texture / clarity │
│ spot removal │
│ sharpen / NR │
│ lens corrections │
│ look (HaldCLUT) │ FR-DEV-3f
├─────────────────────┤
│ geometry │ crop, straighten, rotate
├─────────────────────┤
│ output transform │ → display or export profile
└─────────────────────┘
│
├─→ display texture (composited by Slint — never read back)
└─→ histogram reduction (compute → small buffer, §5.5)
Working precision is f16 in a linear wide-gamut space, quantising once at the output transform.
5.3 Tiling and scheduling
Work decomposes into tiles (default 256×256) scheduled by priority class:
| Class | Work | Preempts |
|---|---|---|
Interactive |
Visible tiles at current viewport | everything |
Prefetch |
Tiles just outside viewport; next cull image | Background |
Background |
Export, thumbnail generation, proxy building | — |
Interactive strictly preempts Background. Without this, moving a slider during a batch
export misses its frame budget — the common case, not an edge case.
Tile results cache keyed by (VersionId, tile, zoom, graph_hash_prefix), where the prefix covers
operations up to the first Affects change. Adjusting exposure reuses cached demosaic and camera
profile output for every tile.
5.4 Mask rasterisation
All masks rasterise on the GPU, including drawn brush strokes (§6.11). Strokes arrive as parameters — points, radius, hardness, flow — and a compute shader rasterises them. The mask never exists in CPU memory.
This is a direct response to darktable, where drawn masks are CPU-rasterised and users consistently describe brush lag as making them "unworkable". The problem is architectural, not performance tuning, and not fixable with a faster CPU.
5.5 Histogram without readback
The histogram is a compute-shader reduction into a small storage buffer, read once per frame at most, and only the bins — never image data. A per-frame CPU readback of pixels would reintroduce exactly the stall §6.1 exists to prevent.
Raw-domain histograms for culling (FR-CULL-3) reduce over the pre-demosaic texture, which is why they can report headroom the embedded JPEG's histogram cannot.
5.6 Device loss
Handled as expected, not exceptional (§6.10):
- Detect
wgpu::SurfaceError::Lostor device-lost callback - Discard all GPU-side state — textures, buffers, pipelines
- Recreate device, recompile from the shader cache
- Re-drive the current render from the edit graph
No user edit is lost, because the graph is CPU-side. Shader pipelines cache on disk keyed by device, driver version, and shader hash (NFR-P12).
6. Data architecture
6.1 Sidecars are authoritative
The durability model inverts the usual arrangement (§6.12):
image.CR3 ← never written to
image.CR3.drsc ← authoritative edit state, plain text
(or app-managed store where the location is read-only)
catalog.sqlite ← index; deletable and rebuildable at any time
Sidecar format is versioned plain text holding all Versions for an image:
schema = 1
image_hash = "blake3:..."
[versions.default]
name = "Default"
created = 2026-08-08T14:02:00Z
revision = 7
device = "…"
[versions.default.ops.exposure]
enabled = true
exposure = 0.35
contrast = 12.0
Writing discipline, each point addressing a documented failure elsewhere:
- Atomic — temp plus rename, never partial
- Debounced — not per slider tick
- Only on real content change — darktable rewrites sidecars without edits, breaking backup deduplication and mtime-based sync
- Never embedded in the RAW — would force re-backup of the whole file
6.2 The catalog as index
SQLite in WAL mode, holding what is needed for interactive query over 50k images and nothing authoritative:
roots(id, kind, grant_blob, label, last_seen)
images(id, root_id, source_ref, content_hash, format, w, h,
captured_at, camera, lens, iso, aperture, shutter,
availability, sidecar_mtime)
versions(id, image_id, name, is_default, graph_hash, rating, label, flag)
keywords(version_id, keyword)
folders(id, root_id, path, etag, parent_id) -- ETag pruning (§6.6)
remote(image_id, file_id, etag, sync_state, remote_path)
cache(version_id, kind, resolution, graph_hash, path, bytes, last_used)
jobs(id, kind, state, payload, attempts) -- resumable across process death
folders.etag must exist from schema v1 — sync pruning cannot work without persisted folder ETags,
and adding it later means a migration plus a full re-scan of every library (§6.6).
Rebuild reads sidecars and re-derives everything else. This makes NFR-R6's recovery path the normal mechanism rather than a last resort.
6.3 Identity
| Entity | Key | Stable across |
|---|---|---|
| Image | content hash | rename, move, re-import |
| Version | UUID | everything |
| Remote file | oc:fileid |
server-side rename/move |
| Cache entry | (version, kind, resolution, graph_hash) |
— |
| Person | UUID | rename, merge, re-index |
Content-hash identity is what makes FR-CAT-9's reconnection work: a moved file is recognised rather than re-imported, and a server-side move is not a re-download of 80 MB.
A person's identity is their UUID, never their name: renaming "Mum" to "Sarah" must not create a second person, and two devices that name the same face independently must be mergeable rather than duplicated (FR-CULL-12). Same reasoning as collections, same mechanism.
6.4 People and faces
Specified by FR-CULL-8 … FR-CULL-12 and NFR-SEC-5. Sketched here because the split across the trust boundary is an architectural decision, not a schema detail.
Three entities. A face is a detection: an image, a box, landmarks, a detector confidence, and
an embedding. A person is a UUID and a name. face_person links them, carrying a calibrated
probability and — critically — a confirmed flag separating what the user asserted from what the
system guessed.
That flag is the whole design. Suggestions are derived data and may be recomputed at will; a confirmation is a user judgement and is never overwritten by a later inference pass. Conflating them would mean a model upgrade silently rewriting the user's own labelling, which is the kind of loss this architecture exists to prevent.
Where each part lives, and why they differ:
| Data | Home | Rebuildable | Rationale |
|---|---|---|---|
| Embeddings, boxes, landmarks | Catalog only | Yes — re-index | Expensive but reproducible. ARCH §6.12: derived data belongs in the disposable index. |
| Cluster assignments, suggestions | Catalog only | Yes | Inference output; changes whenever the model or calibration does. |
| Confirmed person name on an image | Sidecar | No | A human judgement, same class as a rating or keyword (FR-CAT-8). Must survive catalog deletion. |
| Person entity (UUID, name) | Catalog, synced | No | The merge identity; travels with collections under FR-CAT-7's rules. |
The asymmetry is deliberate: deleting the catalog costs an afternoon of re-indexing and loses nothing the user typed. That is the same bargain §6.12 already makes everywhere else.
Detection runs on proxies, not originals. FR-CULL-8 pins this to the FR-CULL-2 preview ladder, so face indexing consumes the same artefacts the grid already built rather than forcing RAW decodes. The consequence for §5's GPU budget is that face inference competes with thumbnailing, not with rendering, and NFR-ARCH-2's priority classes already express that.
Embeddings never enter the diagnostics or crash paths (NFR-SEC-5). This is a structural exclusion, not a redaction rule: NFR-OPS-1's bundle is assembled from an allowlist, so a new table does not silently become uploadable by existing.
7. Concurrency
7.1 Executors
| Executor | Threads | Work |
|---|---|---|
| UI | 1 | Slint event loop. Never blocks. |
| GPU submit | 1 | Command encoding and queue submission |
| Decode | cores − 2 | RAW decode, preview extraction |
| I/O | 4 | Catalog, sidecar, cache, filesystem |
| Network | 2 | Nextcloud transfer |
The UI executor never blocks — this is the mechanism behind R4 and NFR-P9, which the requirements state as outcomes without saying how.
7.2 Cancellation
Cooperative, with tokens threaded through every long operation. Observed within 100 ms (NFR-ARCH-3), including in-flight GPU submissions. Cancelling a scan, export, or sync leaves no partial state beyond what is separately resumable.
7.3 Errors
No worker panics the process. Errors are typed and attach to the affected image or job:
pub enum ImageError { Decode(DecodeError), SourceOffline, Gpu(GpuError), … }
A decode failure marks one image and continues the batch (FR-RAW-4). A GPU error triggers §5.6 recovery. Panics in decode are caught at the boundary, since RAW parsing handles untrusted input.
8. Sync architecture
Sync is pluggable. dr-sync defines a RemoteBackend trait; only the Nextcloud connector is
implemented, but the boundary is designed so S3, generic WebDAV, or a self-hosted photo server can
be added without touching the sync engine.
8.1 Why capability negotiation, not a common denominator
The obvious abstraction — a trait exposing only what every backend can do — would be a mistake here, and it is worth being explicit about why.
Nextcloud's fast path depends on a behaviour that is not a WebDAV guarantee: it propagates ETag
changes up the folder tree, so an unchanged root ETag proves nothing anywhere in the library
changed. That single property is what turns a 50k-image no-op sync into one HTTP request. A
plain WebDAV server (Apache mod_dav, an SFTP mount, a raw S3 bucket) offers no such guarantee, and
a trait built to their common subset would force full enumeration on every sync — the same 50k
requests the design exists to avoid.
So backends declare capabilities, and the sync engine picks the best available strategy:
pub struct Capabilities {
/// How the backend reports what changed. Determines sync cost.
pub change_detection: ChangeDetection,
/// Stable identity across server-side rename/move.
pub stable_ids: bool,
/// Byte-range reads — required for embedded-preview extraction.
pub range_reads: bool,
/// Resumable upload for large files.
pub chunked_upload: Option<ChunkConstraints>,
/// Many-small-files upload in one request (sidecars).
pub bulk_upload: bool,
/// Conditional write for optimistic concurrency.
pub conditional_write: bool,
/// Server-rendered thumbnails, if any.
pub server_previews: ServerPreviews,
}
pub enum ChangeDetection {
/// Backend hands us a cursor; we ask "what changed since?".
/// Cheapest possible. (No current backend, but the shape most
/// object stores and delta APIs fit.)
DeltaCursor,
/// Directory ETags propagate upward — unchanged parent proves
/// unchanged subtree. Nextcloud. One request for a no-op sync.
PropagatingEtags,
/// ETags exist but only on the entry itself. Must enumerate the
/// tree; ETags then avoid re-downloading unchanged content.
LocalEtags,
/// Nothing but modification times. Enumerate and compare.
Timestamps,
}
8.2 Strategy per capability tier
The engine has one strategy per ChangeDetection variant, chosen at connect time:
| Tier | Detection | No-op sync cost, 50k images | Notes |
|---|---|---|---|
| 1 | DeltaCursor |
1 request | Cursor persisted in remote_state |
| 2 | PropagatingEtags |
1 request | Nextcloud. Root ETag unchanged → done |
| 3 | LocalEtags |
~1 per folder | Enumerate tree; ETags prevent re-download |
| 4 | Timestamps |
~1 per folder + clock-skew risk | Degraded; warn the user |
Tiers 3 and 4 are honest degradations, not silent ones — the UI reports the sync strategy in use so a slow backend is visibly slow rather than mysteriously slow.
Capability absence never breaks correctness, only speed — with two exceptions the engine must handle explicitly:
- No
range_reads→ embedded-preview extraction is impossible, so remote browsing must fall back to server previews or full download. On a metered connection the engine refuses full downloads for browsing and reports why. - No
conditional_write→ sidecar conflict detection loses its guarantee. The engine falls back to revision-counter comparison inside the sidecar body, which narrows but does not close the race. This is reported as a reduced-safety mode.
8.3 The trait
#[async_trait]
pub trait RemoteBackend: Send + Sync {
fn capabilities(&self) -> &Capabilities;
fn connect(&mut self, auth: AuthHandle) -> Result<Identity, RemoteError>;
// ---- discovery -------------------------------------------------
/// One level. `since` carries a per-entry validator (ETag/mtime)
/// so unchanged entries can be skipped by the backend where possible.
async fn list(&self, dir: &RemotePath, since: Option<&Validator>)
-> Result<Vec<RemoteEntry>, RemoteError>;
/// Tier 1 only — capability-gated, returns Unsupported otherwise.
async fn delta(&self, cursor: &Cursor)
-> Result<(Vec<RemoteChange>, Cursor), RemoteError>;
/// Tier 2 only — cheap validator for a directory, without listing it.
async fn dir_validator(&self, dir: &RemotePath)
-> Result<Validator, RemoteError>;
// ---- transfer --------------------------------------------------
async fn get(&self, id: &RemoteId, range: Option<Range<u64>>)
-> Result<Bytes, RemoteError>;
async fn put(&self, path: &RemotePath, body: Body, precond: Option<Precondition>)
-> Result<Validator, RemoteError>;
async fn put_many(&self, items: Vec<(RemotePath, Bytes)>)
-> Result<Vec<Result<Validator, RemoteError>>, RemoteError>;
async fn delete(&self, id: &RemoteId, precond: Option<Precondition>)
-> Result<(), RemoteError>;
// ---- optional --------------------------------------------------
async fn thumbnail(&self, id: &RemoteId, size: u32)
-> Result<Option<Bytes>, RemoteError>;
}
Design notes worth keeping:
gettakes an optional range rather than having a separateget_range. Callers always express what they need; backends without range support return the whole object and the engine slices, so correctness holds while the capability flag tells the caller whether it was cheap.puttakes aPrecondition, not a bare ETag — soIfMatch,IfNoneMatch, andNoneare all expressible, and a backend without conditional writes can reject at compile-time-visible runtime rather than silently ignoring.put_manyis a required method with a default implementation looping overput. Nextcloud overrides it with bulk upload; other backends get correct behaviour for free.- Chunked upload is not in the trait. It is an implementation detail of
put— the backend decides based on body size and its own constraints. Exposing it would leak Nextcloud's protocol into the interface.
8.4 The Nextcloud connector
The only implementation. Mapping to the trait:
| Trait method | Nextcloud |
|---|---|
capabilities |
PropagatingEtags, stable ids, ranges, chunked (5MB–5GB), bulk, conditional |
list |
PROPFIND Depth:1 with oc:fileid, getetag, nc:has-preview |
dir_validator |
PROPFIND Depth:0 requesting getetag only |
delta |
Unsupported — no RFC 6578 for files (ARCH §6.6) |
get + range |
GET with Range:; detect support by 206 vs 200, never HEAD |
put large |
Chunked v2: MKCOL upload dir → PUT chunks → MOVE .file |
put_many |
POST /remote.php/dav/bulk, multipart/related |
thumbnail |
/core/preview?fileId=…&forceIcon=false — false is mandatory, or a server that cannot render RAW returns a generic mimetype icon that we would cache as a thumbnail |
| auth | Login Flow v2, system browser, app password |
Pruning walk, the Tier 2 strategy in full:
dir_validator(root)
├─ unchanged vs stored → done. One request, whole library.
└─ changed → list(dir, since)
├─ child dirs whose validator differs → recurse
└─ files whose validator differs → queue transfer
folders.etag must exist from schema v1. Adding it later means a migration plus a full re-scan of
every user's library (ARCH §6.6).
Validated against a real library, 2026-08-09. A cold recursive scan of a 17,185-RAW library
(7,836 CR2 + 9,349 DNG) across 334 directories completed in 34.1 s — Depth: 1 per directory,
never Depth: infinity. Two implementation details worth keeping:
- The format filter sees through VFS placeholder suffixes, so a dehydrated
IMG.CR2.nextcloudmatches as the CR2 it stands for rather than being skipped as an unknown type (§9.0). - Pruning is capability-gated, not assumed. With
LocalEtagsa directory probe costs a request and proves nothing about children, so it is pure overhead; a test asserts zero probes in that case. OnlyPropagatingEtagsmakes an unchanged parent prove an unchanged subtree.
8.5 Sidecar conflict resolution
put with Precondition::IfMatch(validator). On precondition failure:
getthe remote sidecar- Merge per Version, per operation — a crop on one device and an exposure change on another both survive, because they touch disjoint operations
- Genuinely conflicting operations resolve by device timestamp
- Retry with the new validator, bounded retry count
No (conflicted copy) files. Sidecars are a few KB of structured data, so read-merge-rewrite is
cheap and preserves intent — a meaningful improvement over file-level conflict copies, which is what
the sidecar-authoritative model (ARCH §6.12) buys.
9. Caching and availability policy
Sync decides what changed. Caching policy decides what is kept locally — a separate concern, and the one that determines whether the app is usable on a tablet with a 2 TB library behind it.
9.0 Why not the Nextcloud client's Virtual Files
An appealing shortcut: sync the library with the official desktop client in VFS mode and let DarkRoom read ordinary paths, inheriting their sync engine, upload, and conflict handling for free.
Measured on this machine (client 4.0.7, 2026-08-09) and rejected. The configured folder uses
virtualFilesMode=suffix, the only mode Linux supports, holding 121,785 placeholders against
10,267 materialised files — including 7,037 CR2 and 9,411 DNG placeholders.
Three findings, each independently disqualifying:
-
A dehydrated file exists only under a different name.
IMG.CR2is absent; onlyIMG.CR2.nextcloudexists, containing exactly one byte. Any extension-based scan sees.nextcloud, so the app needs placeholder-aware code regardless — VFS is not transparent. -
Reads do not hydrate, but hydration can be requested. Reading a stub returns its one byte and triggers nothing — there is no FUSE layer intercepting reads. However the client exposes a local socket at
$XDG_RUNTIME_DIR/Nextcloud/socketspeaking a newline-delimitedCOMMAND:argumentprotocol, andMAKE_AVAILABLE_LOCALLY:<path>does fetch the file. Verified 2026-08-09: a 1-byte.nextcloudstub was replaced by the real 2.7 MB file within seconds.MAKE_ONLINE_ONLYdehydrates again. -
Even with hydration, granularity is wrong. VFS has two states, 1 byte or all bytes. The preview tier — the one that makes remote browsing viable on mobile data — needs a ~256 KB prefix of a 27 MB file. A hydrating VFS would transfer ~100× what FR-NC-3 requires, which is precisely the cost range extraction exists to avoid.
What this changes, and what it does not. Hydration-on-request makes VFS a usable original tier: for an image the user opens in develop or exports, asking the client to fetch it is a legitimate alternative to fetching it ourselves, and it inherits their transfer, resume and conflict handling for free.
It does not rescue the preview tier, which is the one that matters for browsing. Finding 3 stands: hydration is whole-file, so filling a grid still costs the entire library. Range extraction remains the only mechanism that satisfies FR-NC-3.
Design consequence. VFS is supported as an optional source, not as the transfer layer:
- Recognise
*.nextcloudstubs and report them asAvailability::Offline(FR-NC-6c) rather than as corrupt files. - Where a library lives under a VFS-synced folder, offer "download" on a stub by writing
MAKE_AVAILABLE_LOCALLY:<path>to the socket, rather than fetching a second copy over WebDAV and leaving the client's own state inconsistent. - Never depend on it: the socket is Linux-only, absent on Android, and absent when the client is not running. The direct connector remains the primary path.
9.1 The three tiers, restated as policy
| Tier | Content | Default |
|---|---|---|
| Metadata | Catalog rows, sidecars, ratings | Always synced; kilobytes; even on metered |
| Preview | Embedded JPEGs, display proxies | LRU within a size cap |
| Original | Full RAW | Never by default — explicit pin or on-demand |
Never bulk-sync originals. A 2 TB library against a 128 GB tablet makes mirror-style sync unusable (ARCH §6.8) — which is precisely where the official Nextcloud desktop client's model fails.
9.2 Cache rules
A cache rule pins a set of images at a tier. Rules are user-authored, evaluated against the catalog, and re-evaluated as it changes — so "the last three months" stays current without the user touching it.
pub struct CacheRule {
pub id: RuleId,
pub selector: Selector,
pub tier: Tier, // Preview | Original
pub priority: u8, // eviction order; higher survives longer
pub enabled: bool,
}
pub enum Selector {
Collection(CollectionId),
Folder { root: RootId, path: PathBuf, recursive: bool },
/// Absolute or rolling. `Rolling` re-evaluates daily.
DateRange(DateSelector),
Rating { min: u8 },
Label(ColourLabel),
Flag(FlagState),
Keyword(String),
/// Boolean composition, so "5-star OR flagged, in the last year" works.
All(Vec<Selector>),
Any(Vec<Selector>),
Not(Box<Selector>),
}
pub enum DateSelector {
Between { from: Date, to: Date },
/// "Last 90 days" — window moves with the clock.
Rolling { days: u32 },
/// "This trip" — bounded by a collection's own capture range.
CollectionSpan(CollectionId),
}
Worked examples of what this expresses:
| Intent | Rule |
|---|---|
| Trip on the tablet | Collection(trip) → Original |
| Recent work at full quality | Rolling{90} → Original |
| Portfolio always available | Rating{min:5} → Original, high priority |
| Everything browsable | All → Preview (the implicit default rule) |
| Client edit on the road | All([Collection(job), Flag(Pick)]) → Original |
9.3 Evaluation and eviction
Rules produce a desired state per image; the sync engine reconciles actual against desired:
for each image:
desired = max(tier of every matching enabled rule) // most generous wins
if desired > actual → enqueue fetch (priority = rule priority)
if desired < actual → mark evictable (not evicted immediately)
Eviction is lazy and only under pressure — cache cap reached (NFR-RES-4) or platform memory pressure (FR-PLAT-AND-5). Order: unpinned originals by last-used, then proxies, then thumbnails. Never metadata or sidecars — those are authoritative (ARCH §6.12) and tiny.
An image that leaves a rolling window is not deleted at midnight; it becomes the first candidate when space is actually needed. Fetching is likewise opportunistic: queued at the rule's priority behind anything interactive, and on Android gated by the unmetered-and-charging constraints in FR-NC-6.
9.4 Honesty about what is local
Adobe's sync fails on legibility more than transport: Lightroom Classic syncs only 2560px proxies while displaying the original's filename, extension, and size, so users genuinely do not know what they have. DarkRoom shall not repeat this.
- Every image carries a visible availability state:
Original·Preview·Metadata only·Offline - Operations needing data that is absent say so before starting, with the transfer size
- Export from a preview-only image is refused, not silently degraded
- A pinned set reports its true byte cost before the user commits to it
9.5 Storage
cache_rules(id, selector_json, tier, priority, enabled)
image_cache(image_id, tier_actual, tier_desired, bytes, last_used, pinned_by_rule)
tier_desired is materialised rather than recomputed per query, so the grid can render availability
badges without evaluating every rule for every visible cell.
10. Platform abstraction
pub trait Storage: Send + Sync {
fn roots(&self) -> Vec<RootId>;
fn enumerate(&self, root: RootId, cb: &mut dyn FnMut(SourceRef)) -> Result<(), StorageError>;
fn open(&self, src: &SourceRef) -> Result<Box<dyn SeekableRead>, StorageError>;
fn read_range(&self, src: &SourceRef, r: Range<u64>) -> Result<Vec<u8>, StorageError>;
fn writable_sidecar_dir(&self, src: &SourceRef) -> Option<PathBuf>;
}
pub trait Secrets: Send + Sync { /* Secret Service | Android Keystore */ }
pub trait Lifecycle: Send + Sync { /* process death, memory pressure, background */ }
| Concern | Linux | Android |
|---|---|---|
| Storage | Filesystem under granted roots | SAF tree URIs, DocumentsContract |
| Secrets | Secret Service (libsecret) | Keystore + DataStore + Tink |
| Background | Threads | WorkManager, foreground service |
| Lifecycle | Process lives | Death at any moment; onTrimMemory |
writable_sidecar_dir returns None on read-only mounts and most SAF trees, which is when the
app-managed sidecar store is used.
11. Build order
Independent of D12's resolution, the foundations are shared. What differs is what gets built on top first.
Phase 0 — spikes. S1 (Slint+wgpu zero-copy), S2 (Android, two GPU vendors), S10 (SAF at 10k files), S11 (Play permissions), S9 (cross-platform tolerance). These can invalidate the architecture; everything else assumes they pass.
Phase 1 — foundations, needed either way. dr-types, dr-plat + both implementations,
dr-catalog, dr-sidecar, dr-decode (preview path first), dr-gpu device and tile scheduler.
Phase 2 — depends on D12:
- Culler-first: preview ladder, culling mode, raw histogram, focus peaking, ratings, burst grouping. Ships something usable without any develop chain.
- Vertical-slice-first: minimal develop chain (exposure, curve), export, proving the full architecture end to end but not yet useful.
Phase 3+: the remaining develop operations, masking, sync, ingest, AI denoise.
12. Architectural constraints
Decisions treated as fixed because reversing them is expensive. Each is evidence-backed.
6.1 GPU results never round-trip through the CPU
darktable identifies this as their single biggest bottleneck: OpenCL output returns to the GTK thread, which paints it on one CPU core. It constrains the UI framework choice more than any other requirement — whatever renders the UI must composite our GPU texture directly.
Measured on this project (Radeon RX 7900 XTX, cargo run -p dr-gpu --example bench --features readback), comparing the compute pass alone against compute plus a CPU readback:
| Canvas | Compute | With readback | Readback share |
|---|---|---|---|
| 840×692 | 0.06 ms | 0.63 ms | 90% |
| 1920×1080 | 0.11 ms | 2.35 ms | 95% |
| 3840×2160 | 0.28 ms | 7.43 ms | 96% |
At 4K the shader finishes in 0.28 ms and then 7.15 ms is spent moving pixels through the CPU — a 26× overhead that scales with area, which is why an uncapped window resize falls off a cliff. The constraint is not a stylistic preference; it is the dominant cost in the frame.
6.2 Tiling from day one
Mobile GPUs have far less memory. Retrofitting tiling into a whole-image pipeline is a rewrite.
6.3 Non-destructive edit graph, GPU-resident
Recompute from the graph rather than caching CPU bitmaps between stages.
6.4 GPU-first, not GPU-accelerated
RawTherapee shows excellent quality is achievable on CPU — and that retrofitting GPU into a mature CPU pipeline effectively never happens.
6.5 The catalog is the UI's source of truth for queries
The UI never touches storage or network synchronously.
6.5a The core has no UI dependency
Operations describe parameters as data. This buys headless golden-image testing, one-operation-two- presentations, and UI replaceability. CI-enforced.
6.6 Sync cannot rely on server-side change tokens
Verified: Nextcloud's Directory.php does not implement ISyncCollection; sync tokens exist
only for CalDAV/CardDAV (nextcloud/server#22584, open since 2020). Folder ETags must be persisted
from schema v1.
6.7 Server-side RAW previews are not assumable
Verified: no RAW preview provider ships with Nextcloud. Range-based embedded extraction is the
primary path; server previews are opportunistic. preview_max_filesize_image defaults to 50 MB,
excluding many RAWs even where a provider exists.
Measured against a real instance, 2026-08-09, and the case is stronger than assumed.
/core/preview returned HTTP 400 for every parameter combination attempted — including bare
?fileId=N, and including a JPEG the same server reported as nc:has-preview=true. The cause is
unexplained: it is a server-side preview configuration issue, not a request-shape error on our side.
Recorded as unexplained rather than understood, because the distinction matters if someone later tries to depend on this endpoint. Nothing does today — the connector treats a preview failure as a miss and falls through to range extraction, which is the designed behaviour rather than a workaround.
Range extraction validated on the same library. 262 KB read from a 21.5 MB DNG in 119 ms — 1.22% of the file — yielded camera model and ISO. Extrapolated across the 17,185-file test library, cataloguing by whole-file fetch would move roughly 370 GB; the range path moves a few MB. That ratio is the difference between a viable mobile experience and an unusable one, and it is why §6.7 treats range reads as the mechanism and server previews as a bonus.
6.8 Sync is selective, not mirror-style
6.9 Android forbids the filesystem-scan model
Verified: MANAGE_EXTERNAL_STORAGE is not grantable — Play policy explicitly disallows "media
files access" as a permitted use. READ_MEDIA_IMAGES would not help: proprietary RAW is not typed
image/* by the platform scanner, so it does not appear in MediaStore.Images. SAF is the only
path, and it provides no filesystem path — hence SourceRef (§3.1).
6.10 GPU device loss is expected
Routine on Android: backgrounding, driver resets, thermal events. Recovery is tractable because the edit graph is CPU-side.
6.11 Drawn masks are GPU-rasterised
darktable rasterises parametric masks on GPU and drawn masks on CPU; users describe drawn masking as "unworkable" and the lag is not fixable with more cores. Highest-value architectural win available relative to the incumbents, and free if designed in now.
6.12 Sidecars authoritative, catalog disposable
darktable maintains both a database and sidecars while achieving the reliability of neither —
documented lockouts, corruption where all three recovery paths failed, sidecars silently not written
in 5.0.0. RawTherapee's .pp3 is the model: plain text, next to the file, no database in the trust
path.
6.13 Bit-identity applies to integer state only
GPU float results diverge across vendors — transcendental implementations differ, drivers optimise differently, f16 rounding varies. Cache keys and graph hashes are computed over CPU-side integer state, which is exactly deterministic. Cross-platform rendering equality is a bounded tolerance (R1), not a checksum.
13. Decisions
Full rationale in requirements.md §8. Summary:
| # | Decision | Status |
|---|---|---|
| D1 | Rust + Slint + wgpu | Decided |
| D2 | rawler, LibRaw fallback behind a trait | Decided |
| D3 | First milestone | Provisional — depends on D12 |
| D4 | Nextcloud: ETag pruning, chunked upload, Login Flow v2 | Resolved |
| D5 | lcms2 + GPU-side transforms | Decided |
| D6 | Hand-written WGSL | Decided |
| D7 | reqwest + quick-xml | Decided |
| D8 | GPLv3 | Decided |
| D9 | Declarative operation descriptors | Decided |
| D10 | Single adaptive interface | Decided |
| D11 | Product positioning | Decided |
| D12 | Scope versus pace | Open |
| D13 | Face inference runtime and model licensing | Runtime answered, licensing open |
| D14 | Segmentation source for local masking | Decided — arm C (docs/segmentation.md §14) |
14. Open questions
D12 — scope versus pace. The selected feature set implies years of full-time work against a stated evenings-and-weekends budget. Resolving it sets D3's milestone and §11's Phase 2.
NFR-R8 — CPU fallback extent. §6.4 says GPU-first; NFR-RES-2 assumes a CPU fallback on allocation failure. Decide whether that means a full CPU pipeline (a second implementation) or only tile-spill staging. Recommend the latter.
Slint accessibility on Android. Unverified; may be a toolkit gap (NFR-A11Y-2). Cheaper to discover before the UI is built.
Play Store versus GPLv3. Generally workable but should be confirmed, with F-Droid as fallback.