Document trash, collections, thumbnails, and the UI direction

Specs for the work that follows: soft delete via a MOVE that preserves the
remote id, the collection tree and smart collections, the thumbnail store,
and the derived-state folder.

Adds two design documents. ui-refinement.md names the structural gaps
between the v0.1 UI and something that feels like a photo editor.
view-composition.md proposes a view controller for the display layer,
against the 500-line run() that has become one by accretion.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-08-09 20:36:23 +02:00
co-authored by Claude Opus 5
parent 5365123d92
commit 02b66ddfcf
6 changed files with 1335 additions and 61 deletions
+180 -1
View File
@@ -752,6 +752,101 @@ originals — a substantially smaller sync problem than develop parity.
*Design note:* pinch-zoom accidentally triggering ratings is a documented defect in Lightroom
mobile. Gesture and rating targets must not overlap.
### 3.9.1 People
Face recognition was deferred in §7 through the 2026-08-08 calibration. It is undeferred here in a
narrower form, and the narrowing is the point.
**What changed.** The deferral treated "face recognition" as an AI feature adjacent to subject
masking. It is not the same kind of thing. Masking is a *taste* operation applied to one image;
grouping photographs by who is in them is a **mechanical grouping problem over the whole library**,
which is the category FR-CULL-5 already commits to and already justifies: grouping is the automated
capability photographers consistently praise, because it organises without deciding. Every argument
FR-CULL-5 makes for burst grouping applies unchanged to people grouping. Answering "where are the
frames with the bride in them" across a 4,000-image wedding is a culling operation, and culling is
the differentiator.
**What is deliberately not in scope**, because it is the failure FR-CULL-5 names: no automated
*selection*. Nothing here rejects a frame, ranks a face, scores a smile, or detects a blink. The
feature produces a **filter**, never a judgement. The user's rating axes remain the only thing that
rejects a photograph.
**FR-CULL-8 — Face detection.** The app shall detect faces in library images as a background job,
producing per-face a bounding box, five-point landmarks, a detector confidence, and a 512-dimension
embedding.
Detection runs against the **thumbnail or proxy tier, never a full decode** (FR-CULL-2's ladder).
This is what makes indexing affordable: a library that has been browsed has already paid for its
proxies, so face indexing adds no RAW decodes that were not already happening. Where no proxy
exists, the job requests one at background priority rather than decoding inline.
Detection is a job in the FR-CAT-3 queue and inherits its properties without exception: coalesced
per image, interruptible, resumable across process death (FR-PLAT-AND-3), and strictly preempted by
visible work (NFR-ARCH-2). A library indexes while idle or it does not index; it never competes with
the grid.
*Acceptance:* indexing a 10k-image library completes without the grid dropping below NFR-P9's
interaction target at any point, and survives being killed and restarted with no repeated work
beyond the in-flight image.
**FR-CULL-9 — Calibrated identity.** Face similarity shall be expressed as a **calibrated
probability that two faces are the same person**, not as a raw embedding distance. Every threshold
in the subsystem — clustering, suggestion, auto-confirmation — shall be stated in that probability
space, and no code path may threshold a bare cosine similarity.
This is a hard requirement rather than an implementation detail because the failure mode is
invisible. A raw cosine means something different for every model, every population, and every face
size; a threshold tuned on one library silently misbehaves on another, and an uncalibrated
similarity still *looks* like a plausible number all the way to the user interface. A displayed
confidence that does not mean what it says is worse than no confidence, because it is trusted.
The calibration shall be fitted per library from that library's own faces, and shall report whether
it is valid. Where it is not — too few examples to fit — the app shall say the confidence is
unavailable rather than present an untuned default as though it were measured.
*Acceptance:* on a labelled corpus, the stated probability is within a documented tolerance of the
observed match rate across the probability range (a reliability-diagram check, not a single
accuracy figure).
**FR-CULL-10 — Clustering and naming.** Detected faces shall be clustered into unnamed groups. The
user names a group, and that name applies to its members. A person is thereafter a first-class
catalog entity with a stable UUID, independent of any name given to them.
The user shall be able to **merge** two groups that are the same person, **split** a group that is
not, **remove** a face from a person, and **rename** a person, at any time and without re-indexing.
Splitting must be as easy as merging: clustering will over-merge on siblings, on parents and
children, and on the same person a decade apart, and a tool that can only merge makes its own errors
permanent.
Confirmation is explicit. A face is either **suggested** (the system's inference) or **confirmed**
(the user's judgement), and the two are never conflated in storage or in display. Suggestions may be
recomputed freely; confirmations are user data and are never overwritten by a later inference pass.
**FR-CULL-11 — People as a selector term.** A person shall be a term in the §5 selector language,
composable with every other term.
This is the requirement that pays for the subsystem, and it is nearly free once FR-CULL-10 exists:
because one predicate language serves the library filter, smart collections, and cache rules, a
person term yields all three at once — filter the grid to a person, save "every photo of Anna rated
three or higher" as a smart collection, and pin "every photo of my children" to stay local on the
tablet. The last is a genuinely new capability, not a restatement of the first two.
Selectors shall distinguish confirmed from suggested membership, defaulting to confirmed-only, so a
saved collection does not silently change membership when a later indexing pass revises a guess.
**FR-CULL-12 — Names are user data; embeddings are not.** A confirmed person name is a user
judgement of the same class as a rating or a keyword, and shall be written to the sidecar
(FR-CAT-8), so it survives catalog deletion and travels with the photograph.
Embeddings, detections, cluster assignments, and unconfirmed suggestions are **derived data**. They
live in the catalog only, are rebuildable by re-indexing, and are never written to a sidecar. This
follows ARCH §6.12 exactly: the expensive-but-reproducible artefact stays in the disposable index,
and only the irreplaceable human judgement enters the trust path.
The person UUID is what a cross-device merge keys on, in the same way collections merge (FR-CAT-7).
Two devices that independently name the same cluster produce two people; merging them is the
ordinary FR-CULL-10 merge, not a special case.
---
## 4. Non-functional requirements
@@ -864,6 +959,37 @@ validation in release builds.
**NFR-SEC-4** — No telemetry without explicit opt-in.
**NFR-SEC-5 — Face data stays on the user's own hardware.** Face embeddings (FR-CULL-8) are handled
under a stricter rule than the rest of the catalog.
This is a personal tool for personal libraries (§1.1, D11) — the people in these photographs are the
user's family and friends. That is the reason for the rule, not a reason to relax it: the data is
sensitive precisely because it is personal, and the user is the only party with any claim on it.
- **Never leave the device by default.** Embeddings, face crops, and cluster assignments shall not be
transmitted, uploaded, or included in any diagnostics bundle (NFR-OPS-1) or crash report
(NFR-OPS-2), under any configuration. The diagnostics path has no opt-in for this; it is excluded
outright.
- **Sync is opt-in and separately consented.** Syncing embeddings to the user's own Nextcloud is
permitted — it is their server and their photographs, and it saves re-indexing a library per device
— but it is off by default, is not implied by enabling photo sync, and the consent states in plain
language what is being uploaded and why. Person *names*, being sidecar data (FR-CULL-12), sync with
the sidecar as ordinary metadata.
- **No third-party inference.** Face detection and embedding run locally. No image, crop, or
embedding is sent to a remote inference service, and the app ships no capability to do so.
- **Deletable, in one action.** The user shall be able to delete all face data — embeddings,
detections, clusters, and people — from a single control, without deleting the catalog or any
photograph, and to disable face indexing entirely so that no such data is produced.
- **Model weights are inspectable.** The models used shall be named and versioned in the about
screen, with their licences, so a user can determine what is running on their photographs.
*Rationale:* the rest of this document treats privacy as a property of the network boundary — TLS,
credentials in secure storage, opt-in telemetry. Face data needs more than a well-defended boundary,
because it is not revocable once it has crossed one, and because it describes people who are not the
user. The prohibition is therefore structural rather than configurable: the code paths that would
upload an embedding to anyone but the user's own server do not exist. A setting can be changed by
accident, or by a future maintainer who has forgotten why it was there; an absent code path cannot.
### 4.6 Execution model
**NFR-ARCH-1 — Named executors.** The app defines distinct executors — UI, GPU submission, decode
@@ -1028,6 +1154,56 @@ both a smaller build and a usable one — it needs no develop chain.
Resolving D12 sets D3 and [architecture.md §10](architecture.md)'s Phase 2.
### D13 — face inference runtime and model licensing · **OPEN**
§3.9.1 needs to run two neural networks locally. That collides with two settled positions, and
neither collision is small enough to leave implicit.
**1. The pure-Rust dependency policy.** Every dependency choice in this project has gone the same
way, for the same stated reason: rustls over aws-lc-rs, bundled SQLite over the system library, a
Rust Lensfun port over liblensfun, zune-jpeg over libjpeg — no C dependency to satisfy under the
Android NDK (D1's whole premise). The obvious way to run ONNX models is the ONNX Runtime C++ library,
which would be the largest exception to that policy in the codebase, and it would land on the
platform the policy exists to protect.
The options, in the order I would try them:
| Option | Cost |
|---|---|
| **wgpu compute**, models hand-ported to WGSL | No new dependency at all — the GPU device and shader infrastructure already exist (ARCH §5). Highest implementation effort, and a ViT is a lot of shader. |
| **`burn`** with the wgpu backend | Pure Rust, uses the existing GPU. Young, and ONNX import maturity needs checking against these two specific graphs. |
| **`ort`** (ONNX Runtime bindings) | Fastest to working code, best operator coverage. Reintroduces the C dependency and the NDK cross-compilation problem the policy avoids. |
The tension is real: the cheapest path is the one that breaks the rule. This is worth an explicit
decision rather than a default, and S14 is what informs it.
**2. Model licensing is a distribution blocker, not a detail.** The obvious pretrained weights are
not redistributable under GPLv3. The InsightFace "buffalo" family — ArcFace and the SCRFD detector,
the standard choices — are **licensed for non-commercial research use only**, which is incompatible
with this project's licence and with Flatpak, F-Droid, and Play distribution (NFR-COMPAT-2). Other
candidate weights need their licences read individually rather than assumed.
Two ways out, both with costs:
- **Find permissively-licensed weights** and ship them in-tree. Clean, offline-first, consistent with
how the Lensfun database ships. Requires that suitable weights exist at acceptable accuracy.
- **Download models on first use**, with the user accepting the upstream licence. Sidesteps
redistribution but adds a network dependency to a feature that is otherwise entirely local, needs a
hosting story, and sits badly with the local-first posture of NFR-SEC-5.
**This must be resolved before implementation, not during it.** Discovering at packaging time that
the feature cannot ship is the expensive failure, and it is entirely avoidable — it is a licence-
reading exercise, not a research question. S14 therefore puts it first.
*Prior art available:* `../scene-actor-extraction` is a working implementation of this pipeline
(SCRFD detect → 5-point align → 512-d embedding → Platt-calibrated similarity), benchmarked at 67.4%
macro-F1 on held-out films. Its C++ does not port — different language, OpenCV and TensorRT
dependencies — but its **design decisions do**, and they are the expensive part: the calibrated
probability space that FR-CULL-9 requires, the discipline of never thresholding a bare cosine, and
the practice of leaving an uncertain face honestly unnamed. Personal libraries should also score
better than its film benchmark: cooperative subjects, better lighting, and a closed gallery of dozens
rather than thousands.
---
## 7. Out of scope for v1
@@ -1042,7 +1218,7 @@ note where deferring now constrains the design later.
| Focus stacking | Same provenance consideration. |
| Print layout | — |
| Soft proofing | Parameterise the output colour stage by an arbitrary profile so this becomes a UI addition, not a pipeline change. FR-EXP-3's print-dimension mode already half-commits to print workflows. |
| Face recognition | — |
| ~~Face recognition~~ | **Undeferred 2026-08-09**, in the narrower form specified in §3.9.1 (FR-CULL-8 … FR-CULL-12): people *grouping and search*, no automated selection. Reclassified as culling rather than AI — it is the same mechanical-grouping category as FR-CULL-5, not the taste operation AI masking is. Gated on spike S14 and decision D13. |
| AI subject masking | Deferred per D11. Note darktable shipped this in 5.6 (June 2026), so the gap is now visible. **Conditions for deferring safely:** AI denoise ships in v1 (FR-DEV-3g ✓), manual masking is excellent including GPU-rasterised drawn masks (ARCH §6.11 ✓), and the product has a clear differentiator (culling, §3.9 ✓). When it does land, copy darktable's shape — prompt-point segmentation producing an *editable* mask that behaves like a hand-drawn one — not Adobe's opaque version. |
| AI upscaling | Deferred. Lower priority than denoise, which has no manual fallback. |
| Video | — |
@@ -1077,6 +1253,8 @@ note where deferring now constrains the design later.
| Adaptive layout (§3.5) | Snapshot tests at each breakpoint, and a resize test asserting no state loss across a layout-class transition. |
| Touch targets (FR-UI-3) | Automated check that interactive elements meet the 44pt minimum in touch modality. |
| Export sizing (FR-EXP-3) | Per-mode dimension assertions, including aspect preservation, fill-crop centring, and the upscale-disabled fallback. |
| Identity calibration (FR-CULL-9) | Reliability diagram over a hand-labelled corpus: stated probability against observed match rate, asserted within tolerance across the range — not a single accuracy figure, which would hide exactly the miscalibration this tests for. Plus a static assertion that no comparison thresholds a raw similarity. |
| Face data confinement (NFR-SEC-5) | Assert that a generated diagnostics bundle contains no embedding or face crop, and that with sync disabled no face data appears in any outbound request. Verified by inspecting what the code *can* emit, since the requirement is the absence of a path. |
---
@@ -1108,6 +1286,7 @@ stacks.
| **S8** | **Chunked upload v2** round-trip of a 100MB RAW, including resume after process kill | FR-NC-7 correctness | FR-NC-7 |
| **S12** | **GPU device loss recovery:** induce `VK_ERROR_DEVICE_LOST` mid-render, verify recreation from the edit graph with no lost edits | Whether ARCH §6.10 and NFR-R7 hold | ARCH §6.10 |
| **S13** | **Slint accessibility on Android:** verify TalkBack exposure of names, roles, and values | Whether NFR-A11Y-2 is achievable in the chosen toolkit | NFR-A11Y-2 |
| **S14** | **Face pipeline in Rust, on a real personal library:** run a detector plus an embedder over ~2,000 images through a Rust ONNX runtime, at proxy resolution, on the reference desktop. Measure per-image latency, cluster purity against hand-labelled truth, and fit the FR-CULL-9 calibration to see whether it converges on a library-sized sample. **Resolve the model licence question before writing any of it** | Whether §3.9.1 is buildable without breaking the pure-Rust dependency policy, and whether the accuracy is worth the subsystem | D13, FR-CULL-8, FR-CULL-9 |
### Why this order