Files
dtourolle 427a572aad Record where the named executors stand
NFR-ARCH-1's register note says what b1d1c472 and 7be1efff met and what
they left, but outstanding.md, which exists to show the distance between
the register and the binary, had no entry, and catalog.md §6 still said
interactive work runs "on the decode pool with the I/O pool behind it"
when there are no pools.

outstanding.md §4 gains the entry: named and guarded, not bounded — the
counts are a budget, the guard covers block_on only, the two mask
workers and the core crates' threads are outside the module. catalog.md
names the executors the thumbnail and metadata work starts on and says
the counts are not yet enforced.
2026-09-27 08:00:19 -04:00

1034 lines
56 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# DarkRoom — Catalog, library view, and background work
**Status:** Draft v0.1 · 2026-08-09
**Companion to:** [requirements.md](requirements.md), [architecture.md](architecture.md)
Specifies `dr-catalog`: the index the library view queries, how it stays current without rescanning
everything, and how thumbnails get made. [architecture.md §6.2](architecture.md) sketches the schema
in eight lines; this expands it to the point of implementability and fills the two gaps that sketch
leaves open — **incremental local scan** and **the job queue**.
Sync's remote side is already designed ([architecture.md §8](architecture.md)): ETag pruning turns a
no-op sync of 50k images into one request. Nothing equivalent existed for a local root, which is the
central problem this document solves.
---
## 1. What this must not do
Stated first because every design choice below follows from it.
| Must not | Why |
|---|---|
| Stat 50k files to open the catalog | NFR-P1: catalog open < 2 s desktop, < 4 s Android. SAF `DocumentsContract` queries are far slower than `stat` (spike S10). |
| Re-derive thumbnails for unchanged images | NFR-P3 throughput is for *new* work; redoing it on every connect makes first paint unbounded. |
| Fetch previews for remote images nobody looks at | A 50k remote library at 1–3 MB per range-extract is 50–150 GB. FR-NC-6 forbids bulk transfer by default. |
| Evaluate cache rules per grid cell | ARCH §9.5 already answers this: `tier_desired` is materialised. |
| Block the UI executor on any of it | NFR-P9, NFR-ARCH-1. |
The unifying principle: **work is proportional to what changed, or to what the user is looking at —
never to library size.**
---
## 2. Schema
Extends [architecture.md §6.2](architecture.md). Additions beyond that sketch are marked ⊕.
```sql
-- Roots -----------------------------------------------------------------
roots(
id INTEGER PRIMARY KEY,
kind TEXT, -- 'local' | 'saf' | 'remote'
grant_blob BLOB, -- SAF persisted permission; NULL on Linux
label TEXT,
last_seen INTEGER,
scan_generation INTEGER -- ⊕ bumped per completed scan; see §3.4
);
-- Folders: the unit of change detection, local and remote alike ---------
folders(
id INTEGER PRIMARY KEY,
root_id INTEGER NOT NULL REFERENCES roots(id),
parent_id INTEGER REFERENCES folders(id),
path TEXT NOT NULL,
etag TEXT, -- remote: propagating ETag (ARCH §8.4)
mtime INTEGER, -- ⊕ local: directory mtime
entry_count INTEGER, -- ⊕ local: direct children, mtime's blind spot
scanned_generation INTEGER, -- ⊕ deletion sweep; see §3.4
UNIQUE(root_id, path)
);
-- Images ----------------------------------------------------------------
images(
id INTEGER PRIMARY KEY,
root_id INTEGER NOT NULL REFERENCES roots(id),
folder_id INTEGER REFERENCES folders(id), -- ⊕ folder filter without LIKE
source_ref TEXT NOT NULL,
content_hash TEXT, -- NULL until hashed; see §3.5
format TEXT,
w INTEGER, h INTEGER,
captured_at INTEGER, -- UTC seconds; NULL if EXIF absent
captured_offset INTEGER, -- ⊕ minutes east of UTC; see §4.2
camera TEXT, lens TEXT,
iso INTEGER, aperture REAL, shutter REAL,
availability INTEGER,
file_size INTEGER, -- ⊕ cheap change signal alongside mtime
file_mtime INTEGER, -- ⊕
metadata_state INTEGER, -- ⊕ 0=none 1=stat-only 2=full EXIF; §3.5
sidecar_mtime INTEGER,
UNIQUE(root_id, source_ref)
);
-- Versions, keywords, remote, cache: per ARCH §6.2, unchanged -----------
-- Collections ⊕ ---------------------------------------------------------
collections(
id INTEGER PRIMARY KEY,
name TEXT NOT NULL,
parent_id INTEGER REFERENCES collections(id), -- collection sets
kind INTEGER NOT NULL, -- 0 = manual, 1 = smart
selector_json TEXT, -- smart only; the §5 Selector
created INTEGER
);
collection_members(
collection_id INTEGER NOT NULL REFERENCES collections(id) ON DELETE CASCADE,
image_id INTEGER NOT NULL REFERENCES images(id) ON DELETE CASCADE,
position INTEGER, -- manual ordering; NULL = by capture time
PRIMARY KEY(collection_id, image_id)
);
-- Jobs ⊕ ----------------------------------------------------------------
jobs(
id INTEGER PRIMARY KEY,
kind INTEGER NOT NULL,
subject_id INTEGER, -- image or folder, per kind
priority INTEGER NOT NULL,
state INTEGER NOT NULL, -- 0=pending 1=running 2=failed
attempts INTEGER NOT NULL DEFAULT 0,
not_before INTEGER, -- retry backoff
payload TEXT,
UNIQUE(kind, subject_id) -- coalescing; see §6.2
);
```
Indices that exist for a stated query, not speculatively:
```sql
CREATE INDEX images_captured ON images(captured_at); -- §4 timeline
CREATE INDEX images_folder ON images(folder_id);
CREATE INDEX images_hash ON images(content_hash) WHERE content_hash IS NOT NULL;
CREATE INDEX folders_parent ON folders(parent_id);
CREATE INDEX jobs_ready ON jobs(state, priority DESC, not_before);
CREATE INDEX versions_image ON versions(image_id);
CREATE INDEX members_image ON collection_members(image_id);
```
`content_hash` is indexed *partially*. It is NULL for most rows most of the time (§3.5), and a
partial index over the non-NULL subset is both smaller and what FR-CAT-9's reconnection-by-hash
and FR-CAT-11's duplicate detection actually query.
**Made on first use, not by a migration.** A new `user_version` makes every older build refuse this
catalog's snapshot at sync (`sync::remote_is_mergeable` compares it and nothing else), and a tablet
a release behind would stop merging collections, keywords and people for a feature it does not
have. So what later releases added without needing old rows rewritten is created with `IF NOT
EXISTS` where it is first used, and an older build that meets it ignores it:
- `dedup_probes` (FR-CAT-11a, §3.5);
- the albums (FR-EXP-10, `core/dr-catalog/src/albums.rs`): `albums`, `album_exports` — one row per
file written into an album, keyed on the file name, since two crops of one photograph are two
files — and `album_folders`, this device's folder for each (§8.2);
- `keywords_term_version (keyword, version_id)`, made by `keywords::list`, which counts each word's
photographs on every selection change and without it read a `keywords` row per assignment to
learn its version;
- `faces_box (image_id, model_id, x, y, w, h)`, made by the merge's `match_faces`, which reads every
local face's box and model and without it opened each ~8 KB `faces` row to do so.
A failure to make one of the indexes — a read-only or busy catalog — is logged and the query runs
without it, as it did before.
**Opening does not repeat the backfill.** `schema::backfill` repairs what a write left owing — an
image a scan inserted without its default version, the RAW/JPEG pairing, a merged version's uuid, a
keyword assignment whose word has no term — and it used to run in every `Catalog::open`. Every
worker opens its own connection, so a develop landing paid it five times (~80 ms of CPU on the
reference catalog) to learn that nothing had changed. `core/dr-catalog/src/backfilled.rs` now
records, per path and per process, a stamp read before each backfill: the schema version, the
file's device and inode, and the newest image, version and keyword assignment by content as well as
id — because none of those tables is `AUTOINCREMENT`, a freed newest id is handed out again, and the
row that takes it is exactly one the backfill is owed. An open whose stamp matches skips it (~1 ms);
the first open in a process, a migration, a pull and a replaced file always run it. Kept in memory
rather than in the catalog, so nothing about it travels in the sync snapshot.
---
## 3. Incremental scan
### 3.1 The local analogue of ETag pruning
Nextcloud propagates ETags up the tree, so one request proves a whole library unchanged
([architecture.md §8.4](architecture.md)). A filesystem offers no such guarantee — a directory's
mtime changes when its *direct* entries change, and not when a grandchild does. There is no
cheap "did anything below here change" probe.
So local scan prunes at each level rather than at the root:
```
scan(folder):
(mtime, count) = stat(folder)
if (mtime, count) == stored:
# This directory's own entries are unchanged. Its files need no
# examination at all — but subdirectories may still have changed
# internally, so recurse into known children without listing.
for child in stored_children(folder):
scan(child)
else:
entries = list(folder) # the expensive call
reconcile(folder, entries) # §3.3
for child in entries.dirs: scan(child)
mark scanned(folder, current_generation)
```
Cost is **one `stat` per directory** when nothing changed, versus one per *file*. A 50k-image
library in ~2k folders costs 2k stats — a few milliseconds locally, and the difference between
meeting and missing NFR-P1 on SAF.
The recursion into unchanged directories is not redundant: it is what makes a change to one deep
file detectable at all, given no upward propagation. What it avoids is the *listing* — on SAF a
`DocumentsContract` query returning 200 rows costs far more than a metadata probe on the directory
itself.
### 3.2 Why entry-count as well as mtime
Directory mtime alone misses a real case: delete one file and create another within the same
timestamp granularity, and mtime can be unchanged while contents differ. Some filesystems and most
SAF providers report coarse timestamps, which widens the window.
Storing `(mtime, entry_count)` closes the common form of this — a paired add and remove changes
neither, but that is rarer than a bare add or remove, and both of those move the count. It is a
cheap narrowing, not a proof.
**Where correctness must not depend on it,** the user gets an explicit *Rescan folder* action
(FR-CAT-1), and reconnection matches by content hash (FR-CAT-9). Sync's remote path is unaffected —
ETags are authoritative there.
### 3.3 Reconciling a changed directory
For each entry in a listing:
| Situation | Action |
|---|---|
| Not in catalog | Insert with `metadata_state = 1`; enqueue `ExtractMetadata` |
| In catalog, `(size, mtime)` match | Nothing — the common case |
| In catalog, `(size, mtime)` differ | Re-enqueue `ExtractMetadata` and `Thumbnail`; clear `content_hash` |
| In catalog, absent from listing | Deletion candidate — §3.4 |
| Placeholder (`*.nextcloud`) | Catalogue as the image it stands for; `Availability::Offline` (ARCH §9.0) |
Sidecars are examined in the same pass: a `.drsc` whose mtime exceeds `images.sidecar_mtime` enqueues
a `ReadSidecar` job. This is how an edit made on another device — landed by the Nextcloud client,
not by us — reaches the catalog.
### 3.4 Deletion without a full sweep
A file removed outside the app appears only as an *absence*, which a pruned scan cannot see: the
folder it vanished from has a changed mtime and is listed, but a folder never visited is never
compared.
Generation counting handles this without a full pass. Each scan bumps `roots.scan_generation`, and
every folder reached — whether listed or skipped — records it. After the walk:
```sql
-- Folders never reached: their parent no longer lists them.
DELETE FROM folders
WHERE root_id = ?1 AND scanned_generation < ?2;
```
Images under a deleted folder cascade. Images missing from a *listed* folder are caught directly in
§3.3. Together these cover deletion with no additional traversal.
Deletion here means **removing the catalog row for a source proven absent**, which FR-CAT-9 sharply
distinguishes from a source merely unreachable. A root that fails to open at all — unplugged drive,
revoked SAF grant — aborts the scan and marks the root offline. It never runs the sweep, because
every folder would look unreached and the sweep would delete the entire library.
That guard is the single most dangerous line in this design, and it is stated as an invariant:
**the deletion sweep runs only after a scan that completed without a root-level access error.**
### 3.5 Metadata in two passes
Full EXIF extraction requires opening and parsing each file. At 50k images that is minutes, and it
must not stand between the user and a usable grid.
`metadata_state` records how far each image has got:
| State | Holds | Cost |
|---|---|---|
| 0 — none | Row exists, nothing read | — |
| 1 — stat-only | Name, size, mtime, format from extension | Free, from the listing |
| 2 — full | EXIF: capture time, camera, lens, exposure, dimensions | One open + parse |
The grid is usable at state 1: it can show filenames, sort by filename or file mtime, and display
placeholder cells. Promotion to state 2 runs as background jobs, prioritised by what is on screen
(§6.3), so visible images get real capture times within a frame or two of being scrolled to.
**Capture-time filtering (§4) needs state 2**, so a freshly scanned library's timeline is incomplete
until the pass finishes. The UI states this plainly — a progress affordance on the timeline, not a
silently wrong filter. Which is the FR-NC-6c principle applied to metadata rather than pixels: say
what you actually have.
`content_hash` is a *third*, still lazier tier. It requires reading the whole file, so it is computed
only when something needs it: import duplicate detection (FR-CAT-11), or reconnecting a moved source
(FR-CAT-9). Never during a routine scan.
Consolidating the duplicates a library already holds (FR-CAT-11a, 0.16.0) does not wait for it. It
proves a group the same by `content_hash` where every copy has one, and otherwise by a digest of
each copy's first and last megabyte, read by range through the backend and kept in `dedup_probes`
keyed on the size and mtime it was taken at, so a second review reads nothing.
`dedup_probes` is created on first use rather than by a migration, because a schema bump would make
an older build refuse this catalog's snapshot at sync (`core/dr-catalog/src/duplicates.rs`).
---
## 4. The library view
### 4.1 Query model
The UI never assembles SQL. It hands the catalog a `Query` and receives a stable, windowable result:
```rust
pub struct Query {
pub filter: Selector, // §5 — same type cache rules use
pub sort: Sort,
pub descending: bool,
}
pub enum Sort {
CapturedAt,
Added,
FileName,
Rating,
/// Manual order within a collection; falls back to CapturedAt elsewhere.
CollectionPosition,
}
```
Results are fetched by window, never wholesale — FR-CAT-4 requires memory bounded independently of
catalog size:
```rust
impl Catalog {
fn count(&self, q: &Query) -> Result<usize, CatalogError>;
fn window(&self, q: &Query, range: Range<usize>) -> Result<Vec<GridRow>, CatalogError>;
}
```
`GridRow` carries exactly what a cell draws — id, thumbnail key, availability, rating, flag, capture
time — and nothing that would require a join per cell. Availability badges read `tier_desired`
directly (ARCH §9.5), so no rule evaluation happens on the render path.
A `LIMIT/OFFSET` window degrades at high offsets, since SQLite must walk the skipped rows. Scrolling
is overwhelmingly *sequential*, so the catalog keeps a keyset cursor for forward and backward paging
and falls back to OFFSET only for a scrollbar jump. Jumps are rare and single; scrolling is
continuous.
### 4.2 Time
Capture time is the spine of a photo library, and it has one persistent trap: **a photograph's
timestamp is local to where it was taken.** Store UTC alone and a shoot that ran 09:00–17:00 in
Tokyo displays as spanning two days in Paris. Store local time alone and ordering across a timezone
change is wrong.
So both: `captured_at` in UTC for ordering, `captured_offset` in minutes for display and for
day-bucketing. EXIF `OffsetTimeOriginal` supplies it where present; where absent — common on older
bodies — the offset is NULL and the catalog falls back to the library's configured display timezone,
flagged so the UI can show it as inferred.
Day, month, and year buckets are computed against **local** time. "Everything from 3 August" means
the photographer's 3 August.
The timeline affordance is a histogram of counts per bucket, which the grid uses for scrubbing:
```rust
pub enum Granularity { Year, Month, Day, Hour }
pub struct TimeBucket {
pub start: i64, // UTC seconds, bucket start
pub count: u32,
}
fn timeline(&self, q: &Query, g: Granularity) -> Result<Vec<TimeBucket>, CatalogError>;
```
This is one grouped aggregate over the `images_captured` index, not 50k rows into the UI. It is what
makes "drag across two years to find the trip" work, and it is the cheapest useful thing a library
view can offer over a flat grid.
### 4.3 Filtering interactively
FR-CAT-6 requires filter results to update interactively on 50k images. Three things make that hold:
1. **Filters compile to indexed predicates.** A `Selector` becomes a WHERE clause over indexed
columns. Keyword and collection membership become `EXISTS` subqueries against their own indices.
2. **Count and first window are one round trip.** The grid needs a row count to size its scrollbar
and the first screenful to paint; the catalog returns both together.
3. **A filter change cancels the one in flight.** Typing in a search box issues a query per
keystroke; each supersedes the last (NFR-ARCH-3). Without this the UI queues work it will discard.
---
## 5. Selectors: one type, three uses
[architecture.md §9.2](architecture.md) defines `Selector` for cache rules. The same type expresses
library filters and smart collections. This is deliberate and worth stating as a design decision,
because three near-identical predicate languages is a classic way for a catalog to rot.
| Use | Meaning |
|---|---|
| Library filter | What the grid shows now |
| Smart collection | A saved, named filter (FR-CAT-7) |
| Cache rule | What is kept locally, at which tier (FR-NC-6a) |
One consequence is directly useful: any filter the user has narrowed to can be saved as a smart
collection, and any collection can be pinned offline, with no conversion step. "Show me 5-star images
from the last 90 days" → save as a collection → pin it for the trip. Three features, one mechanism.
`Selector` moves to `dr-types` so `dr-catalog` and `dr-sync` share it without either depending on the
other. It gains variants the cache-rule sketch did not need:
```rust
pub enum Selector {
All, // ⊕ the empty filter
Collection(CollectionId),
Folder { root: RootId, path: String, recursive: bool },
DateRange(DateSelector),
Rating { min: u8 },
Label(ColourLabel),
Flag(FlagState),
Keyword(String),
Camera(String), // ⊕ FR-CAT-6 indexed field
Lens(String), // ⊕
IsoRange { min: u32, max: u32 }, // ⊕
Availability(Availability), // ⊕ "what can I edit right now"
Text(String), // ⊕ filename/keyword substring
Person { id: PersonId, include_suggested: bool }, // ⊕ §10 (FR-CULL-11)
All_(Vec<Selector>),
Any(Vec<Selector>),
Not(Box<Selector>),
}
```
`Person` carries `include_suggested` rather than defaulting silently. A saved collection built from
confirmed faces must not quietly change membership because a later indexing pass guessed at another
face; the user chose "photos of Anna", not "photos the model currently believes contain Anna". The
default is `false`, and the interactive filter offers the looser form explicitly as a way to *find*
faces to confirm.
`Availability` as a selector earns its place: on a tablet the most useful filter is often "what do I
actually have here", and it is also the natural thing to *pin* — "keep everything I've flagged that
isn't already local".
Compilation is a straightforward recursive walk producing SQL with bound parameters. **Nothing
user-supplied is ever interpolated into SQL text.** `Text` becomes a bound `LIKE` pattern with `%`,
`_`, and the escape character escaped.
---
## 6. Background work
### 6.1 Job kinds
```rust
pub enum JobKind {
ScanFolder, // §3, recursive from a folder
ExtractMetadata, // state 1 → 2
Thumbnail, // §7
ReadSidecar, // external sidecar change detected
WriteSidecar, // local edit → disk, debounced (ARCH §6.1)
ContentHash, // on demand only
FetchPreview, // remote range-extract (FR-NC-3)
FetchOriginal, // pinned or explicitly requested
DetectFaces, // §10, on the proxy tier (FR-CULL-8)
}
```
**Thumbnails are not queued (decided 2026-09-26, #73).** `Thumbnail` is kept only so its number
stays taken, and is listed in `JobKind::RETIRED`. Up to 0.16.0 every remote scan enqueued one per
photograph, and nothing claimed the kind — no `JobHandler` was ever registered for it, on desktop or
Android (the same `dr-ui`), and the catalog snapshot carries `jobs` but the merge never reads them
(§8.2). The reference catalog held 23,582 such rows, about 1 MB with its indexes. Thumbnails are
made another way, and by the right source of truth:
- the grid asks a worker for the cells it is drawing, which serves them from the thumbnail store or
range-fetches the embedded preview (§7.2);
- the thumbnail sweep's work list is *what the store does not hold* (`thumbnails_outstanding`).
The store is shared between devices (§7.3), so it is the only thing that knows another device already
made a thumbnail; a per-device queue row cannot. A queue row would be a second, staler record of the
same debt. Metadata is owed the same way — `metadata_state < 2` is the sweep's work list — so the local
walk no longer enqueues `ExtractMetadata` either.
The rows already queued are dropped by `jobs::drop_retired`, from `runner::recover` at every catalog
open, rather than by a migration: a schema bump would make a device still on an older build refuse the
synced snapshot, and "every open" rather than "once" because an older build sharing the catalog
queues them again on its next scan. With the rows gone it is one probe of the `(kind, subject_id)`
index. `every_queued_kind_has_a_consumer` (dr-catalog) holds the rule: it reads the shipping sources
and fails if any kind is enqueued that no handler or claim names.
### 6.2 Coalescing is the point
`UNIQUE(kind, subject_id)` on `jobs` means enqueueing is idempotent: an image touched five times
has one job of a kind, not five. Enqueue is
`INSERT … ON CONFLICT DO UPDATE SET priority = max(priority, excluded.priority)`, so a re-request at
higher priority promotes the existing row rather than duplicating it.
This is what makes "regenerate on update" safe to call liberally. Every code path that notices a
change can just enqueue; the table absorbs the redundancy.
### 6.3 Priority
Reuses the existing GPU scheduler classes ([architecture.md §5.3](architecture.md)) so one notion of
priority governs the whole app:
| Class | Jobs | Preempts |
|---|---|---|
| `Interactive` | Metadata and thumbnails for visible cells; preview for the open image | everything |
| `Prefetch` | The scroll margin; next image in culling | Background |
| `Background` | Bulk metadata, rule-driven fetches, hashing | — |
Visible-cell work is enqueued by the grid as it scrolls, at `Interactive`. The effect is that a
freshly scanned library fills in *where the user is looking* first, and grinds through the rest
behind them. (As built, thumbnails and metadata get this ordering without the queue — the grid
requests its visible cells directly and the sweeps take what is left; see §6.1.)
### 6.4 Durability and failure
Jobs live in the catalog, so they survive process death — which on Android is routine, not
exceptional (FR-PLAT-AND-3). On startup, rows in state `running` revert to `pending`: the process
that owned them is gone.
Failures increment `attempts` and set `not_before` to an exponential backoff. After a bounded retry
count the job is marked failed and attached to its image as a typed error (NFR-ARCH-4) — one
corrupt file does not stall the queue, and the user can see which files failed and why.
**A job runner never touches the UI executor**, and `Interactive` work runs on the decode executor
with the I/O executor behind it (ARCH §7.1). Those are named rather than pooled today: thumbnails
on demand start as `decode:thumbs` and the sweeps as `decode:thumb-sweep` and `decode:metadata`,
through `dr_ui::executors::spawn` and each on a thread of its own, and the thread counts §7.1 gives
are a budget nothing yet enforces (NFR-ARCH-1).
---
## 7. Thumbnails
### 7.1 When
Not "on first connect" as a bulk operation. Thumbnails are generated:
- **On demand**, for cells entering the viewport plus the prefetch margin — at `Interactive`
- **On change**, when §3.3 sees a differing `(size, mtime)`
- **On rule**, for images a cache rule pins at `Preview` or above — at `Background`
- **Never** for a remote image nobody has looked at and no rule covers
For a local library this converges on "everything, eventually", because scrolling reaches everything
and the background pass has nothing else to do. For a remote library it converges on "what you
actually browsed".
**Measured on a real 17,185-RAW library, 2026-08-09:** cataloguing it by whole-file fetch would move
roughly **370 GB**; the range-extract path moves a few MB for the images actually viewed. This is
the single largest cost difference in the design, and it is why §7.1 is a list of narrow triggers
rather than "generate them all on connect".
### 7.2 How, by availability
| Availability | Source | Cost |
|---|---|---|
| `Original`, local | Embedded JPEG via `dr-decode` preview path | ~200 KB read, no demosaic |
| `Original`, no embedded preview | Full decode, downscale | Expensive — `Background` only |
| Remote | Range-extract embedded JPEG (FR-NC-3) | 1–3 MB vs 25–100 MB — **measured: 262 KB of a 21.5 MB DNG, 119 ms, 1.22% of the file** |
| Placeholder / `Offline` | None — render the offline affordance | 0 |
The remote path deliberately does **not** ask the Nextcloud client to hydrate the file. ARCH §9.0
established hydration is whole-file, so it costs ~100× what the range extract does. Hydration stays
reserved for the original tier, where the user has asked for the actual image.
Server previews (`/core/preview`) are tried only where PROPFIND reported `nc:has-preview`. ARCH §6.7
verified stock Nextcloud ships no RAW preview provider, so for RAW this is nearly always absent — it
is an opportunistic saving, never the mechanism.
### 7.3 Storage: sharded, shared, synced
Decided 2026-08-09, implemented in `dr-thumbs`. Thumbnails live in **sharded SQLite databases that
sync to Nextcloud**, so a second device gets a full grid without re-fetching a byte of RAW.
```text
thumbs/
index.sqlite fileid → shard, size accounting, client id, adoption ledger
shard-0000.sqlite ≤ 25 MB, sealed
shard-0001.sqlite ≤ 25 MB, active
```
**Why a thumbnail is worth syncing when the catalog mostly is not.** It is expensive to produce — a
range fetch plus a decode, per image — and byte-identical for every client looking at the same file.
This does not make it authoritative: losing the store costs regeneration and nothing else, so §6.12
is untouched.
**Why shards, and why small.** The 25 MB cap is about *sync granularity*, not SQLite's limits. One
growing database means every client re-downloads all of it whenever a single thumbnail is added.
With sequential fill only the newest shard is ever dirty, so an up-to-date client transfers one small
file. Sealed shards are immutable, which makes them safe to cache forever and cheap to skip.
At ~20 KB per 256px JPEG a shard holds roughly 1,200 thumbnails, so the 17,185-image reference
library lands in ~14 shards.
**Keyed on `oc:fileid`** — stable across server-side rename and move (FR-NC-5), and already in hand
from PROPFIND. Accepted consequence: shards are account-scoped, so the same photograph on two
servers is thumbnailed twice.
**Stored as JPEG, not raw pixels.** A 256×170 RGBA buffer is ~174 KB against ~15 KB encoded. Since
shards sync, that 11× is transfer cost paid by every client, not just disk.
Three invariants, each tested:
| Invariant | Why it matters |
|---|---|
| A sealed shard never reopens | Reopening one forces every client that holds it to re-download |
| Re-storing an existing id updates in place, never migrates | Migrating would rewrite a sealed shard |
| Merging another client's shard is insert-only and idempotent | Both copies derive from the same bytes by the same code, so neither is better; preferring ours avoids dirtying a shard others have synced |
**When the exchange runs.** Corrected 2026-09-20. It fired only after the metadata sweep — hours
on a large library — so a fresh device re-derived every thumbnail it looked at, re-detected faces
and re-read every header before adopting the shards and snapshot that held all of it. It now also
fires the moment the scan completes, which is the first moment the rows the merges key on exist,
and the sweep starts behind it. In steady state that pass is one listing. The catalog merge also
takes **capture metadata** (`captured_at`, offset, camera, lens, ISO) for images still at
`metadata_state < 2`, matched by `oc:fileid` — a date is a fact about the file's bytes, not local
state, and the snapshot already carried it; the sweep's per-chunk query then finds nothing left.
**The transfer**, in `dr-ui`'s `derived_sync`, exchanges shards with `.darkroom-derived/` under the
library root. `ThumbStore::shards()` reports which are sealed, so an up-to-date client's whole pass
is one listing plus whichever shard is still open.
**Why a remote name carries a client id.** Corrected 2026-08-16. Shard ids are *per store* — every
client fills its own numbering from 0 — so the flat `shard-NNNN.sqlite` namespace the transfer first
used had two clients writing one name. Two failures followed from it, and both were live: the second
client's upload **overwrote** content the first still believed was published, and no client could
distinguish a peer's shard 3 from its own, so the only safe reading of "I already hold 3" was to skip
it. Between them, two populated clients exchanged almost nothing — only shards numbered above the
other's highest. A fresh device worked, which is why it went unnoticed: with no local shards there is
nothing to collide with.
The name is now `shard-<client>-NNNN.sqlite`, where `<client>` is minted per store in `index.sqlite`
beside the numbering it qualifies — a store deleted and rebuilt restarts at shard 0 and must not
claim its predecessor's names. Since a client's own ids no longer say anything about what it has
taken from others, `index.sqlite` also keeps an **adoption ledger** of merged remote names and the
size each had. Size, not a flag: a peer's sealed shard never returns, but its open one grows, and
re-merging the grown copy is how the thumbnails it gained arrive.
Flat names left on servers by earlier builds are still read — they report no owner, so each client
adopts them once — and nothing is written under that form again. A flat name whose id and byte size
match a local shard is that client's own earlier upload by the same identity argument used for
sealed shards, so the rename does not cost every client a re-download of its whole store. Older
builds ignore the new names, so they stop receiving shards until updated; nothing is lost, since
their own uploads are still adopted.
Two size classes remain planned — grid (256px) and filmstrip/loupe (1024px). Only the grid class is
implemented. The cache is LRU-capped per NFR-RES-4, and thumbnails evict before proxies and long
after sidecars, which never evict at all (FR-NC-6b).
---
## 7a. Editing collections
Decided 2026-08-09. The schema for collections landed with §2 and the cross-device merge rules with
§8; this is the layer between them — the operations a user actually performs, in
`dr_catalog::collections`.
### 7a.1 Hierarchy and membership are independent
Two structures, deliberately not entangled:
| | Mechanism | Meaning |
|---|---|---|
| Hierarchy | `collections.parent_id` | A collection inside a collection (Lightroom's "collection set"). A parent is an ordinary collection, not a separate kind, so a set can hold images of its own |
| Membership | `collection_members` | An image is in as many collections as the user likes. Nothing moves on disk; no collection owns an image |
**Adding an image to a child does not write a row for the parent.** A parent's contents are the union
of its own members and its descendants', computed on read. Materialising it instead would make one
add touch every ancestor, and a reparent rewrite membership — both of which §8's row-level merge
would then have to reconcile. The read path pays a bounded tree walk instead, which at sidebar scale
is nothing.
The consequence the UI depends on: dragging images onto a collection is **additive**. It does not
remove them from anywhere, which is why the gesture's default action is `copy` and not `move`.
### 7a.2 Rules that exist to prevent silent damage
| Rule | Why |
|---|---|
| Every mutation bumps `revision` | §8.4 resolves conflicts by revision. An edit that updates `modified` alone is invisible to the merge, so the *other* device silently wins and the user's work vanishes |
| A no-op add does **not** bump it | Otherwise an idle device that re-dropped the same images outranks one that did real work |
| Deleting a parent **promotes** its children | The schema's `ON DELETE CASCADE` would take the whole subtree. Losing a nested collection because its container was tidied away is not recoverable |
| Deletion leaves a tombstone | Without it, merging with a device that still holds the collection resurrects it (§8.4) |
| Cycles are refused at the write | Both kinds — parenting under a descendant, and a smart collection whose selector reaches itself. A cycle is unbounded recursion in the tree walk, so it must not be *representable*, not merely handled when drawn |
| Tree walks are depth-guarded anyway | A merge can deliver a row this device never validated. The read path must terminate, so it truncates and logs rather than hanging the UI thread |
| A drop onto a smart collection is refused | Its membership *is* its selector; member rows would be a second source of truth that nothing reads |
| Deep counts are `count(DISTINCT image_id)` | An image in both a parent and a child is one photograph. A count that disagrees with the number of cells drawn makes both untrustworthy |
`collections.uuid` is generated from the OS CSPRNG. A collision fuses two unrelated collections at
the next merge, so the fallback path (used only if `/dev/urandom` cannot be read) logs loudly rather
than degrading identity quality in silence.
### 7a.3 Drag and drop is Slint's, not ours
The first implementation hand-rolled the gesture on `TouchArea` — tracking the press, measuring
travel to distinguish a click from a drag, and deciding the drop target from the last row hovered.
**It did not work**, for a reason worth recording: an interactive `Flickable` claims any drag
beginning inside it for scrolling and *cancels* the child `TouchArea`'s press, so the gesture could
never leave the grid. It also had a correctness hole — a tree rebuilt mid-drag could redirect the
drop, since a captured pointer is invisible to every other element.
Slint 1.17's `DragArea`/`DropArea` own all of it: capture, the click-versus-drag threshold,
arbitration against the `Flickable`, the image under the cursor, and hit-testing the release. What
remains in `collections_ui` is only what Slint cannot know — the payload (which images, read from the
selection when the drag starts) and the **spring**: a dwell timer that opens a collapsed collection
so a nested child can be reached mid-drag, and closes again whatever the drag merely passed over.
One hazard survives the change and is easy to reintroduce. Every consequence of a drop — rebuilding
the tree, refreshing the badges, rereading the grid — *replaces a Slint model*, and doing that inside
the `dropped` handler destroys the elements Slint is still using to deliver that event. So the drop
records its target and `drag-finished` acts on it. This is the same hazard `sync_rows` in `lib.rs`
documents for the adjust panel, and it presents as a control that works once and then goes dead.
---
## 8. Syncing the catalog file
Decided 2026-08-09. **This qualifies [architecture.md §6.12](architecture.md)** — the catalog
remains a rebuildable index, but the file itself now travels to Nextcloud. The qualification is
worth stating precisely, because the sidecar-authoritative model is load-bearing and this is the
one place it bends.
### 8.1 Why collections forced this
Every other thing the catalog holds has authoritative backing outside it. Ratings, labels,
keywords, and edit graphs live in sidecars next to the images, so a rebuild recovers them.
**Collections do not.** A manual collection is a set of images the user assembled by hand; nothing
in the filesystem records it. Losing the catalog loses them, and no rescan brings them back.
So collections need to be durable across devices somehow. Syncing the catalog file is the chosen
mechanism.
### 8.2 What the file sync does and does not carry
Collections were the first thing merged, and the rules in §8.4 were written for them. What else
merges reuses those rules or keys on the same identities, and each has no other home:
- **Collections and their membership** — by uuid and revision, membership as a set union.
- **Keywords** — the vocabulary by the same verdict, the assignments as a union.
- **People and identity judgements** — people by uuid and revision, and the confirmed and rejected
face assignments matched to local faces (`merge::match_faces`): by box first, and — since 0.18.0,
only on the photographs where a remote face is left over — by embedding, a pair being accepted
at cosine ≥ 0.7 when each is the other's best by a lead of ≥ 0.2 (#77; [faces.md §18.2](faces.md)). After every merge,
`dedup_people` folds people of one name whose confirmed faces agree, and a face held twice in
one photograph, through the ordinary `merged_into` redirect, which older builds already honour
(#78; [faces.md §19](faces.md)).
- **Albums** (FR-EXP-10, 0.17.0) — by uuid and revision with tombstones, and what went into each as
a set union keyed on the server's file id (a content hash on a folder library). An album's
server folder is a column of its row and travels with it; a folder on *this device* is in
`album_folders`, which the merge never reads and the upload snapshot drops (§8.3), because a
path or a SAF grant on one device means nothing on another.
- **Capture metadata** — the one exception inside `images`: a date, a camera, a lens and an ISO are
facts about the file's bytes, so a row this device has not yet read takes them from a peer that
has (`merge_metadata`), matched by `oc:fileid`.
The rest of a catalog describes *local* state — folder mtimes, cache file paths, job rows,
`tier_actual` — and importing another device's version of those would be actively wrong; the
downloaded remote is read for the tables above and discarded.
This is what keeps §6.12 substantially intact. The catalog is still deletable; what a rebuild from
sidecars cannot recover — collections and albums, which nothing in the filesystem records — is what
the sync exists to carry.
### 8.3 Two hazards the implementation must handle
**A WAL database is not one file.** Committed transactions can sit in `catalog.sqlite-wal` with the
main file lagging, so copying `catalog.sqlite` alone uploads a torn snapshot — internally consistent
as of some older point, silently missing everything since. The upload therefore never copies the
live file. It builds the snapshot in an empty file. It attaches the catalog, creates each table
from the catalog's own schema, and fills it with `INSERT … SELECT`, all inside one transaction.
That transaction holds a single read snapshot of the catalog, so concurrent writers are serialised
rather than raced, as the backup API did before. The file is then checked with `quick_check` before
it goes anywhere.
**The face crops stay out of the upload.** A crop is a ~5 KB JPEG on each `faces` row. On a 19k-face
library they are 96 MB of a 158 MB catalog. The face shards carry them to other devices, once each.
The merge reads a remote face's box and model to match it to a local one — and, where the boxes cannot decide, its embedding — never its pixels. No
device adopts a downloaded catalog as its own: a fresh device starts empty and takes faces, crops
included, from the shards. So the snapshot's `crop` is NULL, and a merge never writes a local
crop. They were first stripped (2026-08) by copying the whole file with the backup API, setting
`crop` to NULL and `VACUUM`ing. That wrote the file about three times to upload 50 MB. Since #71
the snapshot is built without them. Its header is kept as it was (`user_version`, page size and
the WAL flag), so every earlier build merges it unchanged. NFR-R2 backups still use the backup API
and keep the crops, because a backup is a file the user may have to live on. `album_folders` is
dropped from the built file for the reason §8.2 gives.
**Integer primary keys are not identities.** Two devices each allocate `collections.id = 1` for
different collections, so a row-level merge keyed on the integer id would collide them. Collections
therefore carry a **UUID**, and membership maps across devices by **image content hash**. The
integer ids stay local and are never compared across catalogs.
### 8.4 Merge rules
| Concern | Rule | Why |
|---|---|---|
| Which collection wins | Higher `revision` — a counter bumped per local edit. `modified` only breaks an exact tie | A device with a skewed clock cannot silently overwrite real work. The same reason FR-NC-9 avoids mtime for sidecars |
| Membership | **Set union**, not last-writer-wins | Two devices adding different images to one collection keep both. The exception — a removal racing an addition — resolves toward the addition, which is recoverable by removing it again. A lost addition is not |
| Deletion | Tombstone (`deleted = 1`) carrying a revision | Without it, merging against a device that still holds the collection resurrects it. With a revision, deletion competes on equal footing with a rename |
| An image the remote has and we do not | Skip the membership row | It joins on a later merge, once a scan has catalogued the file. Not an error |
| A remote from a newer schema | Decline before attaching | Attempting it would fail mid-transaction rather than declining cleanly |
| People with the same name | Folded after each merge when their faces agree ([faces.md §19](faces.md)) | Names typed separately on two devices otherwise stay two people for ever |
Merging is idempotent: running it twice reports no changes the second time. That property is tested,
because a merge that oscillates would upload on every sync forever.
### 8.5 What was rejected
**Replace-if-newer.** The literal reading of "sync the file and take the newer one". Rejected
because it is not a merge: whichever device syncs second loses every collection the first did not
have. Binary SQLite files do not merge, so "newer wins" means "older is destroyed".
**A `collections.drsc` sidecar at the library root.** The alternative that would have kept §6.12
untouched, merging as text the way edit sidecars do. Viable, and cheaper in machinery, but it means
a second serialisation format and a second merge implementation for the same data. Recorded here
because if the SQLite path proves troublesome, this is the fallback with a known shape.
---
## 9. What this document does not settle
- **FTS.** `Selector::Text` is a `LIKE` scan over filename and keywords. Adequate at 50k; if free
text over description and title becomes a real workflow, an FTS5 table is the answer, and it is
additive.
- **Smart collection materialisation.** Currently evaluated on read. If a smart collection's
membership needs to be *stable* — for manual ordering, or for a pinned set that must not shift
under the user — it needs materialising with an invalidation rule. Deferred until there is a
concrete need.
- **Multi-root capture-time collisions.** FR-CAT-11 detects duplicates on import, and FR-CAT-11a
consolidates the copies one root holds in several folders; the same image catalogued under two
roots is a related but distinct case, not yet specified.
- **Timeline granularity selection.** Which bucket size the UI picks for a given zoom is a UI
concern, but the catalog should probably suggest one from the query's date span rather than have
the UI guess.
---
## 10. People and faces
Specified by FR-CULL-8 … FR-CULL-12, NFR-SEC-5, [architecture.md §6.4](architecture.md). Gated on
spike S14 and decision D13 — the runtime and the model licences are unresolved, so this is the shape
of the subsystem, not a build order.
[faces.md](faces.md) names the models this shape is filled in with, and adds one column to §10.1's
`faces` table (`crop_px`) that the calibration in its §8 depends on.
### 10.1 Schema (a v5 migration)
```sql
CREATE TABLE people (
id INTEGER PRIMARY KEY,
uuid TEXT NOT NULL UNIQUE, -- merge identity, not the name (ARCH §6.3)
name TEXT NOT NULL,
-- Tombstone-by-redirect. A merged person must outlive its merge, or a
-- device that still has it resurrects it — same hazard collections have.
merged_into INTEGER REFERENCES people(id) ON DELETE SET NULL,
created INTEGER NOT NULL,
revision INTEGER NOT NULL DEFAULT 1,
modified INTEGER NOT NULL
);
CREATE TABLE faces (
id INTEGER PRIMARY KEY,
image_id INTEGER NOT NULL REFERENCES images(id) ON DELETE CASCADE,
-- Normalised to the image's long edge, so a face survives the proxy it was
-- found on being regenerated at another resolution.
x REAL NOT NULL, y REAL NOT NULL, w REAL NOT NULL, h REAL NOT NULL,
landmarks BLOB, -- 5 × (x, y) f32, the alignment input
detector_confidence REAL NOT NULL,
embedding BLOB NOT NULL, -- 512 × f16, the raw model output; re-normalised on load
-- Length of that vector: the model's own reading of how recognisable the
-- crop was, and the gate on whether this face may be compared *against*
-- (faces.md §6, §9). NULL for a face stored as a unit vector before it
-- was kept.
quality REAL,
-- What the eyes are doing (FR-CULL-8a, faces.md §17): per eye P(open),
-- the source pixels across its box and the sharpness of the patch the
-- classifier saw; and P(sunglasses). All seven or none; NULL is "never
-- read", which every filter treats as unknown rather than as closed.
-- The verdict -- open, closed, sunglasses, unclear -- is a rule in
-- dr_face::eyes, not a column.
eye_right REAL,
eye_right_px REAL,
eye_right_sharp REAL,
eye_left REAL,
eye_left_px REAL,
eye_left_sharp REAL,
sunglasses REAL,
-- The 106 dense landmarks the eyes were read from, packed as 16-bit
-- fixed point over the frame: 424 bytes (schema V18). Kept so the next
-- per-face pass runs from the catalog rather than from the original.
landmarks_dense BLOB,
-- Which model produced this. An embedding is only comparable to others
-- from the same model; mixing them silently yields nonsense similarities.
model_id TEXT NOT NULL,
detected_at INTEGER NOT NULL
);
CREATE INDEX faces_image ON faces(image_id);
-- Covers the eyes-open filter's subquery. Without it every check read the
-- whole face row -- the eye columns sit after the blobs -- and one count
-- took 24 s on the reference library (schema V17).
CREATE INDEX faces_eyes ON faces(image_id, eye_right, eye_right_px, eye_right_sharp,
eye_left, eye_left_px, eye_left_sharp, sunglasses);
CREATE TABLE face_person (
face_id INTEGER PRIMARY KEY REFERENCES faces(id) ON DELETE CASCADE,
person_id INTEGER NOT NULL REFERENCES people(id) ON DELETE CASCADE,
-- Calibrated P(this face is this person), never a raw cosine (FR-CULL-9).
probability REAL NOT NULL,
-- The user said so. Never overwritten by a later inference pass.
confirmed INTEGER NOT NULL DEFAULT 0
);
CREATE INDEX face_person_person ON face_person(person_id, confirmed);
```
Three things in that schema are load-bearing:
**`model_id` on every face.** Embeddings from different models are not comparable — this is the one
mistake that produces plausible-looking garbage rather than an error. Storing the model with the
embedding means a model change is detectable and re-indexable, instead of quietly poisoning every
similarity in the library.
**Normalised bounding boxes.** Detection runs on whichever proxy exists (FR-CULL-8). Storing pixel
coordinates would bind a face to a resolution that the cache is entitled to evict and regenerate
differently.
**`confirmed` as a column, not a probability of 1.0.** A confirmation is a different kind of fact
from a confident guess, and collapsing them loses the ability to recompute suggestions without
touching user data.
### 10.2 Why clustering is not a job kind
Detection is per-image and parallel, so it is a job (`DetectFaces`, coalesced per image like any
other). Clustering is a *whole-library* operation over the embeddings detection produced — it has no
natural `subject_id`, and running it per-image would rebuild the world on every photograph.
It therefore runs as a debounced library-level pass, triggered when detection has been idle and the
face count has moved materially since the last clustering. The same reasoning as sidecar writes: the
work is cheap to defer, expensive to repeat, and nobody is waiting on it.
### 10.3 The calibration lives with the library
FR-CULL-9 requires similarity to be a calibrated probability, fitted from this library's own faces.
That fit is a property of the catalog and its model, so it is stored alongside — a small table
holding the fit parameters, its validity flag, and a hash of the face set it was derived from, so a
materially changed library recomputes rather than trusting a stale fit.
When the fit is not valid — a library with too few faces to have positive pairs — the UI says the
confidence is unavailable. It does not fall back to an untuned default dressed up as a measurement.
### 10.4 What this does not settle
- **Which model, and which runtime.** D13. Everything above holds regardless of the answer, which is
why it is specified in terms of "a 512-d embedding from a stated model" rather than a named one.
- **The clustering algorithm.** Density-based over the calibrated distance is the obvious starting
point, but the parameters are an S14 question, not a design-time one.
- **Whether embeddings sync.** NFR-SEC-5 permits it, opt-in. The shard mechanism in §7.3 is the
obvious carrier if they do, but nothing here depends on that decision.
- **Faces in trashed images.** FR-CAT-15's trash moves files; whether their faces stay indexed and
keep contributing to clusters is unspecified. Probably they should be excluded from suggestions but
not deleted, so a restore does not re-index.
---
## 10a. Bursts and near-duplicates
Specified by FR-CULL-5, implemented in `dr_catalog::bursts` (a v11 migration) with the pass that
feeds it in `dr_ui::bursts`.
A burst is a run of frames that are **adjacent in time and look like the frame before them**. Both
halves are load-bearing. Time alone groups a whole wedding ceremony, because a photographer working
steadily never leaves the gap that would end the run. Similarity alone groups a studio setup shot
across two days, which is a project rather than a moment. The bounds are two seconds and eight bits
of a 64-bit difference hash, and the reasoning for each figure is in the module.
**Two seconds, for a burst that fires ten frames in one.** `images.captured_at` is whole seconds:
EXIF's `DateTimeOriginal` has no sub-second field, and `SubSecTimeOriginal` is optional and widely
omitted. Ten frames of a burst therefore arrive sharing a timestamp, and any threshold finer than a
second is a threshold on information the catalog does not have. Where the pace really is faster than
two seconds, the similarity bound is what separates the frames.
**The signal is a perceptual hash of the thumbnail, not of the original.** `images.perceptual_hash`
is filled from the 256px thumbnails §7 already stores — vastly more resolution than a 9×8 reduction
uses — so a library that has been browsed has already paid for its signatures and no RAW is decoded
for this. The consequence is stated rather than hidden: an image with no thumbnail gets no
signature, and a frame with no signature never joins a burst. It is picked up by the next pass.
**It is a pass, not a job kind**, for exactly the reason §10.2 gives for face clustering: a burst is
a property of a *run* of frames and has no natural `subject_id`, so a per-image job would rebuild
the world once per photograph. It runs when the thumbnail sweep finishes, which is the first moment
the signatures can all be computed.
**A newly found burst arrives open.** The pass marks frames; it never takes them off the screen.
Collapsing on discovery would be tidier and would also mean a background pass removing photographs
from under someone part way through a cull. Folding a burst up is the user's act, it is remembered
(`burst_expanded`), and a burst that is already known keeps whatever state it is in — so the pass
that follows the next import does not spring open a morning's work.
**Nothing here ranks a frame.** The representative of a collapsed burst is its *earliest* frame,
which is a fact about the clock rather than a judgement about the photograph. FR-CULL-5 names the
failure this avoids — rejecting the only frame of an important moment because someone blinked — and
the only judgement in the subsystem is the user's own choice of representative, which lives in its
own table (`burst_pick`) so that rebuilding the grouping cannot erase it. Same argument as
`people.ignored` in §10.
**What the collapse costs the grid.** Which rows a collapsed burst hides has to be decided by the
query rather than by the cells, because the grid is a window (`LIMIT n OFFSET k`) and the frames it
hides are mostly not loaded. So the predicate joins `VISIBLE` in every query that lists or counts
cells, under the same discipline: present in four places of five, the header's count, the
scrollbar, the shift-click range and the scrub's ordinal stop describing the same list.
**What this does not settle.** Bursts are local: the tables ride along in the uploaded catalog
snapshot and nothing on the far side reads them, so a second device rebuilds its own grouping from
its own signatures. Making `burst_pick` cross-device is a merge question of the same shape as §8.4's
and is not answered here.
---
## 11. Requirements touched
| ID | How this document addresses it |
|---|---|
| FR-CAT-1 | §3 incremental scan, cancellable and resumable via §6 jobs |
| FR-CAT-3 | §7 thumbnail pyramid, two size classes, embedded-preview fast path |
| FR-CAT-4 | §4.1 windowed queries, memory independent of catalog size |
| FR-CAT-5 | §3.5 two-pass metadata |
| FR-CAT-6 | §4.3 indexed filter compilation, §5 selectors |
| FR-CAT-7 | §2 collections schema, §5 manual and smart, §7a hierarchy, membership and editing |
| FR-CAT-9 | §3.4 the offline/deleted distinction and the sweep guard |
| FR-CAT-11 | §3.5 lazy content hashing |
| FR-NC-3 | §7.2 range-extract for remote thumbnails |
| FR-NC-6a | §5 shared selector type |
| FR-NC-6c | §3.5 metadata honesty, §7.2 availability-driven sourcing |
| NFR-P1 | §3.1 one stat per directory, not per file |
| NFR-P3 | §7.1 on-demand generation |
| NFR-ARCH-2 | §6.3 priority classes shared with the GPU scheduler |
| FR-CULL-5 | §10a burst grouping: capture-time proximity and image similarity, collapse without selection |
| NFR-ARCH-3 | §4.3 query cancellation, §6 job cancellation |
| NFR-RES-4 | §7.3 LRU cap, eviction order |
| FR-CULL-8 | §10.1 `faces` schema, §6.1 `DetectFaces` job kind on the proxy tier |
| FR-CULL-9 | §10.3 per-library calibration, stored with its validity and source hash |
| FR-CULL-10 | §10.1 `people` and `face_person`, merge-by-redirect, §10.2 clustering as a library pass |
| FR-CULL-11 | §5 `Selector::Person`, confirmed-only by default |
| FR-CULL-12 | §10.1 derived data in the catalog; names to the sidecar, UUID as merge identity |
| FR-CULL-8a | §10.1 the seven eye columns, derived like the embedding; the filter term is `RatingFilter::eyes_open` |
| NFR-SEC-5 | §10.4 sync left undecided and off; nothing in §10 emits an embedding |