Add folder scan with format selection; validate A3 on a real library

Library setup as the user described it: pick a folder, choose which RAW
types to look for, scan recursively.

  dr-types::FormatFilter  the tick-box selection, seeing through VFS
                          placeholder suffixes so a dehydrated CR2 still
                          matches as a CR2
  dr-sync::scan           recursive walk, Depth:1 per directory, pruning
                          unchanged subtrees where the backend propagates
                          directory ETags

Verified against nextcloud.tourolle.paris (34.0.2) on a real library:

  browse root      32 entries, 98ms
  scan PhotosRaw   17,185 RAW files in 334 directories, 34.1s
                   (7,836 CR2 + 9,349 DNG)
  range read       262KB of a 21.5MB DNG in 119ms — 1.22% of the file,
                   and enough to read "Canon EOS 6D | ISO 100"

That last line is assumption A3 validated on real data. Cataloguing this
library by whole-file fetch would move roughly 370GB; the range path
moves a few MB.

Pruning is capability-gated rather than assumed: with per-entry ETags a
probe costs a request and proves nothing about children, so it is skipped
entirely. A test asserts zero probes in that case.

Still unresolved: /core/preview returns 400 for every parameter
combination tried, including on a JPEG the server reports as having a
preview. Not a request-shape bug — it fails identically bare. Recorded
rather than worked around; ARCH §6.7 already treats server previews as
opportunistic, so nothing depends on it.
This commit is contained in:
2026-08-09 12:22:31 +02:00
parent fbadf9afc8
commit c8bb08e661
29 changed files with 7193 additions and 234 deletions
+529
View File
@@ -0,0 +1,529 @@
# DarkRoom — Catalog, library view, and background work
**Status:** Draft v0.1 · 2026-08-09
**Companion to:** [requirements.md](requirements.md), [architecture.md](architecture.md)
Specifies `dr-catalog`: the index the library view queries, how it stays current without rescanning
everything, and how thumbnails get made. [architecture.md §6.2](architecture.md) sketches the schema
in eight lines; this expands it to the point of implementability and fills the two gaps that sketch
leaves open — **incremental local scan** and **the job queue**.
Sync's remote side is already designed ([architecture.md §8](architecture.md)): ETag pruning turns a
no-op sync of 50k images into one request. Nothing equivalent existed for a local root, which is the
central problem this document solves.
---
## 1. What this must not do
Stated first because every design choice below follows from it.
| Must not | Why |
|---|---|
| Stat 50k files to open the catalog | NFR-P1: catalog open < 2 s desktop, < 4 s Android. SAF `DocumentsContract` queries are far slower than `stat` (spike S10). |
| Re-derive thumbnails for unchanged images | NFR-P3 throughput is for *new* work; redoing it on every connect makes first paint unbounded. |
| Fetch previews for remote images nobody looks at | A 50k remote library at 1–3 MB per range-extract is 50–150 GB. FR-NC-6 forbids bulk transfer by default. |
| Evaluate cache rules per grid cell | ARCH §9.5 already answers this: `tier_desired` is materialised. |
| Block the UI executor on any of it | NFR-P9, NFR-ARCH-1. |
The unifying principle: **work is proportional to what changed, or to what the user is looking at —
never to library size.**
---
## 2. Schema
Extends [architecture.md §6.2](architecture.md). Additions beyond that sketch are marked ⊕.
```sql
-- Roots -----------------------------------------------------------------
roots(
id INTEGER PRIMARY KEY,
kind TEXT, -- 'local' | 'saf' | 'remote'
grant_blob BLOB, -- SAF persisted permission; NULL on Linux
label TEXT,
last_seen INTEGER,
scan_generation INTEGER -- ⊕ bumped per completed scan; see §3.4
);
-- Folders: the unit of change detection, local and remote alike ---------
folders(
id INTEGER PRIMARY KEY,
root_id INTEGER NOT NULL REFERENCES roots(id),
parent_id INTEGER REFERENCES folders(id),
path TEXT NOT NULL,
etag TEXT, -- remote: propagating ETag (ARCH §8.4)
mtime INTEGER, -- ⊕ local: directory mtime
entry_count INTEGER, -- ⊕ local: direct children, mtime's blind spot
scanned_generation INTEGER, -- ⊕ deletion sweep; see §3.4
UNIQUE(root_id, path)
);
-- Images ----------------------------------------------------------------
images(
id INTEGER PRIMARY KEY,
root_id INTEGER NOT NULL REFERENCES roots(id),
folder_id INTEGER REFERENCES folders(id), -- ⊕ folder filter without LIKE
source_ref TEXT NOT NULL,
content_hash TEXT, -- NULL until hashed; see §3.5
format TEXT,
w INTEGER, h INTEGER,
captured_at INTEGER, -- UTC seconds; NULL if EXIF absent
captured_offset INTEGER, -- ⊕ minutes east of UTC; see §4.2
camera TEXT, lens TEXT,
iso INTEGER, aperture REAL, shutter REAL,
availability INTEGER,
file_size INTEGER, -- ⊕ cheap change signal alongside mtime
file_mtime INTEGER, -- ⊕
metadata_state INTEGER, -- ⊕ 0=none 1=stat-only 2=full EXIF; §3.5
sidecar_mtime INTEGER,
UNIQUE(root_id, source_ref)
);
-- Versions, keywords, remote, cache: per ARCH §6.2, unchanged -----------
-- Collections ⊕ ---------------------------------------------------------
collections(
id INTEGER PRIMARY KEY,
name TEXT NOT NULL,
parent_id INTEGER REFERENCES collections(id), -- collection sets
kind INTEGER NOT NULL, -- 0 = manual, 1 = smart
selector_json TEXT, -- smart only; the §5 Selector
created INTEGER
);
collection_members(
collection_id INTEGER NOT NULL REFERENCES collections(id) ON DELETE CASCADE,
image_id INTEGER NOT NULL REFERENCES images(id) ON DELETE CASCADE,
position INTEGER, -- manual ordering; NULL = by capture time
PRIMARY KEY(collection_id, image_id)
);
-- Jobs ⊕ ----------------------------------------------------------------
jobs(
id INTEGER PRIMARY KEY,
kind INTEGER NOT NULL,
subject_id INTEGER, -- image or folder, per kind
priority INTEGER NOT NULL,
state INTEGER NOT NULL, -- 0=pending 1=running 2=failed
attempts INTEGER NOT NULL DEFAULT 0,
not_before INTEGER, -- retry backoff
payload TEXT,
UNIQUE(kind, subject_id) -- coalescing; see §6.2
);
```
Indices that exist for a stated query, not speculatively:
```sql
CREATE INDEX images_captured ON images(captured_at); -- §4 timeline
CREATE INDEX images_folder ON images(folder_id);
CREATE INDEX images_hash ON images(content_hash) WHERE content_hash IS NOT NULL;
CREATE INDEX folders_parent ON folders(parent_id);
CREATE INDEX jobs_ready ON jobs(state, priority DESC, not_before);
CREATE INDEX versions_image ON versions(image_id);
CREATE INDEX members_image ON collection_members(image_id);
```
`content_hash` is indexed *partially*. It is NULL for most rows most of the time (§3.5), and a
partial index over the non-NULL subset is both smaller and what FR-CAT-9's reconnection-by-hash
and FR-CAT-11's duplicate detection actually query.
---
## 3. Incremental scan
### 3.1 The local analogue of ETag pruning
Nextcloud propagates ETags up the tree, so one request proves a whole library unchanged
([architecture.md §8.4](architecture.md)). A filesystem offers no such guarantee — a directory's
mtime changes when its *direct* entries change, and not when a grandchild does. There is no
cheap "did anything below here change" probe.
So local scan prunes at each level rather than at the root:
```
scan(folder):
(mtime, count) = stat(folder)
if (mtime, count) == stored:
# This directory's own entries are unchanged. Its files need no
# examination at all — but subdirectories may still have changed
# internally, so recurse into known children without listing.
for child in stored_children(folder):
scan(child)
else:
entries = list(folder) # the expensive call
reconcile(folder, entries) # §3.3
for child in entries.dirs: scan(child)
mark scanned(folder, current_generation)
```
Cost is **one `stat` per directory** when nothing changed, versus one per *file*. A 50k-image
library in ~2k folders costs 2k stats — a few milliseconds locally, and the difference between
meeting and missing NFR-P1 on SAF.
The recursion into unchanged directories is not redundant: it is what makes a change to one deep
file detectable at all, given no upward propagation. What it avoids is the *listing* — on SAF a
`DocumentsContract` query returning 200 rows costs far more than a metadata probe on the directory
itself.
### 3.2 Why entry-count as well as mtime
Directory mtime alone misses a real case: delete one file and create another within the same
timestamp granularity, and mtime can be unchanged while contents differ. Some filesystems and most
SAF providers report coarse timestamps, which widens the window.
Storing `(mtime, entry_count)` closes the common form of this — a paired add and remove changes
neither, but that is rarer than a bare add or remove, and both of those move the count. It is a
cheap narrowing, not a proof.
**Where correctness must not depend on it,** the user gets an explicit *Rescan folder* action
(FR-CAT-1), and reconnection matches by content hash (FR-CAT-9). Sync's remote path is unaffected —
ETags are authoritative there.
### 3.3 Reconciling a changed directory
For each entry in a listing:
| Situation | Action |
|---|---|
| Not in catalog | Insert with `metadata_state = 1`; enqueue `ExtractMetadata` |
| In catalog, `(size, mtime)` match | Nothing — the common case |
| In catalog, `(size, mtime)` differ | Re-enqueue `ExtractMetadata` and `Thumbnail`; clear `content_hash` |
| In catalog, absent from listing | Deletion candidate — §3.4 |
| Placeholder (`*.nextcloud`) | Catalogue as the image it stands for; `Availability::Offline` (ARCH §9.0) |
Sidecars are examined in the same pass: a `.drsc` whose mtime exceeds `images.sidecar_mtime` enqueues
a `ReadSidecar` job. This is how an edit made on another device — landed by the Nextcloud client,
not by us — reaches the catalog.
### 3.4 Deletion without a full sweep
A file removed outside the app appears only as an *absence*, which a pruned scan cannot see: the
folder it vanished from has a changed mtime and is listed, but a folder never visited is never
compared.
Generation counting handles this without a full pass. Each scan bumps `roots.scan_generation`, and
every folder reached — whether listed or skipped — records it. After the walk:
```sql
-- Folders never reached: their parent no longer lists them.
DELETE FROM folders
WHERE root_id = ?1 AND scanned_generation < ?2;
```
Images under a deleted folder cascade. Images missing from a *listed* folder are caught directly in
§3.3. Together these cover deletion with no additional traversal.
Deletion here means **removing the catalog row for a source proven absent**, which FR-CAT-9 sharply
distinguishes from a source merely unreachable. A root that fails to open at all — unplugged drive,
revoked SAF grant — aborts the scan and marks the root offline. It never runs the sweep, because
every folder would look unreached and the sweep would delete the entire library.
That guard is the single most dangerous line in this design, and it is stated as an invariant:
**the deletion sweep runs only after a scan that completed without a root-level access error.**
### 3.5 Metadata in two passes
Full EXIF extraction requires opening and parsing each file. At 50k images that is minutes, and it
must not stand between the user and a usable grid.
`metadata_state` records how far each image has got:
| State | Holds | Cost |
|---|---|---|
| 0 — none | Row exists, nothing read | — |
| 1 — stat-only | Name, size, mtime, format from extension | Free, from the listing |
| 2 — full | EXIF: capture time, camera, lens, exposure, dimensions | One open + parse |
The grid is usable at state 1: it can show filenames, sort by filename or file mtime, and display
placeholder cells. Promotion to state 2 runs as background jobs, prioritised by what is on screen
(§6.3), so visible images get real capture times within a frame or two of being scrolled to.
**Capture-time filtering (§4) needs state 2**, so a freshly scanned library's timeline is incomplete
until the pass finishes. The UI states this plainly — a progress affordance on the timeline, not a
silently wrong filter. Which is the FR-NC-6c principle applied to metadata rather than pixels: say
what you actually have.
`content_hash` is a *third*, still lazier tier. It requires reading the whole file, so it is computed
only when something needs it: import duplicate detection (FR-CAT-11), or reconnecting a moved source
(FR-CAT-9). Never during a routine scan.
---
## 4. The library view
### 4.1 Query model
The UI never assembles SQL. It hands the catalog a `Query` and receives a stable, windowable result:
```rust
pub struct Query {
pub filter: Selector, // §5 — same type cache rules use
pub sort: Sort,
pub descending: bool,
}
pub enum Sort {
CapturedAt,
Added,
FileName,
Rating,
/// Manual order within a collection; falls back to CapturedAt elsewhere.
CollectionPosition,
}
```
Results are fetched by window, never wholesale — FR-CAT-4 requires memory bounded independently of
catalog size:
```rust
impl Catalog {
fn count(&self, q: &Query) -> Result<usize, CatalogError>;
fn window(&self, q: &Query, range: Range<usize>) -> Result<Vec<GridRow>, CatalogError>;
}
```
`GridRow` carries exactly what a cell draws — id, thumbnail key, availability, rating, flag, capture
time — and nothing that would require a join per cell. Availability badges read `tier_desired`
directly (ARCH §9.5), so no rule evaluation happens on the render path.
A `LIMIT/OFFSET` window degrades at high offsets, since SQLite must walk the skipped rows. Scrolling
is overwhelmingly *sequential*, so the catalog keeps a keyset cursor for forward and backward paging
and falls back to OFFSET only for a scrollbar jump. Jumps are rare and single; scrolling is
continuous.
### 4.2 Time
Capture time is the spine of a photo library, and it has one persistent trap: **a photograph's
timestamp is local to where it was taken.** Store UTC alone and a shoot that ran 09:00–17:00 in
Tokyo displays as spanning two days in Paris. Store local time alone and ordering across a timezone
change is wrong.
So both: `captured_at` in UTC for ordering, `captured_offset` in minutes for display and for
day-bucketing. EXIF `OffsetTimeOriginal` supplies it where present; where absent — common on older
bodies — the offset is NULL and the catalog falls back to the library's configured display timezone,
flagged so the UI can show it as inferred.
Day, month, and year buckets are computed against **local** time. "Everything from 3 August" means
the photographer's 3 August.
The timeline affordance is a histogram of counts per bucket, which the grid uses for scrubbing:
```rust
pub enum Granularity { Year, Month, Day, Hour }
pub struct TimeBucket {
pub start: i64, // UTC seconds, bucket start
pub count: u32,
}
fn timeline(&self, q: &Query, g: Granularity) -> Result<Vec<TimeBucket>, CatalogError>;
```
This is one grouped aggregate over the `images_captured` index, not 50k rows into the UI. It is what
makes "drag across two years to find the trip" work, and it is the cheapest useful thing a library
view can offer over a flat grid.
### 4.3 Filtering interactively
FR-CAT-6 requires filter results to update interactively on 50k images. Three things make that hold:
1. **Filters compile to indexed predicates.** A `Selector` becomes a WHERE clause over indexed
columns. Keyword and collection membership become `EXISTS` subqueries against their own indices.
2. **Count and first window are one round trip.** The grid needs a row count to size its scrollbar
and the first screenful to paint; the catalog returns both together.
3. **A filter change cancels the one in flight.** Typing in a search box issues a query per
keystroke; each supersedes the last (NFR-ARCH-3). Without this the UI queues work it will discard.
---
## 5. Selectors: one type, three uses
[architecture.md §9.2](architecture.md) defines `Selector` for cache rules. The same type expresses
library filters and smart collections. This is deliberate and worth stating as a design decision,
because three near-identical predicate languages is a classic way for a catalog to rot.
| Use | Meaning |
|---|---|
| Library filter | What the grid shows now |
| Smart collection | A saved, named filter (FR-CAT-7) |
| Cache rule | What is kept locally, at which tier (FR-NC-6a) |
One consequence is directly useful: any filter the user has narrowed to can be saved as a smart
collection, and any collection can be pinned offline, with no conversion step. "Show me 5-star images
from the last 90 days" → save as a collection → pin it for the trip. Three features, one mechanism.
`Selector` moves to `dr-types` so `dr-catalog` and `dr-sync` share it without either depending on the
other. It gains variants the cache-rule sketch did not need:
```rust
pub enum Selector {
All, // ⊕ the empty filter
Collection(CollectionId),
Folder { root: RootId, path: String, recursive: bool },
DateRange(DateSelector),
Rating { min: u8 },
Label(ColourLabel),
Flag(FlagState),
Keyword(String),
Camera(String), // ⊕ FR-CAT-6 indexed field
Lens(String), // ⊕
IsoRange { min: u32, max: u32 }, // ⊕
Availability(Availability), // ⊕ "what can I edit right now"
Text(String), // ⊕ filename/keyword substring
All_(Vec<Selector>),
Any(Vec<Selector>),
Not(Box<Selector>),
}
```
`Availability` as a selector earns its place: on a tablet the most useful filter is often "what do I
actually have here", and it is also the natural thing to *pin* — "keep everything I've flagged that
isn't already local".
Compilation is a straightforward recursive walk producing SQL with bound parameters. **Nothing
user-supplied is ever interpolated into SQL text.** `Text` becomes a bound `LIKE` pattern with `%`,
`_`, and the escape character escaped.
---
## 6. Background work
### 6.1 Job kinds
```rust
pub enum JobKind {
ScanFolder, // §3, recursive from a folder
ExtractMetadata, // state 1 → 2
Thumbnail, // §7
ReadSidecar, // external sidecar change detected
WriteSidecar, // local edit → disk, debounced (ARCH §6.1)
ContentHash, // on demand only
FetchPreview, // remote range-extract (FR-NC-3)
FetchOriginal, // pinned or explicitly requested
}
```
### 6.2 Coalescing is the point
`UNIQUE(kind, subject_id)` on `jobs` means enqueueing is idempotent: an image touched five times
during a scan has one thumbnail job, not five. Enqueue is
`INSERT … ON CONFLICT DO UPDATE SET priority = max(priority, excluded.priority)`, so a re-request at
higher priority promotes the existing row rather than duplicating it.
This is what makes "regenerate on update" safe to call liberally. Every code path that notices a
change can just enqueue; the table absorbs the redundancy.
### 6.3 Priority
Reuses the existing GPU scheduler classes ([architecture.md §5.3](architecture.md)) so one notion of
priority governs the whole app:
| Class | Jobs | Preempts |
|---|---|---|
| `Interactive` | Metadata and thumbnails for visible cells; preview for the open image | everything |
| `Prefetch` | The scroll margin; next image in culling | Background |
| `Background` | Bulk metadata, rule-driven fetches, hashing | — |
Visible-cell work is enqueued by the grid as it scrolls, at `Interactive`. The effect is that a
freshly scanned library fills in *where the user is looking* first, and grinds through the rest
behind them.
### 6.4 Durability and failure
Jobs live in the catalog, so they survive process death — which on Android is routine, not
exceptional (FR-PLAT-AND-3). On startup, rows in state `running` revert to `pending`: the process
that owned them is gone.
Failures increment `attempts` and set `not_before` to an exponential backoff. After a bounded retry
count the job is marked failed and attached to its image as a typed error (NFR-ARCH-4) — one
corrupt file does not stall the queue, and the user can see which files failed and why.
**A job runner never touches the UI executor**, and `Interactive` work runs on the decode pool with
the I/O pool behind it (ARCH §7.1).
---
## 7. Thumbnails
### 7.1 When
Not "on first connect" as a bulk operation. Thumbnails are generated:
- **On demand**, for cells entering the viewport plus the prefetch margin — at `Interactive`
- **On change**, when §3.3 sees a differing `(size, mtime)`
- **On rule**, for images a cache rule pins at `Preview` or above — at `Background`
- **Never** for a remote image nobody has looked at and no rule covers
For a local library this converges on "everything, eventually", because scrolling reaches everything
and the background pass has nothing else to do. For a remote library it converges on "what you
actually browsed", which is the difference between a few hundred megabytes and a hundred gigabytes.
### 7.2 How, by availability
| Availability | Source | Cost |
|---|---|---|
| `Original`, local | Embedded JPEG via `dr-decode` preview path | ~200 KB read, no demosaic |
| `Original`, no embedded preview | Full decode, downscale | Expensive — `Background` only |
| Remote | Range-extract embedded JPEG (FR-NC-3) | 1–3 MB vs 25–100 MB |
| Placeholder / `Offline` | None — render the offline affordance | 0 |
The remote path deliberately does **not** ask the Nextcloud client to hydrate the file. ARCH §9.0
established hydration is whole-file, so it costs ~100× what the range extract does. Hydration stays
reserved for the original tier, where the user has asked for the actual image.
Server previews (`/core/preview`) are tried only where PROPFIND reported `nc:has-preview`. ARCH §6.7
verified stock Nextcloud ships no RAW preview provider, so for RAW this is nearly always absent — it
is an opportunistic saving, never the mechanism.
### 7.3 Storage
Thumbnails are content-addressed by `(content_hash | source_ref, size_class)` and stored as files
under the platform cache directory, with the `cache` table holding the index. Files, not BLOBs:
SQLite handles small blobs well but a 50k-image thumbnail cache is gigabytes, and mixing it into the
catalog would bloat the file the app must open in under two seconds.
Two size classes at v1 — grid (256px) and filmstrip/loupe (1024px) — both long-edge, both JPEG. The
cache is LRU-capped per NFR-RES-4, and thumbnails evict before proxies and long after sidecars,
which never evict at all (FR-NC-6b).
---
## 8. What this document does not settle
- **FTS.** `Selector::Text` is a `LIKE` scan over filename and keywords. Adequate at 50k; if free
text over description and title becomes a real workflow, an FTS5 table is the answer, and it is
additive.
- **Smart collection materialisation.** Currently evaluated on read. If a smart collection's
membership needs to be *stable* — for manual ordering, or for a pinned set that must not shift
under the user — it needs materialising with an invalidation rule. Deferred until there is a
concrete need.
- **Multi-root capture-time collisions.** FR-CAT-11 detects duplicates on import; the same image
catalogued under two roots is a related but distinct case, not yet specified.
- **Timeline granularity selection.** Which bucket size the UI picks for a given zoom is a UI
concern, but the catalog should probably suggest one from the query's date span rather than have
the UI guess.
---
## 9. Requirements touched
| ID | How this document addresses it |
|---|---|
| FR-CAT-1 | §3 incremental scan, cancellable and resumable via §6 jobs |
| FR-CAT-3 | §7 thumbnail pyramid, two size classes, embedded-preview fast path |
| FR-CAT-4 | §4.1 windowed queries, memory independent of catalog size |
| FR-CAT-5 | §3.5 two-pass metadata |
| FR-CAT-6 | §4.3 indexed filter compilation, §5 selectors |
| FR-CAT-7 | §2 collections schema, §5 manual and smart |
| FR-CAT-9 | §3.4 the offline/deleted distinction and the sweep guard |
| FR-CAT-11 | §3.5 lazy content hashing |
| FR-NC-3 | §7.2 range-extract for remote thumbnails |
| FR-NC-6a | §5 shared selector type |
| FR-NC-6c | §3.5 metadata honesty, §7.2 availability-driven sourcing |
| NFR-P1 | §3.1 one stat per directory, not per file |
| NFR-P3 | §7.1 on-demand generation |
| NFR-ARCH-2 | §6.3 priority classes shared with the GPU scheduler |
| NFR-ARCH-3 | §4.3 query cancellation, §6 job cancellation |
| NFR-RES-4 | §7.3 LRU cap, eviction order |