Commit Graph
47 Commits
Author SHA1 Message Date
dtourolleandClaude Opus 5 421a47f1eb Untag FR-CAT-13, which no XMP is read or written to satisfy
FR-CAT-13 asks for standard XMP sidecars read and written — ratings, colour
labels, keywords and hierarchical subjects, title, description, copyright, GPS,
in `xmp:`/`dc:`/`lr:` schemas — so that other tools interoperate. Its one tag
was the module header of `dr-catalog/src/keywords.rs`.

That module stores keywords in SQLite. It names `dc:subject` twice, both times
in prose explaining why a keyword's text is the fact rather than its row id,
which is a good reason to have written it that way and not evidence of an XMP
implementation. Nothing in the tree parses or emits XMP: `dr-export`'s metadata
module writes EXIF and says in its own header that IPTC and XMP are named by
FR-EXP-8 and neither is read.

`dr-preset-xmp` is the crate whose name most invites the mistake. It reads
Lightroom `.xmp` *presets* — develop settings — under FR-DEV-6, and knows
nothing about the metadata schemas FR-CAT-13 is about.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 10:18:10 +02:00
dtourolleandClaude Opus 5 758436cc28 Keep the runner's borrow alive as long as the connection it reads
The first compile this branch had. One borrow error, in the four-thread
contention test: the `Runner` was the block's tail expression, and a
tail's temporaries are dropped after the block's locals, so it outlived
the `conn` it borrowed. Bound to a local, with the ordering rule written
down beside it -- it is exactly the shape someone tidies back.

Everything else stood: clippy clean at -D warnings, and all 18 runner
tests pass, including the four-thread four-connection claim and the
`UPDATE ... RETURNING` rewrite the author flagged as the riskiest line
in the diff.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 23:48:15 +02:00
dtourolleandClaude Opus 5 846a249156 Drain the queue that nothing has ever drained
`jobs` has been a complete durable work queue since the catalog was
written, and nothing has ever taken a job out of it. `claim_next`,
`complete`, `fail` and `recover_orphaned` had no callers outside their own
tests; `enqueue` had three. So the table grew one row per photograph and
kept it forever, and FR-PLAT-AND-3's resumability was a property of code
that never ran.

`runner` is the missing half. It owns no thread, no clock and no policy,
and that is the whole design: on Android the process does not decide when
background work may run. WorkManager does, subject to Doze, battery saver
and FR-NC-6's network constraints, and it revokes permission mid-job by
calling onStopped(). So the runner exposes `run_one` — claim, run, record —
and `drain`, which repeats it against a budget, a deadline and a
cancellation flag the host owns. A `Worker.doWork()` with ten minutes calls
drain with a deadline; a desktop idle pass calls it with none. That is the
seam the Android service plugs into, and it needs no Android to test.

Handlers are supplied from above, because the catalog knows what needs
doing and nothing about how: a thumbnail needs a decoder and a fetch needs
a network stack, neither of which belongs under core/dr-catalog. A runner
claims only kinds some handler declares, so a queue holding work this
device cannot do is left alone rather than failed five times.

Four outcomes, and only two of them are the job's fault. Done deletes the
row; Retry backs off; Abandon gives up now, for a failure no retry can fix;
Interrupted releases the claim with its attempt refunded and ends the
drain, because the host stopped rather than the job — five backgroundings
in a row must not mark good work as failed. Process death is the fifth and
cannot report itself, which is what `recover` is for.

Recovery is called from `show_catalog_now`, which is the one place a
catalog is opened for a session and already returns early if one is open.
It has to be exactly once and before any worker starts: there is no owner
column, so a second pass while a worker held a claim would take it away.
The attempt a dead claim consumed is deliberately kept — a job that takes
the process down with it is indistinguishable from one that fails, and the
attempt counter is the only evidence that survives a death.

The tests cover claiming under contention twice over: sequentially across
two connections, and with four threads on four connections against one
catalog on disk, asserting every job ran exactly once. Plus completion,
backoff, giving up, abandoning, interruption, budget, deadline,
cancellation, and a job orphaned by a simulated crash being reclaimed and
run once rather than lost or repeated.

Not wired to a handler yet, and deliberately not: the only enqueue site
the app actually reaches is the remote scan's, whose thumbnails are already
served by the async grid worker, and `walk`'s two sites are reachable only
from the scan_local example. Inventing a handler to make the plumbing look
used is how a requirement comes to read as covered by code that does not
implement it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 23:31:47 +02:00
dtourolleandClaude Opus 5 d63872e5a9 Make a claim one statement, and give the queue what a runner needs
The claim was a deferred transaction around a SELECT and an UPDATE, and
under a single connection that is fine. Under two it is not what it looks
like: the SELECT takes only a read lock, the UPDATE tries to upgrade, and
in WAL a worker that read the same snapshot as another gets
SQLITE_BUSY_SNAPSHOT on its write. That is not an error a busy handler can
retry away — the fix is to roll back and start over — so the queue was
"safe" only in the sense that the loser failed loudly instead of taking a
job someone else was holding.

`UPDATE jobs SET state = 1, attempts = attempts + 1 WHERE id = (SELECT ...)
RETURNING ...` is one statement and so one implicit transaction that takes
the write lock immediately. Two workers serialise, the loser waits out its
busy timeout, and neither can see a row the other already holds. The
existing tests are unchanged by it, because from one connection the two
forms are indistinguishable — which is exactly why it was never noticed.

The rest is the surface a runner has to have and did not:

- `claim_next_matching` takes only kinds a worker can actually do. Without
  it a device with no connector claims `FetchOriginal`, fails it, and pays
  five wakeups and five backoffs per photograph to reach a conclusion known
  before it started. Filtering after a claim cannot work: the claim has
  already marked the row running.
- `abandon` gives up now, for failures no retry can fix. `fail` uses it for
  its own MAX_ATTEMPTS branch, so there is one statement that ends a job.
- `release` hands a claim back with its attempt refunded, for a worker that
  is being stopped rather than a job that is going wrong. `attempts` stands
  in for the owner column the table does not have: it is bumped by every
  claim, so a stale worker's release matches nothing and changes nothing.
- `reap_orphan_subjects` deletes jobs whose photograph is gone. Coalescing
  keeps the table one row per unit of work and nothing ever shrank it when
  the work stopped existing. `ScanFolder` is excluded because its subject
  is a folder id, and joining that against `images` deletes by coincidence
  of numbering — hence `JobKind::subject_is_image`, and `JobKind::ALL` so
  the next kind added cannot quietly fall out of the filter.
- `counts` is the number a foreground service's notification is built from.

One behaviour change worth stating: a kind this build does not recognise is
now parked with an error rather than read as `ExtractMetadata`. The old
`unwrap_or` would have run a job of an unknown kind as some arbitrary known
one, which is worse than not running it at all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 23:31:27 +02:00
dtourolleandClaude Opus 5 b9891e2c04 Merge master into tablet-selection
Two real conflicts, both from work that landed either side of the same
lines rather than against them.

`lib.rs`: the settings controller was hoisted above the People screen's
wiring, and Android's thumbnail-tier eviction registered itself at the
same point. Independent, so both stay.

`library.rs`: manual collection ordering and burst folding each added a
clause to the same two queries. The scoped range read now carries both —
the folding matters there for one step further on than it does in the
grid, because a collapsed burst is one cell, so an ordinal counted over a
list still holding every frame names a photograph several places away
from the one the user pointed at.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 22:52:42 +02:00
dtourolleandClaude Opus 5 82d9d077b0 Merge: answer Android's memory warnings, and stop reporting a lost root as an empty library
FR-PLAT-AND-5 in full, FR-PLAT-AND-2 in part -- the recovery is built and
live for Nextcloud roots, the SAF cause it names does not exist yet.

FR-PLAT-AND-4 and FR-PLAT-AND-6 are not here, both blocked behind the
same gap: assemble-apk.sh compiles no Java, so the APK cannot carry a
Service or a FileProvider. The container has JDK 17 and build-tools 36;
the build step is what is missing.

Verified: fmt, clippy --workspace --all-targets -D warnings, and 1043
tests across dr-catalog, dr-sync, dr-sync-folder, dr-sync-nextcloud,
dr-plat and dr-ui. The aarch64 target was checked before the branch was
finished but not after; no device was available.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 22:19:46 +02:00
dtourolleandClaude Opus 5 75fd5619ca Refuse a scan whose root has gone, instead of reporting it empty
FR-PLAT-AND-2, and a silent failure on both platforms. `dr_sync::scan`
stepped over a NotFound or PermissionDenied the way it does for a child
that vanished mid-walk -- correct for a child, wrong for the root, where
it ended the walk, returned Ok with nothing in it, and reported a
successful scan of a library that was no longer there.

A lost root is now its own error. The images under it are marked
Availability::Offline per FR-CAT-9 and no catalog row is deleted;
`library::persist` clears the mark per file as each one is listed again,
so a root that comes back needs no repair step.

Partly satisfied rather than closed, and the gap is worth stating.
The recovery half is real and reachable on Android today, because
`map_status` turns Nextcloud's 403 and 404 into it and Nextcloud is how
a phone actually gets a library in this build. The causes the
requirement names -- revocation, reinstall, a removed card -- are
properties of a persisted tree permission, and there is none: SAF does
not exist here, `SourceRef::Document` is constructed only in test
modules, and `LocalStorage` rejects the variant outright. When SAF
lands it becomes a third producer of this error and nothing above it
changes, which is why the discovery belongs in the connector.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 22:19:38 +02:00
dtourolleandClaude Opus 5 fc4157a1e0 Keep the test scene's own arithmetic from overflowing a u64
The first compile this branch ever had. `cargo fmt` reflowed four files
and clippy passed at -D warnings untouched, but one test panicked:
`the_signature_does_not_change_with_scale`, on "attempt to multiply with
overflow".

It is the fixture, not the feature. `scene()`'s little LCG multiplied the
block's y by the golden-ratio constant with a plain `*` while the term
beside it already used `wrapping_mul`, so any scene taller than about 104
pixels overflowed in debug. Only the scale test builds one that large,
which is why 345 of 346 passed around it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 22:07:05 +02:00
dtourolleandClaude Opus 5 d3dadbd725 Group the frames of one moment, by when they were taken and what they look like
A burst is the commonest thing in a cull and the least interesting: twelve
frames of the same gull at 10 fps occupy twelve cells, are scrolled past
twelve times, and end with the photographer keeping one. FR-CULL-5 asks for
them to collapse to one representative and be judged as a unit.

Two signals, because neither alone survives a real library. Time alone groups
a whole wedding ceremony -- a photographer working steadily never leaves the
gap that would end the run. Similarity alone groups a studio setup shot across
two days, which is a project rather than a moment. Together they are specific:
adjacent in time *and* looks like the frame before it.

Two seconds is the time bound, and the reason is worth recording because the
figure looks absurd next to a 10 fps camera. `images.captured_at` is whole
seconds -- EXIF's DateTimeOriginal has no sub-second field and
SubSecTimeOriginal is optional and widely omitted -- so a burst arrives in the
catalog as ten frames sharing one timestamp. Any threshold finer than a second
is a threshold on information that is not there. Where the pace really is
faster, the similarity bound is what separates the frames.

Similarity is a 64-bit difference hash over a 9x8 box-averaged reduction,
compared between *adjacent* frames only. Chained rather than anchored on the
first frame, because by frame twenty a camera following a bird has nothing in
common with frame one while no two neighbours differ by much; the time bound is
what stops the chain running away. There is no all-pairs step and there must
never be one -- that is what turns a grouping pass into something nobody can
afford to run over 50k images.

Nothing here ranks a frame. FR-CULL-5 names the failure it is avoiding, which
is rejecting the only frame of an important moment because somebody blinked, so
there is no sharpness score and no best-of-burst. The representative is the
earliest frame -- a fact about the clock, not a judgement about the photograph
-- and the user's own choice lives in its own table so that rebuilding the
grouping cannot erase it. Same argument `people.ignored` makes one subsystem
over: nothing short of remembering a decision survives re-clustering.

A newly found burst is recorded *open*. Collapsing on discovery would be
tidier, and would also mean a background pass taking photographs off the screen
part way through a cull. The pass marks; the user folds.

It is a pass rather than a job kind for the reason catalog.md 10.2 gives for
face clustering: a burst is a property of a run of frames and has no natural
subject_id, so a per-image job would rebuild the world once per photograph.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 20:37:54 +02:00
dtourolleandClaude Opus 5 574107bc39 Let a manual collection be put in the order it is meant to be seen in
`collection_members.position` and `Sort::CollectionPosition` have been in the
catalog since collections were, and nothing above dr-catalog has ever written
or read either: `collections::set_order` had no callers, and the grid ordered
everything by capture time whatever it was scoped to — dr-ui does not construct
a `Query` at all, it has its own `GRID_ORDER` constant. So a manual collection
was a set with an order nobody could see or change.

Three pieces, because it could not be fewer:

`grid_order_for` decides the ordering from the scope, and both readers take it
from there. That is the load-bearing part. An ordinal only names a photograph
relative to an ordering, so the window read and the span read have to agree —
a shift-click resolved through a different ORDER BY than the cells were drawn
with selects a different run than the one on screen, and the user finds out
when the export runs. `read_ids_span` already stated that invariant about
`GRID_ORDER`; this widens it to an ordering that depends on the scope.

Only a single manual collection has one. A set draws its descendants' images
too, and two children's positions are unrelated integers that interleave
arbitrarily; a smart collection has no member rows to carry a position at all.
Both fall back to capture time and refuse the drop rather than pretending.

The drop is on the cell, on whichever half of it the finger landed — the
trailing edge is the only way to name the last place in a collection, since
there is no cell beyond the last one to drop in front of.

`reordered` is pure and the membership is rewritten whole. `set_order` sets the
positions it is given and leaves the rest, so a partial write would interleave
the moved run with rows nobody touched; and it is read unfiltered, so what the
filter is hiding keeps its place relative to what the user can see.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 20:04:17 +02:00
dtourolleandClaude Opus 5 0407fb8d2d Format the two new examples
They were written after the last fmt run and CI gates on --check.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 12:33:23 +02:00
dtourolleandClaude Opus 5 da20d42d33 Merge master: pluggable storage, and a name that anchors
Conflicts were docs/traceability.md alone, and it is generated — so it
was regenerated rather than hand-merged. dr-face was untouched on the
other side; ui/dr-ui/src/faces.rs and identity_ui.rs auto-merged, the
first around recluster's anchoring and the second around load_faces.

Worth recording because the two branches met on the same problem from
different ends. Master's "Let a name hold a group together" is the fix
for the sixteen Catherines — fourteen of them empty — that this branch
found while measuring the library and reported without fixing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 12:30:01 +02:00
dtourolleandClaude Opus 5 b2250cc460 Measure a regroup on the tablet, not just on the desktop
The GPU question needed a number nobody had: how a regroup divides on
the hardware whose CPU is weakest. dr-face carries no weights and
touches no display, and dr-catalog's example needs only a catalog file,
so both run under adb shell against a copy of a real library.

On the same 18,143 faces — desktop against the tablet — scan 0.96s /
2.61s, agglomerate 1.69s / 2.16s, score 0.26s / 0.40s. The scan is half
the pass on the tablet and under a third on the desktop, because twenty
cores of AVX2 pull ahead of NEON much further than the merge engine's
single-threaded hashing does. So a GPU GEMM is worth roughly 2× a
regroup on the tablet and 1.5× here, and it is the tablet that should
decide whether it is built.

The two architectures agree exactly: the same 1,531,969 evidence pairs,
the same 2,518 groups holding the same 16,246 faces, the same
reliability table. That is a better check on the NEON kernel than the
unit test can be.

Two instruments, both read-only: the example now prints its phases, and
dr-face gains scan_bench, which needs no library at all and so can
answer "how fast is this machine" on a device with nothing on it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 12:16:48 +02:00
dtourolleandClaude Opus 5 ebb7d3cf5c Score a suggestion against the people the user has named
The number beside a suggestion was the mean calibrated probability
between the face and the rest of its group, which measures the wrong
thing twice. It punishes coverage: a person with two hundred faces over
fifteen years is *meant* to have members a given photograph is
orthogonal to, so a correct suggestion onto a well-photographed person
scored low for being well photographed. And it never asked who else the
face might be — a face matching Anna at 0.95 and nobody else, and one
matching Anna at 0.95 and her sister at 0.93, came out identical, when
the second is the only one worth the user's attention.

dr_face::assign answers both, and multiplies them: the mean of the best
ten calibrated matches into the identity (the old mean, capped, which is
what stops coverage counting against it), times that identity's share of
the evidence against every *named* rival.

Only named people compete, and per person rather than per group. Both
halves of that had to be measured on a real 18,000-face library rather
than reasoned about. Normalising across every group made the number
useless — median suggestion 21%, four in five under half — because
clustering leaves one person spread over many groups, so a face competed
against itself; and keying rivals by group left Catherine competing with
Catherine, median 39%. Per named person: median 99.5%.

Rivals are gathered below the merge threshold, down to even odds: a
named person matching at 0.6 will never be merged into but is exactly
the competition to discount for. That would be a second similarity scan,
the expensive half of regrouping a library, so cluster_scored scans once
at the looser floor and hands the merge engine the subset at or above
the threshold — pair for pair what it would have scanned for itself,
held to that by a test.

Leave-one-out over that library's 2,702 confirmations across 54 named
people: 99.33% of faces placed on the right person against the old
mean's 99.15%, and the number shown for the right person moves from a
median of 90.4% to 99.3%. It errs low — 100% correct wherever it states
80% or more — which is the safe direction, and docs/faces.md §9.1 says
plainly that the low bands are not calibrated.

The example that measures it comes too: this is a claim about a
library's numbers, and nobody should have to take it on faith.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 10:44:01 +02:00
dtourolle 5768100816 Borrow the library to index it, and give it back
The passes that need every photograph's bytes — thumbnails, face
indexing — now borrow each one and release it at the end. On a
placeholder library that is the difference between peak disk being the
working set and being the whole library.

Including on cancellation, which was nearly missed: the face sweep
returns mid-loop when the user presses Stop, and without releasing there
the disk is spent and nothing is delivered for it.

`materialise` now answers whether *it* fetched the content. The pool used
to work that out by listing a file's parent directory — one listing per
file across a library — when the backend already had to `stat` it to
decide whether to ask. One syscall instead of a directory walk, and it
removes the bug class the tests found earlier: a file at the library root
has no `parent()`, so every one of them read as already-downloaded.

**Pinning is the retention control**, and it drives the model the catalog
already had rather than a second one. `tier_desired` is what the user
asked to keep hydrated, `pending_pins` is the resumable work list, and a
pinned collection is never dehydrated for the same reason it was never
evicted. It was in fact *broken* here before: `get` on a stub failed, and
the pin worker logged "one unreadable file must not abandon the whole
pin" and silently did nothing.

Pinned originals on such a library are recorded with `path = NULL`
(`Cache::record_in_place`) rather than copied under `originals/`. Two
reasons, and the second is the important one. A copy would hold every
pinned photograph twice, with the budget able to evict the half that was
not costing the disk. And `release` deletes the file a row names — so a
row that names none cannot delete anything, which puts the one
catastrophic operation out of reach by construction rather than by
remembering not to call it. Deleting a materialised file inside a synced
tree removes the photograph from the server and every other device.

Handing disk back is `spawn_dehydrate`, which asks the client.

Two gaps written down rather than papered over (docs/storage.md §7): a
hydrating pass cannot yet quote its cost, because a stub reports no size;
and the two sweeps hold separate pools, so a library indexed for both
fetches twice.
2026-08-29 09:57:53 +02:00
dtourolleandClaude Opus 5 f41b3f6e8e Answer "what would the other device end up with" without the other device
Build and test / Desktop (Linux) (push) Successful in 2h7m52s
Build and test / Layer separation (push) Successful in 59s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 4s
Traceability / Requirement traces (push) Successful in 37s
Build and test / Android (aarch64) (push) Successful in 22m22s
A tablet showed 200 of a person's 611 faces after syncing, and the obvious
suspects — that suggestions deliberately do not travel, that the cross-device
face match was too strict — were both wrong. Finding that out meant reading a
catalog on a release-signed Android build, which cannot be done.

So this stands the second device up locally: an empty catalog, given the images
a scan would have found, the shards adopted into it exactly as a sync does, and
the real catalog merged in as the remote. Then it counts, per person, against
what the source holds.

    person                     source     here
    Catherine                     611      611
    Me                            242      242
    Ian                           219      219

Which settled it: the merge carries everything, and the shortfall was transfer —
shards that never finished arriving. Worth keeping, because "did the sync lose
this or has it not got here yet" is a question that will come up again, and
guessing at it cost most of an evening.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 23:09:28 +02:00
dtourolleandClaude Opus 5 2c84aa1224 Write face shards the way the rest of the catalog writes
The face store was the one part of the catalog still on SQLite's default
rollback journal at `synchronous = FULL`. The catalog itself runs WAL at
`NORMAL` (`schema::configure`) and so does the thumbnail store; nothing decided
this one should differ, it was simply never set.

Measured on this project's own filesystem, that is **21.3 ms per commit against
0.05 ms** — four hundred times. And an export commits four times per
photograph: the shard's transaction, then three separate autocommitting writes
to the index. Ten thousand images is on the order of fourteen minutes spent
doing nothing but waiting for fsync, before a byte goes to the server. That is
the "checking faces…" that appeared to hang.

So: WAL and `synchronous = NORMAL`, matching the rest, and the three index
writes fold into one transaction. `NORMAL` is the same trade the catalog makes —
a shard is derived data, and losing the last commit to a power cut costs one
image re-exported.

WAL brings an obligation with it, because **a shard is uploaded by reading its
file**: the newest commits live in a `-wal` sidecar that no upload sends, so
without a checkpoint the server would receive a database missing exactly the
faces just written, and a peer would adopt it and see nothing wrong. `checkpoint`
folds the logs back in, with `TRUNCATE` rather than the default passive mode,
which gives up when a reader holds the log and would leave the same gap while
reporting success.

Two tests: that the store is in WAL like everything else, and — the one that
matters — that a checkpointed shard copied *without* its `-wal` still holds
every face. That second one fails without the checkpoint, which is how it was
confirmed to be testing something.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 23:08:10 +02:00
dtourolleandClaude Opus 5 a1e8f494b9 Say what the face sync is doing while it does it
The face pass set the status to "checking faces…" once and then said nothing
until it was finished. On a library whose first export after a re-index is 9,849
photographs and 85 MB of shards, that is eight minutes of a progress bar sitting
still — which is indistinguishable from a hang, and was reported as one twice.

Nothing was wrong with the sync. The only fault was that it was silent.

Three places now report, which are the three that take real time:

- **Preparing**, per image with a count, since this is the long one and the only
  one whose length the user cannot guess from anything on screen.
- **Sending**, per shard with its size, because a face shard carries crops and
  runs to tens of megabytes — one of them is a visible wait on any connection.
  Announced before the upload rather than after, since the wait *is* the upload.
- **Taking in** a peer's shard, which is a download and then a row-by-row merge.

The export reports every 25 images rather than every one, so the channel behind
it stays lost in the write it accompanies. `export_to_shards` keeps its old
signature and delegates, so the callers that do not want progress do not grow a
parameter for it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 21:56:16 +02:00
dtourolleandClaude Opus 5 0729dfa359 Stop the face sync opening a database per photograph
"Checking faces…" never finished. `export_to_shards` asks, for every indexed
image in the library, whether the shard store already holds that image at that
index time — and `indexed_at` answered by opening the shard database, running
its six-statement schema batch and two `pragma_table_info` queries, then
querying. Once per image. 9,849 times for this library, on every sync pass,
before a single face had been written.

The index time now lives in the store's `index.sqlite` alongside the shard
number, so the question is one indexed lookup on a connection that is already
open. It stays in the shard as well — that copy is the one that travels — but
nothing reads it from there on the hot path.

The write side had the same shape: `put_image` opened the shard afresh for each
image, which mattered little when exports were a handful of new photographs and
matters a great deal now that a re-index sends thousands. The handle is kept and
reused, invalidated by shard id so sealing a full one and moving to the next
drops it without anything having to remember to.

`INDEX_SCHEMA` is `CREATE ... IF NOT EXISTS` like the shard schema, so the new
column is added on open for an index already on disk — the same trap, caught the
same way.

Two tests: that the index time survives reopening the store, and that an index
written before the column can still be opened and written to.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 21:19:56 +02:00
dtourolleandClaude Opus 5 76eece8500 Add the columns an existing shard never got
`SHARD_SCHEMA` is entirely `CREATE ... IF NOT EXISTS`, which does exactly
nothing to a table that already exists. So `crop` and `indexed_at`, both added
to that batch, never appeared in any shard that had been written before — and
the `INSERT` naming them failed with "no such column".

Which took face export down completely, on every library that had ever synced a
face. Silently: `export_to_shards` returns the error, `sync_face_shards` logs it
at warn, and the sync goes on looking successful while the catalog fills with
faces no other device will ever see. This library's shard sat frozen at 1,807
faces with 15,194 in the catalog, and the reason was this rather than anything
in the export logic.

Shards are upgraded on open now: both columns are additive and nullable, so
catching up is one `ALTER` each. There is deliberately no version counter —
"does this column exist" is the question actually being asked, and asking it
directly cannot fall out of step the way a counter can.

A peer's shard is opened read-only and cannot be repaired, so one written before
crops is read as it stands, with a `NULL` standing in for the column. An adopted
face simply has no crop, which is the truth about it.

Four tests, built against the pre-crop schema written out in full rather than
derived from the current one — the point being that it is *not* the current
schema and must not track it. Verified against the real 1,807-face shard on this
machine: the ALTERs apply, writes succeed, and nothing already in it is lost.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 09:50:55 +02:00
dtourolleandClaude Opus 5 4ed10f7b23 Let a person cross from one device to another
The face shards carry boxes, landmarks and embeddings. What they deliberately
do not carry is who anybody **is** — the person rows, their names, and the
assignments joining the two. Those travel in the catalog snapshot, which is a
whole-file copy and does contain them.

But the snapshot is *merged*, not adopted, and this merge only ever looked at
collections and keywords. `face_shard`'s own module note says people travel in
the snapshot; nothing implemented it. So a second device received every face and
no people at all, and drew an empty People screen over a full catalog. Exactly
what a tablet showed after syncing thousands of faces from a laptop.

What travels is what the user decided, following the rule the rest of this
module already follows — judgements travel, inference is rebuilt:

- **People**, by uuid on `revision`, exactly as a collection is: the name, and
  whether the group was set aside.
- **Confirmations**, and **rejections** — "this is not her" is a fact too, and
  is why re-clustering does not put it back.
- **The suggestions inside an ignored group**, which are otherwise ordinary
  inference but are what anchors the ignore. Without them a group set aside on
  one device reappears on the other, the same fault that made "Not interested"
  not stick locally.

Ordinary suggestions are not carried. Both devices hold the same embeddings and
clustering is deterministic, so each recomputes them and arrives at the same
answer; shipping them would double the merge for no new information.

**A face has no cross-device identity**, and unlike a collection there is no
uuid to give it one. Both devices do agree on `oc:fileid` and roughly on the
box, so a remote face is matched to the local face on the same photograph whose
box overlaps it most, above 0.5 IoU. That is not a new rule — it is the one
`record_detections` already uses to carry a confirmation across a re-index, and
it is loose on purpose: the question is "the same face in the frame", not "the
same rectangle".

A local confirmation is never overwritten. Two devices confirming one face as
different people is a real disagreement and an assignment carries no revision to
settle it with; taking the remote's answer would let a sync undo what the user
just did on the device in their hands.

The remote's schema is probed rather than assumed: `remote_is_mergeable` admits
any catalog at or below this version, so one written before faces existed, or
before V10 added `ignored`, is ordinary. An absent table skips this half instead
of aborting a merge that would otherwise have succeeded.

Nine tests, including that the name lands on the overlapping face and not its
neighbour in the same frame, that a set-aside group stays set aside, that an
ordinary suggestion does not travel, and that merging twice changes nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 09:27:17 +02:00
dtourolleandClaude Opus 5 cf614efa61 Send a re-indexed image's faces to the other devices
`export_to_shards` asked `store.contains(file_id)` and skipped anything the
shard store had already heard of. So an image was exported exactly once, and
re-indexing it updated the catalog and nothing else — every other device kept
the first answer for ever.

That is not hypothetical. This library was re-indexed after the detection floors
changed and crops were added, going from 1,807 faces to 15,194; the shard store
still held the original 1,807, written before any of it. Nothing the re-index
produced could reach another device.

The shard's `indexed` table now carries the catalog's own `indexed_at`, and the
export compares against it. A re-indexed image goes again; an unchanged one
still costs nothing. Copied from the catalog rather than stamped when the shard
is written, because a shard-local write time advances even when nothing changed
and could not answer the question.

The column is nullable so a shard written before it still reads: absent means
"cannot vouch for it", which forces one re-export and then settles.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 09:27:16 +02:00
dtourolleandClaude Opus 5 79c0520506 Keep the face, not just a way to find it again
A face was drawn by decoding the 1024px proxy it was found on and cutting the
box out again, every time the People screen opened. That made the screen a
derivative of the thumbnail cache: evict a proxy — which the cache may do at any
moment — and the cell goes blank, with no way back short of re-fetching the
original over the network and re-detecting it. It also cost a full JPEG decode
per image, per visit, to show a 96px cell.

So the crop is cut once, when the pixels are already in hand at detection time,
and kept. A 160px JPEG is a few KB against the ~250 KB proxy it replaces reading.

Where it lives is the interesting part. The catalog snapshot is uploaded *whole*
on every sync and downloaded by every device, so a crop column there would put
tens of MB on every round trip — the exact cost `face_shard`'s 25 MB cap exists
to bound, and the reason bulk per-face data lives in shards already. Crops
therefore travel in the face shards, beside the embeddings, and
`snapshot_for_upload` strips them from the copy it writes. Nothing reads a crop
out of a merged remote catalog — the merge touches collections and keywords only
— so a receiving device loses nothing. A shard carrying crops holds around 3,500
faces rather than 22,000, which is the price of a second device showing People
immediately instead of re-fetching every proxy.

The column is nullable and the reader falls back to the proxy, so a face indexed
before this still works and the next indexing pass fills it in.

V10 also adds `people.ignored`, for a person the user has looked at and does not
want to identify. Most clusters in a real library are strangers — passers-by,
other people's guests, a face on a poster — and there is no way to tell "not yet
looked at" from "looked at, don't care" without recording the second. It is a
column rather than a deletion because a deleted cluster comes straight back on
the next Regroup: the faces are still there and still similar, and nothing short
of remembering the judgement survives re-clustering. Same argument
`face_person_rejected` makes one level down.

And `prune_empty_unnamed`, for what clustering leaves behind. Regroup creates a
person per unanchored group and never removed the previous run's now-empty ones,
so pressing it twice added a rail entry per group it no longer believed in.
Named people are never touched however empty — a name is user data — nor is a
merge tombstone, which must outlive its faces to keep redirecting.

298 tests pass, including that the snapshot carries no crops while the live
catalog keeps them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 21:17:44 +02:00
dtourolleandClaude Opus 5 b846b312b8 Run the formatter over the face branch before it reaches CI
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 2s
Build and test / Desktop (Linux) (push) Successful in 1h21m32s
Build and test / Layer separation (push) Successful in 37s
Traceability / Requirement traces (push) Successful in 25s
Build and test / Android (aarch64) (push) Failing after 33m58s
The merge of the SCRFD/MobileFaceNet work brought 69 rustfmt diffs across
dr-catalog, dr-face and dr-ui with it, so `cargo fmt --all -- --check` fails
on master and the Desktop job stops at its Format step — before clippy, the
tests or the release build have run at all. That makes the whole desktop
half of CI blind: a real compile error behind this would look exactly the
same from the outside. There was nothing behind it, as it turns out — with
the formatting fixed, clippy, the test suite and the release build all pass.

Every .rs hunk is `cargo fmt --all` on the pinned 1.92.0 toolchain, not a
hand edit, but it is worth being precise about what that moved, because it
is more than whitespace. Besides reflowing signatures and call chains,
rustfmt reordered the `pub mod` and `pub use` items in dr-face/src/lib.rs so
the `#[cfg(feature = "inference")]` entries sort in place, added the trailing
semicolon inside `let ... else { return }` bodies in identity_ui.rs, wrapped
a bare closure body in braces in cluster.rs, adjusted trailing commas, and
dropped a stray blank line at the end of identity_ui.rs. All of it is
semantically inert; none of it changes behaviour.

docs/traceability.md rides along because it has to. The matrix records each
TRACES tag by line number, and reflowing develop.rs, lib.rs, faces.rs,
identity.rs and identity_ui.rs moved them — FR-CAT-8, FR-CAT-9, FR-CULL-10,
FR-DEV-3, FR-DEV-3a and FR-DEV-3c all shift by a line or two. The matrix was
verified up to date on d777f7f before this commit, so this is drift these
formatting changes introduced, not pre-existing staleness being swept up.
Leaving it for a follow-up commit would hand traceability-check.yml a
failure caused entirely by a whitespace change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 13:12:49 +02:00
dtourolleandClaude Opus 5 96d07da15f Sync face data as sealed shards, so a second device does not re-index
Indexing 23,500 images is about two hours of CPU, and the result is
byte-identical on every device: the same model over the same proxy
produces the same embedding. Paying for it once per account rather than
once per device is the point.

Shards rather than the catalog snapshot, because the snapshot goes up
whole on every sync and a fully indexed library carries roughly 30 MB of
embeddings. That is exactly the cost the thumbnail store's 25 MB cap
exists to bound, so face shards use the same cap -- imported from
dr_thumbs rather than restated, since the number is a statement about
sync cost and the two must not drift apart.

The split follows the one already there: bulk immutable data in sealed
shards, small mutable data in the catalog snapshot. Faces, landmarks,
embeddings and run markers shard; people, names and assignments ride the
catalog and merge by uuid.

Keyed on oc:fileid throughout, never on image_id, because a row id means
nothing on another device.

The run marker travels with the faces it describes. Without it a
receiving device cannot tell an image with no faces from one never
examined, and would re-detect every landscape it had just adopted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 22:41:50 +02:00
dtourolleandClaude Opus 5 26a1eb7e28 Record that face detection has run, not just what it found
An image with no faces in it was indistinguishable from one that had
never been looked at, so every landscape, still life and document scan in
the library was re-detected on every pass, for ever. In a real library
that is most of it: on the 23,527-image test library, 64 of the first 110
images indexed contain no face at all.

Schema v9 adds face_index, a run marker per (image, model) carrying the
face count and the proxy edge it read. Keyed on the model, so a model
change puts every image back in the queue by itself.

That makes a coverage figure possible, which is the thing a user actually
wants to see. The audit also splits the outstanding set by whether a
proxy exists, because 23,417 awaiting a proxy and 110 ready to index are
different problems, and telling the user to run indexing again would not
fix the first.

The Identity screen gains Index faces, Stop, and the coverage line.
examples/face_index.rs is the same check and sweep without a window,
which is the right shape for an overnight pass.

Measured on the real library in release: 3.5 images/second, 110 images
and 125 faces in 30 seconds, and a second run correctly finds nothing
left to do.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 22:13:41 +02:00
dtourolleandClaude Opus 5 2ac069a6b3 Index the library's faces, and group them into people
Wires dr-face to dr-catalog: a background sweep that reads the proxy the
grid already built, detects, aligns, embeds and stores, then a clustering
pass that turns those embeddings into suggested people.

Detection runs on the Large thumbnail tier and nowhere else. That is what
makes the feature affordable -- a browsed library has already paid for
its proxies, so face indexing adds no RAW decode that was not already
happening -- and it is why an image whose proxy is missing is skipped
rather than fetched: requesting one here would put face indexing on the
network path FR-CULL-8 keeps it off.

The sweep keeps no cursor. It asks the catalog what is missing, so it
resumes after process death with no repeated work beyond the in-flight
image, and cancelling is dropping the receiver.

recluster writes only the suggested half. Confirmed faces go in as
anchors and come back untouched, and a cluster of one stays nameless --
naming every stray face would fill the People view with noise the user
then has to dismiss.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 20:49:30 +02:00
dtourolleandClaude Opus 5 00e78dc2ac Cluster faces into people, and calibrate what a similarity means
FR-CULL-9 forbids thresholding a bare cosine anywhere in the subsystem,
so calibrate fits P(same person) per library and reports whether the fit
is trustworthy. Two details carry most of the weight.

The fit runs against a 200-bin histogram rather than a pair list: a
25,000-face library has ~3e8 pairs and no gradient descent is running
over that. And a fresh library has no valid calibration, because the
positives have to come from user confirmations or burst siblings --
bootstrapping them from high cosine would fit the calibration to the
belief it was supposed to test.

Clustering defends against the over-merging FR-CULL-10 warns about with
constraints rather than a better threshold: two faces in one photograph
never merge, and two groups confirmed as different people never merge.
Average link rather than single link, so one strong edge cannot weld two
families together.

Calibration is defined once, in dr-face, and dr-catalog re-exports it.
Two implementations of one probability model is exactly how a number
comes to mean the wrong thing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 20:10:33 +02:00
dtourolleandClaude Opus 5 aac3136407 Store faces and the people they belong to
Schema v8: people, faces, face_person, face_person_rejected, and the
per-library calibration. Follows catalog.md 10.1 with two additions the
spec work turned up.

crop_px, because at the 1024px proxy tier a group shot reaches the
embedder at ~50 source pixels upsampled to 112 and a portrait at 340.
FR-CULL-9 names face size as an axis along which an uncalibrated
similarity misbehaves, so it is a stored feature rather than a UI hint.

face_person_rejected, because rejection is not the absence of an
assignment. Without it the next clustering pass re-suggests exactly the
face the user just pushed away, and the tool feels broken.

record_detections replaces rather than appends, since DetectFaces is
coalesced per image -- and carries confirmations across the replacement
by box overlap, so re-indexing with a better model cannot discard the
user's own labelling.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 20:02:52 +02:00
dtourolleandClaude Opus 5 6d6ef8d34b Page the grid along an index instead of sorting the library each time
Scrolling jittered, and this was the largest single reason. Every window the
grid loads is `ORDER BY ... LIMIT n OFFSET k`, and neither half of that was
being answered the cheap way.

**The sort.** `GRID_ORDER` leads with `captured_at IS NULL`, so undated frames
fall to the end. No ordinary index answers that — the leading term is an
expression, not a column — so SQLite sorted the whole library into a temp
b-tree on every window read, then threw away the first `k` rows of it. Schema
V7 indexes the expression exactly as the query writes it, partial on the same
`shadowed_by IS NULL AND trashed_at IS NULL` the grid filters by, so the read
becomes a walk along the index.

**The join.** `LEFT JOIN remote` was paged *after* it was joined, so reading
280 cells at offset 20,000 first seeked into `remote` for all 24,000 rows and
then discarded 23,720 of them. The file ids are now fetched for the 280 rows
that survived — the shape the badge and rating reads already use, one query for
the window rather than one per cell.

Measured together on 24,000 images at offset 20,000: **15.2 ms → 0.36 ms**,
inside a scroll handler that has 16.7 ms to draw a frame.

The test asserts on the query plan rather than on a duration, because there is
no other symptom. A `GRID_ORDER` edited out of step with the index, or a column
added back that drags `remote` in again, both still return exactly the right
cells — just after sorting the library — and the jitter would come back with
nothing to point at.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-25 20:01:04 +02:00
dtourolleandClaude Opus 5 fa4ad6e2d6 Drag the date range on the axis it is chosen from
The range could be turned on with a finger and not aimed with one. Its two
ends were typed as `YYYY-MM-DD` into 108px fields behind a soft keyboard, to
name days already drawn on the axis a thumb away; and the chip that seeds them
takes its span from the timeline's zoom and pan, which are a wheel and a middle
button. A touch screen has neither, so on Android the filter was a switch with
no aim.

The band is now on the timeline. Two ends with grips, dragged along the bars,
released to filter — the histogram was already how a period is found, and this
makes it how a period is stated. Both ends snap to whole days, which is what
the typed fields mean, what `show_range` reads back out, and a floor under a
range dragged shut. The fields stay for what dragging cannot do: name an exact
day, and say in words what the range is.

For that to work the axis had to stop following the range. Redrawn to the band,
it moved the ground under the very handles doing the narrowing, and there was
nothing outside the range left to widen back into.

While there: a fixed number of equal bins instead of calendar buckets. Between
one calendar unit and the next the bar count is free to wander by a factor of
twelve, so zooming in halved it two steps out of three — the same picture drawn
wider until it jumped back to fine. Equal bins also include the empty ones, so
a bar's position on the track and the date under it are finally the same
quantity; before, a library with gaps drew a February six months wide and the
marker, the band and a click all pointed somewhere else. The count is a
setting, 32 or 64, because the right answer is a question about the screen.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-23 18:42:42 +02:00
dtourolleandClaude Opus 5 4b7648082e Let a date range be stated, and draw the axis at the scale it deserves
"Limit to range" did nothing, and the reason was not visible from the button.
It took its span from the timeline's zoom, which is zero until someone zooms —
so `zoomed_span` returned the whole library and the filter narrowed to
everything. The chip lit up and the grid did not change.

The range has ends now, shown and typed as `YYYY-MM-DD`. Seeding them from the
timeline is kept, because zooming to a fortnight and pressing the chip is the
fast path; the fields say which fortnight it landed on and let it be corrected.
Ends given backwards are swapped rather than refused — there is exactly one
range between two days — and the closing day is included, since "to the 5th"
means the whole of the 5th and a range ending at its midnight contains none of
it. `parse_date` refuses anything that is not a date rather than guessing at an
order, because the alternative is a library silently filtered to a span nobody
asked for.

The axis then follows the range. It used to keep drawing the full extent while
a range was on, because it was the only way back out; the typed ends are the
way back out now, so it is free to show what was asked about.

And bucket size is chosen by how many bars it makes rather than by fixed
cut-offs. Each zoom step halves the span, so under thresholds the bar count
halved with it until a boundary was crossed: fifteen years went 15 bars, 8,
then 46, 23, 11, and finally 6. Zooming in made the picture coarser, which is
the opposite of what zooming is for. Aiming at forty bars keeps the count in
the same neighbourhood at every level, and the test asserts the property
directly — halving a span never coarsens the bucket.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-23 12:40:14 +02:00
dtourolleandClaude Opus 5 67f15bffb7 Key collection membership on the identity that exists
Collections synced their names and arrived empty on every device. The names are
keyed on a uuid and worked; the membership union was keyed on
`images.content_hash`, and the schema says plainly what that column is:
"computed only when something needs it (import dedup, reconnect-by-hash), never
in a scan". A library that has only ever been scanned has one for no image at
all, so the join matched nothing and `WHERE ri.content_hash IS NOT NULL`
discarded whatever survived. The union could never have moved a single row.

Measured on a real catalog: 23,174 images, content hashes for 0 of them,
`oc:fileid` for all 23,174, twelve collections, zero members.

So membership now resolves through the file id first, exactly as keyword
assignment already did — `ASSIGN_BY_FILE_ID` was added for this same reason and
its doc comment even notes that membership was still on the hash. It is
recorded for every image the moment a remote scan sees it, survives server-side
rename and move (FR-NC-5), is the same integer on every device pointed at one
Nextcloud, and is already what the thumbnail shards are keyed by.

The content-hash union is kept rather than replaced: a local-only library has no
`remote` rows, and where a hash has been computed it is a true identity that
survives a library moving between servers. Both statements run; `INSERT OR
IGNORE` against the primary key makes the overlap free.

This repairs the merge. It cannot invent membership that no device recorded —
where the rows were never written, collections stay empty until they are filled
in again.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 21:19:13 +02:00
dtourolle 56dd3187f1 Merge integration into wip/ingest
Second pass, against the detail-stage and thumbnail work that has landed since
the first. Resolved and verified here rather than in the shared merge worktree,
so what goes back is a fast-forward.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

# Conflicts:
#	docs/traceability.md
2026-08-22 19:44:42 +02:00
dtourolleandClaude Opus 5 62188ec740 Keep both devices' keywords when the catalogs meet
Keywords are catalog state, and the catalog syncs. Without this, two devices
keywording the same library would resolve to whichever synced last, and an
afternoon of work would vanish with no sign it had ever happened.

The vocabulary merges per row on the rule collections already use: revision
first, timestamp only to break a tie, so a device with a skewed clock cannot win
by having the wrong idea of the time. Assignments merge as a set union, which is
FR-NC-9's principle applied to metadata instead of edit nodes — disjoint work
survives on both sides.

Three things needed care and are commented where they happen:

A deletion travels *by name*, not by identity. Both devices may have minted
their own uuid for one word before they ever synced, so deleting by uuid would
tombstone a row nothing was assigned to and leave every photograph still
carrying the word. The union then refuses to readmit a word a winning tombstone
has just removed — without that filter the remote's live assignments would
resurrect it on the very same pass.

Images are resolved by the server's file id first and the content hash second.
Membership has always used the hash alone, but the hash is computed only when
import dedup or a reconnect asks for it, which for most libraries is never — so
a hash-only union would have quietly done nothing for the ordinary photograph.

A word lands on the local default version. Version uuids do not reconcile in the
catalog at all: ensure_default_versions mints a fresh one per device, so a
uuid-keyed join would have unioned nothing.

Removal still does not propagate. That is the trade collection membership
already makes, for the same reason — an unwanted keyword is removed again in a
second, and a silently lost afternoon is not recoverable at all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 15:51:00 +02:00
dtourolleandClaude Opus 5 2147eaa6a5 Put a keyword on a photograph, not only search for one
The catalog has been able to *find* by keyword since v1 — query.rs joins the
keywords table, matches it exactly, and substring-matches it for free text — and
nothing anywhere could ever put a word there. A user could filter to a keyword
they had no way to apply.

This is the missing half: create, rename, delete, list, assign, unassign, and
the two reads a panel needs. Bulk-only for assignment, because keywording a
selection is the common case rather than the exception — the photographer picks
out the frames with the puffin in them and applies "puffin" once, in one
transaction.

Schema v6 adds `keyword_terms`, and deliberately does *not* touch the v1 join.
The assignment keeps the word as text because the catalog is a rebuildable index
and the durable copies of that fact — the sidecar, XMP dc:subject — both carry a
string; a foreign key would mean a catalog rebuilt from sidecars had to invent
identity rows before it could record anything, and would break the query path
that already works. So the text is the fact, and the new table is only the
identity a rename and a deletion can be keyed on.

`keyword_terms.name` carries no unique index, which looks like an oversight and
is not: two devices that each type "Iceland" are both right until they meet, and
a constraint would abort the merge at that moment. Uniqueness is converged upon
instead — create resolves an existing name, fuse_duplicates collapses a
cross-device pair onto the smaller uuid.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 15:50:49 +02:00
dtourolle 735683b849 wip: ingest 2026-08-22 14:12:40 +02:00
dtourolleandClaude Opus 5 02d629922f Draw the date histogram over the collection you are looking at
Build and test / Desktop (Linux) (push) Failing after 57m25s
Build and test / Layer separation (push) Successful in 33s
Traceability / Requirement traces (push) Failing after 27s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
Build and test / Android (aarch64) (push) Failing after 9m45s
The timeline counted the whole library whatever the grid was showing, so
opening a collection left a fortnight in Arosa as one column of a
fifteen-year axis — an axis describing photographs that were not on
screen.

Scope the buckets and the span to the same collection and rating filter
the grid uses. `timeline_range` counts `images` alone and cannot express
the membership join, so the scoped query lives beside the other scoped
readers in the UI and shares their descendants-of-scope rule.

`catalog_span` now delegates to the same scoped reader. Zoom and scrub
measured the full library while the bars were scoped, so a scrub could
land on an instant the collection did not contain and send the view
somewhere the user had not asked to go.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 22:15:59 +02:00
dtourolleandClaude Opus 5 b0206cbc7a Let a sub-collection stay under its parent through a sync
Build and test / Desktop (Linux) (push) Failing after 57m14s
Build and test / Layer separation (push) Successful in 33s
Traceability / Requirement traces (push) Failing after 29s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
Build and test / Android (aarch64) (push) Failing after 9m47s
The merge inserted every incoming collection with `parent_id = NULL` and
never set it on update, so the hierarchy flattened on each round trip: a
collection nested on one device came back from the server at the top
level. `r.parent_id` was selected and then not read.

The id could not be copied — row ids are local, and the remote's integer
names a different collection here, or none. So carry the parent's uuid
and resolve it locally, in a second pass: rows arrive in whatever order
the query returns, and a child can precede its parent.

Guard the resolution against cycles. Each tree is acyclic alone, but the
union need not be — we may hold A above B while the remote holds B above
A — and closing that loop would make every tree walk spin.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 21:55:09 +02:00
dtourolleandClaude Opus 5 02ae92ba0d Tidy what the seven-branch merge left behind
Build and test / Desktop (Linux) (push) Failing after 57m16s
Build and test / Layer separation (push) Successful in 34s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
Traceability / Requirement traces (push) Successful in 1m7s
Build and test / Android (aarch64) (push) Failing after 9m41s
Three lints, all from merged work rather than from any one branch:
`terrace` and `disc` were steps on the way to the ramp the plateau test now
uses, and the reasoning that discarded them lives in docs/segmentation.md §12
rather than needing the code; two mechanical clippy suggestions in segment and
cache.

1176 tests pass, clippy and fmt clean, traceability regenerated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 12:36:15 +02:00
dtourolleandClaude Opus 5 d913e50948 Select photographs with a finger, and take a collection with you
Two things a tablet could not do. Both existed for a pointer and had no
touch form at all, which on Android meant the collection sidebar was
somewhere to look at rather than somewhere to file into.

**Selecting more than one.** Ctrl-click and shift-click are the only ways
into a multi-selection, and touch has neither. Holding a cell now enters
selection mode, where a tap toggles — reported to Rust as a ctrl-press, so
it goes through the same `apply_press` as everything else rather than
growing a second copy of the selection rules. A double tap takes the run
between where selecting began and there: the touch form of shift-click,
and the reason the anchor from *before* the double tap has to be
remembered, since both of its taps move the anchor onto the cell being
tapped. A "Select" button does the same thing where a gesture would go
undiscovered (FR-UI-4).

**Filing without a drag.** A one-finger drag beginning in the grid belongs
to the Flickable that scrolls it — that is the arbitration working, not a
bug to route around — so the selection can now be filed from a sheet
listing the sidebar's own rows. Copy by default, as the drag has always
been; moving out of the collection being shown is a switch, because it is
the one that takes something away.

**Taking a collection offline.** The machinery was there and reachable only
by scoping the grid to a collection and finding a button behind a
disclosure. Holding a collection's name now asks the question directly, and
the tray on a row and the header button ask the same one — three
affordances doing two different things is how a user comes to avoid all
three. The question is asked rather than a toggle flipped because both
answers are expensive: one downloads gigabytes, the other deletes them, and
the counts and sizes go in the buttons where they are read before the tap.

`Cache::release` is new and is the destructive half `unpin` deliberately is
not. "Remove the local copies" is asked by someone whose device is full,
and withdrawing a promise while leaving the bytes for a future eviction to
notice is not an answer to it. It unpins before forgetting, or the next pin
fetch would dutifully download everything it just deleted.

The sidebar's trays read `tier_actual`, never `tier_desired`: the question
is whether these will open on the aeroplane, and a pin whose download has
not run yet answers no.

TRACES: FR-CAT-7 | FR-NC-6a | FR-NC-6b | FR-NC-6c | FR-UI-2 | FR-UI-3 | FR-UI-4

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 12:27:29 +02:00
dtourolleandClaude Opus 5 4b36ca66aa Render an export into the colour space its file will claim
The colour-managed export branch left one call site deliberately unfixed, and
this is it. `render_for_export` composed with the default sRGB shader, so a
Display P3 export failed with an accurate error rather than producing a
mislabelled file — the right way to leave a half-finished path, and no way to
leave it.

The space is chosen at render time because that is the only time it can be:
the conversion happens in the shader, before the clip to 0..1, so by the time
pixels reach an encoder they are in exactly one space and the only honest
thing left is to label them. `Frame::in_space` carries which, and a mismatch
between what was rendered and what was asked for stays a typed error.

Also regenerates the traceability matrix over the four merged branches.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 09:34:57 +02:00
dtourolleandClaude Opus 5 bb71f141e7 Let a folder on this machine be scanned into the catalog
`dr_catalog::scan` has known since it was written what a changed directory
means — when to prune, when to list, and the one question that decides whether
a deletion sweep is safe. It was fully tested and nothing called it, because
walking a real directory "belongs to the platform layer" and the platform layer
was eleven lines re-exporting `secrets`. So every photograph in DarkRoom
arrived over WebDAV, and a user without a Nextcloud account saw nothing at all.

This is the missing half: a `Storage` trait, a filesystem implementation of it,
and the driver that pours one into the other.

The trait is shaped by the platform it does *not* yet support. Android's SAF
gives no filesystem path, which is why `SourceRef` exists; less obviously, it
gives no way to *compose* one either — a document id is opaque, and the only
way to learn a child's id is the children query that returned it. So a listing
hands back the reference to each entry rather than a name for the caller to
join onto a parent, and there is deliberately no "path + name" helper anywhere
above `LocalStorage`. That single restriction is what makes SAF a second
implementation rather than a second set of call sites. A reference is otherwise
an opaque `(RootId, key)` pair the catalog stores verbatim and rebuilds later,
which a persisted tree grant supports exactly as a relative path does.

A `Path` now appears in one place: `LocalStorage::grant`, where the folder the
user picked is handed in. Everything above it addresses a `RootId`.

`dr_catalog::walk` is the seam. It probes a directory, asks `scan` what that
means, lists only when told to, and reconciles what it found against the rows
it holds. Two things it does are worth saying out loud, because both are ways
to lose a library:

Absence only counts where absence was observed. A listed folder proves its
missing images are gone; a pruned one proves nothing about its contents, and a
scan that was cancelled or that failed part-way proves nothing about folders it
never reached. So the file sweep runs per listed folder, the folder sweep runs
once at the end and only after a complete scan, and a root that cannot be
reached at all marks its images offline and deletes nothing — FR-CAT-9's line
between proven-absent and merely-unreachable, which is the difference between
unplugging a drive and losing everything on it.

A trashed image is absent from its folder on purpose. It is exempt from both
sweeps, and detached from a folder about to be deleted rather than cascaded
away with it, or a soft delete would come undone the first time the folder it
came from was rescanned.

Two things the tests taught, both changes to what was there before:

Modification times are now milliseconds, not seconds. Change detection asks
whether a timestamp moved, so the unit's granularity is the width of the window
in which a change is invisible — and a second is long enough to copy a card and
start a scan. The test that caught it looked like a test bug; it was not. SAF
reports milliseconds natively, so this is also the unit that needs no
conversion on the platform with the coarser clock.

And an in-place rewrite of an existing file is invisible to directory-level
pruning, because writing to a file moves neither its directory's mtime nor its
entry count. That is a real limit, now documented and held by a test rather
than left to be discovered. It bites less than it reads: an export, a restore,
`mv`, and every editor that saves safely write beside the file and rename over
it, which does move both.

Narrowing the format filter no longer deletes what it stops matching, which
fell out of the same principle: unticking JPEG says stop looking for new ones,
not discard the hundred already rated. The files are sitting right there.

`DirState` and `DirEntry` move to `dr-types`. They are the sentence the
platform says to the catalog and both crates need the same one; `scan`
re-exports them so nothing that used them has changed.

Not done: the UI. The launch screen's "Open library" flow is account-shaped
from the first field to the thumbnail worker, and giving it a local branch is
its own piece of work rather than a button. `cargo run -p dr-catalog --example
scan_local -- ~/Pictures` scans a real folder and reports what it cost; run it
twice to see the second run list nothing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 08:59:17 +02:00
dtourolleandClaude Opus 5 03326242a1 Make the CI checks say what they mean, and format the workspace
Build and test / Desktop (Linux) (push) Successful in 1h23m26s
Build and test / Android (aarch64) (push) Failing after 2s
Build and test / Layer separation (push) Successful in 50s
Traceability / Requirement traces (push) Failing after 1m9s
The Android job's "Verify minimum API level" step has never verified the
minimum API level. It took the first `*.so` anywhere under the target
directory, which is a host proc-macro from debug/deps — an x86-64 object
built by the runner's gcc, whose .comment section cannot mention Android
and so can never contradict the expected value. It now reads the artifact
under the target triple, compares against MIN_API parsed from the
Dockerfile rather than a second copy of the number, and fails on a
mismatch. Both sides are checked non-empty first: two failed parses would
otherwise compare equal and pass, which is the same silent success in a
new costume.

The Android image installs one SDK package per layer and keeps the
output. sdkmanager is a JVM program that aborts when it cannot get memory,
and the single `> /dev/null` step reported that as a bare "exit code 134"
while a retry re-downloaded everything that had already succeeded.

tools/ci-local.sh runs all four jobs — desktop, android, layering,
traceability — against the host toolchain, which is pinned to the same
1.92.0 CI installs. Its matrix check compares regeneration against the
working tree rather than against HEAD: CI starts from a clean checkout, so
git's answer is the right one there and reports every local run stale here.

The rest is rustfmt across the workspace, and the clippy findings that
surfaced once it did: manual_contains in dr-thumbs and collections_ui, a
map iterated as pairs for its keys, an index loop over a slice, and two
runtime assertions on a constant now made at compile time.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 12:02:51 +02:00
dtourolleandClaude Opus 5 fa12afed18 Keep originals on this device, by pin and by use
Build and test / Desktop (Linux) (push) Failing after 1s
Build and test / Android (aarch64) (push) Failing after 0s
Build and test / Layer separation (push) Failing after 1s
Traceability / Requirement traces (push) Failing after 2s
Fills in `image_cache`, which the previous commit's "On this device" filter
read but nothing wrote. Also carries in-flight work that shared these files:
the Android TLS root store, the settings page, and a regenerated
traceability report.

# Two populations, deliberately separate

An original is kept here for one of two reasons, and conflating them produces
the exact failure the feature exists to prevent.

**Pinned** originals were asked for. Pinning a collection before a trip is a
promise, so pinned rows are never evicted and never counted against the
budget — a cap that could silently delete a pinned trip would make pinning
worthless, because it could not be relied on without checking.

**Passively cached** originals are a side effect of working: develop already
downloads the whole file, so keeping it costs no bandwidth and saves the
entire transfer next time. This population is what the budget bounds, evicted
least-recently-used, because it otherwise grows until a day of culling fills
a disk.

Sharing one budget would let a large pin starve the passive cache, or let
browsing evict a pin. They are separate.

# What was built

`dr_catalog::cache` owns the bookkeeping — held tier, size, last use, pinned
— and writes the bytes; deciding to download stays with the caller, which is
what keeps a crate with no network out of the network's business. Files are
written to a temporary and renamed, so a dropped connection cannot leave a
truncated file recorded as a complete original. They are named by image id,
not filename: `Photos/IMG_0001.CR2` and `Trips/IMG_0001.CR2` are different
photographs, and a flat cache keyed on the name would serve one for the other.

`spawn_full_fetch` became read-through. A hit is a disk read; a miss stores
what it downloads and enforces the budget. A cache that cannot be opened is a
miss, not a failure to open the photograph.

Pinning writes intent — `tier_desired` — without downloading, so the button
responds immediately, and `spawn_pin_fetch` fills it in sequentially
afterwards. Sequential because these are tens of megabytes each: the lanes
that make the thumbnail sweep fast buy little against one connection's
bandwidth and cost a great deal of memory. A pin interrupted by a lost
connection resumes from where it stopped.

Schema v5 adds `pinned` and `path`. `pinned` is a column rather than something
inferred from `pinned_by_rule`, which is ON DELETE SET NULL and so cannot
answer for an image whose rule was deleted. A v4 catalog migrates in place;
existing rows default to unpinned, the safe direction.

The budget and "keep opened originals" come from the settings page rather than
a constant, and are applied at startup rather than only on change — a cache
capped at 2 GB last session would otherwise spend this one filling to the
default. Turning off keeping leaves what is already cached readable: those
bytes are paid for, and refusing them would re-download images sitting right
there, including pinned ones.

Also removes a doubled `#[test]` introduced in the previous commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 21:12:01 +02:00
dtourolle d5b1f6bff5 Add collections, ratings, and soft delete to the catalog
Three features over a shared schema migration.

Collections: a tree of manual collections plus smart collections whose
membership *is* their stored selector. Dropping images onto a smart
collection is refused rather than silently discarded, so the UI can say why
the drop did nothing — member rows there would be a second source of truth
that nothing reads.

Ratings: the star and pick/reject axes, kept independent.

Trash: soft delete to a folder, then permanent delete.

Catalog::open now backfills after migrating. A migration adds a column but
cannot know what the value should be for rows that already existed;
backfilling on open is what stops those rows being silently partial.

Timeline queries exclude shadowed JPEGs, which would otherwise double every
paired shot in the histogram, and gain a range-bounded variant so zooming in
returns finer buckets rather than the same coarse ones with the ends cropped.

Assisted-by: LLM
2026-08-09 20:42:13 +02:00
dtourolle c8bb08e661 Add folder scan with format selection; validate A3 on a real library
Library setup as the user described it: pick a folder, choose which RAW
types to look for, scan recursively.

  dr-types::FormatFilter  the tick-box selection, seeing through VFS
                          placeholder suffixes so a dehydrated CR2 still
                          matches as a CR2
  dr-sync::scan           recursive walk, Depth:1 per directory, pruning
                          unchanged subtrees where the backend propagates
                          directory ETags

Verified against nextcloud.tourolle.paris (34.0.2) on a real library:

  browse root      32 entries, 98ms
  scan PhotosRaw   17,185 RAW files in 334 directories, 34.1s
                   (7,836 CR2 + 9,349 DNG)
  range read       262KB of a 21.5MB DNG in 119ms — 1.22% of the file,
                   and enough to read "Canon EOS 6D | ISO 100"

That last line is assumption A3 validated on real data. Cataloguing this
library by whole-file fetch would move roughly 370GB; the range path
moves a few MB.

Pruning is capability-gated rather than assumed: with per-entry ETags a
probe costs a request and proves nothing about children, so it is skipped
entirely. A test asserts zero probes in that case.

Still unresolved: /core/preview returns 400 for every parameter
combination tried, including on a JPEG the server reports as having a
preview. Not a request-shape bug — it fails identically bare. Recorded
rather than worked around; ARCH §6.7 already treats server previews as
opportunistic, so nothing depends on it.
2026-08-09 12:22:31 +02:00