Commit Graph
61 Commits
Author SHA1 Message Date
dtourolle adf5d6cdd9 Drop a rival pipeline's marker when an image is re-indexed
record_detections replaces every face on an image whatever model found
them, but left the other models' face_index rows standing. With one
model that was unobservable. With a second pipeline it leaves an image
marked "done" under the first with none of its faces behind the marker
— the state the V12 repair existed to undo — and a user who switched
back would find those photographs permanently empty.

An image now holds the faces of whichever pipeline looked at it last,
and only that pipeline's marker. Confirmed names still carry across by
box overlap, since they were read before the replacement.
2026-09-11 22:12:40 +02:00
dtourolle 16f3fb41a3 Measure the faces already found rather than finding them again
Benchmarks / CPU and I/O (per commit) (push) Successful in 3m13s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 54s
Build and test / Layer separation (push) Failing after 1s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 2s
Traceability / Requirement traces (push) Successful in 52s
Build and test / Android (aarch64) (push) Successful in 30m6s
Every face stored before its quality was kept holds a unit vector, and
V14 forgot the run marker of each image holding one so that the next
sweep would look again. Looking again meant detecting again: a whole
re-detection per image, with every suggestion on it thrown away and the
confirmations carried across by box overlap, to recover one number.

The sweep now has a measuring pass between the proxy repair and the
un-indexed images. It lists every image holding an unmeasured face,
fetches the original once, warps each stored face from the landmarks it
already has, embeds it, and writes the raw vector and its length over
the old row. Ids, boxes and identities are untouched; the marker is
re-written fresh so the sync exports the measured vectors. A face whose
landmarks no longer make a warp is dropped, as detection would have
refused to store it. `faces_unindexed` leaves those images to the
measuring pass, so the V14 deletion no longer costs a second detection.
2026-09-11 21:50:12 +02:00
dtourolle 8b3abdb787 Keep each face's quality, and never compare against a poor one
The embedder's raw output has a length, and the length is a reading of
how recognisable the crop was: a blur, an occlusion or a hard profile
comes out short. Normalising threw it away. A short vector sits near
the middle of the sphere and matches a little of everyone, which is how
one bad crop bridges two people in a grouping pass.

So the length is kept — the store now holds the raw vector, re-normalised
on load, with the length beside it as `faces.quality` — and a face under
MIN_GALLERY_QUALITY (14) is a probe: measured against the gallery and
placed where it fits, but never what another face is measured against.
Two probes are never paired, and a probe is nobody's evidence for a
confidence. The People screen shows the number as "Quality 17.3", dimmed
below the floor.

Faces indexed before this stored unit vectors and have no reading; they
are admitted to the gallery, and schema V14 forgets the run marker of
every image holding one so the next indexing pass measures them. A
peer's unmeasured shard faces are not adopted, or a sync would write
that marker back.
2026-09-11 21:50:12 +02:00
dtourolle bac5801618 Ask which collections a selection is filed in, and how much of it
`collections_for_image` answers this for one photograph and has no
counts, which is enough to badge a cell and not enough to offer a
removal: with forty selected and three of them in "Iceland", a sheet
that says only "Iceland" invites the user to take all forty out of a
collection thirty-seven were never in. `membership_of` returns the
count alongside the name so the row can say "3 of 40".

Chunked over the image list rather than one `IN (...)`, because the
list is a selection and a select-all makes it as large as the library —
past SQLite's bound-parameter cap on exactly the gesture most likely to
produce it. Counts are summed across chunks, so the answer is the one
the unchunked query would have given.

Smart collections are excluded by construction: they have no member
rows, so there is nothing a removal could do.
2026-09-07 19:59:44 +02:00
dtourolleandClaude Opus 5 1ad35e2b87 Read the library's sidecars, so a cull done elsewhere arrives
Judgements only ever travelled outward. A rating went to the catalog and to the
photograph's sidecar, the sidecar reached the server, and there it stopped: the
scan indexes files, `derived_sync` exchanges thumbnails, face shards and
collections, `dr_catalog::merge` reconciles everything in a catalog except
`versions.rating` and `versions.flag`, and the one sidecar reader that existed
ran when a single photograph was opened in develop and handed its answer to the
develop graph. `JobKind::ReadSidecar` was declared for exactly this when the job
queue was written and was never enqueued or handled anywhere.

The grid draws `versions.rating`. So a day of culling on the tablet could not
reach the laptop by any path the application had, and the laptop's catalog says
so plainly: 23,568 images, one of them judged.

`pull_sidecars` closes it, off the back of work the scan already does.
`dr_sync::scan` reports the `.drsc` files it meets in listings it was making
anyway — no extra request, and a directory whose ETag is unchanged is still
pruned before it is listed at all. A new `sidecars` table records the ETag of
each one this device has taken in, so the fetch is one GET per sidecar that
genuinely changed rather than one per photograph. A library nobody has edited
costs nothing.

The judgement is taken rather than maximised. The sidecar is the authoritative
store and the fuse has already settled any contest between devices on
`revision`, so lowering a rating from four to one on the tablet lowers it here —
taking the larger would have refused every demotion the photographer ever made,
which is most of what a second pass over a shoot is. A zero is the exception: it
means *never judged*, not "judged zero", so a sidecar carrying none cannot erase
a star this device holds. That is `merge_judgement`'s asymmetry and it carries
the same known cost — clearing a rating does not propagate.

A sidecar names a stem, so both halves of a RAW-and-JPEG pair are judged: they
are one photograph (FR-CAT-11) sharing one document, and judging only one of
them would leave the grid disagreeing with itself over which it drew. The `LIKE`
that finds them is a filter, not the decision — `sidecar_path` is applied to
every candidate, because a folder is entitled to contain a `%` and a rating
landing on the wrong frame would be silent and permanent.

Failing to read one is not a failure to scan: the ETag goes unrecorded, the
ratings already here stay where they are, and the next scan tries again. The
count is reported to the status line as well as the log, because a grid that
silently gains three hundred stars is indistinguishable from one that has gone
wrong — and because while this number was structurally zero there was nothing to
tell the photographer their cull had not arrived.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-05 16:14:03 +02:00
dtourolleandClaude Opus 5 ca2a135e28 Let two devices name the same photograph's version the same way
A version's uuid is the identity a cross-device merge keys on, and it was
minted at random, per catalog, per image. Two devices indexing one Nextcloud
library therefore held two different uuids for the same photograph — so the
sidecar they shared collected a `default = 1` block each, `Version::merge` was
never handed a matching pair to reconcile, and an afternoon's culling on the
tablet did not exist as far as the laptop was concerned.

`crate::merge` has said so in a comment since it was written: version uuids do
not reconcile across devices, a uuid-keyed join unions nothing, so keywords are
landed on the local default version instead. It named the problem and worked
around it. `rating`'s own comment asserted the opposite — that generating the
uuid here was what made it a cross-device identity — and `library::amend`
repeated the claim. Uniqueness was never the difficulty; agreement was.

`derived_version_uuid` computes it from `oc:fileid` instead. The server assigns
that integer, every client pointed at the library sees the same one, and it
survives a server-side rename and move — the three properties that already made
`ASSIGN_BY_FILE_ID` prefer it to a content hash. The layout is a UUIDv8 (RFC
9562, an application-defined form) carrying all sixty-four bits verbatim across
the variable fields with a fixed tag in the node field, so the mapping is
injective by construction rather than by a hash's good behaviour, and a uuid in
a sidecar can be read back to the file it belongs to by eye.

A library with no server behind it has no shared identity to derive and keeps a
generated one. The split is still reachable there if the folder is synced by
something else; `Sidecar::fuse_default_versions` repairs that case rather than
preventing it.

Deriving it for new rows alone would have fixed nothing — every image in an
existing library already has a version, so every one of them would have carried
on writing to its own rival identity. `align_default_version_uuids` moves them,
and runs from `schema::backfill` on every catalog open. It selects on the tag
in SQL, so a catalog already realigned matches no rows and writes nothing, and
it declines rather than fails where a virtual copy already holds the target.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-05 16:13:58 +02:00
dtourolleandClaude Opus 5 d44bffa4a8 Remember where the photographer was
Opening the application was always a fresh arrival at the beginning of
the library, whatever you had been doing when you closed it.

What is written down is the view, the scope, the rating filter and the
photograph on screen -- the open one in develop, the first visible one in
the grid. Not just a scroll position: a position without the filter that
produced it names a row of a list that no longer exists. Restoring them
has an order for the same reason -- scope, then filter, then position,
then the view -- because each step changes what an ordinal *means*.

Addressed by remote path and collection UUID, never by an ordinal or a
row id. `images.id` and `collections.id` are local to one catalog, and a
grid ordinal is local to one ordering; a record naming either would land
somewhere arbitrary on a second device and after any filter change on
this one. Where the ordinal is needed, `library::ordinal_of_path`
computes it through the grid's own `ORDER BY`, taken verbatim by a window
function rather than spelled a second time as an inequality -- which is
the mistake `grid_order_for` already warns about, and which a manually
ordered collection would make unreadable.

Every failure degrades rather than reports. A collection this device has
not merged leaves the scope at the whole library; a photograph that has
since been deleted falls back to when it was taken, which puts the grid
in the right week; a torn file yields no place and the library opens at
the top. Reopening develop is the one thing that requires an exact match,
because a canvas on a path that no longer resolves is a filename over an
empty frame.

The record lives in `dr-types` beside `Settings` and the store lives here
beside `SettingsStore`, for the reason `dr-types`' manifest gives: a JSON
serialiser in `core/` would be paid for by every crate there. Two files
and two lifetimes, though -- resetting preferences must not forget where
you were.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 20:39:50 +02:00
dtourolleandClaude Opus 5 e7b526c550 Specify face indexing at native resolution, and say what the proxy cost
FR-CULL-8 said detection runs against the thumbnail or proxy tier and
never a full decode, and faces.md §5 said the aligned crop is sampled
from that same proxy. Both are wrong in the same place: they treat
detection and cropping as one resolution problem when they are two, with
opposite answers.

Detection does not care. §4.1 fixes the graph's input at 640x640 and
letterboxes whatever arrives, so a face filling 2% of the frame reaches
the model at 12px whether the buffer handed over is 1024px or 6000px.
Every pixel above the detector's own input is discarded before inference.

The crop cares about nothing else. §5's warp produces the fixed 112x112
ArcFace sees, so source resolution converts directly into whether those
112 pixels were photographed or interpolated. Reading crop_px across the
18,671 faces the proxy-tier implementation stored: 47.3% were upsampled
to reach the embedder, 314 of them by more than 2x, the smallest from 34
source pixels. An upsampled crop does not fail loudly -- it yields a
confident embedding of detail that was never there, and the damage
appears three stages later as clusters that will not separate.

So FR-CULL-8 now specifies four stages with the resolutions named
separately: render native through FR-EXP-9's pipeline, downscale for the
detector, map boxes and landmarks back to native, crop and align from
the native render. The affordability the old rule bought is met instead
by when the pass runs -- background, preempted, resumable -- and the
requirement says plainly what it now costs on a remote library: the
original rather than FR-NC-3's byte range, 412 GB across the reference
library's 19,107 images, so a whole-library pass is a transfer under
FR-NC-6 rather than something that may start on its own.

MIN_CROP_EDGE replaces the MIN_DETECT_EDGE this branch briefly had. Same
number, guarding the quantity that turned out to matter.

faces.md §7b records both measurements, and marks the second as
unexplained rather than dressing it as a finding. Grouped by the buffer
detection ran against, faces per image was 0.078 at 1024 or below and
1.82 at 2048 or better, controlled for file type and size. That gap is
real and reproducible and I cannot account for it, because the letterbox
above says detector input should not matter. M4 is where it gets
settled. The crop measurement does not depend on it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 19:41:20 +02:00
dtourolleandClaude Opus 5 d0ebc9f571 Count the same images in the progress figure that the sweeps index
The Identity screen said 4,593 images were left to index and stayed
there for hours across repeated runs, which is what a stuck job looks
like. It was not stuck. 4,424 of those 4,593 are shadowed -- the JPEG
half of a RAW+JPEG pair -- and no sweep will ever index one, because
every work list is built on VISIBLE, which excludes them. They are not
separate photographs and the grid does not show them either.

But faces::coverage counted them: its denominator was "images WHERE
trashed_at IS NULL", with no shadowed_by clause. So the outstanding
figure had a floor of 4,424 that no amount of work could bring down, and
Coverage::is_complete could never once return true no matter how
completely the library had been indexed. A progress number that cannot
reach its own target is worse than no progress number.

The fix is to count the population the sweeps actually draw from, in all
three places that were describing it differently: coverage's denominator
and its indexed join, and audit's split of the outstanding set, which
had the same gap and fed the same status line.

On the reference library the denominator goes from 23,531 to 19,107 and
outstanding from 4,593 to 169 -- the second of which is a number the
user can watch go down, and which turns out to be a real and separate
fetch failure worth chasing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 19:41:04 +02:00
dtourolleandClaude Opus 5 0a6509de93 Forget the face runs made on proxies too small to see a face
The floor added in the previous commit stops this happening again; it
does nothing about the 1,824 images in the reference library that
already carry a face_index row written against a proxy of 1024 or less.
Those rows are why the damage is permanent rather than merely past. The
work list is "images with no row for this model", so an image examined
against a 1024px proxy -- 0.078 faces per image, nine in ten finding
nothing -- is indistinguishable from one examined properly, and no
later pass will ever offer it to the detector again.

V12 deletes exactly those markers, and nothing else. The faces those
runs did find stay in place and keep drawing the People screen until a
better pass replaces them, and record_detections re-attaches the user's
confirmed names across that replacement by box overlap, so a library
somebody has spent an evening naming does not lose that evening. The
cost is a re-fetch of the affected images.

Deleting the marker rather than teaching the work-list query to select
on source_edge, which was the other option and is worse. A standing
`source_edge < floor` predicate never lets go: an image whose largest
embedded preview is genuinely smaller than the floor would be re-fetched
on every sweep for ever, because the next pass cannot do any better than
the last one did. A one-off deletion gives each affected image exactly
one more attempt through the good path and then lets the ordinary
"has a row" rule settle it.

The threshold is written out in the SQL instead of referring to
dr_face::MIN_DETECT_EDGE. A migration has to keep meaning what it meant
when it ran; binding it to a constant someone may raise later would
quietly change what an old catalog gets migrated to.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 19:40:26 +02:00
dtourolle 19b56ad85c Merge: collection ordering, and a range that says where it ends
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

# Conflicts:
#	docs/traceability.md
#	ui/dr-ui/src/collections_ui.rs
#	ui/dr-ui/src/library.rs
#	ui/dr-ui/ui/app.slint
#	ui/dr-ui/ui/library.slint
#	ui/dr-ui/ui/widgets.slint
2026-08-30 16:38:08 +02:00
dtourolleandClaude Opus 5 d70dcf78d1 Merge: recover a damaged catalog, and capture a crash locally
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 13:45:18 +02:00
dtourolleandClaude Opus 5 8eeb9ba0f6 Offer the backup, and then the rebuild, when the index turns out to be damaged
NFR-R6 asks for an integrity check at startup and two offers behind it, and
none of it existed. `PRAGMA integrity_check` appeared nowhere in the tree,
`Catalog::open` was `open` → `configure` → `migrate` → `backfill` and nothing
else, and corruption therefore surfaced as whatever rusqlite error the first
unlucky query happened to produce — "database disk image is malformed"
attached to a thumbnail refresh, elided into a 34px banner, over an empty
grid saying "No images found · Check the library folder". Two messages that
disagreed, and no way forward but deleting catalog.sqlite by hand.

The property that makes the second offer real was already here and load-
bearing: the catalog is an index, not a source of truth, rebuildable from
sources plus sidecars (invariant §5.2.4, cited by schema.rs, trash.rs and
lib.rs). And sync.rs already knew how to take a coherent snapshot of a WAL
database. What was missing was the check, the type, and the conversation.

Four pieces:

**The type.** `CatalogError::Corrupt`, and — the part that makes it worth
having — a hand-written `From<rusqlite::Error>` that classifies rather than
wraps. `SQLITE_CORRUPT` and `SQLITE_NOTADB` become `Corrupt` wherever they
arise, so a background job that trips over the damage first reports the same
thing the startup check would have. `SQLITE_IOERR` and `SQLITE_BUSY`
deliberately do not: a dropped network mount is a different problem, and
telling someone to rebuild their index would be a wrong answer delivered
confidently.

**The check.** `Catalog::open_verified`, `quick_check` before the open rather
than after, because opening runs migrations and a damaged catalog with an
intact header would otherwise have structure rewritten on top of structure
that is already wrong. Bound to `open_verified` and not to `open`: the check
reads every page, which is affordable once at startup where a user can answer
a question, and not affordable on the dozens of opens a session's background
tasks make.

**The backup.** NFR-R2's second clause, taken between `configure` and
`migrate` in `Catalog::open`. A migration is the one routine operation that
rewrites table structure, so it is the likeliest way this file becomes
unreadable, and it is the last moment the pre-migration state exists to be
copied. Three generations, through SQLite's backup API after a TRUNCATE
checkpoint — never `fs::copy`, which on a WAL database backs up a state older
than the catalog and possibly torn. A failure to take the copy is logged, not
raised: a full disk must not be what makes a library unopenable.

**The conversation.** The first line of the dialogue is that the photographs
and the edits are safe, before the diagnosis, because that is the question the
user is actually asking. Then the two offers, which are *not* interchangeable
and are not presented as if they were: a restore keeps collections, and a
rebuild cannot, because a manual collection is a set of images assembled by
hand and nothing in the filesystem records it (docs/catalog.md §8.1). The
labels say so, and the rebuild does not take the affirmative styling while a
restore is on the table.

One thing that is a fix rather than a feature: `show_catalog_now` now gates
the scan. `Catalog::open` succeeds on a file whose header survived, so the
scan that used to start immediately afterwards would write folder ETags and
image rows into damaged pages in the seconds while the user was still reading
the question — turning a file that had a backup into one where the backup is
the only copy left.

Restore also deletes the damaged catalog's `-wal` and `-shm`. That step is
easy to leave out and fatal to leave out: a journal belonging to the old file,
sitting beside the new one under the same name, is replayed into it on the
next open. That is not a restore, it is a fresh corruption with the evidence
gone.

Tested by corrupting a fixture catalog — 500 images and a collection, then
every page past the second overwritten — and driving both branches. The
restore is asserted on the collection, because a collection is precisely what
distinguishes the two paths; the rebuild on the damaged file being kept and
the next open producing an empty catalog at the current schema. Plus the
`SQLITE_NOTADB` presentation, a damaged backup being refused rather than
installed, and a v1 catalog whose pre-migration backup comes back reading
v1 rather than v11.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 10:34:12 +02:00
dtourolleandClaude Opus 5 421a47f1eb Untag FR-CAT-13, which no XMP is read or written to satisfy
FR-CAT-13 asks for standard XMP sidecars read and written — ratings, colour
labels, keywords and hierarchical subjects, title, description, copyright, GPS,
in `xmp:`/`dc:`/`lr:` schemas — so that other tools interoperate. Its one tag
was the module header of `dr-catalog/src/keywords.rs`.

That module stores keywords in SQLite. It names `dc:subject` twice, both times
in prose explaining why a keyword's text is the fact rather than its row id,
which is a good reason to have written it that way and not evidence of an XMP
implementation. Nothing in the tree parses or emits XMP: `dr-export`'s metadata
module writes EXIF and says in its own header that IPTC and XMP are named by
FR-EXP-8 and neither is read.

`dr-preset-xmp` is the crate whose name most invites the mistake. It reads
Lightroom `.xmp` *presets* — develop settings — under FR-DEV-6, and knows
nothing about the metadata schemas FR-CAT-13 is about.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 10:18:10 +02:00
dtourolleandClaude Opus 5 758436cc28 Keep the runner's borrow alive as long as the connection it reads
The first compile this branch had. One borrow error, in the four-thread
contention test: the `Runner` was the block's tail expression, and a
tail's temporaries are dropped after the block's locals, so it outlived
the `conn` it borrowed. Bound to a local, with the ordering rule written
down beside it -- it is exactly the shape someone tidies back.

Everything else stood: clippy clean at -D warnings, and all 18 runner
tests pass, including the four-thread four-connection claim and the
`UPDATE ... RETURNING` rewrite the author flagged as the riskiest line
in the diff.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 23:48:15 +02:00
dtourolleandClaude Opus 5 846a249156 Drain the queue that nothing has ever drained
`jobs` has been a complete durable work queue since the catalog was
written, and nothing has ever taken a job out of it. `claim_next`,
`complete`, `fail` and `recover_orphaned` had no callers outside their own
tests; `enqueue` had three. So the table grew one row per photograph and
kept it forever, and FR-PLAT-AND-3's resumability was a property of code
that never ran.

`runner` is the missing half. It owns no thread, no clock and no policy,
and that is the whole design: on Android the process does not decide when
background work may run. WorkManager does, subject to Doze, battery saver
and FR-NC-6's network constraints, and it revokes permission mid-job by
calling onStopped(). So the runner exposes `run_one` — claim, run, record —
and `drain`, which repeats it against a budget, a deadline and a
cancellation flag the host owns. A `Worker.doWork()` with ten minutes calls
drain with a deadline; a desktop idle pass calls it with none. That is the
seam the Android service plugs into, and it needs no Android to test.

Handlers are supplied from above, because the catalog knows what needs
doing and nothing about how: a thumbnail needs a decoder and a fetch needs
a network stack, neither of which belongs under core/dr-catalog. A runner
claims only kinds some handler declares, so a queue holding work this
device cannot do is left alone rather than failed five times.

Four outcomes, and only two of them are the job's fault. Done deletes the
row; Retry backs off; Abandon gives up now, for a failure no retry can fix;
Interrupted releases the claim with its attempt refunded and ends the
drain, because the host stopped rather than the job — five backgroundings
in a row must not mark good work as failed. Process death is the fifth and
cannot report itself, which is what `recover` is for.

Recovery is called from `show_catalog_now`, which is the one place a
catalog is opened for a session and already returns early if one is open.
It has to be exactly once and before any worker starts: there is no owner
column, so a second pass while a worker held a claim would take it away.
The attempt a dead claim consumed is deliberately kept — a job that takes
the process down with it is indistinguishable from one that fails, and the
attempt counter is the only evidence that survives a death.

The tests cover claiming under contention twice over: sequentially across
two connections, and with four threads on four connections against one
catalog on disk, asserting every job ran exactly once. Plus completion,
backoff, giving up, abandoning, interruption, budget, deadline,
cancellation, and a job orphaned by a simulated crash being reclaimed and
run once rather than lost or repeated.

Not wired to a handler yet, and deliberately not: the only enqueue site
the app actually reaches is the remote scan's, whose thumbnails are already
served by the async grid worker, and `walk`'s two sites are reachable only
from the scan_local example. Inventing a handler to make the plumbing look
used is how a requirement comes to read as covered by code that does not
implement it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 23:31:47 +02:00
dtourolleandClaude Opus 5 d63872e5a9 Make a claim one statement, and give the queue what a runner needs
The claim was a deferred transaction around a SELECT and an UPDATE, and
under a single connection that is fine. Under two it is not what it looks
like: the SELECT takes only a read lock, the UPDATE tries to upgrade, and
in WAL a worker that read the same snapshot as another gets
SQLITE_BUSY_SNAPSHOT on its write. That is not an error a busy handler can
retry away — the fix is to roll back and start over — so the queue was
"safe" only in the sense that the loser failed loudly instead of taking a
job someone else was holding.

`UPDATE jobs SET state = 1, attempts = attempts + 1 WHERE id = (SELECT ...)
RETURNING ...` is one statement and so one implicit transaction that takes
the write lock immediately. Two workers serialise, the loser waits out its
busy timeout, and neither can see a row the other already holds. The
existing tests are unchanged by it, because from one connection the two
forms are indistinguishable — which is exactly why it was never noticed.

The rest is the surface a runner has to have and did not:

- `claim_next_matching` takes only kinds a worker can actually do. Without
  it a device with no connector claims `FetchOriginal`, fails it, and pays
  five wakeups and five backoffs per photograph to reach a conclusion known
  before it started. Filtering after a claim cannot work: the claim has
  already marked the row running.
- `abandon` gives up now, for failures no retry can fix. `fail` uses it for
  its own MAX_ATTEMPTS branch, so there is one statement that ends a job.
- `release` hands a claim back with its attempt refunded, for a worker that
  is being stopped rather than a job that is going wrong. `attempts` stands
  in for the owner column the table does not have: it is bumped by every
  claim, so a stale worker's release matches nothing and changes nothing.
- `reap_orphan_subjects` deletes jobs whose photograph is gone. Coalescing
  keeps the table one row per unit of work and nothing ever shrank it when
  the work stopped existing. `ScanFolder` is excluded because its subject
  is a folder id, and joining that against `images` deletes by coincidence
  of numbering — hence `JobKind::subject_is_image`, and `JobKind::ALL` so
  the next kind added cannot quietly fall out of the filter.
- `counts` is the number a foreground service's notification is built from.

One behaviour change worth stating: a kind this build does not recognise is
now parked with an error rather than read as `ExtractMetadata`. The old
`unwrap_or` would have run a job of an unknown kind as some arbitrary known
one, which is worse than not running it at all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 23:31:27 +02:00
dtourolleandClaude Opus 5 b9891e2c04 Merge master into tablet-selection
Two real conflicts, both from work that landed either side of the same
lines rather than against them.

`lib.rs`: the settings controller was hoisted above the People screen's
wiring, and Android's thumbnail-tier eviction registered itself at the
same point. Independent, so both stay.

`library.rs`: manual collection ordering and burst folding each added a
clause to the same two queries. The scoped range read now carries both —
the folding matters there for one step further on than it does in the
grid, because a collapsed burst is one cell, so an ordinal counted over a
list still holding every frame names a photograph several places away
from the one the user pointed at.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 22:52:42 +02:00
dtourolleandClaude Opus 5 82d9d077b0 Merge: answer Android's memory warnings, and stop reporting a lost root as an empty library
FR-PLAT-AND-5 in full, FR-PLAT-AND-2 in part -- the recovery is built and
live for Nextcloud roots, the SAF cause it names does not exist yet.

FR-PLAT-AND-4 and FR-PLAT-AND-6 are not here, both blocked behind the
same gap: assemble-apk.sh compiles no Java, so the APK cannot carry a
Service or a FileProvider. The container has JDK 17 and build-tools 36;
the build step is what is missing.

Verified: fmt, clippy --workspace --all-targets -D warnings, and 1043
tests across dr-catalog, dr-sync, dr-sync-folder, dr-sync-nextcloud,
dr-plat and dr-ui. The aarch64 target was checked before the branch was
finished but not after; no device was available.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 22:19:46 +02:00
dtourolleandClaude Opus 5 75fd5619ca Refuse a scan whose root has gone, instead of reporting it empty
FR-PLAT-AND-2, and a silent failure on both platforms. `dr_sync::scan`
stepped over a NotFound or PermissionDenied the way it does for a child
that vanished mid-walk -- correct for a child, wrong for the root, where
it ended the walk, returned Ok with nothing in it, and reported a
successful scan of a library that was no longer there.

A lost root is now its own error. The images under it are marked
Availability::Offline per FR-CAT-9 and no catalog row is deleted;
`library::persist` clears the mark per file as each one is listed again,
so a root that comes back needs no repair step.

Partly satisfied rather than closed, and the gap is worth stating.
The recovery half is real and reachable on Android today, because
`map_status` turns Nextcloud's 403 and 404 into it and Nextcloud is how
a phone actually gets a library in this build. The causes the
requirement names -- revocation, reinstall, a removed card -- are
properties of a persisted tree permission, and there is none: SAF does
not exist here, `SourceRef::Document` is constructed only in test
modules, and `LocalStorage` rejects the variant outright. When SAF
lands it becomes a third producer of this error and nothing above it
changes, which is why the discovery belongs in the connector.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 22:19:38 +02:00
dtourolleandClaude Opus 5 fc4157a1e0 Keep the test scene's own arithmetic from overflowing a u64
The first compile this branch ever had. `cargo fmt` reflowed four files
and clippy passed at -D warnings untouched, but one test panicked:
`the_signature_does_not_change_with_scale`, on "attempt to multiply with
overflow".

It is the fixture, not the feature. `scene()`'s little LCG multiplied the
block's y by the golden-ratio constant with a plain `*` while the term
beside it already used `wrapping_mul`, so any scene taller than about 104
pixels overflowed in debug. Only the scale test builds one that large,
which is why 345 of 346 passed around it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 22:07:05 +02:00
dtourolleandClaude Opus 5 7ad75dd905 Let a manual collection be put in the order the photographer wants
`collections::set_order` and `Sort::CollectionPosition` have been in the
catalog since collections were, and nothing above dr-catalog has ever
called either. A manual order existed, could not be seen, and could not be
set. This is the half that was missing.

Three pieces, because it needed all three to be visible at all.

The catalog gains `orders_manually` and `members_in_order`. The first is
the rule about *when* a manual order means anything, kept in one place with
one name: a collection must be manual, and it must have no children. A set
shows its descendants' images, and positions are only ever assigned within
one collection — so two children's positions are unrelated integers, and
ordering by them would sort the grid by a coincidence. The existing comment
on `read_cells_scoped` already argued this; now something enforces it.

`members_in_order` returns the *whole* membership rather than the filtered
view, because `set_order` renumbers exactly what it is handed. Reordering a
filtered list would renumber those and leave every hidden image on a stale
position — two images sharing one, and a grid that rearranges itself the
moment the filter comes off.

The grid reads position where the scope qualifies and capture time
everywhere else. Manual order joins the member row rather than testing
membership with `IN`, which is safe from fanning out rows *because* that
branch is a single collection.

The gesture is a DropArea over the viewport, drawn only where a reorder
means something, with a caret in the gap the photographs would go into —
a line between two images rather than a highlight on one, because lighting
up a cell would say the drop replaces it.

The trap worth naming: `DropEvent.position` is in **window** coordinates.
Slint maps it through `map_to_window` when the drag begins and hands every
target the same event untranslated, so a target inside a Flickable has to
subtract its own `absolute-position`. Getting that wrong is invisible until
the grid is scrolled, because at the top the two frames coincide.

`reordered` is pure and names its destination by the image it goes before
rather than by an index, because the grid can only name a gap in what it is
showing and the ids are what survive a window swap. A drop that changes
nothing returns the order untouched: that counter is what a cross-device
merge resolves by, and spending a revision on a no-op makes this device win
an argument it did not have.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-29 20:49:00 +02:00
dtourolleandClaude Opus 5 d3dadbd725 Group the frames of one moment, by when they were taken and what they look like
A burst is the commonest thing in a cull and the least interesting: twelve
frames of the same gull at 10 fps occupy twelve cells, are scrolled past
twelve times, and end with the photographer keeping one. FR-CULL-5 asks for
them to collapse to one representative and be judged as a unit.

Two signals, because neither alone survives a real library. Time alone groups
a whole wedding ceremony -- a photographer working steadily never leaves the
gap that would end the run. Similarity alone groups a studio setup shot across
two days, which is a project rather than a moment. Together they are specific:
adjacent in time *and* looks like the frame before it.

Two seconds is the time bound, and the reason is worth recording because the
figure looks absurd next to a 10 fps camera. `images.captured_at` is whole
seconds -- EXIF's DateTimeOriginal has no sub-second field and
SubSecTimeOriginal is optional and widely omitted -- so a burst arrives in the
catalog as ten frames sharing one timestamp. Any threshold finer than a second
is a threshold on information that is not there. Where the pace really is
faster, the similarity bound is what separates the frames.

Similarity is a 64-bit difference hash over a 9x8 box-averaged reduction,
compared between *adjacent* frames only. Chained rather than anchored on the
first frame, because by frame twenty a camera following a bird has nothing in
common with frame one while no two neighbours differ by much; the time bound is
what stops the chain running away. There is no all-pairs step and there must
never be one -- that is what turns a grouping pass into something nobody can
afford to run over 50k images.

Nothing here ranks a frame. FR-CULL-5 names the failure it is avoiding, which
is rejecting the only frame of an important moment because somebody blinked, so
there is no sharpness score and no best-of-burst. The representative is the
earliest frame -- a fact about the clock, not a judgement about the photograph
-- and the user's own choice lives in its own table so that rebuilding the
grouping cannot erase it. Same argument `people.ignored` makes one subsystem
over: nothing short of remembering a decision survives re-clustering.

A newly found burst is recorded *open*. Collapsing on discovery would be
tidier, and would also mean a background pass taking photographs off the screen
part way through a cull. The pass marks; the user folds.

It is a pass rather than a job kind for the reason catalog.md 10.2 gives for
face clustering: a burst is a property of a run of frames and has no natural
subject_id, so a per-image job would rebuild the world once per photograph.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 20:37:54 +02:00
dtourolleandClaude Opus 5 574107bc39 Let a manual collection be put in the order it is meant to be seen in
`collection_members.position` and `Sort::CollectionPosition` have been in the
catalog since collections were, and nothing above dr-catalog has ever written
or read either: `collections::set_order` had no callers, and the grid ordered
everything by capture time whatever it was scoped to — dr-ui does not construct
a `Query` at all, it has its own `GRID_ORDER` constant. So a manual collection
was a set with an order nobody could see or change.

Three pieces, because it could not be fewer:

`grid_order_for` decides the ordering from the scope, and both readers take it
from there. That is the load-bearing part. An ordinal only names a photograph
relative to an ordering, so the window read and the span read have to agree —
a shift-click resolved through a different ORDER BY than the cells were drawn
with selects a different run than the one on screen, and the user finds out
when the export runs. `read_ids_span` already stated that invariant about
`GRID_ORDER`; this widens it to an ordering that depends on the scope.

Only a single manual collection has one. A set draws its descendants' images
too, and two children's positions are unrelated integers that interleave
arbitrarily; a smart collection has no member rows to carry a position at all.
Both fall back to capture time and refuse the drop rather than pretending.

The drop is on the cell, on whichever half of it the finger landed — the
trailing edge is the only way to name the last place in a collection, since
there is no cell beyond the last one to drop in front of.

`reordered` is pure and the membership is rewritten whole. `set_order` sets the
positions it is given and leaves the rest, so a partial write would interleave
the moved run with rows nobody touched; and it is read unfiltered, so what the
filter is hiding keeps its place relative to what the user can see.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 20:04:17 +02:00
dtourolleandClaude Opus 5 0407fb8d2d Format the two new examples
They were written after the last fmt run and CI gates on --check.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 12:33:23 +02:00
dtourolleandClaude Opus 5 da20d42d33 Merge master: pluggable storage, and a name that anchors
Conflicts were docs/traceability.md alone, and it is generated — so it
was regenerated rather than hand-merged. dr-face was untouched on the
other side; ui/dr-ui/src/faces.rs and identity_ui.rs auto-merged, the
first around recluster's anchoring and the second around load_faces.

Worth recording because the two branches met on the same problem from
different ends. Master's "Let a name hold a group together" is the fix
for the sixteen Catherines — fourteen of them empty — that this branch
found while measuring the library and reported without fixing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 12:30:01 +02:00
dtourolleandClaude Opus 5 b2250cc460 Measure a regroup on the tablet, not just on the desktop
The GPU question needed a number nobody had: how a regroup divides on
the hardware whose CPU is weakest. dr-face carries no weights and
touches no display, and dr-catalog's example needs only a catalog file,
so both run under adb shell against a copy of a real library.

On the same 18,143 faces — desktop against the tablet — scan 0.96s /
2.61s, agglomerate 1.69s / 2.16s, score 0.26s / 0.40s. The scan is half
the pass on the tablet and under a third on the desktop, because twenty
cores of AVX2 pull ahead of NEON much further than the merge engine's
single-threaded hashing does. So a GPU GEMM is worth roughly 2× a
regroup on the tablet and 1.5× here, and it is the tablet that should
decide whether it is built.

The two architectures agree exactly: the same 1,531,969 evidence pairs,
the same 2,518 groups holding the same 16,246 faces, the same
reliability table. That is a better check on the NEON kernel than the
unit test can be.

Two instruments, both read-only: the example now prints its phases, and
dr-face gains scan_bench, which needs no library at all and so can
answer "how fast is this machine" on a device with nothing on it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 12:16:48 +02:00
dtourolleandClaude Opus 5 ebb7d3cf5c Score a suggestion against the people the user has named
The number beside a suggestion was the mean calibrated probability
between the face and the rest of its group, which measures the wrong
thing twice. It punishes coverage: a person with two hundred faces over
fifteen years is *meant* to have members a given photograph is
orthogonal to, so a correct suggestion onto a well-photographed person
scored low for being well photographed. And it never asked who else the
face might be — a face matching Anna at 0.95 and nobody else, and one
matching Anna at 0.95 and her sister at 0.93, came out identical, when
the second is the only one worth the user's attention.

dr_face::assign answers both, and multiplies them: the mean of the best
ten calibrated matches into the identity (the old mean, capped, which is
what stops coverage counting against it), times that identity's share of
the evidence against every *named* rival.

Only named people compete, and per person rather than per group. Both
halves of that had to be measured on a real 18,000-face library rather
than reasoned about. Normalising across every group made the number
useless — median suggestion 21%, four in five under half — because
clustering leaves one person spread over many groups, so a face competed
against itself; and keying rivals by group left Catherine competing with
Catherine, median 39%. Per named person: median 99.5%.

Rivals are gathered below the merge threshold, down to even odds: a
named person matching at 0.6 will never be merged into but is exactly
the competition to discount for. That would be a second similarity scan,
the expensive half of regrouping a library, so cluster_scored scans once
at the looser floor and hands the merge engine the subset at or above
the threshold — pair for pair what it would have scanned for itself,
held to that by a test.

Leave-one-out over that library's 2,702 confirmations across 54 named
people: 99.33% of faces placed on the right person against the old
mean's 99.15%, and the number shown for the right person moves from a
median of 90.4% to 99.3%. It errs low — 100% correct wherever it states
80% or more — which is the safe direction, and docs/faces.md §9.1 says
plainly that the low bands are not calibrated.

The example that measures it comes too: this is a claim about a
library's numbers, and nobody should have to take it on faith.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 10:44:01 +02:00
dtourolle 5768100816 Borrow the library to index it, and give it back
The passes that need every photograph's bytes — thumbnails, face
indexing — now borrow each one and release it at the end. On a
placeholder library that is the difference between peak disk being the
working set and being the whole library.

Including on cancellation, which was nearly missed: the face sweep
returns mid-loop when the user presses Stop, and without releasing there
the disk is spent and nothing is delivered for it.

`materialise` now answers whether *it* fetched the content. The pool used
to work that out by listing a file's parent directory — one listing per
file across a library — when the backend already had to `stat` it to
decide whether to ask. One syscall instead of a directory walk, and it
removes the bug class the tests found earlier: a file at the library root
has no `parent()`, so every one of them read as already-downloaded.

**Pinning is the retention control**, and it drives the model the catalog
already had rather than a second one. `tier_desired` is what the user
asked to keep hydrated, `pending_pins` is the resumable work list, and a
pinned collection is never dehydrated for the same reason it was never
evicted. It was in fact *broken* here before: `get` on a stub failed, and
the pin worker logged "one unreadable file must not abandon the whole
pin" and silently did nothing.

Pinned originals on such a library are recorded with `path = NULL`
(`Cache::record_in_place`) rather than copied under `originals/`. Two
reasons, and the second is the important one. A copy would hold every
pinned photograph twice, with the budget able to evict the half that was
not costing the disk. And `release` deletes the file a row names — so a
row that names none cannot delete anything, which puts the one
catastrophic operation out of reach by construction rather than by
remembering not to call it. Deleting a materialised file inside a synced
tree removes the photograph from the server and every other device.

Handing disk back is `spawn_dehydrate`, which asks the client.

Two gaps written down rather than papered over (docs/storage.md §7): a
hydrating pass cannot yet quote its cost, because a stub reports no size;
and the two sweeps hold separate pools, so a library indexed for both
fetches twice.
2026-08-29 09:57:53 +02:00
dtourolleandClaude Opus 5 f41b3f6e8e Answer "what would the other device end up with" without the other device
Build and test / Desktop (Linux) (push) Successful in 2h7m52s
Build and test / Layer separation (push) Successful in 59s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 4s
Traceability / Requirement traces (push) Successful in 37s
Build and test / Android (aarch64) (push) Successful in 22m22s
A tablet showed 200 of a person's 611 faces after syncing, and the obvious
suspects — that suggestions deliberately do not travel, that the cross-device
face match was too strict — were both wrong. Finding that out meant reading a
catalog on a release-signed Android build, which cannot be done.

So this stands the second device up locally: an empty catalog, given the images
a scan would have found, the shards adopted into it exactly as a sync does, and
the real catalog merged in as the remote. Then it counts, per person, against
what the source holds.

    person                     source     here
    Catherine                     611      611
    Me                            242      242
    Ian                           219      219

Which settled it: the merge carries everything, and the shortfall was transfer —
shards that never finished arriving. Worth keeping, because "did the sync lose
this or has it not got here yet" is a question that will come up again, and
guessing at it cost most of an evening.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 23:09:28 +02:00
dtourolleandClaude Opus 5 2c84aa1224 Write face shards the way the rest of the catalog writes
The face store was the one part of the catalog still on SQLite's default
rollback journal at `synchronous = FULL`. The catalog itself runs WAL at
`NORMAL` (`schema::configure`) and so does the thumbnail store; nothing decided
this one should differ, it was simply never set.

Measured on this project's own filesystem, that is **21.3 ms per commit against
0.05 ms** — four hundred times. And an export commits four times per
photograph: the shard's transaction, then three separate autocommitting writes
to the index. Ten thousand images is on the order of fourteen minutes spent
doing nothing but waiting for fsync, before a byte goes to the server. That is
the "checking faces…" that appeared to hang.

So: WAL and `synchronous = NORMAL`, matching the rest, and the three index
writes fold into one transaction. `NORMAL` is the same trade the catalog makes —
a shard is derived data, and losing the last commit to a power cut costs one
image re-exported.

WAL brings an obligation with it, because **a shard is uploaded by reading its
file**: the newest commits live in a `-wal` sidecar that no upload sends, so
without a checkpoint the server would receive a database missing exactly the
faces just written, and a peer would adopt it and see nothing wrong. `checkpoint`
folds the logs back in, with `TRUNCATE` rather than the default passive mode,
which gives up when a reader holds the log and would leave the same gap while
reporting success.

Two tests: that the store is in WAL like everything else, and — the one that
matters — that a checkpointed shard copied *without* its `-wal` still holds
every face. That second one fails without the checkpoint, which is how it was
confirmed to be testing something.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 23:08:10 +02:00
dtourolleandClaude Opus 5 a1e8f494b9 Say what the face sync is doing while it does it
The face pass set the status to "checking faces…" once and then said nothing
until it was finished. On a library whose first export after a re-index is 9,849
photographs and 85 MB of shards, that is eight minutes of a progress bar sitting
still — which is indistinguishable from a hang, and was reported as one twice.

Nothing was wrong with the sync. The only fault was that it was silent.

Three places now report, which are the three that take real time:

- **Preparing**, per image with a count, since this is the long one and the only
  one whose length the user cannot guess from anything on screen.
- **Sending**, per shard with its size, because a face shard carries crops and
  runs to tens of megabytes — one of them is a visible wait on any connection.
  Announced before the upload rather than after, since the wait *is* the upload.
- **Taking in** a peer's shard, which is a download and then a row-by-row merge.

The export reports every 25 images rather than every one, so the channel behind
it stays lost in the write it accompanies. `export_to_shards` keeps its old
signature and delegates, so the callers that do not want progress do not grow a
parameter for it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 21:56:16 +02:00
dtourolleandClaude Opus 5 0729dfa359 Stop the face sync opening a database per photograph
"Checking faces…" never finished. `export_to_shards` asks, for every indexed
image in the library, whether the shard store already holds that image at that
index time — and `indexed_at` answered by opening the shard database, running
its six-statement schema batch and two `pragma_table_info` queries, then
querying. Once per image. 9,849 times for this library, on every sync pass,
before a single face had been written.

The index time now lives in the store's `index.sqlite` alongside the shard
number, so the question is one indexed lookup on a connection that is already
open. It stays in the shard as well — that copy is the one that travels — but
nothing reads it from there on the hot path.

The write side had the same shape: `put_image` opened the shard afresh for each
image, which mattered little when exports were a handful of new photographs and
matters a great deal now that a re-index sends thousands. The handle is kept and
reused, invalidated by shard id so sealing a full one and moving to the next
drops it without anything having to remember to.

`INDEX_SCHEMA` is `CREATE ... IF NOT EXISTS` like the shard schema, so the new
column is added on open for an index already on disk — the same trap, caught the
same way.

Two tests: that the index time survives reopening the store, and that an index
written before the column can still be opened and written to.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 21:19:56 +02:00
dtourolleandClaude Opus 5 76eece8500 Add the columns an existing shard never got
`SHARD_SCHEMA` is entirely `CREATE ... IF NOT EXISTS`, which does exactly
nothing to a table that already exists. So `crop` and `indexed_at`, both added
to that batch, never appeared in any shard that had been written before — and
the `INSERT` naming them failed with "no such column".

Which took face export down completely, on every library that had ever synced a
face. Silently: `export_to_shards` returns the error, `sync_face_shards` logs it
at warn, and the sync goes on looking successful while the catalog fills with
faces no other device will ever see. This library's shard sat frozen at 1,807
faces with 15,194 in the catalog, and the reason was this rather than anything
in the export logic.

Shards are upgraded on open now: both columns are additive and nullable, so
catching up is one `ALTER` each. There is deliberately no version counter —
"does this column exist" is the question actually being asked, and asking it
directly cannot fall out of step the way a counter can.

A peer's shard is opened read-only and cannot be repaired, so one written before
crops is read as it stands, with a `NULL` standing in for the column. An adopted
face simply has no crop, which is the truth about it.

Four tests, built against the pre-crop schema written out in full rather than
derived from the current one — the point being that it is *not* the current
schema and must not track it. Verified against the real 1,807-face shard on this
machine: the ALTERs apply, writes succeed, and nothing already in it is lost.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 09:50:55 +02:00
dtourolleandClaude Opus 5 4ed10f7b23 Let a person cross from one device to another
The face shards carry boxes, landmarks and embeddings. What they deliberately
do not carry is who anybody **is** — the person rows, their names, and the
assignments joining the two. Those travel in the catalog snapshot, which is a
whole-file copy and does contain them.

But the snapshot is *merged*, not adopted, and this merge only ever looked at
collections and keywords. `face_shard`'s own module note says people travel in
the snapshot; nothing implemented it. So a second device received every face and
no people at all, and drew an empty People screen over a full catalog. Exactly
what a tablet showed after syncing thousands of faces from a laptop.

What travels is what the user decided, following the rule the rest of this
module already follows — judgements travel, inference is rebuilt:

- **People**, by uuid on `revision`, exactly as a collection is: the name, and
  whether the group was set aside.
- **Confirmations**, and **rejections** — "this is not her" is a fact too, and
  is why re-clustering does not put it back.
- **The suggestions inside an ignored group**, which are otherwise ordinary
  inference but are what anchors the ignore. Without them a group set aside on
  one device reappears on the other, the same fault that made "Not interested"
  not stick locally.

Ordinary suggestions are not carried. Both devices hold the same embeddings and
clustering is deterministic, so each recomputes them and arrives at the same
answer; shipping them would double the merge for no new information.

**A face has no cross-device identity**, and unlike a collection there is no
uuid to give it one. Both devices do agree on `oc:fileid` and roughly on the
box, so a remote face is matched to the local face on the same photograph whose
box overlaps it most, above 0.5 IoU. That is not a new rule — it is the one
`record_detections` already uses to carry a confirmation across a re-index, and
it is loose on purpose: the question is "the same face in the frame", not "the
same rectangle".

A local confirmation is never overwritten. Two devices confirming one face as
different people is a real disagreement and an assignment carries no revision to
settle it with; taking the remote's answer would let a sync undo what the user
just did on the device in their hands.

The remote's schema is probed rather than assumed: `remote_is_mergeable` admits
any catalog at or below this version, so one written before faces existed, or
before V10 added `ignored`, is ordinary. An absent table skips this half instead
of aborting a merge that would otherwise have succeeded.

Nine tests, including that the name lands on the overlapping face and not its
neighbour in the same frame, that a set-aside group stays set aside, that an
ordinary suggestion does not travel, and that merging twice changes nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 09:27:17 +02:00
dtourolleandClaude Opus 5 cf614efa61 Send a re-indexed image's faces to the other devices
`export_to_shards` asked `store.contains(file_id)` and skipped anything the
shard store had already heard of. So an image was exported exactly once, and
re-indexing it updated the catalog and nothing else — every other device kept
the first answer for ever.

That is not hypothetical. This library was re-indexed after the detection floors
changed and crops were added, going from 1,807 faces to 15,194; the shard store
still held the original 1,807, written before any of it. Nothing the re-index
produced could reach another device.

The shard's `indexed` table now carries the catalog's own `indexed_at`, and the
export compares against it. A re-indexed image goes again; an unchanged one
still costs nothing. Copied from the catalog rather than stamped when the shard
is written, because a shard-local write time advances even when nothing changed
and could not answer the question.

The column is nullable so a shard written before it still reads: absent means
"cannot vouch for it", which forces one re-export and then settles.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 09:27:16 +02:00
dtourolleandClaude Opus 5 79c0520506 Keep the face, not just a way to find it again
A face was drawn by decoding the 1024px proxy it was found on and cutting the
box out again, every time the People screen opened. That made the screen a
derivative of the thumbnail cache: evict a proxy — which the cache may do at any
moment — and the cell goes blank, with no way back short of re-fetching the
original over the network and re-detecting it. It also cost a full JPEG decode
per image, per visit, to show a 96px cell.

So the crop is cut once, when the pixels are already in hand at detection time,
and kept. A 160px JPEG is a few KB against the ~250 KB proxy it replaces reading.

Where it lives is the interesting part. The catalog snapshot is uploaded *whole*
on every sync and downloaded by every device, so a crop column there would put
tens of MB on every round trip — the exact cost `face_shard`'s 25 MB cap exists
to bound, and the reason bulk per-face data lives in shards already. Crops
therefore travel in the face shards, beside the embeddings, and
`snapshot_for_upload` strips them from the copy it writes. Nothing reads a crop
out of a merged remote catalog — the merge touches collections and keywords only
— so a receiving device loses nothing. A shard carrying crops holds around 3,500
faces rather than 22,000, which is the price of a second device showing People
immediately instead of re-fetching every proxy.

The column is nullable and the reader falls back to the proxy, so a face indexed
before this still works and the next indexing pass fills it in.

V10 also adds `people.ignored`, for a person the user has looked at and does not
want to identify. Most clusters in a real library are strangers — passers-by,
other people's guests, a face on a poster — and there is no way to tell "not yet
looked at" from "looked at, don't care" without recording the second. It is a
column rather than a deletion because a deleted cluster comes straight back on
the next Regroup: the faces are still there and still similar, and nothing short
of remembering the judgement survives re-clustering. Same argument
`face_person_rejected` makes one level down.

And `prune_empty_unnamed`, for what clustering leaves behind. Regroup creates a
person per unanchored group and never removed the previous run's now-empty ones,
so pressing it twice added a rail entry per group it no longer believed in.
Named people are never touched however empty — a name is user data — nor is a
merge tombstone, which must outlive its faces to keep redirecting.

298 tests pass, including that the snapshot carries no crops while the live
catalog keeps them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 21:17:44 +02:00
dtourolleandClaude Opus 5 b846b312b8 Run the formatter over the face branch before it reaches CI
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 2s
Build and test / Desktop (Linux) (push) Successful in 1h21m32s
Build and test / Layer separation (push) Successful in 37s
Traceability / Requirement traces (push) Successful in 25s
Build and test / Android (aarch64) (push) Failing after 33m58s
The merge of the SCRFD/MobileFaceNet work brought 69 rustfmt diffs across
dr-catalog, dr-face and dr-ui with it, so `cargo fmt --all -- --check` fails
on master and the Desktop job stops at its Format step — before clippy, the
tests or the release build have run at all. That makes the whole desktop
half of CI blind: a real compile error behind this would look exactly the
same from the outside. There was nothing behind it, as it turns out — with
the formatting fixed, clippy, the test suite and the release build all pass.

Every .rs hunk is `cargo fmt --all` on the pinned 1.92.0 toolchain, not a
hand edit, but it is worth being precise about what that moved, because it
is more than whitespace. Besides reflowing signatures and call chains,
rustfmt reordered the `pub mod` and `pub use` items in dr-face/src/lib.rs so
the `#[cfg(feature = "inference")]` entries sort in place, added the trailing
semicolon inside `let ... else { return }` bodies in identity_ui.rs, wrapped
a bare closure body in braces in cluster.rs, adjusted trailing commas, and
dropped a stray blank line at the end of identity_ui.rs. All of it is
semantically inert; none of it changes behaviour.

docs/traceability.md rides along because it has to. The matrix records each
TRACES tag by line number, and reflowing develop.rs, lib.rs, faces.rs,
identity.rs and identity_ui.rs moved them — FR-CAT-8, FR-CAT-9, FR-CULL-10,
FR-DEV-3, FR-DEV-3a and FR-DEV-3c all shift by a line or two. The matrix was
verified up to date on d777f7f before this commit, so this is drift these
formatting changes introduced, not pre-existing staleness being swept up.
Leaving it for a follow-up commit would hand traceability-check.yml a
failure caused entirely by a whitespace change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 13:12:49 +02:00
dtourolleandClaude Opus 5 96d07da15f Sync face data as sealed shards, so a second device does not re-index
Indexing 23,500 images is about two hours of CPU, and the result is
byte-identical on every device: the same model over the same proxy
produces the same embedding. Paying for it once per account rather than
once per device is the point.

Shards rather than the catalog snapshot, because the snapshot goes up
whole on every sync and a fully indexed library carries roughly 30 MB of
embeddings. That is exactly the cost the thumbnail store's 25 MB cap
exists to bound, so face shards use the same cap -- imported from
dr_thumbs rather than restated, since the number is a statement about
sync cost and the two must not drift apart.

The split follows the one already there: bulk immutable data in sealed
shards, small mutable data in the catalog snapshot. Faces, landmarks,
embeddings and run markers shard; people, names and assignments ride the
catalog and merge by uuid.

Keyed on oc:fileid throughout, never on image_id, because a row id means
nothing on another device.

The run marker travels with the faces it describes. Without it a
receiving device cannot tell an image with no faces from one never
examined, and would re-detect every landscape it had just adopted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 22:41:50 +02:00
dtourolleandClaude Opus 5 26a1eb7e28 Record that face detection has run, not just what it found
An image with no faces in it was indistinguishable from one that had
never been looked at, so every landscape, still life and document scan in
the library was re-detected on every pass, for ever. In a real library
that is most of it: on the 23,527-image test library, 64 of the first 110
images indexed contain no face at all.

Schema v9 adds face_index, a run marker per (image, model) carrying the
face count and the proxy edge it read. Keyed on the model, so a model
change puts every image back in the queue by itself.

That makes a coverage figure possible, which is the thing a user actually
wants to see. The audit also splits the outstanding set by whether a
proxy exists, because 23,417 awaiting a proxy and 110 ready to index are
different problems, and telling the user to run indexing again would not
fix the first.

The Identity screen gains Index faces, Stop, and the coverage line.
examples/face_index.rs is the same check and sweep without a window,
which is the right shape for an overnight pass.

Measured on the real library in release: 3.5 images/second, 110 images
and 125 faces in 30 seconds, and a second run correctly finds nothing
left to do.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 22:13:41 +02:00
dtourolleandClaude Opus 5 2ac069a6b3 Index the library's faces, and group them into people
Wires dr-face to dr-catalog: a background sweep that reads the proxy the
grid already built, detects, aligns, embeds and stores, then a clustering
pass that turns those embeddings into suggested people.

Detection runs on the Large thumbnail tier and nowhere else. That is what
makes the feature affordable -- a browsed library has already paid for
its proxies, so face indexing adds no RAW decode that was not already
happening -- and it is why an image whose proxy is missing is skipped
rather than fetched: requesting one here would put face indexing on the
network path FR-CULL-8 keeps it off.

The sweep keeps no cursor. It asks the catalog what is missing, so it
resumes after process death with no repeated work beyond the in-flight
image, and cancelling is dropping the receiver.

recluster writes only the suggested half. Confirmed faces go in as
anchors and come back untouched, and a cluster of one stays nameless --
naming every stray face would fill the People view with noise the user
then has to dismiss.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 20:49:30 +02:00
dtourolleandClaude Opus 5 00e78dc2ac Cluster faces into people, and calibrate what a similarity means
FR-CULL-9 forbids thresholding a bare cosine anywhere in the subsystem,
so calibrate fits P(same person) per library and reports whether the fit
is trustworthy. Two details carry most of the weight.

The fit runs against a 200-bin histogram rather than a pair list: a
25,000-face library has ~3e8 pairs and no gradient descent is running
over that. And a fresh library has no valid calibration, because the
positives have to come from user confirmations or burst siblings --
bootstrapping them from high cosine would fit the calibration to the
belief it was supposed to test.

Clustering defends against the over-merging FR-CULL-10 warns about with
constraints rather than a better threshold: two faces in one photograph
never merge, and two groups confirmed as different people never merge.
Average link rather than single link, so one strong edge cannot weld two
families together.

Calibration is defined once, in dr-face, and dr-catalog re-exports it.
Two implementations of one probability model is exactly how a number
comes to mean the wrong thing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 20:10:33 +02:00
dtourolleandClaude Opus 5 aac3136407 Store faces and the people they belong to
Schema v8: people, faces, face_person, face_person_rejected, and the
per-library calibration. Follows catalog.md 10.1 with two additions the
spec work turned up.

crop_px, because at the 1024px proxy tier a group shot reaches the
embedder at ~50 source pixels upsampled to 112 and a portrait at 340.
FR-CULL-9 names face size as an axis along which an uncalibrated
similarity misbehaves, so it is a stored feature rather than a UI hint.

face_person_rejected, because rejection is not the absence of an
assignment. Without it the next clustering pass re-suggests exactly the
face the user just pushed away, and the tool feels broken.

record_detections replaces rather than appends, since DetectFaces is
coalesced per image -- and carries confirmations across the replacement
by box overlap, so re-indexing with a better model cannot discard the
user's own labelling.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 20:02:52 +02:00
dtourolleandClaude Opus 5 6d6ef8d34b Page the grid along an index instead of sorting the library each time
Scrolling jittered, and this was the largest single reason. Every window the
grid loads is `ORDER BY ... LIMIT n OFFSET k`, and neither half of that was
being answered the cheap way.

**The sort.** `GRID_ORDER` leads with `captured_at IS NULL`, so undated frames
fall to the end. No ordinary index answers that — the leading term is an
expression, not a column — so SQLite sorted the whole library into a temp
b-tree on every window read, then threw away the first `k` rows of it. Schema
V7 indexes the expression exactly as the query writes it, partial on the same
`shadowed_by IS NULL AND trashed_at IS NULL` the grid filters by, so the read
becomes a walk along the index.

**The join.** `LEFT JOIN remote` was paged *after* it was joined, so reading
280 cells at offset 20,000 first seeked into `remote` for all 24,000 rows and
then discarded 23,720 of them. The file ids are now fetched for the 280 rows
that survived — the shape the badge and rating reads already use, one query for
the window rather than one per cell.

Measured together on 24,000 images at offset 20,000: **15.2 ms → 0.36 ms**,
inside a scroll handler that has 16.7 ms to draw a frame.

The test asserts on the query plan rather than on a duration, because there is
no other symptom. A `GRID_ORDER` edited out of step with the index, or a column
added back that drags `remote` in again, both still return exactly the right
cells — just after sorting the library — and the jitter would come back with
nothing to point at.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-25 20:01:04 +02:00
dtourolleandClaude Opus 5 fa4ad6e2d6 Drag the date range on the axis it is chosen from
The range could be turned on with a finger and not aimed with one. Its two
ends were typed as `YYYY-MM-DD` into 108px fields behind a soft keyboard, to
name days already drawn on the axis a thumb away; and the chip that seeds them
takes its span from the timeline's zoom and pan, which are a wheel and a middle
button. A touch screen has neither, so on Android the filter was a switch with
no aim.

The band is now on the timeline. Two ends with grips, dragged along the bars,
released to filter — the histogram was already how a period is found, and this
makes it how a period is stated. Both ends snap to whole days, which is what
the typed fields mean, what `show_range` reads back out, and a floor under a
range dragged shut. The fields stay for what dragging cannot do: name an exact
day, and say in words what the range is.

For that to work the axis had to stop following the range. Redrawn to the band,
it moved the ground under the very handles doing the narrowing, and there was
nothing outside the range left to widen back into.

While there: a fixed number of equal bins instead of calendar buckets. Between
one calendar unit and the next the bar count is free to wander by a factor of
twelve, so zooming in halved it two steps out of three — the same picture drawn
wider until it jumped back to fine. Equal bins also include the empty ones, so
a bar's position on the track and the date under it are finally the same
quantity; before, a library with gaps drew a February six months wide and the
marker, the band and a click all pointed somewhere else. The count is a
setting, 32 or 64, because the right answer is a question about the screen.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-23 18:42:42 +02:00
dtourolleandClaude Opus 5 4b7648082e Let a date range be stated, and draw the axis at the scale it deserves
"Limit to range" did nothing, and the reason was not visible from the button.
It took its span from the timeline's zoom, which is zero until someone zooms —
so `zoomed_span` returned the whole library and the filter narrowed to
everything. The chip lit up and the grid did not change.

The range has ends now, shown and typed as `YYYY-MM-DD`. Seeding them from the
timeline is kept, because zooming to a fortnight and pressing the chip is the
fast path; the fields say which fortnight it landed on and let it be corrected.
Ends given backwards are swapped rather than refused — there is exactly one
range between two days — and the closing day is included, since "to the 5th"
means the whole of the 5th and a range ending at its midnight contains none of
it. `parse_date` refuses anything that is not a date rather than guessing at an
order, because the alternative is a library silently filtered to a span nobody
asked for.

The axis then follows the range. It used to keep drawing the full extent while
a range was on, because it was the only way back out; the typed ends are the
way back out now, so it is free to show what was asked about.

And bucket size is chosen by how many bars it makes rather than by fixed
cut-offs. Each zoom step halves the span, so under thresholds the bar count
halved with it until a boundary was crossed: fifteen years went 15 bars, 8,
then 46, 23, 11, and finally 6. Zooming in made the picture coarser, which is
the opposite of what zooming is for. Aiming at forty bars keeps the count in
the same neighbourhood at every level, and the test asserts the property
directly — halving a span never coarsens the bucket.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-23 12:40:14 +02:00
dtourolleandClaude Opus 5 67f15bffb7 Key collection membership on the identity that exists
Collections synced their names and arrived empty on every device. The names are
keyed on a uuid and worked; the membership union was keyed on
`images.content_hash`, and the schema says plainly what that column is:
"computed only when something needs it (import dedup, reconnect-by-hash), never
in a scan". A library that has only ever been scanned has one for no image at
all, so the join matched nothing and `WHERE ri.content_hash IS NOT NULL`
discarded whatever survived. The union could never have moved a single row.

Measured on a real catalog: 23,174 images, content hashes for 0 of them,
`oc:fileid` for all 23,174, twelve collections, zero members.

So membership now resolves through the file id first, exactly as keyword
assignment already did — `ASSIGN_BY_FILE_ID` was added for this same reason and
its doc comment even notes that membership was still on the hash. It is
recorded for every image the moment a remote scan sees it, survives server-side
rename and move (FR-NC-5), is the same integer on every device pointed at one
Nextcloud, and is already what the thumbnail shards are keyed by.

The content-hash union is kept rather than replaced: a local-only library has no
`remote` rows, and where a hash has been computed it is a true identity that
survives a library moving between servers. Both statements run; `INSERT OR
IGNORE` against the primary key makes the overlap free.

This repairs the merge. It cannot invent membership that no device recorded —
where the rows were never written, collections stay empty until they are filled
in again.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 21:19:13 +02:00
dtourolle 56dd3187f1 Merge integration into wip/ingest
Second pass, against the detail-stage and thumbnail work that has landed since
the first. Resolved and verified here rather than in the shared merge worktree,
so what goes back is a fast-forward.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

# Conflicts:
#	docs/traceability.md
2026-08-22 19:44:42 +02:00
dtourolleandClaude Opus 5 62188ec740 Keep both devices' keywords when the catalogs meet
Keywords are catalog state, and the catalog syncs. Without this, two devices
keywording the same library would resolve to whichever synced last, and an
afternoon of work would vanish with no sign it had ever happened.

The vocabulary merges per row on the rule collections already use: revision
first, timestamp only to break a tie, so a device with a skewed clock cannot win
by having the wrong idea of the time. Assignments merge as a set union, which is
FR-NC-9's principle applied to metadata instead of edit nodes — disjoint work
survives on both sides.

Three things needed care and are commented where they happen:

A deletion travels *by name*, not by identity. Both devices may have minted
their own uuid for one word before they ever synced, so deleting by uuid would
tombstone a row nothing was assigned to and leave every photograph still
carrying the word. The union then refuses to readmit a word a winning tombstone
has just removed — without that filter the remote's live assignments would
resurrect it on the very same pass.

Images are resolved by the server's file id first and the content hash second.
Membership has always used the hash alone, but the hash is computed only when
import dedup or a reconnect asks for it, which for most libraries is never — so
a hash-only union would have quietly done nothing for the ordinary photograph.

A word lands on the local default version. Version uuids do not reconcile in the
catalog at all: ensure_default_versions mints a fresh one per device, so a
uuid-keyed join would have unioned nothing.

Removal still does not propagate. That is the trade collection membership
already makes, for the same reason — an unwanted keyword is removed again in a
second, and a silently lost afternoon is not recoverable at all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 15:51:00 +02:00
dtourolleandClaude Opus 5 2147eaa6a5 Put a keyword on a photograph, not only search for one
The catalog has been able to *find* by keyword since v1 — query.rs joins the
keywords table, matches it exactly, and substring-matches it for free text — and
nothing anywhere could ever put a word there. A user could filter to a keyword
they had no way to apply.

This is the missing half: create, rename, delete, list, assign, unassign, and
the two reads a panel needs. Bulk-only for assignment, because keywording a
selection is the common case rather than the exception — the photographer picks
out the frames with the puffin in them and applies "puffin" once, in one
transaction.

Schema v6 adds `keyword_terms`, and deliberately does *not* touch the v1 join.
The assignment keeps the word as text because the catalog is a rebuildable index
and the durable copies of that fact — the sidecar, XMP dc:subject — both carry a
string; a foreign key would mean a catalog rebuilt from sidecars had to invent
identity rows before it could record anything, and would break the query path
that already works. So the text is the fact, and the new table is only the
identity a rename and a deletion can be keyed on.

`keyword_terms.name` carries no unique index, which looks like an oversight and
is not: two devices that each type "Iceland" are both right until they meet, and
a constraint would abort the merge at that moment. Uniqueness is converged upon
instead — create resolves an existing name, fuse_duplicates collapses a
cross-device pair onto the smaller uuid.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 15:50:49 +02:00