rustfmt over the files the albums work touched, and the album merge's
incoming row as a named struct rather than an eight-field tuple, which
clippy's type_complexity refused.
An album is a named export destination. Its folder holds only the
exported files; the catalog records, per file, the image it was
rendered from, so an album can show the originals behind its JPEGs
(FR-EXP-10).
The tables are created on first use (CREATE TABLE IF NOT EXISTS), the
way dedup_probes is, rather than by a schema migration: a new
user_version makes every older build refuse this catalog's snapshot at
sync, and the 0.16.0 tablet would stop merging collections, keywords
and people for a feature it does not have.
Albums merge as collections do: by uuid and revision, tombstones on
delete, exports as a set union keyed on the server's file id (content
hash for a folder library). A folder on the server lives on the album
row and syncs; a folder on this device lives in album_folders, which
the merge never reads and the upload snapshot drops, because a path or
a SAF grant on one device means nothing on another.
Exports are keyed on the file name, not the image: two crops of one
photograph are two files and two rows, and an overwrite re-points the
name at whatever wrote it last.
After 9cff677 the loop over the other device's confirmed and ignored faces
(13,000 on the reference library) still asked three cached statements per
face -- the person by uuid, the face's current assignment, and whether this
pair was rejected here. It was 51 ms of a steady-state merge.
The person is now resolved in the statement that reads the incoming rows,
by the local `people.uuid` key:
SCAN fp
SEARCH p USING INTEGER PRIMARY KEY (rowid=?)
SEARCH lp USING COVERING INDEX sqlite_autoindex_people_1 (uuid=?)
and the local `face_person` (16,800 rows) and `face_person_rejected` are
each read once into memory and looked up there. A write goes to the table
and to the map, so a second remote face matched to the same local face sees
what the first left, as it did when each face re-read the table. The
incoming rows are ordered by face id -- the order the table was already
walked in -- since which of two such faces is applied last decides the
answer. An inner join to `people` drops the rows the old loop skipped for
want of a local person, and the counts in the report are unchanged.
After: the loop 10-12 ms. The merge as a whole, with the two changes before
this, went from 228-231 ms to 135 ms best of 5, and every catalog table
checksums the same after the bench as after the old build's run.
`merge::match_faces` reads every local face's box and model to pair the
other device's faces with ours. It took 54 ms of a steady-state merge on the
reference library (19,000 faces).
A `faces` row is eight kilobytes -- the embedding, the crop, the dense
landmarks -- and `model_id` sits past the embedding, so reading it opened
each row's overflow pages:
SCAN f
SEARCH r USING INTEGER PRIMARY KEY (rowid=?)
`faces_box (image_id, model_id, x, y, w, h)` holds every column the scan
asks for:
SCAN f USING COVERING INDEX faces_box
SEARCH r USING INTEGER PRIMARY KEY (rowid=?)
The local scan went from 38 ms to 8 ms (sqlite3 on a copy, aggregated so
output formatting is not timed), and `match_faces` from 54 ms to 30-37 ms;
what remains is the other device's half. That is read from its snapshot,
which has whatever indexes its build made -- this one will carry
`faces_box` in its uploads -- and whose rows have had their crops stripped.
The bench merges a full copy with crops, so it overstates that half.
Created on first use in `match_faces`, with CREATE INDEX IF NOT EXISTS,
rather than by a migration, for the reason `keywords::ensure_term_index`
gives: a schema version bump makes older builds refuse the snapshot, and an
extra index is invisible to them. The first merge after the upgrade builds
it (about a second, once). Its prefix duplicates `faces_image_model`, which
is left alone; the planner takes either for an (image_id, model_id) probe.
Tables checksum the same after the bench run as after the old build's.
`merge_remote_catalog` on the reference library (catalog_bench, a copy
merged with itself: the steady state of a sync pass) cost 228-231 ms best
of 5. Timing its phases put 91 ms in the keyword half, not in the faces the
issue named.
Both assignment unions refuse a word this device holds only as a tombstone,
with a correlated `NOT EXISTS (... deleted = 1) OR EXISTS (... deleted = 0)`
per incoming assignment. The `deleted = 1` half has no index to use --
`keyword_terms_name` is partial on `deleted = 0` -- so it scanned the whole
vocabulary for each of the 10,800 rows:
SCAN rk
CORRELATED SCALAR SUBQUERY 1
SCAN t
CORRELATED SCALAR SUBQUERY 2
SEARCH t USING COVERING INDEX keyword_terms_name (name=?)
The refused words are one set for the whole statement, so it is asked once:
`rk.keyword NOT IN (tombstoned names EXCEPT live names)`, which is the same
condition -- refused exactly when deleted under some identity and live under
none -- and which SQLite builds as a list before the walk:
SCAN rk
LIST SUBQUERY 2
MERGE (EXCEPT) ...
The file-id union alone went from 72 ms to 11 ms (sqlite3 on a copy), and
the keyword phase of the merge from 91 ms to 28-35 ms. Every table of the
catalog checksums the same after the bench as after the old build's run,
and the merge tests for tombstones and renames pass unchanged.
A sync pass that brought nothing new cost 450-540 ms of CPU in
`merge_remote_catalog` on the reference library (24k images, 19k faces),
measured by catalog_bench merging a copy of the catalog with itself.
Most of it was the loop over the other device's confirmed faces and the
faces under its ignored groups -- 13,000 rows. For each one it prepared
three statements from scratch (`query_row`/`execute` with a SQL string
compile the statement every call) and then rewrote the `face_person` row
with the values it already held, dirtying a page per face on every pass.
The rejection loop prepared three more per row.
The statements are now `prepare_cached`, the local assignment is read once
per face (whether it is confirmed, and what it holds, come from the same
row), and the upsert is skipped when the row already says exactly that.
`faces_assigned` is still counted for those rows, so the report is the one
the old code gave, and nothing else reads the difference: the row is
byte-for-byte what the upsert would have written.
After: 279 ms (best of 5, CPU), with every catalog table identical after
the run to the old build's.
docs/ had 26 developer documents flat beside the manual, and the two
audiences are very differently sized: most readers want the manual and
the gesture reference, a few want the register, the designs and the
measurements. The manual and gestures.md stay at the top; everything for
someone changing the code moves to docs/dev/, and the two documents that
name their own successors — the v0.1 milestone and the UI-refinement plan
— go to docs/dev/archive/ rather than being deleted, since both are still
cited. docs/README.md is the index, users first.
Every reference follows: code comments, Cargo manifests, the workflows,
the pre-commit hook, the bench and traceability tools (which locate the
repo root by docs/dev/requirements.md now), packaging, the Docker READMEs,
CLAUDE.md, CONTRIBUTING.md and the README. The matrix links one level
deeper and is regenerated. Links out of the moved documents into the tree
gain a level; a link checker over every Markdown file finds none broken.
The derived sync fired only after the metadata sweep, so a fresh device
re-derived every thumbnail it scrolled past, re-detected faces and re-read
every header for hours before adopting the shards and snapshot that held
all of it. It now fires as soon as the scan completes — the first moment
the rows the merges key on exist — and the sweep starts behind it. In
steady state that pass is one listing.
The catalog merge gains a fourth half: capture metadata (captured_at,
offset, camera, lens, ISO) for images still at metadata_state < 2, matched
by oc:fileid from a remote row at 2. A date is a fact about the file's
bytes, not local state, and the snapshot already carried it. The sweep's
per-chunk query then finds nothing left, and the timeline is whole on a
fresh device without a header fetch.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Choosing "Thorough" made the library look empty. The detector setting
writes under its own faces.model_id, and every reader of "the faces"
keyed on that exact id: the clustering pass, the coverage figure, the
sweep's work list, the shard export and import, and the sync merge's
face matching. On the reference library that restarted coverage at
1,834 of 19,140, drew a People rail of 36 faces for a person with 520,
queued a ~400 GB re-fetch on each device, and stranded the desktop's
3,583 confirmations under the old id: the tablet held the same faces
under the new one and the merge refused to match them. Same photograph,
same box, same embedder, two ids — that is one face, not two libraries.
The embedder half of the id is now the key. embedder_of and embedder_sql
give it to every query; writes keep the full id, so which detector drew
a box stays on record. record_detections is unchanged and is where the
generations meet: an image holds one pipeline's faces at a time, and a
re-detection carries confirmations across by box overlap. The merge's
match_faces applies the same rule within an embedder. The calibration
is keyed on the embedder too, since the similarity space did not change.
Shards travel every generation, each under its own id, and a peer adopts
whichever it is sent — including a stronger detector's pass over an
image it indexed itself with a weaker one, which is the re-detection its
own sweep would otherwise queue, already done. Never downwards: a tablet
on Fast keeps the desktop's Thorough faces. The sweep gains the same
tail — images a weaker detector indexed, after the ones nothing has —
driven by FaceDetector::supersedes, so choosing a stronger detector still
improves the library over time without first making it disappear.
The face shards carry boxes, landmarks and embeddings. What they deliberately
do not carry is who anybody **is** — the person rows, their names, and the
assignments joining the two. Those travel in the catalog snapshot, which is a
whole-file copy and does contain them.
But the snapshot is *merged*, not adopted, and this merge only ever looked at
collections and keywords. `face_shard`'s own module note says people travel in
the snapshot; nothing implemented it. So a second device received every face and
no people at all, and drew an empty People screen over a full catalog. Exactly
what a tablet showed after syncing thousands of faces from a laptop.
What travels is what the user decided, following the rule the rest of this
module already follows — judgements travel, inference is rebuilt:
- **People**, by uuid on `revision`, exactly as a collection is: the name, and
whether the group was set aside.
- **Confirmations**, and **rejections** — "this is not her" is a fact too, and
is why re-clustering does not put it back.
- **The suggestions inside an ignored group**, which are otherwise ordinary
inference but are what anchors the ignore. Without them a group set aside on
one device reappears on the other, the same fault that made "Not interested"
not stick locally.
Ordinary suggestions are not carried. Both devices hold the same embeddings and
clustering is deterministic, so each recomputes them and arrives at the same
answer; shipping them would double the merge for no new information.
**A face has no cross-device identity**, and unlike a collection there is no
uuid to give it one. Both devices do agree on `oc:fileid` and roughly on the
box, so a remote face is matched to the local face on the same photograph whose
box overlaps it most, above 0.5 IoU. That is not a new rule — it is the one
`record_detections` already uses to carry a confirmation across a re-index, and
it is loose on purpose: the question is "the same face in the frame", not "the
same rectangle".
A local confirmation is never overwritten. Two devices confirming one face as
different people is a real disagreement and an assignment carries no revision to
settle it with; taking the remote's answer would let a sync undo what the user
just did on the device in their hands.
The remote's schema is probed rather than assumed: `remote_is_mergeable` admits
any catalog at or below this version, so one written before faces existed, or
before V10 added `ignored`, is ordinary. An absent table skips this half instead
of aborting a merge that would otherwise have succeeded.
Nine tests, including that the name lands on the overlapping face and not its
neighbour in the same frame, that a set-aside group stays set aside, that an
ordinary suggestion does not travel, and that merging twice changes nothing.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Collections synced their names and arrived empty on every device. The names are
keyed on a uuid and worked; the membership union was keyed on
`images.content_hash`, and the schema says plainly what that column is:
"computed only when something needs it (import dedup, reconnect-by-hash), never
in a scan". A library that has only ever been scanned has one for no image at
all, so the join matched nothing and `WHERE ri.content_hash IS NOT NULL`
discarded whatever survived. The union could never have moved a single row.
Measured on a real catalog: 23,174 images, content hashes for 0 of them,
`oc:fileid` for all 23,174, twelve collections, zero members.
So membership now resolves through the file id first, exactly as keyword
assignment already did — `ASSIGN_BY_FILE_ID` was added for this same reason and
its doc comment even notes that membership was still on the hash. It is
recorded for every image the moment a remote scan sees it, survives server-side
rename and move (FR-NC-5), is the same integer on every device pointed at one
Nextcloud, and is already what the thumbnail shards are keyed by.
The content-hash union is kept rather than replaced: a local-only library has no
`remote` rows, and where a hash has been computed it is a true identity that
survives a library moving between servers. Both statements run; `INSERT OR
IGNORE` against the primary key makes the overlap free.
This repairs the merge. It cannot invent membership that no device recorded —
where the rows were never written, collections stay empty until they are filled
in again.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Keywords are catalog state, and the catalog syncs. Without this, two devices
keywording the same library would resolve to whichever synced last, and an
afternoon of work would vanish with no sign it had ever happened.
The vocabulary merges per row on the rule collections already use: revision
first, timestamp only to break a tie, so a device with a skewed clock cannot win
by having the wrong idea of the time. Assignments merge as a set union, which is
FR-NC-9's principle applied to metadata instead of edit nodes — disjoint work
survives on both sides.
Three things needed care and are commented where they happen:
A deletion travels *by name*, not by identity. Both devices may have minted
their own uuid for one word before they ever synced, so deleting by uuid would
tombstone a row nothing was assigned to and leave every photograph still
carrying the word. The union then refuses to readmit a word a winning tombstone
has just removed — without that filter the remote's live assignments would
resurrect it on the very same pass.
Images are resolved by the server's file id first and the content hash second.
Membership has always used the hash alone, but the hash is computed only when
import dedup or a reconnect asks for it, which for most libraries is never — so
a hash-only union would have quietly done nothing for the ordinary photograph.
A word lands on the local default version. Version uuids do not reconcile in the
catalog at all: ensure_default_versions mints a fresh one per device, so a
uuid-keyed join would have unioned nothing.
Removal still does not propagate. That is the trade collection membership
already makes, for the same reason — an unwanted keyword is removed again in a
second, and a silently lost afternoon is not recoverable at all.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The merge inserted every incoming collection with `parent_id = NULL` and
never set it on update, so the hierarchy flattened on each round trip: a
collection nested on one device came back from the server at the top
level. `r.parent_id` was selected and then not read.
The id could not be copied — row ids are local, and the remote's integer
names a different collection here, or none. So carry the parent's uuid
and resolve it locally, in a second pass: rows arrive in whatever order
the query returns, and a child can precede its parent.
Guard the resolution against cycles. Each tree is acyclic alone, but the
union need not be — we may hold A above B while the remote holds B above
A — and closing that loop would make every tree walk spin.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Library setup as the user described it: pick a folder, choose which RAW
types to look for, scan recursively.
dr-types::FormatFilter the tick-box selection, seeing through VFS
placeholder suffixes so a dehydrated CR2 still
matches as a CR2
dr-sync::scan recursive walk, Depth:1 per directory, pruning
unchanged subtrees where the backend propagates
directory ETags
Verified against nextcloud.tourolle.paris (34.0.2) on a real library:
browse root 32 entries, 98ms
scan PhotosRaw 17,185 RAW files in 334 directories, 34.1s
(7,836 CR2 + 9,349 DNG)
range read 262KB of a 21.5MB DNG in 119ms — 1.22% of the file,
and enough to read "Canon EOS 6D | ISO 100"
That last line is assumption A3 validated on real data. Cataloguing this
library by whole-file fetch would move roughly 370GB; the range path
moves a few MB.
Pruning is capability-gated rather than assumed: with per-entry ETags a
probe costs a request and proves nothing about children, so it is skipped
entirely. A test asserts zero probes in that case.
Still unresolved: /core/preview returns 400 for every parameter
combination tried, including on a JPEG the server reports as having a
preview. Not a request-shape bug — it fails identically bare. Recorded
rather than worked around; ARCH §6.7 already treats server previews as
opportunistic, so nothing depends on it.