Keywords are catalog state, and the catalog syncs. Without this, two devices
keywording the same library would resolve to whichever synced last, and an
afternoon of work would vanish with no sign it had ever happened.
The vocabulary merges per row on the rule collections already use: revision
first, timestamp only to break a tie, so a device with a skewed clock cannot win
by having the wrong idea of the time. Assignments merge as a set union, which is
FR-NC-9's principle applied to metadata instead of edit nodes — disjoint work
survives on both sides.
Three things needed care and are commented where they happen:
A deletion travels *by name*, not by identity. Both devices may have minted
their own uuid for one word before they ever synced, so deleting by uuid would
tombstone a row nothing was assigned to and leave every photograph still
carrying the word. The union then refuses to readmit a word a winning tombstone
has just removed — without that filter the remote's live assignments would
resurrect it on the very same pass.
Images are resolved by the server's file id first and the content hash second.
Membership has always used the hash alone, but the hash is computed only when
import dedup or a reconnect asks for it, which for most libraries is never — so
a hash-only union would have quietly done nothing for the ordinary photograph.
A word lands on the local default version. Version uuids do not reconcile in the
catalog at all: ensure_default_versions mints a fresh one per device, so a
uuid-keyed join would have unioned nothing.
Removal still does not propagate. That is the trade collection membership
already makes, for the same reason — an unwanted keyword is removed again in a
second, and a silently lost afternoon is not recoverable at all.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The merge inserted every incoming collection with `parent_id = NULL` and
never set it on update, so the hierarchy flattened on each round trip: a
collection nested on one device came back from the server at the top
level. `r.parent_id` was selected and then not read.
The id could not be copied — row ids are local, and the remote's integer
names a different collection here, or none. So carry the parent's uuid
and resolve it locally, in a second pass: rows arrive in whatever order
the query returns, and a child can precede its parent.
Guard the resolution against cycles. Each tree is acyclic alone, but the
union need not be — we may hold A above B while the remote holds B above
A — and closing that loop would make every tree walk spin.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Library setup as the user described it: pick a folder, choose which RAW
types to look for, scan recursively.
dr-types::FormatFilter the tick-box selection, seeing through VFS
placeholder suffixes so a dehydrated CR2 still
matches as a CR2
dr-sync::scan recursive walk, Depth:1 per directory, pruning
unchanged subtrees where the backend propagates
directory ETags
Verified against nextcloud.tourolle.paris (34.0.2) on a real library:
browse root 32 entries, 98ms
scan PhotosRaw 17,185 RAW files in 334 directories, 34.1s
(7,836 CR2 + 9,349 DNG)
range read 262KB of a 21.5MB DNG in 119ms — 1.22% of the file,
and enough to read "Canon EOS 6D | ISO 100"
That last line is assumption A3 validated on real data. Cataloguing this
library by whole-file fetch would move roughly 370GB; the range path
moves a few MB.
Pruning is capability-gated rather than assumed: with per-entry ETags a
probe costs a request and proves nothing about children, so it is skipped
entirely. A test asserts zero probes in that case.
Still unresolved: /core/preview returns 400 for every parameter
combination tried, including on a JPEG the server reports as having a
preview. Not a request-shape bug — it fails identically bare. Recorded
rather than worked around; ARCH §6.7 already treats server previews as
opportunistic, so nothing depends on it.