23a2f13b462013c5d624b67e84a6877a9a49b8d8
5
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
96d07da15f |
Sync face data as sealed shards, so a second device does not re-index
Indexing 23,500 images is about two hours of CPU, and the result is byte-identical on every device: the same model over the same proxy produces the same embedding. Paying for it once per account rather than once per device is the point. Shards rather than the catalog snapshot, because the snapshot goes up whole on every sync and a fully indexed library carries roughly 30 MB of embeddings. That is exactly the cost the thumbnail store's 25 MB cap exists to bound, so face shards use the same cap -- imported from dr_thumbs rather than restated, since the number is a statement about sync cost and the two must not drift apart. The split follows the one already there: bulk immutable data in sealed shards, small mutable data in the catalog snapshot. Faces, landmarks, embeddings and run markers shard; people, names and assignments ride the catalog and merge by uuid. Keyed on oc:fileid throughout, never on image_id, because a row id means nothing on another device. The run marker travels with the faces it describes. Without it a receiving device cannot tell an image with no faces from one never examined, and would re-detect every landscape it had just adopted. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
00e78dc2ac |
Cluster faces into people, and calibrate what a similarity means
FR-CULL-9 forbids thresholding a bare cosine anywhere in the subsystem, so calibrate fits P(same person) per library and reports whether the fit is trustworthy. Two details carry most of the weight. The fit runs against a 200-bin histogram rather than a pair list: a 25,000-face library has ~3e8 pairs and no gradient descent is running over that. And a fresh library has no valid calibration, because the positives have to come from user confirmations or burst siblings -- bootstrapping them from high cosine would fit the calibration to the belief it was supposed to test. Clustering defends against the over-merging FR-CULL-10 warns about with constraints rather than a better threshold: two faces in one photograph never merge, and two groups confirmed as different people never merge. Average link rather than single link, so one strong edge cannot weld two families together. Calibration is defined once, in dr-face, and dr-catalog re-exports it. Two implementations of one probability model is exactly how a number comes to mean the wrong thing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
bb71f141e7 |
Let a folder on this machine be scanned into the catalog
`dr_catalog::scan` has known since it was written what a changed directory means — when to prune, when to list, and the one question that decides whether a deletion sweep is safe. It was fully tested and nothing called it, because walking a real directory "belongs to the platform layer" and the platform layer was eleven lines re-exporting `secrets`. So every photograph in DarkRoom arrived over WebDAV, and a user without a Nextcloud account saw nothing at all. This is the missing half: a `Storage` trait, a filesystem implementation of it, and the driver that pours one into the other. The trait is shaped by the platform it does *not* yet support. Android's SAF gives no filesystem path, which is why `SourceRef` exists; less obviously, it gives no way to *compose* one either — a document id is opaque, and the only way to learn a child's id is the children query that returned it. So a listing hands back the reference to each entry rather than a name for the caller to join onto a parent, and there is deliberately no "path + name" helper anywhere above `LocalStorage`. That single restriction is what makes SAF a second implementation rather than a second set of call sites. A reference is otherwise an opaque `(RootId, key)` pair the catalog stores verbatim and rebuilds later, which a persisted tree grant supports exactly as a relative path does. A `Path` now appears in one place: `LocalStorage::grant`, where the folder the user picked is handed in. Everything above it addresses a `RootId`. `dr_catalog::walk` is the seam. It probes a directory, asks `scan` what that means, lists only when told to, and reconciles what it found against the rows it holds. Two things it does are worth saying out loud, because both are ways to lose a library: Absence only counts where absence was observed. A listed folder proves its missing images are gone; a pruned one proves nothing about its contents, and a scan that was cancelled or that failed part-way proves nothing about folders it never reached. So the file sweep runs per listed folder, the folder sweep runs once at the end and only after a complete scan, and a root that cannot be reached at all marks its images offline and deletes nothing — FR-CAT-9's line between proven-absent and merely-unreachable, which is the difference between unplugging a drive and losing everything on it. A trashed image is absent from its folder on purpose. It is exempt from both sweeps, and detached from a folder about to be deleted rather than cascaded away with it, or a soft delete would come undone the first time the folder it came from was rescanned. Two things the tests taught, both changes to what was there before: Modification times are now milliseconds, not seconds. Change detection asks whether a timestamp moved, so the unit's granularity is the width of the window in which a change is invisible — and a second is long enough to copy a card and start a scan. The test that caught it looked like a test bug; it was not. SAF reports milliseconds natively, so this is also the unit that needs no conversion on the platform with the coarser clock. And an in-place rewrite of an existing file is invisible to directory-level pruning, because writing to a file moves neither its directory's mtime nor its entry count. That is a real limit, now documented and held by a test rather than left to be discovered. It bites less than it reads: an export, a restore, `mv`, and every editor that saves safely write beside the file and rename over it, which does move both. Narrowing the format filter no longer deletes what it stops matching, which fell out of the same principle: unticking JPEG says stop looking for new ones, not discard the hundred already rated. The files are sitting right there. `DirState` and `DirEntry` move to `dr-types`. They are the sentence the platform says to the catalog and both crates need the same one; `scan` re-exports them so nothing that used them has changed. Not done: the UI. The launch screen's "Open library" flow is account-shaped from the first field to the thumbnail worker, and giving it a local branch is its own piece of work rather than a button. `cargo run -p dr-catalog --example scan_local -- ~/Pictures` scans a real folder and reports what it cost; run it twice to see the second run list nothing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d5b1f6bff5 |
Add collections, ratings, and soft delete to the catalog
Three features over a shared schema migration. Collections: a tree of manual collections plus smart collections whose membership *is* their stored selector. Dropping images onto a smart collection is refused rather than silently discarded, so the UI can say why the drop did nothing — member rows there would be a second source of truth that nothing reads. Ratings: the star and pick/reject axes, kept independent. Trash: soft delete to a folder, then permanent delete. Catalog::open now backfills after migrating. A migration adds a column but cannot know what the value should be for rows that already existed; backfilling on open is what stops those rows being silently partial. Timeline queries exclude shadowed JPEGs, which would otherwise double every paired shot in the histogram, and gain a range-bounded variant so zooming in returns finer buckets rather than the same coarse ones with the ends cropped. Assisted-by: LLM |
||
|
|
c8bb08e661 |
Add folder scan with format selection; validate A3 on a real library
Library setup as the user described it: pick a folder, choose which RAW
types to look for, scan recursively.
dr-types::FormatFilter the tick-box selection, seeing through VFS
placeholder suffixes so a dehydrated CR2 still
matches as a CR2
dr-sync::scan recursive walk, Depth:1 per directory, pruning
unchanged subtrees where the backend propagates
directory ETags
Verified against nextcloud.tourolle.paris (34.0.2) on a real library:
browse root 32 entries, 98ms
scan PhotosRaw 17,185 RAW files in 334 directories, 34.1s
(7,836 CR2 + 9,349 DNG)
range read 262KB of a 21.5MB DNG in 119ms — 1.22% of the file,
and enough to read "Canon EOS 6D | ISO 100"
That last line is assumption A3 validated on real data. Cataloguing this
library by whole-file fetch would move roughly 370GB; the range path
moves a few MB.
Pruning is capability-gated rather than assumed: with per-entry ETags a
probe costs a request and proves nothing about children, so it is skipped
entirely. A test asserts zero probes in that case.
Still unresolved: /core/preview returns 400 for every parameter
combination tried, including on a JPEG the server reports as having a
preview. Not a request-shape bug — it fails identically bare. Recorded
rather than worked around; ARCH §6.7 already treats server previews as
opportunistic, so nothing depends on it.
|