76bc5652d7574490ea7585d2abcaf39481c24947
71
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
facb44cb55 |
Keep the dense landmarks behind each eye reading, packed
The 106 points the eye boxes were cut from, stored beside the reading as 16-bit fixed point over the frame: 424 bytes a face, a seventh of a pixel on a 6000-pixel frame, where f16 at the same size would have been six. Derived data like the embedding, kept for the same reason — it cost a fetch and a model run, and the next per-face pass should run from the catalog. Shards carry it; a peer's shard from before it is still read. |
||
|
|
85cc2b1dcc |
Trace the eye reading to FR-CULL-8a and the chip to FR-CULL-13
The register grew both clauses the same day this was built: FR-CULL-8a is the per-face state the reading is, and FR-CULL-13 is the rule that a signal is shown and filtered and never writes a judgement. The tags, faces.md §17 and catalog.md now say which is which; FR-CULL-8a records what of it is built, and that its third model is under the InsightFace grant by the same decision as the pair. |
||
|
|
d706c12d77 |
Cover the eyes-open subquery with an index
The people filter was served from faces_image without touching a row; reading the eye columns in the same subquery touched every one, and ALTER TABLE had put those seven floats after the embedding and the crop blob. One count took 24 seconds on the reference library, thirteen of them system time. faces_eyes covers the subquery again: five milliseconds. |
||
|
|
54b543fb77 |
Store seven eye numbers per face rather than three
Per eye P(open), the pixels across its box and the sharpness of the patch; and P(sunglasses). The verdict — open, closed, sunglasses, unclear — stays a rule in dr_face::eyes so the floors can move without re-measuring twenty thousand faces. Shards carry the same seven, and a peer's shard from before any of them is still read. |
||
|
|
b908d861e0 |
Keep each face's eye reading in the catalog and in its shard
Three nullable columns beside quality — P(open) for each eye and P(sunglasses) — because the verdict is a rule with thresholds in it and a rule belongs in code, not in rows that would have to be re-measured. NULL is "never read": a face from before the models, or from a device without them, and every reader treats it as unknown rather than as closed. The measuring pass V14 built for the embedding's length is what fills them, so the sweep's work list now also names faces with no eye reading — but only on a device that has the models, or it would fetch every original to do nothing to it. A peer's shard without the reading is still adopted, unlike one without the quality: the pass finds this work by the NULL rather than by the run marker, so adoption costs it nothing. |
||
|
|
327decfab1 |
Fuse every detector's faces into one population per embedder
Choosing "Thorough" made the library look empty. The detector setting writes under its own faces.model_id, and every reader of "the faces" keyed on that exact id: the clustering pass, the coverage figure, the sweep's work list, the shard export and import, and the sync merge's face matching. On the reference library that restarted coverage at 1,834 of 19,140, drew a People rail of 36 faces for a person with 520, queued a ~400 GB re-fetch on each device, and stranded the desktop's 3,583 confirmations under the old id: the tablet held the same faces under the new one and the merge refused to match them. Same photograph, same box, same embedder, two ids — that is one face, not two libraries. The embedder half of the id is now the key. embedder_of and embedder_sql give it to every query; writes keep the full id, so which detector drew a box stays on record. record_detections is unchanged and is where the generations meet: an image holds one pipeline's faces at a time, and a re-detection carries confirmations across by box overlap. The merge's match_faces applies the same rule within an embedder. The calibration is keyed on the embedder too, since the similarity space did not change. Shards travel every generation, each under its own id, and a peer adopts whichever it is sent — including a stronger detector's pass over an image it indexed itself with a weaker one, which is the re-detection its own sweep would otherwise queue, already done. Never downwards: a tablet on Fast keeps the desktop's Thorough faces. The sweep gains the same tail — images a weaker detector indexed, after the ones nothing has — driven by FaceDetector::supersedes, so choosing a stronger detector still improves the library over time without first making it disappear. |
||
|
|
9b627e7713 |
Let a catalog writer wait for its turn instead of losing its work
SQLite's busy timeout defaults to zero, and nothing ever set one: the loser of a write race got SQLITE_BUSY at the moment it asked. WAL does not cover this — it makes one writer and many readers free, and this application constantly has two writers, the face sweep committing a batch while the derived sync imports shards or reclustering reads. The cost was not a retry but lost work. A sweep that had already paid for the detection and the embedding — seconds per image, the expensive part — discarded the result on "storing faces for 214: database is locked" and moved on to the next image. Both the desktop and the tablet logged runs of those on consecutive images, which is a face sweep quietly failing to store the faces it had just computed. Ten seconds, on every connection, set in configure() so that nothing can open the catalog without it — the figure the job runner's own tests have used for this reason since they were written. It is far longer than any transaction here, so it bounds pathology rather than making anyone wait. |
||
|
|
eeee3d920a |
Back the catalog up daily, not only before migrations
NFR-R2 asks for the catalog to be backed up on a schedule and before schema migrations. Only the second half existed: every backup on disk was a pre-migration copy, and a library that never migrated was never backed up at all. A backup is now also taken at the end of a library sweep when the newest one is more than a day old — the moment the catalog is quiet and a day's collection and people edits have just been folded in — on its own thread and its own connection, so the copy of a 130 MB file is not spent on the UI. Whether one is due is read from the backup directory, not the catalog, so the ordinary case costs nothing. An empty catalog is skipped: there is nothing in it a rescan would not rebuild. Pruning to KEEP_BACKUPS applies as before. |
||
|
|
f100db89ca |
Verify the catalog snapshot before it is sent, and after it lands
Two checks around the upload, both cheap next to what they prevent. Before: the snapshot is quick_checked before it leaves. It is the copy every other device merges from, and a damaged one costs each of them a download, a failed merge and a refusal to push. After: the staged upload's size on the server is compared to the bytes sent before it is rotated into place. A chunked upload is assembled server-side, and an assembly that goes wrong is a file of plausible size no device can open — caught here, on the device that caused it, for one listing; otherwise on every other device, after the fact. A mismatch, or a size the server will not confirm, discards the upload and leaves the current copy and its generations untouched. |
||
|
|
896188a489 |
Read the sidecars other editors write, and write them back on request
Benchmarks / CPU and I/O (per commit) (push) Successful in 10m59s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Successful in 1h33m36s
Build and test / Layer separation (push) Successful in 1m2s
Traceability / Requirement traces (push) Successful in 1m25s
🐳 Android image / Build and push (push) Successful in 9s
Build and test / android-image (push) Successful in 9s
Build and test / Android (aarch64) (push) Successful in 56m59s
FR-CAT-13 asked for standard XMP and `core/dr-xmp` answered the file: it
has read and written `dc:subject`, `xmp:Rating`, `xmp:Label` and the IPTC
core since
|
||
|
|
adf5d6cdd9 |
Drop a rival pipeline's marker when an image is re-indexed
record_detections replaces every face on an image whatever model found them, but left the other models' face_index rows standing. With one model that was unobservable. With a second pipeline it leaves an image marked "done" under the first with none of its faces behind the marker — the state the V12 repair existed to undo — and a user who switched back would find those photographs permanently empty. An image now holds the faces of whichever pipeline looked at it last, and only that pipeline's marker. Confirmed names still carry across by box overlap, since they were read before the replacement. |
||
|
|
16f3fb41a3 |
Measure the faces already found rather than finding them again
Benchmarks / CPU and I/O (per commit) (push) Successful in 3m13s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 54s
Build and test / Layer separation (push) Failing after 1s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 2s
Traceability / Requirement traces (push) Successful in 52s
Build and test / Android (aarch64) (push) Successful in 30m6s
Every face stored before its quality was kept holds a unit vector, and V14 forgot the run marker of each image holding one so that the next sweep would look again. Looking again meant detecting again: a whole re-detection per image, with every suggestion on it thrown away and the confirmations carried across by box overlap, to recover one number. The sweep now has a measuring pass between the proxy repair and the un-indexed images. It lists every image holding an unmeasured face, fetches the original once, warps each stored face from the landmarks it already has, embeds it, and writes the raw vector and its length over the old row. Ids, boxes and identities are untouched; the marker is re-written fresh so the sync exports the measured vectors. A face whose landmarks no longer make a warp is dropped, as detection would have refused to store it. `faces_unindexed` leaves those images to the measuring pass, so the V14 deletion no longer costs a second detection. |
||
|
|
8b3abdb787 |
Keep each face's quality, and never compare against a poor one
The embedder's raw output has a length, and the length is a reading of how recognisable the crop was: a blur, an occlusion or a hard profile comes out short. Normalising threw it away. A short vector sits near the middle of the sphere and matches a little of everyone, which is how one bad crop bridges two people in a grouping pass. So the length is kept — the store now holds the raw vector, re-normalised on load, with the length beside it as `faces.quality` — and a face under MIN_GALLERY_QUALITY (14) is a probe: measured against the gallery and placed where it fits, but never what another face is measured against. Two probes are never paired, and a probe is nobody's evidence for a confidence. The People screen shows the number as "Quality 17.3", dimmed below the floor. Faces indexed before this stored unit vectors and have no reading; they are admitted to the gallery, and schema V14 forgets the run marker of every image holding one so the next indexing pass measures them. A peer's unmeasured shard faces are not adopted, or a sync would write that marker back. |
||
|
|
bac5801618 |
Ask which collections a selection is filed in, and how much of it
`collections_for_image` answers this for one photograph and has no counts, which is enough to badge a cell and not enough to offer a removal: with forty selected and three of them in "Iceland", a sheet that says only "Iceland" invites the user to take all forty out of a collection thirty-seven were never in. `membership_of` returns the count alongside the name so the row can say "3 of 40". Chunked over the image list rather than one `IN (...)`, because the list is a selection and a select-all makes it as large as the library — past SQLite's bound-parameter cap on exactly the gesture most likely to produce it. Counts are summed across chunks, so the answer is the one the unchunked query would have given. Smart collections are excluded by construction: they have no member rows, so there is nothing a removal could do. |
||
|
|
1ad35e2b87 |
Read the library's sidecars, so a cull done elsewhere arrives
Judgements only ever travelled outward. A rating went to the catalog and to the photograph's sidecar, the sidecar reached the server, and there it stopped: the scan indexes files, `derived_sync` exchanges thumbnails, face shards and collections, `dr_catalog::merge` reconciles everything in a catalog except `versions.rating` and `versions.flag`, and the one sidecar reader that existed ran when a single photograph was opened in develop and handed its answer to the develop graph. `JobKind::ReadSidecar` was declared for exactly this when the job queue was written and was never enqueued or handled anywhere. The grid draws `versions.rating`. So a day of culling on the tablet could not reach the laptop by any path the application had, and the laptop's catalog says so plainly: 23,568 images, one of them judged. `pull_sidecars` closes it, off the back of work the scan already does. `dr_sync::scan` reports the `.drsc` files it meets in listings it was making anyway — no extra request, and a directory whose ETag is unchanged is still pruned before it is listed at all. A new `sidecars` table records the ETag of each one this device has taken in, so the fetch is one GET per sidecar that genuinely changed rather than one per photograph. A library nobody has edited costs nothing. The judgement is taken rather than maximised. The sidecar is the authoritative store and the fuse has already settled any contest between devices on `revision`, so lowering a rating from four to one on the tablet lowers it here — taking the larger would have refused every demotion the photographer ever made, which is most of what a second pass over a shoot is. A zero is the exception: it means *never judged*, not "judged zero", so a sidecar carrying none cannot erase a star this device holds. That is `merge_judgement`'s asymmetry and it carries the same known cost — clearing a rating does not propagate. A sidecar names a stem, so both halves of a RAW-and-JPEG pair are judged: they are one photograph (FR-CAT-11) sharing one document, and judging only one of them would leave the grid disagreeing with itself over which it drew. The `LIKE` that finds them is a filter, not the decision — `sidecar_path` is applied to every candidate, because a folder is entitled to contain a `%` and a rating landing on the wrong frame would be silent and permanent. Failing to read one is not a failure to scan: the ETag goes unrecorded, the ratings already here stay where they are, and the next scan tries again. The count is reported to the status line as well as the log, because a grid that silently gains three hundred stars is indistinguishable from one that has gone wrong — and because while this number was structurally zero there was nothing to tell the photographer their cull had not arrived. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ca2a135e28 |
Let two devices name the same photograph's version the same way
A version's uuid is the identity a cross-device merge keys on, and it was minted at random, per catalog, per image. Two devices indexing one Nextcloud library therefore held two different uuids for the same photograph — so the sidecar they shared collected a `default = 1` block each, `Version::merge` was never handed a matching pair to reconcile, and an afternoon's culling on the tablet did not exist as far as the laptop was concerned. `crate::merge` has said so in a comment since it was written: version uuids do not reconcile across devices, a uuid-keyed join unions nothing, so keywords are landed on the local default version instead. It named the problem and worked around it. `rating`'s own comment asserted the opposite — that generating the uuid here was what made it a cross-device identity — and `library::amend` repeated the claim. Uniqueness was never the difficulty; agreement was. `derived_version_uuid` computes it from `oc:fileid` instead. The server assigns that integer, every client pointed at the library sees the same one, and it survives a server-side rename and move — the three properties that already made `ASSIGN_BY_FILE_ID` prefer it to a content hash. The layout is a UUIDv8 (RFC 9562, an application-defined form) carrying all sixty-four bits verbatim across the variable fields with a fixed tag in the node field, so the mapping is injective by construction rather than by a hash's good behaviour, and a uuid in a sidecar can be read back to the file it belongs to by eye. A library with no server behind it has no shared identity to derive and keeps a generated one. The split is still reachable there if the folder is synced by something else; `Sidecar::fuse_default_versions` repairs that case rather than preventing it. Deriving it for new rows alone would have fixed nothing — every image in an existing library already has a version, so every one of them would have carried on writing to its own rival identity. `align_default_version_uuids` moves them, and runs from `schema::backfill` on every catalog open. It selects on the tag in SQL, so a catalog already realigned matches no rows and writes nothing, and it declines rather than fails where a virtual copy already holds the target. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d44bffa4a8 |
Remember where the photographer was
Opening the application was always a fresh arrival at the beginning of the library, whatever you had been doing when you closed it. What is written down is the view, the scope, the rating filter and the photograph on screen -- the open one in develop, the first visible one in the grid. Not just a scroll position: a position without the filter that produced it names a row of a list that no longer exists. Restoring them has an order for the same reason -- scope, then filter, then position, then the view -- because each step changes what an ordinal *means*. Addressed by remote path and collection UUID, never by an ordinal or a row id. `images.id` and `collections.id` are local to one catalog, and a grid ordinal is local to one ordering; a record naming either would land somewhere arbitrary on a second device and after any filter change on this one. Where the ordinal is needed, `library::ordinal_of_path` computes it through the grid's own `ORDER BY`, taken verbatim by a window function rather than spelled a second time as an inequality -- which is the mistake `grid_order_for` already warns about, and which a manually ordered collection would make unreadable. Every failure degrades rather than reports. A collection this device has not merged leaves the scope at the whole library; a photograph that has since been deleted falls back to when it was taken, which puts the grid in the right week; a torn file yields no place and the library opens at the top. Reopening develop is the one thing that requires an exact match, because a canvas on a path that no longer resolves is a filename over an empty frame. The record lives in `dr-types` beside `Settings` and the store lives here beside `SettingsStore`, for the reason `dr-types`' manifest gives: a JSON serialiser in `core/` would be paid for by every crate there. Two files and two lifetimes, though -- resetting preferences must not forget where you were. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
e7b526c550 |
Specify face indexing at native resolution, and say what the proxy cost
FR-CULL-8 said detection runs against the thumbnail or proxy tier and never a full decode, and faces.md §5 said the aligned crop is sampled from that same proxy. Both are wrong in the same place: they treat detection and cropping as one resolution problem when they are two, with opposite answers. Detection does not care. §4.1 fixes the graph's input at 640x640 and letterboxes whatever arrives, so a face filling 2% of the frame reaches the model at 12px whether the buffer handed over is 1024px or 6000px. Every pixel above the detector's own input is discarded before inference. The crop cares about nothing else. §5's warp produces the fixed 112x112 ArcFace sees, so source resolution converts directly into whether those 112 pixels were photographed or interpolated. Reading crop_px across the 18,671 faces the proxy-tier implementation stored: 47.3% were upsampled to reach the embedder, 314 of them by more than 2x, the smallest from 34 source pixels. An upsampled crop does not fail loudly -- it yields a confident embedding of detail that was never there, and the damage appears three stages later as clusters that will not separate. So FR-CULL-8 now specifies four stages with the resolutions named separately: render native through FR-EXP-9's pipeline, downscale for the detector, map boxes and landmarks back to native, crop and align from the native render. The affordability the old rule bought is met instead by when the pass runs -- background, preempted, resumable -- and the requirement says plainly what it now costs on a remote library: the original rather than FR-NC-3's byte range, 412 GB across the reference library's 19,107 images, so a whole-library pass is a transfer under FR-NC-6 rather than something that may start on its own. MIN_CROP_EDGE replaces the MIN_DETECT_EDGE this branch briefly had. Same number, guarding the quantity that turned out to matter. faces.md §7b records both measurements, and marks the second as unexplained rather than dressing it as a finding. Grouped by the buffer detection ran against, faces per image was 0.078 at 1024 or below and 1.82 at 2048 or better, controlled for file type and size. That gap is real and reproducible and I cannot account for it, because the letterbox above says detector input should not matter. M4 is where it gets settled. The crop measurement does not depend on it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d0ebc9f571 |
Count the same images in the progress figure that the sweeps index
The Identity screen said 4,593 images were left to index and stayed there for hours across repeated runs, which is what a stuck job looks like. It was not stuck. 4,424 of those 4,593 are shadowed -- the JPEG half of a RAW+JPEG pair -- and no sweep will ever index one, because every work list is built on VISIBLE, which excludes them. They are not separate photographs and the grid does not show them either. But faces::coverage counted them: its denominator was "images WHERE trashed_at IS NULL", with no shadowed_by clause. So the outstanding figure had a floor of 4,424 that no amount of work could bring down, and Coverage::is_complete could never once return true no matter how completely the library had been indexed. A progress number that cannot reach its own target is worse than no progress number. The fix is to count the population the sweeps actually draw from, in all three places that were describing it differently: coverage's denominator and its indexed join, and audit's split of the outstanding set, which had the same gap and fed the same status line. On the reference library the denominator goes from 23,531 to 19,107 and outstanding from 4,593 to 169 -- the second of which is a number the user can watch go down, and which turns out to be a real and separate fetch failure worth chasing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
0a6509de93 |
Forget the face runs made on proxies too small to see a face
The floor added in the previous commit stops this happening again; it does nothing about the 1,824 images in the reference library that already carry a face_index row written against a proxy of 1024 or less. Those rows are why the damage is permanent rather than merely past. The work list is "images with no row for this model", so an image examined against a 1024px proxy -- 0.078 faces per image, nine in ten finding nothing -- is indistinguishable from one examined properly, and no later pass will ever offer it to the detector again. V12 deletes exactly those markers, and nothing else. The faces those runs did find stay in place and keep drawing the People screen until a better pass replaces them, and record_detections re-attaches the user's confirmed names across that replacement by box overlap, so a library somebody has spent an evening naming does not lose that evening. The cost is a re-fetch of the affected images. Deleting the marker rather than teaching the work-list query to select on source_edge, which was the other option and is worse. A standing `source_edge < floor` predicate never lets go: an image whose largest embedded preview is genuinely smaller than the floor would be re-fetched on every sweep for ever, because the next pass cannot do any better than the last one did. A one-off deletion gives each affected image exactly one more attempt through the good path and then lets the ordinary "has a row" rule settle it. The threshold is written out in the SQL instead of referring to dr_face::MIN_DETECT_EDGE. A migration has to keep meaning what it meant when it ran; binding it to a constant someone may raise later would quietly change what an old catalog gets migrated to. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
19b56ad85c |
Merge: collection ordering, and a range that says where it ends
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> # Conflicts: # docs/traceability.md # ui/dr-ui/src/collections_ui.rs # ui/dr-ui/src/library.rs # ui/dr-ui/ui/app.slint # ui/dr-ui/ui/library.slint # ui/dr-ui/ui/widgets.slint |
||
|
|
d70dcf78d1 |
Merge: recover a damaged catalog, and capture a crash locally
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
8eeb9ba0f6 |
Offer the backup, and then the rebuild, when the index turns out to be damaged
NFR-R6 asks for an integrity check at startup and two offers behind it, and none of it existed. `PRAGMA integrity_check` appeared nowhere in the tree, `Catalog::open` was `open` → `configure` → `migrate` → `backfill` and nothing else, and corruption therefore surfaced as whatever rusqlite error the first unlucky query happened to produce — "database disk image is malformed" attached to a thumbnail refresh, elided into a 34px banner, over an empty grid saying "No images found · Check the library folder". Two messages that disagreed, and no way forward but deleting catalog.sqlite by hand. The property that makes the second offer real was already here and load- bearing: the catalog is an index, not a source of truth, rebuildable from sources plus sidecars (invariant §5.2.4, cited by schema.rs, trash.rs and lib.rs). And sync.rs already knew how to take a coherent snapshot of a WAL database. What was missing was the check, the type, and the conversation. Four pieces: **The type.** `CatalogError::Corrupt`, and — the part that makes it worth having — a hand-written `From<rusqlite::Error>` that classifies rather than wraps. `SQLITE_CORRUPT` and `SQLITE_NOTADB` become `Corrupt` wherever they arise, so a background job that trips over the damage first reports the same thing the startup check would have. `SQLITE_IOERR` and `SQLITE_BUSY` deliberately do not: a dropped network mount is a different problem, and telling someone to rebuild their index would be a wrong answer delivered confidently. **The check.** `Catalog::open_verified`, `quick_check` before the open rather than after, because opening runs migrations and a damaged catalog with an intact header would otherwise have structure rewritten on top of structure that is already wrong. Bound to `open_verified` and not to `open`: the check reads every page, which is affordable once at startup where a user can answer a question, and not affordable on the dozens of opens a session's background tasks make. **The backup.** NFR-R2's second clause, taken between `configure` and `migrate` in `Catalog::open`. A migration is the one routine operation that rewrites table structure, so it is the likeliest way this file becomes unreadable, and it is the last moment the pre-migration state exists to be copied. Three generations, through SQLite's backup API after a TRUNCATE checkpoint — never `fs::copy`, which on a WAL database backs up a state older than the catalog and possibly torn. A failure to take the copy is logged, not raised: a full disk must not be what makes a library unopenable. **The conversation.** The first line of the dialogue is that the photographs and the edits are safe, before the diagnosis, because that is the question the user is actually asking. Then the two offers, which are *not* interchangeable and are not presented as if they were: a restore keeps collections, and a rebuild cannot, because a manual collection is a set of images assembled by hand and nothing in the filesystem records it (docs/catalog.md §8.1). The labels say so, and the rebuild does not take the affirmative styling while a restore is on the table. One thing that is a fix rather than a feature: `show_catalog_now` now gates the scan. `Catalog::open` succeeds on a file whose header survived, so the scan that used to start immediately afterwards would write folder ETags and image rows into damaged pages in the seconds while the user was still reading the question — turning a file that had a backup into one where the backup is the only copy left. Restore also deletes the damaged catalog's `-wal` and `-shm`. That step is easy to leave out and fatal to leave out: a journal belonging to the old file, sitting beside the new one under the same name, is replayed into it on the next open. That is not a restore, it is a fresh corruption with the evidence gone. Tested by corrupting a fixture catalog — 500 images and a collection, then every page past the second overwritten — and driving both branches. The restore is asserted on the collection, because a collection is precisely what distinguishes the two paths; the rebuild on the damaged file being kept and the next open producing an empty catalog at the current schema. Plus the `SQLITE_NOTADB` presentation, a damaged backup being refused rather than installed, and a v1 catalog whose pre-migration backup comes back reading v1 rather than v11. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
421a47f1eb |
Untag FR-CAT-13, which no XMP is read or written to satisfy
FR-CAT-13 asks for standard XMP sidecars read and written — ratings, colour labels, keywords and hierarchical subjects, title, description, copyright, GPS, in `xmp:`/`dc:`/`lr:` schemas — so that other tools interoperate. Its one tag was the module header of `dr-catalog/src/keywords.rs`. That module stores keywords in SQLite. It names `dc:subject` twice, both times in prose explaining why a keyword's text is the fact rather than its row id, which is a good reason to have written it that way and not evidence of an XMP implementation. Nothing in the tree parses or emits XMP: `dr-export`'s metadata module writes EXIF and says in its own header that IPTC and XMP are named by FR-EXP-8 and neither is read. `dr-preset-xmp` is the crate whose name most invites the mistake. It reads Lightroom `.xmp` *presets* — develop settings — under FR-DEV-6, and knows nothing about the metadata schemas FR-CAT-13 is about. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
758436cc28 |
Keep the runner's borrow alive as long as the connection it reads
The first compile this branch had. One borrow error, in the four-thread contention test: the `Runner` was the block's tail expression, and a tail's temporaries are dropped after the block's locals, so it outlived the `conn` it borrowed. Bound to a local, with the ordering rule written down beside it -- it is exactly the shape someone tidies back. Everything else stood: clippy clean at -D warnings, and all 18 runner tests pass, including the four-thread four-connection claim and the `UPDATE ... RETURNING` rewrite the author flagged as the riskiest line in the diff. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
846a249156 |
Drain the queue that nothing has ever drained
`jobs` has been a complete durable work queue since the catalog was written, and nothing has ever taken a job out of it. `claim_next`, `complete`, `fail` and `recover_orphaned` had no callers outside their own tests; `enqueue` had three. So the table grew one row per photograph and kept it forever, and FR-PLAT-AND-3's resumability was a property of code that never ran. `runner` is the missing half. It owns no thread, no clock and no policy, and that is the whole design: on Android the process does not decide when background work may run. WorkManager does, subject to Doze, battery saver and FR-NC-6's network constraints, and it revokes permission mid-job by calling onStopped(). So the runner exposes `run_one` — claim, run, record — and `drain`, which repeats it against a budget, a deadline and a cancellation flag the host owns. A `Worker.doWork()` with ten minutes calls drain with a deadline; a desktop idle pass calls it with none. That is the seam the Android service plugs into, and it needs no Android to test. Handlers are supplied from above, because the catalog knows what needs doing and nothing about how: a thumbnail needs a decoder and a fetch needs a network stack, neither of which belongs under core/dr-catalog. A runner claims only kinds some handler declares, so a queue holding work this device cannot do is left alone rather than failed five times. Four outcomes, and only two of them are the job's fault. Done deletes the row; Retry backs off; Abandon gives up now, for a failure no retry can fix; Interrupted releases the claim with its attempt refunded and ends the drain, because the host stopped rather than the job — five backgroundings in a row must not mark good work as failed. Process death is the fifth and cannot report itself, which is what `recover` is for. Recovery is called from `show_catalog_now`, which is the one place a catalog is opened for a session and already returns early if one is open. It has to be exactly once and before any worker starts: there is no owner column, so a second pass while a worker held a claim would take it away. The attempt a dead claim consumed is deliberately kept — a job that takes the process down with it is indistinguishable from one that fails, and the attempt counter is the only evidence that survives a death. The tests cover claiming under contention twice over: sequentially across two connections, and with four threads on four connections against one catalog on disk, asserting every job ran exactly once. Plus completion, backoff, giving up, abandoning, interruption, budget, deadline, cancellation, and a job orphaned by a simulated crash being reclaimed and run once rather than lost or repeated. Not wired to a handler yet, and deliberately not: the only enqueue site the app actually reaches is the remote scan's, whose thumbnails are already served by the async grid worker, and `walk`'s two sites are reachable only from the scan_local example. Inventing a handler to make the plumbing look used is how a requirement comes to read as covered by code that does not implement it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d63872e5a9 |
Make a claim one statement, and give the queue what a runner needs
The claim was a deferred transaction around a SELECT and an UPDATE, and under a single connection that is fine. Under two it is not what it looks like: the SELECT takes only a read lock, the UPDATE tries to upgrade, and in WAL a worker that read the same snapshot as another gets SQLITE_BUSY_SNAPSHOT on its write. That is not an error a busy handler can retry away — the fix is to roll back and start over — so the queue was "safe" only in the sense that the loser failed loudly instead of taking a job someone else was holding. `UPDATE jobs SET state = 1, attempts = attempts + 1 WHERE id = (SELECT ...) RETURNING ...` is one statement and so one implicit transaction that takes the write lock immediately. Two workers serialise, the loser waits out its busy timeout, and neither can see a row the other already holds. The existing tests are unchanged by it, because from one connection the two forms are indistinguishable — which is exactly why it was never noticed. The rest is the surface a runner has to have and did not: - `claim_next_matching` takes only kinds a worker can actually do. Without it a device with no connector claims `FetchOriginal`, fails it, and pays five wakeups and five backoffs per photograph to reach a conclusion known before it started. Filtering after a claim cannot work: the claim has already marked the row running. - `abandon` gives up now, for failures no retry can fix. `fail` uses it for its own MAX_ATTEMPTS branch, so there is one statement that ends a job. - `release` hands a claim back with its attempt refunded, for a worker that is being stopped rather than a job that is going wrong. `attempts` stands in for the owner column the table does not have: it is bumped by every claim, so a stale worker's release matches nothing and changes nothing. - `reap_orphan_subjects` deletes jobs whose photograph is gone. Coalescing keeps the table one row per unit of work and nothing ever shrank it when the work stopped existing. `ScanFolder` is excluded because its subject is a folder id, and joining that against `images` deletes by coincidence of numbering — hence `JobKind::subject_is_image`, and `JobKind::ALL` so the next kind added cannot quietly fall out of the filter. - `counts` is the number a foreground service's notification is built from. One behaviour change worth stating: a kind this build does not recognise is now parked with an error rather than read as `ExtractMetadata`. The old `unwrap_or` would have run a job of an unknown kind as some arbitrary known one, which is worse than not running it at all. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b9891e2c04 |
Merge master into tablet-selection
Two real conflicts, both from work that landed either side of the same lines rather than against them. `lib.rs`: the settings controller was hoisted above the People screen's wiring, and Android's thumbnail-tier eviction registered itself at the same point. Independent, so both stay. `library.rs`: manual collection ordering and burst folding each added a clause to the same two queries. The scoped range read now carries both — the folding matters there for one step further on than it does in the grid, because a collapsed burst is one cell, so an ordinal counted over a list still holding every frame names a photograph several places away from the one the user pointed at. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
82d9d077b0 |
Merge: answer Android's memory warnings, and stop reporting a lost root as an empty library
FR-PLAT-AND-5 in full, FR-PLAT-AND-2 in part -- the recovery is built and live for Nextcloud roots, the SAF cause it names does not exist yet. FR-PLAT-AND-4 and FR-PLAT-AND-6 are not here, both blocked behind the same gap: assemble-apk.sh compiles no Java, so the APK cannot carry a Service or a FileProvider. The container has JDK 17 and build-tools 36; the build step is what is missing. Verified: fmt, clippy --workspace --all-targets -D warnings, and 1043 tests across dr-catalog, dr-sync, dr-sync-folder, dr-sync-nextcloud, dr-plat and dr-ui. The aarch64 target was checked before the branch was finished but not after; no device was available. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
75fd5619ca |
Refuse a scan whose root has gone, instead of reporting it empty
FR-PLAT-AND-2, and a silent failure on both platforms. `dr_sync::scan` stepped over a NotFound or PermissionDenied the way it does for a child that vanished mid-walk -- correct for a child, wrong for the root, where it ended the walk, returned Ok with nothing in it, and reported a successful scan of a library that was no longer there. A lost root is now its own error. The images under it are marked Availability::Offline per FR-CAT-9 and no catalog row is deleted; `library::persist` clears the mark per file as each one is listed again, so a root that comes back needs no repair step. Partly satisfied rather than closed, and the gap is worth stating. The recovery half is real and reachable on Android today, because `map_status` turns Nextcloud's 403 and 404 into it and Nextcloud is how a phone actually gets a library in this build. The causes the requirement names -- revocation, reinstall, a removed card -- are properties of a persisted tree permission, and there is none: SAF does not exist here, `SourceRef::Document` is constructed only in test modules, and `LocalStorage` rejects the variant outright. When SAF lands it becomes a third producer of this error and nothing above it changes, which is why the discovery belongs in the connector. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
fc4157a1e0 |
Keep the test scene's own arithmetic from overflowing a u64
The first compile this branch ever had. `cargo fmt` reflowed four files and clippy passed at -D warnings untouched, but one test panicked: `the_signature_does_not_change_with_scale`, on "attempt to multiply with overflow". It is the fixture, not the feature. `scene()`'s little LCG multiplied the block's y by the golden-ratio constant with a plain `*` while the term beside it already used `wrapping_mul`, so any scene taller than about 104 pixels overflowed in debug. Only the scale test builds one that large, which is why 345 of 346 passed around it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
7ad75dd905 |
Let a manual collection be put in the order the photographer wants
`collections::set_order` and `Sort::CollectionPosition` have been in the catalog since collections were, and nothing above dr-catalog has ever called either. A manual order existed, could not be seen, and could not be set. This is the half that was missing. Three pieces, because it needed all three to be visible at all. The catalog gains `orders_manually` and `members_in_order`. The first is the rule about *when* a manual order means anything, kept in one place with one name: a collection must be manual, and it must have no children. A set shows its descendants' images, and positions are only ever assigned within one collection — so two children's positions are unrelated integers, and ordering by them would sort the grid by a coincidence. The existing comment on `read_cells_scoped` already argued this; now something enforces it. `members_in_order` returns the *whole* membership rather than the filtered view, because `set_order` renumbers exactly what it is handed. Reordering a filtered list would renumber those and leave every hidden image on a stale position — two images sharing one, and a grid that rearranges itself the moment the filter comes off. The grid reads position where the scope qualifies and capture time everywhere else. Manual order joins the member row rather than testing membership with `IN`, which is safe from fanning out rows *because* that branch is a single collection. The gesture is a DropArea over the viewport, drawn only where a reorder means something, with a caret in the gap the photographs would go into — a line between two images rather than a highlight on one, because lighting up a cell would say the drop replaces it. The trap worth naming: `DropEvent.position` is in **window** coordinates. Slint maps it through `map_to_window` when the drag begins and hands every target the same event untranslated, so a target inside a Flickable has to subtract its own `absolute-position`. Getting that wrong is invisible until the grid is scrolled, because at the top the two frames coincide. `reordered` is pure and names its destination by the image it goes before rather than by an index, because the grid can only name a gap in what it is showing and the ids are what survive a window swap. A drop that changes nothing returns the order untouched: that counter is what a cross-device merge resolves by, and spending a revision on a no-op makes this device win an argument it did not have. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d3dadbd725 |
Group the frames of one moment, by when they were taken and what they look like
A burst is the commonest thing in a cull and the least interesting: twelve frames of the same gull at 10 fps occupy twelve cells, are scrolled past twelve times, and end with the photographer keeping one. FR-CULL-5 asks for them to collapse to one representative and be judged as a unit. Two signals, because neither alone survives a real library. Time alone groups a whole wedding ceremony -- a photographer working steadily never leaves the gap that would end the run. Similarity alone groups a studio setup shot across two days, which is a project rather than a moment. Together they are specific: adjacent in time *and* looks like the frame before it. Two seconds is the time bound, and the reason is worth recording because the figure looks absurd next to a 10 fps camera. `images.captured_at` is whole seconds -- EXIF's DateTimeOriginal has no sub-second field and SubSecTimeOriginal is optional and widely omitted -- so a burst arrives in the catalog as ten frames sharing one timestamp. Any threshold finer than a second is a threshold on information that is not there. Where the pace really is faster, the similarity bound is what separates the frames. Similarity is a 64-bit difference hash over a 9x8 box-averaged reduction, compared between *adjacent* frames only. Chained rather than anchored on the first frame, because by frame twenty a camera following a bird has nothing in common with frame one while no two neighbours differ by much; the time bound is what stops the chain running away. There is no all-pairs step and there must never be one -- that is what turns a grouping pass into something nobody can afford to run over 50k images. Nothing here ranks a frame. FR-CULL-5 names the failure it is avoiding, which is rejecting the only frame of an important moment because somebody blinked, so there is no sharpness score and no best-of-burst. The representative is the earliest frame -- a fact about the clock, not a judgement about the photograph -- and the user's own choice lives in its own table so that rebuilding the grouping cannot erase it. Same argument `people.ignored` makes one subsystem over: nothing short of remembering a decision survives re-clustering. A newly found burst is recorded *open*. Collapsing on discovery would be tidier, and would also mean a background pass taking photographs off the screen part way through a cull. The pass marks; the user folds. It is a pass rather than a job kind for the reason catalog.md 10.2 gives for face clustering: a burst is a property of a run of frames and has no natural subject_id, so a per-image job would rebuild the world once per photograph. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
574107bc39 |
Let a manual collection be put in the order it is meant to be seen in
`collection_members.position` and `Sort::CollectionPosition` have been in the catalog since collections were, and nothing above dr-catalog has ever written or read either: `collections::set_order` had no callers, and the grid ordered everything by capture time whatever it was scoped to — dr-ui does not construct a `Query` at all, it has its own `GRID_ORDER` constant. So a manual collection was a set with an order nobody could see or change. Three pieces, because it could not be fewer: `grid_order_for` decides the ordering from the scope, and both readers take it from there. That is the load-bearing part. An ordinal only names a photograph relative to an ordering, so the window read and the span read have to agree — a shift-click resolved through a different ORDER BY than the cells were drawn with selects a different run than the one on screen, and the user finds out when the export runs. `read_ids_span` already stated that invariant about `GRID_ORDER`; this widens it to an ordering that depends on the scope. Only a single manual collection has one. A set draws its descendants' images too, and two children's positions are unrelated integers that interleave arbitrarily; a smart collection has no member rows to carry a position at all. Both fall back to capture time and refuse the drop rather than pretending. The drop is on the cell, on whichever half of it the finger landed — the trailing edge is the only way to name the last place in a collection, since there is no cell beyond the last one to drop in front of. `reordered` is pure and the membership is rewritten whole. `set_order` sets the positions it is given and leaves the rest, so a partial write would interleave the moved run with rows nobody touched; and it is read unfiltered, so what the filter is hiding keeps its place relative to what the user can see. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
0407fb8d2d |
Format the two new examples
They were written after the last fmt run and CI gates on --check. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
da20d42d33 |
Merge master: pluggable storage, and a name that anchors
Conflicts were docs/traceability.md alone, and it is generated — so it was regenerated rather than hand-merged. dr-face was untouched on the other side; ui/dr-ui/src/faces.rs and identity_ui.rs auto-merged, the first around recluster's anchoring and the second around load_faces. Worth recording because the two branches met on the same problem from different ends. Master's "Let a name hold a group together" is the fix for the sixteen Catherines — fourteen of them empty — that this branch found while measuring the library and reported without fixing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b2250cc460 |
Measure a regroup on the tablet, not just on the desktop
The GPU question needed a number nobody had: how a regroup divides on the hardware whose CPU is weakest. dr-face carries no weights and touches no display, and dr-catalog's example needs only a catalog file, so both run under adb shell against a copy of a real library. On the same 18,143 faces — desktop against the tablet — scan 0.96s / 2.61s, agglomerate 1.69s / 2.16s, score 0.26s / 0.40s. The scan is half the pass on the tablet and under a third on the desktop, because twenty cores of AVX2 pull ahead of NEON much further than the merge engine's single-threaded hashing does. So a GPU GEMM is worth roughly 2× a regroup on the tablet and 1.5× here, and it is the tablet that should decide whether it is built. The two architectures agree exactly: the same 1,531,969 evidence pairs, the same 2,518 groups holding the same 16,246 faces, the same reliability table. That is a better check on the NEON kernel than the unit test can be. Two instruments, both read-only: the example now prints its phases, and dr-face gains scan_bench, which needs no library at all and so can answer "how fast is this machine" on a device with nothing on it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ebb7d3cf5c |
Score a suggestion against the people the user has named
The number beside a suggestion was the mean calibrated probability between the face and the rest of its group, which measures the wrong thing twice. It punishes coverage: a person with two hundred faces over fifteen years is *meant* to have members a given photograph is orthogonal to, so a correct suggestion onto a well-photographed person scored low for being well photographed. And it never asked who else the face might be — a face matching Anna at 0.95 and nobody else, and one matching Anna at 0.95 and her sister at 0.93, came out identical, when the second is the only one worth the user's attention. dr_face::assign answers both, and multiplies them: the mean of the best ten calibrated matches into the identity (the old mean, capped, which is what stops coverage counting against it), times that identity's share of the evidence against every *named* rival. Only named people compete, and per person rather than per group. Both halves of that had to be measured on a real 18,000-face library rather than reasoned about. Normalising across every group made the number useless — median suggestion 21%, four in five under half — because clustering leaves one person spread over many groups, so a face competed against itself; and keying rivals by group left Catherine competing with Catherine, median 39%. Per named person: median 99.5%. Rivals are gathered below the merge threshold, down to even odds: a named person matching at 0.6 will never be merged into but is exactly the competition to discount for. That would be a second similarity scan, the expensive half of regrouping a library, so cluster_scored scans once at the looser floor and hands the merge engine the subset at or above the threshold — pair for pair what it would have scanned for itself, held to that by a test. Leave-one-out over that library's 2,702 confirmations across 54 named people: 99.33% of faces placed on the right person against the old mean's 99.15%, and the number shown for the right person moves from a median of 90.4% to 99.3%. It errs low — 100% correct wherever it states 80% or more — which is the safe direction, and docs/faces.md §9.1 says plainly that the low bands are not calibrated. The example that measures it comes too: this is a claim about a library's numbers, and nobody should have to take it on faith. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5768100816 |
Borrow the library to index it, and give it back
The passes that need every photograph's bytes — thumbnails, face indexing — now borrow each one and release it at the end. On a placeholder library that is the difference between peak disk being the working set and being the whole library. Including on cancellation, which was nearly missed: the face sweep returns mid-loop when the user presses Stop, and without releasing there the disk is spent and nothing is delivered for it. `materialise` now answers whether *it* fetched the content. The pool used to work that out by listing a file's parent directory — one listing per file across a library — when the backend already had to `stat` it to decide whether to ask. One syscall instead of a directory walk, and it removes the bug class the tests found earlier: a file at the library root has no `parent()`, so every one of them read as already-downloaded. **Pinning is the retention control**, and it drives the model the catalog already had rather than a second one. `tier_desired` is what the user asked to keep hydrated, `pending_pins` is the resumable work list, and a pinned collection is never dehydrated for the same reason it was never evicted. It was in fact *broken* here before: `get` on a stub failed, and the pin worker logged "one unreadable file must not abandon the whole pin" and silently did nothing. Pinned originals on such a library are recorded with `path = NULL` (`Cache::record_in_place`) rather than copied under `originals/`. Two reasons, and the second is the important one. A copy would hold every pinned photograph twice, with the budget able to evict the half that was not costing the disk. And `release` deletes the file a row names — so a row that names none cannot delete anything, which puts the one catastrophic operation out of reach by construction rather than by remembering not to call it. Deleting a materialised file inside a synced tree removes the photograph from the server and every other device. Handing disk back is `spawn_dehydrate`, which asks the client. Two gaps written down rather than papered over (docs/storage.md §7): a hydrating pass cannot yet quote its cost, because a stub reports no size; and the two sweeps hold separate pools, so a library indexed for both fetches twice. |
||
|
|
f41b3f6e8e |
Answer "what would the other device end up with" without the other device
Build and test / Desktop (Linux) (push) Successful in 2h7m52s
Build and test / Layer separation (push) Successful in 59s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 4s
Traceability / Requirement traces (push) Successful in 37s
Build and test / Android (aarch64) (push) Successful in 22m22s
A tablet showed 200 of a person's 611 faces after syncing, and the obvious
suspects — that suggestions deliberately do not travel, that the cross-device
face match was too strict — were both wrong. Finding that out meant reading a
catalog on a release-signed Android build, which cannot be done.
So this stands the second device up locally: an empty catalog, given the images
a scan would have found, the shards adopted into it exactly as a sync does, and
the real catalog merged in as the remote. Then it counts, per person, against
what the source holds.
person source here
Catherine 611 611
Me 242 242
Ian 219 219
Which settled it: the merge carries everything, and the shortfall was transfer —
shards that never finished arriving. Worth keeping, because "did the sync lose
this or has it not got here yet" is a question that will come up again, and
guessing at it cost most of an evening.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
2c84aa1224 |
Write face shards the way the rest of the catalog writes
The face store was the one part of the catalog still on SQLite's default rollback journal at `synchronous = FULL`. The catalog itself runs WAL at `NORMAL` (`schema::configure`) and so does the thumbnail store; nothing decided this one should differ, it was simply never set. Measured on this project's own filesystem, that is **21.3 ms per commit against 0.05 ms** — four hundred times. And an export commits four times per photograph: the shard's transaction, then three separate autocommitting writes to the index. Ten thousand images is on the order of fourteen minutes spent doing nothing but waiting for fsync, before a byte goes to the server. That is the "checking faces…" that appeared to hang. So: WAL and `synchronous = NORMAL`, matching the rest, and the three index writes fold into one transaction. `NORMAL` is the same trade the catalog makes — a shard is derived data, and losing the last commit to a power cut costs one image re-exported. WAL brings an obligation with it, because **a shard is uploaded by reading its file**: the newest commits live in a `-wal` sidecar that no upload sends, so without a checkpoint the server would receive a database missing exactly the faces just written, and a peer would adopt it and see nothing wrong. `checkpoint` folds the logs back in, with `TRUNCATE` rather than the default passive mode, which gives up when a reader holds the log and would leave the same gap while reporting success. Two tests: that the store is in WAL like everything else, and — the one that matters — that a checkpointed shard copied *without* its `-wal` still holds every face. That second one fails without the checkpoint, which is how it was confirmed to be testing something. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a1e8f494b9 |
Say what the face sync is doing while it does it
The face pass set the status to "checking faces…" once and then said nothing until it was finished. On a library whose first export after a re-index is 9,849 photographs and 85 MB of shards, that is eight minutes of a progress bar sitting still — which is indistinguishable from a hang, and was reported as one twice. Nothing was wrong with the sync. The only fault was that it was silent. Three places now report, which are the three that take real time: - **Preparing**, per image with a count, since this is the long one and the only one whose length the user cannot guess from anything on screen. - **Sending**, per shard with its size, because a face shard carries crops and runs to tens of megabytes — one of them is a visible wait on any connection. Announced before the upload rather than after, since the wait *is* the upload. - **Taking in** a peer's shard, which is a download and then a row-by-row merge. The export reports every 25 images rather than every one, so the channel behind it stays lost in the write it accompanies. `export_to_shards` keeps its old signature and delegates, so the callers that do not want progress do not grow a parameter for it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
0729dfa359 |
Stop the face sync opening a database per photograph
"Checking faces…" never finished. `export_to_shards` asks, for every indexed image in the library, whether the shard store already holds that image at that index time — and `indexed_at` answered by opening the shard database, running its six-statement schema batch and two `pragma_table_info` queries, then querying. Once per image. 9,849 times for this library, on every sync pass, before a single face had been written. The index time now lives in the store's `index.sqlite` alongside the shard number, so the question is one indexed lookup on a connection that is already open. It stays in the shard as well — that copy is the one that travels — but nothing reads it from there on the hot path. The write side had the same shape: `put_image` opened the shard afresh for each image, which mattered little when exports were a handful of new photographs and matters a great deal now that a re-index sends thousands. The handle is kept and reused, invalidated by shard id so sealing a full one and moving to the next drops it without anything having to remember to. `INDEX_SCHEMA` is `CREATE ... IF NOT EXISTS` like the shard schema, so the new column is added on open for an index already on disk — the same trap, caught the same way. Two tests: that the index time survives reopening the store, and that an index written before the column can still be opened and written to. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
76eece8500 |
Add the columns an existing shard never got
`SHARD_SCHEMA` is entirely `CREATE ... IF NOT EXISTS`, which does exactly nothing to a table that already exists. So `crop` and `indexed_at`, both added to that batch, never appeared in any shard that had been written before — and the `INSERT` naming them failed with "no such column". Which took face export down completely, on every library that had ever synced a face. Silently: `export_to_shards` returns the error, `sync_face_shards` logs it at warn, and the sync goes on looking successful while the catalog fills with faces no other device will ever see. This library's shard sat frozen at 1,807 faces with 15,194 in the catalog, and the reason was this rather than anything in the export logic. Shards are upgraded on open now: both columns are additive and nullable, so catching up is one `ALTER` each. There is deliberately no version counter — "does this column exist" is the question actually being asked, and asking it directly cannot fall out of step the way a counter can. A peer's shard is opened read-only and cannot be repaired, so one written before crops is read as it stands, with a `NULL` standing in for the column. An adopted face simply has no crop, which is the truth about it. Four tests, built against the pre-crop schema written out in full rather than derived from the current one — the point being that it is *not* the current schema and must not track it. Verified against the real 1,807-face shard on this machine: the ALTERs apply, writes succeed, and nothing already in it is lost. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4ed10f7b23 |
Let a person cross from one device to another
The face shards carry boxes, landmarks and embeddings. What they deliberately do not carry is who anybody **is** — the person rows, their names, and the assignments joining the two. Those travel in the catalog snapshot, which is a whole-file copy and does contain them. But the snapshot is *merged*, not adopted, and this merge only ever looked at collections and keywords. `face_shard`'s own module note says people travel in the snapshot; nothing implemented it. So a second device received every face and no people at all, and drew an empty People screen over a full catalog. Exactly what a tablet showed after syncing thousands of faces from a laptop. What travels is what the user decided, following the rule the rest of this module already follows — judgements travel, inference is rebuilt: - **People**, by uuid on `revision`, exactly as a collection is: the name, and whether the group was set aside. - **Confirmations**, and **rejections** — "this is not her" is a fact too, and is why re-clustering does not put it back. - **The suggestions inside an ignored group**, which are otherwise ordinary inference but are what anchors the ignore. Without them a group set aside on one device reappears on the other, the same fault that made "Not interested" not stick locally. Ordinary suggestions are not carried. Both devices hold the same embeddings and clustering is deterministic, so each recomputes them and arrives at the same answer; shipping them would double the merge for no new information. **A face has no cross-device identity**, and unlike a collection there is no uuid to give it one. Both devices do agree on `oc:fileid` and roughly on the box, so a remote face is matched to the local face on the same photograph whose box overlaps it most, above 0.5 IoU. That is not a new rule — it is the one `record_detections` already uses to carry a confirmation across a re-index, and it is loose on purpose: the question is "the same face in the frame", not "the same rectangle". A local confirmation is never overwritten. Two devices confirming one face as different people is a real disagreement and an assignment carries no revision to settle it with; taking the remote's answer would let a sync undo what the user just did on the device in their hands. The remote's schema is probed rather than assumed: `remote_is_mergeable` admits any catalog at or below this version, so one written before faces existed, or before V10 added `ignored`, is ordinary. An absent table skips this half instead of aborting a merge that would otherwise have succeeded. Nine tests, including that the name lands on the overlapping face and not its neighbour in the same frame, that a set-aside group stays set aside, that an ordinary suggestion does not travel, and that merging twice changes nothing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
cf614efa61 |
Send a re-indexed image's faces to the other devices
`export_to_shards` asked `store.contains(file_id)` and skipped anything the shard store had already heard of. So an image was exported exactly once, and re-indexing it updated the catalog and nothing else — every other device kept the first answer for ever. That is not hypothetical. This library was re-indexed after the detection floors changed and crops were added, going from 1,807 faces to 15,194; the shard store still held the original 1,807, written before any of it. Nothing the re-index produced could reach another device. The shard's `indexed` table now carries the catalog's own `indexed_at`, and the export compares against it. A re-indexed image goes again; an unchanged one still costs nothing. Copied from the catalog rather than stamped when the shard is written, because a shard-local write time advances even when nothing changed and could not answer the question. The column is nullable so a shard written before it still reads: absent means "cannot vouch for it", which forces one re-export and then settles. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
79c0520506 |
Keep the face, not just a way to find it again
A face was drawn by decoding the 1024px proxy it was found on and cutting the box out again, every time the People screen opened. That made the screen a derivative of the thumbnail cache: evict a proxy — which the cache may do at any moment — and the cell goes blank, with no way back short of re-fetching the original over the network and re-detecting it. It also cost a full JPEG decode per image, per visit, to show a 96px cell. So the crop is cut once, when the pixels are already in hand at detection time, and kept. A 160px JPEG is a few KB against the ~250 KB proxy it replaces reading. Where it lives is the interesting part. The catalog snapshot is uploaded *whole* on every sync and downloaded by every device, so a crop column there would put tens of MB on every round trip — the exact cost `face_shard`'s 25 MB cap exists to bound, and the reason bulk per-face data lives in shards already. Crops therefore travel in the face shards, beside the embeddings, and `snapshot_for_upload` strips them from the copy it writes. Nothing reads a crop out of a merged remote catalog — the merge touches collections and keywords only — so a receiving device loses nothing. A shard carrying crops holds around 3,500 faces rather than 22,000, which is the price of a second device showing People immediately instead of re-fetching every proxy. The column is nullable and the reader falls back to the proxy, so a face indexed before this still works and the next indexing pass fills it in. V10 also adds `people.ignored`, for a person the user has looked at and does not want to identify. Most clusters in a real library are strangers — passers-by, other people's guests, a face on a poster — and there is no way to tell "not yet looked at" from "looked at, don't care" without recording the second. It is a column rather than a deletion because a deleted cluster comes straight back on the next Regroup: the faces are still there and still similar, and nothing short of remembering the judgement survives re-clustering. Same argument `face_person_rejected` makes one level down. And `prune_empty_unnamed`, for what clustering leaves behind. Regroup creates a person per unanchored group and never removed the previous run's now-empty ones, so pressing it twice added a rail entry per group it no longer believed in. Named people are never touched however empty — a name is user data — nor is a merge tombstone, which must outlive its faces to keep redirecting. 298 tests pass, including that the snapshot carries no crops while the live catalog keeps them. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b846b312b8 |
Run the formatter over the face branch before it reaches CI
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 2s
Build and test / Desktop (Linux) (push) Successful in 1h21m32s
Build and test / Layer separation (push) Successful in 37s
Traceability / Requirement traces (push) Successful in 25s
Build and test / Android (aarch64) (push) Failing after 33m58s
The merge of the SCRFD/MobileFaceNet work brought 69 rustfmt diffs across
dr-catalog, dr-face and dr-ui with it, so `cargo fmt --all -- --check` fails
on master and the Desktop job stops at its Format step — before clippy, the
tests or the release build have run at all. That makes the whole desktop
half of CI blind: a real compile error behind this would look exactly the
same from the outside. There was nothing behind it, as it turns out — with
the formatting fixed, clippy, the test suite and the release build all pass.
Every .rs hunk is `cargo fmt --all` on the pinned 1.92.0 toolchain, not a
hand edit, but it is worth being precise about what that moved, because it
is more than whitespace. Besides reflowing signatures and call chains,
rustfmt reordered the `pub mod` and `pub use` items in dr-face/src/lib.rs so
the `#[cfg(feature = "inference")]` entries sort in place, added the trailing
semicolon inside `let ... else { return }` bodies in identity_ui.rs, wrapped
a bare closure body in braces in cluster.rs, adjusted trailing commas, and
dropped a stray blank line at the end of identity_ui.rs. All of it is
semantically inert; none of it changes behaviour.
docs/traceability.md rides along because it has to. The matrix records each
TRACES tag by line number, and reflowing develop.rs, lib.rs, faces.rs,
identity.rs and identity_ui.rs moved them — FR-CAT-8, FR-CAT-9, FR-CULL-10,
FR-DEV-3, FR-DEV-3a and FR-DEV-3c all shift by a line or two. The matrix was
verified up to date on
|
||
|
|
96d07da15f |
Sync face data as sealed shards, so a second device does not re-index
Indexing 23,500 images is about two hours of CPU, and the result is byte-identical on every device: the same model over the same proxy produces the same embedding. Paying for it once per account rather than once per device is the point. Shards rather than the catalog snapshot, because the snapshot goes up whole on every sync and a fully indexed library carries roughly 30 MB of embeddings. That is exactly the cost the thumbnail store's 25 MB cap exists to bound, so face shards use the same cap -- imported from dr_thumbs rather than restated, since the number is a statement about sync cost and the two must not drift apart. The split follows the one already there: bulk immutable data in sealed shards, small mutable data in the catalog snapshot. Faces, landmarks, embeddings and run markers shard; people, names and assignments ride the catalog and merge by uuid. Keyed on oc:fileid throughout, never on image_id, because a row id means nothing on another device. The run marker travels with the faces it describes. Without it a receiving device cannot tell an image with no faces from one never examined, and would re-detect every landscape it had just adopted. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
26a1eb7e28 |
Record that face detection has run, not just what it found
An image with no faces in it was indistinguishable from one that had never been looked at, so every landscape, still life and document scan in the library was re-detected on every pass, for ever. In a real library that is most of it: on the 23,527-image test library, 64 of the first 110 images indexed contain no face at all. Schema v9 adds face_index, a run marker per (image, model) carrying the face count and the proxy edge it read. Keyed on the model, so a model change puts every image back in the queue by itself. That makes a coverage figure possible, which is the thing a user actually wants to see. The audit also splits the outstanding set by whether a proxy exists, because 23,417 awaiting a proxy and 110 ready to index are different problems, and telling the user to run indexing again would not fix the first. The Identity screen gains Index faces, Stop, and the coverage line. examples/face_index.rs is the same check and sweep without a window, which is the right shape for an overnight pass. Measured on the real library in release: 3.5 images/second, 110 images and 125 faces in 30 seconds, and a second run correctly finds nothing left to do. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |