379dd1afcc2bf07fd045ec0e313290d7d3936364
18
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c78b798cf0 |
Fold people of one name whose faces agree, and faces held twice (#78)
Seven names are two or three live people on both devices: Ian (756 confirmed faces, and a second Ian with none), Jessie three times, Claudine, Mathias, Noemi, Pascal and PJ. Each was typed on its own device and carried across by sync, which keys people on their uuid and so keeps both. Each half of a person shows half their photographs. dedup_people::run, in one transaction: - Same-name people (trimmed, case-folded as the Identity screen folds them) merge into the one with the most confirmed faces, ties to the smaller uuid, through faces::merge_people_within, so confirmations, rejections and the survivor's name are kept. A person holding no faces at all merges: there is nothing to compare or to carry. Anyone else needs >= 2 confirmed faces per shared embedder on both sides and centroids at cosine >= 0.7 in each. A face confirmed as one and rejected as the other keeps them apart. Unnamed and set-aside people are never merged by name. - Faces held twice (one image, one embedder, IoU >= 0.5, cosine >= 0.7) keep the stronger detector's row (FaceDetector::outranks), then the confirmed one, then the older. The survivor takes the confirmed assignment and both rows' rejections. A pair confirmed as two different people is left and counted. - Judgements still on a merged-away person move to the person at the end of its redirects, and a redirect cycle (two devices merging one pair in opposite directions) is broken at the smaller uuid. Measured on copies of the desktop catalog and the tablet's server snapshot, w600k_mbf, confirmed faces only: - Centroids of differently named people: 2,699 pairs, median 0.02, 99.9th percentile 0.41. One pair reaches 0.70 (0.700 desktop, 0.705 tablet), "Michelle Casanonve" and "Michelle Casanova", one person typed two ways. Next is 0.62/0.64, "Boris Jost" and "Boris". The highest pair that is plainly two people is 0.43/0.44. - One person split in random halves: minimum 0.69, median 0.91 over 72 people. Four faces against twenty-two reach 0.7 in 97% of draws. One face against twenty of somebody else's reached 0.74 in 3,000 draws, and two faces reached 0.61, hence the two-face minimum. - Pascal (22 and 4 confirmed) is at 0.57 and PJ (14 and 7) at 0.50, under 0.7 on both devices, so both pairs stay apart and are logged. The desktop's second Ian holds 4 suggestions and no confirmations, at 0.38 against Ian's centroid, and stays apart. On the tablet it holds nothing and merges. Why a merge made here survives a peer on 0.17.0: the merged-away person stays as a merged_into redirect with a bumped revision, which the catalog merge has always taken on revision. The peer hides the duplicate and never sends it back as a live person. Its own confirmations of that person stay on the redirect, because a merge never overwrites a local confirmation. The manual merge has always left them there too. They follow the redirect when the peer runs this job. A test syncs two catalog files through the previous merge code and back, and the people converge and stay converged. Once a catalog is clean the job reads 80 redirects, the named people, and the face boxes from the covering faces_box index. That is ~10 ms on the reference library. There is no schema change. The index is created IF NOT EXISTS, as the merge already does. |
||
|
|
1479e45637 |
Keep one person to one face per photograph in the catalog merge
The grouping pass never puts two faces of one photograph in the same group (the cannot-link in dr_face::cluster). The merge did not check this. When the two devices disagree about which face in a frame is a person, merge_people_within applied the remote's confirmation, or the anchor of a set-aside group, to face X. This device already held the same person on face Y of the same photograph, so the person ended up on both faces. The reference library has 80 such person/photograph pairs on the desktop and 89 on the tablet: 79/88 unnamed set-aside groups and one named person confirmed on two faces. There are no duplicate faces (no pair of faces in one image and embedder with IoU >= 0.5). An incoming assignment is now refused when another local face of the same photograph already holds that person. The one exception is an incoming confirmation against a local suggestion: the suggestion is withdrawn and the confirmation is applied. Two confirmations stay as this device has them, the same rule as a local confirmation outranking a remote one. A face that already holds the person is not a rival to itself, so a steady-state pass is unaffected. On the reference pair this refuses 0 assignments and writes the same 8,954 as before; it only changes what a future disagreement does. The refusals are counted in MergeReport::faces_one_per_photograph. Existing pairs are left alone. They are two different faces (cosine 0.31 for the named one), not one face twice, so there is nothing to fuse, and which face is the wrong one is not the merge's to guess. |
||
|
|
b5b30e3750 |
Match synced faces by embedding where the boxes cannot (#77)
On the reference library, 631 of the faces in the tablet's snapshot match no desktop face by box (IoU >= 0.5), so a name on them stays on one device. Twenty of those are the same face with the box drawn somewhere else. Whole photographs sit at IoU 0-0.48 with cosines of 0.72-0.96 between the two devices' vectors, and ten of them already carry the same person on both sides. The merge never read the 1 KB embedding every face row carries. match_faces now runs two passes. The box pass is unchanged except that a pair must now be unique on both sides: a remote face with two overlapping local faces, or a local face overlapped by two remote ones, is no longer settled by whichever overlap is larger. Only for the photographs where a remote face is left over, and a local face is still free, does it read vectors: one json_each statement per side, keyed by row id. That was 357 photographs on the reference library, not all 19 MB of vectors. A left-over face pairs with the local face it resembles most when: - the cosine is >= 0.7, - the two are each other's best, - each leads its runner-up by >= 0.2, and - a box has not already claimed the local face. Anything less decisive stays unmatched, so a new face stays new. Why the threshold is safe, measured on both catalogs (w600k_mbf): - Of 169,548 pairs of different faces in one photograph, 4 reach 0.7 (lookalikes in one frame) and the maximum is 0.82. - At 0.6 the rule would claim two pairs that carry different people on the two devices. At 0.7 it claims 20, none contradicted and 10 corroborated, each leading its runner-up by more than 0.5. - It only compares faces the boxes left unmatched on both sides: 737 such pairs, so about 0.02 false pairs expected. - A low cosine never overrules a box. About 150 box-matched pairs fall below 0.45, because two detectors cut the same tiny face differently. 73 of them carry the same person on both devices. - Faces are compared only within one file_id and one embedder, because the same person in another photograph reaches cosine 1.0. - `Embedding::cosine` refuses a comparison across models. Before -> after on the reference pair: matched by box 18,348 -> 18,348, by embedding 0 -> 20, unmatched 631 -> 611, ambiguous 0 -> 0. The report counts the embedding matches. No schema change. |
||
|
|
9b580c3720 |
Satisfy rustfmt and clippy on the album and folder picker changes
rustfmt over the files the albums work touched, and the album merge's incoming row as a named struct rather than an eight-field tuple, which clippy's type_complexity refused. |
||
|
|
94542371f6 |
Keep albums in the catalog: export folders and what went into them
An album is a named export destination. Its folder holds only the exported files; the catalog records, per file, the image it was rendered from, so an album can show the originals behind its JPEGs (FR-EXP-10). The tables are created on first use (CREATE TABLE IF NOT EXISTS), the way dedup_probes is, rather than by a schema migration: a new user_version makes every older build refuse this catalog's snapshot at sync, and the 0.16.0 tablet would stop merging collections, keywords and people for a feature it does not have. Albums merge as collections do: by uuid and revision, tombstones on delete, exports as a set union keyed on the server's file id (content hash for a folder library). A folder on the server lives on the album row and syncs; a folder on this device lives in album_folders, which the merge never reads and the upload snapshot drops, because a path or a SAF grant on one device means nothing on another. Exports are keyed on the file name, not the image: two crops of one photograph are two files and two rows, and an overwrite re-points the name at whatever wrote it last. |
||
|
|
ce6705be89 |
Merge synced face assignments against what is held, read once per pass
After
|
||
|
|
ae0281fedd |
Match synced faces from an index of their boxes, not from their rows
`merge::match_faces` reads every local face's box and model to pair the other device's faces with ours. It took 54 ms of a steady-state merge on the reference library (19,000 faces). A `faces` row is eight kilobytes -- the embedding, the crop, the dense landmarks -- and `model_id` sits past the embedding, so reading it opened each row's overflow pages: SCAN f SEARCH r USING INTEGER PRIMARY KEY (rowid=?) `faces_box (image_id, model_id, x, y, w, h)` holds every column the scan asks for: SCAN f USING COVERING INDEX faces_box SEARCH r USING INTEGER PRIMARY KEY (rowid=?) The local scan went from 38 ms to 8 ms (sqlite3 on a copy, aggregated so output formatting is not timed), and `match_faces` from 54 ms to 30-37 ms; what remains is the other device's half. That is read from its snapshot, which has whatever indexes its build made -- this one will carry `faces_box` in its uploads -- and whose rows have had their crops stripped. The bench merges a full copy with crops, so it overstates that half. Created on first use in `match_faces`, with CREATE INDEX IF NOT EXISTS, rather than by a migration, for the reason `keywords::ensure_term_index` gives: a schema version bump makes older builds refuse the snapshot, and an extra index is invisible to them. The first merge after the upgrade builds it (about a second, once). Its prefix duplicates `faces_image_model`, which is left alone; the planner takes either for an (image_id, model_id) probe. Tables checksum the same after the bench run as after the old build's. |
||
|
|
981022ab1d |
Ask a synced keyword's tombstone once per merge, not once per assignment
`merge_remote_catalog` on the reference library (catalog_bench, a copy
merged with itself: the steady state of a sync pass) cost 228-231 ms best
of 5. Timing its phases put 91 ms in the keyword half, not in the faces the
issue named.
Both assignment unions refuse a word this device holds only as a tombstone,
with a correlated `NOT EXISTS (... deleted = 1) OR EXISTS (... deleted = 0)`
per incoming assignment. The `deleted = 1` half has no index to use --
`keyword_terms_name` is partial on `deleted = 0` -- so it scanned the whole
vocabulary for each of the 10,800 rows:
SCAN rk
CORRELATED SCALAR SUBQUERY 1
SCAN t
CORRELATED SCALAR SUBQUERY 2
SEARCH t USING COVERING INDEX keyword_terms_name (name=?)
The refused words are one set for the whole statement, so it is asked once:
`rk.keyword NOT IN (tombstoned names EXCEPT live names)`, which is the same
condition -- refused exactly when deleted under some identity and live under
none -- and which SQLite builds as a list before the walk:
SCAN rk
LIST SUBQUERY 2
MERGE (EXCEPT) ...
The file-id union alone went from 72 ms to 11 ms (sqlite3 on a copy), and
the keyword phase of the merge from 91 ms to 28-35 ms. Every table of the
catalog checksums the same after the bench as after the old build's run,
and the merge tests for tombstones and renames pass unchanged.
|
||
|
|
9cff677392 |
Merge a synced catalog's face assignments without re-preparing per face
A sync pass that brought nothing new cost 450-540 ms of CPU in `merge_remote_catalog` on the reference library (24k images, 19k faces), measured by catalog_bench merging a copy of the catalog with itself. Most of it was the loop over the other device's confirmed faces and the faces under its ignored groups -- 13,000 rows. For each one it prepared three statements from scratch (`query_row`/`execute` with a SQL string compile the statement every call) and then rewrote the `face_person` row with the values it already held, dirtying a page per face on every pass. The rejection loop prepared three more per row. The statements are now `prepare_cached`, the local assignment is read once per face (whether it is confirmed, and what it holds, come from the same row), and the upsert is skipped when the row already says exactly that. `faces_assigned` is still counted for those rows, so the report is the one the old code gave, and nothing else reads the difference: the row is byte-for-byte what the upsert would have written. After: 279 ms (best of 5, CPU), with every catalog table identical after the run to the old build's. |
||
|
|
af162dd010 | Merge: origin's sync ordering and adoption work, its scan-complete trigger ported into library_ui/ | ||
|
|
84fade99ec |
Put the developer docs under docs/dev and index the folder for users first
docs/ had 26 developer documents flat beside the manual, and the two audiences are very differently sized: most readers want the manual and the gesture reference, a few want the register, the designs and the measurements. The manual and gestures.md stay at the top; everything for someone changing the code moves to docs/dev/, and the two documents that name their own successors — the v0.1 milestone and the UI-refinement plan — go to docs/dev/archive/ rather than being deleted, since both are still cited. docs/README.md is the index, users first. Every reference follows: code comments, Cargo manifests, the workflows, the pre-commit hook, the bench and traceability tools (which locate the repo root by docs/dev/requirements.md now), packaging, the Docker READMEs, CLAUDE.md, CONTRIBUTING.md and the README. The matrix links one level deeper and is regenerated. Links out of the moved documents into the tree gain a level; a link checker over every Markdown file finds none broken. |
||
|
|
f71d7bacc6 |
Take the server's shards and dates when the scan completes, not after the sweep
Benchmarks / CPU and I/O (per commit) (push) Successful in 8m28s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 36s
Build and test / Layer separation (push) Successful in 28s
Traceability / Requirement traces (push) Failing after 39s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 2s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Successful in 43m11s
Build and test / Windows (x86_64, cross) (push) Failing after 41m4s
The derived sync fired only after the metadata sweep, so a fresh device re-derived every thumbnail it scrolled past, re-detected faces and re-read every header for hours before adopting the shards and snapshot that held all of it. It now fires as soon as the scan completes — the first moment the rows the merges key on exist — and the sweep starts behind it. In steady state that pass is one listing. The catalog merge gains a fourth half: capture metadata (captured_at, offset, camera, lens, ISO) for images still at metadata_state < 2, matched by oc:fileid from a remote row at 2. A date is a fact about the file's bytes, not local state, and the snapshot already carried it. The sweep's per-chunk query then finds nothing left, and the timeline is whole on a fresh device without a header fetch. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
327decfab1 |
Fuse every detector's faces into one population per embedder
Choosing "Thorough" made the library look empty. The detector setting writes under its own faces.model_id, and every reader of "the faces" keyed on that exact id: the clustering pass, the coverage figure, the sweep's work list, the shard export and import, and the sync merge's face matching. On the reference library that restarted coverage at 1,834 of 19,140, drew a People rail of 36 faces for a person with 520, queued a ~400 GB re-fetch on each device, and stranded the desktop's 3,583 confirmations under the old id: the tablet held the same faces under the new one and the merge refused to match them. Same photograph, same box, same embedder, two ids — that is one face, not two libraries. The embedder half of the id is now the key. embedder_of and embedder_sql give it to every query; writes keep the full id, so which detector drew a box stays on record. record_detections is unchanged and is where the generations meet: an image holds one pipeline's faces at a time, and a re-detection carries confirmations across by box overlap. The merge's match_faces applies the same rule within an embedder. The calibration is keyed on the embedder too, since the similarity space did not change. Shards travel every generation, each under its own id, and a peer adopts whichever it is sent — including a stronger detector's pass over an image it indexed itself with a weaker one, which is the re-detection its own sweep would otherwise queue, already done. Never downwards: a tablet on Fast keeps the desktop's Thorough faces. The sweep gains the same tail — images a weaker detector indexed, after the ones nothing has — driven by FaceDetector::supersedes, so choosing a stronger detector still improves the library over time without first making it disappear. |
||
|
|
4ed10f7b23 |
Let a person cross from one device to another
The face shards carry boxes, landmarks and embeddings. What they deliberately do not carry is who anybody **is** — the person rows, their names, and the assignments joining the two. Those travel in the catalog snapshot, which is a whole-file copy and does contain them. But the snapshot is *merged*, not adopted, and this merge only ever looked at collections and keywords. `face_shard`'s own module note says people travel in the snapshot; nothing implemented it. So a second device received every face and no people at all, and drew an empty People screen over a full catalog. Exactly what a tablet showed after syncing thousands of faces from a laptop. What travels is what the user decided, following the rule the rest of this module already follows — judgements travel, inference is rebuilt: - **People**, by uuid on `revision`, exactly as a collection is: the name, and whether the group was set aside. - **Confirmations**, and **rejections** — "this is not her" is a fact too, and is why re-clustering does not put it back. - **The suggestions inside an ignored group**, which are otherwise ordinary inference but are what anchors the ignore. Without them a group set aside on one device reappears on the other, the same fault that made "Not interested" not stick locally. Ordinary suggestions are not carried. Both devices hold the same embeddings and clustering is deterministic, so each recomputes them and arrives at the same answer; shipping them would double the merge for no new information. **A face has no cross-device identity**, and unlike a collection there is no uuid to give it one. Both devices do agree on `oc:fileid` and roughly on the box, so a remote face is matched to the local face on the same photograph whose box overlaps it most, above 0.5 IoU. That is not a new rule — it is the one `record_detections` already uses to carry a confirmation across a re-index, and it is loose on purpose: the question is "the same face in the frame", not "the same rectangle". A local confirmation is never overwritten. Two devices confirming one face as different people is a real disagreement and an assignment carries no revision to settle it with; taking the remote's answer would let a sync undo what the user just did on the device in their hands. The remote's schema is probed rather than assumed: `remote_is_mergeable` admits any catalog at or below this version, so one written before faces existed, or before V10 added `ignored`, is ordinary. An absent table skips this half instead of aborting a merge that would otherwise have succeeded. Nine tests, including that the name lands on the overlapping face and not its neighbour in the same frame, that a set-aside group stays set aside, that an ordinary suggestion does not travel, and that merging twice changes nothing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
67f15bffb7 |
Key collection membership on the identity that exists
Collections synced their names and arrived empty on every device. The names are keyed on a uuid and worked; the membership union was keyed on `images.content_hash`, and the schema says plainly what that column is: "computed only when something needs it (import dedup, reconnect-by-hash), never in a scan". A library that has only ever been scanned has one for no image at all, so the join matched nothing and `WHERE ri.content_hash IS NOT NULL` discarded whatever survived. The union could never have moved a single row. Measured on a real catalog: 23,174 images, content hashes for 0 of them, `oc:fileid` for all 23,174, twelve collections, zero members. So membership now resolves through the file id first, exactly as keyword assignment already did — `ASSIGN_BY_FILE_ID` was added for this same reason and its doc comment even notes that membership was still on the hash. It is recorded for every image the moment a remote scan sees it, survives server-side rename and move (FR-NC-5), is the same integer on every device pointed at one Nextcloud, and is already what the thumbnail shards are keyed by. The content-hash union is kept rather than replaced: a local-only library has no `remote` rows, and where a hash has been computed it is a true identity that survives a library moving between servers. Both statements run; `INSERT OR IGNORE` against the primary key makes the overlap free. This repairs the merge. It cannot invent membership that no device recorded — where the rows were never written, collections stay empty until they are filled in again. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
62188ec740 |
Keep both devices' keywords when the catalogs meet
Keywords are catalog state, and the catalog syncs. Without this, two devices keywording the same library would resolve to whichever synced last, and an afternoon of work would vanish with no sign it had ever happened. The vocabulary merges per row on the rule collections already use: revision first, timestamp only to break a tie, so a device with a skewed clock cannot win by having the wrong idea of the time. Assignments merge as a set union, which is FR-NC-9's principle applied to metadata instead of edit nodes — disjoint work survives on both sides. Three things needed care and are commented where they happen: A deletion travels *by name*, not by identity. Both devices may have minted their own uuid for one word before they ever synced, so deleting by uuid would tombstone a row nothing was assigned to and leave every photograph still carrying the word. The union then refuses to readmit a word a winning tombstone has just removed — without that filter the remote's live assignments would resurrect it on the very same pass. Images are resolved by the server's file id first and the content hash second. Membership has always used the hash alone, but the hash is computed only when import dedup or a reconnect asks for it, which for most libraries is never — so a hash-only union would have quietly done nothing for the ordinary photograph. A word lands on the local default version. Version uuids do not reconcile in the catalog at all: ensure_default_versions mints a fresh one per device, so a uuid-keyed join would have unioned nothing. Removal still does not propagate. That is the trade collection membership already makes, for the same reason — an unwanted keyword is removed again in a second, and a silently lost afternoon is not recoverable at all. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b0206cbc7a |
Let a sub-collection stay under its parent through a sync
Build and test / Desktop (Linux) (push) Failing after 57m14s
Build and test / Layer separation (push) Successful in 33s
Traceability / Requirement traces (push) Failing after 29s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
Build and test / Android (aarch64) (push) Failing after 9m47s
The merge inserted every incoming collection with `parent_id = NULL` and never set it on update, so the hierarchy flattened on each round trip: a collection nested on one device came back from the server at the top level. `r.parent_id` was selected and then not read. The id could not be copied — row ids are local, and the remote's integer names a different collection here, or none. So carry the parent's uuid and resolve it locally, in a second pass: rows arrive in whatever order the query returns, and a child can precede its parent. Guard the resolution against cycles. Each tree is acyclic alone, but the union need not be — we may hold A above B while the remote holds B above A — and closing that loop would make every tree walk spin. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c8bb08e661 |
Add folder scan with format selection; validate A3 on a real library
Library setup as the user described it: pick a folder, choose which RAW
types to look for, scan recursively.
dr-types::FormatFilter the tick-box selection, seeing through VFS
placeholder suffixes so a dehydrated CR2 still
matches as a CR2
dr-sync::scan recursive walk, Depth:1 per directory, pruning
unchanged subtrees where the backend propagates
directory ETags
Verified against nextcloud.tourolle.paris (34.0.2) on a real library:
browse root 32 entries, 98ms
scan PhotosRaw 17,185 RAW files in 334 directories, 34.1s
(7,836 CR2 + 9,349 DNG)
range read 262KB of a 21.5MB DNG in 119ms — 1.22% of the file,
and enough to read "Canon EOS 6D | ISO 100"
That last line is assumption A3 validated on real data. Cataloguing this
library by whole-file fetch would move roughly 370GB; the range path
moves a few MB.
Pruning is capability-gated rather than assumed: with per-entry ETags a
probe costs a request and proves nothing about children, so it is skipped
entirely. A test asserts zero probes in that case.
Still unresolved: /core/preview returns 400 for every parameter
combination tried, including on a JPEG the server reports as having a
preview. Not a request-shape bug — it fails identically bare. Recorded
rather than worked around; ARCH §6.7 already treats server previews as
opportunistic, so nothing depends on it.
|