cb7ad0bbe7a7cfbb6d5f5dd033be0c875d9f410c
9
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
78df4211b0 |
Run the people deduplication after every catalog merge (#78)
A merge is where two devices' people meet: the same name typed on each, or a redirect one of them made. So dedup_people::run now follows every successful sync::merge_remote. It runs on the sync worker, never the UI thread, before the snapshot is pushed, so what it folds reaches the server on the same pass. It runs in its own transaction, and a failure is logged, not returned. What the merge took is committed and valid either way, and the next pass tries again. Once a catalog is clean it costs 8-15 ms on the reference library (19k faces, 26k people rows). On copies of the two real catalogs, a round trip with a peer running the previous merge converges on 75 listed named people on both sides and stays there over a second round. merge_remote with the job takes 75-90 ms there. |
||
|
|
ffdd640170 |
Backfill the catalog once per state, not on every open
Catalog::open ran schema::backfill every time, and every worker thread opens its own connection. A develop landing made five opens, and each paid the RAW/JPEG pairing, the default-version anti-join over every image, the uuid pass over every default version and the keyword check: 17 ms of CPU an open on a copy of the reference catalog, ~80 ms a landing, to confirm that nothing had changed since the open before. Everything the backfill repairs is a row some write added: an image a scan inserted, a version or keyword assignment a merge brought in. So the open now reads a stamp - user_version, max(id) of images and versions, max(rowid) of keywords, and the file's device and inode - and skips the backfill when the stamp matches the one recorded at this path's last backfill in this process. The maxima are each the last page of a b-tree; an open that skips costs ~1 ms. The backfill still runs: - on the first open in a process (nothing recorded yet); - on any open that migrated the schema, unconditionally; - after a pull: merge_remote forgets the path, so the next open backfills even when every incoming row collided and nothing moved; - when the file is replaced under its name: the inode is in the stamp, and recovery::set_aside, the first step of a restore and a rebuild, forgets the path; - when another process or thread adds rows, because the stamp is read from the file, not from anything this process did. The stamp is taken before the backfill, not after. Read after, it would describe the backfill's own inserts, and could record an image another connection inserted in between as covered when it was not. Read before, the worst case is one redundant pass after a backfill that did real work. Kept in memory rather than in the catalog: a stamp row would need a table an older build does not have and would travel in the sync snapshot, where a flag from another device's catalog says nothing about this one. No schema version bump, so the tablet on 0.16.0 still reads the snapshot. Tests cover the skip, a scan's new image, a migration, a pull and a replaced file. |
||
|
|
94542371f6 |
Keep albums in the catalog: export folders and what went into them
An album is a named export destination. Its folder holds only the exported files; the catalog records, per file, the image it was rendered from, so an album can show the originals behind its JPEGs (FR-EXP-10). The tables are created on first use (CREATE TABLE IF NOT EXISTS), the way dedup_probes is, rather than by a schema migration: a new user_version makes every older build refuse this catalog's snapshot at sync, and the 0.16.0 tablet would stop merging collections, keywords and people for a feature it does not have. Albums merge as collections do: by uuid and revision, tombstones on delete, exports as a set union keyed on the server's file id (content hash for a folder library). A folder on the server lives on the album row and syncs; a folder on this device lives in album_folders, which the merge never reads and the upload snapshot drops, because a path or a SAF grant on one device means nothing on another. Exports are keyed on the file name, not the image: two crops of one photograph are two files and two rows, and an overwrite re-points the name at whatever wrote it last. |
||
|
|
49b7bc2f9d |
Build the upload snapshot without the face crops instead of stripping them
Each sync pass spent 0.8-2.0 s of CPU and 1.0-5.4 s wall on the upload snapshot of the reference catalog (24k images, 18,871 faces), ahead of the rest of the pass. The upload itself had been crop-less since the crops moved to the face shards. The cost was in how it got that way. The backup API copied all 158 MB of the catalog, 96 MB of it the ~5 KB JPEG crop on every faces row. Then `UPDATE faces SET crop = NULL` rewrote 18.9k rows and freed their overflow chains, and VACUUM rebuilt the file again. That wrote the catalog about three times over to upload 50 MB. The snapshot is now built rather than copied. An empty file attaches the catalog, creates each table from the catalog's own sqlite_master and fills it with INSERT ... SELECT, with faces.crop selected as NULL. Indexes, triggers and views follow, and user_version, application_id, page size and the WAL header flag are carried over. It all runs in one transaction on the snapshot's connection, so the catalog is read as of one moment and concurrent writers are serialised, not raced, as the backup API did. The build journal is in memory with synchronous off, because the file is scratch that is rebuilt every pass and quick_check'd before upload. Foreign keys are off on that connection. The bundled SQLite enables them, and then a multi-row INSERT into images scans images for children of each new row (shadowed_by is a self-reference with no index), which cost 1.2 s alone. Measured on a .backup copy of the reference catalog with catalog_bench, old and new binaries back to back on a loaded machine: before best 1.0-5.4 s wall, 0.84-1.98 s cpu, 49.8 MB after best 0.40-2.1 s wall, 0.39-0.96 s cpu, 50.4 MB With the machine quiet the new build takes 0.31-0.43 s. What a receiving device gets is unchanged. It is the same schema, the same rows and a NULL crop, which is what 0.16.0 already uploads and merges. The merge reads only a remote face's box and model (merge::match_faces) and never writes a local crop. No device adopts a downloaded catalog as its own, and a fresh one takes faces and crops from the shards. There is no schema bump, so older builds still merge it. NFR-R2 backups keep using the backup API and keep their crops. Tests: the snapshot matches the catalog in schema, row counts, pragmas and WAL header. A leftover file is replaced. Merging a crop-less snapshot carries a confirmed name across by box and leaves the local crop untouched, and does so idempotently. |
||
|
|
f100db89ca |
Verify the catalog snapshot before it is sent, and after it lands
Two checks around the upload, both cheap next to what they prevent. Before: the snapshot is quick_checked before it leaves. It is the copy every other device merges from, and a damaged one costs each of them a download, a failed merge and a refusal to push. After: the staged upload's size on the server is compared to the bytes sent before it is rotated into place. A chunked upload is assembled server-side, and an assembly that goes wrong is a file of plausible size no device can open — caught here, on the device that caused it, for one listing; otherwise on every other device, after the fact. A mismatch, or a size the server will not confirm, discards the upload and leaves the current copy and its generations untouched. |
||
|
|
8eeb9ba0f6 |
Offer the backup, and then the rebuild, when the index turns out to be damaged
NFR-R6 asks for an integrity check at startup and two offers behind it, and none of it existed. `PRAGMA integrity_check` appeared nowhere in the tree, `Catalog::open` was `open` → `configure` → `migrate` → `backfill` and nothing else, and corruption therefore surfaced as whatever rusqlite error the first unlucky query happened to produce — "database disk image is malformed" attached to a thumbnail refresh, elided into a 34px banner, over an empty grid saying "No images found · Check the library folder". Two messages that disagreed, and no way forward but deleting catalog.sqlite by hand. The property that makes the second offer real was already here and load- bearing: the catalog is an index, not a source of truth, rebuildable from sources plus sidecars (invariant §5.2.4, cited by schema.rs, trash.rs and lib.rs). And sync.rs already knew how to take a coherent snapshot of a WAL database. What was missing was the check, the type, and the conversation. Four pieces: **The type.** `CatalogError::Corrupt`, and — the part that makes it worth having — a hand-written `From<rusqlite::Error>` that classifies rather than wraps. `SQLITE_CORRUPT` and `SQLITE_NOTADB` become `Corrupt` wherever they arise, so a background job that trips over the damage first reports the same thing the startup check would have. `SQLITE_IOERR` and `SQLITE_BUSY` deliberately do not: a dropped network mount is a different problem, and telling someone to rebuild their index would be a wrong answer delivered confidently. **The check.** `Catalog::open_verified`, `quick_check` before the open rather than after, because opening runs migrations and a damaged catalog with an intact header would otherwise have structure rewritten on top of structure that is already wrong. Bound to `open_verified` and not to `open`: the check reads every page, which is affordable once at startup where a user can answer a question, and not affordable on the dozens of opens a session's background tasks make. **The backup.** NFR-R2's second clause, taken between `configure` and `migrate` in `Catalog::open`. A migration is the one routine operation that rewrites table structure, so it is the likeliest way this file becomes unreadable, and it is the last moment the pre-migration state exists to be copied. Three generations, through SQLite's backup API after a TRUNCATE checkpoint — never `fs::copy`, which on a WAL database backs up a state older than the catalog and possibly torn. A failure to take the copy is logged, not raised: a full disk must not be what makes a library unopenable. **The conversation.** The first line of the dialogue is that the photographs and the edits are safe, before the diagnosis, because that is the question the user is actually asking. Then the two offers, which are *not* interchangeable and are not presented as if they were: a restore keeps collections, and a rebuild cannot, because a manual collection is a set of images assembled by hand and nothing in the filesystem records it (docs/catalog.md §8.1). The labels say so, and the rebuild does not take the affirmative styling while a restore is on the table. One thing that is a fix rather than a feature: `show_catalog_now` now gates the scan. `Catalog::open` succeeds on a file whose header survived, so the scan that used to start immediately afterwards would write folder ETags and image rows into damaged pages in the seconds while the user was still reading the question — turning a file that had a backup into one where the backup is the only copy left. Restore also deletes the damaged catalog's `-wal` and `-shm`. That step is easy to leave out and fatal to leave out: a journal belonging to the old file, sitting beside the new one under the same name, is replayed into it on the next open. That is not a restore, it is a fresh corruption with the evidence gone. Tested by corrupting a fixture catalog — 500 images and a collection, then every page past the second overwritten — and driving both branches. The restore is asserted on the collection, because a collection is precisely what distinguishes the two paths; the rebuild on the damaged file being kept and the next open producing an empty catalog at the current schema. Plus the `SQLITE_NOTADB` presentation, a damaged backup being refused rather than installed, and a v1 catalog whose pre-migration backup comes back reading v1 rather than v11. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
79c0520506 |
Keep the face, not just a way to find it again
A face was drawn by decoding the 1024px proxy it was found on and cutting the box out again, every time the People screen opened. That made the screen a derivative of the thumbnail cache: evict a proxy — which the cache may do at any moment — and the cell goes blank, with no way back short of re-fetching the original over the network and re-detecting it. It also cost a full JPEG decode per image, per visit, to show a 96px cell. So the crop is cut once, when the pixels are already in hand at detection time, and kept. A 160px JPEG is a few KB against the ~250 KB proxy it replaces reading. Where it lives is the interesting part. The catalog snapshot is uploaded *whole* on every sync and downloaded by every device, so a crop column there would put tens of MB on every round trip — the exact cost `face_shard`'s 25 MB cap exists to bound, and the reason bulk per-face data lives in shards already. Crops therefore travel in the face shards, beside the embeddings, and `snapshot_for_upload` strips them from the copy it writes. Nothing reads a crop out of a merged remote catalog — the merge touches collections and keywords only — so a receiving device loses nothing. A shard carrying crops holds around 3,500 faces rather than 22,000, which is the price of a second device showing People immediately instead of re-fetching every proxy. The column is nullable and the reader falls back to the proxy, so a face indexed before this still works and the next indexing pass fills it in. V10 also adds `people.ignored`, for a person the user has looked at and does not want to identify. Most clusters in a real library are strangers — passers-by, other people's guests, a face on a poster — and there is no way to tell "not yet looked at" from "looked at, don't care" without recording the second. It is a column rather than a deletion because a deleted cluster comes straight back on the next Regroup: the faces are still there and still similar, and nothing short of remembering the judgement survives re-clustering. Same argument `face_person_rejected` makes one level down. And `prune_empty_unnamed`, for what clustering leaves behind. Regroup creates a person per unanchored group and never removed the previous run's now-empty ones, so pressing it twice added a rail entry per group it no longer believed in. Named people are never touched however empty — a name is user data — nor is a merge tombstone, which must outlive its faces to keep redirecting. 298 tests pass, including that the snapshot carries no crops while the live catalog keeps them. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
62188ec740 |
Keep both devices' keywords when the catalogs meet
Keywords are catalog state, and the catalog syncs. Without this, two devices keywording the same library would resolve to whichever synced last, and an afternoon of work would vanish with no sign it had ever happened. The vocabulary merges per row on the rule collections already use: revision first, timestamp only to break a tie, so a device with a skewed clock cannot win by having the wrong idea of the time. Assignments merge as a set union, which is FR-NC-9's principle applied to metadata instead of edit nodes — disjoint work survives on both sides. Three things needed care and are commented where they happen: A deletion travels *by name*, not by identity. Both devices may have minted their own uuid for one word before they ever synced, so deleting by uuid would tombstone a row nothing was assigned to and leave every photograph still carrying the word. The union then refuses to readmit a word a winning tombstone has just removed — without that filter the remote's live assignments would resurrect it on the very same pass. Images are resolved by the server's file id first and the content hash second. Membership has always used the hash alone, but the hash is computed only when import dedup or a reconnect asks for it, which for most libraries is never — so a hash-only union would have quietly done nothing for the ordinary photograph. A word lands on the local default version. Version uuids do not reconcile in the catalog at all: ensure_default_versions mints a fresh one per device, so a uuid-keyed join would have unioned nothing. Removal still does not propagate. That is the trade collection membership already makes, for the same reason — an unwanted keyword is removed again in a second, and a silently lost afternoon is not recoverable at all. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c8bb08e661 |
Add folder scan with format selection; validate A3 on a real library
Library setup as the user described it: pick a folder, choose which RAW
types to look for, scan recursively.
dr-types::FormatFilter the tick-box selection, seeing through VFS
placeholder suffixes so a dehydrated CR2 still
matches as a CR2
dr-sync::scan recursive walk, Depth:1 per directory, pruning
unchanged subtrees where the backend propagates
directory ETags
Verified against nextcloud.tourolle.paris (34.0.2) on a real library:
browse root 32 entries, 98ms
scan PhotosRaw 17,185 RAW files in 334 directories, 34.1s
(7,836 CR2 + 9,349 DNG)
range read 262KB of a 21.5MB DNG in 119ms — 1.22% of the file,
and enough to read "Canon EOS 6D | ISO 100"
That last line is assumption A3 validated on real data. Cataloguing this
library by whole-file fetch would move roughly 370GB; the range path
moves a few MB.
Pruning is capability-gated rather than assumed: with per-entry ETags a
probe costs a request and proves nothing about children, so it is skipped
entirely. A test asserts zero probes in that case.
Still unresolved: /core/preview returns 400 for every parameter
combination tried, including on a JPEG the server reports as having a
preview. Not a request-shape bug — it fails identically bare. Recorded
rather than worked around; ARCH §6.7 already treats server previews as
opportunistic, so nothing depends on it.
|