6f33517b35555d0134b5f191aecd2e0d14098448
24
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1d7115437b |
Date a photograph from its name when its header has none
WhatsApp strips every EXIF tag and names the file "WhatsApp Image 2023-06-15 at 07.00.42.jpeg"; Windows Phone, Android cameras and darktable's import put the date in the name too. Those images sorted after everything else and were absent from the timeline. name_dates reads a date (and a time, when one follows) from the file name, then from the innermost folder that states one. A sequence number after a date is not read as a time, and a bare year folder is not a date. EXIF always wins: only examined rows still undated are filled. The sweep and the metadata repair fill as they mark an image examined, and the open backfill fills catalogs examined by earlier builds. On the reference library that takes 274 undated images to 10; the no-op case is a seek on images_captured, 0.6 ms an open. |
||
|
|
c78b798cf0 |
Fold people of one name whose faces agree, and faces held twice (#78)
Seven names are two or three live people on both devices: Ian (756 confirmed faces, and a second Ian with none), Jessie three times, Claudine, Mathias, Noemi, Pascal and PJ. Each was typed on its own device and carried across by sync, which keys people on their uuid and so keeps both. Each half of a person shows half their photographs. dedup_people::run, in one transaction: - Same-name people (trimmed, case-folded as the Identity screen folds them) merge into the one with the most confirmed faces, ties to the smaller uuid, through faces::merge_people_within, so confirmations, rejections and the survivor's name are kept. A person holding no faces at all merges: there is nothing to compare or to carry. Anyone else needs >= 2 confirmed faces per shared embedder on both sides and centroids at cosine >= 0.7 in each. A face confirmed as one and rejected as the other keeps them apart. Unnamed and set-aside people are never merged by name. - Faces held twice (one image, one embedder, IoU >= 0.5, cosine >= 0.7) keep the stronger detector's row (FaceDetector::outranks), then the confirmed one, then the older. The survivor takes the confirmed assignment and both rows' rejections. A pair confirmed as two different people is left and counted. - Judgements still on a merged-away person move to the person at the end of its redirects, and a redirect cycle (two devices merging one pair in opposite directions) is broken at the smaller uuid. Measured on copies of the desktop catalog and the tablet's server snapshot, w600k_mbf, confirmed faces only: - Centroids of differently named people: 2,699 pairs, median 0.02, 99.9th percentile 0.41. One pair reaches 0.70 (0.700 desktop, 0.705 tablet), "Michelle Casanonve" and "Michelle Casanova", one person typed two ways. Next is 0.62/0.64, "Boris Jost" and "Boris". The highest pair that is plainly two people is 0.43/0.44. - One person split in random halves: minimum 0.69, median 0.91 over 72 people. Four faces against twenty-two reach 0.7 in 97% of draws. One face against twenty of somebody else's reached 0.74 in 3,000 draws, and two faces reached 0.61, hence the two-face minimum. - Pascal (22 and 4 confirmed) is at 0.57 and PJ (14 and 7) at 0.50, under 0.7 on both devices, so both pairs stay apart and are logged. The desktop's second Ian holds 4 suggestions and no confirmations, at 0.38 against Ian's centroid, and stays apart. On the tablet it holds nothing and merges. Why a merge made here survives a peer on 0.17.0: the merged-away person stays as a merged_into redirect with a bumped revision, which the catalog merge has always taken on revision. The peer hides the duplicate and never sends it back as a live person. Its own confirmations of that person stay on the redirect, because a merge never overwrites a local confirmation. The manual merge has always left them there too. They follow the redirect when the peer runs this job. A test syncs two catalog files through the previous merge code and back, and the people converge and stay converged. Once a catalog is clean the job reads 80 redirects, the named people, and the face boxes from the covering faces_box index. That is ~10 ms on the reference library. There is no schema change. The index is created IF NOT EXISTS, as the merge already does. |
||
|
|
ffdd640170 |
Backfill the catalog once per state, not on every open
Catalog::open ran schema::backfill every time, and every worker thread opens its own connection. A develop landing made five opens, and each paid the RAW/JPEG pairing, the default-version anti-join over every image, the uuid pass over every default version and the keyword check: 17 ms of CPU an open on a copy of the reference catalog, ~80 ms a landing, to confirm that nothing had changed since the open before. Everything the backfill repairs is a row some write added: an image a scan inserted, a version or keyword assignment a merge brought in. So the open now reads a stamp - user_version, max(id) of images and versions, max(rowid) of keywords, and the file's device and inode - and skips the backfill when the stamp matches the one recorded at this path's last backfill in this process. The maxima are each the last page of a b-tree; an open that skips costs ~1 ms. The backfill still runs: - on the first open in a process (nothing recorded yet); - on any open that migrated the schema, unconditionally; - after a pull: merge_remote forgets the path, so the next open backfills even when every incoming row collided and nothing moved; - when the file is replaced under its name: the inode is in the stamp, and recovery::set_aside, the first step of a restore and a rebuild, forgets the path; - when another process or thread adds rows, because the stamp is read from the file, not from anything this process did. The stamp is taken before the backfill, not after. Read after, it would describe the backfill's own inserts, and could record an image another connection inserted in between as covered when it was not. Read before, the worst case is one redundant pass after a backfill that did real work. Kept in memory rather than in the catalog: a stamp row would need a table an older build does not have and would travel in the sync snapshot, where a flag from another device's catalog says nothing about this one. No schema version bump, so the tablet on 0.16.0 still reads the snapshot. Tests cover the skip, a scan's new image, a migration, a pull and a replaced file. |
||
|
|
94542371f6 |
Keep albums in the catalog: export folders and what went into them
An album is a named export destination. Its folder holds only the exported files; the catalog records, per file, the image it was rendered from, so an album can show the originals behind its JPEGs (FR-EXP-10). The tables are created on first use (CREATE TABLE IF NOT EXISTS), the way dedup_probes is, rather than by a schema migration: a new user_version makes every older build refuse this catalog's snapshot at sync, and the 0.16.0 tablet would stop merging collections, keywords and people for a feature it does not have. Albums merge as collections do: by uuid and revision, tombstones on delete, exports as a set union keyed on the server's file id (content hash for a folder library). A folder on the server lives on the album row and syncs; a folder on this device lives in album_folders, which the merge never reads and the upload snapshot drops, because a path or a SAF grant on one device means nothing on another. Exports are keyed on the file name, not the image: two crops of one photograph are two files and two rows, and an overwrite re-points the name at whatever wrote it last. |
||
|
|
3c2eacbf3f |
Find catalog duplicates and fold a group onto one copy in one transaction
The library holds the same RAW in several folders: a dated folder, a bck/ beside it, a renamed Darktable export tree. dr_catalog::duplicates is the catalog half of consolidating them (#67). candidates() is one grouped query over root, camera, capture instant and size, joined back for the rows; count() is the same grouping under COUNT. On a copy of the reference catalog (23,582 images) both take 10-35 ms and find 1,836 groups holding 3,379 spare copies. survivor() prefers a copy outside a backup-looking folder, then one still named the way the camera named it, then the oldest, then the lowest id. consolidate() re-checks the plan against the catalog, merges the copies' judgements onto the survivor (highest rating, keywords unioned, collections unioned with the survivor keeping its place, a flag or label the copies agree on, faces via faces::carry_onto_copy) and records the copies as trashed, all in one transaction, so a failure part way leaves the group untouched. preview() runs the same code and rolls it back. Sameness probes are kept in dedup_probes, created on first use rather than by a migration: a schema bump would make older builds refuse this catalog's snapshot at sync. trash::record_trashed_within lets the trash write share the merge's transaction. |
||
|
|
5c00942b84 |
One completeness job over a registry of repairs, and a re-index button
A library's records are never all complete at once. A face found before its quality was kept has no quality; one found before the eye models existed has no reading; one adopted from a peer's shard has no crop; an image the fast detector examined on a 1024 px proxy has boxes the current detector would not have drawn; an image the scan stat'ed has no capture date. On the reference library that is 17,762 faces under the bare w600k_mbf id with no quality, no reading and no dense landmarks, 4,144 of them without a crop, beside 12,217 images the fast detector examined and found nothing in. Every one of those gaps was its own pass — V14's measuring pass, §17.5's eye pass, the sweep's proxy repair, the sweep's detector upgrade — with its own work list, its own count and its own idea of done, and adding a per-face field meant adding a pass. There was no pass at all for the case the library is actually in: boxes and landmarks drawn by a weaker detector on a proxy, which every later per-face pass would have read from. dr_ui::repairs replaces them with one job over a registry. A Repair names one thing a record can lack — the predicate that says which images still owe it, the input its handler needs (a header, the original, or a native render), the handler, and what to record for an image that can never be done. The job unions the predicates into one work list, fetches each image once at the most any claimant asks for, renders it at most once, and runs every handler whose predicate that image still matches, checked again before each because a detection writes every field a per-face handler would fill. The registry today: face-proxy, face-quality, face-eyes, face-crop, face-detection, face-upgrade, metadata — the last there to say that this is not a face job. Adding a field is one entry. A repair's predicate is the only definition of its work: the count the settings page shows, the list the job fetches and the check before its handler run are one predicate, so the job converges. That is why the registry is cut to what the device can do rather than listing what it skips — an entry is a count and a set of originals to fetch — and why an eye reading that cannot be cut is not a criterion. The catalog side is generic to match: record_updates writes whichever fields a FaceUpdate carries and re-marks the image so the shards export it; faces_needing and count_needing answer a predicate the caller supplies, replacing the measuring pass's three special cases. Two buttons on the settings page run the job and differ in one predicate. "Index faces" converges on coverage: has anything examined this image. "Re-index every face" converges on provenance: face-detection claims every image with no marker under the chosen detector, in either of its forms (FaceDetector::model_ids, so a desktop in f32 and a tablet on the Hexagon do not re-index each other's work), and a marker saying a weaker one looked is not that. An original over the fetch budget is left exactly as it was under the re-index, where the sweep marks it examined: a re-detection with nothing found would delete the faces, and "cannot fetch" is not "no faces". |
||
|
|
16f3fb41a3 |
Measure the faces already found rather than finding them again
Benchmarks / CPU and I/O (per commit) (push) Successful in 3m13s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 54s
Build and test / Layer separation (push) Failing after 1s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 2s
Traceability / Requirement traces (push) Successful in 52s
Build and test / Android (aarch64) (push) Successful in 30m6s
Every face stored before its quality was kept holds a unit vector, and V14 forgot the run marker of each image holding one so that the next sweep would look again. Looking again meant detecting again: a whole re-detection per image, with every suggestion on it thrown away and the confirmations carried across by box overlap, to recover one number. The sweep now has a measuring pass between the proxy repair and the un-indexed images. It lists every image holding an unmeasured face, fetches the original once, warps each stored face from the landmarks it already has, embeds it, and writes the raw vector and its length over the old row. Ids, boxes and identities are untouched; the marker is re-written fresh so the sync exports the measured vectors. A face whose landmarks no longer make a warp is dropped, as detection would have refused to store it. `faces_unindexed` leaves those images to the measuring pass, so the V14 deletion no longer costs a second detection. |
||
|
|
8eeb9ba0f6 |
Offer the backup, and then the rebuild, when the index turns out to be damaged
NFR-R6 asks for an integrity check at startup and two offers behind it, and none of it existed. `PRAGMA integrity_check` appeared nowhere in the tree, `Catalog::open` was `open` → `configure` → `migrate` → `backfill` and nothing else, and corruption therefore surfaced as whatever rusqlite error the first unlucky query happened to produce — "database disk image is malformed" attached to a thumbnail refresh, elided into a 34px banner, over an empty grid saying "No images found · Check the library folder". Two messages that disagreed, and no way forward but deleting catalog.sqlite by hand. The property that makes the second offer real was already here and load- bearing: the catalog is an index, not a source of truth, rebuildable from sources plus sidecars (invariant §5.2.4, cited by schema.rs, trash.rs and lib.rs). And sync.rs already knew how to take a coherent snapshot of a WAL database. What was missing was the check, the type, and the conversation. Four pieces: **The type.** `CatalogError::Corrupt`, and — the part that makes it worth having — a hand-written `From<rusqlite::Error>` that classifies rather than wraps. `SQLITE_CORRUPT` and `SQLITE_NOTADB` become `Corrupt` wherever they arise, so a background job that trips over the damage first reports the same thing the startup check would have. `SQLITE_IOERR` and `SQLITE_BUSY` deliberately do not: a dropped network mount is a different problem, and telling someone to rebuild their index would be a wrong answer delivered confidently. **The check.** `Catalog::open_verified`, `quick_check` before the open rather than after, because opening runs migrations and a damaged catalog with an intact header would otherwise have structure rewritten on top of structure that is already wrong. Bound to `open_verified` and not to `open`: the check reads every page, which is affordable once at startup where a user can answer a question, and not affordable on the dozens of opens a session's background tasks make. **The backup.** NFR-R2's second clause, taken between `configure` and `migrate` in `Catalog::open`. A migration is the one routine operation that rewrites table structure, so it is the likeliest way this file becomes unreadable, and it is the last moment the pre-migration state exists to be copied. Three generations, through SQLite's backup API after a TRUNCATE checkpoint — never `fs::copy`, which on a WAL database backs up a state older than the catalog and possibly torn. A failure to take the copy is logged, not raised: a full disk must not be what makes a library unopenable. **The conversation.** The first line of the dialogue is that the photographs and the edits are safe, before the diagnosis, because that is the question the user is actually asking. Then the two offers, which are *not* interchangeable and are not presented as if they were: a restore keeps collections, and a rebuild cannot, because a manual collection is a set of images assembled by hand and nothing in the filesystem records it (docs/catalog.md §8.1). The labels say so, and the rebuild does not take the affirmative styling while a restore is on the table. One thing that is a fix rather than a feature: `show_catalog_now` now gates the scan. `Catalog::open` succeeds on a file whose header survived, so the scan that used to start immediately afterwards would write folder ETags and image rows into damaged pages in the seconds while the user was still reading the question — turning a file that had a backup into one where the backup is the only copy left. Restore also deletes the damaged catalog's `-wal` and `-shm`. That step is easy to leave out and fatal to leave out: a journal belonging to the old file, sitting beside the new one under the same name, is replayed into it on the next open. That is not a restore, it is a fresh corruption with the evidence gone. Tested by corrupting a fixture catalog — 500 images and a collection, then every page past the second overwritten — and driving both branches. The restore is asserted on the collection, because a collection is precisely what distinguishes the two paths; the rebuild on the damaged file being kept and the next open producing an empty catalog at the current schema. Plus the `SQLITE_NOTADB` presentation, a damaged backup being refused rather than installed, and a v1 catalog whose pre-migration backup comes back reading v1 rather than v11. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
846a249156 |
Drain the queue that nothing has ever drained
`jobs` has been a complete durable work queue since the catalog was written, and nothing has ever taken a job out of it. `claim_next`, `complete`, `fail` and `recover_orphaned` had no callers outside their own tests; `enqueue` had three. So the table grew one row per photograph and kept it forever, and FR-PLAT-AND-3's resumability was a property of code that never ran. `runner` is the missing half. It owns no thread, no clock and no policy, and that is the whole design: on Android the process does not decide when background work may run. WorkManager does, subject to Doze, battery saver and FR-NC-6's network constraints, and it revokes permission mid-job by calling onStopped(). So the runner exposes `run_one` — claim, run, record — and `drain`, which repeats it against a budget, a deadline and a cancellation flag the host owns. A `Worker.doWork()` with ten minutes calls drain with a deadline; a desktop idle pass calls it with none. That is the seam the Android service plugs into, and it needs no Android to test. Handlers are supplied from above, because the catalog knows what needs doing and nothing about how: a thumbnail needs a decoder and a fetch needs a network stack, neither of which belongs under core/dr-catalog. A runner claims only kinds some handler declares, so a queue holding work this device cannot do is left alone rather than failed five times. Four outcomes, and only two of them are the job's fault. Done deletes the row; Retry backs off; Abandon gives up now, for a failure no retry can fix; Interrupted releases the claim with its attempt refunded and ends the drain, because the host stopped rather than the job — five backgroundings in a row must not mark good work as failed. Process death is the fifth and cannot report itself, which is what `recover` is for. Recovery is called from `show_catalog_now`, which is the one place a catalog is opened for a session and already returns early if one is open. It has to be exactly once and before any worker starts: there is no owner column, so a second pass while a worker held a claim would take it away. The attempt a dead claim consumed is deliberately kept — a job that takes the process down with it is indistinguishable from one that fails, and the attempt counter is the only evidence that survives a death. The tests cover claiming under contention twice over: sequentially across two connections, and with four threads on four connections against one catalog on disk, asserting every job ran exactly once. Plus completion, backoff, giving up, abandoning, interruption, budget, deadline, cancellation, and a job orphaned by a simulated crash being reclaimed and run once rather than lost or repeated. Not wired to a handler yet, and deliberately not: the only enqueue site the app actually reaches is the remote scan's, whose thumbnails are already served by the async grid worker, and `walk`'s two sites are reachable only from the scan_local example. Inventing a handler to make the plumbing look used is how a requirement comes to read as covered by code that does not implement it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
82d9d077b0 |
Merge: answer Android's memory warnings, and stop reporting a lost root as an empty library
FR-PLAT-AND-5 in full, FR-PLAT-AND-2 in part -- the recovery is built and live for Nextcloud roots, the SAF cause it names does not exist yet. FR-PLAT-AND-4 and FR-PLAT-AND-6 are not here, both blocked behind the same gap: assemble-apk.sh compiles no Java, so the APK cannot carry a Service or a FileProvider. The container has JDK 17 and build-tools 36; the build step is what is missing. Verified: fmt, clippy --workspace --all-targets -D warnings, and 1043 tests across dr-catalog, dr-sync, dr-sync-folder, dr-sync-nextcloud, dr-plat and dr-ui. The aarch64 target was checked before the branch was finished but not after; no device was available. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
75fd5619ca |
Refuse a scan whose root has gone, instead of reporting it empty
FR-PLAT-AND-2, and a silent failure on both platforms. `dr_sync::scan` stepped over a NotFound or PermissionDenied the way it does for a child that vanished mid-walk -- correct for a child, wrong for the root, where it ended the walk, returned Ok with nothing in it, and reported a successful scan of a library that was no longer there. A lost root is now its own error. The images under it are marked Availability::Offline per FR-CAT-9 and no catalog row is deleted; `library::persist` clears the mark per file as each one is listed again, so a root that comes back needs no repair step. Partly satisfied rather than closed, and the gap is worth stating. The recovery half is real and reachable on Android today, because `map_status` turns Nextcloud's 403 and 404 into it and Nextcloud is how a phone actually gets a library in this build. The causes the requirement names -- revocation, reinstall, a removed card -- are properties of a persisted tree permission, and there is none: SAF does not exist here, `SourceRef::Document` is constructed only in test modules, and `LocalStorage` rejects the variant outright. When SAF lands it becomes a third producer of this error and nothing above it changes, which is why the discovery belongs in the connector. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d3dadbd725 |
Group the frames of one moment, by when they were taken and what they look like
A burst is the commonest thing in a cull and the least interesting: twelve frames of the same gull at 10 fps occupy twelve cells, are scrolled past twelve times, and end with the photographer keeping one. FR-CULL-5 asks for them to collapse to one representative and be judged as a unit. Two signals, because neither alone survives a real library. Time alone groups a whole wedding ceremony -- a photographer working steadily never leaves the gap that would end the run. Similarity alone groups a studio setup shot across two days, which is a project rather than a moment. Together they are specific: adjacent in time *and* looks like the frame before it. Two seconds is the time bound, and the reason is worth recording because the figure looks absurd next to a 10 fps camera. `images.captured_at` is whole seconds -- EXIF's DateTimeOriginal has no sub-second field and SubSecTimeOriginal is optional and widely omitted -- so a burst arrives in the catalog as ten frames sharing one timestamp. Any threshold finer than a second is a threshold on information that is not there. Where the pace really is faster, the similarity bound is what separates the frames. Similarity is a 64-bit difference hash over a 9x8 box-averaged reduction, compared between *adjacent* frames only. Chained rather than anchored on the first frame, because by frame twenty a camera following a bird has nothing in common with frame one while no two neighbours differ by much; the time bound is what stops the chain running away. There is no all-pairs step and there must never be one -- that is what turns a grouping pass into something nobody can afford to run over 50k images. Nothing here ranks a frame. FR-CULL-5 names the failure it is avoiding, which is rejecting the only frame of an important moment because somebody blinked, so there is no sharpness score and no best-of-burst. The representative is the earliest frame -- a fact about the clock, not a judgement about the photograph -- and the user's own choice lives in its own table so that rebuilding the grouping cannot erase it. Same argument `people.ignored` makes one subsystem over: nothing short of remembering a decision survives re-clustering. A newly found burst is recorded *open*. Collapsing on discovery would be tidier, and would also mean a background pass taking photographs off the screen part way through a cull. The pass marks; the user folds. It is a pass rather than a job kind for the reason catalog.md 10.2 gives for face clustering: a burst is a property of a run of frames and has no natural subject_id, so a per-image job would rebuild the world once per photograph. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
96d07da15f |
Sync face data as sealed shards, so a second device does not re-index
Indexing 23,500 images is about two hours of CPU, and the result is byte-identical on every device: the same model over the same proxy produces the same embedding. Paying for it once per account rather than once per device is the point. Shards rather than the catalog snapshot, because the snapshot goes up whole on every sync and a fully indexed library carries roughly 30 MB of embeddings. That is exactly the cost the thumbnail store's 25 MB cap exists to bound, so face shards use the same cap -- imported from dr_thumbs rather than restated, since the number is a statement about sync cost and the two must not drift apart. The split follows the one already there: bulk immutable data in sealed shards, small mutable data in the catalog snapshot. Faces, landmarks, embeddings and run markers shard; people, names and assignments ride the catalog and merge by uuid. Keyed on oc:fileid throughout, never on image_id, because a row id means nothing on another device. The run marker travels with the faces it describes. Without it a receiving device cannot tell an image with no faces from one never examined, and would re-detect every landscape it had just adopted. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
aac3136407 |
Store faces and the people they belong to
Schema v8: people, faces, face_person, face_person_rejected, and the per-library calibration. Follows catalog.md 10.1 with two additions the spec work turned up. crop_px, because at the 1024px proxy tier a group shot reaches the embedder at ~50 source pixels upsampled to 112 and a portrait at 340. FR-CULL-9 names face size as an axis along which an uncalibrated similarity misbehaves, so it is a stored feature rather than a UI hint. face_person_rejected, because rejection is not the absence of an assignment. Without it the next clustering pass re-suggests exactly the face the user just pushed away, and the tool feels broken. record_detections replaces rather than appends, since DetectFaces is coalesced per image -- and carries confirmations across the replacement by box overlap, so re-indexing with a better model cannot discard the user's own labelling. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
fa4ad6e2d6 |
Drag the date range on the axis it is chosen from
The range could be turned on with a finger and not aimed with one. Its two ends were typed as `YYYY-MM-DD` into 108px fields behind a soft keyboard, to name days already drawn on the axis a thumb away; and the chip that seeds them takes its span from the timeline's zoom and pan, which are a wheel and a middle button. A touch screen has neither, so on Android the filter was a switch with no aim. The band is now on the timeline. Two ends with grips, dragged along the bars, released to filter — the histogram was already how a period is found, and this makes it how a period is stated. Both ends snap to whole days, which is what the typed fields mean, what `show_range` reads back out, and a floor under a range dragged shut. The fields stay for what dragging cannot do: name an exact day, and say in words what the range is. For that to work the axis had to stop following the range. Redrawn to the band, it moved the ground under the very handles doing the narrowing, and there was nothing outside the range left to widen back into. While there: a fixed number of equal bins instead of calendar buckets. Between one calendar unit and the next the bar count is free to wander by a factor of twelve, so zooming in halved it two steps out of three — the same picture drawn wider until it jumped back to fine. Equal bins also include the empty ones, so a bar's position on the track and the date under it are finally the same quantity; before, a library with gaps drew a February six months wide and the marker, the band and a click all pointed somewhere else. The count is a setting, 32 or 64, because the right answer is a question about the screen. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4b7648082e |
Let a date range be stated, and draw the axis at the scale it deserves
"Limit to range" did nothing, and the reason was not visible from the button. It took its span from the timeline's zoom, which is zero until someone zooms — so `zoomed_span` returned the whole library and the filter narrowed to everything. The chip lit up and the grid did not change. The range has ends now, shown and typed as `YYYY-MM-DD`. Seeding them from the timeline is kept, because zooming to a fortnight and pressing the chip is the fast path; the fields say which fortnight it landed on and let it be corrected. Ends given backwards are swapped rather than refused — there is exactly one range between two days — and the closing day is included, since "to the 5th" means the whole of the 5th and a range ending at its midnight contains none of it. `parse_date` refuses anything that is not a date rather than guessing at an order, because the alternative is a library silently filtered to a span nobody asked for. The axis then follows the range. It used to keep drawing the full extent while a range was on, because it was the only way back out; the typed ends are the way back out now, so it is free to show what was asked about. And bucket size is chosen by how many bars it makes rather than by fixed cut-offs. Each zoom step halves the span, so under thresholds the bar count halved with it until a boundary was crossed: fifteen years went 15 bars, 8, then 46, 23, 11, and finally 6. Zooming in made the picture coarser, which is the opposite of what zooming is for. Aiming at forty bars keeps the count in the same neighbourhood at every level, and the test asserts the property directly — halving a span never coarsens the bucket. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
56dd3187f1 |
Merge integration into wip/ingest
Second pass, against the detail-stage and thumbnail work that has landed since the first. Resolved and verified here rather than in the shared merge worktree, so what goes back is a fast-forward. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> # Conflicts: # docs/traceability.md |
||
|
|
2147eaa6a5 |
Put a keyword on a photograph, not only search for one
The catalog has been able to *find* by keyword since v1 — query.rs joins the keywords table, matches it exactly, and substring-matches it for free text — and nothing anywhere could ever put a word there. A user could filter to a keyword they had no way to apply. This is the missing half: create, rename, delete, list, assign, unassign, and the two reads a panel needs. Bulk-only for assignment, because keywording a selection is the common case rather than the exception — the photographer picks out the frames with the puffin in them and applies "puffin" once, in one transaction. Schema v6 adds `keyword_terms`, and deliberately does *not* touch the v1 join. The assignment keeps the word as text because the catalog is a rebuildable index and the durable copies of that fact — the sidecar, XMP dc:subject — both carry a string; a foreign key would mean a catalog rebuilt from sidecars had to invent identity rows before it could record anything, and would break the query path that already works. So the text is the fact, and the new table is only the identity a rename and a deletion can be keyed on. `keyword_terms.name` carries no unique index, which looks like an oversight and is not: two devices that each type "Iceland" are both right until they meet, and a constraint would abort the merge at that moment. Uniqueness is converged upon instead — create resolves an existing name, fuse_duplicates collapses a cross-device pair onto the smaller uuid. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
735683b849 | wip: ingest | ||
|
|
02d629922f |
Draw the date histogram over the collection you are looking at
Build and test / Desktop (Linux) (push) Failing after 57m25s
Build and test / Layer separation (push) Successful in 33s
Traceability / Requirement traces (push) Failing after 27s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
Build and test / Android (aarch64) (push) Failing after 9m45s
The timeline counted the whole library whatever the grid was showing, so opening a collection left a fortnight in Arosa as one column of a fifteen-year axis — an axis describing photographs that were not on screen. Scope the buckets and the span to the same collection and rating filter the grid uses. `timeline_range` counts `images` alone and cannot express the membership join, so the scoped query lives beside the other scoped readers in the UI and shares their descendants-of-scope rule. `catalog_span` now delegates to the same scoped reader. Zoom and scrub measured the full library while the bars were scoped, so a scrub could land on an instant the collection did not contain and send the view somewhere the user had not asked to go. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
bb71f141e7 |
Let a folder on this machine be scanned into the catalog
`dr_catalog::scan` has known since it was written what a changed directory means — when to prune, when to list, and the one question that decides whether a deletion sweep is safe. It was fully tested and nothing called it, because walking a real directory "belongs to the platform layer" and the platform layer was eleven lines re-exporting `secrets`. So every photograph in DarkRoom arrived over WebDAV, and a user without a Nextcloud account saw nothing at all. This is the missing half: a `Storage` trait, a filesystem implementation of it, and the driver that pours one into the other. The trait is shaped by the platform it does *not* yet support. Android's SAF gives no filesystem path, which is why `SourceRef` exists; less obviously, it gives no way to *compose* one either — a document id is opaque, and the only way to learn a child's id is the children query that returned it. So a listing hands back the reference to each entry rather than a name for the caller to join onto a parent, and there is deliberately no "path + name" helper anywhere above `LocalStorage`. That single restriction is what makes SAF a second implementation rather than a second set of call sites. A reference is otherwise an opaque `(RootId, key)` pair the catalog stores verbatim and rebuilds later, which a persisted tree grant supports exactly as a relative path does. A `Path` now appears in one place: `LocalStorage::grant`, where the folder the user picked is handed in. Everything above it addresses a `RootId`. `dr_catalog::walk` is the seam. It probes a directory, asks `scan` what that means, lists only when told to, and reconciles what it found against the rows it holds. Two things it does are worth saying out loud, because both are ways to lose a library: Absence only counts where absence was observed. A listed folder proves its missing images are gone; a pruned one proves nothing about its contents, and a scan that was cancelled or that failed part-way proves nothing about folders it never reached. So the file sweep runs per listed folder, the folder sweep runs once at the end and only after a complete scan, and a root that cannot be reached at all marks its images offline and deletes nothing — FR-CAT-9's line between proven-absent and merely-unreachable, which is the difference between unplugging a drive and losing everything on it. A trashed image is absent from its folder on purpose. It is exempt from both sweeps, and detached from a folder about to be deleted rather than cascaded away with it, or a soft delete would come undone the first time the folder it came from was rescanned. Two things the tests taught, both changes to what was there before: Modification times are now milliseconds, not seconds. Change detection asks whether a timestamp moved, so the unit's granularity is the width of the window in which a change is invisible — and a second is long enough to copy a card and start a scan. The test that caught it looked like a test bug; it was not. SAF reports milliseconds natively, so this is also the unit that needs no conversion on the platform with the coarser clock. And an in-place rewrite of an existing file is invisible to directory-level pruning, because writing to a file moves neither its directory's mtime nor its entry count. That is a real limit, now documented and held by a test rather than left to be discovered. It bites less than it reads: an export, a restore, `mv`, and every editor that saves safely write beside the file and rename over it, which does move both. Narrowing the format filter no longer deletes what it stops matching, which fell out of the same principle: unticking JPEG says stop looking for new ones, not discard the hundred already rated. The files are sitting right there. `DirState` and `DirEntry` move to `dr-types`. They are the sentence the platform says to the catalog and both crates need the same one; `scan` re-exports them so nothing that used them has changed. Not done: the UI. The launch screen's "Open library" flow is account-shaped from the first field to the thumbnail worker, and giving it a local branch is its own piece of work rather than a button. `cargo run -p dr-catalog --example scan_local -- ~/Pictures` scans a real folder and reports what it cost; run it twice to see the second run list nothing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
fa12afed18 |
Keep originals on this device, by pin and by use
Fills in `image_cache`, which the previous commit's "On this device" filter read but nothing wrote. Also carries in-flight work that shared these files: the Android TLS root store, the settings page, and a regenerated traceability report. # Two populations, deliberately separate An original is kept here for one of two reasons, and conflating them produces the exact failure the feature exists to prevent. **Pinned** originals were asked for. Pinning a collection before a trip is a promise, so pinned rows are never evicted and never counted against the budget — a cap that could silently delete a pinned trip would make pinning worthless, because it could not be relied on without checking. **Passively cached** originals are a side effect of working: develop already downloads the whole file, so keeping it costs no bandwidth and saves the entire transfer next time. This population is what the budget bounds, evicted least-recently-used, because it otherwise grows until a day of culling fills a disk. Sharing one budget would let a large pin starve the passive cache, or let browsing evict a pin. They are separate. # What was built `dr_catalog::cache` owns the bookkeeping — held tier, size, last use, pinned — and writes the bytes; deciding to download stays with the caller, which is what keeps a crate with no network out of the network's business. Files are written to a temporary and renamed, so a dropped connection cannot leave a truncated file recorded as a complete original. They are named by image id, not filename: `Photos/IMG_0001.CR2` and `Trips/IMG_0001.CR2` are different photographs, and a flat cache keyed on the name would serve one for the other. `spawn_full_fetch` became read-through. A hit is a disk read; a miss stores what it downloads and enforces the budget. A cache that cannot be opened is a miss, not a failure to open the photograph. Pinning writes intent — `tier_desired` — without downloading, so the button responds immediately, and `spawn_pin_fetch` fills it in sequentially afterwards. Sequential because these are tens of megabytes each: the lanes that make the thumbnail sweep fast buy little against one connection's bandwidth and cost a great deal of memory. A pin interrupted by a lost connection resumes from where it stopped. Schema v5 adds `pinned` and `path`. `pinned` is a column rather than something inferred from `pinned_by_rule`, which is ON DELETE SET NULL and so cannot answer for an image whose rule was deleted. A v4 catalog migrates in place; existing rows default to unpinned, the safe direction. The budget and "keep opened originals" come from the settings page rather than a constant, and are applied at startup rather than only on change — a cache capped at 2 GB last session would otherwise spend this one filling to the default. Turning off keeping leaves what is already cached readable: those bytes are paid for, and refusing them would re-download images sitting right there, including pinned ones. Also removes a doubled `#[test]` introduced in the previous commit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d5b1f6bff5 |
Add collections, ratings, and soft delete to the catalog
Three features over a shared schema migration. Collections: a tree of manual collections plus smart collections whose membership *is* their stored selector. Dropping images onto a smart collection is refused rather than silently discarded, so the UI can say why the drop did nothing — member rows there would be a second source of truth that nothing reads. Ratings: the star and pick/reject axes, kept independent. Trash: soft delete to a folder, then permanent delete. Catalog::open now backfills after migrating. A migration adds a column but cannot know what the value should be for rows that already existed; backfilling on open is what stops those rows being silently partial. Timeline queries exclude shadowed JPEGs, which would otherwise double every paired shot in the histogram, and gain a range-bounded variant so zooming in returns finer buckets rather than the same coarse ones with the ends cropped. Assisted-by: LLM |
||
|
|
c8bb08e661 |
Add folder scan with format selection; validate A3 on a real library
Library setup as the user described it: pick a folder, choose which RAW
types to look for, scan recursively.
dr-types::FormatFilter the tick-box selection, seeing through VFS
placeholder suffixes so a dehydrated CR2 still
matches as a CR2
dr-sync::scan recursive walk, Depth:1 per directory, pruning
unchanged subtrees where the backend propagates
directory ETags
Verified against nextcloud.tourolle.paris (34.0.2) on a real library:
browse root 32 entries, 98ms
scan PhotosRaw 17,185 RAW files in 334 directories, 34.1s
(7,836 CR2 + 9,349 DNG)
range read 262KB of a 21.5MB DNG in 119ms — 1.22% of the file,
and enough to read "Canon EOS 6D | ISO 100"
That last line is assumption A3 validated on real data. Cataloguing this
library by whole-file fetch would move roughly 370GB; the range path
moves a few MB.
Pruning is capability-gated rather than assumed: with per-entry ETags a
probe costs a request and proves nothing about children, so it is skipped
entirely. A test asserts zero probes in that case.
Still unresolved: /core/preview returns 400 for every parameter
combination tried, including on a JPEG the server reports as having a
preview. Not a request-shape bug — it fails identically bare. Recorded
rather than worked around; ARCH §6.7 already treats server previews as
opportunistic, so nothing depends on it.
|