The library holds the same RAW in several folders: a dated folder, a
bck/ beside it, a renamed Darktable export tree. dr_catalog::duplicates
is the catalog half of consolidating them (#67).
candidates() is one grouped query over root, camera, capture instant and
size, joined back for the rows; count() is the same grouping under
COUNT. On a copy of the reference catalog (23,582 images) both take
10-35 ms and find 1,836 groups holding 3,379 spare copies.
survivor() prefers a copy outside a backup-looking folder, then one
still named the way the camera named it, then the oldest, then the
lowest id.
consolidate() re-checks the plan against the catalog, merges the copies'
judgements onto the survivor (highest rating, keywords unioned,
collections unioned with the survivor keeping its place, a flag or label
the copies agree on, faces via faces::carry_onto_copy) and records the
copies as trashed, all in one transaction, so a failure part way leaves
the group untouched. preview() runs the same code and rolls it back.
Sameness probes are kept in dedup_probes, created on first use rather
than by a migration: a schema bump would make older builds refuse this
catalog's snapshot at sync. trash::record_trashed_within lets the trash
write share the merge's transaction.
`persist` runs after every scan, for every photograph the scan listed. On
a settled library that is the folders whose ETag changed -- a sidecar
written there by a rating is enough -- so one relisted folder of 1,600
images is an ordinary pass, and a first scan is all 24,000.
Per photograph it prepared four statements from their SQL (a folder
lookup, the image upsert, the id read-back, the remote upsert) and then,
after the commit, found the image again by path and enqueued its
thumbnail job as an autocommitting statement of its own -- a commit per
photograph, for rows that were almost all already queued.
Now the statements are prepared once per pass, a folder's id is looked up
once per folder rather than once per photograph in it, and the job is
enqueued inside the transaction with the id already in hand. That also
makes the job atomic with the row it points at, which is what the old
ordering after the commit was trying to guarantee. `jobs::enqueue` uses a
cached statement for the same reason.
persist_bench on a copy of the reference catalog, CPU, best of runs:
largest folder (1,589 images) 102-118 ms -> 10-13 ms
whole library (23,582 images) 1.55-2.19 s -> 188-192 ms
The fingerprint of images, remote, jobs and folders after the run is the
same for both builds.
`adopt_orphan_terms` runs in the backfill on every catalog open. Its
check -- is there a vocabulary row for this word, tombstones included --
cannot use `keyword_terms_name`, which is partial on `deleted = 0`, so the
correlated subquery scanned the vocabulary once for each of the 10,800
assignment rows before `DISTINCT` threw the repeats away: 3.5 ms per open
on the reference library.
The distinct words are taken first and the check runs once per word -- a
few dozen scans of a few dozen rows. Same rows out, since `DISTINCT` over
the assignments is exactly the set of words.
catalog_bench, best of 20: 3.45 ms -> 0.30 ms.
`Catalog::open` runs the backfill every time, and every worker thread
opens its own catalog: the develop view does it to fetch each original and
again for each neighbour it prefetches, and the sync, sweep, burst and
thumbnail workers each do it too. On the reference library (24k images)
an open cost 26 ms of CPU, and most of it was `pair_raw_and_jpeg` reading
all 17,000 RAWs into a map of lowercased stems to find partners for the
1,900 JPEGs that have none -- the same 1,900 on every open.
It now starts from the small side. The unpaired JPEGs are read first, and
it stops there if there are none; otherwise it reads the RAWs in the
folders those JPEGs sit in (plus the unfiled ones when an unfiled JPEG is
waiting), which is 142 on the reference library. A pair is same-folder by
definition, so no pairing is lost; the RAWs are read in id order, so where
two share a stem the later one still wins as it did in the table scan; and
a pass with nothing to pair no longer opens and commits an empty write
transaction.
catalog_bench, best of 20, CPU: `Catalog::open` 26 ms -> 12 ms together
with the next commit (the backfill 24 ms -> 11 ms; this step is ~10 ms of
that). A test covers pairs found among other folders and unfiled images.
Every sync pass exports this device's faces to the shard store and imports
what peers sent, and both walked the whole library asking the store's index
about one image at a time: the export 19,000 `indexed_at` lookups (one per
face marker), the import 23,000 `held_model` lookups (one per image with a
server id), each a statement prepared and run against the index. With
nothing new either way -- the usual pass -- that was all they did.
Measured with catalog_bench against copies of the reference catalog and
face store, best of 5, CPU:
export_to_shards (steady) 119 ms -> 27 ms
import_from_shards (steady) 250 ms -> 87 ms
Each now reads the index in one statement into a map. The import's query
is `held_model`'s, ordered the same way, keeping the first row per file,
and nothing in the loop changes which pipeline a file is held under
(`set_indexed_at` touches only a file already decided; candidates are
distinct files). The export's `put_image_at` does rewrite entries -- but
only its own file's, its generation and the siblings it supersedes -- so a
file already written in this pass is asked of the store again, and every
other answer is the one the lookup would have given. An index that cannot
be read gives an empty map, which is what each failed lookup returned.
The store index and the catalog are identical after the old and new
builds' runs.
A sync pass that brought nothing new cost 450-540 ms of CPU in
`merge_remote_catalog` on the reference library (24k images, 19k faces),
measured by catalog_bench merging a copy of the catalog with itself.
Most of it was the loop over the other device's confirmed faces and the
faces under its ignored groups -- 13,000 rows. For each one it prepared
three statements from scratch (`query_row`/`execute` with a SQL string
compile the statement every call) and then rewrote the `face_person` row
with the values it already held, dirtying a page per face on every pass.
The rejection loop prepared three more per row.
The statements are now `prepare_cached`, the local assignment is read once
per face (whether it is confirmed, and what it holds, come from the same
row), and the upsert is skipped when the row already says exactly that.
`faces_assigned` is still counted for those rows, so the report is the one
the old code gave, and nothing else reads the difference: the row is
byte-for-byte what the upsert would have written.
After: 279 ms (best of 5, CPU), with every catalog table identical after
the run to the old build's.
Two benches for reading side by side before and after a change, against a
copy of a real catalog, in the manner of identity_bench:
- `dr-catalog --example catalog_bench CATALOG [FACES_DIR]` times
`Catalog::open` and the backfill inside it step by step, the upload
snapshot, a merge of the catalog with a copy of itself, and the face
shard export and import in the steady state where nothing is new.
- `persist_bench`, an ignored test in dr-ui's scan module because
`persist` and `apply_judgement` are private to it, replays the
catalog's own rows through `persist` (the largest folder, and the whole
library) and looks up every `.drsc` sidecar the catalog has read. It
works on a scratch copy and prints a fingerprint of what `persist` left,
so two builds can be shown to agree.
Both print best, median and CPU time; the CPU figure is the one to compare
while other builds share the machine.
Colour labels could be read from a Lightroom sidecar and queried by the
selector, but nothing drew one or set one, so the only labels a library
held were ones another program had written.
Every mark carries its label's initial on its colour — R, Y, G, B, P —
so a label is read without telling red from green, which is what
NFR-A11Y-3 asks of colour labels by name. A grid cell shows the mark
before its filename. In the grid, 6, 7, 8 and 9 set red, yellow, green and
blue as Lightroom's keys do, on the photograph under the pointer or on
the selection by the rule the star keys follow; the same key again takes
the label off, and over a mixed selection it sets it on all. The
selection bar gains Label, which opens the six choices — each a mark and
a name — and purple, which has no key, is there. In develop the top bar
says "Label: Green" beside the mark, opens the same choices, and 6-9
label the open photograph.
Each gesture is one catalog transaction, then the grid, the counts and
both sidecars are written as a rating's are. The filter bar gains a chip
per label, its mark and its name with a count, one at a time; the filter
is one SQL term, travels in the place record, and "All" clears it.
Colour labels reached `versions.label` only from an XMP sidecar: nothing
in the catalog could set one, clear one, or read it back alongside the
stars, so there was nothing for an interface to call.
`set_label` and `set_label_many` write it the way ratings are written,
the bulk form in one transaction so a key over a selection is one commit.
`toggled_label` holds Lightroom's rule for a label key: it clears only
when every image already carries that label, and otherwise sets it on all
of them, so a half-red selection comes out red rather than inverted.
`Judgement` carries the label, so the grid's one window query brings it
with the stars, and `label_histogram` counts each label in one grouped
statement for the filter chips. A label does not make a frame "judged":
it is a pile of the photographer's own, not a cull decision.
The doc comment for `default_version_id` had been stranded above
`label_code` when that was inserted; it is back on its function.
The shard import recorded each adopted image in its own transaction:
fourteen thousand commits, and fourteen thousand turns at the write lock
that every read on the UI thread queued behind — the sync was felt as a
laggy grid and as "database is locked" from whichever writer lost the
wait. `record_detections_within` takes the caller's transaction, and the
import commits every hundred images.
The store carried every detector generation of an image — 24,123 entries
for 19,089 images on the reference library, a third of its 293 MB — when
only the strongest is ever adopted. A put now skips a pass a held one
outranks, and retires the passes it outranks from the index; sealed
shards keep their bytes, but nothing is written twice from here.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
baed1c4 landed it over rustfmt's width; `cargo fmt --check` is the
first gate the Desktop job runs.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adopting an image from a peer's face shard stamped its run marker as now,
and the export reads a catalog marker newer than the shard's as a
re-index. So every adopted image went straight back out under this
device's client id: 14,100 adopted, 15,457 "newly indexed" on the next
pass, twenty-two shards of a peer's faces uploaded a second time.
merge_shard now carries the peer's indexed_at into the local index, and
the import writes that marker into face_index; where an older peer's shard
carries none, the store takes the catalog's, so the two agree either way
and the export finds nothing to send.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
docs/ had 26 developer documents flat beside the manual, and the two
audiences are very differently sized: most readers want the manual and
the gesture reference, a few want the register, the designs and the
measurements. The manual and gestures.md stay at the top; everything for
someone changing the code moves to docs/dev/, and the two documents that
name their own successors — the v0.1 milestone and the UI-refinement plan
— go to docs/dev/archive/ rather than being deleted, since both are still
cited. docs/README.md is the index, users first.
Every reference follows: code comments, Cargo manifests, the workflows,
the pre-commit hook, the bench and traceability tools (which locate the
repo root by docs/dev/requirements.md now), packaging, the Docker READMEs,
CLAUDE.md, CONTRIBUTING.md and the README. The matrix links one level
deeper and is regenerated. Links out of the moved documents into the tree
gain a level; a link checker over every Markdown file finds none broken.
The derived sync fired only after the metadata sweep, so a fresh device
re-derived every thumbnail it scrolled past, re-detected faces and re-read
every header for hours before adopting the shards and snapshot that held
all of it. It now fires as soon as the scan completes — the first moment
the rows the merges key on exist — and the sweep starts behind it. In
steady state that pass is one listing.
The catalog merge gains a fourth half: capture metadata (captured_at,
offset, camera, lens, ISO) for images still at metadata_state < 2, matched
by oc:fileid from a remote row at 2. A date is a fact about the file's
bytes, not local state, and the snapshot already carried it. The sweep's
per-chunk query then finds nothing left, and the timeline is whole on a
fresh device without a header fetch.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The markers the previous commit stops writing are already in the
catalogs — 2 on the desktop, 429 on the tablet — and in the shards
both have exchanged. Renaming them to the faces' own id with a fresh
time is what makes the export send each image again, under an entry
newer than the empty one `held_model` would otherwise pick. Where the
old write had inserted its marker beside the right one, the wrong one
goes and the right one is refreshed for the same reason: its entry in
the shards is older than the empty one.
Images V14 left with faces and no marker are not touched. That state is
the quality pass's cue, and the fixed write marks them correctly when
it reaches them.
Checked against copies of both real catalogs: the desktop renames 2,
the tablet deletes 429, both in under 200 ms.
`record_updates` — the write behind the quality, eye and crop passes —
re-marked the image as indexed under the pipeline the pass ran as, and
left the faces it had updated under the id of the detector that found
them. On a desktop set to Thorough that put `scrfd_10g+w600k_mbf` over
faces spelled `w600k_mbf`; on the tablet, `scrfd_10g_i8+w600k_mbf` over
faces it had adopted from the desktop's thorough pass.
Every reader takes the marker and the faces to agree. `marker_under`
reads the marker as the detector having examined the image, so the
upgrade repair never revisits it. The shard store keys each face by
its pipeline id, so `export_to_shards` selects an image's faces by the
marker's id, finds none, and sends an entry that says the thorough
detector looked and found nothing — over photographs with named faces
on them. The desktop's shard index holds 54 such entries beside real
faces; the tablet's eye pass over the faces it had adopted made 430
more, and both devices have exchanged them. `held_model` takes the
newest entry for an image, which is the empty one. Nothing has been
lost yet only because the two spellings of the thorough detector rank
equal and neither side adopts the other's; a third device, or either
one after a reinstall, would adopt "nothing here" for 484 images. And
the desktop's eye pass is 4,739 images from doing the same to every
face from before V14 — which are the ones that only exist on the
desktop, and would then never reach anywhere.
The marker now takes the id the faces carry; the pass's own id is used
only when it dropped the last of them and there is no detector left to
name. A stale marker under another spelling of the same embedder is
removed in the same transaction, so one embedder has one marker.
The tablet showed a fraction of each person: 681 of the desktop's 3,851
confirmations, and none of Ian's 746, Catherine's 626 or my own 480.
Every face that existed on both devices agreed on who it was, and the
people rows were identical — the merge was fine. The missing 3,170
confirmations were on faces the tablet did not hold at all: the
desktop's 16,080 faces from the original detector, on 4,310 images,
detected before schema V14 kept the quality reading.
Those faces were in shards the tablet had already downloaded, in
August's export. `import_from_shards` looked at them on every sync pass
and declined each one, because a face without a quality reading was
"work this device cannot finish": adopting it would write the run
marker, and the marker was what stopped an image being looked at again.
That was true when it was written and has not been since the quality
repair existed — that pass lists its work by `f.quality IS NULL`, not by
the marker, exactly as the eye pass does, and faces without an eye
reading were already adopted on that reasoning.
The refusal had no exit. V14 had deleted the markers of every image
holding such faces so the quality pass would find them, and
`export_to_shards` walks the markers, so the desktop never re-exported
them either; the unmeasured August copies were the only ones there
would ever be. The tablet's answer was to queue all 17,727 images for a
re-detection of its own, a fetch of the whole library, while holding
the faces on disk.
Adopt them. The receiving device's quality pass measures them when it
reaches them, and the desktop's confirmations match onto them by box
overlap on the next catalog merge. The test that asserted the refusal
now asserts the adoption and that the image is still owed to the pass.
"How many images still owe a quality reading" was a correlated EXISTS per
image over `faces`, and the face row is 8 KB of embedding and crop before
the column it looks at, so each count opened every row. Six such counts
run on every open of the Identity screen and at the end of every sweep:
160 ms on the reference library.
V19 adds three partial indexes holding only the faces still owing each
pass, keyed on the image and carrying the model id the predicate reads,
and replaces `faces_image` with `(image_id, model_id)` so "does this image
hold this embedder's faces" is answered from the index too. The planner
takes a partial index when the count is driven from `faces` and ignores it
inside the EXISTS, so `Needs::Face` carries the per-face fragment and
`repairs::count` spells the query from the faces' side; the list and the
per-image check keep the EXISTS. A test holds the two spellings to the
same answer for every repair.
`faces::people` grouped `face_person` after a LEFT JOIN over every person
and sorted the lot by name; the rail then discarded the empty, unnamed
groups a regrouping pass leaves behind — 17,000 of 19,000 rows on the
reference library. `people_in_use` filters them in the WHERE and joins
`people` to face counts aggregated first (2,000 groups), so the sort sees
only the rows that will be drawn. `count_unassigned` replaces fetching
2,400 ids to take their length. `load_people` 22 ms → 10 ms.
`confirm_all` called `faces::confirm` per face, and `split_off` called
`reject` then `confirm` per face: each opens and commits its own
transaction, so a click on a group of several hundred was several hundred
commits. `faces::confirm_all` is two statements — clear the rejections the
confirmations override, then flip the rows — and `faces::reassign` does a
split's reject-and-confirm for every face under one commit. 16 ms → 2 ms
and 22 ms → 4 ms on the largest group.
A library's records are never all complete at once. A face found before
its quality was kept has no quality; one found before the eye models
existed has no reading; one adopted from a peer's shard has no crop; an
image the fast detector examined on a 1024 px proxy has boxes the current
detector would not have drawn; an image the scan stat'ed has no capture
date. On the reference library that is 17,762 faces under the bare
w600k_mbf id with no quality, no reading and no dense landmarks, 4,144 of
them without a crop, beside 12,217 images the fast detector examined and
found nothing in. Every one of those gaps was its own pass — V14's
measuring pass, §17.5's eye pass, the sweep's proxy repair, the sweep's
detector upgrade — with its own work list, its own count and its own idea
of done, and adding a per-face field meant adding a pass. There was no
pass at all for the case the library is actually in: boxes and landmarks
drawn by a weaker detector on a proxy, which every later per-face pass
would have read from.
dr_ui::repairs replaces them with one job over a registry. A Repair names
one thing a record can lack — the predicate that says which images still
owe it, the input its handler needs (a header, the original, or a native
render), the handler, and what to record for an image that can never be
done. The job unions the predicates into one work list, fetches each
image once at the most any claimant asks for, renders it at most once,
and runs every handler whose predicate that image still matches, checked
again before each because a detection writes every field a per-face
handler would fill. The registry today: face-proxy, face-quality,
face-eyes, face-crop, face-detection, face-upgrade, metadata — the last
there to say that this is not a face job. Adding a field is one entry.
A repair's predicate is the only definition of its work: the count the
settings page shows, the list the job fetches and the check before its
handler run are one predicate, so the job converges. That is why the
registry is cut to what the device can do rather than listing what it
skips — an entry is a count and a set of originals to fetch — and why an
eye reading that cannot be cut is not a criterion.
The catalog side is generic to match: record_updates writes whichever
fields a FaceUpdate carries and re-marks the image so the shards export
it; faces_needing and count_needing answer a predicate the caller
supplies, replacing the measuring pass's three special cases.
Two buttons on the settings page run the job and differ in one
predicate. "Index faces" converges on coverage: has anything examined
this image. "Re-index every face" converges on provenance: face-detection
claims every image with no marker under the chosen detector, in either
of its forms (FaceDetector::model_ids, so a desktop in f32 and a tablet
on the Hexagon do not re-index each other's work), and a marker saying a
weaker one looked is not that. An original over the fetch budget is left
exactly as it was under the re-index, where the sweep marks it examined:
a re-detection with nothing found would delete the faces, and "cannot
fetch" is not "no faces".
record_detections replaces an image's faces and carried only the user's
confirmations onto the new ones, by box overlap above 0.5 IoU. Everything
else on the old faces was dropped: the suggestions the last grouping pass
made, and the people the user had said a face was not. On the reference
library that is 13,011 suggestions and 77 rejections beside 3,778
confirmations — a re-detection of it would have been correct by
FR-CULL-12's letter, since suggestions are derived data, and would have
handed back a People screen of strangers.
Now every old face is read before the delete — box, vector, assignment,
rejections — and matched to the new faces one-to-one, best pair first. A
pair qualifies when the boxes overlap at all and either the overlap alone
says so (IoU above 0.5, the old rule) or the embeddings do (cosine above
SAME_FACE_COSINE, 0.45, the reference library's P≈0.95 line). The
embedding route claims the box a low-resolution pass drew badly enough
that overlap alone would not; the vector is also what breaks the tie in a
group photograph, where two neighbouring faces overlap both new boxes.
Overlap is required on both routes, because the same vector elsewhere in
the frame — a mirror, a print on the wall — is not the same face and must
not take its name. Onto the matched face go the assignment as it was,
confirmed or suggested with its probability, and every rejection.
The merge's match_faces still matches by overlap alone across devices; it
is the same question and is not changed here.
The 106 points the eye boxes were cut from, stored beside the reading as
16-bit fixed point over the frame: 424 bytes a face, a seventh of a pixel
on a 6000-pixel frame, where f16 at the same size would have been six.
Derived data like the embedding, kept for the same reason — it cost a
fetch and a model run, and the next per-face pass should run from the
catalog. Shards carry it; a peer's shard from before it is still read.
The register grew both clauses the same day this was built: FR-CULL-8a is
the per-face state the reading is, and FR-CULL-13 is the rule that a
signal is shown and filtered and never writes a judgement. The tags,
faces.md §17 and catalog.md now say which is which; FR-CULL-8a records
what of it is built, and that its third model is under the InsightFace
grant by the same decision as the pair.
The people filter was served from faces_image without touching a row;
reading the eye columns in the same subquery touched every one, and
ALTER TABLE had put those seven floats after the embedding and the crop
blob. One count took 24 seconds on the reference library, thirteen of
them system time. faces_eyes covers the subquery again: five
milliseconds.
Per eye P(open), the pixels across its box and the sharpness of the
patch; and P(sunglasses). The verdict — open, closed, sunglasses,
unclear — stays a rule in dr_face::eyes so the floors can move without
re-measuring twenty thousand faces. Shards carry the same seven, and a
peer's shard from before any of them is still read.
Three nullable columns beside quality — P(open) for each eye and
P(sunglasses) — because the verdict is a rule with thresholds in it and a
rule belongs in code, not in rows that would have to be re-measured. NULL
is "never read": a face from before the models, or from a device without
them, and every reader treats it as unknown rather than as closed.
The measuring pass V14 built for the embedding's length is what fills
them, so the sweep's work list now also names faces with no eye reading
— but only on a device that has the models, or it would fetch every
original to do nothing to it. A peer's shard without the reading is still
adopted, unlike one without the quality: the pass finds this work by the
NULL rather than by the run marker, so adoption costs it nothing.
Choosing "Thorough" made the library look empty. The detector setting
writes under its own faces.model_id, and every reader of "the faces"
keyed on that exact id: the clustering pass, the coverage figure, the
sweep's work list, the shard export and import, and the sync merge's
face matching. On the reference library that restarted coverage at
1,834 of 19,140, drew a People rail of 36 faces for a person with 520,
queued a ~400 GB re-fetch on each device, and stranded the desktop's
3,583 confirmations under the old id: the tablet held the same faces
under the new one and the merge refused to match them. Same photograph,
same box, same embedder, two ids — that is one face, not two libraries.
The embedder half of the id is now the key. embedder_of and embedder_sql
give it to every query; writes keep the full id, so which detector drew
a box stays on record. record_detections is unchanged and is where the
generations meet: an image holds one pipeline's faces at a time, and a
re-detection carries confirmations across by box overlap. The merge's
match_faces applies the same rule within an embedder. The calibration
is keyed on the embedder too, since the similarity space did not change.
Shards travel every generation, each under its own id, and a peer adopts
whichever it is sent — including a stronger detector's pass over an
image it indexed itself with a weaker one, which is the re-detection its
own sweep would otherwise queue, already done. Never downwards: a tablet
on Fast keeps the desktop's Thorough faces. The sweep gains the same
tail — images a weaker detector indexed, after the ones nothing has —
driven by FaceDetector::supersedes, so choosing a stronger detector still
improves the library over time without first making it disappear.
SQLite's busy timeout defaults to zero, and nothing ever set one: the
loser of a write race got SQLITE_BUSY at the moment it asked. WAL does
not cover this — it makes one writer and many readers free, and this
application constantly has two writers, the face sweep committing a
batch while the derived sync imports shards or reclustering reads.
The cost was not a retry but lost work. A sweep that had already paid
for the detection and the embedding — seconds per image, the expensive
part — discarded the result on "storing faces for 214: database is
locked" and moved on to the next image. Both the desktop and the tablet
logged runs of those on consecutive images, which is a face sweep
quietly failing to store the faces it had just computed.
Ten seconds, on every connection, set in configure() so that nothing
can open the catalog without it — the figure the job runner's own tests
have used for this reason since they were written. It is far longer
than any transaction here, so it bounds pathology rather than making
anyone wait.
NFR-R2 asks for the catalog to be backed up on a schedule and before
schema migrations. Only the second half existed: every backup on disk
was a pre-migration copy, and a library that never migrated was never
backed up at all.
A backup is now also taken at the end of a library sweep when the newest
one is more than a day old — the moment the catalog is quiet and a day's
collection and people edits have just been folded in — on its own
thread and its own connection, so the copy of a 130 MB file is not spent
on the UI. Whether one is due is read from the backup directory, not
the catalog, so the ordinary case costs nothing. An empty catalog is
skipped: there is nothing in it a rescan would not rebuild. Pruning to
KEEP_BACKUPS applies as before.
Two checks around the upload, both cheap next to what they prevent.
Before: the snapshot is quick_checked before it leaves. It is the copy
every other device merges from, and a damaged one costs each of them a
download, a failed merge and a refusal to push.
After: the staged upload's size on the server is compared to the bytes
sent before it is rotated into place. A chunked upload is assembled
server-side, and an assembly that goes wrong is a file of plausible
size no device can open — caught here, on the device that caused it,
for one listing; otherwise on every other device, after the fact. A
mismatch, or a size the server will not confirm, discards the upload
and leaves the current copy and its generations untouched.
FR-CAT-13 asked for standard XMP and `core/dr-xmp` answered the file: it
has read and written `dc:subject`, `xmp:Rating`, `xmp:Label` and the IPTC
core since 5fa4c07, under an ownership rule that leaves everything else in
the document untouched. What nothing did was call it. No scan found an
`.xmp` beside a raw, no catalog row was filled from one, no judgement
wrote one back, and the "external modification detected, reload offered"
clause had no mechanism. A library imported from Lightroom came in and
could not go back out.
The scan collects `.xmp` beside `.drsc` from the listings it was already
paying for, and the pull reads each one whose ETag has moved. Both
namings resolve: darktable's `IMG_0001.CR3.xmp` names its file exactly,
Lightroom's `IMG_0001.xmp` names the stem, and under the stem the JPEG
beside a RAW is the same photograph and takes the same document, as
DarkRoom's own sidecar already does. Each is reconciled with the catalog
winning — keywords union, a rating or label taken only where the catalog
has none — because a standard XMP carries nothing that could say whether
its value is newer. A genuine disagreement is not resolved; it is written
to a table, and the settings page offers the sidecars' values against it.
That button is the reload the requirement asks to be offered, and the
ETag that moved is the detection it asks for: an `.xmp` edited elsewhere
is exactly a file the pull's ordinary incrementality re-reads.
Writing goes the other way behind a setting that starts off, since NFR-R4
makes writes beside somebody's originals theirs to switch on. With it on,
a judgement or a keyword rewrites the sidecar of whichever spelling
exists, or creates Lightroom's. The record is read from the catalog
whole at that moment rather than carried from the gesture, so a rating
and a keyword a second apart are two writes of one file that agree. And
the file's own title, caption, copyright and hierarchy come through the
rewrite: the catalog has no columns for them, `rewrite` replaces the
owned set wholesale, and a record that said nothing about them would have
deleted them from a Lightroom sidecar on every star.
The rating's two axes cross the format's one field both ways: a
rejection is Adobe's `-1` and stars are stars, and stars arriving on a
rejected frame lift the rejection, since the file said it was worth a
number. An unrated file says nothing and clears nothing, on the rule the
`.drsc` merge keeps. `versions.label` finally has a reader and a writer,
with the code table moved out of the query so the two cannot drift.
record_detections replaces every face on an image whatever model found
them, but left the other models' face_index rows standing. With one
model that was unobservable. With a second pipeline it leaves an image
marked "done" under the first with none of its faces behind the marker
— the state the V12 repair existed to undo — and a user who switched
back would find those photographs permanently empty.
An image now holds the faces of whichever pipeline looked at it last,
and only that pipeline's marker. Confirmed names still carry across by
box overlap, since they were read before the replacement.
Every face stored before its quality was kept holds a unit vector, and
V14 forgot the run marker of each image holding one so that the next
sweep would look again. Looking again meant detecting again: a whole
re-detection per image, with every suggestion on it thrown away and the
confirmations carried across by box overlap, to recover one number.
The sweep now has a measuring pass between the proxy repair and the
un-indexed images. It lists every image holding an unmeasured face,
fetches the original once, warps each stored face from the landmarks it
already has, embeds it, and writes the raw vector and its length over
the old row. Ids, boxes and identities are untouched; the marker is
re-written fresh so the sync exports the measured vectors. A face whose
landmarks no longer make a warp is dropped, as detection would have
refused to store it. `faces_unindexed` leaves those images to the
measuring pass, so the V14 deletion no longer costs a second detection.
The embedder's raw output has a length, and the length is a reading of
how recognisable the crop was: a blur, an occlusion or a hard profile
comes out short. Normalising threw it away. A short vector sits near
the middle of the sphere and matches a little of everyone, which is how
one bad crop bridges two people in a grouping pass.
So the length is kept — the store now holds the raw vector, re-normalised
on load, with the length beside it as `faces.quality` — and a face under
MIN_GALLERY_QUALITY (14) is a probe: measured against the gallery and
placed where it fits, but never what another face is measured against.
Two probes are never paired, and a probe is nobody's evidence for a
confidence. The People screen shows the number as "Quality 17.3", dimmed
below the floor.
Faces indexed before this stored unit vectors and have no reading; they
are admitted to the gallery, and schema V14 forgets the run marker of
every image holding one so the next indexing pass measures them. A
peer's unmeasured shard faces are not adopted, or a sync would write
that marker back.
`collections_for_image` answers this for one photograph and has no
counts, which is enough to badge a cell and not enough to offer a
removal: with forty selected and three of them in "Iceland", a sheet
that says only "Iceland" invites the user to take all forty out of a
collection thirty-seven were never in. `membership_of` returns the
count alongside the name so the row can say "3 of 40".
Chunked over the image list rather than one `IN (...)`, because the
list is a selection and a select-all makes it as large as the library —
past SQLite's bound-parameter cap on exactly the gesture most likely to
produce it. Counts are summed across chunks, so the answer is the one
the unchunked query would have given.
Smart collections are excluded by construction: they have no member
rows, so there is nothing a removal could do.
Judgements only ever travelled outward. A rating went to the catalog and to the
photograph's sidecar, the sidecar reached the server, and there it stopped: the
scan indexes files, `derived_sync` exchanges thumbnails, face shards and
collections, `dr_catalog::merge` reconciles everything in a catalog except
`versions.rating` and `versions.flag`, and the one sidecar reader that existed
ran when a single photograph was opened in develop and handed its answer to the
develop graph. `JobKind::ReadSidecar` was declared for exactly this when the job
queue was written and was never enqueued or handled anywhere.
The grid draws `versions.rating`. So a day of culling on the tablet could not
reach the laptop by any path the application had, and the laptop's catalog says
so plainly: 23,568 images, one of them judged.
`pull_sidecars` closes it, off the back of work the scan already does.
`dr_sync::scan` reports the `.drsc` files it meets in listings it was making
anyway — no extra request, and a directory whose ETag is unchanged is still
pruned before it is listed at all. A new `sidecars` table records the ETag of
each one this device has taken in, so the fetch is one GET per sidecar that
genuinely changed rather than one per photograph. A library nobody has edited
costs nothing.
The judgement is taken rather than maximised. The sidecar is the authoritative
store and the fuse has already settled any contest between devices on
`revision`, so lowering a rating from four to one on the tablet lowers it here —
taking the larger would have refused every demotion the photographer ever made,
which is most of what a second pass over a shoot is. A zero is the exception: it
means *never judged*, not "judged zero", so a sidecar carrying none cannot erase
a star this device holds. That is `merge_judgement`'s asymmetry and it carries
the same known cost — clearing a rating does not propagate.
A sidecar names a stem, so both halves of a RAW-and-JPEG pair are judged: they
are one photograph (FR-CAT-11) sharing one document, and judging only one of
them would leave the grid disagreeing with itself over which it drew. The `LIKE`
that finds them is a filter, not the decision — `sidecar_path` is applied to
every candidate, because a folder is entitled to contain a `%` and a rating
landing on the wrong frame would be silent and permanent.
Failing to read one is not a failure to scan: the ETag goes unrecorded, the
ratings already here stay where they are, and the next scan tries again. The
count is reported to the status line as well as the log, because a grid that
silently gains three hundred stars is indistinguishable from one that has gone
wrong — and because while this number was structurally zero there was nothing to
tell the photographer their cull had not arrived.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A version's uuid is the identity a cross-device merge keys on, and it was
minted at random, per catalog, per image. Two devices indexing one Nextcloud
library therefore held two different uuids for the same photograph — so the
sidecar they shared collected a `default = 1` block each, `Version::merge` was
never handed a matching pair to reconcile, and an afternoon's culling on the
tablet did not exist as far as the laptop was concerned.
`crate::merge` has said so in a comment since it was written: version uuids do
not reconcile across devices, a uuid-keyed join unions nothing, so keywords are
landed on the local default version instead. It named the problem and worked
around it. `rating`'s own comment asserted the opposite — that generating the
uuid here was what made it a cross-device identity — and `library::amend`
repeated the claim. Uniqueness was never the difficulty; agreement was.
`derived_version_uuid` computes it from `oc:fileid` instead. The server assigns
that integer, every client pointed at the library sees the same one, and it
survives a server-side rename and move — the three properties that already made
`ASSIGN_BY_FILE_ID` prefer it to a content hash. The layout is a UUIDv8 (RFC
9562, an application-defined form) carrying all sixty-four bits verbatim across
the variable fields with a fixed tag in the node field, so the mapping is
injective by construction rather than by a hash's good behaviour, and a uuid in
a sidecar can be read back to the file it belongs to by eye.
A library with no server behind it has no shared identity to derive and keeps a
generated one. The split is still reachable there if the folder is synced by
something else; `Sidecar::fuse_default_versions` repairs that case rather than
preventing it.
Deriving it for new rows alone would have fixed nothing — every image in an
existing library already has a version, so every one of them would have carried
on writing to its own rival identity. `align_default_version_uuids` moves them,
and runs from `schema::backfill` on every catalog open. It selects on the tag
in SQL, so a catalog already realigned matches no rows and writes nothing, and
it declines rather than fails where a virtual copy already holds the target.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Opening the application was always a fresh arrival at the beginning of
the library, whatever you had been doing when you closed it.
What is written down is the view, the scope, the rating filter and the
photograph on screen -- the open one in develop, the first visible one in
the grid. Not just a scroll position: a position without the filter that
produced it names a row of a list that no longer exists. Restoring them
has an order for the same reason -- scope, then filter, then position,
then the view -- because each step changes what an ordinal *means*.
Addressed by remote path and collection UUID, never by an ordinal or a
row id. `images.id` and `collections.id` are local to one catalog, and a
grid ordinal is local to one ordering; a record naming either would land
somewhere arbitrary on a second device and after any filter change on
this one. Where the ordinal is needed, `library::ordinal_of_path`
computes it through the grid's own `ORDER BY`, taken verbatim by a window
function rather than spelled a second time as an inequality -- which is
the mistake `grid_order_for` already warns about, and which a manually
ordered collection would make unreadable.
Every failure degrades rather than reports. A collection this device has
not merged leaves the scope at the whole library; a photograph that has
since been deleted falls back to when it was taken, which puts the grid
in the right week; a torn file yields no place and the library opens at
the top. Reopening develop is the one thing that requires an exact match,
because a canvas on a path that no longer resolves is a filename over an
empty frame.
The record lives in `dr-types` beside `Settings` and the store lives here
beside `SettingsStore`, for the reason `dr-types`' manifest gives: a JSON
serialiser in `core/` would be paid for by every crate there. Two files
and two lifetimes, though -- resetting preferences must not forget where
you were.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
FR-CULL-8 said detection runs against the thumbnail or proxy tier and
never a full decode, and faces.md §5 said the aligned crop is sampled
from that same proxy. Both are wrong in the same place: they treat
detection and cropping as one resolution problem when they are two, with
opposite answers.
Detection does not care. §4.1 fixes the graph's input at 640x640 and
letterboxes whatever arrives, so a face filling 2% of the frame reaches
the model at 12px whether the buffer handed over is 1024px or 6000px.
Every pixel above the detector's own input is discarded before inference.
The crop cares about nothing else. §5's warp produces the fixed 112x112
ArcFace sees, so source resolution converts directly into whether those
112 pixels were photographed or interpolated. Reading crop_px across the
18,671 faces the proxy-tier implementation stored: 47.3% were upsampled
to reach the embedder, 314 of them by more than 2x, the smallest from 34
source pixels. An upsampled crop does not fail loudly -- it yields a
confident embedding of detail that was never there, and the damage
appears three stages later as clusters that will not separate.
So FR-CULL-8 now specifies four stages with the resolutions named
separately: render native through FR-EXP-9's pipeline, downscale for the
detector, map boxes and landmarks back to native, crop and align from
the native render. The affordability the old rule bought is met instead
by when the pass runs -- background, preempted, resumable -- and the
requirement says plainly what it now costs on a remote library: the
original rather than FR-NC-3's byte range, 412 GB across the reference
library's 19,107 images, so a whole-library pass is a transfer under
FR-NC-6 rather than something that may start on its own.
MIN_CROP_EDGE replaces the MIN_DETECT_EDGE this branch briefly had. Same
number, guarding the quantity that turned out to matter.
faces.md §7b records both measurements, and marks the second as
unexplained rather than dressing it as a finding. Grouped by the buffer
detection ran against, faces per image was 0.078 at 1024 or below and
1.82 at 2048 or better, controlled for file type and size. That gap is
real and reproducible and I cannot account for it, because the letterbox
above says detector input should not matter. M4 is where it gets
settled. The crop measurement does not depend on it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The Identity screen said 4,593 images were left to index and stayed
there for hours across repeated runs, which is what a stuck job looks
like. It was not stuck. 4,424 of those 4,593 are shadowed -- the JPEG
half of a RAW+JPEG pair -- and no sweep will ever index one, because
every work list is built on VISIBLE, which excludes them. They are not
separate photographs and the grid does not show them either.
But faces::coverage counted them: its denominator was "images WHERE
trashed_at IS NULL", with no shadowed_by clause. So the outstanding
figure had a floor of 4,424 that no amount of work could bring down, and
Coverage::is_complete could never once return true no matter how
completely the library had been indexed. A progress number that cannot
reach its own target is worse than no progress number.
The fix is to count the population the sweeps actually draw from, in all
three places that were describing it differently: coverage's denominator
and its indexed join, and audit's split of the outstanding set, which
had the same gap and fed the same status line.
On the reference library the denominator goes from 23,531 to 19,107 and
outstanding from 4,593 to 169 -- the second of which is a number the
user can watch go down, and which turns out to be a real and separate
fetch failure worth chasing.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The floor added in the previous commit stops this happening again; it
does nothing about the 1,824 images in the reference library that
already carry a face_index row written against a proxy of 1024 or less.
Those rows are why the damage is permanent rather than merely past. The
work list is "images with no row for this model", so an image examined
against a 1024px proxy -- 0.078 faces per image, nine in ten finding
nothing -- is indistinguishable from one examined properly, and no
later pass will ever offer it to the detector again.
V12 deletes exactly those markers, and nothing else. The faces those
runs did find stay in place and keep drawing the People screen until a
better pass replaces them, and record_detections re-attaches the user's
confirmed names across that replacement by box overlap, so a library
somebody has spent an evening naming does not lose that evening. The
cost is a re-fetch of the affected images.
Deleting the marker rather than teaching the work-list query to select
on source_edge, which was the other option and is worse. A standing
`source_edge < floor` predicate never lets go: an image whose largest
embedded preview is genuinely smaller than the floor would be re-fetched
on every sweep for ever, because the next pass cannot do any better than
the last one did. A one-off deletion gives each affected image exactly
one more attempt through the good path and then lets the ordinary
"has a row" rule settle it.
The threshold is written out in the SQL instead of referring to
dr_face::MIN_DETECT_EDGE. A migration has to keep meaning what it meant
when it ran; binding it to a constant someone may raise later would
quietly change what an old catalog gets migrated to.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
NFR-R6 asks for an integrity check at startup and two offers behind it, and
none of it existed. `PRAGMA integrity_check` appeared nowhere in the tree,
`Catalog::open` was `open` → `configure` → `migrate` → `backfill` and nothing
else, and corruption therefore surfaced as whatever rusqlite error the first
unlucky query happened to produce — "database disk image is malformed"
attached to a thumbnail refresh, elided into a 34px banner, over an empty
grid saying "No images found · Check the library folder". Two messages that
disagreed, and no way forward but deleting catalog.sqlite by hand.
The property that makes the second offer real was already here and load-
bearing: the catalog is an index, not a source of truth, rebuildable from
sources plus sidecars (invariant §5.2.4, cited by schema.rs, trash.rs and
lib.rs). And sync.rs already knew how to take a coherent snapshot of a WAL
database. What was missing was the check, the type, and the conversation.
Four pieces:
**The type.** `CatalogError::Corrupt`, and — the part that makes it worth
having — a hand-written `From<rusqlite::Error>` that classifies rather than
wraps. `SQLITE_CORRUPT` and `SQLITE_NOTADB` become `Corrupt` wherever they
arise, so a background job that trips over the damage first reports the same
thing the startup check would have. `SQLITE_IOERR` and `SQLITE_BUSY`
deliberately do not: a dropped network mount is a different problem, and
telling someone to rebuild their index would be a wrong answer delivered
confidently.
**The check.** `Catalog::open_verified`, `quick_check` before the open rather
than after, because opening runs migrations and a damaged catalog with an
intact header would otherwise have structure rewritten on top of structure
that is already wrong. Bound to `open_verified` and not to `open`: the check
reads every page, which is affordable once at startup where a user can answer
a question, and not affordable on the dozens of opens a session's background
tasks make.
**The backup.** NFR-R2's second clause, taken between `configure` and
`migrate` in `Catalog::open`. A migration is the one routine operation that
rewrites table structure, so it is the likeliest way this file becomes
unreadable, and it is the last moment the pre-migration state exists to be
copied. Three generations, through SQLite's backup API after a TRUNCATE
checkpoint — never `fs::copy`, which on a WAL database backs up a state older
than the catalog and possibly torn. A failure to take the copy is logged, not
raised: a full disk must not be what makes a library unopenable.
**The conversation.** The first line of the dialogue is that the photographs
and the edits are safe, before the diagnosis, because that is the question the
user is actually asking. Then the two offers, which are *not* interchangeable
and are not presented as if they were: a restore keeps collections, and a
rebuild cannot, because a manual collection is a set of images assembled by
hand and nothing in the filesystem records it (docs/catalog.md §8.1). The
labels say so, and the rebuild does not take the affirmative styling while a
restore is on the table.
One thing that is a fix rather than a feature: `show_catalog_now` now gates
the scan. `Catalog::open` succeeds on a file whose header survived, so the
scan that used to start immediately afterwards would write folder ETags and
image rows into damaged pages in the seconds while the user was still reading
the question — turning a file that had a backup into one where the backup is
the only copy left.
Restore also deletes the damaged catalog's `-wal` and `-shm`. That step is
easy to leave out and fatal to leave out: a journal belonging to the old file,
sitting beside the new one under the same name, is replayed into it on the
next open. That is not a restore, it is a fresh corruption with the evidence
gone.
Tested by corrupting a fixture catalog — 500 images and a collection, then
every page past the second overwritten — and driving both branches. The
restore is asserted on the collection, because a collection is precisely what
distinguishes the two paths; the rebuild on the damaged file being kept and
the next open producing an empty catalog at the current schema. Plus the
`SQLITE_NOTADB` presentation, a damaged backup being refused rather than
installed, and a v1 catalog whose pre-migration backup comes back reading
v1 rather than v11.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
FR-CAT-13 asks for standard XMP sidecars read and written — ratings, colour
labels, keywords and hierarchical subjects, title, description, copyright, GPS,
in `xmp:`/`dc:`/`lr:` schemas — so that other tools interoperate. Its one tag
was the module header of `dr-catalog/src/keywords.rs`.
That module stores keywords in SQLite. It names `dc:subject` twice, both times
in prose explaining why a keyword's text is the fact rather than its row id,
which is a good reason to have written it that way and not evidence of an XMP
implementation. Nothing in the tree parses or emits XMP: `dr-export`'s metadata
module writes EXIF and says in its own header that IPTC and XMP are named by
FR-EXP-8 and neither is read.
`dr-preset-xmp` is the crate whose name most invites the mistake. It reads
Lightroom `.xmp` *presets* — develop settings — under FR-DEV-6, and knows
nothing about the metadata schemas FR-CAT-13 is about.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The first compile this branch had. One borrow error, in the four-thread
contention test: the `Runner` was the block's tail expression, and a
tail's temporaries are dropped after the block's locals, so it outlived
the `conn` it borrowed. Bound to a local, with the ordering rule written
down beside it -- it is exactly the shape someone tidies back.
Everything else stood: clippy clean at -D warnings, and all 18 runner
tests pass, including the four-thread four-connection claim and the
`UPDATE ... RETURNING` rewrite the author flagged as the riskiest line
in the diff.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`jobs` has been a complete durable work queue since the catalog was
written, and nothing has ever taken a job out of it. `claim_next`,
`complete`, `fail` and `recover_orphaned` had no callers outside their own
tests; `enqueue` had three. So the table grew one row per photograph and
kept it forever, and FR-PLAT-AND-3's resumability was a property of code
that never ran.
`runner` is the missing half. It owns no thread, no clock and no policy,
and that is the whole design: on Android the process does not decide when
background work may run. WorkManager does, subject to Doze, battery saver
and FR-NC-6's network constraints, and it revokes permission mid-job by
calling onStopped(). So the runner exposes `run_one` — claim, run, record —
and `drain`, which repeats it against a budget, a deadline and a
cancellation flag the host owns. A `Worker.doWork()` with ten minutes calls
drain with a deadline; a desktop idle pass calls it with none. That is the
seam the Android service plugs into, and it needs no Android to test.
Handlers are supplied from above, because the catalog knows what needs
doing and nothing about how: a thumbnail needs a decoder and a fetch needs
a network stack, neither of which belongs under core/dr-catalog. A runner
claims only kinds some handler declares, so a queue holding work this
device cannot do is left alone rather than failed five times.
Four outcomes, and only two of them are the job's fault. Done deletes the
row; Retry backs off; Abandon gives up now, for a failure no retry can fix;
Interrupted releases the claim with its attempt refunded and ends the
drain, because the host stopped rather than the job — five backgroundings
in a row must not mark good work as failed. Process death is the fifth and
cannot report itself, which is what `recover` is for.
Recovery is called from `show_catalog_now`, which is the one place a
catalog is opened for a session and already returns early if one is open.
It has to be exactly once and before any worker starts: there is no owner
column, so a second pass while a worker held a claim would take it away.
The attempt a dead claim consumed is deliberately kept — a job that takes
the process down with it is indistinguishable from one that fails, and the
attempt counter is the only evidence that survives a death.
The tests cover claiming under contention twice over: sequentially across
two connections, and with four threads on four connections against one
catalog on disk, asserting every job ran exactly once. Plus completion,
backoff, giving up, abandoning, interruption, budget, deadline,
cancellation, and a job orphaned by a simulated crash being reclaimed and
run once rather than lost or repeated.
Not wired to a handler yet, and deliberately not: the only enqueue site
the app actually reaches is the remote scan's, whose thumbnails are already
served by the async grid worker, and `walk`'s two sites are reachable only
from the scan_local example. Inventing a handler to make the plumbing look
used is how a requirement comes to read as covered by code that does not
implement it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The claim was a deferred transaction around a SELECT and an UPDATE, and
under a single connection that is fine. Under two it is not what it looks
like: the SELECT takes only a read lock, the UPDATE tries to upgrade, and
in WAL a worker that read the same snapshot as another gets
SQLITE_BUSY_SNAPSHOT on its write. That is not an error a busy handler can
retry away — the fix is to roll back and start over — so the queue was
"safe" only in the sense that the loser failed loudly instead of taking a
job someone else was holding.
`UPDATE jobs SET state = 1, attempts = attempts + 1 WHERE id = (SELECT ...)
RETURNING ...` is one statement and so one implicit transaction that takes
the write lock immediately. Two workers serialise, the loser waits out its
busy timeout, and neither can see a row the other already holds. The
existing tests are unchanged by it, because from one connection the two
forms are indistinguishable — which is exactly why it was never noticed.
The rest is the surface a runner has to have and did not:
- `claim_next_matching` takes only kinds a worker can actually do. Without
it a device with no connector claims `FetchOriginal`, fails it, and pays
five wakeups and five backoffs per photograph to reach a conclusion known
before it started. Filtering after a claim cannot work: the claim has
already marked the row running.
- `abandon` gives up now, for failures no retry can fix. `fail` uses it for
its own MAX_ATTEMPTS branch, so there is one statement that ends a job.
- `release` hands a claim back with its attempt refunded, for a worker that
is being stopped rather than a job that is going wrong. `attempts` stands
in for the owner column the table does not have: it is bumped by every
claim, so a stale worker's release matches nothing and changes nothing.
- `reap_orphan_subjects` deletes jobs whose photograph is gone. Coalescing
keeps the table one row per unit of work and nothing ever shrank it when
the work stopped existing. `ScanFolder` is excluded because its subject
is a folder id, and joining that against `images` deletes by coincidence
of numbering — hence `JobKind::subject_is_image`, and `JobKind::ALL` so
the next kind added cannot quietly fall out of the filter.
- `counts` is the number a foreground service's notification is built from.
One behaviour change worth stating: a kind this build does not recognise is
now parked with an error rather than read as `ExtractMetadata`. The old
`unwrap_or` would have run a job of an unknown kind as some arbitrary known
one, which is worse than not running it at all.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>