Stopping the enqueue leaves the rows already queued: 23,582 on the
reference catalog, about 1 MB of table and indexes that every query over
jobs pays for.
A migration would be the usual tool and is the wrong one here. A schema
bump makes an older build refuse the synced catalog snapshot, and the
tablet is on 0.16.0. So the rows are dropped at runtime instead, by
jobs::drop_retired over a new JobKind::RETIRED list, from runner::recover
- which already runs exactly once per catalog open, before any worker.
It runs every open rather than once because an older build sharing the
catalog queues them again on its next scan. kind leads the
UNIQUE(kind, subject_id) index, so with nothing left it is one index
probe. Measured on a copy of the reference catalog: 23,582 rows dropped
in 40 ms on the first open, 0.07 ms after.
Thumbnail stays in the enum so its number is never reused for a kind
that would then inherit old rows. The runner tests that call recover
move to a live kind; the jobs.rs tests of queue mechanics never call
it and are unchanged.
Refs #73
`persist` runs after every scan, for every photograph the scan listed. On
a settled library that is the folders whose ETag changed -- a sidecar
written there by a rating is enough -- so one relisted folder of 1,600
images is an ordinary pass, and a first scan is all 24,000.
Per photograph it prepared four statements from their SQL (a folder
lookup, the image upsert, the id read-back, the remote upsert) and then,
after the commit, found the image again by path and enqueued its
thumbnail job as an autocommitting statement of its own -- a commit per
photograph, for rows that were almost all already queued.
Now the statements are prepared once per pass, a folder's id is looked up
once per folder rather than once per photograph in it, and the job is
enqueued inside the transaction with the id already in hand. That also
makes the job atomic with the row it points at, which is what the old
ordering after the commit was trying to guarantee. `jobs::enqueue` uses a
cached statement for the same reason.
persist_bench on a copy of the reference catalog, CPU, best of runs:
largest folder (1,589 images) 102-118 ms -> 10-13 ms
whole library (23,582 images) 1.55-2.19 s -> 188-192 ms
The fingerprint of images, remote, jobs and folders after the run is the
same for both builds.
The claim was a deferred transaction around a SELECT and an UPDATE, and
under a single connection that is fine. Under two it is not what it looks
like: the SELECT takes only a read lock, the UPDATE tries to upgrade, and
in WAL a worker that read the same snapshot as another gets
SQLITE_BUSY_SNAPSHOT on its write. That is not an error a busy handler can
retry away — the fix is to roll back and start over — so the queue was
"safe" only in the sense that the loser failed loudly instead of taking a
job someone else was holding.
`UPDATE jobs SET state = 1, attempts = attempts + 1 WHERE id = (SELECT ...)
RETURNING ...` is one statement and so one implicit transaction that takes
the write lock immediately. Two workers serialise, the loser waits out its
busy timeout, and neither can see a row the other already holds. The
existing tests are unchanged by it, because from one connection the two
forms are indistinguishable — which is exactly why it was never noticed.
The rest is the surface a runner has to have and did not:
- `claim_next_matching` takes only kinds a worker can actually do. Without
it a device with no connector claims `FetchOriginal`, fails it, and pays
five wakeups and five backoffs per photograph to reach a conclusion known
before it started. Filtering after a claim cannot work: the claim has
already marked the row running.
- `abandon` gives up now, for failures no retry can fix. `fail` uses it for
its own MAX_ATTEMPTS branch, so there is one statement that ends a job.
- `release` hands a claim back with its attempt refunded, for a worker that
is being stopped rather than a job that is going wrong. `attempts` stands
in for the owner column the table does not have: it is bumped by every
claim, so a stale worker's release matches nothing and changes nothing.
- `reap_orphan_subjects` deletes jobs whose photograph is gone. Coalescing
keeps the table one row per unit of work and nothing ever shrank it when
the work stopped existing. `ScanFolder` is excluded because its subject
is a folder id, and joining that against `images` deletes by coincidence
of numbering — hence `JobKind::subject_is_image`, and `JobKind::ALL` so
the next kind added cannot quietly fall out of the filter.
- `counts` is the number a foreground service's notification is built from.
One behaviour change worth stating: a kind this build does not recognise is
now parked with an error rather than read as `ExtractMetadata`. The old
`unwrap_or` would have run a job of an unknown kind as some arbitrary known
one, which is worse than not running it at all.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Wires dr-face to dr-catalog: a background sweep that reads the proxy the
grid already built, detects, aligns, embeds and stores, then a clustering
pass that turns those embeddings into suggested people.
Detection runs on the Large thumbnail tier and nowhere else. That is what
makes the feature affordable -- a browsed library has already paid for
its proxies, so face indexing adds no RAW decode that was not already
happening -- and it is why an image whose proxy is missing is skipped
rather than fetched: requesting one here would put face indexing on the
network path FR-CULL-8 keeps it off.
The sweep keeps no cursor. It asks the catalog what is missing, so it
resumes after process death with no repeated work beyond the in-flight
image, and cancelling is dropping the receiver.
recluster writes only the suggested half. Confirmed faces go in as
anchors and come back untouched, and a cluster of one stays nameless --
naming every stray face would fill the People view with noise the user
then has to dismiss.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Library setup as the user described it: pick a folder, choose which RAW
types to look for, scan recursively.
dr-types::FormatFilter the tick-box selection, seeing through VFS
placeholder suffixes so a dehydrated CR2 still
matches as a CR2
dr-sync::scan recursive walk, Depth:1 per directory, pruning
unchanged subtrees where the backend propagates
directory ETags
Verified against nextcloud.tourolle.paris (34.0.2) on a real library:
browse root 32 entries, 98ms
scan PhotosRaw 17,185 RAW files in 334 directories, 34.1s
(7,836 CR2 + 9,349 DNG)
range read 262KB of a 21.5MB DNG in 119ms — 1.22% of the file,
and enough to read "Canon EOS 6D | ISO 100"
That last line is assumption A3 validated on real data. Cataloguing this
library by whole-file fetch would move roughly 370GB; the range path
moves a few MB.
Pruning is capability-gated rather than assumed: with per-entry ETags a
probe costs a request and proves nothing about children, so it is skipped
entirely. A test asserts zero probes in that case.
Still unresolved: /core/preview returns 400 for every parameter
combination tried, including on a JPEG the server reports as having a
preview. Not a request-shape bug — it fails identically bare. Recorded
rather than worked around; ARCH §6.7 already treats server previews as
opportunistic, so nothing depends on it.