Files
DarkRoom/core/dr-catalog
dtourolleandClaude Opus 5 d63872e5a9 Make a claim one statement, and give the queue what a runner needs
The claim was a deferred transaction around a SELECT and an UPDATE, and
under a single connection that is fine. Under two it is not what it looks
like: the SELECT takes only a read lock, the UPDATE tries to upgrade, and
in WAL a worker that read the same snapshot as another gets
SQLITE_BUSY_SNAPSHOT on its write. That is not an error a busy handler can
retry away — the fix is to roll back and start over — so the queue was
"safe" only in the sense that the loser failed loudly instead of taking a
job someone else was holding.

`UPDATE jobs SET state = 1, attempts = attempts + 1 WHERE id = (SELECT ...)
RETURNING ...` is one statement and so one implicit transaction that takes
the write lock immediately. Two workers serialise, the loser waits out its
busy timeout, and neither can see a row the other already holds. The
existing tests are unchanged by it, because from one connection the two
forms are indistinguishable — which is exactly why it was never noticed.

The rest is the surface a runner has to have and did not:

- `claim_next_matching` takes only kinds a worker can actually do. Without
  it a device with no connector claims `FetchOriginal`, fails it, and pays
  five wakeups and five backoffs per photograph to reach a conclusion known
  before it started. Filtering after a claim cannot work: the claim has
  already marked the row running.
- `abandon` gives up now, for failures no retry can fix. `fail` uses it for
  its own MAX_ATTEMPTS branch, so there is one statement that ends a job.
- `release` hands a claim back with its attempt refunded, for a worker that
  is being stopped rather than a job that is going wrong. `attempts` stands
  in for the owner column the table does not have: it is bumped by every
  claim, so a stale worker's release matches nothing and changes nothing.
- `reap_orphan_subjects` deletes jobs whose photograph is gone. Coalescing
  keeps the table one row per unit of work and nothing ever shrank it when
  the work stopped existing. `ScanFolder` is excluded because its subject
  is a folder id, and joining that against `images` deletes by coincidence
  of numbering — hence `JobKind::subject_is_image`, and `JobKind::ALL` so
  the next kind added cannot quietly fall out of the filter.
- `counts` is the number a foreground service's notification is built from.

One behaviour change worth stating: a kind this build does not recognise is
now parked with an error rather than read as `ExtractMetadata`. The old
`unwrap_or` would have run a job of an unknown kind as some arbitrary known
one, which is worse than not running it at all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 23:31:27 +02:00
..
2026-08-29 12:33:23 +02:00