Commit Graph
2 Commits
Author SHA1 Message Date
dtourolleandClaude Opus 5 758436cc28 Keep the runner's borrow alive as long as the connection it reads
The first compile this branch had. One borrow error, in the four-thread
contention test: the `Runner` was the block's tail expression, and a
tail's temporaries are dropped after the block's locals, so it outlived
the `conn` it borrowed. Bound to a local, with the ordering rule written
down beside it -- it is exactly the shape someone tidies back.

Everything else stood: clippy clean at -D warnings, and all 18 runner
tests pass, including the four-thread four-connection claim and the
`UPDATE ... RETURNING` rewrite the author flagged as the riskiest line
in the diff.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 23:48:15 +02:00
dtourolleandClaude Opus 5 846a249156 Drain the queue that nothing has ever drained
`jobs` has been a complete durable work queue since the catalog was
written, and nothing has ever taken a job out of it. `claim_next`,
`complete`, `fail` and `recover_orphaned` had no callers outside their own
tests; `enqueue` had three. So the table grew one row per photograph and
kept it forever, and FR-PLAT-AND-3's resumability was a property of code
that never ran.

`runner` is the missing half. It owns no thread, no clock and no policy,
and that is the whole design: on Android the process does not decide when
background work may run. WorkManager does, subject to Doze, battery saver
and FR-NC-6's network constraints, and it revokes permission mid-job by
calling onStopped(). So the runner exposes `run_one` — claim, run, record —
and `drain`, which repeats it against a budget, a deadline and a
cancellation flag the host owns. A `Worker.doWork()` with ten minutes calls
drain with a deadline; a desktop idle pass calls it with none. That is the
seam the Android service plugs into, and it needs no Android to test.

Handlers are supplied from above, because the catalog knows what needs
doing and nothing about how: a thumbnail needs a decoder and a fetch needs
a network stack, neither of which belongs under core/dr-catalog. A runner
claims only kinds some handler declares, so a queue holding work this
device cannot do is left alone rather than failed five times.

Four outcomes, and only two of them are the job's fault. Done deletes the
row; Retry backs off; Abandon gives up now, for a failure no retry can fix;
Interrupted releases the claim with its attempt refunded and ends the
drain, because the host stopped rather than the job — five backgroundings
in a row must not mark good work as failed. Process death is the fifth and
cannot report itself, which is what `recover` is for.

Recovery is called from `show_catalog_now`, which is the one place a
catalog is opened for a session and already returns early if one is open.
It has to be exactly once and before any worker starts: there is no owner
column, so a second pass while a worker held a claim would take it away.
The attempt a dead claim consumed is deliberately kept — a job that takes
the process down with it is indistinguishable from one that fails, and the
attempt counter is the only evidence that survives a death.

The tests cover claiming under contention twice over: sequentially across
two connections, and with four threads on four connections against one
catalog on disk, asserting every job ran exactly once. Plus completion,
backoff, giving up, abandoning, interruption, budget, deadline,
cancellation, and a job orphaned by a simulated crash being reclaimed and
run once rather than lost or repeated.

Not wired to a handler yet, and deliberately not: the only enqueue site
the app actually reaches is the remote scan's, whose thumbnails are already
served by the async grid worker, and `walk`'s two sites are reachable only
from the scan_local example. Inventing a handler to make the plumbing look
used is how a requirement comes to read as covered by code that does not
implement it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 23:31:47 +02:00