Drain the queue that nothing has ever drained
`jobs` has been a complete durable work queue since the catalog was written, and nothing has ever taken a job out of it. `claim_next`, `complete`, `fail` and `recover_orphaned` had no callers outside their own tests; `enqueue` had three. So the table grew one row per photograph and kept it forever, and FR-PLAT-AND-3's resumability was a property of code that never ran. `runner` is the missing half. It owns no thread, no clock and no policy, and that is the whole design: on Android the process does not decide when background work may run. WorkManager does, subject to Doze, battery saver and FR-NC-6's network constraints, and it revokes permission mid-job by calling onStopped(). So the runner exposes `run_one` — claim, run, record — and `drain`, which repeats it against a budget, a deadline and a cancellation flag the host owns. A `Worker.doWork()` with ten minutes calls drain with a deadline; a desktop idle pass calls it with none. That is the seam the Android service plugs into, and it needs no Android to test. Handlers are supplied from above, because the catalog knows what needs doing and nothing about how: a thumbnail needs a decoder and a fetch needs a network stack, neither of which belongs under core/dr-catalog. A runner claims only kinds some handler declares, so a queue holding work this device cannot do is left alone rather than failed five times. Four outcomes, and only two of them are the job's fault. Done deletes the row; Retry backs off; Abandon gives up now, for a failure no retry can fix; Interrupted releases the claim with its attempt refunded and ends the drain, because the host stopped rather than the job — five backgroundings in a row must not mark good work as failed. Process death is the fifth and cannot report itself, which is what `recover` is for. Recovery is called from `show_catalog_now`, which is the one place a catalog is opened for a session and already returns early if one is open. It has to be exactly once and before any worker starts: there is no owner column, so a second pass while a worker held a claim would take it away. The attempt a dead claim consumed is deliberately kept — a job that takes the process down with it is indistinguishable from one that fails, and the attempt counter is the only evidence that survives a death. The tests cover claiming under contention twice over: sequentially across two connections, and with four threads on four connections against one catalog on disk, asserting every job ran exactly once. Plus completion, backoff, giving up, abandoning, interruption, budget, deadline, cancellation, and a job orphaned by a simulated crash being reclaimed and run once rather than lost or repeated. Not wired to a handler yet, and deliberately not: the only enqueue site the app actually reaches is the remote scan's, whose thumbnails are already served by the async grid worker, and `walk`'s two sites are reachable only from the scan_local example. Inventing a handler to make the plumbing look used is how a requirement comes to read as covered by code that does not implement it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -1936,6 +1936,29 @@ fn show_catalog_now(
|
||||
}
|
||||
};
|
||||
|
||||
// Before anything else can see this catalog, and exactly once per open —
|
||||
// the early return above is what makes it once. The job queue is durable,
|
||||
// so a run that was killed mid-job left its row marked `Running` with
|
||||
// nobody holding it; recovery hands those back to be resumed rather than
|
||||
// lost, and drops jobs naming photographs that have since been deleted.
|
||||
//
|
||||
// Here rather than wherever a runner starts, because there is no owner
|
||||
// column in `jobs`: a second recovery pass while a worker held a claim
|
||||
// would take that claim away from it.
|
||||
match dr_catalog::runner::recover(cat.connection()) {
|
||||
Ok(r) if r.did_anything() => log::info!(
|
||||
"job queue: {} interrupted job(s) resumed, \
|
||||
{} for deleted photographs dropped",
|
||||
r.reclaimed,
|
||||
r.reaped
|
||||
),
|
||||
Ok(_) => {}
|
||||
// Not surfaced. The queue is rebuildable like everything else in the
|
||||
// catalog, and a grid that refuses to open because a background queue
|
||||
// could not be tidied is the worse failure by some way.
|
||||
Err(e) => log::warn!("job queue could not be recovered: {e}"),
|
||||
}
|
||||
|
||||
// The sidebar before the grid, because the grid's badges read collection
|
||||
// membership — the same order the scan's completion uses.
|
||||
crate::collections_ui::refresh_tree(window, coll_ctl, &cat);
|
||||
|
||||
Reference in New Issue
Block a user