Drop retired Thumbnail jobs whenever a catalog is opened

Stopping the enqueue leaves the rows already queued: 23,582 on the
reference catalog, about 1 MB of table and indexes that every query over
jobs pays for.

A migration would be the usual tool and is the wrong one here. A schema
bump makes an older build refuse the synced catalog snapshot, and the
tablet is on 0.16.0. So the rows are dropped at runtime instead, by
jobs::drop_retired over a new JobKind::RETIRED list, from runner::recover
- which already runs exactly once per catalog open, before any worker.

It runs every open rather than once because an older build sharing the
catalog queues them again on its next scan. kind leads the
UNIQUE(kind, subject_id) index, so with nothing left it is one index
probe. Measured on a copy of the reference catalog: 23,582 rows dropped
in 40 ms on the first open, 0.07 ms after.

Thumbnail stays in the enum so its number is never reused for a kind
that would then inherit old rows. The runner tests that call recover
move to a live kind; the jobs.rs tests of queue mechanics never call
it and are unchanged.

Refs #73
This commit is contained in:
2026-09-26 13:15:18 -04:00
parent 5fcd3752d7
commit 6e67ef4467
3 changed files with 108 additions and 17 deletions
+57 -11
View File
@@ -77,7 +77,7 @@ pub enum Outcome {
/// Something that can actually do the work a job describes.
///
/// The catalog knows what needs doing and nothing about how — a thumbnail
/// The catalog knows what needs doing and nothing about how — face detection
/// needs a decoder, a fetch needs a network stack, and neither belongs under
/// `core/dr-catalog` (ARCH §4.1: calls go downward). So the queue lives here
/// and the handlers are supplied from above.
@@ -187,17 +187,19 @@ pub struct Recovered {
pub reclaimed: usize,
/// Jobs deleted because the photograph they name no longer exists.
pub reaped: usize,
/// Jobs deleted because their kind is retired ([`JobKind::RETIRED`]).
pub retired: usize,
}
impl Recovered {
pub fn did_anything(&self) -> bool {
self.reclaimed > 0 || self.reaped > 0
self.reclaimed > 0 || self.reaped > 0 || self.retired > 0
}
}
/// Ready the queue for a fresh run, before any worker touches it.
///
/// Two distinct cleanups, and both are startup-only:
/// Three distinct cleanups, and all are startup-only:
///
/// - **Reclaim.** A `Running` row has no owner; the process that claimed it is
/// gone. On Android that is a routine morning, not a crash (FR-PLAT-AND-3).
@@ -205,19 +207,27 @@ impl Recovered {
/// process down with it three times running should not be retried forever,
/// and the attempt counter is the only evidence of that we have.
/// - **Reap.** Jobs naming an image the catalog no longer has. A library that
/// has been culled leaves thumbnail jobs for photographs that were deleted
/// has been culled leaves jobs for photographs that were deleted
/// months ago, and every one of them would be claimed, run and failed.
/// - **Retire.** Rows of a kind nothing enqueues or claims any more
/// ([`jobs::drop_retired`]). Here rather than in a migration so that no
/// schema bump locks an older device out of the synced catalog, and every
/// time rather than once because an older build sharing the catalog will
/// queue them again.
///
/// Reclaim runs first so its count is the honest number of interrupted jobs,
/// Retiring runs first, so the other two never touch rows about to go.
/// Reclaim runs next so its count is the honest number of interrupted jobs,
/// before reaping removes whichever of them pointed at nothing.
///
/// **Call this exactly once per catalog, at startup.** It cannot distinguish a
/// job a dead process was holding from one a live runner is holding right now,
/// because there is no owner column — the queue is durable, not distributed.
pub fn recover(conn: &Connection) -> Result<Recovered, CatalogError> {
let retired = jobs::drop_retired(conn)?;
Ok(Recovered {
reclaimed: jobs::recover_orphaned(conn)?,
reaped: jobs::reap_orphan_subjects(conn)?,
retired,
})
}
@@ -729,7 +739,7 @@ mod tests {
// which is the window a durable queue exists to survive: no `complete`,
// no `fail`, just a row marked `Running` with nobody holding it.
let c = db();
queued(&c, JobKind::Thumbnail, 1);
queued(&c, JobKind::ContentHash, 1);
// The dead process. It claimed the job and never came back.
let claimed = jobs::claim_next(&c, 0).unwrap().expect("claimable");
@@ -738,7 +748,7 @@ mod tests {
// A fresh runner, before it starts, finds the queue empty — the row is
// `Running` and no claim will touch it.
let seen = Arc::new(Mutex::new(Vec::new()));
let mut runner = Runner::new(&c).with(recording(vec![JobKind::Thumbnail], seen.clone()));
let mut runner = Runner::new(&c).with(recording(vec![JobKind::ContentHash], seen.clone()));
assert_eq!(
runner.drain_all(0).unwrap().ran(),
0,
@@ -766,7 +776,7 @@ mod tests {
// the only evidence we keep across a death. Without this a poison-pill
// job would be reclaimed and re-run forever.
let c = db();
queued(&c, JobKind::Thumbnail, 1);
queued(&c, JobKind::ContentHash, 1);
for _ in 0..MAX_ATTEMPTS {
jobs::claim_next(&c, 0).unwrap().expect("claimable");
@@ -782,13 +792,49 @@ mod tests {
assert_eq!(state, JobState::Failed as i64);
}
#[test]
fn recovery_drops_retired_kinds_every_time_and_nothing_else() {
// What 0.16.0 left behind: a thumbnail job per photograph that nothing
// would ever claim, beside live work that must survive.
let c = db();
queued(&c, JobKind::Thumbnail, 1);
queued(&c, JobKind::Thumbnail, 2);
enqueue(
&c,
JobKind::DetectFaces,
Some(1),
Priority::Background,
None,
)
.unwrap();
let first = recover(&c).unwrap();
assert_eq!(first.retired, 2);
assert!(first.did_anything());
let kinds: Vec<i64> = c
.prepare("SELECT kind FROM jobs")
.unwrap()
.query_map([], |r| r.get(0))
.unwrap()
.map(Result::unwrap)
.collect();
assert_eq!(kinds, vec![JobKind::DetectFaces as i64]);
// An older build opening the same catalog queues them again on its
// next scan. The next open by this one clears them again.
enqueue(&c, JobKind::Thumbnail, Some(1), Priority::Background, None).unwrap();
assert_eq!(recover(&c).unwrap().retired, 1);
assert_eq!(recover(&c).unwrap(), Recovered::default());
}
#[test]
fn recovery_drops_jobs_whose_photograph_is_gone() {
// A culled library leaves thumbnail jobs for images deleted months
// ago. Every one would be claimed, run and failed.
let c = db();
queued(&c, JobKind::Thumbnail, 1);
queued(&c, JobKind::Thumbnail, 2);
queued(&c, JobKind::ContentHash, 1);
queued(&c, JobKind::ContentHash, 2);
c.execute("DELETE FROM images WHERE id = 2", []).unwrap();
let recovered = recover(&c).unwrap();
@@ -804,7 +850,7 @@ mod tests {
#[test]
fn a_quiet_startup_recovers_nothing() {
let c = db();
queued(&c, JobKind::Thumbnail, 1);
queued(&c, JobKind::ContentHash, 1);
assert_eq!(recover(&c).unwrap(), Recovered::default());
assert!(!recover(&c).unwrap().did_anything());
}