The backfill stamp read max(id) of images and versions and max(rowid)
of keywords. None of those tables is AUTOINCREMENT, so SQLite hands a
freed newest id out again: empty the trash of the newest photograph and
scan a new one, or let a local folder's walk delete a renamed file's row
and insert the new name in the same pass, and the new image takes the
old id. max(id) does not move, nor does count(*), and when a newer
version elsewhere keeps max(versions.id) still too, the stamp matched
and the open skipped the backfill.
That row is exactly one that needs it. Neither scan path creates the
default version: scan::persist and walk insert the image and leave the
version, the RAW/JPEG pairing and the keyword terms to the next open.
Skipped, the image went without them until the app restarted, so a
rating or a pulled sidecar judgement had no version to land on and a
JPEG beside its RAW showed twice.
The stamp now carries the newest row's content: the newest image's id,
path, added time and whether it has a version; the newest version's id
and image; the newest assignment's rowid, version and word. Whether the
newest image has a version is the part that cannot be fooled - after a
backfill every image has one, and a row that has just taken a freed id
has none - so the two stamps differ even when the same file comes back
at the same id in the same second. Still one statement: three reverse
rowid scans that stop at the first row, and one probe of versions_image.
An open that skips still costs ~1 ms on the reference catalog copy.
This closes the hole in the stamp itself rather than by a forget() at
each delete site, so a delete path added later, or one in another
process, cannot reopen it. Two tests delete the newest image and insert
another at the freed id on a separate connection, with a newer version
elsewhere holding max(versions.id); both fail against the old stamp.
Catalog::open ran schema::backfill every time, and every worker thread
opens its own connection. A develop landing made five opens, and each
paid the RAW/JPEG pairing, the default-version anti-join over every
image, the uuid pass over every default version and the keyword check:
17 ms of CPU an open on a copy of the reference catalog, ~80 ms a
landing, to confirm that nothing had changed since the open before.
Everything the backfill repairs is a row some write added: an image a
scan inserted, a version or keyword assignment a merge brought in. So
the open now reads a stamp - user_version, max(id) of images and
versions, max(rowid) of keywords, and the file's device and inode - and
skips the backfill when the stamp matches the one recorded at this
path's last backfill in this process. The maxima are each the last page
of a b-tree; an open that skips costs ~1 ms.
The backfill still runs:
- on the first open in a process (nothing recorded yet);
- on any open that migrated the schema, unconditionally;
- after a pull: merge_remote forgets the path, so the next open
backfills even when every incoming row collided and nothing moved;
- when the file is replaced under its name: the inode is in the stamp,
and recovery::set_aside, the first step of a restore and a rebuild,
forgets the path;
- when another process or thread adds rows, because the stamp is read
from the file, not from anything this process did.
The stamp is taken before the backfill, not after. Read after, it would
describe the backfill's own inserts, and could record an image another
connection inserted in between as covered when it was not. Read before,
the worst case is one redundant pass after a backfill that did real work.
Kept in memory rather than in the catalog: a stamp row would need a
table an older build does not have and would travel in the sync
snapshot, where a flag from another device's catalog says nothing about
this one. No schema version bump, so the tablet on 0.16.0 still reads
the snapshot. Tests cover the skip, a scan's new image, a migration, a
pull and a replaced file.