The reference catalog held 23,582 Thumbnail jobs, one per image, and
every scan re-coalesced all of them. Nothing has ever claimed that kind:
no JobHandler is registered for it on desktop or Android, and
dr_catalog::sync never merges another device's jobs in.
Thumbnails are owed by the store, not the queue. The grid's worker and
the thumbnail sweep both find their work by asking ThumbStore what it
lacks, and the store is shared between devices, so it is the only record
that knows another device already made one. A queue row was a second,
staler copy of that debt that grew with the library and was read by
nothing.
persist still writes the images and their remote identities in the one
transaction; it just no longer adds a row to jobs for each of them. The
two tests that asserted the rows existed become one that asserts a
repeated scan queues nothing.
Refs #73
When a scan pull takes in sidecars another device wrote, `apply_judgement`
finds the photographs each one describes with
`source_ref LIKE '{stem}.%'`. SQLite's LIKE folds ASCII case, and nothing
indexes `source_ref` case-insensitively, so every lookup read all 24,000
names of the root through the `(root_id, source_ref)` index: 1.5-2 ms per
sidecar, 520-690 ms for the 342 `.drsc` files the reference catalog has
read. Another device culling a shoot is several hundred of them.
Every name beginning `{stem}.` lies in the half-open range
`[{stem}., {stem}/)` -- `/` is the byte after `.` -- which the unique key
serves as a seek. The rows LIKE matched beyond these differed only in
case, and the check that decides, `sidecar_path(source) == sidecar`, has
always compared exactly and refused them; the escaping of `%` and `_`
goes too, since a range has no wildcards.
persist_bench: 342 lookups 634-691 ms -> 9 ms CPU. Checked against the
reference catalog directly as well: for all 18,430 distinct sidecar
names its images imply, the range and the old LIKE, each filtered by
`sidecar_path`, pick the same photographs. A test pins the neighbours of
the range: a case variant, a longer stem, a subfolder named like the
stem, and a folder whose name holds `%` and `_`.
The XMP reader's LIKE (`xmp_sync::images_for`) is left alone: its check
is case-insensitive, so the range would not be a superset there, and it
only runs when the exact darktable-style name is not found.
`persist` runs after every scan, for every photograph the scan listed. On
a settled library that is the folders whose ETag changed -- a sidecar
written there by a rating is enough -- so one relisted folder of 1,600
images is an ordinary pass, and a first scan is all 24,000.
Per photograph it prepared four statements from their SQL (a folder
lookup, the image upsert, the id read-back, the remote upsert) and then,
after the commit, found the image again by path and enqueued its
thumbnail job as an autocommitting statement of its own -- a commit per
photograph, for rows that were almost all already queued.
Now the statements are prepared once per pass, a folder's id is looked up
once per folder rather than once per photograph in it, and the job is
enqueued inside the transaction with the id already in hand. That also
makes the job atomic with the row it points at, which is what the old
ordering after the commit was trying to guarantee. `jobs::enqueue` uses a
cached statement for the same reason.
persist_bench on a copy of the reference catalog, CPU, best of runs:
largest folder (1,589 images) 102-118 ms -> 10-13 ms
whole library (23,582 images) 1.55-2.19 s -> 188-192 ms
The fingerprint of images, remote, jobs and folders after the run is the
same for both builds.
Two benches for reading side by side before and after a change, against a
copy of a real catalog, in the manner of identity_bench:
- `dr-catalog --example catalog_bench CATALOG [FACES_DIR]` times
`Catalog::open` and the backfill inside it step by step, the upload
snapshot, a merge of the catalog with a copy of itself, and the face
shard export and import in the steady state where nothing is new.
- `persist_bench`, an ignored test in dr-ui's scan module because
`persist` and `apply_judgement` are private to it, replays the
catalog's own rows through `persist` (the largest folder, and the whole
library) and looks up every `.drsc` sidecar the catalog has read. It
works on a scratch copy and prints a fingerprint of what `persist` left,
so two builds can be shown to agree.
Both print best, median and CPU time; the CPU figure is the one to compare
while other builds share the machine.
Colour labels could be read from a Lightroom sidecar and queried by the
selector, but nothing drew one or set one, so the only labels a library
held were ones another program had written.
Every mark carries its label's initial on its colour — R, Y, G, B, P —
so a label is read without telling red from green, which is what
NFR-A11Y-3 asks of colour labels by name. A grid cell shows the mark
before its filename. In the grid, 6, 7, 8 and 9 set red, yellow, green and
blue as Lightroom's keys do, on the photograph under the pointer or on
the selection by the rule the star keys follow; the same key again takes
the label off, and over a mixed selection it sets it on all. The
selection bar gains Label, which opens the six choices — each a mark and
a name — and purple, which has no key, is there. In develop the top bar
says "Label: Green" beside the mark, opens the same choices, and 6-9
label the open photograph.
Each gesture is one catalog transaction, then the grid, the counts and
both sidecars are written as a rating's are. The filter bar gains a chip
per label, its mark and its name with a count, one at a time; the filter
is one SQL term, travels in the place record, and "All" clears it.
A rating and a flag are written to DarkRoom's sidecar as well as the
catalog, because the catalog is a disposable index and the sidecar is
how a judgement reaches the photographer's other devices. A label had no
place there, so once labels could be set, one would have lived only in
the catalog of the device it was set on and gone with it.
The sidecar version now carries `label` (0 none, 1-5 as the catalog
codes it), written only when set. It merges under the rating's rule, so a
device that never labelled a frame cannot clear another device's label,
and a code this build does not know reads as none rather than as some
other colour. A judgement write carries the catalog's label with the
stars, and the scan takes a sidecar's label into the catalog when it has
one. An older build keeps the line as an unknown key and writes it back.
library.rs was 7,729 lines wiring together everything "open a remote
library" touches: scanning, pulling other devices' judgements out of
sidecars found along the way, writing local edits back out to the
sidecar outbox, pushing/reloading XMP by hand, fetching and prefetching
thumbnails and originals, generating thumbnails locally, the metadata
and thumbnail background sweeps, on-disk paths for the catalog and
model files, and reading the grid's cells, spans and rating filter.
Same motivation as the develop.rs split (docs/dev/code-health.md CH-1):
a pure, no-behaviour-change move into one file per area, each under
about 1,500 lines.
Tracing actual call sites rather than trusting the file's physical
layout mattered here: `persist`, `load_folder_etags`, `pull_sidecars`,
`load_sidecar_etags`, `record_sidecar_read` and `apply_judgement` sit
textually beside the XMP push/reload functions but are called only
from `run_scan` (pulling a device's own past judgements out of the
sidecars a scan just walked), so they went to scan.rs and not xmp.rs.
`cells` came out at over 1,800 lines once its tests moved with it and
split further into cells.rs (windowed reads, trash, ordinals) and
spans.rs (collection scope, manual reordering, the capture-time
histogram) -- ten submodules rather than the nine first planned.
Previously-private items reached from a sibling module became
`pub(super)`, narrower than the whole-crate reachability one file gave
them. Tests moved with the code they test; the two test fixtures used
across more than one file (`scanned`, and develop.rs's
`session_with_a_left_half_subject` in the matching commit) joined the
shared `test_support` module alongside the existing `entry`/
`with_images`/`image_ids` helpers. `mod.rs` re-exports every module's
public items under `library::`, including the `pub(crate)`
`test_support` module `repairs.rs` reads its fixtures from, so no file
outside `library` needed a change.
The previous commit split develop.rs the same way; taken alone it left
dr-ui without library.rs, so that intermediate commit does not build on
its own. This one restores it.