WhatsApp strips every EXIF tag and names the file "WhatsApp Image
2023-06-15 at 07.00.42.jpeg"; Windows Phone, Android cameras and
darktable's import put the date in the name too. Those images sorted
after everything else and were absent from the timeline.
name_dates reads a date (and a time, when one follows) from the file
name, then from the innermost folder that states one. A sequence number
after a date is not read as a time, and a bare year folder is not a date.
EXIF always wins: only examined rows still undated are filled.
The sweep and the metadata repair fill as they mark an image examined,
and the open backfill fills catalogs examined by earlier builds. On the
reference library that takes 274 undated images to 10; the no-op case is
a seek on images_captured, 0.6 ms an open.
A panorama the merge did not make, a stitch from another program or a
phone's sweep, never had a width or height in the catalog, so it could
never be given a wide cell. The header read that dates a photograph now
records its size as it is seen, orientation applied, alongside the date.
A 4:1 composite drawn in one square cell is a strip a few pixels high.
A photograph about twice as wide as it is tall (1.9 and up) now spans
two columns, three from 2.9, with a thumbnail class of its own whose
long edge is sized for that width; on the tablet, or where the columns
are too few to put it beside anything, it takes the whole row.
Rows are computed in one place, library_ui::layout. The grid is a
lattice of slots: each cell is drawn at the slot Rust gives it, and a
wide cell that would not fit in what is left of a row starts the next
one, leaving the gap empty so the grid still reads in capture order.
The scrollbar spans the slots, a scroll reports a slot that the layout
turns back into an ordinal, and scrubs, restores and the cursor go
through the same conversion. Up and down step by rows through the
layout rather than by a row's worth of ordinals; left and right, a
shift-click's run, burst folding and the timeline are ordinal-based
and unchanged.
The window's own read carries each photograph's w and h, so the cells
know their shape with no query per cell. Where the wide ones sit in
the whole list is one query, run when what the grid lists changes or
when a window finds the layout out of date, and a library with no
panorama answers it from a partial index created on first use
(images_wide), not a schema bump. The merge makes the wide thumbnail
for a wide composite along with the others.
A finished merge drained the outbox and started a rescan beside it. The
scan raced the upload: on a folder library the 800 MB copy was still
running when the folder was listed, on Nextcloud the upload takes
minutes, and either way the listing lacked the composite, recorded the
folder's validator, and nothing looked again until the next sync pass.
The job now reports what the catalog needs (the name, the size of the
picture it opens on, the capture time it wrote into the DNG) and the
library writes the row at once, in one transaction, keyed where the scan
will list the file; a composite with no time of its own takes its
sources' earliest. The grid reloads and shows it beside its sources.
The upload then gives the row what only the server knows: after sending
a file the catalog already has a row for, it lists the folder once, takes
the file id the server assigned, and records it (and, once the merge
makes them, the thumbnails waiting beside the payload) under that id.
Every drain now rescans when something landed in the library, after the
upload rather than beside it. A second merge of the same frames is no
longer named over the first: the name is checked against the catalog's
names in that folder, which is all that knows it once the outbox is empty.
Thirty-nine spawn sites in dr-ui, and one in the Android entry point,
called std::thread::spawn or a Builder of their own, and most of the
threads they started were <unnamed> in a panic message or a profiler.
Each now calls executors::spawn with its executor and a role, so the thread is
named <executor>:<role> — net:sync, decode:thumbs, io:catalog-open —
and knows which executor it is on. The three that already set a name
(automation, import, prefetch) keep their name as the role.
Behaviour is unchanged: each job still gets a thread of its own when it
starts, and spawn panics where std::thread::spawn did.
The module's documentation now says how a job is assigned: by what it
spends its time on, so a sweep that fetches bytes and then decodes them
is Decode, and a sidecar write that touches the catalog is Network.
Left as they were: the segmentation and refine workers in masks_ui.rs,
which another change is reworking, and test-only threads.
holds_original, the prefetch worker's check that a neighbour's original
is already cached, opened the catalog for every neighbour it asked about.
Its own comment called it a row check; the open around it was four of
the five opens a develop landing made.
The worker now keeps one Catalog for the batch it is serving, opened at
the first check and reopened only if the batch names another catalog
file. The connection runs in autocommit, so each check still sees what
fetch_original committed in between. fetch_original is unchanged.
With the backfill no longer run on every open, a landing whose
neighbours are all cached goes from five opens to two, and from ~80 ms
of CPU to ~1-2 ms on a copy of the reference catalog.
rustfmt over the files the albums work touched, and the album merge's
incoming row as a named struct rather than an eight-field tuple, which
clippy's type_complexity refused.
Export took a path typed into the settings page, or a folder inside the
library on the server. The first is how exports end up somewhere nobody
looks; the second put JPEGs into the tree a scan catalogues, where they
came back as photographs beside the RAWs they were made from.
The destination is now an album (FR-EXP-10), chosen by name in the
export sheet. Albums are listed under the collections in the sidebar;
"+" there, or "New album…" in the sheet, opens a sheet for its name and
its folder — on this device through the platform's dialogue, or on the
server through the browser with "New folder". A server folder inside
the library is refused, and the sheet says why. Selecting an album
narrows the grid to the photographs behind its files: library::Scope
is Collection or Album, and scope_clause is the one place the two are
spelled, which also retires the two copies of the collection predicate
total_images_scoped and read_cells_scoped had inlined.
A batch resolves the album when it starts, and refuses in words when
none is chosen, it has gone, or its folder is local to another device.
Each item reports the image it came from, and the files written are
recorded against the album in one transaction when the batch ends.
A server album lives outside the library, so its queued uploads are
relative to the account root. That is a third line in the outbox's
.dest record rather than a leading slash, because a record written
before albums may carry a stray slash and must keep the meaning it was
written with.
An export folder set before albums becomes an album called "Exports"
on first open, so upgrading does not lose where exports were going.
The old destination fields stay in ExportSettings so older settings
files still read.
A preset could not choose a film stock. The stock is a choice of material
rather than a parameter, so `Preset` — a map of `op.param = value` — had
nowhere to hold it, and "Portra 400, printed" could not be saved, copied
or shipped as a look. Worse, the film node's own sliders *were*
parameters: a paste moved one stock's exposure and push onto whatever
stock the target was on, and left the target's tables baked from the
values it had just replaced.
A preset now carries a `FilmRef` beside its parameters. It travels under
whichever scope carries the film node, so the stock and its sliders are
never split, and by the replacement rule every other parameter follows:
applied at that scope, a preset without a film develops the target
without one. `Preset::apply` returns the `FilmRebake` it owes, as
`EditGraph::set_state` already did, because this crate cannot bake a
stock; the develop session pays it before recording the step, and the
batch paste writes the stock into each sidecar through `film_for`. The
library file spells it `film =` / `film_print =`, as a sidecar does, and
an older build keeps those lines as ones it does not understand.
`EditState` keeps the film in its own field only: the parameters it
captures leave it out, so one edit has one place to say which stock it
is on.
Second, a preset now has a reach. Replacement is right for a copy of a
whole edit — "make these match" — and wrong for a look: a stock-only
"Portra 400" applied that way would put the photograph's exposure, white
balance and noise reduction back to default. `Reach::Named` replaces only
the operations a preset names (whole operations, so a look that sets the
blacks resets the whites beside them) and the film only if it names one.
Saved edits and the clipboard keep `Reach::Whole`; the line `reach =
named` is written only for the other, so existing libraries write the
same bytes.
The grid's total is read on every scroll reload (`load_window` compares it
to notice a delete). On the reference library it cost 1.3-1.5 ms best-of-50
by catalog_bench, 2-3.6 ms on a busy machine, and the issue measured 4 ms.
`uncollapsed` asked every visible image whether a collapsed burst stands in
for it -- two primary-key probes per image, 19,000 times, on a library with
no bursts at all:
SCAN i USING INDEX images_grid_order
CORRELATED SCALAR SUBQUERY
SEARCH bm USING INTEGER PRIMARY KEY (rowid=?)
CORRELATED SCALAR SUBQUERY
SEARCH be USING INTEGER PRIMARY KEY (rowid=?)
`total_images_filtered` now counts what the filter keeps and subtracts the
frames `bursts::collapsed_away_frames` lists, under the same filter:
SCALAR SUBQUERY: SCAN i USING INDEX images_grid_order
SCALAR SUBQUERY: SCAN bm; SEARCH be ...; SEARCH i USING INTEGER PRIMARY KEY
The second half walks only `burst_members`. Each image is in it at most
once (it is the key), and the filter is applied to both halves, so the
subtraction removes exactly the rows the predicate used to drop. The new
fragment sits beside `not_collapsed_away` in bursts.rs, and a test holds
the two to the same rows with bursts open and closed.
After: 0.3 ms, the same count (19,152). The cells query keeps the predicate:
it is a window with a LIMIT and needs the rows, not their number. The
rated grid count (3.5-4 ms with a one-star filter) is unchanged: its cost
is the rating subquery per image, and changing how `RatingFilter` spells
it changes every grid and timeline query, which is left for its own change.
`library::local_original_count` feeds the "On this device" chip and runs
beside the rating counts on every star keystroke. On the reference library
it cost 1.3-1.4 ms best-of-50 (3 ms on a busy machine) to find 254
originals among 19,000 visible images.
It was a correlated EXISTS per visible image:
SCAN i USING INDEX images_grid_order
SEARCH ic EXISTS USING INTEGER PRIMARY KEY (rowid=?)
`image_cache` holds a row only for what has been fetched, so the question
is driven from it: `i.id IN (SELECT image_id FROM image_cache WHERE
tier_actual >= Original)`, which SQLite plans as the list first and a probe
of `images` by id for each entry:
SEARCH i USING INTEGER PRIMARY KEY (rowid=?)
LIST SUBQUERY 1
SCAN image_cache
`image_id` is the cache's primary key, so each image is in the list at most
once and the count is the one the EXISTS gave (254). After: 0.05 ms. On a
library whose every original is cached this is as much work as before,
which is the proportion the rule asks for.
The count stays on the keystroke path: dropping it there would leave the
chip stale after a background download until something else refreshed it,
and at this cost there is nothing left to save. catalog_bench spells the
query as dr-ui does, so its copy changes with it.
The reference catalog held 23,582 Thumbnail jobs, one per image, and
every scan re-coalesced all of them. Nothing has ever claimed that kind:
no JobHandler is registered for it on desktop or Android, and
dr_catalog::sync never merges another device's jobs in.
Thumbnails are owed by the store, not the queue. The grid's worker and
the thumbnail sweep both find their work by asking ThumbStore what it
lacks, and the store is shared between devices, so it is the only record
that knows another device already made one. A queue row was a second,
staler copy of that debt that grew with the library and was read by
nothing.
persist still writes the images and their remote identities in the one
transaction; it just no longer adds a row to jobs for each of them. The
two tests that asserted the rows existed become one that asserts a
repeated scan queues nothing.
Refs #73
The develop view reported a remote original on its way through the
error message, so it read "Could not load image" over "Downloading…".
It did so on every step along the roll, including a cached frame that
was ready within a tick, so each step flashed the error.
Waiting is now its own state. On the step, the grid's thumbnail of the
photograph stands in at once. Only when a transfer is really on the
wire does it dim under "Not on this device yet", with a line like
"Downloading — 12.4 of 38.0 MB" and a progress bar.
The bytes come from a new RemoteBackend::get_reporting. The Nextcloud
backend overrides it to read the body chunk by chunk; the default
reports once at the end. Progress is kept in the in-flight registry by
path, because a step usually lands on a frame the prefetcher is already
fetching. The catalog's file length stands in when the server sends no
Content-Length.
When a scan pull takes in sidecars another device wrote, `apply_judgement`
finds the photographs each one describes with
`source_ref LIKE '{stem}.%'`. SQLite's LIKE folds ASCII case, and nothing
indexes `source_ref` case-insensitively, so every lookup read all 24,000
names of the root through the `(root_id, source_ref)` index: 1.5-2 ms per
sidecar, 520-690 ms for the 342 `.drsc` files the reference catalog has
read. Another device culling a shoot is several hundred of them.
Every name beginning `{stem}.` lies in the half-open range
`[{stem}., {stem}/)` -- `/` is the byte after `.` -- which the unique key
serves as a seek. The rows LIKE matched beyond these differed only in
case, and the check that decides, `sidecar_path(source) == sidecar`, has
always compared exactly and refused them; the escaping of `%` and `_`
goes too, since a range has no wildcards.
persist_bench: 342 lookups 634-691 ms -> 9 ms CPU. Checked against the
reference catalog directly as well: for all 18,430 distinct sidecar
names its images imply, the range and the old LIKE, each filtered by
`sidecar_path`, pick the same photographs. A test pins the neighbours of
the range: a case variant, a longer stem, a subfolder named like the
stem, and a folder whose name holds `%` and `_`.
The XMP reader's LIKE (`xmp_sync::images_for`) is left alone: its check
is case-insensitive, so the range would not be a superset there, and it
only runs when the exact darktable-style name is not found.
`persist` runs after every scan, for every photograph the scan listed. On
a settled library that is the folders whose ETag changed -- a sidecar
written there by a rating is enough -- so one relisted folder of 1,600
images is an ordinary pass, and a first scan is all 24,000.
Per photograph it prepared four statements from their SQL (a folder
lookup, the image upsert, the id read-back, the remote upsert) and then,
after the commit, found the image again by path and enqueued its
thumbnail job as an autocommitting statement of its own -- a commit per
photograph, for rows that were almost all already queued.
Now the statements are prepared once per pass, a folder's id is looked up
once per folder rather than once per photograph in it, and the job is
enqueued inside the transaction with the id already in hand. That also
makes the job atomic with the row it points at, which is what the old
ordering after the commit was trying to guarantee. `jobs::enqueue` uses a
cached statement for the same reason.
persist_bench on a copy of the reference catalog, CPU, best of runs:
largest folder (1,589 images) 102-118 ms -> 10-13 ms
whole library (23,582 images) 1.55-2.19 s -> 188-192 ms
The fingerprint of images, remote, jobs and folders after the run is the
same for both builds.
Two benches for reading side by side before and after a change, against a
copy of a real catalog, in the manner of identity_bench:
- `dr-catalog --example catalog_bench CATALOG [FACES_DIR]` times
`Catalog::open` and the backfill inside it step by step, the upload
snapshot, a merge of the catalog with a copy of itself, and the face
shard export and import in the steady state where nothing is new.
- `persist_bench`, an ignored test in dr-ui's scan module because
`persist` and `apply_judgement` are private to it, replays the
catalog's own rows through `persist` (the largest folder, and the whole
library) and looks up every `.drsc` sidecar the catalog has read. It
works on a scratch copy and prints a fingerprint of what `persist` left,
so two builds can be shown to agree.
Both print best, median and CPU time; the CPU figure is the one to compare
while other builds share the machine.
Colour labels could be read from a Lightroom sidecar and queried by the
selector, but nothing drew one or set one, so the only labels a library
held were ones another program had written.
Every mark carries its label's initial on its colour — R, Y, G, B, P —
so a label is read without telling red from green, which is what
NFR-A11Y-3 asks of colour labels by name. A grid cell shows the mark
before its filename. In the grid, 6, 7, 8 and 9 set red, yellow, green and
blue as Lightroom's keys do, on the photograph under the pointer or on
the selection by the rule the star keys follow; the same key again takes
the label off, and over a mixed selection it sets it on all. The
selection bar gains Label, which opens the six choices — each a mark and
a name — and purple, which has no key, is there. In develop the top bar
says "Label: Green" beside the mark, opens the same choices, and 6-9
label the open photograph.
Each gesture is one catalog transaction, then the grid, the counts and
both sidecars are written as a rating's are. The filter bar gains a chip
per label, its mark and its name with a count, one at a time; the filter
is one SQL term, travels in the place record, and "All" clears it.
A rating and a flag are written to DarkRoom's sidecar as well as the
catalog, because the catalog is a disposable index and the sidecar is
how a judgement reaches the photographer's other devices. A label had no
place there, so once labels could be set, one would have lived only in
the catalog of the device it was set on and gone with it.
The sidecar version now carries `label` (0 none, 1-5 as the catalog
codes it), written only when set. It merges under the rating's rule, so a
device that never labelled a frame cannot clear another device's label,
and a code this build does not know reads as none rather than as some
other colour. A judgement write carries the catalog's label with the
stars, and the scan takes a sidecar's label into the catalog when it has
one. An older build keeps the line as an unknown key and writes it back.
With the trait in place the claim still meant nothing while every caller
named dr_decode's free functions: a second decoder would have had to be
threaded through the scan, the thumbnail ladder, import, the viewer,
export, merge and repairs at the moment it arrived.
Each of those now takes a &dyn Decoder and reads headers, previews,
orientation and sensor data through it, including the header budget a
remote fetch asks for (header_bytes) and where it finds the embedded
preview (locate_preview). Only the places that start a job name
dr_decode::default(): the thumbnail, sweep and thumbnail-sweep threads,
the viewer's open handlers, and the request structs a job is handed
(BatchRequest, MergeRequest, the import Request, the repairs Toolkit),
so a caller can be given another decoder by changing what it is handed.
The default is rawler through the same free functions as before, so
nothing a user sees changes. The trait gains Debug as a supertrait so
request structs that derive Debug can carry one.
Rating keys in the grid follow darktable's rule: with the pointer over a
photograph outside the selection, 0-5, P, X and U judge that photograph
alone; over one inside it, the whole selection, as before; off the grid,
the selection. The hover is cleared when the grid scrolls, so a key after
a wheel turn cannot judge whatever used to be under the pointer.
Holding F and tapping digits filters by stars: one digit for exactly
that many, two for everything between them, F alone to show every
rating again. The filter gains a ceiling to do it (`max_rating`, one
BETWEEN in the query). The place record carries it, and a record from an
older build reads as having none. The star chips light across a capped
range and the bar says "2-3★ only" beside them.
The grid also takes Ctrl+E and Ctrl+Shift+E for the selection, Ctrl+V to
paste onto it and Ctrl+A to select all. The "Gestures" button is now
"Help", its sheet "Controls and shortcuts", and F1 opens it.
a_second_claim_waits_for_the_first_to_be_released failed on CI: the
second claim came back Some. The waiter thread signalled the main thread
before calling claim, so the main thread could drop the first guard
before the waiter reached the lock. The path was free by then, and the
waiter claimed it outright.
The registry now keeps a test-only count of threads parked in claim,
bumped under the lock just before the condvar wait. The test spins until
that count is one before releasing. The release needs the same lock, so
it can only reach a waiter that is already waiting. Passed 500 runs in a
row.
library.rs was 7,729 lines wiring together everything "open a remote
library" touches: scanning, pulling other devices' judgements out of
sidecars found along the way, writing local edits back out to the
sidecar outbox, pushing/reloading XMP by hand, fetching and prefetching
thumbnails and originals, generating thumbnails locally, the metadata
and thumbnail background sweeps, on-disk paths for the catalog and
model files, and reading the grid's cells, spans and rating filter.
Same motivation as the develop.rs split (docs/dev/code-health.md CH-1):
a pure, no-behaviour-change move into one file per area, each under
about 1,500 lines.
Tracing actual call sites rather than trusting the file's physical
layout mattered here: `persist`, `load_folder_etags`, `pull_sidecars`,
`load_sidecar_etags`, `record_sidecar_read` and `apply_judgement` sit
textually beside the XMP push/reload functions but are called only
from `run_scan` (pulling a device's own past judgements out of the
sidecars a scan just walked), so they went to scan.rs and not xmp.rs.
`cells` came out at over 1,800 lines once its tests moved with it and
split further into cells.rs (windowed reads, trash, ordinals) and
spans.rs (collection scope, manual reordering, the capture-time
histogram) -- ten submodules rather than the nine first planned.
Previously-private items reached from a sibling module became
`pub(super)`, narrower than the whole-crate reachability one file gave
them. Tests moved with the code they test; the two test fixtures used
across more than one file (`scanned`, and develop.rs's
`session_with_a_left_half_subject` in the matching commit) joined the
shared `test_support` module alongside the existing `entry`/
`with_images`/`image_ids` helpers. `mod.rs` re-exports every module's
public items under `library::`, including the `pub(crate)`
`test_support` module `repairs.rs` reads its fixtures from, so no file
outside `library` needed a change.
The previous commit split develop.rs the same way; taken alone it left
dr-ui without library.rs, so that intermediate commit does not build on
its own. This one restores it.