Commit Graph
270 Commits
Author SHA1 Message Date
dtourolleandClaude Opus 5 7981718d83 Time the phases of a launch, because the tablet has no profiler
Benchmarks / CPU and I/O (per commit) (push) Successful in 3m8s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 1h19m46s
Build and test / Layer separation (push) Successful in 47s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
Traceability / Requirement traces (push) Failing after 46s
Build and test / Android (aarch64) (push) Successful in 1h1m21s
Two things are now off the launch path and the rest of it is unmeasured.
There is no way to attach a profiler to an Android launch, and the
window that matters — from `android_main` to the first `poll_events` —
is over before anything on the device can be asked a question about it.
So a line in the log file is the only measurement anybody gets.

Three of them: the GPU open, the window build, and the total to the
event loop. The last is the one that matters, because it is the figure
the input dispatcher is counting against — anything approaching five
seconds there is the next ANR whatever the phases above it say.

The GPU open is timed rather than moved. It is a Vulkan instance, an
adapter enumeration and a device request, and on the desktop it cannot
be deferred at all: it selects the Slint backend, and creating a window
selects one for us. On Android it could be, because nothing shares that
device with the compositor (TD-1) — but "could be deferred" is not
"costs enough to be worth deferring", and there is no number yet that
says which. This is the line that will produce one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 18:41:34 +02:00
dtourolleandClaude Opus 5 d48e9f6033 Open the catalog on a worker, so a launch is not a page-by-page read
The second thing standing between `android_main` and the first
`poll_events`, and the one that grows with the library rather than with
the APK.

`library_ui::open` called `show_catalog_now`, which called
`Catalog::open_verified`. That runs `PRAGMA quick_check`, which reads
every page of the database, and then `Catalog::open`, which takes a full
SQLite backup of the file before a migration and rewrites its structure
afterwards. On a 50,000-image library that is tens of megabytes of I/O
on a tablet's flash, and it happened before the window had painted
anything — so on Android it was counted against the five seconds the
input dispatcher allows, and on the desktop it was a launch that sat on
a blank window.

`dr_catalog::recovery`'s own module documentation says the check is
affordable "at startup, where a failure has a user in front of it who
can answer a question". That was the intent and it was not true: there
was no interface yet in which to ask. Now there is, because the open
happens on a worker and the answer arrives on a channel drained by a
timer — the same shape the scan, the thumbnails and the login already
use.

The gate the synchronous call provided is kept, and is the reason the
scan moved with it. `Catalog::open` succeeds on a damaged file whose
header survived, so a scan running beside an unanswered recovery
question writes ETags and image rows into damaged pages and turns a
catalog that had a backup into one where the backup is the only copy
left. So the scan now starts from the drain, on the two answers that
permit it, and not at all on `Corrupt`. `library-scanning` stays true
throughout, which hides the Rescan button and stops the gate being
merely advisory.

What the user sees while it runs is a third empty state. The grid
already refused to conflate "still scanning" with "scanned, found
nothing"; "opening the library" is a third answer and it gets its own
sentence, because a grid saying "Scanning…" while nothing is on the
network is the same kind of lie the other two were separated to avoid.

`show_catalog_now` stays, unchanged and blocking, for `recovery_ui`.
That call site has the event loop running, has just replaced the file
under a `forget_catalog`, and has `recovery-busy` on screen — the same
reasoning `recovery_ui::answer` already gives for doing its file copy in
place. The part both paths share is now `adopt_catalog`.

One consequence worth naming: the cache-usage figure on the settings
page was read at startup from a catalog that is no longer open by then.
It moves to the page's `on_open` closure, beside the face coverage,
which is read there for exactly the same reason — it is only ever looked
at while that page is on screen.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 18:41:34 +02:00
dtourolleandClaude Opus 5 1c5849ebe8 Unpack the bundled models on a worker, not on the way to the first frame
Launching v0.10.0 on the tablet produces an ANR: "Waited 5000ms for
MotionEvent", 5,827 ms on one input sequence, over a window that has
never painted. The app recovers and sits at 0% afterwards, so it is a
startup cost rather than a hang.

The structural fact behind it is that `android_main` runs with the
activity's input channel unserviced. Nothing drains it until Slint
reaches `poll_events`, and Slint does not reach `poll_events` until
`dr_ui::run` calls `window.run()` on its last line. Every millisecond
before that is a millisecond the input dispatcher waits on, so five
thousand of them is an ANR whatever the work happens to be.

The largest single piece of that work was here. `install_bundled_models`
copies 41 MB on the first launch after an install — 24.9 MB of scene
model, 13.6 MB of embedder, 2.5 MB of detector — each read whole out of
the APK into a `Vec` and written to `/data`, in a loop, on that thread.
v0.10.0 is the release that added the scene model, which is 60% of that
total, and it is the release the ANR appeared in. The 8,010 minor faults
in the report are about what 41 MB of freshly touched pages costs.

So it moves to a detached thread and the function returns as soon as the
thread is running. Nothing on the launch path wanted the result: the
only two things that read these files are the People screen and the
scene tab, both of which are reached by hand, minutes later, from
workers of their own.

What that costs is a window in which a model looks absent.
`library::face_models` and `library::scene_model` decide availability on
`is_file()`, so during the copy both report their feature unavailable —
which is the same answer they give a build shipping no weights at all,
the ordinary case both were written around. Briefly pessimistic rather
than wrong, and the temporary-name-then-rename that was already there is
what keeps it from being worse than that: a lookup never sees a
half-written file, only an absent one. Both call sites now say so.

A completion line reports the bytes copied and the milliseconds taken,
including when it is zero, so the second launch after an install can be
told from the first in a log rather than by inference.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 18:41:34 +02:00
dtourolleandClaude Opus 5 4efab496c2 Put the refinement on a slider, per layer
Benchmarks / CPU and I/O (per commit) (push) Successful in 3m27s
Benchmarks / Frame budget (on demand) (push) Skipped
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
Build and test / Desktop (Linux) (push) Failing after 1h14m30s
Build and test / Layer separation (push) Successful in 43s
Traceability / Requirement traces (push) Failing after 46s
Build and test / Android (aarch64) (push) Successful in 1h8m56s
The evidence was being gathered and spent immediately at one strictness
nobody could see or change. This makes it a control.

`MaskLayer::refine` is shaping, like the feather, and it lives on the layer
for the layer's reason: two layers may sit on the same category and want
different amounts of it, and the model ran once for both.

## What the segmentation now stores

`CategorySummary` keeps the **coarse** mask and the `Refinement` beside it,
rather than a refined mask. That is what gives the control an off position
that is bit-for-bit the model's own weighting, and what stops a strictness
change from needing the model again.

`category_mask_at` borrows at zero and wherever no refinement could be
fitted, so a layer nobody has touched costs nothing over the old path.

The fit moved to the far side of the orientation permutation. The verdict is
a per-pixel field over the same grid as the mask it gates, so fitting it
upright would mean permuting a proxy-sized buffer afterwards to match — a
second rotation, and a second chance to get one wrong. One logit cell is the
same number of pixels either way: the letterbox scales by the longer edge and
a permutation does not change which edge that is.

The two frames became a `Frames` struct rather than six parameters. This
function reads the picture twice for opposite purposes — the model needs it
upright or it recognises far less, the refinement needs the sensor's grid —
and a transposed pair produces a plausible mask over slightly the wrong
pixels, which is the failure this module is most prone to.

## It rebuilds the field, and the signature says so

`refine` is mixed into `subject_signature`. Unlike a feather, which is read
off a field that is already correct, this changes which pixels are in the
mask at all — so it changes the coverage the field is measured from. Omitting
it is the bug where the slider moves and nothing happens until some unrelated
control invalidates the cache.

That puts it in the same cost class as a close or an open, which is why the
row takes `SliderRow::changed` — already once-per-gesture, since that row
takes `SliderTrack`'s `committed` internally — rather than a live stream.

The slider is offered only where there is something to move: a category
source, *and* a refinement the frame actually gave enough to fit. A control
that moves and does nothing is worse than an absent one.

## Two defaults that are deliberately different

A layer added from the panel starts at 4.0, because a category's edges are
twenty proxy pixels wide before anything is done to them and a photographer
adding a sky mask wants the sky rather than the sky plus every chimney in it.

A layer read from a sidecar with no `refine` key starts at **zero**. A file
written before this control existed has to render as it did then, and a
default of 4 on absence would quietly re-grade every stored category mask in
the catalogue. `a_categorys_refine_strictness_survives_and_defaults_off`
holds both halves, and `an_out_of_range_refine_is_clamped` holds the file to
the scale — past the top of it every colour fails and the mask deletes
itself, which reads as lost work rather than as a bad file.

`MAX_REFINE` is dr-pipeline's own constant mirroring
`dr_segment::STRICTNESS_MAX`, following `Falloff` and `Morphology`: this
crate holds the description of an edit and must not depend on the crate that
runs a model. dr-ui is where the two meet, and the only place that converts.

Verified: fmt clean, clippy --workspace -D warnings clean, 488 dr-pipeline
and 60 dr-segment tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 18:30:40 +02:00
dtourolleandClaude Opus 5 4f4abd335f Cut a scene category back to the pixels that agree with it
A flag in the sky came out weighted as sky, and no feather setting fixed it.

The scene model's logits are `[1, 150, 80, 80]`, so one cell is eight input
pixels; at the 1600px proxy the letterbox scale is 0.4 and **one cell is 20
proxy pixels**, which `rasterise`'s bilinear then spreads across one more
either side. A flag is a handful of cells whose softmax is dominated by the
sky around it. The information was never in the grid, so nothing downstream
of the grid can recover it.

Tiling is the answer for an instance and is not available here: a category
has no bounding box to tile over — sky is wherever the sky is. But the
photograph is at full proxy resolution even though the weights are not, and
it knows exactly where the flag is. So the model says *what*, and the
pixels say *which of them*, which is the division of labour arm C already
draws between the instance model and the watershed.

## Seeds, and why the erosion radius is not a guess

Threshold the weights high, take `signed_distance`, and keep what is more
than 1.5 cells inside. One cell *is* the model's resolution and the bilinear
spreads it across one more, so the band either side of the boundary is smear
rather than evidence. Deriving the radius from `Scene::cell_pixels` rather
than picking a pixel count means it stays right if the proxy edge or the
export changes.

The mirror of that set is a confident *exterior*, free from the same field.

## Dropping small modes is the step that makes it work

Four k-means modes per side, not one Gaussian: sky is blue at the zenith,
white where the cloud is and pale at the horizon, and one blob over all
three rejects two of them.

Then modes holding under 3% of a side are discarded, and without that step
the whole thing fails on the case it was built for. A small flag deep in the
sky has both a high weight and a large distance from the boundary, so it
lands in the interior sample and teaches the model its own colour. It cannot
be excluded geometrically. It can be excluded by share.

Luminance is weighted at a quarter against chrominance for the same reason
the watershed's gradient is. Sky's variance is dominated by luminance, so at
equal weight the distribution is a long bright streak that a mid-grey flag
sits comfortably inside. A flag is separated by chrominance; a cloud is
separated by luminance alone. Not zero, or a dark bird against a bright sky
survives.

## Two tests, because either alone is wrong

Absolute — is this colour plausible under the category, as a chi-square on
the Mahalanobis distance. Comparative — is it likelier inside than outside.
A pixel must pass both.

The absolute test is what catches the flag, whose colour is far from *both*
sides and which the comparative test alone would leave at even odds. The
comparative test is what stops the absolute one needing a constant tuned per
category.

## What this cannot do, written down rather than left to be discovered

An intruder large enough to hold its own mode is kept. By share, a flag over
a fifth of the sky and a cloud bank over a fifth of the sky are the same
object, and colour does not separate them either — a white cloud is as far
from blue sky in chrominance as many intruders are.

So `min_cluster` is not a threshold with a correct value waiting to be
found; it is the trade-off itself, set where a photographic intruder falls.
Both ends are pinned by tests — `a_flag_in_the_sky_is_removed` and
`an_intruder_larger_than_min_cluster_survives` — so that moving the number
reads as moving the trade-off rather than as fixing a bug. The case left
open is a large unrecognised object in a clean category, which wants the
boundary snapped to watershed basins and is a different mechanism.

## Safe to apply without a control

It is subtractive: the output is the input times a factor in `0..=1`. The
worst failure available to it is losing part of a real sky, never gaining a
region, so a blue car below the horizon that was never in the mask cannot be
pulled into it. And a factor in `0..=1` cannot raise a sum, so `scene.rs`'s
partition still holds when every category is refined independently — the
weight taken off the flag lands in the unlisted remainder, which is where a
flag belongs, ADE20K having no class for one.

Every path without the evidence to judge returns the weights untouched and
says which path it took. A refinement that silently did nothing is
indistinguishable from the feature being off, and an empty seed set fitted
to a distribution would reject every pixel.

The signature is deliberately unchanged: categories are addressed by name,
not by index, so a sharper mask cannot create the stale-index hazard the
signature exists to guard against.

The example writes `<prefix>-<category>-refined.ppm` beside the coarse one,
never instead of it — whether this is an improvement is a comparative
judgement and one image cannot answer it.

Verified: fmt clean, clippy -D warnings clean, 57 dr-segment tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 18:30:16 +02:00
dtourolleandClaude Opus 5 bfdb6d5e4b Cut the People rail's portraits off the blocking path
Opening Identity cut a portrait for every person before the screen was
allowed to appear. Measured against the reference library: 3.6 seconds,
of which 3.2 is 931 full 1024px proxy decodes — the fallback for faces
indexed before crops were stored beside them. The rest of the wait was
the rail building 14,268 rows, fixed in the commit before this.

`refresh` now draws with the portraits already cached and hands the rest
to `fill_covers`, which cuts them 8ms at a time — under half a frame —
behind a screen that is already up. Blocking work on that path goes from
4,965ms to 16ms on this library.

A timer rather than a thread. The work is a catalog query and a decode
against a `ThumbStore`, and both handles live on the UI thread; a worker
would need its own connection to the same file, which is what the sweep
and the regrouping pass do because they run for minutes and would
otherwise be unbounded. This is seconds of small, independent pieces, so
slicing answers the same question more cheaply.

`pending` is in rail order, and the rail is sorted by confirmed faces, so
the portraits the user is looking at are cut first. It patches single
rows rather than reloading: a reload would rebuild the model on every
tick, and a model replaced underneath the `ListView` is what the slicing
exists to avoid. Each patch checks the row still holds the person it was
started for — a stale index would draw a face beside somebody else's
name — and stops the fill when it does not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 18:30:16 +02:00
dtourolleandClaude Opus 5 d7b851f622 Build the People rail rows you can see, not the ones you cannot
`Flickable { VerticalLayout { for ... } }` instantiates every row. On the
reference library that is 14,268 row subtrees, which measures at 4.2
seconds before a pixel is drawn — and it was paid again on every action,
because every action reloads the model. Roughly half the wait between
clicking "Identity" and the screen appearing was this.

Slint's compiler has a virtualising path for a `for`; it is what makes
`std-widgets`' `ListView` cheap. It keys on the parent element's base
being *named* `ListView` and exposing the five lengths its layouting code
writes back, and a custom base is explicitly allowed. So `widgets.slint`
grows one, and `std-widgets` stays out of the file that establishes our
style. Measured on the real screen: 200,000 rail rows now render in 62ms.

Three things had to be true together, and each was silent on its own.

**The row height must be constant.** The rail hid a set-aside person with
`height: cond ? 52px : 0px`, a height that reads the model — so the
layout cannot place row N without building rows 0..N, and Slint builds
them all. The filtering moves to Rust, where the toggle was already
reloading anyway.

**The list must not be wrapped.** It carries its own stretch and preferred
size, having no natural height to offer; a Rectangle in between hands the
layout that Rectangle's constraints, which are taken from the list and
are therefore nothing.

**The panel must let it fill.** `Panel` lays its children out with
`alignment: start`, which gives each its preferred height — right for a
column of sliders, wrong for anything that scrolls. Hence `Panel.fill`.

Get any of them wrong and the rail renders empty, with no error and a
model full of people. All three were, in turn, before a headless render
of the real screen showed a blank rail.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 18:30:16 +02:00
dtourolleandClaude Opus 5 f26f1ab694 Withhold the empty groups that fill the People rail
`load_people` returned every live person, and on the reference library
that was 14,268 rows of which 11,739 held no faces at all — 85 per cent
of the rail naming nobody and able to do nothing.

They are not a mystery. A regrouping pass creates a person per cluster;
the next pass moves those faces elsewhere and leaves the person it
emptied behind. `faces::prune_empty_unnamed` exists for exactly this and
runs only at the end of a pass, so nothing clears what accumulates
between them, and the Identity screen never prunes at all.

Each one cost a `for_person` query and a built row on every reload. This
withholds precisely the set the prune already treats as disposable —
empty, unnamed, not set aside — and no more.

Filtered rather than deleted: a screen is being drawn, not a catalog
repaired. Nothing is lost, a sync cannot resurrect what was never
removed, and the prune stays the one place that decides these can go.

An empty group with a *name* still shows. That one is not debris but the
symptom of a real failure — a named person whose faces were regrouped out
from under them — and hiding it would take away the only way to merge
them back.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 18:30:16 +02:00
dtourolleandClaude Opus 5 94b1b9e0cc Reconcile what four branches each built separately
Three seams, found by the first compile after the merge.

Two branches implemented 'can this scope be reordered' independently:
one from the collections model, one from the catalog through
orders_manually, on every re-read and excluding the trash. The second is
the better answer and is what survives; it only needed to set the
property app.slint declares.

Two lints from scene-mask-ui, which was merged mid-flight and had never
been through -D warnings: an is_none check spelled out where clippy wants
?, and a return in a cfg block's tail.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 16:51:41 +02:00
dtourolle 19b56ad85c Merge: collection ordering, and a range that says where it ends
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

# Conflicts:
#	docs/traceability.md
#	ui/dr-ui/src/collections_ui.rs
#	ui/dr-ui/src/library.rs
#	ui/dr-ui/ui/app.slint
#	ui/dr-ui/ui/library.slint
#	ui/dr-ui/ui/widgets.slint
2026-08-30 16:38:08 +02:00
dtourolle 328fda6f7c Merge: mask a whole category, not just one instance
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

# Conflicts:
#	apps/darkroom-desktop/Cargo.toml
#	docs/traceability.md
2026-08-30 16:37:31 +02:00
dtourolleandClaude Opus 5 86260b5028 Bind a category layer to its own distance field
A category mask showed nothing and its adjustment covered the whole
photograph. Both from one line: the loop in `MaskPass::rasterise` picks a
distance field by matching `layer.source`, that match named only `Subject`,
and a `Category` layer fell through to `_ => (&self.empty_subject, 0)` — a
1x1 placeholder. No field, so nothing to draw and nothing to confine the
adjustment.

The comment three lines above the arm I missed describes the failure I then
shipped:

    an absent mask that defaults to "everything" would apply the
    adjustment to the whole photograph

There are *two* matches on `layer.source` in that loop — one choosing the
field, one building the params. Adding the category to the second and not
the first compiles, runs, and is wrong in exactly the way the first one
warns about.

## Also: a missing mask must still be the right size

Both model-backed arms of `ensure_subject_fields` used `unwrap_or_default`,
which yields an empty `Vec` when the coverage is gone. `SubjectMasks::upload`
rejects a wrong-sized field and fails the whole batch, so `self.subjects`
becomes `None` and *every* layer in the stack loses its mask — one stale
reference silently unmasking the others.

Pre-existing, and it mattered less when the only model-backed source was a
subject: an instance index goes missing rarely. A category name goes missing
whenever the descriptor is edited, which is a thing the descriptor exists to
allow. A full-size empty field costs one layer instead of all of them.

Neither of these is reachable from a test on this machine — both live past a
GPU adapter and a real segmentation — so they surfaced the only way they
could, by someone opening the app and looking.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 14:28:27 +02:00
dtourolleandClaude Opus 5 9994bb4ce7 Regenerate the matrix over the third wave
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 14:02:40 +02:00
dtourolleandClaude Opus 5 7f350e8c4c Offer the categories in the masks panel
Everything before this was reachable only from an example that writes PPMs.
This is the part a photographer can touch: a list under the subjects, click
one, get a mask layer for every pixel of that category.

## Under the subjects, and the order is the argument

Clicking the photograph is how a local adjustment usually starts, so the
things the model *found* come first. A category is the move you reach for
deliberately — grade the sky, not this one bird — and putting it second says
so without a word of explanation.

Coverage is shown for the same reason a subject's score is: it tells the
photographer whether a category is worth a click before they spend one
finding out.

## Not a new tab, which is what was asked for

A develop tab is derived from `Attribute`, not declared — the tabs exist
because operations claim an attribute, and no amount of Slint adds one. A
seventh attribute would have meant duplicating every adjustment once per
category, and the combinatorics get silly by the third.

Reached through masks instead, a category composes with every adjustment
that already exists, and inherits feather and falloff rather than needing
its own. What was described as "per-category sliders with feather and decay"
is exactly what this is; only the door is different.

## Verified as far as it can be here

Compiles, populates, round-trips, 560 dr-ui tests green. **Not clicked** —
synthetic input is blocked on this setup, so how it looks and feels is
unverified and wants a human at the window.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 14:02:34 +02:00
dtourolleandClaude Opus 5 763dfd353a Weigh the categories in the same precompute, and mask with them
The scene model shipped with a decoder and no caller. This runs it.

## Beside the instance pass, not instead of it

`compute` now does both on the same upright frame and lays both back down
the same way, so instance masks and category masks index into one grid —
the sensor's. A failure in the scene half is logged and dropped rather than
propagated: no scene model is an ordinary state, and a photograph that can
still be masked by subject should not become unopenable because the
categories are missing.

Categories under half a percent of the frame never reach the cache. A
control that does nothing when moved is worse than an absent one, and each
one it skips is a proxy-sized buffer not allocated.

## The shader needed nothing

A category reaches `dr-gpu` as a soft coverage buffer at proxy resolution,
turned into a distance field — which is exactly what a subject is. So they
share `MODE_SUBJECT`. That is not a shortcut taken for speed: the shader has
no way to tell them apart and no reason to want one. What differs is only
which model produced the coverage, and that has already happened by then.

Feather, falloff, dilation and erosion therefore work on a category on the
day it arrives, because they were never subject-specific.

## Where the weights come from

`scene-model` compiles the graph in and the desktop app takes it; Android
leaves it off and reads the copy `install_bundled_models` unpacks, because
24 MB of constant is worth avoiding in a mobile install and not worth the
plumbing to avoid on a desktop one. Embedded is tried first — a build that
has the weights compiled in should not be silently overridden by a stale
file in a data directory.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 14:02:22 +02:00
dtourolleandClaude Opus 5 d70dcf78d1 Merge: recover a damaged catalog, and capture a crash locally
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 13:45:18 +02:00
dtourolle c48896bd95 Merge: touch selection and drag, from the gallery-selection branch
Verified before merge: fmt clean, clippy -D warnings clean, 563 dr-ui tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

# Conflicts:
#	docs/traceability.md
2026-08-30 13:45:18 +02:00
dtourolleandClaude Opus 5 b34e786f01 Give the drag a pick-up, so it stops losing to the scroll
Dragging a photograph out of the grid worked about half the time, and
nothing on the screen explained the other half.

`DragArea` and `Flickable` do arbitrate, but not evenly. The Flickable
claims any press that travels more than eight pixels along its own axis
within half a second of landing, and holds that claim until the finger
lifts. So a drag toward the sidebar only ever began two ways: a flick
sideways clean enough that the finger never wandered eight pixels
vertically, or a wait of half a second before moving at all. Both are
real gestures and neither was written down.

The wait is now the gesture, and it has a mark. The long press that
already turns on selection mode also picks the photograph up: a ring
opens around the cell and the grid stops scrolling under it, so from
that moment the drag is the only thing the finger can be doing. The cue
can only arrive after the ambiguity has passed, which is the right way
round — when the photograph lifts, dragging it works.

Two details worth naming. The hold is now armed even when selection mode
is already on; it used to be skipped there, on the grounds that there
was no mode left to switch on — but that is precisely the state a
forty-image drag starts from, so the one gesture that most needed a
pick-up was the one with none. And the ring is drawn after the cell
loop rather than on the cell: z-order inside a `for` is loop order, so a
cell grown past its bounds would stand over two neighbours and be cut
off by the other two.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 13:20:56 +02:00
dtourolle 7b5f62019b Merge branch 'master' into fix/gallery-selection
# Conflicts:
#	docs/traceability.md
2026-08-30 11:04:13 +02:00
dtourolle 458025c607 Merge the scene model: a second, ADE20K-trained graph for per-category grades
Five commits. `models/` becomes one tree at the repository root so the
weight the application carries is a single `du -sh`; the export script
stops building its multi-gigabyte venv in RAM; `yolo26s-sem-ade20k`
joins the instance model rather than replacing it; the decoder turns its
logits into a partition of unity over eight photographic categories; and
four packaging routes put the file somewhere each platform can find it.

The instance model stays exactly where it was. A semantic model merges
every pixel of a class into one region, so it cannot separate two people,
and separating two people is what clicking a subject needs. The scene tab
grades whole categories and does not care. docs/segmentation.md §16.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

# Conflicts:
#	docs/traceability.md
2026-08-30 10:50:49 +02:00
dtourolleandClaude Opus 5 8e4e24ad09 Say what the drag payload actually is: the arming, not the cargo
`drag-payload` was documented as "called when a drag starts, so it
always reflects the selection as it is at that moment". Neither half is
true, and the comment is the only reason anyone would believe the drop
reads it.

`DragArea` tests `data.is_empty()` in its event filter — on every
pointer event, the first one included, which arrives long before there
is a drag. And a Slint binding that calls a callback has no dependency
to be invalidated on, so it is evaluated once, when a finger first lands
on that cell, and cached for the life of the cell. What it answers is
therefore always an empty selection.

None of the drop handlers read it; every one of them reads `dragging`,
which `drag-started` fills in at the moment that matters. What this
callback does is keep the `DragArea` armed, and it manages that only
because `set_user_data` is called unconditionally — an empty `Vec` is
still user data. Guarding that call, which reads as an obvious tidy-up,
would silently stop the grid dragging at all.

Comments only.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 10:50:10 +02:00
dtourolleandClaude Opus 5 1a59fb33c1 Stop the press that starts a drag from deselecting what it grabbed
Dragging a selection of forty photographs onto a collection filed one.

Selection mode reports every press as a ctrl-press — deliberately, so
touch and pointer go through one set of rules rather than two — and ctrl
toggled. So the press that took hold of one of the forty took it *out*
of the selection on the way down. `drag-started` then looked at the cell
under the finger, found it unselected, and did exactly what it is meant
to do with an unselected cell: made it the whole selection and carried
it alone. The only sign was the grabbed cell's ring blinking out at the
moment the user began to move.

A plain press has never had this problem, because pressing an
already-selected cell has always been documented to leave the selection
alone — for precisely this reason. Ctrl now does the same: adding still
happens on the press, since the drag reads the selection immediately,
but *removing* is handed back as `Press::Deferred` and applied by the
click. Slint reports a click only for a press that stayed within
`tap-slop`, so a tap still toggles and a drag never does.

The unit tests now go through a `click` helper — a press and the release
that follows it — because that is the only thing a user can perform, and
calling `apply_press` alone would assert against half the policy.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 10:49:51 +02:00
dtourolleandClaude Opus 5 7312aceded Get the scene model onto the devices that need it
The decoder can load from a path; nothing yet put a file at one. Four
packaging routes, and one lookup that finds the result.

## Not `include_bytes!`, unlike the instance model

The instance model is 11 MB and compiled in, which was the right call for
it: Android hands the app no filesystem path (ARCH §6.9) and 11 MB is
tolerable. The scene model is 24 MB, and 35 MB of constants in the binary
is paid by every install whether or not the tab is ever opened.

So it follows `models/face/` instead — carried as an APK asset, unpacked
once at first launch into the shared directory a desktop install already
uses, after which every lookup finds it where it finds a desktop user's.
Assets are stored rather than deflated in the APK, so unpacking is a copy
rather than an inflate.

`embedded-scene-model` exists for the desktop build with nowhere else to
read from, and for tests wanting the real graph. Off by default, which is
the asymmetry with `embedded-model` and the reason for a separate
feature.

## Three files, all or none

`scene_model` insists on the graph, its vocabulary and the category
descriptor together, for the reason `face_models` insists on its pair: a
graph alone decodes to 150 anonymous channels. Reporting the set missing
beats starting and failing at the first inference.

## The two model sets are not the same kind of thing

`install_bundled_models` now carries both, and the distinction is worth
keeping in view. Face weights are absent from the repository *by design*
— the InsightFace grant is research-only (docs/faces.md §2) — so a build
carrying none is ordinary. The scene model is committed, so a build
carrying none means a checkout without `git lfs pull`.

Neither is fatal. A photo editor that refuses to start over a missing
grading feature is worse than one that starts without it, so both report
themselves unavailable exactly as face indexing already did.

The LFS-pointer guards apply to the `.onnx` only. The vocabulary and the
descriptor are legitimately a few kilobytes, and a size check that fails
on them would be a guard against the wrong thing.

`scene_model` is exported ahead of the tab that will consume it so the
packaging added here has something to be verified against — assets
written where no lookup looks would be a silent mistake for as long as
the tab took to arrive.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 10:48:54 +02:00
dtourolleandClaude Opus 5 1ff52102b6 Drop the anchor bookkeeping the double tap took with it
`previous_anchor` existed for one gesture: a double tap in selection
mode took the range from where selecting began, and both taps had
already moved the anchor onto the cell being tapped, so the origin the
user meant had to be remembered separately.

That gesture is gone — "Select to…" says what it is about to do instead
of hiding a forty-image range behind a thing a hand does by accident —
and what is left is a field that four places write, `PressUndo` carries,
`cancel_press` restores, and nothing at all reads. `apply_press` is
`select_row`'s only call now that there is no anchor to remember, so the
wrapper goes with it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 10:40:07 +02:00
dtourolleandClaude Opus 5 8eeb9ba0f6 Offer the backup, and then the rebuild, when the index turns out to be damaged
NFR-R6 asks for an integrity check at startup and two offers behind it, and
none of it existed. `PRAGMA integrity_check` appeared nowhere in the tree,
`Catalog::open` was `open` → `configure` → `migrate` → `backfill` and nothing
else, and corruption therefore surfaced as whatever rusqlite error the first
unlucky query happened to produce — "database disk image is malformed"
attached to a thumbnail refresh, elided into a 34px banner, over an empty
grid saying "No images found · Check the library folder". Two messages that
disagreed, and no way forward but deleting catalog.sqlite by hand.

The property that makes the second offer real was already here and load-
bearing: the catalog is an index, not a source of truth, rebuildable from
sources plus sidecars (invariant §5.2.4, cited by schema.rs, trash.rs and
lib.rs). And sync.rs already knew how to take a coherent snapshot of a WAL
database. What was missing was the check, the type, and the conversation.

Four pieces:

**The type.** `CatalogError::Corrupt`, and — the part that makes it worth
having — a hand-written `From<rusqlite::Error>` that classifies rather than
wraps. `SQLITE_CORRUPT` and `SQLITE_NOTADB` become `Corrupt` wherever they
arise, so a background job that trips over the damage first reports the same
thing the startup check would have. `SQLITE_IOERR` and `SQLITE_BUSY`
deliberately do not: a dropped network mount is a different problem, and
telling someone to rebuild their index would be a wrong answer delivered
confidently.

**The check.** `Catalog::open_verified`, `quick_check` before the open rather
than after, because opening runs migrations and a damaged catalog with an
intact header would otherwise have structure rewritten on top of structure
that is already wrong. Bound to `open_verified` and not to `open`: the check
reads every page, which is affordable once at startup where a user can answer
a question, and not affordable on the dozens of opens a session's background
tasks make.

**The backup.** NFR-R2's second clause, taken between `configure` and
`migrate` in `Catalog::open`. A migration is the one routine operation that
rewrites table structure, so it is the likeliest way this file becomes
unreadable, and it is the last moment the pre-migration state exists to be
copied. Three generations, through SQLite's backup API after a TRUNCATE
checkpoint — never `fs::copy`, which on a WAL database backs up a state older
than the catalog and possibly torn. A failure to take the copy is logged, not
raised: a full disk must not be what makes a library unopenable.

**The conversation.** The first line of the dialogue is that the photographs
and the edits are safe, before the diagnosis, because that is the question the
user is actually asking. Then the two offers, which are *not* interchangeable
and are not presented as if they were: a restore keeps collections, and a
rebuild cannot, because a manual collection is a set of images assembled by
hand and nothing in the filesystem records it (docs/catalog.md §8.1). The
labels say so, and the rebuild does not take the affirmative styling while a
restore is on the table.

One thing that is a fix rather than a feature: `show_catalog_now` now gates
the scan. `Catalog::open` succeeds on a file whose header survived, so the
scan that used to start immediately afterwards would write folder ETags and
image rows into damaged pages in the seconds while the user was still reading
the question — turning a file that had a backup into one where the backup is
the only copy left.

Restore also deletes the damaged catalog's `-wal` and `-shm`. That step is
easy to leave out and fatal to leave out: a journal belonging to the old file,
sitting beside the new one under the same name, is replayed into it on the
next open. That is not a restore, it is a fresh corruption with the evidence
gone.

Tested by corrupting a fixture catalog — 500 images and a collection, then
every page past the second overwritten — and driving both branches. The
restore is asserted on the collection, because a collection is precisely what
distinguishes the two paths; the rebuild on the damaged file being kept and
the next open producing an empty catalog at the current schema. Plus the
`SQLITE_NOTADB` presentation, a damaged backup being refused rather than
installed, and a v1 catalog whose pre-migration backup comes back reading
v1 rather than v11.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 10:34:12 +02:00
dtourolleandClaude Opus 5 e22945a62d Tag the colour-independent status that was already built and unrecorded
NFR-A11Y-3 — no status conveyed by hue alone — read as untagged, and
outstanding.md said "no compliance work found". Both were wrong. Five places
already implement it, and four of them name the requirement in a comment
explaining the design; what none of them had was a TRACES line.

- `histogram.rs::percentage` states clipping as a figure and keeps `<0.1%`
  distinct from `0%`, so the text cannot say "none" while the marker beside it
  is lit.
- `histogram.slint`'s ClipReadout is the other half: a marker that appears and
  disappears rather than changing tint, and the figure next to it. Either alone
  reads.
- `library.slint`'s star strip is a solid star against an outline, differing in
  shape and luminance, over an achromatic palette.
- `library.slint`'s FlagMark is a tick against a cross, and a reject also dims
  its whole cell.
- `peaking.slint`'s colour chips say "Red" and "Cyan". A control for choosing
  between hues, presented only as hues, is unusable by exactly the person most
  likely to need it.

The tag is honest about being wider than the evidence, and outstanding.md now
records both gaps. Only the clipping clause has a test that would fail if the
behaviour were removed; the three Slint components are argued rather than
asserted. And the requirement's first named example — catalog colour labels —
has no interface at all: `label` is a nullable column nothing writes or shows.
That clause is untestable rather than satisfied, and closes when the label UI
is built with a shape from the start.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 10:19:48 +02:00
dtourolleandClaude Opus 5 ef07e6ca3e Give the canvas tools a rail of their own, and the column one width
Build and test / Desktop (Linux) (push) Failing after 1h14m38s
Build and test / Layer separation (push) Successful in 48s
🐳 Android image / Build and push (push) Successful in 16m30s
Build and test / android-image (push) Successful in 16m31s
Traceability / Requirement traces (push) Successful in 1m47s
Build and test / Android (aarch64) (push) Successful in 1h0m21s
Crop, Local and Repair were chips at the head of the develop column, sharing a
row with the adjustment groups and told apart from them by the shape of their
highlight. Three things followed from that, and only the last is cosmetic: the
column closes, so the way out of a mode went away with the way in — hence the
duplicate "Done Cropping" over the canvas; the chips are generated from the
operation set, so the widest thing in the sidebar was a row nobody had chosen
the contents of; and a mode and a filter are different kinds of state wearing
one control.

They are a fixed 60px rail down the left now, generated from a single table in
toolrail.slint. A tool is one row of it plus a drawing plus a ViewMode variant;
nothing in app.slint is touched to add one. What is left of the strip is the
group filters, so it is GroupStrip.

The column stops measuring itself. Every panel published a content-width and
declared it as min-width, and the column took the largest — which spent the
photograph's pixels on whatever happened to be widest, and moved the image
sideways when switching tools swapped one set of panels for another. It is
panel-width now, one number in style.yaml.

That number is 360 and it is measured, not picked: the contents report a
minimum of 344 in every mode, and they do not compress below it because a Text
that does not elide reports the same minimum as preferred. 320 was tried and
sliced Paste down the middle. The Flickable's viewport is floored at the
layout's minimum rather than its preferred width for the same reason — content
that is never told how much room it has cannot adapt to having less.

Removing the eight content-width declarations repairs three comments an
earlier edit had spliced sentences into. The raw histogram's note on keeping
its hint short is rewritten rather than dropped: an over-long hint no longer
widens the column, it pushes the column's minimum past the width it has and
clips the panel, which makes that constraint sharper rather than obsolete.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 09:59:50 +02:00
dtourolleandClaude Opus 5 9e519eb8a6 Make the thin ring the only mark a selected photograph carries
Two treatments said "selected" and neither said it well.

A selected cell got a 2px border and a lifted fill. The border is drawn
on the outside of a cell whose content sits 6px in, so it ate into the
thumbnail: selecting appeared to nudge the photograph. Inside it, a
second thin ring marked the anchor — the end a shift-click measures from
— and that ring was the clearest thing on the cell, so it read as *the*
selection to everyone who had not written it.

Worse, the ring outlived what it described. An anchor survives a
deselection, so a thin box sat around the last photograph touched with
nothing selected at all, indistinguishable from a cell that had stayed
behind. That is the one thing a selection cue must never be: ambiguous
about whether something is selected.

So the ring is now what it already looked like. One mark, drawn inside
the cell and 4px clear of its edge, so it never touches the thumbnail and
never changes a dimension — selecting adds ink and moves nothing. Two
pixels rather than one, because it is now carrying the whole cue across
forty cells at arm's length against a thumbnail of any brightness. The
outer border is hover alone.

The anchor keeps no mark, and loses nothing it was earning: the bar says
"Tap the last photograph" while a range is armed, which answers the
question the ring existed to answer. The ordinal still lives in the
controller and still decides where a range extends from; what is gone is
the claim that the user needs to see it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 09:47:32 +02:00
dtourolleandClaude Opus 5 b120ce48ec Float the selection bar over the grid instead of above it
Selecting the first photograph inserted a 40px row into the view's
vertical flow, so every cell in the grid moved down by it. The act of
selecting shifted the thing being selected out from under the finger —
and a second tap aimed at the neighbour landed on the row below it,
which is the worst possible response to a gesture whose whole job is to
say "this one".

The bar is a floating one now, at the foot of the view. Nothing above it
is re-laid out, so selecting changes what is drawn and never where.

The grid's viewport grows by the same 40px while the bar is there rather
than the Flickable shrinking, which is what keeps that change invisible
too: every cell stays exactly where it was and there is simply further to
scroll, so the last row can be brought clear of the bar instead of being
trapped under it.

It also swallows presses that land on it. A bar floating over the grid is
a bar a thumb can reach for and miss into.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 09:47:20 +02:00
dtourolle f41edc03ff Merge master into wave-2
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

# Conflicts:
#	docs/traceability.md
2026-08-30 09:24:37 +02:00
dtourolleandClaude Opus 5 1cd5ab6815 Put the gesture reference in the application
The document the previous commit generates is for somebody reading the
repository. The person who needs it most is holding a tablet, has just
discovered that a hold does something, and has nowhere to ask what else
does.

So the same scan writes a table the application draws: a "Gestures"
button beside Settings, a sheet with the same scrim and dismissal as the
ones that file and name, and every gesture grouped by where it applies
with its touch, pointer and keyboard routes side by side. Not the `why` —
that is the argument for the design and belongs in the document; on a
phone-sized card it would bury the one line the sheet was opened to read.

The sheet's file knows nothing about what a gesture is. It draws the rows
it is handed, and the rows come from the generated table, because a help
screen with its text typed into it is a second description of one
behaviour — and the second description is always the one that goes stale.
The commit before this deleted a gesture; a hand-kept sheet would still
be describing it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 00:30:21 +02:00
dtourolle 4b212089d2 Merge master into wave-2
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

# Conflicts:
#	docs/traceability.md
2026-08-30 00:25:42 +02:00
dtourolleandClaude Opus 5 964e72bca2 Merge: the formatter's pass over the raw histogram
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 00:25:21 +02:00
dtourolleandClaude Opus 5 e38b730aa6 Reflow what rustfmt wanted in the raw histogram
The author could not run cargo, so this is the formatter's first pass
over the new module and its presentation half.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 00:25:16 +02:00
dtourolleandClaude Opus 5 df90dd95a2 Merge: a histogram that reads the sensor, beside the one that reads the frame
FR-CULL-3's other two bullets. What existed was a display histogram
tagged FR-DSP-7, counting AdjustPass's 8-bit output with r == 255
clipping counters -- it says a highlight is gone precisely where this
requirement needs it to say the highlight is recoverable.

The new reduction runs over the demosaiced scene-linear texture on a
stops-below-saturation axis: camera-native, unbalanced, unmatrixed,
uncurved, normalised by the sensor's own black and white levels, so 1.0
is saturation by construction. Four series, and the fourth is the
brightest channel rather than luma, because a weighted sum of unbalanced
values is a number about nothing. Cached per photograph, not per frame:
nothing downstream of the demosaic can move a count.

Both readings are legitimate and answer different questions, so the
panel offers a choice rather than replacing one with the other.

ARCH 5.5 is amended to match. It specified a pre-demosaic reduction;
retaining the CFA samples costs 48 MB at 24 MP and 120 MB at 60 MP
resident on every photograph opened, whether or not anyone looks at the
histogram, on the platform ARCH 6.2 exists for. The spec now records two
reductions, why the more complete one was not worth its cost, and what
the cheaper one cannot answer: it counts pixels not photosites, it
cannot see above white, and it is measured after the CFA pattern is gone.

Verified: clippy -D warnings clean, 98 dr-gpu tests, 556 dr-ui tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 00:14:34 +02:00
dtourolleandClaude Opus 5 31a3580f9d Extract the gesture vocabulary from the code that implements it
Every gesture the application has was documented in the comment beside
the `TouchArea` that implements it. Excellent comments, and unreachable
by anyone not reading the source — which is the FR-UI-4 failure in a
different costume: a gesture nobody can find is a feature only its author
knows about.

Writing them out again in a hand-kept help page is the failure this
avoids. Two descriptions of one gesture drift, and it is always the prose
that drifts: the code is exercised every time somebody uses the
application and the page is exercised never. A help screen confidently
describing a double tap the grid stopped honouring last week is worse
than no help screen — and the grid did stop honouring one, in the commit
before this.

So the comment beside the implementation stays the only copy, and a
`GESTURE:` block beside it is scanned into two artefacts: `docs/gestures.md`
for a reader, and a Rust table for the application to draw a help sheet
from. Both committed, both gated, so neither can quietly stop describing
the code.

It lives in the traceability crate because it is the same operation on
the same input — walk the tree, pull structured tags out of comments,
render, fail if the committed artefact has moved. Only the vocabulary is
new. It scans `ui` and `apps` alone: a gesture needs an interface to be
performed on, and excluding `tools` is also what stops the scanner
extracting its own worked examples as broken gestures.

Fifteen gestures so far, across the library grid and the People screen.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 00:09:49 +02:00
dtourolleandClaude Opus 5 20c368d3fc Give each touch gesture one meaning, and give tap-to-open back
I broke opening a photograph. The dwell added in "Tell a tap on a
photograph from a hand going past" required a finger to stay down 120 ms,
and a deliberate tap is routinely quicker than that — so the grid stopped
opening anything. Duration was the wrong discriminator: a tap and a brush
are the same length.

**Travel is what separates them, and a graze is by definition a moving
contact.** A press now records where it landed and the release compares:
within 12px it is a tap, beyond that the hand was going somewhere else.
No dwell, so no deliberate tap can be refused, and the rule is the same
for a finger and a mouse — one rule instead of two, and the `touch`
argument the dwell needed goes away with it.

Two real conflicts went with it, because a gesture set that overlaps
itself is unlearnable however each half is documented.

**A drag was also a hold.** Grabbing a cell and moving inside 450 ms left
the hold timer armed underneath the drag, so it fired mid-gesture and put
the grid into selection mode nobody asked for — the drag finished into a
mode that changed what every later tap meant. Starting a drag now cancels
it, exactly as a pinch already did.

**A double tap was also a range.** In selection mode two taps on one cell
selected everything back to where selecting began: no visible state, no
warning, from a thing a hand does by accident. "Select to…" does that job
and announces itself first, so the double tap is gone and two taps are
now two toggles that land where they started. `extend_to_row` went with
it — a second range implementation that only the double tap reached,
where every other range goes through `apply_press`.

The resulting vocabulary, one meaning each: tap opens, tap-and-slide does
nothing, hold starts selecting, drag files, two fingers resize, and while
selecting a tap only ever toggles.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 00:09:37 +02:00
dtourolle da3b1487f9 Merge: drain the job queue that nothing was draining
FR-PLAT-AND-4's Rust half, and FR-PLAT-AND-3's resumability with it. The
queue's claim_next, complete, fail and recover_orphaned had no callers
outside their own tests, so the jobs table accumulated rows nothing ever
ran.

It also fixes a claim that was not safe across two connections: the
deferred transaction took a read lock for the SELECT and only tried to
upgrade at the UPDATE, so in WAL the second worker got
SQLITE_BUSY_SNAPSHOT, which a busy handler cannot retry away. It never
double-claimed, but the loser errored. Now one UPDATE ... RETURNING.

No handler is wired, deliberately. The only enqueue site reachable in
the shipping app produces remote thumbnail jobs already served by the
async grid worker, and inventing a second network path blind is not
worth a requirement reading as covered on the strength of plumbing.

Verified: clippy -D warnings clean, 376 dr-catalog and 548 dr-ui tests,
18 runner tests including four-thread contention and crash recovery.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

# Conflicts:
#	docs/traceability.md
#	ui/dr-ui/src/library.rs
#	ui/dr-ui/src/settings_ui.rs
2026-08-29 23:49:38 +02:00
dtourolle 4b6c110816 Count the sensor's own numbers, so a cull can see headroom the render hides
FR-CULL-3's remaining two bullets. What existed was a *display* histogram
tagged FR-DSP-7: it binds AdjustPass's Rgba8Unorm output, recovers an 8-bit
code value, and counts clipping as `r == 255`. Its own documentation says a
clipped bin means "a highlight that is actually gone rather than one the
transform might still recover", which is the opposite of what a culling
decision needs. FR-CULL-3 asks for the histogram of the sensor data, on the
explicit grounds that a rendered image "systematically lies about what is
recoverable in the raw", and a readout that measures the render cannot answer
that however it is presented.

So this is a second instrument beside the first rather than a setting on it.
Both are true; they are true about different things; the panel offers both
behind a chip row and the words travel with the numbers, because a raw
saturation figure drawn under a heading saying Highlights would be mislabelled
exactly where the difference matters.

**What is reduced over, and what it cost to decide.** ARCH §5.5 specified the
pre-demosaic CFA samples. This reduces over the demosaiced scene-linear
texture instead, and §5.5 is amended to record the choice rather than let the
specification and the code disagree in silence. The texture is camera-native —
unbalanced, unmatrixed, uncurved — and normalised by the sensor's own black
and white levels, so 1.0 is saturation by construction and the distribution
below it is the headroom question with no calibration to carry. Retaining the
CFA samples would mean keeping the packed u32 buffer Demosaicer::run currently
drops: 48 MB at 24 MP, 120 MB at 60 MP, resident per open photograph whether
or not anyone looks at the histogram, on a platform §6.2 exists because memory
is scarce on.

Three things it therefore cannot say, written into the module docs and into
§5.5 rather than left to be discovered: it counts pixels not photosites, so a
saturated site drags its interpolated neighbours up and per-channel clipping
is smeared by about a demosaic kernel; it cannot see above white, because
demosaic.wgsl clamps each photosite at 1.0 for its own good reasons (a Canon
6D reads to 16383 against a declared 15070) so "at saturation" and "a stop
past it" share a bin; and it is measured after the CFA pattern is gone, so it
can name which colour clipped in the reconstructed image but not which
photosite went first.

The axis is stops below saturation, 16 bins per stop over 256 bins — the same
bin count the display reduction uses, so the fold into drawable columns is
shared and a divergence between the two plots would have to be deliberate. A
linear axis spends half its width on the top stop, which is why nobody has
ever drawn a useful linear raw histogram. The fourth series is the brightest
channel rather than luma: these values are unbalanced, so any weighted sum of
them is a number about nothing, and the brightest channel is the one that
saturates first and so the one the headroom question is actually about.

It is a property of the file and not of the render, which has two
consequences. It is computed once per photograph and cached — nothing
downstream of the demosaic can move a count in it — so a cull does not pay the
display histogram's per-frame cost three thousand times. And it describes the
whole frame rather than the visible region, deliberately opposite to
DevelopSession::histogram: a crop changes what is on screen and changes
nothing about what the sensor recorded.

Tags are on the reduction, the type, its constructor and the presentation
arithmetic, each of which has a test that fails if the behaviour goes. The
Slint panel and the push from lib.rs keep their reasoning as prose: nothing
asserts them, and a tag would claim coverage the assertions are not making.
2026-08-29 23:36:21 +02:00
dtourolleandClaude Opus 5 846a249156 Drain the queue that nothing has ever drained
`jobs` has been a complete durable work queue since the catalog was
written, and nothing has ever taken a job out of it. `claim_next`,
`complete`, `fail` and `recover_orphaned` had no callers outside their own
tests; `enqueue` had three. So the table grew one row per photograph and
kept it forever, and FR-PLAT-AND-3's resumability was a property of code
that never ran.

`runner` is the missing half. It owns no thread, no clock and no policy,
and that is the whole design: on Android the process does not decide when
background work may run. WorkManager does, subject to Doze, battery saver
and FR-NC-6's network constraints, and it revokes permission mid-job by
calling onStopped(). So the runner exposes `run_one` — claim, run, record —
and `drain`, which repeats it against a budget, a deadline and a
cancellation flag the host owns. A `Worker.doWork()` with ten minutes calls
drain with a deadline; a desktop idle pass calls it with none. That is the
seam the Android service plugs into, and it needs no Android to test.

Handlers are supplied from above, because the catalog knows what needs
doing and nothing about how: a thumbnail needs a decoder and a fetch needs
a network stack, neither of which belongs under core/dr-catalog. A runner
claims only kinds some handler declares, so a queue holding work this
device cannot do is left alone rather than failed five times.

Four outcomes, and only two of them are the job's fault. Done deletes the
row; Retry backs off; Abandon gives up now, for a failure no retry can fix;
Interrupted releases the claim with its attempt refunded and ends the
drain, because the host stopped rather than the job — five backgroundings
in a row must not mark good work as failed. Process death is the fifth and
cannot report itself, which is what `recover` is for.

Recovery is called from `show_catalog_now`, which is the one place a
catalog is opened for a session and already returns early if one is open.
It has to be exactly once and before any worker starts: there is no owner
column, so a second pass while a worker held a claim would take it away.
The attempt a dead claim consumed is deliberately kept — a job that takes
the process down with it is indistinguishable from one that fails, and the
attempt counter is the only evidence that survives a death.

The tests cover claiming under contention twice over: sequentially across
two connections, and with four threads on four connections against one
catalog on disk, asserting every job ran exactly once. Plus completion,
backoff, giving up, abandoning, interruption, budget, deadline,
cancellation, and a job orphaned by a simulated crash being reclaimed and
run once rather than lost or repeated.

Not wired to a handler yet, and deliberately not: the only enqueue site
the app actually reaches is the remote scan's, whose thumbnails are already
served by the async grid worker, and `walk`'s two sites are reachable only
from the scan_local example. Inventing a handler to make the plumbing look
used is how a requirement comes to read as covered by code that does not
implement it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 23:31:47 +02:00
dtourolleandClaude Opus 5 7418b040da Refuse a scan whose root has gone, instead of reporting it empty
FR-PLAT-AND-2, and a silent failure on both platforms. `dr_sync::scan`
stepped over a NotFound or PermissionDenied the way it does for a child
that vanished mid-walk -- correct for a child, wrong for the root, where
it ended the walk, returned Ok with nothing in it, and reported a
successful scan of a library that was no longer there.

A lost root is now its own error. The images under it are marked
Availability::Offline per FR-CAT-9 and no catalog row is deleted;
`library::persist` clears the mark per file as each one is listed again,
so a root that comes back needs no repair step.

Partly satisfied rather than closed, and the gap is worth stating.
The recovery half is real and reachable on Android today, because
`map_status` turns Nextcloud's 403 and 404 into it and Nextcloud is how
a phone actually gets a library in this build. The causes the
requirement names -- revocation, reinstall, a removed card -- are
properties of a persisted tree permission, and there is none: SAF does
not exist here, `SourceRef::Document` is constructed only in test
modules, and `LocalStorage` rejects the variant outright. When SAF
lands it becomes a third producer of this error and nothing above it
changes, which is why the discovery belongs in the connector.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 23:20:57 +02:00
dtourolleandClaude Opus 5 085ab766b3 Give memory back in the order the user will miss it least
FR-PLAT-AND-5. Android asks for memory back through onTrimMemory and
kills the process if it is not given; until now nothing listened, so the
answer was always "no".

A tiered registry answers instead: GPU caches first, then proxies, then
thumbnails, driven from android_main on MainEvent::LowMemory and
MainEvent::Stop. The order is the argument. A backgrounded app has no
window to draw and therefore no use for a render pipeline, while its
thumbnails are exactly what the user will be looking at half a second
after they come back -- so going into the background frees only the GPU
tier, and only being measured against death frees everything.

Sinks register beside the cache they free and hold weak handles, so the
registry cannot keep a controller -- and every decoded portrait in it --
alive past the interface it belonged to. `try_borrow_mut` and skip: a
warning can land mid-render, freeing textures under the code drawing
with them is worse than missing one, and a warning not acted on is
always followed by another.

The GPU test is the one that matters: an eviction must change no pixel.
A freed intermediate pool whose `colour_key` promise still stands
renders an empty texture, and nothing else would have caught it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 23:20:57 +02:00
dtourolleandClaude Opus 5 55cb9b5b30 Keep the test scene's own arithmetic from overflowing a u64
The first compile this branch ever had. `cargo fmt` reflowed four files
and clippy passed at -D warnings untouched, but one test panicked:
`the_signature_does_not_change_with_scale`, on "attempt to multiply with
overflow".

It is the fixture, not the feature. `scene()`'s little LCG multiplied the
block's y by the golden-ratio constant with a plain `*` while the term
beside it already used `wrapping_mul`, so any scene taller than about 104
pixels overflowed in debug. Only the scale test builds one that large,
which is why 345 of 346 passed around it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 23:20:26 +02:00
dtourolleandClaude Opus 5 e5db849f04 Mark a burst in the grid, and let it be folded away
The counterpart to the grouping: where the signatures come from, and how a
group reaches a cell.

Signatures are computed from the 256px thumbnails dr-thumbs already holds --
vastly more resolution than a 9x8 reduction can use -- so a library that has
been browsed, or that has synced somebody else's shards, has already paid for
them and no RAW is decoded for this. The consequence is stated rather than
hidden: an image with no thumbnail gets no signature and never joins a burst.
That is self-correcting, and it is why the pass runs when the thumbnail sweep
finishes rather than on a timer. Nothing happens at import and nothing happens
at query time.

The mark is drawn as a child of the cell's TouchArea, for the same reason the
star strip is: a click on it must not also reach `cell-clicked` and throw the
user into develop, and children are hit-tested before the element they sit in.
It is never hidden on hover the way the stars are -- a collapsed burst stands
in for frames that are not on screen, and something has to say so whether or
not a pointer is nearby.

Folding changes what the grid's *query* returns rather than what its cells
draw, because the grid is a window over an ordered query and the frames a fold
hides are mostly not loaded. So the predicate joins VISIBLE in every query
that lists or counts cells -- the window, the header's count, the run a
shift-click resolves, and the ordinal a scrub lands on -- under the discipline
VISIBLE's own comment sets out: present in four places of five is worse than
absent, because the counts disagree with the cells and neither looks wrong on
its own. There is a test for exactly that.

`the_window_read_walks_the_ordering_index` now includes the burst clause. It
asserts on the query plan while holding its own copy of the query, so left
alone it would have gone on reporting green against a query the grid no longer
runs. If the clause costs `images_grid_order` and puts the sort back, that
fails here rather than becoming jitter someone measures in six months.

The pass keeps its own drain timer in a thread-local instead of taking fields
on the library controller, so everything the feature needs to run lives in one
file and the screen that starts it holds nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 23:20:26 +02:00
dtourolleandClaude Opus 5 d0b671d4db Put the grouping dials where the regrouping is
The merge probability was `dr_face`'s constant and the smallest group was
a bare `< 2` in the clustering pass. Both were tuned on one library —
1,813 faces of one photographer's family — and the quantity they optimise
is a property of the population, not of the model. A household at close
family resemblance and two thousand strangers at a wedding want different
answers, and neither of them is the reference library. The doc comment
already conceded the point and pointed at `face_index --tune`; a
photographer does not have a terminal.

So they are `FaceSettings` now, saved per device beside the cache budgets
and edited from the People screen — beside the Regroup button that
applies them and the rail that shows what they did, because a value
changed three screens away from its effect is one nobody can tune.

Moving them is safe by construction, which is why nothing asks for
confirmation: a regroup writes only the suggested half, and
confirmations, names and ignores enter as anchors and come back
unchanged. The smallest-group rule is applied only to groups the system
invented — a group the user named or set aside survives it whatever its
size, because a display preference does not overrule a judgement.

**Withdrawal, without which the setting does nothing visible.** Raising
the smallest group stops the pass creating small groups; it does not
remove the ones a previous pass made, because those still hold their
suggestions, so they are not empty, so the prune leaves them. The pass
now releases every unanchored face it did not place before pruning.

And a dial you cannot see the effect of is not a dial. "What would this
do?" runs the same population through the clusterer without opening a
transaction and reports groups, faces grouped and largest group — one row
of `--tune`'s table, on the user's own library, on a worker thread. The
line leads with the group count because that is the number that says
which side of the right setting you are on: it climbs as fragments are
gathered into people and falls as separate people start being welded,
while the grouped-face count rises straight through both.

The preview parks its poll timer in a slot of its own. A preview and a
regroup are allowed to be in flight together, and sharing the sweep's
single slot would have the second to start drop the first's timer —
visible as a Regroup that finished on its worker and never said so.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 23:18:46 +02:00
dtourolleandClaude Opus 5 e465ff0c80 Put the people filter where the filters are
Narrowing the grid to two people at once has worked since people became
a selector term, and it was effectively unreachable. The only control
that could add a second person lived on the People screen, behind
selecting them there, and it appeared only once the grid was already
narrowed to somebody — so "photographs with both of them" needed a
two-screen round trip the user had to guess at.

A filter belongs on the filter bar. A "People" chip there opens a tray of
everyone the library knows; tapping a name adds or removes them, and the
any/all chip beside it — already there, and already the thing nobody
found — now has something to sit next to that explains it. The caption
leads the row so a pair of chips means something before either is
pressed.

The tray is a strip under the bar rather than a popup, the way the
develop column's film picker is: the view scrolls as one, so an inline
strip is taller content and not a second overlay to dismiss. It scrolls
horizontally for the same hard reason the bar above it does — a layout
cannot be narrower than its children's minimums, and forty people would
otherwise set the minimum width of the whole view.

The roster is built on open, not kept in step: indexing and regrouping
change who exists, and a list cached at startup would be stale for
exactly the user who has just been naming people. Named first, then by
how much of them the library holds — the catalog orders by face count
alone, which puts a dozen unnamed strangers ahead of the two people the
user actually cares about.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 23:18:07 +02:00
dtourolleandClaude Opus 5 6d5de21fb7 Tell a tap on a photograph from a hand going past
A brush across the grid opened whichever photograph was under it. Travel
was already answered — the Flickable claims the pointer and the press is
cancelled — but a contact that neither travels nor lasts reaches a
TouchArea as an ordinary press and release, and it was landing the user
in develop.

So a finger now has to stay down for `TAP_MIN_MS` before letting go
counts as opening anything. That is the floor under a tap where the 450 ms
`HOLD_DELAY_MS` is the ceiling: below is a graze, between is a tap, above
is a hold that starts a selection. One scale, three gestures.

Only a finger is held to it. A mouse click is a discrete decision made by
a button and is routinely over in thirty milliseconds, so `cell-pressed`
now reports whether a finger did it — the same finger-id convention the
pinch arbitration beside it already uses — and the dwell applies to touch
alone.

A graze still *selects* the cell it landed on, because the press already
did that. That is the right failure mode: something visible and
reversible rather than a silent nothing, and rather than develop.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 23:18:07 +02:00
dtourolleandClaude Opus 5 f5edd49b6b Let a manual collection be put in the order it is meant to be seen in
`collection_members.position` and `Sort::CollectionPosition` have been in the
catalog since collections were, and nothing above dr-catalog has ever written
or read either: `collections::set_order` had no callers, and the grid ordered
everything by capture time whatever it was scoped to — dr-ui does not construct
a `Query` at all, it has its own `GRID_ORDER` constant. So a manual collection
was a set with an order nobody could see or change.

Three pieces, because it could not be fewer:

`grid_order_for` decides the ordering from the scope, and both readers take it
from there. That is the load-bearing part. An ordinal only names a photograph
relative to an ordering, so the window read and the span read have to agree —
a shift-click resolved through a different ORDER BY than the cells were drawn
with selects a different run than the one on screen, and the user finds out
when the export runs. `read_ids_span` already stated that invariant about
`GRID_ORDER`; this widens it to an ordering that depends on the scope.

Only a single manual collection has one. A set draws its descendants' images
too, and two children's positions are unrelated integers that interleave
arbitrarily; a smart collection has no member rows to carry a position at all.
Both fall back to capture time and refuse the drop rather than pretending.

The drop is on the cell, on whichever half of it the finger landed — the
trailing edge is the only way to name the last place in a collection, since
there is no cell beyond the last one to drop in front of.

`reordered` is pure and the membership is rewritten whole. `set_order` sets the
positions it is given and leaves the rest, so a partial write would interleave
the moved run with rows nobody touched; and it is read unfiltered, so what the
filter is hiding keeps its place relative to what the user can see.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 23:18:06 +02:00
dtourolleandClaude Opus 5 9d11554c71 Make the range gesture visible, and offer the whole grid at once
Touch has had a range gesture for as long as selection mode has: double-tap the
far end. It is invisible, it is unreliable on a grid that scrolls under the
second tap, and it extends from the anchor *before* the two taps moved it — a
rule subtle enough that the code needs two paragraphs to explain it to itself.
Nobody who was not told about it has ever used it.

"Select to…" is the same operation with state you can see. Press it, the strip
stops reporting and says "Tap the last photograph", and the next cell taken is
the far end. It reaches Rust as shift on `cell-pressed`, so it lands in
`apply_press` as the ctrl+shift it already is, and there is no third selection
policy to keep in step with the other two.

This is deliberately not the sweep gesture. A drag that paints cells can only
reach what is on screen, and the ranges that hurt on a tablet are longer than a
screenful — between the two taps here the user may scroll as far as they like,
and the run is resolved by the catalog rather than by what happened to be
loaded. A sweep is still worth having for short runs; it is not what this
should have rested on.

"Select all" beside it, asked of the catalog for the same reason: a select-all
that quietly meant "the hundred cells that happen to be loaded" is a lie the
user cannot see until the export runs.

The double-tap stays. It is tested, and an accelerator that costs nothing is
worth keeping for whoever has already learnt it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 23:18:01 +02:00
dtourolleandClaude Opus 5 d1f3b97245 Ask for the collection's name where the keyboard can reach it
"New collection from selection" created the collection under a placeholder
name and then opened the rename field in the sidebar tree.

On a tablet the sidebar is not on screen. It is instantiated all the same —
app.slint collapses it to zero width and `visible: false` rather than using an
`if`, because an `if` there is a layout loop Slint panics on — so the rename
field was created, its `init` took focus, and Android raised the on-screen
keyboard over a box nobody could see. Nothing else on the screen is focusable,
so the keyboard had nowhere to go: it stayed, the name could not be typed, and
the collection was already written under the name the user did not want.

Asked in a sheet instead, on the same card as the filing and keywording sheets,
before anything is written. That also fixes what was hiding behind it: an
abandoned rename used to leave a "New collection" in the tree, because the
collection existed before the name did.

`Field` gains `take-focus()` so a sheet whose field is the only thing to do in
it can answer the keyboard for the user — a function rather than a property,
because focus is an event and a bound property would re-take it on every
unrelated re-evaluation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 23:17:40 +02:00