Merge: touch selection and drag, from the gallery-selection branch

Verified before merge: fmt clean, clippy -D warnings clean, 563 dr-ui tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

# Conflicts:
#	docs/traceability.md
This commit is contained in:
2026-08-30 13:45:18 +02:00
29 changed files with 1839 additions and 304 deletions
+18 -9
View File
@@ -5,7 +5,7 @@
Every entry here is extracted from the comment beside the code that implements it, so this file cannot describe a gesture the application does not have. Add one by writing a `GESTURE:` block next to the implementation; there is nowhere else to write it.
15 gestures, in 2 places.
16 gestures, in 2 places.
## People
@@ -73,6 +73,15 @@ While selecting, a tap never opens. That is the whole point of the mode: one mea
<sub>`ui/dr-ui/ui/library.slint:1291`</sub>
### Pick a photograph up to drag it
- **Touch** — Press and hold it until a ring opens around it, then drag
- **Pointer** — Drag it
A finger on a photograph might be starting a scroll, and for the first half-second the grid assumes it is. Holding says otherwise, and the ring is the grid saying it heard — from there the drag cannot be lost to a scroll. A mouse never waits: the cursor is precise enough that a sideways drag is unambiguous from the first pixel.
<sub>`ui/dr-ui/ui/library.slint:1321`</sub>
### Select a range
- **Touch** — While selecting, press "Select to…", then tap the last photograph of the run
@@ -80,7 +89,7 @@ While selecting, a tap never opens. That is the whole point of the mode: one mea
This replaced a double tap, which had no visible state and could take forty photographs by accident. The run is resolved by the catalog rather than by what is on screen, so the grid can scroll between the two taps — the ranges that hurt on a tablet are longer than a screenful, which is exactly where a finger sweep runs out.
<sub>`ui/dr-ui/ui/library.slint:1352`</sub>
<sub>`ui/dr-ui/ui/library.slint:1386`</sub>
### Find photographs with two people in them
@@ -89,7 +98,7 @@ This replaced a double tap, which had no visible state and could take forty phot
"Any of them" is a union and "all of them" is an intersection. The tray is where both terms and the choice between them live, because a filter belongs on the filter bar.
<sub>`ui/dr-ui/ui/library.slint:2093`</sub>
<sub>`ui/dr-ui/ui/library.slint:2141`</sub>
### Resize the thumbnails
@@ -98,7 +107,7 @@ This replaced a double tap, which had no visible state and could take forty phot
There is no wheel on a tablet, so without the pinch the cell size could only be changed by a control a finger cannot reach.
<sub>`ui/dr-ui/ui/library.slint:2719`</sub>
<sub>`ui/dr-ui/ui/library.slint:2783`</sub>
### File photographs in a collection
@@ -107,7 +116,7 @@ There is no wheel on a tablet, so without the pinch the cell size could only be
The selection is what the drag carries, which is why selecting several is worth the mode: forty photographs file in one gesture.
<sub>`ui/dr-ui/ui/library.slint:2878`</sub>
<sub>`ui/dr-ui/ui/library.slint:2942`</sub>
### Open a photograph
@@ -116,7 +125,7 @@ The selection is what the drag carries, which is why selecting several is worth
A tap opens; a tap that *moved* does not. Travel is what separates a deliberate tap from a hand brushing past, and it is the only thing that does: the two are the same length. An earlier version required the finger to dwell 120 ms instead, and that rejected ordinary taps — a real tap is often quicker than a brush.
<sub>`ui/dr-ui/ui/library.slint:3133`</sub>
<sub>`ui/dr-ui/ui/library.slint:3210`</sub>
### Rate a photograph without opening it
@@ -126,7 +135,7 @@ A tap opens; a tap that *moved* does not. Travel is what separates a deliberate
A star has to take the press without it also reaching the cell, or every rating throws the user into develop.
<sub>`ui/dr-ui/ui/library.slint:3245`</sub>
<sub>`ui/dr-ui/ui/library.slint:3330`</sub>
### Drop the selection but keep selecting
@@ -135,7 +144,7 @@ A star has to take the press without it also reaching the cell, or every rating
Distinct from Done, which leaves the mode entirely. Clearing keeps it, so the next selection can start straight away.
<sub>`ui/dr-ui/ui/library.slint:3808`</sub>
<sub>`ui/dr-ui/ui/library.slint:3938`</sub>
### Select everything the grid is showing
@@ -144,4 +153,4 @@ Distinct from Done, which leaves the mode entirely. Clearing keeps it, so the ne
A scoped grid of two hundred frames is two hundred taps otherwise, and "all of them, except those three" is a far more common shape than the taps it took to say it.
<sub>`ui/dr-ui/ui/library.slint:3825`</sub>
<sub>`ui/dr-ui/ui/library.slint:3955`</sub>
+66
View File
@@ -435,3 +435,69 @@ dependency detail (`core/dr-segment/models/LICENCE.md`).
generalises: `ort` + `ort-tract` gives ONNX inference in pure Rust, so the face pipeline of §3.9.1
needs no C dependency either. The *model licensing* half of D13 is untouched — the InsightFace
weights are still non-commercial and still unusable here.
---
## 16. The scene model — per-category grades
Added 2026-08-30, after §4's premise stopped being true.
### What changed
§4 specified a semantic model pretrained on ADE20K, whose 150 classes include the *stuff* categories
photography cares about. §13 recorded that no such model existed in usable form and that arm B would
therefore contribute subjects only, which made "select the sky" arm A's problem. Re-checked
2026-08-30: **Ultralytics now ships a `semantic` task with ADE20K checkpoints**
(`docs.ultralytics.com/tasks/semantic`). `yolo26s-sem-ade20k` is in `models/scene/`.
### It is an addition, not a correction to arm B
The instance model stays exactly where it was, and the reason is the one §13 already gave and was
right about: a semantic model merges every pixel of a class into one region, so it cannot separate
two people, and separating two people is what clicking a subject requires. Swapping arm B for this
would regress the primary interaction to fix a secondary one.
So the two divide by *what the user is doing*, not by which is better:
| | `models/segment/` (COCO instances) | `models/scene/` (ADE20K semantics) |
|---|---|---|
| Question | which pixels are *that* dog | how much of this pixel is sky |
| Granularity | per instance | per category, whole frame |
| Drives | local adjustments, subject selection | the scene tab's per-category sliders |
| Vocabulary | 80 things | 150 classes, stuff included |
### The export is truncated, and both reasons matter
Ultralytics ends the graph with `Resize → ArgMax → Cast`, returning a `[1, 640, 640]` u8 label map.
`tools/export-seg-model.sh` cuts that tail and ships the classifier's `[1, 150, 80, 80]` f32 logits.
**Cost.** The `Resize` materialises 150 × 640 × 640 × f32 — 246 MB — and the `ArgMax` then reduces
across the channel axis, striding 409,600 elements per comparison. Measured under load it was
roughly four fifths of total runtime, spent on work the application discards.
**Softness, which is the more important one.** `ArgMax` destroys the per-class scores, and the whole
design of the scene tab rests on keeping them. Softmax over the 150 channels, summed within each
category, produces per-category weights that sum to one at every pixel — a partition of unity.
Feathering that cannot double-grade a boundary. Feathering *hard labels* outward from two adjacent
categories paints both grades into the overlap, and every horizon in the frame acquires a seam.
### The resolution is 80×80, and no setting changes that
The discarded upsample was never information. `Scene` keeps the native grid and resamples on demand,
so the coarseness is visible in the type rather than hidden. A graduated grade over sky or water is
untroubled by it; a rooftop against sky at 100% zoom will show it. This is the constraint most likely
to decide whether the tab feels good, and it is not addressable by choosing a larger checkpoint —
`yolo26n-sem` and `yolo26s-sem` have the same output grid.
### Licence
Unchanged. Same AGPL-3.0 grant as the instance model, same GPLv3 §13 permission, same consequence
already accepted in D14 — so this needed no new licence decision, which is most of why it was cheap.
See `models/LICENCE.md`.
### Measurement
Timings taken while this was chosen came off a laptop compiling other things and are upper bounds
only. `cargo run -p dr-segment --example scene --release --features embedded-scene-model` reports a
median over N runs with the first excluded; a number worth quoting should come from that, on an idle
machine.