Commit Graph
9 Commits
Author SHA1 Message Date
dtourolleandClaude Sonnet 5 b3641b5307 Let a subject's mask be refined, and let several masks be edited at once
Two related gaps in the mask panel, from the same conversation: a subject's
outline is only ever as sharp as the whole-frame pass that found it, and a
change meant for several layers had to be dragged once per layer.

## Refine mask

A "Refine mask" button on a subject layer re-runs detection on a padded crop
around that instance's own box instead of the whole frame — the subject
reaches the model at its own size rather than squeezed into the model's fixed
640x640 window alongside everything else in the photograph. `RefineJob`
mirrors `SegmentationJob`'s split (built on the session, run off it, adopted
back), and the crop itself is rendered through `Framing::set_view` — the same
ephemeral viewport the interactive zoom already uses to render a region above
proxy resolution, so no new render path and no change to the model's own
input size was needed. `dr-segment` is untouched: `Tiling::Whole` already
treats whatever buffer it is handed as the one window.

The result is still downsampled onto the shared proxy grid every instance's
mask lives on, but from a sharper source than the whole-frame pass ever saw
for that subject, which is what the edge actually reads out of.

## Multi-select

`active_mask: Option<String>` is now `active_masks: Vec<String>`. A plain
click still replaces the selection; a control- or command-click toggles one
layer in or out of it. `set_param` and `reset_op` fan out to every selected
layer, each set to the exact value the slider now shows rather than offset by
however far it already was — one slider, one reading, applied everywhere
selected. Dragging a gradient's on-canvas handle is deliberately not
extended to multi-select: several gradients have no single geometry a shared
handle could move, so `gradient_handles` stays empty unless exactly one
layer is selected.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-23 13:17:07 +02:00
dtourolleandClaude Opus 5 7fcb8dc107 Give the develop column room, and the mask a way out of the way
Three faults reported together, and they share a shape: each is something the
panel decided on the photographer's behalf.

**The column was 280px on every screen.** That width was chosen for a tablet,
where the column is a large fraction of the display and every pixel of it is
taken from the photograph. On a desktop window the mode strip alone — two
modes, a separator, "All", and a chip per attribute the operation set declares
— does not fit, so it scrolled sideways. A control you have to pan to reach is
one you do not know is there. 380px when the window is classed expanded, 280px
when it is not; driven off `layout-class` because reading the window width
inside the layout that sets it is a binding loop.

**The overlay could not be hidden.** An earlier "Overlay" button was removed
for a good reason — it *armed* the overlay, so local mode could be entered and
still show nothing. Hiding is the opposite need and was never served: a mask is
judged against the photograph beneath it, and that photograph is exactly what
the overlay covers. `overlay-hidden` is kept separate from `overlay-on` so a
recompute cannot switch the overlay back on under someone who just turned it
off.

**The masks were coarse because the model saw the subject small.** The graph's
input is a fixed 640x640 and every frame is letterboxed into it, so a bird
200px across in a 1600px proxy reaches the model at 80px. `Tiling::Grid` has
been implemented and tested since the segmentation spike and defaulted off,
because it costs one inference per tile — 2.8s for a 3x2 grid against 470ms.
Now offered as "Look closer (slower)", which says what it costs, rather than
spending it on every image or on none.

The tiling choice enters the segmentation signature. A mask stores the
signature its region ids index into, and a tiled run finds different instances
in a different order; sharing a signature would silently reinterpret a layer
built against the coarse pass — a wrong mask rather than a stale one, and
nothing announces it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-23 10:54:35 +02:00
dtourolleandClaude Opus 5 c0e1179936 Find the subjects without stopping the window
"Find subjects" took the UI thread for two thirds of a second on a 22 MP
frame — a proxy render, a readback and a YOLO pass through `ort` — and for
that time the interface was simply gone. The panel apologised for it rather
than hiding it: a "Looking…" label, and a 16 ms `single_shot` so the label
reached the screen before the freeze began, with a comment saying the obvious
fix needed the develop session restructured and was not being taken.

The obstacle was never `Send`. `DevelopSession` is `Send` — the device, the
source texture and the passes all are. What cannot go to a worker is the
`Rc<RefCell<Option<DevelopSession>>>` that every callback in the window
reaches through, and the window has to keep reaching through it while the work
runs. Handing the session over would freeze the interface exactly as
thoroughly as blocking on it did.

So the work takes a copy of what it needs instead. A `SegmentationJob` is the
device, the demosaiced source behind an `Arc`, and the name of the session
that asked. Taking one is two `Arc` bumps; running one is 495 ms on this
desktop; none of it touches the session, and there is deliberately no
`&mut DevelopSession` in scope for a caller to hold across it. The proxy
render travels with it rather than staying behind — a `GpuContext` and a
texture handle are both `Send`, and the model was never the only expensive
half. So does building the mask rasteriser, which is a shader compile:
adopting the result was costing 23 ms, a dropped frame on the one redraw the
user is waiting for, and the rasteriser is needed exactly when the subjects
arrive and never before. What is left on the UI thread is a microsecond.

The answer comes back through a channel a `slint::Timer` polls, which is the
shape `apply_when_ready` already uses for a sidecar fetch.

**A result can outlive the photograph it describes.** Two thirds of a second
is long enough to press the button, think better of it and swipe to the next
frame — and the result landing then would fill the panel with subjects that
are not in the picture, drawing outlines around a dog two photographs back.
Nothing downstream can tell: the masks rasterise and the overlay draws either
way. So every session is minted with an id, a job carries the id it was taken
from, and `delivery` compares the two before anything is applied. An id
rather than a counter beside the session slot, because that slot is written
from four places in `lib.rs` and the fifth would be the one that forgot.

A discard touches nothing on the way out. `segmenting` belongs to whichever
photograph is open now, which may well have a run of its own going, and
clearing it would re-enable a button that is correctly insensitive.

One run at a time, and abandonment is what stops that being a trap. A job
left over from a photograph the user has left is displaced rather than waited
for — otherwise the next frame's "Find subjects" would do nothing for the
length of a run nobody wants, which is the wait this exists to remove. `ort`
offers no way into the inference, so abandoning is checked at the seams there
are: before the job starts, and between the readback and the model. Abandoned
early it costs nothing, abandoned mid-inference it costs the run it was
already committed to, and either way the answer is dropped at the channel.

`DevelopSession::segment` survives as a test-only convenience. Left public it
is precisely the shape that put two thirds of a second on the UI thread in the
first place, and the next caller would reach for it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 12:26:11 +02:00
dtourolle 37f63edbf9 Take the watershed out of the product path
It does not work on a photograph, so nothing should offer it. `Segmentation`
is now one model pass and what it recognised: no region field, no merge
tree, no label upload, no granularity slider, and no readback of the whole
proxy to build a graph that collapses.

A click means "the object under the cursor". The region-selection path went
with the hierarchy it indexed — including the shift-click add/subtract,
which has no meaning for a whole object and would have been a modifier that
silently did nothing.

The passes, the hierarchy and the semantic prior stay in `dr-gpu` and
`dr-segment`, tested and documented. It is the *merge criterion* that
fails — the saddle is the minimum gradient along a boundary, so one weak
pixel merges two regions and real gradient noise puts a weak pixel on every
boundary. That is one function to replace, and the evidence for replacing it
is worth keeping. What is gone is the wiring, the option, and the control
that offered a user a choice with no outcome.

`MaskSource::Regions` remains in the pipeline: it is tested, it round-trips
through the sidecar, and a stored layer that names regions must still load
and be reported stale rather than failing to parse.
2026-08-22 08:39:17 +02:00
dtourolle ec713585a5 Measure the distance to the edge, and get four controls for one transform
Feathering, growing, shrinking, closing and opening are the same number
read differently. With the signed distance from the boundary in hand,
dilation is the set where d >= -r, erosion where d >= +r, and a feather of
any shape is a function of d. So the field is computed once and the
controls are arithmetic on it.

The **field** is what reaches the GPU, not a finished alpha, and that is
the point: growing a mask or changing its falloff then costs a uniform
upload and no recomputation, which is what makes them live controls rather
than ones that stall on every drag. Only closing and opening rebuild,
because after the first threshold the shape has changed and the old
distances describe the old one.

Exact Euclidean, via Felzenszwalb's separable transform — not a chamfer
approximation, which leaves a mask visibly octagonal once grown more than
a few pixels. A test asserts the diagonal is √2 rather than 1 or 2.

It runs on the CPU, which ARCH §5.4 forbids for masks. The rule is about
brush lag — a stroke rasterised per frame — and this is a different
operation: once per mask edit, on input the model already produced here,
producing a field the GPU then samples for free. What it buys is exact
determinism, which matters because masks reach the sidecar as indices and a
field that varied by vendor would mean a mask meaning one thing on the
desktop and another on the phone.

The half-pixel in `signed_distance` is not a detail, and a test caught it.
Measuring to the nearest opposite pixel *centre* puts the smallest
magnitude at 1 either side, so the boundary is nowhere and **eroding by
less than a pixel removes nothing**. A control whose first notch does
nothing is a broken control. Half a pixel off each side puts the boundary
where it physically is, and eroding by 1 takes exactly the outermost ring.

Every falloff curve is 0.5 at the boundary by construction, asserted for
all five: changing the curve should change how the transition looks and
never where it sits.
2026-08-22 08:39:17 +02:00
dtourolle ee10097435 Mask the subject the model found, not the regions underneath it
The watershed hierarchy does not survive a photograph, so local masking
stops depending on it. A layer can now be one recognised object, and the
object's own coverage is the mask.

`Options::watershed` defaults off. It costs ~80 ms plus a full-resolution
readback to produce a ladder that collapses, and paying that on every
photograph buys a control that misleads. Kept switchable rather than
deleted: the passes and the hierarchy are correct in themselves and it is
the merge criterion that fails, which is a change to one function.

Masks now rasterise in **source** space at proxy resolution and are sampled
by the composed shader after the framing map. That fixes a real bug: they
were rasterised in output space, so zooming slid the photograph underneath
a mask that stayed pinned to the viewport, and cropping moved every
adjustment to a different part of the picture. Doing it this way also
leaves the framing map in exactly one place — a second copy in the mask
shader would have been a second thing to keep in step, failing only when
straightened.

A subject is stored as identity, not pixels: the mask is megabytes and is
reproducible by running the same model over the same image, so the sidecar
carries the index, the class and the score, and the session carries the
pixels. The class is there to be checked — if instance 3 comes back a "car"
where it was a "dog", something changed and the layer is stale rather than
silently masking the wrong thing.

The overlay now draws instances and is transparent everywhere else. The
region version covered every pixel and so hid the photograph it was drawn
over; the question it exists to answer is whether an outline follows the
subject, which you can only answer by seeing both.

`examples/local.rs` is the worked example: subject in colour with the rest
monochrome, and the subject lifted out of its background. Run on a 5472x3648
CR2 it finds two people and two cars, and the colour-pop keeps her hat and
hair while the wall and grass behind go grey.
2026-08-22 08:39:17 +02:00
dtourolle b1433ad4a9 Give a mask an edge treatment, and find out the watershed has none worth having
Two things, and the second is why the first matters more than expected.

Mask layers gain a feather, a falloff curve and a morphology, all defined
against a signed distance from the boundary rather than as separate
features — one exact distance field answers "how soft" and "how far" at
once, so dilation is a threshold at -r, erosion one at +r, and closing and
opening are one of each in sequence. The compound pair costs a second
distance field, which is why they are named rather than presented as a
radius that happens to be signed. Types, defaults and sidecar round-trip
only; the field itself is next.

`edge-feather` and `edge-falloff`, not `feather` and `falloff`, because a
radial mask already writes `feather` for the fraction of its radius it
ramps over. Same word, different quantity, different units — sharing the
key would have made an existing file ambiguous.

The diagnostic that provoked this is committed as an ignored test, because
"does the ladder land on things a person means" is the question S15 exists
to answer and it should not depend on whoever still has the script. On
bus.jpg it answers badly: 35,075 regions at blur 2 over an 810x1080 frame,
and cutting that to 400 gives *one* region covering nearly the whole
picture plus 399 noise specks. Not over-segmentation — collapse. Almost
every saddle is near zero, so the merge order joins everything meaningful
before it joins anything spurious, and a global cut spends its entire
budget on grain.

So the granularity ladder does not currently work on a photograph, and the
region masks built on it inherit that. Recorded rather than worked around:
the next commits move local masking onto the model's instances, where the
edge treatment above is what makes a quarter-resolution mask usable.
2026-08-22 08:39:17 +02:00
dtourolle 12d320cf33 Record what the spec got wrong about the model that exists
docs/segmentation.md §4 priced arm B as costing a C dependency under the
NDK and treated that as most of the difference between the arms. It is not
a cost that has to be paid: `ort`'s `alternative-backend` disables its
linking entirely and `ort-tract` supplies the API from tract, which is pure
Rust. D13's "largest exception the policy would tolerate" turns out not to
be needed, and the answer generalises to the face pipeline — so D13's
runtime half is now answered and only its licensing half is open.

Three findings contradict §4 outright and are recorded as F4-F6 rather than
quietly designed around. There is no ADE20K-trained YOLO, so the shipped
vocabulary selects subjects and not stuff — "select the sky" comes from the
watershed or from nowhere. It is instance segmentation, so it partitions
nothing and two people come back as two instances. And tract cannot parse a
dynamic-shape export, which fixes the input at 640 square and makes tiling
the only route to more semantic resolution.

Arm C ships, but §8's criteria are not what decided it, and saying so
matters more than claiming the process worked. §8 asked for a two-
interaction margin over arm A on a traced corpus. That comparison was never
run: F4 and F5 changed what the arms are, and a model that recognises
subjects but has no word for sky cannot be a selection tool alone, while a
watershed cannot tell a person from the wall behind them. They stopped
being candidates and became complements.

What is *not* done is written down as plainly: the 24-image corpus is
untraced, so M1-M4 have no numbers and "this feels right" has not become
one. M5 is answered on one device only, and region ids now reach the
sidecar — so a cross-vendor divergence would mean a mask written on the
desktop meaning something else on Android. F3 stands.
2026-08-22 08:39:17 +02:00
dtourolle 5ecb35864f Put the region map behind the sliders that were already there
A mask layer holds a real develop chain, so the develop panel can edit one
with no new controls: select a layer and the same sliders read and write
its chain instead of the graph's. An operation declared in `ops/` tomorrow
becomes locally adjustable by existing, which is the payoff for making a
layer a chain rather than a handful of special-cased parameters.

`segmentation.rs` joins the two arms into the one thing the view needs.
The model reads the image through a neutral graph rather than the edited
one, so a segmentation survives an exposure change instead of being
invalidated by every slider. Arm B failing is not fatal: a missing or
unreadable model leaves a working watershed map, because refusing to
segment at all would trade a working feature for a strict one.

The overlay colours groups by a golden-angle walk over hue. Deterministic
rather than random, so a region keeps its colour across a level change and
the eye can track it; boundaries drawn black over the fill, because two
adjacent groups landing on near hues read as one region and telling them
apart is the whole reason to look at it.

Clicking the photograph creates the layer if none is selected — that is how
a local adjustment begins, and making the user press "add layer" first
would be a step with no decision in it. Shift-click extends, and clicking a
region already selected removes it, so one gesture both adds and corrects.

`segment-readback` is a new dr-gpu feature and not a loosening of
`readback`. The region-graph transfer is once per image on a worker; the
one AC-8 forbids is per frame in the render loop. Sharing a switch would
have forced a build wanting local masking to unlock the other. F3 still
stands and the feature name says so.
2026-08-22 08:39:17 +02:00