Commit Graph
16 Commits
Author SHA1 Message Date
dtourolleandClaude Opus 5 4f4abd335f Cut a scene category back to the pixels that agree with it
A flag in the sky came out weighted as sky, and no feather setting fixed it.

The scene model's logits are `[1, 150, 80, 80]`, so one cell is eight input
pixels; at the 1600px proxy the letterbox scale is 0.4 and **one cell is 20
proxy pixels**, which `rasterise`'s bilinear then spreads across one more
either side. A flag is a handful of cells whose softmax is dominated by the
sky around it. The information was never in the grid, so nothing downstream
of the grid can recover it.

Tiling is the answer for an instance and is not available here: a category
has no bounding box to tile over — sky is wherever the sky is. But the
photograph is at full proxy resolution even though the weights are not, and
it knows exactly where the flag is. So the model says *what*, and the
pixels say *which of them*, which is the division of labour arm C already
draws between the instance model and the watershed.

## Seeds, and why the erosion radius is not a guess

Threshold the weights high, take `signed_distance`, and keep what is more
than 1.5 cells inside. One cell *is* the model's resolution and the bilinear
spreads it across one more, so the band either side of the boundary is smear
rather than evidence. Deriving the radius from `Scene::cell_pixels` rather
than picking a pixel count means it stays right if the proxy edge or the
export changes.

The mirror of that set is a confident *exterior*, free from the same field.

## Dropping small modes is the step that makes it work

Four k-means modes per side, not one Gaussian: sky is blue at the zenith,
white where the cloud is and pale at the horizon, and one blob over all
three rejects two of them.

Then modes holding under 3% of a side are discarded, and without that step
the whole thing fails on the case it was built for. A small flag deep in the
sky has both a high weight and a large distance from the boundary, so it
lands in the interior sample and teaches the model its own colour. It cannot
be excluded geometrically. It can be excluded by share.

Luminance is weighted at a quarter against chrominance for the same reason
the watershed's gradient is. Sky's variance is dominated by luminance, so at
equal weight the distribution is a long bright streak that a mid-grey flag
sits comfortably inside. A flag is separated by chrominance; a cloud is
separated by luminance alone. Not zero, or a dark bird against a bright sky
survives.

## Two tests, because either alone is wrong

Absolute — is this colour plausible under the category, as a chi-square on
the Mahalanobis distance. Comparative — is it likelier inside than outside.
A pixel must pass both.

The absolute test is what catches the flag, whose colour is far from *both*
sides and which the comparative test alone would leave at even odds. The
comparative test is what stops the absolute one needing a constant tuned per
category.

## What this cannot do, written down rather than left to be discovered

An intruder large enough to hold its own mode is kept. By share, a flag over
a fifth of the sky and a cloud bank over a fifth of the sky are the same
object, and colour does not separate them either — a white cloud is as far
from blue sky in chrominance as many intruders are.

So `min_cluster` is not a threshold with a correct value waiting to be
found; it is the trade-off itself, set where a photographic intruder falls.
Both ends are pinned by tests — `a_flag_in_the_sky_is_removed` and
`an_intruder_larger_than_min_cluster_survives` — so that moving the number
reads as moving the trade-off rather than as fixing a bug. The case left
open is a large unrecognised object in a clean category, which wants the
boundary snapped to watershed basins and is a different mechanism.

## Safe to apply without a control

It is subtractive: the output is the input times a factor in `0..=1`. The
worst failure available to it is losing part of a real sky, never gaining a
region, so a blue car below the horizon that was never in the mask cannot be
pulled into it. And a factor in `0..=1` cannot raise a sum, so `scene.rs`'s
partition still holds when every category is refined independently — the
weight taken off the flag lands in the unlisted remainder, which is where a
flag belongs, ADE20K having no class for one.

Every path without the evidence to judge returns the weights untouched and
says which path it took. A refinement that silently did nothing is
indistinguishable from the feature being off, and an empty seed set fitted
to a distribution would reject every pixel.

The signature is deliberately unchanged: categories are addressed by name,
not by index, so a sharper mask cannot create the stale-index hazard the
signature exists to guard against.

The example writes `<prefix>-<category>-refined.ppm` beside the coarse one,
never instead of it — whether this is an improvement is a comparative
judgement and one image cannot answer it.

Verified: fmt clean, clippy -D warnings clean, 57 dr-segment tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 18:30:16 +02:00
dtourolleandClaude Opus 5 94b1b9e0cc Reconcile what four branches each built separately
Three seams, found by the first compile after the merge.

Two branches implemented 'can this scope be reordered' independently:
one from the collections model, one from the catalog through
orders_manually, on every re-read and excluding the trash. The second is
the better answer and is what survives; it only needed to set the
property app.slint declares.

Two lints from scene-mask-ui, which was merged mid-flight and had never
been through -D warnings: an is_none check spelled out where clippy wants
?, and a return in a cfg block's tail.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 16:51:41 +02:00
dtourolleandClaude Opus 5 763dfd353a Weigh the categories in the same precompute, and mask with them
The scene model shipped with a decoder and no caller. This runs it.

## Beside the instance pass, not instead of it

`compute` now does both on the same upright frame and lays both back down
the same way, so instance masks and category masks index into one grid —
the sensor's. A failure in the scene half is logged and dropped rather than
propagated: no scene model is an ordinary state, and a photograph that can
still be masked by subject should not become unopenable because the
categories are missing.

Categories under half a percent of the frame never reach the cache. A
control that does nothing when moved is worse than an absent one, and each
one it skips is a proxy-sized buffer not allocated.

## The shader needed nothing

A category reaches `dr-gpu` as a soft coverage buffer at proxy resolution,
turned into a distance field — which is exactly what a subject is. So they
share `MODE_SUBJECT`. That is not a shortcut taken for speed: the shader has
no way to tell them apart and no reason to want one. What differs is only
which model produced the coverage, and that has already happened by then.

Feather, falloff, dilation and erosion therefore work on a category on the
day it arrives, because they were never subject-specific.

## Where the weights come from

`scene-model` compiles the graph in and the desktop app takes it; Android
leaves it off and reads the copy `install_bundled_models` unpacks, because
24 MB of constant is worth avoiding in a mobile install and not worth the
plumbing to avoid on a desktop one. Embedded is tried first — a build that
has the weights compiled in should not be silently overridden by a stale
file in a data directory.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 14:02:22 +02:00
dtourolle d777f7f44d Merge branch 'master' into worktree-faces-scrfd-mbf
Build and test / Desktop (Linux) (push) Failing after 25s
Build and test / Layer separation (push) Successful in 22s
Traceability / Requirement traces (push) Successful in 58s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Failing after 33m26s
# Conflicts:
#	docs/traceability.md
#	ui/dr-ui/src/develop.rs
#	ui/dr-ui/src/segmentation.rs
2026-08-27 11:57:38 +02:00
dtourolleandClaude Opus 5 ce201c7dd6 Name the two spaces a photograph lives in, so a turn cannot go the wrong way
Every orientation bug this codebase has had has been the same bug: a turn of
the right size applied in the wrong direction. That failure is worth naming
precisely, because it does not look like one — a quarter turn applied
backwards lands 180 degrees from right, so the result is a plausible
transform of the picture rather than anything obviously broken, and on
landscape frames it is not wrong at all. It was the straighten shear, and it
was the segmentation overlay, and each time it was found by eye rather than
by a test.

The reason it keeps happening is that "rotate 90 degrees clockwise" cannot be
checked by reading it. The reader has to hold in their head which of the two
images is being rotated and which way the y axis runs, and there were four
hand-written copies of the permutation to hold it for: the shader prologue,
its CPU twin, the thumbnail path, and the segmentation.

So nothing added here says clockwise, anticlockwise, horizontal or vertical.
The functions say *which space they take and which space they return* —
`into_shown` and `into_stored`, `source_pixel` and `shown_pixel`,
`into_shown_rect` and `into_stored_rect` — and each takes the dimensions of
the space it reads from, so no caller has to work out which pair it is
holding. `StoredRect` and `ShownRect` are separate types because they are the
same four numbers meaning different things, which is exactly the case where a
mistake is silent: a shown rect measured against stored dimensions produces a
rectangle in the wrong place, not an error.

Underneath there is one permutation. `source_pixel` was already shared by the
prologue and the thumbnails; `source_point` is its normalised twin, written
beside it so the two cannot drift, and everything else is those two read
forwards or backwards. `Orientation::inverse` is the group inverse rather
than `4 - turns`: mirrors apply after the turn, so undoing means undoing them
first, and a mirror seen from the far side of an odd turn is about the other
axis. That is the diagonal-mirror case, tags 5 and 7, and getting it wrong
renders as — again — 180 degrees.

Three call sites lose their own copy: the thumbnail path, `dr-ui`'s
segmentation, and `dr-gpu`'s `local` example. "Upright" now means one thing
across the application rather than one thing per caller.

The gate that matters most is `the_render_and_the_orientation_map_agree`. The
shader prologue and `Orientation` answer the same question by different
routes, and until now nothing checked that they answered it the same way. It
now checks every EXIF tag against every user rotation and mirror on top of
it, because the composition is where the two could agree singly and disagree
together.

The rest earn their place by having caught something. Writing these found two
real errors in this commit's own new code before it ran anywhere: `shown_pixel`
was handed the dimensions of the wrong space and overflowed, and the rect map
turned the wrong way for the diagonal mirrors. A round trip that returns what
went in is the only check worth having here, since every wrong answer is
still a picture.

No behaviour changes. The permutations are the ones that were already being
applied; they are simply applied from one place now.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 09:46:24 +02:00
dtourolleandClaude Opus 5 b55812812a Name segmented people from the faces already recognised in them
The segmenter knows it found a person; the face index knows which person.
Joining them turns "person" in the mask list into "Anna", which is the
difference between a vocabulary of eighty COCO classes and one that
includes the user's family. Selecting a subject in a group photograph
stops being a guessing game between three identical rows.

Containment, not IoU. A face is a small part of the person it belongs to,
so a correct pairing has an IoU near zero and anything IoU-based would
reject every true match.

Confirmed names only. A suggestion is the system's guess, and printing a
guessed name onto a mask region would launder it into a fact.

Writing the tests corrected the design once: a tight head-and-shoulders
portrait, where the face fills most of the person box, is the case where
naming is most certain, not least. An earlier guard rejected exactly that
and has been removed, with the reasoning left as a test because it is
easy to get backwards a second time.

The names hang on the develop session, set when the image opens because
that is the one moment the catalog and the image id are both in reach.
Every segmentation run afterwards picks them up for free, and a library
with no face indexing behaves exactly as it did before.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 23:10:46 +02:00
dtourolleandClaude Opus 5 4a82753d22 Show the detector the photograph, not the sensor's scanlines
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
Build and test / Desktop (Linux) (push) Successful in 19m7s
Build and test / Layer separation (push) Successful in 25s
Traceability / Requirement traces (push) Successful in 23s
Build and test / Android (aarch64) (push) Failing after 33m10s
"Find subjects" was handed the proxy in the sensor's own orientation, so
every frame shot on a body held sideways reached the model lying on its
side — and a model trained on upright photographs is very bad at those.
Measured end to end on a 22 MP frame of two people and a dog: `person
0.36` and nothing else, against `dog 0.82, person 0.61, person 0.49` for
the same pixels stood up. Nothing failed; the panel simply offered one
poor subject where there were three good ones.

The orientation was never dropped on purpose. The proxy is deliberately
rendered through a *neutral* graph — the detection has to survive an
exposure change, or every slider would invalidate the masks built on it
— and neutral took the file's orientation with it along with everything
else. Landscape frames were unaffected, which is why it stood for as
long as it did.

The turn is `Orientation::source_pixel`, the same function the grid's
thumbnails already go through, so the detector and the thumbnailer now
agree about which way is up rather than holding two opinions. What it is
turned by is `Framing::effective_orientation` — the file's EXIF tag and
the photographer's own rotations composed into one permutation, by the
group law rather than by adding the turns, which is a distinction
`Framing` already had to make and had already tested. Rotating the
picture and pressing the button again therefore does what it looks like
it does.

The proxy stays in sensor space and the masks come back into it. That is
not a detail to be tidied later: the generated shader samples the mask
array at `uv_src`, *after* the framing map, so a mask stored upright
would sit a quarter turn off the subject it was drawn around. That is a
wrong mask rather than a weak one, and nothing announces it. So the
picture is stood up for the model and laid back down for everything
else, and `upright`/`lay_down` are returned as a pair because calling
one and forgetting the other is silent.

Both directions are the one function: `upright` gathers through
`source_pixel` and `lay_down` scatters through it. A quarter turn is a
bijection of the pixel grid, so the round trip is exact — no filter, no
resampling, and no hole to fill — and an inverse written out by hand
would be a second thing to keep in step, whose way of being wrong is a
mask mirrored about the wrong axis, which still looks like a mask.

The orientation joins the confidence and the tiling flag in the
segmentation signature, and for the same reason: turning the photograph
changes what the model recognises, so two runs either side of a rotation
are different instance lists. Two that happened to come out the same
length would otherwise share a signature and a stored layer would be
silently re-indexed from one into the other.

The refine pass had it too — it re-runs the model over a crop rendered
in the same sensor space — so it makes the same turn, and would
otherwise have handed back a worse mask than the one it was asked to
improve, on the subject the photographer had just pointed at.

`dr-gpu`'s `local` example is fixed with it. It exists to be the
shipping path with pictures attached, and a diagnostic that reproduces
the bug it is meant to catch is a trap for whoever reads it next.

Seven tests. The round trip is the identity over all eight EXIF tags on
a non-square asymmetric grid; a turn carries whole pixels rather than
shearing the channels apart; a sideways frame reaches the model
upright; a box comes back in sensor pixels, worked out by hand for the
one turn a portrait frame actually writes; a restored box still reads
low-to-high for every tag, since the rest of the pipeline takes
`x1 - x0` without checking the sign; and the eight tags cannot collapse
into one signature key. The existing composition test now runs against
`effective_orientation` itself, over all 8 x 16 baseline-and-user pairs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 22:40:58 +02:00
dtourolleandClaude Sonnet 5 b3641b5307 Let a subject's mask be refined, and let several masks be edited at once
Two related gaps in the mask panel, from the same conversation: a subject's
outline is only ever as sharp as the whole-frame pass that found it, and a
change meant for several layers had to be dragged once per layer.

## Refine mask

A "Refine mask" button on a subject layer re-runs detection on a padded crop
around that instance's own box instead of the whole frame — the subject
reaches the model at its own size rather than squeezed into the model's fixed
640x640 window alongside everything else in the photograph. `RefineJob`
mirrors `SegmentationJob`'s split (built on the session, run off it, adopted
back), and the crop itself is rendered through `Framing::set_view` — the same
ephemeral viewport the interactive zoom already uses to render a region above
proxy resolution, so no new render path and no change to the model's own
input size was needed. `dr-segment` is untouched: `Tiling::Whole` already
treats whatever buffer it is handed as the one window.

The result is still downsampled onto the shared proxy grid every instance's
mask lives on, but from a sharper source than the whole-frame pass ever saw
for that subject, which is what the edge actually reads out of.

## Multi-select

`active_mask: Option<String>` is now `active_masks: Vec<String>`. A plain
click still replaces the selection; a control- or command-click toggles one
layer in or out of it. `set_param` and `reset_op` fan out to every selected
layer, each set to the exact value the slider now shows rather than offset by
however far it already was — one slider, one reading, applied everywhere
selected. Dragging a gradient's on-canvas handle is deliberately not
extended to multi-select: several gradients have no single geometry a shared
handle could move, so `gradient_handles` stays empty unless exactly one
layer is selected.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-23 13:17:07 +02:00
dtourolleandClaude Opus 5 7fcb8dc107 Give the develop column room, and the mask a way out of the way
Three faults reported together, and they share a shape: each is something the
panel decided on the photographer's behalf.

**The column was 280px on every screen.** That width was chosen for a tablet,
where the column is a large fraction of the display and every pixel of it is
taken from the photograph. On a desktop window the mode strip alone — two
modes, a separator, "All", and a chip per attribute the operation set declares
— does not fit, so it scrolled sideways. A control you have to pan to reach is
one you do not know is there. 380px when the window is classed expanded, 280px
when it is not; driven off `layout-class` because reading the window width
inside the layout that sets it is a binding loop.

**The overlay could not be hidden.** An earlier "Overlay" button was removed
for a good reason — it *armed* the overlay, so local mode could be entered and
still show nothing. Hiding is the opposite need and was never served: a mask is
judged against the photograph beneath it, and that photograph is exactly what
the overlay covers. `overlay-hidden` is kept separate from `overlay-on` so a
recompute cannot switch the overlay back on under someone who just turned it
off.

**The masks were coarse because the model saw the subject small.** The graph's
input is a fixed 640x640 and every frame is letterboxed into it, so a bird
200px across in a 1600px proxy reaches the model at 80px. `Tiling::Grid` has
been implemented and tested since the segmentation spike and defaulted off,
because it costs one inference per tile — 2.8s for a 3x2 grid against 470ms.
Now offered as "Look closer (slower)", which says what it costs, rather than
spending it on every image or on none.

The tiling choice enters the segmentation signature. A mask stores the
signature its region ids index into, and a tiled run finds different instances
in a different order; sharing a signature would silently reinterpret a layer
built against the coarse pass — a wrong mask rather than a stale one, and
nothing announces it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-23 10:54:35 +02:00
dtourolleandClaude Opus 5 c0e1179936 Find the subjects without stopping the window
"Find subjects" took the UI thread for two thirds of a second on a 22 MP
frame — a proxy render, a readback and a YOLO pass through `ort` — and for
that time the interface was simply gone. The panel apologised for it rather
than hiding it: a "Looking…" label, and a 16 ms `single_shot` so the label
reached the screen before the freeze began, with a comment saying the obvious
fix needed the develop session restructured and was not being taken.

The obstacle was never `Send`. `DevelopSession` is `Send` — the device, the
source texture and the passes all are. What cannot go to a worker is the
`Rc<RefCell<Option<DevelopSession>>>` that every callback in the window
reaches through, and the window has to keep reaching through it while the work
runs. Handing the session over would freeze the interface exactly as
thoroughly as blocking on it did.

So the work takes a copy of what it needs instead. A `SegmentationJob` is the
device, the demosaiced source behind an `Arc`, and the name of the session
that asked. Taking one is two `Arc` bumps; running one is 495 ms on this
desktop; none of it touches the session, and there is deliberately no
`&mut DevelopSession` in scope for a caller to hold across it. The proxy
render travels with it rather than staying behind — a `GpuContext` and a
texture handle are both `Send`, and the model was never the only expensive
half. So does building the mask rasteriser, which is a shader compile:
adopting the result was costing 23 ms, a dropped frame on the one redraw the
user is waiting for, and the rasteriser is needed exactly when the subjects
arrive and never before. What is left on the UI thread is a microsecond.

The answer comes back through a channel a `slint::Timer` polls, which is the
shape `apply_when_ready` already uses for a sidecar fetch.

**A result can outlive the photograph it describes.** Two thirds of a second
is long enough to press the button, think better of it and swipe to the next
frame — and the result landing then would fill the panel with subjects that
are not in the picture, drawing outlines around a dog two photographs back.
Nothing downstream can tell: the masks rasterise and the overlay draws either
way. So every session is minted with an id, a job carries the id it was taken
from, and `delivery` compares the two before anything is applied. An id
rather than a counter beside the session slot, because that slot is written
from four places in `lib.rs` and the fifth would be the one that forgot.

A discard touches nothing on the way out. `segmenting` belongs to whichever
photograph is open now, which may well have a run of its own going, and
clearing it would re-enable a button that is correctly insensitive.

One run at a time, and abandonment is what stops that being a trap. A job
left over from a photograph the user has left is displaced rather than waited
for — otherwise the next frame's "Find subjects" would do nothing for the
length of a run nobody wants, which is the wait this exists to remove. `ort`
offers no way into the inference, so abandoning is checked at the seams there
are: before the job starts, and between the readback and the model. Abandoned
early it costs nothing, abandoned mid-inference it costs the run it was
already committed to, and either way the answer is dropped at the channel.

`DevelopSession::segment` survives as a test-only convenience. Left public it
is precisely the shape that put two thirds of a second on the UI thread in the
first place, and the next caller would reach for it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 12:26:11 +02:00
dtourolle 37f63edbf9 Take the watershed out of the product path
It does not work on a photograph, so nothing should offer it. `Segmentation`
is now one model pass and what it recognised: no region field, no merge
tree, no label upload, no granularity slider, and no readback of the whole
proxy to build a graph that collapses.

A click means "the object under the cursor". The region-selection path went
with the hierarchy it indexed — including the shift-click add/subtract,
which has no meaning for a whole object and would have been a modifier that
silently did nothing.

The passes, the hierarchy and the semantic prior stay in `dr-gpu` and
`dr-segment`, tested and documented. It is the *merge criterion* that
fails — the saddle is the minimum gradient along a boundary, so one weak
pixel merges two regions and real gradient noise puts a weak pixel on every
boundary. That is one function to replace, and the evidence for replacing it
is worth keeping. What is gone is the wiring, the option, and the control
that offered a user a choice with no outcome.

`MaskSource::Regions` remains in the pipeline: it is tested, it round-trips
through the sidecar, and a stored layer that names regions must still load
and be reported stale rather than failing to parse.
2026-08-22 08:39:17 +02:00
dtourolle ec713585a5 Measure the distance to the edge, and get four controls for one transform
Feathering, growing, shrinking, closing and opening are the same number
read differently. With the signed distance from the boundary in hand,
dilation is the set where d >= -r, erosion where d >= +r, and a feather of
any shape is a function of d. So the field is computed once and the
controls are arithmetic on it.

The **field** is what reaches the GPU, not a finished alpha, and that is
the point: growing a mask or changing its falloff then costs a uniform
upload and no recomputation, which is what makes them live controls rather
than ones that stall on every drag. Only closing and opening rebuild,
because after the first threshold the shape has changed and the old
distances describe the old one.

Exact Euclidean, via Felzenszwalb's separable transform — not a chamfer
approximation, which leaves a mask visibly octagonal once grown more than
a few pixels. A test asserts the diagonal is √2 rather than 1 or 2.

It runs on the CPU, which ARCH §5.4 forbids for masks. The rule is about
brush lag — a stroke rasterised per frame — and this is a different
operation: once per mask edit, on input the model already produced here,
producing a field the GPU then samples for free. What it buys is exact
determinism, which matters because masks reach the sidecar as indices and a
field that varied by vendor would mean a mask meaning one thing on the
desktop and another on the phone.

The half-pixel in `signed_distance` is not a detail, and a test caught it.
Measuring to the nearest opposite pixel *centre* puts the smallest
magnitude at 1 either side, so the boundary is nowhere and **eroding by
less than a pixel removes nothing**. A control whose first notch does
nothing is a broken control. Half a pixel off each side puts the boundary
where it physically is, and eroding by 1 takes exactly the outermost ring.

Every falloff curve is 0.5 at the boundary by construction, asserted for
all five: changing the curve should change how the transition looks and
never where it sits.
2026-08-22 08:39:17 +02:00
dtourolle ee10097435 Mask the subject the model found, not the regions underneath it
The watershed hierarchy does not survive a photograph, so local masking
stops depending on it. A layer can now be one recognised object, and the
object's own coverage is the mask.

`Options::watershed` defaults off. It costs ~80 ms plus a full-resolution
readback to produce a ladder that collapses, and paying that on every
photograph buys a control that misleads. Kept switchable rather than
deleted: the passes and the hierarchy are correct in themselves and it is
the merge criterion that fails, which is a change to one function.

Masks now rasterise in **source** space at proxy resolution and are sampled
by the composed shader after the framing map. That fixes a real bug: they
were rasterised in output space, so zooming slid the photograph underneath
a mask that stayed pinned to the viewport, and cropping moved every
adjustment to a different part of the picture. Doing it this way also
leaves the framing map in exactly one place — a second copy in the mask
shader would have been a second thing to keep in step, failing only when
straightened.

A subject is stored as identity, not pixels: the mask is megabytes and is
reproducible by running the same model over the same image, so the sidecar
carries the index, the class and the score, and the session carries the
pixels. The class is there to be checked — if instance 3 comes back a "car"
where it was a "dog", something changed and the layer is stale rather than
silently masking the wrong thing.

The overlay now draws instances and is transparent everywhere else. The
region version covered every pixel and so hid the photograph it was drawn
over; the question it exists to answer is whether an outline follows the
subject, which you can only answer by seeing both.

`examples/local.rs` is the worked example: subject in colour with the rest
monochrome, and the subject lifted out of its background. Run on a 5472x3648
CR2 it finds two people and two cars, and the colour-pop keeps her hat and
hair while the wall and grass behind go grey.
2026-08-22 08:39:17 +02:00
dtourolle b1433ad4a9 Give a mask an edge treatment, and find out the watershed has none worth having
Two things, and the second is why the first matters more than expected.

Mask layers gain a feather, a falloff curve and a morphology, all defined
against a signed distance from the boundary rather than as separate
features — one exact distance field answers "how soft" and "how far" at
once, so dilation is a threshold at -r, erosion one at +r, and closing and
opening are one of each in sequence. The compound pair costs a second
distance field, which is why they are named rather than presented as a
radius that happens to be signed. Types, defaults and sidecar round-trip
only; the field itself is next.

`edge-feather` and `edge-falloff`, not `feather` and `falloff`, because a
radial mask already writes `feather` for the fraction of its radius it
ramps over. Same word, different quantity, different units — sharing the
key would have made an existing file ambiguous.

The diagnostic that provoked this is committed as an ignored test, because
"does the ladder land on things a person means" is the question S15 exists
to answer and it should not depend on whoever still has the script. On
bus.jpg it answers badly: 35,075 regions at blur 2 over an 810x1080 frame,
and cutting that to 400 gives *one* region covering nearly the whole
picture plus 399 noise specks. Not over-segmentation — collapse. Almost
every saddle is near zero, so the merge order joins everything meaningful
before it joins anything spurious, and a global cut spends its entire
budget on grain.

So the granularity ladder does not currently work on a photograph, and the
region masks built on it inherit that. Recorded rather than worked around:
the next commits move local masking onto the model's instances, where the
edge treatment above is what makes a quarter-resolution mask usable.
2026-08-22 08:39:17 +02:00
dtourolle 12d320cf33 Record what the spec got wrong about the model that exists
docs/segmentation.md §4 priced arm B as costing a C dependency under the
NDK and treated that as most of the difference between the arms. It is not
a cost that has to be paid: `ort`'s `alternative-backend` disables its
linking entirely and `ort-tract` supplies the API from tract, which is pure
Rust. D13's "largest exception the policy would tolerate" turns out not to
be needed, and the answer generalises to the face pipeline — so D13's
runtime half is now answered and only its licensing half is open.

Three findings contradict §4 outright and are recorded as F4-F6 rather than
quietly designed around. There is no ADE20K-trained YOLO, so the shipped
vocabulary selects subjects and not stuff — "select the sky" comes from the
watershed or from nowhere. It is instance segmentation, so it partitions
nothing and two people come back as two instances. And tract cannot parse a
dynamic-shape export, which fixes the input at 640 square and makes tiling
the only route to more semantic resolution.

Arm C ships, but §8's criteria are not what decided it, and saying so
matters more than claiming the process worked. §8 asked for a two-
interaction margin over arm A on a traced corpus. That comparison was never
run: F4 and F5 changed what the arms are, and a model that recognises
subjects but has no word for sky cannot be a selection tool alone, while a
watershed cannot tell a person from the wall behind them. They stopped
being candidates and became complements.

What is *not* done is written down as plainly: the 24-image corpus is
untraced, so M1-M4 have no numbers and "this feels right" has not become
one. M5 is answered on one device only, and region ids now reach the
sidecar — so a cross-vendor divergence would mean a mask written on the
desktop meaning something else on Android. F3 stands.
2026-08-22 08:39:17 +02:00
dtourolle 5ecb35864f Put the region map behind the sliders that were already there
A mask layer holds a real develop chain, so the develop panel can edit one
with no new controls: select a layer and the same sliders read and write
its chain instead of the graph's. An operation declared in `ops/` tomorrow
becomes locally adjustable by existing, which is the payoff for making a
layer a chain rather than a handful of special-cased parameters.

`segmentation.rs` joins the two arms into the one thing the view needs.
The model reads the image through a neutral graph rather than the edited
one, so a segmentation survives an exposure change instead of being
invalidated by every slider. Arm B failing is not fatal: a missing or
unreadable model leaves a working watershed map, because refusing to
segment at all would trade a working feature for a strict one.

The overlay colours groups by a golden-angle walk over hue. Deterministic
rather than random, so a region keeps its colour across a level change and
the eye can track it; boundaries drawn black over the fill, because two
adjacent groups landing on near hues read as one region and telling them
apart is the whole reason to look at it.

Clicking the photograph creates the layer if none is selected — that is how
a local adjustment begins, and making the user press "add layer" first
would be a step with no decision in it. Shift-click extends, and clicking a
region already selected removes it, so one gesture both adds and corrects.

`segment-readback` is a new dr-gpu feature and not a loosening of
`readback`. The region-graph transfer is once per image on a worker; the
one AC-8 forbids is per frame in the render loop. Sharing a switch would
have forced a build wanting local masking to unlock the other. F3 still
stands and the feature name says so.
2026-08-22 08:39:17 +02:00