7a436e2549468e171f16a2b6a5130740743abd5c
21
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
696bafa9d5 |
Undefer AI subject masking, which shipped, and give it a clause
§7 still listed "AI subject masking — deferred per D11" while MaskSource::Subject and MaskSource::Category, backed by dr-segment's instance and semantic models, had been the primary way a local adjustment is made for weeks. The code was tagged FR-DEV-3, which names gradients and brushes and says nothing about a model. FR-DEV-3i now states what exists: a subject or a category found by a local model, stored as identity with the run's signature so that it merges per field and reads as stale rather than wrong, then treated as any other layer by the edge, stroke, composition and reveal clauses. The one place it departs from FR-DEV-19 — coverage written run-length coded beside the layer, so a stored subject renders without a model — is recorded in the clause instead of left for the next audit to find. The segmentation crate and the UI's selection module are tagged to it. |
||
|
|
193b35a249 |
Start a category mask where the photograph can bear it
Clicking "architecture" made a layer whose mask was gone. Every category layer began at STRICTNESS_DEFAULT, and that constant was fitted on the synthetic sky the refine tests build — its own note warns that a real photograph's noise "moves every crossing down together", which turns out to be a considerable understatement. Measured over seven ordinary frames, half scale removes 76% to 99.5% of `architecture`, 36% to 93% of `ground` and 18% to 91% of `vegetation`. Only sky, the category the number was calibrated against, survives it. An empty mask is indistinguishable from a broken one: the layer is listed, the adjustment moves, and no pixel changes. So what this looks like from outside is that the segmentation does not make masks at all. No smaller constant fixes it either, because a nat of evidence means different things over a smooth sky and over a stone facade — the useful position is above 5 on one frame and below 1 on the next. So the frame is asked instead: `Refinement::gentle` walks down from half scale and takes the first rung whose gate removes no more than a sixth of the category's weight, and the model's own outline when none of them does. One `apply` on a friendly photograph and four on an unfriendly one, paid when a layer is made rather than for eight categories nobody masked. The slider's reset went to 4 as well, so taking the control back to its "default" emptied the mask. It goes to zero now, which is the one position documented to mean something: exactly what the model weighted. |
||
|
|
89c4ff1820 |
Store what the model found, so a reopened photograph keeps its masks
A subject or category layer was written to the sidecar as identity alone —
which run, which instance, which category — on the reasoning that the pixels
are reproducible by running the same model over the same image. They are, but
only by *running the model*, and nothing runs one except a photographer
pressing "find subjects". So on every path that did not already have a run in
memory the layer resolved to no coverage, `MaskPass::render` logged "has no
distance field; skipping", and the adjustment was silently absent:
- reopening an edited photograph rendered it without its local adjustments,
and then saved that state back on the way out;
- a batch export from the grid could not have them at any point, because
`render_from_library` opens a session, applies a version and renders, and
there is no model anywhere on that path. Three hundred files written
without the edits their photographer made, over a log warning.
Neither failure announced itself. The generated shader still emits the layer's
block and the empty placeholder multiplies it by zero, so the result is a
well-formed frame that is simply missing an edit — `mask_is_stale` already
named the state and called it "not stale, just unrenderable".
The coverage now travels in the file, as one `coverage = w h levels payload`
line at the end of the layer's block.
Two levels, and that is not a compromise. The model hands out a byte per pixel
but `Shaped::build` measures its distance field from `coverage >= 128` and
throws the shoulder away on the first line; everything soft about the rendered
edge comes afterwards from the layer's feather and falloff, which are read off
the distance. So one bit per pixel is not an approximation of what the model
said — it is exactly the part of it that reaches a pixel, and the stored mask
renders the identical frame. Storing all 256 levels would have stored 1.7 MB
of bilinear interpolation to reconstruct a predicate, and would not even have
compressed: a model mask is a bilinear upsample of a coarse grid, so almost no
two adjacent bytes are alike. Measured on a simulated sky and a simulated
figure at 1600x1067, against 1.71 MB raw: 4.0 kB and 6.5 kB at two levels,
46 kB and 76 kB at sixteen, 835 kB and 1.43 MB at all 256. The level count is
still written into the line, so a later build that finds a use for the
shoulder can write sixteen and this one will read them rather than misreading
a stream of lengths as pairs.
The coder is hand-rolled — run-length pairs in a base-64 varint — because
`dr-pipeline` links nothing, which is the property that lets the descriptor
and codegen logic be tested without a device. `flate2` would have been fewer
lines and a dependency in the one crate that has none.
Where it lives matters more than how it is coded. The raster sits on
`MaskLayer` beside the source, not inside `MaskSource::Subject`: the source is
*identity*, which is what makes it diff as a handful of numbers and merge per
field under FR-NC-9, and a raster in there would have given the merge a binary
blob to arbitrate. It takes no part in `MaskLayer`'s equality for the same
reason — a device that has run the model and one that has not hold the same
edit, and counting the difference would raise a conflict over a cache and let
`remote_wins` answer it by discarding the only copy of the pixels.
Encoding happens in `masks_for_storage`, on the save path, rather than in
`ensure_subject_fields` where every coverage already funnels through.
`ensure_subject_fields` runs on a drag — dilating a mask with a compound
morphology rebuilds the field every frame — and encoding a megapixel raster
per frame is the kind of work NFR-P5 exists to keep off a gesture. Saving
happens once, when the photograph stops being the open one, and already costs
a network round trip.
Version skew holds both ways. A file with no `coverage` line reads exactly as
it did before, which is a layer that needs the model run; an unreadable one
costs the pixels and not the layer, because the layer is the edit and the
raster is a cache of it. An old build reading a new file drops the key it does
not understand, which costs a model run and no work. And a payload that will
not compress is refused rather than truncated: a checkerboard would encode to
twice the raster it came from, so past 64 kB nothing is stored and the
behaviour falls back to what it was — half a mask would render as a mask that
is confidently wrong, which is the failure that tells nobody.
|
||
|
|
1c5849ebe8 |
Unpack the bundled models on a worker, not on the way to the first frame
Launching v0.10.0 on the tablet produces an ANR: "Waited 5000ms for MotionEvent", 5,827 ms on one input sequence, over a window that has never painted. The app recovers and sits at 0% afterwards, so it is a startup cost rather than a hang. The structural fact behind it is that `android_main` runs with the activity's input channel unserviced. Nothing drains it until Slint reaches `poll_events`, and Slint does not reach `poll_events` until `dr_ui::run` calls `window.run()` on its last line. Every millisecond before that is a millisecond the input dispatcher waits on, so five thousand of them is an ANR whatever the work happens to be. The largest single piece of that work was here. `install_bundled_models` copies 41 MB on the first launch after an install — 24.9 MB of scene model, 13.6 MB of embedder, 2.5 MB of detector — each read whole out of the APK into a `Vec` and written to `/data`, in a loop, on that thread. v0.10.0 is the release that added the scene model, which is 60% of that total, and it is the release the ANR appeared in. The 8,010 minor faults in the report are about what 41 MB of freshly touched pages costs. So it moves to a detached thread and the function returns as soon as the thread is running. Nothing on the launch path wanted the result: the only two things that read these files are the People screen and the scene tab, both of which are reached by hand, minutes later, from workers of their own. What that costs is a window in which a model looks absent. `library::face_models` and `library::scene_model` decide availability on `is_file()`, so during the copy both report their feature unavailable — which is the same answer they give a build shipping no weights at all, the ordinary case both were written around. Briefly pessimistic rather than wrong, and the temporary-name-then-rename that was already there is what keeps it from being worse than that: a lookup never sees a half-written file, only an absent one. Both call sites now say so. A completion line reports the bytes copied and the milliseconds taken, including when it is zero, so the second launch after an install can be told from the first in a log rather than by inference. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4efab496c2 |
Put the refinement on a slider, per layer
Benchmarks / CPU and I/O (per commit) (push) Successful in 3m27s
Benchmarks / Frame budget (on demand) (push) Skipped
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
Build and test / Desktop (Linux) (push) Failing after 1h14m30s
Build and test / Layer separation (push) Successful in 43s
Traceability / Requirement traces (push) Failing after 46s
Build and test / Android (aarch64) (push) Successful in 1h8m56s
The evidence was being gathered and spent immediately at one strictness nobody could see or change. This makes it a control. `MaskLayer::refine` is shaping, like the feather, and it lives on the layer for the layer's reason: two layers may sit on the same category and want different amounts of it, and the model ran once for both. ## What the segmentation now stores `CategorySummary` keeps the **coarse** mask and the `Refinement` beside it, rather than a refined mask. That is what gives the control an off position that is bit-for-bit the model's own weighting, and what stops a strictness change from needing the model again. `category_mask_at` borrows at zero and wherever no refinement could be fitted, so a layer nobody has touched costs nothing over the old path. The fit moved to the far side of the orientation permutation. The verdict is a per-pixel field over the same grid as the mask it gates, so fitting it upright would mean permuting a proxy-sized buffer afterwards to match — a second rotation, and a second chance to get one wrong. One logit cell is the same number of pixels either way: the letterbox scales by the longer edge and a permutation does not change which edge that is. The two frames became a `Frames` struct rather than six parameters. This function reads the picture twice for opposite purposes — the model needs it upright or it recognises far less, the refinement needs the sensor's grid — and a transposed pair produces a plausible mask over slightly the wrong pixels, which is the failure this module is most prone to. ## It rebuilds the field, and the signature says so `refine` is mixed into `subject_signature`. Unlike a feather, which is read off a field that is already correct, this changes which pixels are in the mask at all — so it changes the coverage the field is measured from. Omitting it is the bug where the slider moves and nothing happens until some unrelated control invalidates the cache. That puts it in the same cost class as a close or an open, which is why the row takes `SliderRow::changed` — already once-per-gesture, since that row takes `SliderTrack`'s `committed` internally — rather than a live stream. The slider is offered only where there is something to move: a category source, *and* a refinement the frame actually gave enough to fit. A control that moves and does nothing is worse than an absent one. ## Two defaults that are deliberately different A layer added from the panel starts at 4.0, because a category's edges are twenty proxy pixels wide before anything is done to them and a photographer adding a sky mask wants the sky rather than the sky plus every chimney in it. A layer read from a sidecar with no `refine` key starts at **zero**. A file written before this control existed has to render as it did then, and a default of 4 on absence would quietly re-grade every stored category mask in the catalogue. `a_categorys_refine_strictness_survives_and_defaults_off` holds both halves, and `an_out_of_range_refine_is_clamped` holds the file to the scale — past the top of it every colour fails and the mask deletes itself, which reads as lost work rather than as a bad file. `MAX_REFINE` is dr-pipeline's own constant mirroring `dr_segment::STRICTNESS_MAX`, following `Falloff` and `Morphology`: this crate holds the description of an edit and must not depend on the crate that runs a model. dr-ui is where the two meet, and the only place that converts. Verified: fmt clean, clippy --workspace -D warnings clean, 488 dr-pipeline and 60 dr-segment tests. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4f4abd335f |
Cut a scene category back to the pixels that agree with it
A flag in the sky came out weighted as sky, and no feather setting fixed it. The scene model's logits are `[1, 150, 80, 80]`, so one cell is eight input pixels; at the 1600px proxy the letterbox scale is 0.4 and **one cell is 20 proxy pixels**, which `rasterise`'s bilinear then spreads across one more either side. A flag is a handful of cells whose softmax is dominated by the sky around it. The information was never in the grid, so nothing downstream of the grid can recover it. Tiling is the answer for an instance and is not available here: a category has no bounding box to tile over — sky is wherever the sky is. But the photograph is at full proxy resolution even though the weights are not, and it knows exactly where the flag is. So the model says *what*, and the pixels say *which of them*, which is the division of labour arm C already draws between the instance model and the watershed. ## Seeds, and why the erosion radius is not a guess Threshold the weights high, take `signed_distance`, and keep what is more than 1.5 cells inside. One cell *is* the model's resolution and the bilinear spreads it across one more, so the band either side of the boundary is smear rather than evidence. Deriving the radius from `Scene::cell_pixels` rather than picking a pixel count means it stays right if the proxy edge or the export changes. The mirror of that set is a confident *exterior*, free from the same field. ## Dropping small modes is the step that makes it work Four k-means modes per side, not one Gaussian: sky is blue at the zenith, white where the cloud is and pale at the horizon, and one blob over all three rejects two of them. Then modes holding under 3% of a side are discarded, and without that step the whole thing fails on the case it was built for. A small flag deep in the sky has both a high weight and a large distance from the boundary, so it lands in the interior sample and teaches the model its own colour. It cannot be excluded geometrically. It can be excluded by share. Luminance is weighted at a quarter against chrominance for the same reason the watershed's gradient is. Sky's variance is dominated by luminance, so at equal weight the distribution is a long bright streak that a mid-grey flag sits comfortably inside. A flag is separated by chrominance; a cloud is separated by luminance alone. Not zero, or a dark bird against a bright sky survives. ## Two tests, because either alone is wrong Absolute — is this colour plausible under the category, as a chi-square on the Mahalanobis distance. Comparative — is it likelier inside than outside. A pixel must pass both. The absolute test is what catches the flag, whose colour is far from *both* sides and which the comparative test alone would leave at even odds. The comparative test is what stops the absolute one needing a constant tuned per category. ## What this cannot do, written down rather than left to be discovered An intruder large enough to hold its own mode is kept. By share, a flag over a fifth of the sky and a cloud bank over a fifth of the sky are the same object, and colour does not separate them either — a white cloud is as far from blue sky in chrominance as many intruders are. So `min_cluster` is not a threshold with a correct value waiting to be found; it is the trade-off itself, set where a photographic intruder falls. Both ends are pinned by tests — `a_flag_in_the_sky_is_removed` and `an_intruder_larger_than_min_cluster_survives` — so that moving the number reads as moving the trade-off rather than as fixing a bug. The case left open is a large unrecognised object in a clean category, which wants the boundary snapped to watershed basins and is a different mechanism. ## Safe to apply without a control It is subtractive: the output is the input times a factor in `0..=1`. The worst failure available to it is losing part of a real sky, never gaining a region, so a blue car below the horizon that was never in the mask cannot be pulled into it. And a factor in `0..=1` cannot raise a sum, so `scene.rs`'s partition still holds when every category is refined independently — the weight taken off the flag lands in the unlisted remainder, which is where a flag belongs, ADE20K having no class for one. Every path without the evidence to judge returns the weights untouched and says which path it took. A refinement that silently did nothing is indistinguishable from the feature being off, and an empty seed set fitted to a distribution would reject every pixel. The signature is deliberately unchanged: categories are addressed by name, not by index, so a sharper mask cannot create the stale-index hazard the signature exists to guard against. The example writes `<prefix>-<category>-refined.ppm` beside the coarse one, never instead of it — whether this is an improvement is a comparative judgement and one image cannot answer it. Verified: fmt clean, clippy -D warnings clean, 57 dr-segment tests. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
94b1b9e0cc |
Reconcile what four branches each built separately
Three seams, found by the first compile after the merge. Two branches implemented 'can this scope be reordered' independently: one from the collections model, one from the catalog through orders_manually, on every re-read and excluding the trash. The second is the better answer and is what survives; it only needed to set the property app.slint declares. Two lints from scene-mask-ui, which was merged mid-flight and had never been through -D warnings: an is_none check spelled out where clippy wants ?, and a return in a cfg block's tail. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
763dfd353a |
Weigh the categories in the same precompute, and mask with them
The scene model shipped with a decoder and no caller. This runs it. ## Beside the instance pass, not instead of it `compute` now does both on the same upright frame and lays both back down the same way, so instance masks and category masks index into one grid — the sensor's. A failure in the scene half is logged and dropped rather than propagated: no scene model is an ordinary state, and a photograph that can still be masked by subject should not become unopenable because the categories are missing. Categories under half a percent of the frame never reach the cache. A control that does nothing when moved is worse than an absent one, and each one it skips is a proxy-sized buffer not allocated. ## The shader needed nothing A category reaches `dr-gpu` as a soft coverage buffer at proxy resolution, turned into a distance field — which is exactly what a subject is. So they share `MODE_SUBJECT`. That is not a shortcut taken for speed: the shader has no way to tell them apart and no reason to want one. What differs is only which model produced the coverage, and that has already happened by then. Feather, falloff, dilation and erosion therefore work on a category on the day it arrives, because they were never subject-specific. ## Where the weights come from `scene-model` compiles the graph in and the desktop app takes it; Android leaves it off and reads the copy `install_bundled_models` unpacks, because 24 MB of constant is worth avoiding in a mobile install and not worth the plumbing to avoid on a desktop one. Embedded is tried first — a build that has the weights compiled in should not be silently overridden by a stale file in a data directory. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d777f7f44d |
Merge branch 'master' into worktree-faces-scrfd-mbf
Build and test / Desktop (Linux) (push) Failing after 25s
Build and test / Layer separation (push) Successful in 22s
Traceability / Requirement traces (push) Successful in 58s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Failing after 33m26s
# Conflicts: # docs/traceability.md # ui/dr-ui/src/develop.rs # ui/dr-ui/src/segmentation.rs |
||
|
|
ce201c7dd6 |
Name the two spaces a photograph lives in, so a turn cannot go the wrong way
Every orientation bug this codebase has had has been the same bug: a turn of the right size applied in the wrong direction. That failure is worth naming precisely, because it does not look like one — a quarter turn applied backwards lands 180 degrees from right, so the result is a plausible transform of the picture rather than anything obviously broken, and on landscape frames it is not wrong at all. It was the straighten shear, and it was the segmentation overlay, and each time it was found by eye rather than by a test. The reason it keeps happening is that "rotate 90 degrees clockwise" cannot be checked by reading it. The reader has to hold in their head which of the two images is being rotated and which way the y axis runs, and there were four hand-written copies of the permutation to hold it for: the shader prologue, its CPU twin, the thumbnail path, and the segmentation. So nothing added here says clockwise, anticlockwise, horizontal or vertical. The functions say *which space they take and which space they return* — `into_shown` and `into_stored`, `source_pixel` and `shown_pixel`, `into_shown_rect` and `into_stored_rect` — and each takes the dimensions of the space it reads from, so no caller has to work out which pair it is holding. `StoredRect` and `ShownRect` are separate types because they are the same four numbers meaning different things, which is exactly the case where a mistake is silent: a shown rect measured against stored dimensions produces a rectangle in the wrong place, not an error. Underneath there is one permutation. `source_pixel` was already shared by the prologue and the thumbnails; `source_point` is its normalised twin, written beside it so the two cannot drift, and everything else is those two read forwards or backwards. `Orientation::inverse` is the group inverse rather than `4 - turns`: mirrors apply after the turn, so undoing means undoing them first, and a mirror seen from the far side of an odd turn is about the other axis. That is the diagonal-mirror case, tags 5 and 7, and getting it wrong renders as — again — 180 degrees. Three call sites lose their own copy: the thumbnail path, `dr-ui`'s segmentation, and `dr-gpu`'s `local` example. "Upright" now means one thing across the application rather than one thing per caller. The gate that matters most is `the_render_and_the_orientation_map_agree`. The shader prologue and `Orientation` answer the same question by different routes, and until now nothing checked that they answered it the same way. It now checks every EXIF tag against every user rotation and mirror on top of it, because the composition is where the two could agree singly and disagree together. The rest earn their place by having caught something. Writing these found two real errors in this commit's own new code before it ran anywhere: `shown_pixel` was handed the dimensions of the wrong space and overflowed, and the rect map turned the wrong way for the diagonal mirrors. A round trip that returns what went in is the only check worth having here, since every wrong answer is still a picture. No behaviour changes. The permutations are the ones that were already being applied; they are simply applied from one place now. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b55812812a |
Name segmented people from the faces already recognised in them
The segmenter knows it found a person; the face index knows which person. Joining them turns "person" in the mask list into "Anna", which is the difference between a vocabulary of eighty COCO classes and one that includes the user's family. Selecting a subject in a group photograph stops being a guessing game between three identical rows. Containment, not IoU. A face is a small part of the person it belongs to, so a correct pairing has an IoU near zero and anything IoU-based would reject every true match. Confirmed names only. A suggestion is the system's guess, and printing a guessed name onto a mask region would launder it into a fact. Writing the tests corrected the design once: a tight head-and-shoulders portrait, where the face fills most of the person box, is the case where naming is most certain, not least. An earlier guard rejected exactly that and has been removed, with the reasoning left as a test because it is easy to get backwards a second time. The names hang on the develop session, set when the image opens because that is the one moment the catalog and the image id are both in reach. Every segmentation run afterwards picks them up for free, and a library with no face indexing behaves exactly as it did before. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4a82753d22 |
Show the detector the photograph, not the sensor's scanlines
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
Build and test / Desktop (Linux) (push) Successful in 19m7s
Build and test / Layer separation (push) Successful in 25s
Traceability / Requirement traces (push) Successful in 23s
Build and test / Android (aarch64) (push) Failing after 33m10s
"Find subjects" was handed the proxy in the sensor's own orientation, so every frame shot on a body held sideways reached the model lying on its side — and a model trained on upright photographs is very bad at those. Measured end to end on a 22 MP frame of two people and a dog: `person 0.36` and nothing else, against `dog 0.82, person 0.61, person 0.49` for the same pixels stood up. Nothing failed; the panel simply offered one poor subject where there were three good ones. The orientation was never dropped on purpose. The proxy is deliberately rendered through a *neutral* graph — the detection has to survive an exposure change, or every slider would invalidate the masks built on it — and neutral took the file's orientation with it along with everything else. Landscape frames were unaffected, which is why it stood for as long as it did. The turn is `Orientation::source_pixel`, the same function the grid's thumbnails already go through, so the detector and the thumbnailer now agree about which way is up rather than holding two opinions. What it is turned by is `Framing::effective_orientation` — the file's EXIF tag and the photographer's own rotations composed into one permutation, by the group law rather than by adding the turns, which is a distinction `Framing` already had to make and had already tested. Rotating the picture and pressing the button again therefore does what it looks like it does. The proxy stays in sensor space and the masks come back into it. That is not a detail to be tidied later: the generated shader samples the mask array at `uv_src`, *after* the framing map, so a mask stored upright would sit a quarter turn off the subject it was drawn around. That is a wrong mask rather than a weak one, and nothing announces it. So the picture is stood up for the model and laid back down for everything else, and `upright`/`lay_down` are returned as a pair because calling one and forgetting the other is silent. Both directions are the one function: `upright` gathers through `source_pixel` and `lay_down` scatters through it. A quarter turn is a bijection of the pixel grid, so the round trip is exact — no filter, no resampling, and no hole to fill — and an inverse written out by hand would be a second thing to keep in step, whose way of being wrong is a mask mirrored about the wrong axis, which still looks like a mask. The orientation joins the confidence and the tiling flag in the segmentation signature, and for the same reason: turning the photograph changes what the model recognises, so two runs either side of a rotation are different instance lists. Two that happened to come out the same length would otherwise share a signature and a stored layer would be silently re-indexed from one into the other. The refine pass had it too — it re-runs the model over a crop rendered in the same sensor space — so it makes the same turn, and would otherwise have handed back a worse mask than the one it was asked to improve, on the subject the photographer had just pointed at. `dr-gpu`'s `local` example is fixed with it. It exists to be the shipping path with pictures attached, and a diagnostic that reproduces the bug it is meant to catch is a trap for whoever reads it next. Seven tests. The round trip is the identity over all eight EXIF tags on a non-square asymmetric grid; a turn carries whole pixels rather than shearing the channels apart; a sideways frame reaches the model upright; a box comes back in sensor pixels, worked out by hand for the one turn a portrait frame actually writes; a restored box still reads low-to-high for every tag, since the rest of the pipeline takes `x1 - x0` without checking the sign; and the eight tags cannot collapse into one signature key. The existing composition test now runs against `effective_orientation` itself, over all 8 x 16 baseline-and-user pairs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b3641b5307 |
Let a subject's mask be refined, and let several masks be edited at once
Two related gaps in the mask panel, from the same conversation: a subject's outline is only ever as sharp as the whole-frame pass that found it, and a change meant for several layers had to be dragged once per layer. ## Refine mask A "Refine mask" button on a subject layer re-runs detection on a padded crop around that instance's own box instead of the whole frame — the subject reaches the model at its own size rather than squeezed into the model's fixed 640x640 window alongside everything else in the photograph. `RefineJob` mirrors `SegmentationJob`'s split (built on the session, run off it, adopted back), and the crop itself is rendered through `Framing::set_view` — the same ephemeral viewport the interactive zoom already uses to render a region above proxy resolution, so no new render path and no change to the model's own input size was needed. `dr-segment` is untouched: `Tiling::Whole` already treats whatever buffer it is handed as the one window. The result is still downsampled onto the shared proxy grid every instance's mask lives on, but from a sharper source than the whole-frame pass ever saw for that subject, which is what the edge actually reads out of. ## Multi-select `active_mask: Option<String>` is now `active_masks: Vec<String>`. A plain click still replaces the selection; a control- or command-click toggles one layer in or out of it. `set_param` and `reset_op` fan out to every selected layer, each set to the exact value the slider now shows rather than offset by however far it already was — one slider, one reading, applied everywhere selected. Dragging a gradient's on-canvas handle is deliberately not extended to multi-select: several gradients have no single geometry a shared handle could move, so `gradient_handles` stays empty unless exactly one layer is selected. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
7fcb8dc107 |
Give the develop column room, and the mask a way out of the way
Three faults reported together, and they share a shape: each is something the panel decided on the photographer's behalf. **The column was 280px on every screen.** That width was chosen for a tablet, where the column is a large fraction of the display and every pixel of it is taken from the photograph. On a desktop window the mode strip alone — two modes, a separator, "All", and a chip per attribute the operation set declares — does not fit, so it scrolled sideways. A control you have to pan to reach is one you do not know is there. 380px when the window is classed expanded, 280px when it is not; driven off `layout-class` because reading the window width inside the layout that sets it is a binding loop. **The overlay could not be hidden.** An earlier "Overlay" button was removed for a good reason — it *armed* the overlay, so local mode could be entered and still show nothing. Hiding is the opposite need and was never served: a mask is judged against the photograph beneath it, and that photograph is exactly what the overlay covers. `overlay-hidden` is kept separate from `overlay-on` so a recompute cannot switch the overlay back on under someone who just turned it off. **The masks were coarse because the model saw the subject small.** The graph's input is a fixed 640x640 and every frame is letterboxed into it, so a bird 200px across in a 1600px proxy reaches the model at 80px. `Tiling::Grid` has been implemented and tested since the segmentation spike and defaulted off, because it costs one inference per tile — 2.8s for a 3x2 grid against 470ms. Now offered as "Look closer (slower)", which says what it costs, rather than spending it on every image or on none. The tiling choice enters the segmentation signature. A mask stores the signature its region ids index into, and a tiled run finds different instances in a different order; sharing a signature would silently reinterpret a layer built against the coarse pass — a wrong mask rather than a stale one, and nothing announces it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c0e1179936 |
Find the subjects without stopping the window
"Find subjects" took the UI thread for two thirds of a second on a 22 MP frame — a proxy render, a readback and a YOLO pass through `ort` — and for that time the interface was simply gone. The panel apologised for it rather than hiding it: a "Looking…" label, and a 16 ms `single_shot` so the label reached the screen before the freeze began, with a comment saying the obvious fix needed the develop session restructured and was not being taken. The obstacle was never `Send`. `DevelopSession` is `Send` — the device, the source texture and the passes all are. What cannot go to a worker is the `Rc<RefCell<Option<DevelopSession>>>` that every callback in the window reaches through, and the window has to keep reaching through it while the work runs. Handing the session over would freeze the interface exactly as thoroughly as blocking on it did. So the work takes a copy of what it needs instead. A `SegmentationJob` is the device, the demosaiced source behind an `Arc`, and the name of the session that asked. Taking one is two `Arc` bumps; running one is 495 ms on this desktop; none of it touches the session, and there is deliberately no `&mut DevelopSession` in scope for a caller to hold across it. The proxy render travels with it rather than staying behind — a `GpuContext` and a texture handle are both `Send`, and the model was never the only expensive half. So does building the mask rasteriser, which is a shader compile: adopting the result was costing 23 ms, a dropped frame on the one redraw the user is waiting for, and the rasteriser is needed exactly when the subjects arrive and never before. What is left on the UI thread is a microsecond. The answer comes back through a channel a `slint::Timer` polls, which is the shape `apply_when_ready` already uses for a sidecar fetch. **A result can outlive the photograph it describes.** Two thirds of a second is long enough to press the button, think better of it and swipe to the next frame — and the result landing then would fill the panel with subjects that are not in the picture, drawing outlines around a dog two photographs back. Nothing downstream can tell: the masks rasterise and the overlay draws either way. So every session is minted with an id, a job carries the id it was taken from, and `delivery` compares the two before anything is applied. An id rather than a counter beside the session slot, because that slot is written from four places in `lib.rs` and the fifth would be the one that forgot. A discard touches nothing on the way out. `segmenting` belongs to whichever photograph is open now, which may well have a run of its own going, and clearing it would re-enable a button that is correctly insensitive. One run at a time, and abandonment is what stops that being a trap. A job left over from a photograph the user has left is displaced rather than waited for — otherwise the next frame's "Find subjects" would do nothing for the length of a run nobody wants, which is the wait this exists to remove. `ort` offers no way into the inference, so abandoning is checked at the seams there are: before the job starts, and between the readback and the model. Abandoned early it costs nothing, abandoned mid-inference it costs the run it was already committed to, and either way the answer is dropped at the channel. `DevelopSession::segment` survives as a test-only convenience. Left public it is precisely the shape that put two thirds of a second on the UI thread in the first place, and the next caller would reach for it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
37f63edbf9 |
Take the watershed out of the product path
It does not work on a photograph, so nothing should offer it. `Segmentation` is now one model pass and what it recognised: no region field, no merge tree, no label upload, no granularity slider, and no readback of the whole proxy to build a graph that collapses. A click means "the object under the cursor". The region-selection path went with the hierarchy it indexed — including the shift-click add/subtract, which has no meaning for a whole object and would have been a modifier that silently did nothing. The passes, the hierarchy and the semantic prior stay in `dr-gpu` and `dr-segment`, tested and documented. It is the *merge criterion* that fails — the saddle is the minimum gradient along a boundary, so one weak pixel merges two regions and real gradient noise puts a weak pixel on every boundary. That is one function to replace, and the evidence for replacing it is worth keeping. What is gone is the wiring, the option, and the control that offered a user a choice with no outcome. `MaskSource::Regions` remains in the pipeline: it is tested, it round-trips through the sidecar, and a stored layer that names regions must still load and be reported stale rather than failing to parse. |
||
|
|
ec713585a5 |
Measure the distance to the edge, and get four controls for one transform
Feathering, growing, shrinking, closing and opening are the same number read differently. With the signed distance from the boundary in hand, dilation is the set where d >= -r, erosion where d >= +r, and a feather of any shape is a function of d. So the field is computed once and the controls are arithmetic on it. The **field** is what reaches the GPU, not a finished alpha, and that is the point: growing a mask or changing its falloff then costs a uniform upload and no recomputation, which is what makes them live controls rather than ones that stall on every drag. Only closing and opening rebuild, because after the first threshold the shape has changed and the old distances describe the old one. Exact Euclidean, via Felzenszwalb's separable transform — not a chamfer approximation, which leaves a mask visibly octagonal once grown more than a few pixels. A test asserts the diagonal is √2 rather than 1 or 2. It runs on the CPU, which ARCH §5.4 forbids for masks. The rule is about brush lag — a stroke rasterised per frame — and this is a different operation: once per mask edit, on input the model already produced here, producing a field the GPU then samples for free. What it buys is exact determinism, which matters because masks reach the sidecar as indices and a field that varied by vendor would mean a mask meaning one thing on the desktop and another on the phone. The half-pixel in `signed_distance` is not a detail, and a test caught it. Measuring to the nearest opposite pixel *centre* puts the smallest magnitude at 1 either side, so the boundary is nowhere and **eroding by less than a pixel removes nothing**. A control whose first notch does nothing is a broken control. Half a pixel off each side puts the boundary where it physically is, and eroding by 1 takes exactly the outermost ring. Every falloff curve is 0.5 at the boundary by construction, asserted for all five: changing the curve should change how the transition looks and never where it sits. |
||
|
|
ee10097435 |
Mask the subject the model found, not the regions underneath it
The watershed hierarchy does not survive a photograph, so local masking stops depending on it. A layer can now be one recognised object, and the object's own coverage is the mask. `Options::watershed` defaults off. It costs ~80 ms plus a full-resolution readback to produce a ladder that collapses, and paying that on every photograph buys a control that misleads. Kept switchable rather than deleted: the passes and the hierarchy are correct in themselves and it is the merge criterion that fails, which is a change to one function. Masks now rasterise in **source** space at proxy resolution and are sampled by the composed shader after the framing map. That fixes a real bug: they were rasterised in output space, so zooming slid the photograph underneath a mask that stayed pinned to the viewport, and cropping moved every adjustment to a different part of the picture. Doing it this way also leaves the framing map in exactly one place — a second copy in the mask shader would have been a second thing to keep in step, failing only when straightened. A subject is stored as identity, not pixels: the mask is megabytes and is reproducible by running the same model over the same image, so the sidecar carries the index, the class and the score, and the session carries the pixels. The class is there to be checked — if instance 3 comes back a "car" where it was a "dog", something changed and the layer is stale rather than silently masking the wrong thing. The overlay now draws instances and is transparent everywhere else. The region version covered every pixel and so hid the photograph it was drawn over; the question it exists to answer is whether an outline follows the subject, which you can only answer by seeing both. `examples/local.rs` is the worked example: subject in colour with the rest monochrome, and the subject lifted out of its background. Run on a 5472x3648 CR2 it finds two people and two cars, and the colour-pop keeps her hat and hair while the wall and grass behind go grey. |
||
|
|
b1433ad4a9 |
Give a mask an edge treatment, and find out the watershed has none worth having
Two things, and the second is why the first matters more than expected. Mask layers gain a feather, a falloff curve and a morphology, all defined against a signed distance from the boundary rather than as separate features — one exact distance field answers "how soft" and "how far" at once, so dilation is a threshold at -r, erosion one at +r, and closing and opening are one of each in sequence. The compound pair costs a second distance field, which is why they are named rather than presented as a radius that happens to be signed. Types, defaults and sidecar round-trip only; the field itself is next. `edge-feather` and `edge-falloff`, not `feather` and `falloff`, because a radial mask already writes `feather` for the fraction of its radius it ramps over. Same word, different quantity, different units — sharing the key would have made an existing file ambiguous. The diagnostic that provoked this is committed as an ignored test, because "does the ladder land on things a person means" is the question S15 exists to answer and it should not depend on whoever still has the script. On bus.jpg it answers badly: 35,075 regions at blur 2 over an 810x1080 frame, and cutting that to 400 gives *one* region covering nearly the whole picture plus 399 noise specks. Not over-segmentation — collapse. Almost every saddle is near zero, so the merge order joins everything meaningful before it joins anything spurious, and a global cut spends its entire budget on grain. So the granularity ladder does not currently work on a photograph, and the region masks built on it inherit that. Recorded rather than worked around: the next commits move local masking onto the model's instances, where the edge treatment above is what makes a quarter-resolution mask usable. |
||
|
|
12d320cf33 |
Record what the spec got wrong about the model that exists
docs/segmentation.md §4 priced arm B as costing a C dependency under the NDK and treated that as most of the difference between the arms. It is not a cost that has to be paid: `ort`'s `alternative-backend` disables its linking entirely and `ort-tract` supplies the API from tract, which is pure Rust. D13's "largest exception the policy would tolerate" turns out not to be needed, and the answer generalises to the face pipeline — so D13's runtime half is now answered and only its licensing half is open. Three findings contradict §4 outright and are recorded as F4-F6 rather than quietly designed around. There is no ADE20K-trained YOLO, so the shipped vocabulary selects subjects and not stuff — "select the sky" comes from the watershed or from nowhere. It is instance segmentation, so it partitions nothing and two people come back as two instances. And tract cannot parse a dynamic-shape export, which fixes the input at 640 square and makes tiling the only route to more semantic resolution. Arm C ships, but §8's criteria are not what decided it, and saying so matters more than claiming the process worked. §8 asked for a two- interaction margin over arm A on a traced corpus. That comparison was never run: F4 and F5 changed what the arms are, and a model that recognises subjects but has no word for sky cannot be a selection tool alone, while a watershed cannot tell a person from the wall behind them. They stopped being candidates and became complements. What is *not* done is written down as plainly: the 24-image corpus is untraced, so M1-M4 have no numbers and "this feels right" has not become one. M5 is answered on one device only, and region ids now reach the sidecar — so a cross-vendor divergence would mean a mask written on the desktop meaning something else on Android. F3 stands. |
||
|
|
5ecb35864f |
Put the region map behind the sliders that were already there
A mask layer holds a real develop chain, so the develop panel can edit one with no new controls: select a layer and the same sliders read and write its chain instead of the graph's. An operation declared in `ops/` tomorrow becomes locally adjustable by existing, which is the payoff for making a layer a chain rather than a handful of special-cased parameters. `segmentation.rs` joins the two arms into the one thing the view needs. The model reads the image through a neutral graph rather than the edited one, so a segmentation survives an exposure change instead of being invalidated by every slider. Arm B failing is not fatal: a missing or unreadable model leaves a working watershed map, because refusing to segment at all would trade a working feature for a strict one. The overlay colours groups by a golden-angle walk over hue. Deterministic rather than random, so a region keeps its colour across a level change and the eye can track it; boundaries drawn black over the fill, because two adjacent groups landing on near hues read as one region and telling them apart is the whole reason to look at it. Clicking the photograph creates the layer if none is selected — that is how a local adjustment begins, and making the user press "add layer" first would be a step with no decision in it. Shift-click extends, and clicking a region already selected removes it, so one gesture both adds and corrects. `segment-readback` is a new dr-gpu feature and not a loosening of `readback`. The region-graph transfer is once per image on a worker; the one AC-8 forbids is per frame in the render loop. Sharing a switch would have forced a build wanting local masking to unlock the other. F3 still stands and the feature name says so. |