d2af6a398141aef99616c042780ff8beea9feecf
13
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
d777f7f44d |
Merge branch 'master' into worktree-faces-scrfd-mbf
Build and test / Desktop (Linux) (push) Failing after 25s
Build and test / Layer separation (push) Successful in 22s
Traceability / Requirement traces (push) Successful in 58s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Failing after 33m26s
# Conflicts: # docs/traceability.md # ui/dr-ui/src/develop.rs # ui/dr-ui/src/segmentation.rs |
||
|
|
ce201c7dd6 |
Name the two spaces a photograph lives in, so a turn cannot go the wrong way
Every orientation bug this codebase has had has been the same bug: a turn of the right size applied in the wrong direction. That failure is worth naming precisely, because it does not look like one — a quarter turn applied backwards lands 180 degrees from right, so the result is a plausible transform of the picture rather than anything obviously broken, and on landscape frames it is not wrong at all. It was the straighten shear, and it was the segmentation overlay, and each time it was found by eye rather than by a test. The reason it keeps happening is that "rotate 90 degrees clockwise" cannot be checked by reading it. The reader has to hold in their head which of the two images is being rotated and which way the y axis runs, and there were four hand-written copies of the permutation to hold it for: the shader prologue, its CPU twin, the thumbnail path, and the segmentation. So nothing added here says clockwise, anticlockwise, horizontal or vertical. The functions say *which space they take and which space they return* — `into_shown` and `into_stored`, `source_pixel` and `shown_pixel`, `into_shown_rect` and `into_stored_rect` — and each takes the dimensions of the space it reads from, so no caller has to work out which pair it is holding. `StoredRect` and `ShownRect` are separate types because they are the same four numbers meaning different things, which is exactly the case where a mistake is silent: a shown rect measured against stored dimensions produces a rectangle in the wrong place, not an error. Underneath there is one permutation. `source_pixel` was already shared by the prologue and the thumbnails; `source_point` is its normalised twin, written beside it so the two cannot drift, and everything else is those two read forwards or backwards. `Orientation::inverse` is the group inverse rather than `4 - turns`: mirrors apply after the turn, so undoing means undoing them first, and a mirror seen from the far side of an odd turn is about the other axis. That is the diagonal-mirror case, tags 5 and 7, and getting it wrong renders as — again — 180 degrees. Three call sites lose their own copy: the thumbnail path, `dr-ui`'s segmentation, and `dr-gpu`'s `local` example. "Upright" now means one thing across the application rather than one thing per caller. The gate that matters most is `the_render_and_the_orientation_map_agree`. The shader prologue and `Orientation` answer the same question by different routes, and until now nothing checked that they answered it the same way. It now checks every EXIF tag against every user rotation and mirror on top of it, because the composition is where the two could agree singly and disagree together. The rest earn their place by having caught something. Writing these found two real errors in this commit's own new code before it ran anywhere: `shown_pixel` was handed the dimensions of the wrong space and overflowed, and the rect map turned the wrong way for the diagonal mirrors. A round trip that returns what went in is the only check worth having here, since every wrong answer is still a picture. No behaviour changes. The permutations are the ones that were already being applied; they are simply applied from one place now. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b55812812a |
Name segmented people from the faces already recognised in them
The segmenter knows it found a person; the face index knows which person. Joining them turns "person" in the mask list into "Anna", which is the difference between a vocabulary of eighty COCO classes and one that includes the user's family. Selecting a subject in a group photograph stops being a guessing game between three identical rows. Containment, not IoU. A face is a small part of the person it belongs to, so a correct pairing has an IoU near zero and anything IoU-based would reject every true match. Confirmed names only. A suggestion is the system's guess, and printing a guessed name onto a mask region would launder it into a fact. Writing the tests corrected the design once: a tight head-and-shoulders portrait, where the face fills most of the person box, is the case where naming is most certain, not least. An earlier guard rejected exactly that and has been removed, with the reasoning left as a test because it is easy to get backwards a second time. The names hang on the develop session, set when the image opens because that is the one moment the catalog and the image id are both in reach. Every segmentation run afterwards picks them up for free, and a library with no face indexing behaves exactly as it did before. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4a82753d22 |
Show the detector the photograph, not the sensor's scanlines
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
Build and test / Desktop (Linux) (push) Successful in 19m7s
Build and test / Layer separation (push) Successful in 25s
Traceability / Requirement traces (push) Successful in 23s
Build and test / Android (aarch64) (push) Failing after 33m10s
"Find subjects" was handed the proxy in the sensor's own orientation, so every frame shot on a body held sideways reached the model lying on its side — and a model trained on upright photographs is very bad at those. Measured end to end on a 22 MP frame of two people and a dog: `person 0.36` and nothing else, against `dog 0.82, person 0.61, person 0.49` for the same pixels stood up. Nothing failed; the panel simply offered one poor subject where there were three good ones. The orientation was never dropped on purpose. The proxy is deliberately rendered through a *neutral* graph — the detection has to survive an exposure change, or every slider would invalidate the masks built on it — and neutral took the file's orientation with it along with everything else. Landscape frames were unaffected, which is why it stood for as long as it did. The turn is `Orientation::source_pixel`, the same function the grid's thumbnails already go through, so the detector and the thumbnailer now agree about which way is up rather than holding two opinions. What it is turned by is `Framing::effective_orientation` — the file's EXIF tag and the photographer's own rotations composed into one permutation, by the group law rather than by adding the turns, which is a distinction `Framing` already had to make and had already tested. Rotating the picture and pressing the button again therefore does what it looks like it does. The proxy stays in sensor space and the masks come back into it. That is not a detail to be tidied later: the generated shader samples the mask array at `uv_src`, *after* the framing map, so a mask stored upright would sit a quarter turn off the subject it was drawn around. That is a wrong mask rather than a weak one, and nothing announces it. So the picture is stood up for the model and laid back down for everything else, and `upright`/`lay_down` are returned as a pair because calling one and forgetting the other is silent. Both directions are the one function: `upright` gathers through `source_pixel` and `lay_down` scatters through it. A quarter turn is a bijection of the pixel grid, so the round trip is exact — no filter, no resampling, and no hole to fill — and an inverse written out by hand would be a second thing to keep in step, whose way of being wrong is a mask mirrored about the wrong axis, which still looks like a mask. The orientation joins the confidence and the tiling flag in the segmentation signature, and for the same reason: turning the photograph changes what the model recognises, so two runs either side of a rotation are different instance lists. Two that happened to come out the same length would otherwise share a signature and a stored layer would be silently re-indexed from one into the other. The refine pass had it too — it re-runs the model over a crop rendered in the same sensor space — so it makes the same turn, and would otherwise have handed back a worse mask than the one it was asked to improve, on the subject the photographer had just pointed at. `dr-gpu`'s `local` example is fixed with it. It exists to be the shipping path with pictures attached, and a diagnostic that reproduces the bug it is meant to catch is a trap for whoever reads it next. Seven tests. The round trip is the identity over all eight EXIF tags on a non-square asymmetric grid; a turn carries whole pixels rather than shearing the channels apart; a sideways frame reaches the model upright; a box comes back in sensor pixels, worked out by hand for the one turn a portrait frame actually writes; a restored box still reads low-to-high for every tag, since the rest of the pipeline takes `x1 - x0` without checking the sign; and the eight tags cannot collapse into one signature key. The existing composition test now runs against `effective_orientation` itself, over all 8 x 16 baseline-and-user pairs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b3641b5307 |
Let a subject's mask be refined, and let several masks be edited at once
Two related gaps in the mask panel, from the same conversation: a subject's outline is only ever as sharp as the whole-frame pass that found it, and a change meant for several layers had to be dragged once per layer. ## Refine mask A "Refine mask" button on a subject layer re-runs detection on a padded crop around that instance's own box instead of the whole frame — the subject reaches the model at its own size rather than squeezed into the model's fixed 640x640 window alongside everything else in the photograph. `RefineJob` mirrors `SegmentationJob`'s split (built on the session, run off it, adopted back), and the crop itself is rendered through `Framing::set_view` — the same ephemeral viewport the interactive zoom already uses to render a region above proxy resolution, so no new render path and no change to the model's own input size was needed. `dr-segment` is untouched: `Tiling::Whole` already treats whatever buffer it is handed as the one window. The result is still downsampled onto the shared proxy grid every instance's mask lives on, but from a sharper source than the whole-frame pass ever saw for that subject, which is what the edge actually reads out of. ## Multi-select `active_mask: Option<String>` is now `active_masks: Vec<String>`. A plain click still replaces the selection; a control- or command-click toggles one layer in or out of it. `set_param` and `reset_op` fan out to every selected layer, each set to the exact value the slider now shows rather than offset by however far it already was — one slider, one reading, applied everywhere selected. Dragging a gradient's on-canvas handle is deliberately not extended to multi-select: several gradients have no single geometry a shared handle could move, so `gradient_handles` stays empty unless exactly one layer is selected. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
7fcb8dc107 |
Give the develop column room, and the mask a way out of the way
Three faults reported together, and they share a shape: each is something the panel decided on the photographer's behalf. **The column was 280px on every screen.** That width was chosen for a tablet, where the column is a large fraction of the display and every pixel of it is taken from the photograph. On a desktop window the mode strip alone — two modes, a separator, "All", and a chip per attribute the operation set declares — does not fit, so it scrolled sideways. A control you have to pan to reach is one you do not know is there. 380px when the window is classed expanded, 280px when it is not; driven off `layout-class` because reading the window width inside the layout that sets it is a binding loop. **The overlay could not be hidden.** An earlier "Overlay" button was removed for a good reason — it *armed* the overlay, so local mode could be entered and still show nothing. Hiding is the opposite need and was never served: a mask is judged against the photograph beneath it, and that photograph is exactly what the overlay covers. `overlay-hidden` is kept separate from `overlay-on` so a recompute cannot switch the overlay back on under someone who just turned it off. **The masks were coarse because the model saw the subject small.** The graph's input is a fixed 640x640 and every frame is letterboxed into it, so a bird 200px across in a 1600px proxy reaches the model at 80px. `Tiling::Grid` has been implemented and tested since the segmentation spike and defaulted off, because it costs one inference per tile — 2.8s for a 3x2 grid against 470ms. Now offered as "Look closer (slower)", which says what it costs, rather than spending it on every image or on none. The tiling choice enters the segmentation signature. A mask stores the signature its region ids index into, and a tiled run finds different instances in a different order; sharing a signature would silently reinterpret a layer built against the coarse pass — a wrong mask rather than a stale one, and nothing announces it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c0e1179936 |
Find the subjects without stopping the window
"Find subjects" took the UI thread for two thirds of a second on a 22 MP frame — a proxy render, a readback and a YOLO pass through `ort` — and for that time the interface was simply gone. The panel apologised for it rather than hiding it: a "Looking…" label, and a 16 ms `single_shot` so the label reached the screen before the freeze began, with a comment saying the obvious fix needed the develop session restructured and was not being taken. The obstacle was never `Send`. `DevelopSession` is `Send` — the device, the source texture and the passes all are. What cannot go to a worker is the `Rc<RefCell<Option<DevelopSession>>>` that every callback in the window reaches through, and the window has to keep reaching through it while the work runs. Handing the session over would freeze the interface exactly as thoroughly as blocking on it did. So the work takes a copy of what it needs instead. A `SegmentationJob` is the device, the demosaiced source behind an `Arc`, and the name of the session that asked. Taking one is two `Arc` bumps; running one is 495 ms on this desktop; none of it touches the session, and there is deliberately no `&mut DevelopSession` in scope for a caller to hold across it. The proxy render travels with it rather than staying behind — a `GpuContext` and a texture handle are both `Send`, and the model was never the only expensive half. So does building the mask rasteriser, which is a shader compile: adopting the result was costing 23 ms, a dropped frame on the one redraw the user is waiting for, and the rasteriser is needed exactly when the subjects arrive and never before. What is left on the UI thread is a microsecond. The answer comes back through a channel a `slint::Timer` polls, which is the shape `apply_when_ready` already uses for a sidecar fetch. **A result can outlive the photograph it describes.** Two thirds of a second is long enough to press the button, think better of it and swipe to the next frame — and the result landing then would fill the panel with subjects that are not in the picture, drawing outlines around a dog two photographs back. Nothing downstream can tell: the masks rasterise and the overlay draws either way. So every session is minted with an id, a job carries the id it was taken from, and `delivery` compares the two before anything is applied. An id rather than a counter beside the session slot, because that slot is written from four places in `lib.rs` and the fifth would be the one that forgot. A discard touches nothing on the way out. `segmenting` belongs to whichever photograph is open now, which may well have a run of its own going, and clearing it would re-enable a button that is correctly insensitive. One run at a time, and abandonment is what stops that being a trap. A job left over from a photograph the user has left is displaced rather than waited for — otherwise the next frame's "Find subjects" would do nothing for the length of a run nobody wants, which is the wait this exists to remove. `ort` offers no way into the inference, so abandoning is checked at the seams there are: before the job starts, and between the readback and the model. Abandoned early it costs nothing, abandoned mid-inference it costs the run it was already committed to, and either way the answer is dropped at the channel. `DevelopSession::segment` survives as a test-only convenience. Left public it is precisely the shape that put two thirds of a second on the UI thread in the first place, and the next caller would reach for it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
37f63edbf9 |
Take the watershed out of the product path
It does not work on a photograph, so nothing should offer it. `Segmentation` is now one model pass and what it recognised: no region field, no merge tree, no label upload, no granularity slider, and no readback of the whole proxy to build a graph that collapses. A click means "the object under the cursor". The region-selection path went with the hierarchy it indexed — including the shift-click add/subtract, which has no meaning for a whole object and would have been a modifier that silently did nothing. The passes, the hierarchy and the semantic prior stay in `dr-gpu` and `dr-segment`, tested and documented. It is the *merge criterion* that fails — the saddle is the minimum gradient along a boundary, so one weak pixel merges two regions and real gradient noise puts a weak pixel on every boundary. That is one function to replace, and the evidence for replacing it is worth keeping. What is gone is the wiring, the option, and the control that offered a user a choice with no outcome. `MaskSource::Regions` remains in the pipeline: it is tested, it round-trips through the sidecar, and a stored layer that names regions must still load and be reported stale rather than failing to parse. |
||
|
|
ec713585a5 |
Measure the distance to the edge, and get four controls for one transform
Feathering, growing, shrinking, closing and opening are the same number read differently. With the signed distance from the boundary in hand, dilation is the set where d >= -r, erosion where d >= +r, and a feather of any shape is a function of d. So the field is computed once and the controls are arithmetic on it. The **field** is what reaches the GPU, not a finished alpha, and that is the point: growing a mask or changing its falloff then costs a uniform upload and no recomputation, which is what makes them live controls rather than ones that stall on every drag. Only closing and opening rebuild, because after the first threshold the shape has changed and the old distances describe the old one. Exact Euclidean, via Felzenszwalb's separable transform — not a chamfer approximation, which leaves a mask visibly octagonal once grown more than a few pixels. A test asserts the diagonal is √2 rather than 1 or 2. It runs on the CPU, which ARCH §5.4 forbids for masks. The rule is about brush lag — a stroke rasterised per frame — and this is a different operation: once per mask edit, on input the model already produced here, producing a field the GPU then samples for free. What it buys is exact determinism, which matters because masks reach the sidecar as indices and a field that varied by vendor would mean a mask meaning one thing on the desktop and another on the phone. The half-pixel in `signed_distance` is not a detail, and a test caught it. Measuring to the nearest opposite pixel *centre* puts the smallest magnitude at 1 either side, so the boundary is nowhere and **eroding by less than a pixel removes nothing**. A control whose first notch does nothing is a broken control. Half a pixel off each side puts the boundary where it physically is, and eroding by 1 takes exactly the outermost ring. Every falloff curve is 0.5 at the boundary by construction, asserted for all five: changing the curve should change how the transition looks and never where it sits. |
||
|
|
ee10097435 |
Mask the subject the model found, not the regions underneath it
The watershed hierarchy does not survive a photograph, so local masking stops depending on it. A layer can now be one recognised object, and the object's own coverage is the mask. `Options::watershed` defaults off. It costs ~80 ms plus a full-resolution readback to produce a ladder that collapses, and paying that on every photograph buys a control that misleads. Kept switchable rather than deleted: the passes and the hierarchy are correct in themselves and it is the merge criterion that fails, which is a change to one function. Masks now rasterise in **source** space at proxy resolution and are sampled by the composed shader after the framing map. That fixes a real bug: they were rasterised in output space, so zooming slid the photograph underneath a mask that stayed pinned to the viewport, and cropping moved every adjustment to a different part of the picture. Doing it this way also leaves the framing map in exactly one place — a second copy in the mask shader would have been a second thing to keep in step, failing only when straightened. A subject is stored as identity, not pixels: the mask is megabytes and is reproducible by running the same model over the same image, so the sidecar carries the index, the class and the score, and the session carries the pixels. The class is there to be checked — if instance 3 comes back a "car" where it was a "dog", something changed and the layer is stale rather than silently masking the wrong thing. The overlay now draws instances and is transparent everywhere else. The region version covered every pixel and so hid the photograph it was drawn over; the question it exists to answer is whether an outline follows the subject, which you can only answer by seeing both. `examples/local.rs` is the worked example: subject in colour with the rest monochrome, and the subject lifted out of its background. Run on a 5472x3648 CR2 it finds two people and two cars, and the colour-pop keeps her hat and hair while the wall and grass behind go grey. |
||
|
|
b1433ad4a9 |
Give a mask an edge treatment, and find out the watershed has none worth having
Two things, and the second is why the first matters more than expected. Mask layers gain a feather, a falloff curve and a morphology, all defined against a signed distance from the boundary rather than as separate features — one exact distance field answers "how soft" and "how far" at once, so dilation is a threshold at -r, erosion one at +r, and closing and opening are one of each in sequence. The compound pair costs a second distance field, which is why they are named rather than presented as a radius that happens to be signed. Types, defaults and sidecar round-trip only; the field itself is next. `edge-feather` and `edge-falloff`, not `feather` and `falloff`, because a radial mask already writes `feather` for the fraction of its radius it ramps over. Same word, different quantity, different units — sharing the key would have made an existing file ambiguous. The diagnostic that provoked this is committed as an ignored test, because "does the ladder land on things a person means" is the question S15 exists to answer and it should not depend on whoever still has the script. On bus.jpg it answers badly: 35,075 regions at blur 2 over an 810x1080 frame, and cutting that to 400 gives *one* region covering nearly the whole picture plus 399 noise specks. Not over-segmentation — collapse. Almost every saddle is near zero, so the merge order joins everything meaningful before it joins anything spurious, and a global cut spends its entire budget on grain. So the granularity ladder does not currently work on a photograph, and the region masks built on it inherit that. Recorded rather than worked around: the next commits move local masking onto the model's instances, where the edge treatment above is what makes a quarter-resolution mask usable. |
||
|
|
12d320cf33 |
Record what the spec got wrong about the model that exists
docs/segmentation.md §4 priced arm B as costing a C dependency under the NDK and treated that as most of the difference between the arms. It is not a cost that has to be paid: `ort`'s `alternative-backend` disables its linking entirely and `ort-tract` supplies the API from tract, which is pure Rust. D13's "largest exception the policy would tolerate" turns out not to be needed, and the answer generalises to the face pipeline — so D13's runtime half is now answered and only its licensing half is open. Three findings contradict §4 outright and are recorded as F4-F6 rather than quietly designed around. There is no ADE20K-trained YOLO, so the shipped vocabulary selects subjects and not stuff — "select the sky" comes from the watershed or from nowhere. It is instance segmentation, so it partitions nothing and two people come back as two instances. And tract cannot parse a dynamic-shape export, which fixes the input at 640 square and makes tiling the only route to more semantic resolution. Arm C ships, but §8's criteria are not what decided it, and saying so matters more than claiming the process worked. §8 asked for a two- interaction margin over arm A on a traced corpus. That comparison was never run: F4 and F5 changed what the arms are, and a model that recognises subjects but has no word for sky cannot be a selection tool alone, while a watershed cannot tell a person from the wall behind them. They stopped being candidates and became complements. What is *not* done is written down as plainly: the 24-image corpus is untraced, so M1-M4 have no numbers and "this feels right" has not become one. M5 is answered on one device only, and region ids now reach the sidecar — so a cross-vendor divergence would mean a mask written on the desktop meaning something else on Android. F3 stands. |
||
|
|
5ecb35864f |
Put the region map behind the sliders that were already there
A mask layer holds a real develop chain, so the develop panel can edit one with no new controls: select a layer and the same sliders read and write its chain instead of the graph's. An operation declared in `ops/` tomorrow becomes locally adjustable by existing, which is the payoff for making a layer a chain rather than a handful of special-cased parameters. `segmentation.rs` joins the two arms into the one thing the view needs. The model reads the image through a neutral graph rather than the edited one, so a segmentation survives an exposure change instead of being invalidated by every slider. Arm B failing is not fatal: a missing or unreadable model leaves a working watershed map, because refusing to segment at all would trade a working feature for a strict one. The overlay colours groups by a golden-angle walk over hue. Deterministic rather than random, so a region keeps its colour across a level change and the eye can track it; boundaries drawn black over the fill, because two adjacent groups landing on near hues read as one region and telling them apart is the whole reason to look at it. Clicking the photograph creates the layer if none is selected — that is how a local adjustment begins, and making the user press "add layer" first would be a step with no decision in it. Shift-click extends, and clicking a region already selected removes it, so one gesture both adds and corrects. `segment-readback` is a new dr-gpu feature and not a loosening of `readback`. The region-graph transfer is once per image on a worker; the one AC-8 forbids is per frame in the render loop. Sharing a switch would have forced a build wanting local masking to unlock the other. F3 still stands and the feature name says so. |