00663a870b091ff9407e6cfc03c37d0fa989d94a
3
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c396a22dfd |
Paint a mask without ever rasterising one on the CPU
The last line of FR-DEV-3, and the mask ARCH §5.4 was written for. darktable rasterises drawn masks on the CPU and users call the result unworkable; the architecture's answer is that a stroke arrives as *parameters* and the device draws it. This is that, from the model through the sidecar to the pixels — but not the finger: the canvas is somebody else's change, and this leaves it a seam rather than reaching into it. **A stroke is a swept disc along a polyline**, plus erase, radius, hardness and flow. `MaskSource::Brush` holds an ordered list of them, and the order is the mask: an erase after an add takes it away and the same pair reversed does not. Nothing about it is pixels, which is what makes a mask that costs a line of text, diffs by the gesture, and survives a crop, a straighten and an export at any size — the properties a stored raster has none of, and the same argument the region ids were chosen for. Two things keep the point count honest. While the finger is down, a position closer to the last than an eighth of the radius is dropped: a touch screen reports 120 a second, so a finger held still for five seconds is six hundred points in the same place, and simplification would only remove them once the gesture had ended — after every frame in between had drawn all of them. When it ends, Douglas–Peucker at an eighth of the radius removes what a disc that wide cannot express: a swept circle moved by r/8 moves its own edge by r/8, which is inside the soft part of any brush. Coordinates snap to a ten-thousandth of the frame on the way in *and* are written at that precision, so a round trip is exact rather than nearly exact — a file that drifts in the sixth decimal every save is a per-field merge conflict a day, over nothing. **Cost is why the strokes are not drawn by the full-screen triangle the other masks use.** A swept disc is the minimum distance to any of its segments, so a stroke over the whole frame costs `pixels × segments` and both terms grow together — the quadratic that is darktable's problem moved onto the GPU rather than solved. Each stroke is instead drawn over its own bounding box, grown by the radius, so the rasteriser never invokes the shader for a pixel the stroke cannot reach: `area(box) × segments`, which for a dab or a swipe is a small fraction of the frame. A gesture past 256 points continues as a second stroke for the same reason, since a shorter stroke has a smaller box. Add and erase are `dst + a(1 - dst)` and `dst(1 - a)`, which are exactly a source-over and a one-minus-source blend — so they are blend state, not arithmetic, and no pass ever reads the slice it is writing. That is what permits one draw per stroke at all. Within a stroke the coverage is the *minimum* distance over its segments rather than a sum: a path that crosses itself must not build up where it did, or every circle and every scribble would be blotchy wherever consecutive dabs overlap, which is everywhere. Not a distance field, deliberately. `dr-segment`'s transform documents the two conditions that make CPU work right there — once per mask edit, over input already CPU-side — and a stroke fails both: it changes while the finger moves, and its input is a handful of coordinates that never needed to be pixels. It also needs no transform, because the distance to a swept disc is closed form. A stroke is the one mask whose distance field is known without computing one. An unpainted brush layer is inactive rather than empty, which is not an optimisation: `invert` turns empty into everything, so a layer created with invert already set would apply its adjustment to the whole photograph before a single stroke was made. That is the loud, confident kind of wrong this codebase refuses everywhere else a mask can go missing, and there is a rendered test for it. The tests read pixels back off a device rather than checking that the two halves agree with each other. What they pin down is what is silent when wrong: the y flip between mask space and clip space, which a centred stroke would not notice; a bounding box not grown by the radius, which makes a tap draw nothing at all; an aspect ratio ignored, which makes a dab an ellipse on any frame that is not square; a stroke doubling back and building up; and an erase that lost its place in the order and put back paint the user had taken off. Not done here: the interaction. The canvas needs to begin, extend and end a stroke on the active layer, and `DevelopSession::rasterise_masks` still returns early without a segmentation — it takes the proxy size from one, and a brush needs no model to have run over the photograph first. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ee10097435 |
Mask the subject the model found, not the regions underneath it
The watershed hierarchy does not survive a photograph, so local masking stops depending on it. A layer can now be one recognised object, and the object's own coverage is the mask. `Options::watershed` defaults off. It costs ~80 ms plus a full-resolution readback to produce a ladder that collapses, and paying that on every photograph buys a control that misleads. Kept switchable rather than deleted: the passes and the hierarchy are correct in themselves and it is the merge criterion that fails, which is a change to one function. Masks now rasterise in **source** space at proxy resolution and are sampled by the composed shader after the framing map. That fixes a real bug: they were rasterised in output space, so zooming slid the photograph underneath a mask that stayed pinned to the viewport, and cropping moved every adjustment to a different part of the picture. Doing it this way also leaves the framing map in exactly one place — a second copy in the mask shader would have been a second thing to keep in step, failing only when straightened. A subject is stored as identity, not pixels: the mask is megabytes and is reproducible by running the same model over the same image, so the sidecar carries the index, the class and the score, and the session carries the pixels. The class is there to be checked — if instance 3 comes back a "car" where it was a "dog", something changed and the layer is stale rather than silently masking the wrong thing. The overlay now draws instances and is transparent everywhere else. The region version covered every pixel and so hid the photograph it was drawn over; the question it exists to answer is whether an outline follows the subject, which you can only answer by seeing both. `examples/local.rs` is the worked example: subject in colour with the rest monochrome, and the subject lifted out of its background. Run on a 5472x3648 CR2 it finds two people and two cars, and the colour-pop keeps her hat and hair while the wall and grass behind go grey. |
||
|
|
c6a846a1f9 |
Brighten her face without touching the sky behind her
A mask layer is an ordinary develop chain plus a rule about where it applies. Nothing in the chain knows it is being masked, so every operation that works globally now works locally and a newly declared op in `ops/` arrives with local support already done. The composer emits each layer after the global chain and before the conversion out of camera space, which is what a photographer means by "and *then* lift the shadows on her face". Op fragments write to a `c` they expect to own, so a layer block shadows it and copies the result back out through a carrier — assigning the outer one from inside is impossible precisely because it is shadowed. The fused dispatch survives: three global adjustments and two masked ones remain one shader, one read, one write. Masks rasterise on the GPU and never exist in CPU memory (ARCH §5.4). That is the whole reason darktable's brush masks lag, and it is architectural rather than tuning, so it is not a thing to inherit and fix later. The rasteriser is a render pass rather than the compute shader it obviously wants to be, and the format is why: R8Unorm is not a core storage format, so a compute path has to widen masks to four bytes per pixel — 768 MB across eight layers of a 24 MP export, against 192 MB at one byte. A colour attachment takes R8Unorm happily. The array slice comes from the attached view, so no slot uniform exists to disagree with where the pass writes. Region masks index a compacted label field rather than the watershed's raw basin roots, because a root is a sparse index into pixel space and indexing a per-region array by one would need a table the size of the image. Changing a selection then costs a few kilobytes, not a re-upload. Stored as region ids, not as pixels: diffable, mergeable per-field under FR-NC-9, and cheap in a sidecar. The ids only mean anything alongside the segmentation that produced them, so each layer carries that signature and is treated as stale rather than applied when it does not match — a confidently wrong mask being much worse than an absent one. Seven device tests render actual frames and read them back. The unit tests either side check halves that would both pass if the two agreed with each other and were both wrong; a mask sampled with x and y swapped satisfies them and fails these. |