Every neighbourhood pass so far has been a convolution, whose whole
description fits in the uniform block because its structure fixes how many
numbers it needs. Spot removal is not that shape: sixty-four repairs and
one repair are the same shader with a different buffer behind it.
So a pass may declare `storage`, which arrives at binding 3 as
`array<vec4<f32>>` with `arrayLength` in scope. The alternative — packing
the list into uniforms — needs a fixed maximum paid for on every frame, a
composer that can emit vec4 fields because a uniform array's stride is 16
whatever it holds, and it gives the next operation that wants a table
nothing to build on.
The property worth having is what stays out of the generated source: the
count is in the buffer, so placing the tenth spot uploads 512 bytes and
reuses the compiled pipeline, exactly as moving a slider does for the
fused pass. `changing_the_list_does_not_recompile` is that, asserted.
One bind group entry rather than two more layouts, and one placeholder
buffer allocated in `new` rather than sixteen bytes per pass per frame —
a zero-length storage buffer cannot be bound, and per-frame allocation is
what this module's documentation exists to refuse.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
An emulsion is a suspension of crystals. Light sensitises some; development
turns a sensitised one opaque, all or nothing. So a patch of film's density
is a *count* of developed grains, and a count of independent yes/no events
has a variance whether or not anyone wanted texture:
mean = D
variance = D * (Dmax - u * D) / N
That expression is the whole feature. It peaks in the middle of the density
range and vanishes at both ends -- clear film has nothing developed to vary,
black film has nothing left to develop -- so grain lives in the midtones as a
consequence rather than as a "midtone bias" slider.
I was wrong earlier that this needs the detail stage. Nothing in it reads a
neighbouring pixel; the only reason to move it was that grain must be fixed in
film space rather than screen space, and that solves itself: N is grains *per
pixel*, so it scales with the film a pixel covers. Zoom out, each pixel
averages more grains, less variance -- correct, with nothing super-sampled and
nothing filtered. It stays in the fused pass.
Grain goes on the density and *before* the dye, which is the physical order
and not cosmetic. Perturbing the finished colour -- what an effect does --
tints highlights wrong, because that noise never passes through the dye.
Crystal habit lives in `rms_granularity`, the number every datasheet
publishes, now a profile field. It measures exactly what differs between a
cubic emulsion and a tabular one: at equal speed, tabular crystals present
more area per unit silver, so the film reads finer. Delta 100 is quoted near 9
where HP5 is near 12, and that gap *is* the habit. Adding a stock whose grain
is its whole reputation is therefore editing one line, not writing a model.
Three things this cost, all of them worth writing down:
- The default granularity is a colour negative's, blue coarsest. Applied to
Tri-X it put *colour* speckle on a black and white photograph. Monochrome
stocks collapse it at parse, where every other per-layer table is already
replicated from the one measured channel.
- Helpers cannot read uniforms. The composer prefixes a uniform with its
operation's id and rewrites references inside a fragment body only;
helpers are shared and deduplicated, so a bare `gn0` names nothing.
`film_lut` already took its size as an argument for this reason, and now
says so.
- The end-to-end test compares the shader against the CPU model, and grain
is stochastic, so that comparison now runs with grain off. Which means a
grain that never left the CPU would look exactly like a passing suite --
hence a second test that grain off is bit-identical, one grain per pixel
moves it, and ten thousand move it less.
Not here, deliberately: no grain slider. The parameters are physical and
`rms_granularity` is the honest place to scale one from, but its range wants
choosing rather than guessing. Nor a film format -- 35 mm is assumed, and
medium format at the same stock is far less grainy per unit of picture.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Nothing here is film simulation. These are lints that fail master today,
under the -D warnings CI runs with, mostly from a toolchain that learned
new ones rather than from anybody's code -- `is_multiple_of` and the
derivable `Default` did not exist as lints when this was written.
They are fixed rather than allowed, and by hand rather than by trusting
`cargo clippy --fix` wholesale: its automatic pass split a derive in two
and left a stray blank line, which is the sort of thing that is correct
and still wrong to commit.
The four that needed a decision rather than a rewrite:
- The distance transform's inner loop writes through its iterator now.
`q` stays, because it is the position the parabola is evaluated at as
well as the index it is written to -- the lint is about the write.
- `to_source` and `to_proto` take `self` by value. Their receiver is
`Copy`, so this is the same machine code and the honest signature.
- The export path's return type is five levels deep and now has a name,
plus a line saying why the `Option` wraps the `Result`: `None` is
cancellation, which is not a failure and has no error to report.
- A test fills a range instead of looping over one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CI runs cargo fmt --check and clippy -D warnings, and this branch had
never been through either. Both would have failed it.
The bulk was the generated colour tables: eight significant figures where
an f32 carries about 7.2, so the eighth is noise that rounds away at
compile time and clippy's excessive_precision says so 109 times over.
Fixed in the generator rather than only in the file, so it stays fixed --
and the file is trimmed in place rather than re-derived, because
regenerating it needs a colour-science stack that has nothing to do with
the defect.
The format! in the composer is mine too, from extracting the rendering
tail: the braces in it were escaped because the text used to live inside a
larger template, and once extracted the escapes are noise and the call
formats nothing.
Also here, and clearly not mine: an unused import and a shadowed binding
in dr-gpu, and an unused import in a test. They are pre-existing --
clippy has been failing on master before this branch existed, on lints
like is_multiple_of that arrived with a toolchain rather than with
anyone's code. Fixed because CI cannot go green around them, and called
out because a merge commit is a bad place to quietly edit someone else's
crate.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The stock model rendered correctly and nothing could ask for it. This is
the picker, and the sidecar key that makes the choice outlive the session.
How the choice persists was the open question, and the answer was already
written down twice in sidecar.rs: `rating` is a top-level key "because a
rating is not an edit", and `masks` are one "because a layer is not a
scalar". A stock is that kind of thing -- a choice of material, not a
number a slider moves -- so it is a top-level key too.
It stores the **id**, not an index. Stocks are files that users add, so an
index would mean installing a profile silently changed which film every
existing photograph had been developed on. A name this build has no
profile for still round-trips untouched, because the alternative is that
syncing to an older phone quietly un-develops the picture.
Only the names travel. Turning one back into tables needs the profile
database, which dr-pipeline deliberately does not link, so `Version::apply`
clears the film and the session re-bakes -- after the parameters, because
the bake reads the film's own exposure sliders and the print balance is
solved against them. That is also why moving those sliders rebuilds the
lookup where no other control in the panel does: an enlarger's filtration
depends on how the negative was exposed.
The panel keeps its rule. It still names no operation and still generates
every control from a declared parameter kind; the stock gets a bespoke
control beside those, exactly as the mask stack does, and for the same
reason. The film's exposure and print exposure arrive as ordinary
generated sliders.
Two defaults worth stating. Picking a colour negative prints it, because
an unprinted one is an orange strip and offering that as the first thing
somebody sees after choosing Portra reads as a bug rather than as a
choice -- the toggle is there for anyone who wants the scan. And a paste
carries no film: a preset is a parameter map, and a stock is not a
parameter, so pasting one would paste a choice the clipboard never took.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The stock model landed in dr-film with no way to see it. This is the
pipeline node, the two texture bindings it reads, and the end-to-end test
that proves the shader agrees with the model.
The design point is that a film simulation is not an adjustment. Every
other node changes a picture; this one makes it. A stock's characteristic
curve does the camera profile's base curve's job -- from measurements
rather than from a curve somebody drew -- so running both renders the
scene twice: the camera's rendering, and then a film's rendering of that.
It looks like neither, and it reads as a colour-management bug with no
colour-management bug to find.
So `Operation::renders` is new. A node declaring it takes camera RGB and
hands back linear sRGB, and the composer emits neither the base curve nor
the conversion out of camera space. Both halves move together, and the
composer keeps them as one string precisely so that getting half of it
right is impossible.
The tables are not parameters, for the reason vignetting's coefficients
are not: they are measurements. dr-pipeline declares the layout as a plain
struct and keeps its no-dependency property; the two crates share no types
on purpose. `EditGraph::set_film_tables` offers them to every node rather
than to the one that wants them, because knowing which concrete type is
which is what the graph is organised not to know.
Bindings 4 and 5 follow the masks precedent: declared unconditionally so
one bind group layout serves every generated shader, bound to 1x1
placeholders when no stock is loaded. Both are interpolated by hand with
textureLoad -- this pipeline binds no sampler, and adding one for two
lookups would cost a binding in every shader. Uploads are keyed on content
so an unchanged stock does not push half a megabyte across the bus per
frame.
The end-to-end test earned its place immediately: it found the density
lookup being filled z-fastest while a 3D texture upload wants x-fastest,
so the red and blue axes were transposed. Green matched exactly, which is
what that bug looks like -- a plausible photograph of the wrong colour,
and one that every unit test on either side of the seam passes. dr-film
now pins the layout in a test that needs no device, and states it where
the field is declared.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`cargo fmt --check` is a required step and had drifted across 45 files. Most of
it arrived this week: several operations were written in parallel worktrees and
merged by hand, and a hand-merge resolves conflicts without ever running the
formatter over the result.
No behaviour changes — this is `cargo fmt --all` and nothing else, kept as its
own commit so the next reader can skip it wholesale rather than search it for
one that matters.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`texture_contributes_nothing_where_its_scale_does_not_exist` asserted that the
detail chain composed *nothing* when texture's kernel rounded away, and its
comment recorded that empty chain as a gap: the fused pass had already decided
to hand on linear working values, so an empty chain left the output transform
undone and the render was rejected. It said fixing it meant composing both
halves together, at the composition boundary rather than in that file.
Noise reduction closed it there in the same round, by emitting a bodyless
`detail/resolve` pass for exactly this case. So the assertion was describing a
defect that no longer exists, and failing because the defect was fixed.
Now asserts the property it was always about — texture contributes no kernel,
`radius() == 0` — while the chain carries the one pass that finishes the
render. Two agents working in parallel each saw one half of this; it is only
visible with both merged.
A test that cannot be run cannot be checked by running it, and a number
copied out of a test run agrees with whatever the code did on the day.
Each asserted kernel width, tolerance and overshoot bound now carries the
arithmetic that produces it -- the shorter edge, the sigma, the
truncation at two sigmas, and where the rounding falls -- so a reader can
verify the expectation against the recipe without a GPU or a compiler.
Also records the two places where a bound is a bound and not a
measurement: the tolerance in the frame-fraction test is exactly what
rounding a kernel to a whole pixel costs on the smallest frame it uses,
and the halo test's floor and ceiling bracket a peak derived from the
step, the soft limit and the midtone taper rather than from a run.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`the_whole_chain_at_once_compiles` counted one `---- ` block per
operation in the chain. That was true while every operation was a point
operation, and stopped being true the moment a neighbourhood one existed:
clarity and texture are active in that test, and still emit no fused
block, because `compose_full` filters them out and the detail stage
dispatches them separately.
Counted now by asking each operation whether it has a detail stage --
the same question the composer's own filter asks -- rather than by
subtracting a number someone has to remember to update. Capture
sharpening and noise reduction are covered by this without another edit.
The render at the end is now `render_detailed`, which is not a
concession but the stronger test: with a neighbourhood operation active
the fused pass hands on linear working values and the last detail pass
performs the output transform, so rendering the fused half alone is the
mismatch `render_detailed` exists to reject -- and the detail passes are
generated WGSL with uniform blocks of their own, which is exactly what
"everything at once" is here to collide. It renders at 512 rather than
32 because a compositional radius is a fraction of the frame, and on a
32-pixel target every detail kernel rounds away to nothing.
Also records, in `texture_contributes_nothing_where_its_scale_does_not_exist`,
the seam this uncovered: an active detail operation whose kernel rounds
away composes an empty chain while the fused pass has already been
composed to hand on linear values, and nothing can then encode the
result. That test now asserts the property on the composed chain instead
of driving the unrenderable configuration.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A colour square wave built from equal, opposite swings of red and blue is
not a pure colour pattern: Rec. 709 weights them 0.2126 and 0.0722, so it
carries a luminance square wave of about a seventh of the swing underneath.
Correlating the raw red channel therefore reads a constant floor that the
chroma filter is not meant to remove, which compressed every ratio towards
one - enough that the resolution test could no longer tell a correct kernel
from one twice the size. Correlate the colour difference instead, and write
the derivation of each expected value into the test.
Capture sharpening as a two-pass unsharp mask in the detail stage: blur
along x, then along y, each pass applying a one-dimensional high-pass to
luminance so the composite preserves a flat field exactly and matches the
textbook kernel on any locally one-dimensional edge.
The radius is stated in source pixels and converted once per render, so a
radius tuned on a fit view is the radius the exported file gets. Below one
render pixel the operation declines to draw rather than showing sharpening
the file will not contain, and emits a single pass-through that still
carries the output transform.
The develop session now renders through render_detailed, which is what
lets an active neighbourhood operation reach the screen at all.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Checkpoint committed by the coordinator, not by the authoring agent: the
session hit its API limit mid-task and left this work uncommitted. Committed
so it survives, NOT because it is finished - expect failing tests and
half-applied changes. The agent resumes from here.
Checkpoint committed by the coordinator, not by the authoring agent: the
session hit its API limit mid-task and left this work uncommitted. Committed
so it survives, NOT because it is finished - expect failing tests and
half-applied changes. The agent resumes from here.
Checkpoint committed by the coordinator, not by the authoring agent: the
session hit its API limit mid-task and left this work uncommitted. Committed
so it survives, NOT because it is finished - expect failing tests and
half-applied changes. The agent resumes from here.
The declaration in ops/tone_curve.yaml has claimed per-channel curves since
it was written — it is the justification for the operation carrying both
`tone` and `colour`. Only the master curve existed. This is the other three.
The master runs first and the channels grade its result. Both orders are
real images and they differ visibly, so the choice is made and written down
rather than left to the loop: a point placed on the blue curve should act on
the tone the photographer can see, which is what the master has already
produced. The other order anchors the grade to tones the master is about to
move, so adjusting contrast slides a warm shadow up into the midtones.
Every id that existed before today is spelled exactly as it was. The master
curve keeps `p2_y` and the new curves take `r_`, `g_` and `b_` prefixes, so
a sidecar written when there was one curve loads, means what it meant, and
renders the same shader — asserted on the generated source, not on the
parameter values. Nothing needed a version check because nothing was
renamed.
Each curve reaches the shader only when it has been moved off the diagonal,
so an S-curve and no colour work generates what it generated when this
operation held ten parameters instead of forty, down to the uniform names.
The monotonicity guarantee is enforced per curve: a coincident pair on blue
divides by zero exactly as thoroughly as one on the master.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The unit tests either side of the base curve check halves — that the shipped
database parses and lifts its midtones, and that the generated WGSL evaluates
a curve in the right place. Neither would notice if the two agreed with each
other and both were wrong: a curve packed into the wrong uniform slots, or a
flag read from the wrong component, satisfies both and renders nothing.
So render real pixels. A flat frame through a neutral edit, once with the
Canon EOS 6D's curve looked up by name from the YAML and once with the
identity, asserting what a base curve is actually for — midtones lifted,
black still black, white still white, monotone the whole way — plus the
number an unprofiled body must still produce, so "never worse than today"
is a value rather than a promise.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The fused pass hands a fragment a colour and no coordinate. That is what buys
one dispatch for a whole edit, and it is also a wall: sharpening, noise
reduction, clarity, texture, dehaze and spot removal are each defined by what
the neighbours are doing, and FR-DEV-3 and FR-DEV-8 ask for all six. None of
them could be written at any price.
So there is now a detail stage. An operation implements `Operation` for its
parameters exactly as before — the panel, the sidecar, the history and the
presets all work unchanged — and additionally returns `Affects::Detail` and a
`DetailStage` yielding one pass per dispatch. `Affects` grows the third variant
`docs/requirements.md:250` designed and nothing had cut.
Where the stage sits is a colour-science decision, not an arrangement of
convenience. It runs after every point operation and every mask layer, so an
amount chosen against a tone curve survives the curve moving; in linear sRGB
after the camera matrix, because camera RGB has no luminance to sharpen
against; and before the output transform and the clip, because FR-DEV-2 allows
one quantisation and a highlight clipped before a convolution grows a dark
ring. The fused pass therefore ends one of two ways, and when a detail stage
follows it hands on unclipped f16 and the last detail pass encodes.
At render resolution rather than on the source, which is the whole of FR-DSP-1:
a pass before the framing prologue would cost 24 MP to draw a 2 MP preview.
`RenderScale` is what makes that survivable — a radius is stored as a fraction
of the frame's shorter edge, exactly as a mask feather already is, or as a
count of source pixels, and converted per render. It also reports when a radius
is smaller than a proxy pixel rather than drawing a plausible lie; zooming to
1:1 makes the preview exact with no second path.
`Invalidation` gives FR-DEV-3d something to mean. Moving a detail parameter
leaves the colour key alone, so `AdjustPass` keeps the linear intermediate and
skips the fused dispatch: dragging a sharpening slider costs a convolution.
Moving exposure does re-run the detail passes, because they read what the
colour pass wrote, and there is no arrangement of keys that avoids it while
keeping sharpening after tone.
Validated by a separable box blur that is not a develop operation, behind the
`detail-probe` feature and absent from a shipping build. An abstraction with no
consumer is a guess; a box blur's answer is known in closed form, so the tests
assert every byte of the ramp rather than that the edge got softer.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A linear mask at 45° was not at 45°, and a radial with equal radii was an
ellipse. Both on every photograph that is not square, which is all of them.
The geometry is stored in normalised coordinates so that a mask survives a
crop, a zoom and an export at another size — that part was right. What was
wrong is that a *distance* was being measured in those coordinates too, and a
fraction of the width is not the same length as a fraction of the height. So
`dot(uv - centre, axis)` measured the ramp in a space one of whose axes is
squashed against the other by the aspect ratio, and the iso-lines came out
sheared: on a 3:2 frame a ramp asked for at 45° arrives at about 34°.
Nothing announces it. The stored numbers are exactly what was written, the
shader is doing exactly what it says, and the only place the fault exists is
between the photographer's intent and the picture. It has been invisible so far
because there is no way yet to place a gradient by eye — the handles that make
it visible are what turned it up.
So distances and angles move into the frame's own isotropic units: y spans
`0..1` and x spans `0..aspect`, which makes a circle round and 45° a real
diagonal. The centre stays a plain fraction of each axis, because it is a point
and a point has no such problem — and because that is the space a click arrives
in. `frame_delta` is the one conversion and must stay the only one; the mask
array's own dimensions carry the aspect, so it costs no uniform.
The sidecar format does not change. What changes is what the numbers mean, and
the only geometry in the wild is a default that has never been movable.
The two tests are at 96×64 rather than square, which is the whole point: on a
square target this bug cannot be reproduced, and every existing mask test was
square. Both fail without the conversion — the radial reaching 28px sideways
where it reaches 19px down, and the diagonal landing on the wrong side of the
line it is supposed to lie along.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The last line of FR-DEV-3, and the mask ARCH §5.4 was written for. darktable
rasterises drawn masks on the CPU and users call the result unworkable; the
architecture's answer is that a stroke arrives as *parameters* and the device
draws it. This is that, from the model through the sidecar to the pixels — but
not the finger: the canvas is somebody else's change, and this leaves it a
seam rather than reaching into it.
**A stroke is a swept disc along a polyline**, plus erase, radius, hardness and
flow. `MaskSource::Brush` holds an ordered list of them, and the order is the
mask: an erase after an add takes it away and the same pair reversed does not.
Nothing about it is pixels, which is what makes a mask that costs a line of
text, diffs by the gesture, and survives a crop, a straighten and an export at
any size — the properties a stored raster has none of, and the same argument
the region ids were chosen for.
Two things keep the point count honest. While the finger is down, a position
closer to the last than an eighth of the radius is dropped: a touch screen
reports 120 a second, so a finger held still for five seconds is six hundred
points in the same place, and simplification would only remove them once the
gesture had ended — after every frame in between had drawn all of them. When
it ends, Douglas–Peucker at an eighth of the radius removes what a disc that
wide cannot express: a swept circle moved by r/8 moves its own edge by r/8,
which is inside the soft part of any brush. Coordinates snap to a
ten-thousandth of the frame on the way in *and* are written at that precision,
so a round trip is exact rather than nearly exact — a file that drifts in the
sixth decimal every save is a per-field merge conflict a day, over nothing.
**Cost is why the strokes are not drawn by the full-screen triangle the other
masks use.** A swept disc is the minimum distance to any of its segments, so a
stroke over the whole frame costs `pixels × segments` and both terms grow
together — the quadratic that is darktable's problem moved onto the GPU rather
than solved. Each stroke is instead drawn over its own bounding box, grown by
the radius, so the rasteriser never invokes the shader for a pixel the stroke
cannot reach: `area(box) × segments`, which for a dab or a swipe is a small
fraction of the frame. A gesture past 256 points continues as a second stroke
for the same reason, since a shorter stroke has a smaller box.
Add and erase are `dst + a(1 - dst)` and `dst(1 - a)`, which are exactly a
source-over and a one-minus-source blend — so they are blend state, not
arithmetic, and no pass ever reads the slice it is writing. That is what
permits one draw per stroke at all. Within a stroke the coverage is the
*minimum* distance over its segments rather than a sum: a path that crosses
itself must not build up where it did, or every circle and every scribble
would be blotchy wherever consecutive dabs overlap, which is everywhere.
Not a distance field, deliberately. `dr-segment`'s transform documents the two
conditions that make CPU work right there — once per mask edit, over input
already CPU-side — and a stroke fails both: it changes while the finger moves,
and its input is a handful of coordinates that never needed to be pixels. It
also needs no transform, because the distance to a swept disc is closed form.
A stroke is the one mask whose distance field is known without computing one.
An unpainted brush layer is inactive rather than empty, which is not an
optimisation: `invert` turns empty into everything, so a layer created with
invert already set would apply its adjustment to the whole photograph before a
single stroke was made. That is the loud, confident kind of wrong this codebase
refuses everywhere else a mask can go missing, and there is a rendered test for
it.
The tests read pixels back off a device rather than checking that the two
halves agree with each other. What they pin down is what is silent when wrong:
the y flip between mask space and clip space, which a centred stroke would not
notice; a bounding box not grown by the radius, which makes a tap draw nothing
at all; an aspect ratio ignored, which makes a dab an ellipse on any frame that
is not square; a stroke doubling back and building up; and an erase that lost
its place in the order and put back paint the user had taken off.
Not done here: the interaction. The canvas needs to begin, extend and end a
stroke on the active layer, and `DevelopSession::rasterise_masks` still returns
early without a segmentation — it takes the proxy size from one, and a brush
needs no model to have run over the photograph first.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Reported as a thumbnail bug; the export had it too, which is the serious
half. You would have exported a photograph missing every local adjustment.
Both called the unmasked `render`, and the failure is silent by
construction: the generated shader always declares the mask binding and
always emits a block per active layer, so binding the empty placeholder
multiplies each of them by zero. No error, no warning, no missing texture —
the adjustments are simply not there. From inside either path there is
nothing to see.
Every path that produces pixels now goes through one helper that binds the
array, and that is the point of it being one helper rather than three
correct call sites. The array is rasterised in source space at proxy size
and sampled through the framing map, so one array serves every output size:
a 256px thumbnail and a 24 MP export bind the same texture.
Three tests, and the first is the fault stated directly — render the same
edit with and without the array and assert they *differ*. If binding it ever
stops mattering, the masks have stopped reaching the shader. The third
checks the masked share of the frame is the same at 32px and 128px, because
"both non-empty" would pass while a mask that scaled wrongly still ruined
every thumbnail.
The watershed hierarchy does not survive a photograph, so local masking
stops depending on it. A layer can now be one recognised object, and the
object's own coverage is the mask.
`Options::watershed` defaults off. It costs ~80 ms plus a full-resolution
readback to produce a ladder that collapses, and paying that on every
photograph buys a control that misleads. Kept switchable rather than
deleted: the passes and the hierarchy are correct in themselves and it is
the merge criterion that fails, which is a change to one function.
Masks now rasterise in **source** space at proxy resolution and are sampled
by the composed shader after the framing map. That fixes a real bug: they
were rasterised in output space, so zooming slid the photograph underneath
a mask that stayed pinned to the viewport, and cropping moved every
adjustment to a different part of the picture. Doing it this way also
leaves the framing map in exactly one place — a second copy in the mask
shader would have been a second thing to keep in step, failing only when
straightened.
A subject is stored as identity, not pixels: the mask is megabytes and is
reproducible by running the same model over the same image, so the sidecar
carries the index, the class and the score, and the session carries the
pixels. The class is there to be checked — if instance 3 comes back a "car"
where it was a "dog", something changed and the layer is stale rather than
silently masking the wrong thing.
The overlay now draws instances and is transparent everywhere else. The
region version covered every pixel and so hid the photograph it was drawn
over; the question it exists to answer is whether an outline follows the
subject, which you can only answer by seeing both.
`examples/local.rs` is the worked example: subject in colour with the rest
monochrome, and the subject lifted out of its background. Run on a 5472x3648
CR2 it finds two people and two cars, and the colour-pop keeps her hat and
hair while the wall and grass behind go grey.
A mask layer is an ordinary develop chain plus a rule about where it
applies. Nothing in the chain knows it is being masked, so every operation
that works globally now works locally and a newly declared op in `ops/`
arrives with local support already done.
The composer emits each layer after the global chain and before the
conversion out of camera space, which is what a photographer means by "and
*then* lift the shadows on her face". Op fragments write to a `c` they
expect to own, so a layer block shadows it and copies the result back out
through a carrier — assigning the outer one from inside is impossible
precisely because it is shadowed. The fused dispatch survives: three global
adjustments and two masked ones remain one shader, one read, one write.
Masks rasterise on the GPU and never exist in CPU memory (ARCH §5.4). That
is the whole reason darktable's brush masks lag, and it is architectural
rather than tuning, so it is not a thing to inherit and fix later.
The rasteriser is a render pass rather than the compute shader it obviously
wants to be, and the format is why: R8Unorm is not a core storage format,
so a compute path has to widen masks to four bytes per pixel — 768 MB
across eight layers of a 24 MP export, against 192 MB at one byte. A colour
attachment takes R8Unorm happily. The array slice comes from the attached
view, so no slot uniform exists to disagree with where the pass writes.
Region masks index a compacted label field rather than the watershed's raw
basin roots, because a root is a sparse index into pixel space and
indexing a per-region array by one would need a table the size of the
image. Changing a selection then costs a few kilobytes, not a re-upload.
Stored as region ids, not as pixels: diffable, mergeable per-field under
FR-NC-9, and cheap in a sidecar. The ids only mean anything alongside the
segmentation that produced them, so each layer carries that signature and
is treated as stale rather than applied when it does not match — a
confidently wrong mask being much worse than an absent one.
Seven device tests render actual frames and read them back. The unit tests
either side check halves that would both pass if the two agreed with each
other and were both wrong; a mask sampled with x and y swapped satisfies
them and fails these.