FR-CULL-3's peaking half. The raw histogram and raw clipping indicators
remain unbuilt -- what exists is a display histogram tagged FR-DSP-7,
counting AdjustPass's 8-bit output, which reports a highlight as gone
precisely where FR-CULL-3 needs it to report the highlight recoverable.
Verified before merge: fmt clean, clippy --workspace --all-targets
-D warnings green, 11 focus GPU tests, 79 baseline dr-gpu tests, 511
dr-ui tests. The cfg(target_os = "android") arm is unverified -- the
host-target clippy never compiled it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
# Conflicts:
# ui/dr-ui/src/lib.rs
# ui/dr-ui/ui/app.slint
FR-CULL-3's focus peaking. One compute dispatch measures local contrast
in WGSL and writes an overlay texture; on desktop it reaches Slint
through the same zero-copy wgpu import the canvas uses, so nothing
per-pixel touches the CPU on the frame path.
With peaking off the cost is zero and structurally so: focus_overlay
opens with `let settings = self.peaking?;` before the frame is touched,
and clearing drops both overlay textures, so no VRAM is held either.
NFR-P14 is met by construction rather than by measurement -- one
dispatch, no second render, no pipeline compile after session open, and
a test asserting allocations stay at 2 over eight frames. The budget
test asserts 50ms at 4K rather than a tight bound, deliberately: a tight
bound fails on a loaded machine and gets deleted, which is worse than a
loose one that still catches the regression that matters.
TD-1 is amended rather than joined by a TD-6: on Android the overlay
rides the readback that already exists there, roughly doubling that
transfer while peaking is on, and TD-1's own "Done when" removes both
because both are the same missing capability.
Verified: cargo fmt clean; clippy --workspace --all-targets -D warnings
green, which also compiles peaking.slint through dr-ui's build.rs; 11
focus GPU tests and 79 baseline dr-gpu tests pass; 511 dr-ui tests pass.
Not verified: the cfg(target_os = "android") arm, which the host-target
clippy never compiled.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Thirteen requirements were surveyed as built but untagged. Eight of them
were: R3, R6, FR-DEV-1, FR-UI-6, FR-NC-6d, NFR-OPS-3, NFR-PORT-2 and
NFR-SEC-3. Each was read against its full text in requirements.md and
against the code before the tag was added, because a tag that is wrong is
worse than an absent one — it turns a visible gap into an invisible one.
The five that were refused, and why, because the reasoning is the part
worth keeping:
R2 carries "(figure TBD)" in its own acceptance criterion and asks for a
stated prefetch margin and cache-hit rate; neither figure exists anywhere
in the tree and neither quantity is measured, while TD-2 and TD-3 both
describe the thumbnail path falling short of it.
R5 asks for three things and the code does one. The display pipeline does
run at viewport resolution, but "only visible tiles are computed" and
"panning recomputes only newly exposed tiles" need a tile scheduler that
does not exist — and frame_budget.rs currently argues for striking tiled
computation from the interactive path rather than building it.
FR-RAW-2 asks for a trait taking a SourceRef, so that a second decoder can
be added without changing callers. What exists is free functions over
&[u8]. That meets the requirement's stated *purpose* — the same decoder
serves a local file, a SAF document and a byte range, which is exactly why
it takes bytes — but there is no trait and no second implementation seam,
so the requirement should probably be amended rather than tagged.
NFR-ARCH-1 asks for named executors with stated thread counts.
architecture.md §7.1 states the table; nothing implements it. Workers are
twenty-odd ad-hoc std::thread::spawn sites, each building its own
one-worker tokio runtime, with no decode pool, no GPU-submit executor and
no I/O pool. The requirement's own text says R4 and NFR-P9 "assert an
outcome with no stated means", and that is still true.
NFR-SEC-4 is satisfied by absence — there is no telemetry — and absence
has no module to tag. A tag would point at nothing.
NFR-OPS-3 was the closest call of the eight taken. The store is single,
separate from the catalog, survives a catalog rebuild and does not sync
between devices; it has no version *field*, deliberately, and
settings.rs argues why and names the condition that would need one. The
substance is met and the reasoning is recorded where it belongs.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Clarity's Gaussian sigma is 1.2% of the frame's shorter edge, so its radius
is a property of the viewport: 52 render pixels at 4K, two separable passes
of 105 taps each over 8.3 M pixels. That measured 33.9 ms — seven times the
entire fused point chain, for one slider — and is docs/technical-debt.md TD-4.
A detail pass may now declare `output_scale`, and clarity's base is computed
on a grid a quarter the size on each axis.
The pass that combines needs the blur *and* the full-resolution colour, and a
colour that has been through a quarter-scale target is no longer full
resolution. So a scaled pass cannot simply join the ping-pong: there are two
chains now. The full-resolution one carries the colour and no scaled pass
touches it; the reduced one carries the base and reaches the combining pass
through a second binding as `reduced_at()`.
The reduce is a dispatch of its own rather than something the first blur half
does on the way past, and that is the whole difference between this and the
strided kernel the module documentation rules out. A stride samples an image
that is not band-limited and aliases high-frequency content down into the
base, which is then subtracted, and arrives in the output as mottling across
smooth gradients. This band-limits first and samples after. What is discarded
is content the base could not represent at any resolution, because a Gaussian
at sigma = 26 px holds nothing above one cycle per 26 px and the quarter-scale
grid carries one per 8 — so the reduced base is not an approximation of the
full-resolution one, it is the same function sampled where it is still
determined.
Which is also why the scale belongs to the band rather than to the stage.
Texture's sigma is a decade finer, so the reduce pass's own box would be wider
than the Gaussian it was prefiltering; texture never reduces. And clarity
steps 4 -> 2 -> 1 as sigma falls, because a quarter of a small sigma is not a
Gaussian either — the case that gives up is the one that was already cheap.
`radius` stays in each pass's own pixels and `ComposedDetail::radius` multiplies
it back up, so 13 reduced pixels at scale 4 still report the 52 render pixels a
tile would have to be grown by. The halo a scheduler sees does not move.
The halo tests pass unchanged, which was TD-4's stated bar; they render at
1024 px and so exercise the reduced path rather than stepping around it. Added
`crossing_the_reduction_threshold_does_not_change_the_picture`, because
nothing yet compared the reduced form against a *less* reduced one — every
other test measures one form against itself. It renders the same edit either
side of the 4 -> 2 step-down and holds the peak excursion to 0.03 stops and
the reach to 2% of the frame.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`request_adapter` with `HighPerformance` returns one adapter and no
second chance. That is right on a healthy machine and wrong on one with
a sick GPU, which is not rare: observed 2026-08-29 on a laptop whose
discrete card had hit an NVRM assertion failure and a fullchip reset.
The driver still advertised it, wgpu dutifully picked it as the highest
performing, and the process died on it — while a working integrated GPU
and a working external card sat unused in the same enumeration. A photo
editor that will not start because the *fastest* GPU is broken, on a
machine holding two that are not, is worse than a slow one.
So: enumerate, order by preference, take the first that yields a device.
The ordering reproduces what `HighPerformance` meant, so a healthy
machine picks what it always picked and pays one enumeration for it. A
CPU adapter sorts last rather than being excluded — software rendering
is a poor experience and a working one.
Which GPU to prefer is now a policy rather than an assumption, because
the fastest is not obviously the right one. A 24 MP frame is ~96 MB of
RGBA and every upload and export readback crosses PCIe on a discrete
card, where an integrated GPU shares memory and crosses nothing — and
does not empty a battery.
Measured before choosing a default, on this machine's Iris Xe against
its RX 5700 XT. The fused colour pass is within 1.5x, which is the
shape shared memory suits. The neighbourhood stage is 5-8x slower, and
that decides it: clarity at 1920x1200 costs 20 ms on the iGPU, over the
budget on its own at the smallest size tested. So `Performance` stays
the default and `Efficiency` is offered rather than chosen
(`DARKROOM_GPU=integrated`).
docs/frame-budget.md carries the table, and says what it does *not*
show: the harness renders from a resident texture and never uploads or
reads back, so the transfer cost an iGPU avoids appears in none of it.
Import, export and the thumbnail sweeps may well go the other way.
What this cannot fix: a GPU sick enough to accept `request_device` and
segfault afterwards, which arrives as a driver crash rather than an
error. It moves the boundary from "the preferred adapter is unusable" to
"unusable and dishonest about it".
Two agents worked in parallel and neither could see this. The frame-budget
instrument matches `ParamKind::Enum { variants }` by value, which was free when
a descriptor was `&'static` and everything in it was borrowed for the life of
the program. Descriptors are owned now — a declaration parsed at run time
cannot hand out a `&'static` — so `variants` is a `Vec` and the arm was moving
out of a shared reference.
Bound by reference instead. The arm only ever reads the length.
The kind of conflict that survives a clean textual merge: git had nothing to
report, and the two changes are only incompatible once they are in the same
tree.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Whitespace only. `cargo fmt --check` is a required step and the FR-DSP-5 test
arrived disagreeing with it — kept as its own commit so it can be skipped
wholesale rather than read for a change that matters.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
FR-DSP-5 has been satisfied for some time and untagged. `Framing::view` shrinks
the sampled region while the render target keeps its size, so a zoom raises the
resolution the pipeline works at rather than magnifying pixels already drawn —
there is no second full-resolution path because the zoom is that path.
Tagging it on that basis alone is what §7 of the display spec warns against:
traceability counts a requirement as covered when a comment names it, and
checks nothing about the code under the tag. So the tag goes on tests instead,
and the tests are built so that removing the behaviour breaks them. Both
failure modes were checked by hand: deleting the view from `visible_rect`
leaves the 1:1 render flat, and dropping only its offset leaves the render
exactly inverted. The assertion message names both, since those are the two
ways this can go wrong and the numbers alone do not say which.
The fixture is one-pixel black-and-white stripes — the highest frequency an
image can hold, and precisely what a proxy discards. A 1024 px source in a
128 px viewport reads source column `8x + 4` for every output column `x`, all
the same parity, so the fit render comes out uniform; that is asserted first,
because a 1:1 render showing detail proves nothing unless the proxy is known to
carry none. What remains is an equality against the source bytes rather than a
claim that something looks sharper.
The third test takes the arbitrary zoom the requirement also names, and pins
`RenderScale` beside the pixels: a zoom that moved the pixels but not the scale
would sharpen at the wrong radius, which stays invisible until somebody
compares a preview against an export.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
FR-DSP-3 states a latency requirement and nothing checked it, which makes it a
wish. This adds the check and the measurements it guards.
`docs/frame-budget.md` is the bench's output with the reading of §2's decision
rule attached. The short version: every point-operation chain at every viewport
size, fit and at 1:1, is inside 16 ms at the 99th percentile — the widest is
4.5 ms of GPU at 4K — so FR-DSP-2 should be rewritten rather than implemented.
The measurement did find a stage that misses the budget, and it is the one §2
predicted: clarity's 52-pixel separable kernel costs 34 ms at 4K. Tiles make
that worse rather than better, since a tiled convolution reads a halo per tile;
the fix `local_contrast` already names for itself is a base computed at reduced
resolution.
The test guards the fused path and says so, at length, rather than quietly
excluding the expensive stage and letting the tag imply otherwise (§7). What it
asserts is exactly the claim the recommendation rests on: one dispatch over a
viewport-sized target, at a full chain, is comfortably inside a frame.
Two things the numbers forced:
- The two cases are one `#[test]`. As two they ran on a thread each, contended
for the same device, and took the 1:1 case from 2.5 ms to 14.9 ms — a
measurement of the harness that would have flickered either side of the
budget forever.
- The CPU half of the frame is judged only in an optimised build. Composition
is real per-frame work on the UI thread and belongs in the budget, but the
workspace builds its own crates at `opt-level = 0` in dev and `cargo test` is
a dev build, so measuring it there measures rustc. The GPU half is asserted
either way.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`docs/display-and-extension.md` §2 fixes a decision rule in advance: if the
99th percentile of a frame sits inside 16 ms, tiled computation is rewritten
as a scheduling concern for export rather than built on the interactive path.
Nothing in the tree could answer that, so the rule had nothing to act on.
This is the instrument. It renders a 60 MP synthetic source through the real
`render_detailed` at three viewport sizes and four chain lengths, fit and
zoomed to 1:1, and reports nearest-rank percentiles rather than means — a
slider drag is judged by its worst frame.
Three things it does that a simpler timer would not:
- It separates the fused pass from the neighbourhood stage. "Every operation
active" mixes one dispatch together with a chain of convolutions, and §2's
question is about the first of those. `point` is every operation that
contributes a fragment to the fused shader; `all` adds the four with
kernels, and M3 times those alone by moving only a detail parameter so
`render_detailed`'s colour reuse skips the fused dispatch. The reuse is
reported rather than assumed — the `colour` column counts fused dispatches
and must be zero for an M3 row to mean what it says.
- It times the CPU half separately. Composition runs per frame in
`DevelopSession::render`, so it is inside the budget whether or not anyone
has looked at it, and if shader assembly were the expensive half then no
tile scheduler could help.
- It builds the "every operation" chain from `EditGraph::capabilities` rather
than from a list, so declaring a new node does not quietly turn that row
into a shorter chain wearing a longer chain's label.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every orientation bug this codebase has had has been the same bug: a turn of
the right size applied in the wrong direction. That failure is worth naming
precisely, because it does not look like one — a quarter turn applied
backwards lands 180 degrees from right, so the result is a plausible
transform of the picture rather than anything obviously broken, and on
landscape frames it is not wrong at all. It was the straighten shear, and it
was the segmentation overlay, and each time it was found by eye rather than
by a test.
The reason it keeps happening is that "rotate 90 degrees clockwise" cannot be
checked by reading it. The reader has to hold in their head which of the two
images is being rotated and which way the y axis runs, and there were four
hand-written copies of the permutation to hold it for: the shader prologue,
its CPU twin, the thumbnail path, and the segmentation.
So nothing added here says clockwise, anticlockwise, horizontal or vertical.
The functions say *which space they take and which space they return* —
`into_shown` and `into_stored`, `source_pixel` and `shown_pixel`,
`into_shown_rect` and `into_stored_rect` — and each takes the dimensions of
the space it reads from, so no caller has to work out which pair it is
holding. `StoredRect` and `ShownRect` are separate types because they are the
same four numbers meaning different things, which is exactly the case where a
mistake is silent: a shown rect measured against stored dimensions produces a
rectangle in the wrong place, not an error.
Underneath there is one permutation. `source_pixel` was already shared by the
prologue and the thumbnails; `source_point` is its normalised twin, written
beside it so the two cannot drift, and everything else is those two read
forwards or backwards. `Orientation::inverse` is the group inverse rather
than `4 - turns`: mirrors apply after the turn, so undoing means undoing them
first, and a mirror seen from the far side of an odd turn is about the other
axis. That is the diagonal-mirror case, tags 5 and 7, and getting it wrong
renders as — again — 180 degrees.
Three call sites lose their own copy: the thumbnail path, `dr-ui`'s
segmentation, and `dr-gpu`'s `local` example. "Upright" now means one thing
across the application rather than one thing per caller.
The gate that matters most is `the_render_and_the_orientation_map_agree`. The
shader prologue and `Orientation` answer the same question by different
routes, and until now nothing checked that they answered it the same way. It
now checks every EXIF tag against every user rotation and mirror on top of
it, because the composition is where the two could agree singly and disagree
together.
The rest earn their place by having caught something. Writing these found two
real errors in this commit's own new code before it ran anywhere: `shown_pixel`
was handed the dimensions of the wrong space and overflowed, and the rect map
turned the wrong way for the diagonal mirrors. A round trip that returns what
went in is the only check worth having here, since every wrong answer is
still a picture.
No behaviour changes. The permutations are the ones that were already being
applied; they are simply applied from one place now.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
"Find subjects" was handed the proxy in the sensor's own orientation, so
every frame shot on a body held sideways reached the model lying on its
side — and a model trained on upright photographs is very bad at those.
Measured end to end on a 22 MP frame of two people and a dog: `person
0.36` and nothing else, against `dog 0.82, person 0.61, person 0.49` for
the same pixels stood up. Nothing failed; the panel simply offered one
poor subject where there were three good ones.
The orientation was never dropped on purpose. The proxy is deliberately
rendered through a *neutral* graph — the detection has to survive an
exposure change, or every slider would invalidate the masks built on it
— and neutral took the file's orientation with it along with everything
else. Landscape frames were unaffected, which is why it stood for as
long as it did.
The turn is `Orientation::source_pixel`, the same function the grid's
thumbnails already go through, so the detector and the thumbnailer now
agree about which way is up rather than holding two opinions. What it is
turned by is `Framing::effective_orientation` — the file's EXIF tag and
the photographer's own rotations composed into one permutation, by the
group law rather than by adding the turns, which is a distinction
`Framing` already had to make and had already tested. Rotating the
picture and pressing the button again therefore does what it looks like
it does.
The proxy stays in sensor space and the masks come back into it. That is
not a detail to be tidied later: the generated shader samples the mask
array at `uv_src`, *after* the framing map, so a mask stored upright
would sit a quarter turn off the subject it was drawn around. That is a
wrong mask rather than a weak one, and nothing announces it. So the
picture is stood up for the model and laid back down for everything
else, and `upright`/`lay_down` are returned as a pair because calling
one and forgetting the other is silent.
Both directions are the one function: `upright` gathers through
`source_pixel` and `lay_down` scatters through it. A quarter turn is a
bijection of the pixel grid, so the round trip is exact — no filter, no
resampling, and no hole to fill — and an inverse written out by hand
would be a second thing to keep in step, whose way of being wrong is a
mask mirrored about the wrong axis, which still looks like a mask.
The orientation joins the confidence and the tiling flag in the
segmentation signature, and for the same reason: turning the photograph
changes what the model recognises, so two runs either side of a rotation
are different instance lists. Two that happened to come out the same
length would otherwise share a signature and a stored layer would be
silently re-indexed from one into the other.
The refine pass had it too — it re-runs the model over a crop rendered
in the same sensor space — so it makes the same turn, and would
otherwise have handed back a worse mask than the one it was asked to
improve, on the subject the photographer had just pointed at.
`dr-gpu`'s `local` example is fixed with it. It exists to be the
shipping path with pictures attached, and a diagnostic that reproduces
the bug it is meant to catch is a trap for whoever reads it next.
Seven tests. The round trip is the identity over all eight EXIF tags on
a non-square asymmetric grid; a turn carries whole pixels rather than
shearing the channels apart; a sideways frame reaches the model
upright; a box comes back in sensor pixels, worked out by hand for the
one turn a portrait frame actually writes; a restored box still reads
low-to-high for every tag, since the rest of the pipeline takes
`x1 - x0` without checking the sign; and the eight tags cannot collapse
into one signature key. The existing composition test now runs against
`effective_orientation` itself, over all 8 x 16 baseline-and-user pairs.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A clone gets the texture right and the level wrong. Dust on a gradient sky
is copied from a patch a little lighter than the hole it fills, and the
repair reads as a disc even though every grain in it is correct — which is
why FR-DEV-8 asks for heal and not only for clone.
Heal adds the membrane: the difference between the two neighbourhoods,
sampled at twenty-four points around the rim and interpolated across the
disc by inverse square distance. Solving the Poisson problem properly is
tens of Jacobi iterations, and an iteration here is a dispatch — sixty
dispatches to remove a dust spot is not a frame budget. The closed form
costs one loop over the rim, no state, and no second pass.
The spec called for mean-value weights; inverse squares are two
transcendentals per sample cheaper and agree wherever the boundary
difference varies smoothly, which is every repair anyone makes. What
decides whether that trade holds is the measurement, so the measurement is
the test: on a ramp steep enough to leave a clone wrong by 38 levels out
of 255, the heal is wrong by 0. docs/spot-removal.md §6.1 records what
shipped and what it would take to go back.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A spot set now composes detail passes of its own, one per round, and they
go ahead of every operation's kernel. That placement is the decision worth
recording: a sharpening pass reads a neighbourhood, so sharpening a dust
mark before removing it smears its edge into pixels the repair's disc does
not cover, and what survives is a faint over-sharpened ring around an
otherwise perfect patch. It also disagrees with ARCH §5.2, which draws
spot removal after clarity — docs/spot-removal.md §5.1 is where that is
argued out.
Every length reaching the shader is in render pixels, converted here where
the framing is in scope. Both the centre and the source go through
`Framing::output_at` — the same map the fused pass applies to every pixel
— so a rotated photograph rotates the offset with no trigonometry, and the
radius is found by mapping a point one radius above the centre and
measuring, rather than by multiplying by a ratio this function has no
business knowing about. The tests turn and crop the frame and expect the
mark to stay gone, which is the property that arrangement buys.
compose_full now takes the spot set, because a photograph with a repair
and no sharpening still has a detail stage: a fused pass that encoded its
own output there would quantise twice and bind to a texture of the wrong
format. compose_detail_for takes the source size for the same kind of
reason — a RenderScale describes the region on screen, and a spot is
stored against the photograph.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Every neighbourhood pass so far has been a convolution, whose whole
description fits in the uniform block because its structure fixes how many
numbers it needs. Spot removal is not that shape: sixty-four repairs and
one repair are the same shader with a different buffer behind it.
So a pass may declare `storage`, which arrives at binding 3 as
`array<vec4<f32>>` with `arrayLength` in scope. The alternative — packing
the list into uniforms — needs a fixed maximum paid for on every frame, a
composer that can emit vec4 fields because a uniform array's stride is 16
whatever it holds, and it gives the next operation that wants a table
nothing to build on.
The property worth having is what stays out of the generated source: the
count is in the buffer, so placing the tenth spot uploads 512 bytes and
reuses the compiled pipeline, exactly as moving a slider does for the
fused pass. `changing_the_list_does_not_recompile` is that, asserted.
One bind group entry rather than two more layouts, and one placeholder
buffer allocated in `new` rather than sixteen bytes per pass per frame —
a zero-length storage buffer cannot be bound, and per-frame allocation is
what this module's documentation exists to refuse.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
An emulsion is a suspension of crystals. Light sensitises some; development
turns a sensitised one opaque, all or nothing. So a patch of film's density
is a *count* of developed grains, and a count of independent yes/no events
has a variance whether or not anyone wanted texture:
mean = D
variance = D * (Dmax - u * D) / N
That expression is the whole feature. It peaks in the middle of the density
range and vanishes at both ends -- clear film has nothing developed to vary,
black film has nothing left to develop -- so grain lives in the midtones as a
consequence rather than as a "midtone bias" slider.
I was wrong earlier that this needs the detail stage. Nothing in it reads a
neighbouring pixel; the only reason to move it was that grain must be fixed in
film space rather than screen space, and that solves itself: N is grains *per
pixel*, so it scales with the film a pixel covers. Zoom out, each pixel
averages more grains, less variance -- correct, with nothing super-sampled and
nothing filtered. It stays in the fused pass.
Grain goes on the density and *before* the dye, which is the physical order
and not cosmetic. Perturbing the finished colour -- what an effect does --
tints highlights wrong, because that noise never passes through the dye.
Crystal habit lives in `rms_granularity`, the number every datasheet
publishes, now a profile field. It measures exactly what differs between a
cubic emulsion and a tabular one: at equal speed, tabular crystals present
more area per unit silver, so the film reads finer. Delta 100 is quoted near 9
where HP5 is near 12, and that gap *is* the habit. Adding a stock whose grain
is its whole reputation is therefore editing one line, not writing a model.
Three things this cost, all of them worth writing down:
- The default granularity is a colour negative's, blue coarsest. Applied to
Tri-X it put *colour* speckle on a black and white photograph. Monochrome
stocks collapse it at parse, where every other per-layer table is already
replicated from the one measured channel.
- Helpers cannot read uniforms. The composer prefixes a uniform with its
operation's id and rewrites references inside a fragment body only;
helpers are shared and deduplicated, so a bare `gn0` names nothing.
`film_lut` already took its size as an argument for this reason, and now
says so.
- The end-to-end test compares the shader against the CPU model, and grain
is stochastic, so that comparison now runs with grain off. Which means a
grain that never left the CPU would look exactly like a passing suite --
hence a second test that grain off is bit-identical, one grain per pixel
moves it, and ten thousand move it less.
Not here, deliberately: no grain slider. The parameters are physical and
`rms_granularity` is the honest place to scale one from, but its range wants
choosing rather than guessing. Nor a film format -- 35 mm is assumed, and
medium format at the same stock is far less grainy per unit of picture.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Nothing here is film simulation. These are lints that fail master today,
under the -D warnings CI runs with, mostly from a toolchain that learned
new ones rather than from anybody's code -- `is_multiple_of` and the
derivable `Default` did not exist as lints when this was written.
They are fixed rather than allowed, and by hand rather than by trusting
`cargo clippy --fix` wholesale: its automatic pass split a derive in two
and left a stray blank line, which is the sort of thing that is correct
and still wrong to commit.
The four that needed a decision rather than a rewrite:
- The distance transform's inner loop writes through its iterator now.
`q` stays, because it is the position the parabola is evaluated at as
well as the index it is written to -- the lint is about the write.
- `to_source` and `to_proto` take `self` by value. Their receiver is
`Copy`, so this is the same machine code and the honest signature.
- The export path's return type is five levels deep and now has a name,
plus a line saying why the `Option` wraps the `Result`: `None` is
cancellation, which is not a failure and has no error to report.
- A test fills a range instead of looping over one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CI runs cargo fmt --check and clippy -D warnings, and this branch had
never been through either. Both would have failed it.
The bulk was the generated colour tables: eight significant figures where
an f32 carries about 7.2, so the eighth is noise that rounds away at
compile time and clippy's excessive_precision says so 109 times over.
Fixed in the generator rather than only in the file, so it stays fixed --
and the file is trimmed in place rather than re-derived, because
regenerating it needs a colour-science stack that has nothing to do with
the defect.
The format! in the composer is mine too, from extracting the rendering
tail: the braces in it were escaped because the text used to live inside a
larger template, and once extracted the escapes are noise and the call
formats nothing.
Also here, and clearly not mine: an unused import and a shadowed binding
in dr-gpu, and an unused import in a test. They are pre-existing --
clippy has been failing on master before this branch existed, on lints
like is_multiple_of that arrived with a toolchain rather than with
anyone's code. Fixed because CI cannot go green around them, and called
out because a merge commit is a bad place to quietly edit someone else's
crate.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The stock model rendered correctly and nothing could ask for it. This is
the picker, and the sidecar key that makes the choice outlive the session.
How the choice persists was the open question, and the answer was already
written down twice in sidecar.rs: `rating` is a top-level key "because a
rating is not an edit", and `masks` are one "because a layer is not a
scalar". A stock is that kind of thing -- a choice of material, not a
number a slider moves -- so it is a top-level key too.
It stores the **id**, not an index. Stocks are files that users add, so an
index would mean installing a profile silently changed which film every
existing photograph had been developed on. A name this build has no
profile for still round-trips untouched, because the alternative is that
syncing to an older phone quietly un-develops the picture.
Only the names travel. Turning one back into tables needs the profile
database, which dr-pipeline deliberately does not link, so `Version::apply`
clears the film and the session re-bakes -- after the parameters, because
the bake reads the film's own exposure sliders and the print balance is
solved against them. That is also why moving those sliders rebuilds the
lookup where no other control in the panel does: an enlarger's filtration
depends on how the negative was exposed.
The panel keeps its rule. It still names no operation and still generates
every control from a declared parameter kind; the stock gets a bespoke
control beside those, exactly as the mask stack does, and for the same
reason. The film's exposure and print exposure arrive as ordinary
generated sliders.
Two defaults worth stating. Picking a colour negative prints it, because
an unprinted one is an orange strip and offering that as the first thing
somebody sees after choosing Portra reads as a bug rather than as a
choice -- the toggle is there for anyone who wants the scan. And a paste
carries no film: a preset is a parameter map, and a stock is not a
parameter, so pasting one would paste a choice the clipboard never took.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The stock model landed in dr-film with no way to see it. This is the
pipeline node, the two texture bindings it reads, and the end-to-end test
that proves the shader agrees with the model.
The design point is that a film simulation is not an adjustment. Every
other node changes a picture; this one makes it. A stock's characteristic
curve does the camera profile's base curve's job -- from measurements
rather than from a curve somebody drew -- so running both renders the
scene twice: the camera's rendering, and then a film's rendering of that.
It looks like neither, and it reads as a colour-management bug with no
colour-management bug to find.
So `Operation::renders` is new. A node declaring it takes camera RGB and
hands back linear sRGB, and the composer emits neither the base curve nor
the conversion out of camera space. Both halves move together, and the
composer keeps them as one string precisely so that getting half of it
right is impossible.
The tables are not parameters, for the reason vignetting's coefficients
are not: they are measurements. dr-pipeline declares the layout as a plain
struct and keeps its no-dependency property; the two crates share no types
on purpose. `EditGraph::set_film_tables` offers them to every node rather
than to the one that wants them, because knowing which concrete type is
which is what the graph is organised not to know.
Bindings 4 and 5 follow the masks precedent: declared unconditionally so
one bind group layout serves every generated shader, bound to 1x1
placeholders when no stock is loaded. Both are interpolated by hand with
textureLoad -- this pipeline binds no sampler, and adding one for two
lookups would cost a binding in every shader. Uploads are keyed on content
so an unchanged stock does not push half a megabyte across the bus per
frame.
The end-to-end test earned its place immediately: it found the density
lookup being filled z-fastest while a 3D texture upload wants x-fastest,
so the red and blue axes were transposed. Green matched exactly, which is
what that bug looks like -- a plausible photograph of the wrong colour,
and one that every unit test on either side of the seam passes. dr-film
now pins the layout in a test that needs no device, and states it where
the field is declared.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`cargo fmt --check` is a required step and had drifted across 45 files. Most of
it arrived this week: several operations were written in parallel worktrees and
merged by hand, and a hand-merge resolves conflicts without ever running the
formatter over the result.
No behaviour changes — this is `cargo fmt --all` and nothing else, kept as its
own commit so the next reader can skip it wholesale rather than search it for
one that matters.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`texture_contributes_nothing_where_its_scale_does_not_exist` asserted that the
detail chain composed *nothing* when texture's kernel rounded away, and its
comment recorded that empty chain as a gap: the fused pass had already decided
to hand on linear working values, so an empty chain left the output transform
undone and the render was rejected. It said fixing it meant composing both
halves together, at the composition boundary rather than in that file.
Noise reduction closed it there in the same round, by emitting a bodyless
`detail/resolve` pass for exactly this case. So the assertion was describing a
defect that no longer exists, and failing because the defect was fixed.
Now asserts the property it was always about — texture contributes no kernel,
`radius() == 0` — while the chain carries the one pass that finishes the
render. Two agents working in parallel each saw one half of this; it is only
visible with both merged.
Sharpening, noise reduction and clarity were written in parallel and each
rewrote the same two tests, which had counted one fused block per operation —
true only while every operation was a point function.
Kept the exclusive-or formulation: each operation must reach exactly one of
the two stages. A count cannot tell "moved to the detail stage" from
"vanished from both", and that ambiguity is what broke these tests three
times over.
The merge left two fragments of the versions it replaced — a loop over a set
that no longer exists, and the tail of an assertion whose head was gone.
The loop is not restored: `point ^ neighbourhood` already asserts per
operation what it checked over the set. The assertion is, because it catches
a different fault from the exclusive-or — a block in the shader that nothing
in the chain asked for, rather than an operation in the wrong stage.
Both tests composed only the fused half and rendered it through the plain
path. That was correct while every operation was a point operation; with a
kernel in the chain the fused pass stops short of the output transform, so
the render was rejected and the operation-block count was one too high.
Compose both halves and dispatch them together, and assert that each
operation reaches exactly one of the two stages rather than counting blocks
- so the next kernel added extends the coverage instead of breaking it.
every_operation_generates_compilable_wgsl and the_whole_chain_at_once_compiles
both rendered through the fused half only. A neighbourhood operation
contributes no fused fragment, so its kernels went uncompiled — and once one
is active the fused pass hands on linear working values, which plain render
refuses. Both now compose both halves from the one graph, and the fused
block count excludes the operations the detail chain names.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A test that cannot be run cannot be checked by running it, and a number
copied out of a test run agrees with whatever the code did on the day.
Each asserted kernel width, tolerance and overshoot bound now carries the
arithmetic that produces it -- the shorter edge, the sigma, the
truncation at two sigmas, and where the rounding falls -- so a reader can
verify the expectation against the recipe without a GPU or a compiler.
Also records the two places where a bound is a bound and not a
measurement: the tolerance in the frame-fraction test is exactly what
rounding a kernel to a whole pixel costs on the smallest frame it uses,
and the halo test's floor and ceiling bracket a peak derived from the
step, the soft limit and the midtone taper rather than from a run.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`the_whole_chain_at_once_compiles` counted one `---- ` block per
operation in the chain. That was true while every operation was a point
operation, and stopped being true the moment a neighbourhood one existed:
clarity and texture are active in that test, and still emit no fused
block, because `compose_full` filters them out and the detail stage
dispatches them separately.
Counted now by asking each operation whether it has a detail stage --
the same question the composer's own filter asks -- rather than by
subtracting a number someone has to remember to update. Capture
sharpening and noise reduction are covered by this without another edit.
The render at the end is now `render_detailed`, which is not a
concession but the stronger test: with a neighbourhood operation active
the fused pass hands on linear working values and the last detail pass
performs the output transform, so rendering the fused half alone is the
mismatch `render_detailed` exists to reject -- and the detail passes are
generated WGSL with uniform blocks of their own, which is exactly what
"everything at once" is here to collide. It renders at 512 rather than
32 because a compositional radius is a fraction of the frame, and on a
32-pixel target every detail kernel rounds away to nothing.
Also records, in `texture_contributes_nothing_where_its_scale_does_not_exist`,
the seam this uncovered: an active detail operation whose kernel rounds
away composes an empty chain while the fused pass has already been
composed to hand on linear values, and nothing can then encode the
result. That test now asserts the property on the composed chain instead
of driving the unrenderable configuration.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A colour square wave built from equal, opposite swings of red and blue is
not a pure colour pattern: Rec. 709 weights them 0.2126 and 0.0722, so it
carries a luminance square wave of about a seventh of the swing underneath.
Correlating the raw red channel therefore reads a constant floor that the
chroma filter is not meant to remove, which compressed every ratio towards
one - enough that the resolution test could no longer tell a correct kernel
from one twice the size. Correlate the colour difference instead, and write
the derivation of each expected value into the test.
Capture sharpening as a two-pass unsharp mask in the detail stage: blur
along x, then along y, each pass applying a one-dimensional high-pass to
luminance so the composite preserves a flat field exactly and matches the
textbook kernel on any locally one-dimensional edge.
The radius is stated in source pixels and converted once per render, so a
radius tuned on a fit view is the radius the exported file gets. Below one
render pixel the operation declines to draw rather than showing sharpening
the file will not contain, and emits a single pass-through that still
carries the output transform.
The develop session now renders through render_detailed, which is what
lets an active neighbourhood operation reach the screen at all.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Checkpoint committed by the coordinator, not by the authoring agent: the
session hit its API limit mid-task and left this work uncommitted. Committed
so it survives, NOT because it is finished - expect failing tests and
half-applied changes. The agent resumes from here.
Checkpoint committed by the coordinator, not by the authoring agent: the
session hit its API limit mid-task and left this work uncommitted. Committed
so it survives, NOT because it is finished - expect failing tests and
half-applied changes. The agent resumes from here.
Checkpoint committed by the coordinator, not by the authoring agent: the
session hit its API limit mid-task and left this work uncommitted. Committed
so it survives, NOT because it is finished - expect failing tests and
half-applied changes. The agent resumes from here.
The declaration in ops/tone_curve.yaml has claimed per-channel curves since
it was written — it is the justification for the operation carrying both
`tone` and `colour`. Only the master curve existed. This is the other three.
The master runs first and the channels grade its result. Both orders are
real images and they differ visibly, so the choice is made and written down
rather than left to the loop: a point placed on the blue curve should act on
the tone the photographer can see, which is what the master has already
produced. The other order anchors the grade to tones the master is about to
move, so adjusting contrast slides a warm shadow up into the midtones.
Every id that existed before today is spelled exactly as it was. The master
curve keeps `p2_y` and the new curves take `r_`, `g_` and `b_` prefixes, so
a sidecar written when there was one curve loads, means what it meant, and
renders the same shader — asserted on the generated source, not on the
parameter values. Nothing needed a version check because nothing was
renamed.
Each curve reaches the shader only when it has been moved off the diagonal,
so an S-curve and no colour work generates what it generated when this
operation held ten parameters instead of forty, down to the uniform names.
The monotonicity guarantee is enforced per curve: a coincident pair on blue
divides by zero exactly as thoroughly as one on the master.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The unit tests either side of the base curve check halves — that the shipped
database parses and lifts its midtones, and that the generated WGSL evaluates
a curve in the right place. Neither would notice if the two agreed with each
other and both were wrong: a curve packed into the wrong uniform slots, or a
flag read from the wrong component, satisfies both and renders nothing.
So render real pixels. A flat frame through a neutral edit, once with the
Canon EOS 6D's curve looked up by name from the YAML and once with the
identity, asserting what a base curve is actually for — midtones lifted,
black still black, white still white, monotone the whole way — plus the
number an unprofiled body must still produce, so "never worse than today"
is a value rather than a promise.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The fused pass hands a fragment a colour and no coordinate. That is what buys
one dispatch for a whole edit, and it is also a wall: sharpening, noise
reduction, clarity, texture, dehaze and spot removal are each defined by what
the neighbours are doing, and FR-DEV-3 and FR-DEV-8 ask for all six. None of
them could be written at any price.
So there is now a detail stage. An operation implements `Operation` for its
parameters exactly as before — the panel, the sidecar, the history and the
presets all work unchanged — and additionally returns `Affects::Detail` and a
`DetailStage` yielding one pass per dispatch. `Affects` grows the third variant
`docs/requirements.md:250` designed and nothing had cut.
Where the stage sits is a colour-science decision, not an arrangement of
convenience. It runs after every point operation and every mask layer, so an
amount chosen against a tone curve survives the curve moving; in linear sRGB
after the camera matrix, because camera RGB has no luminance to sharpen
against; and before the output transform and the clip, because FR-DEV-2 allows
one quantisation and a highlight clipped before a convolution grows a dark
ring. The fused pass therefore ends one of two ways, and when a detail stage
follows it hands on unclipped f16 and the last detail pass encodes.
At render resolution rather than on the source, which is the whole of FR-DSP-1:
a pass before the framing prologue would cost 24 MP to draw a 2 MP preview.
`RenderScale` is what makes that survivable — a radius is stored as a fraction
of the frame's shorter edge, exactly as a mask feather already is, or as a
count of source pixels, and converted per render. It also reports when a radius
is smaller than a proxy pixel rather than drawing a plausible lie; zooming to
1:1 makes the preview exact with no second path.
`Invalidation` gives FR-DEV-3d something to mean. Moving a detail parameter
leaves the colour key alone, so `AdjustPass` keeps the linear intermediate and
skips the fused dispatch: dragging a sharpening slider costs a convolution.
Moving exposure does re-run the detail passes, because they read what the
colour pass wrote, and there is no arrangement of keys that avoids it while
keeping sharpening after tone.
Validated by a separable box blur that is not a develop operation, behind the
`detail-probe` feature and absent from a shipping build. An abstraction with no
consumer is a guess; a box blur's answer is known in closed form, so the tests
assert every byte of the ramp rather than that the edge got softer.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Colour came from whichever matrix rawler happened to key `D65`, the second
one was discarded, and the rendering was left linear. That is the dcraw
default, and FR-DEV-3e names it as the reason people abandon a converter in
the first hour: correct in the abstract, flat and poor on skin in practice.
The decoder now builds a camera profile.
- `ColorMatrix1/2` and `CalibrationIlluminant1/2`. rawler surfaces these as
an illuminant-keyed map — for DNGs from the tags, and for native formats
from its own camera database — so a Canon CR2 arrives with a tungsten
matrix and a daylight matrix exactly as an Adobe DNG of the same frame
would. Dual-illuminant support is therefore not a DNG feature here.
- `ForwardMatrix1/2`, read straight from the root IFD, because rawler parses
them and never surfaces them. Where a file carries both, they replace the
inverted colour matrix: the same relationship measured in the direction
rendering actually wants, rather than an inversion that amplifies the
measurement error exactly where skin lives.
- `AsShotNeutral`, used to estimate what the scene was lit by and to
interpolate between the two calibrations in mireds. The estimate is
circular — the temperature needs a matrix and the matrix needs the
temperature — so it is a fixed point, three rounds, as Adobe's SDK does it.
Bodies calibrated at neither D65 nor A stopped rendering uncalibrated as a
side effect: a Phase One IQ3 carries D55 and D75 and used to get no matrix
at all.
And a base curve, applied per channel in camera RGB between the last
adjustment and the conversion out of camera space — a toe, a steep midtone
and a shoulder, which is the difference between a photograph and a scan of
one. It is not an edit: no slider, nothing in the sidecar, because it
belongs to the body rather than to anything anyone decided, and a sidecar is
shared between bodies. It is not a develop node either, and `ops/README.md`
now records why. It evaluates on the tone curve's own spline rather than a
second copy, so a profile author placing a control point and a photographer
dragging one mean the same thing by it.
The curves are data. `core/dr-decode/profiles/base_curves.yaml` ships inside
the binary as a floor and is superseded by any copy on disk carrying a
higher `version:`, so a body can be added and distributed without a release
— and, under the GPL, contributed. The comparison runs both ways: a stale
pack cannot hold an upgraded binary back at last year's rendering.
Canon EOS 6D and R6, Nikon Z 6 and D750, Sony A7 III and Fujifilm X-T3 ship
with their own curves. Every other body gets a conservative default, which
is much closer to right than the identity is for any of them. A JPEG gets
none — it has already been rendered once, by the camera.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A linear mask at 45° was not at 45°, and a radial with equal radii was an
ellipse. Both on every photograph that is not square, which is all of them.
The geometry is stored in normalised coordinates so that a mask survives a
crop, a zoom and an export at another size — that part was right. What was
wrong is that a *distance* was being measured in those coordinates too, and a
fraction of the width is not the same length as a fraction of the height. So
`dot(uv - centre, axis)` measured the ramp in a space one of whose axes is
squashed against the other by the aspect ratio, and the iso-lines came out
sheared: on a 3:2 frame a ramp asked for at 45° arrives at about 34°.
Nothing announces it. The stored numbers are exactly what was written, the
shader is doing exactly what it says, and the only place the fault exists is
between the photographer's intent and the picture. It has been invisible so far
because there is no way yet to place a gradient by eye — the handles that make
it visible are what turned it up.
So distances and angles move into the frame's own isotropic units: y spans
`0..1` and x spans `0..aspect`, which makes a circle round and 45° a real
diagonal. The centre stays a plain fraction of each axis, because it is a point
and a point has no such problem — and because that is the space a click arrives
in. `frame_delta` is the one conversion and must stay the only one; the mask
array's own dimensions carry the aspect, so it costs no uniform.
The sidecar format does not change. What changes is what the numbers mean, and
the only geometry in the wild is a default that has never been movable.
The two tests are at 96×64 rather than square, which is the whole point: on a
square target this bug cannot be reproduced, and every existing mask test was
square. Both fail without the conversion — the radial reaching 28px sideways
where it reaches 19px down, and the diagonal landing on the wrong side of the
line it is supposed to lie along.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The last line of FR-DEV-3, and the mask ARCH §5.4 was written for. darktable
rasterises drawn masks on the CPU and users call the result unworkable; the
architecture's answer is that a stroke arrives as *parameters* and the device
draws it. This is that, from the model through the sidecar to the pixels — but
not the finger: the canvas is somebody else's change, and this leaves it a
seam rather than reaching into it.
**A stroke is a swept disc along a polyline**, plus erase, radius, hardness and
flow. `MaskSource::Brush` holds an ordered list of them, and the order is the
mask: an erase after an add takes it away and the same pair reversed does not.
Nothing about it is pixels, which is what makes a mask that costs a line of
text, diffs by the gesture, and survives a crop, a straighten and an export at
any size — the properties a stored raster has none of, and the same argument
the region ids were chosen for.
Two things keep the point count honest. While the finger is down, a position
closer to the last than an eighth of the radius is dropped: a touch screen
reports 120 a second, so a finger held still for five seconds is six hundred
points in the same place, and simplification would only remove them once the
gesture had ended — after every frame in between had drawn all of them. When
it ends, Douglas–Peucker at an eighth of the radius removes what a disc that
wide cannot express: a swept circle moved by r/8 moves its own edge by r/8,
which is inside the soft part of any brush. Coordinates snap to a
ten-thousandth of the frame on the way in *and* are written at that precision,
so a round trip is exact rather than nearly exact — a file that drifts in the
sixth decimal every save is a per-field merge conflict a day, over nothing.
**Cost is why the strokes are not drawn by the full-screen triangle the other
masks use.** A swept disc is the minimum distance to any of its segments, so a
stroke over the whole frame costs `pixels × segments` and both terms grow
together — the quadratic that is darktable's problem moved onto the GPU rather
than solved. Each stroke is instead drawn over its own bounding box, grown by
the radius, so the rasteriser never invokes the shader for a pixel the stroke
cannot reach: `area(box) × segments`, which for a dab or a swipe is a small
fraction of the frame. A gesture past 256 points continues as a second stroke
for the same reason, since a shorter stroke has a smaller box.
Add and erase are `dst + a(1 - dst)` and `dst(1 - a)`, which are exactly a
source-over and a one-minus-source blend — so they are blend state, not
arithmetic, and no pass ever reads the slice it is writing. That is what
permits one draw per stroke at all. Within a stroke the coverage is the
*minimum* distance over its segments rather than a sum: a path that crosses
itself must not build up where it did, or every circle and every scribble
would be blotchy wherever consecutive dabs overlap, which is everywhere.
Not a distance field, deliberately. `dr-segment`'s transform documents the two
conditions that make CPU work right there — once per mask edit, over input
already CPU-side — and a stroke fails both: it changes while the finger moves,
and its input is a handful of coordinates that never needed to be pixels. It
also needs no transform, because the distance to a swept disc is closed form.
A stroke is the one mask whose distance field is known without computing one.
An unpainted brush layer is inactive rather than empty, which is not an
optimisation: `invert` turns empty into everything, so a layer created with
invert already set would apply its adjustment to the whole photograph before a
single stroke was made. That is the loud, confident kind of wrong this codebase
refuses everywhere else a mask can go missing, and there is a rendered test for
it.
The tests read pixels back off a device rather than checking that the two
halves agree with each other. What they pin down is what is silent when wrong:
the y flip between mask space and clip space, which a centred stroke would not
notice; a bounding box not grown by the radius, which makes a tap draw nothing
at all; an aspect ratio ignored, which makes a dab an ellipse on any frame that
is not square; a stroke doubling back and building up; and an erase that lost
its place in the order and put back paint the user had taken off.
Not done here: the interaction. The canvas needs to begin, extend and end a
stroke on the active layer, and `DevelopSession::rasterise_masks` still returns
early without a segmentation — it takes the proxy size from one, and a brush
needs no model to have run over the photograph first.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Reported as a thumbnail bug; the export had it too, which is the serious
half. You would have exported a photograph missing every local adjustment.
Both called the unmasked `render`, and the failure is silent by
construction: the generated shader always declares the mask binding and
always emits a block per active layer, so binding the empty placeholder
multiplies each of them by zero. No error, no warning, no missing texture —
the adjustments are simply not there. From inside either path there is
nothing to see.
Every path that produces pixels now goes through one helper that binds the
array, and that is the point of it being one helper rather than three
correct call sites. The array is rasterised in source space at proxy size
and sampled through the framing map, so one array serves every output size:
a 256px thumbnail and a 24 MP export bind the same texture.
Three tests, and the first is the fault stated directly — render the same
edit with and without the array and assert they *differ*. If binding it ever
stops mattering, the masks have stopped reaching the shader. The third
checks the masked share of the frame is the same at 32px and 128px, because
"both non-empty" would pass while a mask that scaled wrongly still ruined
every thumbnail.
Feathering, growing, shrinking, closing and opening are the same number
read differently. With the signed distance from the boundary in hand,
dilation is the set where d >= -r, erosion where d >= +r, and a feather of
any shape is a function of d. So the field is computed once and the
controls are arithmetic on it.
The **field** is what reaches the GPU, not a finished alpha, and that is
the point: growing a mask or changing its falloff then costs a uniform
upload and no recomputation, which is what makes them live controls rather
than ones that stall on every drag. Only closing and opening rebuild,
because after the first threshold the shape has changed and the old
distances describe the old one.
Exact Euclidean, via Felzenszwalb's separable transform — not a chamfer
approximation, which leaves a mask visibly octagonal once grown more than
a few pixels. A test asserts the diagonal is √2 rather than 1 or 2.
It runs on the CPU, which ARCH §5.4 forbids for masks. The rule is about
brush lag — a stroke rasterised per frame — and this is a different
operation: once per mask edit, on input the model already produced here,
producing a field the GPU then samples for free. What it buys is exact
determinism, which matters because masks reach the sidecar as indices and a
field that varied by vendor would mean a mask meaning one thing on the
desktop and another on the phone.
The half-pixel in `signed_distance` is not a detail, and a test caught it.
Measuring to the nearest opposite pixel *centre* puts the smallest
magnitude at 1 either side, so the boundary is nowhere and **eroding by
less than a pixel removes nothing**. A control whose first notch does
nothing is a broken control. Half a pixel off each side puts the boundary
where it physically is, and eroding by 1 takes exactly the outermost ring.
Every falloff curve is 0.5 at the boundary by construction, asserted for
all five: changing the curve should change how the transition looks and
never where it sits.