Commit Graph
25 Commits
Author SHA1 Message Date
dtourolle e98b1def98 Keep blown highlights grey in a panorama
A clipped photosite reaches the merge as (1, 1, 1), which the as-shot
balance turns magenta. The preview balanced it with no highlight rule at
all, so every blown cloud was pink on the alignment page. The DNG had the
quieter form of the same fault: a frame's gain below one moved a blown
sample off the white level, and a feather mixed it into a neighbour's
real sky, so the develop's own desaturation no longer recognised it.

The merge shader and the preview now write a blown sample, before the
gain, as the camera value the composite's balance calls grey — the
develop pipeline's neutral, fading in from CLIP_ONSET.
2026-09-27 16:29:28 -04:00
dtourolle c50d96e949 Repair hot and dead photosites before the demosaic
A hot photosite went into the demosaic as it was read, and came out as a
coloured cross three pixels wide that nothing later could take back
out. Night and long exposures showed them; the defect-map reader added
for FR-RAW-3 was never wired in, and a CR2 carries no map anyway.

A pass over the mosaic now runs ahead of the demosaic, into a second
buffer. A photosite is hot when it reads more than twice every
same-colour photosite in its 5x5 window plus 2% of the range, and more
than twice each of its eight immediate neighbours of any colour. The
second half keeps stars and glints: real light reaches the sensor
through a lens and an anti-aliasing filter and lights a patch, so the
photosites beside it are lit too, where a hot photosite's are dark. It
is replaced by its brightest same-colour neighbour, which invents
nothing. Dead photosites are the mirror case, judged only where the
neighbourhood is above 5%, so shadow noise clipped at black is left
alone.

The colour of each photosite comes from a 6x6 sensor-anchored tile, so
Bayer and X-Trans share the pass. Export and every other path that
demosaics get it too, and there is no setting: the repair only fires
where a single photosite disagrees with everything around it.

Cost, warm, on a Canon 6D frame (RTX 3050): 91-99 ms to demosaic
before, 94-98 ms after; the extra pass is inside the run-to-run noise.

Tests render a frame with and without the defect and compare the
finished pixels. Without the repair a hot photosite showed by 230 and a
dead one by 168; with it neither shows, and a 3x3 highlight at white
survives.
2026-09-27 06:20:23 -04:00
dtourolle 84fade99ec Put the developer docs under docs/dev and index the folder for users first
docs/ had 26 developer documents flat beside the manual, and the two
audiences are very differently sized: most readers want the manual and
the gesture reference, a few want the register, the designs and the
measurements. The manual and gestures.md stay at the top; everything for
someone changing the code moves to docs/dev/, and the two documents that
name their own successors — the v0.1 milestone and the UI-refinement plan
— go to docs/dev/archive/ rather than being deleted, since both are still
cited. docs/README.md is the index, users first.

Every reference follows: code comments, Cargo manifests, the workflows,
the pre-commit hook, the bench and traceability tools (which locate the
repo root by docs/dev/requirements.md now), packaging, the Docker READMEs,
CLAUDE.md, CONTRIBUTING.md and the README. The matrix links one level
deeper and is regenerated. Links out of the moved documents into the tree
gain a level; a link checker over every Markdown file finds none broken.
2026-09-20 21:16:03 +02:00
dtourolle c6cfb2a02a Put the -1 on the greens along the chroma axis, not across it
The Malvar "R at green in R row" kernel weights the two greens two
sites away along the row at -1 and the pair up and down the column at
+1/2. The shader had the two swapped, in the comment as well as the
code, so the transcription checked against itself. Both sum to zero
and reconstruct a flat patch exactly, which is all the tests fed it.

On an edge the correction at green sites is half strength and the
false colour doubles: 0.375 against 0.19 on a grey step, and a
blue/yellow zipper around every clipped highlight at 1:1. The other
three kernels and the CFA tables were right.

A grey vertical step now runs through the pass; the transposed kernel
fails it at 0.375.
2026-09-19 22:05:39 +02:00
dtourolle 2fd7690b6f Mark a pixel the lens correction pushed off the sensor with alpha 0 in the camera-space tap
The fused shader stored black with alpha 1 for a pixel whose source
coordinate left the frame, and the merge's warp averaged it in like any
other: a dark, badly interpolated fringe along every frame's edge, visible
as a seam wherever a frame ended and, later, as the edge the border fill
continued. The display keeps its opaque black; CameraLinear stores alpha 0
and the warp weights each sample by the alpha it interpolated, dropping a
sample that has none.
2026-09-19 20:41:20 +02:00
dtourolle 44ea763c61 dr-gpu: the merge pass — warp, accumulate, resolve, chunk by chunk
merge.wgsl warps one camera-space tile into one output chunk — output
pixel to direction (the projection maths of dr_pano::projection, verbatim),
direction to the frame's camera, camera to source pixel, bilinear by hand
from four textureLoads because rgba32float is not filterable — and adds it
into a storage-buffer accumulator weighted by its distance from the
frame's edge. A resolve pass divides by the weights and packs sixteen-bit
samples at the sensor's scale with a coverage bit.

MergePass::merge drives it: bands of rows, chunks across a band, and for
each chunk only the frames whose footprint meets it, each rendered as the
source rectangle the chunk needs and nothing more. The working set is one
chunk, one tile and one band (FR-MRG-11); the frame textures are the
caller's to cache. Feathered, not seamed; gain a scalar per frame — the
blend quality is panorama.md §10's step 5, after the path writes a file.
2026-09-19 15:24:12 +02:00
dtourolle df741a8a49 Let one mask be built from more than one selection, and paint into it
A mask the model draws arrives approximately right — stopping inside a
shoulder, leaking into the hair — and FR-DEV-3's edge controls move the
*whole* boundary, so no value of feather or dilation fixes two errors that
go opposite ways. What fixes them is a second selection joined to the first,
and a layer that held exactly one source had nowhere to put one. The brush
the core has had all along was reachable from no control in the application.

A layer is now an ordered list of parts. Each names a source and how it
joins the mask before it — added to it, or taken out of it — and carries its
own edge treatment, because a model's soft coverage and a stroke painted
where it stopped short do not want the same feather. Invert and opacity stay
on the layer, where the composed shader already reads them.

The sidecar grows `[part]` blocks and nothing else. A layer of one part
writes exactly the bytes it always did; a mask block with no part blocks
after it reads back as one part; and a stroke, a join or a source this build
cannot read costs that part rather than the layer. So every sidecar in every
library still parses to the edit it always was.

On the device the parts fold into the layer's one slice, so eight layers
still cost eight channels: union is a `max` blend and subtraction is the
erase blend the brush already used. A part is drawn into a scratch texture
before it is joined, and that is not incidental — an erase stroke means a
hole in *that part*, not a hole in the mask, and drawn straight onto the
accumulator it would punch through the subject underneath. A layer of one
part skips all of it and takes the path it always took.

In the interface: a part list under the selected layer with a chip saying
which way each joins, Add and Subtract beside it, a Select/Paint/Erase strip
with the brush's size, hardness and flow, and a drag on the photograph that
paints. Pressing Paint on a mask that cannot hold a stroke joins a part that
can, rather than explaining that a subject is not a brush. A whole stroke is
one step in the history.

The edge controls now shape the part that is selected rather than the layer,
which is the one behaviour change to an existing control: with a correction
selected, the feather slider softens the correction and leaves the model's
mask alone.
2026-09-07 20:00:40 +02:00
dtourolle 68ebf5d78b Let a mask start from a tone or a colour, not only a shape
Every local adjustment began from a shape: painted, drawn with a handle, or
found by a model. So the only way to hold back a sky was to draw a line near
where it ended, and the only way to warm skin was to paint round it — both of
which put the edit's edge where the photographer put a gesture rather than
where the picture changes. A gradient across a treeline halos, and an
adjustment traced round a face stops on the outline of a hand.

MaskSource grows two variants that select by what a pixel *is*. Luminance
carries two bounds on the perceptual tone scale plus a softness; Colour carries
an arc of hue, a range of chroma, and one softness for every edge of both. Five
floats and three, so they diff, sync and merge per field under FR-NC-9 exactly
as a gradient's geometry does — the property a stored raster has none of, and
the reason the model's coverage had to sit beside its source rather than inside
it.

The pixels are the shader's business and nowhere else's. `mask.wgsl` takes the
demosaiced source as a sixth binding and two new modes read it: decode, balance,
pull a clipped photosite back to neutral, apply the camera matrix, then weigh
the band. Nothing crosses to the CPU but the numbers and the matrix, and each
mask texel averages its own footprint in the source, so a band lands on the tone
an area is rather than on whichever texel a proxy grid happened to land on.

The photograph it measures is the one the camera recorded, before this edit. A
band over the edited result would slide out from under the edit as the edit was
made — raising the highlights would change which pixels counted as highlights,
and the slider would chase its own mask.

Feather, falloff and morphology stay off a range layer, which is what
`shapeable` already meant. All three are functions of the signed distance from
a boundary, and a range has no boundary to be at a distance from; its edge is
the softness of its own band, in the band's units. Offering them would be four
controls that move and change nothing.
2026-09-06 19:01:48 +02:00
dtourolle 4b6c110816 Count the sensor's own numbers, so a cull can see headroom the render hides
FR-CULL-3's remaining two bullets. What existed was a *display* histogram
tagged FR-DSP-7: it binds AdjustPass's Rgba8Unorm output, recovers an 8-bit
code value, and counts clipping as `r == 255`. Its own documentation says a
clipped bin means "a highlight that is actually gone rather than one the
transform might still recover", which is the opposite of what a culling
decision needs. FR-CULL-3 asks for the histogram of the sensor data, on the
explicit grounds that a rendered image "systematically lies about what is
recoverable in the raw", and a readout that measures the render cannot answer
that however it is presented.

So this is a second instrument beside the first rather than a setting on it.
Both are true; they are true about different things; the panel offers both
behind a chip row and the words travel with the numbers, because a raw
saturation figure drawn under a heading saying Highlights would be mislabelled
exactly where the difference matters.

**What is reduced over, and what it cost to decide.** ARCH §5.5 specified the
pre-demosaic CFA samples. This reduces over the demosaiced scene-linear
texture instead, and §5.5 is amended to record the choice rather than let the
specification and the code disagree in silence. The texture is camera-native —
unbalanced, unmatrixed, uncurved — and normalised by the sensor's own black
and white levels, so 1.0 is saturation by construction and the distribution
below it is the headroom question with no calibration to carry. Retaining the
CFA samples would mean keeping the packed u32 buffer Demosaicer::run currently
drops: 48 MB at 24 MP, 120 MB at 60 MP, resident per open photograph whether
or not anyone looks at the histogram, on a platform §6.2 exists because memory
is scarce on.

Three things it therefore cannot say, written into the module docs and into
§5.5 rather than left to be discovered: it counts pixels not photosites, so a
saturated site drags its interpolated neighbours up and per-channel clipping
is smeared by about a demosaic kernel; it cannot see above white, because
demosaic.wgsl clamps each photosite at 1.0 for its own good reasons (a Canon
6D reads to 16383 against a declared 15070) so "at saturation" and "a stop
past it" share a bin; and it is measured after the CFA pattern is gone, so it
can name which colour clipped in the reconstructed image but not which
photosite went first.

The axis is stops below saturation, 16 bins per stop over 256 bins — the same
bin count the display reduction uses, so the fold into drawable columns is
shared and a divergence between the two plots would have to be deliberate. A
linear axis spends half its width on the top stop, which is why nobody has
ever drawn a useful linear raw histogram. The fourth series is the brightest
channel rather than luma: these values are unbalanced, so any weighted sum of
them is a number about nothing, and the brightest channel is the one that
saturates first and so the one the headroom question is actually about.

It is a property of the file and not of the render, which has two
consequences. It is computed once per photograph and cached — nothing
downstream of the demosaic can move a count in it — so a cull does not pay the
display histogram's per-frame cost three thousand times. And it describes the
whole frame rather than the visible region, deliberately opposite to
DevelopSession::histogram: a crop changes what is on screen and changes
nothing about what the sensor recorded.

Tags are on the reduction, the type, its constructor and the presentation
arithmetic, each of which has a test that fails if the behaviour goes. The
Slint panel and the push from lib.rs keep their reasoning as prose: nothing
asserts them, and a tag would claim coverage the assertions are not making.
2026-08-29 23:36:21 +02:00
dtourolleandClaude Opus 5 2168cdd1c4 Mark what is in focus, so a frame can be judged without zooming to 100%
FR-CULL-3's focus peaking. One compute dispatch measures local contrast
in WGSL and writes an overlay texture; on desktop it reaches Slint
through the same zero-copy wgpu import the canvas uses, so nothing
per-pixel touches the CPU on the frame path.

With peaking off the cost is zero and structurally so: focus_overlay
opens with `let settings = self.peaking?;` before the frame is touched,
and clearing drops both overlay textures, so no VRAM is held either.

NFR-P14 is met by construction rather than by measurement -- one
dispatch, no second render, no pipeline compile after session open, and
a test asserting allocations stay at 2 over eight frames. The budget
test asserts 50ms at 4K rather than a tight bound, deliberately: a tight
bound fails on a loaded machine and gets deleted, which is worse than a
loose one that still catches the regression that matters.

TD-1 is amended rather than joined by a TD-6: on Android the overlay
rides the readback that already exists there, roughly doubling that
transfer while peaking is on, and TD-1's own "Done when" removes both
because both are the same missing capability.

Verified: cargo fmt clean; clippy --workspace --all-targets -D warnings
green, which also compiles peaking.slint through dr-ui's build.rs; 11
focus GPU tests and 79 baseline dr-gpu tests pass; 511 dr-ui tests pass.
Not verified: the cfg(target_os = "android") arm, which the host-target
clippy never compiled.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 21:43:08 +02:00
dtourolle 5bcd0e0269 Merge branch 'worktree-agent-a22a049c461818dbe' into integration
# Conflicts:
#	core/dr-pipeline/tests/mask_sidecar.rs
2026-08-22 13:23:34 +02:00
dtourolleandClaude Opus 5 c75863c93f Give a gradient the angle it was asked for
A linear mask at 45° was not at 45°, and a radial with equal radii was an
ellipse. Both on every photograph that is not square, which is all of them.

The geometry is stored in normalised coordinates so that a mask survives a
crop, a zoom and an export at another size — that part was right. What was
wrong is that a *distance* was being measured in those coordinates too, and a
fraction of the width is not the same length as a fraction of the height. So
`dot(uv - centre, axis)` measured the ramp in a space one of whose axes is
squashed against the other by the aspect ratio, and the iso-lines came out
sheared: on a 3:2 frame a ramp asked for at 45° arrives at about 34°.

Nothing announces it. The stored numbers are exactly what was written, the
shader is doing exactly what it says, and the only place the fault exists is
between the photographer's intent and the picture. It has been invisible so far
because there is no way yet to place a gradient by eye — the handles that make
it visible are what turned it up.

So distances and angles move into the frame's own isotropic units: y spans
`0..1` and x spans `0..aspect`, which makes a circle round and 45° a real
diagonal. The centre stays a plain fraction of each axis, because it is a point
and a point has no such problem — and because that is the space a click arrives
in. `frame_delta` is the one conversion and must stay the only one; the mask
array's own dimensions carry the aspect, so it costs no uniform.

The sidecar format does not change. What changes is what the numbers mean, and
the only geometry in the wild is a default that has never been movable.

The two tests are at 96×64 rather than square, which is the whole point: on a
square target this bug cannot be reproduced, and every existing mask test was
square. Both fail without the conversion — the radial reaching 28px sideways
where it reaches 19px down, and the diagonal landing on the wrong side of the
line it is supposed to lie along.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 13:19:55 +02:00
dtourolleandClaude Opus 5 c396a22dfd Paint a mask without ever rasterising one on the CPU
The last line of FR-DEV-3, and the mask ARCH §5.4 was written for. darktable
rasterises drawn masks on the CPU and users call the result unworkable; the
architecture's answer is that a stroke arrives as *parameters* and the device
draws it. This is that, from the model through the sidecar to the pixels — but
not the finger: the canvas is somebody else's change, and this leaves it a
seam rather than reaching into it.

**A stroke is a swept disc along a polyline**, plus erase, radius, hardness and
flow. `MaskSource::Brush` holds an ordered list of them, and the order is the
mask: an erase after an add takes it away and the same pair reversed does not.
Nothing about it is pixels, which is what makes a mask that costs a line of
text, diffs by the gesture, and survives a crop, a straighten and an export at
any size — the properties a stored raster has none of, and the same argument
the region ids were chosen for.

Two things keep the point count honest. While the finger is down, a position
closer to the last than an eighth of the radius is dropped: a touch screen
reports 120 a second, so a finger held still for five seconds is six hundred
points in the same place, and simplification would only remove them once the
gesture had ended — after every frame in between had drawn all of them. When
it ends, Douglas–Peucker at an eighth of the radius removes what a disc that
wide cannot express: a swept circle moved by r/8 moves its own edge by r/8,
which is inside the soft part of any brush. Coordinates snap to a
ten-thousandth of the frame on the way in *and* are written at that precision,
so a round trip is exact rather than nearly exact — a file that drifts in the
sixth decimal every save is a per-field merge conflict a day, over nothing.

**Cost is why the strokes are not drawn by the full-screen triangle the other
masks use.** A swept disc is the minimum distance to any of its segments, so a
stroke over the whole frame costs `pixels × segments` and both terms grow
together — the quadratic that is darktable's problem moved onto the GPU rather
than solved. Each stroke is instead drawn over its own bounding box, grown by
the radius, so the rasteriser never invokes the shader for a pixel the stroke
cannot reach: `area(box) × segments`, which for a dab or a swipe is a small
fraction of the frame. A gesture past 256 points continues as a second stroke
for the same reason, since a shorter stroke has a smaller box.

Add and erase are `dst + a(1 - dst)` and `dst(1 - a)`, which are exactly a
source-over and a one-minus-source blend — so they are blend state, not
arithmetic, and no pass ever reads the slice it is writing. That is what
permits one draw per stroke at all. Within a stroke the coverage is the
*minimum* distance over its segments rather than a sum: a path that crosses
itself must not build up where it did, or every circle and every scribble
would be blotchy wherever consecutive dabs overlap, which is everywhere.

Not a distance field, deliberately. `dr-segment`'s transform documents the two
conditions that make CPU work right there — once per mask edit, over input
already CPU-side — and a stroke fails both: it changes while the finger moves,
and its input is a handful of coordinates that never needed to be pixels. It
also needs no transform, because the distance to a swept disc is closed form.
A stroke is the one mask whose distance field is known without computing one.

An unpainted brush layer is inactive rather than empty, which is not an
optimisation: `invert` turns empty into everything, so a layer created with
invert already set would apply its adjustment to the whole photograph before a
single stroke was made. That is the loud, confident kind of wrong this codebase
refuses everywhere else a mask can go missing, and there is a rendered test for
it.

The tests read pixels back off a device rather than checking that the two
halves agree with each other. What they pin down is what is silent when wrong:
the y flip between mask space and clip space, which a centred stroke would not
notice; a bounding box not grown by the radius, which makes a tap draw nothing
at all; an aspect ratio ignored, which makes a dab an ellipse on any frame that
is not square; a stroke doubling back and building up; and an erase that lost
its place in the order and put back paint the user had taken off.

Not done here: the interaction. The canvas needs to begin, extend and end a
stroke on the active layer, and `DevelopSession::rasterise_masks` still returns
early without a segmentation — it takes the proxy size from one, and a brush
needs no model to have run over the photograph first.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 12:37:42 +02:00
dtourolle ec713585a5 Measure the distance to the edge, and get four controls for one transform
Feathering, growing, shrinking, closing and opening are the same number
read differently. With the signed distance from the boundary in hand,
dilation is the set where d >= -r, erosion where d >= +r, and a feather of
any shape is a function of d. So the field is computed once and the
controls are arithmetic on it.

The **field** is what reaches the GPU, not a finished alpha, and that is
the point: growing a mask or changing its falloff then costs a uniform
upload and no recomputation, which is what makes them live controls rather
than ones that stall on every drag. Only closing and opening rebuild,
because after the first threshold the shape has changed and the old
distances describe the old one.

Exact Euclidean, via Felzenszwalb's separable transform — not a chamfer
approximation, which leaves a mask visibly octagonal once grown more than
a few pixels. A test asserts the diagonal is √2 rather than 1 or 2.

It runs on the CPU, which ARCH §5.4 forbids for masks. The rule is about
brush lag — a stroke rasterised per frame — and this is a different
operation: once per mask edit, on input the model already produced here,
producing a field the GPU then samples for free. What it buys is exact
determinism, which matters because masks reach the sidecar as indices and a
field that varied by vendor would mean a mask meaning one thing on the
desktop and another on the phone.

The half-pixel in `signed_distance` is not a detail, and a test caught it.
Measuring to the nearest opposite pixel *centre* puts the smallest
magnitude at 1 either side, so the boundary is nowhere and **eroding by
less than a pixel removes nothing**. A control whose first notch does
nothing is a broken control. Half a pixel off each side puts the boundary
where it physically is, and eroding by 1 takes exactly the outermost ring.

Every falloff curve is 0.5 at the boundary by construction, asserted for
all five: changing the curve should change how the transition looks and
never where it sits.
2026-08-22 08:39:17 +02:00
dtourolle ee10097435 Mask the subject the model found, not the regions underneath it
The watershed hierarchy does not survive a photograph, so local masking
stops depending on it. A layer can now be one recognised object, and the
object's own coverage is the mask.

`Options::watershed` defaults off. It costs ~80 ms plus a full-resolution
readback to produce a ladder that collapses, and paying that on every
photograph buys a control that misleads. Kept switchable rather than
deleted: the passes and the hierarchy are correct in themselves and it is
the merge criterion that fails, which is a change to one function.

Masks now rasterise in **source** space at proxy resolution and are sampled
by the composed shader after the framing map. That fixes a real bug: they
were rasterised in output space, so zooming slid the photograph underneath
a mask that stayed pinned to the viewport, and cropping moved every
adjustment to a different part of the picture. Doing it this way also
leaves the framing map in exactly one place — a second copy in the mask
shader would have been a second thing to keep in step, failing only when
straightened.

A subject is stored as identity, not pixels: the mask is megabytes and is
reproducible by running the same model over the same image, so the sidecar
carries the index, the class and the score, and the session carries the
pixels. The class is there to be checked — if instance 3 comes back a "car"
where it was a "dog", something changed and the layer is stale rather than
silently masking the wrong thing.

The overlay now draws instances and is transparent everywhere else. The
region version covered every pixel and so hid the photograph it was drawn
over; the question it exists to answer is whether an outline follows the
subject, which you can only answer by seeing both.

`examples/local.rs` is the worked example: subject in colour with the rest
monochrome, and the subject lifted out of its background. Run on a 5472x3648
CR2 it finds two people and two cars, and the colour-pop keeps her hat and
hair while the wall and grass behind go grey.
2026-08-22 08:39:17 +02:00
dtourolle c6a846a1f9 Brighten her face without touching the sky behind her
A mask layer is an ordinary develop chain plus a rule about where it
applies. Nothing in the chain knows it is being masked, so every operation
that works globally now works locally and a newly declared op in `ops/`
arrives with local support already done.

The composer emits each layer after the global chain and before the
conversion out of camera space, which is what a photographer means by "and
*then* lift the shadows on her face". Op fragments write to a `c` they
expect to own, so a layer block shadows it and copies the result back out
through a carrier — assigning the outer one from inside is impossible
precisely because it is shadowed. The fused dispatch survives: three global
adjustments and two masked ones remain one shader, one read, one write.

Masks rasterise on the GPU and never exist in CPU memory (ARCH §5.4). That
is the whole reason darktable's brush masks lag, and it is architectural
rather than tuning, so it is not a thing to inherit and fix later.

The rasteriser is a render pass rather than the compute shader it obviously
wants to be, and the format is why: R8Unorm is not a core storage format,
so a compute path has to widen masks to four bytes per pixel — 768 MB
across eight layers of a 24 MP export, against 192 MB at one byte. A colour
attachment takes R8Unorm happily. The array slice comes from the attached
view, so no slot uniform exists to disagree with where the pass writes.

Region masks index a compacted label field rather than the watershed's raw
basin roots, because a root is a sparse index into pixel space and
indexing a per-region array by one would need a table the size of the
image. Changing a selection then costs a few kilobytes, not a re-upload.

Stored as region ids, not as pixels: diffable, mergeable per-field under
FR-NC-9, and cheap in a sidecar. The ids only mean anything alongside the
segmentation that produced them, so each layer carries that signature and
is treated as stale rather than applied when it does not match — a
confidently wrong mask being much worse than an absent one.

Seven device tests render actual frames and read them back. The unit tests
either side check halves that would both pass if the two agreed with each
other and were both wrong; a mask sampled with x and y swapped satisfies
them and fails these.
2026-08-22 08:39:16 +02:00
dtourolleandClaude Opus 5 31e20399c8 Read the true white level, and clamp the sensor stage at both ends
Two corrections to the sensor stage, found while chasing magenta highlights.
Neither is the cause of that — see below — but both are wrong on their own
terms.

`white_level` took the *first* of rawler's per-channel saturation points. On a
Canon 6D that reports 15070 while the data reaches 16383, so every sample
above it was treated as brighter than white. It takes the maximum now.

The normalisation clamped its floor and not its ceiling, so those over-white
samples passed through as values above 1.0. Clamped at both ends.

**This does not fix the pink.** Measured on _MG_8596.CR2, exported and looked
at: the subject renders correctly and only the blown sky is magenta. A fully
clipped pixel is (1,1,1) in raw, the as-shot balance multiplies it to
(1.93, 1.00, 1.68), and the camera matrix turns that into R 2.88, G 0.51,
B 2.03 — red and blue clip at one, green does not, and the result is magenta.
It is correct white balance applied to already-saturated data, which is the
classic highlight-clipping cast and needs highlight desaturation to fix: a
pixel at saturation carries no colour information and must be rendered
neutral, not balanced.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 13:13:17 +02:00
dtourolle 76bb6b2847 Merge branch 'worktree-watershed-plateaux' 2026-08-17 12:25:41 +02:00
dtourolleandClaude Opus 5 0b20436445 Prove the plateau pass does nothing, and stop paying for it
Picks up the lower-completion work a crashed session left mid-debug, with one
failing test and no diagnosis.

The diagnosis is that the pass is a no-op. Not "does not reduce basin count" —
it changes *no pixel's basin at all*, zero of 9216, comparing one plateau
iteration against sixty-four. That assertion is the substance of this commit:
the original test asserted a consequence (fewer basins) which a working pass
need not produce, so it could have been satisfied by weakening it. A no-op
check cannot pass vacuously, and it is what turned an opinion into a fact.

Three candidate causes were tried and none was it. Exact float equality is
genuinely wrong and is fixed regardless — a gradient computed from 8-bit
samples is never exactly equal across a region the eye calls flat, so `==`
never fires and `<` fires everywhere; `LEVEL_EPS` now sits behind all three
comparisons. The test image is not it either: a flat disc, a terraced disc and
a constant-slope ramp all behave the same.

The finding worth keeping is about the domain rather than the code. On a
gradient-magnitude watershed every flat region of the picture is at gradient
zero, the global minimum, and a plateau with no descending exit is a minimum —
one basin already, nothing to resolve. The plateaux lower-completion is defined
for are regions of constant non-zero gradient, which are rarer in a photograph
than F1's phrasing implies. That may be the whole answer, or it may be hiding
a fourth cause; I could not close it.

So `plateau_iterations` defaults to 0. The implementation stays, correct as
far as it goes and costing nothing until someone finishes it; the test stays,
ignored with its reason; docs/segmentation.md §12 records what was ruled out so
the next attempt starts further along than this one did.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 12:25:28 +02:00
dtourolle 9fc8721fa8 Merge branch 'histogram'
# Conflicts:
#	ui/dr-ui/src/lib.rs
2026-08-17 10:00:27 +02:00
dtourolleandClaude Opus 5 0233df4bf2 See what the highlights are doing: a live histogram (FR-DSP-7)
Exposure, blacks and whites were set by eye. Nothing said a highlight had
blown — the canvas shows white where a channel is at 250 and white where it is
at 255, and the difference is the whole question.

**Counted on the GPU, not on the readback.** There is a full frame sitting in
CPU memory on every canvas update right now — `AdjustPass::read_output`, the
bridge spike S1 removes — and walking it would have been thirty lines and no
shader. FR-DSP-7 states the mechanism and not just the feature: "these derive
from a GPU-side reduction into a small buffer. Per-frame CPU readback of image
data is prohibited." A histogram founded on the bridge would be correct today
and deleted by S1, and would meanwhile be the reason the bridge could not go.
What crosses the bus here is 4104 bytes whatever the image size.

The reduction tallies into workgroup memory first and merges once per
workgroup. A photograph is not noise: a clear sky puts tens of thousands of
adjacent pixels in one bin, and contending for that single global atomic
serialises the dispatch.

**On the settled frame only.** `render_now` already knows whether a gesture is
still moving — `draft` is the flag `redraw` derives from `was_coalesced` — so
the dispatch and its transfer happen once when the slider stops rather than on
each of the forty frames a drag emits. Nothing is lost: a histogram flickering
past under a finger is not a reading anyone takes. FR-DSP-7 requires exactly
this, that it not extend the FR-DSP-3 frame budget.

Luma is weighted in 8.8 fixed point — 54, 183, 19, summing to 256 exactly —
rather than in floats. Not thrift: it makes the shader's arithmetic
reproducible bit for bit, which is what lets the test below be an `assert_eq`
against a CPU count rather than a tolerance. ARCH §6.13's line about integer
state, applied where it happens to also be free.

**What the numbers were checked against.** A flat frame must put all 4096
pixels in one bin and one only. A 256-wide ramp must occupy every level with
exactly the same count, which is what catches an off-by-one in the
quantisation — a `floor` where a rounding was needed shifts the whole
photograph one bin left and looks like nothing at all. And a 101x37 frame of
seeded pseudo-random pixels — deliberately not a multiple of the 16x16
workgroup, so the edge tiles run off the image — is compared slot for slot
against a second, obvious CPU implementation. Exact equality, no tolerance.
The CPU version is a deliberate reimplementation rather than shared code: the
bugs worth catching here are ones shared code would commit identically on both
sides.

Above that, the presentation arithmetic is unit-tested headless, because it is
where a wrong answer is invisible. A histogram of the wrong shape looks exactly
as plausible as one of the right shape. So: 64 columns because it divides 256
and an uneven fold draws an even ramp as a comb; the peak excludes the end
columns, or a night scene scaled against its own black spike is a flat line
with no information in it; heights are clamped into the plot; and "0%" is kept
distinct from "<0.1%" and from "—", since an indicator reading "clipped" over
a figure reading "none" is a panel contradicting itself.

Clipping counts a *pixel* with any channel at an extreme, not a channel. Any,
because a blown red has no gradation left in it however much green and blue
still hold — and it is the saturated highlight, the sunset and the red jersey,
that clips first and recovers worst. Per pixel, because counting channels can
report 200% of a frame clipped, and a percentage above 100 is a readout nobody
trusts again.

Two affordances for it, which NFR-A11Y-3 asks for: a bar standing at the end
of the plot the tones are piling against, and a figure saying how much. Either
alone reads.

The panel sits directly under the capture metadata and above every control,
because it is what the controls are judged against. It is hand-built rather
than generated, and ARCH §4.3a is untroubled: a histogram is not an operation
— no parameters, changes nothing, answers a question rather than asking one —
and nothing in it reads a parameter out of a descriptor.

Three plot colours and a neutral luma trace join the palette. That is the
swatch's exception rather than a second one: a per-channel histogram has to
say which channel, and no achromatic treatment distinguishes red from blue, so
the hue is data exactly as the image beside it is. Held well back from full
strength for the reason the theme preamble gives.

The bounded, non-parking map wait moves out of `AdjustPass` into
`readback::await_mapping`, shared with the histogram's transfer. Thirty lines
of load-bearing reasoning about frozen interfaces and lost devices, and two
copies of it would have drifted.

The histogram describes the frame on the canvas, so it is in the output colour
space FR-DSP-7 asks for, and when zoomed it describes the visible region — a
photographer inspecting a highlight at 4x is asking about that highlight. A
device that cannot build the reduction loses the histogram and keeps the
photograph.

Still to do for FR-DSP-7: the pixel colour readout under the cursor.

324 tests pass, clippy and fmt clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 09:55:40 +02:00
dtourolleandClaude Opus 5 1c0994c807 Demosaic a Fujifilm sensor instead of refusing it
Every RAF stopped at the embedded preview, because the demosaicer had
one kernel and it was a Bayer kernel. D11 makes Fujifilm first-class and
FR-RAW-5 asks for it by name, so a hard error there was a promise we had
not kept.

X-Trans is a 6x6 tile, and nothing in the Bayer path survives that: the
missing channels sit at different offsets at all 36 positions, so there
is no fixed kernel to write. The new shader fits a weighted plane
through each channel's samples in a 5x5 window and carries the other two
channels across as the difference between those planes, keeping the
pixel's own measured value untouched. A plane rather than a mean because
the three channels are sampled at different places in the tile: a mean
compares a red taken slightly left of the pixel with a green taken
slightly right of it, and that offset is a colour cast that follows
every gradient in the frame. The fit is done in white-balanced space,
where the constant-colour-difference model it rests on is actually true
of a neutral subject; that alone halves the error at a luminance edge.

Two compromises, both deliberate.

It is not Markesteijn. There are no directional hypotheses and no
homogeneity map, so it does not resolve detail finer than the CFA period
and a hard edge arrives about two pixels wide. It cannot ring — the
output is bounded by the local sample range — so it does not produce the
worms FR-RAW-5 exists to avoid, but the quality that requirement asks
for is still owed.

The tile's phase is guessed rather than known. rawler has each body's
pattern exactly, as a 36-character string, but CfaPattern::XTrans throws
it away before dr-gpu sees the file, and it is not a constant to
hard-code: the bodies in that database start the tile at four different
origins. So the phase is read back out of the pixels, by grouping the 36
per-position means and taking the grouping with the least spread. That
part needs nothing from the scene. Telling red from blue does — shifting
the tile by half a tile turns it into itself with red and blue swapped,
so no geometry can decide it — and the as-shot white balance is what
breaks the tie. A frame that is almost entirely one colour can defeat
that; widening dr-decode to carry the pattern string would retire the
guess altogether.

The tests assert reconstruction, not success: a flat patch comes back
exactly at all six phases tested, and a linear ramp comes back exactly
too, which is the property the plane fit exists for and the one a mean
would fail.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 08:46:28 +02:00
dtourolleandClaude Opus 5 cb1d2be240 Choose the export folder by walking the server, not by typing it
The destination for a Nextcloud export was a text field. Nobody recalls the
exact spelling of a path three levels down, and getting it wrong does not
fail — `create_dir` makes whatever was typed, so a misremembered folder
becomes a new one at the root and the exports are somewhere nobody looks.

So it is picked the way the library root is picked, using the same
`FolderBrowser` model the launch screen drives: up, into, and "use this
folder", confirming the folder currently *shown* rather than one selected in
the list. Same rule in both places, so the phrase means one thing.

The model is shared; the worker is not. `settings_ui::spawn_folder_list` is a
near-twin of the launch screen's, because that one reaches into the
`LaunchController` for its session and reports onto the launch screen's error
line, while this one is handed credentials and writes to the settings page.
Factoring them together needs a function taking both controllers or a trait
implemented twice to abstract two call sites — more machinery than the twenty
lines it saves. What matters is shared already: navigation behaves identically
because both drive the same model.

The callbacks are wired in `lib.rs` rather than in `settings_ui::wire`,
because listing a remote folder needs credentials and the settings page holds
no session on purpose — it is reachable before a library is opened and must
not depend on one existing. With no account the picker says to sign in first,
rather than showing an empty list that reads as a server with no folders.

Details that are decisions rather than accidents: the picker opens at the
library root rather than at whatever half-typed path is in the field, which
would list nothing and look broken. The listing area is a fixed 180px, since a
folder with sixty children would otherwise push the rest of the settings page
off the bottom. "Up" is disabled at the root rather than hidden, so the row
does not jump as the user navigates. A failed listing leaves the picker open
on the folder it was showing — where the user had got to is not something to
discard over a dropped request. And the chosen folder saves immediately like
every other setting on a page that has no Save button.

The poll timer lives on the controller for the reason `LaunchController` keeps
its own there: a `slint::Timer` stops when dropped, so one local to the
function that starts it would be collected before the listing arrived.

Carries in-flight work from a parallel session — a segmentation pass in
dr-gpu, a sidecar cache, and the develop panel's continuing changes.

1020 tests pass, fmt clean. One clippy warning remains and is not mine:
`sidecar_cache::dir` is unused while that work is in progress.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 07:12:01 +02:00
dtourolleandClaude Opus 5 78e3e6b846 Add the develop pipeline: demosaic and seven raw adjustments
Decode through display, on the GPU: black/white normalisation, Bayer
demosaic, camera colour transform, and the first seven adjustment
operations — white balance, exposure, highlights/shadows, blacks/whites,
brilliance, vibrance, saturation.

Composable shaders. Each operation contributes a WGSL fragment rather
than owning a pass, and dr-pipeline fuses the *active* ones into a single
compute shader. One texture read and one write per frame regardless of
how many adjustments are in play, while the operations stay independent
in Rust — adding one is a new file, with no central shader to edit. An
operation at neutral settings contributes no code, no uniform and no
branch. Uniforms are prefixed per operation so two may both declare
`amount`; helpers dedupe by name from a single source of truth.

Pipelines cache on a structure hash covering the op-set and its order but
not the values, so dragging a slider uploads uniforms and reuses the
compiled pipeline. Measured on a 24 MP CR2: 0.60 ms re-render, one
pipeline compiled across ten slider positions.

The UI is generated, not written. EditGraph::capabilities() reports
parameters with their kinds, ranges, defaults and current values; the
panel builds one control per entry chosen by ParamKind. No file in ui/
names an operation, and dr-pipeline has no wgpu dependency, so codegen is
testable without a device (ARCH §6.5a).

Three defects found against real files, each silent:

- rawler 0.7.2's `xyz_to_cam` is all zeros — deprecated and no longer
  populated. The live matrices are in `color_matrix`, keyed by
  illuminant. Reading the old field yields no colour transform at all.
- `cam_to_xyz_normalized()` returns all NaN on any Bayer sensor: it
  divides each of four rows by its own sum, and the unused fourth
  (emerald) row sums to zero. Inverting the 3x3 ourselves avoids it.
  `wb_coeffs[3]` is NaN for the same reason and is normalised at decode.
- As-shot white balance reached the uniform block but no shader read it,
  so the first render of a real CR2 came out violently green. Green
  photosites collect roughly twice the signal of red and blue. Now
  applied unconditionally before any operation, with tests on ordering.

Demosaic is Malvar-He-Cutler rather than bilinear: gradient-corrected
interpolation at one 5x5 neighbourhood per pixel, where bilinear leaves
visible zippering on any high-contrast edge at 1:1. Two of the four
packed CFA constants were wrong on the first attempt, so all four layouts
are asserted to reconstruct the same colour. Crop origins at odd
coordinates re-phase the pattern; without that, red and blue swap.

X-Trans reports GpuError::UnsupportedCfa rather than approximating with
the Bayer path, which would look like a corrupt file.

206 tests, including GPU tests proving every operation and the full
seven-operation chain generate compilable WGSL.

Known gaps: the display path still reads back to the CPU each frame,
which ARCH §6.1 forbids and AC-8 asserts against — it is gated behind the
`readback` feature and waits on spike S1 wiring Slint's texture import.
Curve shapes are a first draft and want tuning against real photographs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-09 11:37:58 +02:00
dtourolle 82a5e21ec6 Initial workspace: GPU context, compute pass, adaptive Slint shell
Establishes the v0.1 foundations on both platforms:

- dr-types: SourceRef (never a filesystem path — Android SAF has none),
  Format, Availability, Validator with ETag quote normalisation
- dr-gpu: wgpu device, compute pass writing a storage texture, resize
- dr-ui: Slint shell with FR-UI-1 adaptive layout, computed in Rust to
  avoid a binding loop
- docker/android: pinned toolchain, verified producing API 28 ARM binaries

Measured the cost of the temporary CPU readback path (dr-gpu bench):
compute is 0.06-0.28ms across sizes while readback is 0.63-7.43ms, so
readback is 90-96% of frame time and scales with area. Recorded in
ARCH §6.1 — this is why spike S1 is the priority.

Mitigations pending S1: reuse the staging buffer, apply at most one
resize per frame, and cap render resolution at 2048 on the long edge.

10 tests passing; core crates cross-compile for aarch64-linux-android.
2026-08-09 07:42:05 +02:00