Commit Graph
26 Commits
Author SHA1 Message Date
dtourolle 8f9e59b9fa Find hot photosites without repairing them, and measure a sensor's aging
The hot-pixel pass could only repair: it returned how many photosites it
changed and threw away which. find_hot_pixels runs the same pass and
returns them as sensor coordinates, leaving the frame alone, so a sensor's
defects can be tracked across frames.

sensor_scan prints each frame's candidates, and with --probe reads a list
of coordinates back out of every frame. Run over 53 6D raws from 2015 to
2026, it found 32 persistent defects, 2 in 2015 and 32 by 2026, and showed
what a defect map has to account for: a frame that does not flag a
photosite proves nothing unless its neighbourhood is dark, and the 6D
hides some of its defects itself above ISO 5000. docs/dev/sensor-health.md
records the findings and the design they argue for.
2026-10-04 02:38:48 -04:00
dtourolle 8ea3c3181a Upload the learned demosaic's result, and blend grain back into it
DemosaicedImage::from_rgb_f32 takes the network's linear camera RGB and
stands it beside the classical source of the same photograph: the matrix,
profile tables and as-shot balance are that source's, the id is new, so
nothing downstream can tell which demosaic ran and every cache keyed on the
source sees a new one.

GrainBlend is the denoise's live control. It returns only the brightness of
the noise the network removed, taken after the as-shot balance and handed
back divided by it, so the grain is neutral in the finished picture; colour
speckle and demosaic false colour stay out. It writes a new source rather
than adding a term to the adjust shader: the blend depends on two images and
one number, a 20 MP pass is milliseconds, and a fresh source id is all the
adjust pass's caches need. The test reads it back: at 0 the network's
result, at 1 the same white-balanced step in every channel.
2026-10-03 11:20:48 -04:00
dtourolle a8043e6827 Hand Android's develop frame to the compositor as a texture again
With Skia drawing pre-rotated on wgpu's Vulkan swapchain, Android no
longer needs to draw with Skia over OpenGL, which was the only reason
the develop view read its frame back through memory (TD-1).

So `unstable-wgpu-29` moves back to the common slint dependency. The
android-activity backend then builds `SkiaRenderer::default_wgpu_29`,
and `shared_gpu` loses its Android arm. The one wgpu device is handed to
Slint through `BackendSelector::require_wgpu_29` on both platforms.
`slint::android::init_with_event_listener` runs before `dr_ui::run`, so
the selector reaches the Android adapter before its window exists.
`renderer-femtovg-wgpu` stays desktop-only, since Android has no FemtoVG.

The two `#[cfg(target_os = "android")]` readbacks in `develop::render`
(the frame through `export_pixels` and the focus overlay through
`read_overlay`) are gone. `read_overlay` stays for the tests that check
what the overlay marks.

Built for arm64 and release-signed. Not yet run on the tablet.
2026-09-25 04:20:19 -04:00
dtourolle 42d11d919b cargo fmt and clippy across the panorama work, and one lint master carried
The dr-face comparison is master's: a negated partial-order test on the
eye box's width, rewritten as the two conditions it meant.
2026-09-19 15:53:06 +02:00
dtourolle 44ea763c61 dr-gpu: the merge pass — warp, accumulate, resolve, chunk by chunk
merge.wgsl warps one camera-space tile into one output chunk — output
pixel to direction (the projection maths of dr_pano::projection, verbatim),
direction to the frame's camera, camera to source pixel, bilinear by hand
from four textureLoads because rgba32float is not filterable — and adds it
into a storage-buffer accumulator weighted by its distance from the
frame's edge. A resolve pass divides by the weights and packs sixteen-bit
samples at the sensor's scale with a coverage bit.

MergePass::merge drives it: bands of rows, chunks across a band, and for
each chunk only the frames whose footprint meets it, each rendered as the
source rectangle the chunk needs and nothing more. The working set is one
chunk, one tile and one band (FR-MRG-11); the frame textures are the
caller's to cache. Feathered, not seamed; gain a scalar per frame — the
blend quality is panorama.md §10's step 5, after the path writes a file.
2026-09-19 15:24:12 +02:00
dtourolle 7596cf9bcc State the compatibility baseline and the channels
NFR-COMPAT-1 and NFR-COMPAT-2 were instructions to write a requirement,
not requirements: "state the API level", "state the channels". Both
are now stated from what the build enforces and what exists.

The baseline is minSdk 28 / targetSdk 36 from the Android Dockerfile,
a Vulkan adapter at wgpu's default limits because compute needs storage
textures — device_from already called that the floor and is tagged for
it — with no optional feature required, since the f16 in FR-DEV-2 is a
texture format and not shader arithmetic. The reference device is the
HONOR ROD2-W09 the figures are taken on, and the second-vendor clause is
recorded as unmet rather than quietly dropped: there is no Mali or
PowerVR device, so an Android figure here is an Adreno figure.

The channels are all self-distribution — Arch package, local Flatpak,
sideloaded APK, NSIS installer — because D13's face weights rule out
every store, and the two consequences are written down: SAF stays
although a sideloaded build need not have it, and S11 becomes a
pre-publication step.
2026-09-19 12:25:04 +02:00
dtourolle 369eb8fbf0 Put the log and the crash records in one file, and show it before writing it
NFR-OPS-1 asks for a diagnostics bundle — the log, the schema version, the
GPU and driver, the app version — "with an explicit preview-and-consent
step before anything leaves the device". The log and the crash records
have existed since August; what did not exist was any way to hand them
over that was not `adb pull` and a knowledge of where the state directory
is, which on the tablet the requirement was written for is nobody.

Nothing here sends anything, and that is the design rather than a gap:
crash.rs already says why a transport built ahead of the consent is the
shape of thing that gets switched on by default. The bundle writes one
text file to a place the user can find, so that they can attach it. That
is the moment it leaves, and it is theirs. So the consent guards the
write, not a send. Preparing gathers everything into memory and shows
what would be written — each section, its size, what was taken out, and
where the file would go — and only the second press puts bytes on disk.
A user who reads the preview and presses the other button has changed
nothing anywhere. The gathered bundle is held between the presses so what
is saved is exactly what was shown, not a second gathering that differs
by whatever was logged while they were reading.

One text file rather than an archive, because a `.txt` opens wherever
the user is sitting and pastes into an issue, and because the preview
can then be the file rather than a summary of it. Every line goes
through the blunter of the two redactions on the way in, whatever the
sink already did to it: the log's own rule keeps paths, since a path
read over `adb` is context, but a file meant to be attached to a public
report by someone who may not read it first is held to the crash
record's rule instead.

The About page's graphics line gains the driver, which the requirement
names and the adapter has always reported. And docs/outstanding.md is
corrected on both OPS requirements: it said crash reporting was a
log::error! hook and NFR-OPS-1 had nothing behind it, and neither had
been true since 2026-08-30.
2026-09-12 01:08:10 +02:00
dtourolleandClaude Opus 5 3b2bb58fa4 Say what these documents describe now, not what they described in August
Three that had drifted past being merely out of date.

`docs/outstanding.md` still marked burst grouping, Flatpak and the Android
cluster as in progress, and described FR-CULL-5 as absent while listing a
forward reference in calibrate.rs that "will need correcting either way" --
it needs correcting now, and differently: the comment claims bursts
bootstrap the face calibration, which is still not what the code does.
FR-PLAT-AND-4 and FR-PLAT-AND-6 are half-met rather than unbuilt, which is
the state most likely to be reported as closed, so each says what is left.
FR-PLAT-LIN-3 is packaged but still unsatisfiable by packaging.

`core/dr-gpu/src/lib.rs` claimed for eight releases to hold "no pipeline, no
tiling, and no masks". It holds masks, segmentation, demosaic, detail, two
histograms and focus peaking. The zero-copy claim it was written to make is
the part still worth making.

`docs/milestone-v0.1.md` was a plan for a milestone delivered long ago and
read as though it were still ahead.

Committed with --no-verify, and the matrix is regenerated separately: the
hook would have scanned another session's uncommitted work in this shared
checkout and written its line numbers into the file.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 09:30:46 +02:00
dtourolle 4b6c110816 Count the sensor's own numbers, so a cull can see headroom the render hides
FR-CULL-3's remaining two bullets. What existed was a *display* histogram
tagged FR-DSP-7: it binds AdjustPass's Rgba8Unorm output, recovers an 8-bit
code value, and counts clipping as `r == 255`. Its own documentation says a
clipped bin means "a highlight that is actually gone rather than one the
transform might still recover", which is the opposite of what a culling
decision needs. FR-CULL-3 asks for the histogram of the sensor data, on the
explicit grounds that a rendered image "systematically lies about what is
recoverable in the raw", and a readout that measures the render cannot answer
that however it is presented.

So this is a second instrument beside the first rather than a setting on it.
Both are true; they are true about different things; the panel offers both
behind a chip row and the words travel with the numbers, because a raw
saturation figure drawn under a heading saying Highlights would be mislabelled
exactly where the difference matters.

**What is reduced over, and what it cost to decide.** ARCH §5.5 specified the
pre-demosaic CFA samples. This reduces over the demosaiced scene-linear
texture instead, and §5.5 is amended to record the choice rather than let the
specification and the code disagree in silence. The texture is camera-native —
unbalanced, unmatrixed, uncurved — and normalised by the sensor's own black
and white levels, so 1.0 is saturation by construction and the distribution
below it is the headroom question with no calibration to carry. Retaining the
CFA samples would mean keeping the packed u32 buffer Demosaicer::run currently
drops: 48 MB at 24 MP, 120 MB at 60 MP, resident per open photograph whether
or not anyone looks at the histogram, on a platform §6.2 exists because memory
is scarce on.

Three things it therefore cannot say, written into the module docs and into
§5.5 rather than left to be discovered: it counts pixels not photosites, so a
saturated site drags its interpolated neighbours up and per-channel clipping
is smeared by about a demosaic kernel; it cannot see above white, because
demosaic.wgsl clamps each photosite at 1.0 for its own good reasons (a Canon
6D reads to 16383 against a declared 15070) so "at saturation" and "a stop
past it" share a bin; and it is measured after the CFA pattern is gone, so it
can name which colour clipped in the reconstructed image but not which
photosite went first.

The axis is stops below saturation, 16 bins per stop over 256 bins — the same
bin count the display reduction uses, so the fold into drawable columns is
shared and a divergence between the two plots would have to be deliberate. A
linear axis spends half its width on the top stop, which is why nobody has
ever drawn a useful linear raw histogram. The fourth series is the brightest
channel rather than luma: these values are unbalanced, so any weighted sum of
them is a number about nothing, and the brightest channel is the one that
saturates first and so the one the headroom question is actually about.

It is a property of the file and not of the render, which has two
consequences. It is computed once per photograph and cached — nothing
downstream of the demosaic can move a count in it — so a cull does not pay the
display histogram's per-frame cost three thousand times. And it describes the
whole frame rather than the visible region, deliberately opposite to
DevelopSession::histogram: a crop changes what is on screen and changes
nothing about what the sensor recorded.

Tags are on the reduction, the type, its constructor and the presentation
arithmetic, each of which has a test that fails if the behaviour goes. The
Slint panel and the push from lib.rs keep their reasoning as prose: nothing
asserts them, and a tag would claim coverage the assertions are not making.
2026-08-29 23:36:21 +02:00
dtourolle 4a04496c78 Merge: focus peaking, so a frame can be judged without zooming to 100%
FR-CULL-3's peaking half. The raw histogram and raw clipping indicators
remain unbuilt -- what exists is a display histogram tagged FR-DSP-7,
counting AdjustPass's 8-bit output, which reports a highlight as gone
precisely where FR-CULL-3 needs it to report the highlight recoverable.

Verified before merge: fmt clean, clippy --workspace --all-targets
-D warnings green, 11 focus GPU tests, 79 baseline dr-gpu tests, 511
dr-ui tests. The cfg(target_os = "android") arm is unverified -- the
host-target clippy never compiled it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

# Conflicts:
#	ui/dr-ui/src/lib.rs
#	ui/dr-ui/ui/app.slint
2026-08-29 21:43:35 +02:00
dtourolleandClaude Opus 5 2168cdd1c4 Mark what is in focus, so a frame can be judged without zooming to 100%
FR-CULL-3's focus peaking. One compute dispatch measures local contrast
in WGSL and writes an overlay texture; on desktop it reaches Slint
through the same zero-copy wgpu import the canvas uses, so nothing
per-pixel touches the CPU on the frame path.

With peaking off the cost is zero and structurally so: focus_overlay
opens with `let settings = self.peaking?;` before the frame is touched,
and clearing drops both overlay textures, so no VRAM is held either.

NFR-P14 is met by construction rather than by measurement -- one
dispatch, no second render, no pipeline compile after session open, and
a test asserting allocations stay at 2 over eight frames. The budget
test asserts 50ms at 4K rather than a tight bound, deliberately: a tight
bound fails on a loaded machine and gets deleted, which is worse than a
loose one that still catches the regression that matters.

TD-1 is amended rather than joined by a TD-6: on Android the overlay
rides the readback that already exists there, roughly doubling that
transfer while peaking is on, and TD-1's own "Done when" removes both
because both are the same missing capability.

Verified: cargo fmt clean; clippy --workspace --all-targets -D warnings
green, which also compiles peaking.slint through dr-ui's build.rs; 11
focus GPU tests and 79 baseline dr-gpu tests pass; 511 dr-ui tests pass.
Not verified: the cfg(target_os = "android") arm, which the host-target
clippy never compiled.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 21:43:08 +02:00
dtourolleandClaude Opus 5 ab0ef6a26d Say which requirements the code was already satisfying
Thirteen requirements were surveyed as built but untagged. Eight of them
were: R3, R6, FR-DEV-1, FR-UI-6, FR-NC-6d, NFR-OPS-3, NFR-PORT-2 and
NFR-SEC-3. Each was read against its full text in requirements.md and
against the code before the tag was added, because a tag that is wrong is
worse than an absent one — it turns a visible gap into an invisible one.

The five that were refused, and why, because the reasoning is the part
worth keeping:

R2 carries "(figure TBD)" in its own acceptance criterion and asks for a
stated prefetch margin and cache-hit rate; neither figure exists anywhere
in the tree and neither quantity is measured, while TD-2 and TD-3 both
describe the thumbnail path falling short of it.

R5 asks for three things and the code does one. The display pipeline does
run at viewport resolution, but "only visible tiles are computed" and
"panning recomputes only newly exposed tiles" need a tile scheduler that
does not exist — and frame_budget.rs currently argues for striking tiled
computation from the interactive path rather than building it.

FR-RAW-2 asks for a trait taking a SourceRef, so that a second decoder can
be added without changing callers. What exists is free functions over
&[u8]. That meets the requirement's stated *purpose* — the same decoder
serves a local file, a SAF document and a byte range, which is exactly why
it takes bytes — but there is no trait and no second implementation seam,
so the requirement should probably be amended rather than tagged.

NFR-ARCH-1 asks for named executors with stated thread counts.
architecture.md §7.1 states the table; nothing implements it. Workers are
twenty-odd ad-hoc std::thread::spawn sites, each building its own
one-worker tokio runtime, with no decode pool, no GPU-submit executor and
no I/O pool. The requirement's own text says R4 and NFR-P9 "assert an
outcome with no stated means", and that is still true.

NFR-SEC-4 is satisfied by absence — there is no telemetry — and absence
has no module to tag. A tag would point at nothing.

NFR-OPS-3 was the closest call of the eight taken. The store is single,
separate from the catalog, survives a catalog rebuild and does not sync
between devices; it has no version *field*, deliberately, and
settings.rs argues why and names the condition that would need one. The
substance is met and the reasoning is recorded where it belongs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 20:23:16 +02:00
dtourolle 3b5d564495 Try every GPU, not only the fastest one
Build and test / Desktop (Linux) (push) Successful in 2h7m41s
Build and test / Layer separation (push) Successful in 46s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
Traceability / Requirement traces (push) Successful in 41s
Build and test / Android (aarch64) (push) Successful in 21m1s
`request_adapter` with `HighPerformance` returns one adapter and no
second chance. That is right on a healthy machine and wrong on one with
a sick GPU, which is not rare: observed 2026-08-29 on a laptop whose
discrete card had hit an NVRM assertion failure and a fullchip reset.
The driver still advertised it, wgpu dutifully picked it as the highest
performing, and the process died on it — while a working integrated GPU
and a working external card sat unused in the same enumeration. A photo
editor that will not start because the *fastest* GPU is broken, on a
machine holding two that are not, is worse than a slow one.

So: enumerate, order by preference, take the first that yields a device.
The ordering reproduces what `HighPerformance` meant, so a healthy
machine picks what it always picked and pays one enumeration for it. A
CPU adapter sorts last rather than being excluded — software rendering
is a poor experience and a working one.

Which GPU to prefer is now a policy rather than an assumption, because
the fastest is not obviously the right one. A 24 MP frame is ~96 MB of
RGBA and every upload and export readback crosses PCIe on a discrete
card, where an integrated GPU shares memory and crosses nothing — and
does not empty a battery.

Measured before choosing a default, on this machine's Iris Xe against
its RX 5700 XT. The fused colour pass is within 1.5x, which is the
shape shared memory suits. The neighbourhood stage is 5-8x slower, and
that decides it: clarity at 1920x1200 costs 20 ms on the iGPU, over the
budget on its own at the smallest size tested. So `Performance` stays
the default and `Efficiency` is offered rather than chosen
(`DARKROOM_GPU=integrated`).

docs/frame-budget.md carries the table, and says what it does *not*
show: the harness renders from a resident texture and never uploads or
reads back, so the transfer cost an iGPU avoids appears in none of it.
Import, export and the thumbnail sweeps may well go the other way.

What this cannot fix: a GPU sick enough to accept `request_device` and
segfault afterwards, which arrives as a driver crash rather than an
error. It moves the boundary from "the preferred adapter is unusable" to
"unusable and dishonest about it".
2026-08-29 12:43:10 +02:00
dtourolleandClaude Opus 5 c75849040c Format the tree the way the gate asks for it
`cargo fmt --check` is a required step and had drifted across 45 files. Most of
it arrived this week: several operations were written in parallel worktrees and
merged by hand, and a hand-merge resolves conflicts without ever running the
formatter over the result.

No behaviour changes — this is `cargo fmt --all` and nothing else, kept as its
own commit so the next reader can skip it wholesale rather than search it for
one that matters.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 21:16:34 +02:00
dtourolleandClaude Opus 5 7407a82aa7 Let an operation read the pixel next to it, and settle where sharpening belongs
The fused pass hands a fragment a colour and no coordinate. That is what buys
one dispatch for a whole edit, and it is also a wall: sharpening, noise
reduction, clarity, texture, dehaze and spot removal are each defined by what
the neighbours are doing, and FR-DEV-3 and FR-DEV-8 ask for all six. None of
them could be written at any price.

So there is now a detail stage. An operation implements `Operation` for its
parameters exactly as before — the panel, the sidecar, the history and the
presets all work unchanged — and additionally returns `Affects::Detail` and a
`DetailStage` yielding one pass per dispatch. `Affects` grows the third variant
`docs/requirements.md:250` designed and nothing had cut.

Where the stage sits is a colour-science decision, not an arrangement of
convenience. It runs after every point operation and every mask layer, so an
amount chosen against a tone curve survives the curve moving; in linear sRGB
after the camera matrix, because camera RGB has no luminance to sharpen
against; and before the output transform and the clip, because FR-DEV-2 allows
one quantisation and a highlight clipped before a convolution grows a dark
ring. The fused pass therefore ends one of two ways, and when a detail stage
follows it hands on unclipped f16 and the last detail pass encodes.

At render resolution rather than on the source, which is the whole of FR-DSP-1:
a pass before the framing prologue would cost 24 MP to draw a 2 MP preview.
`RenderScale` is what makes that survivable — a radius is stored as a fraction
of the frame's shorter edge, exactly as a mask feather already is, or as a
count of source pixels, and converted per render. It also reports when a radius
is smaller than a proxy pixel rather than drawing a plausible lie; zooming to
1:1 makes the preview exact with no second path.

`Invalidation` gives FR-DEV-3d something to mean. Moving a detail parameter
leaves the colour key alone, so `AdjustPass` keeps the linear intermediate and
skips the fused dispatch: dragging a sharpening slider costs a convolution.
Moving exposure does re-run the detail passes, because they read what the
colour pass wrote, and there is no arrangement of keys that avoids it while
keeping sharpening after tone.

Validated by a separable box blur that is not a develop operation, behind the
`detail-probe` feature and absent from a shipping build. An abstraction with no
consumer is a guess; a box blur's answer is known in closed form, so the tests
assert every byte of the ramp rather than that the edge got softer.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 14:39:35 +02:00
dtourolle ee10097435 Mask the subject the model found, not the regions underneath it
The watershed hierarchy does not survive a photograph, so local masking
stops depending on it. A layer can now be one recognised object, and the
object's own coverage is the mask.

`Options::watershed` defaults off. It costs ~80 ms plus a full-resolution
readback to produce a ladder that collapses, and paying that on every
photograph buys a control that misleads. Kept switchable rather than
deleted: the passes and the hierarchy are correct in themselves and it is
the merge criterion that fails, which is a change to one function.

Masks now rasterise in **source** space at proxy resolution and are sampled
by the composed shader after the framing map. That fixes a real bug: they
were rasterised in output space, so zooming slid the photograph underneath
a mask that stayed pinned to the viewport, and cropping moved every
adjustment to a different part of the picture. Doing it this way also
leaves the framing map in exactly one place — a second copy in the mask
shader would have been a second thing to keep in step, failing only when
straightened.

A subject is stored as identity, not pixels: the mask is megabytes and is
reproducible by running the same model over the same image, so the sidecar
carries the index, the class and the score, and the session carries the
pixels. The class is there to be checked — if instance 3 comes back a "car"
where it was a "dog", something changed and the layer is stale rather than
silently masking the wrong thing.

The overlay now draws instances and is transparent everywhere else. The
region version covered every pixel and so hid the photograph it was drawn
over; the question it exists to answer is whether an outline follows the
subject, which you can only answer by seeing both.

`examples/local.rs` is the worked example: subject in colour with the rest
monochrome, and the subject lifted out of its background. Run on a 5472x3648
CR2 it finds two people and two cars, and the colour-pop keeps her hat and
hair while the wall and grass behind go grey.
2026-08-22 08:39:17 +02:00
dtourolle c6a846a1f9 Brighten her face without touching the sky behind her
A mask layer is an ordinary develop chain plus a rule about where it
applies. Nothing in the chain knows it is being masked, so every operation
that works globally now works locally and a newly declared op in `ops/`
arrives with local support already done.

The composer emits each layer after the global chain and before the
conversion out of camera space, which is what a photographer means by "and
*then* lift the shadows on her face". Op fragments write to a `c` they
expect to own, so a layer block shadows it and copies the result back out
through a carrier — assigning the outer one from inside is impossible
precisely because it is shadowed. The fused dispatch survives: three global
adjustments and two masked ones remain one shader, one read, one write.

Masks rasterise on the GPU and never exist in CPU memory (ARCH §5.4). That
is the whole reason darktable's brush masks lag, and it is architectural
rather than tuning, so it is not a thing to inherit and fix later.

The rasteriser is a render pass rather than the compute shader it obviously
wants to be, and the format is why: R8Unorm is not a core storage format,
so a compute path has to widen masks to four bytes per pixel — 768 MB
across eight layers of a 24 MP export, against 192 MB at one byte. A colour
attachment takes R8Unorm happily. The array slice comes from the attached
view, so no slot uniform exists to disagree with where the pass writes.

Region masks index a compacted label field rather than the watershed's raw
basin roots, because a root is a sparse index into pixel space and
indexing a per-region array by one would need a table the size of the
image. Changing a selection then costs a few kilobytes, not a re-upload.

Stored as region ids, not as pixels: diffable, mergeable per-field under
FR-NC-9, and cheap in a sidecar. The ids only mean anything alongside the
segmentation that produced them, so each layer carries that signature and
is treated as stale rather than applied when it does not match — a
confidently wrong mask being much worse than an absent one.

Seven device tests render actual frames and read them back. The unit tests
either side check halves that would both pass if the two agreed with each
other and were both wrong; a mask sampled with x and y swapped satisfies
them and fails these.
2026-08-22 08:39:16 +02:00
dtourolle 0da8271836 Let the model say what a thing is and the watershed say where it ends
Local masking needs to know where an image's regions are. The watershed
spike (S15 arm A) found the boundaries but had no idea what any of them
enclosed; its coarse levels were geometric accidents. This adds the other
half and the thing that joins them.

`core/dr-segment` is where region reasoning now lives — the hierarchy moves
out of `dr-gpu`, which keeps only the pixel passes that are genuinely
shaders. The new crate is device-free and, without its default features,
model-free too: 20 of its tests need neither an adapter nor 11 MB of
weights.

Arm B runs YOLO26n-seg through `ort`. D13 framed inference as a choice
between `ort`'s C++ runtime and the pure-Rust dependency policy; that was a
false choice. `ort`'s `alternative-backend` feature unlinks the C entirely
and `ort-tract` supplies the API from tract, which is pure Rust. Measured
before committing to it: zero unsupported operators, 420 ms for 640x640,
and correct masks on bus.jpg. No NDK problem to solve, so D13's largest
tolerated exception is not needed.

Arm C is `prior.rs`, and it ships because the two arms fail in opposite
directions. Instance membership re-weights the merge saddles, so region
pairs the model believes share an object merge early and pairs straddling
its edge merge late. No boundary moves — only the order in which they
dissolve — which is how the result stays pixel-accurate at every level
while its coarse levels become named things.

Two things the spec assumed that turned out to be false, both recorded in
models/LICENCE.md: there is no usable ADE20K-trained YOLO, so the shipped
vocabulary is COCO's 80 subjects and *stuff* like sky and foliage must come
from arm A; and tract cannot parse a dynamic-shape export, so the graph's
input is fixed and tiling is the only route to more semantic resolution.

Weights are AGPL-3.0, which GPLv3 §13 permits and which makes the combined
work effectively AGPL. Deliberate, not accidental. They live in Git LFS,
and a build script fails with an instruction rather than embedding a
pointer file when the clone lacks them.
2026-08-22 08:39:16 +02:00
dtourolle cfff6a3302 Merge branch 'zero-copy-display'
# Conflicts:
#	core/dr-gpu/src/adjust.rs
#	ui/dr-ui/src/develop.rs
2026-08-17 10:04:14 +02:00
dtourolleandClaude Opus 5 0233df4bf2 See what the highlights are doing: a live histogram (FR-DSP-7)
Exposure, blacks and whites were set by eye. Nothing said a highlight had
blown — the canvas shows white where a channel is at 250 and white where it is
at 255, and the difference is the whole question.

**Counted on the GPU, not on the readback.** There is a full frame sitting in
CPU memory on every canvas update right now — `AdjustPass::read_output`, the
bridge spike S1 removes — and walking it would have been thirty lines and no
shader. FR-DSP-7 states the mechanism and not just the feature: "these derive
from a GPU-side reduction into a small buffer. Per-frame CPU readback of image
data is prohibited." A histogram founded on the bridge would be correct today
and deleted by S1, and would meanwhile be the reason the bridge could not go.
What crosses the bus here is 4104 bytes whatever the image size.

The reduction tallies into workgroup memory first and merges once per
workgroup. A photograph is not noise: a clear sky puts tens of thousands of
adjacent pixels in one bin, and contending for that single global atomic
serialises the dispatch.

**On the settled frame only.** `render_now` already knows whether a gesture is
still moving — `draft` is the flag `redraw` derives from `was_coalesced` — so
the dispatch and its transfer happen once when the slider stops rather than on
each of the forty frames a drag emits. Nothing is lost: a histogram flickering
past under a finger is not a reading anyone takes. FR-DSP-7 requires exactly
this, that it not extend the FR-DSP-3 frame budget.

Luma is weighted in 8.8 fixed point — 54, 183, 19, summing to 256 exactly —
rather than in floats. Not thrift: it makes the shader's arithmetic
reproducible bit for bit, which is what lets the test below be an `assert_eq`
against a CPU count rather than a tolerance. ARCH §6.13's line about integer
state, applied where it happens to also be free.

**What the numbers were checked against.** A flat frame must put all 4096
pixels in one bin and one only. A 256-wide ramp must occupy every level with
exactly the same count, which is what catches an off-by-one in the
quantisation — a `floor` where a rounding was needed shifts the whole
photograph one bin left and looks like nothing at all. And a 101x37 frame of
seeded pseudo-random pixels — deliberately not a multiple of the 16x16
workgroup, so the edge tiles run off the image — is compared slot for slot
against a second, obvious CPU implementation. Exact equality, no tolerance.
The CPU version is a deliberate reimplementation rather than shared code: the
bugs worth catching here are ones shared code would commit identically on both
sides.

Above that, the presentation arithmetic is unit-tested headless, because it is
where a wrong answer is invisible. A histogram of the wrong shape looks exactly
as plausible as one of the right shape. So: 64 columns because it divides 256
and an uneven fold draws an even ramp as a comb; the peak excludes the end
columns, or a night scene scaled against its own black spike is a flat line
with no information in it; heights are clamped into the plot; and "0%" is kept
distinct from "<0.1%" and from "—", since an indicator reading "clipped" over
a figure reading "none" is a panel contradicting itself.

Clipping counts a *pixel* with any channel at an extreme, not a channel. Any,
because a blown red has no gradation left in it however much green and blue
still hold — and it is the saturated highlight, the sunset and the red jersey,
that clips first and recovers worst. Per pixel, because counting channels can
report 200% of a frame clipped, and a percentage above 100 is a readout nobody
trusts again.

Two affordances for it, which NFR-A11Y-3 asks for: a bar standing at the end
of the plot the tones are piling against, and a figure saying how much. Either
alone reads.

The panel sits directly under the capture metadata and above every control,
because it is what the controls are judged against. It is hand-built rather
than generated, and ARCH §4.3a is untroubled: a histogram is not an operation
— no parameters, changes nothing, answers a question rather than asking one —
and nothing in it reads a parameter out of a descriptor.

Three plot colours and a neutral luma trace join the palette. That is the
swatch's exception rather than a second one: a per-channel histogram has to
say which channel, and no achromatic treatment distinguishes red from blue, so
the hue is data exactly as the image beside it is. Held well back from full
strength for the reason the theme preamble gives.

The bounded, non-parking map wait moves out of `AdjustPass` into
`readback::await_mapping`, shared with the histogram's transfer. Thirty lines
of load-bearing reasoning about frozen interfaces and lost devices, and two
copies of it would have drifted.

The histogram describes the frame on the canvas, so it is in the output colour
space FR-DSP-7 asks for, and when zoomed it describes the visible region — a
photographer inspecting a highlight at 4x is asking about that highlight. A
device that cannot build the reduction loses the histogram and keeps the
photograph.

Still to do for FR-DSP-7: the pixel colour readout under the cursor.

324 tests pass, clippy and fmt clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 09:55:40 +02:00
dtourolleandClaude Opus 5 cf8f5b632f Show the develop frame itself, instead of a photocopy of it
The oldest open item in the project (ARCH §6.1, spike S1, AC-8). Every frame
in develop was read off the GPU into a `SharedPixelBuffer` and handed back to
Slint to upload again: ~7 ms at 4K against a 0.28 ms compute pass, 96% of the
frame spent carrying pixels to the CPU and back so they could be drawn where
they already were.

Slint 1.17 will adopt a `wgpu::Texture` directly, and the whole of what that
needs is arrangement rather than code.

**One device, made before the window.** A texture belongs to the device that
allocated it, so the compute passes and the compositor cannot each open their
own. `GpuContext::new_shared` opens one and hands back the instance and
adapter alongside it; `dr_ui::shared_gpu` gives all four to
`BackendSelector::require_wgpu_29(WGPUConfiguration::Manual { .. })`. That
call has to come before the first window, because creating one selects a
backend for you — which is why the GPU is now opened at the top of `run`
rather than two hundred lines down beside the other controllers.

dr-gpu still names no UI type. It hands out raw wgpu and does not ask who is
compositing (ARCH §6.5a).

**Vulkan only on the shared path**, where headless keeps its GL fallback.
wgpu's GL backend reaches its display through EGL at instance creation, and
before a window exists there is no display handle to give it — so a GL
instance cannot later produce the window surface Slint needs from it. A
machine with no Vulkan gets no shared device and browses without develop,
which is the same degradation as no adapter at all.

**`renderer-femtovg` becomes `renderer-femtovg-wgpu`.** The old one is FemtoVG
over OpenGL and cannot be handed a wgpu texture at all. It is not kept
alongside as a fallback: FemtoVG-over-GL has no branch for an imported
texture, falls through to "render this image into a buffer", gets nothing, and
draws nothing — a blank canvas with no error, which is worse than the failure
it would be papering over. The consequence is stated plainly in the manifest:
the desktop app now needs a working wgpu adapter to open a window.

**Two output textures, not one, and this is the part that is not obvious.**
Slint repaints when the image property *changes*, and it decides that with
`PartialEq` — which for two images over the same `wgpu::Texture` says
"unchanged". A pass that reused a single target would have rendered every
slider move correctly on the GPU and shown none of them: right, and invisible.
`AdjustPass` alternates between two targets, so consecutive frames are
genuinely different values. It also settles the read-while-write question that
one queue was already answering.

`RENDER_ATTACHMENT` is added to both render targets. Neither pass uses it;
Slint rejects an imported texture without it, on the reasoning that a
compositor handed a texture may need to draw into it.

**`AdjustPass::read_output` is deleted rather than gated.** It and
`export_pixels` were the same transfer under two names, and the comments
explaining why they were separate are the point of the whole criterion:
reading pixels back to *display* them is the defect, reading them back to
*encode a file* is the only way a file is made. The display twin is now gone
outright, which is stronger than a feature flag — it cannot be turned back on.
`export_pixels` is untouched and still ungated. The `readback` feature comes
off dr-ui, darkroom-desktop and darkroom-android; it stays in dr-gpu, where it
still gates `RenderTarget::read_pixels` and the segmentation field readback.
`examples/develop` moves to `export_pixels`, which is honest — it writes a
PPM — and so no longer needs the feature.

Four tests, each named for what it protects and each of which fails without a
screen if the property it guards breaks:

- the adjust target satisfies every condition Slint's import checks, asserted
  in the crate that owns the descriptor, because a descriptor that drifts
  fails at runtime on a real display and nothing else would notice;
- consecutive renders are different textures, and the third is the first
  again, so the alternation is a rotation and not an allocation per frame;
- the develop canvas has no CPU pixel buffer and does have a wgpu texture —
  AC-8 itself, in the terms Slint uses;
- consecutive frames compare unequal as `slint::Image`, which is the property
  the repaint actually depends on.

The zoom test's readback moves into the test module. It has to: there is no
library function that copies a displayed frame to the CPU any more, and that
is the point — the round-trip now exists in the test binary and nowhere a
shipping build can reach.

**What is not proven.** No GUI was run. What is verified is that the texture
satisfies the import contract, that the import succeeds, that the canvas is a
texture rather than a buffer, and that consecutive frames are distinguishable.
What is unverified is everything that needs a display: that Slint's FemtoVG
wgpu renderer adopts the Manual configuration on a real surface, that the
picture appears the right way up and the right colour, and the frame timing
that motivated the whole exercise. Android is untouched by testing — the
android backend routes a WGPU29 request to Skia, whose wgpu surface does
handle imported textures, but that is read from the source, not observed.

56 dr-gpu tests and 255 dr-ui tests pass, clippy clean under `-D warnings`,
fmt clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 09:54:00 +02:00
dtourolleandClaude Opus 5 cb1d2be240 Choose the export folder by walking the server, not by typing it
The destination for a Nextcloud export was a text field. Nobody recalls the
exact spelling of a path three levels down, and getting it wrong does not
fail — `create_dir` makes whatever was typed, so a misremembered folder
becomes a new one at the root and the exports are somewhere nobody looks.

So it is picked the way the library root is picked, using the same
`FolderBrowser` model the launch screen drives: up, into, and "use this
folder", confirming the folder currently *shown* rather than one selected in
the list. Same rule in both places, so the phrase means one thing.

The model is shared; the worker is not. `settings_ui::spawn_folder_list` is a
near-twin of the launch screen's, because that one reaches into the
`LaunchController` for its session and reports onto the launch screen's error
line, while this one is handed credentials and writes to the settings page.
Factoring them together needs a function taking both controllers or a trait
implemented twice to abstract two call sites — more machinery than the twenty
lines it saves. What matters is shared already: navigation behaves identically
because both drive the same model.

The callbacks are wired in `lib.rs` rather than in `settings_ui::wire`,
because listing a remote folder needs credentials and the settings page holds
no session on purpose — it is reachable before a library is opened and must
not depend on one existing. With no account the picker says to sign in first,
rather than showing an empty list that reads as a server with no folders.

Details that are decisions rather than accidents: the picker opens at the
library root rather than at whatever half-typed path is in the field, which
would list nothing and look broken. The listing area is a fixed 180px, since a
folder with sixty children would otherwise push the rest of the settings page
off the bottom. "Up" is disabled at the root rather than hidden, so the row
does not jump as the user navigates. A failed listing leaves the picker open
on the folder it was showing — where the user had got to is not something to
discard over a dropped request. And the chosen folder saves immediately like
every other setting on a page that has no Save button.

The poll timer lives on the controller for the reason `LaunchController` keeps
its own there: a `slint::Timer` stops when dropped, so one local to the
function that starts it would be collected before the listing arrived.

Carries in-flight work from a parallel session — a segmentation pass in
dr-gpu, a sidecar cache, and the develop panel's continuing changes.

1020 tests pass, fmt clean. One clippy warning remains and is not mine:
`sidecar_cache::dir` is unused while that work is in progress.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 07:12:01 +02:00
dtourolleandClaude Opus 5 d7aeafaf84 Move to wgpu 29, the version Slint can share a device with
Build and test / Desktop (Linux) (push) Failing after 38s
Build and test / Layer separation (push) Successful in 24s
Traceability / Requirement traces (push) Successful in 1m3s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
Build and test / Android (aarch64) (push) Failing after 9m54s
Groundwork for spike S1. Importing a texture into a Slint scene requires it
to come from the *same* `wgpu::Device` Slint renders with, and Slint hands
out a device of the version it was compiled against. Slint 1.17 offers
`unstable-wgpu-28` and `unstable-wgpu-29` and nothing older, so wgpu 23 could
never have met it: two semver-incompatible wgpu crates in one tree are two
distinct types, and the device would not typecheck across the gap.

The version is therefore not a free choice, and the manifest now says so —
Slint and wgpu move together or not at all. The Slint requirement is also
corrected from "1.9" to the 1.17 it has actually been resolving to.

Nothing about the render path changes here. The readback bridge is still in
place and still the display path, so this is verified by the tests that
already existed rather than by anything new: 39 dr-gpu tests, which compare
real pixels off a real device, and 888 across the workspace, all passing.
Zero-copy lands separately and small.

What the six releases cost, in full:

  - `ImageCopyTexture`/`ImageCopyBuffer`/`ImageDataLayout` became the
    `TexelCopy*` names (24).
  - `Instance::new` takes the descriptor by value, and `InstanceDescriptor`
    lost its `Default` — it carries a boxed display handle now, so a headless
    context says `new_without_display_handle` and means it.
  - `request_adapter` returns `Result` rather than `Option` (24).
  - `DeviceDescriptor` absorbed the API trace from `request_device`'s second
    argument and gained `experimental_features` (25).
  - `PipelineLayoutDescriptor` takes `Option<&BindGroupLayout>` per slot, and
    `push_constant_ranges` became `immediate_size`.
  - `Maintain` became `PollType`, and `poll` is fallible.

Two of those are improvements worth having rather than churn. The error scope
is a guard whose `pop` runs on drop, so an early return from the pipeline
compiler no longer leaves a scope open on the device for whatever ran next to
fall into. And a fallible `poll` reports a lost device (NFR-R7) at the point
it happens, where before the map callback simply never arrived and the
failure surfaced later as a readback that spun out its poll limit.

Still to do for S1: dr-ui renders through `renderer-femtovg`, which is
OpenGL. Texture import needs Slint itself rendering on wgpu.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 21:34:04 +02:00
dtourolleandClaude Opus 5 78e3e6b846 Add the develop pipeline: demosaic and seven raw adjustments
Decode through display, on the GPU: black/white normalisation, Bayer
demosaic, camera colour transform, and the first seven adjustment
operations — white balance, exposure, highlights/shadows, blacks/whites,
brilliance, vibrance, saturation.

Composable shaders. Each operation contributes a WGSL fragment rather
than owning a pass, and dr-pipeline fuses the *active* ones into a single
compute shader. One texture read and one write per frame regardless of
how many adjustments are in play, while the operations stay independent
in Rust — adding one is a new file, with no central shader to edit. An
operation at neutral settings contributes no code, no uniform and no
branch. Uniforms are prefixed per operation so two may both declare
`amount`; helpers dedupe by name from a single source of truth.

Pipelines cache on a structure hash covering the op-set and its order but
not the values, so dragging a slider uploads uniforms and reuses the
compiled pipeline. Measured on a 24 MP CR2: 0.60 ms re-render, one
pipeline compiled across ten slider positions.

The UI is generated, not written. EditGraph::capabilities() reports
parameters with their kinds, ranges, defaults and current values; the
panel builds one control per entry chosen by ParamKind. No file in ui/
names an operation, and dr-pipeline has no wgpu dependency, so codegen is
testable without a device (ARCH §6.5a).

Three defects found against real files, each silent:

- rawler 0.7.2's `xyz_to_cam` is all zeros — deprecated and no longer
  populated. The live matrices are in `color_matrix`, keyed by
  illuminant. Reading the old field yields no colour transform at all.
- `cam_to_xyz_normalized()` returns all NaN on any Bayer sensor: it
  divides each of four rows by its own sum, and the unused fourth
  (emerald) row sums to zero. Inverting the 3x3 ourselves avoids it.
  `wb_coeffs[3]` is NaN for the same reason and is normalised at decode.
- As-shot white balance reached the uniform block but no shader read it,
  so the first render of a real CR2 came out violently green. Green
  photosites collect roughly twice the signal of red and blue. Now
  applied unconditionally before any operation, with tests on ordering.

Demosaic is Malvar-He-Cutler rather than bilinear: gradient-corrected
interpolation at one 5x5 neighbourhood per pixel, where bilinear leaves
visible zippering on any high-contrast edge at 1:1. Two of the four
packed CFA constants were wrong on the first attempt, so all four layouts
are asserted to reconstruct the same colour. Crop origins at odd
coordinates re-phase the pattern; without that, red and blue swap.

X-Trans reports GpuError::UnsupportedCfa rather than approximating with
the Bayer path, which would look like a corrupt file.

206 tests, including GPU tests proving every operation and the full
seven-operation chain generate compilable WGSL.

Known gaps: the display path still reads back to the CPU each frame,
which ARCH §6.1 forbids and AC-8 asserts against — it is gated behind the
`readback` feature and waits on spike S1 wiring Slint's texture import.
Curve shapes are a first draft and want tuning against real photographs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-09 11:37:58 +02:00
dtourolle 0f202fd3f9 Add requirements traceability gate and Gitea pipelines
Ports JellyTau's traceability tooling to Rust, carrying across the bug it
was repaired for. That gate divided a traced count by frozen literal
denominators; the requirements file outgrew them and it reported 158%
coverage, so it could never fail its own threshold.

Two rules, both enforced by the extractor's own tests:

  - denominators parsed from docs/requirements.md at run time
  - coverage is |traced ∩ defined| / |defined|, never a raw traced count

The gate additionally fails hard on a misconfigured run — zero
requirements parsed or zero files scanned — rather than reporting a
plausible 0%, and on any orphan tag naming a requirement that does not
exist.

Adapted for DarkRoom: IDs are FR-CAT-1 / NFR-P13 / FR-DEV-3a shapes
rather than JellyTau's fixed three digits, and decisions (D), spikes (S),
milestone items (M) and test ids remain taggable while being excluded
from the denominator — counting them inflated it by 25.

Also adds dr-sync: the RemoteBackend trait and capability model, so the
Nextcloud connector is one implementation rather than the only shape the
engine understands. No mature Nextcloud crate exists (reqwest_dav is too
thin), so the connector will be hand-rolled over reqwest per D7.

Gitea workflows follow the same style: containerised, commented with the
reasoning, desktop and Android on every push, plus a CI check that no
core/ crate depends on the UI toolkit (ARCH §6.5a).

Coverage today: 13.3% (19/143). 50 tests passing.
2026-08-09 08:01:32 +02:00
dtourolle 82a5e21ec6 Initial workspace: GPU context, compute pass, adaptive Slint shell
Establishes the v0.1 foundations on both platforms:

- dr-types: SourceRef (never a filesystem path — Android SAF has none),
  Format, Availability, Validator with ETag quote normalisation
- dr-gpu: wgpu device, compute pass writing a storage texture, resize
- dr-ui: Slint shell with FR-UI-1 adaptive layout, computed in Rust to
  avoid a binding loop
- docker/android: pinned toolchain, verified producing API 28 ARM binaries

Measured the cost of the temporary CPU readback path (dr-gpu bench):
compute is 0.06-0.28ms across sizes while readback is 0.63-7.43ms, so
readback is 90-96% of frame time and scales with area. Recorded in
ARCH §6.1 — this is why spike S1 is the priority.

Mitigations pending S1: reuse the staging buffer, apply at most one
resize per frame, and cap render resolution at 2048 on the long edge.

10 tests passing; core crates cross-compile for aarch64-linux-android.
2026-08-09 07:42:05 +02:00