Commit Graph
8 Commits
Author SHA1 Message Date
dtourolle cfff6a3302 Merge branch 'zero-copy-display'
# Conflicts:
#	core/dr-gpu/src/adjust.rs
#	ui/dr-ui/src/develop.rs
2026-08-17 10:04:14 +02:00
dtourolleandClaude Opus 5 0233df4bf2 See what the highlights are doing: a live histogram (FR-DSP-7)
Exposure, blacks and whites were set by eye. Nothing said a highlight had
blown — the canvas shows white where a channel is at 250 and white where it is
at 255, and the difference is the whole question.

**Counted on the GPU, not on the readback.** There is a full frame sitting in
CPU memory on every canvas update right now — `AdjustPass::read_output`, the
bridge spike S1 removes — and walking it would have been thirty lines and no
shader. FR-DSP-7 states the mechanism and not just the feature: "these derive
from a GPU-side reduction into a small buffer. Per-frame CPU readback of image
data is prohibited." A histogram founded on the bridge would be correct today
and deleted by S1, and would meanwhile be the reason the bridge could not go.
What crosses the bus here is 4104 bytes whatever the image size.

The reduction tallies into workgroup memory first and merges once per
workgroup. A photograph is not noise: a clear sky puts tens of thousands of
adjacent pixels in one bin, and contending for that single global atomic
serialises the dispatch.

**On the settled frame only.** `render_now` already knows whether a gesture is
still moving — `draft` is the flag `redraw` derives from `was_coalesced` — so
the dispatch and its transfer happen once when the slider stops rather than on
each of the forty frames a drag emits. Nothing is lost: a histogram flickering
past under a finger is not a reading anyone takes. FR-DSP-7 requires exactly
this, that it not extend the FR-DSP-3 frame budget.

Luma is weighted in 8.8 fixed point — 54, 183, 19, summing to 256 exactly —
rather than in floats. Not thrift: it makes the shader's arithmetic
reproducible bit for bit, which is what lets the test below be an `assert_eq`
against a CPU count rather than a tolerance. ARCH §6.13's line about integer
state, applied where it happens to also be free.

**What the numbers were checked against.** A flat frame must put all 4096
pixels in one bin and one only. A 256-wide ramp must occupy every level with
exactly the same count, which is what catches an off-by-one in the
quantisation — a `floor` where a rounding was needed shifts the whole
photograph one bin left and looks like nothing at all. And a 101x37 frame of
seeded pseudo-random pixels — deliberately not a multiple of the 16x16
workgroup, so the edge tiles run off the image — is compared slot for slot
against a second, obvious CPU implementation. Exact equality, no tolerance.
The CPU version is a deliberate reimplementation rather than shared code: the
bugs worth catching here are ones shared code would commit identically on both
sides.

Above that, the presentation arithmetic is unit-tested headless, because it is
where a wrong answer is invisible. A histogram of the wrong shape looks exactly
as plausible as one of the right shape. So: 64 columns because it divides 256
and an uneven fold draws an even ramp as a comb; the peak excludes the end
columns, or a night scene scaled against its own black spike is a flat line
with no information in it; heights are clamped into the plot; and "0%" is kept
distinct from "<0.1%" and from "—", since an indicator reading "clipped" over
a figure reading "none" is a panel contradicting itself.

Clipping counts a *pixel* with any channel at an extreme, not a channel. Any,
because a blown red has no gradation left in it however much green and blue
still hold — and it is the saturated highlight, the sunset and the red jersey,
that clips first and recovers worst. Per pixel, because counting channels can
report 200% of a frame clipped, and a percentage above 100 is a readout nobody
trusts again.

Two affordances for it, which NFR-A11Y-3 asks for: a bar standing at the end
of the plot the tones are piling against, and a figure saying how much. Either
alone reads.

The panel sits directly under the capture metadata and above every control,
because it is what the controls are judged against. It is hand-built rather
than generated, and ARCH §4.3a is untroubled: a histogram is not an operation
— no parameters, changes nothing, answers a question rather than asking one —
and nothing in it reads a parameter out of a descriptor.

Three plot colours and a neutral luma trace join the palette. That is the
swatch's exception rather than a second one: a per-channel histogram has to
say which channel, and no achromatic treatment distinguishes red from blue, so
the hue is data exactly as the image beside it is. Held well back from full
strength for the reason the theme preamble gives.

The bounded, non-parking map wait moves out of `AdjustPass` into
`readback::await_mapping`, shared with the histogram's transfer. Thirty lines
of load-bearing reasoning about frozen interfaces and lost devices, and two
copies of it would have drifted.

The histogram describes the frame on the canvas, so it is in the output colour
space FR-DSP-7 asks for, and when zoomed it describes the visible region — a
photographer inspecting a highlight at 4x is asking about that highlight. A
device that cannot build the reduction loses the histogram and keeps the
photograph.

Still to do for FR-DSP-7: the pixel colour readout under the cursor.

324 tests pass, clippy and fmt clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 09:55:40 +02:00
dtourolleandClaude Opus 5 cf8f5b632f Show the develop frame itself, instead of a photocopy of it
The oldest open item in the project (ARCH §6.1, spike S1, AC-8). Every frame
in develop was read off the GPU into a `SharedPixelBuffer` and handed back to
Slint to upload again: ~7 ms at 4K against a 0.28 ms compute pass, 96% of the
frame spent carrying pixels to the CPU and back so they could be drawn where
they already were.

Slint 1.17 will adopt a `wgpu::Texture` directly, and the whole of what that
needs is arrangement rather than code.

**One device, made before the window.** A texture belongs to the device that
allocated it, so the compute passes and the compositor cannot each open their
own. `GpuContext::new_shared` opens one and hands back the instance and
adapter alongside it; `dr_ui::shared_gpu` gives all four to
`BackendSelector::require_wgpu_29(WGPUConfiguration::Manual { .. })`. That
call has to come before the first window, because creating one selects a
backend for you — which is why the GPU is now opened at the top of `run`
rather than two hundred lines down beside the other controllers.

dr-gpu still names no UI type. It hands out raw wgpu and does not ask who is
compositing (ARCH §6.5a).

**Vulkan only on the shared path**, where headless keeps its GL fallback.
wgpu's GL backend reaches its display through EGL at instance creation, and
before a window exists there is no display handle to give it — so a GL
instance cannot later produce the window surface Slint needs from it. A
machine with no Vulkan gets no shared device and browses without develop,
which is the same degradation as no adapter at all.

**`renderer-femtovg` becomes `renderer-femtovg-wgpu`.** The old one is FemtoVG
over OpenGL and cannot be handed a wgpu texture at all. It is not kept
alongside as a fallback: FemtoVG-over-GL has no branch for an imported
texture, falls through to "render this image into a buffer", gets nothing, and
draws nothing — a blank canvas with no error, which is worse than the failure
it would be papering over. The consequence is stated plainly in the manifest:
the desktop app now needs a working wgpu adapter to open a window.

**Two output textures, not one, and this is the part that is not obvious.**
Slint repaints when the image property *changes*, and it decides that with
`PartialEq` — which for two images over the same `wgpu::Texture` says
"unchanged". A pass that reused a single target would have rendered every
slider move correctly on the GPU and shown none of them: right, and invisible.
`AdjustPass` alternates between two targets, so consecutive frames are
genuinely different values. It also settles the read-while-write question that
one queue was already answering.

`RENDER_ATTACHMENT` is added to both render targets. Neither pass uses it;
Slint rejects an imported texture without it, on the reasoning that a
compositor handed a texture may need to draw into it.

**`AdjustPass::read_output` is deleted rather than gated.** It and
`export_pixels` were the same transfer under two names, and the comments
explaining why they were separate are the point of the whole criterion:
reading pixels back to *display* them is the defect, reading them back to
*encode a file* is the only way a file is made. The display twin is now gone
outright, which is stronger than a feature flag — it cannot be turned back on.
`export_pixels` is untouched and still ungated. The `readback` feature comes
off dr-ui, darkroom-desktop and darkroom-android; it stays in dr-gpu, where it
still gates `RenderTarget::read_pixels` and the segmentation field readback.
`examples/develop` moves to `export_pixels`, which is honest — it writes a
PPM — and so no longer needs the feature.

Four tests, each named for what it protects and each of which fails without a
screen if the property it guards breaks:

- the adjust target satisfies every condition Slint's import checks, asserted
  in the crate that owns the descriptor, because a descriptor that drifts
  fails at runtime on a real display and nothing else would notice;
- consecutive renders are different textures, and the third is the first
  again, so the alternation is a rotation and not an allocation per frame;
- the develop canvas has no CPU pixel buffer and does have a wgpu texture —
  AC-8 itself, in the terms Slint uses;
- consecutive frames compare unequal as `slint::Image`, which is the property
  the repaint actually depends on.

The zoom test's readback moves into the test module. It has to: there is no
library function that copies a displayed frame to the CPU any more, and that
is the point — the round-trip now exists in the test binary and nowhere a
shipping build can reach.

**What is not proven.** No GUI was run. What is verified is that the texture
satisfies the import contract, that the import succeeds, that the canvas is a
texture rather than a buffer, and that consecutive frames are distinguishable.
What is unverified is everything that needs a display: that Slint's FemtoVG
wgpu renderer adopts the Manual configuration on a real surface, that the
picture appears the right way up and the right colour, and the frame timing
that motivated the whole exercise. Android is untouched by testing — the
android backend routes a WGPU29 request to Skia, whose wgpu surface does
handle imported textures, but that is read from the source, not observed.

56 dr-gpu tests and 255 dr-ui tests pass, clippy clean under `-D warnings`,
fmt clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 09:54:00 +02:00
dtourolleandClaude Opus 5 cb1d2be240 Choose the export folder by walking the server, not by typing it
The destination for a Nextcloud export was a text field. Nobody recalls the
exact spelling of a path three levels down, and getting it wrong does not
fail — `create_dir` makes whatever was typed, so a misremembered folder
becomes a new one at the root and the exports are somewhere nobody looks.

So it is picked the way the library root is picked, using the same
`FolderBrowser` model the launch screen drives: up, into, and "use this
folder", confirming the folder currently *shown* rather than one selected in
the list. Same rule in both places, so the phrase means one thing.

The model is shared; the worker is not. `settings_ui::spawn_folder_list` is a
near-twin of the launch screen's, because that one reaches into the
`LaunchController` for its session and reports onto the launch screen's error
line, while this one is handed credentials and writes to the settings page.
Factoring them together needs a function taking both controllers or a trait
implemented twice to abstract two call sites — more machinery than the twenty
lines it saves. What matters is shared already: navigation behaves identically
because both drive the same model.

The callbacks are wired in `lib.rs` rather than in `settings_ui::wire`,
because listing a remote folder needs credentials and the settings page holds
no session on purpose — it is reachable before a library is opened and must
not depend on one existing. With no account the picker says to sign in first,
rather than showing an empty list that reads as a server with no folders.

Details that are decisions rather than accidents: the picker opens at the
library root rather than at whatever half-typed path is in the field, which
would list nothing and look broken. The listing area is a fixed 180px, since a
folder with sixty children would otherwise push the rest of the settings page
off the bottom. "Up" is disabled at the root rather than hidden, so the row
does not jump as the user navigates. A failed listing leaves the picker open
on the folder it was showing — where the user had got to is not something to
discard over a dropped request. And the chosen folder saves immediately like
every other setting on a page that has no Save button.

The poll timer lives on the controller for the reason `LaunchController` keeps
its own there: a `slint::Timer` stops when dropped, so one local to the
function that starts it would be collected before the listing arrived.

Carries in-flight work from a parallel session — a segmentation pass in
dr-gpu, a sidecar cache, and the develop panel's continuing changes.

1020 tests pass, fmt clean. One clippy warning remains and is not mine:
`sidecar_cache::dir` is unused while that work is in progress.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 07:12:01 +02:00
dtourolleandClaude Opus 5 d7aeafaf84 Move to wgpu 29, the version Slint can share a device with
Build and test / Desktop (Linux) (push) Failing after 38s
Build and test / Layer separation (push) Successful in 24s
Traceability / Requirement traces (push) Successful in 1m3s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
Build and test / Android (aarch64) (push) Failing after 9m54s
Groundwork for spike S1. Importing a texture into a Slint scene requires it
to come from the *same* `wgpu::Device` Slint renders with, and Slint hands
out a device of the version it was compiled against. Slint 1.17 offers
`unstable-wgpu-28` and `unstable-wgpu-29` and nothing older, so wgpu 23 could
never have met it: two semver-incompatible wgpu crates in one tree are two
distinct types, and the device would not typecheck across the gap.

The version is therefore not a free choice, and the manifest now says so —
Slint and wgpu move together or not at all. The Slint requirement is also
corrected from "1.9" to the 1.17 it has actually been resolving to.

Nothing about the render path changes here. The readback bridge is still in
place and still the display path, so this is verified by the tests that
already existed rather than by anything new: 39 dr-gpu tests, which compare
real pixels off a real device, and 888 across the workspace, all passing.
Zero-copy lands separately and small.

What the six releases cost, in full:

  - `ImageCopyTexture`/`ImageCopyBuffer`/`ImageDataLayout` became the
    `TexelCopy*` names (24).
  - `Instance::new` takes the descriptor by value, and `InstanceDescriptor`
    lost its `Default` — it carries a boxed display handle now, so a headless
    context says `new_without_display_handle` and means it.
  - `request_adapter` returns `Result` rather than `Option` (24).
  - `DeviceDescriptor` absorbed the API trace from `request_device`'s second
    argument and gained `experimental_features` (25).
  - `PipelineLayoutDescriptor` takes `Option<&BindGroupLayout>` per slot, and
    `push_constant_ranges` became `immediate_size`.
  - `Maintain` became `PollType`, and `poll` is fallible.

Two of those are improvements worth having rather than churn. The error scope
is a guard whose `pop` runs on drop, so an early return from the pipeline
compiler no longer leaves a scope open on the device for whatever ran next to
fall into. And a fallible `poll` reports a lost device (NFR-R7) at the point
it happens, where before the map callback simply never arrived and the
failure surfaced later as a readback that spun out its poll limit.

Still to do for S1: dr-ui renders through `renderer-femtovg`, which is
OpenGL. Texture import needs Slint itself rendering on wgpu.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 21:34:04 +02:00
dtourolleandClaude Opus 5 78e3e6b846 Add the develop pipeline: demosaic and seven raw adjustments
Decode through display, on the GPU: black/white normalisation, Bayer
demosaic, camera colour transform, and the first seven adjustment
operations — white balance, exposure, highlights/shadows, blacks/whites,
brilliance, vibrance, saturation.

Composable shaders. Each operation contributes a WGSL fragment rather
than owning a pass, and dr-pipeline fuses the *active* ones into a single
compute shader. One texture read and one write per frame regardless of
how many adjustments are in play, while the operations stay independent
in Rust — adding one is a new file, with no central shader to edit. An
operation at neutral settings contributes no code, no uniform and no
branch. Uniforms are prefixed per operation so two may both declare
`amount`; helpers dedupe by name from a single source of truth.

Pipelines cache on a structure hash covering the op-set and its order but
not the values, so dragging a slider uploads uniforms and reuses the
compiled pipeline. Measured on a 24 MP CR2: 0.60 ms re-render, one
pipeline compiled across ten slider positions.

The UI is generated, not written. EditGraph::capabilities() reports
parameters with their kinds, ranges, defaults and current values; the
panel builds one control per entry chosen by ParamKind. No file in ui/
names an operation, and dr-pipeline has no wgpu dependency, so codegen is
testable without a device (ARCH §6.5a).

Three defects found against real files, each silent:

- rawler 0.7.2's `xyz_to_cam` is all zeros — deprecated and no longer
  populated. The live matrices are in `color_matrix`, keyed by
  illuminant. Reading the old field yields no colour transform at all.
- `cam_to_xyz_normalized()` returns all NaN on any Bayer sensor: it
  divides each of four rows by its own sum, and the unused fourth
  (emerald) row sums to zero. Inverting the 3x3 ourselves avoids it.
  `wb_coeffs[3]` is NaN for the same reason and is normalised at decode.
- As-shot white balance reached the uniform block but no shader read it,
  so the first render of a real CR2 came out violently green. Green
  photosites collect roughly twice the signal of red and blue. Now
  applied unconditionally before any operation, with tests on ordering.

Demosaic is Malvar-He-Cutler rather than bilinear: gradient-corrected
interpolation at one 5x5 neighbourhood per pixel, where bilinear leaves
visible zippering on any high-contrast edge at 1:1. Two of the four
packed CFA constants were wrong on the first attempt, so all four layouts
are asserted to reconstruct the same colour. Crop origins at odd
coordinates re-phase the pattern; without that, red and blue swap.

X-Trans reports GpuError::UnsupportedCfa rather than approximating with
the Bayer path, which would look like a corrupt file.

206 tests, including GPU tests proving every operation and the full
seven-operation chain generate compilable WGSL.

Known gaps: the display path still reads back to the CPU each frame,
which ARCH §6.1 forbids and AC-8 asserts against — it is gated behind the
`readback` feature and waits on spike S1 wiring Slint's texture import.
Curve shapes are a first draft and want tuning against real photographs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-09 11:37:58 +02:00
dtourolle 0f202fd3f9 Add requirements traceability gate and Gitea pipelines
Ports JellyTau's traceability tooling to Rust, carrying across the bug it
was repaired for. That gate divided a traced count by frozen literal
denominators; the requirements file outgrew them and it reported 158%
coverage, so it could never fail its own threshold.

Two rules, both enforced by the extractor's own tests:

  - denominators parsed from docs/requirements.md at run time
  - coverage is |traced ∩ defined| / |defined|, never a raw traced count

The gate additionally fails hard on a misconfigured run — zero
requirements parsed or zero files scanned — rather than reporting a
plausible 0%, and on any orphan tag naming a requirement that does not
exist.

Adapted for DarkRoom: IDs are FR-CAT-1 / NFR-P13 / FR-DEV-3a shapes
rather than JellyTau's fixed three digits, and decisions (D), spikes (S),
milestone items (M) and test ids remain taggable while being excluded
from the denominator — counting them inflated it by 25.

Also adds dr-sync: the RemoteBackend trait and capability model, so the
Nextcloud connector is one implementation rather than the only shape the
engine understands. No mature Nextcloud crate exists (reqwest_dav is too
thin), so the connector will be hand-rolled over reqwest per D7.

Gitea workflows follow the same style: containerised, commented with the
reasoning, desktop and Android on every push, plus a CI check that no
core/ crate depends on the UI toolkit (ARCH §6.5a).

Coverage today: 13.3% (19/143). 50 tests passing.
2026-08-09 08:01:32 +02:00
dtourolle 82a5e21ec6 Initial workspace: GPU context, compute pass, adaptive Slint shell
Establishes the v0.1 foundations on both platforms:

- dr-types: SourceRef (never a filesystem path — Android SAF has none),
  Format, Availability, Validator with ETag quote normalisation
- dr-gpu: wgpu device, compute pass writing a storage texture, resize
- dr-ui: Slint shell with FR-UI-1 adaptive layout, computed in Rust to
  avoid a binding loop
- docker/android: pinned toolchain, verified producing API 28 ARM binaries

Measured the cost of the temporary CPU readback path (dr-gpu bench):
compute is 0.06-0.28ms across sizes while readback is 0.63-7.43ms, so
readback is 90-96% of frame time and scales with area. Recorded in
ARCH §6.1 — this is why spike S1 is the priority.

Mitigations pending S1: reuse the staging buffer, apply at most one
resize per frame, and cap render resolution at 2048 on the long edge.

10 tests passing; core crates cross-compile for aarch64-linux-android.
2026-08-09 07:42:05 +02:00