c50d96e94977be56b2aaa36a0adab2c41ff6141a
9
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c50d96e949 |
Repair hot and dead photosites before the demosaic
A hot photosite went into the demosaic as it was read, and came out as a coloured cross three pixels wide that nothing later could take back out. Night and long exposures showed them; the defect-map reader added for FR-RAW-3 was never wired in, and a CR2 carries no map anyway. A pass over the mosaic now runs ahead of the demosaic, into a second buffer. A photosite is hot when it reads more than twice every same-colour photosite in its 5x5 window plus 2% of the range, and more than twice each of its eight immediate neighbours of any colour. The second half keeps stars and glints: real light reaches the sensor through a lens and an anti-aliasing filter and lights a patch, so the photosites beside it are lit too, where a hot photosite's are dark. It is replaced by its brightest same-colour neighbour, which invents nothing. Dead photosites are the mirror case, judged only where the neighbourhood is above 5%, so shadow noise clipped at black is left alone. The colour of each photosite comes from a 6x6 sensor-anchored tile, so Bayer and X-Trans share the pass. Export and every other path that demosaics get it too, and there is no setting: the repair only fires where a single photosite disagrees with everything around it. Cost, warm, on a Canon 6D frame (RTX 3050): 91-99 ms to demosaic before, 94-98 ms after; the extra pass is inside the run-to-run noise. Tests render a frame with and without the defect and compare the finished pixels. Without the repair a hot photosite showed by 230 and a dead one by 168; with it neither shows, and a 3x3 highlight at white survives. |
||
|
|
1dc7b45cfe |
Read the fit view's source gather once per framing, not once per frame
At fit, every output pixel of the fused pass loads one texel from a source three or four times its width, on a stride. The memory system fetches the texels it skips along with the one it wanted, so on a 60 MP rgba16float source that gather was most of what the fused pass cost: 10.6 ms of a 2560x1600 frame against 3.8 ms for the same shader reading a contiguous window (the 1:1 view). At 3840x2160 it was 21.1 ms. Those are the laptop RTX 3050 with its clocks held at 420/810 MHz by the power cap; unthrottled the same frames were about 2.0 and 3.2 ms, and the gather is the same share of them. Which texel an output pixel reads depends only on the framing prologue, the framing and warp uniforms, the source and the render size. None of those move during a slider drag, so the gather is the same work every frame. The fused shader now takes a render-sized rgba16float cache of it (bindings 6 and 7, declared in every generated shader like the masks) and a pair of uniform flags: write what was gathered, or read it back at the pixel's own coordinate. AdjustPass keeps the cache and decides per dispatch. The composer supplies `ComposedShader::sample_key`, a hash of the prologue and those uniforms, and AdjustPass adds the image and the size; an image gets a process-unique id for this rather than being held alive by the key. The picture is bit-for-bit the same. The source is rgba16float and so is the cache, so the stored texel is the texel, and only the path that reads a texel whole takes part: an interpolated sample (straightening, lens warps, CA) is a blend that f16 could not hold exactly, so the composer gives it no key and it reads directly as before. The cache is written on the second frame with a given key, not the first: a crop or zoom drag changes the key every frame, and writing then would add a render-sized write to exactly the gestures that can afford it least. It is kept only up to 3840x2400, so an export never parks a full-frame copy on the device, and `release_caches` drops it. Measured with a scratch probe rendering the synthetic 60 MP frame from examples/frame_budget.rs, forty frames per run after six warm-up, five runs of each binary alternated, median of the per-run p50 (GPU idle apart from the power cap): scene before after neutral 2560x1600 fit 10.62 ms 3.88 ms exposure 2560x1600 fit 10.83 ms 3.87 ms nr chroma 2560x1600 fit 19.84 ms 12.69 ms neutral 3840x2160 fit 21.05 ms 7.11 ms exposure 3840x2160 fit 21.08 ms 6.94 ms clarity 3840x2160 fit 42.20 ms 27.88 ms neutral 2560x1600 1:1 3.83 ms 3.84 ms (control: nothing to gain) The rgba8 output of every scene hashed identically before and after, in isolated runs and across all 38 scene/size/view combinations of the probe. New tests walk a pass through direct, write and read frames, a slider move, a neighbourhood operation and a framing change, and compare every frame with a fresh pass that can only have read directly. |
||
|
|
c6cfb2a02a |
Put the -1 on the greens along the chroma axis, not across it
The Malvar "R at green in R row" kernel weights the two greens two sites away along the row at -1 and the pair up and down the column at +1/2. The shader had the two swapped, in the comment as well as the code, so the transcription checked against itself. Both sum to zero and reconstruct a flat patch exactly, which is all the tests fed it. On an edge the correction at green sites is half strength and the false colour doubles: 0.375 against 0.19 on a grey step, and a blue/yellow zipper around every clipped highlight at 1:1. The other three kernels and the CFA tables were right. A grey vertical step now runs through the pass; the transposed kernel fails it at 0.375. |
||
|
|
acab0d7abb |
A linear DNG in and out: the writer, and a three-sample RawImage
dr-export gains write_linear_dng — LinearRaw, DNG 1.4, u16 samples at the sensor's scale, the body's matrices with their illuminants, the as-shot neutral, the EXIF block an export writes — streamed strip by strip through a closure so the composite is never held (FR-MRG-11). The tiff crate's directory is a map, so PhotometricInterpretation is written over what new_image set, which is the trick the S15.1 spike thought it had to hand-roll around. The test reads the file back through rawler. dr-decode's RawImage carries samples_per_pixel (a linear DNG is 3), the body's profile with its calibrations mapped back to EXIF illuminant codes, and the cleaned make and model. The GPU uploads a three-sample image as it is, normalised by black and white like a photosite, through a full f16 conversion — subnormals kept, because a 14-bit LSB sits at f16's smallest normal and rounding it to zero would crush exactly the shadows the file was written to keep. |
||
|
|
743fefe7f1 |
Render each body through the profile its own files describe
Colour came from whichever matrix rawler happened to key `D65`, the second one was discarded, and the rendering was left linear. That is the dcraw default, and FR-DEV-3e names it as the reason people abandon a converter in the first hour: correct in the abstract, flat and poor on skin in practice. The decoder now builds a camera profile. - `ColorMatrix1/2` and `CalibrationIlluminant1/2`. rawler surfaces these as an illuminant-keyed map — for DNGs from the tags, and for native formats from its own camera database — so a Canon CR2 arrives with a tungsten matrix and a daylight matrix exactly as an Adobe DNG of the same frame would. Dual-illuminant support is therefore not a DNG feature here. - `ForwardMatrix1/2`, read straight from the root IFD, because rawler parses them and never surfaces them. Where a file carries both, they replace the inverted colour matrix: the same relationship measured in the direction rendering actually wants, rather than an inversion that amplifies the measurement error exactly where skin lives. - `AsShotNeutral`, used to estimate what the scene was lit by and to interpolate between the two calibrations in mireds. The estimate is circular — the temperature needs a matrix and the matrix needs the temperature — so it is a fixed point, three rounds, as Adobe's SDK does it. Bodies calibrated at neither D65 nor A stopped rendering uncalibrated as a side effect: a Phase One IQ3 carries D55 and D75 and used to get no matrix at all. And a base curve, applied per channel in camera RGB between the last adjustment and the conversion out of camera space — a toe, a steep midtone and a shoulder, which is the difference between a photograph and a scan of one. It is not an edit: no slider, nothing in the sidecar, because it belongs to the body rather than to anything anyone decided, and a sidecar is shared between bodies. It is not a develop node either, and `ops/README.md` now records why. It evaluates on the tone curve's own spline rather than a second copy, so a profile author placing a control point and a photographer dragging one mean the same thing by it. The curves are data. `core/dr-decode/profiles/base_curves.yaml` ships inside the binary as a floor and is superseded by any copy on disk carrying a higher `version:`, so a body can be added and distributed without a release — and, under the GPL, contributed. The comparison runs both ways: a stale pack cannot hold an upgraded binary back at last year's rendering. Canon EOS 6D and R6, Nikon Z 6 and D750, Sony A7 III and Fujifilm X-T3 ship with their own curves. Every other body gets a conservative default, which is much closer to right than the identity is for any of them. A JPEG gets none — it has already been rendered once, by the camera. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1c0994c807 |
Demosaic a Fujifilm sensor instead of refusing it
Every RAF stopped at the embedded preview, because the demosaicer had one kernel and it was a Bayer kernel. D11 makes Fujifilm first-class and FR-RAW-5 asks for it by name, so a hard error there was a promise we had not kept. X-Trans is a 6x6 tile, and nothing in the Bayer path survives that: the missing channels sit at different offsets at all 36 positions, so there is no fixed kernel to write. The new shader fits a weighted plane through each channel's samples in a 5x5 window and carries the other two channels across as the difference between those planes, keeping the pixel's own measured value untouched. A plane rather than a mean because the three channels are sampled at different places in the tile: a mean compares a red taken slightly left of the pixel with a green taken slightly right of it, and that offset is a colour cast that follows every gradient in the frame. The fit is done in white-balanced space, where the constant-colour-difference model it rests on is actually true of a neutral subject; that alone halves the error at a luminance edge. Two compromises, both deliberate. It is not Markesteijn. There are no directional hypotheses and no homogeneity map, so it does not resolve detail finer than the CFA period and a hard edge arrives about two pixels wide. It cannot ring — the output is bounded by the local sample range — so it does not produce the worms FR-RAW-5 exists to avoid, but the quality that requirement asks for is still owed. The tile's phase is guessed rather than known. rawler has each body's pattern exactly, as a 36-character string, but CfaPattern::XTrans throws it away before dr-gpu sees the file, and it is not a constant to hard-code: the bodies in that database start the tile at four different origins. So the phase is read back out of the pixels, by grouping the 36 per-position means and taking the grouping with the least spread. That part needs nothing from the scene. Telling red from blue does — shifting the tile by half a tile turns it into itself with red and blue swapped, so no geometry can decide it — and the as-shot white balance is what breaks the tie. A frame that is almost entirely one colour can defeat that; widening dr-decode to carry the pattern string would retire the guess altogether. The tests assert reconstruction, not success: a flat patch comes back exactly at all six phases tested, and a linear ramp comes back exactly too, which is the property the plane fit exists for and the one a mean would fail. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d7aeafaf84 |
Move to wgpu 29, the version Slint can share a device with
Build and test / Desktop (Linux) (push) Failing after 38s
Build and test / Layer separation (push) Successful in 24s
Traceability / Requirement traces (push) Successful in 1m3s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
Build and test / Android (aarch64) (push) Failing after 9m54s
Groundwork for spike S1. Importing a texture into a Slint scene requires it
to come from the *same* `wgpu::Device` Slint renders with, and Slint hands
out a device of the version it was compiled against. Slint 1.17 offers
`unstable-wgpu-28` and `unstable-wgpu-29` and nothing older, so wgpu 23 could
never have met it: two semver-incompatible wgpu crates in one tree are two
distinct types, and the device would not typecheck across the gap.
The version is therefore not a free choice, and the manifest now says so —
Slint and wgpu move together or not at all. The Slint requirement is also
corrected from "1.9" to the 1.17 it has actually been resolving to.
Nothing about the render path changes here. The readback bridge is still in
place and still the display path, so this is verified by the tests that
already existed rather than by anything new: 39 dr-gpu tests, which compare
real pixels off a real device, and 888 across the workspace, all passing.
Zero-copy lands separately and small.
What the six releases cost, in full:
- `ImageCopyTexture`/`ImageCopyBuffer`/`ImageDataLayout` became the
`TexelCopy*` names (24).
- `Instance::new` takes the descriptor by value, and `InstanceDescriptor`
lost its `Default` — it carries a boxed display handle now, so a headless
context says `new_without_display_handle` and means it.
- `request_adapter` returns `Result` rather than `Option` (24).
- `DeviceDescriptor` absorbed the API trace from `request_device`'s second
argument and gained `experimental_features` (25).
- `PipelineLayoutDescriptor` takes `Option<&BindGroupLayout>` per slot, and
`push_constant_ranges` became `immediate_size`.
- `Maintain` became `PollType`, and `poll` is fallible.
Two of those are improvements worth having rather than churn. The error scope
is a guard whose `pop` runs on drop, so an early return from the pipeline
compiler no longer leaves a scope open on the device for whatever ran next to
fall into. And a fallible `poll` reports a lost device (NFR-R7) at the point
it happens, where before the map callback simply never arrived and the
failure surfaced later as a readback that spun out its poll limit.
Still to do for S1: dr-ui renders through `renderer-femtovg`, which is
OpenGL. Texture import needs Slint itself rendering on wgpu.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
5786977a51 |
Develop a JPEG through the same pipeline as a RAW
DemosaicedImage gains a second producer, from_rgba8, alongside the CFA path. Nothing about the type is CFA-specific — it is "an image on the GPU, ready to adjust" — which is what lets develop mode work on a JPEG without the edit graph or any operation knowing the source was not a RAW file. The one real difference is the transfer function: sensor data is linear, a JPEG is gamma-encoded. Every operation assumes linear scene-referred colour (exposure is a multiply, and doubling a gamma-encoded value is not a stop), so the shader prologue linearises once, at the only point where the two source kinds still differ. The flag rides in as_shot_wb.w, which was padding. For a JPEG the white balance uniform is neutral and the colour matrix is identity, so both stay unconditional multiplies rather than becoming branches. max_dimension is exposed because it is a hardware limit the caller must plan around, not a failure to report afterwards: a 13728x8928 film scan exceeds the common 8192 texture limit, and fitting it first is the only way to develop it at all. Assisted-by: LLM |
||
|
|
78e3e6b846 |
Add the develop pipeline: demosaic and seven raw adjustments
Decode through display, on the GPU: black/white normalisation, Bayer demosaic, camera colour transform, and the first seven adjustment operations — white balance, exposure, highlights/shadows, blacks/whites, brilliance, vibrance, saturation. Composable shaders. Each operation contributes a WGSL fragment rather than owning a pass, and dr-pipeline fuses the *active* ones into a single compute shader. One texture read and one write per frame regardless of how many adjustments are in play, while the operations stay independent in Rust — adding one is a new file, with no central shader to edit. An operation at neutral settings contributes no code, no uniform and no branch. Uniforms are prefixed per operation so two may both declare `amount`; helpers dedupe by name from a single source of truth. Pipelines cache on a structure hash covering the op-set and its order but not the values, so dragging a slider uploads uniforms and reuses the compiled pipeline. Measured on a 24 MP CR2: 0.60 ms re-render, one pipeline compiled across ten slider positions. The UI is generated, not written. EditGraph::capabilities() reports parameters with their kinds, ranges, defaults and current values; the panel builds one control per entry chosen by ParamKind. No file in ui/ names an operation, and dr-pipeline has no wgpu dependency, so codegen is testable without a device (ARCH §6.5a). Three defects found against real files, each silent: - rawler 0.7.2's `xyz_to_cam` is all zeros — deprecated and no longer populated. The live matrices are in `color_matrix`, keyed by illuminant. Reading the old field yields no colour transform at all. - `cam_to_xyz_normalized()` returns all NaN on any Bayer sensor: it divides each of four rows by its own sum, and the unused fourth (emerald) row sums to zero. Inverting the 3x3 ourselves avoids it. `wb_coeffs[3]` is NaN for the same reason and is normalised at decode. - As-shot white balance reached the uniform block but no shader read it, so the first render of a real CR2 came out violently green. Green photosites collect roughly twice the signal of red and blue. Now applied unconditionally before any operation, with tests on ordering. Demosaic is Malvar-He-Cutler rather than bilinear: gradient-corrected interpolation at one 5x5 neighbourhood per pixel, where bilinear leaves visible zippering on any high-contrast edge at 1:1. Two of the four packed CFA constants were wrong on the first attempt, so all four layouts are asserted to reconstruct the same colour. Crop origins at odd coordinates re-phase the pattern; without that, red and blue swap. X-Trans reports GpuError::UnsupportedCfa rather than approximating with the Bayer path, which would look like a corrupt file. 206 tests, including GPU tests proving every operation and the full seven-operation chain generate compilable WGSL. Known gaps: the display path still reads back to the CPU each frame, which ARCH §6.1 forbids and AC-8 asserts against — it is gated behind the `readback` feature and waits on spike S1 wiring Slint's texture import. Curve shapes are a first draft and want tuning against real photographs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |