Add the develop pipeline: demosaic and seven raw adjustments
Decode through display, on the GPU: black/white normalisation, Bayer demosaic, camera colour transform, and the first seven adjustment operations — white balance, exposure, highlights/shadows, blacks/whites, brilliance, vibrance, saturation. Composable shaders. Each operation contributes a WGSL fragment rather than owning a pass, and dr-pipeline fuses the *active* ones into a single compute shader. One texture read and one write per frame regardless of how many adjustments are in play, while the operations stay independent in Rust — adding one is a new file, with no central shader to edit. An operation at neutral settings contributes no code, no uniform and no branch. Uniforms are prefixed per operation so two may both declare `amount`; helpers dedupe by name from a single source of truth. Pipelines cache on a structure hash covering the op-set and its order but not the values, so dragging a slider uploads uniforms and reuses the compiled pipeline. Measured on a 24 MP CR2: 0.60 ms re-render, one pipeline compiled across ten slider positions. The UI is generated, not written. EditGraph::capabilities() reports parameters with their kinds, ranges, defaults and current values; the panel builds one control per entry chosen by ParamKind. No file in ui/ names an operation, and dr-pipeline has no wgpu dependency, so codegen is testable without a device (ARCH §6.5a). Three defects found against real files, each silent: - rawler 0.7.2's `xyz_to_cam` is all zeros — deprecated and no longer populated. The live matrices are in `color_matrix`, keyed by illuminant. Reading the old field yields no colour transform at all. - `cam_to_xyz_normalized()` returns all NaN on any Bayer sensor: it divides each of four rows by its own sum, and the unused fourth (emerald) row sums to zero. Inverting the 3x3 ourselves avoids it. `wb_coeffs[3]` is NaN for the same reason and is normalised at decode. - As-shot white balance reached the uniform block but no shader read it, so the first render of a real CR2 came out violently green. Green photosites collect roughly twice the signal of red and blue. Now applied unconditionally before any operation, with tests on ordering. Demosaic is Malvar-He-Cutler rather than bilinear: gradient-corrected interpolation at one 5x5 neighbourhood per pixel, where bilinear leaves visible zippering on any high-contrast edge at 1:1. Two of the four packed CFA constants were wrong on the first attempt, so all four layouts are asserted to reconstruct the same colour. Crop origins at odd coordinates re-phase the pattern; without that, red and blue swap. X-Trans reports GpuError::UnsupportedCfa rather than approximating with the Bayer path, which would look like a corrupt file. 206 tests, including GPU tests proving every operation and the full seven-operation chain generate compilable WGSL. Known gaps: the display path still reads back to the CPU each frame, which ARCH §6.1 forbids and AC-8 asserts against — it is gated behind the `readback` feature and waits on spike S1 wiring Slint's texture import. Curve shapes are a first draft and want tuning against real photographs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,768 @@
|
||||
//! The adjust pass — runs `dr-pipeline`'s generated shader.
|
||||
//!
|
||||
//! Takes the demosaiced texture, applies the composed operation chain, and
|
||||
//! writes a display-ready RGBA8 texture. One dispatch, whatever the number of
|
||||
//! active operations, because the operations were fused into one shader
|
||||
//! before they got here.
|
||||
//!
|
||||
//! # The pipeline cache
|
||||
//!
|
||||
//! Compiling a shader takes milliseconds — fine once, ruinous per frame while
|
||||
//! a slider is moving. Pipelines are therefore cached by the composed
|
||||
//! shader's `structure_hash`, which covers the operation set and their order
|
||||
//! but not their values. Dragging a slider re-uploads a uniform buffer and
|
||||
//! reuses the compiled pipeline; enabling an operation compiles once and then
|
||||
//! also reuses.
|
||||
|
||||
use std::collections::HashMap;
|
||||
|
||||
use dr_pipeline::ComposedShader;
|
||||
use wgpu::util::DeviceExt;
|
||||
|
||||
use crate::{DemosaicedImage, GpuContext, GpuError};
|
||||
|
||||
/// Number of leading floats in the generated uniform block that the composer
|
||||
/// reserves for base parameters — three padded matrix rows and the as-shot
|
||||
/// white balance. Must match `BASE_UNIFORM_FIELDS` in dr-pipeline.
|
||||
const BASE_FIELDS: usize = 16;
|
||||
|
||||
/// Runs composed operation chains against demosaiced images.
|
||||
pub struct AdjustPass {
|
||||
ctx: GpuContext,
|
||||
bind_group_layout: wgpu::BindGroupLayout,
|
||||
pipeline_layout: wgpu::PipelineLayout,
|
||||
/// Compiled pipelines by structure hash (ARCH §5.6).
|
||||
cache: HashMap<u64, wgpu::ComputePipeline>,
|
||||
/// Output texture, reallocated only when the size changes.
|
||||
target: Option<Target>,
|
||||
}
|
||||
|
||||
struct Target {
|
||||
texture: wgpu::Texture,
|
||||
view: wgpu::TextureView,
|
||||
width: u32,
|
||||
height: u32,
|
||||
}
|
||||
|
||||
impl AdjustPass {
|
||||
pub const FORMAT: wgpu::TextureFormat = wgpu::TextureFormat::Rgba8Unorm;
|
||||
|
||||
pub fn new(ctx: &GpuContext) -> Self {
|
||||
let bind_group_layout =
|
||||
ctx.device
|
||||
.create_bind_group_layout(&wgpu::BindGroupLayoutDescriptor {
|
||||
label: Some("adjust-bgl"),
|
||||
entries: &[
|
||||
// The demosaiced source.
|
||||
wgpu::BindGroupLayoutEntry {
|
||||
binding: 0,
|
||||
visibility: wgpu::ShaderStages::COMPUTE,
|
||||
ty: wgpu::BindingType::Texture {
|
||||
sample_type: wgpu::TextureSampleType::Float { filterable: true },
|
||||
view_dimension: wgpu::TextureViewDimension::D2,
|
||||
multisampled: false,
|
||||
},
|
||||
count: None,
|
||||
},
|
||||
wgpu::BindGroupLayoutEntry {
|
||||
binding: 1,
|
||||
visibility: wgpu::ShaderStages::COMPUTE,
|
||||
ty: wgpu::BindingType::Buffer {
|
||||
ty: wgpu::BufferBindingType::Uniform,
|
||||
has_dynamic_offset: false,
|
||||
min_binding_size: None,
|
||||
},
|
||||
count: None,
|
||||
},
|
||||
wgpu::BindGroupLayoutEntry {
|
||||
binding: 2,
|
||||
visibility: wgpu::ShaderStages::COMPUTE,
|
||||
ty: wgpu::BindingType::StorageTexture {
|
||||
access: wgpu::StorageTextureAccess::WriteOnly,
|
||||
format: Self::FORMAT,
|
||||
view_dimension: wgpu::TextureViewDimension::D2,
|
||||
},
|
||||
count: None,
|
||||
},
|
||||
],
|
||||
});
|
||||
|
||||
let pipeline_layout = ctx
|
||||
.device
|
||||
.create_pipeline_layout(&wgpu::PipelineLayoutDescriptor {
|
||||
label: Some("adjust-layout"),
|
||||
bind_group_layouts: &[&bind_group_layout],
|
||||
push_constant_ranges: &[],
|
||||
});
|
||||
|
||||
Self {
|
||||
ctx: ctx.clone(),
|
||||
bind_group_layout,
|
||||
pipeline_layout,
|
||||
cache: HashMap::new(),
|
||||
target: None,
|
||||
}
|
||||
}
|
||||
|
||||
/// Compile a composed shader, or return the cached pipeline.
|
||||
///
|
||||
/// Compilation errors carry the generated source, since a stray line
|
||||
/// number against code nobody wrote is otherwise very hard to act on.
|
||||
fn pipeline(&mut self, shader: &ComposedShader) -> Result<&wgpu::ComputePipeline, GpuError> {
|
||||
if !self.cache.contains_key(&shader.structure_hash) {
|
||||
// A validation error here is a codegen bug, not a user error.
|
||||
// Push an error scope so it surfaces as a Result rather than a
|
||||
// panic from wgpu's default handler.
|
||||
self.ctx
|
||||
.device
|
||||
.push_error_scope(wgpu::ErrorFilter::Validation);
|
||||
|
||||
let module = self
|
||||
.ctx
|
||||
.device
|
||||
.create_shader_module(wgpu::ShaderModuleDescriptor {
|
||||
label: Some("adjust-generated"),
|
||||
source: wgpu::ShaderSource::Wgsl(shader.source.as_str().into()),
|
||||
});
|
||||
|
||||
let pipeline =
|
||||
self.ctx
|
||||
.device
|
||||
.create_compute_pipeline(&wgpu::ComputePipelineDescriptor {
|
||||
label: Some("adjust-pipeline"),
|
||||
layout: Some(&self.pipeline_layout),
|
||||
module: &module,
|
||||
entry_point: Some("main"),
|
||||
compilation_options: Default::default(),
|
||||
cache: None,
|
||||
});
|
||||
|
||||
if let Some(err) = pollster::block_on(self.ctx.device.pop_error_scope()) {
|
||||
return Err(GpuError::ShaderCompilation(format!(
|
||||
"{err}\n\n--- generated source ---\n{}",
|
||||
numbered(&shader.source)
|
||||
)));
|
||||
}
|
||||
|
||||
self.cache.insert(shader.structure_hash, pipeline);
|
||||
}
|
||||
|
||||
Ok(self
|
||||
.cache
|
||||
.get(&shader.structure_hash)
|
||||
.expect("just inserted"))
|
||||
}
|
||||
|
||||
/// Ensure the output texture matches the requested size.
|
||||
fn ensure_target(&mut self, width: u32, height: u32) {
|
||||
let matches = self
|
||||
.target
|
||||
.as_ref()
|
||||
.is_some_and(|t| t.width == width && t.height == height);
|
||||
if matches {
|
||||
return;
|
||||
}
|
||||
|
||||
let texture = self.ctx.device.create_texture(&wgpu::TextureDescriptor {
|
||||
label: Some("adjust-output"),
|
||||
size: wgpu::Extent3d {
|
||||
width,
|
||||
height,
|
||||
depth_or_array_layers: 1,
|
||||
},
|
||||
mip_level_count: 1,
|
||||
sample_count: 1,
|
||||
dimension: wgpu::TextureDimension::D2,
|
||||
format: Self::FORMAT,
|
||||
usage: wgpu::TextureUsages::STORAGE_BINDING
|
||||
| wgpu::TextureUsages::TEXTURE_BINDING
|
||||
| wgpu::TextureUsages::COPY_SRC,
|
||||
view_formats: &[],
|
||||
});
|
||||
let view = texture.create_view(&Default::default());
|
||||
self.target = Some(Target {
|
||||
texture,
|
||||
view,
|
||||
width,
|
||||
height,
|
||||
});
|
||||
}
|
||||
|
||||
/// Render one frame at the requested output size.
|
||||
///
|
||||
/// `width`/`height` are the *display* size, which is normally far smaller
|
||||
/// than the image. Rendering at viewport resolution rather than sensor
|
||||
/// resolution is what keeps slider interaction inside the frame budget
|
||||
/// (FR-DSP-1).
|
||||
pub fn render(
|
||||
&mut self,
|
||||
source: &DemosaicedImage,
|
||||
shader: &ComposedShader,
|
||||
width: u32,
|
||||
height: u32,
|
||||
) -> Result<&wgpu::Texture, GpuError> {
|
||||
let (width, height) = (width.max(1), height.max(1));
|
||||
self.ensure_target(width, height);
|
||||
|
||||
// Base uniforms: the camera matrix and as-shot white balance, which
|
||||
// every generated shader reads regardless of which operations are
|
||||
// active.
|
||||
let mut uniforms = shader.uniforms.clone();
|
||||
if uniforms.len() < BASE_FIELDS {
|
||||
uniforms.resize(BASE_FIELDS, 0.0);
|
||||
}
|
||||
let m = source.color_matrix();
|
||||
let wb = source.as_shot_wb();
|
||||
// Rows padded to vec4 for std140 alignment.
|
||||
uniforms[0..4].copy_from_slice(&[m[0], m[1], m[2], 0.0]);
|
||||
uniforms[4..8].copy_from_slice(&[m[3], m[4], m[5], 0.0]);
|
||||
uniforms[8..12].copy_from_slice(&[m[6], m[7], m[8], 0.0]);
|
||||
uniforms[12..16].copy_from_slice(&[wb[0], wb[1], wb[2], 0.0]);
|
||||
|
||||
let params_buf = self
|
||||
.ctx
|
||||
.device
|
||||
.create_buffer_init(&wgpu::util::BufferInitDescriptor {
|
||||
label: Some("adjust-params"),
|
||||
contents: bytemuck::cast_slice(&uniforms),
|
||||
usage: wgpu::BufferUsages::UNIFORM,
|
||||
});
|
||||
|
||||
// Borrow order: compile first, since `pipeline` takes &mut self.
|
||||
let _ = self.pipeline(shader)?;
|
||||
let pipeline = self
|
||||
.cache
|
||||
.get(&shader.structure_hash)
|
||||
.expect("compiled above");
|
||||
let target = self.target.as_ref().expect("ensured above");
|
||||
|
||||
let bind_group = self
|
||||
.ctx
|
||||
.device
|
||||
.create_bind_group(&wgpu::BindGroupDescriptor {
|
||||
label: Some("adjust-bg"),
|
||||
layout: &self.bind_group_layout,
|
||||
entries: &[
|
||||
wgpu::BindGroupEntry {
|
||||
binding: 0,
|
||||
resource: wgpu::BindingResource::TextureView(source.view()),
|
||||
},
|
||||
wgpu::BindGroupEntry {
|
||||
binding: 1,
|
||||
resource: params_buf.as_entire_binding(),
|
||||
},
|
||||
wgpu::BindGroupEntry {
|
||||
binding: 2,
|
||||
resource: wgpu::BindingResource::TextureView(&target.view),
|
||||
},
|
||||
],
|
||||
});
|
||||
|
||||
let mut enc = self
|
||||
.ctx
|
||||
.device
|
||||
.create_command_encoder(&wgpu::CommandEncoderDescriptor {
|
||||
label: Some("adjust-encoder"),
|
||||
});
|
||||
{
|
||||
let mut pass = enc.begin_compute_pass(&wgpu::ComputePassDescriptor {
|
||||
label: Some("adjust-pass"),
|
||||
timestamp_writes: None,
|
||||
});
|
||||
pass.set_pipeline(pipeline);
|
||||
pass.set_bind_group(0, &bind_group, &[]);
|
||||
pass.dispatch_workgroups(width.div_ceil(8), height.div_ceil(8), 1);
|
||||
}
|
||||
self.ctx.queue.submit(Some(enc.finish()));
|
||||
|
||||
Ok(&self.target.as_ref().expect("ensured above").texture)
|
||||
}
|
||||
|
||||
/// How many distinct pipelines are compiled. Exposed for tests asserting
|
||||
/// that slider movement does not recompile.
|
||||
pub fn cached_pipelines(&self) -> usize {
|
||||
self.cache.len()
|
||||
}
|
||||
|
||||
pub fn output(&self) -> Option<&wgpu::Texture> {
|
||||
self.target.as_ref().map(|t| &t.texture)
|
||||
}
|
||||
|
||||
/// Copy the output to the CPU as tightly packed RGBA8.
|
||||
///
|
||||
/// **A temporary bridge, not the display path.** ARCH §6.1 forbids this
|
||||
/// round-trip in production and AC-8 asserts it does not happen; it
|
||||
/// exists only because Slint's texture-import path is unwired until
|
||||
/// spike S1. Measured cost at 4K is ~7 ms against a 0.28 ms compute pass
|
||||
/// — 96% of the frame — so this must go, and the `readback` feature gate
|
||||
/// keeps it out of a shipping build.
|
||||
#[cfg(any(test, feature = "readback"))]
|
||||
pub fn read_output(&self) -> Result<(Vec<u8>, u32, u32), GpuError> {
|
||||
let Some(target) = self.target.as_ref() else {
|
||||
return Err(GpuError::Readback("nothing rendered yet".into()));
|
||||
};
|
||||
let (w, h) = (target.width, target.height);
|
||||
|
||||
let unpadded = w * 4;
|
||||
let align = wgpu::COPY_BYTES_PER_ROW_ALIGNMENT;
|
||||
let padded = unpadded.div_ceil(align) * align;
|
||||
|
||||
let buf = self.ctx.device.create_buffer(&wgpu::BufferDescriptor {
|
||||
label: Some("adjust-readback"),
|
||||
size: (padded * h) as u64,
|
||||
usage: wgpu::BufferUsages::COPY_DST | wgpu::BufferUsages::MAP_READ,
|
||||
mapped_at_creation: false,
|
||||
});
|
||||
|
||||
let mut enc = self.ctx.device.create_command_encoder(&Default::default());
|
||||
enc.copy_texture_to_buffer(
|
||||
wgpu::ImageCopyTexture {
|
||||
texture: &target.texture,
|
||||
mip_level: 0,
|
||||
origin: wgpu::Origin3d::ZERO,
|
||||
aspect: wgpu::TextureAspect::All,
|
||||
},
|
||||
wgpu::ImageCopyBuffer {
|
||||
buffer: &buf,
|
||||
layout: wgpu::ImageDataLayout {
|
||||
offset: 0,
|
||||
bytes_per_row: Some(padded),
|
||||
rows_per_image: Some(h),
|
||||
},
|
||||
},
|
||||
wgpu::Extent3d {
|
||||
width: w,
|
||||
height: h,
|
||||
depth_or_array_layers: 1,
|
||||
},
|
||||
);
|
||||
self.ctx.queue.submit(Some(enc.finish()));
|
||||
|
||||
let slice = buf.slice(..);
|
||||
let (tx, rx) = std::sync::mpsc::channel();
|
||||
slice.map_async(wgpu::MapMode::Read, move |r| {
|
||||
let _ = tx.send(r);
|
||||
});
|
||||
self.ctx.device.poll(wgpu::Maintain::Wait);
|
||||
rx.recv()
|
||||
.map_err(|e| GpuError::Readback(e.to_string()))?
|
||||
.map_err(|e| GpuError::Readback(e.to_string()))?;
|
||||
|
||||
let data = slice.get_mapped_range();
|
||||
let mut out = Vec::with_capacity((unpadded * h) as usize);
|
||||
for row in 0..h {
|
||||
let start = (row * padded) as usize;
|
||||
out.extend_from_slice(&data[start..start + unpadded as usize]);
|
||||
}
|
||||
drop(data);
|
||||
buf.unmap();
|
||||
Ok((out, w, h))
|
||||
}
|
||||
}
|
||||
|
||||
/// Number the lines of generated source, so a compiler error can be located.
|
||||
fn numbered(src: &str) -> String {
|
||||
src.lines()
|
||||
.enumerate()
|
||||
.map(|(i, l)| format!("{:>4} | {l}", i + 1))
|
||||
.collect::<Vec<_>>()
|
||||
.join("\n")
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
use dr_decode::{CfaPattern, CropRect, RawImage};
|
||||
use dr_pipeline::ops::{colour, exposure, tone, white_balance};
|
||||
use dr_pipeline::EditGraph;
|
||||
|
||||
use crate::Demosaicer;
|
||||
|
||||
fn ctx() -> Option<GpuContext> {
|
||||
match pollster::block_on(GpuContext::new_headless()) {
|
||||
Ok(c) => Some(c),
|
||||
Err(e) => {
|
||||
eprintln!("skipping: no GPU adapter ({e})");
|
||||
None
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// A flat mid-grey image, so an operation's effect is unambiguous.
|
||||
fn grey_image(ctx: &GpuContext, level: u16) -> DemosaicedImage {
|
||||
let size = 16u32;
|
||||
let mut data = vec![0u16; (size * size) as usize];
|
||||
for v in data.iter_mut() {
|
||||
*v = level;
|
||||
}
|
||||
let raw = RawImage {
|
||||
width: size,
|
||||
height: size,
|
||||
data,
|
||||
cfa_pattern: CfaPattern::Rggb,
|
||||
black_level: [0; 4],
|
||||
white_level: 16383,
|
||||
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
|
||||
// Identity, so the test reasons about the operations alone
|
||||
// rather than about a camera's colour response.
|
||||
color_matrix: Some([1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0]),
|
||||
crop: CropRect {
|
||||
x: 0,
|
||||
y: 0,
|
||||
width: size,
|
||||
height: size,
|
||||
},
|
||||
};
|
||||
Demosaicer::new(ctx)
|
||||
.expect("demosaicer")
|
||||
.run(&raw)
|
||||
.expect("demosaic")
|
||||
}
|
||||
|
||||
fn read_centre(ctx: &GpuContext, tex: &wgpu::Texture) -> [u8; 4] {
|
||||
let w = tex.width();
|
||||
let h = tex.height();
|
||||
let unpadded = w * 4;
|
||||
let align = wgpu::COPY_BYTES_PER_ROW_ALIGNMENT;
|
||||
let padded = unpadded.div_ceil(align) * align;
|
||||
|
||||
let buf = ctx.device.create_buffer(&wgpu::BufferDescriptor {
|
||||
label: Some("adjust-readback"),
|
||||
size: (padded * h) as u64,
|
||||
usage: wgpu::BufferUsages::COPY_DST | wgpu::BufferUsages::MAP_READ,
|
||||
mapped_at_creation: false,
|
||||
});
|
||||
|
||||
let mut enc = ctx.device.create_command_encoder(&Default::default());
|
||||
enc.copy_texture_to_buffer(
|
||||
wgpu::ImageCopyTexture {
|
||||
texture: tex,
|
||||
mip_level: 0,
|
||||
origin: wgpu::Origin3d::ZERO,
|
||||
aspect: wgpu::TextureAspect::All,
|
||||
},
|
||||
wgpu::ImageCopyBuffer {
|
||||
buffer: &buf,
|
||||
layout: wgpu::ImageDataLayout {
|
||||
offset: 0,
|
||||
bytes_per_row: Some(padded),
|
||||
rows_per_image: Some(h),
|
||||
},
|
||||
},
|
||||
wgpu::Extent3d {
|
||||
width: w,
|
||||
height: h,
|
||||
depth_or_array_layers: 1,
|
||||
},
|
||||
);
|
||||
ctx.queue.submit(Some(enc.finish()));
|
||||
|
||||
let slice = buf.slice(..);
|
||||
let (tx, rx) = std::sync::mpsc::channel();
|
||||
slice.map_async(wgpu::MapMode::Read, move |r| {
|
||||
let _ = tx.send(r);
|
||||
});
|
||||
ctx.device.poll(wgpu::Maintain::Wait);
|
||||
rx.recv().expect("map").expect("map ok");
|
||||
|
||||
let data = slice.get_mapped_range();
|
||||
let off = ((h / 2) * padded + (w / 2) * 4) as usize;
|
||||
let px = [data[off], data[off + 1], data[off + 2], data[off + 3]];
|
||||
drop(data);
|
||||
buf.unmap();
|
||||
px
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_neutral_graph_produces_a_compilable_shader() {
|
||||
// The first thing that could go wrong with codegen: the empty case.
|
||||
let Some(ctx) = ctx() else { return };
|
||||
let mut pass = AdjustPass::new(&ctx);
|
||||
let img = grey_image(&ctx, 4000);
|
||||
let shader = EditGraph::default_chain().compose();
|
||||
|
||||
pass.render(&img, &shader, 16, 16)
|
||||
.expect("a neutral chain must compile");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn every_operation_generates_compilable_wgsl() {
|
||||
// The test that justifies the whole codegen approach. Each operation
|
||||
// is compiled on its own, so a WGSL error names the operation that
|
||||
// caused it rather than surfacing only in some combination.
|
||||
let Some(ctx) = ctx() else { return };
|
||||
let mut pass = AdjustPass::new(&ctx);
|
||||
let img = grey_image(&ctx, 4000);
|
||||
|
||||
// A name paired with the edit that activates one operation. Boxed
|
||||
// closures rather than a type alias: the list is read top to bottom
|
||||
// as a table of what is covered.
|
||||
#[allow(clippy::type_complexity)]
|
||||
let cases: Vec<(&str, Box<dyn Fn(&mut EditGraph)>)> = vec![
|
||||
(
|
||||
"white_balance.temperature",
|
||||
Box::new(|g: &mut EditGraph| {
|
||||
g.set_param(white_balance::ID, white_balance::TEMPERATURE, 60.0)
|
||||
}),
|
||||
),
|
||||
(
|
||||
"white_balance.tint",
|
||||
Box::new(|g: &mut EditGraph| {
|
||||
g.set_param(white_balance::ID, white_balance::TINT, -40.0)
|
||||
}),
|
||||
),
|
||||
(
|
||||
"exposure",
|
||||
Box::new(|g: &mut EditGraph| g.set_param(exposure::ID, exposure::EXPOSURE, 1.5)),
|
||||
),
|
||||
(
|
||||
"highlights",
|
||||
Box::new(|g: &mut EditGraph| {
|
||||
g.set_param(tone::HIGHLIGHTS_SHADOWS_ID, tone::HIGHLIGHTS, -70.0)
|
||||
}),
|
||||
),
|
||||
(
|
||||
"shadows",
|
||||
Box::new(|g: &mut EditGraph| {
|
||||
g.set_param(tone::HIGHLIGHTS_SHADOWS_ID, tone::SHADOWS, 70.0)
|
||||
}),
|
||||
),
|
||||
(
|
||||
"blacks",
|
||||
Box::new(|g: &mut EditGraph| {
|
||||
g.set_param(tone::BLACKS_WHITES_ID, tone::BLACKS, -50.0)
|
||||
}),
|
||||
),
|
||||
(
|
||||
"whites",
|
||||
Box::new(|g: &mut EditGraph| {
|
||||
g.set_param(tone::BLACKS_WHITES_ID, tone::WHITES, 50.0)
|
||||
}),
|
||||
),
|
||||
(
|
||||
"brilliance",
|
||||
Box::new(|g: &mut EditGraph| {
|
||||
g.set_param(colour::BRILLIANCE_ID, colour::BRILLIANCE, 60.0)
|
||||
}),
|
||||
),
|
||||
(
|
||||
"vibrance",
|
||||
Box::new(|g: &mut EditGraph| {
|
||||
g.set_param(colour::VIBRANCE_ID, colour::VIBRANCE, 60.0)
|
||||
}),
|
||||
),
|
||||
(
|
||||
"saturation",
|
||||
Box::new(|g: &mut EditGraph| {
|
||||
g.set_param(colour::SATURATION_ID, colour::SATURATION, 60.0)
|
||||
}),
|
||||
),
|
||||
];
|
||||
|
||||
for (name, apply) in cases {
|
||||
let mut g = EditGraph::default_chain();
|
||||
apply(&mut g);
|
||||
let shader = g.compose();
|
||||
pass.render(&img, &shader, 16, 16)
|
||||
.unwrap_or_else(|e| panic!("{name} generated invalid WGSL:\n{e}"));
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn the_whole_chain_at_once_compiles() {
|
||||
// Individually-valid fragments can still collide when combined —
|
||||
// duplicate helpers, clashing locals, a malformed uniform block.
|
||||
let Some(ctx) = ctx() else { return };
|
||||
let mut pass = AdjustPass::new(&ctx);
|
||||
let img = grey_image(&ctx, 4000);
|
||||
|
||||
let mut g = EditGraph::default_chain();
|
||||
g.set_param(white_balance::ID, white_balance::TEMPERATURE, 30.0);
|
||||
g.set_param(white_balance::ID, white_balance::TINT, -20.0);
|
||||
g.set_param(exposure::ID, exposure::EXPOSURE, 0.8);
|
||||
g.set_param(tone::HIGHLIGHTS_SHADOWS_ID, tone::HIGHLIGHTS, -60.0);
|
||||
g.set_param(tone::HIGHLIGHTS_SHADOWS_ID, tone::SHADOWS, 40.0);
|
||||
g.set_param(tone::BLACKS_WHITES_ID, tone::BLACKS, -25.0);
|
||||
g.set_param(tone::BLACKS_WHITES_ID, tone::WHITES, 35.0);
|
||||
g.set_param(colour::BRILLIANCE_ID, colour::BRILLIANCE, 45.0);
|
||||
g.set_param(colour::VIBRANCE_ID, colour::VIBRANCE, 55.0);
|
||||
g.set_param(colour::SATURATION_ID, colour::SATURATION, 15.0);
|
||||
|
||||
let shader = g.compose();
|
||||
assert_eq!(shader.source.matches("---- ").count(), 7);
|
||||
pass.render(&img, &shader, 32, 32)
|
||||
.expect("the full chain must compile");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn exposure_brightens_the_image() {
|
||||
// Proves the uniforms actually reach the shader, not merely that it
|
||||
// compiles.
|
||||
let Some(ctx) = ctx() else { return };
|
||||
let mut pass = AdjustPass::new(&ctx);
|
||||
let img = grey_image(&ctx, 2000);
|
||||
|
||||
let neutral = EditGraph::default_chain().compose();
|
||||
let before = {
|
||||
let t = pass.render(&img, &neutral, 16, 16).expect("render");
|
||||
read_centre(&ctx, t)
|
||||
};
|
||||
|
||||
let mut g = EditGraph::default_chain();
|
||||
g.set_param(exposure::ID, exposure::EXPOSURE, 2.0);
|
||||
let brighter = g.compose();
|
||||
let after = {
|
||||
let t = pass.render(&img, &brighter, 16, 16).expect("render");
|
||||
read_centre(&ctx, t)
|
||||
};
|
||||
|
||||
assert!(
|
||||
after[0] > before[0],
|
||||
"+2 stops should brighten: {before:?} -> {after:?}"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn negative_exposure_darkens_the_image() {
|
||||
let Some(ctx) = ctx() else { return };
|
||||
let mut pass = AdjustPass::new(&ctx);
|
||||
let img = grey_image(&ctx, 8000);
|
||||
|
||||
let neutral = EditGraph::default_chain().compose();
|
||||
let before = {
|
||||
let t = pass.render(&img, &neutral, 16, 16).expect("render");
|
||||
read_centre(&ctx, t)
|
||||
};
|
||||
|
||||
let mut g = EditGraph::default_chain();
|
||||
g.set_param(exposure::ID, exposure::EXPOSURE, -2.0);
|
||||
let darker = g.compose();
|
||||
let after = {
|
||||
let t = pass.render(&img, &darker, 16, 16).expect("render");
|
||||
read_centre(&ctx, t)
|
||||
};
|
||||
|
||||
assert!(after[0] < before[0], "-2 stops should darken");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn full_negative_saturation_produces_grey() {
|
||||
// A neutral grey source cannot show this, so use a coloured one:
|
||||
// a strongly red-weighted image must come out with equal channels.
|
||||
let Some(ctx) = ctx() else { return };
|
||||
let mut pass = AdjustPass::new(&ctx);
|
||||
|
||||
let size = 16u32;
|
||||
let mut data = vec![0u16; (size * size) as usize];
|
||||
for y in 0..size {
|
||||
for x in 0..size {
|
||||
// RGGB: make red photosites bright, others dim.
|
||||
let c = CfaPattern::Rggb.colour_at(x, y);
|
||||
data[(y * size + x) as usize] = if c == 0 { 12000 } else { 3000 };
|
||||
}
|
||||
}
|
||||
let raw = RawImage {
|
||||
width: size,
|
||||
height: size,
|
||||
data,
|
||||
cfa_pattern: CfaPattern::Rggb,
|
||||
black_level: [0; 4],
|
||||
white_level: 16383,
|
||||
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
|
||||
color_matrix: Some([1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0]),
|
||||
crop: CropRect {
|
||||
x: 0,
|
||||
y: 0,
|
||||
width: size,
|
||||
height: size,
|
||||
},
|
||||
};
|
||||
let img = Demosaicer::new(&ctx)
|
||||
.expect("demosaicer")
|
||||
.run(&raw)
|
||||
.expect("demosaic");
|
||||
|
||||
let mut g = EditGraph::default_chain();
|
||||
g.set_param(colour::SATURATION_ID, colour::SATURATION, -100.0);
|
||||
let shader = g.compose();
|
||||
let px = {
|
||||
let t = pass.render(&img, &shader, 16, 16).expect("render");
|
||||
read_centre(&ctx, t)
|
||||
};
|
||||
|
||||
let spread = px[0].abs_diff(px[1]).max(px[1].abs_diff(px[2]));
|
||||
assert!(
|
||||
spread <= 2,
|
||||
"-100 saturation must produce grey, got {px:?} (spread {spread})"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn moving_a_slider_does_not_recompile() {
|
||||
// The property the pipeline cache exists for. Recompiling per frame
|
||||
// would make slider interaction unusable regardless of shader cost.
|
||||
let Some(ctx) = ctx() else { return };
|
||||
let mut pass = AdjustPass::new(&ctx);
|
||||
let img = grey_image(&ctx, 4000);
|
||||
|
||||
let mut g = EditGraph::default_chain();
|
||||
for i in 1..=10 {
|
||||
g.set_param(exposure::ID, exposure::EXPOSURE, i as f32 * 0.2);
|
||||
let shader = g.compose();
|
||||
pass.render(&img, &shader, 16, 16).expect("render");
|
||||
}
|
||||
|
||||
assert_eq!(
|
||||
pass.cached_pipelines(),
|
||||
1,
|
||||
"ten slider positions must share one compiled pipeline"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_different_operation_set_compiles_its_own_pipeline() {
|
||||
let Some(ctx) = ctx() else { return };
|
||||
let mut pass = AdjustPass::new(&ctx);
|
||||
let img = grey_image(&ctx, 4000);
|
||||
|
||||
let mut g = EditGraph::default_chain();
|
||||
g.set_param(exposure::ID, exposure::EXPOSURE, 1.0);
|
||||
pass.render(&img, &g.compose(), 16, 16).expect("render");
|
||||
assert_eq!(pass.cached_pipelines(), 1);
|
||||
|
||||
g.set_param(colour::SATURATION_ID, colour::SATURATION, 40.0);
|
||||
pass.render(&img, &g.compose(), 16, 16).expect("render");
|
||||
assert_eq!(pass.cached_pipelines(), 2);
|
||||
|
||||
// Returning to the earlier state must reuse, not compile a third.
|
||||
g.set_param(colour::SATURATION_ID, colour::SATURATION, 0.0);
|
||||
pass.render(&img, &g.compose(), 16, 16).expect("render");
|
||||
assert_eq!(pass.cached_pipelines(), 2);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn output_is_opaque_everywhere() {
|
||||
// A zero alpha would composite as an invisible image, which reads as
|
||||
// "nothing rendered" rather than as a bug in this pass.
|
||||
let Some(ctx) = ctx() else { return };
|
||||
let mut pass = AdjustPass::new(&ctx);
|
||||
let img = grey_image(&ctx, 4000);
|
||||
let shader = EditGraph::default_chain().compose();
|
||||
let t = pass.render(&img, &shader, 16, 16).expect("render");
|
||||
assert_eq!(read_centre(&ctx, t)[3], 255);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn the_output_resizes_with_the_viewport() {
|
||||
let Some(ctx) = ctx() else { return };
|
||||
let mut pass = AdjustPass::new(&ctx);
|
||||
let img = grey_image(&ctx, 4000);
|
||||
let shader = EditGraph::default_chain().compose();
|
||||
|
||||
let t = pass.render(&img, &shader, 64, 48).expect("render");
|
||||
assert_eq!((t.width(), t.height()), (64, 48));
|
||||
|
||||
let t = pass.render(&img, &shader, 32, 96).expect("render");
|
||||
assert_eq!((t.width(), t.height()), (32, 96));
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,749 @@
|
||||
//! Raw upload, black/white normalisation, and Bayer demosaic.
|
||||
//!
|
||||
//! The first real pipeline stage (ARCH §5.2). It takes CFA sensor data from
|
||||
//! `dr-decode`, uploads it once, and produces a linear scene-referred
|
||||
//! RGBA16Float texture in *camera* colour space. Everything downstream — the
|
||||
//! camera matrix, white balance, the tone operations — works on that texture
|
||||
//! and never sees the CFA pattern.
|
||||
//!
|
||||
//! Uploaded once per image, not per frame. Moving a slider re-runs the adjust
|
||||
//! pass over this texture; it does not re-demosaic, which is what keeps the
|
||||
//! interaction budget (NFR-P9) reachable on a 24 MP file.
|
||||
|
||||
use dr_decode::{CfaPattern, RawImage};
|
||||
use wgpu::util::DeviceExt;
|
||||
|
||||
use crate::{GpuContext, GpuError};
|
||||
|
||||
/// Uniform block for the demosaic pass. Layout must match `demosaic.wgsl`.
|
||||
#[repr(C)]
|
||||
#[derive(Copy, Clone, Debug, bytemuck::Pod, bytemuck::Zeroable)]
|
||||
struct DemosaicParams {
|
||||
width: u32,
|
||||
height: u32,
|
||||
crop_x: u32,
|
||||
crop_y: u32,
|
||||
stride: u32,
|
||||
pattern: u32,
|
||||
_pad0: u32,
|
||||
_pad1: u32,
|
||||
black: [f32; 4],
|
||||
inv_range: [f32; 4],
|
||||
}
|
||||
|
||||
/// A demosaiced image living on the GPU.
|
||||
///
|
||||
/// Linear, scene-referred, camera colour space, RGBA16Float. This is the
|
||||
/// input every adjustment operates on, and the reason the ops need no
|
||||
/// knowledge of sensors or CFA patterns.
|
||||
pub struct DemosaicedImage {
|
||||
texture: wgpu::Texture,
|
||||
view: wgpu::TextureView,
|
||||
width: u32,
|
||||
height: u32,
|
||||
/// Carried through for the camera→sRGB transform in the adjust pass.
|
||||
color_matrix: [f32; 9],
|
||||
/// As-shot white balance, the neutral starting point for the WB control.
|
||||
as_shot_wb: [f32; 3],
|
||||
}
|
||||
|
||||
impl DemosaicedImage {
|
||||
pub const FORMAT: wgpu::TextureFormat = wgpu::TextureFormat::Rgba16Float;
|
||||
|
||||
pub fn texture(&self) -> &wgpu::Texture {
|
||||
&self.texture
|
||||
}
|
||||
|
||||
pub fn view(&self) -> &wgpu::TextureView {
|
||||
&self.view
|
||||
}
|
||||
|
||||
pub fn size(&self) -> (u32, u32) {
|
||||
(self.width, self.height)
|
||||
}
|
||||
|
||||
/// Camera RGB → linear sRGB, row-major. Identity where the body is
|
||||
/// uncalibrated, so the image renders uncalibrated rather than black.
|
||||
pub fn color_matrix(&self) -> [f32; 9] {
|
||||
self.color_matrix
|
||||
}
|
||||
|
||||
/// As-shot white balance multipliers, green-normalised.
|
||||
///
|
||||
/// The white balance control is expressed *relative* to these, so its
|
||||
/// neutral position reproduces what the camera chose.
|
||||
pub fn as_shot_wb(&self) -> [f32; 3] {
|
||||
self.as_shot_wb
|
||||
}
|
||||
}
|
||||
|
||||
/// Runs the demosaic pass. Holds the pipeline so repeated images reuse it.
|
||||
pub struct Demosaicer {
|
||||
ctx: GpuContext,
|
||||
pipeline: wgpu::ComputePipeline,
|
||||
bind_group_layout: wgpu::BindGroupLayout,
|
||||
}
|
||||
|
||||
impl Demosaicer {
|
||||
pub fn new(ctx: &GpuContext) -> Result<Self, GpuError> {
|
||||
let shader = ctx
|
||||
.device
|
||||
.create_shader_module(wgpu::ShaderModuleDescriptor {
|
||||
label: Some("demosaic"),
|
||||
source: wgpu::ShaderSource::Wgsl(include_str!("shaders/demosaic.wgsl").into()),
|
||||
});
|
||||
|
||||
let bind_group_layout =
|
||||
ctx.device
|
||||
.create_bind_group_layout(&wgpu::BindGroupLayoutDescriptor {
|
||||
label: Some("demosaic-bgl"),
|
||||
entries: &[
|
||||
// Raw samples, packed two u16 per u32.
|
||||
wgpu::BindGroupLayoutEntry {
|
||||
binding: 0,
|
||||
visibility: wgpu::ShaderStages::COMPUTE,
|
||||
ty: wgpu::BindingType::Buffer {
|
||||
ty: wgpu::BufferBindingType::Storage { read_only: true },
|
||||
has_dynamic_offset: false,
|
||||
min_binding_size: None,
|
||||
},
|
||||
count: None,
|
||||
},
|
||||
wgpu::BindGroupLayoutEntry {
|
||||
binding: 1,
|
||||
visibility: wgpu::ShaderStages::COMPUTE,
|
||||
ty: wgpu::BindingType::Buffer {
|
||||
ty: wgpu::BufferBindingType::Uniform,
|
||||
has_dynamic_offset: false,
|
||||
min_binding_size: None,
|
||||
},
|
||||
count: None,
|
||||
},
|
||||
wgpu::BindGroupLayoutEntry {
|
||||
binding: 2,
|
||||
visibility: wgpu::ShaderStages::COMPUTE,
|
||||
ty: wgpu::BindingType::StorageTexture {
|
||||
access: wgpu::StorageTextureAccess::WriteOnly,
|
||||
format: DemosaicedImage::FORMAT,
|
||||
view_dimension: wgpu::TextureViewDimension::D2,
|
||||
},
|
||||
count: None,
|
||||
},
|
||||
],
|
||||
});
|
||||
|
||||
let layout = ctx
|
||||
.device
|
||||
.create_pipeline_layout(&wgpu::PipelineLayoutDescriptor {
|
||||
label: Some("demosaic-layout"),
|
||||
bind_group_layouts: &[&bind_group_layout],
|
||||
push_constant_ranges: &[],
|
||||
});
|
||||
|
||||
let pipeline = ctx
|
||||
.device
|
||||
.create_compute_pipeline(&wgpu::ComputePipelineDescriptor {
|
||||
label: Some("demosaic-pipeline"),
|
||||
layout: Some(&layout),
|
||||
module: &shader,
|
||||
entry_point: Some("main"),
|
||||
compilation_options: Default::default(),
|
||||
cache: None,
|
||||
});
|
||||
|
||||
Ok(Self {
|
||||
ctx: ctx.clone(),
|
||||
pipeline,
|
||||
bind_group_layout,
|
||||
})
|
||||
}
|
||||
|
||||
/// Upload and demosaic one image.
|
||||
pub fn run(&self, raw: &RawImage) -> Result<DemosaicedImage, GpuError> {
|
||||
if raw.cfa_pattern.is_xtrans() {
|
||||
// Better to say so than to render a maze of artefacts that reads
|
||||
// as a corrupt file (FR-RAW-5).
|
||||
return Err(GpuError::UnsupportedCfa(
|
||||
"X-Trans demosaic not yet implemented".into(),
|
||||
));
|
||||
}
|
||||
let pattern = match raw.cfa_pattern {
|
||||
CfaPattern::Rggb => 0u32,
|
||||
CfaPattern::Bggr => 1,
|
||||
CfaPattern::Grbg => 2,
|
||||
CfaPattern::Gbrg => 3,
|
||||
other => {
|
||||
return Err(GpuError::UnsupportedCfa(format!("{other:?}")));
|
||||
}
|
||||
};
|
||||
|
||||
let (width, height) = (raw.crop.width.max(1), raw.crop.height.max(1));
|
||||
|
||||
let limits = self.ctx.device.limits();
|
||||
if width > limits.max_texture_dimension_2d || height > limits.max_texture_dimension_2d {
|
||||
return Err(GpuError::TooLarge(format!(
|
||||
"{width}×{height} exceeds the device limit of {}",
|
||||
limits.max_texture_dimension_2d
|
||||
)));
|
||||
}
|
||||
|
||||
// Pack the u16 samples two per u32. WGSL has no u16 storage type, so
|
||||
// unpacking happens in the shader.
|
||||
let packed = pack_samples(&raw.data);
|
||||
let raw_buf = self
|
||||
.ctx
|
||||
.device
|
||||
.create_buffer_init(&wgpu::util::BufferInitDescriptor {
|
||||
label: Some("raw-samples"),
|
||||
contents: bytemuck::cast_slice(&packed),
|
||||
usage: wgpu::BufferUsages::STORAGE,
|
||||
});
|
||||
|
||||
let params = DemosaicParams {
|
||||
width,
|
||||
height,
|
||||
crop_x: raw.crop.x,
|
||||
crop_y: raw.crop.y,
|
||||
stride: raw.width,
|
||||
pattern,
|
||||
_pad0: 0,
|
||||
_pad1: 0,
|
||||
black: black_per_cell(raw),
|
||||
inv_range: inv_range_per_cell(raw),
|
||||
};
|
||||
let params_buf = self
|
||||
.ctx
|
||||
.device
|
||||
.create_buffer_init(&wgpu::util::BufferInitDescriptor {
|
||||
label: Some("demosaic-params"),
|
||||
contents: bytemuck::bytes_of(¶ms),
|
||||
usage: wgpu::BufferUsages::UNIFORM,
|
||||
});
|
||||
|
||||
let texture = self.ctx.device.create_texture(&wgpu::TextureDescriptor {
|
||||
label: Some("demosaiced"),
|
||||
size: wgpu::Extent3d {
|
||||
width,
|
||||
height,
|
||||
depth_or_array_layers: 1,
|
||||
},
|
||||
mip_level_count: 1,
|
||||
sample_count: 1,
|
||||
dimension: wgpu::TextureDimension::D2,
|
||||
format: DemosaicedImage::FORMAT,
|
||||
// STORAGE to write here, TEXTURE_BINDING so the adjust pass can
|
||||
// sample it. COPY_SRC only for tests.
|
||||
usage: wgpu::TextureUsages::STORAGE_BINDING
|
||||
| wgpu::TextureUsages::TEXTURE_BINDING
|
||||
| wgpu::TextureUsages::COPY_SRC,
|
||||
view_formats: &[],
|
||||
});
|
||||
let view = texture.create_view(&Default::default());
|
||||
|
||||
let bind_group = self
|
||||
.ctx
|
||||
.device
|
||||
.create_bind_group(&wgpu::BindGroupDescriptor {
|
||||
label: Some("demosaic-bg"),
|
||||
layout: &self.bind_group_layout,
|
||||
entries: &[
|
||||
wgpu::BindGroupEntry {
|
||||
binding: 0,
|
||||
resource: raw_buf.as_entire_binding(),
|
||||
},
|
||||
wgpu::BindGroupEntry {
|
||||
binding: 1,
|
||||
resource: params_buf.as_entire_binding(),
|
||||
},
|
||||
wgpu::BindGroupEntry {
|
||||
binding: 2,
|
||||
resource: wgpu::BindingResource::TextureView(&view),
|
||||
},
|
||||
],
|
||||
});
|
||||
|
||||
let mut enc = self
|
||||
.ctx
|
||||
.device
|
||||
.create_command_encoder(&wgpu::CommandEncoderDescriptor {
|
||||
label: Some("demosaic-encoder"),
|
||||
});
|
||||
{
|
||||
let mut pass = enc.begin_compute_pass(&wgpu::ComputePassDescriptor {
|
||||
label: Some("demosaic-pass"),
|
||||
timestamp_writes: None,
|
||||
});
|
||||
pass.set_pipeline(&self.pipeline);
|
||||
pass.set_bind_group(0, &bind_group, &[]);
|
||||
pass.dispatch_workgroups(width.div_ceil(8), height.div_ceil(8), 1);
|
||||
}
|
||||
self.ctx.queue.submit(Some(enc.finish()));
|
||||
|
||||
Ok(DemosaicedImage {
|
||||
texture,
|
||||
view,
|
||||
width,
|
||||
height,
|
||||
// Identity where the body is uncalibrated: the image renders with
|
||||
// no colour transform rather than not at all.
|
||||
color_matrix: raw.color_matrix.unwrap_or(IDENTITY_3X3),
|
||||
as_shot_wb: [raw.wb_coeffs[0], raw.wb_coeffs[1], raw.wb_coeffs[2]],
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
const IDENTITY_3X3: [f32; 9] = [1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0];
|
||||
|
||||
/// Pack u16 samples two per u32, little-endian within the word.
|
||||
///
|
||||
/// WGSL has no 16-bit storage type without an optional feature, so the shader
|
||||
/// unpacks. An odd sample count pads with a zero, which is never addressed:
|
||||
/// the shader indexes by pixel, not by word.
|
||||
fn pack_samples(data: &[u16]) -> Vec<u32> {
|
||||
let mut out = Vec::with_capacity(data.len().div_ceil(2));
|
||||
let mut chunks = data.chunks_exact(2);
|
||||
for pair in &mut chunks {
|
||||
out.push(u32::from(pair[0]) | (u32::from(pair[1]) << 16));
|
||||
}
|
||||
if let Some(&last) = chunks.remainder().first() {
|
||||
out.push(u32::from(last));
|
||||
}
|
||||
out
|
||||
}
|
||||
|
||||
/// Black level per CFA cell position, indexed `(y & 1) * 2 + (x & 1)`.
|
||||
///
|
||||
/// `dr-decode` reports four levels in CFA order, which is already this
|
||||
/// layout. Bodies reporting a single level get it broadcast.
|
||||
fn black_per_cell(raw: &RawImage) -> [f32; 4] {
|
||||
let b = raw.black_level;
|
||||
if b[1] == 0 && b[2] == 0 && b[3] == 0 {
|
||||
return [f32::from(b[0]); 4];
|
||||
}
|
||||
[
|
||||
f32::from(b[0]),
|
||||
f32::from(b[1]),
|
||||
f32::from(b[2]),
|
||||
f32::from(b[3]),
|
||||
]
|
||||
}
|
||||
|
||||
/// Reciprocal of the usable range per cell, so the shader avoids a division.
|
||||
///
|
||||
/// A white level at or below black would divide by zero; such a file is
|
||||
/// malformed, and falling back to full scale renders something inspectable
|
||||
/// rather than a NaN texture.
|
||||
fn inv_range_per_cell(raw: &RawImage) -> [f32; 4] {
|
||||
let white = f32::from(raw.white_level);
|
||||
let black = black_per_cell(raw);
|
||||
let mut out = [0.0f32; 4];
|
||||
for (i, &b) in black.iter().enumerate() {
|
||||
let range = white - b;
|
||||
out[i] = if range > 1.0 {
|
||||
1.0 / range
|
||||
} else {
|
||||
1.0 / 65535.0
|
||||
};
|
||||
}
|
||||
out
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
use dr_decode::CropRect;
|
||||
|
||||
fn raw_for(black: [u16; 4], white: u16) -> RawImage {
|
||||
RawImage {
|
||||
width: 4,
|
||||
height: 4,
|
||||
data: vec![0; 16],
|
||||
cfa_pattern: CfaPattern::Rggb,
|
||||
black_level: black,
|
||||
white_level: white,
|
||||
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
|
||||
color_matrix: None,
|
||||
crop: CropRect {
|
||||
x: 0,
|
||||
y: 0,
|
||||
width: 4,
|
||||
height: 4,
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn samples_pack_two_per_word() {
|
||||
let packed = pack_samples(&[0x1234, 0xABCD]);
|
||||
assert_eq!(packed, vec![0xABCD_1234]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn an_odd_sample_count_does_not_lose_the_last_value() {
|
||||
// A sensor with an odd sample count would otherwise drop its final
|
||||
// photosite, or worse, read past the buffer.
|
||||
let packed = pack_samples(&[0x0001, 0x0002, 0x0003]);
|
||||
assert_eq!(packed.len(), 2);
|
||||
assert_eq!(packed[1] & 0xFFFF, 3);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn packing_preserves_every_sample() {
|
||||
let data: Vec<u16> = (0..64).map(|i| i * 1000).collect();
|
||||
let packed = pack_samples(&data);
|
||||
for (i, &expected) in data.iter().enumerate() {
|
||||
let word = packed[i / 2];
|
||||
let got = if i % 2 == 0 {
|
||||
word & 0xFFFF
|
||||
} else {
|
||||
word >> 16
|
||||
};
|
||||
assert_eq!(got as u16, expected, "sample {i}");
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_single_black_level_is_broadcast_to_every_cell() {
|
||||
// Many bodies report one level rather than four; treating the absent
|
||||
// three as zero would leave three quarters of the image lifted.
|
||||
let raw = raw_for([512, 0, 0, 0], 16383);
|
||||
assert_eq!(black_per_cell(&raw), [512.0; 4]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn per_cell_black_levels_are_kept_distinct() {
|
||||
let raw = raw_for([2047, 2048, 2048, 2049], 15070);
|
||||
assert_eq!(black_per_cell(&raw), [2047.0, 2048.0, 2048.0, 2049.0]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn normalisation_maps_white_to_one() {
|
||||
let raw = raw_for([2048, 2048, 2048, 2048], 15070);
|
||||
let inv = inv_range_per_cell(&raw);
|
||||
let normalised = (15070.0 - 2048.0) * inv[0];
|
||||
assert!(
|
||||
(normalised - 1.0).abs() < 1e-5,
|
||||
"the white level must land on 1.0, got {normalised}"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_degenerate_range_does_not_divide_by_zero() {
|
||||
// A malformed file reporting white <= black must not produce NaN
|
||||
// across the whole texture.
|
||||
let raw = raw_for([5000, 5000, 5000, 5000], 4000);
|
||||
let inv = inv_range_per_cell(&raw);
|
||||
assert!(inv.iter().all(|v| v.is_finite() && *v > 0.0));
|
||||
}
|
||||
|
||||
// ---- GPU tests -----------------------------------------------------
|
||||
//
|
||||
// These exercise the shader itself. The CPU tests above cover the
|
||||
// parameter maths; only running the pass proves the kernels and the CFA
|
||||
// indexing are right.
|
||||
|
||||
fn ctx() -> Option<GpuContext> {
|
||||
match pollster::block_on(GpuContext::new_headless()) {
|
||||
Ok(c) => Some(c),
|
||||
Err(e) => {
|
||||
eprintln!("skipping: no GPU adapter ({e})");
|
||||
None
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Build a synthetic CFA image of a uniform colour.
|
||||
///
|
||||
/// Each photosite carries its own channel's value, which is what a
|
||||
/// sensor looking at a flat patch would record. A correct demosaic must
|
||||
/// return that colour at every pixel.
|
||||
fn flat_cfa(pattern: CfaPattern, size: u32, rgb: [u16; 3], black: u16, white: u16) -> RawImage {
|
||||
let mut data = vec![0u16; (size * size) as usize];
|
||||
for y in 0..size {
|
||||
for x in 0..size {
|
||||
let c = pattern.colour_at(x, y) as usize;
|
||||
data[(y * size + x) as usize] = black + rgb[c];
|
||||
}
|
||||
}
|
||||
RawImage {
|
||||
width: size,
|
||||
height: size,
|
||||
data,
|
||||
cfa_pattern: pattern,
|
||||
black_level: [black; 4],
|
||||
white_level: white,
|
||||
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
|
||||
color_matrix: None,
|
||||
crop: CropRect {
|
||||
x: 0,
|
||||
y: 0,
|
||||
width: size,
|
||||
height: size,
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
/// Read back the demosaiced texture as f32 RGBA.
|
||||
fn read_rgba(ctx: &GpuContext, img: &DemosaicedImage) -> Vec<[f32; 4]> {
|
||||
let (w, h) = img.size();
|
||||
let unpadded = w * 8; // RGBA16Float = 8 bytes per pixel
|
||||
let align = wgpu::COPY_BYTES_PER_ROW_ALIGNMENT;
|
||||
let padded = unpadded.div_ceil(align) * align;
|
||||
|
||||
let buf = ctx.device.create_buffer(&wgpu::BufferDescriptor {
|
||||
label: Some("demosaic-readback"),
|
||||
size: (padded * h) as u64,
|
||||
usage: wgpu::BufferUsages::COPY_DST | wgpu::BufferUsages::MAP_READ,
|
||||
mapped_at_creation: false,
|
||||
});
|
||||
|
||||
let mut enc = ctx.device.create_command_encoder(&Default::default());
|
||||
enc.copy_texture_to_buffer(
|
||||
wgpu::ImageCopyTexture {
|
||||
texture: img.texture(),
|
||||
mip_level: 0,
|
||||
origin: wgpu::Origin3d::ZERO,
|
||||
aspect: wgpu::TextureAspect::All,
|
||||
},
|
||||
wgpu::ImageCopyBuffer {
|
||||
buffer: &buf,
|
||||
layout: wgpu::ImageDataLayout {
|
||||
offset: 0,
|
||||
bytes_per_row: Some(padded),
|
||||
rows_per_image: Some(h),
|
||||
},
|
||||
},
|
||||
wgpu::Extent3d {
|
||||
width: w,
|
||||
height: h,
|
||||
depth_or_array_layers: 1,
|
||||
},
|
||||
);
|
||||
ctx.queue.submit(Some(enc.finish()));
|
||||
|
||||
let slice = buf.slice(..);
|
||||
let (tx, rx) = std::sync::mpsc::channel();
|
||||
slice.map_async(wgpu::MapMode::Read, move |r| {
|
||||
let _ = tx.send(r);
|
||||
});
|
||||
ctx.device.poll(wgpu::Maintain::Wait);
|
||||
rx.recv().expect("map").expect("map ok");
|
||||
|
||||
let data = slice.get_mapped_range();
|
||||
let mut out = Vec::with_capacity((w * h) as usize);
|
||||
for y in 0..h {
|
||||
let row = (y * padded) as usize;
|
||||
for x in 0..w {
|
||||
let px = row + (x * 8) as usize;
|
||||
let mut c = [0.0f32; 4];
|
||||
for (i, slot) in c.iter_mut().enumerate() {
|
||||
let o = px + i * 2;
|
||||
let bits = u16::from_le_bytes([data[o], data[o + 1]]);
|
||||
*slot = half_to_f32(bits);
|
||||
}
|
||||
out.push(c);
|
||||
}
|
||||
}
|
||||
drop(data);
|
||||
buf.unmap();
|
||||
out
|
||||
}
|
||||
|
||||
/// Decode an IEEE 754 binary16 value.
|
||||
fn half_to_f32(bits: u16) -> f32 {
|
||||
let sign = f32::from_bits(u32::from(bits & 0x8000) << 16);
|
||||
let exp = (bits >> 10) & 0x1F;
|
||||
let mant = bits & 0x03FF;
|
||||
let magnitude = match exp {
|
||||
0 => f32::from(mant) * 2f32.powi(-24),
|
||||
0x1F => {
|
||||
if mant == 0 {
|
||||
f32::INFINITY
|
||||
} else {
|
||||
f32::NAN
|
||||
}
|
||||
}
|
||||
_ => (1.0 + f32::from(mant) / 1024.0) * 2f32.powi(i32::from(exp) - 15),
|
||||
};
|
||||
magnitude.copysign(sign)
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_flat_patch_demosaics_to_that_colour() {
|
||||
// The fundamental correctness property: a sensor looking at a uniform
|
||||
// colour must reconstruct that colour everywhere. Any error in the
|
||||
// kernels, the CFA indexing, or the normalisation breaks it.
|
||||
let Some(ctx) = ctx() else { return };
|
||||
let d = Demosaicer::new(&ctx).expect("demosaicer");
|
||||
|
||||
let black = 2048u16;
|
||||
let white = 16383u16;
|
||||
let range = f32::from(white - black);
|
||||
// A colour with three clearly distinct channels, so a swap is loud.
|
||||
let rgb = [8000u16, 4000, 2000];
|
||||
let expected = [
|
||||
f32::from(rgb[0]) / range,
|
||||
f32::from(rgb[1]) / range,
|
||||
f32::from(rgb[2]) / range,
|
||||
];
|
||||
|
||||
let raw = flat_cfa(CfaPattern::Rggb, 32, rgb, black, white);
|
||||
let img = d.run(&raw).expect("demosaic");
|
||||
let px = read_rgba(&ctx, &img);
|
||||
|
||||
// Interior pixels only: the border reflects, and a 2-pixel margin is
|
||||
// where that shows.
|
||||
let (w, _) = img.size();
|
||||
for y in 2..30u32 {
|
||||
for x in 2..30u32 {
|
||||
let p = px[(y * w + x) as usize];
|
||||
for (ch, &want) in expected.iter().enumerate() {
|
||||
assert!(
|
||||
(p[ch] - want).abs() < 0.01,
|
||||
"pixel ({x},{y}) channel {ch}: got {}, want {want}",
|
||||
p[ch]
|
||||
);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn every_bayer_layout_reconstructs_the_same_colour() {
|
||||
// The packed CFA constants in the shader are easy to get wrong — two
|
||||
// of the four were wrong on the first attempt, and a wrong one swaps
|
||||
// red and blue. Each layout describes the same scene, so each must
|
||||
// produce the same output.
|
||||
let Some(ctx) = ctx() else { return };
|
||||
let d = Demosaicer::new(&ctx).expect("demosaicer");
|
||||
|
||||
let (black, white) = (0u16, 16383u16);
|
||||
let rgb = [9000u16, 5000, 1500];
|
||||
let expected = [
|
||||
f32::from(rgb[0]) / f32::from(white),
|
||||
f32::from(rgb[1]) / f32::from(white),
|
||||
f32::from(rgb[2]) / f32::from(white),
|
||||
];
|
||||
|
||||
for pattern in [
|
||||
CfaPattern::Rggb,
|
||||
CfaPattern::Bggr,
|
||||
CfaPattern::Grbg,
|
||||
CfaPattern::Gbrg,
|
||||
] {
|
||||
let raw = flat_cfa(pattern, 32, rgb, black, white);
|
||||
let img = d.run(&raw).expect("demosaic");
|
||||
let px = read_rgba(&ctx, &img);
|
||||
let (w, _) = img.size();
|
||||
|
||||
let p = px[(16 * w + 16) as usize];
|
||||
for (ch, &want) in expected.iter().enumerate() {
|
||||
assert!(
|
||||
(p[ch] - want).abs() < 0.01,
|
||||
"{pattern:?} channel {ch}: got {}, want {want} — \
|
||||
a wrong CFA constant swaps channels",
|
||||
p[ch]
|
||||
);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn an_odd_crop_origin_still_reconstructs_correctly() {
|
||||
// The case CropRect::shifts_cfa_phase exists for. Cropping to an odd
|
||||
// origin re-phases the pattern; if the shader is given the unshifted
|
||||
// one, red and blue swap.
|
||||
let Some(ctx) = ctx() else { return };
|
||||
let d = Demosaicer::new(&ctx).expect("demosaicer");
|
||||
|
||||
let (black, white) = (0u16, 16383u16);
|
||||
let rgb = [9000u16, 5000, 1500];
|
||||
let mut raw = flat_cfa(CfaPattern::Rggb, 34, rgb, black, white);
|
||||
|
||||
// Crop one photosite in on both axes, as a body with an odd active
|
||||
// area would. The visible top-left is now green-on-a-red-row.
|
||||
raw.crop = CropRect {
|
||||
x: 1,
|
||||
y: 1,
|
||||
width: 32,
|
||||
height: 32,
|
||||
};
|
||||
let (dx, dy) = raw.crop.shifts_cfa_phase();
|
||||
assert!(dx && dy, "the fixture must actually shift the phase");
|
||||
raw.cfa_pattern = raw.cfa_pattern.shifted(dx, dy);
|
||||
|
||||
let img = d.run(&raw).expect("demosaic");
|
||||
let px = read_rgba(&ctx, &img);
|
||||
let (w, _) = img.size();
|
||||
let p = px[(16 * w + 16) as usize];
|
||||
|
||||
let expected = [
|
||||
f32::from(rgb[0]) / f32::from(white),
|
||||
f32::from(rgb[1]) / f32::from(white),
|
||||
f32::from(rgb[2]) / f32::from(white),
|
||||
];
|
||||
for (ch, &want) in expected.iter().enumerate() {
|
||||
assert!(
|
||||
(p[ch] - want).abs() < 0.01,
|
||||
"channel {ch}: got {}, want {want} — the crop origin \
|
||||
re-phases the CFA and the shader must see the shifted pattern",
|
||||
p[ch]
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn output_is_free_of_nan_and_negatives() {
|
||||
// f16 NaN propagates silently through every later stage; a negative
|
||||
// value breaks the ratio-based operations downstream.
|
||||
let Some(ctx) = ctx() else { return };
|
||||
let d = Demosaicer::new(&ctx).expect("demosaicer");
|
||||
|
||||
// A high-contrast checkerboard, which is where the gradient
|
||||
// correction overshoots hardest.
|
||||
let size = 32u32;
|
||||
let mut data = vec![0u16; (size * size) as usize];
|
||||
for y in 0..size {
|
||||
for x in 0..size {
|
||||
data[(y * size + x) as usize] = if (x / 2 + y / 2) % 2 == 0 { 16000 } else { 40 };
|
||||
}
|
||||
}
|
||||
let raw = RawImage {
|
||||
width: size,
|
||||
height: size,
|
||||
data,
|
||||
cfa_pattern: CfaPattern::Rggb,
|
||||
black_level: [32; 4],
|
||||
white_level: 16383,
|
||||
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
|
||||
color_matrix: None,
|
||||
crop: CropRect {
|
||||
x: 0,
|
||||
y: 0,
|
||||
width: size,
|
||||
height: size,
|
||||
},
|
||||
};
|
||||
|
||||
let img = d.run(&raw).expect("demosaic");
|
||||
for (i, p) in read_rgba(&ctx, &img).iter().enumerate() {
|
||||
for (ch, v) in p.iter().take(3).enumerate() {
|
||||
assert!(v.is_finite(), "pixel {i} channel {ch} is {v}");
|
||||
assert!(*v >= 0.0, "pixel {i} channel {ch} is negative: {v}");
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn xtrans_is_refused_rather_than_rendered_wrong() {
|
||||
let Some(ctx) = ctx() else { return };
|
||||
let d = Demosaicer::new(&ctx).expect("demosaicer");
|
||||
let mut raw = flat_cfa(CfaPattern::Rggb, 8, [100, 100, 100], 0, 1000);
|
||||
raw.cfa_pattern = CfaPattern::XTrans;
|
||||
|
||||
assert!(
|
||||
matches!(d.run(&raw), Err(GpuError::UnsupportedCfa(_))),
|
||||
"an X-Trans file must report the gap, not render artefacts"
|
||||
);
|
||||
}
|
||||
}
|
||||
@@ -21,4 +21,13 @@ pub enum GpuError {
|
||||
|
||||
#[error("readback failed: {0}")]
|
||||
Readback(String),
|
||||
|
||||
/// The X-Trans demosaic is not implemented (FR-RAW-5). Reported rather
|
||||
/// than approximated with the Bayer path, which would produce a maze of
|
||||
/// colour artefacts and look like a corrupt file.
|
||||
#[error("unsupported CFA pattern: {0}")]
|
||||
UnsupportedCfa(String),
|
||||
|
||||
#[error("image too large for this device: {0}")]
|
||||
TooLarge(String),
|
||||
}
|
||||
|
||||
@@ -11,7 +11,11 @@ use std::sync::Arc;
|
||||
|
||||
use wgpu::util::DeviceExt;
|
||||
|
||||
mod adjust;
|
||||
mod demosaic;
|
||||
mod error;
|
||||
pub use adjust::AdjustPass;
|
||||
pub use demosaic::{DemosaicedImage, Demosaicer};
|
||||
pub use error::GpuError;
|
||||
|
||||
/// Owns the wgpu device and queue.
|
||||
|
||||
@@ -0,0 +1,183 @@
|
||||
// Black/white normalisation and Bayer demosaic, in one pass.
|
||||
//
|
||||
// Input is the raw sensor readout as packed u16 samples — one per photosite,
|
||||
// in CFA order. Output is linear scene-referred RGBA16Float in *camera*
|
||||
// colour space; the camera→sRGB matrix belongs to the adjust pass, so this
|
||||
// stage is purely about reconstructing three channels from one.
|
||||
//
|
||||
// The algorithm is Malvar-He-Cutler (ICASSP 2004): bilinear interpolation
|
||||
// plus a Laplacian correction taken from the channel that *is* sampled at
|
||||
// each site. One 5x5 neighbourhood per pixel, and dramatically better than
|
||||
// bilinear on edges — bilinear leaves visible zippering on any high-contrast
|
||||
// boundary, which on a 24 MP file is the first thing seen at 1:1.
|
||||
//
|
||||
// Kernel coefficients below are the paper's, all over 8.
|
||||
|
||||
struct DemosaicParams {
|
||||
// Dimensions of the *cropped* output, in pixels.
|
||||
width: u32,
|
||||
height: u32,
|
||||
// Origin of the crop within the sensor readout, in photosites. Added to
|
||||
// every read so the masked border is never sampled.
|
||||
crop_x: u32,
|
||||
crop_y: u32,
|
||||
// Row stride of the input, in samples.
|
||||
stride: u32,
|
||||
// CFA layout of the *cropped* image, already re-phased for the crop
|
||||
// origin by dr-decode: 0=RGGB, 1=BGGR, 2=GRBG, 3=GBRG.
|
||||
pattern: u32,
|
||||
_pad0: u32,
|
||||
_pad1: u32,
|
||||
// Per-CFA-position black levels, indexed by (y&1)*2 + (x&1).
|
||||
black: vec4<f32>,
|
||||
// Reciprocal of (white - black) per position, precomputed on the CPU so
|
||||
// the shader does no division.
|
||||
inv_range: vec4<f32>,
|
||||
}
|
||||
|
||||
@group(0) @binding(0) var<storage, read> raw: array<u32>;
|
||||
@group(0) @binding(1) var<uniform> params: DemosaicParams;
|
||||
@group(0) @binding(2) var output: texture_storage_2d<rgba16float, write>;
|
||||
|
||||
// Colour of the photosite at (x, y): 0=R, 1=G, 2=B.
|
||||
//
|
||||
// Each pattern is its 2x2 cell read row-major, packed two bits per entry so
|
||||
// the lookup is an index and a shift rather than a branch.
|
||||
fn colour_at(x: u32, y: u32) -> u32 {
|
||||
let cell = (y & 1u) * 2u + (x & 1u);
|
||||
// RGGB = R,G,G,B -> 0,1,1,2 ; BGGR = 2,1,1,0 ; GRBG = 1,0,2,1 ; GBRG = 1,2,0,1
|
||||
// Entry i occupies bits [2i, 2i+1], so the cell order reads
|
||||
// right-to-left in hex. Verified against a table rather than derived by
|
||||
// eye — two of these were wrong on the first attempt.
|
||||
var packed: u32;
|
||||
switch params.pattern {
|
||||
case 0u: { packed = 0x94u; } // RGGB -> [0,1,1,2]
|
||||
case 1u: { packed = 0x16u; } // BGGR -> [2,1,1,0]
|
||||
case 2u: { packed = 0x61u; } // GRBG -> [1,0,2,1]
|
||||
default: { packed = 0x49u; } // GBRG -> [1,2,0,1]
|
||||
}
|
||||
return (packed >> (cell * 2u)) & 3u;
|
||||
}
|
||||
|
||||
// Whether the row through (x, y) is one carrying red photosites.
|
||||
//
|
||||
// Needed at green sites, where red lies along one axis and blue along the
|
||||
// other, and which is which depends on the pattern.
|
||||
fn red_is_horizontal(x: u32, y: u32) -> bool {
|
||||
// The horizontal neighbour of a green site.
|
||||
return colour_at(x + 1u, y) == 0u;
|
||||
}
|
||||
|
||||
// Read one photosite, normalised to [0, 1] against its own black level.
|
||||
//
|
||||
// Coordinates are relative to the crop origin. A 5x5 window at the image edge
|
||||
// reflects rather than reading masked photosites or running off the buffer.
|
||||
fn sample(ix: i32, iy: i32) -> f32 {
|
||||
let w = i32(params.width);
|
||||
let h = i32(params.height);
|
||||
|
||||
// Reflect at the borders, preserving CFA parity: reflecting by an even
|
||||
// distance keeps the mirrored sample the same colour as the one it
|
||||
// stands in for. Clamping instead would flatten the correction term and
|
||||
// leave a visible one-pixel seam along each edge.
|
||||
var cx = ix;
|
||||
var cy = iy;
|
||||
if (cx < 0) { cx = -cx; }
|
||||
if (cy < 0) { cy = -cy; }
|
||||
if (cx > w - 1) { cx = 2 * (w - 1) - cx; }
|
||||
if (cy > h - 1) { cy = 2 * (h - 1) - cy; }
|
||||
cx = clamp(cx, 0, w - 1);
|
||||
cy = clamp(cy, 0, h - 1);
|
||||
|
||||
let sx = u32(cx) + params.crop_x;
|
||||
let sy = u32(cy) + params.crop_y;
|
||||
let index = sy * params.stride + sx;
|
||||
|
||||
// Samples are u16, packed two per u32 word.
|
||||
let word = raw[index >> 1u];
|
||||
let raw_value = select(word & 0xFFFFu, word >> 16u, (index & 1u) == 1u);
|
||||
|
||||
// Black level and range are per CFA position. Subtracting black can go
|
||||
// negative on sensor noise — real signal below the black point — so the
|
||||
// result is clamped rather than allowed to wrap.
|
||||
let cell = (u32(cy) & 1u) * 2u + (u32(cx) & 1u);
|
||||
let value = (f32(raw_value) - params.black[cell]) * params.inv_range[cell];
|
||||
return max(value, 0.0);
|
||||
}
|
||||
|
||||
@compute @workgroup_size(8, 8, 1)
|
||||
fn main(@builtin(global_invocation_id) gid: vec3<u32>) {
|
||||
if (gid.x >= params.width || gid.y >= params.height) {
|
||||
return;
|
||||
}
|
||||
|
||||
let x = i32(gid.x);
|
||||
let y = i32(gid.y);
|
||||
|
||||
let c = sample(x, y);
|
||||
|
||||
// 5x5 neighbourhood.
|
||||
let n1 = sample(x, y - 1);
|
||||
let s1 = sample(x, y + 1);
|
||||
let w1 = sample(x - 1, y);
|
||||
let e1 = sample(x + 1, y);
|
||||
|
||||
let n2 = sample(x, y - 2);
|
||||
let s2 = sample(x, y + 2);
|
||||
let w2 = sample(x - 2, y);
|
||||
let e2 = sample(x + 2, y);
|
||||
|
||||
let nw = sample(x - 1, y - 1);
|
||||
let ne = sample(x + 1, y - 1);
|
||||
let sw = sample(x - 1, y + 1);
|
||||
let se = sample(x + 1, y + 1);
|
||||
|
||||
let axial1 = n1 + s1 + w1 + e1;
|
||||
let diag1 = nw + ne + sw + se;
|
||||
let vert2 = n2 + s2;
|
||||
let horiz2 = w2 + e2;
|
||||
|
||||
let colour = colour_at(gid.x, gid.y);
|
||||
var rgb: vec3<f32>;
|
||||
|
||||
if (colour == 1u) {
|
||||
// ---- Green site ----------------------------------------------
|
||||
// Green is measured. Red and blue are interpolated from their own
|
||||
// axis, with a correction from the green Laplacian.
|
||||
//
|
||||
// Malvar "G at R/B locations" kernels, transposed per axis:
|
||||
// chroma along the row: (5c + 4(w1+e1) - (nw+ne+sw+se) - (n2+s2) + 0.5(w2+e2)) / 8
|
||||
let along_row =
|
||||
(5.0 * c + 4.0 * (w1 + e1) - diag1 - vert2 + 0.5 * horiz2) * 0.125;
|
||||
let along_col =
|
||||
(5.0 * c + 4.0 * (n1 + s1) - diag1 - horiz2 + 0.5 * vert2) * 0.125;
|
||||
|
||||
let red_horizontal = red_is_horizontal(gid.x, gid.y);
|
||||
let r = select(along_col, along_row, red_horizontal);
|
||||
let b = select(along_row, along_col, red_horizontal);
|
||||
rgb = vec3<f32>(r, c, b);
|
||||
} else {
|
||||
// ---- Red or blue site ----------------------------------------
|
||||
// Green at an R/B site: bilinear on the axial neighbours, corrected
|
||||
// by the centre channel's Laplacian.
|
||||
// (4c + 2(n1+s1+w1+e1) - (n2+s2+w2+e2)) / 8
|
||||
let green = (4.0 * c + 2.0 * axial1 - (vert2 + horiz2)) * 0.125;
|
||||
|
||||
// The opposite chroma sits on the diagonals.
|
||||
// (6c + 2(nw+ne+sw+se) - 1.5(n2+s2+w2+e2)) / 8
|
||||
let opposite = (6.0 * c + 2.0 * diag1 - 1.5 * (vert2 + horiz2)) * 0.125;
|
||||
|
||||
if (colour == 0u) {
|
||||
rgb = vec3<f32>(c, green, opposite);
|
||||
} else {
|
||||
rgb = vec3<f32>(opposite, green, c);
|
||||
}
|
||||
}
|
||||
|
||||
// The correction term can overshoot below zero near clipped highlights.
|
||||
// Negative light is not meaningful, and carrying it forward makes the
|
||||
// ratio-based operations downstream (white balance, saturation) misbehave.
|
||||
rgb = max(rgb, vec3<f32>(0.0));
|
||||
|
||||
textureStore(output, vec2<i32>(x, y), vec4<f32>(rgb, 1.0));
|
||||
}
|
||||
Reference in New Issue
Block a user