Brighten her face without touching the sky behind her

A mask layer is an ordinary develop chain plus a rule about where it
applies. Nothing in the chain knows it is being masked, so every operation
that works globally now works locally and a newly declared op in `ops/`
arrives with local support already done.

The composer emits each layer after the global chain and before the
conversion out of camera space, which is what a photographer means by "and
*then* lift the shadows on her face". Op fragments write to a `c` they
expect to own, so a layer block shadows it and copies the result back out
through a carrier — assigning the outer one from inside is impossible
precisely because it is shadowed. The fused dispatch survives: three global
adjustments and two masked ones remain one shader, one read, one write.

Masks rasterise on the GPU and never exist in CPU memory (ARCH §5.4). That
is the whole reason darktable's brush masks lag, and it is architectural
rather than tuning, so it is not a thing to inherit and fix later.

The rasteriser is a render pass rather than the compute shader it obviously
wants to be, and the format is why: R8Unorm is not a core storage format,
so a compute path has to widen masks to four bytes per pixel — 768 MB
across eight layers of a 24 MP export, against 192 MB at one byte. A colour
attachment takes R8Unorm happily. The array slice comes from the attached
view, so no slot uniform exists to disagree with where the pass writes.

Region masks index a compacted label field rather than the watershed's raw
basin roots, because a root is a sparse index into pixel space and
indexing a per-region array by one would need a table the size of the
image. Changing a selection then costs a few kilobytes, not a re-upload.

Stored as region ids, not as pixels: diffable, mergeable per-field under
FR-NC-9, and cheap in a sidecar. The ids only mean anything alongside the
segmentation that produced them, so each layer carries that signature and
is treated as stale rather than applied when it does not match — a
confidently wrong mask being much worse than an absent one.

Seven device tests render actual frames and read them back. The unit tests
either side check halves that would both pass if the two agreed with each
other and were both wrong; a mask sampled with x and y swapped satisfies
them and fails these.
This commit is contained in:
2026-08-22 08:39:16 +02:00
parent 0da8271836
commit c6a846a1f9
9 changed files with 1853 additions and 1 deletions
+65
View File
@@ -60,6 +60,8 @@ pub struct AdjustPass {
targets: [Option<Target>; 2],
/// Which of [`Self::targets`] the last render wrote.
current: usize,
/// Bound at `@binding(3)` when the edit carries no mask layers.
empty_masks: wgpu::TextureView,
}
struct Target {
@@ -109,6 +111,20 @@ impl AdjustPass {
},
count: None,
},
// The local-adjustment masks. Present in every layout
// whether or not the edit has any, because the layout
// is built once here and the generated shader declares
// the binding unconditionally for exactly that reason.
wgpu::BindGroupLayoutEntry {
binding: 3,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::Texture {
sample_type: wgpu::TextureSampleType::Float { filterable: true },
view_dimension: wgpu::TextureViewDimension::D2Array,
multisampled: false,
},
count: None,
},
],
});
@@ -120,6 +136,29 @@ impl AdjustPass {
immediate_size: 0,
});
// A 1x1 single-layer mask, bound when the edit has no local
// adjustments. The generated shader never samples it — no layer block
// is emitted — but a bind group must still satisfy the layout.
let empty = ctx.device.create_texture(&wgpu::TextureDescriptor {
label: Some("adjust-empty-masks"),
size: wgpu::Extent3d {
width: 1,
height: 1,
depth_or_array_layers: 1,
},
mip_level_count: 1,
sample_count: 1,
dimension: wgpu::TextureDimension::D2,
format: crate::MaskArray::FORMAT,
usage: wgpu::TextureUsages::TEXTURE_BINDING,
view_formats: &[],
});
let empty_masks = empty.create_view(&wgpu::TextureViewDescriptor {
label: Some("adjust-empty-masks-view"),
dimension: Some(wgpu::TextureViewDimension::D2Array),
..Default::default()
});
Self {
ctx: ctx.clone(),
bind_group_layout,
@@ -127,6 +166,7 @@ impl AdjustPass {
cache: HashMap::new(),
targets: [None, None],
current: 0,
empty_masks,
}
}
@@ -250,6 +290,25 @@ impl AdjustPass {
shader: &ComposedShader,
width: u32,
height: u32,
) -> Result<&wgpu::Texture, GpuError> {
self.render_masked(source, shader, width, height, None)
}
/// TRACES: FR-DEV-3
/// Render one frame with local adjustments applied.
///
/// `masks` must be the array [`crate::MaskPass`] rasterised for *this*
/// edit: the generated shader addresses slices by index, and an array
/// built from a different stack applies each layer's adjustment through
/// another layer's mask. Passing `None` is correct only for an edit with
/// no active mask layers.
pub fn render_masked(
&mut self,
source: &DemosaicedImage,
shader: &ComposedShader,
width: u32,
height: u32,
masks: Option<&crate::MaskArray>,
) -> Result<&wgpu::Texture, GpuError> {
let (width, height) = (width.max(1), height.max(1));
self.ensure_target(width, height);
@@ -310,6 +369,12 @@ impl AdjustPass {
binding: 2,
resource: wgpu::BindingResource::TextureView(&target.view),
},
wgpu::BindGroupEntry {
binding: 3,
resource: wgpu::BindingResource::TextureView(
masks.map_or(&self.empty_masks, |m| m.view()),
),
},
],
});
+10
View File
@@ -31,4 +31,14 @@ pub enum GpuError {
#[error("image too large for this device: {0}")]
TooLarge(String),
/// A mask input that cannot describe the image it claims to — a label
/// field whose length disagrees with its own dimensions, most often.
///
/// Its own variant rather than a panic because the caller assembles this
/// from a segmentation and a render size that are computed in different
/// places, and a mismatch between them is a bug worth reporting with its
/// numbers rather than an abort.
#[error("invalid mask input: {0}")]
InvalidMask(String),
}
+2
View File
@@ -21,6 +21,7 @@ mod adjust;
mod demosaic;
mod error;
mod histogram;
mod mask;
mod readback;
mod segment;
pub use adjust::AdjustPass;
@@ -29,6 +30,7 @@ pub use error::GpuError;
// Renamed on the way out: `BINS` says enough inside `histogram`, and nothing
// at all at a crate root shared with demosaic and segmentation.
pub use histogram::{Histogram, HistogramPass, BINS as HISTOGRAM_BINS};
pub use mask::{LabelField, MaskArray, MaskPass};
pub use segment::{SegmentOptions, SegmentPass, Segmentation};
/// Owns the wgpu device and queue.
+645
View File
@@ -0,0 +1,645 @@
//! Rasterising local-adjustment masks (ARCH §5.4).
//!
//! Turns a [`MaskStack`]'s rules into an r8unorm texture array, one slice per
//! active layer, which the composed adjust shader samples. Nothing here reads
//! back, and no mask ever exists in CPU memory.
//!
//! # What runs when
//!
//! Rasterising is **not** on the slider path. Dragging exposure on a masked
//! layer changes uniforms only; the mask array is reused untouched. This pass
//! runs when a mask's *shape* changes — a different selection, a moved
//! gradient, a resized output — which is what keeps a local adjustment as
//! responsive as a global one.
//!
//! # The label field
//!
//! Region masks index a compacted label field uploaded once per segmentation.
//! Compacted, rather than the watershed's raw basin roots, because a root is a
//! sparse index into pixel space: indexing a per-region array by one would
//! need a table the size of the image, where compacted ids index an array of
//! `region_count`. The compaction is CPU-side and once per image, which is the
//! same place and cadence the region adjacency graph is already built at.
use dr_pipeline::mask::{MaskSource, MaskStack, MAX_LAYERS};
use wgpu::util::DeviceExt;
use crate::{GpuContext, GpuError};
/// Modes understood by `mask.wgsl`. Kept beside the shader's `switch`.
const MODE_REGIONS: u32 = 0;
const MODE_LINEAR: u32 = 1;
const MODE_RADIAL: u32 = 2;
#[repr(C)]
#[derive(Copy, Clone, bytemuck::Pod, bytemuck::Zeroable)]
struct MaskParams {
width: u32,
height: u32,
label_width: u32,
label_height: u32,
mode: u32,
region_count: u32,
feather: f32,
_pad0: f32,
centre: [f32; 2],
axis: [f32; 2],
softness: f32,
angle: f32,
_pad1: [f32; 2],
}
/// The segmentation a region mask indexes into, resident on the GPU.
///
/// Uploaded once per image. Holds the compacted label field and nothing else —
/// the hierarchy that produced the ids stays on the CPU, where the interactive
/// operations (walk up a level, add a region) are cheap graph work.
pub struct LabelField {
buffer: wgpu::Buffer,
width: u32,
height: u32,
region_count: u32,
}
impl LabelField {
/// Upload a compacted label field.
///
/// `labels` is one region id per pixel, every value below `region_count` —
/// exactly [`dr_segment::RegionField::labels`].
pub fn upload(
ctx: &GpuContext,
labels: &[u32],
width: u32,
height: u32,
region_count: u32,
) -> Result<Self, GpuError> {
if labels.len() != (width * height) as usize {
return Err(GpuError::InvalidMask(format!(
"label field is {} entries, expected {}x{}",
labels.len(),
width,
height
)));
}
let buffer = ctx
.device
.create_buffer_init(&wgpu::util::BufferInitDescriptor {
label: Some("mask-labels"),
contents: bytemuck::cast_slice(labels),
usage: wgpu::BufferUsages::STORAGE,
});
Ok(Self {
buffer,
width,
height,
region_count,
})
}
pub fn region_count(&self) -> u32 {
self.region_count
}
pub fn size(&self) -> (u32, u32) {
(self.width, self.height)
}
}
/// The rasterised masks for one edit.
pub struct MaskArray {
texture: wgpu::Texture,
view: wgpu::TextureView,
width: u32,
height: u32,
layers: u32,
}
impl MaskArray {
pub const FORMAT: wgpu::TextureFormat = wgpu::TextureFormat::R8Unorm;
/// The view the adjust shader binds at `@binding(3)`.
pub fn view(&self) -> &wgpu::TextureView {
&self.view
}
pub fn layers(&self) -> u32 {
self.layers
}
pub fn size(&self) -> (u32, u32) {
(self.width, self.height)
}
fn matches(&self, width: u32, height: u32, layers: u32) -> bool {
self.width == width && self.height == height && self.layers == layers
}
}
/// Rasterises mask layers.
pub struct MaskPass {
ctx: GpuContext,
layout: wgpu::BindGroupLayout,
pipeline: wgpu::RenderPipeline,
array: Option<MaskArray>,
/// How many times the array texture has been (re)allocated.
///
/// Exists to be asserted on. Reallocating per frame instead of per resize
/// is the kind of regression that costs a lot of bandwidth and shows up
/// nowhere in the output, so the cheap reuse path is worth a test that
/// can actually see it.
allocations: usize,
/// A one-region, always-unselected field, for a stack with no region mask.
///
/// The shader's bindings are fixed, so *something* must be bound at the
/// label slots even when rasterising a gradient. A placeholder is cheaper
/// and far simpler than two pipelines differing only in what they ignore.
placeholder: LabelField,
}
impl MaskPass {
pub fn new(ctx: &GpuContext) -> Result<Self, GpuError> {
let scope = ctx.device.push_error_scope(wgpu::ErrorFilter::Validation);
let module = ctx
.device
.create_shader_module(wgpu::ShaderModuleDescriptor {
label: Some("mask"),
source: wgpu::ShaderSource::Wgsl(include_str!("shaders/mask.wgsl").into()),
});
let layout = ctx
.device
.create_bind_group_layout(&wgpu::BindGroupLayoutDescriptor {
label: Some("mask-bgl"),
entries: &[uniform_entry(0), storage_entry(1), storage_entry(2)],
});
let pipeline_layout = ctx
.device
.create_pipeline_layout(&wgpu::PipelineLayoutDescriptor {
label: Some("mask-layout"),
bind_group_layouts: &[Some(&layout)],
immediate_size: 0,
});
let pipeline = ctx
.device
.create_render_pipeline(&wgpu::RenderPipelineDescriptor {
label: Some("mask-pipeline"),
layout: Some(&pipeline_layout),
vertex: wgpu::VertexState {
module: &module,
entry_point: Some("vs"),
compilation_options: Default::default(),
buffers: &[],
},
fragment: Some(wgpu::FragmentState {
module: &module,
entry_point: Some("fs"),
compilation_options: Default::default(),
targets: &[Some(MaskArray::FORMAT.into())],
}),
primitive: wgpu::PrimitiveState::default(),
depth_stencil: None,
multisample: wgpu::MultisampleState::default(),
multiview_mask: None,
cache: None,
});
if let Some(err) = pollster::block_on(scope.pop()) {
return Err(GpuError::ShaderCompilation(err.to_string()));
}
let placeholder = LabelField::upload(ctx, &[0], 1, 1, 0)?;
Ok(Self {
ctx: ctx.clone(),
layout,
pipeline,
array: None,
allocations: 0,
placeholder,
})
}
/// Rasterise every active layer, returning the array to bind.
///
/// `labels` may be `None` when no layer is a region mask; a region layer
/// without one is skipped rather than drawn wrong, since a mask that
/// silently covers the whole frame would apply an edit everywhere.
pub fn render(
&mut self,
stack: &MaskStack,
labels: Option<&LabelField>,
width: u32,
height: u32,
) -> Result<&MaskArray, GpuError> {
// At least one layer, because a zero-layer texture array is invalid
// and the shader binds this slot unconditionally.
let active = stack.active_count().clamp(1, MAX_LAYERS) as u32;
self.ensure_array(width, height, active)?;
let mut encoder = self
.ctx
.device
.create_command_encoder(&wgpu::CommandEncoderDescriptor {
label: Some("mask-encoder"),
});
for (slot, layer) in stack.active().enumerate().take(MAX_LAYERS) {
let field = match (&layer.source, labels) {
(MaskSource::Regions { .. }, None) => {
log::warn!(
"mask layer {} is a region mask with no segmentation loaded; skipping",
layer.id
);
continue;
}
(MaskSource::Regions { .. }, Some(f)) => f,
(_, _) => &self.placeholder,
};
let params = self.params(layer, field, width, height);
let selected = self.selection_buffer(layer, field);
self.draw(&mut encoder, slot as u32, &params, field, &selected);
}
self.ctx.queue.submit([encoder.finish()]);
Ok(self.array.as_ref().expect("array was just ensured"))
}
/// The currently rasterised array, if any.
pub fn array(&self) -> Option<&MaskArray> {
self.array.as_ref()
}
/// How many times the array texture has been allocated. For tests.
pub fn allocations(&self) -> usize {
self.allocations
}
fn params(
&self,
layer: &dr_pipeline::mask::MaskLayer,
field: &LabelField,
width: u32,
height: u32,
) -> MaskParams {
let base = MaskParams {
width,
height,
label_width: field.width,
label_height: field.height,
mode: MODE_REGIONS,
region_count: field.region_count,
feather: 0.0,
_pad0: 0.0,
centre: [0.5, 0.5],
axis: [1.0, 0.0],
softness: 0.0,
angle: 0.0,
_pad1: [0.0, 0.0],
};
match &layer.source {
MaskSource::Regions { .. } => MaskParams {
// A pixel of softening at the proxy-to-output ratio, so the
// edge is equally soft whatever size the render is.
feather: (width as f32 / field.width.max(1) as f32).clamp(0.0, 4.0),
..base
},
MaskSource::Linear {
centre,
angle,
width: ramp,
} => MaskParams {
mode: MODE_LINEAR,
centre: [centre.0, centre.1],
axis: [angle.cos(), angle.sin()],
softness: *ramp,
..base
},
MaskSource::Radial {
centre,
radii,
angle,
feather,
} => MaskParams {
mode: MODE_RADIAL,
centre: [centre.0, centre.1],
axis: [radii.0.max(1e-6), radii.1.max(1e-6)],
softness: *feather,
angle: *angle,
..base
},
}
}
/// One byte-flag per region, or a single zero for a non-region layer.
fn selection_buffer(
&self,
layer: &dr_pipeline::mask::MaskLayer,
field: &LabelField,
) -> wgpu::Buffer {
let mut flags = vec![0u32; field.region_count.max(1) as usize];
if let MaskSource::Regions { ids, .. } = &layer.source {
for &id in ids {
if let Some(slot) = flags.get_mut(id as usize) {
*slot = 1;
}
}
}
self.ctx
.device
.create_buffer_init(&wgpu::util::BufferInitDescriptor {
label: Some("mask-selection"),
contents: bytemuck::cast_slice(&flags),
usage: wgpu::BufferUsages::STORAGE,
})
}
#[allow(clippy::too_many_arguments)]
fn draw(
&self,
encoder: &mut wgpu::CommandEncoder,
slot: u32,
params: &MaskParams,
field: &LabelField,
selected: &wgpu::Buffer,
) {
let params_buf = self
.ctx
.device
.create_buffer_init(&wgpu::util::BufferInitDescriptor {
label: Some("mask-params"),
contents: bytemuck::bytes_of(params),
usage: wgpu::BufferUsages::UNIFORM,
});
let bind_group = self
.ctx
.device
.create_bind_group(&wgpu::BindGroupDescriptor {
label: Some("mask-bind"),
layout: &self.layout,
entries: &[
wgpu::BindGroupEntry {
binding: 0,
resource: params_buf.as_entire_binding(),
},
wgpu::BindGroupEntry {
binding: 1,
resource: field.buffer.as_entire_binding(),
},
wgpu::BindGroupEntry {
binding: 2,
resource: selected.as_entire_binding(),
},
],
});
// The array slice is selected by the attachment rather than by a
// uniform the shader reads — one fewer value that can disagree with
// where the pass actually writes.
let array = self.array.as_ref().expect("array ensured by caller");
let view = array.texture.create_view(&wgpu::TextureViewDescriptor {
label: Some("mask-slice"),
dimension: Some(wgpu::TextureViewDimension::D2),
base_array_layer: slot,
array_layer_count: Some(1),
..Default::default()
});
let mut pass = encoder.begin_render_pass(&wgpu::RenderPassDescriptor {
label: Some("mask-pass"),
color_attachments: &[Some(wgpu::RenderPassColorAttachment {
view: &view,
depth_slice: None,
resolve_target: None,
ops: wgpu::Operations {
// Cleared rather than loaded: every pixel is written by the
// triangle below, and declaring that lets a tiler skip
// reading the previous contents in.
load: wgpu::LoadOp::Clear(wgpu::Color::BLACK),
store: wgpu::StoreOp::Store,
},
})],
depth_stencil_attachment: None,
timestamp_writes: None,
occlusion_query_set: None,
multiview_mask: None,
});
pass.set_pipeline(&self.pipeline);
pass.set_bind_group(0, &bind_group, &[]);
pass.draw(0..3, 0..1);
}
fn ensure_array(&mut self, width: u32, height: u32, layers: u32) -> Result<(), GpuError> {
if self.array.as_ref().is_some_and(|a| a.matches(width, height, layers)) {
return Ok(());
}
let texture = self.ctx.device.create_texture(&wgpu::TextureDescriptor {
label: Some("mask-array"),
size: wgpu::Extent3d {
width,
height,
depth_or_array_layers: layers,
},
mip_level_count: 1,
sample_count: 1,
dimension: wgpu::TextureDimension::D2,
format: MaskArray::FORMAT,
usage: wgpu::TextureUsages::RENDER_ATTACHMENT | wgpu::TextureUsages::TEXTURE_BINDING,
view_formats: &[],
});
let view = texture.create_view(&wgpu::TextureViewDescriptor {
label: Some("mask-array-view"),
dimension: Some(wgpu::TextureViewDimension::D2Array),
..Default::default()
});
self.allocations += 1;
self.array = Some(MaskArray {
texture,
view,
width,
height,
layers,
});
Ok(())
}
}
fn uniform_entry(binding: u32) -> wgpu::BindGroupLayoutEntry {
wgpu::BindGroupLayoutEntry {
binding,
visibility: wgpu::ShaderStages::FRAGMENT,
ty: wgpu::BindingType::Buffer {
ty: wgpu::BufferBindingType::Uniform,
has_dynamic_offset: false,
min_binding_size: None,
},
count: None,
}
}
fn storage_entry(binding: u32) -> wgpu::BindGroupLayoutEntry {
wgpu::BindGroupLayoutEntry {
binding,
visibility: wgpu::ShaderStages::FRAGMENT,
ty: wgpu::BindingType::Buffer {
ty: wgpu::BufferBindingType::Storage { read_only: true },
has_dynamic_offset: false,
min_binding_size: None,
},
count: None,
}
}
#[cfg(test)]
mod tests {
use super::*;
use dr_pipeline::descriptor::ParamId;
use dr_pipeline::mask::MaskLayer;
fn ctx() -> Option<GpuContext> {
pollster::block_on(GpuContext::new_headless()).ok()
}
/// A 4x2 label field: regions 0 and 1 left, 2 and 3 right.
fn labels() -> (Vec<u32>, u32, u32, u32) {
(vec![0, 0, 2, 2, 1, 1, 3, 3], 4, 2, 4)
}
fn lit(source: MaskSource) -> MaskLayer {
let mut layer = MaskLayer::new("m1", source);
layer.set_param("exposure", ParamId("exposure"), 1.0);
layer
}
#[test]
fn a_label_field_of_the_wrong_size_is_rejected() {
let Some(ctx) = ctx() else {
eprintln!("no adapter; skipping");
return;
};
assert!(LabelField::upload(&ctx, &[0, 1, 2], 4, 2, 4).is_err());
}
#[test]
fn region_masks_rasterise_to_the_selected_regions() {
let Some(ctx) = ctx() else {
eprintln!("no adapter; skipping");
return;
};
let (data, w, h, n) = labels();
let field = LabelField::upload(&ctx, &data, w, h, n).expect("upload");
let mut stack = MaskStack::new();
stack.push(lit(MaskSource::Regions {
signature: 1,
level: 4,
ids: vec![0, 1],
}));
let mut pass = MaskPass::new(&ctx).expect("mask pass");
let array = pass.render(&stack, Some(&field), w, h).expect("render");
assert_eq!(array.size(), (w, h));
assert_eq!(array.layers(), 1);
}
/// A region layer with no segmentation must produce nothing rather than
/// an all-covering mask, which would apply the edit to the whole frame.
#[test]
fn a_region_layer_without_labels_is_skipped() {
let Some(ctx) = ctx() else {
eprintln!("no adapter; skipping");
return;
};
let mut stack = MaskStack::new();
stack.push(lit(MaskSource::Regions {
signature: 1,
level: 4,
ids: vec![0],
}));
let mut pass = MaskPass::new(&ctx).expect("mask pass");
assert!(pass.render(&stack, None, 8, 8).is_ok());
}
#[test]
fn gradients_need_no_segmentation() {
let Some(ctx) = ctx() else {
eprintln!("no adapter; skipping");
return;
};
let mut stack = MaskStack::new();
stack.push(lit(MaskSource::Linear {
centre: (0.5, 0.5),
angle: 0.0,
width: 0.2,
}));
stack.push(lit(MaskSource::Radial {
centre: (0.5, 0.5),
radii: (0.3, 0.2),
angle: 0.0,
feather: 0.5,
}));
let mut pass = MaskPass::new(&ctx).expect("mask pass");
let array = pass.render(&stack, None, 16, 16).expect("render");
assert_eq!(array.layers(), 2, "one slice per active layer");
}
#[test]
fn an_empty_stack_still_yields_a_bindable_array() {
let Some(ctx) = ctx() else {
eprintln!("no adapter; skipping");
return;
};
let mut pass = MaskPass::new(&ctx).expect("mask pass");
let array = pass
.render(&MaskStack::new(), None, 8, 8)
.expect("render");
assert_eq!(
array.layers(),
1,
"the adjust shader binds this slot whether or not it reads it"
);
}
#[test]
fn the_array_is_reused_when_nothing_changed() {
let Some(ctx) = ctx() else {
eprintln!("no adapter; skipping");
return;
};
let mut stack = MaskStack::new();
stack.push(lit(MaskSource::Linear {
centre: (0.5, 0.5),
angle: 0.0,
width: 0.2,
}));
let mut pass = MaskPass::new(&ctx).expect("mask pass");
pass.render(&stack, None, 32, 32).expect("render");
assert_eq!(pass.allocations(), 1);
pass.render(&stack, None, 32, 32).expect("render");
assert_eq!(
pass.allocations(),
1,
"same size and layer count should not reallocate"
);
pass.render(&stack, None, 64, 64).expect("render");
assert_eq!(pass.allocations(), 2, "a resize must reallocate");
}
}
+158
View File
@@ -0,0 +1,158 @@
// Rasterise one local-adjustment mask into a layer of the mask array.
//
// ARCH §5.4: every mask becomes pixels here and never in CPU memory. One draw
// per layer, each targeting its own array slice, run only when a mask's
// *shape* changes — moving a slider on a masked layer re-runs the adjust
// shader and not this one.
//
// # Why this is a render pass and not a compute one
//
// The natural shape for this is a compute shader writing a storage texture,
// and the format is what rules that out: **R8Unorm is not a core storage
// format**, so a compute path has to widen the mask to R32Float or RGBA8 —
// four bytes per pixel per layer. At eight layers over a 24 MP export that is
// 768 MB of masks, against 192 MB at one byte. A colour attachment takes
// R8Unorm happily, so the mask stays one byte and the pass becomes a
// full-screen triangle.
//
// The array slice is chosen by the *view* the caller attaches, so there is no
// slot uniform here — one less thing that can disagree with the shader.
struct MaskParams {
// Output size, which is the render size rather than the segmentation's.
width: u32,
height: u32,
// Label field size. Different from the above: the watershed runs at a
// proxy resolution, and the mask is drawn at whatever the display or the
// export asked for.
label_width: u32,
label_height: u32,
// 0 = regions, 1 = linear, 2 = radial.
mode: u32,
// How many regions the label field holds, so an out-of-range label is
// caught rather than read past the end of `selected`.
region_count: u32,
// Softening applied to a region mask, in output pixels.
feather: f32,
_pad0: f32,
// Geometry, in normalised output coordinates. Meaning depends on `mode`.
centre: vec2<f32>,
// Linear: (cos, sin) of the ramp direction. Radial: semi-axes.
axis: vec2<f32>,
// Linear: ramp width. Radial: edge falloff as a fraction of the radius.
softness: f32,
// Radial only: rotation of the ellipse.
angle: f32,
_pad1: vec2<f32>,
}
@group(0) @binding(0) var<uniform> p: MaskParams;
// Compacted region id per pixel of the label field. Compacted rather than the
// watershed's raw basin roots: the roots are sparse indices into pixel space,
// so indexing a per-region array by one would need a table as large as the
// image. The compaction happens once, when the segmentation is built.
@group(0) @binding(1) var<storage, read> labels: array<u32>;
// One entry per region: non-zero if the region is in this mask. Small — a few
// thousand bytes — which is what makes changing a selection cheap.
@group(0) @binding(2) var<storage, read> selected: array<u32>;
// A full-screen triangle rather than a quad: three vertices instead of six,
// no shared edge for the rasteriser to crack along, and no vertex buffer.
@vertex
fn vs(@builtin(vertex_index) i: u32) -> @builtin(position) vec4<f32> {
let x = f32(i32(i) / 2) * 4.0 - 1.0;
let y = f32(i32(i) & 1) * 4.0 - 1.0;
return vec4<f32>(x, y, 0.0, 1.0);
}
fn region_at(px: vec2<i32>) -> u32 {
// Nearest-neighbour from output space into the label field. Deliberately
// not bilinear: region ids are *names*, and the average of region 4 and
// region 9 is not region 6.
let fx = (f32(px.x) + 0.5) / f32(p.width);
let fy = (f32(px.y) + 0.5) / f32(p.height);
let lx = clamp(i32(fx * f32(p.label_width)), 0, i32(p.label_width) - 1);
let ly = clamp(i32(fy * f32(p.label_height)), 0, i32(p.label_height) - 1);
return labels[u32(ly) * p.label_width + u32(lx)];
}
fn in_selection(px: vec2<i32>) -> f32 {
let r = region_at(px);
if (r >= p.region_count) {
return 0.0;
}
return select(0.0, 1.0, selected[r] != 0u);
}
fn region_mask(px: vec2<i32>) -> f32 {
let hard = in_selection(px);
if (p.feather <= 0.0) {
return hard;
}
// Box-average the binary selection over the feather radius. Cheap, and it
// is the whole reason a region mask does not look cut out with scissors:
// the watershed boundary is pixel-exact, which is correct and also harsher
// than any edit wants at a subject's edge.
let r = i32(ceil(p.feather));
var total = 0.0;
var n = 0.0;
for (var dy = -r; dy <= r; dy = dy + 1) {
for (var dx = -r; dx <= r; dx = dx + 1) {
let q = clamp(
px + vec2<i32>(dx, dy),
vec2<i32>(0, 0),
vec2<i32>(i32(p.width) - 1, i32(p.height) - 1),
);
total = total + in_selection(q);
n = n + 1.0;
}
}
return total / n;
}
fn linear_mask(uv: vec2<f32>) -> f32 {
// Signed distance along the ramp direction, from the centre.
let d = dot(uv - p.centre, p.axis);
if (p.softness <= 0.0) {
return select(0.0, 1.0, d >= 0.0);
}
return smoothstep(-p.softness * 0.5, p.softness * 0.5, d);
}
fn radial_mask(uv: vec2<f32>) -> f32 {
let ca = cos(-p.angle);
let sa = sin(-p.angle);
let d = uv - p.centre;
// Into the ellipse's own frame, then normalised by its semi-axes so the
// problem becomes a unit circle.
let local = vec2<f32>(d.x * ca - d.y * sa, d.x * sa + d.y * ca);
let r = length(local / max(p.axis, vec2<f32>(1e-6)));
let edge = clamp(p.softness, 0.0, 1.0);
if (edge <= 0.0) {
return select(0.0, 1.0, r <= 1.0);
}
return 1.0 - smoothstep(1.0 - edge, 1.0, r);
}
@fragment
fn fs(@builtin(position) pos: vec4<f32>) -> @location(0) vec4<f32> {
let px = vec2<i32>(i32(pos.x), i32(pos.y));
// Normalised, so a gradient's geometry survives a crop or an export at
// another size — the mask is defined on the frame, not on a pixel count.
let uv = vec2<f32>(pos.x / f32(p.width), pos.y / f32(p.height));
var m = 0.0;
switch p.mode {
case 0u: { m = region_mask(px); }
case 1u: { m = linear_mask(uv); }
case 2u: { m = radial_mask(uv); }
default: { m = 0.0; }
}
return vec4<f32>(clamp(m, 0.0, 1.0), 0.0, 0.0, 1.0);
}