Let a detail pass carry a list, not only a kernel

Every neighbourhood pass so far has been a convolution, whose whole
description fits in the uniform block because its structure fixes how many
numbers it needs. Spot removal is not that shape: sixty-four repairs and
one repair are the same shader with a different buffer behind it.

So a pass may declare `storage`, which arrives at binding 3 as
`array<vec4<f32>>` with `arrayLength` in scope. The alternative — packing
the list into uniforms — needs a fixed maximum paid for on every frame, a
composer that can emit vec4 fields because a uniform array's stride is 16
whatever it holds, and it gives the next operation that wants a table
nothing to build on.

The property worth having is what stays out of the generated source: the
count is in the buffer, so placing the tenth spot uploads 512 bytes and
reuses the compiled pipeline, exactly as moving a slider does for the
fused pass. `changing_the_list_does_not_recompile` is that, asserted.

One bind group entry rather than two more layouts, and one placeholder
buffer allocated in `new` rather than sixteen bytes per pass per frame —
a zero-length storage buffer cannot be bound, and per-frame allocation is
what this module's documentation exists to refuse.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-26 20:00:24 +02:00
co-authored by Claude Opus 5
parent 44444f7768
commit 6997c0f7ac
8 changed files with 287 additions and 5 deletions
+63
View File
@@ -149,6 +149,8 @@ pub(crate) struct DetailRunner {
/// Compiled pipelines by pass structure hash.
cache: HashMap<u64, wgpu::ComputePipeline>,
pool: Intermediates,
/// See [`placeholder_instances`].
no_instances: wgpu::Buffer,
}
struct Layout {
@@ -156,6 +158,22 @@ struct Layout {
pipeline: wgpu::PipelineLayout,
}
/// What binding 3 holds for a pass that declared no instance list.
///
/// One zeroed element, allocated once. Zero-length storage buffers cannot be
/// bound, and the passes that read this binding are exactly the ones that
/// uploaded something of their own, so nothing ever reads the placeholder's
/// contents — it exists to keep one bind group layout serving both kinds of
/// pass.
fn placeholder_instances(ctx: &GpuContext) -> wgpu::Buffer {
ctx.device
.create_buffer_init(&wgpu::util::BufferInitDescriptor {
label: Some("detail-instances-placeholder"),
contents: bytemuck::cast_slice(&[[0.0f32; 4]]),
usage: wgpu::BufferUsages::STORAGE,
})
}
impl DetailRunner {
pub(crate) fn new(ctx: &GpuContext) -> Self {
Self {
@@ -164,6 +182,7 @@ impl DetailRunner {
to_output: Layout::new(ctx, crate::AdjustPass::FORMAT, "detail-output"),
cache: HashMap::new(),
pool: Intermediates::new(),
no_instances: placeholder_instances(ctx),
}
}
@@ -230,6 +249,24 @@ impl DetailRunner {
usage: wgpu::BufferUsages::UNIFORM,
});
// TRACES: FR-DEV-8
// The instance list, uploaded only by the passes that have one. A
// kernel pass — which is every pass that is a convolution — is
// handed the placeholder allocated once in `new`, because a storage
// buffer of length zero is not bindable and allocating a fresh
// sixteen bytes per pass per frame is the per-frame allocation this
// module's documentation exists to refuse.
let instances = (!pass.storage.is_empty()).then(|| {
self.ctx
.device
.create_buffer_init(&wgpu::util::BufferInitDescriptor {
label: Some("detail-instances"),
contents: bytemuck::cast_slice(pass.storage.as_slice()),
usage: wgpu::BufferUsages::STORAGE,
})
});
let instances = instances.as_ref().unwrap_or(&self.no_instances);
let bind_group = self
.ctx
.device
@@ -249,6 +286,10 @@ impl DetailRunner {
binding: 2,
resource: wgpu::BindingResource::TextureView(destination),
},
wgpu::BindGroupEntry {
binding: 3,
resource: instances.as_entire_binding(),
},
],
});
@@ -338,6 +379,20 @@ impl DetailRunner {
}
}
/// A read-only storage buffer entry, as `mask.rs` declares its strokes.
fn storage_entry(binding: u32) -> wgpu::BindGroupLayoutEntry {
wgpu::BindGroupLayoutEntry {
binding,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::Buffer {
ty: wgpu::BufferBindingType::Storage { read_only: true },
has_dynamic_offset: false,
min_binding_size: None,
},
count: None,
}
}
impl Layout {
fn new(ctx: &GpuContext, format: wgpu::TextureFormat, label: &str) -> Self {
let bind_group = ctx
@@ -376,6 +431,14 @@ impl Layout {
},
count: None,
},
// TRACES: FR-DEV-8
// The instance list, for a pass whose work is a list rather
// than a kernel (`DetailPass::storage`). Every other pass
// gets `Intermediates`' placeholder here — one entry on both
// layouts rather than two more layouts, since a convolution
// that never reads the buffer costs nothing for it being
// bound.
storage_entry(3),
],
});