Let a detail pass carry a list, not only a kernel
Every neighbourhood pass so far has been a convolution, whose whole description fits in the uniform block because its structure fixes how many numbers it needs. Spot removal is not that shape: sixty-four repairs and one repair are the same shader with a different buffer behind it. So a pass may declare `storage`, which arrives at binding 3 as `array<vec4<f32>>` with `arrayLength` in scope. The alternative — packing the list into uniforms — needs a fixed maximum paid for on every frame, a composer that can emit vec4 fields because a uniform array's stride is 16 whatever it holds, and it gives the next operation that wants a table nothing to build on. The property worth having is what stays out of the generated source: the count is in the buffer, so placing the tenth spot uploads 512 bytes and reuses the compiled pipeline, exactly as moving a slider does for the fused pass. `changing_the_list_does_not_recompile` is that, asserted. One bind group entry rather than two more layouts, and one placeholder buffer allocated in `new` rather than sixteen bytes per pass per frame — a zero-length storage buffer cannot be bound, and per-frame allocation is what this module's documentation exists to refuse. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -149,6 +149,8 @@ pub(crate) struct DetailRunner {
|
||||
/// Compiled pipelines by pass structure hash.
|
||||
cache: HashMap<u64, wgpu::ComputePipeline>,
|
||||
pool: Intermediates,
|
||||
/// See [`placeholder_instances`].
|
||||
no_instances: wgpu::Buffer,
|
||||
}
|
||||
|
||||
struct Layout {
|
||||
@@ -156,6 +158,22 @@ struct Layout {
|
||||
pipeline: wgpu::PipelineLayout,
|
||||
}
|
||||
|
||||
/// What binding 3 holds for a pass that declared no instance list.
|
||||
///
|
||||
/// One zeroed element, allocated once. Zero-length storage buffers cannot be
|
||||
/// bound, and the passes that read this binding are exactly the ones that
|
||||
/// uploaded something of their own, so nothing ever reads the placeholder's
|
||||
/// contents — it exists to keep one bind group layout serving both kinds of
|
||||
/// pass.
|
||||
fn placeholder_instances(ctx: &GpuContext) -> wgpu::Buffer {
|
||||
ctx.device
|
||||
.create_buffer_init(&wgpu::util::BufferInitDescriptor {
|
||||
label: Some("detail-instances-placeholder"),
|
||||
contents: bytemuck::cast_slice(&[[0.0f32; 4]]),
|
||||
usage: wgpu::BufferUsages::STORAGE,
|
||||
})
|
||||
}
|
||||
|
||||
impl DetailRunner {
|
||||
pub(crate) fn new(ctx: &GpuContext) -> Self {
|
||||
Self {
|
||||
@@ -164,6 +182,7 @@ impl DetailRunner {
|
||||
to_output: Layout::new(ctx, crate::AdjustPass::FORMAT, "detail-output"),
|
||||
cache: HashMap::new(),
|
||||
pool: Intermediates::new(),
|
||||
no_instances: placeholder_instances(ctx),
|
||||
}
|
||||
}
|
||||
|
||||
@@ -230,6 +249,24 @@ impl DetailRunner {
|
||||
usage: wgpu::BufferUsages::UNIFORM,
|
||||
});
|
||||
|
||||
// TRACES: FR-DEV-8
|
||||
// The instance list, uploaded only by the passes that have one. A
|
||||
// kernel pass — which is every pass that is a convolution — is
|
||||
// handed the placeholder allocated once in `new`, because a storage
|
||||
// buffer of length zero is not bindable and allocating a fresh
|
||||
// sixteen bytes per pass per frame is the per-frame allocation this
|
||||
// module's documentation exists to refuse.
|
||||
let instances = (!pass.storage.is_empty()).then(|| {
|
||||
self.ctx
|
||||
.device
|
||||
.create_buffer_init(&wgpu::util::BufferInitDescriptor {
|
||||
label: Some("detail-instances"),
|
||||
contents: bytemuck::cast_slice(pass.storage.as_slice()),
|
||||
usage: wgpu::BufferUsages::STORAGE,
|
||||
})
|
||||
});
|
||||
let instances = instances.as_ref().unwrap_or(&self.no_instances);
|
||||
|
||||
let bind_group = self
|
||||
.ctx
|
||||
.device
|
||||
@@ -249,6 +286,10 @@ impl DetailRunner {
|
||||
binding: 2,
|
||||
resource: wgpu::BindingResource::TextureView(destination),
|
||||
},
|
||||
wgpu::BindGroupEntry {
|
||||
binding: 3,
|
||||
resource: instances.as_entire_binding(),
|
||||
},
|
||||
],
|
||||
});
|
||||
|
||||
@@ -338,6 +379,20 @@ impl DetailRunner {
|
||||
}
|
||||
}
|
||||
|
||||
/// A read-only storage buffer entry, as `mask.rs` declares its strokes.
|
||||
fn storage_entry(binding: u32) -> wgpu::BindGroupLayoutEntry {
|
||||
wgpu::BindGroupLayoutEntry {
|
||||
binding,
|
||||
visibility: wgpu::ShaderStages::COMPUTE,
|
||||
ty: wgpu::BindingType::Buffer {
|
||||
ty: wgpu::BufferBindingType::Storage { read_only: true },
|
||||
has_dynamic_offset: false,
|
||||
min_binding_size: None,
|
||||
},
|
||||
count: None,
|
||||
}
|
||||
}
|
||||
|
||||
impl Layout {
|
||||
fn new(ctx: &GpuContext, format: wgpu::TextureFormat, label: &str) -> Self {
|
||||
let bind_group = ctx
|
||||
@@ -376,6 +431,14 @@ impl Layout {
|
||||
},
|
||||
count: None,
|
||||
},
|
||||
// TRACES: FR-DEV-8
|
||||
// The instance list, for a pass whose work is a list rather
|
||||
// than a kernel (`DetailPass::storage`). Every other pass
|
||||
// gets `Intermediates`' placeholder here — one entry on both
|
||||
// layouts rather than two more layouts, since a convolution
|
||||
// that never reads the buffer costs nothing for it being
|
||||
// bound.
|
||||
storage_entry(3),
|
||||
],
|
||||
});
|
||||
|
||||
|
||||
Reference in New Issue
Block a user