Files
DarkRoom/core/dr-gpu/src/lib.rs
T
dtourolle 4b6c110816 Count the sensor's own numbers, so a cull can see headroom the render hides
FR-CULL-3's remaining two bullets. What existed was a *display* histogram
tagged FR-DSP-7: it binds AdjustPass's Rgba8Unorm output, recovers an 8-bit
code value, and counts clipping as `r == 255`. Its own documentation says a
clipped bin means "a highlight that is actually gone rather than one the
transform might still recover", which is the opposite of what a culling
decision needs. FR-CULL-3 asks for the histogram of the sensor data, on the
explicit grounds that a rendered image "systematically lies about what is
recoverable in the raw", and a readout that measures the render cannot answer
that however it is presented.

So this is a second instrument beside the first rather than a setting on it.
Both are true; they are true about different things; the panel offers both
behind a chip row and the words travel with the numbers, because a raw
saturation figure drawn under a heading saying Highlights would be mislabelled
exactly where the difference matters.

**What is reduced over, and what it cost to decide.** ARCH §5.5 specified the
pre-demosaic CFA samples. This reduces over the demosaiced scene-linear
texture instead, and §5.5 is amended to record the choice rather than let the
specification and the code disagree in silence. The texture is camera-native —
unbalanced, unmatrixed, uncurved — and normalised by the sensor's own black
and white levels, so 1.0 is saturation by construction and the distribution
below it is the headroom question with no calibration to carry. Retaining the
CFA samples would mean keeping the packed u32 buffer Demosaicer::run currently
drops: 48 MB at 24 MP, 120 MB at 60 MP, resident per open photograph whether
or not anyone looks at the histogram, on a platform §6.2 exists because memory
is scarce on.

Three things it therefore cannot say, written into the module docs and into
§5.5 rather than left to be discovered: it counts pixels not photosites, so a
saturated site drags its interpolated neighbours up and per-channel clipping
is smeared by about a demosaic kernel; it cannot see above white, because
demosaic.wgsl clamps each photosite at 1.0 for its own good reasons (a Canon
6D reads to 16383 against a declared 15070) so "at saturation" and "a stop
past it" share a bin; and it is measured after the CFA pattern is gone, so it
can name which colour clipped in the reconstructed image but not which
photosite went first.

The axis is stops below saturation, 16 bins per stop over 256 bins — the same
bin count the display reduction uses, so the fold into drawable columns is
shared and a divergence between the two plots would have to be deliberate. A
linear axis spends half its width on the top stop, which is why nobody has
ever drawn a useful linear raw histogram. The fourth series is the brightest
channel rather than luma: these values are unbalanced, so any weighted sum of
them is a number about nothing, and the brightest channel is the one that
saturates first and so the one the headroom question is actually about.

It is a property of the file and not of the render, which has two
consequences. It is computed once per photograph and cached — nothing
downstream of the demosaic can move a count in it — so a cull does not pay the
display histogram's per-frame cost three thousand times. And it describes the
whole frame rather than the visible region, deliberately opposite to
DevelopSession::histogram: a crop changes what is on screen and changes
nothing about what the sensor recorded.

Tags are on the reduction, the type, its constructor and the presentation
arithmetic, each of which has a test that fails if the behaviour goes. The
Slint panel and the push from lib.rs keep their reasoning as prose: nothing
asserts them, and a tag would claim coverage the assertions are not making.
2026-08-29 23:36:21 +02:00

813 lines
32 KiB
Rust

//! TRACES: NFR-PORT-2
//! GPU device and compute for DarkRoom.
//!
//! In v0.1 this exists to prove one thing: a compute shader can write a
//! texture that reaches the screen without a CPU round-trip (ARCH §6.1). It
//! holds no pipeline, no tiling, and no masks — those arrive in v0.2.
//!
//! Deliberately free of UI dependencies (ARCH §6.5a). The texture is handed
//! out as a `wgpu::Texture`; who composites it is not this crate's concern.
//!
//! That independence is why [`GpuContext::new_shared`] hands back the raw
//! instance and adapter rather than talking to a compositor itself: the
//! compositor will only sample a texture that came from the device *it* draws
//! with, so somebody has to make one device for both — but it does not have to
//! be this crate, and this crate must not know who it is.
use std::sync::Arc;
use wgpu::util::DeviceExt;
mod adjust;
mod demosaic;
mod detail;
mod error;
mod focus;
mod histogram;
mod mask;
mod raw_histogram;
mod readback;
mod segment;
pub use adjust::AdjustPass;
// The format the neighbourhood stage works in. Public because it is a promise
// rather than an implementation detail: a detail pass is guaranteed linear,
// unclipped, full internal precision (FR-DEV-2), and anyone reasoning about
// VRAM at 24 MP needs to know what an intermediate costs.
pub use demosaic::{DemosaicedImage, Demosaicer};
pub use detail::INTERMEDIATE_FORMAT as DETAIL_INTERMEDIATE_FORMAT;
pub use error::GpuError;
pub use focus::{FocusPeakPass, FocusPeaking, PeakColour, PeakSensitivity};
// Renamed on the way out: `BINS` says enough inside `histogram`, and nothing
// at all at a crate root shared with demosaic and segmentation.
pub use histogram::{Histogram, HistogramPass, BINS as HISTOGRAM_BINS};
pub use mask::{LabelField, MaskArray, MaskPass, SubjectMasks};
// Renamed on the way out on the same terms as the display histogram's
// constants above, and kept distinct from them because the two axes are
// different quantities: one counts output code values, the other counts stops
// below sensor saturation. A caller that confused them would draw a correct
// plot against the wrong scale.
pub use raw_histogram::{
RawHistogram, RawHistogramPass, BINS as RAW_HISTOGRAM_BINS,
BINS_PER_STOP as RAW_HISTOGRAM_BINS_PER_STOP, STOPS as RAW_HISTOGRAM_STOPS,
};
pub use segment::{SegmentOptions, SegmentPass, Segmentation};
/// Owns the wgpu device and queue.
///
/// One device is shared by the compute pipeline and the UI, which is what
/// allows compositing with no interop layer. Cloning is cheap and shares the
/// same underlying device.
#[derive(Clone)]
pub struct GpuContext {
pub device: Arc<wgpu::Device>,
pub queue: Arc<wgpu::Queue>,
adapter_info: wgpu::AdapterInfo,
}
/// TRACES: FR-DSP-1 | AC-8
/// One device, opened so that a compositor can be made to share it.
///
/// The texture the adjust pass writes only reaches the screen without a copy
/// if the compositor is drawing with the *same* `wgpu::Device` — two devices
/// are two address spaces, and a texture from one is not a texture the other
/// can sample. So the device cannot be an implementation detail of either
/// side; it has to be made once and handed to both.
///
/// [`Self::ctx`] is what the compute passes want. The instance and adapter are
/// what a compositor wants in order to adopt the same setup — Slint's
/// `WGPUConfiguration::Manual` asks for all four pieces — and they are handed
/// out raw rather than wrapped, because naming Slint here would put a UI
/// dependency in the one crate that must not have one (ARCH §6.5a).
pub struct SharedGpu {
/// The context every compute pass in this crate runs on.
pub ctx: GpuContext,
/// The instance the compositor will create its window surface from.
pub instance: wgpu::Instance,
/// The adapter [`Self::ctx`]'s device came from.
pub adapter: wgpu::Adapter,
}
/// TRACES: FR-DSP-1 | NFR-RES-4
/// Which GPU to prefer, on a machine with more than one.
///
/// **Not obviously the fastest one**, which is why this is a choice rather
/// than a constant. A discrete card wins on raw compute and loses on every
/// byte that has to reach it: a 24 MP frame is ~96 MB of RGBA, and each
/// upload and each export readback crosses PCIe. An integrated GPU shares
/// memory with the CPU, so those transfers are not transfers. It also does not
/// empty a laptop battery.
///
/// Which of those dominates depends on the work — a slider drag over a
/// resident texture is compute-bound and favours the discrete card, while
/// import, export and thumbnailing are transfer-heavy — so the honest thing is
/// to let it be set rather than to assume.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Default)]
pub enum AdapterPreference {
/// The most capable GPU. What this has always done, and the default: it is
/// the right answer for interactive editing, which is the frame budget
/// that FR-DSP-3 actually measures.
#[default]
Performance,
/// An integrated GPU where there is one — shared memory, no bus crossing,
/// and far less power.
Efficiency,
}
impl AdapterPreference {
/// Read the override, defaulting to [`Performance`](Self::Performance).
///
/// An environment variable rather than a setting, *for now*: this belongs
/// on the settings page beside the cache budget, and putting it there
/// needs a control and a restart prompt, because the device is opened once
/// at startup and shared with the compositor. The variable is what makes
/// the choice testable and gives someone with a broken primary GPU a way
/// out today.
pub fn from_env() -> Self {
match std::env::var("DARKROOM_GPU").as_deref() {
Ok("integrated") | Ok("efficiency") | Ok("igpu") => Self::Efficiency,
_ => Self::Performance,
}
}
/// How much we want an adapter, lowest first.
///
/// A CPU adapter sorts last under both policies rather than being
/// excluded: software rendering is a poor experience and a working one,
/// and on a machine where every real GPU has failed it is the difference
/// between a slow editor and no editor.
fn rank(self, device_type: wgpu::DeviceType) -> u8 {
use wgpu::DeviceType as D;
match (self, device_type) {
(_, D::Cpu) => 4,
(Self::Performance, D::DiscreteGpu) => 0,
(Self::Performance, D::IntegratedGpu) => 1,
(Self::Efficiency, D::IntegratedGpu) => 0,
(Self::Efficiency, D::DiscreteGpu) => 1,
(_, D::VirtualGpu) => 2,
(_, D::Other) => 3,
}
}
}
impl GpuContext {
/// Create a headless context — no surface, no window.
///
/// Used by tests, by the examples, and by anything that only needs to
/// compute. A context opened this way cannot be shared with a compositor:
/// see [`Self::new_shared`] for that, and for why the difference matters.
pub async fn new_headless() -> Result<Self, GpuError> {
// GL is allowed alongside Vulkan here and nowhere else: a machine with
// no Vulkan loader should still run the tests, and a headless context
// never has to produce a window surface — which is precisely the thing
// the GL backend cannot do from an instance opened without a display
// handle.
Self::open(wgpu::Backends::VULKAN | wgpu::Backends::GL)
.await
.map(|shared| shared.ctx)
}
/// TRACES: FR-DSP-1 | AC-8
/// Open a device intended to be shared with the compositor.
///
/// Vulkan only, unlike [`Self::new_headless`]. The caller will hand the
/// instance to a compositor that has to create a *window surface* from it,
/// and wgpu's GL backend reaches its display through EGL at instance
/// creation — an instance opened without a display handle, which is the
/// only kind available before a window exists, cannot then produce a GL
/// surface. Vulkan takes the window handle at surface creation instead, so
/// it is the only backend this order of operations permits.
///
/// A machine with no Vulkan therefore gets no shared device, and the
/// caller is expected to carry on without the develop path rather than
/// refuse to start.
pub async fn new_shared() -> Result<SharedGpu, GpuError> {
// Vulkan on both targets (D1), and here it is not merely the
// preference — see above.
Self::open(wgpu::Backends::VULKAN).await
}
/// TRACES: FR-DSP-1 | NFR-R1
/// Open a device, trying every adapter rather than only the best one.
///
/// # Why this is not `request_adapter`
///
/// `request_adapter` with `HighPerformance` returns *one* adapter and no
/// second chance. That is the right answer on a healthy machine and the
/// wrong one on a machine with a sick GPU, which is not a rare state:
/// observed 2026-08-29 on a laptop whose discrete card had hit an NVRM
/// assertion failure and a fullchip reset. The driver still advertised the
/// adapter, `request_adapter` dutifully picked it as the highest
/// performing, and the process died on it — while a working integrated GPU
/// and a working external card sat unused in the same enumeration.
///
/// A photo editor that will not start because the *fastest* GPU is broken,
/// on a machine holding two that are not, is worse than a slow one.
///
/// So: enumerate, order by how much we want each, and take the first that
/// actually yields a device. The ordering reproduces what
/// `HighPerformance` means — discrete, then integrated, then anything —
/// so the healthy case picks exactly what it picked before and pays one
/// extra enumeration for it.
///
/// # What this cannot do
///
/// A GPU sick enough to accept `request_device` and fail later is still
/// fatal, because the failure arrives as a segfault inside the driver
/// rather than as an error we could catch. This moves the boundary from
/// "the preferred adapter is unusable" to "the preferred adapter is
/// unusable *and* dishonest about it"; it does not remove it. Device loss
/// after a successful open is a different problem with a different answer
/// (ARCH §5.6).
async fn open(backends: wgpu::Backends) -> Result<SharedGpu, GpuError> {
// `new_without_display_handle` rather than a struct literal: the
// descriptor carries a boxed display handle and so has no `Default`,
// and there is no window yet to take one from in either case.
let mut descriptor = wgpu::InstanceDescriptor::new_without_display_handle();
descriptor.backends = backends;
let instance = wgpu::Instance::new(descriptor);
let mut adapters: Vec<wgpu::Adapter> = instance.enumerate_adapters(backends).await;
if adapters.is_empty() {
return Err(GpuError::NoAdapter);
}
let policy = AdapterPreference::from_env();
adapters.sort_by_key(|a| policy.rank(a.get_info().device_type));
// Kept so a total failure can say what it tried. "No suitable GPU
// adapter found" on a machine with three of them sends the reader to
// look for a driver that is installed and loaded.
let mut refusals: Vec<String> = Vec::new();
for adapter in adapters {
let adapter_info = adapter.get_info();
match Self::device_from(&adapter).await {
Ok((device, queue)) => {
log::info!(
"gpu: {} ({:?}, {:?})",
adapter_info.name,
adapter_info.device_type,
adapter_info.backend
);
if !refusals.is_empty() {
// At `info`, not `debug`: the user is now running on
// their second-choice GPU and any performance
// complaint that follows begins here.
log::info!(
"gpu: fell back after {} unusable adapter(s): {}",
refusals.len(),
refusals.join("; ")
);
}
return Ok(SharedGpu {
ctx: Self {
device: Arc::new(device),
queue: Arc::new(queue),
adapter_info,
},
instance,
adapter,
});
}
Err(e) => refusals.push(format!("{} ({e})", adapter_info.name)),
}
}
Err(GpuError::DeviceRequest(format!(
"every adapter refused a device: {}",
refusals.join("; ")
)))
}
/// Ask one adapter for a device, with the limits the pipeline needs.
async fn device_from(adapter: &wgpu::Adapter) -> Result<(wgpu::Device, wgpu::Queue), GpuError> {
adapter
.request_device(&wgpu::DeviceDescriptor {
label: Some("darkroom-device"),
required_features: wgpu::Features::empty(),
// Defaults, not `downlevel_defaults`: storage textures
// in compute shaders are required, and the downlevel tier
// does not guarantee them. This is effectively our GPU
// floor (NFR-COMPAT-1).
//
// `using_resolution` raises only the texture-dimension limits,
// to whatever this adapter actually offers. That matters once
// a compositor shares this device: the default ceiling is
// 8192, and a swapchain image for a large or scaled display
// can exceed it — a limit we chose for our own compute passes
// would otherwise silently cap somebody else's window.
required_limits: wgpu::Limits::default().using_resolution(adapter.limits()),
memory_hints: wgpu::MemoryHints::Performance,
// Nothing behind a feature flag wgpu itself calls unstable —
// the pipeline is ordinary compute and storage textures.
experimental_features: wgpu::ExperimentalFeatures::disabled(),
// The API trace, absorbed into the descriptor in wgpu 25 from
// the second argument this call used to take.
trace: wgpu::Trace::Off,
})
.await
.map_err(|e| GpuError::DeviceRequest(e.to_string()))
}
/// Build a context from a device and queue owned by someone else — the
/// path used when Slint has already created them.
pub fn from_parts(
device: Arc<wgpu::Device>,
queue: Arc<wgpu::Queue>,
adapter_info: wgpu::AdapterInfo,
) -> Self {
Self {
device,
queue,
adapter_info,
}
}
pub fn adapter_name(&self) -> &str {
&self.adapter_info.name
}
pub fn backend(&self) -> wgpu::Backend {
self.adapter_info.backend
}
}
#[repr(C)]
#[derive(Copy, Clone, Debug, bytemuck::Pod, bytemuck::Zeroable)]
struct Params {
width: u32,
height: u32,
phase: f32,
_pad: f32,
}
/// A compute pass writing into a storage texture.
///
/// Stands in for the develop pipeline in v0.1. What matters is the shape:
/// compute writes a texture, the texture is handed to the compositor, and
/// pixels never travel back through the CPU.
/// TRACES: FR-DEV-4 | R4
pub struct RenderTarget {
ctx: GpuContext,
texture: wgpu::Texture,
view: wgpu::TextureView,
pipeline: wgpu::ComputePipeline,
bind_group_layout: wgpu::BindGroupLayout,
bind_group: wgpu::BindGroup,
params_buf: wgpu::Buffer,
width: u32,
height: u32,
/// Reused staging buffer for the temporary readback path. Allocating one
/// per frame is a significant cost at large window sizes.
#[cfg(any(test, feature = "readback"))]
readback_buf: std::cell::RefCell<Option<(wgpu::Buffer, u32)>>,
}
impl RenderTarget {
pub const FORMAT: wgpu::TextureFormat = wgpu::TextureFormat::Rgba8Unorm;
pub fn new(ctx: &GpuContext, width: u32, height: u32) -> Result<Self, GpuError> {
let (width, height) = (width.max(1), height.max(1));
let shader = ctx
.device
.create_shader_module(wgpu::ShaderModuleDescriptor {
label: Some("gradient"),
source: wgpu::ShaderSource::Wgsl(include_str!("shaders/gradient.wgsl").into()),
});
let bind_group_layout =
ctx.device
.create_bind_group_layout(&wgpu::BindGroupLayoutDescriptor {
label: Some("render-target-bgl"),
entries: &[
wgpu::BindGroupLayoutEntry {
binding: 0,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::StorageTexture {
access: wgpu::StorageTextureAccess::WriteOnly,
format: Self::FORMAT,
view_dimension: wgpu::TextureViewDimension::D2,
},
count: None,
},
wgpu::BindGroupLayoutEntry {
binding: 1,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::Buffer {
ty: wgpu::BufferBindingType::Uniform,
has_dynamic_offset: false,
min_binding_size: None,
},
count: None,
},
],
});
let layout = ctx
.device
.create_pipeline_layout(&wgpu::PipelineLayoutDescriptor {
label: Some("render-target-layout"),
bind_group_layouts: &[Some(&bind_group_layout)],
immediate_size: 0,
});
let pipeline = ctx
.device
.create_compute_pipeline(&wgpu::ComputePipelineDescriptor {
label: Some("gradient-pipeline"),
layout: Some(&layout),
module: &shader,
entry_point: Some("main"),
compilation_options: Default::default(),
cache: None,
});
let params_buf = ctx
.device
.create_buffer_init(&wgpu::util::BufferInitDescriptor {
label: Some("params"),
contents: bytemuck::bytes_of(&Params {
width,
height,
phase: 0.0,
_pad: 0.0,
}),
usage: wgpu::BufferUsages::UNIFORM | wgpu::BufferUsages::COPY_DST,
});
let (texture, view) = Self::create_texture(ctx, width, height);
let bind_group = Self::create_bind_group(ctx, &bind_group_layout, &view, &params_buf);
Ok(Self {
ctx: ctx.clone(),
texture,
view,
pipeline,
bind_group_layout,
bind_group,
params_buf,
width,
height,
#[cfg(any(test, feature = "readback"))]
readback_buf: std::cell::RefCell::new(None),
})
}
fn create_texture(
ctx: &GpuContext,
width: u32,
height: u32,
) -> (wgpu::Texture, wgpu::TextureView) {
let texture = ctx.device.create_texture(&wgpu::TextureDescriptor {
label: Some("render-target"),
size: wgpu::Extent3d {
width,
height,
depth_or_array_layers: 1,
},
mip_level_count: 1,
sample_count: 1,
dimension: wgpu::TextureDimension::D2,
format: Self::FORMAT,
// STORAGE_BINDING to write from compute; TEXTURE_BINDING so the
// compositor can sample it. COPY_SRC exists only for tests —
// production never reads this back (ARCH §6.1).
//
// RENDER_ATTACHMENT is not something this pass ever uses. It is
// there because Slint refuses to import a texture without it
// (`TextureImportError::InvalidUsage`), the compositor having to
// assume it may need to draw into what it was given. Declaring an
// unused capability costs an allocation flag and buys the whole
// zero-copy path, so it is a cheap price for AC-8.
usage: wgpu::TextureUsages::STORAGE_BINDING
| wgpu::TextureUsages::TEXTURE_BINDING
| wgpu::TextureUsages::RENDER_ATTACHMENT
| wgpu::TextureUsages::COPY_SRC,
view_formats: &[],
});
let view = texture.create_view(&Default::default());
(texture, view)
}
fn create_bind_group(
ctx: &GpuContext,
layout: &wgpu::BindGroupLayout,
view: &wgpu::TextureView,
params: &wgpu::Buffer,
) -> wgpu::BindGroup {
ctx.device.create_bind_group(&wgpu::BindGroupDescriptor {
label: Some("render-target-bg"),
layout,
entries: &[
wgpu::BindGroupEntry {
binding: 0,
resource: wgpu::BindingResource::TextureView(view),
},
wgpu::BindGroupEntry {
binding: 1,
resource: params.as_entire_binding(),
},
],
})
}
/// Resize, reallocating the texture. No-op when unchanged.
pub fn resize(&mut self, width: u32, height: u32) {
let (width, height) = (width.max(1), height.max(1));
if width == self.width && height == self.height {
return;
}
let (texture, view) = Self::create_texture(&self.ctx, width, height);
self.bind_group =
Self::create_bind_group(&self.ctx, &self.bind_group_layout, &view, &self.params_buf);
self.texture = texture;
self.view = view;
self.width = width;
self.height = height;
#[cfg(any(test, feature = "readback"))]
{
// Size changed, so the staging buffer no longer fits.
*self.readback_buf.borrow_mut() = None;
}
}
/// Run the compute pass. Results stay on the GPU.
pub fn render(&self, phase: f32) {
self.ctx.queue.write_buffer(
&self.params_buf,
0,
bytemuck::bytes_of(&Params {
width: self.width,
height: self.height,
phase,
_pad: 0.0,
}),
);
let mut enc = self
.ctx
.device
.create_command_encoder(&wgpu::CommandEncoderDescriptor {
label: Some("render-encoder"),
});
{
let mut pass = enc.begin_compute_pass(&wgpu::ComputePassDescriptor {
label: Some("gradient-pass"),
timestamp_writes: None,
});
pass.set_pipeline(&self.pipeline);
pass.set_bind_group(0, &self.bind_group, &[]);
// 8x8 workgroups, rounded up so edge pixels are covered.
pass.dispatch_workgroups(self.width.div_ceil(8), self.height.div_ceil(8), 1);
}
self.ctx.queue.submit(Some(enc.finish()));
}
pub fn texture(&self) -> &wgpu::Texture {
&self.texture
}
pub fn view(&self) -> &wgpu::TextureView {
&self.view
}
pub fn size(&self) -> (u32, u32) {
(self.width, self.height)
}
/// Read pixels back to the CPU.
///
/// **Tests only.** Production code must never call this — it is exactly
/// the round-trip ARCH §6.1 forbids, and AC-8 asserts it does not happen.
#[cfg(any(test, feature = "readback"))]
pub async fn read_pixels(&self) -> Result<Vec<u8>, GpuError> {
// Buffer rows must be aligned to COPY_BYTES_PER_ROW_ALIGNMENT (256).
let unpadded = self.width * 4;
let align = wgpu::COPY_BYTES_PER_ROW_ALIGNMENT;
let padded = unpadded.div_ceil(align) * align;
let needed = (padded * self.height) as u64;
let mut slot = self.readback_buf.borrow_mut();
if slot.as_ref().map(|(_, p)| *p) != Some(padded) {
*slot = Some((
self.ctx.device.create_buffer(&wgpu::BufferDescriptor {
label: Some("readback"),
size: needed,
usage: wgpu::BufferUsages::COPY_DST | wgpu::BufferUsages::MAP_READ,
mapped_at_creation: false,
}),
padded,
));
}
let buf = &slot.as_ref().unwrap().0;
let mut enc = self.ctx.device.create_command_encoder(&Default::default());
enc.copy_texture_to_buffer(
wgpu::TexelCopyTextureInfo {
texture: &self.texture,
mip_level: 0,
origin: wgpu::Origin3d::ZERO,
aspect: wgpu::TextureAspect::All,
},
wgpu::TexelCopyBufferInfo {
buffer: buf,
layout: wgpu::TexelCopyBufferLayout {
offset: 0,
bytes_per_row: Some(padded),
rows_per_image: Some(self.height),
},
},
wgpu::Extent3d {
width: self.width,
height: self.height,
depth_or_array_layers: 1,
},
);
self.ctx.queue.submit(Some(enc.finish()));
let slice = buf.slice(..);
let (tx, rx) = std::sync::mpsc::channel();
slice.map_async(wgpu::MapMode::Read, move |r| {
let _ = tx.send(r);
});
// Fallible since wgpu 26, and worth propagating rather than ignoring:
// the failure it reports is a lost device (NFR-R7), and without this
// the map callback below simply never arrives and the error surfaces
// as a timeout somewhere less informative.
self.ctx
.device
.poll(wgpu::PollType::wait_indefinitely())
.map_err(|e| GpuError::Readback(e.to_string()))?;
rx.recv()
.map_err(|e| GpuError::Readback(e.to_string()))?
.map_err(|e| GpuError::Readback(e.to_string()))?;
// Strip row padding.
let data = slice.get_mapped_range();
let mut out = Vec::with_capacity((unpadded * self.height) as usize);
for row in 0..self.height {
let start = (row * padded) as usize;
out.extend_from_slice(&data[start..start + unpadded as usize]);
}
drop(data);
buf.unmap();
Ok(out)
}
}
#[cfg(test)]
mod tests {
use super::*;
fn ctx() -> Option<GpuContext> {
// CI runners and headless machines may have no usable adapter. Skip
// rather than fail — the device-dependent assertions still run
// wherever a GPU exists.
match pollster::block_on(GpuContext::new_headless()) {
Ok(c) => Some(c),
Err(e) => {
eprintln!("skipping: no GPU adapter ({e})");
None
}
}
}
#[test]
fn compute_writes_the_texture() {
let Some(ctx) = ctx() else { return };
let rt = RenderTarget::new(&ctx, 64, 64).expect("render target");
rt.render(0.0);
let px = pollster::block_on(rt.read_pixels()).expect("readback");
assert_eq!(px.len(), 64 * 64 * 4);
// The shader writes opaque pixels everywhere; an all-zero buffer would
// mean the dispatch silently did nothing.
assert!(
px.chunks_exact(4).all(|p| p[3] == 255),
"every pixel should be opaque"
);
assert!(
px.iter().any(|&b| b != 0),
"texture should not be uniformly zero"
);
}
#[test]
fn phase_changes_output() {
let Some(ctx) = ctx() else { return };
let rt = RenderTarget::new(&ctx, 32, 32).expect("render target");
rt.render(0.0);
let a = pollster::block_on(rt.read_pixels()).expect("readback");
rt.render(std::f32::consts::PI);
let b = pollster::block_on(rt.read_pixels()).expect("readback");
assert_ne!(a, b, "moving the highlight should change the image");
}
#[test]
fn resize_reallocates() {
let Some(ctx) = ctx() else { return };
let mut rt = RenderTarget::new(&ctx, 16, 16).expect("render target");
assert_eq!(rt.size(), (16, 16));
rt.resize(48, 24);
assert_eq!(rt.size(), (48, 24));
rt.render(0.0);
let px = pollster::block_on(rt.read_pixels()).expect("readback");
assert_eq!(px.len(), 48 * 24 * 4);
}
#[test]
fn zero_size_is_clamped() {
let Some(ctx) = ctx() else { return };
// A minimised window reports zero; texture creation would panic.
let rt = RenderTarget::new(&ctx, 0, 0).expect("render target");
assert_eq!(rt.size(), (1, 1));
}
}
#[cfg(test)]
mod adapter_choice_tests {
//! Which GPU gets picked, and what happens when it will not open.
//!
//! These are about the *ordering*, which is pure — opening a device needs
//! hardware and is covered by every other test in this crate implicitly.
use super::*;
/// Only the two fields the ordering reads are set; the rest come from
/// `Default`, so a wgpu upgrade that adds another does not break this.
/// The order adapters would be tried in, named so a failure reads as the
/// hardware it stands for.
fn order(policy: AdapterPreference, mut gpus: Vec<(&str, wgpu::DeviceType)>) -> Vec<&str> {
gpus.sort_by_key(|(_, t)| policy.rank(*t));
gpus.into_iter().map(|(name, _)| name).collect()
}
fn a_laptop() -> Vec<(&'static str, wgpu::DeviceType)> {
vec![
("Iris Xe", wgpu::DeviceType::IntegratedGpu),
("RTX 3050", wgpu::DeviceType::DiscreteGpu),
("llvmpipe", wgpu::DeviceType::Cpu),
]
}
#[test]
fn performance_takes_the_discrete_card() {
// What this has always done, and what an interactive slider drag wants:
// the texture is already resident, so the work is compute and the bus
// does not come into it.
assert_eq!(
order(AdapterPreference::Performance, a_laptop()),
["RTX 3050", "Iris Xe", "llvmpipe"]
);
}
#[test]
fn efficiency_takes_the_integrated_one() {
// Shared memory, so a 96 MB frame upload is not a transfer, and a
// laptop battery that lasts. The discrete card stays as the fallback
// rather than being excluded.
assert_eq!(
order(AdapterPreference::Efficiency, a_laptop()),
["Iris Xe", "RTX 3050", "llvmpipe"]
);
}
#[test]
fn software_rendering_is_last_but_never_dropped() {
// On a machine where every real GPU has failed this is the difference
// between a slow editor and no editor.
for policy in [
AdapterPreference::Performance,
AdapterPreference::Efficiency,
] {
assert_eq!(
*order(policy, a_laptop()).last().unwrap(),
"llvmpipe",
"{policy:?}"
);
}
}
#[test]
fn a_machine_with_one_gpu_is_unaffected_by_the_policy() {
// The common case: no choice to make, and no behaviour to change.
let one = vec![("Iris Xe", wgpu::DeviceType::IntegratedGpu)];
assert_eq!(
order(AdapterPreference::Performance, one.clone()),
order(AdapterPreference::Efficiency, one)
);
}
#[test]
fn the_default_is_what_it_did_before() {
// Changing which GPU an existing user lands on is not something to do
// by accident.
assert_eq!(AdapterPreference::default(), AdapterPreference::Performance);
}
}