Poll the adjust readback instead of parking the UI thread

`read_output` waited on the copy with `Maintain::Wait`, which parks the
calling thread until the GPU is done. That call is made from the UI
thread, so the interface was frozen for the length of the copy — the
note on this function measures it at ~7 ms at 4K against a 0.28 ms
compute pass, so nearly all of it was the wait.

`Maintain::Poll` drives the same callbacks without sleeping. The mapping
still completes and the pixels are identical; the thread simply is not
parked while it happens.

The poll loop is bounded. A lost device never delivers the map callback,
and spinning forever on that would hang the app rather than report the
error the caller already handles.

This does not remove the round-trip itself, which ARCH §6.1 forbids and
spike S1 replaces by importing the texture into Slint directly. It stops
the round-trip from blocking input until then.

dr-gpu tests pass, including those comparing readback pixels.
This commit is contained in:
2026-08-09 21:52:59 +02:00
parent 31dd86d8c0
commit 250ff1e327
+35 -3
View File
@@ -31,6 +31,15 @@ use crate::{DemosaicedImage, GpuContext, GpuError};
/// reads them.
const RESERVED_FIELDS: usize = dr_pipeline::RESERVED_UNIFORM_FIELDS;
/// How many non-blocking polls a readback gets before it is called failed.
///
/// A bound rather than a spin forever: if the device is lost the map callback
/// never arrives, and an unbounded loop would hang the interface rather than
/// surfacing the error. Set far above any plausible completion — the copy this
/// waits on is milliseconds — so it is reached only when something is wrong.
#[cfg(any(test, feature = "readback"))]
const READBACK_POLL_LIMIT: u32 = 100_000;
/// Runs composed operation chains against demosaiced images.
pub struct AdjustPass {
ctx: GpuContext,
@@ -353,9 +362,32 @@ impl AdjustPass {
slice.map_async(wgpu::MapMode::Read, move |r| {
let _ = tx.send(r);
});
self.ctx.device.poll(wgpu::Maintain::Wait);
rx.recv()
.map_err(|e| GpuError::Readback(e.to_string()))?
// **Polled without blocking, then checked.**
//
// `Maintain::Wait` parks the calling thread until the GPU has finished,
// and this is called from the UI thread — so that park was a frozen
// interface for the duration of the copy (~7 ms at 4K, per the note
// above). `Poll` drives the same callbacks without sleeping, so the
// loop below stays interruptible and the mapping still completes.
//
// The bounded spin matters: a lost device would otherwise never
// deliver the callback and this would hang the app instead of
// reporting an error.
let mut mapped = None;
for _ in 0..READBACK_POLL_LIMIT {
self.ctx.device.poll(wgpu::Maintain::Poll);
match rx.try_recv() {
Ok(r) => {
mapped = Some(r);
break;
}
Err(std::sync::mpsc::TryRecvError::Empty) => continue,
Err(e) => return Err(GpuError::Readback(e.to_string())),
}
}
mapped
.ok_or_else(|| GpuError::Readback("readback did not complete".into()))?
.map_err(|e| GpuError::Readback(e.to_string()))?;
let data = slice.get_mapped_range();