Show the develop frame itself, instead of a photocopy of it
The oldest open item in the project (ARCH §6.1, spike S1, AC-8). Every frame
in develop was read off the GPU into a `SharedPixelBuffer` and handed back to
Slint to upload again: ~7 ms at 4K against a 0.28 ms compute pass, 96% of the
frame spent carrying pixels to the CPU and back so they could be drawn where
they already were.
Slint 1.17 will adopt a `wgpu::Texture` directly, and the whole of what that
needs is arrangement rather than code.
**One device, made before the window.** A texture belongs to the device that
allocated it, so the compute passes and the compositor cannot each open their
own. `GpuContext::new_shared` opens one and hands back the instance and
adapter alongside it; `dr_ui::shared_gpu` gives all four to
`BackendSelector::require_wgpu_29(WGPUConfiguration::Manual { .. })`. That
call has to come before the first window, because creating one selects a
backend for you — which is why the GPU is now opened at the top of `run`
rather than two hundred lines down beside the other controllers.
dr-gpu still names no UI type. It hands out raw wgpu and does not ask who is
compositing (ARCH §6.5a).
**Vulkan only on the shared path**, where headless keeps its GL fallback.
wgpu's GL backend reaches its display through EGL at instance creation, and
before a window exists there is no display handle to give it — so a GL
instance cannot later produce the window surface Slint needs from it. A
machine with no Vulkan gets no shared device and browses without develop,
which is the same degradation as no adapter at all.
**`renderer-femtovg` becomes `renderer-femtovg-wgpu`.** The old one is FemtoVG
over OpenGL and cannot be handed a wgpu texture at all. It is not kept
alongside as a fallback: FemtoVG-over-GL has no branch for an imported
texture, falls through to "render this image into a buffer", gets nothing, and
draws nothing — a blank canvas with no error, which is worse than the failure
it would be papering over. The consequence is stated plainly in the manifest:
the desktop app now needs a working wgpu adapter to open a window.
**Two output textures, not one, and this is the part that is not obvious.**
Slint repaints when the image property *changes*, and it decides that with
`PartialEq` — which for two images over the same `wgpu::Texture` says
"unchanged". A pass that reused a single target would have rendered every
slider move correctly on the GPU and shown none of them: right, and invisible.
`AdjustPass` alternates between two targets, so consecutive frames are
genuinely different values. It also settles the read-while-write question that
one queue was already answering.
`RENDER_ATTACHMENT` is added to both render targets. Neither pass uses it;
Slint rejects an imported texture without it, on the reasoning that a
compositor handed a texture may need to draw into it.
**`AdjustPass::read_output` is deleted rather than gated.** It and
`export_pixels` were the same transfer under two names, and the comments
explaining why they were separate are the point of the whole criterion:
reading pixels back to *display* them is the defect, reading them back to
*encode a file* is the only way a file is made. The display twin is now gone
outright, which is stronger than a feature flag — it cannot be turned back on.
`export_pixels` is untouched and still ungated. The `readback` feature comes
off dr-ui, darkroom-desktop and darkroom-android; it stays in dr-gpu, where it
still gates `RenderTarget::read_pixels` and the segmentation field readback.
`examples/develop` moves to `export_pixels`, which is honest — it writes a
PPM — and so no longer needs the feature.
Four tests, each named for what it protects and each of which fails without a
screen if the property it guards breaks:
- the adjust target satisfies every condition Slint's import checks, asserted
in the crate that owns the descriptor, because a descriptor that drifts
fails at runtime on a real display and nothing else would notice;
- consecutive renders are different textures, and the third is the first
again, so the alternation is a rotation and not an allocation per frame;
- the develop canvas has no CPU pixel buffer and does have a wgpu texture —
AC-8 itself, in the terms Slint uses;
- consecutive frames compare unequal as `slint::Image`, which is the property
the repaint actually depends on.
The zoom test's readback moves into the test module. It has to: there is no
library function that copies a displayed frame to the CPU any more, and that
is the point — the round-trip now exists in the test binary and nowhere a
shipping build can reach.
**What is not proven.** No GUI was run. What is verified is that the texture
satisfies the import contract, that the import succeeds, that the canvas is a
texture rather than a buffer, and that consecutive frames are distinguishable.
What is unverified is everything that needs a display: that Slint's FemtoVG
wgpu renderer adopts the Manual configuration on a real surface, that the
picture appears the right way up and the right colour, and the frame timing
that motivated the whole exercise. Android is untouched by testing — the
android backend routes a WGPU29 request to Skia, whose wgpu surface does
handle imported textures, but that is read from the source, not observed.
56 dr-gpu tests and 255 dr-ui tests pass, clippy clean under `-D warnings`,
fmt clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
+90
-26
@@ -3,12 +3,17 @@
|
||||
//! A viewer with a develop panel: open a folder of RAW files, decode and
|
||||
//! demosaic on the GPU, and adjust.
|
||||
//!
|
||||
//! **Read before assuming A1 is proven.** Slint's public API for adopting an
|
||||
//! externally created wgpu texture is not wired up here; this build uploads
|
||||
//! through `SharedPixelBuffer`, which *is* a CPU round-trip — explicitly the
|
||||
//! thing ARCH §6.1 forbids in production. Spike S1 replaces it. Until then A1
|
||||
//! is unvalidated, and the develop path pays a readback per frame that the
|
||||
//! finished one will not.
|
||||
//! **The develop frame never leaves the GPU** (ARCH §6.1, AC-8). Spike S1
|
||||
//! wired Slint's texture import: [`shared_gpu`] opens one wgpu device and
|
||||
//! gives it to *both* the compute passes and Slint's renderer, and
|
||||
//! `DevelopSession::render` then hands the compositor the very texture the
|
||||
//! adjust pass wrote. What used to be a readback and an upload per frame is
|
||||
//! now a refcount.
|
||||
//!
|
||||
//! `SharedPixelBuffer` still appears in this file and in the library grid, and
|
||||
//! that is not a relapse: an embedded JPEG preview and a thumbnail are decoded
|
||||
//! on the CPU and have no texture to hand over. AC-8 is about pixels that were
|
||||
//! *computed on the GPU* travelling to the CPU and back to be looked at.
|
||||
//!
|
||||
//! **The develop panel is generated, not written.** [`develop`] asks the
|
||||
//! pipeline what parameters it has and builds a control per answer; no code
|
||||
@@ -576,6 +581,64 @@ enum PointsUpdate {
|
||||
Unchanged,
|
||||
}
|
||||
|
||||
/// TRACES: FR-DSP-1 | AC-8
|
||||
/// Open the one wgpu device the compute passes and the compositor share.
|
||||
///
|
||||
/// **This is the whole of the zero-copy display path, and it is four lines of
|
||||
/// configuration.** A `wgpu::Texture` belongs to the device that allocated it;
|
||||
/// handing one to a compositor drawing on a *different* device is meaningless,
|
||||
/// and the two would have to meet through system memory — which is the round
|
||||
/// trip ARCH §6.1 forbids. So there is exactly one device, made here, before
|
||||
/// anything else needs it.
|
||||
///
|
||||
/// **Called before the window exists, and it must be.** `BackendSelector`
|
||||
/// installs the Slint platform, and Slint installs a default one the first
|
||||
/// time a window is created; selecting afterwards is too late. That is why the
|
||||
/// GPU is opened at the top of [`run`] rather than beside the other
|
||||
/// controllers, where it used to sit.
|
||||
///
|
||||
/// `None` means develop is unavailable and the viewer falls back to embedded
|
||||
/// previews — the same degradation as a machine with no adapter at all.
|
||||
fn shared_gpu() -> Option<dr_gpu::GpuContext> {
|
||||
let shared = match pollster::block_on(dr_gpu::GpuContext::new_shared()) {
|
||||
Ok(shared) => shared,
|
||||
Err(e) => {
|
||||
log::warn!("no shareable GPU: {e}");
|
||||
return None;
|
||||
}
|
||||
};
|
||||
let dr_gpu::SharedGpu {
|
||||
ctx,
|
||||
instance,
|
||||
adapter,
|
||||
} = shared;
|
||||
|
||||
// `Manual` is the variant that means "render with these, do not open your
|
||||
// own". The two clones are of wgpu handles, which are refcounts over the
|
||||
// one device and the one queue — not copies of either.
|
||||
let configuration = slint::wgpu_29::WGPUConfiguration::Manual {
|
||||
instance,
|
||||
adapter,
|
||||
device: (*ctx.device).clone(),
|
||||
queue: (*ctx.queue).clone(),
|
||||
};
|
||||
|
||||
if let Err(e) = slint::BackendSelector::new()
|
||||
.require_wgpu_29(configuration)
|
||||
.select()
|
||||
{
|
||||
// Dropping the context rather than keeping it: Slint has fallen back
|
||||
// to a renderer that did not adopt our device, so every texture this
|
||||
// context produces is one the compositor cannot sample. A disabled
|
||||
// develop panel is a visible, explicable failure; a texture handed
|
||||
// across devices is undefined behaviour on a good day.
|
||||
log::warn!("Slint would not adopt the GPU device, develop disabled: {e}");
|
||||
return None;
|
||||
}
|
||||
|
||||
Some(ctx)
|
||||
}
|
||||
|
||||
/// TRACES: M-13 | M-14
|
||||
/// Build and run the viewer.
|
||||
pub fn run(paths: Vec<PathBuf>) -> Result<()> {
|
||||
@@ -587,7 +650,21 @@ pub fn run(paths: Vec<PathBuf>) -> Result<()> {
|
||||
let entries = Rc::new(RefCell::new(collect(&paths)));
|
||||
log::info!("{} image(s) to browse", entries.borrow().len());
|
||||
|
||||
// Before the window, and it has to be: this selects the Slint backend, and
|
||||
// creating a window selects one for us. See `shared_gpu`. The device is
|
||||
// shared by demosaic, the adjust pass and the compositor; without one the
|
||||
// app still browses through the preview path, just without develop.
|
||||
let gpu = shared_gpu();
|
||||
|
||||
let window = AppWindow::new()?;
|
||||
match &gpu {
|
||||
Some(ctx) => {
|
||||
log::info!("adapter: {} ({:?})", ctx.adapter_name(), ctx.backend());
|
||||
window.set_adapter(ctx.adapter_name().into());
|
||||
window.set_backend(format!("{:?}", ctx.backend()).to_uppercase().into());
|
||||
}
|
||||
None => window.set_backend("NO GPU".into()),
|
||||
}
|
||||
|
||||
// Every background job reports here, and this draws the bar across the top
|
||||
// of the shell and fills the settings page's list. Built before the
|
||||
@@ -913,22 +990,6 @@ pub fn run(paths: Vec<PathBuf>) -> Result<()> {
|
||||
});
|
||||
}
|
||||
|
||||
// The device is shared by demosaic and the adjust pass. Without one the
|
||||
// app still browses through the preview path, just without develop.
|
||||
let gpu = match pollster::block_on(dr_gpu::GpuContext::new_headless()) {
|
||||
Ok(ctx) => {
|
||||
log::info!("adapter: {} ({:?})", ctx.adapter_name(), ctx.backend());
|
||||
window.set_adapter(ctx.adapter_name().into());
|
||||
window.set_backend(format!("{:?}", ctx.backend()).to_uppercase().into());
|
||||
Some(ctx)
|
||||
}
|
||||
Err(e) => {
|
||||
log::warn!("no GPU adapter: {e}");
|
||||
window.set_backend("NO GPU".into());
|
||||
None
|
||||
}
|
||||
};
|
||||
|
||||
window.set_total(entries.borrow().len() as i32);
|
||||
let index = Rc::new(RefCell::new(0usize));
|
||||
// The current develop session, if the file yielded sensor data.
|
||||
@@ -965,9 +1026,11 @@ pub fn run(paths: Vec<PathBuf>) -> Result<()> {
|
||||
|
||||
// **Half resolution while the gesture is still moving.**
|
||||
//
|
||||
// The adjust pass and the readback both scale with pixel count, so
|
||||
// halving each edge is roughly a quarter of the work — the
|
||||
// difference between keeping up with a drag and lagging behind it.
|
||||
// The adjust pass scales with pixel count, so halving each edge is
|
||||
// roughly a quarter of the work — the difference between keeping
|
||||
// up with a drag and lagging behind it. Less dramatic since S1
|
||||
// removed the readback that scaled the same way and cost far more,
|
||||
// but a dispatch is still not free at 4K.
|
||||
// A draft frame is visible for one gesture and is replaced by a
|
||||
// full-resolution one the moment motion stops, so the cost is a
|
||||
// little softness exactly while the image is moving too fast to
|
||||
@@ -1013,7 +1076,8 @@ pub fn run(paths: Vec<PathBuf>) -> Result<()> {
|
||||
|
||||
// **Rendering is decoupled from input, and this is why.**
|
||||
//
|
||||
// A render is a blocking GPU round-trip (see `AdjustPass::read_output`).
|
||||
// A render used to be a blocking GPU round-trip — S1 removed the block,
|
||||
// but not the reason for this, so read it as history that still applies.
|
||||
// Running one straight from a `moved` handler put that stall *inside* the
|
||||
// gesture: touch events arrive far faster than a render completes, so the
|
||||
// input queue backed up, positions arrived stale, and Android — seeing the
|
||||
|
||||
Reference in New Issue
Block a user