//! The frame budget, asserted rather than hoped for. //! //! FR-DSP-3 says a slider updates the visible region within one frame budget at //! proxy resolution. Until this file existed nothing checked it, which made it //! a wish — `docs/display-and-extension.md` §3 is blunt about that, and §7 is //! blunt about what tagging an unchecked requirement does to the coverage //! figure. //! //! The measurements this guards are in [`docs/frame-budget.md`], produced by //! `examples/frame_budget.rs`. This file is the part of them that has to keep //! being true: it renders the **whole point-operation chain** through the real //! `render_detailed` for a hundred frames, moving a slider between each, and //! fails if the 99th percentile leaves the budget. //! //! # What it does not cover, said out loud //! //! **The neighbourhood stage is deliberately not in the asserted chain.** It is //! over the budget today — clarity alone is 34 ms at 4K, because its kernel is //! a fraction of the frame and reaches a 52-pixel radius there — and //! `docs/frame-budget.md` records that, names the fix (a base computed at //! reduced resolution) and does not pretend otherwise. Asserting a budget the //! code does not meet would produce a red suite that everyone learns to ignore; //! asserting it on a chain that quietly excluded the expensive stage *without //! saying so* would be the coverage overstatement §7 warns about. So it is //! excluded, loudly, here. //! //! What is asserted is exactly the claim the FR-DSP-2 recommendation rests on: //! that **one fused dispatch over a viewport-sized target is comfortably inside //! the budget**, at a full chain, fit and at 1:1. If that stops being true, the //! recommendation to strike tiled computation from the interactive path stops //! being supported, and this test is what says so. //! //! # Why the percentile and not the mean //! //! A drag is judged by its worst frame. Nearest-rank over 100 frames puts the //! 99th percentile at the second-worst, which is strict enough to catch a //! stutter and forgiving enough that one scheduler hiccup from an unrelated //! process does not decide the verdict. //! //! # Why the CPU half is asserted only in an optimised build //! //! Composing the shader is per-frame work on the UI thread and belongs in the //! budget — `DevelopSession::render` calls `compose` on every frame, and on a //! full chain it is milliseconds of string formatting. But the workspace builds //! its own crates at `opt-level = 0` in dev (see the root `Cargo.toml`), and //! `cargo test` is a dev build, so that formatting runs unoptimised here and //! measures rustc rather than the pipeline. The GPU half is unaffected: a //! shader is compiled by the driver either way. //! //! So the GPU half is always asserted, and the composition is folded in only //! when `debug_assertions` is off. Running `cargo test --release -p dr-gpu` //! therefore checks strictly more than the default run does, and the numbers //! printed on failure say which of the two halves was over. use dr_gpu::{AdjustPass, DemosaicedImage, GpuContext}; use dr_pipeline::descriptor::ParamKind; use dr_pipeline::ops::exposure; use dr_pipeline::{Affects, Attribute, CropRect, EditGraph, OpId, ParamId}; use std::time::Instant; /// 60 Hz. FR-DSP-3 does not name a number; this is the one every interactive /// application means by "one frame". const BUDGET_MS: f64 = 16.0; /// Measured frames per case. Nearest-rank p99 of 100 is the second-worst. const FRAMES: usize = 100; /// Discarded before measurement: the first frame at a size allocates a render /// target and the first frame of a chain compiles a pipeline. Neither recurs /// during a drag, so neither belongs in a drag's percentile. const WARMUP: usize = 12; /// A 24 MP source — a full-frame camera, and large enough that a 1:1 view of it /// is a genuine zoom rather than a rounding error. /// /// Smaller than the bench's 60 MP on purpose. The fused pass costs what the /// *output* costs, so the source size barely moves these numbers, and 24 MP /// keeps the fixture inside a second even at `opt-level = 0`. const SOURCE: (u32, u32) = (6000, 4000); /// The viewport the budget is asserted at: a 16:10 desktop display. /// /// Not 4K, and the reason is worth stating. At 4K the fused chain still passes /// with room to spare (4.5 ms of GPU; see `docs/frame-budget.md`), but a test /// that renders 8.3 M pixels a hundred times twice over is four seconds of /// suite time to re-establish a conclusion 4.1 M pixels already establishes. const VIEWPORT: (u32, u32) = (2560, 1600); fn ctx() -> Option { // CI runners and headless machines may have no usable adapter. Skip rather // than fail, exactly as the rest of this crate's device tests do. match pollster::block_on(GpuContext::new_headless()) { Ok(c) => Some(c), Err(e) => { eprintln!("skipping: no GPU adapter ({e})"); None } } } /// TRACES: FR-DSP-3 | FR-DSP-5 /// A slider drag on the full point-operation chain stays inside one frame — /// fit, and at 1:1. /// /// The develop view's ordinary case, at a full chain rather than a flattering /// one: every operation that contributes a fragment to the fused shader is /// active, and exposure moves between frames exactly as a drag moves it. /// /// The 1:1 case is the one FR-DSP-5 names and the one FR-DSP-3's asynchronous /// clause was written for. It is asserted here because the measurement found /// that clause unnecessary rather than merely unimplemented: a 1:1 view is /// *cheaper* than a fit view of the same file, since the dispatch is the same /// size and the reads are contiguous rather than strided. If that ever inverts, /// the argument for striking the clause weakens, and this is what would notice. /// /// # One test and not two, deliberately /// /// The two cases were two `#[test]` functions until the numbers said otherwise. /// Cargo runs a binary's tests on a thread each, both of these want the same /// GPU, and contending for it took the 1:1 case from 2.5 ms to 14.9 ms — a /// measurement of the test harness that would have flickered either side of the /// budget forever. A timing assertion has to own the device while it runs, and /// the only way to say that in a test binary is to be the only test in it. #[test] fn a_slider_drag_stays_inside_the_frame_budget() { let Some(ctx) = ctx() else { return }; let source = synthetic_source(&ctx); let mut fit = full_point_chain(); drag(&ctx, &source, &mut fit, VIEWPORT).assert_inside_budget("proxy resolution, fit", VIEWPORT); let mut zoomed = full_point_chain(); zoomed.framing_mut().set_view(one_to_one(VIEWPORT)); drag(&ctx, &source, &mut zoomed, VIEWPORT) .assert_inside_budget("1:1 on a 24 MP source", VIEWPORT); } // --------------------------------------------------------------------------- // The measurement // --------------------------------------------------------------------------- struct Run { /// Per frame: compose, compose the detail chain, hash the invalidation. cpu_p99: f64, /// Per frame: submit and wait for the device to go idle. gpu_p99: f64, /// `cpu + gpu` summed within each frame, then ranked. Not the sum of the /// two percentiles above, which would be a frame that never happened. total_p99: f64, } impl Run { /// Fail if the budget was missed, saying which half missed it. /// /// In a dev build only the GPU half is judged — see the module /// documentation for why — and the CPU figure is still printed, because a /// reader looking at a failure wants both numbers even when only one of /// them is the verdict. fn assert_inside_budget(&self, case: &str, viewport: (u32, u32)) { let judged = if cfg!(debug_assertions) { self.gpu_p99 } else { self.total_p99 }; assert!( judged <= BUDGET_MS, "{case} at {}x{}: p99 of {FRAMES} frames was {judged:.2} ms, over the \ {BUDGET_MS:.0} ms budget (cpu {:.2} ms, gpu {:.2} ms, total {:.2} ms). \ FR-DSP-3 is what this violates; docs/frame-budget.md holds the \ numbers it used to be.", viewport.0, viewport.1, self.cpu_p99, self.gpu_p99, self.total_p99, ); } } /// Render `FRAMES` frames with the exposure slider moving between each. /// /// The three calls before the dispatch are the three `DevelopSession::render` /// makes, in the same order, so this is the develop view's frame rather than an /// idealisation of it. fn drag( ctx: &GpuContext, source: &DemosaicedImage, graph: &mut EditGraph, viewport: (u32, u32), ) -> Run { // Stands for the `VersionId` the app mixes in. Constant because every frame // here is the same photograph. const PHOTOGRAPH: u64 = 0x0dd_ba11; let mut adjust = AdjustPass::new(ctx); let src = source.size(); let (w, h) = viewport; let mut cpu = Vec::with_capacity(FRAMES); let mut gpu = Vec::with_capacity(FRAMES); let mut total = Vec::with_capacity(FRAMES); for i in 0..WARMUP + FRAMES { // A hundredth of a stop per frame: what a drag does, and what stops // `render_detailed` reusing the previous frame's colour result and // turning this into a measurement of nothing. graph.set_param(exposure::ID, exposure::EXPOSURE, 0.30 + i as f32 * 0.01); let t0 = Instant::now(); let shader = graph.compose(); let detail = graph.compose_detail(src, viewport); let colour_key = graph.invalidation().through(Affects::Colour) ^ PHOTOGRAPH; let cpu_elapsed = t0.elapsed(); let t1 = Instant::now(); adjust .render_detailed(source, &shader, w, h, None, &detail, colour_key) .expect("render"); ctx.device .poll(wgpu::PollType::wait_indefinitely()) .expect("poll"); let gpu_elapsed = t1.elapsed(); if i >= WARMUP { cpu.push(cpu_elapsed.as_secs_f64() * 1e3); gpu.push(gpu_elapsed.as_secs_f64() * 1e3); total.push((cpu_elapsed + gpu_elapsed).as_secs_f64() * 1e3); } } Run { cpu_p99: p99(cpu), gpu_p99: p99(gpu), total_p99: p99(total), } } /// Nearest-rank 99th percentile. /// /// Nearest-rank because the samples *are* the population: there is no /// distribution being estimated, only a hundred frames that either fitted in /// the budget or did not. fn p99(mut samples: Vec) -> f64 { samples.sort_by(f64::total_cmp); let n = samples.len(); samples[((0.99 * n as f64).ceil() as usize).clamp(1, n) - 1] } // --------------------------------------------------------------------------- // Fixtures // --------------------------------------------------------------------------- /// Every operation that contributes a fragment to the fused shader, active. /// /// Built from [`EditGraph::capabilities`] rather than from a list of operation /// names, for the same reason the develop panel is: declaring a new node must /// not silently shrink what this test calls "the full chain". A quarter of the /// way from each parameter's default towards whichever end is further from it, /// which is a plausible setting and — the part that matters — is never the /// neutral, since a neutral operation contributes nothing at all. /// /// The film stock is not reachable this way (it is a choice of material, not a /// slider) and is left off. It is one texture lookup and two curve reads; the /// bench includes it and it is worth about a millisecond at this size. fn full_point_chain() -> EditGraph { let mut graph = EditGraph::default_chain(); let moves: Vec<(OpId, ParamId, f32)> = graph .capabilities() .iter() // The neighbourhood operations. See the module documentation for why // they are not here. .filter(|op| !op.attributes.contains(&Attribute::Detail)) .flat_map(|op| { op.params.iter().map(move |p| { let value = match p.kind { ParamKind::Scalar { min, max, .. } => { let far = if (max - p.default).abs() >= (p.default - min).abs() { max } else { min }; p.default + (far - p.default) * 0.25 } ParamKind::Bool => 1.0, // Bound by reference: a descriptor's variants became an // owned `Vec` when descriptors stopped being `&'static`, // and this arm only ever reads the length. ParamKind::Enum { ref variants } => { if variants.len() > 1 { 1.0 } else { 0.0 } } }; (op.id, p.id, value) }) }) .collect(); for (op, param, value) in moves { graph.set_param(op, param, value); } graph } /// The view rect that puts one render pixel on one source pixel, centred. fn one_to_one(render: (u32, u32)) -> CropRect { let w = render.0 as f32 / SOURCE.0 as f32; let h = render.1 as f32 / SOURCE.1 as f32; CropRect { x: (1.0 - w) * 0.5, y: (1.0 - h) * 0.5, width: w, height: h, } } /// A source with structure at every scale. /// /// Not flat: a flat frame lets the memory system serve every sample of every /// pixel from one cache line, which flatters a bandwidth-bound pass by an amount /// that has nothing to do with photographs. fn synthetic_source(ctx: &GpuContext) -> DemosaicedImage { let (w, h) = SOURCE; let mut rgba = vec![0u8; (w as usize) * (h as usize) * 4]; for y in 0..h as usize { let row = y * (w as usize) * 4; for x in 0..w as usize { let n = (x.wrapping_mul(2_654_435_761) ^ y.wrapping_mul(1_640_531_527)) >> 13; let dither = (n & 0x1f) as u32; let gx = (x * 200 / w as usize) as u32; let gy = (y * 55 / h as usize) as u32; let px = &mut rgba[row + x * 4..row + x * 4 + 4]; px[0] = (30 + gx + dither).min(255) as u8; px[1] = (40 + gy + dither).min(255) as u8; px[2] = (60 + gx / 2 + gy + dither).min(255) as u8; px[3] = 255; } } DemosaicedImage::from_rgba8(ctx, &rgba, w, h).expect("upload") }