docs/ had 26 developer documents flat beside the manual, and the two audiences are very differently sized: most readers want the manual and the gesture reference, a few want the register, the designs and the measurements. The manual and gestures.md stay at the top; everything for someone changing the code moves to docs/dev/, and the two documents that name their own successors — the v0.1 milestone and the UI-refinement plan — go to docs/dev/archive/ rather than being deleted, since both are still cited. docs/README.md is the index, users first. Every reference follows: code comments, Cargo manifests, the workflows, the pre-commit hook, the bench and traceability tools (which locate the repo root by docs/dev/requirements.md now), packaging, the Docker READMEs, CLAUDE.md, CONTRIBUTING.md and the README. The matrix links one level deeper and is regenerated. Links out of the moved documents into the tree gain a level; a link checker over every Markdown file finds none broken.
342 lines
14 KiB
Rust
342 lines
14 KiB
Rust
//! The frame budget, asserted rather than hoped for.
|
|
//!
|
|
//! FR-DSP-3 says a slider updates the visible region within one frame budget at
|
|
//! proxy resolution. Until this file existed nothing checked it, which made it
|
|
//! a wish — `docs/dev/display-and-extension.md` §3 is blunt about that, and §7 is
|
|
//! blunt about what tagging an unchecked requirement does to the coverage
|
|
//! figure.
|
|
//!
|
|
//! The measurements this guards are in [`docs/dev/frame-budget.md`], produced by
|
|
//! `examples/frame_budget.rs`. This file is the part of them that has to keep
|
|
//! being true: it renders the **whole point-operation chain** through the real
|
|
//! `render_detailed` for a hundred frames, moving a slider between each, and
|
|
//! fails if the 99th percentile leaves the budget.
|
|
//!
|
|
//! # What it does not cover, said out loud
|
|
//!
|
|
//! **The neighbourhood stage is deliberately not in the asserted chain.** It is
|
|
//! over the budget today — clarity alone is 34 ms at 4K, because its kernel is
|
|
//! a fraction of the frame and reaches a 52-pixel radius there — and
|
|
//! `docs/dev/frame-budget.md` records that, names the fix (a base computed at
|
|
//! reduced resolution) and does not pretend otherwise. Asserting a budget the
|
|
//! code does not meet would produce a red suite that everyone learns to ignore;
|
|
//! asserting it on a chain that quietly excluded the expensive stage *without
|
|
//! saying so* would be the coverage overstatement §7 warns about. So it is
|
|
//! excluded, loudly, here.
|
|
//!
|
|
//! What is asserted is exactly the claim the FR-DSP-2 recommendation rests on:
|
|
//! that **one fused dispatch over a viewport-sized target is comfortably inside
|
|
//! the budget**, at a full chain, fit and at 1:1. If that stops being true, the
|
|
//! recommendation to strike tiled computation from the interactive path stops
|
|
//! being supported, and this test is what says so.
|
|
//!
|
|
//! # Why the percentile and not the mean
|
|
//!
|
|
//! A drag is judged by its worst frame. Nearest-rank over 100 frames puts the
|
|
//! 99th percentile at the second-worst, which is strict enough to catch a
|
|
//! stutter and forgiving enough that one scheduler hiccup from an unrelated
|
|
//! process does not decide the verdict.
|
|
//!
|
|
//! # Why the CPU half is asserted only in an optimised build
|
|
//!
|
|
//! Composing the shader is per-frame work on the UI thread and belongs in the
|
|
//! budget — `DevelopSession::render` calls `compose` on every frame, and on a
|
|
//! full chain it is milliseconds of string formatting. But the workspace builds
|
|
//! its own crates at `opt-level = 0` in dev (see the root `Cargo.toml`), and
|
|
//! `cargo test` is a dev build, so that formatting runs unoptimised here and
|
|
//! measures rustc rather than the pipeline. The GPU half is unaffected: a
|
|
//! shader is compiled by the driver either way.
|
|
//!
|
|
//! So the GPU half is always asserted, and the composition is folded in only
|
|
//! when `debug_assertions` is off. Running `cargo test --release -p dr-gpu`
|
|
//! therefore checks strictly more than the default run does, and the numbers
|
|
//! printed on failure say which of the two halves was over.
|
|
|
|
use dr_gpu::{AdjustPass, DemosaicedImage, GpuContext};
|
|
use dr_pipeline::descriptor::ParamKind;
|
|
use dr_pipeline::ops::exposure;
|
|
use dr_pipeline::{Affects, Attribute, CropRect, EditGraph, OpId, ParamId};
|
|
use std::time::Instant;
|
|
|
|
/// 60 Hz. FR-DSP-3 does not name a number; this is the one every interactive
|
|
/// application means by "one frame".
|
|
const BUDGET_MS: f64 = 16.0;
|
|
|
|
/// Measured frames per case. Nearest-rank p99 of 100 is the second-worst.
|
|
const FRAMES: usize = 100;
|
|
|
|
/// Discarded before measurement: the first frame at a size allocates a render
|
|
/// target and the first frame of a chain compiles a pipeline. Neither recurs
|
|
/// during a drag, so neither belongs in a drag's percentile.
|
|
const WARMUP: usize = 12;
|
|
|
|
/// A 24 MP source — a full-frame camera, and large enough that a 1:1 view of it
|
|
/// is a genuine zoom rather than a rounding error.
|
|
///
|
|
/// Smaller than the bench's 60 MP on purpose. The fused pass costs what the
|
|
/// *output* costs, so the source size barely moves these numbers, and 24 MP
|
|
/// keeps the fixture inside a second even at `opt-level = 0`.
|
|
const SOURCE: (u32, u32) = (6000, 4000);
|
|
|
|
/// The viewport the budget is asserted at: a 16:10 desktop display.
|
|
///
|
|
/// Not 4K, and the reason is worth stating. At 4K the fused chain still passes
|
|
/// with room to spare (4.5 ms of GPU; see `docs/dev/frame-budget.md`), but a test
|
|
/// that renders 8.3 M pixels a hundred times twice over is four seconds of
|
|
/// suite time to re-establish a conclusion 4.1 M pixels already establishes.
|
|
const VIEWPORT: (u32, u32) = (2560, 1600);
|
|
|
|
fn ctx() -> Option<GpuContext> {
|
|
// CI runners and headless machines may have no usable adapter. Skip rather
|
|
// than fail, exactly as the rest of this crate's device tests do.
|
|
match pollster::block_on(GpuContext::new_headless()) {
|
|
Ok(c) => Some(c),
|
|
Err(e) => {
|
|
eprintln!("skipping: no GPU adapter ({e})");
|
|
None
|
|
}
|
|
}
|
|
}
|
|
|
|
/// TRACES: FR-DSP-3 | FR-DSP-5
|
|
/// A slider drag on the full point-operation chain stays inside one frame —
|
|
/// fit, and at 1:1.
|
|
///
|
|
/// The develop view's ordinary case, at a full chain rather than a flattering
|
|
/// one: every operation that contributes a fragment to the fused shader is
|
|
/// active, and exposure moves between frames exactly as a drag moves it.
|
|
///
|
|
/// The 1:1 case is the one FR-DSP-5 names and the one FR-DSP-3's asynchronous
|
|
/// clause was written for. It is asserted here because the measurement found
|
|
/// that clause unnecessary rather than merely unimplemented: a 1:1 view is
|
|
/// *cheaper* than a fit view of the same file, since the dispatch is the same
|
|
/// size and the reads are contiguous rather than strided. If that ever inverts,
|
|
/// the argument for striking the clause weakens, and this is what would notice.
|
|
///
|
|
/// # One test and not two, deliberately
|
|
///
|
|
/// The two cases were two `#[test]` functions until the numbers said otherwise.
|
|
/// Cargo runs a binary's tests on a thread each, both of these want the same
|
|
/// GPU, and contending for it took the 1:1 case from 2.5 ms to 14.9 ms — a
|
|
/// measurement of the test harness that would have flickered either side of the
|
|
/// budget forever. A timing assertion has to own the device while it runs, and
|
|
/// the only way to say that in a test binary is to be the only test in it.
|
|
#[test]
|
|
fn a_slider_drag_stays_inside_the_frame_budget() {
|
|
let Some(ctx) = ctx() else { return };
|
|
let source = synthetic_source(&ctx);
|
|
|
|
let mut fit = full_point_chain();
|
|
drag(&ctx, &source, &mut fit, VIEWPORT).assert_inside_budget("proxy resolution, fit", VIEWPORT);
|
|
|
|
let mut zoomed = full_point_chain();
|
|
zoomed.framing_mut().set_view(one_to_one(VIEWPORT));
|
|
drag(&ctx, &source, &mut zoomed, VIEWPORT)
|
|
.assert_inside_budget("1:1 on a 24 MP source", VIEWPORT);
|
|
}
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// The measurement
|
|
// ---------------------------------------------------------------------------
|
|
|
|
struct Run {
|
|
/// Per frame: compose, compose the detail chain, hash the invalidation.
|
|
cpu_p99: f64,
|
|
/// Per frame: submit and wait for the device to go idle.
|
|
gpu_p99: f64,
|
|
/// `cpu + gpu` summed within each frame, then ranked. Not the sum of the
|
|
/// two percentiles above, which would be a frame that never happened.
|
|
total_p99: f64,
|
|
}
|
|
|
|
impl Run {
|
|
/// Fail if the budget was missed, saying which half missed it.
|
|
///
|
|
/// In a dev build only the GPU half is judged — see the module
|
|
/// documentation for why — and the CPU figure is still printed, because a
|
|
/// reader looking at a failure wants both numbers even when only one of
|
|
/// them is the verdict.
|
|
fn assert_inside_budget(&self, case: &str, viewport: (u32, u32)) {
|
|
let judged = if cfg!(debug_assertions) {
|
|
self.gpu_p99
|
|
} else {
|
|
self.total_p99
|
|
};
|
|
assert!(
|
|
judged <= BUDGET_MS,
|
|
"{case} at {}x{}: p99 of {FRAMES} frames was {judged:.2} ms, over the \
|
|
{BUDGET_MS:.0} ms budget (cpu {:.2} ms, gpu {:.2} ms, total {:.2} ms). \
|
|
FR-DSP-3 is what this violates; docs/dev/frame-budget.md holds the \
|
|
numbers it used to be.",
|
|
viewport.0,
|
|
viewport.1,
|
|
self.cpu_p99,
|
|
self.gpu_p99,
|
|
self.total_p99,
|
|
);
|
|
}
|
|
}
|
|
|
|
/// Render `FRAMES` frames with the exposure slider moving between each.
|
|
///
|
|
/// The three calls before the dispatch are the three `DevelopSession::render`
|
|
/// makes, in the same order, so this is the develop view's frame rather than an
|
|
/// idealisation of it.
|
|
fn drag(
|
|
ctx: &GpuContext,
|
|
source: &DemosaicedImage,
|
|
graph: &mut EditGraph,
|
|
viewport: (u32, u32),
|
|
) -> Run {
|
|
// Stands for the `VersionId` the app mixes in. Constant because every frame
|
|
// here is the same photograph.
|
|
const PHOTOGRAPH: u64 = 0x0dd_ba11;
|
|
|
|
let mut adjust = AdjustPass::new(ctx);
|
|
let src = source.size();
|
|
let (w, h) = viewport;
|
|
|
|
let mut cpu = Vec::with_capacity(FRAMES);
|
|
let mut gpu = Vec::with_capacity(FRAMES);
|
|
let mut total = Vec::with_capacity(FRAMES);
|
|
|
|
for i in 0..WARMUP + FRAMES {
|
|
// A hundredth of a stop per frame: what a drag does, and what stops
|
|
// `render_detailed` reusing the previous frame's colour result and
|
|
// turning this into a measurement of nothing.
|
|
graph.set_param(exposure::ID, exposure::EXPOSURE, 0.30 + i as f32 * 0.01);
|
|
|
|
let t0 = Instant::now();
|
|
let shader = graph.compose();
|
|
let detail = graph.compose_detail(src, viewport);
|
|
let colour_key = graph.invalidation().through(Affects::Colour) ^ PHOTOGRAPH;
|
|
let cpu_elapsed = t0.elapsed();
|
|
|
|
let t1 = Instant::now();
|
|
adjust
|
|
.render_detailed(source, &shader, w, h, None, &detail, colour_key)
|
|
.expect("render");
|
|
ctx.device
|
|
.poll(wgpu::PollType::wait_indefinitely())
|
|
.expect("poll");
|
|
let gpu_elapsed = t1.elapsed();
|
|
|
|
if i >= WARMUP {
|
|
cpu.push(cpu_elapsed.as_secs_f64() * 1e3);
|
|
gpu.push(gpu_elapsed.as_secs_f64() * 1e3);
|
|
total.push((cpu_elapsed + gpu_elapsed).as_secs_f64() * 1e3);
|
|
}
|
|
}
|
|
|
|
Run {
|
|
cpu_p99: p99(cpu),
|
|
gpu_p99: p99(gpu),
|
|
total_p99: p99(total),
|
|
}
|
|
}
|
|
|
|
/// Nearest-rank 99th percentile.
|
|
///
|
|
/// Nearest-rank because the samples *are* the population: there is no
|
|
/// distribution being estimated, only a hundred frames that either fitted in
|
|
/// the budget or did not.
|
|
fn p99(mut samples: Vec<f64>) -> f64 {
|
|
samples.sort_by(f64::total_cmp);
|
|
let n = samples.len();
|
|
samples[((0.99 * n as f64).ceil() as usize).clamp(1, n) - 1]
|
|
}
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Fixtures
|
|
// ---------------------------------------------------------------------------
|
|
|
|
/// Every operation that contributes a fragment to the fused shader, active.
|
|
///
|
|
/// Built from [`EditGraph::capabilities`] rather than from a list of operation
|
|
/// names, for the same reason the develop panel is: declaring a new node must
|
|
/// not silently shrink what this test calls "the full chain". A quarter of the
|
|
/// way from each parameter's default towards whichever end is further from it,
|
|
/// which is a plausible setting and — the part that matters — is never the
|
|
/// neutral, since a neutral operation contributes nothing at all.
|
|
///
|
|
/// The film stock is not reachable this way (it is a choice of material, not a
|
|
/// slider) and is left off. It is one texture lookup and two curve reads; the
|
|
/// bench includes it and it is worth about a millisecond at this size.
|
|
fn full_point_chain() -> EditGraph {
|
|
let mut graph = EditGraph::default_chain();
|
|
let moves: Vec<(OpId, ParamId, f32)> = graph
|
|
.capabilities()
|
|
.iter()
|
|
// The neighbourhood operations. See the module documentation for why
|
|
// they are not here.
|
|
.filter(|op| !op.attributes.contains(&Attribute::Detail))
|
|
.flat_map(|op| {
|
|
op.params.iter().map(move |p| {
|
|
let value = match p.kind {
|
|
ParamKind::Scalar { min, max, .. } => {
|
|
let far = if (max - p.default).abs() >= (p.default - min).abs() {
|
|
max
|
|
} else {
|
|
min
|
|
};
|
|
p.default + (far - p.default) * 0.25
|
|
}
|
|
ParamKind::Bool => 1.0,
|
|
// Bound by reference: a descriptor's variants became an
|
|
// owned `Vec` when descriptors stopped being `&'static`,
|
|
// and this arm only ever reads the length.
|
|
ParamKind::Enum { ref variants } => {
|
|
if variants.len() > 1 {
|
|
1.0
|
|
} else {
|
|
0.0
|
|
}
|
|
}
|
|
};
|
|
(op.id, p.id, value)
|
|
})
|
|
})
|
|
.collect();
|
|
for (op, param, value) in moves {
|
|
graph.set_param(op, param, value);
|
|
}
|
|
graph
|
|
}
|
|
|
|
/// The view rect that puts one render pixel on one source pixel, centred.
|
|
fn one_to_one(render: (u32, u32)) -> CropRect {
|
|
let w = render.0 as f32 / SOURCE.0 as f32;
|
|
let h = render.1 as f32 / SOURCE.1 as f32;
|
|
CropRect {
|
|
x: (1.0 - w) * 0.5,
|
|
y: (1.0 - h) * 0.5,
|
|
width: w,
|
|
height: h,
|
|
}
|
|
}
|
|
|
|
/// A source with structure at every scale.
|
|
///
|
|
/// Not flat: a flat frame lets the memory system serve every sample of every
|
|
/// pixel from one cache line, which flatters a bandwidth-bound pass by an amount
|
|
/// that has nothing to do with photographs.
|
|
fn synthetic_source(ctx: &GpuContext) -> DemosaicedImage {
|
|
let (w, h) = SOURCE;
|
|
let mut rgba = vec![0u8; (w as usize) * (h as usize) * 4];
|
|
for y in 0..h as usize {
|
|
let row = y * (w as usize) * 4;
|
|
for x in 0..w as usize {
|
|
let n = (x.wrapping_mul(2_654_435_761) ^ y.wrapping_mul(1_640_531_527)) >> 13;
|
|
let dither = (n & 0x1f) as u32;
|
|
let gx = (x * 200 / w as usize) as u32;
|
|
let gy = (y * 55 / h as usize) as u32;
|
|
let px = &mut rgba[row + x * 4..row + x * 4 + 4];
|
|
px[0] = (30 + gx + dither).min(255) as u8;
|
|
px[1] = (40 + gy + dither).min(255) as u8;
|
|
px[2] = (60 + gx / 2 + gy + dither).min(255) as u8;
|
|
px[3] = 255;
|
|
}
|
|
}
|
|
DemosaicedImage::from_rgba8(ctx, &rgba, w, h).expect("upload")
|
|
}
|