Files
dtourolle 84fade99ec Put the developer docs under docs/dev and index the folder for users first
docs/ had 26 developer documents flat beside the manual, and the two
audiences are very differently sized: most readers want the manual and
the gesture reference, a few want the register, the designs and the
measurements. The manual and gestures.md stay at the top; everything for
someone changing the code moves to docs/dev/, and the two documents that
name their own successors — the v0.1 milestone and the UI-refinement plan
— go to docs/dev/archive/ rather than being deleted, since both are still
cited. docs/README.md is the index, users first.

Every reference follows: code comments, Cargo manifests, the workflows,
the pre-commit hook, the bench and traceability tools (which locate the
repo root by docs/dev/requirements.md now), packaging, the Docker READMEs,
CLAUDE.md, CONTRIBUTING.md and the README. The matrix links one level
deeper and is regenerated. Links out of the moved documents into the tree
gain a level; a link checker over every Markdown file finds none broken.
2026-09-20 21:16:03 +02:00

342 lines
14 KiB
Rust

//! The frame budget, asserted rather than hoped for.
//!
//! FR-DSP-3 says a slider updates the visible region within one frame budget at
//! proxy resolution. Until this file existed nothing checked it, which made it
//! a wish — `docs/dev/display-and-extension.md` §3 is blunt about that, and §7 is
//! blunt about what tagging an unchecked requirement does to the coverage
//! figure.
//!
//! The measurements this guards are in [`docs/dev/frame-budget.md`], produced by
//! `examples/frame_budget.rs`. This file is the part of them that has to keep
//! being true: it renders the **whole point-operation chain** through the real
//! `render_detailed` for a hundred frames, moving a slider between each, and
//! fails if the 99th percentile leaves the budget.
//!
//! # What it does not cover, said out loud
//!
//! **The neighbourhood stage is deliberately not in the asserted chain.** It is
//! over the budget today — clarity alone is 34 ms at 4K, because its kernel is
//! a fraction of the frame and reaches a 52-pixel radius there — and
//! `docs/dev/frame-budget.md` records that, names the fix (a base computed at
//! reduced resolution) and does not pretend otherwise. Asserting a budget the
//! code does not meet would produce a red suite that everyone learns to ignore;
//! asserting it on a chain that quietly excluded the expensive stage *without
//! saying so* would be the coverage overstatement §7 warns about. So it is
//! excluded, loudly, here.
//!
//! What is asserted is exactly the claim the FR-DSP-2 recommendation rests on:
//! that **one fused dispatch over a viewport-sized target is comfortably inside
//! the budget**, at a full chain, fit and at 1:1. If that stops being true, the
//! recommendation to strike tiled computation from the interactive path stops
//! being supported, and this test is what says so.
//!
//! # Why the percentile and not the mean
//!
//! A drag is judged by its worst frame. Nearest-rank over 100 frames puts the
//! 99th percentile at the second-worst, which is strict enough to catch a
//! stutter and forgiving enough that one scheduler hiccup from an unrelated
//! process does not decide the verdict.
//!
//! # Why the CPU half is asserted only in an optimised build
//!
//! Composing the shader is per-frame work on the UI thread and belongs in the
//! budget — `DevelopSession::render` calls `compose` on every frame, and on a
//! full chain it is milliseconds of string formatting. But the workspace builds
//! its own crates at `opt-level = 0` in dev (see the root `Cargo.toml`), and
//! `cargo test` is a dev build, so that formatting runs unoptimised here and
//! measures rustc rather than the pipeline. The GPU half is unaffected: a
//! shader is compiled by the driver either way.
//!
//! So the GPU half is always asserted, and the composition is folded in only
//! when `debug_assertions` is off. Running `cargo test --release -p dr-gpu`
//! therefore checks strictly more than the default run does, and the numbers
//! printed on failure say which of the two halves was over.
use dr_gpu::{AdjustPass, DemosaicedImage, GpuContext};
use dr_pipeline::descriptor::ParamKind;
use dr_pipeline::ops::exposure;
use dr_pipeline::{Affects, Attribute, CropRect, EditGraph, OpId, ParamId};
use std::time::Instant;
/// 60 Hz. FR-DSP-3 does not name a number; this is the one every interactive
/// application means by "one frame".
const BUDGET_MS: f64 = 16.0;
/// Measured frames per case. Nearest-rank p99 of 100 is the second-worst.
const FRAMES: usize = 100;
/// Discarded before measurement: the first frame at a size allocates a render
/// target and the first frame of a chain compiles a pipeline. Neither recurs
/// during a drag, so neither belongs in a drag's percentile.
const WARMUP: usize = 12;
/// A 24 MP source — a full-frame camera, and large enough that a 1:1 view of it
/// is a genuine zoom rather than a rounding error.
///
/// Smaller than the bench's 60 MP on purpose. The fused pass costs what the
/// *output* costs, so the source size barely moves these numbers, and 24 MP
/// keeps the fixture inside a second even at `opt-level = 0`.
const SOURCE: (u32, u32) = (6000, 4000);
/// The viewport the budget is asserted at: a 16:10 desktop display.
///
/// Not 4K, and the reason is worth stating. At 4K the fused chain still passes
/// with room to spare (4.5 ms of GPU; see `docs/dev/frame-budget.md`), but a test
/// that renders 8.3 M pixels a hundred times twice over is four seconds of
/// suite time to re-establish a conclusion 4.1 M pixels already establishes.
const VIEWPORT: (u32, u32) = (2560, 1600);
fn ctx() -> Option<GpuContext> {
// CI runners and headless machines may have no usable adapter. Skip rather
// than fail, exactly as the rest of this crate's device tests do.
match pollster::block_on(GpuContext::new_headless()) {
Ok(c) => Some(c),
Err(e) => {
eprintln!("skipping: no GPU adapter ({e})");
None
}
}
}
/// TRACES: FR-DSP-3 | FR-DSP-5
/// A slider drag on the full point-operation chain stays inside one frame —
/// fit, and at 1:1.
///
/// The develop view's ordinary case, at a full chain rather than a flattering
/// one: every operation that contributes a fragment to the fused shader is
/// active, and exposure moves between frames exactly as a drag moves it.
///
/// The 1:1 case is the one FR-DSP-5 names and the one FR-DSP-3's asynchronous
/// clause was written for. It is asserted here because the measurement found
/// that clause unnecessary rather than merely unimplemented: a 1:1 view is
/// *cheaper* than a fit view of the same file, since the dispatch is the same
/// size and the reads are contiguous rather than strided. If that ever inverts,
/// the argument for striking the clause weakens, and this is what would notice.
///
/// # One test and not two, deliberately
///
/// The two cases were two `#[test]` functions until the numbers said otherwise.
/// Cargo runs a binary's tests on a thread each, both of these want the same
/// GPU, and contending for it took the 1:1 case from 2.5 ms to 14.9 ms — a
/// measurement of the test harness that would have flickered either side of the
/// budget forever. A timing assertion has to own the device while it runs, and
/// the only way to say that in a test binary is to be the only test in it.
#[test]
fn a_slider_drag_stays_inside_the_frame_budget() {
let Some(ctx) = ctx() else { return };
let source = synthetic_source(&ctx);
let mut fit = full_point_chain();
drag(&ctx, &source, &mut fit, VIEWPORT).assert_inside_budget("proxy resolution, fit", VIEWPORT);
let mut zoomed = full_point_chain();
zoomed.framing_mut().set_view(one_to_one(VIEWPORT));
drag(&ctx, &source, &mut zoomed, VIEWPORT)
.assert_inside_budget("1:1 on a 24 MP source", VIEWPORT);
}
// ---------------------------------------------------------------------------
// The measurement
// ---------------------------------------------------------------------------
struct Run {
/// Per frame: compose, compose the detail chain, hash the invalidation.
cpu_p99: f64,
/// Per frame: submit and wait for the device to go idle.
gpu_p99: f64,
/// `cpu + gpu` summed within each frame, then ranked. Not the sum of the
/// two percentiles above, which would be a frame that never happened.
total_p99: f64,
}
impl Run {
/// Fail if the budget was missed, saying which half missed it.
///
/// In a dev build only the GPU half is judged — see the module
/// documentation for why — and the CPU figure is still printed, because a
/// reader looking at a failure wants both numbers even when only one of
/// them is the verdict.
fn assert_inside_budget(&self, case: &str, viewport: (u32, u32)) {
let judged = if cfg!(debug_assertions) {
self.gpu_p99
} else {
self.total_p99
};
assert!(
judged <= BUDGET_MS,
"{case} at {}x{}: p99 of {FRAMES} frames was {judged:.2} ms, over the \
{BUDGET_MS:.0} ms budget (cpu {:.2} ms, gpu {:.2} ms, total {:.2} ms). \
FR-DSP-3 is what this violates; docs/dev/frame-budget.md holds the \
numbers it used to be.",
viewport.0,
viewport.1,
self.cpu_p99,
self.gpu_p99,
self.total_p99,
);
}
}
/// Render `FRAMES` frames with the exposure slider moving between each.
///
/// The three calls before the dispatch are the three `DevelopSession::render`
/// makes, in the same order, so this is the develop view's frame rather than an
/// idealisation of it.
fn drag(
ctx: &GpuContext,
source: &DemosaicedImage,
graph: &mut EditGraph,
viewport: (u32, u32),
) -> Run {
// Stands for the `VersionId` the app mixes in. Constant because every frame
// here is the same photograph.
const PHOTOGRAPH: u64 = 0x0dd_ba11;
let mut adjust = AdjustPass::new(ctx);
let src = source.size();
let (w, h) = viewport;
let mut cpu = Vec::with_capacity(FRAMES);
let mut gpu = Vec::with_capacity(FRAMES);
let mut total = Vec::with_capacity(FRAMES);
for i in 0..WARMUP + FRAMES {
// A hundredth of a stop per frame: what a drag does, and what stops
// `render_detailed` reusing the previous frame's colour result and
// turning this into a measurement of nothing.
graph.set_param(exposure::ID, exposure::EXPOSURE, 0.30 + i as f32 * 0.01);
let t0 = Instant::now();
let shader = graph.compose();
let detail = graph.compose_detail(src, viewport);
let colour_key = graph.invalidation().through(Affects::Colour) ^ PHOTOGRAPH;
let cpu_elapsed = t0.elapsed();
let t1 = Instant::now();
adjust
.render_detailed(source, &shader, w, h, None, &detail, colour_key)
.expect("render");
ctx.device
.poll(wgpu::PollType::wait_indefinitely())
.expect("poll");
let gpu_elapsed = t1.elapsed();
if i >= WARMUP {
cpu.push(cpu_elapsed.as_secs_f64() * 1e3);
gpu.push(gpu_elapsed.as_secs_f64() * 1e3);
total.push((cpu_elapsed + gpu_elapsed).as_secs_f64() * 1e3);
}
}
Run {
cpu_p99: p99(cpu),
gpu_p99: p99(gpu),
total_p99: p99(total),
}
}
/// Nearest-rank 99th percentile.
///
/// Nearest-rank because the samples *are* the population: there is no
/// distribution being estimated, only a hundred frames that either fitted in
/// the budget or did not.
fn p99(mut samples: Vec<f64>) -> f64 {
samples.sort_by(f64::total_cmp);
let n = samples.len();
samples[((0.99 * n as f64).ceil() as usize).clamp(1, n) - 1]
}
// ---------------------------------------------------------------------------
// Fixtures
// ---------------------------------------------------------------------------
/// Every operation that contributes a fragment to the fused shader, active.
///
/// Built from [`EditGraph::capabilities`] rather than from a list of operation
/// names, for the same reason the develop panel is: declaring a new node must
/// not silently shrink what this test calls "the full chain". A quarter of the
/// way from each parameter's default towards whichever end is further from it,
/// which is a plausible setting and — the part that matters — is never the
/// neutral, since a neutral operation contributes nothing at all.
///
/// The film stock is not reachable this way (it is a choice of material, not a
/// slider) and is left off. It is one texture lookup and two curve reads; the
/// bench includes it and it is worth about a millisecond at this size.
fn full_point_chain() -> EditGraph {
let mut graph = EditGraph::default_chain();
let moves: Vec<(OpId, ParamId, f32)> = graph
.capabilities()
.iter()
// The neighbourhood operations. See the module documentation for why
// they are not here.
.filter(|op| !op.attributes.contains(&Attribute::Detail))
.flat_map(|op| {
op.params.iter().map(move |p| {
let value = match p.kind {
ParamKind::Scalar { min, max, .. } => {
let far = if (max - p.default).abs() >= (p.default - min).abs() {
max
} else {
min
};
p.default + (far - p.default) * 0.25
}
ParamKind::Bool => 1.0,
// Bound by reference: a descriptor's variants became an
// owned `Vec` when descriptors stopped being `&'static`,
// and this arm only ever reads the length.
ParamKind::Enum { ref variants } => {
if variants.len() > 1 {
1.0
} else {
0.0
}
}
};
(op.id, p.id, value)
})
})
.collect();
for (op, param, value) in moves {
graph.set_param(op, param, value);
}
graph
}
/// The view rect that puts one render pixel on one source pixel, centred.
fn one_to_one(render: (u32, u32)) -> CropRect {
let w = render.0 as f32 / SOURCE.0 as f32;
let h = render.1 as f32 / SOURCE.1 as f32;
CropRect {
x: (1.0 - w) * 0.5,
y: (1.0 - h) * 0.5,
width: w,
height: h,
}
}
/// A source with structure at every scale.
///
/// Not flat: a flat frame lets the memory system serve every sample of every
/// pixel from one cache line, which flatters a bandwidth-bound pass by an amount
/// that has nothing to do with photographs.
fn synthetic_source(ctx: &GpuContext) -> DemosaicedImage {
let (w, h) = SOURCE;
let mut rgba = vec![0u8; (w as usize) * (h as usize) * 4];
for y in 0..h as usize {
let row = y * (w as usize) * 4;
for x in 0..w as usize {
let n = (x.wrapping_mul(2_654_435_761) ^ y.wrapping_mul(1_640_531_527)) >> 13;
let dither = (n & 0x1f) as u32;
let gx = (x * 200 / w as usize) as u32;
let gy = (y * 55 / h as usize) as u32;
let px = &mut rgba[row + x * 4..row + x * 4 + 4];
px[0] = (30 + gx + dither).min(255) as u8;
px[1] = (40 + gy + dither).min(255) as u8;
px[2] = (60 + gx / 2 + gy + dither).min(255) as u8;
px[3] = 255;
}
}
DemosaicedImage::from_rgba8(ctx, &rgba, w, h).expect("upload")
}