Mark what is in focus, so a frame can be judged without zooming to 100%
FR-CULL-3's focus peaking. One compute dispatch measures local contrast in WGSL and writes an overlay texture; on desktop it reaches Slint through the same zero-copy wgpu import the canvas uses, so nothing per-pixel touches the CPU on the frame path. With peaking off the cost is zero and structurally so: focus_overlay opens with `let settings = self.peaking?;` before the frame is touched, and clearing drops both overlay textures, so no VRAM is held either. NFR-P14 is met by construction rather than by measurement -- one dispatch, no second render, no pipeline compile after session open, and a test asserting allocations stay at 2 over eight frames. The budget test asserts 50ms at 4K rather than a tight bound, deliberately: a tight bound fails on a loaded machine and gets deleted, which is worse than a loose one that still catches the regression that matters. TD-1 is amended rather than joined by a TD-6: on Android the overlay rides the readback that already exists there, roughly doubling that transfer while peaking is on, and TD-1's own "Done when" removes both because both are the same missing capability. Verified: cargo fmt clean; clippy --workspace --all-targets -D warnings green, which also compiles peaking.slint through dr-ui's build.rs; 11 focus GPU tests and 79 baseline dr-gpu tests pass; 511 dr-ui tests pass. Not verified: the cfg(target_os = "android") arm, which the host-target clippy never compiled. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,141 @@
|
||||
// TRACES: FR-CULL-3 | NFR-P14
|
||||
// Marking what is sharp, in a layer laid over the frame rather than into it.
|
||||
//
|
||||
// # Why the top octave, and not a gradient
|
||||
//
|
||||
// The obvious detector is a gradient magnitude — Sobel, or a central
|
||||
// difference — and it is the wrong one, for a reason that decides whether the
|
||||
// overlay is useful at all. A gradient answers "is there an edge here", and a
|
||||
// defocused edge is still an edge: blur a 100-code step with a two-pixel
|
||||
// Gaussian and the peak gradient is still around 20 codes per pixel, larger
|
||||
// than a genuinely sharp edge across a low-contrast texture. Peaking built on
|
||||
// gradients lights up the out-of-focus background of every portrait ever
|
||||
// taken, which is the frame it exists to reject.
|
||||
//
|
||||
// What separates sharp from soft is *scale*, not amplitude. Defocus is a
|
||||
// low-pass: it removes the top octave and leaves everything below it intact.
|
||||
// So the detector is a high-pass — this pixel against the mean of its eight
|
||||
// neighbours, a discrete Laplacian — which by construction responds only to
|
||||
// the frequencies defocus destroys.
|
||||
//
|
||||
// The arithmetic, on a one-dimensional step of height D:
|
||||
//
|
||||
// | profile | abs(centre - mean of 8) |
|
||||
// |--------------------------|-------------------------|
|
||||
// | hard step, 1 px | 0.375 D |
|
||||
// | Gaussian blur, sigma 1 | ~0.10 D |
|
||||
// | Gaussian blur, sigma 2 | ~0.03 D |
|
||||
// | linear ramp, any slope | 0 |
|
||||
//
|
||||
// The ramp row is the property being bought: the smooth luminance falloff
|
||||
// across an out-of-focus highlight scores zero however bright it is.
|
||||
//
|
||||
// # Why luma, and why the histogram's luma
|
||||
//
|
||||
// One channel rather than three, because a colour edge carrying no luminance
|
||||
// difference is both rare and, at the acuity an overlay is read at, invisible.
|
||||
// The weights are `histogram.wgsl`'s 54/183/19 over 256 — the same Rec.709
|
||||
// weighting on the same encoded values — so the two instruments in this
|
||||
// application agree about what "luma" means. Two definitions of brightness in
|
||||
// one panel is the kind of disagreement nobody finds until it has already
|
||||
// misled someone.
|
||||
//
|
||||
// # Why the frame is read where it is encoded, and not in linear light
|
||||
//
|
||||
// This runs on the output of the display transform, on encoded values, and
|
||||
// that is deliberate: a fixed difference in sRGB code values is roughly
|
||||
// equally visible wherever it sits in the range, which is what a transfer
|
||||
// curve is for. Measured in linear light the same detector would need a
|
||||
// threshold that varied with exposure, and a shadow texture the photographer
|
||||
// can plainly see would score a hundredth of the identical texture in the
|
||||
// highlights. The encoding has already done the normalisation, so the
|
||||
// threshold is one number.
|
||||
//
|
||||
// # Why this writes a layer and not the picture
|
||||
//
|
||||
// The frame the compositor is handed is also what the histogram counts and
|
||||
// what an export renders (`app.slint`, on the region overlay: a diagnostic
|
||||
// "must not reach the histogram, an export, or the texture the develop pass
|
||||
// hands the compositor"). So the marks go in their own texture — transparent
|
||||
// everywhere except where something is in focus — and the compositor blends
|
||||
// them. Nothing about the photograph changes, and the peaking overlay cannot
|
||||
// leak into a measurement or a file.
|
||||
//
|
||||
// Alpha is written as exactly 0 or exactly 1, never between. The importing
|
||||
// compositor's convention for whether colour arrives premultiplied is not
|
||||
// something this shader can see, and at those two values the two conventions
|
||||
// agree — which is a cheaper guarantee than being right about which one it is.
|
||||
|
||||
struct Params {
|
||||
width: u32,
|
||||
height: u32,
|
||||
// Luma difference at which a pixel is called in focus. See
|
||||
// `PeakSensitivity::threshold` for where the three values come from.
|
||||
threshold: f32,
|
||||
// std140 rounds the scalar block up to 16 bytes before the vec4; named so
|
||||
// the Rust struct's padding is visibly the same shape.
|
||||
pad_0: u32,
|
||||
// The mark's colour, fully saturated. Its alpha is ignored — see above.
|
||||
marker: vec4<f32>,
|
||||
}
|
||||
|
||||
@group(0) @binding(0) var frame: texture_2d<f32>;
|
||||
@group(0) @binding(1) var<uniform> params: Params;
|
||||
@group(0) @binding(2) var marks: texture_storage_2d<rgba8unorm, write>;
|
||||
|
||||
/// Rec.709 luma of an encoded triple, weighted exactly as `histogram.wgsl`
|
||||
/// weights it. 54 + 183 + 19 is 256, so the weights sum to unity.
|
||||
fn luma(c: vec3<f32>) -> f32 {
|
||||
return dot(c, vec3<f32>(54.0, 183.0, 19.0) / 256.0);
|
||||
}
|
||||
|
||||
/// A neighbour, with the frame edge held rather than wrapped.
|
||||
///
|
||||
/// Clamping duplicates the edge pixel into the missing half of the
|
||||
/// neighbourhood, which pulls the mean towards the centre and so biases the
|
||||
/// response *down* on the outermost row and column. That is the right
|
||||
/// direction to be wrong in: the failure is a missing mark at the frame edge,
|
||||
/// where nobody is judging focus, rather than a false mark produced by
|
||||
/// folding the opposite side of the picture into the kernel.
|
||||
fn neighbour(x: i32, y: i32) -> f32 {
|
||||
let cx = clamp(x, 0i, i32(params.width) - 1i);
|
||||
let cy = clamp(y, 0i, i32(params.height) - 1i);
|
||||
return luma(textureLoad(frame, vec2<i32>(cx, cy), 0).rgb);
|
||||
}
|
||||
|
||||
// 8x8, matching the detail stage's dispatch. Each texel is loaded by nine
|
||||
// invocations and no workgroup-memory tile is built to avoid that: at viewport
|
||||
// resolution the reads are perfectly coherent and the texture cache serves
|
||||
// eight of the nine. The budget is NFR-P14's 100 ms against a dispatch
|
||||
// measured in tenths of a millisecond, so there is nothing here worth the
|
||||
// complexity of a tiled load.
|
||||
@compute @workgroup_size(8, 8, 1)
|
||||
fn main(@builtin(global_invocation_id) gid: vec3<u32>) {
|
||||
if (gid.x >= params.width || gid.y >= params.height) {
|
||||
return;
|
||||
}
|
||||
let x = i32(gid.x);
|
||||
let y = i32(gid.y);
|
||||
|
||||
// The eight neighbours, centre excluded. Excluded rather than folded in
|
||||
// because it makes the response readable: `abs(c - mean8)` is the height
|
||||
// of this pixel above its surroundings in the same units as the step it
|
||||
// sits on, so the threshold can be quoted as a luma difference rather than
|
||||
// as eight-ninths of one.
|
||||
var sum = 0.0;
|
||||
for (var dy = -1; dy <= 1; dy = dy + 1) {
|
||||
for (var dx = -1; dx <= 1; dx = dx + 1) {
|
||||
if (dx != 0 || dy != 0) {
|
||||
sum = sum + neighbour(x + dx, y + dy);
|
||||
}
|
||||
}
|
||||
}
|
||||
let centre = luma(textureLoad(frame, vec2<i32>(x, y), 0).rgb);
|
||||
let response = abs(centre - sum / 8.0);
|
||||
|
||||
if (response >= params.threshold) {
|
||||
textureStore(marks, vec2<i32>(x, y), vec4<f32>(params.marker.rgb, 1.0));
|
||||
} else {
|
||||
textureStore(marks, vec2<i32>(x, y), vec4<f32>(0.0, 0.0, 0.0, 0.0));
|
||||
}
|
||||
}
|
||||
Reference in New Issue
Block a user