Merge integration into wip/ingest

Brings in the lens-profile and neighbourhood-operation work so the card
import is verified against what it will actually be merged into, rather
than against the tree it was written on.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-22 15:49:51 +02:00
co-authored by Claude Opus 5
21 changed files with 5601 additions and 163 deletions
Generated
+2
View File
@@ -1423,6 +1423,8 @@ dependencies = [
"env_logger",
"log",
"rawler",
"serde",
"serde_norway",
"thiserror 2.0.20",
"zune-jpeg 0.4.21",
]
+7
View File
@@ -8,6 +8,13 @@ license.workspace = true
[dependencies]
dr-types.workspace = true
rawler.workspace = true
# The camera profile database is data, not code (FR-DEV-3e): a YAML file that
# ships with the binary and is superseded by a newer one on disk. serde_norway
# is the workspace's YAML crate — the fork still receiving releases — and it is
# already in the tree for `dr-pipeline`'s node declarations and `dr-ui`'s style
# tokens. Pure Rust, so it costs nothing under the Android NDK.
serde = { workspace = true }
serde_norway.workspace = true
zune-jpeg.workspace = true
thiserror.workspace = true
log.workspace = true
+160
View File
@@ -0,0 +1,160 @@
# DarkRoom camera base curves (FR-DEV-3e).
#
# ---------------------------------------------------------------------------
# Adding a body is editing this file. It is not a code change.
# ---------------------------------------------------------------------------
#
# The copy you are reading is compiled into the binary as a floor. At startup
# `dr_decode::base_curve::load` also looks for `base_curves.yaml` in:
#
# 1. $DARKROOM_PROFILES/ (set it while you are tuning)
# 2. $XDG_DATA_HOME/darkroom/profiles/
# or $HOME/.local/share/darkroom/profiles/
#
# and uses the first one it finds *whose `version:` is higher than this one's*.
# So: bump `version`, drop the file in that directory, restart. A body added
# this afternoon renders correctly this afternoon, with no release and no
# rebuild — which is what the requirement asks for, and what makes these
# contributable under the GPL.
#
# The version check runs both ways on purpose. A file older than the built-in
# copy is ignored with a log line, so upgrading DarkRoom cannot silently lose
# curves to a pack somebody downloaded a year ago.
#
# ---------------------------------------------------------------------------
# What the numbers mean
# ---------------------------------------------------------------------------
#
# Five `[x, y]` control points on a monotone spline (Fritsch-Carlson, the same
# one the tone curve widget draws). Both axes are **linear**:
#
# x scene-referred camera RGB after white balance, 1.0 = sensor saturation
# y display-referred linear; the sRGB transfer function is applied later,
# at the end of the shader, so do not pre-apply a gamma here
#
# The identity is y = x, and it is what an unrecognised body gets if `default:`
# is removed. It is also the wrong answer for almost every photograph: linear
# scene data has middle grey at about 13% and a camera JPEG puts it near 18%,
# so an uncurved render is roughly half a stop dark through the midtones and
# has no highlight rolloff at all.
#
# A curve that works has three parts, and it is worth naming them because they
# are what you are actually tuning:
#
# the toe the first span, slope near or below 1. Deep shadows stay
# deep. Lift it and blacks go milky; crush it and shadow
# detail the sensor recorded disappears.
# the midtones the middle spans, slope well above 1. This is the contrast
# and the brightness people read as "the camera's look".
# the shoulder the last span, slope well below 1. Highlights compress
# toward white instead of arriving there and clipping. It is
# the difference between a rolled-off sky and a white hole.
#
# Two invariants are enforced in code and tested, so a mistake here fails the
# build rather than the photograph: x must strictly increase, y must not
# decrease, and everything must lie inside the unit square.
#
# ---------------------------------------------------------------------------
# Honesty about these values
# ---------------------------------------------------------------------------
#
# These are hand-tuned shapes, not measurements. They encode what every camera
# JPEG rendering has in common — the toe/midtone/shoulder structure above —
# plus each maker's well-known house differences: Canon's gentler shoulder and
# warmer-reading midtones, Nikon's slightly higher midtone contrast, Sony's
# flatter and more conservative default, Fujifilm's markedly contrastier
# Provia-derived rendering.
#
# FR-DEV-3e's acceptance criterion is subjective comparison against each body's
# own JPEG, and meeting it properly needs a frame from that body in front of
# you. Where that has not been done, the entry is still much closer to right
# than the identity — which is the bar these have to clear, and do.
version: 1
# The rendering for a body with no entry of its own.
#
# **Deliberately not the identity.** The failure this requirement exists to fix
# is the flat render, and a conservative curve is far closer to right for every
# body than no curve is for any of them. It is gentler than the per-body
# entries below — a shallower midtone and an earlier, softer shoulder — because
# it has to be safe on a sensor nobody has looked at, and the cost of being too
# tame is a photograph that wants a little contrast rather than one that has
# lost its highlights.
default:
points:
- [0.00, 0.000]
- [0.04, 0.043]
- [0.13, 0.175]
- [0.45, 0.690]
- [1.00, 1.000]
bodies:
# Canon. A soft toe and a long, gradual shoulder — the reason Canon files
# are described as forgiving in highlights and a little low in contrast
# straight out of camera.
- make: Canon
model: EOS 6D
points:
- [0.00, 0.000]
- [0.04, 0.045]
- [0.13, 0.190]
- [0.45, 0.720]
- [1.00, 1.000]
- make: Canon
model: EOS R6
points:
- [0.00, 0.000]
- [0.04, 0.044]
- [0.13, 0.195]
- [0.45, 0.730]
- [1.00, 1.000]
# Nikon. A slightly deeper toe and more midtone slope than Canon, which is
# the "punchier out of camera" difference people describe between the two.
- make: Nikon
model: Z 6
points:
- [0.00, 0.000]
- [0.04, 0.038]
- [0.13, 0.200]
- [0.46, 0.750]
- [1.00, 1.000]
- make: Nikon
model: D750
points:
- [0.00, 0.000]
- [0.04, 0.039]
- [0.13, 0.198]
- [0.46, 0.745]
- [1.00, 1.000]
# Sony. The flattest default of the four, and intentionally so — Sony's own
# rendering leaves more headroom than it uses, which is why Sony files are
# the ones people describe as needing the most work.
- make: Sony
model: ILCE-7M3
points:
- [0.00, 0.000]
- [0.04, 0.048]
- [0.13, 0.185]
- [0.44, 0.700]
- [1.00, 1.000]
# Fujifilm. Provia, the default film simulation: a firm toe, the steepest
# midtones here, and a hard shoulder. It is the most distinctive rendering of
# the four and the one where a flat render looks most obviously wrong.
#
# This entry does *not* read the in-RAF film simulation tag — that is
# FR-DEV-3f, and until it lands every Fujifilm file gets the Provia shape
# whatever the camera was set to.
- make: Fujifilm
model: X-T3
points:
- [0.00, 0.000]
- [0.045, 0.040]
- [0.14, 0.215]
- [0.47, 0.775]
- [1.00, 1.000]
+763
View File
@@ -0,0 +1,763 @@
//! TRACES: FR-DEV-3e
//! Base curves — the per-body rendering that turns a correct exposure into a
//! photograph.
//!
//! # What this is for
//!
//! A camera matrix gets the *colours* right and leaves the picture flat. Sensor
//! data is scene-referred and very nearly linear; a print, a screen and a
//! camera's own JPEG are none of those things. Rendering linear data straight
//! out is the dcraw default, and FR-DEV-3e names it precisely: "the flat,
//! poor-skin-tone rendering characteristic of dcraw defaults, which is the
//! documented reason people abandon darktable in the first hour."
//!
//! The fix is a tone curve applied as part of *reading* the file rather than as
//! an edit — a toe, a steep midtone, and a shoulder that rolls highlights off
//! instead of clipping them. Every raw converter has one. Adobe calls it the
//! camera profile's tone curve, darktable calls it the base curve, and the name
//! here follows darktable's because the placement does too: it runs in camera
//! RGB, after white balance and the user's adjustments, immediately before the
//! conversion out to a working space.
//!
//! # Why it is not an edit
//!
//! It never reaches the sidecar and there is no slider for it, for the same
//! reason the EXIF orientation is not an edit (FR-DEV-3h): it is a property of
//! the body that took the frame, not of what anyone decided about the frame.
//! Sidecars are shared between devices and bodies (FR-NC-9), and one camera's
//! rendering must not follow an edit onto another camera's file.
//!
//! # Why it is data
//!
//! FR-DEV-3e requires the profile database to be "versioned independently of
//! the app binary so bodies and curves can be added without a release — and,
//! under D8's GPLv3, contributed by users". So the curves live in
//! `profiles/base_curves.yaml`, a file that is compiled in as a floor and
//! *overridden* by a copy on disk carrying a higher `version:`. Adding a body
//! is adding ten numbers to a YAML file; shipping that body to users is
//! publishing the file. Neither is a code change and neither needs a release.
//!
//! See [`load`] for the search path and [`Curves::body`] for the matching.
use std::path::{Path, PathBuf};
use std::sync::OnceLock;
/// How many control points a base curve has.
///
/// Five, which is not a coincidence: it is what the tone curve widget uses
/// (`dr_pipeline::ops::curve::POINTS`), so the shader evaluates a profile's
/// curve and a photographer's curve through exactly the same spline. A profile
/// author and a photographer dragging a point mean the same thing by it, and
/// the generated shader carries one implementation rather than two that could
/// disagree.
pub const POINTS: usize = 5;
/// TRACES: FR-DEV-3e
/// A base curve: five points on a monotone spline through the unit square.
///
/// `xs` is scene-linear camera RGB, normalised so that 1.0 is the sensor's
/// saturation point. `ys` is display-referred linear — *not* gamma-encoded,
/// because the sRGB transfer function is applied at the very end of the
/// generated shader and applying it twice would wash the image out.
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct BaseCurve {
pub xs: [f32; POINTS],
pub ys: [f32; POINTS],
}
impl BaseCurve {
/// The curve that does nothing — the identity diagonal.
///
/// What an unrecognised body gets if the database carries no default, and
/// what a JPEG gets always: an already-rendered image must not be rendered
/// a second time.
pub const IDENTITY: Self = Self {
xs: [0.0, 0.25, 0.5, 0.75, 1.0],
ys: [0.0, 0.25, 0.5, 0.75, 1.0],
};
/// Whether this curve would leave the image alone.
///
/// The shader is told to skip the stage entirely when it would, so an
/// unprofiled body costs a branch that is uniform across the dispatch
/// rather than a spline evaluation per channel per pixel.
pub fn is_identity(&self) -> bool {
self.xs
.iter()
.zip(self.ys.iter())
.all(|(x, y)| (x - y).abs() < 1e-6)
}
/// Build from raw pairs, rejecting anything that is not a curve.
///
/// A profile file is data a user may have edited, so this is the boundary
/// where "ten numbers" becomes "a curve": the x coordinates must increase,
/// the y coordinates must not decrease, and both must lie in the unit
/// square. A non-monotone x sends the spline's span search backwards and
/// divides by a negative width; a decreasing y inverts tones locally,
/// which reads as a dark halo through smooth gradients rather than as a
/// bad profile.
///
/// Endpoints are not forced to (0,0) and (1,1). A curve that lifts black
/// slightly, or that places the shoulder below white, is a legitimate
/// rendering choice and several bodies make it.
pub fn from_points(points: &[[f32; 2]]) -> Option<Self> {
if points.len() != POINTS {
return None;
}
let mut xs = [0.0f32; POINTS];
let mut ys = [0.0f32; POINTS];
for (i, p) in points.iter().enumerate() {
if !p[0].is_finite() || !p[1].is_finite() {
return None;
}
if !(0.0..=1.0).contains(&p[0]) || !(0.0..=1.0).contains(&p[1]) {
return None;
}
xs[i] = p[0];
ys[i] = p[1];
}
for i in 1..POINTS {
// Strictly increasing in x — the spline divides by the span width.
if xs[i] <= xs[i - 1] {
return None;
}
// Non-decreasing in y. Flat is allowed: a curve that holds a
// highlight range at white is clipping deliberately.
if ys[i] < ys[i - 1] {
return None;
}
}
Some(Self { xs, ys })
}
}
/// One body's entry in the database.
#[derive(Debug, Clone, PartialEq)]
pub struct BodyCurve {
/// The manufacturer, as the file writes it — "Canon", "NIKON CORPORATION".
pub make: String,
/// The model, as the file writes it — "EOS 6D", "ILCE-7M3".
pub model: String,
pub curve: BaseCurve,
}
/// TRACES: FR-DEV-3e
/// The base curve database.
///
/// Versioned as a whole rather than per body, because that is the unit a user
/// downloads and the unit that has to beat the built-in copy. See [`load`].
#[derive(Debug, Clone, PartialEq)]
pub struct Curves {
version: u32,
default: Option<BaseCurve>,
bodies: Vec<BodyCurve>,
}
impl Curves {
/// TRACES: FR-DEV-3e
/// The curve to render a frame from this body with.
///
/// Falls back, in order, to the database's `default:` and then to the
/// identity. **The default is deliberately not the identity**: an
/// unrecognised body rendered flat is the failure this requirement exists
/// to prevent, and a gentle, conservative curve is much closer to right for
/// every body than no curve is for any of them. A body with its own entry
/// gets that instead.
///
/// # What "this body" has to survive
///
/// The same camera names itself three ways depending on which program last
/// touched the file. A native NEF says make "NIKON CORPORATION", model
/// "NIKON Z 6"; rawler's own database cleans that to "Nikon" and "Z 6"; an
/// Adobe-converted DNG keeps the uncleaned pair. A database that had to
/// spell every variant would go stale the first time a maker changed its
/// mind about its own name, so the matching does the folding instead:
///
/// - Case, punctuation and runs of whitespace are flattened, so
/// "ILCE-7M3", "ILCE 7M3" and "ilce-7m3" are one body.
/// - The make is compared on its **first word only**. Every maker's
/// trailing corporate boilerplate — "CORPORATION", "IMAGING CORP" — is
/// noise, and no two camera manufacturers share a first word.
/// - The model is tried both as written and with a leading copy of the
/// make removed, which is what lets one "Canon"/"EOS 6D" entry cover
/// "Canon EOS 6D" as well.
pub fn body(&self, make: &str, model: &str) -> BaseCurve {
let (make, model) = (make_key(make), normalise(model));
// The model with a leading copy of the maker's name removed.
let bare = model.strip_prefix(&format!("{make} ")).unwrap_or(&model);
self.bodies
.iter()
.find(|b| {
let entry_model = normalise(&b.model);
make_key(&b.make) == make && (entry_model == model || entry_model == bare)
})
.map(|b| b.curve)
.or(self.default)
.unwrap_or(BaseCurve::IDENTITY)
}
/// The database version. Higher wins; see [`load`].
pub fn version(&self) -> u32 {
self.version
}
/// How many bodies have their own curve, excluding the default.
pub fn len(&self) -> usize {
self.bodies.len()
}
pub fn is_empty(&self) -> bool {
self.bodies.is_empty()
}
/// Parse a database from YAML.
///
/// Entries that are not curves are dropped with a warning rather than
/// failing the parse. A user-contributed file with one bad body should
/// cost that body's rendering, not every body's — and the alternative is an
/// application that will not open a photograph because somebody typed a
/// comma.
pub fn parse(yaml: &str) -> Result<Self, String> {
let file: File = serde_norway::from_str(yaml).map_err(|e| e.to_string())?;
let default = file.default.and_then(|d| {
BaseCurve::from_points(&d.points).or_else(|| {
log::warn!("base curves: the default entry is not a monotone curve; ignoring it");
None
})
});
let bodies = file
.bodies
.into_iter()
.filter_map(|b| match BaseCurve::from_points(&b.points) {
Some(curve) => Some(BodyCurve {
make: b.make,
model: b.model,
curve,
}),
None => {
log::warn!(
"base curves: {} {} is not a monotone curve; ignoring it",
b.make,
b.model
);
None
}
})
.collect();
Ok(Self {
version: file.version,
default,
bodies,
})
}
}
/// The copy that ships inside the binary.
///
/// A floor, not the answer: [`load`] prefers a newer file on disk. Compiled in
/// so that a fresh install with no profile directory — and every Android build,
/// where there is no such directory to speak of — still renders properly.
const BUILT_IN: &str = include_str!("../profiles/base_curves.yaml");
/// TRACES: FR-DEV-3e
/// The base curve database, loaded once.
///
/// # The search path, and why it is a version comparison
///
/// 1. `$DARKROOM_PROFILES`, a directory, when set. The escape hatch: a profile
/// author iterating on a curve points this at their working copy and does
/// not have to install anything.
/// 2. `$XDG_DATA_HOME/darkroom/profiles/`, else `$HOME/.local/share/darkroom/profiles/`.
/// The same base directory the catalog uses, chosen there for the same
/// reason — it is data, not cache, and must survive a storage sweep.
/// 3. The copy compiled into the binary.
///
/// The first file that parses *and carries a higher `version:` than the
/// built-in copy* wins. The version check is the whole mechanism the
/// requirement asks for, and it runs in both directions:
///
/// - A downloaded pack at version 7 supersedes a binary shipping version 3, so
/// a body added after the release renders correctly with no release.
/// - A stale pack at version 2 does **not** supersede a binary shipping version
/// 3, so upgrading the application cannot silently lose curves to a file
/// somebody downloaded a year ago and forgot.
///
/// Failures are warnings, never errors. A malformed profile file must cost the
/// user their curves, not their photographs.
pub fn load() -> &'static Curves {
static LOADED: OnceLock<Curves> = OnceLock::new();
LOADED.get_or_init(|| {
let built_in = Curves::parse(BUILT_IN).unwrap_or_else(|e| {
// Unreachable in a build that ran its tests — `the_shipped_database_parses`
// asserts exactly this — but a panic here would mean an
// application that cannot open a photograph because of a typo in a
// data file, which is never the right trade.
log::error!("base curves: the built-in database does not parse: {e}");
Curves {
version: 0,
default: None,
bodies: Vec::new(),
}
});
choose(built_in, &search_path())
})
}
/// The version comparison, separated from where the directories come from.
///
/// Split out so it can be tested against real files in a real directory
/// without the process-wide `OnceLock` and the environment `load` reads. The
/// rule this implements is the whole of what FR-DEV-3e asks for, so it is
/// worth being able to state it as a test rather than as a comment.
fn choose(built_in: Curves, dirs: &[PathBuf]) -> Curves {
for dir in dirs {
let path = dir.join("base_curves.yaml");
let Ok(text) = std::fs::read_to_string(&path) else {
continue;
};
match Curves::parse(&text) {
Ok(external) if external.version > built_in.version => {
log::info!(
"base curves: using {} (version {}, {} bodies) over the built-in version {}",
path.display(),
external.version,
external.len(),
built_in.version
);
return external;
}
Ok(external) => log::info!(
"base curves: ignoring {} at version {}; the built-in database is version {}",
path.display(),
external.version,
built_in.version
),
Err(e) => log::warn!("base curves: {} does not parse: {e}", path.display()),
}
}
built_in
}
/// TRACES: FR-DEV-3e
/// The curve for a body, from the loaded database.
///
/// The one call site the decoder needs; everything above is reachable for
/// tests and for a future profile editor.
pub fn for_body(make: &str, model: &str) -> BaseCurve {
load().body(make, model)
}
/// Directories that may hold a `base_curves.yaml`, most specific first.
fn search_path() -> Vec<PathBuf> {
let mut dirs = Vec::new();
if let Some(explicit) = std::env::var_os("DARKROOM_PROFILES") {
dirs.push(PathBuf::from(explicit));
}
// The same resolution `dr_ui::library::catalog_path` uses, and for the
// same reason: this is data a user may have installed, not a cache. It is
// duplicated rather than shared because `dr-decode` sits far below the UI
// and must not acquire a dependency on it to find a directory.
let base = std::env::var_os("XDG_DATA_HOME")
.map(PathBuf::from)
.or_else(|| std::env::var_os("HOME").map(|h| Path::new(&h).join(".local/share")));
if let Some(base) = base {
dirs.push(base.join("darkroom").join("profiles"));
}
dirs
}
/// A manufacturer's first word, folded.
///
/// "NIKON CORPORATION", "Nikon" and "nikon" all become `NIKON`. The corporate
/// suffixes are not information — they appear or not depending on whether the
/// file went through a DNG converter — and no two camera manufacturers share a
/// first word, so nothing is lost by dropping them.
fn make_key(s: &str) -> String {
normalise(s)
.split(' ')
.next()
.unwrap_or_default()
.to_string()
}
/// Fold a make or model into something two files can agree on.
///
/// Upper-cased, with every run of non-alphanumeric characters collapsed to one
/// space and the ends trimmed, so that "ILCE-7M3", "ILCE 7M3" and "ilce-7m3"
/// become one.
fn normalise(s: &str) -> String {
let mut out = String::with_capacity(s.len());
let mut pending_space = false;
for c in s.chars() {
if c.is_ascii_alphanumeric() {
if pending_space && !out.is_empty() {
out.push(' ');
}
pending_space = false;
out.push(c.to_ascii_uppercase());
} else {
pending_space = true;
}
}
out
}
// ---- The on-disk shape, kept apart from the in-memory one ----------------
//
// Deliberately separate types. The file is data a user edits and is allowed to
// be wrong; `Curves` is a parsed database whose every entry is known to be a
// monotone curve. Deriving `Deserialize` on `BaseCurve` directly would delete
// that boundary and let an unchecked five-point array reach the shader.
//
// Unknown fields are **accepted**, which is not laziness. The database is
// versioned independently of the binary and moves in both directions: a pack
// published after this release may carry keys this build has never heard of —
// a hue twist, a look table (FR-DEV-3f) — and it must still deliver its curves
// to an older DarkRoom rather than failing to parse and leaving every body
// flat. `deny_unknown_fields` would trade that for a diagnostic nobody needs.
#[derive(serde::Deserialize)]
struct File {
version: u32,
#[serde(default)]
default: Option<Entry>,
#[serde(default)]
bodies: Vec<BodyEntry>,
}
#[derive(serde::Deserialize)]
struct Entry {
points: Vec<[f32; 2]>,
}
#[derive(serde::Deserialize)]
struct BodyEntry {
make: String,
model: String,
points: Vec<[f32; 2]>,
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn the_shipped_database_parses_and_carries_a_default() {
// The one test that must never be allowed to fail quietly: `load`
// degrades to an empty database rather than panicking, so without this
// a typo in the YAML would ship as "every photograph renders flat"
// rather than as a build failure.
let curves = Curves::parse(BUILT_IN).expect("the shipped database parses");
assert!(curves.version() >= 1);
assert!(!curves.is_empty(), "the database ships bodies");
assert!(
!curves.body("Nobody", "Nothing").is_identity(),
"an unknown body must still get the default rendering"
);
}
#[test]
fn every_shipped_curve_lifts_the_midtones_and_rolls_the_highlights() {
// What makes a base curve a base curve rather than a decoration. If a
// shipped curve failed either half it would be a worse rendering than
// the flat one it replaced, which is the one outcome forbidden.
let curves = Curves::parse(BUILT_IN).expect("parses");
let all = curves
.bodies
.iter()
.map(|b| (format!("{} {}", b.make, b.model), b.curve))
.chain(curves.default.map(|c| ("default".to_string(), c)));
for (name, curve) in all {
// The midtone point sits above the diagonal: a linear midtone is
// roughly a stop and a half darker than any camera renders it.
let mid = 2;
assert!(
curve.ys[mid] > curve.xs[mid],
"{name} does not lift its midtones ({} -> {})",
curve.xs[mid],
curve.ys[mid]
);
// And the last span is shallower than the one before it, which is
// what a shoulder *is*. Without one the curve clips highlights
// harder than the linear rendering did.
let slope = |i: usize| (curve.ys[i + 1] - curve.ys[i]) / (curve.xs[i + 1] - curve.xs[i]);
assert!(
slope(POINTS - 2) < slope(POINTS - 3),
"{name} has no highlight shoulder"
);
}
}
#[test]
fn a_curve_that_is_not_monotone_is_refused() {
// The profile file is user-editable, so this is a real boundary and
// not a formality. A decreasing y inverts tones locally and shows up
// as a dark halo in a gradient, which reads as a rendering fault
// rather than as a bad profile.
assert_eq!(
BaseCurve::from_points(&[
[0.0, 0.0],
[0.25, 0.4],
[0.5, 0.3],
[0.75, 0.8],
[1.0, 1.0]
]),
None
);
}
#[test]
fn a_curve_whose_x_does_not_advance_is_refused() {
// The spline divides by the span width; a repeated x is a division by
// zero in the shader, which is a NaN pixel rather than an error.
assert_eq!(
BaseCurve::from_points(&[
[0.0, 0.0],
[0.25, 0.3],
[0.25, 0.5],
[0.75, 0.8],
[1.0, 1.0]
]),
None
);
}
#[test]
fn a_curve_of_the_wrong_length_is_refused() {
assert_eq!(BaseCurve::from_points(&[[0.0, 0.0], [1.0, 1.0]]), None);
}
#[test]
fn values_outside_the_unit_square_are_refused() {
// The shader clamps its output at the very end anyway, but a control
// point above 1.0 would put the shoulder outside the range the curve
// is defined over and silently flatten everything below it.
assert_eq!(
BaseCurve::from_points(&[
[0.0, 0.0],
[0.25, 0.3],
[0.5, 1.4],
[0.75, 1.5],
[1.0, 1.6]
]),
None
);
}
#[test]
fn a_body_with_its_own_entry_beats_the_default() {
let curves = Curves::parse(
"version: 2
default:
points: [[0.0, 0.0], [0.25, 0.3], [0.5, 0.6], [0.75, 0.85], [1.0, 1.0]]
bodies:
- make: Canon
model: EOS 6D
points: [[0.0, 0.0], [0.25, 0.35], [0.5, 0.7], [0.75, 0.9], [1.0, 1.0]]
",
)
.expect("parses");
assert_eq!(curves.body("Canon", "EOS 6D").ys[1], 0.35);
assert_eq!(curves.body("Canon", "EOS 5D").ys[1], 0.30);
}
#[test]
fn the_make_may_be_repeated_in_the_model() {
// Canon writes "Canon" as the make and "Canon EOS 6D" as the model;
// rawler's cleaned strings drop the repetition and both reach here.
// One entry has to cover both or half the files on a card miss.
let curves = Curves::parse(
"version: 1
bodies:
- make: Canon
model: EOS 6D
points: [[0.0, 0.0], [0.25, 0.35], [0.5, 0.7], [0.75, 0.9], [1.0, 1.0]]
",
)
.expect("parses");
assert_eq!(curves.body("Canon", "Canon EOS 6D").ys[1], 0.35);
assert_eq!(curves.body("Canon", "EOS 6D").ys[1], 0.35);
assert_eq!(curves.body("CANON", "eos 6d").ys[1], 0.35);
}
#[test]
fn a_corporate_suffix_does_not_hide_a_body() {
// The same Z 6 arrives as "Nikon"/"Z 6" from rawler's camera database
// and as "NIKON CORPORATION"/"NIKON Z 6" from a DNG converted out of
// the same file. Both must find the entry, or converting a file to
// DNG would silently change how it renders.
let curves = Curves::parse(
"version: 1
bodies:
- make: Nikon
model: Z 6
points: [[0.0, 0.0], [0.25, 0.35], [0.5, 0.7], [0.75, 0.9], [1.0, 1.0]]
",
)
.expect("parses");
assert_eq!(curves.body("Nikon", "Z 6").ys[1], 0.35);
assert_eq!(curves.body("NIKON CORPORATION", "NIKON Z 6").ys[1], 0.35);
}
#[test]
fn punctuation_and_spacing_do_not_decide_whether_a_body_is_known() {
let curves = Curves::parse(
"version: 1
bodies:
- make: Sony
model: ILCE-7M3
points: [[0.0, 0.0], [0.25, 0.35], [0.5, 0.7], [0.75, 0.9], [1.0, 1.0]]
",
)
.expect("parses");
assert_eq!(curves.body("SONY", "ILCE 7M3").ys[1], 0.35);
assert_eq!(curves.body("sony", "ilce-7m3").ys[1], 0.35);
}
#[test]
fn one_bad_entry_does_not_cost_the_rest() {
// A user-contributed file with one typo should cost that body's
// rendering, not every body's.
let curves = Curves::parse(
"version: 1
bodies:
- make: Broken
model: Body
points: [[0.0, 0.0], [0.25, 0.9], [0.5, 0.1], [0.75, 0.9], [1.0, 1.0]]
- make: Canon
model: EOS 6D
points: [[0.0, 0.0], [0.25, 0.35], [0.5, 0.7], [0.75, 0.9], [1.0, 1.0]]
",
)
.expect("parses");
assert_eq!(curves.len(), 1);
assert_eq!(curves.body("Canon", "EOS 6D").ys[1], 0.35);
assert!(curves.body("Broken", "Body").is_identity());
}
#[test]
fn a_pack_from_the_future_still_delivers_its_curves() {
// The database is versioned independently of the binary, so a pack
// published after this build may carry keys this build has never heard
// of. It must still hand over the curves it does understand — failing
// the parse would leave every body flat, which is the exact failure
// FR-DEV-3e exists to prevent, delivered by the mechanism meant to
// prevent it.
let curves = Curves::parse(
"version: 9
look_table: ambitious
bodies:
- make: Canon
model: EOS 6D
hue_twist: [1, 2, 3]
points: [[0.0, 0.0], [0.25, 0.35], [0.5, 0.7], [0.75, 0.9], [1.0, 1.0]]
",
)
.expect("an unfamiliar key must not fail the parse");
assert_eq!(curves.version(), 9);
assert_eq!(curves.body("Canon", "EOS 6D").ys[1], 0.35);
}
#[test]
fn an_unknown_body_with_no_default_gets_the_identity() {
// Graceful fallback, stated as a property: never worse than a flat
// render, and never a curve tuned for somebody else's sensor when the
// database declines to offer one.
let curves = Curves::parse("version: 1\nbodies: []\n").expect("parses");
assert!(curves.body("Nobody", "Nothing").is_identity());
}
/// A directory holding one `base_curves.yaml`, unique to the caller.
fn a_pack_dir(name: &str, yaml: &str) -> PathBuf {
let dir = std::env::temp_dir().join(format!("darkroom-base-curves-{name}"));
let _ = std::fs::remove_dir_all(&dir);
std::fs::create_dir_all(&dir).expect("a writable temp directory");
std::fs::write(dir.join("base_curves.yaml"), yaml).expect("write");
dir
}
const A_CANON_ENTRY: &str = "bodies:
- make: Canon
model: EOS 6D
points: [[0.0, 0.0], [0.25, 0.42], [0.5, 0.7], [0.75, 0.9], [1.0, 1.0]]
";
#[test]
fn a_newer_pack_on_disk_supersedes_the_built_in_database() {
// **This is the requirement.** FR-DEV-3e asks for a profile database
// versioned independently of the app binary "so bodies and curves can
// be added without a release". A file with a higher version, dropped
// in the profile directory, is what that means in practice.
let built_in = Curves::parse(BUILT_IN).expect("parses");
let newer = format!("version: {}\n{A_CANON_ENTRY}", built_in.version() + 1);
let dir = a_pack_dir("newer", &newer);
let chosen = choose(built_in.clone(), &[dir]);
assert_eq!(chosen.version(), built_in.version() + 1);
assert_eq!(chosen.body("Canon", "EOS 6D").ys[1], 0.42);
}
#[test]
fn a_stale_pack_does_not_survive_an_upgrade() {
// The other direction, and the one that protects the user. Somebody
// downloads a pack, a release later ships better curves for the same
// bodies, and the forgotten file must not quietly hold the application
// back at last year's rendering.
let built_in = Curves::parse(BUILT_IN).expect("parses");
let stale = format!("version: {}\n{A_CANON_ENTRY}", built_in.version());
let dir = a_pack_dir("stale", &stale);
let chosen = choose(built_in.clone(), &[dir]);
assert_eq!(chosen.version(), built_in.version());
assert_ne!(
chosen.body("Canon", "EOS 6D").ys[1],
0.42,
"an equal version must not displace the built-in database"
);
}
#[test]
fn a_broken_pack_costs_the_curves_and_not_the_photographs() {
// A malformed profile file must degrade to the built-in database, not
// to an error. The user came here to look at a photograph.
let built_in = Curves::parse(BUILT_IN).expect("parses");
let dir = a_pack_dir("broken", "version: [this is not a number\n");
let chosen = choose(built_in.clone(), &[dir]);
assert_eq!(chosen.version(), built_in.version());
assert_eq!(chosen.len(), built_in.len());
}
#[test]
fn a_directory_with_no_pack_in_it_is_simply_skipped() {
// The ordinary case on every machine: the search path exists, the file
// does not. It must not be a warning, an error, or a slow path.
let built_in = Curves::parse(BUILT_IN).expect("parses");
let missing = std::env::temp_dir().join("darkroom-base-curves-nothing-here");
let _ = std::fs::remove_dir_all(&missing);
assert_eq!(choose(built_in.clone(), &[missing]), built_in);
}
#[test]
fn the_identity_is_recognised_as_doing_nothing() {
assert!(BaseCurve::IDENTITY.is_identity());
assert!(!Curves::parse(BUILT_IN)
.expect("parses")
.body("Canon", "EOS 6D")
.is_identity());
}
}
+52 -66
View File
@@ -12,10 +12,13 @@
//! Fusing them would force a full decode where a header read suffices, which
//! is exactly why Lightroom stalls ~2 s per image during culling.
pub mod base_curve;
mod error;
mod locate;
mod preview;
pub mod profile;
pub use base_curve::BaseCurve;
pub use error::DecodeError;
pub use locate::{
defects, is_complete_jpeg, locate_preview, BadLine, BadPixel, Defects, PreviewLocation,
@@ -25,6 +28,7 @@ pub use preview::{
decode_jpeg, extract_embedded_preview, extract_preview, Preview, PreviewSize,
PREVIEW_PROBE_BYTES,
};
pub use profile::CameraProfile;
use dr_types::{Format, Orientation};
@@ -89,7 +93,24 @@ pub struct RawImage {
/// `None` where the body is unknown to the decoder, in which case the
/// pipeline falls back to identity and the result is uncalibrated rather
/// than wrong-by-a-guess.
///
/// Where the body carries two calibrations this is already *interpolated*
/// for the light the frame was shot under; see [`profile::CameraProfile`].
pub color_matrix: Option<[f32; 9]>,
/// TRACES: FR-DEV-3e
/// The per-body rendering curve, the other half of the camera profile.
///
/// The matrix above decides what the colours *are*; this decides what the
/// picture looks like. Carried on the decoded image rather than looked up
/// downstream because this is the only point in the system that knows
/// which body took the frame, and because it is not an edit: it belongs to
/// the file in the same way the masked-photosite crop does, and must never
/// reach a sidecar (FR-NC-9).
///
/// [`BaseCurve::IDENTITY`] for an unknown body with no default in the
/// database, which renders exactly as this decoder did before profiles
/// existed.
pub base_curve: BaseCurve,
/// The usable region of `data`, excluding masked and border photosites.
pub crop: CropRect,
}
@@ -421,15 +442,38 @@ pub fn decode(bytes: &[u8]) -> Result<RawImage, DecodeError> {
let source = RawSource::new_from_slice(bytes);
let decoder =
rawler::get_decoder(&source).map_err(|e| DecodeError::Unsupported(e.to_string()))?;
// Read before decoding, while the decoder is still the cheapest thing in
// the room. These are the DNG tags rawler parses into its IFD and then
// never surfaces — `ForwardMatrix1/2` above all — and they are empty for
// every non-DNG file, which is not a failure (FR-DEV-3e).
let dng = profile::read_dng_matrices(decoder.as_ref());
let image = decoder
.raw_image(&source, &Default::default(), false)
.map_err(|e| DecodeError::Decode(e.to_string()))?;
// Both derived before the match below moves `image.data`, and from the
// same matrix: the balance and the conversion must agree about which white
// is neutral or the frame carries a cast that looks like a decode fault.
let color_matrix = cam_to_srgb(&image);
let wb_coeffs = sane_wb(image.wb_coeffs, xyz_to_cam_of(&image).as_ref());
// The camera profile, and both of the things derived from it, are built
// before the match below moves `image.data`.
//
// They come from *one* profile deliberately: the balance and the
// conversion must agree about which white is neutral, or the frame carries
// a cast that looks like a decode fault. That agreement used to be
// maintained by hand — two functions reading the same illuminant key — and
// is now structural, because there is only one interpolated matrix and
// both callers ask the same object for it.
let profile = profile::CameraProfile::extract(&image, &dng);
let color_matrix = profile.as_ref().and_then(|p| p.cam_to_srgb());
let wb_coeffs = sane_wb(image.wb_coeffs, profile.as_ref().map(|p| p.xyz_to_cam()).as_ref());
// The rendering half of the profile (FR-DEV-3e). rawler's cleaned strings
// are preferred where it has them — they are what the shipped database is
// written against — and the matching folds the variants either way, so a
// DNG naming the same body differently still finds its curve.
let base_curve = base_curve::for_body(
image.camera.clean_make.as_str(),
image.camera.clean_model.as_str(),
);
let data = match image.data {
rawler::RawImageData::Integer(v) => v,
@@ -498,68 +542,10 @@ pub fn decode(bytes: &[u8]) -> Result<RawImage, DecodeError> {
.unwrap_or(u16::MAX),
wb_coeffs,
color_matrix,
base_curve,
})
}
/// TRACES: FR-DEV-3e
/// Compose the camera→sRGB-linear matrix from rawler's XYZ→camera.
///
/// **Two rawler traps this avoids**, both measured on a Canon 6D CR2
/// (2026-08-09):
///
/// 1. `RawImage::xyz_to_cam` is **all zeros** — it carries an upstream
/// deprecation note and 0.7.2 no longer fills it. The live data is
/// `color_matrix`, keyed by illuminant. Reading the old field silently
/// yields no colour transform at all.
/// 2. `cam_to_xyz_normalized()` divides each of four rows by its own sum, and
/// the fourth row (emerald/white, unused on any Bayer body) sums to zero.
/// Every element came back `NaN`. Inverting the 3×3 ourselves avoids the
/// fourth channel entirely.
fn cam_to_srgb(image: &rawler::RawImage) -> Option<[f32; 9]> {
use rawler::imgop::xyz::Illuminant;
// Prefer D65 — it matches sRGB's white point, so no chromatic adaptation
// is needed. Illuminant A (tungsten) is a distant fallback for bodies
// that ship only one matrix; adapting it properly is a v0.2 colour-
// management concern (ARCH §5.2), not something to fake here.
let flat = image
.color_matrix
.get(&Illuminant::D65)
.or_else(|| image.color_matrix.get(&Illuminant::A))?;
if flat.len() < 9 {
return None;
}
let xyz_to_cam: [[f32; 3]; 3] = [
[flat[0], flat[1], flat[2]],
[flat[3], flat[4], flat[5]],
[flat[6], flat[7], flat[8]],
];
cam_to_srgb_from(&xyz_to_cam)
}
/// The camera's XYZ→camera matrix, as rawler holds it.
///
/// Split out so the white-balance fallback and the colour matrix read the same
/// data through the same illuminant preference; two readers disagreeing about
/// which matrix a body uses would balance to one white and convert from
/// another.
fn xyz_to_cam_of(image: &rawler::RawImage) -> Option<[[f32; 3]; 3]> {
use rawler::imgop::xyz::Illuminant;
let flat = image
.color_matrix
.get(&Illuminant::D65)
.or_else(|| image.color_matrix.get(&Illuminant::A))?;
if flat.len() < 9 {
return None;
}
Some([
[flat[0], flat[1], flat[2]],
[flat[3], flat[4], flat[5]],
[flat[6], flat[7], flat[8]],
])
}
/// The matrix maths, split out so it can be tested without a RAW file.
//
// The constants below are quoted at their published precision rather than
@@ -567,7 +553,7 @@ fn xyz_to_cam_of(image: &rawler::RawImage) -> Option<[[f32; 3]; 3]> {
// a linter makes it harder to check against the specification, and the
// rounding happens identically either way.
#[allow(clippy::excessive_precision)]
fn cam_to_srgb_from(xyz_to_cam: &[[f32; 3]; 3]) -> Option<[f32; 9]> {
pub(crate) fn cam_to_srgb_from(xyz_to_cam: &[[f32; 3]; 3]) -> Option<[f32; 9]> {
// XYZ (D65) → linear sRGB, the standard primaries.
const XYZ_TO_SRGB: [[f32; 3]; 3] = [
[3.2404542, -1.5371385, -0.4985314],
@@ -722,7 +708,7 @@ pub fn daylight_wb(xyz_to_cam: &[[f32; 3]; 3]) -> Option<[f32; 3]> {
}
/// Invert a 3×3 matrix, or `None` if it is singular.
fn invert3(m: &[[f32; 3]; 3]) -> Option<[[f32; 3]; 3]> {
pub(crate) fn invert3(m: &[[f32; 3]; 3]) -> Option<[[f32; 3]; 3]> {
let det = m[0][0] * (m[1][1] * m[2][2] - m[1][2] * m[2][1])
- m[0][1] * (m[1][0] * m[2][2] - m[1][2] * m[2][0])
+ m[0][2] * (m[1][0] * m[2][1] - m[1][1] * m[2][0]);
File diff suppressed because it is too large Load Diff
+6
View File
@@ -25,6 +25,12 @@ pollster.workspace = true
[dev-dependencies]
env_logger.workspace = true
# The detail stage's test consumer — a box blur that is not a develop operation
# and never reaches the panel. An abstraction with no consumers is a guess, and
# this is the one that proves the neighbourhood passes compile, ping-pong,
# encode once, and scale between a proxy and an export. A dev-dependency, so a
# shipping `dr-gpu` does not carry it.
dr-pipeline = { workspace = true, features = ["detail-probe"] }
# The local-adjustment example needs the model, which the library half of this
# crate deliberately does not: `dr-gpu` holds the shaders, and the inference
# runtime belongs to whoever is asking a question about the picture.
+455 -75
View File
@@ -16,9 +16,11 @@
use std::collections::HashMap;
use dr_pipeline::ComposedShader;
use dr_pipeline::detail::ComposedDetail;
use dr_pipeline::{ComposedShader, OutputMode};
use wgpu::util::DeviceExt;
use crate::detail::DetailRunner;
use crate::readback::await_mapping;
use crate::{DemosaicedImage, GpuContext, GpuError};
@@ -32,6 +34,16 @@ use crate::{DemosaicedImage, GpuContext, GpuError};
/// reads them.
const RESERVED_FIELDS: usize = dr_pipeline::RESERVED_UNIFORM_FIELDS;
/// TRACES: FR-DEV-3e
/// The two crates must agree on how many points a base curve has.
///
/// `dr-decode` reads them from the profile database and `dr-pipeline` declares
/// the uniform slots; this file is the only place the two meet, and it packs
/// them by index. A disagreement would not fail to compile — it would upload a
/// curve with a point missing or a stale float in it, which renders as a
/// plausible-looking wrong tone response. Cheaper to catch here, at build time.
const _: () = assert!(dr_decode::base_curve::POINTS == dr_pipeline::BASE_CURVE_POINTS);
/// Runs composed operation chains against demosaiced images.
pub struct AdjustPass {
ctx: GpuContext,
@@ -62,6 +74,44 @@ pub struct AdjustPass {
current: usize,
/// Bound at `@binding(3)` when the edit carries no mask layers.
empty_masks: wgpu::TextureView,
/// TRACES: FR-DEV-3 | FR-DEV-3d
/// The neighbourhood stage — sharpening, noise reduction, clarity and the
/// rest of FR-DEV-3's detail set, which cannot be fused into the shader
/// above because they read pixels they are not writing.
///
/// It lives here rather than beside this pass because the two are one
/// render: when a detail chain is present the fused pass writes a linear
/// intermediate the runner owns, and the runner's last pass writes
/// [`Self::targets`]. Kept as separate objects, a caller could hold a
/// stale intermediate against a fresh colour result with nothing to tell
/// it apart.
detail: DetailRunner,
/// The bind group layout for a fused pass writing a linear intermediate.
///
/// A second layout rather than a second pass: the only difference is the
/// storage texture's format, which is part of the layout and cannot be
/// varied per bind group. Built once here, so a detail operation being
/// switched on does not build a pipeline layout mid-frame.
linear_bind_group_layout: wgpu::BindGroupLayout,
linear_pipeline_layout: wgpu::PipelineLayout,
/// TRACES: FR-DEV-3d
/// What the linear intermediate currently holds, and at what size.
///
/// **This is where `Affects::Detail` stops being bookkeeping.** The key is
/// everything the fused dispatch depends on — the caller's
/// `Invalidation::through(Affects::Colour)`, the compiled structure, the
/// uniform values and the output size. When it matches, the colour pass is
/// skipped and only the detail passes run, so dragging a sharpening slider
/// costs a convolution and not a re-render of the whole chain (FR-DEV-3d).
///
/// Cleared by any render that does not write it, so a stale intermediate
/// cannot survive a change of image and be handed to a later detail chain.
colour_key: Option<(u64, u32, u32)>,
/// Fused dispatches actually encoded. Exposed so a test can see the reuse
/// above happening rather than take it on trust.
colour_dispatches: usize,
/// Detail dispatches encoded.
detail_dispatches: usize,
}
struct Target {
@@ -75,58 +125,7 @@ impl AdjustPass {
pub const FORMAT: wgpu::TextureFormat = wgpu::TextureFormat::Rgba8Unorm;
pub fn new(ctx: &GpuContext) -> Self {
let bind_group_layout =
ctx.device
.create_bind_group_layout(&wgpu::BindGroupLayoutDescriptor {
label: Some("adjust-bgl"),
entries: &[
// The demosaiced source.
wgpu::BindGroupLayoutEntry {
binding: 0,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::Texture {
sample_type: wgpu::TextureSampleType::Float { filterable: true },
view_dimension: wgpu::TextureViewDimension::D2,
multisampled: false,
},
count: None,
},
wgpu::BindGroupLayoutEntry {
binding: 1,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::Buffer {
ty: wgpu::BufferBindingType::Uniform,
has_dynamic_offset: false,
min_binding_size: None,
},
count: None,
},
wgpu::BindGroupLayoutEntry {
binding: 2,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::StorageTexture {
access: wgpu::StorageTextureAccess::WriteOnly,
format: Self::FORMAT,
view_dimension: wgpu::TextureViewDimension::D2,
},
count: None,
},
// The local-adjustment masks. Present in every layout
// whether or not the edit has any, because the layout
// is built once here and the generated shader declares
// the binding unconditionally for exactly that reason.
wgpu::BindGroupLayoutEntry {
binding: 3,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::Texture {
sample_type: wgpu::TextureSampleType::Float { filterable: true },
view_dimension: wgpu::TextureViewDimension::D2Array,
multisampled: false,
},
count: None,
},
],
});
let bind_group_layout = Self::layout_writing(ctx, Self::FORMAT, "adjust-bgl");
let pipeline_layout = ctx
.device
@@ -136,6 +135,23 @@ impl AdjustPass {
immediate_size: 0,
});
// The same layout with an `Rgba16Float` storage texture, for the fused
// pass when a detail stage follows it and it hands on linear working
// values instead of encoding (see `dr_pipeline::OutputMode`). The
// format is part of a bind group layout and cannot be varied per bind
// group, so this is a second layout rather than a second binding —
// built here, once, so that switching sharpening on does not construct
// a pipeline layout in the middle of a frame.
let linear_bind_group_layout =
Self::layout_writing(ctx, crate::detail::INTERMEDIATE_FORMAT, "adjust-linear-bgl");
let linear_pipeline_layout =
ctx.device
.create_pipeline_layout(&wgpu::PipelineLayoutDescriptor {
label: Some("adjust-linear-layout"),
bind_group_layouts: &[Some(&linear_bind_group_layout)],
immediate_size: 0,
});
// A 1x1 single-layer mask, bound when the edit has no local
// adjustments. The generated shader never samples it — no layer block
// is emitted — but a bind group must still satisfy the layout.
@@ -167,9 +183,81 @@ impl AdjustPass {
targets: [None, None],
current: 0,
empty_masks,
detail: DetailRunner::new(ctx),
linear_bind_group_layout,
linear_pipeline_layout,
colour_key: None,
colour_dispatches: 0,
detail_dispatches: 0,
}
}
/// The fused pass's bind group layout, for a given storage format.
///
/// Two of these exist — one writing `Rgba8Unorm` and one writing
/// `Rgba16Float` — and they differ in exactly one field. Written once and
/// parameterised rather than copied, because two copies of a four-entry
/// layout is how the mask binding comes to be present in one and absent
/// from the other, and a bind group that satisfies neither is a validation
/// error a long way from its cause.
fn layout_writing(
ctx: &GpuContext,
format: wgpu::TextureFormat,
label: &str,
) -> wgpu::BindGroupLayout {
ctx.device
.create_bind_group_layout(&wgpu::BindGroupLayoutDescriptor {
label: Some(label),
entries: &[
// The demosaiced source.
wgpu::BindGroupLayoutEntry {
binding: 0,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::Texture {
sample_type: wgpu::TextureSampleType::Float { filterable: true },
view_dimension: wgpu::TextureViewDimension::D2,
multisampled: false,
},
count: None,
},
wgpu::BindGroupLayoutEntry {
binding: 1,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::Buffer {
ty: wgpu::BufferBindingType::Uniform,
has_dynamic_offset: false,
min_binding_size: None,
},
count: None,
},
wgpu::BindGroupLayoutEntry {
binding: 2,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::StorageTexture {
access: wgpu::StorageTextureAccess::WriteOnly,
format,
view_dimension: wgpu::TextureViewDimension::D2,
},
count: None,
},
// The local-adjustment masks. Present in every layout
// whether or not the edit has any, because the layout is
// built once here and the generated shader declares the
// binding unconditionally for exactly that reason.
wgpu::BindGroupLayoutEntry {
binding: 3,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::Texture {
sample_type: wgpu::TextureSampleType::Float { filterable: true },
view_dimension: wgpu::TextureViewDimension::D2Array,
multisampled: false,
},
count: None,
},
],
})
}
/// Compile a composed shader, or return the cached pipeline.
///
/// Compilation errors carry the generated source, since a stray line
@@ -197,12 +285,22 @@ impl AdjustPass {
source: wgpu::ShaderSource::Wgsl(shader.source.as_str().into()),
});
// The layout matching what this shader was composed to write. The
// structure hash covers the generated source and the source
// carries the storage format, so the two can never disagree — a
// cached pipeline is always paired with the layout it was built
// against.
let layout = match shader.output_mode {
OutputMode::Encoded => &self.pipeline_layout,
OutputMode::LinearWorking => &self.linear_pipeline_layout,
};
let pipeline =
self.ctx
.device
.create_compute_pipeline(&wgpu::ComputePipelineDescriptor {
label: Some("adjust-pipeline"),
layout: Some(&self.pipeline_layout),
layout: Some(layout),
module: &module,
entry_point: Some("main"),
compilation_options: Default::default(),
@@ -310,28 +408,31 @@ impl AdjustPass {
height: u32,
masks: Option<&crate::MaskArray>,
) -> Result<&wgpu::Texture, GpuError> {
if shader.output_mode != OutputMode::Encoded {
// Composed for a detail stage and dispatched without one. The
// shader writes `rgba16float` and this path binds an `rgba8unorm`
// storage texture, which wgpu rejects — but well after the point
// where the mistake is legible. Saying so here names the actual
// error: the edit has a neighbourhood operation and needs
// `render_detailed`.
return Err(GpuError::ShaderCompilation(
"this shader was composed with a detail stage and writes linear \
working values; render it with `render_detailed` and the \
matching chain from `EditGraph::compose_detail`"
.into(),
));
}
// Any render that does not write the linear intermediate leaves
// whatever is in it belonging to some other edit — or some other
// photograph. Forgetting this is how a detail chain comes to be run
// over a stale colour result, so the key is dropped rather than
// reasoned about.
self.colour_key = None;
let (width, height) = (width.max(1), height.max(1));
self.ensure_target(width, height);
// Base uniforms: the camera matrix and as-shot white balance, which
// every generated shader reads regardless of which operations are
// active. Framing's slots follow them and are filled by the composer,
// which is why only the first sixteen are written here.
let mut uniforms = shader.uniforms.clone();
if uniforms.len() < RESERVED_FIELDS {
uniforms.resize(RESERVED_FIELDS, 0.0);
}
let m = source.color_matrix();
let wb = source.as_shot_wb();
// Rows padded to vec4 for std140 alignment.
uniforms[0..4].copy_from_slice(&[m[0], m[1], m[2], 0.0]);
uniforms[4..8].copy_from_slice(&[m[3], m[4], m[5], 0.0]);
uniforms[8..12].copy_from_slice(&[m[6], m[7], m[8], 0.0]);
// The fourth slot is the non-linear flag, not padding: it tells the
// shader whether to linearise the sampled texel before any operation
// runs. See `DemosaicedImage::is_non_linear`.
let non_linear = if source.is_non_linear() { 1.0 } else { 0.0 };
uniforms[12..16].copy_from_slice(&[wb[0], wb[1], wb[2], non_linear]);
let uniforms = Self::fused_uniforms(source, shader);
let params_buf = self
.ctx
@@ -394,6 +495,7 @@ impl AdjustPass {
pass.dispatch_workgroups(width.div_ceil(8), height.div_ceil(8), 1);
}
self.ctx.queue.submit(Some(enc.finish()));
self.colour_dispatches += 1;
Ok(&self.targets[self.current]
.as_ref()
@@ -401,12 +503,286 @@ impl AdjustPass {
.texture)
}
/// TRACES: FR-DEV-3 | FR-DEV-3d | FR-DEV-4 | FR-DSP-1
/// Render one frame with a neighbourhood stage.
///
/// `shader` and `detail` must be the two halves of **one** composition —
/// `EditGraph::compose_for` and `EditGraph::compose_detail_for` on the same
/// graph, at the same output space. The fused pass stops at linear working
/// values when a detail stage exists and the last detail pass performs the
/// output transform, so a mismatched pair either encodes twice or not at
/// all.
///
/// An empty `detail` falls through to [`Self::render_masked`], which is
/// the honest thing to do rather than an optimisation: an edit with no
/// active sharpening *is* an ordinary edit, and it should cost exactly
/// what one costs.
///
/// # `colour_key`, and why the caller supplies it
///
/// It is `Invalidation::through(Affects::Colour)` for this edit, mixed
/// with whatever names the photograph — a `VersionId`, typically. When it
/// is unchanged, and the size and the composed shader and its uniforms are
/// unchanged with it, the fused dispatch is **skipped** and the linear
/// intermediate from the previous frame is convolved again. Dragging a
/// sharpening slider then costs the detail passes alone, which is the
/// reuse FR-DEV-3d asks for and the operational meaning of
/// `Affects::Detail`.
///
/// The caller supplies it rather than this pass deriving it because only
/// the caller knows which *image* is on screen. Everything else that goes
/// into the fused dispatch — the shader's structure, its uniform values,
/// the output size — is mixed in here, so a caller cannot make the reuse
/// unsound by supplying a key that is merely coarse. It can only do so by
/// supplying one that fails to distinguish two photographs, which is why
/// the identity of the image is spelled out as its job.
// Eight arguments, and every one of them is a distinct thing the render
// depends on: the image, both halves of the composition, the size, the
// masks and the cache key. Bundling them into a struct would move the
// problem rather than solve it — the caller would fill in the same eight
// fields — and would hide that composing the two halves apart is the one
// mistake this signature exists to make visible.
#[allow(clippy::too_many_arguments)]
pub fn render_detailed(
&mut self,
source: &DemosaicedImage,
shader: &ComposedShader,
width: u32,
height: u32,
masks: Option<&crate::MaskArray>,
detail: &ComposedDetail,
colour_key: u64,
) -> Result<&wgpu::Texture, GpuError> {
if detail.is_empty() {
return self.render_masked(source, shader, width, height, masks);
}
if shader.output_mode != OutputMode::LinearWorking {
return Err(GpuError::ShaderCompilation(
"this detail chain expects a fused pass composed to hand on \
linear working values, but the shader given encodes its own \
output; compose both halves from the same graph"
.into(),
));
}
let (width, height) = (width.max(1), height.max(1));
self.ensure_target(width, height);
let uniforms = Self::fused_uniforms(source, shader);
let key = Self::colour_signature(colour_key, shader, &uniforms, masks);
let reuse = self.colour_key == Some((key, width, height));
// Compile before borrowing anything: `pipeline` and `colour_target`
// both want `&mut self`, and the second holds its borrow across the
// encode below.
self.pipeline(shader)?;
let colour_view = self
.detail
.colour_target(detail.len(), width, height)
.clone();
let mut enc = self
.ctx
.device
.create_command_encoder(&wgpu::CommandEncoderDescriptor {
label: Some("adjust-detail-encoder"),
});
if !reuse {
let params_buf = self
.ctx
.device
.create_buffer_init(&wgpu::util::BufferInitDescriptor {
label: Some("adjust-params"),
contents: bytemuck::cast_slice(&uniforms),
usage: wgpu::BufferUsages::UNIFORM,
});
let bind_group = self
.ctx
.device
.create_bind_group(&wgpu::BindGroupDescriptor {
label: Some("adjust-linear-bg"),
layout: &self.linear_bind_group_layout,
entries: &[
wgpu::BindGroupEntry {
binding: 0,
resource: wgpu::BindingResource::TextureView(source.view()),
},
wgpu::BindGroupEntry {
binding: 1,
resource: params_buf.as_entire_binding(),
},
wgpu::BindGroupEntry {
binding: 2,
resource: wgpu::BindingResource::TextureView(&colour_view),
},
wgpu::BindGroupEntry {
binding: 3,
resource: wgpu::BindingResource::TextureView(
masks.map_or(&self.empty_masks, |m| m.view()),
),
},
],
});
let pipeline = self
.cache
.get(&shader.structure_hash)
.expect("compiled above");
let mut pass = enc.begin_compute_pass(&wgpu::ComputePassDescriptor {
label: Some("adjust-pass"),
timestamp_writes: None,
});
pass.set_pipeline(pipeline);
pass.set_bind_group(0, &bind_group, &[]);
pass.dispatch_workgroups(width.div_ceil(8), height.div_ceil(8), 1);
drop(pass);
self.colour_dispatches += 1;
}
// One encoder for the colour pass and every detail pass, submitted
// once — the shape `MaskPass::render` established. Submission order is
// the whole of the synchronisation: each pass reads what the previous
// one wrote, through the same queue.
let target_view = self.targets[self.current]
.as_ref()
.expect("ensured above")
.view
.clone();
let ran = self
.detail
.encode(&mut enc, detail, &target_view, width, height)?;
self.ctx.queue.submit(Some(enc.finish()));
self.detail_dispatches += ran;
self.colour_key = Some((key, width, height));
Ok(&self.targets[self.current]
.as_ref()
.expect("ensured above")
.texture)
}
/// The fused pass's uniform block, with the source's own values written in.
///
/// Split out because both render paths need exactly this and a second copy
/// would eventually disagree about where the camera matrix goes — which is
/// silent, and corrupts every operation's uniforms downstream of it.
fn fused_uniforms(source: &DemosaicedImage, shader: &ComposedShader) -> Vec<f32> {
// Base uniforms: the camera matrix and as-shot white balance, which
// every generated shader reads regardless of which operations are
// active. Framing's slots follow them and are filled by the composer,
// which is why only the first sixteen are written here.
let mut uniforms = shader.uniforms.clone();
if uniforms.len() < RESERVED_FIELDS {
uniforms.resize(RESERVED_FIELDS, 0.0);
}
let m = source.color_matrix();
let wb = source.as_shot_wb();
// Rows padded to vec4 for std140 alignment.
uniforms[0..4].copy_from_slice(&[m[0], m[1], m[2], 0.0]);
uniforms[4..8].copy_from_slice(&[m[3], m[4], m[5], 0.0]);
uniforms[8..12].copy_from_slice(&[m[6], m[7], m[8], 0.0]);
// The fourth slot is the non-linear flag, not padding: it tells the
// shader whether to linearise the sampled texel before any operation
// runs. See `DemosaicedImage::is_non_linear`.
let non_linear = if source.is_non_linear() { 1.0 } else { 0.0 };
uniforms[12..16].copy_from_slice(&[wb[0], wb[1], wb[2], non_linear]);
// TRACES: FR-DEV-3e
// The camera profile's base curve, packed the way the generated block
// declares it: four x, four y, then the fifth point and the flag. The
// flag is what lets one compiled shader serve a profiled body and an
// unprofiled one, so the pipeline cache is not split in two by which
// camera took the frame.
//
// Written here rather than at the call site so that *both* callers —
// the plain render and the masked one — carry the profile. Filling it
// at one of them was how the two halves of this merge each had it.
let curve = source.base_curve();
let on = if curve.is_identity() { 0.0 } else { 1.0 };
let b = dr_pipeline::BASE_CURVE_UNIFORM_OFFSET;
uniforms[b..b + 4].copy_from_slice(&curve.xs[0..4]);
uniforms[b + 4..b + 8].copy_from_slice(&curve.ys[0..4]);
uniforms[b + 8..b + 12].copy_from_slice(&[curve.xs[4], curve.ys[4], on, 0.0]);
uniforms
}
/// TRACES: FR-DEV-3d
/// Everything the fused dispatch depends on, in one integer.
///
/// The caller's edit key, plus the three things the caller does not know
/// about: which pipeline was compiled, what was uploaded to it, and which
/// mask array was bound. Hashing the uniforms rather than trusting the
/// caller's key to cover them is what makes the reuse safe against a
/// caller whose key is coarser than it should be — and the uniforms are
/// parameters and matrix coefficients from the CPU, never rendered floats,
/// so hashing their bit patterns satisfies ARCH §6.13.
fn colour_signature(
caller: u64,
shader: &ComposedShader,
uniforms: &[f32],
masks: Option<&crate::MaskArray>,
) -> u64 {
let mut h: u64 = 0xcbf2_9ce4_8422_2325;
let mut mix = |v: u64| {
for byte in v.to_le_bytes() {
h ^= u64::from(byte);
h = h.wrapping_mul(0x100_0000_01b3);
}
};
mix(caller);
mix(shader.structure_hash);
for v in uniforms {
// Negative zero folded onto zero: the two render identically, and
// a slider that reached zero from below must not miss the cache.
mix(u64::from(if *v == 0.0 { 0 } else { v.to_bits() }));
}
match masks {
None => mix(0),
Some(m) => {
let (w, h) = m.size();
mix(1);
mix(u64::from(w));
mix(u64::from(h));
mix(u64::from(m.layers()));
}
}
h
}
/// How many distinct pipelines are compiled. Exposed for tests asserting
/// that slider movement does not recompile.
pub fn cached_pipelines(&self) -> usize {
self.cache.len()
}
/// How many detail-pass pipelines are compiled. As above, for the stage
/// that runs after this one.
pub fn cached_detail_pipelines(&self) -> usize {
self.detail.cached_pipelines()
}
/// TRACES: FR-DEV-3d
/// Fused colour dispatches encoded since this pass was created.
///
/// Exists to be asserted on. The saving `Affects::Detail` buys — a
/// sharpening slider that does not re-run the colour chain — is invisible
/// in the output by construction, since the picture is meant to be
/// identical either way. A counter is the only thing that can see it.
pub fn colour_dispatches(&self) -> usize {
self.colour_dispatches
}
/// Detail dispatches encoded since this pass was created.
pub fn detail_dispatches(&self) -> usize {
self.detail_dispatches
}
/// How many linear intermediates have been allocated. For tests: see
/// [`crate::MaskPass::allocations`] for the regression this catches.
pub fn detail_allocations(&self) -> usize {
self.detail.allocations()
}
/// The texture the last render wrote, if there has been one.
pub fn output(&self) -> Option<&wgpu::Texture> {
self.targets[self.current].as_ref().map(|t| &t.texture)
@@ -501,7 +877,7 @@ impl AdjustPass {
}
/// Number the lines of generated source, so a compiler error can be located.
fn numbered(src: &str) -> String {
pub(crate) fn numbered(src: &str) -> String {
src.lines()
.enumerate()
.map(|(i, l)| format!("{:>4} | {l}", i + 1))
@@ -512,7 +888,7 @@ fn numbered(src: &str) -> String {
#[cfg(test)]
mod tests {
use super::*;
use dr_decode::{CfaPattern, CropRect, RawImage};
use dr_decode::{BaseCurve, CfaPattern, CropRect, RawImage};
use dr_pipeline::ops::{colour_mixer, exposure, saturation};
use dr_pipeline::EditGraph;
@@ -546,6 +922,7 @@ mod tests {
// Identity, so the test reasons about the operations alone
// rather than about a camera's colour response.
color_matrix: Some([1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0]),
base_curve: BaseCurve::IDENTITY,
crop: CropRect {
x: 0,
y: 0,
@@ -719,6 +1096,7 @@ mod tests {
white_level: 16383,
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
color_matrix: Some([1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0]),
base_curve: BaseCurve::IDENTITY,
crop: CropRect {
x: 0,
y: 0,
@@ -1141,6 +1519,7 @@ mod tests {
white_level: 16383,
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
color_matrix: Some([1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0]),
base_curve: BaseCurve::IDENTITY,
crop: CropRect {
x: 0,
y: 0,
@@ -1240,6 +1619,7 @@ mod tests {
white_level: 16383,
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
color_matrix: Some([1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0]),
base_curve: BaseCurve::IDENTITY,
crop: CropRect {
x: 0,
y: 0,
+34 -1
View File
@@ -10,7 +10,7 @@
//! pass over this texture; it does not re-demosaic, which is what keeps the
//! interaction budget (NFR-P9) reachable on a 24 MP file.
use dr_decode::{CfaPattern, RawImage};
use dr_decode::{BaseCurve, CfaPattern, RawImage};
use wgpu::util::DeviceExt;
use crate::{GpuContext, GpuError};
@@ -83,6 +83,16 @@ pub struct DemosaicedImage {
color_matrix: [f32; 9],
/// As-shot white balance, the neutral starting point for the WB control.
as_shot_wb: [f32; 3],
/// TRACES: FR-DEV-3e
/// The camera profile's rendering curve, carried through for the adjust
/// pass exactly as `color_matrix` is.
///
/// It rides on the image rather than on the edit graph because it is not
/// an edit: it belongs to the body that took the frame, the way the
/// masked-photosite crop and the EXIF orientation do, and a sidecar shared
/// between two bodies must never carry one body's rendering onto the
/// other's file (FR-NC-9).
base_curve: BaseCurve,
/// Whether the texture holds gamma-encoded rather than linear values.
non_linear: bool,
}
@@ -108,6 +118,15 @@ impl DemosaicedImage {
self.color_matrix
}
/// TRACES: FR-DEV-3e
/// The camera profile's base curve, as five `(x, y)` points.
///
/// [`BaseCurve::IDENTITY`] where the body is unprofiled or the source was
/// never raw, in which case the adjust pass skips the stage entirely.
pub fn base_curve(&self) -> BaseCurve {
self.base_curve
}
/// As-shot white balance multipliers, green-normalised.
///
/// The white balance control is expressed *relative* to these, so its
@@ -220,6 +239,12 @@ impl DemosaicedImage {
height,
color_matrix: IDENTITY_3X3,
as_shot_wb: [1.0, 1.0, 1.0],
// **The identity, and this is the whole reason the field is here
// rather than resolved further down.** A JPEG has already had its
// camera's base curve baked in by the camera; applying one again
// would render the rendering, crushing the shadows and flattening
// the highlights of an image that was already finished.
base_curve: BaseCurve::IDENTITY,
non_linear: true,
})
}
@@ -495,6 +520,10 @@ impl Demosaicer {
// no colour transform rather than not at all.
color_matrix: raw.color_matrix.unwrap_or(IDENTITY_3X3),
as_shot_wb: [raw.wb_coeffs[0], raw.wb_coeffs[1], raw.wb_coeffs[2]],
// Whatever the profile database had for this body (FR-DEV-3e),
// resolved at decode because that is the only place the make and
// model are known.
base_curve: raw.base_curve,
// Sensor data is linear by construction — the demosaic shader
// normalises against black and white levels and applies no
// transfer function.
@@ -808,6 +837,7 @@ mod tests {
white_level: white,
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
color_matrix: None,
base_curve: BaseCurve::IDENTITY,
crop: CropRect {
x: 0,
y: 0,
@@ -919,6 +949,7 @@ mod tests {
white_level: white,
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
color_matrix: None,
base_curve: BaseCurve::IDENTITY,
crop: CropRect {
x: 0,
y: 0,
@@ -1164,6 +1195,7 @@ mod tests {
white_level: 16383,
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
color_matrix: None,
base_curve: BaseCurve::IDENTITY,
crop: CropRect {
x: 0,
y: 0,
@@ -1244,6 +1276,7 @@ mod tests {
1.0,
],
color_matrix: None,
base_curve: BaseCurve::IDENTITY,
crop: CropRect {
x: 0,
y: 0,
+396
View File
@@ -0,0 +1,396 @@
//! The detail stage — running `dr-pipeline`'s neighbourhood passes.
//!
//! Where [`crate::AdjustPass`] fuses every point operation into one dispatch,
//! this runs the operations that cannot be fused because they read pixels they
//! are not writing: sharpening, noise reduction, clarity, texture, dehaze,
//! spot removal (FR-DEV-3, FR-DEV-8). `dr_pipeline::detail` decides *what* they
//! are and generates their WGSL; this compiles it, finds it somewhere to
//! write, and dispatches it.
//!
//! # Nothing round-trips
//!
//! Every intermediate here is a `wgpu::Texture` and none of them is ever
//! mapped. The chain is `demosaiced -> fused -> f16 -> f16 -> ... -> rgba8`,
//! all of it on the device, and the last write lands in the same texture the
//! compositor was already being handed. ARCH §6.1 and FR-DEV-4 are satisfied
//! by there being no code here that could violate them, which is the only
//! guarantee worth having.
//!
//! # Following the mask pass rather than inventing a second pattern
//!
//! `mask.rs` established how multi-target work is done in this crate, and this
//! copies it deliberately:
//!
//! - **One encoder for the whole chain.** The mask pass rasterises every layer
//! into one command buffer and submits once; this does the same for every
//! pass. Submission order is the only synchronisation either needs, because
//! both write and then read through the same queue.
//! - **Textures reallocated on size change, never per frame.** `ensure_array`
//! there, [`Intermediates::ensure`] here. Steady-state rendering at one
//! viewport size allocates nothing.
//! - **An allocation counter that exists to be asserted on.** Reallocating per
//! frame instead of per resize costs a great deal of bandwidth and shows up
//! nowhere in the output, which is exactly the kind of regression that needs
//! a test that can see it.
//! - **Pipelines cached by structure hash**, as `AdjustPass` caches its own.
//! Moving a slider re-uploads a uniform buffer; it does not recompile.
//!
//! # The ping-pong, and why there are at most three textures
//!
//! Slot 0 holds what the fused colour pass wrote. It is kept **across frames**,
//! which is what makes [`dr_pipeline::Affects::Detail`] mean something: when
//! only a detail parameter has moved, the colour key is unchanged, the fused
//! dispatch is skipped, and dragging a sharpening slider costs the detail
//! passes alone (FR-DEV-3d).
//!
//! The remaining passes alternate between slots 1 and 2, and the last one
//! writes the display texture directly rather than an intermediate — so a
//! chain of *N* passes costs *N* dispatches and not *N* + 1, and there is no
//! resolve pass to pay for. That leaves the allocation at `1 + min(N-1, 2)`
//! textures: one for a single-pass operation, two for a separable blur, three
//! however long the chain gets after that.
use std::collections::HashMap;
use dr_pipeline::detail::{ComposedDetail, ComposedDetailPass};
use wgpu::util::DeviceExt as _;
use crate::{GpuContext, GpuError};
/// The format every intermediate carries.
///
/// The same `Rgba16Float` the demosaicer produces and the same one ARCH §5.2
/// names as the working precision (FR-DEV-2). It is not a free choice: the
/// stage exists between the colour pass and the output transform precisely so
/// that a kernel runs on linear values at full internal precision, and an
/// 8-bit intermediate would quantise twice and convolve display-encoded
/// numbers — which is how sharpening comes to band a clear sky.
pub const INTERMEDIATE_FORMAT: wgpu::TextureFormat = wgpu::TextureFormat::Rgba16Float;
/// One linear working texture.
struct Slot {
#[allow(dead_code)]
texture: wgpu::Texture,
view: wgpu::TextureView,
}
/// The pool of linear intermediates, sized to the chain and the viewport.
struct Intermediates {
slots: Vec<Slot>,
width: u32,
height: u32,
allocations: usize,
}
impl Intermediates {
fn new() -> Self {
Self {
slots: Vec::new(),
width: 0,
height: 0,
allocations: 0,
}
}
/// Make sure `count` textures of this size exist.
///
/// Grows but never shrinks within a size: an edit that briefly had a
/// three-pass chain and then a one-pass one keeps the spare texture rather
/// than freeing and reallocating it the next time the user turns the
/// operation back on. A size change drops the lot, because none of them
/// fits any more.
fn ensure(&mut self, ctx: &GpuContext, count: usize, width: u32, height: u32) {
if self.width != width || self.height != height {
self.slots.clear();
self.width = width;
self.height = height;
}
while self.slots.len() < count {
let texture = ctx.device.create_texture(&wgpu::TextureDescriptor {
label: Some("detail-intermediate"),
size: wgpu::Extent3d {
width,
height,
depth_or_array_layers: 1,
},
mip_level_count: 1,
sample_count: 1,
dimension: wgpu::TextureDimension::D2,
format: INTERMEDIATE_FORMAT,
// STORAGE_BINDING to be written by a compute pass and
// TEXTURE_BINDING to be read by the next one. Nothing else:
// no RENDER_ATTACHMENT, because unlike the adjust pass's
// output these are never handed to a compositor, and no
// COPY_SRC, because nothing reads them back — that is the
// point (ARCH §6.1).
usage: wgpu::TextureUsages::STORAGE_BINDING
| wgpu::TextureUsages::TEXTURE_BINDING,
view_formats: &[],
});
let view = texture.create_view(&Default::default());
self.slots.push(Slot { texture, view });
self.allocations += 1;
}
}
}
/// Runs the detail stage.
///
/// Owned by [`crate::AdjustPass`] rather than standing alone, because the two
/// halves are one render: the fused pass writes slot 0, this reads it, and the
/// last pass writes the adjust pass's own output texture. Splitting them into
/// two objects with two lifetimes would mean a caller could hold a stale
/// intermediate against a fresh colour result and never be told.
pub(crate) struct DetailRunner {
ctx: GpuContext,
/// Layout for a pass writing another linear intermediate.
to_linear: Layout,
/// Layout for the last pass, which writes the display texture.
to_output: Layout,
/// Compiled pipelines by pass structure hash.
cache: HashMap<u64, wgpu::ComputePipeline>,
pool: Intermediates,
}
struct Layout {
bind_group: wgpu::BindGroupLayout,
pipeline: wgpu::PipelineLayout,
}
impl DetailRunner {
pub(crate) fn new(ctx: &GpuContext) -> Self {
Self {
ctx: ctx.clone(),
to_linear: Layout::new(ctx, INTERMEDIATE_FORMAT, "detail-linear"),
to_output: Layout::new(ctx, crate::AdjustPass::FORMAT, "detail-output"),
cache: HashMap::new(),
pool: Intermediates::new(),
}
}
/// The view the fused colour pass should write, given a chain of `passes`.
///
/// Slot 0, always — it is the one that survives between frames so that a
/// detail-only change can skip the colour dispatch entirely.
pub(crate) fn colour_target(
&mut self,
passes: usize,
width: u32,
height: u32,
) -> &wgpu::TextureView {
// One for the colour pass's result, then one per hand-off between
// detail passes, capped at two because a ping-pong needs no more: the
// last pass writes the display texture rather than an intermediate.
let needed = 1 + passes.saturating_sub(1).min(2);
self.pool.ensure(&self.ctx, needed, width, height);
&self.pool.slots[0].view
}
/// Encode every pass of `chain`, the last one writing `output`.
///
/// The caller must already have run the fused colour pass into
/// [`Self::colour_target`] — or established that a previous frame's is
/// still valid, which is the whole point of keeping slot 0.
pub(crate) fn encode(
&mut self,
encoder: &mut wgpu::CommandEncoder,
chain: &ComposedDetail,
output: &wgpu::TextureView,
width: u32,
height: u32,
) -> Result<usize, GpuError> {
for pass in &chain.passes {
self.compile(pass)?;
}
for (index, pass) in chain.passes.iter().enumerate() {
// Read what the previous pass wrote; write the next slot, or the
// display texture if this is the last one. `index % 2` alternates
// between slots 1 and 2, so a pass never reads the texture it is
// writing — which on a compute pass is not an error the driver
// reports, merely a picture that depends on scheduling.
let source_slot = if index == 0 { 0 } else { 2 - (index % 2) };
let source = &self.pool.slots[source_slot].view;
let destination = if pass.writes_output {
output
} else {
&self.pool.slots[1 + (index % 2)].view
};
let layout = if pass.writes_output {
&self.to_output
} else {
&self.to_linear
};
let params = self
.ctx
.device
.create_buffer_init(&wgpu::util::BufferInitDescriptor {
label: Some("detail-params"),
contents: bytemuck::cast_slice(&pass.uniforms),
usage: wgpu::BufferUsages::UNIFORM,
});
let bind_group = self
.ctx
.device
.create_bind_group(&wgpu::BindGroupDescriptor {
label: Some("detail-bg"),
layout: &layout.bind_group,
entries: &[
wgpu::BindGroupEntry {
binding: 0,
resource: wgpu::BindingResource::TextureView(source),
},
wgpu::BindGroupEntry {
binding: 1,
resource: params.as_entire_binding(),
},
wgpu::BindGroupEntry {
binding: 2,
resource: wgpu::BindingResource::TextureView(destination),
},
],
});
let pipeline = self
.cache
.get(&pass.structure_hash)
.expect("compiled above");
let mut compute = encoder.begin_compute_pass(&wgpu::ComputePassDescriptor {
label: Some(pass.label.as_str()),
timestamp_writes: None,
});
compute.set_pipeline(pipeline);
compute.set_bind_group(0, &bind_group, &[]);
compute.dispatch_workgroups(width.div_ceil(8), height.div_ceil(8), 1);
}
Ok(chain.passes.len())
}
/// Compile one pass, or leave the cached pipeline in place.
///
/// A validation error here is a codegen bug rather than anything the user
/// did, so it is caught in an error scope and returned with the generated
/// source and the pass's label attached — a line number against code
/// nobody wrote, from one of several passes, is otherwise close to
/// unactionable.
fn compile(&mut self, pass: &ComposedDetailPass) -> Result<(), GpuError> {
if self.cache.contains_key(&pass.structure_hash) {
return Ok(());
}
let scope = self
.ctx
.device
.push_error_scope(wgpu::ErrorFilter::Validation);
let module = self
.ctx
.device
.create_shader_module(wgpu::ShaderModuleDescriptor {
label: Some(pass.label.as_str()),
source: wgpu::ShaderSource::Wgsl(pass.source.as_str().into()),
});
let layout = if pass.writes_output {
&self.to_output
} else {
&self.to_linear
};
let pipeline = self
.ctx
.device
.create_compute_pipeline(&wgpu::ComputePipelineDescriptor {
label: Some(pass.label.as_str()),
layout: Some(&layout.pipeline),
module: &module,
entry_point: Some("main"),
compilation_options: Default::default(),
cache: None,
});
if let Some(err) = pollster::block_on(scope.pop()) {
return Err(GpuError::ShaderCompilation(format!(
"detail pass {}: {err}\n\n--- generated source ---\n{}",
pass.label,
crate::adjust::numbered(&pass.source)
)));
}
self.cache.insert(pass.structure_hash, pipeline);
Ok(())
}
/// How many distinct detail pipelines are compiled. For tests asserting
/// that slider movement does not recompile.
pub(crate) fn cached_pipelines(&self) -> usize {
self.cache.len()
}
/// How many intermediate textures have been allocated since this pass was
/// created. For tests — see [`crate::MaskPass::allocations`] for the
/// regression this shape of counter exists to catch.
pub(crate) fn allocations(&self) -> usize {
self.pool.allocations
}
}
impl Layout {
fn new(ctx: &GpuContext, format: wgpu::TextureFormat, label: &str) -> Self {
let bind_group = ctx
.device
.create_bind_group_layout(&wgpu::BindGroupLayoutDescriptor {
label: Some(label),
entries: &[
// The previous stage's result.
wgpu::BindGroupLayoutEntry {
binding: 0,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::Texture {
sample_type: wgpu::TextureSampleType::Float { filterable: true },
view_dimension: wgpu::TextureViewDimension::D2,
multisampled: false,
},
count: None,
},
wgpu::BindGroupLayoutEntry {
binding: 1,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::Buffer {
ty: wgpu::BufferBindingType::Uniform,
has_dynamic_offset: false,
min_binding_size: None,
},
count: None,
},
wgpu::BindGroupLayoutEntry {
binding: 2,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::StorageTexture {
access: wgpu::StorageTextureAccess::WriteOnly,
format,
view_dimension: wgpu::TextureViewDimension::D2,
},
count: None,
},
],
});
let pipeline = ctx
.device
.create_pipeline_layout(&wgpu::PipelineLayoutDescriptor {
label: Some(label),
bind_group_layouts: &[Some(&bind_group)],
immediate_size: 0,
});
Self {
bind_group,
pipeline,
}
}
}
+6
View File
@@ -19,12 +19,18 @@ use wgpu::util::DeviceExt;
mod adjust;
mod demosaic;
mod detail;
mod error;
mod histogram;
mod mask;
mod readback;
mod segment;
pub use adjust::AdjustPass;
// The format the neighbourhood stage works in. Public because it is a promise
// rather than an implementation detail: a detail pass is guaranteed linear,
// unclipped, full internal precision (FR-DEV-2), and anyone reasoning about
// VRAM at 24 MP needs to know what an intermediate costs.
pub use detail::INTERMEDIATE_FORMAT as DETAIL_INTERMEDIATE_FORMAT;
pub use demosaic::{DemosaicedImage, Demosaicer};
pub use error::GpuError;
// Renamed on the way out: `BINS` says enough inside `histogram`, and nothing
+179
View File
@@ -0,0 +1,179 @@
//! TRACES: FR-DEV-3e
//! The camera profile's base curve, end to end on a device.
//!
//! The unit tests either side of this one check halves. `dr-decode` asserts
//! that the shipped database parses and that every curve in it lifts its
//! midtones; `dr-pipeline` asserts that the generated WGSL evaluates a curve
//! in the right place. Neither would notice if the two agreed with each other
//! and both were wrong — a curve packed into the wrong uniform slots, or a
//! flag read from the wrong component, satisfies both and renders nothing.
//!
//! So this renders real pixels twice, once with a profiled body's curve and
//! once with the identity, and asserts the difference is the one a base curve
//! is for: midtones lifted, black still black, white still white.
use dr_decode::{BaseCurve, CfaPattern, CropRect, RawImage};
use dr_gpu::{AdjustPass, Demosaicer, GpuContext};
use dr_pipeline::EditGraph;
const SIZE: u32 = 16;
fn ctx() -> Option<GpuContext> {
pollster::block_on(GpuContext::new_headless()).ok()
}
/// A flat RGGB frame at `level` out of 65535, carrying `curve`.
///
/// Every photosite the same value, so the demosaic result is a uniform grey
/// and the only thing that can move a pixel is the curve. The colour matrix is
/// the identity and the balance is neutral for the same reason: this test is
/// about one stage, and a real body's matrix would make every assertion below
/// a statement about that body instead.
fn flat_raw(level: u16, curve: BaseCurve) -> RawImage {
RawImage {
width: SIZE,
height: SIZE,
data: vec![level; (SIZE * SIZE) as usize],
cfa_pattern: CfaPattern::Rggb,
black_level: [0; 4],
white_level: u16::MAX,
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
color_matrix: Some([1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0]),
base_curve: curve,
crop: CropRect {
x: 0,
y: 0,
width: SIZE,
height: SIZE,
},
}
}
/// Render a neutral edit over a flat frame and return the centre pixel's red.
///
/// The centre rather than a corner: a demosaic has to invent its edges, and
/// the interpolated border of a 16×16 frame is not where anyone should be
/// reading a tone off.
fn rendered_level(ctx: &GpuContext, level: u16, curve: BaseCurve) -> u8 {
let raw = flat_raw(level, curve);
let source = Demosaicer::new(ctx)
.expect("demosaicer")
.run(&raw)
.expect("demosaic");
let shader = EditGraph::default_chain().compose();
let mut adjust = AdjustPass::new(ctx);
adjust
.render(&source, &shader, SIZE, SIZE)
.expect("render");
let (pixels, _, _) = adjust.export_pixels().expect("readback");
let centre = ((SIZE / 2) * SIZE + SIZE / 2) * 4;
pixels[centre as usize]
}
/// The Canon EOS 6D's curve, from the shipped profile database.
///
/// Looked up by name rather than written out, so this also asserts the thing
/// no other test can: that a curve travels from the YAML, through the body
/// match, onto the decoded image and into the uniform block that the shader
/// actually reads.
fn six_d() -> BaseCurve {
let curve = dr_decode::base_curve::for_body("Canon", "EOS 6D");
assert!(
!curve.is_identity(),
"the shipped database must have a curve for the EOS 6D"
);
curve
}
#[test]
fn a_profiled_body_renders_brighter_midtones_than_a_flat_one() {
// **The whole requirement, in one assertion.** A linear midtone renders
// roughly half a stop dark, which is the flat, lifeless look FR-DEV-3e
// exists to get away from. If the curve did not reach the shader — wrong
// slot, wrong flag, wrong stage — this is the only test that would fail.
let Some(ctx) = ctx() else {
eprintln!("skipping: no GPU adapter");
return;
};
// 13% of full scale: roughly where a camera places middle grey, leaving
// about two and a half stops of highlight headroom above it.
let level = (0.13 * 65535.0) as u16;
let flat = rendered_level(&ctx, level, BaseCurve::IDENTITY);
let profiled = rendered_level(&ctx, level, six_d());
assert!(
profiled > flat + 8,
"the profile lifted middle grey from {flat} only to {profiled}"
);
}
#[test]
fn the_curve_leaves_black_black_and_white_white() {
// A base curve renders the range between the endpoints; it must not move
// the endpoints themselves. A curve that lifted black would put a grey
// veil over every night photograph, and one that pulled white down would
// make a correctly exposed frame look underexposed.
let Some(ctx) = ctx() else {
eprintln!("skipping: no GPU adapter");
return;
};
let curve = six_d();
assert_eq!(rendered_level(&ctx, 0, curve), 0, "black moved");
assert_eq!(rendered_level(&ctx, u16::MAX, curve), 255, "white moved");
}
#[test]
fn an_unprofiled_body_renders_exactly_as_it_did_before_profiles_existed() {
// The graceful fallback, asserted as a number rather than as a promise.
// With no curve the pipeline must still be a pass-through: black level
// out, white level in, sRGB encoding on the way to the screen and nothing
// else. "Never worse than today" is the one property this change was not
// allowed to trade away, and the way it would break is silently — a flag
// read from the wrong component would apply a curve nobody asked for.
let Some(ctx) = ctx() else {
eprintln!("skipping: no GPU adapter");
return;
};
for level in [0u16, 4_000, 8_520, 32_768, 60_000, u16::MAX] {
let scene = f32::from(level) / f32::from(u16::MAX);
let expected = (dr_types::Transfer::Srgb.encode(scene) * 255.0).round() as i32;
let got = i32::from(rendered_level(&ctx, level, BaseCurve::IDENTITY));
// Two 8-bit steps: the texture holding the demosaiced frame is
// `Rgba16Float`, so a value round-trips through eleven mantissa bits
// before it is encoded. That is well under one step at any level, and
// the tolerance is for the rounding either side of it rather than for
// the transform being approximate.
assert!(
(got - expected).abs() <= 2,
"raw {level} rendered as {got}, expected about {expected}"
);
}
}
#[test]
fn the_curve_is_monotone_through_the_whole_range() {
// The property the spline's tangent limiting exists to guarantee, checked
// where it actually matters: on the device, through the real uniform
// packing. A curve that dipped anywhere would put a dark band across a
// smooth gradient — a sky, most visibly — and it would read as a
// rendering fault rather than as a bad profile.
let Some(ctx) = ctx() else {
eprintln!("skipping: no GPU adapter");
return;
};
let curve = six_d();
let mut previous = 0u8;
for step in 0..=16u32 {
let level = (step * 65535 / 16) as u16;
let value = rendered_level(&ctx, level, curve);
assert!(
value >= previous,
"the curve fell from {previous} to {value} at raw level {level}"
);
previous = value;
}
}
+408
View File
@@ -0,0 +1,408 @@
//! The neighbourhood stage, end to end on a real device.
//!
//! `dr-pipeline`'s own tests assert what the composer *generates*; nothing
//! there can tell whether the WGSL compiles, whether pass two is handed what
//! pass one wrote, or whether the output transform happens exactly once. Those
//! are questions only a GPU answers, and they are the ones that decide whether
//! a future sharpening operation works or draws nonsense.
//!
//! The consumer is `detail_probe`, a separable box blur that is not a develop
//! operation (see `dr_pipeline::detail::probe`). A box blur is used because its
//! answer is known in closed form: over a step edge it produces a ramp exactly
//! `2r + 1` pixels wide with a computable value at every step, so these tests
//! assert **pixels** rather than "something changed".
//!
//! # Reading the expected values
//!
//! The source is uploaded through `DemosaicedImage::from_rgba8`, which flags it
//! non-linear, so the generated shader decodes sRGB before any operation runs.
//! A black/white step therefore reaches the detail stage as linear 0.0 and 1.0
//! exactly. The blur averages those, and the last detail pass re-encodes. So
//! the expected byte at a column is `srgb_encode(white_taps / (2r + 1))`, with
//! taps clamped at the border — which is exactly what `expected_profile`
//! computes.
use dr_gpu::{AdjustPass, DemosaicedImage, GpuContext};
use dr_pipeline::descriptor::{OpId, ParamId};
use dr_pipeline::detail::probe::BoxBlur;
use dr_pipeline::{Affects, EditGraph, OutputMode};
use dr_types::ColourSpace;
const PROBE: OpId = OpId("detail_probe");
const RADIUS: ParamId = ParamId("radius");
fn ctx() -> Option<GpuContext> {
// CI runners and headless machines may have no usable adapter. Skip rather
// than fail, exactly as the rest of this crate's device tests do.
match pollster::block_on(GpuContext::new_headless()) {
Ok(c) => Some(c),
Err(e) => {
eprintln!("skipping: no GPU adapter ({e})");
None
}
}
}
/// A vertical step edge: black to the left of `size / 2`, white to the right.
///
/// The one image whose blur is worth checking by hand. A gradient would
/// average to itself and hide a kernel that is off by one; a step does not.
fn step_edge(ctx: &GpuContext, size: u32) -> DemosaicedImage {
let data: Vec<u8> = (0..size * size)
.flat_map(|i| {
let x = i % size;
let v = if x < size / 2 { 0u8 } else { 255 };
[v, v, v, 255]
})
.collect();
DemosaicedImage::from_rgba8(ctx, &data, size, size).expect("upload")
}
/// One row of the rendered image, red channel, as bytes.
fn row(pixels: &[u8], size: u32, y: u32) -> Vec<u8> {
(0..size)
.map(|x| pixels[((y * size + x) * 4) as usize])
.collect()
}
fn srgb_encode(v: f32) -> u8 {
let e = if v <= 0.003_130_8 {
v * 12.92
} else {
1.055 * v.powf(1.0 / 2.4) - 0.055
};
(e.clamp(0.0, 1.0) * 255.0).round() as u8
}
/// What a separable box blur of radius `r` must produce over the step edge.
fn expected_profile(size: u32, r: i32) -> Vec<u8> {
let last = size as i32 - 1;
let edge = (size / 2) as i32;
(0..size as i32)
.map(|x| {
let white = (-r..=r)
.filter(|i| (x + i).clamp(0, last) >= edge)
.count();
srgb_encode(white as f32 / (2 * r + 1) as f32)
})
.collect()
}
/// Render one graph, with its detail stage, and read the pixels back.
///
/// This is the whole calling convention a frontend has to adopt, in five
/// lines: compose both halves from one graph at one output space, ask the
/// graph for the scale, and pass the invalidation key through.
fn render(
ctx: &GpuContext,
pass: &mut AdjustPass,
graph: &EditGraph,
source: &DemosaicedImage,
out: u32,
) -> Vec<u8> {
let _ = ctx;
let shader = graph.compose_for(ColourSpace::Srgb);
let scale = graph.render_scale(source.size(), (out, out));
let detail = graph.compose_detail_for(scale, ColourSpace::Srgb);
let key = graph.invalidation().through(Affects::Colour);
pass.render_detailed(source, &shader, out, out, None, &detail, key)
.expect("render");
pass.export_pixels().expect("readback").0
}
#[test]
fn a_neighbourhood_pass_produces_the_pixels_it_should() {
// The whole seam, proved once: an operation that reads its neighbours runs
// on the GPU, and the values it writes are the ones a box blur is defined
// to write. Not "the edge got softer" — every byte of the ramp.
let Some(ctx) = ctx() else { return };
const SIZE: u32 = 64;
let mut graph = EditGraph::with_detail_probe();
graph.set_param(PROBE, RADIUS, 0.0625); // 4 px on a 64 px edge
let source = step_edge(&ctx, SIZE);
let mut pass = AdjustPass::new(&ctx);
let pixels = render(&ctx, &mut pass, &graph, &source, SIZE);
let got = row(&pixels, SIZE, SIZE / 2);
let r = BoxBlur::with_radius(0.0625).kernel(graph.render_scale((SIZE, SIZE), (SIZE, SIZE)));
assert_eq!(r, 4, "5/64 of the shorter edge, rounded");
let want = expected_profile(SIZE, r as i32);
for (x, (a, b)) in got.iter().zip(&want).enumerate() {
assert!(
a.abs_diff(*b) <= 2,
"column {x}: got {a}, expected {b}\ngot: {got:?}\nwant: {want:?}"
);
}
}
#[test]
fn the_second_pass_reads_what_the_first_one_wrote() {
// The ping-pong, stated as a property of the picture rather than of the
// plumbing. A separable blur is symmetric: applied to a *horizontal* step
// it must also soften a horizontal edge in the other direction. Wire the
// second pass to read the original again and the vertical smear vanishes,
// which is exactly what this sees.
let Some(ctx) = ctx() else { return };
const SIZE: u32 = 64;
// A quadrant image: the vertical pass has something to do only if it is
// reading the horizontal pass's output rather than the source.
let data: Vec<u8> = (0..SIZE * SIZE)
.flat_map(|i| {
let (x, y) = (i % SIZE, i / SIZE);
let v = if (x < SIZE / 2) == (y < SIZE / 2) {
0u8
} else {
255
};
[v, v, v, 255]
})
.collect();
let source = DemosaicedImage::from_rgba8(&ctx, &data, SIZE, SIZE).expect("upload");
let mut graph = EditGraph::with_detail_probe();
graph.set_param(PROBE, RADIUS, 0.0625);
let mut pass = AdjustPass::new(&ctx);
let pixels = render(&ctx, &mut pass, &graph, &source, SIZE);
// Two separable passes compose into a true two-dimensional box average —
// but only if the second reads the first's output. Computed in closed form
// over the same window the shader uses, so this is an assertion about
// values rather than about direction.
let r = 4i32;
let last = SIZE as i32 - 1;
let half = (SIZE / 2) as i32;
let quadrant_is_black = |x: i32, y: i32| (x < half) == (y < half);
let want: Vec<u8> = (0..SIZE as i32)
.map(|x| {
let y = half;
let mut white = 0usize;
for dy in -r..=r {
for dx in -r..=r {
let (sx, sy) = ((x + dx).clamp(0, last), (y + dy).clamp(0, last));
if !quadrant_is_black(sx, sy) {
white += 1;
}
}
}
srgb_encode(white as f32 / ((2 * r + 1) * (2 * r + 1)) as f32)
})
.collect();
let got = row(&pixels, SIZE, SIZE / 2);
for (x, (a, b)) in got.iter().zip(&want).enumerate() {
// A second pass reading the *source* instead would leave column 20 at
// 255 where a real 2D average puts it near 196 — so the failure this
// catches is loud, not marginal.
assert!(
a.abs_diff(*b) <= 2,
"column {x}: got {a}, expected {b}\ngot: {got:?}\nwant: {want:?}"
);
}
}
#[test]
fn an_inactive_detail_operation_costs_exactly_nothing() {
// The rule the whole pipeline rests on, carried into this stage. A
// photograph with no sharpening must render through the single fused
// dispatch it always did, allocate no intermediate, and — the part worth
// checking — produce byte-identical pixels to a graph that has no
// neighbourhood operation in it at all.
let Some(ctx) = ctx() else { return };
const SIZE: u32 = 32;
let source = step_edge(&ctx, SIZE);
let probe = EditGraph::with_detail_probe();
assert_eq!(
probe.compose_for(ColourSpace::Srgb).output_mode,
OutputMode::Encoded,
"a neutral detail operation must not change how the fused pass ends"
);
let mut with_probe = AdjustPass::new(&ctx);
let a = render(&ctx, &mut with_probe, &probe, &source, SIZE);
assert_eq!(with_probe.colour_dispatches(), 1);
assert_eq!(with_probe.detail_dispatches(), 0);
assert_eq!(with_probe.detail_allocations(), 0, "nothing was allocated");
let plain = EditGraph::default_chain();
let mut without = AdjustPass::new(&ctx);
let b = render(&ctx, &mut without, &plain, &source, SIZE);
assert_eq!(a, b, "an operation at its defaults must not touch the image");
}
#[test]
fn moving_a_detail_parameter_does_not_re_run_the_colour_pass() {
// TRACES: FR-DEV-3d, and the operational point of `Affects::Detail`.
//
// Invisible in the output by construction — the picture is meant to be
// whatever the sharpening says whichever way it was computed — so a
// dispatch counter is the only thing that can see it. Without this, the
// whole invalidation story is a comment.
let Some(ctx) = ctx() else { return };
const SIZE: u32 = 64;
let source = step_edge(&ctx, SIZE);
let mut pass = AdjustPass::new(&ctx);
let mut graph = EditGraph::with_detail_probe();
graph.set_param(PROBE, RADIUS, 0.0625);
render(&ctx, &mut pass, &graph, &source, SIZE);
assert_eq!(pass.colour_dispatches(), 1);
assert_eq!(pass.detail_dispatches(), 2, "a separable blur is two passes");
// Drag the sharpening slider. The colour chain is untouched, so the linear
// intermediate it wrote is still exactly right.
graph.set_param(PROBE, RADIUS, 0.09);
render(&ctx, &mut pass, &graph, &source, SIZE);
assert_eq!(
pass.colour_dispatches(),
1,
"the fused colour pass re-ran for a change it does not depend on"
);
assert_eq!(pass.detail_dispatches(), 4);
// Now move exposure. The detail stage reads what the colour pass wrote, so
// this one genuinely does have to re-run both — anything else would show a
// sharpened version of the previous exposure.
graph.set_param(
dr_pipeline::ops::exposure::ID,
dr_pipeline::ops::exposure::EXPOSURE,
1.0,
);
render(&ctx, &mut pass, &graph, &source, SIZE);
assert_eq!(pass.colour_dispatches(), 2);
assert_eq!(pass.detail_dispatches(), 6);
}
#[test]
fn dragging_a_slider_recompiles_nothing_and_reallocates_nothing() {
// The two costs that are ruinous per frame and invisible in the output.
// Both are the same rule the rest of the crate follows: values ride in a
// uniform buffer, and textures are reallocated on resize rather than on
// change.
let Some(ctx) = ctx() else { return };
const SIZE: u32 = 48;
let source = step_edge(&ctx, SIZE);
let mut pass = AdjustPass::new(&ctx);
let mut graph = EditGraph::with_detail_probe();
graph.set_param(PROBE, RADIUS, 0.05);
render(&ctx, &mut pass, &graph, &source, SIZE);
let pipelines = pass.cached_detail_pipelines();
let allocations = pass.detail_allocations();
assert_eq!(pipelines, 2, "one per pass of the separable blur");
assert_eq!(allocations, 2, "the colour result, and one hand-off");
for radius in [0.06, 0.07, 0.08, 0.09] {
graph.set_param(PROBE, RADIUS, radius);
render(&ctx, &mut pass, &graph, &source, SIZE);
}
assert_eq!(
pass.cached_detail_pipelines(),
pipelines,
"a radius is a uniform, not a shader"
);
assert_eq!(
pass.detail_allocations(),
allocations,
"a steady viewport must allocate nothing"
);
// A resize is the one thing that legitimately reallocates.
render(&ctx, &mut pass, &graph, &source, SIZE / 2);
assert!(pass.detail_allocations() > allocations);
}
#[test]
fn a_proxy_and_an_export_agree_about_where_the_effect_lands() {
// TRACES: FR-DSP-1 — the subtle one, and the reason `RenderScale` exists.
//
// The same edit, rendered at two resolutions. A radius stored as a
// fraction of the shorter edge must produce a transition covering the same
// *proportion* of the frame at both, or a sharpening tuned on screen is a
// different sharpening in the exported file.
//
// The tolerance is a pixel's worth at the smaller size, because the kernel
// is an integer count and 6.25% of 64 pixels is not 6.25% of 128. That
// rounding is the whole of the error, and it is bounded by half a render
// pixel by construction.
let Some(ctx) = ctx() else { return };
const SOURCE: u32 = 128;
let source = step_edge(&ctx, SOURCE);
let mut graph = EditGraph::with_detail_probe();
graph.set_param(PROBE, RADIUS, 0.0625);
let spread = |out: u32| -> f32 {
let mut pass = AdjustPass::new(&ctx);
let pixels = render(&ctx, &mut pass, &graph, &source, out);
let line = row(&pixels, out, out / 2);
// Where the ramp starts and ends, in fractions of the frame.
let first = line.iter().position(|&v| v > 4).expect("a ramp") as f32;
let last = line.iter().rposition(|&v| v < 251).expect("a ramp") as f32;
(last - first) / out as f32
};
let proxy = spread(SOURCE / 2);
let export = spread(SOURCE);
assert!(
(proxy - export).abs() < 0.03,
"the effect covers {proxy:.3} of the proxy and {export:.3} of the \
export; a radius tuned on screen must land in the file"
);
// And it is a real transition in both, not two flat images agreeing.
assert!(proxy > 0.08 && export > 0.08, "{proxy:.3} / {export:.3}");
}
#[test]
fn the_two_halves_of_one_composition_must_be_dispatched_together() {
// The failure this guards is a bad one to debug: a shader composed to hand
// on linear working values, bound to an rgba8 storage texture. wgpu
// rejects it, but the message is about a bind group, a long way from the
// caller that composed one half of an edit and rendered the other.
let Some(ctx) = ctx() else { return };
const SIZE: u32 = 32;
let source = step_edge(&ctx, SIZE);
let mut pass = AdjustPass::new(&ctx);
let mut graph = EditGraph::with_detail_probe();
graph.set_param(PROBE, RADIUS, 0.0625);
let shader = graph.compose_for(ColourSpace::Srgb);
assert_eq!(shader.output_mode, OutputMode::LinearWorking);
let err = pass
.render_masked(&source, &shader, SIZE, SIZE, None)
.expect_err("a linear-working shader has no business in the plain path");
assert!(
format!("{err}").contains("render_detailed"),
"the error should name the way out: {err}"
);
}
#[test]
fn an_empty_chain_falls_through_to_the_ordinary_render() {
// A caller that always goes through `render_detailed` — which is what a
// frontend will do, since it does not want to branch on whether the user
// has sharpening on — must pay exactly nothing for the edits that have
// none.
let Some(ctx) = ctx() else { return };
const SIZE: u32 = 32;
let source = step_edge(&ctx, SIZE);
let graph = EditGraph::default_chain();
let mut pass = AdjustPass::new(&ctx);
let shader = graph.compose_for(ColourSpace::Srgb);
let scale = graph.render_scale((SIZE, SIZE), (SIZE, SIZE));
let detail = graph.compose_detail_for(scale, ColourSpace::Srgb);
assert!(detail.is_empty());
pass.render_detailed(&source, &shader, SIZE, SIZE, None, &detail, 0)
.expect("render");
assert_eq!(pass.colour_dispatches(), 1);
assert_eq!(pass.detail_dispatches(), 0);
assert_eq!(pass.detail_allocations(), 0);
}
+8
View File
@@ -23,3 +23,11 @@ log.workspace = true
# its author wrote them.
[build-dependencies]
serde_norway.workspace = true
[features]
default = []
# The detail stage's test consumer — a separable box blur that is not a develop
# operation and never appears in the panel. See `src/detail/probe.rs` for why an
# abstraction with no consumers gets a fake one, and `dr-gpu`'s dev-dependency
# for who turns this on. Off by default, so a shipping build does not contain it.
detail-probe = []
+72
View File
@@ -216,6 +216,78 @@ carries lens-profile coefficients that are not parameters. `distortion` and
`aberration` are `Warp`s rather than operations: they rewrite coordinates
before sampling rather than transforming a colour after it.
## Nodes that read their neighbours
`wgsl:` above is handed `c`, a colour, and no coordinate. That is what makes
the fused dispatch possible and it is also a wall: sharpening, noise reduction,
clarity, texture, dehaze and spot removal are all defined by what the
*neighbouring* pixels are doing, and none of them can be written as a function
of `c` at any price.
They go in the **detail stage**, which runs after the fused pass, in linear
light, at render resolution, before the output transform — see
[`../src/detail.rs`](../src/detail.rs) for why each of those is a decision
rather than a convenience. A node of this kind:
- is declared here with `rust:`, like any other hand-written node, because a
kernel is not four facts and stretching this schema to cover one would
produce a worse language than Rust;
- implements `Operation` as usual — descriptor, parameters, `is_active` — so
the panel, the sidecar, the history and the presets all work unchanged;
- returns `Affects::Detail` from `affects()` and `Some(self)` from `detail()`;
- implements `DetailStage::passes`, returning one `DetailPass` per dispatch,
each with a WGSL body, its uniforms, and **its kernel radius in render
pixels**, which the tile scheduler needs and nothing can infer.
The `order:` still belongs here, and still orders the node — among the other
detail nodes. Detail runs as a group after every point operation, so an `order:`
that interleaves one with exposure would be a lie the chain cannot tell.
The one thing to get right is the **unit of a radius**. Never store pixels: a
length is either a fraction of the frame's shorter edge (`RenderScale::
frame_fraction` — clarity, texture, the unit a mask feather already uses) or a
count of source pixels (`RenderScale::source_pixels` — capture sharpening,
luminance NR). `passes()` is given the scale and converts on the CPU. A radius
in raw pixels is a different photograph on screen and in the exported file.
## What is not a node, and why
Three things act on every pixel and are deliberately not in this directory:
the as-shot white balance, the camera matrix, and the **base curve**
(FR-DEV-3e). They are emitted by [`../src/operation.rs`](../src/operation.rs)
into the composed shader's fixed preamble, around the block of nodes.
The test is not "does it transform a colour" — all three do. It is **whose
decision is it**. A node is something a photographer chose: it has parameters,
it moves off a neutral, it lands in the sidecar, it can be undone. These three
are properties of the *file*, at the same standing as the masked-photosite crop
(FR-RAW-3) and the stored orientation (FR-DEV-3h). Nobody chose the sensor's
green sensitivity or the body's rendering; they are what reading the file
correctly means.
Making the base curve a node would have said the opposite in four places at
once. It would have appeared in the develop panel as a control, so an
unprofiled body would show a slider that does nothing. Its values would have
gone into the sidecar, and sidecars are shared between devices and bodies
(FR-NC-9) — one camera's rendering would follow an edit onto another camera's
file. Its neutral would have had to be "the identity", so a profiled body would
open reporting itself modified. And there is no seam through which a node could
learn which camera took the frame: the profile arrives on the decoded image,
travels through `DemosaicedImage` beside the matrix it belongs with, and is
written into the uniform block by the same three lines in `dr-gpu` — which is
exactly the path the matrix already took, because it is exactly the same kind
of thing.
What it *does* share with the tone curve node is the spline. The composer asks
`ToneCurve` for its `curve_span`/`curve_eval` helpers rather than emitting a
second copy, so a profile author placing a control point and a photographer
dragging one mean the same thing by it.
The order still reads correctly from this directory: the base curve runs after
every node in the chain and before the conversion out of camera space. That is
the same reasoning `exposure` records under `placement:` — corrections to
capture are only meaningful on linear values, so the rendering goes last.
## Errors
The build script reports failures by naming the key you got wrong, and exits
+897
View File
@@ -0,0 +1,897 @@
//! Neighbourhood operations — the ones that must read a pixel they are not
//! writing.
//!
//! # Why this exists at all
//!
//! Every operation in [`crate::operation`] contributes a fragment taking a
//! `vec3<f32>` and returning one. That contract is what makes the fused
//! dispatch possible, and it is also an absolute wall: a fragment is handed a
//! colour, not a coordinate, so it cannot look left. Sharpening, noise
//! reduction, clarity, texture, dehaze and spot removal are all defined by
//! what the *neighbours* are doing, and none of them can be written as a point
//! function of `c` at any price.
//!
//! FR-DEV-3 asks for all six and FR-DEV-8 for spot removal. So the fused pass
//! is not the whole pipeline; it is the *point-operation* stage of it, and
//! this module is the stage that follows.
//!
//! # Where it sits, and why there
//!
//! ```text
//! demosaiced source (camera space, full sensor resolution)
//! |
//! | <- framing prologue: output pixel -> source position
//! v
//! +------------------------------------------+
//! | the fused point-operation pass | one dispatch
//! | white balance, exposure, tone, colour |
//! | the mask layers |
//! | camera RGB -> linear sRGB |
//! +------------------------------------------+
//! | rgba16float, linear, **unclipped**, at render resolution
//! v
//! +------------------------------------------+
//! | the detail stage - this module | one dispatch per pass
//! | sharpen, NR, clarity, texture, spots |
//! +------------------------------------------+
//! | the last pass applies the output transform
//! v
//! rgba8unorm display or export texture
//! ```
//!
//! Four things about that position are decisions rather than convenience, and
//! each of them could defensibly have gone the other way.
//!
//! **After tone, not before.** Sharpening before a tone curve and sharpening
//! after it are different pictures, not the same picture computed two ways: an
//! S-curve steepens the mid-tones, so a halo introduced before it is amplified
//! by whatever slope the curve happens to have at that luminance, and the
//! amount that looked right stops looking right the moment the curve moves.
//! After the curve, the amount the user chose is the amount they see, and it
//! survives every later change to tone. This is also what ARCH §5.2 draws:
//! texture, clarity, spot removal and sharpen/NR sit below the tone curve and
//! the colour mixer.
//!
//! **In linear light, after the camera matrix.** The fused pass works in
//! *camera* space, because white balance and exposure are physically
//! meaningful there and nowhere else. A detail pass is the opposite case: it
//! wants a luminance, and camera RGB has no luminance — the three channels are
//! whatever the CFA's dyes passed, and weighting them 0.2126/0.7152/0.0722
//! would be numerology. So the split is taken *after* the `cam_to_srgb`
//! multiply, where the working space is linear sRGB and a luminance is a
//! luminance.
//!
//! **Before the output transform, and before the clip.** FR-DEV-2 allows
//! exactly one quantisation, at the display or export stage. A detail pass
//! reading an 8-bit display-encoded texture and writing another one would
//! quantise twice and do its arithmetic in a space where a difference of one
//! code value means different things at different brightnesses — which is how
//! sharpening ends up with visible banding in a sky. The intermediate is
//! therefore `rgba16float` and holds linear values that have **not** been
//! clamped to `0..=1`: a recovered highlight is still above one at this point,
//! and clipping it before the sharpener sees it would put a hard edge exactly
//! where the sharpener is most visible. The last detail pass performs the
//! primaries conversion, the clip and the encode, so the single quantisation
//! stays single.
//!
//! **After framing, at render resolution.** The alternative — running detail
//! on the demosaiced source before the framing prologue — is superficially
//! attractive, because a radius in sensor pixels would then mean exactly what
//! it says. It is unaffordable: the source is the full sensor, so a detail
//! pass there costs 24 MP of work for a 2 MP preview and FR-DSP-1 stops being
//! true. Running at render resolution instead makes the cost proportional to
//! what is on screen, and pushes the whole difficulty into one place — the
//! scale — which [`RenderScale`] exists to make explicit rather than implicit.
//!
//! # What this stage deliberately cannot do
//!
//! **There is no per-mask detail.** A mask layer's chain is fused into the
//! point-operation pass; the detail stage runs once, afterwards, over the
//! whole frame. Local sharpening is therefore not expressible here, and
//! [`crate::mask::MaskLayer::active_ops`] filters detail operations out rather
//! than emitting a block that would silently do nothing. Making it possible
//! means giving a detail pass the mask array and a layer index, which is a
//! change to this module's shader preamble and not to its shape — but it is
//! not done, and a caller should not assume it.
//!
//! # Adding a neighbourhood operation
//!
//! Declare it in `ops/<id>.yaml` with `rust:`, exactly as the tone curve does
//! — the schema in `ops/README.md` describes a point function, and stretching
//! it to cover kernels would be a worse language than Rust aimed at one
//! caller. Then implement [`crate::Operation`] as usual for the parameters,
//! descriptor and sidecar, and additionally:
//!
//! ```ignore
//! impl Operation for Sharpen {
//! fn affects(&self) -> Affects { Affects::Detail }
//! fn detail(&self) -> Option<&dyn DetailStage> { Some(self) }
//! fn wgsl_body(&self) -> String { String::new() } // never called
//! // ... descriptor, set_param, param, is_active exactly as usual
//! }
//!
//! impl DetailStage for Sharpen {
//! fn passes(&self, scale: RenderScale) -> Vec<DetailPass> { /* ... */ }
//! }
//! ```
//!
//! Everything else arrives unchanged and for free: the develop panel builds
//! its controls from the descriptor, the sidecar persists the parameters, the
//! history and the presets carry them, and an operation at its defaults
//! contributes no pass at all.
use std::fmt::Write as _;
use dr_types::ColourSpace;
use crate::operation::{Helper, Operation, Uniform};
/// Floats the generated detail uniform block always carries, before an
/// operation's own.
///
/// One `vec4`, which is also the smallest a WGSL uniform struct can be and
/// stay aligned. See [`compose_detail`] for what the lanes hold.
pub const DETAIL_BASE_UNIFORM_FIELDS: usize = 4;
/// TRACES: FR-DSP-1
/// The relationship between the resolution an edit is being **rendered** at
/// and the resolution it will eventually be **exported** at.
///
/// # The problem this type is the answer to
///
/// A point operation is scale-free. Exposure is a multiply, and multiplying by
/// two is multiplying by two whether the frame is 2 000 pixels wide or 24 000.
/// Every operation in the fused pass has this property, which is why nothing
/// in the pipeline has needed to know its own resolution until now.
///
/// A neighbourhood operation has no such luck. "Sharpen with a radius of one
/// pixel" is a statement about a specific grid, and the develop view is not
/// rendering on that grid — FR-DSP-1 has it rendering at whatever the viewport
/// needs, which for a 60 MP frame in a 2 000 px panel is one render pixel per
/// nine source pixels. Tune a radius there, export at full size, and the
/// exported file is sharpened at a ninth of the strength the photographer
/// chose. That is not a rounding difference; it is a different photograph.
///
/// # The rule
///
/// **A length is stored normalised and converted here.** Never store pixels in
/// an edit. This is not a new idea in this codebase — [`crate::mask`] already
/// does it, storing every feather and morphology radius as a fraction of the
/// frame's shorter edge and multiplying up in `dr-gpu` at whatever size the
/// mask is being rasterised at (see `MaskLayer::feather`, and
/// `field_short_edge` in `dr-gpu`'s mask pass). A detail operation follows the
/// same rule through [`Self::frame_fraction`] and gets the same guarantee: the
/// effect covers the same *proportion* of the picture at every size, so what
/// was tuned on screen is what lands in the file.
///
/// # Two units, because there are two kinds of length
///
/// The mask rule is not quite enough on its own, because detail operations
/// split into two families that mean different things by "radius":
///
/// - **Compositional** — clarity, texture, dehaze. The radius is a fraction of
/// the picture, tens of pixels at any size, and [`Self::frame_fraction`] is
/// exactly right. These preview faithfully at any scale.
///
/// - **Acutance** — capture sharpening, luminance noise reduction. The radius
/// is a property of the *sensor*: it is about the lens's circle of confusion
/// and the demosaic's interpolation, both measured in source pixels and
/// neither of which cares how large the viewport is.
/// [`Self::source_pixels`] converts one of those into render pixels.
///
/// # The honest limit
///
/// For the second family the conversion runs out. At a one-ninth proxy a
/// 1.0-source-pixel radius is 0.11 render pixels, and there is no kernel that
/// represents a ninth of a pixel — the information the sharpener would act on
/// was thrown away by the downscale before the pass ever ran. No arrangement
/// of this stage recovers it, which is why every editor that has shipped tells
/// the photographer to judge sharpening at 1:1, and why Lightroom's detail
/// panel contains a 1:1 loupe rather than a scaled preview.
///
/// [`Self::resolves`] reports that condition instead of hiding it, so an
/// operation can fade itself out and an interface can say "zoom to 100% to
/// judge this" — which is the truth, and better than a preview that lies.
/// Zooming is enough: the framing's view rect shrinks while the render target
/// keeps its size, so [`Self::ratio`] climbs back to 1.0 at 1:1 and the
/// preview becomes exact, with no separate full-resolution path to maintain.
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct RenderScale {
render: (u32, u32),
full: (u32, u32),
}
impl RenderScale {
/// `render` is the size being rendered now; `full` is the size the same
/// framed region would have at source resolution.
///
/// Both describe *the region being looked at*, not the whole photograph —
/// so a crop and a zoom are already accounted for by the time they arrive.
/// [`crate::EditGraph::render_scale`] works both out from the framing, and
/// is what a caller should normally use.
pub fn new(render: (u32, u32), full: (u32, u32)) -> Self {
Self {
render: (render.0.max(1), render.1.max(1)),
full: (full.0.max(1), full.1.max(1)),
}
}
/// A scale that is already at source resolution — an export, or a 1:1
/// view. [`Self::ratio`] is 1.0 and nothing is approximated.
pub fn full(render: (u32, u32)) -> Self {
Self::new(render, render)
}
pub fn render_size(&self) -> (u32, u32) {
self.render
}
pub fn full_size(&self) -> (u32, u32) {
self.full
}
/// Render pixels per source pixel. 1.0 at export, below 1.0 on a proxy.
///
/// Averaged over the two axes rather than taken from one. They agree to
/// within a pixel by construction — both sizes describe the same rectangle
/// — but each is separately rounded to an integer, and taking the mean
/// stops a narrow viewport disagreeing with itself.
pub fn ratio(&self) -> f32 {
let x = self.render.0 as f32 / self.full.0 as f32;
let y = self.render.1 as f32 / self.full.1 as f32;
(x + y) * 0.5
}
/// Whether this render is smaller than the file it stands for.
pub fn is_proxy(&self) -> bool {
self.ratio() < 0.999
}
/// A length stated in **source pixels**, in render pixels.
///
/// For the acutance family — sharpening, luminance NR — whose radius is a
/// property of the sensor rather than of the composition.
pub fn source_pixels(&self, radius: f32) -> f32 {
radius * self.ratio()
}
/// A length stated as a **fraction of the frame's shorter edge**, in
/// render pixels.
///
/// For the compositional family — clarity, texture, dehaze — and the same
/// unit `dr-gpu`'s mask rasteriser already converts feathers in. An edit
/// stored this way is resolution-independent by construction.
pub fn frame_fraction(&self, fraction: f32) -> f32 {
fraction * self.render.0.min(self.render.1) as f32
}
/// Whether a radius stated in source pixels survives this render.
///
/// False means the effect is smaller than a pixel here and whatever is
/// drawn is a guess. Report it; do not paper over it — see the type's
/// documentation for why there is nothing better to do.
pub fn resolves(&self, radius_in_source_pixels: f32) -> bool {
self.source_pixels(radius_in_source_pixels) >= 1.0
}
}
/// One dispatch of a neighbourhood operation.
///
/// An operation returns as many of these as it needs. A separable Gaussian is
/// two — horizontal then vertical — and gets the ping-pong between them for
/// free; an unsharp mask wanting its blur held alongside the original would be
/// more, and is the case this shape exists to leave room for.
#[derive(Debug, Clone, PartialEq)]
pub struct DetailPass {
/// A short name, used to label the GPU pass and to make a shader
/// compilation failure say which of an operation's passes broke.
pub label: &'static str,
/// The furthest this pass reads from the pixel it writes, in **render**
/// pixels.
///
/// Declared rather than inferred from the WGSL, because nothing can infer
/// it from the WGSL: the offsets are computed at runtime from uniforms.
/// It is the halo a tile has to be grown by before this pass can be
/// computed tile-wise (ARCH §5.3), and it is the reason a detail operation
/// is not simply "some more shader code" — the scheduler has to know how
/// far the dependency reaches before it can schedule anything at all.
///
/// An understated radius shows as a seam at every tile boundary, which is
/// the kind of artefact that looks like a driver bug. State it honestly.
pub radius: u32,
/// The WGSL body.
///
/// Reads and writes `c`, a `vec3<f32>` of **linear sRGB**, pre-loaded with
/// this pixel's own value. Also in scope:
///
/// - `coord: vec2<i32>` — this pixel.
/// - `tap(coord, offset) -> vec3<f32>` — a neighbour, clamped to the edge
/// of the image, which is what makes a kernel at the border average the
/// pixels that exist rather than fade into black.
/// - `render_dims: vec2<f32>` and `render_scale: f32` — the size being
/// rendered and [`RenderScale::ratio`], for the rare pass that needs
/// them in the shader. Prefer computing lengths on the CPU in
/// [`DetailStage::passes`], where the units are named methods rather
/// than an untyped float.
///
/// Uniforms are addressed by the bare names declared in [`Self::uniforms`],
/// exactly as a fused fragment addresses its own; the composer rewrites
/// them to their prefixed struct fields.
///
/// **Values are not clipped.** A recovered highlight arrives above 1.0 and
/// an out-of-gamut colour can arrive below 0.0. That is deliberate — see
/// the module documentation — and a kernel that assumes `0..=1` will
/// produce dark rings around specular highlights.
pub wgsl: String,
/// Uniform values this pass's body reads.
pub uniforms: Vec<Uniform>,
}
/// TRACES: FR-DEV-3 | FR-DEV-8
/// An operation that reads pixels other than the one it is writing.
///
/// Implemented *alongside* [`Operation`], never instead of it: the parameters,
/// the descriptor, the panel controls and the sidecar all come from the
/// `Operation` half, and only the execution differs. An operation that
/// implements this must also return [`crate::Affects::Detail`] from
/// `affects()` and `Some(self)` from `Operation::detail()` — the three are
/// checked against each other by a test in [`crate::operation`], because an
/// operation that forgot one of them would be dropped from both stages and
/// simply not happen, with no error anywhere.
pub trait DetailStage: Send + Sync {
/// The passes to run, in order, at this resolution.
///
/// Called per render, so the operation sees the scale it is actually being
/// asked to draw at and converts its own lengths here — in Rust, where
/// [`RenderScale`]'s two conversions are named after the two units, rather
/// than in WGSL where both would be a bare `f32`.
///
/// Returning an empty vector means "nothing to do at this scale", which is
/// the honest answer for an acutance operation on a heavy proxy. It is
/// **not** how an operation says it is neutral: that is `is_active()`, and
/// an inactive operation is never asked.
fn passes(&self, scale: RenderScale) -> Vec<DetailPass>;
}
/// One compile-ready detail pass: complete WGSL and the uniform block for it.
#[derive(Debug, Clone, PartialEq)]
pub struct ComposedDetailPass {
/// `<op id>/<pass label>`, for GPU labels and error messages.
pub label: String,
/// Complete, compilable WGSL.
pub source: String,
/// Uniform values in the order the generated struct declares them.
pub uniforms: Vec<f32>,
/// See [`DetailPass::radius`].
pub radius: u32,
/// Whether this pass writes the display/export texture rather than another
/// linear intermediate.
///
/// True for exactly the last pass in the chain, which carries the output
/// transform — the primaries conversion, the clip and the encode that the
/// fused pass performs when there is no detail stage at all. Folding them
/// into the last pass rather than adding a resolve dispatch keeps the cost
/// of the stage at one dispatch per pass, not one plus one.
pub writes_output: bool,
/// Identifies this pass's *structure*, for the pipeline cache. Covers the
/// generated source, not the uniform values — so moving a slider uploads a
/// buffer and reuses the compiled pipeline, exactly as the fused pass does.
pub structure_hash: u64,
}
/// The detail stage of one edit, at one resolution.
#[derive(Debug, Clone, Default, PartialEq)]
pub struct ComposedDetail {
pub passes: Vec<ComposedDetailPass>,
}
impl ComposedDetail {
/// Whether the edit has no detail stage — the common case, and the one
/// that must cost nothing.
pub fn is_empty(&self) -> bool {
self.passes.is_empty()
}
pub fn len(&self) -> usize {
self.passes.len()
}
/// The widest halo any pass needs, in render pixels (ARCH §5.3).
pub fn radius(&self) -> u32 {
self.passes.iter().map(|p| p.radius).max().unwrap_or(0)
}
}
/// TRACES: FR-DEV-3 | FR-DSP-1
/// Generate the detail stage for a set of operations at one resolution.
///
/// Operations that declare no [`DetailStage`], or that are at their neutral
/// settings, contribute nothing — the same rule the fused composer follows, so
/// an edit with no sharpening produces an empty chain and `dr-gpu` runs the
/// single dispatch it always did.
///
/// `output` is the space the **last** pass encodes into, and it is a parameter
/// for the same reason it is a parameter to [`crate::compose_with_framing`]: a
/// screen render and a Display P3 export are the same edit and different
/// shaders, and neither is more authoritative than the other.
///
/// # The generated uniform block
///
/// A fixed `vec4` first, then the pass's own scalars, prefixed with the
/// operation id so that a pass never has to know what else is in the block.
/// The lanes of the leading `vec4` are, in order: render width, render height,
/// [`RenderScale::ratio`], and the pass's index within its operation. The
/// first three reach the body as `render_dims` and `render_scale`; the fourth
/// is there because a two-pass operation emitting one body for both directions
/// is a reasonable thing to want, and would otherwise need a uniform of its
/// own purely to say which half it is in.
pub fn compose_detail(
ops: &[Box<dyn Operation>],
scale: RenderScale,
output: ColourSpace,
) -> ComposedDetail {
// Every pass of every active detail operation, flattened, carrying the
// operation it came from for the uniform prefix and the helper set.
let mut planned: Vec<(&'static str, &'static [Helper], DetailPass, usize)> = Vec::new();
for op in ops {
if !op.is_active() {
continue;
}
let Some(stage) = op.detail() else {
continue;
};
let id = op.descriptor().id.0;
for (index, pass) in stage.passes(scale).into_iter().enumerate() {
planned.push((id, op.helpers(), pass, index));
}
}
let last = planned.len().saturating_sub(1);
let passes = planned
.into_iter()
.enumerate()
.map(|(position, (id, helpers, pass, index))| {
compose_one(id, helpers, &pass, index, scale, output, position == last)
})
.collect();
ComposedDetail { passes }
}
#[allow(clippy::too_many_arguments)]
fn compose_one(
id: &str,
helpers: &[Helper],
pass: &DetailPass,
index: usize,
scale: RenderScale,
output: ColourSpace,
writes_output: bool,
) -> ComposedDetailPass {
let prefix = format!("{}_{index}", crate::operation::sanitise(id));
let mut uniform_fields = String::from(
" // x, y: the size being rendered. z: render pixels per source\n\
\x20 // pixel — 1.0 at export, less on a proxy (FR-DSP-1). w: which\n\
\x20 // pass of this operation this is.\n\
\x20 detail_base: vec4<f32>,\n",
);
let (rw, rh) = scale.render_size();
let mut uniform_values = vec![rw as f32, rh as f32, scale.ratio(), index as f32];
debug_assert_eq!(uniform_values.len(), DETAIL_BASE_UNIFORM_FIELDS);
if !pass.uniforms.is_empty() {
let _ = writeln!(uniform_fields, " // {id}/{}", pass.label);
}
for u in &pass.uniforms {
let _ = writeln!(uniform_fields, " {prefix}_{}: f32,", u.name);
uniform_values.push(u.value);
}
// A uniform struct whose size is not a multiple of 16 is rejected by the
// WGSL uniform address space rules — the same padding the fused composer
// applies, for the same reason.
let pad = (4 - (uniform_values.len() % 4)) % 4;
for i in 0..pad {
let _ = writeln!(uniform_fields, " _pad{i}: f32,");
uniform_values.push(0.0);
}
let mut body = pass.wgsl.clone();
for u in &pass.uniforms {
body = crate::operation::rewrite_uniform(&body, u.name, &format!("u.{prefix}_{}", u.name));
}
let mut helper_src = String::new();
let mut seen: Vec<&str> = Vec::new();
for h in helpers {
if seen.contains(&h.name) {
continue;
}
seen.push(h.name);
let _ = writeln!(helper_src, "{}\n", h.source.trim_end());
}
// The storage format and the tail are the *only* difference between an
// intermediate pass and the final one. Everything above — the taps, the
// uniforms, the body — is identical, which is what lets an operation write
// one kernel without knowing whether it happens to be last in the chain.
let (store_format, tail) = if writes_output {
(
"rgba8unorm",
format!(
"{} // Clip to the output gamut and encode. The one quantisation\n\
\x20 // the pipeline performs (FR-DEV-2), and it is here rather than\n\
\x20 // in the fused pass because this is now the last thing to run.\n\
\x20 c = clamp(c, vec3<f32>(0.0), vec3<f32>(1.0));\n\
\x20 textureStore(output, coord, vec4<f32>(encode_output(c), 1.0));",
crate::operation::primaries_conversion(output)
),
)
} else {
(
"rgba16float",
" // Another linear intermediate: no clip and no encode, because\n\
\x20 // the pass after this one still has to read real values.\n\
\x20 textureStore(output, coord, vec4<f32>(c, 1.0));"
.to_string(),
)
};
let encode_fn = if writes_output {
crate::operation::encode_output_fn(output)
} else {
String::new()
};
let label = format!("{id}/{}", pass.label);
let indented = body
.lines()
.map(|l| format!(" {l}"))
.collect::<Vec<_>>()
.join("\n");
let source = format!(
"// GENERATED — do not edit.
//
// Detail pass `{label}` — a neighbourhood operation, which is why it is a
// dispatch of its own rather than a block in the fused shader: it reads pixels
// it is not writing, and the fused contract hands a fragment a colour with no
// way back to a coordinate.
//
// In: linear sRGB, scene-referred, **unclipped**, at render resolution.
// Out: {}
struct Params {{
{uniform_fields}}}
@group(0) @binding(0) var source: texture_2d<f32>;
@group(0) @binding(1) var<uniform> u: Params;
@group(0) @binding(2) var output: texture_storage_2d<{store_format}, write>;
// A neighbour, clamped to the edge of the image.
//
// Clamped rather than zero-filled: a kernel straddling the border must average
// the pixels that exist. Returning zero there darkens every edge by a band the
// width of the radius, which reads as a vignette nobody asked for and is the
// classic way a first convolution goes wrong.
fn tap(coord: vec2<i32>, offset: vec2<i32>) -> vec3<f32> {{
let last = vec2<i32>(textureDimensions(source)) - vec2<i32>(1);
return textureLoad(source, clamp(coord + offset, vec2<i32>(0), last), 0).rgb;
}}
{helper_src}{encode_fn}
@compute @workgroup_size(8, 8, 1)
fn main(@builtin(global_invocation_id) gid: vec3<u32>) {{
let dims = textureDimensions(output);
if (gid.x >= dims.x || gid.y >= dims.y) {{
return;
}}
let coord = vec2<i32>(gid.xy);
// What this render is, relative to the export it has to match.
let render_dims = u.detail_base.xy;
let render_scale = u.detail_base.z;
var c = tap(coord, vec2<i32>(0));
{{
{indented}
}}
{tail}
}}
",
if writes_output {
"display-encoded, in the output space."
} else {
"linear sRGB, for the next pass."
},
);
let structure_hash = crate::operation::hash_source(&source);
ComposedDetailPass {
label,
source,
uniforms: uniform_values,
radius: pass.radius,
writes_output,
structure_hash,
}
}
// Compiled for this crate's own tests as well as for the feature, so that
// `cargo test -p dr-pipeline` exercises the seam whether or not anybody
// downstream remembered to turn the feature on. A test that quietly does not
// exist is worse than no test, because the absence looks like a pass.
#[cfg(any(test, feature = "detail-probe"))]
pub mod probe;
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn a_full_render_approximates_nothing() {
let s = RenderScale::full((2000, 1300));
assert!(!s.is_proxy());
assert!((s.ratio() - 1.0).abs() < 1e-6);
// A one-pixel sharpening radius is one pixel at export, always.
assert!((s.source_pixels(1.0) - 1.0).abs() < 1e-6);
assert!(s.resolves(1.0));
}
#[test]
fn a_proxy_shrinks_a_source_length_and_says_so() {
// A 6000px frame in a 1500px panel: four source pixels per render
// pixel, so a 1px capture-sharpening radius is a quarter of a render
// pixel and cannot be drawn. This is the case the whole type exists
// for, and the answer has to be "no", not a plausible-looking number.
let s = RenderScale::new((1500, 1000), (6000, 4000));
assert!(s.is_proxy());
assert!((s.ratio() - 0.25).abs() < 1e-6);
assert!(!s.resolves(1.0), "a quarter of a pixel is not a kernel");
assert!(s.resolves(4.0), "four source pixels do survive");
}
#[test]
fn a_frame_fraction_is_the_same_proportion_at_every_size() {
// The mask rule, restated as a test: 1% of the shorter edge is 1% of
// the shorter edge whether the render is a thumbnail or an export.
// This is what makes a clarity radius tuned on screen correct in the
// exported file.
let proxy = RenderScale::new((2000, 1333), (6000, 4000));
let export = RenderScale::full((6000, 4000));
let as_fraction = |s: &RenderScale| {
let (w, h) = s.render_size();
s.frame_fraction(0.01) / w.min(h) as f32
};
assert!((as_fraction(&proxy) - as_fraction(&export)).abs() < 1e-6);
// And in absolute terms it really does scale with the render.
assert!((proxy.frame_fraction(0.01) - 13.33).abs() < 0.5);
assert!((export.frame_fraction(0.01) - 40.0).abs() < 0.5);
}
#[test]
fn zooming_to_one_to_one_makes_the_preview_exact() {
// The reason there is no separate full-resolution preview path: the
// framing's view rect shrinks while the render target keeps its size,
// so the ratio climbs back to 1.0 and a sharpening radius means
// exactly what it will mean in the file.
let fit = RenderScale::new((2000, 1333), (6000, 4000));
let one_to_one = RenderScale::new((2000, 1333), (2000, 1333));
assert!(!fit.resolves(1.0));
assert!(one_to_one.resolves(1.0));
}
use crate::detail::probe::BoxBlur;
use crate::operation::{compose_full, OutputMode};
fn with_blur(radius: f32) -> Vec<Box<dyn Operation>> {
let mut ops = crate::ops::chain();
ops.push(Box::new(BoxBlur::with_radius(radius)));
ops
}
fn fused(ops: &[Box<dyn Operation>]) -> crate::ComposedShader {
compose_full(
ops,
&crate::Framing::new(),
dr_types::ColourSpace::Srgb,
&crate::mask::MaskStack::new(),
)
}
#[test]
fn a_detail_operation_contributes_nothing_to_the_fused_shader() {
// The seam itself: a neighbourhood operation is in the graph, is
// active, and yet emits no block in the single dispatch — because it
// physically cannot, and asking it for one would produce an empty
// block that reads as an operation doing nothing.
let shader = fused(&with_blur(0.05));
assert!(
!shader.source.contains("---- detail_probe ----"),
"a detail operation must not appear as a fused fragment"
);
assert!(
!shader.source.contains("detail_probe_radius"),
"nor should it occupy a slot in the fused uniform block"
);
}
#[test]
fn an_active_detail_operation_makes_the_fused_pass_hand_on_linear_values() {
// The other half of the same decision. With no detail stage the fused
// pass encodes and quantises, exactly as it always has; with one, it
// stops at linear working values and the detail chain finishes the
// job. Getting this wrong is not a subtle wrong colour — it is a
// storage format that does not match the texture bound to it.
let neutral = fused(&with_blur(0.0));
assert_eq!(neutral.output_mode, OutputMode::Encoded);
assert!(neutral.source.contains("texture_storage_2d<rgba8unorm"));
assert!(neutral.source.contains("encode_output"));
let blurring = fused(&with_blur(0.05));
assert_eq!(blurring.output_mode, OutputMode::LinearWorking);
assert!(blurring.source.contains("texture_storage_2d<rgba16float"));
assert!(
!blurring.source.contains("fn encode_output"),
"the fused pass must not encode when a detail stage follows: \
FR-DEV-2 allows exactly one quantisation"
);
assert!(
!blurring.source.contains("clamp(c, vec3<f32>(0.0)"),
"nor clip, or the sharpener sees a hard edge at every highlight"
);
}
#[test]
fn a_neutral_detail_operation_costs_the_edit_nothing() {
// The rule the whole pipeline is built on, extended to this stage: an
// operation at its defaults contributes no code, no uniform and no
// dispatch. An unedited photograph must not pay for a sharpener it is
// not using.
let ops = with_blur(0.0);
let composed = compose_detail(
&ops,
RenderScale::full((512, 512)),
dr_types::ColourSpace::Srgb,
);
assert!(composed.is_empty());
assert_eq!(fused(&ops).output_mode, OutputMode::Encoded);
}
#[test]
fn a_separable_blur_becomes_two_passes_and_only_the_last_encodes() {
// The multi-pass case, which is the one the ping-pong exists for. The
// first pass writes a linear intermediate and the second writes the
// display texture — so the output transform happens exactly once, at
// the end, wherever the end happens to be.
let ops = with_blur(0.05);
let composed = compose_detail(
&ops,
RenderScale::full((512, 512)),
dr_types::ColourSpace::Srgb,
);
assert_eq!(composed.len(), 2);
let first = &composed.passes[0];
let last = &composed.passes[1];
assert_eq!(first.label, "detail_probe/horizontal");
assert_eq!(last.label, "detail_probe/vertical");
assert!(!first.writes_output);
assert!(first.source.contains("texture_storage_2d<rgba16float"));
assert!(!first.source.contains("fn encode_output"));
assert!(last.writes_output);
assert!(last.source.contains("texture_storage_2d<rgba8unorm"));
assert!(last.source.contains("fn encode_output"));
// Two passes of one operation are two shaders, so they must not share
// a pipeline-cache entry — the classic way a second pass silently runs
// the first one's code.
assert_ne!(first.structure_hash, last.structure_hash);
}
#[test]
fn a_pass_addresses_its_uniforms_without_knowing_the_block() {
// The same contract the fused composer offers: a body writes `radius`
// and the composer rewrites it to a prefixed struct field, so two
// operations may both call a uniform `radius` and neither has to know.
let ops = with_blur(0.05);
let composed = compose_detail(
&ops,
RenderScale::full((512, 512)),
dr_types::ColourSpace::Srgb,
);
let src = &composed.passes[0].source;
assert!(src.contains("detail_probe_0_radius: f32,"));
assert!(src.contains("let r = i32(u.detail_probe_0_radius);"));
// And the second pass carries its own index, so its uniforms cannot be
// uploaded into the first pass's slots.
assert!(composed.passes[1]
.source
.contains("detail_probe_1_radius: f32,"));
}
#[test]
fn every_pass_declares_a_uniform_block_the_gpu_will_accept() {
// A uniform struct whose size is not a multiple of 16 is rejected
// outright by the WGSL uniform address space rules, and the failure
// arrives as a shader compilation error against generated source.
let ops = with_blur(0.05);
for pass in compose_detail(
&ops,
RenderScale::full((512, 512)),
dr_types::ColourSpace::Srgb,
)
.passes
{
assert_eq!(pass.uniforms.len() % 4, 0, "{}", pass.label);
assert!(pass.uniforms.iter().all(|v| v.is_finite()));
// The base block is first and fixed, so a pass never addresses a
// slot by number and the render size is always in the same place.
assert_eq!(pass.uniforms[0], 512.0);
assert_eq!(pass.uniforms[1], 512.0);
}
}
#[test]
fn the_declared_radius_is_the_halo_a_tile_would_need() {
// ARCH §5.3 schedules tiles, and a tile cannot be computed without
// knowing how far outside itself the pass reads. Nothing can infer it
// from the WGSL — the offsets are computed from uniforms at runtime —
// so the operation states it, and this is the assertion that it states
// the truth rather than zero.
let ops = with_blur(0.05);
let scale = RenderScale::full((400, 400));
let composed = compose_detail(&ops, scale, dr_types::ColourSpace::Srgb);
let expected = BoxBlur::with_radius(0.05).kernel(scale);
assert_eq!(expected, 20, "5% of a 400px edge");
assert_eq!(composed.radius(), expected);
assert!(composed.passes.iter().all(|p| p.radius == expected));
}
#[test]
fn a_normalised_radius_is_the_same_effect_at_every_resolution() {
// FR-DSP-1's hard part, at the level this crate can test it: the same
// edit composed at two sizes produces kernels in the same *proportion*
// to the frame. `dr-gpu`'s `proxy_and_export_agree` checks the pixels
// that fall out of it.
let ops = with_blur(0.04);
let sizes = [(200u32, 200u32), (800, 800), (2400, 2400)];
let fractions: Vec<f32> = sizes
.iter()
.map(|&(w, h)| {
let scale = RenderScale::full((w, h));
let composed = compose_detail(&ops, scale, dr_types::ColourSpace::Srgb);
composed.radius() as f32 / w.min(h) as f32
})
.collect();
for f in &fractions {
assert!(
(f - 0.04).abs() < 0.005,
"the kernel drifted from the declared fraction: {fractions:?}"
);
}
}
#[test]
fn an_edit_with_no_detail_operation_composes_no_passes() {
// The property that keeps the cost of this stage at zero for the
// overwhelmingly common edit: no sharpening means no chain, which
// means `dr-gpu` runs the single fused dispatch it always did.
let ops = crate::ops::chain();
let composed = compose_detail(
&ops,
RenderScale::full((64, 64)),
dr_types::ColourSpace::Srgb,
);
assert!(composed.is_empty());
assert_eq!(composed.radius(), 0);
}
}
+175
View File
@@ -0,0 +1,175 @@
//! A separable box blur, for testing the detail stage. **Not a develop
//! operation.**
//!
//! # Why an abstraction gets a fake consumer
//!
//! The detail stage was written before any of the operations it exists for —
//! sharpening, noise reduction, clarity, spot removal are each their own piece
//! of work — and an abstraction with no consumer is a guess. Nothing would
//! have proved that the WGSL it generates compiles, that the ping-pong hands
//! pass two what pass one wrote, that the last pass really does encode, or
//! that a radius stated in one unit survives the trip from a proxy to an
//! export.
//!
//! So the stage has exactly one consumer, and it lives here, behind the
//! `detail-probe` feature. It is deliberately *not* declared in `ops/`: it has
//! no `order:`, it is not in [`crate::ops::chain`], it never reaches
//! [`crate::EditGraph::capabilities`], and so it cannot appear in the develop
//! panel or in a sidecar. A shipping build does not contain it.
//!
//! # Why a box blur specifically
//!
//! Because its answer is known in closed form. A box blur of radius *r* over a
//! step edge produces a ramp exactly `2r + 1` pixels wide with a known value
//! at every step, so a test can assert *pixels*, not "something changed". A
//! Gaussian would need a tolerance chosen to hide whatever the implementation
//! actually did.
//!
//! And because it is **separable**, which is the property the two-pass case
//! was built for: a horizontal pass then a vertical one is mathematically a 2D
//! box average, so if the ping-pong is wired backwards or a pass reads its own
//! output the result is visibly not a box blur rather than subtly wrong.
use crate::descriptor::{
Attribute, LocalizedKey, OpDescriptor, OpId, ParamDescriptor, ParamId, ParamKind, Scale, Unit,
};
use crate::detail::{DetailPass, DetailStage, RenderScale};
use crate::operation::{Affects, Operation, Uniform};
static DESCRIPTOR: OpDescriptor = OpDescriptor {
id: OpId("detail_probe"),
label: LocalizedKey("op.detail_probe"),
params: &[ParamDescriptor {
id: ParamId("radius"),
label: LocalizedKey("param.detail_probe.radius"),
// A fraction of the frame's shorter edge, which is the unit
// `RenderScale::frame_fraction` converts and the unit a mask feather
// is already stored in. Stating it in pixels is the mistake this
// whole stage is arranged to make impossible.
kind: ParamKind::Scalar {
min: 0.0,
max: 0.25,
scale: Scale::Linear,
unit: Unit::None,
precision: 4,
},
default: 0.0,
facet: None,
}],
attributes: &[Attribute::Detail],
};
/// A separable box blur whose radius is a fraction of the frame's shorter edge.
#[derive(Debug, Clone, Copy, Default)]
pub struct BoxBlur {
radius: f32,
}
impl BoxBlur {
pub fn new() -> Self {
Self::default()
}
/// Set the radius directly, in fractions of the shorter edge.
pub fn with_radius(radius: f32) -> Self {
Self { radius }
}
/// The kernel radius this blur would use at `scale`, in render pixels.
///
/// Exposed so a test can state the expected ramp width without repeating
/// the rounding rule — a test that recomputed it would agree with a bug.
pub fn kernel(&self, scale: RenderScale) -> u32 {
scale.frame_fraction(self.radius).round().max(0.0) as u32
}
}
impl Operation for BoxBlur {
fn descriptor(&self) -> &'static OpDescriptor {
&DESCRIPTOR
}
fn set_param(&mut self, _id: ParamId, value: f32) {
self.radius = value;
}
fn param(&self, _id: ParamId) -> f32 {
self.radius
}
fn is_active(&self) -> bool {
self.radius > 0.0
}
/// Never called. A detail operation contributes no fused fragment, and
/// [`crate::operation::compose_full`] filters it out before asking.
fn wgsl_body(&self) -> String {
String::new()
}
fn uniforms(&self) -> Vec<Uniform> {
Vec::new()
}
fn affects(&self) -> Affects {
Affects::Detail
}
fn detail(&self) -> Option<&dyn DetailStage> {
Some(self)
}
}
impl DetailStage for BoxBlur {
fn passes(&self, scale: RenderScale) -> Vec<DetailPass> {
let r = self.kernel(scale);
// A radius that rounded to nothing is not "blur by zero" — it is an
// effect this render is too small to show. Emitting a pass that
// averages one pixel would burn a dispatch to copy the image.
if r == 0 {
return Vec::new();
}
// Two passes, one per axis. The horizontal one reads the fused pass's
// output and the vertical one reads the horizontal one's, which is the
// whole point: if the ping-pong were wired to hand the second pass the
// original again, the result would be a horizontal smear rather than a
// box, and the test asserting a symmetric ramp would say so.
["x", "y"]
.iter()
.enumerate()
.map(|(axis, _)| DetailPass {
label: if axis == 0 { "horizontal" } else { "vertical" },
radius: r,
uniforms: vec![
Uniform {
name: "radius",
value: r as f32,
},
Uniform {
name: "step_x",
value: if axis == 0 { 1.0 } else { 0.0 },
},
Uniform {
name: "step_y",
value: if axis == 0 { 0.0 } else { 1.0 },
},
],
wgsl: "// One axis of a separable box average.
//
// `tap` clamps at the border, so a kernel hanging off the edge averages the
// edge pixel repeatedly rather than averaging in black — which keeps a
// constant image constant, the cheapest property to check and the first one
// a broken border rule breaks.
let r = i32(radius);
let step = vec2<i32>(i32(step_x), i32(step_y));
var sum = vec3<f32>(0.0);
for (var i = -r; i <= r; i = i + 1) {
sum = sum + tap(coord, step * i);
}
c = sum / f32(2 * r + 1);"
.to_string(),
})
.collect()
}
}
+362
View File
@@ -118,6 +118,28 @@ impl EditGraph {
}
}
/// TRACES: FR-DEV-3
/// The default chain with the detail stage's test consumer appended.
///
/// **Not a shipping path.** `detail_probe` is a separable box blur that
/// exists so the neighbourhood stage has something to run (see
/// [`crate::detail::probe`]); it is not declared in `ops/`, has no place
/// in the pipeline order, and is compiled only for tests and behind the
/// `detail-probe` feature.
///
/// It is a constructor rather than a fixture inside one test module
/// because `dr-gpu` needs the same graph: proving the stage works means
/// dispatching it, and dispatching it means composing both halves of the
/// shader from one graph exactly as the interface will.
#[cfg(any(test, feature = "detail-probe"))]
pub fn with_detail_probe() -> Self {
let mut graph = Self::default_chain();
graph
.ops
.push(Box::new(crate::detail::probe::BoxBlur::new()));
graph
}
/// The local adjustment stack.
pub fn masks(&self) -> &MaskStack {
&self.masks
@@ -342,6 +364,130 @@ impl EditGraph {
pub fn compose_for(&self, output: dr_types::ColourSpace) -> ComposedShader {
compose_full(&self.ops, &self.framing, output, &self.masks)
}
/// TRACES: FR-DSP-1
/// How this render relates to the file it stands for.
///
/// `source` is the demosaiced image's size and `render` the size being
/// drawn now. The result describes *the region on screen*, with the crop
/// and the zoom already folded in: cropping to half the frame while the
/// viewport stays the same size genuinely does show twice the detail, and
/// zooming to 1:1 genuinely does make the preview exact. Both fall out of
/// the arithmetic rather than needing a special case.
///
/// Only the detail stage needs this. Every point operation is scale-free
/// — a multiply is a multiply at any resolution — which is why nothing in
/// the pipeline had to know its own size until a kernel arrived.
pub fn render_scale(&self, source: (u32, u32), render: (u32, u32)) -> crate::detail::RenderScale {
let (fw, fh) = self.framing.output_size(source.0, source.1);
let view = self.framing.view();
// The *viewed* part of the framed image, at source resolution. Zoom
// shrinks the view rect while the render target keeps its size, so
// this is what shrinks and the ratio is what climbs.
let full = (
((fw as f32 * view.width).round() as u32).max(1),
((fh as f32 * view.height).round() as u32).max(1),
);
crate::detail::RenderScale::new(render, full)
}
/// TRACES: FR-DEV-3 | FR-DSP-1
/// Generate the detail stage for this edit at one resolution, to sRGB.
///
/// Empty for every edit with no active neighbourhood operation, which is
/// almost all of them — and in that case [`Self::compose`] emits the
/// single encoded dispatch it always has.
pub fn compose_detail(&self, scale: crate::detail::RenderScale) -> crate::detail::ComposedDetail {
self.compose_detail_for(scale, dr_types::ColourSpace::Srgb)
}
/// TRACES: FR-EXP-2
/// The detail stage, encoded into a chosen output space.
///
/// The space belongs here as well as on [`Self::compose_for`] because when
/// a detail stage exists it is the *last* pass that performs the output
/// transform — the fused pass stops at linear working values. Composing
/// the two halves for different spaces would encode the edit twice, or
/// not at all.
pub fn compose_detail_for(
&self,
scale: crate::detail::RenderScale,
output: dr_types::ColourSpace,
) -> crate::detail::ComposedDetail {
crate::detail::compose_detail(&self.ops, scale, output)
}
/// TRACES: FR-DEV-3d
/// The per-stage cache keys for the current edit.
///
/// See [`crate::Invalidation`] for what the keys mean and what may be
/// cached against them. In short: geometry covers the framing, colour
/// covers every fused operation and every mask layer, and detail covers
/// the neighbourhood operations — so moving one slider moves exactly one
/// key, and a consumer can tell which stages it has to redo.
pub fn invalidation(&self) -> crate::Invalidation {
use crate::operation::{hash_bytes, hash_op, mix, Affects, FNV_OFFSET};
// Geometry: the framing. Its own structure key covers the shape of the
// coordinate map; the parameters cover the magnitudes, which the
// structure key deliberately omits because they do not recompile a
// shader. Both matter to a cached *result*, so both are here.
let mut geometry = mix(FNV_OFFSET, self.framing.structure_key());
for p in self.framing.descriptor().params {
geometry = hash_bytes(geometry, p.id.0.as_bytes());
geometry = mix(
geometry,
u64::from(crate::operation::canonical_bits(self.framing.param(p.id))),
);
}
// The view rect is not a parameter and not in the structure key — it
// is not an edit (see `Framing::view`). It is still an input to every
// rendered pixel, so a cache that ignored it would show the wrong part
// of the photograph after a scroll.
let view = self.framing.view();
for v in [view.x, view.y, view.width, view.height] {
geometry = mix(geometry, u64::from(crate::operation::canonical_bits(v)));
}
let mut colour = FNV_OFFSET;
let mut detail = FNV_OFFSET;
for op in &self.ops {
let target = if op.affects() == Affects::Detail {
&mut detail
} else {
&mut colour
};
*target = hash_op(*target, op.as_ref());
}
// The mask layers belong to the colour stage: their chains are fused
// into the same dispatch, and a layer's *shape* decides which pixels
// that dispatch treats differently. Both halves are folded in.
for layer in self.masks.layers() {
colour = hash_bytes(colour, layer.id.as_bytes());
// The source through its `Debug`, deliberately. A gradient's
// centre, a region's id list and a subject's signature are all
// part of where the layer applies, and matching on the variants
// here would be a second copy of `MaskSource`'s shape that falls
// out of step the first time a variant gains a field — silently,
// and showing as a mask that stops updating. `Debug` cannot fall
// out of step, because it is derived from the definition itself.
colour = hash_bytes(colour, format!("{:?}", layer.source).as_bytes());
colour = mix(colour, u64::from(layer.enabled));
colour = mix(colour, u64::from(layer.invert));
colour = hash_bytes(colour, layer.falloff.name().as_bytes());
for v in [layer.opacity, layer.feather, layer.morph_radius] {
colour = mix(colour, u64::from(crate::operation::canonical_bits(v)));
}
for (op_id, param_id, value) in layer.params() {
colour = hash_bytes(colour, op_id.as_bytes());
colour = hash_bytes(colour, param_id.as_bytes());
colour = mix(colour, u64::from(crate::operation::canonical_bits(value)));
}
}
crate::Invalidation::new(geometry, colour, detail)
}
}
impl Default for EditGraph {
@@ -645,4 +791,220 @@ mod tests {
);
assert_ne!(before, g.compose().structure_hash);
}
// ---- invalidation scoping (FR-DEV-3d) --------------------------------
use crate::descriptor::OpId;
use crate::operation::Affects;
const PROBE: OpId = OpId("detail_probe");
const PROBE_RADIUS: ParamId = ParamId("radius");
#[test]
fn moving_a_detail_parameter_leaves_every_earlier_stage_alone() {
// FR-DEV-3d's headline, and the thing `Affects::Detail` was added to
// make true: dragging a sharpening slider must not re-run the
// demosaic, the framing, or the fused colour pass. The demosaic is not
// a key here at all — no parameter in this graph can reach it — and
// the other two must come out unchanged.
let mut g = EditGraph::with_detail_probe();
let before = g.invalidation();
g.set_param(PROBE, PROBE_RADIUS, 0.05);
let after = g.invalidation();
assert_ne!(
before.of(Affects::Detail),
after.of(Affects::Detail),
"the detail stage's own key must move"
);
assert_eq!(
before.through(Affects::Colour),
after.through(Affects::Colour),
"the fused colour pass's result is still valid, so its cached \
linear intermediate must be reusable"
);
assert_eq!(
before.through(Affects::Geometry),
after.through(Affects::Geometry)
);
}
#[test]
fn moving_a_colour_parameter_leaves_geometry_alone_and_redoes_detail() {
// The other direction, and the half that is easy to get wrong by
// wishing. Exposure does not touch the framing — FR-DEV-3d says so in
// as many words. It *does* invalidate the detail stage's output,
// because the detail stage reads what the colour pass wrote, and
// pretending otherwise would show a sharpened version of the previous
// exposure. The stage's own parameters are still untouched, which is
// what `of` reports and `through` does not.
let mut g = EditGraph::with_detail_probe();
g.set_param(PROBE, PROBE_RADIUS, 0.05);
let before = g.invalidation();
g.set_param(exposure::ID, exposure::EXPOSURE, 1.0);
let after = g.invalidation();
assert_eq!(
before.through(Affects::Geometry),
after.through(Affects::Geometry),
"adjusting exposure shall not re-tile geometry (FR-DEV-3d)"
);
assert_ne!(before.of(Affects::Colour), after.of(Affects::Colour));
assert_eq!(
before.of(Affects::Detail),
after.of(Affects::Detail),
"the sharpening settings did not change"
);
assert_ne!(
before.through(Affects::Detail),
after.through(Affects::Detail),
"but its input did, so its cached output is stale"
);
}
#[test]
fn cropping_invalidates_everything_downstream_of_it() {
// Geometry is upstream of both other stages: it decides which source
// pixel every colour is read from, and — because the detail stage runs
// at render resolution — how many render pixels a kernel spans.
let mut g = EditGraph::with_detail_probe();
g.set_param(PROBE, PROBE_RADIUS, 0.05);
let before = g.invalidation();
g.set_crop(CropRect {
x: 0.1,
y: 0.1,
width: 0.5,
height: 0.5,
});
let after = g.invalidation();
assert_ne!(before.of(Affects::Geometry), after.of(Affects::Geometry));
assert_ne!(
before.through(Affects::Colour),
after.through(Affects::Colour)
);
assert_ne!(
before.through(Affects::Detail),
after.through(Affects::Detail)
);
// Scoped, though: neither later stage's *own* settings moved.
assert_eq!(before.of(Affects::Colour), after.of(Affects::Colour));
assert_eq!(before.of(Affects::Detail), after.of(Affects::Detail));
}
#[test]
fn scrolling_the_view_invalidates_the_render_without_being_an_edit() {
// The view rect is not an edit — it is excluded from the sidecar, the
// structure hash and `is_active` — but it absolutely is an input to
// every pixel. A key that ignored it would leave the previous part of
// the photograph on screen after a pan, which looks like a repaint bug
// and is a cache bug.
let mut g = EditGraph::default_chain();
let before = g.invalidation();
g.framing_mut().set_view(CropRect {
x: 0.25,
y: 0.25,
width: 0.5,
height: 0.5,
});
assert_ne!(
before.of(Affects::Geometry),
g.invalidation().of(Affects::Geometry)
);
}
#[test]
fn returning_a_slider_to_where_it_was_returns_the_key() {
// A cache key that drifted with the *path* rather than the state would
// never hit after an undo, which is the moment it is most wanted.
let mut g = EditGraph::with_detail_probe();
let origin = g.invalidation();
g.set_param(exposure::ID, exposure::EXPOSURE, 1.5);
g.set_param(PROBE, PROBE_RADIUS, 0.05);
assert_ne!(origin, g.invalidation());
g.set_param(exposure::ID, exposure::EXPOSURE, 0.0);
g.set_param(PROBE, PROBE_RADIUS, 0.0);
assert_eq!(origin, g.invalidation(), "the state is what is hashed");
}
#[test]
fn a_local_adjustment_belongs_to_the_colour_stage() {
// A mask layer's chain is fused into the same dispatch as the global
// one, so changing it is a colour change and nothing more. Its
// *shape* counts too: which pixels the dispatch treats differently is
// as much a part of the result as by how much.
use crate::mask::{MaskLayer, MaskSource};
let mut g = EditGraph::with_detail_probe();
let before = g.invalidation();
g.masks_mut().push(MaskLayer::new(
"l1",
MaskSource::Linear {
centre: (0.5, 0.5),
angle: 0.0,
width: 0.2,
},
));
let with_layer = g.invalidation();
assert_ne!(before.of(Affects::Colour), with_layer.of(Affects::Colour));
assert_eq!(
before.of(Affects::Geometry),
with_layer.of(Affects::Geometry)
);
assert_eq!(before.of(Affects::Detail), with_layer.of(Affects::Detail));
// Moving the gradient is a different mask, so a different result.
if let Some(layer) = g.masks_mut().get_mut("l1") {
layer.source = MaskSource::Linear {
centre: (0.2, 0.7),
angle: 0.4,
width: 0.2,
};
}
assert_ne!(
with_layer.of(Affects::Colour),
g.invalidation().of(Affects::Colour)
);
}
#[test]
fn the_render_scale_folds_in_the_crop_and_the_zoom() {
// What a detail operation is handed, and the reason it does not need
// to know that a crop or a zoom happened: both arrive already folded
// into one ratio.
let mut g = EditGraph::default_chain();
let source = (6000, 4000);
// Fit: a 1500px panel over a 6000px frame is a quarter scale.
let fit = g.render_scale(source, (1500, 1000));
assert!((fit.ratio() - 0.25).abs() < 1e-3);
// Zoomed to 1:1 — the view rect shrinks to what the panel can hold,
// the render target keeps its size, and the preview becomes exact.
g.framing_mut().set_view(CropRect {
x: 0.25,
y: 0.25,
width: 0.25,
height: 0.25,
});
let one_to_one = g.render_scale(source, (1500, 1000));
assert!((one_to_one.ratio() - 1.0).abs() < 1e-3);
assert!(one_to_one.resolves(1.0));
// A crop shows fewer source pixels in the same panel, which is more
// render pixels each — a sharpening radius genuinely does grow.
let mut cropped = EditGraph::default_chain();
cropped.set_crop(CropRect {
x: 0.25,
y: 0.25,
width: 0.5,
height: 0.5,
});
let after = cropped.render_scale(source, (1500, 1000));
assert!(after.ratio() > fit.ratio());
}
}
+6 -2
View File
@@ -32,6 +32,7 @@
//! data neither would be physically meaningful (ARCH §5.2).
pub mod descriptor;
pub mod detail;
pub mod framing;
pub mod graph;
pub mod history;
@@ -46,13 +47,16 @@ pub use descriptor::{
Attribute, Facet, LocalizedKey, OpDescriptor, OpId, ParamDescriptor, ParamId, ParamKind, Presentation,
Scale, Unit, WidgetDemand, WidgetKind,
};
pub use detail::{
compose_detail, ComposedDetail, ComposedDetailPass, DetailPass, DetailStage, RenderScale,
};
pub use framing::{CropRect, Framing};
pub use graph::{EditGraph, OpCapability, ParamCapability};
pub use history::{Edit, History};
pub use lens::{compose_warps, ComposedWarp, Warp};
pub use operation::{
compose, compose_with_framing, Affects, ComposedShader, Helper, Operation, Uniform,
RESERVED_UNIFORM_FIELDS,
compose, compose_with_framing, Affects, ComposedShader, Helper, Invalidation, Operation,
OutputMode, Uniform, BASE_CURVE_POINTS, BASE_CURVE_UNIFORM_OFFSET, RESERVED_UNIFORM_FIELDS,
};
pub use preset::{Preset, Scope};
pub use sidecar::{Sidecar, Version};
+15 -1
View File
@@ -775,8 +775,22 @@ impl MaskLayer {
}
}
/// The operations in this layer's chain that reach the shader.
///
/// Neighbourhood operations are excluded, and not as an oversight. A
/// layer's chain is *fused into the point-operation pass* and multiplied
/// by the mask afterwards; the detail stage runs once, over the whole
/// frame, after that pass has finished (see [`crate::detail`]). There is
/// nowhere in that arrangement for a sharpening confined to one mask to
/// happen, so a detail operation in a layer would contribute an empty
/// block, count towards [`Self::is_active`], and cost a mask rasterisation
/// to change nothing. Dropping it here means the layer reports honestly
/// that it has no adjustment rather than appearing to have one.
pub fn active_ops(&self) -> impl Iterator<Item = &dyn Operation> {
self.ops.iter().map(|o| o.as_ref()).filter(|o| o.is_active())
self.ops
.iter()
.map(|o| o.as_ref())
.filter(|o| o.is_active() && o.detail().is_none())
}
/// Whether this layer's region ids belong to a different segmentation.
+517 -18
View File
@@ -28,19 +28,182 @@ use crate::descriptor::{OpDescriptor, ParamId, Presentation};
use crate::framing::{Framing, FRAMING_UNIFORM_FIELDS};
use crate::mask::MaskStack;
/// TRACES: FR-DEV-3d
/// What an operation's parameters affect, for cache invalidation scoping.
///
/// Adjusting exposure must not invalidate the demosaic result; this is what
/// lets the tile cache reuse everything up to the first changed stage
/// (ARCH §5.3).
///
/// **The ordering is the pipeline order**, which is why this derives `Ord`
/// rather than merely `Eq`: geometry decides which source pixel a colour comes
/// from, the fused colour pass transforms it, and the detail stage reads the
/// neighbourhood the colour pass produced. A change at one stage invalidates
/// that stage and every later one, and nothing earlier — see [`Invalidation`],
/// which is where that rule is actually written down and tested.
#[derive(Debug, Clone, Copy, PartialEq, Eq, PartialOrd, Ord)]
pub enum Affects {
/// Per-pixel colour only. Everything in this milestone.
Colour,
/// Pixel positions — crop, rotate. Invalidates geometry-dependent caches.
/// Pixel positions — crop, rotate, straighten. The framing prologue, which
/// also decides the resolution everything downstream runs at.
Geometry,
/// Per-pixel colour. Every operation fused into the single adjust
/// dispatch, and every mask layer's chain.
Colour,
/// TRACES: FR-DEV-3d
/// A pixel's *neighbourhood* — sharpening, noise reduction, clarity,
/// texture, dehaze, spot removal.
///
/// The seam `docs/requirements.md` §3.3 designed and nothing cut until
/// [`crate::detail`] existed. It is a separate variant rather than a flavour
/// of `Colour` because it is a separate *dispatch*: a fragment in the fused
/// pass is handed a colour and has no way back to a coordinate, so a
/// kernel cannot be expressed there at any price.
///
/// What the distinction buys, concretely: the fused pass's result is held
/// in a linear intermediate, so dragging a sharpening slider re-runs the
/// detail dispatches and **not** the colour pass — which is exactly the
/// reuse FR-DEV-3d asks for, and it is asserted in `dr-gpu`'s
/// `detail_stage` tests rather than merely hoped for.
Detail,
}
/// TRACES: FR-DEV-3d
/// One cache key per pipeline stage, derived from the edit.
///
/// # The rule
///
/// A cached result for stage *S* stays valid while *S*'s own key and the keys
/// of every stage **before** it are unchanged. [`Self::of`] is the first half;
/// [`Self::through`] folds in the second and is what a cache should actually
/// store.
///
/// That reads as pedantry until it is applied, at which point it settles the
/// two questions FR-DEV-3d asks:
///
/// - **Changing a detail parameter must not re-run demosaic**, or the framing,
/// or the fused colour pass. It does not: `of(Detail)` moves and
/// `through(Colour)` does not, so the linear intermediate the colour pass
/// wrote is still good and only the detail dispatches run again.
///
/// - **Changing exposure must not re-run anything upstream of colour.** It
/// does not: `through(Geometry)` is untouched, so a tile cache keyed on it
/// survives, and the demosaiced texture — which no key here mentions at all
/// — is never in question.
///
/// It also settles what is *not* true, and the temptation is real: changing
/// exposure **does** re-run the detail passes, because the detail stage reads
/// what the colour pass wrote and that changed. There is no arrangement of
/// keys that avoids it while keeping sharpening after the tone curve, and
/// sharpening after the tone curve is the correct place (see
/// [`crate::detail`]). Anyone who wants exposure to leave the detail stage
/// alone is asking for detail to run *before* tone, which is a different
/// pipeline and a worse picture.
///
/// # Why the demosaic is not in here
///
/// Because no parameter in this graph can change it. The demosaiced texture is
/// a function of the file and the decode settings, both of which live outside
/// the edit graph; a caller keying a cache on it mixes in whatever names the
/// photograph — a `VersionId` — and these keys ride on top.
///
/// # Integer state only
///
/// Every value folded in here is a parameter: a slider position or a number
/// from a sidecar, never a float that came back from the GPU. That is what
/// ARCH §6.13 requires of a cache key, and it is why hashing the raw bit
/// patterns is sound rather than reckless. Negative zero is canonicalised on
/// the way in, because `-0.0 == 0.0` while their bit patterns differ, and a
/// slider that arrived at zero from below would otherwise invalidate a cache
/// that is perfectly valid.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub struct Invalidation {
geometry: u64,
colour: u64,
detail: u64,
}
impl Invalidation {
/// Build from the three per-stage hashes. [`crate::EditGraph::invalidation`]
/// is what computes them; this is public so a caller with its own notion
/// of a stage can construct one.
pub fn new(geometry: u64, colour: u64, detail: u64) -> Self {
Self {
geometry,
colour,
detail,
}
}
/// The key for `stage`'s own parameters, ignoring everything upstream.
///
/// Useful for asserting that a change was correctly *scoped* — that moving
/// a detail slider left the colour stage's parameters alone. Not a cache
/// key: a stage whose own parameters are unchanged still has to re-run if
/// its input changed, which is what [`Self::through`] is for.
pub fn of(&self, stage: Affects) -> u64 {
match stage {
Affects::Geometry => self.geometry,
Affects::Colour => self.colour,
Affects::Detail => self.detail,
}
}
/// The key for the **output** of `stage` — this stage and everything
/// upstream of it. What a cached texture should be keyed on.
pub fn through(&self, stage: Affects) -> u64 {
let mut h = FNV_OFFSET;
h = mix(h, self.geometry);
if stage >= Affects::Colour {
h = mix(h, self.colour);
}
if stage >= Affects::Detail {
h = mix(h, self.detail);
}
h
}
}
/// Fold one operation's identity and settings into a running hash.
///
/// Shared by the stage keys so that two stages cannot come to disagree about
/// what "this operation's state" means — which would show as a cache that is
/// occasionally, unreproducibly stale.
pub(crate) fn hash_op(h: u64, op: &dyn Operation) -> u64 {
let desc = op.descriptor();
let mut h = hash_bytes(h, desc.id.0.as_bytes());
for p in desc.params {
h = hash_bytes(h, p.id.0.as_bytes());
h = mix(h, u64::from(canonical_bits(op.param(p.id))));
}
h
}
/// A parameter's bits, with negative zero folded onto zero.
///
/// `-0.0 == 0.0` as far as every operation is concerned — a slider that
/// reached zero from below produces the same shader and the same picture — but
/// the two have different bit patterns. Hashing them apart would invalidate a
/// cache for a change that is not one.
pub(crate) fn canonical_bits(v: f32) -> u32 {
if v == 0.0 {
0
} else {
v.to_bits()
}
}
pub(crate) fn hash_bytes(mut h: u64, bytes: &[u8]) -> u64 {
for byte in bytes {
h ^= u64::from(*byte);
h = h.wrapping_mul(0x100_0000_01b3);
}
h
}
/// FNV-1a's offset basis. No dependency, and stable across runs and platforms,
/// which a cache key requires.
pub(crate) const FNV_OFFSET: u64 = 0xcbf2_9ce4_8422_2325;
/// A single scalar a fragment reads from the generated uniform block.
///
/// Operations declare uniforms by name and value; the composer assigns them
@@ -85,6 +248,10 @@ pub trait Operation: Send + Sync {
///
/// The fragment runs inside its own block, so locals need no unique
/// names.
///
/// Never called on an operation that declares a [`Self::detail`] stage —
/// a neighbourhood operation is a dispatch of its own and contributes
/// nothing to the fused shader, so it returns an empty string.
fn wgsl_body(&self) -> String;
/// Uniform values this operation's fragment reads.
@@ -95,6 +262,26 @@ pub trait Operation: Send + Sync {
Affects::Colour
}
/// TRACES: FR-DEV-3 | FR-DEV-8
/// This operation's neighbourhood stage, if it has one.
///
/// `None` — the default, and true of every operation that is a function of
/// one colour — means the operation is fused into the single adjust
/// dispatch in the ordinary way.
///
/// `Some` means the opposite: the operation reads pixels it is not
/// writing, cannot be a fragment in a fused shader, and runs as its own
/// dispatch or dispatches after the colour pass. See [`crate::detail`] for
/// where that sits and why, and for what a sharpening operation has to
/// write. An operation returning `Some` must also return
/// [`Affects::Detail`] from [`Self::affects`], which
/// `detail_operations_agree_with_themselves` checks — the two saying
/// different things would leave the operation in neither stage, silently
/// doing nothing.
fn detail(&self) -> Option<&dyn crate::detail::DetailStage> {
None
}
/// Any WGSL helper functions the fragment calls.
///
/// Emitted once per *distinct* function name even if several operations
@@ -128,6 +315,32 @@ pub struct Helper {
pub source: &'static str,
}
/// TRACES: FR-DEV-2 | FR-DEV-3d
/// What the fused pass writes, and therefore what has to be bound to it.
///
/// The fused shader ends one of two ways, and the difference is not cosmetic —
/// it decides the storage texture's format, so a shader composed for one and
/// dispatched against the other is a validation failure rather than a wrong
/// picture.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum OutputMode {
/// `rgba8unorm`, display-encoded in the composed output space. What the
/// pass has always written, and still writes for the overwhelmingly common
/// edit that has no detail stage: one dispatch, one read, one write.
Encoded,
/// `rgba16float`, linear sRGB, **unclipped**, scene-referred.
///
/// Emitted when the edit has an active neighbourhood operation. The detail
/// passes read this, and the last of them performs the output transform,
/// so the pipeline still quantises exactly once (FR-DEV-2) — it simply
/// happens two dispatches later.
///
/// Unclipped matters: a recovered highlight is above 1.0 here, and
/// clamping before a sharpener sees it would draw a hard edge at precisely
/// the luminance a sharpener is most visible at.
LinearWorking,
}
/// The result of composing a set of operations into one shader.
#[derive(Debug, Clone, PartialEq)]
pub struct ComposedShader {
@@ -139,13 +352,43 @@ pub struct ComposedShader {
/// not their values. Two edits differing only in slider positions share
/// a compiled pipeline and differ only in the uniform upload.
pub structure_hash: u64,
/// What this shader writes. See [`OutputMode`].
pub output_mode: OutputMode,
}
/// Fields the generated uniform struct always carries, before op uniforms.
///
/// WGSL requires a uniform struct to be non-empty and 16-byte aligned; these
/// are needed by every generated shader in any case.
const BASE_UNIFORM_FIELDS: usize = 16;
///
/// Twelve of the twenty-eight are the camera profile's base curve
/// ([`BASE_CURVE_UNIFORM_FIELDS`]); the rest are the matrix, the as-shot
/// balance and framing's own block.
const BASE_UNIFORM_FIELDS: usize = 16 + BASE_CURVE_UNIFORM_FIELDS;
/// TRACES: FR-DEV-3e
/// Slots the base curve occupies: five `(x, y)` points and an active flag.
///
/// Twelve rather than eleven so the block stays a whole number of `vec4`s,
/// which is what std140 requires of a uniform struct's members. The spare
/// float is left zero rather than repurposed — a uniform slot that means one
/// thing today and two things next year is how a shader comes to read a
/// highlight rolloff out of a crop rectangle.
const BASE_CURVE_UNIFORM_FIELDS: usize = 12;
/// TRACES: FR-DEV-3e
/// Where the base curve's slots begin in the generated uniform block.
///
/// Exported for the same reason [`RESERVED_UNIFORM_FIELDS`] is: `dr-gpu`
/// writes these by index, and an offset computed independently at both ends is
/// an offset that will eventually disagree with itself.
pub const BASE_CURVE_UNIFORM_OFFSET: usize = 16;
/// How many control points a base curve carries.
///
/// The same five the tone curve widget has, deliberately — see the helper
/// selection in [`compose_full`].
pub const BASE_CURVE_POINTS: usize = 5;
/// Where an operation's own uniforms begin in the generated block.
///
@@ -211,12 +454,33 @@ pub fn compose_full(
output: ColourSpace,
masks: &MaskStack,
) -> ComposedShader {
// Active *point* operations. A neighbourhood operation is filtered out
// here rather than asked for a fragment it cannot write: it reads pixels
// it is not writing, so it belongs to the detail stage that runs after
// this one (see `crate::detail`). Filtering on the declared stage rather
// than on `affects()` means the shader and the stage agree by
// construction — there is one place an operation says which it is.
let active: Vec<&dyn Operation> = ops
.iter()
.map(|o| o.as_ref())
.filter(|o| o.is_active())
.filter(|o| o.is_active() && o.detail().is_none())
.collect();
// Whether a detail stage follows. If one does, this pass stops short of
// the output transform and hands on a linear intermediate; the last detail
// pass finishes the job. Decided from the operations themselves rather
// than from a flag the caller sets, because a caller that got the flag
// wrong would produce a shader whose storage format does not match the
// texture bound to it.
let output_mode = if ops
.iter()
.any(|o| o.is_active() && o.detail().is_some())
{
OutputMode::LinearWorking
} else {
OutputMode::Encoded
};
let mut uniform_fields = String::new();
let mut uniform_values: Vec<f32> = Vec::new();
let mut body = String::new();
@@ -235,10 +499,44 @@ pub fn compose_full(
\x20 // `.w` is not padding: it flags a non-linear source (1.0 for a\n\
\x20 // gamma-encoded JPEG, 0.0 for demosaiced sensor data), which the\n\
\x20 // prologue reads to decide whether to linearise.\n\
\x20 as_shot_wb: vec4<f32>,\n",
\x20 as_shot_wb: vec4<f32>,\n\
\x20 // The camera profile's base curve (FR-DEV-3e): five points on a\n\
\x20 // monotone spline, packed as x0..x3, y0..y3, then (x4, y4, on).\n\
\x20 // `.z` of the last is the flag, not padding — it is 0 for a\n\
\x20 // body with no profile and for an already-rendered source.\n\
\x20 base_curve_x: vec4<f32>,\n\
\x20 base_curve_y: vec4<f32>,\n\
\x20 base_curve_last: vec4<f32>,\n",
);
uniform_values.resize(BASE_UNIFORM_FIELDS, 0.0);
// TRACES: FR-DEV-3e
// The spline the base curve is evaluated on is the *tone curve's* spline,
// reached through the trait rather than reimplemented here.
//
// Two reasons, and the second is the one that matters. The obvious one is
// that a shader carrying two `curve_eval`s would not compile, and the
// composer's helper de-duplication is what makes both stages able to ask
// for it. The real one is that a profile author placing a control point
// and a photographer dragging one must mean the same thing by it — down to
// the Fritsch-Carlson tangent limiting, which is what decides how a
// shoulder actually rolls off. Two implementations that agreed today would
// be two that could disagree later, and the disagreement would show up as
// a body whose profile renders subtly differently from the curve someone
// drew to match it.
//
// Emitted unconditionally, unlike an operation's helpers. The base curve
// is active for every RAW frame — an unprofiled body still gets the
// database's default rendering — so making the shader's shape depend on it
// would split the pipeline cache in two for no benefit. The uniform flag
// above turns it off for the cases that are genuinely already rendered,
// and a branch on a uniform is coherent across the whole dispatch.
for h in crate::ops::ToneCurve::new().helpers() {
if matches!(h.name, "curve_span" | "curve_eval") {
helpers.push(*h);
}
}
// Framing's block follows the base one at a fixed offset, for the same
// reason: the prologue is emitted whether or not any operation is active,
// so these slots cannot be positioned by the op loop below.
@@ -328,8 +626,40 @@ pub fn compose_full(
""
};
let to_output = primaries_conversion(output);
let encode_output = encode_output_fn(output);
// The tail, and it is the whole of the difference between the two output
// modes. Everything above — the prologue, the fragments, the mask layers,
// the camera matrix — is emitted identically either way, so an operation
// cannot tell whether a detail stage follows it and does not have to.
let (store_format, to_output, encode_output, store) = match output_mode {
OutputMode::Encoded => (
"rgba8unorm",
primaries_conversion(output),
encode_output_fn(output),
" // Clip to the output gamut and encode. The clip is last for the reason the
// matrix above is: a colour outside sRGB is still inside a wider space, and
// clipping before the conversion would throw it away for no one's benefit.
c = clamp(c, vec3<f32>(0.0), vec3<f32>(1.0));
textureStore(output, vec2<i32>(gid.xy), vec4<f32>(encode_output(c), 1.0));"
.to_string(),
),
OutputMode::LinearWorking => (
"rgba16float",
String::new(),
String::new(),
" // Stop here: a detail stage follows, and it needs linear values it
// can average. No primaries conversion, no clip and no encode — the
// last detail pass performs all three, so the pipeline still quantises
// exactly once (FR-DEV-2).
//
// Deliberately *not* clamped. A recovered highlight is above 1.0 at this
// point and an out-of-gamut colour can be below 0.0; clipping them here
// would put a hard edge into the very neighbourhood the next pass is
// about to convolve, which is how sharpeners come to draw dark rings
// around specular highlights.
textureStore(output, vec2<i32>(gid.xy), vec4<f32>(c, 1.0));"
.to_string(),
),
};
let source = format!(
"// GENERATED — do not edit.
@@ -344,7 +674,7 @@ struct Params {{
@group(0) @binding(0) var source: texture_2d<f32>;
@group(0) @binding(1) var<uniform> u: Params;
@group(0) @binding(2) var output: texture_storage_2d<rgba8unorm, write>;
@group(0) @binding(2) var output: texture_storage_2d<{store_format}, write>;
// The local adjustment masks, one array layer each, rasterised by a separate
// pass (ARCH §5.4). Declared unconditionally even when no layer is active, so
// that every generated shader shares one bind group layout — a layout that
@@ -419,6 +749,58 @@ fn main(@builtin(global_invocation_id) gid: vec3<u32>) {{
c = mix(c, neutral, clipped);
}}
{body}
// ==== camera profile: the base curve (FR-DEV-3e) ====
//
// Marked with `====` and not the `----` an operation block carries: this
// is not one, and the difference is what several tests count on to tell
// an edit apart from the reading of a file.
//
// The stage between demosaic and the working space that turns a correct
// exposure into a photograph. Sensor data is scene-referred and nearly
// linear; nothing anybody looks at is. Rendering it straight out is the
// dcraw default, and it is flat, dark through the midtones and clips its
// highlights instead of rolling them off.
//
// **In camera RGB, and after the adjustments**, which is a deliberate pair
// of choices:
//
// - Before the matrix, because that is where a base curve is defined and
// where every other converter applies one. The curve was tuned against
// this body's own primaries; moving it after the conversion would apply
// a Canon rendering to sRGB values and change what it does.
// - After exposure and the tonal operations, because those are corrections
// to *capture* and are only meaningful on linear values. A stop is a
// doubling; run exposure after a curve and it stops being one.
//
// Per channel rather than on luminance. It desaturates the extremes
// slightly, and that is the point — it is what makes a blown sky roll
// toward white rather than toward a saturated corner of the gamut, and it
// is what the camera's own JPEG does.
//
// The branch is on a uniform, so the whole dispatch takes the same path.
// It is off for a JPEG and any other already-rendered source, which must
// not be rendered twice, and for a body the profile database declines to
// offer any curve for at all.
if (u.base_curve_last.z > 0.5) {{
c = vec3<f32>(
curve_eval(
u.base_curve_x.x, u.base_curve_y.x, u.base_curve_x.y, u.base_curve_y.y,
u.base_curve_x.z, u.base_curve_y.z, u.base_curve_x.w, u.base_curve_y.w,
u.base_curve_last.x, u.base_curve_last.y, c.r,
),
curve_eval(
u.base_curve_x.x, u.base_curve_y.x, u.base_curve_x.y, u.base_curve_y.y,
u.base_curve_x.z, u.base_curve_y.z, u.base_curve_x.w, u.base_curve_y.w,
u.base_curve_last.x, u.base_curve_last.y, c.g,
),
curve_eval(
u.base_curve_x.x, u.base_curve_y.x, u.base_curve_x.y, u.base_curve_y.y,
u.base_curve_x.z, u.base_curve_y.z, u.base_curve_x.w, u.base_curve_y.w,
u.base_curve_last.x, u.base_curve_last.y, c.b,
),
);
}}
// Camera space -> linear sRGB. Applied after the adjustments so white
// balance and exposure act on sensor-native values, which is where they
// are physically meaningful.
@@ -430,11 +812,7 @@ fn main(@builtin(global_invocation_id) gid: vec3<u32>) {{
dot(u.cam_to_srgb_2.rgb, c),
);
{to_output}
// Clip to the output gamut and encode. The clip is last for the reason the
// matrix above is: a colour outside sRGB is still inside a wider space, and
// clipping before the conversion would throw it away for no one's benefit.
c = clamp(c, vec3<f32>(0.0), vec3<f32>(1.0));
textureStore(output, vec2<i32>(gid.xy), vec4<f32>(encode_output(c), 1.0));
{store}
}}
",
active.len()
@@ -473,6 +851,7 @@ fn main(@builtin(global_invocation_id) gid: vec3<u32>) {{
source,
uniforms: uniform_values,
structure_hash,
output_mode,
}
}
@@ -489,7 +868,7 @@ fn main(@builtin(global_invocation_id) gid: vec3<u32>) {{
/// value it already had. The identity is detected rather than special-cased by
/// name, so a space that happens to share sRGB's primaries would be spared
/// too.
fn primaries_conversion(output: ColourSpace) -> String {
pub(crate) fn primaries_conversion(output: ColourSpace) -> String {
let m = output.from_linear_srgb();
const IDENTITY: [f32; 9] = [1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0];
// A tolerance rather than equality: the matrix is an inverse multiplied by
@@ -528,7 +907,7 @@ fn primaries_conversion(output: ColourSpace) -> String {
///
/// Named `encode_output` whatever the space, so the call site at the end of
/// `main` does not have to know which one it got.
fn encode_output_fn(output: ColourSpace) -> String {
pub(crate) fn encode_output_fn(output: ColourSpace) -> String {
let body = match output.transfer() {
Transfer::Srgb => " let lo = c * 12.92;
let hi = 1.055 * pow(max(c, vec3<f32>(0.0031308)), vec3<f32>(1.0 / 2.4)) - 0.055;
@@ -626,7 +1005,7 @@ const BILINEAR_HELPER: &str = "fn sample_bilinear(uv: vec2<f32>, dims: vec2<u32>
";
/// Fold a value into a hash. FNV-1a's mixing step, over eight bytes.
fn mix(mut h: u64, value: u64) -> u64 {
pub(crate) fn mix(mut h: u64, value: u64) -> u64 {
for byte in value.to_le_bytes() {
h ^= u64::from(byte);
h = h.wrapping_mul(0x100_0000_01b3);
@@ -644,7 +1023,7 @@ fn mix(mut h: u64, value: u64) -> u64 {
/// Still integer state hashed on the CPU, as ARCH §6.13 requires of a cache
/// key: the text is generated from parameters that are neutral or not, never
/// from a rendered float.
fn hash_source(source: &str) -> u64 {
pub(crate) fn hash_source(source: &str) -> u64 {
// FNV-1a: no dependency, stable across runs and platforms, which the
// shader cache key requires.
let mut h: u64 = 0xcbf2_9ce4_8422_2325;
@@ -960,6 +1339,84 @@ mod tests {
assert!(op < matrix, "the camera matrix must come after operations");
}
#[test]
fn the_base_curve_runs_after_the_operations_and_before_the_camera_matrix() {
// TRACES: FR-DEV-3e
// Both halves matter and for different reasons.
//
// After the operations: exposure and the tonal controls are
// corrections to capture, and they are only meaningful on linear
// values. A stop is a doubling; run exposure after a curve and it is
// not one any more, and every slider in the panel starts lying about
// what it does.
//
// Before the matrix: the curve was tuned against this body's own
// primaries. Applied after the conversion it would be a Canon
// rendering acting on sRGB values, which is a different curve.
let ops = vec![fake(&DESC_A, 2.0, false)];
let source = compose(&ops).source;
let op = source.find("---- op_a ----").expect("op present");
let curve = source
.find("if (u.base_curve_last.z > 0.5)")
.expect("base curve applied");
let matrix = source.find("u.cam_to_srgb_0").expect("matrix applied");
assert!(op < curve, "the base curve must come after the operations");
assert!(curve < matrix, "and before the camera matrix");
}
#[test]
fn the_base_curve_reaches_a_shader_with_no_operations_at_all() {
// TRACES: FR-DEV-3e
// The same property as as-shot white balance, and for the same reason:
// it is part of interpreting the file, not part of the edit. An
// unedited RAW must open looking like a photograph rather than like a
// scan of one.
let shader = compose(&[]);
assert!(shader.source.contains("u.base_curve_x"));
assert!(
shader.source.contains("fn curve_eval("),
"the spline it is evaluated on must be emitted too"
);
}
#[test]
fn the_base_curve_and_the_tone_curve_share_one_spline() {
// TRACES: FR-DEV-3e
// Two `curve_eval`s in one shader would not compile — but the reason
// the helper is *shared* rather than merely renamed is that a profile
// author placing a control point and a photographer dragging one must
// mean the same thing by it, down to the tangent limiting that decides
// how a shoulder rolls off.
let mut curve = crate::ops::ToneCurve::new();
curve.set_param(crate::ops::curve::P2_Y, 0.7);
assert!(curve.is_active(), "the fixture must actually reach the shader");
let source = compose(&[Box::new(curve)]).source;
assert_eq!(
source.matches("fn curve_eval(").count(),
1,
"the spline must be declared exactly once"
);
assert_eq!(source.matches("fn curve_span(").count(), 1);
}
#[test]
fn the_base_curve_owns_the_slots_dr_gpu_writes() {
// TRACES: FR-DEV-3e
// `dr-gpu` fills these by index. The offset is exported rather than
// recomputed there, and this asserts the exported number still points
// at the block the shader declares — the failure otherwise is a
// highlight rolloff read out of a crop rectangle, which renders as
// nonsense rather than as an error.
assert_eq!(
BASE_CURVE_UNIFORM_OFFSET + BASE_CURVE_UNIFORM_FIELDS,
BASE_UNIFORM_FIELDS,
"the base curve must be the last thing in the base block"
);
assert_eq!(BASE_CURVE_POINTS * 2 + 1, BASE_CURVE_UNIFORM_FIELDS - 1);
assert!(compose(&[]).uniforms.len() >= BASE_UNIFORM_FIELDS);
}
/// Compose with neutral framing into a chosen output space.
fn compose_to(ops: &[Box<dyn Operation>], output: ColourSpace) -> ComposedShader {
compose_with_framing(ops, &Framing::new(), output)
@@ -1051,4 +1508,46 @@ mod tests {
let shader = compose(&[fake(&DESC_A, 1.0, false)]);
assert!(shader.source.starts_with("// GENERATED"));
}
#[test]
fn detail_operations_agree_with_themselves() {
// An operation says which stage it belongs to in two places — through
// `affects()` and through `detail()` — and the two must say the same
// thing. Disagreement is the worst possible failure mode here, because
// it is silent: an operation claiming `Affects::Detail` while
// returning `None` from `detail()` is fused as a point op and asked
// for a fragment it does not have, and one returning `Some` while
// claiming `Affects::Colour` is filtered out of the fused pass and put
// in the wrong invalidation bucket. Either way the slider moves and
// nothing happens.
//
// Checked over the real chain, plus the test consumer, so that a
// sharpening operation added later is covered by this without anyone
// remembering to extend it.
let mut ops = crate::ops::chain();
ops.push(Box::new(crate::detail::probe::BoxBlur::new()));
for op in &ops {
let id = op.descriptor().id;
assert_eq!(
op.detail().is_some(),
op.affects() == Affects::Detail,
"{id} disagrees with itself about whether it is a \
neighbourhood operation"
);
}
}
#[test]
fn a_detail_operation_never_contributes_a_fused_uniform() {
// Slot order in the generated block is emission order, and nothing
// addresses a slot by number — so an operation that contributed a
// uniform without contributing the fragment that reads it would shift
// every later operation's uniforms out from under its shader. The
// filter in `compose_full` prevents it; this is the assertion that the
// filter is on the right side of the loop.
let mut ops = crate::ops::chain();
ops.push(Box::new(crate::detail::probe::BoxBlur::with_radius(0.05)));
let before = compose(&crate::ops::chain()).uniforms.len();
assert_eq!(compose(&ops).uniforms.len(), before);
}
}