Compare commits

..
Author SHA1 Message Date
dtourolle 8aa10cd249 Release 0.24.0
Benchmarks / CPU and I/O (per commit) (push) Successful in 8m56s
Benchmarks / Frame budget (on demand) (push) Skipped
Traceability / Requirement traces (push) Successful in 1m27s
Build and test / Desktop (Linux) (push) Successful in 1h33m5s
Build and test / Layer separation (push) Successful in 32s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 4s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 2s
Build and test / Android (aarch64) (push) Successful in 45m54s
Build and test / Windows (x86_64, cross) (push) Successful in 55m18s
Build and test / Publish the release (push) Successful in 1m33s
2026-10-07 07:27:52 -04:00
dtourolle c2cfacd7d3 Record the new Best in the spec and the manual
denoise.md §15: why Best became one network, how it compares with the
mixture and Medium on real photographs and the chart, the candidates that
fell short, the file names, the saved-edit numbering, and the timings on
the 3050 -- 0.51-0.54 s whole-frame, 0.95 s in tiles, against 2.60 s for
the mixture in tiles. The manual lists Bilinear, Fast and Best, says an
edit made with Medium opens with Best, and gives the new time; Medium's
close-up goes.
2026-10-07 07:14:52 -04:00
dtourolle a9271c4850 Make Best one network, and retire Medium and the mixture
Best was a mixture of two experts and a gate, 110 GMAC a megapixel;
Medium a single network at 48 that was softer on real edges. fb-combo
(darkroom-denoise, 20 000 steps from fb-edges2, taught by the mixture
with a quarter of its crops from the edge-rich parts of the frames) is
Medium's shape and holds the mixture's edges on real photographs:
edge PSNR within 0.04-0.06 dB at ISO 1600/6400/25600, more sharpness
kept at all three, the chart's edge 0.89 photosites wide against 0.82.
It is 0.27 dB short on smooth areas at ISO 25600. It becomes Best, and
the methods are Bilinear, Fast and Best.

Saved edits keep their numbers: 2, which was Medium, is now Best, and
3, which was Best, is past the end and reads as the default, Best.
The network ships as mosaic-hq, a new name: the result cache keys a
model by name and size, and this one is byte for byte the old Medium's
size. Its tablet form (A16W16) lost 0.00 dB in simulated QDQ at every
ISO and at most 0.09 dB across the noise bracket.
2026-10-07 06:58:24 -04:00
dtourolle 4b71ef0947 Record whole-frame denoise in the spec and the model licences
denoise.md §14: why the tiles waste half of Best's work, where the any-size
networks run and why only there, why the limit is the card's memory, and
the measurement on _MG_8862 — 2.60 s in 1408 tiles, 1.37 s in two
4160 x 3248 tiles, the outputs within fp16's own spread. models/LICENCE.md
lists the three re-exports.
2026-10-06 23:14:11 -04:00
dtourolle 1e1aa1442b Size the whole-frame profile for a 6 GB card
TensorRT plans its memory for the profile's largest shape, and up to a
whole 6D frame with Best's border (4608 x 6656) it asked for 4.9-5.9 GB
and would not build on the RTX 3050. The profile now ends at 4608 x 3328
(15 MP), tuned for 4160 x 3248, and the tiler cuts a 6D frame into two
such tiles: 27 MP of work for 20 MP kept, against 49 MP in 1408 tiles.
The engine's directory names the profile, so a later range never loads
an engine built for this one.
2026-10-06 23:14:10 -04:00
dtourolle 3761281dd1 Ship the denoise networks with any height and width
mosaic-{fast,medium,best}.onnx are the shipped networks re-exported by
darkroom-denoise tools/export_whole.py (22ea648) from the checkpoints
the 1408 files came from: identical to them at 1408 (max |d| = 0), to
torch at 592 x 848, and to tiled inference over the reflected frame in
f64. 42 MB together. The Arch package and the Windows installer carry
them beside the fixed files; the APK leaves them out, since the Hexagon
takes fixed shapes only.
2026-10-06 21:47:20 -04:00
dtourolle 9cba420fd5 Denoise a whole frame in one call where the GPU takes any size
A fixed 1408 tile is exact only in its centre, and Best keeps 896 of
every 1408 it computes: 2.47 photosites of work for each one kept. The
tiler now takes a network of any size as well as a square one, and
plans the frame as the fewest equal tiles under the rung's limit --
one tile, the whole frame and its reflected border, whenever it fits.
If the first call of a plan fails, as a GPU out of memory does, the
kept centre is halved and the frame planned again.

Each shipped network names its any-size sibling (mosaic-best.onnx
beside mosaic-best-1408.onnx). OnnxNet::open takes it where the engine
runs whole frames and the file is installed, and the 1408 tiles
otherwise; open_tiled forces the tiles, and denoise_raw's DR_PLAN=tiles
uses it to compare. The cache key stays on the fixed model: the output
is the same network's. Tests hold any-size tiles, a grid of them and a
plan rebuilt after a failure to the square tiles' answer in every Bayer
phase.
2026-10-06 21:47:19 -04:00
dtourolle 56f4180347 Run a denoiser of any input size on TensorRT and CUDA
Role::WholeDenoiser is the denoise network exported with any height and
width, for a whole frame instead of 1408 tiles whose borders are thrown
away. It is served only where a new size costs nothing: the CUDA
provider, and TensorRT through an optimisation profile from 256 to
4608 x 6656, tuned for the 6D's frame with Best's border. Everywhere
else whole_frame_limit() says None and the fixed tiles run.

ort's TensorRT builder has no profile options, so the engine registers
through the runtime's V2 options with the names 1.30 reads
(trt_profile_{min,opt,max}_shapes). Without a profile a dynamic input
compiled an engine per size at run time, 156 s on the first frame. The
engine lives in its own directory per model: ORT's cache key leaves the
shape out, and the fixed 1408 export and its any-size sibling are the
same graph.
2026-10-06 21:41:31 -04:00
dtourolle 417cba8b4d Silence the hardware module where no runtime is loaded from disk
hardware::detect is read only by api::install_best, which exists with
the native feature; a build of a crate that takes the engine without it
(dr-denoise's own tests) warned that all of it was unused.
2026-10-06 21:41:13 -04:00
dtourolle 7966bf2dd8 Release 0.23.0
Benchmarks / CPU and I/O (per commit) (push) Successful in 8m42s
Benchmarks / Frame budget (on demand) (push) Skipped
Traceability / Requirement traces (push) Successful in 1m22s
Build and test / Android (aarch64) (push) Successful in 47m15s
Build and test / android-image (push) Successful in 3s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / Desktop (Linux) (push) Successful in 1h31m5s
Build and test / windows-image (push) Successful in 3s
🐳 Windows image / Build and push (push) Successful in 2s
Build and test / Layer separation (push) Successful in 53s
Build and test / Windows (x86_64, cross) (push) Successful in 55m3s
Build and test / Publish the release (push) Successful in 1m58s
2026-10-04 22:28:19 -04:00
dtourolle 4822991bec Read an unknown SoC as Qualcomm's
The generic runtime is opened only when the QNN build does not fit, and
the fit rests on ro.soc.manufacturer. A property the app cannot read, or
an Android older than 12 that has none, read as "not Qualcomm" would
put a Qualcomm device on the generic rung and off its Hexagon — what
0.22.1 had just fixed. Only a device that names another vendor is now
not Qualcomm's; the tablet reports QTI.
2026-10-04 22:26:56 -04:00
dtourolle 4b33754482 Warm a rung up before the probe times it
One warm-up run and the median of three: on the Iris Xe the OpenVINO
rung lost to the CPU on the smallest detector in two probes of three,
because an idle integrated GPU takes a few runs to raise its clock —
warm, it is 5.8 ms against 9.5. Three warm-ups and the median of seven
took it in five probes of five (5.6–7.2 ms against 7.7–13.1). The extra
runs cost tens of milliseconds, once per fingerprint.
2026-10-04 21:00:15 -04:00
dtourolle b3999bbc0c Record the Intel and generic rungs in the inference spec
§1.6 is the Iris Xe measurement: OpenVINO fp16 1.3–5.8× the CPU
provider on every shipped model, WebGPU behind it everywhere but the
denoiser and MI-GAN. §2's ladder gains the Intel and generic rows, the
generic one footnoted as unmeasured where it is meant to help. §3.1
lists the two bundled runtimes' licences; §3.2 is how one runtime of
several is chosen per process. D13 notes the bundling.
2026-10-04 21:00:15 -04:00
dtourolle 43402bfe5b Give a phone without a Qualcomm SoC the generic WebGPU runtime
The APK's ONNX Runtime is the QNN build, which carries no WebGPU, so a
non-Qualcomm phone had nothing above the CPU provider. The APK now also
carries Microsoft's stock onnxruntime-android 1.29.0 (32 MB) as
libonnxruntime_generic.so, and the app offers it after the QNN build.

The runtime search stops at a perfect fit, so on a Qualcomm device the
QNN build — listed first — is all that is opened, and the generic build
never loads beside it. A Qualcomm SoC is read from ro.soc.manufacturer
or, before Android 12, from Qualcomm's FastRPC library being present:
a Qualcomm device mistaken for another would trade its Hexagon for the
generic rung.
2026-10-04 21:00:15 -04:00
dtourolle cb97ebe7ac Ship the OpenVINO and WebGPU runtimes in every desktop package
The Windows installer and the Flatpak carried no ONNX Runtime, so they
ran every model on tract's one core; the Arch package left it to an
optional dependency. Each now installs two builds under runtimes/ —
Intel's OpenVINO build and the generic WebGPU one, both with the CPU
provider — fetched by tools/fetch-bundled-runtimes.sh from PyPI wheels
pinned by SHA-256, pruned to the native libraries (81 + 31 MB on Linux,
67 + 42 MB on Windows), licence texts beside them.

darkroom-desktop searches runtimes/openvino and runtimes/webgpu under
each place a package installs to; the engine opens all it finds and
keeps the one that fits the GPU, so a CUDA or ROCm runtime installed
beside them still wins on its vendor's card. On Windows the chosen
runtime's directory goes on PATH, because Intel's build leaves OpenVINO's
DLLs for the loader to find there.

The Windows image gains unzip; the installer smoke test checks both
runtimes landed.
2026-10-04 21:00:15 -04:00
dtourolle 2ced6f114f Name the rung that lost, not one the runtime lacks
A device left on the CPU gave the first failure as the reason, which on
any runtime but NVIDIA's is "TensorRT execution provider is not enabled
in this build". The reason is now the last rung that was tried and lost,
"WebGPU 150.8 ms, slower than the CPU's 26.6 ms"; the full list is still
in the status's failures.
2026-10-04 21:00:15 -04:00
dtourolle 85dee4375b Load the runtime that fits the GPU, not the first one found
A runtime carries one vendor's providers, only one loads per process,
and a device can now hold several: the package's OpenVINO or WebGPU
build, a CUDA build the user fetched, the distribution's ROCm build.
`api::install` opens each it finds, lists its providers with
GetAvailableProviders, and installs the one scoring highest against the
GPUs `hardware::detect` reads from files — a vendor rung on its own
vendor's GPU above OpenVINO on an Intel one above the generic WebGPU
rung above a CPU-only build. Equal scores keep the old first-found
order, and DARKROOM_ORT_DIR still wins outright. The losers stay mapped
rather than unloaded.

The Linux fingerprint now names the OpenCL drivers too, so installing
Intel's re-probes. `ladder` takes DARKROOM_ORT_DIRS to show the choice.
2026-10-04 21:00:15 -04:00
dtourolle 87c405eb46 Add OpenVINO and WebGPU rungs to the inference ladder
OpenVINO is the Intel rung: the integrated or Arc GPU, fp16 for every
role but the embedder, a compiled program per model kept in a directory
per model, precision and runtime version. On the Iris Xe it beats ONNX
Runtime's CPU provider on every shipped model — scrfd_10g 23 ms against
58, the scene model 17 against 57, MI-GAN 57 against 330, a denoise tile
40 against 158.

WebGPU is the generic rung for a GPU no vendor rung covers. It was slower
than the CPU on the Iris Xe, the RTX 3050 and the Adreno, so it is on the
ladder for the GPUs it has not been timed on, behind the probe's clock.

MIGraphX's registration becomes one generic key/value helper that all
three share, with option names read from each runtime's own source.
2026-10-04 21:00:15 -04:00
dtourolle 555ec0efb3 Let ep_probe name the OpenVINO GPU
On the hybrid laptop OpenVINO's `GPU` was the RTX 3050 through NVIDIA's
OpenCL, not the Iris Xe; DARKROOM_OV_GPU picks GPU.0, GPU.1 and so on.
2026-10-04 21:00:15 -04:00
dtourolle 0291b80672 Feed every input in ep_probe
The denoiser takes `mosaic` and `sigma`; with only the first fed, every
provider reported the same failure and the model went unmeasured.
2026-10-04 21:00:15 -04:00
dtourolle ef71bb3289 Time OpenVINO and WebGPU in ep_probe
The Intel and vendor-neutral rungs need a measurement before they join
the ladder (docs/dev/inference.md §2). Both register through the generic
key/value entry point with the option names ONNX Runtime reads at the
wheel's version: OpenVINO 1.24 (`openvino_provider_factory.cc`), WebGPU
1.27 (`webgpu_provider_options.h`, prefixed by the runtime).
DARKROOM_EPS narrows the list to the families a runtime carries.
2026-10-04 21:00:15 -04:00
dtourolle 08b7d23e86 Release 0.22.1
Benchmarks / CPU and I/O (per commit) (push) Successful in 8m35s
Benchmarks / Frame budget (on demand) (push) Skipped
Traceability / Requirement traces (push) Successful in 1m22s
Build and test / Android (aarch64) (push) Successful in 46m43s
Build and test / android-image (push) Successful in 2s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / Desktop (Linux) (push) Successful in 1h4m43s
Build and test / windows-image (push) Successful in 2s
🐳 Windows image / Build and push (push) Successful in 2s
Build and test / Layer separation (push) Successful in 34s
Build and test / Windows (x86_64, cross) (push) Successful in 29m53s
Build and test / Publish the release (push) Successful in 1m24s
2026-10-04 20:35:01 -04:00
dtourolle a116325991 Declare the DSP's RPC library, so the app can reach the Hexagon
QNN's Hexagon stub loads libcdsprpc.so, a vendor library, and from API 31
an app's linker namespace refuses a vendor library its manifest does not
name. QNN then fails to create its device - QNN_DEVICE_ERROR_INVALID_CONFIG,
before it reaches the DSP - and every model ran on the CPU on 0.22.0: AI
denoise took 30-131 s a photograph on the tablet.

The same engine code ran on the HTP from adb's shell, whose namespace has
no such rule, which is what hid it. With the declaration the app's probe
chose the Hexagon on the tablet and compiled every model for it. Not
required, so a device without the library still installs and runs on the
CPU. A test holds the line in the manifest.
2026-10-04 20:19:04 -04:00
dtourolle f20e481358 Probe the Hexagon strictly, and probe again after falling back to the CPU
0.22.0's first launch on the tablet: QNN could not create its device
(QNN_DEVICE_ERROR_INVALID_CONFIG), the session built anyway with every
node on the CPU behind the provider, and the probe timed that - 28.5 ms
against the CPU's own 19.4 - and rejected the Hexagon. The verdict was
cached under the fingerprint, so every model stayed on the CPU on every
later launch: AI denoise took 30-131 s a photograph instead of seconds.
The same A16W8 detector with the APK's own libraries runs on the HTP in
4.4 ms.

The probe's Hexagon session now sets session.disable_cpu_ep_fallback, so
a device that cannot take the graph fails the probe instead of being timed
as the CPU. Only the probe: shipped graphs may keep nodes on the CPU on
purpose. And a selection that fell back to the CPU after an accelerator
failed or lost is probed again on the next launches, up to three probes
per fingerprint; a cache written by 0.22.0 reads as never retried, so the
tablet probes again once this is installed.
2026-10-04 19:42:15 -04:00
dtourolle 16a5957aa7 Compile dr-ui on sixteen codegen units in release builds
Slint expands the .slint files into ~27 MB of Rust, and at the workspace's
single codegen unit LLVM optimised all of it on one thread: 13.5 minutes
of a release build with the other cores idle. The override applies to
dr-ui alone; the image crates keep one unit, and thin LTO still runs at
link time.
2026-10-04 10:45:43 -04:00
dtourolle 5736a21a3a Release 0.22.0
Benchmarks / CPU and I/O (per commit) (push) Successful in 8m36s
Benchmarks / Frame budget (on demand) (push) Skipped
Traceability / Requirement traces (push) Successful in 1m31s
Build and test / Android (aarch64) (push) Successful in 48m34s
Build and test / android-image (push) Successful in 4s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / Desktop (Linux) (push) Successful in 1h32m33s
Build and test / windows-image (push) Successful in 2s
🐳 Windows image / Build and push (push) Successful in 2s
Build and test / Layer separation (push) Successful in 29s
Build and test / Windows (x86_64, cross) (push) Successful in 31m14s
Build and test / Publish the release (push) Successful in 1m9s
2026-10-04 08:13:35 -04:00
dtourolle b6ca7be185 Show each denoise method in the manual as a close-up
The manual's AI denoise section names the four methods and their measured
times, and shows the lamp and railing of the ISO 8000 frame at 1:1 by each
in place of the film and the before/after pair. The scene clicks each
method and waits for that network's result: the repair now logs its own
"learned denoise:" line first, so the wait matches the result's.
2026-10-04 08:11:24 -04:00
dtourolle 06422a07db Offer three denoise networks and a method to choose between them
AI Denoise's Apply switch becomes Method: Bilinear, Fast, Medium, Best,
default Best, so an untouched raw writes nothing and develops through the
mixture. `apply` is still read and never written: 0 is Bilinear, 1 keeps
a network already chosen.

- Best is the mixture of a flat and an edge expert with a learned gate;
  Medium and Fast are students distilled from it. 2.48 s, 0.79 s and
  0.57 s for a 20 MP frame on TensorRT fp16.
- Each network carries its own tile border (256 for the mixture, 192 for
  the students) through `dr_denoise::Shipped` and `TileNet::halo`.
- The file is hashed once at open and each network keys its own cached
  result; Bilinear keeps the result in memory for the way back.
- Each has an .a16w16 sibling for the Hexagon: 0.00 dB on the 6D gate,
  at most 0.11 dB with the noise scaled x0.5 to x4.
- APK BUNDLED 19 -> 23; the PKGBUILD installs all three.
2026-10-04 08:02:25 -04:00
dtourolle 14f08a565f Feed the denoise network without making it wait for the CPU
A 20 MP frame spent 0.32 s outside the network: each tile's mosaic and
sigma gathered on one thread, then its 24 MB output copied out of the
runtime and back into the frame, all in series with the device. Tiles are
now gathered on every core by a producer thread one tile ahead, so the
gather overlaps the run; the centre is written back across cores; and the
tile interface hands its inputs over and lends its output, so neither
side is copied. With a stand-in network that does nothing, the tiler's own
time falls to 0.14 s at the 1408 tile and 0.09 s at 2048. The exactness
and Bayer-phase tests are unchanged and pass.
2026-10-04 07:30:09 -04:00
dtourolle 0c9d564586 Repair photosites beyond 8 sigma of every neighbour before the network
The app's hot-pixel pass takes gross defects only; at ISO 6400-25600 a 6D
frame keeps 1000-2000 photosites more than 8 sigma beyond all their
same-colour and adjacent neighbours, which the network turned into specks.
The same two tests with the threshold in the photosite's own sigma, plus
the factor of two that keeps a bright point of light (where 8 sigma is a
sliver of the signal). The next model is trained behind exactly this; on
an ISO 25600 frame the Rust and training code both repair 935.
2026-10-04 07:30:09 -04:00
dtourolle 75e12441fd Share Lightroom's saturation bands across ours at measured strengths
Photographs opened with an earlier Lightroom edit now import its HSL
saturation as fitted against the library's own Lightroom 6 exports, rather
than one band to one band.

Measured on two looks' exports and their raws (darkroom-lrfit, hsl_map_fit),
by encoded hue: Lightroom's saturation bands act about 45 degrees either
side on our wheel, wider than ours, and not all at our strength. Each is now
shared between two or three of our bands — Aqua mostly cyan and azure, where
skies are; Blue mostly blue and violet; Orange, where skin is, at about 0.4
of its value. Values add when two of Lightroom's bands share one of ours.
On the measured skies the import now lifts muted sky blues about 1.9× against
Lightroom's 2.1×, where it gave 1.15×. Hue and luminance still go one band to
the band of the same hue; they were not measured.
2026-10-04 05:28:33 -04:00
dtourolle 3f8f909e41 Make the colour mixer's saturation reach muted colours
Every photograph with a colour-mixer saturation edit now renders differently:
a raised band is stronger, most of all on muted colours.

The mixer matched bands and judged saturation on scene-linear values, and
scaled chroma by the same factor whatever a colour started at. Against the
photographer's earlier exports of two looks (~90 photographs, their raws, by
encoded hue), a sky band raised by 58 there lifted muted sky blues about
2.1×; here the mixer gave 1.15×, and less in the muted tones that carry most
of a sky or a shadowed snowfield.

Bands are now matched and saturation judged on display-encoded values. A
raised band pushes muted colours hardest and tapers to nothing at full
saturation, at a gain of 3.0, which at the same value lifts muted sky blues
about as those exports did. Lowering saturation still scales every colour
alike. Hue shifts work on the same encoded colour; luminance still scales in
linear light.

The bundled presets that use the mixer, and those whose colour was tuned
against the default rendering, are rescaled to the amount of colour they
had: Vivid 1.30, Vivid warm 1.30, Vivid landscape 1.38, Vivid, strong 1.45,
Vivid portrait 1.15, Punch 1.12, Blue sky 1.08, Deep blue sky 1.12, Polariser
1.26, Blue sky, golden land 1.13 — mean CIELAB chroma over the default
rendering, on 30 raws from the library. Negative values (skin protection)
are left as written.
2026-10-04 05:28:22 -04:00
dtourolle 5ffd54ba43 Leave the profile's look table off by default
Every raw rendered through a camera profile — the library's DNGs with an
embedded profile, and CR2s given one — now renders differently: more
colourful in near-neutral tones. The profile's look table is no longer
applied unless its slider is raised; PROFILE_LOOK names the strength the
profile states.

Against the photographer's earlier exports with no look applied, the default
rendering scores the same with the look table at 100, 50 or 0 (held-out MSE
140, 140, 143), and is 9 % more colourful at 0: the table lowers the
saturation of near-neutral tones, which is exactly where the default
rendering was short of those exports. The user chose more colour.
2026-10-04 05:28:15 -04:00
dtourolle 5a8c3e4c40 Run each model on the Hexagon in the form measured to hold it
The engine knew f32 and int8, and gave the Hexagon int8 for every role it
served. Measured on the tablet itself (inference.md §1.5), int8 lost
5% of the detector's faces at 40-80 px, moved the landmarks 1.5 px,
emptied the segmenter's scores and cost the denoiser 5-9 dB; fp16 the HTP
refuses outright. `Form` gains A16W8 and A16W16, and `Rung::form` now
names one per role: detectors and landmarks A16W8, the segmenter, scene
model, border filler and denoiser A16W16, XFeat int8. The embedder and
the eye classifiers stay on the CPU.

Each loader resolves its `<stem>.<form>.onnx` sibling; the segmenter and
XFeat, compiled into the binary, embed their quantised forms on Android
only and pick through `choose_embedded`. The probe, the compile step and
the cache fingerprint follow the form instead of assuming int8. Detectors
on the new form write `scrfd_*_a16+w600k_mbf`, and `model_ids` answers
for all three spellings.

On the tablet (ORT 1.29 + QNN 2.42), each shipped file against f32 on the
same inputs, and against the CPU's f32 time:
  SCRFD 500m/2.5g/10g  A16W8   100% of faces in every band   4.2/5.1/9.0 ms vs 17/56/198
  landmarks            A16W8   0.25 px in the 192 crop        0.5 ms vs 2.8
  YOLO26n-seg          A16W16  98.2% found, mask IoU 0.994    12.9 ms vs 90
  scene model          A16W16  98.9% of cells agree           15 ms vs 151
  MI-GAN               A16W16  41 dB from f32 in the fill     87 ms vs 488
  XFeat                int8    pano alignment 0.45 px (f32's own spread 0.41)  6.5 ms vs 58
  denoiser             A16W16  0.00 dB at every ISO            95 ms vs 1510 a tile
Face numbers are over public COCO val2017 photographs, not a library.

The APK carries the siblings (BUNDLED 15 -> 19; the old int8 detectors
removed), about 43 MB more. The Windows installer and its CI count skip
them; the Arch and Flatpak packages list their files and never had them.
The ladder example takes a role per model, which is how the per-role
forms above were seen landing on the NPU from the real probe.
2026-10-04 03:45:46 -04:00
dtourolle 0e6ac09fd5 Quantise for the Hexagon with QNN's config and the app's own inputs
tools/quantise-models.sh now writes each model's Hexagon form from a
per-model table: the form its role takes on the NPU (int8, A16W8 or
A16W16), the exact graph rewrites it needs, and the nodes that must stay
float. Ranges are min/max over photographs fed exactly as the app feeds
each model -- the detector and segmenter letterboxes with their own pads
and normalisation, landmark crops from the detector's boxes, MI-GAN with a
panorama-like border, XFeat's grey proxy. The old tool used an
antialiased resize, YOLO's pad of 128 and /255 for every model that was
not a face model, none of which is what the app does.

tools/htp_graph.py holds the rewrites, each checked against the input
graph before use: the denoiser's 6-D Bayer pack and XFeat's 224-slice
unfold as SpaceToDepth (QNN stops at rank 5), computed reshape targets
folded, and bilinear Resize as two MatMuls (the HTP refuses
ResizeBilinear at XFeat's sizes). The denoiser takes ranges computed by
darkroom-denoise's gate on a smaller tile of the same network.
2026-10-04 03:45:07 -04:00
dtourolle 948f6c3ed2 Count the denoise model in the Windows installer's smoke test
package.sh stages models/denoise beside face, scene and inpaint, and
the smoke test counted only the other three, so 0.21.0's Windows job
failed with "expected 14 model files, installed 15". The count reads
the same directories package.sh copies, as its comment intends.
2026-10-04 02:50:15 -04:00
dtourolle 8f9e59b9fa Find hot photosites without repairing them, and measure a sensor's aging
The hot-pixel pass could only repair: it returned how many photosites it
changed and threw away which. find_hot_pixels runs the same pass and
returns them as sensor coordinates, leaving the frame alone, so a sensor's
defects can be tracked across frames.

sensor_scan prints each frame's candidates, and with --probe reads a list
of coordinates back out of every frame. Run over 53 6D raws from 2015 to
2026, it found 32 persistent defects, 2 in 2015 and 32 by 2026, and showed
what a defect map has to account for: a frame that does not flag a
photosite proves nothing unless its neighbourhood is dark, and the 6D
hides some of its defects itself above ISO 5000. docs/dev/sensor-health.md
records the findings and the design they argue for.
2026-10-04 02:38:48 -04:00
dtourolle 83f0461ce7 Sync develop presets through the library
Presets were the one piece of the photographer's work that never left
the device: faces, sidecars, albums, collections, keywords and camera
profiles all travel with the sync pass, the preset library did not.

It now goes to <derived>/presets/library.drpl. PresetLibrary::merge
decides each name against the base the last exchange left (kept per
library beside place.json), so presets added on two devices both
survive, a deletion reaches the other device instead of being restored
by it, and an edit outlives a deletion made elsewhere. The upload is
If-Match / If-None-Match on the server's copy, and a 412 reads and
merges again, so two devices exchanging at once cannot save over each
other. A server copy that will not parse (a newer build's) is left
alone, and a local file that will not read stops the exchange rather
than being taken for an empty library.

The develop view's save merges with the file when the sync changed it
since the view read it, and a sync that brought presets reloads and
redraws the list.

Also corrects the register, which still said camera profiles do not
sync.
2026-10-04 00:50:09 -04:00
dtourolle 26e50ae723 Save presets beside the settings, not under a raw HOME
PresetStore::open built its path from XDG_CONFIG_HOME or HOME. Android
sets neither, so the library resolved to /.config/darkroom, which is
read-only, and every preset saved on the tablet failed. Windows sets no
HOME either and got a directory relative to the working directory. The
settings store was moved to dr_sync::account::config_dir for the same
reason in 0.12.1; the presets now follow it. Linux and macOS resolve to
the same file as before.
2026-10-04 00:49:59 -04:00
dtourolle 25dc0d0179 Give the Vivid presets and Punch measured amounts of colour
On the default rendering, Vivid added 22 % more chroma than the rendering
itself, and Punch 5 % — less than the photographer's earlier exports show
with no look applied (14 % over ours) and well below their everyday look
(27 %). Each preset's colour values (vibrance, saturation, the mixer's
saturation bands) are now scaled together, tone values untouched and
negative ones — Vivid portrait's skin protection — left as written, until
the preset measures: Vivid and Vivid warm 1.30, Vivid landscape 1.38, Vivid,
strong 1.45, Vivid portrait 1.15, Punch 1.12. Measured as mean CIELAB chroma
over 30 raws from the library, as a ratio to the default rendering.
2026-10-03 23:04:46 -04:00
dtourolle a36ec98b36 Carry Lightroom's tone sliders across at measured strengths
Contrast2012 and the four recovery sliders were imported one to one. They do
not mean the same thing here: fitted on the library's Lightroom 6 exports and
their raws — each photograph's sliders carried across as slider × factor, one
factor per slider, on about 90 exports with no look applied, on the Camera Raw
default rendering — ours needed contrast at about a tenth (Lightroom's −100
imported as ours flattens a frame to grey), highlights ×1.4, shadows ×1.9 and
blacks ×1.25. Whites fitted below 1 every time without agreeing where; 0.5 is
a hedge, and says so. Vibrance stays one to one: the op itself is now
calibrated to Lightroom's.
2026-10-03 23:04:46 -04:00
dtourolle c343ac79d3 Describe AI Denoise in the manual as it now is
On for every raw, at the top of the Adjust panel, kept once computed,
and eased off with Strength rather than Keep grain. The timing line is
left as it was; the new model's measured figure replaces it when that
branch lands. The animation still shows the Keep grain slider and
wants recording again.
2026-10-03 22:18:15 -04:00
dtourolle 2e7f14dafe Develop every raw through the AI denoise by default, with a strength, cached
The learned demosaic was an option under Detail, off by default. It is
now how a Bayer raw is developed: on by default at full strength on
every device — which hardware runs it is the inference engine's choice
— and first in the Adjust panel, since it decides what every control
below is applied to.

Strength (0-100, default 100) replaces Keep grain: grain = 100 -
strength, the same luminance-only blend, so moving it is one GPU pass
and never a re-run. 0.21.0's sidecars stored grain; it is still read,
as the inverse, and never written.

With it on for every photograph, the result is now kept on disk
(denoise.md §7.1, §12): the network's output as half floats, keyed on
a SHA-256 of the file's bytes and the model, oldest first past a 5 GB
budget, beside the inference engine's cache. A reopened photograph and
an export of one already developed read it back instead of running the
network again; a damaged entry is a miss.
2026-10-03 22:16:36 -04:00
dtourolle ff0effbfe1 Count the denoise model in the APK's bundled-model list
Benchmarks / CPU and I/O (per commit) (push) Successful in 3m17s
Benchmarks / Frame budget (on demand) (push) Skipped
Traceability / Requirement traces (push) Successful in 56s
Build and test / Android (aarch64) (push) Successful in 47m9s
Build and test / android-image (push) Successful in 3s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / Desktop (Linux) (push) Successful in 1h2m3s
Build and test / windows-image (push) Successful in 2s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / Layer separation (push) Successful in 30s
Build and test / Windows (x86_64, cross) (push) Failing after 49m14s
Build and test / Publish the release (push) Skipped
96001480 added mosaic-1408.onnx to BUNDLED as a fifteenth entry and
left the array's declared length at 14, so the Android build failed
and 0.21.0 got no release. The workspace gates never compile the
Android crate, which is why nothing before CI saw it.
2026-10-03 22:14:26 -04:00
dtourolle 1a03cb52b4 Release 0.21.0
Benchmarks / CPU and I/O (per commit) (push) Successful in 8m38s
Benchmarks / Frame budget (on demand) (push) Skipped
Traceability / Requirement traces (push) Successful in 1m27s
Build and test / Android (aarch64) (push) Failing after 29m52s
Build and test / android-image (push) Successful in 2s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / Desktop (Linux) (push) Successful in 1h2m28s
Build and test / windows-image (push) Successful in 2s
🐳 Windows image / Build and push (push) Successful in 2s
Build and test / Layer separation (push) Successful in 30s
Build and test / Windows (x86_64, cross) (push) Failing after 48m48s
Build and test / Publish the release (push) Skipped
2026-10-03 17:06:49 -04:00
dtourolle ff4b30fbaa Link the inference engine for macOS in a zig container
`docker/macos` builds for aarch64-apple-darwin from Linux with
cargo-zigbuild. Zig carries libSystem and the C headers, so tract's SIMD
kernels compile and the engine's test binaries and examples link as Mach-O
arm64 — the check `cargo check --target` could not do, because tract's
build script needs a macOS C compiler. Crates that link an Apple framework
(dr-plat's keyring, and so the app) still need the Xcode SDK and fail at
the link; macos.md says so.
2026-10-03 16:50:37 -04:00
dtourolle c73743394f Add a CoreML rung on macOS
The macOS ladder was the CPU provider alone, with CoreML listed as a gap.
It is now CoreML, then the CPU, then tract — unmeasured, since nobody here
has a Mac, and safe to ship unmeasured because the probe's clock rejects a
CoreML slower than the CPU and `attempt` refuses one that crashes.

- `Rung::CoreMl`, a compiling rung like TensorRT: an ML Program with every
  compute unit allowed, falling back to the CPU until each model's program
  is built. The embedder stays on the CPU, as on the Hexagon (§7).
- The cache is one directory per model and runtime version. CoreML keys a
  model committed from memory on its input and node names, not its
  weights (ONNX Runtime 1.29, coreml_execution_provider.cc), so two
  exports of one architecture would otherwise share a program.
- The fingerprint on macOS is the chip and the OS release, which ships
  CoreML.
- The desktop looks for the runtime in the bundle's Contents/Frameworks
  and Homebrew's prefixes; fetch-desktop-runtime.sh on a Mac downloads
  ONNX Runtime 1.29.0 for Apple silicon, which carries CoreML.

docs/dev/macos.md says what exists, how to build it, and which log lines
to ask a Mac user for.
2026-10-03 16:50:37 -04:00
dtourolle 872e35670c Log like a debug build on macOS, where a Mac user can find it
Nobody working on DarkRoom has a Mac, so every macOS build is in the hands
of someone who can send a log and cannot attach a debugger. Three changes
make that log worth sending:

- The desktop's default filter on macOS is `debug` for every `dr_*` crate,
  the desktop crate and `onnxruntime` (the runtime's own session log).
- The state directory — the log and crash records — is `~/Library/Logs`
  on macOS rather than the `~/.local/state` Finder hides; Console.app
  lists it. Config and data keep the Unix rules.
- A `diagnostic` cargo profile: release plus line tables, so a crash
  record's backtrace reads file:line. On macOS the tables are in the
  `.dSYM` beside the executable, which the bundle must keep.
2026-10-03 16:49:58 -04:00
dtourolle b562d7b1af Stop retrying a provider that took the app down
The probe runs in the app's process, and a provider can fail by aborting
rather than by returning an error — XNNPACK did on SCRFD. A rung that does
that once would do it on every launch, before the first photograph is on
screen.

Every session build above the CPU, the probe's and each background
compile's, now writes what it is attempting to `attempt` in the cache
directory first and removes it after. After two launches in a row that
died inside the same attempt it is refused and recorded — a rung in
`failed`, an engine in the new `refused` — until the fingerprint changes.
Two, not one, because quitting during a TensorRT compile leaves the same
file.
2026-10-03 16:49:58 -04:00
dtourolle 3689b06c35 Send ONNX Runtime's session log to the app's log
A native session's messages went to ONNX Runtime's stdio logger, which is
nowhere once the app is launched from a menu — and what a provider says
while partitioning a graph (nodes taken, operators declined, a library that
failed to load) is most of what a failed rung tells you. Each session now
forwards them to `log` under the target `onnxruntime`: warnings always,
the runtime's info lines at `debug`, its verbose lines at `trace`.
2026-10-03 16:49:58 -04:00
dtourolle 64ea44aefe Stop reading dates at the first sign the server is unreachable
A window of cells whose thumbnails were cached but whose dates were not
sent a header read per cell, and offline each one was three attempts at
a 15 s connect timeout: `is_transient` counts a network error as worth
retrying, and `read_metadata_only` returned a bare bool that could not
say why a read failed. So the grid sat on "reading N dates" for minutes
against a server that was not there, and no banner went up, because
nothing in that loop ever reported the connection.

`read_metadata_only` now returns a `DateRead`: reached, failed, or
offline. An offline error is returned on the first attempt rather than
retried — a dead server answers the second exactly as the first — while
a 423 lock is still retried, which is what the retry was for. The grid's
worker stops on it and sends `Offline`, as its fetch loop already did,
so the banner goes up and the bar stops. The sweep's lanes stop on it
too, one timeout each rather than one per image.
2026-10-03 16:49:49 -04:00
dtourolle 32a4da0e94 Import DEFAULT_CONTRAST only where the tests use it
The calibration commit imported it at module level, where only the
test module reads it; the build warned and clippy's -D warnings
refuses that.
2026-10-03 16:45:02 -04:00
dtourolle 33779a70bd Answer a thumbnail miss with the other stored class before the network
Offline, a grid zoomed past 256px was blank wherever it had not been
zoomed over before. The store was asked only for the exact class the cell
wanted, and the sweep stores only the grid class, so every zoomed cell
missed and went to a server that was not there — with its 256px thumbnail
sitting in the store the whole time. Online it cost the same round trip,
just without the blank cell at the end of it.

The split now tries the other class on a miss. A smaller one stands in
and the fetch for the real class still goes out; a larger one answers
the request outright, since there is nothing a fetch would improve on.
`ThumbnailReady` carries the class its pixels are, and the drain records
that rather than the batch's class, so a stand-in is replaced on the next
reload instead of being counted as served. `already_served` counts a
held large thumbnail as serving the grid class too, so zooming back out
does not re-read the store for pixels already on screen.

The split moves into `split_by_store` so it can be tested without a
worker thread or a network.
2026-10-03 16:44:58 -04:00
dtourolle 185e134ead Describe the measured tone and vibrance in DarkRoom's own terms
The calibration commits named DarkRoom's default curve after another
product and described vibrance as doing what another editor's does at
the same value. The curve is the DNG SDK's reference, so it is called
that; vibrance is scaled to deliver the strength its value names, as
measured against the photographer's earlier exports. Two test names
follow. The measurements and where they came from are unchanged.
2026-10-03 16:41:11 -04:00
dtourolle 7ae1e27810 Make vibrance deliver the strength its value names
Vibrance delivered about a third of its nominal effect. Its falloff measured
saturation on scene-linear values, where an ordinary tan reads as 0.78 and
keeps a twentieth of the effect; and its skin guard halved it wherever red led
green led blue — 38 % of the pixels of the gallery's exports, every warm colour
rather than skin. The Vivid presets lean on vibrance, which is why they added
less colour than their values promised.

Saturation is now judged on display-encoded values, the guard covers skin hues
(about 10-50 degrees, not strongly saturated), and the gain is fitted: on 45
Lightroom exports whose only colour setting was Vibrance (about +24), the
measured-to-nominal scale was 1.9 before and 1.08 at a gain of 1.2, so 1.3.
2026-10-03 15:54:41 -04:00
dtourolle a4ff7ec2b9 Render raws through the DNG reference curve by default, at contrast 1.5
Decides D21 by measurement. The photo gallery holds Lightroom 6 exports of
raws in the library, each carrying its Camera Raw settings; clustered by
those settings, 663 had no look applied. On 60 of them with their raws,
a third held out, the held-out MSE against Lightroom's JPEG was about 1200
for 0.20.0's sigmoid (0.7 EV darker and flatter), 224 for the DNG reference curve
after baseline exposure, and about 140 once its input is bent by 1.5/1.4
about grey.

So the curve choice defaults to the DNG reference, keeping its index (sidecars
record it), and the default contrast is 1.5. Contrast under the DNG reference curve is
now a power relative to REFERENCE_CONTRAST (1.4), where the table is
untouched; the sigmoid at that contrast still matches the retired base
curve. JPEGs are unaffected: the view transform skips a rendered source.
2026-10-03 15:54:21 -04:00
dtourolle b69fb3e191 Give the denoise work's test images a baseline exposure
The learned-denoise branch merged while this one was open, and two of
its test fixtures build a RawImage without the baseline_exposure field
this branch added; zero is the no-op value.
2026-10-03 14:44:48 -04:00
dtourolle 7a09b640d7 Satisfy clippy on the reference tone curve's data
One sample of the ACR3 table is 0.70711, which clippy reads as an
approximation of 1/sqrt(2). It is the curve's published value, so the
lint is allowed on the table with that reason rather than the number
replaced. And the pair-count check uses is_multiple_of.
2026-10-03 14:22:49 -04:00
dtourolle 8294e6b59f Call the reference curve what it is, and drop wording that reads as copying
The view transform's second curve is the DNG SDK's published reference
rendering — the ACR3 default curve applied by RefBaselineRGBTone — so
it is the "DNG Reference" curve in the panel, D21 and the code, not a
name borrowed from another product. Comments and docs that justified a
choice by another editor doing it ("as their Amount", "so a
photographer arriving from it finds the name") now give the actual
reason. The Vivid presets no longer describe themselves as reaching
for another editor's look; they are DarkRoom's own.

Factual mentions stay: which program wrote the library's DNGs, what
was measured against, and preset import. camera-profiles.md gains §15,
on starting a photograph from the edit it already carries.
2026-10-03 14:22:49 -04:00
dtourolle e12783da9f Open a photograph with the edit it already carries, when DarkRoom has none
The library's DNGs carry the photographer's earlier develop settings
in their embedded XMP — the house style their photographs were made
with. A photograph opened with no edit of DarkRoom's now starts from
that earlier edit, translated (HSL bands, highlights, blacks and the
rest), as one undoable step named "Earlier Edit"; from there it is an
ordinary edit, saved with the photograph. Export does the same, so a
photograph never opened exports as opening it would show.

Only on positive evidence that there is no DarkRoom edit: a local file
with no sidecar beside it, or a server that answered "no such file"
with nothing in the cache. The stored-edit fetch now says which
(FetchedSidecar::absent). Offline, unreachable or unreadable never
counts — the earlier edit would otherwise be saved over a real edit
that merely failed to arrive.
2026-10-03 14:22:49 -04:00
dtourolle 013596e1bd Translate Lightroom's HSL panel, and read the edit inside a DNG
The library's DNGs carry a Lightroom house look in their embedded XMP
— Blue +58, Aqua +50, Yellow and Purple +23, Highlights -40, Blacks
-20 on most — and that, not the camera profile, is why the same files
look richer in Lightroom. The importer skipped exactly that part: the
HSL panel was on its list of structures it did not translate.

Lightroom's eight HSL bands now map onto the colour mixer, hue,
saturation and luminance each one for one: Aqua to our cyan and Purple
to our violet, the nearest of our twelve bands by hue; chartreuse,
spring, azure and rose are left alone. The figures are a first
translation that lr-fit's measurement against Lightroom's output may
yet scale.

read_embedded finds the XMP packet in a photograph's bytes by its
delimiters and translates it, or answers None for a file whose XMP has
no Camera Raw settings — darktable's sidecars, a camera's own packet.
A test reads the library's _MG_9080.dng when it is present.
2026-10-03 14:22:48 -04:00
dtourolle 1e8594724e Keep the sigmoid as the default curve; Camera Raw's tone is a choice
The Camera Raw default rested on comparing against Lightroom previews
of photographs that carry the user's Lightroom edits — HSL saturation
Blue +58, Aqua +50 and more, Highlights -40, Blacks -20, in every
DNG's XMP — so it measured the house look, not Camera Raw's base
rendering. Under the ACR3 curve _MG_9080 renders brighter than its
Lightroom preview (mean 0.39 against 0.31).

So the curve choice's first variant, the default, is the sigmoid again
and every raw renders as in 0.20.0 apart from baseline exposure. D21
and camera-profiles.md §12 now say the default is open, to be decided
by measuring against Lightroom exports of unedited photographs. Tests
that are about Camera Raw's tone choose it explicitly.
2026-10-03 14:22:48 -04:00
dtourolle 0730ef1016 Sync camera profiles through the library, and name the tone curves
camera-profiles.md §13: the derived sync pass gains a profiles step,
after the catalog and before the place, that exchanges the profiles
directory with <library>/.darkroom-derived/profiles. Profiles are
immutable and named for what they hold, so name and size decide: it
uploads what the server lacks or holds at another size and downloads
what this device lacks — parsed before it is kept, written beside its
name and renamed — then reloads the set, so a profile copied out of a
DNG on the desktop renders the body's CR2s on the tablet after its
next sync. Like the place it never fails the pass. Not a catalog
table: a schema change would stop an older peer merging at all.

Labels for the view transform's new curve choice: Curve, Camera Raw,
Sigmoid.
2026-10-03 14:22:47 -04:00
dtourolle db7593f3dc Render raws through Camera Raw's tone by default, after baseline exposure
The rendering half of camera-profiles.md §11-§12 (D21). The view
transform gains a curve choice — Camera Raw (the default) or D19's
sigmoid. Camera Raw converts to linear ProPhoto, clips to [0, 1], runs
the curve on the largest and smallest channel and places the middle
one at its old fraction between them (RefBaselineRGBTone), and
converts back: hue kept, saturation raised where the curve is steep,
which is where Adobe Standard's look desaturated. White sets the input
scale (1 at its default, so sensor white is display white) and
contrast bends the input about grey (1 at its default).

The curve rides in the profile buffer after the tables: the profile's
own, else the ACR3 default, which the placeholder every profile-less
source binds also carries — so a CR2 with no .dcp still gets Camera
Raw's tone. Baseline exposure is a gain folded into the rendering
matrix at upload; RawImage::color_matrix stays the file's for the
merge's linear DNG.

camera_raw::apply_reference is the CPU statement; GPU tests hold the
shader to it on 256 colours and on greys against the ACR3 table. The
sigmoid's own tests now choose it explicitly.
2026-10-03 14:22:47 -04:00
dtourolle 38d414912c Read baseline exposure and profile tone curves; carry the ACR3 curve
The decoding half of camera-profiles.md §11-§12. RawImage gains
baseline_exposure: the file's BaselineExposure plus the chosen
profile's BaselineExposureOffset, as the DNG SDK sums them (+0.25 for
the library's 6D DNGs). A profile copied out of a DNG carries that
DNG's baseline as its offset, so the body's CR2s, which have none,
land at the same total.

ProfileTables gains the profile's ProfileToneCurve, resampled at
decode onto 1025 points with a natural cubic spline; an identity curve
counts as none. dr-types now holds Camera Raw's ACR3 default curve,
RawTherapee's adobe_camera_raw_default_curve copied value for value,
for every raw whose profile has no curve. Nothing renders through
either yet.
2026-10-03 14:22:46 -04:00
dtourolle b58873ef57 Spec baseline exposure, Camera Raw tone and profile sync (D21)
camera-profiles.md §11-§14 close what 0.20.0 left open. Baseline
exposure is the file's plus the profile's offset, applied as a gain on
the camera matrix, and a copied profile carries the DNG's baseline so a
CR2 lands at the same brightness. The view transform gains a Camera Raw
curve — the profile's ProfileToneCurve, else the ACR3 default — applied
Camera Raw's way, on the outer channels in linear ProPhoto, and it is
the default for every raw (D21, the user's choice). Profiles sync
through .darkroom-derived/profiles on the server as a step of the
derived sync pass, not as a catalog table.
2026-10-03 14:22:45 -04:00
dtourolle ababd628ed Show AI denoise in the manual, on a night frame with no faces
A section after Looking closer: what it is for, the switch and its wait,
Keep grain, which cameras it takes and where its noise figures come from.
The scene opens the Brooklyn Bridge at ISO 8000 from the face-free demo
set at 1:1, switches it on, waits for the result to land and keeps some
grain, with a still before and after. It waits on the app's own log line
rather than a fixed time: the network takes seconds on a GPU and more on
the CPU, which is where the recording X server leaves it (13 s).
2026-10-03 12:05:26 -04:00
dtourolle 4eb7cf77f5 Record what the learned denoise shipped as, and what was measured
The grain blend replaces the Amount of §7.2, and why its objection to a
blend does not hold for brightness alone; the Hexagon is out (int8 -6 to
-9 dB); §11 holds the data, the noise model taken from the library, the
model, the validation table, the blind estimate's reach and the speed.
2026-10-03 11:51:02 -04:00
dtourolle 960014803a Ship the denoise model in the Arch package, the APK and the Windows installer
Same LFS-pointer guard as the other models; the APK copies it out of its
assets with the rest.
2026-10-03 11:51:01 -04:00
dtourolle dd43f498fb Run the learned denoise in develop, and export with it
A Bayer photograph keeps its mosaic in the session and is offered the AI
Denoise switch. Asked for, the network runs on the decode executor from a
hot-pixel-repaired copy — the app's own pass — with the frame's noise from
its best source, and its progress in the activity bar; the classical
demosaic shows until the result lands, and the finished job says where the
noise figures came from. Keep grain is a GrainBlend of the two, made once
per value; the render draws it as its source and the adjust pass never
knows. demosaiced stays the classical result, so the raw histogram, the
white balance picker, masks and segmentation still read the sensor.

The develop view reconciles on a 250 ms poll rather than on each way an
edit can change (slider, undo, preset, version, a sidecar from another
device): two comparisons when nothing changed, and no path that can forget.
A failure is not retried until the switch is toggled. An export of a
photograph that asks for it waits for a running job or computes it.
2026-10-03 11:51:00 -04:00
dtourolle ad6bb892f3 Carry the learned denoise's switch and grain as edit settings
Whether to use the learned denoise, and how much grain to keep, are what a
photographer sets, so they travel the one road every setting does: published
as a capability, captured by Preset, stored in the sidecar, replayed by the
undo stack (FR-DEV-3c). Published only on a photograph that can take it, for
the lens switch's reason; the availability is derived from the file and is
not in the state. Off by default, grain 0; a reset returns both.
2026-10-03 11:22:29 -04:00
dtourolle 8ea3c3181a Upload the learned demosaic's result, and blend grain back into it
DemosaicedImage::from_rgb_f32 takes the network's linear camera RGB and
stands it beside the classical source of the same photograph: the matrix,
profile tables and as-shot balance are that source's, the id is new, so
nothing downstream can tell which demosaic ran and every cache keyed on the
source sees a new one.

GrainBlend is the denoise's live control. It returns only the brightness of
the noise the network removed, taken after the as-shot balance and handed
back divided by it, so the grain is neutral in the finished picture; colour
speckle and demosaic false colour stay out. It writes a new source rather
than adding a term to the adjust shader: the blend depends on two images and
one number, a 20 MP pass is milliseconds, and a fresh source id is all the
adjust pass's caches need. The test reads it back: at 0 the network's
result, at 1 the same white-balanced step in every channel.
2026-10-03 11:20:48 -04:00
dtourolle d8304d7c82 Add dr-denoise: the learned demosaic and denoise, without the UI
The noise model takes the best source the frame has: the body's measured
table (the Canon EOS 6D's, from the library), the DNG's NoiseProfile, or
the frame itself — read, row and column noise from its masked border, and
only the shot gain estimated, from the quietest flat patches. Checked on
130 6D frames, the estimate is within 10 % from ISO 1000 up; the network
loses under 0.3 dB for a sigma off by 15-20 %, so every Bayer body is
eligible.

Tiles of 1408 keep their central 1024 behind a 192-photosite halo, past the
185-photosite receptive field, and the frame is extended by reflection,
which keeps every photosite's colour; a pattern that starts on another
colour is read from one photosite up or left so the network sees RGGB, and
nothing is cropped. The tests run every Bayer phase, tiled against whole,
with a stand-in network of known reach.

The model ships as models/denoise/mosaic-1408.onnx (LFS), trained in
darkroom-denoise on the maintainer's own photographs, GPL like the code.
denoise_raw runs a file end to end: on a 6D frame at ISO 8000 the result
matches the training repository's own path to 2.5e-4 at worst, and takes
3.1 s on TensorRT fp16 (75 dB from f32) or 14.4 s on the CPU.
2026-10-03 11:15:50 -04:00
dtourolle 20b7bd7663 Feed every input a model declares when probing a rung
The probe built one zero tensor from the first input and ran the session
with it. Every model so far had one input; the denoiser has two (mosaic and
sigma), so every rung failed with "Missing Input: sigma" and the role was
left on the CPU: 14.4 s for a 20 MP frame where TensorRT fp16 takes 3.1 s.
Zeros now go to each input by name.
2026-10-03 11:15:49 -04:00
dtourolle 6b0d29cc15 Read a DNG's NoiseProfile
The converter's measured noise for the body at that ISO, (S, O) per CFA
plane: the learned denoise's best source for a body with no table of its
own (denoise.md §3.3). Read from the header beside the colour tags, and
printed by rawinfo. Checked against tifffile on a 6D DNG at ISO 5000: all
six values agree.
2026-10-03 10:39:15 -04:00
dtourolle eb91fa02c2 Give the inference engine a denoiser role, kept off the Hexagon
The learned demosaic-and-denoise (denoise.md) runs through the engine like
every other model. fp16 cost it nothing measurable (0.00 dB at every ISO on
validation tiles), so it takes TensorRT's and MIGraphX's fp16 like the
detectors. int8 cost it 6 to 9 dB, far past a 0.5 dB gate, so the Hexagon
refuses the role outright rather than relying on no int8 sibling existing,
and the tablet runs it on the CPU.
2026-10-03 10:39:14 -04:00
dtourolle d4248bc0dd Run the app's hot-pixel pass alone, and dump through it
The learned demosaic replaces the classical one and takes its input, the
mosaic hot_pixels.wgsl leaves (denoise.md §2), so its training data and its
input in the app must come through that pass and not a lookalike. The pass
was recorded inline in Demosaicer::run; it is now built by hot_pass and
recorded by record_hot_pass, which run still uses unchanged, and
Demosaicer::repair_hot_pixels runs it on its own and reads the mosaic back.

mosaic_dump moves to dr-gpu to call it, records how many photosites changed,
and keeps --unrepaired for a raw readout.
2026-10-03 10:25:28 -04:00
dtourolle 1f266a4478 Dump RAW mosaics for training the learned denoise
denoise.md §4.4 requires the training repo to read photosites through
dr-decode, not LibRaw, so black and white levels, the active area and the
CFA phase match what the app will feed the network. mosaic_dump reads
`input<TAB>prefix` lines and writes the whole readout as .npy plus a JSON
of what decode and metadata report. The masked border is kept: its
optically black photosites are a free dark frame for the noise profile.
2026-10-03 10:25:27 -04:00
dtourolle 2fad846cd1 Release 0.20.0
Benchmarks / CPU and I/O (per commit) (push) Successful in 9m1s
Benchmarks / Frame budget (on demand) (push) Skipped
Traceability / Requirement traces (push) Successful in 1m25s
Build and test / Android (aarch64) (push) Successful in 48m24s
Build and test / android-image (push) Successful in 4s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / Desktop (Linux) (push) Successful in 1h22m17s
Build and test / windows-image (push) Successful in 3s
🐳 Windows image / Build and push (push) Successful in 2s
Build and test / Layer separation (push) Successful in 37s
Build and test / Windows (x86_64, cross) (push) Successful in 55m30s
Build and test / Publish the release (push) Successful in 1m27s
2026-10-02 23:07:52 -04:00
dtourolle 5c26dc5033 Record what 0.20.0 closes and leaves open in FR-DEV-3e
The DCP half that outstanding.md listed as deferred is built (D20).
What stays open is the profiles directory, which does not sync, and
the profile tone curve and baseline exposure, which are read and not
applied — and which camera-profiles.md §1 measured as where the rest of
the gap to Lightroom's colour is.
2026-10-02 23:06:17 -04:00
dtourolle abefb94daa Give the export-ignores-the-viewport tests the viewport zoom now takes
0b06e31b ("Let a zoomed view fill the viewport...") added the
viewport's size to DevelopSession::zoom_about and left this
integration test calling it with three arguments, so dr-ui's tests did
not compile. The photograph here is 64×64; a 64×64 viewport keeps the
zoomed view the tests were written against.
2026-10-02 22:44:17 -04:00
dtourolle 54772f94d6 Correct what D20 claimed the profile tables would do for colour
Measured after building them, on four of the library's 6D DNGs: Adobe
Standard's tables lower mean saturation by 3-9 % at defaults, and the
look at 200 % lowers it further. The 6D's look table scales saturation
by 0.925 in its darkest value rows; it was tuned to sit under Camera
Raw's default RGB tone curve, which DarkRoom does not apply, and the
dark-tone desaturation is what is left without it. On _MG_9080 the
Lightroom preview measures 0.49, the matrix alone 0.38, the profile
0.35.

So the spec's "that gap is most of why the same file looks flatter"
was wrong: the gap is tone. The tables stay — they put each hue where
Adobe put it — and the spec, D20 and the Vivid file now say so.
"Stronger camera look" is removed: a stronger Adobe Standard look is a
less saturated picture, the opposite of its name. The Vivid presets,
measured at 0.40-0.46 on the same frame, are what answers "more
colourful" today.
2026-10-02 22:38:07 -04:00
dtourolle 65c1f1a468 Let the develop example render the profile off, the look doubled, or a preset
Diagnostic only. "matrix" switches the camera profile off and
"look200" doubles its look, so a DNG's tables can be judged against
the matrix render; "preset:<name>" applies a shipped preset as the
menu does. The example now renders through render_detailed, the path
every frontend takes, because a preset with clarity in it composes a
detail stage that plain render refuses.
2026-10-02 22:38:07 -04:00
dtourolle cb7ad0bbe7 Ship a Vivid section of presets
Five looks for "more colourful than the default": Vivid, Vivid strong,
Vivid landscape, Vivid warm and Vivid portrait. They lean on vibrance,
which lifts muted colours most and holds skin back, and use saturation
sparingly on top; landscape and portrait work the colour mixer's bands
so foliage and sky get richer while skin does not. A sixth, Stronger
camera look, pushes the camera profile's look table to 175 %, about the
step from Adobe Standard to Adobe Vivid, and only moves vibrance where
a photograph has no profile.

All change only what they name, so they keep a corrected exposure or
white balance, and the shipped-preset tests bound every key and value.
2026-10-02 22:38:07 -04:00
dtourolle eae720ce75 Say which camera profile a photograph renders through, and offer to copy it
The info panel gains a line under the lens: "Adobe Standard · in the
file", the .dcp it came from, "· off" when the photographer switched it
off, or "No camera profile · matrix only" — the ordinary case for a
CR2, worded as a fact rather than a failure. A DNG whose embedded
profile may be copied, for a body with no installed profile, also gets
"Use this profile for every Canon EOS 6D →", which saves it into the
profiles directory; the body's CR2s render through it from their next
decode.

The profiles directory is <data>/profiles, read at start-up on desktop
and Android before anything decodes. The library open path never set
the lens line; it now sets both. Labels: Camera Profile, Use Profile,
Look Amount.
2026-10-02 22:38:07 -04:00
dtourolle c02b401a9a Apply a camera profile's HueSatMap and LookTable after exposure
The second half of D20: a camera_profile scene operation at order 25
that converts working colour into linear ProPhoto, runs the DNG SDK's
HSV lookup through the HueSatMap and then the LookTable, and converts
back. Hue and saturation do not change under the uniform gains before
it, so a 2.5-D HueSatMap gives the same answer as straight after the
matrix, and the look sees the photographer's exposure as it does in
the SDK. Two departures for scene-referred values: value is not
clamped on the way out, and a colour outside ProPhoto passes through.

The operation holds only the switch (on by default) and a look
strength of 0-200 %. It is composed while the switch is on — a new
Operation::composes() separates "does something" from "moved from the
defaults", so an untouched raw renders through its profile and still
writes nothing. The tables come from the source: dr-gpu uploads the
ones DemosaicedImage carries into a storage buffer at @binding(8),
whose two-entry header tells the fragment whether there is anything to
apply, and binds a header of zeros for every other source.

apply_reference is the lookup on the CPU. The GPU test holds the
shader to it over 256 colours, through synthetic tables strong enough
that a wrong index shows, and through the library's real Adobe
Standard tables when the 6D DNG is present.
2026-10-02 22:38:06 -04:00
dtourolle f6a3f3f4e2 Read DCP camera profiles: embedded in a DNG, or a .dcp beside the app
The first half of D20. dr-decode now finds a camera profile's HueSatMap
and LookTable in the order camera-profiles.md §4 gives: the profile a
DNG embeds, then a .dcp in the profiles directory whose
UniqueCameraModel names the body, then none. A .dcp brings its own
matrices, since its tables were measured against its forward matrix.

The HueSatMap is blended for the frame's colour temperature with the
same mired weight the matrices use, once per decode, and the result
rides on RawImage as profile_tables beside color_matrix, so every path
that renders a decoded file gets the same profile without a setter to
forget. Nothing applies the tables yet.

A profile whose embed policy allows copying can be written back out as
a .dcp (rawler's TIFF writer with the RC magic patched in), which is how
the library's 6D CR2s will get the Adobe Standard their DNGs carry. The
table type lives in dr-types because decode, pipeline and GPU all need
its layout. Tests read the library's 6D DNG when it is present.
2026-10-02 22:38:06 -04:00
dtourolle 05ac2416c6 Spec DCP camera profiles (D20)
The library's Canon 6D DNGs were written by Lightroom and embed Adobe
Standard with its HueSatMap and LookTable; DarkRoom renders them through
the matrix alone, which is most of why the same file looks flatter here
than in Lightroom.

camera-profiles.md designs the deferred half of FR-DEV-3e: the tables
applied by a camera_profile scene operation after exposure, the profile
taken from the DNG or from a matched .dcp, tables carried with the
decoded image like the matrix, a look-strength control, and copying an
embedded profile out where its policy allows. FR-DEV-3e gains item 4
and D20 records the placement and what was rejected.
2026-10-02 22:37:58 -04:00
dtourolle 0b06e31bf3 Let a zoomed view fill the viewport rather than keep the photograph's shape
The view was the same fraction of each axis, so it kept the frame's
aspect at every zoom: a portrait zoomed on a landscape screen stayed a
portrait strip with the screen's sides empty. Each axis now shows as
much of the frame as the viewport holds at that magnification, capped
at the whole frame, and the render is fitted to the viewed region
rather than to the frame. A redraw re-cuts a zoomed view about its
centre when the viewport or the crop changes shape.
2026-10-02 22:29:48 -04:00
dtourolle d5c93ae795 Step to a library photograph without flashing frames in between
Benchmarks / Frame budget (on demand) (push) Canceled after 0s
Benchmarks / CPU and I/O (per commit) (push) Canceled after 5s
Traceability / Requirement traces (push) Canceled after 0s
Build and test / Desktop (Linux) (push) Successful in 1h21m7s
Build and test / Layer separation (push) Successful in 46s
🐳 Android image / Build and push (push) Successful in 4s
Build and test / android-image (push) Successful in 5s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 2s
Build and test / Android (aarch64) (push) Successful in 47m42s
Build and test / Windows (x86_64, cross) (push) Successful in 54m42s
Build and test / Publish the release (push) Successful in 52s
A step along the roll showed up to four pictures: the grid thumbnail,
the previous photograph again (the thumbnail was dropped when the bytes
landed, while the canvas still held the last texture through the decode
and first render), the new one at its defaults when a cached original
beat the sidecar, and then its edit.

The thumbnail now stays up until render_now draws the new photograph's
first frame, and that first frame waits up to 400 ms for the stored edit
before drawing at the defaults. A later arrival still redraws.
2026-10-02 20:46:05 -04:00
dtourolle 379dd1afcc Keep a late sidecar off the next photograph
A stored edit that arrived after the view had stepped on was applied to
whatever session was open by then — the next photograph's. The wait now
stops once the open it belongs to is no longer the current one.
2026-10-02 20:44:43 -04:00
dtourolle 825c5af20a Release 0.19.4
Benchmarks / Frame budget (on demand) (push) Canceled after 0s
Benchmarks / CPU and I/O (per commit) (push) Canceled after 5m49s
Traceability / Requirement traces (push) Canceled after 0s
Build and test / android-image (push) Canceled after 0s
🐳 Android image / Build and push (push) Canceled after 0s
Build and test / Android (aarch64) (push) Canceled after 0s
Build and test / windows-image (push) Canceled after 0s
🐳 Windows image / Build and push (push) Canceled after 0s
Build and test / Windows (x86_64, cross) (push) Canceled after 0s
Build and test / Layer separation (push) Canceled after 0s
Build and test / Publish the release (push) Canceled after 0s
Build and test / Desktop (Linux) (push) Canceled after 17s
2026-10-02 20:40:01 -04:00
dtourolle 23a2f13b46 Apply the collision policy to exports bound for the server
A queued export could not see the server, so its name check always
answered "free" and the upload PUT over whatever was there: Increment
and Skip behaved as Overwrite on Nextcloud, and two exports of the same
name queued before either uploaded landed on one file.

The batch now names around what the album records of earlier exports
and what the outbox already holds for that folder. The outbox record
carries the policy, and the drain lists each destination folder once
and applies it against what the server holds: Increment steps past a
taken name and re-points the album's row, Skip drops the entry. A
record without a policy (older builds, a merge's composite) is sent as
named, as before. The album is recorded before the drain starts so a
rename has a row to move.
2026-10-02 19:40:27 -04:00
dtourolle 2c947430e6 Release 0.19.3
Benchmarks / CPU and I/O (per commit) (push) Successful in 5m26s
Benchmarks / Frame budget (on demand) (push) Skipped
Traceability / Requirement traces (push) Successful in 1m9s
Build and test / Android (aarch64) (push) Successful in 30m44s
Build and test / android-image (push) Successful in 3s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / Desktop (Linux) (push) Successful in 49m45s
Build and test / windows-image (push) Successful in 1s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / Layer separation (push) Successful in 35s
Build and test / Windows (x86_64, cross) (push) Successful in 36m12s
Build and test / Publish the release (push) Successful in 1m12s
2026-09-30 22:46:17 -04:00
dtourolle 36556729f1 Let Back close the presets menu instead of the application
The presets menu at the foot of the tool rail is a PopupWindow, and
showing a popup takes focus off the develop view until it closes. The
menu held nothing focusable, so Android's Back gesture, pressed to
dismiss it, found no focus item, went unanswered, and the platform
closed the application. Slint closes a popup on Escape by itself but
not on Back.

The menu now holds a key scope, as the film list does, that closes it
on Back or Escape. The next Back leaves develop for the grid.
2026-09-30 22:11:51 -04:00
dtourolle 6f33517b35 Cut panorama overlaps along seams instead of averaging them
The merge weighted every overlap pixel by its distance from each frame's
edge, a 200 px linear cross-fade. Anything the frames disagreed on —
parallax in the near foreground, grass in the wind, a walker — came out
twice at half strength: a soft double edge at 1:1.

dr_pano::seam picks, per output texel at proxy resolution, which frame a
pixel comes from. Where a new frame overlaps the composite the cost is the
gain-corrected difference plus local detail plus nearness to either
frame's edge, taken as the worst over a small window, and the cut is a
dynamic-programming path across the overlap. merge.wgsl weights each frame
by its tent-filtered share of that map, a 64 px blend that follows the
seam, with the edge feather kept as the fallback. The page's preview uses
the same map, and examples/merge.rs takes --feather-only for comparison.
2026-09-30 21:57:44 -04:00
dtourolle 1d7115437b Date a photograph from its name when its header has none
WhatsApp strips every EXIF tag and names the file "WhatsApp Image
2023-06-15 at 07.00.42.jpeg"; Windows Phone, Android cameras and
darktable's import put the date in the name too. Those images sorted
after everything else and were absent from the timeline.

name_dates reads a date (and a time, when one follows) from the file
name, then from the innermost folder that states one. A sequence number
after a date is not read as a time, and a bare year folder is not a date.
EXIF always wins: only examined rows still undated are filled.

The sweep and the metadata repair fill as they mark an image examined,
and the open backfill fills catalogs examined by earlier builds. On the
reference library that takes 274 undated images to 10; the no-op case is
a seek on images_captured, 0.6 ms an open.
2026-09-30 21:45:54 -04:00
dtourolle caae65c78d Import from an SD card or card reader on Android
Import was switched off on Android: `imports_supported` was true only for
`target_os = "linux"`, and its comment said Android has no path to read a
card by and nowhere to write the copies. Neither holds. With "all files
access" (MANAGE_EXTERNAL_STORAGE, API 30) an app reads the root of an SD
card or a USB card reader by path, `/storage/9C33-6BBD`, and the importer
only ever writes into its own staging directory, which is a plain
directory on Android too. So the engine runs unchanged; what was missing
was finding the card and the permission.

- The manifest declares MANAGE_EXTERNAL_STORAGE, and
  READ_EXTERNAL_STORAGE up to API 29 with requestLegacyExternalStorage,
  which is the same access on 28 and 29.
- Cards.java lists the mounted non-primary volumes through
  StorageManager and opens the system "All files access" page for this
  app. dr_ui::cards is the JNI bridge, through saf's helpers.
- The import page on Android asks for the permission with an "Allow
  access" button until it has it, rather than showing an empty list that
  reads as "no card", and watches for the grant so the list fills in when
  the user comes back from settings.

Google Play restricts this permission to file managers and the like;
DarkRoom is sideloaded, so that does not apply.
2026-09-30 21:30:09 -04:00
172 changed files with 15599 additions and 1057 deletions
+11 -4
View File
@@ -487,10 +487,11 @@ jobs:
wine "$SETUP" /S 2>/dev/null
INST=$(echo "$HOME"/.wine/drive_c/users/*/AppData/Local/Programs/DarkRoom)
ls "$INST"
# As many files as package.sh stages: everything but the READMEs in
# the directories it copies. A literal here went stale the first
# time a model was added.
WANT=$(find models/face models/scene models/inpaint -maxdepth 1 -type f ! -name README.md | wc -l)
# As many files as package.sh stages: everything but the READMEs and
# the Hexagon's quantised siblings in the directories it copies. A
# literal here went stale the first time a model was added.
WANT=$(find models/face models/scene models/inpaint models/denoise -maxdepth 1 -type f ! -name README.md \
! -name '*.int8.onnx' ! -name '*.a16w8.onnx' ! -name '*.a16w16.onnx' | wc -l)
GOT=$(ls "$INST/models" | wc -l)
[ "$GOT" = "$WANT" ] || { echo "FAIL: expected $WANT model files, installed $GOT"; exit 1; }
# The manual, and every picture it shows, counted the same way.
@@ -498,6 +499,12 @@ jobs:
WANT=$(ls docs/manual/media | wc -l)
GOT=$(ls "$INST/manual/media" | wc -l)
[ "$GOT" = "$WANT" ] || { echo "FAIL: expected $WANT manual pictures, installed $GOT"; exit 1; }
# Both bundled runtimes, each with its provider beside it
# (tools/fetch-bundled-runtimes.sh).
for f in openvino/onnxruntime.dll openvino/onnxruntime_providers_openvino.dll \
openvino/openvino.dll webgpu/onnxruntime.dll webgpu/dxcompiler.dll; do
[ -f "$INST/runtimes/$f" ] || { echo "FAIL: runtimes/$f not installed"; exit 1; }
done
wine reg query 'HKCU\Software\Microsoft\Windows\CurrentVersion\Uninstall\DarkRoom' 2>/dev/null \
| grep -q DisplayVersion || { echo "FAIL: no uninstall registry key"; exit 1; }
wine "$INST/darkroom.exe" --version 2>/dev/null | grep -q '^darkroom-desktop ' \
Generated
+44 -25
View File
@@ -1265,7 +1265,7 @@ checksum = "f27ae1dd37df86211c42e150270f82743308803d90a6f6e6651cd730d5e1732f"
[[package]]
name = "darkroom-android"
version = "0.19.2"
version = "0.24.0"
dependencies = [
"android_logger",
"dr-plat",
@@ -1278,7 +1278,7 @@ dependencies = [
[[package]]
name = "darkroom-desktop"
version = "0.19.2"
version = "0.24.0"
dependencies = [
"anyhow",
"dr-plat",
@@ -1454,7 +1454,7 @@ checksum = "d8b14ccef22fc6f5a8f4d7d768562a182c04ce9a3b3157b91390b52ddfdf1a76"
[[package]]
name = "dr-bench"
version = "0.19.2"
version = "0.24.0"
dependencies = [
"anyhow",
"dr-catalog",
@@ -1471,7 +1471,7 @@ dependencies = [
[[package]]
name = "dr-catalog"
version = "0.19.2"
version = "0.24.0"
dependencies = [
"dr-face",
"dr-plat",
@@ -1486,7 +1486,7 @@ dependencies = [
[[package]]
name = "dr-decode"
version = "0.19.2"
version = "0.24.0"
dependencies = [
"dr-types",
"env_logger",
@@ -1498,9 +1498,26 @@ dependencies = [
"zune-jpeg 0.4.21",
]
[[package]]
name = "dr-denoise"
version = "0.24.0"
dependencies = [
"dr-decode",
"dr-gpu",
"dr-inference-engine",
"env_logger",
"log",
"ndarray",
"ort",
"pollster",
"serde",
"serde_norway",
"thiserror 2.0.20",
]
[[package]]
name = "dr-export"
version = "0.19.2"
version = "0.24.0"
dependencies = [
"dr-decode",
"dr-gpu",
@@ -1519,7 +1536,7 @@ dependencies = [
[[package]]
name = "dr-face"
version = "0.19.2"
version = "0.24.0"
dependencies = [
"dr-inference-engine",
"env_logger",
@@ -1532,7 +1549,7 @@ dependencies = [
[[package]]
name = "dr-film"
version = "0.19.2"
version = "0.24.0"
dependencies = [
"log",
"serde",
@@ -1541,7 +1558,7 @@ dependencies = [
[[package]]
name = "dr-gpu"
version = "0.19.2"
version = "0.24.0"
dependencies = [
"bytemuck",
"dr-decode",
@@ -1559,7 +1576,7 @@ dependencies = [
[[package]]
name = "dr-inference-engine"
version = "0.19.2"
version = "0.24.0"
dependencies = [
"env_logger",
"libloading",
@@ -1574,7 +1591,7 @@ dependencies = [
[[package]]
name = "dr-ingest"
version = "0.19.2"
version = "0.24.0"
dependencies = [
"dr-plat",
"dr-types",
@@ -1586,7 +1603,7 @@ dependencies = [
[[package]]
name = "dr-lens"
version = "0.19.2"
version = "0.24.0"
dependencies = [
"lensfun",
"log",
@@ -1594,7 +1611,7 @@ dependencies = [
[[package]]
name = "dr-pano"
version = "0.19.2"
version = "0.24.0"
dependencies = [
"dr-decode",
"dr-inference-engine",
@@ -1608,7 +1625,7 @@ dependencies = [
[[package]]
name = "dr-pipeline"
version = "0.19.2"
version = "0.24.0"
dependencies = [
"dr-types",
"log",
@@ -1617,7 +1634,7 @@ dependencies = [
[[package]]
name = "dr-plat"
version = "0.19.2"
version = "0.24.0"
dependencies = [
"android-native-keyring-store",
"dr-types",
@@ -1633,7 +1650,7 @@ dependencies = [
[[package]]
name = "dr-preset-xmp"
version = "0.19.2"
version = "0.24.0"
dependencies = [
"dr-pipeline",
"log",
@@ -1643,7 +1660,7 @@ dependencies = [
[[package]]
name = "dr-segment"
version = "0.19.2"
version = "0.24.0"
dependencies = [
"dr-inference-engine",
"env_logger",
@@ -1656,7 +1673,7 @@ dependencies = [
[[package]]
name = "dr-sync"
version = "0.19.2"
version = "0.24.0"
dependencies = [
"async-trait",
"dr-plat",
@@ -1670,7 +1687,7 @@ dependencies = [
[[package]]
name = "dr-sync-folder"
version = "0.19.2"
version = "0.24.0"
dependencies = [
"async-trait",
"dr-sync",
@@ -1682,7 +1699,7 @@ dependencies = [
[[package]]
name = "dr-sync-nextcloud"
version = "0.19.2"
version = "0.24.0"
dependencies = [
"async-trait",
"dr-decode",
@@ -1704,7 +1721,7 @@ dependencies = [
[[package]]
name = "dr-thumbs"
version = "0.19.2"
version = "0.24.0"
dependencies = [
"dr-types",
"jpeg-encoder",
@@ -1716,7 +1733,7 @@ dependencies = [
[[package]]
name = "dr-types"
version = "0.19.2"
version = "0.24.0"
dependencies = [
"serde",
"serde_json",
@@ -1725,12 +1742,13 @@ dependencies = [
[[package]]
name = "dr-ui"
version = "0.19.2"
version = "0.24.0"
dependencies = [
"anyhow",
"async-trait",
"dr-catalog",
"dr-decode",
"dr-denoise",
"dr-export",
"dr-face",
"dr-film",
@@ -1750,6 +1768,7 @@ dependencies = [
"dr-types",
"dr-xmp",
"env_logger",
"half",
"i-slint-backend-testing",
"jni 0.22.4",
"log",
@@ -1773,7 +1792,7 @@ dependencies = [
[[package]]
name = "dr-xmp"
version = "0.19.2"
version = "0.24.0"
dependencies = [
"dr-types",
"log",
@@ -7107,7 +7126,7 @@ checksum = "8df9b6e13f2d32c91b9bd719c00d1958837bc7dec474d94952798cc8e69eeec3"
[[package]]
name = "traceability"
version = "0.19.2"
version = "0.24.0"
dependencies = [
"anyhow",
"proc-macro2",
+11 -1
View File
@@ -5,6 +5,7 @@ members = [
"core/dr-catalog",
"core/dr-thumbs",
"core/dr-decode",
"core/dr-denoise",
"core/dr-export",
"core/dr-face",
"core/dr-film",
@@ -32,7 +33,7 @@ members = [
exclude = ["third_party"]
[workspace.package]
version = "0.19.2"
version = "0.24.0"
edition = "2021"
rust-version = "1.92"
license = "GPL-3.0-or-later"
@@ -44,6 +45,7 @@ dr-types = { path = "core/dr-types" }
dr-catalog = { path = "core/dr-catalog" }
dr-thumbs = { path = "core/dr-thumbs" }
dr-decode = { path = "core/dr-decode" }
dr-denoise = { path = "core/dr-denoise" }
dr-export = { path = "core/dr-export" }
# Stated explicitly for the same reason as `dr-segment` below: no dependant
# should drag in an ONNX runtime by accident. Members opt in with
@@ -276,6 +278,14 @@ opt-level = 0
lto = "thin"
codegen-units = 1
# Except dr-ui. Slint expands the `.slint` files into ~27 MB of Rust
# (`out/app.rs`), and at one codegen unit LLVM optimises all of it on a single
# thread: 13.5 minutes of a release build with the other cores idle. The code
# it holds is UI glue — property bindings and callbacks — not the image work,
# which lives in the crates above that keep the single unit.
[profile.release.package.dr-ui]
codegen-units = 16
# A release build that can say where it panicked: line tables, so a crash
# record's backtrace (`dr_plat::crash`) reads `file.rs:123` rather than bare
# addresses. The macOS build uses it (docs/dev/macos.md) — no one here can
+1 -1
View File
@@ -201,7 +201,7 @@ controls, its place in the chain and its tests.
## Where it stands
**0.19.2**, thirty-one tagged releases in. 193 numbered requirements in
**0.24.0**, thirty-nine tagged releases in. 193 numbered requirements in
scope, 85% of them claimed by code and [traced to it](docs/dev/traceability.md);
the rest are written down rather than merely absent.
@@ -4,9 +4,10 @@
Deliberately minimal: this packages the viewer for on-device testing (spike
S2 needs Adreno and Mali hardware, which no emulator represents). Nothing
here is a distribution manifest yet. Only network access is declared: file
access needs no manifest permission because the library grid reads through
SAF, which grants per-tree at runtime (ARCH §6.9).
here is a distribution manifest yet. The library grid needs no storage
permission, because it reads through SAF, which grants per-tree at runtime
(ARCH §6.9); the one storage permission declared is for importing from a
camera card, which is read by path.
Minimal is not the same as empty, and the entries below that are not the
activity are the difference. A manifest is the only place a component can be
@@ -21,13 +22,29 @@
WebDAV listing, thumbnail and image fetches. Without it Android refuses
socket creation outright, and the failure is invisible — no panic to
catch, no log line, just a worker thread that stops. Storage is the
separate case that genuinely needs no permission here, because SAF
grants per-tree at runtime (ARCH §6.9). -->
separate case: the library and album folders need no permission
here, because SAF grants per-tree at runtime (ARCH §6.9). -->
<uses-permission android:name="android.permission.INTERNET" />
<!-- Read before deciding whether a sync may run: FR-NC-6 gates background
work on unmetered-and-charging, which means knowing the network type. -->
<uses-permission android:name="android.permission.ACCESS_NETWORK_STATE" />
<!-- FR-CAT-10: importing from a camera card. The importer reads the card
as files, and "all files access" is what makes an SD card or a USB
card reader readable by path on API 30 and up (see Cards.java). It is
granted on a system settings page, not a dialog; the import page
sends the user there when it is missing. READ_EXTERNAL_STORAGE is the
same thing for API 28 and 29, and means nothing above them; on 29 it
reads by path only with requestLegacyExternalStorage, which is why
<application> carries that flag.
Google Play limits MANAGE_EXTERNAL_STORAGE to a short list of app
kinds. DarkRoom is not distributed through Play. -->
<uses-permission android:name="android.permission.MANAGE_EXTERNAL_STORAGE" />
<uses-permission
android:name="android.permission.READ_EXTERNAL_STORAGE"
android:maxSdkVersion="29" />
<!-- Vulkan 1.1 is what wgpu needs; the API 28 floor is where support is
dependable (NFR-COMPAT-1). Marked required so an unsupported device
fails at install rather than at first frame. -->
@@ -53,8 +70,17 @@
android:icon="@mipmap/ic_launcher"
android:hasCode="true"
android:allowBackup="false"
android:requestLegacyExternalStorage="true"
android:supportsRtl="true">
<!-- The DSP's RPC library, which QNN's Hexagon stub loads. From API 31
an app's linker namespace refuses a vendor library the manifest
does not name, and QNN then fails to create its device
(QNN_DEVICE_ERROR_INVALID_CONFIG) before it reaches the DSP:
every model ran on the CPU on 0.22.0. Not required, so a device
without one still installs and stays on the CPU. -->
<uses-native-library android:name="libcdsprpc.so" android:required="false" />
<!-- NativeActivity rather than a Kotlin Activity: android-activity's
glue loads libdarkroom.so and calls android_main. `android.app.lib_name`
is how it learns which library to load, and must match [lib].name.
@@ -0,0 +1,150 @@
package paris.tourolle.darkroom;
import android.Manifest;
import android.content.Context;
import android.content.Intent;
import android.content.pm.PackageManager;
import android.net.Uri;
import android.os.Build;
import android.os.Environment;
import android.os.storage.StorageManager;
import android.os.storage.StorageVolume;
import android.provider.Settings;
import android.util.Log;
import java.io.File;
import java.util.ArrayList;
import java.util.List;
/**
* Finding a camera card, and the permission that makes it readable (FR-CAT-10).
*
* <p>An import reads the card as files: the survey walks it, the probe reads
* each header and the copy streams each original, all through the same
* {@code std::fs} code the desktop uses. Android hands out such paths —
* {@code /storage/9C33-6BBD/DCIM} — to an app holding "all files access"
* ({@code MANAGE_EXTERNAL_STORAGE}, API 30), which covers the root of an SD
* card and of a USB card reader. Below API 30 the same paths are readable
* with {@code READ_EXTERNAL_STORAGE}.
*
* <p>Not the folder picker {@link FolderPicker} uses for albums. A tree
* granted through SAF is {@code content://} URIs, not paths, and since API 30
* the picker refuses the root of a card outright; reading a card through it
* would mean a second storage implementation under the importer, where this
* needs none.
*
* <p>Google Play restricts this permission to file managers and the like.
* DarkRoom is not distributed through Play, so the restriction does not
* apply; it would need revisiting if that changed.
*/
public final class Cards {
private static final String TAG = "DarkRoom";
private Cards() {
}
/** Whether this app may read a card's files by path. */
public static boolean hasAccess(Context context) {
if (Build.VERSION.SDK_INT >= Build.VERSION_CODES.R) {
return Environment.isExternalStorageManager();
}
return context.checkSelfPermission(Manifest.permission.READ_EXTERNAL_STORAGE)
== PackageManager.PERMISSION_GRANTED;
}
/**
* Open the system page where the user grants it.
*
* <p>A settings page rather than a permission dialog because there is no
* dialog for this one on API 30 and up: the user flips "Allow access to
* manage all files" for this app. Below 30 the context is the application
* context, which cannot raise a runtime permission request (that needs an
* Activity's result), so the app's own settings page is the route there
* too. Either way the app learns of the grant by asking
* {@link #hasAccess} again.
*/
public static void requestAccess(Context context) {
Uri self = Uri.parse("package:" + context.getPackageName());
Intent intent;
if (Build.VERSION.SDK_INT >= Build.VERSION_CODES.R) {
intent = new Intent(Settings.ACTION_MANAGE_APP_ALL_FILES_ACCESS_PERMISSION, self);
} else {
intent = new Intent(Settings.ACTION_APPLICATION_DETAILS_SETTINGS, self);
}
// The context is not an Activity; see FolderPicker.start.
intent.addFlags(Intent.FLAG_ACTIVITY_NEW_TASK);
try {
context.startActivity(intent);
} catch (RuntimeException e) {
// Some builds ship without the per-app page; the list of every
// app holding the permission is the fallback that always exists.
Log.w(TAG, "no per-app all-files page; opening the list", e);
if (Build.VERSION.SDK_INT >= Build.VERSION_CODES.R) {
Intent list = new Intent(Settings.ACTION_MANAGE_ALL_FILES_ACCESS_PERMISSION);
list.addFlags(Intent.FLAG_ACTIVITY_NEW_TASK);
context.startActivity(list);
}
}
}
/**
* Every mounted volume other than the device's own storage.
*
* <p>One string per volume, {@code path \t description \t removable},
* where removable is {@code 1} or {@code 0}: the reason {@link Intents}
* gives for keeping the JNI surface to strings. The primary volume is left
* out — it is the device's internal storage, never a card — and so is
* anything not mounted, which is a card being ejected or one the system
* could not read.
*/
public static String[] volumes(Context context) {
List<String> out = new ArrayList<String>();
StorageManager manager = (StorageManager) context.getSystemService(Context.STORAGE_SERVICE);
if (manager == null) {
return new String[0];
}
for (StorageVolume volume : manager.getStorageVolumes()) {
if (volume.isPrimary()) {
continue;
}
String state = volume.getState();
if (!Environment.MEDIA_MOUNTED.equals(state)
&& !Environment.MEDIA_MOUNTED_READ_ONLY.equals(state)) {
continue;
}
String path = path(volume);
if (path == null) {
Log.w(TAG, "a mounted volume with no path: " + volume);
continue;
}
String description = volume.getDescription(context);
if (description == null) {
description = new File(path).getName();
}
out.add(path + "\t" + description.replace('\t', ' ') + "\t"
+ (volume.isRemovable() ? "1" : "0"));
}
return out.toArray(new String[0]);
}
/**
* Where the volume is mounted.
*
* <p>{@code getDirectory} is API 30. Below it the same answer is the
* hidden {@code getPath}, which every release from 24 to 29 has, reached by
* reflection because android.jar does not declare it.
*/
private static String path(StorageVolume volume) {
if (Build.VERSION.SDK_INT >= Build.VERSION_CODES.R) {
File dir = volume.getDirectory();
return dir == null ? null : dir.getPath();
}
try {
Object path = StorageVolume.class.getMethod("getPath").invoke(volume);
return path == null ? null : path.toString();
} catch (ReflectiveOperationException e) {
Log.w(TAG, "StorageVolume.getPath", e);
return null;
}
}
}
+53 -11
View File
@@ -332,27 +332,38 @@ fn unpack_bundled_models(app: &slint::android::AndroidApp) {
// eyes-open filter has something to read, and a tablet has no other way
// to get them either.
//
// The int8 forms beside the three detectors are what the Hexagon runs
// (docs/dev/inference.md §5); the engine loads the sibling when the probe
// chose that rung and ignores it otherwise.
const BUNDLED: [(&std::ffi::CStr, &str); 14] = [
// The quantised siblings — `.a16w8.onnx`, `.a16w16.onnx` — are what the
// Hexagon runs (docs/dev/inference.md §1.5), each in the narrowest form
// that held that model's accuracy on the tablet; the engine loads the
// sibling when the probe chose that rung and ignores it otherwise. The
// segmenter's and XFeat's forms are compiled into the binary instead,
// beside their f32 graphs.
const BUNDLED: [(&std::ffi::CStr, &str); 21] = [
(c"models/scrfd_500m_640.onnx", "scrfd_500m_640.onnx"),
(
c"models/scrfd_500m_640.int8.onnx",
"scrfd_500m_640.int8.onnx",
c"models/scrfd_500m_640.a16w8.onnx",
"scrfd_500m_640.a16w8.onnx",
),
(c"models/scrfd_2.5g_640.onnx", "scrfd_2.5g_640.onnx"),
(
c"models/scrfd_2.5g_640.int8.onnx",
"scrfd_2.5g_640.int8.onnx",
c"models/scrfd_2.5g_640.a16w8.onnx",
"scrfd_2.5g_640.a16w8.onnx",
),
(c"models/scrfd_10g_640.onnx", "scrfd_10g_640.onnx"),
(c"models/scrfd_10g_640.int8.onnx", "scrfd_10g_640.int8.onnx"),
(
c"models/scrfd_10g_640.a16w8.onnx",
"scrfd_10g_640.a16w8.onnx",
),
(c"models/arcface_mbf_b1.onnx", "arcface_mbf_b1.onnx"),
(c"models/2d106det_b1.onnx", "2d106det_b1.onnx"),
(c"models/2d106det_b1.a16w8.onnx", "2d106det_b1.a16w8.onnx"),
(c"models/ocec_s_b1.onnx", "ocec_s_b1.onnx"),
(c"models/sgc_l_48_b1.onnx", "sgc_l_48_b1.onnx"),
(c"models/yolo26s-sem-ade20k.onnx", "yolo26s-sem-ade20k.onnx"),
(
c"models/yolo26s-sem-ade20k.a16w16.onnx",
"yolo26s-sem-ade20k.a16w16.onnx",
),
(
c"models/yolo26s-sem-ade20k.classes.json",
"yolo26s-sem-ade20k.classes.json",
@@ -360,6 +371,19 @@ fn unpack_bundled_models(app: &slint::android::AndroidApp) {
(c"models/categories.txt", "categories.txt"),
// The panorama border filler (FR-MRG-4); MIT, 28 MB.
(c"models/migan-512.onnx", "migan-512.onnx"),
(c"models/migan-512.a16w16.onnx", "migan-512.a16w16.onnx"),
// The learned demosaic and denoise, one network per method
// (FR-DEV-3g), each with the 16-bit form the Hexagon runs.
(c"models/mosaic-fast-1408.onnx", "mosaic-fast-1408.onnx"),
(
c"models/mosaic-fast-1408.a16w16.onnx",
"mosaic-fast-1408.a16w16.onnx",
),
(c"models/mosaic-hq-1408.onnx", "mosaic-hq-1408.onnx"),
(
c"models/mosaic-hq-1408.a16w16.onnx",
"mosaic-hq-1408.a16w16.onnx",
),
];
let dir = dr_ui::shared_face_models_dir();
@@ -422,8 +446,13 @@ fn unpack_bundled_models(app: &slint::android::AndroidApp) {
// on a first launch they were not on disk until this line. The runtime
// is in the APK's native library directory beside `libdarkroom.so`,
// which is also where Qualcomm's DSP loader has to be pointed for the
// Hexagon skel (docs/dev/inference.md §3, §8).
dr_ui::inference::init(native_library_dir().into_iter().collect());
// Hexagon skel (docs/dev/inference.md §3, §8). Two of them: the QNN
// build, and the generic WebGPU build by its file name, which the
// engine opens only when the first does not fit the SoC (§3.2).
let runtimes = native_library_dir()
.map(|dir| vec![dir.clone(), dir.join("libonnxruntime_generic.so")])
.unwrap_or_default();
dr_ui::inference::init(runtimes);
}
/// The directory the system unpacked this APK's native libraries into.
@@ -509,6 +538,19 @@ mod tests {
Some(value.to_string())
}
/// From API 31 the linker refuses a vendor library the manifest does not
/// name, and QNN cannot create its Hexagon device without the DSP's RPC
/// library: 0.22.0 ran every model on the CPU for want of this line.
#[test]
fn the_npu_can_reach_the_dsp() {
assert!(
manifest().contains(
r#"<uses-native-library android:name="libcdsprpc.so" android:required="false" />"#
),
"libcdsprpc.so must be declared, and not required"
);
}
#[test]
fn a_gallery_can_open_a_photograph_in_this_app() {
let manifest = manifest();
+20 -7
View File
@@ -106,24 +106,37 @@ fn main() -> anyhow::Result<()> {
/// providers, or against the wrong cuDNN — and a system copy whose providers
/// do not load is not a problem, only a slower app: the probe builds a real
/// session before believing a provider.
///
/// The order breaks ties only. The engine opens every runtime on this list
/// and loads the one whose providers fit the GPU (inference.md §3.2), so a
/// package's bundled builds — `runtimes/openvino` and `runtimes/webgpu`
/// beside each place a package installs to, from
/// `tools/fetch-bundled-runtimes.sh` — sit beside a CUDA or ROCm runtime
/// without hiding it.
fn runtime_dirs() -> Vec<PathBuf> {
// A place a package installs to, and the bundled runtimes under it.
fn packaged(dirs: &mut Vec<PathBuf>, base: PathBuf) {
dirs.push(base.join("runtimes/openvino"));
dirs.push(base.join("runtimes/webgpu"));
dirs.push(base);
}
let mut dirs = Vec::new();
if let Some(dir) = std::env::var_os("DARKROOM_ORT_DIR") {
dirs.push(PathBuf::from(dir));
}
if let Ok(exe) = std::env::current_exe() {
if let Some(bin) = exe.parent() {
dirs.push(bin.to_path_buf());
dirs.push(bin.join("../lib/darkroom"));
packaged(&mut dirs, bin.to_path_buf());
packaged(&mut dirs, bin.join("../lib/darkroom"));
}
}
dirs.push(dr_ui::inference::user_runtime_dir());
#[cfg(target_os = "linux")]
dirs.extend([
PathBuf::from("/app/lib/darkroom"),
PathBuf::from("/usr/lib/darkroom"),
PathBuf::from("/usr/lib"),
]);
{
packaged(&mut dirs, PathBuf::from("/app/lib/darkroom"));
packaged(&mut dirs, PathBuf::from("/usr/lib/darkroom"));
dirs.push(PathBuf::from("/usr/lib"));
}
// An app bundle keeps its libraries in `Contents/Frameworks`, beside
// the `Contents/MacOS` the executable is in; then Homebrew's
// `onnxruntime`, Apple silicon's prefix before Intel's. Homebrew's build
+4 -1
View File
@@ -27,7 +27,7 @@
use std::path::PathBuf;
use std::time::{Duration, Instant};
use dr_catalog::{keywords, rating, schema, Catalog};
use dr_catalog::{keywords, name_dates, rating, schema, Catalog};
fn main() {
let mut args: Vec<String> = std::env::args().skip(1).collect();
@@ -108,6 +108,9 @@ fn main() {
time(" keywords::adopt_orphan_terms", 20, || {
keywords::adopt_orphan_terms(conn).unwrap();
});
time(" name_dates::fill", 20, || {
name_dates::fill(conn, None).unwrap();
});
interactive(conn);
+64
View File
@@ -306,6 +306,48 @@ pub fn record_exports(
Ok(())
}
/// Every file name an album records, for an export choosing a name to know
/// what it would land on.
///
/// A server album cannot be asked while the export is queued offline, and
/// the names this app put there are the ones a second export of the same
/// photographs will collide with. One read of the album's rows, not one per
/// candidate name.
pub fn file_names(
conn: &Connection,
id: AlbumId,
) -> Result<std::collections::HashSet<String>, CatalogError> {
ensure_tables(conn)?;
let mut stmt = conn.prepare("SELECT file_name FROM album_exports WHERE album_id = ?1")?;
let rows = stmt
.query_map([id.0 as i64], |r| r.get(0))?
.collect::<Result<_, _>>()?;
Ok(rows)
}
/// A file the upload had to give another name: the server held one by the
/// name the export recorded, put there by something this catalog never
/// saw. The album row follows the file to the name it was given.
///
/// By the album's server folder, because that is all an outbox entry knows.
/// `folder` is spelled as [`Place::Server`] spells it, without slashes at
/// either end.
pub fn rename_export(
conn: &Connection,
folder: &str,
from: &str,
to: &str,
) -> Result<(), CatalogError> {
ensure_tables(conn)?;
conn.execute(
"UPDATE OR REPLACE album_exports SET file_name = ?3
WHERE file_name = ?2
AND album_id IN (SELECT id FROM albums WHERE server_path = ?1 AND deleted = 0)",
rusqlite::params![folder.trim_matches('/'), from, to],
)?;
Ok(())
}
/// The photographs behind an album's files, most recently exported first —
/// what the grid shows when the album is opened.
pub fn sources(conn: &Connection, id: AlbumId) -> Result<Vec<ImageId>, CatalogError> {
@@ -463,6 +505,28 @@ mod tests {
assert_eq!(sources(conn, album).unwrap(), vec![b]);
}
#[test]
fn a_renamed_upload_moves_the_row_of_the_server_album_only() {
let cat = catalog();
let conn = cat.connection();
let web = create(conn, "Web", &Place::Server("Albums/Web".into())).unwrap();
let other = create(conn, "Other", &Place::Server("Albums/Other".into())).unwrap();
let a = image(conn, "a.cr3");
record_exports(conn, web, &[(a, "a.jpg".into())]).unwrap();
record_exports(conn, other, &[(a, "a.jpg".into())]).unwrap();
rename_export(conn, "/Albums/Web", "a.jpg", "a-1.jpg").unwrap();
assert_eq!(
file_names(conn, web).unwrap(),
["a-1.jpg".to_string()].into()
);
assert_eq!(
file_names(conn, other).unwrap(),
["a.jpg".to_string()].into()
);
}
#[test]
fn moving_to_the_server_forgets_the_local_folder() {
let cat = catalog();
+1
View File
@@ -50,6 +50,7 @@ pub mod faces;
pub mod jobs;
pub mod keywords;
pub mod merge;
pub mod name_dates;
pub mod query;
pub mod rating;
pub mod recovery;
+387
View File
@@ -0,0 +1,387 @@
//! TRACES: FR-CAT-5
//! A capture time read from the file's name, for an image whose header has
//! none.
//!
//! # Why
//!
//! A photograph with no EXIF date sorts after everything else, so it is lost
//! at the end of the grid and absent from the timeline. The files that end up
//! there are rarely without a date — they are without *EXIF*: WhatsApp strips
//! every tag and names the file `WhatsApp Image 2023-06-15 at 07.00.42.jpeg`,
//! a Windows Phone wrote `WP_20140922_14_16_27_Pro.jpg`, a phone camera
//! `IMG_20190812_153012.jpg`, and darktable's import renames to
//! `20230629_0001.jpeg`. On the reference library 250 of 274 undated images
//! carried their date in the name or in the folder above it.
//!
//! # What is accepted
//!
//! A date is `YYYYMMDD` as a whole run of digits, or `YYYY`, `MM` and `DD`
//! joined by `-`, `_` or `.`. A time may follow it — `HHMMSS` as one run (or
//! nine digits, milliseconds appended), or three two-digit runs joined by
//! `-`, `_`, `.` or `:` — after `_`, `-`, `.`, `T`, a space or ` at `.
//! Anything else after the date leaves it at midnight: `_0059` in
//! `20230628_0059` is a sequence number, not 00:59, and reading it as a time
//! would invent one.
//!
//! The name is tried first and then each folder above it, innermost first —
//! `2016/2016-11-11/IMG_7910.jpg` is dated by its folder. A bare year folder
//! is not a date: putting a photograph at 1 January is a wrong answer, and an
//! undated one at least says it does not know.
//!
//! The reading is wall-clock time with no zone, stored as EXIF's is
//! (`dr_decode::parse_exif_datetime`), and EXIF always wins: this only fills
//! rows whose `captured_at` is still empty.
use rusqlite::Connection;
use crate::CatalogError;
/// The capture time a path's name states, as wall-clock Unix seconds.
pub fn date_from_path(source_ref: &str) -> Option<i64> {
let mut parts = source_ref.rsplit(['/', '\\']);
let name = parts.next()?;
let stem = name.rsplit_once('.').map_or(name, |(stem, _)| stem);
date_in(stem).or_else(|| parts.find_map(date_in))
}
/// Date every examined, undated image whose name states one.
///
/// `only` limits the pass to the images just examined — what the sweep hands
/// in — and `None` visits every undated image, which is the backfill's case.
/// Both read the undated side alone (`images_captured` answers
/// `captured_at IS NULL` with a seek), never the library.
///
/// Returns how many images were dated.
pub fn fill(conn: &Connection, only: Option<&[i64]>) -> Result<usize, CatalogError> {
let rows: Vec<(i64, String)> = match only {
None => {
let mut stmt = conn.prepare(
"SELECT id, source_ref FROM images
WHERE captured_at IS NULL AND metadata_state >= 2",
)?;
let rows = stmt
.query_map([], |r| Ok((r.get(0)?, r.get(1)?)))?
.collect::<Result<_, _>>()?;
rows
}
Some(ids) => {
let mut stmt = conn.prepare_cached(
"SELECT source_ref FROM images
WHERE id = ?1 AND captured_at IS NULL AND metadata_state >= 2",
)?;
let mut rows = Vec::new();
for &id in ids {
let mut q = stmt.query([id])?;
if let Some(r) = q.next()? {
rows.push((id, r.get(0)?));
}
}
rows
}
};
let dated: Vec<(i64, i64)> = rows
.iter()
.filter_map(|(id, path)| date_from_path(path).map(|at| (*id, at)))
.collect();
if dated.is_empty() {
return Ok(0);
}
// A savepoint rather than a transaction, so a caller already inside one
// can still call this: the backfill's 250 rows are one commit, not 250.
conn.execute_batch("SAVEPOINT name_dates")?;
let written = (|| {
let mut stmt = conn.prepare_cached(
"UPDATE images SET captured_at = ?2 WHERE id = ?1 AND captured_at IS NULL",
)?;
let mut n = 0;
for (id, at) in &dated {
n += stmt.execute(rusqlite::params![id, at])?;
}
Ok::<_, CatalogError>(n)
})();
match written {
Ok(n) => {
conn.execute_batch("RELEASE name_dates")?;
Ok(n)
}
Err(e) => {
let _ = conn.execute_batch("ROLLBACK TO name_dates; RELEASE name_dates");
Err(e)
}
}
}
/// The first date, with its time if one follows, in one name component.
fn date_in(s: &str) -> Option<i64> {
let b = s.as_bytes();
let mut i = 0;
while i < b.len() {
// Only at the start of a run of digits: a date inside a longer number
// is a coincidence, not a date.
if b[i].is_ascii_digit() && (i == 0 || !b[i - 1].is_ascii_digit()) {
if let Some(at) = date_at(b, i) {
return Some(at);
}
}
i += 1;
}
None
}
/// A date starting at `i`, and the time after it if there is one.
fn date_at(b: &[u8], i: usize) -> Option<i64> {
let run = digits(b, i);
let ((y, mo, d), after) = match run.len() {
// YYYYMMDD, or YYYYMMDDHHMMSS written as one number.
8 | 14 => ((num(&run[..4]), num(&run[4..6]), num(&run[6..8])), i + 8),
4 => {
let sep = |at: usize| matches!(b.get(at), Some(b'-' | b'_' | b'.'));
let mo_at = i + 4 + 1;
let d_at = mo_at + 2 + 1;
if !(sep(i + 4) && digits(b, mo_at).len() == 2 && sep(mo_at + 2))
|| digits(b, d_at).len() != 2
{
return None;
}
(
(num(run), num(&b[mo_at..mo_at + 2]), num(&b[d_at..d_at + 2])),
d_at + 2,
)
}
_ => return None,
};
let day = civil_days(y, mo, d)?;
let time = if run.len() == 14 {
hms(num(&run[8..10]), num(&run[10..12]), num(&run[12..14]))
} else {
time_at(b, after)
};
Some(day * 86_400 + time.unwrap_or(0))
}
/// The time following a date that ends at `i`, as seconds into the day.
fn time_at(b: &[u8], i: usize) -> Option<i64> {
let rest = &b[i..];
let start = if rest.starts_with(b" at ") {
i + 4
} else if matches!(rest.first(), Some(b'_' | b'-' | b'.' | b'T' | b' ')) {
i + 1
} else {
return None;
};
let run = digits(b, start);
match run.len() {
// HHMMSS, or with milliseconds appended (Pixel's PXL_…_123456789).
6 | 9 => hms(num(&run[..2]), num(&run[2..4]), num(&run[4..6])),
2 => {
let sep = |at: usize| matches!(b.get(at), Some(b'-' | b'_' | b'.' | b':'));
let (m_at, s_at) = (start + 3, start + 6);
if !(sep(start + 2) && digits(b, m_at).len() == 2 && sep(m_at + 2))
|| digits(b, s_at).len() != 2
{
return None;
}
hms(num(run), num(&b[m_at..m_at + 2]), num(&b[s_at..s_at + 2]))
}
_ => None,
}
}
/// The run of ASCII digits starting at `i`.
fn digits(b: &[u8], i: usize) -> &[u8] {
let rest = b.get(i..).unwrap_or(&[]);
let n = rest.iter().take_while(|c| c.is_ascii_digit()).count();
&rest[..n]
}
fn num(d: &[u8]) -> i64 {
d.iter().fold(0, |n, c| n * 10 + i64::from(c - b'0'))
}
fn hms(h: i64, m: i64, s: i64) -> Option<i64> {
((0..24).contains(&h) && (0..60).contains(&m) && (0..61).contains(&s))
.then_some(h * 3_600 + m * 60 + s)
}
/// Days since 1970-01-01 for a valid civil date, `None` for anything else.
///
/// The year range is EXIF's (`parse_exif_datetime`): wide enough for scanned
/// film, narrow enough that a counter such as `12345678` is not a date.
fn civil_days(y: i64, mo: i64, d: i64) -> Option<i64> {
let leap = y % 4 == 0 && (y % 100 != 0 || y % 400 == 0);
let month_len = match mo {
1 | 3 | 5 | 7 | 8 | 10 | 12 => 31,
4 | 6 | 9 | 11 => 30,
2 if leap => 29,
2 => 28,
_ => return None,
};
if !(1900..=2200).contains(&y) || !(1..=month_len).contains(&d) {
return None;
}
let y_adj = if mo <= 2 { y - 1 } else { y };
let era = y_adj.div_euclid(400);
let yoe = y_adj - era * 400;
let mp = (mo + 9) % 12;
let doy = (153 * mp + 2) / 5 + d - 1;
let doe = yoe * 365 + yoe / 4 - yoe / 100 + doy;
Some(era * 146_097 + doe - 719_468)
}
#[cfg(test)]
mod tests {
use super::*;
/// Wall-clock seconds for a date and time, the expected side of each case.
fn at(y: i64, mo: i64, d: i64, h: i64, mi: i64, s: i64) -> Option<i64> {
Some(civil_days(y, mo, d).unwrap() * 86_400 + h * 3_600 + mi * 60 + s)
}
#[test]
fn the_names_in_the_reference_library_are_read() {
// Every shape here is a file that sat undated at the end of the grid.
for (path, want) in [
(
"PhotosRaw/alps trip/alps whatsapp/WhatsApp Image 2023-06-15 at 07.00.42.jpeg",
at(2023, 6, 15, 7, 0, 42),
),
(
"PhotosRaw/alps trip/alps whatsapp/WhatsApp Image 2023-06-17 at 12.45.52 (1).jpeg",
at(2023, 6, 17, 12, 45, 52),
),
(
"PhotosRaw/WP_20140922_14_16_27_Pro.jpg",
at(2014, 9, 22, 14, 16, 27),
),
// A sequence number after the date is not a time.
(
"PhotosRaw/Darktable/20230629_no_name/20230629_0001.jpeg",
at(2023, 6, 29, 0, 0, 0),
),
("PhotosRaw/20230628_0059.jpg", at(2023, 6, 28, 0, 0, 0)),
(
"PhotosRaw/backdrops/IMG_20130625_0021.jpg",
at(2013, 6, 25, 0, 0, 0),
),
(
"PhotosRaw/alps trip/20230628_0641 - 20230628_0661.jpg",
at(2023, 6, 28, 0, 0, 0),
),
] {
assert_eq!(date_from_path(path), want, "{path}");
}
}
#[test]
fn common_camera_and_app_names_are_read() {
for (path, want) in [
("IMG_20190812_153012.jpg", at(2019, 8, 12, 15, 30, 12)),
("PXL_20210101_123456789.jpg", at(2021, 1, 1, 12, 34, 56)),
(
"Screenshot_2021-03-04-12-30-45.png",
at(2021, 3, 4, 12, 30, 45),
),
(
"Screenshot from 2021-03-04 12-30-45.png",
at(2021, 3, 4, 12, 30, 45),
),
("IMG-20210304-WA0001.jpg", at(2021, 3, 4, 0, 0, 0)),
("20210304143012.jpg", at(2021, 3, 4, 14, 30, 12)),
("2019.12.25 party.jpg", at(2019, 12, 25, 0, 0, 0)),
("signal-2022-01-02-101112.jpg", at(2022, 1, 2, 10, 11, 12)),
("2022-01-02T10:11:12.jpg", at(2022, 1, 2, 10, 11, 12)),
] {
assert_eq!(date_from_path(path), want, "{path}");
}
}
#[test]
fn a_folder_dates_a_name_that_does_not() {
assert_eq!(
date_from_path("PhotosRaw/2016/2016-11-11/IMG_7910.jpg"),
at(2016, 11, 11, 0, 0, 0)
);
// The innermost folder that states a date wins.
assert_eq!(
date_from_path("2016-01-01 trip/2016-01-03/_MG_1.jpg"),
at(2016, 1, 3, 0, 0, 0)
);
// The name beats its folder.
assert_eq!(
date_from_path("2016-11-11/IMG_20161112_080000.jpg"),
at(2016, 11, 12, 8, 0, 0)
);
}
#[test]
fn numbers_that_are_not_dates_are_left_alone() {
for path in [
"PhotosRaw/_MG_9002.jpg",
"PhotosRaw/scanning/fau_2.jpg",
// A year folder is not a day.
"PhotosRaw/2016/_MG_1.jpg",
"IMG_1999.jpg",
"DSC_12345678.jpg", // month 56
"20230230_0001.jpg", // 30 February
"120230615.jpg", // the date is inside a longer number
"1612345678901.jpg", // a millisecond epoch, not a civil date
"2023-6-15.jpg", // a one-digit month is too loose to trust
] {
assert_eq!(date_from_path(path), None, "{path}");
}
}
#[test]
fn a_time_that_cannot_be_is_dropped_and_the_date_kept() {
assert_eq!(
date_from_path("20230615_256199.jpg"),
at(2023, 6, 15, 0, 0, 0)
);
}
#[test]
fn fill_dates_only_examined_undated_rows_and_never_overrides_exif() {
let c = Connection::open_in_memory().unwrap();
crate::schema::migrate(&c).unwrap();
c.execute(
"INSERT INTO roots(id, kind, label) VALUES (1, 'remote', 'lib')",
[],
)
.unwrap();
// (id, name, captured_at, metadata_state)
for (id, name, captured, state) in [
(1i64, "IMG_20190812_153012.jpg", None, 2i64),
// EXIF already answered; the name disagrees and loses.
(2, "IMG_20190812_153012b.jpg", Some(42i64), 2),
// Not yet examined: EXIF may still come, so the name waits.
(3, "IMG_20190813_000000.jpg", None, 1),
(4, "_MG_9002.jpg", None, 2),
] {
c.execute(
"INSERT INTO images(id, root_id, source_ref, captured_at, metadata_state, added_at)
VALUES (?1, 1, ?2, ?3, ?4, 0)",
rusqlite::params![id, name, captured, state],
)
.unwrap();
}
let captured = |id: i64| -> Option<i64> {
c.query_row("SELECT captured_at FROM images WHERE id = ?1", [id], |r| {
r.get(0)
})
.unwrap()
};
assert_eq!(fill(&c, Some(&[2, 3, 4])).unwrap(), 0);
assert_eq!(fill(&c, None).unwrap(), 1);
assert_eq!(captured(1), at(2019, 8, 12, 15, 30, 12));
assert_eq!(captured(2), Some(42));
assert_eq!(captured(3), None);
assert_eq!(captured(4), None);
// Nothing left to do is a no-op, not a rewrite.
assert_eq!(fill(&c, None).unwrap(), 0);
}
}
+9
View File
@@ -356,6 +356,15 @@ pub fn backfill(conn: &Connection) -> Result<Vec<(&'static str, usize)>, Catalog
out.push(("keyword_terms", n));
}
// TRACES: FR-CAT-5
// A date from the file's name for every examined image EXIF left undated.
// The sweep does this as it examines each image; this is for the images
// examined by a build that did not, and reads the undated side alone.
let n = crate::name_dates::fill(conn, None)?;
if n > 0 {
out.push(("dates_from_names", n));
}
Ok(out)
}
+4
View File
@@ -32,6 +32,10 @@ fn main() {
println!("black {:?}", raw.black_level);
println!("white {}", raw.white_level);
println!("wb_coeffs {:?}", raw.wb_coeffs);
match dr_decode::noise_profile(&bytes) {
Some(p) => println!("noise profile {p:?} ((S, O) per plane)"),
None => println!("noise profile none"),
}
match raw.color_matrix {
Some(m) => {
+806
View File
@@ -0,0 +1,806 @@
//! TRACES: FR-DEV-3e
//! DNG camera profiles: the tables on top of the matrix (D20).
//!
//! A profile is what [`crate::profile`] already reads — colour and forward
//! matrices per calibration illuminant — plus two lookups over HSV: the
//! `ProfileHueSatMap`, a calibration, and the `ProfileLookTable`, a rendering
//! intent. `docs/dev/camera-profiles.md` is the design; this module finds
//! them, in the order its §4 gives:
//!
//! 1. embedded in the DNG being decoded ([`Dcp::from_ifd`]);
//! 2. a `.dcp` file in the profiles directory whose `UniqueCameraModel`
//! names this body ([`find`]);
//! 3. nowhere, and the matrix renders alone.
//!
//! A `.dcp` is a TIFF whose magic is `RC` (0x4352) rather than 42, holding one
//! IFD of the same tags a DNG carries. rawler's TIFF reader does not check the
//! magic, so both sources go through the one parser and [`Dcp::from_ifd`].
//!
//! Nothing here applies a table. The lookup is the shader's, with its CPU
//! reference in `dr-pipeline`; this module resolves *which* tables, and blends
//! the HueSatMap for the light the frame was shot under, once per decode.
use std::path::{Path, PathBuf};
use std::sync::{Arc, OnceLock, RwLock};
use dr_types::{HueSatTable, ProfileOrigin, ProfileTables};
use rawler::formats::tiff::{
DirectoryWriter, GenericTiffReader, SRational, TiffWriter, Value, IFD,
};
use rawler::imgop::xyz::Illuminant;
use rawler::tags::DngTag;
use crate::profile::{illuminant_temperature, Calibration, CameraProfile};
/// The magic a `.dcp` carries where a TIFF carries 42.
const DCP_MAGIC: u16 = 0x4352;
/// `ProfileEmbedPolicy` values that permit copying a profile out of the file
/// it came in: 0, "allow copying", and 3, "no restrictions". 1 ("embed if
/// used") and 2 ("embed never") do not.
const COPYABLE_POLICIES: [u32; 2] = [0, 3];
/// A camera profile as a DNG or a `.dcp` states it.
///
/// Indexed `[0]`/`[1]` for calibration 1 and 2, positionally, because that is
/// how the file pairs a matrix and a table with its illuminant.
#[derive(Debug, Clone, PartialEq)]
pub struct Dcp {
/// `ProfileName`. Empty where the file names none.
pub name: String,
/// `UniqueCameraModel`: the body the profile was made for.
pub unique_camera_model: Option<String>,
pub copyright: Option<String>,
pub calibration_signature: Option<String>,
/// `ProfileEmbedPolicy`; 0 where absent, as the DNG specification
/// defaults it.
pub embed_policy: u32,
/// `CalibrationIlluminant1/2`, as EXIF light-source codes.
pub illuminants: [Option<u16>; 2],
/// `ColorMatrix1/2`: XYZ → camera.
pub color_matrix: [Option<[[f32; 3]; 3]>; 2],
/// `ForwardMatrix1/2`: white-balanced camera → XYZ (D50).
pub forward_matrix: [Option<[[f32; 3]; 3]>; 2],
/// `ProfileHueSatMapData1/2`, sharing one dimensions tag.
pub hue_sat: [Option<HueSatTable>; 2],
/// `ProfileLookTableData`.
pub look: Option<HueSatTable>,
/// `ProfileToneCurve`, as stored: input/output pairs. Carried so a copy
/// keeps it, never applied — tone is the view transform's (D19, D20).
pub tone_curve: Option<Vec<f32>>,
/// `BaselineExposureOffset`, in stops: the profile's correction to the
/// file's `BaselineExposure` (camera-profiles.md §11). A profile copied
/// out of a DNG carries that DNG's baseline here, so a raw with no
/// baseline of its own lands at the same brightness.
pub baseline_exposure_offset: f32,
}
impl Dcp {
/// TRACES: FR-DEV-3e
/// Read a profile out of an IFD — a DNG's root, or a `.dcp`'s only one.
///
/// `None` where the IFD carries neither table. A DNG always has matrices
/// and the decoder already reads them; what makes a *profile* worth
/// carrying separately is a table, so its absence is "no profile" rather
/// than a profile that says nothing.
pub fn from_ifd(ifd: &IFD) -> Option<Self> {
let hue_sat_dims = dims(ifd, DngTag::ProfileHueSatMapDims);
let hue_sat_srgb = encoding(ifd, DngTag::ProfileHueSatMapEncoding);
let hue_sat = [DngTag::ProfileHueSatMapData1, DngTag::ProfileHueSatMapData2]
.map(|tag| hue_sat_dims.and_then(|d| table(ifd, tag, d, hue_sat_srgb)));
let look = dims(ifd, DngTag::ProfileLookTableDims).and_then(|d| {
table(
ifd,
DngTag::ProfileLookTableData,
d,
encoding(ifd, DngTag::ProfileLookTableEncoding),
)
});
if hue_sat[0].is_none() && hue_sat[1].is_none() && look.is_none() {
return None;
}
Some(Self {
name: string(ifd, DngTag::ProfileName).unwrap_or_default(),
unique_camera_model: string(ifd, DngTag::UniqueCameraModel),
copyright: string(ifd, DngTag::ProfileCopyright),
calibration_signature: string(ifd, DngTag::ProfileCalibrationSignature),
embed_policy: ifd
.get_entry(DngTag::ProfileEmbedPolicy)
.and_then(|e| e.value.get_u32(0).ok().flatten())
.unwrap_or(0),
illuminants: [
DngTag::CalibrationIlluminant1,
DngTag::CalibrationIlluminant2,
]
.map(|tag| {
ifd.get_entry(tag)
.and_then(|e| e.value.get_u16(0).ok().flatten())
}),
color_matrix: [DngTag::ColorMatrix1, DngTag::ColorMatrix2].map(|t| matrix(ifd, t)),
forward_matrix: [DngTag::ForwardMatrix1, DngTag::ForwardMatrix2]
.map(|t| matrix(ifd, t)),
hue_sat,
look,
tone_curve: ifd
.get_entry(DngTag::ProfileToneCurve)
.and_then(|e| floats(&e.value))
.filter(|v| v.len() >= 4 && v.len() % 2 == 0),
baseline_exposure_offset: stops(ifd, DngTag::BaselineExposureOffset),
})
}
/// TRACES: FR-DEV-3e
/// Parse a `.dcp` file's bytes.
pub fn parse(bytes: &[u8]) -> Result<Self, String> {
if bytes.len() < 8 {
return Err("too short to be a camera profile".into());
}
let magic = match &bytes[..2] {
b"II" => u16::from_le_bytes([bytes[2], bytes[3]]),
b"MM" => u16::from_be_bytes([bytes[2], bytes[3]]),
_ => return Err("not a TIFF-structured file".into()),
};
if magic != DCP_MAGIC {
return Err(format!("magic {magic:#x} is not a camera profile's"));
}
let reader =
GenericTiffReader::new_with_buffer(bytes, 0, 0, Some(0)).map_err(|e| e.to_string())?;
use rawler::formats::tiff::reader::TiffReader;
let profile = Self::from_ifd(reader.root_ifd())
.ok_or_else(|| "a profile with no HueSatMap and no LookTable".to_string())?;
if profile.color_matrix[0].is_none() {
return Err("a profile with no ColorMatrix1".into());
}
Ok(profile)
}
/// Whether the file this profile came in allows it to be copied out
/// (camera-profiles.md §4).
pub fn may_copy(&self) -> bool {
COPYABLE_POLICIES.contains(&self.embed_policy)
}
/// TRACES: FR-DEV-3e
/// Whether this profile was made for the body named.
///
/// `unique` is the file's own `UniqueCameraModel`, where a DNG carries
/// one; `make` and `model` are rawler's cleaned names, joined as Adobe
/// spells a body ("Canon EOS 6D"). Case and runs of spaces are ignored,
/// because the two spellings come from different vendors' tables.
pub fn is_for(&self, unique: Option<&str>, make: &str, model: &str) -> bool {
let Some(mine) = self.unique_camera_model.as_deref().map(normalise) else {
return false;
};
let joined = if normalise(model).starts_with(&normalise(make)) {
normalise(model)
} else {
normalise(&format!("{make} {model}"))
};
unique.map(normalise).as_deref() == Some(mine.as_str()) || joined == mine
}
/// TRACES: FR-DEV-3e
/// The matrices this profile was built against, as the decoder's
/// [`CameraProfile`], with the frame's own as-shot neutral.
///
/// A `.dcp` is a whole profile: its tables were measured relative to its
/// forward matrix, so using them over the file's matrices would apply a
/// correction for a different starting point. `None` where no calibration
/// is usable, and the caller keeps the file's.
pub fn camera_profile(&self, neutral: Option<[f32; 3]>) -> Option<CameraProfile> {
let calibrations = (0..2)
.filter_map(|i| {
let xyz_to_cam = self.color_matrix[i]?;
let temperature = self.temperature(i)?;
Some(Calibration {
temperature,
xyz_to_cam,
forward: self.forward_matrix[i],
})
})
.collect();
CameraProfile::new(calibrations, neutral)
}
/// TRACES: FR-DEV-3e
/// The tables to render this frame with: the HueSatMap blended for the
/// scene's colour temperature, by the same mired weight the matrices use,
/// and the LookTable as it is.
///
/// Tables that change nothing are dropped here, so the shader is never
/// asked to look up an identity.
pub fn tables(&self, scene_temperature: f32, origin: ProfileOrigin) -> ProfileTables {
let hue_sat = match (&self.hue_sat, self.temperature(0), self.temperature(1)) {
([Some(a), Some(b)], Some(ta), Some(tb)) => {
let t = mired_weight(ta, tb, scene_temperature);
a.lerp(b, t).or_else(|| Some(a.clone()))
}
([Some(a), _], _, _) => Some(a.clone()),
([None, Some(b)], _, _) => Some(b.clone()),
([None, None], _, _) => None,
};
ProfileTables {
name: self.name.clone(),
origin,
hue_sat: hue_sat.filter(|t| !t.is_identity()),
look: self.look.clone().filter(|t| !t.is_identity()),
tone_curve: self
.tone_curve
.as_deref()
.and_then(dr_types::tone::resample_tone_curve),
}
}
fn temperature(&self, i: usize) -> Option<f32> {
let code = self.illuminants[i]?;
let illuminant: Illuminant = code.try_into().ok()?;
illuminant_temperature(illuminant)
}
/// TRACES: FR-DEV-3e
/// This profile as `.dcp` bytes, for [`save`].
pub fn to_bytes(&self) -> Result<Vec<u8>, String> {
let mut cursor = std::io::Cursor::new(Vec::new());
let writer = TiffWriter::new(&mut cursor).map_err(|e| e.to_string())?;
let mut dir = DirectoryWriter::new();
if let Some(model) = &self.unique_camera_model {
dir.add_tag(DngTag::UniqueCameraModel, model.as_str());
}
dir.add_tag(DngTag::ProfileName, self.name.as_str());
if let Some(c) = &self.copyright {
dir.add_tag(DngTag::ProfileCopyright, c.as_str());
}
if let Some(s) = &self.calibration_signature {
dir.add_tag(DngTag::ProfileCalibrationSignature, s.as_str());
}
dir.add_tag(DngTag::ProfileEmbedPolicy, self.embed_policy);
let illuminant_tags = [
DngTag::CalibrationIlluminant1,
DngTag::CalibrationIlluminant2,
];
for (tag, code) in illuminant_tags.into_iter().zip(self.illuminants) {
if let Some(code) = code {
dir.add_tag(tag, code);
}
}
for (tag, m) in [DngTag::ColorMatrix1, DngTag::ColorMatrix2]
.into_iter()
.zip(self.color_matrix)
.chain(
[DngTag::ForwardMatrix1, DngTag::ForwardMatrix2]
.into_iter()
.zip(self.forward_matrix),
)
{
if let Some(m) = m {
dir.add_value(tag, srational_matrix(&m));
}
}
if let Some(first) = self.hue_sat.iter().flatten().next() {
dir.add_tag(
DngTag::ProfileHueSatMapDims,
[
first.hue_divisions,
first.sat_divisions,
first.val_divisions,
],
);
dir.add_tag(
DngTag::ProfileHueSatMapEncoding,
u32::from(first.srgb_encoded),
);
for (tag, t) in [DngTag::ProfileHueSatMapData1, DngTag::ProfileHueSatMapData2]
.into_iter()
.zip(&self.hue_sat)
{
if let Some(t) = t {
dir.add_value(
tag,
Value::Float(t.entries.iter().flatten().copied().collect()),
);
}
}
}
if let Some(t) = &self.look {
dir.add_tag(
DngTag::ProfileLookTableDims,
[t.hue_divisions, t.sat_divisions, t.val_divisions],
);
dir.add_tag(DngTag::ProfileLookTableEncoding, u32::from(t.srgb_encoded));
dir.add_value(
DngTag::ProfileLookTableData,
Value::Float(t.entries.iter().flatten().copied().collect()),
);
}
if let Some(curve) = &self.tone_curve {
dir.add_value(DngTag::ProfileToneCurve, Value::Float(curve.clone()));
}
if self.baseline_exposure_offset != 0.0 {
dir.add_value(
DngTag::BaselineExposureOffset,
Value::SRational(vec![SRational::new(
(self.baseline_exposure_offset * 100.0).round() as i32,
100,
)]),
);
}
writer.build(dir).map_err(|e| e.to_string())?;
let mut bytes = cursor.into_inner();
// The writer stamps TIFF's 42 in its own byte order; a profile is the
// same structure with its own magic in the same place.
bytes[2..4].copy_from_slice(&DCP_MAGIC.to_ne_bytes());
Ok(bytes)
}
}
/// The weight toward calibration 2, by reciprocal temperature — the same
/// interpolation [`CameraProfile`] gives the matrices, so the tables and the
/// matrix agree about how far between the two lights a frame was shot.
fn mired_weight(t1: f32, t2: f32, scene: f32) -> f32 {
let mired = |k: f32| 1.0e6 / k.max(1.0);
let (a, b) = (mired(t1), mired(t2));
if (a - b).abs() < 1e-6 {
return 0.0;
}
((mired(scene) - a) / (b - a)).clamp(0.0, 1.0)
}
fn normalise(s: &str) -> String {
s.split_whitespace()
.collect::<Vec<_>>()
.join(" ")
.to_lowercase()
}
fn string(ifd: &IFD, tag: DngTag) -> Option<String> {
ifd.get_entry(tag)
.and_then(|e| e.value.as_string().cloned())
.map(|s| s.trim_end_matches('\0').trim().to_string())
.filter(|s| !s.is_empty())
}
/// A single rational tag in stops, zero where absent or unreadable — the
/// DNG specification's default for both exposure tags.
fn stops(ifd: &IFD, tag: DngTag) -> f32 {
ifd.get_entry(tag)
.and_then(|e| e.value.get_f32(0).ok().flatten())
.filter(|v| v.is_finite())
.unwrap_or(0.0)
}
fn floats(value: &Value) -> Option<Vec<f32>> {
(0..value.count())
.map(|i| value.get_f32(i).ok().flatten())
.collect()
}
fn matrix(ifd: &IFD, tag: DngTag) -> Option<[[f32; 3]; 3]> {
let v = floats(&ifd.get_entry(tag)?.value)?;
if v.len() != 9 || v.iter().any(|x| !x.is_finite()) {
return None;
}
Some([[v[0], v[1], v[2]], [v[3], v[4], v[5]], [v[6], v[7], v[8]]])
}
fn dims(ifd: &IFD, tag: DngTag) -> Option<[u32; 3]> {
let e = ifd.get_entry(tag)?;
let at = |i| e.value.get_u32(i).ok().flatten();
Some([at(0)?, at(1)?, at(2)?])
}
fn encoding(ifd: &IFD, tag: DngTag) -> bool {
ifd.get_entry(tag)
.and_then(|e| e.value.get_u32(0).ok().flatten())
== Some(1)
}
fn table(ifd: &IFD, tag: DngTag, [h, s, v]: [u32; 3], srgb: bool) -> Option<HueSatTable> {
let data = floats(&ifd.get_entry(tag)?.value)?;
if data.len() % 3 != 0 {
return None;
}
let entries = data.chunks_exact(3).map(|c| [c[0], c[1], c[2]]).collect();
HueSatTable::new(h, s, v, srgb, entries)
}
fn srational_matrix(m: &[[f32; 3]; 3]) -> Value {
const SCALE: i32 = 10_000;
Value::SRational(
m.iter()
.flatten()
.map(|v| SRational::new((v * SCALE as f32).round() as i32, SCALE))
.collect(),
)
}
// ---- the profiles directory -------------------------------------------------
/// The `.dcp` files the photographer has installed, loaded once per process.
struct Library {
dir: PathBuf,
/// `(file name, profile)`, sorted by file name so that two profiles for
/// one body resolve the same way on every run (camera-profiles.md §4).
profiles: Vec<(String, Arc<Dcp>)>,
}
fn library() -> &'static RwLock<Option<Library>> {
static LIBRARY: OnceLock<RwLock<Option<Library>>> = OnceLock::new();
LIBRARY.get_or_init(|| RwLock::new(None))
}
/// TRACES: FR-DEV-3e
/// Name the profiles directory and read every `.dcp` in it.
///
/// Called once at start-up by the application, with a path under the
/// platform data directory. A decode before this, or in a process that never
/// calls it (a test, a bench), finds no directory profiles, which is the
/// matrix-only render it always had.
pub fn set_profiles_directory(dir: PathBuf) {
let profiles = load(&dir);
if let Ok(mut lib) = library().write() {
*lib = Some(Library { dir, profiles });
}
}
/// The directory [`set_profiles_directory`] named, if any.
pub fn profiles_directory() -> Option<PathBuf> {
library().read().ok()?.as_ref().map(|l| l.dir.clone())
}
fn load(dir: &Path) -> Vec<(String, Arc<Dcp>)> {
let Ok(entries) = std::fs::read_dir(dir) else {
return Vec::new();
};
let mut out: Vec<(String, Arc<Dcp>)> = entries
.flatten()
.filter(|e| {
e.path()
.extension()
.is_some_and(|x| x.eq_ignore_ascii_case("dcp"))
})
.filter_map(|e| {
let name = e.file_name().to_string_lossy().into_owned();
let bytes = std::fs::read(e.path()).ok()?;
match Dcp::parse(&bytes) {
Ok(p) => Some((name, Arc::new(p))),
Err(why) => {
log::warn!("camera profile {name} skipped: {why}");
None
}
}
})
.collect();
out.sort_by(|a, b| a.0.cmp(&b.0));
log::info!(
"camera profiles: {} loaded from {}",
out.len(),
dir.display()
);
out
}
/// TRACES: FR-DEV-3e
/// The first installed profile, by file name, made for this body.
pub fn find(unique: Option<&str>, make: &str, model: &str) -> Option<(String, Arc<Dcp>)> {
let lib = library().read().ok()?;
lib.as_ref()?
.profiles
.iter()
.find(|(_, p)| p.is_for(unique, make, model))
.cloned()
}
/// TRACES: FR-DEV-3e
/// Save a profile copied out of a photograph into the profiles directory, and
/// make it available to the next decode.
///
/// Refuses a profile whose embed policy does not allow copying, and refuses
/// when no directory is set. Named after the body and the profile, so a
/// second copy of the same profile replaces the first rather than piling up.
pub fn save(profile: &Dcp) -> Result<PathBuf, String> {
if !profile.may_copy() {
return Err("this profile's embed policy does not allow copying it".into());
}
let model = profile
.unique_camera_model
.as_deref()
.ok_or("the profile names no camera")?;
let dir = profiles_directory().ok_or("no profiles directory is set")?;
std::fs::create_dir_all(&dir).map_err(|e| e.to_string())?;
let file_name: String = format!("{model} {}.dcp", profile.name)
.chars()
.map(|c| {
if c.is_alphanumeric() || " -_.".contains(c) {
c
} else {
'_'
}
})
.collect();
let path = dir.join(file_name.trim());
std::fs::write(&path, profile.to_bytes()?).map_err(|e| e.to_string())?;
set_profiles_directory(dir);
Ok(path)
}
/// TRACES: FR-DEV-3e
/// The profile embedded in a file, read on demand — for the panel's offer to
/// copy it, which happens long after the decode that rendered it.
///
/// Reads the header only; no photosite is unpacked.
pub fn embedded_in(bytes: &[u8]) -> Option<Dcp> {
let source = rawler::rawsource::RawSource::new_from_slice(bytes);
let decoder = rawler::get_decoder(&source).ok()?;
let root = decoder
.ifd(rawler::decoders::WellKnownIFD::Root)
.ok()
.flatten()?;
// The copy carries the file's baseline as its offset, so a raw from the
// same body that has no baseline of its own — a CR2 — gets the total the
// DNG renders at (camera-profiles.md §11).
let mut profile = Dcp::from_ifd(&root)?;
profile.baseline_exposure_offset += stops(&root, DngTag::BaselineExposure);
Some(profile)
}
/// TRACES: FR-DEV-3e
/// What one decode resolved: the matrices to render through, the tables on
/// top of them, and the embedded profile if the file had one — kept whole so
/// the panel can offer to copy it.
pub struct Resolved {
pub profile: Option<CameraProfile>,
pub tables: Option<Arc<ProfileTables>>,
pub embedded: Option<Arc<Dcp>>,
/// Stops to add at render: the file's `BaselineExposure` plus the
/// chosen profile's `BaselineExposureOffset` (camera-profiles.md §11).
pub baseline_exposure: f32,
}
/// TRACES: FR-DEV-3e
/// Apply camera-profiles.md §4's order to one decoded file.
///
/// `matrices` is the profile the decoder built from the file; `root` the
/// file's root IFD, where a DNG keeps its embedded profile.
pub fn resolve(
matrices: Option<CameraProfile>,
root: Option<&IFD>,
make: &str,
model: &str,
) -> Resolved {
let embedded = root.and_then(Dcp::from_ifd).map(Arc::new);
let file_baseline = root.map_or(0.0, |r| stops(r, DngTag::BaselineExposure));
if let Some(dcp) = &embedded {
let tables = matrices
.as_ref()
.map(|m| dcp.tables(m.scene_temperature(), ProfileOrigin::Embedded))
.filter(|t| !t.is_empty())
.map(Arc::new);
return Resolved {
profile: matrices,
tables,
baseline_exposure: file_baseline + dcp.baseline_exposure_offset,
embedded,
};
}
let unique = root.and_then(|r| string(r, DngTag::UniqueCameraModel));
if let Some((file, dcp)) = find(unique.as_deref(), make, model) {
let neutral = matrices.as_ref().and_then(|m| m.neutral());
if let Some(own) = dcp.camera_profile(neutral) {
let tables = dcp.tables(own.scene_temperature(), ProfileOrigin::File(file));
return Resolved {
tables: (!tables.is_empty()).then(|| Arc::new(tables)),
profile: Some(own),
embedded: None,
baseline_exposure: file_baseline + dcp.baseline_exposure_offset,
};
}
}
Resolved {
profile: matrices,
tables: None,
embedded: None,
baseline_exposure: file_baseline,
}
}
#[cfg(test)]
mod tests {
use super::*;
fn table(h: u32, s: u32, v: u32, fill: [f32; 3]) -> HueSatTable {
HueSatTable::new(h, s, v, false, vec![fill; (h * s * v) as usize]).unwrap()
}
fn sample() -> Dcp {
Dcp {
name: "Test Standard".into(),
unique_camera_model: Some("Canon EOS 6D".into()),
copyright: Some("nobody".into()),
calibration_signature: Some("com.example".into()),
embed_policy: 0,
illuminants: [Some(17), Some(21)],
color_matrix: [
Some([
[0.7546, -0.1435, -0.0929],
[-0.3846, 1.1488, 0.2692],
[-0.0332, 0.1209, 0.637],
]),
Some([
[0.7034, -0.0804, -0.1014],
[-0.442, 1.2564, 0.2058],
[-0.0851, 0.1994, 0.5758],
]),
],
forward_matrix: [
Some([
[0.7763, 0.0065, 0.1815],
[0.2364, 0.8351, -0.0715],
[-0.0059, -0.4228, 1.2538],
]),
Some([
[0.7464, 0.1044, 0.1135],
[0.2648, 0.9173, -0.182],
[0.0113, -0.2154, 1.0292],
]),
],
hue_sat: [
Some(table(6, 3, 1, [2.0, 1.1, 1.0])),
Some(table(6, 3, 1, [-2.0, 0.9, 1.0])),
],
look: Some(table(4, 2, 3, [0.0, 1.2, 0.95])),
tone_curve: Some(vec![0.0, 0.0, 0.5, 0.6, 1.0, 1.0]),
baseline_exposure_offset: 0.25,
}
}
#[test]
fn a_profile_survives_being_written_and_read_back() {
let original = sample();
let bytes = original.to_bytes().unwrap();
assert_eq!(&bytes[2..4], &DCP_MAGIC.to_ne_bytes());
let back = Dcp::parse(&bytes).unwrap();
assert_eq!(
back.hue_sat, original.hue_sat,
"tables are stored as f32 and come back exact"
);
assert_eq!(back.look, original.look);
assert_eq!(back.name, original.name);
assert_eq!(back.unique_camera_model, original.unique_camera_model);
assert_eq!(back.illuminants, original.illuminants);
assert_eq!(back.tone_curve, original.tone_curve);
assert_eq!(back.baseline_exposure_offset, 0.25);
assert_eq!(
back.forward_matrix, original.forward_matrix,
"four decimals, as the file has"
);
}
#[test]
fn a_tiff_is_not_a_profile() {
let mut bytes = sample().to_bytes().unwrap();
bytes[2..4].copy_from_slice(&42u16.to_ne_bytes());
assert!(Dcp::parse(&bytes).is_err());
assert!(Dcp::parse(b"nonsense").is_err());
}
#[test]
fn a_body_matches_by_unique_model_or_by_make_and_model() {
let p = sample();
assert!(p.is_for(None, "Canon", "EOS 6D"));
assert!(p.is_for(None, "canon", "eos 6d"));
assert!(p.is_for(Some("Canon EOS 6D"), "", ""));
assert!(!p.is_for(None, "Canon", "EOS 6D Mark II"));
assert!(!p.is_for(Some("Canon EOS 5D"), "Canon", "EOS 5D"));
// A model that already starts with the make is not doubled.
assert!(p.is_for(None, "Canon", "Canon EOS 6D"));
}
#[test]
fn the_hue_sat_map_follows_the_light_the_frame_was_shot_under() {
let p = sample();
let at = |k| {
p.tables(k, ProfileOrigin::Embedded)
.hue_sat
.unwrap()
.entries[0]
};
assert_eq!(at(2856.0), [2.0, 1.1, 1.0], "tungsten is calibration 1");
assert_eq!(at(6504.0), [-2.0, 0.9, 1.0], "daylight is calibration 2");
let mid = at(4000.0);
assert!(mid[0] > -2.0 && mid[0] < 2.0, "{mid:?}");
}
#[test]
fn a_table_that_changes_nothing_is_not_handed_on() {
let mut p = sample();
p.hue_sat = [Some(table(6, 3, 1, [0.0, 1.0, 1.0])), None];
let t = p.tables(5000.0, ProfileOrigin::Embedded);
assert!(t.hue_sat.is_none());
assert!(t.look.is_some());
}
#[test]
fn only_a_copyable_policy_may_be_copied() {
let mut p = sample();
for (policy, ok) in [(0, true), (1, false), (2, false), (3, true)] {
p.embed_policy = policy;
assert_eq!(p.may_copy(), ok, "policy {policy}");
}
}
/// A Canon 6D DNG from the library, written by Lightroom 6.14 with Adobe
/// Standard embedded. Read from `DR_DCP_SAMPLE`, else the library path the
/// figures in camera-profiles.md §1 came from; skipped where neither
/// exists, because the file is not ours to put in the repository.
fn six_d_dng() -> Option<Vec<u8>> {
let path = std::env::var_os("DR_DCP_SAMPLE")
.map(PathBuf::from)
.or_else(|| {
std::env::var_os("HOME").map(|h| {
PathBuf::from(h).join("Nextcloud/PhotosRaw/2017/2017-08-12/_MG_9080.dng")
})
})?;
let bytes = std::fs::read(&path).ok();
if bytes.is_none() {
eprintln!("skipped: no sample DNG at {}", path.display());
}
bytes
}
#[test]
fn the_libraries_six_d_dngs_carry_adobe_standard() {
let Some(bytes) = six_d_dng() else { return };
let p = embedded_in(&bytes).expect("an embedded profile");
assert_eq!(p.name, "Adobe Standard");
assert_eq!(p.unique_camera_model.as_deref(), Some("Canon EOS 6D"));
assert_eq!(p.embed_policy, 0);
let hs = p.hue_sat[0].as_ref().unwrap();
assert_eq!(
(hs.hue_divisions, hs.sat_divisions, hs.val_divisions),
(90, 30, 1)
);
assert!(p.hue_sat[1].is_some());
let look = p.look.as_ref().unwrap();
assert_eq!(
(look.hue_divisions, look.sat_divisions, look.val_divisions),
(36, 8, 16)
);
assert!(p.tone_curve.is_none());
assert!(p.may_copy());
let back = Dcp::parse(&p.to_bytes().unwrap()).unwrap();
assert_eq!(
back.hue_sat, p.hue_sat,
"a copied profile keeps its tables bit for bit"
);
assert_eq!(back.look, p.look);
assert!(
back.is_for(None, "Canon", "EOS 6D"),
"and so applies to the body's CR2s"
);
}
#[test]
fn decoding_the_six_d_dng_hands_on_its_tables() {
let Some(bytes) = six_d_dng() else { return };
let raw = crate::decode(&bytes).unwrap();
let tables = raw.profile_tables.expect("tables");
assert_eq!(tables.origin, ProfileOrigin::Embedded);
assert_eq!(tables.name, "Adobe Standard");
assert!(tables.hue_sat.is_some() && tables.look.is_some());
assert!(
tables.tone_curve.is_none(),
"Adobe Standard has no curve of its own"
);
assert_eq!(raw.baseline_exposure, 0.25);
}
#[test]
fn a_profile_brings_its_own_matrices() {
let p = sample();
let cam = p.camera_profile(Some([0.5, 1.0, 0.7])).unwrap();
assert_eq!(cam.calibrations().len(), 2);
assert!(cam.calibrations().iter().all(|c| c.forward.is_some()));
assert!(cam.cam_to_srgb().is_some());
}
}
+49 -2
View File
@@ -16,6 +16,7 @@
//! second decoder can be put behind them without changing any of them
//! (FR-RAW-2). [`Rawler`] is the one that ships; [`default`] hands it out.
pub mod dcp;
mod decoder;
mod error;
mod locate;
@@ -137,8 +138,23 @@ pub struct RawImage {
/// carry on: calibrations and the as-shot neutral. `None` for a body the
/// decoder has no matrix for.
pub profile: Option<profile::CameraProfile>,
/// The body, as rawler cleans the names: what `Make`/`Model` say and what
/// the base-curve database matches on.
/// TRACES: FR-DEV-3e
/// The camera profile's HueSatMap and LookTable, resolved for this frame
/// (D20): embedded in the DNG, or from a matched `.dcp`. `None` renders
/// through the matrix alone.
///
/// Carried with the image, as `color_matrix` is, so that every path that
/// renders a decoded file renders it through the same profile without
/// having to be told — see camera-profiles.md §3.
pub profile_tables: Option<std::sync::Arc<dr_types::ProfileTables>>,
/// TRACES: FR-DEV-3e
/// Stops the render adds before anything else: the file's
/// `BaselineExposure` plus the profile's `BaselineExposureOffset`
/// (camera-profiles.md §11). Applied by the GPU side as a gain on the
/// camera matrix; `color_matrix` itself stays the file's.
pub baseline_exposure: f32,
/// The body, as rawler cleans the names: what `Make`/`Model` say, and
/// what a `.dcp`'s `UniqueCameraModel` is matched against.
pub make: String,
pub model: String,
}
@@ -534,6 +550,17 @@ pub(crate) fn parse_exif_offset(s: &str) -> Option<i32> {
Some(sign * (h * 60 + m))
}
/// TRACES: FR-DEV-3g
/// The DNG `NoiseProfile` of a file, if it carries one: `(S, O)` per CFA
/// colour plane, variance `S·x + O` in black-to-white normalised units. See
/// [`profile::read_noise_profile`]. Reads the header, not the image.
pub fn noise_profile(bytes: &[u8]) -> Option<Vec<(f32, f32)>> {
use rawler::rawsource::RawSource;
let source = RawSource::new_from_slice(bytes);
let decoder = rawler::get_decoder(&source).ok()?;
profile::read_noise_profile(decoder.as_ref())
}
/// TRACES: FR-RAW-3 | FR-EXP-9
/// Fully decode sensor data.
///
@@ -571,6 +598,24 @@ fn decode_unguarded(bytes: &[u8]) -> Result<RawImage, DecodeError> {
// is now structural, because there is only one interpolated matrix and
// both callers ask the same object for it.
let profile = profile::CameraProfile::extract(&image, &dng);
// TRACES: FR-DEV-3e
// The tables, and — where a `.dcp` supplies them — the matrices they were
// built against, which then stand in for the file's (D20).
let root = decoder
.ifd(rawler::decoders::WellKnownIFD::Root)
.ok()
.flatten();
let dcp::Resolved {
profile,
tables: profile_tables,
baseline_exposure,
..
} = dcp::resolve(
profile,
root.as_deref(),
&image.camera.clean_make,
&image.camera.clean_model,
);
let color_matrix = profile.as_ref().and_then(|p| p.cam_to_srgb());
let wb_coeffs = sane_wb(
image.wb_coeffs,
@@ -661,6 +706,8 @@ fn decode_unguarded(bytes: &[u8]) -> Result<RawImage, DecodeError> {
color_matrix,
samples_per_pixel,
profile,
profile_tables,
baseline_exposure,
make: image.camera.clean_make.clone(),
model: image.camera.clean_model.clone(),
})
+30 -1
View File
@@ -521,7 +521,7 @@ fn cct_from_xy(x: f32, y: f32) -> f32 {
/// but a profile calibrated under fluorescent light is describing a sensor
/// under fluorescent light, and placing it at roughly the right colour is much
/// better than discarding it.
fn illuminant_temperature(illuminant: Illuminant) -> Option<f32> {
pub(crate) fn illuminant_temperature(illuminant: Illuminant) -> Option<f32> {
Some(match illuminant {
// CIE standard illuminant A: a tungsten filament at 2856 K. The low
// end of essentially every dual-illuminant profile ever written.
@@ -740,6 +740,35 @@ pub fn read_dng_matrices(decoder: &dyn rawler::decoders::Decoder) -> DngMatrices
}
}
/// TRACES: FR-DEV-3g
/// The DNG `NoiseProfile` tag (51041): the converter's measured noise for
/// this body at this ISO, as `(S, O)` per CFA colour plane, so that a
/// photosite's variance is `S·x + O` with `x` normalised black-to-white.
///
/// One pair means all planes share it. `None` where the file has no such
/// tag — every proprietary raw, and DNGs from converters that do not measure
/// — or where a value is not a finite non-negative number. The learned
/// denoise's second-best noise source (denoise.md §3.3), after a measured
/// table for the body.
pub fn read_noise_profile(decoder: &dyn rawler::decoders::Decoder) -> Option<Vec<(f32, f32)>> {
use rawler::decoders::WellKnownIFD;
use rawler::tags::DngTag;
let ifd = decoder.ifd(WellKnownIFD::Root).ok()??;
let entry = ifd.get_entry_recursive(DngTag::NoiseProfile)?;
let n = entry.count() as usize;
if n < 2 || !n.is_multiple_of(2) {
return None;
}
let pairs: Vec<(f32, f32)> = (0..n / 2)
.map(|i| (entry.force_f32(2 * i), entry.force_f32(2 * i + 1)))
.collect();
pairs
.iter()
.all(|(s, o)| s.is_finite() && o.is_finite() && *s >= 0.0 && *o >= 0.0)
.then_some(pairs)
}
#[cfg(test)]
mod tests {
use super::*;
+32
View File
@@ -0,0 +1,32 @@
[package]
name = "dr-denoise"
version.workspace = true
edition.workspace = true
rust-version.workspace = true
license.workspace = true
[dependencies]
dr-decode.workspace = true
serde = { workspace = true }
serde_norway.workspace = true
thiserror.workspace = true
log.workspace = true
# The network runs under the inference engine like every other model
# (docs/dev/inference.md): `ort` is the API, the engine picks the rung.
# Optional so the noise model and the tiling test without a runtime.
ort = { workspace = true, optional = true }
dr-inference-engine = { workspace = true, optional = true }
ndarray = { workspace = true, optional = true }
[features]
default = ["onnx"]
onnx = ["dep:ort", "dep:dr-inference-engine", "dep:ndarray"]
# A real ONNX Runtime from disk rather than tract alone, as the app links it.
native = ["onnx", "dr-inference-engine/native"]
[dev-dependencies]
# The example repairs hot photosites with the app's own pass, as develop will.
dr-gpu.workspace = true
pollster.workspace = true
env_logger.workspace = true
+159
View File
@@ -0,0 +1,159 @@
//! Denoise one RAW file end to end, as develop will, and time it.
//!
//! ```sh
//! DARKROOM_ORT_DIR=~/.local/share/darkroom/runtime \
//! cargo run --release -p dr-denoise --features native --example denoise_raw -- IMG.CR2 out [fast|best]
//! ```
//!
//! Decode, the app's hot-pixel pass, the frame's noise from its best source,
//! then one of the shipped networks (`best` unless named) under the inference engine on whatever rung this
//! machine probes to. Writes `out.npy` — the active area, `h×w×3` f32 linear
//! camera RGB — for comparison with the training repo's own path
//! (`tools/compare_rust.py` in darkroom-denoise). `DARKROOM_ORT_DIR` points
//! at an ONNX Runtime build; the engine's cache goes to `DR_ENGINE_CACHE` or
//! a temporary directory. The whole-frame network (`mosaic-hq.onnx` beside
//! the fixed file) runs where the rung takes any size; `DR_PLAN=tiles` keeps
//! the 1408² tiles anyway, to compare the two.
use std::path::PathBuf;
use std::time::{Duration, Instant};
use dr_denoise::onnx::OnnxNet;
use dr_inference_engine::{Config, Role};
fn main() {
env_logger::Builder::from_env(env_logger::Env::default().default_filter_or("warn")).init();
let mut args = std::env::args().skip(1);
let (Some(input), Some(out)) = (args.next(), args.next()) else {
eprintln!("usage: denoise_raw RAW OUT_PREFIX [fast|best]");
std::process::exit(2);
};
let shipped = match args.next().as_deref() {
None | Some("best") => dr_denoise::BEST,
Some("fast") => dr_denoise::FAST,
Some(other) => {
eprintln!("no network called {other}: fast or best");
std::process::exit(2);
}
};
let model = PathBuf::from(env!("CARGO_MANIFEST_DIR"))
.join("../../models/denoise")
.join(shipped.file);
let whole = model.with_file_name(shipped.whole);
let tiles_only = std::env::var("DR_PLAN").is_ok_and(|p| p == "tiles");
let mut models = vec![(Role::Denoiser, model.clone())];
if whole.is_file() && !tiles_only {
models.push((Role::WholeDenoiser, whole));
}
let cache = std::env::var_os("DR_ENGINE_CACHE")
.map(PathBuf::from)
.unwrap_or_else(|| std::env::temp_dir().join("dr-denoise-engines"));
let started = Instant::now();
dr_inference_engine::init(Config {
runtime_dirs: std::env::var_os("DARKROOM_ORT_DIR")
.map(PathBuf::from)
.into_iter()
.collect(),
cache_dir: cache,
models,
embedded: Vec::new(),
ceiling: None,
threads: 0,
decay: Duration::ZERO,
});
// Wait for the probe and the engine build, so the timing below is the
// rung this machine settles on, not the fallback used while it compiles.
// The probe starts on its own thread; give it a moment to say so.
std::thread::sleep(Duration::from_secs(1));
loop {
let s = dr_inference_engine::status();
if !s.probing && s.engines.0 >= s.engines.1 {
println!(
"engine {} ({:.1} s to settle)",
s.line(),
started.elapsed().as_secs_f64()
);
break;
}
std::thread::sleep(Duration::from_millis(200));
}
let bytes = std::fs::read(&input).expect("read raw");
let t = Instant::now();
let mut raw = dr_decode::decode(&bytes).expect("decode");
let meta = dr_decode::metadata(&bytes).expect("metadata");
let decode = t.elapsed();
let t = Instant::now();
let ctx =
pollster::block_on(dr_gpu::GpuContext::new_headless()).expect("GPU for the hot-pixel pass");
let repaired = dr_gpu::Demosaicer::new(&ctx)
.expect("demosaicer")
.repair_hot_pixels(&mut raw)
.expect("repair");
let repair = t.elapsed();
let noise = dr_denoise::noise::for_frame(&raw, &bytes, meta.iso)
.expect("no noise source for this frame");
println!(
"frame {} {} ISO {:?}, {}×{}, {:?}, {repaired} hot photosites repaired",
raw.make, raw.model, meta.iso, raw.crop.width, raw.crop.height, raw.cfa_pattern
);
println!(
"noise {} — σ at 10 % grey (G) {:.5}, read {:.5}, row {:.5}, col {:.5}",
noise.source.label(),
noise.sigma(1, 0.1),
noise.o[1].sqrt(),
noise.row,
noise.col
);
let mut net = if tiles_only {
OnnxNet::open_tiled(&model, shipped)
} else {
OnnxNet::open(&model, shipped)
}
.expect("model");
println!(
"rung {} · {}",
net.rung().map(|r| r.label()).unwrap_or("?"),
if net.whole_frame() {
"whole frame"
} else {
"1408² tiles"
}
);
let t = Instant::now();
let rgb = dr_denoise::denoise(&raw, &noise, &mut net, &mut |done, total| {
eprint!("\rtile {done}/{total}");
true
})
.expect("denoise")
.expect("not cancelled");
let run = t.elapsed();
eprintln!();
println!(
"time decode {:.2} s · hot pixels {:.2} s · network {:.2} s ({:.1} MP)",
decode.as_secs_f64(),
repair.as_secs_f64(),
run.as_secs_f64(),
(raw.crop.width * raw.crop.height) as f64 / 1e6
);
let (h, w) = (raw.crop.height as usize, raw.crop.width as usize);
let mut npy = Vec::with_capacity(rgb.len() * 4 + 128);
let mut header =
format!("{{'descr': '<f4', 'fortran_order': False, 'shape': ({h}, {w}, 3), }}");
while (10 + header.len() + 1) % 64 != 0 {
header.push(' ');
}
header.push('\n');
npy.extend_from_slice(b"\x93NUMPY\x01\x00");
npy.extend_from_slice(&(header.len() as u16).to_le_bytes());
npy.extend_from_slice(header.as_bytes());
for v in &rgb {
npy.extend_from_slice(&v.to_le_bytes());
}
std::fs::write(format!("{out}.npy"), npy).expect("write");
println!("wrote {out}.npy");
}
+119
View File
@@ -0,0 +1,119 @@
//! TRACES: FR-DEV-3g
//! Learned demosaic and denoise on the raw mosaic (docs/dev/denoise.md).
//!
//! A network trained on the library's own base-ISO raws with the 6D's
//! measured noise added takes the repaired, normalised mosaic and a σ for
//! every photosite, and returns linear camera RGB at full resolution — the
//! texture the classical demosaic would have produced, with the noise gone.
//! It replaces the demosaic box; nothing downstream changes (§2).
//!
//! - [`noise`] says how noisy each photosite is, from the best source the
//! frame has.
//! - [`tile`] runs a fixed-shape network over a whole frame, exactly.
//! - [`onnx`] is that network under the inference engine.
//!
//! The input must already have been through the app's hot-pixel pass
//! (`dr_gpu::Demosaicer::repair_hot_pixels`): the noise model was fitted
//! with what that pass removes left out.
pub mod noise;
#[cfg(feature = "onnx")]
pub mod onnx;
pub mod repair;
pub mod tile;
use dr_decode::RawImage;
pub use noise::{NoiseModel, Source};
pub use tile::{Sizes, TileNet, HALO};
/// TRACES: FR-DEV-3g
/// A network the app ships in `models/denoise/`: its fixed-tile file, the
/// same network with any height and width for a whole frame (§14), and the
/// context it needs past a tile's kept centre (§13).
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub struct Shipped {
pub file: &'static str,
pub whole: &'static str,
pub halo: usize,
}
/// The smallest student: 0.9 M parameters, 11 GMAC a megapixel.
pub const FAST: Shipped = Shipped {
file: "mosaic-fast-1408.onnx",
whole: "mosaic-fast.onnx",
halo: HALO,
};
/// One network of the first release's shape, 3.2 M parameters and 48 GMAC a
/// megapixel, taught by the mixture of experts that was Best until 0.24:
/// its edges at a third of its work (denoise.md §15). A new file name, not
/// the old Medium's or Best's: the result cache keys a model by its name
/// and size, and this one is byte for byte the old Medium's size.
pub const BEST: Shipped = Shipped {
file: "mosaic-hq-1408.onnx",
whole: "mosaic-hq.onnx",
halo: HALO,
};
#[derive(Debug, thiserror::Error)]
pub enum DenoiseError {
#[error("the network cannot take this photograph: {0}")]
Unsupported(String),
#[error("the denoise model misbehaved: {0}")]
Model(String),
#[error("could not read the denoise model: {0}")]
ModelRead(#[from] std::io::Error),
#[cfg(feature = "onnx")]
#[error(transparent)]
Engine(#[from] dr_inference_engine::Error),
#[cfg(feature = "onnx")]
#[error(transparent)]
Ort(#[from] ort::Error),
}
/// Whether the learned stage can take this frame at all: a Bayer mosaic.
/// X-Trans needs its own model (§9); a linear DNG has no photosites.
pub fn eligible(raw: &RawImage) -> bool {
raw.samples_per_pixel == 1 && tile::rggb_offset(raw.cfa_pattern).is_some()
}
/// The active area of `raw`, denoised and demosaiced: `crop.height ×
/// crop.width` interleaved RGB, linear camera space, normalised black 0 and
/// white 1 per photosite as the classical demosaic normalises.
///
/// `raw` must be hot-pixel repaired. `None` when `progress` stopped it.
pub fn denoise(
raw: &RawImage,
noise: &NoiseModel,
net: &mut dyn TileNet,
progress: &mut dyn FnMut(usize, usize) -> bool,
) -> Result<Option<Vec<f32>>, DenoiseError> {
if !eligible(raw) {
return Err(DenoiseError::Unsupported(format!(
"{:?} with {} samples per photosite",
raw.cfa_pattern, raw.samples_per_pixel
)));
}
let active = noise::active(raw);
let (h, w) = (active.h, active.w);
// The active area laid out once, then the noise-aware repair the model
// was trained behind (see `repair`).
let mut mosaic: Vec<f32> = (0..h * w).map(|i| active.at(i / w, i % w)).collect();
let pattern = raw.cfa_pattern;
let repaired = repair::repair(&mut mosaic, h, w, repair::REPAIR_K, &|y, x, v| {
noise.sigma(pattern.colour_at(x as u32, y as u32) as usize, v)
});
log::info!(
"learned denoise: {repaired} photosites beyond {}σ of every neighbour repaired",
repair::REPAIR_K
);
tile::run_tiled(
net,
h,
w,
raw.cfa_pattern,
&|y, x| mosaic[y * w + x],
&|c, v| noise.sigma(c, v),
progress,
)
}
+407
View File
@@ -0,0 +1,407 @@
//! TRACES: FR-DEV-3g
//! How noisy each photosite is: the network is told, not left to guess
//! (denoise.md §3.3).
//!
//! The model is `σ² = S·x + O + row² + col²` per photosite, `x` the signal
//! normalised black-to-white the way the demosaic normalises it. Three
//! sources, best first:
//!
//! 1. **A measured table** for the body ([`Source::Table`]) — the Canon EOS 6D
//! today, from the library's own frames.
//! 2. **The DNG's `NoiseProfile`** ([`Source::DngProfile`]) — what Adobe's
//! converter measured for the body at that ISO.
//! 3. **The frame itself** ([`Source::Measured`]) — read, row and column
//! noise from its masked border, which is a dark frame taken in the same
//! instant, and only the shot gain estimated, from the quietest flat
//! patches. Checked against the 6D's table on 130 frames: within ±10 % at
//! ISO 1000 and above, scattered below; the network loses under 0.3 dB for
//! a σ off by 15–20 %, and over-estimating costs half what
//! under-estimating does, so the estimate leans high.
//!
//! Row and column noise come from the masked border whenever the frame has
//! one, whatever the source of the rest.
use dr_decode::{CfaPattern, RawImage};
use serde::Deserialize;
/// Where a frame's noise figures came from, for develop to say.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub enum Source {
Table,
DngProfile,
Measured,
}
impl Source {
pub fn label(self) -> &'static str {
match self {
Source::Table => "measured for this camera",
Source::DngProfile => "from the DNG's noise profile",
Source::Measured => "estimated from this photograph",
}
}
}
/// Per-photosite noise in the frame's own normalisation (black 0, white 1).
#[derive(Clone, Debug, PartialEq)]
pub struct NoiseModel {
/// Shot gain per colour, R G B.
pub s: [f32; 3],
/// Read variance per colour, R G B.
pub o: [f32; 3],
/// Standard deviation shared by a whole row, and by a whole column.
pub row: f32,
pub col: f32,
pub source: Source,
}
impl NoiseModel {
/// σ for a photosite of colour `c` (0 R, 1 G, 2 B) reading `x`.
#[inline]
pub fn sigma(&self, c: usize, x: f32) -> f32 {
(self.s[c] * x.max(0.0) + self.o[c] + self.row * self.row + self.col * self.col).sqrt()
}
/// The same figures scaled for the Amount the spec describes (§3.3):
/// above 1 tells the network there is more noise than there is.
pub fn scaled(&self, amount: f32) -> NoiseModel {
let a2 = amount * amount;
NoiseModel {
s: self.s.map(|v| v * a2),
o: self.o.map(|v| v * a2),
row: self.row * amount,
col: self.col * amount,
source: self.source,
}
}
}
/// The frame's noise, from the best source it has.
///
/// `bytes` is the file (for a DNG's `NoiseProfile`), `iso` its EXIF ISO.
/// `None` only for a frame with no masked border, no profile and no table
/// that is also too dark or too busy to measure.
pub fn for_frame(raw: &RawImage, bytes: &[u8], iso: Option<u32>) -> Option<NoiseModel> {
for_frame_with(raw, dr_decode::noise_profile(bytes).as_deref(), iso)
}
/// [`for_frame`], given the file's `NoiseProfile` already read
/// ([`dr_decode::noise_profile`]) rather than the file, for a caller that
/// keeps the header's answer and not the bytes.
pub fn for_frame_with(
raw: &RawImage,
profile: Option<&[(f32, f32)]>,
iso: Option<u32>,
) -> Option<NoiseModel> {
let dark = dark_border(raw);
let mut model = iso
.and_then(|iso| from_table(raw, iso))
.or_else(|| profile.and_then(|p| from_dng_profile(raw, p)))
.or_else(|| measured(raw, dark.as_ref()))?;
if let Some(d) = dark {
// The border saw this exposure's row and column noise directly.
if model.source != Source::Table {
model.row = d.row;
model.col = d.col;
}
}
Some(model)
}
#[derive(Deserialize)]
struct Table {
make: String,
model: String,
rows: Vec<TableRow>,
}
#[derive(Deserialize)]
struct TableRow {
iso: u32,
s_dn: [f32; 4],
o_dn: [f32; 4],
row_dn: f32,
col_dn: f32,
}
const TABLES: &[&str] = &[include_str!("../tables/canon-eos-6d.yaml")];
/// The body's measured table at the nearest ISO it holds, converted from DN
/// to this frame's normalisation.
pub fn from_table(raw: &RawImage, iso: u32) -> Option<NoiseModel> {
let table = TABLES.iter().find_map(|t| {
let t: Table = serde_norway::from_str(t).ok()?;
(t.make.eq_ignore_ascii_case(&raw.make) && t.model.eq_ignore_ascii_case(&raw.model))
.then_some(t)
})?;
let row = table.rows.iter().min_by(|a, b| {
let d = |r: &TableRow| ((r.iso as f32).ln() - (iso as f32).ln()).abs();
d(a).total_cmp(&d(b))
})?;
let span = span(raw);
// RGGB positions → colours: the greens share.
let s = [row.s_dn[0], 0.5 * (row.s_dn[1] + row.s_dn[2]), row.s_dn[3]].map(|v| v / span);
let o =
[row.o_dn[0], 0.5 * (row.o_dn[1] + row.o_dn[2]), row.o_dn[3]].map(|v| v / (span * span));
Some(NoiseModel {
s,
o,
row: row.row_dn / span,
col: row.col_dn / span,
source: Source::Table,
})
}
/// A DNG's `NoiseProfile`: one pair for every plane, or one per colour plane
/// (R, G, B for a Bayer DNG), already in the file's black-to-white units —
/// which are the units `dr-decode` normalises by.
pub fn from_dng_profile(raw: &RawImage, pairs: &[(f32, f32)]) -> Option<NoiseModel> {
if raw.cfa_pattern.is_xtrans() || raw.samples_per_pixel != 1 {
return None;
}
let (s, o) = match pairs {
[(s, o)] => ([*s; 3], [*o; 3]),
[r, g, b, ..] => ([r.0, g.0, b.0], [r.1, g.1, b.1]),
_ => return None,
};
Some(NoiseModel {
s,
o,
row: 0.0,
col: 0.0,
source: Source::DngProfile,
})
}
/// Read, row and column noise measured on the masked border, normalised.
#[derive(Clone, Copy, Debug)]
pub struct Dark {
pub read: f32,
pub row: f32,
pub col: f32,
}
/// The optically black photosites beside and above the active area.
///
/// Keeps well clear of the active area: on the 6D the dozen columns nearest
/// it see light. Photosites over 8σ are the strip's own hot photosites — the
/// same ones in every frame — and are left out, as the app's hot-pixel pass
/// removes their kin before the network sees them.
pub fn dark_border(raw: &RawImage) -> Option<Dark> {
let (x0, y0, w, h) = (
raw.crop.x as usize,
raw.crop.y as usize,
raw.crop.width as usize,
raw.crop.height as usize,
);
let stride = raw.width as usize;
let span = span(raw);
if x0 < 40 || raw.samples_per_pixel != 1 {
return None;
}
let cols = 4..x0 - 16;
let nc = cols.len() as f32;
// Residual after removing each row's mean and each column's mean.
let mut row_means = Vec::with_capacity(h);
let mut col_sum = vec![0.0f64; cols.len()];
for y in y0..y0 + h {
let line = &raw.data[y * stride..y * stride + x0];
let m = cols.clone().map(|x| line[x] as f32).sum::<f32>() / nc;
row_means.push(m);
for (k, x) in cols.clone().enumerate() {
col_sum[k] += (line[x] as f32 - m) as f64;
}
}
let col_mean: Vec<f32> = col_sum.iter().map(|s| (*s / h as f64) as f32).collect();
let resid = |y: usize, k: usize, x: usize| {
raw.data[y * stride + x] as f32 - row_means[y - y0] - col_mean[k]
};
let (mut s1, mut n) = (0.0f64, 0usize);
for y in y0..y0 + h {
for (k, x) in cols.clone().enumerate() {
s1 += (resid(y, k, x) as f64).powi(2);
n += 1;
}
}
let rough = (s1 / n as f64).sqrt() as f32;
let (mut s2, mut n2) = (0.0f64, 0usize);
for y in y0..y0 + h {
for (k, x) in cols.clone().enumerate() {
let r = resid(y, k, x);
if r.abs() < 8.0 * rough {
s2 += (r as f64).powi(2);
n2 += 1;
}
}
}
let read = (s2 / n2.max(1) as f64).sqrt() as f32;
let rm = row_means.iter().sum::<f32>() / h as f32;
let row_var = row_means.iter().map(|m| (m - rm).powi(2)).sum::<f32>() / h as f32;
let row = (row_var - read * read / nc).max(0.0).sqrt();
// Columns: the masked rows above the image span every column.
let col = if y0 >= 24 {
let rows = 4..y0 - 12;
let nr = rows.len() as f32;
let means: Vec<f32> = (x0..x0 + w)
.map(|x| {
rows.clone()
.map(|y| raw.data[y * stride + x] as f32)
.sum::<f32>()
/ nr
})
.collect();
let mm = means.iter().sum::<f32>() / means.len() as f32;
let var = means.iter().map(|m| (m - mm).powi(2)).sum::<f32>() / means.len() as f32;
(var - read * read / nr).max(0.0).sqrt()
} else {
0.0
};
Some(Dark {
read: read / span,
row: row / span,
col: col / span,
})
}
/// The quietest-third bias of the patch variance, and the residual bias the
/// estimate showed against the 6D's table (0.91 at the median), in one: the
/// estimate is divided by this.
const QUIET_FACTOR: f32 = 0.85 * 0.91;
/// The frame's own noise: read noise from the border (or, lacking one, the
/// floor of the quietest patches), shot gain from flat patches of one green
/// plane, the same for every colour, as a sensor's gain is.
pub fn measured(raw: &RawImage, dark: Option<&Dark>) -> Option<NoiseModel> {
if raw.cfa_pattern.is_xtrans() || raw.samples_per_pixel != 1 {
return None;
}
let m = active(raw);
let (h, w) = (m.h, m.w);
// One green plane at a two-photosite pitch.
let (gy, gx) = green_offset(raw.cfa_pattern)?;
let ph = (h - gy) / 2;
let pw = (w - gx) / 2;
let g = |y: usize, x: usize| m.at(gy + 2 * y, gx + 2 * x);
const B: usize = 8;
let mut patches: Vec<(f32, f32)> = Vec::new(); // (level, variance)
for by in 0..ph / B {
for bx in 0..(pw - 2) / B {
let (mut s, mut s2, mut lv) = (0.0f32, 0.0f32, 0.0f32);
for y in by * B..by * B + B {
for x in bx * B..bx * B + B {
// Second difference: cancels any gradient; var = 6σ².
let d = g(y, x + 2) - 2.0 * g(y, x + 1) + g(y, x);
s += d;
s2 += d * d;
lv += g(y, x + 1);
}
}
let n = (B * B) as f32;
let var = (s2 / n - (s / n).powi(2)) / 6.0;
patches.push((lv / n, var));
}
}
let floor = dark.map(|d| d.read);
let lo = 4.0 * floor.unwrap_or(0.002);
patches.retain(|(l, _)| *l > lo && *l < 0.7);
if patches.len() < 500 {
return None;
}
patches.sort_by(|a, b| a.0.total_cmp(&b.0));
let bins = 12;
let per = patches.len() / bins;
let mut ests = Vec::new();
let mut floors = Vec::new();
for b in 0..bins {
let mut bin: Vec<(f32, f32)> = patches[b * per..(b + 1) * per].to_vec();
if bin.len() < 60 {
continue;
}
bin.sort_by(|a, b| a.1.total_cmp(&b.1));
let quiet = &bin[..bin.len() / 3];
let read2 = floor.map(|r| r * r);
let mut e: Vec<f32> = quiet
.iter()
.map(|(l, v)| (v / QUIET_FACTOR - read2.unwrap_or(0.0)) / l)
.collect();
e.sort_by(f32::total_cmp);
ests.push(e[e.len() / 2]);
floors.push(quiet[quiet.len() / 2]);
}
ests.sort_by(f32::total_cmp);
let s = *ests.get(ests.len() / 2)?;
if !(s.is_finite() && s > 0.0) {
return None;
}
// No border: the read variance is what the darkest bin leaves unexplained.
let read2 = match floor {
Some(r) => r * r,
None => {
let (l, v) = floors.first().copied()?;
(v / QUIET_FACTOR - s * l).max(1e-9)
}
};
Some(NoiseModel {
s: [s; 3],
o: [read2; 3],
row: dark.map_or(0.0, |d| d.row),
col: dark.map_or(0.0, |d| d.col),
source: Source::Measured,
})
}
/// Black-to-white range of the frame, as the demosaic normalises it.
pub(crate) fn span(raw: &RawImage) -> f32 {
let black = raw.black_level.iter().map(|&b| b as f32).sum::<f32>() / 4.0;
(raw.white_level as f32 - black).max(1.0)
}
/// Where a green photosite sits in the pattern's 2×2 cell, (dy, dx).
fn green_offset(p: CfaPattern) -> Option<(usize, usize)> {
match p {
CfaPattern::Rggb | CfaPattern::Bggr => Some((0, 1)),
CfaPattern::Grbg | CfaPattern::Gbrg => Some((0, 0)),
_ => None,
}
}
/// The active area, normalised, read lazily.
pub(crate) struct Active<'a> {
raw: &'a RawImage,
black: [f32; 4],
inv: [f32; 4],
pub h: usize,
pub w: usize,
}
impl Active<'_> {
/// Photosite (y, x) of the active area, black 0, white 1.
#[inline]
pub fn at(&self, y: usize, x: usize) -> f32 {
let c = (y & 1) * 2 + (x & 1);
let v = self.raw.data[(self.raw.crop.y as usize + y) * self.raw.width as usize
+ self.raw.crop.x as usize
+ x];
(v as f32 - self.black[c]) * self.inv[c]
}
}
/// Black levels per position of the crop's 2×2 cell, as the demosaic reads
/// them: one reported level is broadcast.
pub(crate) fn active(raw: &RawImage) -> Active<'_> {
let b = raw.black_level;
let black = if b[1] == 0 && b[2] == 0 && b[3] == 0 {
[b[0] as f32; 4]
} else {
b.map(|v| v as f32)
};
let inv = black.map(|bl| 1.0 / (raw.white_level as f32 - bl).max(1.0));
Active {
raw,
black,
inv,
h: raw.crop.height as usize,
w: raw.crop.width as usize,
}
}
+142
View File
@@ -0,0 +1,142 @@
//! TRACES: FR-DEV-3g
//! The denoise network under the inference engine.
//!
//! The shipped export takes `mosaic` and `sigma`, `1×1×1408×1408`, and
//! returns `rgb`, `1×3×1408×1408` (darkroom-denoise `denoise/export.py`,
//! fixed shape because every model the engine runs is). The engine picks the
//! rung: fp16 on TensorRT and MIGraphX, which measured 0.00 dB from f32; f32
//! on CUDA and the CPU; on the Hexagon the `.a16w16.onnx` sibling, 16-bit
//! activations and weights, 0.00 dB from f32 on the tablet itself where int8
//! lost 5–9 dB (docs/dev/inference.md §1.5). That sibling is the same network
//! with the Bayer packing spelled `SpaceToDepth`, which QNN can hold and the
//! 6-D reshape it replaces it cannot.
//!
//! Each network also ships with any height and width (`mosaic-hq.onnx`
//! beside `mosaic-hq-1408.onnx`, darkroom-denoise `tools/export_whole.py`,
//! identical to the fixed file at 1408²). Where the rung takes any size, the
//! frame runs whole instead of in tiles whose borders are thrown away — a
//! 1408² tile keeps 1024², 1.89 photosites computed for each one kept
//! (denoise.md §14).
use crate::tile::{Sizes, TileNet};
use crate::{DenoiseError, Shipped};
use dr_inference_engine::{Form, Model, Role};
/// The edge of the tile the shipped fixed-shape export takes.
pub const TILE: usize = 1408;
/// What a whole-frame input's sides must be multiples of: the networks pack
/// 2×2 and halve three times, so a side is a whole number of positions at
/// their coarsest level only in steps of 16.
pub const ALIGN: usize = 16;
pub struct OnnxNet {
model: Model,
sizes: Sizes,
halo: usize,
}
impl OnnxNet {
/// The network `shipped`, whose fixed-tile file is at `path`.
///
/// On a rung that runs any input size (TensorRT, the CUDA provider —
/// [`dr_inference_engine::whole_frame_limit`]) and with the any-size
/// export installed beside it, the whole-frame network: the frame in one
/// call, or the fewest large tiles that fit (§14). Its output is the
/// fixed tiles' to rounding. Everywhere else, and if the whole-frame
/// model will not open, the 1408² tiles.
pub fn open(path: &std::path::Path, shipped: Shipped) -> Result<Self, DenoiseError> {
if let Some(max) = dr_inference_engine::whole_frame_limit() {
let whole = path.with_file_name(shipped.whole);
if whole.is_file() {
let opened = std::fs::read(&whole)
.map_err(DenoiseError::from)
.and_then(|bytes| {
Ok(dr_inference_engine::open(
Role::WholeDenoiser,
Form::F32,
&bytes,
)?)
});
match opened {
Ok(model) => {
return Ok(OnnxNet {
model,
sizes: Sizes::Any { align: ALIGN, max },
halo: shipped.halo,
})
}
Err(e) => log::warn!(
"learned denoise: {} will not open ({e}); running 1408² tiles",
whole.display()
),
}
}
}
Self::open_tiled(path, shipped)
}
/// The fixed-tile network at `path`, whatever the rung: 1408² tiles.
pub fn open_tiled(path: &std::path::Path, shipped: Shipped) -> Result<Self, DenoiseError> {
let (path, form) = dr_inference_engine::resolve_model(Role::Denoiser, path);
let bytes = std::fs::read(&path)?;
Ok(OnnxNet {
model: dr_inference_engine::open(Role::Denoiser, form, &bytes)?,
sizes: Sizes::Square(TILE),
halo: shipped.halo,
})
}
/// Where it runs, for a status line.
pub fn rung(&self) -> Result<dr_inference_engine::Rung, DenoiseError> {
Ok(self.model.acquire()?.rung())
}
/// Whether this is the whole-frame network.
pub fn whole_frame(&self) -> bool {
matches!(self.sizes, Sizes::Any { .. })
}
}
impl TileNet for OnnxNet {
fn sizes(&self) -> Sizes {
self.sizes
}
fn halo(&self) -> usize {
self.halo
}
fn run(
&mut self,
rows: usize,
cols: usize,
mosaic: Vec<f32>,
sigma: Vec<f32>,
write: &mut dyn FnMut(&[f32]),
) -> Result<(), DenoiseError> {
let shape = ndarray::IxDyn(&[1, 1, rows, cols]);
// The vectors become the tensors: no copy on the way in.
let m = ort::value::Tensor::from_array(
ndarray::Array::from_shape_vec(shape.clone(), mosaic)
.map_err(|e| DenoiseError::Model(e.to_string()))?,
)?;
let s = ort::value::Tensor::from_array(
ndarray::Array::from_shape_vec(shape, sigma)
.map_err(|e| DenoiseError::Model(e.to_string()))?,
)?;
let acquired = self.model.acquire()?;
let mut session = acquired.lock();
let outputs = session.run(ort::inputs!["mosaic" => m, "sigma" => s])?;
let (shape, data) = outputs[0].try_extract_tensor::<f32>()?;
let dims: Vec<i64> = shape.iter().copied().collect();
if dims != [1, 3, rows as i64, cols as i64] {
return Err(DenoiseError::Model(format!(
"output is {dims:?}, expected [1, 3, {rows}, {cols}]"
)));
}
// And none on the way out: the frame is written from the runtime's buffer.
write(data);
Ok(())
}
}
+145
View File
@@ -0,0 +1,145 @@
//! TRACES: FR-DEV-3g
//! Hot and dead photosites, judged against the noise, before the network.
//!
//! The app's own pass (`dr_gpu::Demosaicer::repair_hot_pixels`) runs first and
//! takes the gross defects. At high ISO it leaves thousands of photosites per
//! 6D frame more than 8σ beyond every neighbour, which the network turns into
//! specks. This second pass uses that pass's two tests with the threshold in
//! units of the photosite's own σ from the noise model:
//!
//! - beyond every same-colour neighbour (two photosites away, the 3×3 of its
//! plane) by more than `k·σ`, and
//! - beyond every adjacent photosite, whatever its colour, by more than
//! `k·σ` **and** by a factor of two — what keeps a real point of light,
//! which lights its neighbours through the lens and the anti-aliasing
//! filter. A margin in σ alone is not enough: on a bright star 8σ is a
//! sliver of the signal, and the star would be flattened.
//!
//! A hot one becomes its brightest same-colour neighbour, a dead one its
//! darkest. The shipped model was trained on input repaired exactly so
//! (darkroom-denoise `denoise/repair.py`, `--repair-k 8`): the threshold
//! belongs to the model, and changes with it. Neighbours off the frame are
//! the nearest photosite on it, as the training code reads them.
/// The threshold the shipped model was trained with, in σ.
pub const REPAIR_K: f32 = 8.0;
/// Repair `mosaic` (`h×w`, row-major, normalised) in place; `sigma(y, x, v)`
/// is the photosite's σ. Returns how many photosites changed.
pub fn repair(
mosaic: &mut [f32],
h: usize,
w: usize,
k: f32,
sigma: &(dyn Fn(usize, usize, f32) -> f32 + Sync),
) -> usize {
let copy = mosaic.to_vec();
let original = &copy;
let at = |y: isize, x: isize| {
let y = y.clamp(0, h as isize - 1) as usize;
let x = x.clamp(0, w as isize - 1) as usize;
original[y * w + x]
};
let threads = std::thread::available_parallelism().map_or(1, |n| n.get());
let rows_per = h.div_ceil(threads).max(1);
let mut counts = vec![0usize; h.div_ceil(rows_per)];
std::thread::scope(|scope| {
for ((chunk, rows), count) in mosaic
.chunks_mut(rows_per * w)
.enumerate()
.zip(counts.iter_mut())
{
let at = &at;
scope.spawn(move || {
for (i, row) in rows.chunks_mut(w).enumerate() {
let y = chunk * rows_per + i;
for (x, out) in row.iter_mut().enumerate() {
let v = original[y * w + x];
let (yi, xi) = (y as isize, x as isize);
let (mut s_hi, mut s_lo) = (f32::MIN, f32::MAX);
let (mut a_hi, mut a_lo) = (f32::MIN, f32::MAX);
for dy in -1isize..=1 {
for dx in -1isize..=1 {
if dy == 0 && dx == 0 {
continue;
}
let s = at(yi + 2 * dy, xi + 2 * dx);
s_hi = s_hi.max(s);
s_lo = s_lo.min(s);
let a = at(yi + dy, xi + dx);
a_hi = a_hi.max(a);
a_lo = a_lo.min(a);
}
}
let t = k * sigma(y, x, v);
if v - s_hi > t && v - a_hi > t && a_hi < 0.5 * v {
*out = s_hi;
*count += 1;
} else if s_lo - v > t && a_lo - v > t && v < 0.5 * a_lo {
*out = s_lo;
*count += 1;
}
}
}
});
}
});
counts.iter().sum()
}
#[cfg(test)]
mod tests {
use super::*;
const N: usize = 16;
fn flat(level: f32) -> Vec<f32> {
vec![level; N * N]
}
fn run(m: &mut [f32]) -> usize {
repair(m, N, N, REPAIR_K, &|_, _, _| 0.01)
}
#[test]
fn a_hot_photosite_becomes_its_brightest_same_colour_neighbour() {
let mut m = flat(0.1);
m[8 * N + 8] = 0.5; // 40σ above everything around it
m[8 * N + 10] = 0.12; // a same-colour neighbour, a little brighter
assert_eq!(run(&mut m), 1);
assert_eq!(m[8 * N + 8], 0.12);
}
#[test]
fn a_dead_photosite_in_a_lit_area_is_repaired() {
let mut m = flat(0.5);
m[5 * N + 5] = 0.0;
assert_eq!(run(&mut m), 1);
assert_eq!(m[5 * N + 5], 0.5);
}
#[test]
fn a_point_of_real_light_is_kept() {
// Light through a lens lands on a patch: its adjacent photosites are
// lit too, so the second test refuses it.
let mut m = flat(0.1);
for dy in 0..3 {
for dx in 0..3 {
m[(7 + dy) * N + 7 + dx] = if (dy, dx) == (1, 1) { 0.9 } else { 0.6 };
}
}
let before = m.clone();
assert_eq!(run(&mut m), 0);
assert_eq!(m, before);
}
#[test]
fn noise_within_the_threshold_is_left_alone() {
let mut m: Vec<f32> = (0..N * N)
.map(|i| 0.1 + 0.005 * ((i * 7919 % 13) as f32 - 6.0) / 6.0)
.collect();
let before = m.clone();
assert_eq!(run(&mut m), 0);
assert_eq!(m, before);
}
}
+729
View File
@@ -0,0 +1,729 @@
//! TRACES: FR-DEV-3g
//! A whole frame through a network, in tiles, exactly (denoise.md §3.4, §14).
//!
//! A tile's output is exact in its centre: past a halo wider than the
//! network's receptive field (185 photosites for a single network, more for
//! the mixture), a tile's centre equals the whole frame's at the same place.
//! The frame is extended by reflection about its edge photosites, which
//! keeps every photosite's CFA colour, so edge tiles see real context too.
//!
//! **Tile sizes.** A fixed-shape network takes one square ([`Sizes::Square`],
//! 1408², of which Best keeps 896²). A network exported with any height and
//! width ([`Sizes::Any`]) takes the frame whole when it is small enough, and
//! otherwise the fewest equal tiles that are: [`plan`] picks the grid that
//! computes the fewest photosites. If the first tile of a plan fails — a
//! GPU out of memory — the limit is halved and the frame planned again.
//!
//! **Phase.** The network was trained on RGGB. A frame whose pattern starts
//! on another colour is read from one photosite up and/or left — the
//! reflection supplies that row or column — so its top-left is red, and the
//! output is read back from the same offset. Nothing is cropped.
use dr_decode::CfaPattern;
/// Photosites of context beyond a tile's kept centre, on every side, for a
/// single network; a mixture reaches further and says so through
/// [`TileNet::halo`].
pub const HALO: usize = 192;
/// The tiles a network takes.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum Sizes {
/// One square, `n` photosites a side.
Square(usize),
/// Any rectangle whose sides are multiples of `align`, at most `max`
/// (rows, columns).
Any { align: usize, max: (usize, usize) },
}
/// A network: `mosaic` and `sigma`, `rows×cols` RGGB, in; `3×rows×cols`
/// planar linear camera RGB out.
///
/// The inputs are handed over, and the output is lent to `write` rather than
/// returned: a 1408² tile is 24 MB of output and a whole frame 300 MB, and
/// copying it out of the runtime's buffer and back into the frame was a
/// measurable share of a frame's time.
pub trait TileNet {
/// The tile sizes it takes.
fn sizes(&self) -> Sizes;
/// Photosites of context it needs past a tile's kept centre: at least
/// its receptive field. [`HALO`] unless the network says otherwise.
fn halo(&self) -> usize {
HALO
}
fn run(
&mut self,
rows: usize,
cols: usize,
mosaic: Vec<f32>,
sigma: Vec<f32>,
write: &mut dyn FnMut(&[f32]),
) -> Result<(), crate::DenoiseError>;
}
/// How a frame is cut: every tile `rows × cols` in, keeping its centre
/// `core.0 × core.1` past the halo, on a `grid.0 × grid.1` grid.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub struct Plan {
pub rows: usize,
pub cols: usize,
pub core: (usize, usize),
pub grid: (usize, usize),
}
impl Plan {
/// Photosites the network computes for the frame.
pub fn work(&self) -> usize {
self.grid.0 * self.grid.1 * self.rows * self.cols
}
}
/// The tiles for an `uh × uw` frame (in the network's phase) at `sizes`, or
/// `None` when no tile fits.
///
/// Square tiles are today's grid. Any-size tiles are equal on each axis, so
/// one call shape serves the frame — TensorRT's profile tunes for one, and
/// the CUDA provider searches its algorithms once per shape — and the grid
/// is the one with the least work: one tile whenever the frame and its
/// halo fit under `max`.
pub fn plan(uh: usize, uw: usize, halo: usize, sizes: Sizes) -> Option<Plan> {
match sizes {
Sizes::Square(n) => {
if n <= 2 * halo || !(n - 2 * halo).is_multiple_of(2) {
return None;
}
let core = n - 2 * halo;
Some(Plan {
rows: n,
cols: n,
core: (core, core),
grid: (uh.div_ceil(core), uw.div_ceil(core)),
})
}
Sizes::Any { align, max } => {
// An even align keeps every tile origin on an even photosite,
// so every tile starts on red.
let align = align.max(2).next_multiple_of(2);
let axis = |extent: usize, tiles: usize, limit: usize| {
let size = (extent.div_ceil(tiles) + 2 * halo).next_multiple_of(align);
let core = size.checked_sub(2 * halo)?;
(size <= limit && core > 0 && core.is_multiple_of(2)).then_some((size, core))
};
let mut best: Option<Plan> = None;
for gy in 1..=16 {
let Some((rows, cy)) = axis(uh, gy, max.0) else {
continue;
};
for gx in 1..=16 {
let Some((cols, cx)) = axis(uw, gx, max.1) else {
continue;
};
let p = Plan {
rows,
cols,
core: (cy, cx),
grid: (uh.div_ceil(cy), uw.div_ceil(cx)),
};
if best.is_none_or(|b| p.work() < b.work()) {
best = Some(p);
}
}
}
best
}
}
}
/// Index into `0..n` by reflection about the end photosites, any distance
/// out: …2 1 [0 1 2 … n−1] n−2 n−3…, period `2(n−1)`. Parity is kept, which
/// is what keeps a CFA colour.
#[inline]
pub fn reflect(i: isize, n: usize) -> usize {
if n == 1 {
return 0;
}
let p = 2 * (n as isize - 1);
let m = i.rem_euclid(p);
(if m < n as isize { m } else { p - m }) as usize
}
/// How far up and left to start reading so the first photosite is red.
pub fn rggb_offset(p: CfaPattern) -> Option<(usize, usize)> {
match p {
CfaPattern::Rggb => Some((0, 0)),
CfaPattern::Grbg => Some((0, 1)),
CfaPattern::Gbrg => Some((1, 0)),
CfaPattern::Bggr => Some((1, 1)),
_ => None,
}
}
/// Run `net` over an `h×w` mosaic given by `at(y, x)`, with σ from
/// `sigma(colour, value)`, and return `h×w` interleaved RGB.
///
/// `progress(done, total)` is called after each tile and stops the run by
/// returning `false`, in which case the result is `Ok(None)`. An any-size
/// network whose first tile fails is planned again with tiles half that
/// size, until a tile would keep no centre; then the failure is returned.
#[allow(clippy::too_many_arguments)]
pub fn run_tiled(
net: &mut dyn TileNet,
h: usize,
w: usize,
pattern: CfaPattern,
at: &(dyn Fn(usize, usize) -> f32 + Sync),
sigma: &(dyn Fn(usize, f32) -> f32 + Sync),
progress: &mut dyn FnMut(usize, usize) -> bool,
) -> Result<Option<Vec<f32>>, crate::DenoiseError> {
let (dy, dx) = rggb_offset(pattern).ok_or_else(|| {
crate::DenoiseError::Unsupported(format!("{pattern:?} is not a Bayer pattern"))
})?;
let halo = net.halo();
let (uh, uw) = (h + dy, w + dx);
let mut sizes = net.sizes();
loop {
let plan = plan(uh, uw, halo, sizes).ok_or_else(|| {
crate::DenoiseError::Model(format!(
"no tile of {sizes:?} keeps a centre past a {halo} halo"
))
})?;
match run_plan(net, plan, h, w, (dy, dx), halo, at, sigma, progress) {
Err(Failed { error, first: true }) => {
// The first call of a size is where a GPU runs out of
// memory. Halve the larger kept centre of the tile that
// failed — not the limit, which may be far above it, and not
// the tile, half of which may be all halo — and plan again,
// until no smaller tile keeps a centre.
let Sizes::Any { align, .. } = sizes else {
return Err(error);
};
let (cr, cc) = plan.core;
let smaller = if cr >= cc {
(cr / 2 + 2 * halo, plan.cols)
} else {
(plan.rows, cc / 2 + 2 * halo)
};
let next = Sizes::Any {
align,
max: smaller,
};
if self::plan(uh, uw, halo, next).is_none() {
return Err(error);
}
log::warn!(
"learned denoise: a {}×{} tile failed ({error}); trying tiles up to {}×{}",
plan.rows,
plan.cols,
smaller.0,
smaller.1
);
sizes = next;
}
Err(Failed { error, .. }) => return Err(error),
Ok(done) => return Ok(done),
}
}
}
/// A run that stopped on an error, and whether it was the plan's first call.
struct Failed {
error: crate::DenoiseError,
first: bool,
}
#[allow(clippy::too_many_arguments)]
fn run_plan(
net: &mut dyn TileNet,
plan: Plan,
h: usize,
w: usize,
(dy, dx): (usize, usize),
halo: usize,
at: &(dyn Fn(usize, usize) -> f32 + Sync),
sigma: &(dyn Fn(usize, f32) -> f32 + Sync),
progress: &mut dyn FnMut(usize, usize) -> bool,
) -> Result<Option<Vec<f32>>, Failed> {
let Plan {
rows: nr,
cols: nc,
core: (cr, cc),
grid: (ty, tx),
} = plan;
// In unified coordinates the frame spans u ∈ [dy, dy + h), v ∈ [dx, dx + w).
let (uh, uw) = (h + dy, w + dx);
let total = ty * tx;
let origins: Vec<(usize, usize)> = (0..ty)
.flat_map(|i| (0..tx).map(move |j| (i * cr, j * cc)))
.collect();
let threads = std::thread::available_parallelism().map_or(1, |n| n.get());
// One tile's mosaic and σ, gathered on every core: rows are independent.
let gather = |u0: usize, v0: usize| {
let mut mos = vec![0.0f32; nr * nc];
let mut sig = vec![0.0f32; nr * nc];
let rows_per = nr.div_ceil(threads).max(1);
std::thread::scope(|scope| {
for (chunk, (m, s)) in mos
.chunks_mut(rows_per * nc)
.zip(sig.chunks_mut(rows_per * nc))
.enumerate()
{
scope.spawn(move || {
for (i, (mrow, srow)) in m.chunks_mut(nc).zip(s.chunks_mut(nc)).enumerate() {
let r = chunk * rows_per + i;
// Unified row u = u0 + r − halo; frame row y = u − dy, reflected.
let u = u0 as isize + r as isize - halo as isize;
let y = reflect(u - dy as isize, h);
for c in 0..nc {
let v = v0 as isize + c as isize - halo as isize;
let x = reflect(v - dx as isize, w);
let val = at(y, x);
mrow[c] = val;
// RGGB colour of the tile position (r, c).
srow[c] = sigma([[0, 1], [1, 2]][r & 1][c & 1], val);
}
}
});
}
});
(mos, sig)
};
// Pipelined: the next tile is gathered while the network runs this one,
// so the device does not wait on the CPU. A channel of one keeps at
// most two tiles' inputs alive.
let mut out = vec![0.0f32; h * w * 3];
let stop = std::sync::atomic::AtomicBool::new(false);
let fail = |error, k: usize, stop: &std::sync::atomic::AtomicBool| {
stop.store(true, std::sync::atomic::Ordering::Relaxed);
Failed {
error,
first: k == 0,
}
};
std::thread::scope(|scope| -> Result<Option<()>, Failed> {
let (tx_tiles, rx_tiles) = std::sync::mpsc::sync_channel(1);
let (origins, stop, gather) = (&origins, &stop, &gather);
scope.spawn(move || {
for &(u0, v0) in origins {
if stop.load(std::sync::atomic::Ordering::Relaxed) {
break;
}
if tx_tiles.send((u0, v0, gather(u0, v0))).is_err() {
break;
}
}
});
for k in 0..total {
let Ok((u0, v0, (mos, sig))) = rx_tiles.recv() else {
break;
};
let mut wrong = None;
let ran = net.run(nr, nc, mos, sig, &mut |rgb: &[f32]| {
if rgb.len() != 3 * nr * nc {
wrong = Some(rgb.len());
return;
}
// The tile's centre back into the frame: the frame rows it covers,
// split across cores (each row is written by one thread only).
let (y_lo, y_hi) = (u0.max(dy) - dy, (u0 + cr).min(uh) - dy);
let (x_lo, x_hi) = (v0.max(dx) - dx, (v0 + cc).min(uw) - dx);
if y_hi > y_lo && x_hi > x_lo {
let rows = &mut out[y_lo * w * 3..y_hi * w * 3];
let per = (y_hi - y_lo).div_ceil(threads).max(1);
std::thread::scope(|scope| {
for (chunk, block) in rows.chunks_mut(per * w * 3).enumerate() {
scope.spawn(move || {
for (i, row) in block.chunks_mut(w * 3).enumerate() {
let y = y_lo + chunk * per + i;
// Tile row of frame row y: u = y + dy = u0 + r − halo.
let r = y + dy + halo - u0;
for x in x_lo..x_hi {
let c = x + dx + halo - v0;
for ch in 0..3 {
row[x * 3 + ch] = rgb[ch * nr * nc + r * nc + c];
}
}
}
});
}
});
}
});
if let Err(e) = ran {
while rx_tiles.try_recv().is_ok() {}
return Err(fail(e, k, stop));
}
if let Some(len) = wrong {
while rx_tiles.try_recv().is_ok() {}
return Err(fail(
crate::DenoiseError::Model(format!(
"network returned {len} values for a {nr}×{nc} tile"
)),
k,
stop,
));
}
if !progress(k + 1, total) {
stop.store(true, std::sync::atomic::Ordering::Relaxed);
// Drain so the producer is not left blocked on a full channel.
while rx_tiles.try_recv().is_ok() {}
return Ok(None);
}
}
Ok(Some(()))
})
.map(|done| done.map(|()| out))
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn reflection_keeps_parity_any_distance_out() {
let n = 7;
for i in -40isize..40 {
let r = reflect(i, n);
assert!(r < n);
assert_eq!(
r % 2,
i.rem_euclid(2) as usize,
"index {i} reflected to {r}"
);
}
assert_eq!(reflect(-1, n), 1);
assert_eq!(reflect(7, n), 5);
}
/// A stand-in network with a known, finite reach: each output photosite
/// is its 2×2 quad's (R, mean G, B), averaged over the quads within
/// `reach` quads. Purely a function of the tile, like the real one.
struct BoxNet {
sizes: Sizes,
reach: usize,
/// Fails any call with more photosites than this, as a GPU out of
/// memory does.
fails_above: usize,
calls: Vec<(usize, usize)>,
}
fn square(n: usize, reach: usize) -> BoxNet {
BoxNet {
sizes: Sizes::Square(n),
reach,
fails_above: usize::MAX,
calls: Vec::new(),
}
}
fn any(max: (usize, usize), reach: usize) -> BoxNet {
BoxNet {
sizes: Sizes::Any { align: 16, max },
reach,
fails_above: usize::MAX,
calls: Vec::new(),
}
}
impl TileNet for BoxNet {
fn sizes(&self) -> Sizes {
self.sizes
}
fn run(
&mut self,
rows: usize,
cols: usize,
m: Vec<f32>,
_s: Vec<f32>,
write: &mut dyn FnMut(&[f32]),
) -> Result<(), crate::DenoiseError> {
self.calls.push((rows, cols));
if rows * cols > self.fails_above {
return Err(crate::DenoiseError::Model("out of memory".into()));
}
let (qr, qc) = (rows / 2, cols / 2);
let quad = |qy: usize, qx: usize| {
let (y, x) = (2 * qy, 2 * qx);
[
m[y * cols + x],
0.5 * (m[y * cols + x + 1] + m[(y + 1) * cols + x]),
m[(y + 1) * cols + x + 1],
]
};
let plane = rows * cols;
let mut out = vec![0.0; 3 * plane];
for qy in 0..qr {
for qx in 0..qc {
let mut acc = [0.0f32; 3];
let mut cnt = 0.0;
for a in qy.saturating_sub(self.reach)..(qy + self.reach + 1).min(qr) {
for b in qx.saturating_sub(self.reach)..(qx + self.reach + 1).min(qc) {
let v = quad(a, b);
for c in 0..3 {
acc[c] += v[c];
}
cnt += 1.0;
}
}
for (dy, dx) in [(0, 0), (0, 1), (1, 0), (1, 1)] {
for c in 0..3 {
out[c * plane + (2 * qy + dy) * cols + 2 * qx + dx] = acc[c] / cnt;
}
}
}
}
write(&out);
Ok(())
}
}
/// The mosaic of a smooth colour field in `pattern`, read at (y, x).
fn field(pattern: CfaPattern) -> impl Fn(usize, usize) -> f32 {
move |y, x| {
let rgb = [0.2 + 0.0004 * x as f32, 0.5, 0.1 + 0.0003 * y as f32];
rgb[pattern.colour_at(x as u32, y as u32) as usize]
}
}
#[test]
fn every_bayer_phase_comes_back_as_its_own_colours() {
// A frame of each pattern, its colours known: the network must see
// red where the frame's red photosites are, whatever the phase.
for p in [
CfaPattern::Rggb,
CfaPattern::Grbg,
CfaPattern::Gbrg,
CfaPattern::Bggr,
] {
let (h, w) = (300, 410);
let at = field(p);
let mut net = square(2 * HALO + 64, 0);
let out = run_tiled(&mut net, h, w, p, &at, &|_, _| 0.01, &mut |_, _| true)
.unwrap()
.unwrap();
for (y, x) in [(10, 10), (150, 201), (299, 409), (0, 0), (77, 333)] {
let o = &out[(y * w + x) * 3..(y * w + x) * 3 + 3];
let want = [0.2 + 0.0004 * x as f32, 0.5, 0.1 + 0.0003 * y as f32];
for c in 0..3 {
// Within the quad the binned value is at most a photosite away.
assert!(
(o[c] - want[c]).abs() < 0.0012,
"{p:?} at ({y},{x}) channel {c}: {} vs {}",
o[c],
want[c]
);
}
}
}
}
#[test]
fn tiles_reproduce_one_pass_over_the_reflected_frame() {
// A network whose reach is inside the halo gives the same answer
// tiled small as in one tile covering everything.
let (h, w) = (230, 170);
for p in [CfaPattern::Rggb, CfaPattern::Bggr] {
let at = |y: usize, x: usize| ((y * 7919 + x * 104729) % 1000) as f32 / 1000.0;
let mut small = square(2 * HALO + 32, 20);
let mut big = square(2 * HALO + 256, 20);
let a = run_tiled(&mut small, h, w, p, &at, &|_, _| 0.0, &mut |_, _| true)
.unwrap()
.unwrap();
let b = run_tiled(&mut big, h, w, p, &at, &|_, _| 0.0, &mut |_, _| true)
.unwrap()
.unwrap();
let worst = a
.iter()
.zip(&b)
.map(|(x, y)| (x - y).abs())
.fold(0.0f32, f32::max);
assert!(worst < 1e-5, "{p:?}: tiled and whole differ by {worst}");
}
}
/// The same frame through square tiles, one whole-frame call, a grid of
/// any-size tiles, and a network that runs out of memory on the whole
/// frame and is planned again: one answer.
#[test]
fn any_size_tiles_give_the_square_tiles_answer() {
let (h, w) = (230, 170);
for p in [
CfaPattern::Rggb,
CfaPattern::Grbg,
CfaPattern::Gbrg,
CfaPattern::Bggr,
] {
let at = |y: usize, x: usize| ((y * 7919 + x * 104729) % 1000) as f32 / 1000.0;
let run = |net: &mut BoxNet| {
run_tiled(net, h, w, p, &at, &|_, _| 0.0, &mut |_, _| true)
.unwrap()
.unwrap()
};
// A reach of 6 quads is well inside the halo, and keeps a
// debug-build test of four phases short.
let want = run(&mut square(2 * HALO + 32, 6));
let mut whole = any((4096, 4096), 6);
let got = run(&mut whole);
assert_eq!(whole.calls.len(), 1, "the frame fits: one call");
assert_eq!(got, want, "{p:?}: whole frame");
let mut grid = any((2 * HALO + 96, 2 * HALO + 64), 6);
let got = run(&mut grid);
assert!(grid.calls.len() > 1);
assert!(
grid.calls.windows(2).all(|c| c[0] == c[1]),
"one call shape"
);
assert_eq!(got, want, "{p:?}: a grid of any-size tiles");
let mut tight = any((4096, 4096), 6);
tight.fails_above = (2 * HALO + 200) * (2 * HALO + 200);
let got = run(&mut tight);
assert_eq!(got, want, "{p:?}: planned again after a failure");
assert!(tight.calls.len() > 2, "the whole frame failed, then tiles");
}
}
#[test]
fn the_plan_is_one_tile_when_the_frame_fits_and_the_least_work_when_not() {
// A 6D frame with Best's halo, under the whole-frame limit: one call.
let one = plan(
3648,
5472,
256,
Sizes::Any {
align: 16,
max: (4608, 6656),
},
)
.unwrap();
assert_eq!(one.grid, (1, 1));
assert_eq!((one.rows, one.cols), (4160, 5984));
assert!(one.core.0 >= 3648 && one.core.1 >= 5472);
// Too wide for one: the cheapest grid, every tile within the limit.
let two = plan(
3648,
8192,
256,
Sizes::Any {
align: 16,
max: (4608, 6656),
},
)
.unwrap();
assert!(two.cols <= 6656 && two.rows <= 4608);
assert_eq!(two.grid, (1, 2));
// And always less work than today's 1408 squares.
let squares = plan(3648, 5472, 256, Sizes::Square(1408)).unwrap();
assert_eq!(squares.grid, (5, 7));
assert!(one.work() * 2 < squares.work());
// The whole-frame engine's limit on a 6 GB card: two tiles, each
// within it, and still under half the work of the 1408 squares.
let halves = plan(
3648,
5472,
256,
Sizes::Any {
align: 16,
max: (4608, 3328),
},
)
.unwrap();
assert_eq!(halves.grid, (1, 2));
assert_eq!((halves.rows, halves.cols), (4160, 3248));
assert!(halves.work() * 2 < squares.work());
// A limit no tile fits under.
assert!(plan(
3648,
5472,
256,
Sizes::Any {
align: 16,
max: (400, 400)
}
)
.is_none());
}
#[test]
fn a_cancelled_run_returns_nothing() {
let mut net = square(2 * HALO + 32, 0);
let r = run_tiled(
&mut net,
100,
100,
CfaPattern::Rggb,
&|_, _| 0.5,
&|_, _| 0.0,
&mut |done, _| done < 2,
)
.unwrap();
assert!(r.is_none());
}
}
#[cfg(test)]
mod timing {
use super::*;
/// A network that answers instantly with an output of the right size,
/// so what is timed is the tiler alone: gathering each tile's mosaic and
/// σ, and writing its centre back.
struct Null(usize, Vec<f32>);
impl TileNet for Null {
fn sizes(&self) -> Sizes {
Sizes::Square(self.0)
}
fn run(
&mut self,
_rows: usize,
_cols: usize,
m: Vec<f32>,
_s: Vec<f32>,
write: &mut dyn FnMut(&[f32]),
) -> Result<(), crate::DenoiseError> {
// Stands for the runtime's own output buffer: allocated once.
if self.1.len() != 3 * m.len() {
self.1 = vec![m[0]; 3 * m.len()];
}
write(&self.1);
Ok(())
}
}
/// `cargo test --release -p dr-denoise tiler_overhead -- --ignored --nocapture`
#[test]
#[ignore]
fn tiler_overhead_on_a_6d_frame() {
let (h, w) = (3648, 5472);
let frame: Vec<f32> = (0..h * w).map(|i| (i % 977) as f32 / 977.0).collect();
let at = |y: usize, x: usize| frame[y * w + x];
let sigma = |_c: usize, v: f32| (0.001 * v + 1e-5).sqrt();
for n in [1408usize, 2048] {
let mut net = Null(n, Vec::new());
let t = std::time::Instant::now();
let mut tiles = 0;
run_tiled(
&mut net,
h,
w,
CfaPattern::Rggb,
&at,
&sigma,
&mut |_, total| {
tiles = total;
true
},
)
.unwrap();
let s = t.elapsed().as_secs_f64();
println!(
"tile {n}: {tiles} tiles, tiler alone {s:.2} s ({:.0} ms a tile)",
s / tiles as f64 * 1e3
);
}
}
}
+36
View File
@@ -0,0 +1,36 @@
# Canon EOS 6D noise, measured from the library's own frames (denoise.md §5).
# Shot gain S and read variance O per RGGB position from Adobe's NoiseProfile in
# converted DNGs, in DN at the ISO's own white level; read noise checked against
# the masked border (within 2-3 %); row and column noise from the masked border.
# ISO 50 and 100 are extrapolated (S proportional to ISO). Generated by
# darkroom-denoise tools/profile.py; regenerate there, never edit by hand.
make: Canon
model: EOS 6D
black: 2048
rows:
- {iso: 50, white: 15000, s_dn: [0.0854021, 0.085467, 0.085467, 0.0839724], o_dn: [38.2741, 38.6675, 38.6675, 38.9649], row_dn: 0.3423, col_dn: 0.505}
- {iso: 100, white: 15000, s_dn: [0.170804, 0.170934, 0.170934, 0.167945], o_dn: [38.339, 38.7332, 38.7332, 39.031], row_dn: 0.3423, col_dn: 0.505}
- {iso: 125, white: 15035, s_dn: [0.228108, 0.230361, 0.230361, 0.228345], o_dn: [36.9455, 37.8995, 37.8995, 38.2358], row_dn: 0.3423, col_dn: 0.505}
- {iso: 160, white: 12373, s_dn: [0.289653, 0.294915, 0.294915, 0.286887], o_dn: [15.3717, 16.1346, 16.1346, 16.0413], row_dn: 0.212, col_dn: 0.07151}
- {iso: 200, white: 15035, s_dn: [0.370969, 0.369922, 0.369922, 0.361443], o_dn: [24.2761, 24.0847, 24.0847, 24.2414], row_dn: 0.2692, col_dn: 0}
- {iso: 250, white: 15035, s_dn: [0.461889, 0.457975, 0.457975, 0.449318], o_dn: [38.0975, 37.5041, 37.5041, 37.7424], row_dn: 0.3345, col_dn: 0.4786}
- {iso: 320, white: 12323, s_dn: [0.590765, 0.59755, 0.59755, 0.576843], o_dn: [18.5426, 18.9163, 18.9163, 19.1202], row_dn: 0.3158, col_dn: 0.5174}
- {iso: 400, white: 15035, s_dn: [0.753591, 0.740586, 0.740586, 0.729874], o_dn: [29.3028, 29.8496, 29.8496, 29.7961], row_dn: 0.4378, col_dn: 0.2691}
- {iso: 500, white: 15035, s_dn: [0.937458, 0.920836, 0.920836, 0.899293], o_dn: [45.2196, 46.3954, 46.3954, 46.1012], row_dn: 0.5473, col_dn: 0.4328}
- {iso: 640, white: 12323, s_dn: [1.12726, 1.13527, 1.13527, 1.10159], o_dn: [24.9951, 25.1029, 25.1029, 25.7183], row_dn: 0.316, col_dn: 0.4544}
- {iso: 800, white: 15035, s_dn: [1.44048, 1.42299, 1.42299, 1.40795], o_dn: [38.7891, 39.302, 39.302, 40.079], row_dn: 0.3877, col_dn: 0.2132}
- {iso: 1000, white: 15000, s_dn: [1.77595, 1.75662, 1.75662, 1.74584], o_dn: [63.9499, 64.4203, 64.4203, 65.0739], row_dn: 0.4593, col_dn: 0.3307}
- {iso: 1250, white: 12346, s_dn: [2.18211, 2.18313, 2.18313, 2.11979], o_dn: [41.6124, 42.9483, 42.9483, 43.2075], row_dn: 0.3979, col_dn: 0.4496}
- {iso: 1600, white: 15035, s_dn: [2.75544, 2.74633, 2.74633, 2.69951], o_dn: [66.3905, 66.4104, 66.4104, 67.253], row_dn: 0.4944, col_dn: 0.4593}
- {iso: 2000, white: 15035, s_dn: [3.42754, 3.40445, 3.40445, 3.36808], o_dn: [104.349, 103.648, 103.648, 106.404], row_dn: 0.6094, col_dn: 0.3602}
- {iso: 2500, white: 12330, s_dn: [4.17112, 4.17551, 4.17551, 4.17175], o_dn: [94.4289, 91.8508, 91.8508, 96.3598], row_dn: 0.5671, col_dn: 0}
- {iso: 3200, white: 15035, s_dn: [5.30088, 5.25742, 5.25742, 5.21782], o_dn: [147.421, 147.302, 147.302, 147.01], row_dn: 0.748, col_dn: 0.8611}
- {iso: 4000, white: 15035, s_dn: [6.62037, 6.59922, 6.59922, 6.60871], o_dn: [224.765, 232.408, 232.408, 231.419], row_dn: 0.9335, col_dn: 1.125}
- {iso: 5000, white: 12323, s_dn: [8.49542, 8.48265, 8.48265, 8.41176], o_dn: [232.672, 233.922, 233.922, 256.059], row_dn: 1.085, col_dn: 1.852}
- {iso: 6400, white: 15035, s_dn: [10.6956, 10.7417, 10.7417, 10.6503], o_dn: [360.311, 368.198, 368.198, 362.848], row_dn: 1.326, col_dn: 2.277}
- {iso: 8000, white: 15035, s_dn: [13.1307, 13.3864, 13.3864, 13.147], o_dn: [615.02, 566.666, 566.666, 611.738], row_dn: 1.768, col_dn: 3.141}
- {iso: 10000, white: 12365, s_dn: [16.5338, 16.7603, 16.7603, 16.4739], o_dn: [914.064, 904.583, 904.583, 938.024], row_dn: 2.214, col_dn: 3.605}
- {iso: 12800, white: 15000, s_dn: [18.4717, 20.9315, 20.9315, 19.3821], o_dn: [1431.85, 1432.33, 1432.33, 1477.82], row_dn: 2.568, col_dn: 4.661}
- {iso: 16000, white: 15000, s_dn: [20.527, 26.0841, 26.0841, 21.8866], o_dn: [2203.77, 2357.78, 2357.78, 2193.1], row_dn: 3.521, col_dn: 5.805}
- {iso: 20000, white: 13000, s_dn: [25.1517, 32.5307, 32.5307, 26.2303], o_dn: [3490.34, 3647.93, 3647.93, 3423.33], row_dn: 4.336, col_dn: 7.143}
- {iso: 25600, white: 15000, s_dn: [22.8743, 40.4641, 40.4641, 23.5537], o_dn: [5184.57, 5690.43, 5690.43, 5286.07], row_dn: 5.682, col_dn: 9.193}
+3 -3
View File
@@ -137,8 +137,8 @@ impl Detection {
/// A loaded SCRFD graph.
pub struct Detector {
session: Model,
/// f32 or int8 — the int8 form finds a different set of faces and is a
/// different detector in `model_id` (docs/dev/inference.md §7).
/// f32 or a quantised form — which finds a different set of faces and is
/// a different detector in `model_id` (docs/dev/inference.md §7).
form: Form,
/// Feature-map count: 3 for strides {8,16,32}, 4 for {8,16,32,64}.
///
@@ -155,7 +155,7 @@ impl Detector {
}
/// Load the canonical f32 file at `path`, or the form the device's
/// backend wants instead — the `.int8.onnx` beside it on a Hexagon —
/// backend wants instead — the `.a16w8.onnx` beside it on a Hexagon —
/// which [`Detector::form`] then reports.
pub fn from_path(path: impl AsRef<std::path::Path>) -> Result<Self, FaceError> {
let (path, form) = dr_inference_engine::resolve_model(Role::Detector, path.as_ref());
+11 -2
View File
@@ -123,13 +123,22 @@ pub struct Landmarker {
}
impl Landmarker {
/// The graph at `path`, or the `.a16w8.onnx` sibling beside it when the
/// device's backend runs that (the Hexagon, inference.md §1.5: 0.25 px
/// from f32 in the 192 crop, where int8 moved the points by 1.5).
pub fn from_path(path: impl AsRef<std::path::Path>) -> Result<Self, FaceError> {
let (path, form) = dr_inference_engine::resolve_model(Role::Landmarks, path.as_ref());
let bytes = std::fs::read(path).map_err(FaceError::ModelRead)?;
Self::from_bytes(&bytes)
Self::from_bytes_in(&bytes, form)
}
pub fn from_bytes(bytes: &[u8]) -> Result<Self, FaceError> {
let model = dr_inference_engine::open(Role::Landmarks, Form::F32, bytes)?;
Self::from_bytes_in(bytes, Form::F32)
}
/// `bytes` in a stated numeric form; the output keeps its meaning.
pub fn from_bytes_in(bytes: &[u8], form: Form) -> Result<Self, FaceError> {
let model = dr_inference_engine::open(Role::Landmarks, form, bytes)?;
let acquired = model.acquire()?;
let session = acquired.lock();
+39 -3
View File
@@ -24,7 +24,7 @@ fn main() {
let mut args = std::env::args().skip(1);
let Some(input) = args.next() else {
eprintln!("usage: develop <file.cr2> [out.ppm] [preset]");
eprintln!(" preset: neutral (default) | punchy | recover");
eprintln!(" preset: neutral (default) | matrix | look200 | punchy | recover | …");
std::process::exit(2);
};
let output = args.next().unwrap_or_else(|| "develop.ppm".into());
@@ -78,6 +78,24 @@ fn main() {
graph.set_param(brilliance::ID, brilliance::BRILLIANCE, 40.0);
graph.set_param(white_balance::ID, white_balance::TEMPERATURE, 15.0);
}
// The camera profile switched off: the matrix alone, as every
// photograph rendered before D20. Beside "neutral" on a DNG that
// embeds a profile, the difference is the profile's tables.
"matrix" => {
graph.set_param(
dr_pipeline::ops::camera_profile::ID,
dr_pipeline::ops::camera_profile::APPLY,
0.0,
);
}
// The profile's look table at twice its strength.
"look200" => {
graph.set_param(
dr_pipeline::ops::camera_profile::ID,
dr_pipeline::ops::camera_profile::LOOK,
200.0,
);
}
// Contrast alone, so its effect can be judged without anything else
// moving.
"contrast" => {
@@ -109,6 +127,14 @@ fn main() {
graph.set_param(curve::ID, curve::P0_Y, 0.12);
graph.set_param(curve::ID, curve::P1_Y, 0.32);
}
// A shipped preset by name — `preset:Vivid landscape` — applied as
// the presets menu applies it, so a look can be judged on a real file.
named if named.starts_with("preset:") => {
let name = &named["preset:".len()..];
let preset = dr_pipeline::bundled::lookup(&Default::default(), name)
.unwrap_or_else(|| panic!("no shipped preset called {name:?}"));
let _ = preset.apply(&mut graph, dr_pipeline::Scope::adjustments());
}
_ => {}
}
@@ -123,8 +149,15 @@ fn main() {
let mut adjust = AdjustPass::new(&ctx);
let (w, h) = image.size();
// Through the detail stage when the edit has one — clarity, sharpening
// — which is the path every frontend takes; `render` alone refuses such
// a shader.
let t2 = std::time::Instant::now();
adjust.render(&image, &shader, w, h).expect("adjust");
let detail = graph.compose_detail(image.size(), (w, h));
let key = graph.invalidation().through(dr_pipeline::Affects::Colour);
adjust
.render_detailed(&image, &shader, w, h, None, &detail, key)
.expect("adjust");
ctx.device
.poll(wgpu::PollType::wait_indefinitely())
.expect("poll");
@@ -134,8 +167,11 @@ fn main() {
// path, and it must not recompile.
graph.set_param(exposure::ID, exposure::EXPOSURE, 0.31);
let again = graph.compose();
let key = graph.invalidation().through(dr_pipeline::Affects::Colour);
let t3 = std::time::Instant::now();
adjust.render(&image, &again, w, h).expect("adjust");
adjust
.render_detailed(&image, &again, w, h, None, &detail, key)
.expect("adjust");
ctx.device
.poll(wgpu::PollType::wait_indefinitely())
.expect("poll");
+115
View File
@@ -0,0 +1,115 @@
//! Dump RAW files' mosaics for training the learned denoise (FR-DEV-3g).
//!
//! The training repo must read photosites the way the app reads them —
//! same black and white levels, same active area, same CFA phase — or a
//! network trained on one phase runs on another and paints moiré everywhere
//! (denoise.md §4.4). So it reads this, not LibRaw.
//!
//! The photosites are those the demosaic reads: hot and dead ones repaired by
//! the app's own pass ([`Demosaicer::repair_hot_pixels`], the same shader
//! `run` dispatches), because the learned stage replaces the demosaic and
//! takes its input (denoise.md §2). `--unrepaired` skips it.
//!
//! Reads `input<TAB>output-prefix` lines on stdin and writes, per line,
//! `prefix.npy` (the whole readout, masked border included, `u16`, row-major)
//! and `prefix.json` (what `decode` and `metadata` say about it). The border
//! is kept, and the repair never touches it, because its optically black
//! photosites are a dark frame for free: read noise and row noise at that ISO.
//!
//! ```sh
//! printf 'IMG_0001.CR2\tout/IMG_0001\n' |
//! cargo run --release -p dr-gpu --example mosaic_dump
//! ```
use std::io::{BufRead, Write};
use dr_gpu::{Demosaicer, GpuContext};
fn main() {
let repair = !std::env::args().any(|a| a == "--unrepaired");
let ctx = pollster::block_on(GpuContext::new_headless()).expect("a GPU for the hot-pixel pass");
let demosaicer = Demosaicer::new(&ctx).expect("demosaicer");
let mut failed = 0;
for line in std::io::stdin().lock().lines() {
let line = line.expect("stdin");
let Some((input, prefix)) = line.split_once('\t') else {
continue;
};
match dump(input, prefix, repair.then_some(&demosaicer)) {
Ok(()) => println!("ok\t{input}"),
Err(e) => {
failed += 1;
println!("fail\t{input}\t{e}");
}
}
std::io::stdout().flush().ok();
}
std::process::exit(if failed > 0 { 1 } else { 0 });
}
fn dump(input: &str, prefix: &str, repair: Option<&Demosaicer>) -> Result<(), String> {
let bytes = std::fs::read(input).map_err(|e| e.to_string())?;
let mut raw = dr_decode::decode(&bytes).map_err(|e| e.to_string())?;
if raw.samples_per_pixel != 1 {
return Err("linear DNG: no photosites".into());
}
let repaired = match repair {
Some(d) => d.repair_hot_pixels(&mut raw).map_err(|e| e.to_string())? as i64,
None => -1,
};
let meta = dr_decode::metadata(&bytes).map_err(|e| e.to_string())?;
let mut npy = Vec::with_capacity(raw.data.len() * 2 + 128);
let mut header = format!(
"{{'descr': '<u2', 'fortran_order': False, 'shape': ({}, {}), }}",
raw.height, raw.width
);
// The header, its magic and length are padded to a multiple of 64.
while (10 + header.len() + 1) % 64 != 0 {
header.push(' ');
}
header.push('\n');
npy.extend_from_slice(b"\x93NUMPY\x01\x00");
npy.extend_from_slice(&(header.len() as u16).to_le_bytes());
npy.extend_from_slice(header.as_bytes());
for v in &raw.data {
npy.extend_from_slice(&v.to_le_bytes());
}
std::fs::write(format!("{prefix}.npy"), npy).map_err(|e| e.to_string())?;
let opt = |v: Option<f32>| v.map_or("null".to_string(), |v| v.to_string());
let matrix = raw
.color_matrix
.map_or("null".to_string(), |m| format!("{m:?}"));
let json = format!(
concat!(
"{{\"source\": {:?}, \"make\": {:?}, \"model\": {:?}, ",
"\"width\": {}, \"height\": {}, ",
"\"crop\": [{}, {}, {}, {}], \"cfa\": {:?}, ",
"\"black\": {:?}, \"white\": {}, \"wb\": {:?}, \"cam_to_srgb\": {}, ",
"\"iso\": {}, \"shutter\": {}, \"aperture\": {}, \"captured_at\": {}, ",
"\"hot_repaired\": {}}}\n"
),
input,
raw.make,
raw.model,
raw.width,
raw.height,
raw.crop.x,
raw.crop.y,
raw.crop.width,
raw.crop.height,
format!("{:?}", raw.cfa_pattern),
raw.black_level,
raw.white_level,
raw.wb_coeffs,
matrix,
meta.iso.map_or("null".to_string(), |v| v.to_string()),
opt(meta.shutter),
opt(meta.aperture),
meta.captured_at
.map_or("null".to_string(), |v| v.to_string()),
repaired,
);
std::fs::write(format!("{prefix}.json"), json).map_err(|e| e.to_string())
}
+117
View File
@@ -0,0 +1,117 @@
//! List each RAW file's hot and dead photosite candidates (docs/dev/sensor-health.md).
//!
//! Reads paths on stdin and prints one JSON line per file: its capture
//! conditions and every photosite [`Demosaicer::find_hot_pixels`] flags, as
//! `[x, y, value, hot]` in sensor coordinates. Which candidates are defects
//! is a question across frames, so this answers nothing on its own.
//!
//! ```sh
//! find ~/Pictures -name '*.CR2' | cargo run --release -p dr-gpu --example sensor_scan
//! ```
use std::io::{BufRead, Write};
use dr_gpu::{Demosaicer, GpuContext};
fn main() {
if let Some(list) = std::env::args().skip_while(|a| a != "--probe").nth(1) {
return probe(&list);
}
let ctx = pollster::block_on(GpuContext::new_headless()).expect("a GPU for the hot-pixel pass");
let demosaicer = Demosaicer::new(&ctx).expect("demosaicer");
for line in std::io::stdin().lock().lines() {
let path = line.expect("stdin");
match scan(&path, &demosaicer) {
Ok(json) => println!("{json}"),
Err(e) => eprintln!("fail\t{path}\t{e}"),
}
std::io::stdout().flush().ok();
}
}
fn scan(path: &str, demosaicer: &Demosaicer) -> Result<String, String> {
let bytes = std::fs::read(path).map_err(|e| e.to_string())?;
let raw = dr_decode::decode(&bytes).map_err(|e| e.to_string())?;
let meta = dr_decode::metadata(&bytes).map_err(|e| e.to_string())?;
let sites = demosaicer
.find_hot_pixels(&raw)
.map_err(|e| e.to_string())?;
let opt = |v: Option<f64>| v.map_or("null".to_string(), |v| v.to_string());
let list: Vec<String> = sites
.iter()
.map(|s| {
let v = raw.data[(s.y * raw.width + s.x) as usize];
format!("[{},{},{},{}]", s.x, s.y, v, u8::from(s.hot))
})
.collect();
Ok(format!(
"{{\"path\":{:?},\"model\":{:?},\"captured\":{},\"iso\":{},\"shutter\":{},\"white\":{},\"black\":{:?},\"crop\":[{},{},{},{}],\"sites\":[{}]}}",
path,
format!("{} {}", raw.make, raw.model),
meta.captured_at.map_or("null".to_string(), |t| t.to_string()),
opt(meta.iso.map(f64::from)),
opt(meta.shutter.map(f64::from)),
raw.white_level,
raw.black_level,
raw.crop.x,
raw.crop.y,
raw.crop.width,
raw.crop.height,
list.join(","),
))
}
/// `--probe COORDS`: for each path on stdin, each `x y` line of COORDS as
/// `[value, same-colour neighbour max, median]` over black, on the CPU. A
/// probe asks whether a photosite stood out in a frame where it would have
/// been visible, which the scan's verdict cannot say: a frame that does not
/// flag a defect may only have been too bright around it.
fn probe(list: &str) {
let coords: Vec<(u32, u32)> = std::fs::read_to_string(list)
.expect("coords")
.lines()
.filter_map(|l| {
let mut it = l.split_whitespace().map(|v| v.parse().ok());
Some((it.next()??, it.next()??))
})
.collect();
for line in std::io::stdin().lock().lines() {
let path = line.expect("stdin");
let Ok(bytes) = std::fs::read(&path) else {
continue;
};
let (Ok(raw), Ok(meta)) = (dr_decode::decode(&bytes), dr_decode::metadata(&bytes)) else {
continue;
};
let w = raw.width as i64;
let at = |x: i64, y: i64| {
let cell = (((y - raw.crop.y as i64) & 1) * 2 + ((x - raw.crop.x as i64) & 1)) as usize;
raw.data[(y * w + x) as usize].saturating_sub(raw.black_level[cell])
};
let rows: Vec<String> = coords
.iter()
.map(|&(x, y)| {
let (x, y) = (x as i64, y as i64);
let mut n: Vec<u16> = Vec::new();
for dy in [-2i64, 0, 2] {
for dx in [-2i64, 0, 2] {
if (dx, dy) != (0, 0) {
n.push(at(x + dx, y + dy));
}
}
}
n.sort_unstable();
format!("[{},{},{}]", at(x, y), n[n.len() - 1], n[n.len() / 2])
})
.collect();
println!(
"{{\"path\":{:?},\"captured\":{},\"iso\":{},\"shutter\":{},\"range\":{},\"p\":[{}]}}",
path,
meta.captured_at.unwrap_or(0),
meta.iso.unwrap_or(0),
meta.shutter.unwrap_or(0.0),
raw.white_level - raw.black_level[0],
rows.join(","),
);
}
}
+86 -1
View File
@@ -71,6 +71,14 @@ pub struct AdjustPass {
empty_film_lut: wgpu::TextureView,
/// The loaded stock's tables, once uploaded. See [`Self::set_film`].
film: Option<FilmTextures>,
/// TRACES: FR-DEV-3e
/// Bound at `@binding(8)` for a source with no camera profile tables: the
/// two-entry header of zeros that tells the fragment there is nothing to
/// apply (D20).
empty_profile: wgpu::Buffer,
/// The current source's tables, uploaded, keyed by
/// [`DemosaicedImage::id`] — one upload per source rather than per frame.
profile: Option<(u64, wgpu::Buffer)>,
/// TRACES: FR-DEV-3 | FR-DEV-3d
/// The neighbourhood stage — sharpening, noise reduction, clarity and the
/// rest of FR-DEV-3's detail set, which cannot be fused into the shader
@@ -512,6 +520,37 @@ impl AdjustPass {
}
/// The curve texture to bind: the loaded stock's, or the placeholder.
/// TRACES: FR-DEV-3e
/// A source's camera profile tables as the storage buffer
/// `@binding(8)` reads, laid out by `dr_pipeline`'s `profile_buffer`.
fn upload_profile(ctx: &GpuContext, tables: Option<&dr_types::ProfileTables>) -> wgpu::Buffer {
let data = dr_pipeline::ops::camera_profile::profile_buffer(tables);
ctx.device
.create_buffer_init(&wgpu::util::BufferInitDescriptor {
label: Some("adjust-profile-tables"),
contents: bytemuck::cast_slice(&data),
usage: wgpu::BufferUsages::STORAGE,
})
}
/// TRACES: FR-DEV-3e
/// The buffer to bind for `source`: its tables, uploaded once per source,
/// or the empty header. A cheap handle, cloned out so a caller holding
/// other borrows of `self` can bind it.
fn profile_buffer(&mut self, source: &DemosaicedImage) -> wgpu::Buffer {
let Some(tables) = source.profile_tables() else {
return self.empty_profile.clone();
};
if let Some((id, buffer)) = &self.profile {
if *id == source.id() {
return buffer.clone();
}
}
let buffer = Self::upload_profile(&self.ctx, Some(tables));
self.profile = Some((source.id(), buffer.clone()));
buffer
}
fn film_curves_view(&self) -> &wgpu::TextureView {
self.film
.as_ref()
@@ -645,6 +684,8 @@ impl AdjustPass {
empty_film_curves,
empty_film_lut,
film: None,
empty_profile: Self::upload_profile(ctx, None),
profile: None,
detail: DetailRunner::new(ctx),
linear_bind_group_layout,
linear_pipeline_layout,
@@ -775,6 +816,17 @@ impl AdjustPass {
},
count: None,
},
// The camera profile's tables (FR-DEV-3e, D20).
wgpu::BindGroupLayoutEntry {
binding: 8,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::Buffer {
ty: wgpu::BufferBindingType::Storage { read_only: true },
has_dynamic_offset: false,
min_binding_size: None,
},
count: None,
},
],
})
}
@@ -973,6 +1025,7 @@ impl AdjustPass {
usage: wgpu::BufferUsages::UNIFORM,
});
let profile = self.profile_buffer(source);
let pipeline = self
.cache
.get(&shader.structure_hash)
@@ -1020,6 +1073,10 @@ impl AdjustPass {
binding: 7,
resource: wgpu::BindingResource::TextureView(&sample_out),
},
wgpu::BindGroupEntry {
binding: 8,
resource: profile.as_entire_binding(),
},
],
});
@@ -1139,6 +1196,7 @@ impl AdjustPass {
// be read off `self` at the point the bind group is built.
let film_curves = self.film_curves_view().clone();
let film_lut = self.film_lut_view().clone();
let profile = self.profile_buffer(source);
let mut enc = self
.ctx
@@ -1202,6 +1260,10 @@ impl AdjustPass {
binding: 7,
resource: wgpu::BindingResource::TextureView(&sample_out),
},
wgpu::BindGroupEntry {
binding: 8,
resource: profile.as_entire_binding(),
},
],
});
let pipeline = self
@@ -1298,6 +1360,10 @@ impl AdjustPass {
binding: 7,
resource: wgpu::BindingResource::TextureView(&no_sample_out),
},
wgpu::BindGroupEntry {
binding: 8,
resource: profile.as_entire_binding(),
},
],
});
{
@@ -1609,6 +1675,10 @@ impl AdjustPass {
binding: 7,
resource: wgpu::BindingResource::TextureView(&self.sample.no_sample_out),
},
wgpu::BindGroupEntry {
binding: 8,
resource: self.empty_profile.as_entire_binding(),
},
],
});
@@ -1818,6 +1888,8 @@ mod tests {
color_matrix: Some([1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0]),
samples_per_pixel: 1,
profile: None,
profile_tables: None,
baseline_exposure: 0.0,
make: String::new(),
model: String::new(),
crop: CropRect {
@@ -2021,6 +2093,8 @@ mod tests {
color_matrix: Some([1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0]),
samples_per_pixel: 1,
profile: None,
profile_tables: None,
baseline_exposure: 0.0,
make: String::new(),
model: String::new(),
crop: CropRect {
@@ -2629,6 +2703,8 @@ mod tests {
color_matrix: Some([1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0]),
samples_per_pixel: 1,
profile: None,
profile_tables: None,
baseline_exposure: 0.0,
make: String::new(),
model: String::new(),
crop: CropRect {
@@ -2732,6 +2808,8 @@ mod tests {
color_matrix: Some([1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0]),
samples_per_pixel: 1,
profile: None,
profile_tables: None,
baseline_exposure: 0.0,
make: String::new(),
model: String::new(),
crop: CropRect {
@@ -3342,8 +3420,15 @@ mod tests {
read_centre(&ctx, t)
};
// The default rendering is the DNG reference curve (D21); for a grey its
// ProPhoto round trip is the identity, so the reference applies as is.
let scene = 3537.0 / 16383.0;
let viewed = dr_pipeline::view::Sigmoid::default_curve().channel(scene);
let viewed = dr_pipeline::camera_raw::apply_reference(
&dr_types::tone::ACR3_DEFAULT,
[scene; 3],
dr_pipeline::view::DEFAULT_CONTRAST,
dr_pipeline::view::DEFAULT_WHITE,
)[0];
let expected = (dr_types::Transfer::Srgb.encode(viewed) * 255.0).round() as i32;
let delta = (i32::from(from_sensor[0]) - expected).abs();
assert!(
+351 -64
View File
@@ -107,6 +107,11 @@ pub struct DemosaicedImage {
height: u32,
/// Carried through for the camera→sRGB transform in the adjust pass.
color_matrix: [f32; 9],
/// TRACES: FR-DEV-3e
/// The camera profile's tables, carried through with the matrix for the
/// adjust pass to upload (D20). `None` for a JPEG and for a raw with no
/// profile.
profile_tables: Option<std::sync::Arc<dr_types::ProfileTables>>,
/// As-shot white balance, the neutral starting point for the WB control.
as_shot_wb: [f32; 3],
/// Whether the texture holds gamma-encoded rather than linear values.
@@ -207,6 +212,12 @@ impl DemosaicedImage {
self.color_matrix
}
/// TRACES: FR-DEV-3e
/// The camera profile's tables this source renders through, if any.
pub fn profile_tables(&self) -> Option<&std::sync::Arc<dr_types::ProfileTables>> {
self.profile_tables.as_ref()
}
/// As-shot white balance multipliers, green-normalised.
///
/// The white balance control is expressed *relative* to these, so its
@@ -318,6 +329,7 @@ impl DemosaicedImage {
width,
height,
color_matrix: IDENTITY_3X3,
profile_tables: None,
as_shot_wb: [1.0, 1.0, 1.0],
// **The identity, and this is the whole reason the field is here
// rather than resolved further down.** A JPEG has already been
@@ -334,6 +346,95 @@ impl DemosaicedImage {
}
impl DemosaicedImage {
/// TRACES: FR-DEV-3g
/// The learned demosaic's output for the photograph `like` was
/// demosaiced from: `width × height` interleaved RGB, linear camera
/// space, normalised as the demosaic normalises — the same texture the
/// classical path made, with the noise gone (denoise.md §2).
///
/// Everything that describes the photograph rather than its pixels —
/// matrix, profile tables, as-shot balance — is `like`'s, so nothing
/// downstream can tell which demosaic ran. A new [`Self::id`], so every
/// cache keyed on the source sees a new source.
pub fn from_rgb_f32(
ctx: &GpuContext,
like: &DemosaicedImage,
width: u32,
height: u32,
rgb: &[f32],
) -> Result<Self, GpuError> {
let n = width as usize * height as usize;
if rgb.len() != n * 3 {
return Err(GpuError::TooLarge(format!(
"{} values for a {width}×{height} RGB image",
rgb.len()
)));
}
let limits = ctx.device.limits();
if width > limits.max_texture_dimension_2d || height > limits.max_texture_dimension_2d {
return Err(GpuError::TooLarge(format!(
"{width}×{height} exceeds the device limit of {}",
limits.max_texture_dimension_2d
)));
}
let mut half = vec![0u16; n * 4];
let one = f32_to_f16_bits(1.0);
let threads = std::thread::available_parallelism().map_or(1, |n| n.get());
let per = n.div_ceil(threads).max(1);
std::thread::scope(|scope| {
for (k, out) in half.chunks_mut(per * 4).enumerate() {
scope.spawn(move || {
for (i, texel) in out.chunks_mut(4).enumerate() {
let src = &rgb[(k * per + i) * 3..(k * per + i) * 3 + 3];
for c in 0..3 {
texel[c] = f32_to_f16_bits_unclamped(src[c]);
}
texel[3] = one;
}
});
}
});
let texture = ctx.device.create_texture_with_data(
&ctx.queue,
&wgpu::TextureDescriptor {
label: Some("learned-demosaic-source"),
size: wgpu::Extent3d {
width,
height,
depth_or_array_layers: 1,
},
mip_level_count: 1,
sample_count: 1,
dimension: wgpu::TextureDimension::D2,
format: Self::FORMAT,
usage: wgpu::TextureUsages::TEXTURE_BINDING | wgpu::TextureUsages::COPY_SRC,
view_formats: &[],
},
wgpu::util::TextureDataOrder::LayerMajor,
bytemuck::cast_slice(&half),
);
Ok(like.sibling(texture, width, height))
}
/// A new source standing for the same photograph as `self`: its
/// description kept, its pixels `texture`, a fresh id.
pub(crate) fn sibling(&self, texture: wgpu::Texture, width: u32, height: u32) -> Self {
let view = texture.create_view(&Default::default());
Self {
texture,
view,
width,
height,
color_matrix: self.color_matrix,
profile_tables: self.profile_tables.clone(),
as_shot_wb: self.as_shot_wb,
non_linear: self.non_linear,
id: next_image_id(),
frame: self.frame,
window: self.window,
}
}
/// TRACES: FR-MRG-3
/// A source that is already RGB in camera space: a linear DNG, which is
/// what a merge writes. No demosaic; the samples are normalised by the
@@ -477,7 +578,8 @@ impl DemosaicedImage {
view,
width,
height,
color_matrix: raw.color_matrix.unwrap_or(IDENTITY_3X3),
color_matrix: rendering_matrix(raw),
profile_tables: raw.profile_tables.clone(),
as_shot_wb: [raw.wb_coeffs[0], raw.wb_coeffs[1], raw.wb_coeffs[2]],
non_linear: false,
id: next_image_id(),
@@ -487,6 +589,23 @@ impl DemosaicedImage {
}
}
/// TRACES: FR-DEV-3e
/// The camera matrix a raw renders through: the file's — identity where the
/// body is uncalibrated, so the image renders with no colour transform
/// rather than not at all — times the baseline exposure as a gain
/// (camera-profiles.md §11).
///
/// A uniform gain commutes with every scene operation before the view
/// transform, so folding it in here is the same as an exposure step at the
/// head of the chain, at no cost. The camera-space tap and the white-balance
/// probe read camera RGB before this matrix and are unaffected.
/// `RawImage::color_matrix` stays the file's: a merge writes a linear DNG
/// from it and must not bake a gain into its pixels.
fn rendering_matrix(raw: &RawImage) -> [f32; 9] {
let gain = raw.baseline_exposure.exp2();
raw.color_matrix.unwrap_or(IDENTITY_3X3).map(|v| v * gain)
}
/// Convert an f32 to half-precision bits, the general case: sign,
/// subnormals, round-to-nearest-even, saturation at the largest finite.
///
@@ -748,62 +867,14 @@ impl Demosaicer {
// TRACES: FR-RAW-3
// The mosaic the demosaic actually reads: the readout with its hot and
// dead photosites repaired. A second buffer rather than in place,
// because every photosite's verdict reads its neighbours' originals.
let repaired = self.ctx.device.create_buffer(&wgpu::BufferDescriptor {
label: Some("raw-repaired"),
size: raw_buf.size(),
usage: wgpu::BufferUsages::STORAGE,
mapped_at_creation: false,
});
let words = packed.len() as u32;
let groups = words.div_ceil(HOT_PIXEL_GROUP).max(1);
// A 24 MP readout is 190,000 workgroups, past the 65,535 one
// dispatch dimension may hold, so the grid folds into rows.
let groups_x = groups.min(
self.ctx
.device
.limits()
.max_compute_workgroups_per_dimension,
);
let groups_y = groups.div_ceil(groups_x);
let hot_params = hot_pixel_params(
// dead photosites repaired.
let hot = self.hot_pass(
raw,
(width, height),
words,
groups_x * HOT_PIXEL_GROUP,
&raw_buf,
packed.len() as u32,
xtrans_tile,
);
let hot_params_buf =
self.ctx
.device
.create_buffer_init(&wgpu::util::BufferInitDescriptor {
label: Some("hot-pixel-params"),
contents: bytemuck::bytes_of(&hot_params),
usage: wgpu::BufferUsages::UNIFORM,
});
let hot_bind_group = self
.ctx
.device
.create_bind_group(&wgpu::BindGroupDescriptor {
label: Some("hot-pixel-bg"),
layout: &self.hot_pixel_layout,
entries: &[
wgpu::BindGroupEntry {
binding: 0,
resource: raw_buf.as_entire_binding(),
},
wgpu::BindGroupEntry {
binding: 1,
resource: hot_params_buf.as_entire_binding(),
},
wgpu::BindGroupEntry {
binding: 2,
resource: repaired.as_entire_binding(),
},
],
});
let params_buf = self
.ctx
.device
@@ -842,7 +913,7 @@ impl Demosaicer {
entries: &[
wgpu::BindGroupEntry {
binding: 0,
resource: repaired.as_entire_binding(),
resource: hot.repaired.as_entire_binding(),
},
wgpu::BindGroupEntry {
binding: 1,
@@ -864,15 +935,7 @@ impl Demosaicer {
// Two passes in one submission. wgpu orders a storage write in one
// pass before a read of the same buffer in the next, so the demosaic
// sees every repair.
{
let mut pass = enc.begin_compute_pass(&wgpu::ComputePassDescriptor {
label: Some("hot-pixel-pass"),
timestamp_writes: None,
});
pass.set_pipeline(&self.hot_pixel_pipeline);
pass.set_bind_group(0, &hot_bind_group, &[]);
pass.dispatch_workgroups(groups_x, groups_y, 1);
}
self.record_hot_pass(&mut enc, &hot);
{
let mut pass = enc.begin_compute_pass(&wgpu::ComputePassDescriptor {
label: Some("demosaic-pass"),
@@ -891,7 +954,8 @@ impl Demosaicer {
height,
// Identity where the body is uncalibrated: the image renders with
// no colour transform rather than not at all.
color_matrix: raw.color_matrix.unwrap_or(IDENTITY_3X3),
color_matrix: rendering_matrix(raw),
profile_tables: raw.profile_tables.clone(),
as_shot_wb: [raw.wb_coeffs[0], raw.wb_coeffs[1], raw.wb_coeffs[2]],
// Whatever the profile database had for this body (FR-DEV-3e),
// resolved at decode because that is the only place the make and
@@ -907,6 +971,221 @@ impl Demosaicer {
}
}
/// The hot-pixel pass's resources for one frame, ready to record.
struct HotPass {
repaired: wgpu::Buffer,
bind_group: wgpu::BindGroup,
groups: (u32, u32),
}
impl Demosaicer {
/// Buffers and bindings for the hot and dead photosite repair of `raw`,
/// whose packed samples are in `raw_buf`.
fn hot_pass(
&self,
raw: &RawImage,
(width, height): (u32, u32),
raw_buf: &wgpu::Buffer,
words: u32,
xtrans_tile: Option<[u32; 4]>,
) -> HotPass {
// A second buffer rather than in place, because every photosite's
// verdict reads its neighbours' originals.
let repaired = self.ctx.device.create_buffer(&wgpu::BufferDescriptor {
label: Some("raw-repaired"),
size: raw_buf.size(),
usage: wgpu::BufferUsages::STORAGE | wgpu::BufferUsages::COPY_SRC,
mapped_at_creation: false,
});
let groups = words.div_ceil(HOT_PIXEL_GROUP).max(1);
// A 24 MP readout is 190,000 workgroups, past the 65,535 one
// dispatch dimension may hold, so the grid folds into rows.
let groups_x = groups.min(
self.ctx
.device
.limits()
.max_compute_workgroups_per_dimension,
);
let groups_y = groups.div_ceil(groups_x);
let hot_params = hot_pixel_params(
raw,
(width, height),
words,
groups_x * HOT_PIXEL_GROUP,
xtrans_tile,
);
let hot_params_buf =
self.ctx
.device
.create_buffer_init(&wgpu::util::BufferInitDescriptor {
label: Some("hot-pixel-params"),
contents: bytemuck::bytes_of(&hot_params),
usage: wgpu::BufferUsages::UNIFORM,
});
let bind_group = self
.ctx
.device
.create_bind_group(&wgpu::BindGroupDescriptor {
label: Some("hot-pixel-bg"),
layout: &self.hot_pixel_layout,
entries: &[
wgpu::BindGroupEntry {
binding: 0,
resource: raw_buf.as_entire_binding(),
},
wgpu::BindGroupEntry {
binding: 1,
resource: hot_params_buf.as_entire_binding(),
},
wgpu::BindGroupEntry {
binding: 2,
resource: repaired.as_entire_binding(),
},
],
});
HotPass {
repaired,
bind_group,
groups: (groups_x, groups_y),
}
}
fn record_hot_pass(&self, enc: &mut wgpu::CommandEncoder, hot: &HotPass) {
let mut pass = enc.begin_compute_pass(&wgpu::ComputePassDescriptor {
label: Some("hot-pixel-pass"),
timestamp_writes: None,
});
pass.set_pipeline(&self.hot_pixel_pipeline);
pass.set_bind_group(0, &hot.bind_group, &[]);
pass.dispatch_workgroups(hot.groups.0, hot.groups.1, 1);
}
/// TRACES: FR-RAW-3 | FR-DEV-3g
/// Repair `raw`'s hot and dead photosites in place, exactly as [`Self::run`]
/// does before it demosaics, and return how many changed.
///
/// For the learned demosaic (denoise.md §2), which reads the same repaired
/// mosaic the classical one does: its training data and its input in the
/// app must have been through this one pass, not a lookalike.
pub fn repair_hot_pixels(&self, raw: &mut RawImage) -> Result<usize, GpuError> {
if raw.samples_per_pixel != 1 {
return Ok(0);
}
let words = self.hot_pixel_words(raw)?;
let mut changed = 0;
for (i, v) in raw.data.iter_mut().enumerate() {
let new = unpack_sample(&words, i);
changed += usize::from(new != *v);
*v = new;
}
Ok(changed)
}
/// The photosites [`Self::repair_hot_pixels`] would replace, in sensor
/// coordinates, without replacing them.
///
/// For the sensor health record (docs/dev/sensor-health.md): one frame's
/// verdict is a candidate list, not a defect map — a single photosite of a
/// star that passes both tests reads the same as a hot one. Which of them
/// is the sensor is decided across frames, by who keeps coming back.
pub fn find_hot_pixels(&self, raw: &RawImage) -> Result<Vec<Photosite>, GpuError> {
if raw.samples_per_pixel != 1 {
return Ok(Vec::new());
}
let words = self.hot_pixel_words(raw)?;
let stride = raw.width.max(1);
Ok(raw
.data
.iter()
.enumerate()
.filter_map(|(i, &v)| {
let new = unpack_sample(&words, i);
(new != v).then(|| Photosite {
x: i as u32 % stride,
y: i as u32 / stride,
hot: new < v,
})
})
.collect())
}
/// The hot-pixel pass over `raw`, read back as packed words.
fn hot_pixel_words(&self, raw: &RawImage) -> Result<Vec<u32>, GpuError> {
let (width, height) = (raw.crop.width.max(1), raw.crop.height.max(1));
let xtrans_tile = raw
.cfa_pattern
.is_xtrans()
.then(|| xtrans_params_for(raw, width, height).tile);
let packed = pack_samples(&raw.data);
let raw_buf = self
.ctx
.device
.create_buffer_init(&wgpu::util::BufferInitDescriptor {
label: Some("raw-samples"),
contents: bytemuck::cast_slice(&packed),
usage: wgpu::BufferUsages::STORAGE,
});
let hot = self.hot_pass(
raw,
(width, height),
&raw_buf,
packed.len() as u32,
xtrans_tile,
);
let readback = self.ctx.device.create_buffer(&wgpu::BufferDescriptor {
label: Some("raw-repaired-readback"),
size: hot.repaired.size(),
usage: wgpu::BufferUsages::COPY_DST | wgpu::BufferUsages::MAP_READ,
mapped_at_creation: false,
});
let mut enc = self
.ctx
.device
.create_command_encoder(&wgpu::CommandEncoderDescriptor {
label: Some("hot-pixel-encoder"),
});
self.record_hot_pass(&mut enc, &hot);
enc.copy_buffer_to_buffer(&hot.repaired, 0, &readback, 0, hot.repaired.size());
self.ctx.queue.submit(Some(enc.finish()));
let slice = readback.slice(..);
let (tx, rx) = std::sync::mpsc::channel();
slice.map_async(wgpu::MapMode::Read, move |r| {
let _ = tx.send(r);
});
self.ctx
.device
.poll(wgpu::PollType::wait_indefinitely())
.map_err(|e| GpuError::Readback(e.to_string()))?;
rx.recv()
.map_err(|e| GpuError::Readback(e.to_string()))?
.map_err(|e| GpuError::Readback(e.to_string()))?;
let words: Vec<u32> = bytemuck::cast_slice(&slice.get_mapped_range()).to_vec();
readback.unmap();
Ok(words)
}
}
/// One photosite the hot-pixel pass judged defective.
#[derive(Copy, Clone, Debug, PartialEq, Eq, Hash)]
pub struct Photosite {
/// Sensor coordinates: the full readout, masked border included.
pub x: u32,
pub y: u32,
/// Read far above its neighbourhood; otherwise far below (dead).
pub hot: bool,
}
/// Sample `i` of a readout packed by [`pack_samples`].
fn unpack_sample(words: &[u32], i: usize) -> u16 {
let w = words[i / 2];
(if i.is_multiple_of(2) {
w & 0xFFFF
} else {
w >> 16
}) as u16
}
const IDENTITY_3X3: [f32; 9] = [1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0];
/// Pack u16 samples two per u32, little-endian within the word.
@@ -1399,6 +1678,8 @@ mod tests {
color_matrix: None,
samples_per_pixel: 1,
profile: None,
profile_tables: None,
baseline_exposure: 0.0,
make: String::new(),
model: String::new(),
crop: CropRect {
@@ -1514,6 +1795,8 @@ mod tests {
color_matrix: None,
samples_per_pixel: 1,
profile: None,
profile_tables: None,
baseline_exposure: 0.0,
make: String::new(),
model: String::new(),
crop: CropRect {
@@ -1810,6 +2093,8 @@ mod tests {
color_matrix: None,
samples_per_pixel: 1,
profile: None,
profile_tables: None,
baseline_exposure: 0.0,
make: String::new(),
model: String::new(),
crop: CropRect {
@@ -1894,6 +2179,8 @@ mod tests {
color_matrix: None,
samples_per_pixel: 1,
profile: None,
profile_tables: None,
baseline_exposure: 0.0,
make: String::new(),
model: String::new(),
crop: CropRect {
+219
View File
@@ -0,0 +1,219 @@
//! TRACES: FR-DEV-3g
//! Grain back into a denoised photograph, as brightness only.
//!
//! The learned denoise's one live control. The network's result and the
//! classical demosaic of the same mosaic differ by the noise the network
//! removed — plus the classical path's colour speckle and demosaic false
//! colour, which nobody wants back. So only the brightness of the difference
//! is returned, in proportion to `grain`:
//!
//! `out = denoised + grain · ΔY / wb`, with `ΔY = Y(wb · (classical − denoised))`
//!
//! `Y` is taken after the as-shot balance and handed back divided by it, so
//! the grain is neutral in the finished picture rather than tinted the
//! colour of the sensor's raw response. At 0 the result is the network's
//! exactly; at 1 the brightness noise is all back, the colour noise none.
//!
//! A pass of its own producing a new source rather than a term in the
//! adjust shader: the blend depends only on the two images and one number,
//! a 20 MP pass is a few milliseconds, and a new source id is all the
//! adjust pass's caches need to know it changed.
use std::sync::Arc;
use crate::demosaic::DemosaicedImage;
use crate::{GpuContext, GpuError};
const SHADER: &str = r#"
struct Params {
grain: f32,
_pad0: f32,
_pad1: f32,
_pad2: f32,
wb: vec4<f32>,
}
@group(0) @binding(0) var denoised: texture_2d<f32>;
@group(0) @binding(1) var classical: texture_2d<f32>;
@group(0) @binding(2) var<uniform> p: Params;
@group(0) @binding(3) var out: texture_storage_2d<rgba16float, write>;
@compute @workgroup_size(8, 8)
fn main(@builtin(global_invocation_id) gid: vec3<u32>) {
let dims = textureDimensions(denoised);
if (gid.x >= dims.x || gid.y >= dims.y) {
return;
}
let xy = vec2<i32>(gid.xy);
let d = textureLoad(denoised, xy, 0).rgb;
let c = textureLoad(classical, xy, 0).rgb;
let wb = p.wb.rgb;
let dy = p.grain * dot(vec3<f32>(0.2126, 0.7152, 0.0722), wb * (c - d));
textureStore(out, xy, vec4<f32>(d + dy / wb, 1.0));
}
"#;
#[repr(C)]
#[derive(Copy, Clone, bytemuck::Pod, bytemuck::Zeroable)]
struct Params {
grain: f32,
_pad: [f32; 3],
wb: [f32; 4],
}
pub struct GrainBlend {
ctx: GpuContext,
pipeline: wgpu::ComputePipeline,
layout: wgpu::BindGroupLayout,
}
impl GrainBlend {
pub fn new(ctx: &GpuContext) -> Self {
let device = &ctx.device;
let module = device.create_shader_module(wgpu::ShaderModuleDescriptor {
label: Some("grain-blend"),
source: wgpu::ShaderSource::Wgsl(SHADER.into()),
});
let texture = |binding| wgpu::BindGroupLayoutEntry {
binding,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::Texture {
sample_type: wgpu::TextureSampleType::Float { filterable: false },
view_dimension: wgpu::TextureViewDimension::D2,
multisampled: false,
},
count: None,
};
let layout = device.create_bind_group_layout(&wgpu::BindGroupLayoutDescriptor {
label: Some("grain-blend-layout"),
entries: &[
texture(0),
texture(1),
wgpu::BindGroupLayoutEntry {
binding: 2,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::Buffer {
ty: wgpu::BufferBindingType::Uniform,
has_dynamic_offset: false,
min_binding_size: None,
},
count: None,
},
wgpu::BindGroupLayoutEntry {
binding: 3,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::StorageTexture {
access: wgpu::StorageTextureAccess::WriteOnly,
format: DemosaicedImage::FORMAT,
view_dimension: wgpu::TextureViewDimension::D2,
},
count: None,
},
],
});
let pipeline_layout = device.create_pipeline_layout(&wgpu::PipelineLayoutDescriptor {
label: Some("grain-blend-pipeline-layout"),
bind_group_layouts: &[Some(&layout)],
immediate_size: 0,
});
let pipeline = device.create_compute_pipeline(&wgpu::ComputePipelineDescriptor {
label: Some("grain-blend"),
layout: Some(&pipeline_layout),
module: &module,
entry_point: Some("main"),
compilation_options: Default::default(),
cache: None,
});
Self {
ctx: ctx.clone(),
pipeline,
layout,
}
}
/// `denoised` with `grain` (0–1) of `classical`'s brightness noise back.
/// Both must be the same photograph at the same size.
pub fn blend(
&self,
denoised: &DemosaicedImage,
classical: &DemosaicedImage,
grain: f32,
) -> Result<Arc<DemosaicedImage>, GpuError> {
let (w, h) = (denoised.texture().width(), denoised.texture().height());
if (classical.texture().width(), classical.texture().height()) != (w, h) {
return Err(GpuError::TooLarge(format!(
"grain from a {}×{} source into a {w}×{h} one",
classical.texture().width(),
classical.texture().height()
)));
}
let device = &self.ctx.device;
let texture = device.create_texture(&wgpu::TextureDescriptor {
label: Some("grain-blended-source"),
size: wgpu::Extent3d {
width: w,
height: h,
depth_or_array_layers: 1,
},
mip_level_count: 1,
sample_count: 1,
dimension: wgpu::TextureDimension::D2,
format: DemosaicedImage::FORMAT,
usage: wgpu::TextureUsages::STORAGE_BINDING
| wgpu::TextureUsages::TEXTURE_BINDING
| wgpu::TextureUsages::COPY_SRC,
view_formats: &[],
});
let out_view = texture.create_view(&Default::default());
let wb = denoised.as_shot_wb();
let g = wb[1].max(1e-6);
let params = Params {
grain: grain.clamp(0.0, 1.0),
_pad: [0.0; 3],
// Green-normalised, and never zero: the shader divides by it.
wb: [(wb[0] / g).max(1e-3), 1.0, (wb[2] / g).max(1e-3), 1.0],
};
use wgpu::util::DeviceExt;
let buffer = device.create_buffer_init(&wgpu::util::BufferInitDescriptor {
label: Some("grain-blend-params"),
contents: bytemuck::bytes_of(&params),
usage: wgpu::BufferUsages::UNIFORM,
});
let bind = device.create_bind_group(&wgpu::BindGroupDescriptor {
label: Some("grain-blend-bg"),
layout: &self.layout,
entries: &[
wgpu::BindGroupEntry {
binding: 0,
resource: wgpu::BindingResource::TextureView(denoised.view()),
},
wgpu::BindGroupEntry {
binding: 1,
resource: wgpu::BindingResource::TextureView(classical.view()),
},
wgpu::BindGroupEntry {
binding: 2,
resource: buffer.as_entire_binding(),
},
wgpu::BindGroupEntry {
binding: 3,
resource: wgpu::BindingResource::TextureView(&out_view),
},
],
});
let mut enc = device.create_command_encoder(&wgpu::CommandEncoderDescriptor {
label: Some("grain-blend"),
});
{
let mut pass = enc.begin_compute_pass(&wgpu::ComputePassDescriptor {
label: Some("grain-blend"),
timestamp_writes: None,
});
pass.set_pipeline(&self.pipeline);
pass.set_bind_group(0, &bind, &[]);
pass.dispatch_workgroups(w.div_ceil(8), h.div_ceil(8), 1);
}
self.ctx.queue.submit(Some(enc.finish()));
Ok(Arc::new(denoised.sibling(texture, w, h)))
}
}
+3 -1
View File
@@ -26,6 +26,7 @@ mod demosaic;
mod detail;
mod error;
mod focus;
mod grain;
mod histogram;
mod mask;
mod merge;
@@ -37,10 +38,11 @@ pub use adjust::AdjustPass;
// rather than an implementation detail: a detail pass is guaranteed linear,
// unclipped, full internal precision (FR-DEV-2), and anyone reasoning about
// VRAM at 24 MP needs to know what an intermediate costs.
pub use demosaic::{DemosaicedImage, Demosaicer};
pub use demosaic::{DemosaicedImage, Demosaicer, Photosite};
pub use detail::INTERMEDIATE_FORMAT as DETAIL_INTERMEDIATE_FORMAT;
pub use error::GpuError;
pub use focus::{FocusPeakPass, FocusPeaking, PeakColour, PeakSensitivity};
pub use grain::GrainBlend;
pub use merge::{Band, MergeFrame, MergeOutput, MergePass};
// Renamed on the way out: `BINS` says enough inside `histogram`, and nothing
// at all at a crate root shared with demosaic and segmentation.
+109 -8
View File
@@ -30,18 +30,27 @@
//! are the caller's to provide and cache — `source` is asked for frame `k`
//! as it is needed, and a caller short of memory may demosaic on demand.
//!
//! # The blend
//!
//! With a seam map (`dr_pano::seam`), a frame's weight at a pixel is its
//! share of the map about that pixel — whole on its own side of a seam,
//! nothing on the other, and a ramp across a window `seam_blend` pixels
//! wide that follows the seam. Without one, or where the map has nothing
//! to say, the weight is the distance to the frame's edge over `feather`,
//! which hides exposure steps and does not hide parallax: the average draws
//! anything the frames disagree on twice.
//!
//! # What is not here yet
//!
//! A feathered blend, not seams and a Laplacian pyramid: the weight is the
//! distance to the frame's edge, which hides exposure steps and small
//! misalignments and does not hide parallax. Gain is a scalar per frame
//! the caller supplies. Both are panorama.md §10's step 5, after the path
//! writes a file end to end.
//! A Laplacian pyramid, which would let the seam's blend be narrow for
//! detail and wide for exposure at once. Gain is a scalar per frame the
//! caller supplies.
use std::sync::Arc;
use dr_pano::bundle::Cameras;
use dr_pano::projection::{Bounds, Projection};
use dr_pano::seam::SeamMap;
use wgpu::util::DeviceExt;
use crate::readback::await_mapping;
@@ -58,7 +67,7 @@ pub struct MergeFrame {
}
/// The output the merge produces.
#[derive(Debug, Clone, Copy, PartialEq)]
#[derive(Debug, Clone, PartialEq)]
pub struct MergeOutput {
pub projection: Projection,
/// The projection's scale in output pixels: the cylinder's radius, the
@@ -69,6 +78,11 @@ pub struct MergeOutput {
pub bounds: Bounds,
/// Pixels over which a frame's weight ramps up from its edge.
pub feather: f32,
/// Which frame each part of the output is taken from, laid out at the
/// proxies' scale; `None` for the feathered average everywhere.
pub seams: Option<Arc<SeamMap>>,
/// The width, in output pixels, of the blend across a seam.
pub seam_blend: f32,
/// Chunk size: the unit of GPU work and of memory.
pub chunk: (u32, u32),
/// Multiplies a normalised sample (1.0 = white) to the sensor's scale.
@@ -115,6 +129,12 @@ struct WarpParams {
feather: f32,
clip_onset: f32,
balance: [f32; 4],
seam_origin: [f32; 2],
seam_size: [u32; 2],
seam_px: f32,
seam_radius: f32,
frame_index: u32,
seam_on: u32,
}
#[repr(C)]
@@ -182,6 +202,16 @@ impl MergePass {
count: None,
},
storage(2, false),
wgpu::BindGroupLayoutEntry {
binding: 3,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::Texture {
sample_type: wgpu::TextureSampleType::Uint,
view_dimension: wgpu::TextureViewDimension::D2,
multisampled: false,
},
count: None,
},
],
});
let resolve_layout =
@@ -279,6 +309,32 @@ impl MergePass {
let mut band_cov = vec![false; (out_w * ch) as usize];
let mut chunk_px: Vec<u32> = Vec::new();
// The seam map, once for the whole output, and where it sits in
// this output's coordinates. A one-texel stand-in when there is
// none, because the binding is not optional.
let (seam_tex, seam_origin, seam_px, seam_radius, seam_size) = match &output.seams {
Some(m) => {
let ((ou, ov), px) = m.at_scale(output.scale);
let radius = m.blend_radius(output.scale, f64::from(output.seam_blend));
(
self.label_texture(m.width as u32, m.height as u32, &m.labels),
[ou as f32, ov as f32],
px as f32,
radius as f32,
[m.width as u32, m.height as u32],
)
}
None => (
self.label_texture(1, 1, &[dr_pano::seam::NONE]),
[0.0; 2],
1.0,
1.0,
[1, 1],
),
};
let seam_view = seam_tex.create_view(&Default::default());
let seam_on = u32::from(output.seams.is_some());
let mut y = 0u32;
while y < out_h {
let rows = ch.min(out_h - y);
@@ -346,8 +402,14 @@ impl MergePass {
output.balance[2].max(1e-3),
0.0,
],
seam_origin,
seam_size,
seam_px,
seam_radius,
frame_index: k as u32,
seam_on,
};
self.accumulate(&params, tile);
self.accumulate(&params, tile, &seam_view);
}
self.resolve_chunk((cols, rows), output.sample_scale, &mut chunk_px)?;
@@ -385,7 +447,42 @@ impl MergePass {
self.ctx.queue.submit(Some(enc.finish()));
}
fn accumulate(&mut self, params: &WarpParams, tile: &wgpu::Texture) {
/// The seam map's labels as an `r8uint` texture.
fn label_texture(&self, width: u32, height: u32, labels: &[u8]) -> wgpu::Texture {
let size = wgpu::Extent3d {
width,
height,
depth_or_array_layers: 1,
};
let tex = self.ctx.device.create_texture(&wgpu::TextureDescriptor {
label: Some("merge-seams"),
size,
mip_level_count: 1,
sample_count: 1,
dimension: wgpu::TextureDimension::D2,
format: wgpu::TextureFormat::R8Uint,
usage: wgpu::TextureUsages::TEXTURE_BINDING | wgpu::TextureUsages::COPY_DST,
view_formats: &[],
});
self.ctx.queue.write_texture(
wgpu::TexelCopyTextureInfo {
texture: &tex,
mip_level: 0,
origin: wgpu::Origin3d::ZERO,
aspect: wgpu::TextureAspect::All,
},
labels,
wgpu::TexelCopyBufferLayout {
offset: 0,
bytes_per_row: Some(width),
rows_per_image: Some(height),
},
size,
);
tex
}
fn accumulate(&mut self, params: &WarpParams, tile: &wgpu::Texture, seams: &wgpu::TextureView) {
let chunk = (params.chunk_size[0], params.chunk_size[1]);
let uniforms = self
.ctx
@@ -417,6 +514,10 @@ impl MergePass {
binding: 2,
resource: acc.as_entire_binding(),
},
wgpu::BindGroupEntry {
binding: 3,
resource: wgpu::BindingResource::TextureView(seams),
},
],
});
let mut enc = self.ctx.device.create_command_encoder(&Default::default());
+85 -3
View File
@@ -5,8 +5,9 @@
// pixel it asks which direction that pixel looks along, turns the
// direction into the frame's camera, projects it to a source pixel, and
// if that pixel is inside the tile that was rendered for this chunk,
// samples it and adds it — weighted by its distance from the frame's edge
// — into the accumulator. `resolve` runs once per chunk after every frame
// samples it and adds it — weighted by the frame's share of the seam map
// there, or by its distance from the frame's edge where there is no map —
// into the accumulator. `resolve` runs once per chunk after every frame
// has been added: divides the sums by the weights and packs the result as
// sixteen-bit samples at the sensor's scale (FR-MRG-3).
//
@@ -50,12 +51,83 @@ struct Params {
// white balance the composite will be developed with.
clip_onset: f32,
balance: vec4<f32>,
// The seam map (`dr_pano::seam`): where its texel (0, 0)'s corner sits
// in this output's centred coordinates, its size, output pixels per
// texel, the blend's radius in texels, which frame this dispatch is,
// and whether there is a map at all.
seam_origin: vec2<f32>,
seam_size: vec2<u32>,
seam_px: f32,
seam_radius: f32,
frame_index: u32,
seam_on: u32,
};
@group(0) @binding(0) var<uniform> p: Params;
@group(0) @binding(1) var tile: texture_2d<f32>;
// rgb·w summed, then w: four floats per chunk pixel.
@group(0) @binding(2) var<storage, read_write> acc: array<vec4<f32>>;
// One frame index per texel, 255 for none.
@group(0) @binding(3) var seams: texture_2d<u32>;
const NO_FRAME: u32 = 255u;
fn label(i: i32, j: i32) -> u32 {
if (i < 0 || j < 0 || i >= i32(p.seam_size.x) || j >= i32(p.seam_size.y)) {
return NO_FRAME;
}
return textureLoad(seams, vec2<i32>(i, j), 0).r;
}
// This frame's share of the seam map about output point (u, v): the
// tent-weighted fraction of the texels within the radius that it owns, and
// the weight of the texels owned by anyone (zero where the map has nothing
// to say). `SeamMap::share` verbatim.
fn seam_share(u: f32, v: f32) -> vec2<f32> {
let x = (u - p.seam_origin.x) / p.seam_px - 0.5;
let y = (v - p.seam_origin.y) / p.seam_px - 0.5;
let r = max(p.seam_radius, 1.0);
let x0 = i32(ceil(x - r));
let x1 = i32(floor(x + r));
let y0 = i32(ceil(y - r));
let y1 = i32(floor(y + r));
// Most pixels are nowhere near a seam: if the window's corners, edge
// midpoints and centre agree, so does the window. A seam crossing it
// has to cross its border, between two of those.
let xm = i32(round(x));
let ym = i32(round(y));
let c = label(xm, ym);
if (label(x0, y0) == c && label(x1, y0) == c && label(x0, y1) == c && label(x1, y1) == c
&& label(xm, y0) == c && label(xm, y1) == c && label(x0, ym) == c && label(x1, ym) == c) {
if (c == NO_FRAME) {
return vec2<f32>(0.0, 0.0);
}
return vec2<f32>(select(0.0, 1.0, c == p.frame_index), 1.0);
}
var mine = 0.0;
var owned = 0.0;
for (var j = y0; j <= y1; j = j + 1) {
let wy = 1.0 - abs(y - f32(j)) / r;
if (wy <= 0.0) {
continue;
}
for (var i = x0; i <= x1; i = i + 1) {
let wx = 1.0 - abs(x - f32(i)) / r;
let l = label(i, j);
if (wx <= 0.0 || l == NO_FRAME) {
continue;
}
owned = owned + wx * wy;
if (l == p.frame_index) {
mine = mine + wx * wy;
}
}
}
if (owned <= 0.0) {
return vec2<f32>(0.0, 0.0);
}
return vec2<f32>(mine / owned, 1.0);
}
fn to_direction(u: f32, v: f32) -> vec3<f32> {
let s = p.proj_scale;
@@ -98,7 +170,17 @@ fn warp(@builtin(global_invocation_id) gid: vec3<u32>) {
if (edge <= 0.0) {
return;
}
let w = clamp(edge / max(p.feather, 1.0), 0.0, 1.0);
var w = clamp(edge / max(p.feather, 1.0), 0.0, 1.0);
// With seams, the share of the map scales it. The small floor keeps
// the feather underneath as the answer wherever no frame that reaches
// this pixel owns it — the map is coarser than the output, so at the
// frames' outer edges it can name a frame that falls just short.
if (p.seam_on != 0u) {
let s = seam_share(u, v);
if (s.y > 0.0) {
w = w * (s.x + 1e-4);
}
}
// Into the tile.
let tx = sx - p.tile_origin.x;
let ty = sy - p.tile_origin.y;
+303
View File
@@ -0,0 +1,303 @@
//! TRACES: FR-DEV-3e
//! The camera profile's tables, end to end on a device (D20).
//!
//! `dr-pipeline` holds the lookup to the DNG SDK's algorithm on the CPU
//! (`ops::camera_profile::apply_reference`). Nothing there would notice a
//! shader that disagreed with it — a transposed constant matrix, an index
//! off by one column, a buffer bound in the wrong order — so this renders a
//! frame of 256 different colours through tables that move every one of them
//! a long way, and holds each pixel to the reference.
//!
//! The source is a linear three-sample frame, so the colours arrive exactly
//! as written with no demosaic between, and an identity stands in the view
//! transform's place so the readback is the scene colour, display-encoded.
use std::sync::Arc;
use dr_decode::{CfaPattern, CropRect, RawImage};
use dr_gpu::{AdjustPass, Demosaicer, GpuContext};
use dr_pipeline::descriptor::{Attribute, LocalizedKey, OpDescriptor, OpId, ParamId};
use dr_pipeline::operation::{Operation, Stage, Uniform};
use dr_pipeline::ops::camera_profile::{apply_reference, CameraProfile, APPLY, LOOK, PROFILE_LOOK};
use dr_types::{HueSatTable, ProfileOrigin, ProfileTables, Transfer};
const SIZE: u32 = 16;
fn ctx() -> Option<GpuContext> {
pollster::block_on(GpuContext::new_headless()).ok()
}
/// An identity in the view transform's place.
struct IdentityView;
impl Operation for IdentityView {
fn descriptor(&self) -> Arc<OpDescriptor> {
Arc::new(OpDescriptor {
id: OpId("identity_view"),
label: LocalizedKey("identity_view"),
params: Vec::new(),
attributes: vec![Attribute::Tone],
})
}
fn set_param(&mut self, _: ParamId, _: f32) {}
fn param(&self, _: ParamId) -> f32 {
0.0
}
fn is_active(&self) -> bool {
true
}
fn stage(&self) -> Stage {
Stage::View
}
fn renders(&self) -> bool {
true
}
fn wgsl_body(&self) -> String {
String::new()
}
fn uniforms(&self) -> Vec<Uniform> {
Vec::new()
}
}
/// 256 colours across hue, saturation and value, kept under the prologue's
/// highlight desaturation and above black.
fn colours() -> Vec<[f32; 3]> {
(0..SIZE * SIZE)
.map(|i| {
let f = |k: u32| {
let x = (i.wrapping_mul(2_654_435_761).rotate_left(k * 7) >> 8) % 1000;
0.04 + 0.86 * x as f32 / 1000.0
};
[f(1), f(2), f(3)]
})
.collect()
}
fn frame(tables: Option<ProfileTables>) -> RawImage {
let data = colours()
.iter()
.flat_map(|c| c.map(|v| (v * 65535.0).round() as u16))
.collect();
RawImage {
width: SIZE,
height: SIZE,
data,
cfa_pattern: CfaPattern::Rggb,
black_level: [0; 4],
white_level: u16::MAX,
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
color_matrix: Some([1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0]),
samples_per_pixel: 3,
profile: None,
profile_tables: tables.map(Arc::new),
baseline_exposure: 0.0,
make: String::new(),
model: String::new(),
crop: CropRect {
x: 0,
y: 0,
width: SIZE,
height: SIZE,
},
}
}
/// Tables that move every colour by a different amount: hue shifts of tens
/// of degrees, saturation scales either side of one, and a 3-D, sRGB-indexed
/// look whose value scale varies down the value axis.
fn strong_tables() -> ProfileTables {
let (hd, sd) = (12u32, 5u32);
let hue_sat = (0..hd * sd)
.map(|i| {
let (h, s) = (i / sd, i % sd);
let a = h as f32 / hd as f32 * std::f32::consts::TAU;
[25.0 * a.sin(), 1.0 + 0.3 * a.cos() * s as f32 / 4.0, 1.0]
})
.collect();
let (lh, ls, lv) = (8u32, 4u32, 5u32);
let look = (0..lh * ls * lv)
.map(|i| {
let v = i / (lh * ls);
let h = (i / ls) % lh;
[
-15.0 + 4.0 * h as f32,
1.25 - 0.05 * v as f32,
0.85 + 0.06 * v as f32,
]
})
.collect();
let mut look = HueSatTable::new(lh, ls, lv, true, look).unwrap();
look.srgb_encoded = true;
ProfileTables {
name: "strong".into(),
origin: ProfileOrigin::Embedded,
hue_sat: Some(HueSatTable::new(hd, sd, 1, false, hue_sat).unwrap()),
look: Some(look),
tone_curve: None,
}
}
fn render(ctx: &GpuContext, raw: &RawImage, op: CameraProfile) -> Vec<[u8; 3]> {
let source = Demosaicer::new(ctx)
.expect("demosaicer")
.run(raw)
.expect("upload");
let ops: Vec<Box<dyn Operation>> = vec![Box::new(op), Box::new(IdentityView)];
let shader = dr_pipeline::compose(&ops);
let mut adjust = AdjustPass::new(ctx);
adjust.render(&source, &shader, SIZE, SIZE).expect("render");
let (pixels, _, _) = adjust.export_pixels().expect("readback");
pixels.chunks_exact(4).map(|p| [p[0], p[1], p[2]]).collect()
}
/// The profile at the strength it states — the look table on, as the
/// reference applies it at 1.0. Not the default, which leaves it off (D21).
fn as_stated() -> CameraProfile {
let mut op = CameraProfile::new();
op.set_param(LOOK, PROFILE_LOOK);
op
}
fn encode(c: [f32; 3]) -> [i32; 3] {
c.map(|v| (Transfer::Srgb.encode(v.clamp(0.0, 1.0)) * 255.0).round() as i32)
}
fn assert_agrees(got: &[[u8; 3]], expected: impl Fn([f32; 3]) -> [f32; 3], what: &str) {
let mut moved = 0;
for (i, (c, g)) in colours().into_iter().zip(got).enumerate() {
let want = encode(expected(c));
let g = g.map(i32::from);
// Two 8-bit steps: the half-float source and intermediate, and the
// rounding either side of the encode.
assert!(
want.iter().zip(g).all(|(w, g)| (w - g).abs() <= 2),
"{what}: pixel {i} {c:?} rendered {g:?}, the reference says {want:?}"
);
if want != encode(c) {
moved += 1;
}
}
assert!(
moved > 200,
"{what}: only {moved} of 256 colours moved; the test proves little"
);
}
#[test]
fn the_shader_agrees_with_the_cpu_reference() {
let Some(ctx) = ctx() else {
eprintln!("skipping: no GPU adapter");
return;
};
let tables = strong_tables();
let got = render(&ctx, &frame(Some(tables.clone())), as_stated());
assert_agrees(
&got,
|c| apply_reference(&tables, c, 1.0),
"as the profile states it",
);
let mut doubled = CameraProfile::new();
doubled.set_param(LOOK, 200.0);
let got = render(&ctx, &frame(Some(tables.clone())), doubled);
assert_agrees(&got, |c| apply_reference(&tables, c, 2.0), "look at 200%");
}
#[test]
fn camera_raw_tone_agrees_with_its_cpu_reference() {
// TRACES: FR-DEV-3j
// D21's rendering on 256 colours: the ProPhoto round trip, the clip, the
// curve from the profile buffer's placeholder, and RGBTone's placement
// of the middle channel, against `camera_raw::apply_reference`.
let Some(ctx) = ctx() else {
eprintln!("skipping: no GPU adapter");
return;
};
let source = Demosaicer::new(&ctx)
.expect("demosaicer")
.run(&frame(None))
.expect("upload");
let mut view = dr_pipeline::ops::ViewTransform::new();
view.set_param(
dr_pipeline::ops::view_transform::CURVE,
dr_pipeline::ops::view_transform::CAMERA_RAW,
);
let ops: Vec<Box<dyn Operation>> = vec![Box::new(view)];
let shader = dr_pipeline::compose(&ops);
let mut adjust = AdjustPass::new(&ctx);
adjust.render(&source, &shader, SIZE, SIZE).expect("render");
let (pixels, _, _) = adjust.export_pixels().expect("readback");
let got: Vec<[u8; 3]> = pixels.chunks_exact(4).map(|p| [p[0], p[1], p[2]]).collect();
let curve = &dr_types::tone::ACR3_DEFAULT;
assert_agrees(
&got,
|c| {
dr_pipeline::camera_raw::apply_reference(
curve,
c,
dr_pipeline::view::DEFAULT_CONTRAST,
dr_pipeline::view::DEFAULT_WHITE,
)
},
"DNG reference tone",
);
}
#[test]
fn switched_off_or_absent_the_render_is_unchanged() {
let Some(ctx) = ctx() else {
eprintln!("skipping: no GPU adapter");
return;
};
let bare = render(&ctx, &frame(None), CameraProfile::new());
let mut off = CameraProfile::new();
off.set_param(APPLY, 0.0);
let switched_off = render(&ctx, &frame(Some(strong_tables())), off);
assert_eq!(bare, switched_off, "the switch off is the matrix alone");
// Against the source colours, one 8-bit step for the half-float texture
// the source is uploaded in; the exact comparison is the one above.
for (c, g) in colours().into_iter().zip(&bare) {
let want = encode(c);
assert!(
want.iter()
.zip(g)
.all(|(w, g)| (w - i32::from(*g)).abs() <= 1),
"no tables, no change: {c:?} rendered {g:?}"
);
}
}
#[test]
fn the_libraries_adobe_standard_renders_as_the_reference_does() {
// The real tables, when the library's 6D DNG is on this machine: a 90×30
// HueSatMap and a 36×8×16 LookTable, at the sizes no synthetic test
// reaches.
let Some(ctx) = ctx() else {
eprintln!("skipping: no GPU adapter");
return;
};
let path = std::env::var_os("DR_DCP_SAMPLE")
.map(std::path::PathBuf::from)
.or_else(|| {
std::env::var_os("HOME").map(|h| {
std::path::PathBuf::from(h).join("Nextcloud/PhotosRaw/2017/2017-08-12/_MG_9080.dng")
})
});
let Some(bytes) = path.and_then(|p| std::fs::read(p).ok()) else {
eprintln!("skipping: no sample DNG");
return;
};
let tables = dr_decode::dcp::embedded_in(&bytes)
.expect("Adobe Standard")
.tables(5000.0, ProfileOrigin::Embedded);
let got = render(&ctx, &frame(Some(tables.clone())), as_stated());
for (i, (c, g)) in colours().into_iter().zip(&got).enumerate() {
let want = encode(apply_reference(&tables, c, 1.0));
let g = g.map(i32::from);
assert!(
want.iter().zip(g).all(|(w, g)| (w - g).abs() <= 2),
"pixel {i} {c:?} rendered {g:?}, the reference says {want:?}"
);
}
}
+2
View File
@@ -50,6 +50,8 @@ fn flat_raw(level: u16) -> RawImage {
// film. `dr-pipeline` asserts the suppression on the generated source.
samples_per_pixel: 1,
profile: None,
profile_tables: None,
baseline_exposure: 0.0,
make: String::new(),
model: String::new(),
crop: CropRect {
+146
View File
@@ -0,0 +1,146 @@
//! TRACES: FR-DEV-3g
//! The grain blend, read back off the device.
use dr_decode::{CfaPattern, CropRect, RawImage};
use dr_gpu::{DemosaicedImage, Demosaicer, GpuContext, GrainBlend};
const W: u32 = 16;
const H: u32 = 8;
fn ctx() -> Option<GpuContext> {
pollster::block_on(GpuContext::new_headless()).ok()
}
/// A photograph to stand the uploads beside: its as-shot balance is what
/// the grain is made neutral under.
fn like(ctx: &GpuContext) -> DemosaicedImage {
let raw = RawImage {
width: W,
height: H,
data: vec![400; (W * H) as usize],
cfa_pattern: CfaPattern::Rggb,
black_level: [0; 4],
white_level: 4095,
wb_coeffs: [2.0, 1.0, 1.5, 1.0],
color_matrix: Some([1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0]),
samples_per_pixel: 1,
profile: None,
profile_tables: None,
baseline_exposure: 0.0,
make: String::new(),
model: String::new(),
crop: CropRect {
x: 0,
y: 0,
width: W,
height: H,
},
};
Demosaicer::new(ctx).unwrap().run(&raw).unwrap()
}
fn read(ctx: &GpuContext, img: &DemosaicedImage) -> Vec<[f32; 4]> {
let (w, h) = (img.texture().width(), img.texture().height());
let padded =
(w * 8).div_ceil(wgpu::COPY_BYTES_PER_ROW_ALIGNMENT) * wgpu::COPY_BYTES_PER_ROW_ALIGNMENT;
let buf = ctx.device.create_buffer(&wgpu::BufferDescriptor {
label: None,
size: (padded * h) as u64,
usage: wgpu::BufferUsages::COPY_DST | wgpu::BufferUsages::MAP_READ,
mapped_at_creation: false,
});
let mut enc = ctx.device.create_command_encoder(&Default::default());
enc.copy_texture_to_buffer(
img.texture().as_image_copy(),
wgpu::TexelCopyBufferInfo {
buffer: &buf,
layout: wgpu::TexelCopyBufferLayout {
offset: 0,
bytes_per_row: Some(padded),
rows_per_image: Some(h),
},
},
wgpu::Extent3d {
width: w,
height: h,
depth_or_array_layers: 1,
},
);
ctx.queue.submit(Some(enc.finish()));
let slice = buf.slice(..);
slice.map_async(wgpu::MapMode::Read, |_| {});
ctx.device
.poll(wgpu::PollType::wait_indefinitely())
.unwrap();
let bytes = slice.get_mapped_range();
let mut out = Vec::new();
for y in 0..h as usize {
let row: &[u16] =
bytemuck::cast_slice(&bytes[y * padded as usize..y * padded as usize + w as usize * 8]);
for t in row.chunks(4) {
out.push([0, 1, 2, 3].map(|c| half_to_f32(t[c])));
}
}
out
}
fn half_to_f32(h: u16) -> f32 {
let s = if h & 0x8000 != 0 { -1.0 } else { 1.0 };
let e = ((h >> 10) & 0x1f) as i32;
let m = (h & 0x3ff) as f32;
if e == 0 {
s * m * 2f32.powi(-24)
} else {
s * (1.0 + m / 1024.0) * 2f32.powi(e - 15)
}
}
#[test]
fn grain_returns_only_neutral_brightness() {
let Some(ctx) = ctx() else {
eprintln!("no GPU adapter; skipping");
return;
};
let base = like(&ctx);
let n = (W * H) as usize;
let d: Vec<f32> = (0..n).flat_map(|_| [0.20, 0.30, 0.10]).collect();
// The classical result: the same colour plus noise, coloured noise too.
let c: Vec<f32> = (0..n)
.flat_map(|i| {
let a = ((i * 37) % 11) as f32 / 110.0 - 0.05;
let b = ((i * 53) % 7) as f32 / 140.0 - 0.025;
[0.20 + a, 0.30 + b, 0.10 - a]
})
.collect();
let denoised = DemosaicedImage::from_rgb_f32(&ctx, &base, W, H, &d).unwrap();
let classical = DemosaicedImage::from_rgb_f32(&ctx, &base, W, H, &c).unwrap();
let blend = GrainBlend::new(&ctx);
let wb = [2.0f32, 1.0, 1.5];
let none = read(&ctx, &blend.blend(&denoised, &classical, 0.0).unwrap());
for p in &none {
for ch in 0..3 {
assert!(
(p[ch] - d[ch]).abs() < 1e-3,
"grain 0 must be the network's result: {p:?}"
);
}
}
let all = read(&ctx, &blend.blend(&denoised, &classical, 1.0).unwrap());
for (i, p) in all.iter().enumerate() {
let want_dy: f32 = [0.2126f32, 0.7152, 0.0722]
.iter()
.enumerate()
.map(|(ch, k)| k * wb[ch] * (c[i * 3 + ch] - d[ch]))
.sum();
// After white balance every channel moved by the same amount.
for ch in 0..3 {
let moved = wb[ch] * (p[ch] - d[ch]);
assert!(
(moved - want_dy).abs() < 2e-3,
"pixel {i} channel {ch}: moved {moved}, want {want_dy}"
);
}
}
}
+73 -1
View File
@@ -7,7 +7,7 @@
//! see of a defect it missed is the coloured cross the demosaic makes of it.
use dr_decode::{CfaPattern, CropRect, RawImage};
use dr_gpu::{AdjustPass, Demosaicer, GpuContext};
use dr_gpu::{AdjustPass, Demosaicer, GpuContext, Photosite};
use dr_pipeline::EditGraph;
const SIZE: u32 = 36;
@@ -34,6 +34,8 @@ fn frame(pattern: CfaPattern, level: u16, set: &[(u32, u32, u16)]) -> RawImage {
color_matrix: Some([1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0]),
samples_per_pixel: 1,
profile: None,
profile_tables: None,
baseline_exposure: 0.0,
make: String::new(),
model: String::new(),
crop: CropRect {
@@ -139,3 +141,73 @@ fn a_hot_photosite_on_x_trans_is_invisible() {
let diff = worst(&clean, &hot);
assert!(diff <= 1, "a hot X-Trans photosite still shows, by {diff}");
}
/// The repair alone, read back (FR-DEV-3g): the learned demosaic takes the
/// mosaic this pass leaves, so it must be the same pass and nothing more —
/// the hot photosite replaced, a real highlight and every other photosite
/// untouched.
#[test]
fn the_repaired_mosaic_reads_back_with_only_the_defect_changed() {
let Some(ctx) = ctx() else {
eprintln!("no GPU adapter; skipping");
return;
};
let d = Demosaicer::new(&ctx).expect("demosaicer");
let mut star = vec![(MIDDLE, MIDDLE, WHITE)];
for dy in 0..3 {
for dx in 0..3 {
star.push((4 + dx, 4 + dy, WHITE));
}
}
let before = frame(CfaPattern::Rggb, 40, &star);
let mut raw = before.clone();
let changed = d.repair_hot_pixels(&mut raw).expect("repair");
assert_eq!(changed, 1, "only the lone hot photosite should change");
let at = (MIDDLE * SIZE + MIDDLE) as usize;
assert_eq!(
raw.data[at], 40,
"repaired to its brightest same-colour neighbour"
);
let others = (0..raw.data.len()).filter(|&i| i != at);
assert!(others.into_iter().all(|i| raw.data[i] == before.data[i]));
}
/// Finding without repairing (docs/dev/sensor-health.md): the same verdict as
/// the repair, as sensor coordinates, with the frame left as it was. The
/// sensor health record builds on this, so it must name exactly the
/// photosites the repair would change — the hot one and the dead one, and
/// not the star.
#[test]
fn finding_names_what_the_repair_would_change_and_changes_nothing() {
let Some(ctx) = ctx() else {
eprintln!("no GPU adapter; skipping");
return;
};
let d = Demosaicer::new(&ctx).expect("demosaicer");
let mut set = vec![(MIDDLE, MIDDLE, WHITE), (9, 25, 0)];
for dy in 0..3 {
for dx in 0..3 {
set.push((4 + dx, 4 + dy, WHITE));
}
}
let raw = frame(CfaPattern::Rggb, 1600, &set);
let mut found = d.find_hot_pixels(&raw).expect("find");
found.sort_by_key(|p| (p.y, p.x));
assert_eq!(
found,
vec![
Photosite {
x: MIDDLE,
y: MIDDLE,
hot: true
},
Photosite {
x: 9,
y: 25,
hot: false
},
]
);
let mut repaired = raw.clone();
assert_eq!(d.repair_hot_pixels(&mut repaired).expect("repair"), 2);
}
+2
View File
@@ -94,6 +94,8 @@ fn flat(ctx: &GpuContext, level: f32) -> dr_gpu::DemosaicedImage {
color_matrix: Some([1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0]),
samples_per_pixel: 1,
profile: None,
profile_tables: None,
baseline_exposure: 0.0,
make: String::new(),
model: String::new(),
crop: CropRect {
+2
View File
@@ -58,6 +58,8 @@ fn linear_frame(w: u32, h: u32, noise: bool) -> RawImage {
color_matrix: Some([1.6, -0.5, -0.1, -0.2, 1.4, -0.2, 0.0, -0.4, 1.4]),
samples_per_pixel: 3,
profile: None,
profile_tables: None,
baseline_exposure: 0.0,
make: String::new(),
model: String::new(),
crop: CropRect {
+51 -2
View File
@@ -33,6 +33,8 @@ fn flat_raw(level: u16) -> RawImage {
color_matrix: Some([1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0]),
samples_per_pixel: 1,
profile: None,
profile_tables: None,
baseline_exposure: 0.0,
make: String::new(),
model: String::new(),
crop: CropRect {
@@ -44,6 +46,53 @@ fn flat_raw(level: u16) -> RawImage {
}
}
/// The default chain with D19's sigmoid chosen explicitly, so these tests
/// stay about the sigmoid whichever curve is the default (D21).
fn sigmoid_chain() -> EditGraph {
let mut g = EditGraph::default_chain();
g.set_param(
dr_pipeline::ops::view_transform::ID,
dr_pipeline::ops::view_transform::CURVE,
dr_pipeline::ops::view_transform::SIGMOID,
);
g
}
#[test]
fn camera_raw_tone_agrees_with_the_acr3_curve() {
// TRACES: FR-DEV-3j
// D21: a raw with no profile, the DNG reference curve chosen, renders a grey
// through the ACR3 default curve, which the profile buffer's placeholder
// carries.
let Some(ctx) = ctx() else {
eprintln!("skipping: no GPU adapter");
return;
};
let mut graph = EditGraph::default_chain();
graph.set_param(
dr_pipeline::ops::view_transform::ID,
dr_pipeline::ops::view_transform::CURVE,
dr_pipeline::ops::view_transform::CAMERA_RAW,
);
// The table itself, so its own contrast: the default bends the input a
// little past it (D21).
graph.set_param(
dr_pipeline::ops::view_transform::ID,
dr_pipeline::ops::view_transform::CONTRAST,
dr_pipeline::view::REFERENCE_CONTRAST,
);
for level in [500u16, 4_000, 8_520, 20_000, 40_000] {
let scene = f32::from(level) / f32::from(u16::MAX);
let display = dr_types::tone::evaluate(&dr_types::tone::ACR3_DEFAULT, scene);
let expected = (dr_types::Transfer::Srgb.encode(display) * 255.0).round() as i32;
let got = i32::from(rendered(&ctx, level, &graph));
assert!(
(got - expected).abs() <= 2,
"raw {level} rendered as {got}, the ACR3 curve says {expected}"
);
}
}
/// Render `graph` over a flat frame and return the centre pixel's red.
///
/// The centre rather than a corner: a demosaic has to invent its edges.
@@ -68,7 +117,7 @@ fn the_shader_agrees_with_the_cpu_reference() {
return;
};
let curve = Sigmoid::default_curve();
let graph = EditGraph::default_chain();
let graph = sigmoid_chain();
for level in [0u16, 500, 4_000, 8_520, 32_768, 60_000, u16::MAX] {
let scene = f32::from(level) / f32::from(u16::MAX);
let display = curve.channel(scene).min(1.0);
@@ -94,7 +143,7 @@ fn highlights_above_one_stay_distinct() {
eprintln!("skipping: no GPU adapter");
return;
};
let mut graph = EditGraph::default_chain();
let mut graph = sigmoid_chain();
graph.set_param(
dr_pipeline::ops::exposure::ID,
dr_pipeline::ops::exposure::EXPOSURE,
+151 -18
View File
@@ -24,6 +24,14 @@ enum Ep {
Cpu,
MiGraphX,
MiGraphXFp16,
OpenVinoCpu,
OpenVinoGpu,
OpenVinoGpuFp16,
OpenVinoNpu,
/// Dawn's low-power adapter: the integrated GPU on a hybrid machine.
WebGpuLow,
/// Dawn's high-performance adapter: the discrete one, if there is one.
WebGpuHigh,
}
impl Ep {
@@ -32,20 +40,113 @@ impl Ep {
Ep::Cpu => "CPU",
Ep::MiGraphX => "MIGraphX f32",
Ep::MiGraphXFp16 => "MIGraphX fp16",
Ep::OpenVinoCpu => "OpenVINO CPU",
Ep::OpenVinoGpu => "OpenVINO GPU",
Ep::OpenVinoGpuFp16 => "OpenVINO GPU16",
Ep::OpenVinoNpu => "OpenVINO NPU",
Ep::WebGpuLow => "WebGPU low",
Ep::WebGpuHigh => "WebGPU high",
}
}
/// Whether a second build reads what the first one compiled.
fn caches(self) -> bool {
matches!(
self,
Ep::MiGraphX
| Ep::MiGraphXFp16
| Ep::OpenVinoGpu
| Ep::OpenVinoGpuFp16
| Ep::OpenVinoNpu
)
}
}
fn build(ep: Ep, bytes: &[u8], threads: usize, cache: &Path) -> ort::Result<ort::session::Session> {
let mut b = ort::session::Session::builder()?.with_intra_threads(threads)?;
let dir = |sub: &str| {
let d = cache.join(sub);
let _ = std::fs::create_dir_all(&d);
d.to_string_lossy().into_owned()
};
// `GPU` is OpenVINO's first OpenCL GPU, which on a hybrid laptop can be
// the discrete NVIDIA one; DARKROOM_OV_GPU=GPU.1 names another.
let gpu = std::env::var("DARKROOM_OV_GPU").unwrap_or_else(|_| "GPU".into());
match ep {
Ep::Cpu => {}
Ep::MiGraphX => migraphx(&mut b, false, &cache.join("f32"))?,
Ep::MiGraphXFp16 => migraphx(&mut b, true, &cache.join("fp16"))?,
// Option names as `openvino_provider_factory.cc` reads them at 1.24.
Ep::OpenVinoCpu => append(&mut b, c"OpenVINO", &[("device_type", "CPU".into())])?,
Ep::OpenVinoGpu => append(
&mut b,
c"OpenVINO",
&[
("device_type", gpu.clone()),
("precision", "FP32".into()),
("cache_dir", dir("ov-gpu-f32")),
],
)?,
Ep::OpenVinoGpuFp16 => append(
&mut b,
c"OpenVINO",
&[
("device_type", gpu.clone()),
("precision", "FP16".into()),
("cache_dir", dir("ov-gpu-fp16")),
],
)?,
Ep::OpenVinoNpu => append(
&mut b,
c"OpenVINO",
&[("device_type", "NPU".into()), ("cache_dir", dir("ov-npu"))],
)?,
// `webgpu_provider_options.h` at 1.27; the runtime prefixes the key.
Ep::WebGpuLow => append(
&mut b,
c"WebGPU",
&[("powerPreference", "low-power".into())],
)?,
Ep::WebGpuHigh => append(
&mut b,
c"WebGPU",
&[("powerPreference", "high-performance".into())],
)?,
}
b.commit_from_memory(bytes)
}
/// Any provider through the generic key/value entry point.
fn append(
b: &mut ort::session::builder::SessionBuilder,
name: &std::ffi::CStr,
options: &[(&str, String)],
) -> ort::Result<()> {
use ort::AsPointer;
use std::ffi::CString;
let keys: Vec<CString> = options
.iter()
.map(|(k, _)| CString::new(*k).unwrap())
.collect();
let values: Vec<CString> = options
.iter()
.map(|(_, v)| CString::new(v.as_bytes()).unwrap())
.collect();
let key_ptrs: Vec<_> = keys.iter().map(|k| k.as_ptr()).collect();
let value_ptrs: Vec<_> = values.iter().map(|v| v.as_ptr()).collect();
// SAFETY: as `migraphx` below.
unsafe {
let status = (ort::api().SessionOptionsAppendExecutionProvider)(
b.ptr_mut(),
name.as_ptr(),
key_ptrs.as_ptr(),
value_ptrs.as_ptr(),
keys.len(),
);
ort::Error::result_from_status(status)
}
}
/// Register MIGraphX through the generic key/value API. `ort`'s own
/// builder fills the legacy `OrtMIGraphXProviderOptions`, which 1.29 reads
/// for its precision flags and nothing else: the model cache directory —
@@ -83,19 +184,29 @@ fn migraphx(
/// Median of `runs` timed runs over zeros, in milliseconds, after warm-ups.
fn time(session: &mut ort::session::Session, warmups: usize, runs: usize) -> Result<f64, String> {
let shape: Vec<usize> = session.inputs()[0]
.dtype()
.tensor_shape()
.ok_or("input is not a tensor")?
.iter()
.map(|&d| if d > 0 { d as usize } else { 1 })
.collect();
let zeros = vec![0f32; shape.iter().product()];
// Zeros for every input, not just the first: the denoiser takes
// `mosaic` and `sigma`. A dynamic dimension is read as 1.
let mut inputs = Vec::new();
for input in session.inputs() {
let shape: Vec<usize> = input
.dtype()
.tensor_shape()
.ok_or("input is not a tensor")?
.iter()
.map(|&d| if d > 0 { d as usize } else { 1 })
.collect();
let zeros = vec![0f32; shape.iter().product()];
inputs.push((input.name().to_string(), shape, zeros));
}
let once = |s: &mut ort::session::Session| -> Result<f64, String> {
let input = ort::value::Tensor::from_array((shape.clone(), zeros.clone()))
.map_err(|e| e.to_string())?;
let mut values = Vec::with_capacity(inputs.len());
for (name, shape, zeros) in &inputs {
let value = ort::value::Tensor::from_array((shape.clone(), zeros.clone()))
.map_err(|e| e.to_string())?;
values.push((name.clone(), ort::session::SessionInputValue::from(value)));
}
let t = Instant::now();
let out = s.run(ort::inputs![input]).map_err(|e| e.to_string())?;
let out = s.run(values).map_err(|e| e.to_string())?;
let _ = out[0]
.try_extract_tensor::<f32>()
.map_err(|e| e.to_string())?;
@@ -155,13 +266,35 @@ fn main() {
// A compiling provider is built twice: the second build reads the
// program the first wrote, and its time is what a launch after the
// first costs.
let plan = [
(Ep::Cpu, false),
(Ep::MiGraphX, false),
(Ep::MiGraphX, true),
(Ep::MiGraphXFp16, false),
(Ep::MiGraphXFp16, true),
];
// DARKROOM_EPS narrows the list (`cpu,openvino,webgpu,migraphx`);
// a runtime without a provider fails its build in a millisecond
// anyway, so the default is all of them.
let wanted = std::env::var("DARKROOM_EPS").unwrap_or_default();
let on = |family: &str| wanted.is_empty() || wanted.split(',').any(|w| w == family);
let mut plan = Vec::new();
for (family, eps) in [
("cpu", &[Ep::Cpu][..]),
("migraphx", &[Ep::MiGraphX, Ep::MiGraphXFp16][..]),
(
"openvino",
&[
Ep::OpenVinoCpu,
Ep::OpenVinoGpu,
Ep::OpenVinoGpuFp16,
Ep::OpenVinoNpu,
][..],
),
("webgpu", &[Ep::WebGpuLow, Ep::WebGpuHigh][..]),
] {
if on(family) {
for &ep in eps {
plan.push((ep, false));
if ep.caches() {
plan.push((ep, true));
}
}
}
}
for (ep, cached) in plan {
let started = Instant::now();
match build(ep, &bytes, threads, &cache) {
+49 -19
View File
@@ -4,12 +4,17 @@
//!
//! DARKROOM_ORT_DIR=/usr/lib \
//! cargo run --release -p dr-inference-engine --features native,tract \
//! --example ladder -- CACHE_DIR models/face/scrfd_500m_640.onnx [MODEL.onnx ...]
//! --example ladder -- CACHE_DIR models/face/scrfd_500m_640.onnx [ROLE=MODEL.onnx ...]
//!
//! Every model named is a `Detector` for the config's purposes, which is
//! enough to see the rung taken, the engines compiled and a session land
//! on it. Delete `CACHE_DIR` to see the first run again; keep it to see the
//! second.
//! `DARKROOM_ORT_DIRS=a:b:c` offers several runtimes, as the app's search
//! list does, and shows which the engine chose for this device's GPU.
//!
//! A bare path is a `Detector`; `denoiser=…`, `scene=…`, `inpainter=…`,
//! `landmarks=…` (any `Role`, lower case) says otherwise, so a device can
//! show each role taking its own form (inference.md §1.5). Each is opened
//! through `resolve_model`, as the app opens it, and the line says which
//! form and which rung it landed on. Delete `CACHE_DIR` to see the first
//! run again; keep it to see the second.
use std::path::PathBuf;
use std::time::{Duration, Instant};
@@ -17,7 +22,10 @@ use std::time::{Duration, Instant};
fn main() {
env_logger::Builder::from_env(env_logger::Env::default().default_filter_or("info")).init();
let mut args = std::env::args_os().skip(1).map(PathBuf::from);
let (Some(cache_dir), models) = (args.next(), args.collect::<Vec<_>>()) else {
let (Some(cache_dir), models) = (
args.next(),
args.map(|a| role_and_path(&a)).collect::<Vec<_>>(),
) else {
eprintln!("usage: ladder CACHE_DIR MODEL.onnx [MODEL.onnx ...]");
std::process::exit(2);
};
@@ -26,18 +34,22 @@ fn main() {
std::process::exit(2);
}
// DARKROOM_ORT_DIRS lists several, colon-separated, as the app's search
// does: the engine loads the one that fits the GPU (§3.2).
let runtime_dirs: Vec<PathBuf> = std::env::var_os("DARKROOM_ORT_DIR")
.map(PathBuf::from)
.into_iter()
.chain(
std::env::var_os("DARKROOM_ORT_DIRS")
.map(|v| std::env::split_paths(&v).collect::<Vec<_>>())
.unwrap_or_default(),
)
.collect();
let started = Instant::now();
dr_inference_engine::init(dr_inference_engine::Config {
runtime_dirs,
cache_dir: cache_dir.clone(),
models: models
.iter()
.map(|p| (dr_inference_engine::Role::Detector, p.clone()))
.collect(),
models: models.clone(),
embedded: Vec::new(),
ceiling: None,
threads: 0,
@@ -80,21 +92,39 @@ fn main() {
std::thread::sleep(Duration::from_millis(500));
}
for path in &models {
let bytes = std::fs::read(path).expect("read model");
for (role, path) in &models {
let (path, form) = dr_inference_engine::resolve_model(*role, path);
let bytes = std::fs::read(&path).expect("read model");
let t = Instant::now();
let model = dr_inference_engine::open(
dr_inference_engine::Role::Detector,
dr_inference_engine::Form::F32,
&bytes,
)
.expect("open model");
let model = dr_inference_engine::open(*role, form, &bytes).expect("open model");
let acquired = model.acquire().expect("acquire session");
println!(
"{} on {} in {:.2} s",
"{role:?}: {} ({form:?}) on {} in {:.2} s",
path.file_name().unwrap().to_string_lossy(),
acquired.rung().label(),
t.elapsed().as_secs_f64()
);
}
}
/// `denoiser=path` → (Denoiser, path); a bare path is a detector.
fn role_and_path(arg: &std::path::Path) -> (dr_inference_engine::Role, PathBuf) {
use dr_inference_engine::Role::*;
let s = arg.to_string_lossy();
let Some((name, path)) = s.split_once('=') else {
return (Detector, arg.to_path_buf());
};
let role = match name {
"detector" => Detector,
"embedder" => Embedder,
"segmenter" => Segmenter,
"scene" => Scene,
"landmarks" => Landmarks,
"eyes" => EyeClassifier,
"keypoints" => Keypoints,
"inpainter" => Inpainter,
"denoiser" => Denoiser,
other => panic!("no role {other:?}"),
};
(role, PathBuf::from(path))
}
+161 -28
View File
@@ -51,17 +51,25 @@ pub fn ensure_installed() {
}
}
/// Look for `libonnxruntime` in `dirs`, in order, and hand `ort` the first
/// table that loads; otherwise tract. Once per process.
/// Find every `libonnxruntime` in `dirs`, hand `ort` the table of the one
/// that best fits this device's GPUs, and fall to tract if none loads.
/// Once per process.
///
/// Best fit, not first found (§3.2): a device can hold several runtimes —
/// the package's OpenVINO build, a CUDA build the user fetched, the
/// distribution's ROCm build — and each carries one vendor's providers.
/// Between equals, the earlier directory wins, as it always has, and a
/// runtime that fits perfectly ends the search: the APK's QNN build on a
/// Qualcomm tablet is found first, and the generic build beside it is
/// never opened there.
/// `DARKROOM_ORT_DIR`, when it loads, wins outright: it is how a person
/// says which runtime they mean.
pub fn install(dirs: &[PathBuf]) -> Runtime {
RUNTIME
.get_or_init(|| {
#[cfg(feature = "native")]
for dir in dirs {
match load_native(dir) {
Ok(rt) => return rt,
Err(e) => log::info!("inference: no runtime in {}: {e}", dir.display()),
}
if let Some(rt) = install_best(dirs) {
return rt;
}
#[cfg(not(feature = "native"))]
let _ = dirs;
@@ -70,6 +78,103 @@ pub fn install(dirs: &[PathBuf]) -> Runtime {
.clone()
}
/// A runtime opened to read its providers, not yet handed to `ort`.
#[cfg(feature = "native")]
struct Found {
lib: libloading::Library,
api: *const ort_sys::OrtApi,
path: PathBuf,
version: String,
providers: Vec<String>,
}
#[cfg(feature = "native")]
fn install_best(dirs: &[PathBuf]) -> Option<Runtime> {
let named = std::env::var_os("DARKROOM_ORT_DIR").map(PathBuf::from);
let gpus = crate::hardware::detect();
let mut found: Vec<Found> = Vec::new();
let mut seen = std::collections::HashSet::new();
for dir in dirs {
match open_native(dir) {
Ok(f) => {
// `bin/../lib/darkroom` and `/usr/lib/darkroom` are one file.
if !seen.insert(std::fs::canonicalize(&f.path).unwrap_or(f.path.clone())) {
std::mem::forget(f.lib);
continue;
}
log::info!(
"inference: ONNX Runtime {} at {} offers {}",
f.version,
f.path.display(),
f.providers.join(", ")
);
if named.as_deref() == Some(dir.as_path()) {
found.clear();
found.push(f);
break;
}
let perfect = gpus.score(&f.providers) >= crate::hardware::PERFECT;
found.push(f);
if perfect {
break;
}
}
Err(e) => log::info!("inference: no runtime in {}: {e}", dir.display()),
}
}
let best = (0..found.len())
.max_by_key(|&i| (gpus.score(&found[i].providers), std::cmp::Reverse(i)))?;
let chosen = found.swap_remove(best);
// The others stay mapped. Unloading a C++ runtime after its static
// constructors ran is a crash at exit waiting to happen, and an
// unused mapping costs address space, not memory.
for other in found {
std::mem::forget(other.lib);
}
log::info!("inference: chose {} for {gpus:?}", chosen.path.display());
// SAFETY: the table came from this library's `OrtGetApiBase`, and the
// library is leaked below, so every pointer in the copy stays valid for
// the life of the process.
if !ort::set_api(unsafe { (*chosen.api).clone() }) {
log::warn!("inference: an API table was already installed");
std::mem::forget(chosen.lib);
return None;
}
std::mem::forget(chosen.lib);
// Qualcomm's DSP loader finds the Hexagon skel through this variable,
// and only through it; the runtime's own directory is where the APK
// put it. Harmless anywhere else.
#[cfg(target_os = "android")]
if let Some(dir) = chosen.path.parent().filter(|d| !d.as_os_str().is_empty()) {
std::env::set_var("ADSP_LIBRARY_PATH", dir);
}
// Windows looks for a provider's own dependencies — OpenVINO's DLLs,
// which Intel's build leaves beside it — on the DLL search path, not in
// the provider's directory. Intel's Python shim prepends to `PATH` for
// the same reason; so does this, before any provider loads.
#[cfg(target_os = "windows")]
if let Some(dir) = chosen.path.parent() {
let old = std::env::var_os("PATH").unwrap_or_default();
let dirs = std::iter::once(dir.to_path_buf()).chain(std::env::split_paths(&old));
if let Ok(path) = std::env::join_paths(dirs) {
std::env::set_var("PATH", path);
}
}
log::info!(
"inference: ONNX Runtime {} from {}",
chosen.version,
chosen.path.display()
);
Some(Runtime::OnnxRuntime {
path: chosen.path,
version: chosen.version,
})
}
#[cfg(feature = "tract")]
fn install_tract() -> Runtime {
let _ = ort::set_api(ort_tract::api());
@@ -85,8 +190,12 @@ fn install_tract() -> Runtime {
Runtime::Tract
}
/// Open the runtime in `dir` and read what it offers. `dir` may also name
/// the library itself — Android has two runtimes and one directory, so the
/// second goes by its file name — and an empty path is the bare name
/// through the system loader, which on Android is the APK's own copy.
#[cfg(feature = "native")]
fn load_native(dir: &std::path::Path) -> Result<Runtime, String> {
fn open_native(dir: &std::path::Path) -> Result<Found, String> {
let name = if cfg!(target_os = "windows") {
"onnxruntime.dll"
} else if cfg!(any(target_os = "macos", target_os = "ios")) {
@@ -94,18 +203,21 @@ fn load_native(dir: &std::path::Path) -> Result<Runtime, String> {
} else {
"libonnxruntime.so"
};
// An empty dir means the bare name: the system loader's search, which on
// Android includes the APK's own native libraries.
let is_library = dir.file_name().and_then(|n| n.to_str()).is_some_and(|n| {
n.contains("onnxruntime")
&& (n.ends_with(".so") || n.ends_with(".dll") || n.ends_with(".dylib"))
});
let path = if dir.as_os_str().is_empty() {
PathBuf::from(name)
} else if is_library {
dir.to_path_buf()
} else {
find_library(dir, name).ok_or("not present")?
};
// SAFETY: the library's initialisers are ONNX Runtime's own; the symbol
// is the documented entry point with the documented signature; the table
// is copied out and the library handle is leaked, so every pointer in
// the copy stays valid for the life of the process.
// is the documented entry point with the documented signature. The
// table pointer is valid while `lib` is, which the caller keeps.
unsafe {
let lib = libloading::Library::new(&path).map_err(|e| e.to_string())?;
let get_base: libloading::Symbol<
@@ -125,24 +237,45 @@ fn load_native(dir: &std::path::Path) -> Result<Runtime, String> {
ort_sys::ORT_API_VERSION
));
}
if !ort::set_api((*api).clone()) {
return Err("an API table was already installed".into());
}
std::mem::forget(lib);
// Qualcomm's DSP loader finds the Hexagon skel through this variable,
// and only through it; the runtime's own directory is where the APK
// put it. Harmless anywhere else.
#[cfg(target_os = "android")]
if !dir.as_os_str().is_empty() {
std::env::set_var("ADSP_LIBRARY_PATH", dir);
}
log::info!("inference: ONNX Runtime {version} from {}", path.display());
Ok(Runtime::OnnxRuntime { path, version })
let providers = available_providers(api);
Ok(Found {
lib,
api,
path,
version,
providers,
})
}
}
/// The providers compiled into the runtime behind `api` — not the ones this
/// device can run, which is the probe's question.
///
/// # Safety
/// `api` must be a live table from `GetApi`.
#[cfg(feature = "native")]
unsafe fn available_providers(api: *const ort_sys::OrtApi) -> Vec<String> {
let mut list: *mut *mut std::ffi::c_char = std::ptr::null_mut();
let mut n: std::ffi::c_int = 0;
let status = ((*api).GetAvailableProviders)(&mut list, &mut n);
if !status.0.is_null() {
((*api).ReleaseStatus)(status.0);
return Vec::new();
}
let names = (0..n.max(0) as usize)
.map(|i| {
std::ffi::CStr::from_ptr(*list.add(i))
.to_string_lossy()
.into_owned()
})
.collect();
let status = ((*api).ReleaseAvailableProviders)(list, n);
if !status.0.is_null() {
((*api).ReleaseStatus)(status.0);
}
names
}
/// `libonnxruntime.so` in `dir`, or a versioned spelling of it —
/// `libonnxruntime.so.1.30.0` is what the Python wheel ships, and a package
/// that installs only the versioned file is not wrong.
+41 -6
View File
@@ -8,7 +8,7 @@
use std::path::PathBuf;
use crate::{state, Config, Form, Rung};
use crate::{state, Config, Rung};
enum Source {
File(PathBuf),
@@ -48,12 +48,47 @@ pub fn context_path(cfg: &Config, bytes: &[u8]) -> PathBuf {
/// CoreML's own cache key leaves out the weights of a model loaded from
/// memory (`session::coreml`), and one per runtime version, which wrote it.
pub fn coreml_dir(cfg: &Config, bytes: &[u8]) -> PathBuf {
model_dir(cfg, "coreml", bytes)
}
/// Where OpenVINO compiles `bytes` to, at one precision. OpenVINO hashes
/// the model it is given, weights included, but a key that leaves out
/// what is being varied has cost a day before (CLAUDE.md, "Providers"),
/// and a directory per model and precision costs nothing: the precision
/// is a compile option, and the two forms are different programs.
pub fn openvino_dir(cfg: &Config, bytes: &[u8], fp16: bool) -> PathBuf {
model_dir(
cfg,
if fp16 {
"openvino/fp16"
} else {
"openvino/f32"
},
bytes,
)
}
/// Where TensorRT keeps the engine for a whole-frame model. Its own
/// directory per model: ONNX Runtime's engine cache key leaves the input
/// shape out, and served one export's engine to another of the same graph
/// with a different shape when the denoiser was first cut into pieces
/// (2026-10-04) — the fixed 1408² denoiser and its any-size sibling are
/// exactly that pair. The profile's largest shape is in the name for the
/// same reason: an engine built for one range is not the next one's.
pub fn tensorrt_whole_dir(cfg: &Config, bytes: &[u8]) -> PathBuf {
let (h, w) = crate::WHOLE_FRAME_MAX;
model_dir(cfg, &format!("tensorrt-whole-{h}x{w}"), bytes)
}
/// `<cache>/<provider>/<runtime version>/<hash of the bytes>`: one per
/// model, and one per runtime version, which wrote it.
fn model_dir(cfg: &Config, provider: &str, bytes: &[u8]) -> PathBuf {
let runtime = match crate::api::runtime() {
crate::Runtime::OnnxRuntime { version, .. } => version,
crate::Runtime::Tract => "tract".into(),
};
cfg.cache_dir
.join("coreml")
.join(provider)
.join(runtime)
.join(format!("{:016x}", hash(bytes)))
}
@@ -82,10 +117,10 @@ pub fn run() {
(*role, Source::File(path), size)
})
})
.chain(cfg.embedded.iter().filter_map(|(role, bytes)| {
// An embedded model has no int8 sibling to offer a rung that
// wants one; it runs on that rung's fallback.
(rung.serves(*role) && rung.form(*role) == Form::F32).then_some((
.chain(cfg.embedded.iter().filter_map(|(role, form, bytes)| {
// The embedded form the rung wants, if the build carries it;
// a build without it runs that model on the rung's fallback.
(rung.serves(*role) && rung.form(*role) == *form).then_some((
*role,
Source::Bytes(bytes),
bytes.len() as u64,
+176
View File
@@ -0,0 +1,176 @@
//! Which GPUs this device has, as far as choosing a runtime needs to know
//! (docs/dev/inference.md §3.2).
//!
//! A runtime carries one vendor's providers — Intel's build has OpenVINO,
//! the `onnxruntime-gpu` wheel CUDA and TensorRT, a ROCm build MIGraphX,
//! Microsoft's WebGPU build the generic rung — and only one runtime loads
//! per process. These checks are what lets `api` load the one that fits
//! when a device has several installed. They read files, never a driver:
//! a wrong answer costs a slower rung, which the probe still measures, and
//! a driver call at start-up could cost the launch.
/// What a runtime's providers are scored against.
#[derive(Clone, Copy, Debug, Default, PartialEq, Eq)]
pub struct Gpus {
pub nvidia: bool,
/// An AMD GPU with the ROCm kernel interface, which MIGraphX needs.
pub amd_rocm: bool,
pub intel: bool,
pub qualcomm: bool,
}
/// The score of a runtime whose vendor rung matches the device's GPU.
/// Nothing beats it, so the search stops there.
pub const PERFECT: u32 = 3;
impl Gpus {
/// How well a runtime offering `providers` fits this device. The vendor
/// rungs score above OpenVINO because a machine with an Intel iGPU and
/// an NVIDIA or AMD card wants the card; the generic rung scores above
/// a CPU-only build because it carries the same CPU provider and might
/// beat it.
pub fn score(&self, providers: &[String]) -> u32 {
providers
.iter()
.map(|p| match p.as_str() {
"TensorrtExecutionProvider" | "CUDAExecutionProvider" if self.nvidia => PERFECT,
"MIGraphXExecutionProvider" if self.amd_rocm => PERFECT,
"QNNExecutionProvider" if self.qualcomm => PERFECT,
"CoreMLExecutionProvider" => PERFECT,
"OpenVINOExecutionProvider" if self.intel => 2,
"WebGpuExecutionProvider" => 1,
_ => 0,
})
.max()
.unwrap_or(0)
}
}
#[cfg(target_os = "linux")]
pub fn detect() -> Gpus {
use std::path::Path;
// Every DRM card's PCI vendor: an Intel iGPU is `0x8086` whether or
// not its compute driver is installed, which the probe finds out.
let vendors: Vec<String> = std::fs::read_dir("/sys/class/drm")
.into_iter()
.flatten()
.filter_map(|e| e.ok())
.filter(|e| {
let name = e.file_name();
let name = name.to_string_lossy();
name.starts_with("card") && !name.contains('-')
})
.filter_map(|e| std::fs::read_to_string(e.path().join("device/vendor")).ok())
.map(|v| v.trim().to_string())
.collect();
Gpus {
nvidia: Path::new("/proc/driver/nvidia/version").exists(),
amd_rocm: Path::new("/dev/kfd").exists(),
intel: vendors.iter().any(|v| v == "0x8086"),
qualcomm: false,
}
}
#[cfg(target_os = "windows")]
pub fn detect() -> Gpus {
use std::path::PathBuf;
let root = std::env::var_os("SystemRoot")
.map(PathBuf::from)
.unwrap_or_else(|| PathBuf::from(r"C:\Windows"));
let system32 = root.join("System32");
// Intel's DCH graphics driver, integrated and Arc alike, installs
// from `iigd_dch.inf`; its package directory is the evidence.
let intel = std::fs::read_dir(system32.join(r"DriverStore\FileRepository"))
.into_iter()
.flatten()
.filter_map(|e| e.ok())
.any(|e| e.file_name().to_string_lossy().starts_with("iigd_dch"));
Gpus {
nvidia: system32.join("nvcuda.dll").exists(),
amd_rocm: false,
intel,
qualcomm: false,
}
}
#[cfg(target_os = "android")]
pub fn detect() -> Gpus {
// Fail-safe: only a device that names another vendor is not Qualcomm.
// `ro.soc.manufacturer` exists from Android 12, and a property or file
// the app cannot read reads as nothing; nothing keeps the QNN build
// first, as 0.22 had it, where a Qualcomm device mistaken for another
// would trade its NPU for the generic rung. Qualcomm's FastRPC library,
// which the Hexagon path loads anyway, overrules a name.
let soc = crate::probe::system_property("ro.soc.manufacturer");
let fastrpc = [
"/vendor/lib64/libcdsprpc.so",
"/system/vendor/lib64/libcdsprpc.so",
]
.iter()
.any(|p| std::path::Path::new(p).exists());
Gpus {
qualcomm: qualcomm_soc(&soc) || fastrpc,
..Gpus::default()
}
}
/// Whether `ro.soc.manufacturer` leaves the device Qualcomm's: it says so,
/// or it says nothing.
#[cfg(any(target_os = "android", test))]
fn qualcomm_soc(manufacturer: &str) -> bool {
let m = manufacturer.trim();
m.is_empty() || m.eq_ignore_ascii_case("QTI") || m.eq_ignore_ascii_case("Qualcomm")
}
#[cfg(not(any(target_os = "linux", target_os = "windows", target_os = "android")))]
pub fn detect() -> Gpus {
Gpus::default()
}
#[cfg(test)]
mod tests {
use super::*;
fn offers(p: &[&str]) -> Vec<String> {
p.iter().map(|s| s.to_string()).collect()
}
#[test]
fn only_a_named_other_vendor_is_not_qualcomm() {
assert!(qualcomm_soc("QTI"));
assert!(qualcomm_soc("Qualcomm"));
// Unreadable, or older than Android 12: the QNN build stays first.
assert!(qualcomm_soc(""));
assert!(!qualcomm_soc("Mediatek"));
assert!(!qualcomm_soc("Google"));
assert!(!qualcomm_soc("Samsung"));
}
#[test]
fn the_card_beats_the_integrated_gpu_and_both_beat_the_generic_rung() {
let cpu = offers(&["CPUExecutionProvider"]);
let nvidia = offers(&[
"TensorrtExecutionProvider",
"CUDAExecutionProvider",
"CPUExecutionProvider",
]);
let intel = offers(&["OpenVINOExecutionProvider", "CPUExecutionProvider"]);
let webgpu = offers(&["WebGpuExecutionProvider", "CPUExecutionProvider"]);
let laptop = Gpus {
nvidia: true,
intel: true,
..Gpus::default()
};
assert!(laptop.score(&nvidia) > laptop.score(&intel));
assert!(laptop.score(&intel) > laptop.score(&webgpu));
assert!(laptop.score(&webgpu) > laptop.score(&cpu));
// No Intel GPU: Intel's build is worth no more than a CPU build to
// this device, and the generic rung is worth more.
let amd_on_windows = Gpus::default();
assert_eq!(amd_on_windows.score(&intel), amd_on_windows.score(&cpu));
assert!(amd_on_windows.score(&webgpu) > amd_on_windows.score(&intel));
// A ROCm build on a machine without ROCm is a CPU build.
let rocm = offers(&["MIGraphXExecutionProvider", "CPUExecutionProvider"]);
assert_eq!(amd_on_windows.score(&rocm), 0);
}
}
+237 -58
View File
@@ -21,6 +21,9 @@ use serde::{Deserialize, Serialize};
mod api;
mod engines;
// Read only when a runtime is loaded from disk (`api::install_best`).
#[cfg_attr(not(feature = "native"), allow(dead_code))]
mod hardware;
mod probe;
mod session;
@@ -42,20 +45,76 @@ pub enum Role {
/// XFeat, the panorama keypoint detector (docs/dev/panorama.md).
Keypoints,
/// MI-GAN, the panorama border filler (docs/dev/panorama.md §12). Plain
/// convolutions, so any rung serves it; fp16 on TensorRT and int8 on
/// the Hexagon are the point of it.
/// convolutions, so any rung serves it; fp16 on TensorRT and 16-bit
/// activations on the Hexagon (int8 changes the fill, §1.5).
Inpainter,
/// The learned demosaic and denoise on the raw mosaic (docs/dev/denoise.md).
/// fp16 costs it nothing measurable; int8 costs 6–9 dB, because 256
/// levels cannot hold the shadow steps it exists to recover — so the
/// Hexagon takes it with 16-bit activations and weights (§1.5).
Denoiser,
/// The same denoise networks exported with any height and width, run
/// over a whole frame — or the fewest large tiles that fit — instead of
/// 1408² tiles whose borders are thrown away (docs/dev/denoise.md §14).
/// Served only where a size the graph was not compiled for costs
/// nothing: TensorRT, through an optimisation profile up to
/// [`WHOLE_FRAME_MAX`], and the CUDA provider. Everywhere else the
/// fixed-tile [`Role::Denoiser`] runs; see [`whole_frame_limit`].
WholeDenoiser,
}
/// The largest input, rows × columns, a [`Role::WholeDenoiser`] session
/// takes: TensorRT's optimisation profile is built up to it, and the tiler
/// cuts a larger frame into the fewest tiles no bigger.
///
/// Sized for a 6 GB card. TensorRT plans its memory for the profile's
/// largest shape, and at 4608 × 6656 (a whole 6D frame with Best's border
/// and room to spare) it asked for 4.9–5.9 GB and could not build on the
/// RTX 3050. At 15 MP a 6D frame is two tiles of 4160 × 3248: 27 MP of
/// work for 20 MP kept, against 49 MP in 1408² tiles.
pub const WHOLE_FRAME_MAX: (usize, usize) = (4608, 3328);
/// The input size TensorRT tunes a whole-frame engine for: half a 6D frame
/// with Best's border, the tile the reference measurements run.
pub const WHOLE_FRAME_OPT: (usize, usize) = (4160, 3248);
/// Whether the selected rung runs [`Role::WholeDenoiser`], and if so the
/// largest input it takes. `None` means run the fixed tiles.
pub fn whole_frame_limit() -> Option<(usize, usize)> {
let rung = current_rung(&state().lock().unwrap());
rung.serves(Role::WholeDenoiser).then_some(WHOLE_FRAME_MAX)
}
/// Which numeric form of a model a session was built from.
///
/// `Int8` is a different network from `F32` for a detector — it finds a
/// different set of faces — which is why [`form_suffix`] exists and why a
/// The quantised forms are QDQ graphs, per-channel weights, as QNN's HTP
/// takes them (docs/dev/inference.md §1.5): `Int8` is 8-bit activations and
/// weights, `A16W8` 16-bit activations with 8-bit weights, `A16W16` 16-bit
/// both. The Hexagon accepts no float tensor at all, so these are the
/// whole menu; which one a role gets is [`Rung::form`], measured per model.
///
/// A quantised detector is a different network from the f32 one — it finds
/// a different set of faces — which is why [`form_suffix`] exists and why a
/// caller appends it to `model_id`.
#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash, Serialize, Deserialize)]
pub enum Form {
F32,
Int8,
A16W8,
A16W16,
}
impl Form {
/// The infix of the sibling file that holds this form:
/// `scrfd_500m_640.a16w8.onnx` beside `scrfd_500m_640.onnx`.
pub fn file_tag(self) -> Option<&'static str> {
match self {
Form::F32 => None,
Form::Int8 => Some("int8"),
Form::A16W8 => Some("a16w8"),
Form::A16W16 => Some("a16w16"),
}
}
}
/// A rung of the ladder (§2). Ordered: a user override names the highest rung
@@ -76,7 +135,7 @@ pub enum Rung {
/// removed in ONNX Runtime 1.23, so there is no non-compiling AMD rung
/// to fall back to: this one falls back to the CPU.
MiGraphX,
/// Qualcomm's Hexagon NPU through QNN, int8 models only. Android only.
/// Qualcomm's Hexagon NPU through QNN, quantised models only. Android only.
Hexagon,
/// Apple, through CoreML: the Neural Engine, the GPU or the CPU, as
/// CoreML schedules it. macOS only. Compiles an ML Program per model on
@@ -84,6 +143,16 @@ pub enum Rung {
/// embedder stays on the CPU, as on the Hexagon: the Neural Engine
/// computes in fp16 (§7).
CoreMl,
/// Intel, through OpenVINO on the integrated or Arc GPU. Desktop only.
/// Compiles a program per model, as MIGraphX does, so the CPU is its
/// fallback; fp16 on the same terms as TensorRT (§7).
OpenVino,
/// Any other GPU, through ONNX Runtime's WebGPU provider: Dawn on
/// Vulkan, D3D12 or Metal. The generic rung, for a GPU no vendor rung
/// covers. Measured slower than the CPU on every GPU it has been timed
/// on (§1), so it is on the ladder for the GPUs it has not, and the
/// probe's clock is what keeps it off the rest.
WebGpu,
}
impl Rung {
@@ -95,6 +164,8 @@ impl Rung {
Rung::MiGraphX => "MIGraphX",
Rung::Hexagon => "Hexagon NPU",
Rung::CoreMl => "CoreML",
Rung::OpenVino => "OpenVINO",
Rung::WebGpu => "WebGPU",
}
}
@@ -103,7 +174,13 @@ impl Rung {
fn fallback(self) -> Rung {
match self {
Rung::TensorRt => Rung::Cuda,
Rung::MiGraphX | Rung::Hexagon | Rung::CoreMl | Rung::Cuda | Rung::Cpu => Rung::Cpu,
Rung::MiGraphX
| Rung::Hexagon
| Rung::CoreMl
| Rung::OpenVino
| Rung::WebGpu
| Rung::Cuda
| Rung::Cpu => Rung::Cpu,
}
}
@@ -111,26 +188,52 @@ impl Rung {
fn compiles(self) -> bool {
matches!(
self,
Rung::TensorRt | Rung::MiGraphX | Rung::Hexagon | Rung::CoreMl
Rung::TensorRt | Rung::MiGraphX | Rung::Hexagon | Rung::CoreMl | Rung::OpenVino
)
}
/// The model form this rung wants for a role.
fn form(self, _role: Role) -> Form {
///
/// On the Hexagon, the narrowest form that held each model's accuracy
/// on the tablet itself (§1.5): int8 lost 5% of the detector's faces at
/// 40–80 px, moved the landmarks by 1.5 px and the segmenter's scores
/// to nothing, and the denoiser by 6–9 dB, so those take 16-bit
/// activations; the segmenter, scene model, filler and denoiser also
/// needed 16-bit weights. Only XFeat keeps int8: its panorama alignment
/// moved by no more than f32's own refits do.
pub fn form(self, role: Role) -> Form {
match self {
Rung::Hexagon => Form::Int8,
Rung::Hexagon => match role {
Role::Keypoints => Form::Int8,
Role::Detector | Role::Landmarks => Form::A16W8,
Role::Segmenter | Role::Scene | Role::Inpainter | Role::Denoiser => Form::A16W16,
// Not served there at all: the Hexagon takes fixed shapes.
Role::Embedder | Role::EyeClassifier | Role::WholeDenoiser => Form::F32,
},
_ => Form::F32,
}
}
/// Whether this rung runs `role` at all. The Hexagon takes int8 graphs
/// only, and the embedder is never int8 (§7) — it runs on the CPU
/// beside a detector on the NPU, so its vectors compare across devices.
/// CoreML is kept off the embedder for the same reason: the Neural
/// Engine is fp16, and which unit runs a graph is CoreML's choice.
/// Whether this rung runs `role` at all. The Hexagon takes quantised
/// graphs only, and the embedder is never quantised (§7) — it runs on
/// the CPU beside a detector on the NPU, so its vectors compare across
/// devices; at A16W16 it still missed the 0.999 cosine gate. The eye
/// classifiers stay on the CPU too: a millisecond there, and the two
/// share one role while only one of them held its readings quantised.
/// CoreML is kept off the embedder for the same reason as the Hexagon:
/// the Neural Engine is fp16, and which unit runs a graph is CoreML's
/// choice.
fn serves(self, role: Role) -> bool {
// Any input size only where a new size costs nothing. MIGraphX,
// OpenVINO and CoreML compile per shape, the Hexagon takes fixed
// shapes only, and the CPU could but would hold gigabytes of f32
// activations for a whole frame of Best.
if role == Role::WholeDenoiser {
return matches!(self, Rung::TensorRt | Rung::Cuda);
}
match self {
Rung::Hexagon | Rung::CoreMl => role != Role::Embedder,
Rung::Hexagon => !matches!(role, Role::Embedder | Role::EyeClassifier),
Rung::CoreMl => role != Role::Embedder,
_ => true,
}
}
@@ -153,8 +256,10 @@ pub struct Config {
/// The canonical model files on this device, so engines can be compiled
/// ahead of the first request for them.
pub models: Vec<(Role, PathBuf)>,
/// Models compiled into the binary, for the same reason.
pub embedded: Vec<(Role, &'static [u8])>,
/// Models compiled into the binary, for the same reason, each with the
/// form it is. A build that embeds a quantised sibling lists it here
/// beside the f32 graph, and the compile step takes the one the rung wants.
pub embedded: Vec<(Role, Form, &'static [u8])>,
/// The highest rung the user allows; `None` is "the best that works".
pub ceiling: Option<Rung>,
/// ONNX Runtime's intra-op pool; 0 picks from the core count.
@@ -180,11 +285,11 @@ pub struct Status {
}
impl Status {
/// "Hexagon NPU · int8 · ONNX Runtime 1.29" — the settings row's text.
/// "Hexagon NPU · quantised · ONNX Runtime 1.29" — the settings row's text.
pub fn line(&self) -> String {
let form = match self.rung {
Rung::Hexagon => " · int8",
Rung::TensorRt | Rung::MiGraphX => " · fp16",
Rung::Hexagon => " · quantised",
Rung::TensorRt | Rung::MiGraphX | Rung::OpenVino => " · fp16",
_ => "",
};
format!("{}{} · {}", self.rung.label(), form, self.runtime.label())
@@ -366,6 +471,12 @@ struct Cache {
/// still reads.
#[serde(default)]
refused: BTreeSet<String>,
/// Probes run under this fingerprint (`probe::run`): a fall-back to the
/// CPU is re-probed until there have been `RETRIES`. Defaulted, so a
/// cache from 0.22.0 or before — which may hold exactly such a verdict —
/// probes again.
#[serde(default)]
attempts: u32,
}
struct State {
@@ -456,26 +567,44 @@ fn current_rung(s: &State) -> Rung {
/// The file to load for `role` under the current selection, and its form.
///
/// A rung that wants int8 gets the `.int8.onnx` sibling of the canonical file
/// if it exists; otherwise the canonical file, on the rung's fallback. A
/// caller adds [`form_suffix`] to the `model_id` it records.
/// A rung that wants a quantised form gets that sibling of the canonical
/// file (`<stem>.a16w8.onnx` and so on, [`Form::file_tag`]) if it exists;
/// otherwise the canonical file, on the rung's fallback. A caller adds
/// [`form_suffix`] to the `model_id` it records.
pub fn resolve_model(role: Role, canonical: &Path) -> (PathBuf, Form) {
let rung = current_rung(&state().lock().unwrap());
if rung.serves(role) && rung.form(role) == Form::Int8 {
let sibling = int8_sibling(canonical);
let want = rung.form(role);
if rung.serves(role) && want != Form::F32 {
let sibling = form_sibling(canonical, want);
if sibling.is_file() {
return (sibling, Form::Int8);
return (sibling, want);
}
}
(canonical.to_path_buf(), Form::F32)
}
fn int8_sibling(canonical: &Path) -> PathBuf {
/// The same choice for a model compiled into the binary: of the forms
/// `offered`, the one the current rung wants for `role`, else the f32 one.
/// `offered` must hold an `F32` entry.
pub fn choose_embedded(role: Role, offered: &[(Form, &'static [u8])]) -> (&'static [u8], Form) {
let rung = current_rung(&state().lock().unwrap());
let want = rung.form(role);
let pick = |form| offered.iter().find(|(f, _)| *f == form);
let (form, bytes) = (rung.serves(role).then(|| pick(want)).flatten())
.or_else(|| pick(Form::F32))
.expect("an embedded model offers its f32 form");
(bytes, *form)
}
fn form_sibling(canonical: &Path, form: Form) -> PathBuf {
let stem = canonical
.file_stem()
.map(|s| s.to_string_lossy().into_owned())
.unwrap_or_default();
canonical.with_file_name(format!("{stem}.int8.onnx"))
match form.file_tag() {
Some(tag) => canonical.with_file_name(format!("{stem}.{tag}.onnx")),
None => canonical.to_path_buf(),
}
}
/// What a form appends to a detector's `model_id` (§7).
@@ -483,6 +612,8 @@ pub fn form_suffix(form: Form) -> &'static str {
match form {
Form::F32 => "",
Form::Int8 => "_i8",
Form::A16W8 => "_a16",
Form::A16W16 => "_a16w16",
}
}
@@ -512,8 +643,8 @@ pub fn open(role: Role, form: Form, bytes: &[u8]) -> Result<Model, Error> {
fn effective_rung(s: &State, selected: Rung, role: Role, form: Form, hash: u64) -> Rung {
let mut rung = selected;
if !rung.serves(role) || rung.form(role) != form {
// The embedder on a Hexagon device, or an f32 detector where the int8
// sibling was missing: neither can go to the NPU.
// The embedder on a Hexagon device, or an f32 detector where the
// quantised sibling was missing: neither can go to the NPU.
rung = rung.fallback();
}
if rung.compiles() && !s.cache.compiled.contains(&engines::key_of(rung, hash)) {
@@ -604,8 +735,10 @@ mod tests {
#[test]
fn the_hexagon_never_takes_the_embedder() {
assert!(!Rung::Hexagon.serves(Role::Embedder));
assert!(!Rung::Hexagon.serves(Role::EyeClassifier));
assert!(Rung::Hexagon.serves(Role::Detector));
assert_eq!(Rung::Hexagon.form(Role::Detector), Form::Int8);
assert!(Rung::Hexagon.serves(Role::Denoiser));
assert_eq!(Rung::Hexagon.form(Role::Detector), Form::A16W8);
// A detector offered in f32 on a Hexagon device lands on the CPU.
let s = State {
config: Config::default(),
@@ -616,36 +749,59 @@ mod tests {
probing: false,
wanted: 0,
};
let on = |role, form| effective_rung(&s, Rung::Hexagon, role, form, engines::hash(b""));
assert_eq!(on(Role::Embedder, Form::F32), Rung::Cpu);
assert_eq!(on(Role::Detector, Form::F32), Rung::Cpu);
// A form other than the one the role wants is not the NPU's either:
// an int8 detector left over from an older install stays off it.
assert_eq!(on(Role::Detector, Form::Int8), Rung::Cpu);
// The wanted form whose context is not compiled yet: also the CPU.
assert_eq!(on(Role::Detector, Form::A16W8), Rung::Cpu);
}
/// The form each role gets on the Hexagon is the one measured to hold
/// its accuracy there (§1.5); a change to this table is a change to
/// what the tablet computes, and must come with a measurement.
#[test]
fn each_role_has_its_measured_form_on_the_hexagon() {
use Form::*;
for (role, form) in [
(Role::Detector, A16W8),
(Role::Landmarks, A16W8),
(Role::Segmenter, A16W16),
(Role::Scene, A16W16),
(Role::Inpainter, A16W16),
(Role::Denoiser, A16W16),
(Role::Keypoints, Int8),
(Role::Embedder, F32),
(Role::EyeClassifier, F32),
] {
assert_eq!(Rung::Hexagon.form(role), form, "{role:?}");
}
for rung in [
Rung::Cpu,
Rung::Cuda,
Rung::TensorRt,
Rung::MiGraphX,
Rung::CoreMl,
Rung::OpenVino,
Rung::WebGpu,
] {
assert_eq!(rung.form(Role::Detector), F32);
}
}
#[test]
fn a_form_lives_in_its_tagged_sibling() {
let canonical = Path::new("/m/scrfd_500m_640.onnx");
assert_eq!(form_sibling(canonical, Form::F32), canonical);
assert_eq!(
effective_rung(
&s,
Rung::Hexagon,
Role::Embedder,
Form::F32,
engines::hash(b"")
),
Rung::Cpu
form_sibling(canonical, Form::A16W8),
Path::new("/m/scrfd_500m_640.a16w8.onnx")
);
assert_eq!(
effective_rung(
&s,
Rung::Hexagon,
Role::Detector,
Form::F32,
engines::hash(b"")
),
Rung::Cpu
);
// An int8 detector whose context is not compiled yet: also the CPU.
assert_eq!(
effective_rung(
&s,
Rung::Hexagon,
Role::Detector,
Form::Int8,
engines::hash(b"")
),
Rung::Cpu
form_sibling(canonical, Form::Int8),
Path::new("/m/scrfd_500m_640.int8.onnx")
);
}
@@ -670,6 +826,29 @@ mod tests {
assert_eq!(on(&s, Role::Embedder), Rung::Cpu);
}
/// OpenVINO compiles a program per model, so a request waits on the CPU
/// until the engine thread has built it; WebGPU builds in the session
/// and serves at once. Both take every role in f32 graphs.
#[test]
fn openvino_waits_for_its_program_and_webgpu_does_not() {
let hash = engines::hash(b"detector");
let mut s = State {
config: Config::default(),
cache: Cache::default(),
probing: false,
wanted: 0,
};
let on = |s: &State, rung| effective_rung(s, rung, Role::Detector, Form::F32, hash);
assert_eq!(on(&s, Rung::OpenVino), Rung::Cpu);
s.cache
.compiled
.insert(engines::key_of(Rung::OpenVino, hash));
assert_eq!(on(&s, Rung::OpenVino), Rung::OpenVino);
assert_eq!(on(&s, Rung::WebGpu), Rung::WebGpu);
let embedder = effective_rung(&s, Rung::WebGpu, Role::Embedder, Form::F32, hash);
assert_eq!(embedder, Rung::WebGpu);
}
#[test]
fn the_status_reports_only_the_rungs_above_the_selection() {
let _serial = serial();
+191 -60
View File
@@ -14,23 +14,64 @@ use crate::{api::Runtime, state, Cache, Config, Form, Role, Rung};
/// The rungs to try on this platform, best first, under the user's ceiling.
fn ladder(ceiling: Option<Rung>) -> Vec<Rung> {
// WebGPU is the generic rung (§2): it is reached only on a runtime that
// carries it, which `api` loads where no vendor's runtime fits the
// device, and kept only where it beats the CPU.
#[cfg(target_os = "android")]
let all = [Rung::Hexagon];
let all = [Rung::Hexagon, Rung::WebGpu];
// Unmeasured (§2 ⁵): it is on the ladder because the probe's clock and
// `attempt` make a wrong guess cost one slow or failed probe, not a
// slow or crashing app.
#[cfg(target_os = "macos")]
let all = [Rung::CoreMl];
// A desktop has one vendor's GPU; the other vendor's providers are
// "not enabled in this build" or a library that fails to load, and
// either answer arrives in milliseconds.
// A runtime carries one vendor's providers, chosen for this device's
// GPU (`api`); the others are "not enabled in this build", and that
// answer arrives in milliseconds.
#[cfg(not(any(target_os = "android", target_os = "macos")))]
let all = [Rung::TensorRt, Rung::Cuda, Rung::MiGraphX];
let all = [
Rung::TensorRt,
Rung::Cuda,
Rung::MiGraphX,
Rung::OpenVino,
Rung::WebGpu,
];
all.into_iter()
.filter(|r| ceiling.is_none_or(|c| *r <= c))
.collect()
}
/// Probes under one fingerprint that may end on the CPU after an
/// accelerator failed or lost, before that answer is kept.
const RETRIES: u32 = 3;
/// What a cached probe result is good for.
#[derive(Debug, PartialEq)]
enum Reuse {
/// Use it as it is.
Keep,
/// Probe again: it fell back to the CPU after this many probes.
Again(u32),
/// Another device, runtime or model set: probe from the start.
Fresh,
}
/// The CPU because an accelerator failed or lost is asked again on the next
/// launches, a few times: a failure can be a moment's (QNN could not create
/// its device on 0.22.0's first launch after the update), and keeping it for
/// good left the tablet's every model on the CPU. Bounded, so a wedged
/// driver costs a few launches, not all.
fn reuse(cached: &Cache, fingerprint: &str) -> Reuse {
if cached.fingerprint != fingerprint || cached.rung.is_none() {
return Reuse::Fresh;
}
let fell_back = cached.rung == Some(Rung::Cpu) && !cached.failed.is_empty();
if fell_back && cached.attempts < RETRIES {
Reuse::Again(cached.attempts)
} else {
Reuse::Keep
}
}
/// The probe body. Sets the cache and clears `probing` when done; never
/// panics out, because a failed probe is a result (the floor) and not an
/// error.
@@ -38,20 +79,33 @@ pub fn run(runtime: Runtime) {
let cfg = state().lock().unwrap().config.clone();
let fingerprint = fingerprint(&runtime, &cfg);
let mut attempts = 0;
if let Some(cached) = read_cache(&cfg) {
if cached.fingerprint == fingerprint && cached.rung.is_some() {
log::info!(
"inference: cached selection {} ({})",
cached.rung.unwrap().label(),
cached.reason
);
finish(cached);
return;
match reuse(&cached, &fingerprint) {
Reuse::Keep => {
log::info!(
"inference: cached selection {} ({})",
cached.rung.map_or("?", |r| r.label()),
cached.reason
);
finish(cached);
return;
}
Reuse::Again(n) => {
attempts = n;
log::info!(
"inference: probing again after falling back to the CPU ({}), attempt {} of {RETRIES}",
cached.reason,
n + 1
);
}
Reuse::Fresh => {}
}
}
let mut cache = Cache {
fingerprint,
attempts: attempts + 1,
..Cache::default()
};
@@ -112,7 +166,16 @@ pub fn run(runtime: Runtime) {
}
if cache.rung.is_none() {
cache.rung = Some(Rung::Cpu);
cache.reason = match cache.failed.first() {
// The rung that tried and lost, not the first one the runtime was
// never built with: "WebGPU 150 ms, slower than the CPU" says why
// this device is on the CPU, "TensorRT not enabled" does not.
let tried = cache
.failed
.iter()
.rev()
.find(|(_, why)| !why.contains("in this build"))
.or(cache.failed.first());
cache.reason = match tried {
Some((r, why)) => format!("{} {}", r.label(), first_line(why)),
None => "the only rung on this platform".into(),
};
@@ -174,10 +237,9 @@ pub fn attempt<T>(cfg: &Config, what: &str, f: impl FnOnce() -> T) -> Result<T,
}
/// The smallest detector, or the smallest model of any role if there is
/// none. A ~2 MB detector is the cheapest real test of a provider, and the
/// detector is the role the int8 forms exist for — the eye classifiers are
/// smaller still, and a Hexagon probed with one would fail for want of a
/// form nobody ships.
/// none. A ~2 MB detector is the cheapest real test of a provider, and
/// every rung serves it — the eye classifiers are smaller still, but the
/// Hexagon does not take them, and a probe with one would fail it for that.
fn probe_model(cfg: &Config) -> Option<(Role, PathBuf)> {
let smallest = |want: Option<Role>| {
cfg.models
@@ -193,9 +255,14 @@ fn probe_model(cfg: &Config) -> Option<(Role, PathBuf)> {
smallest(Some(Role::Detector)).or_else(|| smallest(None))
}
/// Build, run once for the engine, then time three runs; the median in
/// milliseconds and, for a compiling rung, the cache key of the engine this
/// just built.
/// Build, warm up, then time seven runs; the median in milliseconds and,
/// for a compiling rung, the cache key of the engine this just built.
///
/// Three warm-ups, not one: an idle integrated GPU takes a few runs to
/// raise its clock. With one, the Iris Xe's OpenVINO lost to the CPU on
/// the smallest detector in two probes of three, where warm it is 5.8 ms
/// against 9.5 (§1.6). The smallest detector is a GPU's worst case; the
/// clock must not also be.
fn time_rung(
rung: Rung,
role: Role,
@@ -203,51 +270,65 @@ fn time_rung(
cfg: &Config,
) -> Result<(f64, Option<String>), String> {
let want = rung.form(role);
let path = match want {
Form::Int8 => {
let p = crate::int8_sibling(canonical);
if !p.is_file() {
return Err(format!("no int8 form of {}", canonical.display()));
}
p
}
Form::F32 => canonical.to_path_buf(),
};
let path = crate::form_sibling(canonical, want);
if want != Form::F32 && !path.is_file() {
return Err(format!(
"no {} form of {}",
want.file_tag().unwrap_or("f32"),
canonical.display()
));
}
let bytes = std::fs::read(&path).map_err(|e| e.to_string())?;
let started = Instant::now();
let mut session =
crate::session::build(rung, role, &bytes, cfg).map_err(|e| first_line(&e.to_string()))?;
let mut session = crate::session::build_probe(rung, role, &bytes, cfg)
.map_err(|e| first_line(&e.to_string()))?;
log::info!(
"inference: {} session built in {:.1} s",
rung.label(),
started.elapsed().as_secs_f64()
);
let shape: Vec<usize> = session.inputs()[0]
.dtype()
.tensor_shape()
.ok_or("model input is not a tensor")?
// Zeros for every input the model declares, by name — the denoiser
// takes two (mosaic and σ), and a probe that fed only the first failed
// every rung and left it on the CPU.
let feeds: Vec<(String, Vec<usize>)> = session
.inputs()
.iter()
.map(|&d| if d > 0 { d as usize } else { 1 })
.collect();
let zeros = vec![0f32; shape.iter().product()];
.map(|i| {
let shape = i
.dtype()
.tensor_shape()
.ok_or("model input is not a tensor")?
.iter()
.map(|&d| if d > 0 { d as usize } else { 1 })
.collect();
Ok((i.name().to_string(), shape))
})
.collect::<Result<_, &str>>()?;
let run = |session: &mut ort::session::Session| -> Result<f64, String> {
let input = ort::value::Tensor::from_array((shape.clone(), zeros.clone()))
.map_err(|e| e.to_string())?;
let mut inputs: Vec<(String, ort::session::SessionInputValue)> = Vec::new();
for (name, shape) in &feeds {
let zeros = vec![0f32; shape.iter().product()];
let t = ort::value::Tensor::from_array((shape.clone(), zeros))
.map_err(|e| e.to_string())?;
inputs.push((name.clone(), t.into()));
}
let t = Instant::now();
let out = session
.run(ort::inputs![input])
.map_err(|e| e.to_string())?;
let out = session.run(inputs).map_err(|e| e.to_string())?;
let _ = out[0]
.try_extract_tensor::<f32>()
.map_err(|e| e.to_string())?;
Ok(t.elapsed().as_secs_f64() * 1e3)
};
run(&mut session)?;
let mut times = [run(&mut session)?, run(&mut session)?, run(&mut session)?];
for _ in 0..3 {
run(&mut session)?;
}
let mut times = (0..7)
.map(|_| run(&mut session))
.collect::<Result<Vec<_>, _>>()?;
times.sort_by(|a, b| a.partial_cmp(b).unwrap());
let key = rung.compiles().then(|| crate::engines::key(rung, &bytes));
Ok((times[1], key))
Ok((times[times.len() / 2], key))
}
/// The part of a provider's error a person can act on. ONNX Runtime's
@@ -283,9 +364,9 @@ fn fingerprint(runtime: &Runtime, cfg: &Config) -> String {
},
device_identity(),
];
for (role, bytes) in &cfg.embedded {
for (role, form, bytes) in &cfg.embedded {
parts.push(format!(
"{role:?} embedded {:016x}",
"{role:?} embedded {form:?} {:016x}",
crate::engines::hash(bytes)
));
}
@@ -294,9 +375,13 @@ fn fingerprint(runtime: &Runtime, cfg: &Config) -> String {
.map(|b| crate::engines::hash(&b))
.unwrap_or(0);
parts.push(format!("{role:?} {hash:016x}"));
let int8 = crate::int8_sibling(path);
if let Ok(b) = std::fs::read(&int8) {
parts.push(format!("{role:?} int8 {:016x}", crate::engines::hash(&b)));
for form in [Form::Int8, Form::A16W8, Form::A16W16] {
if let Ok(b) = std::fs::read(crate::form_sibling(path, form)) {
parts.push(format!(
"{role:?} {form:?} {:016x}",
crate::engines::hash(&b)
));
}
}
}
parts.join("\n")
@@ -324,19 +409,35 @@ fn providers_beside(runtime: &Path) -> String {
#[cfg(target_os = "linux")]
fn device_identity() -> String {
// The NVIDIA driver's version line, or the ROCm release the AMD stack
// came from (`rocm-core` writes it; the kernel driver has no version
// of its own). Absent means neither.
// The NVIDIA driver's version line, the ROCm release the AMD stack came
// from (`rocm-core` writes it; the kernel driver has no version of its
// own), and the OpenCL drivers registered — OpenVINO reaches the GPU
// through one, and installing Intel's is what makes the Iris Xe a rung.
let mut parts = Vec::new();
if let Some(line) = std::fs::read_to_string("/proc/driver/nvidia/version")
.ok()
.and_then(|s| s.lines().next().map(str::to_string))
{
return line;
parts.push(line);
}
if let Ok(rocm) = std::fs::read_to_string("/opt/rocm/.info/version") {
return format!("rocm {}", rocm.trim());
parts.push(format!("rocm {}", rocm.trim()));
}
let mut icds: Vec<String> = std::fs::read_dir("/etc/OpenCL/vendors")
.into_iter()
.flatten()
.filter_map(|e| e.ok())
.map(|e| e.file_name().to_string_lossy().into_owned())
.collect();
icds.sort();
if !icds.is_empty() {
parts.push(format!("opencl {}", icds.join(" ")));
}
if parts.is_empty() {
"no nvidia driver, no rocm, no opencl".into()
} else {
parts.join("; ")
}
"no nvidia driver, no rocm".into()
}
#[cfg(target_os = "android")]
@@ -351,7 +452,7 @@ fn device_identity() -> String {
}
#[cfg(target_os = "android")]
fn system_property(name: &str) -> String {
pub(crate) fn system_property(name: &str) -> String {
extern "C" {
fn __system_property_get(
name: *const std::ffi::c_char,
@@ -481,4 +582,34 @@ mod tests {
died_inside(&cfg, "probe TensorRT", 2);
assert_eq!(attempt(&cfg, "probe CUDA", || 7), Ok(7));
}
/// The tablet's cache after 0.22.0's first launch, as 0.22.0 wrote it:
/// no `attempts`, the Hexagon "rejected", the CPU selected.
const TABLET: &str = r#"{"fingerprint":"f","rung":"Cpu","reason":"Hexagon NPU 28.5 ms, slower than the CPU's 19.4 ms","compiled":[],"failed":[["Hexagon","28.5 ms, slower than the CPU's 19.4 ms"]]}"#;
#[test]
fn a_fall_back_to_the_cpu_is_probed_again_a_few_times() {
let mut cache: Cache = serde_json::from_str(TABLET).unwrap();
assert_eq!(cache.attempts, 0, "a 0.22.0 cache reads as never retried");
assert_eq!(reuse(&cache, "f"), Reuse::Again(0));
cache.attempts = RETRIES - 1;
assert_eq!(reuse(&cache, "f"), Reuse::Again(RETRIES - 1));
cache.attempts = RETRIES;
assert_eq!(reuse(&cache, "f"), Reuse::Keep, "then it is kept");
}
#[test]
fn an_accelerator_chosen_or_a_cpu_only_device_is_kept() {
let mut cache: Cache = serde_json::from_str(TABLET).unwrap();
cache.rung = Some(Rung::Hexagon);
assert_eq!(reuse(&cache, "f"), Reuse::Keep);
cache.rung = Some(Rung::Cpu);
cache.failed.clear();
assert_eq!(
reuse(&cache, "f"),
Reuse::Keep,
"nothing failed: the only rung"
);
assert_eq!(reuse(&cache, "other"), Reuse::Fresh);
}
}
+198 -10
View File
@@ -13,6 +13,30 @@ use crate::{Config, Role, Rung};
/// its clock instead (§4): a provider that hands real work to the CPU is
/// slower than the CPU floor and rejected by the same measurement.
pub fn build(rung: Rung, role: Role, bytes: &[u8], cfg: &Config) -> ort::Result<Session> {
build_with(rung, role, bytes, cfg, false)
}
/// [`build`] for the probe: on the Hexagon, a session that cannot put the
/// whole graph on the NPU fails instead of running the rest on the CPU.
///
/// The probe times a rung by its session, and a QNN provider that could not
/// create its device still builds one — with every node on the CPU behind
/// it. 0.22.0's first launch on the tablet timed that (28.5 ms against the
/// CPU's own 19.4) and put every model on the CPU. Only the probe is strict:
/// some shipped graphs keep a few nodes on the CPU on purpose
/// (`tools/quantise-models.py`, `float_nodes`), and the probe's detector is
/// not one of them.
pub fn build_probe(rung: Rung, role: Role, bytes: &[u8], cfg: &Config) -> ort::Result<Session> {
build_with(rung, role, bytes, cfg, rung == Rung::Hexagon)
}
fn build_with(
rung: Rung,
role: Role,
bytes: &[u8],
cfg: &Config,
strict: bool,
) -> ort::Result<Session> {
// No optimisation level named. ONNX Runtime's default is already its
// fullest, and on tract any level but "disabled" means `into_optimized`,
// whose optimiser divides by zero inside yolo26n-seg (tract-data
@@ -22,15 +46,22 @@ pub fn build(rung: Rung, role: Role, bytes: &[u8], cfg: &Config) -> ort::Result<
if crate::api::runtime().is_native() {
b = with_runtime_log(b)?;
}
if strict {
b = b.with_config_entry("session.disable_cpu_ep_fallback", "1")?;
}
// A Hexagon session loads the compiled context when there is one and
// compiles it from the model when there is not; the engine thread is
// what makes the second case rare (§6).
let context = (rung == Rung::Hexagon).then(|| crate::engines::context_path(cfg, bytes));
let ready = context.as_ref().is_some_and(|p| p.is_file());
// What the rung keeps for this model: the context the Hexagon is to
// write, or the directory CoreML compiles into.
// write, or the directory CoreML or OpenVINO compiles into.
let per_model = match rung {
Rung::CoreMl => Some(crate::engines::coreml_dir(cfg, bytes)),
Rung::TensorRt if role == Role::WholeDenoiser => {
Some(crate::engines::tensorrt_whole_dir(cfg, bytes))
}
Rung::OpenVino => Some(crate::engines::openvino_dir(cfg, bytes, fp16(role))),
_ if ready => None,
_ => context.clone(),
};
@@ -74,6 +105,13 @@ fn with_runtime_log(
.with_log_level(level)?)
}
/// Whether `role` runs in fp16 on a rung that offers it: everything but the
/// embedder, whose comparability across devices is worth more than its
/// fraction of a millisecond (§7).
fn fp16(role: Role) -> bool {
role != Role::Embedder
}
/// The intra-op pool: what the config says, else the cores less two for
/// the compositor and the decoder (§9). tract ignores it.
fn threads(cfg: &Config) -> usize {
@@ -100,6 +138,11 @@ fn providers(
Rung::Cuda => {
Ok(b.with_execution_providers([ep::CUDA::default().build().error_on_failure()])?)
}
Rung::TensorRt if role == Role::WholeDenoiser => {
let mut b = b;
tensorrt_whole(&mut b, per_model.expect("a whole-frame engine directory"))?;
Ok(b.with_execution_providers([ep::CUDA::default().build()])?)
}
Rung::TensorRt => {
let cache = cfg.cache_dir.join("tensorrt");
let _ = std::fs::create_dir_all(&cache);
@@ -110,7 +153,7 @@ fn providers(
// (NFR-RES-2). CUDA behind it takes any node TensorRT declines.
Ok(b.with_execution_providers([
ep::TensorRT::default()
.with_fp16(role != Role::Embedder)
.with_fp16(fp16(role))
.with_engine_cache(true)
.with_engine_cache_path(&cache)
.with_timing_cache(true)
@@ -127,7 +170,7 @@ fn providers(
// directory, keyed on the graph, the GPU and its own version
// but not the precision: hence one directory per precision.
// The CPU takes any node it declines.
let fp16 = role != Role::Embedder;
let fp16 = fp16(role);
let cache = cfg
.cache_dir
.join("migraphx")
@@ -137,6 +180,16 @@ fn providers(
migraphx(&mut b, fp16, &cache)?;
Ok(b)
}
Rung::OpenVino => {
let mut b = b;
openvino(&mut b, fp16(role), per_model)?;
Ok(b)
}
Rung::WebGpu => {
let mut b = b;
webgpu(&mut b)?;
Ok(b)
}
Rung::Hexagon => unreachable!("the Hexagon rung is not on a desktop ladder"),
}
}
@@ -192,15 +245,145 @@ fn migraphx(
b: &mut ort::session::builder::SessionBuilder,
fp16: bool,
cache: &std::path::Path,
) -> ort::Result<()> {
append(
b,
c"MIGraphX",
&[
("migraphx_fp16_enable", if fp16 { "1" } else { "0" }.into()),
(
"migraphx_model_cache_dir",
cache.to_string_lossy().into_owned(),
),
],
)
}
/// TensorRT for a whole-frame model: one engine for every input size up to
/// [`crate::WHOLE_FRAME_MAX`], kept in its own directory.
///
/// `ort`'s builder has no profile options, so this registers through the
/// runtime's TensorRT V2 options, with the names 1.30 reads
/// (`tensorrt_execution_provider_info.cc`): `trt_profile_{min,opt,max}_shapes`.
/// Without a profile a dynamic input compiles a new engine per size at run
/// time — 156 s on the first frame, measured — so the profile is the
/// difference between a whole-frame engine and a stall. fp16, as for every
/// role but the embedder (§7); the denoiser measured 0.00 dB from f32.
#[cfg(not(target_os = "android"))]
fn tensorrt_whole(
b: &mut ort::session::builder::SessionBuilder,
cache: &std::path::Path,
) -> ort::Result<()> {
use ort::AsPointer;
use std::ffi::CString;
let keys = [c"migraphx_fp16_enable", c"migraphx_model_cache_dir"];
let values = [
CString::new(if fp16 { "1" } else { "0" }).unwrap(),
CString::new(cache.to_string_lossy().as_bytes())
.map_err(|e| ort::Error::new(e.to_string()))?,
let _ = std::fs::create_dir_all(cache);
let shapes = |(h, w): (usize, usize)| format!("mosaic:1x1x{h}x{w},sigma:1x1x{h}x{w}");
let dir = cache.to_string_lossy().into_owned();
let options = [
("trt_fp16_enable", "1".to_string()),
("trt_engine_cache_enable", "1".to_string()),
("trt_engine_cache_path", dir.clone()),
("trt_timing_cache_enable", "1".to_string()),
("trt_timing_cache_path", dir),
("trt_max_workspace_size", (1u64 << 30).to_string()),
("trt_profile_min_shapes", shapes((256, 256))),
("trt_profile_opt_shapes", shapes(crate::WHOLE_FRAME_OPT)),
("trt_profile_max_shapes", shapes(crate::WHOLE_FRAME_MAX)),
];
let cstr = |s: &str| CString::new(s).map_err(|e| ort::Error::new(e.to_string()));
let keys = options
.iter()
.map(|(k, _)| cstr(k))
.collect::<ort::Result<Vec<_>>>()?;
let values = options
.iter()
.map(|(_, v)| cstr(v))
.collect::<ort::Result<Vec<_>>>()?;
let key_ptrs: Vec<_> = keys.iter().map(|k| k.as_ptr()).collect();
let value_ptrs: Vec<_> = values.iter().map(|v| v.as_ptr()).collect();
let api = ort::api();
// SAFETY: the documented create / update / append / release sequence
// `ort`'s own TensorRT builder makes, over arrays that outlive it; the
// runtime copies the options into the session before the release.
unsafe {
let mut trt: *mut ort::sys::OrtTensorRTProviderOptionsV2 = std::ptr::null_mut();
ort::Error::result_from_status((api.CreateTensorRTProviderOptions)(&mut trt))?;
let result = ort::Error::result_from_status((api.UpdateTensorRTProviderOptions)(
trt,
key_ptrs.as_ptr(),
value_ptrs.as_ptr(),
keys.len(),
))
.and_then(|()| {
ort::Error::result_from_status((api.SessionOptionsAppendExecutionProvider_TensorRT_V2)(
b.ptr_mut(),
trt,
))
});
(api.ReleaseTensorRTProviderOptions)(trt);
result
}
}
/// OpenVINO on the GPU, compiling into `cache`.
///
/// The option names are those `openvino_provider_factory.cc` reads at 1.24,
/// the version of Intel's `onnxruntime-openvino` build. `GPU` is OpenVINO's
/// first OpenCL GPU: the Intel one on a hybrid laptop with both drivers
/// installed, but an NVIDIA card through its OpenCL when Intel's is absent
/// — slower than the CPU there, and rejected by the probe's clock. The
/// precision is always named: the GPU plugin's own default is fp16, and
/// the embedder must not get it (§7).
#[cfg(not(target_os = "android"))]
fn openvino(
b: &mut ort::session::builder::SessionBuilder,
fp16: bool,
cache: Option<&std::path::Path>,
) -> ort::Result<()> {
let mut options = vec![
("device_type", "GPU".to_string()),
("precision", if fp16 { "FP16" } else { "FP32" }.into()),
];
if let Some(dir) = cache {
let _ = std::fs::create_dir_all(dir);
options.push(("cache_dir", dir.to_string_lossy().into_owned()));
}
append(b, c"OpenVINO", &options)
}
/// WebGPU on the high-performance adapter: the discrete GPU where there is
/// one, since the integrated one on a machine with both is the one this
/// rung is least likely to beat the CPU on. The key is as
/// `webgpu_provider_options.h` spells it, without the `ep.<name>.` prefix
/// the runtime adds.
fn webgpu(b: &mut ort::session::builder::SessionBuilder) -> ort::Result<()> {
append(
b,
c"WebGPU",
&[("powerPreference", "high-performance".to_string())],
)
}
/// Register the provider `name` with `options` through the runtime's
/// generic key/value entry point, which takes every provider by its short
/// name and reads options at the runtime's own version — not at the
/// version `ort`'s builders were written against (CLAUDE.md, "Providers").
fn append(
b: &mut ort::session::builder::SessionBuilder,
name: &std::ffi::CStr,
options: &[(&str, String)],
) -> ort::Result<()> {
use ort::AsPointer;
use std::ffi::CString;
let cstr = |s: &str| CString::new(s).map_err(|e| ort::Error::new(e.to_string()));
let keys = options
.iter()
.map(|(k, _)| cstr(k))
.collect::<ort::Result<Vec<_>>>()?;
let values = options
.iter()
.map(|(_, v)| cstr(v))
.collect::<ort::Result<Vec<_>>>()?;
let key_ptrs: Vec<_> = keys.iter().map(|k| k.as_ptr()).collect();
let value_ptrs: Vec<_> = values.iter().map(|v| v.as_ptr()).collect();
// SAFETY: the documented C call over arrays that outlive it; the
@@ -208,7 +391,7 @@ fn migraphx(
unsafe {
let status = (ort::api().SessionOptionsAppendExecutionProvider)(
b.ptr_mut(),
c"MIGraphX".as_ptr(),
name.as_ptr(),
key_ptrs.as_ptr(),
value_ptrs.as_ptr(),
keys.len(),
@@ -250,7 +433,12 @@ fn providers(
.build()
.error_on_failure()])?)
}
Rung::Cuda | Rung::TensorRt | Rung::MiGraphX | Rung::CoreMl => {
Rung::WebGpu => {
let mut b = b;
webgpu(&mut b)?;
Ok(b)
}
Rung::Cuda | Rung::TensorRt | Rung::MiGraphX | Rung::CoreMl | Rung::OpenVino => {
unreachable!("no desktop rung on Android")
}
}
+12 -1
View File
@@ -13,8 +13,13 @@ const MODELS: &[&str] = &[
"../../models/keypoints/xfeat-768.onnx",
];
const QUANTISED: &[&str] = &[
"../../models/keypoints/xfeat-1024.int8.onnx",
"../../models/keypoints/xfeat-768.int8.onnx",
];
fn main() {
for m in MODELS {
for m in MODELS.iter().chain(QUANTISED) {
println!("cargo:rerun-if-changed={m}");
}
println!("cargo:rerun-if-changed=build.rs");
@@ -26,6 +31,12 @@ fn main() {
for model in MODELS.iter().copied() {
check(model);
}
// The Hexagon's int8 forms ride only in an Android build.
if std::env::var("CARGO_CFG_TARGET_OS").as_deref() == Ok("android") {
for model in QUANTISED.iter().copied() {
check(model);
}
}
}
fn check(model: &str) {
+3
View File
@@ -22,6 +22,7 @@
//! - [`align`] — the whole thing, from features to cameras, honest about
//! what it could not place.
//! - [`projection`] — perspective, cylindrical, spherical.
//! - [`seam`] — which frame each output pixel is taken from.
//! - [`linalg`] — the small dense algebra all of it uses.
//!
//! # What it depends on
@@ -43,6 +44,7 @@ pub mod matching;
#[cfg(feature = "xfeat")]
pub mod migan;
pub mod projection;
pub mod seam;
#[cfg(feature = "xfeat")]
pub mod xfeat;
@@ -52,6 +54,7 @@ pub use features::{Features, Keypoint};
pub use fill::{fill_border, Inpainter, Observer, Params as FillParams};
pub use image::Gray;
pub use projection::Projection;
pub use seam::{SeamMap, SeamOptions};
#[derive(Debug, thiserror::Error)]
pub enum PanoError {
+2 -1
View File
@@ -30,7 +30,8 @@ pub struct MiGan {
impl MiGan {
/// From the model file, in whichever form the engine's rung wants
/// (`resolve_model` picks an int8 sibling for the Hexagon).
/// (`resolve_model` picks the `.a16w16.onnx` sibling on the Hexagon:
/// int8 moved the fill 16 dB from f32's, 16-bit about 41).
pub fn from_path(path: &std::path::Path) -> Result<Self, PanoError> {
use dr_inference_engine::{resolve_model, Role};
let (path, form) = resolve_model(Role::Inpainter, path);
+691
View File
@@ -0,0 +1,691 @@
//! TRACES: FR-MRG-10
//! Where each frame gives way to the next.
//!
//! The first merges averaged every overlap: each frame weighted by its
//! distance from its own edge, so that across two hundred pixels one frame
//! faded into the other. That hides an exposure step and does not hide
//! anything that differs between the frames — parallax on a near slope, a
//! walker, a branch in the wind — which the average draws twice, half as
//! bright, a soft double edge at 1:1.
//!
//! A seam answers it the way every stitcher does: in an overlap, each output
//! pixel is taken from *one* frame, and the line where the choice changes is
//! put where the frames agree and the picture is smooth — through sky,
//! along a shadow, round the walker rather than through him — and away from
//! either frame's edge, where vignetting and the lens correction's fringe
//! live. The blend is then narrow and only across that line.
//!
//! # How
//!
//! At proxy resolution, on the output surface, which fits (panorama.md §5:
//! "it is a mask, not an image"):
//!
//! 1. Frames are laid down one at a time, each next to one already placed.
//! The composite so far is a label per texel and the value its owner saw.
//! 2. Where a new frame overlaps the composite, a cost per texel: the
//! difference between the two (after the gains), how much detail either
//! has there, and how near either frame's edge it is — smoothed over a
//! few texels, because "agree" means locally, not at one pixel.
//! 3. The cut is a path across the overlap, perpendicular to the line from
//! the composite's frames to the new one, found by dynamic programming
//! one row at a time: the per-column seam panorama.md §4 chose over a
//! graph cut because it is the GPU-friendly shape. Texels on the new
//! frame's side of the path become its own.
//!
//! What the merge reads is [`SeamMap::share`]: the fraction of a small
//! window about a point that is labelled with a frame, tent-weighted, which
//! is a narrow blend that follows the seam. `merge.wgsl` computes the same
//! thing on the GPU from the same labels.
use crate::bundle::Cameras;
use crate::image::Gray;
use crate::projection::{self, Projection};
/// No frame owns this texel.
pub const NONE: u8 = 255;
/// The most frames a map can label: one less than [`NONE`].
pub const MAX_FRAMES: usize = NONE as usize;
/// Which frame each texel of the output takes its pixels from.
#[derive(Debug, Clone, PartialEq)]
pub struct SeamMap {
pub width: usize,
pub height: usize,
/// The projection scale the map was laid out at: the proxies' focal
/// length. Output coordinates at any other scale are this times the
/// ratio of the scales.
pub scale: f64,
/// Centred output coordinates, at `scale`, of texel (0, 0)'s top-left
/// corner.
pub origin: (f64, f64),
/// Output units per texel, at `scale`.
pub px: f64,
/// Row-major, one per texel: the frame's index, or [`NONE`].
pub labels: Vec<u8>,
}
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct SeamOptions {
/// The widest the map is laid out, in texels. Wider than the proxies'
/// own resolution buys nothing.
pub max_width: usize,
/// How much detail costs against disagreement: a seam through texture
/// shows even where the frames agree, because the blend across it
/// softens it.
pub detail: f32,
/// How much a frame's edge costs, and how far in from it the cost
/// reaches, in proxy pixels. Frame edges are where vignetting is
/// darkest and the lens correction ran out of sensor.
pub edge: f32,
pub edge_margin: f32,
/// The radius, in texels, a texel's cost looks about it for the worst
/// of its neighbours: at least the radius the merge blends across.
pub smoothing: usize,
}
impl Default for SeamOptions {
fn default() -> Self {
SeamOptions {
max_width: 2048,
detail: 0.5,
edge: 0.5,
edge_margin: 24.0,
smoothing: 4,
}
}
}
/// The most texels a blend reaches either side of a seam. The merge's
/// shader loads the square of twice this per pixel per frame near a seam.
pub const MAX_BLEND_RADIUS: f64 = 4.0;
/// Cost of a texel outside the overlap: high enough that the path keeps to
/// the overlap wherever there is one, finite so that a row with a gap in it
/// still has an answer.
const OUTSIDE: f32 = 1.0e3;
impl SeamMap {
/// The map's origin and texel size in the coordinates of an output
/// laid out at `scale` (the full-resolution focal length, or a fraction
/// of it).
pub fn at_scale(&self, scale: f64) -> ((f64, f64), f64) {
let r = scale / self.scale;
((self.origin.0 * r, self.origin.1 * r), self.px * r)
}
/// The radius, in texels, of a blend `blend_px` output pixels wide in an
/// output laid out at `scale`: what [`Self::share`] and the shader are
/// given, so that the preview and the merge blend alike.
pub fn blend_radius(&self, scale: f64, blend_px: f64) -> f64 {
let (_, px) = self.at_scale(scale);
(blend_px / 2.0 / px).clamp(1.0, MAX_BLEND_RADIUS)
}
/// The share frame `k` has of output point `(u, v)` given at `scale`:
/// the tent-weighted fraction of the texels within `radius` (in texels)
/// that it owns. `None` where no texel in reach is owned at all — the
/// map has nothing to say there, and the caller falls back to its
/// feather.
///
/// This is the function `merge.wgsl`'s `seam_share` repeats; the two
/// must agree.
pub fn share(&self, k: usize, u: f64, v: f64, scale: f64, radius: f64) -> Option<f32> {
let ((ou, ov), px) = self.at_scale(scale);
let x = (u - ou) / px - 0.5;
let y = (v - ov) / px - 0.5;
let r = radius.max(1.0);
let (x0, x1) = ((x - r).ceil() as i64, (x + r).floor() as i64);
let (y0, y1) = ((y - r).ceil() as i64, (y + r).floor() as i64);
let (mut mine, mut all) = (0.0f64, 0.0f64);
for j in y0.max(0)..=y1.min(self.height as i64 - 1) {
let wy = 1.0 - (y - j as f64).abs() / r;
if wy <= 0.0 {
continue;
}
for i in x0.max(0)..=x1.min(self.width as i64 - 1) {
let wx = 1.0 - (x - i as f64).abs() / r;
if wx <= 0.0 {
continue;
}
let l = self.labels[j as usize * self.width + i as usize];
if l == NONE {
continue;
}
all += wx * wy;
if usize::from(l) == k {
mine += wx * wy;
}
}
}
(all > 0.0).then(|| (mine / all) as f32)
}
}
/// One frame warped onto the map: its gain-corrected value and its distance
/// from its own edge (in proxy pixels) per texel, NaN where it does not
/// reach.
struct Warped {
value: Vec<f32>,
edge: Vec<f32>,
}
/// Lay seams across the overlaps of `proxies`, aligned by `cameras` (at the
/// proxies' scale), with `gains` the linear multipliers the merge will
/// apply. `None` if the frames project nowhere or there are more than
/// [`MAX_FRAMES`].
pub fn find(
proxies: &[&Gray],
cameras: &Cameras,
gains: &[f32],
projection: Projection,
opts: &SeamOptions,
) -> Option<SeamMap> {
let n = proxies.len();
if n == 0 || n > MAX_FRAMES || cameras.rotations.len() != n || gains.len() != n {
return None;
}
let (fw, fh) = (proxies[0].width as f64, proxies[0].height as f64);
let scale = cameras.focal;
let bounds = projection::bounds(projection, scale, cameras, (fw, fh))?;
let width = opts.max_width.min(bounds.width().ceil() as usize).max(1);
let px = bounds.width() / width as f64;
let height = ((bounds.height() / px).ceil() as usize).max(1);
let mut map = SeamMap {
width,
height,
scale,
origin: (bounds.min_u, bounds.min_v),
px,
labels: vec![NONE; width * height],
};
// Where each frame's centre lands, in texels: what orders the frames
// and orients each cut.
let centres: Vec<(f64, f64)> = (0..n)
.map(|k| {
let d = cameras.bearing(k, (0.0, 0.0));
projection
.from_direction(scale, d)
.map(|(u, v)| ((u - bounds.min_u) / px, (v - bounds.min_v) / px))
.unwrap_or((width as f64 / 2.0, height as f64 / 2.0))
})
.collect();
// The composite so far: what its owner saw, and how far from the
// owner's edge.
let mut value = vec![f32::NAN; width * height];
let mut edge = vec![f32::NAN; width * height];
for k in order(&centres, (width as f64 / 2.0, height as f64 / 2.0)) {
let w = warp(&map, proxies[k], cameras, k, gains[k], projection);
let overlap: Vec<usize> = (0..width * height)
.filter(|&i| map.labels[i] != NONE && !w.value[i].is_nan())
.collect();
// Texels nobody owns yet are the new frame's without a cut.
let mut take: Vec<bool> = map
.labels
.iter()
.zip(&w.value)
.map(|(&l, v)| l == NONE && !v.is_nan())
.collect();
if !overlap.is_empty() {
cut(
&map, &value, &edge, &w, &overlap, &centres, k, opts, &mut take,
);
}
for i in 0..width * height {
if take[i] {
map.labels[i] = k as u8;
value[i] = w.value[i];
edge[i] = w.edge[i];
}
}
}
Some(map)
}
/// The order frames are laid down in: the one nearest the middle first,
/// then always the unplaced frame nearest any placed one, so that each new
/// frame meets the composite along an overlap rather than across a gap.
fn order(centres: &[(f64, f64)], middle: (f64, f64)) -> Vec<usize> {
let d2 = |a: (f64, f64), b: (f64, f64)| (a.0 - b.0).powi(2) + (a.1 - b.1).powi(2);
let n = centres.len();
let mut placed = vec![false; n];
let mut out = Vec::with_capacity(n);
let first = (0..n)
.min_by(|&a, &b| d2(centres[a], middle).total_cmp(&d2(centres[b], middle)))
.expect("at least one frame");
placed[first] = true;
out.push(first);
while out.len() < n {
let next = (0..n)
.filter(|&k| !placed[k])
.min_by(|&a, &b| {
let near = |k: usize| {
out.iter()
.map(|&p| d2(centres[k], centres[p]))
.fold(f64::MAX, f64::min)
};
near(a).total_cmp(&near(b))
})
.expect("an unplaced frame");
placed[next] = true;
out.push(next);
}
out
}
/// Frame `k` sampled at every texel's centre, bilinearly. The proxy is
/// gamma-encoded grey, so the gain (linear) becomes `gain^(1/2.2)` on it.
fn warp(
map: &SeamMap,
g: &Gray,
cameras: &Cameras,
k: usize,
gain: f32,
projection: Projection,
) -> Warped {
let (fw, fh) = (g.width as f64, g.height as f64);
let gain = gain.max(1e-6).powf(1.0 / 2.2);
let mut value = vec![f32::NAN; map.width * map.height];
let mut edge = vec![f32::NAN; map.width * map.height];
for ty in 0..map.height {
let v = map.origin.1 + (ty as f64 + 0.5) * map.px;
for tx in 0..map.width {
let u = map.origin.0 + (tx as f64 + 0.5) * map.px;
let d = projection.to_direction(map.scale, u, v);
let Some((x, y)) = cameras.project(k, d) else {
continue;
};
let (x, y) = (x + fw / 2.0 - 0.5, y + fh / 2.0 - 0.5);
let e = x.min(fw - 1.0 - x).min(y).min(fh - 1.0 - y);
if e < 0.0 {
continue;
}
let (x0, y0) = (x.floor() as usize, y.floor() as usize);
let (x1, y1) = ((x0 + 1).min(g.width - 1), (y0 + 1).min(g.height - 1));
let (ax, ay) = ((x - x0 as f64) as f32, (y - y0 as f64) as f32);
let at = |xx: usize, yy: usize| g.data[yy * g.width + xx];
let top = at(x0, y0) * (1.0 - ax) + at(x1, y0) * ax;
let bot = at(x0, y1) * (1.0 - ax) + at(x1, y1) * ax;
let i = ty * map.width + tx;
value[i] = (top * (1.0 - ay) + bot * ay) * gain;
edge[i] = e as f32;
}
}
Warped { value, edge }
}
/// Central-difference gradient magnitude of `plane` at texel `i`, from the
/// neighbours that exist.
fn detail(plane: &[f32], width: usize, height: usize, i: usize) -> f32 {
let (x, y) = (i % width, i / width);
let c = plane[i];
let mut g = 0.0f32;
let mut diff = |j: usize| {
let n = plane[j];
if !n.is_nan() {
g = g.max((n - c).abs());
}
};
if x > 0 {
diff(i - 1);
}
if x + 1 < width {
diff(i + 1);
}
if y > 0 {
diff(i - width);
}
if y + 1 < height {
diff(i + width);
}
g
}
/// Cut the overlap between the composite and frame `k`, marking in `take`
/// the overlap texels that go to `k`.
#[allow(clippy::too_many_arguments)]
fn cut(
map: &SeamMap,
value: &[f32],
edge: &[f32],
new: &Warped,
overlap: &[usize],
centres: &[(f64, f64)],
k: usize,
opts: &SeamOptions,
take: &mut [bool],
) {
let (w, h) = (map.width, map.height);
// The raw cost per overlap texel.
let mut raw = vec![f32::NAN; w * h];
let margin = opts.edge_margin.max(1.0);
for &i in overlap {
let differ = (value[i] - new.value[i]).abs();
let detail = detail(value, w, h, i).max(detail(&new.value, w, h, i));
let near = (1.0 - edge[i].min(new.edge[i]) / margin).max(0.0);
raw[i] = differ + opts.detail * detail + opts.edge * near * near + 1e-3;
}
// The worst over a small window: a texel is only cheap if its whole
// neighbourhood agrees, so the path keeps at least the blend's radius
// clear of a difference rather than threading the one lucky texel
// beside it — the blend straddles the path by that much and would
// otherwise reach the difference anyway.
let r = opts.smoothing as isize;
let mut cost = vec![OUTSIDE; w * h];
for &i in overlap {
let (x, y) = ((i % w) as isize, (i / w) as isize);
let mut worst = 0.0f32;
for dy in -r..=r {
for dx in -r..=r {
let (xx, yy) = (x + dx, y + dy);
if xx < 0 || yy < 0 || xx >= w as isize || yy >= h as isize {
continue;
}
let c = raw[yy as usize * w + xx as usize];
if !c.is_nan() {
worst = worst.max(c);
}
}
}
cost[i] = worst;
}
// The axis the cut crosses: from the composite's frames, weighted by how
// much of the overlap each owns, to the new frame.
let mut from = (0.0f64, 0.0f64);
for &i in overlap {
let c = centres[usize::from(map.labels[i])];
from = (from.0 + c.0, from.1 + c.1);
}
let m = overlap.len() as f64;
from = (from.0 / m, from.1 / m);
let to = centres[k];
let (mut ax, mut ay) = (to.0 - from.0, to.1 - from.1);
let len = (ax * ax + ay * ay).sqrt();
if len < 1e-6 {
(ax, ay) = (1.0, 0.0);
} else {
(ax, ay) = (ax / len, ay / len);
}
// Along the cut: perpendicular to the axis.
let (bx, by) = (-ay, ax);
// The overlap's extent in (s along the cut, t across it).
let st = |i: usize| {
let (x, y) = ((i % w) as f64 + 0.5, (i / w) as f64 + 0.5);
(x * bx + y * by, x * ax + y * ay)
};
let (mut s0, mut s1, mut t0, mut t1) = (f64::MAX, f64::MIN, f64::MAX, f64::MIN);
for &i in overlap {
let (s, t) = st(i);
s0 = s0.min(s);
s1 = s1.max(s);
t0 = t0.min(t);
t1 = t1.max(t);
}
let rows = (s1 - s0).round() as usize + 1;
let cols = (t1 - t0).round() as usize + 1;
// The grid in (s, t), each cell sampled from the texel it falls in, so
// that a rotated overlap has no holes.
let mut grid = vec![OUTSIDE; rows * cols];
let mut any = vec![false; rows];
for si in 0..rows {
for ti in 0..cols {
let (s, t) = (s0 + si as f64, t0 + ti as f64);
let x = s * bx + t * ax;
let y = s * by + t * ay;
if x < 0.0 || y < 0.0 {
continue;
}
let (x, y) = (x as usize, y as usize);
if x >= w || y >= h {
continue;
}
let c = cost[y * w + x];
if c < OUTSIDE {
grid[si * cols + ti] = c;
any[si] = true;
}
}
}
// Dynamic programming down the rows: the path moves at most one column
// per row, and starts afresh after a row with no overlap in it.
let mut acc = grid.clone();
let mut from_col = vec![0u32; rows * cols];
for si in 1..rows {
if !any[si] {
continue;
}
let prev = &acc[(si - 1) * cols..si * cols].to_vec();
if !any[si - 1] {
continue;
}
for ti in 0..cols {
let mut best = (prev[ti], ti);
if ti > 0 && prev[ti - 1] < best.0 {
best = (prev[ti - 1], ti - 1);
}
if ti + 1 < cols && prev[ti + 1] < best.0 {
best = (prev[ti + 1], ti + 1);
}
acc[si * cols + ti] += best.0;
from_col[si * cols + ti] = best.1 as u32;
}
}
// Back up from the end of each run of rows with overlap.
let mut seam = vec![usize::MAX; rows];
let mut si = rows;
while si > 0 {
si -= 1;
if !any[si] {
continue;
}
let row = &acc[si * cols..(si + 1) * cols];
let mut t = (0..cols)
.min_by(|&a, &b| row[a].total_cmp(&row[b]))
.unwrap_or(0);
loop {
seam[si] = t;
if si == 0 || !any[si - 1] {
break;
}
t = from_col[si * cols + t] as usize;
si -= 1;
}
}
// The new frame takes the side of the path its centre is on.
for &i in overlap {
let (s, t) = st(i);
let si = ((s - s0).round() as usize).min(rows - 1);
let ti = (t - t0).round();
if seam[si] != usize::MAX && ti >= seam[si] as f64 {
take[i] = true;
}
}
}
#[cfg(test)]
mod tests {
use super::*;
use crate::linalg::{Mat3, Vec3};
/// A scene as a function of direction, and frames of it rendered by the
/// same cameras the seam reads.
fn render(
cameras: &Cameras,
k: usize,
size: (usize, usize),
scene: impl Fn(Vec3) -> f32,
) -> Gray {
let (w, h) = size;
let mut data = vec![0.0; w * h];
for y in 0..h {
for x in 0..w {
let p = (
x as f64 + 0.5 - w as f64 / 2.0,
y as f64 + 0.5 - h as f64 / 2.0,
);
data[y * w + x] = scene(cameras.bearing(k, p));
}
}
Gray {
width: w,
height: h,
data,
}
}
fn yaw(a: f64) -> Mat3 {
let (s, c) = a.sin_cos();
Mat3([[c, 0.0, s], [0.0, 1.0, 0.0], [-s, 0.0, c]])
}
/// Smooth, with a little texture: what a sky over a slope looks like to
/// the cost.
fn landscape(d: Vec3) -> f32 {
let (x, y) = (d.x() / d.z(), d.y() / d.z());
let texture = if y > 0.1 { 0.1 * (y * 40.0).sin() } else { 0.0 };
(0.5 + 0.2 * (x * 3.0).sin() + texture).clamp(0.0, 1.0) as f32
}
fn pair() -> Cameras {
Cameras {
rotations: vec![Mat3::IDENTITY, yaw(0.35)],
focal: 300.0,
}
}
#[test]
fn one_frame_owns_everything_it_reaches() {
let cameras = Cameras {
rotations: vec![Mat3::IDENTITY],
focal: 300.0,
};
let g = render(&cameras, 0, (320, 240), landscape);
let map = find(
&[&g],
&cameras,
&[1.0],
Projection::Perspective,
&Default::default(),
)
.unwrap();
let owned = map.labels.iter().filter(|&&l| l == 0).count();
assert!(owned as f64 > 0.95 * (map.width * map.height) as f64);
}
#[test]
fn each_frame_keeps_its_own_side() {
let cameras = pair();
let frames: Vec<Gray> = (0..2)
.map(|k| render(&cameras, k, (320, 240), landscape))
.collect();
let refs: Vec<&Gray> = frames.iter().collect();
let map = find(
&refs,
&cameras,
&[1.0, 1.0],
Projection::Cylindrical,
&Default::default(),
)
.unwrap();
let mid = map.height / 2 * map.width;
assert_eq!(map.labels[mid + 2], 0, "the left edge is frame 0's alone");
assert_eq!(
map.labels[mid + map.width - 3],
1,
"the right edge is frame 1's"
);
// One change of owner along every row that both frames cross.
for y in 0..map.height {
let row = &map.labels[y * map.width..(y + 1) * map.width];
let owned: Vec<u8> = row.iter().copied().filter(|&l| l != NONE).collect();
let changes = owned.windows(2).filter(|p| p[0] != p[1]).count();
assert!(changes <= 1, "row {y} changes owner {changes} times");
}
}
#[test]
fn the_seam_goes_round_what_only_one_frame_saw() {
// Frame 1 saw something frame 0 did not — a figure that walked into
// the overlap — in the middle of where the two meet.
let cameras = pair();
let figure = Vec3::new(0.175f64.sin(), 0.0, 0.175f64.cos());
let walker = |d: Vec3| {
let near = (d.x() - figure.x()).abs() < 0.04 && (d.y() - figure.y()).abs() < 0.15;
if near {
0.95
} else {
landscape(d)
}
};
let frames = [
render(&cameras, 0, (320, 240), landscape),
render(&cameras, 1, (320, 240), walker),
];
let refs: Vec<&Gray> = frames.iter().collect();
let map = find(
&refs,
&cameras,
&[1.0, 1.0],
Projection::Cylindrical,
&Default::default(),
)
.unwrap();
// Every texel of the figure is taken from the same frame, with a
// blend radius of room to spare, so it is either all there or not at
// all — never half.
let (u, v) = Projection::Cylindrical
.from_direction(map.scale, figure)
.unwrap();
let mut owners = std::collections::HashSet::new();
// The figure's extent on the surface, plus the blend's radius.
let radius = 3.0;
let reach = |half: f64| half * map.scale + radius * map.px;
let (ru, rv) = (reach(0.04), reach(0.15));
let mut dv = -rv;
while dv <= rv {
let mut du = -ru;
while du <= ru {
let s = map.share(1, u + du, v + dv, map.scale, radius);
owners.insert((s.unwrap() * 100.0).round() as i32);
du += map.px;
}
dv += map.px;
}
assert_eq!(owners.len(), 1, "the figure is split: shares {owners:?}");
}
#[test]
fn share_is_a_blend_across_the_seam_and_whole_away_from_it() {
let map = SeamMap {
width: 8,
height: 1,
scale: 1.0,
origin: (0.0, 0.0),
px: 1.0,
labels: vec![0, 0, 0, 0, 1, 1, 1, 1],
};
assert_eq!(map.share(0, 1.5, 0.5, 1.0, 2.0), Some(1.0));
assert_eq!(map.share(1, 6.5, 0.5, 1.0, 2.0), Some(1.0));
let at_seam = map.share(0, 4.0, 0.5, 1.0, 2.0).unwrap();
assert!((at_seam - 0.5).abs() < 1e-6, "{at_seam}");
// And at twice the scale, the same point is twice as far out.
assert_eq!(
map.share(0, 8.0, 1.0, 2.0, 2.0),
map.share(0, 4.0, 0.5, 1.0, 2.0)
);
let empty = SeamMap {
labels: vec![NONE; 8],
..map
};
assert_eq!(empty.share(0, 4.0, 0.5, 1.0, 2.0), None);
}
}
+37 -6
View File
@@ -36,18 +36,49 @@ pub struct XFeat {
pub options: DecodeOptions,
}
/// The bytes of both exports compiled into the binary, for whoever compiles
/// engines ahead of the first request (docs/dev/inference.md §6).
/// The Hexagon's forms (docs/dev/inference.md §1.5): int8, from the same
/// network spelled for the HTP (the unfold as SpaceToDepth, the bilinear
/// resizes as matrix products). Only Android has a Hexagon.
#[cfg(all(feature = "embedded-model", target_os = "android"))]
const EMBEDDED_LANDSCAPE_INT8: &[u8] =
include_bytes!("../../../models/keypoints/xfeat-1024.int8.onnx");
#[cfg(all(feature = "embedded-model", target_os = "android"))]
const EMBEDDED_PORTRAIT_INT8: &[u8] =
include_bytes!("../../../models/keypoints/xfeat-768.int8.onnx");
/// Every form of both exports compiled into the binary, landscape then
/// portrait, for whoever compiles engines ahead of the first request
/// (docs/dev/inference.md §6).
#[cfg(feature = "embedded-model")]
pub fn embedded_model_bytes() -> [&'static [u8]; 2] {
[EMBEDDED_LANDSCAPE, EMBEDDED_PORTRAIT]
pub fn embedded_models() -> [Vec<(dr_inference_engine::Form, &'static [u8])>; 2] {
use dr_inference_engine::Form;
#[allow(unused_mut)]
let mut forms = [
vec![(Form::F32, EMBEDDED_LANDSCAPE)],
vec![(Form::F32, EMBEDDED_PORTRAIT)],
];
#[cfg(target_os = "android")]
{
forms[0].push((Form::Int8, EMBEDDED_LANDSCAPE_INT8));
forms[1].push((Form::Int8, EMBEDDED_PORTRAIT_INT8));
}
forms
}
impl XFeat {
/// The weights compiled into the binary.
/// The weights compiled into the binary, in the form the device's
/// backend runs.
#[cfg(feature = "embedded-model")]
pub fn embedded() -> Result<Self, PanoError> {
Self::from_bytes(EMBEDDED_LANDSCAPE, EMBEDDED_PORTRAIT)
use dr_inference_engine::{choose_embedded, open, Role};
let [l, p] = embedded_models();
let (l, lf) = choose_embedded(Role::Keypoints, &l);
let (p, pf) = choose_embedded(Role::Keypoints, &p);
Ok(XFeat {
landscape: open(Role::Keypoints, lf, l)?,
portrait: open(Role::Keypoints, pf, p)?,
options: DecodeOptions::default(),
})
}
/// From the two exports on disk.
+19
View File
@@ -0,0 +1,19 @@
id: camera_profile
order: 25
# A `rust:` node publishes its own descriptor; its attributes are on the type
# in `../src/ops/camera_profile.rs`.
rust: CameraProfile
why_rust: |
It reads the source's profile tables from a storage buffer no declaration can
name, and it is composed at its defaults — a raw whose profile is on is
rendered through it without the photographer having touched anything —
which a declared node cannot say (D20).
placement: |
After exposure, before contrast (D20, camera-profiles.md §3). Hue and
saturation do not change under the uniform gains before it, so a 2.5-D
HueSatMap gives the same answer here as straight after the matrix; and the
LookTable sees the exposure the photographer chose, as the DNG SDK's does.
Contrast, tone and the colour controls then act on the profiled colour, as
they do in Camera Raw.
+36 -10
View File
@@ -21,24 +21,41 @@ params:
kind: amount
uniforms:
amount: vibrance / 100
amount:
value: vibrance / 100 * 1.3
doc: |
Scaled so that a value delivers the strength it names. Measured, not
chosen: fitted on 45 of the photographer's earlier exports whose only
colour setting was a vibrance of about +24, against their raws.
helpers: [luminance, tone_position, colour_saturation]
helpers: [luminance]
wgsl: |
let luma = luminance(c);
let sat = colour_saturation(c);
// How saturated a colour *looks*, so measured on display-encoded values.
// In scene-linear light an ordinary tan reads as 0.78 saturated and the
// falloff below would leave it a twentieth of the effect; encoded, it reads
// as 0.5, which is what the eye sees.
let e = pow(max(c, vec3<f32>(0.0)), vec3<f32>(1.0 / 2.2));
let e_hi = max(e.r, max(e.g, e.b));
let e_lo = min(e.r, min(e.g, e.b));
let sat = select(0.0, (e_hi - e_lo) / e_hi, e_hi > 0.00001);
// The vibrance curve: full effect on grey, tapering to nothing on colours
// that are already saturated. Squaring the falloff keeps the mid-range
// responsive while still protecting the extremes.
let falloff = (1.0 - sat) * (1.0 - sat);
// Skin protection. Skin sits in a narrow band of hue where red leads green
// leads blue; pushing it is what makes vibrance look wrong on portraits.
// Detected by channel ordering rather than a hue angle, which costs a
// conversion and buys nothing here.
let is_skin = f32(c.r > c.g && c.g > c.b);
// Skin protection, for skin: hues between about 10 and 50 degrees (red
// leading, green between red and blue) that are not strongly saturated.
// Red-over-green-over-blue alone is every warm colour in a photograph —
// wood, sand, brick, sunlit grass — and halving all of them is most of why
// vibrance used to do so little.
let span = max(e_hi - e_lo, 0.00001);
let skin_hue = select(0.0, 60.0 * (e.g - e.b) / span, e.r >= e.g && e.g >= e.b);
let in_band = smoothstep(4.0, 12.0, skin_hue) * (1.0 - smoothstep(42.0, 52.0, skin_hue));
let is_skin = in_band * (1.0 - smoothstep(0.45, 0.7, sat)) * f32(e.r >= e.g && e.g >= e.b);
let skin_guard = 1.0 - is_skin * 0.5;
let strength = amount * falloff * skin_guard;
@@ -49,9 +66,18 @@ tests:
- name: it_starts_neutral
expect_active: false
- name: the_amount_is_normalised_to_unit_range
- name: the_amount_is_the_measured_scale
why: |
Fitted against the photographer's earlier exports, so a value delivers
the strength it names.
set: { vibrance: 100 }
expect: { amount: 1.0 }
expect: { amount: 1.3 }
- name: saturation_is_judged_as_displayed
why: |
Judged in scene-linear light, ordinary warm colours read as nearly
saturated and get almost none of the effect.
expect_wgsl: ["let e = pow(max(c, vec3<f32>(0.0)), vec3<f32>(1.0 / 2.2));"]
- name: muted_colours_get_more_than_saturated_ones
why: |
+1 -1
View File
@@ -25,7 +25,7 @@ vibrance.vibrance = 10
blacks_whites.blacks = -8
clarity.amount = 12
contrast.contrast = 18
vibrance.vibrance = 18
vibrance.vibrance = 36
[preset Recover the sky]
blacks_whites.whites = -10
+13 -13
View File
@@ -15,38 +15,38 @@ drpl 1
[preset Blue sky]
colour_mixer.azure_lum = -20
colour_mixer.azure_sat = 25
colour_mixer.azure_sat = 18
colour_mixer.blue_lum = -15
colour_mixer.blue_sat = 20
colour_mixer.blue_sat = 14
highlights_shadows.highlights = -15
[preset Deep blue sky]
colour_mixer.azure_hue = 10
colour_mixer.azure_lum = -30
colour_mixer.azure_sat = 35
colour_mixer.azure_sat = 22
colour_mixer.blue_lum = -25
colour_mixer.blue_sat = 30
colour_mixer.cyan_sat = 10
colour_mixer.blue_sat = 19
colour_mixer.cyan_sat = 6
highlights_shadows.highlights = -30
[preset Polariser]
colour_mixer.azure_hue = 10
colour_mixer.azure_lum = -35
colour_mixer.azure_sat = 40
colour_mixer.azure_sat = 29
colour_mixer.blue_lum = -30
colour_mixer.blue_sat = 35
colour_mixer.blue_sat = 26
colour_mixer.cyan_lum = -10
colour_mixer.cyan_sat = 15
colour_mixer.cyan_sat = 11
dehaze.amount = 20
highlights_shadows.highlights = -35
vibrance.vibrance = 10
vibrance.vibrance = 7
[preset Blue sky, golden land]
colour_mixer.azure_lum = -20
colour_mixer.azure_sat = 25
colour_mixer.azure_sat = 16
colour_mixer.blue_lum = -15
colour_mixer.blue_sat = 20
colour_mixer.orange_sat = 12
colour_mixer.blue_sat = 13
colour_mixer.orange_sat = 8
colour_mixer.yellow_hue = -10
colour_mixer.yellow_sat = 15
colour_mixer.yellow_sat = 10
highlights_shadows.highlights = -20
+74
View File
@@ -0,0 +1,74 @@
drpl 1
# Vivid: more colour than the default rendering (camera-profiles.md §9).
#
# These do the work themselves, and work on every photograph — a JPEG, a body with no profile. They lean on
# vibrance before saturation: vibrance lifts muted colours most and holds
# skin back, so a frame gets richer before anything in it looks painted.
# Saturation, which moves every colour alike, is used sparingly on top.
#
# Each changes only what it names (FR-DEV-6), so a corrected exposure or
# white balance survives applying one.
#
# How much colour each adds is measured, not guessed: mean CIELAB chroma on
# raws rendered with the default (DNG reference) rendering, as a ratio to that
# rendering. For scale, the photographer's earlier exports of the same kind of
# raws sit at 1.14 with no look applied and 1.27 with their everyday look.
# Vivid 1.30 and Vivid warm 1.30 sit just above that; Vivid landscape 1.38;
# Vivid, strong 1.45; Vivid portrait 1.15, with its skin bands held down as
# written. Tuned by scaling each preset's colour values together, never its
# tone ones.
[preset Vivid]
contrast.contrast = 10
saturation.saturation = 11
vibrance.vibrance = 42
[preset Vivid, strong]
blacks_whites.blacks = -10
clarity.amount = 8
contrast.contrast = 18
saturation.saturation = 19
vibrance.vibrance = 57
# Foliage and sky: green and chartreuse for leaves and grass, azure and blue
# for sky and water, a little yellow for dry grass and stone. The skin bands
# — red and orange — are left where they are, so a figure in a landscape
# keeps a human complexion.
[preset Vivid landscape]
clarity.amount = 10
colour_mixer.azure_lum = -10
colour_mixer.azure_sat = 25
colour_mixer.blue_lum = -10
colour_mixer.blue_sat = 20
colour_mixer.chartreuse_sat = 20
colour_mixer.green_sat = 25
colour_mixer.yellow_sat = 13
contrast.contrast = 12
saturation.saturation = 7
vibrance.vibrance = 32
# Golden hour: oranges and yellows up and a warm cast laid over the
# highlights only, so shadows stay clean rather than muddy.
[preset Vivid warm]
colour_grading.highlight_hue = 45
colour_grading.highlight_strength = 12
colour_mixer.orange_sat = 17
colour_mixer.red_sat = 9
colour_mixer.yellow_sat = 17
contrast.contrast = 8
vibrance.vibrance = 29
# People: everything around the subject gets richer while skin does not.
# Vibrance already protects skin; the orange and red bands are then held a
# little below where they started, because a face is the one colour every
# viewer knows the right value of.
[preset Vivid portrait]
colour_mixer.azure_sat = 14
colour_mixer.blue_sat = 17
colour_mixer.green_sat = 17
colour_mixer.orange_sat = -10
colour_mixer.red_sat = -5
contrast.contrast = 6
saturation.saturation = -5
vibrance.vibrance = 35
+1
View File
@@ -65,6 +65,7 @@ const SECTIONS: &[(&str, &str, &str)] = &[
include_str!("../presets/essentials.drpl"),
),
("skies", "Skies", include_str!("../presets/skies.drpl")),
("vivid", "Vivid", include_str!("../presets/vivid.drpl")),
(
"colour_film",
"Film/Colour",
+168
View File
@@ -0,0 +1,168 @@
//! TRACES: FR-DEV-3j | FR-DEV-3e
//! The DNG SDK's reference tone, as a rendering the view transform can
//! choose (D21).
//!
//! The DNG specification's reference rendering runs a raw through the
//! profile's `ProfileToneCurve`, or the ACR3 default for a profile with none.
//! Half of what the curve does is *how* it is applied. The SDK's
//! `RefBaselineRGBTone` runs it on the largest and the smallest channel, and
//! places the middle channel at the fraction between them it had before. Hue
//! is kept; saturation rises wherever the curve is steeper than the
//! diagonal, which for the ACR3 curve is the shadows and the midtones.
//!
//! It runs in linear ProPhoto, as the SDK does, on values clipped to
//! `[0, 1]`; its output is linear and goes to the output transform as the
//! sigmoid's does. The curve is read from the profile buffer
//! (`ops::camera_profile::profile_buffer`), which always carries one.
//!
//! [`apply_reference`] is the arithmetic on the CPU; the GPU test holds the
//! shader to it.
use crate::ops::camera_profile::{mul, working_prophoto};
use crate::view::{DEFAULT_WHITE, REFERENCE_CONTRAST, SCENE_GREY};
/// The input scale for a white point: 1 at the default, so sensor white is
/// display white as in the SDK's reference; each stop of `white` above it halves the
/// input.
pub fn input_scale(white: f32) -> f32 {
(DEFAULT_WHITE - white).exp2()
}
/// The power the input is bent by about middle grey: 1 at
/// [`REFERENCE_CONTRAST`], where the curve is the reference's untouched.
///
/// The default contrast sits above it, so a photograph out of the camera is
/// bent by `DEFAULT_CONTRAST / REFERENCE_CONTRAST` — the extra contrast
/// Lightroom's exports showed over the bare reference curve (D21 addendum).
pub fn contrast_power(contrast: f32) -> f32 {
contrast / REFERENCE_CONTRAST
}
/// The curve, its scale and its contrast applied to one ProPhoto colour.
fn rgb_tone(curve: &[f32], p: [f32; 3]) -> [f32; 3] {
let p = p.map(|v| v.clamp(0.0, 1.0));
let hi = p[0].max(p[1]).max(p[2]);
let lo = p[0].min(p[1]).min(p[2]);
let (c_hi, c_lo) = (
dr_types::tone::evaluate(curve, hi),
dr_types::tone::evaluate(curve, lo),
);
if hi - lo <= 1e-7 {
return [c_hi; 3];
}
p.map(|v| c_lo + (c_hi - c_lo) * (v - lo) / (hi - lo))
}
/// TRACES: FR-DEV-3j
/// The view transform's DNG reference rendering of one working-space colour.
pub fn apply_reference(curve: &[f32], c: [f32; 3], contrast: f32, white: f32) -> [f32; 3] {
let (to, back) = working_prophoto();
let scale = input_scale(white);
let power = contrast_power(contrast);
let mut p = mul(to, c).map(|v| v * scale);
if power != 1.0 {
p = p.map(|v| SCENE_GREY * (v.max(0.0) / SCENE_GREY).powf(power));
}
mul(back, rgb_tone(curve, p))
}
/// The WGSL, a helper the view transform asks for after
/// `ops::camera_profile`'s ProPhoto constants. Mirrors [`apply_reference`].
pub const CAMERA_RAW_WGSL: &str = "
fn camera_raw_curve(x: f32) -> f32 {
let base = profile_curve_base();
let n = u32(profile_table[2].x);
let s = clamp(x, 0.0, 1.0) * f32(n - 1u);
let i = min(u32(s), n - 2u);
return mix(profile_table[base + i].x, profile_table[base + i + 1u].x, s - f32(i));
}
// The SDK's RGBTone: the curve on the largest and smallest channel, the
// middle one kept at its fraction between them, so hue survives.
fn camera_raw_tone(c: vec3<f32>, scale: f32, power: f32, grey: f32) -> vec3<f32> {
var p = PROFILE_FROM_WORKING * c * scale;
if (power != 1.0) {
p = grey * pow(max(p, vec3<f32>(0.0)) / grey, vec3<f32>(power));
}
p = clamp(p, vec3<f32>(0.0), vec3<f32>(1.0));
let hi = max(p.r, max(p.g, p.b));
let lo = min(p.r, min(p.g, p.b));
let c_hi = camera_raw_curve(hi);
let c_lo = camera_raw_curve(lo);
var out = vec3<f32>(c_hi);
if (hi - lo > 1e-7) {
out = vec3<f32>(c_lo) + (c_hi - c_lo) * (p - vec3<f32>(lo)) / (hi - lo);
}
return PROFILE_TO_WORKING * out;
}
";
#[cfg(test)]
mod tests {
use super::*;
use crate::view::DEFAULT_CONTRAST;
use dr_types::tone::{evaluate, ACR3_DEFAULT};
fn identity() -> Vec<f32> {
(0..1025).map(|i| i as f32 / 1024.0).collect()
}
#[test]
fn at_the_reference_the_input_is_untouched() {
assert_eq!(input_scale(DEFAULT_WHITE), 1.0);
assert_eq!(contrast_power(REFERENCE_CONTRAST), 1.0);
}
#[test]
fn the_default_adds_the_measured_contrast() {
// Fitted on Lightroom exports with neutral settings (D21 addendum):
// the bare reference curve is a little flat against them.
let p = contrast_power(DEFAULT_CONTRAST);
assert!((1.05..1.12).contains(&p), "{p}");
}
#[test]
fn grey_goes_through_the_curve_and_stays_grey() {
for v in [0.02, 0.13, 0.5] {
let out = apply_reference(&ACR3_DEFAULT, [v; 3], REFERENCE_CONTRAST, DEFAULT_WHITE);
let want = evaluate(&ACR3_DEFAULT, v);
assert!(
out.iter().all(|o| (o - want).abs() < 1e-4),
"{v}: {out:?} vs {want}"
);
}
}
#[test]
fn an_identity_curve_changes_nothing_inside_the_range() {
let c = [0.4, 0.2, 0.1];
let out = apply_reference(&identity(), c, REFERENCE_CONTRAST, DEFAULT_WHITE);
assert!(
out.iter().zip(c).all(|(o, c)| (o - c).abs() < 1e-4),
"{out:?}"
);
}
#[test]
fn the_middle_channel_keeps_its_place_between_the_other_two() {
let p = [0.3, 0.12, 0.05];
let out = rgb_tone(&ACR3_DEFAULT, p);
let before = (p[1] - p[2]) / (p[0] - p[2]);
let after = (out[1] - out[2]) / (out[0] - out[2]);
assert!((before - after).abs() < 1e-5, "{before} {after}");
}
#[test]
fn the_acr_curve_raises_saturation_in_the_midtones() {
let p = [0.15, 0.08, 0.05];
let out = rgb_tone(&ACR3_DEFAULT, p);
let sat = |c: [f32; 3]| (c[0] - c[2]) / c[0];
assert!(sat(out) > sat(p), "{p:?} -> {out:?}");
}
#[test]
fn white_halves_the_input_per_stop() {
assert_eq!(input_scale(DEFAULT_WHITE + 1.0), 0.5);
assert!(contrast_power(2.8) > 1.0);
}
}
+9
View File
@@ -464,6 +464,15 @@ impl ParamDescriptor {
}
}
/// The same choice with another variant as its default.
///
/// For a choice whose variants were numbered before its default was
/// settled: a sidecar records the index, so reordering the variants to
/// put the default first would change what saved edits mean.
pub fn with_default(self, default: f32) -> Self {
Self { default, ..self }
}
/// A 0…1 fraction — a proportion of something, rather than an amount.
///
/// Its own constructor because the crop rect needs four of them and the
+242 -9
View File
@@ -170,6 +170,18 @@ pub struct EditGraph {
/// correction the photograph asked for — see
/// [`crate::descriptor::ParamDescriptor::switch_on`].
lens_profile_applied: bool,
/// TRACES: FR-DEV-3g
/// Whether this photograph can take the learned denoise — a Bayer
/// mosaic — set by whoever opened it. Derived from the file like the
/// lens profile, so not in the state; it only decides whether the
/// switch below is offered.
denoise_available: bool,
/// Which demosaic develops the photograph: a network, or the classical
/// one. An edit: published as [`crate::learned_denoise`], captured,
/// stored and undone with the rest (FR-DEV-3c).
denoise_method: crate::learned_denoise::Method,
/// How strongly to denoise, 0–100; what is not taken goes back as grain.
denoise_strength: f32,
}
/// TRACES: FR-DEV-3f
@@ -226,6 +238,9 @@ impl EditGraph {
],
lens_profile: None,
lens_profile_applied: true,
denoise_available: false,
denoise_method: crate::learned_denoise::Method::DEFAULT,
denoise_strength: 100.0,
}
}
@@ -399,6 +414,32 @@ impl EditGraph {
self.lens_profile.as_ref()
}
/// TRACES: FR-DEV-3g
/// Offer the learned denoise, or not: true for a Bayer mosaic.
pub fn set_denoise_available(&mut self, available: bool) {
self.denoise_available = available;
}
/// TRACES: FR-DEV-3g
/// Whether the learned denoise is asked for. A setting kept on a
/// photograph that cannot take it is harmless and does nothing, as a
/// lens switch with no profile does.
pub fn denoise_applied(&self) -> bool {
self.denoise_method.learned()
}
/// TRACES: FR-DEV-3g
/// Which demosaic is asked for.
pub fn denoise_method(&self) -> crate::learned_denoise::Method {
self.denoise_method
}
/// TRACES: FR-DEV-3g
/// The grain to keep, 0–1: what the strength does not take.
pub fn denoise_grain(&self) -> f32 {
(100.0 - self.denoise_strength) / 100.0
}
/// TRACES: FR-DEV-3
/// Whether the matched profile is being applied.
pub fn lens_profile_applied(&self) -> bool {
@@ -581,8 +622,38 @@ impl EditGraph {
}
});
switch
// TRACES: FR-DEV-3g
// Offered only where the photograph can take it, for the lens
// switch's reason: a control that can do nothing must not look as if
// it could.
let denoise = self.denoise_available.then(|| {
let desc = crate::learned_denoise::descriptor();
OpCapability {
id: desc.id,
label: desc.label,
active: self.denoise_applied(),
params: desc
.params
.iter()
.map(|p| ParamCapability {
id: p.id,
label: p.label,
kind: p.kind.clone(),
default: p.default,
value: self.param(desc.id, p.id).unwrap_or(p.default),
facet: p.facet,
})
.collect(),
presentation: None,
attributes: desc.attributes.clone(),
}
});
// The learned denoise first: it decides what every control below
// is applied to, so it heads the panel (docs/dev/denoise.md §7).
denoise
.into_iter()
.chain(switch)
.chain(warps)
.chain(ops)
.chain(std::iter::once(framing))
@@ -675,6 +746,11 @@ impl EditGraph {
// `capabilities`, with the operations and the warps and for the
// same reason (FR-DEV-3c).
lens_profile_applied: _,
// Derived from the file, like the profile above.
denoise_available: _,
// Edits, in the state through `capabilities` like the lens switch.
denoise_method: _,
denoise_strength: _,
masks,
film,
spots,
@@ -746,6 +822,32 @@ impl EditGraph {
}
pub fn set_param(&mut self, op: OpId, param: ParamId, value: f32) {
if op == crate::learned_denoise::ID {
match param {
p if p == crate::learned_denoise::METHOD => {
self.denoise_method = crate::learned_denoise::Method::from_index(value)
}
// 0.21 and 0.22's switch (see `APPLY`): off is the classical
// demosaic, on is a network — the one already chosen, if any.
p if p == crate::learned_denoise::APPLY => {
use crate::learned_denoise::Method;
if value == 0.0 {
self.denoise_method = Method::Bilinear;
} else if !self.denoise_method.learned() {
self.denoise_method = Method::DEFAULT;
}
}
p if p == crate::learned_denoise::STRENGTH => {
self.denoise_strength = value.clamp(0.0, 100.0)
}
// 0.21.0's grain, the strength's inverse (see `GRAIN`).
p if p == crate::learned_denoise::GRAIN => {
self.denoise_strength = 100.0 - value.clamp(0.0, 100.0)
}
_ => log::warn!("unknown parameter {param} on {op}; ignoring"),
}
return;
}
if op == crate::lens::profile_switch::ID {
if param != crate::lens::profile_switch::APPLY {
log::warn!("unknown parameter {param} on {op}; ignoring");
@@ -803,6 +905,17 @@ impl EditGraph {
/// Read a parameter back.
pub fn param(&self, op: OpId, param: ParamId) -> Option<f32> {
if op == crate::learned_denoise::ID {
return match param {
p if p == crate::learned_denoise::METHOD => Some(self.denoise_method.index()),
p if p == crate::learned_denoise::APPLY => {
Some(if self.denoise_applied() { 1.0 } else { 0.0 })
}
p if p == crate::learned_denoise::STRENGTH => Some(self.denoise_strength),
p if p == crate::learned_denoise::GRAIN => Some(100.0 - self.denoise_strength),
_ => None,
};
}
if op == crate::lens::profile_switch::ID {
return (param == crate::lens::profile_switch::APPLY)
.then_some(if self.lens_profile_applied { 1.0 } else { 0.0 });
@@ -849,6 +962,10 @@ impl EditGraph {
// a reset does not change which lens took the photograph. What returns
// to default is the answer to whether to use it, which is on.
self.set_lens_profile_applied(true);
// The learned denoise returns to its default network; whether it is
// available is the file's and stays.
self.denoise_method = crate::learned_denoise::Method::DEFAULT;
self.denoise_strength = 100.0;
}
/// Set the crop rectangle. Clamped to keep it inside the frame.
@@ -1170,19 +1287,21 @@ mod tests {
// Opening an unedited image must produce the image, not an
// interpretation of it.
//
// One block, and it is the view transform: a view operation is
// composed at its defaults, because a photograph with no view
// transform is a scan rather than a picture (FR-DEV-3j). It is still
// neutral in the sense that matters here — nothing moved, nothing is
// written — and every adjustment is absent.
// Two blocks, the view transform and the camera profile: both are
// composed at their defaults, because a photograph with no view
// transform is a scan rather than a picture (FR-DEV-3j) and a raw
// with a profile is rendered through it (D20). They are still neutral
// in the sense that matters here — nothing moved, nothing is written
// — and every adjustment is absent.
let g = EditGraph::default_chain();
assert!(g.is_neutral());
let source = g.compose().source;
assert_eq!(
source.matches("---- ").count(),
1,
2,
"a neutral graph must generate no adjustment blocks"
);
assert!(source.contains("---- camera_profile ----"));
assert!(source.contains("---- view_transform ----"));
}
@@ -1263,13 +1382,14 @@ mod tests {
fn only_active_operations_reach_the_shader() {
// The composition property, end to end: two adjustments out of seven
// available must generate a shader doing exactly two things — and
// the view transform, which every render has (FR-DEV-3j).
// the view transform and camera profile, which every render has
// (FR-DEV-3j, D20).
let mut g = EditGraph::default_chain();
g.set_param(exposure::ID, exposure::EXPOSURE, 1.0);
g.set_param(white_balance::ID, white_balance::TINT, 25.0);
let shader = g.compose();
assert_eq!(shader.source.matches("---- ").count(), 3);
assert_eq!(shader.source.matches("---- ").count(), 4);
assert!(shader.source.contains("---- view_transform ----"));
assert!(shader.source.contains("---- exposure ----"));
assert!(shader.source.contains("---- white_balance ----"));
@@ -1992,4 +2112,117 @@ mod tests {
let after = cropped.render_scale(source, (1500, 1000));
assert!(after.ratio() > fit.ratio());
}
#[test]
fn the_learned_denoise_is_offered_only_where_it_can_run() {
use crate::learned_denoise;
let mut g = EditGraph::default_chain();
assert!(!g.capabilities().iter().any(|c| c.id == learned_denoise::ID));
g.set_denoise_available(true);
let cap = g
.capabilities()
.into_iter()
.find(|c| c.id == learned_denoise::ID)
.expect("offered");
assert!(cap.active, "on by default");
assert_eq!(g.denoise_method(), learned_denoise::Method::Best);
assert_eq!(g.denoise_grain(), 0.0, "at full strength");
assert_eq!(cap.id, g.capabilities()[0].id, "and first in the panel");
g.set_param(
learned_denoise::ID,
learned_denoise::METHOD,
learned_denoise::Method::Bilinear.index(),
);
g.set_param(learned_denoise::ID, learned_denoise::STRENGTH, 70.0);
assert!(!g.denoise_applied());
assert!((g.denoise_grain() - 0.3).abs() < 1e-6);
g.reset();
assert!(g.denoise_applied(), "reset is back to on");
assert_eq!(g.denoise_grain(), 0.0);
}
#[test]
fn an_untouched_raw_writes_nothing_and_develops_through_the_best() {
// TRACES: FR-DEV-3g
use crate::learned_denoise::{self, Method};
let mut g = EditGraph::default_chain();
g.set_denoise_available(true);
assert_eq!(g.denoise_method(), Method::Best);
let stored = |g: &EditGraph| {
crate::Preset::capture_params(g)
.params()
.keys()
.any(|(op, _)| op == learned_denoise::ID.0)
};
assert!(!stored(&g), "the default is not written");
g.set_param(
learned_denoise::ID,
learned_denoise::METHOD,
Method::Fast.index(),
);
assert!(stored(&g), "a choice is");
assert!(
!crate::Preset::capture_params(&g).params().contains_key(&(
learned_denoise::ID.0.into(),
learned_denoise::APPLY.0.into()
)),
"and the old switch never is"
);
}
#[test]
fn an_edit_saved_with_the_switch_keeps_its_look() {
// TRACES: FR-DEV-3g
// 0.21 and 0.22 stored on or off; off is the classical demosaic, and
// on keeps a network already chosen.
use crate::learned_denoise::{self, Method};
let mut g = EditGraph::default_chain();
g.set_param(learned_denoise::ID, learned_denoise::APPLY, 0.0);
assert_eq!(g.denoise_method(), Method::Bilinear);
g.set_param(learned_denoise::ID, learned_denoise::APPLY, 1.0);
assert_eq!(g.denoise_method(), Method::DEFAULT);
g.set_param(
learned_denoise::ID,
learned_denoise::METHOD,
Method::Fast.index(),
);
g.set_param(learned_denoise::ID, learned_denoise::APPLY, 1.0);
assert_eq!(g.denoise_method(), Method::Fast);
// A number from a newer build with more methods is the default.
g.set_param(learned_denoise::ID, learned_denoise::METHOD, 9.0);
assert_eq!(g.denoise_method(), Method::DEFAULT);
}
#[test]
fn an_edit_saved_with_grain_keeps_its_look() {
// TRACES: FR-DEV-3g
// 0.21.0 stored the grain kept rather than the strength.
use crate::learned_denoise;
let mut g = EditGraph::default_chain();
g.set_param(learned_denoise::ID, learned_denoise::GRAIN, 25.0);
assert_eq!(
g.param(learned_denoise::ID, learned_denoise::STRENGTH),
Some(75.0)
);
assert!((g.denoise_grain() - 0.25).abs() < 1e-6);
}
#[test]
fn the_learned_denoise_travels_in_the_state() {
use crate::learned_denoise;
let mut g = EditGraph::default_chain();
g.set_denoise_available(true);
g.set_param(learned_denoise::ID, learned_denoise::STRENGTH, 60.0);
g.set_param(
learned_denoise::ID,
learned_denoise::METHOD,
learned_denoise::Method::Fast.index(),
);
let state = g.state();
let mut h = EditGraph::default_chain();
h.set_denoise_available(true);
let _ = h.set_state(&state);
assert_eq!(h.denoise_method(), learned_denoise::Method::Fast);
assert!((h.denoise_grain() - 0.4).abs() < 1e-6);
}
}
+139
View File
@@ -0,0 +1,139 @@
//! TRACES: FR-DEV-3g
//! The learned denoise's settings: which network develops the photograph,
//! if any, and how much grain to keep.
//!
//! Not an [`crate::operation::Operation`]: the learned stage replaces the
//! demosaic and runs once per photograph, off the render path
//! (docs/dev/denoise.md §2, §7), and the grain is a blend of its result with
//! the classical one, done where the source is chosen. But what a
//! photographer sets travels the one road every setting travels — the
//! capability list feeds the panel, [`crate::Preset`] captures it, the
//! sidecar stores it, the undo stack replays it (FR-DEV-3c) — so it is
//! published as a capability, like the lens profile switch.
use std::sync::{Arc, LazyLock};
use crate::descriptor::{Attribute, LocalizedKey, OpDescriptor, ParamDescriptor, Scale, Unit};
use crate::{OpId, ParamId};
pub const ID: OpId = OpId("learned_denoise");
/// TRACES: FR-DEV-3g
/// Which demosaic develops the photograph, a [`Method`] by index.
pub const METHOD: ParamId = ParamId("method");
/// What 0.21 and 0.22 stored instead of [`METHOD`]: on or off. Still read —
/// off is [`Method::Bilinear`], on is the default network — so an edit saved
/// by those releases keeps its look; never written, and not offered.
pub const APPLY: ParamId = ParamId("apply");
/// TRACES: FR-DEV-3g
/// How strongly to denoise, 0–100: 100 is the network's result as it is, and
/// lower puts the removed noise's brightness back as grain.
pub const STRENGTH: ParamId = ParamId("strength");
/// What 0.21.0 stored instead of [`STRENGTH`]: the grain kept, its inverse.
/// Still read, so an edit saved by that release keeps its look; never
/// written, and not offered as a control.
pub const GRAIN: ParamId = ParamId("grain");
/// TRACES: FR-DEV-3g
/// The demosaics a photograph can be developed with, in the order the
/// sidecar numbers them. Two networks that trade time for quality
/// (docs/dev/denoise.md §15) and the classical demosaic, which is no network
/// at all.
///
/// Until 0.24 there were four — Bilinear, Fast, Medium, Best — and the
/// sidecar keeps their numbers: 2, which was Medium, is now Best, and 3,
/// which was Best, is past the end and reads as the default, which is
/// Best. Both land on the network that replaced them, with no migration.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)]
pub enum Method {
/// The classical demosaic: the noise stays.
Bilinear,
/// The smallest student: a quarter of Best's work.
Fast,
/// One network of the first release's size, taught by the mixture of
/// experts it replaced: the mixture's edges at a third of its work.
Best,
}
impl Method {
pub const ALL: [Method; 3] = [Method::Bilinear, Method::Fast, Method::Best];
pub const DEFAULT: Method = Method::Best;
/// The sidecar's number for it.
pub fn index(self) -> f32 {
Self::ALL.iter().position(|m| *m == self).unwrap_or(0) as f32
}
/// The method a stored number names; out of range is the default, as
/// from a newer build with more of them.
pub fn from_index(value: f32) -> Method {
let i = value.round();
if i >= 0.0 && (i as usize) < Self::ALL.len() {
Self::ALL[i as usize]
} else {
Self::DEFAULT
}
}
/// Whether a network runs at all.
pub fn learned(self) -> bool {
self != Method::Bilinear
}
}
/// The best network by default, at full strength: every Bayer raw is
/// developed from the learned demosaic, and the choice and the slider are
/// there to take it back, trade it for time, or ease it off. It costs seconds per photograph the first time, while
/// the classical demosaic shows; the result is cached, so a photograph
/// reopened or exported does not pay again (docs/dev/denoise.md §7).
pub(crate) static DESCRIPTOR: LazyLock<Arc<OpDescriptor>> = LazyLock::new(|| {
Arc::new(OpDescriptor {
id: ID,
label: LocalizedKey("op.learned_denoise"),
params: vec![
ParamDescriptor::choice(
"method",
"param.learned_denoise.method",
vec![
LocalizedKey("param.learned_denoise.method.bilinear"),
LocalizedKey("param.learned_denoise.method.fast"),
LocalizedKey("param.learned_denoise.method.best"),
],
)
.with_default(Method::DEFAULT.index()),
ParamDescriptor::scalar(
"strength",
"param.learned_denoise.strength",
0.0,
100.0,
100.0,
Unit::Percent,
Scale::Linear,
0,
),
],
// With the classical noise reduction, which is what a photographer
// looks for it beside.
attributes: vec![Attribute::Detail],
})
});
pub fn descriptor() -> Arc<OpDescriptor> {
DESCRIPTOR.clone()
}
#[cfg(test)]
mod tests {
use super::*;
/// An edit saved before 0.24 stored Medium as 2 and Best as 3. Both
/// now name the network that replaced them, and nothing reads as Fast
/// or Bilinear that did not before.
#[test]
fn the_retired_methods_read_as_best() {
assert_eq!(Method::from_index(0.0), Method::Bilinear);
assert_eq!(Method::from_index(1.0), Method::Fast);
assert_eq!(Method::from_index(2.0), Method::Best, "Medium, before 0.24");
assert_eq!(Method::from_index(3.0), Method::Best, "Best, before 0.24");
assert_eq!(Method::Best.index(), 2.0);
}
}
+12
View File
@@ -33,6 +33,7 @@
//! data neither would be physically meaningful (ARCH §5.2).
pub mod bundled;
pub mod camera_raw;
pub mod coverage;
pub mod declared;
pub mod descriptor;
@@ -40,6 +41,7 @@ pub mod detail;
pub mod framing;
pub mod graph;
pub mod history;
pub mod learned_denoise;
pub mod lens;
pub mod mask;
pub mod neutral;
@@ -115,6 +117,16 @@ mod tests {
}
}
// Flipping every switch turns the camera profile *off*, which is
// active — moved from the default — and composes nothing. Put it
// back on; its look strength stays moved, so it is still active and
// now doing something, which is what "fully active" means here (D20).
g.set_param(
crate::ops::camera_profile::ID,
crate::ops::camera_profile::APPLY,
1.0,
);
// `film_sim` is the one node a moved parameter cannot activate: it
// needs a stock's measured tables, which are not parameters and which
// no slider produces. So it is loaded explicitly here.
+20 -2
View File
@@ -270,6 +270,19 @@ pub trait Operation: Send + Sync {
/// what is actually used.
fn is_active(&self) -> bool;
/// TRACES: FR-DEV-3e
/// Whether the composer emits this operation's fragment.
///
/// Default: exactly when it [`Self::is_active`]. The exception is an
/// operation that *is* part of the rendering at its defaults — the
/// camera profile, which an untouched raw is rendered through (D20) —
/// where "moved from the defaults" and "does something" come apart. Such
/// an operation keeps `is_active` meaning the former, so a sidecar still
/// stores nothing for it, and answers this with the latter.
fn composes(&self) -> bool {
self.is_active()
}
/// The WGSL body of this operation's transform.
///
/// Receives `c` (a `vec3<f32>` of linear RGB) and must produce the
@@ -843,7 +856,7 @@ fn compose_inner(
let active: Vec<&dyn Operation> = ops
.iter()
.map(|o| o.as_ref())
.filter(|o| o.is_active() && o.detail().is_none())
.filter(|o| o.composes() && o.detail().is_none())
.collect();
// Whether a detail stage follows. If one does, this pass stops short of
@@ -1040,7 +1053,7 @@ fn compose_inner(
let local: Vec<&crate::mask::LocalOp> = layers.ops.iter().filter(|l| l.op == id).collect();
// The global side of the blend. A view operation always has one; see
// `Stage::View`.
let global = op.is_active() || op.stage() == Stage::View;
let global = op.composes() || op.stage() == Stage::View;
if !global && local.is_empty() {
continue;
}
@@ -1333,6 +1346,11 @@ struct Params {{
// bound to 1x1 placeholders whenever the flags say not to touch them.
@group(0) @binding(6) var sampled: texture_2d<f32>;
@group(0) @binding(7) var sample_out: texture_storage_2d<rgba16float, write>;
// The source's camera profile tables (FR-DEV-3e, D20): a two-entry header,
// then the entries (`ops::camera_profile::profile_buffer`). Declared
// unconditionally like the masks, and bound to a header of zeros — no
// tables — for every source without a profile.
@group(0) @binding(8) var<storage, read> profile_table: array<vec4<f32>>;
{WINDOW_HELPER}{sampler_helper}{helper_src}{encode_output}
// Display-encoded sRGB back to linear, for sources that arrive that way.
+732
View File
@@ -0,0 +1,732 @@
//! TRACES: FR-DEV-3e
//! The camera profile's tables as an operation (D20).
//!
//! The matrix turns camera RGB into colour; a DNG camera profile adds two
//! lookups over hue, saturation and value on top of it — the `HueSatMap`, a
//! calibration, and the `LookTable`, a rendering intent. This operation
//! applies them. `docs/dev/camera-profiles.md` is the design.
//!
//! # Where the tables come from
//!
//! Not from here. They belong to the *source*, like the matrix: `dr-decode`
//! resolves them per file and `dr-gpu` uploads them to the storage buffer
//! every generated shader declares at `@binding(8)`, laid out by
//! [`profile_buffer`]. This operation holds only the photographer's two
//! settings — whether to use the profile, and how strongly to apply its look
//! — so a render path never has to remember to hand it anything.
//!
//! # Why it is composed at its defaults
//!
//! A profile that is on is the rendering, not an edit: an untouched raw
//! renders through it and writes no parameters. So [`Operation::composes`]
//! answers "is the switch on", not "has anything moved". The fragment then
//! branches on the buffer's header, which says whether this source has tables
//! at all; a JPEG, or a raw with no profile, reads two zeros and passes
//! through.
//!
//! # The lookup
//!
//! The DNG SDK's `RefBaselineHueSatMap`, with the two departures §2 of the
//! design gives for scene-referred values: value is not clamped on the way
//! out, and a colour with a negative ProPhoto component passes through.
//! [`apply_reference`] is the same arithmetic on the CPU, and the GPU tests
//! hold the shader to it.
use std::sync::{Arc, LazyLock};
use dr_types::{HueSatTable, ProfileTables};
use crate::descriptor::{
Attribute, LocalizedKey, OpDescriptor, OpId, ParamDescriptor, ParamId, Scale, Unit,
};
use crate::operation::{Helper, Operation, Uniform};
pub const ID: OpId = OpId("camera_profile");
pub const APPLY: ParamId = ParamId("apply");
pub const LOOK: ParamId = ParamId("look");
/// The look's strength at which the LookTable is applied as the profile
/// states it, in percent.
pub const PROFILE_LOOK: f32 = 100.0;
/// The look's default strength: off.
///
/// Measured, not chosen. Against the photographer's earlier exports with no
/// look applied, the default rendering scores the same with the table at 100,
/// 50 or 0 (held-out MSE 140, 140, 143), and is 9 % more colourful without
/// it: the table desaturates near-neutral tones, which is exactly where the
/// default rendering was short of those exports. The table stays one slider
/// away for anyone who wants the profile's look.
pub const DEFAULT_LOOK: f32 = 0.0;
/// Twice the profile's look.
pub const MAX_LOOK: f32 = 2.0 * PROFILE_LOOK;
/// Entries of the buffer's header, before the entries themselves: one
/// `vec4` describing each table — `(hue divisions, saturation divisions,
/// value divisions, sRGB-encoded)`, zero hue divisions meaning absent — and
/// a third whose `.x` is the tone curve's length (camera-profiles.md §12).
pub const HEADER_ENTRIES: usize = 3;
static DESCRIPTOR: LazyLock<Arc<OpDescriptor>> = LazyLock::new(|| {
Arc::new(OpDescriptor {
attributes: vec![Attribute::Colour],
id: ID,
label: LocalizedKey("op.camera_profile"),
params: vec![
ParamDescriptor::switch_on("apply", "param.camera_profile.apply"),
ParamDescriptor::scalar(
"look",
"param.camera_profile.look",
0.0,
MAX_LOOK,
DEFAULT_LOOK,
Unit::Percent,
Scale::Linear,
0,
),
],
})
});
/// Linear sRGB (the working space) to linear ProPhoto, and back, row-major,
/// each row scaled to sum to one so that working white is ProPhoto white
/// exactly and a neutral reaches the tables with zero saturation.
pub(crate) fn working_prophoto() -> &'static ([f32; 9], [f32; 9]) {
static M: LazyLock<([f32; 9], [f32; 9])> = LazyLock::new(|| {
let to = normalise_rows(dr_types::ColourSpace::ProPhoto.from_linear_srgb());
let back = normalise_rows(invert(&to).expect("ProPhoto's matrix is invertible"));
(to, back)
});
&M
}
fn normalise_rows(mut m: [f32; 9]) -> [f32; 9] {
for row in m.chunks_exact_mut(3) {
let sum: f32 = row.iter().sum();
row.iter_mut().for_each(|v| *v /= sum);
}
m
}
fn invert(m: &[f32; 9]) -> Option<[f32; 9]> {
let [a, b, c, d, e, f, g, h, i] = m.map(f64::from);
let det = a * (e * i - f * h) - b * (d * i - f * g) + c * (d * h - e * g);
if det.abs() < 1e-12 {
return None;
}
let inv = [
(e * i - f * h) / det,
(c * h - b * i) / det,
(b * f - c * e) / det,
(f * g - d * i) / det,
(a * i - c * g) / det,
(c * d - a * f) / det,
(d * h - e * g) / det,
(b * g - a * h) / det,
(a * e - b * d) / det,
];
Some(inv.map(|v| v as f32))
}
pub(crate) fn mul(m: &[f32; 9], c: [f32; 3]) -> [f32; 3] {
std::array::from_fn(|r| m[r * 3] * c[0] + m[r * 3 + 1] * c[1] + m[r * 3 + 2] * c[2])
}
/// A row-major matrix as a WGSL `mat3x3`, whose constructor takes columns.
fn wgsl_mat(m: &[f32; 9]) -> String {
let col = |j: usize| format!("vec3<f32>({:e}, {:e}, {:e})", m[j], m[3 + j], m[6 + j]);
format!("mat3x3<f32>({}, {}, {})", col(0), col(1), col(2))
}
/// The working space to ProPhoto and back, as WGSL constants, and where
/// the profile buffer's sections begin. A helper of its own because the
/// view transform's DNG reference curve needs it too, and helpers are emitted
/// once each, in the order first asked for.
pub(crate) static PROPHOTO_HELPER: LazyLock<Helper> = LazyLock::new(|| {
let (to, back) = working_prophoto();
let source = format!(
"const PROFILE_FROM_WORKING = {};\nconst PROFILE_TO_WORKING = {};\n{SECTIONS_WGSL}",
wgsl_mat(to),
wgsl_mat(back)
);
Helper {
name: "profile_curve_base",
source: Box::leak(source.into_boxed_str()),
}
});
static HELPERS: LazyLock<[Helper; 2]> = LazyLock::new(|| {
[
*PROPHOTO_HELPER,
Helper {
name: "profile_apply",
source: LOOKUP_WGSL,
},
]
});
/// Where each section of the profile buffer starts, from its header.
const SECTIONS_WGSL: &str = "
fn profile_entries(dims: vec4<f32>) -> u32 {
return u32(dims.x * dims.y * dims.z);
}
fn profile_look_base() -> u32 {
return 3u + profile_entries(profile_table[0]);
}
fn profile_curve_base() -> u32 {
return profile_look_base() + profile_entries(profile_table[1]);
}
";
/// The lookup, in WGSL. Mirrors [`apply_reference`] line for line.
const LOOKUP_WGSL: &str = r#"
fn profile_srgb_encode(v: f32) -> f32 {
if (v <= 0.0031308) { return v * 12.92; }
return 1.055 * pow(v, 1.0 / 2.4) - 0.055;
}
fn profile_srgb_decode(v: f32) -> f32 {
if (v <= 0.04045) { return v / 12.92; }
return pow((v + 0.055) / 1.055, 2.4);
}
// The DNG SDK's HSV: hue in [0, 6), saturation (max - min) / max, value max.
fn profile_rgb_to_hsv(c: vec3<f32>) -> vec3<f32> {
let v = max(c.r, max(c.g, c.b));
let gap = v - min(c.r, min(c.g, c.b));
if (gap <= 0.0) {
return vec3<f32>(0.0, 0.0, v);
}
var h: f32;
if (c.r == v) {
h = (c.g - c.b) / gap;
if (h < 0.0) { h += 6.0; }
} else if (c.g == v) {
h = 2.0 + (c.b - c.r) / gap;
} else {
h = 4.0 + (c.r - c.g) / gap;
}
return vec3<f32>(h, gap / v, v);
}
fn profile_hsv_to_rgb(hsv: vec3<f32>) -> vec3<f32> {
let s = hsv.y;
let v = hsv.z;
if (s <= 0.0) {
return vec3<f32>(v);
}
let h = hsv.x - 6.0 * floor(hsv.x / 6.0);
let i = min(floor(h), 5.0);
let f = h - i;
let p = v * (1.0 - s);
let q = v * (1.0 - s * f);
let t = v * (1.0 - s * (1.0 - f));
switch (i32(i)) {
case 0: { return vec3<f32>(v, t, p); }
case 1: { return vec3<f32>(q, v, p); }
case 2: { return vec3<f32>(p, v, t); }
case 3: { return vec3<f32>(p, q, v); }
case 4: { return vec3<f32>(t, p, v); }
default: { return vec3<f32>(v, p, q); }
}
}
fn profile_entry(base: u32, at: u32) -> vec3<f32> {
return profile_table[base + at].xyz;
}
// (hue shift in degrees, saturation scale, value scale) at `hsv`: bilinear
// over hue and saturation, hue wrapping, and linear over value for a 3-D
// table. Indices are the SDK's.
fn profile_lookup(dims: vec4<f32>, base: u32, hsv: vec3<f32>) -> vec3<f32> {
let hd = u32(dims.x);
let sd = u32(dims.y);
let vd = u32(dims.z);
var h0 = 0u;
var h1 = 0u;
var hf = 0.0;
if (hd > 1u) {
let hs = hsv.x * f32(hd) / 6.0;
h0 = min(u32(hs), hd - 1u);
hf = hs - f32(h0);
h1 = h0 + 1u;
if (h1 >= hd) { h1 = 0u; }
}
let ss = hsv.y * f32(sd - 1u);
let s0 = min(u32(ss), sd - 2u);
let sf = ss - f32(s0);
var v0 = 0u;
var vf = 0.0;
if (vd > 1u) {
var ve = clamp(hsv.z, 0.0, 1.0);
if (dims.w > 0.5) { ve = profile_srgb_encode(ve); }
let vs = ve * f32(vd - 1u);
v0 = min(u32(vs), vd - 2u);
vf = vs - f32(v0);
}
let val_step = hd * sd;
let lo = v0 * val_step;
var d = mix(
mix(profile_entry(base, lo + h0 * sd + s0), profile_entry(base, lo + h1 * sd + s0), hf),
mix(profile_entry(base, lo + h0 * sd + s0 + 1u), profile_entry(base, lo + h1 * sd + s0 + 1u), hf),
sf);
if (vd > 1u) {
let hi = lo + val_step;
let e = mix(
mix(profile_entry(base, hi + h0 * sd + s0), profile_entry(base, hi + h1 * sd + s0), hf),
mix(profile_entry(base, hi + h0 * sd + s0 + 1u), profile_entry(base, hi + h1 * sd + s0 + 1u), hf),
sf);
d = mix(d, e, vf);
}
return d;
}
// One table applied to a ProPhoto colour, its deltas scaled by `amount`.
fn profile_apply(dims: vec4<f32>, base: u32, c: vec3<f32>, amount: f32) -> vec3<f32> {
let hsv = profile_rgb_to_hsv(c);
var d = profile_lookup(dims, base, hsv);
d = vec3<f32>(d.x * amount, max(1.0 + (d.y - 1.0) * amount, 0.0), max(1.0 + (d.z - 1.0) * amount, 0.0));
let h = hsv.x + d.x * (6.0 / 360.0);
let s = min(hsv.y * d.y, 1.0);
var v = hsv.z * d.z;
if (dims.w > 0.5) {
// The scale is defined on the encoded value; applied as the ratio it
// makes at min(v, 1), so a value above 1.0 is scaled, not clipped.
let vc = min(hsv.z, 1.0);
v = hsv.z;
if (vc > 0.0) {
v = hsv.z * profile_srgb_decode(profile_srgb_encode(vc) * d.z) / vc;
}
}
return profile_hsv_to_rgb(vec3<f32>(h, s, v));
}
"#;
#[derive(Debug, Clone)]
pub struct CameraProfile {
apply: bool,
look: f32,
}
impl Default for CameraProfile {
fn default() -> Self {
Self {
apply: true,
look: DEFAULT_LOOK,
}
}
}
impl CameraProfile {
pub fn new() -> Self {
Self::default()
}
}
impl Operation for CameraProfile {
fn descriptor(&self) -> Arc<OpDescriptor> {
DESCRIPTOR.clone()
}
fn set_param(&mut self, id: ParamId, value: f32) {
match id {
APPLY => self.apply = value != 0.0,
LOOK => self.look = value,
_ => log::warn!("camera_profile: unknown parameter {id}"),
}
}
fn param(&self, id: ParamId) -> f32 {
match id {
APPLY => f32::from(u8::from(self.apply)),
LOOK => self.look,
_ => 0.0,
}
}
fn is_active(&self) -> bool {
!self.apply || self.look != DEFAULT_LOOK
}
fn composes(&self) -> bool {
self.apply
}
fn wgsl_body(&self) -> String {
"\
let hue_sat_dims = profile_table[0];
let look_dims = profile_table[1];
if (hue_sat_dims.x > 0.0 || look_dims.x > 0.0) {
var p = PROFILE_FROM_WORKING * c;
// A colour outside ProPhoto has no HSV the tables were made for; it
// passes through rather than being floored, which would clip it (D19).
if (min(p.r, min(p.g, p.b)) >= 0.0) {
if (hue_sat_dims.x > 0.0) {
p = profile_apply(hue_sat_dims, 3u, p, 1.0);
}
if (look_dims.x > 0.0 && look > 0.0) {
p = profile_apply(look_dims, profile_look_base(), p, look);
}
c = PROFILE_TO_WORKING * p;
}
}"
.into()
}
fn uniforms(&self) -> Vec<Uniform> {
vec![Uniform {
name: "look",
value: self.look / 100.0,
}]
}
fn helpers(&self) -> &[Helper] {
HELPERS.as_slice()
}
}
/// TRACES: FR-DEV-3e | FR-DEV-3j
/// The storage buffer a source's profile is uploaded as: the three header
/// `vec4`s, the HueSatMap's entries, the LookTable's, each entry
/// `(hue shift, saturation scale, value scale, 0)`, then the tone curve's
/// samples in `.x`.
///
/// The curve is always there: the profile's own where it has one, Camera
/// Raw's ACR3 default otherwise — including in the placeholder every source
/// without a profile binds, whose tables are absent, so a raw with no
/// profile still has the reference tone curve when it is chosen (D21).
pub fn profile_buffer(tables: Option<&ProfileTables>) -> Vec<[f32; 4]> {
let header = |t: Option<&HueSatTable>| match t {
Some(t) => [
t.hue_divisions as f32,
t.sat_divisions as f32,
t.val_divisions as f32,
if t.srgb_encoded { 1.0 } else { 0.0 },
],
None => [0.0; 4],
};
let hue_sat = tables.and_then(|t| t.hue_sat.as_ref());
let look = tables.and_then(|t| t.look.as_ref());
let curve: &[f32] = tables
.and_then(|t| t.tone_curve.as_deref())
.unwrap_or(&dr_types::tone::ACR3_DEFAULT);
let mut out = vec![
header(hue_sat),
header(look),
[curve.len() as f32, 0.0, 0.0, 0.0],
];
for t in [hue_sat, look].into_iter().flatten() {
out.extend(t.entries.iter().map(|e| [e[0], e[1], e[2], 0.0]));
}
out.extend(curve.iter().map(|&v| [v, 0.0, 0.0, 0.0]));
out
}
/// TRACES: FR-DEV-3e
/// The fragment's arithmetic on the CPU: a working-space colour through the
/// source's tables, the look at `look` (1.0 = as the profile states it).
///
/// The reference the shader is tested against, and the statement of the
/// algorithm a reader can step through.
pub fn apply_reference(tables: &ProfileTables, c: [f32; 3], look: f32) -> [f32; 3] {
let (to, back) = working_prophoto();
let mut p = mul(to, c);
if p.iter().any(|v| *v < 0.0) {
return c;
}
if let Some(t) = &tables.hue_sat {
p = apply_table(t, p, 1.0);
}
if let Some(t) = tables.look.as_ref().filter(|_| look > 0.0) {
p = apply_table(t, p, look);
}
mul(back, p)
}
fn srgb_encode(v: f32) -> f32 {
if v <= 0.003_130_8 {
v * 12.92
} else {
1.055 * v.powf(1.0 / 2.4) - 0.055
}
}
fn srgb_decode(v: f32) -> f32 {
if v <= 0.040_45 {
v / 12.92
} else {
((v + 0.055) / 1.055).powf(2.4)
}
}
/// The SDK's `DNG_RGBtoHSV`: hue in `[0, 6)`.
pub fn rgb_to_hsv([r, g, b]: [f32; 3]) -> [f32; 3] {
let v = r.max(g).max(b);
let gap = v - r.min(g).min(b);
if gap <= 0.0 {
return [0.0, 0.0, v];
}
let h = if r == v {
let h = (g - b) / gap;
if h < 0.0 {
h + 6.0
} else {
h
}
} else if g == v {
2.0 + (b - r) / gap
} else {
4.0 + (r - g) / gap
};
[h, gap / v, v]
}
pub fn hsv_to_rgb([h, s, v]: [f32; 3]) -> [f32; 3] {
if s <= 0.0 {
return [v; 3];
}
let h = h - 6.0 * (h / 6.0).floor();
let i = h.floor().min(5.0);
let f = h - i;
let p = v * (1.0 - s);
let q = v * (1.0 - s * f);
let t = v * (1.0 - s * (1.0 - f));
match i as i32 {
0 => [v, t, p],
1 => [q, v, p],
2 => [p, v, t],
3 => [p, q, v],
4 => [t, p, v],
_ => [v, p, q],
}
}
fn lookup(t: &HueSatTable, [h, s, v]: [f32; 3]) -> [f32; 3] {
let (hd, sd, vd) = (t.hue_divisions, t.sat_divisions, t.val_divisions);
let (mut h0, mut h1, mut hf) = (0u32, 0u32, 0.0f32);
if hd > 1 {
let hs = h * hd as f32 / 6.0;
h0 = (hs as u32).min(hd - 1);
hf = hs - h0 as f32;
h1 = if h0 + 1 >= hd { 0 } else { h0 + 1 };
}
let ss = s * (sd - 1) as f32;
let s0 = (ss as u32).min(sd - 2);
let sf = ss - s0 as f32;
let (mut v0, mut vf) = (0u32, 0.0f32);
if vd > 1 {
let mut ve = v.clamp(0.0, 1.0);
if t.srgb_encoded {
ve = srgb_encode(ve);
}
let vs = ve * (vd - 1) as f32;
v0 = (vs as u32).min(vd - 2);
vf = vs - v0 as f32;
}
let mix = |a: [f32; 3], b: [f32; 3], w: f32| -> [f32; 3] {
std::array::from_fn(|i| a[i] + (b[i] - a[i]) * w)
};
let at = |v: u32, h: u32, s: u32| t.entries[t.index(h, s, v)];
let plane = |v: u32| {
mix(
mix(at(v, h0, s0), at(v, h1, s0), hf),
mix(at(v, h0, s0 + 1), at(v, h1, s0 + 1), hf),
sf,
)
};
let d = plane(v0);
if vd > 1 {
mix(d, plane(v0 + 1), vf)
} else {
d
}
}
fn apply_table(t: &HueSatTable, c: [f32; 3], amount: f32) -> [f32; 3] {
let hsv = rgb_to_hsv(c);
let d = lookup(t, hsv);
let d = [
d[0] * amount,
(1.0 + (d[1] - 1.0) * amount).max(0.0),
(1.0 + (d[2] - 1.0) * amount).max(0.0),
];
let h = hsv[0] + d[0] * (6.0 / 360.0);
let s = (hsv[1] * d[1]).min(1.0);
let v = if t.srgb_encoded {
let vc = hsv[2].min(1.0);
if vc > 0.0 {
hsv[2] * srgb_decode(srgb_encode(vc) * d[2]) / vc
} else {
hsv[2]
}
} else {
hsv[2] * d[2]
};
hsv_to_rgb([h, s, v])
}
#[cfg(test)]
mod tests {
use super::*;
use dr_types::ProfileOrigin;
fn uniform(h: u32, s: u32, v: u32, e: [f32; 3]) -> HueSatTable {
HueSatTable::new(h, s, v, false, vec![e; (h * s * v) as usize]).unwrap()
}
fn tables(hue_sat: Option<HueSatTable>, look: Option<HueSatTable>) -> ProfileTables {
ProfileTables {
name: "test".into(),
origin: ProfileOrigin::Embedded,
hue_sat,
look,
tone_curve: None,
}
}
fn close(a: [f32; 3], b: [f32; 3], tol: f32) -> bool {
a.iter()
.zip(b)
.all(|(x, y)| (x - y).abs() <= tol * y.abs().max(1.0))
}
#[test]
fn it_starts_neutral_and_composed() {
let op = CameraProfile::new();
assert!(!op.is_active(), "an untouched photograph writes nothing");
assert!(op.composes(), "and still renders through its profile");
let mut off = CameraProfile::new();
off.set_param(APPLY, 0.0);
assert!(off.is_active() && !off.composes());
}
#[test]
fn the_working_space_round_trips_through_prophoto() {
let (to, back) = working_prophoto();
for c in [[1.0, 1.0, 1.0], [0.2, 0.5, 0.1], [4.0, 0.3, 0.02]] {
assert!(close(mul(back, mul(to, c)), c, 1e-5), "{c:?}");
}
let white = mul(to, [1.0; 3]);
assert!(white.iter().all(|v| (v - 1.0).abs() < 1e-6), "{white:?}");
}
#[test]
fn hsv_round_trips() {
for c in [
[0.9, 0.2, 0.1],
[0.1, 0.7, 0.3],
[0.2, 0.3, 0.8],
[0.5, 0.5, 0.5],
[3.0, 1.0, 2.0],
] {
assert!(close(hsv_to_rgb(rgb_to_hsv(c)), c, 1e-6), "{c:?}");
}
}
#[test]
fn grey_passes_through() {
let t = tables(
Some(uniform(6, 3, 1, [30.0, 1.5, 1.0])),
Some(uniform(6, 3, 1, [-20.0, 1.3, 1.0])),
);
for v in [0.0, 0.18, 1.0, 8.0] {
let out = apply_reference(&t, [v; 3], 1.0);
assert!(close(out, [v; 3], 1e-5), "{v}: {out:?}");
}
}
#[test]
fn an_identity_table_changes_nothing() {
let t = tables(
Some(uniform(90, 30, 1, [0.0, 1.0, 1.0])),
Some(uniform(36, 8, 16, [0.0, 1.0, 1.0])),
);
for c in [[0.9, 0.2, 0.1], [0.05, 0.4, 0.2], [2.0, 0.5, 0.3]] {
assert!(close(apply_reference(&t, c, 1.0), c, 1e-5), "{c:?}");
}
}
#[test]
fn a_saturation_scale_scales_saturation() {
let t = tables(Some(uniform(6, 3, 1, [0.0, 1.2, 1.0])), None);
let (to, _) = working_prophoto();
let c = [0.6, 0.3, 0.2];
let before = rgb_to_hsv(mul(to, c));
let after = rgb_to_hsv(mul(to, apply_reference(&t, c, 1.0)));
assert!(
(after[1] - before[1] * 1.2).abs() < 1e-4,
"{before:?} {after:?}"
);
assert!((after[0] - before[0]).abs() < 1e-4);
assert!((after[2] - before[2]).abs() < 1e-4);
}
#[test]
fn hue_interpolation_wraps_from_the_last_column_to_the_first() {
// Four hue columns: a shift only in the first. A hue just short of
// 6.0 (red, from the magenta side) sits between the last column and
// the first, and must take most of the first's shift.
let mut e = vec![[0.0, 1.0, 1.0]; 4 * 2];
e[0] = [40.0, 1.0, 1.0];
e[1] = [40.0, 1.0, 1.0];
let t = HueSatTable::new(4, 2, 1, false, e).unwrap();
let d = lookup(&t, [5.9, 0.5, 0.5]);
assert!(d[0] > 30.0, "{d:?}");
// Columns sit at hue 0, 1.5, 3 and 4.5; between the third and the
// fourth, neither of which shifts, nothing moves.
let d = lookup(&t, [3.7, 0.5, 0.5]);
assert!(d[0].abs() < 1e-6, "{d:?}");
}
#[test]
fn a_value_above_one_stays_above_one() {
let t = tables(None, Some(uniform(6, 3, 4, [5.0, 1.1, 0.9])));
let out = apply_reference(&t, [6.0, 3.0, 2.0], 1.0);
assert!(out.iter().any(|v| *v > 1.0), "{out:?}");
let mut srgb = uniform(6, 3, 4, [0.0, 1.0, 0.9]);
srgb.srgb_encoded = true;
let out = apply_reference(&tables(None, Some(srgb)), [6.0, 3.0, 2.0], 1.0);
assert!(out.iter().all(|v| v.is_finite()) && out[0] > 1.0, "{out:?}");
}
#[test]
fn the_look_strength_scales_the_look_alone() {
let hs = uniform(6, 3, 1, [0.0, 1.1, 1.0]);
let look = uniform(6, 3, 1, [0.0, 1.2, 1.0]);
let t = tables(Some(hs.clone()), Some(look));
let c = [0.5, 0.3, 0.2];
let none = apply_reference(&t, c, 0.0);
assert!(close(
none,
apply_reference(&tables(Some(hs), None), c, 1.0),
1e-6
));
let (to, _) = working_prophoto();
let s = |x| rgb_to_hsv(mul(to, x))[1];
assert!(s(apply_reference(&t, c, 2.0)) > s(apply_reference(&t, c, 1.0)));
}
#[test]
fn the_buffer_puts_the_header_first_and_the_look_after_the_hue_sat_map() {
let bare = profile_buffer(None);
assert_eq!(bare[..2], [[0.0; 4]; 2], "no tables");
assert_eq!(bare[2][0], 1025.0, "and the reference default curve");
assert_eq!(bare.len(), HEADER_ENTRIES + 1025);
let t = tables(
Some(uniform(2, 2, 1, [1.0, 2.0, 3.0])),
Some(uniform(3, 2, 2, [4.0, 5.0, 6.0])),
);
let b = profile_buffer(Some(&t));
assert_eq!(b[0], [2.0, 2.0, 1.0, 0.0]);
assert_eq!(b[1], [3.0, 2.0, 2.0, 0.0]);
assert_eq!(b.len(), HEADER_ENTRIES + 4 + 12 + 1025);
assert_eq!(b[HEADER_ENTRIES], [1.0, 2.0, 3.0, 0.0]);
assert_eq!(b[HEADER_ENTRIES + 4], [4.0, 5.0, 6.0, 0.0]);
assert_eq!(b[HEADER_ENTRIES + 16][0], dr_types::tone::ACR3_DEFAULT[0]);
}
}
+32 -6
View File
@@ -39,6 +39,16 @@ use crate::ops::helpers;
pub const ID: OpId = OpId("colour_mixer");
/// How much a raised saturation band adds to a muted colour, per unit of its
/// value; the push tapers linearly to nothing at full saturation.
///
/// Measured, not chosen. Fitted against the photographer's earlier exports
/// (two looks, ~90 photographs), a sky band raised by 58 there lifted muted
/// sky blues about 2.1×; at 3.0 a band here does about the same at the same
/// value, so imported values stay inside the slider's range (dr-preset-xmp
/// carries the per-band factors). Lowering saturation is unaffected.
pub const SAT_GAIN: f32 = 3.0;
/// The twelve bands, in hue order starting at red.
///
/// Twelve rather than Lightroom's eight: the extra bands fall between the
@@ -333,9 +343,16 @@ impl Operation for ColourMixer {
// Only the bands the user actually touched contribute code. A single
// adjusted band therefore costs one weight evaluation rather than
// twelve — the composition property applied within an operation.
let mut lines = String::from(
let lines = String::from(
"\
let hcl = rgb_to_hcl(c);
// Bands are matched, and saturation judged, on display-encoded values: in
// scene-linear light a muted colour reads as strongly saturated and its hue
// sits away from where it is seen, so a band set on what the photograph
// shows would land on other colours.
let lin = max(c, vec3<f32>(0.0));
let e = pow(lin, vec3<f32>(1.0 / 2.2));
let SAT_GAIN = @SAT_GAIN@;
let hcl = rgb_to_hcl(e);
let hue = hcl.x;
let chroma = hcl.y;
let hi = hcl.z;
@@ -350,6 +367,8 @@ if (chroma > 0.0001) {
",
);
let mut lines = lines.replace("@SAT_GAIN@", &format!("{SAT_GAIN:.4}"));
for (b, band) in BANDS.iter().enumerate() {
let v = self.values[b];
if v.iter().all(|x| *x == 0.0) {
@@ -384,10 +403,17 @@ if (chroma > 0.0001) {
// yellow-green to green, not enough to turn it blue by accident.
let new_hue = hue + d_hue * 30.0;
// Saturation scales chroma; luminance scales the whole colour.
let new_chroma = clamp(chroma * (1.0 + d_sat), 0.0, hi);
c = hue_to_rgb_scale(new_hue, new_chroma, hi);
c = c * exp2(d_lum);
// Saturation. Raising it pushes muted colours hardest and tapers to
// nothing at full saturation, as the eye expects a mixer to; the
// gain makes a value deliver the strength it names, measured
// against the photographer's earlier exports. Lowering it scales
// every colour alike, so -100 is grey.
let sat = chroma / max(hi, 0.00001);
let gain = select(1.0 + d_sat, 1.0 + d_sat * SAT_GAIN * (1.0 - sat), d_sat > 0.0);
let new_chroma = clamp(chroma * gain, 0.0, hi);
let shifted = hue_to_rgb_scale(new_hue, new_chroma, hi);
// Back to linear light, where luminance scales the whole colour.
c = pow(shifted, vec3<f32>(2.2)) * exp2(d_lum);
}
}
c = max(c, vec3<f32>(0.0));",
+2
View File
@@ -66,6 +66,7 @@
// Hand-written nodes. Each is listed in `ops/` with `rust:`, which is what
// places it in the chain; these are the implementations that entry points at.
pub mod aberration;
pub mod camera_profile;
pub mod capture_sharpen;
pub mod colour_mixer;
pub mod curve;
@@ -78,6 +79,7 @@ pub mod view_transform;
pub mod vignetting;
pub use aberration::Aberration;
pub use camera_profile::CameraProfile;
pub use capture_sharpen::CaptureSharpen;
pub use colour_mixer::ColourMixer;
pub use curve::ToneCurve;
+64 -8
View File
@@ -33,11 +33,33 @@ use crate::view::{Sigmoid, CONTRAST_RANGE, DEFAULT_CONTRAST, DEFAULT_WHITE, WHIT
pub const ID: OpId = OpId("view_transform");
pub const CONTRAST: ParamId = ParamId("contrast");
pub const WHITE: ParamId = ParamId("white");
/// TRACES: FR-DEV-3j
/// Which curve renders: the DNG reference (D21, the default) or D19's sigmoid.
///
/// The reference because it is the one that matches what the photographs were
/// first developed with: on 60 Lightroom exports with neutral settings it
/// renders their raws within MSE ~150 of Lightroom's own JPEGs at the default
/// contrast, where 0.20.0's sigmoid was ~1200 (darker by about 0.7 EV and
/// flatter). The sigmoid keeps index 0 because sidecars record the index.
pub const CURVE: ParamId = ParamId("curve");
static HELPERS: [Helper; 1] = [Helper {
name: "view_sigmoid",
source: crate::view::VIEW_SIGMOID_WGSL,
}];
/// [`CURVE`]'s values, in the order of its variants.
pub const SIGMOID: f32 = 0.0;
pub const CAMERA_RAW: f32 = 1.0;
static HELPERS: LazyLock<[Helper; 3]> = LazyLock::new(|| {
[
Helper {
name: "view_sigmoid",
source: crate::view::VIEW_SIGMOID_WGSL,
},
*crate::ops::camera_profile::PROPHOTO_HELPER,
Helper {
name: "camera_raw_tone",
source: crate::camera_raw::CAMERA_RAW_WGSL,
},
]
});
static DESCRIPTOR: LazyLock<Arc<OpDescriptor>> = LazyLock::new(|| {
Arc::new(OpDescriptor {
@@ -67,6 +89,15 @@ static DESCRIPTOR: LazyLock<Arc<OpDescriptor>> = LazyLock::new(|| {
Scale::Linear,
1,
),
ParamDescriptor::choice(
"curve",
"param.view_transform.curve",
vec![
LocalizedKey("param.view_transform.curve.sigmoid"),
LocalizedKey("param.view_transform.curve.camera_raw"),
],
)
.with_default(CAMERA_RAW),
],
})
});
@@ -75,6 +106,7 @@ static DESCRIPTOR: LazyLock<Arc<OpDescriptor>> = LazyLock::new(|| {
pub struct ViewTransform {
contrast: f32,
white: f32,
curve: f32,
}
impl Default for ViewTransform {
@@ -82,6 +114,7 @@ impl Default for ViewTransform {
Self {
contrast: DEFAULT_CONTRAST,
white: DEFAULT_WHITE,
curve: CAMERA_RAW,
}
}
}
@@ -106,6 +139,7 @@ impl Operation for ViewTransform {
match id {
CONTRAST => self.contrast = value,
WHITE => self.white = value,
CURVE => self.curve = value.round(),
_ => log::warn!("view_transform: unknown parameter {id}"),
}
}
@@ -114,12 +148,13 @@ impl Operation for ViewTransform {
match id {
CONTRAST => self.contrast,
WHITE => self.white,
CURVE => self.curve,
_ => 0.0,
}
}
fn is_active(&self) -> bool {
self.contrast != DEFAULT_CONTRAST || self.white != DEFAULT_WHITE
self.contrast != DEFAULT_CONTRAST || self.white != DEFAULT_WHITE || self.curve != CAMERA_RAW
}
fn stage(&self) -> Stage {
@@ -130,8 +165,13 @@ impl Operation for ViewTransform {
"\
// Skipped for an already-rendered source: a JPEG is a display rendering
// already, and rendering it again would compress it twice.
// D19's sigmoid by default, the DNG reference tone by choice (D21).
if (!non_linear) {
c = view_sigmoid(c, slope, inv_k, peak);
if (mode > 0.5) {
c = camera_raw_tone(c, cr_scale, cr_power, cr_grey);
} else {
c = view_sigmoid(c, slope, inv_k, peak);
}
}"
.into()
}
@@ -151,11 +191,27 @@ if (!non_linear) {
name: "peak",
value: s.w,
},
Uniform {
name: "mode",
value: self.curve,
},
Uniform {
name: "cr_scale",
value: crate::camera_raw::input_scale(self.white),
},
Uniform {
name: "cr_power",
value: crate::camera_raw::contrast_power(self.contrast),
},
Uniform {
name: "cr_grey",
value: crate::view::SCENE_GREY,
},
]
}
fn helpers(&self) -> &[Helper] {
&HELPERS
HELPERS.as_slice()
}
}
@@ -191,7 +247,7 @@ mod tests {
let s = Sigmoid::new(2.0, 6.0);
let u = op.uniforms();
assert_eq!(
u.iter().map(|u| u.value).collect::<Vec<_>>(),
u.iter().take(3).map(|u| u.value).collect::<Vec<_>>(),
vec![s.n, s.inv_k, s.w]
);
}
+141
View File
@@ -675,6 +675,56 @@ impl PresetLibrary {
self.unknown.values().map(Vec::len).sum()
}
/// TRACES: FR-DEV-6
/// Merge two copies of a library that both descend from `base`.
///
/// For keeping one library on several devices: `base` is what the last
/// exchange left both sides holding, `ours` this device's copy now and
/// `theirs` the server's. Each name is decided on its own:
///
/// - Changed on one side only — added, edited or deleted — that side's
/// answer stands. This is why the base is needed at all: without it a
/// preset deleted here and one added there look the same, and a
/// deletion would come back on every exchange.
/// - Changed on both sides to the same thing, nothing to decide.
/// - Deleted on one side and edited on the other, the edit stands. A
/// preset is work, and an absence is not.
/// - Edited on both sides differently, ours stands. Either answer loses
/// one edit; this one at least converges, since the other device takes
/// ours on its next exchange as an edit made on one side only.
///
/// A missing `base` is an empty one, which can only add: a device's first
/// exchange is a union of the two libraries, never a deletion.
pub fn merge(base: &Self, ours: &Self, theirs: &Self) -> Self {
let mut out = Self::default();
let names: std::collections::BTreeSet<&str> = ours.names().chain(theirs.names()).collect();
for name in names {
let (b, o, t) = (base.entry(name), ours.entry(name), theirs.entry(name));
let chosen = if o == t || t == b {
o
} else if o == b {
t
} else {
o.or(t)
};
if let Some((preset, unknown)) = chosen {
out.presets.insert(name.to_string(), preset.clone());
if let Some(lines) = unknown {
out.unknown.insert(name.to_string(), lines.clone());
}
}
}
out
}
/// One name's preset together with the lines kept beside it, which are
/// part of what that preset is when two copies are compared.
fn entry(&self, name: &str) -> Option<(&Preset, Option<&Vec<String>>)> {
self.presets
.get(name)
.map(|preset| (preset, self.unknown.get(name)))
}
/// Serialise to the on-disk form.
///
/// Deterministic, like the sidecar's: the same library always produces the
@@ -1479,6 +1529,97 @@ mod tests {
assert_eq!(other.to_text(), named().to_text());
}
/// A one-parameter preset, so two of them differ by their value.
fn exposure(ev: f32) -> Preset {
let mut params = BTreeMap::new();
params.insert(("exposure".to_string(), "exposure".to_string()), ev);
Preset::from_params(params)
}
fn library_of(entries: &[(&str, f32)]) -> PresetLibrary {
let mut lib = PresetLibrary::default();
for (name, ev) in entries {
lib.insert(name, exposure(*ev)).unwrap();
}
lib
}
#[test]
fn a_first_merge_is_the_union_of_both_libraries() {
let ours = library_of(&[("Mine", 1.0), ("Both", 0.5)]);
let theirs = library_of(&[("Theirs", 2.0), ("Both", 0.5)]);
let merged = PresetLibrary::merge(&PresetLibrary::default(), &ours, &theirs);
assert_eq!(
merged.names().collect::<Vec<_>>(),
vec!["Both", "Mine", "Theirs"]
);
}
#[test]
fn a_deletion_on_either_side_is_kept_rather_than_undone() {
// The case the base exists for: without it, the deleted preset is
// indistinguishable from one the other side has just added.
let base = library_of(&[("Gone here", 1.0), ("Gone there", 2.0)]);
let ours = library_of(&[("Gone there", 2.0)]);
let theirs = library_of(&[("Gone here", 1.0)]);
assert!(PresetLibrary::merge(&base, &ours, &theirs).is_empty());
}
#[test]
fn an_edit_on_one_side_reaches_the_other() {
let base = library_of(&[("Warm", 1.0)]);
let ours = library_of(&[("Warm", 1.0)]);
let theirs = library_of(&[("Warm", 1.5)]);
let merged = PresetLibrary::merge(&base, &ours, &theirs);
assert_eq!(merged.get("Warm"), Some(&exposure(1.5)));
// And the other way round.
let merged = PresetLibrary::merge(&base, &theirs, &ours);
assert_eq!(merged.get("Warm"), Some(&exposure(1.5)));
}
#[test]
fn an_edit_outlives_a_deletion_made_elsewhere() {
let base = library_of(&[("Warm", 1.0)]);
let edited = library_of(&[("Warm", 1.5)]);
let deleted = PresetLibrary::default();
assert_eq!(
PresetLibrary::merge(&base, &edited, &deleted).get("Warm"),
Some(&exposure(1.5))
);
assert_eq!(
PresetLibrary::merge(&base, &deleted, &edited).get("Warm"),
Some(&exposure(1.5))
);
}
#[test]
fn two_different_edits_keep_ours_and_then_converge() {
let base = library_of(&[("Warm", 1.0)]);
let here = library_of(&[("Warm", 1.5)]);
let there = library_of(&[("Warm", 0.5)]);
let pushed = PresetLibrary::merge(&base, &here, &there);
assert_eq!(pushed.get("Warm"), Some(&exposure(1.5)));
// The other device's next exchange: its base is what it last pushed,
// its own copy is unchanged since, and the server holds ours.
let settled = PresetLibrary::merge(&there, &there, &pushed);
assert_eq!(settled, pushed);
}
#[test]
fn lines_this_build_cannot_read_travel_with_their_preset() {
let text = format!(
"drpl {LIBRARY_FORMAT_VERSION}\n\n[preset Future]\nexposure.exposure = 0.5\n\
something_new_entirely\n"
);
let theirs = PresetLibrary::parse(&text).unwrap();
let merged = PresetLibrary::merge(
&PresetLibrary::default(),
&PresetLibrary::default(),
&theirs,
);
assert!(merged.to_text().contains("something_new_entirely"));
}
#[test]
fn a_neutral_preset_is_storable_and_survives_the_round_trip() {
// The empty preset is the "clear these forty frames" action, so it has
+19 -6
View File
@@ -51,8 +51,19 @@ pub const SCENE_GREY: f32 = 0.13;
/// Display-linear middle grey — what a camera JPEG shows a grey card as.
pub const DISPLAY_GREY: f32 = 0.18;
/// The default contrast, the sigmoid's log-log slope parameter `n`.
pub const DEFAULT_CONTRAST: f32 = 1.4;
/// The default contrast: the sigmoid's log-log slope `n`, and for the DNG
/// reference curve a power of `DEFAULT_CONTRAST / REFERENCE_CONTRAST` about grey.
///
/// Fitted, not chosen: on Lightroom 6 exports whose look settings were neutral,
/// the DNG reference curve matched the exports best with the input bent by about
/// 1.08 (held-out MSE 224 at 1.4, ~150 at 1.5). The sigmoid, which is no longer
/// the default, fitted best near 1.7 and is better at 1.5 than at 1.4.
pub const DEFAULT_CONTRAST: f32 = 1.5;
/// The contrast at which each curve is its own reference: the sigmoid's match
/// to the retired base curve (see the module note), and the reference table
/// untouched.
pub const REFERENCE_CONTRAST: f32 = 1.4;
/// The default white point, in stops above [`SCENE_GREY`].
pub const DEFAULT_WHITE: f32 = 4.0;
@@ -223,11 +234,13 @@ mod tests {
}
#[test]
fn the_default_stays_close_to_the_retired_curve() {
fn the_reference_stays_close_to_the_retired_curve() {
// TRACES: FR-DEV-3j | FR-DEV-3e
// D19's promise to every existing photograph: the midtones do not
// move by more than a third of a stop.
let s = Sigmoid::default_curve();
// D19's promise, now kept by the sigmoid at its reference contrast
// rather than by the default (D21 moved the default to the DNG reference
// curve): the midtones do not move by more than a third of a stop
// from the retired base curve.
let s = Sigmoid::new(REFERENCE_CONTRAST, DEFAULT_WHITE);
let mut x = 0.03_f32;
while x <= 1.0 {
let ev = (s.channel(x) / retired_default(x)).log2();
+2 -1
View File
@@ -33,7 +33,7 @@
//!
//! A `rust:` node — `tone_curve`, `colour_mixer`, `film_sim`,
//! `capture_sharpen`, `noise_reduction`, `clarity`, `texture`, `dehaze`,
//! `view_transform` —
//! `view_transform`, `camera_profile` —
//! names a hand-written type and has no declaration to interpret. It is not skipped
//! silently: [`every_declared_node_is_checked`] asserts the two sets partition
//! `ops/` between them, so a node that stops being declared cannot quietly
@@ -396,6 +396,7 @@ fn every_declared_node_is_checked() {
assert_eq!(
hand,
[
"camera_profile",
"capture_sharpen",
"clarity",
"colour_mixer",
+390 -28
View File
@@ -35,7 +35,7 @@
//! wrong on most images and invisibly so, which is worse than an honest gap,
//! so these keys are counted as skipped and reported.
//!
//! **Tone curves, colour mixing, masks, lens profiles and grain.** Each is a
//! **Tone curves, masks, lens profiles and grain.** Each is a
//! structure rather than a number, and each would need its own argument about
//! whether the two applications mean the same thing. They are skipped by
//! omission — a key not in the table is simply not understood — and the
@@ -83,35 +83,43 @@ const MAPPINGS: &[Mapping] = &[
param: "exposure",
convert: Convert::Direct,
},
// The five tone sliders do not mean the same thing in the two applications,
// whatever their shared ±100 suggests. The factors were fitted against the
// library's own Lightroom 6 exports and their raws (darkroom-lrfit, 2026-10):
// each photograph's sliders carried across as `slider × factor`, one factor
// per slider, on ~90 exports with no look applied. Contrast is ours at a
// tenth — at −100 ours flattens a frame to grey — and our shadows need
// nearly twice Lightroom's number. Whites barely appears in those exports;
// every fit put it under 1 but none agreed where, so 0.5 is a hedge.
Mapping {
crs: "Contrast2012",
op: "contrast",
param: "contrast",
convert: Convert::Direct,
convert: Convert::Scale(0.1),
},
Mapping {
crs: "Highlights2012",
op: "highlights_shadows",
param: "highlights",
convert: Convert::Direct,
convert: Convert::Scale(1.4),
},
Mapping {
crs: "Shadows2012",
op: "highlights_shadows",
param: "shadows",
convert: Convert::Direct,
convert: Convert::Scale(1.9),
},
Mapping {
crs: "Whites2012",
op: "blacks_whites",
param: "whites",
convert: Convert::Direct,
convert: Convert::Scale(0.5),
},
Mapping {
crs: "Blacks2012",
op: "blacks_whites",
param: "blacks",
convert: Convert::Direct,
convert: Convert::Scale(1.25),
},
Mapping {
crs: "Clarity2012",
@@ -137,6 +145,263 @@ const MAPPINGS: &[Mapping] = &[
param: "saturation",
convert: Convert::Direct,
},
// TRACES: FR-DEV-6
// Lightroom's HSL panel: eight bands, each ±100 for hue, saturation and
// luminance, onto the colour mixer's twelve.
//
// Hue and luminance go to the band of the same hue, one for one: Aqua
// (180°) is our cyan, Purple (270°) our violet.
//
// Saturation is measured. Fitted against the library's Lightroom 6 exports
// of two looks and their raws (darkroom-lrfit, hsl_map_fit), Lightroom's
// saturation bands are about 45° wide either side on our hue wheel, wider
// than ours, and do not all have our strength: each is shared between
// two or three of our bands with the factors below. Aqua sits at 187°
// and reaches into azure, where skies are; Blue at 251° reaches violet;
// Orange, where skin is, carries only ~0.4 — Lightroom's Orange is
// gentle. Red, Purple and Magenta barely appear in those exports and
// take the common gain; Green is capped where its few pixels would push
// it further. Values add when two Lightroom bands share one of ours.
Mapping {
crs: "HueAdjustmentRed",
op: "colour_mixer",
param: "red_hue",
convert: Convert::Direct,
},
Mapping {
crs: "SaturationAdjustmentRed",
op: "colour_mixer",
param: "red_sat",
convert: Convert::Scale(1.02),
},
Mapping {
crs: "SaturationAdjustmentRed",
op: "colour_mixer",
param: "orange_sat",
convert: Convert::Scale(0.34),
},
Mapping {
crs: "SaturationAdjustmentRed",
op: "colour_mixer",
param: "rose_sat",
convert: Convert::Scale(0.34),
},
Mapping {
crs: "LuminanceAdjustmentRed",
op: "colour_mixer",
param: "red_lum",
convert: Convert::Direct,
},
Mapping {
crs: "HueAdjustmentOrange",
op: "colour_mixer",
param: "orange_hue",
convert: Convert::Direct,
},
Mapping {
crs: "SaturationAdjustmentOrange",
op: "colour_mixer",
param: "orange_sat",
convert: Convert::Scale(0.39),
},
Mapping {
crs: "SaturationAdjustmentOrange",
op: "colour_mixer",
param: "red_sat",
convert: Convert::Scale(0.13),
},
Mapping {
crs: "SaturationAdjustmentOrange",
op: "colour_mixer",
param: "yellow_sat",
convert: Convert::Scale(0.13),
},
Mapping {
crs: "LuminanceAdjustmentOrange",
op: "colour_mixer",
param: "orange_lum",
convert: Convert::Direct,
},
Mapping {
crs: "HueAdjustmentYellow",
op: "colour_mixer",
param: "yellow_hue",
convert: Convert::Direct,
},
Mapping {
crs: "SaturationAdjustmentYellow",
op: "colour_mixer",
param: "yellow_sat",
convert: Convert::Scale(0.87),
},
Mapping {
crs: "SaturationAdjustmentYellow",
op: "colour_mixer",
param: "orange_sat",
convert: Convert::Scale(0.29),
},
Mapping {
crs: "SaturationAdjustmentYellow",
op: "colour_mixer",
param: "chartreuse_sat",
convert: Convert::Scale(0.29),
},
Mapping {
crs: "LuminanceAdjustmentYellow",
op: "colour_mixer",
param: "yellow_lum",
convert: Convert::Direct,
},
Mapping {
crs: "HueAdjustmentGreen",
op: "colour_mixer",
param: "green_hue",
convert: Convert::Direct,
},
Mapping {
crs: "SaturationAdjustmentGreen",
op: "colour_mixer",
param: "green_sat",
convert: Convert::Scale(1.2),
},
Mapping {
crs: "SaturationAdjustmentGreen",
op: "colour_mixer",
param: "chartreuse_sat",
convert: Convert::Scale(0.4),
},
Mapping {
crs: "SaturationAdjustmentGreen",
op: "colour_mixer",
param: "spring_sat",
convert: Convert::Scale(0.4),
},
Mapping {
crs: "LuminanceAdjustmentGreen",
op: "colour_mixer",
param: "green_lum",
convert: Convert::Direct,
},
Mapping {
crs: "HueAdjustmentAqua",
op: "colour_mixer",
param: "cyan_hue",
convert: Convert::Direct,
},
Mapping {
crs: "SaturationAdjustmentAqua",
op: "colour_mixer",
param: "cyan_sat",
convert: Convert::Scale(0.9),
},
Mapping {
crs: "SaturationAdjustmentAqua",
op: "colour_mixer",
param: "azure_sat",
convert: Convert::Scale(0.51),
},
Mapping {
crs: "SaturationAdjustmentAqua",
op: "colour_mixer",
param: "spring_sat",
convert: Convert::Scale(0.19),
},
Mapping {
crs: "LuminanceAdjustmentAqua",
op: "colour_mixer",
param: "cyan_lum",
convert: Convert::Direct,
},
Mapping {
crs: "HueAdjustmentBlue",
op: "colour_mixer",
param: "blue_hue",
convert: Convert::Direct,
},
Mapping {
crs: "SaturationAdjustmentBlue",
op: "colour_mixer",
param: "blue_sat",
convert: Convert::Scale(0.75),
},
Mapping {
crs: "SaturationAdjustmentBlue",
op: "colour_mixer",
param: "violet_sat",
convert: Convert::Scale(0.56),
},
Mapping {
crs: "SaturationAdjustmentBlue",
op: "colour_mixer",
param: "azure_sat",
convert: Convert::Scale(0.09),
},
Mapping {
crs: "LuminanceAdjustmentBlue",
op: "colour_mixer",
param: "blue_lum",
convert: Convert::Direct,
},
Mapping {
crs: "HueAdjustmentPurple",
op: "colour_mixer",
param: "violet_hue",
convert: Convert::Direct,
},
Mapping {
crs: "SaturationAdjustmentPurple",
op: "colour_mixer",
param: "violet_sat",
convert: Convert::Scale(1.26),
},
Mapping {
crs: "SaturationAdjustmentPurple",
op: "colour_mixer",
param: "blue_sat",
convert: Convert::Scale(0.42),
},
Mapping {
crs: "SaturationAdjustmentPurple",
op: "colour_mixer",
param: "magenta_sat",
convert: Convert::Scale(0.42),
},
Mapping {
crs: "LuminanceAdjustmentPurple",
op: "colour_mixer",
param: "violet_lum",
convert: Convert::Direct,
},
Mapping {
crs: "HueAdjustmentMagenta",
op: "colour_mixer",
param: "magenta_hue",
convert: Convert::Direct,
},
Mapping {
crs: "SaturationAdjustmentMagenta",
op: "colour_mixer",
param: "magenta_sat",
convert: Convert::Scale(1.05),
},
Mapping {
crs: "SaturationAdjustmentMagenta",
op: "colour_mixer",
param: "violet_sat",
convert: Convert::Scale(0.35),
},
Mapping {
crs: "SaturationAdjustmentMagenta",
op: "colour_mixer",
param: "rose_sat",
convert: Convert::Scale(0.35),
},
Mapping {
crs: "LuminanceAdjustmentMagenta",
op: "colour_mixer",
param: "magenta_lum",
convert: Convert::Direct,
},
// Adobe's sharpening runs 0…150 where ours runs 0…100, so a preset asking
// for its maximum gets ours rather than being clamped there silently.
Mapping {
@@ -192,6 +457,27 @@ pub enum ImportError {
NoSettings,
}
/// TRACES: FR-DEV-6
/// The Lightroom edit stored inside a photograph — the XMP packet Lightroom
/// writes into a DNG — translated as a preset is.
///
/// `None` where the file carries no packet, or one with no Camera Raw
/// settings in it (darktable's sidecars, a camera's own XMP). The packet is
/// found by its delimiters rather than by walking the TIFF structure: it is
/// plain text by specification, and the same search serves any container.
pub fn read_embedded(bytes: &[u8]) -> Option<Import> {
const OPEN: &[u8] = b"<x:xmpmeta";
const CLOSE: &[u8] = b"</x:xmpmeta>";
let start = find(bytes, OPEN)?;
let end = start + find(&bytes[start..], CLOSE)? + CLOSE.len();
let text = std::str::from_utf8(&bytes[start..end]).ok()?;
read_xmp(text).ok().filter(|i| !i.preset.is_empty())
}
fn find(haystack: &[u8], needle: &[u8]) -> Option<usize> {
haystack.windows(needle.len()).position(|w| w == needle)
}
/// Read one Lightroom `.xmp` preset.
///
/// Tolerant in the same direction the sidecar parser is: a value that will not
@@ -285,7 +571,12 @@ pub fn read_xmp(text: &str) -> Result<Import, ImportError> {
// Out-of-range values are left as they are: `EditGraph::set_param`
// clamps when the preset is applied, and clamping here as well would
// mean two places to be wrong about a range.
params.insert((mapping.op.to_string(), mapping.param.to_string()), value);
//
// Added rather than set: one of Lightroom's HSL bands is shared
// between two or three of ours, and two of its bands can share one.
*params
.entry((mapping.op.to_string(), mapping.param.to_string()))
.or_insert(0.0) += value;
}
let skipped = KNOWN_UNSUPPORTED
@@ -375,6 +666,65 @@ mod tests {
.copied()
}
#[test]
fn a_dngs_embedded_lightroom_edit_comes_across_with_its_hsl() {
// TRACES: FR-DEV-6
// The shape Lightroom 6 writes into a DNG, trimmed: the library's
// house look, as camera-profiles.md's D21 note records it.
let mut file = b"II*\0 binary header bytes ".to_vec();
file.extend_from_slice(
br#"<?xpacket begin="" id="W5M0MpCehiHzreSzNTczkc9d"?><x:xmpmeta xmlns:x="adobe:ns:meta/"><rdf:RDF xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#"><rdf:Description rdf:about="" xmlns:crs="http://ns.adobe.com/camera-raw-settings/1.0/" crs:ProcessVersion="6.7" crs:Exposure2012="0.00" crs:Highlights2012="-40" crs:Blacks2012="-20" crs:SaturationAdjustmentBlue="+58" crs:SaturationAdjustmentAqua="+50" crs:SaturationAdjustmentPurple="+23" crs:HueAdjustmentRed="-5" crs:LuminanceAdjustmentGreen="+7"/></rdf:RDF></x:xmpmeta><?xpacket end="w"?>"#,
);
file.extend_from_slice(b"\0 more binary");
let import = read_embedded(&file).expect("an edit");
let p = import.preset.params();
let get = |op: &str, param: &str| p.get(&(op.to_string(), param.to_string())).copied();
// Saturation is shared out by the measured factors; values add.
let near = |got: Option<f32>, want: f32| {
let got = got.expect("set");
assert!((got - want).abs() < 1e-3, "{got} vs {want}");
};
near(get("colour_mixer", "blue_sat"), 58.0 * 0.75 + 23.0 * 0.42);
near(get("colour_mixer", "cyan_sat"), 50.0 * 0.9);
near(get("colour_mixer", "azure_sat"), 50.0 * 0.51 + 58.0 * 0.09);
near(get("colour_mixer", "violet_sat"), 58.0 * 0.56 + 23.0 * 1.26);
assert_eq!(get("colour_mixer", "red_hue"), Some(-5.0));
assert_eq!(get("colour_mixer", "green_lum"), Some(7.0));
assert_eq!(get("highlights_shadows", "highlights"), Some(-40.0 * 1.4));
assert_eq!(get("blacks_whites", "blacks"), Some(-20.0 * 1.25));
}
#[test]
fn the_libraries_dngs_carry_the_house_look() {
// TRACES: FR-DEV-6
// The library's own Lightroom 6 DNG, where it is on this machine.
let Some(home) = std::env::var_os("HOME") else {
return;
};
let path =
std::path::Path::new(&home).join("Nextcloud/PhotosRaw/2017/2017-08-12/_MG_9080.dng");
let Ok(bytes) = std::fs::read(&path) else {
eprintln!("skipped: no sample DNG at {}", path.display());
return;
};
let import = read_embedded(&bytes).expect("Lightroom's edit");
let p = import.preset.params();
let get = |op: &str, param: &str| p.get(&(op.to_string(), param.to_string())).copied();
// The house look's sky: Aqua and Blue land in cyan, azure and blue.
for (param, at_least) in [("cyan_sat", 40.0), ("azure_sat", 25.0), ("blue_sat", 40.0)] {
let v = get("colour_mixer", param).unwrap_or(0.0);
assert!(v >= at_least, "{param} {v}");
}
assert_eq!(get("highlights_shadows", "highlights"), Some(-40.0 * 1.4));
}
#[test]
fn a_file_with_no_camera_raw_settings_has_no_edit() {
assert!(read_embedded(b"no packet at all").is_none());
let darktable = br#"<x:xmpmeta xmlns:x="adobe:ns:meta/"><rdf:RDF xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#"><rdf:Description rdf:about="" xmlns:xmp="http://ns.adobe.com/xap/1.0/" xmp:Rating="3"/></rdf:RDF></x:xmpmeta>"#;
assert!(read_embedded(darktable).is_none());
}
#[test]
fn every_mapping_names_a_parameter_this_build_actually_has() {
// The test that keeps the table honest. Adobe's half cannot be checked
@@ -397,36 +747,48 @@ mod tests {
}
#[test]
fn no_two_mappings_claim_the_same_key_or_the_same_target() {
let mut keys: Vec<&str> = MAPPINGS.iter().map(|m| m.crs).collect();
keys.sort_unstable();
let before = keys.len();
keys.dedup();
assert_eq!(before, keys.len(), "two mappings read the same crs key");
fn no_two_mappings_repeat_a_key_and_target() {
// A saturation band may be shared between several of ours, and two of
// Lightroom's may share one of ours (their values add); but the same
// key written twice to the same target would count it twice.
let mut pairs: Vec<(&str, &str, &str)> =
MAPPINGS.iter().map(|m| (m.crs, m.op, m.param)).collect();
pairs.sort_unstable();
let before = pairs.len();
pairs.dedup();
assert_eq!(before, pairs.len(), "a key is written twice to one target");
let mut targets: Vec<(&str, &str)> = MAPPINGS.iter().map(|m| (m.op, m.param)).collect();
targets.sort_unstable();
let before = targets.len();
targets.dedup();
assert_eq!(
before,
targets.len(),
"two mappings write the same parameter"
);
// Only the HSL saturation bands are shared; every other key has one
// home, so a slip in the table cannot fan a slider out unnoticed.
let mut single: Vec<&str> = MAPPINGS
.iter()
.map(|m| m.crs)
.filter(|k| !k.starts_with("SaturationAdjustment"))
.collect();
single.sort_unstable();
let before = single.len();
single.dedup();
assert_eq!(before, single.len(), "two mappings read the same crs key");
}
#[test]
fn the_settings_that_share_a_convention_come_across_unchanged() {
let import = read_xmp(ATTRIBUTE_FORM).unwrap();
assert_eq!(value(&import, "exposure", "exposure"), Some(0.75));
assert_eq!(value(&import, "contrast", "contrast"), Some(25.0));
assert_eq!(value(&import, "contrast", "contrast"), Some(25.0 * 0.1));
assert_eq!(
value(&import, "highlights_shadows", "highlights"),
Some(-40.0)
Some(-40.0 * 1.4)
);
assert_eq!(
value(&import, "highlights_shadows", "shadows"),
Some(30.0 * 1.9)
);
assert_eq!(value(&import, "blacks_whites", "whites"), Some(10.0 * 0.5));
assert_eq!(
value(&import, "blacks_whites", "blacks"),
Some(-15.0 * 1.25)
);
assert_eq!(value(&import, "highlights_shadows", "shadows"), Some(30.0));
assert_eq!(value(&import, "blacks_whites", "whites"), Some(10.0));
assert_eq!(value(&import, "blacks_whites", "blacks"), Some(-15.0));
assert_eq!(value(&import, "clarity", "amount"), Some(12.0));
assert_eq!(value(&import, "texture", "amount"), Some(8.0));
assert_eq!(value(&import, "vibrance", "vibrance"), Some(20.0));
@@ -467,7 +829,7 @@ mod tests {
// does not say which shape it used.
let import = read_xmp(ELEMENT_FORM).unwrap();
assert_eq!(value(&import, "exposure", "exposure"), Some(0.75));
assert_eq!(value(&import, "contrast", "contrast"), Some(25.0));
assert_eq!(value(&import, "contrast", "contrast"), Some(25.0 * 0.1));
}
#[test]
+14 -4
View File
@@ -13,9 +13,11 @@
use std::path::Path;
const MODEL: &str = "../../models/segment/yolo26n-seg.onnx";
const QUANTISED: &str = "../../models/segment/yolo26n-seg.a16w16.onnx";
fn main() {
println!("cargo:rerun-if-changed={MODEL}");
println!("cargo:rerun-if-changed={QUANTISED}");
println!("cargo:rerun-if-changed=build.rs");
// Only the embedded path needs the file present; a build without it is
@@ -24,10 +26,18 @@ fn main() {
return;
}
let path = Path::new(MODEL);
check(MODEL);
// The Hexagon's quantised form rides only in an Android build.
if std::env::var("CARGO_CFG_TARGET_OS").as_deref() == Ok("android") {
check(QUANTISED);
}
}
fn check(model: &str) {
let path = Path::new(model);
let Ok(bytes) = std::fs::read(path) else {
panic!(
"\n\n{MODEL} is missing.\n\
"\n\n{model} is missing.\n\
It ships in Git LFS. Run `git lfs install && git lfs pull`, or build \
with `--no-default-features` for a watershed-only build.\n"
);
@@ -40,7 +50,7 @@ fn main() {
// happens in practice.
if bytes.starts_with(b"version https://git-lfs") {
panic!(
"\n\n{MODEL} is a Git LFS pointer, not the model ({} bytes).\n\
"\n\n{model} is a Git LFS pointer, not the model ({} bytes).\n\
Run `git lfs install && git lfs pull` to fetch the real file.\n",
bytes.len()
);
@@ -50,7 +60,7 @@ fn main() {
// export is ~11 MB; anything under a megabyte is a truncated checkout.
if bytes.len() < 1_000_000 {
panic!(
"\n\n{MODEL} is only {} bytes — expected ~11 MB.\n\
"\n\n{model} is only {} bytes — expected several MB.\n\
The checkout looks incomplete; try `git lfs pull`.\n",
bytes.len()
);
+1 -1
View File
@@ -72,7 +72,7 @@ pub use refine::{
#[cfg(feature = "semantic")]
pub use scene::{Category, Scene, SceneModel};
#[cfg(feature = "embedded-model")]
pub use semantic::embedded_model_bytes;
pub use semantic::embedded_models;
#[cfg(feature = "semantic")]
pub use semantic::{Instance, SemanticModel, SemanticOptions, Tiling};
+15 -7
View File
@@ -123,21 +123,29 @@ impl SceneModel {
classes: impl AsRef<std::path::Path>,
categories: impl AsRef<std::path::Path>,
) -> Result<Self, SegmentError> {
// The form the device's backend runs: the `.a16w16.onnx` sibling on
// the Hexagon (attention left in float, inference.md §1.5), else this.
let (model, form) =
dr_inference_engine::resolve_model(dr_inference_engine::Role::Scene, model.as_ref());
let bytes = std::fs::read(model).map_err(SegmentError::ModelRead)?;
let classes = std::fs::read_to_string(classes).map_err(SegmentError::ModelRead)?;
let categories = std::fs::read_to_string(categories).map_err(SegmentError::ModelRead)?;
let classes = crate::semantic::parse_classes(&classes);
let categories = parse_categories(&categories, &classes)?;
Self::from_bytes(&bytes, categories)
Self::from_bytes_in(&bytes, form, categories)
}
pub fn from_bytes(bytes: &[u8], categories: Vec<Category>) -> Result<Self, SegmentError> {
// f32, as for `SemanticModel`; see there.
let session = dr_inference_engine::open(
dr_inference_engine::Role::Scene,
dr_inference_engine::Form::F32,
bytes,
)?;
Self::from_bytes_in(bytes, dr_inference_engine::Form::F32, categories)
}
/// `bytes` in a stated numeric form; the outputs keep their shape.
pub fn from_bytes_in(
bytes: &[u8],
form: dr_inference_engine::Form,
categories: Vec<Category>,
) -> Result<Self, SegmentError> {
let session = dr_inference_engine::open(dr_inference_engine::Role::Scene, form, bytes)?;
Ok(Self {
session,
+31 -13
View File
@@ -208,18 +208,32 @@ const EMBEDDED_MODEL: &[u8] = include_bytes!("../../../models/segment/yolo26n-se
#[cfg(feature = "embedded-model")]
const EMBEDDED_CLASSES: &str = include_str!("../../../models/segment/yolo26n-seg.classes.json");
/// The bytes of the model that ships with this crate, for whoever compiles
/// The Hexagon's form (docs/dev/inference.md §1.5): 16-bit activations and
/// weights, the rows' tail left in float. Only Android has a Hexagon, so only
/// Android carries it.
#[cfg(all(feature = "embedded-model", target_os = "android"))]
const EMBEDDED_A16W16: &[u8] = include_bytes!("../../../models/segment/yolo26n-seg.a16w16.onnx");
/// Every form of the model that ships with this crate, for whoever compiles
/// engines ahead of the first request (docs/dev/inference.md §6).
#[cfg(feature = "embedded-model")]
pub fn embedded_model_bytes() -> &'static [u8] {
EMBEDDED_MODEL
pub fn embedded_models() -> Vec<(dr_inference_engine::Form, &'static [u8])> {
#[allow(unused_mut)]
let mut forms = vec![(dr_inference_engine::Form::F32, EMBEDDED_MODEL)];
#[cfg(target_os = "android")]
forms.push((dr_inference_engine::Form::A16W16, EMBEDDED_A16W16));
forms
}
impl SemanticModel {
/// Load the model that ships with this crate.
/// Load the model that ships with this crate, in the form the device's
/// backend runs.
#[cfg(feature = "embedded-model")]
pub fn embedded() -> Result<Self, SegmentError> {
Self::from_bytes(EMBEDDED_MODEL, parse_classes(EMBEDDED_CLASSES))
let forms = embedded_models();
let (bytes, form) =
dr_inference_engine::choose_embedded(dr_inference_engine::Role::Segmenter, &forms);
Self::from_bytes_in(bytes, form, parse_classes(EMBEDDED_CLASSES))
}
/// Load a model from an ONNX file, with `classes` supplying its vocabulary.
@@ -236,14 +250,18 @@ impl SemanticModel {
}
pub fn from_bytes(bytes: &[u8], classes: Vec<Arc<str>>) -> Result<Self, SegmentError> {
// The f32 graph on whatever the device's backend is. An int8 form
// for the Hexagon waits on docs/dev/inference.md §10 M7 — the mask
// boundary has to be measured before it moves.
let session = dr_inference_engine::open(
dr_inference_engine::Role::Segmenter,
dr_inference_engine::Form::F32,
bytes,
)?;
Self::from_bytes_in(bytes, dr_inference_engine::Form::F32, classes)
}
/// `bytes` in a stated numeric form. The quantised one keeps the same
/// outputs (the rows' tail stays float), so decoding does not change; the
/// masks it draws were measured against f32's (inference.md §1.5).
pub fn from_bytes_in(
bytes: &[u8],
form: dr_inference_engine::Form,
classes: Vec<Arc<str>>,
) -> Result<Self, SegmentError> {
let session = dr_inference_engine::open(dr_inference_engine::Role::Segmenter, form, bytes)?;
Ok(Self { session, classes })
}
+179
View File
@@ -0,0 +1,179 @@
//! TRACES: FR-DEV-3e
//! A camera profile's hue/saturation/value tables, as the renderer receives
//! them.
//!
//! # Why this lives in the types crate
//!
//! Three crates handle these and none depends on the next: `dr-decode` reads
//! them out of a DNG or a `.dcp`, `dr-pipeline` emits the shader that indexes
//! them, and `dr-gpu` uploads them in between. The layout — saturation
//! fastest, then hue, then value — is the one fact all three must agree on, so
//! it is stated once, here, by [`HueSatTable::index`].
//!
//! See `docs/dev/camera-profiles.md` for the model and D20 for where the
//! tables run.
/// One `ProfileHueSatMap` or `ProfileLookTable`: a grid over HSV whose every
/// entry is `(hue shift in degrees, saturation scale, value scale)`.
#[derive(Debug, Clone, PartialEq)]
pub struct HueSatTable {
pub hue_divisions: u32,
pub sat_divisions: u32,
/// 1 for a "2.5-D" table, which ignores value.
pub val_divisions: u32,
/// The value axis is indexed by the sRGB-encoded value rather than the
/// linear one (`ProfileHueSatMapEncoding` / `ProfileLookTableEncoding`
/// = 1).
pub srgb_encoded: bool,
/// `hue_divisions × sat_divisions × val_divisions` entries, in
/// [`Self::index`] order.
pub entries: Vec<[f32; 3]>,
}
impl HueSatTable {
/// Build a table, refusing one whose shape cannot be indexed.
///
/// Saturation needs two samples to interpolate between, and a table with
/// a zero dimension or the wrong number of entries is a file that lies
/// about itself; either is `None` rather than a lookup that reads past
/// its end.
pub fn new(
hue_divisions: u32,
sat_divisions: u32,
val_divisions: u32,
srgb_encoded: bool,
entries: Vec<[f32; 3]>,
) -> Option<Self> {
let count = (hue_divisions as usize)
.checked_mul(sat_divisions as usize)?
.checked_mul(val_divisions as usize)?;
let sane = hue_divisions >= 1
&& sat_divisions >= 2
&& val_divisions >= 1
&& entries.len() == count
// Large enough for any real profile (Adobe's largest are
// 90×30×1 and 36×8×16); small enough that a corrupt dimension
// cannot ask the GPU for gigabytes.
&& count <= 1 << 20
&& entries.iter().flatten().all(|v| v.is_finite());
sane.then_some(Self {
hue_divisions,
sat_divisions,
val_divisions,
srgb_encoded,
entries,
})
}
/// Where the entry for `(hue, sat, val)` sits: saturation fastest, then
/// hue, then value, as the DNG specification stores it.
pub fn index(&self, hue: u32, sat: u32, val: u32) -> usize {
((val * self.hue_divisions + hue) * self.sat_divisions + sat) as usize
}
/// Entry-by-entry blend toward `other`, for a two-illuminant HueSatMap.
///
/// `None` where the two are not the same shape, which a well-formed
/// profile never produces — both data tags share one dimensions tag.
pub fn lerp(&self, other: &Self, t: f32) -> Option<Self> {
if (self.hue_divisions, self.sat_divisions, self.val_divisions)
!= (
other.hue_divisions,
other.sat_divisions,
other.val_divisions,
)
{
return None;
}
let entries = self
.entries
.iter()
.zip(&other.entries)
.map(|(a, b)| std::array::from_fn(|i| a[i] + (b[i] - a[i]) * t))
.collect();
Some(Self {
entries,
..self.clone()
})
}
/// Whether every entry is `(0°, 1, 1)`, so the table changes nothing.
pub fn is_identity(&self) -> bool {
self.entries
.iter()
.all(|e| e[0] == 0.0 && e[1] == 1.0 && e[2] == 1.0)
}
}
/// What a source hands the renderer: the tables already resolved for this
/// frame, the HueSatMap blended for the light it was shot under.
#[derive(Debug, Clone, PartialEq)]
pub struct ProfileTables {
/// The profile's name, for the panel (`ProfileName`).
pub name: String,
/// Where it came from, for the panel.
pub origin: ProfileOrigin,
pub hue_sat: Option<HueSatTable>,
pub look: Option<HueSatTable>,
/// The profile's `ProfileToneCurve`, resampled onto
/// [`crate::tone::TONE_SAMPLES`] points; `None` where it has none, and
/// the view transform's DNG reference curve then uses
/// [`crate::tone::ACR3_DEFAULT`] (D21).
pub tone_curve: Option<Vec<f32>>,
}
/// Where a profile was found (camera-profiles.md §4).
#[derive(Debug, Clone, PartialEq, Eq)]
pub enum ProfileOrigin {
/// Embedded in the DNG being rendered.
Embedded,
/// A `.dcp` in the profiles directory, by file name.
File(String),
}
impl ProfileTables {
/// Whether there is anything to apply.
pub fn is_empty(&self) -> bool {
self.hue_sat.is_none() && self.look.is_none() && self.tone_curve.is_none()
}
}
#[cfg(test)]
mod tests {
use super::*;
fn identity(h: u32, s: u32, v: u32) -> HueSatTable {
HueSatTable::new(h, s, v, false, vec![[0.0, 1.0, 1.0]; (h * s * v) as usize]).unwrap()
}
#[test]
fn saturation_varies_fastest_then_hue_then_value() {
let t = identity(4, 3, 2);
assert_eq!(t.index(0, 1, 0), 1);
assert_eq!(t.index(1, 0, 0), 3);
assert_eq!(t.index(0, 0, 1), 12);
assert_eq!(t.index(3, 2, 1), 23);
}
#[test]
fn a_table_that_lies_about_its_size_is_refused() {
assert!(HueSatTable::new(4, 3, 1, false, vec![[0.0, 1.0, 1.0]; 11]).is_none());
assert!(HueSatTable::new(4, 1, 1, false, vec![[0.0, 1.0, 1.0]; 4]).is_none());
assert!(HueSatTable::new(0, 3, 1, false, vec![]).is_none());
let mut bad = vec![[0.0, 1.0, 1.0]; 12];
bad[5][1] = f32::NAN;
assert!(HueSatTable::new(4, 3, 1, false, bad).is_none());
}
#[test]
fn blending_two_illuminants_is_entry_by_entry() {
let a = identity(2, 2, 1);
let mut b = identity(2, 2, 1);
b.entries[3] = [10.0, 2.0, 0.5];
let mid = a.lerp(&b, 0.5).unwrap();
assert_eq!(mid.entries[3], [5.0, 1.5, 0.75]);
assert_eq!(a.lerp(&b, 0.0).unwrap(), a);
assert_eq!(a.lerp(&b, 1.0).unwrap(), b);
assert!(a.lerp(&identity(3, 2, 1), 0.5).is_none());
}
}
+3
View File
@@ -9,12 +9,15 @@ use std::fmt;
use std::ops::Range;
pub mod colour;
pub mod hue_sat;
pub mod place;
pub mod selector;
pub mod settings;
pub mod time;
pub mod tone;
pub use colour::{Chromaticities, Transfer};
pub use hue_sat::{HueSatTable, ProfileOrigin, ProfileTables};
pub use place::{Place, PlaceScope, Screen, StoredFilter};
pub use selector::{ColourLabel, DateSelector, FlagState, Selector, Tier};
pub use settings::{
+28 -9
View File
@@ -275,13 +275,26 @@ impl FaceDetector {
}
}
/// Both ids this detector writes under — the f32 form and the int8 one —
/// for a question that is about the detector and not about which form
/// of it a device happened to run: "has the chosen detector been over
/// this image", asked by a re-index that must not ping-pong between a
/// desktop that runs it in f32 and a tablet that runs it on the Hexagon.
pub fn model_ids(self) -> [&'static str; 2] {
[self.model_id(), self.model_id_int8()]
/// The id when the detector runs with 16-bit activations and 8-bit
/// weights, the Hexagon's form since the int8 one lost faces at 40–80 px
/// (docs/dev/inference.md §1.5). Different again from both, for the same
/// reason as [`Self::model_id_int8`]; the embedder half is unchanged.
pub fn model_id_a16w8(self) -> &'static str {
match self {
FaceDetector::Scrfd500m => "scrfd_500m_a16+w600k_mbf",
FaceDetector::Scrfd2_5g => "scrfd_2.5g_a16+w600k_mbf",
FaceDetector::Scrfd10g => "scrfd_10g_a16+w600k_mbf",
}
}
/// Every id this detector writes under — the f32 form and each quantised
/// one a device has run — for a question that is about the detector and
/// not about which form of it a device happened to run: "has the chosen
/// detector been over this image", asked by a re-index that must not
/// ping-pong between a desktop that runs it in f32 and a tablet that
/// runs it on the Hexagon.
pub fn model_ids(self) -> [&'static str; 3] {
[self.model_id(), self.model_id_int8(), self.model_id_a16w8()]
}
/// The detector that writes under a pipeline id, if it is one of these.
@@ -1769,11 +1782,16 @@ mod tests {
#[test]
fn a_detectors_two_spellings_share_its_embedder_and_nothing_else() {
for d in FaceDetector::ALL {
let [f32_id, int8_id] = d.model_ids();
let [f32_id, int8_id, a16_id] = d.model_ids();
assert_eq!(f32_id, d.model_id());
assert_eq!(int8_id, d.model_id_int8());
assert_eq!(a16_id, d.model_id_a16w8());
assert_ne!(f32_id, int8_id);
assert_eq!(f32_id.rsplit('+').next(), int8_id.rsplit('+').next());
assert_ne!(f32_id, a16_id);
assert_ne!(int8_id, a16_id);
for q in [int8_id, a16_id] {
assert_eq!(f32_id.rsplit('+').next(), q.rsplit('+').next());
}
}
}
@@ -1797,6 +1815,7 @@ mod tests {
}
assert_eq!(FaceDetector::for_model_id(d.model_id()), Some(d));
assert_eq!(FaceDetector::for_model_id(d.model_id_int8()), Some(d));
assert_eq!(FaceDetector::for_model_id(d.model_id_a16w8()), Some(d));
}
assert_eq!(FaceDetector::for_model_id("scrfd_10g+other"), None);
}
+284
View File
@@ -0,0 +1,284 @@
//! TRACES: FR-DEV-3j | FR-DEV-3e
//! The DNG SDK's reference tone curve, as data (D21).
//!
//! # Where the numbers come from
//!
//! [`ACR3_DEFAULT`] is the "ACR3 default" tone curve of Adobe's DNG SDK
//! (`dng_tone_curve_acr3_default`): what the SDK's reference renders a raw through
//! when its camera profile carries no `ProfileToneCurve` of its own, which is
//! true of Adobe Standard. The values are RawTherapee's
//! `adobe_camera_raw_default_curve` (`rtengine/dcp.cc`, GPLv3), copied
//! digit for digit; `the_default_curve_is_rawtherapees` checks a sample of
//! them.
//!
//! It maps linear values to linear values, 1025 samples evenly over
//! `[0, 1]`, interpolated linearly between them. See
//! `docs/dev/camera-profiles.md` §12 for how it is applied — on the largest
//! and smallest channel, not on each — which is half of what it does.
/// Samples in a resolved tone curve: the ACR3 table's own resolution, and
/// what a profile's curve is resampled onto.
pub const TONE_SAMPLES: usize = 1025;
/// The ACR3 default tone curve, linear in, linear out.
///
/// `approx_constant` is allowed because one sample is 0.70711, which clippy
/// takes for an approximation of 1/√2. It is a measured value of the curve,
/// copied as published; replacing it with the constant would change it.
#[rustfmt::skip]
#[allow(clippy::approx_constant)]
pub const ACR3_DEFAULT: [f32; TONE_SAMPLES] = [
0.00000, 0.00078, 0.00160, 0.00242, 0.00314, 0.00385, 0.00460, 0.00539,
0.00623, 0.00712, 0.00806, 0.00906, 0.01012, 0.01122, 0.01238, 0.01359,
0.01485, 0.01616, 0.01751, 0.01890, 0.02033, 0.02180, 0.02331, 0.02485,
0.02643, 0.02804, 0.02967, 0.03134, 0.03303, 0.03475, 0.03648, 0.03824,
0.04002, 0.04181, 0.04362, 0.04545, 0.04730, 0.04916, 0.05103, 0.05292,
0.05483, 0.05675, 0.05868, 0.06063, 0.06259, 0.06457, 0.06655, 0.06856,
0.07057, 0.07259, 0.07463, 0.07668, 0.07874, 0.08081, 0.08290, 0.08499,
0.08710, 0.08921, 0.09134, 0.09348, 0.09563, 0.09779, 0.09996, 0.10214,
0.10433, 0.10652, 0.10873, 0.11095, 0.11318, 0.11541, 0.11766, 0.11991,
0.12218, 0.12445, 0.12673, 0.12902, 0.13132, 0.13363, 0.13595, 0.13827,
0.14061, 0.14295, 0.14530, 0.14765, 0.15002, 0.15239, 0.15477, 0.15716,
0.15956, 0.16197, 0.16438, 0.16680, 0.16923, 0.17166, 0.17410, 0.17655,
0.17901, 0.18148, 0.18395, 0.18643, 0.18891, 0.19141, 0.19391, 0.19641,
0.19893, 0.20145, 0.20398, 0.20651, 0.20905, 0.21160, 0.21416, 0.21672,
0.21929, 0.22185, 0.22440, 0.22696, 0.22950, 0.23204, 0.23458, 0.23711,
0.23963, 0.24215, 0.24466, 0.24717, 0.24967, 0.25216, 0.25465, 0.25713,
0.25961, 0.26208, 0.26454, 0.26700, 0.26945, 0.27189, 0.27433, 0.27676,
0.27918, 0.28160, 0.28401, 0.28641, 0.28881, 0.29120, 0.29358, 0.29596,
0.29833, 0.30069, 0.30305, 0.30540, 0.30774, 0.31008, 0.31241, 0.31473,
0.31704, 0.31935, 0.32165, 0.32395, 0.32623, 0.32851, 0.33079, 0.33305,
0.33531, 0.33756, 0.33981, 0.34205, 0.34428, 0.34650, 0.34872, 0.35093,
0.35313, 0.35532, 0.35751, 0.35969, 0.36187, 0.36404, 0.36620, 0.36835,
0.37050, 0.37264, 0.37477, 0.37689, 0.37901, 0.38112, 0.38323, 0.38533,
0.38742, 0.38950, 0.39158, 0.39365, 0.39571, 0.39777, 0.39982, 0.40186,
0.40389, 0.40592, 0.40794, 0.40996, 0.41197, 0.41397, 0.41596, 0.41795,
0.41993, 0.42191, 0.42388, 0.42584, 0.42779, 0.42974, 0.43168, 0.43362,
0.43554, 0.43747, 0.43938, 0.44129, 0.44319, 0.44509, 0.44698, 0.44886,
0.45073, 0.45260, 0.45447, 0.45632, 0.45817, 0.46002, 0.46186, 0.46369,
0.46551, 0.46733, 0.46914, 0.47095, 0.47275, 0.47454, 0.47633, 0.47811,
0.47989, 0.48166, 0.48342, 0.48518, 0.48693, 0.48867, 0.49041, 0.49214,
0.49387, 0.49559, 0.49730, 0.49901, 0.50072, 0.50241, 0.50410, 0.50579,
0.50747, 0.50914, 0.51081, 0.51247, 0.51413, 0.51578, 0.51742, 0.51906,
0.52069, 0.52232, 0.52394, 0.52556, 0.52717, 0.52878, 0.53038, 0.53197,
0.53356, 0.53514, 0.53672, 0.53829, 0.53986, 0.54142, 0.54297, 0.54452,
0.54607, 0.54761, 0.54914, 0.55067, 0.55220, 0.55371, 0.55523, 0.55673,
0.55824, 0.55973, 0.56123, 0.56271, 0.56420, 0.56567, 0.56715, 0.56861,
0.57007, 0.57153, 0.57298, 0.57443, 0.57587, 0.57731, 0.57874, 0.58017,
0.58159, 0.58301, 0.58443, 0.58583, 0.58724, 0.58864, 0.59003, 0.59142,
0.59281, 0.59419, 0.59556, 0.59694, 0.59830, 0.59966, 0.60102, 0.60238,
0.60373, 0.60507, 0.60641, 0.60775, 0.60908, 0.61040, 0.61173, 0.61305,
0.61436, 0.61567, 0.61698, 0.61828, 0.61957, 0.62087, 0.62216, 0.62344,
0.62472, 0.62600, 0.62727, 0.62854, 0.62980, 0.63106, 0.63232, 0.63357,
0.63482, 0.63606, 0.63730, 0.63854, 0.63977, 0.64100, 0.64222, 0.64344,
0.64466, 0.64587, 0.64708, 0.64829, 0.64949, 0.65069, 0.65188, 0.65307,
0.65426, 0.65544, 0.65662, 0.65779, 0.65897, 0.66013, 0.66130, 0.66246,
0.66362, 0.66477, 0.66592, 0.66707, 0.66821, 0.66935, 0.67048, 0.67162,
0.67275, 0.67387, 0.67499, 0.67611, 0.67723, 0.67834, 0.67945, 0.68055,
0.68165, 0.68275, 0.68385, 0.68494, 0.68603, 0.68711, 0.68819, 0.68927,
0.69035, 0.69142, 0.69249, 0.69355, 0.69461, 0.69567, 0.69673, 0.69778,
0.69883, 0.69988, 0.70092, 0.70196, 0.70300, 0.70403, 0.70506, 0.70609,
0.70711, 0.70813, 0.70915, 0.71017, 0.71118, 0.71219, 0.71319, 0.71420,
0.71520, 0.71620, 0.71719, 0.71818, 0.71917, 0.72016, 0.72114, 0.72212,
0.72309, 0.72407, 0.72504, 0.72601, 0.72697, 0.72794, 0.72890, 0.72985,
0.73081, 0.73176, 0.73271, 0.73365, 0.73460, 0.73554, 0.73647, 0.73741,
0.73834, 0.73927, 0.74020, 0.74112, 0.74204, 0.74296, 0.74388, 0.74479,
0.74570, 0.74661, 0.74751, 0.74842, 0.74932, 0.75021, 0.75111, 0.75200,
0.75289, 0.75378, 0.75466, 0.75555, 0.75643, 0.75730, 0.75818, 0.75905,
0.75992, 0.76079, 0.76165, 0.76251, 0.76337, 0.76423, 0.76508, 0.76594,
0.76679, 0.76763, 0.76848, 0.76932, 0.77016, 0.77100, 0.77183, 0.77267,
0.77350, 0.77432, 0.77515, 0.77597, 0.77680, 0.77761, 0.77843, 0.77924,
0.78006, 0.78087, 0.78167, 0.78248, 0.78328, 0.78408, 0.78488, 0.78568,
0.78647, 0.78726, 0.78805, 0.78884, 0.78962, 0.79040, 0.79118, 0.79196,
0.79274, 0.79351, 0.79428, 0.79505, 0.79582, 0.79658, 0.79735, 0.79811,
0.79887, 0.79962, 0.80038, 0.80113, 0.80188, 0.80263, 0.80337, 0.80412,
0.80486, 0.80560, 0.80634, 0.80707, 0.80780, 0.80854, 0.80926, 0.80999,
0.81072, 0.81144, 0.81216, 0.81288, 0.81360, 0.81431, 0.81503, 0.81574,
0.81645, 0.81715, 0.81786, 0.81856, 0.81926, 0.81996, 0.82066, 0.82135,
0.82205, 0.82274, 0.82343, 0.82412, 0.82480, 0.82549, 0.82617, 0.82685,
0.82753, 0.82820, 0.82888, 0.82955, 0.83022, 0.83089, 0.83155, 0.83222,
0.83288, 0.83354, 0.83420, 0.83486, 0.83552, 0.83617, 0.83682, 0.83747,
0.83812, 0.83877, 0.83941, 0.84005, 0.84069, 0.84133, 0.84197, 0.84261,
0.84324, 0.84387, 0.84450, 0.84513, 0.84576, 0.84639, 0.84701, 0.84763,
0.84825, 0.84887, 0.84949, 0.85010, 0.85071, 0.85132, 0.85193, 0.85254,
0.85315, 0.85375, 0.85436, 0.85496, 0.85556, 0.85615, 0.85675, 0.85735,
0.85794, 0.85853, 0.85912, 0.85971, 0.86029, 0.86088, 0.86146, 0.86204,
0.86262, 0.86320, 0.86378, 0.86435, 0.86493, 0.86550, 0.86607, 0.86664,
0.86720, 0.86777, 0.86833, 0.86889, 0.86945, 0.87001, 0.87057, 0.87113,
0.87168, 0.87223, 0.87278, 0.87333, 0.87388, 0.87443, 0.87497, 0.87552,
0.87606, 0.87660, 0.87714, 0.87768, 0.87821, 0.87875, 0.87928, 0.87981,
0.88034, 0.88087, 0.88140, 0.88192, 0.88244, 0.88297, 0.88349, 0.88401,
0.88453, 0.88504, 0.88556, 0.88607, 0.88658, 0.88709, 0.88760, 0.88811,
0.88862, 0.88912, 0.88963, 0.89013, 0.89063, 0.89113, 0.89163, 0.89212,
0.89262, 0.89311, 0.89360, 0.89409, 0.89458, 0.89507, 0.89556, 0.89604,
0.89653, 0.89701, 0.89749, 0.89797, 0.89845, 0.89892, 0.89940, 0.89987,
0.90035, 0.90082, 0.90129, 0.90176, 0.90222, 0.90269, 0.90316, 0.90362,
0.90408, 0.90454, 0.90500, 0.90546, 0.90592, 0.90637, 0.90683, 0.90728,
0.90773, 0.90818, 0.90863, 0.90908, 0.90952, 0.90997, 0.91041, 0.91085,
0.91130, 0.91173, 0.91217, 0.91261, 0.91305, 0.91348, 0.91392, 0.91435,
0.91478, 0.91521, 0.91564, 0.91606, 0.91649, 0.91691, 0.91734, 0.91776,
0.91818, 0.91860, 0.91902, 0.91944, 0.91985, 0.92027, 0.92068, 0.92109,
0.92150, 0.92191, 0.92232, 0.92273, 0.92314, 0.92354, 0.92395, 0.92435,
0.92475, 0.92515, 0.92555, 0.92595, 0.92634, 0.92674, 0.92713, 0.92753,
0.92792, 0.92831, 0.92870, 0.92909, 0.92947, 0.92986, 0.93025, 0.93063,
0.93101, 0.93139, 0.93177, 0.93215, 0.93253, 0.93291, 0.93328, 0.93366,
0.93403, 0.93440, 0.93478, 0.93515, 0.93551, 0.93588, 0.93625, 0.93661,
0.93698, 0.93734, 0.93770, 0.93807, 0.93843, 0.93878, 0.93914, 0.93950,
0.93986, 0.94021, 0.94056, 0.94092, 0.94127, 0.94162, 0.94197, 0.94231,
0.94266, 0.94301, 0.94335, 0.94369, 0.94404, 0.94438, 0.94472, 0.94506,
0.94540, 0.94573, 0.94607, 0.94641, 0.94674, 0.94707, 0.94740, 0.94774,
0.94807, 0.94839, 0.94872, 0.94905, 0.94937, 0.94970, 0.95002, 0.95035,
0.95067, 0.95099, 0.95131, 0.95163, 0.95194, 0.95226, 0.95257, 0.95289,
0.95320, 0.95351, 0.95383, 0.95414, 0.95445, 0.95475, 0.95506, 0.95537,
0.95567, 0.95598, 0.95628, 0.95658, 0.95688, 0.95718, 0.95748, 0.95778,
0.95808, 0.95838, 0.95867, 0.95897, 0.95926, 0.95955, 0.95984, 0.96013,
0.96042, 0.96071, 0.96100, 0.96129, 0.96157, 0.96186, 0.96214, 0.96242,
0.96271, 0.96299, 0.96327, 0.96355, 0.96382, 0.96410, 0.96438, 0.96465,
0.96493, 0.96520, 0.96547, 0.96574, 0.96602, 0.96629, 0.96655, 0.96682,
0.96709, 0.96735, 0.96762, 0.96788, 0.96815, 0.96841, 0.96867, 0.96893,
0.96919, 0.96945, 0.96971, 0.96996, 0.97022, 0.97047, 0.97073, 0.97098,
0.97123, 0.97149, 0.97174, 0.97199, 0.97223, 0.97248, 0.97273, 0.97297,
0.97322, 0.97346, 0.97371, 0.97395, 0.97419, 0.97443, 0.97467, 0.97491,
0.97515, 0.97539, 0.97562, 0.97586, 0.97609, 0.97633, 0.97656, 0.97679,
0.97702, 0.97725, 0.97748, 0.97771, 0.97794, 0.97817, 0.97839, 0.97862,
0.97884, 0.97907, 0.97929, 0.97951, 0.97973, 0.97995, 0.98017, 0.98039,
0.98061, 0.98082, 0.98104, 0.98125, 0.98147, 0.98168, 0.98189, 0.98211,
0.98232, 0.98253, 0.98274, 0.98295, 0.98315, 0.98336, 0.98357, 0.98377,
0.98398, 0.98418, 0.98438, 0.98458, 0.98478, 0.98498, 0.98518, 0.98538,
0.98558, 0.98578, 0.98597, 0.98617, 0.98636, 0.98656, 0.98675, 0.98694,
0.98714, 0.98733, 0.98752, 0.98771, 0.98789, 0.98808, 0.98827, 0.98845,
0.98864, 0.98882, 0.98901, 0.98919, 0.98937, 0.98955, 0.98973, 0.98991,
0.99009, 0.99027, 0.99045, 0.99063, 0.99080, 0.99098, 0.99115, 0.99133,
0.99150, 0.99167, 0.99184, 0.99201, 0.99218, 0.99235, 0.99252, 0.99269,
0.99285, 0.99302, 0.99319, 0.99335, 0.99351, 0.99368, 0.99384, 0.99400,
0.99416, 0.99432, 0.99448, 0.99464, 0.99480, 0.99495, 0.99511, 0.99527,
0.99542, 0.99558, 0.99573, 0.99588, 0.99603, 0.99619, 0.99634, 0.99649,
0.99664, 0.99678, 0.99693, 0.99708, 0.99722, 0.99737, 0.99751, 0.99766,
0.99780, 0.99794, 0.99809, 0.99823, 0.99837, 0.99851, 0.99865, 0.99879,
0.99892, 0.99906, 0.99920, 0.99933, 0.99947, 0.99960, 0.99974, 0.99987,
1.00000,
];
/// TRACES: FR-DEV-3e
/// A profile's `ProfileToneCurve` — `(x, y)` pairs in `[0, 1]` — resampled
/// onto [`TONE_SAMPLES`] even points with a natural cubic spline, the DNG
/// SDK's `dng_spline_solver`.
///
/// `None` where the pairs do not describe a curve (fewer than two points,
/// x not increasing, values outside `[0, 1]`) or describe the identity,
/// which RawTherapee also treats as no curve.
pub fn resample_tone_curve(pairs: &[f32]) -> Option<Vec<f32>> {
if pairs.len() < 4 || !pairs.len().is_multiple_of(2) {
return None;
}
let xs: Vec<f64> = pairs.iter().step_by(2).map(|&v| f64::from(v)).collect();
let ys: Vec<f64> = pairs
.iter()
.skip(1)
.step_by(2)
.map(|&v| f64::from(v))
.collect();
let sane = xs.windows(2).all(|w| w[1] > w[0])
&& xs
.iter()
.chain(&ys)
.all(|v| v.is_finite() && (-1e-6..=1.0 + 1e-6).contains(v));
if !sane {
return None;
}
if xs.iter().zip(&ys).all(|(x, y)| (x - y).abs() < 1e-6) {
return None;
}
let n = xs.len();
// Natural cubic spline: second derivatives zero at both ends, solved by
// the tridiagonal (Thomas) algorithm.
let mut m = vec![0.0f64; n];
if n > 2 {
let h: Vec<f64> = xs.windows(2).map(|w| w[1] - w[0]).collect();
let mut a = vec![0.0; n];
let mut b = vec![1.0; n];
let mut c = vec![0.0; n];
let mut d = vec![0.0; n];
for i in 1..n - 1 {
a[i] = h[i - 1];
b[i] = 2.0 * (h[i - 1] + h[i]);
c[i] = h[i];
d[i] = 6.0 * ((ys[i + 1] - ys[i]) / h[i] - (ys[i] - ys[i - 1]) / h[i - 1]);
}
for i in 1..n {
let w = a[i] / b[i - 1];
b[i] -= w * c[i - 1];
d[i] -= w * d[i - 1];
}
m[n - 1] = d[n - 1] / b[n - 1];
for i in (0..n - 1).rev() {
m[i] = (d[i] - c[i] * m[i + 1]) / b[i];
}
}
let eval = |x: f64| -> f64 {
if x <= xs[0] {
return ys[0];
}
if x >= xs[n - 1] {
return ys[n - 1];
}
let i = xs.windows(2).position(|w| x <= w[1]).unwrap_or(n - 2);
let h = xs[i + 1] - xs[i];
let t = (x - xs[i]) / h;
let u = 1.0 - t;
u * ys[i]
+ t * ys[i + 1]
+ ((u * u * u - u) * m[i] + (t * t * t - t) * m[i + 1]) * h * h / 6.0
};
Some(
(0..TONE_SAMPLES)
.map(|k| eval(k as f64 / (TONE_SAMPLES - 1) as f64).clamp(0.0, 1.0) as f32)
.collect(),
)
}
/// A resolved curve at `x`, linear interpolation between samples, `x`
/// clamped to `[0, 1]` — the shader's lookup, for the CPU reference.
pub fn evaluate(curve: &[f32], x: f32) -> f32 {
let last = curve.len() - 1;
let s = x.clamp(0.0, 1.0) * last as f32;
let i = (s as usize).min(last - 1);
let f = s - i as f32;
curve[i] + (curve[i + 1] - curve[i]) * f
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn the_default_curve_is_rawtherapees() {
assert_eq!(ACR3_DEFAULT.len(), 1025);
assert_eq!(ACR3_DEFAULT[0], 0.0);
assert_eq!(ACR3_DEFAULT[1024], 1.0);
assert_eq!(ACR3_DEFAULT[1], 0.00078);
assert_eq!(ACR3_DEFAULT[256], 0.52069);
assert_eq!(ACR3_DEFAULT[512], 0.80486);
assert_eq!(ACR3_DEFAULT[768], 0.93986);
assert!(ACR3_DEFAULT.windows(2).all(|w| w[1] >= w[0]), "monotone");
}
#[test]
fn a_resampled_curve_passes_through_its_points() {
let c = resample_tone_curve(&[0.0, 0.0, 0.5, 0.6, 1.0, 1.0]).unwrap();
assert_eq!(c.len(), TONE_SAMPLES);
assert!((evaluate(&c, 0.5) - 0.6).abs() < 1e-4);
assert!(evaluate(&c, 0.0).abs() < 1e-6 && (evaluate(&c, 1.0) - 1.0).abs() < 1e-6);
assert!(
evaluate(&c, 0.25) > 0.25,
"a lifted curve lifts between its points"
);
}
#[test]
fn an_identity_or_broken_curve_is_no_curve() {
assert!(resample_tone_curve(&[0.0, 0.0, 1.0, 1.0]).is_none());
assert!(resample_tone_curve(&[0.0, 0.0, 0.5]).is_none());
assert!(resample_tone_curve(&[0.0, 0.0, 0.6, 0.5, 0.4, 0.9, 1.0, 1.0]).is_none());
}
}
+10 -4
View File
@@ -263,10 +263,12 @@ cp "${DEX}" "${OUT}/staging/classes.dex"
# native library directory. The build links none of it — the app dlopens
# `libonnxruntime.so` at launch and runs on tract if it is not there — so an
# APK without these is a slower app, not a broken one, and `RUNTIME_DIR=none`
# builds exactly that. 174 MB for the default set; the script says which
# Hexagon generations that buys.
# builds exactly that. 206 MB for the default set — 32 MB of it the generic
# WebGPU build for a phone without a Qualcomm SoC; the script says which
# Hexagon generations the rest buys.
if [[ "${RUNTIME_DIR}" != "none" ]]; then
if [[ ! -f "${RUNTIME_DIR}/lib/libonnxruntime.so" ]]; then
if [[ ! -f "${RUNTIME_DIR}/lib/libonnxruntime.so" \
|| ! -f "${RUNTIME_DIR}/lib/libonnxruntime_generic.so" ]]; then
"${REPO}/tools/fetch-android-runtime.sh" "${RUNTIME_DIR}"
fi
cp "${RUNTIME_DIR}"/lib/*.so "${OUT}/staging/lib/${ABI}/"
@@ -299,7 +301,7 @@ fi
rm -rf "${OUT}/staging/assets/models"
mkdir -p "${OUT}/staging/assets/models"
_bundled=""
for _dir in face scene inpaint; do
for _dir in face scene inpaint denoise; do
ASSETS="${REPO}/models/${_dir}"
compgen -G "${ASSETS}/*.onnx" >/dev/null || continue
# An LFS pointer is ~130 bytes and looks exactly like a model to `cp`. Left
@@ -321,6 +323,10 @@ for _dir in face scene inpaint; do
for f in "${ASSETS}"/*; do
case "$(basename "${f}")" in
README.md) continue ;;
# The denoisers' any-size exports run whole frames on TensorRT
# and CUDA (denoise.md §14); the Hexagon takes fixed shapes, and
# 16 MB of graphs it never loads stay out of the APK.
mosaic-fast.onnx | mosaic-hq.onnx) continue ;;
esac
cp "${f}" "${OUT}/staging/assets/models/"
_bundled="${_bundled} $(basename "${f}")"
+3
View File
@@ -31,6 +31,9 @@ ENV DEBIAN_FRONTEND=noninteractive \
# ---------------------------------------------------------------------------
RUN apt-get update && apt-get install -y --no-install-recommends \
ca-certificates curl git git-lfs \
# The bundled ONNX Runtime builds arrive as wheels, which are zips
# (tools/fetch-bundled-runtimes.sh).
unzip \
# A *host* C compiler as well as the cross one: build scripts and
# proc-macros are compiled for Linux and linked with `cc`, whatever
# the target. Without it the very first build script fails with
+14 -3
View File
@@ -51,13 +51,17 @@ sed 's/$/\r/' "${REPO}/LICENSE" > "${STAGE}/LICENSE"
#
# The directories are the ones the APK stages (assemble-apk.sh) and the Arch
# package installs: the face pair and its eye-state models, the scene model
# with its two descriptors, and the panorama border filler. The installer
# with its two descriptors, the panorama border filler and the denoiser. The installer
# smoke test counts the same directories, so a model added here is expected
# there without a number to update.
for dir in face scene inpaint; do
#
# Not the quantised siblings (`*.int8.onnx`, `*.a16w8.onnx`, `*.a16w16.onnx`):
# they are the Hexagon's forms (docs/dev/inference.md §1.5), and a Windows
# machine has no Hexagon to load them.
for dir in face scene inpaint denoise; do
for f in "${REPO}/models/${dir}"/*; do
case "$(basename "${f}")" in
README.md) continue ;;
README.md | *.int8.onnx | *.a16w8.onnx | *.a16w16.onnx) continue ;;
esac
if [[ "${f}" == *.onnx && "$(stat -c%s "${f}")" -lt 100000 ]]; then
echo "error: $(basename "${f}") is $(stat -c%s "${f}") bytes — an LFS pointer, not a model." >&2
@@ -86,6 +90,13 @@ for f in "${REPO}/docs/manual/media"/*; do
done
echo "==> staged the manual and $(ls "${STAGE}/manual/media" | wc -l) picture(s)"
# The two ONNX Runtime builds the app chooses between at launch
# (docs/dev/inference.md §3.2): Intel's OpenVINO build for an Intel GPU and
# the WebGPU build — D3D12 — for any other, both carrying the CPU provider.
# Beside the executable under `runtimes\`, where darkroom-desktop looks.
"${REPO}/tools/fetch-bundled-runtimes.sh" windows "${STAGE}/runtimes"
echo "==> staged $(ls "${STAGE}/runtimes"/* | wc -l) runtime file(s)"
# One installer in the output directory, the one just built. The directory
# is cached between CI runs, so after a version bump a glob over it would find
# two and the smoke test would hand Wine both names as one path.
+391
View File
@@ -0,0 +1,391 @@
# Camera profiles — DCP tables on top of the matrix
Design for the deferred half of **FR-DEV-3e** ([requirements.md](requirements.md)): the
`HueSatMap` and `LookTable` of a DNG camera profile, read from the DNG that carries one or from a
`.dcp` file, and applied after the matrix. Draft of 2026-10-02, recorded as **D20**.
---
## 1. What we are matching
Lightroom renders a raw through a *profile* before any slider moves. A profile is the matrix
DarkRoom already applies, plus two lookup tables indexed by hue, saturation and value:
- **`ProfileHueSatMap`** — a calibration. It corrects what a 3×3 cannot: a sensor whose reds and
oranges sit in the wrong place relative to its blues, which no linear map fixes. Two copies, one
per calibration illuminant, interpolated like the matrices.
- **`ProfileLookTable`** — a rendering intent: hue shifts of up to ±18° by hue, and saturation
and value scales that vary with brightness. The difference between Adobe Standard, Adobe Color
and Adobe Vivid is largely this table, together with each profile's tone curve.
Without them a raw renders through the matrix alone, which is accurate on a ColorChecker and is
not what Lightroom showed for the same file.
**What the tables do *not* do is make a photograph more saturated.** Measured after the build
(2026-10-02), on four of the library's 6D DNGs rendered at defaults: Adobe Standard's tables
*lower* mean saturation by 3–9 %, and the look at 200 % lowers it further. The 6D's look table
scales saturation by 0.925 in its darkest value rows and by 1.0 from about a fifth of full scale
up, and its HueSatMap adds about 1 %. Adobe Standard was tuned to sit under Camera Raw's default
RGB tone curve, which raises saturation in the shadows and midtones, and the look's dark-tone
desaturation offsets it. Without that curve (§6) the offset is all that is left. On `_MG_9080`, Lightroom 6's own preview
measures 0.49 mean HSV saturation; the matrix alone renders 0.38, Adobe Standard's tables 0.35.
That preview also carries whatever was edited in Lightroom, so it is not a clean reference — but
the direction is unambiguous: **the gap to Lightroom's colour is mostly tone, not the profile's
tables.** The tables still matter for hue: they are what puts each body's reds, skin and foliage
where Adobe put them.
**What the library holds** (catalog of 2026-10-02): 17,286 of its raws are Canon EOS 6D. The
9,348 DNGs were written by Lightroom 6.14 and every one sampled embeds *Adobe Standard* with both
tables — `HueSatMapDims 90 30 1`, `LookTableDims 36 8 16`, `ProfileEmbedPolicy 0` ("allow
copying"), no `ProfileToneCurve`. The 7,938 CR2s from the same body carry no profile. So the
tables Lightroom used are already on disk for half the library, and their licence lets them be
applied to the other half.
## 2. The model, as the DNG specification states it
Each table is a grid of `(hueShift°, satScale, valScale)` triples over HSV, stored with
saturation varying fastest, then hue, then value:
```
index = v · (hueDivs · satDivs) + h · satDivs + s
```
Lookup follows the DNG SDK's `RefBaselineHueSatMap`:
- **HSV** is the SDK's: `v = max(r,g,b)`, `s = (v − min)/v`, `h ∈ [0, 6)` from which channel
leads. Grey has `s = 0` and is untouched by construction.
- **Hue wraps**: `hueDivs` samples over 360°, the last interpolating to the first.
- **Saturation** samples `0..=1` at `satDivs` points; linear between.
- **Value** samples `0..=1` at `valDivs` points when `valDivs > 1`; a table with `valDivs = 1` is
"2.5-D" and ignores value. `ProfileHueSatMapEncoding` / `ProfileLookTableEncoding` = 1 means the
value axis is indexed by the sRGB-encoded value; 0 (the default, and the 6D's) means linear.
- **Apply**: `h += hueShift · 6/360`, `s = min(s · satScale, 1)`, `v ·= valScale`; back to RGB.
- **Space**: linear ProPhoto (ROMM) primaries, D50 white — the space the forward matrix lands in.
- **Two illuminants**: `HueSatMapData1/2` are interpolated entry by entry, with the same mired
weight the matrices use. `LookTable` is single.
Two departures, both forced by D19's unbounded scene-linear values (the SDK runs these on `[0, 1]`):
1. **Value is not clamped.** The SDK writes `min(v · valScale, 1)`; here `v · valScale`, unbounded.
For lookup only, the value axis reads `min(v, 1)`, so a highlight above 1.0 uses the table's
brightest row. Where the encoding is sRGB the scale is defined on the encoded value; it is
applied as the ratio `decode(enc(v′)·valScale)/v′` at `v′ = min(v, 1)`, so a value above 1.0
gets the brightest row's ratio rather than a clip.
2. **A colour outside ProPhoto passes through.** A negative component has no SDK HSV. Such a
colour is outside every surface colour a camera records under normal light. It is left
unmodified rather than floored, because flooring it clips a value D19 says nothing may clip.
## 3. Where it sits
```
… camera matrix ─► vignetting(5) ─► exposure(20) ─► camera_profile(25) ─► contrast(30) ─► … ─► view transform
│
working → ProPhoto ─► HueSatMap ─► LookTable ─► ProPhoto → working
```
**A scene operation at order 25, not part of the matrix snippet.** Three reasons:
- **The matrix stays what it is.** `cam_to_srgb` is unchanged and still runs where D19 put it, and
so does every reader of it: the mask pass's copy, the white-balance picker's, the camera-space
tap. The tables add a conversion into ProPhoto and back *inside* their own fragment, through two
constant matrices (§3.1). A photograph with no profile composes exactly the shader it does today.
- **It commutes with what runs before it.** HSV hue and saturation are invariant under a uniform
gain, and vignetting and exposure are uniform gains. So a 2.5-D table — every Adobe HueSatMap
seen, and the 6D's — gives the same answer before or after them. That lets one position serve
both tables, which is where the second reason matters:
- **The look sees exposure.** The SDK applies `LookTable` after its exposure ramp, so a look that
desaturates highlights finds the highlights the photographer chose. At 25 it does too. Contrast,
tone and the colour controls come after it, as they do in the DNG SDK's reference rendering.
**Tables are per source, like the matrix.** They are decoded with the raw, interpolated once at
decode (the HueSatMap blend uses the as-shot neutral, as the matrix does) and carried on
`DemosaicedImage` next to `color_matrix`. `dr-gpu` uploads them to a storage buffer at
`@binding(8)` and writes their dimensions into the base uniform block. Every render path that
reaches `AdjustPass` therefore gets them without being told: develop, export, previews, the tablet.
A path that had to call a setter on the graph would be a path that one day forgot to, and an export
that differed from the screen would be the result.
### 3.1 The two constants
`P⁻¹` is `ColourSpace::ProPhoto.from_linear_srgb()` — the conversion `dr-types` already derives
from the two spaces' chromaticities, adapting D65 to D50 by Bradford, which the export path uses to
write ProPhoto files — and `P` is its inverse. The fragment uses `P⁻¹` going in and `P` coming
out. Each row of both is scaled to sum to one, so working-space white is ProPhoto white exactly and
a neutral reaches the tables at `s = 0`. For a profile with forward matrices this recovers the
SDK's ProPhoto colour to within the difference between that derivation and `forward_to_srgb`'s
published Bradford constants, which is rounding.
## 4. Where a profile comes from
In this order, first match wins:
1. **The profile embedded in the DNG being opened.** It is what the file says, and it was made for
the matrices the file carries. Read from the root IFD through rawler's parsed `IFD`, as
`read_dng_matrices` already reads the forward matrices — no second TIFF parser.
2. **A `.dcp` in the profiles directory** whose `UniqueCameraModel` matches the body — compared
case-insensitively against the DNG's `UniqueCameraModel` where there is one, and against
`make + " " + model` otherwise ("Canon EOS 6D"). A DCP is a whole profile: its matrices replace
the file's, because its tables were built against its forward matrix. The file's as-shot neutral
is kept. If several match, the first by file name wins, so the choice is stable.
3. **None.** The matrix alone, as today.
The profiles directory is `profiles/` under the platform data directory (`dr_plat::dirs`), loaded
once per process. A DCP is a TIFF with the magic `IIRC` (0x4352) in place of 42; rawler's
`GenericTiffReader` already accepts it.
**Copying an embedded profile out.** A DNG whose profile has `ProfileEmbedPolicy` 0 ("allow
copying") or 3 ("no restrictions") can have that profile saved as a `.dcp` into the profiles
directory. That is how the 6D's CR2s get Adobe Standard: open a 6D DNG, choose *Use this profile
for every Canon EOS 6D*. Policies 1 ("embed if used") and 2 ("embed never") offer no such action.
The written file carries the profile's name, copyright and policy unchanged.
**Nothing is shipped.** Adobe's profiles are Adobe's; the application ships no `.dcp` and copies
none on its own. A profile reaches the directory because the photographer put it there or asked
for it to be copied from their own file.
## 5. The control
A develop operation, `camera_profile`, `[colour]`, order 25, hand-written (`rust:`) because it
reads a buffer no declaration can name:
| Parameter | Kind | Default | Meaning |
|---|---|---|---|
| `apply` | Bool | on | Use the profile's tables, or the matrix alone |
| `look` | Scalar 0–200 | 100 | Strength of the `LookTable` |
`look` scales the look's deltas: `hueShift · a`, `1 + (satScale − 1)·a`, `1 + (valScale − 1)·a`,
with `a = look/100`, scales floored at 0. At 200 the look is twice as strong, which is the
"more vivid than Adobe Standard" this started from. The HueSatMap is a calibration and is not
scaled: `apply` is its only switch.
**Always composed while `apply` is on**, as the view transform is: a profile at its defaults *is*
the rendering, not an edit, so an untouched photograph writes no parameters and still renders
through its profile. The fragment branches on the uniform that says whether the source has tables,
so a JPEG, or a raw with none, pays one uniform read. The branch is uniform across the dispatch.
**Mask layers.** A layer may offset `look` (blended as a setting, which is linear) but not `apply`;
the photograph has one profile.
**The panel says which profile is in use**, as the lens line does: *Adobe Standard (in the file)*,
*Adobe Standard (Canon EOS 6D.dcp)*, or *No profile for this camera — matrix only*. The copy-out
action sits on that line.
## 6. Not done, and why
- ~~**`ProfileToneCurve` is read and ignored.**~~ *Done after 0.20.0: §12, D21.* It came back as
an option of the view transform, as this bullet said it would, and that option is the default
for raws.
- ~~**`BaselineExposure` is not applied.**~~ *Done after 0.20.0: §11.*
- **The interpolation follows the as-shot neutral, not the white-balance slider**, as the matrix
does. Camera Raw re-blends on every temperature change; doing so here means the matrix moves too,
which is its own change.
- **Masks select on the matrix's colour.** A colour-range mask sees colour before the profile, as
it sees colour before every other operation. Deterministic, and a mask is drawn on the picture
the user sees only approximately anyway.
- ~~**The profiles directory does not sync.**~~ *Done after 0.20.0: §13.*
- **Rec.2020 working primaries** stay deferred (D19); nothing here depends on them.
## 7. What it costs
- **Every DNG with an embedded profile renders differently** — more saturated, which is the point.
Previews rendered before the change keep the old look until rendered again, as with D19.
- **Tablet and desktop must be released together.** No schema change, and the sidecar gains only
ordinary parameters, but two builds render the same DNG differently.
- **One storage-buffer binding** in every generated shader's layout (a one-entry placeholder when
there are no tables), and two vec4 slots in the base uniform block.
- **Per pixel**: two 3×3 multiplies, two HSV round trips, and 4 + 8 buffer reads (bilinear
HueSatMap, trilinear LookTable). Small next to the fused pass it joins.
## 8. Acceptance
- **Parsing.** The 6D DNG's embedded profile parses to `90×30×1` and `36×8×16`, its policy to 0,
its name to "Adobe Standard"; a `.dcp` written from it parses back to the same tables bit for
bit.
- **The CPU reference matches the SDK's algorithm**: grey passes through; a table of
`(0°, 1, 1)` everywhere is the identity to 1e-6; a uniform `satScale` of 1.2 scales HSV
saturation by 1.2; hue interpolation wraps between the last and first column.
- **The shader agrees with the CPU reference** on a device, within two 8-bit codes of the
display-encoded readback (the only readback the adjust pass has), over 256 colours that tables
of tens of degrees and ±30 % saturation move, and over the library's real Adobe Standard tables
(`dr-gpu/tests/camera_profile.rs`).
- **Scene-referred.** A value above 1.0 leaves the stage above 1.0 (`scene_referred_until_the_view`
covers the operation).
- **Neutral.** `apply` off renders to the bit what a source with no tables renders.
- **Two illuminants.** A HueSatMap at blend weight 0 is Data1, at 1 is Data2.
- **Matching.** An embedded profile beats a directory one; a DCP for "Canon EOS 6D" matches a CR2
whose rawler make/model is "Canon"/"EOS 6D"; no match leaves the matrix.
- **Subjective.** A 6D DNG rendered here at defaults is visibly closer to the same file in
Lightroom 6 with Adobe Standard than the matrix-only render, side by side.
## 9. Vivid presets
Independent of the tables, and — given §1's measurement — the part of this change that actually
answers "more colourful". Shipped in the same change: a *Vivid* section of read-only presets
(`presets/vivid.drpl`) for the "more colourful than the default" request. They use only operations
every photograph has — vibrance, saturation, the colour mixer, colour grading, contrast — so they
work on JPEGs and on bodies with no profile, and change only what they name (FR-DEV-6):
- **Vivid** — vibrance and a little saturation and contrast: the general-purpose one.
- **Vivid, strong** — the same, pushed, with deeper blacks.
- **Vivid landscape** — greens, blues and azure skies, skin bands left alone.
- **Vivid warm** — oranges and yellows up, a warm highlight cast: golden hour.
- **Vivid portrait** — vibrance (which protects skin) with the orange and red bands held back.
Measured on `_MG_9080` (mean HSV saturation; Lightroom's preview 0.49, DarkRoom's default 0.35):
Vivid 0.40, Vivid strong 0.44, Vivid landscape 0.46, Vivid warm 0.38, Vivid portrait 0.37 — the
last two move particular bands, not the whole frame. Rendered with `cargo run --release -p dr-gpu
--example develop -- FILE.dng out.ppm "preset:Vivid"`.
They are bounded by the existing `bundled.rs` tests: every key names a real parameter, every value
is inside its control's range, and every preset changes something.
## 10. Build order
1. `dr-decode`: parse the tables (embedded and `.dcp`), the profiles directory, matching, blending,
the `.dcp` writer. CPU reference of the lookup. Unit tests against the library's 6D DNG,
skipped when it is absent.
2. `dr-pipeline`: the `camera_profile` operation, the base-block slots, `@binding(8)`, the WGSL
lookup; composition tests.
3. `dr-gpu`: carry the tables on `DemosaicedImage`, upload and bind them; the shader-versus-CPU
test on a device.
4. `dr-ui`: profile line, copy-out action, labels; the profiles directory set at start-up on
desktop and Android.
5. The *Vivid* presets.
---
# After 0.20.0: tone, exposure and sync
0.20.0 shipped the tables and §1's measurement showed they were not the gap to Lightroom's colour.
The three items §6 left open are closed here. Draft of 2026-10-03.
## 11. Baseline exposure
`BaselineExposure` (DNG tag 50730) is the stops a converter adds so a camera's middle grey lands
where its maker meant it; the 6D's DNGs say +0.25. `BaselineExposureOffset` (51109) is a profile's
correction to it. The DNG SDK's total is their sum, and so is this one's.
- **Applied as a gain on the camera matrix** in `dr-gpu` when a raw is uploaded:
`cam_to_srgb · 2^total`. A uniform gain commutes with every scene operation before the view
transform, and the camera-space tap and the white-balance probe read camera RGB before the
matrix, so neither changes. `RawImage::color_matrix` itself stays the file's: a merge writes a
linear DNG from it and must not bake a gain into the pixels it also declares in a tag.
- **A copied profile carries the DNG's baseline.** When an embedded profile is saved as a `.dcp`
(§4) its `BaselineExposureOffset` is written as the DNG's `BaselineExposure` plus the profile's
own offset. A CR2 has no baseline of its own, so its total is then the DNG's: the two files of
one body render at one brightness. The cost: a DNG that embeds no profile, carries its own
baseline, and matches a copied `.dcp` counts the baseline twice. Every Adobe-written DNG embeds
its profile, so that DNG is a hand-made one.
## 12. DNG reference tone (D21)
The DNG SDK's reference rendering runs a raw through the profile's
`ProfileToneCurve`, or, for a profile that has none — Adobe Standard among them — through the
*ACR3 default curve*, a 1025-point table published in the DNG SDK and carried by RawTherapee
(GPLv3) as `adobe_camera_raw_default_curve`.
**How it is applied** is half of what it does. The reference does not run the curve on each channel:
`RefBaselineRGBTone` runs it on the largest and smallest channel and places the middle one at the
same fraction between them as before. Hue is kept; saturation rises where the curve is steep —
the shadows and midtones — which is exactly where Adobe Standard's look table desaturated to
compensate. It runs in linear ProPhoto, on values clipped to `[0, 1]`, and its output is linear.
**As the view transform**, not as a stage. D19 has one rendering, last; this is a second kind of
that rendering, chosen by a new parameter on `view_transform`:
| `curve` | What it is | Default for |
|---|---|---|
| Sigmoid | D19's log-logistic curve | — (a choice) |
| DNG reference | the profile's curve, else ACR3, via RGBTone in ProPhoto | every raw (D21) |
*Decided 2026-10-03:* the DNG reference is the default, at contrast 1.5 — measured against
Lightroom exports of photographs with neutral look settings (D21 has the table).
*Amended earlier on 2026-10-03:* the DNG reference was the default in the first draft. The measurement it rested on
compared against Lightroom renders of *edited* photographs; see D21. The default is decided by
measuring against Lightroom exports of unedited ones.
A JPEG is still not rendered again (FR-DEV-3j). Film simulation still replaces the view transform
when a stock is chosen.
**The two sliders keep meaning something** under the DNG reference curve:
- `white` (stops above grey at which the scene reaches display white) sets the input scale:
`2^(4 − white)`. At its default of 4 the scale is 1 — sensor white is display white, as in
the SDK's reference.
- `contrast` bends the input about middle grey before the curve, as a power of
`contrast / 1.4`: 1 at its default, so the curve is the reference's untouched.
**What the ACR curve gives up** is D19's shoulder. Values above display white clip, as they do
in the SDK's reference; highlight recovery is the highlights slider's job before it. Sigmoid stays one
click away for a photograph that wants the shoulder.
**A profile's own curve** is a list of `(x, y)` pairs. It is resampled at decode onto the same
1025 points with a natural cubic spline, the DNG SDK's `dng_spline_solver`. A curve that is the
identity is treated as absent, as RawTherapee does.
**On the GPU** the curve rides in the profile buffer (`@binding(8)`) after the tables: a third
header entry gives its length, then the samples. The placeholder bound for a source without a
profile carries the ACR3 curve, so a CR2 with no `.dcp` renders through the reference tone when that curve is chosen.
## 13. Profiles sync
The profiles directory travels with the library, in the server folder that already carries what
every device must agree on: `<library root>/.darkroom-derived/profiles/`. The scanner excludes
its parent, as it excludes the trash.
- **A step of the derived sync pass** (`derived_sync::run`), after the catalog and before the
place file, and like the place file it never fails the pass. It lists the server folder and the
local one; uploads every local `.dcp` the server lacks, or holds at a different size; downloads
every one the device lacks, reading through a placeholder where the library is a synced folder,
and writing `.tmp` then renaming so a half-written file is never parsed. If it fetched
anything, it reloads the profile set.
- **Files are immutable and named for what they hold** (`<camera> <profile>.dcp`), so a name and a
size say whether two copies are the same. Two devices that copy the same profile write the same
name; neither wins over anything.
- **Not a catalog table.** A schema change stops an older peer merging the catalog at all
(see the memory of 0.13.3), and a blob of ~120 KB would ride in every catalog upload.
- **One directory per install**, the union of every library's profiles. A profile describes a
camera, not a library, so a profile one library brought is right for the same camera in
another.
- **Not handled: deleting.** There is no way to remove a profile from the app; a file removed by
hand on one device comes back from the server on the next pass. A tombstone list is the
follow-up if deleting is added.
## 14. Acceptance for §11–§13
- The ACR3 table is 1025 points from 0 to 1, monotone, and matches RawTherapee's values.
- RGBTone: grey goes through the curve unchanged in hue; a colour keeps its hue (the middle
channel's fraction between the outer two is unchanged); a curve that is the identity changes
nothing; the shader agrees with the CPU reference on a device.
- Sigmoid is the default curve; at its defaults it renders to the bit what 0.20.0 rendered,
apart from baseline exposure.
- A DNG with `BaselineExposure` +0.25 renders a flat grey 0.25 EV brighter than the same pixels
with none; a `.dcp` copied from it carries `BaselineExposureOffset` 0.25 and gives a CR2 the same
total.
- A profile resampled from `(0,0) (0.5,0.6) (1,1)` passes through its points.
- Sync: a `.dcp` present only locally is uploaded; one present only on the server is downloaded,
parsed and matched on the next decode; a file of the same name and size is left alone.
- Measured again on `_MG_9080`: mean saturation at defaults closer to Lightroom's 0.49 than 0.20.0's
0.35.
## 15. The photographer's earlier edit
What §1 and D21 first took for a difference in rendering is an edit. Every DNG in the library
carries, in its embedded XMP, the develop settings it was given before it came to DarkRoom — a
consistent house style: per-colour saturation (blue +58, aqua +50, yellow and purple +23, orange
+13, green +10), highlights −40, blacks −20, with a second variant (vibrance −10, blue +31). Those
settings, not the profile and not the tone curve, are why the same photographs looked richer
before.
- **Translated on open** (`dr_preset_xmp::read_embedded`), the HSL bands onto the colour mixer —
aqua to cyan and purple to violet, the nearest of its twelve by hue — and the rest as the preset
importer already did.
- **Applied only to a photograph DarkRoom has no edit of**, and only on positive evidence: no
sidecar beside a local file, or a server that answered "no such file" with nothing cached
(`FetchedSidecar::absent`). An edit that failed to arrive is not an absent one, and this would
otherwise be saved over it.
- **One undoable step, "Earlier Edit"**, then an ordinary edit, saved with the photograph. Export
applies it the same way, so a photograph never opened exports as opening it would show.
- **The translation is one for one for now.** Measurement against the earlier exports (the
`lr-fit` work) may scale individual bands.
+280 -9
View File
@@ -1,8 +1,9 @@
# Learned denoise — joint demosaic and denoise on the mosaic
Design for **FR-DEV-3g** ([requirements.md](requirements.md)), the learned stage
[outstanding.md §3](outstanding.md) says is missing. Draft of 2026-09-27: nothing here is built,
and every figure marked *estimate* is waiting for the measurement that replaces it.
[outstanding.md §3](outstanding.md) says is missing. Drafted 2026-09-27; a first version shipped
in 0.21.0, and §11 records what was built and measured. Figures still marked *estimate* are
waiting for the measurement that replaces them.
---
@@ -288,17 +289,27 @@ is ~120 MB, and a derived file inside a synced tree is exactly what
with progress over the canvas — the same pattern as a photograph that is only on the server.
- Export needs the result and computes it if the cache has lost it.
### 7.2 The Amount control
### 7.2 The grain control
A Denoise toggle and one Amount slider in develop. Moving the slider runs inference on the
**visible viewport only** (~1 MP, a fraction of a second — *estimate*) so the photographer judges
on the real result; releasing it queues the whole frame. There is no per-frame blend between the
two paths: blending the classical output back in re-adds the noise the network removed.
What shipped is a switch and a **Keep grain** slider, not the Amount described first. The slider
blends the two demosaics per pixel — but only the *brightness* of their difference: `out =
denoised + grain · ΔY / wb`, with `ΔY` the luminance of `wb · (classical − denoised)`. Taken after
the as-shot balance and handed back divided by it, the grain is neutral in the finished picture.
The objection that stood here — that blending the classical output back in re-adds the noise —
holds for a plain mix, which also brings back the classical path's colour speckle and false
colour. A luminance-only blend returns film-like grain and nothing else, and it needs no
inference: one elementwise GPU pass (`dr_gpu::GrainBlend`) per slider value, producing a new
source the adjust pass draws. Comparing the two on real 6D frames, the user chose this one.
The σ-map Amount (§3.3) still works — `NoiseModel::scaled` — and stays available for a later
"strength" control; its cost is a re-run of the network.
### 7.3 Runtime
Through `dr-inference-engine`, as the other models run ([inference.md](inference.md)): TensorRT or
CUDA fp16 on the laptop, MIGraphX on the desktop, ORT CPU everywhere, QNN on the tablet. Work is
Through `dr-inference-engine`, as the other models run ([inference.md](inference.md)), as
`Role::Denoiser`: TensorRT or CUDA fp16 on the laptop, MIGraphX on the desktop, ORT CPU everywhere.
**Not the Hexagon** — see §11 — so the tablet runs it on its CPU. Work is
scheduled in the `Background` class so a slider never waits on it (architecture §5.3).
## 8. Speed and the tablet
@@ -320,6 +331,17 @@ If neither holds S's quality within 0.5 dB of fp32 on the real pairs, **v1 is de
tablet shows the classical path. The sidecar still records the intent, so a desktop can render the
learned result for a photograph edited on the tablet.
**Measured 2026-10-04 (inference.md §1.5): the second way holds, without the first.** The shipped
network, with its Bayer packing re-spelled as `SpaceToDepth` so QNN can hold it (the 6-D reshape
it replaces is exact but past the HTP's rank limit), at A16W16 — 16-bit activations and weights —
scores within 0.00 dB of f32 at ISO 400–25600 on the tablet's own HTP, and within 0.09 dB with the
6D's noise model scaled ×0.5, ×2 and ×4 to stand in for other sensors. A16W8 holds the 6D (worst
−0.19 dB at ISO 25600) but not ×4 noise at 25600 (−0.52 dB), so A16W16 is what ships. int8 loses
4.7–9.2 dB and fp16 is refused outright. A 1408 tile takes 95 ms on the Hexagon against 1510 ms on
the tablet's CPU: about 2.3 s for a 20 MP frame. Calibration ranges come from 96 training-day
tiles across every ISO, a third of them with that scaled noise; coverage of other bodies is that
synthetic bracket, not their raws.
## 9. X-Trans
The requirements tie this stage to FR-RAW-5, and the library has no Fuji raws. What we can do
@@ -356,3 +378,252 @@ MIT architecture, so this model adds no third-party licence to D13.
4. Whether a Lightroom or DxO comparison is available for §6.2.
5. A borrowed X-Trans body, or X-Trans experimental in v1.
## 11. What shipped in 0.21.0, and what was measured
**Data.** 500 distinct ISO 50–100 6D frames from the library, over 121 shooting days (bursts and
near-duplicate perceptual hashes dropped; 55 frames from held-out days for validation). Read through
`dr-gpu`'s `mosaic_dump` example — `dr-decode` and the app's own hot-pixel pass — so the network's
input is the mosaic the classical demosaic reads. Truth by 2×2 binning with a Catmull-Rom quarter-pixel
shift of red and blue (§4.2). Training lives in `darkroom-denoise`, beside `darkroom-infill`.
**Noise model (§5), from the library instead of a capture.** Shot gain and read variance per ISO from
Adobe's `NoiseProfile` in the converted DNGs; read noise checked against each frame's masked border
(agreement within 2–3 % from ISO 125 to 25600); read-noise *shape* taken from the border as quantiles
on a tail-dense grid (excess kurtosis up to ~10 at high ISO), with only the photosites the app's
hot-pixel rule would remove left out; row noise from the border's row means; **column noise** from
the masked rows above the image — about a third of its variance is this sensor's fixed pattern.
Third stops are their own rows: ISO 160, 320 and 640 are quieter than their neighbours, as §5.1
expected. Training without the column noise left the 6D's vertical stripes in (0.90 DN of 1.01);
with it, 0.17 DN.
**Model.** Not NAFNet: its channel attention averages over the whole input, which breaks exact
tiling. A U-Net of 3×3 convolutions, ReLU, strided and transposed convolutions and additive skips —
3.2 M parameters, 48 GMAC per raw megapixel, receptive field 185 photosites (counted from the
layers; a perturbation probe under-read it as 157 because a switched-off ReLU hides a path).
Tiles of 1408 keep their central 1024 behind a 192 halo, exactly.
**Results.** PSNR after the display transform, held-out days, step 60 000:
| ISO | Network | Bilinear | Bilinear on a clean mosaic |
|---|---|---|---|
| 400 | 41.6 | 36.8 | 40.0 |
| 1600 | 40.8 | 33.8 | 40.0 |
| 6400 | 39.5 | 29.3 | 40.0 |
| 25600 | 37.8 | 24.6 | 40.0 |
Unbiased in linear light on real frames (shadow level within 1 % of a heavily averaged bilinear).
Checked against the app's own render for channel and axis order (`tools/check_against_app.py`).
**Precision (§8).** fp16: 0.00 dB at every ISO. int8 QDQ, calibrated on training tiles: −6 to −9 dB
— the shadow steps §8 feared losing are lost. So the Hexagon refuses the role and the tablet runs f32
on its CPU; the residual head of §8 is the route back.
**Noise for any Bayer body (§3.3).** Table, then `NoiseProfile`, then the frame itself: read, row and
column noise from its masked border, the shot gain alone estimated from the quietest flat patches.
On 130 6D frames the estimate is within ±10 % of the table from ISO 1000 up and scattered below. The
network loses under 0.3 dB for σ off by 15–20 % and twice as much for under- as for over-estimating;
the estimate leans high. Every Bayer body is offered the switch; develop says which source was used.
**Speed, a whole 6D frame (20 MP).** TensorRT fp16 3.1 s, ONNX Runtime CPU 14.4 s, on the laptop —
measured while the GPU sat power-capped at an 810 MHz memory clock; uncapped is expected to be
about four times faster. The Rust path reproduces the training repository's output to 2.5e-4 at
worst; TensorRT fp16 is 75 dB from f32.
**Not yet:** ~~the result is not cached across sessions (§7.1) — reopening recomputes;~~ done after
0.21.0, §12; the tripod real
pairs of §6.1; X-Trans (§9); the hand-written WGSL path, for which `export.py` already writes the
weights blob and a manifest a shader can follow.
## 12. On by default, with a strength, and cached (after 0.21.0)
The photographer asked for the learned demosaic to be how a raw is developed, not an option found
under Detail. So:
- **On by default, at full strength, on every device.** The switch is `switch_on`, so an untouched
photograph writes nothing and is developed from the network everywhere; turning it off is the
edit. Which hardware runs it is the inference engine's choice (inference.md), not this setting's:
the default does not depend on what a device is believed to manage.
- **Strength replaces Keep grain.** 0–100, default 100, and grain = 100 − strength, so it is the
same luminance-only blend of §7.2 and moving it is one GPU pass, never a re-run. An edit saved by
0.21.0 stored `grain`; it is still read, as its inverse, and never written.
- **First in the panel**, above the lens corrections: it decides what every control below is
applied to. Its attribute is still Detail, so it also stays where the Detail tab shows it.
- **Cached on disk** (§7.1): the network's output for a file, as half floats (about 120 MB for
20 MP — no compressor to link on Android), keyed on a SHA-256 of the file's bytes and the model
file's name and size, oldest first past a 5 GB budget, beside the inference engine's cache under
the data root. The strength is applied afterwards and is not in the key. A reopened photograph,
and an export of one already developed, read it back instead of recomputing.
What it costs: every raw opened runs the network once, with the classical demosaic shown until the
result lands, and a first export of an unopened raw runs it too. Every raw renders differently from
0.21.0 unless switched off.
## 13. Three networks and a method (after 0.22.0)
The photographer asked for a choice between quality and time. `Method` replaces the Apply switch:
`Bilinear`, `Fast`, `Medium`, `Best`, by index in that order, default `Best`. An untouched raw
writes nothing and develops through `Best`. `apply` is still read and never written: 0 is
`Bilinear`, 1 keeps a network already chosen or is the default. A number past the list, from a newer
build, reads as the default. A build before this one ignores `method` and develops through its own
network, which is the most an older peer can do.
**The networks** (darkroom-denoise, every one trained on the same data and noise as §11, plus 1,201
further frames cropped from the library and 6,000 drawn scenes — polygons, lines of one to four
photosites, text, gratings — rendered at 4× through a random affine and smooth displacement, so
edges fall off the photosite grid):
| Method | File | Network | Parameters | GMAC / MP | Halo |
|---|---|---|---|---|---|
| Best | `mosaic-best-1408.onnx` | two U-Nets of §11's shape (a flat expert from `m2`, an edge expert from the ×100 edge-weighted run) and a 128 k-parameter gate that blends them per photosite | 6.4 M | 110 | 256 |
| Medium | `mosaic-medium-1408.onnx` | §11's U-Net, distilled from Best (75 % its output, 25 % the truth) | 3.2 M | 48 | 192 |
| Fast | `mosaic-fast-1408.onnx` | widths 16-32-64-128, blocks 1-1-1-2, distilled the same way | 0.93 M | 11 | 192 |
The gate learned on its own to trust the edge expert at 0.77–0.88 on edges and not at all on flat
areas. The mixture's receptive field is the experts' plus the gate's, so it keeps the centre of a
1408 tile past a 256 halo, where the single networks keep 1024 past 192. `dr_denoise::Shipped`
carries each file's halo, and `TileNet::halo` hands it to the tiler.
**Quality.** PSNR after the display transform on 1,842 held-out crops, and the width of a hard
edge on the drawn chart at ISO 6400 (truth 0.80 photosites; lower is sharper):
| | ISO 400 | 1600 | 6400 | 25600 | Edge width |
|---|---|---|---|---|---|
| §11's network | 40.61 | 39.77 | 38.43 | 36.67 | 1.77 |
| Best | 40.69 | 39.84 | 38.47 | 36.71 | 0.82 |
| Medium | 40.59 | 39.75 | 38.40 | 36.65 | 1.30 |
| Fast | 39.90 | 39.06 | 37.54 | 35.33 | 1.84 |
| Bilinear | 36.15 | 33.26 | 28.72 | 23.83 | 2.15 |
On photographs the three are close; on hard edges Best is half as wide as §11's network and
Medium most of the way there. Fast costs a dB at high ISO and edges as soft as §11's.
**Speed**, a whole 20 MP 6D frame, the network alone, TensorRT fp16 on the laptop's RTX 3050
(uncapped: memory at 5 GHz), engine already built: Best 2.48 s, Medium 0.79 s, Fast 0.57 s. Decode
and the hot-pixel pass add about 0.5 s. The first build of each TensorRT engine takes 80 s (Fast) to
190 s (Best), in the background at first launch, cached after.
**Before the network, two passes changed since §11.**
- *A noise-aware repair* (`dr_denoise::repair`) after the app's hot-pixel pass: a photosite more
than 8σ beyond every same-colour neighbour *and* every adjacent photosite, and more than twice
each adjacent one, is clamped to the brightest of its same-colour neighbours; a dead one, to the
darkest. The ratio test is what spares a point of light, whose neighbours are lit too. The networks
were trained behind the same pass (the Python and Rust agree: 935 repairs on an ISO 25600 frame).
- *The tiler feeds the network without waiting*: tiles are gathered on every core by a producer
thread one tile ahead, and the output is written back in parallel from the runtime's own buffer.
0.14 s of tiler for a frame, which is what keeps Fast under a second.
**The Hexagon.** Each network has an `.a16w16.onnx` sibling made by `tools/quantise-models.sh
--ranges`, the ranges from darkroom-3e's gate (96 training tiles, a third at noise ×2 and ×4). On
the 6D gate A16W16 loses 0.00 dB for all three; with the noise scaled ×0.5–×4 at most 0.11 dB.
A16W8 holds the gate (≤ 0.27 dB) but loses 0.63 dB on Medium at ×4, so A16W16 stays the form.
**Cache.** Each network keys its own results (§7.1 keys on the model's file name), and the file is
hashed once at open, so changing the method never re-reads it. Choosing `Bilinear` keeps the
network's result in memory for the way back; changing to another network drops it, and coming back
reads the cache.
**Packaging.** All six files in the APK (`BUNDLED`, 23 entries, +44.6 MB, ~41 MB compressed); the
three f32 networks in the Arch package and the Windows installer, which stage `models/denoise` by
directory.
## 14. A whole frame, not 1408² tiles (after 0.23.0)
A fixed 1408² tile is exact only past its halo, and Best's halo is 256: of every 1408² it computes
it keeps 896², 2.47 photosites of work for each one kept (Medium and Fast keep 1024², 1.89×). On a
GPU the network can instead run over the whole frame and its reflected border in one call, which is
exact by the same argument (§3.4) and wastes only the border.
**The networks** are re-exported with any height and width (`mosaic-{best,medium,fast}.onnx` beside
the 1408 files; darkroom-denoise `tools/export_whole.py`), from the checkpoints the shipped files
came from. The tool refuses unless each matches its 1408 file at 1408² (max |Δ| = 0 for all three),
matches torch at 592 × 848, and equals tiled inference over the reflected frame in f64 (≤ 7e-16).
The APK leaves them out: the Hexagon takes fixed shapes.
**Where they run.** `Role::WholeDenoiser` is served by TensorRT and the CUDA provider only, the rungs
where a new input size costs nothing at run time; MIGraphX, OpenVINO and CoreML compile per shape,
the Hexagon takes fixed shapes, and the CPU would hold gigabytes of f32 activations. Everywhere else
`whole_frame_limit()` is `None` and the 1408² tiles run as before. TensorRT gets an optimisation
profile up to `WHOLE_FRAME_MAX` (4608 × 3328) — without one a dynamic input compiles a new engine per
size at run time — through the runtime's V2 options, since `ort`'s builder has none, and keeps the
engine in a directory per model and profile (ONNX Runtime's cache key leaves the shape out).
**The limit is the card's memory.** TensorRT plans its memory for the profile's largest shape. A
profile up to a whole 6D frame with Best's border (4608 × 6656) asked for 4.9–5.9 GB and would not
build on the 6 GB RTX 3050. At 15 MP the tiler (`tile::plan`) cuts the frame into the fewest equal
tiles under the limit: a 6D frame is two of 4160 × 3248, 27 MP of work for 20 MP kept, against 49 MP
in 1408² tiles. If a plan's first call fails, as a GPU out of memory does, its kept centre is halved
and the frame planned again.
**Measured** 2026-10-06 on `_MG_8862` (6D, ISO 8000, 20 MP), RTX 3050 Laptop, TensorRT fp16, P3 /
5001 MHz, another session's paused training holding 1.3 GB:
| Best | Network time | Against the tiles |
|---|---|---|
| 1408² tiles | 2.60 s | — |
| Whole frame, two 4160 × 3248 tiles | **1.37 s** | max \|Δ\| 0.0029, mean 1.1e-5 — fp16's own spread (GPU tiles against CPU tiles: 0.0025) |
The first build of the whole-frame engine took 28 minutes, in the background at first launch, with
the 1408² tiles serving meanwhile — against about 3 minutes for the fixed one; the profile's range
is what it tunes across. A cached engine loads in about a second.
In PyTorch fp16 the same network over the whole 20 MP frame in one call took 3.5× less than in
tiles, so a card that holds a whole frame gains more than the 6 GB one does; `WHOLE_FRAME_MAX` is a
constant sized for 6 GB until the limit follows the card's memory.
## 15. Best becomes one network (0.24)
The photographer's goal for 0.24 was Best's quality in under a second on the laptop. Whole frames
(§14) took the mixture from 2.60 s to 1.37 s and no further on a 6 GB card, so the other half was
a single network that holds the mixture's quality at a third of its work. Methods are now
`Bilinear`, `Fast` and `Best`; Medium and the mixture are retired.
**The network** is `fb-combo` (darkroom-denoise, 2026-10-07): §11's shape (32-64-128-192, blocks
1-1-2-2, 3.2 M parameters, 48 GMAC/MP, halo 192), 20 000 steps from `fb-edges2` ← `student-m`,
taught by the mixture at a half share, with 10 % drawn scenes and 25 % crops from the edge-rich
cells of the training frames (branch `edge-sampling`). Scored on real photographs — the chart
overstated the mixture's lead (a chart-sharp network was softer than Medium on real edges) — on the
validation crops in the top quarter for sharp detail:
| | Edge PSNR, ISO 1600 / 6400 / 25600 | Sharpness kept | Smooth areas | Held-out PSNR, ISO 400 / 1600 / 6400 / 25600 | Chart edge |
|---|---|---|---|---|---|
| Mixture (Best to 0.23) | 30.71 / 29.93 / 28.55 | 0.899 / 0.868 / 0.782 | 42.61 / 41.71 / 40.12 | 40.69 / 39.84 / 38.47 / 36.71 | 0.82 |
| `fb-combo` (Best from 0.24) | 30.67 / 29.87 / 28.49 | 0.902 / 0.874 / 0.792 | 42.54 / 41.57 / 39.85 | 40.64 / 39.78 / 38.38 / 36.54 | 0.89 |
| Medium (to 0.23) | 30.43 / 29.68 / 28.38 | 0.896 / 0.862 / 0.776 | 42.59 / 41.69 / 40.08 | 40.59 / 39.75 / 38.40 / 36.65 | 1.30 |
Edges within 0.04–0.06 dB and more sharpness kept at every ISO; the known shortfall is smooth areas
at ISO 25600, 0.27 dB. The photographer took it as it stood at 20 000 of a planned 30 000 steps.
Others tried on the way, each short of the mixture on real photographs: `fb-sharp` (drawn scenes,
chart-sharp but Medium's real edges), `fb-edges` (half edge-rich crops: edges close, ISO 25600
flats −0.24 dB), `fb-edges2` (a quarter: 0.03–0.14 dB short everywhere, chart 1.33–1.47), and a
from-scratch 24-48-96-128 between Fast and Medium.
**Files.** `mosaic-hq-1408.onnx`, `mosaic-hq.onnx` (any size) and `mosaic-hq-1408.a16w16.onnx` for
the Hexagon. A new name, not Medium's or Best's: the result cache keys a model by name and size,
and this one is byte for byte Medium's size. The tablet form lost 0.00 dB in simulated QDQ at every
ISO and at most 0.09 dB across the ×0.5–×4 noise bracket (A16W8 0.08 / 0.26 dB; int8 −10.6 dB);
not yet confirmed on the tablet itself.
**Saved edits** keep their numbers: 2, which was Medium, is now Best; 3, which was Best, is past the
end and reads as the default, Best. Both land on the new network with no migration.
**Measured** 2026-10-07, `_MG_8862`, RTX 3050 Laptop, TensorRT fp16, P3 / 5001 MHz, nothing else on
the card:
| Best | Network time | Peak GPU memory |
|---|---|---|
| mixture, 1408² tiles (0.23) | 2.60 s | — |
| mixture, whole frame (§14) | 1.37 s | — |
| `fb-combo`, 1408² tiles | 0.95 s | 0.55 GB |
| `fb-combo`, whole frame (two 4160 × 3248) | **0.51–0.54 s** | 1.75 GB |
Decode and the hot-pixel pass add 0.4–0.5 s, so a photograph is about a second end to end. Whole
frame against tiles: max |Δ| 0.0029, 90 dB apart — fp16's spread. The whole-frame engine's first
build took 12 minutes (the mixture's 28); from the cache it loads in about a second, so the session
keeps the engine's ordinary 30 s idle decay rather than unloading after each photograph: at 1.75 GB
it fits beside the develop view on a 6 GB card, and an unload would cost the next photograph a
second.
The manual's close-up for Best is still the mixture's render, which this network matches to within
the table above; it is re-recorded with the next pass of `tools/manual/record.sh`.
+152 -10
View File
@@ -127,7 +127,8 @@ Three things the table settles.
five of six models. A whole-library face index on the tablet goes from ~100 ms + 39 ms per face
to ~1.4 ms + 12 ms per face, and the "Thorough" detector — 3× the cost of "Fast" today — becomes
free. Its price is that the models must be **quantised to int8**, which is an accuracy question
§5 has to answer before it is believed.
§5 has to answer before it is believed. (§1.5 answered it: int8 lost faces, and every model but
XFeat ships with 16-bit activations, at about three times these timings.)
- **The embedder does not gain from either accelerator.** 112×112 input, per-op overhead
dominates; it is 9 ms on the tablet's CPU and 12 ms on its NPU. It stays float, which §7 turns
from a performance footnote into a correctness rule.
@@ -137,8 +138,101 @@ Three things the table settles.
- **On AMD, MIGraphX fp16 is 4–17× the CPU provider** on the detectors and 60× on the
inpainter, with the same first-run compile cost as TensorRT and no rung between it and the CPU.
### 1.5 The Hexagon at every bit width · 2026-10-04
§1.1's Hexagon column is int8 calibrated on noise: timing only. This is the follow-up — every
model, every bit width the HTP offers, calibrated on real photographs and **scored on the tablet
itself** (ORT 1.29 + QNN 2.42, `htp_arch` 73), against the f32 model on the same inputs. The
"Form shipped" column is the files in `models/`, re-scored on the tablet after
`tools/quantise-models.sh` wrote them. The
calibration and scoring photographs are 800 from the public COCO val2017 set (CC-BY); the face
models' numbers are over the faces in them of at least 32 px. The tools are `tools/quantise-models.sh`
and the scratch harness described with it.
**What the HTP accepts.** fp16: nothing — every fp16 operator fails validation (3110), on QNN
2.42 and 2.50, with `htp_arch` and every `soc_model` tried; the fp16 rung stays off the table until
someone has Qualcomm's own SDK to say why. 4-bit weights (A8W4, A16W4): load, and wreck accuracy
(SCRFD finds 25–35% of f32's faces). What is left: **A8W8 (int8), A16W8 and A16W16**, all running
the whole graph. 16-bit activations cost about 3× int8's time, A16W16 about 4×.
| Model | ORT CPU f32 | Form shipped | Hexagon | On the tablet, against f32 | int8 for comparison |
|---|---|---|---|---|---|
| scrfd_500m / 2.5g / 10g | 17 / 56 / 198 ms | **A16W8** | 4.2 / 5.1 / 9.0 ms | 100% of faces found in every size band; keypoints 0.3–0.6% of the box | 94–95% of faces at 40–80 px |
| 2d106det (landmarks) | 2.8 ms | **A16W8** | 0.5 ms | 0.25 px in the 192 crop (eye points 0.20) | 1.5 px, and 29 partitions at 7.3 ms |
| yolo26n-seg | 90 ms | **A16W16**, tail in float | 12.9 ms | 98.2% of objects, mask IoU 0.994 | 74% (simulated) |
| yolo26s-sem-ade20k | 151 ms | **A16W16**, attention in float | 15 ms | 98.9% of cells agree on the class, TV 0.009 | 67% |
| migan-512 | 488 ms | **A16W16** | 87 ms | 41 dB from f32 in the fill (worst 1%: 30 dB) | 16 dB (simulated) |
| xfeat-1024 / 768 | 58 ms | **int8**, rewritten graph | 6.5 ms | panorama alignment 0.45 px from f32's — f32's own refit on 90% of its matches is 0.41 | — |
| mosaic-1408 (denoiser) | 1510 ms a tile | **A16W16**, rewritten graph | 95 ms a tile | 0.00 dB at every ISO; ≤ 0.09 dB with the noise scaled ×0.5–×4 | −4.7 to −9.2 dB |
| arcface_mbf (embedder) | 8.5 ms | f32, CPU | — | A16W16: cosine 0.9995, p1 0.9967 — misses §7's 0.999 gate | — |
| ocec, sgc (eyes) | 1, 1.7 ms | f32, CPU | — | sgc flips 1.45% of views even at A16W16; not worth a millisecond | — |
Four things the table needed that the f32 graphs did not have, all in `tools/htp_graph.py` and all
checked exact against the f32 graph before they are used:
- **Rank ≤ 5.** QNN's tensors stop at rank 5, and the denoiser packs the mosaic through a 6-D
reshape (6007 at compose). For one channel that reshape is `SpaceToDepth(2)`. XFeat's 8×8 unfold
is 224 Slices and 6-D Concats; it is `SpaceToDepth(8)` (736 nodes to 60).
- **No bilinear Resize at XFeat's sizes** (3110). A half-pixel bilinear resize between fixed sizes
is two constant matrices, so it is two MatMuls.
- **One scale per tensor.** The segmenter's output rows carry boxes in pixels beside scores in
0..1; quantised as one tensor the scores vanish. Everything from the Concat that builds the rows
stays float, on the CPU, where the top-300 selection costs nothing.
- **Float where the HTP's 16-bit arithmetic drifts.** The scene model's one attention block
(two MatMuls and a Softmax over 400 tokens) moved its agreement from 98.7% to 96.7%; it stays
float.
**ORT's CPU simulation of a QDQ graph is not the tablet.** It matched to the hundredth of a dB for
the denoiser and to rounding for the detectors, landmarks and XFeat, and it overstated MI-GAN by
27 dB and the scene model by three points. Every number above is the device's; a new form is not
measured until it has run there.
**XFeat's int8 loses keypoints and not the panorama.** 83% of f32's keypoints come back within
1.5 px; but over the twelve-frame `fixtures/pano/2025-08-05` sweep, the homographies fitted from
int8's matches land 0.45 px from f32's in the overlaps — the spread f32 shows against itself
(0.41).
---
### 1.6 Intel Iris Xe, and the generic rung · 2026-10-04
The RTX 3050 laptop's other GPU: Raptor Lake-P's Iris Xe (96 EU), Intel's `onnxruntime-openvino`
1.24.1 (OpenVINO 2025.4.1) and Microsoft's `onnxruntime-webgpu` 1.27.0, both PyPI wheels, through
`ep_probe`. Three warm-ups, the median of 15 runs. Another build shared the CPU during the run, so
the CPU columns are a little pessimistic; the GPU columns are not.
| Model | ORT CPU f32 | OpenVINO CPU | OpenVINO GPU f32 | **OpenVINO GPU fp16** | WebGPU (Iris Xe) |
|---|---|---|---|---|---|
| scrfd_500m (Fast) | 9.5 | 10.8 | 7.3 | **5.8** | 24.2 |
| scrfd_2.5g (Balanced) | 18.5 | 16.2 | 15.5 | **11.1** | 40.8 |
| scrfd_10g (Thorough) | 58.5 | 72.4 | 38.7 | **23.0** | 84.5 |
| arcface_mbf (per face) | 9.5 | 11.7 | **3.0** | 2.4 | 56.6 |
| 2d106det (landmarks) | 10.8 | 2.0 | 2.4 | **1.9** | 42.7 |
| yolo26s-sem-ade20k | 57.0 | 47.4 | 26.0 | **16.7** | 53.0 |
| xfeat-1024 | 23.5 | 18.5 | 19.9 | **17.5** | 34.2 |
| migan-512 (per tile) | 330 | ✗ ¹ | 89.7 | **57.2** | 275 |
| mosaic-fast-1408 (per tile) | 159 | 107 | 68.1 | **40.0** | 188 ² |
| mosaic-best-1408 (per tile) | 1109 | 1670 | 947 | **604** | 1034 ² |
¹ OpenVINO's CPU plugin refuses the graph at initialisation. Not shipped (§3.2), so moot.
² A later run, after `ep_probe` learned to feed the denoiser's two inputs, under heavier load: the
CPU provider took 256 and 1034 ms in that run, so WebGPU beat it by a quarter on mosaic-fast and tied
on mosaic-best — the only rows where it is not well behind.
- **OpenVINO on the Iris Xe beats ONNX Runtime's CPU provider on every model**, 1.3× on XFeat to
5.8× on MI-GAN, with a 1–3 s compile per graph and 0.1–0.4 s from its cache. It is the Intel rung.
fp16 is worth 1.3–1.7× over f32 here, against 1.1–1.35× on MIGraphX.
- **Its "GPU" is OpenCL's first GPU, not Intel's.** Before `intel-compute-runtime` was installed
the only OpenCL driver was NVIDIA's, and `device_type=GPU` ran on the RTX 3050 — slower than the
CPU, which is the probe's to catch. Read the process's maps for `libigdrcl` before believing a
number is the iGPU's.
- **WebGPU is slower than the CPU on the Iris Xe** on everything but MI-GAN, as it was on the
Adreno (§1.1), and on the RTX 3050 through Vulkan too. It is on the ladder anyway, as the generic
rung (§2): for GPUs no vendor rung covers — an AMD card on Windows or without ROCm, a Mali — where
it is unmeasured, and the probe's clock decides.
- **OpenVINO's CPU plugin is not a better floor.** It wins on some graphs and loses on scrfd_10g,
the embedder and mosaic-best, and refuses MI-GAN.
## 2. The shape of the answer
A **ladder per platform**, walked at start-up, with the first rung that builds a real session
@@ -146,11 +240,12 @@ winning:
| Platform | 1st | 2nd | 3rd | Floor |
|---|---|---|---|---|
| Android, Qualcomm with a Hexagon the shipped QNN skel covers (V68–V81) | QNN HTP, int8 model | ORT CPU, f32 model | — | tract |
| Android, any other SoC | ORT CPU, f32 | — | — | tract |
| Android, Qualcomm with a Hexagon the shipped QNN skel covers (V68–V81) | QNN HTP, each model's quantised form (§1.5) | ORT CPU, f32 model | — | tract |
| Android, any other SoC ⁶ | WebGPU (Vulkan), f32 | ORT CPU, f32 | — | tract |
| Linux / Windows, NVIDIA GPU | TensorRT, f32 model, fp16 engine | CUDA provider, f32 | ORT CPU, f32 | tract |
| Linux, AMD GPU with ROCm | MIGraphX, f32 model, fp16 program | ORT CPU, f32 | — | tract |
| Linux / Windows, no GPU stack | ORT CPU, f32 | — | — | tract |
| Linux / Windows, Intel GPU | OpenVINO, f32 model, fp16 program (§1.6) | ORT CPU, f32 | — | tract |
| Linux / Windows, any other GPU ⁶ | WebGPU (Vulkan / D3D12), f32 | ORT CPU, f32 | — | tract |
| macOS ⁵ | CoreML, f32 model, ML Program | ORT CPU, f32 | — | tract |
⁵ **Unmeasured**, and the one exception to the rule below: nobody here has a Mac. The rung is on
@@ -160,11 +255,17 @@ down is refused on the third launch (§4, `attempt`). The embedder stays on the
macOS log that shows a probe line is this row's measurement; [macos.md](macos.md) says what to
ask for.
⁶ **The generic rung, unmeasured where it is meant to help.** WebGPU lost to the CPU on every GPU
it has been timed on — the Adreno, the Iris Xe, the RTX 3050 (§1.1, §1.6) — none of which it serves
here, since each has its own rung. It is on the ladder for the GPUs that have none, on the same
terms as CoreML: a WebGPU that is slower than the CPU is rejected by §4's clock, one that errors is
recorded as failed. Its first measurement on an AMD card without ROCm, or a Mali, is this row's.
Deliberately **not** on any ladder, with the measurement that excluded each: NNAPI (no driver),
XNNPACK (slower than CPU, aborts on SCRFD), WebGPU (slower than CPU), the Adreno through QNN (works,
but never where the Hexagon does not also), CUDA int8 (slower than CUDA f32), the ROCm provider
(gone: §1.3). A rung is added to this table by a measurement on this page, not by a provider
existing.
XNNPACK (slower than CPU, aborts on SCRFD), the Adreno through QNN (works, but never where the
Hexagon does not also), CUDA int8 (slower than CUDA f32), the ROCm provider (gone: §1.3),
OpenVINO's CPU plugin as a floor (§1.6). A rung is added to this table by a measurement on this
page, not by a provider existing — the two footnoted rows are the exceptions, and say so.
The AMD ladder has no middle rung. TensorRT falls back to the CUDA provider while its engines
compile; MIGraphX has no such twin, so its fallback is the CPU provider, and the minute or two of
@@ -240,6 +341,40 @@ it.
Both positions are D13 territory and are recorded there (§12).
Two more runtimes ship in every desktop package since 0.23 (§3.2), and their parts are all
redistributable:
| Component | Licence | Shipped |
|---|---|---|
| Intel's `onnxruntime-openvino` build, with OpenVINO 2025.4.1 and oneTBB | MIT; Apache-2.0; Apache-2.0 | Linux and Windows packages, texts beside the libraries |
| Microsoft's WebGPU build (Dawn inside); on Windows the DirectX shader compiler | MIT; LLVM / MIT | Linux and Windows packages |
| Microsoft's stock `onnxruntime-android` (WebGPU) | MIT | The APK, as `libonnxruntime_generic.so` |
### 3.2 Several runtimes, one per process
A runtime carries one vendor's providers: Intel's build has OpenVINO, the `onnxruntime-gpu` wheel
CUDA and TensorRT, a ROCm build MIGraphX, Microsoft's WebGPU build the generic rung, the APK's QNN
build the Hexagon. No prebuilt carries two vendors, and `set_api` takes one table per process.
So a device that may hold several — the package's OpenVINO and WebGPU builds, a CUDA build the user
fetched, the distribution's ROCm build — has to choose which to load *before* the probe, and
cannot choose by trying.
`api::install` opens every runtime on the search list, asks each for `GetAvailableProviders`, and
loads the one that scores highest against the GPUs `hardware::detect` reads from files: a vendor
rung on its own vendor's GPU (NVIDIA driver, `/dev/kfd`, a Qualcomm SoC, macOS) above OpenVINO on
an Intel GPU (PCI vendor `0x8086`; on Windows Intel's DCH driver package) above WebGPU above a
CPU-only build. Equal scores keep the search order, a perfect fit ends the search — the APK's QNN
build is listed first, so on a Qualcomm device the generic build is never opened — and
`DARKROOM_ORT_DIR` wins outright. The losers stay mapped: unloading a C++ runtime whose static
constructors ran is a crash at exit waiting to happen.
The desktop packages install the two bundled builds under `runtimes/openvino` and
`runtimes/webgpu` beside each place a package installs to, from
`tools/fetch-bundled-runtimes.sh` (PyPI wheels pinned by SHA-256, pruned to the native libraries:
81 + 31 MB on Linux, 67 + 42 MB on Windows). On Windows the chosen runtime's directory is put on
`PATH`, because Intel's build leaves OpenVINO's DLLs for the loader to find there. The Flatpak has
no Intel OpenCL driver in its sandbox, so an Intel machine there settles on the CPU.
---
## 4. Selection — the probe, its cache, and what it may not do
@@ -298,10 +433,10 @@ of which form they load:
| Form | Who produces it | When | Needed by |
|---|---|---|---|
| f32 ONNX, shape-fixed, **opset ≥ 13** | `tools/fix-face-model-shapes.sh`, `tools/export-seg-model.sh` | Release time, once | Every rung except Hexagon |
| int8 QDQ ONNX, per-channel, uint8 activations | `tools/quantise-models.sh` (new) | Release time, once, **calibrated on real photographs** | Hexagon |
| QDQ ONNX, per-channel — int8, A16W8 or A16W16 per model (§1.5) | `tools/quantise-models.sh` | Release time, once, **calibrated on real photographs**, scored on the tablet | Hexagon |
| TensorRT engine (`.engine`, per GPU architecture and TensorRT version) | The app, from the f32 file | First run on that device, in the background | TensorRT rung |
| MIGraphX program (`.mxr`, per GPU architecture, MIGraphX version and precision) | The app, from the f32 file | First run on that device, in the background | MIGraphX rung |
| QNN context binary | The app, from the int8 file | First run on that device, in the background | Hexagon rung |
| QNN context binary | The app, from the quantised file | First run on that device, in the background | Hexagon rung |
Two rules.
@@ -380,6 +515,11 @@ already the rule for the detector and because §5 is the gate on whether the int
enough to be *offered* at all. f32 on tract, ORT CPU, CUDA and TensorRT-f32 are one identity: the
same graph, the same arithmetic, differences at the last bit.
After §1.5 the Hexagon runs the detectors in **A16W8**, and that is a third spelling:
`scrfd_500m_a16+w600k_mbf` and its two siblings. Same rule, same reconciliation; a tablet that
indexed under `_i8` keeps those rows, and `FaceDetector::model_ids` answers "has this detector been
over this image" for all three forms.
**The embedder** is where comparability across devices is the whole point, and it is the one
model that no accelerator helps (§1.4). So: **the embedder runs in f32 on every rung.** On TensorRT
that means the embedder's engine is built without fp16 while the detector's is built with it; on
@@ -543,6 +683,8 @@ device are comparable. *Acceptance:* M3.
**D13 — updated.** The runtime half is reopened to the extent of §3: the Rust build stays C-free
under `alternative-backend`; packages may install a dynamically loaded ONNX Runtime and, per §3.1,
the Qualcomm QNN runtime; the NVIDIA libraries are not bundled. The licensing half is unchanged.
Since 0.23 every desktop package bundles two runtimes — Intel's OpenVINO build and the WebGPU
build — and the APK a second, generic one; the engine loads the one that fits the GPU (§3.2).
---
+7
View File
@@ -53,6 +53,13 @@ left outstanding, and its DCP half stays deferred as before. §4's FR-DSP-2 and
record the one case that now tiles, a linear DNG larger than one texture, and §11 the merge's frame
choice, which changes how FR-MRG-5 is met rather than whether.
**And for 0.20.0.** FR-DEV-3e's DCP half is built (D20, [camera-profiles.md](camera-profiles.md)):
the HueSatMap and LookTable from a DNG's embedded profile or a matched `.dcp`. Two pieces stay
open, both named in that design's §6: the profiles directory does not sync, so a CR2 can render
with a copied profile on one device and without it on another; and `ProfileToneCurve` and
`BaselineExposure` are read and not applied. The measurement in its §1 says the second is where the
remaining gap to Lightroom's colour lies.
---
## 1. Plugins — post-v1 since 2026-09-19
+40 -7
View File
@@ -367,7 +367,7 @@ Built, on branch `merge/panorama`, in the order §10 gave:
| The camera-space tap | `OutputMode::CameraLinear`, `AdjustPass::render_camera_linear` | Done, `rgba32float`, tiles by view rect |
| Linear DNG writer, streamed | `dr_export::write_linear_dng` | Done; rawler reads it back |
| A three-sample `RawImage` re-entering the pipeline | `dr-decode`, `DemosaicedImage::from_linear_rgb16` | Done |
| Warp, accumulate, resolve, chunk by chunk | `dr_gpu::MergePass`, `merge.wgsl` | Done; feathered blend, scalar gain |
| Warp, accumulate, resolve, chunk by chunk | `dr_gpu::MergePass`, `merge.wgsl` | Done; seams (§11.1) over a feather, scalar gain |
| The job: load, proxies, align, gains, confirm, merge, provenance | `dr_ui::merge` | Done; `examples/merge.rs` drives it headless |
| The page: table, preview, projection, Merge/Stop/Back; the grid's button | `merge.slint`, `merge_ui.rs` | Done; `DARKROOM_START_MERGE=a.CR2,b.CR2` lands on it |
| Placement beside the sources through the outbox, rescan | `merge_ui.rs` | Done, untested against a server |
@@ -385,9 +385,11 @@ half.
the file carries the black border. The largest inscribed rectangle over
the coverage, then the DNG's `DefaultCropOrigin`/`DefaultCropSize`, so
nothing is thrown away and the develop view opens on the picture.
2. **Seams and the pyramid** (§10 step 5). The feather hides exposure and
small misalignment; parallax on the near slope will show as a soft
double edge at 1:1.
2. **The pyramid** (§10 step 5). Seams landed 2026-09-30 (§11.1); the
blend across them is one width for every frequency, so an exposure step
the gains leave is narrowed to the seam's 64 px rather than hidden over
the old 200. A Laplacian pyramid would blend low frequencies wide and
detail narrow.
3. **Vignetting in the tap.** The lens profile's distortion is applied
before the fetch; its vignetting is an operation and is not. Frame edges
are darker than their centres by the lens's falloff, and the feather
@@ -402,6 +404,35 @@ half.
catalog's `content_hash` is null for most images most of the time. The
hash can join it when the catalog has one.
### 11.1 Seams — 2026-09-30
The feather averaged every overlap over 200 px, so anything the frames
disagreed on — parallax on the near slope, a walker, wind in a branch — came
out twice at half strength: a soft double edge at 1:1, reported as a glitch.
`dr_pano::seam` now chooses, per output texel at proxy resolution, which
frame it is taken from. Frames are laid down nearest-first; where a new one
overlaps the composite, each texel costs the gain-corrected difference
between the two, plus the detail either has there, plus nearness to either
frame's edge (vignetting, the lens correction's fringe), taken as the
**worst** over a 4-texel window so the path stays a blend radius clear of a
difference rather than grazing it. The cut is a dynamic-programming path
across the overlap, perpendicular to the line from the composite's frames to
the new one: §4's per-column seam, not a graph cut. The map is computed per
projection, for the page's preview and again for the merge.
`merge.wgsl` weights a frame by its tent-filtered share of the label map
about each pixel (`SeamMap::share`, repeated verbatim), over a window
`seam_blend_px` wide (64, capped at 4 texels either side). The edge feather
remains underneath as a factor and, with a 1e-4 floor, as the answer where
the map names no frame that reaches the pixel. `--feather-only` on
`examples/merge.rs` merges the old way, for comparison.
Known limits: one axis per new frame, so in a multi-row set a frame
overlapping its left neighbour and the row above is cut along a compromise
direction; the cost reads grey proxies, so a difference in hue alone is
invisible to it.
## 12. Filling the border instead of cropping it — MI-GAN, read and measured 2026-09-19
Raised after the first merges: the ragged border a cylinder leaves could be
@@ -436,9 +467,11 @@ the tablet. Three ways to make it viable, none built:
1. **Fill at a quarter of the resolution and upsample.** Sky and scree
tolerate it; twenty-odd tiles, about three minutes on the desktop CPU. A
background job with the outbox's patience, not an interactive one.
2. **int8 on the tablet's Hexagon through QNN**, where the plain-conv design
is the point and the whole graph should run in milliseconds. The setup
exists from the eye-state work; MI-GAN is a candidate for the same path.
2. **The tablet's Hexagon through QNN**, where the plain-conv design is the
point. Measured 2026-10-04 (inference.md §1.5): int8 changes the fill
(16 dB from f32's), so it ships with 16-bit activations and weights —
87 ms a tile against 488 ms on the tablet's CPU, the whole graph on the
NPU, 41 dB from f32 in the hole.
3. **A WGSL runtime for those six operators.** A project of its own, and
the only route that would make it interactive on the desktop.
+103 -2
View File
@@ -433,9 +433,18 @@ camera RGB, where its multipliers are defined, and every other operation receive
colour. Before D19 the edits ran in camera RGB and the matrix came after them, so a hue in the
colour mixer and the weights in `luminance()` meant something different on every body.
**Deferred but not foreclosed:** full `.dcp` support with `HueSatDeltas`, `ProfileLookTable`, and
~~**Deferred but not foreclosed:** full `.dcp` support with `HueSatDeltas`, `ProfileLookTable`, and
dual-illuminant interpolation. The stage shall be structured so these are additions rather than a
pipeline reordering.
pipeline reordering.~~ *Amended 2026-10-02 (D20):* dual-illuminant interpolation of the matrices
was built with item 1. The tables follow, designed in [camera-profiles.md](camera-profiles.md):
4. **DCP tables.** `ProfileHueSatMap` (both illuminants, blended as the matrices are) and
`ProfileLookTable`, read from the profile embedded in a DNG or from a `.dcp` file in the
profiles directory matched by `UniqueCameraModel`, the embedded one first. They are applied by a
`camera_profile` scene operation after exposure, with a switch and a look strength (0–200 %),
on by default where a profile exists. `ProfileToneCurve` is read and not applied: tone is the
view transform's (D19). An embedded profile whose `ProfileEmbedPolicy` allows copying can be
saved as a `.dcp` for other files from the same body. The application ships no profile.
Rationale for the reduced scope: a bare 3×3 matrix produces the flat, poor-skin-tone rendering
characteristic of dcraw defaults, which is the documented reason people abandon darktable in the
@@ -448,6 +457,9 @@ measurements, and their provenance was not known well enough to keep them as def
*Acceptance:* the default render is subjectively comparable to the camera's own JPEG — through
FR-DEV-3j's default, for every body. ΔE2000 validation against ColorChecker references applies
once DCP support lands.
For item 4: the lookup follows the DNG SDK's on `[0, 1]` and leaves values above 1.0 above it;
grey and an identity table pass through unchanged; the shader agrees with the CPU reference; with
the switch off the render is to the bit the one with no profile (camera-profiles.md §8).
**FR-DEV-3f — Look emulation.** Support HaldCLUT import, which inherits the existing free film
simulation ecosystem at near-zero implementation cost, plus reading the in-RAF film simulation tag
@@ -593,6 +605,12 @@ until deleted or renamed. Shipped and imported presets change only the operation
name, so a look applied to a corrected photograph keeps the correction; a copy or a saved
edit replaces everything in scope.
The user's presets sync between devices through each library they open, as one file beside the
camera profiles. Each name merges on its own against what the last exchange left both sides
holding, so presets added on two devices both survive, a deletion on one reaches the other rather
than being restored by it, and an edit outlives a deletion made elsewhere. The write is
conditional on the server's copy, so two devices exchanging at once cannot save over each other.
**FR-DEV-7 — Before/after.** Compare current edit state against the unedited original or against
a chosen history state.
@@ -2425,6 +2443,8 @@ Rationale, evidence, and the eliminated alternatives are recorded in
| D12 | Scope versus pace | **DECIDED 2026-09-19** — settled by events; full scope stands, no v1 date |
| D18 | Derived images | **DECIDED 2026-09-19** — a merge writes a new source file; no multi-source Version |
| D19 | Scene-referred pipeline | **DECIDED 2026-09-27** — edits on unbounded scene-linear colour; one view transform, last; per-body base curves retired |
| D21 | DNG reference tone for raws | **DECIDED 2026-10-03** — the view transform's DNG reference curve (profile's, else ACR3 default, via RGBTone in ProPhoto) is the default for every raw, at contrast 1.5 (a ×1.07 power about grey); measured against Lightroom exports of photographs with neutral look settings; the sigmoid stays a choice |
| D20 | DCP camera profiles | **DECIDED 2026-10-02** — HueSatMap and LookTable as a scene operation after exposure; embedded profile first, then a matched `.dcp`; tone curve not applied; none shipped |
### D11 — product positioning
@@ -2730,6 +2750,87 @@ unbounded, but several fragments floor at zero, which clips a colour outside sRG
mixer's bands and the colour grading wheel would need their hues re-measured. Gamut compression
beyond the output transform's clip goes with it.
### D20 — DCP camera profiles · **DECIDED 2026-10-02**
**A camera profile's `HueSatMap` and `LookTable` are applied by a `camera_profile` scene
operation at order 25, after exposure, converting into linear ProPhoto and back inside its own
fragment.** Design and the full argument: [camera-profiles.md](camera-profiles.md).
*Why now.* The library's 9,348 Canon 6D DNGs carry Adobe Standard's tables, which Lightroom
rendered them through, and DarkRoom ignored them, so every hue on those files sat somewhere other
than where Lightroom put it. *Measured after building it:* the tables are not why Lightroom's
rendering looks richer — at defaults they lower mean saturation by 3–9 %, because Adobe Standard's
look desaturates dark tones to sit under Camera Raw's tone curve, which DarkRoom does not apply.
The richer colour is tone, and the Vivid presets (FR-DEV-6) are what answers it today
(camera-profiles.md §1).
*Why there.* The matrix snippet stays what D19 made it, and every copy of it (masks, picker,
camera-space tap) stays correct without changing. Hue and saturation are invariant under the
uniform gains that precede order 25, so a 2.5-D HueSatMap gives the same answer there as straight
after the matrix, and the LookTable sees the photographer's exposure, as it does in the SDK.
*Rejected.* Extending the matrix snippet: every duplicate of it would have had to follow. Two
operations, one per table: the HueSatMap has no control of its own and commutes to the same place.
Applying `ProfileToneCurve`: a per-body tone curve is what D19 retired. Shipping Adobe's profiles:
they are not ours to ship. Handing the tables to the graph through a setter, as lens profiles are:
every render path would have to remember to call it. They travel with the decoded image, as the
matrix does.
*What it costs.* Every DNG with an embedded profile renders differently; previews refresh only when
rendered again; tablet and desktop release together. The profiles directory syncs through the
library's derived folder (camera-profiles.md §13).
### D21 — DNG reference tone for raws · **DECIDED 2026-10-03**
*Decided by measurement, later the same day.* The library's photo gallery holds Lightroom 6
exports of raws that are in the library, each carrying its Camera Raw settings. Clustered by
those settings, 663 exports had none of the house look (Linear curve, no HSL, no parametric
curve, no split toning); 60 of them with their raws, two thirds fitted and one third held out,
rendered by DarkRoom against Lightroom's JPEG (MSE, sRGB 8-bit):
| Rendering | Held-out MSE |
|---|---|
| 0.20.0's sigmoid at its defaults | ~1200 — about 0.7 EV darker, and flatter |
| Sigmoid, exposure, contrast and white fitted | ~150 (contrast 1.73, +0.73 EV) |
| DNG reference curve after baseline exposure, at contrast 1.4 | 224 |
| DNG reference curve, contrast 1.5 | ~150 |
| DNG reference curve, exposure, contrast and white fitted | 143 |
So the DNG reference curve is the default for every raw, and the default contrast is 1.5 — under that
curve a power of 1.5/1.4 about grey (`REFERENCE_CONTRAST` is where the curve is untouched). The
brightness needs nothing: baseline exposure and the curve together land where the earlier exports do. The
profile's look strength, vibrance and saturation bought nothing measurable on those exports. The
user chose to change every photograph rather than keep edited ones on the old rendering. The
fitting tools live outside the repository (`darkroom-lrfit`). *Amended 2026-10-04:* the look strength now defaults to 0. It scored the same at 100, 50 and 0 (held-out MSE 140, 140, 143) and the rendering is 9 % more colourful without it — the table desaturates near-neutral tones, where the default was short of those exports; the user chose more colour.
*Amended earlier the same day:* the default was **not** decided. The measurement below was against
Lightroom previews of photographs carrying the user's Lightroom edits — a house look of HSL
saturation (Blue +58, Aqua +50, …), Highlights −40 and Blacks −20 in every DNG's XMP — so it said
nothing about Camera Raw's base rendering. Under the DNG reference curve `_MG_9080` renders brighter
than its Lightroom preview (mean 0.39 against 0.31). The sigmoid stays the default; the curve
below is a choice; the default is decided by measurement against Lightroom exports of unedited
photographs (the `lr-fit` work). What follows is the original text.
**The view transform has two curves, and Camera Raw's is the default for every raw.** It is the
profile's `ProfileToneCurve`, or the ACR3 default curve where the profile has none or there is no
profile, applied Camera Raw's way — on the largest and smallest channel in linear ProPhoto, the
middle placed proportionally — after `BaselineExposure`. D19's sigmoid stays as the other choice.
Design: [camera-profiles.md](camera-profiles.md) §11–§13.
*Why.* Measured on the library's 6D DNGs after D20 (camera-profiles.md §1): the profile tables
lowered saturation, because Adobe's look tables were tuned to sit under this curve. The user's
complaint was that Lightroom's rendering is more colourful, and this curve is most of the reason.
Chosen by the user over limiting it to raws with a profile, or making it opt-in.
*What it reverses in D19.* D19 rejected per-body curves as defaults because their provenance was
unknown. A DCP's curve and the ACR3 table have known provenance — Adobe's, published — and so the
objection that retired the base curves does not apply. D19's other half stands: nothing before
the view transform clamps, and the curve is the view transform, last.
*What it costs.* Every raw renders differently again, and highlights above display white clip
where the sigmoid rolled them off; Sigmoid is one click away. Previews refresh only when rendered
again; tablet and desktop release together.
### D16 — plugin licensing · **OPEN, post-v1**
> Deferred with §3.10 on 2026-09-19. Still to be answered before the format is published as
+82
View File
@@ -0,0 +1,82 @@
# DarkRoom — Sensor health: a dated defect map per body
**Status:** Spike · 2026-10-04 · not built
**Companion to:** [requirements.md](requirements.md) FR-RAW-3, [catalog.md](catalog.md)
A sensor gains defective photosites as it ages, and a photosite that has gone bad does not recover.
This records, per camera body, which photosites are defective and since when, so that the library
can show how a sensor has aged and the hot-pixel repair can fix the defects a body is known to have
rather than only those that stand out in the frame at hand.
What exists is the measuring tool: `Demosaicer::find_hot_pixels` (the photosites the repair pass
would replace, without replacing them) and `core/dr-gpu/examples/sensor_scan.rs`, which prints them
per frame and, with `--probe`, reads a list of coordinates back out of each frame. The rest of this
document is what a spike with them on the 6D found, and the design it argues for.
---
## 1. What is wanted
- **Settings → Bodies**, one entry per body, with a graph of the defective share over time in two
series: photosites (the raw mosaic) and 2×2 cells holding at least one defective photosite (what
reaches a pixel of the output).
- **A dated defect map** per body, synced with the library like any other catalog data, and
cumulative: an entry is never removed.
- **The repair reads the map** whose date is nearest the frame's, and fixes every defect the body had
by then, whether or not it stands out in that frame.
- One body per model for now: the catalog stores `make model`, not a body serial.
## 2. What the spike found (6D, 2026-10-03)
53 CR2s, up to four per quarter at the highest ISO of a day, 2015 to 2026. The library holds almost
no 6D raws from 2016–2022, so onsets in that span are dated to the span, not the year. `sensor_scan`
ran at 0.8 s per frame, decode included.
**Persistence separates the sensor from the scene.** 4,179 photosites were flagged at least once;
3,554 on one day only (stars, glints, noise). A defect is a photosite that keeps coming back.
**A frame that does not flag a photosite is not evidence it was clean.** The repair's test is
relative to the neighbourhood, so a defect in a lit area does not stand out. Confirmed defects were
flagged in a median 20 % of the frames after their first sighting. Only a frame whose neighbourhood
at that photosite is dark counts, either way.
**The 6D hides some defects itself at high ISO.** (2517, 3172) reads 8,000–13,000 over neighbours
near 200 at ISO 2000–5000, and does not stand out at all at ISO 6400–12800 (82 over 149 on
2025-03-15). The camera appears to map out photosites it knows at those gains. So evidence for this
body comes from ISO ≤ 5000; the cut-off must be learned per body, not fixed.
**Long exposures light everything.** A 9.8 s frame saturated every candidate; it confirms, it does
not date.
**The curve.** Counting a photosite as defective from the first frame where it stands out, provided
it stands out in at least 60 % of the observable frames after that (32 defects; 31 with a clean
observable frame before onset to bracket it):
| Year | Defects | Share of photosites |
|---|---|---|
| 2015 | 2 | 0.1 ppm |
| 2016–2021 | 2 | 0.1 ppm |
| 2022 | 7 | 0.3 ppm |
| 2023 | 26 | 1.3 ppm |
| 2026 | 32 | 1.6 ppm |
(2517, 3172) is the shape every entry should have: clean at ISO 800–1000 in 2015 and at ISO 100–200
in 2016, then 338 over 72 at ISO 100 on 2022-08-13 and in every comparable frame since.
Weak defects exist too, about twice their neighbours ((1814, 3039)); the blind repair misses them in
most frames. A known map would catch them.
## 3. Design it argues for
- **Evidence per frame, per known photosite**: observable (dark neighbourhood, ISO inside the
body's band) and, if so, lit or clean. Not just the frame's flagged list.
- **A map entry is a bracket**: last clean observation, first lit observation, strength, kind. Onset
lies between the two; the graph plots it at the first, and can show the bracket.
- **The repair**: every entry whose first lit date is on or before the frame's capture date; for an
entry whose bracket contains the date, probe the photosite in the frame itself.
- **Storage and sync**: catalog tables created on first use (as `albums` does), so no schema bump
breaks an older peer. They travel in the snapshot by default. Merge is a set union of photosites
per body, the earlier first-lit and the later last-clean winning, which makes it commutative and
keeps the map cumulative.
- **Work**: a sample, not the library. Frames are picked for what they can reveal (dark, mid ISO,
long exposures), a few per body per month, and after the first pass only new imports are read.
+120 -120
View File
File diff suppressed because one or more lines are too long
+2 -2
View File
@@ -287,7 +287,6 @@ $LOCALAPPDATA\Programs\DarkRoom\
models\
scrfd_500m_640.onnx scrfd_2.5g_640.onnx scrfd_10g_640.onnx arcface_mbf_b1.onnx
2d106det_b1.onnx ocec_s_b1.onnx sgc_l_48_b1.onnx
scrfd_500m_640.int8.onnx scrfd_2.5g_640.int8.onnx scrfd_10g_640.int8.onnx
yolo26s-sem-ade20k.onnx yolo26s-sem-ade20k.classes.json categories.txt
migan-512.onnx
manual\
@@ -297,7 +296,8 @@ $LOCALAPPDATA\Programs\DarkRoom\
```
Plus a Start Menu shortcut, and nothing on the desktop unless the user ticks it. The models are
the same ten files the APK bundles and the PKGBUILD installs; `models\` beside the executable is
the f32 files the APK bundles and the PKGBUILD installs — not the APK's quantised siblings, which
only a Hexagon runs; `models\` beside the executable is
where §3.2's lookup finds them. **No `LICENSE` yet**: the repository has no licence file at its
root (the Arch package points at the system's shared GPL text), so the installer has no licence
page until one is added — a one-file change, and the `.nsi` says where the page then goes. The face weights carry the research-only grant that
+24 -24
View File
@@ -39,7 +39,7 @@ The list is longer than it is tall, so a way to walk it that cannot be lost to t
Anchored on the fingers' midpoint, and on the pointer, so the gesture reads as magnifying the picture rather than sliding it about. Double-tap is the way to an exact 1:1; this is the way to everything in between. Past 1:1 the pixels are shown as they are, square and unsmoothed; below it, filtered.
<sub>`ui/dr-ui/ui/app.slint:2025`</sub>
<sub>`ui/dr-ui/ui/app.slint:2029`</sub>
### Move a magnified photograph about
@@ -50,7 +50,7 @@ Anchored on the fingers' midpoint, and on the pointer, so the gesture reads as m
Only once there is something outside the viewport to reach, which is why the cursor becomes a hand exactly then. The view is clamped to the frame: panning past the edge would show undefined area beside the photograph, and that reads as a rendering fault rather than as the end of the picture.
<sub>`ui/dr-ui/ui/app.slint:2121`</sub>
<sub>`ui/dr-ui/ui/app.slint:2125`</sub>
### Paint a mask by hand
@@ -60,7 +60,7 @@ Only once there is something outside the viewport to reach, which is why the cur
A model's mask stops inside a shoulder and leaks into the hair, and no single edge control fixes two errors that go opposite ways. The whole stroke is one step in the history, so taking a mark back costs one press however long it took to make.
<sub>`ui/dr-ui/ui/app.slint:2212`</sub>
<sub>`ui/dr-ui/ui/app.slint:2216`</sub>
### Open this list
@@ -70,7 +70,7 @@ A model's mask stops inside a shoulder and leaks into the hair, and no single ed
Most of the keys are develop's, and a reference that could only be opened from the grid had to be looked up before opening the photograph they were wanted for.
<sub>`ui/dr-ui/ui/app.slint:2438`</sub>
<sub>`ui/dr-ui/ui/app.slint:2442`</sub>
### Take back the last change
@@ -81,7 +81,7 @@ Most of the keys are develop's, and a reference that could only be opened from t
A whole drag is one step, so undo takes back a decision rather than a frame of a gesture. The list is there because arriving six steps back costs what arriving from one does.
<sub>`ui/dr-ui/ui/app.slint:2468`</sub>
<sub>`ui/dr-ui/ui/app.slint:2472`</sub>
### Do it again after taking it back
@@ -90,7 +90,7 @@ A whole drag is one step, so undo takes back a decision rather than a frame of a
- **Keyboard** — `Ctrl+Shift+Z`, or `Ctrl+Y`
- **See it** — [in the manual](manual/README.md#history-snapshots-presets)
<sub>`ui/dr-ui/ui/app.slint:2482`</sub>
<sub>`ui/dr-ui/ui/app.slint:2486`</sub>
### Remove a repair
@@ -98,7 +98,7 @@ A whole drag is one step, so undo takes back a decision rather than a frame of a
- **Pointer** — Click it, then Delete Repair
- **Keyboard** — `Delete` or `Backspace`, while repairing
<sub>`ui/dr-ui/ui/app.slint:2502`</sub>
<sub>`ui/dr-ui/ui/app.slint:2506`</sub>
### Copy the settings from this photograph
@@ -109,7 +109,7 @@ A whole drag is one step, so undo takes back a decision rather than a frame of a
The button is the copy that has to work: a tablet has no modifier key to hold and no menu bar to hang the action from. The shortcut is an accelerator for a control that is on screen either way.
<sub>`ui/dr-ui/ui/app.slint:2521`</sub>
<sub>`ui/dr-ui/ui/app.slint:2525`</sub>
### Paste the settings onto this photograph
@@ -120,7 +120,7 @@ The button is the copy that has to work: a tablet has no modifier key to hold an
The button names what would be pasted — "3 adjustments", and whether the crop is coming with it — which the shortcut cannot say. Both paste the same scope.
<sub>`ui/dr-ui/ui/app.slint:2534`</sub>
<sub>`ui/dr-ui/ui/app.slint:2538`</sub>
### Choose which kinds of edit a copy carries
@@ -131,7 +131,7 @@ The button names what would be pasted — "3 adjustments", and whether the crop
Lightroom's Copy Settings. Pasting a look across a shoot usually means leaving each frame's crop and rotation alone, and that is a choice to make at the moment of copying.
<sub>`ui/dr-ui/ui/app.slint:2552`</sub>
<sub>`ui/dr-ui/ui/app.slint:2556`</sub>
### Export this photograph as the last one was
@@ -142,7 +142,7 @@ Lightroom's Copy Settings. Pasting a look across a shoot usually means leaving e
Every export runs on the defaults in Settings, so "as the last one was" is what the button already does. The chord is Lightroom's and darktable's, kept so hands that learned it there need not learn it again.
<sub>`ui/dr-ui/ui/app.slint:2577`</sub>
<sub>`ui/dr-ui/ui/app.slint:2581`</sub>
### Choose how to export, then export
@@ -153,7 +153,7 @@ Every export runs on the defaults in Settings, so "as the last one was" is what
The export sheet is the export defaults alone with an Export button. What is chosen there is kept, so it is also what the next Ctrl+Shift+E uses.
<sub>`ui/dr-ui/ui/app.slint:2590`</sub>
<sub>`ui/dr-ui/ui/app.slint:2594`</sub>
### Keep a crop that leaves a mask outside
@@ -161,7 +161,7 @@ The export sheet is the export defaults alone with an Export button. What is cho
- **Pointer** — Press "Keep crop" on the notice, or "Undo crop" to take it back
- **Keyboard** — `Enter` keeps it; `Ctrl+Z` takes the crop back, like any other step
<sub>`ui/dr-ui/ui/app.slint:2655`</sub>
<sub>`ui/dr-ui/ui/app.slint:2659`</sub>
### Go back to the grid
@@ -171,7 +171,7 @@ The export sheet is the export defaults alone with an Export button. What is cho
Lightroom's key for the grid. Escape gets there too, but a step at a time — out of a mode, then out of a zoom — where this goes straight back.
<sub>`ui/dr-ui/ui/app.slint:2672`</sub>
<sub>`ui/dr-ui/ui/app.slint:2676`</sub>
### Nudge the control last moved
@@ -181,7 +181,7 @@ Lightroom's key for the grid. Escape gets there too, but a step at a time — ou
Lightroom's keys for the selected slider. There is no focus ring on a slider here, so "selected" is the last one moved — the same control `R` puts back — which covers the framing sliders, perspective included, as well as the adjustments.
<sub>`ui/dr-ui/ui/app.slint:2701`</sub>
<sub>`ui/dr-ui/ui/app.slint:2705`</sub>
### Change which group of adjustments is on screen
@@ -192,7 +192,7 @@ Lightroom's keys for the selected slider. There is no focus ring on a slider her
The groups are whatever the operation set declares itself to be about, so there are as many as the pipeline has and no key can be assigned to one of them by name. Stepping is the binding that survives a node being added.
<sub>`ui/dr-ui/ui/app.slint:2729`</sub>
<sub>`ui/dr-ui/ui/app.slint:2733`</sub>
### Look at the photograph at 1:1
@@ -203,7 +203,7 @@ The groups are whatever the operation set declares itself to be about, so there
Noise reduction and capture sharpening are judgements about single pixels, and a fitted view averages several of the file's into each one on screen — so the frame looks softer than it is and the correction goes too far. The point and the magnification survive opening the next photograph, which is what makes checking the same eye across forty portraits forty keystrokes rather than forty pans. From 1:1 on the photograph is drawn as its own pixels, each a hard-edged square, rather than smoothed into a blur.
<sub>`ui/dr-ui/ui/app.slint:2765`</sub>
<sub>`ui/dr-ui/ui/app.slint:2769`</sub>
### Rate this photograph
@@ -211,7 +211,7 @@ Noise reduction and capture sharpening are judgements about single pixels, and a
- **Pointer** — Click a star in the top bar
- **Keyboard** — `0`–`5`
<sub>`ui/dr-ui/ui/app.slint:2822`</sub>
<sub>`ui/dr-ui/ui/app.slint:2826`</sub>
### Pick or reject this photograph
@@ -221,7 +221,7 @@ Noise reduction and capture sharpening are judgements about single pixels, and a
The grid's keys, on the photograph that is open (FR-UI-5, 2026-09-19). Judging here does not move on to the next frame: that belongs to culling, and in develop the photograph in front of you is the one being worked on.
<sub>`ui/dr-ui/ui/app.slint:2828`</sub>
<sub>`ui/dr-ui/ui/app.slint:2832`</sub>
### Give this photograph a colour label
@@ -232,7 +232,7 @@ The grid's keys, on the photograph that is open (FR-UI-5, 2026-09-19). Judging h
The grid's keys, on the photograph that is open, so labelling while stepping through a folder is one hand's work. The bar names the label in words beside its mark.
<sub>`ui/dr-ui/ui/app.slint:2858`</sub>
<sub>`ui/dr-ui/ui/app.slint:2862`</sub>
### Move to the next or previous photograph
@@ -243,7 +243,7 @@ The grid's keys, on the photograph that is open, so labelling while stepping thr
The edit on screen is saved on the way out, so stepping through a folder is as much a departure as going back to the grid and loses nothing. A and D as well as the arrows, so the left hand steps along the roll while the right stays on the mouse.
<sub>`ui/dr-ui/ui/app.slint:2883`</sub>
<sub>`ui/dr-ui/ui/app.slint:2887`</sub>
### See the photograph before you edited it
@@ -254,7 +254,7 @@ The edit on screen is saved on the way out, so stepping through a folder is as m
Held rather than toggled, and no split screen: a split halves the working image on the tablet the column was sized for, and the comparison photographers describe making is a flick back and forth. It takes no history step, so checking whether a frame is overcooked costs nothing to undo afterwards.
<sub>`ui/dr-ui/ui/app.slint:3008`</sub>
<sub>`ui/dr-ui/ui/app.slint:3012`</sub>
### Put one control back to its default
@@ -338,7 +338,7 @@ The question a correction raises is whether it did what it was for — whether t
One key for "up one", innermost first: a question before the sheet under it, a sheet before the view, a view before the library. Nothing is left behind a dialogue that the key walked straight past.
<sub>`ui/dr-ui/ui/app.slint:1024`</sub>
<sub>`ui/dr-ui/ui/app.slint:1026`</sub>
### Do what a sheet offers
@@ -346,7 +346,7 @@ One key for "up one", innermost first: a question before the sheet under it, a s
- **Pointer** — Press its button — Export, or Copy
- **Keyboard** — `Enter`, on the export and copy sheets
<sub>`ui/dr-ui/ui/app.slint:1034`</sub>
<sub>`ui/dr-ui/ui/app.slint:1036`</sub>
### Scroll by the scrollbar
+47
View File
@@ -205,6 +205,53 @@ the sensor recorded.
![Zooming to 1:1 with a double-click, panning, then further in with the wheel](media/develop-zoom.gif)
### AI denoise
How every raw is developed. `AI Denoise`, at the top of the Adjust panel,
replaces how the camera's raw data is turned into colour: a network trained
on this library's own photographs removes the noise and the blotches of
colour that come with it, while keeping the fine detail. Look at it at
1:1, where noise lives.
`Method` chooses how:
- `Best`, the default: clean skies and sharp lettering, edges kept as
crisp as the camera recorded them.
- `Fast`: a smaller network, taught the same way. Visibly noisier at very
high ISO than `Best`, but still far cleaner than none, and quicker.
- `Bilinear`: the camera's ordinary conversion, noise and all.
A photograph last edited with `Medium`, which earlier versions offered,
opens with `Best`.
The photograph shows the camera's ordinary conversion while the network
works, with its progress in the bar at the top, and changes when it is
done — on a laptop's graphics card, about a second for a 20-megapixel
photograph with `Best`, reading the file included; longer on a processor
alone or on the tablet. After installing, the graphics card spends up to a
quarter of an hour preparing each network, once, in the background; the
photographs developed meanwhile take a little longer. The result is kept, so a photograph opened again,
or exported, does not wait a second time, and switching back to a method
already used is quick.
`Strength` eases it off: below 100 % it puts back some of what was removed,
as grain without colour, for a picture that does not look too smooth.
The lamp and railing of a night frame at ISO 8000, at 1:1, by each method:
| Bilinear | Fast |
|---|---|
| ![The railing and the lamp at ISO 8000, as the camera recorded them](media/develop-denoise-bilinear.png) | ![The same, with the Fast network](media/develop-denoise-fast.png) |
| **Best** | |
| ![The same, with the Best network](media/develop-denoise-best.png) | |
It works on raw files from any camera with the usual colour pattern of
red, green and blue squares — not on JPEGs, and not yet on Fujifilm's
X-Trans. How noisy the camera is at each ISO was measured for the Canon
EOS 6D; for other cameras it is read from a DNG's own figures or
estimated from the photograph, and the finished job in the activity list
says which. An export uses
the method the photograph has.
### Moving between photographs
The roll along the foot of the canvas holds the photographs the grid was
+41
View File
@@ -117,6 +117,7 @@ th { color: var(--ink-dim); font-weight: 600; }
<ul>
<li><a href="#light">Light</a></li>
<li><a href="#looking-closer">Looking closer</a></li>
<li><a href="#ai-denoise">AI denoise</a></li>
<li><a href="#moving-between-photographs">Moving between photographs</a></li>
<li><a href="#white-balance-from-the-photograph">White balance from the photograph</a></li>
<li><a href="#composing">Composing</a></li>
@@ -287,6 +288,46 @@ wheel zooms to any amount in between. Past 1:1 the file's own pixels are
drawn as hard-edged blocks rather than smoothed, so what you see is what
the sensor recorded.</p>
<figure><img loading="lazy" src="media/develop-zoom.gif" alt="Zooming to 1:1 with a double-click, panning, then further in with the wheel"><figcaption>Zooming to 1:1 with a double-click, panning, then further in with the wheel</figcaption></figure>
<h3 id="ai-denoise">AI denoise</h3>
<p>How every raw is developed. <code>AI Denoise</code>, at the top of the Adjust panel,
replaces how the camera's raw data is turned into colour: a network trained
on this library's own photographs removes the noise and the blotches of
colour that come with it, while keeping the fine detail. Look at it at
1:1, where noise lives.</p>
<p><code>Method</code> chooses how:</p>
<ul>
<li><code>Best</code>, the default: clean skies and sharp lettering, edges kept as
crisp as the camera recorded them.</li>
<li><code>Fast</code>: a smaller network, taught the same way. Visibly noisier at very
high ISO than <code>Best</code>, but still far cleaner than none, and quicker.</li>
<li><code>Bilinear</code>: the camera's ordinary conversion, noise and all.</li>
</ul>
<p>A photograph last edited with <code>Medium</code>, which earlier versions offered,
opens with <code>Best</code>.</p>
<p>The photograph shows the camera's ordinary conversion while the network
works, with its progress in the bar at the top, and changes when it is
done — on a laptop's graphics card, about a second for a 20-megapixel
photograph with <code>Best</code>, reading the file included; longer on a processor
alone or on the tablet. After installing, the graphics card spends up to a
quarter of an hour preparing each network, once, in the background; the
photographs developed meanwhile take a little longer. The result is kept, so a photograph opened again,
or exported, does not wait a second time, and switching back to a method
already used is quick.
<code>Strength</code> eases it off: below 100 % it puts back some of what was removed,
as grain without colour, for a picture that does not look too smooth.</p>
<p>The lamp and railing of a night frame at ISO 8000, at 1:1, by each method:</p>
<table><thead><tr><th>Bilinear</th><th>Fast</th></tr></thead><tbody>
<tr><td><img src="media/develop-denoise-bilinear.png" alt="The railing and the lamp at ISO 8000, as the camera recorded them" /></td><td><img src="media/develop-denoise-fast.png" alt="The same, with the Fast network" /></td></tr>
<tr><td><strong>Best</strong></td><td></td></tr>
<tr><td><img src="media/develop-denoise-best.png" alt="The same, with the Best network" /></td><td></td></tr>
</tbody></table>
<p>It works on raw files from any camera with the usual colour pattern of
red, green and blue squares — not on JPEGs, and not yet on Fujifilm's
X-Trans. How noisy the camera is at each ISO was measured for the Canon
EOS 6D; for other cameras it is read from a DNG's own figures or
estimated from the photograph, and the finished job in the activity list
says which. An export uses
the method the photograph has.</p>
<h3 id="moving-between-photographs">Moving between photographs</h3>
<p>The roll along the foot of the canvas holds the photographs the grid was
showing; click one to open it. The right arrow, <code>D</code> or space opens the next,
Binary file not shown.

Some files were not shown because too many files have changed in this diff Show More