Compare commits

..
22 Commits
Author SHA1 Message Date
dtourolle 08b7d23e86 Release 0.22.1
Benchmarks / CPU and I/O (per commit) (push) Successful in 8m35s
Benchmarks / Frame budget (on demand) (push) Skipped
Traceability / Requirement traces (push) Successful in 1m22s
Build and test / Android (aarch64) (push) Successful in 46m43s
Build and test / android-image (push) Successful in 2s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / Desktop (Linux) (push) Successful in 1h4m43s
Build and test / windows-image (push) Successful in 2s
🐳 Windows image / Build and push (push) Successful in 2s
Build and test / Layer separation (push) Successful in 34s
Build and test / Windows (x86_64, cross) (push) Successful in 29m53s
Build and test / Publish the release (push) Successful in 1m24s
2026-10-04 20:35:01 -04:00
dtourolle a116325991 Declare the DSP's RPC library, so the app can reach the Hexagon
QNN's Hexagon stub loads libcdsprpc.so, a vendor library, and from API 31
an app's linker namespace refuses a vendor library its manifest does not
name. QNN then fails to create its device - QNN_DEVICE_ERROR_INVALID_CONFIG,
before it reaches the DSP - and every model ran on the CPU on 0.22.0: AI
denoise took 30-131 s a photograph on the tablet.

The same engine code ran on the HTP from adb's shell, whose namespace has
no such rule, which is what hid it. With the declaration the app's probe
chose the Hexagon on the tablet and compiled every model for it. Not
required, so a device without the library still installs and runs on the
CPU. A test holds the line in the manifest.
2026-10-04 20:19:04 -04:00
dtourolle f20e481358 Probe the Hexagon strictly, and probe again after falling back to the CPU
0.22.0's first launch on the tablet: QNN could not create its device
(QNN_DEVICE_ERROR_INVALID_CONFIG), the session built anyway with every
node on the CPU behind the provider, and the probe timed that - 28.5 ms
against the CPU's own 19.4 - and rejected the Hexagon. The verdict was
cached under the fingerprint, so every model stayed on the CPU on every
later launch: AI denoise took 30-131 s a photograph instead of seconds.
The same A16W8 detector with the APK's own libraries runs on the HTP in
4.4 ms.

The probe's Hexagon session now sets session.disable_cpu_ep_fallback, so
a device that cannot take the graph fails the probe instead of being timed
as the CPU. Only the probe: shipped graphs may keep nodes on the CPU on
purpose. And a selection that fell back to the CPU after an accelerator
failed or lost is probed again on the next launches, up to three probes
per fingerprint; a cache written by 0.22.0 reads as never retried, so the
tablet probes again once this is installed.
2026-10-04 19:42:15 -04:00
dtourolle 16a5957aa7 Compile dr-ui on sixteen codegen units in release builds
Slint expands the .slint files into ~27 MB of Rust, and at the workspace's
single codegen unit LLVM optimised all of it on one thread: 13.5 minutes
of a release build with the other cores idle. The override applies to
dr-ui alone; the image crates keep one unit, and thin LTO still runs at
link time.
2026-10-04 10:45:43 -04:00
dtourolle 5736a21a3a Release 0.22.0
Benchmarks / CPU and I/O (per commit) (push) Successful in 8m36s
Benchmarks / Frame budget (on demand) (push) Skipped
Traceability / Requirement traces (push) Successful in 1m31s
Build and test / Android (aarch64) (push) Successful in 48m34s
Build and test / android-image (push) Successful in 4s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / Desktop (Linux) (push) Successful in 1h32m33s
Build and test / windows-image (push) Successful in 2s
🐳 Windows image / Build and push (push) Successful in 2s
Build and test / Layer separation (push) Successful in 29s
Build and test / Windows (x86_64, cross) (push) Successful in 31m14s
Build and test / Publish the release (push) Successful in 1m9s
2026-10-04 08:13:35 -04:00
dtourolle b6ca7be185 Show each denoise method in the manual as a close-up
The manual's AI denoise section names the four methods and their measured
times, and shows the lamp and railing of the ISO 8000 frame at 1:1 by each
in place of the film and the before/after pair. The scene clicks each
method and waits for that network's result: the repair now logs its own
"learned denoise:" line first, so the wait matches the result's.
2026-10-04 08:11:24 -04:00
dtourolle 06422a07db Offer three denoise networks and a method to choose between them
AI Denoise's Apply switch becomes Method: Bilinear, Fast, Medium, Best,
default Best, so an untouched raw writes nothing and develops through the
mixture. `apply` is still read and never written: 0 is Bilinear, 1 keeps
a network already chosen.

- Best is the mixture of a flat and an edge expert with a learned gate;
  Medium and Fast are students distilled from it. 2.48 s, 0.79 s and
  0.57 s for a 20 MP frame on TensorRT fp16.
- Each network carries its own tile border (256 for the mixture, 192 for
  the students) through `dr_denoise::Shipped` and `TileNet::halo`.
- The file is hashed once at open and each network keys its own cached
  result; Bilinear keeps the result in memory for the way back.
- Each has an .a16w16 sibling for the Hexagon: 0.00 dB on the 6D gate,
  at most 0.11 dB with the noise scaled x0.5 to x4.
- APK BUNDLED 19 -> 23; the PKGBUILD installs all three.
2026-10-04 08:02:25 -04:00
dtourolle 14f08a565f Feed the denoise network without making it wait for the CPU
A 20 MP frame spent 0.32 s outside the network: each tile's mosaic and
sigma gathered on one thread, then its 24 MB output copied out of the
runtime and back into the frame, all in series with the device. Tiles are
now gathered on every core by a producer thread one tile ahead, so the
gather overlaps the run; the centre is written back across cores; and the
tile interface hands its inputs over and lends its output, so neither
side is copied. With a stand-in network that does nothing, the tiler's own
time falls to 0.14 s at the 1408 tile and 0.09 s at 2048. The exactness
and Bayer-phase tests are unchanged and pass.
2026-10-04 07:30:09 -04:00
dtourolle 0c9d564586 Repair photosites beyond 8 sigma of every neighbour before the network
The app's hot-pixel pass takes gross defects only; at ISO 6400-25600 a 6D
frame keeps 1000-2000 photosites more than 8 sigma beyond all their
same-colour and adjacent neighbours, which the network turned into specks.
The same two tests with the threshold in the photosite's own sigma, plus
the factor of two that keeps a bright point of light (where 8 sigma is a
sliver of the signal). The next model is trained behind exactly this; on
an ISO 25600 frame the Rust and training code both repair 935.
2026-10-04 07:30:09 -04:00
dtourolle 75e12441fd Share Lightroom's saturation bands across ours at measured strengths
Photographs opened with an earlier Lightroom edit now import its HSL
saturation as fitted against the library's own Lightroom 6 exports, rather
than one band to one band.

Measured on two looks' exports and their raws (darkroom-lrfit, hsl_map_fit),
by encoded hue: Lightroom's saturation bands act about 45 degrees either
side on our wheel, wider than ours, and not all at our strength. Each is now
shared between two or three of our bands — Aqua mostly cyan and azure, where
skies are; Blue mostly blue and violet; Orange, where skin is, at about 0.4
of its value. Values add when two of Lightroom's bands share one of ours.
On the measured skies the import now lifts muted sky blues about 1.9× against
Lightroom's 2.1×, where it gave 1.15×. Hue and luminance still go one band to
the band of the same hue; they were not measured.
2026-10-04 05:28:33 -04:00
dtourolle 3f8f909e41 Make the colour mixer's saturation reach muted colours
Every photograph with a colour-mixer saturation edit now renders differently:
a raised band is stronger, most of all on muted colours.

The mixer matched bands and judged saturation on scene-linear values, and
scaled chroma by the same factor whatever a colour started at. Against the
photographer's earlier exports of two looks (~90 photographs, their raws, by
encoded hue), a sky band raised by 58 there lifted muted sky blues about
2.1×; here the mixer gave 1.15×, and less in the muted tones that carry most
of a sky or a shadowed snowfield.

Bands are now matched and saturation judged on display-encoded values. A
raised band pushes muted colours hardest and tapers to nothing at full
saturation, at a gain of 3.0, which at the same value lifts muted sky blues
about as those exports did. Lowering saturation still scales every colour
alike. Hue shifts work on the same encoded colour; luminance still scales in
linear light.

The bundled presets that use the mixer, and those whose colour was tuned
against the default rendering, are rescaled to the amount of colour they
had: Vivid 1.30, Vivid warm 1.30, Vivid landscape 1.38, Vivid, strong 1.45,
Vivid portrait 1.15, Punch 1.12, Blue sky 1.08, Deep blue sky 1.12, Polariser
1.26, Blue sky, golden land 1.13 — mean CIELAB chroma over the default
rendering, on 30 raws from the library. Negative values (skin protection)
are left as written.
2026-10-04 05:28:22 -04:00
dtourolle 5ffd54ba43 Leave the profile's look table off by default
Every raw rendered through a camera profile — the library's DNGs with an
embedded profile, and CR2s given one — now renders differently: more
colourful in near-neutral tones. The profile's look table is no longer
applied unless its slider is raised; PROFILE_LOOK names the strength the
profile states.

Against the photographer's earlier exports with no look applied, the default
rendering scores the same with the look table at 100, 50 or 0 (held-out MSE
140, 140, 143), and is 9 % more colourful at 0: the table lowers the
saturation of near-neutral tones, which is exactly where the default
rendering was short of those exports. The user chose more colour.
2026-10-04 05:28:15 -04:00
dtourolle 5a8c3e4c40 Run each model on the Hexagon in the form measured to hold it
The engine knew f32 and int8, and gave the Hexagon int8 for every role it
served. Measured on the tablet itself (inference.md §1.5), int8 lost
5% of the detector's faces at 40-80 px, moved the landmarks 1.5 px,
emptied the segmenter's scores and cost the denoiser 5-9 dB; fp16 the HTP
refuses outright. `Form` gains A16W8 and A16W16, and `Rung::form` now
names one per role: detectors and landmarks A16W8, the segmenter, scene
model, border filler and denoiser A16W16, XFeat int8. The embedder and
the eye classifiers stay on the CPU.

Each loader resolves its `<stem>.<form>.onnx` sibling; the segmenter and
XFeat, compiled into the binary, embed their quantised forms on Android
only and pick through `choose_embedded`. The probe, the compile step and
the cache fingerprint follow the form instead of assuming int8. Detectors
on the new form write `scrfd_*_a16+w600k_mbf`, and `model_ids` answers
for all three spellings.

On the tablet (ORT 1.29 + QNN 2.42), each shipped file against f32 on the
same inputs, and against the CPU's f32 time:
  SCRFD 500m/2.5g/10g  A16W8   100% of faces in every band   4.2/5.1/9.0 ms vs 17/56/198
  landmarks            A16W8   0.25 px in the 192 crop        0.5 ms vs 2.8
  YOLO26n-seg          A16W16  98.2% found, mask IoU 0.994    12.9 ms vs 90
  scene model          A16W16  98.9% of cells agree           15 ms vs 151
  MI-GAN               A16W16  41 dB from f32 in the fill     87 ms vs 488
  XFeat                int8    pano alignment 0.45 px (f32's own spread 0.41)  6.5 ms vs 58
  denoiser             A16W16  0.00 dB at every ISO            95 ms vs 1510 a tile
Face numbers are over public COCO val2017 photographs, not a library.

The APK carries the siblings (BUNDLED 15 -> 19; the old int8 detectors
removed), about 43 MB more. The Windows installer and its CI count skip
them; the Arch and Flatpak packages list their files and never had them.
The ladder example takes a role per model, which is how the per-role
forms above were seen landing on the NPU from the real probe.
2026-10-04 03:45:46 -04:00
dtourolle 0e6ac09fd5 Quantise for the Hexagon with QNN's config and the app's own inputs
tools/quantise-models.sh now writes each model's Hexagon form from a
per-model table: the form its role takes on the NPU (int8, A16W8 or
A16W16), the exact graph rewrites it needs, and the nodes that must stay
float. Ranges are min/max over photographs fed exactly as the app feeds
each model -- the detector and segmenter letterboxes with their own pads
and normalisation, landmark crops from the detector's boxes, MI-GAN with a
panorama-like border, XFeat's grey proxy. The old tool used an
antialiased resize, YOLO's pad of 128 and /255 for every model that was
not a face model, none of which is what the app does.

tools/htp_graph.py holds the rewrites, each checked against the input
graph before use: the denoiser's 6-D Bayer pack and XFeat's 224-slice
unfold as SpaceToDepth (QNN stops at rank 5), computed reshape targets
folded, and bilinear Resize as two MatMuls (the HTP refuses
ResizeBilinear at XFeat's sizes). The denoiser takes ranges computed by
darkroom-denoise's gate on a smaller tile of the same network.
2026-10-04 03:45:07 -04:00
dtourolle 948f6c3ed2 Count the denoise model in the Windows installer's smoke test
package.sh stages models/denoise beside face, scene and inpaint, and
the smoke test counted only the other three, so 0.21.0's Windows job
failed with "expected 14 model files, installed 15". The count reads
the same directories package.sh copies, as its comment intends.
2026-10-04 02:50:15 -04:00
dtourolle 8f9e59b9fa Find hot photosites without repairing them, and measure a sensor's aging
The hot-pixel pass could only repair: it returned how many photosites it
changed and threw away which. find_hot_pixels runs the same pass and
returns them as sensor coordinates, leaving the frame alone, so a sensor's
defects can be tracked across frames.

sensor_scan prints each frame's candidates, and with --probe reads a list
of coordinates back out of every frame. Run over 53 6D raws from 2015 to
2026, it found 32 persistent defects, 2 in 2015 and 32 by 2026, and showed
what a defect map has to account for: a frame that does not flag a
photosite proves nothing unless its neighbourhood is dark, and the 6D
hides some of its defects itself above ISO 5000. docs/dev/sensor-health.md
records the findings and the design they argue for.
2026-10-04 02:38:48 -04:00
dtourolle 83f0461ce7 Sync develop presets through the library
Presets were the one piece of the photographer's work that never left
the device: faces, sidecars, albums, collections, keywords and camera
profiles all travel with the sync pass, the preset library did not.

It now goes to <derived>/presets/library.drpl. PresetLibrary::merge
decides each name against the base the last exchange left (kept per
library beside place.json), so presets added on two devices both
survive, a deletion reaches the other device instead of being restored
by it, and an edit outlives a deletion made elsewhere. The upload is
If-Match / If-None-Match on the server's copy, and a 412 reads and
merges again, so two devices exchanging at once cannot save over each
other. A server copy that will not parse (a newer build's) is left
alone, and a local file that will not read stops the exchange rather
than being taken for an empty library.

The develop view's save merges with the file when the sync changed it
since the view read it, and a sync that brought presets reloads and
redraws the list.

Also corrects the register, which still said camera profiles do not
sync.
2026-10-04 00:50:09 -04:00
dtourolle 26e50ae723 Save presets beside the settings, not under a raw HOME
PresetStore::open built its path from XDG_CONFIG_HOME or HOME. Android
sets neither, so the library resolved to /.config/darkroom, which is
read-only, and every preset saved on the tablet failed. Windows sets no
HOME either and got a directory relative to the working directory. The
settings store was moved to dr_sync::account::config_dir for the same
reason in 0.12.1; the presets now follow it. Linux and macOS resolve to
the same file as before.
2026-10-04 00:49:59 -04:00
dtourolle 25dc0d0179 Give the Vivid presets and Punch measured amounts of colour
On the default rendering, Vivid added 22 % more chroma than the rendering
itself, and Punch 5 % — less than the photographer's earlier exports show
with no look applied (14 % over ours) and well below their everyday look
(27 %). Each preset's colour values (vibrance, saturation, the mixer's
saturation bands) are now scaled together, tone values untouched and
negative ones — Vivid portrait's skin protection — left as written, until
the preset measures: Vivid and Vivid warm 1.30, Vivid landscape 1.38, Vivid,
strong 1.45, Vivid portrait 1.15, Punch 1.12. Measured as mean CIELAB chroma
over 30 raws from the library, as a ratio to the default rendering.
2026-10-03 23:04:46 -04:00
dtourolle a36ec98b36 Carry Lightroom's tone sliders across at measured strengths
Contrast2012 and the four recovery sliders were imported one to one. They do
not mean the same thing here: fitted on the library's Lightroom 6 exports and
their raws — each photograph's sliders carried across as slider × factor, one
factor per slider, on about 90 exports with no look applied, on the Camera Raw
default rendering — ours needed contrast at about a tenth (Lightroom's −100
imported as ours flattens a frame to grey), highlights ×1.4, shadows ×1.9 and
blacks ×1.25. Whites fitted below 1 every time without agreeing where; 0.5 is
a hedge, and says so. Vibrance stays one to one: the op itself is now
calibrated to Lightroom's.
2026-10-03 23:04:46 -04:00
dtourolle c343ac79d3 Describe AI Denoise in the manual as it now is
On for every raw, at the top of the Adjust panel, kept once computed,
and eased off with Strength rather than Keep grain. The timing line is
left as it was; the new model's measured figure replaces it when that
branch lands. The animation still shows the Keep grain slider and
wants recording again.
2026-10-03 22:18:15 -04:00
dtourolle 2e7f14dafe Develop every raw through the AI denoise by default, with a strength, cached
The learned demosaic was an option under Detail, off by default. It is
now how a Bayer raw is developed: on by default at full strength on
every device — which hardware runs it is the inference engine's choice
— and first in the Adjust panel, since it decides what every control
below is applied to.

Strength (0-100, default 100) replaces Keep grain: grain = 100 -
strength, the same luminance-only blend, so moving it is one GPU pass
and never a re-run. 0.21.0's sidecars stored grain; it is still read,
as the inverse, and never written.

With it on for every photograph, the result is now kept on disk
(denoise.md §7.1, §12): the network's output as half floats, keyed on
a SHA-256 of the file's bytes and the model, oldest first past a 5 GB
budget, beside the inference engine's cache. A reopened photograph and
an export of one already developed read it back instead of running the
network again; a damaged entry is a miss.
2026-10-03 22:16:36 -04:00
96 changed files with 3720 additions and 796 deletions
+5 -4
View File
@@ -487,10 +487,11 @@ jobs:
wine "$SETUP" /S 2>/dev/null
INST=$(echo "$HOME"/.wine/drive_c/users/*/AppData/Local/Programs/DarkRoom)
ls "$INST"
# As many files as package.sh stages: everything but the READMEs in
# the directories it copies. A literal here went stale the first
# time a model was added.
WANT=$(find models/face models/scene models/inpaint -maxdepth 1 -type f ! -name README.md | wc -l)
# As many files as package.sh stages: everything but the READMEs and
# the Hexagon's quantised siblings in the directories it copies. A
# literal here went stale the first time a model was added.
WANT=$(find models/face models/scene models/inpaint models/denoise -maxdepth 1 -type f ! -name README.md \
! -name '*.int8.onnx' ! -name '*.a16w8.onnx' ! -name '*.a16w16.onnx' | wc -l)
GOT=$(ls "$INST/models" | wc -l)
[ "$GOT" = "$WANT" ] || { echo "FAIL: expected $WANT model files, installed $GOT"; exit 1; }
# The manual, and every picture it shows, counted the same way.
Generated
+27 -26
View File
@@ -1265,7 +1265,7 @@ checksum = "f27ae1dd37df86211c42e150270f82743308803d90a6f6e6651cd730d5e1732f"
[[package]]
name = "darkroom-android"
version = "0.21.0"
version = "0.22.1"
dependencies = [
"android_logger",
"dr-plat",
@@ -1278,7 +1278,7 @@ dependencies = [
[[package]]
name = "darkroom-desktop"
version = "0.21.0"
version = "0.22.1"
dependencies = [
"anyhow",
"dr-plat",
@@ -1454,7 +1454,7 @@ checksum = "d8b14ccef22fc6f5a8f4d7d768562a182c04ce9a3b3157b91390b52ddfdf1a76"
[[package]]
name = "dr-bench"
version = "0.21.0"
version = "0.22.1"
dependencies = [
"anyhow",
"dr-catalog",
@@ -1471,7 +1471,7 @@ dependencies = [
[[package]]
name = "dr-catalog"
version = "0.21.0"
version = "0.22.1"
dependencies = [
"dr-face",
"dr-plat",
@@ -1486,7 +1486,7 @@ dependencies = [
[[package]]
name = "dr-decode"
version = "0.21.0"
version = "0.22.1"
dependencies = [
"dr-types",
"env_logger",
@@ -1500,7 +1500,7 @@ dependencies = [
[[package]]
name = "dr-denoise"
version = "0.21.0"
version = "0.22.1"
dependencies = [
"dr-decode",
"dr-gpu",
@@ -1517,7 +1517,7 @@ dependencies = [
[[package]]
name = "dr-export"
version = "0.21.0"
version = "0.22.1"
dependencies = [
"dr-decode",
"dr-gpu",
@@ -1536,7 +1536,7 @@ dependencies = [
[[package]]
name = "dr-face"
version = "0.21.0"
version = "0.22.1"
dependencies = [
"dr-inference-engine",
"env_logger",
@@ -1549,7 +1549,7 @@ dependencies = [
[[package]]
name = "dr-film"
version = "0.21.0"
version = "0.22.1"
dependencies = [
"log",
"serde",
@@ -1558,7 +1558,7 @@ dependencies = [
[[package]]
name = "dr-gpu"
version = "0.21.0"
version = "0.22.1"
dependencies = [
"bytemuck",
"dr-decode",
@@ -1576,7 +1576,7 @@ dependencies = [
[[package]]
name = "dr-inference-engine"
version = "0.21.0"
version = "0.22.1"
dependencies = [
"env_logger",
"libloading",
@@ -1591,7 +1591,7 @@ dependencies = [
[[package]]
name = "dr-ingest"
version = "0.21.0"
version = "0.22.1"
dependencies = [
"dr-plat",
"dr-types",
@@ -1603,7 +1603,7 @@ dependencies = [
[[package]]
name = "dr-lens"
version = "0.21.0"
version = "0.22.1"
dependencies = [
"lensfun",
"log",
@@ -1611,7 +1611,7 @@ dependencies = [
[[package]]
name = "dr-pano"
version = "0.21.0"
version = "0.22.1"
dependencies = [
"dr-decode",
"dr-inference-engine",
@@ -1625,7 +1625,7 @@ dependencies = [
[[package]]
name = "dr-pipeline"
version = "0.21.0"
version = "0.22.1"
dependencies = [
"dr-types",
"log",
@@ -1634,7 +1634,7 @@ dependencies = [
[[package]]
name = "dr-plat"
version = "0.21.0"
version = "0.22.1"
dependencies = [
"android-native-keyring-store",
"dr-types",
@@ -1650,7 +1650,7 @@ dependencies = [
[[package]]
name = "dr-preset-xmp"
version = "0.21.0"
version = "0.22.1"
dependencies = [
"dr-pipeline",
"log",
@@ -1660,7 +1660,7 @@ dependencies = [
[[package]]
name = "dr-segment"
version = "0.21.0"
version = "0.22.1"
dependencies = [
"dr-inference-engine",
"env_logger",
@@ -1673,7 +1673,7 @@ dependencies = [
[[package]]
name = "dr-sync"
version = "0.21.0"
version = "0.22.1"
dependencies = [
"async-trait",
"dr-plat",
@@ -1687,7 +1687,7 @@ dependencies = [
[[package]]
name = "dr-sync-folder"
version = "0.21.0"
version = "0.22.1"
dependencies = [
"async-trait",
"dr-sync",
@@ -1699,7 +1699,7 @@ dependencies = [
[[package]]
name = "dr-sync-nextcloud"
version = "0.21.0"
version = "0.22.1"
dependencies = [
"async-trait",
"dr-decode",
@@ -1721,7 +1721,7 @@ dependencies = [
[[package]]
name = "dr-thumbs"
version = "0.21.0"
version = "0.22.1"
dependencies = [
"dr-types",
"jpeg-encoder",
@@ -1733,7 +1733,7 @@ dependencies = [
[[package]]
name = "dr-types"
version = "0.21.0"
version = "0.22.1"
dependencies = [
"serde",
"serde_json",
@@ -1742,7 +1742,7 @@ dependencies = [
[[package]]
name = "dr-ui"
version = "0.21.0"
version = "0.22.1"
dependencies = [
"anyhow",
"async-trait",
@@ -1768,6 +1768,7 @@ dependencies = [
"dr-types",
"dr-xmp",
"env_logger",
"half",
"i-slint-backend-testing",
"jni 0.22.4",
"log",
@@ -1791,7 +1792,7 @@ dependencies = [
[[package]]
name = "dr-xmp"
version = "0.21.0"
version = "0.22.1"
dependencies = [
"dr-types",
"log",
@@ -7125,7 +7126,7 @@ checksum = "8df9b6e13f2d32c91b9bd719c00d1958837bc7dec474d94952798cc8e69eeec3"
[[package]]
name = "traceability"
version = "0.21.0"
version = "0.22.1"
dependencies = [
"anyhow",
"proc-macro2",
+9 -1
View File
@@ -33,7 +33,7 @@ members = [
exclude = ["third_party"]
[workspace.package]
version = "0.21.0"
version = "0.22.1"
edition = "2021"
rust-version = "1.92"
license = "GPL-3.0-or-later"
@@ -278,6 +278,14 @@ opt-level = 0
lto = "thin"
codegen-units = 1
# Except dr-ui. Slint expands the `.slint` files into ~27 MB of Rust
# (`out/app.rs`), and at one codegen unit LLVM optimises all of it on a single
# thread: 13.5 minutes of a release build with the other cores idle. The code
# it holds is UI glue — property bindings and callbacks — not the image work,
# which lives in the crates above that keep the single unit.
[profile.release.package.dr-ui]
codegen-units = 16
# A release build that can say where it panicked: line tables, so a crash
# record's backtrace (`dr_plat::crash`) reads `file.rs:123` rather than bare
# addresses. The macOS build uses it (docs/dev/macos.md) — no one here can
+1 -1
View File
@@ -201,7 +201,7 @@ controls, its place in the chain and its tests.
## Where it stands
**0.21.0**, thirty-five tagged releases in. 193 numbered requirements in
**0.22.1**, thirty-seven tagged releases in. 193 numbered requirements in
scope, 85% of them claimed by code and [traced to it](docs/dev/traceability.md);
the rest are written down rather than merely absent.
@@ -73,6 +73,14 @@
android:requestLegacyExternalStorage="true"
android:supportsRtl="true">
<!-- The DSP's RPC library, which QNN's Hexagon stub loads. From API 31
an app's linker namespace refuses a vendor library the manifest
does not name, and QNN then fails to create its device
(QNN_DEVICE_ERROR_INVALID_CONFIG) before it reaches the DSP:
every model ran on the CPU on 0.22.0. Not required, so a device
without one still installs and stays on the CPU. -->
<uses-native-library android:name="libcdsprpc.so" android:required="false" />
<!-- NativeActivity rather than a Kotlin Activity: android-activity's
glue loads libdarkroom.so and calls android_main. `android.app.lib_name`
is how it learns which library to load, and must match [lib].name.
+51 -10
View File
@@ -332,27 +332,38 @@ fn unpack_bundled_models(app: &slint::android::AndroidApp) {
// eyes-open filter has something to read, and a tablet has no other way
// to get them either.
//
// The int8 forms beside the three detectors are what the Hexagon runs
// (docs/dev/inference.md §5); the engine loads the sibling when the probe
// chose that rung and ignores it otherwise.
const BUNDLED: [(&std::ffi::CStr, &str); 15] = [
// The quantised siblings — `.a16w8.onnx`, `.a16w16.onnx` — are what the
// Hexagon runs (docs/dev/inference.md §1.5), each in the narrowest form
// that held that model's accuracy on the tablet; the engine loads the
// sibling when the probe chose that rung and ignores it otherwise. The
// segmenter's and XFeat's forms are compiled into the binary instead,
// beside their f32 graphs.
const BUNDLED: [(&std::ffi::CStr, &str); 23] = [
(c"models/scrfd_500m_640.onnx", "scrfd_500m_640.onnx"),
(
c"models/scrfd_500m_640.int8.onnx",
"scrfd_500m_640.int8.onnx",
c"models/scrfd_500m_640.a16w8.onnx",
"scrfd_500m_640.a16w8.onnx",
),
(c"models/scrfd_2.5g_640.onnx", "scrfd_2.5g_640.onnx"),
(
c"models/scrfd_2.5g_640.int8.onnx",
"scrfd_2.5g_640.int8.onnx",
c"models/scrfd_2.5g_640.a16w8.onnx",
"scrfd_2.5g_640.a16w8.onnx",
),
(c"models/scrfd_10g_640.onnx", "scrfd_10g_640.onnx"),
(c"models/scrfd_10g_640.int8.onnx", "scrfd_10g_640.int8.onnx"),
(
c"models/scrfd_10g_640.a16w8.onnx",
"scrfd_10g_640.a16w8.onnx",
),
(c"models/arcface_mbf_b1.onnx", "arcface_mbf_b1.onnx"),
(c"models/2d106det_b1.onnx", "2d106det_b1.onnx"),
(c"models/2d106det_b1.a16w8.onnx", "2d106det_b1.a16w8.onnx"),
(c"models/ocec_s_b1.onnx", "ocec_s_b1.onnx"),
(c"models/sgc_l_48_b1.onnx", "sgc_l_48_b1.onnx"),
(c"models/yolo26s-sem-ade20k.onnx", "yolo26s-sem-ade20k.onnx"),
(
c"models/yolo26s-sem-ade20k.a16w16.onnx",
"yolo26s-sem-ade20k.a16w16.onnx",
),
(
c"models/yolo26s-sem-ade20k.classes.json",
"yolo26s-sem-ade20k.classes.json",
@@ -360,7 +371,24 @@ fn unpack_bundled_models(app: &slint::android::AndroidApp) {
(c"models/categories.txt", "categories.txt"),
// The panorama border filler (FR-MRG-4); MIT, 28 MB.
(c"models/migan-512.onnx", "migan-512.onnx"),
(c"models/mosaic-1408.onnx", "mosaic-1408.onnx"),
(c"models/migan-512.a16w16.onnx", "migan-512.a16w16.onnx"),
// The learned demosaic and denoise, one network per method
// (FR-DEV-3g), each with the 16-bit form the Hexagon runs.
(c"models/mosaic-fast-1408.onnx", "mosaic-fast-1408.onnx"),
(
c"models/mosaic-fast-1408.a16w16.onnx",
"mosaic-fast-1408.a16w16.onnx",
),
(c"models/mosaic-medium-1408.onnx", "mosaic-medium-1408.onnx"),
(
c"models/mosaic-medium-1408.a16w16.onnx",
"mosaic-medium-1408.a16w16.onnx",
),
(c"models/mosaic-best-1408.onnx", "mosaic-best-1408.onnx"),
(
c"models/mosaic-best-1408.a16w16.onnx",
"mosaic-best-1408.a16w16.onnx",
),
];
let dir = dr_ui::shared_face_models_dir();
@@ -510,6 +538,19 @@ mod tests {
Some(value.to_string())
}
/// From API 31 the linker refuses a vendor library the manifest does not
/// name, and QNN cannot create its Hexagon device without the DSP's RPC
/// library: 0.22.0 ran every model on the CPU for want of this line.
#[test]
fn the_npu_can_reach_the_dsp() {
assert!(
manifest().contains(
r#"<uses-native-library android:name="libcdsprpc.so" android:required="false" />"#
),
"libcdsprpc.so must be declared, and not required"
);
}
#[test]
fn a_gallery_can_open_a_photograph_in_this_app() {
let manifest = manifest();
+16 -6
View File
@@ -2,11 +2,11 @@
//!
//! ```sh
//! DARKROOM_ORT_DIR=~/.local/share/darkroom/runtime \
//! cargo run --release -p dr-denoise --features native --example denoise_raw -- IMG.CR2 out
//! cargo run --release -p dr-denoise --features native --example denoise_raw -- IMG.CR2 out [fast|medium|best]
//! ```
//!
//! Decode, the app's hot-pixel pass, the frame's noise from its best source,
//! then the shipped network under the inference engine on whatever rung this
//! then one of the shipped networks (`best` unless named) under the inference engine on whatever rung this
//! machine probes to. Writes `out.npy` — the active area, `h×w×3` f32 linear
//! camera RGB — for comparison with the training repo's own path
//! (`tools/compare_rust.py` in darkroom-denoise). `DARKROOM_ORT_DIR` points
@@ -23,11 +23,21 @@ fn main() {
env_logger::Builder::from_env(env_logger::Env::default().default_filter_or("warn")).init();
let mut args = std::env::args().skip(1);
let (Some(input), Some(out)) = (args.next(), args.next()) else {
eprintln!("usage: denoise_raw RAW OUT_PREFIX");
eprintln!("usage: denoise_raw RAW OUT_PREFIX [fast|medium|best]");
std::process::exit(2);
};
let model =
PathBuf::from(env!("CARGO_MANIFEST_DIR")).join("../../models/denoise/mosaic-1408.onnx");
let shipped = match args.next().as_deref() {
None | Some("best") => dr_denoise::BEST,
Some("medium") => dr_denoise::MEDIUM,
Some("fast") => dr_denoise::FAST,
Some(other) => {
eprintln!("no network called {other}: fast, medium or best");
std::process::exit(2);
}
};
let model = PathBuf::from(env!("CARGO_MANIFEST_DIR"))
.join("../../models/denoise")
.join(shipped.file);
let cache = std::env::var_os("DR_ENGINE_CACHE")
.map(PathBuf::from)
.unwrap_or_else(|| std::env::temp_dir().join("dr-denoise-engines"));
@@ -91,7 +101,7 @@ fn main() {
noise.col
);
let mut net = OnnxNet::from_path(&model).expect("model");
let mut net = OnnxNet::from_path(&model, shipped.halo).expect("model");
println!(
"rung {}",
net.rung().map(|r| r.label()).unwrap_or("?")
+40 -1
View File
@@ -19,6 +19,7 @@
pub mod noise;
#[cfg(feature = "onnx")]
pub mod onnx;
pub mod repair;
pub mod tile;
use dr_decode::RawImage;
@@ -26,6 +27,33 @@ use dr_decode::RawImage;
pub use noise::{NoiseModel, Source};
pub use tile::{TileNet, HALO};
/// TRACES: FR-DEV-3g
/// A network the app ships in `models/denoise/`: its file, and the context
/// it needs past a tile's kept centre (docs/dev/denoise.md §13).
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub struct Shipped {
pub file: &'static str,
pub halo: usize,
}
/// The smallest student: 0.9 M parameters, 11 GMAC a megapixel.
pub const FAST: Shipped = Shipped {
file: "mosaic-fast-1408.onnx",
halo: HALO,
};
/// A student of the mixture with the first release's shape: 3.2 M
/// parameters, 48 GMAC a megapixel.
pub const MEDIUM: Shipped = Shipped {
file: "mosaic-medium-1408.onnx",
halo: HALO,
};
/// The mixture: a flat expert, an edge expert and the gate that blends them.
/// It reaches further than either, so it keeps a smaller centre of each tile.
pub const BEST: Shipped = Shipped {
file: "mosaic-best-1408.onnx",
halo: 256,
};
#[derive(Debug, thiserror::Error)]
pub enum DenoiseError {
#[error("the network cannot take this photograph: {0}")]
@@ -67,12 +95,23 @@ pub fn denoise(
}
let active = noise::active(raw);
let (h, w) = (active.h, active.w);
// The active area laid out once, then the noise-aware repair the model
// was trained behind (see `repair`).
let mut mosaic: Vec<f32> = (0..h * w).map(|i| active.at(i / w, i % w)).collect();
let pattern = raw.cfa_pattern;
let repaired = repair::repair(&mut mosaic, h, w, repair::REPAIR_K, &|y, x, v| {
noise.sigma(pattern.colour_at(x as u32, y as u32) as usize, v)
});
log::info!(
"learned denoise: {repaired} photosites beyond {}σ of every neighbour repaired",
repair::REPAIR_K
);
tile::run_tiled(
net,
h,
w,
raw.cfa_pattern,
&|y, x| active.at(y, x),
&|y, x| mosaic[y * w + x],
&|c, v| noise.sigma(c, v),
progress,
)
+26 -6
View File
@@ -5,7 +5,11 @@
//! returns `rgb`, `1×3×1408×1408` (darkroom-denoise `denoise/export.py`,
//! fixed shape because every model the engine runs is). The engine picks the
//! rung: fp16 on TensorRT and MIGraphX, which measured 0.00 dB from f32; f32
//! on CUDA and the CPU; never the Hexagon, where int8 lost 6–9 dB.
//! on CUDA and the CPU; on the Hexagon the `.a16w16.onnx` sibling, 16-bit
//! activations and weights, 0.00 dB from f32 on the tablet itself where int8
//! lost 5–9 dB (docs/dev/inference.md §1.5). That sibling is the same network
//! with the Bayer packing spelled `SpaceToDepth`, which QNN can hold and the
//! 6-D reshape it replaces it cannot.
use crate::tile::TileNet;
use crate::DenoiseError;
@@ -17,15 +21,19 @@ pub const TILE: usize = 1408;
pub struct OnnxNet {
model: Model,
tile: usize,
halo: usize,
}
impl OnnxNet {
pub fn from_path(path: &std::path::Path) -> Result<Self, DenoiseError> {
/// The network at `path`, which needs `halo` photosites of context
/// ([`crate::Shipped::halo`]).
pub fn from_path(path: &std::path::Path, halo: usize) -> Result<Self, DenoiseError> {
let (path, form) = dr_inference_engine::resolve_model(Role::Denoiser, path);
let bytes = std::fs::read(&path)?;
Ok(OnnxNet {
model: dr_inference_engine::open(Role::Denoiser, form, &bytes)?,
tile: TILE,
halo,
})
}
@@ -40,15 +48,25 @@ impl TileNet for OnnxNet {
self.tile
}
fn run(&mut self, mosaic: &[f32], sigma: &[f32]) -> Result<Vec<f32>, DenoiseError> {
fn halo(&self) -> usize {
self.halo
}
fn run(
&mut self,
mosaic: Vec<f32>,
sigma: Vec<f32>,
write: &mut dyn FnMut(&[f32]),
) -> Result<(), DenoiseError> {
let n = self.tile;
let shape = ndarray::IxDyn(&[1, 1, n, n]);
// The vectors become the tensors: no copy on the way in.
let m = ort::value::Tensor::from_array(
ndarray::Array::from_shape_vec(shape.clone(), mosaic.to_vec())
ndarray::Array::from_shape_vec(shape.clone(), mosaic)
.map_err(|e| DenoiseError::Model(e.to_string()))?,
)?;
let s = ort::value::Tensor::from_array(
ndarray::Array::from_shape_vec(shape, sigma.to_vec())
ndarray::Array::from_shape_vec(shape, sigma)
.map_err(|e| DenoiseError::Model(e.to_string()))?,
)?;
let acquired = self.model.acquire()?;
@@ -61,6 +79,8 @@ impl TileNet for OnnxNet {
"output is {dims:?}, expected [1, 3, {n}, {n}]"
)));
}
Ok(data.to_vec())
// And none on the way out: the frame is written from the runtime's buffer.
write(data);
Ok(())
}
}
+145
View File
@@ -0,0 +1,145 @@
//! TRACES: FR-DEV-3g
//! Hot and dead photosites, judged against the noise, before the network.
//!
//! The app's own pass (`dr_gpu::Demosaicer::repair_hot_pixels`) runs first and
//! takes the gross defects. At high ISO it leaves thousands of photosites per
//! 6D frame more than 8σ beyond every neighbour, which the network turns into
//! specks. This second pass uses that pass's two tests with the threshold in
//! units of the photosite's own σ from the noise model:
//!
//! - beyond every same-colour neighbour (two photosites away, the 3×3 of its
//! plane) by more than `k·σ`, and
//! - beyond every adjacent photosite, whatever its colour, by more than
//! `k·σ` **and** by a factor of two — what keeps a real point of light,
//! which lights its neighbours through the lens and the anti-aliasing
//! filter. A margin in σ alone is not enough: on a bright star 8σ is a
//! sliver of the signal, and the star would be flattened.
//!
//! A hot one becomes its brightest same-colour neighbour, a dead one its
//! darkest. The shipped model was trained on input repaired exactly so
//! (darkroom-denoise `denoise/repair.py`, `--repair-k 8`): the threshold
//! belongs to the model, and changes with it. Neighbours off the frame are
//! the nearest photosite on it, as the training code reads them.
/// The threshold the shipped model was trained with, in σ.
pub const REPAIR_K: f32 = 8.0;
/// Repair `mosaic` (`h×w`, row-major, normalised) in place; `sigma(y, x, v)`
/// is the photosite's σ. Returns how many photosites changed.
pub fn repair(
mosaic: &mut [f32],
h: usize,
w: usize,
k: f32,
sigma: &(dyn Fn(usize, usize, f32) -> f32 + Sync),
) -> usize {
let copy = mosaic.to_vec();
let original = &copy;
let at = |y: isize, x: isize| {
let y = y.clamp(0, h as isize - 1) as usize;
let x = x.clamp(0, w as isize - 1) as usize;
original[y * w + x]
};
let threads = std::thread::available_parallelism().map_or(1, |n| n.get());
let rows_per = h.div_ceil(threads).max(1);
let mut counts = vec![0usize; h.div_ceil(rows_per)];
std::thread::scope(|scope| {
for ((chunk, rows), count) in mosaic
.chunks_mut(rows_per * w)
.enumerate()
.zip(counts.iter_mut())
{
let at = &at;
scope.spawn(move || {
for (i, row) in rows.chunks_mut(w).enumerate() {
let y = chunk * rows_per + i;
for (x, out) in row.iter_mut().enumerate() {
let v = original[y * w + x];
let (yi, xi) = (y as isize, x as isize);
let (mut s_hi, mut s_lo) = (f32::MIN, f32::MAX);
let (mut a_hi, mut a_lo) = (f32::MIN, f32::MAX);
for dy in -1isize..=1 {
for dx in -1isize..=1 {
if dy == 0 && dx == 0 {
continue;
}
let s = at(yi + 2 * dy, xi + 2 * dx);
s_hi = s_hi.max(s);
s_lo = s_lo.min(s);
let a = at(yi + dy, xi + dx);
a_hi = a_hi.max(a);
a_lo = a_lo.min(a);
}
}
let t = k * sigma(y, x, v);
if v - s_hi > t && v - a_hi > t && a_hi < 0.5 * v {
*out = s_hi;
*count += 1;
} else if s_lo - v > t && a_lo - v > t && v < 0.5 * a_lo {
*out = s_lo;
*count += 1;
}
}
}
});
}
});
counts.iter().sum()
}
#[cfg(test)]
mod tests {
use super::*;
const N: usize = 16;
fn flat(level: f32) -> Vec<f32> {
vec![level; N * N]
}
fn run(m: &mut [f32]) -> usize {
repair(m, N, N, REPAIR_K, &|_, _, _| 0.01)
}
#[test]
fn a_hot_photosite_becomes_its_brightest_same_colour_neighbour() {
let mut m = flat(0.1);
m[8 * N + 8] = 0.5; // 40σ above everything around it
m[8 * N + 10] = 0.12; // a same-colour neighbour, a little brighter
assert_eq!(run(&mut m), 1);
assert_eq!(m[8 * N + 8], 0.12);
}
#[test]
fn a_dead_photosite_in_a_lit_area_is_repaired() {
let mut m = flat(0.5);
m[5 * N + 5] = 0.0;
assert_eq!(run(&mut m), 1);
assert_eq!(m[5 * N + 5], 0.5);
}
#[test]
fn a_point_of_real_light_is_kept() {
// Light through a lens lands on a patch: its adjacent photosites are
// lit too, so the second test refuses it.
let mut m = flat(0.1);
for dy in 0..3 {
for dx in 0..3 {
m[(7 + dy) * N + 7 + dx] = if (dy, dx) == (1, 1) { 0.9 } else { 0.6 };
}
}
let before = m.clone();
assert_eq!(run(&mut m), 0);
assert_eq!(m, before);
}
#[test]
fn noise_within_the_threshold_is_left_alone() {
let mut m: Vec<f32> = (0..N * N)
.map(|i| 0.1 + 0.005 * ((i * 7919 % 13) as f32 - 6.0) / 6.0)
.collect();
let before = m.clone();
assert_eq!(run(&mut m), 0);
assert_eq!(m, before);
}
}
+204 -57
View File
@@ -15,15 +15,32 @@
use dr_decode::CfaPattern;
/// Photosites of context beyond a tile's kept centre, on every side.
/// Photosites of context beyond a tile's kept centre, on every side, for a
/// single network; a mixture reaches further and says so through
/// [`TileNet::halo`].
pub const HALO: usize = 192;
/// A fixed-shape network: `mosaic` and `sigma`, `n×n` RGGB, in; `3×n×n`
/// planar linear camera RGB out.
///
/// The inputs are handed over, and the output is lent to `write` rather than
/// returned: a 1408² tile is 24 MB of output, and copying it out of the
/// runtime's buffer and back into the frame was a measurable share of a
/// frame's time.
pub trait TileNet {
/// The edge `n` of the square tile the network takes.
fn tile(&self) -> usize;
fn run(&mut self, mosaic: &[f32], sigma: &[f32]) -> Result<Vec<f32>, crate::DenoiseError>;
/// Photosites of context it needs past a tile's kept centre: at least
/// its receptive field. [`HALO`] unless the network says otherwise.
fn halo(&self) -> usize {
HALO
}
fn run(
&mut self,
mosaic: Vec<f32>,
sigma: Vec<f32>,
write: &mut dyn FnMut(&[f32]),
) -> Result<(), crate::DenoiseError>;
}
/// Index into `0..n` by reflection about the end photosites, any distance
@@ -61,76 +78,138 @@ pub fn run_tiled(
h: usize,
w: usize,
pattern: CfaPattern,
at: &dyn Fn(usize, usize) -> f32,
sigma: &dyn Fn(usize, f32) -> f32,
at: &(dyn Fn(usize, usize) -> f32 + Sync),
sigma: &(dyn Fn(usize, f32) -> f32 + Sync),
progress: &mut dyn FnMut(usize, usize) -> bool,
) -> Result<Option<Vec<f32>>, crate::DenoiseError> {
let (dy, dx) = rggb_offset(pattern).ok_or_else(|| {
crate::DenoiseError::Unsupported(format!("{pattern:?} is not a Bayer pattern"))
})?;
let n = net.tile();
if n <= 2 * HALO || !(n - 2 * HALO).is_multiple_of(2) {
let (n, halo) = (net.tile(), net.halo());
if n <= 2 * halo || !(n - 2 * halo).is_multiple_of(2) {
return Err(crate::DenoiseError::Model(format!(
"tile {n} leaves no even centre past a {HALO} halo"
"tile {n} leaves no even centre past a {halo} halo"
)));
}
let core = n - 2 * HALO;
let core = n - 2 * halo;
// In unified coordinates the frame spans u ∈ [dy, dy + h), v ∈ [dx, dx + w).
let (uh, uw) = (h + dy, w + dx);
let (ty, tx) = (uh.div_ceil(core), uw.div_ceil(core));
let total = ty * tx;
let origins: Vec<(usize, usize)> = (0..ty)
.flat_map(|i| (0..tx).map(move |j| (i * core, j * core)))
.collect();
let threads = std::thread::available_parallelism().map_or(1, |n| n.get());
// One tile's mosaic and σ, gathered on every core: rows are independent.
let gather = |u0: usize, v0: usize| {
let mut mos = vec![0.0f32; n * n];
let mut sig = vec![0.0f32; n * n];
let rows_per = n.div_ceil(threads).max(1);
std::thread::scope(|scope| {
for (chunk, (m, s)) in mos
.chunks_mut(rows_per * n)
.zip(sig.chunks_mut(rows_per * n))
.enumerate()
{
scope.spawn(move || {
for (i, (mrow, srow)) in m.chunks_mut(n).zip(s.chunks_mut(n)).enumerate() {
let r = chunk * rows_per + i;
// Unified row u = u0 + r − halo; frame row y = u − dy, reflected.
let u = u0 as isize + r as isize - halo as isize;
let y = reflect(u - dy as isize, h);
for c in 0..n {
let v = v0 as isize + c as isize - halo as isize;
let x = reflect(v - dx as isize, w);
let val = at(y, x);
mrow[c] = val;
// RGGB colour of the tile position (r, c).
srow[c] = sigma([[0, 1], [1, 2]][r & 1][c & 1], val);
}
}
});
}
});
(mos, sig)
};
// Pipelined: the next tile is gathered while the network runs this one,
// so the device does not wait on the CPU. A channel of one keeps at
// most two tiles' inputs alive.
let mut out = vec![0.0f32; h * w * 3];
let mut mos = vec![0.0f32; n * n];
let mut sig = vec![0.0f32; n * n];
// RGGB colour of unified position (u, v).
let colour = |u: usize, v: usize| [[0, 1], [1, 2]][u & 1][v & 1];
for (k, (i, j)) in (0..ty)
.flat_map(|i| (0..tx).map(move |j| (i, j)))
.enumerate()
{
let (u0, v0) = (i * core, j * core);
for r in 0..n {
// Unified row u = u0 + r − HALO; frame row y = u − dy, reflected.
let u = u0 as isize + r as isize - HALO as isize;
let y = reflect(u - dy as isize, h);
for c in 0..n {
let v = v0 as isize + c as isize - HALO as isize;
let x = reflect(v - dx as isize, w);
let val = at(y, x);
mos[r * n + c] = val;
sig[r * n + c] = sigma(colour(r, c), val);
}
}
let rgb = net.run(&mos, &sig)?;
if rgb.len() != 3 * n * n {
return Err(crate::DenoiseError::Model(format!(
"network returned {} values for a {n}² tile",
rgb.len()
)));
}
for r in HALO..HALO + core {
let u = u0 + r - HALO;
if u < dy || u >= uh {
continue;
}
let y = u - dy;
for c in HALO..HALO + core {
let v = v0 + c - HALO;
if v < dx || v >= uw {
continue;
let stop = std::sync::atomic::AtomicBool::new(false);
std::thread::scope(|scope| -> Result<Option<()>, crate::DenoiseError> {
let (tx_tiles, rx_tiles) = std::sync::mpsc::sync_channel(1);
let (origins, stop, gather) = (&origins, &stop, &gather);
scope.spawn(move || {
for &(u0, v0) in origins {
if stop.load(std::sync::atomic::Ordering::Relaxed) {
break;
}
let x = v - dx;
let o = (y * w + x) * 3;
for ch in 0..3 {
out[o + ch] = rgb[ch * n * n + r * n + c];
if tx_tiles.send((u0, v0, gather(u0, v0))).is_err() {
break;
}
}
});
for k in 0..total {
let Ok((u0, v0, (mos, sig))) = rx_tiles.recv() else {
break;
};
let mut wrong = None;
let ran = net.run(mos, sig, &mut |rgb: &[f32]| {
if rgb.len() != 3 * n * n {
wrong = Some(rgb.len());
return;
}
// The tile's centre back into the frame: the frame rows it covers,
// split across cores (each row is written by one thread only).
let (y_lo, y_hi) = (
(u0 + dy.saturating_sub(u0)).max(dy) - dy,
(u0 + core).min(uh) - dy,
);
let (x_lo, x_hi) = ((v0.max(dx)) - dx, (v0 + core).min(uw) - dx);
if y_hi > y_lo && x_hi > x_lo {
let rows = &mut out[y_lo * w * 3..y_hi * w * 3];
let per = (y_hi - y_lo).div_ceil(threads).max(1);
std::thread::scope(|scope| {
for (chunk, block) in rows.chunks_mut(per * w * 3).enumerate() {
scope.spawn(move || {
for (i, row) in block.chunks_mut(w * 3).enumerate() {
let y = y_lo + chunk * per + i;
// Tile row of frame row y: u = y + dy = u0 + r − halo.
let r = y + dy + halo - u0;
for x in x_lo..x_hi {
let c = x + dx + halo - v0;
for ch in 0..3 {
row[x * 3 + ch] = rgb[ch * n * n + r * n + c];
}
}
}
});
}
});
}
});
if let Err(e) = ran {
stop.store(true, std::sync::atomic::Ordering::Relaxed);
return Err(e);
}
if let Some(len) = wrong {
stop.store(true, std::sync::atomic::Ordering::Relaxed);
return Err(crate::DenoiseError::Model(format!(
"network returned {len} values for a {n}² tile"
)));
}
if !progress(k + 1, total) {
stop.store(true, std::sync::atomic::Ordering::Relaxed);
// Drain so the producer is not left blocked on a full channel.
while rx_tiles.try_recv().is_ok() {}
return Ok(None);
}
}
if !progress(k + 1, total) {
return Ok(None);
}
}
Ok(Some(out))
Ok(Some(()))
})
.map(|done| done.map(|()| out))
}
#[cfg(test)]
@@ -165,7 +244,12 @@ mod tests {
fn tile(&self) -> usize {
self.n
}
fn run(&mut self, m: &[f32], _s: &[f32]) -> Result<Vec<f32>, crate::DenoiseError> {
fn run(
&mut self,
m: Vec<f32>,
_s: Vec<f32>,
write: &mut dyn FnMut(&[f32]),
) -> Result<(), crate::DenoiseError> {
let n = self.n;
let q = n / 2;
let quad = |qy: usize, qx: usize| {
@@ -197,7 +281,8 @@ mod tests {
}
}
}
Ok(out)
write(&out);
Ok(())
}
}
@@ -293,3 +378,65 @@ mod tests {
assert!(r.is_none());
}
}
#[cfg(test)]
mod timing {
use super::*;
/// A network that answers instantly with an output of the right size,
/// so what is timed is the tiler alone: gathering each tile's mosaic and
/// σ, and writing its centre back.
struct Null(usize, Vec<f32>);
impl TileNet for Null {
fn tile(&self) -> usize {
self.0
}
fn run(
&mut self,
m: Vec<f32>,
_s: Vec<f32>,
write: &mut dyn FnMut(&[f32]),
) -> Result<(), crate::DenoiseError> {
// Stands for the runtime's own output buffer: allocated once.
if self.1.len() != 3 * m.len() {
self.1 = vec![m[0]; 3 * m.len()];
}
write(&self.1);
Ok(())
}
}
/// `cargo test --release -p dr-denoise tiler_overhead -- --ignored --nocapture`
#[test]
#[ignore]
fn tiler_overhead_on_a_6d_frame() {
let (h, w) = (3648, 5472);
let frame: Vec<f32> = (0..h * w).map(|i| (i % 977) as f32 / 977.0).collect();
let at = |y: usize, x: usize| frame[y * w + x];
let sigma = |_c: usize, v: f32| (0.001 * v + 1e-5).sqrt();
for n in [1408usize, 2048] {
let mut net = Null(n, Vec::new());
let t = std::time::Instant::now();
let mut tiles = 0;
run_tiled(
&mut net,
h,
w,
CfaPattern::Rggb,
&at,
&sigma,
&mut |_, total| {
tiles = total;
true
},
)
.unwrap();
let s = t.elapsed().as_secs_f64();
println!(
"tile {n}: {tiles} tiles, tiler alone {s:.2} s ({:.0} ms a tile)",
s / tiles as f64 * 1e3
);
}
}
}
+3 -3
View File
@@ -137,8 +137,8 @@ impl Detection {
/// A loaded SCRFD graph.
pub struct Detector {
session: Model,
/// f32 or int8 — the int8 form finds a different set of faces and is a
/// different detector in `model_id` (docs/dev/inference.md §7).
/// f32 or a quantised form — which finds a different set of faces and is
/// a different detector in `model_id` (docs/dev/inference.md §7).
form: Form,
/// Feature-map count: 3 for strides {8,16,32}, 4 for {8,16,32,64}.
///
@@ -155,7 +155,7 @@ impl Detector {
}
/// Load the canonical f32 file at `path`, or the form the device's
/// backend wants instead — the `.int8.onnx` beside it on a Hexagon —
/// backend wants instead — the `.a16w8.onnx` beside it on a Hexagon —
/// which [`Detector::form`] then reports.
pub fn from_path(path: impl AsRef<std::path::Path>) -> Result<Self, FaceError> {
let (path, form) = dr_inference_engine::resolve_model(Role::Detector, path.as_ref());
+11 -2
View File
@@ -123,13 +123,22 @@ pub struct Landmarker {
}
impl Landmarker {
/// The graph at `path`, or the `.a16w8.onnx` sibling beside it when the
/// device's backend runs that (the Hexagon, inference.md §1.5: 0.25 px
/// from f32 in the 192 crop, where int8 moved the points by 1.5).
pub fn from_path(path: impl AsRef<std::path::Path>) -> Result<Self, FaceError> {
let (path, form) = dr_inference_engine::resolve_model(Role::Landmarks, path.as_ref());
let bytes = std::fs::read(path).map_err(FaceError::ModelRead)?;
Self::from_bytes(&bytes)
Self::from_bytes_in(&bytes, form)
}
pub fn from_bytes(bytes: &[u8]) -> Result<Self, FaceError> {
let model = dr_inference_engine::open(Role::Landmarks, Form::F32, bytes)?;
Self::from_bytes_in(bytes, Form::F32)
}
/// `bytes` in a stated numeric form; the output keeps its meaning.
pub fn from_bytes_in(bytes: &[u8], form: Form) -> Result<Self, FaceError> {
let model = dr_inference_engine::open(Role::Landmarks, form, bytes)?;
let acquired = model.acquire()?;
let session = acquired.lock();
+117
View File
@@ -0,0 +1,117 @@
//! List each RAW file's hot and dead photosite candidates (docs/dev/sensor-health.md).
//!
//! Reads paths on stdin and prints one JSON line per file: its capture
//! conditions and every photosite [`Demosaicer::find_hot_pixels`] flags, as
//! `[x, y, value, hot]` in sensor coordinates. Which candidates are defects
//! is a question across frames, so this answers nothing on its own.
//!
//! ```sh
//! find ~/Pictures -name '*.CR2' | cargo run --release -p dr-gpu --example sensor_scan
//! ```
use std::io::{BufRead, Write};
use dr_gpu::{Demosaicer, GpuContext};
fn main() {
if let Some(list) = std::env::args().skip_while(|a| a != "--probe").nth(1) {
return probe(&list);
}
let ctx = pollster::block_on(GpuContext::new_headless()).expect("a GPU for the hot-pixel pass");
let demosaicer = Demosaicer::new(&ctx).expect("demosaicer");
for line in std::io::stdin().lock().lines() {
let path = line.expect("stdin");
match scan(&path, &demosaicer) {
Ok(json) => println!("{json}"),
Err(e) => eprintln!("fail\t{path}\t{e}"),
}
std::io::stdout().flush().ok();
}
}
fn scan(path: &str, demosaicer: &Demosaicer) -> Result<String, String> {
let bytes = std::fs::read(path).map_err(|e| e.to_string())?;
let raw = dr_decode::decode(&bytes).map_err(|e| e.to_string())?;
let meta = dr_decode::metadata(&bytes).map_err(|e| e.to_string())?;
let sites = demosaicer
.find_hot_pixels(&raw)
.map_err(|e| e.to_string())?;
let opt = |v: Option<f64>| v.map_or("null".to_string(), |v| v.to_string());
let list: Vec<String> = sites
.iter()
.map(|s| {
let v = raw.data[(s.y * raw.width + s.x) as usize];
format!("[{},{},{},{}]", s.x, s.y, v, u8::from(s.hot))
})
.collect();
Ok(format!(
"{{\"path\":{:?},\"model\":{:?},\"captured\":{},\"iso\":{},\"shutter\":{},\"white\":{},\"black\":{:?},\"crop\":[{},{},{},{}],\"sites\":[{}]}}",
path,
format!("{} {}", raw.make, raw.model),
meta.captured_at.map_or("null".to_string(), |t| t.to_string()),
opt(meta.iso.map(f64::from)),
opt(meta.shutter.map(f64::from)),
raw.white_level,
raw.black_level,
raw.crop.x,
raw.crop.y,
raw.crop.width,
raw.crop.height,
list.join(","),
))
}
/// `--probe COORDS`: for each path on stdin, each `x y` line of COORDS as
/// `[value, same-colour neighbour max, median]` over black, on the CPU. A
/// probe asks whether a photosite stood out in a frame where it would have
/// been visible, which the scan's verdict cannot say: a frame that does not
/// flag a defect may only have been too bright around it.
fn probe(list: &str) {
let coords: Vec<(u32, u32)> = std::fs::read_to_string(list)
.expect("coords")
.lines()
.filter_map(|l| {
let mut it = l.split_whitespace().map(|v| v.parse().ok());
Some((it.next()??, it.next()??))
})
.collect();
for line in std::io::stdin().lock().lines() {
let path = line.expect("stdin");
let Ok(bytes) = std::fs::read(&path) else {
continue;
};
let (Ok(raw), Ok(meta)) = (dr_decode::decode(&bytes), dr_decode::metadata(&bytes)) else {
continue;
};
let w = raw.width as i64;
let at = |x: i64, y: i64| {
let cell = (((y - raw.crop.y as i64) & 1) * 2 + ((x - raw.crop.x as i64) & 1)) as usize;
raw.data[(y * w + x) as usize].saturating_sub(raw.black_level[cell])
};
let rows: Vec<String> = coords
.iter()
.map(|&(x, y)| {
let (x, y) = (x as i64, y as i64);
let mut n: Vec<u16> = Vec::new();
for dy in [-2i64, 0, 2] {
for dx in [-2i64, 0, 2] {
if (dx, dy) != (0, 0) {
n.push(at(x + dx, y + dy));
}
}
}
n.sort_unstable();
format!("[{},{},{}]", at(x, y), n[n.len() - 1], n[n.len() / 2])
})
.collect();
println!(
"{{\"path\":{:?},\"captured\":{},\"iso\":{},\"shutter\":{},\"range\":{},\"p\":[{}]}}",
path,
meta.captured_at.unwrap_or(0),
meta.iso.unwrap_or(0),
meta.shutter.unwrap_or(0.0),
raw.white_level - raw.black_level[0],
rows.join(","),
);
}
}
+61 -9
View File
@@ -1071,6 +1071,46 @@ impl Demosaicer {
if raw.samples_per_pixel != 1 {
return Ok(0);
}
let words = self.hot_pixel_words(raw)?;
let mut changed = 0;
for (i, v) in raw.data.iter_mut().enumerate() {
let new = unpack_sample(&words, i);
changed += usize::from(new != *v);
*v = new;
}
Ok(changed)
}
/// The photosites [`Self::repair_hot_pixels`] would replace, in sensor
/// coordinates, without replacing them.
///
/// For the sensor health record (docs/dev/sensor-health.md): one frame's
/// verdict is a candidate list, not a defect map — a single photosite of a
/// star that passes both tests reads the same as a hot one. Which of them
/// is the sensor is decided across frames, by who keeps coming back.
pub fn find_hot_pixels(&self, raw: &RawImage) -> Result<Vec<Photosite>, GpuError> {
if raw.samples_per_pixel != 1 {
return Ok(Vec::new());
}
let words = self.hot_pixel_words(raw)?;
let stride = raw.width.max(1);
Ok(raw
.data
.iter()
.enumerate()
.filter_map(|(i, &v)| {
let new = unpack_sample(&words, i);
(new != v).then(|| Photosite {
x: i as u32 % stride,
y: i as u32 / stride,
hot: new < v,
})
})
.collect())
}
/// The hot-pixel pass over `raw`, read back as packed words.
fn hot_pixel_words(&self, raw: &RawImage) -> Result<Vec<u32>, GpuError> {
let (width, height) = (raw.crop.width.max(1), raw.crop.height.max(1));
let xtrans_tile = raw
.cfa_pattern
@@ -1122,18 +1162,30 @@ impl Demosaicer {
.map_err(|e| GpuError::Readback(e.to_string()))?;
let words: Vec<u32> = bytemuck::cast_slice(&slice.get_mapped_range()).to_vec();
readback.unmap();
let mut changed = 0;
for (i, v) in raw.data.iter_mut().enumerate() {
let w = words[i / 2];
let new = if i % 2 == 0 { w & 0xFFFF } else { w >> 16 } as u16;
changed += usize::from(new != *v);
*v = new;
}
Ok(changed)
Ok(words)
}
}
/// One photosite the hot-pixel pass judged defective.
#[derive(Copy, Clone, Debug, PartialEq, Eq, Hash)]
pub struct Photosite {
/// Sensor coordinates: the full readout, masked border included.
pub x: u32,
pub y: u32,
/// Read far above its neighbourhood; otherwise far below (dead).
pub hot: bool,
}
/// Sample `i` of a readout packed by [`pack_samples`].
fn unpack_sample(words: &[u32], i: usize) -> u16 {
let w = words[i / 2];
(if i.is_multiple_of(2) {
w & 0xFFFF
} else {
w >> 16
}) as u16
}
const IDENTITY_3X3: [f32; 9] = [1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0];
/// Pack u16 samples two per u32, little-endian within the word.
+1 -1
View File
@@ -38,7 +38,7 @@ pub use adjust::AdjustPass;
// rather than an implementation detail: a detail pass is guaranteed linear,
// unclipped, full internal precision (FR-DEV-2), and anyone reasoning about
// VRAM at 24 MP needs to know what an intermediate costs.
pub use demosaic::{DemosaicedImage, Demosaicer};
pub use demosaic::{DemosaicedImage, Demosaicer, Photosite};
pub use detail::INTERMEDIATE_FORMAT as DETAIL_INTERMEDIATE_FORMAT;
pub use error::GpuError;
pub use focus::{FocusPeakPass, FocusPeaking, PeakColour, PeakSensitivity};
+16 -4
View File
@@ -18,7 +18,7 @@ use dr_decode::{CfaPattern, CropRect, RawImage};
use dr_gpu::{AdjustPass, Demosaicer, GpuContext};
use dr_pipeline::descriptor::{Attribute, LocalizedKey, OpDescriptor, OpId, ParamId};
use dr_pipeline::operation::{Operation, Stage, Uniform};
use dr_pipeline::ops::camera_profile::{apply_reference, CameraProfile, APPLY, LOOK};
use dr_pipeline::ops::camera_profile::{apply_reference, CameraProfile, APPLY, LOOK, PROFILE_LOOK};
use dr_types::{HueSatTable, ProfileOrigin, ProfileTables, Transfer};
const SIZE: u32 = 16;
@@ -151,6 +151,14 @@ fn render(ctx: &GpuContext, raw: &RawImage, op: CameraProfile) -> Vec<[u8; 3]> {
pixels.chunks_exact(4).map(|p| [p[0], p[1], p[2]]).collect()
}
/// The profile at the strength it states — the look table on, as the
/// reference applies it at 1.0. Not the default, which leaves it off (D21).
fn as_stated() -> CameraProfile {
let mut op = CameraProfile::new();
op.set_param(LOOK, PROFILE_LOOK);
op
}
fn encode(c: [f32; 3]) -> [i32; 3] {
c.map(|v| (Transfer::Srgb.encode(v.clamp(0.0, 1.0)) * 255.0).round() as i32)
}
@@ -183,8 +191,12 @@ fn the_shader_agrees_with_the_cpu_reference() {
return;
};
let tables = strong_tables();
let got = render(&ctx, &frame(Some(tables.clone())), CameraProfile::new());
assert_agrees(&got, |c| apply_reference(&tables, c, 1.0), "at defaults");
let got = render(&ctx, &frame(Some(tables.clone())), as_stated());
assert_agrees(
&got,
|c| apply_reference(&tables, c, 1.0),
"as the profile states it",
);
let mut doubled = CameraProfile::new();
doubled.set_param(LOOK, 200.0);
@@ -279,7 +291,7 @@ fn the_libraries_adobe_standard_renders_as_the_reference_does() {
let tables = dr_decode::dcp::embedded_in(&bytes)
.expect("Adobe Standard")
.tables(5000.0, ProfileOrigin::Embedded);
let got = render(&ctx, &frame(Some(tables.clone())), CameraProfile::new());
let got = render(&ctx, &frame(Some(tables.clone())), as_stated());
for (i, (c, g)) in colours().into_iter().zip(&got).enumerate() {
let want = encode(apply_reference(&tables, c, 1.0));
let g = g.map(i32::from);
+41 -1
View File
@@ -7,7 +7,7 @@
//! see of a defect it missed is the coloured cross the demosaic makes of it.
use dr_decode::{CfaPattern, CropRect, RawImage};
use dr_gpu::{AdjustPass, Demosaicer, GpuContext};
use dr_gpu::{AdjustPass, Demosaicer, GpuContext, Photosite};
use dr_pipeline::EditGraph;
const SIZE: u32 = 36;
@@ -171,3 +171,43 @@ fn the_repaired_mosaic_reads_back_with_only_the_defect_changed() {
let others = (0..raw.data.len()).filter(|&i| i != at);
assert!(others.into_iter().all(|i| raw.data[i] == before.data[i]));
}
/// Finding without repairing (docs/dev/sensor-health.md): the same verdict as
/// the repair, as sensor coordinates, with the frame left as it was. The
/// sensor health record builds on this, so it must name exactly the
/// photosites the repair would change — the hot one and the dead one, and
/// not the star.
#[test]
fn finding_names_what_the_repair_would_change_and_changes_nothing() {
let Some(ctx) = ctx() else {
eprintln!("no GPU adapter; skipping");
return;
};
let d = Demosaicer::new(&ctx).expect("demosaicer");
let mut set = vec![(MIDDLE, MIDDLE, WHITE), (9, 25, 0)];
for dy in 0..3 {
for dx in 0..3 {
set.push((4 + dx, 4 + dy, WHITE));
}
}
let raw = frame(CfaPattern::Rggb, 1600, &set);
let mut found = d.find_hot_pixels(&raw).expect("find");
found.sort_by_key(|p| (p.y, p.x));
assert_eq!(
found,
vec![
Photosite {
x: MIDDLE,
y: MIDDLE,
hot: true
},
Photosite {
x: 9,
y: 25,
hot: false
},
]
);
let mut repaired = raw.clone();
assert_eq!(d.repair_hot_pixels(&mut repaired).expect("repair"), 2);
}
+39 -19
View File
@@ -4,12 +4,14 @@
//!
//! DARKROOM_ORT_DIR=/usr/lib \
//! cargo run --release -p dr-inference-engine --features native,tract \
//! --example ladder -- CACHE_DIR models/face/scrfd_500m_640.onnx [MODEL.onnx ...]
//! --example ladder -- CACHE_DIR models/face/scrfd_500m_640.onnx [ROLE=MODEL.onnx ...]
//!
//! Every model named is a `Detector` for the config's purposes, which is
//! enough to see the rung taken, the engines compiled and a session land
//! on it. Delete `CACHE_DIR` to see the first run again; keep it to see the
//! second.
//! A bare path is a `Detector`; `denoiser=…`, `scene=…`, `inpainter=…`,
//! `landmarks=…` (any `Role`, lower case) says otherwise, so a device can
//! show each role taking its own form (inference.md §1.5). Each is opened
//! through `resolve_model`, as the app opens it, and the line says which
//! form and which rung it landed on. Delete `CACHE_DIR` to see the first
//! run again; keep it to see the second.
use std::path::PathBuf;
use std::time::{Duration, Instant};
@@ -17,7 +19,10 @@ use std::time::{Duration, Instant};
fn main() {
env_logger::Builder::from_env(env_logger::Env::default().default_filter_or("info")).init();
let mut args = std::env::args_os().skip(1).map(PathBuf::from);
let (Some(cache_dir), models) = (args.next(), args.collect::<Vec<_>>()) else {
let (Some(cache_dir), models) = (
args.next(),
args.map(|a| role_and_path(&a)).collect::<Vec<_>>(),
) else {
eprintln!("usage: ladder CACHE_DIR MODEL.onnx [MODEL.onnx ...]");
std::process::exit(2);
};
@@ -34,10 +39,7 @@ fn main() {
dr_inference_engine::init(dr_inference_engine::Config {
runtime_dirs,
cache_dir: cache_dir.clone(),
models: models
.iter()
.map(|p| (dr_inference_engine::Role::Detector, p.clone()))
.collect(),
models: models.clone(),
embedded: Vec::new(),
ceiling: None,
threads: 0,
@@ -80,21 +82,39 @@ fn main() {
std::thread::sleep(Duration::from_millis(500));
}
for path in &models {
let bytes = std::fs::read(path).expect("read model");
for (role, path) in &models {
let (path, form) = dr_inference_engine::resolve_model(*role, path);
let bytes = std::fs::read(&path).expect("read model");
let t = Instant::now();
let model = dr_inference_engine::open(
dr_inference_engine::Role::Detector,
dr_inference_engine::Form::F32,
&bytes,
)
.expect("open model");
let model = dr_inference_engine::open(*role, form, &bytes).expect("open model");
let acquired = model.acquire().expect("acquire session");
println!(
"{} on {} in {:.2} s",
"{role:?}: {} ({form:?}) on {} in {:.2} s",
path.file_name().unwrap().to_string_lossy(),
acquired.rung().label(),
t.elapsed().as_secs_f64()
);
}
}
/// `denoiser=path` → (Denoiser, path); a bare path is a detector.
fn role_and_path(arg: &std::path::Path) -> (dr_inference_engine::Role, PathBuf) {
use dr_inference_engine::Role::*;
let s = arg.to_string_lossy();
let Some((name, path)) = s.split_once('=') else {
return (Detector, arg.to_path_buf());
};
let role = match name {
"detector" => Detector,
"embedder" => Embedder,
"segmenter" => Segmenter,
"scene" => Scene,
"landmarks" => Landmarks,
"eyes" => EyeClassifier,
"keypoints" => Keypoints,
"inpainter" => Inpainter,
"denoiser" => Denoiser,
other => panic!("no role {other:?}"),
};
(role, PathBuf::from(path))
}
+5 -5
View File
@@ -8,7 +8,7 @@
use std::path::PathBuf;
use crate::{state, Config, Form, Rung};
use crate::{state, Config, Rung};
enum Source {
File(PathBuf),
@@ -82,10 +82,10 @@ pub fn run() {
(*role, Source::File(path), size)
})
})
.chain(cfg.embedded.iter().filter_map(|(role, bytes)| {
// An embedded model has no int8 sibling to offer a rung that
// wants one; it runs on that rung's fallback.
(rung.serves(*role) && rung.form(*role) == Form::F32).then_some((
.chain(cfg.embedded.iter().filter_map(|(role, form, bytes)| {
// The embedded form the rung wants, if the build carries it;
// a build without it runs that model on the rung's fallback.
(rung.serves(*role) && rung.form(*role) == *form).then_some((
*role,
Source::Bytes(bytes),
bytes.len() as u64,
+145 -59
View File
@@ -42,25 +42,46 @@ pub enum Role {
/// XFeat, the panorama keypoint detector (docs/dev/panorama.md).
Keypoints,
/// MI-GAN, the panorama border filler (docs/dev/panorama.md §12). Plain
/// convolutions, so any rung serves it; fp16 on TensorRT and int8 on
/// the Hexagon are the point of it.
/// convolutions, so any rung serves it; fp16 on TensorRT and 16-bit
/// activations on the Hexagon (int8 changes the fill, §1.5).
Inpainter,
/// The learned demosaic and denoise on the raw mosaic (docs/dev/denoise.md).
/// fp16 costs it nothing measurable; int8 costs 6–9 dB, because 256
/// levels cannot hold the shadow steps it exists to recover — so the
/// Hexagon does not take it.
/// Hexagon takes it with 16-bit activations and weights (§1.5).
Denoiser,
}
/// Which numeric form of a model a session was built from.
///
/// `Int8` is a different network from `F32` for a detector — it finds a
/// different set of faces — which is why [`form_suffix`] exists and why a
/// The quantised forms are QDQ graphs, per-channel weights, as QNN's HTP
/// takes them (docs/dev/inference.md §1.5): `Int8` is 8-bit activations and
/// weights, `A16W8` 16-bit activations with 8-bit weights, `A16W16` 16-bit
/// both. The Hexagon accepts no float tensor at all, so these are the
/// whole menu; which one a role gets is [`Rung::form`], measured per model.
///
/// A quantised detector is a different network from the f32 one — it finds
/// a different set of faces — which is why [`form_suffix`] exists and why a
/// caller appends it to `model_id`.
#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash, Serialize, Deserialize)]
pub enum Form {
F32,
Int8,
A16W8,
A16W16,
}
impl Form {
/// The infix of the sibling file that holds this form:
/// `scrfd_500m_640.a16w8.onnx` beside `scrfd_500m_640.onnx`.
pub fn file_tag(self) -> Option<&'static str> {
match self {
Form::F32 => None,
Form::Int8 => Some("int8"),
Form::A16W8 => Some("a16w8"),
Form::A16W16 => Some("a16w16"),
}
}
}
/// A rung of the ladder (§2). Ordered: a user override names the highest rung
@@ -81,7 +102,7 @@ pub enum Rung {
/// removed in ONNX Runtime 1.23, so there is no non-compiling AMD rung
/// to fall back to: this one falls back to the CPU.
MiGraphX,
/// Qualcomm's Hexagon NPU through QNN, int8 models only. Android only.
/// Qualcomm's Hexagon NPU through QNN, quantised models only. Android only.
Hexagon,
/// Apple, through CoreML: the Neural Engine, the GPU or the CPU, as
/// CoreML schedules it. macOS only. Compiles an ML Program per model on
@@ -121,23 +142,38 @@ impl Rung {
}
/// The model form this rung wants for a role.
fn form(self, _role: Role) -> Form {
///
/// On the Hexagon, the narrowest form that held each model's accuracy
/// on the tablet itself (§1.5): int8 lost 5% of the detector's faces at
/// 40–80 px, moved the landmarks by 1.5 px and the segmenter's scores
/// to nothing, and the denoiser by 6–9 dB, so those take 16-bit
/// activations; the segmenter, scene model, filler and denoiser also
/// needed 16-bit weights. Only XFeat keeps int8: its panorama alignment
/// moved by no more than f32's own refits do.
pub fn form(self, role: Role) -> Form {
match self {
Rung::Hexagon => Form::Int8,
Rung::Hexagon => match role {
Role::Keypoints => Form::Int8,
Role::Detector | Role::Landmarks => Form::A16W8,
Role::Segmenter | Role::Scene | Role::Inpainter | Role::Denoiser => Form::A16W16,
Role::Embedder | Role::EyeClassifier => Form::F32,
},
_ => Form::F32,
}
}
/// Whether this rung runs `role` at all. The Hexagon takes int8 graphs
/// only, and the embedder is never int8 (§7) — it runs on the CPU
/// beside a detector on the NPU, so its vectors compare across devices.
/// Nor is the denoiser: its int8 form failed the 0.5 dB gate by 6–9 dB
/// (denoise.md §8), so it runs on the CPU there too. CoreML is kept off
/// the embedder for the same reason as the Hexagon: the Neural Engine is
/// fp16, and which unit runs a graph is CoreML's choice.
/// Whether this rung runs `role` at all. The Hexagon takes quantised
/// graphs only, and the embedder is never quantised (§7) — it runs on
/// the CPU beside a detector on the NPU, so its vectors compare across
/// devices; at A16W16 it still missed the 0.999 cosine gate. The eye
/// classifiers stay on the CPU too: a millisecond there, and the two
/// share one role while only one of them held its readings quantised.
/// CoreML is kept off the embedder for the same reason as the Hexagon:
/// the Neural Engine is fp16, and which unit runs a graph is CoreML's
/// choice.
fn serves(self, role: Role) -> bool {
match self {
Rung::Hexagon => !matches!(role, Role::Embedder | Role::Denoiser),
Rung::Hexagon => !matches!(role, Role::Embedder | Role::EyeClassifier),
Rung::CoreMl => role != Role::Embedder,
_ => true,
}
@@ -161,8 +197,10 @@ pub struct Config {
/// The canonical model files on this device, so engines can be compiled
/// ahead of the first request for them.
pub models: Vec<(Role, PathBuf)>,
/// Models compiled into the binary, for the same reason.
pub embedded: Vec<(Role, &'static [u8])>,
/// Models compiled into the binary, for the same reason, each with the
/// form it is. A build that embeds a quantised sibling lists it here
/// beside the f32 graph, and the compile step takes the one the rung wants.
pub embedded: Vec<(Role, Form, &'static [u8])>,
/// The highest rung the user allows; `None` is "the best that works".
pub ceiling: Option<Rung>,
/// ONNX Runtime's intra-op pool; 0 picks from the core count.
@@ -188,10 +226,10 @@ pub struct Status {
}
impl Status {
/// "Hexagon NPU · int8 · ONNX Runtime 1.29" — the settings row's text.
/// "Hexagon NPU · quantised · ONNX Runtime 1.29" — the settings row's text.
pub fn line(&self) -> String {
let form = match self.rung {
Rung::Hexagon => " · int8",
Rung::Hexagon => " · quantised",
Rung::TensorRt | Rung::MiGraphX => " · fp16",
_ => "",
};
@@ -374,6 +412,12 @@ struct Cache {
/// still reads.
#[serde(default)]
refused: BTreeSet<String>,
/// Probes run under this fingerprint (`probe::run`): a fall-back to the
/// CPU is re-probed until there have been `RETRIES`. Defaulted, so a
/// cache from 0.22.0 or before — which may hold exactly such a verdict —
/// probes again.
#[serde(default)]
attempts: u32,
}
struct State {
@@ -464,26 +508,44 @@ fn current_rung(s: &State) -> Rung {
/// The file to load for `role` under the current selection, and its form.
///
/// A rung that wants int8 gets the `.int8.onnx` sibling of the canonical file
/// if it exists; otherwise the canonical file, on the rung's fallback. A
/// caller adds [`form_suffix`] to the `model_id` it records.
/// A rung that wants a quantised form gets that sibling of the canonical
/// file (`<stem>.a16w8.onnx` and so on, [`Form::file_tag`]) if it exists;
/// otherwise the canonical file, on the rung's fallback. A caller adds
/// [`form_suffix`] to the `model_id` it records.
pub fn resolve_model(role: Role, canonical: &Path) -> (PathBuf, Form) {
let rung = current_rung(&state().lock().unwrap());
if rung.serves(role) && rung.form(role) == Form::Int8 {
let sibling = int8_sibling(canonical);
let want = rung.form(role);
if rung.serves(role) && want != Form::F32 {
let sibling = form_sibling(canonical, want);
if sibling.is_file() {
return (sibling, Form::Int8);
return (sibling, want);
}
}
(canonical.to_path_buf(), Form::F32)
}
fn int8_sibling(canonical: &Path) -> PathBuf {
/// The same choice for a model compiled into the binary: of the forms
/// `offered`, the one the current rung wants for `role`, else the f32 one.
/// `offered` must hold an `F32` entry.
pub fn choose_embedded(role: Role, offered: &[(Form, &'static [u8])]) -> (&'static [u8], Form) {
let rung = current_rung(&state().lock().unwrap());
let want = rung.form(role);
let pick = |form| offered.iter().find(|(f, _)| *f == form);
let (form, bytes) = (rung.serves(role).then(|| pick(want)).flatten())
.or_else(|| pick(Form::F32))
.expect("an embedded model offers its f32 form");
(bytes, *form)
}
fn form_sibling(canonical: &Path, form: Form) -> PathBuf {
let stem = canonical
.file_stem()
.map(|s| s.to_string_lossy().into_owned())
.unwrap_or_default();
canonical.with_file_name(format!("{stem}.int8.onnx"))
match form.file_tag() {
Some(tag) => canonical.with_file_name(format!("{stem}.{tag}.onnx")),
None => canonical.to_path_buf(),
}
}
/// What a form appends to a detector's `model_id` (§7).
@@ -491,6 +553,8 @@ pub fn form_suffix(form: Form) -> &'static str {
match form {
Form::F32 => "",
Form::Int8 => "_i8",
Form::A16W8 => "_a16",
Form::A16W16 => "_a16w16",
}
}
@@ -520,8 +584,8 @@ pub fn open(role: Role, form: Form, bytes: &[u8]) -> Result<Model, Error> {
fn effective_rung(s: &State, selected: Rung, role: Role, form: Form, hash: u64) -> Rung {
let mut rung = selected;
if !rung.serves(role) || rung.form(role) != form {
// The embedder on a Hexagon device, or an f32 detector where the int8
// sibling was missing: neither can go to the NPU.
// The embedder on a Hexagon device, or an f32 detector where the
// quantised sibling was missing: neither can go to the NPU.
rung = rung.fallback();
}
if rung.compiles() && !s.cache.compiled.contains(&engines::key_of(rung, hash)) {
@@ -612,9 +676,10 @@ mod tests {
#[test]
fn the_hexagon_never_takes_the_embedder() {
assert!(!Rung::Hexagon.serves(Role::Embedder));
assert!(!Rung::Hexagon.serves(Role::Denoiser));
assert!(!Rung::Hexagon.serves(Role::EyeClassifier));
assert!(Rung::Hexagon.serves(Role::Detector));
assert_eq!(Rung::Hexagon.form(Role::Detector), Form::Int8);
assert!(Rung::Hexagon.serves(Role::Denoiser));
assert_eq!(Rung::Hexagon.form(Role::Detector), Form::A16W8);
// A detector offered in f32 on a Hexagon device lands on the CPU.
let s = State {
config: Config::default(),
@@ -625,36 +690,57 @@ mod tests {
probing: false,
wanted: 0,
};
let on = |role, form| effective_rung(&s, Rung::Hexagon, role, form, engines::hash(b""));
assert_eq!(on(Role::Embedder, Form::F32), Rung::Cpu);
assert_eq!(on(Role::Detector, Form::F32), Rung::Cpu);
// A form other than the one the role wants is not the NPU's either:
// an int8 detector left over from an older install stays off it.
assert_eq!(on(Role::Detector, Form::Int8), Rung::Cpu);
// The wanted form whose context is not compiled yet: also the CPU.
assert_eq!(on(Role::Detector, Form::A16W8), Rung::Cpu);
}
/// The form each role gets on the Hexagon is the one measured to hold
/// its accuracy there (§1.5); a change to this table is a change to
/// what the tablet computes, and must come with a measurement.
#[test]
fn each_role_has_its_measured_form_on_the_hexagon() {
use Form::*;
for (role, form) in [
(Role::Detector, A16W8),
(Role::Landmarks, A16W8),
(Role::Segmenter, A16W16),
(Role::Scene, A16W16),
(Role::Inpainter, A16W16),
(Role::Denoiser, A16W16),
(Role::Keypoints, Int8),
(Role::Embedder, F32),
(Role::EyeClassifier, F32),
] {
assert_eq!(Rung::Hexagon.form(role), form, "{role:?}");
}
for rung in [
Rung::Cpu,
Rung::Cuda,
Rung::TensorRt,
Rung::MiGraphX,
Rung::CoreMl,
] {
assert_eq!(rung.form(Role::Detector), F32);
}
}
#[test]
fn a_form_lives_in_its_tagged_sibling() {
let canonical = Path::new("/m/scrfd_500m_640.onnx");
assert_eq!(form_sibling(canonical, Form::F32), canonical);
assert_eq!(
effective_rung(
&s,
Rung::Hexagon,
Role::Embedder,
Form::F32,
engines::hash(b"")
),
Rung::Cpu
form_sibling(canonical, Form::A16W8),
Path::new("/m/scrfd_500m_640.a16w8.onnx")
);
assert_eq!(
effective_rung(
&s,
Rung::Hexagon,
Role::Detector,
Form::F32,
engines::hash(b"")
),
Rung::Cpu
);
// An int8 detector whose context is not compiled yet: also the CPU.
assert_eq!(
effective_rung(
&s,
Rung::Hexagon,
Role::Detector,
Form::Int8,
engines::hash(b"")
),
Rung::Cpu
form_sibling(canonical, Form::Int8),
Path::new("/m/scrfd_500m_640.int8.onnx")
);
}
+105 -29
View File
@@ -31,6 +31,38 @@ fn ladder(ceiling: Option<Rung>) -> Vec<Rung> {
.collect()
}
/// Probes under one fingerprint that may end on the CPU after an
/// accelerator failed or lost, before that answer is kept.
const RETRIES: u32 = 3;
/// What a cached probe result is good for.
#[derive(Debug, PartialEq)]
enum Reuse {
/// Use it as it is.
Keep,
/// Probe again: it fell back to the CPU after this many probes.
Again(u32),
/// Another device, runtime or model set: probe from the start.
Fresh,
}
/// The CPU because an accelerator failed or lost is asked again on the next
/// launches, a few times: a failure can be a moment's (QNN could not create
/// its device on 0.22.0's first launch after the update), and keeping it for
/// good left the tablet's every model on the CPU. Bounded, so a wedged
/// driver costs a few launches, not all.
fn reuse(cached: &Cache, fingerprint: &str) -> Reuse {
if cached.fingerprint != fingerprint || cached.rung.is_none() {
return Reuse::Fresh;
}
let fell_back = cached.rung == Some(Rung::Cpu) && !cached.failed.is_empty();
if fell_back && cached.attempts < RETRIES {
Reuse::Again(cached.attempts)
} else {
Reuse::Keep
}
}
/// The probe body. Sets the cache and clears `probing` when done; never
/// panics out, because a failed probe is a result (the floor) and not an
/// error.
@@ -38,20 +70,33 @@ pub fn run(runtime: Runtime) {
let cfg = state().lock().unwrap().config.clone();
let fingerprint = fingerprint(&runtime, &cfg);
let mut attempts = 0;
if let Some(cached) = read_cache(&cfg) {
if cached.fingerprint == fingerprint && cached.rung.is_some() {
log::info!(
"inference: cached selection {} ({})",
cached.rung.unwrap().label(),
cached.reason
);
finish(cached);
return;
match reuse(&cached, &fingerprint) {
Reuse::Keep => {
log::info!(
"inference: cached selection {} ({})",
cached.rung.map_or("?", |r| r.label()),
cached.reason
);
finish(cached);
return;
}
Reuse::Again(n) => {
attempts = n;
log::info!(
"inference: probing again after falling back to the CPU ({}), attempt {} of {RETRIES}",
cached.reason,
n + 1
);
}
Reuse::Fresh => {}
}
}
let mut cache = Cache {
fingerprint,
attempts: attempts + 1,
..Cache::default()
};
@@ -174,10 +219,9 @@ pub fn attempt<T>(cfg: &Config, what: &str, f: impl FnOnce() -> T) -> Result<T,
}
/// The smallest detector, or the smallest model of any role if there is
/// none. A ~2 MB detector is the cheapest real test of a provider, and the
/// detector is the role the int8 forms exist for — the eye classifiers are
/// smaller still, and a Hexagon probed with one would fail for want of a
/// form nobody ships.
/// none. A ~2 MB detector is the cheapest real test of a provider, and
/// every rung serves it — the eye classifiers are smaller still, but the
/// Hexagon does not take them, and a probe with one would fail it for that.
fn probe_model(cfg: &Config) -> Option<(Role, PathBuf)> {
let smallest = |want: Option<Role>| {
cfg.models
@@ -203,20 +247,18 @@ fn time_rung(
cfg: &Config,
) -> Result<(f64, Option<String>), String> {
let want = rung.form(role);
let path = match want {
Form::Int8 => {
let p = crate::int8_sibling(canonical);
if !p.is_file() {
return Err(format!("no int8 form of {}", canonical.display()));
}
p
}
Form::F32 => canonical.to_path_buf(),
};
let path = crate::form_sibling(canonical, want);
if want != Form::F32 && !path.is_file() {
return Err(format!(
"no {} form of {}",
want.file_tag().unwrap_or("f32"),
canonical.display()
));
}
let bytes = std::fs::read(&path).map_err(|e| e.to_string())?;
let started = Instant::now();
let mut session =
crate::session::build(rung, role, &bytes, cfg).map_err(|e| first_line(&e.to_string()))?;
let mut session = crate::session::build_probe(rung, role, &bytes, cfg)
.map_err(|e| first_line(&e.to_string()))?;
log::info!(
"inference: {} session built in {:.1} s",
rung.label(),
@@ -295,9 +337,9 @@ fn fingerprint(runtime: &Runtime, cfg: &Config) -> String {
},
device_identity(),
];
for (role, bytes) in &cfg.embedded {
for (role, form, bytes) in &cfg.embedded {
parts.push(format!(
"{role:?} embedded {:016x}",
"{role:?} embedded {form:?} {:016x}",
crate::engines::hash(bytes)
));
}
@@ -306,9 +348,13 @@ fn fingerprint(runtime: &Runtime, cfg: &Config) -> String {
.map(|b| crate::engines::hash(&b))
.unwrap_or(0);
parts.push(format!("{role:?} {hash:016x}"));
let int8 = crate::int8_sibling(path);
if let Ok(b) = std::fs::read(&int8) {
parts.push(format!("{role:?} int8 {:016x}", crate::engines::hash(&b)));
for form in [Form::Int8, Form::A16W8, Form::A16W16] {
if let Ok(b) = std::fs::read(crate::form_sibling(path, form)) {
parts.push(format!(
"{role:?} {form:?} {:016x}",
crate::engines::hash(&b)
));
}
}
}
parts.join("\n")
@@ -493,4 +539,34 @@ mod tests {
died_inside(&cfg, "probe TensorRT", 2);
assert_eq!(attempt(&cfg, "probe CUDA", || 7), Ok(7));
}
/// The tablet's cache after 0.22.0's first launch, as 0.22.0 wrote it:
/// no `attempts`, the Hexagon "rejected", the CPU selected.
const TABLET: &str = r#"{"fingerprint":"f","rung":"Cpu","reason":"Hexagon NPU 28.5 ms, slower than the CPU's 19.4 ms","compiled":[],"failed":[["Hexagon","28.5 ms, slower than the CPU's 19.4 ms"]]}"#;
#[test]
fn a_fall_back_to_the_cpu_is_probed_again_a_few_times() {
let mut cache: Cache = serde_json::from_str(TABLET).unwrap();
assert_eq!(cache.attempts, 0, "a 0.22.0 cache reads as never retried");
assert_eq!(reuse(&cache, "f"), Reuse::Again(0));
cache.attempts = RETRIES - 1;
assert_eq!(reuse(&cache, "f"), Reuse::Again(RETRIES - 1));
cache.attempts = RETRIES;
assert_eq!(reuse(&cache, "f"), Reuse::Keep, "then it is kept");
}
#[test]
fn an_accelerator_chosen_or_a_cpu_only_device_is_kept() {
let mut cache: Cache = serde_json::from_str(TABLET).unwrap();
cache.rung = Some(Rung::Hexagon);
assert_eq!(reuse(&cache, "f"), Reuse::Keep);
cache.rung = Some(Rung::Cpu);
cache.failed.clear();
assert_eq!(
reuse(&cache, "f"),
Reuse::Keep,
"nothing failed: the only rung"
);
assert_eq!(reuse(&cache, "other"), Reuse::Fresh);
}
}
+27
View File
@@ -13,6 +13,30 @@ use crate::{Config, Role, Rung};
/// its clock instead (§4): a provider that hands real work to the CPU is
/// slower than the CPU floor and rejected by the same measurement.
pub fn build(rung: Rung, role: Role, bytes: &[u8], cfg: &Config) -> ort::Result<Session> {
build_with(rung, role, bytes, cfg, false)
}
/// [`build`] for the probe: on the Hexagon, a session that cannot put the
/// whole graph on the NPU fails instead of running the rest on the CPU.
///
/// The probe times a rung by its session, and a QNN provider that could not
/// create its device still builds one — with every node on the CPU behind
/// it. 0.22.0's first launch on the tablet timed that (28.5 ms against the
/// CPU's own 19.4) and put every model on the CPU. Only the probe is strict:
/// some shipped graphs keep a few nodes on the CPU on purpose
/// (`tools/quantise-models.py`, `float_nodes`), and the probe's detector is
/// not one of them.
pub fn build_probe(rung: Rung, role: Role, bytes: &[u8], cfg: &Config) -> ort::Result<Session> {
build_with(rung, role, bytes, cfg, rung == Rung::Hexagon)
}
fn build_with(
rung: Rung,
role: Role,
bytes: &[u8],
cfg: &Config,
strict: bool,
) -> ort::Result<Session> {
// No optimisation level named. ONNX Runtime's default is already its
// fullest, and on tract any level but "disabled" means `into_optimized`,
// whose optimiser divides by zero inside yolo26n-seg (tract-data
@@ -22,6 +46,9 @@ pub fn build(rung: Rung, role: Role, bytes: &[u8], cfg: &Config) -> ort::Result<
if crate::api::runtime().is_native() {
b = with_runtime_log(b)?;
}
if strict {
b = b.with_config_entry("session.disable_cpu_ep_fallback", "1")?;
}
// A Hexagon session loads the compiled context when there is one and
// compiles it from the model when there is not; the engine thread is
// what makes the second case rare (§6).
+12 -1
View File
@@ -13,8 +13,13 @@ const MODELS: &[&str] = &[
"../../models/keypoints/xfeat-768.onnx",
];
const QUANTISED: &[&str] = &[
"../../models/keypoints/xfeat-1024.int8.onnx",
"../../models/keypoints/xfeat-768.int8.onnx",
];
fn main() {
for m in MODELS {
for m in MODELS.iter().chain(QUANTISED) {
println!("cargo:rerun-if-changed={m}");
}
println!("cargo:rerun-if-changed=build.rs");
@@ -26,6 +31,12 @@ fn main() {
for model in MODELS.iter().copied() {
check(model);
}
// The Hexagon's int8 forms ride only in an Android build.
if std::env::var("CARGO_CFG_TARGET_OS").as_deref() == Ok("android") {
for model in QUANTISED.iter().copied() {
check(model);
}
}
}
fn check(model: &str) {
+2 -1
View File
@@ -30,7 +30,8 @@ pub struct MiGan {
impl MiGan {
/// From the model file, in whichever form the engine's rung wants
/// (`resolve_model` picks an int8 sibling for the Hexagon).
/// (`resolve_model` picks the `.a16w16.onnx` sibling on the Hexagon:
/// int8 moved the fill 16 dB from f32's, 16-bit about 41).
pub fn from_path(path: &std::path::Path) -> Result<Self, PanoError> {
use dr_inference_engine::{resolve_model, Role};
let (path, form) = resolve_model(Role::Inpainter, path);
+37 -6
View File
@@ -36,18 +36,49 @@ pub struct XFeat {
pub options: DecodeOptions,
}
/// The bytes of both exports compiled into the binary, for whoever compiles
/// engines ahead of the first request (docs/dev/inference.md §6).
/// The Hexagon's forms (docs/dev/inference.md §1.5): int8, from the same
/// network spelled for the HTP (the unfold as SpaceToDepth, the bilinear
/// resizes as matrix products). Only Android has a Hexagon.
#[cfg(all(feature = "embedded-model", target_os = "android"))]
const EMBEDDED_LANDSCAPE_INT8: &[u8] =
include_bytes!("../../../models/keypoints/xfeat-1024.int8.onnx");
#[cfg(all(feature = "embedded-model", target_os = "android"))]
const EMBEDDED_PORTRAIT_INT8: &[u8] =
include_bytes!("../../../models/keypoints/xfeat-768.int8.onnx");
/// Every form of both exports compiled into the binary, landscape then
/// portrait, for whoever compiles engines ahead of the first request
/// (docs/dev/inference.md §6).
#[cfg(feature = "embedded-model")]
pub fn embedded_model_bytes() -> [&'static [u8]; 2] {
[EMBEDDED_LANDSCAPE, EMBEDDED_PORTRAIT]
pub fn embedded_models() -> [Vec<(dr_inference_engine::Form, &'static [u8])>; 2] {
use dr_inference_engine::Form;
#[allow(unused_mut)]
let mut forms = [
vec![(Form::F32, EMBEDDED_LANDSCAPE)],
vec![(Form::F32, EMBEDDED_PORTRAIT)],
];
#[cfg(target_os = "android")]
{
forms[0].push((Form::Int8, EMBEDDED_LANDSCAPE_INT8));
forms[1].push((Form::Int8, EMBEDDED_PORTRAIT_INT8));
}
forms
}
impl XFeat {
/// The weights compiled into the binary.
/// The weights compiled into the binary, in the form the device's
/// backend runs.
#[cfg(feature = "embedded-model")]
pub fn embedded() -> Result<Self, PanoError> {
Self::from_bytes(EMBEDDED_LANDSCAPE, EMBEDDED_PORTRAIT)
use dr_inference_engine::{choose_embedded, open, Role};
let [l, p] = embedded_models();
let (l, lf) = choose_embedded(Role::Keypoints, &l);
let (p, pf) = choose_embedded(Role::Keypoints, &p);
Ok(XFeat {
landscape: open(Role::Keypoints, lf, l)?,
portrait: open(Role::Keypoints, pf, p)?,
options: DecodeOptions::default(),
})
}
/// From the two exports on disk.
+1 -1
View File
@@ -25,7 +25,7 @@ vibrance.vibrance = 10
blacks_whites.blacks = -8
clarity.amount = 12
contrast.contrast = 18
vibrance.vibrance = 18
vibrance.vibrance = 36
[preset Recover the sky]
blacks_whites.whites = -10
+13 -13
View File
@@ -15,38 +15,38 @@ drpl 1
[preset Blue sky]
colour_mixer.azure_lum = -20
colour_mixer.azure_sat = 25
colour_mixer.azure_sat = 18
colour_mixer.blue_lum = -15
colour_mixer.blue_sat = 20
colour_mixer.blue_sat = 14
highlights_shadows.highlights = -15
[preset Deep blue sky]
colour_mixer.azure_hue = 10
colour_mixer.azure_lum = -30
colour_mixer.azure_sat = 35
colour_mixer.azure_sat = 22
colour_mixer.blue_lum = -25
colour_mixer.blue_sat = 30
colour_mixer.cyan_sat = 10
colour_mixer.blue_sat = 19
colour_mixer.cyan_sat = 6
highlights_shadows.highlights = -30
[preset Polariser]
colour_mixer.azure_hue = 10
colour_mixer.azure_lum = -35
colour_mixer.azure_sat = 40
colour_mixer.azure_sat = 29
colour_mixer.blue_lum = -30
colour_mixer.blue_sat = 35
colour_mixer.blue_sat = 26
colour_mixer.cyan_lum = -10
colour_mixer.cyan_sat = 15
colour_mixer.cyan_sat = 11
dehaze.amount = 20
highlights_shadows.highlights = -35
vibrance.vibrance = 10
vibrance.vibrance = 7
[preset Blue sky, golden land]
colour_mixer.azure_lum = -20
colour_mixer.azure_sat = 25
colour_mixer.azure_sat = 16
colour_mixer.blue_lum = -15
colour_mixer.blue_sat = 20
colour_mixer.orange_sat = 12
colour_mixer.blue_sat = 13
colour_mixer.orange_sat = 8
colour_mixer.yellow_hue = -10
colour_mixer.yellow_sat = 15
colour_mixer.yellow_sat = 10
highlights_shadows.highlights = -20
+28 -19
View File
@@ -9,18 +9,27 @@ drpl 1
#
# Each changes only what it names (FR-DEV-6), so a corrected exposure or
# white balance survives applying one.
#
# How much colour each adds is measured, not guessed: mean CIELAB chroma on
# raws rendered with the default (DNG reference) rendering, as a ratio to that
# rendering. For scale, the photographer's earlier exports of the same kind of
# raws sit at 1.14 with no look applied and 1.27 with their everyday look.
# Vivid 1.30 and Vivid warm 1.30 sit just above that; Vivid landscape 1.38;
# Vivid, strong 1.45; Vivid portrait 1.15, with its skin bands held down as
# written. Tuned by scaling each preset's colour values together, never its
# tone ones.
[preset Vivid]
contrast.contrast = 10
saturation.saturation = 8
vibrance.vibrance = 30
saturation.saturation = 11
vibrance.vibrance = 42
[preset Vivid, strong]
blacks_whites.blacks = -10
clarity.amount = 8
contrast.contrast = 18
saturation.saturation = 15
vibrance.vibrance = 45
saturation.saturation = 19
vibrance.vibrance = 57
# Foliage and sky: green and chartreuse for leaves and grass, azure and blue
# for sky and water, a little yellow for dry grass and stone. The skin bands
@@ -29,37 +38,37 @@ vibrance.vibrance = 45
[preset Vivid landscape]
clarity.amount = 10
colour_mixer.azure_lum = -10
colour_mixer.azure_sat = 20
colour_mixer.azure_sat = 25
colour_mixer.blue_lum = -10
colour_mixer.blue_sat = 15
colour_mixer.chartreuse_sat = 15
colour_mixer.green_sat = 20
colour_mixer.yellow_sat = 10
colour_mixer.blue_sat = 20
colour_mixer.chartreuse_sat = 20
colour_mixer.green_sat = 25
colour_mixer.yellow_sat = 13
contrast.contrast = 12
saturation.saturation = 5
vibrance.vibrance = 25
saturation.saturation = 7
vibrance.vibrance = 32
# Golden hour: oranges and yellows up and a warm cast laid over the
# highlights only, so shadows stay clean rather than muddy.
[preset Vivid warm]
colour_grading.highlight_hue = 45
colour_grading.highlight_strength = 12
colour_mixer.orange_sat = 15
colour_mixer.red_sat = 8
colour_mixer.yellow_sat = 15
colour_mixer.orange_sat = 17
colour_mixer.red_sat = 9
colour_mixer.yellow_sat = 17
contrast.contrast = 8
vibrance.vibrance = 25
vibrance.vibrance = 29
# People: everything around the subject gets richer while skin does not.
# Vibrance already protects skin; the orange and red bands are then held a
# little below where they started, because a face is the one colour every
# viewer knows the right value of.
[preset Vivid portrait]
colour_mixer.azure_sat = 10
colour_mixer.blue_sat = 12
colour_mixer.green_sat = 12
colour_mixer.azure_sat = 14
colour_mixer.blue_sat = 17
colour_mixer.green_sat = 17
colour_mixer.orange_sat = -10
colour_mixer.red_sat = -5
contrast.contrast = 6
saturation.saturation = -5
vibrance.vibrance = 25
vibrance.vibrance = 35
+135 -32
View File
@@ -176,12 +176,12 @@ pub struct EditGraph {
/// lens profile, so not in the state; it only decides whether the
/// switch below is offered.
denoise_available: bool,
/// Whether the learned denoise replaces the demosaic. An edit: published
/// as [`crate::learned_denoise`], captured, stored and undone with the
/// rest (FR-DEV-3c).
denoise_applied: bool,
/// How much of the removed noise's brightness to put back, 0–100.
denoise_grain: f32,
/// Which demosaic develops the photograph: a network, or the classical
/// one. An edit: published as [`crate::learned_denoise`], captured,
/// stored and undone with the rest (FR-DEV-3c).
denoise_method: crate::learned_denoise::Method,
/// How strongly to denoise, 0–100; what is not taken goes back as grain.
denoise_strength: f32,
}
/// TRACES: FR-DEV-3f
@@ -239,8 +239,8 @@ impl EditGraph {
lens_profile: None,
lens_profile_applied: true,
denoise_available: false,
denoise_applied: false,
denoise_grain: 0.0,
denoise_method: crate::learned_denoise::Method::DEFAULT,
denoise_strength: 100.0,
}
}
@@ -425,13 +425,19 @@ impl EditGraph {
/// photograph that cannot take it is harmless and does nothing, as a
/// lens switch with no profile does.
pub fn denoise_applied(&self) -> bool {
self.denoise_applied
self.denoise_method.learned()
}
/// TRACES: FR-DEV-3g
/// The grain to keep, 0–1.
/// Which demosaic is asked for.
pub fn denoise_method(&self) -> crate::learned_denoise::Method {
self.denoise_method
}
/// TRACES: FR-DEV-3g
/// The grain to keep, 0–1: what the strength does not take.
pub fn denoise_grain(&self) -> f32 {
self.denoise_grain / 100.0
(100.0 - self.denoise_strength) / 100.0
}
/// TRACES: FR-DEV-3
@@ -625,7 +631,7 @@ impl EditGraph {
OpCapability {
id: desc.id,
label: desc.label,
active: self.denoise_applied,
active: self.denoise_applied(),
params: desc
.params
.iter()
@@ -643,10 +649,12 @@ impl EditGraph {
}
});
switch
// The learned denoise first: it decides what every control below
// is applied to, so it heads the panel (docs/dev/denoise.md §7).
denoise
.into_iter()
.chain(switch)
.chain(warps)
.chain(denoise)
.chain(ops)
.chain(std::iter::once(framing))
.collect()
@@ -741,8 +749,8 @@ impl EditGraph {
// Derived from the file, like the profile above.
denoise_available: _,
// Edits, in the state through `capabilities` like the lens switch.
denoise_applied: _,
denoise_grain: _,
denoise_method: _,
denoise_strength: _,
masks,
film,
spots,
@@ -816,9 +824,25 @@ impl EditGraph {
pub fn set_param(&mut self, op: OpId, param: ParamId, value: f32) {
if op == crate::learned_denoise::ID {
match param {
p if p == crate::learned_denoise::APPLY => self.denoise_applied = value != 0.0,
p if p == crate::learned_denoise::METHOD => {
self.denoise_method = crate::learned_denoise::Method::from_index(value)
}
// 0.21 and 0.22's switch (see `APPLY`): off is the classical
// demosaic, on is a network — the one already chosen, if any.
p if p == crate::learned_denoise::APPLY => {
use crate::learned_denoise::Method;
if value == 0.0 {
self.denoise_method = Method::Bilinear;
} else if !self.denoise_method.learned() {
self.denoise_method = Method::DEFAULT;
}
}
p if p == crate::learned_denoise::STRENGTH => {
self.denoise_strength = value.clamp(0.0, 100.0)
}
// 0.21.0's grain, the strength's inverse (see `GRAIN`).
p if p == crate::learned_denoise::GRAIN => {
self.denoise_grain = value.clamp(0.0, 100.0)
self.denoise_strength = 100.0 - value.clamp(0.0, 100.0)
}
_ => log::warn!("unknown parameter {param} on {op}; ignoring"),
}
@@ -883,10 +907,12 @@ impl EditGraph {
pub fn param(&self, op: OpId, param: ParamId) -> Option<f32> {
if op == crate::learned_denoise::ID {
return match param {
p if p == crate::learned_denoise::METHOD => Some(self.denoise_method.index()),
p if p == crate::learned_denoise::APPLY => {
Some(if self.denoise_applied { 1.0 } else { 0.0 })
Some(if self.denoise_applied() { 1.0 } else { 0.0 })
}
p if p == crate::learned_denoise::GRAIN => Some(self.denoise_grain),
p if p == crate::learned_denoise::STRENGTH => Some(self.denoise_strength),
p if p == crate::learned_denoise::GRAIN => Some(100.0 - self.denoise_strength),
_ => None,
};
}
@@ -936,10 +962,10 @@ impl EditGraph {
// a reset does not change which lens took the photograph. What returns
// to default is the answer to whether to use it, which is on.
self.set_lens_profile_applied(true);
// The learned denoise returns to off; whether it is available is the
// file's and stays.
self.denoise_applied = false;
self.denoise_grain = 0.0;
// The learned denoise returns to its default network; whether it is
// available is the file's and stays.
self.denoise_method = crate::learned_denoise::Method::DEFAULT;
self.denoise_strength = 100.0;
}
/// Set the crop rectangle. Clamped to keep it inside the frame.
@@ -2098,28 +2124,105 @@ mod tests {
.into_iter()
.find(|c| c.id == learned_denoise::ID)
.expect("offered");
assert!(!cap.active, "off until asked for");
g.set_param(learned_denoise::ID, learned_denoise::APPLY, 1.0);
g.set_param(learned_denoise::ID, learned_denoise::GRAIN, 30.0);
assert!(g.denoise_applied());
assert!(cap.active, "on by default");
assert_eq!(g.denoise_method(), learned_denoise::Method::Best);
assert_eq!(g.denoise_grain(), 0.0, "at full strength");
assert_eq!(cap.id, g.capabilities()[0].id, "and first in the panel");
g.set_param(
learned_denoise::ID,
learned_denoise::METHOD,
learned_denoise::Method::Bilinear.index(),
);
g.set_param(learned_denoise::ID, learned_denoise::STRENGTH, 70.0);
assert!(!g.denoise_applied());
assert!((g.denoise_grain() - 0.3).abs() < 1e-6);
g.reset();
assert!(!g.denoise_applied());
assert!(g.denoise_applied(), "reset is back to on");
assert_eq!(g.denoise_grain(), 0.0);
}
#[test]
fn an_untouched_raw_writes_nothing_and_develops_through_the_best() {
// TRACES: FR-DEV-3g
use crate::learned_denoise::{self, Method};
let mut g = EditGraph::default_chain();
g.set_denoise_available(true);
assert_eq!(g.denoise_method(), Method::Best);
let stored = |g: &EditGraph| {
crate::Preset::capture_params(g)
.params()
.keys()
.any(|(op, _)| op == learned_denoise::ID.0)
};
assert!(!stored(&g), "the default is not written");
g.set_param(
learned_denoise::ID,
learned_denoise::METHOD,
Method::Fast.index(),
);
assert!(stored(&g), "a choice is");
assert!(
!crate::Preset::capture_params(&g).params().contains_key(&(
learned_denoise::ID.0.into(),
learned_denoise::APPLY.0.into()
)),
"and the old switch never is"
);
}
#[test]
fn an_edit_saved_with_the_switch_keeps_its_look() {
// TRACES: FR-DEV-3g
// 0.21 and 0.22 stored on or off; off is the classical demosaic, and
// on keeps a network already chosen.
use crate::learned_denoise::{self, Method};
let mut g = EditGraph::default_chain();
g.set_param(learned_denoise::ID, learned_denoise::APPLY, 0.0);
assert_eq!(g.denoise_method(), Method::Bilinear);
g.set_param(learned_denoise::ID, learned_denoise::APPLY, 1.0);
assert_eq!(g.denoise_method(), Method::DEFAULT);
g.set_param(
learned_denoise::ID,
learned_denoise::METHOD,
Method::Fast.index(),
);
g.set_param(learned_denoise::ID, learned_denoise::APPLY, 1.0);
assert_eq!(g.denoise_method(), Method::Fast);
// A number from a newer build with more methods is the default.
g.set_param(learned_denoise::ID, learned_denoise::METHOD, 9.0);
assert_eq!(g.denoise_method(), Method::DEFAULT);
}
#[test]
fn an_edit_saved_with_grain_keeps_its_look() {
// TRACES: FR-DEV-3g
// 0.21.0 stored the grain kept rather than the strength.
use crate::learned_denoise;
let mut g = EditGraph::default_chain();
g.set_param(learned_denoise::ID, learned_denoise::GRAIN, 25.0);
assert_eq!(
g.param(learned_denoise::ID, learned_denoise::STRENGTH),
Some(75.0)
);
assert!((g.denoise_grain() - 0.25).abs() < 1e-6);
}
#[test]
fn the_learned_denoise_travels_in_the_state() {
use crate::learned_denoise;
let mut g = EditGraph::default_chain();
g.set_denoise_available(true);
g.set_param(learned_denoise::ID, learned_denoise::APPLY, 1.0);
g.set_param(learned_denoise::ID, learned_denoise::GRAIN, 40.0);
g.set_param(learned_denoise::ID, learned_denoise::STRENGTH, 60.0);
g.set_param(
learned_denoise::ID,
learned_denoise::METHOD,
learned_denoise::Method::Medium.index(),
);
let state = g.state();
let mut h = EditGraph::default_chain();
h.set_denoise_available(true);
let _ = h.set_state(&state);
assert!(h.denoise_applied());
assert_eq!(h.denoise_method(), learned_denoise::Method::Medium);
assert!((h.denoise_grain() - 0.4).abs() < 1e-6);
}
}
+77 -9
View File
@@ -1,6 +1,6 @@
//! TRACES: FR-DEV-3g
//! The learned denoise's settings: whether to use it, and how much grain to
//! keep.
//! The learned denoise's settings: which network develops the photograph,
//! if any, and how much grain to keep.
//!
//! Not an [`crate::operation::Operation`]: the learned stage replaces the
//! demosaic and runs once per photograph, off the render path
@@ -17,24 +17,92 @@ use crate::descriptor::{Attribute, LocalizedKey, OpDescriptor, ParamDescriptor,
use crate::{OpId, ParamId};
pub const ID: OpId = OpId("learned_denoise");
/// TRACES: FR-DEV-3g
/// Which demosaic develops the photograph, a [`Method`] by index.
pub const METHOD: ParamId = ParamId("method");
/// What 0.21 and 0.22 stored instead of [`METHOD`]: on or off. Still read —
/// off is [`Method::Bilinear`], on is the default network — so an edit saved
/// by those releases keeps its look; never written, and not offered.
pub const APPLY: ParamId = ParamId("apply");
/// TRACES: FR-DEV-3g
/// How strongly to denoise, 0–100: 100 is the network's result as it is, and
/// lower puts the removed noise's brightness back as grain.
pub const STRENGTH: ParamId = ParamId("strength");
/// What 0.21.0 stored instead of [`STRENGTH`]: the grain kept, its inverse.
/// Still read, so an edit saved by that release keeps its look; never
/// written, and not offered as a control.
pub const GRAIN: ParamId = ParamId("grain");
/// Off by default: it costs seconds per photograph and replaces the
/// demosaic, which is the photographer's call. Grain 0 is the network's
/// result as it is.
/// TRACES: FR-DEV-3g
/// The demosaics a photograph can be developed with, in the order the
/// sidecar numbers them. Three networks that trade time for quality — the
/// same training, distilled into smaller students (docs/dev/denoise.md §13)
/// — and the classical demosaic, which is no network at all.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)]
pub enum Method {
/// The classical demosaic: the noise stays.
Bilinear,
/// The smallest student: a quarter of the medium network's work.
Fast,
/// One network the size of the first release's.
Medium,
/// Two experts, one for flat areas and one for edges, and a gate.
Best,
}
impl Method {
pub const ALL: [Method; 4] = [Method::Bilinear, Method::Fast, Method::Medium, Method::Best];
pub const DEFAULT: Method = Method::Best;
/// The sidecar's number for it.
pub fn index(self) -> f32 {
Self::ALL.iter().position(|m| *m == self).unwrap_or(0) as f32
}
/// The method a stored number names; out of range is the default, as
/// from a newer build with more of them.
pub fn from_index(value: f32) -> Method {
let i = value.round();
if i >= 0.0 && (i as usize) < Self::ALL.len() {
Self::ALL[i as usize]
} else {
Self::DEFAULT
}
}
/// Whether a network runs at all.
pub fn learned(self) -> bool {
self != Method::Bilinear
}
}
/// The best network by default, at full strength: every Bayer raw is
/// developed from the learned demosaic, and the choice and the slider are
/// there to take it back, trade it for time, or ease it off. It costs seconds per photograph the first time, while
/// the classical demosaic shows; the result is cached, so a photograph
/// reopened or exported does not pay again (docs/dev/denoise.md §7).
pub(crate) static DESCRIPTOR: LazyLock<Arc<OpDescriptor>> = LazyLock::new(|| {
Arc::new(OpDescriptor {
id: ID,
label: LocalizedKey("op.learned_denoise"),
params: vec![
ParamDescriptor::switch("apply", "param.learned_denoise.apply"),
ParamDescriptor::choice(
"method",
"param.learned_denoise.method",
vec![
LocalizedKey("param.learned_denoise.method.bilinear"),
LocalizedKey("param.learned_denoise.method.fast"),
LocalizedKey("param.learned_denoise.method.medium"),
LocalizedKey("param.learned_denoise.method.best"),
],
)
.with_default(Method::DEFAULT.index()),
ParamDescriptor::scalar(
"grain",
"param.learned_denoise.grain",
"strength",
"param.learned_denoise.strength",
0.0,
100.0,
0.0,
100.0,
Unit::Percent,
Scale::Linear,
0,
+11 -2
View File
@@ -47,9 +47,18 @@ pub const LOOK: ParamId = ParamId("look");
/// The look's strength at which the LookTable is applied as the profile
/// states it, in percent.
pub const DEFAULT_LOOK: f32 = 100.0;
pub const PROFILE_LOOK: f32 = 100.0;
/// The look's default strength: off.
///
/// Measured, not chosen. Against the photographer's earlier exports with no
/// look applied, the default rendering scores the same with the table at 100,
/// 50 or 0 (held-out MSE 140, 140, 143), and is 9 % more colourful without
/// it: the table desaturates near-neutral tones, which is exactly where the
/// default rendering was short of those exports. The table stays one slider
/// away for anyone who wants the profile's look.
pub const DEFAULT_LOOK: f32 = 0.0;
/// Twice the profile's look.
pub const MAX_LOOK: f32 = 200.0;
pub const MAX_LOOK: f32 = 2.0 * PROFILE_LOOK;
/// Entries of the buffer's header, before the entries themselves: one
/// `vec4` describing each table — `(hue divisions, saturation divisions,
+32 -6
View File
@@ -39,6 +39,16 @@ use crate::ops::helpers;
pub const ID: OpId = OpId("colour_mixer");
/// How much a raised saturation band adds to a muted colour, per unit of its
/// value; the push tapers linearly to nothing at full saturation.
///
/// Measured, not chosen. Fitted against the photographer's earlier exports
/// (two looks, ~90 photographs), a sky band raised by 58 there lifted muted
/// sky blues about 2.1×; at 3.0 a band here does about the same at the same
/// value, so imported values stay inside the slider's range (dr-preset-xmp
/// carries the per-band factors). Lowering saturation is unaffected.
pub const SAT_GAIN: f32 = 3.0;
/// The twelve bands, in hue order starting at red.
///
/// Twelve rather than Lightroom's eight: the extra bands fall between the
@@ -333,9 +343,16 @@ impl Operation for ColourMixer {
// Only the bands the user actually touched contribute code. A single
// adjusted band therefore costs one weight evaluation rather than
// twelve — the composition property applied within an operation.
let mut lines = String::from(
let lines = String::from(
"\
let hcl = rgb_to_hcl(c);
// Bands are matched, and saturation judged, on display-encoded values: in
// scene-linear light a muted colour reads as strongly saturated and its hue
// sits away from where it is seen, so a band set on what the photograph
// shows would land on other colours.
let lin = max(c, vec3<f32>(0.0));
let e = pow(lin, vec3<f32>(1.0 / 2.2));
let SAT_GAIN = @SAT_GAIN@;
let hcl = rgb_to_hcl(e);
let hue = hcl.x;
let chroma = hcl.y;
let hi = hcl.z;
@@ -350,6 +367,8 @@ if (chroma > 0.0001) {
",
);
let mut lines = lines.replace("@SAT_GAIN@", &format!("{SAT_GAIN:.4}"));
for (b, band) in BANDS.iter().enumerate() {
let v = self.values[b];
if v.iter().all(|x| *x == 0.0) {
@@ -384,10 +403,17 @@ if (chroma > 0.0001) {
// yellow-green to green, not enough to turn it blue by accident.
let new_hue = hue + d_hue * 30.0;
// Saturation scales chroma; luminance scales the whole colour.
let new_chroma = clamp(chroma * (1.0 + d_sat), 0.0, hi);
c = hue_to_rgb_scale(new_hue, new_chroma, hi);
c = c * exp2(d_lum);
// Saturation. Raising it pushes muted colours hardest and tapers to
// nothing at full saturation, as the eye expects a mixer to; the
// gain makes a value deliver the strength it names, measured
// against the photographer's earlier exports. Lowering it scales
// every colour alike, so -100 is grey.
let sat = chroma / max(hi, 0.00001);
let gain = select(1.0 + d_sat, 1.0 + d_sat * SAT_GAIN * (1.0 - sat), d_sat > 0.0);
let new_chroma = clamp(chroma * gain, 0.0, hi);
let shifted = hue_to_rgb_scale(new_hue, new_chroma, hi);
// Back to linear light, where luminance scales the whole colour.
c = pow(shifted, vec3<f32>(2.2)) * exp2(d_lum);
}
}
c = max(c, vec3<f32>(0.0));",
+141
View File
@@ -675,6 +675,56 @@ impl PresetLibrary {
self.unknown.values().map(Vec::len).sum()
}
/// TRACES: FR-DEV-6
/// Merge two copies of a library that both descend from `base`.
///
/// For keeping one library on several devices: `base` is what the last
/// exchange left both sides holding, `ours` this device's copy now and
/// `theirs` the server's. Each name is decided on its own:
///
/// - Changed on one side only — added, edited or deleted — that side's
/// answer stands. This is why the base is needed at all: without it a
/// preset deleted here and one added there look the same, and a
/// deletion would come back on every exchange.
/// - Changed on both sides to the same thing, nothing to decide.
/// - Deleted on one side and edited on the other, the edit stands. A
/// preset is work, and an absence is not.
/// - Edited on both sides differently, ours stands. Either answer loses
/// one edit; this one at least converges, since the other device takes
/// ours on its next exchange as an edit made on one side only.
///
/// A missing `base` is an empty one, which can only add: a device's first
/// exchange is a union of the two libraries, never a deletion.
pub fn merge(base: &Self, ours: &Self, theirs: &Self) -> Self {
let mut out = Self::default();
let names: std::collections::BTreeSet<&str> = ours.names().chain(theirs.names()).collect();
for name in names {
let (b, o, t) = (base.entry(name), ours.entry(name), theirs.entry(name));
let chosen = if o == t || t == b {
o
} else if o == b {
t
} else {
o.or(t)
};
if let Some((preset, unknown)) = chosen {
out.presets.insert(name.to_string(), preset.clone());
if let Some(lines) = unknown {
out.unknown.insert(name.to_string(), lines.clone());
}
}
}
out
}
/// One name's preset together with the lines kept beside it, which are
/// part of what that preset is when two copies are compared.
fn entry(&self, name: &str) -> Option<(&Preset, Option<&Vec<String>>)> {
self.presets
.get(name)
.map(|preset| (preset, self.unknown.get(name)))
}
/// Serialise to the on-disk form.
///
/// Deterministic, like the sidecar's: the same library always produces the
@@ -1479,6 +1529,97 @@ mod tests {
assert_eq!(other.to_text(), named().to_text());
}
/// A one-parameter preset, so two of them differ by their value.
fn exposure(ev: f32) -> Preset {
let mut params = BTreeMap::new();
params.insert(("exposure".to_string(), "exposure".to_string()), ev);
Preset::from_params(params)
}
fn library_of(entries: &[(&str, f32)]) -> PresetLibrary {
let mut lib = PresetLibrary::default();
for (name, ev) in entries {
lib.insert(name, exposure(*ev)).unwrap();
}
lib
}
#[test]
fn a_first_merge_is_the_union_of_both_libraries() {
let ours = library_of(&[("Mine", 1.0), ("Both", 0.5)]);
let theirs = library_of(&[("Theirs", 2.0), ("Both", 0.5)]);
let merged = PresetLibrary::merge(&PresetLibrary::default(), &ours, &theirs);
assert_eq!(
merged.names().collect::<Vec<_>>(),
vec!["Both", "Mine", "Theirs"]
);
}
#[test]
fn a_deletion_on_either_side_is_kept_rather_than_undone() {
// The case the base exists for: without it, the deleted preset is
// indistinguishable from one the other side has just added.
let base = library_of(&[("Gone here", 1.0), ("Gone there", 2.0)]);
let ours = library_of(&[("Gone there", 2.0)]);
let theirs = library_of(&[("Gone here", 1.0)]);
assert!(PresetLibrary::merge(&base, &ours, &theirs).is_empty());
}
#[test]
fn an_edit_on_one_side_reaches_the_other() {
let base = library_of(&[("Warm", 1.0)]);
let ours = library_of(&[("Warm", 1.0)]);
let theirs = library_of(&[("Warm", 1.5)]);
let merged = PresetLibrary::merge(&base, &ours, &theirs);
assert_eq!(merged.get("Warm"), Some(&exposure(1.5)));
// And the other way round.
let merged = PresetLibrary::merge(&base, &theirs, &ours);
assert_eq!(merged.get("Warm"), Some(&exposure(1.5)));
}
#[test]
fn an_edit_outlives_a_deletion_made_elsewhere() {
let base = library_of(&[("Warm", 1.0)]);
let edited = library_of(&[("Warm", 1.5)]);
let deleted = PresetLibrary::default();
assert_eq!(
PresetLibrary::merge(&base, &edited, &deleted).get("Warm"),
Some(&exposure(1.5))
);
assert_eq!(
PresetLibrary::merge(&base, &deleted, &edited).get("Warm"),
Some(&exposure(1.5))
);
}
#[test]
fn two_different_edits_keep_ours_and_then_converge() {
let base = library_of(&[("Warm", 1.0)]);
let here = library_of(&[("Warm", 1.5)]);
let there = library_of(&[("Warm", 0.5)]);
let pushed = PresetLibrary::merge(&base, &here, &there);
assert_eq!(pushed.get("Warm"), Some(&exposure(1.5)));
// The other device's next exchange: its base is what it last pushed,
// its own copy is unchanged since, and the server holds ours.
let settled = PresetLibrary::merge(&there, &there, &pushed);
assert_eq!(settled, pushed);
}
#[test]
fn lines_this_build_cannot_read_travel_with_their_preset() {
let text = format!(
"drpl {LIBRARY_FORMAT_VERSION}\n\n[preset Future]\nexposure.exposure = 0.5\n\
something_new_entirely\n"
);
let theirs = PresetLibrary::parse(&text).unwrap();
let merged = PresetLibrary::merge(
&PresetLibrary::default(),
&PresetLibrary::default(),
&theirs,
);
assert!(merged.to_text().contains("something_new_entirely"));
}
#[test]
fn a_neutral_preset_is_storable_and_survives_the_round_trip() {
// The empty preset is the "clear these forty frames" action, so it has
+188 -49
View File
@@ -83,35 +83,43 @@ const MAPPINGS: &[Mapping] = &[
param: "exposure",
convert: Convert::Direct,
},
// The five tone sliders do not mean the same thing in the two applications,
// whatever their shared ±100 suggests. The factors were fitted against the
// library's own Lightroom 6 exports and their raws (darkroom-lrfit, 2026-10):
// each photograph's sliders carried across as `slider × factor`, one factor
// per slider, on ~90 exports with no look applied. Contrast is ours at a
// tenth — at −100 ours flattens a frame to grey — and our shadows need
// nearly twice Lightroom's number. Whites barely appears in those exports;
// every fit put it under 1 but none agreed where, so 0.5 is a hedge.
Mapping {
crs: "Contrast2012",
op: "contrast",
param: "contrast",
convert: Convert::Direct,
convert: Convert::Scale(0.1),
},
Mapping {
crs: "Highlights2012",
op: "highlights_shadows",
param: "highlights",
convert: Convert::Direct,
convert: Convert::Scale(1.4),
},
Mapping {
crs: "Shadows2012",
op: "highlights_shadows",
param: "shadows",
convert: Convert::Direct,
convert: Convert::Scale(1.9),
},
Mapping {
crs: "Whites2012",
op: "blacks_whites",
param: "whites",
convert: Convert::Direct,
convert: Convert::Scale(0.5),
},
Mapping {
crs: "Blacks2012",
op: "blacks_whites",
param: "blacks",
convert: Convert::Direct,
convert: Convert::Scale(1.25),
},
Mapping {
crs: "Clarity2012",
@@ -139,12 +147,21 @@ const MAPPINGS: &[Mapping] = &[
},
// TRACES: FR-DEV-6
// Lightroom's HSL panel: eight bands, each ±100 for hue, saturation and
// luminance, onto the colour mixer's twelve. Lightroom's bands sit where
// ours do except two, matched to the nearest of ours by hue: Aqua (180°)
// is our cyan, Purple (270°) our violet; chartreuse, spring, azure and
// rose have no Lightroom counterpart and are left alone. One for one, as
// a first translation — the band widths differ, and `lr-fit`'s
// measurement against Lightroom's own output may yet scale these.
// luminance, onto the colour mixer's twelve.
//
// Hue and luminance go to the band of the same hue, one for one: Aqua
// (180°) is our cyan, Purple (270°) our violet.
//
// Saturation is measured. Fitted against the library's Lightroom 6 exports
// of two looks and their raws (darkroom-lrfit, hsl_map_fit), Lightroom's
// saturation bands are about 45° wide either side on our hue wheel, wider
// than ours, and do not all have our strength: each is shared between
// two or three of our bands with the factors below. Aqua sits at 187°
// and reaches into azure, where skies are; Blue at 251° reaches violet;
// Orange, where skin is, carries only ~0.4 — Lightroom's Orange is
// gentle. Red, Purple and Magenta barely appear in those exports and
// take the common gain; Green is capped where its few pixels would push
// it further. Values add when two Lightroom bands share one of ours.
Mapping {
crs: "HueAdjustmentRed",
op: "colour_mixer",
@@ -155,7 +172,19 @@ const MAPPINGS: &[Mapping] = &[
crs: "SaturationAdjustmentRed",
op: "colour_mixer",
param: "red_sat",
convert: Convert::Direct,
convert: Convert::Scale(1.02),
},
Mapping {
crs: "SaturationAdjustmentRed",
op: "colour_mixer",
param: "orange_sat",
convert: Convert::Scale(0.34),
},
Mapping {
crs: "SaturationAdjustmentRed",
op: "colour_mixer",
param: "rose_sat",
convert: Convert::Scale(0.34),
},
Mapping {
crs: "LuminanceAdjustmentRed",
@@ -173,7 +202,19 @@ const MAPPINGS: &[Mapping] = &[
crs: "SaturationAdjustmentOrange",
op: "colour_mixer",
param: "orange_sat",
convert: Convert::Direct,
convert: Convert::Scale(0.39),
},
Mapping {
crs: "SaturationAdjustmentOrange",
op: "colour_mixer",
param: "red_sat",
convert: Convert::Scale(0.13),
},
Mapping {
crs: "SaturationAdjustmentOrange",
op: "colour_mixer",
param: "yellow_sat",
convert: Convert::Scale(0.13),
},
Mapping {
crs: "LuminanceAdjustmentOrange",
@@ -191,7 +232,19 @@ const MAPPINGS: &[Mapping] = &[
crs: "SaturationAdjustmentYellow",
op: "colour_mixer",
param: "yellow_sat",
convert: Convert::Direct,
convert: Convert::Scale(0.87),
},
Mapping {
crs: "SaturationAdjustmentYellow",
op: "colour_mixer",
param: "orange_sat",
convert: Convert::Scale(0.29),
},
Mapping {
crs: "SaturationAdjustmentYellow",
op: "colour_mixer",
param: "chartreuse_sat",
convert: Convert::Scale(0.29),
},
Mapping {
crs: "LuminanceAdjustmentYellow",
@@ -209,7 +262,19 @@ const MAPPINGS: &[Mapping] = &[
crs: "SaturationAdjustmentGreen",
op: "colour_mixer",
param: "green_sat",
convert: Convert::Direct,
convert: Convert::Scale(1.2),
},
Mapping {
crs: "SaturationAdjustmentGreen",
op: "colour_mixer",
param: "chartreuse_sat",
convert: Convert::Scale(0.4),
},
Mapping {
crs: "SaturationAdjustmentGreen",
op: "colour_mixer",
param: "spring_sat",
convert: Convert::Scale(0.4),
},
Mapping {
crs: "LuminanceAdjustmentGreen",
@@ -227,7 +292,19 @@ const MAPPINGS: &[Mapping] = &[
crs: "SaturationAdjustmentAqua",
op: "colour_mixer",
param: "cyan_sat",
convert: Convert::Direct,
convert: Convert::Scale(0.9),
},
Mapping {
crs: "SaturationAdjustmentAqua",
op: "colour_mixer",
param: "azure_sat",
convert: Convert::Scale(0.51),
},
Mapping {
crs: "SaturationAdjustmentAqua",
op: "colour_mixer",
param: "spring_sat",
convert: Convert::Scale(0.19),
},
Mapping {
crs: "LuminanceAdjustmentAqua",
@@ -245,7 +322,19 @@ const MAPPINGS: &[Mapping] = &[
crs: "SaturationAdjustmentBlue",
op: "colour_mixer",
param: "blue_sat",
convert: Convert::Direct,
convert: Convert::Scale(0.75),
},
Mapping {
crs: "SaturationAdjustmentBlue",
op: "colour_mixer",
param: "violet_sat",
convert: Convert::Scale(0.56),
},
Mapping {
crs: "SaturationAdjustmentBlue",
op: "colour_mixer",
param: "azure_sat",
convert: Convert::Scale(0.09),
},
Mapping {
crs: "LuminanceAdjustmentBlue",
@@ -263,7 +352,19 @@ const MAPPINGS: &[Mapping] = &[
crs: "SaturationAdjustmentPurple",
op: "colour_mixer",
param: "violet_sat",
convert: Convert::Direct,
convert: Convert::Scale(1.26),
},
Mapping {
crs: "SaturationAdjustmentPurple",
op: "colour_mixer",
param: "blue_sat",
convert: Convert::Scale(0.42),
},
Mapping {
crs: "SaturationAdjustmentPurple",
op: "colour_mixer",
param: "magenta_sat",
convert: Convert::Scale(0.42),
},
Mapping {
crs: "LuminanceAdjustmentPurple",
@@ -281,7 +382,19 @@ const MAPPINGS: &[Mapping] = &[
crs: "SaturationAdjustmentMagenta",
op: "colour_mixer",
param: "magenta_sat",
convert: Convert::Direct,
convert: Convert::Scale(1.05),
},
Mapping {
crs: "SaturationAdjustmentMagenta",
op: "colour_mixer",
param: "violet_sat",
convert: Convert::Scale(0.35),
},
Mapping {
crs: "SaturationAdjustmentMagenta",
op: "colour_mixer",
param: "rose_sat",
convert: Convert::Scale(0.35),
},
Mapping {
crs: "LuminanceAdjustmentMagenta",
@@ -458,7 +571,12 @@ pub fn read_xmp(text: &str) -> Result<Import, ImportError> {
// Out-of-range values are left as they are: `EditGraph::set_param`
// clamps when the preset is applied, and clamping here as well would
// mean two places to be wrong about a range.
params.insert((mapping.op.to_string(), mapping.param.to_string()), value);
//
// Added rather than set: one of Lightroom's HSL bands is shared
// between two or three of ours, and two of its bands can share one.
*params
.entry((mapping.op.to_string(), mapping.param.to_string()))
.or_insert(0.0) += value;
}
let skipped = KNOWN_UNSUPPORTED
@@ -561,13 +679,19 @@ mod tests {
let import = read_embedded(&file).expect("an edit");
let p = import.preset.params();
let get = |op: &str, param: &str| p.get(&(op.to_string(), param.to_string())).copied();
assert_eq!(get("colour_mixer", "blue_sat"), Some(58.0));
assert_eq!(get("colour_mixer", "cyan_sat"), Some(50.0));
assert_eq!(get("colour_mixer", "violet_sat"), Some(23.0));
// Saturation is shared out by the measured factors; values add.
let near = |got: Option<f32>, want: f32| {
let got = got.expect("set");
assert!((got - want).abs() < 1e-3, "{got} vs {want}");
};
near(get("colour_mixer", "blue_sat"), 58.0 * 0.75 + 23.0 * 0.42);
near(get("colour_mixer", "cyan_sat"), 50.0 * 0.9);
near(get("colour_mixer", "azure_sat"), 50.0 * 0.51 + 58.0 * 0.09);
near(get("colour_mixer", "violet_sat"), 58.0 * 0.56 + 23.0 * 1.26);
assert_eq!(get("colour_mixer", "red_hue"), Some(-5.0));
assert_eq!(get("colour_mixer", "green_lum"), Some(7.0));
assert_eq!(get("highlights_shadows", "highlights"), Some(-40.0));
assert_eq!(get("blacks_whites", "blacks"), Some(-20.0));
assert_eq!(get("highlights_shadows", "highlights"), Some(-40.0 * 1.4));
assert_eq!(get("blacks_whites", "blacks"), Some(-20.0 * 1.25));
}
#[test]
@@ -586,9 +710,12 @@ mod tests {
let import = read_embedded(&bytes).expect("Lightroom's edit");
let p = import.preset.params();
let get = |op: &str, param: &str| p.get(&(op.to_string(), param.to_string())).copied();
assert_eq!(get("colour_mixer", "blue_sat"), Some(58.0));
assert_eq!(get("colour_mixer", "cyan_sat"), Some(50.0));
assert_eq!(get("highlights_shadows", "highlights"), Some(-40.0));
// The house look's sky: Aqua and Blue land in cyan, azure and blue.
for (param, at_least) in [("cyan_sat", 40.0), ("azure_sat", 25.0), ("blue_sat", 40.0)] {
let v = get("colour_mixer", param).unwrap_or(0.0);
assert!(v >= at_least, "{param} {v}");
}
assert_eq!(get("highlights_shadows", "highlights"), Some(-40.0 * 1.4));
}
#[test]
@@ -620,36 +747,48 @@ mod tests {
}
#[test]
fn no_two_mappings_claim_the_same_key_or_the_same_target() {
let mut keys: Vec<&str> = MAPPINGS.iter().map(|m| m.crs).collect();
keys.sort_unstable();
let before = keys.len();
keys.dedup();
assert_eq!(before, keys.len(), "two mappings read the same crs key");
fn no_two_mappings_repeat_a_key_and_target() {
// A saturation band may be shared between several of ours, and two of
// Lightroom's may share one of ours (their values add); but the same
// key written twice to the same target would count it twice.
let mut pairs: Vec<(&str, &str, &str)> =
MAPPINGS.iter().map(|m| (m.crs, m.op, m.param)).collect();
pairs.sort_unstable();
let before = pairs.len();
pairs.dedup();
assert_eq!(before, pairs.len(), "a key is written twice to one target");
let mut targets: Vec<(&str, &str)> = MAPPINGS.iter().map(|m| (m.op, m.param)).collect();
targets.sort_unstable();
let before = targets.len();
targets.dedup();
assert_eq!(
before,
targets.len(),
"two mappings write the same parameter"
);
// Only the HSL saturation bands are shared; every other key has one
// home, so a slip in the table cannot fan a slider out unnoticed.
let mut single: Vec<&str> = MAPPINGS
.iter()
.map(|m| m.crs)
.filter(|k| !k.starts_with("SaturationAdjustment"))
.collect();
single.sort_unstable();
let before = single.len();
single.dedup();
assert_eq!(before, single.len(), "two mappings read the same crs key");
}
#[test]
fn the_settings_that_share_a_convention_come_across_unchanged() {
let import = read_xmp(ATTRIBUTE_FORM).unwrap();
assert_eq!(value(&import, "exposure", "exposure"), Some(0.75));
assert_eq!(value(&import, "contrast", "contrast"), Some(25.0));
assert_eq!(value(&import, "contrast", "contrast"), Some(25.0 * 0.1));
assert_eq!(
value(&import, "highlights_shadows", "highlights"),
Some(-40.0)
Some(-40.0 * 1.4)
);
assert_eq!(
value(&import, "highlights_shadows", "shadows"),
Some(30.0 * 1.9)
);
assert_eq!(value(&import, "blacks_whites", "whites"), Some(10.0 * 0.5));
assert_eq!(
value(&import, "blacks_whites", "blacks"),
Some(-15.0 * 1.25)
);
assert_eq!(value(&import, "highlights_shadows", "shadows"), Some(30.0));
assert_eq!(value(&import, "blacks_whites", "whites"), Some(10.0));
assert_eq!(value(&import, "blacks_whites", "blacks"), Some(-15.0));
assert_eq!(value(&import, "clarity", "amount"), Some(12.0));
assert_eq!(value(&import, "texture", "amount"), Some(8.0));
assert_eq!(value(&import, "vibrance", "vibrance"), Some(20.0));
@@ -690,7 +829,7 @@ mod tests {
// does not say which shape it used.
let import = read_xmp(ELEMENT_FORM).unwrap();
assert_eq!(value(&import, "exposure", "exposure"), Some(0.75));
assert_eq!(value(&import, "contrast", "contrast"), Some(25.0));
assert_eq!(value(&import, "contrast", "contrast"), Some(25.0 * 0.1));
}
#[test]
+14 -4
View File
@@ -13,9 +13,11 @@
use std::path::Path;
const MODEL: &str = "../../models/segment/yolo26n-seg.onnx";
const QUANTISED: &str = "../../models/segment/yolo26n-seg.a16w16.onnx";
fn main() {
println!("cargo:rerun-if-changed={MODEL}");
println!("cargo:rerun-if-changed={QUANTISED}");
println!("cargo:rerun-if-changed=build.rs");
// Only the embedded path needs the file present; a build without it is
@@ -24,10 +26,18 @@ fn main() {
return;
}
let path = Path::new(MODEL);
check(MODEL);
// The Hexagon's quantised form rides only in an Android build.
if std::env::var("CARGO_CFG_TARGET_OS").as_deref() == Ok("android") {
check(QUANTISED);
}
}
fn check(model: &str) {
let path = Path::new(model);
let Ok(bytes) = std::fs::read(path) else {
panic!(
"\n\n{MODEL} is missing.\n\
"\n\n{model} is missing.\n\
It ships in Git LFS. Run `git lfs install && git lfs pull`, or build \
with `--no-default-features` for a watershed-only build.\n"
);
@@ -40,7 +50,7 @@ fn main() {
// happens in practice.
if bytes.starts_with(b"version https://git-lfs") {
panic!(
"\n\n{MODEL} is a Git LFS pointer, not the model ({} bytes).\n\
"\n\n{model} is a Git LFS pointer, not the model ({} bytes).\n\
Run `git lfs install && git lfs pull` to fetch the real file.\n",
bytes.len()
);
@@ -50,7 +60,7 @@ fn main() {
// export is ~11 MB; anything under a megabyte is a truncated checkout.
if bytes.len() < 1_000_000 {
panic!(
"\n\n{MODEL} is only {} bytes — expected ~11 MB.\n\
"\n\n{model} is only {} bytes — expected several MB.\n\
The checkout looks incomplete; try `git lfs pull`.\n",
bytes.len()
);
+1 -1
View File
@@ -72,7 +72,7 @@ pub use refine::{
#[cfg(feature = "semantic")]
pub use scene::{Category, Scene, SceneModel};
#[cfg(feature = "embedded-model")]
pub use semantic::embedded_model_bytes;
pub use semantic::embedded_models;
#[cfg(feature = "semantic")]
pub use semantic::{Instance, SemanticModel, SemanticOptions, Tiling};
+15 -7
View File
@@ -123,21 +123,29 @@ impl SceneModel {
classes: impl AsRef<std::path::Path>,
categories: impl AsRef<std::path::Path>,
) -> Result<Self, SegmentError> {
// The form the device's backend runs: the `.a16w16.onnx` sibling on
// the Hexagon (attention left in float, inference.md §1.5), else this.
let (model, form) =
dr_inference_engine::resolve_model(dr_inference_engine::Role::Scene, model.as_ref());
let bytes = std::fs::read(model).map_err(SegmentError::ModelRead)?;
let classes = std::fs::read_to_string(classes).map_err(SegmentError::ModelRead)?;
let categories = std::fs::read_to_string(categories).map_err(SegmentError::ModelRead)?;
let classes = crate::semantic::parse_classes(&classes);
let categories = parse_categories(&categories, &classes)?;
Self::from_bytes(&bytes, categories)
Self::from_bytes_in(&bytes, form, categories)
}
pub fn from_bytes(bytes: &[u8], categories: Vec<Category>) -> Result<Self, SegmentError> {
// f32, as for `SemanticModel`; see there.
let session = dr_inference_engine::open(
dr_inference_engine::Role::Scene,
dr_inference_engine::Form::F32,
bytes,
)?;
Self::from_bytes_in(bytes, dr_inference_engine::Form::F32, categories)
}
/// `bytes` in a stated numeric form; the outputs keep their shape.
pub fn from_bytes_in(
bytes: &[u8],
form: dr_inference_engine::Form,
categories: Vec<Category>,
) -> Result<Self, SegmentError> {
let session = dr_inference_engine::open(dr_inference_engine::Role::Scene, form, bytes)?;
Ok(Self {
session,
+31 -13
View File
@@ -208,18 +208,32 @@ const EMBEDDED_MODEL: &[u8] = include_bytes!("../../../models/segment/yolo26n-se
#[cfg(feature = "embedded-model")]
const EMBEDDED_CLASSES: &str = include_str!("../../../models/segment/yolo26n-seg.classes.json");
/// The bytes of the model that ships with this crate, for whoever compiles
/// The Hexagon's form (docs/dev/inference.md §1.5): 16-bit activations and
/// weights, the rows' tail left in float. Only Android has a Hexagon, so only
/// Android carries it.
#[cfg(all(feature = "embedded-model", target_os = "android"))]
const EMBEDDED_A16W16: &[u8] = include_bytes!("../../../models/segment/yolo26n-seg.a16w16.onnx");
/// Every form of the model that ships with this crate, for whoever compiles
/// engines ahead of the first request (docs/dev/inference.md §6).
#[cfg(feature = "embedded-model")]
pub fn embedded_model_bytes() -> &'static [u8] {
EMBEDDED_MODEL
pub fn embedded_models() -> Vec<(dr_inference_engine::Form, &'static [u8])> {
#[allow(unused_mut)]
let mut forms = vec![(dr_inference_engine::Form::F32, EMBEDDED_MODEL)];
#[cfg(target_os = "android")]
forms.push((dr_inference_engine::Form::A16W16, EMBEDDED_A16W16));
forms
}
impl SemanticModel {
/// Load the model that ships with this crate.
/// Load the model that ships with this crate, in the form the device's
/// backend runs.
#[cfg(feature = "embedded-model")]
pub fn embedded() -> Result<Self, SegmentError> {
Self::from_bytes(EMBEDDED_MODEL, parse_classes(EMBEDDED_CLASSES))
let forms = embedded_models();
let (bytes, form) =
dr_inference_engine::choose_embedded(dr_inference_engine::Role::Segmenter, &forms);
Self::from_bytes_in(bytes, form, parse_classes(EMBEDDED_CLASSES))
}
/// Load a model from an ONNX file, with `classes` supplying its vocabulary.
@@ -236,14 +250,18 @@ impl SemanticModel {
}
pub fn from_bytes(bytes: &[u8], classes: Vec<Arc<str>>) -> Result<Self, SegmentError> {
// The f32 graph on whatever the device's backend is. An int8 form
// for the Hexagon waits on docs/dev/inference.md §10 M7 — the mask
// boundary has to be measured before it moves.
let session = dr_inference_engine::open(
dr_inference_engine::Role::Segmenter,
dr_inference_engine::Form::F32,
bytes,
)?;
Self::from_bytes_in(bytes, dr_inference_engine::Form::F32, classes)
}
/// `bytes` in a stated numeric form. The quantised one keeps the same
/// outputs (the rows' tail stays float), so decoding does not change; the
/// masks it draws were measured against f32's (inference.md §1.5).
pub fn from_bytes_in(
bytes: &[u8],
form: dr_inference_engine::Form,
classes: Vec<Arc<str>>,
) -> Result<Self, SegmentError> {
let session = dr_inference_engine::open(dr_inference_engine::Role::Segmenter, form, bytes)?;
Ok(Self { session, classes })
}
+28 -9
View File
@@ -275,13 +275,26 @@ impl FaceDetector {
}
}
/// Both ids this detector writes under — the f32 form and the int8 one —
/// for a question that is about the detector and not about which form
/// of it a device happened to run: "has the chosen detector been over
/// this image", asked by a re-index that must not ping-pong between a
/// desktop that runs it in f32 and a tablet that runs it on the Hexagon.
pub fn model_ids(self) -> [&'static str; 2] {
[self.model_id(), self.model_id_int8()]
/// The id when the detector runs with 16-bit activations and 8-bit
/// weights, the Hexagon's form since the int8 one lost faces at 40–80 px
/// (docs/dev/inference.md §1.5). Different again from both, for the same
/// reason as [`Self::model_id_int8`]; the embedder half is unchanged.
pub fn model_id_a16w8(self) -> &'static str {
match self {
FaceDetector::Scrfd500m => "scrfd_500m_a16+w600k_mbf",
FaceDetector::Scrfd2_5g => "scrfd_2.5g_a16+w600k_mbf",
FaceDetector::Scrfd10g => "scrfd_10g_a16+w600k_mbf",
}
}
/// Every id this detector writes under — the f32 form and each quantised
/// one a device has run — for a question that is about the detector and
/// not about which form of it a device happened to run: "has the chosen
/// detector been over this image", asked by a re-index that must not
/// ping-pong between a desktop that runs it in f32 and a tablet that
/// runs it on the Hexagon.
pub fn model_ids(self) -> [&'static str; 3] {
[self.model_id(), self.model_id_int8(), self.model_id_a16w8()]
}
/// The detector that writes under a pipeline id, if it is one of these.
@@ -1769,11 +1782,16 @@ mod tests {
#[test]
fn a_detectors_two_spellings_share_its_embedder_and_nothing_else() {
for d in FaceDetector::ALL {
let [f32_id, int8_id] = d.model_ids();
let [f32_id, int8_id, a16_id] = d.model_ids();
assert_eq!(f32_id, d.model_id());
assert_eq!(int8_id, d.model_id_int8());
assert_eq!(a16_id, d.model_id_a16w8());
assert_ne!(f32_id, int8_id);
assert_eq!(f32_id.rsplit('+').next(), int8_id.rsplit('+').next());
assert_ne!(f32_id, a16_id);
assert_ne!(int8_id, a16_id);
for q in [int8_id, a16_id] {
assert_eq!(f32_id.rsplit('+').next(), q.rsplit('+').next());
}
}
}
@@ -1797,6 +1815,7 @@ mod tests {
}
assert_eq!(FaceDetector::for_model_id(d.model_id()), Some(d));
assert_eq!(FaceDetector::for_model_id(d.model_id_int8()), Some(d));
assert_eq!(FaceDetector::for_model_id(d.model_id_a16w8()), Some(d));
}
assert_eq!(FaceDetector::for_model_id("scrfd_10g+other"), None);
}
+5 -1
View File
@@ -54,10 +54,14 @@ sed 's/$/\r/' "${REPO}/LICENSE" > "${STAGE}/LICENSE"
# with its two descriptors, the panorama border filler and the denoiser. The installer
# smoke test counts the same directories, so a model added here is expected
# there without a number to update.
#
# Not the quantised siblings (`*.int8.onnx`, `*.a16w8.onnx`, `*.a16w16.onnx`):
# they are the Hexagon's forms (docs/dev/inference.md §1.5), and a Windows
# machine has no Hexagon to load them.
for dir in face scene inpaint denoise; do
for f in "${REPO}/models/${dir}"/*; do
case "$(basename "${f}")" in
README.md) continue ;;
README.md | *.int8.onnx | *.a16w8.onnx | *.a16w16.onnx) continue ;;
esac
if [[ "${f}" == *.onnx && "$(stat -c%s "${f}")" -lt 100000 ]]; then
echo "error: $(basename "${f}") is $(stat -c%s "${f}") bytes — an LFS pointer, not a model." >&2
+106 -1
View File
@@ -331,6 +331,17 @@ If neither holds S's quality within 0.5 dB of fp32 on the real pairs, **v1 is de
tablet shows the classical path. The sidecar still records the intent, so a desktop can render the
learned result for a photograph edited on the tablet.
**Measured 2026-10-04 (inference.md §1.5): the second way holds, without the first.** The shipped
network, with its Bayer packing re-spelled as `SpaceToDepth` so QNN can hold it (the 6-D reshape
it replaces is exact but past the HTP's rank limit), at A16W16 — 16-bit activations and weights —
scores within 0.00 dB of f32 at ISO 400–25600 on the tablet's own HTP, and within 0.09 dB with the
6D's noise model scaled ×0.5, ×2 and ×4 to stand in for other sensors. A16W8 holds the 6D (worst
−0.19 dB at ISO 25600) but not ×4 noise at 25600 (−0.52 dB), so A16W16 is what ships. int8 loses
4.7–9.2 dB and fp16 is refused outright. A 1408 tile takes 95 ms on the Hexagon against 1510 ms on
the tablet's CPU: about 2.3 s for a 20 MP frame. Calibration ranges come from 96 training-day
tiles across every ISO, a third of them with that scaled noise; coverage of other bodies is that
synthetic bracket, not their raws.
## 9. X-Trans
The requirements tie this stage to FR-RAW-5, and the library has no Fuji raws. What we can do
@@ -418,6 +429,100 @@ measured while the GPU sat power-capped at an 810 MHz memory clock; uncapped is
about four times faster. The Rust path reproduces the training repository's output to 2.5e-4 at
worst; TensorRT fp16 is 75 dB from f32.
**Not yet:** the result is not cached across sessions (§7.1) — reopening recomputes; the tripod real
**Not yet:** ~~the result is not cached across sessions (§7.1) — reopening recomputes;~~ done after
0.21.0, §12; the tripod real
pairs of §6.1; X-Trans (§9); the hand-written WGSL path, for which `export.py` already writes the
weights blob and a manifest a shader can follow.
## 12. On by default, with a strength, and cached (after 0.21.0)
The photographer asked for the learned demosaic to be how a raw is developed, not an option found
under Detail. So:
- **On by default, at full strength, on every device.** The switch is `switch_on`, so an untouched
photograph writes nothing and is developed from the network everywhere; turning it off is the
edit. Which hardware runs it is the inference engine's choice (inference.md), not this setting's:
the default does not depend on what a device is believed to manage.
- **Strength replaces Keep grain.** 0–100, default 100, and grain = 100 − strength, so it is the
same luminance-only blend of §7.2 and moving it is one GPU pass, never a re-run. An edit saved by
0.21.0 stored `grain`; it is still read, as its inverse, and never written.
- **First in the panel**, above the lens corrections: it decides what every control below is
applied to. Its attribute is still Detail, so it also stays where the Detail tab shows it.
- **Cached on disk** (§7.1): the network's output for a file, as half floats (about 120 MB for
20 MP — no compressor to link on Android), keyed on a SHA-256 of the file's bytes and the model
file's name and size, oldest first past a 5 GB budget, beside the inference engine's cache under
the data root. The strength is applied afterwards and is not in the key. A reopened photograph,
and an export of one already developed, read it back instead of recomputing.
What it costs: every raw opened runs the network once, with the classical demosaic shown until the
result lands, and a first export of an unopened raw runs it too. Every raw renders differently from
0.21.0 unless switched off.
## 13. Three networks and a method (after 0.22.0)
The photographer asked for a choice between quality and time. `Method` replaces the Apply switch:
`Bilinear`, `Fast`, `Medium`, `Best`, by index in that order, default `Best`. An untouched raw
writes nothing and develops through `Best`. `apply` is still read and never written: 0 is
`Bilinear`, 1 keeps a network already chosen or is the default. A number past the list, from a newer
build, reads as the default. A build before this one ignores `method` and develops through its own
network, which is the most an older peer can do.
**The networks** (darkroom-denoise, every one trained on the same data and noise as §11, plus 1,201
further frames cropped from the library and 6,000 drawn scenes — polygons, lines of one to four
photosites, text, gratings — rendered at 4× through a random affine and smooth displacement, so
edges fall off the photosite grid):
| Method | File | Network | Parameters | GMAC / MP | Halo |
|---|---|---|---|---|---|
| Best | `mosaic-best-1408.onnx` | two U-Nets of §11's shape (a flat expert from `m2`, an edge expert from the ×100 edge-weighted run) and a 128 k-parameter gate that blends them per photosite | 6.4 M | 110 | 256 |
| Medium | `mosaic-medium-1408.onnx` | §11's U-Net, distilled from Best (75 % its output, 25 % the truth) | 3.2 M | 48 | 192 |
| Fast | `mosaic-fast-1408.onnx` | widths 16-32-64-128, blocks 1-1-1-2, distilled the same way | 0.93 M | 11 | 192 |
The gate learned on its own to trust the edge expert at 0.77–0.88 on edges and not at all on flat
areas. The mixture's receptive field is the experts' plus the gate's, so it keeps the centre of a
1408 tile past a 256 halo, where the single networks keep 1024 past 192. `dr_denoise::Shipped`
carries each file's halo, and `TileNet::halo` hands it to the tiler.
**Quality.** PSNR after the display transform on 1,842 held-out crops, and the width of a hard
edge on the drawn chart at ISO 6400 (truth 0.80 photosites; lower is sharper):
| | ISO 400 | 1600 | 6400 | 25600 | Edge width |
|---|---|---|---|---|---|
| §11's network | 40.61 | 39.77 | 38.43 | 36.67 | 1.77 |
| Best | 40.69 | 39.84 | 38.47 | 36.71 | 0.82 |
| Medium | 40.59 | 39.75 | 38.40 | 36.65 | 1.30 |
| Fast | 39.90 | 39.06 | 37.54 | 35.33 | 1.84 |
| Bilinear | 36.15 | 33.26 | 28.72 | 23.83 | 2.15 |
On photographs the three are close; on hard edges Best is half as wide as §11's network and
Medium most of the way there. Fast costs a dB at high ISO and edges as soft as §11's.
**Speed**, a whole 20 MP 6D frame, the network alone, TensorRT fp16 on the laptop's RTX 3050
(uncapped: memory at 5 GHz), engine already built: Best 2.48 s, Medium 0.79 s, Fast 0.57 s. Decode
and the hot-pixel pass add about 0.5 s. The first build of each TensorRT engine takes 80 s (Fast) to
190 s (Best), in the background at first launch, cached after.
**Before the network, two passes changed since §11.**
- *A noise-aware repair* (`dr_denoise::repair`) after the app's hot-pixel pass: a photosite more
than 8σ beyond every same-colour neighbour *and* every adjacent photosite, and more than twice
each adjacent one, is clamped to the brightest of its same-colour neighbours; a dead one, to the
darkest. The ratio test is what spares a point of light, whose neighbours are lit too. The networks
were trained behind the same pass (the Python and Rust agree: 935 repairs on an ISO 25600 frame).
- *The tiler feeds the network without waiting*: tiles are gathered on every core by a producer
thread one tile ahead, and the output is written back in parallel from the runtime's own buffer.
0.14 s of tiler for a frame, which is what keeps Fast under a second.
**The Hexagon.** Each network has an `.a16w16.onnx` sibling made by `tools/quantise-models.sh
--ranges`, the ranges from darkroom-3e's gate (96 training tiles, a third at noise ×2 and ×4). On
the 6D gate A16W16 loses 0.00 dB for all three; with the noise scaled ×0.5–×4 at most 0.11 dB.
A16W8 holds the gate (≤ 0.27 dB) but loses 0.63 dB on Medium at ×4, so A16W16 stays the form.
**Cache.** Each network keys its own results (§7.1 keys on the model's file name), and the file is
hashed once at open, so changing the method never re-reads it. Choosing `Bilinear` keeps the
network's result in memory for the way back; changing to another network drops it, and coming back
reads the cache.
**Packaging.** All six files in the APK (`BUNDLED`, 23 entries, +44.6 MB, ~41 MB compressed); the
three f32 networks in the Arch package and the Windows installer, which stage `models/denoise` by
directory.
+64 -4
View File
@@ -127,7 +127,8 @@ Three things the table settles.
five of six models. A whole-library face index on the tablet goes from ~100 ms + 39 ms per face
to ~1.4 ms + 12 ms per face, and the "Thorough" detector — 3× the cost of "Fast" today — becomes
free. Its price is that the models must be **quantised to int8**, which is an accuracy question
§5 has to answer before it is believed.
§5 has to answer before it is believed. (§1.5 answered it: int8 lost faces, and every model but
XFeat ships with 16-bit activations, at about three times these timings.)
- **The embedder does not gain from either accelerator.** 112×112 input, per-op overhead
dominates; it is 9 ms on the tablet's CPU and 12 ms on its NPU. It stays float, which §7 turns
from a performance footnote into a correctness rule.
@@ -137,6 +138,60 @@ Three things the table settles.
- **On AMD, MIGraphX fp16 is 4–17× the CPU provider** on the detectors and 60× on the
inpainter, with the same first-run compile cost as TensorRT and no rung between it and the CPU.
### 1.5 The Hexagon at every bit width · 2026-10-04
§1.1's Hexagon column is int8 calibrated on noise: timing only. This is the follow-up — every
model, every bit width the HTP offers, calibrated on real photographs and **scored on the tablet
itself** (ORT 1.29 + QNN 2.42, `htp_arch` 73), against the f32 model on the same inputs. The
"Form shipped" column is the files in `models/`, re-scored on the tablet after
`tools/quantise-models.sh` wrote them. The
calibration and scoring photographs are 800 from the public COCO val2017 set (CC-BY); the face
models' numbers are over the faces in them of at least 32 px. The tools are `tools/quantise-models.sh`
and the scratch harness described with it.
**What the HTP accepts.** fp16: nothing — every fp16 operator fails validation (3110), on QNN
2.42 and 2.50, with `htp_arch` and every `soc_model` tried; the fp16 rung stays off the table until
someone has Qualcomm's own SDK to say why. 4-bit weights (A8W4, A16W4): load, and wreck accuracy
(SCRFD finds 25–35% of f32's faces). What is left: **A8W8 (int8), A16W8 and A16W16**, all running
the whole graph. 16-bit activations cost about 3× int8's time, A16W16 about 4×.
| Model | ORT CPU f32 | Form shipped | Hexagon | On the tablet, against f32 | int8 for comparison |
|---|---|---|---|---|---|
| scrfd_500m / 2.5g / 10g | 17 / 56 / 198 ms | **A16W8** | 4.2 / 5.1 / 9.0 ms | 100% of faces found in every size band; keypoints 0.3–0.6% of the box | 94–95% of faces at 40–80 px |
| 2d106det (landmarks) | 2.8 ms | **A16W8** | 0.5 ms | 0.25 px in the 192 crop (eye points 0.20) | 1.5 px, and 29 partitions at 7.3 ms |
| yolo26n-seg | 90 ms | **A16W16**, tail in float | 12.9 ms | 98.2% of objects, mask IoU 0.994 | 74% (simulated) |
| yolo26s-sem-ade20k | 151 ms | **A16W16**, attention in float | 15 ms | 98.9% of cells agree on the class, TV 0.009 | 67% |
| migan-512 | 488 ms | **A16W16** | 87 ms | 41 dB from f32 in the fill (worst 1%: 30 dB) | 16 dB (simulated) |
| xfeat-1024 / 768 | 58 ms | **int8**, rewritten graph | 6.5 ms | panorama alignment 0.45 px from f32's — f32's own refit on 90% of its matches is 0.41 | — |
| mosaic-1408 (denoiser) | 1510 ms a tile | **A16W16**, rewritten graph | 95 ms a tile | 0.00 dB at every ISO; ≤ 0.09 dB with the noise scaled ×0.5–×4 | −4.7 to −9.2 dB |
| arcface_mbf (embedder) | 8.5 ms | f32, CPU | — | A16W16: cosine 0.9995, p1 0.9967 — misses §7's 0.999 gate | — |
| ocec, sgc (eyes) | 1, 1.7 ms | f32, CPU | — | sgc flips 1.45% of views even at A16W16; not worth a millisecond | — |
Four things the table needed that the f32 graphs did not have, all in `tools/htp_graph.py` and all
checked exact against the f32 graph before they are used:
- **Rank ≤ 5.** QNN's tensors stop at rank 5, and the denoiser packs the mosaic through a 6-D
reshape (6007 at compose). For one channel that reshape is `SpaceToDepth(2)`. XFeat's 8×8 unfold
is 224 Slices and 6-D Concats; it is `SpaceToDepth(8)` (736 nodes to 60).
- **No bilinear Resize at XFeat's sizes** (3110). A half-pixel bilinear resize between fixed sizes
is two constant matrices, so it is two MatMuls.
- **One scale per tensor.** The segmenter's output rows carry boxes in pixels beside scores in
0..1; quantised as one tensor the scores vanish. Everything from the Concat that builds the rows
stays float, on the CPU, where the top-300 selection costs nothing.
- **Float where the HTP's 16-bit arithmetic drifts.** The scene model's one attention block
(two MatMuls and a Softmax over 400 tokens) moved its agreement from 98.7% to 96.7%; it stays
float.
**ORT's CPU simulation of a QDQ graph is not the tablet.** It matched to the hundredth of a dB for
the denoiser and to rounding for the detectors, landmarks and XFeat, and it overstated MI-GAN by
27 dB and the scene model by three points. Every number above is the device's; a new form is not
measured until it has run there.
**XFeat's int8 loses keypoints and not the panorama.** 83% of f32's keypoints come back within
1.5 px; but over the twelve-frame `fixtures/pano/2025-08-05` sweep, the homographies fitted from
int8's matches land 0.45 px from f32's in the overlaps — the spread f32 shows against itself
(0.41).
---
## 2. The shape of the answer
@@ -146,7 +201,7 @@ winning:
| Platform | 1st | 2nd | 3rd | Floor |
|---|---|---|---|---|
| Android, Qualcomm with a Hexagon the shipped QNN skel covers (V68–V81) | QNN HTP, int8 model | ORT CPU, f32 model | — | tract |
| Android, Qualcomm with a Hexagon the shipped QNN skel covers (V68–V81) | QNN HTP, each model's quantised form (§1.5) | ORT CPU, f32 model | — | tract |
| Android, any other SoC | ORT CPU, f32 | — | — | tract |
| Linux / Windows, NVIDIA GPU | TensorRT, f32 model, fp16 engine | CUDA provider, f32 | ORT CPU, f32 | tract |
| Linux, AMD GPU with ROCm | MIGraphX, f32 model, fp16 program | ORT CPU, f32 | — | tract |
@@ -298,10 +353,10 @@ of which form they load:
| Form | Who produces it | When | Needed by |
|---|---|---|---|
| f32 ONNX, shape-fixed, **opset ≥ 13** | `tools/fix-face-model-shapes.sh`, `tools/export-seg-model.sh` | Release time, once | Every rung except Hexagon |
| int8 QDQ ONNX, per-channel, uint8 activations | `tools/quantise-models.sh` (new) | Release time, once, **calibrated on real photographs** | Hexagon |
| QDQ ONNX, per-channel — int8, A16W8 or A16W16 per model (§1.5) | `tools/quantise-models.sh` | Release time, once, **calibrated on real photographs**, scored on the tablet | Hexagon |
| TensorRT engine (`.engine`, per GPU architecture and TensorRT version) | The app, from the f32 file | First run on that device, in the background | TensorRT rung |
| MIGraphX program (`.mxr`, per GPU architecture, MIGraphX version and precision) | The app, from the f32 file | First run on that device, in the background | MIGraphX rung |
| QNN context binary | The app, from the int8 file | First run on that device, in the background | Hexagon rung |
| QNN context binary | The app, from the quantised file | First run on that device, in the background | Hexagon rung |
Two rules.
@@ -380,6 +435,11 @@ already the rule for the detector and because §5 is the gate on whether the int
enough to be *offered* at all. f32 on tract, ORT CPU, CUDA and TensorRT-f32 are one identity: the
same graph, the same arithmetic, differences at the last bit.
After §1.5 the Hexagon runs the detectors in **A16W8**, and that is a third spelling:
`scrfd_500m_a16+w600k_mbf` and its two siblings. Same rule, same reconciliation; a tablet that
indexed under `_i8` keeps those rows, and `FaceDetector::model_ids` answers "has this detector been
over this image" for all three forms.
**The embedder** is where comparability across devices is the whole point, and it is the one
model that no accelerator helps (§1.4). So: **the embedder runs in f32 on every rung.** On TensorRT
that means the embedder's engine is built without fp16 while the detector's is built with it; on
+5 -3
View File
@@ -467,9 +467,11 @@ the tablet. Three ways to make it viable, none built:
1. **Fill at a quarter of the resolution and upsample.** Sky and scree
tolerate it; twenty-odd tiles, about three minutes on the desktop CPU. A
background job with the outbox's patience, not an interactive one.
2. **int8 on the tablet's Hexagon through QNN**, where the plain-conv design
is the point and the whole graph should run in milliseconds. The setup
exists from the eye-state work; MI-GAN is a candidate for the same path.
2. **The tablet's Hexagon through QNN**, where the plain-conv design is the
point. Measured 2026-10-04 (inference.md §1.5): int8 changes the fill
(16 dB from f32's), so it ships with 16-bit activations and weights —
87 ms a tile against 488 ms on the tablet's CPU, the whole graph on the
NPU, 41 dB from f32 in the hole.
3. **A WGSL runtime for those six operators.** A project of its own, and
the only route that would make it interactive on the desktop.
+9 -2
View File
@@ -605,6 +605,12 @@ until deleted or renamed. Shipped and imported presets change only the operation
name, so a look applied to a corrected photograph keeps the correction; a copy or a saved
edit replaces everything in scope.
The user's presets sync between devices through each library they open, as one file beside the
camera profiles. Each name merges on its own against what the last exchange left both sides
holding, so presets added on two devices both survive, a deletion on one reaches the other rather
than being restored by it, and an edit outlives a deletion made elsewhere. The write is
conditional on the server's copy, so two devices exchanging at once cannot save over each other.
**FR-DEV-7 — Before/after.** Compare current edit state against the unedited original or against
a chosen history state.
@@ -2771,7 +2777,8 @@ every render path would have to remember to call it. They travel with the decode
matrix does.
*What it costs.* Every DNG with an embedded profile renders differently; previews refresh only when
rendered again; tablet and desktop release together. The profiles directory does not sync yet.
rendered again; tablet and desktop release together. The profiles directory syncs through the
library's derived folder (camera-profiles.md §13).
### D21 — DNG reference tone for raws · **DECIDED 2026-10-03**
@@ -2794,7 +2801,7 @@ curve a power of 1.5/1.4 about grey (`REFERENCE_CONTRAST` is where the curve is
brightness needs nothing: baseline exposure and the curve together land where the earlier exports do. The
profile's look strength, vibrance and saturation bought nothing measurable on those exports. The
user chose to change every photograph rather than keep edited ones on the old rendering. The
fitting tools live outside the repository (`darkroom-lrfit`).
fitting tools live outside the repository (`darkroom-lrfit`). *Amended 2026-10-04:* the look strength now defaults to 0. It scored the same at 100, 50 and 0 (held-out MSE 140, 140, 143) and the rendering is 9 % more colourful without it — the table desaturates near-neutral tones, where the default was short of those exports; the user chose more colour.
*Amended earlier the same day:* the default was **not** decided. The measurement below was against
Lightroom previews of photographs carrying the user's Lightroom edits — a house look of HSL
+82
View File
@@ -0,0 +1,82 @@
# DarkRoom — Sensor health: a dated defect map per body
**Status:** Spike · 2026-10-04 · not built
**Companion to:** [requirements.md](requirements.md) FR-RAW-3, [catalog.md](catalog.md)
A sensor gains defective photosites as it ages, and a photosite that has gone bad does not recover.
This records, per camera body, which photosites are defective and since when, so that the library
can show how a sensor has aged and the hot-pixel repair can fix the defects a body is known to have
rather than only those that stand out in the frame at hand.
What exists is the measuring tool: `Demosaicer::find_hot_pixels` (the photosites the repair pass
would replace, without replacing them) and `core/dr-gpu/examples/sensor_scan.rs`, which prints them
per frame and, with `--probe`, reads a list of coordinates back out of each frame. The rest of this
document is what a spike with them on the 6D found, and the design it argues for.
---
## 1. What is wanted
- **Settings → Bodies**, one entry per body, with a graph of the defective share over time in two
series: photosites (the raw mosaic) and 2×2 cells holding at least one defective photosite (what
reaches a pixel of the output).
- **A dated defect map** per body, synced with the library like any other catalog data, and
cumulative: an entry is never removed.
- **The repair reads the map** whose date is nearest the frame's, and fixes every defect the body had
by then, whether or not it stands out in that frame.
- One body per model for now: the catalog stores `make model`, not a body serial.
## 2. What the spike found (6D, 2026-10-03)
53 CR2s, up to four per quarter at the highest ISO of a day, 2015 to 2026. The library holds almost
no 6D raws from 2016–2022, so onsets in that span are dated to the span, not the year. `sensor_scan`
ran at 0.8 s per frame, decode included.
**Persistence separates the sensor from the scene.** 4,179 photosites were flagged at least once;
3,554 on one day only (stars, glints, noise). A defect is a photosite that keeps coming back.
**A frame that does not flag a photosite is not evidence it was clean.** The repair's test is
relative to the neighbourhood, so a defect in a lit area does not stand out. Confirmed defects were
flagged in a median 20 % of the frames after their first sighting. Only a frame whose neighbourhood
at that photosite is dark counts, either way.
**The 6D hides some defects itself at high ISO.** (2517, 3172) reads 8,000–13,000 over neighbours
near 200 at ISO 2000–5000, and does not stand out at all at ISO 6400–12800 (82 over 149 on
2025-03-15). The camera appears to map out photosites it knows at those gains. So evidence for this
body comes from ISO ≤ 5000; the cut-off must be learned per body, not fixed.
**Long exposures light everything.** A 9.8 s frame saturated every candidate; it confirms, it does
not date.
**The curve.** Counting a photosite as defective from the first frame where it stands out, provided
it stands out in at least 60 % of the observable frames after that (32 defects; 31 with a clean
observable frame before onset to bracket it):
| Year | Defects | Share of photosites |
|---|---|---|
| 2015 | 2 | 0.1 ppm |
| 2016–2021 | 2 | 0.1 ppm |
| 2022 | 7 | 0.3 ppm |
| 2023 | 26 | 1.3 ppm |
| 2026 | 32 | 1.6 ppm |
(2517, 3172) is the shape every entry should have: clean at ISO 800–1000 in 2015 and at ISO 100–200
in 2016, then 338 over 72 at ISO 100 on 2022-08-13 and in every comparable frame since.
Weak defects exist too, about twice their neighbours ((1814, 3039)); the blind repair misses them in
most frames. A known map would catch them.
## 3. Design it argues for
- **Evidence per frame, per known photosite**: observable (dark neighbourhood, ISO inside the
body's band) and, if so, lit or clean. Not just the frame's flagged list.
- **A map entry is a bracket**: last clean observation, first lit observation, strength, kind. Onset
lies between the two; the graph plots it at the first, and can show the bracket.
- **The repair**: every entry whose first lit date is on or before the frame's capture date; for an
entry whose bracket contains the date, probe the photosite in the frame itself.
- **Storage and sync**: catalog tables created on first use (as `albums` does), so no schema bump
breaks an older peer. They travel in the snapshot by default. Merge is a set union of photosites
per body, the earlier first-lit and the later last-clean winning, which makes it commutative and
keeps the map cumulative.
- **Work**: a sample, not the library. Frames are picked for what they can reveal (dark, mid ISO,
long exposures), a few per body per month, and after the first pass only new imports are read.
File diff suppressed because one or more lines are too long
+2 -2
View File
@@ -287,7 +287,6 @@ $LOCALAPPDATA\Programs\DarkRoom\
models\
scrfd_500m_640.onnx scrfd_2.5g_640.onnx scrfd_10g_640.onnx arcface_mbf_b1.onnx
2d106det_b1.onnx ocec_s_b1.onnx sgc_l_48_b1.onnx
scrfd_500m_640.int8.onnx scrfd_2.5g_640.int8.onnx scrfd_10g_640.int8.onnx
yolo26s-sem-ade20k.onnx yolo26s-sem-ade20k.classes.json categories.txt
migan-512.onnx
manual\
@@ -297,7 +296,8 @@ $LOCALAPPDATA\Programs\DarkRoom\
```
Plus a Start Menu shortcut, and nothing on the desktop unless the user ticks it. The models are
the same ten files the APK bundles and the PKGBUILD installs; `models\` beside the executable is
the f32 files the APK bundles and the PKGBUILD installs — not the APK's quantised siblings, which
only a Hexagon runs; `models\` beside the executable is
where §3.2's lookup finds them. **No `LICENSE` yet**: the repository has no licence file at its
root (the Arch package points at the system's shared GPL text), so the installer has no licence
page until one is added — a one-file change, and the `.nsi` says where the page then goes. The face weights carry the research-only grant that
+33 -15
View File
@@ -207,24 +207,42 @@ the sensor recorded.
### AI denoise
For a photograph taken in poor light at a high ISO. `AI Denoise`, in the
Detail group, replaces how the camera's raw data is turned into colour: a
network trained on this library's own photographs removes the noise and
the blotches of colour that come with it, while keeping the fine detail.
Look at it at 1:1, where noise lives.
How every raw is developed. `AI Denoise`, at the top of the Adjust panel,
replaces how the camera's raw data is turned into colour: a network trained
on this library's own photographs removes the noise and the blotches of
colour that come with it, while keeping the fine detail. Look at it at
1:1, where noise lives.
Switch it on with `Apply`. The photograph keeps showing as it was while
the network works, with its progress in the bar at the top, and changes
when it is done — a few seconds on a computer with a graphics card, about
fifteen on its processor alone, longer on the tablet. `Keep grain` puts
back some of what was removed, as grain without colour, for a picture that
does not look too smooth.
`Method` chooses how:
![An ISO 8000 night frame at 1:1, AI Denoise switched on, then some grain kept](media/develop-denoise.gif)
- `Best`, the default: two networks, one for smooth areas and one for
edges, blended where each is better. The cleanest skies and the sharpest
lettering, and the slowest.
- `Medium`: one network taught by `Best`. Nearly as clean in smooth areas,
a little softer on hard edges, in about a third of the time.
- `Fast`: a smaller one, taught the same way. Visibly noisier at very high
ISO than the other two, but still far cleaner than none, and quick.
- `Bilinear`: the camera's ordinary conversion, noise and all.
| Before | After |
The photograph shows the camera's ordinary conversion while the network
works, with its progress in the bar at the top, and changes when it is
done — on a laptop's graphics card, about two and a half seconds for a
20-megapixel photograph with `Best` and under one with the other two;
longer on a processor alone or on the tablet. The first photograph after
installing waits a few minutes more while the graphics card prepares each
network, once. The result is kept, so a photograph opened again,
or exported, does not wait a second time, and switching back to a method
already used is quick.
`Strength` eases it off: below 100 % it puts back some of what was removed,
as grain without colour, for a picture that does not look too smooth.
The lamp and railing of a night frame at ISO 8000, at 1:1, by each method:
| Bilinear | Fast |
|---|---|
| ![The railing and the lamp at ISO 8000, as the camera recorded them](media/develop-denoise-before.png) | ![The same, with AI Denoise](media/develop-denoise-after.png) |
| ![The railing and the lamp at ISO 8000, as the camera recorded them](media/develop-denoise-bilinear.png) | ![The same, with the Fast network](media/develop-denoise-fast.png) |
| **Medium** | **Best** |
| ![The same, with the Medium network](media/develop-denoise-medium.png) | ![The same, with the Best network](media/develop-denoise-best.png) |
It works on raw files from any camera with the usual colour pattern of
red, green and blue squares — not on JPEGs, and not yet on Fujifilm's
@@ -232,7 +250,7 @@ X-Trans. How noisy the camera is at each ISO was measured for the Canon
EOS 6D; for other cameras it is read from a DNG's own figures or
estimated from the photograph, and the finished job in the activity list
says which. An export uses
it whenever the photograph has it switched on.
the method the photograph has.
### Moving between photographs
+33 -15
View File
@@ -289,20 +289,38 @@ drawn as hard-edged blocks rather than smoothed, so what you see is what
the sensor recorded.</p>
<figure><img loading="lazy" src="media/develop-zoom.gif" alt="Zooming to 1:1 with a double-click, panning, then further in with the wheel"><figcaption>Zooming to 1:1 with a double-click, panning, then further in with the wheel</figcaption></figure>
<h3 id="ai-denoise">AI denoise</h3>
<p>For a photograph taken in poor light at a high ISO. <code>AI Denoise</code>, in the
Detail group, replaces how the camera's raw data is turned into colour: a
network trained on this library's own photographs removes the noise and
the blotches of colour that come with it, while keeping the fine detail.
Look at it at 1:1, where noise lives.</p>
<p>Switch it on with <code>Apply</code>. The photograph keeps showing as it was while
the network works, with its progress in the bar at the top, and changes
when it is done — a few seconds on a computer with a graphics card, about
fifteen on its processor alone, longer on the tablet. <code>Keep grain</code> puts
back some of what was removed, as grain without colour, for a picture that
does not look too smooth.</p>
<figure><img loading="lazy" src="media/develop-denoise.gif" alt="An ISO 8000 night frame at 1:1, AI Denoise switched on, then some grain kept"><figcaption>An ISO 8000 night frame at 1:1, AI Denoise switched on, then some grain kept</figcaption></figure>
<table><thead><tr><th>Before</th><th>After</th></tr></thead><tbody>
<tr><td><img src="media/develop-denoise-before.png" alt="The railing and the lamp at ISO 8000, as the camera recorded them" /></td><td><img src="media/develop-denoise-after.png" alt="The same, with AI Denoise" /></td></tr>
<p>How every raw is developed. <code>AI Denoise</code>, at the top of the Adjust panel,
replaces how the camera's raw data is turned into colour: a network trained
on this library's own photographs removes the noise and the blotches of
colour that come with it, while keeping the fine detail. Look at it at
1:1, where noise lives.</p>
<p><code>Method</code> chooses how:</p>
<ul>
<li><code>Best</code>, the default: two networks, one for smooth areas and one for
edges, blended where each is better. The cleanest skies and the sharpest
lettering, and the slowest.</li>
<li><code>Medium</code>: one network taught by <code>Best</code>. Nearly as clean in smooth areas,
a little softer on hard edges, in about a third of the time.</li>
<li><code>Fast</code>: a smaller one, taught the same way. Visibly noisier at very high
ISO than the other two, but still far cleaner than none, and quick.</li>
<li><code>Bilinear</code>: the camera's ordinary conversion, noise and all.</li>
</ul>
<p>The photograph shows the camera's ordinary conversion while the network
works, with its progress in the bar at the top, and changes when it is
done — on a laptop's graphics card, about two and a half seconds for a
20-megapixel photograph with <code>Best</code> and under one with the other two;
longer on a processor alone or on the tablet. The first photograph after
installing waits a few minutes more while the graphics card prepares each
network, once. The result is kept, so a photograph opened again,
or exported, does not wait a second time, and switching back to a method
already used is quick.
<code>Strength</code> eases it off: below 100 % it puts back some of what was removed,
as grain without colour, for a picture that does not look too smooth.</p>
<p>The lamp and railing of a night frame at ISO 8000, at 1:1, by each method:</p>
<table><thead><tr><th>Bilinear</th><th>Fast</th></tr></thead><tbody>
<tr><td><img src="media/develop-denoise-bilinear.png" alt="The railing and the lamp at ISO 8000, as the camera recorded them" /></td><td><img src="media/develop-denoise-fast.png" alt="The same, with the Fast network" /></td></tr>
<tr><td><strong>Medium</strong></td><td><strong>Best</strong></td></tr>
<tr><td><img src="media/develop-denoise-medium.png" alt="The same, with the Medium network" /></td><td><img src="media/develop-denoise-best.png" alt="The same, with the Best network" /></td></tr>
</tbody></table>
<p>It works on raw files from any camera with the usual colour pattern of
red, green and blue squares — not on JPEGs, and not yet on Fujifilm's
@@ -310,7 +328,7 @@ X-Trans. How noisy the camera is at each ISO was measured for the Canon
EOS 6D; for other cameras it is read from a DNG's own figures or
estimated from the photograph, and the finished job in the activity list
says which. An export uses
it whenever the photograph has it switched on.</p>
the method the photograph has.</p>
<h3 id="moving-between-photographs">Moving between photographs</h3>
<p>The roll along the foot of the canvas holds the photographs the grid was
showing; click one to open it. The right arrow, <code>D</code> or space opens the next,
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
+16 -4
View File
@@ -14,6 +14,13 @@ Both come from `https://huggingface.co/Ultralytics/YOLO26`. The face weights in
The keypoint weights in `keypoints/` and the border filler in `inpaint/` are
the other two, and the easiest — see the last two sections.
Every quantised sibling — `*.int8.onnx`, `*.a16w8.onnx`, `*.a16w16.onnx`, the
forms the tablet's Hexagon runs (`tools/quantise-models.sh`) — is the same
weights rounded, and carries exactly the grant of the file it was made from.
What calibration adds is one minimum and maximum per tensor: from public COCO
val2017 photographs (CC-BY 4.0) for the image models, and for the denoiser from
the same training tiles its weights were learned from. No image is in the files.
## The grant
**Ultralytics releases YOLO under AGPL-3.0**, and the weights carry the same
@@ -127,11 +134,16 @@ declined, and the InsightFace grant of D13).
| File | Source | Trained on | Used by |
|---|---|---|---|
| `denoise/mosaic-1408.onnx` | trained from scratch in the `darkroom-denoise` repository (2026-10-03, run `m2`, 60 000 steps) | 427 of the maintainer's own base-ISO Canon EOS 6D raws, with the 6D's measured noise added | the learned demosaic and denoise (FR-DEV-3g) |
| `denoise/mosaic-best-1408.onnx` | trained in the `darkroom-denoise` repository (2026-10-04, run `final`, 30 000 steps, from the experts of runs `m2` and `edges-100`) | 1,701 of the maintainer's own base-ISO raws and 6,000 synthetic scenes the repository draws itself, with the Canon EOS 6D's measured noise added | the learned demosaic and denoise, Best (FR-DEV-3g) |
| `denoise/mosaic-medium-1408.onnx` | distilled from `final` in the same repository (2026-10-04, run `student-m`, 20 000 steps, from `m2`) | the same | Medium |
| `denoise/mosaic-fast-1408.onnx` | distilled from `final` (2026-10-04, run `student-s`, 30 000 steps, from scratch) | the same | Fast |
A U-Net of plain 3×3 convolutions, ReLU, strided and transposed
U-Nets of plain 3×3 convolutions, ReLU, strided and transposed
convolutions and additive skips — no third-party architecture code or
weights — at a fixed `1×1×1408×1408` for `mosaic` and `sigma`, exported by
`python -m denoise.export` in `darkroom-denoise`. Trained only on
weights — and, for Best, two of them blended per photosite by a small gate
network of the same parts. Each at a fixed `1×1×1408×1408` for `mosaic`
and `sigma`, exported by `python -m denoise.export` in `darkroom-denoise`;
the `.a16w16.onnx` siblings are the same networks quantised for the
Hexagon by `tools/quantise-models.sh`. Trained only on
photographs the maintainer owns, so the weights carry no grant but the
project's own: GPL-3.0-or-later, like the code (denoise.md §10).
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
+14 -8
View File
@@ -28,15 +28,21 @@ every reader treats "never read" as unknown, never as closed. They are found in
directory as the pair, so a hand-placed pair does not pick up a package's eye models from a
directory it otherwise outranks.
scrfd_500m_640.int8.onnx 0.8 MB the same three, in the form the Hexagon NPU takes
scrfd_2.5g_640.int8.onnx 0.9 MB (docs/dev/inference.md §5) — opset 17, per-channel int8
scrfd_10g_640.int8.onnx 4.3 MB weights, uint8 activations, calibrated on 96 photographs
scrfd_500m_640.a16w8.onnx the same three, in the form the Hexagon NPU takes
scrfd_2.5g_640.a16w8.onnx (docs/dev/inference.md §1.5) — 16-bit activations,
scrfd_10g_640.a16w8.onnx per-channel 8-bit weights, calibrated on 300 photographs
2d106det_b1.a16w8.onnx the landmarks, likewise
The int8 files are **derived** by `tools/quantise-models.sh` from the f32 ones beside them and
travel with them: the engine loads the `.int8.onnx` sibling when the device's backend wants it and
the canonical file otherwise, and a library indexed on the int8 form records it as a different
detector (`scrfd_500m_i8+w600k_mbf`), because it finds a different set of faces. Every other
platform ignores them. The embedder has no int8 form and never will (§7 of the same document).
These are **derived** by `tools/quantise-models.sh` from the f32 files beside them and travel with
them: the engine loads the `.a16w8.onnx` sibling when the device's backend wants it and the
canonical file otherwise, and a library indexed on that form records it as a different detector
(`scrfd_500m_a16+w600k_mbf`), because it finds a different set of faces. Every other platform
ignores them, and the Windows installer leaves them out. The embedder and the eye classifiers have
no quantised form: the embedder's vectors must compare across devices (inference.md §7), and the
classifiers cost a millisecond on the CPU.
The detectors were int8 until §1.5 measured them on the tablet: int8 found 94–95% of f32's faces
at 40–80 px, A16W8 all of them. A tablet that indexed under the int8 ids keeps those rows.
**A clone without git-lfs gets a ~130-byte pointer where each model should be.** Both packagers check
for exactly that and refuse, rather than shipping the pointer and failing inside tract on the user's
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
+12 -9
View File
@@ -4,7 +4,7 @@
# makes `makepkg -si` in this directory install what you are actually working
# on. Swap `source` for a tagged tarball when there is something to release.
pkgname=darkroom
pkgver=0.21.0
pkgver=0.22.1
# Back to 1 with the version: a new pkgver is a new archive name, so there is
# nothing for makepkg to reuse and nothing for a release number to disambiguate.
pkgrel=1
@@ -124,12 +124,15 @@ package() {
fi
install -Dm644 "${_src}" "${pkgdir}/usr/share/darkroom/models/migan-512.onnx"
# The learned demosaic and denoise (the project's own weights, GPL —
# models/LICENCE.md). Same pointer check, same directory.
_src="models/denoise/mosaic-1408.onnx"
if [[ "$(stat -c%s "${_src}")" -lt 100000 ]]; then
echo "error: the denoise model is an LFS pointer — run: git lfs pull" >&2
return 1
fi
install -Dm644 "${_src}" "${pkgdir}/usr/share/darkroom/models/mosaic-1408.onnx"
# The learned demosaic and denoise, one network per method (the project's
# own weights, GPL — models/LICENCE.md). Same pointer check, same
# directory.
for _net in fast medium best; do
_src="models/denoise/mosaic-${_net}-1408.onnx"
if [[ "$(stat -c%s "${_src}")" -lt 100000 ]]; then
echo "error: the ${_net} denoise model is an LFS pointer — run: git lfs pull" >&2
return 1
fi
install -Dm644 "${_src}" "${pkgdir}/usr/share/darkroom/models/mosaic-${_net}-1408.onnx"
done
}
+202
View File
@@ -0,0 +1,202 @@
"""Exact rewrites that make a graph one QNN's HTP can hold (docs/dev/inference.md §1.5).
Each rewrite spells the same arithmetic in operators the Hexagon runs, and
`rewrite` checks the result against the input graph on ONNX Runtime's CPU
before returning it — a rewrite that changes an output by more than float
rounding is refused, not shipped.
- **6-D Bayer pack → SpaceToDepth(2).** The denoiser packs the mosaic with
`Reshape(1,C,H/2,2,W/2,2) → Transpose(0,1,3,5,2,4) → Reshape(1,4C,…)`.
QNN's tensors stop at rank 5 (error 6007 at compose); for C = 1 that
sequence *is* SpaceToDepth.
- **Unfold → SpaceToDepth(8).** XFeat spells its 8×8 unfold as 224 Slices,
225 Transposes and two 6-D Concats; it is SpaceToDepth(8) of the
normalised image, 736 nodes down to 60.
- **Computed reshape targets → constants.** `Shape → Slice → Concat` feeding
a Reshape, where shape inference proves the answer.
- **Bilinear Resize → two MatMuls.** The HTP refuses `ResizeBilinear` at the
sizes XFeat uses (3110); a half-pixel bilinear resize between fixed sizes
is `X · Rxᵀ` then `Ry ·`, with ONNX's own edge-clamped weights.
- **InstanceNormalization → primitives** is not here: it is exact, but the
two full-image reductions it becomes cost the HTP more than it saves. XFeat
keeps its normalisation in int8, which the HTP takes.
"""
import numpy as np
import onnx
import onnxruntime as ort
from onnx import helper, numpy_helper, shape_inference
TOLERANCE = 1e-4 # relative to each output's largest magnitude
def _prune(g):
"""Drop nodes nobody reads and initialisers nobody uses."""
while True:
used = {i for n in g.node for i in n.input} | {o.name for o in g.output}
keep = [n for n in g.node if any(o in used for o in n.output)]
if len(keep) == len(g.node):
break
del g.node[:]
g.node.extend(keep)
used = {i for n in g.node for i in n.input}
keep = [i for i in g.initializer if i.name in used]
del g.initializer[:]
g.initializer.extend(keep)
def bayer_pack(m):
g = m.graph
prod = {o: n for n in g.node for o in n.output}
users = {}
for n in g.node:
for i in n.input:
users.setdefault(i, []).append(n)
swaps = {}
for n in g.node:
if n.op_type != "Transpose":
continue
perm = next(a.ints for a in n.attribute if a.name == "perm")
r1, u = prod.get(n.input[0]), users.get(n.output[0], [])
if list(perm) != [0, 1, 3, 5, 2, 4] or not r1 or r1.op_type != "Reshape":
continue
if len(u) != 1 or u[0].op_type != "Reshape":
continue
s2d = helper.make_node("SpaceToDepth", [r1.input[0]], [u[0].output[0]], name=n.name + "_s2d", blocksize=2)
swaps[id(r1)] = s2d
swaps[id(n)] = swaps[id(u[0])] = None
nodes = [swaps.get(id(n), n) for n in g.node if swaps.get(id(n), n) is not None]
del g.node[:]
g.node.extend(nodes)
return m
def fold_reshapes(m):
m = shape_inference.infer_shapes(m)
g = m.graph
vi = {v.name: v for v in list(g.value_info) + list(g.output) + list(g.input)}
inits = {i.name for i in g.initializer}
dims = lambda t: [d.dim_value for d in vi[t].type.tensor_type.shape.dim] if t in vi else []
for n in g.node:
if n.op_type != "Reshape" or n.input[1] in inits:
continue
shape = dims(n.output[0])
if not shape or 0 in shape:
src = dims(n.input[0])
if len(src) != 5 or 0 in src: # (1,C,k,H,W) -> (1,C·k,H,W), the denoiser's tile
continue
shape = [src[0], src[1] * src[2], src[3], src[4]]
name = n.output[0] + "_shape"
g.initializer.append(numpy_helper.from_array(np.array(shape, np.int64), name))
n.input[1] = name
_prune(g)
del g.value_info[:]
return m
def unfold(m, block=8):
"""Replace the Slice/Transpose/Concat region ending in the Reshape that
produces the (1, block², H/block, W/block) tensor with SpaceToDepth."""
g = m.graph
prod = {o: n for n in g.node for o in n.output}
region_ops = ("Slice", "Transpose", "Concat", "Unsqueeze", "Reshape")
inits = {i.name for i in g.initializer}
m_inf = shape_inference.infer_shapes(m)
vi = {v.name: [d.dim_value for d in v.type.tensor_type.shape.dim] for v in m_inf.graph.value_info}
for end in g.node:
out = vi.get(end.output[0], [])
if end.op_type != "Reshape" or len(out) != 4 or out[1] != block * block:
continue
seen, stack, leaves = set(), [end.input[0]], set()
while stack:
t = stack.pop()
if t in seen or t in inits:
continue
seen.add(t)
n = prod.get(t)
if n is not None and n.op_type in region_ops:
stack += list(n.input)
else:
leaves.add(t)
if len(leaves) != 1 or len(seen) < 100: # the hand-written unfold, not an ordinary reshape
continue
region = {id(end)} | {id(prod[t]) for t in seen if t in prod and prod[t].op_type in region_ops}
nodes = []
for n in g.node:
if id(n) in region:
if n is end:
nodes.append(helper.make_node("SpaceToDepth", [next(iter(leaves))], [end.output[0]],
name=end.name + "_s2d", blocksize=block))
continue
nodes.append(n)
del g.node[:]
g.node.extend(nodes)
_prune(g)
del g.value_info[:]
return m
return m
def _bilinear(n_in, n_out):
r = np.zeros((n_out, n_in), np.float32)
for o in range(n_out):
x = min(max((o + 0.5) * n_in / n_out - 0.5, 0), n_in - 1)
i0 = int(np.floor(x))
f = x - i0
r[o, i0] += 1 - f
r[o, min(i0 + 1, n_in - 1)] += f
return r
def resize_matmul(m):
"""Every linear, half-pixel Resize between fixed NCHW sizes."""
m = shape_inference.infer_shapes(m)
g = m.graph
vi = {v.name: [d.dim_value for d in v.type.tensor_type.shape.dim] for v in list(g.value_info) + list(g.output)}
nodes = []
for n in g.node:
a = {x.name: helper.get_attribute_value(x) for x in n.attribute}
ok = (n.op_type == "Resize" and a.get("mode") == b"linear"
and a.get("coordinate_transformation_mode", b"half_pixel") == b"half_pixel"
and len(vi.get(n.input[0], [])) == 4 and len(vi.get(n.output[0], [])) == 4)
if not ok:
nodes.append(n)
continue
(_, _, h, w), (_, _, h2, w2) = vi[n.input[0]], vi[n.output[0]]
t = n.name
rx = numpy_helper.from_array(_bilinear(w, w2).T.copy(), t + "_rxT")
ry = numpy_helper.from_array(_bilinear(h, h2).T.copy(), t + "_ryT")
g.initializer.extend([rx, ry])
nodes += [
helper.make_node("MatMul", [n.input[0], rx.name], [t + "_w"], name=t + "_mw"),
helper.make_node("Transpose", [t + "_w"], [t + "_t"], name=t + "_t1", perm=[0, 1, 3, 2]),
helper.make_node("MatMul", [t + "_t", ry.name], [t + "_h"], name=t + "_mh"),
helper.make_node("Transpose", [t + "_h"], [n.output[0]], name=t + "_t2", perm=[0, 1, 3, 2]),
]
del g.node[:]
g.node.extend(nodes)
_prune(g)
del g.value_info[:]
return m
REWRITES = {"bayer": [bayer_pack, fold_reshapes], "unfold": [unfold], "resize": [resize_matmul]}
def rewrite(path, names):
"""The graph at `path` with the named rewrites applied, checked exact."""
m = onnx.load(path)
before = len(m.graph.node)
for name in names:
for step in REWRITES[name]:
m = step(m)
onnx.checker.check_model(m)
a = ort.InferenceSession(path, providers=["CPUExecutionProvider"])
b = ort.InferenceSession(m.SerializeToString(), providers=["CPUExecutionProvider"])
rng = np.random.default_rng(3)
feed = {i.name: rng.random([d if isinstance(d, int) else 1 for d in i.shape], dtype=np.float32) for i in a.get_inputs()}
for o, x, y in zip(a.get_outputs(), a.run(None, feed), b.run(None, feed)):
err = float(np.abs(x - y).max()) / max(float(np.abs(x).max()), 1e-6)
if err > TOLERANCE:
raise SystemExit(f"{path}: rewrite {names} moved output {o.name} by {err:.2e} (relative)")
print(f" rewrites {'+'.join(names)}: {before} -> {len(m.graph.node)} nodes, exact")
return m
+35 -26
View File
@@ -842,13 +842,21 @@ def develop_zoom():
pause(1.2)
@scene(media=['develop-denoise.gif', 'develop-denoise-before.png', 'develop-denoise-after.png'],
# The methods in the order the scene visits them: `Best` is what the
# photograph opens with, then each smaller network, then none.
DENOISE_METHODS = ['Best', 'Medium', 'Fast', 'Bilinear']
DENOISE_CLOSE_UP = 560 # pixels of canvas, square, around the lamp at 1:1
@scene(media=[f'develop-denoise-{m.lower()}.png' for m in DENOISE_METHODS],
sources=DEVELOP_SRC + ['ui/dr-ui/src/develop/denoise.rs', 'core/dr-denoise/**',
'core/dr-gpu/src/grain.rs', 'models/denoise/**'])
'core/dr-pipeline/src/learned_denoise.rs', 'models/denoise/**'])
def develop_denoise():
"""A night frame at ISO 8000 at 1:1, AI Denoise switched on and landed,
then some grain kept. Waits for the network rather than for a fixed
time: on the CPU it takes several times what it does on a GPU."""
"""The lit lamp and railing of an ISO 8000 night frame at 1:1, once per
AI Denoise method, each cut to the same square of the canvas. Waits for
each network rather than for a fixed time: on the CPU, which the demo
profile uses, Best takes several times what Fast does."""
mark = log_size()
at_develop(DENOISE)
a = dr.photo(0.45, 0.55) # the lit lamp, the railing and the skyline over the water
dr.move(*a)
@@ -857,32 +865,31 @@ def develop_denoise():
pause(1.5)
group('Detail')
in_column('AI Denoise@Text')
shot('develop-denoise-before')
rec('develop-denoise')
pause(0.8)
mark = log_size()
dr.click(*denoise_switch())
t0 = time.time()
while time.time() - t0 < 300 and not denoise_landed(mark):
pause(0.5)
pause(1.5)
shot('develop-denoise-after')
slide('Keep grain', 60)
pause(2.0)
cut()
for method in DENOISE_METHODS:
if method != 'Best':
mark = log_size()
dr.click(*in_column(f'{method}@RadioButton'))
if method != 'Bilinear':
t0 = time.time()
while time.time() - t0 < 600 and not denoise_landed(mark):
pause(0.5)
pause(1.5)
close_up(f'develop-denoise-{method.lower()}', a)
undo_all()
dr.move(*a)
dr.x('click', '--repeat', 2, '--delay', 80, 1)
pause(1.2)
def denoise_switch():
"""The `Apply` box under the AI Denoise heading — the lens profile's
switch is also called Apply, so it is found by where it sits."""
head = dr.matches('AI Denoise@Text', within=column())[0]
below = [e for e in dr.matches('Apply@CheckBox', within=column()) if e['y'] > head['y']]
e = min(below, key=lambda e: e['y'])
return int(e['x'] + e['w'] / 2), int(e['y'] + e['h'] / 2)
def close_up(name, p, size=DENOISE_CLOSE_UP):
"""A square of the canvas centred on `p`, kept inside the canvas."""
shot(name)
x0, y0, x1, y1 = dr.rect('id:canvas-image')
half = size // 2
cx = min(max(p[0], x0 + half), x1 - half)
cy = min(max(p[1], y0 + half), y1 - half)
subprocess.run(['mogrify', '-crop', f'{size}x{size}+{cx - half}+{cy - half}', '+repage',
f'{OUT}/{name}.png'], check=True)
def log_size():
@@ -899,7 +906,9 @@ def denoise_landed(since):
try:
with open(f'{dr.HOME}/app.log', 'rb') as f:
f.seek(since)
return b'learned denoise:' in f.read()
# The result's line, "learned denoise: W×H on …", not the
# repair's, which comes first.
return re.search(rb'learned denoise: \d+\xc3\x97', f.read()) is not None
except OSError:
return False
+231 -112
View File
@@ -3,143 +3,262 @@
import glob
import os
import sys
from pathlib import Path
import numpy as np
import onnx
import onnxruntime as ort
from onnx import version_converter
from onnxruntime.quantization import (
CalibrationDataReader,
CalibrationMethod,
QuantFormat,
QuantType,
quantize_static,
)
from onnxruntime.quantization.calibrate import create_calibrator
from onnxruntime.quantization.calibrate import save_tensors_data
from onnxruntime.quantization.shape_inference import quant_pre_process
from pathlib import Path
from PIL import Image, ImageOps
from onnxruntime.quantization import CalibrationDataReader, CalibrationMethod, QuantType, quantize_static
from onnxruntime.quantization.calibrate import create_calibrator, save_tensors_data
from onnxruntime.quantization.execution_providers.qnn import get_qnn_qdq_config
from PIL import Image
PHOTOS = 64 # enough for a stable range; more only costs time
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import htp_graph # noqa: E402
MODELS = Path(__file__).resolve().parents[1] / "models"
PHOTOS = 300 # calibration photographs; face crops and keypoints come from fewer
CHUNK = 4 # inputs whose activations are held at once (scrfd_10g: ~1 GB each)
Q = QuantType
FORMS = {"int8": (Q.QUInt8, Q.QInt8), "a16w8": (Q.QUInt16, Q.QInt8), "a16w16": (Q.QUInt16, Q.QInt16)}
# Per model: where it lives, the form `Rung::form` gives its role, the exact
# rewrites its graph needs, and nodes that stay float on the CPU because one
# scale cannot serve the tensor (the segmenter's rows: boxes in pixels beside
# scores in 0..1) or because the HTP's 16-bit arithmetic drifts there (the
# scene model's attention). Measured, inference.md §1.5.
TABLE = {
"scrfd_500m_640": dict(dir="face", form="a16w8", feed="scrfd"),
"scrfd_2.5g_640": dict(dir="face", form="a16w8", feed="scrfd"),
"scrfd_10g_640": dict(dir="face", form="a16w8", feed="scrfd"),
"2d106det_b1": dict(dir="face", form="a16w8", feed="landmarks"),
"yolo26n-seg": dict(dir="segment", form="a16w16", feed="yolo", float_from="/model.23/Concat_4"),
"yolo26s-sem-ade20k": dict(dir="scene", form="a16w16", feed="yolo",
float_nodes=["/model.10/m/m.0/attn/MatMul", "/model.10/m/m.0/attn/Softmax",
"/model.10/m/m.0/attn/MatMul_1"]),
"migan-512": dict(dir="inpaint", form="a16w16", feed="migan"),
"xfeat-1024": dict(dir="keypoints", form="int8", feed="xfeat", rewrites=["unfold", "resize"]),
"xfeat-768": dict(dir="keypoints", form="int8", feed="xfeat", rewrites=["unfold", "resize"]),
"mosaic-fast-1408": dict(dir="denoise", form="a16w16", feed=None, rewrites=["bayer"]),
"mosaic-medium-1408": dict(dir="denoise", form="a16w16", feed=None, rewrites=["bayer"]),
# Exported with its packs as SpaceToDepth already; the rewrite finds
# nothing to do.
"mosaic-best-1408": dict(dir="denoise", form="a16w16", feed=None, rewrites=["bayer"]),
}
def letterbox(img, edge, pad, norm):
"""The app's Letterbox::sample: fit the long side to `edge`, centre, pad."""
img = ImageOps.exif_transpose(img).convert("RGB")
w, h = img.size
scale = edge / max(w, h)
nw, nh = max(1, round(w * scale)), max(1, round(h * scale))
img = img.resize((nw, nh), Image.BILINEAR)
canvas = Image.new("RGB", (edge, edge), (pad, pad, pad))
canvas.paste(img, ((edge - nw) // 2, (edge - nh) // 2))
x = np.asarray(canvas, dtype=np.float32) # HWC, 0..255
x = norm(x)
return np.ascontiguousarray(x.transpose(2, 0, 1))[None] # NCHW
# ---- the app's samplers (dr-face Letterbox, dr-segment Letterbox::sample, align.rs) ----
def load(p):
return np.asarray(Image.open(p).convert("RGB"), np.float32) / 255.0
def preprocessing(name, shape):
"""Which normalisation this model is fed in the app.
SCRFD (`dr-face::detect`): `(v - 127.5) / 128`, padded with 114.
ArcFace (`dr-face::embed`): the same, on an aligned 112 crop — a
letterboxed photograph is the wrong distribution, but the embedder is
never quantised (§7), so this is only ever a fallback.
YOLO (`dr-segment`): `v / 255`, padded with 0.5.
"""
edge = shape[-1]
if name.startswith("scrfd") or name.startswith("arcface"):
return edge, 114, lambda x: (x - 127.5) / 128.0
return edge, 128, lambda x: x / 255.0
def bilinear(img, sx, sy):
h, w = img.shape[:2]
x0, y0 = np.floor(sx).astype(np.int64), np.floor(sy).astype(np.int64)
fx, fy = (sx - x0)[..., None], (sy - y0)[..., None]
c = lambda a, n: np.clip(a, 0, n - 1) # noqa: E731 - neighbours clamp at the edge
top = img[c(y0, h), c(x0, w)] * (1 - fx) + img[c(y0, h), c(x0 + 1, w)] * fx
bot = img[c(y0 + 1, h), c(x0, w)] * (1 - fx) + img[c(y0 + 1, h), c(x0 + 1, w)] * fx
return top * (1 - fy) + bot * fy
class Photos(CalibrationDataReader):
"""One chunk of photographs, fed as the app would feed them."""
def chw(x):
return np.ascontiguousarray(x.transpose(2, 0, 1))[None].astype(np.float32)
def __init__(self, input_name, paths, edge, pad, norm):
self.name = input_name
self.items = iter(letterbox(Image.open(p), edge, pad, norm) for p in paths)
def letterbox(img, edge, yolo):
h, w = img.shape[:2]
s = min(edge / w, edge / h)
px, py = (edge - w * s) / 2, (edge - h * s) / 2
ix, iy = np.meshgrid(np.arange(edge) + 0.5, np.arange(edge) + 0.5)
if yolo: # semantic.rs: no -0.5, pad 0.5, 0..1
sx, sy = (ix - px) / s, (iy - py) / s
out = bilinear(img, sx, sy)
out[(sx < 0) | (sx >= w) | (sy < 0) | (sy >= h)] = 0.5
else: # detect.rs: -0.5, pad 114, (v·255 − 127.5)/128
sx, sy = (ix - px) / s - 0.5, (iy - py) / s - 0.5
out = (bilinear(img, sx, sy) * 255 - 127.5) / 128
out[(sx < -0.5) | (sx > w - 0.5) | (sy < -0.5) | (sy > h - 0.5)] = (114 - 127.5) / 128
return chw(out), (s, px, py)
def crop_box(img, x0, y0, bw, bh, ow, oh):
"""align.rs crop_box: output (u+.5) → source, −0.5, bilinear, outside black."""
u, v = np.meshgrid(np.arange(ow) + 0.5, np.arange(oh) + 0.5)
sx, sy = x0 + u * bw / ow - 0.5, y0 + v * bh / oh - 0.5
out = bilinear(img, sx, sy)
h, w = img.shape[:2]
out[(sx < -1) | (sx > w) | (sy < -1) | (sy > h)] = 0
return out
def scrfd_boxes(outs, s, px, py):
"""detect.rs decode: score ≥ 0.5, greedy NMS at 0.4, min side 24 px."""
fmc = len(outs) // 3
boxes, scores = [], []
for i, st in enumerate([8, 16, 32, 64][:fmc]):
sc, bx = outs[i].reshape(-1), outs[fmc + i].reshape(-1, 4)
idx = np.nonzero(sc >= 0.5)[0]
cell = idx // 2
cx, cy = (cell % (640 // st)) * st, (cell // (640 // st)) * st
boxes.append(np.stack([cx - bx[idx, 0] * st, cy - bx[idx, 1] * st, cx + bx[idx, 2] * st, cy + bx[idx, 3] * st], 1))
scores.append(sc[idx])
b, sc = np.concatenate(boxes), np.concatenate(scores)
keep = []
for i in np.argsort(-sc):
x0 = np.maximum(b[i, 0], b[keep, 0]); y0 = np.maximum(b[i, 1], b[keep, 1])
x1 = np.minimum(b[i, 2], b[keep, 2]); y1 = np.minimum(b[i, 3], b[keep, 3])
inter = np.clip(x1 - x0, 0, None) * np.clip(y1 - y0, 0, None)
area = lambda r: (r[..., 2] - r[..., 0]) * (r[..., 3] - r[..., 1]) # noqa: E731
if not keep or (inter / (area(b[i]) + area(b[keep]) - inter)).max() <= 0.4:
keep.append(i)
b = (b[keep] - [px, py, px, py]) / s
return b[np.minimum(b[:, 2] - b[:, 0], b[:, 3] - b[:, 1]) >= 32]
# ---- one calibration input per photograph (or per face), as the app makes it ----
def feeds(kind, photos, model):
name = model.get_inputs()[0].name
if kind == "scrfd":
for p in photos:
yield {name: letterbox(load(p), 640, False)[0]}
elif kind == "landmarks": # landmarks.rs: 1.5× the box, square, 0..255
det = ort.InferenceSession(str(MODELS / "face/scrfd_10g_640.onnx"), providers=["CPUExecutionProvider"])
for p in photos:
img = load(p)
x, ctx = letterbox(img, 640, False)
for b in scrfd_boxes(det.run(None, {"input.1": x}), *ctx):
cx, cy, side = (b[0] + b[2]) / 2, (b[1] + b[3]) / 2, 1.5 * max(b[2] - b[0], b[3] - b[1])
yield {name: chw(crop_box(img, cx - side / 2, cy - side / 2, side, side, 192, 192) * 255)}
elif kind == "yolo":
for p in photos:
yield {name: letterbox(load(p), 640, True)[0]}
elif kind == "migan": # migan.rs: ch0 = known − 0.5, ch1–3 = (rgb·2 − 1)·known
rng = np.random.default_rng(7)
for p in photos:
img = load(p)
h, w = img.shape[:2]
e = min(h, w)
sq = Image.fromarray((img[(h - e) // 2:(h + e) // 2, (w - e) // 2:(w + e) // 2] * 255).astype(np.uint8))
img = np.asarray(sq.resize((512, 512), Image.BILINEAR), np.float32) / 255
known = np.ones((512, 512), np.float32)
for _ in range(rng.integers(1, 3)): # a panorama's unknown border: a wedge along one edge
side = rng.integers(4)
depth = np.linspace(rng.integers(20, 110), rng.integers(20, 110), 512).astype(int)
edge = np.arange(512)[:, None] < depth[None, :] # [depth, along]: inside the wedge
wedge = edge if side % 2 == 0 else edge[::-1] # top / bottom of a column
known[wedge if side < 2 else wedge.T] = 0 # or left / right of a row
x = np.concatenate([(known - 0.5)[None], ((img * 2 - 1) * known[..., None]).transpose(2, 0, 1)])[None]
yield {name: x.astype(np.float32)}
elif kind == "xfeat": # xfeat.rs: grey 0..1, shrink to fit, top-left, zero pad
_, _, H, W = [d if isinstance(d, int) else 1 for d in model.get_inputs()[0].shape]
for p in photos:
g = load(p).mean(2) # a display-rendered photograph is already the app's (R+G+B)/3 ^ 1/2.2
h, w = g.shape
s = min(W / w, H / h, 1.0)
if s < 1:
g = np.asarray(Image.fromarray(g).resize((round(w * s), round(h * s)), Image.BOX))
pad = np.zeros((H, W), np.float32)
pad[:g.shape[0], :g.shape[1]] = g
yield {name: pad[None, None]}
class Items(CalibrationDataReader):
def __init__(self, items):
self.it = iter(items)
def get_next(self):
x = next(self.items, None)
return None if x is None else {self.name: x}
return next(self.it, None)
# Photographs whose activations are held in memory at once. Every ONNX
# Runtime calibrator keeps each image's whole set of activations until it
# folds them into a range — a gigabyte an image on the 10g detector at 640²,
# and folded once at the end, an OOM kill with no message. Folding every
# `CHUNK` images gives ranges identical to folding once (checked on
# scrfd_500m, 129 tensors, no difference) at a bounded cost.
CHUNK = 4
def calibrate(path, items, cache):
"""Min/max ranges in chunks — every ORT calibrator holds all activations
until it folds them, and the others measurably degrade the result."""
cal = create_calibrator(Path(path), None, augmented_model_path=f"{path}.aug.onnx",
calibrate_method=CalibrationMethod.MinMax)
batch, n = [], 0
for item in items:
batch.append(item)
n += 1
if len(batch) == CHUNK:
cal.collect_data(Items(batch))
batch = []
if batch:
cal.collect_data(Items(batch))
save_tensors_data(cal.compute_data(), cache)
os.remove(f"{path}.aug.onnx")
return n
def calibrate(pre, name, photos, cache):
"""Min/max ranges over `photos`, written to `cache` for quantize_static.
Plain min/max: the moving average and the strided option of
`quantize_static` both measured worse than this on held-out proxies, and
the percentile method has no memory bound at all.
"""
import onnxruntime as ort
s = ort.InferenceSession(pre, providers=["CPUExecutionProvider"])
i = s.get_inputs()[0]
shape = [d if isinstance(d, int) else 1 for d in i.shape]
edge, pad, norm = preprocessing(name, shape)
calibrator = create_calibrator(
Path(pre),
None,
augmented_model_path=f"{pre}.augmented.onnx",
calibrate_method=CalibrationMethod.MinMax,
)
for start in range(0, len(photos), CHUNK):
calibrator.collect_data(Photos(i.name, photos[start : start + CHUNK], edge, pad, norm))
ranges = calibrator.compute_data()
save_tensors_data(ranges, cache)
os.remove(f"{pre}.augmented.onnx")
def downstream(m, start):
names, live = set(), set()
for n in m.graph.node:
if n.name == start or any(i in live for i in n.input):
names.add(n.name)
live.update(n.output)
return sorted(names)
def main():
photo_dir, models = sys.argv[1], sys.argv[2:]
photos = sorted(
p
for ext in ("jpg", "jpeg", "JPG", "JPEG", "png")
for p in glob.glob(os.path.join(photo_dir, "**", f"*.{ext}"), recursive=True)
)[:PHOTOS]
if len(photos) < 20:
sys.exit(f"only {len(photos)} photographs under {photo_dir}; calibration wants dozens")
print(f"==> calibrating on {len(photos)} photographs")
args = sys.argv[1:]
ranges = None
if args[:1] == ["--ranges"]:
ranges, args = args[1], args[2:]
photo_dir = None
else:
photo_dir, args = args[0], args[1:]
names = args or [n for n in TABLE if TABLE[n]["feed"]]
photos = []
if photo_dir:
photos = sorted(p for e in ("jpg", "jpeg", "JPG", "JPEG", "png")
for p in glob.glob(os.path.join(photo_dir, "**", f"*.{e}"), recursive=True))[:PHOTOS]
if len(photos) < 50:
sys.exit(f"only {len(photos)} photographs under {photo_dir}; calibration wants hundreds")
print(f"==> calibrating on {len(photos)} photographs")
for src in models:
stem, _ = os.path.splitext(src)
name = os.path.basename(src)
out = f"{stem}.int8.onnx"
for stem in names:
spec = TABLE[stem]
src = MODELS / spec["dir"] / f"{stem}.onnx"
out = src.with_name(f"{stem}.{spec['form']}.onnx")
work = src.with_name(f"{stem}.quant-work.onnx")
print(f" {stem} -> {out.name}")
m = onnx.load(src)
opset = next((o.version for o in m.opset_import if o.domain in ("", "ai.onnx")), 0)
work = f"{stem}.quant-work.onnx"
if opset < 13:
print(f" {name}: opset {opset} -> 17")
m = version_converter.convert_version(m, 17)
if next(o.version for o in m.opset_import if o.domain in ("", "ai.onnx")) < 13:
m = version_converter.convert_version(m, 17) # per-channel QDQ needs 13
m.ir_version = 8
onnx.save(m, work)
pre = f"{stem}.quant-pre.onnx"
quant_pre_process(work, pre)
cache = f"{stem}.quant-ranges.json"
calibrate(pre, name, photos, cache)
quantize_static(
pre,
out,
None,
quant_format=QuantFormat.QDQ,
per_channel=True,
activation_type=QuantType.QUInt8,
weight_type=QuantType.QInt8,
calibrate_method=CalibrationMethod.MinMax,
calibration_cache_path=cache,
)
for f in (work, pre, cache):
os.remove(f)
print(f" {out}: {os.path.getsize(out) // 1024} KB")
if spec.get("rewrites"):
onnx.save(htp_graph.rewrite(str(work), spec["rewrites"]), work)
model = ort.InferenceSession(str(work), providers=["CPUExecutionProvider"])
cache = str(work) + ".ranges"
if spec["feed"] is None:
if not ranges:
sys.exit(f"{stem} is calibrated on mosaics, not photographs: pass --ranges")
cache = ranges
else:
n = calibrate(str(work), feeds(spec["feed"], photos, model), cache)
print(f" {n} calibration inputs")
act, wt = FORMS[spec["form"]]
# The config needs a reader only to exist; the ranges come from `cache`.
zeros = {i.name: np.zeros([d if isinstance(d, int) else 1 for d in i.shape], np.float32)
for i in model.get_inputs()}
cfg = get_qnn_qdq_config(str(work), Items([zeros]), activation_type=act, weight_type=wt, per_channel=True)
exclude = list(cfg.nodes_to_exclude or []) + spec.get("float_nodes", [])
if spec.get("float_from"):
exclude += downstream(onnx.load(work), spec["float_from"])
quantize_static(str(work), str(out), None, quant_format=cfg.quant_format,
op_types_to_quantize=cfg.op_types_to_quantize, per_channel=True,
activation_type=act, weight_type=wt, nodes_to_exclude=exclude,
calibrate_method=CalibrationMethod.MinMax, extra_options=cfg.extra_options,
calibration_cache_path=cache)
os.remove(work)
if cache != ranges:
os.remove(cache)
print(f" {out.name}: {out.stat().st_size // 1024} KB")
if __name__ == "__main__":
+22 -18
View File
@@ -1,30 +1,34 @@
#!/usr/bin/env bash
# Produce the int8 form of a model for the Hexagon (docs/dev/inference.md §5).
# Produce the Hexagon's form of each model (docs/dev/inference.md §1.5, §5).
#
# ./tools/quantise-models.sh PHOTO_DIR MODEL.onnx [MODEL.onnx ...]
# ./tools/quantise-models.sh PHOTO_DIR [MODEL ...]
# ./tools/quantise-models.sh --ranges RANGES.json mosaic-medium-1408
#
# Writes `MODEL.int8.onnx` beside each input: a QDQ graph, per-channel int8
# weights, uint8 activations — the form QNN's HTP backend takes whole. The
# activations' ranges come from running the f32 model over the photographs in
# PHOTO_DIR, fed exactly as the app feeds them (letterboxed to the model's
# input, the detector's `(x - 127.5) / 128` normalisation), which is why
# this is a release-time step and not something the device does: it needs
# real photographs and, after it, a person reading §10 M2's numbers.
# Writes `<stem>.<form>.onnx` beside each canonical file under models/: a QDQ
# graph from QNN's own quantisation config, per-channel weights, in the form
# the engine's `Rung::form` names for that role — A16W8, A16W16 or int8, each
# the narrowest that held the model's accuracy on the tablet. The activation
# ranges come from running the f32 model over the photographs in PHOTO_DIR,
# fed exactly as the app feeds them (letterbox maths, pads, normalisation,
# face crops through the app's own similarity), which is why this is a
# release-time step and not something the device does. With no MODEL, every
# model in the table.
#
# The SCRFD and ArcFace exports are opset 11; per-channel QDQ needs 13, so a
# model below 13 is first upgraded to 17. That changes only the graph's
# spelling, not a weight — and it is what `tools/fix-face-model-shapes.sh`
# will do to the canonical files in the same model release.
# The denoiser is calibrated on noisy mosaics, not photographs: its ranges
# come from darkroom-denoise's precision gate (`--ranges`), computed on a
# smaller tile of the same network — activation ranges do not depend on the
# tile's size, and the tensor names match.
#
# Then measure before shipping: a quantised form is a different network, and
# the numbers in inference.md §1.5 are what each one had to hold.
#
# A venv per run, like fix-face-model-shapes.sh: the tools are not a build
# input and nothing in the tree should have them on its path.
set -euo pipefail
if [ "$#" -lt 2 ]; then
sed -n '2,20p' "$0" >&2
if [ "$#" -lt 1 ]; then
sed -n '2,27p' "$0" >&2
exit 2
fi
PHOTOS="$1"; shift
[ -d "${PHOTOS}" ] || { echo "no such directory: ${PHOTOS}" >&2; exit 1; }
WORK="$(mktemp -d -p /var/tmp quantise-models.XXXXXX)"
trap 'rm -rf "${WORK}"' EXIT
@@ -32,4 +36,4 @@ echo "==> venv in ${WORK}"
uv venv --python 3.12 "${WORK}/venv" >/dev/null
VIRTUAL_ENV="${WORK}/venv" uv pip install --quiet onnx onnxruntime pillow numpy sympy
exec "${WORK}/venv/bin/python" "$(dirname "$0")/quantise-models.py" "${PHOTOS}" "$@"
exec "${WORK}/venv/bin/python" "$(dirname "$0")/quantise-models.py" "$@"
+1
View File
@@ -42,6 +42,7 @@ dr-ingest.workspace = true
# The sameness probe of a catalog duplicate (FR-CAT-11a): SHA-256 over the
# ends of each copy, the digest the import already uses for whole files.
sha2 = "0.10"
half = "2.7"
dr-film.workspace = true
# The lens profile database, here for the same reason dr-film is: dr-pipeline
# knows the maths of lens correction and deliberately has no dependency with
+261 -1
View File
@@ -96,6 +96,11 @@ pub struct SyncReport {
/// (camera-profiles.md §13).
pub profiles_uploaded: usize,
pub profiles_downloaded: usize,
/// TRACES: FR-DEV-6
/// The develop presets: whether another device's changes reached this
/// one's library, and whether this one's reached the server.
pub presets_adopted: bool,
pub presets_uploaded: bool,
}
impl SyncReport {
@@ -109,6 +114,7 @@ impl SyncReport {
|| self.face_shards_downloaded > 0
|| self.place_adopted
|| self.profiles_downloaded > 0
|| self.presets_adopted
}
}
@@ -124,12 +130,15 @@ pub enum SyncMessage {
///
/// Runs on its own thread with its own runtime, like every other network path
/// here — the Slint loop must never block (NFR-P9).
// One argument per thing the pass touches; see `run` below.
#[allow(clippy::too_many_arguments)]
pub fn spawn_sync(
conn: Connection,
root: String,
thumbs_dir: PathBuf,
catalog_path: PathBuf,
place_path: PathBuf,
presets: PresetFiles,
scratch: PathBuf,
// TRACES: FR-CULL-8
// Which face pipeline's shards to export and adopt. From the settings
@@ -165,6 +174,7 @@ pub fn spawn_sync(
&thumbs_dir,
&catalog_path,
&place_path,
&presets,
&scratch,
&face_model_id,
&tx,
@@ -184,7 +194,7 @@ pub fn spawn_sync(
rx
}
// Eight, because a sync touches eight distinct things — the same reason
// Nine, because a sync touches nine distinct things — the same reason
// `repairs::spawn` carries the allow: bundling them into a struct would name
// nothing that exists.
#[allow(clippy::too_many_arguments)]
@@ -194,6 +204,7 @@ async fn run(
thumbs_dir: &Path,
catalog_path: &Path,
place_path: &Path,
presets: &PresetFiles,
scratch: &Path,
face_model_id: &str,
tx: &std::sync::mpsc::Sender<SyncMessage>,
@@ -240,6 +251,12 @@ async fn run(
sync_profiles(backend, &base, &dir, &mut report).await;
}
// TRACES: FR-DEV-6
// The develop presets, a few kilobytes. Like the profiles, never fails
// the pass.
let _ = tx.send(SyncMessage::Status("checking presets…".into()));
sync_presets(backend, &base, presets, &mut report).await;
// TRACES: FR-UI-8
// Last, and it costs one small GET plus at most one small PUT. Last because
// it is the only thing here that is not derived state and so the only thing
@@ -1122,6 +1139,156 @@ async fn sync_profiles(
}
}
/// TRACES: FR-DEV-6
/// Where this device keeps the preset library, and what the last exchange
/// of it with this library left both sides holding.
///
/// The library is the device's, shared by every library it opens; the base
/// is per library, because each library's server holds its own copy and has
/// its own history of exchanges with this device.
#[derive(Debug, Clone)]
pub struct PresetFiles {
pub library: PathBuf,
pub base: PathBuf,
}
/// The preset library on the server, in its own folder so finding it is a
/// listing of one file rather than of every shard beside it.
const PRESETS_DIR: &str = "presets";
const PRESETS_NAME: &str = "library.drpl";
/// TRACES: FR-DEV-6
/// Exchange the develop presets with `<derived>/presets/library.drpl`.
///
/// Unlike the place, this is merged rather than replaced: a preset saved on
/// the tablet and another saved on the desktop between two passes must both
/// survive, and a preset deleted on one must not come back from the other.
/// [`PresetLibrary::merge`](dr_pipeline::PresetLibrary::merge) decides each
/// name against the base the last exchange left, which is what tells a
/// deletion here from an addition there.
///
/// The write is conditional on the server still holding what was read, so
/// two devices exchanging at once cannot each save over the other's
/// additions; the one that loses the race reads again and merges again.
/// Never fails the pass.
async fn sync_presets(
backend: &dyn RemoteBackend,
base: &RemotePath,
files: &PresetFiles,
report: &mut SyncReport,
) {
for _ in 0..3 {
match exchange_presets(backend, base, files, report).await {
Err(RemoteError::PreconditionFailed) => {
log::debug!("presets: another device wrote first; reading again");
}
Err(e) => {
log::debug!("not exchanging presets: {e}");
return;
}
Ok(()) => return,
}
}
}
async fn exchange_presets(
backend: &dyn RemoteBackend,
base: &RemotePath,
files: &PresetFiles,
report: &mut SyncReport,
) -> Result<(), RemoteError> {
use crate::preset_store::PresetStore;
use dr_pipeline::PresetLibrary;
let dir = RemotePath::new(format!("{}/{PRESETS_DIR}", base.as_str()));
let target = RemotePath::new(format!("{}/{PRESETS_NAME}", dir.as_str()));
// A listing that fails reads as "not there". That is safe because the
// write below is then `IfAbsent`, which a server holding one refuses.
let listed = backend.list(&dir, None).await.ok().and_then(|entries| {
entries
.into_iter()
.find(|e| e.kind == dr_sync::EntryKind::File && e.path.name() == PRESETS_NAME)
});
let mut theirs = match &listed {
None => PresetLibrary::default(),
Some(_) => {
let bytes = read_derived(backend, &target).await?;
match std::str::from_utf8(&bytes)
.map_err(|e| e.to_string())
.and_then(|t| PresetLibrary::parse(t).map_err(|e| e.to_string()))
{
Ok(library) => library,
Err(e) => {
// A newer build's format, or damage. Either way not ours
// to write over.
log::warn!("the preset library on the server will not read ({e}); left alone");
return Ok(());
}
}
}
};
let store = PresetStore::open_at(files.library.clone());
let read = match store.try_load() {
Ok(library) => library.unwrap_or_default(),
Err(e) => {
// An empty library here would read as every preset deleted.
log::warn!("not exchanging presets: {}: {e}", store.path().display());
return Ok(());
}
};
let base_store = PresetStore::open_at(files.base.clone());
let last = base_store.try_load().ok().flatten().unwrap_or_default();
// Copies an older build seeded of the shipped presets are not the
// photographer's, and are dropped on both sides before anything travels.
let mut ours = read.clone();
dr_pipeline::bundled::forget_unchanged_copies(&mut ours);
dr_pipeline::bundled::forget_unchanged_copies(&mut theirs);
let merged = PresetLibrary::merge(&last, &ours, &theirs);
if listed.is_none() && merged.is_empty() {
return Ok(());
}
if listed.is_none() || merged != theirs {
let _ = backend.create_dir(&dir).await;
let precondition = match &listed {
Some(entry) => dr_sync::Precondition::IfMatch(entry.validator.clone()),
None => dr_sync::Precondition::IfAbsent,
};
backend
.put(&target, merged.to_text().into_bytes(), Some(precondition))
.await?;
report.presets_uploaded = true;
}
if merged != ours {
// The develop view saves this file too. A preset it saved since the
// read above is merged in rather than written over; it reaches the
// server on the next pass, as an addition against the base below.
let local = match store.try_load() {
Ok(Some(now)) if now != read => PresetLibrary::merge(&read, &merged, &now),
_ => merged.clone(),
};
match store.save(&local) {
Ok(()) => report.presets_adopted = true,
Err(e) => {
log::warn!("saving presets from the library: {e}");
return Ok(());
}
}
}
if merged != last {
if let Err(e) = base_store.save(&merged) {
log::debug!("recording the preset exchange: {e}");
}
}
Ok(())
}
/// TRACES: FR-UI-8
/// Fetch just the place, for the handover at launch.
///
@@ -1934,6 +2101,99 @@ mod derived_guard_tests {
let _ = std::fs::remove_dir_all(&root);
}
/// One device's preset files, under `root`.
fn preset_device(root: &Path, name: &str) -> PresetFiles {
PresetFiles {
library: root.join(name).join("presets.drpl"),
base: root.join(name).join("presets.base.drpl"),
}
}
fn preset_names(files: &PresetFiles) -> Vec<String> {
crate::preset_store::PresetStore::open_at(files.library.clone())
.load()
.names()
.map(str::to_string)
.collect()
}
fn save_presets(files: &PresetFiles, names: &[&str]) {
let mut library = dr_pipeline::PresetLibrary::default();
for name in names {
library
.insert(name, dr_pipeline::Preset::default())
.unwrap();
}
crate::preset_store::PresetStore::open_at(files.library.clone())
.save(&library)
.unwrap();
}
#[tokio::test]
async fn presets_reach_every_device_and_so_do_their_deletions() {
// TRACES: FR-DEV-6
let root = std::env::temp_dir().join(format!("dr-preset-sync-{}", std::process::id()));
let _ = std::fs::remove_dir_all(&root);
let server = root.join("server");
std::fs::create_dir_all(server.join(".darkroom-derived")).unwrap();
let backend = dr_sync_folder::FolderBackend::new(&server).unwrap();
let base = RemotePath::new(".darkroom-derived");
let (desk, tablet) = (preset_device(&root, "desk"), preset_device(&root, "tablet"));
save_presets(&desk, &["Warm"]);
save_presets(&tablet, &["Mono"]);
for files in [&desk, &tablet, &desk] {
sync_presets(&backend, &base, files, &mut SyncReport::default()).await;
}
assert_eq!(preset_names(&desk), vec!["Mono", "Warm"]);
assert_eq!(preset_names(&tablet), vec!["Mono", "Warm"]);
// Deleted on the desk: gone from the tablet, not back on the desk.
save_presets(&desk, &["Mono"]);
for files in [&desk, &tablet, &desk] {
sync_presets(&backend, &base, files, &mut SyncReport::default()).await;
}
assert_eq!(preset_names(&desk), vec!["Mono"]);
assert_eq!(preset_names(&tablet), vec!["Mono"]);
// Settled: a pass with nothing new writes nothing.
let mut quiet = SyncReport::default();
sync_presets(&backend, &base, &tablet, &mut quiet).await;
assert!(!quiet.presets_uploaded && !quiet.presets_adopted);
let _ = std::fs::remove_dir_all(&root);
}
#[tokio::test]
async fn a_preset_library_the_server_cannot_read_is_left_alone() {
// TRACES: FR-DEV-6
// A newer build's format reads as unreadable here, and is that build's
// presets: neither written over nor taken as an empty library.
let root = std::env::temp_dir().join(format!("dr-preset-unread-{}", std::process::id()));
let _ = std::fs::remove_dir_all(&root);
let server = root.join("server");
let held = server.join(".darkroom-derived/presets/library.drpl");
std::fs::create_dir_all(held.parent().unwrap()).unwrap();
std::fs::write(&held, "drpl 9999\n").unwrap();
let backend = dr_sync_folder::FolderBackend::new(&server).unwrap();
let desk = preset_device(&root, "desk");
save_presets(&desk, &["Warm"]);
let mut report = SyncReport::default();
sync_presets(
&backend,
&RemotePath::new(".darkroom-derived"),
&desk,
&mut report,
)
.await;
assert!(!report.presets_uploaded);
assert_eq!(std::fs::read_to_string(&held).unwrap(), "drpl 9999\n");
assert_eq!(preset_names(&desk), vec!["Warm"]);
let _ = std::fs::remove_dir_all(&root);
}
#[tokio::test]
async fn a_place_that_could_not_be_read_is_never_written_over() {
// A dehydrated placeholder, and the record on the server may well be
+145 -39
View File
@@ -18,6 +18,7 @@ use std::sync::Arc;
use dr_decode::RawImage;
use dr_gpu::{DemosaicedImage, GrainBlend};
use dr_pipeline::learned_denoise::Method;
use super::session::DevelopSession;
@@ -41,6 +42,12 @@ pub(crate) struct DenoiseState {
failed: Option<String>,
/// Where the noise figures came from, for the panel.
source: Option<dr_denoise::Source>,
/// The file's bytes, hashed for the on-disk cache, which keys each
/// network's result on them and the model (`denoise_cache::FileHash`).
file_hash: Option<super::denoise_cache::FileHash>,
/// The method the result, the job and the failure above are for. A
/// different one asked for discards them.
method: Option<Method>,
}
struct Job {
@@ -89,6 +96,11 @@ impl DevelopSession {
if self.denoise.mosaic.is_some() {
self.denoise.profile = dr_decode::noise_profile(bytes);
self.denoise.iso = meta.iso;
// TRACES: FR-DEV-3g
// The bytes are only here now, so they are hashed now: a
// reopened or exported photograph finds its result on disk,
// under whichever method it asks for.
self.denoise.file_hash = Some(super::denoise_cache::FileHash::of(bytes));
}
}
@@ -97,6 +109,7 @@ impl DevelopSession {
/// Cheap when nothing changed; the develop view calls it after every
/// change to the edit, whatever made it — a slider, undo, a version.
pub fn reconcile_denoise(&mut self) {
self.forget_other_method();
let wanted = self.graph.denoise_applied() && self.denoise.mosaic.is_some();
if !wanted {
if let Some(job) = self.denoise.job.take() {
@@ -112,20 +125,11 @@ impl DevelopSession {
{
return;
}
let Some(model) = crate::library::denoise_model() else {
self.denoise.failed = Some("the denoise model is not installed".into());
let Some(work) = self.work(Arc::new(AtomicBool::new(false))) else {
return;
};
let cancel = work.cancel.clone();
let (tx, rx) = mpsc::channel();
let cancel = Arc::new(AtomicBool::new(false));
let work = Work {
ctx: self.ctx.clone(),
raw: self.denoise.mosaic.clone().expect("checked above"),
profile: self.denoise.profile.clone(),
iso: self.denoise.iso,
model,
cancel: cancel.clone(),
};
crate::executors::spawn(crate::executors::Executor::Decode, "denoise", move || {
let result = work.run(&mut |done, total| {
let _ = tx.send(Msg::Progress(done, total));
@@ -135,6 +139,51 @@ impl DevelopSession {
self.denoise.job = Some(Job { rx, cancel });
}
/// Drop what was computed, or is being computed, for a network the edit
/// no longer asks for — a choice in the panel, an undo, a version.
/// Each network's result stays in the on-disk cache, so going back to
/// one is a read, not a run. The classical demosaic asks for no network
/// and drops nothing: the result is kept for the way back.
fn forget_other_method(&mut self) {
let asked = self.graph.denoise_method();
if !asked.learned() || self.denoise.method == Some(asked) {
return;
}
if self.denoise.method.is_none() {
// The first network asked for: nothing computed is another's.
self.denoise.method = Some(asked);
return;
}
if let Some(job) = self.denoise.job.take() {
job.cancel.store(true, Ordering::Relaxed);
}
self.denoise.result = None;
self.denoise.blended = None;
self.denoise.failed = None;
self.denoise.source = None;
self.denoise.method = Some(asked);
}
/// The job for the method the edit asks for, or `None` — with the reason
/// kept as the failure — where its network is not installed.
fn work(&mut self, cancel: Arc<AtomicBool>) -> Option<Work> {
let raw = self.denoise.mosaic.clone()?;
let Some((model, net)) = crate::library::denoise_model(self.graph.denoise_method()) else {
self.denoise.failed = Some("the denoise model is not installed".into());
return None;
};
Some(Work {
ctx: self.ctx.clone(),
raw,
profile: self.denoise.profile.clone(),
iso: self.denoise.iso,
cache_key: self.denoise.file_hash.as_ref().map(|h| h.key(&model)),
model,
halo: net.halo,
cancel,
})
}
/// Collect what the job sent since the last poll.
pub fn poll_denoise(&mut self) -> DenoiseStatus {
let Some(job) = &self.denoise.job else {
@@ -173,9 +222,10 @@ impl DevelopSession {
if !self.graph.denoise_applied() || self.denoise.result.is_some() {
return Ok(());
}
let Some(raw) = self.denoise.mosaic.clone() else {
self.forget_other_method();
if self.denoise.mosaic.is_none() {
return Ok(());
};
}
// Already under way: wait for it rather than start again.
if let Some(job) = self.denoise.job.take() {
for msg in job.rx.iter() {
@@ -184,14 +234,8 @@ impl DevelopSession {
}
}
}
let model = crate::library::denoise_model().ok_or("the denoise model is not installed")?;
let work = Work {
ctx: self.ctx.clone(),
raw,
profile: self.denoise.profile.clone(),
iso: self.denoise.iso,
model,
cancel: Arc::new(AtomicBool::new(false)),
let Some(work) = self.work(Arc::new(AtomicBool::new(false))) else {
return Err(self.denoise.failed.clone().unwrap_or_default());
};
let finished = work.run(&mut |_, _| {})?;
self.land(finished)
@@ -272,12 +316,33 @@ struct Work {
profile: Option<Vec<(f32, f32)>>,
iso: Option<u32>,
model: std::path::PathBuf,
/// The context `model` needs past a tile's kept centre.
halo: usize,
cancel: Arc<AtomicBool>,
cache_key: Option<String>,
}
impl Work {
fn run(self, progress: &mut dyn FnMut(usize, usize)) -> Result<Finished, String> {
let started = std::time::Instant::now();
// TRACES: FR-DEV-3g
// A result computed before — this photograph opened earlier, or
// developed and now exported — is read back rather than recomputed.
let cache_dir = super::denoise_cache::dir();
if let Some(hit) = self
.cache_key
.as_deref()
.and_then(|key| super::denoise_cache::load(&cache_dir, key))
{
return Ok(Finished {
rgb: hit.rgb,
width: hit.width,
height: hit.height,
source: hit.source,
rung: "the cache".into(),
seconds: started.elapsed().as_secs_f64(),
});
}
// The app's own hot-pixel pass, on a copy: the classical source was
// repaired by the same pass inside `Demosaicer::run`.
let mut raw = (*self.raw).clone();
@@ -286,8 +351,8 @@ impl Work {
.map_err(|e| e.to_string())?;
let noise = dr_denoise::noise::for_frame_with(&raw, self.profile.as_deref(), self.iso)
.ok_or("this photograph gives no way to measure its noise")?;
let mut net =
dr_denoise::onnx::OnnxNet::from_path(&self.model).map_err(|e| e.to_string())?;
let mut net = dr_denoise::onnx::OnnxNet::from_path(&self.model, self.halo)
.map_err(|e| e.to_string())?;
let rung = net
.rung()
.map(|r| r.label().to_string())
@@ -299,14 +364,24 @@ impl Work {
})
.map_err(|e| e.to_string())?
.ok_or("stopped")?;
Ok(Finished {
let finished = Finished {
rgb,
width: raw.crop.width,
height: raw.crop.height,
source: noise.source,
rung,
seconds: started.elapsed().as_secs_f64(),
})
};
if let Some(key) = self.cache_key.as_deref() {
let entry = super::denoise_cache::Entry {
rgb: finished.rgb.clone(),
width: finished.width,
height: finished.height,
source: finished.source,
};
super::denoise_cache::store(&cache_dir, key, &entry, super::denoise_cache::BUDGET);
}
Ok(finished)
}
}
@@ -385,15 +460,24 @@ mod tests {
let result = s.denoise.result.clone().unwrap();
let same = |a: &Arc<DemosaicedImage>, b: &Arc<DemosaicedImage>| Arc::ptr_eq(a, b);
// Off: the classical demosaic, result or no result.
assert!(same(&s.developed_source(), &classical));
// On, no grain: the network's result as it is.
s.graph
.set_param(learned_denoise::ID, learned_denoise::APPLY, 1.0);
// On by default, at full strength: the network's result as it is.
assert!(same(&s.developed_source(), &result));
// Grain: a blend, made once per value and reused until it moves.
// Bilinear: the classical demosaic, result or no result.
s.graph.set_param(
learned_denoise::ID,
learned_denoise::METHOD,
Method::Bilinear.index(),
);
assert!(same(&s.developed_source(), &classical));
s.graph.set_param(
learned_denoise::ID,
learned_denoise::METHOD,
Method::DEFAULT.index(),
);
// Less strength: a blend, made once per value and reused until it
// moves.
s.graph
.set_param(learned_denoise::ID, learned_denoise::GRAIN, 40.0);
.set_param(learned_denoise::ID, learned_denoise::STRENGTH, 60.0);
let blended = s.developed_source();
assert!(!same(&blended, &result) && !same(&blended, &classical));
assert!(
@@ -401,7 +485,7 @@ mod tests {
"the same grain must not blend again"
);
s.graph
.set_param(learned_denoise::ID, learned_denoise::GRAIN, 60.0);
.set_param(learned_denoise::ID, learned_denoise::STRENGTH, 40.0);
assert!(!same(&s.developed_source(), &blended));
// The sensor's own reading stays the classical one throughout.
assert!(same(&s.demosaiced, &classical));
@@ -413,19 +497,41 @@ mod tests {
let mut s =
DevelopSession::open_owned(&ctx, bayer(64, 64), dr_types::Orientation::NORMAL).unwrap();
s.denoise.failed = Some("no model".into());
s.graph
.set_param(learned_denoise::ID, learned_denoise::APPLY, 1.0);
s.reconcile_denoise();
assert!(
s.denoise.job.is_none(),
"a failure is not retried while the switch stays on"
"a failure is not retried while the method stays"
);
s.graph.set_param(
learned_denoise::ID,
learned_denoise::METHOD,
Method::Bilinear.index(),
);
s.graph
.set_param(learned_denoise::ID, learned_denoise::APPLY, 0.0);
s.reconcile_denoise();
assert!(
s.denoise.failed.is_none(),
"toggling off is how a failure is retried"
"choosing bilinear is how a failure is retried"
);
}
#[test]
fn another_network_discards_the_result_and_bilinear_keeps_it() {
let Some(ctx) = headless() else { return };
let mut s =
DevelopSession::open_owned(&ctx, bayer(64, 64), dr_types::Orientation::NORMAL).unwrap();
s.denoise.method = Some(Method::DEFAULT);
landed(&mut s);
let to = |s: &mut DevelopSession, m: Method| {
s.graph
.set_param(learned_denoise::ID, learned_denoise::METHOD, m.index());
s.forget_other_method();
};
to(&mut s, Method::Bilinear);
assert!(s.denoise.result.is_some(), "kept for the way back");
to(&mut s, Method::DEFAULT);
assert!(s.denoise.result.is_some());
to(&mut s, Method::Fast);
assert!(s.denoise.result.is_none(), "another network's picture");
assert_eq!(s.denoise.method, Some(Method::Fast));
}
}
+255
View File
@@ -0,0 +1,255 @@
//! TRACES: FR-DEV-3g
//! The learned denoise's results, kept on disk (docs/dev/denoise.md §7.1).
//!
//! The network takes seconds per photograph and is on by default, so a
//! photograph reopened, or exported after it was developed, must not pay
//! again. A result is the network's output as it is — linear camera RGB at
//! the crop's size — written as half floats: about 120 MB for 20 MP, and no
//! compressor to link on Android. The strength slider is applied afterwards
//! and is not part of the key, so moving it never invalidates anything.
//!
//! **Keyed on the file's bytes and the model**: a SHA-256 of what was
//! decoded, and the model file's name and size. Anything that changes the
//! input or the network changes the key; the edit does not.
//!
//! **Bounded by a budget**, oldest first: a hit refreshes an entry's time, a
//! write evicts what no longer fits. Disposable — a peer of the inference
//! engine's cache, never synced — so an entry that cannot be read is simply
//! recomputed.
use std::io::{Read, Write};
use std::path::{Path, PathBuf};
use sha2::{Digest, Sha256};
/// Bytes the cache may hold before the oldest entries go: about forty 20 MP
/// photographs.
pub const BUDGET: u64 = 5 * 1024 * 1024 * 1024;
const MAGIC: &[u8; 8] = b"DRDN1\0\0\0";
const HEADER: usize = 8 + 4 + 4 + 1;
/// A cached result: the RGB samples, their size, and where the noise
/// figures came from.
pub struct Entry {
pub rgb: Vec<f32>,
pub width: u32,
pub height: u32,
pub source: dr_denoise::Source,
}
/// The cache's directory: beside the inference engine's, under the data
/// root, which is writable on every platform.
pub fn dir() -> PathBuf {
crate::library::inference_cache_dir()
.parent()
.map(|p| p.join("denoise-cache"))
.unwrap_or_else(|| PathBuf::from("denoise-cache"))
}
/// A file's bytes, hashed once at open: each method's network keys its
/// result from this, so changing the method does not read the file again.
#[derive(Clone)]
pub struct FileHash(Sha256);
impl FileHash {
pub fn of(bytes: &[u8]) -> Self {
let mut h = Sha256::new();
h.update(bytes);
FileHash(h)
}
/// The key for these bytes under a model.
pub fn key(&self, model: &Path) -> String {
let mut h = self.0.clone();
if let Some(name) = model.file_name() {
h.update(name.to_string_lossy().as_bytes());
}
let size = std::fs::metadata(model).map(|m| m.len()).unwrap_or(0);
h.update(size.to_le_bytes());
let digest = h.finalize();
digest.iter().map(|b| format!("{b:02x}")).collect()
}
}
fn path_in(dir: &Path, key: &str) -> PathBuf {
dir.join(format!("{key}.drdn"))
}
fn source_code(s: dr_denoise::Source) -> u8 {
match s {
dr_denoise::Source::Table => 0,
dr_denoise::Source::DngProfile => 1,
dr_denoise::Source::Measured => 2,
}
}
fn source_from(code: u8) -> Option<dr_denoise::Source> {
Some(match code {
0 => dr_denoise::Source::Table,
1 => dr_denoise::Source::DngProfile,
2 => dr_denoise::Source::Measured,
_ => return None,
})
}
/// The entry for `key`, if one is held and reads back whole. A hit
/// refreshes its time, so what is in use outlives what is not.
pub fn load(dir: &Path, key: &str) -> Option<Entry> {
let path = path_in(dir, key);
let mut file = std::fs::File::open(&path).ok()?;
let mut head = [0u8; HEADER];
file.read_exact(&mut head).ok()?;
if &head[..8] != MAGIC {
return None;
}
let width = u32::from_le_bytes(head[8..12].try_into().ok()?);
let height = u32::from_le_bytes(head[12..16].try_into().ok()?);
let source = source_from(head[16])?;
let samples = (width as usize)
.checked_mul(height as usize)?
.checked_mul(3)?;
let mut raw = vec![0u8; samples.checked_mul(2)?];
file.read_exact(&mut raw).ok()?;
let rgb = raw
.chunks_exact(2)
.map(|b| half::f16::from_le_bytes([b[0], b[1]]).to_f32())
.collect();
let _ = file.set_modified(std::time::SystemTime::now());
Some(Entry {
rgb,
width,
height,
source,
})
}
/// Keep `entry` under `key`, then evict to `budget`. Written beside its name
/// and renamed, so a reader never sees half a file. Failure only costs a
/// recompute next time, so it is logged and swallowed.
pub fn store(dir: &Path, key: &str, entry: &Entry, budget: u64) {
let result = (|| -> std::io::Result<()> {
std::fs::create_dir_all(dir)?;
let path = path_in(dir, key);
let partial = path.with_extension("part");
let mut out = std::io::BufWriter::new(std::fs::File::create(&partial)?);
out.write_all(MAGIC)?;
out.write_all(&entry.width.to_le_bytes())?;
out.write_all(&entry.height.to_le_bytes())?;
out.write_all(&[source_code(entry.source)])?;
for v in &entry.rgb {
out.write_all(&half::f16::from_f32(*v).to_le_bytes())?;
}
out.into_inner().map_err(|e| e.into_error())?.sync_all()?;
std::fs::rename(&partial, &path)
})();
if let Err(e) = result {
log::warn!("denoise cache: not kept ({e})");
return;
}
evict(dir, budget);
}
/// Remove the oldest entries until what is left fits `budget`.
pub fn evict(dir: &Path, budget: u64) {
let Ok(read) = std::fs::read_dir(dir) else {
return;
};
let mut entries: Vec<(std::time::SystemTime, u64, PathBuf)> = read
.flatten()
.filter(|e| e.path().extension().is_some_and(|x| x == "drdn"))
.filter_map(|e| {
let m = e.metadata().ok()?;
Some((m.modified().ok()?, m.len(), e.path()))
})
.collect();
let mut total: u64 = entries.iter().map(|(_, len, _)| len).sum();
entries.sort_by_key(|(t, _, _)| *t);
for (_, len, path) in entries {
if total <= budget {
break;
}
if std::fs::remove_file(&path).is_ok() {
total -= len;
}
}
}
#[cfg(test)]
mod tests {
use super::*;
fn scratch(name: &str) -> PathBuf {
let d =
std::env::temp_dir().join(format!("dr-denoise-cache-{name}-{}", std::process::id()));
let _ = std::fs::remove_dir_all(&d);
d
}
fn entry(w: u32, h: u32) -> Entry {
Entry {
rgb: (0..w * h * 3).map(|i| i as f32 / 100.0).collect(),
width: w,
height: h,
source: dr_denoise::Source::DngProfile,
}
}
#[test]
fn a_result_comes_back_as_it_went_in_to_half_precision() {
let dir = scratch("roundtrip");
let e = entry(4, 3);
store(&dir, "k", &e, BUDGET);
let back = load(&dir, "k").expect("a hit");
assert_eq!((back.width, back.height), (4, 3));
assert_eq!(back.source, dr_denoise::Source::DngProfile);
for (a, b) in e.rgb.iter().zip(&back.rgb) {
assert!((a - b).abs() <= a.abs() * 1e-3 + 1e-4, "{a} {b}");
}
assert!(load(&dir, "other").is_none());
let _ = std::fs::remove_dir_all(&dir);
}
#[test]
fn the_key_follows_the_bytes_and_the_model() {
let dir = scratch("key");
std::fs::create_dir_all(&dir).unwrap();
let model = dir.join("m.onnx");
std::fs::write(&model, b"weights").unwrap();
let a = FileHash::of(b"photo").key(&model);
assert_eq!(a, FileHash::of(b"photo").key(&model));
assert_ne!(a, FileHash::of(b"photo2").key(&model));
std::fs::write(&model, b"other weights").unwrap();
assert_ne!(a, FileHash::of(b"photo").key(&model));
let _ = std::fs::remove_dir_all(&dir);
}
#[test]
fn the_oldest_entries_go_first_past_the_budget() {
let dir = scratch("evict");
let e = entry(10, 10);
store(&dir, "old", &e, BUDGET);
let one = std::fs::metadata(path_in(&dir, "old")).unwrap().len();
std::thread::sleep(std::time::Duration::from_millis(20));
store(&dir, "mid", &e, BUDGET);
std::thread::sleep(std::time::Duration::from_millis(20));
// A hit makes the oldest the most recent.
load(&dir, "old").unwrap();
std::thread::sleep(std::time::Duration::from_millis(20));
store(&dir, "new", &e, 2 * one);
assert!(load(&dir, "mid").is_none(), "the least recently used went");
assert!(load(&dir, "old").is_some() && load(&dir, "new").is_some());
let _ = std::fs::remove_dir_all(&dir);
}
#[test]
fn a_damaged_entry_is_a_miss() {
let dir = scratch("damaged");
store(&dir, "k", &entry(4, 4), BUDGET);
let p = path_in(&dir, "k");
let bytes = std::fs::read(&p).unwrap();
std::fs::write(&p, &bytes[..bytes.len() / 2]).unwrap();
assert!(load(&dir, "k").is_none());
let _ = std::fs::remove_dir_all(&dir);
}
}
+1
View File
@@ -18,6 +18,7 @@
mod curves;
mod denoise;
mod denoise_cache;
mod framing;
mod history;
mod mask_ops;
+2 -1
View File
@@ -942,7 +942,8 @@ mod tests {
dr_pipeline::ParamId("blue_sat"),
)
};
assert_eq!(blue(), Some(58.0));
// Blue is shared across our bands at measured factors; blue takes 0.75.
assert_eq!(blue(), Some(58.0 * 0.75));
assert!(!s.adopt_earlier_edit(), "taken: a second call does nothing");
assert!(!s.graph.is_neutral());
s.undo();
+16 -9
View File
@@ -34,8 +34,10 @@ pub fn init(runtime_dirs: Vec<PathBuf>) {
(Role::EyeClassifier, crate::library::EYE_MODEL),
(Role::EyeClassifier, crate::library::SUNGLASSES_MODEL),
(Role::Inpainter, crate::library::INPAINT_MODEL),
(Role::Denoiser, crate::library::DENOISE_MODEL),
]);
wanted.extend(
[dr_denoise::FAST, dr_denoise::MEDIUM, dr_denoise::BEST].map(|n| (Role::Denoiser, n.file)),
);
let models: Vec<(Role, PathBuf)> = wanted
.into_iter()
.filter_map(|(role, name)| Some((role, crate::library::shared_model(name)?)))
@@ -46,12 +48,14 @@ pub fn init(runtime_dirs: Vec<PathBuf>) {
cache_dir: crate::library::inference_cache_dir(),
models,
embedded: {
let [landscape, portrait] = dr_pano::xfeat::embedded_model_bytes();
vec![
(Role::Segmenter, dr_segment::embedded_model_bytes()),
(Role::Keypoints, landscape),
(Role::Keypoints, portrait),
]
let [landscape, portrait] = dr_pano::xfeat::embedded_models();
let tag = |role| move |(form, bytes)| (role, form, bytes);
dr_segment::embedded_models()
.into_iter()
.map(tag(Role::Segmenter))
.chain(landscape.into_iter().map(tag(Role::Keypoints)))
.chain(portrait.into_iter().map(tag(Role::Keypoints)))
.collect()
},
ceiling: None,
threads: 0,
@@ -81,7 +85,7 @@ pub fn user_runtime_dir() -> PathBuf {
///
/// Reads the shared and system directories only. An account-private model
/// directory can override the file `library::face_models` loads, but not
/// which form the backend wants, and the int8 sibling is something a
/// which form the backend wants, and the quantised sibling is something a
/// packager ships, not something a user drops in.
pub fn detector_form(detector: FaceDetector) -> Form {
let canonical = crate::library::shared_model(detector.file_name())
@@ -92,8 +96,11 @@ pub fn detector_form(detector: FaceDetector) -> Form {
/// The `faces.model_id` this device indexes under with `detector`.
pub fn model_id(detector: FaceDetector) -> &'static str {
match detector_form(detector) {
Form::F32 => detector.model_id(),
Form::Int8 => detector.model_id_int8(),
Form::A16W8 => detector.model_id_a16w8(),
// No detector is offered in A16W16 (inference.md §1.5); were one, it
// would be the network f32 is to the last bit that a person can see.
Form::F32 | Form::A16W16 => detector.model_id(),
}
}
+9 -1
View File
@@ -223,7 +223,15 @@ fn catalogued(key: &str) -> Option<&'static str> {
// panel under its own name — see `rows_filtered`.
"param.lens_profile.apply" => "Apply",
"param.learned_denoise.apply" => "Apply",
"param.learned_denoise.grain" => "Keep grain",
// Which demosaic: the classical one, or a network by how long it takes.
"param.learned_denoise.method" => "Method",
"param.learned_denoise.method.bilinear" => "Bilinear",
"param.learned_denoise.method.fast" => "Fast",
"param.learned_denoise.method.medium" => "Medium",
"param.learned_denoise.method.best" => "Best",
// How strongly: 100 % is the network's result, and less puts the
// removed noise's brightness back as grain.
"param.learned_denoise.strength" => "Strength",
"param.camera_profile.apply" => "Use Profile",
// The LookTable's strength.
"param.camera_profile.look" => "Look Amount",
+34 -5
View File
@@ -39,6 +39,20 @@ pub fn place_path(account: &Account) -> PathBuf {
data_root().join(account.namespace()).join("place.json")
}
/// TRACES: FR-DEV-6
/// The preset library as this device and this library's server last agreed
/// on it — the base `PresetLibrary::merge` decides deletions against.
///
/// Per library, beside the place, because each library's server keeps its
/// own copy. In the data directory for the reason the place is: a base the
/// system deleted would turn the next exchange into a union, and every
/// preset deleted since the last one would come back.
pub fn presets_base_path(account: &Account) -> PathBuf {
data_root()
.join(account.namespace())
.join("presets.base.drpl")
}
/// The directory every account's data hangs off.
///
/// **Not the cache directory, and on Android that distinction is the whole
@@ -376,13 +390,28 @@ pub fn inpaint_model() -> Option<PathBuf> {
}
/// TRACES: FR-DEV-3g
/// The learned demosaic and denoise, as shipped in `models/denoise/`.
pub const DENOISE_MODEL: &str = "mosaic-1408.onnx";
/// The network a denoise method runs, as shipped in `models/denoise/`;
/// `None` for the classical demosaic, which runs none.
pub fn denoise_network(
method: dr_pipeline::learned_denoise::Method,
) -> Option<dr_denoise::Shipped> {
use dr_pipeline::learned_denoise::Method;
match method {
Method::Bilinear => None,
Method::Fast => Some(dr_denoise::FAST),
Method::Medium => Some(dr_denoise::MEDIUM),
Method::Best => Some(dr_denoise::BEST),
}
}
/// TRACES: FR-DEV-3g
/// Where the denoise model is, by the border filler's search.
pub fn denoise_model() -> Option<PathBuf> {
shared_model(DENOISE_MODEL)
/// Where a method's network is, by the border filler's search, and the
/// context it needs.
pub fn denoise_model(
method: dr_pipeline::learned_denoise::Method,
) -> Option<(PathBuf, dr_denoise::Shipped)> {
let net = denoise_network(method)?;
Some((shared_model(net.file)?, net))
}
#[cfg(test)]
+5
View File
@@ -290,6 +290,10 @@ pub struct LibraryController {
/// old behaviour rather than a crash.
pub(super) coll_ctl:
RefCell<Option<std::rc::Weak<crate::collections_ui::CollectionsController>>>,
/// TRACES: FR-DEV-6
/// The develop view's preset list, so a sync that brought presets from
/// another device can redraw it. `Weak` for the reason `coll_ctl` is.
pub(crate) named_presets: RefCell<Option<std::rc::Weak<crate::presets::NamedPresets>>>,
/// Timeline view state: how far zoomed in, and around what instant.
///
/// Zoom is a level rather than a span so the axis halves and doubles in
@@ -497,6 +501,7 @@ impl LibraryController {
applying_place: std::cell::Cell::new(false),
viewing_trash: std::cell::Cell::new(false),
coll_ctl: RefCell::new(None),
named_presets: RefCell::new(None),
timeline_zoom: RefCell::new(0),
timeline_centre: RefCell::new(None),
pinch_accum: RefCell::new(1.0),
+20
View File
@@ -123,6 +123,12 @@ pub(super) fn start_derived_sync(window: &AppWindow, ctl: &Rc<LibraryController>
library::thumbs_dir(&conn.account),
catalog_path,
library::place_path(&conn.account),
crate::derived_sync::PresetFiles {
library: crate::preset_store::PresetStore::open()
.path()
.to_path_buf(),
base: library::presets_base_path(&conn.account),
},
scratch,
ctl.face_model_id.borrow().clone(),
);
@@ -237,6 +243,20 @@ pub(super) fn start_derived_sync(window: &AppWindow, ctl: &Rc<LibraryController>
}
}
}
// TRACES: FR-DEV-6
// Presets saved on another device, into the list the
// develop view is drawing.
if report.presets_adopted {
let named = ctl_cb
.named_presets
.borrow()
.as_ref()
.and_then(|n| n.upgrade());
if let Some(named) = named {
named.reload();
crate::presets::render_named(&w, &named);
}
}
stop(&ctl_cb.sync_timer);
return;
}
+36 -23
View File
@@ -27,17 +27,14 @@ pub struct PresetStore {
impl PresetStore {
/// Open the store at the platform config location.
///
/// Linux: `$XDG_CONFIG_HOME/darkroom/presets.drpl`, falling back to
/// `~/.config` — the same resolution `SettingsStore` does, so the files sit
/// together and a user backing up one takes all of them.
/// Beside `settings.json`, in [`dr_sync::account::config_dir`], so a user
/// backing up one takes all of them. Never from `HOME` directly: Android
/// declares its data directory rather than setting `HOME`, and a path
/// built from an empty one is `/.config`, which is read-only — every
/// preset saved on the tablet failed. Windows sets no `HOME` either, and
/// got a directory relative to wherever the app was started.
pub fn open() -> Self {
let dir = std::env::var_os("XDG_CONFIG_HOME")
.map(PathBuf::from)
.unwrap_or_else(|| {
PathBuf::from(std::env::var("HOME").unwrap_or_default()).join(".config")
})
.join("darkroom");
Self::open_at(dir.join(format!(
Self::open_at(dr_sync::account::config_dir().join(format!(
"presets.{}",
dr_pipeline::preset::LIBRARY_EXTENSION
)))
@@ -66,19 +63,16 @@ impl PresetStore {
/// a preset, at which point they have chosen to. Nothing here deletes the
/// file, and the warning names the path so it can be recovered by hand.
pub fn load(&self) -> PresetLibrary {
match std::fs::read_to_string(&self.path) {
Ok(text) => match PresetLibrary::parse(&text) {
Ok(library) => library,
Err(e) => {
log::warn!(
"{} is not a readable preset library ({e}); \
starting empty, the file is left alone",
self.path.display()
);
PresetLibrary::default()
}
},
Err(e) if e.kind() == std::io::ErrorKind::NotFound => PresetLibrary::default(),
match self.try_load() {
Ok(library) => library.unwrap_or_default(),
Err(PresetStoreError::Unreadable(e)) => {
log::warn!(
"{} is not a readable preset library ({e}); \
starting empty, the file is left alone",
self.path.display()
);
PresetLibrary::default()
}
Err(e) => {
log::warn!("reading {}: {e}; starting empty", self.path.display());
PresetLibrary::default()
@@ -86,6 +80,23 @@ impl PresetStore {
}
}
/// The stored library, `None` when there is no file, or why it could not
/// be read.
///
/// For a caller that must not mistake a file it could not read for an
/// empty library. The sync is one: to it an empty library next to a
/// non-empty one from the last exchange means "every preset was deleted",
/// and it would carry that to every other device.
pub fn try_load(&self) -> Result<Option<PresetLibrary>, PresetStoreError> {
match std::fs::read_to_string(&self.path) {
Ok(text) => PresetLibrary::parse(&text)
.map(Some)
.map_err(|e| PresetStoreError::Unreadable(e.to_string())),
Err(e) if e.kind() == std::io::ErrorKind::NotFound => Ok(None),
Err(e) => Err(e.into()),
}
}
/// Persist the library, replacing whatever was there.
pub fn save(&self, library: &PresetLibrary) -> Result<(), PresetStoreError> {
if let Some(parent) = self.path.parent() {
@@ -108,6 +119,8 @@ impl PresetStore {
pub enum PresetStoreError {
#[error("preset library io: {0}")]
Io(#[from] std::io::Error),
#[error("{0}")]
Unreadable(String),
}
#[cfg(test)]
+86 -7
View File
@@ -723,6 +723,14 @@ pub fn wire(
pub struct NamedPresets {
store: PresetStore,
library: RefCell<PresetLibrary>,
/// The file as this window last read or wrote it.
///
/// The sync rewrites the file from another thread when presets arrive
/// from another device. Saving the in-memory copy over that would delete
/// them here, and the next exchange would carry the deletion everywhere.
/// So a save first checks whether the file still holds this, and merges
/// with it as the base if not (`PresetLibrary::merge`).
on_disk: RefCell<PresetLibrary>,
/// The category folders the photographer has opened, by key (see
/// [`PresetTree`]). Every folder starts closed, so a list of seventy
/// presets arrives as six lines; for the session, not saved, as the
@@ -752,16 +760,35 @@ impl NamedPresets {
// writes the file without them, and until then nothing is lost,
// because each one is identical to the preset now shipped in its
// place.
let mut library = store.load();
let on_disk = store.load();
let library = Self::shown(&on_disk);
Rc::new(Self {
store,
library: RefCell::new(library),
on_disk: RefCell::new(on_disk),
open: RefCell::default(),
})
}
/// What the window lists from what the file holds.
fn shown(on_disk: &PresetLibrary) -> PresetLibrary {
let mut library = on_disk.clone();
let forgotten = dr_pipeline::bundled::forget_unchanged_copies(&mut library);
if forgotten > 0 {
log::info!("{forgotten} seeded preset copies are now shipped presets");
}
Rc::new(Self {
store,
library: RefCell::new(library),
open: RefCell::default(),
})
library
}
/// TRACES: FR-DEV-6
/// Read the file again, after a sync brought presets from another device.
///
/// Every change is saved as it is made, so there is nothing in memory
/// this could discard.
pub fn reload(&self) {
let on_disk = self.store.load();
*self.library.borrow_mut() = Self::shown(&on_disk);
*self.on_disk.borrow_mut() = on_disk;
}
/// Read presets from a file or a folder and store them.
@@ -930,9 +957,22 @@ impl NamedPresets {
/// the write failed and leave what is on screen matching what is in
/// memory until the next successful save.
fn persist(&self, rollback: impl FnOnce(&mut PresetLibrary)) -> Result<(), SaveError> {
// A file that changed since it was read here was the sync; take what
// it brought rather than saving over it. One that will not read is
// left to the save, which says so.
if let Ok(Some(now)) = self.store.try_load() {
if now != *self.on_disk.borrow() {
let merged =
PresetLibrary::merge(&self.on_disk.borrow(), &self.library.borrow(), &now);
*self.library.borrow_mut() = merged;
}
}
let result = self.store.save(&self.library.borrow());
match result {
Ok(()) => Ok(()),
Ok(()) => {
*self.on_disk.borrow_mut() = self.library.borrow().clone();
Ok(())
}
Err(e) => {
rollback(&mut self.library.borrow_mut());
// Named, the way the settings page names its file: "could not
@@ -1188,6 +1228,7 @@ pub fn wire_named(
open,
} = develop;
*library.named_presets.borrow_mut() = Some(Rc::downgrade(&named));
render_named(window, &named);
// --- open and close a category folder ---------------------------------
@@ -1974,6 +2015,44 @@ mod tests {
assert_eq!(reloaded.names(), vec!["Warm".to_string()]);
}
#[test]
fn a_save_keeps_what_the_sync_wrote_since_the_window_read_the_file() {
// TRACES: FR-DEV-6
// The sync rewrites the file from its own thread. Saving the window's
// copy over it would delete what arrived, and the next exchange would
// carry that deletion to the device that sent it.
let (presets, dir) = named("sync-wrote-meanwhile");
let path = dir.join("presets.drpl");
let mut arrived = PresetLibrary::default();
arrived
.insert("From the tablet", Preset::default())
.unwrap();
PresetStore::open_at(path.clone()).save(&arrived).unwrap();
presets.insert("Warm", Preset::capture(&edited())).unwrap();
let reloaded = NamedPresets::open_at(path);
assert_eq!(
reloaded.names(),
vec!["From the tablet".to_string(), "Warm".to_string()]
);
}
#[test]
fn a_reload_lists_what_the_sync_brought() {
// TRACES: FR-DEV-6
let (presets, dir) = named("reload-after-sync");
let mut arrived = PresetLibrary::default();
arrived
.insert("From the tablet", Preset::default())
.unwrap();
PresetStore::open_at(dir.join("presets.drpl"))
.save(&arrived)
.unwrap();
presets.reload();
assert_eq!(presets.names(), vec!["From the tablet".to_string()]);
}
#[test]
fn a_saved_preset_carries_the_edit_it_captured() {
let (presets, _dir) = named("carries-the-edit");