1206 Commits
Author SHA1 Message Date
dtourolle 8aa10cd249 Release 0.24.0
Benchmarks / CPU and I/O (per commit) (push) Successful in 8m56s
Benchmarks / Frame budget (on demand) (push) Skipped
Traceability / Requirement traces (push) Successful in 1m27s
Build and test / Desktop (Linux) (push) Successful in 1h33m5s
Build and test / Layer separation (push) Successful in 32s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 4s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 2s
Build and test / Android (aarch64) (push) Successful in 45m54s
Build and test / Windows (x86_64, cross) (push) Successful in 55m18s
Build and test / Publish the release (push) Successful in 1m33s
v0.24.0
2026-10-07 07:27:52 -04:00
dtourolle c2cfacd7d3 Record the new Best in the spec and the manual
denoise.md §15: why Best became one network, how it compares with the
mixture and Medium on real photographs and the chart, the candidates that
fell short, the file names, the saved-edit numbering, and the timings on
the 3050 -- 0.51-0.54 s whole-frame, 0.95 s in tiles, against 2.60 s for
the mixture in tiles. The manual lists Bilinear, Fast and Best, says an
edit made with Medium opens with Best, and gives the new time; Medium's
close-up goes.
2026-10-07 07:14:52 -04:00
dtourolle a9271c4850 Make Best one network, and retire Medium and the mixture
Best was a mixture of two experts and a gate, 110 GMAC a megapixel;
Medium a single network at 48 that was softer on real edges. fb-combo
(darkroom-denoise, 20 000 steps from fb-edges2, taught by the mixture
with a quarter of its crops from the edge-rich parts of the frames) is
Medium's shape and holds the mixture's edges on real photographs:
edge PSNR within 0.04-0.06 dB at ISO 1600/6400/25600, more sharpness
kept at all three, the chart's edge 0.89 photosites wide against 0.82.
It is 0.27 dB short on smooth areas at ISO 25600. It becomes Best, and
the methods are Bilinear, Fast and Best.

Saved edits keep their numbers: 2, which was Medium, is now Best, and
3, which was Best, is past the end and reads as the default, Best.
The network ships as mosaic-hq, a new name: the result cache keys a
model by name and size, and this one is byte for byte the old Medium's
size. Its tablet form (A16W16) lost 0.00 dB in simulated QDQ at every
ISO and at most 0.09 dB across the noise bracket.
2026-10-07 06:58:24 -04:00
dtourolle 4b71ef0947 Record whole-frame denoise in the spec and the model licences
denoise.md §14: why the tiles waste half of Best's work, where the any-size
networks run and why only there, why the limit is the card's memory, and
the measurement on _MG_8862 — 2.60 s in 1408 tiles, 1.37 s in two
4160 x 3248 tiles, the outputs within fp16's own spread. models/LICENCE.md
lists the three re-exports.
2026-10-06 23:14:11 -04:00
dtourolle 1e1aa1442b Size the whole-frame profile for a 6 GB card
TensorRT plans its memory for the profile's largest shape, and up to a
whole 6D frame with Best's border (4608 x 6656) it asked for 4.9-5.9 GB
and would not build on the RTX 3050. The profile now ends at 4608 x 3328
(15 MP), tuned for 4160 x 3248, and the tiler cuts a 6D frame into two
such tiles: 27 MP of work for 20 MP kept, against 49 MP in 1408 tiles.
The engine's directory names the profile, so a later range never loads
an engine built for this one.
2026-10-06 23:14:10 -04:00
dtourolle 3761281dd1 Ship the denoise networks with any height and width
mosaic-{fast,medium,best}.onnx are the shipped networks re-exported by
darkroom-denoise tools/export_whole.py (22ea648) from the checkpoints
the 1408 files came from: identical to them at 1408 (max |d| = 0), to
torch at 592 x 848, and to tiled inference over the reflected frame in
f64. 42 MB together. The Arch package and the Windows installer carry
them beside the fixed files; the APK leaves them out, since the Hexagon
takes fixed shapes only.
2026-10-06 21:47:20 -04:00
dtourolle 9cba420fd5 Denoise a whole frame in one call where the GPU takes any size
A fixed 1408 tile is exact only in its centre, and Best keeps 896 of
every 1408 it computes: 2.47 photosites of work for each one kept. The
tiler now takes a network of any size as well as a square one, and
plans the frame as the fewest equal tiles under the rung's limit --
one tile, the whole frame and its reflected border, whenever it fits.
If the first call of a plan fails, as a GPU out of memory does, the
kept centre is halved and the frame planned again.

Each shipped network names its any-size sibling (mosaic-best.onnx
beside mosaic-best-1408.onnx). OnnxNet::open takes it where the engine
runs whole frames and the file is installed, and the 1408 tiles
otherwise; open_tiled forces the tiles, and denoise_raw's DR_PLAN=tiles
uses it to compare. The cache key stays on the fixed model: the output
is the same network's. Tests hold any-size tiles, a grid of them and a
plan rebuilt after a failure to the square tiles' answer in every Bayer
phase.
2026-10-06 21:47:19 -04:00
dtourolle 56f4180347 Run a denoiser of any input size on TensorRT and CUDA
Role::WholeDenoiser is the denoise network exported with any height and
width, for a whole frame instead of 1408 tiles whose borders are thrown
away. It is served only where a new size costs nothing: the CUDA
provider, and TensorRT through an optimisation profile from 256 to
4608 x 6656, tuned for the 6D's frame with Best's border. Everywhere
else whole_frame_limit() says None and the fixed tiles run.

ort's TensorRT builder has no profile options, so the engine registers
through the runtime's V2 options with the names 1.30 reads
(trt_profile_{min,opt,max}_shapes). Without a profile a dynamic input
compiled an engine per size at run time, 156 s on the first frame. The
engine lives in its own directory per model: ORT's cache key leaves the
shape out, and the fixed 1408 export and its any-size sibling are the
same graph.
2026-10-06 21:41:31 -04:00
dtourolle 417cba8b4d Silence the hardware module where no runtime is loaded from disk
hardware::detect is read only by api::install_best, which exists with
the native feature; a build of a crate that takes the engine without it
(dr-denoise's own tests) warned that all of it was unused.
2026-10-06 21:41:13 -04:00
dtourolle 7966bf2dd8 Release 0.23.0
Benchmarks / CPU and I/O (per commit) (push) Successful in 8m42s
Benchmarks / Frame budget (on demand) (push) Skipped
Traceability / Requirement traces (push) Successful in 1m22s
Build and test / Android (aarch64) (push) Successful in 47m15s
Build and test / android-image (push) Successful in 3s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / Desktop (Linux) (push) Successful in 1h31m5s
Build and test / windows-image (push) Successful in 3s
🐳 Windows image / Build and push (push) Successful in 2s
Build and test / Layer separation (push) Successful in 53s
Build and test / Windows (x86_64, cross) (push) Successful in 55m3s
Build and test / Publish the release (push) Successful in 1m58s
v0.23.0
2026-10-04 22:28:19 -04:00
dtourolle 4822991bec Read an unknown SoC as Qualcomm's
The generic runtime is opened only when the QNN build does not fit, and
the fit rests on ro.soc.manufacturer. A property the app cannot read, or
an Android older than 12 that has none, read as "not Qualcomm" would
put a Qualcomm device on the generic rung and off its Hexagon — what
0.22.1 had just fixed. Only a device that names another vendor is now
not Qualcomm's; the tablet reports QTI.
2026-10-04 22:26:56 -04:00
dtourolle 4b33754482 Warm a rung up before the probe times it
One warm-up run and the median of three: on the Iris Xe the OpenVINO
rung lost to the CPU on the smallest detector in two probes of three,
because an idle integrated GPU takes a few runs to raise its clock —
warm, it is 5.8 ms against 9.5. Three warm-ups and the median of seven
took it in five probes of five (5.6–7.2 ms against 7.7–13.1). The extra
runs cost tens of milliseconds, once per fingerprint.
2026-10-04 21:00:15 -04:00
dtourolle b3999bbc0c Record the Intel and generic rungs in the inference spec
§1.6 is the Iris Xe measurement: OpenVINO fp16 1.3–5.8× the CPU
provider on every shipped model, WebGPU behind it everywhere but the
denoiser and MI-GAN. §2's ladder gains the Intel and generic rows, the
generic one footnoted as unmeasured where it is meant to help. §3.1
lists the two bundled runtimes' licences; §3.2 is how one runtime of
several is chosen per process. D13 notes the bundling.
2026-10-04 21:00:15 -04:00
dtourolle 43402bfe5b Give a phone without a Qualcomm SoC the generic WebGPU runtime
The APK's ONNX Runtime is the QNN build, which carries no WebGPU, so a
non-Qualcomm phone had nothing above the CPU provider. The APK now also
carries Microsoft's stock onnxruntime-android 1.29.0 (32 MB) as
libonnxruntime_generic.so, and the app offers it after the QNN build.

The runtime search stops at a perfect fit, so on a Qualcomm device the
QNN build — listed first — is all that is opened, and the generic build
never loads beside it. A Qualcomm SoC is read from ro.soc.manufacturer
or, before Android 12, from Qualcomm's FastRPC library being present:
a Qualcomm device mistaken for another would trade its Hexagon for the
generic rung.
2026-10-04 21:00:15 -04:00
dtourolle cb97ebe7ac Ship the OpenVINO and WebGPU runtimes in every desktop package
The Windows installer and the Flatpak carried no ONNX Runtime, so they
ran every model on tract's one core; the Arch package left it to an
optional dependency. Each now installs two builds under runtimes/ —
Intel's OpenVINO build and the generic WebGPU one, both with the CPU
provider — fetched by tools/fetch-bundled-runtimes.sh from PyPI wheels
pinned by SHA-256, pruned to the native libraries (81 + 31 MB on Linux,
67 + 42 MB on Windows), licence texts beside them.

darkroom-desktop searches runtimes/openvino and runtimes/webgpu under
each place a package installs to; the engine opens all it finds and
keeps the one that fits the GPU, so a CUDA or ROCm runtime installed
beside them still wins on its vendor's card. On Windows the chosen
runtime's directory goes on PATH, because Intel's build leaves OpenVINO's
DLLs for the loader to find there.

The Windows image gains unzip; the installer smoke test checks both
runtimes landed.
2026-10-04 21:00:15 -04:00
dtourolle 2ced6f114f Name the rung that lost, not one the runtime lacks
A device left on the CPU gave the first failure as the reason, which on
any runtime but NVIDIA's is "TensorRT execution provider is not enabled
in this build". The reason is now the last rung that was tried and lost,
"WebGPU 150.8 ms, slower than the CPU's 26.6 ms"; the full list is still
in the status's failures.
2026-10-04 21:00:15 -04:00
dtourolle 85dee4375b Load the runtime that fits the GPU, not the first one found
A runtime carries one vendor's providers, only one loads per process,
and a device can now hold several: the package's OpenVINO or WebGPU
build, a CUDA build the user fetched, the distribution's ROCm build.
`api::install` opens each it finds, lists its providers with
GetAvailableProviders, and installs the one scoring highest against the
GPUs `hardware::detect` reads from files — a vendor rung on its own
vendor's GPU above OpenVINO on an Intel one above the generic WebGPU
rung above a CPU-only build. Equal scores keep the old first-found
order, and DARKROOM_ORT_DIR still wins outright. The losers stay mapped
rather than unloaded.

The Linux fingerprint now names the OpenCL drivers too, so installing
Intel's re-probes. `ladder` takes DARKROOM_ORT_DIRS to show the choice.
2026-10-04 21:00:15 -04:00
dtourolle 87c405eb46 Add OpenVINO and WebGPU rungs to the inference ladder
OpenVINO is the Intel rung: the integrated or Arc GPU, fp16 for every
role but the embedder, a compiled program per model kept in a directory
per model, precision and runtime version. On the Iris Xe it beats ONNX
Runtime's CPU provider on every shipped model — scrfd_10g 23 ms against
58, the scene model 17 against 57, MI-GAN 57 against 330, a denoise tile
40 against 158.

WebGPU is the generic rung for a GPU no vendor rung covers. It was slower
than the CPU on the Iris Xe, the RTX 3050 and the Adreno, so it is on the
ladder for the GPUs it has not been timed on, behind the probe's clock.

MIGraphX's registration becomes one generic key/value helper that all
three share, with option names read from each runtime's own source.
2026-10-04 21:00:15 -04:00
dtourolle 555ec0efb3 Let ep_probe name the OpenVINO GPU
On the hybrid laptop OpenVINO's `GPU` was the RTX 3050 through NVIDIA's
OpenCL, not the Iris Xe; DARKROOM_OV_GPU picks GPU.0, GPU.1 and so on.
2026-10-04 21:00:15 -04:00
dtourolle 0291b80672 Feed every input in ep_probe
The denoiser takes `mosaic` and `sigma`; with only the first fed, every
provider reported the same failure and the model went unmeasured.
2026-10-04 21:00:15 -04:00
dtourolle ef71bb3289 Time OpenVINO and WebGPU in ep_probe
The Intel and vendor-neutral rungs need a measurement before they join
the ladder (docs/dev/inference.md §2). Both register through the generic
key/value entry point with the option names ONNX Runtime reads at the
wheel's version: OpenVINO 1.24 (`openvino_provider_factory.cc`), WebGPU
1.27 (`webgpu_provider_options.h`, prefixed by the runtime).
DARKROOM_EPS narrows the list to the families a runtime carries.
2026-10-04 21:00:15 -04:00
dtourolle 08b7d23e86 Release 0.22.1
Benchmarks / CPU and I/O (per commit) (push) Successful in 8m35s
Benchmarks / Frame budget (on demand) (push) Skipped
Traceability / Requirement traces (push) Successful in 1m22s
Build and test / Android (aarch64) (push) Successful in 46m43s
Build and test / android-image (push) Successful in 2s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / Desktop (Linux) (push) Successful in 1h4m43s
Build and test / windows-image (push) Successful in 2s
🐳 Windows image / Build and push (push) Successful in 2s
Build and test / Layer separation (push) Successful in 34s
Build and test / Windows (x86_64, cross) (push) Successful in 29m53s
Build and test / Publish the release (push) Successful in 1m24s
v0.22.1
2026-10-04 20:35:01 -04:00
dtourolle a116325991 Declare the DSP's RPC library, so the app can reach the Hexagon
QNN's Hexagon stub loads libcdsprpc.so, a vendor library, and from API 31
an app's linker namespace refuses a vendor library its manifest does not
name. QNN then fails to create its device - QNN_DEVICE_ERROR_INVALID_CONFIG,
before it reaches the DSP - and every model ran on the CPU on 0.22.0: AI
denoise took 30-131 s a photograph on the tablet.

The same engine code ran on the HTP from adb's shell, whose namespace has
no such rule, which is what hid it. With the declaration the app's probe
chose the Hexagon on the tablet and compiled every model for it. Not
required, so a device without the library still installs and runs on the
CPU. A test holds the line in the manifest.
2026-10-04 20:19:04 -04:00
dtourolle f20e481358 Probe the Hexagon strictly, and probe again after falling back to the CPU
0.22.0's first launch on the tablet: QNN could not create its device
(QNN_DEVICE_ERROR_INVALID_CONFIG), the session built anyway with every
node on the CPU behind the provider, and the probe timed that - 28.5 ms
against the CPU's own 19.4 - and rejected the Hexagon. The verdict was
cached under the fingerprint, so every model stayed on the CPU on every
later launch: AI denoise took 30-131 s a photograph instead of seconds.
The same A16W8 detector with the APK's own libraries runs on the HTP in
4.4 ms.

The probe's Hexagon session now sets session.disable_cpu_ep_fallback, so
a device that cannot take the graph fails the probe instead of being timed
as the CPU. Only the probe: shipped graphs may keep nodes on the CPU on
purpose. And a selection that fell back to the CPU after an accelerator
failed or lost is probed again on the next launches, up to three probes
per fingerprint; a cache written by 0.22.0 reads as never retried, so the
tablet probes again once this is installed.
2026-10-04 19:42:15 -04:00
dtourolle 16a5957aa7 Compile dr-ui on sixteen codegen units in release builds
Slint expands the .slint files into ~27 MB of Rust, and at the workspace's
single codegen unit LLVM optimised all of it on one thread: 13.5 minutes
of a release build with the other cores idle. The override applies to
dr-ui alone; the image crates keep one unit, and thin LTO still runs at
link time.
2026-10-04 10:45:43 -04:00
dtourolle 5736a21a3a Release 0.22.0
Benchmarks / CPU and I/O (per commit) (push) Successful in 8m36s
Benchmarks / Frame budget (on demand) (push) Skipped
Traceability / Requirement traces (push) Successful in 1m31s
Build and test / Android (aarch64) (push) Successful in 48m34s
Build and test / android-image (push) Successful in 4s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / Desktop (Linux) (push) Successful in 1h32m33s
Build and test / windows-image (push) Successful in 2s
🐳 Windows image / Build and push (push) Successful in 2s
Build and test / Layer separation (push) Successful in 29s
Build and test / Windows (x86_64, cross) (push) Successful in 31m14s
Build and test / Publish the release (push) Successful in 1m9s
v0.22.0
2026-10-04 08:13:35 -04:00
dtourolle b6ca7be185 Show each denoise method in the manual as a close-up
The manual's AI denoise section names the four methods and their measured
times, and shows the lamp and railing of the ISO 8000 frame at 1:1 by each
in place of the film and the before/after pair. The scene clicks each
method and waits for that network's result: the repair now logs its own
"learned denoise:" line first, so the wait matches the result's.
2026-10-04 08:11:24 -04:00
dtourolle 06422a07db Offer three denoise networks and a method to choose between them
AI Denoise's Apply switch becomes Method: Bilinear, Fast, Medium, Best,
default Best, so an untouched raw writes nothing and develops through the
mixture. `apply` is still read and never written: 0 is Bilinear, 1 keeps
a network already chosen.

- Best is the mixture of a flat and an edge expert with a learned gate;
  Medium and Fast are students distilled from it. 2.48 s, 0.79 s and
  0.57 s for a 20 MP frame on TensorRT fp16.
- Each network carries its own tile border (256 for the mixture, 192 for
  the students) through `dr_denoise::Shipped` and `TileNet::halo`.
- The file is hashed once at open and each network keys its own cached
  result; Bilinear keeps the result in memory for the way back.
- Each has an .a16w16 sibling for the Hexagon: 0.00 dB on the 6D gate,
  at most 0.11 dB with the noise scaled x0.5 to x4.
- APK BUNDLED 19 -> 23; the PKGBUILD installs all three.
2026-10-04 08:02:25 -04:00
dtourolle 14f08a565f Feed the denoise network without making it wait for the CPU
A 20 MP frame spent 0.32 s outside the network: each tile's mosaic and
sigma gathered on one thread, then its 24 MB output copied out of the
runtime and back into the frame, all in series with the device. Tiles are
now gathered on every core by a producer thread one tile ahead, so the
gather overlaps the run; the centre is written back across cores; and the
tile interface hands its inputs over and lends its output, so neither
side is copied. With a stand-in network that does nothing, the tiler's own
time falls to 0.14 s at the 1408 tile and 0.09 s at 2048. The exactness
and Bayer-phase tests are unchanged and pass.
2026-10-04 07:30:09 -04:00
dtourolle 0c9d564586 Repair photosites beyond 8 sigma of every neighbour before the network
The app's hot-pixel pass takes gross defects only; at ISO 6400-25600 a 6D
frame keeps 1000-2000 photosites more than 8 sigma beyond all their
same-colour and adjacent neighbours, which the network turned into specks.
The same two tests with the threshold in the photosite's own sigma, plus
the factor of two that keeps a bright point of light (where 8 sigma is a
sliver of the signal). The next model is trained behind exactly this; on
an ISO 25600 frame the Rust and training code both repair 935.
2026-10-04 07:30:09 -04:00
dtourolle 75e12441fd Share Lightroom's saturation bands across ours at measured strengths
Photographs opened with an earlier Lightroom edit now import its HSL
saturation as fitted against the library's own Lightroom 6 exports, rather
than one band to one band.

Measured on two looks' exports and their raws (darkroom-lrfit, hsl_map_fit),
by encoded hue: Lightroom's saturation bands act about 45 degrees either
side on our wheel, wider than ours, and not all at our strength. Each is now
shared between two or three of our bands — Aqua mostly cyan and azure, where
skies are; Blue mostly blue and violet; Orange, where skin is, at about 0.4
of its value. Values add when two of Lightroom's bands share one of ours.
On the measured skies the import now lifts muted sky blues about 1.9× against
Lightroom's 2.1×, where it gave 1.15×. Hue and luminance still go one band to
the band of the same hue; they were not measured.
2026-10-04 05:28:33 -04:00
dtourolle 3f8f909e41 Make the colour mixer's saturation reach muted colours
Every photograph with a colour-mixer saturation edit now renders differently:
a raised band is stronger, most of all on muted colours.

The mixer matched bands and judged saturation on scene-linear values, and
scaled chroma by the same factor whatever a colour started at. Against the
photographer's earlier exports of two looks (~90 photographs, their raws, by
encoded hue), a sky band raised by 58 there lifted muted sky blues about
2.1×; here the mixer gave 1.15×, and less in the muted tones that carry most
of a sky or a shadowed snowfield.

Bands are now matched and saturation judged on display-encoded values. A
raised band pushes muted colours hardest and tapers to nothing at full
saturation, at a gain of 3.0, which at the same value lifts muted sky blues
about as those exports did. Lowering saturation still scales every colour
alike. Hue shifts work on the same encoded colour; luminance still scales in
linear light.

The bundled presets that use the mixer, and those whose colour was tuned
against the default rendering, are rescaled to the amount of colour they
had: Vivid 1.30, Vivid warm 1.30, Vivid landscape 1.38, Vivid, strong 1.45,
Vivid portrait 1.15, Punch 1.12, Blue sky 1.08, Deep blue sky 1.12, Polariser
1.26, Blue sky, golden land 1.13 — mean CIELAB chroma over the default
rendering, on 30 raws from the library. Negative values (skin protection)
are left as written.
2026-10-04 05:28:22 -04:00
dtourolle 5ffd54ba43 Leave the profile's look table off by default
Every raw rendered through a camera profile — the library's DNGs with an
embedded profile, and CR2s given one — now renders differently: more
colourful in near-neutral tones. The profile's look table is no longer
applied unless its slider is raised; PROFILE_LOOK names the strength the
profile states.

Against the photographer's earlier exports with no look applied, the default
rendering scores the same with the look table at 100, 50 or 0 (held-out MSE
140, 140, 143), and is 9 % more colourful at 0: the table lowers the
saturation of near-neutral tones, which is exactly where the default
rendering was short of those exports. The user chose more colour.
2026-10-04 05:28:15 -04:00
dtourolle 5a8c3e4c40 Run each model on the Hexagon in the form measured to hold it
The engine knew f32 and int8, and gave the Hexagon int8 for every role it
served. Measured on the tablet itself (inference.md §1.5), int8 lost
5% of the detector's faces at 40-80 px, moved the landmarks 1.5 px,
emptied the segmenter's scores and cost the denoiser 5-9 dB; fp16 the HTP
refuses outright. `Form` gains A16W8 and A16W16, and `Rung::form` now
names one per role: detectors and landmarks A16W8, the segmenter, scene
model, border filler and denoiser A16W16, XFeat int8. The embedder and
the eye classifiers stay on the CPU.

Each loader resolves its `<stem>.<form>.onnx` sibling; the segmenter and
XFeat, compiled into the binary, embed their quantised forms on Android
only and pick through `choose_embedded`. The probe, the compile step and
the cache fingerprint follow the form instead of assuming int8. Detectors
on the new form write `scrfd_*_a16+w600k_mbf`, and `model_ids` answers
for all three spellings.

On the tablet (ORT 1.29 + QNN 2.42), each shipped file against f32 on the
same inputs, and against the CPU's f32 time:
  SCRFD 500m/2.5g/10g  A16W8   100% of faces in every band   4.2/5.1/9.0 ms vs 17/56/198
  landmarks            A16W8   0.25 px in the 192 crop        0.5 ms vs 2.8
  YOLO26n-seg          A16W16  98.2% found, mask IoU 0.994    12.9 ms vs 90
  scene model          A16W16  98.9% of cells agree           15 ms vs 151
  MI-GAN               A16W16  41 dB from f32 in the fill     87 ms vs 488
  XFeat                int8    pano alignment 0.45 px (f32's own spread 0.41)  6.5 ms vs 58
  denoiser             A16W16  0.00 dB at every ISO            95 ms vs 1510 a tile
Face numbers are over public COCO val2017 photographs, not a library.

The APK carries the siblings (BUNDLED 15 -> 19; the old int8 detectors
removed), about 43 MB more. The Windows installer and its CI count skip
them; the Arch and Flatpak packages list their files and never had them.
The ladder example takes a role per model, which is how the per-role
forms above were seen landing on the NPU from the real probe.
2026-10-04 03:45:46 -04:00
dtourolle 0e6ac09fd5 Quantise for the Hexagon with QNN's config and the app's own inputs
tools/quantise-models.sh now writes each model's Hexagon form from a
per-model table: the form its role takes on the NPU (int8, A16W8 or
A16W16), the exact graph rewrites it needs, and the nodes that must stay
float. Ranges are min/max over photographs fed exactly as the app feeds
each model -- the detector and segmenter letterboxes with their own pads
and normalisation, landmark crops from the detector's boxes, MI-GAN with a
panorama-like border, XFeat's grey proxy. The old tool used an
antialiased resize, YOLO's pad of 128 and /255 for every model that was
not a face model, none of which is what the app does.

tools/htp_graph.py holds the rewrites, each checked against the input
graph before use: the denoiser's 6-D Bayer pack and XFeat's 224-slice
unfold as SpaceToDepth (QNN stops at rank 5), computed reshape targets
folded, and bilinear Resize as two MatMuls (the HTP refuses
ResizeBilinear at XFeat's sizes). The denoiser takes ranges computed by
darkroom-denoise's gate on a smaller tile of the same network.
2026-10-04 03:45:07 -04:00
dtourolle 948f6c3ed2 Count the denoise model in the Windows installer's smoke test
package.sh stages models/denoise beside face, scene and inpaint, and
the smoke test counted only the other three, so 0.21.0's Windows job
failed with "expected 14 model files, installed 15". The count reads
the same directories package.sh copies, as its comment intends.
2026-10-04 02:50:15 -04:00
dtourolle 8f9e59b9fa Find hot photosites without repairing them, and measure a sensor's aging
The hot-pixel pass could only repair: it returned how many photosites it
changed and threw away which. find_hot_pixels runs the same pass and
returns them as sensor coordinates, leaving the frame alone, so a sensor's
defects can be tracked across frames.

sensor_scan prints each frame's candidates, and with --probe reads a list
of coordinates back out of every frame. Run over 53 6D raws from 2015 to
2026, it found 32 persistent defects, 2 in 2015 and 32 by 2026, and showed
what a defect map has to account for: a frame that does not flag a
photosite proves nothing unless its neighbourhood is dark, and the 6D
hides some of its defects itself above ISO 5000. docs/dev/sensor-health.md
records the findings and the design they argue for.
2026-10-04 02:38:48 -04:00
dtourolle 83f0461ce7 Sync develop presets through the library
Presets were the one piece of the photographer's work that never left
the device: faces, sidecars, albums, collections, keywords and camera
profiles all travel with the sync pass, the preset library did not.

It now goes to <derived>/presets/library.drpl. PresetLibrary::merge
decides each name against the base the last exchange left (kept per
library beside place.json), so presets added on two devices both
survive, a deletion reaches the other device instead of being restored
by it, and an edit outlives a deletion made elsewhere. The upload is
If-Match / If-None-Match on the server's copy, and a 412 reads and
merges again, so two devices exchanging at once cannot save over each
other. A server copy that will not parse (a newer build's) is left
alone, and a local file that will not read stops the exchange rather
than being taken for an empty library.

The develop view's save merges with the file when the sync changed it
since the view read it, and a sync that brought presets reloads and
redraws the list.

Also corrects the register, which still said camera profiles do not
sync.
2026-10-04 00:50:09 -04:00
dtourolle 26e50ae723 Save presets beside the settings, not under a raw HOME
PresetStore::open built its path from XDG_CONFIG_HOME or HOME. Android
sets neither, so the library resolved to /.config/darkroom, which is
read-only, and every preset saved on the tablet failed. Windows sets no
HOME either and got a directory relative to the working directory. The
settings store was moved to dr_sync::account::config_dir for the same
reason in 0.12.1; the presets now follow it. Linux and macOS resolve to
the same file as before.
2026-10-04 00:49:59 -04:00
dtourolle 25dc0d0179 Give the Vivid presets and Punch measured amounts of colour
On the default rendering, Vivid added 22 % more chroma than the rendering
itself, and Punch 5 % — less than the photographer's earlier exports show
with no look applied (14 % over ours) and well below their everyday look
(27 %). Each preset's colour values (vibrance, saturation, the mixer's
saturation bands) are now scaled together, tone values untouched and
negative ones — Vivid portrait's skin protection — left as written, until
the preset measures: Vivid and Vivid warm 1.30, Vivid landscape 1.38, Vivid,
strong 1.45, Vivid portrait 1.15, Punch 1.12. Measured as mean CIELAB chroma
over 30 raws from the library, as a ratio to the default rendering.
2026-10-03 23:04:46 -04:00
dtourolle a36ec98b36 Carry Lightroom's tone sliders across at measured strengths
Contrast2012 and the four recovery sliders were imported one to one. They do
not mean the same thing here: fitted on the library's Lightroom 6 exports and
their raws — each photograph's sliders carried across as slider × factor, one
factor per slider, on about 90 exports with no look applied, on the Camera Raw
default rendering — ours needed contrast at about a tenth (Lightroom's −100
imported as ours flattens a frame to grey), highlights ×1.4, shadows ×1.9 and
blacks ×1.25. Whites fitted below 1 every time without agreeing where; 0.5 is
a hedge, and says so. Vibrance stays one to one: the op itself is now
calibrated to Lightroom's.
2026-10-03 23:04:46 -04:00
dtourolle c343ac79d3 Describe AI Denoise in the manual as it now is
On for every raw, at the top of the Adjust panel, kept once computed,
and eased off with Strength rather than Keep grain. The timing line is
left as it was; the new model's measured figure replaces it when that
branch lands. The animation still shows the Keep grain slider and
wants recording again.
2026-10-03 22:18:15 -04:00
dtourolle 2e7f14dafe Develop every raw through the AI denoise by default, with a strength, cached
The learned demosaic was an option under Detail, off by default. It is
now how a Bayer raw is developed: on by default at full strength on
every device — which hardware runs it is the inference engine's choice
— and first in the Adjust panel, since it decides what every control
below is applied to.

Strength (0-100, default 100) replaces Keep grain: grain = 100 -
strength, the same luminance-only blend, so moving it is one GPU pass
and never a re-run. 0.21.0's sidecars stored grain; it is still read,
as the inverse, and never written.

With it on for every photograph, the result is now kept on disk
(denoise.md §7.1, §12): the network's output as half floats, keyed on
a SHA-256 of the file's bytes and the model, oldest first past a 5 GB
budget, beside the inference engine's cache. A reopened photograph and
an export of one already developed read it back instead of running the
network again; a damaged entry is a miss.
2026-10-03 22:16:36 -04:00
dtourolle ff0effbfe1 Count the denoise model in the APK's bundled-model list
Benchmarks / CPU and I/O (per commit) (push) Successful in 3m17s
Benchmarks / Frame budget (on demand) (push) Skipped
Traceability / Requirement traces (push) Successful in 56s
Build and test / Android (aarch64) (push) Successful in 47m9s
Build and test / android-image (push) Successful in 3s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / Desktop (Linux) (push) Successful in 1h2m3s
Build and test / windows-image (push) Successful in 2s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / Layer separation (push) Successful in 30s
Build and test / Windows (x86_64, cross) (push) Failing after 49m14s
Build and test / Publish the release (push) Skipped
96001480 added mosaic-1408.onnx to BUNDLED as a fifteenth entry and
left the array's declared length at 14, so the Android build failed
and 0.21.0 got no release. The workspace gates never compile the
Android crate, which is why nothing before CI saw it.
2026-10-03 22:14:26 -04:00
dtourolle 1a03cb52b4 Release 0.21.0
Benchmarks / CPU and I/O (per commit) (push) Successful in 8m38s
Benchmarks / Frame budget (on demand) (push) Skipped
Traceability / Requirement traces (push) Successful in 1m27s
Build and test / Android (aarch64) (push) Failing after 29m52s
Build and test / android-image (push) Successful in 2s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / Desktop (Linux) (push) Successful in 1h2m28s
Build and test / windows-image (push) Successful in 2s
🐳 Windows image / Build and push (push) Successful in 2s
Build and test / Layer separation (push) Successful in 30s
Build and test / Windows (x86_64, cross) (push) Failing after 48m48s
Build and test / Publish the release (push) Skipped
2026-10-03 17:06:49 -04:00
dtourolle ff4b30fbaa Link the inference engine for macOS in a zig container
`docker/macos` builds for aarch64-apple-darwin from Linux with
cargo-zigbuild. Zig carries libSystem and the C headers, so tract's SIMD
kernels compile and the engine's test binaries and examples link as Mach-O
arm64 — the check `cargo check --target` could not do, because tract's
build script needs a macOS C compiler. Crates that link an Apple framework
(dr-plat's keyring, and so the app) still need the Xcode SDK and fail at
the link; macos.md says so.
2026-10-03 16:50:37 -04:00
dtourolle c73743394f Add a CoreML rung on macOS
The macOS ladder was the CPU provider alone, with CoreML listed as a gap.
It is now CoreML, then the CPU, then tract — unmeasured, since nobody here
has a Mac, and safe to ship unmeasured because the probe's clock rejects a
CoreML slower than the CPU and `attempt` refuses one that crashes.

- `Rung::CoreMl`, a compiling rung like TensorRT: an ML Program with every
  compute unit allowed, falling back to the CPU until each model's program
  is built. The embedder stays on the CPU, as on the Hexagon (§7).
- The cache is one directory per model and runtime version. CoreML keys a
  model committed from memory on its input and node names, not its
  weights (ONNX Runtime 1.29, coreml_execution_provider.cc), so two
  exports of one architecture would otherwise share a program.
- The fingerprint on macOS is the chip and the OS release, which ships
  CoreML.
- The desktop looks for the runtime in the bundle's Contents/Frameworks
  and Homebrew's prefixes; fetch-desktop-runtime.sh on a Mac downloads
  ONNX Runtime 1.29.0 for Apple silicon, which carries CoreML.

docs/dev/macos.md says what exists, how to build it, and which log lines
to ask a Mac user for.
2026-10-03 16:50:37 -04:00
dtourolle 872e35670c Log like a debug build on macOS, where a Mac user can find it
Nobody working on DarkRoom has a Mac, so every macOS build is in the hands
of someone who can send a log and cannot attach a debugger. Three changes
make that log worth sending:

- The desktop's default filter on macOS is `debug` for every `dr_*` crate,
  the desktop crate and `onnxruntime` (the runtime's own session log).
- The state directory — the log and crash records — is `~/Library/Logs`
  on macOS rather than the `~/.local/state` Finder hides; Console.app
  lists it. Config and data keep the Unix rules.
- A `diagnostic` cargo profile: release plus line tables, so a crash
  record's backtrace reads file:line. On macOS the tables are in the
  `.dSYM` beside the executable, which the bundle must keep.
2026-10-03 16:49:58 -04:00
dtourolle b562d7b1af Stop retrying a provider that took the app down
The probe runs in the app's process, and a provider can fail by aborting
rather than by returning an error — XNNPACK did on SCRFD. A rung that does
that once would do it on every launch, before the first photograph is on
screen.

Every session build above the CPU, the probe's and each background
compile's, now writes what it is attempting to `attempt` in the cache
directory first and removes it after. After two launches in a row that
died inside the same attempt it is refused and recorded — a rung in
`failed`, an engine in the new `refused` — until the fingerprint changes.
Two, not one, because quitting during a TensorRT compile leaves the same
file.
2026-10-03 16:49:58 -04:00
dtourolle 3689b06c35 Send ONNX Runtime's session log to the app's log
A native session's messages went to ONNX Runtime's stdio logger, which is
nowhere once the app is launched from a menu — and what a provider says
while partitioning a graph (nodes taken, operators declined, a library that
failed to load) is most of what a failed rung tells you. Each session now
forwards them to `log` under the target `onnxruntime`: warnings always,
the runtime's info lines at `debug`, its verbose lines at `trace`.
2026-10-03 16:49:58 -04:00