denoise.md §15: why Best became one network, how it compares with the
mixture and Medium on real photographs and the chart, the candidates that
fell short, the file names, the saved-edit numbering, and the timings on
the 3050 -- 0.51-0.54 s whole-frame, 0.95 s in tiles, against 2.60 s for
the mixture in tiles. The manual lists Bilinear, Fast and Best, says an
edit made with Medium opens with Best, and gives the new time; Medium's
close-up goes.
Best was a mixture of two experts and a gate, 110 GMAC a megapixel;
Medium a single network at 48 that was softer on real edges. fb-combo
(darkroom-denoise, 20 000 steps from fb-edges2, taught by the mixture
with a quarter of its crops from the edge-rich parts of the frames) is
Medium's shape and holds the mixture's edges on real photographs:
edge PSNR within 0.04-0.06 dB at ISO 1600/6400/25600, more sharpness
kept at all three, the chart's edge 0.89 photosites wide against 0.82.
It is 0.27 dB short on smooth areas at ISO 25600. It becomes Best, and
the methods are Bilinear, Fast and Best.
Saved edits keep their numbers: 2, which was Medium, is now Best, and
3, which was Best, is past the end and reads as the default, Best.
The network ships as mosaic-hq, a new name: the result cache keys a
model by name and size, and this one is byte for byte the old Medium's
size. Its tablet form (A16W16) lost 0.00 dB in simulated QDQ at every
ISO and at most 0.09 dB across the noise bracket.
denoise.md §14: why the tiles waste half of Best's work, where the any-size
networks run and why only there, why the limit is the card's memory, and
the measurement on _MG_8862 — 2.60 s in 1408 tiles, 1.37 s in two
4160 x 3248 tiles, the outputs within fp16's own spread. models/LICENCE.md
lists the three re-exports.
TensorRT plans its memory for the profile's largest shape, and up to a
whole 6D frame with Best's border (4608 x 6656) it asked for 4.9-5.9 GB
and would not build on the RTX 3050. The profile now ends at 4608 x 3328
(15 MP), tuned for 4160 x 3248, and the tiler cuts a 6D frame into two
such tiles: 27 MP of work for 20 MP kept, against 49 MP in 1408 tiles.
The engine's directory names the profile, so a later range never loads
an engine built for this one.
mosaic-{fast,medium,best}.onnx are the shipped networks re-exported by
darkroom-denoise tools/export_whole.py (22ea648) from the checkpoints
the 1408 files came from: identical to them at 1408 (max |d| = 0), to
torch at 592 x 848, and to tiled inference over the reflected frame in
f64. 42 MB together. The Arch package and the Windows installer carry
them beside the fixed files; the APK leaves them out, since the Hexagon
takes fixed shapes only.
A fixed 1408 tile is exact only in its centre, and Best keeps 896 of
every 1408 it computes: 2.47 photosites of work for each one kept. The
tiler now takes a network of any size as well as a square one, and
plans the frame as the fewest equal tiles under the rung's limit --
one tile, the whole frame and its reflected border, whenever it fits.
If the first call of a plan fails, as a GPU out of memory does, the
kept centre is halved and the frame planned again.
Each shipped network names its any-size sibling (mosaic-best.onnx
beside mosaic-best-1408.onnx). OnnxNet::open takes it where the engine
runs whole frames and the file is installed, and the 1408 tiles
otherwise; open_tiled forces the tiles, and denoise_raw's DR_PLAN=tiles
uses it to compare. The cache key stays on the fixed model: the output
is the same network's. Tests hold any-size tiles, a grid of them and a
plan rebuilt after a failure to the square tiles' answer in every Bayer
phase.
Role::WholeDenoiser is the denoise network exported with any height and
width, for a whole frame instead of 1408 tiles whose borders are thrown
away. It is served only where a new size costs nothing: the CUDA
provider, and TensorRT through an optimisation profile from 256 to
4608 x 6656, tuned for the 6D's frame with Best's border. Everywhere
else whole_frame_limit() says None and the fixed tiles run.
ort's TensorRT builder has no profile options, so the engine registers
through the runtime's V2 options with the names 1.30 reads
(trt_profile_{min,opt,max}_shapes). Without a profile a dynamic input
compiled an engine per size at run time, 156 s on the first frame. The
engine lives in its own directory per model: ORT's cache key leaves the
shape out, and the fixed 1408 export and its any-size sibling are the
same graph.
hardware::detect is read only by api::install_best, which exists with
the native feature; a build of a crate that takes the engine without it
(dr-denoise's own tests) warned that all of it was unused.
The generic runtime is opened only when the QNN build does not fit, and
the fit rests on ro.soc.manufacturer. A property the app cannot read, or
an Android older than 12 that has none, read as "not Qualcomm" would
put a Qualcomm device on the generic rung and off its Hexagon — what
0.22.1 had just fixed. Only a device that names another vendor is now
not Qualcomm's; the tablet reports QTI.
One warm-up run and the median of three: on the Iris Xe the OpenVINO
rung lost to the CPU on the smallest detector in two probes of three,
because an idle integrated GPU takes a few runs to raise its clock —
warm, it is 5.8 ms against 9.5. Three warm-ups and the median of seven
took it in five probes of five (5.6–7.2 ms against 7.7–13.1). The extra
runs cost tens of milliseconds, once per fingerprint.
§1.6 is the Iris Xe measurement: OpenVINO fp16 1.3–5.8× the CPU
provider on every shipped model, WebGPU behind it everywhere but the
denoiser and MI-GAN. §2's ladder gains the Intel and generic rows, the
generic one footnoted as unmeasured where it is meant to help. §3.1
lists the two bundled runtimes' licences; §3.2 is how one runtime of
several is chosen per process. D13 notes the bundling.
The APK's ONNX Runtime is the QNN build, which carries no WebGPU, so a
non-Qualcomm phone had nothing above the CPU provider. The APK now also
carries Microsoft's stock onnxruntime-android 1.29.0 (32 MB) as
libonnxruntime_generic.so, and the app offers it after the QNN build.
The runtime search stops at a perfect fit, so on a Qualcomm device the
QNN build — listed first — is all that is opened, and the generic build
never loads beside it. A Qualcomm SoC is read from ro.soc.manufacturer
or, before Android 12, from Qualcomm's FastRPC library being present:
a Qualcomm device mistaken for another would trade its Hexagon for the
generic rung.
The Windows installer and the Flatpak carried no ONNX Runtime, so they
ran every model on tract's one core; the Arch package left it to an
optional dependency. Each now installs two builds under runtimes/ —
Intel's OpenVINO build and the generic WebGPU one, both with the CPU
provider — fetched by tools/fetch-bundled-runtimes.sh from PyPI wheels
pinned by SHA-256, pruned to the native libraries (81 + 31 MB on Linux,
67 + 42 MB on Windows), licence texts beside them.
darkroom-desktop searches runtimes/openvino and runtimes/webgpu under
each place a package installs to; the engine opens all it finds and
keeps the one that fits the GPU, so a CUDA or ROCm runtime installed
beside them still wins on its vendor's card. On Windows the chosen
runtime's directory goes on PATH, because Intel's build leaves OpenVINO's
DLLs for the loader to find there.
The Windows image gains unzip; the installer smoke test checks both
runtimes landed.
A device left on the CPU gave the first failure as the reason, which on
any runtime but NVIDIA's is "TensorRT execution provider is not enabled
in this build". The reason is now the last rung that was tried and lost,
"WebGPU 150.8 ms, slower than the CPU's 26.6 ms"; the full list is still
in the status's failures.
A runtime carries one vendor's providers, only one loads per process,
and a device can now hold several: the package's OpenVINO or WebGPU
build, a CUDA build the user fetched, the distribution's ROCm build.
`api::install` opens each it finds, lists its providers with
GetAvailableProviders, and installs the one scoring highest against the
GPUs `hardware::detect` reads from files — a vendor rung on its own
vendor's GPU above OpenVINO on an Intel one above the generic WebGPU
rung above a CPU-only build. Equal scores keep the old first-found
order, and DARKROOM_ORT_DIR still wins outright. The losers stay mapped
rather than unloaded.
The Linux fingerprint now names the OpenCL drivers too, so installing
Intel's re-probes. `ladder` takes DARKROOM_ORT_DIRS to show the choice.
OpenVINO is the Intel rung: the integrated or Arc GPU, fp16 for every
role but the embedder, a compiled program per model kept in a directory
per model, precision and runtime version. On the Iris Xe it beats ONNX
Runtime's CPU provider on every shipped model — scrfd_10g 23 ms against
58, the scene model 17 against 57, MI-GAN 57 against 330, a denoise tile
40 against 158.
WebGPU is the generic rung for a GPU no vendor rung covers. It was slower
than the CPU on the Iris Xe, the RTX 3050 and the Adreno, so it is on the
ladder for the GPUs it has not been timed on, behind the probe's clock.
MIGraphX's registration becomes one generic key/value helper that all
three share, with option names read from each runtime's own source.
The Intel and vendor-neutral rungs need a measurement before they join
the ladder (docs/dev/inference.md §2). Both register through the generic
key/value entry point with the option names ONNX Runtime reads at the
wheel's version: OpenVINO 1.24 (`openvino_provider_factory.cc`), WebGPU
1.27 (`webgpu_provider_options.h`, prefixed by the runtime).
DARKROOM_EPS narrows the list to the families a runtime carries.
QNN's Hexagon stub loads libcdsprpc.so, a vendor library, and from API 31
an app's linker namespace refuses a vendor library its manifest does not
name. QNN then fails to create its device - QNN_DEVICE_ERROR_INVALID_CONFIG,
before it reaches the DSP - and every model ran on the CPU on 0.22.0: AI
denoise took 30-131 s a photograph on the tablet.
The same engine code ran on the HTP from adb's shell, whose namespace has
no such rule, which is what hid it. With the declaration the app's probe
chose the Hexagon on the tablet and compiled every model for it. Not
required, so a device without the library still installs and runs on the
CPU. A test holds the line in the manifest.
0.22.0's first launch on the tablet: QNN could not create its device
(QNN_DEVICE_ERROR_INVALID_CONFIG), the session built anyway with every
node on the CPU behind the provider, and the probe timed that - 28.5 ms
against the CPU's own 19.4 - and rejected the Hexagon. The verdict was
cached under the fingerprint, so every model stayed on the CPU on every
later launch: AI denoise took 30-131 s a photograph instead of seconds.
The same A16W8 detector with the APK's own libraries runs on the HTP in
4.4 ms.
The probe's Hexagon session now sets session.disable_cpu_ep_fallback, so
a device that cannot take the graph fails the probe instead of being timed
as the CPU. Only the probe: shipped graphs may keep nodes on the CPU on
purpose. And a selection that fell back to the CPU after an accelerator
failed or lost is probed again on the next launches, up to three probes
per fingerprint; a cache written by 0.22.0 reads as never retried, so the
tablet probes again once this is installed.
Slint expands the .slint files into ~27 MB of Rust, and at the workspace's
single codegen unit LLVM optimised all of it on one thread: 13.5 minutes
of a release build with the other cores idle. The override applies to
dr-ui alone; the image crates keep one unit, and thin LTO still runs at
link time.
The manual's AI denoise section names the four methods and their measured
times, and shows the lamp and railing of the ISO 8000 frame at 1:1 by each
in place of the film and the before/after pair. The scene clicks each
method and waits for that network's result: the repair now logs its own
"learned denoise:" line first, so the wait matches the result's.
AI Denoise's Apply switch becomes Method: Bilinear, Fast, Medium, Best,
default Best, so an untouched raw writes nothing and develops through the
mixture. `apply` is still read and never written: 0 is Bilinear, 1 keeps
a network already chosen.
- Best is the mixture of a flat and an edge expert with a learned gate;
Medium and Fast are students distilled from it. 2.48 s, 0.79 s and
0.57 s for a 20 MP frame on TensorRT fp16.
- Each network carries its own tile border (256 for the mixture, 192 for
the students) through `dr_denoise::Shipped` and `TileNet::halo`.
- The file is hashed once at open and each network keys its own cached
result; Bilinear keeps the result in memory for the way back.
- Each has an .a16w16 sibling for the Hexagon: 0.00 dB on the 6D gate,
at most 0.11 dB with the noise scaled x0.5 to x4.
- APK BUNDLED 19 -> 23; the PKGBUILD installs all three.
A 20 MP frame spent 0.32 s outside the network: each tile's mosaic and
sigma gathered on one thread, then its 24 MB output copied out of the
runtime and back into the frame, all in series with the device. Tiles are
now gathered on every core by a producer thread one tile ahead, so the
gather overlaps the run; the centre is written back across cores; and the
tile interface hands its inputs over and lends its output, so neither
side is copied. With a stand-in network that does nothing, the tiler's own
time falls to 0.14 s at the 1408 tile and 0.09 s at 2048. The exactness
and Bayer-phase tests are unchanged and pass.
The app's hot-pixel pass takes gross defects only; at ISO 6400-25600 a 6D
frame keeps 1000-2000 photosites more than 8 sigma beyond all their
same-colour and adjacent neighbours, which the network turned into specks.
The same two tests with the threshold in the photosite's own sigma, plus
the factor of two that keeps a bright point of light (where 8 sigma is a
sliver of the signal). The next model is trained behind exactly this; on
an ISO 25600 frame the Rust and training code both repair 935.
Photographs opened with an earlier Lightroom edit now import its HSL
saturation as fitted against the library's own Lightroom 6 exports, rather
than one band to one band.
Measured on two looks' exports and their raws (darkroom-lrfit, hsl_map_fit),
by encoded hue: Lightroom's saturation bands act about 45 degrees either
side on our wheel, wider than ours, and not all at our strength. Each is now
shared between two or three of our bands — Aqua mostly cyan and azure, where
skies are; Blue mostly blue and violet; Orange, where skin is, at about 0.4
of its value. Values add when two of Lightroom's bands share one of ours.
On the measured skies the import now lifts muted sky blues about 1.9× against
Lightroom's 2.1×, where it gave 1.15×. Hue and luminance still go one band to
the band of the same hue; they were not measured.
Every photograph with a colour-mixer saturation edit now renders differently:
a raised band is stronger, most of all on muted colours.
The mixer matched bands and judged saturation on scene-linear values, and
scaled chroma by the same factor whatever a colour started at. Against the
photographer's earlier exports of two looks (~90 photographs, their raws, by
encoded hue), a sky band raised by 58 there lifted muted sky blues about
2.1×; here the mixer gave 1.15×, and less in the muted tones that carry most
of a sky or a shadowed snowfield.
Bands are now matched and saturation judged on display-encoded values. A
raised band pushes muted colours hardest and tapers to nothing at full
saturation, at a gain of 3.0, which at the same value lifts muted sky blues
about as those exports did. Lowering saturation still scales every colour
alike. Hue shifts work on the same encoded colour; luminance still scales in
linear light.
The bundled presets that use the mixer, and those whose colour was tuned
against the default rendering, are rescaled to the amount of colour they
had: Vivid 1.30, Vivid warm 1.30, Vivid landscape 1.38, Vivid, strong 1.45,
Vivid portrait 1.15, Punch 1.12, Blue sky 1.08, Deep blue sky 1.12, Polariser
1.26, Blue sky, golden land 1.13 — mean CIELAB chroma over the default
rendering, on 30 raws from the library. Negative values (skin protection)
are left as written.
Every raw rendered through a camera profile — the library's DNGs with an
embedded profile, and CR2s given one — now renders differently: more
colourful in near-neutral tones. The profile's look table is no longer
applied unless its slider is raised; PROFILE_LOOK names the strength the
profile states.
Against the photographer's earlier exports with no look applied, the default
rendering scores the same with the look table at 100, 50 or 0 (held-out MSE
140, 140, 143), and is 9 % more colourful at 0: the table lowers the
saturation of near-neutral tones, which is exactly where the default
rendering was short of those exports. The user chose more colour.
The engine knew f32 and int8, and gave the Hexagon int8 for every role it
served. Measured on the tablet itself (inference.md §1.5), int8 lost
5% of the detector's faces at 40-80 px, moved the landmarks 1.5 px,
emptied the segmenter's scores and cost the denoiser 5-9 dB; fp16 the HTP
refuses outright. `Form` gains A16W8 and A16W16, and `Rung::form` now
names one per role: detectors and landmarks A16W8, the segmenter, scene
model, border filler and denoiser A16W16, XFeat int8. The embedder and
the eye classifiers stay on the CPU.
Each loader resolves its `<stem>.<form>.onnx` sibling; the segmenter and
XFeat, compiled into the binary, embed their quantised forms on Android
only and pick through `choose_embedded`. The probe, the compile step and
the cache fingerprint follow the form instead of assuming int8. Detectors
on the new form write `scrfd_*_a16+w600k_mbf`, and `model_ids` answers
for all three spellings.
On the tablet (ORT 1.29 + QNN 2.42), each shipped file against f32 on the
same inputs, and against the CPU's f32 time:
SCRFD 500m/2.5g/10g A16W8 100% of faces in every band 4.2/5.1/9.0 ms vs 17/56/198
landmarks A16W8 0.25 px in the 192 crop 0.5 ms vs 2.8
YOLO26n-seg A16W16 98.2% found, mask IoU 0.994 12.9 ms vs 90
scene model A16W16 98.9% of cells agree 15 ms vs 151
MI-GAN A16W16 41 dB from f32 in the fill 87 ms vs 488
XFeat int8 pano alignment 0.45 px (f32's own spread 0.41) 6.5 ms vs 58
denoiser A16W16 0.00 dB at every ISO 95 ms vs 1510 a tile
Face numbers are over public COCO val2017 photographs, not a library.
The APK carries the siblings (BUNDLED 15 -> 19; the old int8 detectors
removed), about 43 MB more. The Windows installer and its CI count skip
them; the Arch and Flatpak packages list their files and never had them.
The ladder example takes a role per model, which is how the per-role
forms above were seen landing on the NPU from the real probe.
tools/quantise-models.sh now writes each model's Hexagon form from a
per-model table: the form its role takes on the NPU (int8, A16W8 or
A16W16), the exact graph rewrites it needs, and the nodes that must stay
float. Ranges are min/max over photographs fed exactly as the app feeds
each model -- the detector and segmenter letterboxes with their own pads
and normalisation, landmark crops from the detector's boxes, MI-GAN with a
panorama-like border, XFeat's grey proxy. The old tool used an
antialiased resize, YOLO's pad of 128 and /255 for every model that was
not a face model, none of which is what the app does.
tools/htp_graph.py holds the rewrites, each checked against the input
graph before use: the denoiser's 6-D Bayer pack and XFeat's 224-slice
unfold as SpaceToDepth (QNN stops at rank 5), computed reshape targets
folded, and bilinear Resize as two MatMuls (the HTP refuses
ResizeBilinear at XFeat's sizes). The denoiser takes ranges computed by
darkroom-denoise's gate on a smaller tile of the same network.
package.sh stages models/denoise beside face, scene and inpaint, and
the smoke test counted only the other three, so 0.21.0's Windows job
failed with "expected 14 model files, installed 15". The count reads
the same directories package.sh copies, as its comment intends.
The hot-pixel pass could only repair: it returned how many photosites it
changed and threw away which. find_hot_pixels runs the same pass and
returns them as sensor coordinates, leaving the frame alone, so a sensor's
defects can be tracked across frames.
sensor_scan prints each frame's candidates, and with --probe reads a list
of coordinates back out of every frame. Run over 53 6D raws from 2015 to
2026, it found 32 persistent defects, 2 in 2015 and 32 by 2026, and showed
what a defect map has to account for: a frame that does not flag a
photosite proves nothing unless its neighbourhood is dark, and the 6D
hides some of its defects itself above ISO 5000. docs/dev/sensor-health.md
records the findings and the design they argue for.
Presets were the one piece of the photographer's work that never left
the device: faces, sidecars, albums, collections, keywords and camera
profiles all travel with the sync pass, the preset library did not.
It now goes to <derived>/presets/library.drpl. PresetLibrary::merge
decides each name against the base the last exchange left (kept per
library beside place.json), so presets added on two devices both
survive, a deletion reaches the other device instead of being restored
by it, and an edit outlives a deletion made elsewhere. The upload is
If-Match / If-None-Match on the server's copy, and a 412 reads and
merges again, so two devices exchanging at once cannot save over each
other. A server copy that will not parse (a newer build's) is left
alone, and a local file that will not read stops the exchange rather
than being taken for an empty library.
The develop view's save merges with the file when the sync changed it
since the view read it, and a sync that brought presets reloads and
redraws the list.
Also corrects the register, which still said camera profiles do not
sync.
PresetStore::open built its path from XDG_CONFIG_HOME or HOME. Android
sets neither, so the library resolved to /.config/darkroom, which is
read-only, and every preset saved on the tablet failed. Windows sets no
HOME either and got a directory relative to the working directory. The
settings store was moved to dr_sync::account::config_dir for the same
reason in 0.12.1; the presets now follow it. Linux and macOS resolve to
the same file as before.
On the default rendering, Vivid added 22 % more chroma than the rendering
itself, and Punch 5 % — less than the photographer's earlier exports show
with no look applied (14 % over ours) and well below their everyday look
(27 %). Each preset's colour values (vibrance, saturation, the mixer's
saturation bands) are now scaled together, tone values untouched and
negative ones — Vivid portrait's skin protection — left as written, until
the preset measures: Vivid and Vivid warm 1.30, Vivid landscape 1.38, Vivid,
strong 1.45, Vivid portrait 1.15, Punch 1.12. Measured as mean CIELAB chroma
over 30 raws from the library, as a ratio to the default rendering.
Contrast2012 and the four recovery sliders were imported one to one. They do
not mean the same thing here: fitted on the library's Lightroom 6 exports and
their raws — each photograph's sliders carried across as slider × factor, one
factor per slider, on about 90 exports with no look applied, on the Camera Raw
default rendering — ours needed contrast at about a tenth (Lightroom's −100
imported as ours flattens a frame to grey), highlights ×1.4, shadows ×1.9 and
blacks ×1.25. Whites fitted below 1 every time without agreeing where; 0.5 is
a hedge, and says so. Vibrance stays one to one: the op itself is now
calibrated to Lightroom's.
On for every raw, at the top of the Adjust panel, kept once computed,
and eased off with Strength rather than Keep grain. The timing line is
left as it was; the new model's measured figure replaces it when that
branch lands. The animation still shows the Keep grain slider and
wants recording again.
The learned demosaic was an option under Detail, off by default. It is
now how a Bayer raw is developed: on by default at full strength on
every device — which hardware runs it is the inference engine's choice
— and first in the Adjust panel, since it decides what every control
below is applied to.
Strength (0-100, default 100) replaces Keep grain: grain = 100 -
strength, the same luminance-only blend, so moving it is one GPU pass
and never a re-run. 0.21.0's sidecars stored grain; it is still read,
as the inverse, and never written.
With it on for every photograph, the result is now kept on disk
(denoise.md §7.1, §12): the network's output as half floats, keyed on
a SHA-256 of the file's bytes and the model, oldest first past a 5 GB
budget, beside the inference engine's cache. A reopened photograph and
an export of one already developed read it back instead of running the
network again; a damaged entry is a miss.
96001480 added mosaic-1408.onnx to BUNDLED as a fifteenth entry and
left the array's declared length at 14, so the Android build failed
and 0.21.0 got no release. The workspace gates never compile the
Android crate, which is why nothing before CI saw it.
`docker/macos` builds for aarch64-apple-darwin from Linux with
cargo-zigbuild. Zig carries libSystem and the C headers, so tract's SIMD
kernels compile and the engine's test binaries and examples link as Mach-O
arm64 — the check `cargo check --target` could not do, because tract's
build script needs a macOS C compiler. Crates that link an Apple framework
(dr-plat's keyring, and so the app) still need the Xcode SDK and fail at
the link; macos.md says so.
The macOS ladder was the CPU provider alone, with CoreML listed as a gap.
It is now CoreML, then the CPU, then tract — unmeasured, since nobody here
has a Mac, and safe to ship unmeasured because the probe's clock rejects a
CoreML slower than the CPU and `attempt` refuses one that crashes.
- `Rung::CoreMl`, a compiling rung like TensorRT: an ML Program with every
compute unit allowed, falling back to the CPU until each model's program
is built. The embedder stays on the CPU, as on the Hexagon (§7).
- The cache is one directory per model and runtime version. CoreML keys a
model committed from memory on its input and node names, not its
weights (ONNX Runtime 1.29, coreml_execution_provider.cc), so two
exports of one architecture would otherwise share a program.
- The fingerprint on macOS is the chip and the OS release, which ships
CoreML.
- The desktop looks for the runtime in the bundle's Contents/Frameworks
and Homebrew's prefixes; fetch-desktop-runtime.sh on a Mac downloads
ONNX Runtime 1.29.0 for Apple silicon, which carries CoreML.
docs/dev/macos.md says what exists, how to build it, and which log lines
to ask a Mac user for.
Nobody working on DarkRoom has a Mac, so every macOS build is in the hands
of someone who can send a log and cannot attach a debugger. Three changes
make that log worth sending:
- The desktop's default filter on macOS is `debug` for every `dr_*` crate,
the desktop crate and `onnxruntime` (the runtime's own session log).
- The state directory — the log and crash records — is `~/Library/Logs`
on macOS rather than the `~/.local/state` Finder hides; Console.app
lists it. Config and data keep the Unix rules.
- A `diagnostic` cargo profile: release plus line tables, so a crash
record's backtrace reads file:line. On macOS the tables are in the
`.dSYM` beside the executable, which the bundle must keep.
The probe runs in the app's process, and a provider can fail by aborting
rather than by returning an error — XNNPACK did on SCRFD. A rung that does
that once would do it on every launch, before the first photograph is on
screen.
Every session build above the CPU, the probe's and each background
compile's, now writes what it is attempting to `attempt` in the cache
directory first and removes it after. After two launches in a row that
died inside the same attempt it is refused and recorded — a rung in
`failed`, an engine in the new `refused` — until the fingerprint changes.
Two, not one, because quitting during a TensorRT compile leaves the same
file.
A native session's messages went to ONNX Runtime's stdio logger, which is
nowhere once the app is launched from a menu — and what a provider says
while partitioning a graph (nodes taken, operators declined, a library that
failed to load) is most of what a failed rung tells you. Each session now
forwards them to `log` under the target `onnxruntime`: warnings always,
the runtime's info lines at `debug`, its verbose lines at `trace`.