Commit Graph
17 Commits
Author SHA1 Message Date
dtourolle c73743394f Add a CoreML rung on macOS
The macOS ladder was the CPU provider alone, with CoreML listed as a gap.
It is now CoreML, then the CPU, then tract — unmeasured, since nobody here
has a Mac, and safe to ship unmeasured because the probe's clock rejects a
CoreML slower than the CPU and `attempt` refuses one that crashes.

- `Rung::CoreMl`, a compiling rung like TensorRT: an ML Program with every
  compute unit allowed, falling back to the CPU until each model's program
  is built. The embedder stays on the CPU, as on the Hexagon (§7).
- The cache is one directory per model and runtime version. CoreML keys a
  model committed from memory on its input and node names, not its
  weights (ONNX Runtime 1.29, coreml_execution_provider.cc), so two
  exports of one architecture would otherwise share a program.
- The fingerprint on macOS is the chip and the OS release, which ships
  CoreML.
- The desktop looks for the runtime in the bundle's Contents/Frameworks
  and Homebrew's prefixes; fetch-desktop-runtime.sh on a Mac downloads
  ONNX Runtime 1.29.0 for Apple silicon, which carries CoreML.

docs/dev/macos.md says what exists, how to build it, and which log lines
to ask a Mac user for.
2026-10-03 16:50:37 -04:00
dtourolle b562d7b1af Stop retrying a provider that took the app down
The probe runs in the app's process, and a provider can fail by aborting
rather than by returning an error — XNNPACK did on SCRFD. A rung that does
that once would do it on every launch, before the first photograph is on
screen.

Every session build above the CPU, the probe's and each background
compile's, now writes what it is attempting to `attempt` in the cache
directory first and removes it after. After two launches in a row that
died inside the same attempt it is refused and recorded — a rung in
`failed`, an engine in the new `refused` — until the fingerprint changes.
Two, not one, because quitting during a TensorRT compile leaves the same
file.
2026-10-03 16:49:58 -04:00
dtourolle 3689b06c35 Send ONNX Runtime's session log to the app's log
A native session's messages went to ONNX Runtime's stdio logger, which is
nowhere once the app is launched from a menu — and what a provider says
while partitioning a graph (nodes taken, operators declined, a library that
failed to load) is most of what a failed rung tells you. Each session now
forwards them to `log` under the target `onnxruntime`: warnings always,
the runtime's info lines at `debug`, its verbose lines at `trace`.
2026-10-03 16:49:58 -04:00
dtourolle 20b7bd7663 Feed every input a model declares when probing a rung
The probe built one zero tensor from the first input and ran the session
with it. Every model so far had one input; the denoiser has two (mosaic and
sigma), so every rung failed with "Missing Input: sigma" and the role was
left on the CPU: 14.4 s for a 20 MP frame where TensorRT fp16 takes 3.1 s.
Zeros now go to each input by name.
2026-10-03 11:15:49 -04:00
dtourolle eb91fa02c2 Give the inference engine a denoiser role, kept off the Hexagon
The learned demosaic-and-denoise (denoise.md) runs through the engine like
every other model. fp16 cost it nothing measurable (0.00 dB at every ISO on
validation tiles), so it takes TensorRT's and MIGraphX's fp16 like the
detectors. int8 cost it 6 to 9 dB, far past a 0.5 dB gate, so the Hexagon
refuses the role outright rather than relying on no int8 sibling existing,
and the tablet runs it on the CPU.
2026-10-03 10:39:14 -04:00
dtourolle 84fade99ec Put the developer docs under docs/dev and index the folder for users first
docs/ had 26 developer documents flat beside the manual, and the two
audiences are very differently sized: most readers want the manual and
the gesture reference, a few want the register, the designs and the
measurements. The manual and gestures.md stay at the top; everything for
someone changing the code moves to docs/dev/, and the two documents that
name their own successors — the v0.1 milestone and the UI-refinement plan
— go to docs/dev/archive/ rather than being deleted, since both are still
cited. docs/README.md is the index, users first.

Every reference follows: code comments, Cargo manifests, the workflows,
the pre-commit hook, the bench and traceability tools (which locate the
repo root by docs/dev/requirements.md now), packaging, the Docker READMEs,
CLAUDE.md, CONTRIBUTING.md and the README. The matrix links one level
deeper and is regenerated. Links out of the moved documents into the tree
gain a level; a link checker over every Markdown file finds none broken.
2026-09-20 21:16:03 +02:00
dtourolleandClaude Opus 5 39a22875b1 Add the MIGraphX rung for AMD GPUs
Benchmarks / CPU and I/O (per commit) (push) Failing after 6m20s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 45s
Build and test / Layer separation (push) Successful in 26s
Traceability / Requirement traces (push) Failing after 46s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Failing after 2m19s
Build and test / Windows (x86_64, cross) (push) Failing after 3m2s
Measured on a Radeon RX 7900 XT against Arch's onnxruntime-rocm 1.29
(docs/inference.md §1.3): MIGraphX fp16 runs the detectors at 2.4–3.4 ms
against 10–58 ms on the CPU provider, the inpainter at 8 ms against 514,
with a 15–135 s compile per graph the first time and under a second from
its cache after. A compiling rung on TensorRT's terms, wired the same way.

The ROCm execution provider is gone (removed in ONNX Runtime 1.23), so the
AMD ladder is MIGraphX then the CPU, with no non-compiling rung between.

MIGraphX is registered through the runtime's generic key/value entry
point rather than ort's builder: 1.29 reads the legacy options struct for
its precision flags only, and the compiled-program cache directory
(`migraphx_model_cache_dir`) only travels the generic way. The provider's
cache key omits the precision, so f32 and fp16 programs get their own
directories. The probe fingerprint now includes the provider libraries
beside the runtime and the ROCm version, since a distribution's CPU and
ROCm builds are the same file at the same path.

`status().failed` reports only the rungs above the selection, so an AMD
desktop's About line says why MIGraphX won rather than that the NVIDIA
providers are not in the build.

Two examples: `ep_probe` times each provider cold and from cache, and
`ladder` drives `init` as the app does to watch the first-run sequence.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 19:23:00 +02:00
dtourolle ecb648818b Search the user's own runtime directory before the system library
The reference desktop's only system ONNX Runtime is Arch's
onnxruntime-opt-cuda: 1.29, built without TensorRT and against cuDNN 8
on a cuDNN 9 machine. The probe rejects both providers correctly and
the app runs on the CPU provider, which is right and not what anyone
wants. runtime/ beside the models is now searched ahead of /usr/lib,
tools/fetch-desktop-runtime.sh fills it with the four libraries from
the current onnxruntime-gpu wheel (cuDNN 9, TensorRT 10), and the
About caption lists every rung that lost and why, not only the first.
Verified: the app selects TensorRT from that directory with no
environment variable set.
2026-09-19 21:15:55 +02:00
dtourolle 031ba7b77d Hash a model's bytes once, at open, not on every acquire
A border fill acquires the filler once a tile, and each acquire hashed
the 28 MB model twice — 60 ms a tile, a third of the tile's run on a
throttled TensorRT. The Model keeps its hash from open.
2026-09-19 20:41:21 +02:00
dtourolle c5f07f9ced Give the engine an Inpainter role for the panorama border filler
MI-GAN is plain convolutions, so every rung serves it and none needs a
special form; the role exists so resolve_model and the probe's fingerprint
know the model, and so the merge job can open it through the engine rather
than tract, which takes 7.4 s a tile for it.
2026-09-19 20:41:20 +02:00
dtourolle 46af2a0a46 Stop naming an optimisation level: on tract it means into_optimized, which aborts on yolo26n-seg
ONNX Runtime's default is already its fullest level. ort-tract maps any
level but disabled to tract's optimiser, whose slice pass divides by
zero inside the segmenter's graph — a panic across the C API and so an
abort, which is what stopped dr-ui's develop test. The app never asked
tract for that and does not start now.
2026-09-19 16:35:22 +02:00
dtourolle 95c9cffc0d Keep the embedder off the Hexagon, and let the probe example ask for a runtime
On the tablet the engine compiled arcface for the NPU: the routing
compared the form a rung wants with the form on offer, and for the
embedder both are f32, so nothing said no. A rung now says which roles
it serves at all, and the Hexagon does not serve the embedder (§7 —
its vectors must compare across devices). Tested at the routing seam.

dr-segment's onnx_probe example still named ort-tract, which is what
stopped the workspace test build.
2026-09-19 16:21:53 +02:00
dtourolle cbbe67fbd7 Let the probe's clock be its proof, not disable_cpu_ep_fallback
The strict flag refused the Hexagon over the ten quantise/dequantise
nodes at the graph's edges that QNN declines by policy, which cost
microseconds. A provider that hands real work to the CPU is slower than
the CPU floor and the timing already rejects it; the tablet measured
2.3 ms on the NPU against a 29.7 ms floor.
2026-09-19 16:13:19 +02:00
dtourolle 691af96e3e Keep the readable half of a provider's error for the settings row
ONNX Runtime's errors open with a source path and a template signature;
the first 160 characters of a CUDA failure were all signature. The
reason now starts at the first word a person can act on.
2026-09-19 16:07:19 +02:00
dtourolle 7a436e2549 Move the panorama keypoint detector onto the engine, and probe with a detector
XFeat's two exports are a Keypoints role now; the crate no longer names
tract, and the app compiles TensorRT engines for both ahead of the
first merge. The probe picks the smallest *detector* rather than the
smallest file: the tablet's first run chose the 112 KB eye classifier,
which has no int8 form, and reported the Hexagon as failed for want of
one.
2026-09-19 16:05:08 +02:00
dtourolle 05508741af Start the inference engine from both apps and show its choice in Settings
The desktop names where a package may have put libonnxruntime — an
override variable, beside the executable, the package's own library
directory, the Flatpak prefix, the system library directory — and
Android points at the APK's native library directory, which is also
what Qualcomm's DSP loader must be told for the Hexagon skel. Android
starts the engine at the end of the model unpack rather than at launch,
because the probe fingerprints the model files and a first launch has
none until then.

The About panel gains an Inference row beside Graphics, re-read every
two seconds while the probe runs and engines land, and faces.model_id
carries the detector's form: an int8 detector finds a different set of
faces and is a different population (docs/inference.md §7). A
low-memory signal drops every idle session with the GPU caches.

The APK assembly bundles ONNX Runtime and the Qualcomm HTP libraries
from Maven, fetched by tools/fetch-android-runtime.sh with their
published checksums; RUNTIME_DIR=none builds the tract-only APK, which
is a slower app and not a broken one. The desktop packages carry no
runtime yet.

Two probe fixes from the first desktop run: the floor must not be
built with CPU fallback disabled, and a versioned libonnxruntime.so is
a runtime too. On the reference desktop the probe now loads ONNX
Runtime 1.30, measures 30 ms on the CPU provider, and selects TensorRT
at 1.5 ms.
2026-09-19 16:02:37 +02:00
dtourolle d15c41e699 Add dr-inference-engine and route every model session through it
One crate names the runtime, the providers and the devices; dr-face and
dr-segment ask it for a session by role. It hands ort an API table once
per process — from a libonnxruntime it dlopens when the app names a
directory holding one, otherwise from tract — so the Rust build stays
free of C on every target and a package can install the runtime as a
file (docs/inference.md §3).

Sessions live in a registry behind a Model handle that holds the bytes,
not the session: every use refreshes a timestamp and a reaper unloads
whatever sat idle past the decay. A scan that runs the detector on each
image never lets it go idle; a click in the develop view lets the
segmenter go after thirty seconds; a handle used after that reloads,
and reloads on a higher rung if a compiled engine has landed meanwhile.

The probe walks the platform's ladder by building strict sessions and
timing them against the CPU provider, caches the choice against a
fingerprint of the runtime, driver, hardware and models, and compiles
engines for the selected rung in the background, smallest model first.
Nothing in this commit turns the native path on: the apps still run on
tract until they call init with a runtime directory.
2026-09-19 16:02:37 +02:00