Record the Intel and generic rungs in the inference spec

§1.6 is the Iris Xe measurement: OpenVINO fp16 1.3–5.8× the CPU
provider on every shipped model, WebGPU behind it everywhere but the
denoiser and MI-GAN. §2's ladder gains the Intel and generic rows, the
generic one footnoted as unmeasured where it is meant to help. §3.1
lists the two bundled runtimes' licences; §3.2 is how one runtime of
several is chosen per process. D13 notes the bundling.
This commit is contained in:
2026-10-04 21:00:15 -04:00
parent 43402bfe5b
commit b3999bbc0c
+88 -6
View File
@@ -194,6 +194,45 @@ int8's matches land 0.45 px from f32's in the overlaps — the spread f32 shows
---
### 1.6 Intel Iris Xe, and the generic rung · 2026-10-04
The RTX 3050 laptop's other GPU: Raptor Lake-P's Iris Xe (96 EU), Intel's `onnxruntime-openvino`
1.24.1 (OpenVINO 2025.4.1) and Microsoft's `onnxruntime-webgpu` 1.27.0, both PyPI wheels, through
`ep_probe`. Three warm-ups, the median of 15 runs. Another build shared the CPU during the run, so
the CPU columns are a little pessimistic; the GPU columns are not.
| Model | ORT CPU f32 | OpenVINO CPU | OpenVINO GPU f32 | **OpenVINO GPU fp16** | WebGPU (Iris Xe) |
|---|---|---|---|---|---|
| scrfd_500m (Fast) | 9.5 | 10.8 | 7.3 | **5.8** | 24.2 |
| scrfd_2.5g (Balanced) | 18.5 | 16.2 | 15.5 | **11.1** | 40.8 |
| scrfd_10g (Thorough) | 58.5 | 72.4 | 38.7 | **23.0** | 84.5 |
| arcface_mbf (per face) | 9.5 | 11.7 | **3.0** | 2.4 | 56.6 |
| 2d106det (landmarks) | 10.8 | 2.0 | 2.4 | **1.9** | 42.7 |
| yolo26s-sem-ade20k | 57.0 | 47.4 | 26.0 | **16.7** | 53.0 |
| xfeat-1024 | 23.5 | 18.5 | 19.9 | **17.5** | 34.2 |
| migan-512 (per tile) | 330 | ✗ ¹ | 89.7 | **57.2** | 275 |
| mosaic-fast-1408 (per tile) | 159 | 107 | 68.1 | **40.0** | 188 ² |
| mosaic-best-1408 (per tile) | 1109 | 1670 | 947 | **604** | 1034 ² |
¹ OpenVINO's CPU plugin refuses the graph at initialisation. Not shipped (§3.2), so moot.
² A later run, after `ep_probe` learned to feed the denoiser's two inputs, under heavier load: the
CPU provider took 256 and 1034 ms in that run, so WebGPU beat it by a quarter on mosaic-fast and tied
on mosaic-best — the only rows where it is not well behind.
- **OpenVINO on the Iris Xe beats ONNX Runtime's CPU provider on every model**, 1.3× on XFeat to
5.8× on MI-GAN, with a 1–3 s compile per graph and 0.1–0.4 s from its cache. It is the Intel rung.
fp16 is worth 1.3–1.7× over f32 here, against 1.1–1.35× on MIGraphX.
- **Its "GPU" is OpenCL's first GPU, not Intel's.** Before `intel-compute-runtime` was installed
the only OpenCL driver was NVIDIA's, and `device_type=GPU` ran on the RTX 3050 — slower than the
CPU, which is the probe's to catch. Read the process's maps for `libigdrcl` before believing a
number is the iGPU's.
- **WebGPU is slower than the CPU on the Iris Xe** on everything but MI-GAN, as it was on the
Adreno (§1.1), and on the RTX 3050 through Vulkan too. It is on the ladder anyway, as the generic
rung (§2): for GPUs no vendor rung covers — an AMD card on Windows or without ROCm, a Mali — where
it is unmeasured, and the probe's clock decides.
- **OpenVINO's CPU plugin is not a better floor.** It wins on some graphs and loses on scrfd_10g,
the embedder and mosaic-best, and refuses MI-GAN.
## 2. The shape of the answer
A **ladder per platform**, walked at start-up, with the first rung that builds a real session
@@ -202,10 +241,11 @@ winning:
| Platform | 1st | 2nd | 3rd | Floor |
|---|---|---|---|---|
| Android, Qualcomm with a Hexagon the shipped QNN skel covers (V68–V81) | QNN HTP, each model's quantised form (§1.5) | ORT CPU, f32 model | — | tract |
| Android, any other SoC | ORT CPU, f32 | — | — | tract |
| Android, any other SoC ⁶ | WebGPU (Vulkan), f32 | ORT CPU, f32 | — | tract |
| Linux / Windows, NVIDIA GPU | TensorRT, f32 model, fp16 engine | CUDA provider, f32 | ORT CPU, f32 | tract |
| Linux, AMD GPU with ROCm | MIGraphX, f32 model, fp16 program | ORT CPU, f32 | — | tract |
| Linux / Windows, no GPU stack | ORT CPU, f32 | — | — | tract |
| Linux / Windows, Intel GPU | OpenVINO, f32 model, fp16 program (§1.6) | ORT CPU, f32 | — | tract |
| Linux / Windows, any other GPU ⁶ | WebGPU (Vulkan / D3D12), f32 | ORT CPU, f32 | — | tract |
| macOS ⁵ | CoreML, f32 model, ML Program | ORT CPU, f32 | — | tract |
⁵ **Unmeasured**, and the one exception to the rule below: nobody here has a Mac. The rung is on
@@ -215,11 +255,17 @@ down is refused on the third launch (§4, `attempt`). The embedder stays on the
macOS log that shows a probe line is this row's measurement; [macos.md](macos.md) says what to
ask for.
⁶ **The generic rung, unmeasured where it is meant to help.** WebGPU lost to the CPU on every GPU
it has been timed on — the Adreno, the Iris Xe, the RTX 3050 (§1.1, §1.6) — none of which it serves
here, since each has its own rung. It is on the ladder for the GPUs that have none, on the same
terms as CoreML: a WebGPU that is slower than the CPU is rejected by §4's clock, one that errors is
recorded as failed. Its first measurement on an AMD card without ROCm, or a Mali, is this row's.
Deliberately **not** on any ladder, with the measurement that excluded each: NNAPI (no driver),
XNNPACK (slower than CPU, aborts on SCRFD), WebGPU (slower than CPU), the Adreno through QNN (works,
but never where the Hexagon does not also), CUDA int8 (slower than CUDA f32), the ROCm provider
(gone: §1.3). A rung is added to this table by a measurement on this page, not by a provider
existing.
XNNPACK (slower than CPU, aborts on SCRFD), the Adreno through QNN (works, but never where the
Hexagon does not also), CUDA int8 (slower than CUDA f32), the ROCm provider (gone: §1.3),
OpenVINO's CPU plugin as a floor (§1.6). A rung is added to this table by a measurement on this
page, not by a provider existing — the two footnoted rows are the exceptions, and say so.
The AMD ladder has no middle rung. TensorRT falls back to the CUDA provider while its engines
compile; MIGraphX has no such twin, so its fallback is the CPU provider, and the minute or two of
@@ -295,6 +341,40 @@ it.
Both positions are D13 territory and are recorded there (§12).
Two more runtimes ship in every desktop package since 0.23 (§3.2), and their parts are all
redistributable:
| Component | Licence | Shipped |
|---|---|---|
| Intel's `onnxruntime-openvino` build, with OpenVINO 2025.4.1 and oneTBB | MIT; Apache-2.0; Apache-2.0 | Linux and Windows packages, texts beside the libraries |
| Microsoft's WebGPU build (Dawn inside); on Windows the DirectX shader compiler | MIT; LLVM / MIT | Linux and Windows packages |
| Microsoft's stock `onnxruntime-android` (WebGPU) | MIT | The APK, as `libonnxruntime_generic.so` |
### 3.2 Several runtimes, one per process
A runtime carries one vendor's providers: Intel's build has OpenVINO, the `onnxruntime-gpu` wheel
CUDA and TensorRT, a ROCm build MIGraphX, Microsoft's WebGPU build the generic rung, the APK's QNN
build the Hexagon. No prebuilt carries two vendors, and `set_api` takes one table per process.
So a device that may hold several — the package's OpenVINO and WebGPU builds, a CUDA build the user
fetched, the distribution's ROCm build — has to choose which to load *before* the probe, and
cannot choose by trying.
`api::install` opens every runtime on the search list, asks each for `GetAvailableProviders`, and
loads the one that scores highest against the GPUs `hardware::detect` reads from files: a vendor
rung on its own vendor's GPU (NVIDIA driver, `/dev/kfd`, a Qualcomm SoC, macOS) above OpenVINO on
an Intel GPU (PCI vendor `0x8086`; on Windows Intel's DCH driver package) above WebGPU above a
CPU-only build. Equal scores keep the search order, a perfect fit ends the search — the APK's QNN
build is listed first, so on a Qualcomm device the generic build is never opened — and
`DARKROOM_ORT_DIR` wins outright. The losers stay mapped: unloading a C++ runtime whose static
constructors ran is a crash at exit waiting to happen.
The desktop packages install the two bundled builds under `runtimes/openvino` and
`runtimes/webgpu` beside each place a package installs to, from
`tools/fetch-bundled-runtimes.sh` (PyPI wheels pinned by SHA-256, pruned to the native libraries:
81 + 31 MB on Linux, 67 + 42 MB on Windows). On Windows the chosen runtime's directory is put on
`PATH`, because Intel's build leaves OpenVINO's DLLs for the loader to find there. The Flatpak has
no Intel OpenCL driver in its sandbox, so an Intel machine there settles on the CPU.
---
## 4. Selection — the probe, its cache, and what it may not do
@@ -603,6 +683,8 @@ device are comparable. *Acceptance:* M3.
**D13 — updated.** The runtime half is reopened to the extent of §3: the Rust build stays C-free
under `alternative-backend`; packages may install a dynamically loaded ONNX Runtime and, per §3.1,
the Qualcomm QNN runtime; the NVIDIA libraries are not bundled. The licensing half is unchanged.
Since 0.23 every desktop package bundles two runtimes — Intel's OpenVINO build and the WebGPU
build — and the APK a second, generic one; the engine loads the one that fits the GPU (§3.2).
---