Record the Intel and generic rungs in the inference spec
§1.6 is the Iris Xe measurement: OpenVINO fp16 1.3–5.8× the CPU provider on every shipped model, WebGPU behind it everywhere but the denoiser and MI-GAN. §2's ladder gains the Intel and generic rows, the generic one footnoted as unmeasured where it is meant to help. §3.1 lists the two bundled runtimes' licences; §3.2 is how one runtime of several is chosen per process. D13 notes the bundling.
This commit is contained in:
+88
-6
@@ -194,6 +194,45 @@ int8's matches land 0.45 px from f32's in the overlaps — the spread f32 shows
|
||||
|
||||
---
|
||||
|
||||
### 1.6 Intel Iris Xe, and the generic rung · 2026-10-04
|
||||
|
||||
The RTX 3050 laptop's other GPU: Raptor Lake-P's Iris Xe (96 EU), Intel's `onnxruntime-openvino`
|
||||
1.24.1 (OpenVINO 2025.4.1) and Microsoft's `onnxruntime-webgpu` 1.27.0, both PyPI wheels, through
|
||||
`ep_probe`. Three warm-ups, the median of 15 runs. Another build shared the CPU during the run, so
|
||||
the CPU columns are a little pessimistic; the GPU columns are not.
|
||||
|
||||
| Model | ORT CPU f32 | OpenVINO CPU | OpenVINO GPU f32 | **OpenVINO GPU fp16** | WebGPU (Iris Xe) |
|
||||
|---|---|---|---|---|---|
|
||||
| scrfd_500m (Fast) | 9.5 | 10.8 | 7.3 | **5.8** | 24.2 |
|
||||
| scrfd_2.5g (Balanced) | 18.5 | 16.2 | 15.5 | **11.1** | 40.8 |
|
||||
| scrfd_10g (Thorough) | 58.5 | 72.4 | 38.7 | **23.0** | 84.5 |
|
||||
| arcface_mbf (per face) | 9.5 | 11.7 | **3.0** | 2.4 | 56.6 |
|
||||
| 2d106det (landmarks) | 10.8 | 2.0 | 2.4 | **1.9** | 42.7 |
|
||||
| yolo26s-sem-ade20k | 57.0 | 47.4 | 26.0 | **16.7** | 53.0 |
|
||||
| xfeat-1024 | 23.5 | 18.5 | 19.9 | **17.5** | 34.2 |
|
||||
| migan-512 (per tile) | 330 | ✗ ¹ | 89.7 | **57.2** | 275 |
|
||||
| mosaic-fast-1408 (per tile) | 159 | 107 | 68.1 | **40.0** | 188 ² |
|
||||
| mosaic-best-1408 (per tile) | 1109 | 1670 | 947 | **604** | 1034 ² |
|
||||
|
||||
¹ OpenVINO's CPU plugin refuses the graph at initialisation. Not shipped (§3.2), so moot.
|
||||
² A later run, after `ep_probe` learned to feed the denoiser's two inputs, under heavier load: the
|
||||
CPU provider took 256 and 1034 ms in that run, so WebGPU beat it by a quarter on mosaic-fast and tied
|
||||
on mosaic-best — the only rows where it is not well behind.
|
||||
|
||||
- **OpenVINO on the Iris Xe beats ONNX Runtime's CPU provider on every model**, 1.3× on XFeat to
|
||||
5.8× on MI-GAN, with a 1–3 s compile per graph and 0.1–0.4 s from its cache. It is the Intel rung.
|
||||
fp16 is worth 1.3–1.7× over f32 here, against 1.1–1.35× on MIGraphX.
|
||||
- **Its "GPU" is OpenCL's first GPU, not Intel's.** Before `intel-compute-runtime` was installed
|
||||
the only OpenCL driver was NVIDIA's, and `device_type=GPU` ran on the RTX 3050 — slower than the
|
||||
CPU, which is the probe's to catch. Read the process's maps for `libigdrcl` before believing a
|
||||
number is the iGPU's.
|
||||
- **WebGPU is slower than the CPU on the Iris Xe** on everything but MI-GAN, as it was on the
|
||||
Adreno (§1.1), and on the RTX 3050 through Vulkan too. It is on the ladder anyway, as the generic
|
||||
rung (§2): for GPUs no vendor rung covers — an AMD card on Windows or without ROCm, a Mali — where
|
||||
it is unmeasured, and the probe's clock decides.
|
||||
- **OpenVINO's CPU plugin is not a better floor.** It wins on some graphs and loses on scrfd_10g,
|
||||
the embedder and mosaic-best, and refuses MI-GAN.
|
||||
|
||||
## 2. The shape of the answer
|
||||
|
||||
A **ladder per platform**, walked at start-up, with the first rung that builds a real session
|
||||
@@ -202,10 +241,11 @@ winning:
|
||||
| Platform | 1st | 2nd | 3rd | Floor |
|
||||
|---|---|---|---|---|
|
||||
| Android, Qualcomm with a Hexagon the shipped QNN skel covers (V68–V81) | QNN HTP, each model's quantised form (§1.5) | ORT CPU, f32 model | — | tract |
|
||||
| Android, any other SoC | ORT CPU, f32 | — | — | tract |
|
||||
| Android, any other SoC ⁶ | WebGPU (Vulkan), f32 | ORT CPU, f32 | — | tract |
|
||||
| Linux / Windows, NVIDIA GPU | TensorRT, f32 model, fp16 engine | CUDA provider, f32 | ORT CPU, f32 | tract |
|
||||
| Linux, AMD GPU with ROCm | MIGraphX, f32 model, fp16 program | ORT CPU, f32 | — | tract |
|
||||
| Linux / Windows, no GPU stack | ORT CPU, f32 | — | — | tract |
|
||||
| Linux / Windows, Intel GPU | OpenVINO, f32 model, fp16 program (§1.6) | ORT CPU, f32 | — | tract |
|
||||
| Linux / Windows, any other GPU ⁶ | WebGPU (Vulkan / D3D12), f32 | ORT CPU, f32 | — | tract |
|
||||
| macOS ⁵ | CoreML, f32 model, ML Program | ORT CPU, f32 | — | tract |
|
||||
|
||||
⁵ **Unmeasured**, and the one exception to the rule below: nobody here has a Mac. The rung is on
|
||||
@@ -215,11 +255,17 @@ down is refused on the third launch (§4, `attempt`). The embedder stays on the
|
||||
macOS log that shows a probe line is this row's measurement; [macos.md](macos.md) says what to
|
||||
ask for.
|
||||
|
||||
⁶ **The generic rung, unmeasured where it is meant to help.** WebGPU lost to the CPU on every GPU
|
||||
it has been timed on — the Adreno, the Iris Xe, the RTX 3050 (§1.1, §1.6) — none of which it serves
|
||||
here, since each has its own rung. It is on the ladder for the GPUs that have none, on the same
|
||||
terms as CoreML: a WebGPU that is slower than the CPU is rejected by §4's clock, one that errors is
|
||||
recorded as failed. Its first measurement on an AMD card without ROCm, or a Mali, is this row's.
|
||||
|
||||
Deliberately **not** on any ladder, with the measurement that excluded each: NNAPI (no driver),
|
||||
XNNPACK (slower than CPU, aborts on SCRFD), WebGPU (slower than CPU), the Adreno through QNN (works,
|
||||
but never where the Hexagon does not also), CUDA int8 (slower than CUDA f32), the ROCm provider
|
||||
(gone: §1.3). A rung is added to this table by a measurement on this page, not by a provider
|
||||
existing.
|
||||
XNNPACK (slower than CPU, aborts on SCRFD), the Adreno through QNN (works, but never where the
|
||||
Hexagon does not also), CUDA int8 (slower than CUDA f32), the ROCm provider (gone: §1.3),
|
||||
OpenVINO's CPU plugin as a floor (§1.6). A rung is added to this table by a measurement on this
|
||||
page, not by a provider existing — the two footnoted rows are the exceptions, and say so.
|
||||
|
||||
The AMD ladder has no middle rung. TensorRT falls back to the CUDA provider while its engines
|
||||
compile; MIGraphX has no such twin, so its fallback is the CPU provider, and the minute or two of
|
||||
@@ -295,6 +341,40 @@ it.
|
||||
|
||||
Both positions are D13 territory and are recorded there (§12).
|
||||
|
||||
Two more runtimes ship in every desktop package since 0.23 (§3.2), and their parts are all
|
||||
redistributable:
|
||||
|
||||
| Component | Licence | Shipped |
|
||||
|---|---|---|
|
||||
| Intel's `onnxruntime-openvino` build, with OpenVINO 2025.4.1 and oneTBB | MIT; Apache-2.0; Apache-2.0 | Linux and Windows packages, texts beside the libraries |
|
||||
| Microsoft's WebGPU build (Dawn inside); on Windows the DirectX shader compiler | MIT; LLVM / MIT | Linux and Windows packages |
|
||||
| Microsoft's stock `onnxruntime-android` (WebGPU) | MIT | The APK, as `libonnxruntime_generic.so` |
|
||||
|
||||
### 3.2 Several runtimes, one per process
|
||||
|
||||
A runtime carries one vendor's providers: Intel's build has OpenVINO, the `onnxruntime-gpu` wheel
|
||||
CUDA and TensorRT, a ROCm build MIGraphX, Microsoft's WebGPU build the generic rung, the APK's QNN
|
||||
build the Hexagon. No prebuilt carries two vendors, and `set_api` takes one table per process.
|
||||
So a device that may hold several — the package's OpenVINO and WebGPU builds, a CUDA build the user
|
||||
fetched, the distribution's ROCm build — has to choose which to load *before* the probe, and
|
||||
cannot choose by trying.
|
||||
|
||||
`api::install` opens every runtime on the search list, asks each for `GetAvailableProviders`, and
|
||||
loads the one that scores highest against the GPUs `hardware::detect` reads from files: a vendor
|
||||
rung on its own vendor's GPU (NVIDIA driver, `/dev/kfd`, a Qualcomm SoC, macOS) above OpenVINO on
|
||||
an Intel GPU (PCI vendor `0x8086`; on Windows Intel's DCH driver package) above WebGPU above a
|
||||
CPU-only build. Equal scores keep the search order, a perfect fit ends the search — the APK's QNN
|
||||
build is listed first, so on a Qualcomm device the generic build is never opened — and
|
||||
`DARKROOM_ORT_DIR` wins outright. The losers stay mapped: unloading a C++ runtime whose static
|
||||
constructors ran is a crash at exit waiting to happen.
|
||||
|
||||
The desktop packages install the two bundled builds under `runtimes/openvino` and
|
||||
`runtimes/webgpu` beside each place a package installs to, from
|
||||
`tools/fetch-bundled-runtimes.sh` (PyPI wheels pinned by SHA-256, pruned to the native libraries:
|
||||
81 + 31 MB on Linux, 67 + 42 MB on Windows). On Windows the chosen runtime's directory is put on
|
||||
`PATH`, because Intel's build leaves OpenVINO's DLLs for the loader to find there. The Flatpak has
|
||||
no Intel OpenCL driver in its sandbox, so an Intel machine there settles on the CPU.
|
||||
|
||||
---
|
||||
|
||||
## 4. Selection — the probe, its cache, and what it may not do
|
||||
@@ -603,6 +683,8 @@ device are comparable. *Acceptance:* M3.
|
||||
**D13 — updated.** The runtime half is reopened to the extent of §3: the Rust build stays C-free
|
||||
under `alternative-backend`; packages may install a dynamically loaded ONNX Runtime and, per §3.1,
|
||||
the Qualcomm QNN runtime; the NVIDIA libraries are not bundled. The licensing half is unchanged.
|
||||
Since 0.23 every desktop package bundles two runtimes — Intel's OpenVINO build and the WebGPU
|
||||
build — and the APK a second, generic one; the engine loads the one that fits the GPU (§3.2).
|
||||
|
||||
---
|
||||
|
||||
|
||||
Reference in New Issue
Block a user