diff --git a/docs/dev/inference.md b/docs/dev/inference.md index 3d7d1ad..4389ae6 100644 --- a/docs/dev/inference.md +++ b/docs/dev/inference.md @@ -194,6 +194,45 @@ int8's matches land 0.45 px from f32's in the overlaps — the spread f32 shows --- +### 1.6 Intel Iris Xe, and the generic rung · 2026-10-04 + +The RTX 3050 laptop's other GPU: Raptor Lake-P's Iris Xe (96 EU), Intel's `onnxruntime-openvino` +1.24.1 (OpenVINO 2025.4.1) and Microsoft's `onnxruntime-webgpu` 1.27.0, both PyPI wheels, through +`ep_probe`. Three warm-ups, the median of 15 runs. Another build shared the CPU during the run, so +the CPU columns are a little pessimistic; the GPU columns are not. + +| Model | ORT CPU f32 | OpenVINO CPU | OpenVINO GPU f32 | **OpenVINO GPU fp16** | WebGPU (Iris Xe) | +|---|---|---|---|---|---| +| scrfd_500m (Fast) | 9.5 | 10.8 | 7.3 | **5.8** | 24.2 | +| scrfd_2.5g (Balanced) | 18.5 | 16.2 | 15.5 | **11.1** | 40.8 | +| scrfd_10g (Thorough) | 58.5 | 72.4 | 38.7 | **23.0** | 84.5 | +| arcface_mbf (per face) | 9.5 | 11.7 | **3.0** | 2.4 | 56.6 | +| 2d106det (landmarks) | 10.8 | 2.0 | 2.4 | **1.9** | 42.7 | +| yolo26s-sem-ade20k | 57.0 | 47.4 | 26.0 | **16.7** | 53.0 | +| xfeat-1024 | 23.5 | 18.5 | 19.9 | **17.5** | 34.2 | +| migan-512 (per tile) | 330 | ✗ ¹ | 89.7 | **57.2** | 275 | +| mosaic-fast-1408 (per tile) | 159 | 107 | 68.1 | **40.0** | 188 ² | +| mosaic-best-1408 (per tile) | 1109 | 1670 | 947 | **604** | 1034 ² | + +¹ OpenVINO's CPU plugin refuses the graph at initialisation. Not shipped (§3.2), so moot. +² A later run, after `ep_probe` learned to feed the denoiser's two inputs, under heavier load: the +CPU provider took 256 and 1034 ms in that run, so WebGPU beat it by a quarter on mosaic-fast and tied +on mosaic-best — the only rows where it is not well behind. + +- **OpenVINO on the Iris Xe beats ONNX Runtime's CPU provider on every model**, 1.3× on XFeat to + 5.8× on MI-GAN, with a 1–3 s compile per graph and 0.1–0.4 s from its cache. It is the Intel rung. + fp16 is worth 1.3–1.7× over f32 here, against 1.1–1.35× on MIGraphX. +- **Its "GPU" is OpenCL's first GPU, not Intel's.** Before `intel-compute-runtime` was installed + the only OpenCL driver was NVIDIA's, and `device_type=GPU` ran on the RTX 3050 — slower than the + CPU, which is the probe's to catch. Read the process's maps for `libigdrcl` before believing a + number is the iGPU's. +- **WebGPU is slower than the CPU on the Iris Xe** on everything but MI-GAN, as it was on the + Adreno (§1.1), and on the RTX 3050 through Vulkan too. It is on the ladder anyway, as the generic + rung (§2): for GPUs no vendor rung covers — an AMD card on Windows or without ROCm, a Mali — where + it is unmeasured, and the probe's clock decides. +- **OpenVINO's CPU plugin is not a better floor.** It wins on some graphs and loses on scrfd_10g, + the embedder and mosaic-best, and refuses MI-GAN. + ## 2. The shape of the answer A **ladder per platform**, walked at start-up, with the first rung that builds a real session @@ -202,10 +241,11 @@ winning: | Platform | 1st | 2nd | 3rd | Floor | |---|---|---|---|---| | Android, Qualcomm with a Hexagon the shipped QNN skel covers (V68–V81) | QNN HTP, each model's quantised form (§1.5) | ORT CPU, f32 model | — | tract | -| Android, any other SoC | ORT CPU, f32 | — | — | tract | +| Android, any other SoC ⁶ | WebGPU (Vulkan), f32 | ORT CPU, f32 | — | tract | | Linux / Windows, NVIDIA GPU | TensorRT, f32 model, fp16 engine | CUDA provider, f32 | ORT CPU, f32 | tract | | Linux, AMD GPU with ROCm | MIGraphX, f32 model, fp16 program | ORT CPU, f32 | — | tract | -| Linux / Windows, no GPU stack | ORT CPU, f32 | — | — | tract | +| Linux / Windows, Intel GPU | OpenVINO, f32 model, fp16 program (§1.6) | ORT CPU, f32 | — | tract | +| Linux / Windows, any other GPU ⁶ | WebGPU (Vulkan / D3D12), f32 | ORT CPU, f32 | — | tract | | macOS ⁵ | CoreML, f32 model, ML Program | ORT CPU, f32 | — | tract | ⁵ **Unmeasured**, and the one exception to the rule below: nobody here has a Mac. The rung is on @@ -215,11 +255,17 @@ down is refused on the third launch (§4, `attempt`). The embedder stays on the macOS log that shows a probe line is this row's measurement; [macos.md](macos.md) says what to ask for. +⁶ **The generic rung, unmeasured where it is meant to help.** WebGPU lost to the CPU on every GPU +it has been timed on — the Adreno, the Iris Xe, the RTX 3050 (§1.1, §1.6) — none of which it serves +here, since each has its own rung. It is on the ladder for the GPUs that have none, on the same +terms as CoreML: a WebGPU that is slower than the CPU is rejected by §4's clock, one that errors is +recorded as failed. Its first measurement on an AMD card without ROCm, or a Mali, is this row's. + Deliberately **not** on any ladder, with the measurement that excluded each: NNAPI (no driver), -XNNPACK (slower than CPU, aborts on SCRFD), WebGPU (slower than CPU), the Adreno through QNN (works, -but never where the Hexagon does not also), CUDA int8 (slower than CUDA f32), the ROCm provider -(gone: §1.3). A rung is added to this table by a measurement on this page, not by a provider -existing. +XNNPACK (slower than CPU, aborts on SCRFD), the Adreno through QNN (works, but never where the +Hexagon does not also), CUDA int8 (slower than CUDA f32), the ROCm provider (gone: §1.3), +OpenVINO's CPU plugin as a floor (§1.6). A rung is added to this table by a measurement on this +page, not by a provider existing — the two footnoted rows are the exceptions, and say so. The AMD ladder has no middle rung. TensorRT falls back to the CUDA provider while its engines compile; MIGraphX has no such twin, so its fallback is the CPU provider, and the minute or two of @@ -295,6 +341,40 @@ it. Both positions are D13 territory and are recorded there (§12). +Two more runtimes ship in every desktop package since 0.23 (§3.2), and their parts are all +redistributable: + +| Component | Licence | Shipped | +|---|---|---| +| Intel's `onnxruntime-openvino` build, with OpenVINO 2025.4.1 and oneTBB | MIT; Apache-2.0; Apache-2.0 | Linux and Windows packages, texts beside the libraries | +| Microsoft's WebGPU build (Dawn inside); on Windows the DirectX shader compiler | MIT; LLVM / MIT | Linux and Windows packages | +| Microsoft's stock `onnxruntime-android` (WebGPU) | MIT | The APK, as `libonnxruntime_generic.so` | + +### 3.2 Several runtimes, one per process + +A runtime carries one vendor's providers: Intel's build has OpenVINO, the `onnxruntime-gpu` wheel +CUDA and TensorRT, a ROCm build MIGraphX, Microsoft's WebGPU build the generic rung, the APK's QNN +build the Hexagon. No prebuilt carries two vendors, and `set_api` takes one table per process. +So a device that may hold several — the package's OpenVINO and WebGPU builds, a CUDA build the user +fetched, the distribution's ROCm build — has to choose which to load *before* the probe, and +cannot choose by trying. + +`api::install` opens every runtime on the search list, asks each for `GetAvailableProviders`, and +loads the one that scores highest against the GPUs `hardware::detect` reads from files: a vendor +rung on its own vendor's GPU (NVIDIA driver, `/dev/kfd`, a Qualcomm SoC, macOS) above OpenVINO on +an Intel GPU (PCI vendor `0x8086`; on Windows Intel's DCH driver package) above WebGPU above a +CPU-only build. Equal scores keep the search order, a perfect fit ends the search — the APK's QNN +build is listed first, so on a Qualcomm device the generic build is never opened — and +`DARKROOM_ORT_DIR` wins outright. The losers stay mapped: unloading a C++ runtime whose static +constructors ran is a crash at exit waiting to happen. + +The desktop packages install the two bundled builds under `runtimes/openvino` and +`runtimes/webgpu` beside each place a package installs to, from +`tools/fetch-bundled-runtimes.sh` (PyPI wheels pinned by SHA-256, pruned to the native libraries: +81 + 31 MB on Linux, 67 + 42 MB on Windows). On Windows the chosen runtime's directory is put on +`PATH`, because Intel's build leaves OpenVINO's DLLs for the loader to find there. The Flatpak has +no Intel OpenCL driver in its sandbox, so an Intel machine there settles on the CPU. + --- ## 4. Selection — the probe, its cache, and what it may not do @@ -603,6 +683,8 @@ device are comparable. *Acceptance:* M3. **D13 — updated.** The runtime half is reopened to the extent of §3: the Rust build stays C-free under `alternative-backend`; packages may install a dynamically loaded ONNX Runtime and, per §3.1, the Qualcomm QNN runtime; the NVIDIA libraries are not bundled. The licensing half is unchanged. +Since 0.23 every desktop package bundles two runtimes — Intel's OpenVINO build and the WebGPU +build — and the APK a second, generic one; the engine loads the one that fits the GPU (§3.2). ---