75e12441fdbe5d56dd5f25fe108956c436a7e1ba
2
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
5a8c3e4c40 |
Run each model on the Hexagon in the form measured to hold it
The engine knew f32 and int8, and gave the Hexagon int8 for every role it served. Measured on the tablet itself (inference.md §1.5), int8 lost 5% of the detector's faces at 40-80 px, moved the landmarks 1.5 px, emptied the segmenter's scores and cost the denoiser 5-9 dB; fp16 the HTP refuses outright. `Form` gains A16W8 and A16W16, and `Rung::form` now names one per role: detectors and landmarks A16W8, the segmenter, scene model, border filler and denoiser A16W16, XFeat int8. The embedder and the eye classifiers stay on the CPU. Each loader resolves its `<stem>.<form>.onnx` sibling; the segmenter and XFeat, compiled into the binary, embed their quantised forms on Android only and pick through `choose_embedded`. The probe, the compile step and the cache fingerprint follow the form instead of assuming int8. Detectors on the new form write `scrfd_*_a16+w600k_mbf`, and `model_ids` answers for all three spellings. On the tablet (ORT 1.29 + QNN 2.42), each shipped file against f32 on the same inputs, and against the CPU's f32 time: SCRFD 500m/2.5g/10g A16W8 100% of faces in every band 4.2/5.1/9.0 ms vs 17/56/198 landmarks A16W8 0.25 px in the 192 crop 0.5 ms vs 2.8 YOLO26n-seg A16W16 98.2% found, mask IoU 0.994 12.9 ms vs 90 scene model A16W16 98.9% of cells agree 15 ms vs 151 MI-GAN A16W16 41 dB from f32 in the fill 87 ms vs 488 XFeat int8 pano alignment 0.45 px (f32's own spread 0.41) 6.5 ms vs 58 denoiser A16W16 0.00 dB at every ISO 95 ms vs 1510 a tile Face numbers are over public COCO val2017 photographs, not a library. The APK carries the siblings (BUNDLED 15 -> 19; the old int8 detectors removed), about 43 MB more. The Windows installer and its CI count skip them; the Arch and Flatpak packages list their files and never had them. The ladder example takes a role per model, which is how the per-role forms above were seen landing on the NPU from the real probe. |
||
|
|
39a22875b1 |
Add the MIGraphX rung for AMD GPUs
Benchmarks / CPU and I/O (per commit) (push) Failing after 6m20s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 45s
Build and test / Layer separation (push) Successful in 26s
Traceability / Requirement traces (push) Failing after 46s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Failing after 2m19s
Build and test / Windows (x86_64, cross) (push) Failing after 3m2s
Measured on a Radeon RX 7900 XT against Arch's onnxruntime-rocm 1.29 (docs/inference.md §1.3): MIGraphX fp16 runs the detectors at 2.4–3.4 ms against 10–58 ms on the CPU provider, the inpainter at 8 ms against 514, with a 15–135 s compile per graph the first time and under a second from its cache after. A compiling rung on TensorRT's terms, wired the same way. The ROCm execution provider is gone (removed in ONNX Runtime 1.23), so the AMD ladder is MIGraphX then the CPU, with no non-compiling rung between. MIGraphX is registered through the runtime's generic key/value entry point rather than ort's builder: 1.29 reads the legacy options struct for its precision flags only, and the compiled-program cache directory (`migraphx_model_cache_dir`) only travels the generic way. The provider's cache key omits the precision, so f32 and fp16 programs get their own directories. The probe fingerprint now includes the provider libraries beside the runtime and the ROCm version, since a distribution's CPU and ROCm builds are the same file at the same path. `status().failed` reports only the rungs above the selection, so an AMD desktop's About line says why MIGraphX won rather than that the NVIDIA providers are not in the build. Two examples: `ep_probe` times each provider cold and from cache, and `ladder` drives `init` as the app does to watch the first-run sequence. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |