Let the probe's clock be its proof, not disable_cpu_ep_fallback

The strict flag refused the Hexagon over the ten quantise/dequantise
nodes at the graph's edges that QNN declines by policy, which cost
microseconds. A provider that hands real work to the CPU is slower than
the CPU floor and the timing already rejects it; the tablet measured
2.3 ms on the NPU against a 29.7 ms floor.
This commit is contained in:
2026-09-19 16:13:19 +02:00
parent 691af96e3e
commit cbbe67fbd7
5 changed files with 26 additions and 29 deletions
+8 -5
View File
@@ -195,11 +195,14 @@ state where the provider registered and the session then failed, and a provider
took the graph, and rejected every node at partition time. The probe therefore:
1. Loads the runtime library (§3), or falls to tract and stops.
2. For each rung in this platform's ladder, in order: builds a session for the **smallest model
in the set** (`scrfd_500m`) on that provider with `error_on_failure`, runs it once on a fixed
input, and reads back the provider assignment from the session — the rung is taken only if the
provider ran **at least 95% of the graph's nodes**. A provider that silently hands the graph to
the CPU is the CPU rung with extra overhead, and the app should say "CPU".
2. Times the **smallest detector** on the CPU provider first — the floor. Then, for each rung
in this platform's ladder, in order: builds a session for the same model on that provider
with `error_on_failure`, runs it once on a fixed input, and times three more runs. **The rung
is taken only if its median beats the floor.** That one measurement is the proof the provider
took the graph: one that silently hands the work to the CPU is the CPU rung with extra
overhead, slower than the floor, and rejected. (ONNX Runtime's
`session.disable_cpu_ep_fallback` was the first draft of this proof and refuses the Hexagon
over the ten quantise/dequantise nodes at the graph's edges that QNN declines by policy.)
3. Records the outcome — rung, runtime version, provider version, device identity (GPU name and
compute capability; SoC model and Hexagon arch), and the models' content hashes — to a small
file beside `shared_face_models_dir`. The next start-up trusts the file **unless** any of those