Let the probe's clock be its proof, not disable_cpu_ep_fallback
The strict flag refused the Hexagon over the ten quantise/dequantise nodes at the graph's edges that QNN declines by policy, which cost microseconds. A provider that hands real work to the CPU is slower than the CPU floor and the timing already rejects it; the tablet measured 2.3 ms on the NPU against a 29.7 ms floor.
This commit is contained in:
+8
-5
@@ -195,11 +195,14 @@ state where the provider registered and the session then failed, and a provider
|
||||
took the graph, and rejected every node at partition time. The probe therefore:
|
||||
|
||||
1. Loads the runtime library (§3), or falls to tract and stops.
|
||||
2. For each rung in this platform's ladder, in order: builds a session for the **smallest model
|
||||
in the set** (`scrfd_500m`) on that provider with `error_on_failure`, runs it once on a fixed
|
||||
input, and reads back the provider assignment from the session — the rung is taken only if the
|
||||
provider ran **at least 95% of the graph's nodes**. A provider that silently hands the graph to
|
||||
the CPU is the CPU rung with extra overhead, and the app should say "CPU".
|
||||
2. Times the **smallest detector** on the CPU provider first — the floor. Then, for each rung
|
||||
in this platform's ladder, in order: builds a session for the same model on that provider
|
||||
with `error_on_failure`, runs it once on a fixed input, and times three more runs. **The rung
|
||||
is taken only if its median beats the floor.** That one measurement is the proof the provider
|
||||
took the graph: one that silently hands the work to the CPU is the CPU rung with extra
|
||||
overhead, slower than the floor, and rejected. (ONNX Runtime's
|
||||
`session.disable_cpu_ep_fallback` was the first draft of this proof and refuses the Hexagon
|
||||
over the ten quantise/dequantise nodes at the graph's edges that QNN declines by policy.)
|
||||
3. Records the outcome — rung, runtime version, provider version, device identity (GPU name and
|
||||
compute capability; SoC model and Hexagon arch), and the models' content hashes — to a small
|
||||
file beside `shared_face_models_dir`. The next start-up trusts the file **unless** any of those
|
||||
|
||||
Reference in New Issue
Block a user