Let the probe's clock be its proof, not disable_cpu_ep_fallback

The strict flag refused the Hexagon over the ten quantise/dequantise
nodes at the graph's edges that QNN declines by policy, which cost
microseconds. A provider that hands real work to the CPU is slower than
the CPU floor and the timing already rejects it; the tablet measured
2.3 ms on the NPU against a 29.7 ms floor.
This commit is contained in:
2026-09-19 16:13:19 +02:00
parent 691af96e3e
commit cbbe67fbd7
5 changed files with 26 additions and 29 deletions
+1 -1
View File
@@ -89,7 +89,7 @@ pub fn run() {
} }
log::info!("inference: compiling {name} for {}", rung.label()); log::info!("inference: compiling {name} for {}", rung.label());
let started = std::time::Instant::now(); let started = std::time::Instant::now();
match crate::session::build(rung, role, &bytes, &cfg, false) { match crate::session::build(rung, role, &bytes, &cfg) {
Ok(session) => { Ok(session) => {
drop(session); drop(session);
let mut s = state().lock().unwrap(); let mut s = state().lock().unwrap();
+1 -1
View File
@@ -253,7 +253,7 @@ fn acquire(role: Role, form: Form, bytes: &Arc<[u8]>) -> Result<Acquired, Error>
// Built outside the registry lock: a TensorRT engine load is long enough // Built outside the registry lock: a TensorRT engine load is long enough
// that another role's acquire should not wait on it. // that another role's acquire should not wait on it.
let session = session::build(rung, role, bytes, &cfg, false)?; let session = session::build(rung, role, bytes, &cfg)?;
log::debug!("inference: {role:?} loaded on {}", rung.label()); log::debug!("inference: {role:?} loaded on {}", rung.label());
let entry = Arc::new(Loaded { let entry = Arc::new(Loaded {
rung, rung,
+9 -9
View File
@@ -1,9 +1,9 @@
//! Walk the ladder, once, by building real sessions (docs/inference.md §4). //! Walk the ladder, once, by building real sessions (docs/inference.md §4).
//! //!
//! A rung is taken when a strict session builds on it, runs, and is faster //! A rung is taken when a session builds on it, runs, and is faster than
//! than the floor. Both halves matter: a provider can register and then fail //! the floor. Both halves matter: a provider can register and then fail at
//! at partition time, and a provider can take a graph and run it slower than //! partition time, and a provider can take a graph — or quietly hand most
//! the CPU would have. The outcome is cached against a fingerprint of the //! of it back to the CPU — and run it slower than the CPU would have. The outcome is cached against a fingerprint of the
//! runtime, the driver, the hardware and the models, and trusted until any //! runtime, the driver, the hardware and the models, and trusted until any
//! of those changes. //! of those changes.
@@ -135,9 +135,9 @@ fn probe_model(cfg: &Config) -> Option<(Role, PathBuf)> {
smallest(Some(Role::Detector)).or_else(|| smallest(None)) smallest(Some(Role::Detector)).or_else(|| smallest(None))
} }
/// Build strictly, run once for the engine, then time three runs; the /// Build, run once for the engine, then time three runs; the median in
/// median in milliseconds and, for a compiling rung, the cache key of the /// milliseconds and, for a compiling rung, the cache key of the engine this
/// engine this just built. /// just built.
fn time_rung( fn time_rung(
rung: Rung, rung: Rung,
role: Role, role: Role,
@@ -157,8 +157,8 @@ fn time_rung(
}; };
let bytes = std::fs::read(&path).map_err(|e| e.to_string())?; let bytes = std::fs::read(&path).map_err(|e| e.to_string())?;
let started = Instant::now(); let started = Instant::now();
let mut session = crate::session::build(rung, role, &bytes, cfg, true) let mut session =
.map_err(|e| first_line(&e.to_string()))?; crate::session::build(rung, role, &bytes, cfg).map_err(|e| first_line(&e.to_string()))?;
log::info!( log::info!(
"inference: {} session built in {:.1} s", "inference: {} session built in {:.1} s",
rung.label(), rung.label(),
+7 -13
View File
@@ -7,22 +7,16 @@ use crate::{Config, Role, Rung};
/// Build a session for `bytes` on `rung`. /// Build a session for `bytes` on `rung`.
/// ///
/// `strict` is the probe's flag: with it, a provider that would hand any /// Not strict about the CPU: `session.disable_cpu_ep_fallback` was tried as
/// node to the CPU fails the build instead, so "the session built" means /// the probe's proof that a provider took the graph, and it refuses the
/// "the provider took the graph" and not "the provider registered" (§4). /// Hexagon over the ten quantise/dequantise nodes at the graph's edges that
pub fn build( /// QNN declines by policy and that cost microseconds. The probe's proof is
rung: Rung, /// its clock instead (§4): a provider that hands real work to the CPU is
role: Role, /// slower than the CPU floor and rejected by the same measurement.
bytes: &[u8], pub fn build(rung: Rung, role: Role, bytes: &[u8], cfg: &Config) -> ort::Result<Session> {
cfg: &Config,
strict: bool,
) -> ort::Result<Session> {
let mut b = Session::builder()? let mut b = Session::builder()?
.with_optimization_level(GraphOptimizationLevel::Level3)? .with_optimization_level(GraphOptimizationLevel::Level3)?
.with_intra_threads(threads(cfg))?; .with_intra_threads(threads(cfg))?;
if strict && rung != Rung::Cpu {
b = b.with_config_entry("session.disable_cpu_ep_fallback", "1")?;
}
// A Hexagon session loads the compiled context when there is one and // A Hexagon session loads the compiled context when there is one and
// compiles it from the model when there is not; the engine thread is // compiles it from the model when there is not; the engine thread is
// what makes the second case rare (§6). // what makes the second case rare (§6).
+8 -5
View File
@@ -195,11 +195,14 @@ state where the provider registered and the session then failed, and a provider
took the graph, and rejected every node at partition time. The probe therefore: took the graph, and rejected every node at partition time. The probe therefore:
1. Loads the runtime library (§3), or falls to tract and stops. 1. Loads the runtime library (§3), or falls to tract and stops.
2. For each rung in this platform's ladder, in order: builds a session for the **smallest model 2. Times the **smallest detector** on the CPU provider first — the floor. Then, for each rung
in the set** (`scrfd_500m`) on that provider with `error_on_failure`, runs it once on a fixed in this platform's ladder, in order: builds a session for the same model on that provider
input, and reads back the provider assignment from the session — the rung is taken only if the with `error_on_failure`, runs it once on a fixed input, and times three more runs. **The rung
provider ran **at least 95% of the graph's nodes**. A provider that silently hands the graph to is taken only if its median beats the floor.** That one measurement is the proof the provider
the CPU is the CPU rung with extra overhead, and the app should say "CPU". took the graph: one that silently hands the work to the CPU is the CPU rung with extra
overhead, slower than the floor, and rejected. (ONNX Runtime's
`session.disable_cpu_ep_fallback` was the first draft of this proof and refuses the Hexagon
over the ten quantise/dequantise nodes at the graph's edges that QNN declines by policy.)
3. Records the outcome — rung, runtime version, provider version, device identity (GPU name and 3. Records the outcome — rung, runtime version, provider version, device identity (GPU name and
compute capability; SoC model and Hexagon arch), and the models' content hashes — to a small compute capability; SoC model and Hexagon arch), and the models' content hashes — to a small
file beside `shared_face_models_dir`. The next start-up trusts the file **unless** any of those file beside `shared_face_models_dir`. The next start-up trusts the file **unless** any of those