Let the probe's clock be its proof, not disable_cpu_ep_fallback
The strict flag refused the Hexagon over the ten quantise/dequantise nodes at the graph's edges that QNN declines by policy, which cost microseconds. A provider that hands real work to the CPU is slower than the CPU floor and the timing already rejects it; the tablet measured 2.3 ms on the NPU against a 29.7 ms floor.
This commit is contained in:
@@ -89,7 +89,7 @@ pub fn run() {
|
|||||||
}
|
}
|
||||||
log::info!("inference: compiling {name} for {}", rung.label());
|
log::info!("inference: compiling {name} for {}", rung.label());
|
||||||
let started = std::time::Instant::now();
|
let started = std::time::Instant::now();
|
||||||
match crate::session::build(rung, role, &bytes, &cfg, false) {
|
match crate::session::build(rung, role, &bytes, &cfg) {
|
||||||
Ok(session) => {
|
Ok(session) => {
|
||||||
drop(session);
|
drop(session);
|
||||||
let mut s = state().lock().unwrap();
|
let mut s = state().lock().unwrap();
|
||||||
|
|||||||
@@ -253,7 +253,7 @@ fn acquire(role: Role, form: Form, bytes: &Arc<[u8]>) -> Result<Acquired, Error>
|
|||||||
|
|
||||||
// Built outside the registry lock: a TensorRT engine load is long enough
|
// Built outside the registry lock: a TensorRT engine load is long enough
|
||||||
// that another role's acquire should not wait on it.
|
// that another role's acquire should not wait on it.
|
||||||
let session = session::build(rung, role, bytes, &cfg, false)?;
|
let session = session::build(rung, role, bytes, &cfg)?;
|
||||||
log::debug!("inference: {role:?} loaded on {}", rung.label());
|
log::debug!("inference: {role:?} loaded on {}", rung.label());
|
||||||
let entry = Arc::new(Loaded {
|
let entry = Arc::new(Loaded {
|
||||||
rung,
|
rung,
|
||||||
|
|||||||
@@ -1,9 +1,9 @@
|
|||||||
//! Walk the ladder, once, by building real sessions (docs/inference.md §4).
|
//! Walk the ladder, once, by building real sessions (docs/inference.md §4).
|
||||||
//!
|
//!
|
||||||
//! A rung is taken when a strict session builds on it, runs, and is faster
|
//! A rung is taken when a session builds on it, runs, and is faster than
|
||||||
//! than the floor. Both halves matter: a provider can register and then fail
|
//! the floor. Both halves matter: a provider can register and then fail at
|
||||||
//! at partition time, and a provider can take a graph and run it slower than
|
//! partition time, and a provider can take a graph — or quietly hand most
|
||||||
//! the CPU would have. The outcome is cached against a fingerprint of the
|
//! of it back to the CPU — and run it slower than the CPU would have. The outcome is cached against a fingerprint of the
|
||||||
//! runtime, the driver, the hardware and the models, and trusted until any
|
//! runtime, the driver, the hardware and the models, and trusted until any
|
||||||
//! of those changes.
|
//! of those changes.
|
||||||
|
|
||||||
@@ -135,9 +135,9 @@ fn probe_model(cfg: &Config) -> Option<(Role, PathBuf)> {
|
|||||||
smallest(Some(Role::Detector)).or_else(|| smallest(None))
|
smallest(Some(Role::Detector)).or_else(|| smallest(None))
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Build strictly, run once for the engine, then time three runs; the
|
/// Build, run once for the engine, then time three runs; the median in
|
||||||
/// median in milliseconds and, for a compiling rung, the cache key of the
|
/// milliseconds and, for a compiling rung, the cache key of the engine this
|
||||||
/// engine this just built.
|
/// just built.
|
||||||
fn time_rung(
|
fn time_rung(
|
||||||
rung: Rung,
|
rung: Rung,
|
||||||
role: Role,
|
role: Role,
|
||||||
@@ -157,8 +157,8 @@ fn time_rung(
|
|||||||
};
|
};
|
||||||
let bytes = std::fs::read(&path).map_err(|e| e.to_string())?;
|
let bytes = std::fs::read(&path).map_err(|e| e.to_string())?;
|
||||||
let started = Instant::now();
|
let started = Instant::now();
|
||||||
let mut session = crate::session::build(rung, role, &bytes, cfg, true)
|
let mut session =
|
||||||
.map_err(|e| first_line(&e.to_string()))?;
|
crate::session::build(rung, role, &bytes, cfg).map_err(|e| first_line(&e.to_string()))?;
|
||||||
log::info!(
|
log::info!(
|
||||||
"inference: {} session built in {:.1} s",
|
"inference: {} session built in {:.1} s",
|
||||||
rung.label(),
|
rung.label(),
|
||||||
|
|||||||
@@ -7,22 +7,16 @@ use crate::{Config, Role, Rung};
|
|||||||
|
|
||||||
/// Build a session for `bytes` on `rung`.
|
/// Build a session for `bytes` on `rung`.
|
||||||
///
|
///
|
||||||
/// `strict` is the probe's flag: with it, a provider that would hand any
|
/// Not strict about the CPU: `session.disable_cpu_ep_fallback` was tried as
|
||||||
/// node to the CPU fails the build instead, so "the session built" means
|
/// the probe's proof that a provider took the graph, and it refuses the
|
||||||
/// "the provider took the graph" and not "the provider registered" (§4).
|
/// Hexagon over the ten quantise/dequantise nodes at the graph's edges that
|
||||||
pub fn build(
|
/// QNN declines by policy and that cost microseconds. The probe's proof is
|
||||||
rung: Rung,
|
/// its clock instead (§4): a provider that hands real work to the CPU is
|
||||||
role: Role,
|
/// slower than the CPU floor and rejected by the same measurement.
|
||||||
bytes: &[u8],
|
pub fn build(rung: Rung, role: Role, bytes: &[u8], cfg: &Config) -> ort::Result<Session> {
|
||||||
cfg: &Config,
|
|
||||||
strict: bool,
|
|
||||||
) -> ort::Result<Session> {
|
|
||||||
let mut b = Session::builder()?
|
let mut b = Session::builder()?
|
||||||
.with_optimization_level(GraphOptimizationLevel::Level3)?
|
.with_optimization_level(GraphOptimizationLevel::Level3)?
|
||||||
.with_intra_threads(threads(cfg))?;
|
.with_intra_threads(threads(cfg))?;
|
||||||
if strict && rung != Rung::Cpu {
|
|
||||||
b = b.with_config_entry("session.disable_cpu_ep_fallback", "1")?;
|
|
||||||
}
|
|
||||||
// A Hexagon session loads the compiled context when there is one and
|
// A Hexagon session loads the compiled context when there is one and
|
||||||
// compiles it from the model when there is not; the engine thread is
|
// compiles it from the model when there is not; the engine thread is
|
||||||
// what makes the second case rare (§6).
|
// what makes the second case rare (§6).
|
||||||
|
|||||||
+8
-5
@@ -195,11 +195,14 @@ state where the provider registered and the session then failed, and a provider
|
|||||||
took the graph, and rejected every node at partition time. The probe therefore:
|
took the graph, and rejected every node at partition time. The probe therefore:
|
||||||
|
|
||||||
1. Loads the runtime library (§3), or falls to tract and stops.
|
1. Loads the runtime library (§3), or falls to tract and stops.
|
||||||
2. For each rung in this platform's ladder, in order: builds a session for the **smallest model
|
2. Times the **smallest detector** on the CPU provider first — the floor. Then, for each rung
|
||||||
in the set** (`scrfd_500m`) on that provider with `error_on_failure`, runs it once on a fixed
|
in this platform's ladder, in order: builds a session for the same model on that provider
|
||||||
input, and reads back the provider assignment from the session — the rung is taken only if the
|
with `error_on_failure`, runs it once on a fixed input, and times three more runs. **The rung
|
||||||
provider ran **at least 95% of the graph's nodes**. A provider that silently hands the graph to
|
is taken only if its median beats the floor.** That one measurement is the proof the provider
|
||||||
the CPU is the CPU rung with extra overhead, and the app should say "CPU".
|
took the graph: one that silently hands the work to the CPU is the CPU rung with extra
|
||||||
|
overhead, slower than the floor, and rejected. (ONNX Runtime's
|
||||||
|
`session.disable_cpu_ep_fallback` was the first draft of this proof and refuses the Hexagon
|
||||||
|
over the ten quantise/dequantise nodes at the graph's edges that QNN declines by policy.)
|
||||||
3. Records the outcome — rung, runtime version, provider version, device identity (GPU name and
|
3. Records the outcome — rung, runtime version, provider version, device identity (GPU name and
|
||||||
compute capability; SoC model and Hexagon arch), and the models' content hashes — to a small
|
compute capability; SoC model and Hexagon arch), and the models' content hashes — to a small
|
||||||
file beside `shared_face_models_dir`. The next start-up trusts the file **unless** any of those
|
file beside `shared_face_models_dir`. The next start-up trusts the file **unless** any of those
|
||||||
|
|||||||
Reference in New Issue
Block a user