Run each model on the Hexagon in the form measured to hold it
The engine knew f32 and int8, and gave the Hexagon int8 for every role it served. Measured on the tablet itself (inference.md §1.5), int8 lost 5% of the detector's faces at 40-80 px, moved the landmarks 1.5 px, emptied the segmenter's scores and cost the denoiser 5-9 dB; fp16 the HTP refuses outright. `Form` gains A16W8 and A16W16, and `Rung::form` now names one per role: detectors and landmarks A16W8, the segmenter, scene model, border filler and denoiser A16W16, XFeat int8. The embedder and the eye classifiers stay on the CPU. Each loader resolves its `<stem>.<form>.onnx` sibling; the segmenter and XFeat, compiled into the binary, embed their quantised forms on Android only and pick through `choose_embedded`. The probe, the compile step and the cache fingerprint follow the form instead of assuming int8. Detectors on the new form write `scrfd_*_a16+w600k_mbf`, and `model_ids` answers for all three spellings. On the tablet (ORT 1.29 + QNN 2.42), each shipped file against f32 on the same inputs, and against the CPU's f32 time: SCRFD 500m/2.5g/10g A16W8 100% of faces in every band 4.2/5.1/9.0 ms vs 17/56/198 landmarks A16W8 0.25 px in the 192 crop 0.5 ms vs 2.8 YOLO26n-seg A16W16 98.2% found, mask IoU 0.994 12.9 ms vs 90 scene model A16W16 98.9% of cells agree 15 ms vs 151 MI-GAN A16W16 41 dB from f32 in the fill 87 ms vs 488 XFeat int8 pano alignment 0.45 px (f32's own spread 0.41) 6.5 ms vs 58 denoiser A16W16 0.00 dB at every ISO 95 ms vs 1510 a tile Face numbers are over public COCO val2017 photographs, not a library. The APK carries the siblings (BUNDLED 15 -> 19; the old int8 detectors removed), about 43 MB more. The Windows installer and its CI count skip them; the Arch and Flatpak packages list their files and never had them. The ladder example takes a role per model, which is how the per-role forms above were seen landing on the NPU from the real probe.
This commit is contained in:
@@ -487,10 +487,11 @@ jobs:
|
|||||||
wine "$SETUP" /S 2>/dev/null
|
wine "$SETUP" /S 2>/dev/null
|
||||||
INST=$(echo "$HOME"/.wine/drive_c/users/*/AppData/Local/Programs/DarkRoom)
|
INST=$(echo "$HOME"/.wine/drive_c/users/*/AppData/Local/Programs/DarkRoom)
|
||||||
ls "$INST"
|
ls "$INST"
|
||||||
# As many files as package.sh stages: everything but the READMEs in
|
# As many files as package.sh stages: everything but the READMEs and
|
||||||
# the directories it copies. A literal here went stale the first
|
# the Hexagon's quantised siblings in the directories it copies. A
|
||||||
# time a model was added.
|
# literal here went stale the first time a model was added.
|
||||||
WANT=$(find models/face models/scene models/inpaint models/denoise -maxdepth 1 -type f ! -name README.md | wc -l)
|
WANT=$(find models/face models/scene models/inpaint models/denoise -maxdepth 1 -type f ! -name README.md \
|
||||||
|
! -name '*.int8.onnx' ! -name '*.a16w8.onnx' ! -name '*.a16w16.onnx' | wc -l)
|
||||||
GOT=$(ls "$INST/models" | wc -l)
|
GOT=$(ls "$INST/models" | wc -l)
|
||||||
[ "$GOT" = "$WANT" ] || { echo "FAIL: expected $WANT model files, installed $GOT"; exit 1; }
|
[ "$GOT" = "$WANT" ] || { echo "FAIL: expected $WANT model files, installed $GOT"; exit 1; }
|
||||||
# The manual, and every picture it shows, counted the same way.
|
# The manual, and every picture it shows, counted the same way.
|
||||||
|
|||||||
@@ -332,27 +332,38 @@ fn unpack_bundled_models(app: &slint::android::AndroidApp) {
|
|||||||
// eyes-open filter has something to read, and a tablet has no other way
|
// eyes-open filter has something to read, and a tablet has no other way
|
||||||
// to get them either.
|
// to get them either.
|
||||||
//
|
//
|
||||||
// The int8 forms beside the three detectors are what the Hexagon runs
|
// The quantised siblings — `.a16w8.onnx`, `.a16w16.onnx` — are what the
|
||||||
// (docs/dev/inference.md §5); the engine loads the sibling when the probe
|
// Hexagon runs (docs/dev/inference.md §1.5), each in the narrowest form
|
||||||
// chose that rung and ignores it otherwise.
|
// that held that model's accuracy on the tablet; the engine loads the
|
||||||
const BUNDLED: [(&std::ffi::CStr, &str); 15] = [
|
// sibling when the probe chose that rung and ignores it otherwise. The
|
||||||
|
// segmenter's and XFeat's forms are compiled into the binary instead,
|
||||||
|
// beside their f32 graphs.
|
||||||
|
const BUNDLED: [(&std::ffi::CStr, &str); 19] = [
|
||||||
(c"models/scrfd_500m_640.onnx", "scrfd_500m_640.onnx"),
|
(c"models/scrfd_500m_640.onnx", "scrfd_500m_640.onnx"),
|
||||||
(
|
(
|
||||||
c"models/scrfd_500m_640.int8.onnx",
|
c"models/scrfd_500m_640.a16w8.onnx",
|
||||||
"scrfd_500m_640.int8.onnx",
|
"scrfd_500m_640.a16w8.onnx",
|
||||||
),
|
),
|
||||||
(c"models/scrfd_2.5g_640.onnx", "scrfd_2.5g_640.onnx"),
|
(c"models/scrfd_2.5g_640.onnx", "scrfd_2.5g_640.onnx"),
|
||||||
(
|
(
|
||||||
c"models/scrfd_2.5g_640.int8.onnx",
|
c"models/scrfd_2.5g_640.a16w8.onnx",
|
||||||
"scrfd_2.5g_640.int8.onnx",
|
"scrfd_2.5g_640.a16w8.onnx",
|
||||||
),
|
),
|
||||||
(c"models/scrfd_10g_640.onnx", "scrfd_10g_640.onnx"),
|
(c"models/scrfd_10g_640.onnx", "scrfd_10g_640.onnx"),
|
||||||
(c"models/scrfd_10g_640.int8.onnx", "scrfd_10g_640.int8.onnx"),
|
(
|
||||||
|
c"models/scrfd_10g_640.a16w8.onnx",
|
||||||
|
"scrfd_10g_640.a16w8.onnx",
|
||||||
|
),
|
||||||
(c"models/arcface_mbf_b1.onnx", "arcface_mbf_b1.onnx"),
|
(c"models/arcface_mbf_b1.onnx", "arcface_mbf_b1.onnx"),
|
||||||
(c"models/2d106det_b1.onnx", "2d106det_b1.onnx"),
|
(c"models/2d106det_b1.onnx", "2d106det_b1.onnx"),
|
||||||
|
(c"models/2d106det_b1.a16w8.onnx", "2d106det_b1.a16w8.onnx"),
|
||||||
(c"models/ocec_s_b1.onnx", "ocec_s_b1.onnx"),
|
(c"models/ocec_s_b1.onnx", "ocec_s_b1.onnx"),
|
||||||
(c"models/sgc_l_48_b1.onnx", "sgc_l_48_b1.onnx"),
|
(c"models/sgc_l_48_b1.onnx", "sgc_l_48_b1.onnx"),
|
||||||
(c"models/yolo26s-sem-ade20k.onnx", "yolo26s-sem-ade20k.onnx"),
|
(c"models/yolo26s-sem-ade20k.onnx", "yolo26s-sem-ade20k.onnx"),
|
||||||
|
(
|
||||||
|
c"models/yolo26s-sem-ade20k.a16w16.onnx",
|
||||||
|
"yolo26s-sem-ade20k.a16w16.onnx",
|
||||||
|
),
|
||||||
(
|
(
|
||||||
c"models/yolo26s-sem-ade20k.classes.json",
|
c"models/yolo26s-sem-ade20k.classes.json",
|
||||||
"yolo26s-sem-ade20k.classes.json",
|
"yolo26s-sem-ade20k.classes.json",
|
||||||
@@ -360,7 +371,9 @@ fn unpack_bundled_models(app: &slint::android::AndroidApp) {
|
|||||||
(c"models/categories.txt", "categories.txt"),
|
(c"models/categories.txt", "categories.txt"),
|
||||||
// The panorama border filler (FR-MRG-4); MIT, 28 MB.
|
// The panorama border filler (FR-MRG-4); MIT, 28 MB.
|
||||||
(c"models/migan-512.onnx", "migan-512.onnx"),
|
(c"models/migan-512.onnx", "migan-512.onnx"),
|
||||||
|
(c"models/migan-512.a16w16.onnx", "migan-512.a16w16.onnx"),
|
||||||
(c"models/mosaic-1408.onnx", "mosaic-1408.onnx"),
|
(c"models/mosaic-1408.onnx", "mosaic-1408.onnx"),
|
||||||
|
(c"models/mosaic-1408.a16w16.onnx", "mosaic-1408.a16w16.onnx"),
|
||||||
];
|
];
|
||||||
|
|
||||||
let dir = dr_ui::shared_face_models_dir();
|
let dir = dr_ui::shared_face_models_dir();
|
||||||
|
|||||||
@@ -5,7 +5,11 @@
|
|||||||
//! returns `rgb`, `1×3×1408×1408` (darkroom-denoise `denoise/export.py`,
|
//! returns `rgb`, `1×3×1408×1408` (darkroom-denoise `denoise/export.py`,
|
||||||
//! fixed shape because every model the engine runs is). The engine picks the
|
//! fixed shape because every model the engine runs is). The engine picks the
|
||||||
//! rung: fp16 on TensorRT and MIGraphX, which measured 0.00 dB from f32; f32
|
//! rung: fp16 on TensorRT and MIGraphX, which measured 0.00 dB from f32; f32
|
||||||
//! on CUDA and the CPU; never the Hexagon, where int8 lost 6–9 dB.
|
//! on CUDA and the CPU; on the Hexagon the `.a16w16.onnx` sibling, 16-bit
|
||||||
|
//! activations and weights, 0.00 dB from f32 on the tablet itself where int8
|
||||||
|
//! lost 5–9 dB (docs/dev/inference.md §1.5). That sibling is the same network
|
||||||
|
//! with the Bayer packing spelled `SpaceToDepth`, which QNN can hold and the
|
||||||
|
//! 6-D reshape it replaces it cannot.
|
||||||
|
|
||||||
use crate::tile::TileNet;
|
use crate::tile::TileNet;
|
||||||
use crate::DenoiseError;
|
use crate::DenoiseError;
|
||||||
|
|||||||
@@ -137,8 +137,8 @@ impl Detection {
|
|||||||
/// A loaded SCRFD graph.
|
/// A loaded SCRFD graph.
|
||||||
pub struct Detector {
|
pub struct Detector {
|
||||||
session: Model,
|
session: Model,
|
||||||
/// f32 or int8 — the int8 form finds a different set of faces and is a
|
/// f32 or a quantised form — which finds a different set of faces and is
|
||||||
/// different detector in `model_id` (docs/dev/inference.md §7).
|
/// a different detector in `model_id` (docs/dev/inference.md §7).
|
||||||
form: Form,
|
form: Form,
|
||||||
/// Feature-map count: 3 for strides {8,16,32}, 4 for {8,16,32,64}.
|
/// Feature-map count: 3 for strides {8,16,32}, 4 for {8,16,32,64}.
|
||||||
///
|
///
|
||||||
@@ -155,7 +155,7 @@ impl Detector {
|
|||||||
}
|
}
|
||||||
|
|
||||||
/// Load the canonical f32 file at `path`, or the form the device's
|
/// Load the canonical f32 file at `path`, or the form the device's
|
||||||
/// backend wants instead — the `.int8.onnx` beside it on a Hexagon —
|
/// backend wants instead — the `.a16w8.onnx` beside it on a Hexagon —
|
||||||
/// which [`Detector::form`] then reports.
|
/// which [`Detector::form`] then reports.
|
||||||
pub fn from_path(path: impl AsRef<std::path::Path>) -> Result<Self, FaceError> {
|
pub fn from_path(path: impl AsRef<std::path::Path>) -> Result<Self, FaceError> {
|
||||||
let (path, form) = dr_inference_engine::resolve_model(Role::Detector, path.as_ref());
|
let (path, form) = dr_inference_engine::resolve_model(Role::Detector, path.as_ref());
|
||||||
|
|||||||
@@ -123,13 +123,22 @@ pub struct Landmarker {
|
|||||||
}
|
}
|
||||||
|
|
||||||
impl Landmarker {
|
impl Landmarker {
|
||||||
|
/// The graph at `path`, or the `.a16w8.onnx` sibling beside it when the
|
||||||
|
/// device's backend runs that (the Hexagon, inference.md §1.5: 0.25 px
|
||||||
|
/// from f32 in the 192 crop, where int8 moved the points by 1.5).
|
||||||
pub fn from_path(path: impl AsRef<std::path::Path>) -> Result<Self, FaceError> {
|
pub fn from_path(path: impl AsRef<std::path::Path>) -> Result<Self, FaceError> {
|
||||||
|
let (path, form) = dr_inference_engine::resolve_model(Role::Landmarks, path.as_ref());
|
||||||
let bytes = std::fs::read(path).map_err(FaceError::ModelRead)?;
|
let bytes = std::fs::read(path).map_err(FaceError::ModelRead)?;
|
||||||
Self::from_bytes(&bytes)
|
Self::from_bytes_in(&bytes, form)
|
||||||
}
|
}
|
||||||
|
|
||||||
pub fn from_bytes(bytes: &[u8]) -> Result<Self, FaceError> {
|
pub fn from_bytes(bytes: &[u8]) -> Result<Self, FaceError> {
|
||||||
let model = dr_inference_engine::open(Role::Landmarks, Form::F32, bytes)?;
|
Self::from_bytes_in(bytes, Form::F32)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// `bytes` in a stated numeric form; the output keeps its meaning.
|
||||||
|
pub fn from_bytes_in(bytes: &[u8], form: Form) -> Result<Self, FaceError> {
|
||||||
|
let model = dr_inference_engine::open(Role::Landmarks, form, bytes)?;
|
||||||
let acquired = model.acquire()?;
|
let acquired = model.acquire()?;
|
||||||
let session = acquired.lock();
|
let session = acquired.lock();
|
||||||
|
|
||||||
|
|||||||
@@ -4,12 +4,14 @@
|
|||||||
//!
|
//!
|
||||||
//! DARKROOM_ORT_DIR=/usr/lib \
|
//! DARKROOM_ORT_DIR=/usr/lib \
|
||||||
//! cargo run --release -p dr-inference-engine --features native,tract \
|
//! cargo run --release -p dr-inference-engine --features native,tract \
|
||||||
//! --example ladder -- CACHE_DIR models/face/scrfd_500m_640.onnx [MODEL.onnx ...]
|
//! --example ladder -- CACHE_DIR models/face/scrfd_500m_640.onnx [ROLE=MODEL.onnx ...]
|
||||||
//!
|
//!
|
||||||
//! Every model named is a `Detector` for the config's purposes, which is
|
//! A bare path is a `Detector`; `denoiser=…`, `scene=…`, `inpainter=…`,
|
||||||
//! enough to see the rung taken, the engines compiled and a session land
|
//! `landmarks=…` (any `Role`, lower case) says otherwise, so a device can
|
||||||
//! on it. Delete `CACHE_DIR` to see the first run again; keep it to see the
|
//! show each role taking its own form (inference.md §1.5). Each is opened
|
||||||
//! second.
|
//! through `resolve_model`, as the app opens it, and the line says which
|
||||||
|
//! form and which rung it landed on. Delete `CACHE_DIR` to see the first
|
||||||
|
//! run again; keep it to see the second.
|
||||||
|
|
||||||
use std::path::PathBuf;
|
use std::path::PathBuf;
|
||||||
use std::time::{Duration, Instant};
|
use std::time::{Duration, Instant};
|
||||||
@@ -17,7 +19,10 @@ use std::time::{Duration, Instant};
|
|||||||
fn main() {
|
fn main() {
|
||||||
env_logger::Builder::from_env(env_logger::Env::default().default_filter_or("info")).init();
|
env_logger::Builder::from_env(env_logger::Env::default().default_filter_or("info")).init();
|
||||||
let mut args = std::env::args_os().skip(1).map(PathBuf::from);
|
let mut args = std::env::args_os().skip(1).map(PathBuf::from);
|
||||||
let (Some(cache_dir), models) = (args.next(), args.collect::<Vec<_>>()) else {
|
let (Some(cache_dir), models) = (
|
||||||
|
args.next(),
|
||||||
|
args.map(|a| role_and_path(&a)).collect::<Vec<_>>(),
|
||||||
|
) else {
|
||||||
eprintln!("usage: ladder CACHE_DIR MODEL.onnx [MODEL.onnx ...]");
|
eprintln!("usage: ladder CACHE_DIR MODEL.onnx [MODEL.onnx ...]");
|
||||||
std::process::exit(2);
|
std::process::exit(2);
|
||||||
};
|
};
|
||||||
@@ -34,10 +39,7 @@ fn main() {
|
|||||||
dr_inference_engine::init(dr_inference_engine::Config {
|
dr_inference_engine::init(dr_inference_engine::Config {
|
||||||
runtime_dirs,
|
runtime_dirs,
|
||||||
cache_dir: cache_dir.clone(),
|
cache_dir: cache_dir.clone(),
|
||||||
models: models
|
models: models.clone(),
|
||||||
.iter()
|
|
||||||
.map(|p| (dr_inference_engine::Role::Detector, p.clone()))
|
|
||||||
.collect(),
|
|
||||||
embedded: Vec::new(),
|
embedded: Vec::new(),
|
||||||
ceiling: None,
|
ceiling: None,
|
||||||
threads: 0,
|
threads: 0,
|
||||||
@@ -80,21 +82,39 @@ fn main() {
|
|||||||
std::thread::sleep(Duration::from_millis(500));
|
std::thread::sleep(Duration::from_millis(500));
|
||||||
}
|
}
|
||||||
|
|
||||||
for path in &models {
|
for (role, path) in &models {
|
||||||
let bytes = std::fs::read(path).expect("read model");
|
let (path, form) = dr_inference_engine::resolve_model(*role, path);
|
||||||
|
let bytes = std::fs::read(&path).expect("read model");
|
||||||
let t = Instant::now();
|
let t = Instant::now();
|
||||||
let model = dr_inference_engine::open(
|
let model = dr_inference_engine::open(*role, form, &bytes).expect("open model");
|
||||||
dr_inference_engine::Role::Detector,
|
|
||||||
dr_inference_engine::Form::F32,
|
|
||||||
&bytes,
|
|
||||||
)
|
|
||||||
.expect("open model");
|
|
||||||
let acquired = model.acquire().expect("acquire session");
|
let acquired = model.acquire().expect("acquire session");
|
||||||
println!(
|
println!(
|
||||||
"{} on {} in {:.2} s",
|
"{role:?}: {} ({form:?}) on {} in {:.2} s",
|
||||||
path.file_name().unwrap().to_string_lossy(),
|
path.file_name().unwrap().to_string_lossy(),
|
||||||
acquired.rung().label(),
|
acquired.rung().label(),
|
||||||
t.elapsed().as_secs_f64()
|
t.elapsed().as_secs_f64()
|
||||||
);
|
);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// `denoiser=path` → (Denoiser, path); a bare path is a detector.
|
||||||
|
fn role_and_path(arg: &std::path::Path) -> (dr_inference_engine::Role, PathBuf) {
|
||||||
|
use dr_inference_engine::Role::*;
|
||||||
|
let s = arg.to_string_lossy();
|
||||||
|
let Some((name, path)) = s.split_once('=') else {
|
||||||
|
return (Detector, arg.to_path_buf());
|
||||||
|
};
|
||||||
|
let role = match name {
|
||||||
|
"detector" => Detector,
|
||||||
|
"embedder" => Embedder,
|
||||||
|
"segmenter" => Segmenter,
|
||||||
|
"scene" => Scene,
|
||||||
|
"landmarks" => Landmarks,
|
||||||
|
"eyes" => EyeClassifier,
|
||||||
|
"keypoints" => Keypoints,
|
||||||
|
"inpainter" => Inpainter,
|
||||||
|
"denoiser" => Denoiser,
|
||||||
|
other => panic!("no role {other:?}"),
|
||||||
|
};
|
||||||
|
(role, PathBuf::from(path))
|
||||||
|
}
|
||||||
|
|||||||
@@ -8,7 +8,7 @@
|
|||||||
|
|
||||||
use std::path::PathBuf;
|
use std::path::PathBuf;
|
||||||
|
|
||||||
use crate::{state, Config, Form, Rung};
|
use crate::{state, Config, Rung};
|
||||||
|
|
||||||
enum Source {
|
enum Source {
|
||||||
File(PathBuf),
|
File(PathBuf),
|
||||||
@@ -82,10 +82,10 @@ pub fn run() {
|
|||||||
(*role, Source::File(path), size)
|
(*role, Source::File(path), size)
|
||||||
})
|
})
|
||||||
})
|
})
|
||||||
.chain(cfg.embedded.iter().filter_map(|(role, bytes)| {
|
.chain(cfg.embedded.iter().filter_map(|(role, form, bytes)| {
|
||||||
// An embedded model has no int8 sibling to offer a rung that
|
// The embedded form the rung wants, if the build carries it;
|
||||||
// wants one; it runs on that rung's fallback.
|
// a build without it runs that model on the rung's fallback.
|
||||||
(rung.serves(*role) && rung.form(*role) == Form::F32).then_some((
|
(rung.serves(*role) && rung.form(*role) == *form).then_some((
|
||||||
*role,
|
*role,
|
||||||
Source::Bytes(bytes),
|
Source::Bytes(bytes),
|
||||||
bytes.len() as u64,
|
bytes.len() as u64,
|
||||||
|
|||||||
@@ -42,25 +42,46 @@ pub enum Role {
|
|||||||
/// XFeat, the panorama keypoint detector (docs/dev/panorama.md).
|
/// XFeat, the panorama keypoint detector (docs/dev/panorama.md).
|
||||||
Keypoints,
|
Keypoints,
|
||||||
/// MI-GAN, the panorama border filler (docs/dev/panorama.md §12). Plain
|
/// MI-GAN, the panorama border filler (docs/dev/panorama.md §12). Plain
|
||||||
/// convolutions, so any rung serves it; fp16 on TensorRT and int8 on
|
/// convolutions, so any rung serves it; fp16 on TensorRT and 16-bit
|
||||||
/// the Hexagon are the point of it.
|
/// activations on the Hexagon (int8 changes the fill, §1.5).
|
||||||
Inpainter,
|
Inpainter,
|
||||||
/// The learned demosaic and denoise on the raw mosaic (docs/dev/denoise.md).
|
/// The learned demosaic and denoise on the raw mosaic (docs/dev/denoise.md).
|
||||||
/// fp16 costs it nothing measurable; int8 costs 6–9 dB, because 256
|
/// fp16 costs it nothing measurable; int8 costs 6–9 dB, because 256
|
||||||
/// levels cannot hold the shadow steps it exists to recover — so the
|
/// levels cannot hold the shadow steps it exists to recover — so the
|
||||||
/// Hexagon does not take it.
|
/// Hexagon takes it with 16-bit activations and weights (§1.5).
|
||||||
Denoiser,
|
Denoiser,
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Which numeric form of a model a session was built from.
|
/// Which numeric form of a model a session was built from.
|
||||||
///
|
///
|
||||||
/// `Int8` is a different network from `F32` for a detector — it finds a
|
/// The quantised forms are QDQ graphs, per-channel weights, as QNN's HTP
|
||||||
/// different set of faces — which is why [`form_suffix`] exists and why a
|
/// takes them (docs/dev/inference.md §1.5): `Int8` is 8-bit activations and
|
||||||
|
/// weights, `A16W8` 16-bit activations with 8-bit weights, `A16W16` 16-bit
|
||||||
|
/// both. The Hexagon accepts no float tensor at all, so these are the
|
||||||
|
/// whole menu; which one a role gets is [`Rung::form`], measured per model.
|
||||||
|
///
|
||||||
|
/// A quantised detector is a different network from the f32 one — it finds
|
||||||
|
/// a different set of faces — which is why [`form_suffix`] exists and why a
|
||||||
/// caller appends it to `model_id`.
|
/// caller appends it to `model_id`.
|
||||||
#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash, Serialize, Deserialize)]
|
#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash, Serialize, Deserialize)]
|
||||||
pub enum Form {
|
pub enum Form {
|
||||||
F32,
|
F32,
|
||||||
Int8,
|
Int8,
|
||||||
|
A16W8,
|
||||||
|
A16W16,
|
||||||
|
}
|
||||||
|
|
||||||
|
impl Form {
|
||||||
|
/// The infix of the sibling file that holds this form:
|
||||||
|
/// `scrfd_500m_640.a16w8.onnx` beside `scrfd_500m_640.onnx`.
|
||||||
|
pub fn file_tag(self) -> Option<&'static str> {
|
||||||
|
match self {
|
||||||
|
Form::F32 => None,
|
||||||
|
Form::Int8 => Some("int8"),
|
||||||
|
Form::A16W8 => Some("a16w8"),
|
||||||
|
Form::A16W16 => Some("a16w16"),
|
||||||
|
}
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
/// A rung of the ladder (§2). Ordered: a user override names the highest rung
|
/// A rung of the ladder (§2). Ordered: a user override names the highest rung
|
||||||
@@ -81,7 +102,7 @@ pub enum Rung {
|
|||||||
/// removed in ONNX Runtime 1.23, so there is no non-compiling AMD rung
|
/// removed in ONNX Runtime 1.23, so there is no non-compiling AMD rung
|
||||||
/// to fall back to: this one falls back to the CPU.
|
/// to fall back to: this one falls back to the CPU.
|
||||||
MiGraphX,
|
MiGraphX,
|
||||||
/// Qualcomm's Hexagon NPU through QNN, int8 models only. Android only.
|
/// Qualcomm's Hexagon NPU through QNN, quantised models only. Android only.
|
||||||
Hexagon,
|
Hexagon,
|
||||||
/// Apple, through CoreML: the Neural Engine, the GPU or the CPU, as
|
/// Apple, through CoreML: the Neural Engine, the GPU or the CPU, as
|
||||||
/// CoreML schedules it. macOS only. Compiles an ML Program per model on
|
/// CoreML schedules it. macOS only. Compiles an ML Program per model on
|
||||||
@@ -121,23 +142,38 @@ impl Rung {
|
|||||||
}
|
}
|
||||||
|
|
||||||
/// The model form this rung wants for a role.
|
/// The model form this rung wants for a role.
|
||||||
fn form(self, _role: Role) -> Form {
|
///
|
||||||
|
/// On the Hexagon, the narrowest form that held each model's accuracy
|
||||||
|
/// on the tablet itself (§1.5): int8 lost 5% of the detector's faces at
|
||||||
|
/// 40–80 px, moved the landmarks by 1.5 px and the segmenter's scores
|
||||||
|
/// to nothing, and the denoiser by 6–9 dB, so those take 16-bit
|
||||||
|
/// activations; the segmenter, scene model, filler and denoiser also
|
||||||
|
/// needed 16-bit weights. Only XFeat keeps int8: its panorama alignment
|
||||||
|
/// moved by no more than f32's own refits do.
|
||||||
|
pub fn form(self, role: Role) -> Form {
|
||||||
match self {
|
match self {
|
||||||
Rung::Hexagon => Form::Int8,
|
Rung::Hexagon => match role {
|
||||||
|
Role::Keypoints => Form::Int8,
|
||||||
|
Role::Detector | Role::Landmarks => Form::A16W8,
|
||||||
|
Role::Segmenter | Role::Scene | Role::Inpainter | Role::Denoiser => Form::A16W16,
|
||||||
|
Role::Embedder | Role::EyeClassifier => Form::F32,
|
||||||
|
},
|
||||||
_ => Form::F32,
|
_ => Form::F32,
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Whether this rung runs `role` at all. The Hexagon takes int8 graphs
|
/// Whether this rung runs `role` at all. The Hexagon takes quantised
|
||||||
/// only, and the embedder is never int8 (§7) — it runs on the CPU
|
/// graphs only, and the embedder is never quantised (§7) — it runs on
|
||||||
/// beside a detector on the NPU, so its vectors compare across devices.
|
/// the CPU beside a detector on the NPU, so its vectors compare across
|
||||||
/// Nor is the denoiser: its int8 form failed the 0.5 dB gate by 6–9 dB
|
/// devices; at A16W16 it still missed the 0.999 cosine gate. The eye
|
||||||
/// (denoise.md §8), so it runs on the CPU there too. CoreML is kept off
|
/// classifiers stay on the CPU too: a millisecond there, and the two
|
||||||
/// the embedder for the same reason as the Hexagon: the Neural Engine is
|
/// share one role while only one of them held its readings quantised.
|
||||||
/// fp16, and which unit runs a graph is CoreML's choice.
|
/// CoreML is kept off the embedder for the same reason as the Hexagon:
|
||||||
|
/// the Neural Engine is fp16, and which unit runs a graph is CoreML's
|
||||||
|
/// choice.
|
||||||
fn serves(self, role: Role) -> bool {
|
fn serves(self, role: Role) -> bool {
|
||||||
match self {
|
match self {
|
||||||
Rung::Hexagon => !matches!(role, Role::Embedder | Role::Denoiser),
|
Rung::Hexagon => !matches!(role, Role::Embedder | Role::EyeClassifier),
|
||||||
Rung::CoreMl => role != Role::Embedder,
|
Rung::CoreMl => role != Role::Embedder,
|
||||||
_ => true,
|
_ => true,
|
||||||
}
|
}
|
||||||
@@ -161,8 +197,10 @@ pub struct Config {
|
|||||||
/// The canonical model files on this device, so engines can be compiled
|
/// The canonical model files on this device, so engines can be compiled
|
||||||
/// ahead of the first request for them.
|
/// ahead of the first request for them.
|
||||||
pub models: Vec<(Role, PathBuf)>,
|
pub models: Vec<(Role, PathBuf)>,
|
||||||
/// Models compiled into the binary, for the same reason.
|
/// Models compiled into the binary, for the same reason, each with the
|
||||||
pub embedded: Vec<(Role, &'static [u8])>,
|
/// form it is. A build that embeds a quantised sibling lists it here
|
||||||
|
/// beside the f32 graph, and the compile step takes the one the rung wants.
|
||||||
|
pub embedded: Vec<(Role, Form, &'static [u8])>,
|
||||||
/// The highest rung the user allows; `None` is "the best that works".
|
/// The highest rung the user allows; `None` is "the best that works".
|
||||||
pub ceiling: Option<Rung>,
|
pub ceiling: Option<Rung>,
|
||||||
/// ONNX Runtime's intra-op pool; 0 picks from the core count.
|
/// ONNX Runtime's intra-op pool; 0 picks from the core count.
|
||||||
@@ -188,10 +226,10 @@ pub struct Status {
|
|||||||
}
|
}
|
||||||
|
|
||||||
impl Status {
|
impl Status {
|
||||||
/// "Hexagon NPU · int8 · ONNX Runtime 1.29" — the settings row's text.
|
/// "Hexagon NPU · quantised · ONNX Runtime 1.29" — the settings row's text.
|
||||||
pub fn line(&self) -> String {
|
pub fn line(&self) -> String {
|
||||||
let form = match self.rung {
|
let form = match self.rung {
|
||||||
Rung::Hexagon => " · int8",
|
Rung::Hexagon => " · quantised",
|
||||||
Rung::TensorRt | Rung::MiGraphX => " · fp16",
|
Rung::TensorRt | Rung::MiGraphX => " · fp16",
|
||||||
_ => "",
|
_ => "",
|
||||||
};
|
};
|
||||||
@@ -464,26 +502,44 @@ fn current_rung(s: &State) -> Rung {
|
|||||||
|
|
||||||
/// The file to load for `role` under the current selection, and its form.
|
/// The file to load for `role` under the current selection, and its form.
|
||||||
///
|
///
|
||||||
/// A rung that wants int8 gets the `.int8.onnx` sibling of the canonical file
|
/// A rung that wants a quantised form gets that sibling of the canonical
|
||||||
/// if it exists; otherwise the canonical file, on the rung's fallback. A
|
/// file (`<stem>.a16w8.onnx` and so on, [`Form::file_tag`]) if it exists;
|
||||||
/// caller adds [`form_suffix`] to the `model_id` it records.
|
/// otherwise the canonical file, on the rung's fallback. A caller adds
|
||||||
|
/// [`form_suffix`] to the `model_id` it records.
|
||||||
pub fn resolve_model(role: Role, canonical: &Path) -> (PathBuf, Form) {
|
pub fn resolve_model(role: Role, canonical: &Path) -> (PathBuf, Form) {
|
||||||
let rung = current_rung(&state().lock().unwrap());
|
let rung = current_rung(&state().lock().unwrap());
|
||||||
if rung.serves(role) && rung.form(role) == Form::Int8 {
|
let want = rung.form(role);
|
||||||
let sibling = int8_sibling(canonical);
|
if rung.serves(role) && want != Form::F32 {
|
||||||
|
let sibling = form_sibling(canonical, want);
|
||||||
if sibling.is_file() {
|
if sibling.is_file() {
|
||||||
return (sibling, Form::Int8);
|
return (sibling, want);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
(canonical.to_path_buf(), Form::F32)
|
(canonical.to_path_buf(), Form::F32)
|
||||||
}
|
}
|
||||||
|
|
||||||
fn int8_sibling(canonical: &Path) -> PathBuf {
|
/// The same choice for a model compiled into the binary: of the forms
|
||||||
|
/// `offered`, the one the current rung wants for `role`, else the f32 one.
|
||||||
|
/// `offered` must hold an `F32` entry.
|
||||||
|
pub fn choose_embedded(role: Role, offered: &[(Form, &'static [u8])]) -> (&'static [u8], Form) {
|
||||||
|
let rung = current_rung(&state().lock().unwrap());
|
||||||
|
let want = rung.form(role);
|
||||||
|
let pick = |form| offered.iter().find(|(f, _)| *f == form);
|
||||||
|
let (form, bytes) = (rung.serves(role).then(|| pick(want)).flatten())
|
||||||
|
.or_else(|| pick(Form::F32))
|
||||||
|
.expect("an embedded model offers its f32 form");
|
||||||
|
(bytes, *form)
|
||||||
|
}
|
||||||
|
|
||||||
|
fn form_sibling(canonical: &Path, form: Form) -> PathBuf {
|
||||||
let stem = canonical
|
let stem = canonical
|
||||||
.file_stem()
|
.file_stem()
|
||||||
.map(|s| s.to_string_lossy().into_owned())
|
.map(|s| s.to_string_lossy().into_owned())
|
||||||
.unwrap_or_default();
|
.unwrap_or_default();
|
||||||
canonical.with_file_name(format!("{stem}.int8.onnx"))
|
match form.file_tag() {
|
||||||
|
Some(tag) => canonical.with_file_name(format!("{stem}.{tag}.onnx")),
|
||||||
|
None => canonical.to_path_buf(),
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
/// What a form appends to a detector's `model_id` (§7).
|
/// What a form appends to a detector's `model_id` (§7).
|
||||||
@@ -491,6 +547,8 @@ pub fn form_suffix(form: Form) -> &'static str {
|
|||||||
match form {
|
match form {
|
||||||
Form::F32 => "",
|
Form::F32 => "",
|
||||||
Form::Int8 => "_i8",
|
Form::Int8 => "_i8",
|
||||||
|
Form::A16W8 => "_a16",
|
||||||
|
Form::A16W16 => "_a16w16",
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -520,8 +578,8 @@ pub fn open(role: Role, form: Form, bytes: &[u8]) -> Result<Model, Error> {
|
|||||||
fn effective_rung(s: &State, selected: Rung, role: Role, form: Form, hash: u64) -> Rung {
|
fn effective_rung(s: &State, selected: Rung, role: Role, form: Form, hash: u64) -> Rung {
|
||||||
let mut rung = selected;
|
let mut rung = selected;
|
||||||
if !rung.serves(role) || rung.form(role) != form {
|
if !rung.serves(role) || rung.form(role) != form {
|
||||||
// The embedder on a Hexagon device, or an f32 detector where the int8
|
// The embedder on a Hexagon device, or an f32 detector where the
|
||||||
// sibling was missing: neither can go to the NPU.
|
// quantised sibling was missing: neither can go to the NPU.
|
||||||
rung = rung.fallback();
|
rung = rung.fallback();
|
||||||
}
|
}
|
||||||
if rung.compiles() && !s.cache.compiled.contains(&engines::key_of(rung, hash)) {
|
if rung.compiles() && !s.cache.compiled.contains(&engines::key_of(rung, hash)) {
|
||||||
@@ -612,9 +670,10 @@ mod tests {
|
|||||||
#[test]
|
#[test]
|
||||||
fn the_hexagon_never_takes_the_embedder() {
|
fn the_hexagon_never_takes_the_embedder() {
|
||||||
assert!(!Rung::Hexagon.serves(Role::Embedder));
|
assert!(!Rung::Hexagon.serves(Role::Embedder));
|
||||||
assert!(!Rung::Hexagon.serves(Role::Denoiser));
|
assert!(!Rung::Hexagon.serves(Role::EyeClassifier));
|
||||||
assert!(Rung::Hexagon.serves(Role::Detector));
|
assert!(Rung::Hexagon.serves(Role::Detector));
|
||||||
assert_eq!(Rung::Hexagon.form(Role::Detector), Form::Int8);
|
assert!(Rung::Hexagon.serves(Role::Denoiser));
|
||||||
|
assert_eq!(Rung::Hexagon.form(Role::Detector), Form::A16W8);
|
||||||
// A detector offered in f32 on a Hexagon device lands on the CPU.
|
// A detector offered in f32 on a Hexagon device lands on the CPU.
|
||||||
let s = State {
|
let s = State {
|
||||||
config: Config::default(),
|
config: Config::default(),
|
||||||
@@ -625,36 +684,57 @@ mod tests {
|
|||||||
probing: false,
|
probing: false,
|
||||||
wanted: 0,
|
wanted: 0,
|
||||||
};
|
};
|
||||||
|
let on = |role, form| effective_rung(&s, Rung::Hexagon, role, form, engines::hash(b""));
|
||||||
|
assert_eq!(on(Role::Embedder, Form::F32), Rung::Cpu);
|
||||||
|
assert_eq!(on(Role::Detector, Form::F32), Rung::Cpu);
|
||||||
|
// A form other than the one the role wants is not the NPU's either:
|
||||||
|
// an int8 detector left over from an older install stays off it.
|
||||||
|
assert_eq!(on(Role::Detector, Form::Int8), Rung::Cpu);
|
||||||
|
// The wanted form whose context is not compiled yet: also the CPU.
|
||||||
|
assert_eq!(on(Role::Detector, Form::A16W8), Rung::Cpu);
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The form each role gets on the Hexagon is the one measured to hold
|
||||||
|
/// its accuracy there (§1.5); a change to this table is a change to
|
||||||
|
/// what the tablet computes, and must come with a measurement.
|
||||||
|
#[test]
|
||||||
|
fn each_role_has_its_measured_form_on_the_hexagon() {
|
||||||
|
use Form::*;
|
||||||
|
for (role, form) in [
|
||||||
|
(Role::Detector, A16W8),
|
||||||
|
(Role::Landmarks, A16W8),
|
||||||
|
(Role::Segmenter, A16W16),
|
||||||
|
(Role::Scene, A16W16),
|
||||||
|
(Role::Inpainter, A16W16),
|
||||||
|
(Role::Denoiser, A16W16),
|
||||||
|
(Role::Keypoints, Int8),
|
||||||
|
(Role::Embedder, F32),
|
||||||
|
(Role::EyeClassifier, F32),
|
||||||
|
] {
|
||||||
|
assert_eq!(Rung::Hexagon.form(role), form, "{role:?}");
|
||||||
|
}
|
||||||
|
for rung in [
|
||||||
|
Rung::Cpu,
|
||||||
|
Rung::Cuda,
|
||||||
|
Rung::TensorRt,
|
||||||
|
Rung::MiGraphX,
|
||||||
|
Rung::CoreMl,
|
||||||
|
] {
|
||||||
|
assert_eq!(rung.form(Role::Detector), F32);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn a_form_lives_in_its_tagged_sibling() {
|
||||||
|
let canonical = Path::new("/m/scrfd_500m_640.onnx");
|
||||||
|
assert_eq!(form_sibling(canonical, Form::F32), canonical);
|
||||||
assert_eq!(
|
assert_eq!(
|
||||||
effective_rung(
|
form_sibling(canonical, Form::A16W8),
|
||||||
&s,
|
Path::new("/m/scrfd_500m_640.a16w8.onnx")
|
||||||
Rung::Hexagon,
|
|
||||||
Role::Embedder,
|
|
||||||
Form::F32,
|
|
||||||
engines::hash(b"")
|
|
||||||
),
|
|
||||||
Rung::Cpu
|
|
||||||
);
|
);
|
||||||
assert_eq!(
|
assert_eq!(
|
||||||
effective_rung(
|
form_sibling(canonical, Form::Int8),
|
||||||
&s,
|
Path::new("/m/scrfd_500m_640.int8.onnx")
|
||||||
Rung::Hexagon,
|
|
||||||
Role::Detector,
|
|
||||||
Form::F32,
|
|
||||||
engines::hash(b"")
|
|
||||||
),
|
|
||||||
Rung::Cpu
|
|
||||||
);
|
|
||||||
// An int8 detector whose context is not compiled yet: also the CPU.
|
|
||||||
assert_eq!(
|
|
||||||
effective_rung(
|
|
||||||
&s,
|
|
||||||
Rung::Hexagon,
|
|
||||||
Role::Detector,
|
|
||||||
Form::Int8,
|
|
||||||
engines::hash(b"")
|
|
||||||
),
|
|
||||||
Rung::Cpu
|
|
||||||
);
|
);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -174,10 +174,9 @@ pub fn attempt<T>(cfg: &Config, what: &str, f: impl FnOnce() -> T) -> Result<T,
|
|||||||
}
|
}
|
||||||
|
|
||||||
/// The smallest detector, or the smallest model of any role if there is
|
/// The smallest detector, or the smallest model of any role if there is
|
||||||
/// none. A ~2 MB detector is the cheapest real test of a provider, and the
|
/// none. A ~2 MB detector is the cheapest real test of a provider, and
|
||||||
/// detector is the role the int8 forms exist for — the eye classifiers are
|
/// every rung serves it — the eye classifiers are smaller still, but the
|
||||||
/// smaller still, and a Hexagon probed with one would fail for want of a
|
/// Hexagon does not take them, and a probe with one would fail it for that.
|
||||||
/// form nobody ships.
|
|
||||||
fn probe_model(cfg: &Config) -> Option<(Role, PathBuf)> {
|
fn probe_model(cfg: &Config) -> Option<(Role, PathBuf)> {
|
||||||
let smallest = |want: Option<Role>| {
|
let smallest = |want: Option<Role>| {
|
||||||
cfg.models
|
cfg.models
|
||||||
@@ -203,16 +202,14 @@ fn time_rung(
|
|||||||
cfg: &Config,
|
cfg: &Config,
|
||||||
) -> Result<(f64, Option<String>), String> {
|
) -> Result<(f64, Option<String>), String> {
|
||||||
let want = rung.form(role);
|
let want = rung.form(role);
|
||||||
let path = match want {
|
let path = crate::form_sibling(canonical, want);
|
||||||
Form::Int8 => {
|
if want != Form::F32 && !path.is_file() {
|
||||||
let p = crate::int8_sibling(canonical);
|
return Err(format!(
|
||||||
if !p.is_file() {
|
"no {} form of {}",
|
||||||
return Err(format!("no int8 form of {}", canonical.display()));
|
want.file_tag().unwrap_or("f32"),
|
||||||
}
|
canonical.display()
|
||||||
p
|
));
|
||||||
}
|
}
|
||||||
Form::F32 => canonical.to_path_buf(),
|
|
||||||
};
|
|
||||||
let bytes = std::fs::read(&path).map_err(|e| e.to_string())?;
|
let bytes = std::fs::read(&path).map_err(|e| e.to_string())?;
|
||||||
let started = Instant::now();
|
let started = Instant::now();
|
||||||
let mut session =
|
let mut session =
|
||||||
@@ -295,9 +292,9 @@ fn fingerprint(runtime: &Runtime, cfg: &Config) -> String {
|
|||||||
},
|
},
|
||||||
device_identity(),
|
device_identity(),
|
||||||
];
|
];
|
||||||
for (role, bytes) in &cfg.embedded {
|
for (role, form, bytes) in &cfg.embedded {
|
||||||
parts.push(format!(
|
parts.push(format!(
|
||||||
"{role:?} embedded {:016x}",
|
"{role:?} embedded {form:?} {:016x}",
|
||||||
crate::engines::hash(bytes)
|
crate::engines::hash(bytes)
|
||||||
));
|
));
|
||||||
}
|
}
|
||||||
@@ -306,9 +303,13 @@ fn fingerprint(runtime: &Runtime, cfg: &Config) -> String {
|
|||||||
.map(|b| crate::engines::hash(&b))
|
.map(|b| crate::engines::hash(&b))
|
||||||
.unwrap_or(0);
|
.unwrap_or(0);
|
||||||
parts.push(format!("{role:?} {hash:016x}"));
|
parts.push(format!("{role:?} {hash:016x}"));
|
||||||
let int8 = crate::int8_sibling(path);
|
for form in [Form::Int8, Form::A16W8, Form::A16W16] {
|
||||||
if let Ok(b) = std::fs::read(&int8) {
|
if let Ok(b) = std::fs::read(crate::form_sibling(path, form)) {
|
||||||
parts.push(format!("{role:?} int8 {:016x}", crate::engines::hash(&b)));
|
parts.push(format!(
|
||||||
|
"{role:?} {form:?} {:016x}",
|
||||||
|
crate::engines::hash(&b)
|
||||||
|
));
|
||||||
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
parts.join("\n")
|
parts.join("\n")
|
||||||
|
|||||||
+12
-1
@@ -13,8 +13,13 @@ const MODELS: &[&str] = &[
|
|||||||
"../../models/keypoints/xfeat-768.onnx",
|
"../../models/keypoints/xfeat-768.onnx",
|
||||||
];
|
];
|
||||||
|
|
||||||
|
const QUANTISED: &[&str] = &[
|
||||||
|
"../../models/keypoints/xfeat-1024.int8.onnx",
|
||||||
|
"../../models/keypoints/xfeat-768.int8.onnx",
|
||||||
|
];
|
||||||
|
|
||||||
fn main() {
|
fn main() {
|
||||||
for m in MODELS {
|
for m in MODELS.iter().chain(QUANTISED) {
|
||||||
println!("cargo:rerun-if-changed={m}");
|
println!("cargo:rerun-if-changed={m}");
|
||||||
}
|
}
|
||||||
println!("cargo:rerun-if-changed=build.rs");
|
println!("cargo:rerun-if-changed=build.rs");
|
||||||
@@ -26,6 +31,12 @@ fn main() {
|
|||||||
for model in MODELS.iter().copied() {
|
for model in MODELS.iter().copied() {
|
||||||
check(model);
|
check(model);
|
||||||
}
|
}
|
||||||
|
// The Hexagon's int8 forms ride only in an Android build.
|
||||||
|
if std::env::var("CARGO_CFG_TARGET_OS").as_deref() == Ok("android") {
|
||||||
|
for model in QUANTISED.iter().copied() {
|
||||||
|
check(model);
|
||||||
|
}
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
fn check(model: &str) {
|
fn check(model: &str) {
|
||||||
|
|||||||
@@ -30,7 +30,8 @@ pub struct MiGan {
|
|||||||
|
|
||||||
impl MiGan {
|
impl MiGan {
|
||||||
/// From the model file, in whichever form the engine's rung wants
|
/// From the model file, in whichever form the engine's rung wants
|
||||||
/// (`resolve_model` picks an int8 sibling for the Hexagon).
|
/// (`resolve_model` picks the `.a16w16.onnx` sibling on the Hexagon:
|
||||||
|
/// int8 moved the fill 16 dB from f32's, 16-bit about 41).
|
||||||
pub fn from_path(path: &std::path::Path) -> Result<Self, PanoError> {
|
pub fn from_path(path: &std::path::Path) -> Result<Self, PanoError> {
|
||||||
use dr_inference_engine::{resolve_model, Role};
|
use dr_inference_engine::{resolve_model, Role};
|
||||||
let (path, form) = resolve_model(Role::Inpainter, path);
|
let (path, form) = resolve_model(Role::Inpainter, path);
|
||||||
|
|||||||
@@ -36,18 +36,49 @@ pub struct XFeat {
|
|||||||
pub options: DecodeOptions,
|
pub options: DecodeOptions,
|
||||||
}
|
}
|
||||||
|
|
||||||
/// The bytes of both exports compiled into the binary, for whoever compiles
|
/// The Hexagon's forms (docs/dev/inference.md §1.5): int8, from the same
|
||||||
/// engines ahead of the first request (docs/dev/inference.md §6).
|
/// network spelled for the HTP (the unfold as SpaceToDepth, the bilinear
|
||||||
|
/// resizes as matrix products). Only Android has a Hexagon.
|
||||||
|
#[cfg(all(feature = "embedded-model", target_os = "android"))]
|
||||||
|
const EMBEDDED_LANDSCAPE_INT8: &[u8] =
|
||||||
|
include_bytes!("../../../models/keypoints/xfeat-1024.int8.onnx");
|
||||||
|
#[cfg(all(feature = "embedded-model", target_os = "android"))]
|
||||||
|
const EMBEDDED_PORTRAIT_INT8: &[u8] =
|
||||||
|
include_bytes!("../../../models/keypoints/xfeat-768.int8.onnx");
|
||||||
|
|
||||||
|
/// Every form of both exports compiled into the binary, landscape then
|
||||||
|
/// portrait, for whoever compiles engines ahead of the first request
|
||||||
|
/// (docs/dev/inference.md §6).
|
||||||
#[cfg(feature = "embedded-model")]
|
#[cfg(feature = "embedded-model")]
|
||||||
pub fn embedded_model_bytes() -> [&'static [u8]; 2] {
|
pub fn embedded_models() -> [Vec<(dr_inference_engine::Form, &'static [u8])>; 2] {
|
||||||
[EMBEDDED_LANDSCAPE, EMBEDDED_PORTRAIT]
|
use dr_inference_engine::Form;
|
||||||
|
#[allow(unused_mut)]
|
||||||
|
let mut forms = [
|
||||||
|
vec![(Form::F32, EMBEDDED_LANDSCAPE)],
|
||||||
|
vec![(Form::F32, EMBEDDED_PORTRAIT)],
|
||||||
|
];
|
||||||
|
#[cfg(target_os = "android")]
|
||||||
|
{
|
||||||
|
forms[0].push((Form::Int8, EMBEDDED_LANDSCAPE_INT8));
|
||||||
|
forms[1].push((Form::Int8, EMBEDDED_PORTRAIT_INT8));
|
||||||
|
}
|
||||||
|
forms
|
||||||
}
|
}
|
||||||
|
|
||||||
impl XFeat {
|
impl XFeat {
|
||||||
/// The weights compiled into the binary.
|
/// The weights compiled into the binary, in the form the device's
|
||||||
|
/// backend runs.
|
||||||
#[cfg(feature = "embedded-model")]
|
#[cfg(feature = "embedded-model")]
|
||||||
pub fn embedded() -> Result<Self, PanoError> {
|
pub fn embedded() -> Result<Self, PanoError> {
|
||||||
Self::from_bytes(EMBEDDED_LANDSCAPE, EMBEDDED_PORTRAIT)
|
use dr_inference_engine::{choose_embedded, open, Role};
|
||||||
|
let [l, p] = embedded_models();
|
||||||
|
let (l, lf) = choose_embedded(Role::Keypoints, &l);
|
||||||
|
let (p, pf) = choose_embedded(Role::Keypoints, &p);
|
||||||
|
Ok(XFeat {
|
||||||
|
landscape: open(Role::Keypoints, lf, l)?,
|
||||||
|
portrait: open(Role::Keypoints, pf, p)?,
|
||||||
|
options: DecodeOptions::default(),
|
||||||
|
})
|
||||||
}
|
}
|
||||||
|
|
||||||
/// From the two exports on disk.
|
/// From the two exports on disk.
|
||||||
|
|||||||
@@ -13,9 +13,11 @@
|
|||||||
use std::path::Path;
|
use std::path::Path;
|
||||||
|
|
||||||
const MODEL: &str = "../../models/segment/yolo26n-seg.onnx";
|
const MODEL: &str = "../../models/segment/yolo26n-seg.onnx";
|
||||||
|
const QUANTISED: &str = "../../models/segment/yolo26n-seg.a16w16.onnx";
|
||||||
|
|
||||||
fn main() {
|
fn main() {
|
||||||
println!("cargo:rerun-if-changed={MODEL}");
|
println!("cargo:rerun-if-changed={MODEL}");
|
||||||
|
println!("cargo:rerun-if-changed={QUANTISED}");
|
||||||
println!("cargo:rerun-if-changed=build.rs");
|
println!("cargo:rerun-if-changed=build.rs");
|
||||||
|
|
||||||
// Only the embedded path needs the file present; a build without it is
|
// Only the embedded path needs the file present; a build without it is
|
||||||
@@ -24,10 +26,18 @@ fn main() {
|
|||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
|
|
||||||
let path = Path::new(MODEL);
|
check(MODEL);
|
||||||
|
// The Hexagon's quantised form rides only in an Android build.
|
||||||
|
if std::env::var("CARGO_CFG_TARGET_OS").as_deref() == Ok("android") {
|
||||||
|
check(QUANTISED);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
fn check(model: &str) {
|
||||||
|
let path = Path::new(model);
|
||||||
let Ok(bytes) = std::fs::read(path) else {
|
let Ok(bytes) = std::fs::read(path) else {
|
||||||
panic!(
|
panic!(
|
||||||
"\n\n{MODEL} is missing.\n\
|
"\n\n{model} is missing.\n\
|
||||||
It ships in Git LFS. Run `git lfs install && git lfs pull`, or build \
|
It ships in Git LFS. Run `git lfs install && git lfs pull`, or build \
|
||||||
with `--no-default-features` for a watershed-only build.\n"
|
with `--no-default-features` for a watershed-only build.\n"
|
||||||
);
|
);
|
||||||
@@ -40,7 +50,7 @@ fn main() {
|
|||||||
// happens in practice.
|
// happens in practice.
|
||||||
if bytes.starts_with(b"version https://git-lfs") {
|
if bytes.starts_with(b"version https://git-lfs") {
|
||||||
panic!(
|
panic!(
|
||||||
"\n\n{MODEL} is a Git LFS pointer, not the model ({} bytes).\n\
|
"\n\n{model} is a Git LFS pointer, not the model ({} bytes).\n\
|
||||||
Run `git lfs install && git lfs pull` to fetch the real file.\n",
|
Run `git lfs install && git lfs pull` to fetch the real file.\n",
|
||||||
bytes.len()
|
bytes.len()
|
||||||
);
|
);
|
||||||
@@ -50,7 +60,7 @@ fn main() {
|
|||||||
// export is ~11 MB; anything under a megabyte is a truncated checkout.
|
// export is ~11 MB; anything under a megabyte is a truncated checkout.
|
||||||
if bytes.len() < 1_000_000 {
|
if bytes.len() < 1_000_000 {
|
||||||
panic!(
|
panic!(
|
||||||
"\n\n{MODEL} is only {} bytes — expected ~11 MB.\n\
|
"\n\n{model} is only {} bytes — expected several MB.\n\
|
||||||
The checkout looks incomplete; try `git lfs pull`.\n",
|
The checkout looks incomplete; try `git lfs pull`.\n",
|
||||||
bytes.len()
|
bytes.len()
|
||||||
);
|
);
|
||||||
|
|||||||
@@ -72,7 +72,7 @@ pub use refine::{
|
|||||||
#[cfg(feature = "semantic")]
|
#[cfg(feature = "semantic")]
|
||||||
pub use scene::{Category, Scene, SceneModel};
|
pub use scene::{Category, Scene, SceneModel};
|
||||||
#[cfg(feature = "embedded-model")]
|
#[cfg(feature = "embedded-model")]
|
||||||
pub use semantic::embedded_model_bytes;
|
pub use semantic::embedded_models;
|
||||||
#[cfg(feature = "semantic")]
|
#[cfg(feature = "semantic")]
|
||||||
pub use semantic::{Instance, SemanticModel, SemanticOptions, Tiling};
|
pub use semantic::{Instance, SemanticModel, SemanticOptions, Tiling};
|
||||||
|
|
||||||
|
|||||||
@@ -123,21 +123,29 @@ impl SceneModel {
|
|||||||
classes: impl AsRef<std::path::Path>,
|
classes: impl AsRef<std::path::Path>,
|
||||||
categories: impl AsRef<std::path::Path>,
|
categories: impl AsRef<std::path::Path>,
|
||||||
) -> Result<Self, SegmentError> {
|
) -> Result<Self, SegmentError> {
|
||||||
|
// The form the device's backend runs: the `.a16w16.onnx` sibling on
|
||||||
|
// the Hexagon (attention left in float, inference.md §1.5), else this.
|
||||||
|
let (model, form) =
|
||||||
|
dr_inference_engine::resolve_model(dr_inference_engine::Role::Scene, model.as_ref());
|
||||||
let bytes = std::fs::read(model).map_err(SegmentError::ModelRead)?;
|
let bytes = std::fs::read(model).map_err(SegmentError::ModelRead)?;
|
||||||
let classes = std::fs::read_to_string(classes).map_err(SegmentError::ModelRead)?;
|
let classes = std::fs::read_to_string(classes).map_err(SegmentError::ModelRead)?;
|
||||||
let categories = std::fs::read_to_string(categories).map_err(SegmentError::ModelRead)?;
|
let categories = std::fs::read_to_string(categories).map_err(SegmentError::ModelRead)?;
|
||||||
let classes = crate::semantic::parse_classes(&classes);
|
let classes = crate::semantic::parse_classes(&classes);
|
||||||
let categories = parse_categories(&categories, &classes)?;
|
let categories = parse_categories(&categories, &classes)?;
|
||||||
Self::from_bytes(&bytes, categories)
|
Self::from_bytes_in(&bytes, form, categories)
|
||||||
}
|
}
|
||||||
|
|
||||||
pub fn from_bytes(bytes: &[u8], categories: Vec<Category>) -> Result<Self, SegmentError> {
|
pub fn from_bytes(bytes: &[u8], categories: Vec<Category>) -> Result<Self, SegmentError> {
|
||||||
// f32, as for `SemanticModel`; see there.
|
Self::from_bytes_in(bytes, dr_inference_engine::Form::F32, categories)
|
||||||
let session = dr_inference_engine::open(
|
}
|
||||||
dr_inference_engine::Role::Scene,
|
|
||||||
dr_inference_engine::Form::F32,
|
/// `bytes` in a stated numeric form; the outputs keep their shape.
|
||||||
bytes,
|
pub fn from_bytes_in(
|
||||||
)?;
|
bytes: &[u8],
|
||||||
|
form: dr_inference_engine::Form,
|
||||||
|
categories: Vec<Category>,
|
||||||
|
) -> Result<Self, SegmentError> {
|
||||||
|
let session = dr_inference_engine::open(dr_inference_engine::Role::Scene, form, bytes)?;
|
||||||
|
|
||||||
Ok(Self {
|
Ok(Self {
|
||||||
session,
|
session,
|
||||||
|
|||||||
@@ -208,18 +208,32 @@ const EMBEDDED_MODEL: &[u8] = include_bytes!("../../../models/segment/yolo26n-se
|
|||||||
#[cfg(feature = "embedded-model")]
|
#[cfg(feature = "embedded-model")]
|
||||||
const EMBEDDED_CLASSES: &str = include_str!("../../../models/segment/yolo26n-seg.classes.json");
|
const EMBEDDED_CLASSES: &str = include_str!("../../../models/segment/yolo26n-seg.classes.json");
|
||||||
|
|
||||||
/// The bytes of the model that ships with this crate, for whoever compiles
|
/// The Hexagon's form (docs/dev/inference.md §1.5): 16-bit activations and
|
||||||
|
/// weights, the rows' tail left in float. Only Android has a Hexagon, so only
|
||||||
|
/// Android carries it.
|
||||||
|
#[cfg(all(feature = "embedded-model", target_os = "android"))]
|
||||||
|
const EMBEDDED_A16W16: &[u8] = include_bytes!("../../../models/segment/yolo26n-seg.a16w16.onnx");
|
||||||
|
|
||||||
|
/// Every form of the model that ships with this crate, for whoever compiles
|
||||||
/// engines ahead of the first request (docs/dev/inference.md §6).
|
/// engines ahead of the first request (docs/dev/inference.md §6).
|
||||||
#[cfg(feature = "embedded-model")]
|
#[cfg(feature = "embedded-model")]
|
||||||
pub fn embedded_model_bytes() -> &'static [u8] {
|
pub fn embedded_models() -> Vec<(dr_inference_engine::Form, &'static [u8])> {
|
||||||
EMBEDDED_MODEL
|
#[allow(unused_mut)]
|
||||||
|
let mut forms = vec![(dr_inference_engine::Form::F32, EMBEDDED_MODEL)];
|
||||||
|
#[cfg(target_os = "android")]
|
||||||
|
forms.push((dr_inference_engine::Form::A16W16, EMBEDDED_A16W16));
|
||||||
|
forms
|
||||||
}
|
}
|
||||||
|
|
||||||
impl SemanticModel {
|
impl SemanticModel {
|
||||||
/// Load the model that ships with this crate.
|
/// Load the model that ships with this crate, in the form the device's
|
||||||
|
/// backend runs.
|
||||||
#[cfg(feature = "embedded-model")]
|
#[cfg(feature = "embedded-model")]
|
||||||
pub fn embedded() -> Result<Self, SegmentError> {
|
pub fn embedded() -> Result<Self, SegmentError> {
|
||||||
Self::from_bytes(EMBEDDED_MODEL, parse_classes(EMBEDDED_CLASSES))
|
let forms = embedded_models();
|
||||||
|
let (bytes, form) =
|
||||||
|
dr_inference_engine::choose_embedded(dr_inference_engine::Role::Segmenter, &forms);
|
||||||
|
Self::from_bytes_in(bytes, form, parse_classes(EMBEDDED_CLASSES))
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Load a model from an ONNX file, with `classes` supplying its vocabulary.
|
/// Load a model from an ONNX file, with `classes` supplying its vocabulary.
|
||||||
@@ -236,14 +250,18 @@ impl SemanticModel {
|
|||||||
}
|
}
|
||||||
|
|
||||||
pub fn from_bytes(bytes: &[u8], classes: Vec<Arc<str>>) -> Result<Self, SegmentError> {
|
pub fn from_bytes(bytes: &[u8], classes: Vec<Arc<str>>) -> Result<Self, SegmentError> {
|
||||||
// The f32 graph on whatever the device's backend is. An int8 form
|
Self::from_bytes_in(bytes, dr_inference_engine::Form::F32, classes)
|
||||||
// for the Hexagon waits on docs/dev/inference.md §10 M7 — the mask
|
}
|
||||||
// boundary has to be measured before it moves.
|
|
||||||
let session = dr_inference_engine::open(
|
/// `bytes` in a stated numeric form. The quantised one keeps the same
|
||||||
dr_inference_engine::Role::Segmenter,
|
/// outputs (the rows' tail stays float), so decoding does not change; the
|
||||||
dr_inference_engine::Form::F32,
|
/// masks it draws were measured against f32's (inference.md §1.5).
|
||||||
bytes,
|
pub fn from_bytes_in(
|
||||||
)?;
|
bytes: &[u8],
|
||||||
|
form: dr_inference_engine::Form,
|
||||||
|
classes: Vec<Arc<str>>,
|
||||||
|
) -> Result<Self, SegmentError> {
|
||||||
|
let session = dr_inference_engine::open(dr_inference_engine::Role::Segmenter, form, bytes)?;
|
||||||
|
|
||||||
Ok(Self { session, classes })
|
Ok(Self { session, classes })
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -275,13 +275,26 @@ impl FaceDetector {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Both ids this detector writes under — the f32 form and the int8 one —
|
/// The id when the detector runs with 16-bit activations and 8-bit
|
||||||
/// for a question that is about the detector and not about which form
|
/// weights, the Hexagon's form since the int8 one lost faces at 40–80 px
|
||||||
/// of it a device happened to run: "has the chosen detector been over
|
/// (docs/dev/inference.md §1.5). Different again from both, for the same
|
||||||
/// this image", asked by a re-index that must not ping-pong between a
|
/// reason as [`Self::model_id_int8`]; the embedder half is unchanged.
|
||||||
/// desktop that runs it in f32 and a tablet that runs it on the Hexagon.
|
pub fn model_id_a16w8(self) -> &'static str {
|
||||||
pub fn model_ids(self) -> [&'static str; 2] {
|
match self {
|
||||||
[self.model_id(), self.model_id_int8()]
|
FaceDetector::Scrfd500m => "scrfd_500m_a16+w600k_mbf",
|
||||||
|
FaceDetector::Scrfd2_5g => "scrfd_2.5g_a16+w600k_mbf",
|
||||||
|
FaceDetector::Scrfd10g => "scrfd_10g_a16+w600k_mbf",
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Every id this detector writes under — the f32 form and each quantised
|
||||||
|
/// one a device has run — for a question that is about the detector and
|
||||||
|
/// not about which form of it a device happened to run: "has the chosen
|
||||||
|
/// detector been over this image", asked by a re-index that must not
|
||||||
|
/// ping-pong between a desktop that runs it in f32 and a tablet that
|
||||||
|
/// runs it on the Hexagon.
|
||||||
|
pub fn model_ids(self) -> [&'static str; 3] {
|
||||||
|
[self.model_id(), self.model_id_int8(), self.model_id_a16w8()]
|
||||||
}
|
}
|
||||||
|
|
||||||
/// The detector that writes under a pipeline id, if it is one of these.
|
/// The detector that writes under a pipeline id, if it is one of these.
|
||||||
@@ -1769,11 +1782,16 @@ mod tests {
|
|||||||
#[test]
|
#[test]
|
||||||
fn a_detectors_two_spellings_share_its_embedder_and_nothing_else() {
|
fn a_detectors_two_spellings_share_its_embedder_and_nothing_else() {
|
||||||
for d in FaceDetector::ALL {
|
for d in FaceDetector::ALL {
|
||||||
let [f32_id, int8_id] = d.model_ids();
|
let [f32_id, int8_id, a16_id] = d.model_ids();
|
||||||
assert_eq!(f32_id, d.model_id());
|
assert_eq!(f32_id, d.model_id());
|
||||||
assert_eq!(int8_id, d.model_id_int8());
|
assert_eq!(int8_id, d.model_id_int8());
|
||||||
|
assert_eq!(a16_id, d.model_id_a16w8());
|
||||||
assert_ne!(f32_id, int8_id);
|
assert_ne!(f32_id, int8_id);
|
||||||
assert_eq!(f32_id.rsplit('+').next(), int8_id.rsplit('+').next());
|
assert_ne!(f32_id, a16_id);
|
||||||
|
assert_ne!(int8_id, a16_id);
|
||||||
|
for q in [int8_id, a16_id] {
|
||||||
|
assert_eq!(f32_id.rsplit('+').next(), q.rsplit('+').next());
|
||||||
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -1797,6 +1815,7 @@ mod tests {
|
|||||||
}
|
}
|
||||||
assert_eq!(FaceDetector::for_model_id(d.model_id()), Some(d));
|
assert_eq!(FaceDetector::for_model_id(d.model_id()), Some(d));
|
||||||
assert_eq!(FaceDetector::for_model_id(d.model_id_int8()), Some(d));
|
assert_eq!(FaceDetector::for_model_id(d.model_id_int8()), Some(d));
|
||||||
|
assert_eq!(FaceDetector::for_model_id(d.model_id_a16w8()), Some(d));
|
||||||
}
|
}
|
||||||
assert_eq!(FaceDetector::for_model_id("scrfd_10g+other"), None);
|
assert_eq!(FaceDetector::for_model_id("scrfd_10g+other"), None);
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -54,10 +54,14 @@ sed 's/$/\r/' "${REPO}/LICENSE" > "${STAGE}/LICENSE"
|
|||||||
# with its two descriptors, the panorama border filler and the denoiser. The installer
|
# with its two descriptors, the panorama border filler and the denoiser. The installer
|
||||||
# smoke test counts the same directories, so a model added here is expected
|
# smoke test counts the same directories, so a model added here is expected
|
||||||
# there without a number to update.
|
# there without a number to update.
|
||||||
|
#
|
||||||
|
# Not the quantised siblings (`*.int8.onnx`, `*.a16w8.onnx`, `*.a16w16.onnx`):
|
||||||
|
# they are the Hexagon's forms (docs/dev/inference.md §1.5), and a Windows
|
||||||
|
# machine has no Hexagon to load them.
|
||||||
for dir in face scene inpaint denoise; do
|
for dir in face scene inpaint denoise; do
|
||||||
for f in "${REPO}/models/${dir}"/*; do
|
for f in "${REPO}/models/${dir}"/*; do
|
||||||
case "$(basename "${f}")" in
|
case "$(basename "${f}")" in
|
||||||
README.md) continue ;;
|
README.md | *.int8.onnx | *.a16w8.onnx | *.a16w16.onnx) continue ;;
|
||||||
esac
|
esac
|
||||||
if [[ "${f}" == *.onnx && "$(stat -c%s "${f}")" -lt 100000 ]]; then
|
if [[ "${f}" == *.onnx && "$(stat -c%s "${f}")" -lt 100000 ]]; then
|
||||||
echo "error: $(basename "${f}") is $(stat -c%s "${f}") bytes — an LFS pointer, not a model." >&2
|
echo "error: $(basename "${f}") is $(stat -c%s "${f}") bytes — an LFS pointer, not a model." >&2
|
||||||
|
|||||||
@@ -331,6 +331,17 @@ If neither holds S's quality within 0.5 dB of fp32 on the real pairs, **v1 is de
|
|||||||
tablet shows the classical path. The sidecar still records the intent, so a desktop can render the
|
tablet shows the classical path. The sidecar still records the intent, so a desktop can render the
|
||||||
learned result for a photograph edited on the tablet.
|
learned result for a photograph edited on the tablet.
|
||||||
|
|
||||||
|
**Measured 2026-10-04 (inference.md §1.5): the second way holds, without the first.** The shipped
|
||||||
|
network, with its Bayer packing re-spelled as `SpaceToDepth` so QNN can hold it (the 6-D reshape
|
||||||
|
it replaces is exact but past the HTP's rank limit), at A16W16 — 16-bit activations and weights —
|
||||||
|
scores within 0.00 dB of f32 at ISO 400–25600 on the tablet's own HTP, and within 0.09 dB with the
|
||||||
|
6D's noise model scaled ×0.5, ×2 and ×4 to stand in for other sensors. A16W8 holds the 6D (worst
|
||||||
|
−0.19 dB at ISO 25600) but not ×4 noise at 25600 (−0.52 dB), so A16W16 is what ships. int8 loses
|
||||||
|
4.7–9.2 dB and fp16 is refused outright. A 1408 tile takes 95 ms on the Hexagon against 1510 ms on
|
||||||
|
the tablet's CPU: about 2.3 s for a 20 MP frame. Calibration ranges come from 96 training-day
|
||||||
|
tiles across every ISO, a third of them with that scaled noise; coverage of other bodies is that
|
||||||
|
synthetic bracket, not their raws.
|
||||||
|
|
||||||
## 9. X-Trans
|
## 9. X-Trans
|
||||||
|
|
||||||
The requirements tie this stage to FR-RAW-5, and the library has no Fuji raws. What we can do
|
The requirements tie this stage to FR-RAW-5, and the library has no Fuji raws. What we can do
|
||||||
|
|||||||
+64
-4
@@ -127,7 +127,8 @@ Three things the table settles.
|
|||||||
five of six models. A whole-library face index on the tablet goes from ~100 ms + 39 ms per face
|
five of six models. A whole-library face index on the tablet goes from ~100 ms + 39 ms per face
|
||||||
to ~1.4 ms + 12 ms per face, and the "Thorough" detector — 3× the cost of "Fast" today — becomes
|
to ~1.4 ms + 12 ms per face, and the "Thorough" detector — 3× the cost of "Fast" today — becomes
|
||||||
free. Its price is that the models must be **quantised to int8**, which is an accuracy question
|
free. Its price is that the models must be **quantised to int8**, which is an accuracy question
|
||||||
§5 has to answer before it is believed.
|
§5 has to answer before it is believed. (§1.5 answered it: int8 lost faces, and every model but
|
||||||
|
XFeat ships with 16-bit activations, at about three times these timings.)
|
||||||
- **The embedder does not gain from either accelerator.** 112×112 input, per-op overhead
|
- **The embedder does not gain from either accelerator.** 112×112 input, per-op overhead
|
||||||
dominates; it is 9 ms on the tablet's CPU and 12 ms on its NPU. It stays float, which §7 turns
|
dominates; it is 9 ms on the tablet's CPU and 12 ms on its NPU. It stays float, which §7 turns
|
||||||
from a performance footnote into a correctness rule.
|
from a performance footnote into a correctness rule.
|
||||||
@@ -137,6 +138,60 @@ Three things the table settles.
|
|||||||
- **On AMD, MIGraphX fp16 is 4–17× the CPU provider** on the detectors and 60× on the
|
- **On AMD, MIGraphX fp16 is 4–17× the CPU provider** on the detectors and 60× on the
|
||||||
inpainter, with the same first-run compile cost as TensorRT and no rung between it and the CPU.
|
inpainter, with the same first-run compile cost as TensorRT and no rung between it and the CPU.
|
||||||
|
|
||||||
|
### 1.5 The Hexagon at every bit width · 2026-10-04
|
||||||
|
|
||||||
|
§1.1's Hexagon column is int8 calibrated on noise: timing only. This is the follow-up — every
|
||||||
|
model, every bit width the HTP offers, calibrated on real photographs and **scored on the tablet
|
||||||
|
itself** (ORT 1.29 + QNN 2.42, `htp_arch` 73), against the f32 model on the same inputs. The
|
||||||
|
"Form shipped" column is the files in `models/`, re-scored on the tablet after
|
||||||
|
`tools/quantise-models.sh` wrote them. The
|
||||||
|
calibration and scoring photographs are 800 from the public COCO val2017 set (CC-BY); the face
|
||||||
|
models' numbers are over the faces in them of at least 32 px. The tools are `tools/quantise-models.sh`
|
||||||
|
and the scratch harness described with it.
|
||||||
|
|
||||||
|
**What the HTP accepts.** fp16: nothing — every fp16 operator fails validation (3110), on QNN
|
||||||
|
2.42 and 2.50, with `htp_arch` and every `soc_model` tried; the fp16 rung stays off the table until
|
||||||
|
someone has Qualcomm's own SDK to say why. 4-bit weights (A8W4, A16W4): load, and wreck accuracy
|
||||||
|
(SCRFD finds 25–35% of f32's faces). What is left: **A8W8 (int8), A16W8 and A16W16**, all running
|
||||||
|
the whole graph. 16-bit activations cost about 3× int8's time, A16W16 about 4×.
|
||||||
|
|
||||||
|
| Model | ORT CPU f32 | Form shipped | Hexagon | On the tablet, against f32 | int8 for comparison |
|
||||||
|
|---|---|---|---|---|---|
|
||||||
|
| scrfd_500m / 2.5g / 10g | 17 / 56 / 198 ms | **A16W8** | 4.2 / 5.1 / 9.0 ms | 100% of faces found in every size band; keypoints 0.3–0.6% of the box | 94–95% of faces at 40–80 px |
|
||||||
|
| 2d106det (landmarks) | 2.8 ms | **A16W8** | 0.5 ms | 0.25 px in the 192 crop (eye points 0.20) | 1.5 px, and 29 partitions at 7.3 ms |
|
||||||
|
| yolo26n-seg | 90 ms | **A16W16**, tail in float | 12.9 ms | 98.2% of objects, mask IoU 0.994 | 74% (simulated) |
|
||||||
|
| yolo26s-sem-ade20k | 151 ms | **A16W16**, attention in float | 15 ms | 98.9% of cells agree on the class, TV 0.009 | 67% |
|
||||||
|
| migan-512 | 488 ms | **A16W16** | 87 ms | 41 dB from f32 in the fill (worst 1%: 30 dB) | 16 dB (simulated) |
|
||||||
|
| xfeat-1024 / 768 | 58 ms | **int8**, rewritten graph | 6.5 ms | panorama alignment 0.45 px from f32's — f32's own refit on 90% of its matches is 0.41 | — |
|
||||||
|
| mosaic-1408 (denoiser) | 1510 ms a tile | **A16W16**, rewritten graph | 95 ms a tile | 0.00 dB at every ISO; ≤ 0.09 dB with the noise scaled ×0.5–×4 | −4.7 to −9.2 dB |
|
||||||
|
| arcface_mbf (embedder) | 8.5 ms | f32, CPU | — | A16W16: cosine 0.9995, p1 0.9967 — misses §7's 0.999 gate | — |
|
||||||
|
| ocec, sgc (eyes) | 1, 1.7 ms | f32, CPU | — | sgc flips 1.45% of views even at A16W16; not worth a millisecond | — |
|
||||||
|
|
||||||
|
Four things the table needed that the f32 graphs did not have, all in `tools/htp_graph.py` and all
|
||||||
|
checked exact against the f32 graph before they are used:
|
||||||
|
|
||||||
|
- **Rank ≤ 5.** QNN's tensors stop at rank 5, and the denoiser packs the mosaic through a 6-D
|
||||||
|
reshape (6007 at compose). For one channel that reshape is `SpaceToDepth(2)`. XFeat's 8×8 unfold
|
||||||
|
is 224 Slices and 6-D Concats; it is `SpaceToDepth(8)` (736 nodes to 60).
|
||||||
|
- **No bilinear Resize at XFeat's sizes** (3110). A half-pixel bilinear resize between fixed sizes
|
||||||
|
is two constant matrices, so it is two MatMuls.
|
||||||
|
- **One scale per tensor.** The segmenter's output rows carry boxes in pixels beside scores in
|
||||||
|
0..1; quantised as one tensor the scores vanish. Everything from the Concat that builds the rows
|
||||||
|
stays float, on the CPU, where the top-300 selection costs nothing.
|
||||||
|
- **Float where the HTP's 16-bit arithmetic drifts.** The scene model's one attention block
|
||||||
|
(two MatMuls and a Softmax over 400 tokens) moved its agreement from 98.7% to 96.7%; it stays
|
||||||
|
float.
|
||||||
|
|
||||||
|
**ORT's CPU simulation of a QDQ graph is not the tablet.** It matched to the hundredth of a dB for
|
||||||
|
the denoiser and to rounding for the detectors, landmarks and XFeat, and it overstated MI-GAN by
|
||||||
|
27 dB and the scene model by three points. Every number above is the device's; a new form is not
|
||||||
|
measured until it has run there.
|
||||||
|
|
||||||
|
**XFeat's int8 loses keypoints and not the panorama.** 83% of f32's keypoints come back within
|
||||||
|
1.5 px; but over the twelve-frame `fixtures/pano/2025-08-05` sweep, the homographies fitted from
|
||||||
|
int8's matches land 0.45 px from f32's in the overlaps — the spread f32 shows against itself
|
||||||
|
(0.41).
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## 2. The shape of the answer
|
## 2. The shape of the answer
|
||||||
@@ -146,7 +201,7 @@ winning:
|
|||||||
|
|
||||||
| Platform | 1st | 2nd | 3rd | Floor |
|
| Platform | 1st | 2nd | 3rd | Floor |
|
||||||
|---|---|---|---|---|
|
|---|---|---|---|---|
|
||||||
| Android, Qualcomm with a Hexagon the shipped QNN skel covers (V68–V81) | QNN HTP, int8 model | ORT CPU, f32 model | — | tract |
|
| Android, Qualcomm with a Hexagon the shipped QNN skel covers (V68–V81) | QNN HTP, each model's quantised form (§1.5) | ORT CPU, f32 model | — | tract |
|
||||||
| Android, any other SoC | ORT CPU, f32 | — | — | tract |
|
| Android, any other SoC | ORT CPU, f32 | — | — | tract |
|
||||||
| Linux / Windows, NVIDIA GPU | TensorRT, f32 model, fp16 engine | CUDA provider, f32 | ORT CPU, f32 | tract |
|
| Linux / Windows, NVIDIA GPU | TensorRT, f32 model, fp16 engine | CUDA provider, f32 | ORT CPU, f32 | tract |
|
||||||
| Linux, AMD GPU with ROCm | MIGraphX, f32 model, fp16 program | ORT CPU, f32 | — | tract |
|
| Linux, AMD GPU with ROCm | MIGraphX, f32 model, fp16 program | ORT CPU, f32 | — | tract |
|
||||||
@@ -298,10 +353,10 @@ of which form they load:
|
|||||||
| Form | Who produces it | When | Needed by |
|
| Form | Who produces it | When | Needed by |
|
||||||
|---|---|---|---|
|
|---|---|---|---|
|
||||||
| f32 ONNX, shape-fixed, **opset ≥ 13** | `tools/fix-face-model-shapes.sh`, `tools/export-seg-model.sh` | Release time, once | Every rung except Hexagon |
|
| f32 ONNX, shape-fixed, **opset ≥ 13** | `tools/fix-face-model-shapes.sh`, `tools/export-seg-model.sh` | Release time, once | Every rung except Hexagon |
|
||||||
| int8 QDQ ONNX, per-channel, uint8 activations | `tools/quantise-models.sh` (new) | Release time, once, **calibrated on real photographs** | Hexagon |
|
| QDQ ONNX, per-channel — int8, A16W8 or A16W16 per model (§1.5) | `tools/quantise-models.sh` | Release time, once, **calibrated on real photographs**, scored on the tablet | Hexagon |
|
||||||
| TensorRT engine (`.engine`, per GPU architecture and TensorRT version) | The app, from the f32 file | First run on that device, in the background | TensorRT rung |
|
| TensorRT engine (`.engine`, per GPU architecture and TensorRT version) | The app, from the f32 file | First run on that device, in the background | TensorRT rung |
|
||||||
| MIGraphX program (`.mxr`, per GPU architecture, MIGraphX version and precision) | The app, from the f32 file | First run on that device, in the background | MIGraphX rung |
|
| MIGraphX program (`.mxr`, per GPU architecture, MIGraphX version and precision) | The app, from the f32 file | First run on that device, in the background | MIGraphX rung |
|
||||||
| QNN context binary | The app, from the int8 file | First run on that device, in the background | Hexagon rung |
|
| QNN context binary | The app, from the quantised file | First run on that device, in the background | Hexagon rung |
|
||||||
|
|
||||||
Two rules.
|
Two rules.
|
||||||
|
|
||||||
@@ -380,6 +435,11 @@ already the rule for the detector and because §5 is the gate on whether the int
|
|||||||
enough to be *offered* at all. f32 on tract, ORT CPU, CUDA and TensorRT-f32 are one identity: the
|
enough to be *offered* at all. f32 on tract, ORT CPU, CUDA and TensorRT-f32 are one identity: the
|
||||||
same graph, the same arithmetic, differences at the last bit.
|
same graph, the same arithmetic, differences at the last bit.
|
||||||
|
|
||||||
|
After §1.5 the Hexagon runs the detectors in **A16W8**, and that is a third spelling:
|
||||||
|
`scrfd_500m_a16+w600k_mbf` and its two siblings. Same rule, same reconciliation; a tablet that
|
||||||
|
indexed under `_i8` keeps those rows, and `FaceDetector::model_ids` answers "has this detector been
|
||||||
|
over this image" for all three forms.
|
||||||
|
|
||||||
**The embedder** is where comparability across devices is the whole point, and it is the one
|
**The embedder** is where comparability across devices is the whole point, and it is the one
|
||||||
model that no accelerator helps (§1.4). So: **the embedder runs in f32 on every rung.** On TensorRT
|
model that no accelerator helps (§1.4). So: **the embedder runs in f32 on every rung.** On TensorRT
|
||||||
that means the embedder's engine is built without fp16 while the detector's is built with it; on
|
that means the embedder's engine is built without fp16 while the detector's is built with it; on
|
||||||
|
|||||||
@@ -467,9 +467,11 @@ the tablet. Three ways to make it viable, none built:
|
|||||||
1. **Fill at a quarter of the resolution and upsample.** Sky and scree
|
1. **Fill at a quarter of the resolution and upsample.** Sky and scree
|
||||||
tolerate it; twenty-odd tiles, about three minutes on the desktop CPU. A
|
tolerate it; twenty-odd tiles, about three minutes on the desktop CPU. A
|
||||||
background job with the outbox's patience, not an interactive one.
|
background job with the outbox's patience, not an interactive one.
|
||||||
2. **int8 on the tablet's Hexagon through QNN**, where the plain-conv design
|
2. **The tablet's Hexagon through QNN**, where the plain-conv design is the
|
||||||
is the point and the whole graph should run in milliseconds. The setup
|
point. Measured 2026-10-04 (inference.md §1.5): int8 changes the fill
|
||||||
exists from the eye-state work; MI-GAN is a candidate for the same path.
|
(16 dB from f32's), so it ships with 16-bit activations and weights —
|
||||||
|
87 ms a tile against 488 ms on the tablet's CPU, the whole graph on the
|
||||||
|
NPU, 41 dB from f32 in the hole.
|
||||||
3. **A WGSL runtime for those six operators.** A project of its own, and
|
3. **A WGSL runtime for those six operators.** A project of its own, and
|
||||||
the only route that would make it interactive on the desktop.
|
the only route that would make it interactive on the desktop.
|
||||||
|
|
||||||
|
|||||||
+10
-10
File diff suppressed because one or more lines are too long
+2
-2
@@ -287,7 +287,6 @@ $LOCALAPPDATA\Programs\DarkRoom\
|
|||||||
models\
|
models\
|
||||||
scrfd_500m_640.onnx scrfd_2.5g_640.onnx scrfd_10g_640.onnx arcface_mbf_b1.onnx
|
scrfd_500m_640.onnx scrfd_2.5g_640.onnx scrfd_10g_640.onnx arcface_mbf_b1.onnx
|
||||||
2d106det_b1.onnx ocec_s_b1.onnx sgc_l_48_b1.onnx
|
2d106det_b1.onnx ocec_s_b1.onnx sgc_l_48_b1.onnx
|
||||||
scrfd_500m_640.int8.onnx scrfd_2.5g_640.int8.onnx scrfd_10g_640.int8.onnx
|
|
||||||
yolo26s-sem-ade20k.onnx yolo26s-sem-ade20k.classes.json categories.txt
|
yolo26s-sem-ade20k.onnx yolo26s-sem-ade20k.classes.json categories.txt
|
||||||
migan-512.onnx
|
migan-512.onnx
|
||||||
manual\
|
manual\
|
||||||
@@ -297,7 +296,8 @@ $LOCALAPPDATA\Programs\DarkRoom\
|
|||||||
```
|
```
|
||||||
|
|
||||||
Plus a Start Menu shortcut, and nothing on the desktop unless the user ticks it. The models are
|
Plus a Start Menu shortcut, and nothing on the desktop unless the user ticks it. The models are
|
||||||
the same ten files the APK bundles and the PKGBUILD installs; `models\` beside the executable is
|
the f32 files the APK bundles and the PKGBUILD installs — not the APK's quantised siblings, which
|
||||||
|
only a Hexagon runs; `models\` beside the executable is
|
||||||
where §3.2's lookup finds them. **No `LICENSE` yet**: the repository has no licence file at its
|
where §3.2's lookup finds them. **No `LICENSE` yet**: the repository has no licence file at its
|
||||||
root (the Arch package points at the system's shared GPL text), so the installer has no licence
|
root (the Arch package points at the system's shared GPL text), so the installer has no licence
|
||||||
page until one is added — a one-file change, and the `.nsi` says where the page then goes. The face weights carry the research-only grant that
|
page until one is added — a one-file change, and the `.nsi` says where the page then goes. The face weights carry the research-only grant that
|
||||||
|
|||||||
@@ -14,6 +14,13 @@ Both come from `https://huggingface.co/Ultralytics/YOLO26`. The face weights in
|
|||||||
The keypoint weights in `keypoints/` and the border filler in `inpaint/` are
|
The keypoint weights in `keypoints/` and the border filler in `inpaint/` are
|
||||||
the other two, and the easiest — see the last two sections.
|
the other two, and the easiest — see the last two sections.
|
||||||
|
|
||||||
|
Every quantised sibling — `*.int8.onnx`, `*.a16w8.onnx`, `*.a16w16.onnx`, the
|
||||||
|
forms the tablet's Hexagon runs (`tools/quantise-models.sh`) — is the same
|
||||||
|
weights rounded, and carries exactly the grant of the file it was made from.
|
||||||
|
What calibration adds is one minimum and maximum per tensor: from public COCO
|
||||||
|
val2017 photographs (CC-BY 4.0) for the image models, and for the denoiser from
|
||||||
|
the same training tiles its weights were learned from. No image is in the files.
|
||||||
|
|
||||||
## The grant
|
## The grant
|
||||||
|
|
||||||
**Ultralytics releases YOLO under AGPL-3.0**, and the weights carry the same
|
**Ultralytics releases YOLO under AGPL-3.0**, and the weights carry the same
|
||||||
|
|||||||
Binary file not shown.
Binary file not shown.
+14
-8
@@ -28,15 +28,21 @@ every reader treats "never read" as unknown, never as closed. They are found in
|
|||||||
directory as the pair, so a hand-placed pair does not pick up a package's eye models from a
|
directory as the pair, so a hand-placed pair does not pick up a package's eye models from a
|
||||||
directory it otherwise outranks.
|
directory it otherwise outranks.
|
||||||
|
|
||||||
scrfd_500m_640.int8.onnx 0.8 MB the same three, in the form the Hexagon NPU takes
|
scrfd_500m_640.a16w8.onnx the same three, in the form the Hexagon NPU takes
|
||||||
scrfd_2.5g_640.int8.onnx 0.9 MB (docs/dev/inference.md §5) — opset 17, per-channel int8
|
scrfd_2.5g_640.a16w8.onnx (docs/dev/inference.md §1.5) — 16-bit activations,
|
||||||
scrfd_10g_640.int8.onnx 4.3 MB weights, uint8 activations, calibrated on 96 photographs
|
scrfd_10g_640.a16w8.onnx per-channel 8-bit weights, calibrated on 300 photographs
|
||||||
|
2d106det_b1.a16w8.onnx the landmarks, likewise
|
||||||
|
|
||||||
The int8 files are **derived** by `tools/quantise-models.sh` from the f32 ones beside them and
|
These are **derived** by `tools/quantise-models.sh` from the f32 files beside them and travel with
|
||||||
travel with them: the engine loads the `.int8.onnx` sibling when the device's backend wants it and
|
them: the engine loads the `.a16w8.onnx` sibling when the device's backend wants it and the
|
||||||
the canonical file otherwise, and a library indexed on the int8 form records it as a different
|
canonical file otherwise, and a library indexed on that form records it as a different detector
|
||||||
detector (`scrfd_500m_i8+w600k_mbf`), because it finds a different set of faces. Every other
|
(`scrfd_500m_a16+w600k_mbf`), because it finds a different set of faces. Every other platform
|
||||||
platform ignores them. The embedder has no int8 form and never will (§7 of the same document).
|
ignores them, and the Windows installer leaves them out. The embedder and the eye classifiers have
|
||||||
|
no quantised form: the embedder's vectors must compare across devices (inference.md §7), and the
|
||||||
|
classifiers cost a millisecond on the CPU.
|
||||||
|
|
||||||
|
The detectors were int8 until §1.5 measured them on the tablet: int8 found 94–95% of f32's faces
|
||||||
|
at 40–80 px, A16W8 all of them. A tablet that indexed under the int8 ids keeps those rows.
|
||||||
|
|
||||||
**A clone without git-lfs gets a ~130-byte pointer where each model should be.** Both packagers check
|
**A clone without git-lfs gets a ~130-byte pointer where each model should be.** Both packagers check
|
||||||
for exactly that and refuse, rather than shipping the pointer and failing inside tract on the user's
|
for exactly that and refuse, rather than shipping the pointer and failing inside tract on the user's
|
||||||
|
|||||||
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
@@ -46,12 +46,14 @@ pub fn init(runtime_dirs: Vec<PathBuf>) {
|
|||||||
cache_dir: crate::library::inference_cache_dir(),
|
cache_dir: crate::library::inference_cache_dir(),
|
||||||
models,
|
models,
|
||||||
embedded: {
|
embedded: {
|
||||||
let [landscape, portrait] = dr_pano::xfeat::embedded_model_bytes();
|
let [landscape, portrait] = dr_pano::xfeat::embedded_models();
|
||||||
vec![
|
let tag = |role| move |(form, bytes)| (role, form, bytes);
|
||||||
(Role::Segmenter, dr_segment::embedded_model_bytes()),
|
dr_segment::embedded_models()
|
||||||
(Role::Keypoints, landscape),
|
.into_iter()
|
||||||
(Role::Keypoints, portrait),
|
.map(tag(Role::Segmenter))
|
||||||
]
|
.chain(landscape.into_iter().map(tag(Role::Keypoints)))
|
||||||
|
.chain(portrait.into_iter().map(tag(Role::Keypoints)))
|
||||||
|
.collect()
|
||||||
},
|
},
|
||||||
ceiling: None,
|
ceiling: None,
|
||||||
threads: 0,
|
threads: 0,
|
||||||
@@ -81,7 +83,7 @@ pub fn user_runtime_dir() -> PathBuf {
|
|||||||
///
|
///
|
||||||
/// Reads the shared and system directories only. An account-private model
|
/// Reads the shared and system directories only. An account-private model
|
||||||
/// directory can override the file `library::face_models` loads, but not
|
/// directory can override the file `library::face_models` loads, but not
|
||||||
/// which form the backend wants, and the int8 sibling is something a
|
/// which form the backend wants, and the quantised sibling is something a
|
||||||
/// packager ships, not something a user drops in.
|
/// packager ships, not something a user drops in.
|
||||||
pub fn detector_form(detector: FaceDetector) -> Form {
|
pub fn detector_form(detector: FaceDetector) -> Form {
|
||||||
let canonical = crate::library::shared_model(detector.file_name())
|
let canonical = crate::library::shared_model(detector.file_name())
|
||||||
@@ -92,8 +94,11 @@ pub fn detector_form(detector: FaceDetector) -> Form {
|
|||||||
/// The `faces.model_id` this device indexes under with `detector`.
|
/// The `faces.model_id` this device indexes under with `detector`.
|
||||||
pub fn model_id(detector: FaceDetector) -> &'static str {
|
pub fn model_id(detector: FaceDetector) -> &'static str {
|
||||||
match detector_form(detector) {
|
match detector_form(detector) {
|
||||||
Form::F32 => detector.model_id(),
|
|
||||||
Form::Int8 => detector.model_id_int8(),
|
Form::Int8 => detector.model_id_int8(),
|
||||||
|
Form::A16W8 => detector.model_id_a16w8(),
|
||||||
|
// No detector is offered in A16W16 (inference.md §1.5); were one, it
|
||||||
|
// would be the network f32 is to the last bit that a person can see.
|
||||||
|
Form::F32 | Form::A16W16 => detector.model_id(),
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user