S15.4, CPU half: XFeat runs in ~400 ms per frame on the tablet

tools/onnx-probe-on-device.sh cross-builds dr-segment's onnx_probe
without the embedded segmentation model, pushes it with a model to the
attached device and times two runs. The 768×1024 XFeat export takes
~400 ms on the reference tablet's NEON cores against ~300 ms on the
desktop, with identical output ranges — inside NFR-MRG-1's 1 s per frame.
The blend half of S15.4 waits for a chunked blend to exist.
This commit is contained in:
2026-09-19 15:24:12 +02:00
parent 5bf06c5030
commit 2bf0ec8dba
3 changed files with 71 additions and 5 deletions
+8 -3
View File
@@ -183,9 +183,14 @@ standard: `Conv`, `InstanceNormalization`, `AveragePool`, `Resize`, `Slice`,
the 2.8 MB file through the app's own `ort`-over-tract backend with nothing the 2.8 MB file through the app's own `ort`-over-tract backend with nothing
unsupported, in 28 ms, and runs it in **~300 ms on the reference desktop's unsupported, in 28 ms, and runs it in **~300 ms on the reference desktop's
CPU**. The weights ship as `models/keypoints/xfeat-1024.onnx`, recorded in CPU**. The weights ship as `models/keypoints/xfeat-1024.onnx`, recorded in
`models/LICENCE.md`. Still to do: the tablet figure (S15.4), and a `models/LICENCE.md`. **On the tablet** (S15.4's CPU half, same day):
keypoint-level comparison against the PyTorch reference once the Rust decoder `tools/onnx-probe-on-device.sh` cross-builds the probe, and the same file
exists — the probe proves the graph runs, not that the numbers match. runs in **~400 ms per frame** on the reference tablet's NEON cores (ROD2-W09,
SM8635), with output ranges identical to the desktop's — inside NFR-MRG-1's
1 s per frame with room to spare, and 1.3× the desktop rather than the 2×
faces.md §9 measured for its scan. Still to do: a keypoint-level comparison
against the PyTorch reference once the Rust decoder exists — the probe
proves the graph runs, not that the numbers match.
The outputs are three maps at 1/8 resolution, 96×128 for the export size: The outputs are three maps at 1/8 resolution, 96×128 for the export size:
64-channel descriptors, 65-channel keypoint logits (each 8×8 cell's position 64-channel descriptors, 65-channel keypoint logits (each 8×8 cell's position
+3 -2
View File
@@ -1839,8 +1839,9 @@ touched by a resize. That is the list in the criterion, and it is the testable h
**NFR-MRG-1 — Merge latency.** For five 24 MP frames: the alignment preview (FR-MRG-7) within **NFR-MRG-1 — Merge latency.** For five 24 MP frames: the alignment preview (FR-MRG-7) within
5 s on the reference desktop and 15 s on the reference tablet, of which keypoint detection is at 5 s on the reference desktop and 15 s on the reference tablet, of which keypoint detection is at
most 1 s per frame on the tablet's CPU; the full merge written to disk within 60 s on the desktop. most 1 s per frame on the tablet's CPU (measured 2026-09-19 at ~0.4 s, S15.4); the full merge
The tablet's full-merge figure is set by S15, not guessed here. written to disk within 60 s on the desktop. The tablet's full-merge figure is set by S15's blend
measurement, not guessed here.
**Performance regressions fail the build.** §8's benchmark suite runs per-commit; a regression **Performance regressions fail the build.** §8's benchmark suite runs per-commit; a regression
beyond a stated tolerance is a build failure, not a notification. Performance work rots otherwise. beyond a stated tolerance is a build failure, not a notification. Performance work rots otherwise.
+60
View File
@@ -0,0 +1,60 @@
#!/usr/bin/env bash
# Time an ONNX model under tract on a connected Android device.
#
# ./tools/onnx-probe-on-device.sh models/keypoints/xfeat-1024.onnx [1x1x768x1024]
#
# ## Why this exists
#
# Every model the app runs on the CPU — faces, masks, keypoints — runs under
# tract on NEON when it runs on the tablet, and the desktop figure says
# nothing about that: twenty cores of AVX2 against a tablet's big.LITTLE
# cluster is not a ratio anyone should guess (faces.md §9 measured 2×
# on one workload and would have guessed wrong). NFR-MRG-1 puts a per-frame
# budget on keypoint detection *on the tablet*, and this is how the number
# is measured rather than assumed.
#
# Cross-builds `dr-segment`'s `onnx_probe` example — without the embedded
# segmentation model, which the probe does not need and which is 11 MB the
# push would otherwise carry — pushes it with the model to `/data/local/tmp`,
# runs it twice (tract's first run pays for planning) and removes both.
#
# Not part of CI, which has no device attached.
set -euo pipefail
TARGET=aarch64-linux-android
API=26
DEST=/data/local/tmp
MODEL="${1:?usage: onnx-probe-on-device.sh <model.onnx> [NxCxHxW]}"
SHAPE="${2:-1x1x768x1024}"
ndk="${ANDROID_NDK_HOME:-}"
if [[ -z "$ndk" ]]; then
ndk=$(find "${ANDROID_HOME:-$HOME/Android/Sdk}/ndk" -maxdepth 1 -mindepth 1 -type d 2>/dev/null |
sort -V | tail -1)
fi
clang="$ndk/toolchains/llvm/prebuilt/linux-x86_64/bin/$TARGET$API-clang"
[[ -x "$clang" ]] || {
echo "no $TARGET$API-clang in $ndk — set ANDROID_NDK_HOME" >&2
exit 1
}
rustup target list --installed | grep -qx "$TARGET" || rustup target add "$TARGET"
adb devices | grep -qw device || {
echo "no device attached — plug the tablet in and enable USB debugging" >&2
exit 1
}
echo "building onnx_probe for $TARGET…"
CARGO_TARGET_AARCH64_LINUX_ANDROID_LINKER="$clang" \
CC_aarch64_linux_android="$clang" \
cargo build -q -p dr-segment --example onnx_probe --release --target "$TARGET" \
--no-default-features --features semantic
binary="${CARGO_TARGET_DIR:-target}/$TARGET/release/examples/onnx_probe"
echo "pushing $(basename "$binary") and $(basename "$MODEL")…"
adb push "$binary" "$DEST/onnx_probe" >/dev/null
adb push "$MODEL" "$DEST/probe.onnx" >/dev/null
adb shell chmod 755 "$DEST/onnx_probe"
adb shell "$DEST/onnx_probe" "$DEST/probe.onnx" "$SHAPE"
adb shell rm -f "$DEST/onnx_probe" "$DEST/probe.onnx"