From 2bf0ec8dba11b9dc1baa35919fb9cda1c5157aec Mon Sep 17 00:00:00 2001 From: Duncan Tourolle Date: Sat, 19 Sep 2026 13:01:37 +0200 Subject: [PATCH] S15.4, CPU half: XFeat runs in ~400 ms per frame on the tablet MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit tools/onnx-probe-on-device.sh cross-builds dr-segment's onnx_probe without the embedded segmentation model, pushes it with a model to the attached device and times two runs. The 768×1024 XFeat export takes ~400 ms on the reference tablet's NEON cores against ~300 ms on the desktop, with identical output ranges — inside NFR-MRG-1's 1 s per frame. The blend half of S15.4 waits for a chunked blend to exist. --- docs/panorama.md | 11 +++++-- docs/requirements.md | 5 +-- tools/onnx-probe-on-device.sh | 60 +++++++++++++++++++++++++++++++++++ 3 files changed, 71 insertions(+), 5 deletions(-) create mode 100755 tools/onnx-probe-on-device.sh diff --git a/docs/panorama.md b/docs/panorama.md index f96217f..09e1415 100644 --- a/docs/panorama.md +++ b/docs/panorama.md @@ -183,9 +183,14 @@ standard: `Conv`, `InstanceNormalization`, `AveragePool`, `Resize`, `Slice`, the 2.8 MB file through the app's own `ort`-over-tract backend with nothing unsupported, in 28 ms, and runs it in **~300 ms on the reference desktop's CPU**. The weights ship as `models/keypoints/xfeat-1024.onnx`, recorded in -`models/LICENCE.md`. Still to do: the tablet figure (S15.4), and a -keypoint-level comparison against the PyTorch reference once the Rust decoder -exists — the probe proves the graph runs, not that the numbers match. +`models/LICENCE.md`. **On the tablet** (S15.4's CPU half, same day): +`tools/onnx-probe-on-device.sh` cross-builds the probe, and the same file +runs in **~400 ms per frame** on the reference tablet's NEON cores (ROD2-W09, +SM8635), with output ranges identical to the desktop's — inside NFR-MRG-1's +1 s per frame with room to spare, and 1.3× the desktop rather than the 2× +faces.md §9 measured for its scan. Still to do: a keypoint-level comparison +against the PyTorch reference once the Rust decoder exists — the probe +proves the graph runs, not that the numbers match. The outputs are three maps at 1/8 resolution, 96×128 for the export size: 64-channel descriptors, 65-channel keypoint logits (each 8×8 cell's position diff --git a/docs/requirements.md b/docs/requirements.md index d4eceb7..2ec5388 100644 --- a/docs/requirements.md +++ b/docs/requirements.md @@ -1839,8 +1839,9 @@ touched by a resize. That is the list in the criterion, and it is the testable h **NFR-MRG-1 — Merge latency.** For five 24 MP frames: the alignment preview (FR-MRG-7) within 5 s on the reference desktop and 15 s on the reference tablet, of which keypoint detection is at -most 1 s per frame on the tablet's CPU; the full merge written to disk within 60 s on the desktop. -The tablet's full-merge figure is set by S15, not guessed here. +most 1 s per frame on the tablet's CPU (measured 2026-09-19 at ~0.4 s, S15.4); the full merge +written to disk within 60 s on the desktop. The tablet's full-merge figure is set by S15's blend +measurement, not guessed here. **Performance regressions fail the build.** §8's benchmark suite runs per-commit; a regression beyond a stated tolerance is a build failure, not a notification. Performance work rots otherwise. diff --git a/tools/onnx-probe-on-device.sh b/tools/onnx-probe-on-device.sh new file mode 100755 index 0000000..4ea6db1 --- /dev/null +++ b/tools/onnx-probe-on-device.sh @@ -0,0 +1,60 @@ +#!/usr/bin/env bash +# Time an ONNX model under tract on a connected Android device. +# +# ./tools/onnx-probe-on-device.sh models/keypoints/xfeat-1024.onnx [1x1x768x1024] +# +# ## Why this exists +# +# Every model the app runs on the CPU — faces, masks, keypoints — runs under +# tract on NEON when it runs on the tablet, and the desktop figure says +# nothing about that: twenty cores of AVX2 against a tablet's big.LITTLE +# cluster is not a ratio anyone should guess (faces.md §9 measured 2× +# on one workload and would have guessed wrong). NFR-MRG-1 puts a per-frame +# budget on keypoint detection *on the tablet*, and this is how the number +# is measured rather than assumed. +# +# Cross-builds `dr-segment`'s `onnx_probe` example — without the embedded +# segmentation model, which the probe does not need and which is 11 MB the +# push would otherwise carry — pushes it with the model to `/data/local/tmp`, +# runs it twice (tract's first run pays for planning) and removes both. +# +# Not part of CI, which has no device attached. +set -euo pipefail + +TARGET=aarch64-linux-android +API=26 +DEST=/data/local/tmp + +MODEL="${1:?usage: onnx-probe-on-device.sh [NxCxHxW]}" +SHAPE="${2:-1x1x768x1024}" + +ndk="${ANDROID_NDK_HOME:-}" +if [[ -z "$ndk" ]]; then + ndk=$(find "${ANDROID_HOME:-$HOME/Android/Sdk}/ndk" -maxdepth 1 -mindepth 1 -type d 2>/dev/null | + sort -V | tail -1) +fi +clang="$ndk/toolchains/llvm/prebuilt/linux-x86_64/bin/$TARGET$API-clang" +[[ -x "$clang" ]] || { + echo "no $TARGET$API-clang in $ndk — set ANDROID_NDK_HOME" >&2 + exit 1 +} +rustup target list --installed | grep -qx "$TARGET" || rustup target add "$TARGET" + +adb devices | grep -qw device || { + echo "no device attached — plug the tablet in and enable USB debugging" >&2 + exit 1 +} + +echo "building onnx_probe for $TARGET…" +CARGO_TARGET_AARCH64_LINUX_ANDROID_LINKER="$clang" \ +CC_aarch64_linux_android="$clang" \ + cargo build -q -p dr-segment --example onnx_probe --release --target "$TARGET" \ + --no-default-features --features semantic +binary="${CARGO_TARGET_DIR:-target}/$TARGET/release/examples/onnx_probe" + +echo "pushing $(basename "$binary") and $(basename "$MODEL")…" +adb push "$binary" "$DEST/onnx_probe" >/dev/null +adb push "$MODEL" "$DEST/probe.onnx" >/dev/null +adb shell chmod 755 "$DEST/onnx_probe" +adb shell "$DEST/onnx_probe" "$DEST/probe.onnx" "$SHAPE" +adb shell rm -f "$DEST/onnx_probe" "$DEST/probe.onnx"