S15.4, CPU half: XFeat runs in ~400 ms per frame on the tablet

tools/onnx-probe-on-device.sh cross-builds dr-segment's onnx_probe
without the embedded segmentation model, pushes it with a model to the
attached device and times two runs. The 768×1024 XFeat export takes
~400 ms on the reference tablet's NEON cores against ~300 ms on the
desktop, with identical output ranges — inside NFR-MRG-1's 1 s per frame.
The blend half of S15.4 waits for a chunked blend to exist.
This commit is contained in:
2026-09-19 15:24:12 +02:00
parent 5bf06c5030
commit 2bf0ec8dba
3 changed files with 71 additions and 5 deletions
+8 -3
View File
@@ -183,9 +183,14 @@ standard: `Conv`, `InstanceNormalization`, `AveragePool`, `Resize`, `Slice`,
the 2.8 MB file through the app's own `ort`-over-tract backend with nothing
unsupported, in 28 ms, and runs it in **~300 ms on the reference desktop's
CPU**. The weights ship as `models/keypoints/xfeat-1024.onnx`, recorded in
`models/LICENCE.md`. Still to do: the tablet figure (S15.4), and a
keypoint-level comparison against the PyTorch reference once the Rust decoder
exists — the probe proves the graph runs, not that the numbers match.
`models/LICENCE.md`. **On the tablet** (S15.4's CPU half, same day):
`tools/onnx-probe-on-device.sh` cross-builds the probe, and the same file
runs in **~400 ms per frame** on the reference tablet's NEON cores (ROD2-W09,
SM8635), with output ranges identical to the desktop's — inside NFR-MRG-1's
1 s per frame with room to spare, and 1.3× the desktop rather than the 2×
faces.md §9 measured for its scan. Still to do: a keypoint-level comparison
against the PyTorch reference once the Rust decoder exists — the probe
proves the graph runs, not that the numbers match.
The outputs are three maps at 1/8 resolution, 96×128 for the export size:
64-channel descriptors, 65-channel keypoint logits (each 8×8 cell's position