S15.4, CPU half: XFeat runs in ~400 ms per frame on the tablet
tools/onnx-probe-on-device.sh cross-builds dr-segment's onnx_probe without the embedded segmentation model, pushes it with a model to the attached device and times two runs. The 768×1024 XFeat export takes ~400 ms on the reference tablet's NEON cores against ~300 ms on the desktop, with identical output ranges — inside NFR-MRG-1's 1 s per frame. The blend half of S15.4 waits for a chunked blend to exist.
This commit is contained in:
+8
-3
@@ -183,9 +183,14 @@ standard: `Conv`, `InstanceNormalization`, `AveragePool`, `Resize`, `Slice`,
|
||||
the 2.8 MB file through the app's own `ort`-over-tract backend with nothing
|
||||
unsupported, in 28 ms, and runs it in **~300 ms on the reference desktop's
|
||||
CPU**. The weights ship as `models/keypoints/xfeat-1024.onnx`, recorded in
|
||||
`models/LICENCE.md`. Still to do: the tablet figure (S15.4), and a
|
||||
keypoint-level comparison against the PyTorch reference once the Rust decoder
|
||||
exists — the probe proves the graph runs, not that the numbers match.
|
||||
`models/LICENCE.md`. **On the tablet** (S15.4's CPU half, same day):
|
||||
`tools/onnx-probe-on-device.sh` cross-builds the probe, and the same file
|
||||
runs in **~400 ms per frame** on the reference tablet's NEON cores (ROD2-W09,
|
||||
SM8635), with output ranges identical to the desktop's — inside NFR-MRG-1's
|
||||
1 s per frame with room to spare, and 1.3× the desktop rather than the 2×
|
||||
faces.md §9 measured for its scan. Still to do: a keypoint-level comparison
|
||||
against the PyTorch reference once the Rust decoder exists — the probe
|
||||
proves the graph runs, not that the numbers match.
|
||||
|
||||
The outputs are three maps at 1/8 resolution, 96×128 for the export size:
|
||||
64-channel descriptors, 65-channel keypoint logits (each 8×8 cell's position
|
||||
|
||||
Reference in New Issue
Block a user