S15.4, CPU half: XFeat runs in ~400 ms per frame on the tablet
tools/onnx-probe-on-device.sh cross-builds dr-segment's onnx_probe without the embedded segmentation model, pushes it with a model to the attached device and times two runs. The 768×1024 XFeat export takes ~400 ms on the reference tablet's NEON cores against ~300 ms on the desktop, with identical output ranges — inside NFR-MRG-1's 1 s per frame. The blend half of S15.4 waits for a chunked blend to exist.
This commit is contained in:
+8
-3
@@ -183,9 +183,14 @@ standard: `Conv`, `InstanceNormalization`, `AveragePool`, `Resize`, `Slice`,
|
||||
the 2.8 MB file through the app's own `ort`-over-tract backend with nothing
|
||||
unsupported, in 28 ms, and runs it in **~300 ms on the reference desktop's
|
||||
CPU**. The weights ship as `models/keypoints/xfeat-1024.onnx`, recorded in
|
||||
`models/LICENCE.md`. Still to do: the tablet figure (S15.4), and a
|
||||
keypoint-level comparison against the PyTorch reference once the Rust decoder
|
||||
exists — the probe proves the graph runs, not that the numbers match.
|
||||
`models/LICENCE.md`. **On the tablet** (S15.4's CPU half, same day):
|
||||
`tools/onnx-probe-on-device.sh` cross-builds the probe, and the same file
|
||||
runs in **~400 ms per frame** on the reference tablet's NEON cores (ROD2-W09,
|
||||
SM8635), with output ranges identical to the desktop's — inside NFR-MRG-1's
|
||||
1 s per frame with room to spare, and 1.3× the desktop rather than the 2×
|
||||
faces.md §9 measured for its scan. Still to do: a keypoint-level comparison
|
||||
against the PyTorch reference once the Rust decoder exists — the probe
|
||||
proves the graph runs, not that the numbers match.
|
||||
|
||||
The outputs are three maps at 1/8 resolution, 96×128 for the export size:
|
||||
64-channel descriptors, 65-channel keypoint logits (each 8×8 cell's position
|
||||
|
||||
@@ -1839,8 +1839,9 @@ touched by a resize. That is the list in the criterion, and it is the testable h
|
||||
|
||||
**NFR-MRG-1 — Merge latency.** For five 24 MP frames: the alignment preview (FR-MRG-7) within
|
||||
5 s on the reference desktop and 15 s on the reference tablet, of which keypoint detection is at
|
||||
most 1 s per frame on the tablet's CPU; the full merge written to disk within 60 s on the desktop.
|
||||
The tablet's full-merge figure is set by S15, not guessed here.
|
||||
most 1 s per frame on the tablet's CPU (measured 2026-09-19 at ~0.4 s, S15.4); the full merge
|
||||
written to disk within 60 s on the desktop. The tablet's full-merge figure is set by S15's blend
|
||||
measurement, not guessed here.
|
||||
|
||||
**Performance regressions fail the build.** §8's benchmark suite runs per-commit; a regression
|
||||
beyond a stated tolerance is a build failure, not a notification. Performance work rots otherwise.
|
||||
|
||||
Reference in New Issue
Block a user