S15.2: XFeat exports at a fixed shape and loads under tract
tools/export-xfeat.sh exports the convolutional network alone at 768×1024 grayscale, on the pattern of export-seg-model.sh: thirteen standard operator types, no dynamic axes, the keypoint decoding left to Rust. examples/onnx_probe loads it through the ort-over-tract backend the app ships with nothing unsupported and runs it in ~300 ms on the desktop CPU. The weights are Apache-2.0, read from the repository's LICENSE, with no grant on the checkpoint — recorded in models/LICENCE.md before they land, as FR-MRG-8 asks. The probe stays: the next model will need the same check.
This commit is contained in:
@@ -139,6 +139,27 @@ first (D13's lesson, S15.2).
|
||||
| SuperPoint, SuperGlue, R2D2, SiLK, MASt3R | non-commercial | — | Out on licence |
|
||||
| LightGlue | Apache-2.0 | Transformer over a variable keypoint count | Not until mutual-nearest-neighbour matching fails on a real set |
|
||||
|
||||
**S15.2, 2026-09-19: XFeat loads under tract.** `tools/export-xfeat.sh`
|
||||
exports the network alone at 768×1024 — thirteen operator types, all
|
||||
standard: `Conv`, `InstanceNormalization`, `AveragePool`, `Resize`, `Slice`,
|
||||
`Transpose`, `Reshape`, `Concat`, `Add`, `Relu`, `Sigmoid`, `ReduceMean`,
|
||||
`Unsqueeze` — and
|
||||
[`examples/onnx_probe.rs`](../core/dr-segment/examples/onnx_probe.rs) loads
|
||||
the 2.8 MB file through the app's own `ort`-over-tract backend with nothing
|
||||
unsupported, in 28 ms, and runs it in **~300 ms on the reference desktop's
|
||||
CPU**. The weights ship as `models/keypoints/xfeat-1024.onnx`, recorded in
|
||||
`models/LICENCE.md`. Still to do: the tablet figure (S15.4), and a
|
||||
keypoint-level comparison against the PyTorch reference once the Rust decoder
|
||||
exists — the probe proves the graph runs, not that the numbers match.
|
||||
|
||||
The outputs are three maps at 1/8 resolution, 96×128 for the export size:
|
||||
64-channel descriptors, 65-channel keypoint logits (each 8×8 cell's position
|
||||
plus "none"), and a reliability heatmap. The Rust decoder is: softmax over the
|
||||
65, pixel-shuffle the first 64 to full resolution, 5×5 non-maximum
|
||||
suppression, top-k by reliability, bilinear sampling of the descriptor at
|
||||
each keypoint, L2 normalise. That is `detectAndCompute` in the reference,
|
||||
minus the network.
|
||||
|
||||
Without weights: AKAZE (BSD, `akaze` from rust-cv), which is adequate on
|
||||
well-textured overlaps and worse on sky, repeated structure and exposure
|
||||
drift — which is where a learned detector earns its place.
|
||||
|
||||
Reference in New Issue
Block a user