tract runs every model on one core on every platform. Measured against
ONNX Runtime's providers on the MagicPad 2 and the reference desktop:
ORT CPU alone is 3-10x, the Hexagon at int8 runs the detectors in
1-3 ms, TensorRT is ~2x the CUDA provider. NNAPI, XNNPACK, WebGPU and
CUDA int8 were tried and excluded with the numbers that excluded them.
The spec keeps the build C-free: ort::set_api takes a table from a
dlopened runtime or from ort-tract, chosen once per process. Rungs
are chosen by building a real session, cached until an input changes,
and compiled engines are built in the background after the first
frame. The embedder stays f32 everywhere; int8 detectors are a
distinct model_id and are gated on a recall measurement.