tools/quantise-models.sh writes the QDQ form QNN's HTP backend takes whole: opset 17, per-channel int8 weights, uint8 activations, ranges from running the f32 graph over photographs fed exactly as the app feeds them. The calibration is strided, four images at a time, because every ONNX Runtime calibrator holds each image's whole set of activations until it folds them — a gigabyte an image on the 10g detector, and an OOM kill with no message when folded once at the end. Release-time, never on the device (docs/inference.md §5): it needs real photographs and a person reading the recall measurement that gates whether each file is offered.
913 KiBLFS
913 KiBLFS