Quantise for the Hexagon with QNN's config and the app's own inputs

tools/quantise-models.sh now writes each model's Hexagon form from a
per-model table: the form its role takes on the NPU (int8, A16W8 or
A16W16), the exact graph rewrites it needs, and the nodes that must stay
float. Ranges are min/max over photographs fed exactly as the app feeds
each model -- the detector and segmenter letterboxes with their own pads
and normalisation, landmark crops from the detector's boxes, MI-GAN with a
panorama-like border, XFeat's grey proxy. The old tool used an
antialiased resize, YOLO's pad of 128 and /255 for every model that was
not a face model, none of which is what the app does.

tools/htp_graph.py holds the rewrites, each checked against the input
graph before use: the denoiser's 6-D Bayer pack and XFeat's 224-slice
unfold as SpaceToDepth (QNN stops at rank 5), computed reshape targets
folded, and bilinear Resize as two MatMuls (the HTP refuses
ResizeBilinear at XFeat's sizes). The denoiser takes ranges computed by
darkroom-denoise's gate on a smaller tile of the same network.
This commit is contained in:
2026-10-04 03:45:07 -04:00
parent 948f6c3ed2
commit 0e6ac09fd5
3 changed files with 451 additions and 130 deletions
+22 -18
View File
@@ -1,30 +1,34 @@
#!/usr/bin/env bash
# Produce the int8 form of a model for the Hexagon (docs/dev/inference.md §5).
# Produce the Hexagon's form of each model (docs/dev/inference.md §1.5, §5).
#
# ./tools/quantise-models.sh PHOTO_DIR MODEL.onnx [MODEL.onnx ...]
# ./tools/quantise-models.sh PHOTO_DIR [MODEL ...]
# ./tools/quantise-models.sh --ranges RANGES.json mosaic-1408
#
# Writes `MODEL.int8.onnx` beside each input: a QDQ graph, per-channel int8
# weights, uint8 activations — the form QNN's HTP backend takes whole. The
# activations' ranges come from running the f32 model over the photographs in
# PHOTO_DIR, fed exactly as the app feeds them (letterboxed to the model's
# input, the detector's `(x - 127.5) / 128` normalisation), which is why
# this is a release-time step and not something the device does: it needs
# real photographs and, after it, a person reading §10 M2's numbers.
# Writes `<stem>.<form>.onnx` beside each canonical file under models/: a QDQ
# graph from QNN's own quantisation config, per-channel weights, in the form
# the engine's `Rung::form` names for that role — A16W8, A16W16 or int8, each
# the narrowest that held the model's accuracy on the tablet. The activation
# ranges come from running the f32 model over the photographs in PHOTO_DIR,
# fed exactly as the app feeds them (letterbox maths, pads, normalisation,
# face crops through the app's own similarity), which is why this is a
# release-time step and not something the device does. With no MODEL, every
# model in the table.
#
# The SCRFD and ArcFace exports are opset 11; per-channel QDQ needs 13, so a
# model below 13 is first upgraded to 17. That changes only the graph's
# spelling, not a weight — and it is what `tools/fix-face-model-shapes.sh`
# will do to the canonical files in the same model release.
# The denoiser is calibrated on noisy mosaics, not photographs: its ranges
# come from darkroom-denoise's precision gate (`--ranges`), computed on a
# smaller tile of the same network — activation ranges do not depend on the
# tile's size, and the tensor names match.
#
# Then measure before shipping: a quantised form is a different network, and
# the numbers in inference.md §1.5 are what each one had to hold.
#
# A venv per run, like fix-face-model-shapes.sh: the tools are not a build
# input and nothing in the tree should have them on its path.
set -euo pipefail
if [ "$#" -lt 2 ]; then
sed -n '2,20p' "$0" >&2
if [ "$#" -lt 1 ]; then
sed -n '2,27p' "$0" >&2
exit 2
fi
PHOTOS="$1"; shift
[ -d "${PHOTOS}" ] || { echo "no such directory: ${PHOTOS}" >&2; exit 1; }
WORK="$(mktemp -d -p /var/tmp quantise-models.XXXXXX)"
trap 'rm -rf "${WORK}"' EXIT
@@ -32,4 +36,4 @@ echo "==> venv in ${WORK}"
uv venv --python 3.12 "${WORK}/venv" >/dev/null
VIRTUAL_ENV="${WORK}/venv" uv pip install --quiet onnx onnxruntime pillow numpy sympy
exec "${WORK}/venv/bin/python" "$(dirname "$0")/quantise-models.py" "${PHOTOS}" "$@"
exec "${WORK}/venv/bin/python" "$(dirname "$0")/quantise-models.py" "$@"