Files
DarkRoom/tools/fix-face-model-shapes.sh
T
dtourolle 4f31123b0c Let the user choose which SCRFD finds their faces
faces.md §12.3 measured what the cheapest detector costs: the small
faces in every group shot, and a dog embedded a dozen times. Which
trade is right depends on the machine doing the sweep — a desktop left
overnight and a tablet on a battery want different answers — so the
detector is now a per-device setting, Fast / Balanced / Thorough on
the settings page beside the indexing button, persisted with the rest
of the settings file.

A detector is half of a model id. Every face, marker, shard and
calibration is keyed on faces.model_id precisely so that a model change
is a new id and a re-index rather than a silent change under existing
data, and a detector change is a model change: it decides which faces
exist and where the landmarks that align them land. So each choice
names its own pipeline. 500M keeps the bare "w600k_mbf" every existing
library was written under, so an upgrade disturbs nothing; the others
are qualified. Choosing one restarts coverage from zero under the new
id, the sweep re-detects, confirmed names carry across by box overlap,
and the sync shards are keyed by the same id so a peer on another
setting neither adopts nor pollutes them. The library controller
carries the id into the sync the same way it carries the cache budget,
because the sync starts from places that have no settings in reach.

All three shape-fixed exports ship — APK, Arch, Flatpak — since a
tablet has no other way to obtain the one it was not installed with;
the APK grows by twenty megabytes for the choice.
2026-09-11 22:12:53 +02:00

112 lines
4.3 KiB
Bash
Executable File

#!/usr/bin/env bash
# Freeze the input dimensions of the face models so tract can parse them.
#
# ./tools/fix-face-model-shapes.sh IN.onnx OUT.onnx --input NAME=1,3,640,640
# ./tools/fix-face-model-shapes.sh IN.onnx OUT.onnx --dim NAME=1
#
# The two the face pipeline needs, verified 2026-08-26 (docs/faces.md §12 M1):
#
# ... det_500m.onnx scrfd_500m_640.onnx --input input.1=1,3,640,640
# ... det_2.5g.onnx scrfd_2.5g_640.onnx --input input.1=1,3,640,640
# ... det_10g.onnx scrfd_10g_640.onnx --input input.1=1,3,640,640
# ... w600k_mbf.onnx arcface_mbf_b1.onnx --dim None=1
#
# ## Why this exists
#
# The InsightFace exports declare dynamic input dimensions — SCRFD's H and W,
# ArcFace's batch N. **tract cannot parse either graph in that form**, failing
# at the input node and at the first Conv respectively:
#
# scrfd_500m_bnkps.onnx Translating node #0 "input.1" Source ToTypedTranslator
# arcface_w600k_mbf.onnx Failed analyse for node #139 "Conv_0" ConvHir
#
# Both load cleanly once the dims are pinned. This is the same wall dr-segment
# hit, which is why `tools/export-seg-model.sh` passes `dynamic=False`; here we
# cannot re-export from PyTorch, because the weights are InsightFace's and the
# training code is not in the loop, so the dims are rewritten in the ONNX file
# instead.
#
# `make_dynamic_shape_fixed` only edits the declared dimension; it does not
# retrain, requantise, or change a single weight. The output is numerically the
# same graph with one shape pinned. SCRFD's *outputs* were already static — the
# export was made at 640 and only its input forgot to say so — which is why 640
# is not a free choice here.
#
# ## Why it is a script and not a build step
#
# Same reason as the segmentation export: the model is not a build input
# (docs/faces.md §2.2 — the weights are never committed, because InsightFace's
# grant is non-commercial). This runs once, wherever the user's model lives,
# and the app loads the result. It exists so the transformation is reproducible
# rather than a binary someone once produced and nobody can regenerate.
#
# Requires `uv`. Everything else is fetched into a throwaway venv, in /var/tmp
# rather than /tmp — /tmp here is a tmpfs, and onnxruntime is not small.
set -euo pipefail
if [ "$#" -lt 4 ]; then
sed -n '2,10p' "$0" >&2
exit 2
fi
IN="$1"; shift
OUT="$1"; shift
[ -f "$IN" ] || { echo "no such model: $IN" >&2; exit 1; }
WORK="$(mktemp -d -p /var/tmp fix-face-shapes.XXXXXX)"
trap 'rm -rf "${WORK}"' EXIT
echo "==> venv in ${WORK}"
uv venv --python 3.12 "${WORK}/venv" >/dev/null
VIRTUAL_ENV="${WORK}/venv" uv pip install --quiet onnx onnxruntime
# Two forms, because the two models need different ones — and which one a graph
# needs is not a matter of taste:
#
# --dim NAME=VALUE for a *named* symbolic dimension.
# --input NAME=D,D,D,D for a dimension that is dynamic but unnamed.
#
# ArcFace declares its batch as the literal dim_param "None", so `--dim` binds
# it. SCRFD's H and W carry no dim_param at all, so there is no name to bind
# and the whole input shape has to be restated. Reaching for `--dim` first and
# getting a silent no-op is the half-hour worth skipping.
CUR="$IN"
STEP=0
while [ "$#" -gt 0 ]; do
FLAG="$1"; shift
PAIR="${1:-}"; shift || true
NAME="${PAIR%%=*}"
VAL="${PAIR#*=}"
STEP=$((STEP + 1))
NEXT="${WORK}/step${STEP}.onnx"
case "$FLAG" in
--dim)
echo "==> dim_param ${NAME} := ${VAL}"
"${WORK}/venv/bin/python" -m onnxruntime.tools.make_dynamic_shape_fixed \
--dim_param "${NAME}" --dim_value "${VAL}" "${CUR}" "${NEXT}"
;;
--input)
echo "==> input ${NAME} := ${VAL}"
"${WORK}/venv/bin/python" -m onnxruntime.tools.make_dynamic_shape_fixed \
--input_name "${NAME}" --input_shape "${VAL}" "${CUR}" "${NEXT}"
;;
*)
echo "unknown flag ${FLAG} (want --dim or --input)" >&2
exit 2
;;
esac
CUR="${NEXT}"
done
cp "${CUR}" "${OUT}"
echo "==> wrote ${OUT}"
"${WORK}/venv/bin/python" - "$OUT" <<'PY'
import sys, onnx
m = onnx.load(sys.argv[1])
for vi in list(m.graph.input) + list(m.graph.output):
dims = [d.dim_value if d.HasField("dim_value") else (d.dim_param or "?")
for d in vi.type.tensor_type.shape.dim]
print(f" {vi.name:<24} {dims}")
PY