Run each model on the Hexagon in the form measured to hold it
The engine knew f32 and int8, and gave the Hexagon int8 for every role it served. Measured on the tablet itself (inference.md §1.5), int8 lost 5% of the detector's faces at 40-80 px, moved the landmarks 1.5 px, emptied the segmenter's scores and cost the denoiser 5-9 dB; fp16 the HTP refuses outright. `Form` gains A16W8 and A16W16, and `Rung::form` now names one per role: detectors and landmarks A16W8, the segmenter, scene model, border filler and denoiser A16W16, XFeat int8. The embedder and the eye classifiers stay on the CPU. Each loader resolves its `<stem>.<form>.onnx` sibling; the segmenter and XFeat, compiled into the binary, embed their quantised forms on Android only and pick through `choose_embedded`. The probe, the compile step and the cache fingerprint follow the form instead of assuming int8. Detectors on the new form write `scrfd_*_a16+w600k_mbf`, and `model_ids` answers for all three spellings. On the tablet (ORT 1.29 + QNN 2.42), each shipped file against f32 on the same inputs, and against the CPU's f32 time: SCRFD 500m/2.5g/10g A16W8 100% of faces in every band 4.2/5.1/9.0 ms vs 17/56/198 landmarks A16W8 0.25 px in the 192 crop 0.5 ms vs 2.8 YOLO26n-seg A16W16 98.2% found, mask IoU 0.994 12.9 ms vs 90 scene model A16W16 98.9% of cells agree 15 ms vs 151 MI-GAN A16W16 41 dB from f32 in the fill 87 ms vs 488 XFeat int8 pano alignment 0.45 px (f32's own spread 0.41) 6.5 ms vs 58 denoiser A16W16 0.00 dB at every ISO 95 ms vs 1510 a tile Face numbers are over public COCO val2017 photographs, not a library. The APK carries the siblings (BUNDLED 15 -> 19; the old int8 detectors removed), about 43 MB more. The Windows installer and its CI count skip them; the Arch and Flatpak packages list their files and never had them. The ladder example takes a role per model, which is how the per-role forms above were seen landing on the NPU from the real probe.
This commit is contained in:
@@ -331,6 +331,17 @@ If neither holds S's quality within 0.5 dB of fp32 on the real pairs, **v1 is de
|
||||
tablet shows the classical path. The sidecar still records the intent, so a desktop can render the
|
||||
learned result for a photograph edited on the tablet.
|
||||
|
||||
**Measured 2026-10-04 (inference.md §1.5): the second way holds, without the first.** The shipped
|
||||
network, with its Bayer packing re-spelled as `SpaceToDepth` so QNN can hold it (the 6-D reshape
|
||||
it replaces is exact but past the HTP's rank limit), at A16W16 — 16-bit activations and weights —
|
||||
scores within 0.00 dB of f32 at ISO 400–25600 on the tablet's own HTP, and within 0.09 dB with the
|
||||
6D's noise model scaled ×0.5, ×2 and ×4 to stand in for other sensors. A16W8 holds the 6D (worst
|
||||
−0.19 dB at ISO 25600) but not ×4 noise at 25600 (−0.52 dB), so A16W16 is what ships. int8 loses
|
||||
4.7–9.2 dB and fp16 is refused outright. A 1408 tile takes 95 ms on the Hexagon against 1510 ms on
|
||||
the tablet's CPU: about 2.3 s for a 20 MP frame. Calibration ranges come from 96 training-day
|
||||
tiles across every ISO, a third of them with that scaled noise; coverage of other bodies is that
|
||||
synthetic bracket, not their raws.
|
||||
|
||||
## 9. X-Trans
|
||||
|
||||
The requirements tie this stage to FR-RAW-5, and the library has no Fuji raws. What we can do
|
||||
|
||||
+64
-4
@@ -127,7 +127,8 @@ Three things the table settles.
|
||||
five of six models. A whole-library face index on the tablet goes from ~100 ms + 39 ms per face
|
||||
to ~1.4 ms + 12 ms per face, and the "Thorough" detector — 3× the cost of "Fast" today — becomes
|
||||
free. Its price is that the models must be **quantised to int8**, which is an accuracy question
|
||||
§5 has to answer before it is believed.
|
||||
§5 has to answer before it is believed. (§1.5 answered it: int8 lost faces, and every model but
|
||||
XFeat ships with 16-bit activations, at about three times these timings.)
|
||||
- **The embedder does not gain from either accelerator.** 112×112 input, per-op overhead
|
||||
dominates; it is 9 ms on the tablet's CPU and 12 ms on its NPU. It stays float, which §7 turns
|
||||
from a performance footnote into a correctness rule.
|
||||
@@ -137,6 +138,60 @@ Three things the table settles.
|
||||
- **On AMD, MIGraphX fp16 is 4–17× the CPU provider** on the detectors and 60× on the
|
||||
inpainter, with the same first-run compile cost as TensorRT and no rung between it and the CPU.
|
||||
|
||||
### 1.5 The Hexagon at every bit width · 2026-10-04
|
||||
|
||||
§1.1's Hexagon column is int8 calibrated on noise: timing only. This is the follow-up — every
|
||||
model, every bit width the HTP offers, calibrated on real photographs and **scored on the tablet
|
||||
itself** (ORT 1.29 + QNN 2.42, `htp_arch` 73), against the f32 model on the same inputs. The
|
||||
"Form shipped" column is the files in `models/`, re-scored on the tablet after
|
||||
`tools/quantise-models.sh` wrote them. The
|
||||
calibration and scoring photographs are 800 from the public COCO val2017 set (CC-BY); the face
|
||||
models' numbers are over the faces in them of at least 32 px. The tools are `tools/quantise-models.sh`
|
||||
and the scratch harness described with it.
|
||||
|
||||
**What the HTP accepts.** fp16: nothing — every fp16 operator fails validation (3110), on QNN
|
||||
2.42 and 2.50, with `htp_arch` and every `soc_model` tried; the fp16 rung stays off the table until
|
||||
someone has Qualcomm's own SDK to say why. 4-bit weights (A8W4, A16W4): load, and wreck accuracy
|
||||
(SCRFD finds 25–35% of f32's faces). What is left: **A8W8 (int8), A16W8 and A16W16**, all running
|
||||
the whole graph. 16-bit activations cost about 3× int8's time, A16W16 about 4×.
|
||||
|
||||
| Model | ORT CPU f32 | Form shipped | Hexagon | On the tablet, against f32 | int8 for comparison |
|
||||
|---|---|---|---|---|---|
|
||||
| scrfd_500m / 2.5g / 10g | 17 / 56 / 198 ms | **A16W8** | 4.2 / 5.1 / 9.0 ms | 100% of faces found in every size band; keypoints 0.3–0.6% of the box | 94–95% of faces at 40–80 px |
|
||||
| 2d106det (landmarks) | 2.8 ms | **A16W8** | 0.5 ms | 0.25 px in the 192 crop (eye points 0.20) | 1.5 px, and 29 partitions at 7.3 ms |
|
||||
| yolo26n-seg | 90 ms | **A16W16**, tail in float | 12.9 ms | 98.2% of objects, mask IoU 0.994 | 74% (simulated) |
|
||||
| yolo26s-sem-ade20k | 151 ms | **A16W16**, attention in float | 15 ms | 98.9% of cells agree on the class, TV 0.009 | 67% |
|
||||
| migan-512 | 488 ms | **A16W16** | 87 ms | 41 dB from f32 in the fill (worst 1%: 30 dB) | 16 dB (simulated) |
|
||||
| xfeat-1024 / 768 | 58 ms | **int8**, rewritten graph | 6.5 ms | panorama alignment 0.45 px from f32's — f32's own refit on 90% of its matches is 0.41 | — |
|
||||
| mosaic-1408 (denoiser) | 1510 ms a tile | **A16W16**, rewritten graph | 95 ms a tile | 0.00 dB at every ISO; ≤ 0.09 dB with the noise scaled ×0.5–×4 | −4.7 to −9.2 dB |
|
||||
| arcface_mbf (embedder) | 8.5 ms | f32, CPU | — | A16W16: cosine 0.9995, p1 0.9967 — misses §7's 0.999 gate | — |
|
||||
| ocec, sgc (eyes) | 1, 1.7 ms | f32, CPU | — | sgc flips 1.45% of views even at A16W16; not worth a millisecond | — |
|
||||
|
||||
Four things the table needed that the f32 graphs did not have, all in `tools/htp_graph.py` and all
|
||||
checked exact against the f32 graph before they are used:
|
||||
|
||||
- **Rank ≤ 5.** QNN's tensors stop at rank 5, and the denoiser packs the mosaic through a 6-D
|
||||
reshape (6007 at compose). For one channel that reshape is `SpaceToDepth(2)`. XFeat's 8×8 unfold
|
||||
is 224 Slices and 6-D Concats; it is `SpaceToDepth(8)` (736 nodes to 60).
|
||||
- **No bilinear Resize at XFeat's sizes** (3110). A half-pixel bilinear resize between fixed sizes
|
||||
is two constant matrices, so it is two MatMuls.
|
||||
- **One scale per tensor.** The segmenter's output rows carry boxes in pixels beside scores in
|
||||
0..1; quantised as one tensor the scores vanish. Everything from the Concat that builds the rows
|
||||
stays float, on the CPU, where the top-300 selection costs nothing.
|
||||
- **Float where the HTP's 16-bit arithmetic drifts.** The scene model's one attention block
|
||||
(two MatMuls and a Softmax over 400 tokens) moved its agreement from 98.7% to 96.7%; it stays
|
||||
float.
|
||||
|
||||
**ORT's CPU simulation of a QDQ graph is not the tablet.** It matched to the hundredth of a dB for
|
||||
the denoiser and to rounding for the detectors, landmarks and XFeat, and it overstated MI-GAN by
|
||||
27 dB and the scene model by three points. Every number above is the device's; a new form is not
|
||||
measured until it has run there.
|
||||
|
||||
**XFeat's int8 loses keypoints and not the panorama.** 83% of f32's keypoints come back within
|
||||
1.5 px; but over the twelve-frame `fixtures/pano/2025-08-05` sweep, the homographies fitted from
|
||||
int8's matches land 0.45 px from f32's in the overlaps — the spread f32 shows against itself
|
||||
(0.41).
|
||||
|
||||
---
|
||||
|
||||
## 2. The shape of the answer
|
||||
@@ -146,7 +201,7 @@ winning:
|
||||
|
||||
| Platform | 1st | 2nd | 3rd | Floor |
|
||||
|---|---|---|---|---|
|
||||
| Android, Qualcomm with a Hexagon the shipped QNN skel covers (V68–V81) | QNN HTP, int8 model | ORT CPU, f32 model | — | tract |
|
||||
| Android, Qualcomm with a Hexagon the shipped QNN skel covers (V68–V81) | QNN HTP, each model's quantised form (§1.5) | ORT CPU, f32 model | — | tract |
|
||||
| Android, any other SoC | ORT CPU, f32 | — | — | tract |
|
||||
| Linux / Windows, NVIDIA GPU | TensorRT, f32 model, fp16 engine | CUDA provider, f32 | ORT CPU, f32 | tract |
|
||||
| Linux, AMD GPU with ROCm | MIGraphX, f32 model, fp16 program | ORT CPU, f32 | — | tract |
|
||||
@@ -298,10 +353,10 @@ of which form they load:
|
||||
| Form | Who produces it | When | Needed by |
|
||||
|---|---|---|---|
|
||||
| f32 ONNX, shape-fixed, **opset ≥ 13** | `tools/fix-face-model-shapes.sh`, `tools/export-seg-model.sh` | Release time, once | Every rung except Hexagon |
|
||||
| int8 QDQ ONNX, per-channel, uint8 activations | `tools/quantise-models.sh` (new) | Release time, once, **calibrated on real photographs** | Hexagon |
|
||||
| QDQ ONNX, per-channel — int8, A16W8 or A16W16 per model (§1.5) | `tools/quantise-models.sh` | Release time, once, **calibrated on real photographs**, scored on the tablet | Hexagon |
|
||||
| TensorRT engine (`.engine`, per GPU architecture and TensorRT version) | The app, from the f32 file | First run on that device, in the background | TensorRT rung |
|
||||
| MIGraphX program (`.mxr`, per GPU architecture, MIGraphX version and precision) | The app, from the f32 file | First run on that device, in the background | MIGraphX rung |
|
||||
| QNN context binary | The app, from the int8 file | First run on that device, in the background | Hexagon rung |
|
||||
| QNN context binary | The app, from the quantised file | First run on that device, in the background | Hexagon rung |
|
||||
|
||||
Two rules.
|
||||
|
||||
@@ -380,6 +435,11 @@ already the rule for the detector and because §5 is the gate on whether the int
|
||||
enough to be *offered* at all. f32 on tract, ORT CPU, CUDA and TensorRT-f32 are one identity: the
|
||||
same graph, the same arithmetic, differences at the last bit.
|
||||
|
||||
After §1.5 the Hexagon runs the detectors in **A16W8**, and that is a third spelling:
|
||||
`scrfd_500m_a16+w600k_mbf` and its two siblings. Same rule, same reconciliation; a tablet that
|
||||
indexed under `_i8` keeps those rows, and `FaceDetector::model_ids` answers "has this detector been
|
||||
over this image" for all three forms.
|
||||
|
||||
**The embedder** is where comparability across devices is the whole point, and it is the one
|
||||
model that no accelerator helps (§1.4). So: **the embedder runs in f32 on every rung.** On TensorRT
|
||||
that means the embedder's engine is built without fp16 while the detector's is built with it; on
|
||||
|
||||
@@ -467,9 +467,11 @@ the tablet. Three ways to make it viable, none built:
|
||||
1. **Fill at a quarter of the resolution and upsample.** Sky and scree
|
||||
tolerate it; twenty-odd tiles, about three minutes on the desktop CPU. A
|
||||
background job with the outbox's patience, not an interactive one.
|
||||
2. **int8 on the tablet's Hexagon through QNN**, where the plain-conv design
|
||||
is the point and the whole graph should run in milliseconds. The setup
|
||||
exists from the eye-state work; MI-GAN is a candidate for the same path.
|
||||
2. **The tablet's Hexagon through QNN**, where the plain-conv design is the
|
||||
point. Measured 2026-10-04 (inference.md §1.5): int8 changes the fill
|
||||
(16 dB from f32's), so it ships with 16-bit activations and weights —
|
||||
87 ms a tile against 488 ms on the tablet's CPU, the whole graph on the
|
||||
NPU, 41 dB from f32 in the hole.
|
||||
3. **A WGSL runtime for those six operators.** A project of its own, and
|
||||
the only route that would make it interactive on the desktop.
|
||||
|
||||
|
||||
+10
-10
File diff suppressed because one or more lines are too long
+2
-2
@@ -287,7 +287,6 @@ $LOCALAPPDATA\Programs\DarkRoom\
|
||||
models\
|
||||
scrfd_500m_640.onnx scrfd_2.5g_640.onnx scrfd_10g_640.onnx arcface_mbf_b1.onnx
|
||||
2d106det_b1.onnx ocec_s_b1.onnx sgc_l_48_b1.onnx
|
||||
scrfd_500m_640.int8.onnx scrfd_2.5g_640.int8.onnx scrfd_10g_640.int8.onnx
|
||||
yolo26s-sem-ade20k.onnx yolo26s-sem-ade20k.classes.json categories.txt
|
||||
migan-512.onnx
|
||||
manual\
|
||||
@@ -297,7 +296,8 @@ $LOCALAPPDATA\Programs\DarkRoom\
|
||||
```
|
||||
|
||||
Plus a Start Menu shortcut, and nothing on the desktop unless the user ticks it. The models are
|
||||
the same ten files the APK bundles and the PKGBUILD installs; `models\` beside the executable is
|
||||
the f32 files the APK bundles and the PKGBUILD installs — not the APK's quantised siblings, which
|
||||
only a Hexagon runs; `models\` beside the executable is
|
||||
where §3.2's lookup finds them. **No `LICENSE` yet**: the repository has no licence file at its
|
||||
root (the Arch package points at the system's shared GPL text), so the installer has no licence
|
||||
page until one is added — a one-file change, and the `.nsi` says where the page then goes. The face weights carry the research-only grant that
|
||||
|
||||
Reference in New Issue
Block a user