Record what the spec got wrong about the model that exists

docs/segmentation.md §4 priced arm B as costing a C dependency under the
NDK and treated that as most of the difference between the arms. It is not
a cost that has to be paid: `ort`'s `alternative-backend` disables its
linking entirely and `ort-tract` supplies the API from tract, which is pure
Rust. D13's "largest exception the policy would tolerate" turns out not to
be needed, and the answer generalises to the face pipeline — so D13's
runtime half is now answered and only its licensing half is open.

Three findings contradict §4 outright and are recorded as F4-F6 rather than
quietly designed around. There is no ADE20K-trained YOLO, so the shipped
vocabulary selects subjects and not stuff — "select the sky" comes from the
watershed or from nowhere. It is instance segmentation, so it partitions
nothing and two people come back as two instances. And tract cannot parse a
dynamic-shape export, which fixes the input at 640 square and makes tiling
the only route to more semantic resolution.

Arm C ships, but §8's criteria are not what decided it, and saying so
matters more than claiming the process worked. §8 asked for a two-
interaction margin over arm A on a traced corpus. That comparison was never
run: F4 and F5 changed what the arms are, and a model that recognises
subjects but has no word for sky cannot be a selection tool alone, while a
watershed cannot tell a person from the wall behind them. They stopped
being candidates and became complements.

What is *not* done is written down as plainly: the 24-image corpus is
untraced, so M1-M4 have no numbers and "this feels right" has not become
one. M5 is answered on one device only, and region ids now reach the
sidecar — so a cross-vendor divergence would mean a mask written on the
desktop meaning something else on Android. F3 stands.
This commit is contained in:
2026-08-22 08:39:17 +02:00
parent 9b4f0815e5
commit 12d320cf33
6 changed files with 395 additions and 115 deletions
+13 -1
View File
@@ -1192,7 +1192,19 @@ both a smaller build and a usable one — it needs no develop chain.
Resolving D12 sets D3 and [architecture.md §10](architecture.md)'s Phase 2.
### D13 — face inference runtime and model licensing · **OPEN**
### D13 — face inference runtime and model licensing · **RUNTIME ANSWERED, LICENSING OPEN**
> **Updated 2026-08-21.** The runtime half of this decision is settled, and by a route the table
> below does not contain. `ort` 2.0's `alternative-backend` feature *disables its linking entirely*
> and lets another engine supply the `OrtApi`; `ort-tract` supplies it from `tract`, which is pure
> Rust. So the third option's operator coverage comes with the first option's dependency profile —
> no C, no NDK problem, no exception to the policy. Measured on a real graph before being relied on:
> YOLO26n-seg loads with zero unsupported operators and runs 640×640 in ~470 ms of CPU
> (docs/segmentation.md §13). §3.9.1's detector and embedder are different graphs and their coverage
> has not been checked, but the *approach* no longer needs a decision.
>
> **The licensing half is untouched.** The InsightFace weights are still non-commercial and still
> unusable here. That remains what S14 has to resolve first.
§3.9.1 needs to run two neural networks locally. That collides with two settled positions, and
neither collision is small enough to leave implicit.