Record what the spec got wrong about the model that exists
docs/segmentation.md §4 priced arm B as costing a C dependency under the NDK and treated that as most of the difference between the arms. It is not a cost that has to be paid: `ort`'s `alternative-backend` disables its linking entirely and `ort-tract` supplies the API from tract, which is pure Rust. D13's "largest exception the policy would tolerate" turns out not to be needed, and the answer generalises to the face pipeline — so D13's runtime half is now answered and only its licensing half is open. Three findings contradict §4 outright and are recorded as F4-F6 rather than quietly designed around. There is no ADE20K-trained YOLO, so the shipped vocabulary selects subjects and not stuff — "select the sky" comes from the watershed or from nowhere. It is instance segmentation, so it partitions nothing and two people come back as two instances. And tract cannot parse a dynamic-shape export, which fixes the input at 640 square and makes tiling the only route to more semantic resolution. Arm C ships, but §8's criteria are not what decided it, and saying so matters more than claiming the process worked. §8 asked for a two- interaction margin over arm A on a traced corpus. That comparison was never run: F4 and F5 changed what the arms are, and a model that recognises subjects but has no word for sky cannot be a selection tool alone, while a watershed cannot tell a person from the wall behind them. They stopped being candidates and became complements. What is *not* done is written down as plainly: the 24-image corpus is untraced, so M1-M4 have no numbers and "this feels right" has not become one. M5 is answered on one device only, and region ids now reach the sidecar — so a cross-vendor divergence would mean a mask written on the desktop meaning something else on Android. F3 stands.
This commit is contained in:
+13
-1
@@ -1192,7 +1192,19 @@ both a smaller build and a usable one — it needs no develop chain.
|
||||
|
||||
Resolving D12 sets D3 and [architecture.md §10](architecture.md)'s Phase 2.
|
||||
|
||||
### D13 — face inference runtime and model licensing · **OPEN**
|
||||
### D13 — face inference runtime and model licensing · **RUNTIME ANSWERED, LICENSING OPEN**
|
||||
|
||||
> **Updated 2026-08-21.** The runtime half of this decision is settled, and by a route the table
|
||||
> below does not contain. `ort` 2.0's `alternative-backend` feature *disables its linking entirely*
|
||||
> and lets another engine supply the `OrtApi`; `ort-tract` supplies it from `tract`, which is pure
|
||||
> Rust. So the third option's operator coverage comes with the first option's dependency profile —
|
||||
> no C, no NDK problem, no exception to the policy. Measured on a real graph before being relied on:
|
||||
> YOLO26n-seg loads with zero unsupported operators and runs 640×640 in ~470 ms of CPU
|
||||
> (docs/segmentation.md §13). §3.9.1's detector and embedder are different graphs and their coverage
|
||||
> has not been checked, but the *approach* no longer needs a decision.
|
||||
>
|
||||
> **The licensing half is untouched.** The InsightFace weights are still non-commercial and still
|
||||
> unusable here. That remains what S14 has to resolve first.
|
||||
|
||||
§3.9.1 needs to run two neural networks locally. That collides with two settled positions, and
|
||||
neither collision is small enough to leave implicit.
|
||||
|
||||
Reference in New Issue
Block a user