Record what the spec got wrong about the model that exists

docs/segmentation.md §4 priced arm B as costing a C dependency under the
NDK and treated that as most of the difference between the arms. It is not
a cost that has to be paid: `ort`'s `alternative-backend` disables its
linking entirely and `ort-tract` supplies the API from tract, which is pure
Rust. D13's "largest exception the policy would tolerate" turns out not to
be needed, and the answer generalises to the face pipeline — so D13's
runtime half is now answered and only its licensing half is open.

Three findings contradict §4 outright and are recorded as F4-F6 rather than
quietly designed around. There is no ADE20K-trained YOLO, so the shipped
vocabulary selects subjects and not stuff — "select the sky" comes from the
watershed or from nowhere. It is instance segmentation, so it partitions
nothing and two people come back as two instances. And tract cannot parse a
dynamic-shape export, which fixes the input at 640 square and makes tiling
the only route to more semantic resolution.

Arm C ships, but §8's criteria are not what decided it, and saying so
matters more than claiming the process worked. §8 asked for a two-
interaction margin over arm A on a traced corpus. That comparison was never
run: F4 and F5 changed what the arms are, and a model that recognises
subjects but has no word for sky cannot be a selection tool alone, while a
watershed cannot tell a person from the wall behind them. They stopped
being candidates and became complements.

What is *not* done is written down as plainly: the 24-image corpus is
untraced, so M1-M4 have no numbers and "this feels right" has not become
one. M5 is answered on one device only, and region ids now reach the
sidecar — so a cross-vendor divergence would mean a mask written on the
desktop meaning something else on Android. F3 stands.
This commit is contained in:
2026-08-22 08:39:17 +02:00
parent 9b4f0815e5
commit 12d320cf33
6 changed files with 395 additions and 115 deletions
+87 -8
View File
@@ -345,14 +345,93 @@ measurement that can invalidate the region-id representation.
---
## 13. Register entries
## 13. Arms B and C, and what §4 got wrong
To be added when this is accepted:
Built 2026-08-21 on branch `local-adjustments`. `core/dr-segment`, `core/dr-gpu/src/mask.rs`,
`core/dr-pipeline/src/mask.rs`, and the develop panel.
**S15** — *Region segmentation for local masking:* implement arm A and the shared consumers, trace a
24-image corpus, measure M1–M7 across arms. **Resolve the licence question before writing arm B.**
Answers: which segmentation source local masking snaps to, and whether a mask can be stored as region
ids at all. Relates to: D13, D14, FR-DEV-3, ARCH §5.4, §6.13.
**The dependency question dissolved rather than being decided.** §4 and D13 both priced arm B as
costing a C dependency under the Android NDK, and treated that as most of the difference between the
arms. It is not a cost that has to be paid: `ort` 2.0's `alternative-backend` feature disables the
linking entirely and lets another engine supply the `OrtApi`, and `ort-tract` — same authors,
MIT/Apache — supplies it from `tract`, which is pure Rust. So arm B runs through `ort`'s API with no
C anywhere, and D13's "largest exception the policy would tolerate" turns out not to be needed.
**D14** — *Segmentation source for local masking* · **OPEN**. Decided by S15 against the criteria in
§8. Reopens D13's dependency-policy question if arm C wins.
Measured before committing to it, because tract's operator coverage is the thing that could have
sunk it: **yolo26n-seg loads with zero unsupported operators** and runs 640×640 in ~470 ms on the
reference desktop's CPU. Correct masks on the standard `bus.jpg` — one bus and three people, outlines
following the subjects.
Three findings that contradict §4 directly, and all three change the design rather than the schedule.
**F4 — there is no ADE20K-trained YOLO.** §4's whole argument for arm B was ADE20K's 150 classes and
their *stuff* categories: "'that patch of sky' is a class here." Checked 2026-08-21: Ultralytics ships
YOLO26-seg trained on **COCO**, whose 80 classes are all *things*, and the one HuggingFace repository
claiming a YOLO/ADE20K combination is empty. ADE20K models exist as SegFormer/OneFormer/MaskFormer
transformers, not as YOLO.
So the shipped vocabulary selects **subjects**, not **stuff**. "Select the person" works; "select the
sky" does not come from the model at all and must come from the watershed. That is a narrower arm B
than §4 assumed, and it *raises* the importance of arm C rather than lowering it — the model can no
longer be the whole answer for anything.
**F5 — it is instance segmentation, not semantic segmentation.** §4 assumed a flat partition with
class ids that would "populate `RegionField` directly". YOLO-seg does not partition the image; it
finds objects, and most pixels in a landscape belong to no instance. Two consequences, one bad and
one better than expected: nothing populates a `RegionField` on its own, and two people come back as
*two* instances where a semantic model would have returned one "person" area covering both. For
selecting a subject the latter is the behaviour worth having.
**F6 — tract cannot parse a dynamic-shape export.** It fails shape inference on the neck's `Concat`.
The graph therefore ships with its input fixed at 640×640 square, and every image is letterboxed into
it. This is the constraint behind the tiling option in `semantic.rs`: with a fixed window, tiling is
the *only* route to more semantic resolution, and it costs one inference per tile (≈2.8 s for a 3×2
grid over a 1600 px proxy against 470 ms whole-frame). Defaulted off — a photographic subject is
usually large in frame, which is the case whole-frame inference handles best — and left implemented
so §9's corpus can settle it rather than an argument.
**Arm C ships, and the §8 criteria were not what decided it.** §8 asked for a two-interaction margin
over arm A on the scene subset. That comparison was never run, because F4 and F5 changed what the
arms *are*: with a model that recognises subjects and has no word for sky, arm B alone cannot be a
selection tool at all (§8's "arm B never ships alone" holds, for a stronger reason than expected),
and arm A alone cannot tell a person from the wall behind them. They are complements rather than
candidates. Arm C's implementation is `prior.rs`: instance membership re-weights the merge saddles,
so region pairs the model believes share an object merge early and pairs straddling its edge merge
late. **No boundary moves** — only the order in which boundaries dissolve — which is how the result
stays pixel-accurate at every level while its coarse levels become named things.
**M1–M4 remain unmeasured.** The 24-image corpus of §9 has not been traced, so there are no
interaction counts and no boundary-accuracy numbers. What exists is a working feature and the
evidence that each piece does what it claims in isolation. The corpus is still the thing that would
turn "this feels right" into a number, and it is the largest piece of §11 left undone.
**M5 is unchanged and still the risk it was.** Run-to-run on one device is identical, asserted by a
test. Cross-vendor is untouched. Region ids now reach the sidecar, so if the label field proves
non-deterministic across vendors a mask written on the desktop will not mean the same thing on
Android — see `MaskSource::Regions::signature`, which detects a *retuned* segmentation but not a
differently-rounded one.
**F3 still stands.** `Segmentation::read_field` still copies the label and gradient buffers to the
CPU to build the region graph. It is now behind its own `segment-readback` feature rather than
sharing `readback` — this transfer is once per image on a worker, where the one AC-8 forbids is per
frame in the render loop — but the accumulation still belongs GPU-side with atomics.
---
## 14. Register entries
**S15** — *Region segmentation for local masking* · **CLOSED 2026-08-21**. Arms A, B and C built; the
shared consumers built; the licence question resolved (§7, D14). The 24-image corpus was not traced,
so M1–M4 are unmeasured and M5 is answered only on one device. Answers: local masking snaps to arm C,
and a mask is stored as region ids. Relates to: D13, D14, FR-DEV-3, ARCH §5.4, §6.13.
**D14** — *Segmentation source for local masking* · **DECIDED 2026-08-21**: **arm C**, a watershed
hierarchy re-weighted by YOLO26n-seg instance membership, with the model optional and the watershed
sufficient without it. Weights ship in-tree under AGPL-3.0, which GPLv3 §13 permits and which makes
the combined work effectively AGPL — a deliberate change to DarkRoom's licensing posture, not a
dependency detail (`core/dr-segment/models/LICENCE.md`).
**D13** — *inference runtime* · the dependency half is **answered** for segmentation and the answer
generalises: `ort` + `ort-tract` gives ONNX inference in pure Rust, so the face pipeline of §3.9.1
needs no C dependency either. The *model licensing* half of D13 is untouched — the InsightFace
weights are still non-commercial and still unusable here.