Record what the spec got wrong about the model that exists
docs/segmentation.md §4 priced arm B as costing a C dependency under the NDK and treated that as most of the difference between the arms. It is not a cost that has to be paid: `ort`'s `alternative-backend` disables its linking entirely and `ort-tract` supplies the API from tract, which is pure Rust. D13's "largest exception the policy would tolerate" turns out not to be needed, and the answer generalises to the face pipeline — so D13's runtime half is now answered and only its licensing half is open. Three findings contradict §4 outright and are recorded as F4-F6 rather than quietly designed around. There is no ADE20K-trained YOLO, so the shipped vocabulary selects subjects and not stuff — "select the sky" comes from the watershed or from nowhere. It is instance segmentation, so it partitions nothing and two people come back as two instances. And tract cannot parse a dynamic-shape export, which fixes the input at 640 square and makes tiling the only route to more semantic resolution. Arm C ships, but §8's criteria are not what decided it, and saying so matters more than claiming the process worked. §8 asked for a two- interaction margin over arm A on a traced corpus. That comparison was never run: F4 and F5 changed what the arms are, and a model that recognises subjects but has no word for sky cannot be a selection tool alone, while a watershed cannot tell a person from the wall behind them. They stopped being candidates and became complements. What is *not* done is written down as plainly: the 24-image corpus is untraced, so M1-M4 have no numbers and "this feels right" has not become one. M5 is answered on one device only, and region ids now reach the sidecar — so a cross-vendor divergence would mean a mask written on the desktop meaning something else on Android. F3 stands.
This commit is contained in:
+87
-8
@@ -345,14 +345,93 @@ measurement that can invalidate the region-id representation.
|
||||
|
||||
---
|
||||
|
||||
## 13. Register entries
|
||||
## 13. Arms B and C, and what §4 got wrong
|
||||
|
||||
To be added when this is accepted:
|
||||
Built 2026-08-21 on branch `local-adjustments`. `core/dr-segment`, `core/dr-gpu/src/mask.rs`,
|
||||
`core/dr-pipeline/src/mask.rs`, and the develop panel.
|
||||
|
||||
**S15** — *Region segmentation for local masking:* implement arm A and the shared consumers, trace a
|
||||
24-image corpus, measure M1–M7 across arms. **Resolve the licence question before writing arm B.**
|
||||
Answers: which segmentation source local masking snaps to, and whether a mask can be stored as region
|
||||
ids at all. Relates to: D13, D14, FR-DEV-3, ARCH §5.4, §6.13.
|
||||
**The dependency question dissolved rather than being decided.** §4 and D13 both priced arm B as
|
||||
costing a C dependency under the Android NDK, and treated that as most of the difference between the
|
||||
arms. It is not a cost that has to be paid: `ort` 2.0's `alternative-backend` feature disables the
|
||||
linking entirely and lets another engine supply the `OrtApi`, and `ort-tract` — same authors,
|
||||
MIT/Apache — supplies it from `tract`, which is pure Rust. So arm B runs through `ort`'s API with no
|
||||
C anywhere, and D13's "largest exception the policy would tolerate" turns out not to be needed.
|
||||
|
||||
**D14** — *Segmentation source for local masking* · **OPEN**. Decided by S15 against the criteria in
|
||||
§8. Reopens D13's dependency-policy question if arm C wins.
|
||||
Measured before committing to it, because tract's operator coverage is the thing that could have
|
||||
sunk it: **yolo26n-seg loads with zero unsupported operators** and runs 640×640 in ~470 ms on the
|
||||
reference desktop's CPU. Correct masks on the standard `bus.jpg` — one bus and three people, outlines
|
||||
following the subjects.
|
||||
|
||||
Three findings that contradict §4 directly, and all three change the design rather than the schedule.
|
||||
|
||||
**F4 — there is no ADE20K-trained YOLO.** §4's whole argument for arm B was ADE20K's 150 classes and
|
||||
their *stuff* categories: "'that patch of sky' is a class here." Checked 2026-08-21: Ultralytics ships
|
||||
YOLO26-seg trained on **COCO**, whose 80 classes are all *things*, and the one HuggingFace repository
|
||||
claiming a YOLO/ADE20K combination is empty. ADE20K models exist as SegFormer/OneFormer/MaskFormer
|
||||
transformers, not as YOLO.
|
||||
|
||||
So the shipped vocabulary selects **subjects**, not **stuff**. "Select the person" works; "select the
|
||||
sky" does not come from the model at all and must come from the watershed. That is a narrower arm B
|
||||
than §4 assumed, and it *raises* the importance of arm C rather than lowering it — the model can no
|
||||
longer be the whole answer for anything.
|
||||
|
||||
**F5 — it is instance segmentation, not semantic segmentation.** §4 assumed a flat partition with
|
||||
class ids that would "populate `RegionField` directly". YOLO-seg does not partition the image; it
|
||||
finds objects, and most pixels in a landscape belong to no instance. Two consequences, one bad and
|
||||
one better than expected: nothing populates a `RegionField` on its own, and two people come back as
|
||||
*two* instances where a semantic model would have returned one "person" area covering both. For
|
||||
selecting a subject the latter is the behaviour worth having.
|
||||
|
||||
**F6 — tract cannot parse a dynamic-shape export.** It fails shape inference on the neck's `Concat`.
|
||||
The graph therefore ships with its input fixed at 640×640 square, and every image is letterboxed into
|
||||
it. This is the constraint behind the tiling option in `semantic.rs`: with a fixed window, tiling is
|
||||
the *only* route to more semantic resolution, and it costs one inference per tile (≈2.8 s for a 3×2
|
||||
grid over a 1600 px proxy against 470 ms whole-frame). Defaulted off — a photographic subject is
|
||||
usually large in frame, which is the case whole-frame inference handles best — and left implemented
|
||||
so §9's corpus can settle it rather than an argument.
|
||||
|
||||
**Arm C ships, and the §8 criteria were not what decided it.** §8 asked for a two-interaction margin
|
||||
over arm A on the scene subset. That comparison was never run, because F4 and F5 changed what the
|
||||
arms *are*: with a model that recognises subjects and has no word for sky, arm B alone cannot be a
|
||||
selection tool at all (§8's "arm B never ships alone" holds, for a stronger reason than expected),
|
||||
and arm A alone cannot tell a person from the wall behind them. They are complements rather than
|
||||
candidates. Arm C's implementation is `prior.rs`: instance membership re-weights the merge saddles,
|
||||
so region pairs the model believes share an object merge early and pairs straddling its edge merge
|
||||
late. **No boundary moves** — only the order in which boundaries dissolve — which is how the result
|
||||
stays pixel-accurate at every level while its coarse levels become named things.
|
||||
|
||||
**M1–M4 remain unmeasured.** The 24-image corpus of §9 has not been traced, so there are no
|
||||
interaction counts and no boundary-accuracy numbers. What exists is a working feature and the
|
||||
evidence that each piece does what it claims in isolation. The corpus is still the thing that would
|
||||
turn "this feels right" into a number, and it is the largest piece of §11 left undone.
|
||||
|
||||
**M5 is unchanged and still the risk it was.** Run-to-run on one device is identical, asserted by a
|
||||
test. Cross-vendor is untouched. Region ids now reach the sidecar, so if the label field proves
|
||||
non-deterministic across vendors a mask written on the desktop will not mean the same thing on
|
||||
Android — see `MaskSource::Regions::signature`, which detects a *retuned* segmentation but not a
|
||||
differently-rounded one.
|
||||
|
||||
**F3 still stands.** `Segmentation::read_field` still copies the label and gradient buffers to the
|
||||
CPU to build the region graph. It is now behind its own `segment-readback` feature rather than
|
||||
sharing `readback` — this transfer is once per image on a worker, where the one AC-8 forbids is per
|
||||
frame in the render loop — but the accumulation still belongs GPU-side with atomics.
|
||||
|
||||
---
|
||||
|
||||
## 14. Register entries
|
||||
|
||||
**S15** — *Region segmentation for local masking* · **CLOSED 2026-08-21**. Arms A, B and C built; the
|
||||
shared consumers built; the licence question resolved (§7, D14). The 24-image corpus was not traced,
|
||||
so M1–M4 are unmeasured and M5 is answered only on one device. Answers: local masking snaps to arm C,
|
||||
and a mask is stored as region ids. Relates to: D13, D14, FR-DEV-3, ARCH §5.4, §6.13.
|
||||
|
||||
**D14** — *Segmentation source for local masking* · **DECIDED 2026-08-21**: **arm C**, a watershed
|
||||
hierarchy re-weighted by YOLO26n-seg instance membership, with the model optional and the watershed
|
||||
sufficient without it. Weights ship in-tree under AGPL-3.0, which GPLv3 §13 permits and which makes
|
||||
the combined work effectively AGPL — a deliberate change to DarkRoom's licensing posture, not a
|
||||
dependency detail (`core/dr-segment/models/LICENCE.md`).
|
||||
|
||||
**D13** — *inference runtime* · the dependency half is **answered** for segmentation and the answer
|
||||
generalises: `ort` + `ort-tract` gives ONNX inference in pure Rust, so the face pipeline of §3.9.1
|
||||
needs no C dependency either. The *model licensing* half of D13 is untouched — the InsightFace
|
||||
weights are still non-commercial and still unusable here.
|
||||
|
||||
Reference in New Issue
Block a user