Let the model say what a thing is and the watershed say where it ends

Local masking needs to know where an image's regions are. The watershed
spike (S15 arm A) found the boundaries but had no idea what any of them
enclosed; its coarse levels were geometric accidents. This adds the other
half and the thing that joins them.

`core/dr-segment` is where region reasoning now lives — the hierarchy moves
out of `dr-gpu`, which keeps only the pixel passes that are genuinely
shaders. The new crate is device-free and, without its default features,
model-free too: 20 of its tests need neither an adapter nor 11 MB of
weights.

Arm B runs YOLO26n-seg through `ort`. D13 framed inference as a choice
between `ort`'s C++ runtime and the pure-Rust dependency policy; that was a
false choice. `ort`'s `alternative-backend` feature unlinks the C entirely
and `ort-tract` supplies the API from tract, which is pure Rust. Measured
before committing to it: zero unsupported operators, 420 ms for 640x640,
and correct masks on bus.jpg. No NDK problem to solve, so D13's largest
tolerated exception is not needed.

Arm C is `prior.rs`, and it ships because the two arms fail in opposite
directions. Instance membership re-weights the merge saddles, so region
pairs the model believes share an object merge early and pairs straddling
its edge merge late. No boundary moves — only the order in which they
dissolve — which is how the result stays pixel-accurate at every level
while its coarse levels become named things.

Two things the spec assumed that turned out to be false, both recorded in
models/LICENCE.md: there is no usable ADE20K-trained YOLO, so the shipped
vocabulary is COCO's 80 subjects and *stuff* like sky and foliage must come
from arm A; and tract cannot parse a dynamic-shape export, so the graph's
input is fixed and tiling is the only route to more semantic resolution.

Weights are AGPL-3.0, which GPLv3 §13 permits and which makes the combined
work effectively AGPL. Deliberate, not accidental. They live in Git LFS,
and a build script fails with an instruction rather than embedding a
pointer file when the clone lacks them.
This commit is contained in:
2026-08-22 08:39:16 +02:00
parent ecd6df686c
commit 0da8271836
18 changed files with 2287 additions and 19 deletions
+54
View File
@@ -0,0 +1,54 @@
# Model weights — licensing
`yolo26n-seg.onnx` is exported from Ultralytics YOLO26n-seg
(`https://huggingface.co/Ultralytics/YOLO26`, `yolo26n-seg.pt`) by
`tools/export-seg-model.sh`. `yolo26n-seg.classes.json` is that checkpoint's
class vocabulary, written out by the same script.
## The grant
**Ultralytics releases YOLO under AGPL-3.0**, and the weights carry the same
grant as the framework — the HuggingFace repository declares `agpl-3.0` for the
checkpoints themselves, not merely for the training code. A commercial licence
is offered separately; DarkRoom does not use it and does not need it.
## What that means for DarkRoom
DarkRoom is GPL-3.0-or-later. **GPLv3 §13 explicitly permits combination with
AGPL-3.0 code**, so redistributing these weights inside this repository is
allowed — this is *not* the situation the InsightFace "buffalo" weights would
have created, where a non-commercial research grant is simply incompatible with
the project's licence and with F-Droid, Flatpak and Play distribution
(NFR-COMPAT-2, D13).
The consequence, and it is a real one: **the combined work is effectively
AGPL-3.0.** §13's permission runs one way — the AGPL's §13 network-use condition
attaches to the portion under that licence. For a local-first desktop and
Android photo editor that condition has no practical bite, because there is no
network service offering the combined work to remote users. It would acquire
bite the moment any hosted or server-side rendering appeared, and that is the
thing to remember rather than rediscover.
This was decided deliberately (D14), not arrived at by accident, and
`docs/segmentation.md` §7 records the reasoning.
## Class vocabulary — a caveat worth reading
`docs/segmentation.md` §4 specified YOLO **pretrained on ADE20K**, whose 150
classes include the *stuff* categories that matter most in photography — sky,
vegetation, water, wall, mountain.
**No such model exists in usable form.** Checked 2026-08-21: Ultralytics ships
YOLO26-seg trained on **COCO**, whose 80 classes are all *things* — person,
dog, car, bird, potted plant — and the one HuggingFace repository claiming a
YOLO/ADE20K combination (`laxmacl/yolov8-ade20k`) is empty. ADE20K semantic
models do exist, but as SegFormer/OneFormer/MaskFormer transformers, not YOLO.
So the shipped vocabulary selects **subjects**, not **stuff**. "Select the
person" works; "select the sky" does not come from the model and must come from
the watershed hierarchy instead. That is a narrower arm B than §4 assumed, and
it raises rather than lowers the importance of arm C.
The loader treats the vocabulary as model metadata rather than compiled-in
knowledge, so adding a stuff-class model later is a file plus a descriptor, not
a code change.