Files
DarkRoom/models/LICENCE.md
T
dtourolleandClaude Opus 5 9a2b39b8e5 Add the ADE20K scene model beside the instance one
`models/LICENCE.md` recorded, on 2026-08-21, that no YOLO model trained on
ADE20K existed in usable form — the stuff classes photography cares about,
sky and vegetation and water, had no model to come from. Re-checked
2026-08-30: Ultralytics now ships a `semantic` task with ADE20K
checkpoints, so `models/scene/` holds `yolo26s-sem-ade20k`.

This is an addition, not a replacement. A semantic model labels every
pixel but merges same-class pixels into one region, so it cannot tell
three people apart — which is exactly what clicking a subject needs, and
exactly what `segment/`'s COCO instance model already does. The scene tab
grades per category and does not care that instances are merged. Keeping
both is the point.

## The export is truncated, deliberately

Ultralytics ends the graph with `Resize -> ArgMax -> Cast` and hands back
a `[1, 640, 640]` u8 label map. The script cuts that tail and exposes the
classifier's `[1, 150, 80, 80]` f32 logits instead, for two reasons.

Cost: the Resize materialises 150 x 640 x 640 x f32, 246 MB, and ArgMax
then reduces across the channel axis, striding 409,600 elements per
comparison. On one loaded machine the full graph ran ~1160 ms against
~500 ms truncated — roughly four fifths of the time spent on work the
application discards. Those numbers were measured under contention and
are upper bounds, but the ratio is structural.

Softness: ArgMax destroys the per-class scores, and the scene tab needs
them. Softmax over the 150 channels, summed within each photographic
category, yields per-category weights summing to 1 at every pixel.
Feathering a partition of unity cannot double-grade a boundary, whereas
feathering hard labels outward from two adjacent categories paints both
grades into the overlap and haloes every horizon.

The discarded upsample was never information: the graph's true spatial
resolution is the 80x80 logit grid, and the application can resample from
that itself.

The tail is matched by op type and asserted before cutting, so an
upstream graph change fails loudly in the exporter rather than quietly
shipping a differently-shaped model.

Nothing reads these weights yet — the decode path, the category
descriptor grouping 150 classes into ~8 photographic ones, and the scene
tab are still to come. At 24 MB this model also wants the runtime-asset
treatment `models/face/` already gets on Android rather than
`include_bytes!`; embedding it would put ~35 MB of weights in the binary.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 10:05:44 +02:00

72 lines
3.5 KiB
Markdown

# Model weights — licensing
Two Ultralytics checkpoints ship here, both exported by
`tools/export-seg-model.sh`, each with its class vocabulary written out by the
same script:
| File | Checkpoint | Trained on | Used by |
|---|---|---|---|
| `segment/yolo26n-seg.onnx` | `yolo26n-seg.pt` | COCO, 80 *thing* classes | local adjustments, subject selection |
| `scene/yolo26s-sem-ade20k.onnx` | `yolo26s-sem-ade20k.pt` | ADE20K, 150 classes | the scene tab's per-category grades |
Both come from `https://huggingface.co/Ultralytics/YOLO26`. The face weights in
`face/` are a separate matter with a separate grant — see `face/README.md`.
## The grant
**Ultralytics releases YOLO under AGPL-3.0**, and the weights carry the same
grant as the framework — the HuggingFace repository declares `agpl-3.0` for the
checkpoints themselves, not merely for the training code. A commercial licence
is offered separately; DarkRoom does not use it and does not need it.
## What that means for DarkRoom
DarkRoom is GPL-3.0-or-later. **GPLv3 §13 explicitly permits combination with
AGPL-3.0 code**, so redistributing these weights inside this repository is
allowed — this is *not* the situation the InsightFace "buffalo" weights would
have created, where a non-commercial research grant is simply incompatible with
the project's licence and with F-Droid, Flatpak and Play distribution
(NFR-COMPAT-2, D13).
The consequence, and it is a real one: **the combined work is effectively
AGPL-3.0.** §13's permission runs one way — the AGPL's §13 network-use condition
attaches to the portion under that licence. For a local-first desktop and
Android photo editor that condition has no practical bite, because there is no
network service offering the combined work to remote users. It would acquire
bite the moment any hosted or server-side rendering appeared, and that is the
thing to remember rather than rediscover.
This was decided deliberately (D14), not arrived at by accident, and
`docs/segmentation.md` §7 records the reasoning.
## Class vocabulary — a caveat worth reading
`docs/segmentation.md` §4 specified YOLO **pretrained on ADE20K**, whose 150
classes include the *stuff* categories that matter most in photography — sky,
vegetation, water, wall, mountain.
**This was true when written and is not any more.** Checked 2026-08-21, no
YOLO/ADE20K combination existed: Ultralytics shipped YOLO26-seg on **COCO**
only, and the one HuggingFace repository claiming otherwise
(`laxmacl/yolov8-ade20k`) was empty. Re-checked 2026-08-30: Ultralytics now
ships a `semantic` task with ADE20K checkpoints
(`https://docs.ultralytics.com/tasks/semantic`), and `yolo26s-sem-ade20k` is
what `scene/` holds.
So the two vocabularies divide the work rather than compete:
- **`segment/`, COCO, 80 things.** Separates *instances* — clicking one of
three people selects that person. This is what local adjustments need, and a
semantic model cannot do it: it would return one "person" region covering all
three.
- **`scene/`, ADE20K, 150 classes.** Labels every pixel, including the *stuff*
COCO has no word for — sky, vegetation, water, mountain, wall. This is what
the scene tab's per-category grades need, and it does not care that instances
are merged, because a per-category grade applies to the whole category.
Neither replaces the other. Keeping both is the deliberate choice.
The loader treats each vocabulary as model metadata rather than compiled-in
knowledge, which is what made adding the second model a file plus a descriptor
rather than a code change — as this document predicted it would be.