Merge the scene model: a second, ADE20K-trained graph for per-category grades

Five commits. `models/` becomes one tree at the repository root so the
weight the application carries is a single `du -sh`; the export script
stops building its multi-gigabyte venv in RAM; `yolo26s-sem-ade20k`
joins the instance model rather than replacing it; the decoder turns its
logits into a partition of unity over eight photographic categories; and
four packaging routes put the file somewhere each platform can find it.

The instance model stays exactly where it was. A semantic model merges
every pixel of a class into one region, so it cannot separate two people,
and separating two people is what clicking a subject needs. The scene tab
grades whole categories and does not care. docs/segmentation.md §16.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

# Conflicts:
#	docs/traceability.md
This commit is contained in:
2026-08-30 10:50:49 +02:00
24 changed files with 1387 additions and 166 deletions
+66
View File
@@ -435,3 +435,69 @@ dependency detail (`core/dr-segment/models/LICENCE.md`).
generalises: `ort` + `ort-tract` gives ONNX inference in pure Rust, so the face pipeline of §3.9.1
needs no C dependency either. The *model licensing* half of D13 is untouched — the InsightFace
weights are still non-commercial and still unusable here.
---
## 16. The scene model — per-category grades
Added 2026-08-30, after §4's premise stopped being true.
### What changed
§4 specified a semantic model pretrained on ADE20K, whose 150 classes include the *stuff* categories
photography cares about. §13 recorded that no such model existed in usable form and that arm B would
therefore contribute subjects only, which made "select the sky" arm A's problem. Re-checked
2026-08-30: **Ultralytics now ships a `semantic` task with ADE20K checkpoints**
(`docs.ultralytics.com/tasks/semantic`). `yolo26s-sem-ade20k` is in `models/scene/`.
### It is an addition, not a correction to arm B
The instance model stays exactly where it was, and the reason is the one §13 already gave and was
right about: a semantic model merges every pixel of a class into one region, so it cannot separate
two people, and separating two people is what clicking a subject requires. Swapping arm B for this
would regress the primary interaction to fix a secondary one.
So the two divide by *what the user is doing*, not by which is better:
| | `models/segment/` (COCO instances) | `models/scene/` (ADE20K semantics) |
|---|---|---|
| Question | which pixels are *that* dog | how much of this pixel is sky |
| Granularity | per instance | per category, whole frame |
| Drives | local adjustments, subject selection | the scene tab's per-category sliders |
| Vocabulary | 80 things | 150 classes, stuff included |
### The export is truncated, and both reasons matter
Ultralytics ends the graph with `Resize → ArgMax → Cast`, returning a `[1, 640, 640]` u8 label map.
`tools/export-seg-model.sh` cuts that tail and ships the classifier's `[1, 150, 80, 80]` f32 logits.
**Cost.** The `Resize` materialises 150 × 640 × 640 × f32 — 246 MB — and the `ArgMax` then reduces
across the channel axis, striding 409,600 elements per comparison. Measured under load it was
roughly four fifths of total runtime, spent on work the application discards.
**Softness, which is the more important one.** `ArgMax` destroys the per-class scores, and the whole
design of the scene tab rests on keeping them. Softmax over the 150 channels, summed within each
category, produces per-category weights that sum to one at every pixel — a partition of unity.
Feathering that cannot double-grade a boundary. Feathering *hard labels* outward from two adjacent
categories paints both grades into the overlap, and every horizon in the frame acquires a seam.
### The resolution is 80×80, and no setting changes that
The discarded upsample was never information. `Scene` keeps the native grid and resamples on demand,
so the coarseness is visible in the type rather than hidden. A graduated grade over sky or water is
untroubled by it; a rooftop against sky at 100% zoom will show it. This is the constraint most likely
to decide whether the tab feels good, and it is not addressable by choosing a larger checkpoint —
`yolo26n-sem` and `yolo26s-sem` have the same output grid.
### Licence
Unchanged. Same AGPL-3.0 grant as the instance model, same GPLv3 §13 permission, same consequence
already accepted in D14 — so this needed no new licence decision, which is most of why it was cheap.
See `models/LICENCE.md`.
### Measurement
Timings taken while this was chosen came off a laptop compiling other things and are upper bounds
only. `cargo run -p dr-segment --example scene --release --features embedded-scene-model` reports a
median over N runs with the first excluded; a number worth quoting should come from that, on an idle
machine.
+42 -42
View File
File diff suppressed because one or more lines are too long