Merge the scene model: a second, ADE20K-trained graph for per-category grades
Five commits. `models/` becomes one tree at the repository root so the weight the application carries is a single `du -sh`; the export script stops building its multi-gigabyte venv in RAM; `yolo26s-sem-ade20k` joins the instance model rather than replacing it; the decoder turns its logits into a partition of unity over eight photographic categories; and four packaging routes put the file somewhere each platform can find it. The instance model stays exactly where it was. A semantic model merges every pixel of a class into one region, so it cannot separate two people, and separating two people is what clicking a subject needs. The scene tab grades whole categories and does not care. docs/segmentation.md §16. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> # Conflicts: # docs/traceability.md
This commit is contained in:
@@ -435,3 +435,69 @@ dependency detail (`core/dr-segment/models/LICENCE.md`).
|
||||
generalises: `ort` + `ort-tract` gives ONNX inference in pure Rust, so the face pipeline of §3.9.1
|
||||
needs no C dependency either. The *model licensing* half of D13 is untouched — the InsightFace
|
||||
weights are still non-commercial and still unusable here.
|
||||
|
||||
---
|
||||
|
||||
## 16. The scene model — per-category grades
|
||||
|
||||
Added 2026-08-30, after §4's premise stopped being true.
|
||||
|
||||
### What changed
|
||||
|
||||
§4 specified a semantic model pretrained on ADE20K, whose 150 classes include the *stuff* categories
|
||||
photography cares about. §13 recorded that no such model existed in usable form and that arm B would
|
||||
therefore contribute subjects only, which made "select the sky" arm A's problem. Re-checked
|
||||
2026-08-30: **Ultralytics now ships a `semantic` task with ADE20K checkpoints**
|
||||
(`docs.ultralytics.com/tasks/semantic`). `yolo26s-sem-ade20k` is in `models/scene/`.
|
||||
|
||||
### It is an addition, not a correction to arm B
|
||||
|
||||
The instance model stays exactly where it was, and the reason is the one §13 already gave and was
|
||||
right about: a semantic model merges every pixel of a class into one region, so it cannot separate
|
||||
two people, and separating two people is what clicking a subject requires. Swapping arm B for this
|
||||
would regress the primary interaction to fix a secondary one.
|
||||
|
||||
So the two divide by *what the user is doing*, not by which is better:
|
||||
|
||||
| | `models/segment/` (COCO instances) | `models/scene/` (ADE20K semantics) |
|
||||
|---|---|---|
|
||||
| Question | which pixels are *that* dog | how much of this pixel is sky |
|
||||
| Granularity | per instance | per category, whole frame |
|
||||
| Drives | local adjustments, subject selection | the scene tab's per-category sliders |
|
||||
| Vocabulary | 80 things | 150 classes, stuff included |
|
||||
|
||||
### The export is truncated, and both reasons matter
|
||||
|
||||
Ultralytics ends the graph with `Resize → ArgMax → Cast`, returning a `[1, 640, 640]` u8 label map.
|
||||
`tools/export-seg-model.sh` cuts that tail and ships the classifier's `[1, 150, 80, 80]` f32 logits.
|
||||
|
||||
**Cost.** The `Resize` materialises 150 × 640 × 640 × f32 — 246 MB — and the `ArgMax` then reduces
|
||||
across the channel axis, striding 409,600 elements per comparison. Measured under load it was
|
||||
roughly four fifths of total runtime, spent on work the application discards.
|
||||
|
||||
**Softness, which is the more important one.** `ArgMax` destroys the per-class scores, and the whole
|
||||
design of the scene tab rests on keeping them. Softmax over the 150 channels, summed within each
|
||||
category, produces per-category weights that sum to one at every pixel — a partition of unity.
|
||||
Feathering that cannot double-grade a boundary. Feathering *hard labels* outward from two adjacent
|
||||
categories paints both grades into the overlap, and every horizon in the frame acquires a seam.
|
||||
|
||||
### The resolution is 80×80, and no setting changes that
|
||||
|
||||
The discarded upsample was never information. `Scene` keeps the native grid and resamples on demand,
|
||||
so the coarseness is visible in the type rather than hidden. A graduated grade over sky or water is
|
||||
untroubled by it; a rooftop against sky at 100% zoom will show it. This is the constraint most likely
|
||||
to decide whether the tab feels good, and it is not addressable by choosing a larger checkpoint —
|
||||
`yolo26n-sem` and `yolo26s-sem` have the same output grid.
|
||||
|
||||
### Licence
|
||||
|
||||
Unchanged. Same AGPL-3.0 grant as the instance model, same GPLv3 §13 permission, same consequence
|
||||
already accepted in D14 — so this needed no new licence decision, which is most of why it was cheap.
|
||||
See `models/LICENCE.md`.
|
||||
|
||||
### Measurement
|
||||
|
||||
Timings taken while this was chosen came off a laptop compiling other things and are upper bounds
|
||||
only. `cargo run -p dr-segment --example scene --release --features embedded-scene-model` reports a
|
||||
median over N runs with the first excluded; a number worth quoting should come from that, on an idle
|
||||
machine.
|
||||
|
||||
+42
-42
File diff suppressed because one or more lines are too long
Reference in New Issue
Block a user