Decode the scene model into per-category weights
The weights landed last commit with nothing to read them. This is the decoder, and the shape of it follows from one property worth stating before the code: the categories must partition the image. ## Why a partition, and not a mask per category The scene tab applies one grade to every pixel of a category — lift the sky, desaturate foliage — and both grades meet at the horizon. If each category carried an independent mask, feathering them outward would make the boundary band belong to both, so both grades would land there and every horizon would acquire a visible seam. Feathering has to *blend* there, not accumulate. So `marginalise` takes one softmax over all 150 channels and sums within each category. Grouping cannot change a total of one, so the listed categories plus the unlisted remainder sum to one at every pixel, by construction rather than by normalising afterwards. `parse_categories` refuses a descriptor that claims a class twice, because that is the one input that would quietly make the property untrue. ## The descriptor is data, and hand-written `models/scene/categories.txt` groups ADE20K's 150 classes into the eight a photographer would recognise. It is a file rather than a table in Rust for the reason `models/LICENCE.md` predicted — a vocabulary is model metadata — and it is line-oriented with comments rather than JSON like the `.classes.json` beside it, because that file is generated and this one is argued. Why `swimming pool` is water and not architecture belongs next to the line that says so. Classes are named, not indexed. An index is silently wrong after a re-export; a name is loudly wrong, and the loader refuses one the model does not have. ## Resolution, kept visible `Scene` holds the native 80×80 logit grid and resamples on demand rather than upsampling once at load. The coarseness is real — it is what the graph produces — and a type that hides it behind an early resize invites callers to expect detail that was never there. `rasterise` is where the letterbox inverse lives, once. `Letterbox` and `Window` become `pub(crate)` and `to_proto` generalises to `to_grid`, because both dense outputs this crate reads are an even fraction of the same letterboxed square and differ only in the divisor. ## Verified by looking, which is the only way this gets verified `examples/scene.rs` writes the photograph dimmed outside each category. A transposed axis or an off-by-one in the inverse produces perfectly plausible weights over slightly the wrong pixels, and no unit test catches that. On an indoor frame the person mask lands on the person, including the outstretched arm, and sky reads ~5% against a bright ceiling. It doubles as the benchmark, because every timing quoted while this model was chosen came off a laptop compiling other things and none of them belong in a document. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -435,3 +435,69 @@ dependency detail (`core/dr-segment/models/LICENCE.md`).
|
||||
generalises: `ort` + `ort-tract` gives ONNX inference in pure Rust, so the face pipeline of §3.9.1
|
||||
needs no C dependency either. The *model licensing* half of D13 is untouched — the InsightFace
|
||||
weights are still non-commercial and still unusable here.
|
||||
|
||||
---
|
||||
|
||||
## 16. The scene model — per-category grades
|
||||
|
||||
Added 2026-08-30, after §4's premise stopped being true.
|
||||
|
||||
### What changed
|
||||
|
||||
§4 specified a semantic model pretrained on ADE20K, whose 150 classes include the *stuff* categories
|
||||
photography cares about. §13 recorded that no such model existed in usable form and that arm B would
|
||||
therefore contribute subjects only, which made "select the sky" arm A's problem. Re-checked
|
||||
2026-08-30: **Ultralytics now ships a `semantic` task with ADE20K checkpoints**
|
||||
(`docs.ultralytics.com/tasks/semantic`). `yolo26s-sem-ade20k` is in `models/scene/`.
|
||||
|
||||
### It is an addition, not a correction to arm B
|
||||
|
||||
The instance model stays exactly where it was, and the reason is the one §13 already gave and was
|
||||
right about: a semantic model merges every pixel of a class into one region, so it cannot separate
|
||||
two people, and separating two people is what clicking a subject requires. Swapping arm B for this
|
||||
would regress the primary interaction to fix a secondary one.
|
||||
|
||||
So the two divide by *what the user is doing*, not by which is better:
|
||||
|
||||
| | `models/segment/` (COCO instances) | `models/scene/` (ADE20K semantics) |
|
||||
|---|---|---|
|
||||
| Question | which pixels are *that* dog | how much of this pixel is sky |
|
||||
| Granularity | per instance | per category, whole frame |
|
||||
| Drives | local adjustments, subject selection | the scene tab's per-category sliders |
|
||||
| Vocabulary | 80 things | 150 classes, stuff included |
|
||||
|
||||
### The export is truncated, and both reasons matter
|
||||
|
||||
Ultralytics ends the graph with `Resize → ArgMax → Cast`, returning a `[1, 640, 640]` u8 label map.
|
||||
`tools/export-seg-model.sh` cuts that tail and ships the classifier's `[1, 150, 80, 80]` f32 logits.
|
||||
|
||||
**Cost.** The `Resize` materialises 150 × 640 × 640 × f32 — 246 MB — and the `ArgMax` then reduces
|
||||
across the channel axis, striding 409,600 elements per comparison. Measured under load it was
|
||||
roughly four fifths of total runtime, spent on work the application discards.
|
||||
|
||||
**Softness, which is the more important one.** `ArgMax` destroys the per-class scores, and the whole
|
||||
design of the scene tab rests on keeping them. Softmax over the 150 channels, summed within each
|
||||
category, produces per-category weights that sum to one at every pixel — a partition of unity.
|
||||
Feathering that cannot double-grade a boundary. Feathering *hard labels* outward from two adjacent
|
||||
categories paints both grades into the overlap, and every horizon in the frame acquires a seam.
|
||||
|
||||
### The resolution is 80×80, and no setting changes that
|
||||
|
||||
The discarded upsample was never information. `Scene` keeps the native grid and resamples on demand,
|
||||
so the coarseness is visible in the type rather than hidden. A graduated grade over sky or water is
|
||||
untroubled by it; a rooftop against sky at 100% zoom will show it. This is the constraint most likely
|
||||
to decide whether the tab feels good, and it is not addressable by choosing a larger checkpoint —
|
||||
`yolo26n-sem` and `yolo26s-sem` have the same output grid.
|
||||
|
||||
### Licence
|
||||
|
||||
Unchanged. Same AGPL-3.0 grant as the instance model, same GPLv3 §13 permission, same consequence
|
||||
already accepted in D14 — so this needed no new licence decision, which is most of why it was cheap.
|
||||
See `models/LICENCE.md`.
|
||||
|
||||
### Measurement
|
||||
|
||||
Timings taken while this was chosen came off a laptop compiling other things and are upper bounds
|
||||
only. `cargo run -p dr-segment --example scene --release --features embedded-scene-model` reports a
|
||||
median over N runs with the first excluded; a number worth quoting should come from that, on an idle
|
||||
machine.
|
||||
|
||||
Reference in New Issue
Block a user