7ae1e2781071a2a771b75c7cc7a07ae5f16f93b2
2
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
8df6000e4b |
Decode the scene model into per-category weights
The weights landed last commit with nothing to read them. This is the decoder, and the shape of it follows from one property worth stating before the code: the categories must partition the image. ## Why a partition, and not a mask per category The scene tab applies one grade to every pixel of a category — lift the sky, desaturate foliage — and both grades meet at the horizon. If each category carried an independent mask, feathering them outward would make the boundary band belong to both, so both grades would land there and every horizon would acquire a visible seam. Feathering has to *blend* there, not accumulate. So `marginalise` takes one softmax over all 150 channels and sums within each category. Grouping cannot change a total of one, so the listed categories plus the unlisted remainder sum to one at every pixel, by construction rather than by normalising afterwards. `parse_categories` refuses a descriptor that claims a class twice, because that is the one input that would quietly make the property untrue. ## The descriptor is data, and hand-written `models/scene/categories.txt` groups ADE20K's 150 classes into the eight a photographer would recognise. It is a file rather than a table in Rust for the reason `models/LICENCE.md` predicted — a vocabulary is model metadata — and it is line-oriented with comments rather than JSON like the `.classes.json` beside it, because that file is generated and this one is argued. Why `swimming pool` is water and not architecture belongs next to the line that says so. Classes are named, not indexed. An index is silently wrong after a re-export; a name is loudly wrong, and the loader refuses one the model does not have. ## Resolution, kept visible `Scene` holds the native 80×80 logit grid and resamples on demand rather than upsampling once at load. The coarseness is real — it is what the graph produces — and a type that hides it behind an early resize invites callers to expect detail that was never there. `rasterise` is where the letterbox inverse lives, once. `Letterbox` and `Window` become `pub(crate)` and `to_proto` generalises to `to_grid`, because both dense outputs this crate reads are an even fraction of the same letterboxed square and differ only in the divisor. ## Verified by looking, which is the only way this gets verified `examples/scene.rs` writes the photograph dimmed outside each category. A transposed axis or an off-by-one in the inverse produces perfectly plausible weights over slightly the wrong pixels, and no unit test catches that. On an indoor frame the person mask lands on the person, including the outstretched arm, and sky reads ~5% against a bright ceiling. It doubles as the benchmark, because every timing quoted while this model was chosen came off a laptop compiling other things and none of them belong in a document. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9a2b39b8e5 |
Add the ADE20K scene model beside the instance one
`models/LICENCE.md` recorded, on 2026-08-21, that no YOLO model trained on ADE20K existed in usable form — the stuff classes photography cares about, sky and vegetation and water, had no model to come from. Re-checked 2026-08-30: Ultralytics now ships a `semantic` task with ADE20K checkpoints, so `models/scene/` holds `yolo26s-sem-ade20k`. This is an addition, not a replacement. A semantic model labels every pixel but merges same-class pixels into one region, so it cannot tell three people apart — which is exactly what clicking a subject needs, and exactly what `segment/`'s COCO instance model already does. The scene tab grades per category and does not care that instances are merged. Keeping both is the point. ## The export is truncated, deliberately Ultralytics ends the graph with `Resize -> ArgMax -> Cast` and hands back a `[1, 640, 640]` u8 label map. The script cuts that tail and exposes the classifier's `[1, 150, 80, 80]` f32 logits instead, for two reasons. Cost: the Resize materialises 150 x 640 x 640 x f32, 246 MB, and ArgMax then reduces across the channel axis, striding 409,600 elements per comparison. On one loaded machine the full graph ran ~1160 ms against ~500 ms truncated — roughly four fifths of the time spent on work the application discards. Those numbers were measured under contention and are upper bounds, but the ratio is structural. Softness: ArgMax destroys the per-class scores, and the scene tab needs them. Softmax over the 150 channels, summed within each photographic category, yields per-category weights summing to 1 at every pixel. Feathering a partition of unity cannot double-grade a boundary, whereas feathering hard labels outward from two adjacent categories paints both grades into the overlap and haloes every horizon. The discarded upsample was never information: the graph's true spatial resolution is the 80x80 logit grid, and the application can resample from that itself. The tail is matched by op type and asserted before cutting, so an upstream graph change fails loudly in the exporter rather than quietly shipping a differently-shaped model. Nothing reads these weights yet — the decode path, the category descriptor grouping 150 classes into ~8 photographic ones, and the scene tab are still to come. At 24 MB this model also wants the runtime-asset treatment `models/face/` already gets on Android rather than `include_bytes!`; embedding it would put ~35 MB of weights in the binary. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |