`models/LICENCE.md` recorded, on 2026-08-21, that no YOLO model trained on ADE20K existed in usable form — the stuff classes photography cares about, sky and vegetation and water, had no model to come from. Re-checked 2026-08-30: Ultralytics now ships a `semantic` task with ADE20K checkpoints, so `models/scene/` holds `yolo26s-sem-ade20k`. This is an addition, not a replacement. A semantic model labels every pixel but merges same-class pixels into one region, so it cannot tell three people apart — which is exactly what clicking a subject needs, and exactly what `segment/`'s COCO instance model already does. The scene tab grades per category and does not care that instances are merged. Keeping both is the point. ## The export is truncated, deliberately Ultralytics ends the graph with `Resize -> ArgMax -> Cast` and hands back a `[1, 640, 640]` u8 label map. The script cuts that tail and exposes the classifier's `[1, 150, 80, 80]` f32 logits instead, for two reasons. Cost: the Resize materialises 150 x 640 x 640 x f32, 246 MB, and ArgMax then reduces across the channel axis, striding 409,600 elements per comparison. On one loaded machine the full graph ran ~1160 ms against ~500 ms truncated — roughly four fifths of the time spent on work the application discards. Those numbers were measured under contention and are upper bounds, but the ratio is structural. Softness: ArgMax destroys the per-class scores, and the scene tab needs them. Softmax over the 150 channels, summed within each photographic category, yields per-category weights summing to 1 at every pixel. Feathering a partition of unity cannot double-grade a boundary, whereas feathering hard labels outward from two adjacent categories paints both grades into the overlap and haloes every horizon. The discarded upsample was never information: the graph's true spatial resolution is the 80x80 logit grid, and the application can resample from that itself. The tail is matched by op type and asserted before cutting, so an upstream graph change fails loudly in the exporter rather than quietly shipping a differently-shaped model. Nothing reads these weights yet — the decode path, the category descriptor grouping 150 classes into ~8 photographic ones, and the scene tab are still to come. At 24 MB this model also wants the runtime-asset treatment `models/face/` already gets on Android rather than `include_bytes!`; embedding it would put ~35 MB of weights in the binary. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
24 MiBLFS
24 MiBLFS
The file is too large to be shown.
View Raw