The weights landed last commit with nothing to read them. This is the
decoder, and the shape of it follows from one property worth stating
before the code: the categories must partition the image.
## Why a partition, and not a mask per category
The scene tab applies one grade to every pixel of a category — lift the
sky, desaturate foliage — and both grades meet at the horizon. If each
category carried an independent mask, feathering them outward would make
the boundary band belong to both, so both grades would land there and
every horizon would acquire a visible seam. Feathering has to *blend*
there, not accumulate.
So `marginalise` takes one softmax over all 150 channels and sums within
each category. Grouping cannot change a total of one, so the listed
categories plus the unlisted remainder sum to one at every pixel, by
construction rather than by normalising afterwards. `parse_categories`
refuses a descriptor that claims a class twice, because that is the one
input that would quietly make the property untrue.
## The descriptor is data, and hand-written
`models/scene/categories.txt` groups ADE20K's 150 classes into the eight
a photographer would recognise. It is a file rather than a table in Rust
for the reason `models/LICENCE.md` predicted — a vocabulary is model
metadata — and it is line-oriented with comments rather than JSON like
the `.classes.json` beside it, because that file is generated and this
one is argued. Why `swimming pool` is water and not architecture belongs
next to the line that says so.
Classes are named, not indexed. An index is silently wrong after a
re-export; a name is loudly wrong, and the loader refuses one the model
does not have.
## Resolution, kept visible
`Scene` holds the native 80×80 logit grid and resamples on demand rather
than upsampling once at load. The coarseness is real — it is what the
graph produces — and a type that hides it behind an early resize invites
callers to expect detail that was never there. `rasterise` is where the
letterbox inverse lives, once.
`Letterbox` and `Window` become `pub(crate)` and `to_proto` generalises
to `to_grid`, because both dense outputs this crate reads are an even
fraction of the same letterboxed square and differ only in the divisor.
## Verified by looking, which is the only way this gets verified
`examples/scene.rs` writes the photograph dimmed outside each category. A
transposed axis or an off-by-one in the inverse produces perfectly
plausible weights over slightly the wrong pixels, and no unit test
catches that. On an indoor frame the person mask lands on the person,
including the outstretched arm, and sky reads ~5% against a bright
ceiling.
It doubles as the benchmark, because every timing quoted while this model
was chosen came off a laptop compiling other things and none of them
belong in a document.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`cargo fmt --check` is a required step and had drifted across 45 files. Most of
it arrived this week: several operations were written in parallel worktrees and
merged by hand, and a hand-merge resolves conflicts without ever running the
formatter over the result.
No behaviour changes — this is `cargo fmt --all` and nothing else, kept as its
own commit so the next reader can skip it wholesale rather than search it for
one that matters.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Local masking needs to know where an image's regions are. The watershed
spike (S15 arm A) found the boundaries but had no idea what any of them
enclosed; its coarse levels were geometric accidents. This adds the other
half and the thing that joins them.
`core/dr-segment` is where region reasoning now lives — the hierarchy moves
out of `dr-gpu`, which keeps only the pixel passes that are genuinely
shaders. The new crate is device-free and, without its default features,
model-free too: 20 of its tests need neither an adapter nor 11 MB of
weights.
Arm B runs YOLO26n-seg through `ort`. D13 framed inference as a choice
between `ort`'s C++ runtime and the pure-Rust dependency policy; that was a
false choice. `ort`'s `alternative-backend` feature unlinks the C entirely
and `ort-tract` supplies the API from tract, which is pure Rust. Measured
before committing to it: zero unsupported operators, 420 ms for 640x640,
and correct masks on bus.jpg. No NDK problem to solve, so D13's largest
tolerated exception is not needed.
Arm C is `prior.rs`, and it ships because the two arms fail in opposite
directions. Instance membership re-weights the merge saddles, so region
pairs the model believes share an object merge early and pairs straddling
its edge merge late. No boundary moves — only the order in which they
dissolve — which is how the result stays pixel-accurate at every level
while its coarse levels become named things.
Two things the spec assumed that turned out to be false, both recorded in
models/LICENCE.md: there is no usable ADE20K-trained YOLO, so the shipped
vocabulary is COCO's 80 subjects and *stuff* like sky and foliage must come
from arm A; and tract cannot parse a dynamic-shape export, so the graph's
input is fixed and tiling is the only route to more semantic resolution.
Weights are AGPL-3.0, which GPLv3 §13 permits and which makes the combined
work effectively AGPL. Deliberate, not accidental. They live in Git LFS,
and a build script fails with an instruction rather than embedding a
pointer file when the clone lacks them.