The weights landed last commit with nothing to read them. This is the decoder, and the shape of it follows from one property worth stating before the code: the categories must partition the image. ## Why a partition, and not a mask per category The scene tab applies one grade to every pixel of a category — lift the sky, desaturate foliage — and both grades meet at the horizon. If each category carried an independent mask, feathering them outward would make the boundary band belong to both, so both grades would land there and every horizon would acquire a visible seam. Feathering has to *blend* there, not accumulate. So `marginalise` takes one softmax over all 150 channels and sums within each category. Grouping cannot change a total of one, so the listed categories plus the unlisted remainder sum to one at every pixel, by construction rather than by normalising afterwards. `parse_categories` refuses a descriptor that claims a class twice, because that is the one input that would quietly make the property untrue. ## The descriptor is data, and hand-written `models/scene/categories.txt` groups ADE20K's 150 classes into the eight a photographer would recognise. It is a file rather than a table in Rust for the reason `models/LICENCE.md` predicted — a vocabulary is model metadata — and it is line-oriented with comments rather than JSON like the `.classes.json` beside it, because that file is generated and this one is argued. Why `swimming pool` is water and not architecture belongs next to the line that says so. Classes are named, not indexed. An index is silently wrong after a re-export; a name is loudly wrong, and the loader refuses one the model does not have. ## Resolution, kept visible `Scene` holds the native 80×80 logit grid and resamples on demand rather than upsampling once at load. The coarseness is real — it is what the graph produces — and a type that hides it behind an early resize invites callers to expect detail that was never there. `rasterise` is where the letterbox inverse lives, once. `Letterbox` and `Window` become `pub(crate)` and `to_proto` generalises to `to_grid`, because both dense outputs this crate reads are an even fraction of the same letterboxed square and differ only in the divisor. ## Verified by looking, which is the only way this gets verified `examples/scene.rs` writes the photograph dimmed outside each category. A transposed axis or an off-by-one in the inverse produces perfectly plausible weights over slightly the wrong pixels, and no unit test catches that. On an indoor frame the person mask lands on the person, including the outstretched arm, and sky reads ~5% against a bright ceiling. It doubles as the benchmark, because every timing quoted while this model was chosen came off a laptop compiling other things and none of them belong in a document. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
57 lines
2.2 KiB
TOML
57 lines
2.2 KiB
TOML
[package]
|
|
name = "dr-segment"
|
|
version.workspace = true
|
|
edition.workspace = true
|
|
rust-version.workspace = true
|
|
license.workspace = true
|
|
# Guards against a Git LFS pointer being embedded in place of the weights.
|
|
build = "build.rs"
|
|
|
|
[dependencies]
|
|
thiserror.workspace = true
|
|
log.workspace = true
|
|
|
|
# Inference. `ort` is the API; **tract is the engine** — see the workspace
|
|
# manifest for why the C++ ONNX Runtime is not linked here.
|
|
ort = { workspace = true, optional = true }
|
|
ort-tract = { workspace = true, optional = true }
|
|
ndarray = { workspace = true, optional = true }
|
|
|
|
[dev-dependencies]
|
|
# The example reads an ordinary JPEG, because the thing worth looking at is
|
|
# whether detections land on a real photograph. Pure Rust, and already in the
|
|
# tree for embedded previews.
|
|
zune-jpeg.workspace = true
|
|
env_logger.workspace = true
|
|
|
|
[features]
|
|
# On by default: a local adjustment that cannot select a subject is half the
|
|
# feature, and the whole point of the tract backend is that enabling this costs
|
|
# no C dependency on any platform.
|
|
default = ["semantic", "embedded-model"]
|
|
|
|
# Arm B — the ONNX runtime and the instance decoder.
|
|
#
|
|
# Separable because the watershed half is genuinely independent of it: with
|
|
# this off, `dr-segment` is a pure-CPU graph algorithm crate with no model to
|
|
# carry, which is what the headless hierarchy tests want.
|
|
semantic = ["dep:ort", "dep:ort-tract", "dep:ndarray"]
|
|
|
|
# Compile the weights into the binary.
|
|
#
|
|
# Separate from `semantic` because the two answer different questions. Android
|
|
# hands the app no filesystem path to read a model from (ARCH §6.9), so there
|
|
# it must be embedded; a desktop packager pointing at a system model directory,
|
|
# or a test that only needs the decoder, wants the runtime without the 11 MB.
|
|
embedded-model = ["semantic"]
|
|
|
|
# Compile the *scene* model in too, and off by default where `embedded-model`
|
|
# is on.
|
|
#
|
|
# The asymmetry is its size. At 24 MB it is more than twice the instance model,
|
|
# and Android reaches it the way it reaches the face weights — unpacked from
|
|
# APK assets at first launch — rather than by carrying it in the binary. This
|
|
# feature is for a desktop build with nowhere else to read it from, and for
|
|
# tests that want the real graph.
|
|
embedded-scene-model = ["semantic"]
|