Let the model say what a thing is and the watershed say where it ends

Local masking needs to know where an image's regions are. The watershed
spike (S15 arm A) found the boundaries but had no idea what any of them
enclosed; its coarse levels were geometric accidents. This adds the other
half and the thing that joins them.

`core/dr-segment` is where region reasoning now lives — the hierarchy moves
out of `dr-gpu`, which keeps only the pixel passes that are genuinely
shaders. The new crate is device-free and, without its default features,
model-free too: 20 of its tests need neither an adapter nor 11 MB of
weights.

Arm B runs YOLO26n-seg through `ort`. D13 framed inference as a choice
between `ort`'s C++ runtime and the pure-Rust dependency policy; that was a
false choice. `ort`'s `alternative-backend` feature unlinks the C entirely
and `ort-tract` supplies the API from tract, which is pure Rust. Measured
before committing to it: zero unsupported operators, 420 ms for 640x640,
and correct masks on bus.jpg. No NDK problem to solve, so D13's largest
tolerated exception is not needed.

Arm C is `prior.rs`, and it ships because the two arms fail in opposite
directions. Instance membership re-weights the merge saddles, so region
pairs the model believes share an object merge early and pairs straddling
its edge merge late. No boundary moves — only the order in which they
dissolve — which is how the result stays pixel-accurate at every level
while its coarse levels become named things.

Two things the spec assumed that turned out to be false, both recorded in
models/LICENCE.md: there is no usable ADE20K-trained YOLO, so the shipped
vocabulary is COCO's 80 subjects and *stuff* like sky and foliage must come
from arm A; and tract cannot parse a dynamic-shape export, so the graph's
input is fixed and tiling is the only route to more semantic resolution.

Weights are AGPL-3.0, which GPLv3 §13 permits and which makes the combined
work effectively AGPL. Deliberate, not accidental. They live in Git LFS,
and a build script fails with an instruction rather than embedding a
pointer file when the clone lacks them.
This commit is contained in:
2026-08-22 08:39:16 +02:00
parent ecd6df686c
commit 0da8271836
18 changed files with 2287 additions and 19 deletions
+59
View File
@@ -0,0 +1,59 @@
//! Region segmentation for local masking (S15, docs/segmentation.md).
//!
//! Local adjustments need to know where the image's regions are before they
//! can snap a mask to one. This crate is that map, and it is deliberately
//! **device-free**: the watershed's pixel passes live in `dr-gpu` because they
//! are shaders, and everything that reasons about *regions* rather than
//! *pixels* lives here, where it can be tested on hand-built inputs with no
//! adapter present (ARCH §6.5a).
//!
//! # The three arms
//!
//! [`hierarchy`] is **arm A** — a watershed over-segments the image and the
//! recorded merge order becomes a granularity ladder. Deterministic, needs no
//! model, works on any picture, and knows nothing about what anything *is*.
//!
//! [`semantic`] is **arm B** — a YOLO instance-segmentation model naming the
//! subjects it recognises. Knows what things are, and is vague about exactly
//! where their edges fall (its prototypes are quarter-resolution).
//!
//! [`prior`] is **arm C**, and it is the one that ships. Arm B's instances
//! *re-weight* arm A's merge order, so coarse levels of the ladder line up
//! with real objects while every boundary stays exactly where the watershed
//! put it. The model contributes what it is good at — knowing what things are
//! — and the watershed contributes what it is good at, which is knowing where
//! the edge is, to the pixel, at every scale.
//!
//! That combination is also what repairs the vocabulary problem. The shipped
//! model is COCO-trained, so it recognises subjects and has no class for sky,
//! foliage or wall (`models/LICENCE.md`). Selecting those falls to arm A,
//! which never needed a vocabulary to begin with.
pub mod hierarchy;
pub mod prior;
#[cfg(feature = "semantic")]
pub mod semantic;
pub use hierarchy::{Edge, Merge, MergeTree, RegionField};
pub use prior::{Membership, PriorOptions};
#[cfg(feature = "semantic")]
pub use semantic::{Instance, SemanticModel, SemanticOptions, Tiling};
/// What can go wrong between an image and a region map.
#[derive(Debug, thiserror::Error)]
pub enum SegmentError {
#[error("could not read model file: {0}")]
ModelRead(#[source] std::io::Error),
#[cfg(feature = "semantic")]
#[error("inference failed: {0}")]
Inference(#[source] ort::Error),
#[error("image buffer is {got} floats, expected {expected} (RGB, three per pixel)")]
ImageShape { expected: usize, got: usize },
/// The graph produced something the decoder does not recognise — a
/// different model, or a different export of the same one.
#[error("model output '{0}' did not have the expected shape")]
OutputShape(&'static str),
}