Files
DarkRoom/core/dr-segment/src/lib.rs
T
dtourolleandClaude Opus 5 8df6000e4b Decode the scene model into per-category weights
The weights landed last commit with nothing to read them. This is the
decoder, and the shape of it follows from one property worth stating
before the code: the categories must partition the image.

## Why a partition, and not a mask per category

The scene tab applies one grade to every pixel of a category — lift the
sky, desaturate foliage — and both grades meet at the horizon. If each
category carried an independent mask, feathering them outward would make
the boundary band belong to both, so both grades would land there and
every horizon would acquire a visible seam. Feathering has to *blend*
there, not accumulate.

So `marginalise` takes one softmax over all 150 channels and sums within
each category. Grouping cannot change a total of one, so the listed
categories plus the unlisted remainder sum to one at every pixel, by
construction rather than by normalising afterwards. `parse_categories`
refuses a descriptor that claims a class twice, because that is the one
input that would quietly make the property untrue.

## The descriptor is data, and hand-written

`models/scene/categories.txt` groups ADE20K's 150 classes into the eight
a photographer would recognise. It is a file rather than a table in Rust
for the reason `models/LICENCE.md` predicted — a vocabulary is model
metadata — and it is line-oriented with comments rather than JSON like
the `.classes.json` beside it, because that file is generated and this
one is argued. Why `swimming pool` is water and not architecture belongs
next to the line that says so.

Classes are named, not indexed. An index is silently wrong after a
re-export; a name is loudly wrong, and the loader refuses one the model
does not have.

## Resolution, kept visible

`Scene` holds the native 80×80 logit grid and resamples on demand rather
than upsampling once at load. The coarseness is real — it is what the
graph produces — and a type that hides it behind an early resize invites
callers to expect detail that was never there. `rasterise` is where the
letterbox inverse lives, once.

`Letterbox` and `Window` become `pub(crate)` and `to_proto` generalises
to `to_grid`, because both dense outputs this crate reads are an even
fraction of the same letterboxed square and differ only in the divisor.

## Verified by looking, which is the only way this gets verified

`examples/scene.rs` writes the photograph dimmed outside each category. A
transposed axis or an off-by-one in the inverse produces perfectly
plausible weights over slightly the wrong pixels, and no unit test
catches that. On an indoor frame the person mask lands on the person,
including the outstretched arm, and sky reads ~5% against a bright
ceiling.

It doubles as the benchmark, because every timing quoted while this model
was chosen came off a laptop compiling other things and none of them
belong in a document.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 10:48:37 +02:00

81 lines
3.6 KiB
Rust

//! Region segmentation for local masking (S15, docs/segmentation.md).
//!
//! Local adjustments need to know where the image's regions are before they
//! can snap a mask to one. This crate is that map, and it is deliberately
//! **device-free**: the watershed's pixel passes live in `dr-gpu` because they
//! are shaders, and everything that reasons about *regions* rather than
//! *pixels* lives here, where it can be tested on hand-built inputs with no
//! adapter present (ARCH §6.5a).
//!
//! # The three arms
//!
//! [`hierarchy`] is **arm A** — a watershed over-segments the image and the
//! recorded merge order becomes a granularity ladder. Deterministic, needs no
//! model, works on any picture, and knows nothing about what anything *is*.
//!
//! [`semantic`] is **arm B** — a YOLO instance-segmentation model naming the
//! subjects it recognises. Knows what things are, and is vague about exactly
//! where their edges fall (its prototypes are quarter-resolution).
//!
//! [`prior`] is **arm C**, and it is the one that ships. Arm B's instances
//! *re-weight* arm A's merge order, so coarse levels of the ladder line up
//! with real objects while every boundary stays exactly where the watershed
//! put it. The model contributes what it is good at — knowing what things are
//! — and the watershed contributes what it is good at, which is knowing where
//! the edge is, to the pixel, at every scale.
//!
//! That combination is also what repairs the vocabulary problem. The shipped
//! model is COCO-trained, so it recognises subjects and has no class for sky,
//! foliage or wall (`models/LICENCE.md`). Selecting those falls to arm A,
//! which never needed a vocabulary to begin with.
//!
//! # And [`scene`], which is not one of the arms
//!
//! The three arms all serve *local* adjustment: they exist so a mask can be
//! snapped to one region of the picture. [`scene`] serves the opposite move —
//! one grade applied to every pixel of a category at once, sky or foliage or
//! water — and reads a second, ADE20K-trained model to do it. It shares this
//! crate because it shares the runtime and the letterbox, not because it is
//! another way of doing the same thing.
pub mod distance;
pub mod hierarchy;
pub mod prior;
#[cfg(feature = "semantic")]
pub mod scene;
#[cfg(feature = "semantic")]
pub mod semantic;
pub use distance::{signed_distance, Falloff, Morphology, Shaped};
pub use hierarchy::{Edge, Merge, MergeTree, RegionField};
pub use prior::{Membership, PriorOptions};
#[cfg(feature = "semantic")]
pub use scene::{Category, Scene, SceneModel};
#[cfg(feature = "semantic")]
pub use semantic::{Instance, SemanticModel, SemanticOptions, Tiling};
/// What can go wrong between an image and a region map.
#[derive(Debug, thiserror::Error)]
pub enum SegmentError {
#[error("could not read model file: {0}")]
ModelRead(#[source] std::io::Error),
#[cfg(feature = "semantic")]
#[error("inference failed: {0}")]
Inference(#[source] ort::Error),
#[error("image buffer is {got} floats, expected {expected} (RGB, three per pixel)")]
ImageShape { expected: usize, got: usize },
/// The graph produced something the decoder does not recognise — a
/// different model, or a different export of the same one.
#[error("model output '{0}' did not have the expected shape")]
OutputShape(&'static str),
/// `models/scene/categories.txt` and the model disagree, or the descriptor
/// is malformed. Its own variant rather than a parse error because every
/// case carries a specific sentence about what to fix.
#[error("category descriptor: {0}")]
CategoryDescriptor(String),
}