Feathering, growing, shrinking, closing and opening are the same number read differently. With the signed distance from the boundary in hand, dilation is the set where d >= -r, erosion where d >= +r, and a feather of any shape is a function of d. So the field is computed once and the controls are arithmetic on it. The **field** is what reaches the GPU, not a finished alpha, and that is the point: growing a mask or changing its falloff then costs a uniform upload and no recomputation, which is what makes them live controls rather than ones that stall on every drag. Only closing and opening rebuild, because after the first threshold the shape has changed and the old distances describe the old one. Exact Euclidean, via Felzenszwalb's separable transform — not a chamfer approximation, which leaves a mask visibly octagonal once grown more than a few pixels. A test asserts the diagonal is √2 rather than 1 or 2. It runs on the CPU, which ARCH §5.4 forbids for masks. The rule is about brush lag — a stroke rasterised per frame — and this is a different operation: once per mask edit, on input the model already produced here, producing a field the GPU then samples for free. What it buys is exact determinism, which matters because masks reach the sidecar as indices and a field that varied by vendor would mean a mask meaning one thing on the desktop and another on the phone. The half-pixel in `signed_distance` is not a detail, and a test caught it. Measuring to the nearest opposite pixel *centre* puts the smallest magnitude at 1 either side, so the boundary is nowhere and **eroding by less than a pixel removes nothing**. A control whose first notch does nothing is a broken control. Half a pixel off each side puts the boundary where it physically is, and eroding by 1 takes exactly the outermost ring. Every falloff curve is 0.5 at the boundary by construction, asserted for all five: changing the curve should change how the transition looks and never where it sits.
62 lines
2.7 KiB
Rust
62 lines
2.7 KiB
Rust
//! Region segmentation for local masking (S15, docs/segmentation.md).
|
|
//!
|
|
//! Local adjustments need to know where the image's regions are before they
|
|
//! can snap a mask to one. This crate is that map, and it is deliberately
|
|
//! **device-free**: the watershed's pixel passes live in `dr-gpu` because they
|
|
//! are shaders, and everything that reasons about *regions* rather than
|
|
//! *pixels* lives here, where it can be tested on hand-built inputs with no
|
|
//! adapter present (ARCH §6.5a).
|
|
//!
|
|
//! # The three arms
|
|
//!
|
|
//! [`hierarchy`] is **arm A** — a watershed over-segments the image and the
|
|
//! recorded merge order becomes a granularity ladder. Deterministic, needs no
|
|
//! model, works on any picture, and knows nothing about what anything *is*.
|
|
//!
|
|
//! [`semantic`] is **arm B** — a YOLO instance-segmentation model naming the
|
|
//! subjects it recognises. Knows what things are, and is vague about exactly
|
|
//! where their edges fall (its prototypes are quarter-resolution).
|
|
//!
|
|
//! [`prior`] is **arm C**, and it is the one that ships. Arm B's instances
|
|
//! *re-weight* arm A's merge order, so coarse levels of the ladder line up
|
|
//! with real objects while every boundary stays exactly where the watershed
|
|
//! put it. The model contributes what it is good at — knowing what things are
|
|
//! — and the watershed contributes what it is good at, which is knowing where
|
|
//! the edge is, to the pixel, at every scale.
|
|
//!
|
|
//! That combination is also what repairs the vocabulary problem. The shipped
|
|
//! model is COCO-trained, so it recognises subjects and has no class for sky,
|
|
//! foliage or wall (`models/LICENCE.md`). Selecting those falls to arm A,
|
|
//! which never needed a vocabulary to begin with.
|
|
|
|
pub mod distance;
|
|
pub mod hierarchy;
|
|
pub mod prior;
|
|
#[cfg(feature = "semantic")]
|
|
pub mod semantic;
|
|
|
|
pub use distance::{signed_distance, Falloff, Morphology, Shaped};
|
|
pub use hierarchy::{Edge, Merge, MergeTree, RegionField};
|
|
pub use prior::{Membership, PriorOptions};
|
|
#[cfg(feature = "semantic")]
|
|
pub use semantic::{Instance, SemanticModel, SemanticOptions, Tiling};
|
|
|
|
/// What can go wrong between an image and a region map.
|
|
#[derive(Debug, thiserror::Error)]
|
|
pub enum SegmentError {
|
|
#[error("could not read model file: {0}")]
|
|
ModelRead(#[source] std::io::Error),
|
|
|
|
#[cfg(feature = "semantic")]
|
|
#[error("inference failed: {0}")]
|
|
Inference(#[source] ort::Error),
|
|
|
|
#[error("image buffer is {got} floats, expected {expected} (RGB, three per pixel)")]
|
|
ImageShape { expected: usize, got: usize },
|
|
|
|
/// The graph produced something the decoder does not recognise — a
|
|
/// different model, or a different export of the same one.
|
|
#[error("model output '{0}' did not have the expected shape")]
|
|
OutputShape(&'static str),
|
|
}
|