Let the model say what a thing is and the watershed say where it ends

Local masking needs to know where an image's regions are. The watershed
spike (S15 arm A) found the boundaries but had no idea what any of them
enclosed; its coarse levels were geometric accidents. This adds the other
half and the thing that joins them.

`core/dr-segment` is where region reasoning now lives — the hierarchy moves
out of `dr-gpu`, which keeps only the pixel passes that are genuinely
shaders. The new crate is device-free and, without its default features,
model-free too: 20 of its tests need neither an adapter nor 11 MB of
weights.

Arm B runs YOLO26n-seg through `ort`. D13 framed inference as a choice
between `ort`'s C++ runtime and the pure-Rust dependency policy; that was a
false choice. `ort`'s `alternative-backend` feature unlinks the C entirely
and `ort-tract` supplies the API from tract, which is pure Rust. Measured
before committing to it: zero unsupported operators, 420 ms for 640x640,
and correct masks on bus.jpg. No NDK problem to solve, so D13's largest
tolerated exception is not needed.

Arm C is `prior.rs`, and it ships because the two arms fail in opposite
directions. Instance membership re-weights the merge saddles, so region
pairs the model believes share an object merge early and pairs straddling
its edge merge late. No boundary moves — only the order in which they
dissolve — which is how the result stays pixel-accurate at every level
while its coarse levels become named things.

Two things the spec assumed that turned out to be false, both recorded in
models/LICENCE.md: there is no usable ADE20K-trained YOLO, so the shipped
vocabulary is COCO's 80 subjects and *stuff* like sky and foliage must come
from arm A; and tract cannot parse a dynamic-shape export, so the graph's
input is fixed and tiling is the only route to more semantic resolution.

Weights are AGPL-3.0, which GPLv3 §13 permits and which makes the combined
work effectively AGPL. Deliberate, not accidental. They live in Git LFS,
and a build script fails with an instruction rather than embedding a
pointer file when the clone lacks them.
This commit is contained in:
2026-08-22 08:39:16 +02:00
parent ecd6df686c
commit 0da8271836
18 changed files with 2287 additions and 19 deletions
+149
View File
@@ -0,0 +1,149 @@
//! Run the semantic arm over a JPEG and write what it found.
//!
//! The point of S15 step 2 applied to arm B: no amount of unit testing settles
//! whether the decode is right, because a transposed axis or an off-by-one in
//! the letterbox produces perfectly plausible numbers and a mask sitting six
//! pixels to the left. Looking at the overlay settles it in one glance.
//!
//! ```sh
//! cargo run -p dr-segment --example detect --release -- photo.jpg
//! cargo run -p dr-segment --example detect --release -- photo.jpg out 0.25 tiled
//! ```
//!
//! Writes `<prefix>-overlay.ppm` — the image with each instance tinted by a
//! per-instance colour — and prints the detection list. PPM for the same
//! reason the other examples use it: no encoder dependency, and every viewer
//! reads it.
use dr_segment::semantic::{SemanticModel, SemanticOptions, Tiling};
fn main() {
env_logger::init();
let mut args = std::env::args().skip(1);
let Some(path) = args.next() else {
eprintln!("usage: detect <photo.jpg> [out-prefix] [confidence] [tiled]");
std::process::exit(2);
};
let prefix = args.next().unwrap_or_else(|| "detect".into());
let confidence = args
.next()
.and_then(|s| s.parse().ok())
.unwrap_or(SemanticOptions::default().confidence);
let tiled = args.next().is_some_and(|s| s == "tiled");
let (rgb, width, height) = load_jpeg(&path);
println!("image {width}x{height}");
let options = SemanticOptions {
confidence,
tiling: if tiled {
Tiling::Grid { overlap: 0.25 }
} else {
Tiling::Whole
},
..SemanticOptions::default()
};
println!(
"tiling {}",
if tiled { "grid, 25% overlap" } else { "whole frame" }
);
let t0 = std::time::Instant::now();
let mut model = SemanticModel::embedded().expect("load embedded model");
println!("load {:.0} ms", t0.elapsed().as_secs_f32() * 1000.0);
let t1 = std::time::Instant::now();
let instances = model
.detect(&rgb, width, height, &options)
.expect("inference");
println!(
"detect {:.0} ms",
t1.elapsed().as_secs_f32() * 1000.0
);
println!("found {} instances", instances.len());
for (i, inst) in instances.iter().enumerate() {
let covered = inst.mask.iter().filter(|&&m| m >= 0.5).count();
println!(
" [{i:2}] {:<14} {:.2} box ({:.0},{:.0})-({:.0},{:.0}) {:.1}% of frame",
inst.class_name,
inst.score,
inst.bbox.0,
inst.bbox.1,
inst.bbox.2,
inst.bbox.3,
100.0 * covered as f32 / (width * height) as f32,
);
}
// Tint each instance and write the composite. A mask in the wrong place is
// obvious here and invisible in the numbers above.
let mut out = vec![0u8; width * height * 3];
for (p, px) in out.chunks_exact_mut(3).enumerate() {
for c in 0..3 {
px[c] = (rgb[p * 3 + c].clamp(0.0, 1.0) * 255.0) as u8;
}
}
for (i, inst) in instances.iter().enumerate() {
let tint = colour(i);
for (p, &m) in inst.mask.iter().enumerate() {
if m < 0.5 {
continue;
}
let px = &mut out[p * 3..p * 3 + 3];
for c in 0..3 {
px[c] = ((px[c] as f32) * 0.45 + tint[c] as f32 * 0.55) as u8;
}
}
}
let file = format!("{prefix}-overlay.ppm");
write_ppm(&file, &out, width, height);
println!("wrote {file}");
}
/// A distinct colour per instance index — the same golden-angle walk the
/// watershed example uses, so the two overlays are read the same way.
fn colour(i: usize) -> [u8; 3] {
let h = (i as f32 * 137.508) % 360.0;
let (c, x) = (255.0, 255.0 * (1.0 - ((h / 60.0) % 2.0 - 1.0).abs()));
let (r, g, b) = match (h / 60.0) as u32 {
0 => (c, x, 0.0),
1 => (x, c, 0.0),
2 => (0.0, c, x),
3 => (0.0, x, c),
4 => (x, 0.0, c),
_ => (c, 0.0, x),
};
[r as u8, g as u8, b as u8]
}
fn load_jpeg(path: &str) -> (Vec<f32>, usize, usize) {
let bytes = std::fs::read(path).unwrap_or_else(|e| panic!("read {path}: {e}"));
let mut decoder = zune_jpeg::JpegDecoder::new(&bytes);
let pixels = decoder.decode().expect("decode jpeg");
let info = decoder.info().expect("jpeg info");
let (w, h) = (info.width as usize, info.height as usize);
// The model was trained on gamma-encoded sRGB, so the JPEG's own values go
// through unlinearised — this is one of the few places in the codebase
// where *not* linearising is the correct thing to do.
let rgb = match pixels.len() / (w * h) {
3 => pixels.iter().map(|&v| v as f32 / 255.0).collect(),
1 => pixels
.iter()
.flat_map(|&v| [v as f32 / 255.0; 3])
.collect(),
n => panic!("unexpected {n} components per pixel"),
};
(rgb, w, h)
}
fn write_ppm(path: &str, rgb: &[u8], width: usize, height: usize) {
use std::io::Write;
let mut f = std::io::BufWriter::new(std::fs::File::create(path).expect("create ppm"));
write!(f, "P6\n{width} {height}\n255\n").expect("ppm header");
f.write_all(rgb).expect("ppm body");
}