The weights landed last commit with nothing to read them. This is the decoder, and the shape of it follows from one property worth stating before the code: the categories must partition the image. ## Why a partition, and not a mask per category The scene tab applies one grade to every pixel of a category — lift the sky, desaturate foliage — and both grades meet at the horizon. If each category carried an independent mask, feathering them outward would make the boundary band belong to both, so both grades would land there and every horizon would acquire a visible seam. Feathering has to *blend* there, not accumulate. So `marginalise` takes one softmax over all 150 channels and sums within each category. Grouping cannot change a total of one, so the listed categories plus the unlisted remainder sum to one at every pixel, by construction rather than by normalising afterwards. `parse_categories` refuses a descriptor that claims a class twice, because that is the one input that would quietly make the property untrue. ## The descriptor is data, and hand-written `models/scene/categories.txt` groups ADE20K's 150 classes into the eight a photographer would recognise. It is a file rather than a table in Rust for the reason `models/LICENCE.md` predicted — a vocabulary is model metadata — and it is line-oriented with comments rather than JSON like the `.classes.json` beside it, because that file is generated and this one is argued. Why `swimming pool` is water and not architecture belongs next to the line that says so. Classes are named, not indexed. An index is silently wrong after a re-export; a name is loudly wrong, and the loader refuses one the model does not have. ## Resolution, kept visible `Scene` holds the native 80×80 logit grid and resamples on demand rather than upsampling once at load. The coarseness is real — it is what the graph produces — and a type that hides it behind an early resize invites callers to expect detail that was never there. `rasterise` is where the letterbox inverse lives, once. `Letterbox` and `Window` become `pub(crate)` and `to_proto` generalises to `to_grid`, because both dense outputs this crate reads are an even fraction of the same letterboxed square and differ only in the divisor. ## Verified by looking, which is the only way this gets verified `examples/scene.rs` writes the photograph dimmed outside each category. A transposed axis or an off-by-one in the inverse produces perfectly plausible weights over slightly the wrong pixels, and no unit test catches that. On an indoor frame the person mask lands on the person, including the outstretched arm, and sky reads ~5% against a bright ceiling. It doubles as the benchmark, because every timing quoted while this model was chosen came off a laptop compiling other things and none of them belong in a document. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
175 lines
6.8 KiB
Rust
175 lines
6.8 KiB
Rust
//! Run the scene model over a JPEG, time it, and write what it saw.
|
||
//!
|
||
//! Two jobs in one example because they need the same setup and answering
|
||
//! either one alone leaves the other open.
|
||
//!
|
||
//! **Looking.** Same argument as `detect`: no unit test settles whether the
|
||
//! letterbox inverse in `Scene::rasterise` is right, because an off-by-one
|
||
//! produces perfectly plausible weights over slightly the wrong pixels. A sky
|
||
//! mask laid over the photograph settles it in one glance.
|
||
//!
|
||
//! **Timing.** Every number quoted while this model was being chosen came off a
|
||
//! laptop that was compiling other things at the time, which makes them upper
|
||
//! bounds and nothing better. This exists so the figure that ends up in a
|
||
//! document came from a quiet machine and can be reproduced on another one.
|
||
//!
|
||
//! ```sh
|
||
//! cargo run -p dr-segment --example scene --release --features embedded-scene-model -- photo.jpg
|
||
//! cargo run -p dr-segment --example scene --release -- photo.jpg out 20 \
|
||
//! models/scene/yolo26s-sem-ade20k.onnx
|
||
//! ```
|
||
//!
|
||
//! Writes `<prefix>-<category>.ppm` per category — the photograph darkened
|
||
//! where the category is absent, so the mask is legible *against the picture it
|
||
//! came from* rather than as an abstract grey field. PPM for the same reason
|
||
//! the other examples use it: no encoder dependency, and every viewer reads it.
|
||
//!
|
||
//! Timings are reported as a median over the requested run count, with the
|
||
//! first run excluded. That first pass pays for tract's lazy allocation and is
|
||
//! not representative of the second image a session decodes.
|
||
|
||
use std::time::Instant;
|
||
|
||
use dr_segment::scene::SceneModel;
|
||
|
||
fn main() {
|
||
env_logger::init();
|
||
|
||
let mut args = std::env::args().skip(1);
|
||
let Some(path) = args.next() else {
|
||
eprintln!(
|
||
"usage: scene <photo.jpg> [out-prefix] [runs] [model.onnx classes.json categories.txt]"
|
||
);
|
||
eprintln!(" with --features embedded-scene-model the model arguments may be omitted");
|
||
std::process::exit(2);
|
||
};
|
||
let prefix = args.next().unwrap_or_else(|| "scene".into());
|
||
let runs: usize = args
|
||
.next()
|
||
.and_then(|r| r.parse().ok())
|
||
.unwrap_or(10)
|
||
.max(1);
|
||
|
||
let (rgb, width, height) = read_jpeg(&path);
|
||
println!("{path}: {width}×{height}");
|
||
|
||
let mut model = match (args.next(), args.next(), args.next()) {
|
||
(Some(m), Some(c), Some(g)) => {
|
||
SceneModel::from_path(m, c, g).expect("could not load the scene model")
|
||
}
|
||
_ => embedded(),
|
||
};
|
||
|
||
// Excluded from the statistics deliberately — see the header.
|
||
let warm = Instant::now();
|
||
let scene = model
|
||
.analyse(&rgb, width, height)
|
||
.expect("inference failed");
|
||
println!("first run: {:?} (allocation included)", warm.elapsed());
|
||
|
||
let mut times: Vec<f64> = Vec::with_capacity(runs);
|
||
for _ in 0..runs {
|
||
let start = Instant::now();
|
||
let _ = model
|
||
.analyse(&rgb, width, height)
|
||
.expect("inference failed");
|
||
times.push(start.elapsed().as_secs_f64() * 1000.0);
|
||
}
|
||
times.sort_by(f64::total_cmp);
|
||
println!(
|
||
"{runs} runs: median {:.0} ms (min {:.0}, max {:.0})",
|
||
times[times.len() / 2],
|
||
times[0],
|
||
times[times.len() - 1],
|
||
);
|
||
|
||
let (gw, gh) = scene.grid_size();
|
||
println!("logit grid: {gw}×{gh}");
|
||
println!();
|
||
|
||
// Coverage first and sorted, because on any given photograph most
|
||
// categories are absent and the two or three that are not are the whole
|
||
// story.
|
||
let mut ranked: Vec<(usize, f32)> = (0..scene.categories().len())
|
||
.map(|k| (k, scene.coverage(k)))
|
||
.collect();
|
||
ranked.sort_by(|a, b| b.1.total_cmp(&a.1));
|
||
|
||
for (k, coverage) in ranked {
|
||
let name = &scene.categories()[k];
|
||
println!("{name:>14} {:5.1}%", coverage * 100.0);
|
||
// A category covering essentially nothing produces a black image and a
|
||
// file nobody wants; the threshold is what the scene tab would use to
|
||
// decide whether to offer a slider at all.
|
||
if coverage < 0.005 {
|
||
continue;
|
||
}
|
||
let mask = scene
|
||
.rasterise(k, width, height)
|
||
.expect("category index came from the same Scene");
|
||
write_overlay(&format!("{prefix}-{name}.ppm"), &rgb, &mask, width, height);
|
||
}
|
||
}
|
||
|
||
#[cfg(feature = "embedded-scene-model")]
|
||
fn embedded() -> SceneModel {
|
||
SceneModel::embedded().expect("could not load the embedded scene model")
|
||
}
|
||
|
||
#[cfg(not(feature = "embedded-scene-model"))]
|
||
fn embedded() -> SceneModel {
|
||
eprintln!(
|
||
"no model given, and this build has no embedded one.\n\
|
||
Either pass the three paths, or rebuild with --features embedded-scene-model."
|
||
);
|
||
std::process::exit(2);
|
||
}
|
||
|
||
/// The photograph, dimmed where the category is not.
|
||
///
|
||
/// Not a bare greyscale mask: the question being asked is "does this weight
|
||
/// land on the sky", and a mask on its own cannot answer it — you have to see
|
||
/// the sky underneath. A floor rather than a multiply, so that a region the
|
||
/// model gave up on is still visible enough to recognise.
|
||
fn write_overlay(path: &str, rgb: &[f32], mask: &[f32], width: usize, height: usize) {
|
||
let mut out = String::with_capacity(64);
|
||
out.push_str(&format!("P3\n{width} {height}\n255\n"));
|
||
let mut bytes = out.into_bytes();
|
||
|
||
for i in 0..width * height {
|
||
let w = mask[i].clamp(0.0, 1.0);
|
||
let gain = 0.15 + 0.85 * w;
|
||
for c in 0..3 {
|
||
let v = (rgb[i * 3 + c] * gain * 255.0).clamp(0.0, 255.0) as u8;
|
||
bytes.extend_from_slice(v.to_string().as_bytes());
|
||
bytes.push(if c == 2 { b'\n' } else { b' ' });
|
||
}
|
||
}
|
||
|
||
match std::fs::write(path, bytes) {
|
||
Ok(()) => println!(" wrote {path}"),
|
||
Err(e) => eprintln!(" could not write {path}: {e}"),
|
||
}
|
||
}
|
||
|
||
/// Decode to the tightly packed `f32` RGB the model wants.
|
||
fn read_jpeg(path: &str) -> (Vec<f32>, usize, usize) {
|
||
let bytes = std::fs::read(path).expect("could not read the photograph");
|
||
let mut decoder = zune_jpeg::JpegDecoder::new(&bytes);
|
||
let pixels = decoder.decode().expect("could not decode the photograph");
|
||
let info = decoder.info().expect("decoded image has no dimensions");
|
||
let (width, height) = (info.width as usize, info.height as usize);
|
||
|
||
// zune hands back whatever the file had. Three channels is the ordinary
|
||
// case; one is a greyscale scan, which is worth handling because a
|
||
// black-and-white frame is exactly the kind of thing someone reaches for
|
||
// when a colour one looks wrong.
|
||
let components = pixels.len() / (width * height);
|
||
let rgb = match components {
|
||
3 => pixels.iter().map(|&p| p as f32 / 255.0).collect(),
|
||
1 => pixels.iter().flat_map(|&p| [p as f32 / 255.0; 3]).collect(),
|
||
n => panic!("unsupported component count: {n}"),
|
||
};
|
||
(rgb, width, height)
|
||
}
|