Decode the scene model into per-category weights

The weights landed last commit with nothing to read them. This is the
decoder, and the shape of it follows from one property worth stating
before the code: the categories must partition the image.

## Why a partition, and not a mask per category

The scene tab applies one grade to every pixel of a category — lift the
sky, desaturate foliage — and both grades meet at the horizon. If each
category carried an independent mask, feathering them outward would make
the boundary band belong to both, so both grades would land there and
every horizon would acquire a visible seam. Feathering has to *blend*
there, not accumulate.

So `marginalise` takes one softmax over all 150 channels and sums within
each category. Grouping cannot change a total of one, so the listed
categories plus the unlisted remainder sum to one at every pixel, by
construction rather than by normalising afterwards. `parse_categories`
refuses a descriptor that claims a class twice, because that is the one
input that would quietly make the property untrue.

## The descriptor is data, and hand-written

`models/scene/categories.txt` groups ADE20K's 150 classes into the eight
a photographer would recognise. It is a file rather than a table in Rust
for the reason `models/LICENCE.md` predicted — a vocabulary is model
metadata — and it is line-oriented with comments rather than JSON like
the `.classes.json` beside it, because that file is generated and this
one is argued. Why `swimming pool` is water and not architecture belongs
next to the line that says so.

Classes are named, not indexed. An index is silently wrong after a
re-export; a name is loudly wrong, and the loader refuses one the model
does not have.

## Resolution, kept visible

`Scene` holds the native 80×80 logit grid and resamples on demand rather
than upsampling once at load. The coarseness is real — it is what the
graph produces — and a type that hides it behind an early resize invites
callers to expect detail that was never there. `rasterise` is where the
letterbox inverse lives, once.

`Letterbox` and `Window` become `pub(crate)` and `to_proto` generalises
to `to_grid`, because both dense outputs this crate reads are an even
fraction of the same letterboxed square and differ only in the divisor.

## Verified by looking, which is the only way this gets verified

`examples/scene.rs` writes the photograph dimmed outside each category. A
transposed axis or an off-by-one in the inverse produces perfectly
plausible weights over slightly the wrong pixels, and no unit test
catches that. On an indoor frame the person mask lands on the person,
including the outstretched arm, and sky reads ~5% against a bright
ceiling.

It doubles as the benchmark, because every timing quoted while this model
was chosen came off a laptop compiling other things and none of them
belong in a document.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-30 10:48:37 +02:00
co-authored by Claude Opus 5
parent 79b7f54e04
commit 8df6000e4b
8 changed files with 922 additions and 56 deletions
+174
View File
@@ -0,0 +1,174 @@
//! Run the scene model over a JPEG, time it, and write what it saw.
//!
//! Two jobs in one example because they need the same setup and answering
//! either one alone leaves the other open.
//!
//! **Looking.** Same argument as `detect`: no unit test settles whether the
//! letterbox inverse in `Scene::rasterise` is right, because an off-by-one
//! produces perfectly plausible weights over slightly the wrong pixels. A sky
//! mask laid over the photograph settles it in one glance.
//!
//! **Timing.** Every number quoted while this model was being chosen came off a
//! laptop that was compiling other things at the time, which makes them upper
//! bounds and nothing better. This exists so the figure that ends up in a
//! document came from a quiet machine and can be reproduced on another one.
//!
//! ```sh
//! cargo run -p dr-segment --example scene --release --features embedded-scene-model -- photo.jpg
//! cargo run -p dr-segment --example scene --release -- photo.jpg out 20 \
//! models/scene/yolo26s-sem-ade20k.onnx
//! ```
//!
//! Writes `<prefix>-<category>.ppm` per category — the photograph darkened
//! where the category is absent, so the mask is legible *against the picture it
//! came from* rather than as an abstract grey field. PPM for the same reason
//! the other examples use it: no encoder dependency, and every viewer reads it.
//!
//! Timings are reported as a median over the requested run count, with the
//! first run excluded. That first pass pays for tract's lazy allocation and is
//! not representative of the second image a session decodes.
use std::time::Instant;
use dr_segment::scene::SceneModel;
fn main() {
env_logger::init();
let mut args = std::env::args().skip(1);
let Some(path) = args.next() else {
eprintln!(
"usage: scene <photo.jpg> [out-prefix] [runs] [model.onnx classes.json categories.txt]"
);
eprintln!(" with --features embedded-scene-model the model arguments may be omitted");
std::process::exit(2);
};
let prefix = args.next().unwrap_or_else(|| "scene".into());
let runs: usize = args
.next()
.and_then(|r| r.parse().ok())
.unwrap_or(10)
.max(1);
let (rgb, width, height) = read_jpeg(&path);
println!("{path}: {width}×{height}");
let mut model = match (args.next(), args.next(), args.next()) {
(Some(m), Some(c), Some(g)) => {
SceneModel::from_path(m, c, g).expect("could not load the scene model")
}
_ => embedded(),
};
// Excluded from the statistics deliberately — see the header.
let warm = Instant::now();
let scene = model
.analyse(&rgb, width, height)
.expect("inference failed");
println!("first run: {:?} (allocation included)", warm.elapsed());
let mut times: Vec<f64> = Vec::with_capacity(runs);
for _ in 0..runs {
let start = Instant::now();
let _ = model
.analyse(&rgb, width, height)
.expect("inference failed");
times.push(start.elapsed().as_secs_f64() * 1000.0);
}
times.sort_by(f64::total_cmp);
println!(
"{runs} runs: median {:.0} ms (min {:.0}, max {:.0})",
times[times.len() / 2],
times[0],
times[times.len() - 1],
);
let (gw, gh) = scene.grid_size();
println!("logit grid: {gw}×{gh}");
println!();
// Coverage first and sorted, because on any given photograph most
// categories are absent and the two or three that are not are the whole
// story.
let mut ranked: Vec<(usize, f32)> = (0..scene.categories().len())
.map(|k| (k, scene.coverage(k)))
.collect();
ranked.sort_by(|a, b| b.1.total_cmp(&a.1));
for (k, coverage) in ranked {
let name = &scene.categories()[k];
println!("{name:>14} {:5.1}%", coverage * 100.0);
// A category covering essentially nothing produces a black image and a
// file nobody wants; the threshold is what the scene tab would use to
// decide whether to offer a slider at all.
if coverage < 0.005 {
continue;
}
let mask = scene
.rasterise(k, width, height)
.expect("category index came from the same Scene");
write_overlay(&format!("{prefix}-{name}.ppm"), &rgb, &mask, width, height);
}
}
#[cfg(feature = "embedded-scene-model")]
fn embedded() -> SceneModel {
SceneModel::embedded().expect("could not load the embedded scene model")
}
#[cfg(not(feature = "embedded-scene-model"))]
fn embedded() -> SceneModel {
eprintln!(
"no model given, and this build has no embedded one.\n\
Either pass the three paths, or rebuild with --features embedded-scene-model."
);
std::process::exit(2);
}
/// The photograph, dimmed where the category is not.
///
/// Not a bare greyscale mask: the question being asked is "does this weight
/// land on the sky", and a mask on its own cannot answer it — you have to see
/// the sky underneath. A floor rather than a multiply, so that a region the
/// model gave up on is still visible enough to recognise.
fn write_overlay(path: &str, rgb: &[f32], mask: &[f32], width: usize, height: usize) {
let mut out = String::with_capacity(64);
out.push_str(&format!("P3\n{width} {height}\n255\n"));
let mut bytes = out.into_bytes();
for i in 0..width * height {
let w = mask[i].clamp(0.0, 1.0);
let gain = 0.15 + 0.85 * w;
for c in 0..3 {
let v = (rgb[i * 3 + c] * gain * 255.0).clamp(0.0, 255.0) as u8;
bytes.extend_from_slice(v.to_string().as_bytes());
bytes.push(if c == 2 { b'\n' } else { b' ' });
}
}
match std::fs::write(path, bytes) {
Ok(()) => println!(" wrote {path}"),
Err(e) => eprintln!(" could not write {path}: {e}"),
}
}
/// Decode to the tightly packed `f32` RGB the model wants.
fn read_jpeg(path: &str) -> (Vec<f32>, usize, usize) {
let bytes = std::fs::read(path).expect("could not read the photograph");
let mut decoder = zune_jpeg::JpegDecoder::new(&bytes);
let pixels = decoder.decode().expect("could not decode the photograph");
let info = decoder.info().expect("decoded image has no dimensions");
let (width, height) = (info.width as usize, info.height as usize);
// zune hands back whatever the file had. Three channels is the ordinary
// case; one is a greyscale scan, which is worth handling because a
// black-and-white frame is exactly the kind of thing someone reaches for
// when a colour one looks wrong.
let components = pixels.len() / (width * height);
let rgb = match components {
3 => pixels.iter().map(|&p| p as f32 / 255.0).collect(),
1 => pixels.iter().flat_map(|&p| [p as f32 / 255.0; 3]).collect(),
n => panic!("unsupported component count: {n}"),
};
(rgb, width, height)
}