Files
DarkRoom/core/dr-segment/examples/scene.rs
T
dtourolleandClaude Opus 5 4f4abd335f Cut a scene category back to the pixels that agree with it
A flag in the sky came out weighted as sky, and no feather setting fixed it.

The scene model's logits are `[1, 150, 80, 80]`, so one cell is eight input
pixels; at the 1600px proxy the letterbox scale is 0.4 and **one cell is 20
proxy pixels**, which `rasterise`'s bilinear then spreads across one more
either side. A flag is a handful of cells whose softmax is dominated by the
sky around it. The information was never in the grid, so nothing downstream
of the grid can recover it.

Tiling is the answer for an instance and is not available here: a category
has no bounding box to tile over — sky is wherever the sky is. But the
photograph is at full proxy resolution even though the weights are not, and
it knows exactly where the flag is. So the model says *what*, and the
pixels say *which of them*, which is the division of labour arm C already
draws between the instance model and the watershed.

## Seeds, and why the erosion radius is not a guess

Threshold the weights high, take `signed_distance`, and keep what is more
than 1.5 cells inside. One cell *is* the model's resolution and the bilinear
spreads it across one more, so the band either side of the boundary is smear
rather than evidence. Deriving the radius from `Scene::cell_pixels` rather
than picking a pixel count means it stays right if the proxy edge or the
export changes.

The mirror of that set is a confident *exterior*, free from the same field.

## Dropping small modes is the step that makes it work

Four k-means modes per side, not one Gaussian: sky is blue at the zenith,
white where the cloud is and pale at the horizon, and one blob over all
three rejects two of them.

Then modes holding under 3% of a side are discarded, and without that step
the whole thing fails on the case it was built for. A small flag deep in the
sky has both a high weight and a large distance from the boundary, so it
lands in the interior sample and teaches the model its own colour. It cannot
be excluded geometrically. It can be excluded by share.

Luminance is weighted at a quarter against chrominance for the same reason
the watershed's gradient is. Sky's variance is dominated by luminance, so at
equal weight the distribution is a long bright streak that a mid-grey flag
sits comfortably inside. A flag is separated by chrominance; a cloud is
separated by luminance alone. Not zero, or a dark bird against a bright sky
survives.

## Two tests, because either alone is wrong

Absolute — is this colour plausible under the category, as a chi-square on
the Mahalanobis distance. Comparative — is it likelier inside than outside.
A pixel must pass both.

The absolute test is what catches the flag, whose colour is far from *both*
sides and which the comparative test alone would leave at even odds. The
comparative test is what stops the absolute one needing a constant tuned per
category.

## What this cannot do, written down rather than left to be discovered

An intruder large enough to hold its own mode is kept. By share, a flag over
a fifth of the sky and a cloud bank over a fifth of the sky are the same
object, and colour does not separate them either — a white cloud is as far
from blue sky in chrominance as many intruders are.

So `min_cluster` is not a threshold with a correct value waiting to be
found; it is the trade-off itself, set where a photographic intruder falls.
Both ends are pinned by tests — `a_flag_in_the_sky_is_removed` and
`an_intruder_larger_than_min_cluster_survives` — so that moving the number
reads as moving the trade-off rather than as fixing a bug. The case left
open is a large unrecognised object in a clean category, which wants the
boundary snapped to watershed basins and is a different mechanism.

## Safe to apply without a control

It is subtractive: the output is the input times a factor in `0..=1`. The
worst failure available to it is losing part of a real sky, never gaining a
region, so a blue car below the horizon that was never in the mask cannot be
pulled into it. And a factor in `0..=1` cannot raise a sum, so `scene.rs`'s
partition still holds when every category is refined independently — the
weight taken off the flag lands in the unlisted remainder, which is where a
flag belongs, ADE20K having no class for one.

Every path without the evidence to judge returns the weights untouched and
says which path it took. A refinement that silently did nothing is
indistinguishable from the feature being off, and an empty seed set fitted
to a distribution would reject every pixel.

The signature is deliberately unchanged: categories are addressed by name,
not by index, so a sharper mask cannot create the stale-index hazard the
signature exists to guard against.

The example writes `<prefix>-<category>-refined.ppm` beside the coarse one,
never instead of it — whether this is an improvement is a comparative
judgement and one image cannot answer it.

Verified: fmt clean, clippy -D warnings clean, 57 dr-segment tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 18:30:16 +02:00

217 lines
8.6 KiB
Rust
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
//! Run the scene model over a JPEG, time it, and write what it saw.
//!
//! Two jobs in one example because they need the same setup and answering
//! either one alone leaves the other open.
//!
//! **Looking.** Same argument as `detect`: no unit test settles whether the
//! letterbox inverse in `Scene::rasterise` is right, because an off-by-one
//! produces perfectly plausible weights over slightly the wrong pixels. A sky
//! mask laid over the photograph settles it in one glance.
//!
//! **Timing.** Every number quoted while this model was being chosen came off a
//! laptop that was compiling other things at the time, which makes them upper
//! bounds and nothing better. This exists so the figure that ends up in a
//! document came from a quiet machine and can be reproduced on another one.
//!
//! ```sh
//! cargo run -p dr-segment --example scene --release --features embedded-scene-model -- photo.jpg
//! cargo run -p dr-segment --example scene --release -- photo.jpg out 20 \
//! models/scene/yolo26s-sem-ade20k.onnx
//! ```
//!
//! Writes `<prefix>-<category>.ppm` per category — the photograph darkened
//! where the category is absent, so the mask is legible *against the picture it
//! came from* rather than as an abstract grey field. PPM for the same reason
//! the other examples use it: no encoder dependency, and every viewer reads it.
//!
//! And `<prefix>-<category>-refined.ppm` beside it, which is the same category
//! after [`dr_segment::refine_category`] has cut it back to the pixels whose
//! colour agrees with it. Both, never one: whether that refinement is an
//! improvement is a comparative judgement — did the flag come out of the sky,
//! and is the sky still there afterwards — and a single image cannot answer
//! it. The percentage printed beside each is how much weight came off, which
//! is the number to be suspicious of when it is large.
//!
//! Timings are reported as a median over the requested run count, with the
//! first run excluded. That first pass pays for tract's lazy allocation and is
//! not representative of the second image a session decodes.
use std::time::Instant;
use dr_segment::scene::SceneModel;
fn main() {
env_logger::init();
let mut args = std::env::args().skip(1);
let Some(path) = args.next() else {
eprintln!(
"usage: scene <photo.jpg> [out-prefix] [runs] [model.onnx classes.json categories.txt]"
);
eprintln!(" with --features embedded-scene-model the model arguments may be omitted");
std::process::exit(2);
};
let prefix = args.next().unwrap_or_else(|| "scene".into());
let runs: usize = args
.next()
.and_then(|r| r.parse().ok())
.unwrap_or(10)
.max(1);
let (rgb, width, height) = read_jpeg(&path);
println!("{path}: {width}×{height}");
let mut model = match (args.next(), args.next(), args.next()) {
(Some(m), Some(c), Some(g)) => {
SceneModel::from_path(m, c, g).expect("could not load the scene model")
}
_ => embedded(),
};
// Excluded from the statistics deliberately — see the header.
let warm = Instant::now();
let scene = model
.analyse(&rgb, width, height)
.expect("inference failed");
println!("first run: {:?} (allocation included)", warm.elapsed());
let mut times: Vec<f64> = Vec::with_capacity(runs);
for _ in 0..runs {
let start = Instant::now();
let _ = model
.analyse(&rgb, width, height)
.expect("inference failed");
times.push(start.elapsed().as_secs_f64() * 1000.0);
}
times.sort_by(f64::total_cmp);
println!(
"{runs} runs: median {:.0} ms (min {:.0}, max {:.0})",
times[times.len() / 2],
times[0],
times[times.len() - 1],
);
let (gw, gh) = scene.grid_size();
println!("logit grid: {gw}×{gh}");
println!();
// Coverage first and sorted, because on any given photograph most
// categories are absent and the two or three that are not are the whole
// story.
let mut ranked: Vec<(usize, f32)> = (0..scene.categories().len())
.map(|k| (k, scene.coverage(k)))
.collect();
ranked.sort_by(|a, b| b.1.total_cmp(&a.1));
for (k, coverage) in ranked {
let name = &scene.categories()[k];
println!("{name:>14} {:5.1}%", coverage * 100.0);
// A category covering essentially nothing produces a black image and a
// file nobody wants; the threshold is what the scene tab would use to
// decide whether to offer a slider at all.
if coverage < 0.005 {
continue;
}
let mask = scene
.rasterise(k, width, height)
.expect("category index came from the same Scene");
write_overlay(&format!("{prefix}-{name}.ppm"), &rgb, &mask, width, height);
// The same category cut back to the pixels whose colour agrees with
// it, written *beside* the coarse one rather than instead of it. The
// judgement this example exists to support is comparative — is the
// flag out, and is the sky still there — and it cannot be made from
// one image.
let start = Instant::now();
let (refined, what) = dr_segment::refine_category(
&mask,
&rgb,
width,
height,
scene.cell_pixels(),
&dr_segment::RefineOptions::default(),
);
let took = start.elapsed();
match what {
dr_segment::Refined::Applied { removed } => {
println!(
" refined in {took:?}: {:.1}% of the weight cut",
removed * 100.0
);
write_overlay(
&format!("{prefix}-{name}-refined.ppm"),
&rgb,
&refined,
width,
height,
);
}
dr_segment::Refined::Skipped(why) => {
println!(" left coarse: {why:?}");
}
}
}
}
#[cfg(feature = "embedded-scene-model")]
fn embedded() -> SceneModel {
SceneModel::embedded().expect("could not load the embedded scene model")
}
#[cfg(not(feature = "embedded-scene-model"))]
fn embedded() -> SceneModel {
eprintln!(
"no model given, and this build has no embedded one.\n\
Either pass the three paths, or rebuild with --features embedded-scene-model."
);
std::process::exit(2);
}
/// The photograph, dimmed where the category is not.
///
/// Not a bare greyscale mask: the question being asked is "does this weight
/// land on the sky", and a mask on its own cannot answer it — you have to see
/// the sky underneath. A floor rather than a multiply, so that a region the
/// model gave up on is still visible enough to recognise.
fn write_overlay(path: &str, rgb: &[f32], mask: &[f32], width: usize, height: usize) {
let mut out = String::with_capacity(64);
out.push_str(&format!("P3\n{width} {height}\n255\n"));
let mut bytes = out.into_bytes();
for i in 0..width * height {
let w = mask[i].clamp(0.0, 1.0);
let gain = 0.15 + 0.85 * w;
for c in 0..3 {
let v = (rgb[i * 3 + c] * gain * 255.0).clamp(0.0, 255.0) as u8;
bytes.extend_from_slice(v.to_string().as_bytes());
bytes.push(if c == 2 { b'\n' } else { b' ' });
}
}
match std::fs::write(path, bytes) {
Ok(()) => println!(" wrote {path}"),
Err(e) => eprintln!(" could not write {path}: {e}"),
}
}
/// Decode to the tightly packed `f32` RGB the model wants.
fn read_jpeg(path: &str) -> (Vec<f32>, usize, usize) {
let bytes = std::fs::read(path).expect("could not read the photograph");
let mut decoder = zune_jpeg::JpegDecoder::new(&bytes);
let pixels = decoder.decode().expect("could not decode the photograph");
let info = decoder.info().expect("decoded image has no dimensions");
let (width, height) = (info.width as usize, info.height as usize);
// zune hands back whatever the file had. Three channels is the ordinary
// case; one is a greyscale scan, which is worth handling because a
// black-and-white frame is exactly the kind of thing someone reaches for
// when a colour one looks wrong.
let components = pixels.len() / (width * height);
let rgb = match components {
3 => pixels.iter().map(|&p| p as f32 / 255.0).collect(),
1 => pixels.iter().flat_map(|&p| [p as f32 / 255.0; 3]).collect(),
n => panic!("unsupported component count: {n}"),
};
(rgb, width, height)
}