Cut a scene category back to the pixels that agree with it
A flag in the sky came out weighted as sky, and no feather setting fixed it. The scene model's logits are `[1, 150, 80, 80]`, so one cell is eight input pixels; at the 1600px proxy the letterbox scale is 0.4 and **one cell is 20 proxy pixels**, which `rasterise`'s bilinear then spreads across one more either side. A flag is a handful of cells whose softmax is dominated by the sky around it. The information was never in the grid, so nothing downstream of the grid can recover it. Tiling is the answer for an instance and is not available here: a category has no bounding box to tile over — sky is wherever the sky is. But the photograph is at full proxy resolution even though the weights are not, and it knows exactly where the flag is. So the model says *what*, and the pixels say *which of them*, which is the division of labour arm C already draws between the instance model and the watershed. ## Seeds, and why the erosion radius is not a guess Threshold the weights high, take `signed_distance`, and keep what is more than 1.5 cells inside. One cell *is* the model's resolution and the bilinear spreads it across one more, so the band either side of the boundary is smear rather than evidence. Deriving the radius from `Scene::cell_pixels` rather than picking a pixel count means it stays right if the proxy edge or the export changes. The mirror of that set is a confident *exterior*, free from the same field. ## Dropping small modes is the step that makes it work Four k-means modes per side, not one Gaussian: sky is blue at the zenith, white where the cloud is and pale at the horizon, and one blob over all three rejects two of them. Then modes holding under 3% of a side are discarded, and without that step the whole thing fails on the case it was built for. A small flag deep in the sky has both a high weight and a large distance from the boundary, so it lands in the interior sample and teaches the model its own colour. It cannot be excluded geometrically. It can be excluded by share. Luminance is weighted at a quarter against chrominance for the same reason the watershed's gradient is. Sky's variance is dominated by luminance, so at equal weight the distribution is a long bright streak that a mid-grey flag sits comfortably inside. A flag is separated by chrominance; a cloud is separated by luminance alone. Not zero, or a dark bird against a bright sky survives. ## Two tests, because either alone is wrong Absolute — is this colour plausible under the category, as a chi-square on the Mahalanobis distance. Comparative — is it likelier inside than outside. A pixel must pass both. The absolute test is what catches the flag, whose colour is far from *both* sides and which the comparative test alone would leave at even odds. The comparative test is what stops the absolute one needing a constant tuned per category. ## What this cannot do, written down rather than left to be discovered An intruder large enough to hold its own mode is kept. By share, a flag over a fifth of the sky and a cloud bank over a fifth of the sky are the same object, and colour does not separate them either — a white cloud is as far from blue sky in chrominance as many intruders are. So `min_cluster` is not a threshold with a correct value waiting to be found; it is the trade-off itself, set where a photographic intruder falls. Both ends are pinned by tests — `a_flag_in_the_sky_is_removed` and `an_intruder_larger_than_min_cluster_survives` — so that moving the number reads as moving the trade-off rather than as fixing a bug. The case left open is a large unrecognised object in a clean category, which wants the boundary snapped to watershed basins and is a different mechanism. ## Safe to apply without a control It is subtractive: the output is the input times a factor in `0..=1`. The worst failure available to it is losing part of a real sky, never gaining a region, so a blue car below the horizon that was never in the mask cannot be pulled into it. And a factor in `0..=1` cannot raise a sum, so `scene.rs`'s partition still holds when every category is refined independently — the weight taken off the flag lands in the unlisted remainder, which is where a flag belongs, ADE20K having no class for one. Every path without the evidence to judge returns the weights untouched and says which path it took. A refinement that silently did nothing is indistinguishable from the feature being off, and an empty seed set fitted to a distribution would reject every pixel. The signature is deliberately unchanged: categories are addressed by name, not by index, so a sharper mask cannot create the stale-index hazard the signature exists to guard against. The example writes `<prefix>-<category>-refined.ppm` beside the coarse one, never instead of it — whether this is an improvement is a comparative judgement and one image cannot answer it. Verified: fmt clean, clippy -D warnings clean, 57 dr-segment tests. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -24,6 +24,14 @@
|
|||||||
//! came from* rather than as an abstract grey field. PPM for the same reason
|
//! came from* rather than as an abstract grey field. PPM for the same reason
|
||||||
//! the other examples use it: no encoder dependency, and every viewer reads it.
|
//! the other examples use it: no encoder dependency, and every viewer reads it.
|
||||||
//!
|
//!
|
||||||
|
//! And `<prefix>-<category>-refined.ppm` beside it, which is the same category
|
||||||
|
//! after [`dr_segment::refine_category`] has cut it back to the pixels whose
|
||||||
|
//! colour agrees with it. Both, never one: whether that refinement is an
|
||||||
|
//! improvement is a comparative judgement — did the flag come out of the sky,
|
||||||
|
//! and is the sky still there afterwards — and a single image cannot answer
|
||||||
|
//! it. The percentage printed beside each is how much weight came off, which
|
||||||
|
//! is the number to be suspicious of when it is large.
|
||||||
|
//!
|
||||||
//! Timings are reported as a median over the requested run count, with the
|
//! Timings are reported as a median over the requested run count, with the
|
||||||
//! first run excluded. That first pass pays for tract's lazy allocation and is
|
//! first run excluded. That first pass pays for tract's lazy allocation and is
|
||||||
//! not representative of the second image a session decodes.
|
//! not representative of the second image a session decodes.
|
||||||
@@ -108,6 +116,40 @@ fn main() {
|
|||||||
.rasterise(k, width, height)
|
.rasterise(k, width, height)
|
||||||
.expect("category index came from the same Scene");
|
.expect("category index came from the same Scene");
|
||||||
write_overlay(&format!("{prefix}-{name}.ppm"), &rgb, &mask, width, height);
|
write_overlay(&format!("{prefix}-{name}.ppm"), &rgb, &mask, width, height);
|
||||||
|
|
||||||
|
// The same category cut back to the pixels whose colour agrees with
|
||||||
|
// it, written *beside* the coarse one rather than instead of it. The
|
||||||
|
// judgement this example exists to support is comparative — is the
|
||||||
|
// flag out, and is the sky still there — and it cannot be made from
|
||||||
|
// one image.
|
||||||
|
let start = Instant::now();
|
||||||
|
let (refined, what) = dr_segment::refine_category(
|
||||||
|
&mask,
|
||||||
|
&rgb,
|
||||||
|
width,
|
||||||
|
height,
|
||||||
|
scene.cell_pixels(),
|
||||||
|
&dr_segment::RefineOptions::default(),
|
||||||
|
);
|
||||||
|
let took = start.elapsed();
|
||||||
|
match what {
|
||||||
|
dr_segment::Refined::Applied { removed } => {
|
||||||
|
println!(
|
||||||
|
" refined in {took:?}: {:.1}% of the weight cut",
|
||||||
|
removed * 100.0
|
||||||
|
);
|
||||||
|
write_overlay(
|
||||||
|
&format!("{prefix}-{name}-refined.ppm"),
|
||||||
|
&rgb,
|
||||||
|
&refined,
|
||||||
|
width,
|
||||||
|
height,
|
||||||
|
);
|
||||||
|
}
|
||||||
|
dr_segment::Refined::Skipped(why) => {
|
||||||
|
println!(" left coarse: {why:?}");
|
||||||
|
}
|
||||||
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -37,10 +37,25 @@
|
|||||||
//! water — and reads a second, ADE20K-trained model to do it. It shares this
|
//! water — and reads a second, ADE20K-trained model to do it. It shares this
|
||||||
//! crate because it shares the runtime and the letterbox, not because it is
|
//! crate because it shares the runtime and the letterbox, not because it is
|
||||||
//! another way of doing the same thing.
|
//! another way of doing the same thing.
|
||||||
|
//!
|
||||||
|
//! [`refine`] is what the scene model needs and the instance model does not.
|
||||||
|
//! A category's weights come off an 80×80 grid, so one cell is twenty pixels
|
||||||
|
//! of a 1600px proxy and anything smaller than that — a flag in the sky, a
|
||||||
|
//! chimney, a bare branch — is averaged into whatever surrounds it. There is
|
||||||
|
//! no tiling answer here the way there is for an instance, because a category
|
||||||
|
//! has no bounding box to tile over. So the fix is the same one arm C makes:
|
||||||
|
//! the model says *what*, and the photograph's own pixels say *which* of them
|
||||||
|
//! belong to it.
|
||||||
|
//!
|
||||||
|
//! It shares no code with the arms and it is not a fourth one — it sharpens a
|
||||||
|
//! mask that already exists rather than proposing regions — but it is built on
|
||||||
|
//! [`distance`] for the same reason arm C is built on the watershed, which is
|
||||||
|
//! that the useful question is always "how far inside am I".
|
||||||
|
|
||||||
pub mod distance;
|
pub mod distance;
|
||||||
pub mod hierarchy;
|
pub mod hierarchy;
|
||||||
pub mod prior;
|
pub mod prior;
|
||||||
|
pub mod refine;
|
||||||
#[cfg(feature = "semantic")]
|
#[cfg(feature = "semantic")]
|
||||||
pub mod scene;
|
pub mod scene;
|
||||||
#[cfg(feature = "semantic")]
|
#[cfg(feature = "semantic")]
|
||||||
@@ -49,6 +64,7 @@ pub mod semantic;
|
|||||||
pub use distance::{signed_distance, Falloff, Morphology, Shaped};
|
pub use distance::{signed_distance, Falloff, Morphology, Shaped};
|
||||||
pub use hierarchy::{Edge, Merge, MergeTree, RegionField};
|
pub use hierarchy::{Edge, Merge, MergeTree, RegionField};
|
||||||
pub use prior::{Membership, PriorOptions};
|
pub use prior::{Membership, PriorOptions};
|
||||||
|
pub use refine::{refine_category, RefineOptions, Refined, SkipReason};
|
||||||
#[cfg(feature = "semantic")]
|
#[cfg(feature = "semantic")]
|
||||||
pub use scene::{Category, Scene, SceneModel};
|
pub use scene::{Category, Scene, SceneModel};
|
||||||
#[cfg(feature = "semantic")]
|
#[cfg(feature = "semantic")]
|
||||||
|
|||||||
@@ -0,0 +1,870 @@
|
|||||||
|
//! Sharpening a coarse category mask against the photograph's own colours.
|
||||||
|
//!
|
||||||
|
//! # The problem, as a number
|
||||||
|
//!
|
||||||
|
//! The `scene` module keeps the model's native `[1, 150, 80, 80]` logit grid, so
|
||||||
|
//! one cell is eight of the graph's input pixels across. At the 1600px proxy
|
||||||
|
//! the application segments at, the letterbox scale is `0.4` and **one cell is
|
||||||
|
//! 20 proxy pixels** — which the bilinear in `Scene::rasterise` then spreads
|
||||||
|
//! across one more either side.
|
||||||
|
//!
|
||||||
|
//! So a flag in the sky, or a chimney, or a bare branch, sits inside a handful
|
||||||
|
//! of cells whose softmax is dominated by the sky around it, and comes out
|
||||||
|
//! weighted as sky. No feather setting recovers it, because the information
|
||||||
|
//! was never in the grid.
|
||||||
|
//!
|
||||||
|
//! # What this does about it
|
||||||
|
//!
|
||||||
|
//! The photograph is at full proxy resolution even though the weights are not,
|
||||||
|
//! and it knows exactly where the flag is. The move is to let the *model*
|
||||||
|
//! decide what the category is and the *pixels* decide which of them belong to
|
||||||
|
//! it — the same division of labour [`crate::prior`] already draws between the
|
||||||
|
//! instance model and the watershed, applied to categories.
|
||||||
|
//!
|
||||||
|
//! 1. **Seeds.** Threshold the weights high, then erode by
|
||||||
|
//! [`RefineOptions::margin_cells`] logit cells using
|
||||||
|
//! [`crate::distance::signed_distance`]. What survives is confidently
|
||||||
|
//! inside; the mirror of it is confidently outside. The erosion radius is
|
||||||
|
//! derived from the grid rather than picked, because one cell *is* the
|
||||||
|
//! model's resolution and anything within a cell of the boundary is
|
||||||
|
//! precisely what there is no evidence about.
|
||||||
|
//!
|
||||||
|
//! 2. **A distribution per side, with several modes.** A single Gaussian over
|
||||||
|
//! sky is wrong: sky is blue at the zenith, white where the cloud is, and
|
||||||
|
//! pale at the horizon, and one blob over all three rejects two of them. So
|
||||||
|
//! k-means, which is GrabCut's mixture without the EM.
|
||||||
|
//!
|
||||||
|
//! 3. **Small modes are dropped** ([`RefineOptions::min_cluster`]). This is
|
||||||
|
//! the step that makes the whole thing work on the case it was built for.
|
||||||
|
//! A *small* flag deep in the sky has both a high weight and a large
|
||||||
|
//! distance from the mask's boundary, so it lands in the interior sample
|
||||||
|
//! and would teach the model its own colour. It cannot be excluded
|
||||||
|
//! geometrically — but it is a few percent of the sky, and a mode holding a
|
||||||
|
//! few percent of the confident interior is far likelier to be an intruder
|
||||||
|
//! than a real appearance of the category.
|
||||||
|
//!
|
||||||
|
//! 4. **Two tests, and a pixel must pass both.** An absolute one — is this
|
||||||
|
//! colour plausible under the category at all, as a chi-square on the
|
||||||
|
//! Mahalanobis distance — and a comparative one, is it likelier under the
|
||||||
|
//! interior than the exterior. The absolute test is what catches the flag,
|
||||||
|
//! whose colour is far from *both* sides and which the comparative test
|
||||||
|
//! alone would leave at even odds. The comparative test is what stops the
|
||||||
|
//! absolute one from needing a per-category constant.
|
||||||
|
//!
|
||||||
|
//! # Two properties that make it safe to apply blind
|
||||||
|
//!
|
||||||
|
//! **It is subtractive.** The output is the input times a factor in `0..=1`,
|
||||||
|
//! never more. So the worst failure available to it is losing part of a real
|
||||||
|
//! sky; it can never *gain* a region, and a blue car below the horizon that
|
||||||
|
//! was never in the mask cannot be pulled into it.
|
||||||
|
//!
|
||||||
|
//! **The partition survives.** `scene.rs` rests entirely on the
|
||||||
|
//! categories summing to at most one at every pixel — that is what lets two
|
||||||
|
//! adjacent grades be feathered without painting both into the overlap.
|
||||||
|
//! Multiplying by a factor in `0..=1` cannot raise a sum, so refining every
|
||||||
|
//! category independently still leaves a partition. The weight taken off the
|
||||||
|
//! flag lands in the unlisted remainder, which is exactly where a flag
|
||||||
|
//! belongs: ADE20K has no class for one.
|
||||||
|
//!
|
||||||
|
//! # What it cannot do, stated rather than discovered
|
||||||
|
//!
|
||||||
|
//! **An intruder large enough to hold its own mode is kept.** Step 3 separates
|
||||||
|
//! intruder from category by *share*, and by that measure a flag covering a
|
||||||
|
//! fifth of the sky and a bank of cloud covering a fifth of the sky are the
|
||||||
|
//! same object. Colour does not separate them either — a white cloud is as far
|
||||||
|
//! from blue sky in chrominance as many intruders are.
|
||||||
|
//!
|
||||||
|
//! So [`RefineOptions::min_cluster`] is not a threshold with a correct value
|
||||||
|
//! waiting to be found. It is the trade-off itself, and it is set where a
|
||||||
|
//! *photographic* intruder falls: a flag, a chimney or a bird is a fraction of
|
||||||
|
//! a percent of the sky it sits in, where the cloud that must survive is
|
||||||
|
//! usually tens of percent. Both ends of that are pinned by tests
|
||||||
|
//! (`a_flag_in_the_sky_is_removed` and
|
||||||
|
//! `an_intruder_larger_than_min_cluster_survives`) so that moving the number
|
||||||
|
//! reads as moving the trade-off rather than as fixing a bug.
|
||||||
|
//!
|
||||||
|
//! The case this leaves open is a *large* unrecognised object in an otherwise
|
||||||
|
//! clean category — a building filling half the sky. That wants the boundary
|
||||||
|
//! snapped to the watershed's basins, which is a different mechanism and not
|
||||||
|
//! this one.
|
||||||
|
//!
|
||||||
|
//! # And when it cannot tell
|
||||||
|
//!
|
||||||
|
//! Every path that lacks the evidence to judge returns the weights untouched
|
||||||
|
//! and says so in [`Refined::Skipped`], rather than returning a plausible
|
||||||
|
//! answer. An empty seed set fitted to a distribution would reject *every*
|
||||||
|
//! pixel, and a sky that silently vanished from a mask is the kind of failure
|
||||||
|
//! nobody attributes to the right place.
|
||||||
|
|
||||||
|
use crate::distance::signed_distance;
|
||||||
|
|
||||||
|
/// How aggressively a coarse category mask is cut back to the pixels.
|
||||||
|
#[derive(Debug, Clone, Copy, PartialEq)]
|
||||||
|
pub struct RefineOptions {
|
||||||
|
/// Weight at or above which a pixel may seed the interior distribution.
|
||||||
|
///
|
||||||
|
/// High on purpose. This is not the mask's threshold — it is the standard
|
||||||
|
/// for "the model was sure here", and everything downstream is a
|
||||||
|
/// consequence of what these pixels look like.
|
||||||
|
pub seed_weight: f32,
|
||||||
|
|
||||||
|
/// How far inside the boundary a seed must lie, in logit cells.
|
||||||
|
///
|
||||||
|
/// One cell is the model's own resolution and `rasterise`'s bilinear
|
||||||
|
/// spreads a cell's influence across one more, so the first cell and a
|
||||||
|
/// half either side of the boundary is smear rather than evidence.
|
||||||
|
pub margin_cells: f32,
|
||||||
|
|
||||||
|
/// Modes fitted per side.
|
||||||
|
///
|
||||||
|
/// Four covers blue, cloud and horizon haze with one spare. More would fit
|
||||||
|
/// the intruder its own mode reliably enough to keep it, which is the
|
||||||
|
/// opposite of the point.
|
||||||
|
pub clusters: usize,
|
||||||
|
|
||||||
|
/// Share of a side's samples below which a mode is discarded as an
|
||||||
|
/// intruder rather than kept as part of the category.
|
||||||
|
///
|
||||||
|
/// The one number here with a real trade-off in it: too high and a genuine
|
||||||
|
/// wisp of cloud is cut out of the sky, too low and a small flag teaches
|
||||||
|
/// the model to keep it.
|
||||||
|
pub min_cluster: f32,
|
||||||
|
|
||||||
|
/// How much luminance counts against chrominance in judging a colour.
|
||||||
|
///
|
||||||
|
/// Well under one, and that is the difference between this working and not
|
||||||
|
/// on sky. Sky's variance is *dominated* by luminance — the zenith is
|
||||||
|
/// several stops off the horizon and a cloud is brighter than either — so
|
||||||
|
/// at equal weight the distribution is a long bright streak that a
|
||||||
|
/// mid-grey flag sits comfortably inside. A flag is separated by
|
||||||
|
/// chrominance; a cloud is separated by luminance alone. Down-weighting
|
||||||
|
/// keeps the cloud and catches the flag.
|
||||||
|
///
|
||||||
|
/// Not zero, or a dark bird against a bright sky is kept.
|
||||||
|
///
|
||||||
|
/// The same reasoning `dr_gpu::SegmentOptions` applies to the watershed's
|
||||||
|
/// gradient, for the same reason.
|
||||||
|
pub luma_weight: f32,
|
||||||
|
|
||||||
|
/// Squared Mahalanobis distance beyond which a colour is implausible under
|
||||||
|
/// the category.
|
||||||
|
///
|
||||||
|
/// Chi-square with three degrees of freedom: `11.34` is the 99th
|
||||||
|
/// percentile, so under the fitted model one confident pixel in a hundred
|
||||||
|
/// is expected to fail this on its own.
|
||||||
|
pub outlier: f32,
|
||||||
|
|
||||||
|
/// Width of the band, in nats, over which the gate falls from keep to cut.
|
||||||
|
///
|
||||||
|
/// A hard cut would put an aliased edge through the photograph at exactly
|
||||||
|
/// the place a person is looking.
|
||||||
|
pub soften: f32,
|
||||||
|
|
||||||
|
/// Radius, in pixels, the gate is blurred by before it is applied.
|
||||||
|
///
|
||||||
|
/// A per-pixel colour test on a real sky speckles: sensor noise puts
|
||||||
|
/// individual pixels over the line in both directions. Two pixels of blur
|
||||||
|
/// removes the speckle and costs nothing at the boundary, which is already
|
||||||
|
/// soft by [`RefineOptions::soften`].
|
||||||
|
pub smooth_px: f32,
|
||||||
|
|
||||||
|
/// Samples below which a side is too small to fit anything to.
|
||||||
|
pub min_samples: usize,
|
||||||
|
}
|
||||||
|
|
||||||
|
impl Default for RefineOptions {
|
||||||
|
fn default() -> Self {
|
||||||
|
Self {
|
||||||
|
seed_weight: 0.9,
|
||||||
|
margin_cells: 1.5,
|
||||||
|
clusters: 4,
|
||||||
|
min_cluster: 0.03,
|
||||||
|
luma_weight: 0.25,
|
||||||
|
outlier: 11.34,
|
||||||
|
soften: 1.5,
|
||||||
|
smooth_px: 2.0,
|
||||||
|
min_samples: 500,
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Why a refinement declined to do anything.
|
||||||
|
///
|
||||||
|
/// Carried out rather than logged in here, because the caller knows which
|
||||||
|
/// category it asked about and this module does not.
|
||||||
|
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
|
||||||
|
pub enum SkipReason {
|
||||||
|
/// Nothing survived the erosion — the category is present but everywhere
|
||||||
|
/// thinner than the model's own resolution, so there is no pixel it is
|
||||||
|
/// sure about.
|
||||||
|
NoInterior,
|
||||||
|
/// The frame is essentially all this category, leaving nothing to
|
||||||
|
/// contrast it against.
|
||||||
|
NoExterior,
|
||||||
|
/// The buffers handed in do not describe one image.
|
||||||
|
Mismatched,
|
||||||
|
}
|
||||||
|
|
||||||
|
/// What a refinement did.
|
||||||
|
#[derive(Debug, Clone, Copy, PartialEq)]
|
||||||
|
pub enum Refined {
|
||||||
|
/// Applied. `removed` is the fraction of the category's original total
|
||||||
|
/// weight the gate took away — small for a clean sky, a few percent for
|
||||||
|
/// one with a flag in it, and large enough to be worth a look in the log
|
||||||
|
/// if the colour model has gone wrong.
|
||||||
|
Applied {
|
||||||
|
removed: f32,
|
||||||
|
},
|
||||||
|
Skipped(SkipReason),
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Number of features a pixel is judged on: one luminance, two chrominance.
|
||||||
|
const FEATURES: usize = 3;
|
||||||
|
|
||||||
|
/// Floor on a fitted variance.
|
||||||
|
///
|
||||||
|
/// A mode over a flat patch of sky has a variance of nearly nothing, and
|
||||||
|
/// nothing in the denominator makes every other pixel infinitely improbable.
|
||||||
|
/// This is a standard deviation of one part in a hundred, comfortably below
|
||||||
|
/// anything a photograph resolves and comfortably above zero.
|
||||||
|
const VARIANCE_FLOOR: f32 = 1e-4;
|
||||||
|
|
||||||
|
/// Cap on the samples fitted per side.
|
||||||
|
///
|
||||||
|
/// k-means over a million pixels answers the same as k-means over twenty
|
||||||
|
/// thousand of them, and costs fifty times as much. The subsample is a fixed
|
||||||
|
/// stride rather than a random draw so the result is reproducible.
|
||||||
|
const MAX_SAMPLES: usize = 20_000;
|
||||||
|
|
||||||
|
/// Lloyd iterations.
|
||||||
|
///
|
||||||
|
/// Fixed rather than run to convergence: a stopping rule that depends on
|
||||||
|
/// floating-point comparison is a stopping rule that can differ between
|
||||||
|
/// machines, and a mask that differs between machines reaches the sidecar as
|
||||||
|
/// indices meaning one thing on the desktop and another on the phone
|
||||||
|
/// (docs/segmentation.md §6).
|
||||||
|
const ITERATIONS: usize = 12;
|
||||||
|
|
||||||
|
/// TRACES: FR-DEV-3
|
||||||
|
/// Cut a coarse category mask back to the pixels that belong to it.
|
||||||
|
///
|
||||||
|
/// `weights` is one category's coverage at `width * height`, as
|
||||||
|
/// `Scene::rasterise` produces it. `rgb` is the same picture, tightly packed
|
||||||
|
/// `f32` RGB — the proxy the model itself read, so the two describe one frame.
|
||||||
|
///
|
||||||
|
/// `cell_pixels` is how many pixels of *this* buffer one logit cell spans; see
|
||||||
|
/// `Scene::cell_pixels`, which is where the caller should get it rather than
|
||||||
|
/// re-deriving a letterbox inverse.
|
||||||
|
///
|
||||||
|
/// Returns the refined weights and what was done to them. On any
|
||||||
|
/// [`Refined::Skipped`] the weights come back identical to the input.
|
||||||
|
pub fn refine_category(
|
||||||
|
weights: &[f32],
|
||||||
|
rgb: &[f32],
|
||||||
|
width: usize,
|
||||||
|
height: usize,
|
||||||
|
cell_pixels: f32,
|
||||||
|
options: &RefineOptions,
|
||||||
|
) -> (Vec<f32>, Refined) {
|
||||||
|
let pixels = width * height;
|
||||||
|
if weights.len() != pixels || rgb.len() != pixels * 3 || pixels == 0 {
|
||||||
|
return (weights.to_vec(), Refined::Skipped(SkipReason::Mismatched));
|
||||||
|
}
|
||||||
|
|
||||||
|
// The distance field is over the *thresholded* weights, so the boundary it
|
||||||
|
// measures from is where the model stopped being sure — not where the mask
|
||||||
|
// will eventually be cut, which is what this function is deciding.
|
||||||
|
let coverage: Vec<u8> = weights
|
||||||
|
.iter()
|
||||||
|
.map(|&w| (w.clamp(0.0, 1.0) * 255.0).round() as u8)
|
||||||
|
.collect();
|
||||||
|
let threshold = (options.seed_weight.clamp(0.0, 1.0) * 255.0).round() as u8;
|
||||||
|
let distance = signed_distance(&coverage, width, height, threshold);
|
||||||
|
|
||||||
|
let margin = options.margin_cells * cell_pixels.max(1.0);
|
||||||
|
|
||||||
|
let interior = sample(&distance, rgb, options.luma_weight, |d| d >= margin);
|
||||||
|
if interior.len() < options.min_samples {
|
||||||
|
return (weights.to_vec(), Refined::Skipped(SkipReason::NoInterior));
|
||||||
|
}
|
||||||
|
let exterior = sample(&distance, rgb, options.luma_weight, |d| d <= -margin);
|
||||||
|
if exterior.len() < options.min_samples {
|
||||||
|
return (weights.to_vec(), Refined::Skipped(SkipReason::NoExterior));
|
||||||
|
}
|
||||||
|
|
||||||
|
let inside = fit(&interior, options.clusters, options.min_cluster);
|
||||||
|
// The exterior keeps every mode it finds. Pruning there would be the wrong
|
||||||
|
// sign: a small mode outside is a small *object*, and forgetting it only
|
||||||
|
// makes the comparative test more willing to keep its pixels.
|
||||||
|
let outside = fit(&exterior, options.clusters, 0.0);
|
||||||
|
|
||||||
|
let mut gate = vec![0.0f32; pixels];
|
||||||
|
for (p, cell) in gate.iter_mut().enumerate() {
|
||||||
|
let x = feature(rgb, p, options.luma_weight);
|
||||||
|
|
||||||
|
// Squared Mahalanobis for the absolute test, because that is what is
|
||||||
|
// chi-square distributed; the density — the same distance with the
|
||||||
|
// mode's own spread folded in — for the comparative one, so that a
|
||||||
|
// tight mode and a loose one are compared fairly.
|
||||||
|
let (maha_in, nll_in) = nearest(&inside, &x);
|
||||||
|
let (_, nll_out) = nearest(&outside, &x);
|
||||||
|
|
||||||
|
let plausible = 0.5 * (options.outlier - maha_in);
|
||||||
|
let likelier = nll_out - nll_in;
|
||||||
|
let verdict = plausible.min(likelier);
|
||||||
|
|
||||||
|
*cell = smoothstep(-options.soften, options.soften, verdict);
|
||||||
|
}
|
||||||
|
|
||||||
|
if options.smooth_px >= 1.0 {
|
||||||
|
gate = blur(&gate, width, height, options.smooth_px.round() as usize);
|
||||||
|
}
|
||||||
|
|
||||||
|
let before: f32 = weights.iter().sum();
|
||||||
|
let refined: Vec<f32> = weights.iter().zip(&gate).map(|(&w, &g)| w * g).collect();
|
||||||
|
let after: f32 = refined.iter().sum();
|
||||||
|
let removed = if before > 0.0 {
|
||||||
|
((before - after) / before).clamp(0.0, 1.0)
|
||||||
|
} else {
|
||||||
|
0.0
|
||||||
|
};
|
||||||
|
|
||||||
|
(refined, Refined::Applied { removed })
|
||||||
|
}
|
||||||
|
|
||||||
|
/// One pixel's colour, as the three numbers the distributions are fitted over.
|
||||||
|
///
|
||||||
|
/// Opponent axes rather than raw RGB, and no division anywhere: normalising
|
||||||
|
/// chrominance by luminance would be exposure-invariant and would also blow up
|
||||||
|
/// in the shadows, which on a landscape is the whole foreground.
|
||||||
|
fn feature(rgb: &[f32], p: usize, luma_weight: f32) -> [f32; FEATURES] {
|
||||||
|
let (r, g, b) = (rgb[p * 3], rgb[p * 3 + 1], rgb[p * 3 + 2]);
|
||||||
|
let luma = 0.2126 * r + 0.7152 * g + 0.0722 * b;
|
||||||
|
[luma_weight * luma, r - g, b - 0.5 * (r + g)]
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Every pixel whose distance passes `want`, subsampled to [`MAX_SAMPLES`].
|
||||||
|
///
|
||||||
|
/// Two passes — count, then take a fixed stride — rather than one pass
|
||||||
|
/// collecting everything and thinning afterwards, so a sky filling a 22 MP
|
||||||
|
/// proxy never allocates twenty million features to throw most of them away.
|
||||||
|
fn sample(
|
||||||
|
distance: &[f32],
|
||||||
|
rgb: &[f32],
|
||||||
|
luma_weight: f32,
|
||||||
|
want: impl Fn(f32) -> bool,
|
||||||
|
) -> Vec<[f32; FEATURES]> {
|
||||||
|
let total = distance.iter().filter(|&&d| want(d)).count();
|
||||||
|
if total == 0 {
|
||||||
|
return Vec::new();
|
||||||
|
}
|
||||||
|
let stride = total.div_ceil(MAX_SAMPLES).max(1);
|
||||||
|
|
||||||
|
let mut out = Vec::with_capacity(total.div_ceil(stride));
|
||||||
|
let mut seen = 0usize;
|
||||||
|
for (p, &d) in distance.iter().enumerate() {
|
||||||
|
if !want(d) {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
if seen.is_multiple_of(stride) {
|
||||||
|
out.push(feature(rgb, p, luma_weight));
|
||||||
|
}
|
||||||
|
seen += 1;
|
||||||
|
}
|
||||||
|
out
|
||||||
|
}
|
||||||
|
|
||||||
|
/// One mode of a side's colour distribution.
|
||||||
|
#[derive(Debug, Clone, Copy)]
|
||||||
|
struct Mode {
|
||||||
|
mean: [f32; FEATURES],
|
||||||
|
/// Per-feature, not a full covariance. The opponent axes are close enough
|
||||||
|
/// to decorrelated by construction that the off-diagonal terms buy little,
|
||||||
|
/// and a diagonal fit stays well-conditioned on the few hundred samples a
|
||||||
|
/// small mode gets.
|
||||||
|
variance: [f32; FEATURES],
|
||||||
|
/// `0.5 * ln|Σ|`, precomputed because it is constant per mode and the
|
||||||
|
/// inner loop runs once per pixel.
|
||||||
|
half_log_det: f32,
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Fit `k` modes to one side, dropping any that hold less than `min_share`.
|
||||||
|
fn fit(samples: &[[f32; FEATURES]], k: usize, min_share: f32) -> Vec<Mode> {
|
||||||
|
let k = k.max(1).min(samples.len());
|
||||||
|
let mut centres = seed_centres(samples, k);
|
||||||
|
let k = centres.len();
|
||||||
|
|
||||||
|
let mut assignment = vec![0usize; samples.len()];
|
||||||
|
for _ in 0..ITERATIONS {
|
||||||
|
let mut moved = false;
|
||||||
|
for (s, sample) in samples.iter().enumerate() {
|
||||||
|
let mut best = 0usize;
|
||||||
|
let mut best_d = f32::INFINITY;
|
||||||
|
for (c, centre) in centres.iter().enumerate() {
|
||||||
|
let d = squared(sample, centre);
|
||||||
|
if d < best_d {
|
||||||
|
best_d = d;
|
||||||
|
best = c;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if assignment[s] != best {
|
||||||
|
assignment[s] = best;
|
||||||
|
moved = true;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if !moved {
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
|
||||||
|
let mut sums = vec![[0.0f32; FEATURES]; k];
|
||||||
|
let mut counts = vec![0usize; k];
|
||||||
|
for (s, sample) in samples.iter().enumerate() {
|
||||||
|
let c = assignment[s];
|
||||||
|
for f in 0..FEATURES {
|
||||||
|
sums[c][f] += sample[f];
|
||||||
|
}
|
||||||
|
counts[c] += 1;
|
||||||
|
}
|
||||||
|
for c in 0..k {
|
||||||
|
// An emptied centre is left where it was rather than re-seeded.
|
||||||
|
// Re-seeding is the usual advice, but it needs a random draw
|
||||||
|
// inside the loop and the determinism is worth more here than the
|
||||||
|
// empty cluster costs.
|
||||||
|
if counts[c] == 0 {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
for f in 0..FEATURES {
|
||||||
|
centres[c][f] = sums[c][f] / counts[c] as f32;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Mean and spread from the final assignment.
|
||||||
|
let mut sums = vec![[0.0f32; FEATURES]; k];
|
||||||
|
let mut squares = vec![[0.0f32; FEATURES]; k];
|
||||||
|
let mut counts = vec![0usize; k];
|
||||||
|
for (s, sample) in samples.iter().enumerate() {
|
||||||
|
let c = assignment[s];
|
||||||
|
for f in 0..FEATURES {
|
||||||
|
sums[c][f] += sample[f];
|
||||||
|
squares[c][f] += sample[f] * sample[f];
|
||||||
|
}
|
||||||
|
counts[c] += 1;
|
||||||
|
}
|
||||||
|
|
||||||
|
let floor = (min_share * samples.len() as f32) as usize;
|
||||||
|
let mut modes = Vec::new();
|
||||||
|
let mut largest: Option<(usize, Mode)> = None;
|
||||||
|
for c in 0..k {
|
||||||
|
if counts[c] == 0 {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
let n = counts[c] as f32;
|
||||||
|
let mut mean = [0.0f32; FEATURES];
|
||||||
|
let mut variance = [0.0f32; FEATURES];
|
||||||
|
let mut half_log_det = 0.0f32;
|
||||||
|
for f in 0..FEATURES {
|
||||||
|
mean[f] = sums[c][f] / n;
|
||||||
|
variance[f] = (squares[c][f] / n - mean[f] * mean[f]).max(VARIANCE_FLOOR);
|
||||||
|
half_log_det += 0.5 * variance[f].ln();
|
||||||
|
}
|
||||||
|
let mode = Mode {
|
||||||
|
mean,
|
||||||
|
variance,
|
||||||
|
half_log_det,
|
||||||
|
};
|
||||||
|
|
||||||
|
if largest.is_none_or(|(best, _)| counts[c] > best) {
|
||||||
|
largest = Some((counts[c], mode));
|
||||||
|
}
|
||||||
|
if counts[c] >= floor {
|
||||||
|
modes.push(mode);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Pruning everything is possible when `min_cluster` is set high and the
|
||||||
|
// samples split evenly. Keeping the largest is better than returning
|
||||||
|
// nothing, which would make every pixel infinitely improbable and cut the
|
||||||
|
// whole category away.
|
||||||
|
if modes.is_empty() {
|
||||||
|
modes.extend(largest.map(|(_, m)| m));
|
||||||
|
}
|
||||||
|
modes
|
||||||
|
}
|
||||||
|
|
||||||
|
/// k-means++ seeding, with a fixed sequence.
|
||||||
|
///
|
||||||
|
/// The usual seeding is random, and random would make a mask depend on which
|
||||||
|
/// run produced it. The draw is a fixed-seed LCG instead: still spread out,
|
||||||
|
/// still far better than taking the first `k` samples, and the same every time
|
||||||
|
/// on every machine.
|
||||||
|
fn seed_centres(samples: &[[f32; FEATURES]], k: usize) -> Vec<[f32; FEATURES]> {
|
||||||
|
let mut rng = Lcg(0x2545_F491_4F6C_DD1D);
|
||||||
|
let mut centres = vec![samples[0]];
|
||||||
|
let mut nearest = vec![f32::INFINITY; samples.len()];
|
||||||
|
|
||||||
|
while centres.len() < k {
|
||||||
|
let last = *centres.last().expect("seeded with one");
|
||||||
|
let mut total = 0.0f64;
|
||||||
|
for (s, sample) in samples.iter().enumerate() {
|
||||||
|
nearest[s] = nearest[s].min(squared(sample, &last));
|
||||||
|
total += nearest[s] as f64;
|
||||||
|
}
|
||||||
|
if total <= 0.0 {
|
||||||
|
// Every sample coincides with a centre — a perfectly flat side.
|
||||||
|
// More modes would all be the same mode.
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
|
||||||
|
let mut target = rng.unit() as f64 * total;
|
||||||
|
let mut pick = samples.len() - 1;
|
||||||
|
for (s, &d) in nearest.iter().enumerate() {
|
||||||
|
target -= d as f64;
|
||||||
|
if target <= 0.0 {
|
||||||
|
pick = s;
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
centres.push(samples[pick]);
|
||||||
|
}
|
||||||
|
centres
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The closest mode, as `(squared Mahalanobis, negative log density)`.
|
||||||
|
///
|
||||||
|
/// Nearest rather than a proper mixture sum: the largest term dominates a
|
||||||
|
/// well-separated mixture, and taking the max is what makes the two numbers
|
||||||
|
/// this returns describe *the same* mode.
|
||||||
|
fn nearest(modes: &[Mode], x: &[f32; FEATURES]) -> (f32, f32) {
|
||||||
|
let mut best = (f32::INFINITY, f32::INFINITY);
|
||||||
|
for mode in modes {
|
||||||
|
let mut maha = 0.0f32;
|
||||||
|
for ((v, mean), variance) in x.iter().zip(&mode.mean).zip(&mode.variance) {
|
||||||
|
let d = v - mean;
|
||||||
|
maha += d * d / variance;
|
||||||
|
}
|
||||||
|
let nll = 0.5 * maha + mode.half_log_det;
|
||||||
|
if nll < best.1 {
|
||||||
|
best = (maha, nll);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
best
|
||||||
|
}
|
||||||
|
|
||||||
|
fn squared(a: &[f32; FEATURES], b: &[f32; FEATURES]) -> f32 {
|
||||||
|
(0..FEATURES).map(|f| (a[f] - b[f]).powi(2)).sum()
|
||||||
|
}
|
||||||
|
|
||||||
|
fn smoothstep(edge0: f32, edge1: f32, x: f32) -> f32 {
|
||||||
|
if edge1 <= edge0 {
|
||||||
|
return if x >= edge1 { 1.0 } else { 0.0 };
|
||||||
|
}
|
||||||
|
let t = ((x - edge0) / (edge1 - edge0)).clamp(0.0, 1.0);
|
||||||
|
t * t * (3.0 - 2.0 * t)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Separable box blur, clamped at the edges.
|
||||||
|
///
|
||||||
|
/// A box rather than a gaussian because it is being applied to a gate that is
|
||||||
|
/// already soft: what is wanted is the removal of single-pixel speckle, and the
|
||||||
|
/// shape of the kernel that does it does not matter.
|
||||||
|
fn blur(src: &[f32], width: usize, height: usize, radius: usize) -> Vec<f32> {
|
||||||
|
if radius == 0 || width == 0 || height == 0 {
|
||||||
|
return src.to_vec();
|
||||||
|
}
|
||||||
|
let span = (2 * radius + 1) as f32;
|
||||||
|
let mut mid = vec![0.0f32; src.len()];
|
||||||
|
for y in 0..height {
|
||||||
|
for x in 0..width {
|
||||||
|
let mut sum = 0.0f32;
|
||||||
|
for k in 0..=2 * radius {
|
||||||
|
let sx = (x + k).saturating_sub(radius).min(width - 1);
|
||||||
|
sum += src[y * width + sx];
|
||||||
|
}
|
||||||
|
mid[y * width + x] = sum / span;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
let mut out = vec![0.0f32; src.len()];
|
||||||
|
for y in 0..height {
|
||||||
|
for x in 0..width {
|
||||||
|
let mut sum = 0.0f32;
|
||||||
|
for k in 0..=2 * radius {
|
||||||
|
let sy = (y + k).saturating_sub(radius).min(height - 1);
|
||||||
|
sum += mid[sy * width + x];
|
||||||
|
}
|
||||||
|
out[y * width + x] = sum / span;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
out
|
||||||
|
}
|
||||||
|
|
||||||
|
/// A fixed-sequence generator, for the one place a draw is needed.
|
||||||
|
struct Lcg(u64);
|
||||||
|
|
||||||
|
impl Lcg {
|
||||||
|
fn unit(&mut self) -> f32 {
|
||||||
|
self.0 = self
|
||||||
|
.0
|
||||||
|
.wrapping_mul(6364136223846793005)
|
||||||
|
.wrapping_add(1442695040888963407);
|
||||||
|
// The high bits are the well-mixed ones in an LCG; the low bits cycle
|
||||||
|
// with a short period and would make the draw far from uniform.
|
||||||
|
((self.0 >> 40) as f32) / ((1u64 << 24) as f32)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#[cfg(test)]
|
||||||
|
mod tests {
|
||||||
|
use super::*;
|
||||||
|
|
||||||
|
/// One logit cell, in the pixels of the synthetic frames below.
|
||||||
|
const CELL: f32 = 8.0;
|
||||||
|
|
||||||
|
/// Edge of every synthetic frame.
|
||||||
|
///
|
||||||
|
/// Large enough that an intruder can be a realistic *fraction* of the
|
||||||
|
/// category rather than a realistic number of pixels — which is what
|
||||||
|
/// [`RefineOptions::min_cluster`] is measured in, and getting that
|
||||||
|
/// proportion wrong is the difference between this method working and not.
|
||||||
|
const EDGE: usize = 192;
|
||||||
|
|
||||||
|
/// A frame with a blue upper half and a green lower half, and a category
|
||||||
|
/// mask that claims the whole upper half plus a smear over the boundary —
|
||||||
|
/// which is what the coarse grid actually produces.
|
||||||
|
fn landscape(w: usize, h: usize) -> (Vec<f32>, Vec<f32>) {
|
||||||
|
let mut rgb = vec![0.0f32; w * h * 3];
|
||||||
|
let mut weights = vec![0.0f32; w * h];
|
||||||
|
for y in 0..h {
|
||||||
|
for x in 0..w {
|
||||||
|
let p = y * w + x;
|
||||||
|
// A vertical gradient in the sky, because a real one has one
|
||||||
|
// and a single Gaussian over luma is exactly what it breaks.
|
||||||
|
let t = y as f32 / h as f32;
|
||||||
|
if y < h / 2 {
|
||||||
|
rgb[p * 3] = 0.25 + 0.4 * t;
|
||||||
|
rgb[p * 3 + 1] = 0.45 + 0.35 * t;
|
||||||
|
rgb[p * 3 + 2] = 0.85;
|
||||||
|
} else {
|
||||||
|
rgb[p * 3] = 0.20;
|
||||||
|
rgb[p * 3 + 1] = 0.45;
|
||||||
|
rgb[p * 3 + 2] = 0.15;
|
||||||
|
}
|
||||||
|
weights[p] = if y < h / 2 { 1.0 } else { 0.0 };
|
||||||
|
}
|
||||||
|
}
|
||||||
|
(rgb, weights)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Paint a rectangle into the frame, and let the coarse mask claim it —
|
||||||
|
/// the flag in the sky.
|
||||||
|
fn intrude(
|
||||||
|
rgb: &mut [f32],
|
||||||
|
weights: &mut [f32],
|
||||||
|
w: usize,
|
||||||
|
rect: (usize, usize, usize, usize),
|
||||||
|
colour: [f32; 3],
|
||||||
|
) {
|
||||||
|
let (x0, y0, x1, y1) = rect;
|
||||||
|
for y in y0..y1 {
|
||||||
|
for x in x0..x1 {
|
||||||
|
let p = y * w + x;
|
||||||
|
rgb[p * 3..p * 3 + 3].copy_from_slice(&colour);
|
||||||
|
weights[p] = 1.0;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
fn mean(v: &[f32], w: usize, rect: (usize, usize, usize, usize)) -> f32 {
|
||||||
|
let (x0, y0, x1, y1) = rect;
|
||||||
|
let mut sum = 0.0;
|
||||||
|
for y in y0..y1 {
|
||||||
|
for x in x0..x1 {
|
||||||
|
sum += v[y * w + x];
|
||||||
|
}
|
||||||
|
}
|
||||||
|
sum / ((x1 - x0) * (y1 - y0)) as f32
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The case the module exists for: a red flag inside the sky, claimed by
|
||||||
|
/// the coarse mask, must come back out — while the sky around it stays.
|
||||||
|
///
|
||||||
|
/// The flag is 16×16 against an eroded sky of ~16k pixels, so it is 1.6%
|
||||||
|
/// of the category. That proportion is the test, as much as the colour is:
|
||||||
|
/// see [`an_intruder_larger_than_min_cluster_survives`] for the same
|
||||||
|
/// picture with a bigger flag and the opposite result.
|
||||||
|
#[test]
|
||||||
|
fn a_flag_in_the_sky_is_removed() {
|
||||||
|
let (w, h) = (EDGE, EDGE);
|
||||||
|
let (mut rgb, mut weights) = landscape(w, h);
|
||||||
|
intrude(
|
||||||
|
&mut rgb,
|
||||||
|
&mut weights,
|
||||||
|
w,
|
||||||
|
(88, 40, 104, 56),
|
||||||
|
[0.75, 0.10, 0.12],
|
||||||
|
);
|
||||||
|
|
||||||
|
let (out, what) = refine_category(&weights, &rgb, w, h, CELL, &RefineOptions::default());
|
||||||
|
assert!(
|
||||||
|
matches!(what, Refined::Applied { .. }),
|
||||||
|
"should have had the evidence to judge: {what:?}"
|
||||||
|
);
|
||||||
|
|
||||||
|
// Inside the flag, clear of its own blurred edge.
|
||||||
|
let inside = mean(&out, w, (92, 44, 100, 52));
|
||||||
|
assert!(inside < 0.2, "the flag should be cut out, got {inside}");
|
||||||
|
|
||||||
|
// Sky far from the flag and far from the horizon.
|
||||||
|
let sky = mean(&out, w, (8, 8, 60, 60));
|
||||||
|
assert!(sky > 0.8, "the sky around it should survive, got {sky}");
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The method's stated limit, pinned so that it is a decision rather than
|
||||||
|
/// a surprise.
|
||||||
|
///
|
||||||
|
/// An intruder big enough to hold its own k-means mode above
|
||||||
|
/// [`RefineOptions::min_cluster`] is indistinguishable, by this test, from
|
||||||
|
/// a legitimate second appearance of the category — a bank of cloud is
|
||||||
|
/// exactly that, and the two differ only in what a person calls them. So a
|
||||||
|
/// 22% flag is kept, and the knob that would cut it is the same knob that
|
||||||
|
/// would cut the cloud in
|
||||||
|
/// [`a_cloud_is_not_mistaken_for_an_intruder`].
|
||||||
|
///
|
||||||
|
/// This is not a defect to be fixed by raising `min_cluster`; it is the
|
||||||
|
/// trade-off that setting *is*. Recorded here so that a later change which
|
||||||
|
/// makes this test fail is recognised as having moved the trade-off rather
|
||||||
|
/// than as having fixed a bug.
|
||||||
|
#[test]
|
||||||
|
fn an_intruder_larger_than_min_cluster_survives() {
|
||||||
|
let (w, h) = (EDGE, EDGE);
|
||||||
|
let (mut rgb, mut weights) = landscape(w, h);
|
||||||
|
intrude(
|
||||||
|
&mut rgb,
|
||||||
|
&mut weights,
|
||||||
|
w,
|
||||||
|
(60, 20, 120, 80),
|
||||||
|
[0.75, 0.10, 0.12],
|
||||||
|
);
|
||||||
|
|
||||||
|
let (out, _) = refine_category(&weights, &rgb, w, h, CELL, &RefineOptions::default());
|
||||||
|
let inside = mean(&out, w, (75, 35, 105, 65));
|
||||||
|
assert!(
|
||||||
|
inside > 0.7,
|
||||||
|
"a 22% intruder is kept — see this test's comment: {inside}"
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The refinement may only ever take weight away.
|
||||||
|
///
|
||||||
|
/// This is what bounds its damage and what keeps `scene.rs`'s partition
|
||||||
|
/// true when several categories are refined independently, so it is
|
||||||
|
/// asserted rather than left as a property of the arithmetic.
|
||||||
|
#[test]
|
||||||
|
fn refinement_is_subtractive() {
|
||||||
|
let (w, h) = (EDGE, EDGE);
|
||||||
|
let (mut rgb, mut weights) = landscape(w, h);
|
||||||
|
intrude(
|
||||||
|
&mut rgb,
|
||||||
|
&mut weights,
|
||||||
|
w,
|
||||||
|
(88, 40, 104, 56),
|
||||||
|
[0.75, 0.1, 0.12],
|
||||||
|
);
|
||||||
|
|
||||||
|
let (out, _) = refine_category(&weights, &rgb, w, h, CELL, &RefineOptions::default());
|
||||||
|
for (p, (&before, &after)) in weights.iter().zip(&out).enumerate() {
|
||||||
|
assert!(
|
||||||
|
after <= before + 1e-6,
|
||||||
|
"pixel {p} gained weight: {before} -> {after}"
|
||||||
|
);
|
||||||
|
assert!(
|
||||||
|
(0.0..=1.0).contains(&after),
|
||||||
|
"pixel {p} out of range: {after}"
|
||||||
|
);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// A cloud is a legitimate part of the sky and is separated from the blue
|
||||||
|
/// by luminance alone, which is exactly what `luma_weight` is for.
|
||||||
|
///
|
||||||
|
/// Without the down-weighting this test fails and the flag test passes,
|
||||||
|
/// which is why both are here.
|
||||||
|
#[test]
|
||||||
|
fn a_cloud_is_not_mistaken_for_an_intruder() {
|
||||||
|
let (w, h) = (EDGE, EDGE);
|
||||||
|
let (mut rgb, mut weights) = landscape(w, h);
|
||||||
|
// Big and pale — near-neutral, much brighter than the blue.
|
||||||
|
intrude(
|
||||||
|
&mut rgb,
|
||||||
|
&mut weights,
|
||||||
|
w,
|
||||||
|
(30, 12, 150, 60),
|
||||||
|
[0.92, 0.94, 0.96],
|
||||||
|
);
|
||||||
|
|
||||||
|
let (out, _) = refine_category(&weights, &rgb, w, h, CELL, &RefineOptions::default());
|
||||||
|
let cloud = mean(&out, w, (50, 24, 130, 48));
|
||||||
|
assert!(cloud > 0.7, "the cloud should stay sky, got {cloud}");
|
||||||
|
}
|
||||||
|
|
||||||
|
/// No confident interior means no distribution, and no distribution must
|
||||||
|
/// mean "leave it alone" rather than "reject everything".
|
||||||
|
#[test]
|
||||||
|
fn a_mask_thinner_than_the_grid_is_left_alone() {
|
||||||
|
let (w, h) = (EDGE, EDGE);
|
||||||
|
let (rgb, _) = landscape(w, h);
|
||||||
|
// A three-pixel stripe: narrower than one cell, so the erosion at
|
||||||
|
// 1.5 cells empties it.
|
||||||
|
let mut weights = vec![0.0f32; w * h];
|
||||||
|
for y in 0..h {
|
||||||
|
for x in 60..63 {
|
||||||
|
weights[y * w + x] = 1.0;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
let (out, what) = refine_category(&weights, &rgb, w, h, CELL, &RefineOptions::default());
|
||||||
|
assert_eq!(what, Refined::Skipped(SkipReason::NoInterior));
|
||||||
|
assert_eq!(out, weights, "a skip must return the input untouched");
|
||||||
|
}
|
||||||
|
|
||||||
|
/// A frame that is entirely one category has nothing to contrast against.
|
||||||
|
#[test]
|
||||||
|
fn a_frame_of_nothing_but_sky_is_left_alone() {
|
||||||
|
let (w, h) = (EDGE, EDGE);
|
||||||
|
let rgb = vec![0.6f32; w * h * 3];
|
||||||
|
let weights = vec![1.0f32; w * h];
|
||||||
|
|
||||||
|
let (out, what) = refine_category(&weights, &rgb, w, h, CELL, &RefineOptions::default());
|
||||||
|
assert_eq!(what, Refined::Skipped(SkipReason::NoExterior));
|
||||||
|
assert_eq!(out, weights);
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Masks reach the sidecar as indices, so the same input must give the
|
||||||
|
/// same mask on every run and every machine — which is why the k-means
|
||||||
|
/// seeding uses a fixed sequence and the iteration count is fixed.
|
||||||
|
#[test]
|
||||||
|
fn the_same_input_gives_the_same_answer() {
|
||||||
|
let (w, h) = (EDGE, EDGE);
|
||||||
|
let (mut rgb, mut weights) = landscape(w, h);
|
||||||
|
intrude(
|
||||||
|
&mut rgb,
|
||||||
|
&mut weights,
|
||||||
|
w,
|
||||||
|
(88, 40, 104, 56),
|
||||||
|
[0.8, 0.2, 0.2],
|
||||||
|
);
|
||||||
|
|
||||||
|
let opts = RefineOptions::default();
|
||||||
|
let (a, _) = refine_category(&weights, &rgb, w, h, CELL, &opts);
|
||||||
|
let (b, _) = refine_category(&weights, &rgb, w, h, CELL, &opts);
|
||||||
|
assert_eq!(a, b);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn buffers_that_disagree_are_refused() {
|
||||||
|
let (out, what) =
|
||||||
|
refine_category(&[1.0; 4], &[0.5; 6], 2, 2, CELL, &RefineOptions::default());
|
||||||
|
assert_eq!(what, Refined::Skipped(SkipReason::Mismatched));
|
||||||
|
assert_eq!(out, vec![1.0; 4]);
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -292,6 +292,23 @@ impl Scene {
|
|||||||
(self.grid_width, self.grid_height)
|
(self.grid_width, self.grid_height)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// How many pixels of the analysed image one logit cell spans.
|
||||||
|
///
|
||||||
|
/// This is the module header's "resolution, stated plainly" as a number a
|
||||||
|
/// caller can act on: [`Self::rasterise`] will happily hand back a
|
||||||
|
/// full-resolution buffer, and this says how much of that resolution is
|
||||||
|
/// the model's and how much is the bilinear's. At the 1600px proxy the
|
||||||
|
/// application segments at, it is 20.
|
||||||
|
///
|
||||||
|
/// [`crate::refine`] is the one caller, and it needs this rather than a
|
||||||
|
/// pixel count of its own because the erosion that separates evidence from
|
||||||
|
/// smear is *defined* as a multiple of the model's own resolution. Derived
|
||||||
|
/// from the fitted letterbox rather than recomputed from the image size,
|
||||||
|
/// so the two cannot drift.
|
||||||
|
pub fn cell_pixels(&self) -> f32 {
|
||||||
|
GRID_STRIDE as f32 / self.letterbox.scale()
|
||||||
|
}
|
||||||
|
|
||||||
/// One category's weights over the logit grid.
|
/// One category's weights over the logit grid.
|
||||||
pub fn weight(&self, category: usize) -> Option<&[f32]> {
|
pub fn weight(&self, category: usize) -> Option<&[f32]> {
|
||||||
let cells = self.grid_width * self.grid_height;
|
let cells = self.grid_width * self.grid_height;
|
||||||
|
|||||||
@@ -472,6 +472,16 @@ impl Letterbox {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Input pixels per source pixel.
|
||||||
|
///
|
||||||
|
/// An accessor rather than a public field so that the one place that
|
||||||
|
/// converts a model-space size into a source-space one — `Scene::cell_pixels`
|
||||||
|
/// — reads the same number this struct fitted, instead of recomputing a
|
||||||
|
/// letterbox inverse and drifting from it.
|
||||||
|
pub(crate) fn scale(&self) -> f32 {
|
||||||
|
self.scale
|
||||||
|
}
|
||||||
|
|
||||||
/// Resample a source window into the graph's `[1, 3, 640, 640]` input.
|
/// Resample a source window into the graph's `[1, 3, 640, 640]` input.
|
||||||
///
|
///
|
||||||
/// Bilinear, and grey (`0.5`) in the padding — the value the network sees
|
/// Bilinear, and grey (`0.5`) in the padding — the value the network sees
|
||||||
|
|||||||
File diff suppressed because one or more lines are too long
@@ -271,6 +271,22 @@ pub struct Options {
|
|||||||
/// answer. It is the right answer for a bird against sky, so it is offered
|
/// answer. It is the right answer for a bird against sky, so it is offered
|
||||||
/// per-image rather than chosen once for all of them.
|
/// per-image rather than chosen once for all of them.
|
||||||
pub fine: bool,
|
pub fine: bool,
|
||||||
|
|
||||||
|
/// TRACES: FR-DEV-3
|
||||||
|
/// Cut each scene category back to the pixels whose colour agrees with it.
|
||||||
|
///
|
||||||
|
/// The scene model's counterpart to [`Options::fine`], and it exists
|
||||||
|
/// because tiling is not available here: a category has no bounding box to
|
||||||
|
/// tile over — sky is wherever the sky is — so the only route to a sharper
|
||||||
|
/// category edge is the photograph itself. See [`dr_segment::refine`].
|
||||||
|
///
|
||||||
|
/// On by default where `fine` is off, because the two have opposite costs.
|
||||||
|
/// Tiling is six inferences and a five-second wait; this is a distance
|
||||||
|
/// transform and a k-means over a subsample, tens of milliseconds against
|
||||||
|
/// a precompute already measured in hundreds. And it is subtractive, so
|
||||||
|
/// the worst it can do is take too much of a category rather than invent
|
||||||
|
/// one.
|
||||||
|
pub refine: bool,
|
||||||
}
|
}
|
||||||
|
|
||||||
impl Default for Options {
|
impl Default for Options {
|
||||||
@@ -278,6 +294,7 @@ impl Default for Options {
|
|||||||
Self {
|
Self {
|
||||||
confidence: 0.30,
|
confidence: 0.30,
|
||||||
fine: false,
|
fine: false,
|
||||||
|
refine: true,
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -353,7 +370,7 @@ pub fn compute(
|
|||||||
// way. Failures here are logged and dropped rather than propagated: no
|
// way. Failures here are logged and dropped rather than propagated: no
|
||||||
// scene model is an ordinary state, and a photograph that can be masked by
|
// scene model is an ordinary state, and a photograph that can be masked by
|
||||||
// subject should not become unopenable because the categories are absent.
|
// subject should not become unopenable because the categories are absent.
|
||||||
let categories = match scene_categories(&stood_up, uw, uh, orientation) {
|
let categories = match scene_categories(&stood_up, uw, uh, orientation, options.refine) {
|
||||||
Ok(c) => c,
|
Ok(c) => c,
|
||||||
Err(e) => {
|
Err(e) => {
|
||||||
log::info!("no scene categories for this frame: {e}");
|
log::info!("no scene categories for this frame: {e}");
|
||||||
@@ -539,6 +556,7 @@ fn scene_categories(
|
|||||||
width: usize,
|
width: usize,
|
||||||
height: usize,
|
height: usize,
|
||||||
orientation: Orientation,
|
orientation: Orientation,
|
||||||
|
refine: bool,
|
||||||
) -> Result<Vec<CategorySummary>, String> {
|
) -> Result<Vec<CategorySummary>, String> {
|
||||||
let mut model = load_scene_model()?;
|
let mut model = load_scene_model()?;
|
||||||
let scene = model
|
let scene = model
|
||||||
@@ -558,6 +576,42 @@ fn scene_categories(
|
|||||||
let Some(weights) = scene.rasterise(index, width, height) else {
|
let Some(weights) = scene.rasterise(index, width, height) else {
|
||||||
continue;
|
continue;
|
||||||
};
|
};
|
||||||
|
|
||||||
|
// Cut the category back to the pixels whose colour agrees with it.
|
||||||
|
//
|
||||||
|
// Here rather than after `lay_down` because this reads the photograph,
|
||||||
|
// and the photograph is upright at this point — refining on the far
|
||||||
|
// side of the permutation would mean carrying a second, rotated copy
|
||||||
|
// of the proxy across for it to read.
|
||||||
|
let weights = if refine {
|
||||||
|
let (refined, what) = dr_segment::refine_category(
|
||||||
|
&weights,
|
||||||
|
rgb,
|
||||||
|
width,
|
||||||
|
height,
|
||||||
|
scene.cell_pixels(),
|
||||||
|
&dr_segment::RefineOptions::default(),
|
||||||
|
);
|
||||||
|
match what {
|
||||||
|
dr_segment::Refined::Applied { removed } => {
|
||||||
|
log::debug!(
|
||||||
|
"refined '{name}': cut {:.1}% of its weight",
|
||||||
|
removed * 100.0
|
||||||
|
);
|
||||||
|
}
|
||||||
|
// An ordinary state, not a failure — see `dr_segment::refine`.
|
||||||
|
// Logged all the same, because a refinement that silently did
|
||||||
|
// nothing is indistinguishable from the feature being off, and
|
||||||
|
// "the sky looks the way it did before" is what both look like.
|
||||||
|
dr_segment::Refined::Skipped(why) => {
|
||||||
|
log::debug!("left '{name}' coarse: {why:?}");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
refined
|
||||||
|
} else {
|
||||||
|
weights
|
||||||
|
};
|
||||||
|
|
||||||
// Back into sensor space, exactly as an instance mask is: the grid a
|
// Back into sensor space, exactly as an instance mask is: the grid a
|
||||||
// stored layer indexes into has to be the sensor's whatever the model
|
// stored layer indexes into has to be the sensor's whatever the model
|
||||||
// was shown. `lay_down` wants a box too, so it gets the whole frame —
|
// was shown. `lay_down` wants a box too, so it gets the whole frame —
|
||||||
|
|||||||
Reference in New Issue
Block a user