A flag in the sky came out weighted as sky, and no feather setting fixed it. The scene model's logits are `[1, 150, 80, 80]`, so one cell is eight input pixels; at the 1600px proxy the letterbox scale is 0.4 and **one cell is 20 proxy pixels**, which `rasterise`'s bilinear then spreads across one more either side. A flag is a handful of cells whose softmax is dominated by the sky around it. The information was never in the grid, so nothing downstream of the grid can recover it. Tiling is the answer for an instance and is not available here: a category has no bounding box to tile over — sky is wherever the sky is. But the photograph is at full proxy resolution even though the weights are not, and it knows exactly where the flag is. So the model says *what*, and the pixels say *which of them*, which is the division of labour arm C already draws between the instance model and the watershed. ## Seeds, and why the erosion radius is not a guess Threshold the weights high, take `signed_distance`, and keep what is more than 1.5 cells inside. One cell *is* the model's resolution and the bilinear spreads it across one more, so the band either side of the boundary is smear rather than evidence. Deriving the radius from `Scene::cell_pixels` rather than picking a pixel count means it stays right if the proxy edge or the export changes. The mirror of that set is a confident *exterior*, free from the same field. ## Dropping small modes is the step that makes it work Four k-means modes per side, not one Gaussian: sky is blue at the zenith, white where the cloud is and pale at the horizon, and one blob over all three rejects two of them. Then modes holding under 3% of a side are discarded, and without that step the whole thing fails on the case it was built for. A small flag deep in the sky has both a high weight and a large distance from the boundary, so it lands in the interior sample and teaches the model its own colour. It cannot be excluded geometrically. It can be excluded by share. Luminance is weighted at a quarter against chrominance for the same reason the watershed's gradient is. Sky's variance is dominated by luminance, so at equal weight the distribution is a long bright streak that a mid-grey flag sits comfortably inside. A flag is separated by chrominance; a cloud is separated by luminance alone. Not zero, or a dark bird against a bright sky survives. ## Two tests, because either alone is wrong Absolute — is this colour plausible under the category, as a chi-square on the Mahalanobis distance. Comparative — is it likelier inside than outside. A pixel must pass both. The absolute test is what catches the flag, whose colour is far from *both* sides and which the comparative test alone would leave at even odds. The comparative test is what stops the absolute one needing a constant tuned per category. ## What this cannot do, written down rather than left to be discovered An intruder large enough to hold its own mode is kept. By share, a flag over a fifth of the sky and a cloud bank over a fifth of the sky are the same object, and colour does not separate them either — a white cloud is as far from blue sky in chrominance as many intruders are. So `min_cluster` is not a threshold with a correct value waiting to be found; it is the trade-off itself, set where a photographic intruder falls. Both ends are pinned by tests — `a_flag_in_the_sky_is_removed` and `an_intruder_larger_than_min_cluster_survives` — so that moving the number reads as moving the trade-off rather than as fixing a bug. The case left open is a large unrecognised object in a clean category, which wants the boundary snapped to watershed basins and is a different mechanism. ## Safe to apply without a control It is subtractive: the output is the input times a factor in `0..=1`. The worst failure available to it is losing part of a real sky, never gaining a region, so a blue car below the horizon that was never in the mask cannot be pulled into it. And a factor in `0..=1` cannot raise a sum, so `scene.rs`'s partition still holds when every category is refined independently — the weight taken off the flag lands in the unlisted remainder, which is where a flag belongs, ADE20K having no class for one. Every path without the evidence to judge returns the weights untouched and says which path it took. A refinement that silently did nothing is indistinguishable from the feature being off, and an empty seed set fitted to a distribution would reject every pixel. The signature is deliberately unchanged: categories are addressed by name, not by index, so a sharper mask cannot create the stale-index hazard the signature exists to guard against. The example writes `<prefix>-<category>-refined.ppm` beside the coarse one, never instead of it — whether this is an improvement is a comparative judgement and one image cannot answer it. Verified: fmt clean, clippy -D warnings clean, 57 dr-segment tests. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
927 lines
36 KiB
Rust
927 lines
36 KiB
Rust
//! Finding what can be selected in the photograph on screen.
|
|
//!
|
|
//! One model pass, and what it recognised. That is the whole subsystem now.
|
|
//!
|
|
//! # What used to be here
|
|
//!
|
|
//! A watershed over-segmented the image, a merge tree turned that into a
|
|
//! granularity ladder, and a click walked up it (docs/segmentation.md, arms A
|
|
//! and C). It is gone from this path, and the reason is measured rather than
|
|
//! aesthetic: on a real photograph the saddles are near zero almost
|
|
//! everywhere, so the merge order joins everything meaningful before it joins
|
|
//! anything spurious. Cutting a 45,808-basin field of a 22 MP frame to 2,000
|
|
//! regions left **one** region covering nearly the whole picture plus specks
|
|
//! (§15). A ladder that collapses is not a ladder.
|
|
//!
|
|
//! The passes and the hierarchy still exist in `dr-gpu` and `dr-segment`,
|
|
//! tested and documented, because it is the *merge criterion* that fails and
|
|
//! that is one function. What does not exist any more is a product path
|
|
//! through them, or a control offering a choice that does nothing.
|
|
//!
|
|
//! # This is a precompute, and it is slow
|
|
//!
|
|
//! ~700 ms on a 22 MP frame: a proxy render, a readback, and the model. It
|
|
//! runs **once per image, when the user asks**, and never on the frame path.
|
|
//! Every interaction it enables — click a subject, grow a mask, change a
|
|
//! falloff — reads its cached output.
|
|
//!
|
|
//! All of it is on a worker. [`compute`] takes an owned buffer and a
|
|
//! `GpuContext`, which is what makes that possible — nothing here touches the
|
|
//! develop session, and [`crate::develop::SegmentationJob`] is the piece that
|
|
//! carries the proxy render across with it.
|
|
|
|
use std::sync::Arc;
|
|
|
|
use dr_gpu::GpuContext;
|
|
use dr_pipeline::mask::segmentation_signature;
|
|
use dr_types::Orientation;
|
|
|
|
/// One photographic category, over the whole frame.
|
|
///
|
|
/// The counterpart to [`InstanceSummary`] and deliberately thinner: a category
|
|
/// has no box, because it is not one object in one place — sky is wherever the
|
|
/// sky is, in as many disconnected pieces as the frame has windows.
|
|
#[derive(Debug, Clone)]
|
|
pub struct CategorySummary {
|
|
pub name: std::sync::Arc<str>,
|
|
/// Fraction of the frame this category covers, for ordering the list and
|
|
/// for hiding a category that would give the user a control that does
|
|
/// nothing.
|
|
pub coverage: f32,
|
|
/// Coverage at proxy resolution, quantised to a byte — the same
|
|
/// representation, and for the same reasons, as `InstanceSummary::mask`.
|
|
pub mask: Vec<u8>,
|
|
}
|
|
|
|
/// One recognised object.
|
|
#[derive(Debug, Clone)]
|
|
pub struct InstanceSummary {
|
|
pub class_name: Arc<str>,
|
|
pub score: f32,
|
|
/// Coverage at proxy resolution, quantised to a byte.
|
|
///
|
|
/// A byte rather than the `f32` the model produces: 256 levels is far
|
|
/// finer than an edge anyone can see, and at four bytes a pixel a handful
|
|
/// of objects would be most of a hundred megabytes for one photograph.
|
|
///
|
|
/// This is the *source* a distance field is built from, not the mask
|
|
/// itself — `dr_segment::Shaped` turns it into one.
|
|
pub mask: Vec<u8>,
|
|
/// `(x0, y0, x1, y1)` in [`Segmentation::proxy_size`] pixels.
|
|
///
|
|
/// Carried through from `dr_segment::Instance` rather than re-derived
|
|
/// from the mask, so a refine pass knows what region to crop without
|
|
/// scanning a proxy-sized buffer for its own extent.
|
|
pub bbox: (f32, f32, f32, f32),
|
|
}
|
|
|
|
/// One image's recognised objects, ready to mask.
|
|
pub struct Segmentation {
|
|
instances: Vec<InstanceSummary>,
|
|
/// What the scene model made of the same frame, empty when no scene model
|
|
/// could be found.
|
|
///
|
|
/// Empty is an ordinary state, not a failure: a build without the weights
|
|
/// compiled in and without them installed simply offers no categories, the
|
|
/// same way a missing face model turns face indexing off rather than
|
|
/// stopping the app.
|
|
categories: Vec<CategorySummary>,
|
|
/// Identifies this run, so a stored layer can tell whether the index it
|
|
/// holds still means what it meant.
|
|
signature: u64,
|
|
/// The space instance masks are defined in, in **source** proxy pixels.
|
|
///
|
|
/// Everything a mask is built from lives here, which is what lets the
|
|
/// render sample it *after* the framing map rather than before — so a mask
|
|
/// stays on the photograph through a zoom, a pan and a crop.
|
|
proxy: (usize, usize),
|
|
}
|
|
|
|
impl Segmentation {
|
|
pub fn signature(&self) -> u64 {
|
|
self.signature
|
|
}
|
|
|
|
pub fn proxy_size(&self) -> (usize, usize) {
|
|
self.proxy
|
|
}
|
|
|
|
pub fn instances(&self) -> &[InstanceSummary] {
|
|
&self.instances
|
|
}
|
|
|
|
/// One instance's coverage, at [`Self::proxy_size`].
|
|
pub fn instance_mask(&self, index: usize) -> Option<&[u8]> {
|
|
self.instances.get(index).map(|i| i.mask.as_slice())
|
|
}
|
|
|
|
pub fn categories(&self) -> &[CategorySummary] {
|
|
&self.categories
|
|
}
|
|
|
|
/// One category's coverage, at [`Self::proxy_size`].
|
|
///
|
|
/// By name, matching `MaskSource::Category`. A linear scan because there
|
|
/// are eight of them and a map would be more machinery than lookup.
|
|
pub fn category_mask(&self, name: &str) -> Option<&[u8]> {
|
|
self.categories
|
|
.iter()
|
|
.find(|c| &*c.name == name)
|
|
.map(|c| c.mask.as_slice())
|
|
}
|
|
|
|
/// Replace one instance in place, keeping every other index and the
|
|
/// signature unchanged.
|
|
///
|
|
/// What a refine pass calls once it has a sharper mask for the subject at
|
|
/// `index`: the layers pointing at this run by index still mean what they
|
|
/// meant, they just read better pixels now.
|
|
pub fn replace_instance(&mut self, index: usize, instance: InstanceSummary) {
|
|
if let Some(slot) = self.instances.get_mut(index) {
|
|
*slot = instance;
|
|
}
|
|
}
|
|
|
|
/// The strongest instance covering a point in normalised image
|
|
/// coordinates.
|
|
///
|
|
/// Strongest rather than smallest: detections are score-ordered and
|
|
/// overlapping ones are usually the same object found twice, so the more
|
|
/// confident is the better guess. A person in front of a bus wins over the
|
|
/// bus, because the person's mask is the one under the cursor at all.
|
|
pub fn instance_at(&self, x: f32, y: f32) -> Option<usize> {
|
|
if !(0.0..1.0).contains(&x) || !(0.0..1.0).contains(&y) {
|
|
return None;
|
|
}
|
|
let (w, h) = self.proxy;
|
|
if w == 0 || h == 0 {
|
|
return None;
|
|
}
|
|
let px = ((x * w as f32) as usize).min(w - 1);
|
|
let py = ((y * h as f32) as usize).min(h - 1);
|
|
let p = py * w + px;
|
|
|
|
self.instances
|
|
.iter()
|
|
.enumerate()
|
|
.filter(|(_, i)| i.mask.get(p).is_some_and(|&v| v >= 128))
|
|
.max_by(|(_, a), (_, b)| a.score.total_cmp(&b.score))
|
|
.map(|(i, _)| i)
|
|
}
|
|
|
|
/// A false-coloured picture of what a click can select, in source space.
|
|
///
|
|
/// **Transparent where nothing is selectable.** The region map this
|
|
/// replaced covered every pixel and so hid the photograph it was drawn
|
|
/// over; the question an overlay exists to answer is whether an outline
|
|
/// follows the subject, and that can only be answered by seeing both.
|
|
pub fn overlay_rgba(&self) -> (Vec<u8>, u32, u32) {
|
|
let (w, h) = self.proxy;
|
|
let mut out = vec![0u8; w * h * 4];
|
|
|
|
// Weakest first, so where two detections overlap the more confident
|
|
// one is the colour on top — matching which a click would select.
|
|
let mut order: Vec<usize> = (0..self.instances.len()).collect();
|
|
order.sort_by(|&a, &b| self.instances[a].score.total_cmp(&self.instances[b].score));
|
|
|
|
for &i in &order {
|
|
let [r, g, b] = instance_colour(i as u32);
|
|
for (p, &cov) in self.instances[i].mask.iter().enumerate() {
|
|
if cov < 128 || p * 4 + 3 >= out.len() {
|
|
continue;
|
|
}
|
|
out[p * 4] = r;
|
|
out[p * 4 + 1] = g;
|
|
out[p * 4 + 2] = b;
|
|
out[p * 4 + 3] = 255;
|
|
}
|
|
}
|
|
|
|
// The outline drawn opaque white over the fill. It is the part being
|
|
// judged — a fill can look right while its edge sits several pixels
|
|
// off the subject — and it survives the low opacity the fill is
|
|
// composited at.
|
|
let solid = |p: usize| out.get(p * 4 + 3).is_some_and(|&a| a > 0);
|
|
let mut edges = Vec::new();
|
|
for y in 0..h {
|
|
for x in 0..w {
|
|
let p = y * w + x;
|
|
if !solid(p) {
|
|
continue;
|
|
}
|
|
let boundary = (x + 1 == w || !solid(p + 1))
|
|
|| (x == 0 || !solid(p - 1))
|
|
|| (y + 1 == h || !solid(p + w))
|
|
|| (y == 0 || !solid(p - w));
|
|
if boundary {
|
|
edges.push(p);
|
|
}
|
|
}
|
|
}
|
|
for p in edges {
|
|
out[p * 4..p * 4 + 4].copy_from_slice(&[255, 255, 255, 255]);
|
|
}
|
|
|
|
(out, w as u32, h as u32)
|
|
}
|
|
}
|
|
|
|
/// A distinct colour per instance.
|
|
///
|
|
/// Golden-angle hue stepping, deterministic rather than random: the same
|
|
/// object is the same colour every time the overlay is drawn, so the eye can
|
|
/// track it while a mask is being shaped.
|
|
fn instance_colour(index: u32) -> [u8; 3] {
|
|
let h = (index as f32 * 137.508) % 360.0;
|
|
let c = 230.0;
|
|
let x = c * (1.0 - ((h / 60.0) % 2.0 - 1.0).abs());
|
|
let (r, g, b) = match (h / 60.0) as u32 {
|
|
0 => (c, x, 0.0),
|
|
1 => (x, c, 0.0),
|
|
2 => (0.0, c, x),
|
|
3 => (0.0, x, c),
|
|
4 => (x, 0.0, c),
|
|
_ => (c, 0.0, x),
|
|
};
|
|
[r as u8 + 25, g as u8 + 25, b as u8 + 25]
|
|
}
|
|
|
|
/// What to run.
|
|
#[derive(Debug, Clone, Copy, PartialEq)]
|
|
pub struct Options {
|
|
/// Detections below this are dropped.
|
|
///
|
|
/// Deliberately low. A weak detection costs a spurious entry in a list the
|
|
/// user is choosing from, and its score is shown beside it; a missed one
|
|
/// costs a subject that cannot be selected at all, which is the worse
|
|
/// failure for a selection tool.
|
|
pub confidence: f32,
|
|
/// TRACES: FR-DEV-3
|
|
/// Run the model over overlapping tiles instead of the whole frame once.
|
|
///
|
|
/// Off by default and deliberately so. The graph's input is fixed at
|
|
/// 640x640 (docs/segmentation.md F6), so every image is letterboxed into
|
|
/// it and a subject 200px across in a 1600px proxy reaches the model at
|
|
/// 80px — which is where a coarse outline comes from. Tiling is the only
|
|
/// route to more resolution with a fixed window, and it costs one
|
|
/// inference per tile: about 2.8s for a 3x2 grid against 470ms whole-frame.
|
|
///
|
|
/// Six times the wait is the wrong default for the common case, where the
|
|
/// subject is large in frame and whole-frame inference is already the best
|
|
/// answer. It is the right answer for a bird against sky, so it is offered
|
|
/// per-image rather than chosen once for all of them.
|
|
pub fine: bool,
|
|
|
|
/// TRACES: FR-DEV-3
|
|
/// Cut each scene category back to the pixels whose colour agrees with it.
|
|
///
|
|
/// The scene model's counterpart to [`Options::fine`], and it exists
|
|
/// because tiling is not available here: a category has no bounding box to
|
|
/// tile over — sky is wherever the sky is — so the only route to a sharper
|
|
/// category edge is the photograph itself. See [`dr_segment::refine`].
|
|
///
|
|
/// On by default where `fine` is off, because the two have opposite costs.
|
|
/// Tiling is six inferences and a five-second wait; this is a distance
|
|
/// transform and a k-means over a subsample, tens of milliseconds against
|
|
/// a precompute already measured in hundreds. And it is subtractive, so
|
|
/// the worst it can do is take too much of a category rather than invent
|
|
/// one.
|
|
pub refine: bool,
|
|
}
|
|
|
|
impl Default for Options {
|
|
fn default() -> Self {
|
|
Self {
|
|
confidence: 0.30,
|
|
fine: false,
|
|
refine: true,
|
|
}
|
|
}
|
|
}
|
|
|
|
/// Find what can be selected in this photograph.
|
|
///
|
|
/// `rgb` is the rendered proxy — tightly packed RGB floats at
|
|
/// `(width, height)`, in the **sensor's** own orientation. Passed in rather
|
|
/// than derived here because the caller already has it, and re-deriving it
|
|
/// would mean a second readback of something the CPU is holding.
|
|
///
|
|
/// `orientation` is the file's EXIF tag composed with whatever turns the
|
|
/// photographer has since applied — `Framing::effective_orientation`, one
|
|
/// permutation covering both. The model is shown the picture through it and
|
|
/// its answers come back without it, so everything this returns is in sensor
|
|
/// space exactly as it was before the detector was taught to read.
|
|
pub fn compute(
|
|
_ctx: &GpuContext,
|
|
rgb: &[f32],
|
|
width: usize,
|
|
height: usize,
|
|
orientation: Orientation,
|
|
options: &Options,
|
|
) -> Result<Segmentation, String> {
|
|
// The model reads the photograph; everything else here speaks sensor.
|
|
let (stood_up, uw, uh) = upright(rgb, width, height, orientation);
|
|
let found = detect(&stood_up, uw, uh, options.fine)?;
|
|
|
|
let instances: Vec<InstanceSummary> = found
|
|
.iter()
|
|
.filter(|i| i.score >= options.confidence)
|
|
.map(|i| {
|
|
let (mask, bbox) = lay_down(&i.mask, i.bbox, uw, uh, orientation);
|
|
InstanceSummary {
|
|
class_name: i.class_name.clone(),
|
|
score: i.score,
|
|
// Quantised after the permutation, so the byte stored is a
|
|
// rounding of the model's own coverage and not of a copy.
|
|
mask: quantise(&mask),
|
|
bbox,
|
|
}
|
|
})
|
|
.collect();
|
|
|
|
let signature = segmentation_signature(
|
|
width as u32,
|
|
height as u32,
|
|
instances.len() as u32,
|
|
// The tiling choice belongs in the signature as much as the
|
|
// confidence does. A mask stores the signature of the segmentation its
|
|
// region ids index into (`MaskSource::Regions`), and a tiled run finds
|
|
// different instances in a different order — so if the two runs shared
|
|
// a signature, a layer built against the coarse pass would be silently
|
|
// reinterpreted against the fine one. That is a *wrong* mask, which is
|
|
// far worse than a stale one, because nothing announces it.
|
|
//
|
|
// The orientation is in for the same reason and it is not hypothetical:
|
|
// turning the photograph changes what the model recognises, so a run
|
|
// before a quarter turn and a run after it are different instance
|
|
// lists. Two lists that happened to come out the same length would
|
|
// otherwise share a signature, and a layer built against the first
|
|
// would be silently re-indexed into the second.
|
|
options.confidence.to_bits() as u64
|
|
^ if options.fine {
|
|
0x9E37_79B9_7F4A_7C15
|
|
} else {
|
|
0
|
|
}
|
|
^ orientation_key(orientation),
|
|
);
|
|
|
|
// The scene pass, on the same upright frame and laid back down the same
|
|
// way. Failures here are logged and dropped rather than propagated: no
|
|
// scene model is an ordinary state, and a photograph that can be masked by
|
|
// subject should not become unopenable because the categories are absent.
|
|
let categories = match scene_categories(&stood_up, uw, uh, orientation, options.refine) {
|
|
Ok(c) => c,
|
|
Err(e) => {
|
|
log::info!("no scene categories for this frame: {e}");
|
|
Vec::new()
|
|
}
|
|
};
|
|
|
|
Ok(Segmentation {
|
|
instances,
|
|
categories,
|
|
signature,
|
|
// **Sensor space, not the model's.** `lay_down` put every mask back,
|
|
// so the grid a stored layer indexes into is the one it always was —
|
|
// see `upright` for why the model saw a different one.
|
|
proxy: (width, height),
|
|
})
|
|
}
|
|
|
|
/// TRACES: FR-DEV-3 | FR-DEV-3h
|
|
/// Turn the proxy the way the photographer is looking at it.
|
|
///
|
|
/// **Why this exists at all.** A camera held sideways writes its sensor rows
|
|
/// the way it always does, and the render puts them right by way of
|
|
/// `Framing`. The proxy the model reads is deliberately rendered through a
|
|
/// *neutral* graph — the detection has to survive an exposure change, or
|
|
/// every slider would invalidate the masks built on it — and neutral took the
|
|
/// orientation with it. So the detector was handed a portrait frame lying on
|
|
/// its side, and a model trained on upright photographs is very bad at those.
|
|
/// Measured end to end on one 22 MP frame of two people and a dog: `person
|
|
/// 0.36` and nothing else, against `dog 0.82, person 0.61, person 0.49` for
|
|
/// the same pixels stood up.
|
|
///
|
|
/// One line, because the permutation belongs to
|
|
/// [`dr_types::Orientation`] and every other consumer goes through the same
|
|
/// one — the grid's thumbnails included, which is what makes "upright" mean
|
|
/// one thing across the application rather than one thing per caller.
|
|
pub(crate) fn upright(
|
|
rgb: &[f32],
|
|
width: usize,
|
|
height: usize,
|
|
orientation: Orientation,
|
|
) -> (Vec<f32>, usize, usize) {
|
|
let (out, w, h) = orientation.into_shown(rgb, width as u32, height as u32, 3);
|
|
(out, w as usize, h as usize)
|
|
}
|
|
|
|
/// TRACES: FR-DEV-3
|
|
/// Put what the model answered back onto the sensor's grid.
|
|
///
|
|
/// The counterpart of [`upright`], and the two are always used as a pair: a
|
|
/// mask is only ever in the model's frame between those two calls. Returned
|
|
/// together rather than as two functions a caller composes, because calling
|
|
/// one and forgetting the other is silent — the mask lands a quarter turn off
|
|
/// the subject, which reads as a bad detection rather than as a bug.
|
|
///
|
|
/// `dw`/`dh` are the *shown* dimensions, as [`upright`] returned them.
|
|
pub(crate) fn lay_down(
|
|
mask: &[f32],
|
|
bbox: (f32, f32, f32, f32),
|
|
dw: usize,
|
|
dh: usize,
|
|
orientation: Orientation,
|
|
) -> (Vec<f32>, (f32, f32, f32, f32)) {
|
|
let (out, sw, sh) = orientation.into_stored(mask, dw as u32, dh as u32, 1);
|
|
(out, lay_down_bbox(bbox, dw, dh, sw, sh, orientation))
|
|
}
|
|
|
|
/// [`lay_down`] for a box.
|
|
///
|
|
/// Normalised on the way in and scaled on the way out, so the turn itself is
|
|
/// `Orientation::into_stored_rect` rather than a fourth copy of the corner
|
|
/// arithmetic. A box is the one place a permutation can be *nearly* right —
|
|
/// the corners land correctly and `x0 > x1` — so the shared map takes the
|
|
/// extremes and this only has to say what space it is in.
|
|
fn lay_down_bbox(
|
|
bbox: (f32, f32, f32, f32),
|
|
dw: usize,
|
|
dh: usize,
|
|
sw: u32,
|
|
sh: u32,
|
|
orientation: Orientation,
|
|
) -> (f32, f32, f32, f32) {
|
|
if dw == 0 || dh == 0 {
|
|
return bbox;
|
|
}
|
|
let (fw, fh) = (dw as f32, dh as f32);
|
|
let shown = dr_types::ShownRect {
|
|
x: bbox.0 / fw,
|
|
y: bbox.1 / fh,
|
|
width: (bbox.2 - bbox.0) / fw,
|
|
height: (bbox.3 - bbox.1) / fh,
|
|
};
|
|
let stored = orientation.into_stored_rect(shown);
|
|
(
|
|
stored.x * sw as f32,
|
|
stored.y * sh as f32,
|
|
(stored.x + stored.width) * sw as f32,
|
|
(stored.y + stored.height) * sh as f32,
|
|
)
|
|
}
|
|
|
|
/// One of eight transforms, as bits a signature can carry.
|
|
fn orientation_key(orientation: Orientation) -> u64 {
|
|
u64::from(orientation.quarter_turns)
|
|
| (u64::from(orientation.flip_h) << 2)
|
|
| (u64::from(orientation.flip_v) << 3)
|
|
}
|
|
|
|
/// Classes a recognised face may put a name on.
|
|
///
|
|
/// Only these. A face inside a `tv` or a `laptop` is a photograph of someone on
|
|
/// a screen, and renaming the television to "Anna" would be worse than leaving
|
|
/// it as the model found it.
|
|
fn is_person_class(name: &str) -> bool {
|
|
name == "person"
|
|
}
|
|
|
|
impl Segmentation {
|
|
/// Relabel `person` instances with the name of the face inside them.
|
|
///
|
|
/// The segmenter knows it found *a person*; the face index knows *which*
|
|
/// person. Joining them turns "person" in the mask list into "Anna", which
|
|
/// is the difference between COCO's eighty classes and a vocabulary that
|
|
/// includes the user's family — and it is the same click either way, so the
|
|
/// gain is entirely in being able to tell two people apart before clicking.
|
|
///
|
|
/// `faces` are in the same proxy pixels as [`InstanceSummary::bbox`]. The
|
|
/// geometry lives in `dr_face::naming`, where it is testable without a
|
|
/// model or a catalog.
|
|
///
|
|
/// Returns how many instances gained a name. An unrecognised person keeps
|
|
/// the model's own label, which is the right default: "person" is merely
|
|
/// unhelpful, where a wrong name is wrong and the user cannot tell which
|
|
/// they are looking at.
|
|
pub fn apply_names(&mut self, faces: &[dr_face::NamedFace<'_>]) -> usize {
|
|
dr_face::name_instances(
|
|
&mut self.instances,
|
|
faces,
|
|
|i| i.bbox,
|
|
|i| is_person_class(&i.class_name),
|
|
|i, name| i.class_name = name.into(),
|
|
)
|
|
}
|
|
}
|
|
|
|
/// Load the model and run it.
|
|
///
|
|
/// Loading is ~24 ms against the ~470 ms of inference that follows, and this
|
|
/// runs once per image — so caching the session would keep 11 MB of weights
|
|
/// resident for the life of the app to save five percent of a background task.
|
|
fn detect(
|
|
rgb: &[f32],
|
|
width: usize,
|
|
height: usize,
|
|
fine: bool,
|
|
) -> Result<Vec<dr_segment::Instance>, String> {
|
|
let mut model = dr_segment::SemanticModel::embedded().map_err(|e| e.to_string())?;
|
|
let options = dr_segment::SemanticOptions {
|
|
// A quarter shared with each neighbour. It has to exceed zero at all,
|
|
// or a subject sitting on a seam is cut in half by both tiles and
|
|
// recognised by neither; a quarter is enough to carry a whole subject
|
|
// inside one tile at the sizes tiling is reached for.
|
|
tiling: if fine {
|
|
dr_segment::Tiling::Grid { overlap: 0.25 }
|
|
} else {
|
|
dr_segment::Tiling::Whole
|
|
},
|
|
..dr_segment::SemanticOptions::default()
|
|
};
|
|
model
|
|
.detect(rgb, width, height, &options)
|
|
.map_err(|e| e.to_string())
|
|
}
|
|
|
|
/// Weigh the photographic categories, if this build can find a scene model.
|
|
///
|
|
/// Separate from [`detect`] rather than folded into it because the two are
|
|
/// independent: a build with no scene model still segments subjects, and a
|
|
/// frame with no recognisable subject still has sky. Neither failure should
|
|
/// take the other down.
|
|
fn scene_categories(
|
|
rgb: &[f32],
|
|
width: usize,
|
|
height: usize,
|
|
orientation: Orientation,
|
|
refine: bool,
|
|
) -> Result<Vec<CategorySummary>, String> {
|
|
let mut model = load_scene_model()?;
|
|
let scene = model
|
|
.analyse(rgb, width, height)
|
|
.map_err(|e| e.to_string())?;
|
|
|
|
let mut out = Vec::new();
|
|
for (index, name) in scene.categories().iter().enumerate() {
|
|
let coverage = scene.coverage(index);
|
|
// A category the model barely saw is not worth a mask buffer the size
|
|
// of the proxy, and offering it in the list would be offering a
|
|
// control that does nothing when moved. The threshold is the one the
|
|
// example prints against.
|
|
if coverage < 0.005 {
|
|
continue;
|
|
}
|
|
let Some(weights) = scene.rasterise(index, width, height) else {
|
|
continue;
|
|
};
|
|
|
|
// Cut the category back to the pixels whose colour agrees with it.
|
|
//
|
|
// Here rather than after `lay_down` because this reads the photograph,
|
|
// and the photograph is upright at this point — refining on the far
|
|
// side of the permutation would mean carrying a second, rotated copy
|
|
// of the proxy across for it to read.
|
|
let weights = if refine {
|
|
let (refined, what) = dr_segment::refine_category(
|
|
&weights,
|
|
rgb,
|
|
width,
|
|
height,
|
|
scene.cell_pixels(),
|
|
&dr_segment::RefineOptions::default(),
|
|
);
|
|
match what {
|
|
dr_segment::Refined::Applied { removed } => {
|
|
log::debug!(
|
|
"refined '{name}': cut {:.1}% of its weight",
|
|
removed * 100.0
|
|
);
|
|
}
|
|
// An ordinary state, not a failure — see `dr_segment::refine`.
|
|
// Logged all the same, because a refinement that silently did
|
|
// nothing is indistinguishable from the feature being off, and
|
|
// "the sky looks the way it did before" is what both look like.
|
|
dr_segment::Refined::Skipped(why) => {
|
|
log::debug!("left '{name}' coarse: {why:?}");
|
|
}
|
|
}
|
|
refined
|
|
} else {
|
|
weights
|
|
};
|
|
|
|
// Back into sensor space, exactly as an instance mask is: the grid a
|
|
// stored layer indexes into has to be the sensor's whatever the model
|
|
// was shown. `lay_down` wants a box too, so it gets the whole frame —
|
|
// a category has no meaningful extent.
|
|
let (mask, _) = lay_down(
|
|
&weights,
|
|
(0.0, 0.0, width as f32, height as f32),
|
|
width,
|
|
height,
|
|
orientation,
|
|
);
|
|
out.push(CategorySummary {
|
|
name: name.clone(),
|
|
coverage,
|
|
mask: quantise(&mask),
|
|
});
|
|
}
|
|
|
|
// Largest first, which is the order the list is worth reading in.
|
|
out.sort_by(|a, b| b.coverage.total_cmp(&a.coverage));
|
|
Ok(out)
|
|
}
|
|
|
|
/// Find a scene model: compiled in if this build has one, installed otherwise.
|
|
///
|
|
/// The order matters. A build with the weights compiled in should not be
|
|
/// silently overridden by a stale file in a data directory, and a build
|
|
/// without them has nothing to fall back *from* — so "embedded, then
|
|
/// installed" is the only ordering that is not surprising either way.
|
|
fn load_scene_model() -> Result<dr_segment::SceneModel, String> {
|
|
#[cfg(feature = "scene-model")]
|
|
{
|
|
dr_segment::SceneModel::embedded().map_err(|e| e.to_string())
|
|
}
|
|
#[cfg(not(feature = "scene-model"))]
|
|
{
|
|
// Account-independent, like `shared_face_models_dir` and for the same
|
|
// reason: this runs on a worker with no session in hand. Android
|
|
// unpacks the APK's copy to exactly this directory before any store
|
|
// opens.
|
|
let dir = crate::library::shared_face_models_dir();
|
|
dr_segment::SceneModel::from_path(
|
|
dir.join("yolo26s-sem-ade20k.onnx"),
|
|
dir.join("yolo26s-sem-ade20k.classes.json"),
|
|
dir.join("categories.txt"),
|
|
)
|
|
.map_err(|e| e.to_string())
|
|
}
|
|
}
|
|
|
|
/// The model's soft coverage, to a byte per pixel.
|
|
///
|
|
/// Rounded rather than truncated, so a coverage of exactly 0.5 lands on the
|
|
/// threshold the selection tests against instead of one below it.
|
|
fn quantise(mask: &[f32]) -> Vec<u8> {
|
|
mask.iter()
|
|
.map(|&v| (v.clamp(0.0, 1.0) * 255.0).round() as u8)
|
|
.collect()
|
|
}
|
|
|
|
#[cfg(test)]
|
|
mod tests {
|
|
use super::*;
|
|
|
|
/// Two objects: a big weak one on the left, a small strong one that
|
|
/// overlaps it.
|
|
fn overlapping() -> Segmentation {
|
|
let (w, h) = (8usize, 4usize);
|
|
let mut big = vec![0u8; w * h];
|
|
let mut small = vec![0u8; w * h];
|
|
for y in 0..h {
|
|
for x in 0..6 {
|
|
big[y * w + x] = 255;
|
|
}
|
|
for x in 4..8 {
|
|
small[y * w + x] = 255;
|
|
}
|
|
}
|
|
|
|
Segmentation {
|
|
instances: vec![
|
|
InstanceSummary {
|
|
class_name: "bus".into(),
|
|
score: 0.5,
|
|
mask: big,
|
|
bbox: (0.0, 0.0, 6.0, h as f32),
|
|
},
|
|
InstanceSummary {
|
|
class_name: "person".into(),
|
|
score: 0.9,
|
|
mask: small,
|
|
bbox: (4.0, 0.0, 8.0, h as f32),
|
|
},
|
|
],
|
|
categories: Vec::new(),
|
|
signature: 1,
|
|
proxy: (w, h),
|
|
}
|
|
}
|
|
|
|
#[test]
|
|
fn a_click_finds_the_object_under_it() {
|
|
let seg = overlapping();
|
|
assert_eq!(seg.instance_at(0.1, 0.5), Some(0), "only the bus here");
|
|
assert_eq!(seg.instance_at(0.95, 0.5), Some(1), "only the person here");
|
|
}
|
|
|
|
/// The overlap rule, and the one that decides what a click means where two
|
|
/// detections cover the same pixel.
|
|
#[test]
|
|
fn overlapping_objects_resolve_to_the_more_confident() {
|
|
let seg = overlapping();
|
|
assert_eq!(
|
|
seg.instance_at(0.6, 0.5),
|
|
Some(1),
|
|
"the person at 0.9 beats the bus at 0.5"
|
|
);
|
|
}
|
|
|
|
#[test]
|
|
fn a_click_outside_the_frame_selects_nothing() {
|
|
let seg = overlapping();
|
|
assert_eq!(seg.instance_at(-0.1, 0.5), None);
|
|
assert_eq!(seg.instance_at(1.5, 0.5), None);
|
|
assert_eq!(seg.instance_at(0.5, 1.0), None, "the far edge is exclusive");
|
|
}
|
|
|
|
#[test]
|
|
fn a_click_on_nothing_selects_nothing() {
|
|
let seg = Segmentation {
|
|
instances: Vec::new(),
|
|
categories: Vec::new(),
|
|
signature: 1,
|
|
proxy: (4, 4),
|
|
};
|
|
assert_eq!(seg.instance_at(0.5, 0.5), None);
|
|
}
|
|
|
|
#[test]
|
|
fn colours_are_stable_and_distinct() {
|
|
assert_eq!(instance_colour(7), instance_colour(7));
|
|
assert_ne!(instance_colour(0), instance_colour(1));
|
|
assert_ne!(instance_colour(1), instance_colour(2));
|
|
}
|
|
|
|
#[test]
|
|
fn every_colour_is_visible_against_a_photograph() {
|
|
for i in 0..64u32 {
|
|
let [r, g, b] = instance_colour(i);
|
|
assert!(r.max(g).max(b) >= 200, "instance {i} is too dark");
|
|
}
|
|
}
|
|
|
|
/// The overlay must not cover the picture: that is the difference between
|
|
/// this and the region map it replaced.
|
|
#[test]
|
|
fn the_overlay_is_transparent_where_nothing_was_found() {
|
|
let seg = Segmentation {
|
|
instances: Vec::new(),
|
|
categories: Vec::new(),
|
|
signature: 1,
|
|
proxy: (4, 4),
|
|
};
|
|
let (px, w, h) = seg.overlay_rgba();
|
|
assert_eq!((w, h), (4, 4));
|
|
assert!(px.chunks_exact(4).all(|p| p[3] == 0), "nothing to draw");
|
|
}
|
|
|
|
#[test]
|
|
fn the_overlay_outlines_what_it_fills() {
|
|
let seg = overlapping();
|
|
let (px, w, _) = seg.overlay_rgba();
|
|
let at = |x: usize, y: usize| {
|
|
let p = (y * w as usize + x) * 4;
|
|
[px[p], px[p + 1], px[p + 2], px[p + 3]]
|
|
};
|
|
// The rightmost column of the person is against the frame edge, so it
|
|
// is an outline pixel.
|
|
assert_eq!(at(7, 1), [255, 255, 255, 255]);
|
|
// And an interior pixel keeps its fill.
|
|
assert_ne!(at(2, 1)[3], 0, "the bus is filled");
|
|
assert_ne!(at(2, 1), [255, 255, 255, 255], "and not all outline");
|
|
}
|
|
|
|
#[test]
|
|
fn quantising_rounds_rather_than_truncates() {
|
|
// Exactly half must reach the threshold the selection tests against.
|
|
assert_eq!(quantise(&[0.5]), vec![128]);
|
|
assert_eq!(quantise(&[0.0, 1.0]), vec![0, 255]);
|
|
// And values outside the range cannot wrap.
|
|
assert_eq!(quantise(&[-1.0, 2.0]), vec![0, 255]);
|
|
}
|
|
|
|
/// A non-square, wholly asymmetric grid: every pixel is its own index, so
|
|
/// any permutation that is not the intended one shows up as a mismatch
|
|
/// rather than being hidden by a symmetry.
|
|
fn ramp(w: usize, h: usize) -> Vec<f32> {
|
|
(0..w * h).flat_map(|i| [i as f32, 0.0, 0.0]).collect()
|
|
}
|
|
|
|
fn red(rgb: &[f32]) -> Vec<f32> {
|
|
rgb.chunks_exact(3).map(|p| p[0]).collect()
|
|
}
|
|
|
|
/// The property the whole fix rests on: what the model is shown and what
|
|
/// comes back are the same permutation, run in opposite directions. If
|
|
/// they ever disagree, every subject mask lands somewhere other than its
|
|
/// subject — and looks like a mask while doing it.
|
|
#[test]
|
|
fn standing_a_frame_up_and_laying_it_down_is_the_identity() {
|
|
const W: usize = 5;
|
|
const H: usize = 3;
|
|
let source = ramp(W, H);
|
|
|
|
for tag in 1..=8u16 {
|
|
let o = Orientation::from_exif(tag);
|
|
let (up, uw, uh) = upright(&source, W, H, o);
|
|
let (ow, oh) = o.oriented_size(W as u32, H as u32);
|
|
assert_eq!(
|
|
(uw, uh),
|
|
(ow as usize, oh as usize),
|
|
"tag {tag}: the upright size is the oriented one"
|
|
);
|
|
|
|
let (back, _) = lay_down(&red(&up), (0.0, 0.0, 1.0, 1.0), uw, uh, o);
|
|
assert_eq!(back, red(&source), "tag {tag} did not come back");
|
|
}
|
|
}
|
|
|
|
/// The colour channels must travel together. Reading a pixel three times
|
|
/// with one index arithmetic mistake gives a plausible image with its
|
|
/// channels sheared, which the model would still detect *something* in.
|
|
#[test]
|
|
fn a_turn_carries_whole_pixels() {
|
|
let rgb: Vec<f32> = (0..2 * 3)
|
|
.flat_map(|i| [i as f32, i as f32 + 100.0, i as f32 + 200.0])
|
|
.collect();
|
|
let o = Orientation::from_exif(6);
|
|
let (up, uw, uh) = upright(&rgb, 2, 3, o);
|
|
|
|
assert_eq!((uw, uh), (3, 2));
|
|
for p in up.chunks_exact(3) {
|
|
assert_eq!(p[1], p[0] + 100.0, "green left its pixel");
|
|
assert_eq!(p[2], p[0] + 200.0, "blue left its pixel");
|
|
}
|
|
}
|
|
|
|
/// A portrait frame is the case this exists for: the sensor is landscape,
|
|
/// the photograph is not, and the model has to be given the photograph.
|
|
#[test]
|
|
fn a_sideways_frame_reaches_the_model_upright() {
|
|
let o = Orientation::from_exif(6);
|
|
assert!(!o.is_normal());
|
|
|
|
let (_, uw, uh) = upright(&ramp(1600, 1066), 1600, 1066, o);
|
|
assert_eq!((uw, uh), (1066, 1600), "the model still got a landscape");
|
|
}
|
|
|
|
/// Where the model's box ends up, worked out by hand for the one turn a
|
|
/// portrait phone or a sideways body actually writes.
|
|
#[test]
|
|
fn a_box_comes_back_in_sensor_pixels() {
|
|
let o = Orientation::from_exif(6);
|
|
// Shown 4x6; the sensor it came from is 6x4.
|
|
let bbox = lay_down_bbox((0.0, 0.0, 2.0, 3.0), 4, 6, 6, 4, o);
|
|
assert_eq!(bbox, (0.0, 2.0, 3.0, 4.0));
|
|
}
|
|
|
|
/// Whatever a turn does to a box, it must still read low-to-high — a
|
|
/// permutation exchanges which corner is which, and the rest of the mask
|
|
/// pipeline measures `(x1 - x0)` without checking the sign.
|
|
#[test]
|
|
fn a_restored_box_keeps_its_corners_in_order() {
|
|
for tag in 1..=8u16 {
|
|
let o = Orientation::from_exif(tag);
|
|
let (sw, sh) = o.oriented_size(9, 6);
|
|
let (x0, y0, x1, y1) = lay_down_bbox((1.0, 2.0, 7.0, 5.0), 9, 6, sw, sh, o);
|
|
assert!(x0 <= x1, "tag {tag}: x runs backwards");
|
|
assert!(y0 <= y1, "tag {tag}: y runs backwards");
|
|
// A permutation moves a box; it does not resize one.
|
|
let (sw, sh) = o.oriented_size(9, 6);
|
|
assert!(x1 <= sw as f32 && y1 <= sh as f32, "tag {tag}: box escaped");
|
|
assert!((((x1 - x0) * (y1 - y0)) - 18.0).abs() < 1e-3, "tag {tag}");
|
|
}
|
|
}
|
|
|
|
/// Turning the photograph changes what the model recognises, so the two
|
|
/// runs are different instance lists. If they could share a signature, a
|
|
/// layer built against one would be silently re-indexed into the other —
|
|
/// the same failure the tiling flag is in the signature to prevent.
|
|
#[test]
|
|
fn turning_the_photograph_changes_the_signature() {
|
|
let mut seen = std::collections::HashSet::new();
|
|
for tag in 1..=8u16 {
|
|
let o = Orientation::from_exif(tag);
|
|
assert!(
|
|
seen.insert(orientation_key(o)),
|
|
"tag {tag} shares a key with an earlier one"
|
|
);
|
|
}
|
|
assert_eq!(seen.len(), 8, "eight tags, but some collapsed");
|
|
}
|
|
|
|
#[test]
|
|
fn confidence_changes_the_signature() {
|
|
// A different threshold is a different instance list, so the indices a
|
|
// stored layer holds mean something else.
|
|
let a = segmentation_signature(100, 100, 3, Options::default().confidence.to_bits() as u64);
|
|
let b = segmentation_signature(100, 100, 3, 0.9f32.to_bits() as u64);
|
|
assert_ne!(a, b);
|
|
}
|
|
}
|