Files
DarkRoom/ui/dr-ui/src/segmentation.rs
T
dtourolleandClaude Opus 5 4a82753d22
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
Build and test / Desktop (Linux) (push) Successful in 19m7s
Build and test / Layer separation (push) Successful in 25s
Traceability / Requirement traces (push) Successful in 23s
Build and test / Android (aarch64) (push) Failing after 33m10s
Show the detector the photograph, not the sensor's scanlines
"Find subjects" was handed the proxy in the sensor's own orientation, so
every frame shot on a body held sideways reached the model lying on its
side — and a model trained on upright photographs is very bad at those.
Measured end to end on a 22 MP frame of two people and a dog: `person
0.36` and nothing else, against `dog 0.82, person 0.61, person 0.49` for
the same pixels stood up. Nothing failed; the panel simply offered one
poor subject where there were three good ones.

The orientation was never dropped on purpose. The proxy is deliberately
rendered through a *neutral* graph — the detection has to survive an
exposure change, or every slider would invalidate the masks built on it
— and neutral took the file's orientation with it along with everything
else. Landscape frames were unaffected, which is why it stood for as
long as it did.

The turn is `Orientation::source_pixel`, the same function the grid's
thumbnails already go through, so the detector and the thumbnailer now
agree about which way is up rather than holding two opinions. What it is
turned by is `Framing::effective_orientation` — the file's EXIF tag and
the photographer's own rotations composed into one permutation, by the
group law rather than by adding the turns, which is a distinction
`Framing` already had to make and had already tested. Rotating the
picture and pressing the button again therefore does what it looks like
it does.

The proxy stays in sensor space and the masks come back into it. That is
not a detail to be tidied later: the generated shader samples the mask
array at `uv_src`, *after* the framing map, so a mask stored upright
would sit a quarter turn off the subject it was drawn around. That is a
wrong mask rather than a weak one, and nothing announces it. So the
picture is stood up for the model and laid back down for everything
else, and `upright`/`lay_down` are returned as a pair because calling
one and forgetting the other is silent.

Both directions are the one function: `upright` gathers through
`source_pixel` and `lay_down` scatters through it. A quarter turn is a
bijection of the pixel grid, so the round trip is exact — no filter, no
resampling, and no hole to fill — and an inverse written out by hand
would be a second thing to keep in step, whose way of being wrong is a
mask mirrored about the wrong axis, which still looks like a mask.

The orientation joins the confidence and the tiling flag in the
segmentation signature, and for the same reason: turning the photograph
changes what the model recognises, so two runs either side of a rotation
are different instance lists. Two that happened to come out the same
length would otherwise share a signature and a stored layer would be
silently re-indexed from one into the other.

The refine pass had it too — it re-runs the model over a crop rendered
in the same sensor space — so it makes the same turn, and would
otherwise have handed back a worse mask than the one it was asked to
improve, on the subject the photographer had just pointed at.

`dr-gpu`'s `local` example is fixed with it. It exists to be the
shipping path with pictures attached, and a diagnostic that reproduces
the bug it is meant to catch is a trap for whoever reads it next.

Seven tests. The round trip is the identity over all eight EXIF tags on
a non-square asymmetric grid; a turn carries whole pixels rather than
shearing the channels apart; a sideways frame reaches the model
upright; a box comes back in sensor pixels, worked out by hand for the
one turn a portrait frame actually writes; a restored box still reads
low-to-high for every tag, since the rest of the pipeline takes
`x1 - x0` without checking the sign; and the eight tags cannot collapse
into one signature key. The existing composition test now runs against
`effective_orientation` itself, over all 8 x 16 baseline-and-user pairs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 22:40:58 +02:00

754 lines
29 KiB
Rust

//! Finding what can be selected in the photograph on screen.
//!
//! One model pass, and what it recognised. That is the whole subsystem now.
//!
//! # What used to be here
//!
//! A watershed over-segmented the image, a merge tree turned that into a
//! granularity ladder, and a click walked up it (docs/segmentation.md, arms A
//! and C). It is gone from this path, and the reason is measured rather than
//! aesthetic: on a real photograph the saddles are near zero almost
//! everywhere, so the merge order joins everything meaningful before it joins
//! anything spurious. Cutting a 45,808-basin field of a 22 MP frame to 2,000
//! regions left **one** region covering nearly the whole picture plus specks
//! (§15). A ladder that collapses is not a ladder.
//!
//! The passes and the hierarchy still exist in `dr-gpu` and `dr-segment`,
//! tested and documented, because it is the *merge criterion* that fails and
//! that is one function. What does not exist any more is a product path
//! through them, or a control offering a choice that does nothing.
//!
//! # This is a precompute, and it is slow
//!
//! ~700 ms on a 22 MP frame: a proxy render, a readback, and the model. It
//! runs **once per image, when the user asks**, and never on the frame path.
//! Every interaction it enables — click a subject, grow a mask, change a
//! falloff — reads its cached output.
//!
//! All of it is on a worker. [`compute`] takes an owned buffer and a
//! `GpuContext`, which is what makes that possible — nothing here touches the
//! develop session, and [`crate::develop::SegmentationJob`] is the piece that
//! carries the proxy render across with it.
use std::sync::Arc;
use dr_gpu::GpuContext;
use dr_pipeline::mask::segmentation_signature;
use dr_types::Orientation;
/// One recognised object.
#[derive(Debug, Clone)]
pub struct InstanceSummary {
pub class_name: Arc<str>,
pub score: f32,
/// Coverage at proxy resolution, quantised to a byte.
///
/// A byte rather than the `f32` the model produces: 256 levels is far
/// finer than an edge anyone can see, and at four bytes a pixel a handful
/// of objects would be most of a hundred megabytes for one photograph.
///
/// This is the *source* a distance field is built from, not the mask
/// itself — `dr_segment::Shaped` turns it into one.
pub mask: Vec<u8>,
/// `(x0, y0, x1, y1)` in [`Segmentation::proxy_size`] pixels.
///
/// Carried through from `dr_segment::Instance` rather than re-derived
/// from the mask, so a refine pass knows what region to crop without
/// scanning a proxy-sized buffer for its own extent.
pub bbox: (f32, f32, f32, f32),
}
/// One image's recognised objects, ready to mask.
pub struct Segmentation {
instances: Vec<InstanceSummary>,
/// Identifies this run, so a stored layer can tell whether the index it
/// holds still means what it meant.
signature: u64,
/// The space instance masks are defined in, in **source** proxy pixels.
///
/// Everything a mask is built from lives here, which is what lets the
/// render sample it *after* the framing map rather than before — so a mask
/// stays on the photograph through a zoom, a pan and a crop.
proxy: (usize, usize),
}
impl Segmentation {
pub fn signature(&self) -> u64 {
self.signature
}
pub fn proxy_size(&self) -> (usize, usize) {
self.proxy
}
pub fn instances(&self) -> &[InstanceSummary] {
&self.instances
}
/// One instance's coverage, at [`Self::proxy_size`].
pub fn instance_mask(&self, index: usize) -> Option<&[u8]> {
self.instances.get(index).map(|i| i.mask.as_slice())
}
/// Replace one instance in place, keeping every other index and the
/// signature unchanged.
///
/// What a refine pass calls once it has a sharper mask for the subject at
/// `index`: the layers pointing at this run by index still mean what they
/// meant, they just read better pixels now.
pub fn replace_instance(&mut self, index: usize, instance: InstanceSummary) {
if let Some(slot) = self.instances.get_mut(index) {
*slot = instance;
}
}
/// The strongest instance covering a point in normalised image
/// coordinates.
///
/// Strongest rather than smallest: detections are score-ordered and
/// overlapping ones are usually the same object found twice, so the more
/// confident is the better guess. A person in front of a bus wins over the
/// bus, because the person's mask is the one under the cursor at all.
pub fn instance_at(&self, x: f32, y: f32) -> Option<usize> {
if !(0.0..1.0).contains(&x) || !(0.0..1.0).contains(&y) {
return None;
}
let (w, h) = self.proxy;
if w == 0 || h == 0 {
return None;
}
let px = ((x * w as f32) as usize).min(w - 1);
let py = ((y * h as f32) as usize).min(h - 1);
let p = py * w + px;
self.instances
.iter()
.enumerate()
.filter(|(_, i)| i.mask.get(p).is_some_and(|&v| v >= 128))
.max_by(|(_, a), (_, b)| a.score.total_cmp(&b.score))
.map(|(i, _)| i)
}
/// A false-coloured picture of what a click can select, in source space.
///
/// **Transparent where nothing is selectable.** The region map this
/// replaced covered every pixel and so hid the photograph it was drawn
/// over; the question an overlay exists to answer is whether an outline
/// follows the subject, and that can only be answered by seeing both.
pub fn overlay_rgba(&self) -> (Vec<u8>, u32, u32) {
let (w, h) = self.proxy;
let mut out = vec![0u8; w * h * 4];
// Weakest first, so where two detections overlap the more confident
// one is the colour on top — matching which a click would select.
let mut order: Vec<usize> = (0..self.instances.len()).collect();
order.sort_by(|&a, &b| self.instances[a].score.total_cmp(&self.instances[b].score));
for &i in &order {
let [r, g, b] = instance_colour(i as u32);
for (p, &cov) in self.instances[i].mask.iter().enumerate() {
if cov < 128 || p * 4 + 3 >= out.len() {
continue;
}
out[p * 4] = r;
out[p * 4 + 1] = g;
out[p * 4 + 2] = b;
out[p * 4 + 3] = 255;
}
}
// The outline drawn opaque white over the fill. It is the part being
// judged — a fill can look right while its edge sits several pixels
// off the subject — and it survives the low opacity the fill is
// composited at.
let solid = |p: usize| out.get(p * 4 + 3).is_some_and(|&a| a > 0);
let mut edges = Vec::new();
for y in 0..h {
for x in 0..w {
let p = y * w + x;
if !solid(p) {
continue;
}
let boundary = (x + 1 == w || !solid(p + 1))
|| (x == 0 || !solid(p - 1))
|| (y + 1 == h || !solid(p + w))
|| (y == 0 || !solid(p - w));
if boundary {
edges.push(p);
}
}
}
for p in edges {
out[p * 4..p * 4 + 4].copy_from_slice(&[255, 255, 255, 255]);
}
(out, w as u32, h as u32)
}
}
/// A distinct colour per instance.
///
/// Golden-angle hue stepping, deterministic rather than random: the same
/// object is the same colour every time the overlay is drawn, so the eye can
/// track it while a mask is being shaped.
fn instance_colour(index: u32) -> [u8; 3] {
let h = (index as f32 * 137.508) % 360.0;
let c = 230.0;
let x = c * (1.0 - ((h / 60.0) % 2.0 - 1.0).abs());
let (r, g, b) = match (h / 60.0) as u32 {
0 => (c, x, 0.0),
1 => (x, c, 0.0),
2 => (0.0, c, x),
3 => (0.0, x, c),
4 => (x, 0.0, c),
_ => (c, 0.0, x),
};
[r as u8 + 25, g as u8 + 25, b as u8 + 25]
}
/// What to run.
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct Options {
/// Detections below this are dropped.
///
/// Deliberately low. A weak detection costs a spurious entry in a list the
/// user is choosing from, and its score is shown beside it; a missed one
/// costs a subject that cannot be selected at all, which is the worse
/// failure for a selection tool.
pub confidence: f32,
/// TRACES: FR-DEV-3
/// Run the model over overlapping tiles instead of the whole frame once.
///
/// Off by default and deliberately so. The graph's input is fixed at
/// 640x640 (docs/segmentation.md F6), so every image is letterboxed into
/// it and a subject 200px across in a 1600px proxy reaches the model at
/// 80px — which is where a coarse outline comes from. Tiling is the only
/// route to more resolution with a fixed window, and it costs one
/// inference per tile: about 2.8s for a 3x2 grid against 470ms whole-frame.
///
/// Six times the wait is the wrong default for the common case, where the
/// subject is large in frame and whole-frame inference is already the best
/// answer. It is the right answer for a bird against sky, so it is offered
/// per-image rather than chosen once for all of them.
pub fine: bool,
}
impl Default for Options {
fn default() -> Self {
Self {
confidence: 0.30,
fine: false,
}
}
}
/// Find what can be selected in this photograph.
///
/// `rgb` is the rendered proxy — tightly packed RGB floats at
/// `(width, height)`, in the **sensor's** own orientation. Passed in rather
/// than derived here because the caller already has it, and re-deriving it
/// would mean a second readback of something the CPU is holding.
///
/// `orientation` is the file's EXIF tag composed with whatever turns the
/// photographer has since applied — `Framing::effective_orientation`, one
/// permutation covering both. The model is shown the picture through it and
/// its answers come back without it, so everything this returns is in sensor
/// space exactly as it was before the detector was taught to read.
pub fn compute(
_ctx: &GpuContext,
rgb: &[f32],
width: usize,
height: usize,
orientation: Orientation,
options: &Options,
) -> Result<Segmentation, String> {
// The model reads the photograph; everything else here speaks sensor.
let (stood_up, uw, uh) = upright(rgb, width, height, orientation);
let found = detect(&stood_up, uw, uh, options.fine)?;
let instances: Vec<InstanceSummary> = found
.iter()
.filter(|i| i.score >= options.confidence)
.map(|i| {
let (mask, bbox) = lay_down(&i.mask, i.bbox, uw, uh, orientation);
InstanceSummary {
class_name: i.class_name.clone(),
score: i.score,
// Quantised after the permutation, so the byte stored is a
// rounding of the model's own coverage and not of a copy.
mask: quantise(&mask),
bbox,
}
})
.collect();
let signature = segmentation_signature(
width as u32,
height as u32,
instances.len() as u32,
// The tiling choice belongs in the signature as much as the
// confidence does. A mask stores the signature of the segmentation its
// region ids index into (`MaskSource::Regions`), and a tiled run finds
// different instances in a different order — so if the two runs shared
// a signature, a layer built against the coarse pass would be silently
// reinterpreted against the fine one. That is a *wrong* mask, which is
// far worse than a stale one, because nothing announces it.
//
// The orientation is in for the same reason and it is not hypothetical:
// turning the photograph changes what the model recognises, so a run
// before a quarter turn and a run after it are different instance
// lists. Two lists that happened to come out the same length would
// otherwise share a signature, and a layer built against the first
// would be silently re-indexed into the second.
options.confidence.to_bits() as u64
^ if options.fine {
0x9E37_79B9_7F4A_7C15
} else {
0
}
^ orientation_key(orientation),
);
Ok(Segmentation {
instances,
signature,
// **Sensor space, not the model's.** `lay_down` put every mask back,
// so the grid a stored layer indexes into is the one it always was —
// see `upright` for why the model saw a different one.
proxy: (width, height),
})
}
/// TRACES: FR-DEV-3 | FR-DEV-3h
/// Turn the proxy the way the photographer is looking at it.
///
/// **Why this exists at all.** A camera held sideways writes its sensor rows
/// the way it always does, and the render puts them right by way of
/// `Framing`. The proxy the model reads is deliberately rendered through a
/// *neutral* graph — the detection has to survive an exposure change, or
/// every slider would invalidate the masks built on it — and neutral took the
/// orientation with it. So the detector was handed a portrait frame lying on
/// its side, and a model trained on upright photographs is very bad at those.
/// Measured end to end on one 22 MP frame of two people and a dog: `person
/// 0.36` and nothing else, against `dog 0.82, person 0.61, person 0.49` for
/// the same pixels stood up.
///
/// The turn is [`Orientation::source_pixel`], which is the function the grid's
/// thumbnails already go through (`dr_decode::Preview::apply_orientation`) —
/// so the detector now reads exactly the kind of image the thumbnailer makes,
/// rather than a second opinion about what "upright" means.
///
/// A quarter turn and its mirrors are a permutation of the pixel grid, so this
/// is exact: no filter, no resampling, and no edge softened on the way in that
/// would have to be judged on the way out.
pub(crate) fn upright(
rgb: &[f32],
width: usize,
height: usize,
orientation: Orientation,
) -> (Vec<f32>, usize, usize) {
if orientation.is_normal() || width == 0 || height == 0 {
return (rgb.to_vec(), width, height);
}
let (dw, dh) = oriented(width, height, orientation);
let mut out = vec![0.0f32; dw * dh * 3];
for y in 0..dh {
for x in 0..dw {
let (sx, sy) = orientation.source_pixel(x as u32, y as u32, dw as u32, dh as u32);
let s = (sy as usize * width + sx as usize) * 3;
let d = (y * dw + x) * 3;
out[d..d + 3].copy_from_slice(&rgb[s..s + 3]);
}
}
(out, dw, dh)
}
/// TRACES: FR-DEV-3
/// Put what the model answered back onto the sensor's grid.
///
/// The counterpart of [`upright`], and the two are always used as a pair: a
/// mask is only ever in the model's frame between those two calls. Returned
/// together rather than as two functions a caller composes, because calling
/// one and forgetting the other is silent — the mask lands a quarter turn off
/// the subject, which reads as a bad detection rather than as a bug.
///
/// `dw`/`dh` are the *upright* dimensions, as [`upright`] returned them.
pub(crate) fn lay_down(
mask: &[f32],
bbox: (f32, f32, f32, f32),
dw: usize,
dh: usize,
orientation: Orientation,
) -> (Vec<f32>, (f32, f32, f32, f32)) {
if orientation.is_normal() {
return (mask.to_vec(), bbox);
}
(
lay_down_mask(mask, dw, dh, orientation),
lay_down_bbox(bbox, dw, dh, orientation),
)
}
/// The same permutation as [`upright`], read as a scatter rather than a
/// gather: a quarter turn is a bijection of the grid, so writing every upright
/// pixel to where it came from fills the sensor-space mask exactly once and
/// leaves no hole. Running the one function in the one direction is the point
/// — an inverse written out by hand is a second thing to keep in step, and its
/// way of being wrong is a mask mirrored about the wrong axis, which still
/// looks like a mask.
fn lay_down_mask(mask: &[f32], dw: usize, dh: usize, orientation: Orientation) -> Vec<f32> {
let (sw, sh) = oriented(dw, dh, orientation);
let mut out = vec![0.0f32; sw * sh];
for y in 0..dh {
for x in 0..dw {
let (sx, sy) = orientation.source_pixel(x as u32, y as u32, dw as u32, dh as u32);
out[sy as usize * sw + sx as usize] = mask[y * dw + x];
}
}
out
}
/// [`lay_down_mask`] for a box.
///
/// The corners go through the same permutation in *continuous* coordinates —
/// `dw - x` where the pixel map says `dw - 1 - x`, because a pixel centre at
/// `x + 0.5` has to land at `dw - x - 0.5`. Then the extremes, since a turn
/// exchanges which corner is which and a box written `(x0, y0, x1, y1)` has to
/// keep `x0 <= x1`.
fn lay_down_bbox(
bbox: (f32, f32, f32, f32),
dw: usize,
dh: usize,
orientation: Orientation,
) -> (f32, f32, f32, f32) {
let (sw, sh) = oriented(dw, dh, orientation);
let (dw, dh) = (dw as f32, dh as f32);
let (sw, sh) = (sw as f32, sh as f32);
let corner = |x: f32, y: f32| {
let (mut sx, mut sy) = match orientation.quarter_turns {
1 => (y, dw - x),
2 => (dw - x, dh - y),
3 => (dh - y, x),
_ => (x, y),
};
if orientation.flip_h {
sx = sw - sx;
}
if orientation.flip_v {
sy = sh - sy;
}
(sx, sy)
};
let (ax, ay) = corner(bbox.0, bbox.1);
let (bx, by) = corner(bbox.2, bbox.3);
(ax.min(bx), ay.min(by), ax.max(bx), ay.max(by))
}
/// The size those `width x height` pixels have once turned.
fn oriented(width: usize, height: usize, orientation: Orientation) -> (usize, usize) {
if orientation.swaps_axes() {
(height, width)
} else {
(width, height)
}
}
/// One of eight transforms, as bits a signature can carry.
fn orientation_key(orientation: Orientation) -> u64 {
u64::from(orientation.quarter_turns)
| (u64::from(orientation.flip_h) << 2)
| (u64::from(orientation.flip_v) << 3)
}
/// Load the model and run it.
///
/// Loading is ~24 ms against the ~470 ms of inference that follows, and this
/// runs once per image — so caching the session would keep 11 MB of weights
/// resident for the life of the app to save five percent of a background task.
fn detect(
rgb: &[f32],
width: usize,
height: usize,
fine: bool,
) -> Result<Vec<dr_segment::Instance>, String> {
let mut model = dr_segment::SemanticModel::embedded().map_err(|e| e.to_string())?;
let options = dr_segment::SemanticOptions {
// A quarter shared with each neighbour. It has to exceed zero at all,
// or a subject sitting on a seam is cut in half by both tiles and
// recognised by neither; a quarter is enough to carry a whole subject
// inside one tile at the sizes tiling is reached for.
tiling: if fine {
dr_segment::Tiling::Grid { overlap: 0.25 }
} else {
dr_segment::Tiling::Whole
},
..dr_segment::SemanticOptions::default()
};
model
.detect(rgb, width, height, &options)
.map_err(|e| e.to_string())
}
/// The model's soft coverage, to a byte per pixel.
///
/// Rounded rather than truncated, so a coverage of exactly 0.5 lands on the
/// threshold the selection tests against instead of one below it.
fn quantise(mask: &[f32]) -> Vec<u8> {
mask.iter()
.map(|&v| (v.clamp(0.0, 1.0) * 255.0).round() as u8)
.collect()
}
#[cfg(test)]
mod tests {
use super::*;
/// Two objects: a big weak one on the left, a small strong one that
/// overlaps it.
fn overlapping() -> Segmentation {
let (w, h) = (8usize, 4usize);
let mut big = vec![0u8; w * h];
let mut small = vec![0u8; w * h];
for y in 0..h {
for x in 0..6 {
big[y * w + x] = 255;
}
for x in 4..8 {
small[y * w + x] = 255;
}
}
Segmentation {
instances: vec![
InstanceSummary {
class_name: "bus".into(),
score: 0.5,
mask: big,
bbox: (0.0, 0.0, 6.0, h as f32),
},
InstanceSummary {
class_name: "person".into(),
score: 0.9,
mask: small,
bbox: (4.0, 0.0, 8.0, h as f32),
},
],
signature: 1,
proxy: (w, h),
}
}
#[test]
fn a_click_finds_the_object_under_it() {
let seg = overlapping();
assert_eq!(seg.instance_at(0.1, 0.5), Some(0), "only the bus here");
assert_eq!(seg.instance_at(0.95, 0.5), Some(1), "only the person here");
}
/// The overlap rule, and the one that decides what a click means where two
/// detections cover the same pixel.
#[test]
fn overlapping_objects_resolve_to_the_more_confident() {
let seg = overlapping();
assert_eq!(
seg.instance_at(0.6, 0.5),
Some(1),
"the person at 0.9 beats the bus at 0.5"
);
}
#[test]
fn a_click_outside_the_frame_selects_nothing() {
let seg = overlapping();
assert_eq!(seg.instance_at(-0.1, 0.5), None);
assert_eq!(seg.instance_at(1.5, 0.5), None);
assert_eq!(seg.instance_at(0.5, 1.0), None, "the far edge is exclusive");
}
#[test]
fn a_click_on_nothing_selects_nothing() {
let seg = Segmentation {
instances: Vec::new(),
signature: 1,
proxy: (4, 4),
};
assert_eq!(seg.instance_at(0.5, 0.5), None);
}
#[test]
fn colours_are_stable_and_distinct() {
assert_eq!(instance_colour(7), instance_colour(7));
assert_ne!(instance_colour(0), instance_colour(1));
assert_ne!(instance_colour(1), instance_colour(2));
}
#[test]
fn every_colour_is_visible_against_a_photograph() {
for i in 0..64u32 {
let [r, g, b] = instance_colour(i);
assert!(r.max(g).max(b) >= 200, "instance {i} is too dark");
}
}
/// The overlay must not cover the picture: that is the difference between
/// this and the region map it replaced.
#[test]
fn the_overlay_is_transparent_where_nothing_was_found() {
let seg = Segmentation {
instances: Vec::new(),
signature: 1,
proxy: (4, 4),
};
let (px, w, h) = seg.overlay_rgba();
assert_eq!((w, h), (4, 4));
assert!(px.chunks_exact(4).all(|p| p[3] == 0), "nothing to draw");
}
#[test]
fn the_overlay_outlines_what_it_fills() {
let seg = overlapping();
let (px, w, _) = seg.overlay_rgba();
let at = |x: usize, y: usize| {
let p = (y * w as usize + x) * 4;
[px[p], px[p + 1], px[p + 2], px[p + 3]]
};
// The rightmost column of the person is against the frame edge, so it
// is an outline pixel.
assert_eq!(at(7, 1), [255, 255, 255, 255]);
// And an interior pixel keeps its fill.
assert_ne!(at(2, 1)[3], 0, "the bus is filled");
assert_ne!(at(2, 1), [255, 255, 255, 255], "and not all outline");
}
#[test]
fn quantising_rounds_rather_than_truncates() {
// Exactly half must reach the threshold the selection tests against.
assert_eq!(quantise(&[0.5]), vec![128]);
assert_eq!(quantise(&[0.0, 1.0]), vec![0, 255]);
// And values outside the range cannot wrap.
assert_eq!(quantise(&[-1.0, 2.0]), vec![0, 255]);
}
/// A non-square, wholly asymmetric grid: every pixel is its own index, so
/// any permutation that is not the intended one shows up as a mismatch
/// rather than being hidden by a symmetry.
fn ramp(w: usize, h: usize) -> Vec<f32> {
(0..w * h).flat_map(|i| [i as f32, 0.0, 0.0]).collect()
}
fn red(rgb: &[f32]) -> Vec<f32> {
rgb.chunks_exact(3).map(|p| p[0]).collect()
}
/// The property the whole fix rests on: what the model is shown and what
/// comes back are the same permutation, run in opposite directions. If
/// they ever disagree, every subject mask lands somewhere other than its
/// subject — and looks like a mask while doing it.
#[test]
fn standing_a_frame_up_and_laying_it_down_is_the_identity() {
const W: usize = 5;
const H: usize = 3;
let source = ramp(W, H);
for tag in 1..=8u16 {
let o = Orientation::from_exif(tag);
let (up, uw, uh) = upright(&source, W, H, o);
let (ow, oh) = o.oriented_size(W as u32, H as u32);
assert_eq!(
(uw, uh),
(ow as usize, oh as usize),
"tag {tag}: the upright size is the oriented one"
);
let back = lay_down_mask(&red(&up), uw, uh, o);
assert_eq!(back, red(&source), "tag {tag} did not come back");
}
}
/// The colour channels must travel together. Reading a pixel three times
/// with one index arithmetic mistake gives a plausible image with its
/// channels sheared, which the model would still detect *something* in.
#[test]
fn a_turn_carries_whole_pixels() {
let rgb: Vec<f32> = (0..2 * 3)
.flat_map(|i| [i as f32, i as f32 + 100.0, i as f32 + 200.0])
.collect();
let o = Orientation::from_exif(6);
let (up, uw, uh) = upright(&rgb, 2, 3, o);
assert_eq!((uw, uh), (3, 2));
for p in up.chunks_exact(3) {
assert_eq!(p[1], p[0] + 100.0, "green left its pixel");
assert_eq!(p[2], p[0] + 200.0, "blue left its pixel");
}
}
/// A portrait frame is the case this exists for: the sensor is landscape,
/// the photograph is not, and the model has to be given the photograph.
#[test]
fn a_sideways_frame_reaches_the_model_upright() {
let o = Orientation::from_exif(6);
assert!(!o.is_normal());
let (_, uw, uh) = upright(&ramp(1600, 1066), 1600, 1066, o);
assert_eq!((uw, uh), (1066, 1600), "the model still got a landscape");
}
/// Where the model's box ends up, worked out by hand for the one turn a
/// portrait phone or a sideways body actually writes.
#[test]
fn a_box_comes_back_in_sensor_pixels() {
let o = Orientation::from_exif(6);
// Upright 4x6; the sensor it came from is 6x4.
let bbox = lay_down_bbox((0.0, 0.0, 2.0, 3.0), 4, 6, o);
assert_eq!(bbox, (0.0, 2.0, 3.0, 4.0));
}
/// Whatever a turn does to a box, it must still read low-to-high — a
/// permutation exchanges which corner is which, and the rest of the mask
/// pipeline measures `(x1 - x0)` without checking the sign.
#[test]
fn a_restored_box_keeps_its_corners_in_order() {
for tag in 1..=8u16 {
let o = Orientation::from_exif(tag);
let (x0, y0, x1, y1) = lay_down_bbox((1.0, 2.0, 7.0, 5.0), 9, 6, o);
assert!(x0 <= x1, "tag {tag}: x runs backwards");
assert!(y0 <= y1, "tag {tag}: y runs backwards");
// A permutation moves a box; it does not resize one.
let (sw, sh) = oriented(9, 6, o);
assert!(x1 <= sw as f32 && y1 <= sh as f32, "tag {tag}: box escaped");
assert!((((x1 - x0) * (y1 - y0)) - 18.0).abs() < 1e-3, "tag {tag}");
}
}
/// Turning the photograph changes what the model recognises, so the two
/// runs are different instance lists. If they could share a signature, a
/// layer built against one would be silently re-indexed into the other —
/// the same failure the tiling flag is in the signature to prevent.
#[test]
fn turning_the_photograph_changes_the_signature() {
let mut seen = std::collections::HashSet::new();
for tag in 1..=8u16 {
let o = Orientation::from_exif(tag);
assert!(
seen.insert(orientation_key(o)),
"tag {tag} shares a key with an earlier one"
);
}
assert_eq!(seen.len(), 8, "eight tags, but some collapsed");
}
#[test]
fn confidence_changes_the_signature() {
// A different threshold is a different instance list, so the indices a
// stored layer holds mean something else.
let a = segmentation_signature(100, 100, 3, Options::default().confidence.to_bits() as u64);
let b = segmentation_signature(100, 100, 3, 0.9f32.to_bits() as u64);
assert_ne!(a, b);
}
}