Show the detector the photograph, not the sensor's scanlines
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
Build and test / Desktop (Linux) (push) Successful in 19m7s
Build and test / Layer separation (push) Successful in 25s
Traceability / Requirement traces (push) Successful in 23s
Build and test / Android (aarch64) (push) Failing after 33m10s

"Find subjects" was handed the proxy in the sensor's own orientation, so
every frame shot on a body held sideways reached the model lying on its
side — and a model trained on upright photographs is very bad at those.
Measured end to end on a 22 MP frame of two people and a dog: `person
0.36` and nothing else, against `dog 0.82, person 0.61, person 0.49` for
the same pixels stood up. Nothing failed; the panel simply offered one
poor subject where there were three good ones.

The orientation was never dropped on purpose. The proxy is deliberately
rendered through a *neutral* graph — the detection has to survive an
exposure change, or every slider would invalidate the masks built on it
— and neutral took the file's orientation with it along with everything
else. Landscape frames were unaffected, which is why it stood for as
long as it did.

The turn is `Orientation::source_pixel`, the same function the grid's
thumbnails already go through, so the detector and the thumbnailer now
agree about which way is up rather than holding two opinions. What it is
turned by is `Framing::effective_orientation` — the file's EXIF tag and
the photographer's own rotations composed into one permutation, by the
group law rather than by adding the turns, which is a distinction
`Framing` already had to make and had already tested. Rotating the
picture and pressing the button again therefore does what it looks like
it does.

The proxy stays in sensor space and the masks come back into it. That is
not a detail to be tidied later: the generated shader samples the mask
array at `uv_src`, *after* the framing map, so a mask stored upright
would sit a quarter turn off the subject it was drawn around. That is a
wrong mask rather than a weak one, and nothing announces it. So the
picture is stood up for the model and laid back down for everything
else, and `upright`/`lay_down` are returned as a pair because calling
one and forgetting the other is silent.

Both directions are the one function: `upright` gathers through
`source_pixel` and `lay_down` scatters through it. A quarter turn is a
bijection of the pixel grid, so the round trip is exact — no filter, no
resampling, and no hole to fill — and an inverse written out by hand
would be a second thing to keep in step, whose way of being wrong is a
mask mirrored about the wrong axis, which still looks like a mask.

The orientation joins the confidence and the tiling flag in the
segmentation signature, and for the same reason: turning the photograph
changes what the model recognises, so two runs either side of a rotation
are different instance lists. Two that happened to come out the same
length would otherwise share a signature and a stored layer would be
silently re-indexed from one into the other.

The refine pass had it too — it re-runs the model over a crop rendered
in the same sensor space — so it makes the same turn, and would
otherwise have handed back a worse mask than the one it was asked to
improve, on the subject the photographer had just pointed at.

`dr-gpu`'s `local` example is fixed with it. It exists to be the
shipping path with pictures attached, and a diagnostic that reproduces
the bug it is meant to catch is a trap for whoever reads it next.

Seven tests. The round trip is the identity over all eight EXIF tags on
a non-square asymmetric grid; a turn carries whole pixels rather than
shearing the channels apart; a sideways frame reaches the model
upright; a box comes back in sensor pixels, worked out by hand for the
one turn a portrait frame actually writes; a restored box still reads
low-to-high for every tag, since the rest of the pipeline takes
`x1 - x0` without checking the sign; and the eight tags cannot collapse
into one signature key. The existing composition test now runs against
`effective_orientation` itself, over all 8 x 16 baseline-and-user pairs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-26 22:40:58 +02:00
co-authored by Claude Opus 5
parent 336a1fd296
commit 4a82753d22
5 changed files with 449 additions and 51 deletions
+291 -11
View File
@@ -34,6 +34,7 @@ use std::sync::Arc;
use dr_gpu::GpuContext;
use dr_pipeline::mask::segmentation_signature;
use dr_types::Orientation;
/// One recognised object.
#[derive(Debug, Clone)]
@@ -243,27 +244,41 @@ impl Default for Options {
/// Find what can be selected in this photograph.
///
/// `rgb` is the proxy the model reads — tightly packed RGB floats at
/// `(width, height)`. Passed in rather than derived here because the caller
/// already has the rendered proxy, and re-deriving it would mean a second
/// readback of something the CPU is holding.
/// `rgb` is the rendered proxy — tightly packed RGB floats at
/// `(width, height)`, in the **sensor's** own orientation. Passed in rather
/// than derived here because the caller already has it, and re-deriving it
/// would mean a second readback of something the CPU is holding.
///
/// `orientation` is the file's EXIF tag composed with whatever turns the
/// photographer has since applied — `Framing::effective_orientation`, one
/// permutation covering both. The model is shown the picture through it and
/// its answers come back without it, so everything this returns is in sensor
/// space exactly as it was before the detector was taught to read.
pub fn compute(
_ctx: &GpuContext,
rgb: &[f32],
width: usize,
height: usize,
orientation: Orientation,
options: &Options,
) -> Result<Segmentation, String> {
let found = detect(rgb, width, height, options.fine)?;
// The model reads the photograph; everything else here speaks sensor.
let (stood_up, uw, uh) = upright(rgb, width, height, orientation);
let found = detect(&stood_up, uw, uh, options.fine)?;
let instances: Vec<InstanceSummary> = found
.iter()
.filter(|i| i.score >= options.confidence)
.map(|i| InstanceSummary {
class_name: i.class_name.clone(),
score: i.score,
mask: quantise(&i.mask),
bbox: i.bbox,
.map(|i| {
let (mask, bbox) = lay_down(&i.mask, i.bbox, uw, uh, orientation);
InstanceSummary {
class_name: i.class_name.clone(),
score: i.score,
// Quantised after the permutation, so the byte stored is a
// rounding of the model's own coverage and not of a copy.
mask: quantise(&mask),
bbox,
}
})
.collect();
@@ -278,21 +293,177 @@ pub fn compute(
// a signature, a layer built against the coarse pass would be silently
// reinterpreted against the fine one. That is a *wrong* mask, which is
// far worse than a stale one, because nothing announces it.
//
// The orientation is in for the same reason and it is not hypothetical:
// turning the photograph changes what the model recognises, so a run
// before a quarter turn and a run after it are different instance
// lists. Two lists that happened to come out the same length would
// otherwise share a signature, and a layer built against the first
// would be silently re-indexed into the second.
options.confidence.to_bits() as u64
^ if options.fine {
0x9E37_79B9_7F4A_7C15
} else {
0
},
}
^ orientation_key(orientation),
);
Ok(Segmentation {
instances,
signature,
// **Sensor space, not the model's.** `lay_down` put every mask back,
// so the grid a stored layer indexes into is the one it always was —
// see `upright` for why the model saw a different one.
proxy: (width, height),
})
}
/// TRACES: FR-DEV-3 | FR-DEV-3h
/// Turn the proxy the way the photographer is looking at it.
///
/// **Why this exists at all.** A camera held sideways writes its sensor rows
/// the way it always does, and the render puts them right by way of
/// `Framing`. The proxy the model reads is deliberately rendered through a
/// *neutral* graph — the detection has to survive an exposure change, or
/// every slider would invalidate the masks built on it — and neutral took the
/// orientation with it. So the detector was handed a portrait frame lying on
/// its side, and a model trained on upright photographs is very bad at those.
/// Measured end to end on one 22 MP frame of two people and a dog: `person
/// 0.36` and nothing else, against `dog 0.82, person 0.61, person 0.49` for
/// the same pixels stood up.
///
/// The turn is [`Orientation::source_pixel`], which is the function the grid's
/// thumbnails already go through (`dr_decode::Preview::apply_orientation`) —
/// so the detector now reads exactly the kind of image the thumbnailer makes,
/// rather than a second opinion about what "upright" means.
///
/// A quarter turn and its mirrors are a permutation of the pixel grid, so this
/// is exact: no filter, no resampling, and no edge softened on the way in that
/// would have to be judged on the way out.
pub(crate) fn upright(
rgb: &[f32],
width: usize,
height: usize,
orientation: Orientation,
) -> (Vec<f32>, usize, usize) {
if orientation.is_normal() || width == 0 || height == 0 {
return (rgb.to_vec(), width, height);
}
let (dw, dh) = oriented(width, height, orientation);
let mut out = vec![0.0f32; dw * dh * 3];
for y in 0..dh {
for x in 0..dw {
let (sx, sy) = orientation.source_pixel(x as u32, y as u32, dw as u32, dh as u32);
let s = (sy as usize * width + sx as usize) * 3;
let d = (y * dw + x) * 3;
out[d..d + 3].copy_from_slice(&rgb[s..s + 3]);
}
}
(out, dw, dh)
}
/// TRACES: FR-DEV-3
/// Put what the model answered back onto the sensor's grid.
///
/// The counterpart of [`upright`], and the two are always used as a pair: a
/// mask is only ever in the model's frame between those two calls. Returned
/// together rather than as two functions a caller composes, because calling
/// one and forgetting the other is silent — the mask lands a quarter turn off
/// the subject, which reads as a bad detection rather than as a bug.
///
/// `dw`/`dh` are the *upright* dimensions, as [`upright`] returned them.
pub(crate) fn lay_down(
mask: &[f32],
bbox: (f32, f32, f32, f32),
dw: usize,
dh: usize,
orientation: Orientation,
) -> (Vec<f32>, (f32, f32, f32, f32)) {
if orientation.is_normal() {
return (mask.to_vec(), bbox);
}
(
lay_down_mask(mask, dw, dh, orientation),
lay_down_bbox(bbox, dw, dh, orientation),
)
}
/// The same permutation as [`upright`], read as a scatter rather than a
/// gather: a quarter turn is a bijection of the grid, so writing every upright
/// pixel to where it came from fills the sensor-space mask exactly once and
/// leaves no hole. Running the one function in the one direction is the point
/// — an inverse written out by hand is a second thing to keep in step, and its
/// way of being wrong is a mask mirrored about the wrong axis, which still
/// looks like a mask.
fn lay_down_mask(mask: &[f32], dw: usize, dh: usize, orientation: Orientation) -> Vec<f32> {
let (sw, sh) = oriented(dw, dh, orientation);
let mut out = vec![0.0f32; sw * sh];
for y in 0..dh {
for x in 0..dw {
let (sx, sy) = orientation.source_pixel(x as u32, y as u32, dw as u32, dh as u32);
out[sy as usize * sw + sx as usize] = mask[y * dw + x];
}
}
out
}
/// [`lay_down_mask`] for a box.
///
/// The corners go through the same permutation in *continuous* coordinates —
/// `dw - x` where the pixel map says `dw - 1 - x`, because a pixel centre at
/// `x + 0.5` has to land at `dw - x - 0.5`. Then the extremes, since a turn
/// exchanges which corner is which and a box written `(x0, y0, x1, y1)` has to
/// keep `x0 <= x1`.
fn lay_down_bbox(
bbox: (f32, f32, f32, f32),
dw: usize,
dh: usize,
orientation: Orientation,
) -> (f32, f32, f32, f32) {
let (sw, sh) = oriented(dw, dh, orientation);
let (dw, dh) = (dw as f32, dh as f32);
let (sw, sh) = (sw as f32, sh as f32);
let corner = |x: f32, y: f32| {
let (mut sx, mut sy) = match orientation.quarter_turns {
1 => (y, dw - x),
2 => (dw - x, dh - y),
3 => (dh - y, x),
_ => (x, y),
};
if orientation.flip_h {
sx = sw - sx;
}
if orientation.flip_v {
sy = sh - sy;
}
(sx, sy)
};
let (ax, ay) = corner(bbox.0, bbox.1);
let (bx, by) = corner(bbox.2, bbox.3);
(ax.min(bx), ay.min(by), ax.max(bx), ay.max(by))
}
/// The size those `width x height` pixels have once turned.
fn oriented(width: usize, height: usize, orientation: Orientation) -> (usize, usize) {
if orientation.swaps_axes() {
(height, width)
} else {
(width, height)
}
}
/// One of eight transforms, as bits a signature can carry.
fn orientation_key(orientation: Orientation) -> u64 {
u64::from(orientation.quarter_turns)
| (u64::from(orientation.flip_h) << 2)
| (u64::from(orientation.flip_v) << 3)
}
/// Load the model and run it.
///
/// Loading is ~24 ms against the ~470 ms of inference that follows, and this
@@ -462,6 +633,115 @@ mod tests {
assert_eq!(quantise(&[-1.0, 2.0]), vec![0, 255]);
}
/// A non-square, wholly asymmetric grid: every pixel is its own index, so
/// any permutation that is not the intended one shows up as a mismatch
/// rather than being hidden by a symmetry.
fn ramp(w: usize, h: usize) -> Vec<f32> {
(0..w * h).flat_map(|i| [i as f32, 0.0, 0.0]).collect()
}
fn red(rgb: &[f32]) -> Vec<f32> {
rgb.chunks_exact(3).map(|p| p[0]).collect()
}
/// The property the whole fix rests on: what the model is shown and what
/// comes back are the same permutation, run in opposite directions. If
/// they ever disagree, every subject mask lands somewhere other than its
/// subject — and looks like a mask while doing it.
#[test]
fn standing_a_frame_up_and_laying_it_down_is_the_identity() {
const W: usize = 5;
const H: usize = 3;
let source = ramp(W, H);
for tag in 1..=8u16 {
let o = Orientation::from_exif(tag);
let (up, uw, uh) = upright(&source, W, H, o);
let (ow, oh) = o.oriented_size(W as u32, H as u32);
assert_eq!(
(uw, uh),
(ow as usize, oh as usize),
"tag {tag}: the upright size is the oriented one"
);
let back = lay_down_mask(&red(&up), uw, uh, o);
assert_eq!(back, red(&source), "tag {tag} did not come back");
}
}
/// The colour channels must travel together. Reading a pixel three times
/// with one index arithmetic mistake gives a plausible image with its
/// channels sheared, which the model would still detect *something* in.
#[test]
fn a_turn_carries_whole_pixels() {
let rgb: Vec<f32> = (0..2 * 3)
.flat_map(|i| [i as f32, i as f32 + 100.0, i as f32 + 200.0])
.collect();
let o = Orientation::from_exif(6);
let (up, uw, uh) = upright(&rgb, 2, 3, o);
assert_eq!((uw, uh), (3, 2));
for p in up.chunks_exact(3) {
assert_eq!(p[1], p[0] + 100.0, "green left its pixel");
assert_eq!(p[2], p[0] + 200.0, "blue left its pixel");
}
}
/// A portrait frame is the case this exists for: the sensor is landscape,
/// the photograph is not, and the model has to be given the photograph.
#[test]
fn a_sideways_frame_reaches_the_model_upright() {
let o = Orientation::from_exif(6);
assert!(!o.is_normal());
let (_, uw, uh) = upright(&ramp(1600, 1066), 1600, 1066, o);
assert_eq!((uw, uh), (1066, 1600), "the model still got a landscape");
}
/// Where the model's box ends up, worked out by hand for the one turn a
/// portrait phone or a sideways body actually writes.
#[test]
fn a_box_comes_back_in_sensor_pixels() {
let o = Orientation::from_exif(6);
// Upright 4x6; the sensor it came from is 6x4.
let bbox = lay_down_bbox((0.0, 0.0, 2.0, 3.0), 4, 6, o);
assert_eq!(bbox, (0.0, 2.0, 3.0, 4.0));
}
/// Whatever a turn does to a box, it must still read low-to-high — a
/// permutation exchanges which corner is which, and the rest of the mask
/// pipeline measures `(x1 - x0)` without checking the sign.
#[test]
fn a_restored_box_keeps_its_corners_in_order() {
for tag in 1..=8u16 {
let o = Orientation::from_exif(tag);
let (x0, y0, x1, y1) = lay_down_bbox((1.0, 2.0, 7.0, 5.0), 9, 6, o);
assert!(x0 <= x1, "tag {tag}: x runs backwards");
assert!(y0 <= y1, "tag {tag}: y runs backwards");
// A permutation moves a box; it does not resize one.
let (sw, sh) = oriented(9, 6, o);
assert!(x1 <= sw as f32 && y1 <= sh as f32, "tag {tag}: box escaped");
assert!((((x1 - x0) * (y1 - y0)) - 18.0).abs() < 1e-3, "tag {tag}");
}
}
/// Turning the photograph changes what the model recognises, so the two
/// runs are different instance lists. If they could share a signature, a
/// layer built against one would be silently re-indexed into the other —
/// the same failure the tiling flag is in the signature to prevent.
#[test]
fn turning_the_photograph_changes_the_signature() {
let mut seen = std::collections::HashSet::new();
for tag in 1..=8u16 {
let o = Orientation::from_exif(tag);
assert!(
seen.insert(orientation_key(o)),
"tag {tag} shares a key with an earlier one"
);
}
assert_eq!(seen.len(), 8, "eight tags, but some collapsed");
}
#[test]
fn confidence_changes_the_signature() {
// A different threshold is a different instance list, so the indices a