Show the detector the photograph, not the sensor's scanlines
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
Build and test / Desktop (Linux) (push) Successful in 19m7s
Build and test / Layer separation (push) Successful in 25s
Traceability / Requirement traces (push) Successful in 23s
Build and test / Android (aarch64) (push) Failing after 33m10s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
Build and test / Desktop (Linux) (push) Successful in 19m7s
Build and test / Layer separation (push) Successful in 25s
Traceability / Requirement traces (push) Successful in 23s
Build and test / Android (aarch64) (push) Failing after 33m10s
"Find subjects" was handed the proxy in the sensor's own orientation, so every frame shot on a body held sideways reached the model lying on its side — and a model trained on upright photographs is very bad at those. Measured end to end on a 22 MP frame of two people and a dog: `person 0.36` and nothing else, against `dog 0.82, person 0.61, person 0.49` for the same pixels stood up. Nothing failed; the panel simply offered one poor subject where there were three good ones. The orientation was never dropped on purpose. The proxy is deliberately rendered through a *neutral* graph — the detection has to survive an exposure change, or every slider would invalidate the masks built on it — and neutral took the file's orientation with it along with everything else. Landscape frames were unaffected, which is why it stood for as long as it did. The turn is `Orientation::source_pixel`, the same function the grid's thumbnails already go through, so the detector and the thumbnailer now agree about which way is up rather than holding two opinions. What it is turned by is `Framing::effective_orientation` — the file's EXIF tag and the photographer's own rotations composed into one permutation, by the group law rather than by adding the turns, which is a distinction `Framing` already had to make and had already tested. Rotating the picture and pressing the button again therefore does what it looks like it does. The proxy stays in sensor space and the masks come back into it. That is not a detail to be tidied later: the generated shader samples the mask array at `uv_src`, *after* the framing map, so a mask stored upright would sit a quarter turn off the subject it was drawn around. That is a wrong mask rather than a weak one, and nothing announces it. So the picture is stood up for the model and laid back down for everything else, and `upright`/`lay_down` are returned as a pair because calling one and forgetting the other is silent. Both directions are the one function: `upright` gathers through `source_pixel` and `lay_down` scatters through it. A quarter turn is a bijection of the pixel grid, so the round trip is exact — no filter, no resampling, and no hole to fill — and an inverse written out by hand would be a second thing to keep in step, whose way of being wrong is a mask mirrored about the wrong axis, which still looks like a mask. The orientation joins the confidence and the tiling flag in the segmentation signature, and for the same reason: turning the photograph changes what the model recognises, so two runs either side of a rotation are different instance lists. Two that happened to come out the same length would otherwise share a signature and a stored layer would be silently re-indexed from one into the other. The refine pass had it too — it re-runs the model over a crop rendered in the same sensor space — so it makes the same turn, and would otherwise have handed back a worse mask than the one it was asked to improve, on the subject the photographer had just pointed at. `dr-gpu`'s `local` example is fixed with it. It exists to be the shipping path with pictures attached, and a diagnostic that reproduces the bug it is meant to catch is a trap for whoever reads it next. Seven tests. The round trip is the identity over all eight EXIF tags on a non-square asymmetric grid; a turn carries whole pixels rather than shearing the channels apart; a sideways frame reaches the model upright; a box comes back in sensor pixels, worked out by hand for the one turn a portrait frame actually writes; a restored box still reads low-to-high for every tag, since the rest of the pipeline takes `x1 - x0` without checking the sign; and the eight tags cannot collapse into one signature key. The existing composition test now runs against `effective_orientation` itself, over all 8 x 16 baseline-and-user pairs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -56,7 +56,13 @@ fn main() {
|
||||
// ---- the photograph ---------------------------------------------------
|
||||
let bytes = std::fs::read(&path).expect("read file");
|
||||
let raw = dr_decode::decode(&bytes).expect("decode");
|
||||
// The tag, because the model reads photographs and the sensor stores
|
||||
// scanlines. See `stand_up` below: this example exists to be the shipping
|
||||
// path with pictures attached, so it has to make the same turn the
|
||||
// develop session makes.
|
||||
let orientation = dr_decode::orientation(&bytes).unwrap_or_default();
|
||||
println!("source {} × {}", raw.crop.width, raw.crop.height);
|
||||
println!("turns {}", orientation.quarter_turns);
|
||||
let source = Demosaicer::new(&ctx)
|
||||
.expect("demosaicer")
|
||||
.run(&raw)
|
||||
@@ -94,10 +100,16 @@ fn main() {
|
||||
.collect();
|
||||
|
||||
// ---- find the subject -------------------------------------------------
|
||||
//
|
||||
// Stood up first. A model trained on upright photographs is very bad at
|
||||
// sideways ones, and the proxy above is in the sensor's own orientation
|
||||
// — see `stand_up`.
|
||||
let (upright, uw, uh) = stand_up(&rgb, pw as usize, ph as usize, orientation);
|
||||
|
||||
let t = std::time::Instant::now();
|
||||
let mut model = SemanticModel::embedded().expect("model");
|
||||
let instances = model
|
||||
.detect(&rgb, pw as usize, ph as usize, &SemanticOptions::default())
|
||||
.detect(&upright, uw, uh, &SemanticOptions::default())
|
||||
.expect("detect");
|
||||
println!(
|
||||
"detect {} found in {:.0} ms",
|
||||
@@ -119,10 +131,11 @@ fn main() {
|
||||
subject.class_name, subject.score
|
||||
);
|
||||
|
||||
// Quantised exactly as the develop session does, so this example exercises
|
||||
// the shipping path rather than a shortcut around it.
|
||||
let alpha: Vec<u8> = subject
|
||||
.mask
|
||||
// Laid back down, then quantised exactly as the develop session does, so
|
||||
// this example exercises the shipping path rather than a shortcut around
|
||||
// it. The mask has to end up in *source* space: the composed shader
|
||||
// samples the mask array after the framing map.
|
||||
let alpha: Vec<u8> = lay_down(&subject.mask, uw, uh, orientation)
|
||||
.iter()
|
||||
.map(|&v| (v.clamp(0.0, 1.0) * 255.0).round() as u8)
|
||||
.collect();
|
||||
@@ -328,6 +341,56 @@ fn render_stack(
|
||||
.expect("render");
|
||||
}
|
||||
|
||||
/// Turn the proxy the way the photographer is looking at it, so the model
|
||||
/// reads a photograph rather than a scanline order.
|
||||
///
|
||||
/// `Orientation::source_pixel` is the permutation, and it is the same function
|
||||
/// the grid's thumbnails go through — the point being that the detector and
|
||||
/// the thumbnailer agree about which way is up. A quarter turn is a bijection
|
||||
/// of the pixel grid, so nothing is resampled in either direction.
|
||||
fn stand_up(
|
||||
rgb: &[f32],
|
||||
width: usize,
|
||||
height: usize,
|
||||
o: dr_types::Orientation,
|
||||
) -> (Vec<f32>, usize, usize) {
|
||||
if o.is_normal() {
|
||||
return (rgb.to_vec(), width, height);
|
||||
}
|
||||
let (dw, dh) = o.oriented_size(width as u32, height as u32);
|
||||
let (dw, dh) = (dw as usize, dh as usize);
|
||||
|
||||
let mut out = vec![0.0f32; dw * dh * 3];
|
||||
for y in 0..dh {
|
||||
for x in 0..dw {
|
||||
let (sx, sy) = o.source_pixel(x as u32, y as u32, dw as u32, dh as u32);
|
||||
let s = (sy as usize * width + sx as usize) * 3;
|
||||
let d = (y * dw + x) * 3;
|
||||
out[d..d + 3].copy_from_slice(&rgb[s..s + 3]);
|
||||
}
|
||||
}
|
||||
(out, dw, dh)
|
||||
}
|
||||
|
||||
/// [`stand_up`] run backwards: the same map read as a scatter, which fills
|
||||
/// every source pixel exactly once because the map is a bijection.
|
||||
fn lay_down(mask: &[f32], dw: usize, dh: usize, o: dr_types::Orientation) -> Vec<f32> {
|
||||
if o.is_normal() {
|
||||
return mask.to_vec();
|
||||
}
|
||||
let (sw, sh) = o.oriented_size(dw as u32, dh as u32);
|
||||
let (sw, sh) = (sw as usize, sh as usize);
|
||||
|
||||
let mut out = vec![0.0f32; sw * sh];
|
||||
for y in 0..dh {
|
||||
for x in 0..dw {
|
||||
let (sx, sy) = o.source_pixel(x as u32, y as u32, dw as u32, dh as u32);
|
||||
out[sy as usize * sw + sx as usize] = mask[y * dw + x];
|
||||
}
|
||||
}
|
||||
out
|
||||
}
|
||||
|
||||
fn fit(w: u32, h: u32, longest: u32) -> (u32, u32) {
|
||||
let s = (longest as f32 / w.max(h) as f32).min(1.0);
|
||||
(
|
||||
|
||||
Reference in New Issue
Block a user