Show the detector the photograph, not the sensor's scanlines
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
Build and test / Desktop (Linux) (push) Successful in 19m7s
Build and test / Layer separation (push) Successful in 25s
Traceability / Requirement traces (push) Successful in 23s
Build and test / Android (aarch64) (push) Failing after 33m10s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
Build and test / Desktop (Linux) (push) Successful in 19m7s
Build and test / Layer separation (push) Successful in 25s
Traceability / Requirement traces (push) Successful in 23s
Build and test / Android (aarch64) (push) Failing after 33m10s
"Find subjects" was handed the proxy in the sensor's own orientation, so every frame shot on a body held sideways reached the model lying on its side — and a model trained on upright photographs is very bad at those. Measured end to end on a 22 MP frame of two people and a dog: `person 0.36` and nothing else, against `dog 0.82, person 0.61, person 0.49` for the same pixels stood up. Nothing failed; the panel simply offered one poor subject where there were three good ones. The orientation was never dropped on purpose. The proxy is deliberately rendered through a *neutral* graph — the detection has to survive an exposure change, or every slider would invalidate the masks built on it — and neutral took the file's orientation with it along with everything else. Landscape frames were unaffected, which is why it stood for as long as it did. The turn is `Orientation::source_pixel`, the same function the grid's thumbnails already go through, so the detector and the thumbnailer now agree about which way is up rather than holding two opinions. What it is turned by is `Framing::effective_orientation` — the file's EXIF tag and the photographer's own rotations composed into one permutation, by the group law rather than by adding the turns, which is a distinction `Framing` already had to make and had already tested. Rotating the picture and pressing the button again therefore does what it looks like it does. The proxy stays in sensor space and the masks come back into it. That is not a detail to be tidied later: the generated shader samples the mask array at `uv_src`, *after* the framing map, so a mask stored upright would sit a quarter turn off the subject it was drawn around. That is a wrong mask rather than a weak one, and nothing announces it. So the picture is stood up for the model and laid back down for everything else, and `upright`/`lay_down` are returned as a pair because calling one and forgetting the other is silent. Both directions are the one function: `upright` gathers through `source_pixel` and `lay_down` scatters through it. A quarter turn is a bijection of the pixel grid, so the round trip is exact — no filter, no resampling, and no hole to fill — and an inverse written out by hand would be a second thing to keep in step, whose way of being wrong is a mask mirrored about the wrong axis, which still looks like a mask. The orientation joins the confidence and the tiling flag in the segmentation signature, and for the same reason: turning the photograph changes what the model recognises, so two runs either side of a rotation are different instance lists. Two that happened to come out the same length would otherwise share a signature and a stored layer would be silently re-indexed from one into the other. The refine pass had it too — it re-runs the model over a crop rendered in the same sensor space — so it makes the same turn, and would otherwise have handed back a worse mask than the one it was asked to improve, on the subject the photographer had just pointed at. `dr-gpu`'s `local` example is fixed with it. It exists to be the shipping path with pictures attached, and a diagnostic that reproduces the bug it is meant to catch is a trap for whoever reads it next. Seven tests. The round trip is the identity over all eight EXIF tags on a non-square asymmetric grid; a turn carries whole pixels rather than shearing the channels apart; a sideways frame reaches the model upright; a box comes back in sensor pixels, worked out by hand for the one turn a portrait frame actually writes; a restored box still reads low-to-high for every tag, since the rest of the pipeline takes `x1 - x0` without checking the sign; and the eight tags cannot collapse into one signature key. The existing composition test now runs against `effective_orientation` itself, over all 8 x 16 baseline-and-user pairs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -359,6 +359,7 @@ impl Framing {
|
||||
self.baseline = orientation;
|
||||
}
|
||||
|
||||
/// TRACES: FR-DEV-3 | FR-DEV-3h
|
||||
/// The baseline and the user's turns and mirrors, collapsed into one.
|
||||
///
|
||||
/// Everything that renders or measures the frame goes through here; only
|
||||
@@ -370,7 +371,16 @@ impl Framing {
|
||||
/// axes when that turn is odd — which is exactly the case that a naive
|
||||
/// "add the turns, or the flags" gets wrong, and gets wrong silently,
|
||||
/// since the result is still a valid-looking orientation.
|
||||
fn effective(&self) -> (u8, bool, bool) {
|
||||
///
|
||||
/// Returned as an [`dr_types::Orientation`] because that is what it *is*
|
||||
/// — a quarter turn and two mirrors — and because saying so lets a
|
||||
/// caller outside the render reuse
|
||||
/// [`dr_types::Orientation::source_pixel`] rather than write the
|
||||
/// permutation out a second time. `dr-ui`'s segmentation is that caller:
|
||||
/// the detector has to read the photograph the way the photographer does,
|
||||
/// and the thumbnail path already turns its pixels with the same
|
||||
/// function.
|
||||
pub fn effective_orientation(&self) -> dr_types::Orientation {
|
||||
let b = self.baseline;
|
||||
// The user's mirrors, seen from the far side of the baseline's turn.
|
||||
let (ux, uy) = if b.swaps_axes() {
|
||||
@@ -378,11 +388,16 @@ impl Framing {
|
||||
} else {
|
||||
(self.flip_h, self.flip_v)
|
||||
};
|
||||
(
|
||||
(b.quarter_turns + self.quarter_turns) % 4,
|
||||
b.flip_h != ux,
|
||||
b.flip_v != uy,
|
||||
)
|
||||
dr_types::Orientation {
|
||||
quarter_turns: (b.quarter_turns + self.quarter_turns) % 4,
|
||||
flip_h: b.flip_h != ux,
|
||||
flip_v: b.flip_v != uy,
|
||||
}
|
||||
}
|
||||
|
||||
fn effective(&self) -> (u8, bool, bool) {
|
||||
let o = self.effective_orientation();
|
||||
(o.quarter_turns, o.flip_h, o.flip_v)
|
||||
}
|
||||
|
||||
/// Whether this stage currently changes the image.
|
||||
@@ -931,12 +946,10 @@ mod tests {
|
||||
f.set_param(FLIP_H, f32::from(u8::from(user_flip_h)));
|
||||
f.set_param(FLIP_V, f32::from(u8::from(user_flip_v)));
|
||||
|
||||
let (t, fh, fv) = f.effective();
|
||||
let combined = Orientation {
|
||||
quarter_turns: t,
|
||||
flip_h: fh,
|
||||
flip_v: fv,
|
||||
};
|
||||
// The public accessor, not the private tuple: this is
|
||||
// the permutation `dr-ui` turns the detector's input
|
||||
// by, so it is the one that has to be right.
|
||||
let combined = f.effective_orientation();
|
||||
|
||||
// The output size the composed transform produces must
|
||||
// be the one the two stages produce in sequence.
|
||||
|
||||
Reference in New Issue
Block a user