docs/ had 26 developer documents flat beside the manual, and the two audiences are very differently sized: most readers want the manual and the gesture reference, a few want the register, the designs and the measurements. The manual and gestures.md stay at the top; everything for someone changing the code moves to docs/dev/, and the two documents that name their own successors — the v0.1 milestone and the UI-refinement plan — go to docs/dev/archive/ rather than being deleted, since both are still cited. docs/README.md is the index, users first. Every reference follows: code comments, Cargo manifests, the workflows, the pre-commit hook, the bench and traceability tools (which locate the repo root by docs/dev/requirements.md now), packaging, the Docker READMEs, CLAUDE.md, CONTRIBUTING.md and the README. The matrix links one level deeper and is regenerated. Links out of the moved documents into the tree gain a level; a link checker over every Markdown file finds none broken.
1061 lines
42 KiB
Rust
1061 lines
42 KiB
Rust
//! TRACES: FR-DEV-3i
|
|
//! Finding what can be selected in the photograph on screen.
|
|
//!
|
|
//! One model pass, and what it recognised. That is the whole subsystem now.
|
|
//!
|
|
//! # What used to be here
|
|
//!
|
|
//! A watershed over-segmented the image, a merge tree turned that into a
|
|
//! granularity ladder, and a click walked up it (docs/dev/segmentation.md, arms A
|
|
//! and C). It is gone from this path, and the reason is measured rather than
|
|
//! aesthetic: on a real photograph the saddles are near zero almost
|
|
//! everywhere, so the merge order joins everything meaningful before it joins
|
|
//! anything spurious. Cutting a 45,808-basin field of a 22 MP frame to 2,000
|
|
//! regions left **one** region covering nearly the whole picture plus specks
|
|
//! (§15). A ladder that collapses is not a ladder.
|
|
//!
|
|
//! The passes and the hierarchy still exist in `dr-gpu` and `dr-segment`,
|
|
//! tested and documented, because it is the *merge criterion* that fails and
|
|
//! that is one function. What does not exist any more is a product path
|
|
//! through them, or a control offering a choice that does nothing.
|
|
//!
|
|
//! # This is a precompute, and it is slow
|
|
//!
|
|
//! ~700 ms on a 22 MP frame: a proxy render, a readback, and the model. It
|
|
//! runs **once per image, when the user asks**, and never on the frame path.
|
|
//! Every interaction it enables — click a subject, grow a mask, change a
|
|
//! falloff — reads its cached output.
|
|
//!
|
|
//! All of it is on a worker. [`compute`] takes an owned buffer and a
|
|
//! `GpuContext`, which is what makes that possible — nothing here touches the
|
|
//! develop session, and [`crate::develop::SegmentationJob`] is the piece that
|
|
//! carries the proxy render across with it.
|
|
|
|
use std::borrow::Cow;
|
|
use std::sync::Arc;
|
|
|
|
use dr_gpu::GpuContext;
|
|
use dr_pipeline::mask::segmentation_signature;
|
|
use dr_types::Orientation;
|
|
|
|
/// One photographic category, over the whole frame.
|
|
///
|
|
/// The counterpart to [`InstanceSummary`] and deliberately thinner: a category
|
|
/// has no box, because it is not one object in one place — sky is wherever the
|
|
/// sky is, in as many disconnected pieces as the frame has windows.
|
|
#[derive(Debug, Clone)]
|
|
pub struct CategorySummary {
|
|
pub name: std::sync::Arc<str>,
|
|
/// Fraction of the frame this category covers, for ordering the list and
|
|
/// for hiding a category that would give the user a control that does
|
|
/// nothing.
|
|
pub coverage: f32,
|
|
/// Coverage at proxy resolution, quantised to a byte — the same
|
|
/// representation, and for the same reasons, as `InstanceSummary::mask`.
|
|
///
|
|
/// **The model's own weighting, before any refinement.** Storing the
|
|
/// coarse mask rather than a refined one is what lets the refine control
|
|
/// have an off position that is exactly the old behaviour, and what stops
|
|
/// a strictness change from needing the model run again.
|
|
pub mask: Vec<u8>,
|
|
/// The evidence for cutting this category back to the pixels whose colour
|
|
/// agrees with it, when there was enough of the frame to gather any.
|
|
///
|
|
/// `None` is ordinary: a category thinner everywhere than the model's own
|
|
/// resolution, or one filling the whole frame, gives nothing to contrast
|
|
/// against — `dr_segment::refine` says which. A layer on such a category
|
|
/// simply has no refine control, which is the honest presentation of
|
|
/// "there was nothing to judge it with".
|
|
pub refinement: Option<dr_segment::Refinement>,
|
|
}
|
|
|
|
/// One recognised object.
|
|
#[derive(Debug, Clone)]
|
|
pub struct InstanceSummary {
|
|
pub class_name: Arc<str>,
|
|
pub score: f32,
|
|
/// Coverage at proxy resolution, quantised to a byte.
|
|
///
|
|
/// A byte rather than the `f32` the model produces: 256 levels is far
|
|
/// finer than an edge anyone can see, and at four bytes a pixel a handful
|
|
/// of objects would be most of a hundred megabytes for one photograph.
|
|
///
|
|
/// This is the *source* a distance field is built from, not the mask
|
|
/// itself — `dr_segment::Shaped` turns it into one.
|
|
pub mask: Vec<u8>,
|
|
/// `(x0, y0, x1, y1)` in [`Segmentation::proxy_size`] pixels.
|
|
///
|
|
/// Carried through from `dr_segment::Instance` rather than re-derived
|
|
/// from the mask, so a refine pass knows what region to crop without
|
|
/// scanning a proxy-sized buffer for its own extent.
|
|
pub bbox: (f32, f32, f32, f32),
|
|
}
|
|
|
|
/// One image's recognised objects, ready to mask.
|
|
pub struct Segmentation {
|
|
instances: Vec<InstanceSummary>,
|
|
/// What the scene model made of the same frame, empty when no scene model
|
|
/// could be found.
|
|
///
|
|
/// Empty is an ordinary state, not a failure: a build without the weights
|
|
/// compiled in and without them installed simply offers no categories, the
|
|
/// same way a missing face model turns face indexing off rather than
|
|
/// stopping the app.
|
|
categories: Vec<CategorySummary>,
|
|
/// Identifies this run, so a stored layer can tell whether the index it
|
|
/// holds still means what it meant.
|
|
signature: u64,
|
|
/// The space instance masks are defined in, in **source** proxy pixels.
|
|
///
|
|
/// Everything a mask is built from lives here, which is what lets the
|
|
/// render sample it *after* the framing map rather than before — so a mask
|
|
/// stays on the photograph through a zoom, a pan and a crop.
|
|
proxy: (usize, usize),
|
|
}
|
|
|
|
impl Segmentation {
|
|
/// TRACES: FR-DEV-3
|
|
/// A segmentation assembled by hand, for tests that need one without a
|
|
/// model.
|
|
///
|
|
/// Test-only, and it exists because the alternative is worse: the tests
|
|
/// that need coverage to be *a particular shape* — that a stored raster
|
|
/// renders the same pixels as the model's own did — cannot ask a real run
|
|
/// for a shape, and running the model to find out what it happened to
|
|
/// detect in a synthetic frame makes the assertion depend on the weights.
|
|
/// Every test that is genuinely about the model still runs it.
|
|
#[cfg(test)]
|
|
pub(crate) fn for_test(
|
|
instances: Vec<InstanceSummary>,
|
|
categories: Vec<CategorySummary>,
|
|
signature: u64,
|
|
proxy: (usize, usize),
|
|
) -> Self {
|
|
Self {
|
|
instances,
|
|
categories,
|
|
signature,
|
|
proxy,
|
|
}
|
|
}
|
|
|
|
pub fn signature(&self) -> u64 {
|
|
self.signature
|
|
}
|
|
|
|
pub fn proxy_size(&self) -> (usize, usize) {
|
|
self.proxy
|
|
}
|
|
|
|
pub fn instances(&self) -> &[InstanceSummary] {
|
|
&self.instances
|
|
}
|
|
|
|
/// One instance's coverage, at [`Self::proxy_size`].
|
|
pub fn instance_mask(&self, index: usize) -> Option<&[u8]> {
|
|
self.instances.get(index).map(|i| i.mask.as_slice())
|
|
}
|
|
|
|
pub fn categories(&self) -> &[CategorySummary] {
|
|
&self.categories
|
|
}
|
|
|
|
/// One category's coverage, at [`Self::proxy_size`].
|
|
///
|
|
/// By name, matching `MaskSource::Category`. A linear scan because there
|
|
/// are eight of them and a map would be more machinery than lookup.
|
|
pub fn category_mask(&self, name: &str) -> Option<&[u8]> {
|
|
self.categories
|
|
.iter()
|
|
.find(|c| &*c.name == name)
|
|
.map(|c| c.mask.as_slice())
|
|
}
|
|
|
|
/// One category's coverage, cut back at the given strictness.
|
|
///
|
|
/// What the rasteriser reads. Borrowed at
|
|
/// [`dr_segment::STRICTNESS_OFF`] and wherever no refinement could be
|
|
/// fitted, so the common case — a layer whose control has not been moved —
|
|
/// costs nothing at all over [`Self::category_mask`].
|
|
///
|
|
/// The strictness is the layer's, not the segmentation's: two layers may
|
|
/// sit on the same category and want different amounts of it, and the
|
|
/// model ran once for both.
|
|
pub fn category_mask_at(&self, name: &str, strictness: f32) -> Option<Cow<'_, [u8]>> {
|
|
let category = self.categories.iter().find(|c| &*c.name == name)?;
|
|
match &category.refinement {
|
|
Some(refinement) if strictness > dr_segment::STRICTNESS_OFF => Some(Cow::Owned(
|
|
refinement.apply_coverage(&category.mask, strictness),
|
|
)),
|
|
_ => Some(Cow::Borrowed(category.mask.as_slice())),
|
|
}
|
|
}
|
|
|
|
/// Whether this category has evidence a refine control could act on.
|
|
///
|
|
/// The panel asks before offering the slider: a control that moves and
|
|
/// does nothing is worse than an absent one.
|
|
pub fn category_is_refinable(&self, name: &str) -> bool {
|
|
self.categories
|
|
.iter()
|
|
.find(|c| &*c.name == name)
|
|
.is_some_and(|c| c.refinement.is_some())
|
|
}
|
|
|
|
/// Where this category's refine control should start on *this*
|
|
/// photograph.
|
|
///
|
|
/// [`dr_segment::STRICTNESS_OFF`] where nothing was fitted, and otherwise
|
|
/// whatever [`dr_segment::Refinement::gentle`] finds the frame will bear —
|
|
/// see there for why a constant could not do it and what a constant cost.
|
|
///
|
|
/// Costs one to four `apply` passes, tens of milliseconds each. That is
|
|
/// why it is asked here, when a layer is made, rather than for every
|
|
/// category during the precompute: eight categories nobody masked would be
|
|
/// seconds added to a wait, and the answer is only wanted for the one that
|
|
/// was clicked.
|
|
pub fn category_default_refine(&self, name: &str) -> f32 {
|
|
self.categories
|
|
.iter()
|
|
.find(|c| &*c.name == name)
|
|
.and_then(|c| Some((c.refinement.as_ref()?, &c.mask)))
|
|
.map_or(dr_segment::STRICTNESS_OFF, |(refinement, mask)| {
|
|
refinement.gentle(mask)
|
|
})
|
|
}
|
|
|
|
/// Replace one instance in place, keeping every other index and the
|
|
/// signature unchanged.
|
|
///
|
|
/// What a refine pass calls once it has a sharper mask for the subject at
|
|
/// `index`: the layers pointing at this run by index still mean what they
|
|
/// meant, they just read better pixels now.
|
|
pub fn replace_instance(&mut self, index: usize, instance: InstanceSummary) {
|
|
if let Some(slot) = self.instances.get_mut(index) {
|
|
*slot = instance;
|
|
}
|
|
}
|
|
|
|
/// The strongest instance covering a point in normalised image
|
|
/// coordinates.
|
|
///
|
|
/// Strongest rather than smallest: detections are score-ordered and
|
|
/// overlapping ones are usually the same object found twice, so the more
|
|
/// confident is the better guess. A person in front of a bus wins over the
|
|
/// bus, because the person's mask is the one under the cursor at all.
|
|
pub fn instance_at(&self, x: f32, y: f32) -> Option<usize> {
|
|
if !(0.0..1.0).contains(&x) || !(0.0..1.0).contains(&y) {
|
|
return None;
|
|
}
|
|
let (w, h) = self.proxy;
|
|
if w == 0 || h == 0 {
|
|
return None;
|
|
}
|
|
let px = ((x * w as f32) as usize).min(w - 1);
|
|
let py = ((y * h as f32) as usize).min(h - 1);
|
|
let p = py * w + px;
|
|
|
|
self.instances
|
|
.iter()
|
|
.enumerate()
|
|
.filter(|(_, i)| i.mask.get(p).is_some_and(|&v| v >= 128))
|
|
.max_by(|(_, a), (_, b)| a.score.total_cmp(&b.score))
|
|
.map(|(i, _)| i)
|
|
}
|
|
|
|
/// A false-coloured picture of what a click can select, in source space.
|
|
///
|
|
/// **Transparent where nothing is selectable.** The region map this
|
|
/// replaced covered every pixel and so hid the photograph it was drawn
|
|
/// over; the question an overlay exists to answer is whether an outline
|
|
/// follows the subject, and that can only be answered by seeing both.
|
|
pub fn overlay_rgba(&self) -> (Vec<u8>, u32, u32) {
|
|
let (w, h) = self.proxy;
|
|
let mut out = vec![0u8; w * h * 4];
|
|
|
|
// Weakest first, so where two detections overlap the more confident
|
|
// one is the colour on top — matching which a click would select.
|
|
let mut order: Vec<usize> = (0..self.instances.len()).collect();
|
|
order.sort_by(|&a, &b| self.instances[a].score.total_cmp(&self.instances[b].score));
|
|
|
|
for &i in &order {
|
|
let [r, g, b] = instance_colour(i as u32);
|
|
for (p, &cov) in self.instances[i].mask.iter().enumerate() {
|
|
if cov < 128 || p * 4 + 3 >= out.len() {
|
|
continue;
|
|
}
|
|
out[p * 4] = r;
|
|
out[p * 4 + 1] = g;
|
|
out[p * 4 + 2] = b;
|
|
out[p * 4 + 3] = 255;
|
|
}
|
|
}
|
|
|
|
// The outline drawn opaque white over the fill. It is the part being
|
|
// judged — a fill can look right while its edge sits several pixels
|
|
// off the subject — and it survives the low opacity the fill is
|
|
// composited at.
|
|
let solid = |p: usize| out.get(p * 4 + 3).is_some_and(|&a| a > 0);
|
|
let mut edges = Vec::new();
|
|
for y in 0..h {
|
|
for x in 0..w {
|
|
let p = y * w + x;
|
|
if !solid(p) {
|
|
continue;
|
|
}
|
|
let boundary = (x + 1 == w || !solid(p + 1))
|
|
|| (x == 0 || !solid(p - 1))
|
|
|| (y + 1 == h || !solid(p + w))
|
|
|| (y == 0 || !solid(p - w));
|
|
if boundary {
|
|
edges.push(p);
|
|
}
|
|
}
|
|
}
|
|
for p in edges {
|
|
out[p * 4..p * 4 + 4].copy_from_slice(&[255, 255, 255, 255]);
|
|
}
|
|
|
|
(out, w as u32, h as u32)
|
|
}
|
|
}
|
|
|
|
/// A distinct colour per instance.
|
|
///
|
|
/// Golden-angle hue stepping, deterministic rather than random: the same
|
|
/// object is the same colour every time the overlay is drawn, so the eye can
|
|
/// track it while a mask is being shaped.
|
|
fn instance_colour(index: u32) -> [u8; 3] {
|
|
let h = (index as f32 * 137.508) % 360.0;
|
|
let c = 230.0;
|
|
let x = c * (1.0 - ((h / 60.0) % 2.0 - 1.0).abs());
|
|
let (r, g, b) = match (h / 60.0) as u32 {
|
|
0 => (c, x, 0.0),
|
|
1 => (x, c, 0.0),
|
|
2 => (0.0, c, x),
|
|
3 => (0.0, x, c),
|
|
4 => (x, 0.0, c),
|
|
_ => (c, 0.0, x),
|
|
};
|
|
[r as u8 + 25, g as u8 + 25, b as u8 + 25]
|
|
}
|
|
|
|
/// What to run.
|
|
#[derive(Debug, Clone, Copy, PartialEq)]
|
|
pub struct Options {
|
|
/// Detections below this are dropped.
|
|
///
|
|
/// Deliberately low. A weak detection costs a spurious entry in a list the
|
|
/// user is choosing from, and its score is shown beside it; a missed one
|
|
/// costs a subject that cannot be selected at all, which is the worse
|
|
/// failure for a selection tool.
|
|
pub confidence: f32,
|
|
/// TRACES: FR-DEV-3
|
|
/// Run the model over overlapping tiles instead of the whole frame once.
|
|
///
|
|
/// Off by default and deliberately so. The graph's input is fixed at
|
|
/// 640x640 (docs/dev/segmentation.md F6), so every image is letterboxed into
|
|
/// it and a subject 200px across in a 1600px proxy reaches the model at
|
|
/// 80px — which is where a coarse outline comes from. Tiling is the only
|
|
/// route to more resolution with a fixed window, and it costs one
|
|
/// inference per tile: about 2.8s for a 3x2 grid against 470ms whole-frame.
|
|
///
|
|
/// Six times the wait is the wrong default for the common case, where the
|
|
/// subject is large in frame and whole-frame inference is already the best
|
|
/// answer. It is the right answer for a bird against sky, so it is offered
|
|
/// per-image rather than chosen once for all of them.
|
|
pub fine: bool,
|
|
|
|
/// TRACES: FR-DEV-3
|
|
/// Gather the evidence for cutting each scene category back to the pixels
|
|
/// whose colour agrees with it.
|
|
///
|
|
/// The scene model's counterpart to [`Options::fine`], and it exists
|
|
/// because tiling is not available here: a category has no bounding box to
|
|
/// tile over — sky is wherever the sky is — so the only route to a sharper
|
|
/// category edge is the photograph itself. See [`dr_segment::refine`].
|
|
///
|
|
/// This buys the *possibility* of refinement rather than any of it. What
|
|
/// it produces is a `dr_segment::Refinement` per category, which changes
|
|
/// no mask until a layer's refine control is moved off zero — so turning
|
|
/// this off costs the control, not the appearance of anything.
|
|
///
|
|
/// On by default where `fine` is off, because the two have opposite costs.
|
|
/// Tiling is six inferences and a five-second wait; this is a distance
|
|
/// transform and a k-means over a subsample per category, tens of
|
|
/// milliseconds each against a precompute already measured in hundreds.
|
|
pub refine: bool,
|
|
}
|
|
|
|
impl Default for Options {
|
|
fn default() -> Self {
|
|
Self {
|
|
confidence: 0.30,
|
|
fine: false,
|
|
refine: true,
|
|
}
|
|
}
|
|
}
|
|
|
|
/// Find what can be selected in this photograph.
|
|
///
|
|
/// `rgb` is the rendered proxy — tightly packed RGB floats at
|
|
/// `(width, height)`, in the **sensor's** own orientation. Passed in rather
|
|
/// than derived here because the caller already has it, and re-deriving it
|
|
/// would mean a second readback of something the CPU is holding.
|
|
///
|
|
/// `orientation` is the file's EXIF tag composed with whatever turns the
|
|
/// photographer has since applied — `Framing::effective_orientation`, one
|
|
/// permutation covering both. The model is shown the picture through it and
|
|
/// its answers come back without it, so everything this returns is in sensor
|
|
/// space exactly as it was before the detector was taught to read.
|
|
pub fn compute(
|
|
_ctx: &GpuContext,
|
|
rgb: &[f32],
|
|
width: usize,
|
|
height: usize,
|
|
orientation: Orientation,
|
|
options: &Options,
|
|
) -> Result<Segmentation, String> {
|
|
// The model reads the photograph; everything else here speaks sensor.
|
|
let (stood_up, uw, uh) = upright(rgb, width, height, orientation);
|
|
let found = detect(&stood_up, uw, uh, options.fine)?;
|
|
|
|
let instances: Vec<InstanceSummary> = found
|
|
.iter()
|
|
.filter(|i| i.score >= options.confidence)
|
|
.map(|i| {
|
|
let (mask, bbox) = lay_down(&i.mask, i.bbox, uw, uh, orientation);
|
|
InstanceSummary {
|
|
class_name: i.class_name.clone(),
|
|
score: i.score,
|
|
// Quantised after the permutation, so the byte stored is a
|
|
// rounding of the model's own coverage and not of a copy.
|
|
mask: quantise(&mask),
|
|
bbox,
|
|
}
|
|
})
|
|
.collect();
|
|
|
|
let signature = segmentation_signature(
|
|
width as u32,
|
|
height as u32,
|
|
instances.len() as u32,
|
|
// The tiling choice belongs in the signature as much as the
|
|
// confidence does. A mask stores the signature of the segmentation its
|
|
// region ids index into (`MaskSource::Regions`), and a tiled run finds
|
|
// different instances in a different order — so if the two runs shared
|
|
// a signature, a layer built against the coarse pass would be silently
|
|
// reinterpreted against the fine one. That is a *wrong* mask, which is
|
|
// far worse than a stale one, because nothing announces it.
|
|
//
|
|
// The orientation is in for the same reason and it is not hypothetical:
|
|
// turning the photograph changes what the model recognises, so a run
|
|
// before a quarter turn and a run after it are different instance
|
|
// lists. Two lists that happened to come out the same length would
|
|
// otherwise share a signature, and a layer built against the first
|
|
// would be silently re-indexed into the second.
|
|
options.confidence.to_bits() as u64
|
|
^ if options.fine {
|
|
0x9E37_79B9_7F4A_7C15
|
|
} else {
|
|
0
|
|
}
|
|
^ orientation_key(orientation),
|
|
);
|
|
|
|
// The scene pass, on the same upright frame and laid back down the same
|
|
// way. Failures here are logged and dropped rather than propagated: no
|
|
// scene model is an ordinary state, and a photograph that can be masked by
|
|
// subject should not become unopenable because the categories are absent.
|
|
let categories = match scene_categories(
|
|
&Frames {
|
|
upright: &stood_up,
|
|
upright_size: (uw, uh),
|
|
sensor: rgb,
|
|
sensor_size: (width, height),
|
|
},
|
|
orientation,
|
|
options.refine,
|
|
) {
|
|
Ok(c) => c,
|
|
Err(e) => {
|
|
log::info!("no scene categories for this frame: {e}");
|
|
Vec::new()
|
|
}
|
|
};
|
|
|
|
Ok(Segmentation {
|
|
instances,
|
|
categories,
|
|
signature,
|
|
// **Sensor space, not the model's.** `lay_down` put every mask back,
|
|
// so the grid a stored layer indexes into is the one it always was —
|
|
// see `upright` for why the model saw a different one.
|
|
proxy: (width, height),
|
|
})
|
|
}
|
|
|
|
/// TRACES: FR-DEV-3 | FR-DEV-3h
|
|
/// Turn the proxy the way the photographer is looking at it.
|
|
///
|
|
/// **Why this exists at all.** A camera held sideways writes its sensor rows
|
|
/// the way it always does, and the render puts them right by way of
|
|
/// `Framing`. The proxy the model reads is deliberately rendered through a
|
|
/// *neutral* graph — the detection has to survive an exposure change, or
|
|
/// every slider would invalidate the masks built on it — and neutral took the
|
|
/// orientation with it. So the detector was handed a portrait frame lying on
|
|
/// its side, and a model trained on upright photographs is very bad at those.
|
|
/// Measured end to end on one 22 MP frame of two people and a dog: `person
|
|
/// 0.36` and nothing else, against `dog 0.82, person 0.61, person 0.49` for
|
|
/// the same pixels stood up.
|
|
///
|
|
/// One line, because the permutation belongs to
|
|
/// [`dr_types::Orientation`] and every other consumer goes through the same
|
|
/// one — the grid's thumbnails included, which is what makes "upright" mean
|
|
/// one thing across the application rather than one thing per caller.
|
|
pub(crate) fn upright(
|
|
rgb: &[f32],
|
|
width: usize,
|
|
height: usize,
|
|
orientation: Orientation,
|
|
) -> (Vec<f32>, usize, usize) {
|
|
let (out, w, h) = orientation.into_shown(rgb, width as u32, height as u32, 3);
|
|
(out, w as usize, h as usize)
|
|
}
|
|
|
|
/// TRACES: FR-DEV-3
|
|
/// Put what the model answered back onto the sensor's grid.
|
|
///
|
|
/// The counterpart of [`upright`], and the two are always used as a pair: a
|
|
/// mask is only ever in the model's frame between those two calls. Returned
|
|
/// together rather than as two functions a caller composes, because calling
|
|
/// one and forgetting the other is silent — the mask lands a quarter turn off
|
|
/// the subject, which reads as a bad detection rather than as a bug.
|
|
///
|
|
/// `dw`/`dh` are the *shown* dimensions, as [`upright`] returned them.
|
|
pub(crate) fn lay_down(
|
|
mask: &[f32],
|
|
bbox: (f32, f32, f32, f32),
|
|
dw: usize,
|
|
dh: usize,
|
|
orientation: Orientation,
|
|
) -> (Vec<f32>, (f32, f32, f32, f32)) {
|
|
let (out, sw, sh) = orientation.into_stored(mask, dw as u32, dh as u32, 1);
|
|
(out, lay_down_bbox(bbox, dw, dh, sw, sh, orientation))
|
|
}
|
|
|
|
/// [`lay_down`] for a box.
|
|
///
|
|
/// Normalised on the way in and scaled on the way out, so the turn itself is
|
|
/// `Orientation::into_stored_rect` rather than a fourth copy of the corner
|
|
/// arithmetic. A box is the one place a permutation can be *nearly* right —
|
|
/// the corners land correctly and `x0 > x1` — so the shared map takes the
|
|
/// extremes and this only has to say what space it is in.
|
|
fn lay_down_bbox(
|
|
bbox: (f32, f32, f32, f32),
|
|
dw: usize,
|
|
dh: usize,
|
|
sw: u32,
|
|
sh: u32,
|
|
orientation: Orientation,
|
|
) -> (f32, f32, f32, f32) {
|
|
if dw == 0 || dh == 0 {
|
|
return bbox;
|
|
}
|
|
let (fw, fh) = (dw as f32, dh as f32);
|
|
let shown = dr_types::ShownRect {
|
|
x: bbox.0 / fw,
|
|
y: bbox.1 / fh,
|
|
width: (bbox.2 - bbox.0) / fw,
|
|
height: (bbox.3 - bbox.1) / fh,
|
|
};
|
|
let stored = orientation.into_stored_rect(shown);
|
|
(
|
|
stored.x * sw as f32,
|
|
stored.y * sh as f32,
|
|
(stored.x + stored.width) * sw as f32,
|
|
(stored.y + stored.height) * sh as f32,
|
|
)
|
|
}
|
|
|
|
/// One of eight transforms, as bits a signature can carry.
|
|
fn orientation_key(orientation: Orientation) -> u64 {
|
|
u64::from(orientation.quarter_turns)
|
|
| (u64::from(orientation.flip_h) << 2)
|
|
| (u64::from(orientation.flip_v) << 3)
|
|
}
|
|
|
|
/// Classes a recognised face may put a name on.
|
|
///
|
|
/// Only these. A face inside a `tv` or a `laptop` is a photograph of someone on
|
|
/// a screen, and renaming the television to "Anna" would be worse than leaving
|
|
/// it as the model found it.
|
|
fn is_person_class(name: &str) -> bool {
|
|
name == "person"
|
|
}
|
|
|
|
impl Segmentation {
|
|
/// Relabel `person` instances with the name of the face inside them.
|
|
///
|
|
/// The segmenter knows it found *a person*; the face index knows *which*
|
|
/// person. Joining them turns "person" in the mask list into "Anna", which
|
|
/// is the difference between COCO's eighty classes and a vocabulary that
|
|
/// includes the user's family — and it is the same click either way, so the
|
|
/// gain is entirely in being able to tell two people apart before clicking.
|
|
///
|
|
/// `faces` are in the same proxy pixels as [`InstanceSummary::bbox`]. The
|
|
/// geometry lives in `dr_face::naming`, where it is testable without a
|
|
/// model or a catalog.
|
|
///
|
|
/// Returns how many instances gained a name. An unrecognised person keeps
|
|
/// the model's own label, which is the right default: "person" is merely
|
|
/// unhelpful, where a wrong name is wrong and the user cannot tell which
|
|
/// they are looking at.
|
|
pub fn apply_names(&mut self, faces: &[dr_face::NamedFace<'_>]) -> usize {
|
|
dr_face::name_instances(
|
|
&mut self.instances,
|
|
faces,
|
|
|i| i.bbox,
|
|
|i| is_person_class(&i.class_name),
|
|
|i, name| i.class_name = name.into(),
|
|
)
|
|
}
|
|
}
|
|
|
|
/// Load the model and run it.
|
|
///
|
|
/// Loading is ~24 ms against the ~470 ms of inference that follows, and this
|
|
/// runs once per image — so caching the session would keep 11 MB of weights
|
|
/// resident for the life of the app to save five percent of a background task.
|
|
fn detect(
|
|
rgb: &[f32],
|
|
width: usize,
|
|
height: usize,
|
|
fine: bool,
|
|
) -> Result<Vec<dr_segment::Instance>, String> {
|
|
let mut model = dr_segment::SemanticModel::embedded().map_err(|e| e.to_string())?;
|
|
let options = dr_segment::SemanticOptions {
|
|
// A quarter shared with each neighbour. It has to exceed zero at all,
|
|
// or a subject sitting on a seam is cut in half by both tiles and
|
|
// recognised by neither; a quarter is enough to carry a whole subject
|
|
// inside one tile at the sizes tiling is reached for.
|
|
tiling: if fine {
|
|
dr_segment::Tiling::Grid { overlap: 0.25 }
|
|
} else {
|
|
dr_segment::Tiling::Whole
|
|
},
|
|
..dr_segment::SemanticOptions::default()
|
|
};
|
|
model
|
|
.detect(rgb, width, height, &options)
|
|
.map_err(|e| e.to_string())
|
|
}
|
|
|
|
/// The one photograph, in the two frames this pass needs it in.
|
|
///
|
|
/// Both, and not one plus a permutation applied where needed, because this
|
|
/// function reads the picture *twice* for different purposes and the two want
|
|
/// opposite frames. The model must be shown an upright photograph or it
|
|
/// recognises far less (see [`upright`]); the refinement must read the sensor's
|
|
/// own grid, because the field it produces has to line up pixel for pixel with
|
|
/// a mask that has already been laid back down.
|
|
///
|
|
/// Carried as a struct rather than six parameters so that a caller cannot
|
|
/// quietly transpose the pair — which would produce a plausible mask over
|
|
/// slightly the wrong pixels, the failure this whole module is most prone to.
|
|
struct Frames<'a> {
|
|
/// As the photographer sees it. What the model reads.
|
|
upright: &'a [f32],
|
|
upright_size: (usize, usize),
|
|
/// As the sensor wrote it. The space every stored mask lives in.
|
|
sensor: &'a [f32],
|
|
sensor_size: (usize, usize),
|
|
}
|
|
|
|
/// Weigh the photographic categories, if this build can find a scene model.
|
|
///
|
|
/// Separate from [`detect`] rather than folded into it because the two are
|
|
/// independent: a build with no scene model still segments subjects, and a
|
|
/// frame with no recognisable subject still has sky. Neither failure should
|
|
/// take the other down.
|
|
fn scene_categories(
|
|
frames: &Frames<'_>,
|
|
orientation: Orientation,
|
|
refine: bool,
|
|
) -> Result<Vec<CategorySummary>, String> {
|
|
let (upright, (width, height)) = (frames.upright, frames.upright_size);
|
|
let (sensor, (sensor_width, sensor_height)) = (frames.sensor, frames.sensor_size);
|
|
|
|
let mut model = load_scene_model()?;
|
|
let scene = model
|
|
.analyse(upright, width, height)
|
|
.map_err(|e| e.to_string())?;
|
|
|
|
let mut out = Vec::new();
|
|
for (index, name) in scene.categories().iter().enumerate() {
|
|
let coverage = scene.coverage(index);
|
|
// A category the model barely saw is not worth a mask buffer the size
|
|
// of the proxy, and offering it in the list would be offering a
|
|
// control that does nothing when moved. The threshold is the one the
|
|
// example prints against.
|
|
if coverage < 0.005 {
|
|
continue;
|
|
}
|
|
let Some(weights) = scene.rasterise(index, width, height) else {
|
|
continue;
|
|
};
|
|
|
|
// Back into sensor space, exactly as an instance mask is: the grid a
|
|
// stored layer indexes into has to be the sensor's whatever the model
|
|
// was shown. `lay_down` wants a box too, so it gets the whole frame —
|
|
// a category has no meaningful extent.
|
|
let (mask, _) = lay_down(
|
|
&weights,
|
|
(0.0, 0.0, width as f32, height as f32),
|
|
width,
|
|
height,
|
|
orientation,
|
|
);
|
|
|
|
// Gather the evidence for cutting this category back, on the far side
|
|
// of the permutation.
|
|
//
|
|
// Sensor space rather than the model's, and that is the whole reason
|
|
// this happens here: the verdict is a per-pixel field over the same
|
|
// grid as the mask it gates, so fitting it upright would mean
|
|
// permuting it afterwards to match — a second rotation of a
|
|
// proxy-sized buffer, and a second chance to get one wrong.
|
|
//
|
|
// One logit cell is the same number of pixels in either frame: the
|
|
// letterbox scales by the *longer* edge, and a permutation does not
|
|
// change which edge that is.
|
|
let refinement = refine
|
|
.then(|| {
|
|
dr_segment::Refinement::compute(
|
|
&mask,
|
|
sensor,
|
|
sensor_width,
|
|
sensor_height,
|
|
scene.cell_pixels(),
|
|
&dr_segment::RefineOptions::default(),
|
|
)
|
|
})
|
|
.and_then(|result| match result {
|
|
Ok(refinement) => Some(refinement),
|
|
// An ordinary state, not a failure — see `dr_segment::refine`.
|
|
// Logged all the same, because the layer will simply have no
|
|
// refine control and "there is no slider" should be traceable
|
|
// to a reason rather than looking like a missing feature.
|
|
Err(why) => {
|
|
log::debug!("no refinement for '{name}': {why:?}");
|
|
None
|
|
}
|
|
});
|
|
|
|
out.push(CategorySummary {
|
|
name: name.clone(),
|
|
coverage,
|
|
mask: quantise(&mask),
|
|
refinement,
|
|
});
|
|
}
|
|
|
|
// Largest first, which is the order the list is worth reading in.
|
|
out.sort_by(|a, b| b.coverage.total_cmp(&a.coverage));
|
|
Ok(out)
|
|
}
|
|
|
|
/// Find a scene model: compiled in if this build has one, installed otherwise.
|
|
///
|
|
/// The order matters. A build with the weights compiled in should not be
|
|
/// silently overridden by a stale file in a data directory, and a build
|
|
/// without them has nothing to fall back *from* — so "embedded, then
|
|
/// installed" is the only ordering that is not surprising either way.
|
|
fn load_scene_model() -> Result<dr_segment::SceneModel, String> {
|
|
#[cfg(feature = "scene-model")]
|
|
{
|
|
dr_segment::SceneModel::embedded().map_err(|e| e.to_string())
|
|
}
|
|
#[cfg(not(feature = "scene-model"))]
|
|
{
|
|
// Account-independent, like `shared_face_models_dir` and for the same
|
|
// reason: this runs on a worker with no session in hand. Android
|
|
// unpacks the APK's copy to exactly this directory, on a worker of its
|
|
// own started at launch — so on the first launch after an install this
|
|
// can report the model missing for the couple of seconds the 25 MB copy
|
|
// takes. See `shared_face_models_dir` for why it is no longer done
|
|
// before the first frame.
|
|
let dir = crate::library::shared_face_models_dir();
|
|
dr_segment::SceneModel::from_path(
|
|
dir.join("yolo26s-sem-ade20k.onnx"),
|
|
dir.join("yolo26s-sem-ade20k.classes.json"),
|
|
dir.join("categories.txt"),
|
|
)
|
|
.map_err(|e| e.to_string())
|
|
}
|
|
}
|
|
|
|
/// The model's soft coverage, to a byte per pixel.
|
|
///
|
|
/// Rounded rather than truncated, so a coverage of exactly 0.5 lands on the
|
|
/// threshold the selection tests against instead of one below it.
|
|
fn quantise(mask: &[f32]) -> Vec<u8> {
|
|
mask.iter()
|
|
.map(|&v| (v.clamp(0.0, 1.0) * 255.0).round() as u8)
|
|
.collect()
|
|
}
|
|
|
|
#[cfg(test)]
|
|
mod tests {
|
|
use super::*;
|
|
|
|
/// Two objects: a big weak one on the left, a small strong one that
|
|
/// overlaps it.
|
|
fn overlapping() -> Segmentation {
|
|
let (w, h) = (8usize, 4usize);
|
|
let mut big = vec![0u8; w * h];
|
|
let mut small = vec![0u8; w * h];
|
|
for y in 0..h {
|
|
for x in 0..6 {
|
|
big[y * w + x] = 255;
|
|
}
|
|
for x in 4..8 {
|
|
small[y * w + x] = 255;
|
|
}
|
|
}
|
|
|
|
Segmentation {
|
|
instances: vec![
|
|
InstanceSummary {
|
|
class_name: "bus".into(),
|
|
score: 0.5,
|
|
mask: big,
|
|
bbox: (0.0, 0.0, 6.0, h as f32),
|
|
},
|
|
InstanceSummary {
|
|
class_name: "person".into(),
|
|
score: 0.9,
|
|
mask: small,
|
|
bbox: (4.0, 0.0, 8.0, h as f32),
|
|
},
|
|
],
|
|
categories: Vec::new(),
|
|
signature: 1,
|
|
proxy: (w, h),
|
|
}
|
|
}
|
|
|
|
#[test]
|
|
fn a_click_finds_the_object_under_it() {
|
|
let seg = overlapping();
|
|
assert_eq!(seg.instance_at(0.1, 0.5), Some(0), "only the bus here");
|
|
assert_eq!(seg.instance_at(0.95, 0.5), Some(1), "only the person here");
|
|
}
|
|
|
|
/// The overlap rule, and the one that decides what a click means where two
|
|
/// detections cover the same pixel.
|
|
#[test]
|
|
fn overlapping_objects_resolve_to_the_more_confident() {
|
|
let seg = overlapping();
|
|
assert_eq!(
|
|
seg.instance_at(0.6, 0.5),
|
|
Some(1),
|
|
"the person at 0.9 beats the bus at 0.5"
|
|
);
|
|
}
|
|
|
|
#[test]
|
|
fn a_click_outside_the_frame_selects_nothing() {
|
|
let seg = overlapping();
|
|
assert_eq!(seg.instance_at(-0.1, 0.5), None);
|
|
assert_eq!(seg.instance_at(1.5, 0.5), None);
|
|
assert_eq!(seg.instance_at(0.5, 1.0), None, "the far edge is exclusive");
|
|
}
|
|
|
|
#[test]
|
|
fn a_click_on_nothing_selects_nothing() {
|
|
let seg = Segmentation {
|
|
instances: Vec::new(),
|
|
categories: Vec::new(),
|
|
signature: 1,
|
|
proxy: (4, 4),
|
|
};
|
|
assert_eq!(seg.instance_at(0.5, 0.5), None);
|
|
}
|
|
|
|
#[test]
|
|
fn colours_are_stable_and_distinct() {
|
|
assert_eq!(instance_colour(7), instance_colour(7));
|
|
assert_ne!(instance_colour(0), instance_colour(1));
|
|
assert_ne!(instance_colour(1), instance_colour(2));
|
|
}
|
|
|
|
#[test]
|
|
fn every_colour_is_visible_against_a_photograph() {
|
|
for i in 0..64u32 {
|
|
let [r, g, b] = instance_colour(i);
|
|
assert!(r.max(g).max(b) >= 200, "instance {i} is too dark");
|
|
}
|
|
}
|
|
|
|
/// The overlay must not cover the picture: that is the difference between
|
|
/// this and the region map it replaced.
|
|
#[test]
|
|
fn the_overlay_is_transparent_where_nothing_was_found() {
|
|
let seg = Segmentation {
|
|
instances: Vec::new(),
|
|
categories: Vec::new(),
|
|
signature: 1,
|
|
proxy: (4, 4),
|
|
};
|
|
let (px, w, h) = seg.overlay_rgba();
|
|
assert_eq!((w, h), (4, 4));
|
|
assert!(px.chunks_exact(4).all(|p| p[3] == 0), "nothing to draw");
|
|
}
|
|
|
|
#[test]
|
|
fn the_overlay_outlines_what_it_fills() {
|
|
let seg = overlapping();
|
|
let (px, w, _) = seg.overlay_rgba();
|
|
let at = |x: usize, y: usize| {
|
|
let p = (y * w as usize + x) * 4;
|
|
[px[p], px[p + 1], px[p + 2], px[p + 3]]
|
|
};
|
|
// The rightmost column of the person is against the frame edge, so it
|
|
// is an outline pixel.
|
|
assert_eq!(at(7, 1), [255, 255, 255, 255]);
|
|
// And an interior pixel keeps its fill.
|
|
assert_ne!(at(2, 1)[3], 0, "the bus is filled");
|
|
assert_ne!(at(2, 1), [255, 255, 255, 255], "and not all outline");
|
|
}
|
|
|
|
#[test]
|
|
fn quantising_rounds_rather_than_truncates() {
|
|
// Exactly half must reach the threshold the selection tests against.
|
|
assert_eq!(quantise(&[0.5]), vec![128]);
|
|
assert_eq!(quantise(&[0.0, 1.0]), vec![0, 255]);
|
|
// And values outside the range cannot wrap.
|
|
assert_eq!(quantise(&[-1.0, 2.0]), vec![0, 255]);
|
|
}
|
|
|
|
/// A non-square, wholly asymmetric grid: every pixel is its own index, so
|
|
/// any permutation that is not the intended one shows up as a mismatch
|
|
/// rather than being hidden by a symmetry.
|
|
fn ramp(w: usize, h: usize) -> Vec<f32> {
|
|
(0..w * h).flat_map(|i| [i as f32, 0.0, 0.0]).collect()
|
|
}
|
|
|
|
fn red(rgb: &[f32]) -> Vec<f32> {
|
|
rgb.chunks_exact(3).map(|p| p[0]).collect()
|
|
}
|
|
|
|
/// The property the whole fix rests on: what the model is shown and what
|
|
/// comes back are the same permutation, run in opposite directions. If
|
|
/// they ever disagree, every subject mask lands somewhere other than its
|
|
/// subject — and looks like a mask while doing it.
|
|
#[test]
|
|
fn standing_a_frame_up_and_laying_it_down_is_the_identity() {
|
|
const W: usize = 5;
|
|
const H: usize = 3;
|
|
let source = ramp(W, H);
|
|
|
|
for tag in 1..=8u16 {
|
|
let o = Orientation::from_exif(tag);
|
|
let (up, uw, uh) = upright(&source, W, H, o);
|
|
let (ow, oh) = o.oriented_size(W as u32, H as u32);
|
|
assert_eq!(
|
|
(uw, uh),
|
|
(ow as usize, oh as usize),
|
|
"tag {tag}: the upright size is the oriented one"
|
|
);
|
|
|
|
let (back, _) = lay_down(&red(&up), (0.0, 0.0, 1.0, 1.0), uw, uh, o);
|
|
assert_eq!(back, red(&source), "tag {tag} did not come back");
|
|
}
|
|
}
|
|
|
|
/// The colour channels must travel together. Reading a pixel three times
|
|
/// with one index arithmetic mistake gives a plausible image with its
|
|
/// channels sheared, which the model would still detect *something* in.
|
|
#[test]
|
|
fn a_turn_carries_whole_pixels() {
|
|
let rgb: Vec<f32> = (0..2 * 3)
|
|
.flat_map(|i| [i as f32, i as f32 + 100.0, i as f32 + 200.0])
|
|
.collect();
|
|
let o = Orientation::from_exif(6);
|
|
let (up, uw, uh) = upright(&rgb, 2, 3, o);
|
|
|
|
assert_eq!((uw, uh), (3, 2));
|
|
for p in up.chunks_exact(3) {
|
|
assert_eq!(p[1], p[0] + 100.0, "green left its pixel");
|
|
assert_eq!(p[2], p[0] + 200.0, "blue left its pixel");
|
|
}
|
|
}
|
|
|
|
/// A portrait frame is the case this exists for: the sensor is landscape,
|
|
/// the photograph is not, and the model has to be given the photograph.
|
|
#[test]
|
|
fn a_sideways_frame_reaches_the_model_upright() {
|
|
let o = Orientation::from_exif(6);
|
|
assert!(!o.is_normal());
|
|
|
|
let (_, uw, uh) = upright(&ramp(1600, 1066), 1600, 1066, o);
|
|
assert_eq!((uw, uh), (1066, 1600), "the model still got a landscape");
|
|
}
|
|
|
|
/// Where the model's box ends up, worked out by hand for the one turn a
|
|
/// portrait phone or a sideways body actually writes.
|
|
#[test]
|
|
fn a_box_comes_back_in_sensor_pixels() {
|
|
let o = Orientation::from_exif(6);
|
|
// Shown 4x6; the sensor it came from is 6x4.
|
|
let bbox = lay_down_bbox((0.0, 0.0, 2.0, 3.0), 4, 6, 6, 4, o);
|
|
assert_eq!(bbox, (0.0, 2.0, 3.0, 4.0));
|
|
}
|
|
|
|
/// Whatever a turn does to a box, it must still read low-to-high — a
|
|
/// permutation exchanges which corner is which, and the rest of the mask
|
|
/// pipeline measures `(x1 - x0)` without checking the sign.
|
|
#[test]
|
|
fn a_restored_box_keeps_its_corners_in_order() {
|
|
for tag in 1..=8u16 {
|
|
let o = Orientation::from_exif(tag);
|
|
let (sw, sh) = o.oriented_size(9, 6);
|
|
let (x0, y0, x1, y1) = lay_down_bbox((1.0, 2.0, 7.0, 5.0), 9, 6, sw, sh, o);
|
|
assert!(x0 <= x1, "tag {tag}: x runs backwards");
|
|
assert!(y0 <= y1, "tag {tag}: y runs backwards");
|
|
// A permutation moves a box; it does not resize one.
|
|
let (sw, sh) = o.oriented_size(9, 6);
|
|
assert!(x1 <= sw as f32 && y1 <= sh as f32, "tag {tag}: box escaped");
|
|
assert!((((x1 - x0) * (y1 - y0)) - 18.0).abs() < 1e-3, "tag {tag}");
|
|
}
|
|
}
|
|
|
|
/// Turning the photograph changes what the model recognises, so the two
|
|
/// runs are different instance lists. If they could share a signature, a
|
|
/// layer built against one would be silently re-indexed into the other —
|
|
/// the same failure the tiling flag is in the signature to prevent.
|
|
#[test]
|
|
fn turning_the_photograph_changes_the_signature() {
|
|
let mut seen = std::collections::HashSet::new();
|
|
for tag in 1..=8u16 {
|
|
let o = Orientation::from_exif(tag);
|
|
assert!(
|
|
seen.insert(orientation_key(o)),
|
|
"tag {tag} shares a key with an earlier one"
|
|
);
|
|
}
|
|
assert_eq!(seen.len(), 8, "eight tags, but some collapsed");
|
|
}
|
|
|
|
#[test]
|
|
fn confidence_changes_the_signature() {
|
|
// A different threshold is a different instance list, so the indices a
|
|
// stored layer holds mean something else.
|
|
let a = segmentation_signature(100, 100, 3, Options::default().confidence.to_bits() as u64);
|
|
let b = segmentation_signature(100, 100, 3, 0.9f32.to_bits() as u64);
|
|
assert_ne!(a, b);
|
|
}
|
|
}
|