Files
DarkRoom/core/dr-segment/src/lib.rs
T
dtourolle 5a8c3e4c40 Run each model on the Hexagon in the form measured to hold it
The engine knew f32 and int8, and gave the Hexagon int8 for every role it
served. Measured on the tablet itself (inference.md §1.5), int8 lost
5% of the detector's faces at 40-80 px, moved the landmarks 1.5 px,
emptied the segmenter's scores and cost the denoiser 5-9 dB; fp16 the HTP
refuses outright. `Form` gains A16W8 and A16W16, and `Rung::form` now
names one per role: detectors and landmarks A16W8, the segmenter, scene
model, border filler and denoiser A16W16, XFeat int8. The embedder and
the eye classifiers stay on the CPU.

Each loader resolves its `<stem>.<form>.onnx` sibling; the segmenter and
XFeat, compiled into the binary, embed their quantised forms on Android
only and pick through `choose_embedded`. The probe, the compile step and
the cache fingerprint follow the form instead of assuming int8. Detectors
on the new form write `scrfd_*_a16+w600k_mbf`, and `model_ids` answers
for all three spellings.

On the tablet (ORT 1.29 + QNN 2.42), each shipped file against f32 on the
same inputs, and against the CPU's f32 time:
  SCRFD 500m/2.5g/10g  A16W8   100% of faces in every band   4.2/5.1/9.0 ms vs 17/56/198
  landmarks            A16W8   0.25 px in the 192 crop        0.5 ms vs 2.8
  YOLO26n-seg          A16W16  98.2% found, mask IoU 0.994    12.9 ms vs 90
  scene model          A16W16  98.9% of cells agree           15 ms vs 151
  MI-GAN               A16W16  41 dB from f32 in the fill     87 ms vs 488
  XFeat                int8    pano alignment 0.45 px (f32's own spread 0.41)  6.5 ms vs 58
  denoiser             A16W16  0.00 dB at every ISO            95 ms vs 1510 a tile
Face numbers are over public COCO val2017 photographs, not a library.

The APK carries the siblings (BUNDLED 15 -> 19; the old int8 detectors
removed), about 43 MB more. The Windows installer and its CI count skip
them; the Arch and Flatpak packages list their files and never had them.
The ladder example takes a role per model, which is how the per-role
forms above were seen landing on the NPU from the real probe.
2026-10-04 03:45:46 -04:00

113 lines
5.0 KiB
Rust
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
//! TRACES: FR-DEV-3i
//! Region segmentation for local masking (S15, docs/dev/segmentation.md).
//!
//! Local adjustments need to know where the image's regions are before they
//! can snap a mask to one. This crate is that map, and it is deliberately
//! **device-free**: the watershed's pixel passes live in `dr-gpu` because they
//! are shaders, and everything that reasons about *regions* rather than
//! *pixels* lives here, where it can be tested on hand-built inputs with no
//! adapter present (ARCH §6.5a).
//!
//! # The three arms
//!
//! [`hierarchy`] is **arm A** — a watershed over-segments the image and the
//! recorded merge order becomes a granularity ladder. Deterministic, needs no
//! model, works on any picture, and knows nothing about what anything *is*.
//!
//! [`semantic`] is **arm B** — a YOLO instance-segmentation model naming the
//! subjects it recognises. Knows what things are, and is vague about exactly
//! where their edges fall (its prototypes are quarter-resolution).
//!
//! [`prior`] is **arm C**, and it is the one that ships. Arm B's instances
//! *re-weight* arm A's merge order, so coarse levels of the ladder line up
//! with real objects while every boundary stays exactly where the watershed
//! put it. The model contributes what it is good at — knowing what things are
//! — and the watershed contributes what it is good at, which is knowing where
//! the edge is, to the pixel, at every scale.
//!
//! That combination is also what repairs the vocabulary problem. The shipped
//! model is COCO-trained, so it recognises subjects and has no class for sky,
//! foliage or wall (`models/LICENCE.md`). Selecting those falls to arm A,
//! which never needed a vocabulary to begin with.
//!
//! # And [`scene`], which is not one of the arms
//!
//! The three arms all serve *local* adjustment: they exist so a mask can be
//! snapped to one region of the picture. [`scene`] serves the opposite move —
//! one grade applied to every pixel of a category at once, sky or foliage or
//! water — and reads a second, ADE20K-trained model to do it. It shares this
//! crate because it shares the runtime and the letterbox, not because it is
//! another way of doing the same thing.
//!
//! [`refine`] is what the scene model needs and the instance model does not.
//! A category's weights come off an 80×80 grid, so one cell is twenty pixels
//! of a 1600px proxy and anything smaller than that — a flag in the sky, a
//! chimney, a bare branch — is averaged into whatever surrounds it. There is
//! no tiling answer here the way there is for an instance, because a category
//! has no bounding box to tile over. So the fix is the same one arm C makes:
//! the model says *what*, and the photograph's own pixels say *which* of them
//! belong to it.
//!
//! It shares no code with the arms and it is not a fourth one — it sharpens a
//! mask that already exists rather than proposing regions — but it is built on
//! [`distance`] for the same reason arm C is built on the watershed, which is
//! that the useful question is always "how far inside am I".
pub mod distance;
pub mod hierarchy;
pub mod prior;
pub mod refine;
#[cfg(feature = "semantic")]
pub mod scene;
#[cfg(feature = "semantic")]
pub mod semantic;
pub use distance::{signed_distance, Falloff, Morphology, Shaped};
pub use hierarchy::{Edge, Merge, MergeTree, RegionField};
pub use prior::{Membership, PriorOptions};
pub use refine::{
refine_category, RefineOptions, Refined, Refinement, SkipReason, STRICTNESS_DEFAULT,
STRICTNESS_MAX, STRICTNESS_OFF,
};
#[cfg(feature = "semantic")]
pub use scene::{Category, Scene, SceneModel};
#[cfg(feature = "embedded-model")]
pub use semantic::embedded_models;
#[cfg(feature = "semantic")]
pub use semantic::{Instance, SemanticModel, SemanticOptions, Tiling};
/// What can go wrong between an image and a region map.
#[derive(Debug, thiserror::Error)]
pub enum SegmentError {
#[error("could not read model file: {0}")]
ModelRead(#[source] std::io::Error),
#[cfg(feature = "semantic")]
#[error("inference failed: {0}")]
Inference(#[source] ort::Error),
#[error("image buffer is {got} floats, expected {expected} (RGB, three per pixel)")]
ImageShape { expected: usize, got: usize },
/// The graph produced something the decoder does not recognise — a
/// different model, or a different export of the same one.
#[error("model output '{0}' did not have the expected shape")]
OutputShape(&'static str),
/// `models/scene/categories.txt` and the model disagree, or the descriptor
/// is malformed. Its own variant rather than a parse error because every
/// case carries a specific sentence about what to fix.
#[error("category descriptor: {0}")]
CategoryDescriptor(String),
}
#[cfg(feature = "semantic")]
impl From<dr_inference_engine::Error> for SegmentError {
fn from(e: dr_inference_engine::Error) -> Self {
match e {
dr_inference_engine::Error::Inference(e) => SegmentError::Inference(e),
dr_inference_engine::Error::Io(e) => SegmentError::ModelRead(e),
}
}
}