Put the developer docs under docs/dev and index the folder for users first
docs/ had 26 developer documents flat beside the manual, and the two audiences are very differently sized: most readers want the manual and the gesture reference, a few want the register, the designs and the measurements. The manual and gestures.md stay at the top; everything for someone changing the code moves to docs/dev/, and the two documents that name their own successors — the v0.1 milestone and the UI-refinement plan — go to docs/dev/archive/ rather than being deleted, since both are still cited. docs/README.md is the index, users first. Every reference follows: code comments, Cargo manifests, the workflows, the pre-commit hook, the bench and traceability tools (which locate the repo root by docs/dev/requirements.md now), packaging, the Docker READMEs, CLAUDE.md, CONTRIBUTING.md and the README. The matrix links one level deeper and is regenerated. Links out of the moved documents into the tree gain a level; a link checker over every Markdown file finds none broken.
This commit is contained in:
@@ -315,7 +315,7 @@ fn full_library(
|
||||
|
||||
// The three phases, separately, because "a regroup takes n seconds" does
|
||||
// not tell anyone which half to optimise — and the answer differs between
|
||||
// a desktop and a tablet (docs/faces.md §9).
|
||||
// a desktop and a tablet (docs/dev/faces.md §9).
|
||||
{
|
||||
let dim = candidates.first().map(|c| c.embedding.len()).unwrap_or(0);
|
||||
let flat: Vec<f32> = candidates
|
||||
|
||||
@@ -82,7 +82,7 @@
|
||||
//! Grouping has no natural `subject_id`: it is a property of a *run* of frames,
|
||||
//! so a per-image job would rebuild the world once per photograph. It is
|
||||
//! therefore a debounced library-level pass, for exactly the reasons
|
||||
//! docs/catalog.md §10.2 gives for face clustering, and [`regroup`] is the whole
|
||||
//! docs/dev/catalog.md §10.2 gives for face clustering, and [`regroup`] is the whole
|
||||
//! of it — one ordered walk, no per-pair comparison beyond adjacent frames.
|
||||
//!
|
||||
//! # Grouping is not hiding
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
//! Face data as sealed shards, so a second device does not re-index the library.
|
||||
//!
|
||||
//! Indexing a 23,500-image library is on the order of two hours of CPU
|
||||
//! (docs/faces.md §12.2). It is also **byte-identical on every device**: the
|
||||
//! (docs/dev/faces.md §12.2). It is also **byte-identical on every device**: the
|
||||
//! same model over the same proxy produces the same embedding. Paying for it
|
||||
//! once per account rather than once per device is the whole point of this
|
||||
//! module, and it is the same bargain the thumbnail store already makes.
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
//! TRACES: FR-CULL-8 | FR-CULL-9 | FR-CULL-10 | FR-CULL-11 | FR-CULL-12 | NFR-SEC-5
|
||||
//! People and faces: what was detected, who it is, and who said so.
|
||||
//!
|
||||
//! The storage half of docs/faces.md. `dr-face` finds faces and turns them into
|
||||
//! The storage half of docs/dev/faces.md. `dr-face` finds faces and turns them into
|
||||
//! 512 numbers; this module is where those numbers acquire an identity, and
|
||||
//! where the user's corrections outrank the model's guesses.
|
||||
//!
|
||||
@@ -110,7 +110,7 @@ pub struct DetectedFace {
|
||||
/// Raw rather than unit length, so the length ([`Self::quality`]) is in
|
||||
/// the blob and not only beside it. Readers re-normalise on load.
|
||||
pub embedding: Vec<u8>,
|
||||
/// Source pixels across the aligned crop (docs/faces.md §7).
|
||||
/// Source pixels across the aligned crop (docs/dev/faces.md §7).
|
||||
pub crop_px: f32,
|
||||
/// Length of the raw embedding before normalisation — the model's own
|
||||
/// reading of how recognisable the crop was, and the gate on whether
|
||||
@@ -211,7 +211,7 @@ pub use dr_face::Calibration;
|
||||
/// the same face in the same photograph, for carrying an identity across a
|
||||
/// re-detection.
|
||||
///
|
||||
/// Set at the reference library's P≈0.95 line (docs/faces.md §9's table:
|
||||
/// Set at the reference library's P≈0.95 line (docs/dev/faces.md §9's table:
|
||||
/// 0.449), which is far above anything two different people in one frame
|
||||
/// reach and below what one face re-embedded from a better crop of itself
|
||||
/// does. The number is only ever asked about *overlapping* boxes on *one*
|
||||
@@ -2469,7 +2469,7 @@ mod tests {
|
||||
}
|
||||
|
||||
/// The reference implementation's fitted MBF curve puts the P=0.5 boundary
|
||||
/// at cosine 0.267 (docs/faces.md §1). Our own first end-to-end run scored
|
||||
/// at cosine 0.267 (docs/dev/faces.md §1). Our own first end-to-end run scored
|
||||
/// 0.596 between distinct photographs of one person and 0.05 between
|
||||
/// different people, so those two must land either side.
|
||||
#[test]
|
||||
|
||||
@@ -729,7 +729,7 @@ fn attached_has_table(conn: &Connection, schema: &str, table: &str) -> Result<bo
|
||||
/// # What travels, and what is recomputed
|
||||
///
|
||||
/// The rule this module already follows for the rest of the catalog: user
|
||||
/// judgements travel, inference is rebuilt. Concretely (docs/faces.md, and the
|
||||
/// judgements travel, inference is rebuilt. Concretely (docs/dev/faces.md, and the
|
||||
/// asymmetry `crate::faces` opens with):
|
||||
///
|
||||
/// - **People** — uuid, name, and whether the user set them aside. Merged by
|
||||
|
||||
@@ -16,7 +16,7 @@
|
||||
//! # The one thing a rebuild does not recover
|
||||
//!
|
||||
//! **Collections.** A manual collection is a set of images the user assembled
|
||||
//! by hand and nothing in the filesystem records it (`docs/catalog.md` §8.1) —
|
||||
//! by hand and nothing in the filesystem records it (`docs/dev/catalog.md` §8.1) —
|
||||
//! which is the whole reason the catalog file itself syncs. So the two offers
|
||||
//! are not interchangeable, and the interface must not present them as if they
|
||||
//! were: a restore keeps the user's collections, a rebuild does not.
|
||||
@@ -570,7 +570,7 @@ mod tests {
|
||||
// The first NFR-R6 branch, asserted on the thing that distinguishes it
|
||||
// from the second: a collection exists nowhere but the catalog, so it
|
||||
// is the evidence that the *contents* came back and not merely a
|
||||
// readable file (docs/catalog.md §8.1).
|
||||
// readable file (docs/dev/catalog.md §8.1).
|
||||
let dir = tempdir("restore");
|
||||
let path = dir.join("catalog.sqlite");
|
||||
fixture(&path, 500);
|
||||
|
||||
@@ -1009,7 +1009,7 @@ CREATE INDEX face_index_model ON face_index(model_id);
|
||||
|
||||
const V8: &str = r#"
|
||||
-- TRACES: FR-CULL-8 | FR-CULL-9 | FR-CULL-10 | FR-CULL-11 | FR-CULL-12 | NFR-SEC-5
|
||||
-- People and faces (docs/faces.md, docs/catalog.md §10).
|
||||
-- People and faces (docs/dev/faces.md, docs/dev/catalog.md §10).
|
||||
--
|
||||
-- Everything here is **derived data** except one column. Faces, landmarks,
|
||||
-- embeddings, cluster assignments and suggestions are all reproducible by
|
||||
@@ -1046,7 +1046,7 @@ CREATE TABLE faces (
|
||||
landmarks BLOB NOT NULL, -- 5 x (x, y) f32, normalised likewise
|
||||
detector_confidence REAL NOT NULL,
|
||||
embedding BLOB NOT NULL, -- 512 x f16; unit length until V14, raw since
|
||||
-- Source pixels across the aligned 112x112 crop (docs/faces.md §7).
|
||||
-- Source pixels across the aligned 112x112 crop (docs/dev/faces.md §7).
|
||||
--
|
||||
-- Not cosmetic: it is the honest quality signal for the UI, a feature in
|
||||
-- the §8 calibration -- FR-CULL-9 names face size as an axis along which an
|
||||
|
||||
@@ -11,7 +11,7 @@ log.workspace = true
|
||||
|
||||
# Inference. `ort` is the API; **what runs it is `dr-inference-engine`'s
|
||||
# business** — tract, or an ONNX Runtime the app found on disk, on whichever
|
||||
# provider the device has (docs/inference.md). This crate never names either.
|
||||
# provider the device has (docs/dev/inference.md). This crate never names either.
|
||||
ort = { workspace = true, optional = true }
|
||||
dr-inference-engine = { workspace = true, optional = true }
|
||||
ndarray = { workspace = true, optional = true }
|
||||
@@ -37,7 +37,7 @@ required-features = ["inference"]
|
||||
|
||||
[features]
|
||||
# Nothing on by default, and in particular **no `embedded-model`**: the weights
|
||||
# are not a build input and never become one (docs/faces.md §2.2). A feature
|
||||
# are not a build input and never become one (docs/dev/faces.md §2.2). A feature
|
||||
# flag that *could* embed them is a flag someone eventually sets in a packaging
|
||||
# script, and the InsightFace grant does not survive that.
|
||||
default = []
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
//! Detect the faces in a JPEG and read each one's eyes (docs/faces.md §17).
|
||||
//! Detect the faces in a JPEG and read each one's eyes (docs/dev/faces.md §17).
|
||||
//!
|
||||
//! The thing worth looking at is whether the eye boxes land on eyes and
|
||||
//! whether soft ones are refused — so with `--dump DIR` the crops the
|
||||
|
||||
@@ -7,7 +7,7 @@
|
||||
//! DET.onnx EMB.onnx photo.jpg [photo.jpg ...]
|
||||
//!
|
||||
//! The models must have had their input dims frozen first; see
|
||||
//! `tools/fix-face-model-shapes.sh` and docs/faces.md §12 M1.
|
||||
//! `tools/fix-face-model-shapes.sh` and docs/dev/faces.md §12 M1.
|
||||
|
||||
use std::time::Instant;
|
||||
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
//! M1 (docs/faces.md §12) — will tract load these graphs at all?
|
||||
//! M1 (docs/dev/faces.md §12) — will tract load these graphs at all?
|
||||
//!
|
||||
//! The one measurement everything else in the face subsystem is conditional
|
||||
//! on. `det_500m.onnx` has a dynamic H/W input, which is exactly what tract
|
||||
|
||||
@@ -9,7 +9,7 @@
|
||||
//!
|
||||
//! # What it is for
|
||||
//!
|
||||
//! docs/faces.md §9 has the desktop numbers and the question they leave open:
|
||||
//! docs/dev/faces.md §9 has the desktop numbers and the question they leave open:
|
||||
//! a GPU GEMM is worth roughly 1.5× of a regroup on a twenty-core desktop,
|
||||
//! because the scan is under a third of the pass there. On a tablet the CPU is
|
||||
//! several times slower and the GPU is not, so the same optimisation is worth
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
//! Five-point face alignment (docs/faces.md §5).
|
||||
//! Five-point face alignment (docs/dev/faces.md §5).
|
||||
//!
|
||||
//! ArcFace embeddings are trained on faces warped to a canonical 112×112
|
||||
//! arrangement. Feeding the model a plain bounding-box crop *works* — it
|
||||
@@ -208,7 +208,7 @@ impl Similarity {
|
||||
///
|
||||
/// # Why least squares and not RANSAC
|
||||
///
|
||||
/// The reference C++ implementation (docs/faces.md §1.1) fits this with
|
||||
/// The reference C++ implementation (docs/dev/faces.md §1.1) fits this with
|
||||
/// OpenCV's `estimateAffinePartial2D` under RANSAC. RANSAC over five points is
|
||||
/// a strange fit: the minimal sample for a similarity is two, so it can discard
|
||||
/// landmarks it judges outliers and solve from a subset — and on a profile face
|
||||
@@ -427,7 +427,7 @@ fn sample_window(
|
||||
// ── eyes ──────────────────────────────────────────────────────────────────
|
||||
|
||||
/// Width of an eye crop as the classifier reads it, in pixels. Fixed by the
|
||||
/// OCEC input (`docs/faces.md` §17): 40 wide, 24 high.
|
||||
/// OCEC input (`docs/dev/faces.md` §17): 40 wide, 24 high.
|
||||
pub const EYE_PATCH_WIDTH: usize = 40;
|
||||
/// Height of an eye crop as the classifier reads it, in pixels.
|
||||
pub const EYE_PATCH_HEIGHT: usize = 24;
|
||||
@@ -438,7 +438,7 @@ pub const EYE_PATCH_HEIGHT: usize = 24;
|
||||
/// The classifier was trained on a whole-body detector's *eye* boxes — tight
|
||||
/// round the palpebral fissure — and measured on 25 open-eyed faces from the
|
||||
/// reference library, a tight box is what it wants: 22 of 25 read open at
|
||||
/// 0 and 0.1, 18 at 0.4, 14 at 0.6 (docs/faces.md §17.2). A tenth, so a
|
||||
/// 0 and 0.1, 18 at 0.4, 14 at 0.6 (docs/dev/faces.md §17.2). A tenth, so a
|
||||
/// contour landing a pixel short of the lashes still holds them.
|
||||
pub const EYE_BOX_MARGIN: f32 = 0.1;
|
||||
|
||||
@@ -571,7 +571,7 @@ pub const SUNGLASSES_EDGE: usize = 48;
|
||||
/// clear glasses, at 0.68. Erring towards "sunglasses" is the safe direction
|
||||
/// for what this feeds: a face called sunglasses is left alone by the
|
||||
/// eyes-open filter, where a pair of sunglasses missed hands the eye
|
||||
/// classifier a lens to guess at (docs/faces.md §17).
|
||||
/// classifier a lens to guess at (docs/dev/faces.md §17).
|
||||
pub const SUNGLASSES_WINDOWS: [(f32, f32, f32, f32); 2] =
|
||||
[(0.0, 0.0, 112.0, 112.0), (-5.0, -14.0, 122.0, 122.0)];
|
||||
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
//! Cosine to probability (docs/faces.md §8, FR-CULL-9).
|
||||
//! Cosine to probability (docs/dev/faces.md §8, FR-CULL-9).
|
||||
//!
|
||||
//! FR-CULL-9 is a hard requirement rather than an implementation detail: no
|
||||
//! code path may threshold a bare cosine, every threshold in the subsystem is
|
||||
@@ -28,7 +28,7 @@
|
||||
//! calibration to the belief it was supposed to test — and that is the whole
|
||||
//! of the alternative.
|
||||
//!
|
||||
//! docs/faces.md §8.1 names one more that would cost no labelling at all: two
|
||||
//! docs/dev/faces.md §8.1 names one more that would cost no labelling at all: two
|
||||
//! faces in adjacent frames of one burst are near-certainly the same person,
|
||||
//! and FR-CULL-5's grouping is sitting there. Nothing draws on it. This crate
|
||||
//! cannot see a catalog, let alone the bursts in one — it is handed cosines by
|
||||
@@ -83,7 +83,7 @@ pub struct Calibration {
|
||||
}
|
||||
|
||||
impl Default for Calibration {
|
||||
/// The reference implementation's fitted MBF curve (docs/faces.md §1):
|
||||
/// The reference implementation's fitted MBF curve (docs/dev/faces.md §1):
|
||||
/// steepness 16.2, P=0.5 at cosine 0.267.
|
||||
///
|
||||
/// **`valid` is false**, and that is the point. It is a documented
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
//! TRACES: FR-CULL-8a
|
||||
//! The two small classifiers behind a face's eye state (docs/faces.md §17).
|
||||
//! The two small classifiers behind a face's eye state (docs/dev/faces.md §17).
|
||||
//!
|
||||
//! **OCEC** — *open closed eyes classification*, Hyodo 2025 — reads one
|
||||
//! 40×24 eye and answers P(open). **SGC** — *sunglasses classification*,
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
//! Grouping faces into people (docs/faces.md §9, FR-CULL-10).
|
||||
//! Grouping faces into people (docs/dev/faces.md §9, FR-CULL-10).
|
||||
//!
|
||||
//! Model-free: this is arithmetic over embeddings, and it is where the
|
||||
//! subsystem's accuracy actually lives, so it is testable with no weights on
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
//! SCRFD face detection (docs/faces.md §4).
|
||||
//! SCRFD face detection (docs/dev/faces.md §4).
|
||||
//!
|
||||
//! One forward pass produces a box, a confidence and **five landmarks** per
|
||||
//! face — the landmarks being the reason for this detector rather than a
|
||||
@@ -138,7 +138,7 @@ impl Detection {
|
||||
pub struct Detector {
|
||||
session: Model,
|
||||
/// f32 or int8 — the int8 form finds a different set of faces and is a
|
||||
/// different detector in `model_id` (docs/inference.md §7).
|
||||
/// different detector in `model_id` (docs/dev/inference.md §7).
|
||||
form: Form,
|
||||
/// Feature-map count: 3 for strides {8,16,32}, 4 for {8,16,32,64}.
|
||||
///
|
||||
@@ -346,7 +346,7 @@ fn iou(a: &(f32, f32, f32, f32), b: &(f32, f32, f32, f32)) -> f32 {
|
||||
/// How the image is fitted into the graph's fixed square input.
|
||||
///
|
||||
/// The forward and inverse mappings live in one struct on purpose:
|
||||
/// docs/faces.md §4.1 notes that what matters is not *where* the padding goes
|
||||
/// docs/dev/faces.md §4.1 notes that what matters is not *where* the padding goes
|
||||
/// but that the two agree. A mismatch offsets every box and landmark by the
|
||||
/// padding, producing detections that look plausible and embeddings that
|
||||
/// quietly cluster badly three stages later.
|
||||
@@ -372,7 +372,7 @@ impl Letterbox {
|
||||
///
|
||||
/// `(x·255 − 127.5) / 128` — note `/128`, not `/127.5`. The reference
|
||||
/// implementation this is ported from uses `/128` for both models, and
|
||||
/// every measured number in docs/faces.md §1 came from it.
|
||||
/// every measured number in docs/dev/faces.md §1 came from it.
|
||||
///
|
||||
/// Padding is grey, matching the reference's `114`: the value the network
|
||||
/// reads least as an edge, where black would draw a hard border across the
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
//! ArcFace / MobileFaceNet inference (docs/faces.md §6).
|
||||
//! ArcFace / MobileFaceNet inference (docs/dev/faces.md §6).
|
||||
//!
|
||||
//! Takes an aligned crop and returns 512 L2-normalised floats. The alignment is
|
||||
//! not optional and cannot be skipped by accident: [`Embedder::embed`] takes an
|
||||
@@ -66,7 +66,7 @@ impl Embedder {
|
||||
|
||||
pub fn from_bytes(bytes: &[u8], model: ModelId) -> Result<Self, FaceError> {
|
||||
// Always the f32 form: an embedding must compare across devices
|
||||
// (docs/inference.md §7), and the engine pins this role to it.
|
||||
// (docs/dev/inference.md §7), and the engine pins this role to it.
|
||||
let loaded = dr_inference_engine::open(Role::Embedder, Form::F32, bytes)?;
|
||||
let acquired = loaded.acquire()?;
|
||||
let session = acquired.lock();
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
//! What an embedder produces, and how it is stored (docs/faces.md §6).
|
||||
//! What an embedder produces, and how it is stored (docs/dev/faces.md §6).
|
||||
//!
|
||||
//! Deliberately **model-free**: the vector, its identity, its comparison and
|
||||
//! its storage encoding are arithmetic, and `calibrate` and `cluster` are built
|
||||
@@ -273,7 +273,7 @@ mod tests {
|
||||
);
|
||||
}
|
||||
|
||||
/// The claim docs/faces.md §6 makes about the storage format: the f16
|
||||
/// The claim docs/dev/faces.md §6 makes about the storage format: the f16
|
||||
/// round-trip costs ~1e-3 of cosine, three orders below the separation
|
||||
/// between a match and a non-match.
|
||||
#[test]
|
||||
|
||||
@@ -67,14 +67,14 @@ pub const SUNGLASSES_THRESHOLD: f32 = 0.5;
|
||||
/// The classifier was trained on eyes down to about a dozen pixels wide
|
||||
/// (its reference footage averaged 15–21); below that the 40-pixel patch is
|
||||
/// an interpolation of nothing, and the answer is noise that reads as
|
||||
/// "closed". docs/faces.md §17.3 has the measurement behind the number.
|
||||
/// "closed". docs/dev/faces.md §17.3 has the measurement behind the number.
|
||||
pub const MIN_EYE_PX: f32 = 12.0;
|
||||
|
||||
/// Least [`Eye::sharpness`] for the eye to be read.
|
||||
///
|
||||
/// The same measure as the face's `min_sharpness`, over the eye patch, and
|
||||
/// chosen the same way: the value under which the open-eyed faces of the
|
||||
/// reference sample were being called closed. docs/faces.md §17.3.
|
||||
/// reference sample were being called closed. docs/dev/faces.md §17.3.
|
||||
pub const MIN_EYE_SHARPNESS: f32 = 0.02;
|
||||
|
||||
/// An eye narrower than this fraction of its partner is the far eye of a
|
||||
@@ -82,7 +82,7 @@ pub const MIN_EYE_SHARPNESS: f32 = 0.02;
|
||||
///
|
||||
/// A landmark model's contour for a hidden eye collapses towards the nose.
|
||||
/// Measured on twenty native renders of the reference library
|
||||
/// (docs/faces.md §17.4): profiles put the far eye at 0.02–0.43 of the near
|
||||
/// (docs/dev/faces.md §17.4): profiles put the far eye at 0.02–0.43 of the near
|
||||
/// one, two three-quarter faces whose far eye read closed sat at 0.54, and
|
||||
/// every face looking at the camera — winks included, since a shut eye's
|
||||
/// box keeps its width — sat at 0.78 or more. 0.6 splits the gap.
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
//! TRACES: FR-CULL-8a
|
||||
//! Dense facial landmarks — InsightFace's `2d106det` (docs/faces.md §17.2).
|
||||
//! Dense facial landmarks — InsightFace's `2d106det` (docs/dev/faces.md §17.2).
|
||||
//!
|
||||
//! SCRFD's five points place a face; they do not place an eye. Its eye
|
||||
//! point is loose enough that a window centred on it left the eye in a
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
//! Faces and identity (S14, docs/faces.md).
|
||||
//! Faces and identity (S14, docs/dev/faces.md).
|
||||
//!
|
||||
//! Two models, run over the native render, producing per face a box, five
|
||||
//! landmarks, a confidence and a 512-d embedding (FR-CULL-8) — and then the
|
||||
@@ -21,9 +21,9 @@
|
||||
//! for a packaging script to switch on. The application obtains a model at
|
||||
//! runtime; this crate takes bytes and never fetches anything.
|
||||
//!
|
||||
//! docs/faces.md §2 is the full reading, including what would have to change
|
||||
//! docs/dev/faces.md §2 is the full reading, including what would have to change
|
||||
//! for that to stop being true. The eye-state models are the exception: MIT,
|
||||
//! weights and all, and shipped in `models/face/` (docs/faces.md §17).
|
||||
//! weights and all, and shipped in `models/face/` (docs/dev/faces.md §17).
|
||||
//!
|
||||
//! # Why the runtime is split behind a feature
|
||||
//!
|
||||
@@ -55,7 +55,7 @@ pub mod references;
|
||||
/// Smallest long edge a face crop may be sampled from.
|
||||
///
|
||||
/// **A floor on the crop source, not on the detector input.** The distinction
|
||||
/// is the whole of FR-CULL-8 and `docs/faces.md` §7: detection letterboxes
|
||||
/// is the whole of FR-CULL-8 and `docs/dev/faces.md` §7: detection letterboxes
|
||||
/// every buffer into 640×640, so its input resolution decides nothing, while
|
||||
/// [`warp`] samples the 112×112 the embedder sees and so converts source
|
||||
/// resolution directly into embedding quality. FR-CULL-8 requires that crop to
|
||||
|
||||
@@ -115,7 +115,7 @@ pub struct Faces<'a> {
|
||||
/// Source pixels across the aligned crop, for the calibration's size term.
|
||||
pub crop_px: &'a [f32],
|
||||
/// Which photograph each face came from. Two faces in one frame are not
|
||||
/// the same person, so those pairs are never returned (docs/faces.md §9).
|
||||
/// the same person, so those pairs are never returned (docs/dev/faces.md §9).
|
||||
pub images: &'a [u64],
|
||||
/// Which faces may be compared *against* — the gallery
|
||||
/// ([`crate::embedding::MIN_GALLERY_QUALITY`]).
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
//! What a frame actually costs — the measurement FR-DSP-2 is waiting on.
|
||||
//!
|
||||
//! `docs/display-and-extension.md` §2 argues that tiled computation predates
|
||||
//! `docs/dev/display-and-extension.md` §2 argues that tiled computation predates
|
||||
//! the fused-shader design and may not need to exist: the composer folds every
|
||||
//! active operation into **one dispatch over a viewport-sized target**, so the
|
||||
//! problem tiles were invented to solve may already be solved. That argument
|
||||
@@ -28,7 +28,7 @@
|
||||
//! the per-frame CPU half is dominated by shader-source assembly, which is
|
||||
//! string formatting and is several times slower unoptimised.
|
||||
//!
|
||||
//! The committed numbers live in `docs/frame-budget.md`. Rerun this and diff
|
||||
//! The committed numbers live in `docs/dev/frame-budget.md`. Rerun this and diff
|
||||
//! that file; a regression should be a diff rather than somebody's memory.
|
||||
//!
|
||||
//! # Why the 99th percentile and not the mean
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
//! Segment an image and write the granularity ladder as false-coloured PPMs.
|
||||
//!
|
||||
//! The whole point of S15 step 2 (docs/segmentation.md §11): look at the
|
||||
//! The whole point of S15 step 2 (docs/dev/segmentation.md §11): look at the
|
||||
//! ladder and decide whether clicking through it would land on the things a
|
||||
//! person means. No amount of design settles that — the pictures do.
|
||||
//!
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
//! Watershed segmentation — arm A's GPU half (S15, docs/segmentation.md).
|
||||
//! Watershed segmentation — arm A's GPU half (S15, docs/dev/segmentation.md).
|
||||
//!
|
||||
//! Runs the five passes in `shaders/watershed.wgsl` over a demosaiced image
|
||||
//! and leaves a basin label per pixel on the GPU. The hierarchy built from
|
||||
@@ -20,7 +20,7 @@
|
||||
//! one AC-8 forbids is per frame in the render loop, and sharing a switch
|
||||
//! would force a build wanting local masking to unlock the other.
|
||||
//!
|
||||
//! It is still a real cost and still unfinished. F3 in docs/segmentation.md
|
||||
//! It is still a real cost and still unfinished. F3 in docs/dev/segmentation.md
|
||||
//! §12 stands: the adjacency accumulation belongs GPU-side with atomics, and
|
||||
//! until it moves there every segmentation pays a full-resolution transfer.
|
||||
//! Read the feature name as a description of a known gap rather than as
|
||||
@@ -36,7 +36,7 @@ pub struct SegmentOptions {
|
||||
/// Longest proxy edge. The segmentation runs here, not at sensor
|
||||
/// resolution: a 24 MP watershed costs 12× the memory to place boundaries
|
||||
/// a person cannot see, and the boundary refinement that matters at 1:1
|
||||
/// is a separate stage (docs/segmentation.md §4).
|
||||
/// is a separate stage (docs/dev/segmentation.md §4).
|
||||
pub max_edge: u32,
|
||||
/// Pre-smoothing radius in proxy pixels. The caller's to raise with ISO —
|
||||
/// this is the single knob that decides whether a noisy file segments
|
||||
@@ -69,7 +69,7 @@ impl Default for SegmentOptions {
|
||||
w_chroma: 0.5,
|
||||
// **Zero: the pass is off.** It is implemented, dispatched
|
||||
// correctly and measurably changes nothing — see the ignored test
|
||||
// below and §12 of docs/segmentation.md. Until that is understood,
|
||||
// below and §12 of docs/dev/segmentation.md. Until that is understood,
|
||||
// running it would buy 64 dispatches per segmentation and no
|
||||
// improvement, so the default declines to pay.
|
||||
plateau_iterations: 0,
|
||||
@@ -486,7 +486,7 @@ impl Segmentation {
|
||||
/// a region graph of a few thousand nodes that every later interaction
|
||||
/// reads from the CPU anyway.
|
||||
///
|
||||
/// What it is *not* is finished. F3 in docs/segmentation.md §12 stands:
|
||||
/// What it is *not* is finished. F3 in docs/dev/segmentation.md §12 stands:
|
||||
/// the adjacency accumulation belongs on the GPU with atomics, and until
|
||||
/// it moves there a segmentation costs one full-resolution transfer of the
|
||||
/// label and gradient buffers. That is a real cost on a phone and the
|
||||
@@ -724,7 +724,7 @@ mod tests {
|
||||
px
|
||||
}
|
||||
#[test]
|
||||
#[ignore = "the plateau pass is a measured no-op; see docs/segmentation.md §12"]
|
||||
#[ignore = "the plateau pass is a measured no-op; see docs/dev/segmentation.md §12"]
|
||||
fn lower_completion_drains_a_plateau_instead_of_shattering_it() {
|
||||
// F1, asserted rather than eyeballed, and asserted at the level where
|
||||
// it matters.
|
||||
@@ -741,7 +741,7 @@ mod tests {
|
||||
// with no exit anywhere — cannot be drained by a distance that has
|
||||
// nowhere to descend to, and collapsing it fully would need connected
|
||||
// component labelling rather than a local rule. It is not worth it:
|
||||
// see docs/segmentation.md §12.
|
||||
// see docs/dev/segmentation.md §12.
|
||||
let Some(ctx) = ctx() else { return };
|
||||
let (w, h) = (96u32, 96u32);
|
||||
let src = DemosaicedImage::from_rgba8(&ctx, &ramp(w, h), w, h).expect("source");
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
// Watershed segmentation — the passes behind arm A of S15 (docs/segmentation.md).
|
||||
// Watershed segmentation — the passes behind arm A of S15 (docs/dev/segmentation.md).
|
||||
//
|
||||
// Seven entry points forming one chain:
|
||||
//
|
||||
@@ -197,7 +197,7 @@ fn gradient(@builtin(global_invocation_id) gid: vec3<u32>) {
|
||||
// lowest-indexed neighbour, which is up and to the left. Each pixel therefore
|
||||
// walks diagonally until it falls off the plateau, and one flat region becomes
|
||||
// a fan of diagonal chains rather than one basin — visible as hatching across
|
||||
// what should be a single area (docs/segmentation.md §12, F1).
|
||||
// what should be a single area (docs/dev/segmentation.md §12, F1).
|
||||
//
|
||||
// The fix is the standard lower-completion: give each plateau pixel its
|
||||
// geodesic distance to the nearest pixel that *does* have a lower neighbour,
|
||||
|
||||
@@ -2,11 +2,11 @@
|
||||
//!
|
||||
//! FR-DSP-3 says a slider updates the visible region within one frame budget at
|
||||
//! proxy resolution. Until this file existed nothing checked it, which made it
|
||||
//! a wish — `docs/display-and-extension.md` §3 is blunt about that, and §7 is
|
||||
//! a wish — `docs/dev/display-and-extension.md` §3 is blunt about that, and §7 is
|
||||
//! blunt about what tagging an unchecked requirement does to the coverage
|
||||
//! figure.
|
||||
//!
|
||||
//! The measurements this guards are in [`docs/frame-budget.md`], produced by
|
||||
//! The measurements this guards are in [`docs/dev/frame-budget.md`], produced by
|
||||
//! `examples/frame_budget.rs`. This file is the part of them that has to keep
|
||||
//! being true: it renders the **whole point-operation chain** through the real
|
||||
//! `render_detailed` for a hundred frames, moving a slider between each, and
|
||||
@@ -17,7 +17,7 @@
|
||||
//! **The neighbourhood stage is deliberately not in the asserted chain.** It is
|
||||
//! over the budget today — clarity alone is 34 ms at 4K, because its kernel is
|
||||
//! a fraction of the frame and reaches a 52-pixel radius there — and
|
||||
//! `docs/frame-budget.md` records that, names the fix (a base computed at
|
||||
//! `docs/dev/frame-budget.md` records that, names the fix (a base computed at
|
||||
//! reduced resolution) and does not pretend otherwise. Asserting a budget the
|
||||
//! code does not meet would produce a red suite that everyone learns to ignore;
|
||||
//! asserting it on a chain that quietly excluded the expensive stage *without
|
||||
@@ -81,7 +81,7 @@ const SOURCE: (u32, u32) = (6000, 4000);
|
||||
/// The viewport the budget is asserted at: a 16:10 desktop display.
|
||||
///
|
||||
/// Not 4K, and the reason is worth stating. At 4K the fused chain still passes
|
||||
/// with room to spare (4.5 ms of GPU; see `docs/frame-budget.md`), but a test
|
||||
/// with room to spare (4.5 ms of GPU; see `docs/dev/frame-budget.md`), but a test
|
||||
/// that renders 8.3 M pixels a hundred times twice over is four seconds of
|
||||
/// suite time to re-establish a conclusion 4.1 M pixels already establishes.
|
||||
const VIEWPORT: (u32, u32) = (2560, 1600);
|
||||
@@ -166,7 +166,7 @@ impl Run {
|
||||
judged <= BUDGET_MS,
|
||||
"{case} at {}x{}: p99 of {FRAMES} frames was {judged:.2} ms, over the \
|
||||
{BUDGET_MS:.0} ms budget (cpu {:.2} ms, gpu {:.2} ms, total {:.2} ms). \
|
||||
FR-DSP-3 is what this violates; docs/frame-budget.md holds the \
|
||||
FR-DSP-3 is what this violates; docs/dev/frame-budget.md holds the \
|
||||
numbers it used to be.",
|
||||
viewport.0,
|
||||
viewport.1,
|
||||
|
||||
@@ -326,7 +326,7 @@ fn a_proxy_and_an_export_agree_about_the_effect() {
|
||||
|
||||
#[test]
|
||||
fn crossing_the_reduction_threshold_does_not_change_the_picture() {
|
||||
// TRACES: FR-DSP-3 — `docs/technical-debt.md` TD-4, held in pixels.
|
||||
// TRACES: FR-DSP-3 — `docs/dev/technical-debt.md` TD-4, held in pixels.
|
||||
//
|
||||
// Clarity's base is computed on a reduced grid, and how reduced depends on
|
||||
// the viewport: `LocalContrast::reduction` steps 4 -> 2 -> 1 as sigma
|
||||
|
||||
@@ -10,7 +10,7 @@
|
||||
//! code path — the zoom is the full-resolution path — which is why the
|
||||
//! requirement has been satisfied for some time without anyone tagging it.
|
||||
//!
|
||||
//! `docs/display-and-extension.md` §7 is the reason this file exists rather
|
||||
//! `docs/dev/display-and-extension.md` §7 is the reason this file exists rather
|
||||
//! than a tag on `framing.rs`: a requirement counts as covered when a `TRACES`
|
||||
//! comment names it, and nothing checks that the code under the tag does the
|
||||
//! thing. `FR-DEV-8` is tagged against plumbing a future operation would use.
|
||||
|
||||
@@ -6,7 +6,7 @@ rust-version.workspace = true
|
||||
license.workspace = true
|
||||
|
||||
# The one crate that names a runtime, a provider, a vendor library or a
|
||||
# device (docs/inference.md §8). `dr-face` and `dr-segment` ask it for a
|
||||
# device (docs/dev/inference.md §8). `dr-face` and `dr-segment` ask it for a
|
||||
# session by role and never see which of these answered.
|
||||
|
||||
[dependencies]
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
//! The API table `ort` runs on, chosen once (docs/inference.md §3).
|
||||
//! The API table `ort` runs on, chosen once (docs/dev/inference.md §3).
|
||||
//!
|
||||
//! `ort` with `alternative-backend` links no runtime and asks, on first use,
|
||||
//! for an `OrtApi` — a struct of function pointers. Two things can fill it:
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
//! Compiled engines: what a rung builds once per device, and the thread that
|
||||
//! builds them before anyone asks (docs/inference.md §5, §6).
|
||||
//! builds them before anyone asks (docs/dev/inference.md §5, §6).
|
||||
//!
|
||||
//! TensorRT keeps its own engine cache keyed by graph hash; QNN writes a
|
||||
//! context model. Both are opaque to this crate, which tracks only *that* a
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
//! Which runtime, which provider and which model form — decided once per
|
||||
//! device, and the only crate that knows the answer (docs/inference.md).
|
||||
//! device, and the only crate that knows the answer (docs/dev/inference.md).
|
||||
//!
|
||||
//! Consumers ask for a session by [`Role`] and get `ort`'s `Session` back;
|
||||
//! what built it — tract on one core, ONNX Runtime's CPU pool, a TensorRT
|
||||
@@ -35,13 +35,13 @@ pub enum Role {
|
||||
Embedder,
|
||||
Segmenter,
|
||||
Scene,
|
||||
/// The dense landmark model behind the eye reading (docs/faces.md §7c).
|
||||
/// The dense landmark model behind the eye reading (docs/dev/faces.md §7c).
|
||||
Landmarks,
|
||||
/// The eye-state and sunglasses classifiers, a few hundred kilobytes.
|
||||
EyeClassifier,
|
||||
/// XFeat, the panorama keypoint detector (docs/panorama.md).
|
||||
/// XFeat, the panorama keypoint detector (docs/dev/panorama.md).
|
||||
Keypoints,
|
||||
/// MI-GAN, the panorama border filler (docs/panorama.md §12). Plain
|
||||
/// MI-GAN, the panorama border filler (docs/dev/panorama.md §12). Plain
|
||||
/// convolutions, so any rung serves it; fp16 on TensorRT and int8 on
|
||||
/// the Hexagon are the point of it.
|
||||
Inpainter,
|
||||
@@ -516,7 +516,7 @@ mod tests {
|
||||
|
||||
/// The smallest shipped graph, if this checkout has the weights; a test
|
||||
/// suite that needs a research-licensed download is one that does not
|
||||
/// run in CI (docs/faces.md §3), so absence is a skip.
|
||||
/// run in CI (docs/dev/faces.md §3), so absence is a skip.
|
||||
fn probe_bytes() -> Option<Vec<u8>> {
|
||||
let path = concat!(
|
||||
env!("CARGO_MANIFEST_DIR"),
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
//! Walk the ladder, once, by building real sessions (docs/inference.md §4).
|
||||
//! Walk the ladder, once, by building real sessions (docs/dev/inference.md §4).
|
||||
//!
|
||||
//! A rung is taken when a session builds on it, runs, and is faster than
|
||||
//! the floor. Both halves matter: a provider can register and then fail at
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
//! One session builder per rung (docs/inference.md §2, §7, §9).
|
||||
//! One session builder per rung (docs/dev/inference.md §2, §7, §9).
|
||||
|
||||
use ort::session::Session;
|
||||
|
||||
|
||||
@@ -13,7 +13,7 @@ log.workspace = true
|
||||
|
||||
# Inference for the learned keypoint detector, on the same footing as
|
||||
# `dr-segment`: `ort` is the API, `dr-inference-engine` decides what runs
|
||||
# it (docs/inference.md), and both are optional so that the geometry —
|
||||
# it (docs/dev/inference.md), and both are optional so that the geometry —
|
||||
# matching, the rotation solve, the projections — is a dependency-free crate
|
||||
# that tests without a model.
|
||||
ort = { workspace = true, optional = true }
|
||||
|
||||
@@ -95,7 +95,7 @@ pub struct Params {
|
||||
/// and, beyond the band being filled, still unknown. That is what the
|
||||
/// shipped model was trained on (a fine-tune of MI-GAN on voids cut
|
||||
/// from photographs the way a cylindrical merge cuts them, see
|
||||
/// `docs/panorama.md` §14); a ring would give it a fold to continue.
|
||||
/// `docs/dev/panorama.md` §14); a ring would give it a fold to continue.
|
||||
///
|
||||
/// Non-zero is the stock model's crutch: a plain reflection of a deep
|
||||
/// hole pulls in whatever is that far from the edge — a ridge, a peak —
|
||||
|
||||
@@ -9,7 +9,7 @@
|
||||
//! on every rung, and what they cost is the whole story of whether a fill
|
||||
//! is interactive: 7.4 s a tile under tract, 0.4 s under ONNX Runtime's
|
||||
//! CPU pool, 23 ms in fp16 and 13 ms in int8 on a laptop's TensorRT
|
||||
//! (2026-09-19, docs/panorama.md §12).
|
||||
//! (2026-09-19, docs/dev/panorama.md §12).
|
||||
//!
|
||||
//! The model's contract, from the reference `export_inference_model.py`:
|
||||
//! input `1×4×512×512` float — channel 0 is `mask − 0.5` with 1 where the
|
||||
|
||||
@@ -4,7 +4,7 @@
|
||||
//! Apache-2.0 weights (`models/LICENCE.md`), exported at a fixed shape by
|
||||
//! `tools/export-xfeat.sh` and loaded through the same `dr-inference-engine`
|
||||
//! `dr-segment` and `dr-face` use, so this adds no runtime and no C to the
|
||||
//! tree; what runs it is the device's business (docs/inference.md). ~300 ms
|
||||
//! tree; what runs it is the device's business (docs/dev/inference.md). ~300 ms
|
||||
//! per frame on tract on the reference desktop, ~400 ms on the tablet
|
||||
//! (S15.2, S15.4).
|
||||
|
||||
@@ -37,7 +37,7 @@ pub struct XFeat {
|
||||
}
|
||||
|
||||
/// The bytes of both exports compiled into the binary, for whoever compiles
|
||||
/// engines ahead of the first request (docs/inference.md §6).
|
||||
/// engines ahead of the first request (docs/dev/inference.md §6).
|
||||
#[cfg(feature = "embedded-model")]
|
||||
pub fn embedded_model_bytes() -> [&'static [u8]; 2] {
|
||||
[EMBEDDED_LANDSCAPE, EMBEDDED_PORTRAIT]
|
||||
|
||||
@@ -315,7 +315,7 @@ pub struct DetailPass {
|
||||
/// 52 render pixels at 4K — holds no spatial frequency a quarter-scale
|
||||
/// grid cannot represent. Computing it at the render size therefore buys
|
||||
/// nothing and costs everything: 105 taps over 8.3 M pixels, twice, which
|
||||
/// measured at 34 ms and is where `docs/technical-debt.md` TD-4 came from.
|
||||
/// measured at 34 ms and is where `docs/dev/technical-debt.md` TD-4 came from.
|
||||
/// At a quarter it is a sixteenth of the pixels at a quarter of the
|
||||
/// radius, and the result is not an approximation of the full-resolution
|
||||
/// base — it is the same band-limited function, sampled where it is still
|
||||
@@ -568,7 +568,7 @@ pub fn compose_detail(
|
||||
/// photograph the photographer thinks they are sharpening.
|
||||
///
|
||||
/// It also means ARCH §5.2's stage list, which draws spot removal after
|
||||
/// texture and clarity, is not what this does — see `docs/spot-removal.md`
|
||||
/// texture and clarity, is not what this does — see `docs/dev/spot-removal.md`
|
||||
/// §5.1, which is where the disagreement is written down.
|
||||
pub fn compose_detail_with(
|
||||
ops: &[Box<dyn Operation>],
|
||||
|
||||
@@ -113,7 +113,7 @@ pub struct EditGraph {
|
||||
/// a sidecar comes to name one stock while the shader draws another.
|
||||
film: Option<Film>,
|
||||
/// TRACES: FR-DEV-8
|
||||
/// The repairs (`docs/spot-removal.md`).
|
||||
/// The repairs (`docs/dev/spot-removal.md`).
|
||||
///
|
||||
/// Apart from `ops` for the third time and the same reason: a spot is not
|
||||
/// a scalar, and a list of them is not a slider. It sits beside the masks
|
||||
@@ -122,7 +122,7 @@ pub struct EditGraph {
|
||||
/// photograph comes from.
|
||||
spots: SpotSet,
|
||||
/// The lens corrections that rewrite coordinates: distortion and lateral
|
||||
/// chromatic aberration (`docs/architecture.md` §5.2).
|
||||
/// chromatic aberration (`docs/dev/architecture.md` §5.2).
|
||||
///
|
||||
/// Apart from `ops` for the fourth time, and this one is not about shape
|
||||
/// but about direction. Every [`Operation`] is a function from colour to
|
||||
|
||||
@@ -27,7 +27,7 @@
|
||||
//! [`MaskSource::Regions`] stores integers naming regions in the segmentation
|
||||
//! hierarchy (`dr-segment`). That choice is what makes a mask diffable, cheap
|
||||
//! in a sidecar, and mergeable per-field under FR-NC-9 — three properties a
|
||||
//! stored raster has none of (docs/segmentation.md §1). Two devices that
|
||||
//! stored raster has none of (docs/dev/segmentation.md §1). Two devices that
|
||||
//! select the same subject produce the same small sorted list, and a sync
|
||||
//! conflict between them is resolvable rather than a binary blob fight.
|
||||
//!
|
||||
@@ -564,7 +564,7 @@ pub enum MaskSource {
|
||||
/// This is what the watershed and the semantic model exist to produce.
|
||||
/// Selecting a subject means "the regions the model's instance covers",
|
||||
/// and the resulting edge is the watershed's, which is to say the image's
|
||||
/// own (docs/segmentation.md §5).
|
||||
/// own (docs/dev/segmentation.md §5).
|
||||
Regions {
|
||||
/// Which segmentation these ids index into.
|
||||
///
|
||||
@@ -590,7 +590,7 @@ pub enum MaskSource {
|
||||
/// **The primary way a local adjustment is made.** The watershed hierarchy
|
||||
/// this crate was first built around does not survive a photograph: its
|
||||
/// saddles are near zero almost everywhere, so a global cut collapses the
|
||||
/// frame into one region plus noise (docs/segmentation.md §15). A model
|
||||
/// frame into one region plus noise (docs/dev/segmentation.md §15). A model
|
||||
/// instance is a whole object, found as one thing, and needs no ladder.
|
||||
///
|
||||
/// The trade is that the boundary is the model's — a quarter-resolution
|
||||
@@ -625,7 +625,7 @@ pub enum MaskSource {
|
||||
/// reason both exist. A subject is *one* instance — this dog, not that one
|
||||
/// — found by a COCO-trained instance model. A category is *all* the sky,
|
||||
/// or all the foliage, from an ADE20K-trained semantic model that has no
|
||||
/// notion of instances at all (docs/segmentation.md §16).
|
||||
/// notion of instances at all (docs/dev/segmentation.md §16).
|
||||
///
|
||||
/// So this is what a global grade attaches to: lift the sky, desaturate
|
||||
/// the vegetation, warm the architecture. Asking it for "that person
|
||||
@@ -2231,7 +2231,7 @@ fn reveal_block(slot: usize, layer: &MaskLayer, style: RevealStyle, colour: [f32
|
||||
/// **not** a hash of the label field: that would be a readback on a path that
|
||||
/// must not have one (ARCH §6.1), and would also make the signature depend on
|
||||
/// float arithmetic whose cross-vendor determinism is exactly the open
|
||||
/// question (docs/segmentation.md §6, M5).
|
||||
/// question (docs/dev/segmentation.md §6, M5).
|
||||
pub fn segmentation_signature(width: u32, height: u32, regions: u32, tuning: u64) -> u64 {
|
||||
// FNV-1a over the four fields. Small, dependency-free, and adequate: this
|
||||
// guards against accidental mismatch, not against a forged sidecar.
|
||||
|
||||
@@ -54,7 +54,7 @@ pub enum Affects {
|
||||
/// A pixel's *neighbourhood* — sharpening, noise reduction, clarity,
|
||||
/// texture, dehaze, spot removal.
|
||||
///
|
||||
/// The seam `docs/requirements.md` §3.3 designed and nothing cut until
|
||||
/// The seam `docs/dev/requirements.md` §3.3 designed and nothing cut until
|
||||
/// [`crate::detail`] existed. It is a separate variant rather than a flavour
|
||||
/// of `Colour` because it is a separate *dispatch*: a fragment in the fused
|
||||
/// pass is handed a colour and has no way back to a coordinate, so a
|
||||
|
||||
@@ -99,7 +99,7 @@
|
||||
//! A minimum over a patch is separable, as a Gaussian is: minimum along x,
|
||||
//! then along y. That alone is not enough. The patch is 1% of the shorter edge
|
||||
//! — 61 taps across at 4K — and two passes of 61 taps is the arithmetic that
|
||||
//! measured 34 ms for clarity and became `docs/technical-debt.md` TD-4.
|
||||
//! measured 34 ms for clarity and became `docs/dev/technical-debt.md` TD-4.
|
||||
//!
|
||||
//! A minimum has a property a Gaussian does not: **erosions compose by adding
|
||||
//! their structuring elements**. The minimum over a contiguous run of `d`
|
||||
|
||||
@@ -150,7 +150,7 @@
|
||||
//! the artefact this control must not have.
|
||||
//!
|
||||
//! Run at the render size, that measured **34 ms at 4K** — seven times the
|
||||
//! entire fused point chain, for one slider — which is `docs/technical-debt.md`
|
||||
//! entire fused point chain, for one slider — which is `docs/dev/technical-debt.md`
|
||||
//! TD-4 and is what [`Recipe::base_scale`] now answers. The base is computed on
|
||||
//! a grid a quarter the size on each axis: a sixteenth of the pixels at a
|
||||
//! quarter of the radius.
|
||||
@@ -271,7 +271,7 @@ impl Band for Coarse {
|
||||
threshold: 0.35,
|
||||
gain: 1.0,
|
||||
midtone_taper: true,
|
||||
// A quarter, which is what `docs/technical-debt.md` TD-4 bought back.
|
||||
// A quarter, which is what `docs/dev/technical-debt.md` TD-4 bought back.
|
||||
//
|
||||
// σ is 1.2% of the shorter edge — 26 px at 4K — so the base holds no
|
||||
// spatial frequency anywhere near the quarter-scale Nyquist of one
|
||||
|
||||
@@ -217,7 +217,7 @@ pub struct Version {
|
||||
/// graph's film cleared and the caller re-bakes — see `EditGraph::set_film`.
|
||||
pub film: Option<FilmRef>,
|
||||
/// TRACES: FR-DEV-8 | FR-NC-9
|
||||
/// The repairs (`docs/spot-removal.md`).
|
||||
/// The repairs (`docs/dev/spot-removal.md`).
|
||||
///
|
||||
/// A line per spot, keyed `spot.<id>`, rather than a block per spot as a
|
||||
/// mask gets: a spot is eight numbers, and sixty-four blocks would bury the
|
||||
|
||||
@@ -6,7 +6,7 @@
|
||||
//! are blended. No pixels are stored, here or anywhere: the shader draws the
|
||||
//! repair from these numbers every time the photograph is rendered, which is
|
||||
//! what makes it non-destructive, cheap to sync, and undoable
|
||||
//! (`docs/spot-removal.md`).
|
||||
//! (`docs/dev/spot-removal.md`).
|
||||
//!
|
||||
//! # Why this is not an operation
|
||||
//!
|
||||
@@ -61,7 +61,7 @@ pub const MAX_SPOTS: usize = 64;
|
||||
/// bounds something that is otherwise unbounded: a detail pass declares how far
|
||||
/// it reads from the pixel it writes, and for a spot that is the offset plus
|
||||
/// the radius. An unbounded offset is an unbounded halo, which is a pass the
|
||||
/// tile scheduler cannot plan (ARCH §5.3, `docs/spot-removal.md` §5.3).
|
||||
/// tile scheduler cannot plan (ARCH §5.3, `docs/dev/spot-removal.md` §5.3).
|
||||
pub const MAX_SOURCE_DISTANCE: f32 = 0.5;
|
||||
|
||||
/// The radius a new spot starts at, in frame units.
|
||||
@@ -210,7 +210,7 @@ impl Spot {
|
||||
/// **FR-DEV-8 asks for automatic source placement, and this is the cheap
|
||||
/// half of it.** The good half searches the photograph for a patch whose
|
||||
/// surroundings match — a compute dispatch scoring candidate offsets, and
|
||||
/// one small readback when the spot is created (`docs/spot-removal.md`
|
||||
/// one small readback when the spot is created (`docs/dev/spot-removal.md`
|
||||
/// §8). This is what stands in for it, and it is worth having on its own
|
||||
/// terms rather than as a placeholder: dust sits on skies, skies are
|
||||
/// smooth, and a patch two and a half radii away is nearly always the same
|
||||
|
||||
@@ -13,7 +13,7 @@ log.workspace = true
|
||||
|
||||
# Inference. `ort` is the API; **what runs it is `dr-inference-engine`'s
|
||||
# business** — tract, or an ONNX Runtime the app found on disk, on whichever
|
||||
# provider the device has (docs/inference.md). This crate never names either.
|
||||
# provider the device has (docs/dev/inference.md). This crate never names either.
|
||||
ort = { workspace = true, optional = true }
|
||||
dr-inference-engine = { workspace = true, optional = true }
|
||||
ndarray = { workspace = true, optional = true }
|
||||
|
||||
@@ -34,7 +34,7 @@
|
||||
//! And doing it here buys two things a shader could not. It is **exactly
|
||||
//! deterministic**, which matters because masks reach the sidecar as indices
|
||||
//! and a field that varied by vendor would mean a mask meaning one thing on
|
||||
//! the desktop and another on the phone (docs/segmentation.md §6, M5). And it
|
||||
//! the desktop and another on the phone (docs/dev/segmentation.md §6, M5). And it
|
||||
//! is testable against hand-computed distances with no adapter present.
|
||||
//!
|
||||
//! # The transform
|
||||
|
||||
@@ -40,7 +40,7 @@ pub struct Edge {
|
||||
|
||||
/// A partition of the image into labelled regions, plus how they adjoin.
|
||||
///
|
||||
/// The shared interface from docs/segmentation.md §2: arm A produces this
|
||||
/// The shared interface from docs/dev/segmentation.md §2: arm A produces this
|
||||
/// from a watershed, arm B would produce it from a class map, and the
|
||||
/// consumers above cannot tell which.
|
||||
#[derive(Debug, Clone, PartialEq)]
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
//! TRACES: FR-DEV-3i
|
||||
//! Region segmentation for local masking (S15, docs/segmentation.md).
|
||||
//! Region segmentation for local masking (S15, docs/dev/segmentation.md).
|
||||
//!
|
||||
//! Local adjustments need to know where the image's regions are before they
|
||||
//! can snap a mask to one. This crate is that map, and it is deliberately
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
//! Arm C — semantic instances as a prior over the watershed merge order.
|
||||
//!
|
||||
//! docs/segmentation.md §5. The spec calls this the expected winner and it is
|
||||
//! docs/dev/segmentation.md §5. The spec calls this the expected winner and it is
|
||||
//! what ships, for a reason that survives the model turning out to be narrower
|
||||
//! than §4 assumed: the two arms fail in *opposite* directions, so each one
|
||||
//! covers the other's failure.
|
||||
@@ -233,7 +233,7 @@ pub fn apply_semantic_prior(
|
||||
/// This is the interaction the whole spike exists to enable, and the reason it
|
||||
/// returns *region ids* rather than a raster: a mask that is a set of integers
|
||||
/// is diffable, mergeable at node level under FR-NC-9, and cheap in a sidecar
|
||||
/// (docs/segmentation.md §1). A raster is none of those.
|
||||
/// (docs/dev/segmentation.md §1). A raster is none of those.
|
||||
///
|
||||
/// The returned ids are sorted, so the same click always produces the same
|
||||
/// mask — which is what lets it be a cache key.
|
||||
|
||||
@@ -60,7 +60,7 @@
|
||||
//! So the last step is a **marker-based watershed**. The mask is eroded to
|
||||
//! give two markers — confidently inside, confidently outside — and the flood
|
||||
//! runs in the ribbon left between them, meeting along the most expensive line
|
||||
//! it can find. The cost is a sum of terms, as docs/segmentation.md §2 says it
|
||||
//! it can find. The cost is a sum of terms, as docs/dev/segmentation.md §2 says it
|
||||
//! should be: the photograph's own edges, and the colour model's disagreement.
|
||||
//!
|
||||
//! Markers are what make this the right shape rather than the watershed §15
|
||||
@@ -244,7 +244,7 @@ pub struct RefineOptions {
|
||||
/// How much the photograph's own edges count against the colour model in
|
||||
/// the flood's cost, `0.0..=1.0`.
|
||||
///
|
||||
/// docs/segmentation.md §2 specifies the cost as *a sum of terms* — image
|
||||
/// docs/dev/segmentation.md §2 specifies the cost as *a sum of terms* — image
|
||||
/// gradient always available, semantic evidence added when a model is
|
||||
/// present — and this is the mix. At one the boundary lands purely on the
|
||||
/// strongest edge in the band; at zero purely where the colour verdict
|
||||
@@ -378,7 +378,7 @@ const MAX_SAMPLES: usize = 20_000;
|
||||
/// floating-point comparison is a stopping rule that can differ between
|
||||
/// machines, and a mask that differs between machines reaches the sidecar as
|
||||
/// indices meaning one thing on the desktop and another on the phone
|
||||
/// (docs/segmentation.md §6).
|
||||
/// (docs/dev/segmentation.md §6).
|
||||
const ITERATIONS: usize = 12;
|
||||
|
||||
/// Half-width of the verdict scale, in nats.
|
||||
@@ -676,7 +676,7 @@ impl Refinement {
|
||||
/// fronts meet along the most expensive line in the ribbon — which is the
|
||||
/// watershed, and which is where the boundary belongs.
|
||||
///
|
||||
/// The cost is a sum of terms, as docs/segmentation.md §2 says it should
|
||||
/// The cost is a sum of terms, as docs/dev/segmentation.md §2 says it should
|
||||
/// be: the photograph's own edges, and the colour model's disagreement.
|
||||
/// Neither alone is right. An edge with no colour meaning is a texture,
|
||||
/// and a colour change with no edge is a gradient.
|
||||
@@ -889,7 +889,7 @@ fn neighbours(p: usize, w: usize, h: usize) -> impl Iterator<Item = usize> {
|
||||
/// Edge strength over the opponent features, as one byte per pixel.
|
||||
///
|
||||
/// Sobel over the same three numbers the colour model is fitted on, rather
|
||||
/// than over plain luma — docs/segmentation.md §3 is explicit that a
|
||||
/// than over plain luma — docs/dev/segmentation.md §3 is explicit that a
|
||||
/// channel-weighted RGB gradient reads a saturated red edge as weaker than it
|
||||
/// looks, and a flag against sky is exactly that edge.
|
||||
///
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
//! Semantic segmentation — arm B (S15, docs/segmentation.md §4).
|
||||
//! Semantic segmentation — arm B (S15, docs/dev/segmentation.md §4).
|
||||
//!
|
||||
//! Runs a YOLO instance-segmentation graph over a proxy-resolution image and
|
||||
//! returns the instances it found: a class, a score, a box, and a soft mask
|
||||
@@ -209,7 +209,7 @@ const EMBEDDED_MODEL: &[u8] = include_bytes!("../../../models/segment/yolo26n-se
|
||||
const EMBEDDED_CLASSES: &str = include_str!("../../../models/segment/yolo26n-seg.classes.json");
|
||||
|
||||
/// The bytes of the model that ships with this crate, for whoever compiles
|
||||
/// engines ahead of the first request (docs/inference.md §6).
|
||||
/// engines ahead of the first request (docs/dev/inference.md §6).
|
||||
#[cfg(feature = "embedded-model")]
|
||||
pub fn embedded_model_bytes() -> &'static [u8] {
|
||||
EMBEDDED_MODEL
|
||||
@@ -237,7 +237,7 @@ impl SemanticModel {
|
||||
|
||||
pub fn from_bytes(bytes: &[u8], classes: Vec<Arc<str>>) -> Result<Self, SegmentError> {
|
||||
// The f32 graph on whatever the device's backend is. An int8 form
|
||||
// for the Hexagon waits on docs/inference.md §10 M7 — the mask
|
||||
// for the Hexagon waits on docs/dev/inference.md §10 M7 — the mask
|
||||
// boundary has to be measured before it moves.
|
||||
let session = dr_inference_engine::open(
|
||||
dr_inference_engine::Role::Segmenter,
|
||||
|
||||
@@ -177,7 +177,7 @@ pub struct FaceSettings {
|
||||
/// Which SCRFD graph the indexing pass detects with.
|
||||
///
|
||||
/// Three exports of one architecture, differing only in how much computation
|
||||
/// they spend, and docs/faces.md §12.3 is the measurement that made this a
|
||||
/// they spend, and docs/dev/faces.md §12.3 is the measurement that made this a
|
||||
/// choice rather than a constant: over the same photographs the cheapest one
|
||||
/// misses the small faces in a group and reports a dog a dozen times, the
|
||||
/// middle one finds 14% more faces for 12% more time, and the largest a
|
||||
@@ -261,7 +261,7 @@ impl FaceDetector {
|
||||
}
|
||||
}
|
||||
|
||||
/// The id when the detector runs in its int8 form (docs/inference.md §7).
|
||||
/// The id when the detector runs in its int8 form (docs/dev/inference.md §7).
|
||||
///
|
||||
/// A different detector: it finds a different set of faces, so it is a
|
||||
/// different population of detections. The embedder half is unchanged,
|
||||
|
||||
Reference in New Issue
Block a user