Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
9d04ff2154 | ||
|
|
c6cfb2a02a | ||
|
|
b83f192847 | ||
|
|
ecb648818b | ||
|
|
5fbf8944d7 | ||
|
|
b502a8ef90 | ||
|
|
6fbdb1d06f | ||
|
|
d04b8f6044 | ||
|
|
8baa46ff49 | ||
|
|
e43ae10439 | ||
|
|
104e3a106f | ||
|
|
031ba7b77d | ||
|
|
54a80e688c | ||
|
|
2fd7690b6f | ||
|
|
c5f07f9ced | ||
|
|
5c00942b84 | ||
|
|
2a4ac0ed3d | ||
|
|
46af2a0a46 | ||
|
|
95c9cffc0d | ||
|
|
cbbe67fbd7 | ||
|
|
691af96e3e | ||
|
|
7a436e2549 | ||
|
|
76bc5652d7 | ||
|
|
4ed29b9d81 | ||
|
|
05508741af | ||
|
|
d15c41e699 | ||
|
|
caf21bea64 | ||
|
|
6739fdf908 | ||
|
|
42d11d919b | ||
|
|
67f225beba | ||
|
|
39adfd4b75 | ||
|
|
57ed51c1c5 | ||
|
|
30bd276d0b | ||
|
|
75d2ceb23c | ||
|
|
2e9a1eb0f0 | ||
|
|
44ea763c61 | ||
|
|
acab0d7abb | ||
|
|
9b6b4942cf | ||
|
|
54290b9540 | ||
|
|
231b4a54ab | ||
|
|
2bf0ec8dba | ||
|
|
5bf06c5030 | ||
|
|
44fdcbc6f7 | ||
|
|
f9510405c3 | ||
|
|
7e6b25b21b | ||
|
|
e4b6b6c935 | ||
|
|
1ded5afbaa | ||
|
|
c901fc1a0a | ||
|
|
f79a76f2d5 | ||
|
|
facb44cb55 | ||
|
|
85cc2b1dcc | ||
|
|
d706c12d77 | ||
|
|
cd0ca6785f | ||
|
|
83f4253b6a | ||
|
|
6aae4c3eb0 | ||
|
|
54b543fb77 | ||
|
|
f5956707e7 | ||
|
|
b908d861e0 | ||
|
|
6b51726322 | ||
|
|
2481904016 | ||
|
|
a92ae4576f | ||
|
|
ed4460cb9c | ||
|
|
7596cf9bcc | ||
|
|
696bafa9d5 | ||
|
|
d259c0d4bb | ||
|
|
95458356da | ||
|
|
dc9db11033 | ||
|
|
3692306fd3 | ||
|
|
c826fed605 | ||
|
|
c921852d89 | ||
|
|
e6ac31d39d | ||
|
|
7db999c1f6 | ||
|
|
30b89ad70a | ||
|
|
327decfab1 | ||
|
|
f8addbee53 | ||
|
|
c0b1e78f7c | ||
|
|
78cb00634e | ||
|
|
2917b7427d | ||
|
|
33e2e277a2 | ||
|
|
9b627e7713 |
@@ -10,3 +10,13 @@
|
||||
# detects exactly that and fails with an instruction rather than embedding the
|
||||
# pointer and failing at inference time.
|
||||
*.onnx filter=lfs diff=lfs merge=lfs -text
|
||||
|
||||
# Test photographs live in LFS too, and are fetched only by the tests that
|
||||
# need them.
|
||||
#
|
||||
# `fixtures/**` holds real camera files — a twelve-frame panorama set is
|
||||
# 325 MB — and CI's `git lfs pull` excludes the directory, so a checkout
|
||||
# carries pointers there until a merge test asks for the frames. Same
|
||||
# reasoning as the models, with the opposite default: the model is not
|
||||
# optional and the fixtures are.
|
||||
fixtures/** filter=lfs diff=lfs merge=lfs -text
|
||||
|
||||
@@ -148,7 +148,7 @@ jobs:
|
||||
| while read -r key; do git config --local --unset-all "$key"; done || true
|
||||
git config --local lfs.url \
|
||||
"https://x-access-token:${LFS_TOKEN}@gitea.tourolle.paris/dtourolle/DarkRoom.git/info/lfs"
|
||||
git lfs pull
|
||||
git lfs pull --exclude="fixtures/**"
|
||||
|
||||
- name: Cache cargo
|
||||
uses: actions/cache@v4
|
||||
|
||||
@@ -96,7 +96,7 @@ jobs:
|
||||
| while read -r key; do git config --local --unset-all "$key"; done || true
|
||||
git config --local lfs.url \
|
||||
"https://x-access-token:${LFS_TOKEN}@gitea.tourolle.paris/dtourolle/DarkRoom.git/info/lfs"
|
||||
git lfs pull
|
||||
git lfs pull --exclude="fixtures/**"
|
||||
ls -lR models/
|
||||
|
||||
- name: Cache cargo
|
||||
@@ -213,7 +213,7 @@ jobs:
|
||||
| while read -r key; do git config --local --unset-all "$key"; done || true
|
||||
git config --local lfs.url \
|
||||
"https://x-access-token:${LFS_TOKEN}@gitea.tourolle.paris/dtourolle/DarkRoom.git/info/lfs"
|
||||
git lfs pull
|
||||
git lfs pull --exclude="fixtures/**"
|
||||
ls -lR models/
|
||||
|
||||
- name: Cache cargo
|
||||
@@ -406,7 +406,7 @@ jobs:
|
||||
| while read -r key; do git config --local --unset-all "$key"; done || true
|
||||
git config --local lfs.url \
|
||||
"https://x-access-token:${LFS_TOKEN}@gitea.tourolle.paris/dtourolle/DarkRoom.git/info/lfs"
|
||||
git lfs pull
|
||||
git lfs pull --exclude="fixtures/**"
|
||||
ls -l models/face models/scene
|
||||
|
||||
- name: Cache cargo
|
||||
|
||||
@@ -23,3 +23,4 @@ tools/film-profiles/upstream/
|
||||
# checkout, so it is larger than the repository it sits in.
|
||||
/.flatpak-builder/
|
||||
/build/
|
||||
__pycache__/
|
||||
|
||||
Generated
+57
-25
@@ -1221,7 +1221,7 @@ checksum = "f27ae1dd37df86211c42e150270f82743308803d90a6f6e6651cd730d5e1732f"
|
||||
|
||||
[[package]]
|
||||
name = "darkroom-android"
|
||||
version = "0.12.1"
|
||||
version = "0.13.1"
|
||||
dependencies = [
|
||||
"android_logger",
|
||||
"dr-plat",
|
||||
@@ -1234,7 +1234,7 @@ dependencies = [
|
||||
|
||||
[[package]]
|
||||
name = "darkroom-desktop"
|
||||
version = "0.12.1"
|
||||
version = "0.13.1"
|
||||
dependencies = [
|
||||
"anyhow",
|
||||
"dr-plat",
|
||||
@@ -1408,7 +1408,7 @@ checksum = "d8b14ccef22fc6f5a8f4d7d768562a182c04ce9a3b3157b91390b52ddfdf1a76"
|
||||
|
||||
[[package]]
|
||||
name = "dr-bench"
|
||||
version = "0.12.1"
|
||||
version = "0.13.1"
|
||||
dependencies = [
|
||||
"anyhow",
|
||||
"dr-catalog",
|
||||
@@ -1425,7 +1425,7 @@ dependencies = [
|
||||
|
||||
[[package]]
|
||||
name = "dr-catalog"
|
||||
version = "0.12.1"
|
||||
version = "0.13.1"
|
||||
dependencies = [
|
||||
"dr-face",
|
||||
"dr-plat",
|
||||
@@ -1440,7 +1440,7 @@ dependencies = [
|
||||
|
||||
[[package]]
|
||||
name = "dr-decode"
|
||||
version = "0.12.1"
|
||||
version = "0.13.1"
|
||||
dependencies = [
|
||||
"dr-types",
|
||||
"env_logger",
|
||||
@@ -1454,7 +1454,7 @@ dependencies = [
|
||||
|
||||
[[package]]
|
||||
name = "dr-export"
|
||||
version = "0.12.1"
|
||||
version = "0.13.1"
|
||||
dependencies = [
|
||||
"dr-decode",
|
||||
"dr-gpu",
|
||||
@@ -1465,6 +1465,7 @@ dependencies = [
|
||||
"log",
|
||||
"png",
|
||||
"pollster",
|
||||
"rawler",
|
||||
"thiserror 2.0.20",
|
||||
"tiff",
|
||||
"zune-jpeg 0.4.21",
|
||||
@@ -1472,20 +1473,20 @@ dependencies = [
|
||||
|
||||
[[package]]
|
||||
name = "dr-face"
|
||||
version = "0.12.1"
|
||||
version = "0.13.1"
|
||||
dependencies = [
|
||||
"dr-inference-engine",
|
||||
"env_logger",
|
||||
"log",
|
||||
"ndarray",
|
||||
"ort",
|
||||
"ort-tract",
|
||||
"thiserror 2.0.20",
|
||||
"zune-jpeg 0.4.21",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "dr-film"
|
||||
version = "0.12.1"
|
||||
version = "0.13.1"
|
||||
dependencies = [
|
||||
"log",
|
||||
"serde",
|
||||
@@ -1494,11 +1495,12 @@ dependencies = [
|
||||
|
||||
[[package]]
|
||||
name = "dr-gpu"
|
||||
version = "0.12.1"
|
||||
version = "0.13.1"
|
||||
dependencies = [
|
||||
"bytemuck",
|
||||
"dr-decode",
|
||||
"dr-film",
|
||||
"dr-pano",
|
||||
"dr-pipeline",
|
||||
"dr-segment",
|
||||
"dr-types",
|
||||
@@ -1509,9 +1511,23 @@ dependencies = [
|
||||
"wgpu",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "dr-inference-engine"
|
||||
version = "0.13.1"
|
||||
dependencies = [
|
||||
"libloading",
|
||||
"log",
|
||||
"ort",
|
||||
"ort-sys",
|
||||
"ort-tract",
|
||||
"serde",
|
||||
"serde_json",
|
||||
"thiserror 2.0.20",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "dr-ingest"
|
||||
version = "0.12.1"
|
||||
version = "0.13.1"
|
||||
dependencies = [
|
||||
"dr-plat",
|
||||
"dr-types",
|
||||
@@ -1523,15 +1539,29 @@ dependencies = [
|
||||
|
||||
[[package]]
|
||||
name = "dr-lens"
|
||||
version = "0.12.1"
|
||||
version = "0.13.1"
|
||||
dependencies = [
|
||||
"lensfun",
|
||||
"log",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "dr-pano"
|
||||
version = "0.13.1"
|
||||
dependencies = [
|
||||
"dr-decode",
|
||||
"dr-inference-engine",
|
||||
"dr-types",
|
||||
"env_logger",
|
||||
"log",
|
||||
"ndarray",
|
||||
"ort",
|
||||
"thiserror 2.0.20",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "dr-pipeline"
|
||||
version = "0.12.1"
|
||||
version = "0.13.1"
|
||||
dependencies = [
|
||||
"dr-types",
|
||||
"log",
|
||||
@@ -1540,7 +1570,7 @@ dependencies = [
|
||||
|
||||
[[package]]
|
||||
name = "dr-plat"
|
||||
version = "0.12.1"
|
||||
version = "0.13.1"
|
||||
dependencies = [
|
||||
"android-native-keyring-store",
|
||||
"dr-types",
|
||||
@@ -1556,7 +1586,7 @@ dependencies = [
|
||||
|
||||
[[package]]
|
||||
name = "dr-preset-xmp"
|
||||
version = "0.12.1"
|
||||
version = "0.13.1"
|
||||
dependencies = [
|
||||
"dr-pipeline",
|
||||
"log",
|
||||
@@ -1566,20 +1596,20 @@ dependencies = [
|
||||
|
||||
[[package]]
|
||||
name = "dr-segment"
|
||||
version = "0.12.1"
|
||||
version = "0.13.1"
|
||||
dependencies = [
|
||||
"dr-inference-engine",
|
||||
"env_logger",
|
||||
"log",
|
||||
"ndarray",
|
||||
"ort",
|
||||
"ort-tract",
|
||||
"thiserror 2.0.20",
|
||||
"zune-jpeg 0.4.21",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "dr-sync"
|
||||
version = "0.12.1"
|
||||
version = "0.13.1"
|
||||
dependencies = [
|
||||
"async-trait",
|
||||
"dr-plat",
|
||||
@@ -1593,7 +1623,7 @@ dependencies = [
|
||||
|
||||
[[package]]
|
||||
name = "dr-sync-folder"
|
||||
version = "0.12.1"
|
||||
version = "0.13.1"
|
||||
dependencies = [
|
||||
"async-trait",
|
||||
"dr-sync",
|
||||
@@ -1605,7 +1635,7 @@ dependencies = [
|
||||
|
||||
[[package]]
|
||||
name = "dr-sync-nextcloud"
|
||||
version = "0.12.1"
|
||||
version = "0.13.1"
|
||||
dependencies = [
|
||||
"async-trait",
|
||||
"dr-decode",
|
||||
@@ -1627,7 +1657,7 @@ dependencies = [
|
||||
|
||||
[[package]]
|
||||
name = "dr-thumbs"
|
||||
version = "0.12.1"
|
||||
version = "0.13.1"
|
||||
dependencies = [
|
||||
"dr-types",
|
||||
"jpeg-encoder",
|
||||
@@ -1639,7 +1669,7 @@ dependencies = [
|
||||
|
||||
[[package]]
|
||||
name = "dr-types"
|
||||
version = "0.12.1"
|
||||
version = "0.13.1"
|
||||
dependencies = [
|
||||
"serde",
|
||||
"serde_json",
|
||||
@@ -1648,7 +1678,7 @@ dependencies = [
|
||||
|
||||
[[package]]
|
||||
name = "dr-ui"
|
||||
version = "0.12.1"
|
||||
version = "0.13.1"
|
||||
dependencies = [
|
||||
"anyhow",
|
||||
"async-trait",
|
||||
@@ -1658,8 +1688,10 @@ dependencies = [
|
||||
"dr-face",
|
||||
"dr-film",
|
||||
"dr-gpu",
|
||||
"dr-inference-engine",
|
||||
"dr-ingest",
|
||||
"dr-lens",
|
||||
"dr-pano",
|
||||
"dr-pipeline",
|
||||
"dr-plat",
|
||||
"dr-preset-xmp",
|
||||
@@ -1688,7 +1720,7 @@ dependencies = [
|
||||
|
||||
[[package]]
|
||||
name = "dr-xmp"
|
||||
version = "0.12.1"
|
||||
version = "0.13.1"
|
||||
dependencies = [
|
||||
"dr-types",
|
||||
"log",
|
||||
@@ -6989,7 +7021,7 @@ checksum = "8df9b6e13f2d32c91b9bd719c00d1958837bc7dec474d94952798cc8e69eeec3"
|
||||
|
||||
[[package]]
|
||||
name = "traceability"
|
||||
version = "0.12.1"
|
||||
version = "0.13.1"
|
||||
dependencies = [
|
||||
"anyhow",
|
||||
"serde",
|
||||
|
||||
+8
-1
@@ -8,9 +8,11 @@ members = [
|
||||
"core/dr-export",
|
||||
"core/dr-face",
|
||||
"core/dr-film",
|
||||
"core/dr-inference-engine",
|
||||
"core/dr-ingest",
|
||||
"core/dr-gpu",
|
||||
"core/dr-lens",
|
||||
"core/dr-pano",
|
||||
"core/dr-pipeline",
|
||||
"core/dr-preset-xmp",
|
||||
"core/dr-segment",
|
||||
@@ -27,7 +29,7 @@ members = [
|
||||
]
|
||||
|
||||
[workspace.package]
|
||||
version = "0.12.1"
|
||||
version = "0.13.1"
|
||||
edition = "2021"
|
||||
rust-version = "1.92"
|
||||
license = "GPL-3.0-or-later"
|
||||
@@ -45,9 +47,14 @@ dr-export = { path = "core/dr-export" }
|
||||
# `features = ["inference"]`.
|
||||
dr-face = { path = "core/dr-face", default-features = false }
|
||||
dr-film = { path = "core/dr-film" }
|
||||
# `tract` on by default so a test binary can open a session with nothing
|
||||
# installed; the apps add `native` to look for a runtime file (docs/inference.md §3).
|
||||
dr-inference-engine = { path = "core/dr-inference-engine" }
|
||||
dr-ingest = { path = "core/dr-ingest" }
|
||||
dr-gpu = { path = "core/dr-gpu" }
|
||||
dr-lens = { path = "core/dr-lens" }
|
||||
# Optional runtime, like `dr-segment`: the geometry never needs a model.
|
||||
dr-pano = { path = "core/dr-pano", default-features = false }
|
||||
dr-pipeline = { path = "core/dr-pipeline" }
|
||||
dr-preset-xmp = { path = "core/dr-preset-xmp" }
|
||||
# `default-features = false` belongs *here*, not on each dependant: a member
|
||||
|
||||
@@ -324,17 +324,40 @@ fn unpack_bundled_models(app: &slint::android::AndroidApp) {
|
||||
// (`FaceDetector`, docs/faces.md §12.3) and a tablet has no other way to
|
||||
// obtain the one it was not shipped with. Twenty megabytes of APK for
|
||||
// the choice; the embedder is the same for all three.
|
||||
const BUNDLED: [(&std::ffi::CStr, &str); 7] = [
|
||||
//
|
||||
// Then the three eye-state models (docs/faces.md §17): landmarks, open
|
||||
// or closed, sunglasses. The app indexes without them; with them the
|
||||
// eyes-open filter has something to read, and a tablet has no other way
|
||||
// to get them either.
|
||||
//
|
||||
// The int8 forms beside the three detectors are what the Hexagon runs
|
||||
// (docs/inference.md §5); the engine loads the sibling when the probe
|
||||
// chose that rung and ignores it otherwise.
|
||||
const BUNDLED: [(&std::ffi::CStr, &str); 14] = [
|
||||
(c"models/scrfd_500m_640.onnx", "scrfd_500m_640.onnx"),
|
||||
(
|
||||
c"models/scrfd_500m_640.int8.onnx",
|
||||
"scrfd_500m_640.int8.onnx",
|
||||
),
|
||||
(c"models/scrfd_2.5g_640.onnx", "scrfd_2.5g_640.onnx"),
|
||||
(
|
||||
c"models/scrfd_2.5g_640.int8.onnx",
|
||||
"scrfd_2.5g_640.int8.onnx",
|
||||
),
|
||||
(c"models/scrfd_10g_640.onnx", "scrfd_10g_640.onnx"),
|
||||
(c"models/scrfd_10g_640.int8.onnx", "scrfd_10g_640.int8.onnx"),
|
||||
(c"models/arcface_mbf_b1.onnx", "arcface_mbf_b1.onnx"),
|
||||
(c"models/2d106det_b1.onnx", "2d106det_b1.onnx"),
|
||||
(c"models/ocec_s_b1.onnx", "ocec_s_b1.onnx"),
|
||||
(c"models/sgc_l_48_b1.onnx", "sgc_l_48_b1.onnx"),
|
||||
(c"models/yolo26s-sem-ade20k.onnx", "yolo26s-sem-ade20k.onnx"),
|
||||
(
|
||||
c"models/yolo26s-sem-ade20k.classes.json",
|
||||
"yolo26s-sem-ade20k.classes.json",
|
||||
),
|
||||
(c"models/categories.txt", "categories.txt"),
|
||||
// The panorama border filler (FR-MRG-4); MIT, 28 MB.
|
||||
(c"models/migan-512.onnx", "migan-512.onnx"),
|
||||
];
|
||||
|
||||
let dir = dr_ui::shared_face_models_dir();
|
||||
@@ -343,8 +366,8 @@ fn unpack_bundled_models(app: &slint::android::AndroidApp) {
|
||||
|
||||
for (asset_path, name) in BUNDLED {
|
||||
let dest = dir.join(name);
|
||||
// Already unpacked. Not re-read on every launch: this is 61 MB of
|
||||
// copying across the seven entries, and the file does not change without
|
||||
// Already unpacked. Not re-read on every launch: this is 73 MB of
|
||||
// copying across the ten entries, and the file does not change without
|
||||
// the APK changing, at which point the install wiped it anyway. It
|
||||
// matters more now than it did — a launch that skips every entry here
|
||||
// costs nothing at all, which is what makes the second launch after an
|
||||
@@ -392,6 +415,27 @@ fn unpack_bundled_models(app: &slint::android::AndroidApp) {
|
||||
"bundled models ready: {copied} bytes copied in {} ms",
|
||||
started.elapsed().as_millis()
|
||||
);
|
||||
|
||||
// Now, and not at launch: the probe fingerprints the model files, and
|
||||
// on a first launch they were not on disk until this line. The runtime
|
||||
// is in the APK's native library directory beside `libdarkroom.so`,
|
||||
// which is also where Qualcomm's DSP loader has to be pointed for the
|
||||
// Hexagon skel (docs/inference.md §3, §8).
|
||||
dr_ui::inference::init(native_library_dir().into_iter().collect());
|
||||
}
|
||||
|
||||
/// The directory the system unpacked this APK's native libraries into.
|
||||
///
|
||||
/// Read from where the loader put *this* library rather than asked of the
|
||||
/// activity: `android-activity` does not expose `nativeLibraryDir`, and the
|
||||
/// answer is in `/proc/self/maps` for free.
|
||||
#[cfg(target_os = "android")]
|
||||
fn native_library_dir() -> Option<std::path::PathBuf> {
|
||||
let maps = std::fs::read_to_string("/proc/self/maps").ok()?;
|
||||
maps.lines()
|
||||
.filter_map(|l| l.split_whitespace().nth(5))
|
||||
.find(|p| p.ends_with("/libdarkroom.so"))
|
||||
.and_then(|p| std::path::Path::new(p).parent().map(Into::into))
|
||||
}
|
||||
|
||||
/// TRACES: FR-PLAT-AND-6
|
||||
|
||||
@@ -61,6 +61,11 @@ fn main() -> anyhow::Result<()> {
|
||||
eprintln!("usage: darkroom-desktop <file-or-directory>...");
|
||||
}
|
||||
|
||||
// Before the window: the probe runs on its own thread and the first
|
||||
// frame does not wait for it, but the models a background job asks for
|
||||
// should already know where the runtime is (docs/inference.md §4).
|
||||
dr_ui::inference::init(runtime_dirs());
|
||||
|
||||
dr_ui::run(paths)?;
|
||||
|
||||
// Skip Rust's normal static/thread-local teardown on the way out: a
|
||||
@@ -70,3 +75,39 @@ fn main() -> anyhow::Result<()> {
|
||||
// destruction" when the window is closed.
|
||||
std::process::exit(0);
|
||||
}
|
||||
|
||||
/// Where a desktop package may have put `libonnxruntime`, most specific
|
||||
/// first. None of these existing is the tract build, which is a complete
|
||||
/// application and not an error (docs/inference.md §3).
|
||||
///
|
||||
/// `DARKROOM_ORT_DIR` is for a developer pointing at a runtime that is not
|
||||
/// installed — the wheel's `capi` directory, say. Then beside the executable
|
||||
/// and in the package's private library directory, for a package that
|
||||
/// bundles its own; then the user's own `runtime/` beside the models, where
|
||||
/// `tools/fetch-desktop-runtime.sh` puts one; then the Flatpak prefix; then
|
||||
/// the system library directory, for a distribution that ships ONNX Runtime
|
||||
/// as a package of its own. The user's copy outranks the system's because
|
||||
/// the system's is the one most likely to be built without the GPU
|
||||
/// providers, or against the wrong cuDNN — and a system copy whose providers
|
||||
/// do not load is not a problem, only a slower app: the probe builds a real
|
||||
/// session before believing a provider.
|
||||
fn runtime_dirs() -> Vec<PathBuf> {
|
||||
let mut dirs = Vec::new();
|
||||
if let Some(dir) = std::env::var_os("DARKROOM_ORT_DIR") {
|
||||
dirs.push(PathBuf::from(dir));
|
||||
}
|
||||
if let Ok(exe) = std::env::current_exe() {
|
||||
if let Some(bin) = exe.parent() {
|
||||
dirs.push(bin.to_path_buf());
|
||||
dirs.push(bin.join("../lib/darkroom"));
|
||||
}
|
||||
}
|
||||
dirs.push(dr_ui::inference::user_runtime_dir());
|
||||
#[cfg(target_os = "linux")]
|
||||
dirs.extend([
|
||||
PathBuf::from("/app/lib/darkroom"),
|
||||
PathBuf::from("/usr/lib/darkroom"),
|
||||
PathBuf::from("/usr/lib"),
|
||||
]);
|
||||
dirs
|
||||
}
|
||||
|
||||
@@ -48,9 +48,10 @@ pub const SHARD_MAX_BYTES: u64 = dr_thumbs::SHARD_MAX_BYTES;
|
||||
/// Bytes one stored face occupies, near enough to bound a shard by.
|
||||
///
|
||||
/// Counted rather than measured: the embedding is fixed at 512 × f16, the
|
||||
/// landmarks at 5 × 2 × f32, and the rest is a handful of numbers. Measuring
|
||||
/// landmarks at 5 × 2 × f32 and the dense ones at 106 × 2 × u16, and the
|
||||
/// rest is a handful of numbers. Measuring
|
||||
/// the file after each insert would mean a `VACUUM` to get an honest answer.
|
||||
const BYTES_PER_FACE: u64 = 1024 + 40 + 64;
|
||||
const BYTES_PER_FACE: u64 = 1024 + 40 + 424 + 64;
|
||||
|
||||
/// Bytes a stored crop occupies, near enough to bound a shard by.
|
||||
///
|
||||
@@ -79,6 +80,11 @@ pub struct SharedFace {
|
||||
/// See `faces::DetectedFace::quality`. `None` from a shard written before
|
||||
/// the number was kept.
|
||||
pub quality: Option<f32>,
|
||||
/// See `faces::DetectedFace::eyes`. `None` from a peer without the eye
|
||||
/// models, or a shard written before they existed.
|
||||
pub eyes: Option<dr_face::EyeReading>,
|
||||
/// See `faces::DetectedFace::landmarks_dense`; empty where none.
|
||||
pub landmarks_dense: Vec<u8>,
|
||||
/// The face cut out and encoded, or empty where none was kept.
|
||||
///
|
||||
/// Travels with the face rather than in the catalog snapshot, which is the
|
||||
@@ -149,6 +155,29 @@ impl FaceShardStore {
|
||||
.flatten()
|
||||
}
|
||||
|
||||
/// The pipeline this store holds an image under, among those sharing
|
||||
/// `model_id`'s embedder — the most recently indexed where a peer has
|
||||
/// sent more than one.
|
||||
///
|
||||
/// What the import asks: not "has anyone run *this* detector over it" but
|
||||
/// "does anyone hold comparable faces for it". See `faces::embedder_of`.
|
||||
pub fn held_model(&self, file_id: u64, model_id: &str) -> Option<String> {
|
||||
self.index
|
||||
.query_row(
|
||||
&format!(
|
||||
"SELECT model_id FROM entries
|
||||
WHERE file_id = ?1 AND {} = ?2
|
||||
ORDER BY indexed_at DESC NULLS LAST, model_id",
|
||||
crate::faces::embedder_sql("model_id")
|
||||
),
|
||||
rusqlite::params![file_id as i64, crate::faces::embedder_of(model_id)],
|
||||
|r| r.get::<_, String>(0),
|
||||
)
|
||||
.optional()
|
||||
.ok()
|
||||
.flatten()
|
||||
}
|
||||
|
||||
pub fn contains(&self, file_id: u64, model_id: &str) -> bool {
|
||||
self.index
|
||||
.query_row(
|
||||
@@ -227,8 +256,12 @@ impl FaceShardStore {
|
||||
tx.execute(
|
||||
"INSERT INTO faces
|
||||
(file_id, model_id, x, y, w, h, landmarks, confidence,
|
||||
embedding, crop_px, crop, quality)
|
||||
VALUES (?1, ?2, ?3, ?4, ?5, ?6, ?7, ?8, ?9, ?10, ?11, ?12)",
|
||||
embedding, crop_px, crop, quality,
|
||||
eye_right, eye_right_px, eye_right_sharp,
|
||||
eye_left, eye_left_px, eye_left_sharp, sunglasses,
|
||||
landmarks_dense)
|
||||
VALUES (?1, ?2, ?3, ?4, ?5, ?6, ?7, ?8, ?9, ?10, ?11, ?12,
|
||||
?13, ?14, ?15, ?16, ?17, ?18, ?19, ?20)",
|
||||
rusqlite::params![
|
||||
f.file_id as i64,
|
||||
f.model_id,
|
||||
@@ -242,6 +275,14 @@ impl FaceShardStore {
|
||||
f.crop_px as f64,
|
||||
(!f.crop.is_empty()).then_some(f.crop.as_slice()),
|
||||
f.quality.map(f64::from),
|
||||
f.eyes.map(|e| f64::from(e.right.open)),
|
||||
f.eyes.map(|e| f64::from(e.right.px)),
|
||||
f.eyes.map(|e| f64::from(e.right.sharpness)),
|
||||
f.eyes.map(|e| f64::from(e.left.open)),
|
||||
f.eyes.map(|e| f64::from(e.left.px)),
|
||||
f.eyes.map(|e| f64::from(e.left.sharpness)),
|
||||
f.eyes.map(|e| f64::from(e.sunglasses)),
|
||||
(!f.landmarks_dense.is_empty()).then_some(f.landmarks_dense.as_slice()),
|
||||
],
|
||||
)?;
|
||||
}
|
||||
@@ -446,10 +487,16 @@ impl FaceShardStore {
|
||||
}
|
||||
let mut fq = src.prepare(&format!(
|
||||
"SELECT f.file_id, f.model_id, f.x, f.y, f.w, f.h, f.landmarks,
|
||||
f.confidence, f.embedding, f.crop_px, {}, {}
|
||||
f.confidence, f.embedding, f.crop_px, {}, {}, {}, {}
|
||||
FROM faces f WHERE f.file_id = ?1 AND f.model_id = ?2",
|
||||
column_or_null(&src, "crop"),
|
||||
column_or_null(&src, "quality"),
|
||||
crate::schema::EYE_COLUMNS
|
||||
.iter()
|
||||
.map(|c| column_or_null(&src, c))
|
||||
.collect::<Vec<_>>()
|
||||
.join(", "),
|
||||
column_or_null(&src, "landmarks_dense"),
|
||||
))?;
|
||||
let faces: Vec<SharedFace> = fq
|
||||
.query_map(rusqlite::params![file_id, &model_id], read_shared_face)?
|
||||
@@ -490,7 +537,8 @@ impl FaceShardStore {
|
||||
|
||||
let mut q = conn.prepare(
|
||||
"SELECT file_id, model_id, x, y, w, h, landmarks, confidence, embedding, crop_px,
|
||||
crop, quality
|
||||
crop, quality, eye_right, eye_right_px, eye_right_sharp,
|
||||
eye_left, eye_left_px, eye_left_sharp, sunglasses, landmarks_dense
|
||||
FROM faces WHERE file_id = ?1 AND model_id = ?2",
|
||||
)?;
|
||||
let faces: Vec<SharedFace> = q
|
||||
@@ -548,6 +596,14 @@ fn upgrade_shard(conn: &Connection) -> Result<(), CatalogError> {
|
||||
("faces", "crop", "BLOB"),
|
||||
("indexed", "indexed_at", "INTEGER"),
|
||||
("faces", "quality", "REAL"),
|
||||
("faces", "eye_right", "REAL"),
|
||||
("faces", "eye_right_px", "REAL"),
|
||||
("faces", "eye_right_sharp", "REAL"),
|
||||
("faces", "eye_left", "REAL"),
|
||||
("faces", "eye_left_px", "REAL"),
|
||||
("faces", "eye_left_sharp", "REAL"),
|
||||
("faces", "sunglasses", "REAL"),
|
||||
("faces", "landmarks_dense", "BLOB"),
|
||||
] {
|
||||
if !has_column(conn, table, column)? {
|
||||
conn.execute_batch(&format!("ALTER TABLE {table} ADD COLUMN {column} {decl}"))?;
|
||||
@@ -571,7 +627,7 @@ fn has_column(conn: &Connection, table: &str, column: &str) -> Result<bool, Cata
|
||||
///
|
||||
/// `column` is one of this module's own names, never anything read from
|
||||
/// outside, which is what makes formatting it into SQL acceptable.
|
||||
fn column_or_null(conn: &Connection, column: &'static str) -> String {
|
||||
fn column_or_null(conn: &Connection, column: &str) -> String {
|
||||
match has_column(conn, "faces", column) {
|
||||
Ok(true) => format!("f.{column}"),
|
||||
_ => "NULL".to_string(),
|
||||
@@ -608,16 +664,21 @@ pub fn export_to_shards_reporting(
|
||||
model_id: &str,
|
||||
progress: &mut dyn FnMut(usize, usize),
|
||||
) -> Result<usize, CatalogError> {
|
||||
let mut q = conn.prepare(
|
||||
"SELECT r.file_id, fi.image_id, fi.source_edge, fi.indexed_at
|
||||
// Every pipeline sharing this one's embedder, each image under the id
|
||||
// that actually indexed it. A device that switched detectors still holds
|
||||
// most of its library under the previous id, and those faces are exactly
|
||||
// as comparable — and as wanted by a peer — as the new ones.
|
||||
let mut q = conn.prepare(&format!(
|
||||
"SELECT r.file_id, fi.image_id, fi.source_edge, fi.indexed_at, fi.model_id
|
||||
FROM face_index fi
|
||||
JOIN remote r ON r.image_id = fi.image_id
|
||||
WHERE fi.model_id = ?1
|
||||
WHERE {} = ?1
|
||||
ORDER BY fi.image_id",
|
||||
)?;
|
||||
let rows: Vec<(i64, i64, i64, i64)> = q
|
||||
.query_map([model_id], |r| {
|
||||
Ok((r.get(0)?, r.get(1)?, r.get(2)?, r.get(3)?))
|
||||
crate::faces::embedder_sql("fi.model_id")
|
||||
))?;
|
||||
let rows: Vec<(i64, i64, i64, i64, String)> = q
|
||||
.query_map([crate::faces::embedder_of(model_id)], |r| {
|
||||
Ok((r.get(0)?, r.get(1)?, r.get(2)?, r.get(3)?, r.get(4)?))
|
||||
})?
|
||||
.collect::<Result<_, _>>()?;
|
||||
|
||||
@@ -627,7 +688,8 @@ pub fn export_to_shards_reporting(
|
||||
|
||||
let total = rows.len();
|
||||
let mut exported = 0;
|
||||
for (seen, (file_id, image_id, edge, indexed_at)) in rows.into_iter().enumerate() {
|
||||
for (seen, (file_id, image_id, edge, indexed_at, model_id)) in rows.into_iter().enumerate() {
|
||||
let model_id = model_id.as_str();
|
||||
if seen.is_multiple_of(REPORT_EVERY) {
|
||||
progress(seen, total);
|
||||
}
|
||||
@@ -648,7 +710,8 @@ pub fn export_to_shards_reporting(
|
||||
}
|
||||
let mut fq = conn.prepare(
|
||||
"SELECT x, y, w, h, landmarks, detector_confidence, embedding, crop_px, crop,
|
||||
quality
|
||||
quality, eye_right, eye_right_px, eye_right_sharp,
|
||||
eye_left, eye_left_px, eye_left_sharp, sunglasses, landmarks_dense
|
||||
FROM faces WHERE image_id = ?1 AND model_id = ?2",
|
||||
)?;
|
||||
let faces: Vec<SharedFace> = fq
|
||||
@@ -666,6 +729,8 @@ pub fn export_to_shards_reporting(
|
||||
crop_px: r.get::<_, f64>(7)? as f32,
|
||||
crop: r.get::<_, Option<Vec<u8>>>(8)?.unwrap_or_default(),
|
||||
quality: r.get::<_, Option<f64>>(9)?.map(|q| q as f32),
|
||||
eyes: crate::faces::read_eyes(r, 10)?,
|
||||
landmarks_dense: r.get::<_, Option<Vec<u8>>>(17)?.unwrap_or_default(),
|
||||
})
|
||||
})?
|
||||
.collect::<Result<_, _>>()?;
|
||||
@@ -689,10 +754,20 @@ pub fn export_to_shards_reporting(
|
||||
/// adopted rather than re-detected, which is the difference between a new
|
||||
/// device being useful in a minute and in two hours.
|
||||
///
|
||||
/// Skips any image this device has already indexed itself. Local work is not
|
||||
/// second-guessed by a peer's — the two should agree, since the same model over
|
||||
/// the same proxy is deterministic, but where they do not, the copy this device
|
||||
/// computed is the one it can vouch for.
|
||||
/// Skips any image this device has already indexed itself under this
|
||||
/// pipeline or any sharing its embedder — unless the peer ran a detector that
|
||||
/// outranks the one that indexed it here. Local work is not second-guessed
|
||||
/// by a peer's equal: the two should agree, since the same model over the
|
||||
/// same proxy is deterministic, and where they do not, the copy this device
|
||||
/// computed is the one it can vouch for. A peer's *stronger* pass is another
|
||||
/// matter: it is the re-detection this device's own sweep would queue
|
||||
/// (`FaceDetector::supersedes`), already done, and taking it is what spares
|
||||
/// a tablet the fetch. Names survive the replacement by box overlap and
|
||||
/// embedding, as they do a local re-detection (`faces::record_detections`).
|
||||
///
|
||||
/// A peer's faces are taken under whichever compatible detector found them:
|
||||
/// a tablet set to the fast detector adopts the desktop's thorough pass
|
||||
/// rather than re-detecting it worse.
|
||||
///
|
||||
/// Returns how many images were adopted.
|
||||
pub fn import_from_shards(
|
||||
@@ -700,26 +775,49 @@ pub fn import_from_shards(
|
||||
store: &FaceShardStore,
|
||||
model_id: &str,
|
||||
) -> Result<usize, CatalogError> {
|
||||
use dr_types::FaceDetector;
|
||||
|
||||
// Only images this device actually has. A shard covers the whole account,
|
||||
// and a device holding a subset of the library should take only its own
|
||||
// part rather than accumulating faces for photographs it cannot show.
|
||||
let mut q = conn.prepare(
|
||||
"SELECT r.file_id, r.image_id
|
||||
//
|
||||
// With the pipeline that indexed each one here, or NULL: the marker is
|
||||
// what decides whether a peer's copy is a gap filled or an upgrade.
|
||||
let mut q = conn.prepare(&format!(
|
||||
"SELECT r.file_id, r.image_id,
|
||||
(SELECT fi.model_id FROM face_index fi
|
||||
WHERE fi.image_id = r.image_id AND {} = ?1)
|
||||
FROM remote r
|
||||
JOIN images i ON i.id = r.image_id
|
||||
WHERE i.trashed_at IS NULL
|
||||
AND NOT EXISTS (
|
||||
SELECT 1 FROM face_index fi
|
||||
WHERE fi.image_id = r.image_id AND fi.model_id = ?1
|
||||
)",
|
||||
)?;
|
||||
let candidates: Vec<(i64, i64)> = q
|
||||
.query_map([model_id], |r| Ok((r.get(0)?, r.get(1)?)))?
|
||||
WHERE i.trashed_at IS NULL",
|
||||
crate::faces::embedder_sql("fi.model_id")
|
||||
))?;
|
||||
let candidates: Vec<(i64, i64, Option<String>)> = q
|
||||
.query_map([crate::faces::embedder_of(model_id)], |r| {
|
||||
Ok((r.get(0)?, r.get(1)?, r.get(2)?))
|
||||
})?
|
||||
.collect::<Result<_, _>>()?;
|
||||
|
||||
let mut adopted = 0;
|
||||
for (file_id, image_id) in candidates {
|
||||
let Some((faces, edge)) = store.get_image(file_id as u64, model_id)? else {
|
||||
for (file_id, image_id, local) in candidates {
|
||||
let Some(held) = store.held_model(file_id as u64, model_id) else {
|
||||
continue;
|
||||
};
|
||||
if let Some(local) = local {
|
||||
// An unknown detector on either side cannot be ranked, and an
|
||||
// unranked peer is treated as an equal: kept out.
|
||||
let upgrade = match (
|
||||
FaceDetector::for_model_id(&held),
|
||||
FaceDetector::for_model_id(&local),
|
||||
) {
|
||||
(Some(theirs), Some(ours)) => theirs.outranks(ours),
|
||||
_ => false,
|
||||
};
|
||||
if !upgrade {
|
||||
continue;
|
||||
}
|
||||
}
|
||||
let Some((faces, edge)) = store.get_image(file_id as u64, &held)? else {
|
||||
continue;
|
||||
};
|
||||
// A peer that embedded before the quality was kept has done work this
|
||||
@@ -727,6 +825,12 @@ pub fn import_from_shards(
|
||||
// adopting the faces would write the run marker that keeps them from
|
||||
// ever being measured (schema V14). Left for this device's own pass —
|
||||
// or for the peer's, whose re-export replaces these.
|
||||
//
|
||||
// A missing *eye* reading is not the same case and is adopted. The
|
||||
// measuring pass finds those by the NULL, not by the marker, so
|
||||
// adopting the faces costs the reading nothing (schema V16) — and a
|
||||
// peer that has no eye models may be the only one that has done the
|
||||
// detection at all.
|
||||
if faces.iter().any(|f| f.quality.is_none()) {
|
||||
continue;
|
||||
}
|
||||
@@ -742,6 +846,8 @@ pub fn import_from_shards(
|
||||
embedding: f.embedding,
|
||||
crop_px: f.crop_px,
|
||||
quality: f.quality,
|
||||
eyes: f.eyes,
|
||||
landmarks_dense: f.landmarks_dense,
|
||||
model_id: f.model_id,
|
||||
// A peer that indexed before crops existed sends none, and the
|
||||
// reader falls back to the proxy exactly as it does for a face
|
||||
@@ -753,7 +859,7 @@ pub fn import_from_shards(
|
||||
crate::faces::record_detections(
|
||||
conn,
|
||||
dr_types::ImageId(image_id as u64),
|
||||
model_id,
|
||||
&held,
|
||||
edge,
|
||||
&local,
|
||||
)?;
|
||||
@@ -788,6 +894,8 @@ fn read_shared_face(r: &rusqlite::Row<'_>) -> rusqlite::Result<SharedFace> {
|
||||
crop_px: r.get::<_, f64>(9)? as f32,
|
||||
crop: r.get::<_, Option<Vec<u8>>>(10)?.unwrap_or_default(),
|
||||
quality: r.get::<_, Option<f64>>(11)?.map(|q| q as f32),
|
||||
eyes: crate::faces::read_eyes(r, 12)?,
|
||||
landmarks_dense: r.get::<_, Option<Vec<u8>>>(19)?.unwrap_or_default(),
|
||||
})
|
||||
}
|
||||
|
||||
@@ -875,7 +983,20 @@ CREATE TABLE IF NOT EXISTS faces (
|
||||
-- Length of the raw embedding (`faces::DetectedFace::quality`). NULL from
|
||||
-- a build that did not keep it, and a face the receiving device will not
|
||||
-- adopt -- see `import_from_shards`.
|
||||
quality REAL
|
||||
quality REAL,
|
||||
-- The eye reading (`faces::DetectedFace::eyes`), all seven or none. NULL
|
||||
-- from a peer without the eye models; adopted anyway, and read by the
|
||||
-- receiving device's own measuring pass if it has them.
|
||||
eye_right REAL,
|
||||
eye_right_px REAL,
|
||||
eye_right_sharp REAL,
|
||||
eye_left REAL,
|
||||
eye_left_px REAL,
|
||||
eye_left_sharp REAL,
|
||||
sunglasses REAL,
|
||||
-- The dense landmarks behind the reading (`faces::DetectedFace::
|
||||
-- landmarks_dense`), 424 bytes packed; NULL where none.
|
||||
landmarks_dense BLOB
|
||||
);
|
||||
CREATE INDEX IF NOT EXISTS faces_file ON faces(file_id, model_id);
|
||||
|
||||
@@ -935,6 +1056,8 @@ mod tests {
|
||||
embedding: vec![seed; 1024],
|
||||
crop_px: 180.0,
|
||||
quality: Some(17.5),
|
||||
eyes: None,
|
||||
landmarks_dense: Vec::new(),
|
||||
crop: vec![seed; 64],
|
||||
}
|
||||
}
|
||||
@@ -1144,6 +1267,8 @@ mod catalog_round_trip {
|
||||
embedding: vec![seed; 1024],
|
||||
crop_px: 180.0,
|
||||
quality: Some(20.0),
|
||||
eyes: None,
|
||||
landmarks_dense: Vec::new(),
|
||||
model_id: "w600k_mbf".into(),
|
||||
crop: vec![seed; 64],
|
||||
}
|
||||
@@ -1197,6 +1322,49 @@ mod catalog_round_trip {
|
||||
assert!(emb.iter().any(|e| e.embedding[0] == 1));
|
||||
}
|
||||
|
||||
/// The desktop switched to a stronger detector part-way through the
|
||||
/// library, so its faces sit under two pipeline ids. A tablet on the
|
||||
/// original detector must receive *all* of them — each under the id that
|
||||
/// found it — and not re-detect the thorough half worse.
|
||||
#[test]
|
||||
fn every_generation_sharing_an_embedder_travels_and_is_adopted() {
|
||||
let a = device(&[(1, 5001), (2, 5002)]);
|
||||
let b = device(&[(90, 5001), (91, 5002)]);
|
||||
|
||||
faces::record_detections(&a, dr_types::ImageId(1), "w600k_mbf", 1024, &[detected(1)])
|
||||
.unwrap();
|
||||
let mut thorough = detected(2);
|
||||
thorough.model_id = "scrfd_10g+w600k_mbf".into();
|
||||
faces::record_detections(
|
||||
&a,
|
||||
dr_types::ImageId(2),
|
||||
"scrfd_10g+w600k_mbf",
|
||||
1024,
|
||||
&[thorough],
|
||||
)
|
||||
.unwrap();
|
||||
|
||||
let mut store_a = FaceShardStore::open(&tempdir("a")).unwrap();
|
||||
assert_eq!(
|
||||
export_to_shards(&a, &mut store_a, "scrfd_10g+w600k_mbf").unwrap(),
|
||||
2,
|
||||
"the export left the earlier detector's images behind"
|
||||
);
|
||||
|
||||
let mut store_b = FaceShardStore::open(&tempdir("b")).unwrap();
|
||||
store_b.merge_shard(&store_a.shard_path(0)).unwrap();
|
||||
assert_eq!(import_from_shards(&b, &store_b, "w600k_mbf").unwrap(), 2);
|
||||
|
||||
assert_eq!(faces::coverage(&b, "w600k_mbf").unwrap().outstanding(), 0);
|
||||
let old = faces::for_image(&b, dr_types::ImageId(90)).unwrap();
|
||||
let new = faces::for_image(&b, dr_types::ImageId(91)).unwrap();
|
||||
assert_eq!(old[0].model_id, "w600k_mbf");
|
||||
assert_eq!(
|
||||
new[0].model_id, "scrfd_10g+w600k_mbf",
|
||||
"adopted under the wrong id"
|
||||
);
|
||||
}
|
||||
|
||||
/// A face a peer embedded without measuring it is work this device
|
||||
/// cannot finish, and adopting it would write the marker that stops it
|
||||
/// ever being measured. The image stays outstanding instead.
|
||||
@@ -1216,6 +1384,8 @@ mod catalog_round_trip {
|
||||
embedding: vec![1; 1024],
|
||||
crop_px: 180.0,
|
||||
quality,
|
||||
eyes: None,
|
||||
landmarks_dense: Vec::new(),
|
||||
crop: Vec::new(),
|
||||
};
|
||||
store
|
||||
@@ -1257,6 +1427,59 @@ mod catalog_round_trip {
|
||||
assert_eq!(emb[0].embedding[0], 9, "B's own embedding was overwritten");
|
||||
}
|
||||
|
||||
/// A peer's stronger detector is the re-detection this device would
|
||||
/// otherwise queue for itself. Taking it saves the fetch; the name the
|
||||
/// user confirmed here rides across on the box, as it would locally.
|
||||
#[test]
|
||||
fn a_peers_stronger_pass_replaces_a_weaker_local_one_and_keeps_the_name() {
|
||||
let a = device(&[(1, 5001)]);
|
||||
let b = device(&[(50, 5001)]);
|
||||
|
||||
let ids =
|
||||
faces::record_detections(&b, dr_types::ImageId(50), "w600k_mbf", 1024, &[detected(9)])
|
||||
.unwrap();
|
||||
let anna = faces::create_person(&b, "Anna").unwrap();
|
||||
faces::confirm(&b, ids[0], anna).unwrap();
|
||||
|
||||
let mut thorough = detected(7);
|
||||
thorough.model_id = "scrfd_10g+w600k_mbf".into();
|
||||
let mut second = detected(8);
|
||||
second.model_id = "scrfd_10g+w600k_mbf".into();
|
||||
second.x = 0.6;
|
||||
faces::record_detections(
|
||||
&a,
|
||||
dr_types::ImageId(1),
|
||||
"scrfd_10g+w600k_mbf",
|
||||
1024,
|
||||
&[thorough, second],
|
||||
)
|
||||
.unwrap();
|
||||
|
||||
let mut store = FaceShardStore::open(&tempdir("upgrade")).unwrap();
|
||||
export_to_shards(&a, &mut store, "scrfd_10g+w600k_mbf").unwrap();
|
||||
assert_eq!(import_from_shards(&b, &store, "w600k_mbf").unwrap(), 1);
|
||||
|
||||
let got = faces::for_image(&b, dr_types::ImageId(50)).unwrap();
|
||||
assert_eq!(got.len(), 2, "the stronger pass was not adopted");
|
||||
let named = got
|
||||
.iter()
|
||||
.find(|f| f.person == Some(anna))
|
||||
.expect("the name was lost");
|
||||
assert!(named.confirmed);
|
||||
assert_eq!(named.model_id, "scrfd_10g+w600k_mbf");
|
||||
|
||||
// And never downwards: A on the fast detector keeps B's thorough faces.
|
||||
let mut store_b = FaceShardStore::open(&tempdir("downgrade")).unwrap();
|
||||
faces::record_detections(&b, dr_types::ImageId(50), "w600k_mbf", 1024, &[detected(9)])
|
||||
.unwrap();
|
||||
export_to_shards(&b, &mut store_b, "w600k_mbf").unwrap();
|
||||
assert_eq!(
|
||||
import_from_shards(&a, &store_b, "scrfd_10g+w600k_mbf").unwrap(),
|
||||
0
|
||||
);
|
||||
assert_eq!(faces::for_image(&a, dr_types::ImageId(1)).unwrap().len(), 2);
|
||||
}
|
||||
|
||||
/// A device holding a subset of the library takes only its own part.
|
||||
#[test]
|
||||
fn a_device_ignores_faces_for_photographs_it_does_not_have() {
|
||||
@@ -1349,6 +1572,8 @@ mod catalog_round_trip {
|
||||
embedding: vec![seed; 1024],
|
||||
crop_px: 180.0,
|
||||
quality: None,
|
||||
eyes: None,
|
||||
landmarks_dense: Vec::new(),
|
||||
crop: vec![seed; 64],
|
||||
}
|
||||
}
|
||||
|
||||
+913
-137
File diff suppressed because it is too large
Load Diff
@@ -61,7 +61,7 @@ pub use collections::{Collection, CollectionKind, TreeRow};
|
||||
pub use dedup::{seen_by_content, seen_by_metadata, set_content_hash};
|
||||
pub use error::CatalogError;
|
||||
pub use face_shard::{FaceShardStore, SharedFace};
|
||||
pub use faces::{Calibration, DetectedFace, Face, FaceId, Measurement, Person, PersonId};
|
||||
pub use faces::{Calibration, DetectedFace, Face, FaceId, FaceUpdate, Person, PersonId};
|
||||
pub use jobs::{Job, JobKind, Priority};
|
||||
pub use keywords::{Coverage, Keyword, KeywordId, SelectionKeyword};
|
||||
pub use merge::MergeReport;
|
||||
|
||||
@@ -991,9 +991,13 @@ fn remote_has_column(tx: &Connection, table: &str, column: &str) -> Result<bool,
|
||||
/// Remote face row id to local face row id, by photograph and box overlap.
|
||||
///
|
||||
/// See [`merge_people_within`] for why a face has no shared identity and this
|
||||
/// has to be derived. Only faces from the same model are compared: boxes from
|
||||
/// two different detectors are not the same measurement, and matching across
|
||||
/// them would attach a judgement to a face nobody looked at.
|
||||
/// has to be derived. Faces are compared within an *embedder*
|
||||
/// (`faces::embedder_of`), not within an exact pipeline id: two detectors in
|
||||
/// front of the same embedder draw boxes around the same faces, and a
|
||||
/// confirmation made on one device's box is about the face, not the
|
||||
/// rectangle — the same judgement `faces::record_detections` makes when it
|
||||
/// carries a confirmation across a re-detection. Keying on the exact id was
|
||||
/// what let a detector change strand every name on the device that made it.
|
||||
fn match_faces(tx: &Connection) -> Result<std::collections::HashMap<i64, i64>, CatalogError> {
|
||||
/// Loose on purpose — "the same face in the frame", not "the same
|
||||
/// rectangle". The figure `record_detections` uses for the same job.
|
||||
@@ -1026,7 +1030,8 @@ fn match_faces(tx: &Connection) -> Result<std::collections::HashMap<i64, i64>, C
|
||||
})?;
|
||||
for row in rows {
|
||||
let (file_id, model, boxed) = row?;
|
||||
local.entry((file_id, model)).or_default().push(boxed);
|
||||
let embedder = crate::faces::embedder_of(&model).to_string();
|
||||
local.entry((file_id, embedder)).or_default().push(boxed);
|
||||
}
|
||||
}
|
||||
if local.is_empty() {
|
||||
@@ -1056,7 +1061,8 @@ fn match_faces(tx: &Connection) -> Result<std::collections::HashMap<i64, i64>, C
|
||||
|
||||
for row in rows {
|
||||
let (remote_id, file_id, model, rbox) = row?;
|
||||
let Some(candidates) = local.get(&(file_id, model)) else {
|
||||
let embedder = crate::faces::embedder_of(&model).to_string();
|
||||
let Some(candidates) = local.get(&(file_id, embedder)) else {
|
||||
continue;
|
||||
};
|
||||
let best = candidates
|
||||
@@ -1977,6 +1983,54 @@ mod tests {
|
||||
assert_eq!(person_of(&c, local), Some(("Anna".to_string(), true)));
|
||||
}
|
||||
|
||||
/// The bug this rule exists for: the desktop switched to a stronger
|
||||
/// detector and confirmed 3,500 faces under the old pipeline id; the
|
||||
/// tablet held the same faces under the new one, and not one name
|
||||
/// crossed, because the match demanded the exact id. Same photograph,
|
||||
/// same box, same embedder — that is the same face.
|
||||
#[test]
|
||||
fn a_confirmation_crosses_a_detector_change() {
|
||||
let c = two_catalogs();
|
||||
for db in ["main", "remote_cat"] {
|
||||
add_synced_image(&c, db, 1, 5000);
|
||||
}
|
||||
let local = add_face(&c, "main", 7, 1, 0.30);
|
||||
c.execute(
|
||||
"UPDATE main.faces SET model_id = 'scrfd_10g+w600k_mbf' WHERE id = ?1",
|
||||
[local],
|
||||
)
|
||||
.unwrap();
|
||||
let remote = add_face(&c, "remote_cat", 42, 1, 0.31);
|
||||
add_person(&c, "remote_cat", 3, "u-anna", "Anna", false);
|
||||
assign(&c, "remote_cat", remote, 3, true);
|
||||
|
||||
let report = merge_all(&c).unwrap();
|
||||
assert_eq!(report.faces_assigned, 1);
|
||||
assert_eq!(person_of(&c, local), Some(("Anna".to_string(), true)));
|
||||
}
|
||||
|
||||
/// A different embedder is a different space, and a box there is a face
|
||||
/// nobody here has a vector for.
|
||||
#[test]
|
||||
fn a_confirmation_does_not_cross_an_embedder_change() {
|
||||
let c = two_catalogs();
|
||||
for db in ["main", "remote_cat"] {
|
||||
add_synced_image(&c, db, 1, 5000);
|
||||
}
|
||||
let local = add_face(&c, "main", 7, 1, 0.30);
|
||||
c.execute(
|
||||
"UPDATE main.faces SET model_id = 'scrfd_10g+other_embedder' WHERE id = ?1",
|
||||
[local],
|
||||
)
|
||||
.unwrap();
|
||||
let remote = add_face(&c, "remote_cat", 42, 1, 0.31);
|
||||
add_person(&c, "remote_cat", 3, "u-anna", "Anna", false);
|
||||
assign(&c, "remote_cat", remote, 3, true);
|
||||
|
||||
merge_all(&c).unwrap();
|
||||
assert_eq!(person_of(&c, local), None, "matched across embedders");
|
||||
}
|
||||
|
||||
/// Boxes from two devices are close but not identical. Matching has to be
|
||||
/// by overlap, not equality, or nothing ever lines up.
|
||||
#[test]
|
||||
|
||||
+248
-12
@@ -15,7 +15,7 @@ use rusqlite::Connection;
|
||||
use crate::error::CatalogError;
|
||||
|
||||
/// Schema version this build writes and understands.
|
||||
pub const SCHEMA_VERSION: i64 = 15;
|
||||
pub const SCHEMA_VERSION: i64 = 18;
|
||||
|
||||
/// Apply migrations up to [`SCHEMA_VERSION`].
|
||||
///
|
||||
@@ -143,9 +143,61 @@ pub fn migrate(conn: &Connection) -> Result<i64, CatalogError> {
|
||||
tx.commit()?;
|
||||
}
|
||||
|
||||
if from < 16 {
|
||||
let tx = conn.unchecked_transaction()?;
|
||||
// Guarded like V14's column, and for the same reason: `ALTER TABLE
|
||||
// ... ADD COLUMN` has no `IF NOT EXISTS`, and this step must be
|
||||
// re-enterable (NFR-R5).
|
||||
for column in EYE_COLUMNS {
|
||||
let present: bool = tx
|
||||
.prepare("SELECT 1 FROM pragma_table_info('faces') WHERE name = ?1")?
|
||||
.exists([column])?;
|
||||
if !present {
|
||||
tx.execute_batch(&format!("ALTER TABLE faces ADD COLUMN {column} REAL;"))?;
|
||||
}
|
||||
}
|
||||
tx.pragma_update(None, "user_version", 16)?;
|
||||
tx.commit()?;
|
||||
}
|
||||
|
||||
if from < 17 {
|
||||
let tx = conn.unchecked_transaction()?;
|
||||
tx.execute_batch(V17)?;
|
||||
tx.pragma_update(None, "user_version", 17)?;
|
||||
tx.commit()?;
|
||||
}
|
||||
|
||||
if from < 18 {
|
||||
let tx = conn.unchecked_transaction()?;
|
||||
// Guarded like V14's and V16's columns: ALTER has no IF NOT EXISTS
|
||||
// and the step must be re-enterable (NFR-R5).
|
||||
let present: bool = tx
|
||||
.prepare("SELECT 1 FROM pragma_table_info('faces') WHERE name = 'landmarks_dense'")?
|
||||
.exists([])?;
|
||||
if !present {
|
||||
tx.execute_batch("ALTER TABLE faces ADD COLUMN landmarks_dense BLOB;")?;
|
||||
}
|
||||
tx.pragma_update(None, "user_version", 18)?;
|
||||
tx.commit()?;
|
||||
}
|
||||
|
||||
Ok(from)
|
||||
}
|
||||
|
||||
/// The seven columns V16 adds to `faces`, in the order the readers name them.
|
||||
///
|
||||
/// Named once because three places have to agree on them: this migration,
|
||||
/// [`for_attached`], and the face shard's own catch-up (`face_shard`).
|
||||
pub const EYE_COLUMNS: [&str; 7] = [
|
||||
"eye_right",
|
||||
"eye_right_px",
|
||||
"eye_right_sharp",
|
||||
"eye_left",
|
||||
"eye_left_px",
|
||||
"eye_left_sharp",
|
||||
"sunglasses",
|
||||
];
|
||||
|
||||
/// Recompute columns a migration added, for rows that predate it.
|
||||
///
|
||||
/// A migration adds a column with a default; it cannot know what the value
|
||||
@@ -208,11 +260,38 @@ pub fn backfill(conn: &Connection) -> Result<Vec<(&'static str, usize)>, Catalog
|
||||
Ok(out)
|
||||
}
|
||||
|
||||
/// How long a connection waits for a writer to finish before giving up.
|
||||
///
|
||||
/// TRACES: NFR-R1
|
||||
/// SQLite's default is **zero** — the loser of a race gets `SQLITE_BUSY` at
|
||||
/// once rather than a turn — and WAL does not change that for two writers. One
|
||||
/// writer and many readers is the case WAL makes free; this is the other one,
|
||||
/// and this application has it constantly: the face sweep commits a batch while
|
||||
/// reclustering reads, the derived sync imports shards while the sweep writes.
|
||||
///
|
||||
/// Without a timeout that contention was *lost work*, not a retry. A face
|
||||
/// sweep that had already paid for the detection and the embedding — the
|
||||
/// expensive part, seconds per image — threw the result away on
|
||||
/// `storing faces for 214: database is locked` and moved on, and both the
|
||||
/// desktop and the tablet logged runs of those on consecutive images.
|
||||
///
|
||||
/// Ten seconds, matching the figure the job runner's tests already use for the
|
||||
/// same reason. It is far longer than any transaction here (a sweep batch is
|
||||
/// sub-second; the slowest is a WAL checkpoint of a 130 MB catalog), so in
|
||||
/// practice it is a bound on pathology rather than a wait anyone sits through.
|
||||
/// The tension with NFR-P9 is real but one-sided: a query on the UI thread
|
||||
/// would rather wait for its turn than fail, because the failure is what the
|
||||
/// user sees as "cannot open catalog".
|
||||
const BUSY_TIMEOUT: std::time::Duration = std::time::Duration::from_secs(10);
|
||||
|
||||
/// Connection setup applied on every open, migration or not.
|
||||
///
|
||||
/// WAL is required by NFR-R1: it survives power loss without corruption, and
|
||||
/// it lets a background job write while the grid reads.
|
||||
pub fn configure(conn: &Connection) -> Result<(), CatalogError> {
|
||||
// Before the pragmas, so that a connection racing a migration waits for it
|
||||
// rather than failing on the first statement it tries.
|
||||
conn.busy_timeout(BUSY_TIMEOUT)?;
|
||||
conn.pragma_update(None, "journal_mode", "WAL")?;
|
||||
// NORMAL rather than FULL: with WAL this is durable across process death
|
||||
// (which is what FR-PLAT-AND-3 cares about) and only risks the last
|
||||
@@ -276,7 +355,15 @@ pub fn for_attached(schema_name: &str) -> String {
|
||||
"{}\n{}\n{}\n\
|
||||
ALTER TABLE {schema_name}.people ADD COLUMN ignored INTEGER NOT NULL DEFAULT 0;\n\
|
||||
ALTER TABLE {schema_name}.faces ADD COLUMN crop BLOB;\n\
|
||||
ALTER TABLE {schema_name}.faces ADD COLUMN quality REAL;",
|
||||
ALTER TABLE {schema_name}.faces ADD COLUMN quality REAL;\n\
|
||||
ALTER TABLE {schema_name}.faces ADD COLUMN eye_right REAL;\n\
|
||||
ALTER TABLE {schema_name}.faces ADD COLUMN eye_right_px REAL;\n\
|
||||
ALTER TABLE {schema_name}.faces ADD COLUMN eye_right_sharp REAL;\n\
|
||||
ALTER TABLE {schema_name}.faces ADD COLUMN eye_left REAL;\n\
|
||||
ALTER TABLE {schema_name}.faces ADD COLUMN eye_left_px REAL;\n\
|
||||
ALTER TABLE {schema_name}.faces ADD COLUMN eye_left_sharp REAL;\n\
|
||||
ALTER TABLE {schema_name}.faces ADD COLUMN sunglasses REAL;\n\
|
||||
ALTER TABLE {schema_name}.faces ADD COLUMN landmarks_dense BLOB;",
|
||||
rewrite_for_attached(V1, schema_name),
|
||||
rewrite_for_attached(V6, schema_name),
|
||||
rewrite_for_attached(V8, schema_name),
|
||||
@@ -614,20 +701,20 @@ const V14: &str = r#"
|
||||
-- face this rule is not yet protecting anyone from, and the only way to
|
||||
-- measure it is to embed it again.
|
||||
--
|
||||
-- The sweep's measuring pass is what does that: `dr_ui::library::
|
||||
-- faces_unmeasured` lists every image holding a face with no reading, and
|
||||
-- each face is embedded again from the native render with the landmarks it
|
||||
-- already has, the raw vector written over the old one (`record_measurements`)
|
||||
-- and nothing else touched -- not the id, not the box, not who the user said
|
||||
-- it was. The faces keep drawing the People screen throughout.
|
||||
-- The `face-quality` repair is what does that (`dr_ui::repairs`, once the
|
||||
-- sweep's measuring pass): it lists every face with no reading, and each is
|
||||
-- embedded again from the native render with the landmarks it already has,
|
||||
-- the raw vector written over the old one (`record_updates`) and nothing
|
||||
-- else touched -- not the id, not the box, not who the user said it was.
|
||||
-- The faces keep drawing the People screen throughout.
|
||||
--
|
||||
-- The run markers of those images are forgotten too, exactly as V12 forgot
|
||||
-- the runs made against too small a proxy. The build this shipped in had no
|
||||
-- measuring pass yet, and a marker is the one thing that stops a face ever
|
||||
-- being looked at again; with the pass in place `faces_unindexed` leaves
|
||||
-- these images to it rather than detecting them from scratch, so the
|
||||
-- deletion costs nothing -- and an image that was examined and found empty
|
||||
-- keeps its marker, since there is nothing on it to measure.
|
||||
-- being looked at again; with the repair in place, detection leaves an
|
||||
-- image holding this embedder's faces to it rather than detecting from
|
||||
-- scratch, so the deletion costs nothing -- and an image that was examined
|
||||
-- and found empty keeps its marker, since there is nothing on it to measure.
|
||||
--
|
||||
-- The cost is a re-fetch of every image with a face on it, on the next pass
|
||||
-- the user starts. That is a whole-library transfer (FR-NC-6), and it starts
|
||||
@@ -679,6 +766,72 @@ CREATE TABLE IF NOT EXISTS xmp_conflicts (
|
||||
);
|
||||
"#;
|
||||
|
||||
// V18 -- TRACES: FR-CULL-8a | FR-CULL-12
|
||||
//
|
||||
// The 106 dense landmarks the eye pass reads its eye boxes from, kept beside
|
||||
// the reading as `dr_face::Landmarks::to_packed_bytes`: 106 x (x, y) as
|
||||
// 16-bit fixed point over the frame, 424 bytes a face, a seventh of a
|
||||
// pixel on a 6000-pixel frame. Derived data under FR-CULL-12 -- rebuilt by
|
||||
// re-reading, never in a sidecar -- and stored for the same reason the
|
||||
// embedding is: it cost a fetch of the original and a model run, and the
|
||||
// next per-face pass (head pose, expression) should not have to pay either
|
||||
// again. NULL where the face was never read.
|
||||
//
|
||||
// Added in `migrate`, guarded, like every ALTER here (NFR-R5).
|
||||
|
||||
const V17: &str = r#"
|
||||
-- TRACES: FR-CULL-8a | FR-CULL-13 | NFR-P9
|
||||
-- The eyes-open filter's index, and a lesson about where a column lands.
|
||||
--
|
||||
-- The people filter is a correlated EXISTS over `faces` per image, and it
|
||||
-- was fast because `faces_image` *covers* it: the subquery never touched a
|
||||
-- row. Reading V16's seven eye columns in the same subquery did touch the
|
||||
-- row -- and `ALTER TABLE ADD COLUMN` puts a column at the end of the
|
||||
-- record, after the 1 KB embedding and the ~5 KB crop, so every check
|
||||
-- dragged six kilobytes off disk to reach seven floats. Measured on the
|
||||
-- reference library: 24 seconds for one count, thirteen of them system
|
||||
-- time. With this index the same count takes five milliseconds, because
|
||||
-- the subquery is served from the index again and never reads a row.
|
||||
--
|
||||
-- The columns are listed in EYE_COLUMNS' order behind `image_id`, which is
|
||||
-- the key the subquery searches on. Nothing else changed in V17; a catalog
|
||||
-- already at V16 needs only this.
|
||||
CREATE INDEX IF NOT EXISTS faces_eyes ON faces(
|
||||
image_id, eye_right, eye_right_px, eye_right_sharp,
|
||||
eye_left, eye_left_px, eye_left_sharp, sunglasses
|
||||
);
|
||||
"#;
|
||||
|
||||
// V16 -- TRACES: FR-CULL-8a
|
||||
//
|
||||
// What each face's eyes are doing: for each eye P(open), the source pixels
|
||||
// across its box and the sharpness of the patch the classifier saw; and
|
||||
// P(sunglasses) for the head. Seven numbers rather than a verdict, because
|
||||
// the verdict is a rule with thresholds in it (dr_face::eyes::EyeReading::
|
||||
// state) and a rule belongs in code that can be changed, not in rows that
|
||||
// would have to be re-measured.
|
||||
//
|
||||
// The pixels and the sharpness are what stop a smear reading as a blink: an
|
||||
// eye too small or too soft to read is not asked, and a face with no
|
||||
// readable eye is "unclear", which no filter drops. Sunglasses are a column
|
||||
// of their own for the same kind of reason — the eye classifier answers
|
||||
// confidently over dark glass, and its answer means nothing there. A filter
|
||||
// for "eyes open" reads all seven.
|
||||
//
|
||||
// NULL means "never measured" -- a face indexed before this version, or on a
|
||||
// device without the eye models -- and a NULL is left alone by every filter
|
||||
// that reads these, so an old library does not empty its grid the moment the
|
||||
// chip is pressed. The sweep's measuring pass fills them in, from the native
|
||||
// render, with the landmarks already stored: the same pass V14 built for the
|
||||
// embedding's length, extended to ask the eye models too. No run marker is
|
||||
// forgotten here, for the reason V14's note gives -- the measuring pass
|
||||
// finds its own work by the NULL, and deleting markers would only put the
|
||||
// detector back over images it has finished with.
|
||||
//
|
||||
// The columns are added in `migrate`, guarded, because ALTER has no IF NOT
|
||||
// EXISTS and the step has to be re-enterable (NFR-R5). Their names are
|
||||
// `EYE_COLUMNS`.
|
||||
|
||||
const V9: &str = r#"
|
||||
-- TRACES: FR-CULL-8
|
||||
-- A record that face detection has *run* on an image, distinct from what it
|
||||
@@ -1148,6 +1301,60 @@ CREATE INDEX jobs_ready ON jobs(state, priority DESC, not_before);
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
|
||||
#[test]
|
||||
fn a_writer_waits_for_its_turn_rather_than_losing_its_work() {
|
||||
// The failure this exists for: a face sweep that had already paid for
|
||||
// the detection and the embedding threw the result away on
|
||||
// "database is locked" and moved on. WAL does not help here — it makes
|
||||
// one writer and many readers free, and this is two writers.
|
||||
let dir = std::env::temp_dir().join(format!(
|
||||
"dr-busy-{}-{:?}",
|
||||
std::process::id(),
|
||||
std::thread::current().id()
|
||||
));
|
||||
let _ = std::fs::remove_dir_all(&dir);
|
||||
std::fs::create_dir_all(&dir).unwrap();
|
||||
let path = dir.join("catalog.sqlite");
|
||||
|
||||
let held = rusqlite::Connection::open(&path).unwrap();
|
||||
configure(&held).unwrap();
|
||||
migrate(&held).unwrap();
|
||||
|
||||
let other = rusqlite::Connection::open(&path).unwrap();
|
||||
configure(&other).unwrap();
|
||||
|
||||
// Every connection carries the timeout, which is what makes the wait
|
||||
// below a wait rather than an immediate error.
|
||||
let timeout: i64 = other
|
||||
.query_row("PRAGMA busy_timeout", [], |r| r.get(0))
|
||||
.unwrap();
|
||||
assert_eq!(timeout, BUSY_TIMEOUT.as_millis() as i64);
|
||||
|
||||
// A writer holds the database; the other one must still get its turn
|
||||
// once the first commits, rather than failing at the moment it asks.
|
||||
let writing = held.unchecked_transaction().unwrap();
|
||||
held.execute(
|
||||
"INSERT INTO roots(id, kind, label) VALUES (1, 'local', 'lib')",
|
||||
[],
|
||||
)
|
||||
.unwrap();
|
||||
|
||||
let handle = std::thread::spawn(move || {
|
||||
other.execute(
|
||||
"INSERT INTO roots(id, kind, label) VALUES (2, 'local', 'two')",
|
||||
[],
|
||||
)
|
||||
});
|
||||
std::thread::sleep(std::time::Duration::from_millis(150));
|
||||
writing.commit().unwrap();
|
||||
|
||||
assert!(
|
||||
handle.join().unwrap().is_ok(),
|
||||
"the second writer waited and then wrote, rather than erroring"
|
||||
);
|
||||
let _ = std::fs::remove_dir_all(&dir);
|
||||
}
|
||||
use super::*;
|
||||
|
||||
fn mem() -> Connection {
|
||||
@@ -1435,6 +1642,35 @@ mod tests {
|
||||
assert_eq!(migrate(&c).unwrap(), SCHEMA_VERSION);
|
||||
}
|
||||
|
||||
/// V16 adds its columns guarded, so a catalog whose version was rewound
|
||||
/// after the columns landed — the rollback NFR-R5 contemplates — migrates
|
||||
/// again rather than failing on "duplicate column".
|
||||
#[test]
|
||||
fn the_eye_columns_survive_a_rewound_version() {
|
||||
let c = mem();
|
||||
migrate(&c).unwrap();
|
||||
for column in EYE_COLUMNS {
|
||||
let present: bool = c
|
||||
.prepare("SELECT 1 FROM pragma_table_info('faces') WHERE name = ?1")
|
||||
.unwrap()
|
||||
.exists([column])
|
||||
.unwrap();
|
||||
assert!(present, "{column} missing after migration");
|
||||
}
|
||||
c.pragma_update(None, "user_version", 15).unwrap();
|
||||
assert_eq!(migrate(&c).unwrap(), 15);
|
||||
let indexed: bool = c
|
||||
.prepare("SELECT 1 FROM sqlite_master WHERE type = 'index' AND name = 'faces_eyes'")
|
||||
.unwrap()
|
||||
.exists([])
|
||||
.unwrap();
|
||||
assert!(indexed, "V17's covering index is there");
|
||||
let v: i64 = c
|
||||
.query_row("PRAGMA user_version", [], |r| r.get(0))
|
||||
.unwrap();
|
||||
assert_eq!(v, SCHEMA_VERSION);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn refuses_a_catalog_from_a_newer_build() {
|
||||
let c = mem();
|
||||
|
||||
@@ -0,0 +1,221 @@
|
||||
//! TRACES: S15 | FR-MRG-3
|
||||
//! Spike S15.1 — does rawler read back a linear DNG this application writes?
|
||||
//!
|
||||
//! cargo run -p dr-decode --example linear_dng [-- <out.dng>]
|
||||
//!
|
||||
//! Decides FR-MRG-3's container. A panorama composite is three linear samples
|
||||
//! per pixel with a camera matrix attached, which is exactly what a
|
||||
//! `LinearRaw` DNG is; if rawler parses one, the composite re-enters the
|
||||
//! library as `Format::Dng` and the only new decode work is a `cpp == 3`
|
||||
//! branch. If it does not, the container is a float TIFF with a decode path
|
||||
//! of its own.
|
||||
//!
|
||||
//! The file is hand-rolled rather than written with the `tiff` crate, whose
|
||||
//! encoder fixes `PhotometricInterpretation` to RGB and cannot say
|
||||
//! `LinearRaw`. Eighty lines of IFD is the cheaper thing to own than a fork.
|
||||
|
||||
use rawler::rawsource::RawSource;
|
||||
|
||||
const W: u32 = 64;
|
||||
const H: u32 = 48;
|
||||
|
||||
fn main() {
|
||||
let bytes = write_linear_dng(W, H);
|
||||
if let Some(path) = std::env::args().nth(1) {
|
||||
std::fs::write(&path, &bytes).expect("write");
|
||||
println!("wrote {path} ({} bytes)", bytes.len());
|
||||
}
|
||||
|
||||
let source = RawSource::new_from_slice(&bytes);
|
||||
let decoder = match rawler::get_decoder(&source) {
|
||||
Ok(d) => d,
|
||||
Err(e) => {
|
||||
println!("FAIL get_decoder: {e}");
|
||||
std::process::exit(1);
|
||||
}
|
||||
};
|
||||
println!("ok decoder found");
|
||||
|
||||
let image = match decoder.raw_image(&source, &Default::default(), false) {
|
||||
Ok(i) => i,
|
||||
Err(e) => {
|
||||
println!("FAIL raw_image: {e}");
|
||||
std::process::exit(1);
|
||||
}
|
||||
};
|
||||
println!(
|
||||
"ok raw_image: {}×{}, cpp {}, bps {}, {} samples, make {:?} model {:?}",
|
||||
image.width,
|
||||
image.height,
|
||||
image.cpp,
|
||||
image.bps,
|
||||
match &image.data {
|
||||
rawler::RawImageData::Integer(v) => v.len(),
|
||||
rawler::RawImageData::Float(v) => v.len(),
|
||||
},
|
||||
image.make,
|
||||
image.model
|
||||
);
|
||||
println!(
|
||||
" white {:?} black {:?} wb {:?}",
|
||||
image.whitelevel.0,
|
||||
image
|
||||
.blacklevel
|
||||
.levels
|
||||
.iter()
|
||||
.map(|r| r.n as f32 / r.d.max(1) as f32)
|
||||
.collect::<Vec<_>>(),
|
||||
image.wb_coeffs
|
||||
);
|
||||
|
||||
// The pixel at (1, 0) was written as (1000, 2000, 3000): if the samples
|
||||
// come back interleaved in that order, cpp == 3 means what it says.
|
||||
if let rawler::RawImageData::Integer(v) = &image.data {
|
||||
let i = image.cpp;
|
||||
println!(" pixel (1,0) = {:?}", &v[i..i + image.cpp.min(3)]);
|
||||
}
|
||||
|
||||
// What dr-decode itself makes of it: the colour matrix rawler parsed into
|
||||
// the camera definition, and the profile the decoder would build from it.
|
||||
println!(" rawler color_matrix: {:?}", image.camera.color_matrix);
|
||||
let dng = dr_decode::profile::read_dng_matrices(decoder.as_ref());
|
||||
let profile = dr_decode::CameraProfile::extract(&image, &dng);
|
||||
println!(
|
||||
" CameraProfile: {}",
|
||||
profile
|
||||
.as_ref()
|
||||
.map(|p| format!("xyz_to_cam {:?}", p.xyz_to_cam()))
|
||||
.unwrap_or_else(|| "none".into())
|
||||
);
|
||||
|
||||
match dr_decode::decode(&bytes) {
|
||||
Ok(r) => println!(
|
||||
"note dr_decode::decode accepted it as CFA: {}×{}, {} samples — the cpp==3 branch is the work",
|
||||
r.width,
|
||||
r.height,
|
||||
r.data.len()
|
||||
),
|
||||
Err(e) => println!("note dr_decode::decode refused it: {e} — the cpp==3 branch is the work"),
|
||||
}
|
||||
}
|
||||
|
||||
/// A minimal `LinearRaw` DNG: one IFD, uncompressed 16-bit RGB, the tags a
|
||||
/// decoder needs to treat it as a DNG and the matrix a develop chain needs
|
||||
/// to treat it as a camera. Little-endian, one strip.
|
||||
fn write_linear_dng(w: u32, h: u32) -> Vec<u8> {
|
||||
// Pixels first, so their offset is known: a ramp with one marker pixel.
|
||||
let mut pixels: Vec<u16> = Vec::with_capacity((w * h * 3) as usize);
|
||||
for y in 0..h {
|
||||
for x in 0..w {
|
||||
if (x, y) == (1, 0) {
|
||||
pixels.extend([1000, 2000, 3000]);
|
||||
} else {
|
||||
let v = ((x + y) * 512).min(65535) as u16;
|
||||
pixels.extend([v, v / 2, v / 3]);
|
||||
}
|
||||
}
|
||||
}
|
||||
let pixel_bytes: Vec<u8> = pixels.iter().flat_map(|v| v.to_le_bytes()).collect();
|
||||
|
||||
// Layout: header (8) | pixels | extra data | IFD.
|
||||
let pixels_off = 8u32;
|
||||
let extra_off = pixels_off + pixel_bytes.len() as u32;
|
||||
|
||||
// Values that do not fit in four bytes go in `extra`, and the entry
|
||||
// points at them.
|
||||
let mut extra: Vec<u8> = Vec::new();
|
||||
let mut entries: Vec<(u16, u16, u32, [u8; 4])> = Vec::new();
|
||||
|
||||
fn short(tag: u16, v: u16) -> (u16, u16, u32, [u8; 4]) {
|
||||
let mut b = [0u8; 4];
|
||||
b[..2].copy_from_slice(&v.to_le_bytes());
|
||||
(tag, 3, 1, b)
|
||||
}
|
||||
fn long(tag: u16, v: u32) -> (u16, u16, u32, [u8; 4]) {
|
||||
(tag, 4, 1, v.to_le_bytes())
|
||||
}
|
||||
fn ascii(extra: &mut Vec<u8>, extra_off: u32, tag: u16, s: &str) -> (u16, u16, u32, [u8; 4]) {
|
||||
let mut bytes = s.as_bytes().to_vec();
|
||||
bytes.push(0);
|
||||
let off = extra_off + extra.len() as u32;
|
||||
extra.extend(&bytes);
|
||||
(tag, 2, bytes.len() as u32, off.to_le_bytes())
|
||||
}
|
||||
|
||||
entries.push(long(254, 0)); // NewSubfileType: main image
|
||||
entries.push(long(256, w));
|
||||
entries.push(long(257, h));
|
||||
// BitsPerSample ×3 — three shorts, six bytes, so out of line.
|
||||
{
|
||||
let off = extra_off + extra.len() as u32;
|
||||
for _ in 0..3 {
|
||||
extra.extend(16u16.to_le_bytes());
|
||||
}
|
||||
entries.push((258, 3, 3, off.to_le_bytes()));
|
||||
}
|
||||
entries.push(short(259, 1)); // Compression: none
|
||||
entries.push(short(262, 34892)); // PhotometricInterpretation: LinearRaw
|
||||
entries.push(ascii(&mut extra, extra_off, 271, "DarkRoom"));
|
||||
entries.push(ascii(&mut extra, extra_off, 272, "Panorama"));
|
||||
entries.push(long(273, pixels_off)); // StripOffsets
|
||||
entries.push(short(274, 1)); // Orientation
|
||||
entries.push(short(277, 3)); // SamplesPerPixel
|
||||
entries.push(long(278, h)); // RowsPerStrip
|
||||
entries.push(long(279, pixel_bytes.len() as u32)); // StripByteCounts
|
||||
entries.push(short(284, 1)); // PlanarConfiguration: chunky
|
||||
entries.push((50706, 1, 4, [1, 4, 0, 0])); // DNGVersion
|
||||
entries.push((50707, 1, 4, [1, 4, 0, 0])); // DNGBackwardVersion
|
||||
entries.push(ascii(&mut extra, extra_off, 50708, "DarkRoom Panorama")); // UniqueCameraModel
|
||||
entries.push(long(50717, 65535)); // WhiteLevel
|
||||
|
||||
// ColorMatrix1: XYZ → camera, 9 SRATIONALs. A plausible sRGB-ish matrix
|
||||
// (the inverse of the sRGB D65 primaries), scaled to integers.
|
||||
{
|
||||
let m: [(i32, i32); 9] = [
|
||||
(32406, 10000),
|
||||
(-15372, 10000),
|
||||
(-4986, 10000),
|
||||
(-9689, 10000),
|
||||
(18758, 10000),
|
||||
(415, 10000),
|
||||
(557, 10000),
|
||||
(-2040, 10000),
|
||||
(10570, 10000),
|
||||
];
|
||||
let off = extra_off + extra.len() as u32;
|
||||
for (n, d) in m {
|
||||
extra.extend(n.to_le_bytes());
|
||||
extra.extend(d.to_le_bytes());
|
||||
}
|
||||
entries.push((50721, 10, 9, off.to_le_bytes()));
|
||||
}
|
||||
// AsShotNeutral: 3 RATIONALs, neutral.
|
||||
{
|
||||
let off = extra_off + extra.len() as u32;
|
||||
for _ in 0..3 {
|
||||
extra.extend(1u32.to_le_bytes());
|
||||
extra.extend(1u32.to_le_bytes());
|
||||
}
|
||||
entries.push((50728, 5, 3, off.to_le_bytes()));
|
||||
}
|
||||
entries.push(short(50778, 21)); // CalibrationIlluminant1: D65
|
||||
|
||||
entries.sort_by_key(|e| e.0);
|
||||
|
||||
let ifd_off = extra_off + extra.len() as u32;
|
||||
let mut out = Vec::new();
|
||||
out.extend(b"II");
|
||||
out.extend(42u16.to_le_bytes());
|
||||
out.extend(ifd_off.to_le_bytes());
|
||||
out.extend(&pixel_bytes);
|
||||
out.extend(&extra);
|
||||
out.extend((entries.len() as u16).to_le_bytes());
|
||||
for (tag, ty, count, value) in &entries {
|
||||
out.extend(tag.to_le_bytes());
|
||||
out.extend(ty.to_le_bytes());
|
||||
out.extend(count.to_le_bytes());
|
||||
out.extend(value);
|
||||
}
|
||||
out.extend(0u32.to_le_bytes()); // no next IFD
|
||||
out
|
||||
}
|
||||
@@ -25,6 +25,42 @@ pub enum DecodeError {
|
||||
CorruptPreview(String),
|
||||
}
|
||||
|
||||
/// Run a decoder call, and return a panic inside it as an error.
|
||||
///
|
||||
/// TRACES: FR-RAW-4 | NFR-SEC-1 | NFR-R3
|
||||
/// rawler `panic!`s on some malformed input rather than returning `Err` — a
|
||||
/// DNG whose IFD claims a >50000 px image, for one, which is in the reference
|
||||
/// library. A panic on a worker thread ends the thread: the face sweep that
|
||||
/// met that file stopped 13 seconds in, three sweeps running, with "17301
|
||||
/// image(s) to index" as the last word and nothing to say why. FR-RAW-4's
|
||||
/// rule — a malformed file must not abort a batch — is this crate's to keep
|
||||
/// whatever the library beneath it does, so every entry point that calls into
|
||||
/// rawler runs through here, and a file that panics the decoder is one failed
|
||||
/// file like any other.
|
||||
///
|
||||
/// The crash hook still records the panic, because it runs before unwinding
|
||||
/// reaches this frame; that is right — it is a real defect in a dependency
|
||||
/// and the record is how it gets reported upstream — and a repeat is the same
|
||||
/// file being met again rather than a new fault.
|
||||
pub(crate) fn guarded<T>(
|
||||
what: &'static str,
|
||||
f: impl FnOnce() -> Result<T, DecodeError>,
|
||||
) -> Result<T, DecodeError> {
|
||||
match std::panic::catch_unwind(std::panic::AssertUnwindSafe(f)) {
|
||||
Ok(result) => result,
|
||||
Err(payload) => {
|
||||
let msg = payload
|
||||
.downcast_ref::<&str>()
|
||||
.map(|s| s.to_string())
|
||||
.or_else(|| payload.downcast_ref::<String>().cloned())
|
||||
.unwrap_or_else(|| "no message".to_string());
|
||||
Err(DecodeError::Decode(format!(
|
||||
"{what}: the decoder panicked on this file: {msg}"
|
||||
)))
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
impl DecodeError {
|
||||
/// Whether a fallback path might still produce an image.
|
||||
///
|
||||
@@ -49,4 +85,28 @@ mod tests {
|
||||
// A genuinely unsupported file has nowhere to fall through to.
|
||||
assert!(!DecodeError::Unsupported("unknown".into()).has_fallback());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_panic_in_the_decoder_is_an_error_and_the_thread_survives() {
|
||||
// The property the face sweep relies on: one file that panics rawler
|
||||
// is one failed file, not the end of the pass. The message travels,
|
||||
// because "decode failed" alone sends the reader to the crash log.
|
||||
let err = guarded("decode", || -> Result<(), DecodeError> {
|
||||
panic!("rawler: surely there's no such thing as a {}MP image!", 600)
|
||||
})
|
||||
.unwrap_err();
|
||||
let text = err.to_string();
|
||||
assert!(text.contains("panicked"), "{text}");
|
||||
assert!(text.contains("600MP"), "{text}");
|
||||
assert!(!err.has_fallback(), "a panic is not a missing preview");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_result_passes_through_untouched() {
|
||||
assert_eq!(guarded("decode", || Ok::<_, DecodeError>(7)).unwrap(), 7);
|
||||
assert!(matches!(
|
||||
guarded("decode", || Err::<(), _>(DecodeError::NoPreview)),
|
||||
Err(DecodeError::NoPreview)
|
||||
));
|
||||
}
|
||||
}
|
||||
|
||||
@@ -134,6 +134,22 @@ pub struct RawImage {
|
||||
pub base_curve: BaseCurve,
|
||||
/// The usable region of `data`, excluding masked and border photosites.
|
||||
pub crop: CropRect,
|
||||
/// TRACES: FR-MRG-3
|
||||
/// Samples per photosite in `data`: 1 for a colour-filter-array capture,
|
||||
/// 3 for a *linear* DNG — demosaiced RGB, still camera-space, which is
|
||||
/// what a merge writes. With 3, `cfa_pattern` means nothing, `data` is
|
||||
/// `width × height × 3` interleaved, and the GPU uploads it as it is
|
||||
/// rather than demosaicing.
|
||||
pub samples_per_pixel: u8,
|
||||
/// TRACES: FR-MRG-3
|
||||
/// The body's colour profile as the file carried it, for a composite to
|
||||
/// carry on: calibrations and the as-shot neutral. `None` for a body the
|
||||
/// decoder has no matrix for.
|
||||
pub profile: Option<profile::CameraProfile>,
|
||||
/// The body, as rawler cleans the names: what `Make`/`Model` say and what
|
||||
/// the base-curve database matches on.
|
||||
pub make: String,
|
||||
pub model: String,
|
||||
}
|
||||
|
||||
/// TRACES: FR-RAW-3
|
||||
@@ -285,6 +301,10 @@ pub fn probe(header: &[u8]) -> Option<Format> {
|
||||
/// TRACES: FR-CAT-5 | M-12
|
||||
/// Read capture metadata without decoding sensor data.
|
||||
pub fn metadata(bytes: &[u8]) -> Result<Metadata, DecodeError> {
|
||||
error::guarded("metadata", || metadata_unguarded(bytes))
|
||||
}
|
||||
|
||||
fn metadata_unguarded(bytes: &[u8]) -> Result<Metadata, DecodeError> {
|
||||
use rawler::rawsource::RawSource;
|
||||
|
||||
// rawler has no decoder for a plain JPEG, so without this every JPEG in a
|
||||
@@ -510,6 +530,10 @@ pub(crate) fn parse_exif_offset(s: &str) -> Option<i32> {
|
||||
/// Only develop and export should call it; culling and the grid must not
|
||||
/// (FR-CULL-1).
|
||||
pub fn decode(bytes: &[u8]) -> Result<RawImage, DecodeError> {
|
||||
error::guarded("decode", || decode_unguarded(bytes))
|
||||
}
|
||||
|
||||
fn decode_unguarded(bytes: &[u8]) -> Result<RawImage, DecodeError> {
|
||||
use rawler::rawsource::RawSource;
|
||||
|
||||
let source = RawSource::new_from_slice(bytes);
|
||||
@@ -551,6 +575,21 @@ pub fn decode(bytes: &[u8]) -> Result<RawImage, DecodeError> {
|
||||
image.camera.clean_model.as_str(),
|
||||
);
|
||||
|
||||
// TRACES: FR-MRG-3
|
||||
// A linear DNG — three samples per pixel, no colour filter array — is a
|
||||
// composite this application wrote (or any other demosaiced DNG). It
|
||||
// carries the same scale, matrices and neutral as a CFA file and goes
|
||||
// through the same profile; only the demosaic is skipped.
|
||||
let samples_per_pixel = match image.cpp {
|
||||
1 => 1u8,
|
||||
3 => 3,
|
||||
other => {
|
||||
return Err(DecodeError::Unsupported(format!(
|
||||
"{other} samples per pixel; only CFA (1) and linear RGB (3) are handled"
|
||||
)))
|
||||
}
|
||||
};
|
||||
|
||||
let data = match image.data {
|
||||
rawler::RawImageData::Integer(v) => v,
|
||||
rawler::RawImageData::Float(v) => {
|
||||
@@ -619,6 +658,10 @@ pub fn decode(bytes: &[u8]) -> Result<RawImage, DecodeError> {
|
||||
wb_coeffs,
|
||||
color_matrix,
|
||||
base_curve,
|
||||
samples_per_pixel,
|
||||
profile,
|
||||
make: image.camera.clean_make.clone(),
|
||||
model: image.camera.clean_model.clone(),
|
||||
})
|
||||
}
|
||||
|
||||
|
||||
@@ -158,6 +158,10 @@ pub enum PreviewSize {
|
||||
/// Returns [`DecodeError::NoPreview`] where there is none at all: a
|
||||
/// fall-through signal, not a failure (see [`DecodeError::has_fallback`]).
|
||||
pub fn extract_preview(bytes: &[u8], size: PreviewSize) -> Result<Preview, DecodeError> {
|
||||
crate::error::guarded("preview", || extract_preview_unguarded(bytes, size))
|
||||
}
|
||||
|
||||
fn extract_preview_unguarded(bytes: &[u8], size: PreviewSize) -> Result<Preview, DecodeError> {
|
||||
use rawler::rawsource::RawSource;
|
||||
|
||||
// A plain JPEG *is* its own preview — rawler has no decoder for one, and
|
||||
|
||||
@@ -392,6 +392,21 @@ impl CameraProfile {
|
||||
}
|
||||
|
||||
/// The calibrations this profile was built from, coolest first.
|
||||
/// TRACES: FR-MRG-3
|
||||
/// The calibrations as a DNG carries them: `(CalibrationIlluminant,
|
||||
/// ColorMatrix)` with the EXIF light-source code, for a composite to
|
||||
/// write the profile of the body that took its sources.
|
||||
///
|
||||
/// The code is recovered from the temperature, which is lossy only for
|
||||
/// illuminants this profile never kept: `extract` drops calibrations
|
||||
/// whose illuminant has no temperature, so every one here maps back.
|
||||
pub fn dng_calibrations(&self) -> Vec<(u16, [[f32; 3]; 3])> {
|
||||
self.calibrations
|
||||
.iter()
|
||||
.map(|c| (illuminant_code(c.temperature), c.xyz_to_cam))
|
||||
.collect()
|
||||
}
|
||||
|
||||
pub fn calibrations(&self) -> &[Calibration] {
|
||||
&self.calibrations
|
||||
}
|
||||
@@ -528,6 +543,39 @@ fn illuminant_temperature(illuminant: Illuminant) -> Option<f32> {
|
||||
})
|
||||
}
|
||||
|
||||
/// The EXIF `LightSource` code for a calibration temperature — the inverse
|
||||
/// of [`illuminant_temperature`], on the temperatures it produces.
|
||||
fn illuminant_code(temperature: f32) -> u16 {
|
||||
// Nearest of the table, so a temperature that came through a float
|
||||
// round-trip still lands on its illuminant. Where two illuminants share
|
||||
// a temperature (D55 and Daylight, D65 and Cloudy, D75 and Shade) the
|
||||
// CIE standard one is written: it is what every profile database means.
|
||||
const TABLE: &[(f32, u16)] = &[
|
||||
(2856.0, 17), // A
|
||||
(3200.0, 24), // ISO studio tungsten
|
||||
(3500.0, 15), // white fluorescent
|
||||
(4150.0, 14), // cool white fluorescent
|
||||
(4230.0, 2), // fluorescent
|
||||
(4874.0, 18), // B
|
||||
(5000.0, 13), // daylight white fluorescent
|
||||
(5003.0, 23), // D50
|
||||
(5503.0, 20), // D55
|
||||
(6430.0, 12), // daylight fluorescent
|
||||
(6504.0, 21), // D65
|
||||
(6774.0, 19), // C
|
||||
(7504.0, 22), // D75
|
||||
];
|
||||
TABLE
|
||||
.iter()
|
||||
.min_by(|a, b| {
|
||||
(a.0 - temperature)
|
||||
.abs()
|
||||
.total_cmp(&(b.0 - temperature).abs())
|
||||
})
|
||||
.map(|(_, code)| *code)
|
||||
.unwrap_or(255)
|
||||
}
|
||||
|
||||
/// Compose a forward matrix into camera RGB → linear sRGB.
|
||||
///
|
||||
/// `forward` takes white-balanced camera RGB to XYZ under D50, which is the
|
||||
|
||||
@@ -38,4 +38,7 @@ dr-gpu.workspace = true
|
||||
dr-pipeline.workspace = true
|
||||
env_logger.workspace = true
|
||||
pollster.workspace = true
|
||||
# The DNG writer's test reads its output back through the decoder the
|
||||
# library uses, which is the whole claim the writer makes (S15.1).
|
||||
rawler.workspace = true
|
||||
zune-jpeg.workspace = true
|
||||
|
||||
@@ -0,0 +1,323 @@
|
||||
//! TRACES: FR-MRG-3
|
||||
//! A linear DNG: the container a merge writes its composite into.
|
||||
//!
|
||||
//! Decided by S15.1 (2026-09-19): rawler reads back a `LinearRaw` DNG the
|
||||
//! application writes, so a composite re-enters the library as
|
||||
//! `Format::Dng` through the decoder every camera DNG uses. What is written
|
||||
//! is a RAW in every sense a warp can preserve — camera-linear `u16`
|
||||
//! samples at the first source's own scale, its matrices, illuminants,
|
||||
//! as-shot neutral and body name — so the panorama is developed afterwards
|
||||
//! as one photograph, from the sensor's numbers.
|
||||
//!
|
||||
//! # Streamed, not buffered
|
||||
//!
|
||||
//! The composite is larger than any single photograph the pipeline renders
|
||||
//! and larger than the tablet's memory (FR-MRG-11), so the writer never
|
||||
//! holds it. Strips are pulled from the caller one at a time through a
|
||||
//! closure, in order, and written as they arrive; the caller renders a band
|
||||
//! of chunks, hands over its rows, and moves on.
|
||||
//!
|
||||
//! # Why the `tiff` crate after all
|
||||
//!
|
||||
//! S15.1's spike hand-rolled its IFD because the crate's encoder fixes
|
||||
//! `PhotometricInterpretation` to RGB when the image is opened. It does — but
|
||||
//! a directory is a map and a later `write_tag` on the same tag replaces the
|
||||
//! earlier, so `LinearRaw` goes in over the top and everything else the
|
||||
//! crate does (strips, offsets, sub-IFDs, the EXIF block `encode.rs` already
|
||||
//! knows how to write) is kept.
|
||||
|
||||
use std::io::{Seek, Write};
|
||||
|
||||
use tiff::encoder::{colortype, DirectoryEncoder, SRational, TiffEncoder, TiffKind, TiffValue};
|
||||
use tiff::tags::Tag;
|
||||
|
||||
use crate::encode::{sub_directories, tag_metadata, Ascii, Rationals};
|
||||
use crate::{ExportError, SourceMetadata};
|
||||
|
||||
/// What the DNG says about the camera that "took" the composite: the first
|
||||
/// source's profile, carried across so the composite develops through it.
|
||||
#[derive(Debug, Clone, PartialEq)]
|
||||
pub struct DngProfile {
|
||||
/// `UniqueCameraModel`, the name the profile database matches on.
|
||||
pub unique_model: String,
|
||||
/// `(CalibrationIlluminant, ColorMatrix)`: the EXIF light-source code and
|
||||
/// the XYZ → camera matrix measured under it. One or two.
|
||||
pub calibrations: Vec<(u16, [[f32; 3]; 3])>,
|
||||
/// `AsShotNeutral`, camera RGB of the scene's white.
|
||||
pub as_shot_neutral: [f32; 3],
|
||||
/// `WhiteLevel`: the sample value that is clipping. The first source's
|
||||
/// white minus its black, since the samples are black-subtracted.
|
||||
pub white_level: u32,
|
||||
}
|
||||
|
||||
/// Write a linear DNG, pulling `rows_per_strip`-row strips from `strips`.
|
||||
///
|
||||
/// Each call to `strips` receives the strip index and a buffer to fill with
|
||||
/// `width × rows × 3` interleaved RGB `u16` samples (the last strip may be
|
||||
/// shorter). `source` supplies the `Make`, `Model`, dates and EXIF block
|
||||
/// exactly as an export does (FR-EXP-8 sanitising already applied by the
|
||||
/// caller).
|
||||
///
|
||||
/// `PhotometricInterpretation = LinearRaw`, `DNGVersion 1.4`, uncompressed,
|
||||
/// `Orientation = 1` — the composite is written upright (panorama.md §8).
|
||||
///
|
||||
/// `crop` is asked once every strip is in, and its answer — the largest
|
||||
/// rectangle the frames covered, found while the strips went by
|
||||
/// (`Inscribed`) — becomes `DefaultCropOrigin`/`DefaultCropSize`
|
||||
/// (FR-MRG-4): the file opens on the picture, and the border is still in it.
|
||||
// Eight arguments, and each is a different thing: the sink, three
|
||||
// dimensions, the profile, the header, the strip source and the crop. A
|
||||
// struct for them would be a struct with one caller.
|
||||
#[allow(clippy::too_many_arguments)]
|
||||
pub fn write_linear_dng<W, F, C>(
|
||||
out: W,
|
||||
width: u32,
|
||||
height: u32,
|
||||
rows_per_strip: u32,
|
||||
profile: &DngProfile,
|
||||
source: Option<&SourceMetadata>,
|
||||
mut strips: F,
|
||||
crop: C,
|
||||
) -> Result<(), ExportError>
|
||||
where
|
||||
W: Write + Seek,
|
||||
F: FnMut(usize, &mut Vec<u16>) -> Result<(), ExportError>,
|
||||
C: FnOnce() -> Option<crate::Rect>,
|
||||
{
|
||||
let enc = |e: tiff::TiffError| ExportError::Encode(e.to_string());
|
||||
let mut encoder = TiffEncoder::new(out).map_err(enc)?;
|
||||
let sub = sub_directories(&mut encoder, source, width, height)?;
|
||||
let mut image = encoder
|
||||
.new_image::<colortype::RGB16>(width, height)
|
||||
.map_err(enc)?;
|
||||
image.rows_per_strip(rows_per_strip.max(1)).map_err(enc)?;
|
||||
tag_metadata(image.encoder(), source, &sub)?;
|
||||
tag_dng(image.encoder(), profile).map_err(enc)?;
|
||||
|
||||
let rows = rows_per_strip.max(1);
|
||||
let strip_count = height.div_ceil(rows) as usize;
|
||||
let mut buf: Vec<u16> = Vec::with_capacity((width * rows * 3) as usize);
|
||||
for k in 0..strip_count {
|
||||
buf.clear();
|
||||
strips(k, &mut buf)?;
|
||||
let expected_rows = rows.min(height - k as u32 * rows);
|
||||
let expected = (width * expected_rows * 3) as usize;
|
||||
if buf.len() != expected {
|
||||
return Err(ExportError::Encode(format!(
|
||||
"strip {k} has {} samples, expected {expected}",
|
||||
buf.len()
|
||||
)));
|
||||
}
|
||||
image.write_strip(&buf).map_err(enc)?;
|
||||
}
|
||||
if let Some(r) = crop().filter(|r| r.width > 0 && r.height > 0) {
|
||||
let r = crate::Rect {
|
||||
x: r.x.min(width - 1),
|
||||
y: r.y.min(height - 1),
|
||||
width: r.width.min(width - r.x.min(width - 1)),
|
||||
height: r.height.min(height - r.y.min(height - 1)),
|
||||
};
|
||||
image
|
||||
.encoder()
|
||||
.write_tag(Tag::Unknown(tag::DEFAULT_CROP_ORIGIN), &[r.x, r.y][..])
|
||||
.map_err(enc)?;
|
||||
image
|
||||
.encoder()
|
||||
.write_tag(
|
||||
Tag::Unknown(tag::DEFAULT_CROP_SIZE),
|
||||
&[r.width, r.height][..],
|
||||
)
|
||||
.map_err(enc)?;
|
||||
}
|
||||
image.finish().map_err(enc)
|
||||
}
|
||||
|
||||
/// The tags that make a TIFF a DNG, and a linear one.
|
||||
fn tag_dng<W, K>(dir: &mut DirectoryEncoder<'_, W, K>, profile: &DngProfile) -> tiff::TiffResult<()>
|
||||
where
|
||||
W: Write + Seek,
|
||||
K: TiffKind,
|
||||
{
|
||||
// Over the top of what `new_image` wrote: this is the whole trick.
|
||||
dir.write_tag(Tag::PhotometricInterpretation, LINEAR_RAW)?;
|
||||
dir.write_tag(Tag::Orientation, 1u16)?;
|
||||
dir.write_tag(Tag::Unknown(tag::DNG_VERSION), &[1u8, 4, 0, 0][..])?;
|
||||
dir.write_tag(Tag::Unknown(tag::DNG_BACKWARD_VERSION), &[1u8, 4, 0, 0][..])?;
|
||||
dir.write_tag(
|
||||
Tag::Unknown(tag::UNIQUE_CAMERA_MODEL),
|
||||
Ascii(&profile.unique_model),
|
||||
)?;
|
||||
dir.write_tag(
|
||||
Tag::Unknown(tag::WHITE_LEVEL),
|
||||
&[profile.white_level; 3][..],
|
||||
)?;
|
||||
dir.write_tag(Tag::Unknown(tag::BLACK_LEVEL), &[0u32; 3][..])?;
|
||||
|
||||
for (slot, (illuminant, matrix)) in profile.calibrations.iter().take(2).enumerate() {
|
||||
let (ill_tag, mat_tag) = if slot == 0 {
|
||||
(tag::CALIBRATION_ILLUMINANT_1, tag::COLOR_MATRIX_1)
|
||||
} else {
|
||||
(tag::CALIBRATION_ILLUMINANT_2, tag::COLOR_MATRIX_2)
|
||||
};
|
||||
dir.write_tag(Tag::Unknown(ill_tag), *illuminant)?;
|
||||
let flat: Vec<SRational> = matrix
|
||||
.iter()
|
||||
.flatten()
|
||||
.map(|&v| SRational {
|
||||
n: (v * 10_000.0).round() as i32,
|
||||
d: 10_000,
|
||||
})
|
||||
.collect();
|
||||
dir.write_tag(Tag::Unknown(mat_tag), SRationals(&flat))?;
|
||||
}
|
||||
|
||||
let neutral: Vec<(u32, u32)> = profile
|
||||
.as_shot_neutral
|
||||
.iter()
|
||||
.map(|&v| ((v.max(0.0) * 1_000_000.0).round() as u32, 1_000_000))
|
||||
.collect();
|
||||
dir.write_tag(Tag::Unknown(tag::AS_SHOT_NEUTRAL), Rationals(&neutral))?;
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// `PhotometricInterpretation` for demosaiced, un-rendered sensor data.
|
||||
const LINEAR_RAW: u16 = 34892;
|
||||
|
||||
/// DNG tag numbers the `tiff` crate has no names for.
|
||||
mod tag {
|
||||
pub const DNG_VERSION: u16 = 50706;
|
||||
pub const DNG_BACKWARD_VERSION: u16 = 50707;
|
||||
pub const UNIQUE_CAMERA_MODEL: u16 = 50708;
|
||||
pub const BLACK_LEVEL: u16 = 50714;
|
||||
pub const WHITE_LEVEL: u16 = 50717;
|
||||
pub const DEFAULT_CROP_ORIGIN: u16 = 50719;
|
||||
pub const DEFAULT_CROP_SIZE: u16 = 50720;
|
||||
pub const COLOR_MATRIX_1: u16 = 50721;
|
||||
pub const COLOR_MATRIX_2: u16 = 50722;
|
||||
pub const AS_SHOT_NEUTRAL: u16 = 50728;
|
||||
pub const CALIBRATION_ILLUMINANT_1: u16 = 50778;
|
||||
pub const CALIBRATION_ILLUMINANT_2: u16 = 50779;
|
||||
}
|
||||
|
||||
/// A run of `SRATIONAL`s, as `encode::Rationals` is for `RATIONAL`.
|
||||
struct SRationals<'a>(&'a [SRational]);
|
||||
|
||||
impl TiffValue for SRationals<'_> {
|
||||
const BYTE_LEN: u8 = 8;
|
||||
const FIELD_TYPE: tiff::tags::Type = tiff::tags::Type::SRATIONAL;
|
||||
|
||||
fn count(&self) -> usize {
|
||||
self.0.len()
|
||||
}
|
||||
|
||||
fn data(&self) -> std::borrow::Cow<'_, [u8]> {
|
||||
let mut out = Vec::with_capacity(self.0.len() * 8);
|
||||
for r in self.0 {
|
||||
out.extend_from_slice(&r.n.to_ne_bytes());
|
||||
out.extend_from_slice(&r.d.to_ne_bytes());
|
||||
}
|
||||
std::borrow::Cow::Owned(out)
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
fn profile() -> DngProfile {
|
||||
DngProfile {
|
||||
unique_model: "Canon EOS 6D".into(),
|
||||
calibrations: vec![
|
||||
(17, [[0.8, -0.2, 0.1], [-0.3, 1.1, 0.2], [0.0, -0.1, 0.9]]),
|
||||
(21, [[0.7, -0.1, 0.0], [-0.2, 1.0, 0.1], [0.0, -0.2, 0.8]]),
|
||||
],
|
||||
as_shot_neutral: [0.5, 1.0, 0.6],
|
||||
white_level: 13_023,
|
||||
}
|
||||
}
|
||||
|
||||
fn write(width: u32, height: u32, rows: u32) -> Vec<u8> {
|
||||
let mut bytes = std::io::Cursor::new(Vec::new());
|
||||
let source = SourceMetadata {
|
||||
make: Some("Canon".into()),
|
||||
model: Some("Canon EOS 6D".into()),
|
||||
..Default::default()
|
||||
};
|
||||
write_linear_dng(
|
||||
&mut bytes,
|
||||
width,
|
||||
height,
|
||||
rows,
|
||||
&profile(),
|
||||
Some(&source),
|
||||
|k, buf| {
|
||||
let first = k as u32 * rows;
|
||||
let n = rows.min(height - first);
|
||||
for y in first..first + n {
|
||||
for x in 0..width {
|
||||
buf.extend([(x + y * width) as u16, 1000, 2000]);
|
||||
}
|
||||
}
|
||||
Ok(())
|
||||
},
|
||||
|| {
|
||||
Some(crate::Rect {
|
||||
x: 2,
|
||||
y: 1,
|
||||
width: 15,
|
||||
height: 10,
|
||||
})
|
||||
},
|
||||
)
|
||||
.expect("written");
|
||||
bytes.into_inner()
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn rawler_reads_it_back_as_linear_raw() {
|
||||
let bytes = write(20, 13, 4);
|
||||
let source = rawler::rawsource::RawSource::new_from_slice(&bytes);
|
||||
let decoder = rawler::get_decoder(&source).expect("a DNG");
|
||||
let image = decoder
|
||||
.raw_image(&source, &Default::default(), false)
|
||||
.expect("decodes");
|
||||
assert_eq!((image.width, image.height, image.cpp), (20, 13, 3));
|
||||
assert_eq!(image.whitelevel.0[0], 13_023);
|
||||
// Pixel (3, 2) is (3 + 2·20, 1000, 2000) — samples in order, strips
|
||||
// joined without a seam.
|
||||
let rawler::RawImageData::Integer(data) = &image.data else {
|
||||
panic!("integer samples")
|
||||
};
|
||||
let i = (2 * 20 + 3) * 3;
|
||||
assert_eq!(&data[i..i + 3], &[43, 1000, 2000]);
|
||||
// Last row, from the short final strip.
|
||||
let i = (12 * 20 + 19) * 3;
|
||||
assert_eq!(data[i], (19 + 12 * 20) as u16);
|
||||
// The profile came through as the camera's.
|
||||
assert!(!image.camera.color_matrix.is_empty());
|
||||
assert_eq!(image.model, "Canon EOS 6D");
|
||||
// The default crop is what the decoder reports as the picture.
|
||||
let crop = image.crop_area.expect("a crop");
|
||||
assert_eq!((crop.p.x, crop.p.y, crop.d.w, crop.d.h), (2, 1, 15, 10));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_strip_of_the_wrong_length_is_refused() {
|
||||
let mut bytes = std::io::Cursor::new(Vec::new());
|
||||
let err = write_linear_dng(
|
||||
&mut bytes,
|
||||
8,
|
||||
8,
|
||||
8,
|
||||
&profile(),
|
||||
None,
|
||||
|_, buf| {
|
||||
buf.extend([0u16; 10]);
|
||||
Ok(())
|
||||
},
|
||||
|| None,
|
||||
)
|
||||
.unwrap_err();
|
||||
assert!(matches!(err, ExportError::Encode(_)));
|
||||
}
|
||||
}
|
||||
@@ -213,7 +213,7 @@ impl tiff::encoder::TiffValue for Undefined<'_> {
|
||||
/// specification says, `dr-decode` reads them back with `from_utf8_lossy`, and
|
||||
/// a mangled accent is a far better outcome than a refusal. So the bytes go
|
||||
/// through verbatim with the terminating NUL the type requires.
|
||||
struct Ascii<'a>(&'a str);
|
||||
pub(crate) struct Ascii<'a>(pub(crate) &'a str);
|
||||
|
||||
impl tiff::encoder::TiffValue for Ascii<'_> {
|
||||
const BYTE_LEN: u8 = 1;
|
||||
@@ -241,7 +241,7 @@ impl tiff::encoder::TiffValue for Ascii<'_> {
|
||||
/// a value that forced little-endian would be read back byte-swapped on a
|
||||
/// big-endian machine. `exif.rs` builds its own header and so chooses its own
|
||||
/// order; here the container has already chosen.
|
||||
struct Rationals<'a>(&'a [(u32, u32)]);
|
||||
pub(crate) struct Rationals<'a>(pub(crate) &'a [(u32, u32)]);
|
||||
|
||||
impl tiff::encoder::TiffValue for Rationals<'_> {
|
||||
const BYTE_LEN: u8 = 8;
|
||||
@@ -317,7 +317,7 @@ where
|
||||
/// and then no pointer is written either, so the file has no trace of the
|
||||
/// directory rather than a pointer to an empty one.
|
||||
#[derive(Default)]
|
||||
struct SubDirectories {
|
||||
pub(crate) struct SubDirectories {
|
||||
exif: Option<u32>,
|
||||
gps: Option<u32>,
|
||||
}
|
||||
@@ -335,7 +335,7 @@ struct SubDirectories {
|
||||
/// A TIFF gets no separate EXIF *block* — no APP1, no `eXIf` chunk. Its own
|
||||
/// directory is the EXIF structure, and adding a second copy inside it would
|
||||
/// give a reader two answers to every question.
|
||||
fn sub_directories<W>(
|
||||
pub(crate) fn sub_directories<W>(
|
||||
encoder: &mut tiff::encoder::TiffEncoder<W>,
|
||||
source: Option<&SourceMetadata>,
|
||||
width: u32,
|
||||
@@ -465,7 +465,7 @@ where
|
||||
///
|
||||
/// No `Orientation`, for the reason `exif.rs` gives at length: the pixels
|
||||
/// arriving here are already upright.
|
||||
fn tag_metadata<W, K>(
|
||||
pub(crate) fn tag_metadata<W, K>(
|
||||
dir: &mut tiff::encoder::DirectoryEncoder<'_, W, K>,
|
||||
source: Option<&SourceMetadata>,
|
||||
sub: &SubDirectories,
|
||||
|
||||
@@ -0,0 +1,156 @@
|
||||
//! TRACES: FR-MRG-4
|
||||
//! The largest rectangle inside a coverage mask, found a row at a time.
|
||||
//!
|
||||
//! A merged panorama has ragged edges: the frames' footprints under a
|
||||
//! cylinder or a sphere are not rectangles, and the composite carries a
|
||||
//! black border where none of them reached. FR-MRG-4 asks for an auto-crop
|
||||
//! to the largest inscribed rectangle. This finds it as the bands are
|
||||
//! produced, so the composite is never held to be measured (FR-MRG-11):
|
||||
//! each row extends a running histogram of consecutive covered rows above
|
||||
//! it, and the largest rectangle ending on that row is the largest
|
||||
//! rectangle under the histogram — a stack pass, linear in the width.
|
||||
//!
|
||||
//! The crop is written as the DNG's `DefaultCropOrigin`/`DefaultCropSize`,
|
||||
//! which every reader honours and which discards nothing: the pixels
|
||||
//! outside it are still in the file for a photographer who wants them.
|
||||
|
||||
/// The rectangle so far, in pixels from the top left.
|
||||
#[derive(Debug, Clone, Copy, PartialEq, Eq, Default)]
|
||||
pub struct Rect {
|
||||
pub x: u32,
|
||||
pub y: u32,
|
||||
pub width: u32,
|
||||
pub height: u32,
|
||||
}
|
||||
|
||||
impl Rect {
|
||||
pub fn area(&self) -> u64 {
|
||||
u64::from(self.width) * u64::from(self.height)
|
||||
}
|
||||
}
|
||||
|
||||
/// Feed rows top to bottom; ask for the best at any point.
|
||||
#[derive(Debug, Clone)]
|
||||
pub struct Inscribed {
|
||||
width: usize,
|
||||
/// How many consecutive covered rows end at the last row fed, per column.
|
||||
heights: Vec<u32>,
|
||||
rows: u32,
|
||||
best: Rect,
|
||||
}
|
||||
|
||||
impl Inscribed {
|
||||
pub fn new(width: u32) -> Self {
|
||||
Inscribed {
|
||||
width: width as usize,
|
||||
heights: vec![0; width as usize],
|
||||
rows: 0,
|
||||
best: Rect::default(),
|
||||
}
|
||||
}
|
||||
|
||||
/// One more row of coverage, `width` long.
|
||||
pub fn push_row(&mut self, covered: &[bool]) {
|
||||
debug_assert_eq!(covered.len(), self.width);
|
||||
for (h, &c) in self.heights.iter_mut().zip(covered) {
|
||||
*h = if c { *h + 1 } else { 0 };
|
||||
}
|
||||
self.rows += 1;
|
||||
// Largest rectangle under the histogram, with a sentinel column of
|
||||
// height 0 at the end so every bar is popped.
|
||||
let mut stack: Vec<usize> = Vec::new();
|
||||
for i in 0..=self.width {
|
||||
let h = if i < self.width { self.heights[i] } else { 0 };
|
||||
while let Some(&top) = stack.last() {
|
||||
if self.heights[top] <= h {
|
||||
break;
|
||||
}
|
||||
stack.pop();
|
||||
let height = self.heights[top];
|
||||
let left = stack.last().map_or(0, |&l| l + 1);
|
||||
let width = (i - left) as u32;
|
||||
let area = u64::from(width) * u64::from(height);
|
||||
if area > self.best.area() {
|
||||
self.best = Rect {
|
||||
x: left as u32,
|
||||
y: self.rows - height,
|
||||
width,
|
||||
height,
|
||||
};
|
||||
}
|
||||
}
|
||||
stack.push(i);
|
||||
}
|
||||
}
|
||||
|
||||
/// Several rows at once, as a band hands them over.
|
||||
pub fn push_rows(&mut self, covered: &[bool], rows: u32) {
|
||||
for r in 0..rows as usize {
|
||||
self.push_row(&covered[r * self.width..(r + 1) * self.width]);
|
||||
}
|
||||
}
|
||||
|
||||
pub fn best(&self) -> Rect {
|
||||
self.best
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
fn from_art(art: &[&str]) -> Rect {
|
||||
let mut ins = Inscribed::new(art[0].len() as u32);
|
||||
for row in art {
|
||||
let covered: Vec<bool> = row.chars().map(|c| c == '#').collect();
|
||||
ins.push_row(&covered);
|
||||
}
|
||||
ins.best()
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_full_mask_is_its_own_rectangle() {
|
||||
let r = from_art(&["####", "####", "####"]);
|
||||
assert_eq!(
|
||||
r,
|
||||
Rect {
|
||||
x: 0,
|
||||
y: 0,
|
||||
width: 4,
|
||||
height: 3
|
||||
}
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn ragged_edges_are_cut_off() {
|
||||
// A cylinder's footprint: narrower at top and bottom.
|
||||
let r = from_art(&[
|
||||
"..####..", ".######.", "########", "########", ".######.", "..####..",
|
||||
]);
|
||||
// 6 wide × 4 tall = 24 beats 8 × 2 = 16 and 4 × 6 = 24 ties; the
|
||||
// first found wins a tie, which is the wider one here.
|
||||
assert_eq!(r.area(), 24);
|
||||
assert!(r.width == 6 && r.height == 4 || r.width == 4 && r.height == 6);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_hole_is_avoided() {
|
||||
let r = from_art(&["#####", "##.##", "#####", "#####"]);
|
||||
// Left of the hole: 2 × 4 = 8; right: 2 × 4 = 8; below: 5 × 2 = 10.
|
||||
assert_eq!(
|
||||
r,
|
||||
Rect {
|
||||
x: 0,
|
||||
y: 2,
|
||||
width: 5,
|
||||
height: 2
|
||||
}
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn nothing_covered_is_nothing() {
|
||||
assert_eq!(from_art(&["....", "...."]).area(), 0);
|
||||
}
|
||||
}
|
||||
@@ -24,16 +24,20 @@
|
||||
|
||||
use dr_types::{ColourSpace, ExportFormat, ExportSettings};
|
||||
|
||||
mod dng;
|
||||
mod encode;
|
||||
mod error;
|
||||
mod exif;
|
||||
pub mod icc;
|
||||
mod inscribed;
|
||||
mod metadata;
|
||||
mod name;
|
||||
mod sharpen;
|
||||
mod size;
|
||||
|
||||
pub use dng::{write_linear_dng, DngProfile};
|
||||
pub use error::ExportError;
|
||||
pub use inscribed::{Inscribed, Rect};
|
||||
pub use metadata::SourceMetadata;
|
||||
pub use name::{resolve_name, NameContext};
|
||||
pub use size::target_size;
|
||||
|
||||
+10
-5
@@ -9,10 +9,11 @@ license.workspace = true
|
||||
thiserror.workspace = true
|
||||
log.workspace = true
|
||||
|
||||
# Inference. `ort` is the API; **tract is the engine** — see the workspace
|
||||
# manifest, and docs/faces.md §3, for why the C++ ONNX Runtime is not linked.
|
||||
# Inference. `ort` is the API; **what runs it is `dr-inference-engine`'s
|
||||
# business** — tract, or an ONNX Runtime the app found on disk, on whichever
|
||||
# provider the device has (docs/inference.md). This crate never names either.
|
||||
ort = { workspace = true, optional = true }
|
||||
ort-tract = { workspace = true, optional = true }
|
||||
dr-inference-engine = { workspace = true, optional = true }
|
||||
ndarray = { workspace = true, optional = true }
|
||||
|
||||
[dev-dependencies]
|
||||
@@ -20,7 +21,7 @@ zune-jpeg.workspace = true
|
||||
env_logger.workspace = true
|
||||
# The M1 probe drives `ort` directly so it can print the raw load error.
|
||||
ort = { workspace = true }
|
||||
ort-tract = { workspace = true }
|
||||
dr-inference-engine = { workspace = true }
|
||||
|
||||
[[example]]
|
||||
name = "probe"
|
||||
@@ -30,6 +31,10 @@ required-features = ["inference"]
|
||||
name = "faces"
|
||||
required-features = ["inference"]
|
||||
|
||||
[[example]]
|
||||
name = "eyes"
|
||||
required-features = ["inference"]
|
||||
|
||||
[features]
|
||||
# Nothing on by default, and in particular **no `embedded-model`**: the weights
|
||||
# are not a build input and never become one (docs/faces.md §2.2). A feature
|
||||
@@ -44,4 +49,4 @@ default = []
|
||||
# must be testable against synthetic embeddings on a machine with no weights on
|
||||
# it — a test suite that needs a research-licensed download is a test suite
|
||||
# that does not run in CI.
|
||||
inference = ["dep:ort", "dep:ort-tract", "dep:ndarray"]
|
||||
inference = ["dep:ort", "dep:dr-inference-engine", "dep:ndarray"]
|
||||
|
||||
@@ -0,0 +1,145 @@
|
||||
//! Detect the faces in a JPEG and read each one's eyes (docs/faces.md §17).
|
||||
//!
|
||||
//! The thing worth looking at is whether the eye boxes land on eyes and
|
||||
//! whether soft ones are refused — so with `--dump DIR` the crops the
|
||||
//! classifiers were shown are written out as PPMs, one per eye and one per
|
||||
//! head framing, named by image and face, and every line carries the
|
||||
//! numbers the readability floors are set from.
|
||||
//!
|
||||
//! cargo run -p dr-face --features inference --example eyes -- \
|
||||
//! DET.onnx 2D106DET.onnx OCEC.onnx SGC.onnx [--dump DIR] photo.jpg [photo.jpg ...]
|
||||
//!
|
||||
//! All four models must have had their dynamic dims pinned first; see
|
||||
//! `tools/fix-face-model-shapes.sh`.
|
||||
|
||||
use std::path::{Path, PathBuf};
|
||||
use std::time::Instant;
|
||||
|
||||
use dr_face::{align, DetectOptions, Detector, EyeModels, Pixels};
|
||||
|
||||
fn main() {
|
||||
env_logger::init();
|
||||
|
||||
let mut args: Vec<String> = std::env::args().skip(1).collect();
|
||||
let dump = args.iter().position(|a| a == "--dump").map(|i| {
|
||||
args.remove(i);
|
||||
PathBuf::from(args.remove(i))
|
||||
});
|
||||
if args.len() < 5 {
|
||||
eprintln!(
|
||||
"usage: eyes DET.onnx 2D106DET.onnx OCEC.onnx SGC.onnx [--dump DIR] IMAGE.jpg [IMAGE.jpg ...]"
|
||||
);
|
||||
std::process::exit(2);
|
||||
}
|
||||
if let Some(d) = &dump {
|
||||
std::fs::create_dir_all(d).expect("dump dir");
|
||||
}
|
||||
|
||||
let t = Instant::now();
|
||||
let mut detector = Detector::from_path(&args[0]).expect("load detector");
|
||||
let mut models = EyeModels::from_paths(&args[1], &args[2], &args[3]).expect("load eye models");
|
||||
println!("loaded the models in {:?}", t.elapsed());
|
||||
|
||||
let opts = DetectOptions::default();
|
||||
for path in &args[4..] {
|
||||
let (rgb, w, h) = match load_jpeg(path) {
|
||||
Ok(v) => v,
|
||||
Err(e) => {
|
||||
println!("{path}: {e}");
|
||||
continue;
|
||||
}
|
||||
};
|
||||
let dets = detector.detect(&rgb, w, h, &opts).expect("detect");
|
||||
println!("\n{path} ({w}×{h}) {} face(s)", dets.len());
|
||||
|
||||
let stem = Path::new(path)
|
||||
.file_stem()
|
||||
.map(|s| s.to_string_lossy().into_owned())
|
||||
.unwrap_or_default();
|
||||
|
||||
for (i, d) in dets.iter().enumerate() {
|
||||
let px = Pixels::RgbF32(&rgb);
|
||||
let t = Instant::now();
|
||||
let reading = models
|
||||
.read(px, w, h, d.bbox, &d.landmarks)
|
||||
.expect("read eyes");
|
||||
let ms = t.elapsed().as_secs_f64() * 1e3;
|
||||
let Some((r, lm)) = reading else {
|
||||
println!(" [{i}] nothing to cut, skipped");
|
||||
continue;
|
||||
};
|
||||
println!(
|
||||
" [{i}] conf {:.2} box {:.0}×{:.0} right {:.3} ({:.0}px, sharp {:.3}) left {:.3} ({:.0}px, sharp {:.3}) sunglasses {:.3} → {:?} ({ms:.1} ms)",
|
||||
d.confidence,
|
||||
d.width(),
|
||||
d.height(),
|
||||
r.right.open,
|
||||
r.right.px,
|
||||
r.right.sharpness,
|
||||
r.left.open,
|
||||
r.left.px,
|
||||
r.left.sharpness,
|
||||
r.sunglasses,
|
||||
r.state(),
|
||||
);
|
||||
if let Some(dir) = &dump {
|
||||
// The same crops `EyeModels::read` cut, cut again for the
|
||||
// sheet from the landmarks it handed back: the reading itself
|
||||
// carries numbers, not pixels.
|
||||
for (name, contour) in [("right", lm.right_eye()), ("left", lm.left_eye())] {
|
||||
if let Some(patch) =
|
||||
align::eye_box(&contour).and_then(|b| align::eye_patch(px, w, h, b))
|
||||
{
|
||||
write_ppm(
|
||||
&dir.join(format!("{stem}-{i}-{name}.ppm")),
|
||||
patch.pixels(),
|
||||
align::EYE_PATCH_WIDTH,
|
||||
align::EYE_PATCH_HEIGHT,
|
||||
);
|
||||
}
|
||||
}
|
||||
if let Some(head) = align::head_views(px, w, h, &d.landmarks) {
|
||||
for (n, view) in head.views().enumerate() {
|
||||
write_ppm(
|
||||
&dir.join(format!("{stem}-{i}-head{n}.ppm")),
|
||||
view,
|
||||
align::SUNGLASSES_EDGE,
|
||||
align::SUNGLASSES_EDGE,
|
||||
);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
fn write_ppm(path: &Path, rgb: &[f32], w: usize, h: usize) {
|
||||
let mut out = format!("P6\n{w} {h}\n255\n").into_bytes();
|
||||
out.extend(
|
||||
rgb.iter()
|
||||
.map(|v| (v.clamp(0.0, 1.0) * 255.0).round() as u8),
|
||||
);
|
||||
std::fs::write(path, out).expect("write ppm");
|
||||
}
|
||||
|
||||
/// Decode to the tightly packed `f32` RGB `0.0..=1.0` the crate expects.
|
||||
fn load_jpeg(path: &str) -> Result<(Vec<f32>, usize, usize), String> {
|
||||
let bytes = std::fs::read(path).map_err(|e| e.to_string())?;
|
||||
let mut dec = zune_jpeg::JpegDecoder::new(&bytes);
|
||||
let px = dec.decode().map_err(|e| e.to_string())?;
|
||||
let info = dec.info().ok_or("no jpeg header")?;
|
||||
let (w, h) = (info.width as usize, info.height as usize);
|
||||
|
||||
let rgb: Vec<f32> = match px.len() / (w * h) {
|
||||
3 => px.iter().map(|&v| v as f32 / 255.0).collect(),
|
||||
1 => px
|
||||
.iter()
|
||||
.flat_map(|&v| {
|
||||
let g = v as f32 / 255.0;
|
||||
[g, g, g]
|
||||
})
|
||||
.collect(),
|
||||
n => return Err(format!("{n} components per pixel, expected 1 or 3")),
|
||||
};
|
||||
Ok((rgb, w, h))
|
||||
}
|
||||
+467
-56
@@ -117,51 +117,56 @@ impl Aligned112 {
|
||||
/// `face_index --quality` prints the joint distribution so the two are
|
||||
/// chosen together rather than each in ignorance of the other.
|
||||
pub fn sharpness(&self) -> f32 {
|
||||
let e = ALIGNED_EDGE;
|
||||
let luma: Vec<f32> = self
|
||||
.pixels
|
||||
.chunks_exact(3)
|
||||
.map(|p| 0.2126 * p[0] + 0.7152 * p[1] + 0.0722 * p[2])
|
||||
.collect();
|
||||
|
||||
let (mut lap_sum, mut lap_sq) = (0.0_f64, 0.0_f64);
|
||||
let (mut lum_sum, mut lum_sq) = (0.0_f64, 0.0_f64);
|
||||
let mut n = 0.0_f64;
|
||||
|
||||
for y in 1..e - 1 {
|
||||
for x in 1..e - 1 {
|
||||
let i = y * e + x;
|
||||
// Four-neighbour Laplacian. The 8-neighbour form is more
|
||||
// sensitive to diagonal detail and also to noise, which on a
|
||||
// high-ISO frame is exactly the thing that must not read as
|
||||
// sharpness.
|
||||
let lap = 4.0 * luma[i] - luma[i - 1] - luma[i + 1] - luma[i - e] - luma[i + e];
|
||||
let lap = lap as f64;
|
||||
lap_sum += lap;
|
||||
lap_sq += lap * lap;
|
||||
|
||||
let l = luma[i] as f64;
|
||||
lum_sum += l;
|
||||
lum_sq += l * l;
|
||||
n += 1.0;
|
||||
}
|
||||
}
|
||||
|
||||
if n == 0.0 {
|
||||
return 0.0;
|
||||
}
|
||||
let lap_var = (lap_sq / n - (lap_sum / n).powi(2)).max(0.0);
|
||||
let lum_var = (lum_sq / n - (lum_sum / n).powi(2)).max(0.0);
|
||||
|
||||
// A crop with no luma variation has no edges to find either, so the
|
||||
// ratio is 0/0. Zero is the right answer: nothing there is a face.
|
||||
if lum_var <= 1e-9 {
|
||||
return 0.0;
|
||||
}
|
||||
(lap_var / lum_var) as f32
|
||||
laplacian_ratio(&self.pixels, ALIGNED_EDGE, ALIGNED_EDGE)
|
||||
}
|
||||
}
|
||||
|
||||
/// Variance of the four-neighbour Laplacian over the variance of the luma,
|
||||
/// for a `w × h` RGB crop — the measure [`Aligned112::sharpness`] describes,
|
||||
/// shared with [`EyePatch::sharpness`].
|
||||
fn laplacian_ratio(pixels: &[f32], w: usize, h: usize) -> f32 {
|
||||
let luma: Vec<f32> = pixels
|
||||
.chunks_exact(3)
|
||||
.map(|p| 0.2126 * p[0] + 0.7152 * p[1] + 0.0722 * p[2])
|
||||
.collect();
|
||||
|
||||
let (mut lap_sum, mut lap_sq) = (0.0_f64, 0.0_f64);
|
||||
let (mut lum_sum, mut lum_sq) = (0.0_f64, 0.0_f64);
|
||||
let mut n = 0.0_f64;
|
||||
|
||||
for y in 1..h.saturating_sub(1) {
|
||||
for x in 1..w.saturating_sub(1) {
|
||||
let i = y * w + x;
|
||||
// Four-neighbour Laplacian. The 8-neighbour form is more
|
||||
// sensitive to diagonal detail and also to noise, which on a
|
||||
// high-ISO frame is exactly the thing that must not read as
|
||||
// sharpness.
|
||||
let lap = 4.0 * luma[i] - luma[i - 1] - luma[i + 1] - luma[i - w] - luma[i + w];
|
||||
let lap = lap as f64;
|
||||
lap_sum += lap;
|
||||
lap_sq += lap * lap;
|
||||
|
||||
let l = luma[i] as f64;
|
||||
lum_sum += l;
|
||||
lum_sq += l * l;
|
||||
n += 1.0;
|
||||
}
|
||||
}
|
||||
|
||||
if n == 0.0 {
|
||||
return 0.0;
|
||||
}
|
||||
let lap_var = (lap_sq / n - (lap_sum / n).powi(2)).max(0.0);
|
||||
let lum_var = (lum_sq / n - (lum_sum / n).powi(2)).max(0.0);
|
||||
|
||||
// A crop with no luma variation has no edges to find either, so the
|
||||
// ratio is 0/0. Zero is the right answer: nothing there is a face.
|
||||
if lum_var <= 1e-9 {
|
||||
return 0.0;
|
||||
}
|
||||
(lap_var / lum_var) as f32
|
||||
}
|
||||
|
||||
/// A similarity transform: rotation, uniform scale, translation.
|
||||
///
|
||||
/// Stored as the four independent parameters rather than a 2×3 matrix so that
|
||||
@@ -351,27 +356,314 @@ pub fn warp_pixels(
|
||||
let m = fit_similarity(landmarks, &ARCFACE_TEMPLATE)?;
|
||||
|
||||
let e = ALIGNED_EDGE;
|
||||
let mut pixels = vec![0.0_f32; e * e * 3];
|
||||
for v in 0..e {
|
||||
for u in 0..e {
|
||||
// Pixel centres, so the transform is not off by half a pixel —
|
||||
// which is small enough to survive review and large enough to
|
||||
// matter on a 40-pixel face.
|
||||
let (x, y) = m.invert(u as f32 + 0.5, v as f32 + 0.5);
|
||||
let (x, y) = (x - 0.5, y - 0.5);
|
||||
let out = (v * e + u) * 3;
|
||||
sample_bilinear(px, width, height, x, y, &mut pixels[out..out + 3]);
|
||||
}
|
||||
}
|
||||
|
||||
let window = TemplateWindow {
|
||||
x: 0.0,
|
||||
y: 0.0,
|
||||
w: e as f32,
|
||||
h: e as f32,
|
||||
};
|
||||
Some(Aligned112 {
|
||||
pixels,
|
||||
pixels: sample_window(px, width, height, &m, &window, e, e),
|
||||
// The warp maps `scale` source pixels to one destination pixel, so the
|
||||
// crop spans 112/scale of the source.
|
||||
source_px: ALIGNED_EDGE as f32 / m.scale(),
|
||||
})
|
||||
}
|
||||
|
||||
/// A rectangle in **template** coordinates — the 112-unit frame
|
||||
/// [`ARCFACE_TEMPLATE`] is written in — that a crop is sampled from.
|
||||
///
|
||||
/// Every crop this module makes is one of these resampled through the same
|
||||
/// fitted similarity: the aligned face is the window `(0, 0, 112, 112)`, an
|
||||
/// eye is a small window around its template point, a head is a window larger
|
||||
/// than the face. Stating them all in one frame is what lets a second crop be
|
||||
/// added as a constant rather than a second warp, and what keeps them
|
||||
/// consistent with each other — the eye window sits where the eye landmark
|
||||
/// lands *after* alignment, so a tilted face gets an upright eye.
|
||||
#[derive(Debug, Clone, Copy, PartialEq)]
|
||||
struct TemplateWindow {
|
||||
x: f32,
|
||||
y: f32,
|
||||
w: f32,
|
||||
h: f32,
|
||||
}
|
||||
|
||||
/// Resample `window` of the template frame into an `out_w × out_h` RGB buffer.
|
||||
///
|
||||
/// Bilinear, from the source, in one step — the property [`warp`] insists on,
|
||||
/// and every crop through here inherits it. The output pixel `(u, v)` is placed
|
||||
/// at its centre in the window, taken back through `m` to source coordinates,
|
||||
/// and sampled there; the window's aspect is **not** preserved when it differs
|
||||
/// from the output's, which is deliberate for the eye classifier (it was
|
||||
/// trained on detector boxes resized the same way) and moot for the others.
|
||||
fn sample_window(
|
||||
px: Pixels<'_>,
|
||||
width: usize,
|
||||
height: usize,
|
||||
m: &Similarity,
|
||||
window: &TemplateWindow,
|
||||
out_w: usize,
|
||||
out_h: usize,
|
||||
) -> Vec<f32> {
|
||||
let mut pixels = vec![0.0_f32; out_w * out_h * 3];
|
||||
let sx = window.w / out_w as f32;
|
||||
let sy = window.h / out_h as f32;
|
||||
for v in 0..out_h {
|
||||
for u in 0..out_w {
|
||||
// Pixel centres, so the transform is not off by half a pixel —
|
||||
// which is small enough to survive review and large enough to
|
||||
// matter on a 40-pixel face.
|
||||
let tx = window.x + (u as f32 + 0.5) * sx;
|
||||
let ty = window.y + (v as f32 + 0.5) * sy;
|
||||
let (x, y) = m.invert(tx, ty);
|
||||
let (x, y) = (x - 0.5, y - 0.5);
|
||||
let out = (v * out_w + u) * 3;
|
||||
sample_bilinear(px, width, height, x, y, &mut pixels[out..out + 3]);
|
||||
}
|
||||
}
|
||||
pixels
|
||||
}
|
||||
|
||||
// ── eyes ──────────────────────────────────────────────────────────────────
|
||||
|
||||
/// Width of an eye crop as the classifier reads it, in pixels. Fixed by the
|
||||
/// OCEC input (`docs/faces.md` §17): 40 wide, 24 high.
|
||||
pub const EYE_PATCH_WIDTH: usize = 40;
|
||||
/// Height of an eye crop as the classifier reads it, in pixels.
|
||||
pub const EYE_PATCH_HEIGHT: usize = 24;
|
||||
|
||||
/// How much an eye's box is grown beyond its lid contour, as a fraction of
|
||||
/// its width and height on each side.
|
||||
///
|
||||
/// The classifier was trained on a whole-body detector's *eye* boxes — tight
|
||||
/// round the palpebral fissure — and measured on 25 open-eyed faces from the
|
||||
/// reference library, a tight box is what it wants: 22 of 25 read open at
|
||||
/// 0 and 0.1, 18 at 0.4, 14 at 0.6 (docs/faces.md §17.2). A tenth, so a
|
||||
/// contour landing a pixel short of the lashes still holds them.
|
||||
pub const EYE_BOX_MARGIN: f32 = 0.1;
|
||||
|
||||
/// Height a shut eye's box is given, as a fraction of its width.
|
||||
///
|
||||
/// A closed eye's contour has no height. The box is given the height an
|
||||
/// open eye of the same width would have, so the classifier sees the same
|
||||
/// framing either way — which is what it was trained on.
|
||||
pub const EYE_BOX_MIN_ASPECT: f32 = 0.4;
|
||||
|
||||
/// The box round an eye's lid contour, in the contour's own coordinates:
|
||||
/// `(x, y, w, h)`.
|
||||
///
|
||||
/// Model-free: the contour is whatever the landmark model gave for the ten
|
||||
/// (or so) points on the lids, in source pixels. `None` for an empty
|
||||
/// contour or one with no width, which is what a hidden eye's collapsed
|
||||
/// contour can come to.
|
||||
pub fn eye_box(contour: &[(f32, f32)]) -> Option<(f32, f32, f32, f32)> {
|
||||
let (mut x0, mut y0, mut x1, mut y1) = (f32::MAX, f32::MAX, f32::MIN, f32::MIN);
|
||||
for &(x, y) in contour {
|
||||
x0 = x0.min(x);
|
||||
y0 = y0.min(y);
|
||||
x1 = x1.max(x);
|
||||
y1 = y1.max(y);
|
||||
}
|
||||
let w = x1 - x0;
|
||||
if contour.is_empty() || w <= 0.0 || w.is_nan() {
|
||||
return None;
|
||||
}
|
||||
let h = (y1 - y0).max(w * EYE_BOX_MIN_ASPECT);
|
||||
let cy = (y0 + y1) / 2.0;
|
||||
let (mx, my) = (w * EYE_BOX_MARGIN, h * EYE_BOX_MARGIN);
|
||||
Some((x0 - mx, cy - h / 2.0 - my, w + 2.0 * mx, h + 2.0 * my))
|
||||
}
|
||||
|
||||
/// One eye, resampled to the classifier's input.
|
||||
///
|
||||
/// Constructible only by [`eye_patch`], for the reason [`Aligned112`] is
|
||||
/// only constructible by [`warp`]: the classifier accepting a plain buffer
|
||||
/// would accept any 40×24 of anything, and its answer would still be a
|
||||
/// plausible probability.
|
||||
#[derive(Debug, Clone, PartialEq)]
|
||||
pub struct EyePatch {
|
||||
/// `24 × 40 × 3`, row-major RGB in `0.0..=1.0`.
|
||||
pixels: Vec<f32>,
|
||||
/// Source pixels across the box the patch was cut from.
|
||||
source_px: f32,
|
||||
}
|
||||
|
||||
impl EyePatch {
|
||||
pub fn pixels(&self) -> &[f32] {
|
||||
&self.pixels
|
||||
}
|
||||
|
||||
/// Source pixels across the eye box — how much eye there was to read.
|
||||
///
|
||||
/// The classifier was trained down to eyes a dozen pixels wide, and
|
||||
/// below that a crop is an interpolation of nothing; `crate::eyes` draws
|
||||
/// the line. Zero when the box had no width, which is a hidden eye.
|
||||
pub fn source_px(&self) -> f32 {
|
||||
self.source_px
|
||||
}
|
||||
|
||||
/// How sharp the eye the classifier is about to see actually is —
|
||||
/// [`Aligned112::sharpness`]'s measure, over the patch.
|
||||
///
|
||||
/// The reason it exists is the reason the face's does: a soft eye is
|
||||
/// not a closed one, but a classifier shown a smear says "closed" with
|
||||
/// the same confidence it says anything, and the only defence is to
|
||||
/// not ask. A face sharp enough to embed can still hold an eye too soft
|
||||
/// to read — it is a fortieth of the face — so the measure is taken
|
||||
/// here and not inherited from the crop.
|
||||
pub fn sharpness(&self) -> f32 {
|
||||
laplacian_ratio(&self.pixels, EYE_PATCH_WIDTH, EYE_PATCH_HEIGHT)
|
||||
}
|
||||
}
|
||||
|
||||
/// Cut an eye out of the source at the classifier's size, from an
|
||||
/// axis-aligned box in source pixels — [`eye_box`]'s, as a rule.
|
||||
///
|
||||
/// Upright and from the frame, not through the face's alignment: the
|
||||
/// classifier's training crops were detector boxes, and a landmark model's
|
||||
/// contour already says where the eye is on a tilted head. Bilinear in one
|
||||
/// step from the native buffer, so a large face gives real pixels; the
|
||||
/// box's aspect is not preserved, which is what the training resize did.
|
||||
pub fn eye_patch(
|
||||
px: Pixels<'_>,
|
||||
width: usize,
|
||||
height: usize,
|
||||
bbox: (f32, f32, f32, f32),
|
||||
) -> Option<EyePatch> {
|
||||
let pixels = crop_box(px, width, height, bbox, EYE_PATCH_WIDTH, EYE_PATCH_HEIGHT)?;
|
||||
Some(EyePatch {
|
||||
pixels,
|
||||
source_px: bbox.2,
|
||||
})
|
||||
}
|
||||
|
||||
// ── sunglasses ────────────────────────────────────────────────────────────
|
||||
|
||||
/// Edge of the crop the sunglasses classifier reads. Fixed by the SGC input:
|
||||
/// 48×48.
|
||||
pub const SUNGLASSES_EDGE: usize = 48;
|
||||
|
||||
/// The windows read for the sunglasses classifier, in template units:
|
||||
/// `(x, y, w, h)`.
|
||||
///
|
||||
/// **Two framings, and the classifier's answer is the higher of the two.**
|
||||
/// It was trained on a whole-body detector's *head* boxes, and a head box
|
||||
/// is not reproducible from five landmarks: how much hair and hat it took in
|
||||
/// depended on the person. So it is shown the face twice — once as the
|
||||
/// aligned crop itself, once shifted up and widened to take in hair and
|
||||
/// hat at the cost of the chin, which is roughly where a head box falls —
|
||||
/// and a pair of sunglasses counts if it looks like one in either.
|
||||
///
|
||||
/// Measured over 12 faces in sunglasses and 28 with plainly visible eyes
|
||||
/// from the reference library (`examples/eyes.rs --head`), at the 0.5
|
||||
/// threshold:
|
||||
///
|
||||
/// | window | sunglasses found | clear eyes kept |
|
||||
/// |---|---|---|
|
||||
/// | the aligned face, `(0, 0, 112, 112)` | 9 | 28 |
|
||||
/// | a head, `(-5, -14, 122, 122)` | 6 | 27 |
|
||||
/// | a larger head, `(-30, -55, 172, 190)` | 6 | 25 |
|
||||
/// | **the higher of the first two** | **11** | 27 |
|
||||
///
|
||||
/// The face-tight crop alone was the best single framing, which was not the
|
||||
/// expectation; the head framing found the sunglasses under a cap that the
|
||||
/// face crop missed. The one clear-eyed face the pair loses wears a cap and
|
||||
/// clear glasses, at 0.68. Erring towards "sunglasses" is the safe direction
|
||||
/// for what this feeds: a face called sunglasses is left alone by the
|
||||
/// eyes-open filter, where a pair of sunglasses missed hands the eye
|
||||
/// classifier a lens to guess at (docs/faces.md §17).
|
||||
pub const SUNGLASSES_WINDOWS: [(f32, f32, f32, f32); 2] =
|
||||
[(0.0, 0.0, 112.0, 112.0), (-5.0, -14.0, 122.0, 122.0)];
|
||||
|
||||
/// The framings of one face the sunglasses classifier is shown.
|
||||
///
|
||||
/// A newtype for the reason [`EyePatch`] is one.
|
||||
#[derive(Debug, Clone, PartialEq)]
|
||||
pub struct HeadViews {
|
||||
/// Each `48 × 48 × 3`, row-major RGB in `0.0..=1.0`.
|
||||
views: Vec<Vec<f32>>,
|
||||
}
|
||||
|
||||
impl HeadViews {
|
||||
pub fn views(&self) -> impl Iterator<Item = &[f32]> {
|
||||
self.views.iter().map(Vec::as_slice)
|
||||
}
|
||||
}
|
||||
|
||||
/// Cut the [`SUNGLASSES_WINDOWS`] out of the source, aligned, at the
|
||||
/// classifier's size.
|
||||
pub fn head_views(
|
||||
px: Pixels<'_>,
|
||||
width: usize,
|
||||
height: usize,
|
||||
landmarks: &[(f32, f32); 5],
|
||||
) -> Option<HeadViews> {
|
||||
head_views_in(px, width, height, landmarks, &SUNGLASSES_WINDOWS)
|
||||
}
|
||||
|
||||
/// [`head_views`] over windows other than [`SUNGLASSES_WINDOWS`].
|
||||
///
|
||||
/// For measuring them, which is how the constant was chosen
|
||||
/// (`examples/eyes.rs --head`); production callers use the constant.
|
||||
pub fn head_views_in(
|
||||
px: Pixels<'_>,
|
||||
width: usize,
|
||||
height: usize,
|
||||
landmarks: &[(f32, f32); 5],
|
||||
windows: &[(f32, f32, f32, f32)],
|
||||
) -> Option<HeadViews> {
|
||||
if !px.fits(width, height) || windows.is_empty() {
|
||||
return None;
|
||||
}
|
||||
let m = fit_similarity(landmarks, &ARCFACE_TEMPLATE)?;
|
||||
let views = windows
|
||||
.iter()
|
||||
.map(|&(x, y, w, h)| {
|
||||
let window = TemplateWindow { x, y, w, h };
|
||||
sample_window(
|
||||
px,
|
||||
width,
|
||||
height,
|
||||
&m,
|
||||
&window,
|
||||
SUNGLASSES_EDGE,
|
||||
SUNGLASSES_EDGE,
|
||||
)
|
||||
})
|
||||
.collect();
|
||||
Some(HeadViews { views })
|
||||
}
|
||||
|
||||
/// An axis-aligned crop of the source, resampled to `out_w × out_h` RGB.
|
||||
///
|
||||
/// `(x, y, w, h)` in source pixels; the aspect is not preserved when it
|
||||
/// differs from the output's. Bilinear in one step, like every crop here;
|
||||
/// pixels outside the source read black. What a landmark model trained on
|
||||
/// detector boxes wants — upright, from the frame — as against the aligned
|
||||
/// windows above.
|
||||
pub fn crop_box(
|
||||
px: Pixels<'_>,
|
||||
width: usize,
|
||||
height: usize,
|
||||
(x, y, w, h): (f32, f32, f32, f32),
|
||||
out_w: usize,
|
||||
out_h: usize,
|
||||
) -> Option<Vec<f32>> {
|
||||
if !px.fits(width, height) || w <= 0.0 || h <= 0.0 {
|
||||
return None;
|
||||
}
|
||||
let identity = Similarity {
|
||||
a: 1.0,
|
||||
b: 0.0,
|
||||
tx: 0.0,
|
||||
ty: 0.0,
|
||||
};
|
||||
let window = TemplateWindow { x, y, w, h };
|
||||
Some(sample_window(
|
||||
px, width, height, &identity, &window, out_w, out_h,
|
||||
))
|
||||
}
|
||||
|
||||
fn sample_bilinear(px: Pixels<'_>, w: usize, h: usize, x: f32, y: f32, out: &mut [f32]) {
|
||||
let x0 = x.floor();
|
||||
let y0 = y.floor();
|
||||
@@ -489,6 +781,125 @@ mod tests {
|
||||
}
|
||||
}
|
||||
|
||||
/// A source whose red channel is its x coordinate and green its y, so a
|
||||
/// crop's mean colour says where in the source it was taken from.
|
||||
fn coordinate_image(w: usize, h: usize) -> Vec<f32> {
|
||||
let mut rgb = vec![0.0_f32; w * h * 3];
|
||||
for y in 0..h {
|
||||
for x in 0..w {
|
||||
rgb[(y * w + x) * 3] = x as f32 / w as f32;
|
||||
rgb[(y * w + x) * 3 + 1] = y as f32 / h as f32;
|
||||
}
|
||||
}
|
||||
rgb
|
||||
}
|
||||
|
||||
fn mean_channel(px: &[f32], c: usize) -> f32 {
|
||||
let n = px.len() / 3;
|
||||
px.chunks_exact(3).map(|p| p[c]).sum::<f32>() / n as f32
|
||||
}
|
||||
|
||||
/// The box is the contour's bounds, grown by the margin, and a shut
|
||||
/// eye's flat contour is given an open eye's height.
|
||||
#[test]
|
||||
fn an_eye_box_holds_its_contour_with_a_margin() {
|
||||
let open = [(100.0, 50.0), (110.0, 46.0), (120.0, 50.0), (110.0, 54.0)];
|
||||
let (x, y, w, h) = eye_box(&open).unwrap();
|
||||
assert!((w - 20.0 * (1.0 + 2.0 * EYE_BOX_MARGIN)).abs() < 1e-4);
|
||||
assert!((h - 8.0 * (1.0 + 2.0 * EYE_BOX_MARGIN)).abs() < 1e-4);
|
||||
assert!((x + w / 2.0 - 110.0).abs() < 1e-4);
|
||||
assert!((y + h / 2.0 - 50.0).abs() < 1e-4);
|
||||
|
||||
let shut = [(100.0, 50.0), (110.0, 50.0), (120.0, 50.0)];
|
||||
let (_, _, w2, h2) = eye_box(&shut).unwrap();
|
||||
assert!((w2 - w).abs() < 1e-4, "same width");
|
||||
assert!((h2 - 20.0 * EYE_BOX_MIN_ASPECT * (1.0 + 2.0 * EYE_BOX_MARGIN)).abs() < 1e-4);
|
||||
|
||||
assert!(eye_box(&[]).is_none());
|
||||
assert!(eye_box(&[(5.0, 5.0), (5.0, 9.0)]).is_none(), "no width");
|
||||
}
|
||||
|
||||
/// The patch is cut from the box it was given, upright, and knows how
|
||||
/// many source pixels it spans.
|
||||
#[test]
|
||||
fn an_eye_patch_is_the_box_resampled() {
|
||||
let (w, h) = (200, 200);
|
||||
let rgb = coordinate_image(w, h);
|
||||
let bbox = (60.0, 90.0, 30.0, 12.0);
|
||||
let eye = eye_patch(Pixels::RgbF32(&rgb), w, h, bbox).unwrap();
|
||||
assert_eq!(eye.pixels().len(), EYE_PATCH_WIDTH * EYE_PATCH_HEIGHT * 3);
|
||||
assert_eq!(eye.source_px(), 30.0);
|
||||
let cx = mean_channel(eye.pixels(), 0) * w as f32;
|
||||
let cy = mean_channel(eye.pixels(), 1) * h as f32;
|
||||
assert!((cx - 75.0).abs() < 0.6, "{cx}");
|
||||
assert!((cy - 96.0).abs() < 0.6, "{cy}");
|
||||
// No width, or a buffer that is not the size it claims: nothing.
|
||||
assert!(eye_patch(Pixels::RgbF32(&rgb), w, h, (60.0, 90.0, 0.0, 12.0)).is_none());
|
||||
assert!(eye_patch(Pixels::RgbF32(&rgb), 190, 200, bbox).is_none());
|
||||
}
|
||||
|
||||
/// A soft eye scores lower than the same eye sharp, on the patch itself.
|
||||
#[test]
|
||||
fn an_eye_patchs_sharpness_falls_with_blur() {
|
||||
let edge = 120;
|
||||
let sharp = image(
|
||||
edge,
|
||||
|x, y| if (x / 5 + y / 5) % 2 == 0 { 0.9 } else { 0.1 },
|
||||
);
|
||||
let soft = blur(&blur(&sharp, edge), edge);
|
||||
let bbox = (20.0, 40.0, 40.0, 24.0);
|
||||
let a = eye_patch(Pixels::RgbF32(&sharp), edge, edge, bbox)
|
||||
.unwrap()
|
||||
.sharpness();
|
||||
let b = eye_patch(Pixels::RgbF32(&soft), edge, edge, bbox)
|
||||
.unwrap()
|
||||
.sharpness();
|
||||
assert!(a > b * 2.0, "sharp {a} should clearly beat blurred {b}");
|
||||
}
|
||||
|
||||
/// The second sunglasses framing takes in more than the face — it starts
|
||||
/// above the template's top edge and ends below its bottom — and the
|
||||
/// first is the aligned face itself.
|
||||
#[test]
|
||||
fn the_head_views_are_the_face_and_a_wider_framing_of_it() {
|
||||
let (w, h) = (300, 300);
|
||||
let rgb = coordinate_image(w, h);
|
||||
let lm = shifted_scaled(1.0, 100.0, 100.0, 0.0);
|
||||
let head = head_views(Pixels::RgbF32(&rgb), w, h, &lm).unwrap();
|
||||
let views: Vec<&[f32]> = head.views().collect();
|
||||
let face = warp(&rgb, w, h, &lm).unwrap();
|
||||
assert_eq!(views.len(), SUNGLASSES_WINDOWS.len());
|
||||
for v in &views {
|
||||
assert_eq!(v.len(), SUNGLASSES_EDGE * SUNGLASSES_EDGE * 3);
|
||||
}
|
||||
|
||||
// The face view samples the same region as the aligned crop.
|
||||
assert!((mean_channel(views[0], 0) - mean_channel(face.pixels(), 0)).abs() < 0.01);
|
||||
assert!((mean_channel(views[0], 1) - mean_channel(face.pixels(), 1)).abs() < 0.01);
|
||||
|
||||
let (x, y, ww, hh) = SUNGLASSES_WINDOWS[1];
|
||||
assert!(
|
||||
x < 0.0 && y < 0.0,
|
||||
"the window starts outside the face crop"
|
||||
);
|
||||
assert!(x + ww > ALIGNED_EDGE as f32, "and is wider than it");
|
||||
assert!(y + hh < ALIGNED_EDGE as f32, "but stops short of the chin");
|
||||
// Centred horizontally on the face, so the two share a mean x.
|
||||
assert!((mean_channel(views[1], 0) - mean_channel(face.pixels(), 0)).abs() < 0.01);
|
||||
// Its first row lies above the face's first row.
|
||||
assert!(views[1][1] < face.pixels()[1]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn degenerate_landmarks_yield_no_head_crop() {
|
||||
let rgb = vec![0.5_f32; 64 * 64 * 3];
|
||||
let degenerate = [(50.0, 50.0); 5];
|
||||
assert!(head_views(Pixels::RgbF32(&rgb), 64, 64, °enerate).is_none());
|
||||
// And a buffer that is not the size it claims.
|
||||
let lm = shifted_scaled(1.0, 0.0, 0.0, 0.0);
|
||||
assert!(head_views(Pixels::RgbF32(&rgb), 60, 60, &lm).is_none());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn out_of_bounds_samples_read_black_rather_than_wrapping() {
|
||||
let rgb = vec![1.0_f32; 32 * 32 * 3];
|
||||
|
||||
@@ -0,0 +1,263 @@
|
||||
//! TRACES: FR-CULL-8a
|
||||
//! The two small classifiers behind a face's eye state (docs/faces.md §17).
|
||||
//!
|
||||
//! **OCEC** — *open closed eyes classification*, Hyodo 2025 — reads one
|
||||
//! 40×24 eye and answers P(open). **SGC** — *sunglasses classification*,
|
||||
//! Hyodo 2026 — reads a 48×48 head and answers P(sunglasses); it is shown
|
||||
//! two framings of each face and the higher answer stands, for the reason
|
||||
//! [`crate::align::SUNGLASSES_WINDOWS`] gives. Both are
|
||||
//! depthwise-separable CNNs of a few hundred kilobytes, both MIT with their
|
||||
//! weights, and both were exported with BatchNorm already folded, which is
|
||||
//! about the friendliest graph tract can be handed.
|
||||
//!
|
||||
//! Neither takes a plain buffer. [`EyeClassifier::classify`] takes an
|
||||
//! [`EyePatch`] and [`SunglassesClassifier::classify`] a [`HeadViews`], each
|
||||
//! constructible only by the crop in [`crate::align`] that puts the right
|
||||
//! pixels in it — the same defence [`crate::embed::Embedder`] makes with
|
||||
//! [`crate::align::Aligned112`], for the same reason: a classifier handed the
|
||||
//! wrong region returns a confident probability of nothing. Where the eye
|
||||
//! box comes from is [`crate::landmarks`]; [`EyeModels::read`] is the whole
|
||||
//! chain.
|
||||
//!
|
||||
//! # The graphs must have a fixed batch
|
||||
//!
|
||||
//! Both ship with a dynamic batch dimension, which tract will not analyse.
|
||||
//! `tools/fix-face-model-shapes.sh` pins it to 1, exactly as it does for the
|
||||
//! embedder; the shipped files are the pinned ones.
|
||||
//!
|
||||
//! # Pre-processing
|
||||
//!
|
||||
//! Read off the reference demos rather than assumed: RGB, `x / 255`, NCHW,
|
||||
//! the crop resized to the input with bilinear interpolation and **without**
|
||||
//! preserving its aspect. [`crate::align`]'s crops arrive already at the
|
||||
//! input size in `0..=1`, so there is nothing left to do but lay them out.
|
||||
|
||||
use ndarray::Array4;
|
||||
|
||||
use crate::align::{
|
||||
eye_box, eye_patch, head_views, EyePatch, HeadViews, EYE_PATCH_HEIGHT, EYE_PATCH_WIDTH,
|
||||
SUNGLASSES_EDGE,
|
||||
};
|
||||
use crate::eyes::{Eye, EyeReading};
|
||||
use crate::landmarks::{Landmarker, Landmarks};
|
||||
use crate::{FaceError, Pixels};
|
||||
use dr_inference_engine::{Form, Model, Role};
|
||||
|
||||
/// A loaded OCEC graph.
|
||||
pub struct EyeClassifier {
|
||||
session: Model,
|
||||
}
|
||||
|
||||
/// A loaded SGC graph.
|
||||
pub struct SunglassesClassifier {
|
||||
session: Model,
|
||||
}
|
||||
|
||||
/// Open a single-input, single-output classifier and check it is the shape
|
||||
/// the crop feeding it will be.
|
||||
///
|
||||
/// The check is against the *input*, because that is where these two graphs
|
||||
/// differ from each other and from everything else in this crate: an SGC file
|
||||
/// given to the eye classifier would otherwise be resized into by an eye
|
||||
/// patch, and answer. `expected` names the model in the error.
|
||||
fn open_classifier(
|
||||
bytes: &[u8],
|
||||
expected: &'static str,
|
||||
(h, w): (usize, usize),
|
||||
) -> Result<Model, FaceError> {
|
||||
let model = dr_inference_engine::open(Role::EyeClassifier, Form::F32, bytes)?;
|
||||
let acquired = model.acquire()?;
|
||||
let session = acquired.lock();
|
||||
|
||||
let input = session.inputs().first().ok_or(FaceError::WrongModel {
|
||||
expected,
|
||||
detail: "model has no inputs".into(),
|
||||
})?;
|
||||
let shape: Option<Vec<i64>> = input.dtype().tensor_shape().map(|s| s.to_vec());
|
||||
let want = [1, 3, h as i64, w as i64];
|
||||
if shape.as_deref() != Some(&want[..]) {
|
||||
return Err(FaceError::WrongModel {
|
||||
expected,
|
||||
detail: format!(
|
||||
"input '{}' is {:?}, expected {:?} (batch pinned to 1)",
|
||||
input.name(),
|
||||
shape,
|
||||
want
|
||||
),
|
||||
});
|
||||
}
|
||||
if session.outputs().len() != 1 {
|
||||
return Err(FaceError::WrongModel {
|
||||
expected,
|
||||
detail: format!("{} outputs, expected one", session.outputs().len()),
|
||||
});
|
||||
}
|
||||
drop(session);
|
||||
drop(acquired);
|
||||
Ok(model)
|
||||
}
|
||||
|
||||
/// Lay a `h × w` RGB crop out as the `[1, 3, h, w]` tensor both graphs take.
|
||||
fn to_nchw(pixels: &[f32], h: usize, w: usize) -> Array4<f32> {
|
||||
let mut input = Array4::<f32>::zeros((1, 3, h, w));
|
||||
for y in 0..h {
|
||||
for x in 0..w {
|
||||
for c in 0..3 {
|
||||
input[[0, c, y, x]] = pixels[(y * w + x) * 3 + c];
|
||||
}
|
||||
}
|
||||
}
|
||||
input
|
||||
}
|
||||
|
||||
/// Run a one-number classifier and read its sigmoid back, clamped.
|
||||
fn run_scalar(model: &Model, input: Array4<f32>, expected: &'static str) -> Result<f32, FaceError> {
|
||||
let acquired = model.acquire()?;
|
||||
let mut session = acquired.lock();
|
||||
let outputs = session
|
||||
.run(ort::inputs![
|
||||
ort::value::Tensor::from_array(input).map_err(FaceError::Inference)?
|
||||
])
|
||||
.map_err(FaceError::Inference)?;
|
||||
let (_, data) = outputs[0]
|
||||
.try_extract_tensor::<f32>()
|
||||
.map_err(FaceError::Inference)?;
|
||||
let Some(&p) = data.first() else {
|
||||
return Err(FaceError::WrongModel {
|
||||
expected,
|
||||
detail: "empty output".into(),
|
||||
});
|
||||
};
|
||||
// The graph ends in a sigmoid, so this is a clamp against rounding and
|
||||
// nothing more — the reference demo does the same.
|
||||
Ok(p.clamp(0.0, 1.0))
|
||||
}
|
||||
|
||||
impl EyeClassifier {
|
||||
pub fn from_path(path: impl AsRef<std::path::Path>) -> Result<Self, FaceError> {
|
||||
let bytes = std::fs::read(path).map_err(FaceError::ModelRead)?;
|
||||
Self::from_bytes(&bytes)
|
||||
}
|
||||
|
||||
pub fn from_bytes(bytes: &[u8]) -> Result<Self, FaceError> {
|
||||
Ok(Self {
|
||||
session: open_classifier(bytes, "OCEC", (EYE_PATCH_HEIGHT, EYE_PATCH_WIDTH))?,
|
||||
})
|
||||
}
|
||||
|
||||
/// P(open) for one eye.
|
||||
pub fn classify(&mut self, eye: &EyePatch) -> Result<f32, FaceError> {
|
||||
let input = to_nchw(eye.pixels(), EYE_PATCH_HEIGHT, EYE_PATCH_WIDTH);
|
||||
run_scalar(&self.session, input, "OCEC")
|
||||
}
|
||||
}
|
||||
|
||||
impl SunglassesClassifier {
|
||||
pub fn from_path(path: impl AsRef<std::path::Path>) -> Result<Self, FaceError> {
|
||||
let bytes = std::fs::read(path).map_err(FaceError::ModelRead)?;
|
||||
Self::from_bytes(&bytes)
|
||||
}
|
||||
|
||||
pub fn from_bytes(bytes: &[u8]) -> Result<Self, FaceError> {
|
||||
Ok(Self {
|
||||
session: open_classifier(bytes, "SGC", (SUNGLASSES_EDGE, SUNGLASSES_EDGE))?,
|
||||
})
|
||||
}
|
||||
|
||||
/// P(sunglasses) for one head: the highest answer over its framings.
|
||||
pub fn classify(&mut self, head: &HeadViews) -> Result<f32, FaceError> {
|
||||
let mut best = 0.0_f32;
|
||||
for view in head.views() {
|
||||
let input = to_nchw(view, SUNGLASSES_EDGE, SUNGLASSES_EDGE);
|
||||
best = best.max(run_scalar(&self.session, input, "SGC")?);
|
||||
}
|
||||
Ok(best)
|
||||
}
|
||||
}
|
||||
|
||||
/// The three models behind a reading, which is how every caller holds them.
|
||||
///
|
||||
/// One struct rather than three optional parameters, because a partial
|
||||
/// reading is not a reading: an eye state with no sunglasses number behind
|
||||
/// it is exactly the beach-photograph failure [`crate::eyes`] describes, and
|
||||
/// an eye box without the landmarks is the loose one this module replaced.
|
||||
/// The models load together or not at all.
|
||||
pub struct EyeModels {
|
||||
pub landmarks: Landmarker,
|
||||
pub eyes: EyeClassifier,
|
||||
pub sunglasses: SunglassesClassifier,
|
||||
}
|
||||
|
||||
impl EyeModels {
|
||||
pub fn from_paths(
|
||||
landmarks: impl AsRef<std::path::Path>,
|
||||
eyes: impl AsRef<std::path::Path>,
|
||||
sunglasses: impl AsRef<std::path::Path>,
|
||||
) -> Result<Self, FaceError> {
|
||||
Ok(Self {
|
||||
landmarks: Landmarker::from_path(landmarks)?,
|
||||
eyes: EyeClassifier::from_path(eyes)?,
|
||||
sunglasses: SunglassesClassifier::from_path(sunglasses)?,
|
||||
})
|
||||
}
|
||||
|
||||
/// Read one face's eyes, and hand back the dense landmarks it read them
|
||||
/// from.
|
||||
///
|
||||
/// `bbox` is the detector's `(x0, y0, x1, y1)` and `landmarks5` its five
|
||||
/// points, both in source pixels; the buffer is the one the aligned
|
||||
/// crop was taken from, so an eye is read from the same pixels the
|
||||
/// embedder saw the face in. `None` where nothing could be cut — a
|
||||
/// degenerate box or landmarks — which the caller stores as "not read".
|
||||
///
|
||||
/// The landmarks come back because they cost a model run the caller will
|
||||
/// not want to pay twice: stored beside the reading, a later pass over
|
||||
/// faces — head pose, expression — has them without the original.
|
||||
pub fn read(
|
||||
&mut self,
|
||||
px: Pixels<'_>,
|
||||
width: usize,
|
||||
height: usize,
|
||||
bbox: (f32, f32, f32, f32),
|
||||
landmarks5: &[(f32, f32); 5],
|
||||
) -> Result<Option<(EyeReading, Landmarks)>, FaceError> {
|
||||
let Some(lm) = self.landmarks.landmarks(px, width, height, bbox)? else {
|
||||
return Ok(None);
|
||||
};
|
||||
let Some(head) = head_views(px, width, height, landmarks5) else {
|
||||
return Ok(None);
|
||||
};
|
||||
let mut eye = |contour: &[(f32, f32)]| -> Result<Eye, FaceError> {
|
||||
// A hidden eye's contour can collapse to no width. Its numbers
|
||||
// are then zero — no pixels, no sharpness — which is what the
|
||||
// rule in `crate::eyes` reads as "not readable".
|
||||
let Some(b) = eye_box(contour) else {
|
||||
return Ok(Eye {
|
||||
open: 0.0,
|
||||
px: 0.0,
|
||||
sharpness: 0.0,
|
||||
});
|
||||
};
|
||||
let Some(patch) = eye_patch(px, width, height, b) else {
|
||||
return Ok(Eye {
|
||||
open: 0.0,
|
||||
px: 0.0,
|
||||
sharpness: 0.0,
|
||||
});
|
||||
};
|
||||
Ok(Eye {
|
||||
open: self.eyes.classify(&patch)?,
|
||||
px: patch.source_px(),
|
||||
sharpness: patch.sharpness(),
|
||||
})
|
||||
};
|
||||
let right = eye(&lm.right_eye())?;
|
||||
let left = eye(&lm.left_eye())?;
|
||||
let reading = EyeReading {
|
||||
right,
|
||||
left,
|
||||
sunglasses: self.sunglasses.classify(&head)?,
|
||||
};
|
||||
Ok(Some((reading, lm)))
|
||||
}
|
||||
}
|
||||
+36
-14
@@ -14,7 +14,8 @@
|
||||
|
||||
use ndarray::Array4;
|
||||
|
||||
use crate::{install_backend, FaceError};
|
||||
use crate::FaceError;
|
||||
use dr_inference_engine::{Form, Model, Role};
|
||||
|
||||
/// The graph's input edge, in pixels. See the module note: not configurable.
|
||||
pub const INPUT_EDGE: usize = 640;
|
||||
@@ -135,7 +136,10 @@ impl Detection {
|
||||
|
||||
/// A loaded SCRFD graph.
|
||||
pub struct Detector {
|
||||
session: ort::session::Session,
|
||||
session: Model,
|
||||
/// f32 or int8 — the int8 form finds a different set of faces and is a
|
||||
/// different detector in `model_id` (docs/inference.md §7).
|
||||
form: Form,
|
||||
/// Feature-map count: 3 for strides {8,16,32}, 4 for {8,16,32,64}.
|
||||
///
|
||||
/// Discovered from the output count rather than assumed, because both
|
||||
@@ -145,18 +149,29 @@ pub struct Detector {
|
||||
}
|
||||
|
||||
impl Detector {
|
||||
pub fn from_path(path: impl AsRef<std::path::Path>) -> Result<Self, FaceError> {
|
||||
let bytes = std::fs::read(path).map_err(FaceError::ModelRead)?;
|
||||
Self::from_bytes(&bytes)
|
||||
/// Which form this detector was loaded from.
|
||||
pub fn form(&self) -> Form {
|
||||
self.form
|
||||
}
|
||||
|
||||
pub fn from_bytes(bytes: &[u8]) -> Result<Self, FaceError> {
|
||||
install_backend();
|
||||
/// Load the canonical f32 file at `path`, or the form the device's
|
||||
/// backend wants instead — the `.int8.onnx` beside it on a Hexagon —
|
||||
/// which [`Detector::form`] then reports.
|
||||
pub fn from_path(path: impl AsRef<std::path::Path>) -> Result<Self, FaceError> {
|
||||
let (path, form) = dr_inference_engine::resolve_model(Role::Detector, path.as_ref());
|
||||
let bytes = std::fs::read(path).map_err(FaceError::ModelRead)?;
|
||||
Self::from_bytes_in(&bytes, form)
|
||||
}
|
||||
|
||||
let session = ort::session::Session::builder()
|
||||
.map_err(FaceError::Inference)?
|
||||
.commit_from_memory(bytes)
|
||||
.map_err(FaceError::Inference)?;
|
||||
/// An f32 graph from memory.
|
||||
pub fn from_bytes(bytes: &[u8]) -> Result<Self, FaceError> {
|
||||
Self::from_bytes_in(bytes, Form::F32)
|
||||
}
|
||||
|
||||
fn from_bytes_in(bytes: &[u8], form: Form) -> Result<Self, FaceError> {
|
||||
let model = dr_inference_engine::open(Role::Detector, form, bytes)?;
|
||||
let acquired = model.acquire()?;
|
||||
let session = acquired.lock();
|
||||
|
||||
let n_out = session.outputs().len();
|
||||
if n_out % 3 != 0 || !(9..=12).contains(&n_out) {
|
||||
@@ -191,7 +206,13 @@ impl Detector {
|
||||
}
|
||||
}
|
||||
|
||||
Ok(Self { session, fmc })
|
||||
drop(session);
|
||||
drop(acquired);
|
||||
Ok(Self {
|
||||
session: model,
|
||||
form,
|
||||
fmc,
|
||||
})
|
||||
}
|
||||
|
||||
/// Stride levels this graph emits.
|
||||
@@ -223,8 +244,9 @@ impl Detector {
|
||||
let lb = Letterbox::fit(width as f32, height as f32);
|
||||
let input = lb.sample(rgb, width, height);
|
||||
|
||||
let outputs = self
|
||||
.session
|
||||
let acquired = self.session.acquire()?;
|
||||
let mut session = acquired.lock();
|
||||
let outputs = session
|
||||
.run(ort::inputs![
|
||||
ort::value::Tensor::from_array(input).map_err(FaceError::Inference)?
|
||||
])
|
||||
|
||||
+17
-11
@@ -14,7 +14,8 @@ use ndarray::Array4;
|
||||
|
||||
use crate::align::{Aligned112, ALIGNED_EDGE};
|
||||
use crate::embedding::{normalise, Embedding, ModelId, EMBEDDING_DIM};
|
||||
use crate::{install_backend, FaceError};
|
||||
use crate::FaceError;
|
||||
use dr_inference_engine::{Form, Model, Role};
|
||||
|
||||
/// What one pass of the embedder produces: the direction, and the length.
|
||||
///
|
||||
@@ -53,7 +54,7 @@ impl Embedded {
|
||||
|
||||
/// A loaded ArcFace graph.
|
||||
pub struct Embedder {
|
||||
session: ort::session::Session,
|
||||
session: Model,
|
||||
model: ModelId,
|
||||
}
|
||||
|
||||
@@ -64,12 +65,11 @@ impl Embedder {
|
||||
}
|
||||
|
||||
pub fn from_bytes(bytes: &[u8], model: ModelId) -> Result<Self, FaceError> {
|
||||
install_backend();
|
||||
|
||||
let session = ort::session::Session::builder()
|
||||
.map_err(FaceError::Inference)?
|
||||
.commit_from_memory(bytes)
|
||||
.map_err(FaceError::Inference)?;
|
||||
// Always the f32 form: an embedding must compare across devices
|
||||
// (docs/inference.md §7), and the engine pins this role to it.
|
||||
let loaded = dr_inference_engine::open(Role::Embedder, Form::F32, bytes)?;
|
||||
let acquired = loaded.acquire()?;
|
||||
let session = acquired.lock();
|
||||
|
||||
// One output, `[1, 512]`. Checked because an ArcFace variant with a
|
||||
// different embedding width would otherwise be read as a truncated
|
||||
@@ -90,7 +90,12 @@ impl Embedder {
|
||||
});
|
||||
}
|
||||
|
||||
Ok(Self { session, model })
|
||||
drop(session);
|
||||
drop(acquired);
|
||||
Ok(Self {
|
||||
session: loaded,
|
||||
model,
|
||||
})
|
||||
}
|
||||
|
||||
pub fn model(&self) -> &ModelId {
|
||||
@@ -111,8 +116,9 @@ impl Embedder {
|
||||
}
|
||||
}
|
||||
|
||||
let outputs = self
|
||||
.session
|
||||
let acquired = self.session.acquire()?;
|
||||
let mut session = acquired.lock();
|
||||
let outputs = session
|
||||
.run(ort::inputs![
|
||||
ort::value::Tensor::from_array(input).map_err(FaceError::Inference)?
|
||||
])
|
||||
|
||||
@@ -0,0 +1,263 @@
|
||||
//! TRACES: FR-CULL-8a
|
||||
//! What a face's eyes are doing, and how the numbers behind it are read.
|
||||
//!
|
||||
//! Model-free: the models in [`crate::classify`] produce the numbers, and
|
||||
//! everything that interprets them — the catalog's filter, the People
|
||||
//! screen's label — comes through here, so a threshold lives in exactly one
|
||||
//! place.
|
||||
//!
|
||||
//! # Seven numbers, one answer
|
||||
//!
|
||||
//! An eye classifier answers "open or closed" for whatever it is shown, and
|
||||
//! it is shown three things it cannot answer for. **Dark glass**: over
|
||||
//! sunglasses it answers anyway, confidently, for a state that cannot be
|
||||
//! seen — so the reading carries P(sunglasses) from a classifier that looks
|
||||
//! at the whole head, and that takes precedence. **A smear**: a soft eye is
|
||||
//! not a closed one, but shown a blur the classifier says "closed" with the
|
||||
//! same confidence it says anything, and on the reference library that was
|
||||
//! the commonest wrong answer of all — small faces, motion, a proxy where
|
||||
//! the native render should have been. So each eye carries how many source
|
||||
//! pixels it spanned and how sharp the patch was, and an eye under either
|
||||
//! floor is not asked. **A cheek**: a head turned far enough hides its far
|
||||
//! eye, and the landmark contour of a hidden eye collapses to a sliver; an
|
||||
//! eye much narrower than its partner is not asked either.
|
||||
//!
|
||||
//! The two eyes are kept apart rather than averaged. A wink is one eye
|
||||
//! closed, and averaging it lands at 0.5 — the one value that says the least.
|
||||
//! [`EyeState::Open`] requires every eye that *could be read* to be open;
|
||||
//! a face with no readable eye is [`EyeState::Unreadable`], which is not a
|
||||
//! blink and not open, and a filter for either leaves it alone.
|
||||
|
||||
/// One eye's numbers.
|
||||
#[derive(Debug, Clone, Copy, PartialEq)]
|
||||
pub struct Eye {
|
||||
/// P(open), the classifier's sigmoid.
|
||||
pub open: f32,
|
||||
/// Source pixels across the eye box — [`crate::align::EyePatch::source_px`].
|
||||
pub px: f32,
|
||||
/// [`crate::align::EyePatch::sharpness`] of the patch the classifier saw.
|
||||
pub sharpness: f32,
|
||||
}
|
||||
|
||||
/// The numbers the models produced for one face.
|
||||
///
|
||||
/// Stored per face, nullable as a whole: a face indexed before the eye models
|
||||
/// existed, or on a device without them, has no reading rather than a
|
||||
/// reading of zeros.
|
||||
#[derive(Debug, Clone, Copy, PartialEq)]
|
||||
pub struct EyeReading {
|
||||
/// The subject's **right** eye — image-left.
|
||||
pub right: Eye,
|
||||
/// The subject's **left** eye — image-right.
|
||||
pub left: Eye,
|
||||
/// P(the head wears sunglasses).
|
||||
pub sunglasses: f32,
|
||||
}
|
||||
|
||||
/// Above this an eye is open. The classifier's own decision point; its
|
||||
/// training put the two classes either side of a sigmoid and this is where
|
||||
/// the sigmoid crosses.
|
||||
pub const EYES_OPEN_THRESHOLD: f32 = 0.5;
|
||||
|
||||
/// Above this the head wears sunglasses and the eye readings are moot.
|
||||
pub const SUNGLASSES_THRESHOLD: f32 = 0.5;
|
||||
|
||||
/// Fewest source pixels across an eye box for the eye to be read.
|
||||
///
|
||||
/// The classifier was trained on eyes down to about a dozen pixels wide
|
||||
/// (its reference footage averaged 15–21); below that the 40-pixel patch is
|
||||
/// an interpolation of nothing, and the answer is noise that reads as
|
||||
/// "closed". docs/faces.md §17.3 has the measurement behind the number.
|
||||
pub const MIN_EYE_PX: f32 = 12.0;
|
||||
|
||||
/// Least [`Eye::sharpness`] for the eye to be read.
|
||||
///
|
||||
/// The same measure as the face's `min_sharpness`, over the eye patch, and
|
||||
/// chosen the same way: the value under which the open-eyed faces of the
|
||||
/// reference sample were being called closed. docs/faces.md §17.3.
|
||||
pub const MIN_EYE_SHARPNESS: f32 = 0.02;
|
||||
|
||||
/// An eye narrower than this fraction of its partner is the far eye of a
|
||||
/// turned head, out of view behind the nose, and is not read.
|
||||
///
|
||||
/// A landmark model's contour for a hidden eye collapses towards the nose.
|
||||
/// Measured on twenty native renders of the reference library
|
||||
/// (docs/faces.md §17.4): profiles put the far eye at 0.02–0.43 of the near
|
||||
/// one, two three-quarter faces whose far eye read closed sat at 0.54, and
|
||||
/// every face looking at the camera — winks included, since a shut eye's
|
||||
/// box keeps its width — sat at 0.78 or more. 0.6 splits the gap.
|
||||
pub const HIDDEN_EYE_RATIO: f32 = 0.6;
|
||||
|
||||
/// What the reading says, for a screen or a filter.
|
||||
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
|
||||
pub enum EyeState {
|
||||
/// Every eye that could be read is open.
|
||||
Open,
|
||||
/// An eye that could be read is closed — a blink, or a wink.
|
||||
Closed,
|
||||
/// The eyes cannot be seen. Neither open nor closed, and a filter for
|
||||
/// either leaves the face alone.
|
||||
Sunglasses,
|
||||
/// No eye was sharp enough, large enough and in view to read. Neither
|
||||
/// open nor closed, like sunglasses, and left alone by every filter.
|
||||
Unreadable,
|
||||
}
|
||||
|
||||
impl Eye {
|
||||
/// Whether this eye can be read at all: enough pixels, sharp enough,
|
||||
/// and not the collapsed contour of a hidden eye — measured against
|
||||
/// `other`, its partner.
|
||||
pub fn readable(&self, other: &Eye) -> bool {
|
||||
self.px >= MIN_EYE_PX
|
||||
&& self.sharpness >= MIN_EYE_SHARPNESS
|
||||
&& self.px >= other.px * HIDDEN_EYE_RATIO
|
||||
}
|
||||
}
|
||||
|
||||
impl EyeReading {
|
||||
pub fn state(&self) -> EyeState {
|
||||
if self.sunglasses >= SUNGLASSES_THRESHOLD {
|
||||
return EyeState::Sunglasses;
|
||||
}
|
||||
let readable = [
|
||||
self.right.readable(&self.left).then_some(self.right.open),
|
||||
self.left.readable(&self.right).then_some(self.left.open),
|
||||
];
|
||||
let mut any = false;
|
||||
for open in readable.into_iter().flatten() {
|
||||
any = true;
|
||||
if open < EYES_OPEN_THRESHOLD {
|
||||
return EyeState::Closed;
|
||||
}
|
||||
}
|
||||
if any {
|
||||
EyeState::Open
|
||||
} else {
|
||||
EyeState::Unreadable
|
||||
}
|
||||
}
|
||||
|
||||
/// Whether this is a face a "no one blinking" filter should drop.
|
||||
///
|
||||
/// The filter's question, rather than [`EyeState`]'s four-way answer,
|
||||
/// because the two differ on exactly the cases that matter: a face
|
||||
/// behind sunglasses, or one whose eyes could not be read, is not open
|
||||
/// — and it is not a blink either. Only [`EyeState::Closed`] is one.
|
||||
pub fn is_blink(&self) -> bool {
|
||||
self.state() == EyeState::Closed
|
||||
}
|
||||
}
|
||||
|
||||
impl EyeState {
|
||||
/// The word the People screen puts on the face.
|
||||
pub fn label(&self) -> &'static str {
|
||||
match self {
|
||||
EyeState::Open => "Eyes open",
|
||||
EyeState::Closed => "Eyes closed",
|
||||
EyeState::Sunglasses => "Sunglasses",
|
||||
EyeState::Unreadable => "Eyes unclear",
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
fn eye(open: f32) -> Eye {
|
||||
Eye {
|
||||
open,
|
||||
px: 40.0,
|
||||
sharpness: 0.1,
|
||||
}
|
||||
}
|
||||
|
||||
fn reading(right: f32, left: f32, sunglasses: f32) -> EyeReading {
|
||||
EyeReading {
|
||||
right: eye(right),
|
||||
left: eye(left),
|
||||
sunglasses,
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn both_eyes_open_is_open() {
|
||||
assert_eq!(reading(0.9, 0.8, 0.1).state(), EyeState::Open);
|
||||
assert!(!reading(0.9, 0.8, 0.1).is_blink());
|
||||
}
|
||||
|
||||
/// A wink is not "eyes open": one eye closed lands the same place a
|
||||
/// blink does, and a filter for "nobody blinking" should drop it.
|
||||
#[test]
|
||||
fn one_eye_closed_is_closed() {
|
||||
assert_eq!(reading(0.9, 0.2, 0.1).state(), EyeState::Closed);
|
||||
assert_eq!(reading(0.2, 0.9, 0.1).state(), EyeState::Closed);
|
||||
assert!(reading(0.2, 0.9, 0.1).is_blink());
|
||||
}
|
||||
|
||||
/// The whole reason the sunglasses number exists: whatever the eye
|
||||
/// classifier says over dark glass, it is not a reading of the eyes.
|
||||
#[test]
|
||||
fn sunglasses_override_the_eye_readings_either_way() {
|
||||
assert_eq!(reading(0.9, 0.9, 0.8).state(), EyeState::Sunglasses);
|
||||
assert_eq!(reading(0.1, 0.1, 0.8).state(), EyeState::Sunglasses);
|
||||
assert!(!reading(0.1, 0.1, 0.8).is_blink());
|
||||
}
|
||||
|
||||
/// A soft or tiny eye is not asked; if neither can be, the face is
|
||||
/// unreadable rather than closed.
|
||||
#[test]
|
||||
fn a_soft_or_tiny_eye_is_not_read() {
|
||||
let mut r = reading(0.1, 0.9, 0.0);
|
||||
r.right.sharpness = MIN_EYE_SHARPNESS / 2.0;
|
||||
assert_eq!(r.state(), EyeState::Open, "the soft closed eye is ignored");
|
||||
|
||||
let mut r = reading(0.1, 0.9, 0.0);
|
||||
r.right.px = MIN_EYE_PX - 1.0;
|
||||
assert_eq!(r.state(), EyeState::Open, "the tiny closed eye is ignored");
|
||||
|
||||
let mut r = reading(0.1, 0.1, 0.0);
|
||||
r.right.sharpness = 0.0;
|
||||
r.left.px = 3.0;
|
||||
assert_eq!(r.state(), EyeState::Unreadable);
|
||||
assert!(!r.is_blink());
|
||||
assert_eq!(r.state().label(), "Eyes unclear");
|
||||
}
|
||||
|
||||
/// A profile: the far eye's contour collapses, and the sliver is not
|
||||
/// read. The near eye still decides.
|
||||
#[test]
|
||||
fn a_turned_heads_collapsed_far_eye_is_not_read() {
|
||||
let mut r = reading(0.05, 0.95, 0.0);
|
||||
r.right.px = 40.0 * HIDDEN_EYE_RATIO - 1.0;
|
||||
assert!(!r.right.readable(&r.left));
|
||||
assert_eq!(r.state(), EyeState::Open);
|
||||
|
||||
let mut blink = reading(0.95, 0.05, 0.0);
|
||||
blink.right.px = 40.0 * HIDDEN_EYE_RATIO - 1.0;
|
||||
assert_eq!(blink.state(), EyeState::Closed);
|
||||
|
||||
// Both eyes narrow but alike is not a turned head: both count.
|
||||
let mut small = reading(0.05, 0.95, 0.0);
|
||||
small.right.px = 14.0;
|
||||
small.left.px = 14.0;
|
||||
assert_eq!(small.state(), EyeState::Closed);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn the_thresholds_are_inclusive_at_the_decision_point() {
|
||||
assert_eq!(
|
||||
reading(EYES_OPEN_THRESHOLD, EYES_OPEN_THRESHOLD, 0.0).state(),
|
||||
EyeState::Open
|
||||
);
|
||||
assert_eq!(
|
||||
reading(1.0, 1.0, SUNGLASSES_THRESHOLD).state(),
|
||||
EyeState::Sunglasses
|
||||
);
|
||||
let mut r = reading(1.0, 1.0, 0.0);
|
||||
r.right.px = MIN_EYE_PX;
|
||||
r.left.px = MIN_EYE_PX;
|
||||
r.right.sharpness = MIN_EYE_SHARPNESS;
|
||||
assert!(r.right.readable(&r.left));
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,263 @@
|
||||
//! TRACES: FR-CULL-8a
|
||||
//! Dense facial landmarks — InsightFace's `2d106det` (docs/faces.md §17.2).
|
||||
//!
|
||||
//! SCRFD's five points place a face; they do not place an eye. Its eye
|
||||
//! point is loose enough that a window centred on it left the eye in a
|
||||
//! corner on turned and smiling heads, and two model-free ways of
|
||||
//! re-centring it made things worse. So a second model draws the eye's lid
|
||||
//! contour, and the eye box is cut from that.
|
||||
//!
|
||||
//! **Why this one.** Three were measured on the same faces — MediaPipe Face
|
||||
//! Mesh V2, PIPNet and this — and tied on what the eye classifier made of
|
||||
//! their boxes (22 of 25 open eyes read open, against 19 from the SCRFD
|
||||
//! point). This is the cheapest of the three by a wide margin (5 MB, 106
|
||||
//! points, ~24 ms in tract), and it is under the grant the detector and
|
||||
//! embedder already carry rather than a new one to read.
|
||||
//!
|
||||
//! # Pre-processing
|
||||
//!
|
||||
//! Ported from InsightFace's `landmark.py`: a square crop centred on the
|
||||
//! detector box, 1.5× its longer edge, resized to 192; **RGB in 0..255**
|
||||
//! (the graph carries its own `bn_data` normalisation, so `input_mean` is
|
||||
//! 0 and `input_std` 1); 106 `(x, y)` in −1..1 mapped back through
|
||||
//! `(p + 1) · 96`. The graph's batch dimension is the literal `None` and
|
||||
//! is pinned to 1 by `tools/fix-face-model-shapes.sh`, like the embedder's.
|
||||
//!
|
||||
//! # The layout
|
||||
//!
|
||||
//! Checked by drawing the points on the reference faces rather than taken
|
||||
//! from a diagram: the subject's right eye (image-left) is points 33–42,
|
||||
//! the left 87–96, ten each round the lids.
|
||||
|
||||
use ndarray::Array4;
|
||||
|
||||
use crate::align::crop_box;
|
||||
use crate::{FaceError, Pixels};
|
||||
use dr_inference_engine::{Form, Model, Role};
|
||||
|
||||
/// The graph's input edge, in pixels.
|
||||
pub const INPUT_EDGE: usize = 192;
|
||||
|
||||
/// How many points the model returns.
|
||||
pub const POINTS: usize = 106;
|
||||
|
||||
/// The crop's edge as a multiple of the detector box's longer edge.
|
||||
const CROP_SCALE: f32 = 1.5;
|
||||
|
||||
/// The span of the frame, in long-edge units, the packed form covers: a
|
||||
/// quarter of the frame outside each edge.
|
||||
pub const PACKED_RANGE: (f32, f32) = (-0.25, 1.25);
|
||||
|
||||
/// Bytes the packed form of one face's landmarks takes.
|
||||
pub const PACKED_BYTES: usize = POINTS * 4;
|
||||
|
||||
/// Point indices of the subject's right eye's lid contour (image-left).
|
||||
pub const RIGHT_EYE: [usize; 10] = [33, 34, 35, 36, 37, 38, 39, 40, 41, 42];
|
||||
/// Point indices of the subject's left eye's lid contour (image-right).
|
||||
pub const LEFT_EYE: [usize; 10] = [87, 88, 89, 90, 91, 92, 93, 94, 95, 96];
|
||||
|
||||
/// The 106 points of one face, in **source pixels**.
|
||||
#[derive(Debug, Clone, PartialEq)]
|
||||
pub struct Landmarks {
|
||||
pub points: [(f32, f32); POINTS],
|
||||
}
|
||||
|
||||
impl Landmarks {
|
||||
/// Storage form: `106 × (x, y)` as little-endian **`u16` fixed point**
|
||||
/// over the frame, 424 bytes.
|
||||
///
|
||||
/// Each coordinate is normalised by `long_edge` like the five points the
|
||||
/// catalog already keeps, then mapped over [`PACKED_RANGE`] — a quarter
|
||||
/// of the frame either side of it, because a landmark on a face at the
|
||||
/// edge does land outside the image — onto 0..65535. That is 0.14 source
|
||||
/// pixels on a 6000-pixel frame. `f16` would be the same size and worse:
|
||||
/// its three significant figures near 1.0 are six pixels at that scale,
|
||||
/// and the eye contour this is kept for is drawn to the pixel.
|
||||
pub fn to_packed_bytes(&self, long_edge: f32) -> Vec<u8> {
|
||||
let (lo, hi) = PACKED_RANGE;
|
||||
let pack = |v: f32| -> [u8; 2] {
|
||||
let t = ((v / long_edge - lo) / (hi - lo)).clamp(0.0, 1.0);
|
||||
((t * 65535.0).round() as u16).to_le_bytes()
|
||||
};
|
||||
let mut out = Vec::with_capacity(POINTS * 4);
|
||||
for &(x, y) in &self.points {
|
||||
out.extend_from_slice(&pack(x));
|
||||
out.extend_from_slice(&pack(y));
|
||||
}
|
||||
out
|
||||
}
|
||||
|
||||
/// [`Self::to_packed_bytes`] read back, into source pixels of a frame
|
||||
/// with this `long_edge`. `None` for a blob of the wrong length.
|
||||
pub fn from_packed_bytes(bytes: &[u8], long_edge: f32) -> Option<Self> {
|
||||
if bytes.len() != POINTS * 4 {
|
||||
return None;
|
||||
}
|
||||
let (lo, hi) = PACKED_RANGE;
|
||||
let unpack = |b: &[u8]| -> f32 {
|
||||
let t = u16::from_le_bytes([b[0], b[1]]) as f32 / 65535.0;
|
||||
(t * (hi - lo) + lo) * long_edge
|
||||
};
|
||||
let mut points = [(0.0_f32, 0.0_f32); POINTS];
|
||||
for (i, p) in points.iter_mut().enumerate() {
|
||||
let at = i * 4;
|
||||
*p = (unpack(&bytes[at..at + 2]), unpack(&bytes[at + 2..at + 4]));
|
||||
}
|
||||
Some(Self { points })
|
||||
}
|
||||
|
||||
/// The lid contour of the subject's right eye.
|
||||
pub fn right_eye(&self) -> [(f32, f32); 10] {
|
||||
RIGHT_EYE.map(|i| self.points[i])
|
||||
}
|
||||
|
||||
/// The lid contour of the subject's left eye.
|
||||
pub fn left_eye(&self) -> [(f32, f32); 10] {
|
||||
LEFT_EYE.map(|i| self.points[i])
|
||||
}
|
||||
}
|
||||
|
||||
/// A loaded `2d106det` graph.
|
||||
pub struct Landmarker {
|
||||
session: Model,
|
||||
}
|
||||
|
||||
impl Landmarker {
|
||||
pub fn from_path(path: impl AsRef<std::path::Path>) -> Result<Self, FaceError> {
|
||||
let bytes = std::fs::read(path).map_err(FaceError::ModelRead)?;
|
||||
Self::from_bytes(&bytes)
|
||||
}
|
||||
|
||||
pub fn from_bytes(bytes: &[u8]) -> Result<Self, FaceError> {
|
||||
let model = dr_inference_engine::open(Role::Landmarks, Form::F32, bytes)?;
|
||||
let acquired = model.acquire()?;
|
||||
let session = acquired.lock();
|
||||
|
||||
let input = session.inputs().first().ok_or(FaceError::WrongModel {
|
||||
expected: "2d106det",
|
||||
detail: "model has no inputs".into(),
|
||||
})?;
|
||||
let shape: Option<Vec<i64>> = input.dtype().tensor_shape().map(|s| s.to_vec());
|
||||
let want = [1, 3, INPUT_EDGE as i64, INPUT_EDGE as i64];
|
||||
if shape.as_deref() != Some(&want[..]) {
|
||||
return Err(FaceError::WrongModel {
|
||||
expected: "2d106det",
|
||||
detail: format!(
|
||||
"input '{}' is {:?}, expected {:?} (batch pinned to 1)",
|
||||
input.name(),
|
||||
shape,
|
||||
want
|
||||
),
|
||||
});
|
||||
}
|
||||
let out = session.outputs().first().ok_or(FaceError::WrongModel {
|
||||
expected: "2d106det",
|
||||
detail: "model has no outputs".into(),
|
||||
})?;
|
||||
let last: Option<i64> = out.dtype().tensor_shape().and_then(|d| d.last().copied());
|
||||
if last != Some((POINTS * 2) as i64) {
|
||||
return Err(FaceError::WrongModel {
|
||||
expected: "2d106det",
|
||||
detail: format!(
|
||||
"output '{}' is {:?}-wide, expected {}",
|
||||
out.name(),
|
||||
last,
|
||||
POINTS * 2
|
||||
),
|
||||
});
|
||||
}
|
||||
drop(session);
|
||||
drop(acquired);
|
||||
Ok(Self { session: model })
|
||||
}
|
||||
|
||||
/// The landmarks of the face in `bbox` — `(x0, y0, x1, y1)` in source
|
||||
/// pixels, the detector's box — read from the source.
|
||||
///
|
||||
/// `None` for a box with no area or a buffer that is not the size it
|
||||
/// claims, as every crop here.
|
||||
pub fn landmarks(
|
||||
&mut self,
|
||||
px: Pixels<'_>,
|
||||
width: usize,
|
||||
height: usize,
|
||||
bbox: (f32, f32, f32, f32),
|
||||
) -> Result<Option<Landmarks>, FaceError> {
|
||||
let (w, h) = (bbox.2 - bbox.0, bbox.3 - bbox.1);
|
||||
let side = w.max(h) * CROP_SCALE;
|
||||
let (cx, cy) = ((bbox.0 + bbox.2) / 2.0, (bbox.1 + bbox.3) / 2.0);
|
||||
let (x0, y0) = (cx - side / 2.0, cy - side / 2.0);
|
||||
let Some(crop) = crop_box(
|
||||
px,
|
||||
width,
|
||||
height,
|
||||
(x0, y0, side, side),
|
||||
INPUT_EDGE,
|
||||
INPUT_EDGE,
|
||||
) else {
|
||||
return Ok(None);
|
||||
};
|
||||
|
||||
let e = INPUT_EDGE;
|
||||
let mut input = Array4::<f32>::zeros((1, 3, e, e));
|
||||
for y in 0..e {
|
||||
for x in 0..e {
|
||||
for c in 0..3 {
|
||||
input[[0, c, y, x]] = crop[(y * e + x) * 3 + c] * 255.0;
|
||||
}
|
||||
}
|
||||
}
|
||||
let acquired = self.session.acquire()?;
|
||||
let mut session = acquired.lock();
|
||||
let outputs = session
|
||||
.run(ort::inputs![
|
||||
ort::value::Tensor::from_array(input).map_err(FaceError::Inference)?
|
||||
])
|
||||
.map_err(FaceError::Inference)?;
|
||||
let (_, data) = outputs[0]
|
||||
.try_extract_tensor::<f32>()
|
||||
.map_err(FaceError::Inference)?;
|
||||
if data.len() < POINTS * 2 {
|
||||
return Err(FaceError::WrongModel {
|
||||
expected: "2d106det",
|
||||
detail: format!("got {} values, expected {}", data.len(), POINTS * 2),
|
||||
});
|
||||
}
|
||||
|
||||
// −1..1 in the crop → crop pixels → source pixels.
|
||||
let scale = side / e as f32;
|
||||
let half = e as f32 / 2.0;
|
||||
let mut points = [(0.0_f32, 0.0_f32); POINTS];
|
||||
for (i, p) in points.iter_mut().enumerate() {
|
||||
let (u, v) = ((data[2 * i] + 1.0) * half, (data[2 * i + 1] + 1.0) * half);
|
||||
*p = (x0 + u * scale, y0 + v * scale);
|
||||
}
|
||||
Ok(Some(Landmarks { points }))
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
/// Packed and unpacked, every point comes back within a fifth of a
|
||||
/// source pixel on a 6000-pixel frame — including one outside the
|
||||
/// image, which a face at the edge does produce.
|
||||
#[test]
|
||||
fn dense_landmarks_round_trip_through_their_packed_bytes() {
|
||||
let mut points = [(0.0_f32, 0.0_f32); POINTS];
|
||||
for (i, p) in points.iter_mut().enumerate() {
|
||||
*p = (i as f32 * 37.3 - 200.0, 5900.0 - i as f32 * 11.1);
|
||||
}
|
||||
let lm = Landmarks { points };
|
||||
let bytes = lm.to_packed_bytes(6000.0);
|
||||
assert_eq!(bytes.len(), PACKED_BYTES);
|
||||
assert_eq!(PACKED_BYTES, 424);
|
||||
let back = Landmarks::from_packed_bytes(&bytes, 6000.0).unwrap();
|
||||
for (a, b) in lm.points.iter().zip(back.points.iter()) {
|
||||
assert!((a.0 - b.0).abs() < 0.2, "{} vs {}", a.0, b.0);
|
||||
assert!((a.1 - b.1).abs() < 0.2, "{} vs {}", a.1, b.1);
|
||||
}
|
||||
assert!(Landmarks::from_packed_bytes(&bytes[..100], 6000.0).is_none());
|
||||
}
|
||||
}
|
||||
+32
-19
@@ -1,8 +1,10 @@
|
||||
//! Faces and identity (S14, docs/faces.md).
|
||||
//!
|
||||
//! Two models, run over the proxy tier, producing per face a box, five
|
||||
//! Two models, run over the native render, producing per face a box, five
|
||||
//! landmarks, a confidence and a 512-d embedding (FR-CULL-8) — and then the
|
||||
//! arithmetic that turns embeddings into people (FR-CULL-9, FR-CULL-10).
|
||||
//! arithmetic that turns embeddings into people (FR-CULL-9, FR-CULL-10). Two
|
||||
//! more, optional, read each face's eyes and whether sunglasses hide them
|
||||
//! (FR-CULL-8a, [`classify`] and [`eyes`]).
|
||||
//!
|
||||
//! Like `dr-segment`, this crate is **device-free**: no GPU adapter, no
|
||||
//! Slint, nothing that needs a display. Unlike `dr-segment`, it carries **no
|
||||
@@ -20,7 +22,8 @@
|
||||
//! runtime; this crate takes bytes and never fetches anything.
|
||||
//!
|
||||
//! docs/faces.md §2 is the full reading, including what would have to change
|
||||
//! for that to stop being true.
|
||||
//! for that to stop being true. The eye-state models are the exception: MIT,
|
||||
//! weights and all, and shipped in `models/face/` (docs/faces.md §17).
|
||||
//!
|
||||
//! # Why the runtime is split behind a feature
|
||||
//!
|
||||
@@ -34,12 +37,17 @@
|
||||
pub mod align;
|
||||
pub mod assign;
|
||||
pub mod calibrate;
|
||||
#[cfg(feature = "inference")]
|
||||
pub mod classify;
|
||||
pub mod cluster;
|
||||
#[cfg(feature = "inference")]
|
||||
pub mod detect;
|
||||
#[cfg(feature = "inference")]
|
||||
pub mod embed;
|
||||
pub mod embedding;
|
||||
pub mod eyes;
|
||||
#[cfg(feature = "inference")]
|
||||
pub mod landmarks;
|
||||
pub mod naming;
|
||||
pub mod neighbours;
|
||||
|
||||
@@ -65,10 +73,13 @@ pub mod neighbours;
|
||||
pub const MIN_CROP_EDGE: u32 = 1025;
|
||||
|
||||
pub use align::{
|
||||
warp, warp_pixels, Aligned112, Pixels, Similarity, ALIGNED_EDGE, ARCFACE_TEMPLATE,
|
||||
crop_box, eye_box, eye_patch, head_views, warp, warp_pixels, Aligned112, EyePatch, HeadViews,
|
||||
Pixels, Similarity, ALIGNED_EDGE, ARCFACE_TEMPLATE,
|
||||
};
|
||||
pub use assign::{identity_shares, RIVAL_FLOOR, TOP_MATCHES};
|
||||
pub use calibrate::{Calibration, Pairs, ReliabilityBand};
|
||||
#[cfg(feature = "inference")]
|
||||
pub use classify::{EyeClassifier, EyeModels, SunglassesClassifier};
|
||||
pub use cluster::{
|
||||
cluster, cluster_scored, split, Candidate, Cluster, Grouping, DEFAULT_MERGE_PROBABILITY,
|
||||
};
|
||||
@@ -79,6 +90,12 @@ pub use embed::{Embedded, Embedder};
|
||||
pub use embedding::{
|
||||
in_gallery, read_f16_bytes, Embedding, ModelId, EMBEDDING_DIM, MIN_GALLERY_QUALITY,
|
||||
};
|
||||
pub use eyes::{
|
||||
Eye, EyeReading, EyeState, EYES_OPEN_THRESHOLD, HIDDEN_EYE_RATIO, MIN_EYE_PX,
|
||||
MIN_EYE_SHARPNESS, SUNGLASSES_THRESHOLD,
|
||||
};
|
||||
#[cfg(feature = "inference")]
|
||||
pub use landmarks::{Landmarker, Landmarks};
|
||||
pub use naming::{name_for_instance, name_instances, NamedFace};
|
||||
|
||||
/// What can go wrong between an image and a face.
|
||||
@@ -107,25 +124,21 @@ pub enum FaceError {
|
||||
ImageShape { expected: usize, got: usize },
|
||||
}
|
||||
|
||||
/// Install tract as `ort`'s backend.
|
||||
///
|
||||
/// Idempotent, and it must happen before any other `ort` call: with
|
||||
/// `alternative-backend` there is no linked runtime to fall back on, so an
|
||||
/// un-set API is a panic rather than a slow path. Same helper as
|
||||
/// `dr-segment::semantic`, for the same reason.
|
||||
#[cfg(feature = "inference")]
|
||||
pub(crate) fn install_backend() {
|
||||
use std::sync::Once;
|
||||
static ONCE: Once = Once::new();
|
||||
ONCE.call_once(|| {
|
||||
let _ = ort::set_api(ort_tract::api());
|
||||
});
|
||||
impl From<dr_inference_engine::Error> for FaceError {
|
||||
fn from(e: dr_inference_engine::Error) -> Self {
|
||||
match e {
|
||||
dr_inference_engine::Error::Inference(e) => FaceError::Inference(e),
|
||||
dr_inference_engine::Error::Io(e) => FaceError::ModelRead(e),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// [`install_backend`] for the M1 probe example, which drives `ort` directly
|
||||
/// rather than through [`detect::Detector`] so it can report the raw error.
|
||||
/// Make sure `ort` has a backend, for the M1 probe example, which drives
|
||||
/// `ort` directly rather than through [`detect::Detector`] so it can report
|
||||
/// the raw error. Every other path goes through `dr-inference-engine`.
|
||||
#[cfg(feature = "inference")]
|
||||
#[doc(hidden)]
|
||||
pub fn install_backend_for_probe() {
|
||||
install_backend();
|
||||
dr_inference_engine::ensure_runtime();
|
||||
}
|
||||
|
||||
@@ -7,6 +7,11 @@ license.workspace = true
|
||||
|
||||
[dependencies]
|
||||
dr-types.workspace = true
|
||||
# The merge's geometry (FR-MRG-10): rotations, the focal length and the
|
||||
# projections, solved on proxies by dr-pano and consumed here per chunk. The
|
||||
# geometry alone — no keypoint model, no runtime — which is what the
|
||||
# workspace entry turns off.
|
||||
dr-pano.workspace = true
|
||||
dr-decode.workspace = true
|
||||
dr-pipeline.workspace = true
|
||||
# The watershed's pixel passes are here because they are shaders; everything
|
||||
|
||||
+233
-7
@@ -101,6 +101,16 @@ pub struct AdjustPass {
|
||||
/// switched on does not build a pipeline layout mid-frame.
|
||||
linear_bind_group_layout: wgpu::BindGroupLayout,
|
||||
linear_pipeline_layout: wgpu::PipelineLayout,
|
||||
/// TRACES: FR-MRG-2
|
||||
/// A third layout, writing `rgba32float`, for the camera-space tap a
|
||||
/// merge reads (`OutputMode::CameraLinear`). Same reasoning as the
|
||||
/// linear one: the format is in the layout, so a format is a layout.
|
||||
camera_bind_group_layout: wgpu::BindGroupLayout,
|
||||
camera_pipeline_layout: wgpu::PipelineLayout,
|
||||
/// The camera-space texture the last `render_camera_linear` wrote.
|
||||
/// Separate from `targets`: a different format, and a merge reads it
|
||||
/// back or samples it while the display targets go on being swapped.
|
||||
camera_target: Option<Target>,
|
||||
/// TRACES: FR-DEV-3d
|
||||
/// What the linear intermediate currently holds, and at what size.
|
||||
///
|
||||
@@ -303,6 +313,11 @@ impl AdjustPass {
|
||||
|
||||
pub const FORMAT: wgpu::TextureFormat = wgpu::TextureFormat::Rgba8Unorm;
|
||||
|
||||
/// TRACES: FR-MRG-2
|
||||
/// The camera-space tap's format: full precision, because what it holds
|
||||
/// is written back as a RAW at the sensor's own scale (FR-MRG-3).
|
||||
pub const CAMERA_FORMAT: wgpu::TextureFormat = wgpu::TextureFormat::Rgba32Float;
|
||||
|
||||
pub fn new(ctx: &GpuContext) -> Self {
|
||||
let bind_group_layout = Self::layout_writing(ctx, Self::FORMAT, "adjust-bgl");
|
||||
|
||||
@@ -330,6 +345,15 @@ impl AdjustPass {
|
||||
bind_group_layouts: &[Some(&linear_bind_group_layout)],
|
||||
immediate_size: 0,
|
||||
});
|
||||
let camera_bind_group_layout =
|
||||
Self::layout_writing(ctx, Self::CAMERA_FORMAT, "adjust-camera-bgl");
|
||||
let camera_pipeline_layout =
|
||||
ctx.device
|
||||
.create_pipeline_layout(&wgpu::PipelineLayoutDescriptor {
|
||||
label: Some("adjust-camera-layout"),
|
||||
bind_group_layouts: &[Some(&camera_bind_group_layout)],
|
||||
immediate_size: 0,
|
||||
});
|
||||
|
||||
// A 1x1 single-layer mask, bound when the edit has no local
|
||||
// adjustments. The generated shader never samples it — no layer block
|
||||
@@ -412,6 +436,9 @@ impl AdjustPass {
|
||||
detail: DetailRunner::new(ctx),
|
||||
linear_bind_group_layout,
|
||||
linear_pipeline_layout,
|
||||
camera_bind_group_layout,
|
||||
camera_pipeline_layout,
|
||||
camera_target: None,
|
||||
colour_key: None,
|
||||
colour_dispatches: 0,
|
||||
detail_dispatches: 0,
|
||||
@@ -549,6 +576,7 @@ impl AdjustPass {
|
||||
let layout = match shader.output_mode {
|
||||
OutputMode::Encoded => &self.pipeline_layout,
|
||||
OutputMode::LinearWorking => &self.linear_pipeline_layout,
|
||||
OutputMode::CameraLinear => &self.camera_pipeline_layout,
|
||||
};
|
||||
|
||||
let pipeline =
|
||||
@@ -1126,28 +1154,206 @@ impl AdjustPass {
|
||||
self.copy_output()
|
||||
}
|
||||
|
||||
/// TRACES: FR-MRG-2
|
||||
/// Render the camera-space tap: the source after its lens warp and
|
||||
/// nothing else, at full precision.
|
||||
///
|
||||
/// `shader` must come from `EditGraph::compose_camera_linear` — it is
|
||||
/// refused otherwise, for the reason `render_masked` refuses a linear
|
||||
/// one: the storage format is in the layout. The profile uniforms are
|
||||
/// filled neutral here rather than from the source, which is the whole
|
||||
/// point of the mode (`OutputMode::CameraLinear`): unit white balance,
|
||||
/// identity matrix, base curve off. The non-linear flag is kept, so a
|
||||
/// JPEG source is still linearised — camera space for a JPEG is the
|
||||
/// decoded values made linear, which is the best that exists.
|
||||
///
|
||||
/// The texture stays on the device for a merge's warp to sample; see
|
||||
/// [`Self::camera_texture`] and [`Self::read_camera_linear`].
|
||||
pub fn render_camera_linear(
|
||||
&mut self,
|
||||
source: &DemosaicedImage,
|
||||
shader: &ComposedShader,
|
||||
width: u32,
|
||||
height: u32,
|
||||
) -> Result<&wgpu::Texture, GpuError> {
|
||||
if shader.output_mode != OutputMode::CameraLinear {
|
||||
return Err(GpuError::ShaderCompilation(
|
||||
"render_camera_linear takes the shader from EditGraph::compose_camera_linear \
|
||||
and no other; this one writes a different format"
|
||||
.into(),
|
||||
));
|
||||
}
|
||||
self.colour_key = None;
|
||||
let (width, height) = (width.max(1), height.max(1));
|
||||
self.ensure_camera_target(width, height);
|
||||
|
||||
let mut uniforms = Self::fused_uniforms(source, shader);
|
||||
// Neutral profile: the numbers the sensor produced, and only those.
|
||||
let non_linear = uniforms[15];
|
||||
uniforms[0..4].copy_from_slice(&[1.0, 0.0, 0.0, 0.0]);
|
||||
uniforms[4..8].copy_from_slice(&[0.0, 1.0, 0.0, 0.0]);
|
||||
uniforms[8..12].copy_from_slice(&[0.0, 0.0, 1.0, 0.0]);
|
||||
uniforms[12..16].copy_from_slice(&[1.0, 1.0, 1.0, non_linear]);
|
||||
let b = dr_pipeline::BASE_CURVE_UNIFORM_OFFSET;
|
||||
uniforms[b + 10] = 0.0;
|
||||
|
||||
let params_buf = self
|
||||
.ctx
|
||||
.device
|
||||
.create_buffer_init(&wgpu::util::BufferInitDescriptor {
|
||||
label: Some("adjust-camera-params"),
|
||||
contents: bytemuck::cast_slice(&uniforms),
|
||||
usage: wgpu::BufferUsages::UNIFORM,
|
||||
});
|
||||
|
||||
let _ = self.pipeline(shader)?;
|
||||
let pipeline = self
|
||||
.cache
|
||||
.get(&shader.structure_hash)
|
||||
.expect("compiled above");
|
||||
let target = self.camera_target.as_ref().expect("ensured above");
|
||||
|
||||
let bind_group = self
|
||||
.ctx
|
||||
.device
|
||||
.create_bind_group(&wgpu::BindGroupDescriptor {
|
||||
label: Some("adjust-camera-bg"),
|
||||
layout: &self.camera_bind_group_layout,
|
||||
entries: &[
|
||||
wgpu::BindGroupEntry {
|
||||
binding: 0,
|
||||
resource: wgpu::BindingResource::TextureView(source.view()),
|
||||
},
|
||||
wgpu::BindGroupEntry {
|
||||
binding: 1,
|
||||
resource: params_buf.as_entire_binding(),
|
||||
},
|
||||
wgpu::BindGroupEntry {
|
||||
binding: 2,
|
||||
resource: wgpu::BindingResource::TextureView(&target.view),
|
||||
},
|
||||
wgpu::BindGroupEntry {
|
||||
binding: 3,
|
||||
resource: wgpu::BindingResource::TextureView(&self.empty_masks),
|
||||
},
|
||||
wgpu::BindGroupEntry {
|
||||
binding: 4,
|
||||
resource: wgpu::BindingResource::TextureView(self.film_curves_view()),
|
||||
},
|
||||
wgpu::BindGroupEntry {
|
||||
binding: 5,
|
||||
resource: wgpu::BindingResource::TextureView(self.film_lut_view()),
|
||||
},
|
||||
],
|
||||
});
|
||||
|
||||
let mut enc = self
|
||||
.ctx
|
||||
.device
|
||||
.create_command_encoder(&wgpu::CommandEncoderDescriptor {
|
||||
label: Some("adjust-camera-encoder"),
|
||||
});
|
||||
{
|
||||
let mut pass = enc.begin_compute_pass(&wgpu::ComputePassDescriptor {
|
||||
label: Some("adjust-camera-pass"),
|
||||
timestamp_writes: None,
|
||||
});
|
||||
pass.set_pipeline(pipeline);
|
||||
pass.set_bind_group(0, &bind_group, &[]);
|
||||
pass.dispatch_workgroups(width.div_ceil(8), height.div_ceil(8), 1);
|
||||
}
|
||||
self.ctx.queue.submit(Some(enc.finish()));
|
||||
self.colour_dispatches += 1;
|
||||
|
||||
Ok(&self.camera_target.as_ref().expect("ensured above").texture)
|
||||
}
|
||||
|
||||
/// The camera-space texture, if one has been rendered.
|
||||
pub fn camera_texture(&self) -> Option<&wgpu::Texture> {
|
||||
self.camera_target.as_ref().map(|t| &t.texture)
|
||||
}
|
||||
|
||||
/// TRACES: FR-MRG-2
|
||||
/// Read the camera-space tap back: tightly packed RGBA `f32`,
|
||||
/// `width * height * 4` values, alpha 1.0 everywhere.
|
||||
pub fn read_camera_linear(&self) -> Result<(Vec<f32>, u32, u32), GpuError> {
|
||||
let Some(target) = self.camera_target.as_ref() else {
|
||||
return Err(GpuError::Readback("no camera-space render yet".into()));
|
||||
};
|
||||
let (bytes, w, h) = Self::copy_texture(&self.ctx, &target.texture, w_h(target), 16)?;
|
||||
let floats: Vec<f32> = bytes
|
||||
.chunks_exact(4)
|
||||
.map(|b| f32::from_le_bytes([b[0], b[1], b[2], b[3]]))
|
||||
.collect();
|
||||
Ok((floats, w, h))
|
||||
}
|
||||
|
||||
fn ensure_camera_target(&mut self, width: u32, height: u32) {
|
||||
if self
|
||||
.camera_target
|
||||
.as_ref()
|
||||
.is_some_and(|t| t.width == width && t.height == height)
|
||||
{
|
||||
return;
|
||||
}
|
||||
let texture = self.ctx.device.create_texture(&wgpu::TextureDescriptor {
|
||||
label: Some("adjust-camera-output"),
|
||||
size: wgpu::Extent3d {
|
||||
width,
|
||||
height,
|
||||
depth_or_array_layers: 1,
|
||||
},
|
||||
mip_level_count: 1,
|
||||
sample_count: 1,
|
||||
dimension: wgpu::TextureDimension::D2,
|
||||
format: Self::CAMERA_FORMAT,
|
||||
// Written by compute, sampled by a merge's warp, copied out for
|
||||
// the CPU. Never handed to the compositor, so no RENDER_ATTACHMENT.
|
||||
usage: wgpu::TextureUsages::STORAGE_BINDING
|
||||
| wgpu::TextureUsages::TEXTURE_BINDING
|
||||
| wgpu::TextureUsages::COPY_SRC,
|
||||
view_formats: &[],
|
||||
});
|
||||
let view = texture.create_view(&Default::default());
|
||||
self.camera_target = Some(Target {
|
||||
texture,
|
||||
view,
|
||||
width,
|
||||
height,
|
||||
});
|
||||
}
|
||||
|
||||
/// The transfer itself.
|
||||
fn copy_output(&self) -> Result<(Vec<u8>, u32, u32), GpuError> {
|
||||
let Some(target) = self.targets[self.current].as_ref() else {
|
||||
return Err(GpuError::Readback("nothing rendered yet".into()));
|
||||
};
|
||||
let (w, h) = (target.width, target.height);
|
||||
Self::copy_texture(&self.ctx, &target.texture, w_h(target), 4)
|
||||
}
|
||||
|
||||
let unpadded = w * 4;
|
||||
/// Copy a whole texture to the CPU, `bytes_per_pixel` wide, rows
|
||||
/// unpadded. Shared by the display readback and the camera-space one.
|
||||
fn copy_texture(
|
||||
ctx: &GpuContext,
|
||||
texture: &wgpu::Texture,
|
||||
(w, h): (u32, u32),
|
||||
bytes_per_pixel: u32,
|
||||
) -> Result<(Vec<u8>, u32, u32), GpuError> {
|
||||
let unpadded = w * bytes_per_pixel;
|
||||
let align = wgpu::COPY_BYTES_PER_ROW_ALIGNMENT;
|
||||
let padded = unpadded.div_ceil(align) * align;
|
||||
|
||||
let buf = self.ctx.device.create_buffer(&wgpu::BufferDescriptor {
|
||||
let buf = ctx.device.create_buffer(&wgpu::BufferDescriptor {
|
||||
label: Some("adjust-readback"),
|
||||
size: (padded * h) as u64,
|
||||
usage: wgpu::BufferUsages::COPY_DST | wgpu::BufferUsages::MAP_READ,
|
||||
mapped_at_creation: false,
|
||||
});
|
||||
|
||||
let mut enc = self.ctx.device.create_command_encoder(&Default::default());
|
||||
let mut enc = ctx.device.create_command_encoder(&Default::default());
|
||||
enc.copy_texture_to_buffer(
|
||||
wgpu::TexelCopyTextureInfo {
|
||||
texture: &target.texture,
|
||||
texture,
|
||||
mip_level: 0,
|
||||
origin: wgpu::Origin3d::ZERO,
|
||||
aspect: wgpu::TextureAspect::All,
|
||||
@@ -1166,7 +1372,7 @@ impl AdjustPass {
|
||||
depth_or_array_layers: 1,
|
||||
},
|
||||
);
|
||||
self.ctx.queue.submit(Some(enc.finish()));
|
||||
ctx.queue.submit(Some(enc.finish()));
|
||||
|
||||
let slice = buf.slice(..);
|
||||
let (tx, rx) = std::sync::mpsc::channel();
|
||||
@@ -1177,7 +1383,7 @@ impl AdjustPass {
|
||||
// Polled rather than parked, and bounded rather than spun forever —
|
||||
// see `readback::await_mapping`, which the histogram's own transfer
|
||||
// shares for exactly the same reasons.
|
||||
await_mapping(&self.ctx, &rx)?;
|
||||
await_mapping(ctx, &rx)?;
|
||||
|
||||
let data = slice.get_mapped_range();
|
||||
let mut out = Vec::with_capacity((unpadded * h) as usize);
|
||||
@@ -1191,6 +1397,10 @@ impl AdjustPass {
|
||||
}
|
||||
}
|
||||
|
||||
fn w_h(t: &Target) -> (u32, u32) {
|
||||
(t.width, t.height)
|
||||
}
|
||||
|
||||
/// Number the lines of generated source, so a compiler error can be located.
|
||||
pub(crate) fn numbered(src: &str) -> String {
|
||||
src.lines()
|
||||
@@ -1242,6 +1452,10 @@ mod tests {
|
||||
// rather than about a camera's colour response.
|
||||
color_matrix: Some([1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0]),
|
||||
base_curve: BaseCurve::IDENTITY,
|
||||
samples_per_pixel: 1,
|
||||
profile: None,
|
||||
make: String::new(),
|
||||
model: String::new(),
|
||||
crop: CropRect {
|
||||
x: 0,
|
||||
y: 0,
|
||||
@@ -1442,6 +1656,10 @@ mod tests {
|
||||
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
|
||||
color_matrix: Some([1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0]),
|
||||
base_curve: BaseCurve::IDENTITY,
|
||||
samples_per_pixel: 1,
|
||||
profile: None,
|
||||
make: String::new(),
|
||||
model: String::new(),
|
||||
crop: CropRect {
|
||||
x: 0,
|
||||
y: 0,
|
||||
@@ -1969,6 +2187,10 @@ mod tests {
|
||||
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
|
||||
color_matrix: Some([1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0]),
|
||||
base_curve: BaseCurve::IDENTITY,
|
||||
samples_per_pixel: 1,
|
||||
profile: None,
|
||||
make: String::new(),
|
||||
model: String::new(),
|
||||
crop: CropRect {
|
||||
x: 0,
|
||||
y: 0,
|
||||
@@ -2069,6 +2291,10 @@ mod tests {
|
||||
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
|
||||
color_matrix: Some([1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0]),
|
||||
base_curve: BaseCurve::IDENTITY,
|
||||
samples_per_pixel: 1,
|
||||
profile: None,
|
||||
make: String::new(),
|
||||
model: String::new(),
|
||||
crop: CropRect {
|
||||
x: 0,
|
||||
y: 0,
|
||||
|
||||
@@ -250,6 +250,135 @@ impl DemosaicedImage {
|
||||
}
|
||||
}
|
||||
|
||||
impl DemosaicedImage {
|
||||
/// TRACES: FR-MRG-3
|
||||
/// A source that is already RGB in camera space: a linear DNG, which is
|
||||
/// what a merge writes. No demosaic; the samples are normalised by the
|
||||
/// file's black and white levels exactly as the demosaic kernel would
|
||||
/// normalise a photosite, and everything else — the matrix, the
|
||||
/// balance, the body's base curve — is carried through as for a CFA
|
||||
/// file, because the composite is developed as one photograph from the
|
||||
/// body that took its sources.
|
||||
pub fn from_linear_rgb16(ctx: &GpuContext, raw: &RawImage) -> Result<Self, GpuError> {
|
||||
let (width, height) = (raw.crop.width.max(1), raw.crop.height.max(1));
|
||||
let limits = ctx.device.limits();
|
||||
if width > limits.max_texture_dimension_2d || height > limits.max_texture_dimension_2d {
|
||||
return Err(GpuError::TooLarge(format!(
|
||||
"{width}×{height} exceeds the device limit of {}",
|
||||
limits.max_texture_dimension_2d
|
||||
)));
|
||||
}
|
||||
let stride = raw.width as usize * 3;
|
||||
let expected = raw.height as usize * stride;
|
||||
if raw.data.len() < expected {
|
||||
return Err(GpuError::TooLarge(format!(
|
||||
"{} samples is short of the {expected} a {}×{} RGB image needs",
|
||||
raw.data.len(),
|
||||
raw.width,
|
||||
raw.height
|
||||
)));
|
||||
}
|
||||
let black = black_per_cell(raw);
|
||||
let inv = inv_range_per_cell(raw);
|
||||
// Per channel rather than per CFA cell: R, G, B are the first three.
|
||||
let mut half: Vec<u16> = Vec::with_capacity((width * height * 4) as usize);
|
||||
for y in 0..height as usize {
|
||||
let row = (raw.crop.y as usize + y) * stride + raw.crop.x as usize * 3;
|
||||
for x in 0..width as usize {
|
||||
let p = &raw.data[row + x * 3..row + x * 3 + 3];
|
||||
for c in 0..3 {
|
||||
let v = (f32::from(p[c]) - black[c]) * inv[c];
|
||||
half.push(f32_to_f16_bits_unclamped(v));
|
||||
}
|
||||
half.push(f32_to_f16_bits(1.0));
|
||||
}
|
||||
}
|
||||
let texture = ctx.device.create_texture_with_data(
|
||||
&ctx.queue,
|
||||
&wgpu::TextureDescriptor {
|
||||
label: Some("linear-rgb-source"),
|
||||
size: wgpu::Extent3d {
|
||||
width,
|
||||
height,
|
||||
depth_or_array_layers: 1,
|
||||
},
|
||||
mip_level_count: 1,
|
||||
sample_count: 1,
|
||||
dimension: wgpu::TextureDimension::D2,
|
||||
format: Self::FORMAT,
|
||||
usage: wgpu::TextureUsages::TEXTURE_BINDING | wgpu::TextureUsages::COPY_SRC,
|
||||
view_formats: &[],
|
||||
},
|
||||
wgpu::util::TextureDataOrder::LayerMajor,
|
||||
bytemuck::cast_slice(&half),
|
||||
);
|
||||
let view = texture.create_view(&Default::default());
|
||||
Ok(Self {
|
||||
texture,
|
||||
view,
|
||||
width,
|
||||
height,
|
||||
color_matrix: raw.color_matrix.unwrap_or(IDENTITY_3X3),
|
||||
as_shot_wb: [raw.wb_coeffs[0], raw.wb_coeffs[1], raw.wb_coeffs[2]],
|
||||
base_curve: raw.base_curve,
|
||||
non_linear: false,
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
/// Convert an f32 to half-precision bits, the general case: sign,
|
||||
/// subnormals, round-to-nearest-even, saturation at the largest finite.
|
||||
///
|
||||
/// `f32_to_f16_bits` below is the 8-bit special case and says why it can
|
||||
/// be; this one exists because a linear DNG is not that case. A 14-bit
|
||||
/// sensor's least significant step, normalised, is 6.1e-5 — right at f16's
|
||||
/// smallest normal (6.1e-5) — so the deepest shadows of a composite land
|
||||
/// in the subnormal range, and rounding them to zero would crush the
|
||||
/// shadows of exactly the file that was written to keep them. Values below
|
||||
/// zero (black subtraction on a noisy photosite) and above one (a highlight
|
||||
/// past the white level) are legitimate and kept.
|
||||
fn f32_to_f16_bits_unclamped(v: f32) -> u16 {
|
||||
let bits = v.to_bits();
|
||||
let sign = ((bits >> 16) & 0x8000) as u16;
|
||||
let exp = ((bits >> 23) & 0xFF) as i32;
|
||||
let mant = bits & 0x7F_FFFF;
|
||||
if exp == 0xFF {
|
||||
// Infinity or NaN: a NaN sample is a decode fault; store the largest
|
||||
// finite rather than propagate it through a blend.
|
||||
return sign | 0x7BFF;
|
||||
}
|
||||
let e = exp - 127 + 15;
|
||||
if e >= 0x1F {
|
||||
return sign | 0x7BFF;
|
||||
}
|
||||
if e <= 0 {
|
||||
// Subnormal in f16 (or underflow). Shift the full mantissa with its
|
||||
// implicit bit right by the deficit, rounding to nearest even.
|
||||
if e < -10 {
|
||||
return sign;
|
||||
}
|
||||
let m = (mant | 0x80_0000) >> (1 - e);
|
||||
let shift = 13;
|
||||
let rounded = round_shift(m, shift);
|
||||
return sign | rounded as u16;
|
||||
}
|
||||
let rounded = round_shift(mant, 13);
|
||||
// Rounding can carry into the exponent; that is correct.
|
||||
sign | (((e as u32) << 10) + rounded) as u16
|
||||
}
|
||||
|
||||
/// `v >> shift`, rounded to nearest with ties to even.
|
||||
fn round_shift(v: u32, shift: u32) -> u32 {
|
||||
let half = 1u32 << (shift - 1);
|
||||
let mask = (1u32 << shift) - 1;
|
||||
let low = v & mask;
|
||||
let mut out = v >> shift;
|
||||
if low > half || (low == half && (out & 1) == 1) {
|
||||
out += 1;
|
||||
}
|
||||
out
|
||||
}
|
||||
|
||||
/// Convert an f32 to IEEE 754 half-precision bits.
|
||||
///
|
||||
/// Written out rather than pulled in as a dependency: the inputs here are
|
||||
@@ -389,6 +518,9 @@ impl Demosaicer {
|
||||
/// `RawImage`; which of the two CFA families it came off is this
|
||||
/// function's problem, not theirs.
|
||||
pub fn run(&self, raw: &RawImage) -> Result<DemosaicedImage, GpuError> {
|
||||
if raw.samples_per_pixel == 3 {
|
||||
return DemosaicedImage::from_linear_rgb16(&self.ctx, raw);
|
||||
}
|
||||
let (width, height) = (raw.crop.width.max(1), raw.crop.height.max(1));
|
||||
|
||||
let limits = self.ctx.device.limits();
|
||||
@@ -827,6 +959,43 @@ mod tests {
|
||||
use super::*;
|
||||
use dr_decode::CropRect;
|
||||
|
||||
fn f16_to_f32(bits: u16) -> f32 {
|
||||
let sign = if bits & 0x8000 != 0 { -1.0 } else { 1.0 };
|
||||
let e = ((bits >> 10) & 0x1F) as i32;
|
||||
let m = (bits & 0x3FF) as f32;
|
||||
if e == 0 {
|
||||
sign * m * 2f32.powi(-24)
|
||||
} else {
|
||||
sign * (1.0 + m / 1024.0) * 2f32.powi(e - 15)
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn unclamped_half_keeps_shadows_signs_and_highlights() {
|
||||
// A 14-bit LSB, normalised: subnormal in f16, and must not be zero.
|
||||
let lsb = 1.0 / 16383.0;
|
||||
let back = f16_to_f32(f32_to_f16_bits_unclamped(lsb));
|
||||
assert!((back - lsb).abs() / lsb < 0.01, "{back} vs {lsb}");
|
||||
// A quarter of that, still representable.
|
||||
let tiny = lsb / 4.0;
|
||||
let back = f16_to_f32(f32_to_f16_bits_unclamped(tiny));
|
||||
assert!((back - tiny).abs() / tiny < 0.05, "{back} vs {tiny}");
|
||||
// Below zero and above one survive.
|
||||
assert!((f16_to_f32(f32_to_f16_bits_unclamped(-0.01)) + 0.01).abs() < 1e-5);
|
||||
assert!((f16_to_f32(f32_to_f16_bits_unclamped(1.75)) - 1.75).abs() < 1e-3);
|
||||
// Exact values are exact.
|
||||
assert_eq!(f32_to_f16_bits_unclamped(1.0), 0x3C00);
|
||||
assert_eq!(f32_to_f16_bits_unclamped(0.5), 0x3800);
|
||||
assert_eq!(f32_to_f16_bits_unclamped(0.0), 0);
|
||||
// Within one ULP of the clamped one on its domain: that one
|
||||
// truncates the mantissa, this one rounds it.
|
||||
for i in 0..=255 {
|
||||
let v = i as f32 / 255.0;
|
||||
let (a, b) = (f32_to_f16_bits_unclamped(v), f32_to_f16_bits(v));
|
||||
assert!(a.abs_diff(b) <= 1, "{v}: {a} vs {b}");
|
||||
}
|
||||
}
|
||||
|
||||
fn raw_for(black: [u16; 4], white: u16) -> RawImage {
|
||||
RawImage {
|
||||
width: 4,
|
||||
@@ -838,6 +1007,10 @@ mod tests {
|
||||
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
|
||||
color_matrix: None,
|
||||
base_curve: BaseCurve::IDENTITY,
|
||||
samples_per_pixel: 1,
|
||||
profile: None,
|
||||
make: String::new(),
|
||||
model: String::new(),
|
||||
crop: CropRect {
|
||||
x: 0,
|
||||
y: 0,
|
||||
@@ -950,6 +1123,10 @@ mod tests {
|
||||
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
|
||||
color_matrix: None,
|
||||
base_curve: BaseCurve::IDENTITY,
|
||||
samples_per_pixel: 1,
|
||||
profile: None,
|
||||
make: String::new(),
|
||||
model: String::new(),
|
||||
crop: CropRect {
|
||||
x: 0,
|
||||
y: 0,
|
||||
@@ -1170,6 +1347,53 @@ mod tests {
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_grey_step_edge_stays_grey() {
|
||||
// A flat patch cannot tell the Malvar kernels from any other set of
|
||||
// weights that sum to zero. An edge can. A grey vertical step, so
|
||||
// every photosite records the same profile, must come back with the
|
||||
// three channels close together on both sides; any spread is false
|
||||
// colour from interpolating across the edge.
|
||||
//
|
||||
// The bound is set by the paper's kernels, which peak at 0.19 here.
|
||||
// With the ±2 terms of the green-site kernels transposed — the bug
|
||||
// this test was written against — the peak is 0.375.
|
||||
let Some(ctx) = ctx() else { return };
|
||||
let d = Demosaicer::new(&ctx).expect("demosaicer");
|
||||
|
||||
let size = 32u32;
|
||||
let white = 16383u16;
|
||||
let mut raw = flat_cfa(CfaPattern::Rggb, size, [0, 0, 0], 0, white);
|
||||
for y in 0..size {
|
||||
for x in size / 2..size {
|
||||
raw.data[(y * size + x) as usize] = white;
|
||||
}
|
||||
}
|
||||
|
||||
let img = d.run(&raw).expect("demosaic");
|
||||
let px = read_rgba(&ctx, &img);
|
||||
let (w, _) = img.size();
|
||||
|
||||
let mut worst = (0.0f32, 0u32, 0u32);
|
||||
for y in 2..size - 2 {
|
||||
for x in 2..size - 2 {
|
||||
let p = px[(y * w + x) as usize];
|
||||
let spread = (p[0] - p[1]).abs().max((p[2] - p[1]).abs());
|
||||
if spread > worst.0 {
|
||||
worst = (spread, x, y);
|
||||
}
|
||||
}
|
||||
}
|
||||
assert!(
|
||||
worst.0 < 0.25,
|
||||
"false colour of {} at ({}, {}) on a grey edge — the green-site \
|
||||
kernels are interpolating across the edge",
|
||||
worst.0,
|
||||
worst.1,
|
||||
worst.2
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn output_is_free_of_nan_and_negatives() {
|
||||
// f16 NaN propagates silently through every later stage; a negative
|
||||
@@ -1196,6 +1420,10 @@ mod tests {
|
||||
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
|
||||
color_matrix: None,
|
||||
base_curve: BaseCurve::IDENTITY,
|
||||
samples_per_pixel: 1,
|
||||
profile: None,
|
||||
make: String::new(),
|
||||
model: String::new(),
|
||||
crop: CropRect {
|
||||
x: 0,
|
||||
y: 0,
|
||||
@@ -1277,6 +1505,10 @@ mod tests {
|
||||
],
|
||||
color_matrix: None,
|
||||
base_curve: BaseCurve::IDENTITY,
|
||||
samples_per_pixel: 1,
|
||||
profile: None,
|
||||
make: String::new(),
|
||||
model: String::new(),
|
||||
crop: CropRect {
|
||||
x: 0,
|
||||
y: 0,
|
||||
|
||||
@@ -28,6 +28,7 @@ mod error;
|
||||
mod focus;
|
||||
mod histogram;
|
||||
mod mask;
|
||||
mod merge;
|
||||
mod raw_histogram;
|
||||
mod readback;
|
||||
mod segment;
|
||||
@@ -40,6 +41,7 @@ pub use demosaic::{DemosaicedImage, Demosaicer};
|
||||
pub use detail::INTERMEDIATE_FORMAT as DETAIL_INTERMEDIATE_FORMAT;
|
||||
pub use error::GpuError;
|
||||
pub use focus::{FocusPeakPass, FocusPeaking, PeakColour, PeakSensitivity};
|
||||
pub use merge::{Band, MergeFrame, MergeOutput, MergePass};
|
||||
// Renamed on the way out: `BINS` says enough inside `histogram`, and nothing
|
||||
// at all at a crate root shared with demosaic and segmentation.
|
||||
pub use histogram::{Histogram, HistogramPass, BINS as HISTOGRAM_BINS};
|
||||
@@ -281,6 +283,7 @@ impl GpuContext {
|
||||
)))
|
||||
}
|
||||
|
||||
/// TRACES: NFR-COMPAT-1
|
||||
/// Ask one adapter for a device, with the limits the pipeline needs.
|
||||
async fn device_from(adapter: &wgpu::Adapter) -> Result<(wgpu::Device, wgpu::Queue), GpuError> {
|
||||
adapter
|
||||
|
||||
@@ -0,0 +1,547 @@
|
||||
//! TRACES: FR-MRG-10 | FR-MRG-11
|
||||
//! The merge: source frames warped into an output surface, chunk by chunk.
|
||||
//!
|
||||
//! The per-pixel half of a panorama (FR-MRG-10), on the GPU: the warp of a
|
||||
//! source tile into an output chunk, the weighted accumulation across
|
||||
//! frames, and the resolve to sixteen-bit samples. The geometry it is
|
||||
//! given — rotations, focal length, projection — is `dr-pano`'s, solved on
|
||||
//! proxies before any full-resolution pixel exists (panorama.md §5), and
|
||||
//! that is what makes this simple: every output pixel's source coordinates
|
||||
//! are a closed-form function, so a chunk can be produced from the source
|
||||
//! tiles that project into it and nothing else.
|
||||
//!
|
||||
//! # The loop
|
||||
//!
|
||||
//! ```text
|
||||
//! for each band of rows of the output:
|
||||
//! for each chunk across the band:
|
||||
//! zero the accumulator
|
||||
//! for each frame whose footprint meets the chunk:
|
||||
//! the source rectangle the chunk needs, from the geometry
|
||||
//! render it camera-linear through the pipeline (the tile)
|
||||
//! warp the tile into the chunk, accumulate ← GPU
|
||||
//! resolve the chunk to u16 ← GPU
|
||||
//! copy it into the band
|
||||
//! hand the band to the writer (one DNG strip)
|
||||
//! ```
|
||||
//!
|
||||
//! No stage holds the composite (FR-MRG-11): the working set is one
|
||||
//! chunk's accumulator, one tile, one band of u16 rows. The frame textures
|
||||
//! are the caller's to provide and cache — `source` is asked for frame `k`
|
||||
//! as it is needed, and a caller short of memory may demosaic on demand.
|
||||
//!
|
||||
//! # What is not here yet
|
||||
//!
|
||||
//! A feathered blend, not seams and a Laplacian pyramid: the weight is the
|
||||
//! distance to the frame's edge, which hides exposure steps and small
|
||||
//! misalignments and does not hide parallax. Gain is a scalar per frame
|
||||
//! the caller supplies. Both are panorama.md §10's step 5, after the path
|
||||
//! writes a file end to end.
|
||||
|
||||
use std::sync::Arc;
|
||||
|
||||
use dr_pano::bundle::Cameras;
|
||||
use dr_pano::projection::{Bounds, Projection};
|
||||
use wgpu::util::DeviceExt;
|
||||
|
||||
use crate::readback::await_mapping;
|
||||
use crate::{AdjustPass, DemosaicedImage, GpuContext, GpuError};
|
||||
|
||||
/// One frame's part in the merge.
|
||||
pub struct MergeFrame {
|
||||
/// The frame's edit, for its lens corrections — the only part of an
|
||||
/// edit the camera-space tap uses (FR-MRG-2).
|
||||
pub graph: Arc<dr_pipeline::EditGraph>,
|
||||
/// Multiplies the frame's samples, to bring its exposure to the
|
||||
/// reference frame's. 1.0 for no correction.
|
||||
pub gain: f32,
|
||||
}
|
||||
|
||||
/// The output the merge produces.
|
||||
#[derive(Debug, Clone, Copy, PartialEq)]
|
||||
pub struct MergeOutput {
|
||||
pub projection: Projection,
|
||||
/// The projection's scale in output pixels: the cylinder's radius, the
|
||||
/// plane's distance. The source focal length at full resolution gives
|
||||
/// output pixels the size of source pixels at the centre.
|
||||
pub scale: f64,
|
||||
/// The rectangle of the projection to produce, centred coordinates.
|
||||
pub bounds: Bounds,
|
||||
/// Pixels over which a frame's weight ramps up from its edge.
|
||||
pub feather: f32,
|
||||
/// Chunk size: the unit of GPU work and of memory.
|
||||
pub chunk: (u32, u32),
|
||||
/// Multiplies a normalised sample (1.0 = white) to the sensor's scale.
|
||||
pub sample_scale: f32,
|
||||
}
|
||||
|
||||
impl MergeOutput {
|
||||
pub fn width(&self) -> u32 {
|
||||
self.bounds.width().ceil().max(1.0) as u32
|
||||
}
|
||||
pub fn height(&self) -> u32 {
|
||||
self.bounds.height().ceil().max(1.0) as u32
|
||||
}
|
||||
}
|
||||
|
||||
/// A band of finished rows: `rows × width × 3` RGB `u16`, plus a coverage
|
||||
/// mask (`true` where any frame reached the pixel).
|
||||
pub struct Band<'a> {
|
||||
pub first_row: u32,
|
||||
pub rows: u32,
|
||||
pub rgb: &'a [u16],
|
||||
pub covered: &'a [bool],
|
||||
}
|
||||
|
||||
#[repr(C)]
|
||||
#[derive(Clone, Copy, bytemuck::Pod, bytemuck::Zeroable)]
|
||||
struct WarpParams {
|
||||
chunk_origin: [f32; 2],
|
||||
chunk_size: [u32; 2],
|
||||
projection: u32,
|
||||
proj_scale: f32,
|
||||
focal: f32,
|
||||
gain: f32,
|
||||
r0: [f32; 4],
|
||||
r1: [f32; 4],
|
||||
r2: [f32; 4],
|
||||
frame_size: [f32; 2],
|
||||
tile_origin: [f32; 2],
|
||||
tile_size: [u32; 2],
|
||||
feather: f32,
|
||||
_pad: f32,
|
||||
}
|
||||
|
||||
#[repr(C)]
|
||||
#[derive(Clone, Copy, bytemuck::Pod, bytemuck::Zeroable)]
|
||||
struct ResolveParams {
|
||||
chunk_size: [u32; 2],
|
||||
scale: f32,
|
||||
_pad: f32,
|
||||
}
|
||||
|
||||
/// The two pipelines and the chunk buffers.
|
||||
pub struct MergePass {
|
||||
ctx: GpuContext,
|
||||
warp: wgpu::ComputePipeline,
|
||||
warp_layout: wgpu::BindGroupLayout,
|
||||
resolve: wgpu::ComputePipeline,
|
||||
resolve_layout: wgpu::BindGroupLayout,
|
||||
/// Accumulator and packed output for the current chunk size.
|
||||
buffers: Option<(wgpu::Buffer, wgpu::Buffer, wgpu::Buffer, (u32, u32))>,
|
||||
}
|
||||
|
||||
impl MergePass {
|
||||
pub fn new(ctx: &GpuContext) -> Result<Self, GpuError> {
|
||||
let module = ctx
|
||||
.device
|
||||
.create_shader_module(wgpu::ShaderModuleDescriptor {
|
||||
label: Some("merge"),
|
||||
source: wgpu::ShaderSource::Wgsl(include_str!("shaders/merge.wgsl").into()),
|
||||
});
|
||||
let uniform = |binding| wgpu::BindGroupLayoutEntry {
|
||||
binding,
|
||||
visibility: wgpu::ShaderStages::COMPUTE,
|
||||
ty: wgpu::BindingType::Buffer {
|
||||
ty: wgpu::BufferBindingType::Uniform,
|
||||
has_dynamic_offset: false,
|
||||
min_binding_size: None,
|
||||
},
|
||||
count: None,
|
||||
};
|
||||
let storage = |binding, read_only| wgpu::BindGroupLayoutEntry {
|
||||
binding,
|
||||
visibility: wgpu::ShaderStages::COMPUTE,
|
||||
ty: wgpu::BindingType::Buffer {
|
||||
ty: wgpu::BufferBindingType::Storage { read_only },
|
||||
has_dynamic_offset: false,
|
||||
min_binding_size: None,
|
||||
},
|
||||
count: None,
|
||||
};
|
||||
let warp_layout = ctx
|
||||
.device
|
||||
.create_bind_group_layout(&wgpu::BindGroupLayoutDescriptor {
|
||||
label: Some("merge-warp-bgl"),
|
||||
entries: &[
|
||||
uniform(0),
|
||||
wgpu::BindGroupLayoutEntry {
|
||||
binding: 1,
|
||||
visibility: wgpu::ShaderStages::COMPUTE,
|
||||
ty: wgpu::BindingType::Texture {
|
||||
// Unfilterable: rgba32float, loaded by hand.
|
||||
sample_type: wgpu::TextureSampleType::Float { filterable: false },
|
||||
view_dimension: wgpu::TextureViewDimension::D2,
|
||||
multisampled: false,
|
||||
},
|
||||
count: None,
|
||||
},
|
||||
storage(2, false),
|
||||
],
|
||||
});
|
||||
let resolve_layout =
|
||||
ctx.device
|
||||
.create_bind_group_layout(&wgpu::BindGroupLayoutDescriptor {
|
||||
label: Some("merge-resolve-bgl"),
|
||||
entries: &[uniform(0), storage(1, true), storage(2, false)],
|
||||
});
|
||||
let pipeline = |name: &str, layout: &wgpu::BindGroupLayout| {
|
||||
let pl = ctx
|
||||
.device
|
||||
.create_pipeline_layout(&wgpu::PipelineLayoutDescriptor {
|
||||
label: Some(name),
|
||||
bind_group_layouts: &[Some(layout)],
|
||||
immediate_size: 0,
|
||||
});
|
||||
ctx.device
|
||||
.create_compute_pipeline(&wgpu::ComputePipelineDescriptor {
|
||||
label: Some(name),
|
||||
layout: Some(&pl),
|
||||
module: &module,
|
||||
entry_point: Some(name),
|
||||
compilation_options: Default::default(),
|
||||
cache: None,
|
||||
})
|
||||
};
|
||||
Ok(MergePass {
|
||||
ctx: ctx.clone(),
|
||||
warp: pipeline("warp", &warp_layout),
|
||||
warp_layout,
|
||||
resolve: pipeline("resolve", &resolve_layout),
|
||||
resolve_layout,
|
||||
buffers: None,
|
||||
})
|
||||
}
|
||||
|
||||
/// Allocate the chunk buffers for this size if the last ones differ.
|
||||
fn ensure_buffers(&mut self, chunk: (u32, u32)) {
|
||||
if self.buffers.as_ref().is_none_or(|b| b.3 != chunk) {
|
||||
let n = u64::from(chunk.0) * u64::from(chunk.1);
|
||||
let acc = self.ctx.device.create_buffer(&wgpu::BufferDescriptor {
|
||||
label: Some("merge-acc"),
|
||||
size: n * 16,
|
||||
usage: wgpu::BufferUsages::STORAGE | wgpu::BufferUsages::COPY_DST,
|
||||
mapped_at_creation: false,
|
||||
});
|
||||
let out = self.ctx.device.create_buffer(&wgpu::BufferDescriptor {
|
||||
label: Some("merge-out"),
|
||||
size: n * 8,
|
||||
usage: wgpu::BufferUsages::STORAGE | wgpu::BufferUsages::COPY_SRC,
|
||||
mapped_at_creation: false,
|
||||
});
|
||||
let read = self.ctx.device.create_buffer(&wgpu::BufferDescriptor {
|
||||
label: Some("merge-read"),
|
||||
size: n * 8,
|
||||
usage: wgpu::BufferUsages::COPY_DST | wgpu::BufferUsages::MAP_READ,
|
||||
mapped_at_creation: false,
|
||||
});
|
||||
self.buffers = Some((acc, out, read, chunk));
|
||||
}
|
||||
}
|
||||
|
||||
fn chunk_buffers(&self) -> (&wgpu::Buffer, &wgpu::Buffer, &wgpu::Buffer) {
|
||||
let b = self.buffers.as_ref().expect("ensured by the caller");
|
||||
(&b.0, &b.1, &b.2)
|
||||
}
|
||||
|
||||
/// Produce the whole output, band by band, handing each finished band
|
||||
/// to `sink`.
|
||||
///
|
||||
/// `cameras` are in **full-resolution source pixels** (`frame_size`),
|
||||
/// with frame `k` corresponding to `frames[k]` and `source(k)`. `source`
|
||||
/// supplies the demosaiced frame on demand and may cache as it sees fit.
|
||||
#[allow(clippy::too_many_arguments)]
|
||||
pub fn merge<S, F>(
|
||||
&mut self,
|
||||
adjust: &mut AdjustPass,
|
||||
frames: &[MergeFrame],
|
||||
cameras: &Cameras,
|
||||
frame_size: (u32, u32),
|
||||
output: &MergeOutput,
|
||||
mut source: S,
|
||||
mut sink: F,
|
||||
mut cancelled: impl FnMut() -> bool,
|
||||
) -> Result<(), GpuError>
|
||||
where
|
||||
S: FnMut(usize) -> Result<Arc<DemosaicedImage>, GpuError>,
|
||||
F: FnMut(Band<'_>) -> Result<(), GpuError>,
|
||||
{
|
||||
let (out_w, out_h) = (output.width(), output.height());
|
||||
let (cw, ch) = (output.chunk.0.max(8), output.chunk.1.max(8));
|
||||
let (fw, fh) = (frame_size.0 as f64, frame_size.1 as f64);
|
||||
|
||||
let mut band_rgb = vec![0u16; (out_w * ch * 3) as usize];
|
||||
let mut band_cov = vec![false; (out_w * ch) as usize];
|
||||
let mut chunk_px: Vec<u32> = Vec::new();
|
||||
|
||||
let mut y = 0u32;
|
||||
while y < out_h {
|
||||
let rows = ch.min(out_h - y);
|
||||
band_rgb.iter_mut().for_each(|v| *v = 0);
|
||||
band_cov.iter_mut().for_each(|v| *v = false);
|
||||
|
||||
let mut x = 0u32;
|
||||
while x < out_w {
|
||||
if cancelled() {
|
||||
return Err(GpuError::Readback("merge cancelled".into()));
|
||||
}
|
||||
let cols = cw.min(out_w - x);
|
||||
let origin = (
|
||||
output.bounds.min_u + f64::from(x),
|
||||
output.bounds.min_v + f64::from(y),
|
||||
);
|
||||
self.zero_accumulator((cols, rows));
|
||||
|
||||
for (k, frame) in frames.iter().enumerate() {
|
||||
let Some(rect) = source_rect(
|
||||
output.projection,
|
||||
output.scale,
|
||||
cameras,
|
||||
k,
|
||||
origin,
|
||||
(cols, rows),
|
||||
(fw, fh),
|
||||
) else {
|
||||
continue;
|
||||
};
|
||||
let image = source(k)?;
|
||||
// The tile: that rectangle of the frame, camera-linear,
|
||||
// at 1:1.
|
||||
let view = dr_pipeline::CropRect {
|
||||
x: (rect.0 as f32) / fw as f32,
|
||||
y: (rect.1 as f32) / fh as f32,
|
||||
width: (rect.2 as f32) / fw as f32,
|
||||
height: (rect.3 as f32) / fh as f32,
|
||||
};
|
||||
let shader = frame.graph.compose_camera_linear(view);
|
||||
let tile = adjust.render_camera_linear(&image, &shader, rect.2, rect.3)?;
|
||||
let r = cameras.rotations[k].transpose();
|
||||
let params = WarpParams {
|
||||
chunk_origin: [origin.0 as f32, origin.1 as f32],
|
||||
chunk_size: [cols, rows],
|
||||
projection: match output.projection {
|
||||
Projection::Perspective => 0,
|
||||
Projection::Cylindrical => 1,
|
||||
Projection::Spherical => 2,
|
||||
},
|
||||
proj_scale: output.scale as f32,
|
||||
focal: cameras.focal as f32,
|
||||
gain: frame.gain,
|
||||
r0: [r.0[0][0] as f32, r.0[0][1] as f32, r.0[0][2] as f32, 0.0],
|
||||
r1: [r.0[1][0] as f32, r.0[1][1] as f32, r.0[1][2] as f32, 0.0],
|
||||
r2: [r.0[2][0] as f32, r.0[2][1] as f32, r.0[2][2] as f32, 0.0],
|
||||
frame_size: [fw as f32, fh as f32],
|
||||
tile_origin: [rect.0 as f32, rect.1 as f32],
|
||||
tile_size: [rect.2, rect.3],
|
||||
feather: output.feather,
|
||||
_pad: 0.0,
|
||||
};
|
||||
self.accumulate(¶ms, tile);
|
||||
}
|
||||
|
||||
self.resolve_chunk((cols, rows), output.sample_scale, &mut chunk_px)?;
|
||||
// Into the band.
|
||||
for row in 0..rows as usize {
|
||||
for col in 0..cols as usize {
|
||||
let px = chunk_px[(row * cols as usize + col) * 2..][..2].to_vec();
|
||||
let i = row * out_w as usize + (x as usize + col);
|
||||
band_rgb[i * 3] = (px[0] & 0xFFFF) as u16;
|
||||
band_rgb[i * 3 + 1] = (px[0] >> 16) as u16;
|
||||
band_rgb[i * 3 + 2] = (px[1] & 0xFFFF) as u16;
|
||||
band_cov[i] = (px[1] >> 16) != 0;
|
||||
}
|
||||
}
|
||||
x += cols;
|
||||
}
|
||||
|
||||
sink(Band {
|
||||
first_row: y,
|
||||
rows,
|
||||
rgb: &band_rgb[..(out_w * rows * 3) as usize],
|
||||
covered: &band_cov[..(out_w * rows) as usize],
|
||||
})?;
|
||||
y += rows;
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
fn zero_accumulator(&mut self, chunk: (u32, u32)) {
|
||||
self.ensure_buffers(chunk);
|
||||
let (acc, _, _) = self.chunk_buffers();
|
||||
let n = u64::from(chunk.0) * u64::from(chunk.1) * 16;
|
||||
let mut enc = self.ctx.device.create_command_encoder(&Default::default());
|
||||
enc.clear_buffer(acc, 0, Some(n));
|
||||
self.ctx.queue.submit(Some(enc.finish()));
|
||||
}
|
||||
|
||||
fn accumulate(&mut self, params: &WarpParams, tile: &wgpu::Texture) {
|
||||
let chunk = (params.chunk_size[0], params.chunk_size[1]);
|
||||
let uniforms = self
|
||||
.ctx
|
||||
.device
|
||||
.create_buffer_init(&wgpu::util::BufferInitDescriptor {
|
||||
label: Some("merge-warp-params"),
|
||||
contents: bytemuck::bytes_of(params),
|
||||
usage: wgpu::BufferUsages::UNIFORM,
|
||||
});
|
||||
let view = tile.create_view(&Default::default());
|
||||
self.ensure_buffers(chunk);
|
||||
let (acc, _, _) = self.chunk_buffers();
|
||||
let bind = self
|
||||
.ctx
|
||||
.device
|
||||
.create_bind_group(&wgpu::BindGroupDescriptor {
|
||||
label: Some("merge-warp-bg"),
|
||||
layout: &self.warp_layout,
|
||||
entries: &[
|
||||
wgpu::BindGroupEntry {
|
||||
binding: 0,
|
||||
resource: uniforms.as_entire_binding(),
|
||||
},
|
||||
wgpu::BindGroupEntry {
|
||||
binding: 1,
|
||||
resource: wgpu::BindingResource::TextureView(&view),
|
||||
},
|
||||
wgpu::BindGroupEntry {
|
||||
binding: 2,
|
||||
resource: acc.as_entire_binding(),
|
||||
},
|
||||
],
|
||||
});
|
||||
let mut enc = self.ctx.device.create_command_encoder(&Default::default());
|
||||
{
|
||||
let mut pass = enc.begin_compute_pass(&Default::default());
|
||||
pass.set_pipeline(&self.warp);
|
||||
pass.set_bind_group(0, &bind, &[]);
|
||||
pass.dispatch_workgroups(chunk.0.div_ceil(8), chunk.1.div_ceil(8), 1);
|
||||
}
|
||||
self.ctx.queue.submit(Some(enc.finish()));
|
||||
}
|
||||
|
||||
fn resolve_chunk(
|
||||
&mut self,
|
||||
chunk: (u32, u32),
|
||||
scale: f32,
|
||||
out: &mut Vec<u32>,
|
||||
) -> Result<(), GpuError> {
|
||||
let params = ResolveParams {
|
||||
chunk_size: [chunk.0, chunk.1],
|
||||
scale,
|
||||
_pad: 0.0,
|
||||
};
|
||||
let uniforms = self
|
||||
.ctx
|
||||
.device
|
||||
.create_buffer_init(&wgpu::util::BufferInitDescriptor {
|
||||
label: Some("merge-resolve-params"),
|
||||
contents: bytemuck::bytes_of(¶ms),
|
||||
usage: wgpu::BufferUsages::UNIFORM,
|
||||
});
|
||||
let n = u64::from(chunk.0) * u64::from(chunk.1);
|
||||
self.ensure_buffers(chunk);
|
||||
let (acc, packed, read) = self.chunk_buffers();
|
||||
let bind = self
|
||||
.ctx
|
||||
.device
|
||||
.create_bind_group(&wgpu::BindGroupDescriptor {
|
||||
label: Some("merge-resolve-bg"),
|
||||
layout: &self.resolve_layout,
|
||||
entries: &[
|
||||
wgpu::BindGroupEntry {
|
||||
binding: 0,
|
||||
resource: uniforms.as_entire_binding(),
|
||||
},
|
||||
wgpu::BindGroupEntry {
|
||||
binding: 1,
|
||||
resource: acc.as_entire_binding(),
|
||||
},
|
||||
wgpu::BindGroupEntry {
|
||||
binding: 2,
|
||||
resource: packed.as_entire_binding(),
|
||||
},
|
||||
],
|
||||
});
|
||||
let mut enc = self.ctx.device.create_command_encoder(&Default::default());
|
||||
{
|
||||
let mut pass = enc.begin_compute_pass(&Default::default());
|
||||
pass.set_pipeline(&self.resolve);
|
||||
pass.set_bind_group(0, &bind, &[]);
|
||||
pass.dispatch_workgroups(chunk.0.div_ceil(8), chunk.1.div_ceil(8), 1);
|
||||
}
|
||||
enc.copy_buffer_to_buffer(packed, 0, read, 0, n * 8);
|
||||
self.ctx.queue.submit(Some(enc.finish()));
|
||||
|
||||
let slice = read.slice(..n * 8);
|
||||
let (tx, rx) = std::sync::mpsc::channel();
|
||||
slice.map_async(wgpu::MapMode::Read, move |r| {
|
||||
let _ = tx.send(r);
|
||||
});
|
||||
await_mapping(&self.ctx, &rx)?;
|
||||
{
|
||||
let data = slice.get_mapped_range();
|
||||
out.clear();
|
||||
out.extend_from_slice(bytemuck::cast_slice::<u8, u32>(&data));
|
||||
}
|
||||
read.unmap();
|
||||
Ok(())
|
||||
}
|
||||
}
|
||||
|
||||
/// The rectangle of frame `k` (x, y, w, h in source pixels) a chunk reads,
|
||||
/// or `None` if the chunk sees nothing of the frame.
|
||||
///
|
||||
/// Walks the chunk's border, projects each point into the frame, and takes
|
||||
/// the bounding box with a two-pixel margin for the bilinear fetch. The
|
||||
/// border rather than the corners because under a cylinder or sphere the
|
||||
/// extreme of a footprint is not at a corner.
|
||||
fn source_rect(
|
||||
projection: Projection,
|
||||
scale: f64,
|
||||
cameras: &Cameras,
|
||||
k: usize,
|
||||
origin: (f64, f64),
|
||||
size: (u32, u32),
|
||||
frame: (f64, f64),
|
||||
) -> Option<(u32, u32, u32, u32)> {
|
||||
let (w, h) = (f64::from(size.0), f64::from(size.1));
|
||||
let steps = 16;
|
||||
let mut min = (f64::MAX, f64::MAX);
|
||||
let mut max = (f64::MIN, f64::MIN);
|
||||
let mut any = false;
|
||||
let mut visit = |u: f64, v: f64| {
|
||||
let d = projection.to_direction(scale, u, v);
|
||||
if let Some((x, y)) = cameras.project(k, d) {
|
||||
let (x, y) = (x + frame.0 / 2.0, y + frame.1 / 2.0);
|
||||
min = (min.0.min(x), min.1.min(y));
|
||||
max = (max.0.max(x), max.1.max(y));
|
||||
any = true;
|
||||
}
|
||||
};
|
||||
for s in 0..=steps {
|
||||
let t = f64::from(s) / f64::from(steps);
|
||||
visit(origin.0 + w * t, origin.1);
|
||||
visit(origin.0 + w * t, origin.1 + h);
|
||||
visit(origin.0, origin.1 + h * t);
|
||||
visit(origin.0 + w, origin.1 + h * t);
|
||||
}
|
||||
// The interior too, coarsely: a chunk can contain a frame entirely.
|
||||
for i in 1..4 {
|
||||
for j in 1..4 {
|
||||
visit(
|
||||
origin.0 + w * f64::from(i) / 4.0,
|
||||
origin.1 + h * f64::from(j) / 4.0,
|
||||
);
|
||||
}
|
||||
}
|
||||
if !any {
|
||||
return None;
|
||||
}
|
||||
let x0 = (min.0.floor() - 2.0).max(0.0);
|
||||
let y0 = (min.1.floor() - 2.0).max(0.0);
|
||||
let x1 = (max.0.ceil() + 2.0).min(frame.0);
|
||||
let y1 = (max.1.ceil() + 2.0).min(frame.1);
|
||||
if x1 <= x0 || y1 <= y0 {
|
||||
return None;
|
||||
}
|
||||
Some((x0 as u32, y0 as u32, (x1 - x0) as u32, (y1 - y0) as u32))
|
||||
}
|
||||
@@ -160,12 +160,18 @@ fn main(@builtin(global_invocation_id) gid: vec3<u32>) {
|
||||
// Green is measured. Red and blue are interpolated from their own
|
||||
// axis, with a correction from the green Laplacian.
|
||||
//
|
||||
// Malvar "G at R/B locations" kernels, transposed per axis:
|
||||
// chroma along the row: (5c + 4(w1+e1) - (nw+ne+sw+se) - (n2+s2) + 0.5(w2+e2)) / 8
|
||||
// Malvar "R at green in R row" kernel, and its transpose:
|
||||
// chroma along the row: (5c + 4(w1+e1) - (nw+ne+sw+se) - (w2+e2) + 0.5(n2+s2)) / 8
|
||||
//
|
||||
// The -1 goes on the two greens *along* the chroma axis and the +0.5
|
||||
// on the pair across it. Transposed, both kernels still sum to zero
|
||||
// and reconstruct a flat patch exactly, but on an edge the correction
|
||||
// at green sites is half strength and the false colour doubles: a
|
||||
// blue/yellow zipper around every clipped highlight.
|
||||
let along_row =
|
||||
(5.0 * c + 4.0 * (w1 + e1) - diag1 - vert2 + 0.5 * horiz2) * 0.125;
|
||||
(5.0 * c + 4.0 * (w1 + e1) - diag1 - horiz2 + 0.5 * vert2) * 0.125;
|
||||
let along_col =
|
||||
(5.0 * c + 4.0 * (n1 + s1) - diag1 - horiz2 + 0.5 * vert2) * 0.125;
|
||||
(5.0 * c + 4.0 * (n1 + s1) - diag1 - vert2 + 0.5 * horiz2) * 0.125;
|
||||
|
||||
let red_horizontal = red_is_horizontal(gid.x, gid.y);
|
||||
let r = select(along_col, along_row, red_horizontal);
|
||||
|
||||
@@ -0,0 +1,166 @@
|
||||
// TRACES: FR-MRG-10 | FR-MRG-11
|
||||
// The merge: one source tile warped into one output chunk, accumulated.
|
||||
//
|
||||
// Two entry points. `warp` runs once per (chunk, frame): for every chunk
|
||||
// pixel it asks which direction that pixel looks along, turns the
|
||||
// direction into the frame's camera, projects it to a source pixel, and
|
||||
// if that pixel is inside the tile that was rendered for this chunk,
|
||||
// samples it and adds it — weighted by its distance from the frame's edge
|
||||
// — into the accumulator. `resolve` runs once per chunk after every frame
|
||||
// has been added: divides the sums by the weights and packs the result as
|
||||
// sixteen-bit samples at the sensor's scale (FR-MRG-3).
|
||||
//
|
||||
// The accumulator is a buffer and not a storage texture, because WebGPU
|
||||
// allows a read-write storage texture only in the 32-bit single-channel
|
||||
// formats, and this wants four channels. The tile is sampled by hand from
|
||||
// four `textureLoad`s rather than through a sampler, because `rgba32float`
|
||||
// is not filterable without an optional feature, and the tile is
|
||||
// `rgba32float` on purpose (panorama.md §5.1).
|
||||
//
|
||||
// The projection maths is `dr_pano::projection` verbatim; the two must
|
||||
// agree, and a golden test compares them.
|
||||
|
||||
struct Params {
|
||||
// Where the chunk's pixel (0, 0) sits in centred output coordinates,
|
||||
// and the chunk's size.
|
||||
chunk_origin: vec2<f32>,
|
||||
chunk_size: vec2<u32>,
|
||||
// 0 perspective, 1 cylindrical, 2 spherical; and the projection's
|
||||
// scale (the cylinder's radius, the sphere's, the plane's distance) in
|
||||
// output pixels.
|
||||
projection: u32,
|
||||
proj_scale: f32,
|
||||
// The frame's focal length in source pixels, and the gain the frame's
|
||||
// exposure is corrected by.
|
||||
focal: f32,
|
||||
gain: f32,
|
||||
// World → this frame's camera: the transpose of its rotation, one row
|
||||
// per vec4 (padded).
|
||||
r0: vec4<f32>,
|
||||
r1: vec4<f32>,
|
||||
r2: vec4<f32>,
|
||||
// The full frame's size in source pixels (for the edge weight), the
|
||||
// tile's origin within the frame, and the tile's size.
|
||||
frame_size: vec2<f32>,
|
||||
tile_origin: vec2<f32>,
|
||||
tile_size: vec2<u32>,
|
||||
// Pixels over which the weight ramps from the edge to full.
|
||||
feather: f32,
|
||||
_pad: f32,
|
||||
};
|
||||
|
||||
@group(0) @binding(0) var<uniform> p: Params;
|
||||
@group(0) @binding(1) var tile: texture_2d<f32>;
|
||||
// rgb·w summed, then w: four floats per chunk pixel.
|
||||
@group(0) @binding(2) var<storage, read_write> acc: array<vec4<f32>>;
|
||||
|
||||
fn to_direction(u: f32, v: f32) -> vec3<f32> {
|
||||
let s = p.proj_scale;
|
||||
if (p.projection == 0u) {
|
||||
return normalize(vec3<f32>(u, v, s));
|
||||
}
|
||||
if (p.projection == 1u) {
|
||||
let theta = u / s;
|
||||
return normalize(vec3<f32>(sin(theta), v / s, cos(theta)));
|
||||
}
|
||||
let theta = u / s;
|
||||
let phi = v / s;
|
||||
return vec3<f32>(sin(theta) * cos(phi), sin(phi), cos(theta) * cos(phi));
|
||||
}
|
||||
|
||||
fn load(x: i32, y: i32) -> vec4<f32> {
|
||||
return textureLoad(tile, vec2<i32>(x, y), 0);
|
||||
}
|
||||
|
||||
@compute @workgroup_size(8, 8, 1)
|
||||
fn warp(@builtin(global_invocation_id) gid: vec3<u32>) {
|
||||
if (gid.x >= p.chunk_size.x || gid.y >= p.chunk_size.y) {
|
||||
return;
|
||||
}
|
||||
let u = p.chunk_origin.x + f32(gid.x) + 0.5;
|
||||
let v = p.chunk_origin.y + f32(gid.y) + 0.5;
|
||||
let d = to_direction(u, v);
|
||||
let c = vec3<f32>(dot(p.r0.xyz, d), dot(p.r1.xyz, d), dot(p.r2.xyz, d));
|
||||
if (c.z <= 1e-6) {
|
||||
return;
|
||||
}
|
||||
// Source pixel, in the full frame, with the principal point at its
|
||||
// centre. `- 0.5` puts pixel centres on integer coordinates for the
|
||||
// bilinear fetch below.
|
||||
let sx = p.focal * c.x / c.z + p.frame_size.x * 0.5 - 0.5;
|
||||
let sy = p.focal * c.y / c.z + p.frame_size.y * 0.5 - 0.5;
|
||||
// Weight: distance to the nearest frame edge, in pixels, over the
|
||||
// feather. Zero outside the frame.
|
||||
let edge = min(min(sx, p.frame_size.x - 1.0 - sx), min(sy, p.frame_size.y - 1.0 - sy));
|
||||
if (edge <= 0.0) {
|
||||
return;
|
||||
}
|
||||
let w = clamp(edge / max(p.feather, 1.0), 0.0, 1.0);
|
||||
// Into the tile.
|
||||
let tx = sx - p.tile_origin.x;
|
||||
let ty = sy - p.tile_origin.y;
|
||||
let tw = f32(p.tile_size.x);
|
||||
let th = f32(p.tile_size.y);
|
||||
if (tx < 0.0 || ty < 0.0 || tx > tw - 1.0 || ty > th - 1.0) {
|
||||
return;
|
||||
}
|
||||
let x0 = i32(floor(tx));
|
||||
let y0 = i32(floor(ty));
|
||||
let x1 = min(x0 + 1, i32(p.tile_size.x) - 1);
|
||||
let y1 = min(y0 + 1, i32(p.tile_size.y) - 1);
|
||||
let fx = tx - f32(x0);
|
||||
let fy = ty - f32(y0);
|
||||
// The four texels, with their alpha: the tap writes alpha 0 where the
|
||||
// lens correction found no source pixel, and a sample that touches one
|
||||
// of those is a partial pixel — down-weighted by exactly how much of
|
||||
// it is missing, and dropped when all of it is.
|
||||
let s00 = load(x0, y0);
|
||||
let s10 = load(x1, y0);
|
||||
let s01 = load(x0, y1);
|
||||
let s11 = load(x1, y1);
|
||||
let top = mix(s00, s10, fx);
|
||||
let bot = mix(s01, s11, fx);
|
||||
let s = mix(top, bot, fy);
|
||||
if (s.a <= 0.001) {
|
||||
return;
|
||||
}
|
||||
// Colour is the alpha-weighted mean of the texels that exist.
|
||||
let rgb = s.rgb / s.a * p.gain;
|
||||
let wa = w * s.a;
|
||||
let i = gid.y * p.chunk_size.x + gid.x;
|
||||
acc[i] = acc[i] + vec4<f32>(rgb * wa, wa);
|
||||
}
|
||||
|
||||
// Resolve: the accumulated chunk to sixteen-bit samples.
|
||||
struct ResolveParams {
|
||||
chunk_size: vec2<u32>,
|
||||
// Multiplies a normalised value (1.0 = the sensor's white) back to the
|
||||
// sensor's scale: the source's white minus its black (FR-MRG-3).
|
||||
scale: f32,
|
||||
_pad: f32,
|
||||
};
|
||||
|
||||
@group(0) @binding(0) var<uniform> rp: ResolveParams;
|
||||
@group(0) @binding(1) var<storage, read> racc: array<vec4<f32>>;
|
||||
// Two u32 per pixel: (r | g << 16), (b | coverage << 16). Coverage is
|
||||
// 65535 where any frame reached the pixel and 0 where none did, so the
|
||||
// CPU can tell an empty pixel from a black one.
|
||||
@group(0) @binding(2) var<storage, read_write> out: array<vec2<u32>>;
|
||||
|
||||
@compute @workgroup_size(8, 8, 1)
|
||||
fn resolve(@builtin(global_invocation_id) gid: vec3<u32>) {
|
||||
if (gid.x >= rp.chunk_size.x || gid.y >= rp.chunk_size.y) {
|
||||
return;
|
||||
}
|
||||
let i = gid.y * rp.chunk_size.x + gid.x;
|
||||
let a = racc[i];
|
||||
if (a.w <= 0.0) {
|
||||
out[i] = vec2<u32>(0u, 0u);
|
||||
return;
|
||||
}
|
||||
let rgb = clamp(a.rgb / a.w * rp.scale, vec3<f32>(0.0), vec3<f32>(65535.0));
|
||||
let r = u32(round(rgb.r));
|
||||
let g = u32(round(rgb.g));
|
||||
let b = u32(round(rgb.b));
|
||||
out[i] = vec2<u32>(r | (g << 16u), b | (65535u << 16u));
|
||||
}
|
||||
@@ -40,6 +40,10 @@ fn flat_raw(level: u16, curve: BaseCurve) -> RawImage {
|
||||
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
|
||||
color_matrix: Some([1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0]),
|
||||
base_curve: curve,
|
||||
samples_per_pixel: 1,
|
||||
profile: None,
|
||||
make: String::new(),
|
||||
model: String::new(),
|
||||
crop: CropRect {
|
||||
x: 0,
|
||||
y: 0,
|
||||
|
||||
@@ -46,6 +46,10 @@ fn flat_raw(level: u16) -> RawImage {
|
||||
// leaving a curve here would test the suppression rather than the
|
||||
// film. `dr-pipeline` asserts the suppression on the generated source.
|
||||
base_curve: BaseCurve::IDENTITY,
|
||||
samples_per_pixel: 1,
|
||||
profile: None,
|
||||
make: String::new(),
|
||||
model: String::new(),
|
||||
crop: CropRect {
|
||||
x: 0,
|
||||
y: 0,
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
//! TRACES: FR-DSP-5
|
||||
//! TRACES: FR-DSP-5 | R5
|
||||
//! Zooming to 1:1 samples the source, pixel for pixel.
|
||||
//!
|
||||
//! FR-DSP-5: *"Fit, 1:1, and arbitrary zoom levels. At 1:1 and above, the
|
||||
|
||||
@@ -0,0 +1,44 @@
|
||||
[package]
|
||||
name = "dr-inference-engine"
|
||||
version.workspace = true
|
||||
edition.workspace = true
|
||||
rust-version.workspace = true
|
||||
license.workspace = true
|
||||
|
||||
# The one crate that names a runtime, a provider, a vendor library or a
|
||||
# device (docs/inference.md §8). `dr-face` and `dr-segment` ask it for a
|
||||
# session by role and never see which of these answered.
|
||||
|
||||
[dependencies]
|
||||
thiserror.workspace = true
|
||||
log.workspace = true
|
||||
serde.workspace = true
|
||||
serde_json.workspace = true
|
||||
|
||||
# `ort` is the API; what supplies it is decided once per process (§3):
|
||||
# `libonnxruntime` found on disk, or `tract`. Both are behind
|
||||
# `alternative-backend`, so nothing here links C on any target.
|
||||
ort = { workspace = true }
|
||||
ort-tract = { workspace = true, optional = true }
|
||||
# dlopen, and the C types of the table it fetches. Both pure Rust;
|
||||
# `libloading` is already in the tree through wgpu.
|
||||
libloading = { version = "0.8", optional = true }
|
||||
ort-sys = { version = "2.0.0-rc.13", default-features = false, features = ["disable-linking"], optional = true }
|
||||
|
||||
# The NVIDIA rungs exist on the desktop only. These features add `ort`'s
|
||||
# option builders and nothing else — no linking under `alternative-backend` —
|
||||
# but an Android binary has no business carrying even the option names, and
|
||||
# the packaging must never be tempted to (§2, §3.1).
|
||||
[target.'cfg(not(target_os = "android"))'.dependencies]
|
||||
ort = { workspace = true, features = ["cuda", "tensorrt"] }
|
||||
|
||||
[target.'cfg(target_os = "android")'.dependencies]
|
||||
ort = { workspace = true, features = ["qnn"] }
|
||||
|
||||
[features]
|
||||
# The floor: `tract` supplies the API table when no runtime file is found, or
|
||||
# always, in a build without `native`. Tests want this and nothing else.
|
||||
default = ["tract"]
|
||||
tract = ["dep:ort-tract"]
|
||||
# Look for `libonnxruntime` on disk and hand its table to `ort`.
|
||||
native = ["dep:libloading", "dep:ort-sys"]
|
||||
@@ -0,0 +1,169 @@
|
||||
//! The API table `ort` runs on, chosen once (docs/inference.md §3).
|
||||
//!
|
||||
//! `ort` with `alternative-backend` links no runtime and asks, on first use,
|
||||
//! for an `OrtApi` — a struct of function pointers. Two things can fill it:
|
||||
//! a `libonnxruntime` this module `dlopen`s, or `ort-tract`. The Rust build
|
||||
//! is identical either way; the difference is whether a file was found.
|
||||
|
||||
use std::path::PathBuf;
|
||||
use std::sync::OnceLock;
|
||||
|
||||
/// What supplied the table.
|
||||
#[derive(Clone, Debug, PartialEq, Eq)]
|
||||
pub enum Runtime {
|
||||
/// Pure Rust, one core, every operator these graphs use. The floor.
|
||||
Tract,
|
||||
/// The C++ ONNX Runtime, loaded from `path`.
|
||||
OnnxRuntime { path: PathBuf, version: String },
|
||||
}
|
||||
|
||||
impl Runtime {
|
||||
pub fn label(&self) -> String {
|
||||
match self {
|
||||
Runtime::Tract => "tract".into(),
|
||||
Runtime::OnnxRuntime { version, .. } => format!("ONNX Runtime {version}"),
|
||||
}
|
||||
}
|
||||
|
||||
pub fn is_native(&self) -> bool {
|
||||
matches!(self, Runtime::OnnxRuntime { .. })
|
||||
}
|
||||
}
|
||||
|
||||
static RUNTIME: OnceLock<Runtime> = OnceLock::new();
|
||||
|
||||
/// The runtime in use; tract until something installs another.
|
||||
pub fn runtime() -> Runtime {
|
||||
RUNTIME.get().cloned().unwrap_or(Runtime::Tract)
|
||||
}
|
||||
|
||||
/// Install a table if none is installed yet — tract, since no directories
|
||||
/// were named. What a test or an example gets, unless `DARKROOM_ORT_DIR`
|
||||
/// names a runtime: the same variable the desktop honours, so an example
|
||||
/// can be pointed at the runtime the app uses without learning `init`.
|
||||
pub fn ensure_installed() {
|
||||
if RUNTIME.get().is_none() {
|
||||
let dirs: Vec<PathBuf> = std::env::var_os("DARKROOM_ORT_DIR")
|
||||
.map(PathBuf::from)
|
||||
.into_iter()
|
||||
.collect();
|
||||
install(&dirs);
|
||||
}
|
||||
}
|
||||
|
||||
/// Look for `libonnxruntime` in `dirs`, in order, and hand `ort` the first
|
||||
/// table that loads; otherwise tract. Once per process.
|
||||
pub fn install(dirs: &[PathBuf]) -> Runtime {
|
||||
RUNTIME
|
||||
.get_or_init(|| {
|
||||
#[cfg(feature = "native")]
|
||||
for dir in dirs {
|
||||
match load_native(dir) {
|
||||
Ok(rt) => return rt,
|
||||
Err(e) => log::info!("inference: no runtime in {}: {e}", dir.display()),
|
||||
}
|
||||
}
|
||||
#[cfg(not(feature = "native"))]
|
||||
let _ = dirs;
|
||||
install_tract()
|
||||
})
|
||||
.clone()
|
||||
}
|
||||
|
||||
#[cfg(feature = "tract")]
|
||||
fn install_tract() -> Runtime {
|
||||
let _ = ort::set_api(ort_tract::api());
|
||||
Runtime::Tract
|
||||
}
|
||||
|
||||
#[cfg(not(feature = "tract"))]
|
||||
fn install_tract() -> Runtime {
|
||||
// A build with neither tract nor a runtime file has nothing to run
|
||||
// models on; every `open` will report the un-set API rather than panic
|
||||
// somewhere deeper.
|
||||
log::error!("inference: no ONNX Runtime found and tract is not compiled in");
|
||||
Runtime::Tract
|
||||
}
|
||||
|
||||
#[cfg(feature = "native")]
|
||||
fn load_native(dir: &std::path::Path) -> Result<Runtime, String> {
|
||||
let name = if cfg!(target_os = "windows") {
|
||||
"onnxruntime.dll"
|
||||
} else if cfg!(any(target_os = "macos", target_os = "ios")) {
|
||||
"libonnxruntime.dylib"
|
||||
} else {
|
||||
"libonnxruntime.so"
|
||||
};
|
||||
// An empty dir means the bare name: the system loader's search, which on
|
||||
// Android includes the APK's own native libraries.
|
||||
let path = if dir.as_os_str().is_empty() {
|
||||
PathBuf::from(name)
|
||||
} else {
|
||||
find_library(dir, name).ok_or("not present")?
|
||||
};
|
||||
|
||||
// SAFETY: the library's initialisers are ONNX Runtime's own; the symbol
|
||||
// is the documented entry point with the documented signature; the table
|
||||
// is copied out and the library handle is leaked, so every pointer in
|
||||
// the copy stays valid for the life of the process.
|
||||
unsafe {
|
||||
let lib = libloading::Library::new(&path).map_err(|e| e.to_string())?;
|
||||
let get_base: libloading::Symbol<
|
||||
unsafe extern "system" fn() -> *const ort_sys::OrtApiBase,
|
||||
> = lib.get(b"OrtGetApiBase\0").map_err(|e| e.to_string())?;
|
||||
let base = get_base();
|
||||
if base.is_null() {
|
||||
return Err("OrtGetApiBase returned null".into());
|
||||
}
|
||||
let version = std::ffi::CStr::from_ptr(((*base).GetVersionString)())
|
||||
.to_string_lossy()
|
||||
.into_owned();
|
||||
let api = ((*base).GetApi)(ort_sys::ORT_API_VERSION);
|
||||
if api.is_null() {
|
||||
return Err(format!(
|
||||
"ONNX Runtime {version} is older than API version {}",
|
||||
ort_sys::ORT_API_VERSION
|
||||
));
|
||||
}
|
||||
if !ort::set_api((*api).clone()) {
|
||||
return Err("an API table was already installed".into());
|
||||
}
|
||||
std::mem::forget(lib);
|
||||
|
||||
// Qualcomm's DSP loader finds the Hexagon skel through this variable,
|
||||
// and only through it; the runtime's own directory is where the APK
|
||||
// put it. Harmless anywhere else.
|
||||
#[cfg(target_os = "android")]
|
||||
if !dir.as_os_str().is_empty() {
|
||||
std::env::set_var("ADSP_LIBRARY_PATH", dir);
|
||||
}
|
||||
|
||||
log::info!("inference: ONNX Runtime {version} from {}", path.display());
|
||||
Ok(Runtime::OnnxRuntime { path, version })
|
||||
}
|
||||
}
|
||||
|
||||
/// `libonnxruntime.so` in `dir`, or a versioned spelling of it —
|
||||
/// `libonnxruntime.so.1.30.0` is what the Python wheel ships, and a package
|
||||
/// that installs only the versioned file is not wrong.
|
||||
#[cfg(feature = "native")]
|
||||
fn find_library(dir: &std::path::Path, name: &str) -> Option<PathBuf> {
|
||||
let exact = dir.join(name);
|
||||
if exact.is_file() {
|
||||
return Some(exact);
|
||||
}
|
||||
let prefix = format!("{name}.");
|
||||
let mut versioned: Vec<PathBuf> = std::fs::read_dir(dir)
|
||||
.ok()?
|
||||
.filter_map(|e| e.ok())
|
||||
.map(|e| e.path())
|
||||
.filter(|p| {
|
||||
p.is_file()
|
||||
&& p.file_name()
|
||||
.and_then(|n| n.to_str())
|
||||
.is_some_and(|n| n.starts_with(&prefix))
|
||||
})
|
||||
.collect();
|
||||
versioned.sort();
|
||||
versioned.pop()
|
||||
}
|
||||
@@ -0,0 +1,121 @@
|
||||
//! Compiled engines: what a rung builds once per device, and the thread that
|
||||
//! builds them before anyone asks (docs/inference.md §5, §6).
|
||||
//!
|
||||
//! TensorRT keeps its own engine cache keyed by graph hash; QNN writes a
|
||||
//! context model. Both are opaque to this crate, which tracks only *that* a
|
||||
//! model compiled — by the hash of its bytes — so [`crate::open`] can tell a
|
||||
//! request whether to expect the rung or its fallback.
|
||||
|
||||
use std::path::PathBuf;
|
||||
|
||||
use crate::{state, Config, Form, Rung};
|
||||
|
||||
enum Source {
|
||||
File(PathBuf),
|
||||
Bytes(&'static [u8]),
|
||||
}
|
||||
|
||||
/// 64-bit FNV-1a. A cache key, not a checksum: two model files that collide
|
||||
/// here would have to also be the same size and the same role, and the cost
|
||||
/// of that is a rebuilt engine.
|
||||
pub fn hash(bytes: &[u8]) -> u64 {
|
||||
let mut h = 0xcbf2_9ce4_8422_2325u64;
|
||||
for &b in bytes {
|
||||
h ^= b as u64;
|
||||
h = h.wrapping_mul(0x0000_0100_0000_01b3);
|
||||
}
|
||||
h
|
||||
}
|
||||
|
||||
/// The cache entry for `bytes` compiled on `rung`.
|
||||
pub fn key(rung: Rung, bytes: &[u8]) -> String {
|
||||
key_of(rung, hash(bytes))
|
||||
}
|
||||
|
||||
/// The same, from a hash already taken.
|
||||
pub fn key_of(rung: Rung, hash: u64) -> String {
|
||||
format!("{}:{:016x}", rung.label(), hash)
|
||||
}
|
||||
|
||||
/// Where QNN's compiled context for `bytes` lives.
|
||||
pub fn context_path(cfg: &Config, bytes: &[u8]) -> PathBuf {
|
||||
cfg.cache_dir
|
||||
.join("qnn")
|
||||
.join(format!("{:016x}_ctx.onnx", hash(bytes)))
|
||||
}
|
||||
|
||||
/// After the probe: compile every configured model the selected rung can
|
||||
/// take, smallest first, recording each as it lands.
|
||||
pub fn run() {
|
||||
let (rung, cfg) = {
|
||||
let s = state().lock().unwrap();
|
||||
(crate::current_rung(&s), s.config.clone())
|
||||
};
|
||||
if !rung.compiles() {
|
||||
return;
|
||||
}
|
||||
|
||||
// Smallest first, so the detector — the one that runs per image — is
|
||||
// ready soonest (§6 step 3).
|
||||
let mut jobs: Vec<(crate::Role, Source, u64)> = cfg
|
||||
.models
|
||||
.iter()
|
||||
.filter(|(role, _)| rung.serves(*role))
|
||||
.filter_map(|(role, path)| {
|
||||
let (path, form) = crate::resolve_model(*role, path);
|
||||
(form == rung.form(*role)).then(|| {
|
||||
let size = std::fs::metadata(&path).map(|m| m.len()).unwrap_or(0);
|
||||
(*role, Source::File(path), size)
|
||||
})
|
||||
})
|
||||
.chain(cfg.embedded.iter().filter_map(|(role, bytes)| {
|
||||
// An embedded model has no int8 sibling to offer a rung that
|
||||
// wants one; it runs on that rung's fallback.
|
||||
(rung.serves(*role) && rung.form(*role) == Form::F32).then_some((
|
||||
*role,
|
||||
Source::Bytes(bytes),
|
||||
bytes.len() as u64,
|
||||
))
|
||||
}))
|
||||
.collect();
|
||||
jobs.sort_by_key(|j| j.2);
|
||||
state().lock().unwrap().wanted = jobs.len();
|
||||
|
||||
for (role, source, _) in jobs {
|
||||
let (bytes, name) = match &source {
|
||||
Source::File(path) => match std::fs::read(path) {
|
||||
Ok(b) => (b, path.display().to_string()),
|
||||
Err(_) => continue,
|
||||
},
|
||||
Source::Bytes(b) => (b.to_vec(), format!("embedded {role:?}")),
|
||||
};
|
||||
let key = key(rung, &bytes);
|
||||
if state().lock().unwrap().cache.compiled.contains(&key) {
|
||||
continue;
|
||||
}
|
||||
log::info!("inference: compiling {name} for {}", rung.label());
|
||||
let started = std::time::Instant::now();
|
||||
match crate::session::build(rung, role, &bytes, &cfg) {
|
||||
Ok(session) => {
|
||||
drop(session);
|
||||
let mut s = state().lock().unwrap();
|
||||
s.cache.compiled.insert(key);
|
||||
crate::probe::write_cache(&s.config, &s.cache);
|
||||
log::info!(
|
||||
"inference: {name} ready on {} in {:.1} s",
|
||||
rung.label(),
|
||||
started.elapsed().as_secs_f64()
|
||||
);
|
||||
}
|
||||
Err(e) => {
|
||||
// This model stays on the fallback; the others still get
|
||||
// their engine. A corrected model file changes the hash and
|
||||
// is retried.
|
||||
log::warn!(
|
||||
"inference: {name} will not compile for {}: {e}",
|
||||
rung.label()
|
||||
);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,623 @@
|
||||
//! Which runtime, which provider and which model form — decided once per
|
||||
//! device, and the only crate that knows the answer (docs/inference.md).
|
||||
//!
|
||||
//! Consumers ask for a session by [`Role`] and get `ort`'s `Session` back;
|
||||
//! what built it — tract on one core, ONNX Runtime's CPU pool, a TensorRT
|
||||
//! engine, the Hexagon — is this crate's business and shows up in
|
||||
//! [`status`] for the settings row and nowhere else.
|
||||
//!
|
||||
//! The shape follows §3 of the spec: `ort` links nothing (`alternative-backend`),
|
||||
//! and the first call hands it an API table from either a `libonnxruntime`
|
||||
//! found on disk or from `tract`. That choice is once per process, because
|
||||
//! `ort::set_api` is; everything after it — which provider, whether an engine
|
||||
//! has been compiled yet — is per session and may change between two calls.
|
||||
|
||||
use std::collections::{BTreeSet, HashMap};
|
||||
use std::path::{Path, PathBuf};
|
||||
use std::sync::{Arc, Mutex, MutexGuard, OnceLock};
|
||||
use std::time::{Duration, Instant};
|
||||
|
||||
use serde::{Deserialize, Serialize};
|
||||
|
||||
mod api;
|
||||
mod engines;
|
||||
mod probe;
|
||||
mod session;
|
||||
|
||||
pub use api::Runtime;
|
||||
pub use ort::session::Session;
|
||||
|
||||
/// What a model is for. The role fixes the precision rule (§7): an embedder
|
||||
/// runs in f32 on every rung, a detector may run in fp16 or int8.
|
||||
#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash, Serialize, Deserialize)]
|
||||
pub enum Role {
|
||||
Detector,
|
||||
Embedder,
|
||||
Segmenter,
|
||||
Scene,
|
||||
/// The dense landmark model behind the eye reading (docs/faces.md §7c).
|
||||
Landmarks,
|
||||
/// The eye-state and sunglasses classifiers, a few hundred kilobytes.
|
||||
EyeClassifier,
|
||||
/// XFeat, the panorama keypoint detector (docs/panorama.md).
|
||||
Keypoints,
|
||||
/// MI-GAN, the panorama border filler (docs/panorama.md §12). Plain
|
||||
/// convolutions, so any rung serves it; fp16 on TensorRT and int8 on
|
||||
/// the Hexagon are the point of it.
|
||||
Inpainter,
|
||||
}
|
||||
|
||||
/// Which numeric form of a model a session was built from.
|
||||
///
|
||||
/// `Int8` is a different network from `F32` for a detector — it finds a
|
||||
/// different set of faces — which is why [`form_suffix`] exists and why a
|
||||
/// caller appends it to `model_id`.
|
||||
#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash, Serialize, Deserialize)]
|
||||
pub enum Form {
|
||||
F32,
|
||||
Int8,
|
||||
}
|
||||
|
||||
/// A rung of the ladder (§2). Ordered: a user override names the highest rung
|
||||
/// the probe may take, and a compiling rung falls back to the one below it
|
||||
/// until its engine exists.
|
||||
#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash, PartialOrd, Ord, Serialize, Deserialize)]
|
||||
pub enum Rung {
|
||||
/// ONNX Runtime's CPU provider, or tract when no runtime file was found.
|
||||
Cpu,
|
||||
/// NVIDIA, through the CUDA provider. Desktop only.
|
||||
Cuda,
|
||||
/// NVIDIA, through a TensorRT engine compiled on this device. Desktop only.
|
||||
TensorRt,
|
||||
/// Qualcomm's Hexagon NPU through QNN, int8 models only. Android only.
|
||||
Hexagon,
|
||||
}
|
||||
|
||||
impl Rung {
|
||||
pub fn label(self) -> &'static str {
|
||||
match self {
|
||||
Rung::Cpu => "CPU",
|
||||
Rung::Cuda => "CUDA",
|
||||
Rung::TensorRt => "TensorRT",
|
||||
Rung::Hexagon => "Hexagon NPU",
|
||||
}
|
||||
}
|
||||
|
||||
/// The rung a request lands on while this one's engine is still being
|
||||
/// compiled (§6 step 2).
|
||||
fn fallback(self) -> Rung {
|
||||
match self {
|
||||
Rung::TensorRt => Rung::Cuda,
|
||||
Rung::Hexagon | Rung::Cuda | Rung::Cpu => Rung::Cpu,
|
||||
}
|
||||
}
|
||||
|
||||
/// Whether a session on this rung needs an engine built first.
|
||||
fn compiles(self) -> bool {
|
||||
matches!(self, Rung::TensorRt | Rung::Hexagon)
|
||||
}
|
||||
|
||||
/// The model form this rung wants for a role.
|
||||
fn form(self, _role: Role) -> Form {
|
||||
match self {
|
||||
Rung::Hexagon => Form::Int8,
|
||||
_ => Form::F32,
|
||||
}
|
||||
}
|
||||
|
||||
/// Whether this rung runs `role` at all. The Hexagon takes int8 graphs
|
||||
/// only, and the embedder is never int8 (§7) — it runs on the CPU
|
||||
/// beside a detector on the NPU, so its vectors compare across devices.
|
||||
fn serves(self, role: Role) -> bool {
|
||||
match self {
|
||||
Rung::Hexagon => role != Role::Embedder,
|
||||
_ => true,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// How long a session outlives its last use unless [`Config::decay`] says
|
||||
/// otherwise: long enough for the next click, short enough that a session's
|
||||
/// GPU or NPU memory does not sit under the develop view for long.
|
||||
pub const DEFAULT_DECAY: Duration = Duration::from_secs(30);
|
||||
|
||||
/// What [`init`] is told once, at launch.
|
||||
#[derive(Clone, Debug, Default)]
|
||||
pub struct Config {
|
||||
/// Where to look for `libonnxruntime`, in order. An empty path means "the
|
||||
/// bare library name through the system loader", which is how the APK's
|
||||
/// own copy is found on Android.
|
||||
pub runtime_dirs: Vec<PathBuf>,
|
||||
/// Probe cache and compiled engines (§4, §5). Disposable.
|
||||
pub cache_dir: PathBuf,
|
||||
/// The canonical model files on this device, so engines can be compiled
|
||||
/// ahead of the first request for them.
|
||||
pub models: Vec<(Role, PathBuf)>,
|
||||
/// Models compiled into the binary, for the same reason.
|
||||
pub embedded: Vec<(Role, &'static [u8])>,
|
||||
/// The highest rung the user allows; `None` is "the best that works".
|
||||
pub ceiling: Option<Rung>,
|
||||
/// ONNX Runtime's intra-op pool; 0 picks from the core count.
|
||||
pub threads: usize,
|
||||
/// How long an unused session stays loaded. Zero means the default.
|
||||
pub decay: Duration,
|
||||
}
|
||||
|
||||
/// One line for the settings row, and the numbers behind the progress row.
|
||||
#[derive(Clone, Debug)]
|
||||
pub struct Status {
|
||||
pub runtime: Runtime,
|
||||
/// The rung selected, or the floor while the probe is still running.
|
||||
pub rung: Rung,
|
||||
/// Why — "probe passed", or the failure that demoted the rung above.
|
||||
pub reason: String,
|
||||
pub probing: bool,
|
||||
/// Engines compiled and engines wanted, for a compiling rung; `(0, 0)`
|
||||
/// otherwise.
|
||||
pub engines: (usize, usize),
|
||||
/// Every rung above the selected one that was tried, and why it lost.
|
||||
pub failed: Vec<(Rung, String)>,
|
||||
}
|
||||
|
||||
impl Status {
|
||||
/// "Hexagon NPU · int8 · ONNX Runtime 1.29" — the settings row's text.
|
||||
pub fn line(&self) -> String {
|
||||
let form = match self.rung {
|
||||
Rung::Hexagon => " · int8",
|
||||
Rung::TensorRt => " · fp16",
|
||||
_ => "",
|
||||
};
|
||||
format!("{}{} · {}", self.rung.label(), form, self.runtime.label())
|
||||
}
|
||||
}
|
||||
|
||||
/// A model the caller can run, whatever is or is not loaded right now.
|
||||
///
|
||||
/// Holds the bytes, not a session. [`Model::acquire`] finds the loaded copy
|
||||
/// in the registry — shared with every other holder of the same model —
|
||||
/// or loads one, and every acquire refreshes the copy's last-used time.
|
||||
/// The reaper unloads anything idle for [`Config::decay`]; a scan that runs
|
||||
/// the detector on every image never lets it go idle, a click in the
|
||||
/// develop view lets the segmenter go after a quiet spell, and a handle
|
||||
/// used again after that simply loads again. Nobody states a policy.
|
||||
///
|
||||
/// The registry key includes the rung, so a reload after a compiled engine
|
||||
/// has landed moves up to it by itself (§6 step 4).
|
||||
pub struct Model {
|
||||
role: Role,
|
||||
form: Form,
|
||||
bytes: Arc<[u8]>,
|
||||
/// `engines::hash` of the bytes, taken once: an acquire per tile of a
|
||||
/// border fill must not hash 28 MB each time.
|
||||
hash: u64,
|
||||
}
|
||||
|
||||
/// A loaded session, held for one `run` and its output decoding.
|
||||
pub struct Acquired {
|
||||
entry: Arc<Loaded>,
|
||||
}
|
||||
|
||||
struct Loaded {
|
||||
rung: Rung,
|
||||
session: Mutex<Session>,
|
||||
last_used: Mutex<Instant>,
|
||||
}
|
||||
|
||||
impl Model {
|
||||
/// The loaded session, loading it if the reaper took it. Lock it for
|
||||
/// one run; a scan and a develop click can want the same detector at
|
||||
/// once, and the second waits on the first.
|
||||
pub fn acquire(&self) -> Result<Acquired, Error> {
|
||||
acquire(self.role, self.form, &self.bytes, self.hash)
|
||||
}
|
||||
|
||||
pub fn form(&self) -> Form {
|
||||
self.form
|
||||
}
|
||||
}
|
||||
|
||||
impl Acquired {
|
||||
pub fn lock(&self) -> MutexGuard<'_, Session> {
|
||||
self.entry.session.lock().unwrap_or_else(|e| e.into_inner())
|
||||
}
|
||||
|
||||
/// Where this session runs.
|
||||
pub fn rung(&self) -> Rung {
|
||||
self.entry.rung
|
||||
}
|
||||
}
|
||||
|
||||
impl Drop for Acquired {
|
||||
fn drop(&mut self) {
|
||||
// The clock starts when the use ends, not when it began: a long run
|
||||
// is not idle time.
|
||||
*self.entry.last_used.lock().unwrap() = Instant::now();
|
||||
}
|
||||
}
|
||||
|
||||
type Registry = HashMap<String, Arc<Loaded>>;
|
||||
|
||||
static REGISTRY: OnceLock<Mutex<Registry>> = OnceLock::new();
|
||||
|
||||
fn registry() -> &'static Mutex<Registry> {
|
||||
REGISTRY.get_or_init(|| {
|
||||
std::thread::Builder::new()
|
||||
.name("inference-reaper".into())
|
||||
.spawn(|| loop {
|
||||
std::thread::sleep(Duration::from_secs(5));
|
||||
release_idle();
|
||||
})
|
||||
.expect("spawn inference reaper");
|
||||
Mutex::new(HashMap::new())
|
||||
})
|
||||
}
|
||||
|
||||
fn acquire(role: Role, form: Form, bytes: &Arc<[u8]>, hash: u64) -> Result<Acquired, Error> {
|
||||
api::ensure_installed();
|
||||
let (rung, cfg) = {
|
||||
let s = state().lock().unwrap();
|
||||
let selected = current_rung(&s);
|
||||
(
|
||||
effective_rung(&s, selected, role, form, hash),
|
||||
s.config.clone(),
|
||||
)
|
||||
};
|
||||
let key = format!("{role:?}:{}", engines::key_of(rung, hash));
|
||||
|
||||
if let Some(entry) = registry().lock().unwrap().get(&key).cloned() {
|
||||
*entry.last_used.lock().unwrap() = Instant::now();
|
||||
return Ok(Acquired { entry });
|
||||
}
|
||||
|
||||
// Built outside the registry lock: a TensorRT engine load is long enough
|
||||
// that another role's acquire should not wait on it.
|
||||
let session = session::build(rung, role, bytes, &cfg)?;
|
||||
log::debug!("inference: {role:?} loaded on {}", rung.label());
|
||||
let entry = Arc::new(Loaded {
|
||||
rung,
|
||||
session: Mutex::new(session),
|
||||
last_used: Mutex::new(Instant::now()),
|
||||
});
|
||||
let mut reg = registry().lock().unwrap();
|
||||
// Two acquires raced; keep the first, drop this one.
|
||||
let entry = reg.entry(key).or_insert_with(|| entry.clone()).clone();
|
||||
Ok(Acquired { entry })
|
||||
}
|
||||
|
||||
/// Unload every session idle for longer than the decay. The reaper does
|
||||
/// this every five seconds. A session in use survives until its run ends:
|
||||
/// the `Acquired` holds it, the registry merely forgets it.
|
||||
pub fn release_idle() {
|
||||
let decay = match state().lock().unwrap().config.decay {
|
||||
Duration::ZERO => DEFAULT_DECAY,
|
||||
d => d,
|
||||
};
|
||||
let now = Instant::now();
|
||||
registry()
|
||||
.lock()
|
||||
.unwrap()
|
||||
.retain(|_, e| now.duration_since(*e.last_used.lock().unwrap()) < decay);
|
||||
}
|
||||
|
||||
/// Unload every session now, decay or not — what a low-memory signal
|
||||
/// asks for. Sessions mid-run finish first.
|
||||
pub fn release_all() {
|
||||
registry().lock().unwrap().clear();
|
||||
}
|
||||
|
||||
/// Unload every session of `role` now — "I am done segmenting".
|
||||
pub fn unload(role: Role) {
|
||||
let prefix = format!("{role:?}:");
|
||||
registry()
|
||||
.lock()
|
||||
.unwrap()
|
||||
.retain(|k, _| !k.starts_with(&prefix));
|
||||
}
|
||||
|
||||
/// How many sessions are loaded, for the settings row and the tests.
|
||||
pub fn loaded() -> usize {
|
||||
registry().lock().unwrap().len()
|
||||
}
|
||||
|
||||
#[derive(Debug, thiserror::Error)]
|
||||
pub enum Error {
|
||||
#[error(transparent)]
|
||||
Inference(#[from] ort::Error),
|
||||
#[error("reading model: {0}")]
|
||||
Io(#[from] std::io::Error),
|
||||
}
|
||||
|
||||
/// What the probe writes and the next launch reads (§4 step 3).
|
||||
#[derive(Clone, Debug, Default, Serialize, Deserialize)]
|
||||
struct Cache {
|
||||
/// Runtime, driver, hardware and model identity; any change re-probes.
|
||||
fingerprint: String,
|
||||
rung: Option<Rung>,
|
||||
reason: String,
|
||||
/// Model hashes whose engine exists on disk, per compiling rung.
|
||||
compiled: BTreeSet<String>,
|
||||
/// Rungs that failed under this fingerprint, and why. Not retried until
|
||||
/// the fingerprint changes: a wedged driver must not cost every launch
|
||||
/// thirty seconds.
|
||||
failed: Vec<(Rung, String)>,
|
||||
}
|
||||
|
||||
struct State {
|
||||
config: Config,
|
||||
cache: Cache,
|
||||
probing: bool,
|
||||
wanted: usize,
|
||||
}
|
||||
|
||||
static STATE: OnceLock<Mutex<State>> = OnceLock::new();
|
||||
|
||||
fn state() -> &'static Mutex<State> {
|
||||
STATE.get_or_init(|| {
|
||||
Mutex::new(State {
|
||||
config: Config::default(),
|
||||
cache: Cache::default(),
|
||||
probing: false,
|
||||
wanted: 0,
|
||||
})
|
||||
})
|
||||
}
|
||||
|
||||
/// Choose the runtime and start the probe. Idempotent; the first call wins.
|
||||
///
|
||||
/// Returns at once: the probe and any engine compilation run on their own
|
||||
/// low-priority thread, and every request meanwhile is served by the floor
|
||||
/// (§4). Never blocks the first frame.
|
||||
pub fn init(config: Config) {
|
||||
let runtime = api::install(&config.runtime_dirs);
|
||||
{
|
||||
let mut s = state().lock().unwrap();
|
||||
if s.probing || s.cache.rung.is_some() {
|
||||
return;
|
||||
}
|
||||
s.config = config;
|
||||
s.probing = true;
|
||||
}
|
||||
log::info!("inference: runtime {}", runtime.label());
|
||||
std::thread::Builder::new()
|
||||
.name("inference-probe".into())
|
||||
.spawn(move || {
|
||||
probe::run(runtime);
|
||||
engines::run();
|
||||
})
|
||||
.expect("spawn inference probe");
|
||||
}
|
||||
|
||||
/// Make sure `ort` has an API table, for code that drives `ort` directly.
|
||||
/// [`open`] does this itself; only the M1 probe example needs it by name.
|
||||
pub fn ensure_runtime() {
|
||||
api::ensure_installed();
|
||||
}
|
||||
|
||||
/// The line for the settings row.
|
||||
pub fn status() -> Status {
|
||||
let s = state().lock().unwrap();
|
||||
let rung = current_rung(&s);
|
||||
Status {
|
||||
runtime: api::runtime(),
|
||||
rung,
|
||||
reason: s.cache.reason.clone(),
|
||||
failed: s.cache.failed.clone(),
|
||||
probing: s.probing,
|
||||
engines: if rung.compiles() {
|
||||
(s.cache.compiled.len(), s.wanted)
|
||||
} else {
|
||||
(0, 0)
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
fn current_rung(s: &State) -> Rung {
|
||||
if s.probing {
|
||||
Rung::Cpu
|
||||
} else {
|
||||
s.cache.rung.unwrap_or(Rung::Cpu)
|
||||
}
|
||||
}
|
||||
|
||||
/// The file to load for `role` under the current selection, and its form.
|
||||
///
|
||||
/// A rung that wants int8 gets the `.int8.onnx` sibling of the canonical file
|
||||
/// if it exists; otherwise the canonical file, on the rung's fallback. A
|
||||
/// caller adds [`form_suffix`] to the `model_id` it records.
|
||||
pub fn resolve_model(role: Role, canonical: &Path) -> (PathBuf, Form) {
|
||||
let rung = current_rung(&state().lock().unwrap());
|
||||
if rung.serves(role) && rung.form(role) == Form::Int8 {
|
||||
let sibling = int8_sibling(canonical);
|
||||
if sibling.is_file() {
|
||||
return (sibling, Form::Int8);
|
||||
}
|
||||
}
|
||||
(canonical.to_path_buf(), Form::F32)
|
||||
}
|
||||
|
||||
fn int8_sibling(canonical: &Path) -> PathBuf {
|
||||
let stem = canonical
|
||||
.file_stem()
|
||||
.map(|s| s.to_string_lossy().into_owned())
|
||||
.unwrap_or_default();
|
||||
canonical.with_file_name(format!("{stem}.int8.onnx"))
|
||||
}
|
||||
|
||||
/// What a form appends to a detector's `model_id` (§7).
|
||||
pub fn form_suffix(form: Form) -> &'static str {
|
||||
match form {
|
||||
Form::F32 => "",
|
||||
Form::Int8 => "_i8",
|
||||
}
|
||||
}
|
||||
|
||||
/// A handle on the model `bytes` in `role`.
|
||||
///
|
||||
/// Loads it once here, so a graph the runtime rejects fails at
|
||||
/// construction and not on the first image; what happens to that session
|
||||
/// afterwards is the registry's business (see [`Model`]).
|
||||
///
|
||||
/// Works without [`init`] — a test, or the examples — by installing tract
|
||||
/// and using the CPU rung, which is exactly what every consumer did before
|
||||
/// this crate existed.
|
||||
pub fn open(role: Role, form: Form, bytes: &[u8]) -> Result<Model, Error> {
|
||||
let bytes: Arc<[u8]> = Arc::from(bytes);
|
||||
let hash = engines::hash(&bytes);
|
||||
acquire(role, form, &bytes, hash)?;
|
||||
Ok(Model {
|
||||
role,
|
||||
form,
|
||||
bytes,
|
||||
hash,
|
||||
})
|
||||
}
|
||||
|
||||
/// Where a request lands: the selected rung unless the role's precision rule,
|
||||
/// the form on offer, or a missing engine says one lower (§6 step 4).
|
||||
fn effective_rung(s: &State, selected: Rung, role: Role, form: Form, hash: u64) -> Rung {
|
||||
let mut rung = selected;
|
||||
if !rung.serves(role) || rung.form(role) != form {
|
||||
// The embedder on a Hexagon device, or an f32 detector where the int8
|
||||
// sibling was missing: neither can go to the NPU.
|
||||
rung = rung.fallback();
|
||||
}
|
||||
if rung.compiles() && !s.cache.compiled.contains(&engines::key_of(rung, hash)) {
|
||||
rung = rung.fallback();
|
||||
}
|
||||
rung
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
/// The registry is one per process, so these run one at a time.
|
||||
static SERIAL: Mutex<()> = Mutex::new(());
|
||||
fn serial() -> MutexGuard<'static, ()> {
|
||||
SERIAL.lock().unwrap_or_else(|e| e.into_inner())
|
||||
}
|
||||
|
||||
/// The smallest shipped graph, if this checkout has the weights; a test
|
||||
/// suite that needs a research-licensed download is one that does not
|
||||
/// run in CI (docs/faces.md §3), so absence is a skip.
|
||||
fn probe_bytes() -> Option<Vec<u8>> {
|
||||
let path = concat!(
|
||||
env!("CARGO_MANIFEST_DIR"),
|
||||
"/../../models/face/scrfd_500m_640.onnx"
|
||||
);
|
||||
let bytes = std::fs::read(path).ok()?;
|
||||
(bytes.len() > 100_000).then_some(bytes)
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn two_handles_on_one_model_share_one_session() {
|
||||
let _serial = serial();
|
||||
let Some(bytes) = probe_bytes() else { return };
|
||||
release_all();
|
||||
let a = open(Role::Detector, Form::F32, &bytes).unwrap();
|
||||
let b = open(Role::Detector, Form::F32, &bytes).unwrap();
|
||||
assert_eq!(loaded(), 1);
|
||||
let (x, y) = (a.acquire().unwrap(), b.acquire().unwrap());
|
||||
assert!(Arc::ptr_eq(&x.entry, &y.entry));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_released_model_reloads_on_its_next_use() {
|
||||
let _serial = serial();
|
||||
let Some(bytes) = probe_bytes() else { return };
|
||||
release_all();
|
||||
let model = open(Role::Detector, Form::F32, &bytes).unwrap();
|
||||
assert_eq!(loaded(), 1);
|
||||
release_all();
|
||||
assert_eq!(loaded(), 0);
|
||||
let acquired = model.acquire().unwrap();
|
||||
assert_eq!(loaded(), 1);
|
||||
assert_eq!(acquired.lock().inputs().len(), 1);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn an_idle_session_decays_and_a_used_one_does_not() {
|
||||
let _serial = serial();
|
||||
let Some(bytes) = probe_bytes() else { return };
|
||||
release_all();
|
||||
state().lock().unwrap().config.decay = Duration::from_millis(50);
|
||||
let model = open(Role::Detector, Form::F32, &bytes).unwrap();
|
||||
// Used within the decay: stays.
|
||||
std::thread::sleep(Duration::from_millis(30));
|
||||
drop(model.acquire().unwrap());
|
||||
release_idle();
|
||||
assert_eq!(loaded(), 1);
|
||||
// Idle past it: goes.
|
||||
std::thread::sleep(Duration::from_millis(80));
|
||||
release_idle();
|
||||
assert_eq!(loaded(), 0);
|
||||
state().lock().unwrap().config.decay = Duration::ZERO;
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn unload_by_role_leaves_the_other_roles() {
|
||||
let _serial = serial();
|
||||
let Some(bytes) = probe_bytes() else { return };
|
||||
release_all();
|
||||
let _d = open(Role::Detector, Form::F32, &bytes).unwrap();
|
||||
let _s = open(Role::Segmenter, Form::F32, &bytes).unwrap();
|
||||
assert_eq!(loaded(), 2);
|
||||
unload(Role::Segmenter);
|
||||
assert_eq!(loaded(), 1);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn the_hexagon_never_takes_the_embedder() {
|
||||
assert!(!Rung::Hexagon.serves(Role::Embedder));
|
||||
assert!(Rung::Hexagon.serves(Role::Detector));
|
||||
assert_eq!(Rung::Hexagon.form(Role::Detector), Form::Int8);
|
||||
// A detector offered in f32 on a Hexagon device lands on the CPU.
|
||||
let s = State {
|
||||
config: Config::default(),
|
||||
cache: Cache {
|
||||
rung: Some(Rung::Hexagon),
|
||||
..Cache::default()
|
||||
},
|
||||
probing: false,
|
||||
wanted: 0,
|
||||
};
|
||||
assert_eq!(
|
||||
effective_rung(
|
||||
&s,
|
||||
Rung::Hexagon,
|
||||
Role::Embedder,
|
||||
Form::F32,
|
||||
engines::hash(b"")
|
||||
),
|
||||
Rung::Cpu
|
||||
);
|
||||
assert_eq!(
|
||||
effective_rung(
|
||||
&s,
|
||||
Rung::Hexagon,
|
||||
Role::Detector,
|
||||
Form::F32,
|
||||
engines::hash(b"")
|
||||
),
|
||||
Rung::Cpu
|
||||
);
|
||||
// An int8 detector whose context is not compiled yet: also the CPU.
|
||||
assert_eq!(
|
||||
effective_rung(
|
||||
&s,
|
||||
Rung::Hexagon,
|
||||
Role::Detector,
|
||||
Form::Int8,
|
||||
engines::hash(b"")
|
||||
),
|
||||
Rung::Cpu
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn the_status_line_reads_as_the_floor_before_init() {
|
||||
let s = status();
|
||||
assert_eq!(s.rung, Rung::Cpu);
|
||||
assert!(s.line().starts_with("CPU"), "{}", s.line());
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,302 @@
|
||||
//! Walk the ladder, once, by building real sessions (docs/inference.md §4).
|
||||
//!
|
||||
//! A rung is taken when a session builds on it, runs, and is faster than
|
||||
//! the floor. Both halves matter: a provider can register and then fail at
|
||||
//! partition time, and a provider can take a graph — or quietly hand most
|
||||
//! of it back to the CPU — and run it slower than the CPU would have. The outcome is cached against a fingerprint of the
|
||||
//! runtime, the driver, the hardware and the models, and trusted until any
|
||||
//! of those changes.
|
||||
|
||||
use std::path::{Path, PathBuf};
|
||||
use std::time::Instant;
|
||||
|
||||
use crate::{api::Runtime, state, Cache, Config, Form, Role, Rung};
|
||||
|
||||
/// The rungs to try on this platform, best first, under the user's ceiling.
|
||||
fn ladder(ceiling: Option<Rung>) -> Vec<Rung> {
|
||||
#[cfg(target_os = "android")]
|
||||
let all = [Rung::Hexagon];
|
||||
#[cfg(not(target_os = "android"))]
|
||||
let all = [Rung::TensorRt, Rung::Cuda];
|
||||
all.into_iter()
|
||||
.filter(|r| ceiling.is_none_or(|c| *r <= c))
|
||||
.collect()
|
||||
}
|
||||
|
||||
/// The probe body. Sets the cache and clears `probing` when done; never
|
||||
/// panics out, because a failed probe is a result (the floor) and not an
|
||||
/// error.
|
||||
pub fn run(runtime: Runtime) {
|
||||
let cfg = state().lock().unwrap().config.clone();
|
||||
let fingerprint = fingerprint(&runtime, &cfg);
|
||||
|
||||
if let Some(cached) = read_cache(&cfg) {
|
||||
if cached.fingerprint == fingerprint && cached.rung.is_some() {
|
||||
log::info!(
|
||||
"inference: cached selection {} ({})",
|
||||
cached.rung.unwrap().label(),
|
||||
cached.reason
|
||||
);
|
||||
finish(cached);
|
||||
return;
|
||||
}
|
||||
}
|
||||
|
||||
let mut cache = Cache {
|
||||
fingerprint,
|
||||
..Cache::default()
|
||||
};
|
||||
|
||||
if !runtime.is_native() {
|
||||
cache.rung = Some(Rung::Cpu);
|
||||
cache.reason = "no ONNX Runtime found; tract on one core".into();
|
||||
write_cache(&cfg, &cache);
|
||||
finish(cache);
|
||||
return;
|
||||
}
|
||||
|
||||
let Some((role, canonical)) = probe_model(&cfg) else {
|
||||
cache.rung = Some(Rung::Cpu);
|
||||
cache.reason = "no model to probe with".into();
|
||||
write_cache(&cfg, &cache);
|
||||
finish(cache);
|
||||
return;
|
||||
};
|
||||
|
||||
let floor = match time_rung(Rung::Cpu, role, &canonical, &cfg) {
|
||||
Ok((ms, _)) => ms,
|
||||
Err(e) => {
|
||||
// The CPU provider failing is the runtime failing; there is
|
||||
// nothing below it to try, and the reason is worth reading.
|
||||
cache.rung = Some(Rung::Cpu);
|
||||
cache.reason = format!("CPU provider failed: {e}");
|
||||
write_cache(&cfg, &cache);
|
||||
finish(cache);
|
||||
return;
|
||||
}
|
||||
};
|
||||
log::info!("inference: floor {floor:.1} ms on the CPU provider");
|
||||
|
||||
for rung in ladder(cfg.ceiling) {
|
||||
match time_rung(rung, role, &canonical, &cfg) {
|
||||
Ok((ms, key)) if ms < floor => {
|
||||
cache.rung = Some(rung);
|
||||
cache.reason = format!("{ms:.1} ms against {floor:.1} ms on the CPU");
|
||||
if let Some(key) = key {
|
||||
cache.compiled.insert(key);
|
||||
}
|
||||
break;
|
||||
}
|
||||
Ok((ms, _)) => {
|
||||
let why = format!("{ms:.1} ms, slower than the CPU's {floor:.1} ms");
|
||||
log::info!("inference: {} rejected: {why}", rung.label());
|
||||
cache.failed.push((rung, why));
|
||||
}
|
||||
Err(e) => {
|
||||
log::info!("inference: {} failed: {e}", rung.label());
|
||||
cache.failed.push((rung, e));
|
||||
}
|
||||
}
|
||||
}
|
||||
if cache.rung.is_none() {
|
||||
cache.rung = Some(Rung::Cpu);
|
||||
cache.reason = match cache.failed.first() {
|
||||
Some((r, why)) => format!("{} {}", r.label(), first_line(why)),
|
||||
None => "the only rung on this platform".into(),
|
||||
};
|
||||
}
|
||||
write_cache(&cfg, &cache);
|
||||
finish(cache);
|
||||
}
|
||||
|
||||
fn finish(cache: Cache) {
|
||||
let mut s = state().lock().unwrap();
|
||||
s.cache = cache;
|
||||
s.probing = false;
|
||||
}
|
||||
|
||||
/// The smallest detector, or the smallest model of any role if there is
|
||||
/// none. A ~2 MB detector is the cheapest real test of a provider, and the
|
||||
/// detector is the role the int8 forms exist for — the eye classifiers are
|
||||
/// smaller still, and a Hexagon probed with one would fail for want of a
|
||||
/// form nobody ships.
|
||||
fn probe_model(cfg: &Config) -> Option<(Role, PathBuf)> {
|
||||
let smallest = |want: Option<Role>| {
|
||||
cfg.models
|
||||
.iter()
|
||||
.filter(|(role, _)| want.is_none_or(|w| *role == w))
|
||||
.filter_map(|(role, path)| {
|
||||
let size = std::fs::metadata(path).ok()?.len();
|
||||
Some((size, *role, path.clone()))
|
||||
})
|
||||
.min_by_key(|(size, _, _)| *size)
|
||||
.map(|(_, role, path)| (role, path))
|
||||
};
|
||||
smallest(Some(Role::Detector)).or_else(|| smallest(None))
|
||||
}
|
||||
|
||||
/// Build, run once for the engine, then time three runs; the median in
|
||||
/// milliseconds and, for a compiling rung, the cache key of the engine this
|
||||
/// just built.
|
||||
fn time_rung(
|
||||
rung: Rung,
|
||||
role: Role,
|
||||
canonical: &Path,
|
||||
cfg: &Config,
|
||||
) -> Result<(f64, Option<String>), String> {
|
||||
let want = rung.form(role);
|
||||
let path = match want {
|
||||
Form::Int8 => {
|
||||
let p = crate::int8_sibling(canonical);
|
||||
if !p.is_file() {
|
||||
return Err(format!("no int8 form of {}", canonical.display()));
|
||||
}
|
||||
p
|
||||
}
|
||||
Form::F32 => canonical.to_path_buf(),
|
||||
};
|
||||
let bytes = std::fs::read(&path).map_err(|e| e.to_string())?;
|
||||
let started = Instant::now();
|
||||
let mut session =
|
||||
crate::session::build(rung, role, &bytes, cfg).map_err(|e| first_line(&e.to_string()))?;
|
||||
log::info!(
|
||||
"inference: {} session built in {:.1} s",
|
||||
rung.label(),
|
||||
started.elapsed().as_secs_f64()
|
||||
);
|
||||
|
||||
let shape: Vec<usize> = session.inputs()[0]
|
||||
.dtype()
|
||||
.tensor_shape()
|
||||
.ok_or("model input is not a tensor")?
|
||||
.iter()
|
||||
.map(|&d| if d > 0 { d as usize } else { 1 })
|
||||
.collect();
|
||||
let zeros = vec![0f32; shape.iter().product()];
|
||||
let run = |session: &mut ort::session::Session| -> Result<f64, String> {
|
||||
let input = ort::value::Tensor::from_array((shape.clone(), zeros.clone()))
|
||||
.map_err(|e| e.to_string())?;
|
||||
let t = Instant::now();
|
||||
let out = session
|
||||
.run(ort::inputs![input])
|
||||
.map_err(|e| e.to_string())?;
|
||||
let _ = out[0]
|
||||
.try_extract_tensor::<f32>()
|
||||
.map_err(|e| e.to_string())?;
|
||||
Ok(t.elapsed().as_secs_f64() * 1e3)
|
||||
};
|
||||
run(&mut session)?;
|
||||
let mut times = [run(&mut session)?, run(&mut session)?, run(&mut session)?];
|
||||
times.sort_by(|a, b| a.partial_cmp(b).unwrap());
|
||||
let key = rung.compiles().then(|| crate::engines::key(rung, &bytes));
|
||||
Ok((times[1], key))
|
||||
}
|
||||
|
||||
/// The part of a provider's error a person can act on. ONNX Runtime's
|
||||
/// begin with a source path and a C++ template signature; the words —
|
||||
/// "CUDA failure 999: unknown error", "FAIL : Failed to load library" —
|
||||
/// come after, and the settings row has room for one line of them.
|
||||
fn first_line(s: &str) -> String {
|
||||
let line = s.lines().next().unwrap_or("");
|
||||
let start = ["failure", "FAIL :", "Error:", "error:"]
|
||||
.iter()
|
||||
.filter_map(|m| line.find(m))
|
||||
.min()
|
||||
.unwrap_or(0);
|
||||
line[start..].chars().take(200).collect()
|
||||
}
|
||||
|
||||
/// Everything a change of which should re-probe: the runtime and where it
|
||||
/// came from, this crate, the platform, the driver or SoC, and the models.
|
||||
fn fingerprint(runtime: &Runtime, cfg: &Config) -> String {
|
||||
let mut parts = vec![
|
||||
format!("engine {}", env!("CARGO_PKG_VERSION")),
|
||||
format!("{} {}", std::env::consts::OS, std::env::consts::ARCH),
|
||||
match runtime {
|
||||
Runtime::Tract => "tract".to_string(),
|
||||
Runtime::OnnxRuntime { path, version } => format!("ort {version} {}", path.display()),
|
||||
},
|
||||
device_identity(),
|
||||
];
|
||||
for (role, bytes) in &cfg.embedded {
|
||||
parts.push(format!(
|
||||
"{role:?} embedded {:016x}",
|
||||
crate::engines::hash(bytes)
|
||||
));
|
||||
}
|
||||
for (role, path) in &cfg.models {
|
||||
let hash = std::fs::read(path)
|
||||
.map(|b| crate::engines::hash(&b))
|
||||
.unwrap_or(0);
|
||||
parts.push(format!("{role:?} {hash:016x}"));
|
||||
let int8 = crate::int8_sibling(path);
|
||||
if let Ok(b) = std::fs::read(&int8) {
|
||||
parts.push(format!("{role:?} int8 {:016x}", crate::engines::hash(&b)));
|
||||
}
|
||||
}
|
||||
parts.join("\n")
|
||||
}
|
||||
|
||||
#[cfg(target_os = "linux")]
|
||||
fn device_identity() -> String {
|
||||
// The NVIDIA driver's version line; absent means no NVIDIA driver.
|
||||
std::fs::read_to_string("/proc/driver/nvidia/version")
|
||||
.ok()
|
||||
.and_then(|s| s.lines().next().map(str::to_string))
|
||||
.unwrap_or_else(|| "no nvidia driver".into())
|
||||
}
|
||||
|
||||
#[cfg(target_os = "android")]
|
||||
fn device_identity() -> String {
|
||||
// The SoC and the vendor's build: a Hexagon appears or disappears with
|
||||
// either.
|
||||
format!(
|
||||
"{} {}",
|
||||
system_property("ro.soc.model"),
|
||||
system_property("ro.build.version.incremental")
|
||||
)
|
||||
}
|
||||
|
||||
#[cfg(target_os = "android")]
|
||||
fn system_property(name: &str) -> String {
|
||||
extern "C" {
|
||||
fn __system_property_get(
|
||||
name: *const std::ffi::c_char,
|
||||
value: *mut std::ffi::c_char,
|
||||
) -> i32;
|
||||
}
|
||||
let name = std::ffi::CString::new(name).unwrap();
|
||||
let mut buf = [0u8; 92]; // PROP_VALUE_MAX
|
||||
// SAFETY: bionic's documented call; the buffer is PROP_VALUE_MAX bytes.
|
||||
let n = unsafe { __system_property_get(name.as_ptr(), buf.as_mut_ptr().cast()) };
|
||||
String::from_utf8_lossy(&buf[..n.max(0) as usize]).into_owned()
|
||||
}
|
||||
|
||||
#[cfg(not(any(target_os = "linux", target_os = "android")))]
|
||||
fn device_identity() -> String {
|
||||
String::new()
|
||||
}
|
||||
|
||||
fn cache_path(cfg: &Config) -> PathBuf {
|
||||
cfg.cache_dir.join("backend.json")
|
||||
}
|
||||
|
||||
fn read_cache(cfg: &Config) -> Option<Cache> {
|
||||
let text = std::fs::read_to_string(cache_path(cfg)).ok()?;
|
||||
serde_json::from_str(&text).ok()
|
||||
}
|
||||
|
||||
/// Written whole and renamed into place, so a reader never sees half.
|
||||
pub fn write_cache(cfg: &Config, cache: &Cache) {
|
||||
if cfg.cache_dir.as_os_str().is_empty() {
|
||||
return;
|
||||
}
|
||||
let path = cache_path(cfg);
|
||||
let tmp = path.with_extension("json.tmp");
|
||||
let _ = std::fs::create_dir_all(&cfg.cache_dir);
|
||||
if let Ok(text) = serde_json::to_string_pretty(cache) {
|
||||
if std::fs::write(&tmp, text).is_ok() {
|
||||
let _ = std::fs::rename(&tmp, &path);
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,125 @@
|
||||
//! One session builder per rung (docs/inference.md §2, §7, §9).
|
||||
|
||||
use ort::session::Session;
|
||||
|
||||
use crate::{Config, Role, Rung};
|
||||
|
||||
/// Build a session for `bytes` on `rung`.
|
||||
///
|
||||
/// Not strict about the CPU: `session.disable_cpu_ep_fallback` was tried as
|
||||
/// the probe's proof that a provider took the graph, and it refuses the
|
||||
/// Hexagon over the ten quantise/dequantise nodes at the graph's edges that
|
||||
/// QNN declines by policy and that cost microseconds. The probe's proof is
|
||||
/// its clock instead (§4): a provider that hands real work to the CPU is
|
||||
/// slower than the CPU floor and rejected by the same measurement.
|
||||
pub fn build(rung: Rung, role: Role, bytes: &[u8], cfg: &Config) -> ort::Result<Session> {
|
||||
// No optimisation level named. ONNX Runtime's default is already its
|
||||
// fullest, and on tract any level but "disabled" means `into_optimized`,
|
||||
// whose optimiser divides by zero inside yolo26n-seg (tract-data
|
||||
// `stack_tensors`) — a panic across the C API, which is an abort. The
|
||||
// app never asked tract for that and does not start now.
|
||||
let mut b = Session::builder()?.with_intra_threads(threads(cfg))?;
|
||||
// A Hexagon session loads the compiled context when there is one and
|
||||
// compiles it from the model when there is not; the engine thread is
|
||||
// what makes the second case rare (§6).
|
||||
let context = (rung == Rung::Hexagon).then(|| crate::engines::context_path(cfg, bytes));
|
||||
let ready = context.as_ref().is_some_and(|p| p.is_file());
|
||||
b = providers(
|
||||
b,
|
||||
rung,
|
||||
role,
|
||||
cfg,
|
||||
if ready { None } else { context.as_deref() },
|
||||
)?;
|
||||
match (ready, context) {
|
||||
(true, Some(path)) => b.commit_from_file(path),
|
||||
_ => b.commit_from_memory(bytes),
|
||||
}
|
||||
}
|
||||
|
||||
/// The intra-op pool: what the config says, else the cores less two for
|
||||
/// the compositor and the decoder (§9). tract ignores it.
|
||||
fn threads(cfg: &Config) -> usize {
|
||||
if cfg.threads > 0 {
|
||||
return cfg.threads;
|
||||
}
|
||||
std::thread::available_parallelism()
|
||||
.map(|n| n.get().saturating_sub(2).max(1))
|
||||
.unwrap_or(1)
|
||||
}
|
||||
|
||||
#[cfg(not(target_os = "android"))]
|
||||
fn providers(
|
||||
b: ort::session::builder::SessionBuilder,
|
||||
rung: Rung,
|
||||
role: Role,
|
||||
cfg: &Config,
|
||||
_generate_context: Option<&std::path::Path>,
|
||||
) -> ort::Result<ort::session::builder::SessionBuilder> {
|
||||
use ort::ep;
|
||||
match rung {
|
||||
Rung::Cpu => Ok(b),
|
||||
Rung::Cuda => {
|
||||
Ok(b.with_execution_providers([ep::CUDA::default().build().error_on_failure()])?)
|
||||
}
|
||||
Rung::TensorRt => {
|
||||
let cache = cfg.cache_dir.join("tensorrt");
|
||||
let _ = std::fs::create_dir_all(&cache);
|
||||
let cache = cache.to_string_lossy().into_owned();
|
||||
// fp16 for everything but the embedder, whose comparability
|
||||
// across devices is worth more than its 0.2 ms (§7). The
|
||||
// workspace cap keeps the develop view's tiles on the card
|
||||
// (NFR-RES-2). CUDA behind it takes any node TensorRT declines.
|
||||
Ok(b.with_execution_providers([
|
||||
ep::TensorRT::default()
|
||||
.with_fp16(role != Role::Embedder)
|
||||
.with_engine_cache(true)
|
||||
.with_engine_cache_path(&cache)
|
||||
.with_timing_cache(true)
|
||||
.with_timing_cache_path(&cache)
|
||||
.with_max_workspace_size(512 << 20)
|
||||
.build()
|
||||
.error_on_failure(),
|
||||
ep::CUDA::default().build(),
|
||||
])?)
|
||||
}
|
||||
Rung::Hexagon => unreachable!("the Hexagon rung is not on a desktop ladder"),
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(target_os = "android")]
|
||||
fn providers(
|
||||
b: ort::session::builder::SessionBuilder,
|
||||
rung: Rung,
|
||||
_role: Role,
|
||||
_cfg: &Config,
|
||||
generate_context: Option<&std::path::Path>,
|
||||
) -> ort::Result<ort::session::builder::SessionBuilder> {
|
||||
use ort::ep;
|
||||
match rung {
|
||||
Rung::Cpu => Ok(b),
|
||||
Rung::Hexagon => {
|
||||
// The HTP compiles the graph once per device (0.8–1.7 s here).
|
||||
// With `ep.context_enable` ONNX Runtime writes the compiled
|
||||
// context beside the probe cache; the next session loads that
|
||||
// file as its model and skips the compile (§5).
|
||||
let mut b = b;
|
||||
if let Some(ctx) = generate_context {
|
||||
let _ = std::fs::create_dir_all(ctx.parent().unwrap());
|
||||
b = b
|
||||
.with_config_entry("ep.context_enable", "1")?
|
||||
.with_config_entry("ep.context_file_path", ctx.to_string_lossy())?
|
||||
.with_config_entry("ep.context_embed_mode", "0")?;
|
||||
}
|
||||
// Quantise/dequantise at the graph's edges stay on the NPU too,
|
||||
// so a strict build is a whole-graph build.
|
||||
Ok(b.with_execution_providers([ep::QNN::default()
|
||||
.with_backend_path("libQnnHtp.so")
|
||||
.with_performance_mode(ep::qnn::PerformanceMode::Burst)
|
||||
.with_offload_graph_io_quantization(false)
|
||||
.build()
|
||||
.error_on_failure()])?)
|
||||
}
|
||||
Rung::Cuda | Rung::TensorRt => unreachable!("no NVIDIA rung on Android"),
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,38 @@
|
||||
[package]
|
||||
name = "dr-pano"
|
||||
version.workspace = true
|
||||
edition.workspace = true
|
||||
rust-version.workspace = true
|
||||
license.workspace = true
|
||||
# Guards against a Git LFS pointer being embedded in place of the weights.
|
||||
build = "build.rs"
|
||||
|
||||
[dependencies]
|
||||
thiserror.workspace = true
|
||||
log.workspace = true
|
||||
|
||||
# Inference for the learned keypoint detector, on the same footing as
|
||||
# `dr-segment`: `ort` is the API, `dr-inference-engine` decides what runs
|
||||
# it (docs/inference.md), and both are optional so that the geometry —
|
||||
# matching, the rotation solve, the projections — is a dependency-free crate
|
||||
# that tests without a model.
|
||||
ort = { workspace = true, optional = true }
|
||||
dr-inference-engine = { workspace = true, optional = true }
|
||||
ndarray = { workspace = true, optional = true }
|
||||
|
||||
[dev-dependencies]
|
||||
# The example aligns real frames from their embedded previews.
|
||||
dr-decode.workspace = true
|
||||
dr-types.workspace = true
|
||||
env_logger.workspace = true
|
||||
|
||||
[features]
|
||||
default = ["xfeat", "embedded-model"]
|
||||
|
||||
# The XFeat detector (FR-MRG-8) and the MI-GAN filler (FR-MRG-4). Off, the
|
||||
# crate has no model and no runtime — a build that only wants the geometry.
|
||||
xfeat = ["dep:ort", "dep:dr-inference-engine", "dep:ndarray"]
|
||||
|
||||
# Compile the weights into the binary, for the same reason `dr-segment` does:
|
||||
# Android hands the app no path to read a model from (ARCH §6.9).
|
||||
embedded-model = ["xfeat"]
|
||||
@@ -0,0 +1,48 @@
|
||||
//! Check the model is a model and not an LFS pointer.
|
||||
//!
|
||||
//! `models/keypoints/*.onnx` is stored in Git LFS (see `.gitattributes`). A
|
||||
//! clone made without git-lfs, or with `GIT_LFS_SKIP_SMUDGE` set, leaves a
|
||||
//! ~130-byte text pointer at that path instead of the weights, and
|
||||
//! `include_bytes!` would embed it without complaint. Same guard as
|
||||
//! `dr-segment`'s, for the same failure.
|
||||
|
||||
use std::path::Path;
|
||||
|
||||
const MODELS: &[&str] = &[
|
||||
"../../models/keypoints/xfeat-1024.onnx",
|
||||
"../../models/keypoints/xfeat-768.onnx",
|
||||
];
|
||||
|
||||
fn main() {
|
||||
for m in MODELS {
|
||||
println!("cargo:rerun-if-changed={m}");
|
||||
}
|
||||
println!("cargo:rerun-if-changed=build.rs");
|
||||
|
||||
if std::env::var_os("CARGO_FEATURE_EMBEDDED_MODEL").is_none() {
|
||||
return;
|
||||
}
|
||||
|
||||
for model in MODELS.iter().copied() {
|
||||
check(model);
|
||||
}
|
||||
}
|
||||
|
||||
fn check(model: &str) {
|
||||
let path = Path::new(model);
|
||||
let Ok(bytes) = std::fs::read(path) else {
|
||||
panic!(
|
||||
"\n\n{model} is missing.\n\
|
||||
It ships in Git LFS. Run `git lfs install && git lfs pull`, or build \
|
||||
with `--no-default-features` for a geometry-only build.\n"
|
||||
);
|
||||
};
|
||||
|
||||
if bytes.starts_with(b"version https://git-lfs.github.com/spec/") {
|
||||
panic!(
|
||||
"\n\n{model} is a Git LFS pointer, not the model.\n\
|
||||
Run `git lfs install && git lfs pull`, or build with \
|
||||
`--no-default-features` for a geometry-only build.\n"
|
||||
);
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,190 @@
|
||||
//! Align real frames from their embedded previews and draw the result.
|
||||
//!
|
||||
//! ```sh
|
||||
//! cargo run -p dr-pano --example align --release -- fixtures/pano/2025-08-05/*.CR2
|
||||
//! cargo run -p dr-pano --example align --release -- out-prefix frame1.CR2 frame2.CR2 …
|
||||
//! ```
|
||||
//!
|
||||
//! The point of looking rather than asserting: a rotation solve that is
|
||||
//! numerically converged and geometrically wrong — a mirrored axis, a
|
||||
//! transposed homography, an orientation applied the wrong way — produces
|
||||
//! perfectly plausible residuals and a picture that is obviously broken.
|
||||
//! This writes `<prefix>-cyl.ppm`: every frame's preview warped onto a
|
||||
//! cylinder and averaged where they overlap, at a size that fits on a
|
||||
//! screen. Ghosting in the overlaps is the alignment error, made visible.
|
||||
//!
|
||||
//! Previews, not RAW: the alignment runs on proxies in the application too
|
||||
//! (FR-MRG-7), and a camera's embedded JPEG is a proxy the decoder already
|
||||
//! extracts in milliseconds. What is different from the real path is only
|
||||
//! that the pixels are the camera's rendering rather than ours, which the
|
||||
//! geometry does not care about.
|
||||
|
||||
use std::path::PathBuf;
|
||||
use std::time::Instant;
|
||||
|
||||
use dr_pano::bundle::Cameras;
|
||||
use dr_pano::{align, xfeat::XFeat, AlignOptions, Gray, Projection};
|
||||
|
||||
fn main() {
|
||||
env_logger::init();
|
||||
let mut args: Vec<String> = std::env::args().skip(1).collect();
|
||||
if args.is_empty() {
|
||||
eprintln!("usage: align [out-prefix] <frame>...");
|
||||
std::process::exit(2);
|
||||
}
|
||||
let prefix =
|
||||
if args[0].ends_with(".CR2") || args[0].ends_with(".dng") || args[0].ends_with(".jpg") {
|
||||
"align".to_string()
|
||||
} else {
|
||||
args.remove(0)
|
||||
};
|
||||
let paths: Vec<PathBuf> = args.iter().map(PathBuf::from).collect();
|
||||
|
||||
// Previews, oriented, at proxy size.
|
||||
let t = Instant::now();
|
||||
let mut proxies: Vec<Gray> = Vec::new();
|
||||
for p in &paths {
|
||||
let bytes = std::fs::read(p).expect("read");
|
||||
let preview = dr_decode::extract_preview(&bytes, dr_decode::PreviewSize::Full)
|
||||
.expect("embedded preview");
|
||||
let orientation =
|
||||
dr_decode::orientation(&bytes[..bytes.len().min(dr_decode::HEADER_BYTES as usize)])
|
||||
.unwrap_or(dr_types::Orientation::NORMAL);
|
||||
let tag = match orientation.quarter_turns {
|
||||
1 => 6,
|
||||
2 => 3,
|
||||
3 => 8,
|
||||
_ => 1,
|
||||
};
|
||||
let gray = Gray::from_rgba8(
|
||||
&preview.rgba,
|
||||
preview.width as usize,
|
||||
preview.height as usize,
|
||||
)
|
||||
.oriented(tag);
|
||||
let (fitted, _) = gray.fitted(
|
||||
dr_pano::xfeat::INPUT_LONG_EDGE,
|
||||
dr_pano::xfeat::INPUT_LONG_EDGE,
|
||||
);
|
||||
println!(
|
||||
"{:<14} preview {}×{} orientation {} → proxy {}×{}",
|
||||
p.file_name().unwrap().to_string_lossy(),
|
||||
preview.width,
|
||||
preview.height,
|
||||
tag,
|
||||
fitted.width,
|
||||
fitted.height
|
||||
);
|
||||
proxies.push(fitted);
|
||||
}
|
||||
println!("previews in {:?}", t.elapsed());
|
||||
|
||||
// Keypoints.
|
||||
let t = Instant::now();
|
||||
let mut detector = XFeat::embedded().expect("model");
|
||||
let features: Vec<_> = proxies
|
||||
.iter()
|
||||
.map(|g| detector.detect(g).expect("detect"))
|
||||
.collect();
|
||||
for (i, f) in features.iter().enumerate() {
|
||||
println!("frame {i}: {} keypoints", f.len());
|
||||
}
|
||||
println!(
|
||||
"detection in {:?} ({:?} per frame)",
|
||||
t.elapsed(),
|
||||
t.elapsed() / proxies.len() as u32
|
||||
);
|
||||
|
||||
// Alignment.
|
||||
let t = Instant::now();
|
||||
let opts = AlignOptions::default();
|
||||
let alignment = align(&features, &opts).expect("align");
|
||||
println!("alignment in {:?}", t.elapsed());
|
||||
println!(
|
||||
"focal {:.1} px, long edge {} px ({:.1} mm on full frame), rms {:.3} px",
|
||||
alignment.focal,
|
||||
proxies[0].width.max(proxies[0].height),
|
||||
alignment.focal * 36.0 / proxies[0].width.max(proxies[0].height) as f64,
|
||||
alignment.rms_px
|
||||
);
|
||||
for l in &alignment.links {
|
||||
println!(
|
||||
" link {}–{}: {} inliers of {} matches",
|
||||
l.i, l.j, l.inliers, l.matches
|
||||
);
|
||||
}
|
||||
for (k, why) in &alignment.unaligned {
|
||||
println!(" UNALIGNED frame {k}: {why}");
|
||||
}
|
||||
let root = alignment
|
||||
.rotations
|
||||
.iter()
|
||||
.position(|r| *r == Some(dr_pano::linalg::Mat3::IDENTITY))
|
||||
.unwrap_or(0);
|
||||
for (k, r) in alignment.rotations.iter().enumerate() {
|
||||
if let Some(r) = r {
|
||||
// Yaw about y, pitch about x, roll about z, from the matrix's
|
||||
// columns — enough to read a sweep by eye.
|
||||
let yaw = r.0[0][2].atan2(r.0[2][2]).to_degrees();
|
||||
let pitch = (-r.0[1][2]).asin().to_degrees();
|
||||
let roll = r.0[1][0].atan2(r.0[1][1]).to_degrees();
|
||||
println!(
|
||||
" frame {k}: yaw {yaw:7.2}° pitch {pitch:6.2}° roll {roll:6.2}°{}",
|
||||
if k == root { " (reference)" } else { "" }
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
if !alignment.is_complete() {
|
||||
eprintln!("not drawing: the set is not fully aligned");
|
||||
std::process::exit(1);
|
||||
}
|
||||
|
||||
// Draw: a cylinder, averaged where frames overlap.
|
||||
let t = Instant::now();
|
||||
let cameras: Cameras = alignment.cameras();
|
||||
let (fw, fh) = (proxies[0].width as f64, proxies[0].height as f64);
|
||||
let scale = alignment.focal;
|
||||
let bounds = dr_pano::projection::bounds(Projection::Cylindrical, scale, &cameras, (fw, fh))
|
||||
.expect("bounds");
|
||||
// Fit to 3000 px wide.
|
||||
let out_w = 3000usize;
|
||||
let px = bounds.width() / out_w as f64;
|
||||
let out_h = (bounds.height() / px).ceil() as usize;
|
||||
let mut sum = vec![0.0f32; out_w * out_h];
|
||||
let mut count = vec![0u16; out_w * out_h];
|
||||
for oy in 0..out_h {
|
||||
for ox in 0..out_w {
|
||||
let u = bounds.min_u + (ox as f64 + 0.5) * px;
|
||||
let v = bounds.min_v + (oy as f64 + 0.5) * px;
|
||||
let d = Projection::Cylindrical.to_direction(scale, u, v);
|
||||
for (k, g) in proxies.iter().enumerate() {
|
||||
let Some((x, y)) = cameras.project(k, d) else {
|
||||
continue;
|
||||
};
|
||||
let (x, y) = (x + g.width as f64 / 2.0, y + g.height as f64 / 2.0);
|
||||
if x < 0.0 || y < 0.0 || x >= g.width as f64 - 1.0 || y >= g.height as f64 - 1.0 {
|
||||
continue;
|
||||
}
|
||||
let (x0, y0) = (x as usize, y as usize);
|
||||
let (tx, ty) = ((x - x0 as f64) as f32, (y - y0 as f64) as f32);
|
||||
let p = |xx: usize, yy: usize| g.data[yy * g.width + xx];
|
||||
let val = (p(x0, y0) * (1.0 - tx) + p(x0 + 1, y0) * tx) * (1.0 - ty)
|
||||
+ (p(x0, y0 + 1) * (1.0 - tx) + p(x0 + 1, y0 + 1) * tx) * ty;
|
||||
sum[oy * out_w + ox] += val;
|
||||
count[oy * out_w + ox] += 1;
|
||||
}
|
||||
}
|
||||
}
|
||||
let mut ppm = format!("P5\n{out_w} {out_h}\n255\n").into_bytes();
|
||||
ppm.extend(sum.iter().zip(&count).map(|(s, c)| {
|
||||
if *c == 0 {
|
||||
0u8
|
||||
} else {
|
||||
((s / f32::from(*c)).clamp(0.0, 1.0) * 255.0) as u8
|
||||
}
|
||||
}));
|
||||
let out = format!("{prefix}-cyl.pgm");
|
||||
std::fs::write(&out, ppm).expect("write");
|
||||
println!("wrote {out} ({out_w}×{out_h}) in {:?}", t.elapsed());
|
||||
}
|
||||
@@ -0,0 +1,459 @@
|
||||
//! TRACES: FR-MRG-1 | FR-MRG-5
|
||||
//! From features to cameras: the alignment of a whole set.
|
||||
//!
|
||||
//! 1. Match every pair of frames (`matching`).
|
||||
//! 2. For each pair with enough matches, a robust homography
|
||||
//! (`homography::ransac_homography`); a pair is a *link* when its inliers
|
||||
//! pass Brown & Lowe's test, `n_inliers > 8 + 0.3 · n_matches`, which
|
||||
//! is what separates a real overlap from a coincidence of descriptors.
|
||||
//! 3. The focal length: the median of what the links' homographies imply,
|
||||
//! or the caller's hint if none of them implies anything.
|
||||
//! 4. A spanning tree over the links, strongest first, from the
|
||||
//! best-connected frame; rotations chained along it.
|
||||
//! 5. Bundle adjustment over every link's inliers (`bundle`).
|
||||
//!
|
||||
//! What it refuses to do is guess. A frame the tree does not reach is
|
||||
//! reported by index with the reason (FR-MRG-5) and left out of the
|
||||
//! cameras; the caller decides whether a set with a hole is worth
|
||||
//! stitching, and the requirement says it is not.
|
||||
|
||||
use crate::bundle::{self, AdjustOptions, Cameras, Observation};
|
||||
use crate::features::Features;
|
||||
use crate::homography::{self, RobustHomography};
|
||||
use crate::linalg::Mat3;
|
||||
use crate::matching::{match_features, Match};
|
||||
use crate::PanoError;
|
||||
|
||||
#[derive(Debug, Clone, Copy, PartialEq)]
|
||||
pub struct AlignOptions {
|
||||
/// Descriptor similarity floor for a match (`matching`).
|
||||
pub min_similarity: f32,
|
||||
/// RANSAC agreement distance, in pixels of the features' image.
|
||||
pub ransac_px: f64,
|
||||
pub ransac_iterations: usize,
|
||||
/// A pair needs at least this many inliers to be a link, on top of
|
||||
/// Brown & Lowe's ratio test.
|
||||
pub min_inliers: usize,
|
||||
/// Focal length in pixels of the features' image, if the caller knows
|
||||
/// it (EXIF and a sensor width). Used only when the homographies do not
|
||||
/// determine one.
|
||||
pub focal_hint: Option<f64>,
|
||||
pub adjust: AdjustOptions,
|
||||
/// For RANSAC's sampling: the same seed gives the same alignment
|
||||
/// (NFR-MRG-2).
|
||||
pub seed: u64,
|
||||
}
|
||||
|
||||
impl Default for AlignOptions {
|
||||
fn default() -> Self {
|
||||
AlignOptions {
|
||||
min_similarity: 0.82,
|
||||
ransac_px: 3.0,
|
||||
ransac_iterations: 1000,
|
||||
min_inliers: 12,
|
||||
focal_hint: None,
|
||||
adjust: AdjustOptions::default(),
|
||||
seed: 0x5eed,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// An overlap the alignment trusts.
|
||||
#[derive(Debug, Clone, PartialEq)]
|
||||
pub struct Link {
|
||||
pub i: usize,
|
||||
pub j: usize,
|
||||
pub matches: usize,
|
||||
pub inliers: usize,
|
||||
/// Maps centred points of `i` to centred points of `j`.
|
||||
pub h: Mat3,
|
||||
}
|
||||
|
||||
/// Why a frame is not in the alignment.
|
||||
#[derive(Debug, Clone, PartialEq, Eq)]
|
||||
pub enum Unaligned {
|
||||
/// Not enough matches with any other frame to try a geometry.
|
||||
NoMatches,
|
||||
/// Matches existed but none survived RANSAC as a real overlap.
|
||||
NoOverlap,
|
||||
/// Overlaps existed but only with frames that are themselves unaligned.
|
||||
Disconnected,
|
||||
}
|
||||
|
||||
impl std::fmt::Display for Unaligned {
|
||||
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
|
||||
f.write_str(match self {
|
||||
Unaligned::NoMatches => "too few matching features with any other frame",
|
||||
Unaligned::NoOverlap => "no consistent overlap with any other frame",
|
||||
Unaligned::Disconnected => "overlaps only with frames that could not be aligned",
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
/// The result: cameras for the aligned frames, and the rest named.
|
||||
#[derive(Debug, Clone, PartialEq)]
|
||||
pub struct Alignment {
|
||||
/// One rotation per input frame, camera to world, for aligned frames;
|
||||
/// `None` for the unaligned. The reference frame is the best-connected
|
||||
/// one and has the identity.
|
||||
pub rotations: Vec<Option<Mat3>>,
|
||||
/// Focal length in pixels of the features' image.
|
||||
pub focal: f64,
|
||||
pub links: Vec<Link>,
|
||||
pub unaligned: Vec<(usize, Unaligned)>,
|
||||
/// Bundle adjustment's RMS reprojection error, in pixels.
|
||||
pub rms_px: f64,
|
||||
}
|
||||
|
||||
impl Alignment {
|
||||
pub fn is_complete(&self) -> bool {
|
||||
self.unaligned.is_empty()
|
||||
}
|
||||
|
||||
/// The cameras of the aligned frames, indexed as the input — a frame
|
||||
/// that is not aligned is given the identity, so this is only useful
|
||||
/// when [`Self::is_complete`].
|
||||
pub fn cameras(&self) -> Cameras {
|
||||
Cameras {
|
||||
rotations: self
|
||||
.rotations
|
||||
.iter()
|
||||
.map(|r| r.unwrap_or(Mat3::IDENTITY))
|
||||
.collect(),
|
||||
focal: self.focal,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Align a set of frames from their features.
|
||||
///
|
||||
/// Every `Features` must be in its own frame's pixel coordinates with the
|
||||
/// image size filled in; points are centred on the image centre here. The
|
||||
/// frames must all come from the same lens at the same focal length, which
|
||||
/// is the panorama assumption and not checked — the caller has the EXIF.
|
||||
pub fn align(frames: &[Features], opts: &AlignOptions) -> Result<Alignment, PanoError> {
|
||||
let n = frames.len();
|
||||
if n < 2 {
|
||||
return Err(PanoError::Input(
|
||||
"a panorama needs at least two frames".into(),
|
||||
));
|
||||
}
|
||||
|
||||
let centre = |k: usize, i: usize| -> (f64, f64) {
|
||||
let kp = frames[k].keypoints[i];
|
||||
(
|
||||
f64::from(kp.x) - frames[k].width as f64 / 2.0,
|
||||
f64::from(kp.y) - frames[k].height as f64 / 2.0,
|
||||
)
|
||||
};
|
||||
// Scale for the DLT's conditioning: points of order one.
|
||||
let scale = 1.0
|
||||
/ frames
|
||||
.iter()
|
||||
.map(|f| f.width.max(f.height) as f64)
|
||||
.fold(1.0, f64::max);
|
||||
|
||||
// 1 + 2: every pair.
|
||||
let mut links = Vec::new();
|
||||
let mut observations: Vec<Observation> = Vec::new();
|
||||
let mut matched_any = vec![false; n];
|
||||
let t_match = std::time::Instant::now();
|
||||
for i in 0..n {
|
||||
for j in i + 1..n {
|
||||
let matches: Vec<Match> = match_features(&frames[i], &frames[j], opts.min_similarity);
|
||||
log::debug!("pair {i}-{j}: {} matches", matches.len());
|
||||
if matches.len() < 4 {
|
||||
continue;
|
||||
}
|
||||
matched_any[i] = true;
|
||||
matched_any[j] = true;
|
||||
let pairs: Vec<((f64, f64), (f64, f64))> = matches
|
||||
.iter()
|
||||
.map(|m| {
|
||||
let (a, b) = (centre(i, m.a), centre(j, m.b));
|
||||
((a.0 * scale, a.1 * scale), (b.0 * scale, b.1 * scale))
|
||||
})
|
||||
.collect();
|
||||
let Some(RobustHomography { h, inliers }) = homography::ransac_homography(
|
||||
&pairs,
|
||||
opts.ransac_px * scale,
|
||||
opts.ransac_iterations,
|
||||
opts.seed ^ ((i as u64) << 32 | j as u64),
|
||||
) else {
|
||||
continue;
|
||||
};
|
||||
let needed = (8.0 + 0.3 * matches.len() as f64).ceil() as usize;
|
||||
log::debug!("pair {i}-{j}: {} inliers, {needed} needed", inliers.len());
|
||||
if inliers.len() <= needed || inliers.len() < opts.min_inliers {
|
||||
continue;
|
||||
}
|
||||
// Back to pixels: H_px = S⁻¹ H S.
|
||||
let m = h.0;
|
||||
let h_px = Mat3([
|
||||
[m[0][0], m[0][1], m[0][2] / scale],
|
||||
[m[1][0], m[1][1], m[1][2] / scale],
|
||||
[m[2][0] * scale, m[2][1] * scale, m[2][2]],
|
||||
]);
|
||||
for &k in &inliers {
|
||||
let (a, b) = pairs[k];
|
||||
observations.push(Observation {
|
||||
i,
|
||||
j,
|
||||
pi: (a.0 / scale, a.1 / scale),
|
||||
pj: (b.0 / scale, b.1 / scale),
|
||||
});
|
||||
}
|
||||
links.push(Link {
|
||||
i,
|
||||
j,
|
||||
matches: matches.len(),
|
||||
inliers: inliers.len(),
|
||||
h: h_px,
|
||||
});
|
||||
}
|
||||
}
|
||||
|
||||
log::debug!("matching and pairwise geometry in {:?}", t_match.elapsed());
|
||||
|
||||
// 3: the focal length.
|
||||
let mut estimates: Vec<f64> = links
|
||||
.iter()
|
||||
.filter_map(|l| homography::focal_from_homography(&l.h))
|
||||
.filter(|f| f.is_finite() && *f > 0.0)
|
||||
.collect();
|
||||
let longest = frames
|
||||
.iter()
|
||||
.map(|f| f.width.max(f.height) as f64)
|
||||
.fold(0.0, f64::max);
|
||||
let focal = if !estimates.is_empty() {
|
||||
estimates.sort_by(f64::total_cmp);
|
||||
let median = estimates[estimates.len() / 2];
|
||||
// A homography of a nearly pure pan can imply almost anything;
|
||||
// clamp to the range a real lens on this sensor can reach.
|
||||
median.clamp(0.3 * longest, 6.0 * longest)
|
||||
} else if let Some(hint) = opts.focal_hint {
|
||||
hint
|
||||
} else {
|
||||
// No overlap said anything and nobody told us: a normal lens.
|
||||
longest
|
||||
};
|
||||
|
||||
// 4: spanning tree, strongest link first, from the best-connected frame.
|
||||
let mut rotations: Vec<Option<Mat3>> = vec![None; n];
|
||||
let mut unaligned = Vec::new();
|
||||
if links.is_empty() {
|
||||
for (k, &matched) in matched_any.iter().enumerate() {
|
||||
unaligned.push((
|
||||
k,
|
||||
if matched {
|
||||
Unaligned::NoOverlap
|
||||
} else {
|
||||
Unaligned::NoMatches
|
||||
},
|
||||
));
|
||||
}
|
||||
return Ok(Alignment {
|
||||
rotations,
|
||||
focal,
|
||||
links,
|
||||
unaligned,
|
||||
rms_px: 0.0,
|
||||
});
|
||||
}
|
||||
let mut degree = vec![0usize; n];
|
||||
for l in &links {
|
||||
degree[l.i] += l.inliers;
|
||||
degree[l.j] += l.inliers;
|
||||
}
|
||||
let root = (0..n).max_by_key(|&k| degree[k]).unwrap_or(0);
|
||||
rotations[root] = Some(Mat3::IDENTITY);
|
||||
loop {
|
||||
// The strongest link from an aligned frame to an unaligned one.
|
||||
let best = links
|
||||
.iter()
|
||||
.filter(|l| rotations[l.i].is_some() != rotations[l.j].is_some())
|
||||
.max_by_key(|l| l.inliers);
|
||||
let Some(l) = best else { break };
|
||||
let r_ij = homography::rotation_from_homography(&l.h, focal);
|
||||
// H_ij takes points of i to j, so bearings b_j = R_ij b_i, and with
|
||||
// world = R_i · cam_i: R_j = R_i · R_ijᵀ.
|
||||
if let Some(ri) = rotations[l.i] {
|
||||
rotations[l.j] = Some((ri * r_ij.transpose()).orthonormalised());
|
||||
} else if let Some(rj) = rotations[l.j] {
|
||||
rotations[l.i] = Some((rj * r_ij).orthonormalised());
|
||||
}
|
||||
}
|
||||
for k in 0..n {
|
||||
if rotations[k].is_none() {
|
||||
let reason = if !matched_any[k] {
|
||||
Unaligned::NoMatches
|
||||
} else if links.iter().any(|l| l.i == k || l.j == k) {
|
||||
Unaligned::Disconnected
|
||||
} else {
|
||||
Unaligned::NoOverlap
|
||||
};
|
||||
unaligned.push((k, reason));
|
||||
}
|
||||
}
|
||||
|
||||
// 5: adjust the aligned frames together. The reference frame must be
|
||||
// index 0 of the adjustment (it holds frame 0 fixed), so the aligned
|
||||
// frames are renumbered with the root first.
|
||||
let aligned: Vec<usize> = std::iter::once(root)
|
||||
.chain((0..n).filter(|&k| k != root && rotations[k].is_some()))
|
||||
.collect();
|
||||
let index_of = |k: usize| aligned.iter().position(|&a| a == k);
|
||||
let start = Cameras {
|
||||
rotations: aligned.iter().map(|&k| rotations[k].unwrap()).collect(),
|
||||
focal,
|
||||
};
|
||||
let obs: Vec<Observation> = observations
|
||||
.iter()
|
||||
.filter_map(|o| {
|
||||
Some(Observation {
|
||||
i: index_of(o.i)?,
|
||||
j: index_of(o.j)?,
|
||||
pi: o.pi,
|
||||
pj: o.pj,
|
||||
})
|
||||
})
|
||||
.collect();
|
||||
let t_adjust = std::time::Instant::now();
|
||||
let adjusted = bundle::adjust(start, &obs, &opts.adjust)?;
|
||||
log::debug!(
|
||||
"bundle adjustment: {} observations, {} iterations in {:?}",
|
||||
obs.len(),
|
||||
adjusted.iterations,
|
||||
t_adjust.elapsed()
|
||||
);
|
||||
for (slot, &k) in aligned.iter().enumerate() {
|
||||
rotations[k] = Some(adjusted.cameras.rotations[slot]);
|
||||
}
|
||||
|
||||
Ok(Alignment {
|
||||
rotations,
|
||||
focal: adjusted.cameras.focal,
|
||||
links,
|
||||
unaligned,
|
||||
rms_px: adjusted.rms_px,
|
||||
})
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
use crate::features::{Keypoint, DESCRIPTOR_LEN};
|
||||
use crate::linalg::Vec3;
|
||||
|
||||
/// Frames of a synthetic sweep: world directions with random unit
|
||||
/// descriptors, each frame seeing the ones in its field of view.
|
||||
fn synthetic_sweep(
|
||||
n: usize,
|
||||
step: f64,
|
||||
f: f64,
|
||||
w: usize,
|
||||
h: usize,
|
||||
) -> (Vec<Features>, Cameras) {
|
||||
let mut seed = 777u64;
|
||||
let mut rnd = || {
|
||||
seed = seed
|
||||
.wrapping_mul(6364136223846793005)
|
||||
.wrapping_add(1442695040888963407);
|
||||
((seed >> 33) as f64 / (1u64 << 31) as f64) - 0.5
|
||||
};
|
||||
let rotations: Vec<Mat3> = (0..n)
|
||||
.map(|k| {
|
||||
Mat3::rotation(Vec3::new(0.0, 1.0, 0.0), step * k as f64)
|
||||
* Mat3::rotation(Vec3::new(1.0, 0.0, 0.0), 0.02 * ((k % 3) as f64 - 1.0))
|
||||
})
|
||||
.collect();
|
||||
let truth = Cameras {
|
||||
rotations,
|
||||
focal: f,
|
||||
};
|
||||
let total = step * (n as f64 - 1.0);
|
||||
let mut frames: Vec<Features> = (0..n)
|
||||
.map(|_| Features {
|
||||
keypoints: Vec::new(),
|
||||
descriptors: Vec::new(),
|
||||
width: w,
|
||||
height: h,
|
||||
})
|
||||
.collect();
|
||||
for _ in 0..600 * n {
|
||||
let yaw = rnd() * (total + 0.8) + total / 2.0;
|
||||
let pitch = rnd() * 0.5;
|
||||
let d = Vec3::new(
|
||||
yaw.sin() * pitch.cos(),
|
||||
pitch.sin(),
|
||||
yaw.cos() * pitch.cos(),
|
||||
);
|
||||
let desc: Vec<f32> = (0..DESCRIPTOR_LEN).map(|_| rnd() as f32).collect();
|
||||
let norm = desc.iter().map(|v| v * v).sum::<f32>().sqrt();
|
||||
let desc: Vec<f32> = desc.iter().map(|v| v / norm).collect();
|
||||
for (k, frame) in frames.iter_mut().enumerate() {
|
||||
if let Some(p) = truth.project(k, d) {
|
||||
let (x, y) = (p.0 + w as f64 / 2.0, p.1 + h as f64 / 2.0);
|
||||
if x >= 0.0 && x < w as f64 && y >= 0.0 && y < h as f64 {
|
||||
frame.keypoints.push(Keypoint {
|
||||
x: (x + rnd() * 0.6) as f32,
|
||||
y: (y + rnd() * 0.6) as f32,
|
||||
score: 1.0,
|
||||
});
|
||||
frame.descriptors.extend_from_slice(&desc);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
(frames, truth)
|
||||
}
|
||||
|
||||
fn angle_between(a: Mat3, b: Mat3) -> f64 {
|
||||
(a.transpose() * b).log().norm()
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_synthetic_sweep_is_aligned_to_its_truth() {
|
||||
let (frames, truth) = synthetic_sweep(6, 0.3, 1400.0, 1024, 768);
|
||||
let out = align(&frames, &AlignOptions::default()).expect("aligned");
|
||||
assert!(out.is_complete(), "unaligned: {:?}", out.unaligned);
|
||||
assert_eq!(out.links.len(), 5 + 4, "links: {}", out.links.len());
|
||||
assert!((out.focal - 1400.0).abs() < 15.0, "focal {}", out.focal);
|
||||
assert!(out.rms_px < 1.0, "rms {}", out.rms_px);
|
||||
// Relative rotations match the truth's, whichever frame is the root.
|
||||
let root = out
|
||||
.rotations
|
||||
.iter()
|
||||
.position(|r| *r == Some(Mat3::IDENTITY))
|
||||
.unwrap();
|
||||
for k in 0..6 {
|
||||
let rel_truth = truth.rotations[root].transpose() * truth.rotations[k];
|
||||
let rel_out = out.rotations[k].unwrap();
|
||||
let err = angle_between(rel_truth, rel_out);
|
||||
assert!(err < 2e-3, "frame {k} off by {err} rad");
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_frame_from_nowhere_is_named_not_guessed() {
|
||||
let (mut frames, _) = synthetic_sweep(4, 0.3, 1400.0, 1024, 768);
|
||||
// Frame 3 gets descriptors nobody else has.
|
||||
for v in &mut frames[3].descriptors {
|
||||
*v = -*v;
|
||||
}
|
||||
let out = align(&frames, &AlignOptions::default()).expect("aligned");
|
||||
assert_eq!(out.unaligned.len(), 1);
|
||||
assert_eq!(out.unaligned[0].0, 3);
|
||||
assert!(out.rotations[3].is_none());
|
||||
assert!(out.rotations[..3].iter().all(Option::is_some));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn one_frame_is_refused() {
|
||||
let (frames, _) = synthetic_sweep(1, 0.3, 1400.0, 640, 480);
|
||||
assert!(matches!(
|
||||
align(&frames, &AlignOptions::default()),
|
||||
Err(PanoError::Input(_))
|
||||
));
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,446 @@
|
||||
//! Bundle adjustment: every rotation and the focal length, refined together.
|
||||
//!
|
||||
//! The pairwise homographies (`homography.rs`) each know about two frames.
|
||||
//! Chained around a loop they disagree with themselves by the accumulated
|
||||
//! error, and a twelve-frame sweep chained end to end drifts by a visible
|
||||
//! amount. This solves for all the rotations at once, against every inlier
|
||||
//! match of every pair, so the error is spread rather than accumulated —
|
||||
//! Brown & Lowe's step 4, with the camera model reduced to what a panorama
|
||||
//! needs: one rotation per frame and one focal length shared by all.
|
||||
//!
|
||||
//! Levenberg–Marquardt with a numerical Jacobian. Analytic derivatives of a
|
||||
//! rotation's projection are not hard, but they are a second place the
|
||||
//! model is written down, and the model is small: forty parameters, a few
|
||||
//! thousand residuals, a Jacobian that costs forty residual evaluations.
|
||||
//! The whole solve is milliseconds. Correctness over cleverness, and one
|
||||
//! definition of the projection to keep right.
|
||||
|
||||
use crate::linalg::{DMat, Mat3, Vec3};
|
||||
use crate::PanoError;
|
||||
|
||||
/// A point in one image, centred on the principal point, in pixels.
|
||||
pub type Point = (f64, f64);
|
||||
|
||||
/// One inlier match between two frames.
|
||||
#[derive(Debug, Clone, Copy, PartialEq)]
|
||||
pub struct Observation {
|
||||
pub i: usize,
|
||||
pub j: usize,
|
||||
pub pi: Point,
|
||||
pub pj: Point,
|
||||
}
|
||||
|
||||
/// What the adjustment starts from and returns: a rotation per frame
|
||||
/// (camera to world; frame 0 is the world) and the focal length in pixels.
|
||||
#[derive(Debug, Clone, PartialEq)]
|
||||
pub struct Cameras {
|
||||
pub rotations: Vec<Mat3>,
|
||||
pub focal: f64,
|
||||
}
|
||||
|
||||
impl Cameras {
|
||||
/// The unit direction, in world space, that pixel `p` of frame `i` looks
|
||||
/// along.
|
||||
pub fn bearing(&self, i: usize, p: Point) -> Vec3 {
|
||||
self.rotations[i] * Vec3::new(p.0, p.1, self.focal).normalised()
|
||||
}
|
||||
|
||||
/// Where world direction `d` lands in frame `j`, or `None` if it is
|
||||
/// behind the camera.
|
||||
pub fn project(&self, j: usize, d: Vec3) -> Option<Point> {
|
||||
let c = self.rotations[j].transpose() * d;
|
||||
if c.z() <= 1e-9 {
|
||||
return None;
|
||||
}
|
||||
Some((self.focal * c.x() / c.z(), self.focal * c.y() / c.z()))
|
||||
}
|
||||
}
|
||||
|
||||
#[derive(Debug, Clone, Copy, PartialEq)]
|
||||
pub struct AdjustOptions {
|
||||
pub max_iterations: usize,
|
||||
/// Residuals beyond this many pixels are down-weighted (Huber), so a
|
||||
/// mismatch RANSAC let through pulls with bounded force.
|
||||
pub huber_px: f64,
|
||||
/// Whether the focal length is a free parameter. Off, it is held at the
|
||||
/// starting value — for a set whose rotations are all small, the focal
|
||||
/// length is weakly observable and better taken from the homographies'
|
||||
/// median than pulled about by noise.
|
||||
pub refine_focal: bool,
|
||||
}
|
||||
|
||||
impl Default for AdjustOptions {
|
||||
fn default() -> Self {
|
||||
AdjustOptions {
|
||||
max_iterations: 50,
|
||||
huber_px: 3.0,
|
||||
refine_focal: true,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// The adjusted cameras and the fit.
|
||||
#[derive(Debug, Clone, PartialEq)]
|
||||
pub struct Adjusted {
|
||||
pub cameras: Cameras,
|
||||
/// Root-mean-square reprojection error over all observations, in pixels
|
||||
/// (unweighted, so an outlier RANSAC missed shows here rather than
|
||||
/// hiding under its Huber weight).
|
||||
pub rms_px: f64,
|
||||
pub iterations: usize,
|
||||
}
|
||||
|
||||
/// Refine `start` against `observations`.
|
||||
///
|
||||
/// Frame 0's rotation is held fixed: the world frame is arbitrary and
|
||||
/// fixing one camera removes the freedom. Every other frame must appear in
|
||||
/// at least one observation or its rotation is undetermined and the normal
|
||||
/// equations are singular — the caller (`align`) guarantees it by only
|
||||
/// adjusting frames a spanning tree reached.
|
||||
pub fn adjust(
|
||||
start: Cameras,
|
||||
observations: &[Observation],
|
||||
opts: &AdjustOptions,
|
||||
) -> Result<Adjusted, PanoError> {
|
||||
let n_frames = start.rotations.len();
|
||||
if n_frames < 2 || observations.is_empty() {
|
||||
let rms = rms(&start, observations);
|
||||
return Ok(Adjusted {
|
||||
cameras: start,
|
||||
rms_px: rms,
|
||||
iterations: 0,
|
||||
});
|
||||
}
|
||||
// Every adjustable frame must be constrained by something, or its
|
||||
// block of the normal equations is zero and the solve is meaningless —
|
||||
// checked here, by name, rather than left to surface as a step that
|
||||
// fails to lower the cost.
|
||||
let mut seen = vec![false; n_frames];
|
||||
for o in observations {
|
||||
seen[o.i] = true;
|
||||
seen[o.j] = true;
|
||||
}
|
||||
if let Some(k) = (1..n_frames).find(|&k| !seen[k]) {
|
||||
return Err(PanoError::Geometry(format!(
|
||||
"frame {k} has no observations constraining it"
|
||||
)));
|
||||
}
|
||||
let n_rot = 3 * (n_frames - 1);
|
||||
let n_params = n_rot + usize::from(opts.refine_focal);
|
||||
let n_res = 2 * observations.len();
|
||||
|
||||
// Parameters are *increments* on the current cameras, re-applied each
|
||||
// accepted step: rotation k ← exp(δ_k) · rotation k, focal ← f · exp(δ_f).
|
||||
// Composing on the left keeps the increment in world space, where a
|
||||
// small rotation means the same thing for every frame.
|
||||
let apply = |base: &Cameras, x: &[f64]| -> Cameras {
|
||||
let mut rotations = base.rotations.clone();
|
||||
for k in 1..n_frames {
|
||||
let w = Vec3::new(x[3 * (k - 1)], x[3 * (k - 1) + 1], x[3 * (k - 1) + 2]);
|
||||
rotations[k] = (Mat3::exp(w) * base.rotations[k]).orthonormalised();
|
||||
}
|
||||
let focal = if opts.refine_focal {
|
||||
base.focal * x[n_rot].exp()
|
||||
} else {
|
||||
base.focal
|
||||
};
|
||||
Cameras { rotations, focal }
|
||||
};
|
||||
|
||||
let residuals = |c: &Cameras, out: &mut Vec<f64>| {
|
||||
out.clear();
|
||||
for o in observations {
|
||||
let d = c.bearing(o.i, o.pi);
|
||||
match c.project(o.j, d) {
|
||||
Some((x, y)) => {
|
||||
out.push(x - o.pj.0);
|
||||
out.push(y - o.pj.1);
|
||||
}
|
||||
None => {
|
||||
// Behind the camera: as wrong as a residual can be
|
||||
// without being infinite. The Huber weight caps its pull.
|
||||
out.push(1e4);
|
||||
out.push(1e4);
|
||||
}
|
||||
}
|
||||
}
|
||||
};
|
||||
|
||||
let weights = |r: &[f64], out: &mut Vec<f64>| {
|
||||
out.clear();
|
||||
for pair in r.chunks_exact(2) {
|
||||
let m = (pair[0] * pair[0] + pair[1] * pair[1]).sqrt();
|
||||
let w = if m > opts.huber_px {
|
||||
opts.huber_px / m
|
||||
} else {
|
||||
1.0
|
||||
};
|
||||
out.push(w);
|
||||
out.push(w);
|
||||
}
|
||||
};
|
||||
|
||||
// The robust cost itself, not the weighted sum of squares: the weights
|
||||
// above are the IRLS linearisation for one step, and comparing two
|
||||
// steps by sums taken under different weights would accept the wrong
|
||||
// ones. Huber: quadratic within the threshold, linear beyond it.
|
||||
let cost = |r: &[f64]| -> f64 {
|
||||
r.chunks_exact(2)
|
||||
.map(|pair| {
|
||||
let m = (pair[0] * pair[0] + pair[1] * pair[1]).sqrt();
|
||||
if m <= opts.huber_px {
|
||||
m * m
|
||||
} else {
|
||||
2.0 * opts.huber_px * m - opts.huber_px * opts.huber_px
|
||||
}
|
||||
})
|
||||
.sum()
|
||||
};
|
||||
|
||||
let mut cameras = start;
|
||||
let mut r = Vec::with_capacity(n_res);
|
||||
let mut w = Vec::with_capacity(n_res);
|
||||
residuals(&cameras, &mut r);
|
||||
weights(&r, &mut w);
|
||||
let mut current = cost(&r);
|
||||
|
||||
let mut lambda = 1e-3;
|
||||
let mut jac = vec![0.0f64; n_res * n_params];
|
||||
let mut r_plus = Vec::with_capacity(n_res);
|
||||
let zero = vec![0.0f64; n_params];
|
||||
let mut iterations = 0;
|
||||
|
||||
for _ in 0..opts.max_iterations {
|
||||
iterations += 1;
|
||||
|
||||
// Numerical Jacobian about the current cameras (x = 0).
|
||||
const H: f64 = 1e-6;
|
||||
for p in 0..n_params {
|
||||
let mut x = zero.clone();
|
||||
x[p] = H;
|
||||
let c_plus = apply(&cameras, &x);
|
||||
residuals(&c_plus, &mut r_plus);
|
||||
for (k, (rp, r0)) in r_plus.iter().zip(&r).enumerate() {
|
||||
jac[k * n_params + p] = (rp - r0) / H;
|
||||
}
|
||||
}
|
||||
|
||||
// Normal equations, weighted: (JᵀWJ + λ·diag) δ = −JᵀWr.
|
||||
let mut a = DMat::zeros(n_params);
|
||||
let mut b = vec![0.0f64; n_params];
|
||||
for k in 0..n_res {
|
||||
let row = &jac[k * n_params..(k + 1) * n_params];
|
||||
let wk = w[k];
|
||||
for p in 0..n_params {
|
||||
b[p] -= wk * row[p] * r[k];
|
||||
for q in 0..n_params {
|
||||
a[(p, q)] += wk * row[p] * row[q];
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Try steps with increasing damping until one lowers the cost.
|
||||
let mut accepted = false;
|
||||
for _ in 0..10 {
|
||||
let mut damped = a.clone();
|
||||
for p in 0..n_params {
|
||||
let d = a[(p, p)];
|
||||
damped[(p, p)] = d + lambda * d.max(1e-9);
|
||||
}
|
||||
let Some(delta) = damped.solve_spd(&b) else {
|
||||
return Err(PanoError::Geometry(
|
||||
"the adjustment's normal equations are singular: a frame has no \
|
||||
observations constraining it"
|
||||
.into(),
|
||||
));
|
||||
};
|
||||
let candidate = apply(&cameras, &delta);
|
||||
residuals(&candidate, &mut r_plus);
|
||||
let c_new = cost(&r_plus);
|
||||
if c_new < current {
|
||||
let improvement = (current - c_new) / current.max(1e-12);
|
||||
let step: f64 = delta.iter().map(|d| d * d).sum::<f64>().sqrt();
|
||||
cameras = candidate;
|
||||
std::mem::swap(&mut r, &mut r_plus);
|
||||
weights(&r, &mut w);
|
||||
current = c_new;
|
||||
lambda = (lambda / 3.0).max(1e-9);
|
||||
accepted = true;
|
||||
// Converged when a *lightly damped* step no longer helps. A
|
||||
// heavily damped step is small by construction and would
|
||||
// pass an improvement test long before the minimum.
|
||||
if step < 1e-10 || (improvement < 1e-8 && lambda < 1e-2) {
|
||||
return Ok(Adjusted {
|
||||
rms_px: rms(&cameras, observations),
|
||||
cameras,
|
||||
iterations,
|
||||
});
|
||||
}
|
||||
break;
|
||||
}
|
||||
lambda *= 5.0;
|
||||
}
|
||||
if !accepted {
|
||||
break;
|
||||
}
|
||||
}
|
||||
|
||||
Ok(Adjusted {
|
||||
rms_px: rms(&cameras, observations),
|
||||
cameras,
|
||||
iterations,
|
||||
})
|
||||
}
|
||||
|
||||
/// Unweighted RMS reprojection error in pixels.
|
||||
pub fn rms(c: &Cameras, observations: &[Observation]) -> f64 {
|
||||
if observations.is_empty() {
|
||||
return 0.0;
|
||||
}
|
||||
let sum: f64 = observations
|
||||
.iter()
|
||||
.map(|o| match c.project(o.j, c.bearing(o.i, o.pi)) {
|
||||
Some((x, y)) => (x - o.pj.0).powi(2) + (y - o.pj.1).powi(2),
|
||||
None => 1e8,
|
||||
})
|
||||
.sum();
|
||||
(sum / observations.len() as f64).sqrt()
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
/// A synthetic sweep: `n` cameras panned by `step` radians each with a
|
||||
/// little pitch and roll, `f` pixels, and matches between neighbours
|
||||
/// from a cloud of world directions.
|
||||
fn sweep(n: usize, step: f64, f: f64, noise_px: f64) -> (Cameras, Vec<Observation>) {
|
||||
let mut rotations = Vec::new();
|
||||
for k in 0..n {
|
||||
let yaw = step * k as f64;
|
||||
let pitch = 0.01 * ((k * 7) % 3) as f64;
|
||||
let roll = 0.005 * ((k * 5) % 4) as f64;
|
||||
let r = Mat3::rotation(Vec3::new(0.0, 1.0, 0.0), yaw)
|
||||
* Mat3::rotation(Vec3::new(1.0, 0.0, 0.0), pitch)
|
||||
* Mat3::rotation(Vec3::new(0.0, 0.0, 1.0), roll);
|
||||
rotations.push(r);
|
||||
}
|
||||
let truth = Cameras {
|
||||
rotations,
|
||||
focal: f,
|
||||
};
|
||||
|
||||
// World directions: a fan across the whole sweep.
|
||||
let mut obs = Vec::new();
|
||||
let mut seed = 12345u64;
|
||||
let mut rnd = || {
|
||||
seed = seed
|
||||
.wrapping_mul(6364136223846793005)
|
||||
.wrapping_add(1442695040888963407);
|
||||
((seed >> 33) as f64 / (1u64 << 31) as f64) - 0.5
|
||||
};
|
||||
let total = step * (n as f64 - 1.0);
|
||||
for _ in 0..400 * n {
|
||||
let yaw = rnd() * (total + 0.8) + total / 2.0;
|
||||
let pitch = rnd() * 0.5;
|
||||
let d = Vec3::new(
|
||||
yaw.sin() * pitch.cos(),
|
||||
pitch.sin(),
|
||||
yaw.cos() * pitch.cos(),
|
||||
)
|
||||
.normalised();
|
||||
// Visible in which frames? Within ±0.35 f of centre.
|
||||
let mut seen: Vec<(usize, Point)> = Vec::new();
|
||||
for k in 0..n {
|
||||
if let Some(p) = truth.project(k, d) {
|
||||
if p.0.abs() < 0.35 * f && p.1.abs() < 0.25 * f {
|
||||
seen.push((k, (p.0 + rnd() * noise_px, p.1 + rnd() * noise_px)));
|
||||
}
|
||||
}
|
||||
}
|
||||
for a in 0..seen.len() {
|
||||
for b in a + 1..seen.len() {
|
||||
obs.push(Observation {
|
||||
i: seen[a].0,
|
||||
j: seen[b].0,
|
||||
pi: seen[a].1,
|
||||
pj: seen[b].1,
|
||||
});
|
||||
}
|
||||
}
|
||||
}
|
||||
(truth, obs)
|
||||
}
|
||||
|
||||
fn angle_between(a: Mat3, b: Mat3) -> f64 {
|
||||
(a.transpose() * b).log().norm()
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_perturbed_start_converges_back_to_the_truth() {
|
||||
let (truth, obs) = sweep(6, 0.3, 1400.0, 0.0);
|
||||
assert!(obs.len() > 500);
|
||||
// Perturb every rotation but the first by ~1°, and the focal by 5%.
|
||||
let mut start = truth.clone();
|
||||
for k in 1..6 {
|
||||
let w = Vec3::new(0.01, -0.015, 0.008) * (k as f64 / 3.0);
|
||||
start.rotations[k] = Mat3::exp(w) * start.rotations[k];
|
||||
}
|
||||
start.focal *= 1.05;
|
||||
let before = rms(&start, &obs);
|
||||
let out = adjust(start, &obs, &AdjustOptions::default()).expect("solvable");
|
||||
assert!(out.rms_px < 1e-3, "rms {} (was {before})", out.rms_px);
|
||||
assert!(
|
||||
(out.cameras.focal - 1400.0).abs() < 0.5,
|
||||
"focal {}",
|
||||
out.cameras.focal
|
||||
);
|
||||
for k in 0..6 {
|
||||
let err = angle_between(out.cameras.rotations[k], truth.rotations[k]);
|
||||
assert!(err < 1e-5, "frame {k} off by {err} rad");
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn noise_is_averaged_rather_than_accumulated() {
|
||||
let (truth, obs) = sweep(8, 0.25, 1400.0, 1.0);
|
||||
let mut start = truth.clone();
|
||||
for k in 1..8 {
|
||||
start.rotations[k] =
|
||||
Mat3::exp(Vec3::new(0.0, 0.004 * k as f64, 0.0)) * start.rotations[k];
|
||||
}
|
||||
let out = adjust(start, &obs, &AdjustOptions::default()).expect("solvable");
|
||||
// ±0.5 px of uniform noise on every coordinate has an RMS of 0.41 px
|
||||
// per axis, so the fit's RMS over both axes should sit near 0.58 and
|
||||
// cannot be much below it.
|
||||
assert!(out.rms_px < 0.7, "rms {}", out.rms_px);
|
||||
|
||||
// The focal length and the sweep are nearly degenerate for a
|
||||
// single row: only the perspective inside each overlap pins the
|
||||
// focal, and a pixel of noise is worth about a tenth of a percent of
|
||||
// it. What that error does is scale every yaw by the same factor —
|
||||
// a uniform stretch of the panorama, invisible in the result — so the
|
||||
// absolute rotation error grows linearly along the sweep and is not
|
||||
// the measure of the solve. The residual after removing that stretch
|
||||
// is.
|
||||
let f_ratio = out.cameras.focal / 1400.0;
|
||||
assert!((f_ratio - 1.0).abs() < 5e-3, "focal {}", out.cameras.focal);
|
||||
for k in 0..8 {
|
||||
let yaw_k = 0.25 * k as f64;
|
||||
let expected_stretch = (f_ratio - 1.0).abs() * yaw_k;
|
||||
let err = angle_between(out.cameras.rotations[k], truth.rotations[k]);
|
||||
assert!(
|
||||
err < expected_stretch + 1.5e-4,
|
||||
"frame {k} off by {err} rad, {expected_stretch} of it the focal's"
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_frame_without_observations_is_refused() {
|
||||
let (truth, mut obs) = sweep(4, 0.3, 1400.0, 0.0);
|
||||
obs.retain(|o| o.i != 3 && o.j != 3);
|
||||
let err = adjust(truth, &obs, &AdjustOptions::default()).unwrap_err();
|
||||
assert!(matches!(err, PanoError::Geometry(_)));
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,336 @@
|
||||
//! Keypoints with descriptors, and the decoder that reads them out of
|
||||
//! XFeat's dense maps.
|
||||
//!
|
||||
//! The network (S15.2) produces three maps at an eighth of the input
|
||||
//! resolution and stops; everything from there to a list of keypoints is
|
||||
//! this file, in plain Rust, for the reason `dr-segment` decodes yolo26's
|
||||
//! heads itself: the post-processing is cheap, shape-dependent and exactly
|
||||
//! the kind of graph tract parses badly. It is a port of the reference
|
||||
//! `XFeat.detectAndCompute`, step for step, so that a keypoint here is the
|
||||
//! keypoint the paper's numbers were measured on.
|
||||
|
||||
/// One detected point, in the pixel coordinates of the image it was
|
||||
/// detected in, with the detector's confidence.
|
||||
#[derive(Debug, Clone, Copy, PartialEq)]
|
||||
pub struct Keypoint {
|
||||
pub x: f32,
|
||||
pub y: f32,
|
||||
/// The reliability the detector assigned; higher is better, and the
|
||||
/// scale is the detector's own — comparable within one model only.
|
||||
pub score: f32,
|
||||
}
|
||||
|
||||
/// The keypoints of one image and their descriptors.
|
||||
#[derive(Debug, Clone, PartialEq)]
|
||||
pub struct Features {
|
||||
pub keypoints: Vec<Keypoint>,
|
||||
/// `keypoints.len() × DESCRIPTOR_LEN`, each row L2-normalised, so that a
|
||||
/// dot product between two rows is their cosine similarity.
|
||||
pub descriptors: Vec<f32>,
|
||||
/// The image the coordinates are in.
|
||||
pub width: usize,
|
||||
pub height: usize,
|
||||
}
|
||||
|
||||
/// The length of one descriptor. XFeat's is 64; the matcher does not care
|
||||
/// what the number is, only that both sides agree.
|
||||
pub const DESCRIPTOR_LEN: usize = 64;
|
||||
|
||||
impl Features {
|
||||
pub fn len(&self) -> usize {
|
||||
self.keypoints.len()
|
||||
}
|
||||
|
||||
pub fn is_empty(&self) -> bool {
|
||||
self.keypoints.is_empty()
|
||||
}
|
||||
|
||||
pub fn descriptor(&self, i: usize) -> &[f32] {
|
||||
&self.descriptors[i * DESCRIPTOR_LEN..(i + 1) * DESCRIPTOR_LEN]
|
||||
}
|
||||
}
|
||||
|
||||
/// XFeat's three output maps, as the network hands them back.
|
||||
///
|
||||
/// All three are `channels × height × width` at an eighth of the input, in
|
||||
/// the NCHW order the ONNX export declares (`feats [1, 64, H/8, W/8]`,
|
||||
/// `keypoints [1, 65, H/8, W/8]`, `heatmap [1, 1, H/8, W/8]`).
|
||||
pub struct XFeatMaps<'a> {
|
||||
/// 64 channels: the dense descriptor field.
|
||||
pub feats: &'a [f32],
|
||||
/// 65 channels: for each 8×8 cell, a logit per position plus one for
|
||||
/// "no keypoint here".
|
||||
pub keypoints: &'a [f32],
|
||||
/// 1 channel: reliability.
|
||||
pub heatmap: &'a [f32],
|
||||
/// The maps' width and height (the input's, divided by eight).
|
||||
pub width: usize,
|
||||
pub height: usize,
|
||||
}
|
||||
|
||||
/// How the decoder picks keypoints.
|
||||
#[derive(Debug, Clone, Copy, PartialEq)]
|
||||
pub struct DecodeOptions {
|
||||
/// Keep at most this many, by score. The reference default is 4096.
|
||||
pub top_k: usize,
|
||||
/// A cell position's softmax probability must exceed this to be a
|
||||
/// keypoint at all. The reference default is 0.05.
|
||||
pub threshold: f32,
|
||||
/// Ignore keypoints within this many pixels of the map's edge. A frame
|
||||
/// padded into the detector's fixed input (`Gray::padded`) has a hard
|
||||
/// edge where the padding starts, and the detector fires on it.
|
||||
pub border: usize,
|
||||
}
|
||||
|
||||
impl Default for DecodeOptions {
|
||||
fn default() -> Self {
|
||||
DecodeOptions {
|
||||
top_k: 4096,
|
||||
threshold: 0.05,
|
||||
border: 4,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Decode keypoints and descriptors from the network's maps.
|
||||
///
|
||||
/// The reference, step for step:
|
||||
/// 1. softmax over the 65 logits of each cell, keep the 64 positions;
|
||||
/// 2. pixel-shuffle those into a full-resolution keypoint heatmap — channel
|
||||
/// `c` of cell `(cx, cy)` is pixel `(cx·8 + c%8, cy·8 + c/8)`;
|
||||
/// 3. 5×5 non-maximum suppression over that heatmap, above `threshold`;
|
||||
/// 4. score each survivor by its heatmap value times the reliability map
|
||||
/// sampled bilinearly at its position;
|
||||
/// 5. keep the `top_k` by score;
|
||||
/// 6. sample the descriptor field bilinearly at each and L2-normalise.
|
||||
///
|
||||
/// Bilinear where the reference samples the descriptor field bicubically:
|
||||
/// a quarter-pixel's difference in a field that is smooth by construction,
|
||||
/// and one interpolator rather than two to keep correct.
|
||||
pub fn decode_xfeat(maps: &XFeatMaps<'_>, opts: &DecodeOptions) -> Features {
|
||||
let (w8, h8) = (maps.width, maps.height);
|
||||
let (w, h) = (w8 * 8, h8 * 8);
|
||||
let cells = w8 * h8;
|
||||
debug_assert_eq!(maps.keypoints.len(), 65 * cells);
|
||||
debug_assert_eq!(maps.feats.len(), DESCRIPTOR_LEN * cells);
|
||||
debug_assert_eq!(maps.heatmap.len(), cells);
|
||||
|
||||
// 1 + 2: softmax per cell, scattered into the full-resolution heatmap.
|
||||
let mut heat = vec![0.0f32; w * h];
|
||||
for cy in 0..h8 {
|
||||
for cx in 0..w8 {
|
||||
let cell = cy * w8 + cx;
|
||||
let logit = |c: usize| maps.keypoints[c * cells + cell];
|
||||
let max = (0..65).map(logit).fold(f32::MIN, f32::max);
|
||||
let mut sum = 0.0f32;
|
||||
let mut exps = [0.0f32; 65];
|
||||
for (c, e) in exps.iter_mut().enumerate() {
|
||||
*e = (logit(c) - max).exp();
|
||||
sum += *e;
|
||||
}
|
||||
for (c, e) in exps.iter().enumerate().take(64) {
|
||||
let (dx, dy) = (c % 8, c / 8);
|
||||
heat[(cy * 8 + dy) * w + cx * 8 + dx] = e / sum;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// 3: a pixel survives if it is the maximum of its 5×5 neighbourhood and
|
||||
// above threshold. Ties go to every tied pixel, as the reference's
|
||||
// `x == max_pool(x)` does.
|
||||
let border = opts.border.max(2);
|
||||
let mut survivors: Vec<(usize, usize, f32)> = Vec::new();
|
||||
for y in border..h.saturating_sub(border) {
|
||||
for x in border..w.saturating_sub(border) {
|
||||
let v = heat[y * w + x];
|
||||
if v <= opts.threshold {
|
||||
continue;
|
||||
}
|
||||
let mut is_max = true;
|
||||
'nb: for ny in y - 2..=y + 2 {
|
||||
for nx in x - 2..=x + 2 {
|
||||
if heat[ny * w + nx] > v {
|
||||
is_max = false;
|
||||
break 'nb;
|
||||
}
|
||||
}
|
||||
}
|
||||
if is_max {
|
||||
survivors.push((x, y, v));
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// 4: heatmap value × reliability, the latter sampled at the keypoint's
|
||||
// position in map coordinates (`align_corners = False`: pixel `x` of the
|
||||
// full image is `x / 8 - 0.5` in the map).
|
||||
let sample = |field: &[f32], channels: usize, c: usize, x: f32, y: f32| -> f32 {
|
||||
let fx = (x / 8.0 - 0.5).clamp(0.0, (w8 - 1) as f32);
|
||||
let fy = (y / 8.0 - 0.5).clamp(0.0, (h8 - 1) as f32);
|
||||
let x0 = fx as usize;
|
||||
let y0 = fy as usize;
|
||||
let x1 = (x0 + 1).min(w8 - 1);
|
||||
let y1 = (y0 + 1).min(h8 - 1);
|
||||
let tx = fx - x0 as f32;
|
||||
let ty = fy - y0 as f32;
|
||||
let at = |xx: usize, yy: usize| field[c * (w8 * h8) + yy * w8 + xx];
|
||||
let _ = channels;
|
||||
let top = at(x0, y0) * (1.0 - tx) + at(x1, y0) * tx;
|
||||
let bot = at(x0, y1) * (1.0 - tx) + at(x1, y1) * tx;
|
||||
top * (1.0 - ty) + bot * ty
|
||||
};
|
||||
let mut scored: Vec<(usize, usize, f32)> = survivors
|
||||
.into_iter()
|
||||
.map(|(x, y, v)| {
|
||||
let r = sample(maps.heatmap, 1, 0, x as f32, y as f32);
|
||||
(x, y, v * r)
|
||||
})
|
||||
.collect();
|
||||
|
||||
// 5: best first, then cut. `sort_unstable_by` on a total order of the
|
||||
// score; NaN cannot occur — every input is a probability or a sigmoid.
|
||||
scored.sort_unstable_by(|a, b| b.2.total_cmp(&a.2));
|
||||
scored.truncate(opts.top_k);
|
||||
|
||||
// 6: descriptors.
|
||||
let mut keypoints = Vec::with_capacity(scored.len());
|
||||
let mut descriptors = Vec::with_capacity(scored.len() * DESCRIPTOR_LEN);
|
||||
for (x, y, score) in scored {
|
||||
let (xf, yf) = (x as f32, y as f32);
|
||||
let start = descriptors.len();
|
||||
for c in 0..DESCRIPTOR_LEN {
|
||||
descriptors.push(sample(maps.feats, DESCRIPTOR_LEN, c, xf, yf));
|
||||
}
|
||||
let norm = descriptors[start..]
|
||||
.iter()
|
||||
.map(|v| v * v)
|
||||
.sum::<f32>()
|
||||
.sqrt()
|
||||
.max(1e-12);
|
||||
for v in &mut descriptors[start..] {
|
||||
*v /= norm;
|
||||
}
|
||||
keypoints.push(Keypoint {
|
||||
x: xf,
|
||||
y: yf,
|
||||
score,
|
||||
});
|
||||
}
|
||||
|
||||
Features {
|
||||
keypoints,
|
||||
descriptors,
|
||||
width: w,
|
||||
height: h,
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
/// Maps for a `w8 × h8` grid where every cell says "no keypoint" except
|
||||
/// the listed ones, which put all their weight on one position.
|
||||
fn maps(w8: usize, h8: usize, hot: &[(usize, usize, usize)]) -> (Vec<f32>, Vec<f32>, Vec<f32>) {
|
||||
let cells = w8 * h8;
|
||||
let mut kp = vec![0.0f32; 65 * cells];
|
||||
// "None" strongly preferred everywhere.
|
||||
for cell in 0..cells {
|
||||
kp[64 * cells + cell] = 10.0;
|
||||
}
|
||||
for &(cx, cy, c) in hot {
|
||||
let cell = cy * w8 + cx;
|
||||
kp[64 * cells + cell] = 0.0;
|
||||
kp[c * cells + cell] = 10.0;
|
||||
}
|
||||
let heat = vec![0.5f32; cells];
|
||||
// Descriptors: channel c is constant c across the field, so any
|
||||
// sampled descriptor is the same known vector.
|
||||
let mut feats = vec![0.0f32; DESCRIPTOR_LEN * cells];
|
||||
for c in 0..DESCRIPTOR_LEN {
|
||||
for v in &mut feats[c * cells..(c + 1) * cells] {
|
||||
*v = c as f32;
|
||||
}
|
||||
}
|
||||
(feats, kp, heat)
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_hot_cell_position_becomes_a_keypoint_at_the_right_pixel() {
|
||||
// Cell (2, 1), channel 8*3 + 5 = 29 → pixel (2*8 + 5, 1*8 + 3).
|
||||
let (f, k, h) = maps(8, 8, &[(2, 1, 29)]);
|
||||
let out = decode_xfeat(
|
||||
&XFeatMaps {
|
||||
feats: &f,
|
||||
keypoints: &k,
|
||||
heatmap: &h,
|
||||
width: 8,
|
||||
height: 8,
|
||||
},
|
||||
&DecodeOptions::default(),
|
||||
);
|
||||
assert_eq!(out.len(), 1);
|
||||
assert_eq!((out.keypoints[0].x, out.keypoints[0].y), (21.0, 11.0));
|
||||
assert_eq!((out.width, out.height), (64, 64));
|
||||
// Score is the softmax weight (~1) times the reliability (0.5).
|
||||
assert!((out.keypoints[0].score - 0.5).abs() < 5e-3);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn descriptors_are_unit_length() {
|
||||
let (f, k, h) = maps(8, 8, &[(3, 3, 0), (5, 5, 63)]);
|
||||
let out = decode_xfeat(
|
||||
&XFeatMaps {
|
||||
feats: &f,
|
||||
keypoints: &k,
|
||||
heatmap: &h,
|
||||
width: 8,
|
||||
height: 8,
|
||||
},
|
||||
&DecodeOptions::default(),
|
||||
);
|
||||
assert_eq!(out.len(), 2);
|
||||
for i in 0..2 {
|
||||
let n: f32 = out.descriptor(i).iter().map(|v| v * v).sum();
|
||||
assert!((n - 1.0).abs() < 1e-5);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn top_k_keeps_the_best() {
|
||||
let (f, k, mut h) = maps(8, 8, &[(1, 1, 0), (3, 3, 0), (5, 5, 0)]);
|
||||
// Make cell (3, 3) the most reliable.
|
||||
h[3 * 8 + 3] = 0.9;
|
||||
let out = decode_xfeat(
|
||||
&XFeatMaps {
|
||||
feats: &f,
|
||||
keypoints: &k,
|
||||
heatmap: &h,
|
||||
width: 8,
|
||||
height: 8,
|
||||
},
|
||||
&DecodeOptions {
|
||||
top_k: 1,
|
||||
..Default::default()
|
||||
},
|
||||
);
|
||||
assert_eq!(out.len(), 1);
|
||||
assert_eq!((out.keypoints[0].x, out.keypoints[0].y), (24.0, 24.0));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn the_border_is_excluded() {
|
||||
let (f, k, h) = maps(8, 8, &[(0, 0, 0)]);
|
||||
let out = decode_xfeat(
|
||||
&XFeatMaps {
|
||||
feats: &f,
|
||||
keypoints: &k,
|
||||
heatmap: &h,
|
||||
width: 8,
|
||||
height: 8,
|
||||
},
|
||||
&DecodeOptions::default(),
|
||||
);
|
||||
assert!(out.is_empty());
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,793 @@
|
||||
//! TRACES: FR-MRG-4
|
||||
//! Filling a composite's uncovered border, tile by tile, with an inpainter.
|
||||
//!
|
||||
//! A merged panorama has a ragged border where no frame reached. FR-MRG-4
|
||||
//! crops it by default; this fills it instead, when the photographer asks,
|
||||
//! with pixels a model invents from the picture around them. Everything
|
||||
//! here is the geometry of that — which tiles to run, what context to hand
|
||||
//! the model, how to put its answers back — and none of it is the model:
|
||||
//! that is the [`Inpainter`] trait, with MI-GAN behind it in `migan.rs`
|
||||
//! and a fake in the tests.
|
||||
//!
|
||||
//! # Context across the edge
|
||||
//!
|
||||
//! An inpainting model is trained on holes *inside* pictures. A panorama's
|
||||
//! border is a hole at the picture's *edge*: real content on one side,
|
||||
//! nothing on the other, and a model given that invents a structure along
|
||||
//! the open side — streaks of road in the sky, on the first try
|
||||
//! (2026-09-19). So the known content is mirrored across the coverage
|
||||
//! edge, column by column for the top and bottom bands and row by row for
|
||||
//! the sides, into the hole and into a padding ring around the picture,
|
||||
//! and the ring is presented as *known*. The model then interpolates
|
||||
//! between real content and its mirror rather than extrapolating into
|
||||
//! nothing. The ring is cut off at the end.
|
||||
//!
|
||||
//! # Structure from far away, texture from near
|
||||
//!
|
||||
//! One tiled pass at the working resolution was not enough: a 512-px tile
|
||||
//! straddling the coverage edge sees a few hundred pixels of real content
|
||||
//! on one side and invents the rest from that, two neighbouring tiles
|
||||
//! invent differently, and the seams and the merge's own fringe leak into
|
||||
//! the fill. [`fill_border`] therefore runs in two stages. A **coarse**
|
||||
//! pass at a quarter of the size, where the whole border and hundreds of
|
||||
//! pixels of real context sit inside a handful of tiles, decides the
|
||||
//! structure — where the slope goes, where the sky stays sky. Then
|
||||
//! **fine** passes regenerate the hole in bands from the real edge
|
||||
//! outward: each band is the only unknown, with real content (or the band
|
||||
//! before, freshly textured) on its near side and the coarse fill,
|
||||
//! upsampled, on its far side — blurry, but the right structure — so the
|
||||
//! model generates texture and a transition, never a large hole from
|
||||
//! nothing.
|
||||
//!
|
||||
//! Tiles overlap by a third and are blended under a raised-cosine window,
|
||||
//! so the seams between tiles do not show; the model's answer replaces
|
||||
//! only the pixels that were unknown, and the picture itself is untouched.
|
||||
|
||||
use crate::PanoError;
|
||||
|
||||
/// A model that fills a square hole from its surroundings.
|
||||
pub trait Inpainter {
|
||||
/// The square tile it takes, in pixels.
|
||||
fn tile(&self) -> usize;
|
||||
|
||||
/// Fill one tile. `rgb` is `tile × tile × 3`, row-major, 0..1, with the
|
||||
/// unknown pixels' values meaningless; `known` is `tile × tile`. The
|
||||
/// result is `tile × tile × 3`, 0..1, of which only the unknown pixels
|
||||
/// are read.
|
||||
fn fill(&mut self, rgb: &[f32], known: &[bool]) -> Result<Vec<f32>, PanoError>;
|
||||
}
|
||||
|
||||
/// What a caller hears from [`fill_border`]: progress, for a page's bar,
|
||||
/// and — for whoever is looking at why a fill went wrong — each stage's
|
||||
/// picture as it lands. A plain `FnMut(usize, usize)` is an observer that
|
||||
/// hears only the progress.
|
||||
pub trait Observer {
|
||||
/// `(done, total)` tiles, the total an estimate until the last band.
|
||||
fn progress(&mut self, done: usize, total: usize);
|
||||
/// A stage's result, `width × height × 3`: `coarse` (at the coarse
|
||||
/// size), `band-N` after each fine band, `feathered` at the end.
|
||||
fn stage(&mut self, _name: &str, _rgb: &[f32], _width: usize, _height: usize) {}
|
||||
}
|
||||
|
||||
impl<F: FnMut(usize, usize)> Observer for F {
|
||||
fn progress(&mut self, done: usize, total: usize) {
|
||||
self(done, total)
|
||||
}
|
||||
}
|
||||
|
||||
/// How far the picture is extended with mirrored content before tiling.
|
||||
/// Half a tile: enough that a hole at the edge sits well inside a tile.
|
||||
pub const RING: usize = 256;
|
||||
|
||||
/// The fill's knobs, in pixels of the working image. The defaults are
|
||||
/// what the fixture panorama looked best with on 2026-09-19; the merge
|
||||
/// page exposes every one of them while the fill is experimental, so a
|
||||
/// bad corner can be worked on from the picture rather than the code.
|
||||
#[derive(Debug, Clone, Copy, PartialEq)]
|
||||
pub struct Params {
|
||||
/// The coarse pass's reduction: 1 skips it.
|
||||
pub coarse: usize,
|
||||
/// The fine passes' band width.
|
||||
pub band: usize,
|
||||
/// How deep into the picture the mirrored context reaches. A plain
|
||||
/// reflection of a deep hole pulls in whatever is that far from the
|
||||
/// edge — a ridge, a peak — and the model, told that is what lies
|
||||
/// beyond, paints it upside down. Folding the reflection within this
|
||||
/// band keeps the ring looking like the edge it continues (sky beside
|
||||
/// sky, grass beside grass) and nothing further away.
|
||||
pub mirror_depth: usize,
|
||||
/// How far inside the real edge the fill also regenerates, the two
|
||||
/// blended by distance. A hard cut between real pixels and invented
|
||||
/// ones is a line whatever the fill's quality; blended over this many
|
||||
/// pixels it is not. Zero is the hard cut.
|
||||
pub feather: usize,
|
||||
/// The step between tiles, at most the tile; two thirds of it usual.
|
||||
pub stride: usize,
|
||||
}
|
||||
|
||||
impl Default for Params {
|
||||
fn default() -> Self {
|
||||
Params {
|
||||
coarse: 4,
|
||||
band: 96,
|
||||
mirror_depth: 48,
|
||||
feather: 24,
|
||||
stride: 384,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Fill the unknown pixels of `rgb` (`width × height × 3`, 0..1) in place:
|
||||
/// the coarse pass, then the fine bands, then the seam feathered over
|
||||
/// `feather` pixels inside the real edge. Returns the tiles run.
|
||||
///
|
||||
/// `known` is `width × height`. `observer` hears the progress and, if it
|
||||
/// cares, each stage.
|
||||
pub fn fill_border(
|
||||
rgb: &mut [f32],
|
||||
width: usize,
|
||||
height: usize,
|
||||
known: &[bool],
|
||||
model: &mut dyn Inpainter,
|
||||
params: Params,
|
||||
observer: &mut dyn Observer,
|
||||
) -> Result<usize, PanoError> {
|
||||
let Params {
|
||||
coarse: q,
|
||||
band,
|
||||
mirror_depth,
|
||||
feather,
|
||||
stride,
|
||||
} = params;
|
||||
let q = q.max(1);
|
||||
let band = band.max(8);
|
||||
let mirror_depth = mirror_depth.max(1);
|
||||
if width == 0 || height == 0 || rgb.len() != width * height * 3 || known.len() != width * height
|
||||
{
|
||||
return Err(PanoError::Input("fill: buffer sizes disagree".into()));
|
||||
}
|
||||
if known.iter().all(|&k| k) {
|
||||
return Ok(0);
|
||||
}
|
||||
// The fill regenerates a margin inside the real edge too, and the
|
||||
// result is blended with the real pixels across it at the end.
|
||||
let real = rgb.to_vec();
|
||||
let outer = known.to_vec();
|
||||
let mut inner = known.to_vec();
|
||||
erode(&mut inner, width, height, feather);
|
||||
let known = &inner[..];
|
||||
let mut done = 0usize;
|
||||
|
||||
// Coarse: a fraction of the size, unknown where any pixel of the cell was.
|
||||
let (cw, ch) = ((width / q).max(1), (height / q).max(1));
|
||||
let mut coarse = vec![0.0f32; cw * ch * 3];
|
||||
let mut cknown = vec![true; cw * ch];
|
||||
for y in 0..ch {
|
||||
for x in 0..cw {
|
||||
let mut sum = [0.0f32; 3];
|
||||
let mut n = 0.0f32;
|
||||
let mut all_known = true;
|
||||
for dy in 0..q {
|
||||
for dx in 0..q {
|
||||
let (sx, sy) = ((x * q + dx).min(width - 1), (y * q + dy).min(height - 1));
|
||||
let i = sy * width + sx;
|
||||
all_known &= known[i];
|
||||
for c in 0..3 {
|
||||
sum[c] += rgb[i * 3 + c];
|
||||
}
|
||||
n += 1.0;
|
||||
}
|
||||
}
|
||||
for c in 0..3 {
|
||||
coarse[(y * cw + x) * 3 + c] = sum[c] / n;
|
||||
}
|
||||
cknown[y * cw + x] = all_known;
|
||||
}
|
||||
}
|
||||
let estimate = |tiles: usize| tiles * 4;
|
||||
done += fill_once(
|
||||
&mut coarse,
|
||||
cw,
|
||||
ch,
|
||||
&cknown,
|
||||
model,
|
||||
stride,
|
||||
mirror_depth,
|
||||
|n, t| observer.progress(n, estimate(t)),
|
||||
)?;
|
||||
observer.stage("coarse", &coarse, cw, ch);
|
||||
|
||||
// The hole starts as the coarse structure, upsampled.
|
||||
for y in 0..height {
|
||||
for x in 0..width {
|
||||
let i = y * width + x;
|
||||
if known[i] {
|
||||
continue;
|
||||
}
|
||||
let fx = ((x as f32 + 0.5) / q as f32 - 0.5).clamp(0.0, (cw - 1) as f32);
|
||||
let fy = ((y as f32 + 0.5) / q as f32 - 0.5).clamp(0.0, (ch - 1) as f32);
|
||||
let (x0, y0) = (fx as usize, fy as usize);
|
||||
let (x1, y1) = ((x0 + 1).min(cw - 1), (y0 + 1).min(ch - 1));
|
||||
let (tx, ty) = (fx - x0 as f32, fy - y0 as f32);
|
||||
for c in 0..3 {
|
||||
let at = |xx: usize, yy: usize| coarse[(yy * cw + xx) * 3 + c];
|
||||
rgb[i * 3 + c] = (at(x0, y0) * (1.0 - tx) + at(x1, y0) * tx) * (1.0 - ty)
|
||||
+ (at(x0, y1) * (1.0 - tx) + at(x1, y1) * tx) * ty;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Fine, in bands from the edge outward.
|
||||
let dist = distance_to_known(known, width, height);
|
||||
let mut band_known = vec![true; width * height];
|
||||
let mut b = 0usize;
|
||||
loop {
|
||||
let lo = (b * band).saturating_sub(band / 2) as f32;
|
||||
let hi = ((b + 1) * band) as f32;
|
||||
let mut any = false;
|
||||
for i in 0..width * height {
|
||||
let in_band = !known[i] && dist[i] > lo && dist[i] <= hi;
|
||||
band_known[i] = !in_band;
|
||||
any |= in_band;
|
||||
}
|
||||
if !any {
|
||||
break;
|
||||
}
|
||||
let before = done;
|
||||
done += fill_once(
|
||||
rgb,
|
||||
width,
|
||||
height,
|
||||
&band_known,
|
||||
model,
|
||||
stride,
|
||||
mirror_depth,
|
||||
|n, t| observer.progress(before + n, before + estimate(t)),
|
||||
)?;
|
||||
observer.stage(&format!("band-{b}"), rgb, width, height);
|
||||
b += 1;
|
||||
}
|
||||
|
||||
// The seam: across the margin, real on the inside, invented on the
|
||||
// outside, a smooth ramp between by distance from the true hole.
|
||||
if feather > 0 {
|
||||
let to_hole =
|
||||
distance_to_known(&outer.iter().map(|k| !k).collect::<Vec<_>>(), width, height);
|
||||
for i in 0..width * height {
|
||||
if !outer[i] || known[i] {
|
||||
continue;
|
||||
}
|
||||
// In the margin: outer says known, inner says not.
|
||||
let t = (to_hole[i] / feather as f32).clamp(0.0, 1.0);
|
||||
let t = t * t * (3.0 - 2.0 * t);
|
||||
for c in 0..3 {
|
||||
rgb[i * 3 + c] = rgb[i * 3 + c] * (1.0 - t) + real[i * 3 + c] * t;
|
||||
}
|
||||
}
|
||||
}
|
||||
observer.stage("feathered", rgb, width, height);
|
||||
observer.progress(done, done);
|
||||
Ok(done)
|
||||
}
|
||||
|
||||
/// One tiled pass: every unknown pixel regenerated from the tiles that
|
||||
/// touch it, the rest kept. Returns the tiles run.
|
||||
#[allow(clippy::too_many_arguments)]
|
||||
fn fill_once(
|
||||
rgb: &mut [f32],
|
||||
width: usize,
|
||||
height: usize,
|
||||
known: &[bool],
|
||||
model: &mut dyn Inpainter,
|
||||
stride: usize,
|
||||
mirror_depth: usize,
|
||||
mut progress: impl FnMut(usize, usize),
|
||||
) -> Result<usize, PanoError> {
|
||||
let t = model.tile();
|
||||
if t == 0 || known.iter().all(|&k| k) {
|
||||
return Ok(0);
|
||||
}
|
||||
|
||||
// The padded canvas with mirrored context, and the hole within it.
|
||||
let ctx = MirroredContext::build(rgb, width, height, known, mirror_depth);
|
||||
let (pw, ph) = (ctx.width, ctx.height);
|
||||
|
||||
// Tiles that touch the hole, on a grid that reaches both far edges.
|
||||
let starts = |n: usize| -> Vec<usize> {
|
||||
if n <= t {
|
||||
return vec![0];
|
||||
}
|
||||
let mut v: Vec<usize> = (0..=n - t).step_by(stride.clamp(1, t)).collect();
|
||||
if *v.last().unwrap_or(&0) != n - t {
|
||||
v.push(n - t);
|
||||
}
|
||||
v
|
||||
};
|
||||
let ys = starts(ph);
|
||||
let xs = starts(pw);
|
||||
let mut tiles = Vec::new();
|
||||
for &y in &ys {
|
||||
for &x in &xs {
|
||||
if y + t > ph || x + t > pw {
|
||||
continue;
|
||||
}
|
||||
let touches =
|
||||
(y..y + t).any(|yy| ctx.hole[yy * pw + x..yy * pw + x + t].iter().any(|&h| h));
|
||||
if touches {
|
||||
tiles.push((x, y));
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Raised-cosine window, so overlapping tiles blend.
|
||||
let hann: Vec<f32> = (0..t)
|
||||
.map(|i| {
|
||||
let s = ((i as f32 + 1.0) / (t as f32 + 1.0) * std::f32::consts::PI).sin();
|
||||
s * s + 1e-3
|
||||
})
|
||||
.collect();
|
||||
|
||||
let mut acc = vec![0.0f32; pw * ph * 3];
|
||||
let mut wsum = vec![0.0f32; pw * ph];
|
||||
let mut tile_rgb = vec![0.0f32; t * t * 3];
|
||||
let mut tile_known = vec![false; t * t];
|
||||
let total = tiles.len();
|
||||
for (n, &(x, y)) in tiles.iter().enumerate() {
|
||||
progress(n, total);
|
||||
for r in 0..t {
|
||||
let src = ((y + r) * pw + x) * 3;
|
||||
tile_rgb[r * t * 3..(r + 1) * t * 3].copy_from_slice(&ctx.rgb[src..src + t * 3]);
|
||||
let ks = (y + r) * pw + x;
|
||||
for c in 0..t {
|
||||
tile_known[r * t + c] = !ctx.hole[ks + c];
|
||||
}
|
||||
}
|
||||
let out = model.fill(&tile_rgb, &tile_known)?;
|
||||
if out.len() != t * t * 3 {
|
||||
return Err(PanoError::Model(format!(
|
||||
"the inpainter returned {} values for a {t}×{t} tile",
|
||||
out.len()
|
||||
)));
|
||||
}
|
||||
for r in 0..t {
|
||||
for c in 0..t {
|
||||
let w = hann[r] * hann[c];
|
||||
let p = (y + r) * pw + (x + c);
|
||||
for ch in 0..3 {
|
||||
acc[p * 3 + ch] += out[(r * t + c) * 3 + ch] * w;
|
||||
}
|
||||
wsum[p] += w;
|
||||
}
|
||||
}
|
||||
}
|
||||
progress(total, total);
|
||||
|
||||
// Back into the picture: only the unknown pixels change.
|
||||
for yy in 0..height {
|
||||
for xx in 0..width {
|
||||
let i = yy * width + xx;
|
||||
if known[i] {
|
||||
continue;
|
||||
}
|
||||
let p = (yy + RING) * pw + (xx + RING);
|
||||
if wsum[p] > 0.0 {
|
||||
for ch in 0..3 {
|
||||
rgb[i * 3 + ch] = (acc[p * 3 + ch] / wsum[p]).clamp(0.0, 1.0);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
Ok(total)
|
||||
}
|
||||
|
||||
/// Shrink `known` by `iterations` pixels on every side, in place.
|
||||
///
|
||||
/// The merge's coverage edge carries a fringe — the last partly-covered
|
||||
/// pixels of a frame, and whatever the renderer did at the boundary — and
|
||||
/// a fill that stops exactly at the coverage bit leaves it as a dark line
|
||||
/// along the seam. Eight pixels at a quarter of the composite's resolution
|
||||
/// was what it took on the fixture.
|
||||
pub fn erode(known: &mut [bool], width: usize, height: usize, iterations: usize) {
|
||||
let mut next = known.to_vec();
|
||||
for _ in 0..iterations {
|
||||
for y in 0..height {
|
||||
for x in 0..width {
|
||||
let i = y * width + x;
|
||||
if !known[i] {
|
||||
continue;
|
||||
}
|
||||
let edge = x == 0
|
||||
|| y == 0
|
||||
|| x + 1 == width
|
||||
|| y + 1 == height
|
||||
|| !known[i - 1]
|
||||
|| !known[i + 1]
|
||||
|| !known[i - width]
|
||||
|| !known[i + width];
|
||||
next[i] = !edge;
|
||||
}
|
||||
}
|
||||
known.copy_from_slice(&next);
|
||||
}
|
||||
}
|
||||
|
||||
/// Distance from each pixel to the nearest known one, by two chamfer
|
||||
/// sweeps — within a few percent of Euclidean, and enough to cut bands.
|
||||
fn distance_to_known(known: &[bool], width: usize, height: usize) -> Vec<f32> {
|
||||
let inf = (width + height) as f32;
|
||||
let mut d: Vec<f32> = known.iter().map(|&k| if k { 0.0 } else { inf }).collect();
|
||||
let (a, b) = (1.0f32, std::f32::consts::SQRT_2);
|
||||
for y in 0..height {
|
||||
for x in 0..width {
|
||||
let i = y * width + x;
|
||||
let mut v = d[i];
|
||||
if x > 0 {
|
||||
v = v.min(d[i - 1] + a);
|
||||
}
|
||||
if y > 0 {
|
||||
v = v.min(d[i - width] + a);
|
||||
if x > 0 {
|
||||
v = v.min(d[i - width - 1] + b);
|
||||
}
|
||||
if x + 1 < width {
|
||||
v = v.min(d[i - width + 1] + b);
|
||||
}
|
||||
}
|
||||
d[i] = v;
|
||||
}
|
||||
}
|
||||
for y in (0..height).rev() {
|
||||
for x in (0..width).rev() {
|
||||
let i = y * width + x;
|
||||
let mut v = d[i];
|
||||
if x + 1 < width {
|
||||
v = v.min(d[i + 1] + a);
|
||||
}
|
||||
if y + 1 < height {
|
||||
v = v.min(d[i + width] + a);
|
||||
if x + 1 < width {
|
||||
v = v.min(d[i + width + 1] + b);
|
||||
}
|
||||
if x > 0 {
|
||||
v = v.min(d[i + width - 1] + b);
|
||||
}
|
||||
}
|
||||
d[i] = v;
|
||||
}
|
||||
}
|
||||
d
|
||||
}
|
||||
|
||||
/// Distance beyond the edge to distance inside it, folded within `depth`
|
||||
/// ([`Params::mirror_depth`]): a triangle wave, so the band is read
|
||||
/// forward and back rather than clamped to one row.
|
||||
fn fold(d: usize, depth: usize) -> usize {
|
||||
let period = 2 * depth;
|
||||
let r = d % period;
|
||||
if r <= depth {
|
||||
r
|
||||
} else {
|
||||
period - r
|
||||
}
|
||||
}
|
||||
|
||||
/// The picture on a canvas `RING` wider on every side, with the hole and
|
||||
/// the ring filled by mirroring the known content across the coverage
|
||||
/// edge — the nearest `depth` of it, folded — and the hole, the
|
||||
/// original unknown and nothing else, marked.
|
||||
struct MirroredContext {
|
||||
width: usize,
|
||||
height: usize,
|
||||
rgb: Vec<f32>,
|
||||
hole: Vec<bool>,
|
||||
}
|
||||
|
||||
impl MirroredContext {
|
||||
fn build(rgb: &[f32], width: usize, height: usize, known: &[bool], depth: usize) -> Self {
|
||||
let fold = |d: usize| fold(d, depth);
|
||||
let (pw, ph) = (width + 2 * RING, height + 2 * RING);
|
||||
let mut canvas = vec![0.0f32; pw * ph * 3];
|
||||
let mut kn = vec![false; pw * ph];
|
||||
let mut hole = vec![false; pw * ph];
|
||||
for y in 0..height {
|
||||
for x in 0..width {
|
||||
let i = y * width + x;
|
||||
let p = (y + RING) * pw + (x + RING);
|
||||
canvas[p * 3..p * 3 + 3].copy_from_slice(&rgb[i * 3..i * 3 + 3]);
|
||||
kn[p] = known[i];
|
||||
hole[p] = !known[i];
|
||||
}
|
||||
}
|
||||
|
||||
// Per column: mirror across the first and last known row.
|
||||
for x in 0..pw {
|
||||
let first = (0..ph).find(|&y| kn[y * pw + x]);
|
||||
let Some(first) = first else { continue };
|
||||
let last = (0..ph).rev().find(|&y| kn[y * pw + x]).unwrap_or(first);
|
||||
for y in 0..first {
|
||||
let m = (first + fold(first - y)).min(last);
|
||||
let (d, s) = ((y * pw + x) * 3, (m * pw + x) * 3);
|
||||
canvas.copy_within(s..s + 3, d);
|
||||
}
|
||||
for y in last + 1..ph {
|
||||
let m = last.saturating_sub(fold(y - last)).max(first);
|
||||
let (d, s) = ((y * pw + x) * 3, (m * pw + x) * 3);
|
||||
canvas.copy_within(s..s + 3, d);
|
||||
}
|
||||
}
|
||||
// Per row, for the sides, over what is there now.
|
||||
for y in 0..ph {
|
||||
let first = (0..pw).find(|&x| kn[y * pw + x]);
|
||||
let Some(first) = first else { continue };
|
||||
let last = (0..pw).rev().find(|&x| kn[y * pw + x]).unwrap_or(first);
|
||||
for x in 0..first {
|
||||
let m = (first + fold(first - x)).min(last);
|
||||
let (d, s) = ((y * pw + x) * 3, (y * pw + m) * 3);
|
||||
canvas.copy_within(s..s + 3, d);
|
||||
}
|
||||
for x in last + 1..pw {
|
||||
let m = last.saturating_sub(fold(x - last)).max(first);
|
||||
let (d, s) = ((y * pw + x) * 3, (y * pw + m) * 3);
|
||||
canvas.copy_within(s..s + 3, d);
|
||||
}
|
||||
}
|
||||
MirroredContext {
|
||||
width: pw,
|
||||
height: ph,
|
||||
rgb: canvas,
|
||||
hole,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
/// The tests' small pictures: a 48-px stride, a given feather.
|
||||
fn test_params(feather: usize) -> Params {
|
||||
Params {
|
||||
stride: 48,
|
||||
feather,
|
||||
..Params::default()
|
||||
}
|
||||
}
|
||||
|
||||
/// Paints every unknown pixel a fixed grey and copies the known ones,
|
||||
/// and remembers what it was shown.
|
||||
struct Flat {
|
||||
tile: usize,
|
||||
seen: Vec<(Vec<f32>, Vec<bool>)>,
|
||||
}
|
||||
|
||||
impl Inpainter for Flat {
|
||||
fn tile(&self) -> usize {
|
||||
self.tile
|
||||
}
|
||||
fn fill(&mut self, rgb: &[f32], known: &[bool]) -> Result<Vec<f32>, PanoError> {
|
||||
self.seen.push((rgb.to_vec(), known.to_vec()));
|
||||
Ok(rgb
|
||||
.chunks_exact(3)
|
||||
.zip(known)
|
||||
.flat_map(|(p, &k)| if k { [p[0], p[1], p[2]] } else { [0.5; 3] })
|
||||
.collect())
|
||||
}
|
||||
}
|
||||
|
||||
fn picture(w: usize, h: usize, border: usize) -> (Vec<f32>, Vec<bool>) {
|
||||
let mut rgb = vec![0.0; w * h * 3];
|
||||
let mut known = vec![false; w * h];
|
||||
for y in 0..h {
|
||||
for x in 0..w {
|
||||
let i = y * w + x;
|
||||
if y >= border && y < h - border {
|
||||
known[i] = true;
|
||||
rgb[i * 3] = x as f32 / w as f32;
|
||||
rgb[i * 3 + 1] = y as f32 / h as f32;
|
||||
rgb[i * 3 + 2] = 0.25;
|
||||
}
|
||||
}
|
||||
}
|
||||
(rgb, known)
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn unknown_pixels_take_the_model_and_known_ones_do_not_move() {
|
||||
let (mut rgb, known) = picture(300, 200, 20);
|
||||
let before = rgb.clone();
|
||||
let mut model = Flat {
|
||||
tile: 64,
|
||||
seen: Vec::new(),
|
||||
};
|
||||
let tiles = fill_border(
|
||||
&mut rgb,
|
||||
300,
|
||||
200,
|
||||
&known,
|
||||
&mut model,
|
||||
test_params(0),
|
||||
&mut |_, _| {},
|
||||
)
|
||||
.unwrap();
|
||||
assert!(tiles > 0);
|
||||
for i in 0..300 * 200 {
|
||||
if known[i] {
|
||||
assert_eq!(&rgb[i * 3..i * 3 + 3], &before[i * 3..i * 3 + 3]);
|
||||
} else {
|
||||
for c in 0..3 {
|
||||
assert!((rgb[i * 3 + c] - 0.5).abs() < 1e-4, "pixel {i}");
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn the_model_is_shown_mirrored_context_not_black() {
|
||||
let (mut rgb, known) = picture(300, 200, 20);
|
||||
let mut model = Flat {
|
||||
tile: 64,
|
||||
seen: Vec::new(),
|
||||
};
|
||||
fill_border(
|
||||
&mut rgb,
|
||||
300,
|
||||
200,
|
||||
&known,
|
||||
&mut model,
|
||||
test_params(0),
|
||||
&mut |_, _| {},
|
||||
)
|
||||
.unwrap();
|
||||
for (tile_rgb, tile_known) in &model.seen {
|
||||
let known_non_black = tile_rgb
|
||||
.chunks_exact(3)
|
||||
.zip(tile_known)
|
||||
.filter(|(_, &k)| k)
|
||||
.any(|(p, _)| p.iter().any(|v| *v > 0.0));
|
||||
assert!(known_non_black);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn the_fine_passes_run_in_bands_after_the_coarse_one() {
|
||||
// A 150-tall hole above and below a picture: the coarse pass sees
|
||||
// it at a quarter; the fine passes need two bands of BAND pixels.
|
||||
let (mut rgb, known) = picture(200, 500, 150);
|
||||
let mut model = Flat {
|
||||
tile: 64,
|
||||
seen: Vec::new(),
|
||||
};
|
||||
fill_border(
|
||||
&mut rgb,
|
||||
200,
|
||||
500,
|
||||
&known,
|
||||
&mut model,
|
||||
test_params(0),
|
||||
&mut |_, _| {},
|
||||
)
|
||||
.unwrap();
|
||||
assert!(model.seen.len() > 4);
|
||||
// Every unknown pixel was reached.
|
||||
for i in 0..200 * 500 {
|
||||
if !known[i] {
|
||||
assert!((rgb[i * 3] - 0.5).abs() < 1e-4, "pixel {i}");
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn the_seam_ramps_from_real_to_invented_across_the_feather() {
|
||||
let (mut rgb, known) = picture(300, 200, 20);
|
||||
let before = rgb.clone();
|
||||
let mut model = Flat {
|
||||
tile: 64,
|
||||
seen: Vec::new(),
|
||||
};
|
||||
fill_border(
|
||||
&mut rgb,
|
||||
300,
|
||||
200,
|
||||
&known,
|
||||
&mut model,
|
||||
test_params(8),
|
||||
&mut |_, _| {},
|
||||
)
|
||||
.unwrap();
|
||||
// Row 20 is the real edge; the margin runs to row 27. At the edge
|
||||
// the value is the model's grey, eight rows in it is the picture's.
|
||||
let at = |y: usize| rgb[(y * 300 + 150) * 3 + 2];
|
||||
assert!((at(20) - 0.5).abs() < 0.05, "{}", at(20));
|
||||
assert!((at(29) - before[(29 * 300 + 150) * 3 + 2]).abs() < 1e-4);
|
||||
let (lo, hi) = (at(20).min(at(29)), at(20).max(at(29)));
|
||||
assert!(
|
||||
at(23) > lo + 0.02 && at(23) < hi - 0.02,
|
||||
"{} between {lo} and {hi}",
|
||||
at(23)
|
||||
);
|
||||
// The hole itself is the model's.
|
||||
assert!((at(5) - 0.5).abs() < 1e-4);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn erosion_shrinks_the_known_region_from_every_edge() {
|
||||
let (_, mut known) = picture(20, 20, 4);
|
||||
erode(&mut known, 20, 20, 2);
|
||||
assert!(known[8 * 20 + 10]);
|
||||
assert!(!known[5 * 20 + 10]);
|
||||
assert!(!known[8 * 20 + 1]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn distance_counts_pixels_from_the_known_region() {
|
||||
let (_, known) = picture(20, 20, 4);
|
||||
let d = distance_to_known(&known, 20, 20);
|
||||
assert_eq!(d[4 * 20 + 10], 0.0);
|
||||
assert!((d[3 * 20 + 10] - 1.0).abs() < 1e-6);
|
||||
assert!((d[10] - 4.0).abs() < 1e-6);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_fully_covered_picture_runs_nothing() {
|
||||
let (mut rgb, known) = picture(100, 100, 0);
|
||||
let mut model = Flat {
|
||||
tile: 64,
|
||||
seen: Vec::new(),
|
||||
};
|
||||
assert_eq!(
|
||||
fill_border(
|
||||
&mut rgb,
|
||||
100,
|
||||
100,
|
||||
&known,
|
||||
&mut model,
|
||||
test_params(0),
|
||||
&mut |_, _| {}
|
||||
)
|
||||
.unwrap(),
|
||||
0
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn the_context_mirrors_the_top_rows_upward() {
|
||||
let (rgb, known) = picture(40, 30, 5);
|
||||
let ctx = MirroredContext::build(&rgb, 40, 30, &known, 48);
|
||||
let x = RING + 10;
|
||||
let first = RING + 5;
|
||||
for k in 1..=4 {
|
||||
let above = ((first - k) * ctx.width + x) * 3;
|
||||
let mirror = ((first + k) * ctx.width + x) * 3;
|
||||
assert_eq!(&ctx.rgb[above..above + 3], &ctx.rgb[mirror..mirror + 3]);
|
||||
}
|
||||
assert!(ctx.hole[(RING + 2) * ctx.width + x]);
|
||||
assert!(!ctx.hole[(RING - 2) * ctx.width + x]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn the_mirror_reaches_no_deeper_than_its_band() {
|
||||
// A ridge 200 rows in must not appear in the ring: beyond the band
|
||||
// the reflection folds back towards the edge rather than on into
|
||||
// the picture.
|
||||
let (mut rgb, known) = picture(40, 400, 5);
|
||||
let ridge = 5 + 200;
|
||||
for x in 0..40 {
|
||||
rgb[(ridge * 40 + x) * 3..(ridge * 40 + x) * 3 + 3].copy_from_slice(&[0.9, 0.1, 0.1]);
|
||||
}
|
||||
let ctx = MirroredContext::build(&rgb, 40, 400, &known, 48);
|
||||
let x = RING + 10;
|
||||
for y in 0..RING + 5 {
|
||||
let p = (y * ctx.width + x) * 3;
|
||||
assert!(
|
||||
ctx.rgb[p] < 0.5,
|
||||
"row {y} of the ring shows the ridge ({:?})",
|
||||
&ctx.rgb[p..p + 3]
|
||||
);
|
||||
}
|
||||
assert_eq!(fold(0, 48), 0);
|
||||
assert_eq!(fold(48, 48), 48);
|
||||
assert_eq!(fold(58, 48), 38);
|
||||
assert_eq!(fold(96, 48), 0);
|
||||
assert_eq!(fold(99, 48), 3);
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,367 @@
|
||||
//! Pairwise geometry: a homography between two frames, found robustly.
|
||||
//!
|
||||
//! Two frames of a panorama are related by a rotation, and a rotation seen
|
||||
//! through one lens is a homography of the image plane — `H = K R Kᵀ⁻¹`. The
|
||||
//! homography is estimated first, from matches, because it does not need
|
||||
//! the focal length; the focal length is then *read off* it (§ below), and
|
||||
//! the rotation follows from both. This is the order Brown & Lowe (2007)
|
||||
//! and OpenCV's stitcher use, and it is what makes the pipeline work when
|
||||
//! EXIF says nothing about the lens.
|
||||
//!
|
||||
//! Coordinates throughout are **centred**: the principal point is the
|
||||
//! origin. The focal formulae assume it, and centring before the DLT also
|
||||
//! conditions the linear system — Hartley's normalisation, done once by the
|
||||
//! caller rather than inside every solve.
|
||||
|
||||
use crate::linalg::{DMat, Mat3, Vec3};
|
||||
|
||||
/// A point in one image, centred on the principal point.
|
||||
pub type Point = (f64, f64);
|
||||
|
||||
/// Apply a homography to a point.
|
||||
pub fn apply(h: &Mat3, p: Point) -> Option<Point> {
|
||||
let v = *h * Vec3::new(p.0, p.1, 1.0);
|
||||
if v.z().abs() < 1e-12 {
|
||||
return None;
|
||||
}
|
||||
Some((v.x() / v.z(), v.y() / v.z()))
|
||||
}
|
||||
|
||||
/// Least-squares homography from at least four correspondences by the
|
||||
/// direct linear transform, with `h33` fixed at 1.
|
||||
///
|
||||
/// Fixing `h33` turns the homogeneous 8×9 system into an ordinary 8-unknown
|
||||
/// least-squares problem that the normal equations and a Cholesky
|
||||
/// factorisation solve without an SVD. The one homography it cannot
|
||||
/// represent — `h33 = 0`, a point at the origin mapped to infinity — does
|
||||
/// not occur between overlapping frames of one scene.
|
||||
///
|
||||
/// The points should be scaled to order one (divide by the focal length or
|
||||
/// the image size) before calling: the normal equations square the
|
||||
/// conditioning, and pixel coordinates in the thousands make them singular
|
||||
/// in `f64`.
|
||||
pub fn dlt(pairs: &[(Point, Point)]) -> Option<Mat3> {
|
||||
if pairs.len() < 4 {
|
||||
return None;
|
||||
}
|
||||
// Each pair gives two rows of A h = b with h = (h11..h32).
|
||||
// x' = (h11 x + h12 y + h13) / (h31 x + h32 y + 1)
|
||||
// → h11 x + h12 y + h13 - h31 x x' - h32 y x' = x'
|
||||
let mut ata = DMat::zeros(8);
|
||||
let mut atb = [0.0f64; 8];
|
||||
for &((x, y), (xp, yp)) in pairs {
|
||||
let rows: [([f64; 8], f64); 2] = [
|
||||
([x, y, 1.0, 0.0, 0.0, 0.0, -x * xp, -y * xp], xp),
|
||||
([0.0, 0.0, 0.0, x, y, 1.0, -x * yp, -y * yp], yp),
|
||||
];
|
||||
for (a, b) in rows {
|
||||
for i in 0..8 {
|
||||
atb[i] += a[i] * b;
|
||||
for j in 0..8 {
|
||||
ata[(i, j)] += a[i] * a[j];
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
let h = ata.solve_spd(&atb)?;
|
||||
Some(Mat3([
|
||||
[h[0], h[1], h[2]],
|
||||
[h[3], h[4], h[5]],
|
||||
[h[6], h[7], 1.0],
|
||||
]))
|
||||
}
|
||||
|
||||
/// A homography with the correspondences that agree with it.
|
||||
#[derive(Debug, Clone, PartialEq)]
|
||||
pub struct RobustHomography {
|
||||
pub h: Mat3,
|
||||
/// Indices into the input pairs.
|
||||
pub inliers: Vec<usize>,
|
||||
}
|
||||
|
||||
/// RANSAC over [`dlt`] on four-point samples, then a final least-squares
|
||||
/// fit over every inlier.
|
||||
///
|
||||
/// `threshold` is the reprojection distance, in the same units as the
|
||||
/// points, within which a pair counts as agreeing. The iteration count
|
||||
/// adapts to the inlier ratio found so far in the usual way, capped at
|
||||
/// `max_iterations`. `seed` makes a run reproducible (NFR-MRG-2): the
|
||||
/// sampling is a small linear congruential generator, not the system's.
|
||||
pub fn ransac_homography(
|
||||
pairs: &[(Point, Point)],
|
||||
threshold: f64,
|
||||
max_iterations: usize,
|
||||
seed: u64,
|
||||
) -> Option<RobustHomography> {
|
||||
if pairs.len() < 4 {
|
||||
return None;
|
||||
}
|
||||
let n = pairs.len();
|
||||
let thr2 = threshold * threshold;
|
||||
let mut rng = Lcg(seed);
|
||||
let mut best: Option<(Vec<usize>, Mat3)> = None;
|
||||
let mut iterations = max_iterations;
|
||||
let mut i = 0;
|
||||
while i < iterations {
|
||||
i += 1;
|
||||
let sample = rng.distinct4(n);
|
||||
let Some(h) = dlt(&sample.map(|k| pairs[k])) else {
|
||||
continue;
|
||||
};
|
||||
let inliers: Vec<usize> = (0..n).filter(|&k| agrees(&h, pairs[k], thr2)).collect();
|
||||
if best.as_ref().is_none_or(|(b, _)| inliers.len() > b.len()) {
|
||||
// Adapt: enough iterations to have drawn one all-inlier sample
|
||||
// with probability 0.99, given the ratio seen so far.
|
||||
let w = inliers.len() as f64 / n as f64;
|
||||
let p_all = w.powi(4);
|
||||
if p_all > 0.0 && p_all < 1.0 {
|
||||
let needed = ((1.0 - 0.99f64).ln() / (1.0 - p_all).ln()).ceil() as usize;
|
||||
iterations = iterations.min(needed.max(i + 1));
|
||||
}
|
||||
best = Some((inliers, h));
|
||||
}
|
||||
}
|
||||
let (inliers, h) = best?;
|
||||
if inliers.len() < 4 {
|
||||
return None;
|
||||
}
|
||||
// Refit on every inlier, and keep the refit only if it did not lose
|
||||
// support — a least-squares fit over a set with a few borderline points
|
||||
// can be pulled off the consensus the sample found.
|
||||
let refit: Vec<(Point, Point)> = inliers.iter().map(|&k| pairs[k]).collect();
|
||||
let h = match dlt(&refit) {
|
||||
Some(r) => {
|
||||
let count = (0..n).filter(|&k| agrees(&r, pairs[k], thr2)).count();
|
||||
if count >= inliers.len() {
|
||||
r
|
||||
} else {
|
||||
h
|
||||
}
|
||||
}
|
||||
None => h,
|
||||
};
|
||||
let inliers: Vec<usize> = (0..n).filter(|&k| agrees(&h, pairs[k], thr2)).collect();
|
||||
Some(RobustHomography { h, inliers })
|
||||
}
|
||||
|
||||
fn agrees(h: &Mat3, (p, q): (Point, Point), thr2: f64) -> bool {
|
||||
match apply(h, p) {
|
||||
Some((x, y)) => {
|
||||
let (dx, dy) = (x - q.0, y - q.1);
|
||||
dx * dx + dy * dy <= thr2
|
||||
}
|
||||
None => false,
|
||||
}
|
||||
}
|
||||
|
||||
/// The focal length a homography implies, if it implies one.
|
||||
///
|
||||
/// For `H = K R K⁻¹` with `K = diag(f, f, 1)` and the principal point at the
|
||||
/// origin, the orthonormality of `R` gives two independent estimates of `f²`
|
||||
/// from the first two rows and two from the first two columns; each is
|
||||
/// taken where it is positive and the better-conditioned of the pair is
|
||||
/// chosen, as OpenCV's `focalsFromHomography` does. The geometric mean of
|
||||
/// the row and column estimates is returned. `None` when the homography is
|
||||
/// too close to a pure translation to say anything — every estimate is then
|
||||
/// a ratio of small numbers.
|
||||
pub fn focal_from_homography(h: &Mat3) -> Option<f64> {
|
||||
let m = h.0;
|
||||
let (h00, h01, h02) = (m[0][0], m[0][1], m[0][2]);
|
||||
let (h10, h11, h12) = (m[1][0], m[1][1], m[1][2]);
|
||||
let (h20, h21) = (m[2][0], m[2][1]);
|
||||
|
||||
let pick = |mut v1: f64, mut v2: f64, d1: f64, d2: f64| -> Option<f64> {
|
||||
if v1 < v2 {
|
||||
std::mem::swap(&mut v1, &mut v2);
|
||||
}
|
||||
if v1 > 0.0 && v2 > 0.0 {
|
||||
Some((if d1.abs() > d2.abs() { v1 } else { v2 }).sqrt())
|
||||
} else if v1 > 0.0 {
|
||||
Some(v1.sqrt())
|
||||
} else {
|
||||
None
|
||||
}
|
||||
};
|
||||
|
||||
// From the third row.
|
||||
let d1 = h20 * h21;
|
||||
let d2 = (h21 - h20) * (h21 + h20);
|
||||
let f1 = if d1.abs() > 1e-12 || d2.abs() > 1e-12 {
|
||||
let v1 = if d1.abs() > 1e-12 {
|
||||
-(h00 * h01 + h10 * h11) / d1
|
||||
} else {
|
||||
f64::NAN
|
||||
};
|
||||
let v2 = if d2.abs() > 1e-12 {
|
||||
(h00 * h00 + h10 * h10 - h01 * h01 - h11 * h11) / d2
|
||||
} else {
|
||||
f64::NAN
|
||||
};
|
||||
pick(nan_to_neg(v1), nan_to_neg(v2), d1, d2)
|
||||
} else {
|
||||
None
|
||||
};
|
||||
|
||||
// From the third column.
|
||||
let d1 = h00 * h10 + h01 * h11;
|
||||
let d2 = h00 * h00 + h01 * h01 - h10 * h10 - h11 * h11;
|
||||
let f0 = if d1.abs() > 1e-12 || d2.abs() > 1e-12 {
|
||||
let v1 = if d1.abs() > 1e-12 {
|
||||
-h02 * h12 / d1
|
||||
} else {
|
||||
f64::NAN
|
||||
};
|
||||
let v2 = if d2.abs() > 1e-12 {
|
||||
(h12 * h12 - h02 * h02) / d2
|
||||
} else {
|
||||
f64::NAN
|
||||
};
|
||||
pick(nan_to_neg(v1), nan_to_neg(v2), d1, d2)
|
||||
} else {
|
||||
None
|
||||
};
|
||||
|
||||
match (f0, f1) {
|
||||
(Some(a), Some(b)) => Some((a * b).sqrt()),
|
||||
(Some(a), None) | (None, Some(a)) => Some(a),
|
||||
(None, None) => None,
|
||||
}
|
||||
}
|
||||
|
||||
fn nan_to_neg(v: f64) -> f64 {
|
||||
if v.is_finite() {
|
||||
v
|
||||
} else {
|
||||
-1.0
|
||||
}
|
||||
}
|
||||
|
||||
/// The rotation a homography encodes for a known focal length:
|
||||
/// `R = K⁻¹ H K`, re-orthonormalised, with the scale of `H` divided out.
|
||||
pub fn rotation_from_homography(h: &Mat3, f: f64) -> Mat3 {
|
||||
let m = h.0;
|
||||
// K⁻¹ H K with K = diag(f, f, 1): scale the third row by f and the
|
||||
// third column by 1/f.
|
||||
let r = Mat3([
|
||||
[m[0][0], m[0][1], m[0][2] / f],
|
||||
[m[1][0], m[1][1], m[1][2] / f],
|
||||
[m[2][0] * f, m[2][1] * f, m[2][2]],
|
||||
]);
|
||||
r.orthonormalised()
|
||||
}
|
||||
|
||||
/// A small deterministic generator for RANSAC's samples.
|
||||
struct Lcg(u64);
|
||||
|
||||
impl Lcg {
|
||||
fn next(&mut self) -> u64 {
|
||||
// Knuth's MMIX constants.
|
||||
self.0 = self
|
||||
.0
|
||||
.wrapping_mul(6364136223846793005)
|
||||
.wrapping_add(1442695040888963407);
|
||||
self.0 >> 33
|
||||
}
|
||||
|
||||
fn below(&mut self, n: usize) -> usize {
|
||||
(self.next() % n as u64) as usize
|
||||
}
|
||||
|
||||
fn distinct4(&mut self, n: usize) -> [usize; 4] {
|
||||
let mut s = [0usize; 4];
|
||||
for i in 0..4 {
|
||||
loop {
|
||||
let k = self.below(n);
|
||||
if !s[..i].contains(&k) {
|
||||
s[i] = k;
|
||||
break;
|
||||
}
|
||||
}
|
||||
}
|
||||
s
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
/// Points under a known rotation seen through a known focal length,
|
||||
/// in centred image coordinates scaled by that focal length.
|
||||
fn synthetic(f: f64, r: Mat3, n: usize, noise: f64, seed: u64) -> Vec<(Point, Point)> {
|
||||
let mut rng = Lcg(seed);
|
||||
let mut out = Vec::new();
|
||||
while out.len() < n {
|
||||
// A point on the first image plane, within ±0.3 f of centre.
|
||||
let x = (rng.below(6001) as f64 - 3000.0) / 10000.0;
|
||||
let y = (rng.below(4001) as f64 - 2000.0) / 10000.0;
|
||||
let b = Vec3::new(x, y, 1.0);
|
||||
let v = r * b;
|
||||
if v.z() <= 0.2 {
|
||||
continue;
|
||||
}
|
||||
let nx = (rng.below(2001) as f64 - 1000.0) / 1000.0 * noise;
|
||||
let ny = (rng.below(2001) as f64 - 1000.0) / 1000.0 * noise;
|
||||
out.push(((x, y), (v.x() / v.z() + nx, v.y() / v.z() + ny)));
|
||||
}
|
||||
let _ = f;
|
||||
out
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn dlt_recovers_a_known_homography_exactly() {
|
||||
let r = Mat3::exp(Vec3::new(0.05, 0.3, 0.02));
|
||||
let pairs = synthetic(1.0, r, 12, 0.0, 1);
|
||||
let h = dlt(&pairs).expect("solvable");
|
||||
for &(p, q) in &pairs {
|
||||
let (x, y) = apply(&h, p).unwrap();
|
||||
assert!((x - q.0).abs() < 1e-9 && (y - q.1).abs() < 1e-9);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn ransac_finds_the_consensus_among_outliers() {
|
||||
let r = Mat3::exp(Vec3::new(-0.02, 0.25, 0.01));
|
||||
let mut pairs = synthetic(1.0, r, 60, 0.0005, 2);
|
||||
// Forty outliers: wrong second point.
|
||||
let mut rng = Lcg(9);
|
||||
for _ in 0..40 {
|
||||
let k = rng.below(60);
|
||||
let (p, _) = pairs[k];
|
||||
pairs.push((p, ((rng.below(1000) as f64 - 500.0) / 1000.0, 0.1)));
|
||||
}
|
||||
let robust = ransac_homography(&pairs, 0.003, 500, 3).expect("found");
|
||||
assert!(
|
||||
robust.inliers.len() >= 55,
|
||||
"{} inliers",
|
||||
robust.inliers.len()
|
||||
);
|
||||
assert!(robust.inliers.iter().all(|&k| k < 60));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn focal_is_read_off_a_rotation_homography() {
|
||||
// H in *pixel* coordinates for f = 1400: K R K⁻¹.
|
||||
let f = 1400.0;
|
||||
let r = Mat3::exp(Vec3::new(0.03, 0.35, -0.01));
|
||||
let m = r.0;
|
||||
let h = Mat3([
|
||||
[m[0][0], m[0][1], m[0][2] * f],
|
||||
[m[1][0], m[1][1], m[1][2] * f],
|
||||
[m[2][0] / f, m[2][1] / f, m[2][2]],
|
||||
]);
|
||||
let est = focal_from_homography(&h).expect("estimable");
|
||||
assert!((est - f).abs() / f < 1e-6, "{est}");
|
||||
let back = rotation_from_homography(&h, f);
|
||||
for (row, truth) in back.0.iter().zip(&m) {
|
||||
for (a, b) in row.iter().zip(truth) {
|
||||
assert!((a - b).abs() < 1e-9);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn the_identity_implies_no_focal() {
|
||||
assert!(focal_from_homography(&Mat3::IDENTITY).is_none());
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,262 @@
|
||||
//! The grayscale proxy a detector reads.
|
||||
//!
|
||||
//! Alignment runs on proxies (FR-MRG-7) — a detector at 1024 px sees
|
||||
//! everything it needs, and the full-resolution frames never leave the GPU.
|
||||
//! This is that proxy: one channel, `f32` in `0.0..=1.0`, upright, and no
|
||||
//! larger than the detector's fixed input.
|
||||
|
||||
/// A single-channel image, row-major, values in `0.0..=1.0`.
|
||||
#[derive(Debug, Clone, PartialEq)]
|
||||
pub struct Gray {
|
||||
pub width: usize,
|
||||
pub height: usize,
|
||||
pub data: Vec<f32>,
|
||||
}
|
||||
|
||||
impl Gray {
|
||||
/// From tightly packed 8-bit RGBA, by the Rec. 709 luma weights.
|
||||
///
|
||||
/// The proxy is what a detector looks at, not what the photographer
|
||||
/// sees, so which luma is used matters less than that it is the same one
|
||||
/// for every frame — a keypoint's descriptor must not change between two
|
||||
/// frames because they were converted differently.
|
||||
pub fn from_rgba8(rgba: &[u8], width: usize, height: usize) -> Gray {
|
||||
let n = width * height;
|
||||
assert!(
|
||||
rgba.len() >= n * 4,
|
||||
"rgba buffer is short for {width}×{height}"
|
||||
);
|
||||
let data = rgba[..n * 4]
|
||||
.chunks_exact(4)
|
||||
.map(|p| {
|
||||
(0.2126 * f32::from(p[0]) + 0.7152 * f32::from(p[1]) + 0.0722 * f32::from(p[2]))
|
||||
/ 255.0
|
||||
})
|
||||
.collect();
|
||||
Gray {
|
||||
width,
|
||||
height,
|
||||
data,
|
||||
}
|
||||
}
|
||||
|
||||
/// Apply an EXIF orientation so the image is upright.
|
||||
///
|
||||
/// Learned detectors are not rotation-invariant — a descriptor of a
|
||||
/// feature seen sideways is a different descriptor — and a portrait set
|
||||
/// (the 6D fixture is one) would match poorly or not at all fed as
|
||||
/// stored. The camera says which way is up; the proxy is turned before
|
||||
/// anything looks at it, and the composite is written upright.
|
||||
///
|
||||
/// The value is the EXIF `Orientation` tag. Mirrored values (2, 4, 5, 7)
|
||||
/// are not produced by any camera and are treated as their unmirrored
|
||||
/// counterparts.
|
||||
pub fn oriented(&self, orientation: u16) -> Gray {
|
||||
match orientation {
|
||||
3 | 4 => self.rotated_180(),
|
||||
6 | 5 => self.rotated_90_cw(),
|
||||
8 | 7 => self.rotated_90_ccw(),
|
||||
_ => self.clone(),
|
||||
}
|
||||
}
|
||||
|
||||
fn rotated_90_cw(&self) -> Gray {
|
||||
let (w, h) = (self.width, self.height);
|
||||
let mut data = vec![0.0; w * h];
|
||||
for y in 0..h {
|
||||
for x in 0..w {
|
||||
// Source (x, y) lands at (h - 1 - y, x) in an h-wide image.
|
||||
data[x * h + (h - 1 - y)] = self.data[y * w + x];
|
||||
}
|
||||
}
|
||||
Gray {
|
||||
width: h,
|
||||
height: w,
|
||||
data,
|
||||
}
|
||||
}
|
||||
|
||||
fn rotated_90_ccw(&self) -> Gray {
|
||||
let (w, h) = (self.width, self.height);
|
||||
let mut data = vec![0.0; w * h];
|
||||
for y in 0..h {
|
||||
for x in 0..w {
|
||||
// Source (x, y) lands at (y, w - 1 - x) in an h-wide image.
|
||||
data[(w - 1 - x) * h + y] = self.data[y * w + x];
|
||||
}
|
||||
}
|
||||
Gray {
|
||||
width: h,
|
||||
height: w,
|
||||
data,
|
||||
}
|
||||
}
|
||||
|
||||
fn rotated_180(&self) -> Gray {
|
||||
let mut data = self.data.clone();
|
||||
data.reverse();
|
||||
Gray {
|
||||
width: self.width,
|
||||
height: self.height,
|
||||
data,
|
||||
}
|
||||
}
|
||||
|
||||
/// Resample to exactly `width × height` by area averaging on the way
|
||||
/// down and bilinear on the way up.
|
||||
///
|
||||
/// Area averaging, not point sampling, for a reduction: a 5472 px frame
|
||||
/// to 1024 is a factor of five, and picking one source pixel in
|
||||
/// twenty-five aliases every edge the detector is looking for.
|
||||
pub fn resampled(&self, width: usize, height: usize) -> Gray {
|
||||
if width == self.width && height == self.height {
|
||||
return self.clone();
|
||||
}
|
||||
let mut data = vec![0.0f32; width * height];
|
||||
let sx = self.width as f64 / width as f64;
|
||||
let sy = self.height as f64 / height as f64;
|
||||
if sx >= 1.0 && sy >= 1.0 {
|
||||
for oy in 0..height {
|
||||
let y0 = (oy as f64 * sy) as usize;
|
||||
let y1 = (((oy + 1) as f64 * sy) as usize).clamp(y0 + 1, self.height);
|
||||
for ox in 0..width {
|
||||
let x0 = (ox as f64 * sx) as usize;
|
||||
let x1 = (((ox + 1) as f64 * sx) as usize).clamp(x0 + 1, self.width);
|
||||
let mut sum = 0.0f32;
|
||||
for y in y0..y1 {
|
||||
let row = &self.data[y * self.width..(y + 1) * self.width];
|
||||
sum += row[x0..x1].iter().sum::<f32>();
|
||||
}
|
||||
data[oy * width + ox] = sum / ((y1 - y0) * (x1 - x0)) as f32;
|
||||
}
|
||||
}
|
||||
} else {
|
||||
for oy in 0..height {
|
||||
let fy = ((oy as f64 + 0.5) * sy - 0.5).max(0.0);
|
||||
let y0 = (fy as usize).min(self.height - 1);
|
||||
let y1 = (y0 + 1).min(self.height - 1);
|
||||
let ty = (fy - y0 as f64) as f32;
|
||||
for ox in 0..width {
|
||||
let fx = ((ox as f64 + 0.5) * sx - 0.5).max(0.0);
|
||||
let x0 = (fx as usize).min(self.width - 1);
|
||||
let x1 = (x0 + 1).min(self.width - 1);
|
||||
let tx = (fx - x0 as f64) as f32;
|
||||
let p = |x: usize, y: usize| self.data[y * self.width + x];
|
||||
let top = p(x0, y0) * (1.0 - tx) + p(x1, y0) * tx;
|
||||
let bot = p(x0, y1) * (1.0 - tx) + p(x1, y1) * tx;
|
||||
data[oy * width + ox] = top * (1.0 - ty) + bot * ty;
|
||||
}
|
||||
}
|
||||
}
|
||||
Gray {
|
||||
width,
|
||||
height,
|
||||
data,
|
||||
}
|
||||
}
|
||||
|
||||
/// Scale so the image fits inside `max_width × max_height`, preserving
|
||||
/// aspect, never enlarging. Returns the image and the scale applied,
|
||||
/// which is what maps a proxy keypoint back to the source.
|
||||
pub fn fitted(&self, max_width: usize, max_height: usize) -> (Gray, f64) {
|
||||
let scale = (max_width as f64 / self.width as f64)
|
||||
.min(max_height as f64 / self.height as f64)
|
||||
.min(1.0);
|
||||
let w = ((self.width as f64 * scale).round() as usize).max(1);
|
||||
let h = ((self.height as f64 * scale).round() as usize).max(1);
|
||||
(self.resampled(w, h), w as f64 / self.width as f64)
|
||||
}
|
||||
|
||||
/// Copy into the top-left of a `width × height` canvas, zero elsewhere.
|
||||
///
|
||||
/// The detector's input is a fixed shape (S15.2), and a frame that fits
|
||||
/// inside it is padded rather than stretched: stretching changes the
|
||||
/// aspect and with it every descriptor.
|
||||
pub fn padded(&self, width: usize, height: usize) -> Gray {
|
||||
assert!(self.width <= width && self.height <= height);
|
||||
let mut data = vec![0.0; width * height];
|
||||
for y in 0..self.height {
|
||||
data[y * width..y * width + self.width]
|
||||
.copy_from_slice(&self.data[y * self.width..(y + 1) * self.width]);
|
||||
}
|
||||
Gray {
|
||||
width,
|
||||
height,
|
||||
data,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
fn ramp(w: usize, h: usize) -> Gray {
|
||||
Gray {
|
||||
width: w,
|
||||
height: h,
|
||||
data: (0..w * h).map(|i| i as f32).collect(),
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn rotating_four_quarter_turns_is_the_identity() {
|
||||
let g = ramp(5, 3);
|
||||
let mut r = g.clone();
|
||||
for _ in 0..4 {
|
||||
r = r.rotated_90_cw();
|
||||
}
|
||||
assert_eq!(r, g);
|
||||
assert_eq!(g.rotated_90_cw().rotated_90_ccw(), g);
|
||||
assert_eq!(g.rotated_180().rotated_180(), g);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_clockwise_turn_moves_the_top_left_to_the_top_right() {
|
||||
// 2×3 image, pixel values by position.
|
||||
let g = ramp(2, 3);
|
||||
let r = g.rotated_90_cw();
|
||||
assert_eq!((r.width, r.height), (3, 2));
|
||||
// Top-left of source (value 0) is at top-right of result.
|
||||
assert_eq!(r.data[2], 0.0);
|
||||
// Bottom-left of source (value 4) is at top-left of result.
|
||||
assert_eq!(r.data[0], 4.0);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn orientation_8_is_a_counter_clockwise_turn() {
|
||||
let g = ramp(4, 2);
|
||||
assert_eq!(g.oriented(8), g.rotated_90_ccw());
|
||||
assert_eq!(g.oriented(6), g.rotated_90_cw());
|
||||
assert_eq!(g.oriented(1), g);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn downsampling_by_two_averages_blocks() {
|
||||
let g = Gray {
|
||||
width: 4,
|
||||
height: 2,
|
||||
data: vec![0.0, 1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0],
|
||||
};
|
||||
let r = g.resampled(2, 1);
|
||||
assert_eq!(r.data, vec![2.5, 4.5]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn fitting_never_enlarges_and_reports_the_scale() {
|
||||
let g = ramp(100, 50);
|
||||
let (f, s) = g.fitted(1024, 768);
|
||||
assert_eq!((f.width, f.height), (100, 50));
|
||||
assert_eq!(s, 1.0);
|
||||
let (f, s) = g.fitted(50, 50);
|
||||
assert_eq!((f.width, f.height), (50, 25));
|
||||
assert_eq!(s, 0.5);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn padding_places_the_image_at_the_origin() {
|
||||
let g = ramp(2, 2);
|
||||
let p = g.padded(3, 3);
|
||||
assert_eq!(p.data, vec![0.0, 1.0, 0.0, 2.0, 3.0, 0.0, 0.0, 0.0, 0.0]);
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,86 @@
|
||||
//! TRACES: FR-MRG-1 | FR-MRG-10
|
||||
//! Panorama geometry — from several frames to the rotations that relate
|
||||
//! them, and the projections that lay them out.
|
||||
//!
|
||||
//! This is the CPU half of a merge (FR-MRG-10): keypoints, matching, the
|
||||
//! rotation solve and the choice of output surface. The per-pixel half —
|
||||
//! rendering, warping, seams, blending — is the GPU's and lives in
|
||||
//! `dr-gpu`, driven from above; nothing here touches a full-resolution
|
||||
//! pixel. The split is the whole design (panorama.md §4): everything in
|
||||
//! this crate is bounded by the number of frames, not the size of the
|
||||
//! composite, and runs on proxies.
|
||||
//!
|
||||
//! # Layout
|
||||
//!
|
||||
//! - [`image`] — the grayscale proxy a detector reads: oriented, resampled.
|
||||
//! - [`features`] — keypoints with descriptors, and the XFeat decoder.
|
||||
//! - [`xfeat`] — the network under tract (feature `xfeat`).
|
||||
//! - [`matching`] — mutual nearest neighbours.
|
||||
//! - [`homography`] — a robust pairwise homography, the focal length read
|
||||
//! off it, and the rotation it implies.
|
||||
//! - [`bundle`] — every rotation and the focal length refined together.
|
||||
//! - [`align`] — the whole thing, from features to cameras, honest about
|
||||
//! what it could not place.
|
||||
//! - [`projection`] — perspective, cylindrical, spherical.
|
||||
//! - [`linalg`] — the small dense algebra all of it uses.
|
||||
//!
|
||||
//! # What it depends on
|
||||
//!
|
||||
//! Nothing, without the `xfeat` feature: the geometry is pure Rust with
|
||||
//! hand-rolled linear algebra (`linalg` says why) so that it tests without
|
||||
//! a model, a GPU or a device, on synthetic sets whose answer is known
|
||||
//! exactly. With the feature it adds the same `ort`-over-tract runtime the
|
||||
//! rest of the application already carries.
|
||||
|
||||
pub mod align;
|
||||
pub mod bundle;
|
||||
pub mod features;
|
||||
pub mod fill;
|
||||
pub mod homography;
|
||||
pub mod image;
|
||||
pub mod linalg;
|
||||
pub mod matching;
|
||||
#[cfg(feature = "xfeat")]
|
||||
pub mod migan;
|
||||
pub mod projection;
|
||||
#[cfg(feature = "xfeat")]
|
||||
pub mod xfeat;
|
||||
|
||||
pub use align::{align, AlignOptions, Alignment, Link, Unaligned};
|
||||
pub use bundle::Cameras;
|
||||
pub use features::{Features, Keypoint};
|
||||
pub use fill::{fill_border, Inpainter, Observer, Params as FillParams};
|
||||
pub use image::Gray;
|
||||
pub use projection::Projection;
|
||||
|
||||
#[derive(Debug, thiserror::Error)]
|
||||
pub enum PanoError {
|
||||
#[error("bad input: {0}")]
|
||||
Input(String),
|
||||
#[error("geometry: {0}")]
|
||||
Geometry(String),
|
||||
#[error("model: {0}")]
|
||||
Model(String),
|
||||
#[error("could not read the model: {0}")]
|
||||
ModelRead(#[source] std::io::Error),
|
||||
#[cfg(feature = "xfeat")]
|
||||
#[error("inference: {0}")]
|
||||
Inference(#[source] ort::Error),
|
||||
}
|
||||
|
||||
#[cfg(feature = "xfeat")]
|
||||
impl From<ort::Error> for PanoError {
|
||||
fn from(e: ort::Error) -> Self {
|
||||
PanoError::Inference(e)
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(feature = "xfeat")]
|
||||
impl From<dr_inference_engine::Error> for PanoError {
|
||||
fn from(e: dr_inference_engine::Error) -> Self {
|
||||
match e {
|
||||
dr_inference_engine::Error::Inference(e) => PanoError::Inference(e),
|
||||
dr_inference_engine::Error::Io(e) => PanoError::ModelRead(e),
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,371 @@
|
||||
//! The small dense linear algebra the geometry needs, and nothing more.
|
||||
//!
|
||||
//! Hand-rolled rather than pulled in, and the decision was made on purpose
|
||||
//! (2026-09-19): the largest system this crate ever solves is a rotation
|
||||
//! per frame plus one focal length — forty unknowns for a dozen frames —
|
||||
//! and everything else is three-vectors. A general linear-algebra crate
|
||||
//! would be the largest dependency in `dr-pano` by an order of magnitude,
|
||||
//! for a Cholesky factorisation that is thirty lines.
|
||||
//!
|
||||
//! `f64` throughout. The geometry is solved once per merge on a few thousand
|
||||
//! matches; there is no reason to give up precision for speed here, and the
|
||||
//! bundle adjustment's normal equations are poorly conditioned enough near
|
||||
//! convergence that `f32` would stall it.
|
||||
|
||||
use std::ops::{Add, Index, IndexMut, Mul, Neg, Sub};
|
||||
|
||||
/// A vector in three dimensions.
|
||||
#[derive(Debug, Clone, Copy, PartialEq, Default)]
|
||||
pub struct Vec3(pub [f64; 3]);
|
||||
|
||||
impl Vec3 {
|
||||
pub const fn new(x: f64, y: f64, z: f64) -> Self {
|
||||
Vec3([x, y, z])
|
||||
}
|
||||
|
||||
pub fn dot(self, o: Vec3) -> f64 {
|
||||
self.0[0] * o.0[0] + self.0[1] * o.0[1] + self.0[2] * o.0[2]
|
||||
}
|
||||
|
||||
pub fn cross(self, o: Vec3) -> Vec3 {
|
||||
Vec3([
|
||||
self.0[1] * o.0[2] - self.0[2] * o.0[1],
|
||||
self.0[2] * o.0[0] - self.0[0] * o.0[2],
|
||||
self.0[0] * o.0[1] - self.0[1] * o.0[0],
|
||||
])
|
||||
}
|
||||
|
||||
pub fn norm(self) -> f64 {
|
||||
self.dot(self).sqrt()
|
||||
}
|
||||
|
||||
/// The unit vector along `self`, or `self` unchanged if it is zero.
|
||||
pub fn normalised(self) -> Vec3 {
|
||||
let n = self.norm();
|
||||
if n > 0.0 {
|
||||
self * (1.0 / n)
|
||||
} else {
|
||||
self
|
||||
}
|
||||
}
|
||||
|
||||
pub fn x(self) -> f64 {
|
||||
self.0[0]
|
||||
}
|
||||
pub fn y(self) -> f64 {
|
||||
self.0[1]
|
||||
}
|
||||
pub fn z(self) -> f64 {
|
||||
self.0[2]
|
||||
}
|
||||
}
|
||||
|
||||
impl Add for Vec3 {
|
||||
type Output = Vec3;
|
||||
fn add(self, o: Vec3) -> Vec3 {
|
||||
Vec3([self.0[0] + o.0[0], self.0[1] + o.0[1], self.0[2] + o.0[2]])
|
||||
}
|
||||
}
|
||||
|
||||
impl Sub for Vec3 {
|
||||
type Output = Vec3;
|
||||
fn sub(self, o: Vec3) -> Vec3 {
|
||||
Vec3([self.0[0] - o.0[0], self.0[1] - o.0[1], self.0[2] - o.0[2]])
|
||||
}
|
||||
}
|
||||
|
||||
impl Mul<f64> for Vec3 {
|
||||
type Output = Vec3;
|
||||
fn mul(self, s: f64) -> Vec3 {
|
||||
Vec3([self.0[0] * s, self.0[1] * s, self.0[2] * s])
|
||||
}
|
||||
}
|
||||
|
||||
impl Neg for Vec3 {
|
||||
type Output = Vec3;
|
||||
fn neg(self) -> Vec3 {
|
||||
Vec3([-self.0[0], -self.0[1], -self.0[2]])
|
||||
}
|
||||
}
|
||||
|
||||
/// A 3×3 matrix, row-major.
|
||||
#[derive(Debug, Clone, Copy, PartialEq)]
|
||||
pub struct Mat3(pub [[f64; 3]; 3]);
|
||||
|
||||
impl Mat3 {
|
||||
pub const IDENTITY: Mat3 = Mat3([[1.0, 0.0, 0.0], [0.0, 1.0, 0.0], [0.0, 0.0, 1.0]]);
|
||||
|
||||
/// The matrix whose columns are `a`, `b`, `c`.
|
||||
pub fn from_columns(a: Vec3, b: Vec3, c: Vec3) -> Mat3 {
|
||||
Mat3([
|
||||
[a.0[0], b.0[0], c.0[0]],
|
||||
[a.0[1], b.0[1], c.0[1]],
|
||||
[a.0[2], b.0[2], c.0[2]],
|
||||
])
|
||||
}
|
||||
|
||||
pub fn transpose(self) -> Mat3 {
|
||||
let m = self.0;
|
||||
Mat3([
|
||||
[m[0][0], m[1][0], m[2][0]],
|
||||
[m[0][1], m[1][1], m[2][1]],
|
||||
[m[0][2], m[1][2], m[2][2]],
|
||||
])
|
||||
}
|
||||
|
||||
pub fn column(self, i: usize) -> Vec3 {
|
||||
Vec3([self.0[0][i], self.0[1][i], self.0[2][i]])
|
||||
}
|
||||
|
||||
pub fn trace(self) -> f64 {
|
||||
self.0[0][0] + self.0[1][1] + self.0[2][2]
|
||||
}
|
||||
|
||||
/// The rotation about `axis` (any length) by `angle` radians — Rodrigues.
|
||||
pub fn rotation(axis: Vec3, angle: f64) -> Mat3 {
|
||||
let k = axis.normalised();
|
||||
let (s, c) = angle.sin_cos();
|
||||
let t = 1.0 - c;
|
||||
let (x, y, z) = (k.0[0], k.0[1], k.0[2]);
|
||||
Mat3([
|
||||
[t * x * x + c, t * x * y - s * z, t * x * z + s * y],
|
||||
[t * x * y + s * z, t * y * y + c, t * y * z - s * x],
|
||||
[t * x * z - s * y, t * y * z + s * x, t * z * z + c],
|
||||
])
|
||||
}
|
||||
|
||||
/// The rotation whose axis-angle vector is `w` (direction is the axis,
|
||||
/// length is the angle). The exponential map; [`Self::log`] inverts it.
|
||||
pub fn exp(w: Vec3) -> Mat3 {
|
||||
let angle = w.norm();
|
||||
if angle < 1e-12 {
|
||||
// First-order: I + [w]×, which is what the limit is and avoids
|
||||
// dividing by the angle.
|
||||
let (x, y, z) = (w.0[0], w.0[1], w.0[2]);
|
||||
return Mat3([[1.0, -z, y], [z, 1.0, -x], [-y, x, 1.0]]);
|
||||
}
|
||||
Mat3::rotation(w, angle)
|
||||
}
|
||||
|
||||
/// The axis-angle vector of a rotation matrix. Inverse of [`Self::exp`].
|
||||
pub fn log(self) -> Vec3 {
|
||||
let m = self.0;
|
||||
let cos = ((self.trace() - 1.0) * 0.5).clamp(-1.0, 1.0);
|
||||
let axis = Vec3([m[2][1] - m[1][2], m[0][2] - m[2][0], m[1][0] - m[0][1]]);
|
||||
if cos > 1.0 - 1e-6 {
|
||||
// Small angle: `acos` near 1 loses everything below ~1e-8 to
|
||||
// rounding, but the antisymmetric part is `2 sin θ · axis` and
|
||||
// keeps it. First order, exact to the precision that matters.
|
||||
return axis * 0.5;
|
||||
}
|
||||
let angle = cos.acos();
|
||||
if angle > std::f64::consts::PI - 1e-6 {
|
||||
// Near π the antisymmetric part vanishes; take the axis from the
|
||||
// symmetric part instead. Rare for a panorama, but the solver may
|
||||
// pass through it on a bad start and must not return NaN.
|
||||
let d = Vec3([
|
||||
((m[0][0] + 1.0) * 0.5).max(0.0).sqrt(),
|
||||
((m[1][1] + 1.0) * 0.5).max(0.0).sqrt(),
|
||||
((m[2][2] + 1.0) * 0.5).max(0.0).sqrt(),
|
||||
]);
|
||||
return d.normalised() * angle;
|
||||
}
|
||||
axis * (angle / (2.0 * angle.sin()))
|
||||
}
|
||||
|
||||
/// Re-orthonormalise a matrix that has drifted from a rotation through
|
||||
/// accumulated products. Gram–Schmidt on the columns; cheap and adequate
|
||||
/// for drift of the size floating-point products produce.
|
||||
pub fn orthonormalised(self) -> Mat3 {
|
||||
let a = self.column(0).normalised();
|
||||
let b = (self.column(1) - a * a.dot(self.column(1))).normalised();
|
||||
let c = a.cross(b);
|
||||
Mat3::from_columns(a, b, c)
|
||||
}
|
||||
}
|
||||
|
||||
impl Mul<Vec3> for Mat3 {
|
||||
type Output = Vec3;
|
||||
fn mul(self, v: Vec3) -> Vec3 {
|
||||
let m = self.0;
|
||||
Vec3([
|
||||
m[0][0] * v.0[0] + m[0][1] * v.0[1] + m[0][2] * v.0[2],
|
||||
m[1][0] * v.0[0] + m[1][1] * v.0[1] + m[1][2] * v.0[2],
|
||||
m[2][0] * v.0[0] + m[2][1] * v.0[1] + m[2][2] * v.0[2],
|
||||
])
|
||||
}
|
||||
}
|
||||
|
||||
impl Mul for Mat3 {
|
||||
type Output = Mat3;
|
||||
fn mul(self, o: Mat3) -> Mat3 {
|
||||
let mut r = [[0.0; 3]; 3];
|
||||
for (i, row) in r.iter_mut().enumerate() {
|
||||
for (j, cell) in row.iter_mut().enumerate() {
|
||||
*cell = (0..3).map(|k| self.0[i][k] * o.0[k][j]).sum();
|
||||
}
|
||||
}
|
||||
Mat3(r)
|
||||
}
|
||||
}
|
||||
|
||||
/// A dense square matrix, for the normal equations.
|
||||
#[derive(Debug, Clone, PartialEq)]
|
||||
pub struct DMat {
|
||||
n: usize,
|
||||
data: Vec<f64>,
|
||||
}
|
||||
|
||||
impl DMat {
|
||||
pub fn zeros(n: usize) -> DMat {
|
||||
DMat {
|
||||
n,
|
||||
data: vec![0.0; n * n],
|
||||
}
|
||||
}
|
||||
|
||||
pub fn n(&self) -> usize {
|
||||
self.n
|
||||
}
|
||||
|
||||
/// Solve `self · x = b` for a symmetric positive-definite `self` by
|
||||
/// Cholesky factorisation. `None` if the matrix is not positive definite,
|
||||
/// which for the normal equations means the problem is not determined by
|
||||
/// the data — a frame with no matches, for instance — and the caller
|
||||
/// should say so rather than proceed.
|
||||
///
|
||||
/// Destroys neither input: the factor is built in a copy. The systems
|
||||
/// here are at most a few dozen unknowns and the copy is nothing.
|
||||
pub fn solve_spd(&self, b: &[f64]) -> Option<Vec<f64>> {
|
||||
let n = self.n;
|
||||
debug_assert_eq!(b.len(), n);
|
||||
let mut l = vec![0.0; n * n];
|
||||
for j in 0..n {
|
||||
let mut d = self[(j, j)];
|
||||
for k in 0..j {
|
||||
d -= l[j * n + k] * l[j * n + k];
|
||||
}
|
||||
if d <= 0.0 || !d.is_finite() {
|
||||
return None;
|
||||
}
|
||||
let djj = d.sqrt();
|
||||
l[j * n + j] = djj;
|
||||
for i in j + 1..n {
|
||||
let mut s = self[(i, j)];
|
||||
for k in 0..j {
|
||||
s -= l[i * n + k] * l[j * n + k];
|
||||
}
|
||||
l[i * n + j] = s / djj;
|
||||
}
|
||||
}
|
||||
// Forward: L y = b.
|
||||
let mut y = vec![0.0; n];
|
||||
for i in 0..n {
|
||||
let mut s = b[i];
|
||||
for k in 0..i {
|
||||
s -= l[i * n + k] * y[k];
|
||||
}
|
||||
y[i] = s / l[i * n + i];
|
||||
}
|
||||
// Back: Lᵀ x = y.
|
||||
let mut x = vec![0.0; n];
|
||||
for i in (0..n).rev() {
|
||||
let mut s = y[i];
|
||||
for k in i + 1..n {
|
||||
s -= l[k * n + i] * x[k];
|
||||
}
|
||||
x[i] = s / l[i * n + i];
|
||||
}
|
||||
Some(x)
|
||||
}
|
||||
}
|
||||
|
||||
impl Index<(usize, usize)> for DMat {
|
||||
type Output = f64;
|
||||
fn index(&self, (i, j): (usize, usize)) -> &f64 {
|
||||
&self.data[i * self.n + j]
|
||||
}
|
||||
}
|
||||
|
||||
impl IndexMut<(usize, usize)> for DMat {
|
||||
fn index_mut(&mut self, (i, j): (usize, usize)) -> &mut f64 {
|
||||
&mut self.data[i * self.n + j]
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
fn close(a: f64, b: f64) -> bool {
|
||||
(a - b).abs() < 1e-9
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn exp_and_log_are_inverses() {
|
||||
for w in [
|
||||
Vec3::new(0.1, -0.2, 0.3),
|
||||
Vec3::new(1.0, 0.0, 0.0),
|
||||
Vec3::new(0.0, 0.0, 2.5),
|
||||
Vec3::new(1e-9, 0.0, 0.0),
|
||||
] {
|
||||
let back = Mat3::exp(w).log();
|
||||
for i in 0..3 {
|
||||
assert!(close(back.0[i], w.0[i]), "{w:?} -> {back:?}");
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_rotation_is_orthonormal_and_preserves_length() {
|
||||
let r = Mat3::exp(Vec3::new(0.4, 0.5, -0.6));
|
||||
let rt = r.transpose() * r;
|
||||
for i in 0..3 {
|
||||
for j in 0..3 {
|
||||
assert!(close(rt.0[i][j], Mat3::IDENTITY.0[i][j]));
|
||||
}
|
||||
}
|
||||
let v = Vec3::new(1.0, 2.0, 3.0);
|
||||
assert!(close((r * v).norm(), v.norm()));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn rotation_about_z_turns_x_towards_y() {
|
||||
let r = Mat3::rotation(Vec3::new(0.0, 0.0, 1.0), std::f64::consts::FRAC_PI_2);
|
||||
let v = r * Vec3::new(1.0, 0.0, 0.0);
|
||||
assert!(close(v.x(), 0.0) && close(v.y(), 1.0) && close(v.z(), 0.0));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn cholesky_solves_a_small_spd_system() {
|
||||
// A = Bᵀ B for a random-ish B is SPD by construction.
|
||||
let b = [
|
||||
[2.0, 1.0, 0.0],
|
||||
[1.0, 3.0, 1.0],
|
||||
[0.0, 1.0, 4.0],
|
||||
[1.0, 1.0, 1.0],
|
||||
];
|
||||
let mut a = DMat::zeros(3);
|
||||
for i in 0..3 {
|
||||
for j in 0..3 {
|
||||
a[(i, j)] = (0..4).map(|k| b[k][i] * b[k][j]).sum();
|
||||
}
|
||||
}
|
||||
let x_true = [1.0, -2.0, 0.5];
|
||||
let rhs: Vec<f64> = (0..3)
|
||||
.map(|i| (0..3).map(|j| a[(i, j)] * x_true[j]).sum())
|
||||
.collect();
|
||||
let x = a.solve_spd(&rhs).expect("spd");
|
||||
for i in 0..3 {
|
||||
assert!(close(x[i], x_true[i]), "{x:?}");
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn cholesky_refuses_an_indefinite_matrix() {
|
||||
let mut a = DMat::zeros(2);
|
||||
a[(0, 0)] = 1.0;
|
||||
a[(1, 1)] = -1.0;
|
||||
assert!(a.solve_spd(&[1.0, 1.0]).is_none());
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,173 @@
|
||||
//! Descriptor matching between two images.
|
||||
//!
|
||||
//! Mutual nearest neighbour on cosine similarity, with a floor on the
|
||||
//! similarity — the reference XFeat's own matcher (`match_mkpts`,
|
||||
//! `min_cossim = 0.82`). For a panorama that is enough: one lens, one
|
||||
//! scene, near-pure rotation and 20–40 % overlap make the matching problem
|
||||
//! easy, and what is hard — sky, repeated structure, exposure drift — is
|
||||
//! handled by the detector's descriptors and by RANSAC downstream, not by a
|
||||
//! cleverer matcher. A learned matcher (LightGlue) is the step after this
|
||||
//! one fails on a real set, and it has not (panorama.md §6).
|
||||
//!
|
||||
//! Brute force. `4096 × 4096 × 64` multiply-adds is a billion per pair,
|
||||
//! and a twelve-frame set has sixty-six pairs: a minute single-threaded
|
||||
//! and scalar (measured 2026-09-19: 51 s), a few seconds vectorised across
|
||||
//! the cores. Not worth an index, but worth doing properly.
|
||||
|
||||
use crate::features::{Features, DESCRIPTOR_LEN};
|
||||
|
||||
const _: () = assert!(DESCRIPTOR_LEN.is_multiple_of(8));
|
||||
|
||||
/// A correspondence: keypoint `a` in the first image matches keypoint `b`
|
||||
/// in the second, with the cosine similarity of their descriptors.
|
||||
#[derive(Debug, Clone, Copy, PartialEq)]
|
||||
pub struct Match {
|
||||
pub a: usize,
|
||||
pub b: usize,
|
||||
pub similarity: f32,
|
||||
}
|
||||
|
||||
/// Match two sets of features.
|
||||
///
|
||||
/// A pair is kept when each is the other's nearest neighbour and their
|
||||
/// similarity is at least `min_similarity`.
|
||||
pub fn match_features(a: &Features, b: &Features, min_similarity: f32) -> Vec<Match> {
|
||||
if a.is_empty() || b.is_empty() {
|
||||
return Vec::new();
|
||||
}
|
||||
let (na, nb) = (a.len(), b.len());
|
||||
|
||||
// The whole similarity matrix, once. Both nearest-neighbour directions
|
||||
// read it, which halves the multiply-adds against computing each
|
||||
// direction on its own; 4096 × 4096 × f32 is 64 MB, transient.
|
||||
let mut sim = vec![0.0f32; na * nb];
|
||||
let threads = std::thread::available_parallelism()
|
||||
.map(usize::from)
|
||||
.unwrap_or(1)
|
||||
.clamp(1, 16);
|
||||
let rows_per = na.div_ceil(threads);
|
||||
std::thread::scope(|scope| {
|
||||
for (t, chunk) in sim.chunks_mut(rows_per * nb).enumerate() {
|
||||
scope.spawn(move || {
|
||||
let first = t * rows_per;
|
||||
for (r, row) in chunk.chunks_mut(nb).enumerate() {
|
||||
let da = a.descriptor(first + r);
|
||||
for (j, cell) in row.iter_mut().enumerate() {
|
||||
*cell = dot(da, b.descriptor(j));
|
||||
}
|
||||
}
|
||||
});
|
||||
}
|
||||
});
|
||||
|
||||
// Best in `b` for each `a`, and best in `a` for each `b`.
|
||||
let best_ab: Vec<(usize, f32)> = sim
|
||||
.chunks_exact(nb)
|
||||
.map(|row| {
|
||||
row.iter().enumerate().fold(
|
||||
(0usize, f32::MIN),
|
||||
|acc, (j, &s)| if s > acc.1 { (j, s) } else { acc },
|
||||
)
|
||||
})
|
||||
.collect();
|
||||
let mut best_ba = vec![(0usize, f32::MIN); nb];
|
||||
for (i, row) in sim.chunks_exact(nb).enumerate() {
|
||||
for (j, &s) in row.iter().enumerate() {
|
||||
if s > best_ba[j].1 {
|
||||
best_ba[j] = (i, s);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
best_ab
|
||||
.iter()
|
||||
.enumerate()
|
||||
.filter_map(|(ia, &(ib, s))| {
|
||||
(best_ba[ib].0 == ia && s >= min_similarity).then_some(Match {
|
||||
a: ia,
|
||||
b: ib,
|
||||
similarity: s,
|
||||
})
|
||||
})
|
||||
.collect()
|
||||
}
|
||||
|
||||
#[inline]
|
||||
fn dot(a: &[f32], b: &[f32]) -> f32 {
|
||||
// Eight independent accumulators over exact 8-lane chunks: the shape
|
||||
// the compiler turns into one vector multiply-add per chunk, and no
|
||||
// bounds checks inside the loop. `DESCRIPTOR_LEN` is a multiple of 8.
|
||||
let (a, b) = (&a[..DESCRIPTOR_LEN], &b[..DESCRIPTOR_LEN]);
|
||||
let mut acc = [0.0f32; 8];
|
||||
for (ca, cb) in a.chunks_exact(8).zip(b.chunks_exact(8)) {
|
||||
for k in 0..8 {
|
||||
acc[k] += ca[k] * cb[k];
|
||||
}
|
||||
}
|
||||
acc.iter().sum()
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
use crate::features::Keypoint;
|
||||
|
||||
/// Features whose descriptors are unit vectors along the given axes.
|
||||
fn along(axes: &[usize]) -> Features {
|
||||
let mut descriptors = vec![0.0; axes.len() * DESCRIPTOR_LEN];
|
||||
for (i, &ax) in axes.iter().enumerate() {
|
||||
descriptors[i * DESCRIPTOR_LEN + ax] = 1.0;
|
||||
}
|
||||
Features {
|
||||
keypoints: axes
|
||||
.iter()
|
||||
.map(|_| Keypoint {
|
||||
x: 0.0,
|
||||
y: 0.0,
|
||||
score: 1.0,
|
||||
})
|
||||
.collect(),
|
||||
descriptors,
|
||||
width: 1,
|
||||
height: 1,
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn identical_descriptors_match_mutually() {
|
||||
let a = along(&[0, 1, 2]);
|
||||
let b = along(&[2, 0, 1]);
|
||||
let m = match_features(&a, &b, 0.8);
|
||||
let mut pairs: Vec<(usize, usize)> = m.iter().map(|m| (m.a, m.b)).collect();
|
||||
pairs.sort();
|
||||
assert_eq!(pairs, vec![(0, 1), (1, 2), (2, 0)]);
|
||||
assert!(m.iter().all(|m| (m.similarity - 1.0).abs() < 1e-6));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_descriptor_with_no_counterpart_is_unmatched() {
|
||||
let a = along(&[0, 1, 5]);
|
||||
let b = along(&[0, 1]);
|
||||
let m = match_features(&a, &b, 0.8);
|
||||
assert_eq!(m.len(), 2);
|
||||
assert!(m.iter().all(|m| m.a != 2));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn mutuality_breaks_a_one_sided_match() {
|
||||
// b0 is the nearest to both a0 and a1, but a0 is its nearest — a1
|
||||
// must not be matched to it.
|
||||
let mut a = along(&[0, 0]);
|
||||
a.descriptors[DESCRIPTOR_LEN] = 0.9;
|
||||
a.descriptors[DESCRIPTOR_LEN + 1] = (1.0f32 - 0.81).sqrt();
|
||||
let b = along(&[0]);
|
||||
let m = match_features(&a, &b, 0.0);
|
||||
assert_eq!(m.len(), 1);
|
||||
assert_eq!((m[0].a, m[0].b), (0, 0));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn empty_input_is_empty_output() {
|
||||
assert!(match_features(&along(&[]), &along(&[1]), 0.5).is_empty());
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,107 @@
|
||||
//! TRACES: FR-MRG-4
|
||||
//! MI-GAN, the border filler, under the inference engine.
|
||||
//!
|
||||
//! Sargsyan et al., ICCV 2023 (Picsart AI Research): inpainting built for
|
||||
//! phones — about six million parameters of plain convolutions, no FFT and
|
||||
//! no attention, so it quantises and runs on a DSP. MIT, code and weights
|
||||
//! (`models/LICENCE.md`). The bare 512 generator is what ships, exported
|
||||
//! at a fixed shape by `tools/export-migan.sh`; its six operator types load
|
||||
//! on every rung, and what they cost is the whole story of whether a fill
|
||||
//! is interactive: 7.4 s a tile under tract, 0.4 s under ONNX Runtime's
|
||||
//! CPU pool, 23 ms in fp16 and 13 ms in int8 on a laptop's TensorRT
|
||||
//! (2026-09-19, docs/panorama.md §12).
|
||||
//!
|
||||
//! The model's contract, from the reference `export_inference_model.py`:
|
||||
//! input `1×4×512×512` float — channel 0 is `mask − 0.5` with 1 where the
|
||||
//! picture is known, channels 1–3 the RGB in −1..1 with the unknown pixels
|
||||
//! zeroed; output `1×3×512×512` in −1..1, of which the caller keeps the
|
||||
//! unknown pixels. That is [`crate::fill::Inpainter`], and the rest —
|
||||
//! which tiles, what context, how to blend — is `fill.rs`.
|
||||
|
||||
use crate::fill::Inpainter;
|
||||
use crate::PanoError;
|
||||
|
||||
/// The tile the shipped export takes.
|
||||
pub const TILE: usize = 512;
|
||||
|
||||
pub struct MiGan {
|
||||
model: dr_inference_engine::Model,
|
||||
}
|
||||
|
||||
impl MiGan {
|
||||
/// From the model file, in whichever form the engine's rung wants
|
||||
/// (`resolve_model` picks an int8 sibling for the Hexagon).
|
||||
pub fn from_path(path: &std::path::Path) -> Result<Self, PanoError> {
|
||||
use dr_inference_engine::{resolve_model, Role};
|
||||
let (path, form) = resolve_model(Role::Inpainter, path);
|
||||
let bytes = std::fs::read(&path).map_err(PanoError::ModelRead)?;
|
||||
Self::from_bytes(&bytes, form)
|
||||
}
|
||||
|
||||
pub fn from_bytes(bytes: &[u8], form: dr_inference_engine::Form) -> Result<Self, PanoError> {
|
||||
use dr_inference_engine::Role;
|
||||
Ok(MiGan {
|
||||
model: dr_inference_engine::open(Role::Inpainter, form, bytes)?,
|
||||
})
|
||||
}
|
||||
|
||||
/// Where the fill runs, for a status line.
|
||||
pub fn rung(&self) -> Result<dr_inference_engine::Rung, PanoError> {
|
||||
Ok(self.model.acquire()?.rung())
|
||||
}
|
||||
}
|
||||
|
||||
impl Inpainter for MiGan {
|
||||
fn tile(&self) -> usize {
|
||||
TILE
|
||||
}
|
||||
|
||||
fn fill(&mut self, rgb: &[f32], known: &[bool]) -> Result<Vec<f32>, PanoError> {
|
||||
let n = TILE * TILE;
|
||||
if rgb.len() != n * 3 || known.len() != n {
|
||||
return Err(PanoError::Input(format!(
|
||||
"MI-GAN takes a {TILE}×{TILE} tile; given {} values and {} mask entries",
|
||||
rgb.len(),
|
||||
known.len()
|
||||
)));
|
||||
}
|
||||
// NCHW: the mask plane, then the three masked colour planes.
|
||||
let mut input = vec![0.0f32; 4 * n];
|
||||
for i in 0..n {
|
||||
let m = if known[i] { 1.0 } else { 0.0 };
|
||||
input[i] = m - 0.5;
|
||||
for c in 0..3 {
|
||||
input[(c + 1) * n + i] = (rgb[i * 3 + c] * 2.0 - 1.0) * m;
|
||||
}
|
||||
}
|
||||
let tensor = ort::value::Tensor::from_array(
|
||||
ndarray::Array::from_shape_vec(ndarray::IxDyn(&[1, 4, TILE, TILE]), input)
|
||||
.expect("shape matches by construction"),
|
||||
)?;
|
||||
let started = std::time::Instant::now();
|
||||
let acquired = self.model.acquire()?;
|
||||
let acquired_at = started.elapsed();
|
||||
let mut session = acquired.lock();
|
||||
let outputs = session.run(ort::inputs![tensor])?;
|
||||
log::trace!(
|
||||
"migan: tile on {} — acquire {:.1} ms, run {:.1} ms",
|
||||
acquired.rung().label(),
|
||||
acquired_at.as_secs_f64() * 1e3,
|
||||
(started.elapsed() - acquired_at).as_secs_f64() * 1e3
|
||||
);
|
||||
let (shape, data) = outputs[0].try_extract_tensor::<f32>()?;
|
||||
let dims: Vec<i64> = shape.iter().copied().collect();
|
||||
if dims != [1, 3, TILE as i64, TILE as i64] {
|
||||
return Err(PanoError::Model(format!(
|
||||
"MI-GAN output is {dims:?}, expected [1, 3, {TILE}, {TILE}]"
|
||||
)));
|
||||
}
|
||||
let mut out = vec![0.0f32; n * 3];
|
||||
for i in 0..n {
|
||||
for c in 0..3 {
|
||||
out[i * 3 + c] = (data[c * n + i] * 0.5 + 0.5).clamp(0.0, 1.0);
|
||||
}
|
||||
}
|
||||
Ok(out)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,197 @@
|
||||
//! TRACES: FR-MRG-4
|
||||
//! The surface the composite is drawn on.
|
||||
//!
|
||||
//! A panorama is a set of directions; a picture is a plane. The projection
|
||||
//! is the map between them, and the three offered are the three every
|
||||
//! stitcher offers because each is right for a different field of view:
|
||||
//! perspective keeps straight lines straight and cannot reach 180°;
|
||||
//! cylindrical keeps verticals vertical and stretches nothing horizontally,
|
||||
//! for the wide single row; spherical for anything that also looks up.
|
||||
//!
|
||||
//! Every function here is the *inverse* map — output pixel to direction —
|
||||
//! because that is what a gather needs (`lens.rs` in `dr-pipeline` says
|
||||
//! why a warp is written that way), and it is the function the WGSL warp
|
||||
//! will repeat verbatim. The forward map exists for bounds only.
|
||||
|
||||
use crate::linalg::Vec3;
|
||||
|
||||
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
|
||||
pub enum Projection {
|
||||
Perspective,
|
||||
Cylindrical,
|
||||
Spherical,
|
||||
}
|
||||
|
||||
impl Projection {
|
||||
/// Which projection a field of view calls for.
|
||||
///
|
||||
/// Perspective stretches the edges by `1 / cos` of the angle from the
|
||||
/// centre, which is 2× at 60° and unbounded at 90°; the switch is where
|
||||
/// that stretch starts to look like a mistake. Spherical is for a set
|
||||
/// that spans enough vertically that a cylinder would stretch the top
|
||||
/// and bottom the same way.
|
||||
pub fn suggest(horizontal_fov: f64, vertical_fov: f64) -> Projection {
|
||||
if horizontal_fov < 70f64.to_radians() && vertical_fov < 70f64.to_radians() {
|
||||
Projection::Perspective
|
||||
} else if vertical_fov < 100f64.to_radians() {
|
||||
Projection::Cylindrical
|
||||
} else {
|
||||
Projection::Spherical
|
||||
}
|
||||
}
|
||||
|
||||
/// The direction an output point looks along. `scale` is the output's
|
||||
/// focal length in pixels: the radius of the cylinder or sphere, or the
|
||||
/// plane's distance. Coordinates are centred on the projection's origin
|
||||
/// (the direction `+z`).
|
||||
pub fn to_direction(self, scale: f64, u: f64, v: f64) -> Vec3 {
|
||||
match self {
|
||||
Projection::Perspective => Vec3::new(u, v, scale).normalised(),
|
||||
Projection::Cylindrical => {
|
||||
let theta = u / scale;
|
||||
Vec3::new(theta.sin(), v / scale, theta.cos()).normalised()
|
||||
}
|
||||
Projection::Spherical => {
|
||||
let theta = u / scale;
|
||||
let phi = v / scale;
|
||||
Vec3::new(theta.sin() * phi.cos(), phi.sin(), theta.cos() * phi.cos())
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Where a direction lands on the output, or `None` where the
|
||||
/// projection cannot show it (behind a perspective plane, at a
|
||||
/// cylinder's poles).
|
||||
pub fn from_direction(self, scale: f64, d: Vec3) -> Option<(f64, f64)> {
|
||||
let (x, y, z) = (d.x(), d.y(), d.z());
|
||||
match self {
|
||||
Projection::Perspective => (z > 1e-9).then(|| (scale * x / z, scale * y / z)),
|
||||
Projection::Cylindrical => {
|
||||
let r = (x * x + z * z).sqrt();
|
||||
(r > 1e-9).then(|| (scale * x.atan2(z), scale * y / r))
|
||||
}
|
||||
Projection::Spherical => {
|
||||
let r = (x * x + z * z).sqrt();
|
||||
Some((scale * x.atan2(z), scale * y.atan2(r)))
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// The output rectangle a set of frames covers, in centred output pixels.
|
||||
#[derive(Debug, Clone, Copy, PartialEq)]
|
||||
pub struct Bounds {
|
||||
pub min_u: f64,
|
||||
pub min_v: f64,
|
||||
pub max_u: f64,
|
||||
pub max_v: f64,
|
||||
}
|
||||
|
||||
impl Bounds {
|
||||
pub fn width(&self) -> f64 {
|
||||
self.max_u - self.min_u
|
||||
}
|
||||
pub fn height(&self) -> f64 {
|
||||
self.max_v - self.min_v
|
||||
}
|
||||
}
|
||||
|
||||
/// Bounds of the frames' footprints under `projection`, by walking each
|
||||
/// frame's border.
|
||||
///
|
||||
/// `frame_size` is the frames' width and height in the same pixels the
|
||||
/// cameras' focal length is in. The border is sampled rather than only its
|
||||
/// corners because under a cylinder the widest point of a rolled frame is
|
||||
/// not a corner.
|
||||
pub fn bounds(
|
||||
projection: Projection,
|
||||
scale: f64,
|
||||
cameras: &crate::bundle::Cameras,
|
||||
frame_size: (f64, f64),
|
||||
) -> Option<Bounds> {
|
||||
let (w, h) = frame_size;
|
||||
let mut b: Option<Bounds> = None;
|
||||
let steps = 64;
|
||||
for k in 0..cameras.rotations.len() {
|
||||
for s in 0..steps {
|
||||
let t = s as f64 / steps as f64;
|
||||
for p in [
|
||||
(-w / 2.0 + w * t, -h / 2.0),
|
||||
(-w / 2.0 + w * t, h / 2.0),
|
||||
(-w / 2.0, -h / 2.0 + h * t),
|
||||
(w / 2.0, -h / 2.0 + h * t),
|
||||
] {
|
||||
let d = cameras.bearing(k, p);
|
||||
let Some((u, v)) = projection.from_direction(scale, d) else {
|
||||
continue;
|
||||
};
|
||||
b = Some(match b {
|
||||
None => Bounds {
|
||||
min_u: u,
|
||||
min_v: v,
|
||||
max_u: u,
|
||||
max_v: v,
|
||||
},
|
||||
Some(b) => Bounds {
|
||||
min_u: b.min_u.min(u),
|
||||
min_v: b.min_v.min(v),
|
||||
max_u: b.max_u.max(u),
|
||||
max_v: b.max_v.max(v),
|
||||
},
|
||||
});
|
||||
}
|
||||
}
|
||||
}
|
||||
b
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
#[test]
|
||||
fn to_and_from_direction_are_inverses() {
|
||||
for proj in [
|
||||
Projection::Perspective,
|
||||
Projection::Cylindrical,
|
||||
Projection::Spherical,
|
||||
] {
|
||||
for (u, v) in [(0.0, 0.0), (300.0, -200.0), (-900.0, 450.0)] {
|
||||
let d = proj.to_direction(1000.0, u, v);
|
||||
let (bu, bv) = proj.from_direction(1000.0, d).expect("in front");
|
||||
assert!(
|
||||
(bu - u).abs() < 1e-9 && (bv - v).abs() < 1e-9,
|
||||
"{proj:?} {u} {v}"
|
||||
);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn the_origin_looks_down_z_in_every_projection() {
|
||||
for proj in [
|
||||
Projection::Perspective,
|
||||
Projection::Cylindrical,
|
||||
Projection::Spherical,
|
||||
] {
|
||||
let d = proj.to_direction(500.0, 0.0, 0.0);
|
||||
assert!((d.z() - 1.0).abs() < 1e-12);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_cylinder_maps_ninety_degrees_to_a_quarter_turn_of_pixels() {
|
||||
let d = Vec3::new(1.0, 0.0, 0.0);
|
||||
let (u, v) = Projection::Cylindrical.from_direction(100.0, d).unwrap();
|
||||
assert!((u - 100.0 * std::f64::consts::FRAC_PI_2).abs() < 1e-9);
|
||||
assert_eq!(v, 0.0);
|
||||
assert!(Projection::Perspective.from_direction(100.0, d).is_none());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn suggestion_widens_with_the_field() {
|
||||
assert_eq!(Projection::suggest(0.5, 0.5), Projection::Perspective);
|
||||
assert_eq!(Projection::suggest(2.5, 0.8), Projection::Cylindrical);
|
||||
assert_eq!(Projection::suggest(3.0, 2.5), Projection::Spherical);
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,151 @@
|
||||
//! TRACES: FR-MRG-8
|
||||
//! The XFeat detector — the network under tract, and the decoder after it.
|
||||
//!
|
||||
//! Apache-2.0 weights (`models/LICENCE.md`), exported at a fixed shape by
|
||||
//! `tools/export-xfeat.sh` and loaded through the same `dr-inference-engine`
|
||||
//! `dr-segment` and `dr-face` use, so this adds no runtime and no C to the
|
||||
//! tree; what runs it is the device's business (docs/inference.md). ~300 ms
|
||||
//! per frame on tract on the reference desktop, ~400 ms on the tablet
|
||||
//! (S15.2, S15.4).
|
||||
|
||||
use crate::features::{decode_xfeat, DecodeOptions, Features, XFeatMaps, DESCRIPTOR_LEN};
|
||||
use crate::image::Gray;
|
||||
use crate::PanoError;
|
||||
|
||||
/// The two input shapes the shipped exports were made for: one landscape,
|
||||
/// one portrait, the same weights. A frame is fitted into whichever
|
||||
/// matches its aspect, so a portrait set does not spend half the
|
||||
/// detector's width on padding — which is what the 6D fixture did before
|
||||
/// the second export existed (512 × 768 of a 1024 × 768 input). A
|
||||
/// different size is a different file (`tools/export-xfeat.sh`).
|
||||
pub const INPUT_LANDSCAPE: (usize, usize) = (1024, 768);
|
||||
pub const INPUT_PORTRAIT: (usize, usize) = (768, 1024);
|
||||
|
||||
/// The long edge of the detector's input, for callers sizing a proxy.
|
||||
pub const INPUT_LONG_EDGE: usize = 1024;
|
||||
|
||||
#[cfg(feature = "embedded-model")]
|
||||
const EMBEDDED_LANDSCAPE: &[u8] = include_bytes!("../../../models/keypoints/xfeat-1024.onnx");
|
||||
#[cfg(feature = "embedded-model")]
|
||||
const EMBEDDED_PORTRAIT: &[u8] = include_bytes!("../../../models/keypoints/xfeat-768.onnx");
|
||||
|
||||
/// A loaded detector: the network at both shapes.
|
||||
pub struct XFeat {
|
||||
landscape: dr_inference_engine::Model,
|
||||
portrait: dr_inference_engine::Model,
|
||||
pub options: DecodeOptions,
|
||||
}
|
||||
|
||||
/// The bytes of both exports compiled into the binary, for whoever compiles
|
||||
/// engines ahead of the first request (docs/inference.md §6).
|
||||
#[cfg(feature = "embedded-model")]
|
||||
pub fn embedded_model_bytes() -> [&'static [u8]; 2] {
|
||||
[EMBEDDED_LANDSCAPE, EMBEDDED_PORTRAIT]
|
||||
}
|
||||
|
||||
impl XFeat {
|
||||
/// The weights compiled into the binary.
|
||||
#[cfg(feature = "embedded-model")]
|
||||
pub fn embedded() -> Result<Self, PanoError> {
|
||||
Self::from_bytes(EMBEDDED_LANDSCAPE, EMBEDDED_PORTRAIT)
|
||||
}
|
||||
|
||||
/// From the two exports on disk.
|
||||
pub fn from_paths(
|
||||
landscape: &std::path::Path,
|
||||
portrait: &std::path::Path,
|
||||
) -> Result<Self, PanoError> {
|
||||
let l = std::fs::read(landscape).map_err(PanoError::ModelRead)?;
|
||||
let p = std::fs::read(portrait).map_err(PanoError::ModelRead)?;
|
||||
Self::from_bytes(&l, &p)
|
||||
}
|
||||
|
||||
pub fn from_bytes(landscape: &[u8], portrait: &[u8]) -> Result<Self, PanoError> {
|
||||
use dr_inference_engine::{Form, Role};
|
||||
Ok(XFeat {
|
||||
landscape: dr_inference_engine::open(Role::Keypoints, Form::F32, landscape)?,
|
||||
portrait: dr_inference_engine::open(Role::Keypoints, Form::F32, portrait)?,
|
||||
options: DecodeOptions::default(),
|
||||
})
|
||||
}
|
||||
|
||||
/// Detect keypoints in an upright grayscale image.
|
||||
///
|
||||
/// The image is fitted into the network's input of matching aspect —
|
||||
/// scaled down if larger, never up, and padded to the right and bottom
|
||||
/// — and the keypoints come back in the coordinates of `image` itself,
|
||||
/// so a caller that already scaled a frame to a proxy maps them on with
|
||||
/// the scale it used and nothing else.
|
||||
pub fn detect(&mut self, image: &Gray) -> Result<Features, PanoError> {
|
||||
let ((in_w, in_h), model) = if image.height > image.width {
|
||||
(INPUT_PORTRAIT, &self.portrait)
|
||||
} else {
|
||||
(INPUT_LANDSCAPE, &self.landscape)
|
||||
};
|
||||
let acquired = model.acquire()?;
|
||||
let mut session = acquired.lock();
|
||||
let (fitted, scale) = image.fitted(in_w, in_h);
|
||||
let padded = fitted.padded(in_w, in_h);
|
||||
|
||||
let input =
|
||||
ndarray::Array::from_shape_vec(ndarray::IxDyn(&[1, 1, in_h, in_w]), padded.data)
|
||||
.expect("shape matches the buffer by construction");
|
||||
let tensor = ort::value::Tensor::from_array(input).map_err(PanoError::Inference)?;
|
||||
let outputs = session
|
||||
.run(ort::inputs![tensor])
|
||||
.map_err(PanoError::Inference)?;
|
||||
|
||||
let (w8, h8) = (in_w / 8, in_h / 8);
|
||||
let expect = |i: usize, channels: usize| -> Result<Vec<f32>, PanoError> {
|
||||
let (shape, data) = outputs[i]
|
||||
.try_extract_tensor::<f32>()
|
||||
.map_err(PanoError::Inference)?;
|
||||
let dims: Vec<i64> = shape.iter().copied().collect();
|
||||
if dims != [1, channels as i64, h8 as i64, w8 as i64] {
|
||||
return Err(PanoError::Model(format!(
|
||||
"output {i} is {dims:?}, expected [1, {channels}, {h8}, {w8}] — \
|
||||
not the export this decoder was written for"
|
||||
)));
|
||||
}
|
||||
Ok(data.to_vec())
|
||||
};
|
||||
let feats = expect(0, DESCRIPTOR_LEN)?;
|
||||
let keypoints = expect(1, 65)?;
|
||||
let heatmap = expect(2, 1)?;
|
||||
|
||||
let mut features = decode_xfeat(
|
||||
&XFeatMaps {
|
||||
feats: &feats,
|
||||
keypoints: &keypoints,
|
||||
heatmap: &heatmap,
|
||||
width: w8,
|
||||
height: h8,
|
||||
},
|
||||
&self.options,
|
||||
);
|
||||
|
||||
// Back to the caller's image: drop anything the padding produced,
|
||||
// undo the fit.
|
||||
let border = self.options.border as f32;
|
||||
let limit_x = fitted.width as f32 - border;
|
||||
let limit_y = fitted.height as f32 - border;
|
||||
let mut kept_kp = Vec::with_capacity(features.len());
|
||||
let mut kept_desc = Vec::with_capacity(features.descriptors.len());
|
||||
for (i, kp) in features.keypoints.iter().enumerate() {
|
||||
if kp.x >= limit_x || kp.y >= limit_y {
|
||||
continue;
|
||||
}
|
||||
kept_kp.push(crate::features::Keypoint {
|
||||
x: (kp.x / scale as f32),
|
||||
y: (kp.y / scale as f32),
|
||||
score: kp.score,
|
||||
});
|
||||
kept_desc.extend_from_slice(features.descriptor(i));
|
||||
}
|
||||
features.keypoints = kept_kp;
|
||||
features.descriptors = kept_desc;
|
||||
features.width = image.width;
|
||||
features.height = image.height;
|
||||
Ok(features)
|
||||
}
|
||||
}
|
||||
@@ -849,6 +849,14 @@ impl EditGraph {
|
||||
)
|
||||
}
|
||||
|
||||
/// TRACES: FR-MRG-2
|
||||
/// The camera-space tap for a merge: this edit's lens corrections and
|
||||
/// nothing else of it, stored at full precision. See
|
||||
/// [`crate::operation::compose_camera_linear`].
|
||||
pub fn compose_camera_linear(&self, view: crate::framing::CropRect) -> ComposedShader {
|
||||
crate::operation::compose_camera_linear(&self.warps, self.framing.baseline(), view)
|
||||
}
|
||||
|
||||
/// TRACES: FR-DEV-19c
|
||||
/// [`Self::compose_for`], with one layer's mask drawn over the picture.
|
||||
///
|
||||
|
||||
@@ -555,6 +555,7 @@ impl Morphology {
|
||||
];
|
||||
}
|
||||
|
||||
/// TRACES: FR-DEV-3 | FR-DEV-3i
|
||||
/// Where a mask layer applies.
|
||||
#[derive(Debug, Clone, PartialEq)]
|
||||
pub enum MaskSource {
|
||||
@@ -1303,7 +1304,12 @@ impl PartialEq for MaskPart {
|
||||
}
|
||||
}
|
||||
|
||||
/// TRACES: FR-DEV-19
|
||||
/// One local adjustment: a rule about *where*, plus a chain saying *what*.
|
||||
///
|
||||
/// The *where* is editable after the fact — composed from parts (FR-DEV-19a),
|
||||
/// painted into and out of (FR-DEV-19b), shown (FR-DEV-19c) — and every edit
|
||||
/// to it is geometry and parameters in this struct, never a raster in a file.
|
||||
pub struct MaskLayer {
|
||||
/// Stable identity, for the sidecar and for merge (FR-NC-9).
|
||||
pub id: String,
|
||||
|
||||
@@ -424,6 +424,21 @@ pub enum OutputMode {
|
||||
/// clamping before a sharpener sees it would draw a hard edge at precisely
|
||||
/// the luminance a sharpener is most visible at.
|
||||
LinearWorking,
|
||||
/// TRACES: FR-MRG-2
|
||||
/// `rgba32float`, **camera space**: after the lens warp and nothing else.
|
||||
///
|
||||
/// What a merge stitches (FR-MRG-2). The shader is the linear tail with
|
||||
/// no operations, and the caller fills the reserved uniforms neutral —
|
||||
/// unit white balance, identity matrix, base curve off — so what is
|
||||
/// stored is the sensor's own numbers, demosaiced and undistorted. Only
|
||||
/// [`compose_camera_linear`] produces it, and only
|
||||
/// `AdjustPass::render_camera_linear` accepts it, so the neutral
|
||||
/// uniforms cannot be forgotten by a caller that composed it by mistake.
|
||||
///
|
||||
/// Thirty-two bits rather than sixteen because the composite is written
|
||||
/// back as a RAW at the sensor's scale (FR-MRG-3): a 14-bit sensor has
|
||||
/// 16 384 steps to white and `f16` keeps 2 048 of them in the top octave.
|
||||
CameraLinear,
|
||||
}
|
||||
|
||||
/// The result of composing a set of operations into one shader.
|
||||
@@ -573,6 +588,54 @@ pub fn compose_full_revealing(
|
||||
spots: &crate::spot::SpotSet,
|
||||
warps: &[Box<dyn crate::lens::Warp>],
|
||||
reveal: Option<&crate::mask::Reveal>,
|
||||
) -> ComposedShader {
|
||||
compose_inner(ops, framing, output, masks, spots, warps, reveal, None)
|
||||
}
|
||||
|
||||
/// TRACES: FR-MRG-2
|
||||
/// The camera-space tap: the fused pass with no operations, stopping after
|
||||
/// the lens warp and storing `rgba32float` ([`OutputMode::CameraLinear`]).
|
||||
///
|
||||
/// Takes the warps and a view, because that is all the tap uses of an edit:
|
||||
/// no crop (a merge wants the whole frame, and crops the composite), no
|
||||
/// masks, no spots, no operations. `view` is the tile — the fraction of the
|
||||
/// undistorted frame to render, in `Framing::set_view`'s terms — so that a
|
||||
/// merge pulls source tiles on demand (FR-MRG-11) rather than a frame that
|
||||
/// may not fit. The camera profile's uniforms are still declared — the
|
||||
/// prologue is the same — and the GPU side fills them neutral.
|
||||
pub fn compose_camera_linear(
|
||||
warps: &[Box<dyn crate::lens::Warp>],
|
||||
baseline: dr_types::Orientation,
|
||||
view: crate::framing::CropRect,
|
||||
) -> ComposedShader {
|
||||
let mut framing = Framing::default();
|
||||
// The file's orientation and nothing of the user's: a merge aligned
|
||||
// its frames upright (`dr_pano::Gray::oriented`), so its tiles must be
|
||||
// upright too, and `view` is a fraction of the upright frame.
|
||||
framing.set_baseline(baseline);
|
||||
framing.set_view(view);
|
||||
compose_inner(
|
||||
&[],
|
||||
&framing,
|
||||
ColourSpace::Srgb,
|
||||
&MaskStack::new(),
|
||||
&crate::spot::SpotSet::new(),
|
||||
warps,
|
||||
None,
|
||||
Some(OutputMode::CameraLinear),
|
||||
)
|
||||
}
|
||||
|
||||
#[allow(clippy::too_many_arguments)]
|
||||
fn compose_inner(
|
||||
ops: &[Box<dyn Operation>],
|
||||
framing: &Framing,
|
||||
output: ColourSpace,
|
||||
masks: &MaskStack,
|
||||
spots: &crate::spot::SpotSet,
|
||||
warps: &[Box<dyn crate::lens::Warp>],
|
||||
reveal: Option<&crate::mask::Reveal>,
|
||||
forced: Option<OutputMode>,
|
||||
) -> ComposedShader {
|
||||
// The lens corrections, composed into one coordinate transform. Beside
|
||||
// `framing` because they are the other half of the same stage: framing
|
||||
@@ -603,12 +666,18 @@ pub fn compose_full_revealing(
|
||||
// all: a photograph with a spot on it and no sharpening still has a detail
|
||||
// stage, and a fused pass that encoded its own output there would quantise
|
||||
// twice and be bound to a texture of the wrong format.
|
||||
let output_mode =
|
||||
//
|
||||
// `forced` is the one exception, and it is not a caller flag in the
|
||||
// sense above: `compose_camera_linear` is the only function that passes
|
||||
// it, with an empty operation list, and the mode it forces has its own
|
||||
// storage format and its own render entry on the GPU side.
|
||||
let output_mode = forced.unwrap_or(
|
||||
if ops.iter().any(|o| o.is_active() && o.detail().is_some()) || !spots.is_neutral() {
|
||||
OutputMode::LinearWorking
|
||||
} else {
|
||||
OutputMode::Encoded
|
||||
};
|
||||
},
|
||||
);
|
||||
|
||||
// Whether an operation has taken over the rendering. Decided from the
|
||||
// operations for the same reason `output_mode` is: a caller that got it
|
||||
@@ -788,12 +857,22 @@ pub fn compose_full_revealing(
|
||||
""
|
||||
};
|
||||
|
||||
// TRACES: FR-MRG-2
|
||||
// A pixel whose source coordinate leaves the frame — the corners a lens
|
||||
// correction pulls in — is stored black. For the display that is the
|
||||
// right picture; for a merge it is a pixel that does not exist and must
|
||||
// not be averaged in as if it did, so the camera-space tap marks it
|
||||
// with alpha 0 and the warp reads the alpha as validity.
|
||||
let void_alpha = match output_mode {
|
||||
OutputMode::CameraLinear => "0.0",
|
||||
_ => "1.0",
|
||||
};
|
||||
let prologue = format!(
|
||||
"{}{}{}\n{}",
|
||||
framing.wgsl_prologue(),
|
||||
channel_positions,
|
||||
warp.body,
|
||||
sample_source(interpolate, warp.splits_channels)
|
||||
sample_source(interpolate, warp.splits_channels).replace("VOID_ALPHA", void_alpha)
|
||||
);
|
||||
let sampler_helper = if interpolate { BILINEAR_HELPER } else { "" };
|
||||
|
||||
@@ -827,6 +906,16 @@ pub fn compose_full_revealing(
|
||||
// would put a hard edge into the very neighbourhood the next pass is
|
||||
// about to convolve, which is how sharpeners come to draw dark rings
|
||||
// around specular highlights.
|
||||
textureStore(output, vec2<i32>(gid.xy), vec4<f32>(c, 1.0));"
|
||||
.to_string(),
|
||||
),
|
||||
OutputMode::CameraLinear => (
|
||||
"rgba32float",
|
||||
String::new(),
|
||||
String::new(),
|
||||
" // Camera space, for a merge (FR-MRG-2): the sensor's numbers after the
|
||||
// lens warp, with the profile uniforms filled neutral by the caller so the
|
||||
// prologue above changed nothing. Not clamped, not encoded, full precision.
|
||||
textureStore(output, vec2<i32>(gid.xy), vec4<f32>(c, 1.0));"
|
||||
.to_string(),
|
||||
),
|
||||
@@ -1158,7 +1247,7 @@ pub(crate) fn sample_source(interpolate: bool, splits_channels: bool) -> &'stati
|
||||
// remove. `sample_bilinear` clamps its own texel indices, so red and blue
|
||||
// land on the edge pixel rather than out of bounds.
|
||||
if (any(uv_src < vec2<f32>(0.0)) || any(uv_src >= vec2<f32>(1.0))) {
|
||||
textureStore(output, vec2<i32>(gid.xy), vec4<f32>(0.0, 0.0, 0.0, 1.0));
|
||||
textureStore(output, vec2<i32>(gid.xy), vec4<f32>(0.0, 0.0, 0.0, VOID_ALPHA));
|
||||
return;
|
||||
}
|
||||
|
||||
@@ -1186,7 +1275,7 @@ pub(crate) fn sample_source(interpolate: bool, splits_channels: bool) -> &'stati
|
||||
// corners; render them black rather than clamping, which would smear an
|
||||
// edge pixel across them.
|
||||
if (any(uv_src < vec2<f32>(0.0)) || any(uv_src >= vec2<f32>(1.0))) {
|
||||
textureStore(output, vec2<i32>(gid.xy), vec4<f32>(0.0, 0.0, 0.0, 1.0));
|
||||
textureStore(output, vec2<i32>(gid.xy), vec4<f32>(0.0, 0.0, 0.0, VOID_ALPHA));
|
||||
return;
|
||||
}
|
||||
|
||||
@@ -1229,7 +1318,7 @@ pub(crate) fn sample_source(interpolate: bool, splits_channels: bool) -> &'stati
|
||||
// transformed at all. Render it black rather than clamping, which would
|
||||
// smear an edge pixel across the gap.
|
||||
if (any(uv_src < vec2<f32>(0.0)) || any(uv_src >= vec2<f32>(1.0))) {
|
||||
textureStore(output, vec2<i32>(gid.xy), vec4<f32>(0.0, 0.0, 0.0, 1.0));
|
||||
textureStore(output, vec2<i32>(gid.xy), vec4<f32>(0.0, 0.0, 0.0, VOID_ALPHA));
|
||||
return;
|
||||
}
|
||||
|
||||
@@ -2030,6 +2119,35 @@ mod tests {
|
||||
assert!(convert < clip, "the clip must come after the conversion");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn the_camera_space_tap_marks_a_pixel_off_the_sensor_with_alpha_zero() {
|
||||
// TRACES: FR-MRG-2
|
||||
// A lens correction pulls the corners in, and the pixels it leaves
|
||||
// behind have no source. The display stores them black and opaque;
|
||||
// the tap stores them black and *transparent*, so a merge can tell
|
||||
// "nothing here" from "black here" and never averages the fringe in.
|
||||
let tap = compose_camera_linear(
|
||||
&[],
|
||||
dr_types::Orientation::default(),
|
||||
crate::framing::CropRect::default(),
|
||||
)
|
||||
.source;
|
||||
assert!(
|
||||
tap.contains("vec4<f32>(0.0, 0.0, 0.0, 0.0)"),
|
||||
"the tap must store alpha 0 off the sensor:\n{tap}"
|
||||
);
|
||||
assert!(
|
||||
!tap.contains("VOID_ALPHA"),
|
||||
"the placeholder must be substituted:\n{tap}"
|
||||
);
|
||||
let display = compose(&[]).source;
|
||||
assert!(
|
||||
!display.contains("vec4<f32>(0.0, 0.0, 0.0, 0.0)"),
|
||||
"the display keeps its opaque black:\n{display}"
|
||||
);
|
||||
assert!(!display.contains("VOID_ALPHA"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn the_generated_matrix_is_the_one_the_profile_writer_will_use() {
|
||||
// The shader encodes the pixels and `dr-export` describes them, from
|
||||
|
||||
@@ -110,6 +110,21 @@ pub struct Sidecar {
|
||||
///
|
||||
/// Preserved so a older build round-trips a newer file without loss.
|
||||
unknown_blocks: Vec<String>,
|
||||
/// TRACES: FR-MRG-6
|
||||
/// Provenance: the sources a composite was merged from, in order, as
|
||||
/// the library names them. Empty for a photograph the camera took.
|
||||
///
|
||||
/// Top-level rather than per version because it is a fact about the
|
||||
/// file, not about an edit — every version of a panorama is a version
|
||||
/// of the same twelve frames. Written as one `derived_from = …` line per
|
||||
/// source before the first version block, where a build that predates
|
||||
/// the field keeps the lines as unknown and writes them back untouched.
|
||||
pub derived_from: Vec<String>,
|
||||
/// TRACES: FR-MRG-6
|
||||
/// How the composite was made: `panorama cylindrical 49.7mm 12 frames`.
|
||||
/// Free text for a panel; the parameters that matter to a re-merge are
|
||||
/// the sources and the projection, and both are legible in it.
|
||||
pub merge: Option<String>,
|
||||
}
|
||||
|
||||
/// TRACES: FR-DEV-3f
|
||||
@@ -823,6 +838,13 @@ impl Sidecar {
|
||||
/// caller may compare content to decide whether an upload is needed.
|
||||
pub fn to_text(&self) -> String {
|
||||
let mut out = format!("drsc {FORMAT_VERSION}\n");
|
||||
// TRACES: FR-MRG-6
|
||||
for source in &self.derived_from {
|
||||
let _ = writeln!(out, "derived_from = {source}");
|
||||
}
|
||||
if let Some(merge) = &self.merge {
|
||||
let _ = writeln!(out, "merge = {merge}");
|
||||
}
|
||||
for block in &self.unknown_blocks {
|
||||
let _ = writeln!(out, "{block}");
|
||||
}
|
||||
@@ -1008,7 +1030,12 @@ impl Sidecar {
|
||||
}
|
||||
|
||||
let Some(version) = current.as_mut() else {
|
||||
sidecar.unknown_blocks.push(line.to_string());
|
||||
// TRACES: FR-MRG-6
|
||||
match key {
|
||||
"derived_from" => sidecar.derived_from.push(value.to_string()),
|
||||
"merge" => sidecar.merge = Some(value.to_string()),
|
||||
_ => sidecar.unknown_blocks.push(line.to_string()),
|
||||
}
|
||||
continue;
|
||||
};
|
||||
|
||||
@@ -2209,6 +2236,26 @@ mod tests {
|
||||
/// the claim is about *those bytes*: a sidecar generated by this build
|
||||
/// would agree with this build by construction, and would go on agreeing
|
||||
/// with it through a rename that broke every file on disk.
|
||||
#[test]
|
||||
fn provenance_round_trips_at_the_top_level() {
|
||||
// TRACES: FR-MRG-6
|
||||
let text = "drsc 1\nderived_from = 2025/_MG_8320.CR2\nderived_from = 2025/_MG_8321.CR2\n\
|
||||
merge = panorama cylindrical 49.7mm 2 frames\n\n[version u1]\nname = Default\n\
|
||||
revision = 1\nmodified = 0\n";
|
||||
let sidecar = Sidecar::parse(text).expect("parses");
|
||||
assert_eq!(
|
||||
sidecar.derived_from,
|
||||
vec!["2025/_MG_8320.CR2", "2025/_MG_8321.CR2"]
|
||||
);
|
||||
assert_eq!(
|
||||
sidecar.merge.as_deref(),
|
||||
Some("panorama cylindrical 49.7mm 2 frames")
|
||||
);
|
||||
let out = sidecar.to_text();
|
||||
assert!(out.starts_with("drsc 1\nderived_from = 2025/_MG_8320.CR2\n"));
|
||||
assert_eq!(Sidecar::parse(&out).expect("re-parses"), sidecar);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_sidecar_from_before_the_channel_curves_still_names_the_master() {
|
||||
let text = "drsc 1\n\n[version u1]\nname = Default\nrevision = 4\nmodified = 9\n\
|
||||
|
||||
@@ -11,10 +11,11 @@ build = "build.rs"
|
||||
thiserror.workspace = true
|
||||
log.workspace = true
|
||||
|
||||
# Inference. `ort` is the API; **tract is the engine** — see the workspace
|
||||
# manifest for why the C++ ONNX Runtime is not linked here.
|
||||
# Inference. `ort` is the API; **what runs it is `dr-inference-engine`'s
|
||||
# business** — tract, or an ONNX Runtime the app found on disk, on whichever
|
||||
# provider the device has (docs/inference.md). This crate never names either.
|
||||
ort = { workspace = true, optional = true }
|
||||
ort-tract = { workspace = true, optional = true }
|
||||
dr-inference-engine = { workspace = true, optional = true }
|
||||
ndarray = { workspace = true, optional = true }
|
||||
|
||||
[dev-dependencies]
|
||||
@@ -23,6 +24,9 @@ ndarray = { workspace = true, optional = true }
|
||||
# tree for embedded previews.
|
||||
zune-jpeg.workspace = true
|
||||
env_logger.workspace = true
|
||||
# The probe example drives `ort` directly to print the raw load error, and
|
||||
# asks the engine for a runtime by name rather than naming one itself.
|
||||
dr-inference-engine.workspace = true
|
||||
|
||||
[features]
|
||||
# On by default: a local adjustment that cannot select a subject is half the
|
||||
@@ -35,7 +39,7 @@ default = ["semantic", "embedded-model"]
|
||||
# Separable because the watershed half is genuinely independent of it: with
|
||||
# this off, `dr-segment` is a pure-CPU graph algorithm crate with no model to
|
||||
# carry, which is what the headless hierarchy tests want.
|
||||
semantic = ["dep:ort", "dep:ort-tract", "dep:ndarray"]
|
||||
semantic = ["dep:ort", "dep:dr-inference-engine", "dep:ndarray"]
|
||||
|
||||
# Compile the weights into the binary.
|
||||
#
|
||||
|
||||
@@ -0,0 +1,87 @@
|
||||
//! TRACES: S15 | FR-MRG-8
|
||||
//! Load an ONNX file through the application's own runtime and run it once.
|
||||
//!
|
||||
//! ```sh
|
||||
//! cargo run -p dr-segment --example onnx_probe --release -- model.onnx [1x1x768x1024]
|
||||
//! ```
|
||||
//!
|
||||
//! The F6 check, as a tool. tract's operator coverage is the thing that can
|
||||
//! sink a model choice — `segmentation.md` records a dynamic-shape export it
|
||||
//! could not parse at all — and the only way to know is to load the file
|
||||
//! under the backend the app ships and see. This does that for any model,
|
||||
//! before any Rust is written against its outputs: it prints the declared
|
||||
//! inputs and outputs, runs zeros through at the given shape, and times it.
|
||||
//!
|
||||
//! Written for S15.2 (XFeat), kept because the next model will need it too.
|
||||
|
||||
use std::time::Instant;
|
||||
|
||||
fn main() {
|
||||
let mut args = std::env::args().skip(1);
|
||||
let Some(path) = args.next() else {
|
||||
eprintln!("usage: onnx_probe <model.onnx> [NxCxHxW]");
|
||||
std::process::exit(2);
|
||||
};
|
||||
let shape: Vec<usize> = args
|
||||
.next()
|
||||
.map(|s| {
|
||||
s.split('x')
|
||||
.map(|d| d.parse().expect("dimension"))
|
||||
.collect()
|
||||
})
|
||||
.unwrap_or_else(|| vec![1, 1, 768, 1024]);
|
||||
|
||||
let bytes = std::fs::read(&path).expect("read model");
|
||||
println!("{path}: {} bytes", bytes.len());
|
||||
|
||||
dr_inference_engine::ensure_runtime();
|
||||
let t = Instant::now();
|
||||
let mut session =
|
||||
match ort::session::Session::builder().and_then(|mut b| b.commit_from_memory(&bytes)) {
|
||||
Ok(s) => s,
|
||||
Err(e) => {
|
||||
println!("FAIL load: {e}");
|
||||
std::process::exit(1);
|
||||
}
|
||||
};
|
||||
println!("ok loaded in {:?}", t.elapsed());
|
||||
for i in session.inputs().iter() {
|
||||
println!(" input {} {:?}", i.name(), i.dtype());
|
||||
}
|
||||
for o in session.outputs().iter() {
|
||||
println!(" output {} {:?}", o.name(), o.dtype());
|
||||
}
|
||||
|
||||
let n: usize = shape.iter().product();
|
||||
|
||||
// Twice: the first run pays for tract's optimisation and plan, the second
|
||||
// is the number that matters. The tensor is built per run rather than
|
||||
// cloned — `Tensor::clone` under the tract backend panics.
|
||||
for pass in 1..=2 {
|
||||
let input =
|
||||
ndarray::Array::from_shape_vec(ndarray::IxDyn(&shape), vec![0.0f32; n]).expect("shape");
|
||||
let tensor = ort::value::Tensor::from_array(input).expect("tensor");
|
||||
let t = Instant::now();
|
||||
let outputs = match session.run(ort::inputs![tensor]) {
|
||||
Ok(o) => o,
|
||||
Err(e) => {
|
||||
println!("FAIL run: {e}");
|
||||
std::process::exit(1);
|
||||
}
|
||||
};
|
||||
println!("ok run {pass} in {:?}", t.elapsed());
|
||||
if pass == 2 {
|
||||
for i in 0..outputs.len() {
|
||||
match outputs[i].try_extract_tensor::<f32>() {
|
||||
Ok((shape, data)) => {
|
||||
let (lo, hi) = data
|
||||
.iter()
|
||||
.fold((f32::MAX, f32::MIN), |(lo, hi), &v| (lo.min(v), hi.max(v)));
|
||||
println!(" output {i}: shape {shape:?}, range {lo:.4}..{hi:.4}");
|
||||
}
|
||||
Err(e) => println!(" output {i}: not f32 ({e})"),
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -1,3 +1,4 @@
|
||||
//! TRACES: FR-DEV-3i
|
||||
//! Region segmentation for local masking (S15, docs/segmentation.md).
|
||||
//!
|
||||
//! Local adjustments need to know where the image's regions are before they
|
||||
@@ -70,6 +71,8 @@ pub use refine::{
|
||||
};
|
||||
#[cfg(feature = "semantic")]
|
||||
pub use scene::{Category, Scene, SceneModel};
|
||||
#[cfg(feature = "embedded-model")]
|
||||
pub use semantic::embedded_model_bytes;
|
||||
#[cfg(feature = "semantic")]
|
||||
pub use semantic::{Instance, SemanticModel, SemanticOptions, Tiling};
|
||||
|
||||
@@ -97,3 +100,13 @@ pub enum SegmentError {
|
||||
#[error("category descriptor: {0}")]
|
||||
CategoryDescriptor(String),
|
||||
}
|
||||
|
||||
#[cfg(feature = "semantic")]
|
||||
impl From<dr_inference_engine::Error> for SegmentError {
|
||||
fn from(e: dr_inference_engine::Error) -> Self {
|
||||
match e {
|
||||
dr_inference_engine::Error::Inference(e) => SegmentError::Inference(e),
|
||||
dr_inference_engine::Error::Io(e) => SegmentError::ModelRead(e),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -59,7 +59,7 @@ use ndarray::ArrayView3;
|
||||
|
||||
#[cfg(test)]
|
||||
use crate::semantic::INPUT_EDGE;
|
||||
use crate::semantic::{install_backend, Letterbox, Window};
|
||||
use crate::semantic::{Letterbox, Window};
|
||||
use crate::SegmentError;
|
||||
|
||||
/// Classes in the ADE20K vocabulary the scene model was trained on.
|
||||
@@ -84,7 +84,7 @@ pub struct Category {
|
||||
|
||||
/// The scene model, and the categories it has been told to report.
|
||||
pub struct SceneModel {
|
||||
session: ort::session::Session,
|
||||
session: dr_inference_engine::Model,
|
||||
categories: Vec<Category>,
|
||||
}
|
||||
|
||||
@@ -132,12 +132,12 @@ impl SceneModel {
|
||||
}
|
||||
|
||||
pub fn from_bytes(bytes: &[u8], categories: Vec<Category>) -> Result<Self, SegmentError> {
|
||||
install_backend();
|
||||
|
||||
let session = ort::session::Session::builder()
|
||||
.map_err(SegmentError::Inference)?
|
||||
.commit_from_memory(bytes)
|
||||
.map_err(SegmentError::Inference)?;
|
||||
// f32, as for `SemanticModel`; see there.
|
||||
let session = dr_inference_engine::open(
|
||||
dr_inference_engine::Role::Scene,
|
||||
dr_inference_engine::Form::F32,
|
||||
bytes,
|
||||
)?;
|
||||
|
||||
Ok(Self {
|
||||
session,
|
||||
@@ -170,13 +170,9 @@ impl SceneModel {
|
||||
});
|
||||
}
|
||||
|
||||
// Split the borrow: `run` needs the session mutably while
|
||||
// `marginalise` needs the categories, and going through `self` for
|
||||
// both at once is what the borrow checker objects to.
|
||||
let Self {
|
||||
session,
|
||||
categories,
|
||||
} = self;
|
||||
let categories = &self.categories;
|
||||
let acquired = self.session.acquire()?;
|
||||
let mut session = acquired.lock();
|
||||
|
||||
let window = Window {
|
||||
x: 0.0,
|
||||
|
||||
@@ -194,7 +194,7 @@ impl Instance {
|
||||
/// Holds an `ort` session, so it is neither `Clone` nor cheap to build —
|
||||
/// construct once and keep it. Loading is ~50 ms.
|
||||
pub struct SemanticModel {
|
||||
session: ort::session::Session,
|
||||
session: dr_inference_engine::Model,
|
||||
classes: Vec<Arc<str>>,
|
||||
}
|
||||
|
||||
@@ -208,6 +208,13 @@ const EMBEDDED_MODEL: &[u8] = include_bytes!("../../../models/segment/yolo26n-se
|
||||
#[cfg(feature = "embedded-model")]
|
||||
const EMBEDDED_CLASSES: &str = include_str!("../../../models/segment/yolo26n-seg.classes.json");
|
||||
|
||||
/// The bytes of the model that ships with this crate, for whoever compiles
|
||||
/// engines ahead of the first request (docs/inference.md §6).
|
||||
#[cfg(feature = "embedded-model")]
|
||||
pub fn embedded_model_bytes() -> &'static [u8] {
|
||||
EMBEDDED_MODEL
|
||||
}
|
||||
|
||||
impl SemanticModel {
|
||||
/// Load the model that ships with this crate.
|
||||
#[cfg(feature = "embedded-model")]
|
||||
@@ -229,15 +236,14 @@ impl SemanticModel {
|
||||
}
|
||||
|
||||
pub fn from_bytes(bytes: &[u8], classes: Vec<Arc<str>>) -> Result<Self, SegmentError> {
|
||||
// Idempotent, and it must happen before any other `ort` call: with
|
||||
// `alternative-backend` there is no linked runtime to fall back on, so
|
||||
// an un-set API is a panic rather than a slow path.
|
||||
install_backend();
|
||||
|
||||
let session = ort::session::Session::builder()
|
||||
.map_err(SegmentError::Inference)?
|
||||
.commit_from_memory(bytes)
|
||||
.map_err(SegmentError::Inference)?;
|
||||
// The f32 graph on whatever the device's backend is. An int8 form
|
||||
// for the Hexagon waits on docs/inference.md §10 M7 — the mask
|
||||
// boundary has to be measured before it moves.
|
||||
let session = dr_inference_engine::open(
|
||||
dr_inference_engine::Role::Segmenter,
|
||||
dr_inference_engine::Form::F32,
|
||||
bytes,
|
||||
)?;
|
||||
|
||||
Ok(Self { session, classes })
|
||||
}
|
||||
@@ -332,8 +338,9 @@ impl SemanticModel {
|
||||
let letterbox = Letterbox::fit(window.w, window.h);
|
||||
let input = letterbox.sample(rgb, width, height, window);
|
||||
|
||||
let outputs = self
|
||||
.session
|
||||
let acquired = self.session.acquire()?;
|
||||
let mut session = acquired.lock();
|
||||
let outputs = session
|
||||
.run(ort::inputs![
|
||||
ort::value::Tensor::from_array(input).map_err(SegmentError::Inference)?
|
||||
])
|
||||
@@ -674,18 +681,6 @@ fn steps(extent: f32, edge: f32, stride: f32) -> usize {
|
||||
}
|
||||
}
|
||||
|
||||
/// Point `ort` at tract, exactly once per process.
|
||||
pub(crate) fn install_backend() {
|
||||
use std::sync::Once;
|
||||
static ONCE: Once = Once::new();
|
||||
ONCE.call_once(|| {
|
||||
// Returns false if an API was already installed, which is not an error
|
||||
// — it means something else got here first, and there is only one
|
||||
// backend compiled in for it to have chosen.
|
||||
let _ = ort::set_api(ort_tract::api());
|
||||
});
|
||||
}
|
||||
|
||||
/// Read the class list written beside the model by `tools/export-seg-model.sh`.
|
||||
///
|
||||
/// A deliberately small hand-rolled reader for a flat array of strings, rather
|
||||
|
||||
@@ -126,6 +126,11 @@ pub struct StoredFilter {
|
||||
/// does not derive serde — and "all of them" is the only thing the second
|
||||
/// variant means.
|
||||
pub people_all: bool,
|
||||
/// TRACES: FR-CULL-8a | FR-CULL-13
|
||||
/// Whether the grid was narrowed to photographs with nobody blinking.
|
||||
/// Travels: it is a narrowing like `local_only`, and a record without it
|
||||
/// — from a build before it existed — reads as off.
|
||||
pub eyes_open: bool,
|
||||
}
|
||||
|
||||
/// Where the photographer was, at the moment they were there.
|
||||
|
||||
+167
-10
@@ -186,17 +186,27 @@ pub struct FaceSettings {
|
||||
/// and a tablet on a battery want different answers — so it is a setting,
|
||||
/// per device, like the rest of this file.
|
||||
///
|
||||
/// # A detector is half of a model id
|
||||
/// # A detector is half of a model id, and the half that does not split
|
||||
/// # the library
|
||||
///
|
||||
/// Every face row, run marker, sync shard and calibration is keyed by
|
||||
/// `faces.model_id` (catalog.md §10.1), and the schema's whole reason for
|
||||
/// carrying that column is that *a model change is a new id and a re-index*
|
||||
/// rather than a silent change under existing data. A detector change is a
|
||||
/// model change: it decides which faces exist and where the landmarks that
|
||||
/// align them land. So each variant names its own pipeline, and choosing
|
||||
/// another one puts every image back in the queue and shows the People screen
|
||||
/// for the new pipeline — empty until the sweep has run, with confirmed names
|
||||
/// carried across by `record_detections`' box overlap.
|
||||
/// Every face row, run marker, sync shard and calibration carries
|
||||
/// `faces.model_id` (catalog.md §10.1), spelled `detector+embedder`, so which
|
||||
/// pipeline produced a face is always on record. But every reader of "the
|
||||
/// faces" keys on the *embedder* half (`dr_catalog::faces::embedder_of`):
|
||||
/// the embedder is what makes two vectors comparable, and the detector only
|
||||
/// decides where the boxes are. So the three variants here are one
|
||||
/// population, and choosing another one empties nothing — the People screen,
|
||||
/// the clustering and the sync all go on over every face already found.
|
||||
/// What a stronger choice does is queue the images a weaker one indexed for
|
||||
/// re-detection, after the ones nothing has indexed ([`Self::supersedes`]),
|
||||
/// with names carried across by `record_detections`' box and embedding
|
||||
/// matching.
|
||||
///
|
||||
/// The first version of this setting keyed everything on the full id, and
|
||||
/// choosing Thorough emptied both devices' People screens until a whole-
|
||||
/// library re-index had run on each — and, worse, stranded every name
|
||||
/// confirmed on one device, because the other held the same faces under a
|
||||
/// different id and the merge would not match them.
|
||||
///
|
||||
/// The first variant's id is the bare embedder name, because that is the id
|
||||
/// every library indexed before this setting existed was written under;
|
||||
@@ -251,6 +261,64 @@ impl FaceDetector {
|
||||
}
|
||||
}
|
||||
|
||||
/// The id when the detector runs in its int8 form (docs/inference.md §7).
|
||||
///
|
||||
/// A different detector: it finds a different set of faces, so it is a
|
||||
/// different population of detections. The embedder half is unchanged,
|
||||
/// because the embedder never runs in int8, and `embedder_of` keeps the
|
||||
/// two spellings' vectors in one space.
|
||||
pub fn model_id_int8(self) -> &'static str {
|
||||
match self {
|
||||
FaceDetector::Scrfd500m => "scrfd_500m_i8+w600k_mbf",
|
||||
FaceDetector::Scrfd2_5g => "scrfd_2.5g_i8+w600k_mbf",
|
||||
FaceDetector::Scrfd10g => "scrfd_10g_i8+w600k_mbf",
|
||||
}
|
||||
}
|
||||
|
||||
/// Both ids this detector writes under — the f32 form and the int8 one —
|
||||
/// for a question that is about the detector and not about which form
|
||||
/// of it a device happened to run: "has the chosen detector been over
|
||||
/// this image", asked by a re-index that must not ping-pong between a
|
||||
/// desktop that runs it in f32 and a tablet that runs it on the Hexagon.
|
||||
pub fn model_ids(self) -> [&'static str; 2] {
|
||||
[self.model_id(), self.model_id_int8()]
|
||||
}
|
||||
|
||||
/// The detector that writes under a pipeline id, if it is one of these.
|
||||
///
|
||||
/// The inverse of [`Self::model_ids`] -- either spelling. `None` for an
|
||||
/// id from another embedder or a build this one does not know, which a
|
||||
/// caller treats as "cannot rank" rather than as weaker than anything.
|
||||
pub fn for_model_id(model_id: &str) -> Option<FaceDetector> {
|
||||
FaceDetector::ALL
|
||||
.into_iter()
|
||||
.find(|d| d.model_ids().contains(&model_id))
|
||||
}
|
||||
|
||||
/// Whether this detector finds more than `other` does — the measured
|
||||
/// order of §12.3, which is also the order of [`Self::ALL`].
|
||||
pub fn outranks(self, other: FaceDetector) -> bool {
|
||||
let rank = |d: FaceDetector| FaceDetector::ALL.iter().position(|x| *x == d);
|
||||
rank(self) > rank(other)
|
||||
}
|
||||
|
||||
/// The pipeline ids this detector is worth re-running over.
|
||||
///
|
||||
/// Every variant shares one embedder, so a library indexed under any of
|
||||
/// them is one population (`dr_catalog::faces::embedder_of`) and a
|
||||
/// detector change empties nothing. What a stronger detector *adds* is
|
||||
/// the faces a weaker one missed, so a sweep re-detects the images a
|
||||
/// weaker one indexed — after the ones nothing has indexed — and never
|
||||
/// the other way round: a tablet set to Fast keeps the desktop's
|
||||
/// Thorough faces rather than replacing them with fewer.
|
||||
pub fn supersedes(self) -> &'static [&'static str] {
|
||||
match self {
|
||||
FaceDetector::Scrfd500m => &[],
|
||||
FaceDetector::Scrfd2_5g => &["w600k_mbf"],
|
||||
FaceDetector::Scrfd10g => &["w600k_mbf", "scrfd_2.5g+w600k_mbf"],
|
||||
}
|
||||
}
|
||||
|
||||
/// What the picker calls it.
|
||||
pub fn label(self) -> &'static str {
|
||||
match self {
|
||||
@@ -508,6 +576,28 @@ pub struct CacheSettings {
|
||||
/// no bandwidth and saves the whole transfer next time. Off is for metered
|
||||
/// or small-disk devices, where the user would rather re-fetch than store.
|
||||
pub keep_opened_originals: bool,
|
||||
|
||||
/// How many photographs either side of the open one are fetched into the
|
||||
/// cache ahead of being asked for, so a step along the roll is a disk
|
||||
/// read rather than a download. Zero fetches nothing ahead.
|
||||
///
|
||||
/// Each side, so the total is twice this; and closest first, working
|
||||
/// outwards, so a small number still covers the step most likely to be
|
||||
/// taken next. Bounded by [`Self::AHEAD_CHOICES`] because every unit is a
|
||||
/// whole RAW file: 10 each side is a few hundred megabytes per open,
|
||||
/// which a wired desktop shrugs at and a phone on a hotel connection does
|
||||
/// not. Moot while `keep_opened_originals` is off — a fetch the cache
|
||||
/// would discard on arrival is not made.
|
||||
pub fetch_ahead: u32,
|
||||
}
|
||||
|
||||
impl CacheSettings {
|
||||
/// The look-ahead depths the settings page offers, each side.
|
||||
///
|
||||
/// Snapped to, not clamped, for the reason `LibrarySettings::BAR_CHOICES`
|
||||
/// is: the page lights the chip that matches, and a number it does not
|
||||
/// offer would leave every chip dark.
|
||||
pub const AHEAD_CHOICES: [u32; 5] = [0, 2, 5, 10, 20];
|
||||
}
|
||||
|
||||
/// 8 GB of passively cached originals — roughly 250 full-frame RAWs.
|
||||
@@ -532,6 +622,10 @@ impl Default for CacheSettings {
|
||||
original_budget_bytes: Some(DEFAULT_ORIGINAL_BUDGET_BYTES),
|
||||
thumbnail_budget_bytes: Some(DEFAULT_THUMBNAIL_BUDGET_BYTES),
|
||||
keep_opened_originals: true,
|
||||
// Enough that the next few steps in either direction are already
|
||||
// here by the time they are taken, without a click on one frame
|
||||
// committing a phone to half a gigabyte.
|
||||
fetch_ahead: 5,
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -1175,6 +1269,15 @@ impl Settings {
|
||||
.unwrap_or(LibrarySettings::default().timeline_bars);
|
||||
}
|
||||
|
||||
// The same snap, for the same page.
|
||||
if !CacheSettings::AHEAD_CHOICES.contains(&self.cache.fetch_ahead) {
|
||||
let wanted = self.cache.fetch_ahead;
|
||||
self.cache.fetch_ahead = CacheSettings::AHEAD_CHOICES
|
||||
.into_iter()
|
||||
.min_by_key(|n| n.abs_diff(wanted))
|
||||
.unwrap_or(CacheSettings::default().fetch_ahead);
|
||||
}
|
||||
|
||||
// A zero-pixel or zero-percent export produces no image. Nudged to the
|
||||
// smallest thing that does, rather than back to the default: the user
|
||||
// clearly wanted "small", and silently restoring 2048 would ignore
|
||||
@@ -1410,6 +1513,10 @@ mod tests {
|
||||
Some(DEFAULT_THUMBNAIL_BUDGET_BYTES)
|
||||
);
|
||||
assert!(cache.keep_opened_originals);
|
||||
assert!(
|
||||
CacheSettings::AHEAD_CHOICES.contains(&cache.fetch_ahead),
|
||||
"the default look-ahead must be one the page can light"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
@@ -1640,9 +1747,41 @@ mod tests {
|
||||
assert_eq!(s.faces, FaceSettings::default());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_detectors_two_spellings_share_its_embedder_and_nothing_else() {
|
||||
for d in FaceDetector::ALL {
|
||||
let [f32_id, int8_id] = d.model_ids();
|
||||
assert_eq!(f32_id, d.model_id());
|
||||
assert_eq!(int8_id, d.model_id_int8());
|
||||
assert_ne!(f32_id, int8_id);
|
||||
assert_eq!(f32_id.rsplit('+').next(), int8_id.rsplit('+').next());
|
||||
}
|
||||
}
|
||||
|
||||
/// The detector every existing library was indexed with must keep the id
|
||||
/// those libraries were written under, or an upgrade would report every
|
||||
/// one of them un-indexed.
|
||||
#[test]
|
||||
fn detectors_rank_in_the_measured_order_and_only_supersede_downwards() {
|
||||
use FaceDetector::*;
|
||||
assert!(Scrfd10g.outranks(Scrfd2_5g));
|
||||
assert!(Scrfd2_5g.outranks(Scrfd500m));
|
||||
assert!(!Scrfd500m.outranks(Scrfd10g));
|
||||
assert!(!Scrfd10g.outranks(Scrfd10g));
|
||||
for d in FaceDetector::ALL {
|
||||
for weaker in d.supersedes() {
|
||||
let w = FaceDetector::for_model_id(weaker).expect(weaker);
|
||||
assert!(
|
||||
d.outranks(w),
|
||||
"{d:?} lists {weaker} but does not outrank it"
|
||||
);
|
||||
}
|
||||
assert_eq!(FaceDetector::for_model_id(d.model_id()), Some(d));
|
||||
assert_eq!(FaceDetector::for_model_id(d.model_id_int8()), Some(d));
|
||||
}
|
||||
assert_eq!(FaceDetector::for_model_id("scrfd_10g+other"), None);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn the_default_detector_keeps_the_legacy_model_id() {
|
||||
assert_eq!(FaceDetector::default(), FaceDetector::Scrfd500m);
|
||||
@@ -1740,6 +1879,24 @@ mod tests {
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn sanitise_snaps_a_hand_edited_look_ahead_to_an_offered_one() {
|
||||
let mut s = Settings::default();
|
||||
s.cache.fetch_ahead = 7;
|
||||
s.sanitise();
|
||||
assert_eq!(s.cache.fetch_ahead, 5);
|
||||
|
||||
s.cache.fetch_ahead = 100;
|
||||
s.sanitise();
|
||||
assert_eq!(s.cache.fetch_ahead, 20);
|
||||
|
||||
for n in CacheSettings::AHEAD_CHOICES {
|
||||
s.cache.fetch_ahead = n;
|
||||
s.sanitise();
|
||||
assert_eq!(s.cache.fetch_ahead, n, "an offered depth is left alone");
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn sanitise_rescues_a_zero_dimension() {
|
||||
let mut s = Settings::default();
|
||||
|
||||
@@ -17,6 +17,8 @@
|
||||
# JNILIBS where cargo-ndk wrote the .so (default: $TARGET_DIR/jniLibs)
|
||||
# OUT output directory (default: $TARGET_DIR/apk)
|
||||
# KEYSTORE signing keystore (default: $TARGET_DIR/debug.keystore)
|
||||
# RUNTIME_DIR the inference runtime (default: $TARGET_DIR/runtime,
|
||||
# fetched by tools/fetch-android-runtime.sh)
|
||||
# ABI Android ABI (default: arm64-v8a)
|
||||
# RUST_TARGET Rust target triple (default: aarch64-linux-android)
|
||||
#
|
||||
@@ -43,6 +45,7 @@ TARGET_DIR="$(cd "${TARGET_DIR}" && pwd)"
|
||||
JNILIBS="${JNILIBS:-${TARGET_DIR}/jniLibs}"
|
||||
OUT="${OUT:-${TARGET_DIR}/apk}"
|
||||
KEYSTORE="${KEYSTORE:-${TARGET_DIR}/debug.keystore}"
|
||||
RUNTIME_DIR="${RUNTIME_DIR:-${TARGET_DIR}/runtime}"
|
||||
|
||||
# Release signing is selected by supplying a password, not by a flag, so there
|
||||
# is no way to ask for a release build and silently get a debug one.
|
||||
@@ -255,6 +258,25 @@ fi
|
||||
cp "${SO}" "${OUT}/staging/lib/${ABI}/libdarkroom.so"
|
||||
cp "${DEX}" "${OUT}/staging/classes.dex"
|
||||
|
||||
# The inference runtime (docs/inference.md §3): ONNX Runtime and Qualcomm's
|
||||
# Hexagon backend, beside libdarkroom.so so the app finds them in its own
|
||||
# native library directory. The build links none of it — the app dlopens
|
||||
# `libonnxruntime.so` at launch and runs on tract if it is not there — so an
|
||||
# APK without these is a slower app, not a broken one, and `RUNTIME_DIR=none`
|
||||
# builds exactly that. 174 MB for the default set; the script says which
|
||||
# Hexagon generations that buys.
|
||||
if [[ "${RUNTIME_DIR}" != "none" ]]; then
|
||||
if [[ ! -f "${RUNTIME_DIR}/lib/libonnxruntime.so" ]]; then
|
||||
"${REPO}/tools/fetch-android-runtime.sh" "${RUNTIME_DIR}"
|
||||
fi
|
||||
cp "${RUNTIME_DIR}"/lib/*.so "${OUT}/staging/lib/${ABI}/"
|
||||
mkdir -p "${OUT}/staging/assets/licences"
|
||||
cp "${RUNTIME_DIR}"/QNN-*.* "${OUT}/staging/assets/licences/" 2>/dev/null || true
|
||||
echo " runtime: $(ls "${RUNTIME_DIR}/lib" | wc -l) libraries from ${RUNTIME_DIR}/lib"
|
||||
else
|
||||
echo " runtime: none (tract only)"
|
||||
fi
|
||||
|
||||
# The models. Android has no other route to one — app-private storage is not
|
||||
# user-reachable and the in-app fetch is unbuilt (docs/faces.md §2.2a) — so
|
||||
# they go in the APK and `android_main` unpacks them on first launch. The
|
||||
@@ -274,10 +296,10 @@ cp "${DEX}" "${OUT}/staging/classes.dex"
|
||||
#
|
||||
# Cleared first: a previous run that died between staging and cleanup would
|
||||
# otherwise leave models in the APK that are no longer in the tree.
|
||||
rm -rf "${OUT}/staging/assets"
|
||||
rm -rf "${OUT}/staging/assets/models"
|
||||
mkdir -p "${OUT}/staging/assets/models"
|
||||
_bundled=""
|
||||
for _dir in face scene; do
|
||||
for _dir in face scene inpaint; do
|
||||
ASSETS="${REPO}/models/${_dir}"
|
||||
compgen -G "${ASSETS}/*.onnx" >/dev/null || continue
|
||||
# An LFS pointer is ~130 bytes and looks exactly like a model to `cp`. Left
|
||||
@@ -315,7 +337,7 @@ fi
|
||||
# install times sane.
|
||||
cd "${OUT}/staging"
|
||||
cp "${OUT}/base.apk" "${OUT}/unaligned.apk"
|
||||
zip -q -0 -X "${OUT}/unaligned.apk" "lib/${ABI}/libdarkroom.so"
|
||||
zip -q -0 -X "${OUT}/unaligned.apk" lib/"${ABI}"/*.so
|
||||
zip -q -X "${OUT}/unaligned.apk" classes.dex
|
||||
# Stored, not deflated: an ONNX graph is mostly incompressible float data, so
|
||||
# deflating it buys a few percent and costs the whole file being inflated into
|
||||
|
||||
@@ -41,7 +41,7 @@ Measured on the first build, 2026-09-12:
|
||||
| DLL imports | 26 Windows system DLLs; **no** `libwinpthread-1.dll`, `libgcc_s`, `libstdc++` |
|
||||
| `wine darkroom-desktop.exe --version` | `darkroom-desktop 0.12.0`, exit 0, 0.1 s |
|
||||
| `makensis` | 105 MB `DarkRoom-<version>-x86_64-setup.exe`, 64-bit stub |
|
||||
| `wine setup.exe /S` | Installs exe + 7 models to `AppData\Local\Programs\DarkRoom`, writes the `HKCU` uninstall key; the installed exe runs |
|
||||
| `wine setup.exe /S` | Installs exe + 10 models to `AppData\Local\Programs\DarkRoom`, writes the `HKCU` uninstall key; the installed exe runs |
|
||||
| `wine uninstall.exe /S` | Removes the directory and the key |
|
||||
| Start Menu shortcut | **Not verifiable here.** `CreateShortcut` is `IShellLink` and does nothing under a headless Wine; the directory beside it is created. Check on Windows. |
|
||||
|
||||
|
||||
@@ -1099,6 +1099,13 @@ architecture; everything else assumes they pass.
|
||||
|
||||
Decisions treated as fixed because reversing them is expensive. Each is evidence-backed.
|
||||
|
||||
> **On the numbering.** The subsections below are numbered **6.1–6.13**, not 12.x. They were §6
|
||||
> when this document was first written, and every `ARCH §6.n` citation in `requirements.md`, in
|
||||
> `docs/*.md` and in source comments names them by those numbers. Renumbering would break several
|
||||
> hundred citations for no gain, so the numbers stay: `ARCH §6.1` means the first constraint here,
|
||||
> and never §6.1 of the data architecture above, which is cited as *§6 Data architecture* or by
|
||||
> its title.
|
||||
|
||||
### 6.1 GPU results never round-trip through the CPU
|
||||
|
||||
darktable identifies this as their single biggest bottleneck: OpenCL output returns to the GTK
|
||||
@@ -1229,7 +1236,7 @@ Full rationale in [requirements.md §8](requirements.md). Summary:
|
||||
| D10 | Single adaptive interface | Decided |
|
||||
| D11 | Product positioning | Decided |
|
||||
| D12 | Scope versus pace | **Open** |
|
||||
| D13 | Face inference runtime and model licensing | **Runtime answered**, licensing open |
|
||||
| D13 | Face inference runtime and model licensing | **Runtime answered**, reopened for per-device backends (docs/inference.md); licensing open |
|
||||
| D14 | Segmentation source for local masking | Decided — arm C (docs/segmentation.md §14) |
|
||||
| D15 | Target devices — 12-inch tablet and desktop, no phone | Decided (requirements D15) |
|
||||
|
||||
|
||||
@@ -748,12 +748,34 @@ CREATE TABLE faces (
|
||||
-- (faces.md §6, §9). NULL for a face stored as a unit vector before it
|
||||
-- was kept.
|
||||
quality REAL,
|
||||
-- What the eyes are doing (FR-CULL-8a, faces.md §17): per eye P(open),
|
||||
-- the source pixels across its box and the sharpness of the patch the
|
||||
-- classifier saw; and P(sunglasses). All seven or none; NULL is "never
|
||||
-- read", which every filter treats as unknown rather than as closed.
|
||||
-- The verdict -- open, closed, sunglasses, unclear -- is a rule in
|
||||
-- dr_face::eyes, not a column.
|
||||
eye_right REAL,
|
||||
eye_right_px REAL,
|
||||
eye_right_sharp REAL,
|
||||
eye_left REAL,
|
||||
eye_left_px REAL,
|
||||
eye_left_sharp REAL,
|
||||
sunglasses REAL,
|
||||
-- The 106 dense landmarks the eyes were read from, packed as 16-bit
|
||||
-- fixed point over the frame: 424 bytes (schema V18). Kept so the next
|
||||
-- per-face pass runs from the catalog rather than from the original.
|
||||
landmarks_dense BLOB,
|
||||
-- Which model produced this. An embedding is only comparable to others
|
||||
-- from the same model; mixing them silently yields nonsense similarities.
|
||||
model_id TEXT NOT NULL,
|
||||
detected_at INTEGER NOT NULL
|
||||
);
|
||||
CREATE INDEX faces_image ON faces(image_id);
|
||||
-- Covers the eyes-open filter's subquery. Without it every check read the
|
||||
-- whole face row -- the eye columns sit after the blobs -- and one count
|
||||
-- took 24 s on the reference library (schema V17).
|
||||
CREATE INDEX faces_eyes ON faces(image_id, eye_right, eye_right_px, eye_right_sharp,
|
||||
eye_left, eye_left_px, eye_left_sharp, sunglasses);
|
||||
|
||||
CREATE TABLE face_person (
|
||||
face_id INTEGER PRIMARY KEY REFERENCES faces(id) ON DELETE CASCADE,
|
||||
@@ -895,4 +917,5 @@ and is not answered here.
|
||||
| FR-CULL-10 | §10.1 `people` and `face_person`, merge-by-redirect, §10.2 clustering as a library pass |
|
||||
| FR-CULL-11 | §5 `Selector::Person`, confirmed-only by default |
|
||||
| FR-CULL-12 | §10.1 derived data in the catalog; names to the sidecar, UUID as merge identity |
|
||||
| FR-CULL-8a | §10.1 the seven eye columns, derived like the embedding; the filter term is `RatingFilter::eyes_open` |
|
||||
| NFR-SEC-5 | §10.4 sync left undecided and off; nothing in §10 emits an embedding |
|
||||
|
||||
+336
-16
@@ -540,8 +540,8 @@ images indexed, **64 contain no face at all**.
|
||||
|
||||
So `face_index` records the *run*: one row per `(image, model)` carrying the timestamp, the number of
|
||||
faces found — zero is the interesting value — and the long edge of the proxy it read. Keyed on the
|
||||
model, so a model change puts every image back in the queue without anyone having to remember to
|
||||
clear anything.
|
||||
model, so an *embedder* change puts every image back in the queue without anyone having to remember
|
||||
to clear anything; a detector change in front of the same embedder does not (§12.3).
|
||||
|
||||
Three things fall out of it that were not otherwise available:
|
||||
|
||||
@@ -816,10 +816,10 @@ here, and it is the phone and tablet story that should decide whether it gets bu
|
||||
anything downstream sees them. A probe's confidence (§9.1) is computed from the references it
|
||||
matched; a reference's confidence hears nothing from a probe. A face whose quality was never
|
||||
recorded is admitted to the gallery — a rule that cannot be checked admits rather than excludes —
|
||||
and the next indexing pass **measures** it: `faces_unmeasured` lists every image holding one, and
|
||||
each such face is embedded again from the native render with the landmarks it already has, the raw
|
||||
vector written over the old one and its id, box and identity untouched
|
||||
(`faces::record_measurements`). No detector runs and no suggestion is lost — the cost is the
|
||||
and the next indexing pass **measures** it: the `face-quality` repair (§18.1) lists every image
|
||||
holding one, and each such face is embedded again from the native render with the landmarks it
|
||||
already has, the raw vector written over the old one and its id, box and identity untouched
|
||||
(`faces::record_updates`). No detector runs and no suggestion is lost — the cost is the
|
||||
original fetched once more, since the length exists only at the moment of embedding.
|
||||
|
||||
**The algorithm.** Constrained average-link agglomeration over the probability graph, merging while
|
||||
@@ -1100,14 +1100,33 @@ is not a desktop-only feature (NFR-RES-2). It loads in tract with the same fix a
|
||||
decodes through the same nine-output path unchanged. Same licence, same `buffalo_m` release page.
|
||||
|
||||
**Which detector runs is a setting** — `FaceSettings::detector`, per device, on the settings page
|
||||
beside the indexing button as Fast / Balanced / Thorough. All three files ship. Each detector is
|
||||
its own `faces.model_id` (`w600k_mbf` for `500M`, unchanged, so nothing already indexed is
|
||||
disturbed; `scrfd_2.5g+w600k_mbf` and `scrfd_10g+w600k_mbf` for the others), which is the
|
||||
mechanism §2.1 always intended for a model change: the coverage figure restarts at zero under the
|
||||
new id, the sweep re-detects, `record_detections` carries confirmed names across by box overlap and
|
||||
drops the previous pipeline's marker for each image it revisits, and the sync shards are keyed by
|
||||
the same id so a peer on another setting neither adopts nor pollutes them. The default stays `500M`
|
||||
so that an upgrade changes nothing until the user chooses; the recommendation is `2.5G`.
|
||||
beside the indexing button as Fast / Balanced / Thorough. All three files ship. Each detector
|
||||
writes its own `faces.model_id` (`w600k_mbf` for `500M`, unchanged; `scrfd_2.5g+w600k_mbf` and
|
||||
`scrfd_10g+w600k_mbf` for the others), so which pipeline drew a box is always on record. The
|
||||
default stays `500M` so that an upgrade changes nothing until the user chooses; the recommendation
|
||||
is `2.5G`.
|
||||
|
||||
**One population per embedder, not one per detector · 2026-09-19.** The first cut of the setting
|
||||
keyed every reader on the full id — the clustering pass, the coverage figure, the sweep's work
|
||||
list, the shard export and import, and the sync merge's face matching — on the theory that a
|
||||
detector change is a model change. Measured on the reference library it was a disaster: choosing
|
||||
Thorough on both devices restarted coverage at 1,834 of 19,140, the People screen showed only the
|
||||
faces the new pipeline had reached, the desktop's 3,583 confirmations under the old id could not
|
||||
reach the tablet because the merge demanded the same id on both sides, and each device faced a
|
||||
~400 GB re-fetch before the library looked whole again. The embedder is `w600k_mbf` in every
|
||||
variant; its vectors are one space, and the detector only decides where the boxes are.
|
||||
|
||||
So every reader now keys on the embedder half of the id (`faces::embedder_of`, and
|
||||
`embedder_sql` for the queries): all three detectors are one population, and changing between
|
||||
them empties nothing. `record_detections` is unchanged — an image holds one pipeline's faces at a
|
||||
time, and a re-detection carries identities across by box overlap and embedding (§18) — and it is
|
||||
where the generations meet. The merge's `match_faces` matches within an embedder for the same reason. The
|
||||
shards travel every generation, each under its own id, and a peer adopts whichever it is sent.
|
||||
What a stronger choice still does is queue the images a weaker detector indexed for re-detection
|
||||
(`FaceDetector::supersedes`), after the ones nothing has indexed and never downwards, so a tablet
|
||||
on Fast keeps the desktop's Thorough faces rather than replacing them with fewer. The calibration
|
||||
(§8) is keyed on the embedder too: it is a fit over the similarity space, and that space did not
|
||||
change.
|
||||
|
||||
---
|
||||
|
||||
@@ -1238,12 +1257,313 @@ frames to family snapshots as the real unknowns.
|
||||
|
||||
| ID | How this document addresses it |
|
||||
|---|---|
|
||||
| FR-CULL-8 | §4 detection, §6 embedding, §7 the proxy tier and its consequences, §10 the `DetectFaces` job |
|
||||
| FR-CULL-8 | §4 detection, §6 embedding, §7 the proxy tier and its consequences, §10 the `DetectFaces` job, §18 the re-index |
|
||||
| FR-CULL-9 | §8 — pairs, fit, validity, and the reliability-diagram acceptance test |
|
||||
| FR-CULL-10 | §9 constrained agglomeration, confirmations as anchors, split by re-agglomeration |
|
||||
| FR-CULL-10 | §9 constrained agglomeration, confirmations as anchors, split by re-agglomeration; §18 what a re-index carries across |
|
||||
| FR-CULL-11 | §10 the `Person` selector term, confirmed-only by default |
|
||||
| FR-CULL-12 | §10 schema unchanged from catalog.md §10.1: embeddings derived, names to the sidecar |
|
||||
| NFR-SEC-5 | §11 — the obligations restated as structural properties, one of them CI-checkable |
|
||||
| NFR-COMPAT-2 | §2 — why the obvious weights cannot ship, and what does instead |
|
||||
| NFR-RES-2 | §1 the cheap model pair, §12 M2/M3/M10 on a phone |
|
||||
| NFR-ARCH-2 | §10 background priority, preempted by visible work |
|
||||
| FR-CULL-8a | §17 — eye state and sunglasses: the models, the crops, the readability rule, the measurements |
|
||||
| FR-CULL-13 | §17.3 — the reading is evidence: a chip that narrows and a badge that explains, and nothing that writes a judgement |
|
||||
|
||||
---
|
||||
|
||||
## 17. Eyes and sunglasses · 2026-09-19
|
||||
|
||||
FR-CULL-8a's eye state, under FR-CULL-13's rule. Three more models run over every face the
|
||||
pipeline already aligns, and what they produce is a **filter term** — "eyes open", on the people
|
||||
filter — and a badge on the People screen. Nothing acts on it: FR-CULL-13 says the signal is
|
||||
evidence, shown and filterable, and never holds the pen, and §3.9.1's exclusion of blink
|
||||
*detection* was re-read the same day as the exclusion of blink *selection* it always was. The chip
|
||||
narrows the grid the way a star count does, and rates nothing. Head pose, the other half of
|
||||
FR-CULL-8a, is not built; §17.3 says what stands in for it.
|
||||
|
||||
### 17.1 The models, and why these
|
||||
|
||||
| | 2d106det | OCEC | SGC |
|
||||
|---|---|---|---|
|
||||
| Answers | 106 landmarks, ten round each eye's lids | P(this eye is open) | P(this head wears sunglasses) |
|
||||
| Input | 192² RGB 0..255, the detector box at 1.5× | one eye, 40×24 RGB, `x/255` | one head, 48×48 RGB, `x/255` |
|
||||
| Shipped | `buffalo_l`'s, 4.8 MB | S, 483 KB, F1 0.9943 on its own split | L, 6.1 MB, F1 0.9554 on its own split |
|
||||
| Licence | InsightFace's research-only grant, like the pair (§2.2a) | MIT, code and weights | MIT, code and weights |
|
||||
| Training data | InsightFace's | *Open and Closed Eyes* (ODC-By 1.0) + Wholebody34 crops (Apache 2.0) | **not stated** — recorded in `models/face/README.md` |
|
||||
| Cost in tract | ~24 ms per face | ~6 ms per eye | ~6 ms per framing, two framings |
|
||||
|
||||
The two classifiers are from Katsuya Hyodo's ultra-lightweight series — the same author as the
|
||||
whole-body detector §1.1's reference pipeline uses — and are the first weights in `models/face/`
|
||||
that do not come out when the project publishes. The landmark model is under the grant the pair
|
||||
already carries; it was chosen over two permissively licensed alternatives on a measurement
|
||||
(§17.2) after the decision that this project will not be commercial, which is what §2.2a already
|
||||
records for the pair.
|
||||
|
||||
All three load in tract as shipped, dynamic batch and all — the first graphs in this subsystem to
|
||||
do so — and are pinned to a batch of 1 by `tools/fix-face-model-shapes.sh` anyway, because a graph
|
||||
the engine *analyses* and a graph it has been *measured running* are different claims, and the
|
||||
embedder's precedent is the safer one. Six milliseconds per classifier call against 0.4 in the
|
||||
reference README is tract's per-call overhead on a graph this small; the whole reading is under
|
||||
60 ms per face beside an embedding at 160 ms and a native decode in seconds.
|
||||
|
||||
### 17.2 Where the eye box comes from
|
||||
|
||||
**The eye classifier was trained on a whole-body detector's eye boxes, and this pipeline has no
|
||||
eye boxes.** It has five landmarks, and SCRFD's eye point is loose: it is one of five points that
|
||||
place a face, not an eye centre, and on a turned or smiling head the eye sat in a corner of a
|
||||
window centred on it. Everything below was measured on 60 proxies from the reference library with
|
||||
25 plainly open-eyed faces labelled by hand (`examples/eyes.rs --dump`, then a contact sheet), and
|
||||
the count that matters is how many of those 25 the classifier read as open in both eyes.
|
||||
|
||||
**A window on the SCRFD point: 19 of 25.** Windows from 20×10 to 34×17 template units all gave
|
||||
19–20; smaller lost eyes. Two model-free ways of re-centring the window were then tried and both
|
||||
lost eyes: the darkest blob near the landmark is the inner corner's shadow or the lash line
|
||||
(19 → 15), and the most contrasty window is the one that takes in the edge of the nose (19 → 9).
|
||||
The landmark as SCRFD gives it beats either.
|
||||
|
||||
**A box from a landmark model's lid contour: 22 of 25.** Three models were run over the same
|
||||
faces, each fed the crop its reference code feeds it, and the eye box cut as the bounding box of
|
||||
the lid points grown by a margin:
|
||||
|
||||
| model | points | input | tract | per face | open at margin 0.1 |
|
||||
|---|---|---|---|---|---|
|
||||
| MediaPipe Face Mesh V2 (Apache 2.0) | 478, with z | 256² | loads | ~36 ms | 22 |
|
||||
| PIPNet, PINTO's irnet18 export (WFLW, research-only) | 68 | 256² | loads | ~98 ms | 20 |
|
||||
| **InsightFace 2d106det** | **106** | **192²** | **loads** | **~24 ms** | **22** |
|
||||
|
||||
The margin was swept on the two that tied: 22 at 0 and 0.1, 18 at 0.4, 14 at 0.6 — the training
|
||||
crops were tight detector boxes, and a tight box is what the classifier wants (`EYE_BOX_MARGIN`).
|
||||
2d106det ships: it tied the best, costs the least, and is under a grant the project has already
|
||||
accepted. Face Mesh would be the choice if that changed; it also gives z and an iris, neither of
|
||||
which this needs yet.
|
||||
|
||||
The box is cut **upright from the native render**, not through the face's alignment — the
|
||||
training crops were detector boxes, and the contour already says where the eye is on a tilted
|
||||
head (`align::eye_patch`). A shut eye's contour has no height and is given an open eye's
|
||||
(`EYE_BOX_MIN_ASPECT`), so the classifier sees the same framing either way.
|
||||
|
||||
### 17.3 Not asking what cannot be answered
|
||||
|
||||
The 25 open faces were never the real problem. The real problem was the faces that were *not*
|
||||
open-eyed by the classifier's account and were not blinks either, and on the reference sample they
|
||||
were the commonest wrong answer of all: **a soft eye reads as closed.** A face small enough that
|
||||
its eye was seven pixels wide, a motion-blurred face, a face from a 1024 proxy where the native
|
||||
render should have been — each produced a confident "closed" from a classifier shown a smear. The
|
||||
same failure the face's own sharpness gate exists for (§4.3), one stage down, where the face's gate
|
||||
cannot see it: a face sharp enough to embed can hold an eye too soft to read, because the eye is a
|
||||
fortieth of it.
|
||||
|
||||
So the reading is **seven numbers, not a verdict** — per eye P(open), the source pixels across its
|
||||
box and the sharpness of the patch the classifier saw; and P(sunglasses) — stored as such
|
||||
(`faces.eye_right`, `faces.eye_right_px`, `faces.eye_right_sharp`, likewise `eye_left`, and
|
||||
`faces.sunglasses`, schema V16), and the verdict is a rule with thresholds in it,
|
||||
`dr_face::eyes::EyeReading::state`, the only place the thresholds live:
|
||||
|
||||
```
|
||||
sunglasses ≥ 0.5 → Sunglasses (whatever the eyes said)
|
||||
an eye is readable when px ≥ 12
|
||||
and sharpness ≥ 0.02
|
||||
and px ≥ 0.6 × the other eye's px
|
||||
no readable eye → Unreadable
|
||||
a readable eye < 0.5 → Closed (a blink, or a wink)
|
||||
otherwise → Open
|
||||
```
|
||||
|
||||
**Sunglasses take precedence** because the eye classifier answers confidently over dark glass:
|
||||
over a woman in sunglasses on the reference library it read her right eye 0.97 open. **The
|
||||
pixel floor** is where the classifier's own training stopped — its reference footage averaged
|
||||
15–21 pixels an eye. **The sharpness floor** is the face's measure over the patch, set where the
|
||||
sample's open eyes were being called closed: the open set ran from 0.019 (a lens reflection) to
|
||||
5.4, the unreadable ones under 0.02 with the pixels to match. **The width ratio** is the profile:
|
||||
a landmark model's contour for the far eye of a turned head collapses towards the nose. On the
|
||||
twenty native renders of §17.4, profiles put the far eye at 0.02–0.43 of the near one's width,
|
||||
two three-quarter faces whose far eye was reading closed sat at 0.54, and every face looking at
|
||||
the camera sat at 0.78 or more — a shut eye's box keeps its width, so a wink is not mistaken for
|
||||
a turn. 0.6 splits the gap. An eye that fails any of the three is not asked, the near eye still
|
||||
decides, and a face with no readable eye is *unclear* — which is not a blink, and not open, and
|
||||
which no filter drops.
|
||||
|
||||
**The two eyes are kept apart** rather than averaged, because a wink averages to 0.5 — the one
|
||||
value that says the least — and "eyes open" means every eye that could be read.
|
||||
|
||||
With the rule in place, the same 60 proxies read: 33 open, 14 closed, 24 sunglasses, 14 unclear.
|
||||
Of the 25 labelled open faces, 22 open, 2 unclear (eye boxes of 7 and 12 pixels on a child's
|
||||
face), 1 closed — a squinting smile whose contour collapsed to eleven pixels, which the classifier
|
||||
is not wrong to call narrow. The 14 closed are downcast eyes, laughs, two sunglasses the head
|
||||
classifier missed, and the squint. A face 141 pixels across the eye but motion-blurred to a
|
||||
sharpness of 0.016 reads *unclear* where it read *closed* before, which is the change this
|
||||
section is for.
|
||||
|
||||
**The filter drops only *Closed*.** `RatingFilter::eyes_open` compiles the rule above into a
|
||||
predicate on the face row, ANDed into the chosen people's face subquery, so "Anna, eyes open" asks
|
||||
about Anna's face and not about Bob blinking beside her. The chip is offered only while someone is
|
||||
chosen and goes when the last person does — without a name in front of it, it would be a verdict
|
||||
on everyone in the frame. The predicate still handles the empty case, as `NOT EXISTS` over every
|
||||
face, for a filter arriving by another route; a landscape passes because there is no one in it to
|
||||
have blinked. Sunglasses pass. Unclear passes. Never read passes — that last is what keeps an old
|
||||
library from emptying its grid the moment the chip is pressed: until the measuring pass has run,
|
||||
the honest answer is "everything". A test drives the same five readings through the SQL and
|
||||
through `state()` and requires the two to agree, so the badge and the grid cannot say different
|
||||
things.
|
||||
|
||||
**And an index, learned the slow way.** The people filter was always served from a covering index
|
||||
on `faces(image_id)`; the moment its subquery read the eye columns it had to read the face *row*,
|
||||
and `ALTER TABLE ADD COLUMN` had put those seven floats after the embedding and the crop blob —
|
||||
six kilobytes to reach every one. One count took 24 seconds on the reference library, thirteen of
|
||||
them system time. `faces_eyes` (schema V17) covers the subquery again: five milliseconds.
|
||||
|
||||
### 17.4 What the sample says about accuracy, and what it does not
|
||||
|
||||
The 60-proxy sample above was run at proxy resolution, where the production pass reads the native
|
||||
render; the eye box on a 200-pixel face is 40 source pixels from the proxy and 240 from the
|
||||
original. So the shipped configuration was also run over **twenty native renders** from the
|
||||
reference library — a wedding burst of six frames with six or seven faces each, and a dozen
|
||||
singles — exported by `face_native --export` and read by `examples/eyes.rs --dump`, 62 faces in
|
||||
all: 20 open, 28 closed, 12 sunglasses, 2 unclear before the width ratio was moved (below).
|
||||
|
||||
Read off the contact sheet, face by face: the one real blink in the set (`7884.dng`, a man with
|
||||
his eyes shut) is *closed*; the laughing faces with their eyes screwed shut are *closed*, which a
|
||||
photographer would call right; the downcast faces are *closed*, which is arguable; the profiles
|
||||
are judged on the near eye and mostly *open*, which the SCRFD-point pass could not do. Two
|
||||
faces were wrong: three-quarter views whose far eye's box came to 0.54 of the near one's and read
|
||||
closed over a cheek, which is what moved the width ratio from 0.45 to 0.6. One is beyond any
|
||||
floor: a face with a porcelain pot held over the eyes, whose contour is a guess and whose "eyes"
|
||||
are sharp white china.
|
||||
|
||||
Two things no sample so far can say. None holds more than one real blink, so the precision of
|
||||
*Closed* is not measured — every closed verdict on the two sheets but the occlusion and the
|
||||
sunglasses misses is a narrowed or shut eye rather than a wrong one, but that is a reading of a
|
||||
contact sheet, not a number. And the floors were set on a few dozen faces. Both are M11.
|
||||
|
||||
### 17.4a What is kept per face
|
||||
|
||||
The box, the five landmarks, `crop_px`, the embedding and the crop were already there. The eye
|
||||
pass adds the seven numbers of §17.3 and the **106 dense landmarks** it read the eye boxes from,
|
||||
packed as 16-bit fixed point over the frame — 424 bytes a face, a seventh of a pixel on a
|
||||
6000-pixel frame (`dr_face::Landmarks::to_packed_bytes`, schema V18). Kept for the reason the
|
||||
embedding is kept: it cost a fetch of the original and a model run, and the next per-face pass —
|
||||
head pose, expression, whatever FR-CULL-8a grows — should run from the catalog. `f16` would have
|
||||
been the same size and worse: three figures near 1.0 is six pixels at that scale.
|
||||
|
||||
### 17.5 The measuring pass, and shards
|
||||
|
||||
A face indexed before the models existed, or on a device without them, has no reading. The
|
||||
`face-eyes` repair (§18.1; on the day this was written, the sweep's measuring pass — the one V14
|
||||
built to re-embed faces stored as unit vectors) lists those faces, on a device that has the models,
|
||||
and reads their eyes from the same native render with the box and landmarks already stored. No
|
||||
detector runs and no identity moves. A device *without* the models has no such repair, or it would
|
||||
fetch every original in the library to do nothing to it; the repair's predicate is the one both the
|
||||
count and the work list use, so the pass converges. The People screen's coverage line counts
|
||||
these faces as work to read and keeps the button while any remain — reading **Read eye state**
|
||||
once detection is complete and only readings are left, which is the state an already-indexed
|
||||
library is in the day the models arrive.
|
||||
|
||||
Shards carry the seven columns beside `quality`. A peer's faces without a reading are **adopted**,
|
||||
unlike a peer's faces without a quality (§14): the measuring pass finds this work by the NULL and
|
||||
not by the run marker, so adoption costs the reading nothing, and a peer with no eye models may be
|
||||
the only device that has done the detection at all.
|
||||
|
||||
### 17.6 Still to measure
|
||||
|
||||
| # | Measure | Why it decides something |
|
||||
|---|---|---|
|
||||
| **M11** | Open-eye recall and blink precision from the **native** pass with the shipped configuration, on a labelled sample that contains real blinks — a burst with one in it is enough | §17.3's figures are proxy figures from 25 faces, with no precision beside them; this is the number FR-CULL-13's acceptance clause asks for, and where the two readability floors get set on more than 25 faces |
|
||||
| **M12** | Whether the remaining open-eye failure — a lens reflection over an open eye — moves with the sharpness floor, or needs the eye classifier told about spectacles | If the latter, the fix is a classifier trained closer to this domain, and that is a different decision |
|
||||
| **M13** | Sunglasses recall on more than twelve faces, and the false-positive rate on caps and clear glasses | The two sunglasses the head classifier missed on the sample became false blinks; a library of skiers would say whether that is two faces or a class |
|
||||
|
||||
---
|
||||
|
||||
## 18. The completeness job, and what a re-index carries across · 2026-09-19
|
||||
|
||||
The reference library on the day this was written: 17,762 faces under the bare `w600k_mbf` id —
|
||||
found by the fast detector on 1024 px proxies, stored as unit vectors, no quality, no crop on 4,144
|
||||
of them, no eye reading, no dense landmarks — beside 1,177 under `scrfd_10g+w600k_mbf` from the
|
||||
native pass, and 12,217 images the fast detector examined and found nothing in. Over those faces:
|
||||
3,778 confirmations, 13,011 suggestions, 77 rejections, and 17,276 people rows. Every one of those
|
||||
gaps was, until now, its own pass: V14's measuring pass for the quality, §17.5's for the eyes, the
|
||||
sweep's proxy repair, the sweep's detector upgrade, and a re-index that did not exist. Adding a
|
||||
per-face field meant adding a pass, with its own work list, its own count and its own idea of done.
|
||||
|
||||
### 18.1 One job over a registry
|
||||
|
||||
`dr_ui::repairs` replaces them with one job over a **registry**. A `Repair` names one thing a catalog
|
||||
record can lack — the predicate that says which images still owe it, the input its handler needs
|
||||
(the file's header, the whole original, or a native render), the handler that fills it, and, where
|
||||
there is one, what to record for an image that can never be done. The job unions the predicates
|
||||
into one work list, fetches each image once at the most any claimant asks for, renders it at most
|
||||
once, and runs every handler whose predicate that image still matches — checked again before each,
|
||||
because one handler's write satisfies the next's (a detection writes every field a per-face handler
|
||||
would fill). The registry today:
|
||||
|
||||
| Repair | Owed by | Input | Handler |
|
||||
|---|---|---|---|
|
||||
| `face-proxy` (sweep only) | images holding faces whose 1024 px proxy is not in the store | native render | detect again, write the proxy |
|
||||
| `face-quality` | faces with `quality IS NULL` | native render | warp from the stored landmarks, embed, write the raw vector and its length (reads eyes on the same warp where it can) |
|
||||
| `face-eyes` | faces with no eye reading or no dense landmarks, on a device with the eye models | native render | read the eyes and dense landmarks from the stored box and landmarks |
|
||||
| `face-crop` | faces with `crop IS NULL` | native render | cut the crop from the frame |
|
||||
| `face-detection` | sweep: images with no marker under the embedder and no faces; re-index: images with no marker under the **chosen detector**, either spelling | native render | detect, embed, replace the faces, carry identities across (§18.2) |
|
||||
| `face-upgrade` (sweep only) | images whose marker is a weaker detector's | native render | as `face-detection` |
|
||||
| `metadata` | `images.metadata_state < 2` | header | EXIF to the catalog, the dateless marked examined |
|
||||
|
||||
A repair's predicate is the *only* definition of its work. The count the settings page shows
|
||||
(`faces::audit`, per repair), the list the job fetches and the check before each handler are one
|
||||
predicate, so a record the count reports is one the job fetches and one the handler fills, and the
|
||||
job ends. This is why an eye reading that cannot be cut is not a criterion — such a face stays
|
||||
unread however often it is detected, and listing it would fetch its original on every press — and
|
||||
why a degenerate face is dropped rather than left. It is also why the registry is cut to what the
|
||||
device can do (`Capabilities`): a device without the eye models has no `face-eyes` entry, rather
|
||||
than an entry it skips, because an entry is a count and a set of originals to fetch.
|
||||
|
||||
The registry is ordered, and the order is the work list's: an image only the last repair claims
|
||||
comes after one the first does, which is what puts a few hundred proxy repairs ahead of twenty
|
||||
thousand un-indexed images. Within one image the same order runs the handlers, detection before the
|
||||
per-face repairs, since detection fills what they would.
|
||||
|
||||
Adding a field is one entry: a predicate over `faces f` or `images i`, and a handler that fills it
|
||||
from `Fetched`. `metadata` is in the table to say that this is not a face job — the same machinery
|
||||
carries a capture date, and could carry a thumbnail, a perceptual hash or a head pose.
|
||||
|
||||
### 18.1a The two scopes
|
||||
|
||||
Both buttons on the settings page run the job; they differ in one predicate. **Index faces**
|
||||
(`Scope::Outstanding`) converges on *coverage* — has anything examined this image — and treats a
|
||||
face a weaker detector found on a proxy as found, which is the right question for a pass that must
|
||||
not fetch the library twice. **Re-index every face** (`Scope::Reindex`) converges on *provenance*:
|
||||
`face-detection` claims every image with no marker under the chosen detector, in either of its
|
||||
forms (`FaceDetector::model_ids`, so a desktop running it in f32 and a tablet on the Hexagon in int8
|
||||
do not re-index each other's work), and a marker saying a weaker one looked is not that. This is
|
||||
the one place in the subsystem keyed on the exact detector rather than the embedder. Convergent all
|
||||
the same: an image the job has been through leaves the list, a kill costs the images in flight, and
|
||||
a second press resumes.
|
||||
|
||||
An original over the fetch budget is skipped without being fetched. Under the sweep, detection
|
||||
records an examination that found nothing — the honest record for an image that cannot be
|
||||
examined, and what stops the half gigabyte being spent once per sweep. Under the re-index, and
|
||||
under every repair over records that already exist, it is left exactly as it was: a re-detection
|
||||
with nothing found would delete the faces, and "cannot fetch" is not "no faces".
|
||||
|
||||
### 18.2 What is carried across
|
||||
|
||||
`dr_catalog::faces::record_detections` replaces an image's faces and carries identities onto the
|
||||
new ones. Before this section it carried confirmations only, by box overlap above 0.5 IoU, and a
|
||||
re-detection of the library above would have left 13,011 suggestions and 77 rejections on the
|
||||
floor — correct by FR-CULL-12's letter, since suggestions are derived data, and a People screen
|
||||
emptied to strangers by the user's own button.
|
||||
|
||||
Now every old face is read before the delete — box, vector, assignment, rejections — and matched to
|
||||
the new faces one-to-one, best pair first. A pair qualifies when the boxes **overlap at all** and
|
||||
either the overlap alone says so (IoU above 0.5, the old rule) or the embeddings do (cosine above
|
||||
`SAME_FACE_COSINE` = 0.45, the reference library's P≈0.95 line from §9's table). The embedding route
|
||||
is for the box a low-resolution pass drew badly enough that overlap alone would not claim it; the
|
||||
vector is also what breaks the tie in a group photograph, where two neighbouring faces overlap both
|
||||
new boxes. Overlap is required on both routes because the same vector elsewhere in the frame — a
|
||||
mirror, a print on the wall — is not the same face and must not take its name. Onto the matched
|
||||
face go the assignment as it was, confirmed or suggested with its probability, and every
|
||||
rejection.
|
||||
|
||||
It is a match, not an update in place, and that is why the per-face repairs exist beside
|
||||
detection: where nothing about a face but one field needs doing, `record_updates` keeps the id and
|
||||
there is nothing to judge.
|
||||
|
||||
The merge's `match_faces` still matches by overlap alone across devices. It is the same question,
|
||||
and the same answer would serve it; it is not changed here.
|
||||
|
||||
+34
-25
@@ -5,7 +5,7 @@
|
||||
|
||||
Every entry here is extracted from the comment beside the code that implements it, so this file cannot describe a gesture the application does not have. Add one by writing a `GESTURE:` block next to the implementation; there is nowhere else to write it.
|
||||
|
||||
40 gestures, in 4 places.
|
||||
41 gestures, in 4 places.
|
||||
|
||||
## Develop
|
||||
|
||||
@@ -25,7 +25,7 @@ Sampling a neutral is the first move of the tonal pass — every colour judgemen
|
||||
|
||||
Anchored on the fingers' midpoint, and on the pointer, so the gesture reads as magnifying the picture rather than sliding it about. Double-tap is the way to an exact 1:1; this is the way to everything in between.
|
||||
|
||||
<sub>`ui/dr-ui/ui/app.slint:1974`</sub>
|
||||
<sub>`ui/dr-ui/ui/app.slint:2068`</sub>
|
||||
|
||||
### Move a magnified photograph about
|
||||
|
||||
@@ -34,7 +34,7 @@ Anchored on the fingers' midpoint, and on the pointer, so the gesture reads as m
|
||||
|
||||
Only once there is something outside the viewport to reach, which is why the cursor becomes a hand exactly then. The view is clamped to the frame: panning past the edge would show undefined area beside the photograph, and that reads as a rendering fault rather than as the end of the picture.
|
||||
|
||||
<sub>`ui/dr-ui/ui/app.slint:2065`</sub>
|
||||
<sub>`ui/dr-ui/ui/app.slint:2159`</sub>
|
||||
|
||||
### Paint a mask by hand
|
||||
|
||||
@@ -43,7 +43,7 @@ Only once there is something outside the viewport to reach, which is why the cur
|
||||
|
||||
A model's mask stops inside a shoulder and leaks into the hair, and no single edge control fixes two errors that go opposite ways. The whole stroke is one step in the history, so taking a mark back costs one press however long it took to make.
|
||||
|
||||
<sub>`ui/dr-ui/ui/app.slint:2152`</sub>
|
||||
<sub>`ui/dr-ui/ui/app.slint:2246`</sub>
|
||||
|
||||
### Take back the last change
|
||||
|
||||
@@ -53,7 +53,7 @@ A model's mask stops inside a shoulder and leaks into the hair, and no single ed
|
||||
|
||||
A whole drag is one step, so undo takes back a decision rather than a frame of a gesture. The list is there because arriving six steps back costs what arriving from one does.
|
||||
|
||||
<sub>`ui/dr-ui/ui/app.slint:2373`</sub>
|
||||
<sub>`ui/dr-ui/ui/app.slint:2467`</sub>
|
||||
|
||||
### Do it again after taking it back
|
||||
|
||||
@@ -61,7 +61,7 @@ A whole drag is one step, so undo takes back a decision rather than a frame of a
|
||||
- **Pointer** — Click it, or press Redo in the History header
|
||||
- **Keyboard** — Ctrl+Shift+Z
|
||||
|
||||
<sub>`ui/dr-ui/ui/app.slint:2386`</sub>
|
||||
<sub>`ui/dr-ui/ui/app.slint:2480`</sub>
|
||||
|
||||
### Copy the settings from this photograph
|
||||
|
||||
@@ -71,7 +71,7 @@ A whole drag is one step, so undo takes back a decision rather than a frame of a
|
||||
|
||||
The panel is the copy that has to work: a tablet has no modifier key to hold and no menu bar to hang the action from. The shortcut is an accelerator for a control that is on screen either way.
|
||||
|
||||
<sub>`ui/dr-ui/ui/app.slint:2419`</sub>
|
||||
<sub>`ui/dr-ui/ui/app.slint:2513`</sub>
|
||||
|
||||
### Paste the settings onto this photograph
|
||||
|
||||
@@ -81,7 +81,7 @@ The panel is the copy that has to work: a tablet has no modifier key to hold and
|
||||
|
||||
The button names what would be pasted — "3 adjustments", and whether the crop is coming with it — which the shortcut cannot say. Both paste the same scope.
|
||||
|
||||
<sub>`ui/dr-ui/ui/app.slint:2431`</sub>
|
||||
<sub>`ui/dr-ui/ui/app.slint:2525`</sub>
|
||||
|
||||
### Change which group of adjustments is on screen
|
||||
|
||||
@@ -91,7 +91,7 @@ The button names what would be pasted — "3 adjustments", and whether the crop
|
||||
|
||||
The groups are whatever the operation set declares itself to be about, so there are as many as the pipeline has and no key can be assigned to one of them by name. Stepping is the binding that survives a node being added.
|
||||
|
||||
<sub>`ui/dr-ui/ui/app.slint:2459`</sub>
|
||||
<sub>`ui/dr-ui/ui/app.slint:2553`</sub>
|
||||
|
||||
### Look at the photograph at 1:1
|
||||
|
||||
@@ -101,7 +101,7 @@ The groups are whatever the operation set declares itself to be about, so there
|
||||
|
||||
Noise reduction and capture sharpening are judgements about single pixels, and a fitted view averages several of the file's into each one on screen — so the frame looks softer than it is and the correction goes too far. The point and the magnification survive opening the next photograph, which is what makes checking the same eye across forty portraits forty keystrokes rather than forty pans.
|
||||
|
||||
<sub>`ui/dr-ui/ui/app.slint:2494`</sub>
|
||||
<sub>`ui/dr-ui/ui/app.slint:2588`</sub>
|
||||
|
||||
### Move to the next or previous photograph
|
||||
|
||||
@@ -111,7 +111,7 @@ Noise reduction and capture sharpening are judgements about single pixels, and a
|
||||
|
||||
The edit on screen is saved on the way out, so stepping through a folder is as much a departure as going back to the grid and loses nothing.
|
||||
|
||||
<sub>`ui/dr-ui/ui/app.slint:2546`</sub>
|
||||
<sub>`ui/dr-ui/ui/app.slint:2640`</sub>
|
||||
|
||||
### See the photograph before you edited it
|
||||
|
||||
@@ -121,7 +121,7 @@ The edit on screen is saved on the way out, so stepping through a folder is as m
|
||||
|
||||
Held rather than toggled, and no split screen: a split halves the working image on the tablet the column was sized for, and the comparison photographers describe making is a flick back and forth. It takes no history step, so checking whether a frame is overcooked costs nothing to undo afterwards.
|
||||
|
||||
<sub>`ui/dr-ui/ui/app.slint:2670`</sub>
|
||||
<sub>`ui/dr-ui/ui/app.slint:2764`</sub>
|
||||
|
||||
### Put one control back to its default
|
||||
|
||||
@@ -226,7 +226,7 @@ Double-click is what a file manager and a Lightroom panel use for the same thing
|
||||
|
||||
Grouping over-merges on siblings, on parents and children, and on the same person a decade apart, so splitting is as prominent as merging. A tool that can only merge makes its own errors permanent.
|
||||
|
||||
<sub>`ui/dr-ui/ui/identity.slint:161`</sub>
|
||||
<sub>`ui/dr-ui/ui/identity.slint:188`</sub>
|
||||
|
||||
### Rule on a suggested face
|
||||
|
||||
@@ -235,7 +235,7 @@ Grouping over-merges on siblings, on parents and children, and on the same perso
|
||||
|
||||
A face is either the system's guess or the user's judgement, and the two are never conflated. A rejection is remembered, so the face is not suggested for that person again. The gesture note above is the whole label: a tick and a cross are only "confirm" and "reject" to someone who can see the suggestion they sit beside, and `IconButton`'s fallback would announce them as "check" and "cross" — two icon names that say nothing about which person is being ruled on.
|
||||
|
||||
<sub>`ui/dr-ui/ui/identity.slint:181`</sub>
|
||||
<sub>`ui/dr-ui/ui/identity.slint:208`</sub>
|
||||
|
||||
### See a person's photographs
|
||||
|
||||
@@ -244,7 +244,7 @@ A face is either the system's guess or the user's judgement, and the two are nev
|
||||
|
||||
This is the point of having identified anybody. Without it the screen is a filing cabinet with no drawer handles.
|
||||
|
||||
<sub>`ui/dr-ui/ui/identity.slint:602`</sub>
|
||||
<sub>`ui/dr-ui/ui/identity.slint:635`</sub>
|
||||
|
||||
### Change how faces are grouped
|
||||
|
||||
@@ -253,7 +253,7 @@ This is the point of having identified anybody. Without it the screen is a filin
|
||||
|
||||
The right match confidence is a property of your library, not of the model. "What would this do?" answers for this library without writing anything; names, confirmations and the groups you have set aside are kept whatever the dials say.
|
||||
|
||||
<sub>`ui/dr-ui/ui/identity.slint:639`</sub>
|
||||
<sub>`ui/dr-ui/ui/identity.slint:672`</sub>
|
||||
|
||||
## Library grid
|
||||
|
||||
@@ -301,6 +301,15 @@ This replaced a double tap, which had no visible state and could take forty phot
|
||||
|
||||
<sub>`ui/dr-ui/ui/library.slint:1414`</sub>
|
||||
|
||||
### Take the blinks out of a burst
|
||||
|
||||
- **Touch** — Narrow to a person, then tap "Eyes open" beside their name on the filter bar
|
||||
- **Pointer** — Narrow to a person, then click "Eyes open" beside their name on the filter bar
|
||||
|
||||
Face indexing reads each face's eyes. The chip drops frames where the chosen people are caught blinking, and leaves sunglasses and eyes it could not read alone.
|
||||
|
||||
<sub>`ui/dr-ui/ui/library.slint:2106`</sub>
|
||||
|
||||
### Find photographs with two people in them
|
||||
|
||||
- **Touch** — Open the People chip on the filter bar, tap each name, then switch the chip beside them to "all of them"
|
||||
@@ -308,7 +317,7 @@ This replaced a double tap, which had no visible state and could take forty phot
|
||||
|
||||
"Any of them" is a union and "all of them" is an intersection. The tray is where both terms and the choice between them live, because a filter belongs on the filter bar.
|
||||
|
||||
<sub>`ui/dr-ui/ui/library.slint:2102`</sub>
|
||||
<sub>`ui/dr-ui/ui/library.slint:2135`</sub>
|
||||
|
||||
### Resize the thumbnails
|
||||
|
||||
@@ -317,7 +326,7 @@ This replaced a double tap, which had no visible state and could take forty phot
|
||||
|
||||
There is no wheel on a tablet, so without the pinch the cell size could only be changed by a control a finger cannot reach.
|
||||
|
||||
<sub>`ui/dr-ui/ui/library.slint:2755`</sub>
|
||||
<sub>`ui/dr-ui/ui/library.slint:2789`</sub>
|
||||
|
||||
### File photographs in a collection
|
||||
|
||||
@@ -326,7 +335,7 @@ There is no wheel on a tablet, so without the pinch the cell size could only be
|
||||
|
||||
The selection is what the drag carries, which is why selecting several is worth the mode: forty photographs file in one gesture.
|
||||
|
||||
<sub>`ui/dr-ui/ui/library.slint:2952`</sub>
|
||||
<sub>`ui/dr-ui/ui/library.slint:2986`</sub>
|
||||
|
||||
### Open a photograph
|
||||
|
||||
@@ -335,7 +344,7 @@ The selection is what the drag carries, which is why selecting several is worth
|
||||
|
||||
A tap opens; a tap that *moved* does not. Travel is what separates a deliberate tap from a hand brushing past, and it is the only thing that does: the two are the same length. An earlier version required the finger to dwell 120 ms instead, and that rejected ordinary taps — a real tap is often quicker than a brush.
|
||||
|
||||
<sub>`ui/dr-ui/ui/library.slint:3220`</sub>
|
||||
<sub>`ui/dr-ui/ui/library.slint:3254`</sub>
|
||||
|
||||
### Rate a photograph without opening it
|
||||
|
||||
@@ -345,7 +354,7 @@ A tap opens; a tap that *moved* does not. Travel is what separates a deliberate
|
||||
|
||||
A star has to take the press without it also reaching the cell, or every rating throws the user into develop.
|
||||
|
||||
<sub>`ui/dr-ui/ui/library.slint:3340`</sub>
|
||||
<sub>`ui/dr-ui/ui/library.slint:3374`</sub>
|
||||
|
||||
### Choose the frame a folded burst shows
|
||||
|
||||
@@ -354,7 +363,7 @@ A star has to take the press without it also reaching the cell, or every rating
|
||||
|
||||
A folded burst draws its earliest frame, which is a fact about the clock and not a judgement about the photograph — nothing in this application ranks a frame (FR-CULL-5). But the point of a burst is that one of the twelve is better than the other eleven, and the photographer is the only one who knows which. So the choice is offered on the frames themselves, while they are open and side by side, which is the one moment the alternatives are on screen to be compared.
|
||||
|
||||
<sub>`ui/dr-ui/ui/library.slint:3471`</sub>
|
||||
<sub>`ui/dr-ui/ui/library.slint:3505`</sub>
|
||||
|
||||
### Drop the selection but keep selecting
|
||||
|
||||
@@ -363,7 +372,7 @@ A folded burst draws its earliest frame, which is a fact about the clock and not
|
||||
|
||||
Distinct from Done, which leaves the mode entirely. Clearing keeps it, so the next selection can start straight away.
|
||||
|
||||
<sub>`ui/dr-ui/ui/library.slint:4132`</sub>
|
||||
<sub>`ui/dr-ui/ui/library.slint:4166`</sub>
|
||||
|
||||
### Select everything the grid is showing
|
||||
|
||||
@@ -372,7 +381,7 @@ Distinct from Done, which leaves the mode entirely. Clearing keeps it, so the ne
|
||||
|
||||
A scoped grid of two hundred frames is two hundred taps otherwise, and "all of them, except those three" is a far more common shape than the taps it took to say it.
|
||||
|
||||
<sub>`ui/dr-ui/ui/library.slint:4149`</sub>
|
||||
<sub>`ui/dr-ui/ui/library.slint:4183`</sub>
|
||||
|
||||
### Take photographs out of a collection
|
||||
|
||||
@@ -381,4 +390,4 @@ A scoped grid of two hundred frames is two hundred taps otherwise, and "all of t
|
||||
|
||||
The badge on a cell says a photograph is filed in three collections and never which. This is the sheet that names them, and the only way out of one the grid is not currently scoped to.
|
||||
|
||||
<sub>`ui/dr-ui/ui/library.slint:4264`</sub>
|
||||
<sub>`ui/dr-ui/ui/library.slint:4307`</sub>
|
||||
|
||||
@@ -0,0 +1,484 @@
|
||||
# Inference backends — the runtime and the model, chosen per device
|
||||
|
||||
Spec for **S16**, the build that puts the neural models on the hardware each device actually has.
|
||||
|
||||
Every model DarkRoom runs today — the three SCRFD detectors, the ArcFace embedder, YOLO26n-seg and
|
||||
the ADE20K scene model — runs through `tract`, on one CPU core, on every platform. That was the
|
||||
right first answer: D13's runtime half chose it because it costs no C dependency, and
|
||||
[faces.md](faces.md) and [segmentation.md](segmentation.md) were written against it. It is also
|
||||
between 20× and 300× slower than what the same devices can do, and this document is the record of
|
||||
having measured that and the specification of what replaces it.
|
||||
|
||||
**It does not reopen D13's licensing half.** The weights are the same files under the same grant.
|
||||
It does reopen the *runtime* half, and §3 is where it says how far.
|
||||
|
||||
---
|
||||
|
||||
## 1. What was measured · 2026-09-19
|
||||
|
||||
One benchmark, two builds of it: `ort`'s API over `tract` (exactly what the app links) and `ort`'s
|
||||
API over a dynamically loaded ONNX Runtime with each execution provider in turn. Random 640×640
|
||||
input, three warm-ups, the median of 15–30 timed runs, milliseconds. The same input every run, so
|
||||
the numbers are compute cost and nothing else.
|
||||
|
||||
### 1.1 The tablet — Honor MagicPad 2, Snapdragon 8s Gen 3
|
||||
|
||||
SM8635: 1× Cortex-X4, 4× A720, 3× A520, Adreno 735, Hexagon V73. Android 16.
|
||||
|
||||
| Model | **tract** (today) | ORT CPU f32 | ORT CPU int8 | Adreno f32 ¹ | **Hexagon int8** ² |
|
||||
|---|---|---|---|---|---|
|
||||
| scrfd_500m (Fast) | 98 | 16 | 8 | 21 | **1.4** |
|
||||
| scrfd_2.5g (Balanced) | 161 | 59 | 19 | ✗ | **1.8** |
|
||||
| scrfd_10g (Thorough) | 489 | 204 | 48 | ✗ | **3.2** |
|
||||
| arcface_mbf (per face) | 39 | 9 | 13 | 24 | 12 |
|
||||
| yolo26n-seg | 287 | 94 | 38 | 54 | **5.1** |
|
||||
| yolo26s-sem-ade20k | 408 | 154 | 46 | 47 | **3.7** |
|
||||
|
||||
¹ Qualcomm's own GPU backend (`libQnnGpu.so`, OpenCL). Fails on the two larger SCRFD graphs at an
|
||||
`AveragePool` the layout transformer cannot place. ONNX Runtime's WebGPU provider also runs on this
|
||||
GPU and was slower than the CPU on every model; it is not in the table because it is not a
|
||||
candidate.
|
||||
² QNN's HTP backend. The Hexagon **refuses float32 and float16 tensors** in this ORT 1.29 + QNN
|
||||
2.42 pairing (error 3110 on every node, with `enable_htp_fp16_precision` set or not); int8 QDQ
|
||||
graphs run with 99.6% of nodes on the NPU — 1718 of 1725 for SCRFD-500m, the remainder being the
|
||||
quantise/dequantise at the graph's edges — verified from the partition log, not inferred from the
|
||||
timing.
|
||||
|
||||
Also tried and rejected: **NNAPI** — the device registers no neural-networks HAL at all, so the
|
||||
provider has nothing to talk to; Google deprecated it in Android 15 and Qualcomm stopped shipping
|
||||
drivers for it. **XNNPACK** — slower than ORT's default CPU kernels on every model that loaded, and
|
||||
aborts inside its partitioner on the SCRFD graphs.
|
||||
|
||||
### 1.2 The desktop — RTX 3050 Laptop, Raptor Lake, 20 threads
|
||||
|
||||
| Model | **tract** (today) | ORT CPU f32 | ORT CPU int8 | CUDA f32 | CUDA fp16 | TensorRT f32 | TensorRT fp16 | TensorRT int8 |
|
||||
|---|---|---|---|---|---|---|---|---|
|
||||
| scrfd_500m | 104 | 12 | 8 | 5.6 | 3.7 | 2.5 | **1.8** | ✗ ⁴ |
|
||||
| scrfd_2.5g | 162 | 27 | 12 | 6.2 | 5.3 | 2.9 | **1.9** | ✗ ⁴ |
|
||||
| scrfd_10g | 514 | 99 | 32 | 14.5 | 9.3 | 7.6 | **3.3** | ✗ ⁴ |
|
||||
| arcface_mbf | 45 | 15 | 18 | 1.1 | 0.8 | 0.9 | 0.7 | ✗ ⁴ |
|
||||
| yolo26n-seg | 307 | 60 | 38 | 9.3 | ✗ ³ | 6.8 | **5.5** | ✗ ⁴ |
|
||||
| yolo26s-sem-ade20k | 395 | 61 | 31 | 10.7 | ✗ ³ | 7.9 | **3.4** | ✗ ⁴ |
|
||||
|
||||
³ The offline fp16 conversion (`onnxconverter-common`) left a mixed-type node the CUDA provider
|
||||
rejects. TensorRT converts to fp16 itself at engine build and does not have this problem, which is
|
||||
one reason it is the target and the CUDA provider is the fallback.
|
||||
⁴ TensorRT refuses the QDQ form ONNX Runtime's quantiser writes for the Hexagon (uint8
|
||||
activations); it wants symmetric int8. Not pursued: fp16 needs no quantisation, no calibration and
|
||||
no accuracy gate, and it is already 30–60× tract.
|
||||
|
||||
CUDA int8 is deliberately absent: the CUDA provider has no int8 kernels and runs a QDQ graph by
|
||||
dequantising it, which measured *slower* than f32 (7.0 vs 5.6 ms on scrfd_500m). Int8 on NVIDIA is
|
||||
TensorRT's job.
|
||||
|
||||
**TensorRT's first load is 16–116 s per model in f32 and 35–290 s in fp16** (yolo26n-seg the
|
||||
worst: nearly five minutes), because it is compiling an engine for this exact GPU. The engine caches to disk and the second load is milliseconds. That
|
||||
number is what §6 is designed around.
|
||||
|
||||
### 1.3 What the numbers say
|
||||
|
||||
- **`tract` is single-threaded.** The tablet's one X4 core and one Raptor Lake core give the same
|
||||
tract numbers. Replacing it with ONNX Runtime's CPU provider, *no accelerator involved*, is 3–6×
|
||||
on the tablet and 8–10× on the desktop. That is the floor, and it is available on every platform
|
||||
the app builds for.
|
||||
- **The Hexagon is the standout.** A 6 W NPU running int8 beats a discrete RTX 3050 running f32 on
|
||||
five of six models. A whole-library face index on the tablet goes from ~100 ms + 39 ms per face
|
||||
to ~1.4 ms + 12 ms per face, and the "Thorough" detector — 3× the cost of "Fast" today — becomes
|
||||
free. Its price is that the models must be **quantised to int8**, which is an accuracy question
|
||||
§5 has to answer before it is believed.
|
||||
- **The embedder does not gain from either accelerator.** 112×112 input, per-op overhead
|
||||
dominates; it is 9 ms on the tablet's CPU and 12 ms on its NPU. It stays float, which §7 turns
|
||||
from a performance footnote into a correctness rule.
|
||||
- **On NVIDIA, TensorRT fp16 ≈ 3× the CUDA provider**, and the CUDA provider ≈ 2× the
|
||||
multi-threaded CPU; at fp16 the detectors are 1.8–3.3 ms with no quantisation at all. Both leave the twenty cores free for decoding during a batch index, which the table does not
|
||||
show and which matters more than the ratio.
|
||||
|
||||
---
|
||||
|
||||
## 2. The shape of the answer
|
||||
|
||||
A **ladder per platform**, walked at start-up, with the first rung that builds a real session
|
||||
winning:
|
||||
|
||||
| Platform | 1st | 2nd | 3rd | Floor |
|
||||
|---|---|---|---|---|
|
||||
| Android, Qualcomm with a Hexagon the shipped QNN skel covers (V68–V81) | QNN HTP, int8 model | ORT CPU, f32 model | — | tract |
|
||||
| Android, any other SoC | ORT CPU, f32 | — | — | tract |
|
||||
| Linux / Windows, NVIDIA GPU | TensorRT, f32 model, fp16 engine | CUDA provider, f32 | ORT CPU, f32 | tract |
|
||||
| Linux / Windows, no NVIDIA | ORT CPU, f32 | — | — | tract |
|
||||
| macOS ⁵ | ORT CPU, f32 | — | — | tract |
|
||||
|
||||
⁵ CoreML is the obvious rung and is unmeasured; it is listed so its absence is a gap and not an
|
||||
oversight.
|
||||
|
||||
Deliberately **not** on any ladder, with the measurement that excluded each: NNAPI (no driver),
|
||||
XNNPACK (slower than CPU, aborts on SCRFD), WebGPU (slower than CPU), the Adreno through QNN (works,
|
||||
but never where the Hexagon does not also), CUDA int8 (slower than CUDA f32). A rung is added to this
|
||||
table by a measurement on this page, not by a provider existing.
|
||||
|
||||
Two things the ladder is *not*: it is not a per-model choice — one backend serves every model on a
|
||||
device, because §7's identity rule needs the detector and embedder on the same runtime for the
|
||||
same reason `faces.model_id` pairs them; and it is not a per-account choice — it is a property of
|
||||
the hardware, like `shared_face_models_dir` is, and it lives beside it.
|
||||
|
||||
---
|
||||
|
||||
## 3. The dependency policy, and how far this reopens it
|
||||
|
||||
D13 chose `ort` over `tract` because `alternative-backend` made ONNX Runtime's *API* available with
|
||||
none of its *C*. Every rung above the floor needs the C++ ONNX Runtime and, for the two that matter
|
||||
most, vendor libraries on top: Qualcomm's QNN runtime (~60 MB for one Hexagon generation; it is
|
||||
per-SoC) and NVIDIA's TensorRT plus cuDNN (~600 MB with the CUDA libraries, and cuDNN's major
|
||||
version must match what ONNX Runtime was built against — the Arch package on the reference desktop
|
||||
was unusable for exactly that reason).
|
||||
|
||||
The policy protected the **build**: no C to cross-compile under the NDK, no toolchain to keep in
|
||||
step. This document keeps that intact, and the mechanism is the one thing about `ort` that makes it
|
||||
possible:
|
||||
|
||||
**`ort::set_api` accepts any `OrtApi` table.** With `alternative-backend` on, `ort` links nothing
|
||||
and asks for the table once per process. The application can `dlopen` a `libonnxruntime.so` it
|
||||
finds on disk, call `OrtGetApiBase()->GetApi(version)` and hand that table over; or, if there is no
|
||||
such file, hand over `ort_tract::api()`. The Rust build is identical in both cases — pure Rust,
|
||||
`cargo build --target aarch64-linux-android` sees the same dependency graph it sees today. What
|
||||
changes is that the runtime is a **file the package installs**, next to the models, and the app
|
||||
looks for it at start-up.
|
||||
|
||||
Consequences that follow and are accepted:
|
||||
|
||||
- **The runtime is chosen once per process**, because `set_api` is once per process. The ladder
|
||||
in §2 is walked at start-up and the result is what every session in that process uses. There is
|
||||
no "tract for this model, ORT for that one", and there is no falling back to tract *after* ONNX
|
||||
Runtime has loaded — but there does not need to be: once the library loads, its CPU provider is
|
||||
always there, and every fallback the ladder needs is between providers *inside* it.
|
||||
- **Feature flags stay as they are.** `dr-face`'s `inference` and `dr-segment`'s `semantic`
|
||||
continue to mean "compiled against `ort`'s API"; nothing at build time knows or cares which table
|
||||
will be supplied. The one addition is a `native-probe` feature on the new crate (§8) that pulls in
|
||||
`libloading`, which is pure Rust and already in the tree via `wgpu`.
|
||||
- **The packagers ship the runtime, not the build.** The Arch package, the Flatpak manifest, the
|
||||
NSIS installer and `assemble-apk.sh` each gain the ONNX Runtime library for their platform, and
|
||||
the Android and NVIDIA variants gain the vendor libraries — each under the licence the packager
|
||||
reads first (§3.1). A package without them is not broken; it is the tract build, and it says so
|
||||
on the about screen.
|
||||
- **The NDK problem does not come back.** `libonnxruntime.so` for Android is a prebuilt from
|
||||
Maven (`com.microsoft.onnxruntime:onnxruntime-android-qnn`), extracted by `assemble-apk.sh` into
|
||||
`jniLibs/` the way the models are bundled as assets today. Nothing compiles it.
|
||||
|
||||
### 3.1 Licences the packagers read before shipping a runtime
|
||||
|
||||
Written down now, because [segmentation.md §7](segmentation.md) established that reading the grant
|
||||
is cheaper than discovering it at packaging time.
|
||||
|
||||
| Component | Licence | Redistributable in a self-distributed package? |
|
||||
|---|---|---|
|
||||
| ONNX Runtime | MIT | Yes |
|
||||
| Qualcomm QNN runtime (`com.qualcomm.qti:qnn-runtime` on Maven) | Qualcomm AI Engine Direct SDK licence — proprietary, redistribution permitted for applications using it | Yes for the APK, with the licence text shipped; not for a source distribution. **To be read in full, not summarised from memory, before the APK gains it.** |
|
||||
| CUDA runtime, cuDNN, TensorRT | NVIDIA EULAs — redistributable with an application, with the licence text, not modifiable | Yes for a package that bundles them. 600 MB. The alternative is to load them from the user's system install if present and skip the rung otherwise — which is what §4's probe does anyway. |
|
||||
|
||||
The position this takes: the **NVIDIA libraries are not bundled**. The desktop package probes for a
|
||||
system CUDA/TensorRT install and uses it if it is version-compatible; a desktop without one runs on
|
||||
ORT CPU, which is still 8–10× today. Bundling 600 MB for a rung that is 2× again is not a trade
|
||||
worth making unmeasured, and it can be revisited by a measurement on a batch index. The **QNN
|
||||
runtime is bundled** in the APK, because the Hexagon is the difference between a tablet that
|
||||
indexes a library overnight and one that does it over lunch, and the package is 60 MB larger for
|
||||
it.
|
||||
|
||||
Both positions are D13 territory and are recorded there (§12).
|
||||
|
||||
---
|
||||
|
||||
## 4. Selection — the probe, its cache, and what it may not do
|
||||
|
||||
**A rung is chosen by building a real session on it, not by asking whether it exists.** Both
|
||||
failure modes that are not "the provider is absent" were hit on 2026-09-19: a driver in a wedged
|
||||
state where the provider registered and the session then failed, and a provider that registered,
|
||||
took the graph, and rejected every node at partition time. The probe therefore:
|
||||
|
||||
1. Loads the runtime library (§3), or falls to tract and stops.
|
||||
2. Times the **smallest detector** on the CPU provider first — the floor. Then, for each rung
|
||||
in this platform's ladder, in order: builds a session for the same model on that provider
|
||||
with `error_on_failure`, runs it once on a fixed input, and times three more runs. **The rung
|
||||
is taken only if its median beats the floor.** That one measurement is the proof the provider
|
||||
took the graph: one that silently hands the work to the CPU is the CPU rung with extra
|
||||
overhead, slower than the floor, and rejected. (ONNX Runtime's
|
||||
`session.disable_cpu_ep_fallback` was the first draft of this proof and refuses the Hexagon
|
||||
over the ten quantise/dequantise nodes at the graph's edges that QNN declines by policy.)
|
||||
3. Records the outcome — rung, runtime version, provider version, device identity (GPU name and
|
||||
compute capability; SoC model and Hexagon arch), and the models' content hashes — to a small
|
||||
file beside `shared_face_models_dir`. The next start-up trusts the file **unless** any of those
|
||||
inputs changed, in which case it probes again. A driver update, a runtime update, a new model
|
||||
file: each invalidates the cache by construction, and none needs a "reset backend" button.
|
||||
|
||||
What the probe may not do:
|
||||
|
||||
- **Block the first frame.** It runs on the same background as `install_bundled_models` and for
|
||||
the same reason: a TensorRT probe can take thirty seconds cold, and a tablet that stalls that long
|
||||
is an ANR. Until it reports, every model request is answered by the floor the runtime supports
|
||||
(ORT CPU if the library loaded, tract otherwise), and a job that started on the floor finishes
|
||||
on it — a backend does not change under a running index.
|
||||
- **Retry a rung that failed within a session.** A failed probe is cached as a failure with the
|
||||
same inputs; the rung is tried again when an input changes. Otherwise a wedged driver means a
|
||||
thirty-second stall on every launch.
|
||||
- **Choose for the user without saying so.** Settings gains one row, *Inference backend*, showing
|
||||
what was chosen and why in one line ("Hexagon NPU · int8 · QNN 2.42"; "CPU · ONNX Runtime 1.30 ·
|
||||
TensorRT probe failed: cuDNN 8 required"), with an override to force any lower rung. The about
|
||||
screen carries the same line beside the model names NFR-SEC-5 already puts there.
|
||||
|
||||
---
|
||||
|
||||
## 5. Model variants, and who makes them
|
||||
|
||||
Every model exists in one **canonical** form — the f32 ONNX file the app ships or the user supplies
|
||||
today — and, where a rung needs it, a **derived** form. The ladder's rungs are specified in terms
|
||||
of which form they load:
|
||||
|
||||
| Form | Who produces it | When | Needed by |
|
||||
|---|---|---|---|
|
||||
| f32 ONNX, shape-fixed, **opset ≥ 13** | `tools/fix-face-model-shapes.sh`, `tools/export-seg-model.sh` | Release time, once | Every rung except Hexagon |
|
||||
| int8 QDQ ONNX, per-channel, uint8 activations | `tools/quantise-models.sh` (new) | Release time, once, **calibrated on real photographs** | Hexagon |
|
||||
| TensorRT engine (`.engine`, per GPU architecture and TensorRT version) | The app, from the f32 file | First run on that device, in the background | TensorRT rung |
|
||||
| QNN context binary | The app, from the int8 file | First run on that device, in the background | Hexagon rung |
|
||||
|
||||
Two rules.
|
||||
|
||||
**Quantisation is a release-time step, not a device-time one.** The int8 files that produced §1's
|
||||
numbers were calibrated on random noise, which is enough to time and worthless to trust. A real
|
||||
int8 detector is calibrated on a few hundred real photographs and then measured against the f32
|
||||
detector on the reference library by [faces.md §12.3](faces.md)'s method — faces found, per size
|
||||
band, per detector — before it ships. That needs the reference library and a person reading the
|
||||
result, and it happens once per model release, in `tools/`, beside the shape-fixing it already
|
||||
depends on. The device never quantises anything.
|
||||
|
||||
The SCRFD and ArcFace files are **opset 11** as InsightFace exported them, and per-channel QDQ needs
|
||||
13; `tools/fix-face-model-shapes.sh` gains an opset upgrade to 17 (`onnx.version_converter`,
|
||||
`ir_version` 8), which `tract` has been verified to load and which every provider on this page
|
||||
prefers. That is a change to the canonical file and so a change to the shipped models, and it
|
||||
happens in the same model release as the int8 files.
|
||||
|
||||
**Compilation is a device-time step, and it is cached.** A TensorRT engine is specific to the GPU
|
||||
it was built on and the TensorRT that built it; a QNN context binary is specific to the Hexagon
|
||||
generation. Neither can ship. Both are built by the app the first time that rung is selected, in
|
||||
the background (§6), and written beside the probe cache keyed by the same inputs. They are
|
||||
**derived, disposable, and regenerable**: deleting the cache directory costs the next launch a
|
||||
rebuild and nothing else, and the directory is excluded from anything that syncs (it is a peer of
|
||||
`thumbs`, not of the catalog).
|
||||
|
||||
---
|
||||
|
||||
## 6. First run — building engines without the user waiting for them
|
||||
|
||||
The sequence on a device where a compiling rung (TensorRT, Hexagon) is selected:
|
||||
|
||||
1. **Launch.** The runtime loads; the probe (§4) starts in the background; the app serves every
|
||||
model request from the floor. Face indexing, segmentation and scene grading all work, at
|
||||
today's speed or better (ORT CPU).
|
||||
2. **Probe reports** — say, TensorRT. The compiling rung is now *selected* but has **no engines**.
|
||||
Model requests continue on the fallback rung below it (CUDA provider for TensorRT; ORT CPU for
|
||||
Hexagon), which needs no compilation and is already faster than the floor.
|
||||
3. **Engines build**, one model at a time, on a single low-priority background thread, smallest
|
||||
model first so the detector — the one that runs per image — is ready soonest. On the reference
|
||||
desktop that is ~1 minute for the first detector and ~10 minutes for all six at fp16; on the tablet the QNN
|
||||
context binaries take 0.8–1.7 s each and the whole set is ready before the user has opened a
|
||||
library. Each engine is written to a temporary name and renamed into place, so a request never
|
||||
sees a half-written file.
|
||||
4. **Requests move up as engines land.** A model whose engine exists loads it on the selected
|
||||
rung; one whose engine is still building loads on the fallback. **A running job does not
|
||||
switch** — an index that started on the CUDA provider finishes on it — because §7 needs one
|
||||
`model_id` per job, and because a job is the wrong granularity for surprise.
|
||||
5. On Android, the build runs only while the app is in the foreground and the device is not in
|
||||
battery saver (NFR-RES-3): a context binary takes a second, so this costs nothing, and the rule
|
||||
exists for the day a model takes longer.
|
||||
|
||||
Settings shows a one-line progress row while engines build ("Preparing GPU engines · 3 of 6") and
|
||||
nothing when they are done. A build failure demotes the rung — it is recorded in the probe cache
|
||||
as a failure with the model hash as an input, so a corrected model file retries it — and the app
|
||||
carries on one rung down, saying so in the same row.
|
||||
|
||||
---
|
||||
|
||||
## 7. Identity — what changes `model_id` and what may not
|
||||
|
||||
`faces.model_id` exists so that two libraries indexed with different networks are never compared
|
||||
as if they were one ([catalog.md §10.1](catalog.md); the trap is written up in
|
||||
[faces.md §14](faces.md)). A backend that changes what a network *computes* is a different network
|
||||
and must be a different `model_id`; one that changes only *where* it computes it must not be.
|
||||
|
||||
**The detector.** An int8 SCRFD finds a different set of faces from the f32 one — that is what
|
||||
§5's acceptance measures — so **the numeric form is part of the detector's identity**:
|
||||
`scrfd_500m` and `scrfd_500m_i8` are two detectors in `model_id`, and a library indexed on the
|
||||
tablet's Hexagon and continued on the desktop is two populations, which a re-index on either side
|
||||
reconciles the same way a switch from Fast to Thorough does today. That is acceptable because it is
|
||||
already the rule for the detector and because §5 is the gate on whether the int8 form is close
|
||||
enough to be *offered* at all. f32 on tract, ORT CPU, CUDA and TensorRT-f32 are one identity: the
|
||||
same graph, the same arithmetic, differences at the last bit.
|
||||
|
||||
**The embedder** is where comparability across devices is the whole point, and it is the one
|
||||
model that no accelerator helps (§1.3). So: **the embedder runs in f32 on every rung.** On TensorRT
|
||||
that means the embedder's engine is built without fp16 while the detector's is built with it; on
|
||||
the Hexagon it means the embedder is not on the NPU at all — it runs on the ORT CPU rung at 9 ms,
|
||||
and the ladder's "one backend per device" is, precisely, one backend *per model role*, with the
|
||||
embedder pinned. A `w600k_mbf` embedding from any device is comparable with one from any other,
|
||||
which is the property the identity system, the calibration and the cross-device merge all rest
|
||||
on, and it is not for sale for 3 ms.
|
||||
|
||||
If S16 wants fp16 for the embedder later, the gate is written now: over the reference library's
|
||||
faces, the cosine between the f32 and fp16 embedding of the same crop exceeds 0.999 for 99.9% of
|
||||
faces and the calibration's fitted threshold moves by less than its own confidence interval
|
||||
([faces.md §8.3](faces.md)). Until measured, f32.
|
||||
|
||||
**Segmentation and the scene model** carry no identity across devices — their outputs are
|
||||
recomputed per image and never stored beyond the cache — so they take whatever the rung offers,
|
||||
int8 included, subject to §10's own acceptance.
|
||||
|
||||
---
|
||||
|
||||
## 8. Crate shape — `core/dr-inference-engine`
|
||||
|
||||
The seam is the same shape as [storage.md](storage.md)'s: a small crate below the consumers that is
|
||||
the **only** place naming a provider, a library file or a vendor, with the consumers reduced to
|
||||
"give me a session for these bytes in this role".
|
||||
|
||||
```
|
||||
core/dr-inference-engine
|
||||
src/lib.rs Runtime (Tract | Onnx { lib, version }), Backend (rung), Role (Detector | Embedder | Segmenter)
|
||||
src/probe.rs §4 — the ladder per platform, the session-build probe, the cache file
|
||||
src/engines.rs §6 — background compilation, the cache directory, progress
|
||||
src/session.rs open(role, bytes) -> ort::Session, applying the rung and the role's precision rule
|
||||
src/api.rs the one unsafe block: dlopen libonnxruntime, fetch OrtApi, ort::set_api — or ort_tract::api()
|
||||
```
|
||||
|
||||
- `dr-face` and `dr-segment` **delete** their private `install_backend` and their direct
|
||||
`Session::builder()` calls and take an `&dr_inference_engine::Sessions` where they take model bytes today.
|
||||
Their tests keep `tract` — `dr_inference_engine::Sessions::tract()` is a constructor and the test-only path.
|
||||
- `dr-inference-engine` depends on `ort` with the same workspace features as today plus `cuda`, `tensorrt`,
|
||||
`qnn`: those features add option builders, not linking, under `alternative-backend`. **Verified
|
||||
for the QNN, CUDA and TensorRT builders on 2026-09-19** — they go through the API table's generic
|
||||
`SessionOptionsAppendExecutionProvider*`. The NNAPI builder resolves a symbol directly and would
|
||||
not; it is not needed and is not enabled.
|
||||
- `dr-ui` owns the settings row, the about-screen line and the progress row; it holds one
|
||||
`Sessions` per process, created at launch, and passes it down. `dr_ui::library` gains
|
||||
`inference_cache_dir()` beside `shared_face_models_dir()`, on the same account-independent
|
||||
footing and for the same reason.
|
||||
- The Android entry point's `install_bundled_models` also extracts nothing new: `jniLibs/` is
|
||||
loaded by the system loader, and `dr-inference-engine` on Android looks for `libonnxruntime.so` through
|
||||
`dlopen` by bare name first, which resolves to the APK's copy, before any directory.
|
||||
|
||||
---
|
||||
|
||||
## 9. Threads and memory
|
||||
|
||||
- ONNX Runtime's intra-op pool is sized to the physical cores minus two on desktop and to the
|
||||
performance cores on Android (the X4 and the A720s; the A520s are for the compositor). tract's
|
||||
single thread today is the reason a batch index leaves nineteen cores idle; ORT CPU with the
|
||||
pool is the reason it will not. One session per model per process; `Session::run` is
|
||||
`&mut self`-free in `ort` and internally serialised, and the index job is the only caller.
|
||||
- A TensorRT session pins GPU memory for its workspace; the builder is capped at 512 MB on the
|
||||
reference 6 GB card and the cap is a setting, because the develop view's tiles share the card
|
||||
(NFR-RES-2). The engine cache on disk is bounded by the model set — six engines, ~80 MB — and
|
||||
needs no LRU.
|
||||
- The Hexagon rung sets QNN's performance mode to `Burst` for the duration of an index job and
|
||||
`Default` otherwise; a 5 ms detector does not need the NPU clocked up between images.
|
||||
- The probe (§4) and the engine build (§6) run on one dedicated low-priority thread. They never
|
||||
share the index job's pool: a probe that competes with the job it is meant to speed up is the
|
||||
frame-budget trap in a new coat.
|
||||
|
||||
---
|
||||
|
||||
## 10. What S16 measures
|
||||
|
||||
In order, with the gate each is:
|
||||
|
||||
| # | Question | Gate |
|
||||
|---|---|---|
|
||||
| M1 | Does one binary carry both tables? `dlopen` + `set_api` on Linux, Windows and Android; `ort_tract::api()` when the file is absent. | Go / no-go for §3. If `set_api` cannot take a dynamically fetched table on some platform, that platform ships two binaries, and the cost is stated. |
|
||||
| M2 | Do the int8 SCRFD detectors, **calibrated on real photographs**, find the faces? faces.md §12.3's method over the reference library, per size band, against f32. | Ship the int8 form for a detector only if it finds ≥ 97% of the f32 detector's faces above 40 px and the difference is not concentrated in one band. Otherwise that detector's Hexagon rung is ORT CPU int8-free, and the table in §1.1 says what that costs. |
|
||||
| M3 | Does the embedder on ORT CPU beside a detector on the Hexagon (§7) produce embeddings within the f32 gate? | It must — same graph, same arithmetic. This is a check that the plumbing did not quantise it by accident. |
|
||||
| M4 | Is a batch index on the tablet and on the desktop faster by the ratio §1 predicts, end to end, decode included? | The face index over the reference library (18,143 faces): report wall-clock on tract, on the floor and on the selected rung, and where the time went. The prediction is that decode becomes the bottleneck on both; if it does not, say why. |
|
||||
| M5 | Does the first-run sequence (§6) hold: nothing blocks the first frame, engines land, requests move up, a running job does not switch? | Observed on both devices with the app's own progress row, and with the cache directory deleted between runs. |
|
||||
| M6 | What does the APK weigh with the QNN runtime, and does a non-Qualcomm Android device (any one) still launch and index on the floor? | Size reported; launch verified on one non-Qualcomm device or an emulator. |
|
||||
| M7 | The segmentation and scene models at int8 on the Hexagon: does the mask boundary move? segmentation.md §6's IoU against f32 over its corpus. | Ship int8 for a model only above the IoU floor that document set for arm B. |
|
||||
|
||||
M1 and M2 are the ones the rest is conditional on, and M2 is the one that needs a person.
|
||||
|
||||
### 10.1 M2 result · 2026-09-19
|
||||
|
||||
Each int8 detector against its own f32 form, over 400 proxies evenly spaced through the reference
|
||||
library, on ONNX Runtime's CPU provider (the int8 graph is the same file the Hexagon loads;
|
||||
`ui/dr-ui/examples/face_detectors`). Calibrated on 64 proxies from the same library, disjoint
|
||||
from the 400.
|
||||
|
||||
| Detector | f32 faces | int8 faces | both | int8 only | f32 only | found ≥ 32 px | found, all sizes |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| scrfd_500m | 1342 | 1319 | 1287 | 32 | 55 | 95.6% | 95.9% |
|
||||
| scrfd_2.5g | 1525 | 1482 | 1478 | 4 | 47 | 96.1% | 96.9% |
|
||||
| scrfd_10g | 1769 | 1760 | 1749 | 11 | 20 | 100% | 98.9% |
|
||||
|
||||
The 10g form clears the 97% gate; 500m and 2.5g sit one point under it. What they lose is
|
||||
specific: the faces in the "f32 only" column have a **median confidence of 0.52** against a
|
||||
threshold of 0.50 — detections the f32 graph itself barely made, that int8 rounding drops to the
|
||||
other side of the line — and the extra faces int8 finds are the same kind (median 0.51–0.52).
|
||||
Not a size-band failure: the losses are spread across bands in proportion. Shipped as they are,
|
||||
with the number on record; a threshold of 0.48 for the int8 forms would recover most of the
|
||||
margin, and is the first thing to try if a library's count on the tablet reads low.
|
||||
|
||||
Two things the calibration taught, both in `tools/quantise-models.py`: the calibration set has
|
||||
to contain faces (a first attempt on landscape photographs produced a graph that found nothing —
|
||||
the score head's ranges had never seen the face regime), and ONNX Runtime's own strided and
|
||||
moving-average calibration modes both measurably degrade the result on these graphs, while
|
||||
driving the calibrator in chunks by hand reproduces the plain min/max ranges exactly.
|
||||
|
||||
---
|
||||
|
||||
## 11. Order
|
||||
|
||||
1. **`dr-inference-engine` with the two tables and the floor** — `set_api` from a dlopened runtime, tract
|
||||
otherwise, ORT CPU as the only rung. Consumers moved over; tests unchanged. This alone is the
|
||||
3–10× and is the build most of the value sits in. M1.
|
||||
2. **The probe and its cache** (§4), with the settings row and the about line. Still CPU-only;
|
||||
the ladder has one rung. M5's first half.
|
||||
3. **`tools/quantise-models.sh`** and the opset upgrade; the int8 detectors calibrated and
|
||||
measured. M2, M7. This is the step with a person in it and it runs in parallel with 4.
|
||||
4. **The Hexagon rung**, the QNN runtime in the APK, the context-binary cache. M3, M6.
|
||||
5. **The TensorRT and CUDA rungs** on desktop, the engine cache, the first-run sequence. M5's
|
||||
second half.
|
||||
6. **M4** last, on both devices, and the number goes in this document.
|
||||
|
||||
The Windows installer and the Flatpak manifest are touched in steps 1 and 5 only, and only to add
|
||||
a file each; the Arch package likewise.
|
||||
|
||||
---
|
||||
|
||||
## 12. Register entries
|
||||
|
||||
**FR-INF-1 — Runtime selection.** On launch the application shall determine, per device and
|
||||
without blocking the first frame, the fastest inference backend that can build and run a session
|
||||
for the shipped models, by attempting it; shall record and reuse that determination until the
|
||||
runtime, driver, hardware or models change; and shall display the backend in use in Settings and
|
||||
on the about screen. *Acceptance:* §10 M1 and M5.
|
||||
|
||||
**FR-INF-2 — Derived engines.** Backends that require device-specific compilation shall compile in
|
||||
the background after selection, shall serve requests from the next lower backend until each engine
|
||||
is ready, and shall not change the backend of a job in progress. *Acceptance:* M5.
|
||||
|
||||
**FR-INF-3 — Model forms.** Quantised model forms are produced at release time from real
|
||||
calibration data and are shipped only when they meet §10's accuracy gates against the canonical
|
||||
form; the application never quantises on the device. *Acceptance:* M2, M7.
|
||||
|
||||
**NFR-INF-1 — Embedding comparability.** Face embeddings shall be computed at a precision whose
|
||||
deviation from the f32 reference is within §7's gate, on every backend, so that embeddings from any
|
||||
device are comparable. *Acceptance:* M3.
|
||||
|
||||
**D13 — updated.** The runtime half is reopened to the extent of §3: the Rust build stays C-free
|
||||
under `alternative-backend`; packages may install a dynamically loaded ONNX Runtime and, per §3.1,
|
||||
the Qualcomm QNN runtime; the NVIDIA libraries are not bundled. The licensing half is unchanged.
|
||||
|
||||
---
|
||||
|
||||
## 13. Requirements touched
|
||||
|
||||
FR-CULL-8 (indexing time is what this exists to change), FR-CULL-9 (NFR-INF-1 is the guard on its
|
||||
calibration), NFR-RES-2 (TensorRT workspace against the develop view's tiles), NFR-RES-3
|
||||
(background compilation on Android), NFR-SEC-5 (the about screen's line), NFR-COMPAT-2 (each
|
||||
channel gains a runtime file), FR-PLAT-AND-1 (unchanged; the runtime lives in the APK, not in
|
||||
storage), ARCH §6.1 (the once-per-image budget this was sized against no longer binds; what could
|
||||
run per frame is a separate question this document does not open).
|
||||
+39
-11
@@ -30,7 +30,15 @@ one most likely to be reported as closed.
|
||||
|
||||
---
|
||||
|
||||
## 1. Plugins — 21 requirements, and a contradiction to resolve before any of them
|
||||
## 1. Plugins — post-v1 since 2026-09-19
|
||||
|
||||
> **Resolved, in the register.** The contradiction below was settled on 2026-09-19 the way the
|
||||
> last paragraph of this section asked: §3.10 is marked `(post-v1)` clause by clause, §7's row
|
||||
> says so with a reason, D16 defers with it, and the traceability tool lists deferred clauses in
|
||||
> their own table instead of counting them. Coverage went from 72.2% of 194 to 80.6% of 170 on
|
||||
> that edit alone. What follows is kept as the record of what was decided and why; nothing in it
|
||||
> is owed a tag.
|
||||
|
||||
|
||||
**Untagged:** FR-PLG-1, -1a, -2a, -2b, -2c, -3, -3a, -4, -4a, -5, -5a, -5b, -5c, -6, -6a, -7, -8,
|
||||
-9, -10, -11, -12.
|
||||
@@ -153,7 +161,7 @@ What exists is the declaration and not the mechanism: `DetailPass::radius` is do
|
||||
a tile would have to be grown by, with a test that pins it, and there is no scheduler to read it.
|
||||
That is deliberate plumbing, not an oversight.
|
||||
|
||||
**So the open question here is not "when is tiling built" but "is FR-DSP-2 still a requirement".**
|
||||
**So the open question here is not "when is tiling built" but "is FR-DSP-2 still a requirement" — and on 2026-09-19 the answer was: as written, until S6 runs.** FR-DSP-2 now carries a status note saying exactly that, and R5's note no longer claims it was rewritten.
|
||||
Two measurements say it costs more than it saves on the interactive path. Neither says anything
|
||||
about the export path or about a device under memory pressure, which is where the case for it
|
||||
actually lives — and that is spike S6, which has not run.
|
||||
@@ -163,9 +171,11 @@ quarter-resolution base are adjacent and are not it: both are fixed choices abou
|
||||
compute at, where FR-DSP-4 asks for a first frame that is deliberately cheap and a second that
|
||||
replaces it. Nothing tracks a "this frame is provisional" state.
|
||||
|
||||
**NFR-RES-2 — Images larger than GPU memory.** No answer, and §4.3 knows it: the requirement text
|
||||
itself asks the reader to "decide explicitly" how ARCH §6.4 and NFR-RES-2 are reconciled. There is
|
||||
no headroom budget, no allocation-failure fallback, and no spill. Spike S6 — a tiled pipeline on a
|
||||
**NFR-RES-2 — Images larger than GPU memory.** Half answered. NFR-R8's "decide explicitly" was
|
||||
decided on 2026-09-19: there is no CPU render pipeline, the degraded mode is the viewer on
|
||||
embedded previews with develop withheld, and NFR-RES-2 no longer promises a fallback render. What
|
||||
remains unbuilt is the memory half: there is no headroom budget, no allocation-failure staging,
|
||||
and no spill. Spike S6 — a tiled pipeline on a
|
||||
mid-range Android device with an image larger than available GPU memory — is the one that would
|
||||
settle both this and FR-DSP-2, and there is no evidence it has run.
|
||||
|
||||
@@ -445,9 +455,27 @@ whether something *should* be built — which is the opposite of the order §9 a
|
||||
|
||||
---
|
||||
|
||||
## 11. D12, which governs all of the above
|
||||
## 11. Merging — specified 2026-09-19, nothing built
|
||||
|
||||
[Decision D12 — scope versus pace](requirements.md) is still **OPEN**, and says:
|
||||
§3.11 was written on 2026-09-19 under D18, undeferring the panorama from §7 and leaving HDR merge
|
||||
and focus stacking there with their data model decided. Eleven `FR-MRG` clauses and two `NFR-MRG`
|
||||
figures entered the register at once with no code behind any of them, which is why the coverage
|
||||
figure fell from 83.0% to 77.2% on the same day — a specification, not a regression.
|
||||
|
||||
[panorama.md](panorama.md) is the design, and its §10 is the order of work. Nothing starts before
|
||||
**S15**: whether rawler reads back a linear DNG the application writes, whether XFeat loads under
|
||||
tract at a fixed shape, whether the working-space texture can be tapped where FR-MRG-2 needs it,
|
||||
and what a chunked blend of a 100 MP composite costs on the tablet. The first two are a day each
|
||||
and either can change the design, which is the reason they come first.
|
||||
|
||||
## 12. D12, which governed all of the above
|
||||
|
||||
> **Decided 2026-09-19, by events.** The scope stands as calibrated, v1 has no date, and `(post-v1)`
|
||||
> in §7 is the one way a clause leaves the count — used for the plugin API and nothing else. The
|
||||
> argument below is kept because it is what the decision weighed; its prediction about tablet
|
||||
> editing was right, and the cluster was built anyway.
|
||||
|
||||
[Decision D12 — scope versus pace](requirements.md) said, while it was open:
|
||||
|
||||
> The calibration selected an ambitious feature set — full tablet editing, full ingest, culling as a
|
||||
> differentiator, complete GPU masking, AI denoise, Fuji-first colour, deep sync, sidecar durability
|
||||
@@ -461,7 +489,7 @@ execution limits and two GPU vendors to validate (§5), and every one of those i
|
||||
The parts that *were* built — the develop pipeline, sync, faces, the catalog — are the parts that
|
||||
did not need a decision first.
|
||||
|
||||
D12 is not resolved by choosing to work faster. It is resolved by moving requirements across the
|
||||
line into §7, which costs nothing but the admission, and which this document is intended to make
|
||||
easy: every cluster above is a candidate, and each says what it would take to build and what it
|
||||
would cost to drop. Resolving D12 sets D3 and [architecture.md §10](architecture.md)'s Phase 2.
|
||||
D12 was not resolved by choosing to work faster, and in the end not by moving clusters into §7
|
||||
either, except the one: plugins. Every other cluster above stays in scope, and each still says what
|
||||
it would take to build. That is the list. D3 is delivered, and
|
||||
[architecture.md §11](architecture.md)'s build order is what followed.
|
||||
|
||||
@@ -0,0 +1,488 @@
|
||||
# Panorama
|
||||
|
||||
**Status:** Draft · 2026-09-19
|
||||
**Companion to:** [requirements.md](requirements.md) §3.11 FR-MRG-1 … 11, D18, S15 · [architecture.md](architecture.md) §5.2, §6.2
|
||||
|
||||
The first merge (§3.11): several frames, rotated about one point, become one
|
||||
photograph. This document is how that lands on the pipeline that exists now —
|
||||
which stages, where each runs, how the composite is produced in chunks when it
|
||||
is larger than any texture or any memory, what is ported from where, and what
|
||||
the keypoint model may be under D8.
|
||||
|
||||
---
|
||||
|
||||
## 1. Why it is worth the work
|
||||
|
||||
The audience shoots panoramas and leaves the application to stitch them. That
|
||||
is the same workflow break dust was (`spot-removal.md` §1): a RAW editor that
|
||||
does everything but the one thing, and the photographer's work ends up in a
|
||||
JPEG produced by a tool that never saw the RAW.
|
||||
|
||||
It is also the merge whose alignment problem is smallest. A panorama is a
|
||||
rotation — three parameters per frame plus a focal length — with no depth to
|
||||
recover. HDR merge and focus stacking share its data model (D18) and most of
|
||||
its machinery (FR-MRG-3, 5, 6, 7, 10, 11 are written to be general); building
|
||||
the panorama first builds the shared part on the easiest geometry.
|
||||
|
||||
## 2. Non-goals
|
||||
|
||||
- **Not structure-from-motion.** No translation is solved for. A hand-held set
|
||||
with parallax gets its ghosts hidden by seam placement, and a set with real
|
||||
parallax is not a panorama. COLMAP's front end is the right mental model;
|
||||
its back end is the wrong problem.
|
||||
- **Not a multi-source Version.** D18. The composite is a file, and nothing in
|
||||
the catalog, the sidecar format or sync learns about cross-references.
|
||||
- **Not boundary fill.** Painting pixels that were never captured is the pixel
|
||||
editing §1.3 excludes. Auto-crop is the tool.
|
||||
- **Not automatic.** The tool proposes an alignment and writes nothing until
|
||||
the photographer confirms. Same rule as spot removal and D17, for the same
|
||||
reason: a merge that silently omits or misplaces a frame is the failure this
|
||||
application must not have.
|
||||
- **Not HDR-panorama in one pass.** Until HDR merge exists on its own, a
|
||||
bracketed panorama is bracketed frames merged first, then stitched.
|
||||
|
||||
## 3. What is new, precisely
|
||||
|
||||
Nearly all of it, unlike spot removal. The pipeline renders one source to one
|
||||
texture; nothing in the tree detects keypoints, estimates a rotation, warps
|
||||
into a projection, finds a seam, or blends a pyramid. What exists and is
|
||||
reused:
|
||||
|
||||
| Exists | Where | Reused for |
|
||||
|---|---|---|
|
||||
| Render a source through the fused pass, with a linear f16 output mode | `dr-gpu` demosaic → `AdjustPass`, `OutputMode::LinearWorking` | FR-MRG-2's camera-space input, as a compose entry with no operations and the profile uniforms neutral (S15.3) |
|
||||
| Tiled rendering with a priority scheduler | ARCH §5.3 | Pulling source tiles on demand into an output chunk (§5 below) |
|
||||
| A non-CFA source entering the pipeline | `Demosaicer::from_rgba8` | The composite's decode path, if the container is a TIFF (S15.1) |
|
||||
| DNG matrices read through rawler | `dr-decode::profile` | The composite's decode path, if the container is a DNG |
|
||||
| Static-shape ONNX under tract, heads decoded in Rust | `dr-segment` | The keypoint model (§6) |
|
||||
| A batch worker with its own `GpuContext`, activity row, cancel | `dr-ui::export` | FR-MRG-7 verbatim |
|
||||
| The 16-bit TIFF encoder with metadata sub-IFDs | `dr-export::encode` | FR-MRG-3's writer, extended to linear samples |
|
||||
| Multi-select in the grid | `collections_ui::selected` | The entry point |
|
||||
|
||||
New: a `core/dr-pano` crate holding the geometry (keypoints, matching, the
|
||||
rotation solve), a set of WGSL passes in `dr-gpu` (reprojection, gain,
|
||||
seam, pyramid blend), the chunked output driver, the container writer, and
|
||||
the dialog.
|
||||
|
||||
## 4. The stages, and where each runs
|
||||
|
||||
FR-MRG-10 states the rule; this is the table it was written from.
|
||||
|
||||
| Stage | Cost shape | Runs on | Why |
|
||||
|---|---|---|---|
|
||||
| Source to camera-linear | per pixel, full res | GPU, the existing pipeline | It *is* the pipeline, stopped early |
|
||||
| Keypoint detection | once per frame, at 1024 px | CPU, tract (NEON on the tablet) | Bounded by frame count, not output size. Same runtime faces and masks use. Hand-written WGSL convolutions for a model that runs five times would be work with no visible gain. |
|
||||
| Descriptor matching | K² × D per pair | CPU, SIMD | 2048² × 64 × 10 pairs ≈ 3 GFLOP — tens of milliseconds |
|
||||
| Rotation solve, bundle adjustment | 3N + 1 parameters, Levenberg–Marquardt | CPU | Microseconds. Not parallel work. |
|
||||
| Preview reprojection | per pixel, proxy res | GPU, interactive | Projection and horizon changes re-warp N proxies at frame rate |
|
||||
| Full-resolution warp | per output pixel | GPU, chunked (§5) | The heaviest thing in the application |
|
||||
| Gain compensation | per overlap region | GPU reduction, then N scalars | Sums, on the histogram pass's pattern (ARCH §5.5) |
|
||||
| Seam finding | per overlap pixel | GPU-friendly variant | Graph cut resists the GPU; a distance-transform or per-column DP seam does not. The algorithm is chosen for the GPU, not for the paper. |
|
||||
| Multi-band blend | per pixel × levels | GPU, chunked | Laplacian pyramids are separable convolutions — the detail stage's shape |
|
||||
| Encode | per pixel, once | CPU, streamed per chunk row | As export does |
|
||||
|
||||
## 5. Chunked in output space
|
||||
|
||||
FR-MRG-11 forbids holding the composite as one texture, and two facts force it
|
||||
before memory does:
|
||||
|
||||
- `max_texture_dimension_2d` is 8192 on many mobile GPUs and 16384 on desktop.
|
||||
A three-row panorama is routinely 20 000 px wide.
|
||||
- Five 24 MP frames at working precision are ~1 GB together. The tablet does
|
||||
not have it.
|
||||
|
||||
**The geometry is known before any full-resolution pixel exists.** Alignment
|
||||
runs on proxies; what comes out is a rotation per frame, a focal length, a
|
||||
projection and an output rectangle. From those, every output pixel's source
|
||||
coordinates in every frame are a closed-form function. That is what makes
|
||||
chunking simple rather than clever:
|
||||
|
||||
```
|
||||
for each output chunk C (e.g. 2048 × 2048, in output space):
|
||||
frames_in(C) = frames whose projected footprint intersects C
|
||||
for each frame F in frames_in(C):
|
||||
source tiles T(F, C) = tiles of F that project into C, plus a margin
|
||||
render T(F, C) to scene-linear through the pipeline's tile cache
|
||||
warp T(F, C) into C's coordinate frame ← GPU
|
||||
gain-correct, seam, blend within C ← GPU, with overlap
|
||||
read C back, encode its rows ← CPU, streamed
|
||||
```
|
||||
|
||||
The working set is one chunk, its per-frame warped copies, and the source
|
||||
tiles that fed them. It does not grow with the composite.
|
||||
|
||||
**The blend needs a margin.** A Laplacian pyramid of L levels reads
|
||||
2^L pixels beyond the chunk edge; a chunk is therefore rendered with a margin
|
||||
of that width and the margin discarded after the blend. Seams cross chunk
|
||||
boundaries and must agree on both sides: the seam is found once at a reduced
|
||||
resolution over the whole overlap (which fits — it is a mask, not an image),
|
||||
then upsampled into each chunk. The same is true of gain: the scalars are
|
||||
solved once from proxy-resolution overlaps and applied everywhere.
|
||||
|
||||
**Source tiles are the pipeline's tiles.** ARCH §5.3's cache keys by
|
||||
`(VersionId, tile, zoom, graph_hash_prefix)`; the merge asks for tiles of a
|
||||
neutral graph at zoom 1 and gets the same caching every other consumer does.
|
||||
A tile pulled for one chunk is usually needed by the neighbouring chunk, and
|
||||
stays hot for it.
|
||||
|
||||
### 5.1 The tap — S15.3, answered by reading the composer
|
||||
|
||||
The fused shader's order, fixed by `operation.rs`'s own tests: warp → as-shot
|
||||
white balance → operations → base curve → camera matrix → store. The store is
|
||||
either the display encode or, in `OutputMode::LinearWorking`, an unclipped
|
||||
`rgba16float` of linear sRGB. That mode exists for the detail stage and is
|
||||
selected from the operations, never by a caller flag, so that a shader and
|
||||
the texture bound to it cannot disagree.
|
||||
|
||||
The merge wants the values *before* the curve and matrix (FR-MRG-2), and the
|
||||
composer already makes that a matter of uniforms rather than structure: the
|
||||
white balance, the matrix and the curve's active flag are all in the reserved
|
||||
uniform block, and a fused pass with no operations, `as_shot_wb = 1`,
|
||||
`cam_to_srgb = I` and `base_curve_last.z = 0` stores exactly camera-linear
|
||||
RGB after the warp. So the tap is:
|
||||
|
||||
- `EditGraph::compose_camera_linear()` — the `LinearWorking` tail with an
|
||||
empty operation list and identity framing, paired by name with
|
||||
- `AdjustPass::render_camera_linear()` — binds the f16 target, fills the
|
||||
reserved uniforms neutral instead of from the source, returns the texture,
|
||||
- and a float readback beside the existing 8-bit one.
|
||||
|
||||
Nothing in the chain moves. **Precision:** the tap and every chunk buffer
|
||||
after it should be `rgba32float`, not f16. A 14-bit sensor has 16 384 steps
|
||||
to white; f16 has 2 048 in the top octave, and a composite that is going to
|
||||
be re-developed deserves the sensor's precision. The cost is 2× on buffers
|
||||
FR-MRG-11 already bounds.
|
||||
|
||||
**What the DNG carries as a consequence:** the first source's `Make`,
|
||||
`Model` and `UniqueCameraModel` — so `base_curve::for_body` finds the 6D's
|
||||
curve — its `ColorMatrix1`/`2` with illuminants, and its `AsShotNeutral`. The
|
||||
composite then develops through the same profile as its sources, applied
|
||||
once. The spike's 64 × 48 file (§8) already carries the matrix and neutral;
|
||||
the body name is a string.
|
||||
|
||||
## 6. The keypoint model
|
||||
|
||||
FR-MRG-8: works without weights, better with them. The licence read comes
|
||||
first (D13's lesson, S15.2).
|
||||
|
||||
| Model | Licence | Fits tract? | Position |
|
||||
|---|---|---|---|
|
||||
| **XFeat** (CVPR 2024) | Apache-2.0 | Plain convolutions, fully convolutional, the repo ships an ONNX export | **Chosen.** Fixed 1024 px input, dense heatmap and descriptor map out, NMS and top-K in Rust — the yolo26 pattern |
|
||||
| DISK | Apache-2.0 | U-Net, static | Second choice; stronger descriptors, ~3–4× the compute |
|
||||
| ALIKE | BSD-3 | Plain convolutions | Fallback if XFeat's export fails F6 |
|
||||
| ALIKED | BSD-3 | Deformable convolution in the descriptor head | Unlikely to load |
|
||||
| SuperPoint, SuperGlue, R2D2, SiLK, MASt3R | non-commercial | — | Out on licence |
|
||||
| LightGlue | Apache-2.0 | Transformer over a variable keypoint count | Not until mutual-nearest-neighbour matching fails on a real set |
|
||||
|
||||
**S15.2, 2026-09-19: XFeat loads under tract.** `tools/export-xfeat.sh`
|
||||
exports the network alone at 768×1024 — thirteen operator types, all
|
||||
standard: `Conv`, `InstanceNormalization`, `AveragePool`, `Resize`, `Slice`,
|
||||
`Transpose`, `Reshape`, `Concat`, `Add`, `Relu`, `Sigmoid`, `ReduceMean`,
|
||||
`Unsqueeze` — and
|
||||
[`examples/onnx_probe.rs`](../core/dr-segment/examples/onnx_probe.rs) loads
|
||||
the 2.8 MB file through the app's own `ort`-over-tract backend with nothing
|
||||
unsupported, in 28 ms, and runs it in **~300 ms on the reference desktop's
|
||||
CPU**. The weights ship as `models/keypoints/xfeat-1024.onnx`, recorded in
|
||||
`models/LICENCE.md`. **On the tablet** (S15.4's CPU half, same day):
|
||||
`tools/onnx-probe-on-device.sh` cross-builds the probe, and the same file
|
||||
runs in **~400 ms per frame** on the reference tablet's NEON cores (ROD2-W09,
|
||||
SM8635), with output ranges identical to the desktop's — inside NFR-MRG-1's
|
||||
1 s per frame with room to spare, and 1.3× the desktop rather than the 2×
|
||||
faces.md §9 measured for its scan. Still to do: a keypoint-level comparison
|
||||
against the PyTorch reference once the Rust decoder exists — the probe
|
||||
proves the graph runs, not that the numbers match.
|
||||
|
||||
The outputs are three maps at 1/8 resolution, 96×128 for the export size:
|
||||
64-channel descriptors, 65-channel keypoint logits (each 8×8 cell's position
|
||||
plus "none"), and a reliability heatmap. The Rust decoder is: softmax over the
|
||||
65, pixel-shuffle the first 64 to full resolution, 5×5 non-maximum
|
||||
suppression, top-k by reliability, bilinear sampling of the descriptor at
|
||||
each keypoint, L2 normalise. That is `detectAndCompute` in the reference,
|
||||
minus the network.
|
||||
|
||||
Without weights: AKAZE (BSD, `akaze` from rust-cv), which is adequate on
|
||||
well-textured overlaps and worse on sky, repeated structure and exposure
|
||||
drift — which is where a learned detector earns its place.
|
||||
|
||||
Matching is mutual nearest neighbour with a ratio test, then RANSAC on a
|
||||
rotation model. For a panorama — one lens, near-pure rotation, 20–40 %
|
||||
overlap — that is what Hugin and OpenCV's stitcher use, and it is enough.
|
||||
|
||||
## 7. What is ported from where
|
||||
|
||||
Nothing is linked; everything is read.
|
||||
|
||||
| Source | Licence | Taken |
|
||||
|---|---|---|
|
||||
| OpenCV `modules/stitching` | Apache-2.0 | The stage layout — Brown & Lowe (2007) as a set of small classes with one job each — and the warpers' projection maths |
|
||||
| OpenPano (ppwwyyxx) | MIT (verify on read) | The estimation and bundle-adjustment maths, function by function, with outputs diffed against it |
|
||||
| enblend-enfuse | GPLv2+ | Seam-line optimisation and Burt–Adelson multi-band blending |
|
||||
| Hugin `nona` | GPLv2+ | The GLSL remapper, as the reference for the WGSL warp |
|
||||
|
||||
The golden set (§8 of the requirements) is OpenCV's stitcher on the same
|
||||
inputs: a reference output to compare against, within a tolerance calibrated
|
||||
the way S9 calibrates R1.
|
||||
|
||||
## 8. The output file
|
||||
|
||||
FR-MRG-3. A linear DNG at the source's native scale: `u16` samples on the
|
||||
first source's black-subtracted scale, `WhiteLevel` = its white minus its
|
||||
black (13 023 for the 6D set: 15 070 − 2 047), `BlackLevel` = 0. Not rescaled
|
||||
to 65 535 — the sensor had 14 bits and the file says so, and a value the
|
||||
sensor could not have produced is not invented by a multiply. The first
|
||||
source's `Make`, `Model`, `UniqueCameraModel`, `ColorMatrix1/2`,
|
||||
`CalibrationIlluminant1/2`, `AsShotNeutral` and EXIF are carried, so the
|
||||
composite develops through the same profile as its sources. Named from the
|
||||
first source with a `-pano` suffix, beside it.
|
||||
|
||||
Three samples per pixel rather than a CFA: the warp resamples, and there is no
|
||||
sensor grid to mosaic back onto. Nothing else about being a RAW is lost —
|
||||
no white balance, no curve, no matrix, no clip has been applied — and the
|
||||
photographer develops the panorama afterwards as one photograph.
|
||||
|
||||
The sources are portrait frames in the 6D set: `Orientation` is applied
|
||||
before alignment (learned features are not rotation-invariant) and the
|
||||
composite is written upright with `Orientation = 1`.
|
||||
|
||||
Two containers were candidates and S15.1 decided, on 2026-09-19:
|
||||
|
||||
- **Linear DNG.** `PhotometricInterpretation = LinearRaw`, three samples per
|
||||
pixel, `ColorMatrix1` carried from the first source. Re-enters through
|
||||
rawler as `Format::Dng` with no new decode path, *if* rawler reads it back.
|
||||
What Lightroom writes.
|
||||
- **Float TIFF.** `SampleFormat = IEEEFP`, 16 or 32 bits, an ICC profile for
|
||||
the working space. Needs `Format::Tiff` and a decode path, but the writer is
|
||||
the existing encoder with a different sample type, and nothing about it is
|
||||
uncertain.
|
||||
|
||||
**Linear DNG.** [`examples/linear_dng.rs`](../core/dr-decode/examples/linear_dng.rs)
|
||||
hand-rolls a 64 × 48 `LinearRaw` DNG — one IFD, uncompressed 16-bit RGB,
|
||||
`DNGVersion`, `ColorMatrix1`, `AsShotNeutral`, `CalibrationIlluminant1` — and
|
||||
rawler 0.7 reads it back: `cpp 3`, the samples interleaved as written, the
|
||||
matrix parsed into the camera definition, and `CameraProfile::extract` builds
|
||||
the same profile it would for a camera file. ImageMagick's libraw reads the
|
||||
same bytes. What does *not* yet work is `dr_decode::decode`, which accepts the
|
||||
file as CFA and hands the pipeline three times the samples it expects: the
|
||||
`cpp == 3` branch is the work, and it is the only decode work.
|
||||
|
||||
The composite therefore enters the pipeline as a non-CFA, *linear* source —
|
||||
`from_rgba8`'s sibling with `non_linear = false` and the colour matrix carried
|
||||
from the DNG — and is developed as any RAW is. The writer is the example's
|
||||
IFD, grown up: tiled rather than one strip (FR-MRG-11 encodes per chunk), and
|
||||
carrying the first source's EXIF in a sub-IFD as `dr-export` already does.
|
||||
|
||||
## 9. Interaction
|
||||
|
||||
- The entry is the grid's selection: two or more images, one action, "Merge
|
||||
to panorama". One image, or images from different roots, and the action
|
||||
says why it is unavailable.
|
||||
- The dialog shows the aligned proxies in the chosen projection, with the
|
||||
projection, horizon and crop controls of FR-MRG-4, and the per-frame
|
||||
residuals. A frame that failed to align is named there (FR-MRG-5), and the
|
||||
merge cannot be confirmed with it in the set.
|
||||
- Confirm starts the FR-MRG-7 job. The composite appears in the grid when the
|
||||
file is written and catalogued, beside its sources, with the merge as the
|
||||
first entry in its history.
|
||||
|
||||
## 10. Order of work
|
||||
|
||||
1. **S15**, all four, before anything else. (1) and (2) are a day each and
|
||||
either can change the design.
|
||||
2. `dr-pano`: keypoints (AKAZE first, XFeat when S15.2 passes), matching,
|
||||
RANSAC, rotation solve. Unit-tested against synthetic rotations of one
|
||||
frame, where the answer is known exactly.
|
||||
3. The working-space tap, and the preview reprojection pass. At this point the
|
||||
dialog can show an alignment.
|
||||
4. The chunked driver with a feathered blend — the whole path end to end,
|
||||
writing a file, before the blend is good.
|
||||
5. Gain, seams, multi-band.
|
||||
6. The container, the catalog entry, provenance, the history entry.
|
||||
7. Tablet: NFR-MRG-1's figure, and FR-MRG-9's ceiling.
|
||||
|
||||
## 11. Where it stands — 2026-09-19, end of the first day
|
||||
|
||||
Built, on branch `merge/panorama`, in the order §10 gave:
|
||||
|
||||
| Piece | Where | State |
|
||||
|---|---|---|
|
||||
| Geometry: keypoints, matching, homography, focal, bundle adjustment, projections | `core/dr-pano` | Done; 33 tests without a model; the fixture aligns in 4.5 s |
|
||||
| XFeat at two shapes under tract | `models/keypoints`, `dr_pano::xfeat` | Done; 300 ms/frame desktop, 400 ms tablet |
|
||||
| The camera-space tap | `OutputMode::CameraLinear`, `AdjustPass::render_camera_linear` | Done, `rgba32float`, tiles by view rect |
|
||||
| Linear DNG writer, streamed | `dr_export::write_linear_dng` | Done; rawler reads it back |
|
||||
| A three-sample `RawImage` re-entering the pipeline | `dr-decode`, `DemosaicedImage::from_linear_rgb16` | Done |
|
||||
| Warp, accumulate, resolve, chunk by chunk | `dr_gpu::MergePass`, `merge.wgsl` | Done; feathered blend, scalar gain |
|
||||
| The job: load, proxies, align, gains, confirm, merge, provenance | `dr_ui::merge` | Done; `examples/merge.rs` drives it headless |
|
||||
| The page: table, preview, projection, Merge/Stop/Back; the grid's button | `merge.slint`, `merge_ui.rs` | Done; `DARKROOM_START_MERGE=a.CR2,b.CR2` lands on it |
|
||||
| Placement beside the sources through the outbox, rescan | `merge_ui.rs` | Done, untested against a server |
|
||||
|
||||
**Measured on the fixture (desktop, 12 × 20 MP, Intel adapter):** proxies
|
||||
and keypoints 4 s, alignment 4.5–12.6 s (load-sensitive: the matcher is
|
||||
every core), the merge **26 s for a 22 993 × 5 980 composite** in twelve
|
||||
bands of 2048 × 512 chunks, 45 s all told, an 825 MB DNG. NFR-MRG-1's 60 s
|
||||
holds on the desktop with room; the tablet's figure is still S15.4's open
|
||||
half.
|
||||
|
||||
**Open, in the order they matter:**
|
||||
|
||||
1. **Auto-crop (FR-MRG-4).** The merge returns a coverage mask per band and
|
||||
the file carries the black border. The largest inscribed rectangle over
|
||||
the coverage, then the DNG's `DefaultCropOrigin`/`DefaultCropSize`, so
|
||||
nothing is thrown away and the develop view opens on the picture.
|
||||
2. **Seams and the pyramid** (§10 step 5). The feather hides exposure and
|
||||
small misalignment; parallax on the near slope will show as a soft
|
||||
double edge at 1:1.
|
||||
3. **Vignetting in the tap.** The lens profile's distortion is applied
|
||||
before the fetch; its vignetting is an operation and is not. Frame edges
|
||||
are darker than their centres by the lens's falloff, and the feather
|
||||
averages them into the overlaps.
|
||||
4. **The tablet:** memory (twelve 40 MB sensor buffers on the CPU, one
|
||||
demosaiced frame at a time on the GPU), the figure, and FR-MRG-9's
|
||||
ceiling.
|
||||
5. **Horizon and drag-to-correct (FR-MRG-4, the proposed 4a).** The
|
||||
alignment failed on nothing in the fixture; the interaction waits for a
|
||||
set it fails on.
|
||||
6. **`derived_from` names sources by file name**, not content hash: the
|
||||
catalog's `content_hash` is null for most images most of the time. The
|
||||
hash can join it when the catalog has one.
|
||||
|
||||
## 12. Filling the border instead of cropping it — MI-GAN, read and measured 2026-09-19
|
||||
|
||||
Raised after the first merges: the ragged border a cylinder leaves could be
|
||||
*filled* rather than cropped away. FR-MRG-4 says no boundary fill, on
|
||||
§1.3's "not a pixel editor"; this is the evidence for deciding whether to
|
||||
revise that, not a revision.
|
||||
|
||||
**The candidate: MI-GAN** (Sargsyan et al., ICCV 2023, Picsart AI Research).
|
||||
Image inpainting designed for mobile: ~6 M parameters, plain convolutions —
|
||||
no FFT, no attention — so it quantises to int8 and runs on a phone's DSP,
|
||||
with quality close to LaMa and CoModGAN.
|
||||
|
||||
**Licence: MIT, code and weights alike** (`LICENSE` and `LICENSE-WEIGHTS`
|
||||
in the repository, read the same day). The cleanest position of any model
|
||||
in the tree — GPL-compatible, store-compatible, no grant to read around.
|
||||
|
||||
**Export.** The HuggingFace ONNX files are the *pipeline* — uint8 image and
|
||||
mask in, crop-around-mask, resize and blend inside the graph, every
|
||||
dimension dynamic — and tract refuses them (F6 again). The bare generator
|
||||
exports cleanly from the `migan_512_places2.pt` state dict at a fixed
|
||||
`1×4×512×512` (`export_migan.py` in the spike directory; the input is
|
||||
`mask − 0.5` and the masked RGB in −1..1, the output RGB in −1..1, the
|
||||
caller composites). After slimming the graph is **six operator types**:
|
||||
`Add, Clip, Conv, LeakyRelu, Mul, Resize`. 28 MB.
|
||||
|
||||
**Under tract on the reference desktop: loads in 53 ms, runs in 7.4 s per
|
||||
512 × 512 tile, f32.** That is the number. The fixture's border is two
|
||||
ragged bands across 22 993 px — roughly ninety 512-px tiles at full
|
||||
resolution — so a CPU-f32 fill is ten minutes on the desktop and longer on
|
||||
the tablet. Three ways to make it viable, none built:
|
||||
|
||||
1. **Fill at a quarter of the resolution and upsample.** Sky and scree
|
||||
tolerate it; twenty-odd tiles, about three minutes on the desktop CPU. A
|
||||
background job with the outbox's patience, not an interactive one.
|
||||
2. **int8 on the tablet's Hexagon through QNN**, where the plain-conv design
|
||||
is the point and the whole graph should run in milliseconds. The setup
|
||||
exists from the eye-state work; MI-GAN is a candidate for the same path.
|
||||
3. **A WGSL runtime for those six operators.** A project of its own, and
|
||||
the only route that would make it interactive on the desktop.
|
||||
|
||||
Whichever, the fill is a *proposal* under FR-MRG-1's rule — shown, then
|
||||
confirmed — and it would sit beside the crop, not replace it: the crop is
|
||||
free and honest, the fill is invented pixels, and the photographer chooses.
|
||||
|
||||
## 13. The fill, built — 2026-09-19, evening
|
||||
|
||||
Built the same day on `merge/fill`, on the engine (S16) rather than tract,
|
||||
and FR-MRG-4 revised to admit it: the border is *cropped or filled*, the
|
||||
photographer's choice, the crop the default.
|
||||
|
||||
**What runs.** `dr_pano::fill` is the engine-independent half: an
|
||||
`Inpainter` trait (a 512-px tile in, the same tile out) and `fill_border`,
|
||||
which owns everything the model does not — which tiles, what context, how
|
||||
to blend. `dr_pano::migan::MiGan` is the trait over the shipped generator
|
||||
under `dr_inference_engine` with the new `Role::Inpainter`, so it takes
|
||||
whichever rung the device has. The merge job runs the fill at **half the
|
||||
composite's resolution**, in a display-ish space (white balance, camera
|
||||
matrix, gamma — invertible, so the result goes back to camera-linear and
|
||||
into the same linear DNG), and the full-resolution merge samples the fill
|
||||
where no frame reached.
|
||||
|
||||
**What the spike taught, tried in order and kept or dropped.**
|
||||
|
||||
1. *Context across the coverage edge.* MI-GAN was trained on holes inside
|
||||
pictures; given a hole at the picture's edge it invents a structure along
|
||||
the open side (white streaks in the sky, on the first try). The known
|
||||
content is therefore **mirrored** across the coverage edge into the hole
|
||||
and into a 256-px ring, column by column for the top and bottom bands
|
||||
and row by row for the sides; the model interpolates between real and
|
||||
mirrored sky rather than extrapolating into nothing. *Replicated* rows
|
||||
(the edge row continued flat) streaked the grass; a detrended mix (tone
|
||||
replicated, texture mirrored) smeared; a low-pass extrapolation banded.
|
||||
Mirror stays.
|
||||
2. *Coarse to fine.* One pass at the working resolution let the boundary
|
||||
leak in — each 512 tile saw only its own corner of the hole. So a
|
||||
**coarse pass at a quarter** decides the structure with the whole border
|
||||
in a few tiles, and **fine passes in 96-px bands** from the real edge
|
||||
outward regenerate texture, each band the only unknown with the previous
|
||||
band on its near side and the upsampled coarse fill on its far side.
|
||||
3. *The seam.* A hard cut between real and invented showed as a sharpness
|
||||
step. The known mask is eroded by a **24-px feather** (48 at half
|
||||
resolution) and the fill blended in across that margin by distance to
|
||||
the real edge, smoothstep.
|
||||
4. *Partial pixels.* The seams were still visible until the cause was found
|
||||
upstream of the fill: the camera-space tap stored **black with alpha 1**
|
||||
for a pixel the lens correction pushed off the sensor, and the warp
|
||||
averaged it in — a dark, poorly interpolated fringe along every frame's
|
||||
edge that the fill then continued. `OutputMode::CameraLinear` now
|
||||
stores alpha 0 for a pixel that is not there and the merge's warp
|
||||
weights by the sampled alpha, so the fringe never enters the composite.
|
||||
The mask erosion before the fill dropped from 16 px to 4.
|
||||
|
||||
5. *What is still wrong, and why it ships anyway.* With the seams gone the
|
||||
content itself is the problem in the deep corners: the model, trained
|
||||
on Places2, puts bright cloud-and-peak shapes into a sky hole and a
|
||||
water-like band under grass — its prior for "top of a picture" and
|
||||
"bottom of a landscape", not anything in the context (the same shapes
|
||||
appear with the mirror capped, uncapped, and on the CPU as on TensorRT).
|
||||
Thin borders are fine; that is most of a hand-held sweep. So the fill
|
||||
ships **experimental**: opt-in, previewed, its knobs on the page and
|
||||
in the sidecar, and `cargo run -p dr-ui --example fill` re-runs any
|
||||
merge's dumped input (`DR_FILL_DUMP=dir`) stage by stage in seconds so
|
||||
the next attempt is made from the picture, not from a seven-minute
|
||||
merge. Candidates for that attempt: a context that is not a mirror at
|
||||
all in deep holes (the coarse pass's own answer, iterated), a sky
|
||||
detector that fills sky by extrapolating the gradient and leaves the
|
||||
model to texture, or a different model.
|
||||
|
||||
**Measured, the fixture's twelve frames (22 991 × 5 978), 348 tiles at
|
||||
half resolution.** 312 s on ONNX Runtime's CPU pool on the reference
|
||||
desktop (≈ 0.8 s a tile). On TensorRT fp16: **100 s**, of which 60 ms a
|
||||
tile was the engine hashing the 28 MB model on every acquire (fixed, the
|
||||
hash is taken at open) and 150 ms a tile the GPU itself — throttled:
|
||||
`trtexec` on the same engine read 23 ms at noon on a cool machine and
|
||||
152 ms that evening after two hours of builds, nvidia-smi showing SW power
|
||||
cap and thermal slowdown. Cool, the fill is ~10 s. The TensorRT engine
|
||||
compiles once, in 13 minutes, cached under the inference directory.
|
||||
|
||||
**The runtime is a packaging matter.** Arch's `onnxruntime-opt-cuda` has
|
||||
no TensorRT provider ("not enabled in this build") and its CUDA provider
|
||||
does not load against cuDNN 9, so on this machine the app fell to ORT CPU
|
||||
until the official `onnxruntime-linux-x64-gpu_cuda13` tarball (1.30.0,
|
||||
which links the system CUDA 13.4 and TensorRT 10.16) was unpacked and
|
||||
named with `DARKROOM_ORT_DIR`; `/usr/lib/darkroom` is searched too, for a
|
||||
package that ships it. §12's point 2 for the tablet is unchanged.
|
||||
|
||||
**On the page.** A *Border* choice beside the projection — *Crop to the
|
||||
picture* / *Fill the border* — with a caption saying what the fill is; the
|
||||
preview re-renders filled when chosen, at preview resolution, so the
|
||||
choice is seen before it is confirmed (FR-MRG-1). Greyed out with the reason
|
||||
when `migan-512.onnx` is not in the model directory. Under the fill, while
|
||||
it is experimental, its six knobs as sliders — working scale, edge
|
||||
erosion, coarse pass, band width, mirror depth, seam feather — each
|
||||
committing a redraw of the preview. A filled merge's sidecar says `border
|
||||
filled` with the knobs used, and its default crop is still the inscribed
|
||||
rectangle.
|
||||
|
||||
**Ships.** `models/inpaint/migan-512.onnx` (LFS, 28 MB, MIT,
|
||||
`models/LICENCE.md`), installed by the PKGBUILD and unpacked by the APK
|
||||
beside the face and scene models; `tools/export-migan.sh` regenerates it
|
||||
from the upstream checkpoint.
|
||||
+529
-74
@@ -1,6 +1,6 @@
|
||||
# DarkRoom — Requirements Specification
|
||||
|
||||
**Status:** Draft v0.1 · 2026-08-08
|
||||
**Status:** Living document · first written 2026-08-08 · audited 2026-09-19
|
||||
**Owner:** Duncan Tourolle
|
||||
|
||||
A cross-platform, non-destructive RAW photo editor for Linux desktop and Android, in the
|
||||
@@ -22,7 +22,7 @@ metadata, and exports finished images.
|
||||
| Platform | Priority | Notes |
|
||||
|---|---|---|
|
||||
| Linux desktop | Primary | X11 and Wayland. Development and reference platform. |
|
||||
| Android | Primary | Tablet-first; phone supported. Shares the image core. |
|
||||
| Android | Primary | A 12-inch tablet (D15). Phones are not a target: the build runs on one, and nothing is designed for one. Shares the image core. |
|
||||
|
||||
Other platforms (Windows, macOS, iOS) are explicitly out of scope for v1, but the architecture
|
||||
must not foreclose them. In practice this means the GPU abstraction and the image core must not
|
||||
@@ -57,6 +57,15 @@ These are the user's stated requirements, restated as testable criteria.
|
||||
| **R4** | HW acceleration and parallelism | All per-pixel work runs on GPU compute. CPU work (decode, I/O) is parallelised across cores. The UI executor never blocks on image work (NFR-ARCH-1). |
|
||||
| **R5** | Work on downscaled proxies for display | Display pipeline operates at viewport resolution, not source resolution. The tiling clauses this criterion used to carry have been struck — see below. |
|
||||
| **R6** | Nextcloud integration | Browse, download, and upload images and edit metadata against a Nextcloud instance, offline-capable. |
|
||||
| **R7** | Judge anywhere, on evidence | Rating and flag are reachable from every view that shows a photograph, and apply to the one on screen. Everything the app computes about a frame in aid of culling — clipping, focus, burst membership, per-face state — is shown as evidence the photographer reads, and **no code path writes a rating or flag without a user action** (FR-CULL-13). |
|
||||
|
||||
**On R7, added 2026-09-19.** Stated from use rather than from the brief. Two things prompted it.
|
||||
Judgement keys had been built into the grid alone, so a photograph opened in develop — the view a
|
||||
photographer is most sure about — could not be rated without leaving it; the rule is that judging
|
||||
follows the photograph, not the view. And specifying per-face signals (FR-CULL-8a) forced the
|
||||
question of what a signal is *for*, which sharpened what this document already said in FR-CULL-5:
|
||||
the app may know a great deal about a frame and may say all of it, and it never holds the pen.
|
||||
R7 is the user-level statement; FR-CULL-13 is the testable one.
|
||||
|
||||
**On R1's tolerance.** An earlier draft required output to be *bit-identical* across platforms.
|
||||
That is not achievable and the requirement has been corrected. Floating-point compute results
|
||||
@@ -68,14 +77,17 @@ R1 is therefore stated as a bounded tolerance — a defined maximum per-pixel de
|
||||
ΔE2000 for colour or ULPs at the working precision. **The threshold must be fixed before spike S9**,
|
||||
because S9 both validates R1 and calibrates what the achievable tolerance actually is.
|
||||
|
||||
Where genuine bit-identity is required — cache keys, edit-graph hashing (§5.2 invariant 3) — it
|
||||
Where genuine bit-identity is required — cache keys, edit-graph hashing (ARCH §3.4, ARCH §6.13) — it
|
||||
applies to *integer* operations on CPU-side state, which are deterministic, never to GPU float
|
||||
results.
|
||||
|
||||
**On R5's tiling.** An earlier draft added two clauses to R5's criterion: *"only visible tiles are
|
||||
computed; panning recomputes only newly exposed tiles"*. They have been struck, and the reason is
|
||||
the same one FR-DSP-2 was rewritten for rather than implemented:
|
||||
[frame-budget.md](frame-budget.md) measured it.
|
||||
the one [frame-budget.md](frame-budget.md) measured — the same measurement that argues FR-DSP-2
|
||||
should be rewritten rather than implemented. FR-DSP-2 itself has **not** been rewritten: it
|
||||
stands as written until spike S6 has run on constrained Android hardware, because the desktop
|
||||
measurement cannot speak for a device whose GPU memory the image exceeds (decided 2026-09-19;
|
||||
see the note under FR-DSP-2).
|
||||
|
||||
R5's actual demand is met and tested. The display pipeline works at viewport resolution:
|
||||
`Framing::view` shrinks the sampled region while the render target keeps its size, so zooming raises
|
||||
@@ -190,7 +202,7 @@ in the catalog when it was trashed and the path it came from; restore moves it b
|
||||
Permanent delete removes the file first and the catalog row second, and a delete of something
|
||||
already gone counts as success.
|
||||
|
||||
A flag alone would not survive invariant 5.2.4: the catalog is rebuildable from sources, so a
|
||||
A flag alone would not survive ARCH §6.12 (the catalog is a rebuildable index): the catalog is rebuildable from sources, so a
|
||||
rescan would find every "deleted" file still in the library and re-index it. The folder is the
|
||||
durable fact and the row is the convenience — which also means the scanner shall exclude the trash
|
||||
folder, and that a user can recover by hand without DarkRoom. Derived data keyed on the file
|
||||
@@ -298,7 +310,7 @@ pub trait Operation: Send + Sync {
|
||||
/// Parameter values → GPU work. No UI types cross this boundary.
|
||||
fn encode(&self, enc: &mut ComputeEncoder, ctx: &TileContext);
|
||||
|
||||
/// Identity for cache invalidation (see §5.2 invariant 3).
|
||||
/// Identity for cache invalidation (see ARCH §3.4, ARCH §6.13).
|
||||
fn params_hash(&self) -> u64;
|
||||
}
|
||||
|
||||
@@ -333,7 +345,7 @@ pub enum WidgetKind {
|
||||
```
|
||||
|
||||
**The pipeline crate shall not depend on the UI toolkit.** Descriptors carry data, never widgets.
|
||||
This keeps the edit chain testable headless (see §9's golden-image tests, which must link no UI)
|
||||
This keeps the edit chain testable headless (see §8's golden-image tests, which must link no UI)
|
||||
and is what allows one operation to render differently on touch and desktop.
|
||||
|
||||
**FR-DEV-3b — Frontend presentation mapping.** The frontend maps `ParamKind` to a concrete control
|
||||
@@ -401,8 +413,9 @@ within 0.06 in linear sRGB; the baked lookup's interpolation error stays under o
|
||||
value; and the shader agrees with the CPU model, which agrees in turn with an independent
|
||||
reference implementation.
|
||||
|
||||
**Open:** how the chosen stock persists. Sidecar parameters are `f32` and the stock list is
|
||||
data-driven, so neither an index nor a name fits the existing shape.
|
||||
*Resolved 2026-09-19:* the chosen stock persists **by id** in its own sidecar field, not as a
|
||||
parameter — `core/dr-pipeline/src/sidecar.rs` records why an index was rejected (installing a
|
||||
profile would silently change which film every existing photograph was developed on).
|
||||
|
||||
**FR-DEV-3g — AI denoise.** Learned denoising operating in the raw domain, ideally jointly with
|
||||
demosaic.
|
||||
@@ -438,6 +451,29 @@ missing tag into a visibly wrong image, and most files have no tag.
|
||||
in develop with no user action, and its sidecar is byte-identical to that of the same frame shot in
|
||||
landscape.
|
||||
|
||||
**FR-DEV-3i — Model-found masks.** A mask layer's part may be a thing a model found in the
|
||||
photograph, selected by pointing at it: one **subject** — this dog, not that one — from an
|
||||
instance model, or one **category** — all the sky, all the foliage — from a semantic model. The
|
||||
selection is stored as *identity* (which run, which instance or which category name, with the
|
||||
run's signature) so that two devices selecting the same subject hold the same value and merge per
|
||||
field under FR-NC-9, and a layer whose run no longer matches reads as **stale** rather than
|
||||
silently masking something else. The mask arrives approximately right and soft, and it is then an
|
||||
ordinary layer: FR-DEV-3's edge treatment shapes it, FR-DEV-19b's strokes correct it, FR-DEV-19a
|
||||
composes it with a gradient or a range, and FR-DEV-19c shows it. Inference is local
|
||||
(NFR-SEC-4); the models and their grants are named in `models/LICENCE.md` and D14.
|
||||
|
||||
*Added 2026-09-19, because it was built and §7 still said it was deferred.* The deferral treated
|
||||
subject masking as an AI feature to be copied from Adobe or darktable; what was built is the
|
||||
shape §7's own note asked for — a mask that behaves like a hand-drawn one — and it has been the
|
||||
primary way a local adjustment is made since the watershed hierarchy failed on real photographs
|
||||
(`docs/segmentation.md` §15).
|
||||
|
||||
One departure from FR-DEV-19 is recorded rather than hidden. A model-found part's coverage is
|
||||
also written to the sidecar, run-length coded beside the layer, because a stored subject that is
|
||||
"reproducible by running the model again" renders as nothing on a device or in a batch export
|
||||
that never runs a model. It is a materialisation of the identity, not the edit: it takes no part
|
||||
in equality or merge, and the identity remains what the part means.
|
||||
|
||||
**FR-DEV-4 — Ordered, GPU-resident execution.** The pipeline executes as a sequence of GPU
|
||||
compute stages. Intermediate results remain in GPU memory between stages. **Processed pixels
|
||||
shall reach the display without a CPU round-trip.** *(This is a hard architectural constraint —
|
||||
@@ -554,6 +590,17 @@ processes approximately 2000px of data, not 60MP.
|
||||
intersecting the viewport are computed. Panning computes only newly exposed tiles; already-valid
|
||||
tiles are reused.
|
||||
|
||||
> **Status, 2026-09-19: as written, unbuilt, and waiting on S6.** The interactive path does not
|
||||
> tile, and [frame-budget.md](frame-budget.md) argues it should not on the reference desktop: a
|
||||
> fused pass over a 4K viewport costs 4.5 ms of a 16 ms budget, a tile cache could save at most
|
||||
> that, and the one stage over budget is a convolution that tiling makes worse. That argument is
|
||||
> a desktop measurement. The case this clause was written for — an image larger than the GPU
|
||||
> memory of a mid-range Android device (NFR-RES-2, ARCH §6.2) — has not been measured, and S6 is
|
||||
> the spike that measures it. Until it runs the clause stands, so that a decision is taken on a
|
||||
> number from the device the clause is about rather than from the one it is not. If S6 finds the
|
||||
> fused pass inside budget there too, FR-DSP-2 becomes a scheduling concern for export and
|
||||
> thumbnailing as frame-budget.md proposes; if not, S6 names the stage to tile.
|
||||
|
||||
**FR-DSP-3 — Interactive latency.** Moving a slider updates the visible region within one frame
|
||||
budget at proxy resolution. When a full-resolution result is needed it is computed
|
||||
asynchronously, and the proxy result remains on screen until it is ready.
|
||||
@@ -592,7 +639,7 @@ colours on the second display is a correctness defect, not a polish item.
|
||||
|
||||
DarkRoom ships **one adaptive interface**, not separate touch and desktop applications. A single
|
||||
Slint codebase reflows by available space and input modality, guaranteeing feature parity by
|
||||
construction. Phones are out of scope for v1 (§1.3); the layout family spans tablet and desktop.
|
||||
construction. Phones are not a target (D15); the layout family spans tablet and desktop.
|
||||
|
||||
**FR-UI-1 — Layout breakpoints.** The interface adapts across at least two layout classes:
|
||||
|
||||
@@ -613,6 +660,14 @@ mode.
|
||||
> fires only on a desktop window dragged narrow. What portrait needs is the dock, which is the
|
||||
> aspect axis above and not a second class (D-N7).
|
||||
|
||||
> **Amended 2026-09-19.** The expanded row has said "filmstrip" since the table was written, and
|
||||
> the build had the photo roll on demand in both classes — the compact row applied everywhere. In
|
||||
> the expanded class the roll is **open by default** and closable, and it is shown for a set of
|
||||
> files named on the command line as much as for a library, because a set of photographs is a set.
|
||||
> It carries the place FR-UI-8 remembers — how many, which one, its name, and what the filter is
|
||||
> narrowing to — so that develop and the grid read as one interface with two views of the same
|
||||
> set, not two screens joined by a button.
|
||||
|
||||
**FR-UI-2 — Input modality.** The interface detects and adapts to the active input method, which
|
||||
is independent of layout class: a tablet may have a keyboard and pointer attached, and a desktop
|
||||
may have a touchscreen. Modality affects control sizing and affordances (FR-DEV-3b), not layout —
|
||||
@@ -630,6 +685,16 @@ functionality is touch-only.
|
||||
menus, and scroll-wheel adjustment on numeric controls. Keyboard shortcuts cover navigation,
|
||||
rating, and common adjustments. Neither is required for any operation to be reachable.
|
||||
|
||||
> **Amended 2026-09-19.** "Rating" above is not qualified by view and was built as though it were:
|
||||
> the judgement keys lived in the grid alone. They shall work in every view that shows a
|
||||
> photograph, applying to the **one on screen** — in develop, the open photograph, never a
|
||||
> selection left behind in the grid — with a pointer and touch equivalent in the same view, since
|
||||
> a tablet has no number row. Judging in develop does **not** advance to the next frame:
|
||||
> auto-advance belongs to FR-CULL-4's mode, where moving on is the point, and in develop the
|
||||
> photographer is working on the frame in front of them. The current rating and flag are shown
|
||||
> wherever they can be set, and on the roll's cells, so stepping along a set shows what has been
|
||||
> judged.
|
||||
|
||||
**FR-UI-6 — Shared component library.** Touch and desktop presentations are variants of shared
|
||||
components, not parallel implementations. A new operation (FR-DEV-3c) becomes usable on both
|
||||
without frontend work.
|
||||
@@ -1088,6 +1153,11 @@ mandatory.
|
||||
- **One-key reject**, plus the full rating, flag, and colour-label axes
|
||||
- **Filter to unjudged**, so a session resumes where it stopped
|
||||
- Keyboard-driven on desktop; single-thumb reachable on tablet
|
||||
- **Evidence beside the frame** (added 2026-09-19) — FR-CULL-13's signals, readable at a glance
|
||||
without leaving the mode or opening the frame
|
||||
- **The last import as a scope** (added 2026-09-19) — the set a culling session most often starts
|
||||
from, reachable in one step and counted, expressed as a term of the selector language (ARCH §9.2) rather
|
||||
than as a special collection
|
||||
|
||||
**FR-CULL-5 — Burst and near-duplicate grouping.** Group frames by capture-time proximity and image
|
||||
similarity, allowing a burst to collapse to one representative and be judged as a unit.
|
||||
@@ -1106,6 +1176,31 @@ originals — a substantially smaller sync problem than develop parity.
|
||||
*Design note:* pinch-zoom accidentally triggering ratings is a documented defect in Lightroom
|
||||
mobile. Gesture and rating targets must not overlap.
|
||||
|
||||
**FR-CULL-13 — Evidence, never verdicts.** Added 2026-09-19; numbered after the people clauses
|
||||
because it governs them too. Everything the app computes about a frame in aid of culling is
|
||||
**evidence**: shown beside the frame, filterable and sortable through the selector language (ARCH §9.2),
|
||||
and — within a burst — permitted to *propose* which frame represents the group. The signals are
|
||||
the raw histogram, clipping and focus (FR-CULL-3), burst membership (FR-CULL-5), and per-face eye
|
||||
state and head pose (FR-CULL-8a). The list grows by adding to it, never by adding a second kind of
|
||||
thing.
|
||||
|
||||
What evidence may not do: change a rating, a flag, a colour label, or trash membership. There shall
|
||||
be **no code path from a signal to a judgement write** without a user action between them, and a
|
||||
proposal — a representative, "three of these have eyes closed" — is accepted by a press, never by a
|
||||
timeout or a default. This restates FR-CULL-5's ground rather than extending it: automated
|
||||
selection is distrusted because its documented failure is rejecting the only frame of a moment
|
||||
because someone blinked, and the remedy is not a better classifier, it is that the classifier does
|
||||
not hold the pen.
|
||||
|
||||
Evidence says what it measured. A focus figure names its region; a per-face count names the
|
||||
faces; a signal that could not be computed — no faces found, the original not on this device — is
|
||||
shown as absent, never as zero.
|
||||
|
||||
*Acceptance:* a test enumerates every write to the rating and flag axes and shows each reachable
|
||||
only from an input event. The evidence for a frame is visible in FR-CULL-4's mode and on its grid
|
||||
cell without opening it, and a filter on any one signal returns exactly the set whose chips show
|
||||
it.
|
||||
|
||||
### 3.9.1 People
|
||||
|
||||
Face recognition was deferred in §7 through the 2026-08-08 calibration. It is undeferred here in a
|
||||
@@ -1121,10 +1216,17 @@ frames with the bride in them" across a 4,000-image wedding is a culling operati
|
||||
the differentiator.
|
||||
|
||||
**What is deliberately not in scope**, because it is the failure FR-CULL-5 names: no automated
|
||||
*selection*. Nothing here rejects a frame, ranks a face, scores a smile, or detects a blink. The
|
||||
*selection*. Nothing here rejects a frame, ranks a face, scores a smile, ~~or detects a blink~~. The
|
||||
feature produces a **filter**, never a judgement. The user's rating axes remain the only thing that
|
||||
rejects a photograph.
|
||||
|
||||
> **Amended 2026-09-19.** "Detects a blink" is struck. The exclusion was always of *judgement*,
|
||||
> and a blink is a fact about a frame of the same kind as a clipped highlight: reporting it is
|
||||
> evidence, acting on it is the failure. FR-CULL-8a detects eye state and head pose; FR-CULL-13
|
||||
> says what may be done with them — shown, filtered, sorted, proposed — and what may not. The
|
||||
> sentence that survives is the one that matters: the user's rating axes remain the only thing
|
||||
> that rejects a photograph.
|
||||
|
||||
**FR-CULL-8 — Face detection.** The app shall detect faces in library images as a background job,
|
||||
producing per-face a bounding box, five-point landmarks, a detector confidence, and a 512-dimension
|
||||
embedding.
|
||||
@@ -1178,6 +1280,50 @@ interaction target at any point, and survives being killed and restarted with no
|
||||
beyond the in-flight image. No face is stored whose aligned crop was upsampled beyond a stated
|
||||
factor; the crop source resolution is recorded per face (`crop_px`) and is auditable.
|
||||
|
||||
**FR-CULL-8a — Per-face state.** Added 2026-09-19. For every detected face the app shall record,
|
||||
as derived data under FR-CULL-12:
|
||||
|
||||
- **Eye state** — one open-probability per eye, from a classifier over a crop around each eye
|
||||
landmark. Per face this reads as both open, one closed, or both closed; per frame, as a count.
|
||||
- **Head pose** — yaw, pitch and roll, solved from the five landmarks against a generic face
|
||||
template. No model: this is geometry the detector has already paid for. *Facing the camera* is
|
||||
the pose within a stated band, and is the proxy for eye contact this document adopts — because
|
||||
gaze estimation has no redistributable weights (D13), and because in the photographs where it
|
||||
matters, groups and children and events, a turned head is the decision and averted eyes are
|
||||
often the better frame.
|
||||
|
||||
Both are terms in the selector language (ARCH §9.2, FR-CULL-11) — "everyone's eyes open", "facing the
|
||||
camera" — composable with a person and with every other term. Both feed FR-CULL-13's evidence and
|
||||
may propose a burst representative (FR-CULL-5); neither may set a rating or a flag.
|
||||
|
||||
**The weights ship under the same test as every other model**: redistributable under a licence
|
||||
compatible with GPLv3 and with Flatpak, F-Droid and Play, with the training data's terms read as
|
||||
well as the weights' (D13). The candidate identified 2026-09-19 is OCEC — MIT for code and
|
||||
weights, data under ODC-By 1.0 and Apache 2.0, six variants from 112 KB to 6.4 MB, a 24×40 crop
|
||||
per eye, sub-millisecond on CPU, opset 17 with batch-norm already folded. Its published F1 of
|
||||
0.99 is on its own crops; ours are cut from a five-point landmark, so the number is measured on
|
||||
this library before it is believed.
|
||||
|
||||
**Built 2026-09-19, in part** ([faces.md §17](faces.md)). Eye state ships: OCEC over a box cut
|
||||
from the lid contour of InsightFace's `2d106det`, plus a sunglasses classifier (SGC, MIT) whose
|
||||
answer takes precedence over the eyes it hides, and — the part the measurement forced — per eye
|
||||
the source pixels across the box and the sharpness of the patch, so an eye too small, too soft, or
|
||||
hidden by the turn of the head is recorded as *unreadable* rather than read as closed. The filter
|
||||
is the "Eyes open" chip on the library's people filter: with people chosen, it asks about their
|
||||
faces, and it drops a frame only on a closed eye that could be read. Head pose is **not** built;
|
||||
the hidden eye of a turned head is caught by its collapsed contour instead, and the selector-term
|
||||
form ("everyone's eyes open", composable in a saved collection) waits on FR-CULL-11's person term,
|
||||
which the grid filter also predates. The landmark model is under the InsightFace grant, accepted
|
||||
on the same terms as the detector and embedder (faces.md §2.2a, decision 2026-09-19: this project
|
||||
will not be commercial), so the weights test above is met by the two classifiers and not by the
|
||||
third model.
|
||||
|
||||
*Acceptance:* on a labelled set of at least 500 faces from the reference library, eye state is
|
||||
within a stated tolerance of its published F1, reported separately for glasses, profile, and
|
||||
faces under 60 px; facing-the-camera agrees with a hand-labelled split at a stated rate. Both are
|
||||
recomputed by re-indexing, neither is written to a sidecar, and the indexing pass stays inside
|
||||
FR-CULL-8's acceptance.
|
||||
|
||||
**FR-CULL-9 — Calibrated identity.** Face similarity shall be expressed as a **calibrated
|
||||
probability that two faces are the same person**, not as a raw embedding distance. Every threshold
|
||||
in the subsystem — clustering, suggestion, auto-confirmation — shall be stated in that probability
|
||||
@@ -1216,7 +1362,7 @@ Confirmation is explicit. A face is either **suggested** (the system's inference
|
||||
(the user's judgement), and the two are never conflated in storage or in display. Suggestions may be
|
||||
recomputed freely; confirmations are user data and are never overwritten by a later inference pass.
|
||||
|
||||
**FR-CULL-11 — People as a selector term.** A person shall be a term in the §5 selector language,
|
||||
**FR-CULL-11 — People as a selector term.** A person shall be a term in the selector language (ARCH §9.2),
|
||||
composable with every other term.
|
||||
|
||||
This is the requirement that pays for the subsystem, and it is nearly free once FR-CULL-10 exists:
|
||||
@@ -1243,6 +1389,14 @@ ordinary FR-CULL-10 merge, not a special case.
|
||||
|
||||
### 3.10 Extensibility and plugins
|
||||
|
||||
> **Post-v1, decided 2026-09-19.** Every clause in this section, and NFR-SEC-6 which exists for
|
||||
> it, is marked `(post-v1)` on its defining line and is outside the v1 count. The section stays as
|
||||
> the design of record — an extension point that is built as though a plugin might one day reach it
|
||||
> is cheaper than one retrofitted — but nothing here is owed a tag, and §7's row is the statement
|
||||
> of scope. The audit that prompted this found the register saying both things at once: §7 had
|
||||
> deferred "Plugin API" in a bare row since the first draft while these 23 clauses counted against
|
||||
> coverage, which measured the contradiction rather than the software. D16 defers with the section.
|
||||
|
||||
FR-DEV-3c already buys one form of this: an operation is added by writing a declaration, and the
|
||||
develop panel grows its controls without a frontend change. This section extends that property past
|
||||
the compiler. **A plugin is a file dropped into a directory; the application uses it without being
|
||||
@@ -1258,7 +1412,7 @@ lives.
|
||||
So the classes are ordered by expense, and the rule is: **an extension point is declarative unless
|
||||
it is demonstrably impossible to make it so.**
|
||||
|
||||
**FR-PLG-1 — Three plugin classes.** The app shall support exactly three, and no fourth shall be
|
||||
**FR-PLG-1 — Three plugin classes.** *(post-v1)* The app shall support exactly three, and no fourth shall be
|
||||
introduced without a recorded decision.
|
||||
|
||||
| Class | Mechanism | Covers | Executes code |
|
||||
@@ -1277,7 +1431,7 @@ one compiler version and one build of every crate it touches; and an in-process
|
||||
the whole application's privileges, which would make NFR-SEC-4 a promise about third-party code
|
||||
rather than a property of the system.
|
||||
|
||||
**FR-PLG-1a — Plugins orchestrate; the GPU does pixels.** No plugin interface shall pass
|
||||
**FR-PLG-1a — Plugins orchestrate; the GPU does pixels.** *(post-v1)* No plugin interface shall pass
|
||||
full-resolution pixel data across a sandbox boundary. A computational plugin receives handles and
|
||||
small buffers, and expresses per-pixel work as class-1 shader passes it declares. This is what keeps
|
||||
WebAssembly's arithmetic penalty irrelevant, and it is also ARCH §6.1 applied to plugins: a plugin
|
||||
@@ -1285,7 +1439,7 @@ must not be the reason a result round-trips through the CPU.
|
||||
|
||||
#### Declarative plugins
|
||||
|
||||
**FR-PLG-2 — The node declaration is the plugin format.** The schema documented in
|
||||
**FR-PLG-2 — The node declaration is the plugin format.** *(post-v1)* The schema documented in
|
||||
`core/dr-pipeline/ops/README.md` — parameters, uniform expressions, WGSL, helpers, activity,
|
||||
presentation, attributes, order, tests — shall be readable at **load time** as well as build time,
|
||||
from a plugin directory, with no change to what a declaration means.
|
||||
@@ -1300,7 +1454,7 @@ generated `match` is faster than an interpreted one and because their tests must
|
||||
the same way and for the same reason that a declared node is today indistinguishable from a
|
||||
hand-written one.
|
||||
|
||||
**FR-PLG-2a — Fragment nodes and pass nodes.** Two templates, and a plugin author chooses by
|
||||
**FR-PLG-2a — Fragment nodes and pass nodes.** *(post-v1)* Two templates, and a plugin author chooses by
|
||||
answering one question: does this operation need to read a pixel other than its own?
|
||||
|
||||
| | Fragment node | Pass node |
|
||||
@@ -1315,20 +1469,20 @@ A declaration shall state which it is, and the cost shall be visible to the user
|
||||
listing, because a chain of pass nodes is how a fast application becomes a slow one without any
|
||||
single decision having been wrong.
|
||||
|
||||
**FR-PLG-2b — Mask generators are declarative.** `Linear` and `Radial` mask sources are already
|
||||
**FR-PLG-2b — Mask generators are declarative.** *(post-v1)* `Linear` and `Radial` mask sources are already
|
||||
geometry in normalised coordinates rasterised by a shader (ARCH §5.4). A plugin shall be able to
|
||||
contribute a mask generator on the same terms — declared parameters plus a WGSL function from
|
||||
normalised coordinates to coverage — reaching luminosity-range, colour-range, and further gradient
|
||||
forms with no code. Mask sources that are *identity into a segmentation* (`Regions`, `Subject`)
|
||||
are not declarative and belong to class 3.
|
||||
|
||||
**FR-PLG-2c — Scopes split at the existing seam.** A scope plugin is a class-1 compute shader
|
||||
**FR-PLG-2c — Scopes split at the existing seam.** *(post-v1)* A scope plugin is a class-1 compute shader
|
||||
producing a small bin buffer plus a class-2 view drawing it. This is the split already in place for
|
||||
the histogram — `dr-gpu` counts, `dr-ui` shapes, Slint draws — and it holds for waveform,
|
||||
vectorscope and RGB parade without change. The counting half shall not read back full-resolution
|
||||
pixels (ARCH §5.5).
|
||||
|
||||
**FR-PLG-2d — The vocabularies stay closed.** `WidgetKind`, `attributes`, and the parameter `kind`
|
||||
**FR-PLG-2d — The vocabularies stay closed.** *(post-v1)* `WidgetKind`, `attributes`, and the parameter `kind`
|
||||
list remain closed enumerations, and a plugin may use them but shall not extend them. The reason
|
||||
given in the node README strengthens rather than weakens here: a typo that creates a new category
|
||||
containing exactly one control is indistinguishable from a deliberate new category until somebody
|
||||
@@ -1338,7 +1492,7 @@ control that was never written.
|
||||
|
||||
#### View plugins
|
||||
|
||||
**FR-PLG-3 — Declared slots, typed contracts.** A view plugin shall be a Slint component compiled at
|
||||
**FR-PLG-3 — Declared slots, typed contracts.** *(post-v1)* A view plugin shall be a Slint component compiled at
|
||||
runtime and instantiated into a **named slot** the application declares — not a licence to draw
|
||||
anywhere in the window. Each slot states the data it provides and the callbacks it accepts, and a
|
||||
component that does not match its slot's contract shall be rejected at load with a message naming
|
||||
@@ -1347,7 +1501,7 @@ the mismatch.
|
||||
Slots are a closed list under the same reasoning as FR-PLG-2d, and the composition rules of
|
||||
ARCH §4.3a continue to apply: a slot describes what a view *is for*, never how much room it has.
|
||||
|
||||
**FR-PLG-3a — A view plugin cannot be trusted with the UI thread.** Class 2 is the one class with no
|
||||
**FR-PLG-3a — A view plugin cannot be trusted with the UI thread.** *(post-v1)* Class 2 is the one class with no
|
||||
sandbox — an interpreted component runs on the UI executor and can violate NFR-ARCH-1 by looping.
|
||||
The application shall therefore watchdog slot rendering, disable a component that exceeds a stated
|
||||
budget, and report which plugin was disabled. A view plugin that fails shall leave the slot empty
|
||||
@@ -1355,12 +1509,12 @@ and the application usable; it shall never take down the window.
|
||||
|
||||
#### Computational plugins
|
||||
|
||||
**FR-PLG-4 — One interface per extension point, versioned.** Each class-3 extension point shall
|
||||
**FR-PLG-4 — One interface per extension point, versioned.** *(post-v1)* Each class-3 extension point shall
|
||||
define an explicit interface, versioned independently, and a plugin shall declare which version it
|
||||
implements. Interfaces are the only surface a computational plugin can reach: there is no ambient
|
||||
filesystem, no network, and no access to the catalog.
|
||||
|
||||
**FR-PLG-4a — Capabilities are granted, never assumed.** A plugin that needs to read a file or
|
||||
**FR-PLG-4a — Capabilities are granted, never assumed.** *(post-v1)* A plugin that needs to read a file or
|
||||
reach the network shall declare the capability, and the user shall grant it explicitly with the
|
||||
reason shown. A plugin's declared capabilities shall be visible before installation, and a plugin
|
||||
that requests none — which is the expected case for a segmentation strategy — shall be installable
|
||||
@@ -1371,7 +1525,7 @@ socket cannot send a photograph anywhere, and that is a structural property rath
|
||||
|
||||
#### Versioning
|
||||
|
||||
**FR-PLG-5 — Three independent version numbers.** Conflating any two of these produces a wrong
|
||||
**FR-PLG-5 — Three independent version numbers.** *(post-v1)* Conflating any two of these produces a wrong
|
||||
answer in both directions, so the format shall carry all three.
|
||||
|
||||
| Number | Owned by | Governs | Moves when |
|
||||
@@ -1384,11 +1538,11 @@ The common case is a plugin renaming a parameter: sidecars break while the inter
|
||||
moves. The converse also occurs. Binding sidecar compatibility to the interface version would be
|
||||
wrong in both.
|
||||
|
||||
**FR-PLG-5a — Declared, never inferred.** A plugin shall state its interface version explicitly.
|
||||
**FR-PLG-5a — Declared, never inferred.** *(post-v1)* A plugin shall state its interface version explicitly.
|
||||
Deducing it from which keys are present produces files that are ambiguous between two versions, and
|
||||
the ambiguity surfaces years later as a wrong render rather than as an error.
|
||||
|
||||
**FR-PLG-5b — A supported window, and a written policy.** The application shall support the current
|
||||
**FR-PLG-5b — A supported window, and a written policy.** *(post-v1)* The application shall support the current
|
||||
interface version and at least one predecessor, and the deprecation policy shall be stated in the
|
||||
plugin authoring documentation rather than decided per release under pressure.
|
||||
|
||||
@@ -1396,13 +1550,13 @@ plugin authoring documentation rather than decided per release under pressure.
|
||||
means default — the rule `active:`, `presentation:` and `define:` already follow. Holding that
|
||||
discipline is what keeps the version number nearly stationary.
|
||||
|
||||
**FR-PLG-5c — Adapt at the boundary, normalise inward.** A plugin declaring an older interface
|
||||
**FR-PLG-5c — Adapt at the boundary, normalise inward.** *(post-v1)* A plugin declaring an older interface
|
||||
version shall be adapted at load into the current internal representation, and nothing downstream
|
||||
shall be able to tell. Version branches threaded through the pipeline are how this becomes
|
||||
unmaintainable; the single adaptation point is the same discipline that lets a declared node and a
|
||||
hand-written one be one thing by the time anything reads them.
|
||||
|
||||
**FR-PLG-6 — Migrations are data.** Because every parameter is an `f32` addressed by a flat
|
||||
**FR-PLG-6 — Migrations are data.** *(post-v1)* Because every parameter is an `f32` addressed by a flat
|
||||
`op_id.param_id` key, a schema migration is a rewrite table rather than code. A plugin shall be able
|
||||
to declare migrations between consecutive parameter schema versions, supporting at minimum rename,
|
||||
rescale, and default-for-a-new-parameter.
|
||||
@@ -1411,14 +1565,14 @@ The **application** applies them, chained, at load, so that everything downstrea
|
||||
current-schema parameters. Migrations shall be testable through the same declared-test mechanism as
|
||||
the node itself.
|
||||
|
||||
**FR-PLG-6a — `params_version` is a promise, not a hint.** Bumping it locks every older build out of
|
||||
**FR-PLG-6a — `params_version` is a promise, not a hint.** *(post-v1)* Bumping it locks every older build out of
|
||||
the edits that use it (FR-PLG-9). It shall be bumped only when an older build would genuinely
|
||||
*misread* the file, and never merely because a parameter was added — absence already means default,
|
||||
which already means neutral. This obligation belongs in the authoring documentation in as many words.
|
||||
|
||||
#### Sidecars
|
||||
|
||||
**FR-PLG-7 — The sidecar records identity, not location.** Each version block shall record, for
|
||||
**FR-PLG-7 — The sidecar records identity, not location.** *(post-v1)* Each version block shall record, for
|
||||
every plugin it depends on: the plugin id, its version, its parameter schema version, and a content
|
||||
hash of the artefact. These merge key-wise like every other line in the format (FR-NC-9).
|
||||
|
||||
@@ -1432,7 +1586,7 @@ The content hash carries a second benefit: "install filmic 2.1" resolves to exac
|
||||
original edit was rendered with, which makes substitution and silent render drift detectable rather
|
||||
than merely regrettable.
|
||||
|
||||
**FR-PLG-8 — A missing plugin shall never cost an edit.** The sidecar already preserves lines it
|
||||
**FR-PLG-8 — A missing plugin shall never cost an edit.** *(post-v1)* The sidecar already preserves lines it
|
||||
does not understand verbatim and writes them back untouched, so a machine lacking a plugin cannot
|
||||
destroy an edit that uses it. That property is now load-bearing and shall be treated as such.
|
||||
|
||||
@@ -1456,7 +1610,7 @@ Beyond preservation:
|
||||
*Acceptance:* a sidecar written with a plugin installed, opened and saved on a machine without it,
|
||||
is byte-identical to the original.
|
||||
|
||||
**FR-PLG-9 — Forward skew is quarantined, not guessed.** Where a version block names a parameter
|
||||
**FR-PLG-9 — Forward skew is quarantined, not guessed.** *(post-v1)* Where a version block names a parameter
|
||||
schema version **newer** than the installed plugin declares, the application shall divert that
|
||||
operation's parameters into the preserved-verbatim path *before applying any of them*, and mark the
|
||||
version quarantined.
|
||||
@@ -1485,7 +1639,7 @@ with an older one, is byte-identical afterwards.
|
||||
|
||||
#### Distribution
|
||||
|
||||
**FR-PLG-10 — Registry resolution and verified artefacts.** Installation shall resolve a plugin id
|
||||
**FR-PLG-10 — Registry resolution and verified artefacts.** *(post-v1)* Installation shall resolve a plugin id
|
||||
through a registry the user has configured, with a default registry shipped. The application shall
|
||||
verify the artefact against the hash or signature the registry states before loading it.
|
||||
|
||||
@@ -1504,7 +1658,7 @@ dependency on one company's availability.
|
||||
|
||||
#### Authoring and operations
|
||||
|
||||
**FR-PLG-11 — Plugins are validated, and validation is the author's tool.** A declared plugin's
|
||||
**FR-PLG-11 — Plugins are validated, and validation is the author's tool.** *(post-v1)* A declared plugin's
|
||||
`tests:` shall be runnable outside the application, against the same interpreter that loads it, via
|
||||
a command-line validator. The application shall ship a scaffold command producing a minimal working
|
||||
plugin of each class.
|
||||
@@ -1514,7 +1668,7 @@ manner `build.rs` already reports them. **A malformed plugin shall be skipped, n
|
||||
build script's exit-on-error discipline is right for an author with a compiler open and wrong for a
|
||||
user opening their library.
|
||||
|
||||
**FR-PLG-12 — Failure is attributable and revocable.** The application shall record per-plugin
|
||||
**FR-PLG-12 — Failure is attributable and revocable.** *(post-v1)* The application shall record per-plugin
|
||||
timing and error counts, surface them in the plugin listing, and allow any plugin to be disabled
|
||||
without uninstalling it — including on the next launch after a crash, so a plugin that prevents
|
||||
startup can be disabled by someone who cannot start the application.
|
||||
@@ -1524,6 +1678,144 @@ application starts, opens a photograph, names the responsible plugin, and contin
|
||||
|
||||
---
|
||||
|
||||
### 3.11 Merging images
|
||||
|
||||
Several photographs become one. Panorama is the first merge and the only one specified; HDR merge
|
||||
and focus stacking share its data model (D18) and are still deferred in §7. The clauses below are
|
||||
written for the panorama and, where a clause is general to any merge, say so.
|
||||
|
||||
**FR-MRG-1 — Panorama from a selection.** Two or more selected images are aligned and blended
|
||||
into one composite, which is written as a new source file per D18. The tool is never automatic:
|
||||
it proposes an alignment, the photographer sees it and confirms, and nothing is written before
|
||||
that press.
|
||||
|
||||
The stated audience (D11) shoots panoramas and currently leaves the application to stitch them,
|
||||
which is the workflow break FR-DEV-8 was added to close for dust. It is also the first of the
|
||||
three §7 merges, and the one whose alignment problem is smallest — a rotation about one point,
|
||||
with no depth to recover — so it is where the shared machinery is built.
|
||||
|
||||
**FR-MRG-2 — What is stitched.** Each source enters the merge in **camera space**: after black
|
||||
and white levels, demosaic and lens distortion correction, and before everything else — no white
|
||||
balance, no base curve, no camera matrix, no edit. The composite carries the first source's body,
|
||||
colour matrix and as-shot neutral, so that it is developed afterwards exactly as one of its
|
||||
sources would be: the camera profile, the white balance and every operation in §3.3 are applied
|
||||
once, to the composite, in its own develop.
|
||||
|
||||
This is the clause that decides what the output *is*. Stitching the rendered edits is what a JPEG
|
||||
stitcher does; the result cannot be re-developed, and any difference between the frames' edits
|
||||
becomes a seam. Stitching camera-space pixels produces a photograph the camera could have taken,
|
||||
and nothing is applied twice. The cut sits *below* the profile, not above it, for a reason S15.3
|
||||
found in the pipeline: the base curve is part of the profile (FR-DEV-3e) and is applied to every
|
||||
frame of a known body, so a composite that baked it in and then developed as one would render the
|
||||
curve twice. Lens correction alone sits above the cut, because a distorted frame does not align.
|
||||
White balance sits below it because the sensor saw the same light in every frame: un-balanced
|
||||
camera RGB agrees across the overlaps whether or not the camera's auto white balance drifted, and
|
||||
the balanced values would not.
|
||||
|
||||
**FR-MRG-3 — The output file.** *(general to any merge)* A RAW, in every sense that survives a
|
||||
warp: camera-linear samples at the **source's native scale** — integers on the first source's
|
||||
black-subtracted scale with its white level, never rescaled to fill 16 bits — with the first
|
||||
source's body, colour matrices, illuminants and as-shot neutral carried, and its capture
|
||||
metadata. What a merge cannot preserve is the colour filter array: a warped image has no sensor
|
||||
grid to re-mosaic onto, so the composite is three samples per pixel. Named from the first source
|
||||
with a stated suffix and placed beside it. Where the sources' folder is not writable — a
|
||||
remote-only tier, a read-only mount — it goes where an export goes (FR-EXP-6) and the interface
|
||||
says so before the merge starts.
|
||||
|
||||
The container is a linear DNG (S15.1, decided 2026-09-19): rawler reads back what the application
|
||||
writes, and the composite re-enters as `Format::Dng` through the decoder every camera DNG uses.
|
||||
The point of the whole clause is that the panorama is *developed afterwards* — white balance,
|
||||
profile, exposure, everything in §3.3 — as one photograph, from the sensor's own numbers.
|
||||
|
||||
**FR-MRG-4 — Projection and framing.** Cylindrical, spherical or perspective, chosen from the
|
||||
field of view and overridable; the horizon levelled from the estimated rotations, overridable by a
|
||||
drag; the border either cropped to the largest inscribed rectangle or filled, the photographer's
|
||||
choice, the crop the default.
|
||||
|
||||
*The auto-crop is non-destructive* (built 2026-09-19): it is the composite's default crop, not a
|
||||
cut — the whole merge including its border is in the file, and resetting the crop shows it.
|
||||
|
||||
*The fill is generative and opt-in* (revised the same day, having first said "no boundary
|
||||
fill"): MI-GAN (Sargsyan et al., ICCV 2023; MIT code and weights, `models/LICENCE.md`) paints
|
||||
the uncovered border from the picture's own edge, under the inference engine (S16, [inference.md](inference.md)). It is a
|
||||
proposal under FR-MRG-1 — shown on the page, chosen against the crop, confirmed before the merge
|
||||
— never the default, and a merge that used it says so in its sidecar (`border filled`) so the
|
||||
invented pixels are declared, not passed off as captured. §1.3 still holds: the fill touches only
|
||||
pixels no frame reached, never the photograph; a filled merge keeps the inscribed crop as its
|
||||
default so the honest picture is one reset away. The filler is loaded from the model directory
|
||||
when present and the choice is greyed out, with the reason, when it is not.
|
||||
|
||||
*Experimental, as shipped 2026-09-19.* The fill is right in thin borders and wrong in deep
|
||||
corners, where the model invents cloud and water where there is sky and grass
|
||||
(panorama.md §13); it ships opt-in with **every knob on the page** — working scale, edge
|
||||
erosion, coarse pass, band width, mirror depth, seam feather — each redrawing the preview, and
|
||||
the knobs used are written into the sidecar's `merge` line beside `border filled`. The knobs
|
||||
leave the page when the defaults are right; the sidecar record stays.
|
||||
|
||||
**FR-MRG-5 — Honesty of failure.** *(general to any merge)* A frame that cannot be aligned is
|
||||
named, with why — too few matches, no overlap with any other frame, a residual above the stated
|
||||
bound — and the merge stops. Never a silent drop, never a best-effort composite with a frame
|
||||
missing.
|
||||
|
||||
The same rule as `spot-removal.md`'s and D17's: a tool that quietly alters or omits part of a
|
||||
photograph is the failure this application must not have, and here the omission would be an
|
||||
entire frame.
|
||||
|
||||
**FR-MRG-6 — Provenance.** *(general to any merge)* The composite's sidecar carries
|
||||
`derived_from`: the content hashes of its sources in order, and the merge parameters. The history
|
||||
records the merge as the first entry, and export metadata declares the composite as one. Sources
|
||||
trashed later leave the list dangling; the panel says so and nothing is blocked.
|
||||
|
||||
Provenance, not dependency. The composite renders from itself alone; `derived_from` exists so the
|
||||
photographer, and anyone they hand the file to, can see what it is made of. The audience is
|
||||
RAW-literate and a composite declares itself (D17). Nothing here specifies C2PA.
|
||||
|
||||
**FR-MRG-7 — Execution.** *(general to any merge)* A background job on the pattern FR-EXP-7
|
||||
established: its own thread, its own `GpuContext`, a row in the activity panel, cancellable with
|
||||
NFR-ARCH-3's bound. Alignment runs at proxy resolution and drives the preview; the full-resolution
|
||||
warp and blend run only on confirm.
|
||||
|
||||
**FR-MRG-8 — Model-optional.** Keypoint detection works without any model weights and better
|
||||
with them, the convention `dr-segment` set. Weights that ship are recorded in `models/LICENCE.md`
|
||||
before they land, under D8's compatibility test, and their absence degrades quality rather than
|
||||
disabling the feature.
|
||||
|
||||
**FR-MRG-9 — Platform.** Desktop first. Android runs the same code within NFR-RES-2 and
|
||||
FR-MRG-11, with a stated ceiling on frame count and source resolution, refused with a message,
|
||||
rather than an out-of-memory kill.
|
||||
|
||||
**FR-MRG-10 — Where the work runs.** *(general to any merge)* Every per-pixel stage of a merge —
|
||||
rendering the sources, the preview reprojection, the full-resolution warp, gain compensation, the
|
||||
seam and the blend — runs on the GPU as WGSL, under ARCH §6.4. The stages that are not per-pixel
|
||||
— keypoint detection at proxy resolution, descriptor matching, and the rotation solve over a few
|
||||
parameters per frame — run on the CPU, and the specification says so rather than leaving it to be
|
||||
"moved later".
|
||||
|
||||
The per-pixel stages are the whole cost, and the tablet is where the cost is paid: the output is
|
||||
larger than any single photograph the pipeline has rendered, and a CPU blend of it would take
|
||||
minutes there. The CPU stages are bounded by frame count, not by output size — detection is once
|
||||
per frame at 1024 px, on the same runtime faces and masks use — and moving a small model's
|
||||
convolutions to hand-written WGSL is real work for no visible gain. Seam finding is the one
|
||||
classic stage that resists the GPU; the seam algorithm is chosen for the GPU, not for the paper.
|
||||
|
||||
**FR-MRG-11 — Tiled in output space.** *(general to any merge)* No stage may hold the composite as
|
||||
one texture, on any platform. The warp, seam, blend and encode proceed in output-space chunks, each
|
||||
pulling only the source tiles that project into it, so the working set is one chunk plus its
|
||||
sources' tiles regardless of how large the composite is.
|
||||
|
||||
Two facts force this before memory does. `max_texture_dimension_2d` is 8192 on many mobile GPUs
|
||||
and 16384 on desktop, and a three-row panorama is routinely 20 000 px wide — the composite would
|
||||
not fit a texture even with the memory to spare. And five 24 MP frames at the working precision
|
||||
are ~1 GB together, which the tablet does not have. ARCH §6.2 applies to the composite as it
|
||||
applies to a source, and retrofitting it would be the rewrite it warns about.
|
||||
|
||||
**Non-goals, fixed now.** No translation solve or parallax correction — seam placement is the
|
||||
tool for a hand-held set, and a photograph with real parallax is not a panorama. No HDR panorama
|
||||
in one operation until HDR merge exists on its own. No live re-stitch: a different projection or
|
||||
crop after the fact is a new file, not an edit. No video.
|
||||
|
||||
---
|
||||
|
||||
## 4. Non-functional requirements
|
||||
|
||||
### 4.1 Performance targets
|
||||
@@ -1565,7 +1857,13 @@ the class where it was expressed.
|
||||
Everything the photographer produced or navigated to is a different matter, and none of it may be
|
||||
touched by a resize. That is the list in the criterion, and it is the testable half.
|
||||
|
||||
**Performance regressions fail the build.** §9's benchmark suite runs per-commit; a regression
|
||||
**NFR-MRG-1 — Merge latency.** For five 24 MP frames: the alignment preview (FR-MRG-7) within
|
||||
5 s on the reference desktop and 15 s on the reference tablet, of which keypoint detection is at
|
||||
most 1 s per frame on the tablet's CPU (measured 2026-09-19 at ~0.4 s, S15.4); the full merge
|
||||
written to disk within 60 s on the desktop. The tablet's full-merge figure is set by S15's blend
|
||||
measurement, not guessed here.
|
||||
|
||||
**Performance regressions fail the build.** §8's benchmark suite runs per-commit; a regression
|
||||
beyond a stated tolerance is a build failure, not a notification. Performance work rots otherwise.
|
||||
|
||||
### 4.2 Reliability
|
||||
@@ -1590,7 +1888,7 @@ ARCH §6.6 already anticipates one migration (folder ETags); there will be other
|
||||
exist before the first one.
|
||||
|
||||
**NFR-R6 — Corruption recovery.** On failing an integrity check at startup, the app offers restore
|
||||
from the NFR-R2 backup, and failing that, rebuild from sources plus sidecars per invariant 5.2.4.
|
||||
from the NFR-R2 backup, and failing that, rebuild from sources plus sidecars per ARCH §6.12.
|
||||
FR-CAT-8 is what makes that second path real for local-only users.
|
||||
|
||||
**NFR-R7 — GPU device loss.** The GPU layer treats device loss as an expected event (ARCH §6.10): detect
|
||||
@@ -1602,11 +1900,19 @@ loss mid-render.
|
||||
available, the app starts in a stated degraded mode with defined capability limits rather than
|
||||
failing to launch.
|
||||
|
||||
*This requires reconciling ARCH §6.4 and NFR-RES-2:* ARCH §6.4 says the GPU path is primary rather than an
|
||||
optimisation, while NFR-RES-2 assumes a CPU fallback on allocation failure. **Decide explicitly**
|
||||
whether v1 includes a full CPU pipeline, or whether "CPU fallback" means only tile-spill staging
|
||||
with no independent CPU render path. The latter is recommended; the former is a second full
|
||||
implementation.
|
||||
**Decided 2026-09-19: there is no CPU render pipeline.** The degraded mode is the one the
|
||||
viewer already has (`dr_ui::shared_gpu` returning `None`): the library opens, the grid and the
|
||||
culling views run on embedded previews and cached proxies, metadata, ratings and collections are
|
||||
fully editable, and develop and export are unavailable and say so. "CPU fallback" in this
|
||||
document means nothing more than staging through host memory when GPU memory is short; it never
|
||||
means a second implementation of the operations. ARCH §6.4 stands as written, and NFR-RES-2's
|
||||
fallback clause has been reworded to match. A second full pipeline was the alternative, and it was
|
||||
declined for the reason ARCH §6.4 gives: the GPU path is the product, not an optimisation of it.
|
||||
|
||||
**NFR-MRG-2 — Reproducible merges.** The same sources, the same settings and the same device
|
||||
produce a byte-identical composite. Across devices the comparison is tolerance-based, calibrated
|
||||
as S9 calibrates R1: the merge is float work end to end, and ARCH §6.13's bit-identity applies to
|
||||
integer state only.
|
||||
|
||||
### 4.3 Resource behaviour
|
||||
|
||||
@@ -1614,7 +1920,9 @@ implementation.
|
||||
size and image count. Caches are evictable under pressure.
|
||||
|
||||
**NFR-RES-2 — GPU memory.** The pipeline shall handle images larger than available GPU memory by
|
||||
tiling. GPU memory headroom is configurable, with a CPU fallback path if allocation fails.
|
||||
tiling. GPU memory headroom is configurable. Where an allocation fails, work is staged through host
|
||||
memory or refused with a typed error (`GpuError::TooLarge`) — never rendered by a CPU pipeline,
|
||||
which does not exist (NFR-R8).
|
||||
|
||||
**NFR-RES-3 — Mobile power.** On Android the app shall not render continuously when idle. Battery
|
||||
and thermal behaviour are first-class concerns; background sync respects metered-connection and
|
||||
@@ -1677,7 +1985,7 @@ user. The prohibition is therefore structural rather than configurable: the code
|
||||
upload an embedding to anyone but the user's own server do not exist. A setting can be changed by
|
||||
accident, or by a future maintainer who has forgotten why it was there; an absent code path cannot.
|
||||
|
||||
**NFR-SEC-6 — Third-party code runs inside a boundary, not beside the application.** FR-PLG-1
|
||||
**NFR-SEC-6 — Third-party code runs inside a boundary, not beside the application.** *(post-v1)* FR-PLG-1
|
||||
admits code the user did not write and the project did not review. The privacy properties asserted
|
||||
elsewhere in §4.5 are properties of *this* codebase, and none of them survives a plugin that can
|
||||
open a socket.
|
||||
@@ -1743,24 +2051,46 @@ selection.
|
||||
|
||||
### 4.8 Compatibility baseline
|
||||
|
||||
**NFR-COMPAT-1 — Supported hardware.** Every §4.1 Android figure is meaningless without this. State:
|
||||
**NFR-COMPAT-1 — Supported hardware.** Stated 2026-09-19 from what the build enforces and the
|
||||
hardware the figures are taken on. Every §4.1 Android figure is measured against this.
|
||||
|
||||
- Minimum and target Android API level (targetSdk 36 is currently required for Play distribution)
|
||||
- Minimum Vulkan version and the required feature and limit set — including whether `shaderFloat16`
|
||||
and 16-bit storage are required, since **FR-DEV-2's f16 pipeline depends on them and their
|
||||
absence would jeopardise R1**
|
||||
- Minimum device RAM, and minimum desktop Vulkan/Mesa versions
|
||||
- The **specific** reference Android device the §4.1 column is measured on, plus a secondary device
|
||||
from a different GPU vendor
|
||||
| | Baseline |
|
||||
|---|---|
|
||||
| Android API | **minSdk 28**, targetSdk 36 (`docker/android/Dockerfile` `MIN_API`/`ANDROID_API`; the linker targets 28, so the binary runs on the oldest level it claims) |
|
||||
| GPU, both platforms | A **Vulkan** adapter accepted by wgpu at its **default limits** — not the downlevel tier — because storage textures in compute shaders are required; `dr_gpu::GpuContext::device_from` states this as the floor. No optional feature is required: `required_features` is empty. The pipeline stores intermediates in `Rgba16Float` textures, which are core Vulkan; **`shaderFloat16` arithmetic is not required** and FR-DEV-2's "16-bit float" is a storage precision, not a shader-arithmetic one. The GL backend is accepted on desktop as a last resort |
|
||||
| Desktop | Linux with a Mesa or vendor Vulkan driver the reference machine's era or later; Windows under NFR-COMPAT-2's installer. No minimum Mesa version is asserted beyond "wgpu's default limits are met" |
|
||||
| RAM | No figure asserted; NFR-P8's budgets are the constraint, not the device's total |
|
||||
| **Reference Android device** | **HONOR ROD2-W09**, Qualcomm SM8635 (Adreno), 12 GB, Android 16 / API 36, 1920×3000 at 400 dpi. Every Android column in §4.1 is measured on it |
|
||||
| Second-vendor device | **None.** The clause asking for a Mali or PowerVR device is **unmet and known unmet**: no such hardware exists in the project, and S2's two-vendor run is blocked on it (#47). Until one exists, an Android figure in this document is an Adreno figure |
|
||||
|
||||
Adreno, Mali, and PowerVR diverge significantly in compute behaviour and in external-memory interop
|
||||
— exactly what spike S1 tests. The spec already applies this reasoning to desktop drivers; it
|
||||
applies at least as strongly on Android.
|
||||
applies at least as strongly on Android, which is why the missing second vendor is recorded as a
|
||||
gap rather than dropped.
|
||||
|
||||
**NFR-COMPAT-2 — Distribution channels.** State the v1 channels (e.g. Flatpak and AppImage on Linux;
|
||||
Play Store and/or F-Droid on Android). Play distribution is what makes ARCH §6.9's constraints binding —
|
||||
a sideloaded or F-Droid build could use different permissions, so the channel decision and the
|
||||
storage design are coupled.
|
||||
*A figure worth noting.* At 400 dpi the reference tablet's portrait width is roughly 768 logical
|
||||
pixels, which is under D15's assumed ~1024 and under `EXPANDED_MIN_WIDTH`. Whether Slint's scale
|
||||
factor agrees with the platform density is open issue N6; this row records the measurement and
|
||||
does not resolve it.
|
||||
|
||||
**NFR-COMPAT-2 — Distribution channels.** Decided 2026-09-19. Every channel is
|
||||
**self-distribution**: built by the Gitea CI from one commit, installed by hand, and published to
|
||||
no store or catalogue. That is a consequence of D13 — the face weights' grant rules out every
|
||||
public channel — and it is the position until those weights are replaced.
|
||||
|
||||
| Platform | Channel | Built by |
|
||||
|---|---|---|
|
||||
| Linux | Arch package (`packaging/PKGBUILD`) | CI, `packaging/pkg/` |
|
||||
| Linux | Flatpak from the manifest in `packaging/flatpak/`, installed locally — **not** submitted to Flathub | CI |
|
||||
| Android | Signed release APK, sideloaded (`docker/android/assemble-apk.sh`, the release key of 2026-09-11) — **not** Play, **not** F-Droid | CI |
|
||||
| Windows | NSIS per-user installer cross-built from Linux (`docker/windows/`, FR-PLAT-WIN-3) | CI |
|
||||
|
||||
Two consequences the channel decision has on the requirements above it. ARCH §6.9's SAF-only
|
||||
storage model was written because Play would reject anything else; a sideloaded build *could*
|
||||
request broader permissions, and FR-PLAT-AND-1 keeps SAF anyway, because the design is right on
|
||||
its own terms and because the day a public channel opens is not the day to redesign storage.
|
||||
And S11, the Play permissions dry-run, becomes a pre-publication step rather than a Tier-1 spike:
|
||||
there is nothing to submit until D13 is closed.
|
||||
|
||||
### 4.9 Accessibility and internationalisation
|
||||
|
||||
@@ -1787,7 +2117,7 @@ more in a colour-grading application than in most software.
|
||||
|
||||
Entity definitions, invariants, and all architectural constraints are specified in
|
||||
[architecture.md](architecture.md) — §3 (core abstractions), §6 (data architecture), and
|
||||
§11 (architectural constraints).
|
||||
§12 (architectural constraints — whose subsections keep their original 6.x numbers, so `ARCH §6.1` in this document means §12's first constraint, not §6.1 of the data architecture).
|
||||
|
||||
Requirements in this document that depend on an architectural guarantee cite it inline. The
|
||||
constraints most load-bearing for testability are:
|
||||
@@ -1804,13 +2134,13 @@ constraints most load-bearing for testability are:
|
||||
## 6. Decisions
|
||||
|
||||
Rationale, evidence, and the eliminated alternatives are recorded in
|
||||
[architecture.md §12](architecture.md). Outcomes only:
|
||||
[architecture.md §13](architecture.md). Outcomes only:
|
||||
|
||||
| # | Decision | Outcome |
|
||||
|---|---|---|
|
||||
| D1 | Language and UI framework | Rust + Slint, rendering through wgpu |
|
||||
| D2 | RAW decoder | rawler; LibRaw fallback behind a trait |
|
||||
| D3 | First milestone | **Provisional — depends on D12** |
|
||||
| D3 | First milestone | Delivered — [milestone-v0.1.md](milestone-v0.1.md), closed 2026-08-30 |
|
||||
| D4 | Nextcloud sync mechanism | ETag pruning, chunked upload v2, Login Flow v2 |
|
||||
| D5 | Colour management | lcms2 + GPU-side matrix/LUT transforms |
|
||||
| D6 | Shader authoring | Hand-written WGSL |
|
||||
@@ -1819,7 +2149,8 @@ Rationale, evidence, and the eliminated alternatives are recorded in
|
||||
| D9 | Operation UI model | Declarative parameter descriptors |
|
||||
| D10 | Interface strategy | One adaptive UI, tablet + desktop |
|
||||
| D11 | Product positioning | Culling-first differentiator; see below |
|
||||
| D12 | Scope versus pace | **OPEN** |
|
||||
| D12 | Scope versus pace | **DECIDED 2026-09-19** — settled by events; full scope stands, no v1 date |
|
||||
| D18 | Derived images | **DECIDED 2026-09-19** — a merge writes a new source file; no multi-source Version |
|
||||
|
||||
### D11 — product positioning
|
||||
|
||||
@@ -1834,14 +2165,28 @@ Settled by requirements calibration, 2026-08-08.
|
||||
| Ingest | Full workflow — template rename, checksum verify, dual-destination |
|
||||
| Colour defaults | Good, not obsessive — matrices plus per-body base curve |
|
||||
| Film simulation | Fujifilm explicitly targeted |
|
||||
| AI | Denoise in v1; masking deferred |
|
||||
| AI | Denoise in v1; masking deferred. Per-face eye state and head pose are in v1 **as culling evidence, not AI** (FR-CULL-8a, FR-CULL-13); gaze deferred (§7) |
|
||||
| Local adjustments | Full masking, GPU-rasterised |
|
||||
| Sync | The reason the project exists |
|
||||
| Durability | Sidecar-first |
|
||||
| Licence | GPLv3 |
|
||||
| Pace | Evenings and weekends, indefinite |
|
||||
|
||||
### D12 — scope versus pace · **OPEN**
|
||||
### D12 — scope versus pace · **DECIDED 2026-09-19**
|
||||
|
||||
**Settled by events.** The reconciliation this decision asked for never happened as a decision;
|
||||
it happened as a build. Between the calibration and this audit the milestone D3 pointed at was
|
||||
delivered and closed (2026-08-30), and the application went through eleven further releases to
|
||||
0.12.2 carrying culling, develop, masks, faces, sync, Flatpak and a Windows channel. The scope as
|
||||
calibrated stands as written, and the pace is the pace. What was reconciled is the *date*: v1 has
|
||||
none. A requirement in this document is in scope until §7 says otherwise, and "post-v1" in §7 is
|
||||
the only mechanism by which something leaves the count — used on 2026-09-19 for the plugin API,
|
||||
and for nothing else.
|
||||
|
||||
The two tensions below are kept as the record of what the decision weighed. Both resolved
|
||||
themselves the same way: the expensive selection was built anyway, and the differentiator was
|
||||
built alongside the develop chain rather than instead of it.
|
||||
|
||||
|
||||
The calibration selected an ambitious feature set — full tablet editing, full ingest, culling as a
|
||||
differentiator, complete GPU masking, AI denoise, Fuji-first colour, deep sync, sidecar durability
|
||||
@@ -1862,7 +2207,7 @@ much.
|
||||
architecture but is useful to nobody. If culling is what no existing tool does well, a culler is
|
||||
both a smaller build and a usable one — it needs no develop chain.
|
||||
|
||||
Resolving D12 sets D3 and [architecture.md §10](architecture.md)'s Phase 2.
|
||||
D12 was what set D3, and [architecture.md §11](architecture.md)'s build order followed it.
|
||||
|
||||
### D15 — target devices · **DECIDED 2026-08-22**
|
||||
|
||||
@@ -1888,7 +2233,38 @@ desktop window, not as a second interface.
|
||||
|
||||
---
|
||||
|
||||
### D13 — face inference runtime and model licensing · **RUNTIME ANSWERED, LICENSING OPEN**
|
||||
### D13 — face inference runtime and model licensing · **RUNTIME ANSWERED · LICENSING POSITION RECORDED 2026-09-19**
|
||||
|
||||
> **Runtime, reopened 2026-09-19 — to the extent of [inference.md](inference.md) §3.** The
|
||||
> pure-Rust build stands: `ort` still links nothing. What changed is that `ort::set_api` can be
|
||||
> handed the table of a `libonnxruntime` the *package* installs, and the app now looks for one at
|
||||
> launch and runs on tract only when there is none. Measured before it was built: tract runs
|
||||
> every model on one core at the same speed on a tablet and a twenty-core desktop; ONNX Runtime's
|
||||
> CPU provider alone is 3–10× that, the Hexagon at int8 runs the detectors in 1–3 ms, TensorRT
|
||||
> at fp16 in 2–3 ms. The Android APK bundles ONNX Runtime and Qualcomm's HTP libraries (§3.1 —
|
||||
> the QNN licence is read, not summarised, before a release carries them); the desktop packages
|
||||
> bundle nothing NVIDIA and use a system CUDA/TensorRT if the probe finds one that works. The
|
||||
> licensing half is unchanged.
|
||||
|
||||
> **Position, 2026-09-19.** DarkRoom is non-commercial software, built and installed by its
|
||||
> author for personal libraries, and it uses the InsightFace SCRFD detectors and ArcFace embedder
|
||||
> under their **non-commercial research grant** as such. That is the position, and it is taken
|
||||
> with the risks written down rather than around them:
|
||||
>
|
||||
> - The grant is a **use** restriction and binds every user of the app, not only the project. It
|
||||
> is not GPL-compatible and cannot become so; D8's licence covers this codebase and not those
|
||||
> weights, which `models/face/README.md` says in as many words.
|
||||
> - It is incompatible with **every public channel** — Flathub, F-Droid, Play. NFR-COMPAT-2's
|
||||
> channels are therefore all self-distribution (a CI-built APK, a Flatpak and an Arch package
|
||||
> installed by hand, an NSIS installer), and nothing is published to a store or a catalogue
|
||||
> while these weights are in the tree. Publishing is the event that reopens this decision, and
|
||||
> S14's licence search is what would close it: the OCEC eye-state weights showed a clean chain is
|
||||
> possible, and a clean detector and embedder are what is missing.
|
||||
> - The about screen names the models and their grant (NFR-SEC-5), so the person running the app
|
||||
> can read the restriction they are under.
|
||||
>
|
||||
> The runtime half is unchanged below.
|
||||
|
||||
|
||||
> **Updated 2026-08-21.** The runtime half of this decision is settled, and by a route the table
|
||||
> below does not contain. `ort` 2.0's `alternative-backend` feature *disables its linking entirely*
|
||||
@@ -1901,6 +2277,15 @@ desktop window, not as a second interface.
|
||||
>
|
||||
> **The licensing half is untouched.** The InsightFace weights are still non-commercial and still
|
||||
> unusable here. That remains what S14 has to resolve first.
|
||||
>
|
||||
> **Updated 2026-09-19.** Two further models were read for FR-CULL-8a. **Eye state:** OCEC
|
||||
> (PINTO0309) is MIT for code and weights and trained on ODC-By 1.0 and Apache 2.0 data — the
|
||||
> first face-adjacent weights found with a clean chain end to end, and the smallest by two orders
|
||||
> of magnitude. **Gaze:** none. MobileGaze, L2CS-Net and their descendants carry MIT on the
|
||||
> repository and Gaze360 in the weights, and Gaze360's research licence restricts *"models trained
|
||||
> on dataset"* by name; MPIIGaze and ETH-XGaze are no better. Head pose needs no weights at all.
|
||||
> None of this moves the detector and embedder, which are still the buffalo grant and still what
|
||||
> S14 resolves first: clean eye weights behind an unshippable detector ship nothing.
|
||||
|
||||
§3.9.1 needs to run two neural networks locally. That collides with two settled positions, and
|
||||
neither collision is small enough to leave implicit.
|
||||
@@ -1952,7 +2337,70 @@ rather than thousands.
|
||||
|
||||
---
|
||||
|
||||
### D16 — plugin licensing · **OPEN**
|
||||
### D17 — cross-frame face repair · **OPEN**
|
||||
|
||||
Raised 2026-09-19: pair FR-CULL-8a's eye state with the burst, and let a closed-eyed face in the
|
||||
chosen frame be replaced by the same person's open-eyed face from a neighbouring frame. It inverts
|
||||
FR-CULL-5's fear — the tool rescues the only frame of the moment instead of rejecting it — and it
|
||||
collides with four things this document says.
|
||||
|
||||
1. **§1.3, "not a pixel editor."** Survivable by the FR-DEV-8 precedent: a repair is numbers in
|
||||
the graph and no pixels are stored. A face repair is a spot whose source is another frame,
|
||||
aligned by the similarity the face subsystem already fits between two landmark sets and blended
|
||||
by the membrane heal already shipped. It reopens `spot-removal.md`'s non-goal that the source is
|
||||
"a patch from the same photograph", which would be revised deliberately, not quietly.
|
||||
2. **ARCH §6.3, single-source `Image`.** The real one. §7 defers panorama, HDR and focus stacking on the
|
||||
grounds that the schema cannot express an image derived from several sources. This is that case
|
||||
in a milder form — the output is still frame A, but A's sidecar now names B — and it brings the
|
||||
rules the tiers already have: B's original must be present to render A (FR-NC-6c: visible and
|
||||
priced before starting), trashing B must know A depends on it (FR-CAT-15), and `Version::merge`
|
||||
has never seen a cross-reference. Reopening ARCH §6.3 for this reopens it for the three deferred
|
||||
rows at once, which is the argument for doing it once and properly rather than for this alone.
|
||||
*D18 has since done it once, the other way:* a merge writes a new file and ARCH §6.3 stands.
|
||||
That leaves this case as the only one that would still need a cross-reference — the output is
|
||||
frame A, not a new file — so the cost above is now this feature's alone to justify.
|
||||
3. **Tone.** The source patch goes through A's chain, not B's — B demosaiced and run through A's
|
||||
parameters to the head of the detail chain, then sampled. A second small pipeline at proxy
|
||||
resolution; a second full demosaic at export. The heal hides lighting drift between frames; it
|
||||
does not hide a turned head, and no blend does.
|
||||
4. **Never automatic.** `spot-removal.md`'s own rule: a false positive silently alters a
|
||||
photograph, which is the failure this application must not have. The tool may propose — same
|
||||
person, closed here, open two frames on, small alignment residual — and applies on a press.
|
||||
|
||||
And one thing the document does not say: **provenance.** The audience is RAW-literate (D11). A
|
||||
composite declares itself — the source frame in the sidecar and the history, and in export
|
||||
metadata. Nothing here specifies C2PA; that is a decision this one depends on.
|
||||
|
||||
Deferred under D12 until FR-CULL-8a exists and spot removal's disc has, in its own words, been
|
||||
finished and used. The first cut, when it comes, is "clone from a neighbouring frame" as a spot
|
||||
source; the face-aware proposal is a layer over that.
|
||||
|
||||
### D18 — derived images · **DECIDED 2026-09-19**
|
||||
|
||||
**A merge produces a new source file, not a multi-source Version.** The composite is written
|
||||
beside its sources (FR-MRG-3), gets its own sidecar and content-hash identity, and is from then on
|
||||
an ordinary `Image`: developed, synced, exported and trashed like any other. Its sidecar carries
|
||||
`derived_from` (FR-MRG-6) as *provenance*, not as a *dependency* — trashing a source does not break
|
||||
the composite, and rendering it needs nothing but itself.
|
||||
|
||||
This is the ARCH §6.3 question §7 has been keeping open for panorama, HDR and focus stacking,
|
||||
answered once for all three. The alternative — a Version whose inputs are several other images,
|
||||
rendered live — was priced in D17: sidecar cross-references, `Version::merge` seeing a reference
|
||||
for the first time, FR-CAT-15 knowing that trashing B breaks A, FR-NC-6c requiring every source
|
||||
present and priced before A can render, and an export that decodes N files. Every one of those is
|
||||
a change to a subsystem that works today, and none of them buys the photographer anything a file
|
||||
does not. It is also what Lightroom does, and the audience (D11) knows it.
|
||||
|
||||
What it forecloses, so that it reads as a decision: re-merging with different settings is a new
|
||||
file, not an edit to the old one; and the composite occupies disk — a five-frame panorama is a
|
||||
100–200 MB file, which the photographer chose to make. D17 narrows accordingly to the one case
|
||||
where the output is still frame A, and inherits nothing from this decision but the provenance
|
||||
rule.
|
||||
|
||||
### D16 — plugin licensing · **OPEN, post-v1**
|
||||
|
||||
> Deferred with §3.10 on 2026-09-19. Still to be answered before the format is published as
|
||||
> stable, which is now a post-v1 event; nothing in v1 waits on it.
|
||||
|
||||
D8 puts the application under GPLv3. §3.10 admits third-party plugins in three forms, and the
|
||||
derivative-work question is answered differently for each — a YAML-and-WGSL declaration is data of
|
||||
@@ -1986,15 +2434,18 @@ note where deferring now constrains the design later.
|
||||
| Deferred | Note |
|
||||
|---|---|
|
||||
| Tethered shooting | — |
|
||||
| Panorama and HDR merge | **Keep the schema open** — these produce images derived from multiple sources, which §5.1's single-source `Image` cannot express. |
|
||||
| Focus stacking | Same provenance consideration. |
|
||||
| ~~Panorama~~ | **Undeferred 2026-09-19** — §3.11 (FR-MRG-1 … FR-MRG-11), under D18, which answers the schema question this row was holding open: a merge is a new source file, and ARCH §6.3's single-source `Image` does not change. |
|
||||
| HDR merge | Deferred. D18 answers the data model; the merge itself — exposure alignment, ghost handling, the tone of the result — is not specified. FR-MRG-3, 5, 6, 7, 10 and 11 are written to be general to it. |
|
||||
| Focus stacking | Deferred, on the same terms as HDR merge. |
|
||||
| Cross-frame face repair ("best take") | A face from a neighbouring frame of the same burst, aligned by its landmarks and blended by FR-DEV-8's heal. Mechanically a spot whose source is another photograph; **the same multi-source schema question as the two rows above, arriving early** — D17. Deferred rather than refused, with its non-goals fixed now: never automatic, geometry not corrected, source frame declared in sidecar, history and export. |
|
||||
| Gaze and eye-contact estimation | Every open gaze model read on 2026-09-19 is trained on Gaze360, MPIIGaze or ETH-XGaze, all research-only, and Gaze360's licence restricts *models trained on it* by name — the InsightFace situation again (D13). Head pose from the five landmarks is the proxy (FR-CULL-8a). Revisit when weights with a clean data chain exist; iris offset within the eye crop is the licence-free fallback if the proxy proves too weak. |
|
||||
| Print layout | — |
|
||||
| Soft proofing | Parameterise the output colour stage by an arbitrary profile so this becomes a UI addition, not a pipeline change. FR-EXP-3's print-dimension mode already half-commits to print workflows. |
|
||||
| ~~Face recognition~~ | **Undeferred 2026-08-09**, in the narrower form specified in §3.9.1 (FR-CULL-8 … FR-CULL-12): people *grouping and search*, no automated selection. Reclassified as culling rather than AI — it is the same mechanical-grouping category as FR-CULL-5, not the taste operation AI masking is. Gated on spike S14 and decision D13. |
|
||||
| AI subject masking | Deferred per D11. Note darktable shipped this in 5.6 (June 2026), so the gap is now visible. **Conditions for deferring safely:** AI denoise ships in v1 (FR-DEV-3g ✓), manual masking is excellent including GPU-rasterised drawn masks (ARCH §6.11 ✓), and the product has a clear differentiator (culling, §3.9 ✓). When it does land, copy darktable's shape — prompt-point segmentation producing an *editable* mask that behaves like a hand-drawn one — not Adobe's opaque version. |
|
||||
| ~~Face recognition~~ | **Undeferred 2026-08-09**, in the narrower form specified in §3.9.1 (FR-CULL-8 … FR-CULL-13): people *grouping and search*, no automated selection. Reclassified as culling rather than AI — it is the same mechanical-grouping category as FR-CULL-5, not the taste operation AI masking is. Gated on spike S14 and decision D13. |
|
||||
| ~~AI subject masking~~ | **Undeferred 2026-09-19** — it had been built. FR-DEV-3i states what exists: subject and category masks from local models, stored as identity, editable like a drawn mask. The shape this row asked for when it was written — point at a thing, get an editable mask — is the shape that shipped. |
|
||||
| AI upscaling | Deferred. Lower priority than denoise, which has no manual fallback. |
|
||||
| Video | — |
|
||||
| Plugin API | — |
|
||||
| Plugin API | **Post-v1**, decided 2026-09-19. Specified in full in §3.10 as the design of record; every clause there is marked `(post-v1)` and sits outside the coverage denominator. D16 (plugin licensing) defers with it. FR-DEV-3c's compile-time declarations are not plugins and remain in v1. |
|
||||
| Multi-user / server-side catalog | — |
|
||||
| Watermarking | Cheap *if* the export pipeline anticipates a compositing stage; expensive to retrofit otherwise. Consider reserving the stage now. |
|
||||
| Multiple catalogs, catalog merge | Interacts with NFR-OPS-3: preferences must not live in the catalog. |
|
||||
@@ -2022,7 +2473,7 @@ note where deferring now constrains the design later.
|
||||
| Cancellation (NFR-ARCH-3) | Assert every long-running operation observes cancellation within the stated bound, including in-flight GPU work. |
|
||||
| Layer separation (ARCH §6.5a) | CI dependency-tree assertion: no `core/*` crate may transitively depend on a UI toolkit. |
|
||||
| Operation self-description (FR-DEV-3c) | A test operation added to the registry appears in a generated panel with no frontend change. |
|
||||
| Adaptive layout (§3.5) | Snapshot tests at each breakpoint, and a resize test asserting no state loss across a layout-class transition. |
|
||||
| Adaptive layout (§3.5) | Snapshot tests at each breakpoint, and a resize test asserting no loss of *photographic* state across a layout-class transition — NFR-P11's list, with panel disclosure exempt as that clause says. |
|
||||
| Touch targets (FR-UI-3) | Automated check that interactive elements meet the 44pt minimum in touch modality. |
|
||||
| Export sizing (FR-EXP-3) | Per-mode dimension assertions, including aspect preservation, fill-crop centring, and the upscale-disabled fallback. |
|
||||
| Identity calibration (FR-CULL-9) | Reliability diagram over a hand-labelled corpus: stated probability against observed match rate, asserted within tolerance across the range — not a single accuracy figure, which would hide exactly the miscalibration this tests for. Plus a static assertion that no comparison thresholds a raw similarity. |
|
||||
@@ -2060,6 +2511,8 @@ stacks.
|
||||
| **S13** | **Slint accessibility on Android:** verify TalkBack exposure of names, roles, and values | Whether NFR-A11Y-2 is achievable in the chosen toolkit | NFR-A11Y-2 |
|
||||
| **S14** | **Face pipeline in Rust, on a real personal library:** run a detector plus an embedder over ~2,000 images through a Rust ONNX runtime, at proxy resolution, on the reference desktop. Measure per-image latency, cluster purity against hand-labelled truth, and fit the FR-CULL-9 calibration to see whether it converges on a library-sized sample. **Resolve the model licence question before writing any of it** | Whether §3.9.1 is buildable without breaking the pure-Rust dependency policy, and whether the accuracy is worth the subsystem | D13, FR-CULL-8, FR-CULL-9 |
|
||||
|
||||
| **S15** | **Panorama pre-conditions, in order of what can kill it:** (1) write a linear DNG with the `tiff` crate and read it back through rawler — decides FR-MRG-3's container; (2) export XFeat to ONNX at a fixed 1024 px input and load it under tract with zero unsupported operators — the F6 check `segmentation.md` records, and the licence read first; (3) find where in the fused chain FR-MRG-2's camera-space tap sits and what it costs to expose; (4) a tiled multi-band blend of a 100 MP output on the reference tablet, and XFeat's per-frame time on its CPU — the two halves of NFR-MRG-1 | Whether §3.11 is buildable on the pipeline as it stands, and what the tablet figure is | D18, FR-MRG-2, FR-MRG-3, FR-MRG-8, FR-MRG-11, NFR-MRG-1 |
|
||||
|
||||
### Why this order
|
||||
|
||||
**S1, S2, and S10 are the three that can invalidate the architecture.** S1 and S2 test ARCH §6.1 — the
|
||||
@@ -2072,7 +2525,7 @@ this revision it was unstated.
|
||||
this app category before designing around it is far cheaper than discovering it at submission.
|
||||
|
||||
**S9 moved to Tier 1** because it does not merely test R1 — it *calibrates* it. R1's tolerance
|
||||
threshold cannot be fixed sensibly without knowing the real cross-vendor deviation, and the §9
|
||||
threshold cannot be fixed sensibly without knowing the real cross-vendor deviation, and the §8
|
||||
golden-image strategy depends on that number.
|
||||
|
||||
Test S1 on Mesa/AMD, Intel, and NVIDIA proprietary drivers, under both X11 and Wayland. FD-based
|
||||
@@ -2090,4 +2543,6 @@ Android GPU vendors.
|
||||
- **Demosaic** — reconstructing full RGB from a colour-filter-array sensor capture.
|
||||
- **CFA** — colour filter array (Bayer, X-Trans).
|
||||
- **Sidecar** — a small file alongside the source holding edit metadata.
|
||||
- **Composite** — an image produced by a merge (§3.11) from several sources; a source file in its own right under D18.
|
||||
- **Merge** — an operation that produces a composite: panorama, HDR merge, focus stacking.
|
||||
- **Pixel pipeline** — the ordered chain of processing stages from sensor data to output.
|
||||
|
||||
+144
-121
File diff suppressed because one or more lines are too long
+3
-2
@@ -281,13 +281,14 @@ $LOCALAPPDATA\Programs\DarkRoom\
|
||||
darkroom.exe
|
||||
models\
|
||||
scrfd_500m_640.onnx scrfd_2.5g_640.onnx scrfd_10g_640.onnx arcface_mbf_b1.onnx
|
||||
2d106det_b1.onnx ocec_s_b1.onnx sgc_l_48_b1.onnx
|
||||
yolo26s-sem-ade20k.onnx yolo26s-sem-ade20k.classes.json categories.txt
|
||||
LICENSE
|
||||
uninstall.exe
|
||||
```
|
||||
|
||||
Plus a Start Menu shortcut, and nothing on the desktop unless the user ticks it. The models are
|
||||
the same seven files the APK bundles and the PKGBUILD installs; `models\` beside the executable is
|
||||
the same ten files the APK bundles and the PKGBUILD installs; `models\` beside the executable is
|
||||
where §3.2's lookup finds them. **No `LICENSE` yet**: the repository has no licence file at its
|
||||
root (the Arch package points at the system's shared GPL text), so the installer has no licence
|
||||
page until one is added — a one-file change, and the `.nsi` says where the page then goes. The face weights carry the research-only grant that
|
||||
@@ -458,7 +459,7 @@ container was the cheaper way to get a pinned MinGW and a Wine that could be thr
|
||||
| It links | Yes, first attempt once the link flags were right. 115 MB, `PE32+ … (GUI)`. Two dead-code warnings, both `cfg`-shadowed constants, fixed. |
|
||||
| It is a Windows executable | 26 imports, all Windows system DLLs. No MinGW runtime. `.rsrc` carries `PRODUCTVERSION 0,12,0,0`, `ProductName DarkRoom`, the icon. |
|
||||
| It starts | `wine darkroom-desktop.exe --version` → `darkroom-desktop 0.12.0`, exit 0, 0.1 s. |
|
||||
| The installer runs | `makensis` → 105 MB. `/S` installs the exe and seven models to `AppData\Local\Programs\DarkRoom`, writes the `HKCU` uninstall key; the installed exe runs; `uninstall.exe /S` removes directory and key. Shortcut unverifiable (§6). |
|
||||
| The installer runs | `makensis` → 105 MB. `/S` installs the exe and ten models to `AppData\Local\Programs\DarkRoom`, writes the `HKCU` uninstall key; the installed exe runs; `uninstall.exe /S` removes directory and key. Shortcut unverifiable (§6). |
|
||||
| It draws a window | Not attempted. |
|
||||
|
||||
**What the first draft got wrong**, kept in place above with a note rather than rewritten, because
|
||||
|
||||
BIN
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
@@ -11,6 +11,8 @@ same script:
|
||||
|
||||
Both come from `https://huggingface.co/Ultralytics/YOLO26`. The face weights in
|
||||
`face/` are a separate matter with a separate grant — see `face/README.md`.
|
||||
The keypoint weights in `keypoints/` and the border filler in `inpaint/` are
|
||||
the other two, and the easiest — see the last two sections.
|
||||
|
||||
## The grant
|
||||
|
||||
@@ -69,3 +71,49 @@ Neither replaces the other. Keeping both is the deliberate choice.
|
||||
The loader treats each vocabulary as model metadata rather than compiled-in
|
||||
knowledge, which is what made adding the second model a file plus a descriptor
|
||||
rather than a code change — as this document predicted it would be.
|
||||
|
||||
## `keypoints/` — XFeat, Apache-2.0
|
||||
|
||||
| File | Source | Trained on | Used by |
|
||||
|---|---|---|---|
|
||||
| `keypoints/xfeat-1024.onnx` | `weights/xfeat.pt` from `https://github.com/verlab/accelerated_features` | MegaDepth + synthetic warps, by the authors | panorama alignment (FR-MRG-8), landscape frames |
|
||||
| `keypoints/xfeat-768.onnx` | the same weights | — | the same, portrait frames |
|
||||
|
||||
Exported by `tools/export-xfeat.sh` at fixed grayscale inputs of 1024×768
|
||||
and 768×1024 — the same weights twice, because tract needs a static shape
|
||||
and a portrait frame in a landscape input wastes half of it. Only the
|
||||
convolutional network is in each file; the keypoint decoding is Rust.
|
||||
|
||||
**The repository and its weights are Apache-2.0**, read on 2026-09-19 from the
|
||||
`LICENSE` at its root, with no separate grant on the checkpoint and no
|
||||
non-commercial clause anywhere in the tree. Apache-2.0 is GPLv3-compatible
|
||||
one way — code and weights under it may be combined into a GPLv3 work — so
|
||||
this is neither the InsightFace situation (D13, a use restriction that binds
|
||||
every user) nor the Ultralytics one (D14, where the combined work becomes
|
||||
AGPL). It is the licence position this document would have wanted for every
|
||||
model in it, and it was chosen over stronger detectors partly for that reason:
|
||||
SuperPoint and SuperGlue are non-commercial, R2D2 and SiLK are CC BY-NC.
|
||||
|
||||
The training data is the authors' concern, not a licence on the weights:
|
||||
XFeat trains on MegaDepth, which is itself a research dataset, but the weights
|
||||
are released under the repository's licence without a data-derived
|
||||
restriction — unlike the gaze models §7 of the requirements declined, where
|
||||
the dataset licence restricts models trained on it by name.
|
||||
|
||||
## `inpaint/` — MI-GAN, MIT
|
||||
|
||||
| File | Source | Trained on | Used by |
|
||||
|---|---|---|---|
|
||||
| `inpaint/migan-512.onnx` | `migan_512_places2.pt` from `https://github.com/Picsart-AI-Research/MI-GAN` (Sargsyan et al., ICCV 2023) | Places2, by the authors | the panorama border fill (FR-MRG-4) |
|
||||
|
||||
Exported by `tools/export-migan.sh`: the bare 512 generator at a fixed
|
||||
`1×4×512×512`, six operator types. The tiling, the context and the blend are
|
||||
Rust (`dr_pano::fill`).
|
||||
|
||||
**MIT, code and weights alike** — `LICENSE` and `LICENSE-WEIGHTS` in the
|
||||
repository, both read on 2026-09-19, both the plain MIT text with no further
|
||||
grant. GPL-compatible, store-compatible, nothing to read around: the cleanest
|
||||
position of any model here. The training set is Places2, a research dataset,
|
||||
but the weights are released under the repository's licence without a
|
||||
data-derived restriction (contrast the gaze models §7 of the requirements
|
||||
declined, and the InsightFace grant of D13).
|
||||
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user