Let the user choose which SCRFD finds their faces

faces.md §12.3 measured what the cheapest detector costs: the small
faces in every group shot, and a dog embedded a dozen times. Which
trade is right depends on the machine doing the sweep — a desktop left
overnight and a tablet on a battery want different answers — so the
detector is now a per-device setting, Fast / Balanced / Thorough on
the settings page beside the indexing button, persisted with the rest
of the settings file.

A detector is half of a model id. Every face, marker, shard and
calibration is keyed on faces.model_id precisely so that a model change
is a new id and a re-index rather than a silent change under existing
data, and a detector change is a model change: it decides which faces
exist and where the landmarks that align them land. So each choice
names its own pipeline. 500M keeps the bare "w600k_mbf" every existing
library was written under, so an upgrade disturbs nothing; the others
are qualified. Choosing one restarts coverage from zero under the new
id, the sweep re-detects, confirmed names carry across by box overlap,
and the sync shards are keyed by the same id so a peer on another
setting neither adopts nor pollutes them. The library controller
carries the id into the sync the same way it carries the cache budget,
because the sync starts from places that have no settings in reach.

All three shape-fixed exports ship — APK, Arch, Flatpak — since a
tablet has no other way to obtain the one it was not installed with;
the APK grows by twenty megabytes for the choice.
This commit is contained in:
2026-09-11 22:12:53 +02:00
parent adf5d6cdd9
commit 4f31123b0c
20 changed files with 468 additions and 122 deletions
+2 -2
View File
@@ -19,8 +19,8 @@ pub use place::{Place, PlaceScope, Screen, StoredFilter};
pub use selector::{ColourLabel, DateSelector, FlagState, Selector, Tier};
pub use settings::{
CacheSettings, CollisionPolicy, ColourSpace, DevelopSettings, ExportFormat, ExportSettings,
ExportTarget, GroupNavigation, ImportSettings, LibrarySettings, OutputSharpening, ScreenSize,
Settings, SizingMode,
ExportTarget, FaceDetector, GroupNavigation, ImportSettings, LibrarySettings, OutputSharpening,
ScreenSize, Settings, SizingMode,
};
pub use time::{
civil_from_unix, civil_from_unix_at, format_date, parse_date, unix_from_civil, Civil,
+127
View File
@@ -157,6 +157,96 @@ pub struct FaceSettings {
/// dropped by this, whatever their size: those are the user's judgements and
/// a display preference does not overrule them (FR-CULL-12).
pub min_group_size: u32,
/// Which graph finds the faces. See [`FaceDetector`].
pub detector: FaceDetector,
}
/// TRACES: FR-CULL-8
/// Which SCRFD graph the indexing pass detects with.
///
/// Three exports of one architecture, differing only in how much computation
/// they spend, and docs/faces.md §12.3 is the measurement that made this a
/// choice rather than a constant: over the same photographs the cheapest one
/// misses the small faces in a group and reports a dog a dozen times, the
/// middle one finds 14% more faces for 12% more time, and the largest a
/// further 12% for three times the cost. Which of those is the right trade
/// depends on the machine doing the sweep — a desktop left running overnight
/// and a tablet on a battery want different answers — so it is a setting,
/// per device, like the rest of this file.
///
/// # A detector is half of a model id
///
/// Every face row, run marker, sync shard and calibration is keyed by
/// `faces.model_id` (catalog.md §10.1), and the schema's whole reason for
/// carrying that column is that *a model change is a new id and a re-index*
/// rather than a silent change under existing data. A detector change is a
/// model change: it decides which faces exist and where the landmarks that
/// align them land. So each variant names its own pipeline, and choosing
/// another one puts every image back in the queue and shows the People screen
/// for the new pipeline — empty until the sweep has run, with confirmed names
/// carried across by `record_detections`' box overlap.
///
/// The first variant's id is the bare embedder name, because that is the id
/// every library indexed before this setting existed was written under;
/// making it `scrfd_500m+w600k_mbf` would have told those libraries they had
/// never been indexed.
#[derive(Debug, Clone, Copy, Default, PartialEq, Eq, Serialize, Deserialize)]
pub enum FaceDetector {
/// 2.5 GFLOPs. The measured sweet spot: nearly the same speed, and a
/// cleaner set of faces.
#[serde(rename = "scrfd_2.5g")]
Scrfd2_5g,
/// 10 GFLOPs. The most faces, at three times the time per image.
#[serde(rename = "scrfd_10g")]
Scrfd10g,
/// 500 MFLOPs. What every library was indexed with until now.
///
/// Last, because `other` has to be: a file written by a build that knows
/// a fourth detector loads as this one rather than throwing every other
/// setting away with it. [`FaceDetector::ALL`] is the display order.
#[default]
#[serde(rename = "scrfd_500m", other)]
Scrfd500m,
}
impl FaceDetector {
/// In the order the settings page offers them: cheapest first.
pub const ALL: [FaceDetector; 3] = [
FaceDetector::Scrfd500m,
FaceDetector::Scrfd2_5g,
FaceDetector::Scrfd10g,
];
/// The shape-fixed export's file name, as `tools/fix-face-model-shapes.sh`
/// writes it and every packager installs it.
pub fn file_name(self) -> &'static str {
match self {
FaceDetector::Scrfd500m => "scrfd_500m_640.onnx",
FaceDetector::Scrfd2_5g => "scrfd_2.5g_640.onnx",
FaceDetector::Scrfd10g => "scrfd_10g_640.onnx",
}
}
/// The `faces.model_id` this detector's pipeline writes under.
///
/// The embedder is the same `w600k_mbf` in every case; see the type note
/// for why the first is bare and the others are qualified.
pub fn model_id(self) -> &'static str {
match self {
FaceDetector::Scrfd500m => "w600k_mbf",
FaceDetector::Scrfd2_5g => "scrfd_2.5g+w600k_mbf",
FaceDetector::Scrfd10g => "scrfd_10g+w600k_mbf",
}
}
/// What the picker calls it.
pub fn label(self) -> &'static str {
match self {
FaceDetector::Scrfd500m => "Fast",
FaceDetector::Scrfd2_5g => "Balanced",
FaceDetector::Scrfd10g => "Thorough",
}
}
}
impl FaceSettings {
@@ -195,6 +285,7 @@ impl Default for FaceSettings {
// a setting, so an existing library regroups identically until the
// user moves it.
min_group_size: 2,
detector: FaceDetector::default(),
}
}
}
@@ -1537,6 +1628,42 @@ mod tests {
assert_eq!(s.faces, FaceSettings::default());
}
/// The detector every existing library was indexed with must keep the id
/// those libraries were written under, or an upgrade would report every
/// one of them un-indexed.
#[test]
fn the_default_detector_keeps_the_legacy_model_id() {
assert_eq!(FaceDetector::default(), FaceDetector::Scrfd500m);
assert_eq!(FaceDetector::default().model_id(), "w600k_mbf");
}
/// Two pipelines must never share an id: the whole point of the column is
/// that their faces are not interchangeable.
#[test]
fn every_detector_has_its_own_model_id_and_file() {
let ids: std::collections::HashSet<_> =
FaceDetector::ALL.iter().map(|d| d.model_id()).collect();
assert_eq!(ids.len(), FaceDetector::ALL.len());
let files: std::collections::HashSet<_> =
FaceDetector::ALL.iter().map(|d| d.file_name()).collect();
assert_eq!(files.len(), FaceDetector::ALL.len());
}
#[test]
fn the_detector_round_trips_and_an_unknown_one_falls_back() {
let mut s = Settings::default();
s.faces.detector = FaceDetector::Scrfd10g;
let text = serde_json::to_string(&s).unwrap();
assert!(text.contains(r#""detector":"scrfd_10g""#));
let back: Settings = serde_json::from_str(&text).unwrap();
assert_eq!(back.faces.detector, FaceDetector::Scrfd10g);
let newer = r#"{"faces":{"detector":"scrfd_99g"}}"#;
let s: Settings =
serde_json::from_str(newer).expect("unknown detector should not refuse the file");
assert_eq!(s.faces.detector, FaceDetector::Scrfd500m);
}
#[test]
fn sanitise_leaves_a_real_destination_alone() {
let mut s = Settings::default();