Files
scene-actor-extraction/src/config.hpp
T
dtourolle 06373817a2 refactor(config): give the tuned constants a real provenance, and make them reachable
Two problems, both of which made a number look more settled than it is.

The provenance was a dead link. config.hpp cited
docs/rep4-optimizer-results.md for prob_threshold, extinction_sec,
anneal_sec and the expansion default. That file was renamed to
model-bakeoff.md and then rewritten; the comments were never repointed,
so the most consequential constant in the pipeline appeared to have no
source at all.

Following it up produced something worse than a broken link.
prob_threshold=0.754 comes from the ORIGINAL rep4 document (still
readable at `git show d340da7:docs/rep4-optimizer-results.md`). The
rewrite that replaced it reports finding "a real scoring bug in
optimize.py: a candidate whose hardest film's replay timed out was
averaged over survivors instead of penalized, silently rewarding partial
coverage. Affected 3 of 16 training combos". So 0.754 was fitted under
scoring that was later found wrong, the corrected sweep converged
elsewhere, and no corrected prob_threshold is recorded anywhere. The
comment now says that, along with the surviving document's own verdict
that the optimum "generalizes unevenly -- strong on 3 of 5 held-out
films, badly broken on 2".

Four constants were unreachable. ownership_logodds lived on
TrackRegistry::Config, and max_views/admit_below/rho_max on
EvidenceDiscounter::Config, which main built with the one-argument
constructor -- so nothing short of a recompile could move any of them.
rho_max's own comment defers to "the sweep (VR-007)" for where it
belongs, and that sweep could not reach it.

They now live in Config with CLI flags and are exposed to the replay
harness. ownership_logodds is worth singling out: below it a track makes
no presence claim at all, so it decides whether an actor is reported
rather than how confidently -- arguably the most consequential constant
after prob_threshold, and until now unswept and unsettable.

No behaviour change: every default is the value that was compiled in.

TRACES: AR-025, AR-017 | SR-002
2026-08-05 17:43:51 +02:00

272 lines
18 KiB
C++
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
#pragma once
#include "inference/backend_config.hpp"
#include <string>
inline const std::string kDefaultDetectorModel = std::string(SAE_MODELS_DIR) + "/scrfd_500m_bnkps.onnx";
inline const std::string kDefaultArcfaceModel = std::string(SAE_MODELS_DIR) + "/LVFace-B_Glint360K.onnx";
inline const std::string kDefaultSceneModel = std::string(SAE_MODELS_DIR) + "/transnetv2.onnx";
enum class Verbosity {
minimal, // actor names + merged time windows only
standard, // per-frame detail: bbox, similarity, unknowns logged
xray, // Jellyfin-Xray format: {"second": ["Actor", ...], ...}
};
// debug verbosity = compile with -DSAE_DEBUG → scene_analyze_debug binary
struct Config {
// ── Input ─────────────────────────────────────────────────────────────────
std::string movie_path;
std::string gallery_path;
// TRACES: IR-002 | SR-003
// "global" (matched against the whole library) or "limited" (this title's
// credited cast only). The strongest single quality signal when two
// manifests compete for the same cut: identical gallery_size can mean very
// different recall depending on which was used.
std::string gallery_scope{"global"}; // gallery.json produced by build_gallery
// ── Output ───────────────────────────────────────────────────────────────
std::string output_path; // annotations.json
Verbosity verbosity{Verbosity::minimal};
// When set, tee the embedder output to an HDF5 dump (schema:
// scripts/optimizer/SCHEMA.md) for offline threshold-sweep replay via sae_kpn.
std::string dump_embeddings_path;
/// TRACES: VR-015 | PR-004
// When set, write a per-node timing and bottleneck report here (src/
// benchmark.hpp) and print it at shutdown. Costs one background thread
// reading relaxed atomics on a timer, so it is safe to leave on, but a
// measurement run should still be isolated (nothing else on the GPU).
std::string benchmark_path;
int benchmark_interval_ms{100}; // channel-occupancy sampling period
// ── Sampling ─────────────────────────────────────────────────────────────
float sample_fps{1.0f}; // frames to analyse per second of movie
float max_decode_fps{0.f}; // wall-clock cap on source decode rate (0 = uncapped)
double start_sec{0.0}; // seek to this timestamp before sampling
double end_sec{-1.0}; // stop at this timestamp (-1 = end of file)
// ── Detection (SCRFD-500MF via cv::dnn::Net) ──────────────────────────────
std::string detector_model;
std::string detector_engine; // optional path to pre-built TRT engine; bypasses ORT
// TRACES: AR-003 | SR-002
// 0 = no cap, the default. A fixed cap discards the SMALLEST faces first,
// which are exactly the background cast X-Ray still credits with scene
// membership. Per-frame cost is contained by backpressure (AR-004) rather
// than by throwing work away. Set >0 only to bound a pathological source.
int max_faces{0};
float min_face_px{40.f}; // discard detections narrower or shorter than this
float detector_conf{0.5f};
float detector_nms{0.4f};
/// TRACES: GR-004 | SR-001
// Gallery ↔ embedder binding. A gallery built with a different model than the
// one loaded here is a hard error, always. This flag additionally promotes
// "cannot prove they match" (unstamped legacy gallery, or a name-only match
// because the ONNX could not be hashed) from a loud warning to a hard error.
// Also settable via SAE_REQUIRE_GALLERY_STAMP=1. Measurement runs want it on.
bool require_gallery_stamp{false}; // --require-gallery-stamp
// ── Recognition (ArcFace ONNX) ────────────────────────────────────────────
std::string arcface_model;
std::string arcface_engine; // optional path to a pre-built TRT engine; bypasses ORT
int embed_batch_size{4}; // max faces per ORT Run() call — bounds per-call latency
float match_prior{0.5f}; // base-rate prior; 0.5 = use calibrated sigmoid directly
// Tuned by Differential Evolution against Amazon X-Ray per-second presence
// over the 4-film rep4 matrix. Best model+mode: LVFace-B_Glint360K, full
// gallery, expansion on. Supersedes an earlier 9-film scene-union tuning
// (0.76); that metric hid out-of-cast false positives.
//
// **Read the provenance before trusting the value.** Two things about it:
//
// 1. The document it came from no longer exists under that name. It was
// docs/rep4-optimizer-results.md, renamed to docs/model-bakeoff.md and
// then rewritten (0bd2747). This comment pointed at the dead path for
// long enough that the number looked unsourced. The original is still
// readable at `git show d340da7:docs/rep4-optimizer-results.md`, where
// the shipped triple appears as
// `prob_threshold=0.754, anneal_sec=35.5`.
//
// 2. **0.754 predates a scoring bug fix and was never re-derived.** That
// same rewrite reports finding "a real scoring bug in optimize.py: a
// candidate whose hardest film's replay timed out was averaged over
// survivors instead of penalized, silently rewarding partial coverage.
// Affected 3 of 16 training combos". The corrected sweep converged
// somewhere else — the surviving document records anneal_sec=59.2,
// extinction_sec=59.2 against the 35.5/57.4 shipped alongside this
// threshold — and no corrected prob_threshold is recorded anywhere.
// (The other two constants are now withdrawn outright, which is why
// only this one still matters.)
//
// The doc is also candid that the optimum "generalizes unevenly — strong on
// 3 of 5 held-out films, badly broken on 2 (one with a 974-count misID
// blowup)", and that it is shipped anyway because it still beats the old
// defaults on average. That is a defensible call and not a settled,
// film-agnostic optimum; it should be visible here rather than only in a
// document this comment used to point at incorrectly.
float prob_threshold{0.754f}; // posterior P(match | sim, prior) threshold
// TRACES: AR-024 | SR-002
// match_threshold (0.45), match_ratio (0.80) and match_ratio_ceil (0.65) are
// RETIRED, joining track_max_embed_dist, cut_revive_sim, expand_novelty_sim
// and expand_track_spread_max. All were raw cosine distances, and they were
// the accept rule whenever the calibration fit failed — so the one situation
// in which the pipeline knew its probabilities were untrustworthy was the
// one in which it stopped using them. An unfitted sigmoid is now the
// fallback everywhere, which is at least the same wrong number in every
// stage. See identity_matcher_node.hpp.
// ── Cut detection ────────────────────────────────────────────────────────
float cut_threshold{0.70f}; // grayscale histogram correlation below this → hard cut
// ── Scene detection (TransNetV2, opt-in) ─────────────────────────────────
// When enabled, the source decodes densely (native FPS) and a decimator
// splits the stream: full-res 1-FPS frames to the face pipeline, and a
// downscaled dense stream to the TransNetV2 scene detector. Shot boundaries
// it finds are surfaced as Frame::is_scene_boundary. This is separate from
// the always-on histogram cut, which flags intra-scene camera-angle changes.
bool scene_detect{false}; // master switch (--scene-detect)
std::string scene_model; // TransNetV2 .onnx (default set in main)
std::string scene_engine; // optional pre-built TRT .engine; bypasses ORT
float scene_threshold{0.60f}; // sigmoid boundary prob above this → boundary
// (this export's non-boundary baseline sits
// at ~0.50; real boundaries spike to ~0.7+)
int scene_stride{50}; // frames advanced between windows (≤ kWindow)
// Dense-decode knobs (only active with scene_detect). Dense decode of every
// native-rate frame is the pipeline's cost driver, which is what made the
// temporal shortcut below tempting.
/// TRACES: AR-011 | SR-002
// scene_decode_fps: rate the source decodes at in dense mode.
// **0 = native, and native is the only correct setting.** kWindow is 100
// frames: at native 25 fps that window spans ~4 s, which is what
// TransNetV2 was trained on; at the 12 fps this used to default to it
// spans ~8.3 s, so the model saw half-speed motion over twice its
// temporal context. Boundary *timestamps* stay right either way — which
// is exactly why the degradation was invisible, and why the compressed
// separation it produced (~0.50 baseline against ~0.7+ peaks) was read
// as a property of the export rather than of the input. Lowering this
// buys decode time by running the model off-distribution; reach for
// dense_scale or scene_stride instead, which do not.
// dense_scale: downscale factor applied to decoded frames in dense mode
// (0<f≤1; e.g. 0.5 = half size). Cheaper sws_scale + smaller frames
// through the fanout. A spatial reduction, and TransNetV2 downsamples to
// 48×27 regardless, so unlike the above it is a documented, understood
// degradation. NOTE: also shrinks what the face detector sees — keep
// ≥0.5 on 1080p sources so SCRFD still resolves small faces. 1 = off.
float scene_decode_fps{0.f}; // dense decode rate (0 = native)
float dense_scale{1.0f}; // dense-mode frame downscale (1 = off)
// ── Face tracking (frame-to-frame) ───────────────────────────────────────
/// TRACES: AR-007, AR-008, AR-024 | SR-002
// track_alpha is the *base* weight, used on ordinary frames. It is
// frame-dependent (AR-007): on is_cut / is_scene_boundary, and for any track
// that is no longer on screen, it drops to 0 (embedding only), because
// position carries no information across a viewpoint change or a gap.
float track_alpha{0.4f}; // base cost weight: 0=embedding only, 1=spatial only
float track_min_iou{0.1f}; // IoU below which spatial link alone is rejected
// Minimum P(same person) for an association to be admissible on appearance
// alone. This replaces track_max_embed_dist (a raw cosine distance, AR-024).
// 0.5 is not a tuned constant: it is the decision boundary. Below it the pair
// is more likely two people than one, and no amount of IoU makes that a link
// worth asserting on identity grounds.
float track_assoc_min_prob{0.5f};
// How long a track that has gone off screen stays available for association
// before the registry reaps it and emits its presence claim (AR-013).
// Replaces track_max_frames_missing: a frame count silently changed meaning
// with sample_fps, and the same number had to be guessed twice (once for an
// ordinary miss, once for a cut). Seconds mean one thing at any sample rate.
double track_extinction_sec{5.0};
// ── Ownership and evidence accumulation (AR-025) ──────────────────────────
// TRACES: AR-025, AR-017 | SR-002
// These four decided how presence is claimed and were unreachable: they
// lived as in-class initialisers on TrackRegistry::Config and
// EvidenceDiscounter::Config, and main constructed the discounter with the
// one-argument constructor, so nothing short of a recompile could move
// them. rho_max's own comment defers to "the sweep (VR-007)" for where it
// belongs — a sweep that could not reach it.
//
// ownership_logodds is arguably the most consequential constant in the
// pipeline after prob_threshold: below it a track produces no presence
// claim at all, so it decides whether an actor is reported rather than how
// confidently. 2.0 is a posterior of ~0.88. Unswept.
float ownership_logodds{2.0f};
// How much a single observation may move a track's belief. n_eff =
// n / (1 + (n-1)·rho), so rho_max caps what a repeated view can ever be
// worth: 0.5 caps it at two independent observations however long the shot
// runs. It is deliberately below 1 — a held pose still yields a fresh
// detection, alignment and noise realisation, so a little independent
// evidence survives. Setting it to 1 freezes belief after the first frame,
// which is the bug this replaced.
float evidence_rho_max{0.5f};
// P(same view) below this and the observation counts as a genuinely new
// look, so it joins the per-track view set.
float evidence_admit_below{0.6f};
// Distinct views remembered per track, which bounds the novelty comparison.
int evidence_max_views{8};
// ── Scene tracking ────────────────────────────────────────────────────────
// TRACES: AR-012, AR-013 | SR-002
// extinction_sec (57.4) and anneal_sec (35.5) are GONE, along with
// SceneTrackerFunc, which is what read the first of them. docs/SPEC.md
// specified this removal and ended it "grep for both names and expect no
// survivors"; there were about forty, and the register meanwhile recorded
// both as Withdrawn and "deleted rather than retained at zero" on the
// grounds that a field naming a mechanism the pipeline no longer has is
// actively misleading.
//
// Both existed to bridge gaps between isolated accepted frames. A track
// that survives its own gaps leaves them nothing to do: AR-012 makes a
// window the extent of a track an actor owns, and AR-013 ends it at the
// last sighting. The keep-alive answered the same question again and
// answered it worse, by re-opening exactly the trailing cool-down AR-013
// refuses.
//
// track_extinction_sec above is NOT the same knob under a new name. It
// bounds how long a lost track stays available for re-association, which is
// a tracking question; it never extends a presence claim.
// ── Per-film gallery expansion ────────────────────────────────────────────
// Within one uncut track every face is the same physical person — a free
// same-identity label the baked gallery lacks. When a track is confidently
// owned by an actor, its gallery-far (pose-varied) embeddings are validated
// new reference views; they are promoted into a per-film, in-memory annex so
// later frames/tracks of that actor at similar poses recognise. See
// gallery/track_gallery.hpp.
// Default ON: the rep4 matrix (docs/model-bakeoff.md, "Two effects in
// isolation") found expansion helps recall on the full (unrestricted)
// gallery for the winning model/mode — the opposite of the earlier
// assumption that it only helps restricted galleries. The same section is
// explicit that on the full gallery it buys +2.1pp F1 and +3.9pp recall
// "at a real cost" in misIDs, where in restricted mode it is a clean win.
bool expand_gallery{true}; // master switch
int expand_buffer_size{20}; // per-track diversity buffer capacity
// TRACES: AR-018, AR-024 | SR-005
// Banded admission for the per-subject store, in PROBABILITY space. An
// embedding joins only if P(same person) against something already stored
// lands inside [lo, hi]: above hi it is redundant, below lo it is evidence
// the track is not one person. The same lo is re-applied to the whole store
// at promotion time — see track_gallery.hpp. This is the only threshold the
// expansion path has: it replaces the raw-cosine expand_novelty_sim (0.55)
// and expand_track_spread_max (0.60), which are retired (AR-024).
// Working values pending VR-007; sweep both bounds, they fail in opposite
// directions.
float expand_band_lo{0.90f};
float expand_band_hi{0.95f};
int expand_min_anchor_frames{3}; // require ≥N accepted frames naming the actor before
// the track is confirmed and its buffer promoted
std::string expand_debug_dir; // if set, dump promoted mugshots + embeddings here
// ── Inference backend tuning ────────────────────────────────────────────
// Consumed by the compiled-in inference backend (ORT or TRT).
// INT8 is unsafe for ArcFace without a calibration table.
BackendConfig trt{}; // fp16=true, int8=false, cache_dir="./trt_cache"
// ── Debug output (only used when SAE_DEBUG is defined) ───────────────────
#ifdef SAE_DEBUG
std::string debug_dir{"debug_frames"};
float crop_context{1.5f}; // bbox expansion factor for context crop
#endif
};