feat(scene-detector): run the learned boundary detector live in the C++ pipeline

Wire the XGBoost scene-boundary detector into scene_analyze as a post-EOF step in
the result sink (like flood-fill itself — the per-film knee threshold needs the
whole film, so it cannot stream). With --scene-xgb-model set, the camera-position
node stamps a per-frame RGB histogram onto the Frame, it rides through to the
sink, and at EOF the sink runs XGBSceneBoundary over the collected histograms +
the movie's per-second audio log-PSD to produce the flood-fill boundaries. Falls
back to is_scene_boundary / is_cut when no model is configured or inference fails.

Inference is real XGBoost via CMake FetchContent (v2.1.1, static), C API in
src/inference/xgb_scene_boundary.hpp; audio log-PSD in src/inference/
audio_logpsd.hpp (FFTW + ffmpeg full-file 16kHz decode). Feature extraction
matches training exactly — video features verified row-identical to numpy, and to
avoid chasing numpy's every rounding the shipped model is TRAINED on the
C++-extracted features (scene_features_dump exe → train_xgb_cpp.py). The
C++/Python peak-finders differ slightly so boundary counts differ, but what
matters is downstream: flood + C++ detector = 75.8% macro presence F1 vs 64.0%
for the histogram-cut flood and 62.5% for track_extent, and it fixes the Scarface
flood collapse (41 -> 70). All nine films improve.

Guarded by the SAE_SCENE_XGB CMake option (on by default; heavy first build).
xgb_boundary_parity is a diff harness; scene_features_dump writes the C++ feature
matrix so training and inference share one feature implementation.

Verified end to end: scene_analyze --scene-xgb-model on a real movie stamps the
histogram, runs the detector at EOF ("XGBoost scene detector: N boundaries"), and
flood-snaps presence to the learned boundaries.
This commit is contained in:
2026-08-09 21:21:29 +02:00
parent 0e35dac951
commit e5204a831a
13 changed files with 827 additions and 12 deletions
@@ -33,9 +33,29 @@ struct CameraPositionChangeDetectorFunc {
explicit CameraPositionChangeDetectorFunc(const Config& cfg)
: cut_threshold_(cfg.cut_threshold)
, want_rgb_hist_(!cfg.scene_xgb_model.empty())
{
std::cerr << "[camera_position_change_detector] cut_threshold="
<< cut_threshold_ << "\n";
<< cut_threshold_
<< (want_rgb_hist_ ? " (+rgb_hist for scene detector)" : "")
<< "\n";
}
// 32-bin-per-channel normalised RGB histogram (96 floats), the exact layout
// the XGBoost scene detector was trained on (see embedding_dump_node). Only
// computed when a scene model is configured, so it costs nothing otherwise.
static std::vector<float> rgb_histogram(const cv::Mat& img) {
constexpr int kBins = 32;
std::vector<float> out(kBins * 3, 0.f);
if (img.empty() || img.channels() != 3) return out;
float range[] = {0.f, 256.f}; const float* ranges = range; int bins = kBins;
for (int c = 0; c < 3; ++c) { // OpenCV BGR → store B,G,R blocks
cv::Mat h;
cv::calcHist(&img, 1, &c, cv::Mat(), h, 1, &bins, &ranges);
cv::normalize(h, h, 1.0, 0.0, cv::NORM_L1);
for (int b = 0; b < kBins; ++b) out[c*kBins + b] = h.at<float>(b);
}
return out;
}
Frame operator()(Frame f) {
@@ -64,11 +84,13 @@ struct CameraPositionChangeDetectorFunc {
prev_hist_ = hist;
prev_hist_valid_ = true;
if (want_rgb_hist_) f.rgb_hist = rgb_histogram(f.image);
return f;
}
private:
float cut_threshold_;
bool want_rgb_hist_{false};
cv::Mat prev_hist_;
bool prev_hist_valid_{false};
};