Files
scene-actor-extraction/src/nodes/scene_detector_node.hpp
T
dtourolle 88c42573a5 fix(pipeline): finish three changes that had only been half applied
Each of these was recorded as done and was done in one place out of two.

AR-011 -- the TransNetV2 dedup window. The derived window
(dedup_window_sec, median observed interval halved) reached scenes.json
and nothing else. SceneBoundaries, the path that actually feeds
is_scene_boundary to the tracker, kept the literal 0.04 s under a
comment claiming it "matches the dedup scenes.json applies, so the two
views agree". They did not agree. 0.04 is one frame at 25 fps and wider
than a frame at 30, so two cuts on consecutive frames merged into one
and the loss was invisible: the pipeline simply saw fewer boundaries.
The detector now supplies the window it derived.

AR-019 -- ownership. The register says ownership "comes from the
registry, not a second local tally". Both existed: promotion fired on a
local accepted-frame count and fell back to a local per-actor plurality
when the registry had not yet claimed the track. That fallback was
reachable in the live pipeline, not just in tests -- three accepted
frames arrive well before a posterior crosses the ownership threshold --
so in practice the plurality usually decided, and it could not see the
AR-025 correlation discounting it was meant to defer to. The tally is
gone; promotion now requires the registry's verdict, with the accepted
-frame count demoted to an explicit evidence floor.

AR-017 -- the route. DeadTrack carried belief but no route, and the sink
wrote the literal string "live", so a field the schema publishes could
not distinguish anything. AR-017's own verification asks for "deferred
and pooled routes distinguishable". Route is now an enum on the claim.
Only `live` occurs today; `deferred` exists so AR-020's pass has
somewhere to write instead of a serialisation change to make.

Also: TrackGallery::forget had no callers, under a comment asserting the
matcher called it "on a cut or track disappearance". The cut half was
true by another route; the disappearance half was not, so a track that
died quietly kept its diversity buffer until the next cut cleared
everything. Replaced with prune_dead against the registry's own
liveness, the same shape as the tracker's prune_boxes -- a second
opinion about which tracks exist is a second thing that can be wrong.

Removes dead logistic/logit helpers and fixes five TRACES tags that used
a comma where a pipe separates requirement types, which the gate had
been reporting as diagnostics.

TRACES: AR-011, AR-017, AR-019 | IR-002 | SR-002, SR-005
2026-08-05 17:33:12 +02:00

272 lines
12 KiB
C++

#pragma once
#include "types.hpp"
#include "config.hpp"
#include "scene_boundaries.hpp"
#include <memory>
#include "inference/scene_detector.hpp"
#include <nlohmann/json.hpp>
#include <algorithm>
#include <atomic>
#include <deque>
#include <fstream>
#include <iostream>
#include <string>
#include <vector>
// ── SceneDetectorFunc ─────────────────────────────────────────────────────────
// KPN sink node: TransNetV2 shot-boundary detection on the dense frame stream.
//
// Buffers incoming (dense, native-rate) Frames into a rolling window of
// ISceneDetector::kWindow (=100) frames. Every `stride` frames it runs one
// inference and reads back per-frame boundary probabilities, but only trusts the
// central region of each window — TransNetV2 (like most sliding-window boundary
// models) is unreliable near the window edges where it lacks temporal context.
// Overlapping windows by (kWindow - stride) frames means every frame is scored
// from at least one window's trusted centre.
//
// Boundaries (prob > scene_threshold, local maxima) are collected with their
// timestamps and written to scenes.json alongside the main annotations output on
// EOF. This branch is terminal: it produces no pipeline messages, only a file.
struct SceneDetectorFunc {
static constexpr std::string_view label() { return "scene_detector"; }
/// TRACES: AR-010 | SR-002
/// Publish each window's verdict as it is scored, so the face branch — held
/// back by channel depth — can consult it for frames it has not reached yet.
void set_boundaries(std::shared_ptr<SceneBoundaries> b) { shared_ = std::move(b); }
SceneDetectorFunc(const Config& cfg, std::atomic<bool>& done)
: detector_(make_scene_detector(cfg))
, threshold_(cfg.scene_threshold)
, stride_(std::clamp(cfg.scene_stride, 1, ISceneDetector::kWindow))
, output_path_(scenes_path(cfg.output_path))
, movie_path_(cfg.movie_path)
, done_(done)
{
// Trusted centre half of each window. Frames outside [guard, kWindow-guard)
// are re-scored by an adjacent window, so we ignore them here to avoid
// edge artefacts and double-counting.
guard_ = (ISceneDetector::kWindow - stride_) / 2;
std::cerr << "[scene_detector] threshold=" << threshold_
<< " stride=" << stride_
<< " guard=" << guard_
<< " output=" << output_path_ << "\n";
}
void operator()(Frame f) {
if (f.eof) {
flush_remaining();
// Release anyone waiting on the join: the tail frames after the last
// full window will never be covered, so waiting for them would hang.
if (shared_) shared_->finish();
write_output();
done_.store(true, std::memory_order_release);
return;
}
/// TRACES: AR-011 | SR-002
// Learn the cadence of the stream from the stream itself, rather than
// assuming one. See dedup_window_sec().
if (prev_ts_ >= 0.0 && intervals_.size() < kCadenceSamples) {
const double dt = f.timestamp_sec - prev_ts_;
if (dt > 0.0) intervals_.push_back(dt);
}
prev_ts_ = f.timestamp_sec;
images_.push_back(f.image);
times_.push_back(f.timestamp_sec);
// Once we have a full window, score it and slide forward by `stride`.
while (static_cast<int>(images_.size()) >= ISceneDetector::kWindow) {
score_window();
for (int i = 0; i < stride_; ++i) {
images_.pop_front();
times_.pop_front();
}
window_base_ += stride_;
}
}
/// TRACES: AR-011 | SR-002
// How close two boundaries have to be before they are the same boundary,
// derived from the cadence the detector was actually fed.
//
// What this replaces is a literal 0.04 s — one frame at 25 fps, and silently
// wrong at any other rate. On a 30 fps source it spans more than a frame, so
// two cuts on consecutive frames merge into one and a real boundary is lost;
// the output does not show this, it simply contains fewer cuts. Assuming a
// frame rate is the same class of mistake as feeding a model the wrong rate,
// which is why this belongs to AR-011 and not to a tidy-up.
//
// Half a frame, not a whole one, because the only thing being deduplicated is
// one frame scored by two overlapping windows — a gap of zero. Two distinct
// frames are a full interval apart and must both survive. Half an interval
// separates those two cases without putting the decision on the knife-edge
// where floating-point error settles it.
//
// Median, not mean: a seek, or a gap where the decoder dropped a frame,
// contributes one long interval that would drag a mean and cannot move a
// median.
static double dedup_window_sec(std::vector<double> intervals) {
if (intervals.empty()) return 0.0; // <2 frames: nothing to deduplicate
const std::size_t mid = intervals.size() / 2;
std::nth_element(intervals.begin(), intervals.begin() + mid,
intervals.end());
return intervals[mid] * 0.5;
}
private:
// Run TransNetV2 on the leading kWindow frames of the buffer and record any
// boundaries found within the trusted centre region.
void score_window() {
std::vector<cv::Mat> win(images_.begin(),
images_.begin() + ISceneDetector::kWindow);
std::vector<float> probs = detector_->detect_window(win);
// On the very first window there is no preceding window, so trust from 0;
// otherwise skip the leading guard already covered by the previous window.
const int lo = (window_base_ == 0) ? 0 : guard_;
const int hi = ISceneDetector::kWindow - guard_;
std::vector<double> fresh;
for (int i = lo; i < hi; ++i) {
if (probs[i] <= threshold_) continue;
// Local maximum → the boundary frame (avoid a run of high scores
// registering as several adjacent cuts).
const bool peak =
(i == 0 || probs[i] >= probs[i-1]) &&
(i == kLast_() || probs[i] >= probs[i+1]);
if (peak) {
boundaries_.push_back({times_[i], probs[i]});
fresh.push_back(times_[i]);
}
}
/// TRACES: AR-010 | SR-002
// Publish with a watermark: everything up to times_[hi-1] now has a
// final verdict. The face branch consults this for frames it has not
// reached yet, and the watermark is what lets it tell "no boundary
// here" from "not scored yet".
/// TRACES: AR-011 | SR-002
// Hand the join the same dedup window scenes.json uses, derived from the
// observed cadence rather than assumed. Set on every window because the
// median refines as intervals accumulate; it converges within the first
// window and costs a double assignment thereafter.
if (shared_) {
shared_->set_merge_window(dedup_window_sec(intervals_));
if (hi > lo) shared_->publish(fresh, times_[hi - 1]);
}
}
// At EOF the tail (< kWindow frames) never formed a full window. Pad it out
// to kWindow by repeating the last frame so the final real frames still get
// scored, then take only the region past what earlier windows covered.
void flush_remaining() {
const int n = static_cast<int>(images_.size());
if (n == 0) return;
std::vector<cv::Mat> win(images_.begin(), images_.end());
cv::Mat last = win.back();
while (static_cast<int>(win.size()) < ISceneDetector::kWindow)
win.push_back(last);
std::vector<float> probs = detector_->detect_window(win);
const int lo = (window_base_ == 0) ? 0 : guard_;
std::vector<double> fresh;
for (int i = lo; i < n; ++i) { // only real (non-padded) frames
if (probs[i] <= threshold_) continue;
const bool peak =
(i == 0 || probs[i] >= probs[i-1]) &&
(i == n - 1 || probs[i] >= probs[i+1]);
if (peak) {
boundaries_.push_back({times_[i], probs[i]});
fresh.push_back(times_[i]);
}
}
/// TRACES: AR-010 | SR-002
// Publish the tail too. Without this the final frames — everything after
// the last full window — reach the join with no verdict and are treated
// as boundary-free without evidence, which is precisely the ambiguity
// the watermark exists to prevent.
if (shared_ && n > 0) {
shared_->set_merge_window(dedup_window_sec(intervals_));
shared_->publish(fresh, times_[n - 1]);
}
}
void write_output() {
if (written_) return;
written_ = true;
// Merge boundaries closer than one frame apart (dedup across window seams).
const double dedup_sec = dedup_window_sec(intervals_);
std::sort(boundaries_.begin(), boundaries_.end(),
[](const Boundary& a, const Boundary& b) {
return a.t < b.t;
});
nlohmann::json root;
root["schema_version"] = 1;
root["movie"] = movie_path_;
root["model"] = "transnetv2";
root["threshold"] = threshold_;
nlohmann::json cuts = nlohmann::json::array();
double last_t = -1e9;
for (const auto& b : boundaries_) {
if (b.t - last_t < dedup_sec) continue;
cuts.push_back({{"t", b.t}, {"probability", b.prob}});
last_t = b.t;
}
root["cuts"] = std::move(cuts);
std::ofstream f(output_path_);
if (!f.is_open()) {
std::cerr << "\n[scene_detector] ERROR: cannot write "
<< output_path_ << "\n";
return;
}
f << root.dump(2) << "\n";
// Report the derived cadence: VR-006 re-tunes scene_threshold against it,
// and a rate that is not the source's is the first thing to suspect.
std::cerr << "\n[scene_detector] wrote " << root["cuts"].size()
<< " boundaries → " << output_path_
<< " (dedup=" << dedup_sec << "s from "
<< (dedup_sec > 0.0 ? 0.5 / dedup_sec : 0.0) << " fps)\n";
}
static int kLast_() { return ISceneDetector::kWindow - 1; }
// annotations.json → annotations.scenes.json (or scenes.json for bare names)
static std::string scenes_path(const std::string& out) {
auto dot = out.find_last_of('.');
if (dot == std::string::npos) return out + ".scenes.json";
return out.substr(0, dot) + ".scenes.json";
}
struct Boundary { double t; float prob; };
// Enough to establish a rate; bounded so a feature-length film does not
// accumulate one double per frame for a number that stops moving early.
static constexpr std::size_t kCadenceSamples = 512;
std::unique_ptr<ISceneDetector> detector_;
float threshold_;
int stride_;
int guard_{0};
std::string output_path_;
std::string movie_path_;
std::atomic<bool>& done_;
std::deque<cv::Mat> images_;
std::deque<double> times_;
int64_t window_base_{0}; // frame index of images_.front()
std::vector<Boundary> boundaries_;
double prev_ts_{-1.0}; // AR-011: cadence, learned not assumed
std::vector<double> intervals_;
bool written_{false};
std::shared_ptr<SceneBoundaries> shared_; ///< AR-010 join point
};