Files
scene-actor-extraction/docs/scene-boundary-detector.md
T
dtourolle 0feafec7c9 feat(review): GT-aware TP/FP/FN frame annotation + scene-detector examples
dump_error_frames.py drew every identified box green, so a false positive looked
like a true positive and a missed cast member was invisible. Make the annotation
ground-truth aware, matching what the per-second scorer classifies:
  - GREEN  true positive  — a name X-Ray also credits to this scene
  - RED    false positive — a name X-Ray does NOT credit here (the real error)
  - ORANGE unknown detection
  - BLUE   a text panel listing X-Ray cast present with no detected face (the
           structural false-negatives — no box exists to draw)

Add two representative annotated frames to the scene-detector page: a clean
green-TP second, and the face-vs-scene-cast case (a red FP lead + six off-camera
cast in blue) that makes the recall ceiling visual. Frames are generated by the
script from replay.py --raw-out output; the two committed examples are hand-picked
doc assets (bulk experiments/dump_review is regenerable and gitignored).
2026-08-09 22:28:21 +02:00

9.5 KiB
Raw Blame History

The learned scene-boundary detector

Presence uses flood-fill: an actor seen once inside a shot is reported for the whole shot ([prev_boundary, next_boundary]). That only works if the boundaries are good. This page is the story of getting them good — a learned scene-boundary detector that lifts per-second actor-presence F1 from 62.6% to 74.9% across the nine-film X-Ray benchmark, and fixes the film where naive flood-fill was actively harmful.

That 74.9% is the leave-one-out figure: each film is scored by a detector trained on the other eight, so no film's presence is measured with a detector that ever saw it. It is the honest generalisation number, and it is only ~1 point below the all-nine-trained model (75.8%) — the detector barely overfits.

Why the old cut detector wasn't enough

The always-on boundary source was the grayscale histogram-correlation cut detector (camera_position_change_detector): mark a cut when the frame-to-frame grayscale histogram correlation drops below 0.70. It is cheap and it fires on obvious hard cuts, but on a low-contrast, uniformly-graded film it is nearly blind. On Scarface it fired once in 10,204 frames. Flood-fill then snapped every actor across essentially the whole film:

Scarface precision recall
flood + grayscale cuts 26% 95%
track-extent (no flood) 92% 45%

That single failure is what motivated everything below: flood-fill needs a boundary source that works regardless of grade.

What we are detecting, and why it is hard

The training target is Amazon X-Ray scene boundaries (scenes.csv). These are narrative scenes — a new location or beat in the story — not shot cuts. There are only ~2060 of them per film (median scene ~170 s), and many transition within continuous visual style and continuous audio. So the signal is sparse and often genuinely faint: a boundary detector working from audio-visual features can never recall a narrative cut that has no audio-visual signature.

This shapes every result: absolute boundary-F1 is modest by construction. What matters is the downstream number — does snapping flood-fill to these boundaries name the right actors — and there the gain is large.

The features (what worked, measured)

Everything is per second, aligned to the 1-fps presence grid.

  • Delta histograms, not raw histograms. The raw RGB histogram encodes what a frame looks like, not that it changed — measured boundary separability ~1.4×. The symmetric histogram delta |hist(t+k) hist(tk)| separates boundaries 45×. Leading with deltas (k = 1,2,4,8 s) and dropping the raw histogram was the single biggest feature win (LSTM F1 7.5% → 10.8%).
  • A multi-scale "ramp" bank. Antisymmetric matched filters at half-widths H = 2,4,6,8,10 s; the model weights the scales. Different films' boundaries peak at different widths.
  • A time-since-last-boundary "debounce" clock, scaled by the corpus mean scene length (~205 s), encoding that scenes don't restart moments apart.
  • Audio log-PSD (per-second, 4 s window, ~57 log-frequency bins). Measured weak on its own — a standalone audio cutter scored only 36% held-out F1, because narrative boundaries usually have continuous audio — but it is complementary on the films where video is weak (Downton, Sound of Metal), so it is included and the model uses it where it helps.

Detector development: features and model

Dead ends, all measured and discarded: audio-only detection; raw histograms/PSDs as input; a two-tower BiLSTM (no better than the tree, far slower); larger FFT windows / more frequency bins (worse — boundaries are short events); and TransNetV2 (a Conv3D net that will not co-reside with the ROCm/VAAPI stack).

The model

  • XGBoost regressor over a ±3 s window of the features above, predicting a soft Gaussian proximity-to-boundary target (exp(-(d/σ)²), σ = 10 s). Regression to a soft target — rather than a hard 0/1 label — stops a near-miss from being trained as a hard negative, and yields a smooth score whose peaks are the boundaries.
  • Per-film knee threshold. The predicted peak heights form a convex-decreasing curve; the knee (max drop below the endpoints' chord) is where real boundaries give way to noise. Selecting at the knee self-calibrates the boundary count to roughly the true scene count, per film, with no global threshold that would be wrong for every grade.
  • Trained on all nine films for the shipped model. Café Society and Scarface (the low-contrast grades) must be in training — held out, the model cannot generalise to them; in training they reach 7086% boundary-F1.

Boundary detection, held out (leave-one-out, ±20 s tolerance — appropriate given ~170 s scenes): ~34% F1, versus ~27% for the grayscale baseline. The absolute number is capped by the narrative-vs-audiovisual mismatch above; the point is the downstream effect.

The result that matters: actor presence

Per-second X-Ray presence F1, macro over the nine films, at the shipped presence config. The learned column is leave-one-out — each film scored by a detector trained on the other eight:

boundary source for flood-fill presence F1
track-extent (flood off) 62.6%
flood + grayscale cuts 64.0%
flood + learned detector (LOO) 74.9%

Macro presence F1 by flood-fill boundary source

+12.3 points over track-extent, +10.9 over the grayscale-cut flood, and it improves every one of the nine films — under honest leave-one-out. Per film:

Per-film presence F1 by boundary source

film track-extent flood+grayscale flood+learned (LOO)
Benny & Joon 77.3 80.2 78.2
Café Society 59.1 62.2 69.8
Downton Abbey 41.0 51.8 78.6
Lord of War 74.8 77.1 77.8
Lovelace 70.3 74.0 78.2
The Many Saints of Newark 37.5 43.9 53.4
Scarface 62.6 40.9 74.9
Sound of Metal 75.0 78.1 86.8
Valerian 65.6 67.7 76.2

The two headline cases:

  • Scarface: the grayscale-cut flood breaks it (62.6 → 40.9), because it detects one cut in the whole film. The learned detector — on a film it never trained on — takes it to 74.9%. This is the strongest evidence the detector generalises: it fixes the exact failure that motivated it, held out.
  • Downton Abbey: 41.0 (track-extent) → 51.8 (grayscale) → 78.6 — a +37-point swing on the hardest film.

Naive flood-fill barely beat doing nothing (64% vs 62%) and broke a film. With a real boundary detector, flood-fill is decisively the right mode.

What the frames look like

scripts/optimizer/dump_error_frames.py pulls representative seconds and draws each face box coloured against X-Ray's scene cast: green = true positive (a name X-Ray also credits to this scene), red = false positive (a name X-Ray does not credit here — the real error), orange = an unknown detection. Cast X-Ray lists as present but for whom no face was detected — the structural false-negatives a face pipeline can never box — are listed as a blue panel.

A correctly identified second: green true-positive boxes

Above: three faces named correctly (green). Below: the face-vs-scene-cast tension made visual — the one visible face is confidently named (here it is a red false-positive, a lead X-Ray did not credit to this exact scene), while six credited cast members are off-camera with no face to detect (blue). This is why recall against X-Ray has a structural ceiling, not a fixable bug.

A false-positive box (red) with off-screen cast listed (blue)

In the pipeline

Boundary detection is a post-EOF step, like flood-fill itself: the per-film knee needs every peak, so it can only run once the whole film is seen. The camera_position_change_detector stamps a per-frame RGB histogram onto each frame; it rides through to the result sink; at end-of-stream the sink runs the detector over the collected histograms plus the movie's audio log-PSD and snaps the presence windows to the result. Enable it with:

scene_analyze --movie <file> --gallery <gallery.h5> \
    --scene-xgb-model models/scene_boundary_xgb.json

Inference is real XGBoost, built into the binary via CMake (SAE_SCENE_XGB); the audio log-PSD uses FFTW + the existing FFmpeg decode. To keep training and inference on one feature implementation, the shipped model is trained on the C++-extracted features (scene_features_dumptrain_xgb_cpp.py) rather than a re-implementation in Python — parity by construction. Verified end to end through scene_analyze on a movie file and through the Jellyfin work-queue worker.

Reproduce

# per-second audio log-PSD for each film
.venv-rocm/bin/python scripts/scene_detector/extract_audio_features.py \
    --manifest experiments/manifests/films_LVFace_opencv5.json

# C++ feature matrices (same features training and inference share)
build/scene_features_dump <dump.h5> <movie> <features.h5>

# train the shipped model on all nine films
.venv-rocm/bin/python scripts/scene_detector/train_xgb_cpp.py --train-all

# downstream A/B (track-extent vs flood+grayscale vs flood+learned)
scripts/scene_detector/downstream_presence.py