Files
scene-actor-extraction/docs/scene-boundary-detector.md
T
dtourolle edf19ab798 docs(scene-detector): document the learned scene-boundary detector
New docs/scene-boundary-detector.md: why the grayscale cut detector wasn't
enough (Scarface: 1 cut in 10k frames → flood-fill P=26%), what X-Ray boundaries
are and why they're hard, the feature/model design (delta histograms, multi-scale
ramp bank, scene-length debounce, soft-target XGBoost regressor, per-film knee),
and the measured dead ends (audio-only, raw features, LSTM, TransNetV2).

Headline result, honest leave-one-out (each film scored by a detector trained on
the other eight): flood + learned detector = 74.9% macro presence F1, vs 64.0%
for grayscale-cut flood and 62.6% for track-extent — +12.3pp, improving all nine
films. Fixes the Scarface flood collapse (grayscale 40.9 → learned 74.9, on a
film the detector never trained on) and swings Downton +37pp.

Figures are generated by scripts/scene_detector/make_figures.py from the saved
results (experiments/results/scene_boundary/downstream_loo.json); the PNGs
themselves follow the repo convention of not committing regenerable chart assets.
Added to the mkdocs nav.
2026-08-09 22:15:32 +02:00

173 lines
8.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# The learned scene-boundary detector
Presence uses **flood-fill**: an actor seen once inside a shot is reported for the
whole shot (`[prev_boundary, next_boundary]`). That only works if the boundaries
are good. This page is the story of getting them good — a learned scene-boundary
detector that lifts per-second actor-presence F1 from **62.6% to 74.9%** across
the nine-film X-Ray benchmark, and fixes the film where naive flood-fill was
actively harmful.
That 74.9% is the **leave-one-out** figure: each film is scored by a detector
trained on the *other eight*, so no film's presence is measured with a detector
that ever saw it. It is the honest generalisation number, and it is only ~1 point
below the all-nine-trained model (75.8%) — the detector barely overfits.
## Why the old cut detector wasn't enough
The always-on boundary source was the grayscale histogram-correlation cut detector
(`camera_position_change_detector`): mark a cut when the frame-to-frame grayscale
histogram correlation drops below 0.70. It is cheap and it fires on obvious hard
cuts, but on a low-contrast, uniformly-graded film it is nearly blind. On
**Scarface** it fired **once in 10,204 frames**. Flood-fill then snapped every
actor across essentially the whole film:
| Scarface | precision | recall |
| -------- | --------- | ------ |
| flood + grayscale cuts | **26%** | 95% |
| track-extent (no flood) | 92% | 45% |
That single failure is what motivated everything below: flood-fill needs a
boundary source that works regardless of grade.
## What we are detecting, and why it is hard
The training target is **Amazon X-Ray scene boundaries** (`scenes.csv`). These are
*narrative* scenes — a new location or beat in the story — not shot cuts. There
are only ~2060 of them per film (median scene ~170 s), and many transition
*within* continuous visual style and continuous audio. So the signal is sparse and
often genuinely faint: a boundary detector working from audio-visual features can
never recall a narrative cut that has no audio-visual signature.
This shapes every result: absolute boundary-F1 is modest by construction. What
matters is the **downstream** number — does snapping flood-fill to these
boundaries name the right actors — and there the gain is large.
## The features (what worked, measured)
Everything is per second, aligned to the 1-fps presence grid.
- **Delta histograms, not raw histograms.** The raw RGB histogram encodes what a
frame *looks like*, not that it *changed* — measured boundary separability ~1.4×.
The **symmetric histogram delta** `|hist(t+k) hist(tk)|` separates boundaries
**45×**. Leading with deltas (k = 1,2,4,8 s) and dropping the raw histogram was
the single biggest feature win (LSTM F1 7.5% → 10.8%).
- **A multi-scale "ramp" bank.** Antisymmetric matched filters at half-widths
H = 2,4,6,8,10 s; the model weights the scales. Different films' boundaries peak
at different widths.
- **A time-since-last-boundary "debounce" clock**, scaled by the corpus mean scene
length (~205 s), encoding that scenes don't restart moments apart.
- **Audio log-PSD** (per-second, 4 s window, ~57 log-frequency bins). Measured
weak on its own — a standalone audio cutter scored only 36% held-out F1, because
narrative boundaries usually have continuous audio — but it is complementary on
the films where video is weak (Downton, Sound of Metal), so it is included and
the model uses it where it helps.
![Detector development: features and model](assets/images/scene_detector_evolution.png)
Dead ends, all measured and discarded: audio-only detection; raw
histograms/PSDs as input; a two-tower BiLSTM (no better than the tree, far slower);
larger FFT windows / more frequency bins (worse — boundaries are short events);
and TransNetV2 (a Conv3D net that will not co-reside with the ROCm/VAAPI stack).
## The model
- **XGBoost regressor** over a ±3 s window of the features above, predicting a
**soft Gaussian proximity-to-boundary target** (`exp(-(d/σ)²)`, σ = 10 s).
Regression to a soft target — rather than a hard 0/1 label — stops a near-miss
from being trained as a hard negative, and yields a smooth score whose **peaks**
are the boundaries.
- **Per-film knee threshold.** The predicted peak heights form a
convex-decreasing curve; the knee (max drop below the endpoints' chord) is where
real boundaries give way to noise. Selecting at the knee **self-calibrates the
boundary count** to roughly the true scene count, per film, with no global
threshold that would be wrong for every grade.
- **Trained on all nine films** for the shipped model. Café Society and Scarface
(the low-contrast grades) *must* be in training — held out, the model cannot
generalise to them; in training they reach 7086% boundary-F1.
Boundary detection, held out (leave-one-out, ±20 s tolerance — appropriate given
~170 s scenes): **~34% F1, versus ~27% for the grayscale baseline.** The absolute
number is capped by the narrative-vs-audiovisual mismatch above; the point is the
downstream effect.
## The result that matters: actor presence
Per-second X-Ray presence F1, macro over the nine films, at the shipped presence
config. The learned column is **leave-one-out** — each film scored by a detector
trained on the other eight:
| boundary source for flood-fill | presence F1 |
| ------------------------------ | ----------- |
| track-extent (flood off) | 62.6% |
| flood + grayscale cuts | 64.0% |
| **flood + learned detector (LOO)** | **74.9%** |
![Macro presence F1 by flood-fill boundary source](assets/images/scene_presence_macro.png)
**+12.3 points over track-extent, +10.9 over the grayscale-cut flood, and it
improves every one of the nine films — under honest leave-one-out.** Per film:
![Per-film presence F1 by boundary source](assets/images/scene_presence_by_source.png)
| film | track-extent | flood+grayscale | flood+learned (LOO) |
| ---- | -----------: | --------------: | ------------------: |
| Benny & Joon | 77.3 | 80.2 | 78.2 |
| Café Society | 59.1 | 62.2 | 69.8 |
| Downton Abbey | 41.0 | 51.8 | **78.6** |
| Lord of War | 74.8 | 77.1 | 77.8 |
| Lovelace | 70.3 | 74.0 | 78.2 |
| The Many Saints of Newark | 37.5 | 43.9 | 53.4 |
| Scarface | 62.6 | **40.9** | **74.9** |
| Sound of Metal | 75.0 | 78.1 | 86.8 |
| Valerian | 65.6 | 67.7 | 76.2 |
The two headline cases:
- **Scarface**: the grayscale-cut flood *breaks* it (62.6 → 40.9), because it
detects one cut in the whole film. The learned detector — **on a film it never
trained on** — takes it to **74.9%**. This is the strongest evidence the
detector generalises: it fixes the exact failure that motivated it, held out.
- **Downton Abbey**: 41.0 (track-extent) → 51.8 (grayscale) → **78.6** — a
+37-point swing on the hardest film.
Naive flood-fill barely beat doing nothing (64% vs 62%) and broke a film. With a
real boundary detector, flood-fill is decisively the right mode.
## In the pipeline
Boundary detection is a **post-EOF step**, like flood-fill itself: the per-film
knee needs every peak, so it can only run once the whole film is seen. The
`camera_position_change_detector` stamps a per-frame RGB histogram onto each frame;
it rides through to the result sink; at end-of-stream the sink runs the detector
over the collected histograms plus the movie's audio log-PSD and snaps the
presence windows to the result. Enable it with:
```bash
scene_analyze --movie <file> --gallery <gallery.h5> \
--scene-xgb-model models/scene_boundary_xgb.json
```
Inference is real XGBoost, built into the binary via CMake (`SAE_SCENE_XGB`); the
audio log-PSD uses FFTW + the existing FFmpeg decode. To keep training and
inference on one feature implementation, the shipped model is **trained on the
C++-extracted features** (`scene_features_dump``train_xgb_cpp.py`) rather than a
re-implementation in Python — parity by construction. Verified end to end through
`scene_analyze` on a movie file and through the Jellyfin work-queue worker.
## Reproduce
```bash
# per-second audio log-PSD for each film
.venv-rocm/bin/python scripts/scene_detector/extract_audio_features.py \
--manifest experiments/manifests/films_LVFace_opencv5.json
# C++ feature matrices (same features training and inference share)
build/scene_features_dump <dump.h5> <movie> <features.h5>
# train the shipped model on all nine films
.venv-rocm/bin/python scripts/scene_detector/train_xgb_cpp.py --train-all
# downstream A/B (track-extent vs flood+grayscale vs flood+learned)
scripts/scene_detector/downstream_presence.py
```