d753062c6ce6df9b6a7f027ffcdad6d13530fb6c
Source (KPN++ pipeline nodes, ArcFace embedders, SCRFD/YuNet detectors, gallery builder), build scripts, and eval artifacts. - external/KPN as a git submodule (gitea.tourolle.paris/dtourolle/KPN) - ONNX models tracked via Git LFS (models/*.onnx) - generated outputs, TensorRT engines, reference repos, and media ignored
Scene Actor Extraction
Identifies actors in movie files and produces X-ray-style scene annotations compatible with Jellyfin. Built on a KPN++ pipeline with ArcFace embeddings and a tracked-identity matcher.
How it works
- Build a gallery — download actor headshots from TMDB/IMDB, embed them with ArcFace (
build_gallery/scripts/make_gallery.py). - Analyze a movie —
scene_analyzedecodes frames at configurable FPS, detects faces (YuNet/SCRFD), tracks them across cuts, matches identities against the gallery using calibrated similarity, and writes time-window JSON. - Output — minimal mode produces Jellyfin-ready actor name + time-window JSON; standard mode adds per-frame bbox, similarity, and track data.
Dependencies
| Dependency | Role |
|---|---|
| KPN++ | Pipeline backbone (nodes, networks) |
| OpenCV 4 | Video decode, image ops, DNN inference, YuNet face detection |
| ONNX Runtime | SCRFD face detector (dynamic shape nodes unsupported by cv::dnn) |
| nlohmann/json | JSON I/O |
Build
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j$(nproc)
Optional flags:
| Flag | Default | Effect |
|---|---|---|
-DSAE_WEB_DEBUG=ON |
OFF | Enables KPN web debug UI at localhost:9090 |
Models
Download the required ONNX models:
bash scripts/download_models.sh
Models are placed in external/:
arcface_w600k_r50.onnx— primary ArcFace embedderarcface_w600k_mbf.onnx,arcface_r18.onnx— lighter alternativesface_detection_yunet_2023mar.onnx— YuNet face detectorscrfd_500m_bnkps.onnx— SCRFD face detector
Binaries
| Binary | Description |
|---|---|
scene_analyze |
Main analysis pipeline, writes JSON output |
scene_analyze_debug |
Same as above + per-frame annotated JPEGs (SAE_DEBUG=1) |
scene_preview |
Live OpenCV display window while analysing |
build_gallery |
Offline gallery builder from a directory of images |
embed_faces |
Standalone embedder used by gallery scripts |
scene_analyze
./build/scene_analyze --gallery gallery.json --input movie.mp4 [options]
Key options:
| Flag | Default | Description |
|---|---|---|
--fps |
1 | Frames per second to sample (5–10 recommended for tracking) |
--prob-threshold |
0.5 | Minimum calibrated match probability |
--match-threshold |
— | Raw cosine similarity threshold (fallback) |
--extinction |
5s | How long a track persists after last detection |
--track-alpha |
— | IoU vs. embedding weight in Hungarian assignment |
--track-min-iou |
— | Minimum IoU gate for spatial assignment |
--track-max-embed |
— | Maximum embedding distance gate |
--track-max-missing |
— | Frames a track survives without a detection |
--track-min-frames |
3 | Observations before a track's mean embedding is used for matching |
Gallery builder
python3 scripts/make_gallery.py --tmdb-bearer <JWT> --movie-id <TMDB_ID> --output gallery.json
Fetches cast images from TMDB and embeds them via embed_faces.
Output format
Minimal (default) — Jellyfin-ready:
[
{ "actor": "Name", "start": 12.0, "end": 45.5 }
]
Standard — per-frame detail with bounding boxes, similarity scores, and track IDs.
Pipeline topology
frame_source → face_detector → face_aligner → embedder
→ face_tracker → identity_matcher → scene_tracker → result_sink
Debug/preview branches fan out automatically from identity_matcher.
Evaluation
Scripts in eval/ and scripts/movienet_*.py support benchmarking against the MovieNet dataset.
Languages
C++
57.4%
Python
36.8%
Shell
3.9%
CMake
1.9%