dtourolleandClaude Opus 5 f0c7126f80 feat: TrackRegistry — presence follows track extent
The spine of the redesign. Presence is now the extent of a track an actor owns,
[first_seen, last_seen], rather than the subset of frames in which recognition
happened to succeed. An actor recognised only at the end of a long track is
present for all of it, which is what the scene-scoped ground truth actually
records.

AR-013 — `last_seen` as an optional carries the entire liveness state: unset
means on screen, set means went off at that timestamp and still revivable,
reaped means emitted and erased. No missing-frame counter, no expired flag. It
subsumes the tracker's existing two-pool split, so there is no separate revival
path — matching a dormant track is ordinary inter-frame association.

The asymmetry is the point: interior gaps are claimed, the trailing cool-down is
not. A face lost and re-associated within the timeout never closed its track, so
the gap is presence — someone briefly occluded has not left the scene. But a
track that dies ends at its last sighting, never at the death time. That is
precisely the over-claim the retired extinction_sec keep-alive produced, where
presence ran on into the closing credits.

AR-014 — a belief swap A→B closes the track and opens a successor at the swap
frame. Not a correction: two non-twins both clearing the threshold on one face
is not realistic, whereas a track_id carried across a viewpoint change onto a
different person is. Treating it as a swap-and-continue would emit one window
blending two people; treating it as a boundary yields two that are each right.

AR-015 — two live tracks owned by one actor means at least one is wrong, since a
person cannot be in two places at once. A reverse index catches it on the update
that causes it rather than by scanning. This makes identity a third cut
detector, independent of the histogram and TransNetV2 and firing where those
failed.

AR-016 — flush() closes tracks still live at EOF. Without it a film ending
mid-shot silently drops its closing cast, which presents as a recognition miss
rather than a bookkeeping bug.

Reaping hands the dead track to the aggregator and erases it, so the registry
holds only live tracks and its size is bounded by concurrent on-screen faces
rather than growing with the film.

Locking: a frame's association pass is atomic as a unit via FrameScope, since
per-call locking would let another thread observe a half-updated frame.
owner() reads tally and verdict under one lock — separately, a track could be
both unowned and owned within a single promotion decision. A vote for an
already-reaped track is dropped and counted, because a nonzero count means the
timeout is shorter than the matcher's lag.

11 unit tests, driven directly against the registry with no network and no
fixture — the awkward cases are constructed rather than hunted for. Suite: 75
cases, 3236 assertions. Coverage 14/63 to 20/63.

Not yet wired into FaceTrackerFunc; that is AR-007/AR-008.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

TRACES: AR-012, AR-013, AR-014, AR-015, AR-016, AR-017 | SR-002
2026-07-31 09:08:39 +02:00

Scene Actor Extraction

Identifies actors in movie files and produces X-ray-style scene annotations compatible with Jellyfin. Built on a KPN++ pipeline with ArcFace/LVFace embeddings and a tracked-identity matcher.

67.4% macro-F1 against Amazon X-Ray ground truth, on 5 films never seen by the optimizer (89.7% P / 65.4% R training-set; see the generalization-gap discussion in the deep dive). Full benchmark write-up, model comparison, and failure-mode analysis: https://pages.tourolle.paris/dtourolle/scene-actor-extraction/

A perfect X-Ray second on a held-out film A perfect X-Ray second on a held-out film (never used for threshold tuning): every visible face named at 100%, the background extra honestly left unnamed, and the two credited cast without a visible face correctly carried as present off-screen. Bottom panels show the per-second verdict against Amazon X-Ray (green = correct, orange = wrong, blue = missed).

It also doesn't care whether the face is in the room:

Herbie Hancock identified on an in-fiction video-call screen Herbie Hancock at 98% — as a face on a screen inside the movie, under a sci-fi HUD overlay.

How it works

  1. Build a gallery — download actor headshots from TMDB/IMDB, embed them with ArcFace or LVFace (build_gallery / scripts/make_gallery.py).
  2. Analyze a moviescene_analyze decodes frames at configurable FPS, detects faces (SCRFD), tracks them across cuts, matches identities against the gallery using calibrated similarity, and writes time-window JSON.
  3. Output — minimal mode produces Jellyfin-ready actor name + time-window JSON; standard mode adds per-frame bbox, similarity, and track data.

Pipeline topology

Dependencies

Dependency Role
KPN++ Pipeline backbone (nodes, networks)
OpenCV 4 Video decode, image ops, DNN inference
ONNX Runtime SCRFD face detector (dynamic shape nodes unsupported by cv::dnn)
TensorRT + CUDA runtime + cuBLAS Optional TRT engines for SCRFD/ArcFace (--detector-engine/--arcface-engine); identity_matcher's GPU gallery scan
FFmpeg (libav*) NVDEC hardware video decode + colour conversion
nlohmann/json JSON I/O
nanobind Python bindings for sae_embed

Build

cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j$(nproc)

This also builds sae_embed, a Python module (via nanobind) that loads the SCRFD detector and ArcFace embedder once and exposes a reusable embed() method. The gallery-builder scripts (make_gallery.py, make_jellyfin_gallery.py, movienet_eval.py) import it directly — there is no subprocess fallback, so if it's missing they exit with a build instruction:

cmake --build build --target sae_embed

Optional flags:

Flag Default Effect
-DSAE_WEB_DEBUG=ON OFF Enables KPN web debug UI at localhost:9090

Models

The ONNX model weights live in models/ (tracked via Git LFS):

  • LVFace-B_Glint360K.onnx — LVFace embedder (ViT backbone, ICCV 2025), the default (best F1 in the rep4 model bake-off, see docs/rep4-optimizer-results.md)
  • arcface_w600k_r50.onnx — ArcFace embedder, previous default
  • arcface_w600k_mbf.onnx, arcface_r18.onnx — lighter ArcFace alternatives
  • face_detection_yunet_2023mar.onnx — YuNet face detector
  • scrfd_500m_bnkps.onnx — SCRFD face detector

LVFace

LVFace is a Vision-Transformer face recognition model. The LVFace-B_Glint360K.onnx export shares ArcFace's I/O contract (112×112 aligned BGR crop → L2-normalised 512-d embedding) and its (x 127.5)/128 input scaling, so it slots straight into the existing embedder — just point --arcface-model at it:

./build/scene_analyze --arcface-model models/LVFace-B_Glint360K.onnx \
    --gallery gallery.h5 --movie movie.mp4

Important: embeddings from different recognition models are not interchangeable. A gallery (and its calibration cache) must be built with the same embedder used for analysis — rebuild the gallery with --arcface models/LVFace-B_Glint360K.onnx before analysing with LVFace.

If they are missing (e.g. LFS not fetched), re-download them with:

bash scripts/download_models.sh

Model licensing: the model weights carry their own licenses, separate from this project's MIT license, and are redistributed here under those upstream terms. Several — notably the InsightFace "buffalo" models (ArcFace / SCRFD) — are licensed for non-commercial research use only. Review and comply with each model's license before use.

Binaries

Binary Description
scene_analyze Main analysis pipeline, writes JSON output
scene_analyze_debug Same as above + per-frame annotated JPEGs (SAE_DEBUG=1)
scene_preview Live OpenCV display window while analysing
build_gallery Offline gallery builder from a directory of images
embed_faces CLI: image(s) → embedding JSON, used by gallery-builder scripts
sae_embed Python module (nanobind) used by gallery-builder scripts — loads SCRFD+ArcFace once

scene_analyze

./build/scene_analyze --gallery gallery.h5 --movie movie.mp4 [options]

Key options:

Flag Default Description
--fps 1 Frames per second to sample (510 recommended for tracking)
--prob-threshold 0.5 Minimum calibrated match probability
--match-threshold Raw cosine similarity threshold (fallback)
--extinction 5s How long a track persists after last detection
--track-alpha IoU vs. embedding weight in Hungarian assignment
--track-min-iou Minimum IoU gate for spatial assignment
--track-max-embed Maximum embedding distance gate
--track-max-missing Frames a track survives without a detection

Per-movie (TMDB):

python3 scripts/make_gallery.py --tmdb-key <TMDB_KEY> --movie-id <TMDB_ID> --output gallery.h5

Fetches cast images from TMDB and embeds them via sae_embed.

Whole-library (Jellyfin):

python3 scripts/make_jellyfin_gallery.py \
    --jellyfin-url http://jellyfin.local:8096 \
    --api-key <API_KEY> \
    --output gallery.h5

Scans every Movie/Series in Jellyfin, collects the unique cast across the whole library, downloads each actor's headshot directly from Jellyfin (no TMDB key needed), and embeds them via sae_embed into one global gallery.h5. Since identity_matcher scores faces against the entire gallery, scene_analyze can then recognise any actor from your library in any film — not just the cast listed for that one title. Pass --merge on later runs to only embed actors newly added to the library. Pass --tmdb-key to fall back to TMDB profile images for actors with no usable image cached in Jellyfin.

Jellyfin/TMDB lookups and image downloads for different actors run concurrently (--workers, default 8). Embedding is GPU-bound, so it's gated separately via --embed-concurrency (default 1) — only that many embed calls run at once while other actors' downloads continue in the background.

To restrict a single-title run to that title's credited cast (faster, fewer look-alike mismatches), filter the global gallery first:

python3 scripts/filter_gallery.py \
    --gallery gallery.h5 \
    --jellyfin-url http://jellyfin.local:8096 \
    --api-key <API_KEY> \
    --title "The Matrix" \
    --output gallery_matrix.h5

Running directly from Jellyfin

scripts/run_from_jellyfin.py resolves a title to its media file via the Jellyfin API, filters the gallery to that title's cast, and runs scene_analyze in one step. Requires this tool to run on a host that shares Jellyfin's media mount (it uses the item's on-disk Path, not a stream URL):

python3 scripts/run_from_jellyfin.py \
    --jellyfin-url http://jellyfin.local:8096 \
    --api-key <API_KEY> \
    --title "The Matrix" \
    --gallery gallery.h5 \
    -- --fps 5 --verbosity 2

Anything after -- is passed through to scene_analyze unchanged. Pass --no-filter to use the gallery as-is (skip per-title cast filtering), or --item-id instead of --title to skip the search.

After a successful run, the output JSON is pushed to the JRay Jellyfin plugin's Truth endpoint (PUT /Plugins/JRay/Items/{itemId}/Truth) so Jellyfin picks it up immediately, using --api-key (must be an Administrator key for the push to succeed). Pass --no-push to skip this and only write --output locally (e.g. for local debugging).

Worker mode

Pass --worker instead of --item-id/--title to run this as an extraction worker: it polls the JRay plugin's GET /Plugins/JRay/Tasks/Pending endpoint for a random batch of items with no truth data yet, processes each one, and pushes the result back. The endpoint's sampling spreads work across the backlog without any server-side task tracking, so any number of workers can poll the same library concurrently.

python3 scripts/run_from_jellyfin.py \
    --jellyfin-url http://jellyfin.local:8096 \
    --api-key <ADMIN_API_KEY> \
    --gallery whole_gallery.h5 \
    --worker \
    -- --fps 5
  • --poll-limit — batch size requested from Tasks/Pending (default 10, max 100)
  • --poll-interval — seconds to sleep between polls when the backlog is empty (default 60)
  • --once — process a single batch and exit instead of looping forever

A failure on one item (bad path, push rejected, etc.) is logged and the worker moves on to the next item rather than exiting.

Output format

Minimal (default) — Jellyfin-ready:

[
  { "actor": "Name", "start": 12.0, "end": 45.5 }
]

Standard — per-frame detail with bounding boxes, similarity scores, and track IDs.

Evaluation

Scripts in eval/ and scripts/movienet_*.py support benchmarking against the MovieNet dataset.

License

The source code in this repository is licensed under the MIT License.

The MIT license covers only the code. The model weights in models/ (see Models) are redistributed under their own licenses — several for non-commercial research use only. Third-party libraries this software links against (OpenCV, ONNX Runtime, FFmpeg, TensorRT/CUDA, nlohmann/json, nanobind, and others) likewise carry their own licenses.

S
Description
No description provided
Readme MIT
737 MiB
Languages
C++ 57.4%
Python 36.8%
Shell 3.9%
CMake 1.9%