Files
scene-actor-extraction/scripts/validation/README.md
T
dtourolleandClaude Opus 5 042e424961 study(VR-005): minimum face size from downscaled gallery mugshots
Holds out one mugshot per actor, degrades that probe to each candidate
face size and matches it against a gallery held at native resolution,
reporting TPI/FPI per size. Replaces AR-002's 66x66 px working estimate
with a measurement. Needs no video and no ground truth beyond the
mugshot cache already on disk.

LVFace-B over 258 actors, 999 gallery embeddings, threshold 0.754:

    px    12    16    20    24    32    40   48+
   TPI   6.6% 46.5% 81.8% 93.4% 98.1% 99.2% 99.2%

FPI is 0.000 at every size — a face too small to identify degrades to
unidentified, never to a wrong name. rank-1 holds at >=99.6% from 24 px
up, so what fails first is the calibrated probability crossing
threshold, not the ranking.

Two limits on reading this. FPI grows with the number of actors
competing, so 258 understates it against a production library. And
detection and alignment run on the native image with only the resulting
112x112 crop degraded, so landmark error at small face sizes is excluded
by construction and the curve is an upper bound — VR-010 measures the
same question end to end, and lands well above these numbers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

TRACES: VR-005 | AR-002
2026-07-31 15:17:39 +02:00

109 lines
5.0 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# scripts/validation — per-scene actor-presence eval
Validates the pipeline's per-scene "who's on screen" output against external
ground truth, offline. Annealing (`anneal_sec`) means an actor's presence is only
defined *after* the whole file is merged into `[start,end]` windows, so we cannot
score live: process → write the pipeline JSON → **sample timepoints** → compare
predicted vs ground-truth presence sets → micro-sum TP/FP/FN → precision/recall/F1.
## Ground-truth sources
| Source | Semantics | Fair to a face pipeline? | What it measures |
| ------ | --------- | ------------------------ | ---------------- |
| **MovieNet-PS** | on-screen **face** presence per shot | yes — like-for-like | recognition accuracy |
| **Amazon X-Ray** (Zenodo) | **cast-in-scene** (incl. off-camera / non-speaking) | no — penalizes by design | coverage ceiling; recall gap = actors we structurally can't see |
- MovieNet is the honest recognition number.
- X-Ray is an upper bound: its recall gap tells you how much presence is off-camera
cast a face detector can never reach — not a pipeline error.
X-Ray dataset: Zenodo DOI `10.5281/zenodo.17659734` (CC-BY-4.0). Per movie it ships
`people.csv`, `scenes.csv`, `people_in_scenes.csv`.
## Usage
```bash
# against Amazon X-Ray CSVs for one title
python scripts/validation/sample_eval.py \
--pred "Scene in a Mall.json" \
--xray /data/xray/<movie_dir> \
--gallery gallery_arcface_w600k_r50.json \
--step 1.0
# against MovieNet-PS for one title
python scripts/validation/sample_eval.py \
--pred out.json \
--movienet /data/movienet --split Train_app10 --title tt0032138 \
--gallery gallery_arcface_w600k_r50.json
```
### Sampling modes
- `--step S` regular grid every S s (default 1.0) — time-weighted headline number.
- `--random N` N uniform-random timepoints (for confidence intervals).
- `--scene-anchored` one timepoint per GT scene midpoint — the literal X-Ray
"did I get this scene's cast right?" question; neutralizes long-scene bias.
Ground truth is compared **raw** (annealing is *not* applied to GT).
## Matching & masking
Identity is provider-agnostic (`identity.py`): each actor is the *set* of every key
we can derive — `imdb:nm…`, `tmdb:…`, `jf:…`, `name:<normalized>`. Predicted and GT
actors match iff their key-sets intersect, so an output carrying only tmdb/jellyfin
ids still joins X-Ray's `nm` ids via the normalized-name fallback.
Scoring is **masked to `gallery ∩ GT`**: a GT actor absent from the gallery is
ignored (not an FN), so we measure pipeline accuracy, not gallery coverage. Without
`--gallery` the mask falls back to `GT ∩ pred` keys. `--no-mask` disables it.
### Exact id join via the tmdb→imdb crosswalk (recommended)
The gallery/pipeline output key actors by **TMDB** id (no `nm…`), while X-Ray and
MovieNet key on **IMDb**. They only overlap on the fuzzy `name:` key by default.
Build a cached `tmdb→imdb` table once and pass it with `--crosswalk` to turn the
name join into an exact id join:
```bash
# one-time: resolve every gallery tmdb id via TMDB /person/{id}/external_ids
python scripts/validation/tmdb_imdb_map.py \
--gallery gallery_arcface_w600k_r50.json \
--out scripts/validation/tmdb_imdb.json # TMDB_API_KEY from env/.env
# then score with exact ids
python scripts/validation/sample_eval.py --pred out.json --xray <dir> \
--gallery gallery_arcface_w600k_r50.json \
--crosswalk scripts/validation/tmdb_imdb.json
```
The table caches nulls (tmdb ids TMDB has no IMDb id for) and checkpoints, so a
re-run only resolves new ids. TMDB is authoritative for this crosswalk — there is
no clean free bulk `tmdb_person ↔ nm` file, so we query the API once and cache.
## Minimum face size (VR-005)
`min_face_size.py` is a separate, self-contained study: it needs no video and no
ground truth, only the gallery mugshot cache. It holds out one image per actor,
degrades that probe to each candidate face size and matches it against a gallery
held at **native** resolution, reporting TPI/FPI per size — the measurement that
replaces AR-002's 66×66 px estimate.
```bash
python scripts/validation/min_face_size.py \
--images images --gallery gallery_lvface.h5 \
--arcface models/LVFace-B_Glint360K.onnx \
--actors 100 --out experiments/results/vr005_min_face_size
```
FPI grows with the number of actors competing, so a 100-actor run understates it
against a library of thousands: read FPI as relative across sizes, not as an
absolute rate. Re-run per `--arcface` model to see whether `min_face_px` should be
one constant or scale with the embedder (GR-004).
## Files
- `sample_eval.py` — CLI scorer.
- `ground_truth.py``XRayGroundTruth`, `MovieNetGroundTruth` loaders.
- `identity.py` — provider-agnostic match keys.
- `tmdb_imdb_map.py` — build/consult the cached `tmdb→imdb` crosswalk.
- `min_face_size.py` — VR-005 probe-size sweep (see above).
- `test_sample_eval.py` — self-contained tests (`python scripts/validation/test_sample_eval.py`).