Files
scene-actor-extraction/docs/methodology.md
T
dtourolle ea922356f1 docs: archive the July 2026 report; new methodology for the opencv5 run
The July report (4-model ArcFace/LVFace bake-off, pre-opencv5 framework,
3-film training + held-out validation) is superseded by the opencv5 build:
single-model LVFace-B, a 6-knob DE sweep over all 9 films, flood-fill
presence, and the registry/decode fixes. Rather than overwrite it, archive
it date-suffixed and start the current report fresh.

- Rename the six July result pages to *-2026-07.md, rewrite their
  intra-archive cross-links, and add an "Archived (July 2026)" banner to each.
- mkdocs nav: current report at top, the July set under an Archive section.
- New docs/methodology.md for the opencv5 run: corrects the withdrawn
  anneal_sec/extinction_sec presence bridging (windows are now
  [first_seen, last_seen], AR-012/013), documents the two presence modes
  (track_extent / flood), and records that every eval scores all 9 films.

The current experiment log (model-bakeoff.md) and Home rewrite land once
the DE sweep converges and the final optimum is known.
2026-08-09 10:48:45 +02:00

5.7 KiB
Raw Blame History

How we score against X-Ray

Every number in this report comes from one comparison, and that comparison has a mismatch at its core: the ground truth is scene-level, the pipeline's output is per-second, and the two do not mean the same thing. This page documents the comparison once so the findings can rely on it.

What Amazon X-Ray records

X-Ray ships three tables per film: scenes.csv (a list of [start, end] timespans), people_in_scenes.csv (which actors are credited in each scene), and people.csv (actor identities). There is no per-frame or per-second annotation anywhere in X-Ray. A scene might run 45 seconds, and X-Ray records one cast list for the entire span, not "on screen from second 12 to second 30."

To compare against per-second predictions, second_score.py expands every scene into per-second ground truth by copying the whole scene's cast list onto every second inside it:

for sn, (t0, t1) in spans.items():
    cast = scene_cast.get(sn, [])
    for t in range(int(t0), int(t1)):
        timeline[t] = cast

If X-Ray credits five actors to a 30-second scene, all five count as ground truth present for all 30 seconds, including seconds where only one is on screen. This is not a simplification the pipeline introduces; it is the only reading X-Ray's data supports, because X-Ray records nothing finer.

How the pipeline reports presence

A presence claim is one actor owning one time window. How that window is derived is a tunable choice — a knob the optimizer weighs — with two modes:

  • track_extent (default). A claim is exactly [first_seen, last_seen] of a track the actor owned (AR-012), ending at the last sighting and never after (AR-013). There is no keep-alive: the withdrawn anneal_sec and the scene-tracker extinction_sec — which the July report's windows were held open by — are gone. A track that survives its own gaps needs no bridge; a gap after the final sighting is never claimed.
  • flood. Each claim is snapped to the shot it sits in, so an actor seen once anywhere in a shot is reported for the whole shot [prev_boundary, next_boundary]. Boundaries come from TransNetV2 shot detection when available, otherwise from the always-on histogram cut detector (is_cut). This trades precision for recall against X-Ray's scene-level granularity, and the optimizer decides per run whether it pays.

Do not confuse the surviving track_extinction_sec with the withdrawn scene extinction_sec: the former bounds how long a lost track stays available for re-association (a tracking question), and never extends a presence claim.

The two limits this does not resolve

The face-vs-presence ceiling. X-Ray credits scene membership regardless of whether a face is ever visible: background crew, characters shot from behind, voice-only presence. No face pipeline can recover a face that never appears, so recall against X-Ray is a structural ceiling, not a defect.

Flood-fill can overshoot. Snapping to a shot correctly answers "still in this scene" through an intra-scene cut, but a shot boundary is not a scene boundary: on a film with sparse cuts, flood-fill can carry an actor across a long "shot" they only briefly appeared in. This is why flood-fill is a knob, not a default — its value depends on the film's cut density.

Precision, recall, and the misID weighting

Per sampled second t:

TPI (true positive instances): actors both X-Ray and the pipeline agree are present.

FPI (false positive instances): actors the pipeline reports that are not in X-Ray's cast for this second, split into:

  • FPI_incast: the actor is in the film's cast, just not credited to this scene. A timing or boundary slip.
  • FPI_misid: the actor is not in the film's cast at all. A genuine wrong-identity error, weighted 10× in the precision objective, because naming someone not even in the film is categorically worse than a few seconds of scene-boundary slop.

!!! note "Every headline P and F1 is misID-weighted"

Precision puts each `FPI_misid` into the denominator 10 times
(`precision = TPI / (TPI + FPI_incast + 10·FPI_misid)`,
[`second_score.py`](https://REPOLINK/scripts/optimizer/second_score.py)).
This deliberately punishes naming an out-of-film actor far harder than a
boundary slip, so the `P` column is not raw precision and a misID-heavy
film's `P` is depressed super-linearly.

FN (false negatives): actors X-Ray lists that the pipeline never reports, counted only for actors who have a gallery reference embedding. An actor with no reference photo can never be recognized, and counting them as a miss would measure gallery coverage, not recognition accuracy.

Two further numbers accompany F1:

agreement_rate: mean per-second Jaccard overlap (|Pred ∩ GT| / |Pred GT|) — partial credit, so naming 2 of 3 present actors scores 2/3, not 0.

exact_match_rate: the fraction of seconds where the pipeline's named set exactly equals X-Ray's — no partial credit, dominated by recall.

The benchmark set

Unlike the July report — which trained on a 3-film subset and validated on held-out films to keep evaluations fast — this run scores all 9 films on every evaluation. The registry one-clock fix and uncapped dumps made full-set replay affordable, so the reported optimum is tuned against the complete set rather than a training subset.

Reproduce

python3 scripts/optimizer/second_score.py \
    --pred pred.json --xray experiments/xray/.../<xray_dir> \
    --gallery experiments/galleries/gallery_LVFace-B_Glint360K.h5

See the full experiment log for how pred.json is produced and where the shipped src/config.hpp defaults come from.