Replaces narrative claims with verified numbers across all report pages:
- Cross-model held-out validation (LVFace/mbf/r18, all 5 held-out
films): LVFace wins every film outright, not just "consistent with"
the training-set pick. r50 dropped from the detailed comparison
(gallery has ~30% fewer reference images per actor than the other
three models on identical source photos).
- Per-film training breakdown: LVFace does not win every training
film (mbf beats it on Lord of War); the 75.3% macro figure hides a
10.7pp spread.
- Gallery coverage computed per film (20.3%-78.6%) instead of one
flat 67%-missing average.
- Found and fixed a real scoring bug in optimize.py: a candidate
whose hardest film's replay timed out was averaged over survivors
instead of penalized, silently rewarding partial coverage. Affected
3 of 16 training combos; corrected throughout, and optimize.py now
scores an incomplete evaluation f1=0.0 instead of averaging over
whichever films happened to finish.
- Every FPI frame in the deep dive now comes from the proper montage
renderer (Onscreen/Offscreen panel, ghosts never drawn as boxes),
never the bare-box debug overlay used earlier.
- Every distinct out-of-cast name across all 9 films gets its own
frame at its first appearance (9 names, 4 films), not a
single-example spot check: 2 ground-truth gaps, 1 photograph
misread as a person, 6 genuine lookalike confusions.
- New methodology.md: the scene-level-vs-per-second scoring mismatch
that the rest of the report assumes, written out once.
- Cut the deadlock/gdb debugging narrative from the experiment log;
kept the one fact that matters (KPN's node/network split lets the
expensive GPU stage run once and the cheap stage replay against
cached embeddings).
- Plain declarative style throughout, no em dashes, no blog voice.