docs: full data-grounded rewrite of the performance report
Replaces narrative claims with verified numbers across all report pages: - Cross-model held-out validation (LVFace/mbf/r18, all 5 held-out films): LVFace wins every film outright, not just "consistent with" the training-set pick. r50 dropped from the detailed comparison (gallery has ~30% fewer reference images per actor than the other three models on identical source photos). - Per-film training breakdown: LVFace does not win every training film (mbf beats it on Lord of War); the 75.3% macro figure hides a 10.7pp spread. - Gallery coverage computed per film (20.3%-78.6%) instead of one flat 67%-missing average. - Found and fixed a real scoring bug in optimize.py: a candidate whose hardest film's replay timed out was averaged over survivors instead of penalized, silently rewarding partial coverage. Affected 3 of 16 training combos; corrected throughout, and optimize.py now scores an incomplete evaluation f1=0.0 instead of averaging over whichever films happened to finish. - Every FPI frame in the deep dive now comes from the proper montage renderer (Onscreen/Offscreen panel, ghosts never drawn as boxes), never the bare-box debug overlay used earlier. - Every distinct out-of-cast name across all 9 films gets its own frame at its first appearance (9 names, 4 films), not a single-example spot check: 2 ground-truth gaps, 1 photograph misread as a person, 6 genuine lookalike confusions. - New methodology.md: the scene-level-vs-per-second scoring mismatch that the rest of the report assumes, written out once. - Cut the deadlock/gdb debugging narrative from the experiment log; kept the one fact that matters (KPN's node/network split lets the expensive GPU stage run once and the cheap stage replay against cached embeddings). - Plain declarative style throughout, no em dashes, no blog voice.
This commit is contained in:
@@ -14,6 +14,11 @@ compile_commands.json
|
||||
*.so
|
||||
*.dylib
|
||||
*.json
|
||||
# Exception: small, curated result summaries backing specific numbers quoted
|
||||
# in docs/ (cross-model held-out scores, per-film training breakdown, gallery
|
||||
# coverage). Regenerate with scripts/docs/run_holdout_all_models.py and
|
||||
# scripts/docs/gallery_coverage_per_film.py.
|
||||
!docs_data/*.json
|
||||
# Video files
|
||||
*.mp4
|
||||
*.mkv
|
||||
|
||||
+78
-45
@@ -1,70 +1,103 @@
|
||||
# Which embedding model is best?
|
||||
|
||||
Four candidates went into the bake-off: three ArcFace variants (w600k-R50,
|
||||
R18, w600k-MBF) and LVFace-B (Glint360K), a Vision-Transformer embedder that's
|
||||
a drop-in replacement for ArcFace's `[N,3,112,112]` input / 512-d output. The
|
||||
open question: is LVFace (455MB) actually better, or just the biggest?
|
||||
Three ArcFace variants (w600k-R50, R18, w600k-MBF) and LVFace-B (Glint360K,
|
||||
455MB) were compared. r50 is excluded from the training/held-out comparison
|
||||
below; its gallery has roughly 30% fewer reference images per actor than the
|
||||
other three on the identical source photos, which confounds a direct score
|
||||
comparison (see [the full experiment log](model-bakeoff.md) for detail). It
|
||||
remains in the calibration comparison, which does not depend on the gallery
|
||||
image count.
|
||||
|
||||
## First signal: calibration curves
|
||||
|
||||
Each gallery carries a fitted Platt sigmoid `P(match | cosine similarity) =
|
||||
σ(a·sim + b)`, embedded directly in the gallery's HDF5 file
|
||||
([`src/gallery/gallery_calibration.hpp`](https://REPOLINK/src/gallery/gallery_calibration.hpp)). This is a property of the embedding
|
||||
space alone — computed from intra/inter-actor reference-image pairs, no
|
||||
tracking or scene logic involved — so it's a clean first read on discriminative
|
||||
power before running a single benchmark.
|
||||
σ(a·sim + b)`, stored directly in the gallery HDF5
|
||||
([`src/gallery/gallery_calibration.hpp`](https://REPOLINK/src/gallery/gallery_calibration.hpp)).
|
||||
This is a property of the embedding space alone, computed from intra- and
|
||||
inter-actor reference-image pairs with no tracking or scene logic involved,
|
||||
so it is a clean first read on discriminative power before running a
|
||||
benchmark.
|
||||
|
||||

|
||||
|
||||
| model | `a` (steepness) | boundary at P=0.5 |
|
||||
| model | a (steepness) | boundary at P=0.5 |
|
||||
|---|---|---|
|
||||
| **LVFace-B Glint360K** | **17.7** | **sim 0.228** |
|
||||
| LVFace-B Glint360K | 17.7 | sim 0.228 |
|
||||
| ArcFace w600k-MBF | 16.2 | sim 0.267 |
|
||||
| ArcFace w600k-R50 | 15.4 | sim 0.301 |
|
||||
| ArcFace R18 | 15.3 | sim 0.309 |
|
||||
|
||||
LVFace has both the steepest transition and the lowest decision boundary — it
|
||||
separates same-actor from different-actor reference pairs more confidently, at
|
||||
a *lower* similarity threshold, than any ArcFace variant. That's a genuine
|
||||
head start before the tracking/scoring pipeline is even involved.
|
||||
LVFace has both the steepest transition and the lowest decision boundary,
|
||||
separating same-actor from different-actor reference pairs more confidently
|
||||
at a lower similarity than any ArcFace variant.
|
||||
|
||||
## Second signal: F1 on the actual benchmark
|
||||
## Second signal: held-out F1
|
||||
|
||||
Best full-gallery (no cast-restriction) result per model, from the 16-combo
|
||||
bake-off matrix ([full experiment log](model-bakeoff.md)):
|
||||
Each model's own tuned `full_exp` config, replayed against the 5 films the
|
||||
optimizer never saw and scored the same way:
|
||||
|
||||
| film | LVFace F1 | mbf F1 | r18 F1 |
|
||||
|---|---|---|---|
|
||||
| Benny & Joon | 83.0% | 78.5% | 77.1% |
|
||||
| Lovelace | 77.5% | 73.7% | 72.2% |
|
||||
| Valerian and the City of a Thousand Planets | 74.1% | 70.2% | 71.0% |
|
||||
| Downton Abbey: A New Era | 56.2% | 55.0% | 53.0% |
|
||||
| The Many Saints of Newark | 46.3% | 44.5% | 42.1% |
|
||||
| **macro average** | **67.4%** | **64.4%** | **63.1%** |
|
||||
|
||||
LVFace scores highest on all 5 held-out films; the ranking never flips
|
||||
between models. Total misID count across the 5 films: LVFace 1032, mbf
|
||||
2197, r18 1224. LVFace has less than half mbf's misID total and still
|
||||
scores higher on every film.
|
||||
|
||||
Held-out results are stronger evidence than training results, because
|
||||
training numbers can reflect what the optimizer was tuned to fit rather
|
||||
than general performance. On training data, the ordering is not as clean:
|
||||
|
||||
| film | LVFace F1 | mbf F1 | r18 F1 | best |
|
||||
|---|---|---|---|---|
|
||||
| Café Society | 68.1% | 62.2% | 60.1% | LVFace |
|
||||
| Lord of War | 75.6% | 77.2% | 75.6% | mbf |
|
||||
| Scarface | 71.5% | 68.6% | 64.1% | LVFace |
|
||||
| Sound of Metal | 78.8% | 76.5% | 71.6% | LVFace |
|
||||
|
||||
mbf beats LVFace on Lord of War (77.2% vs 75.6%), the only film in either
|
||||
table where LVFace does not score highest. LVFace's training-set macro
|
||||
average (75.3%, see [the full experiment log](model-bakeoff.md)) is not a
|
||||
uniform win across every film it contributes to; the held-out result, where
|
||||
LVFace wins all 5 films outright, is the stronger claim.
|
||||
|
||||
This reverses an earlier, superseded benchmarking pass that used a
|
||||
scene-union metric and found the three models statistically
|
||||
indistinguishable (around 85% each), concluding LVFace was not worth its
|
||||
size. That metric masked out-of-cast false positives behind a
|
||||
gallery-intersect-cast recall filter; the per-second metric used here does
|
||||
not.
|
||||
|
||||
## Full training-matrix picture
|
||||
|
||||

|
||||
|
||||
Best full-gallery combo per model (all three are `full_exp`), from the
|
||||
training matrix in [the full experiment log](model-bakeoff.md):
|
||||
|
||||
| model | F1 | P | R | misID |
|
||||
|---|---|---|---|---|
|
||||
| **LVFace-B Glint360K** | **75.3%** | 89.7% | **65.4%** | 232 |
|
||||
| ArcFace w600k-MBF | 74.2% | 87.4% | 64.4% | 57 |
|
||||
| LVFace-B Glint360K | 75.3% | 89.7% | 65.4% | 232 |
|
||||
| ArcFace w600k-MBF | 72.0% | 87.7% | 61.4% | 240 |
|
||||
| ArcFace R18 | 69.1% | 87.6% | 57.7% | 242 |
|
||||
| ArcFace w600k-R50 | 68.5% | 94.0% | 54.1% | 150 |
|
||||
|
||||
The full 16-combo picture makes the model ordering visible at a glance — LVFace
|
||||
(yellow) tops both the restricted and full columns, and R18 (green) props up
|
||||
the bottom of the full-gallery ranking:
|
||||
LVFace leads within both the restricted and full gallery modes, visible
|
||||
directly in the chart above without reading the table. The three models'
|
||||
misID counts on the full gallery are nearly identical (232/240/242); LVFace's
|
||||
lead here is a precision-and-recall lead, not a misID one.
|
||||
|
||||

|
||||
## Operational note
|
||||
|
||||
LVFace wins outright, with the highest recall of any full-mode combo. This
|
||||
reverses an earlier conclusion from a prior (superseded) benchmarking pass
|
||||
using a scene-union metric, which found the three models statistically
|
||||
indistinguishable (~85% each) and concluded LVFace wasn't worth its size — that
|
||||
metric hid out-of-cast false positives behind a gallery∩cast recall mask (see
|
||||
[the prior optimizer round](optimizer-experiments.md)); the per-second metric
|
||||
used here does not.
|
||||
|
||||
Held-out validation (5 films never seen by the optimizer) confirms LVFace's
|
||||
lead holds up out of sample — see the
|
||||
[LVFace deep dive](lvface-deep-dive.md) for the full breakdown, including
|
||||
where it fails.
|
||||
|
||||
## Caveat: model choice is an operational change
|
||||
|
||||
Switching the default embedder isn't just flipping a config value — the
|
||||
gallery itself is model-specific (embeddings from different models aren't
|
||||
comparable), so any existing gallery built against ArcFace w600k-R50 needs to
|
||||
be rebuilt from source images against LVFace before the new default takes
|
||||
effect. [`scripts/optimizer/reembed_gallery.py`](https://REPOLINK/scripts/optimizer/reembed_gallery.py)
|
||||
Switching the default embedder is not a config change alone; the gallery
|
||||
is model-specific, since embeddings from different models are not
|
||||
comparable. Any existing gallery built against a different model must be
|
||||
rebuilt from source images before the new default takes effect.
|
||||
[`scripts/optimizer/reembed_gallery.py`](https://REPOLINK/scripts/optimizer/reembed_gallery.py)
|
||||
does this from a reference gallery's cached source images without
|
||||
re-downloading anything.
|
||||
|
||||
+53
-49
@@ -1,65 +1,69 @@
|
||||
# Whole gallery vs. limited (cast-restricted) gallery
|
||||
# Whole gallery vs. cast-restricted gallery
|
||||
|
||||
Two ways to run the matcher: **full** scores every detected face against the
|
||||
entire library gallery (2418 actors across the 9-film benchmark set); **restricted**
|
||||
pre-filters each film's gallery down to just its Jellyfin-credited cast (typically
|
||||
~15 top-billed actors) before the matcher ever runs.
|
||||
Two ways to run the matcher. Full mode scores every detected face against
|
||||
the entire 2418-actor gallery. Restricted mode pre-filters each film's
|
||||
gallery down to just its Jellyfin-credited cast (typically around 15
|
||||
top-billed actors) before the matcher runs.
|
||||
|
||||
## The result
|
||||
## Result
|
||||
|
||||
Averaged across all 4 models and both expansion settings, on the 4 bake-off training
|
||||
films:
|
||||
Averaged across the 3 compared models (r50 excluded, see
|
||||
[the full experiment log](model-bakeoff.md)) and both expansion settings, on
|
||||
the 4 training films:
|
||||
|
||||
| scope | F1 | P | R | total misID (8 evals) |
|
||||
| scope | F1 | P | R | total misID |
|
||||
|---|---|---|---|---|
|
||||
| full | 71.2% | 91.1% | 59.0% | 1073 |
|
||||
| **restricted** | **74.5%** | 92.2% | **62.9%** | **329** |
|
||||
| full | 71.1% | 89.6% | 59.6% | 1121 |
|
||||
| restricted | 75.9% | 90.4% | 65.6% | 299 |
|
||||
|
||||
This is not a precision/recall trade — restriction wins on every axis at once:
|
||||
**+3.3pp F1, +3.9pp recall, and less than a third the total misIDs.** Fewer
|
||||
Restriction improves every metric at once, not a precision/recall trade:
|
||||
+4.8pp F1, +6.0pp recall, roughly a quarter the total misIDs. Fewer
|
||||
candidates in the matcher's search space means fewer opportunities for a
|
||||
look-alike false match (an actor who happens to share enough facial structure
|
||||
with someone in the film, but isn't actually in it), and the recall gain shows
|
||||
it isn't costing real detections to get there.
|
||||
lookalike false match, and the recall gain shows this does not cost real
|
||||
detections.
|
||||
|
||||
Per-model, every single model's best-scoring combo in the full 16-way matrix is
|
||||
a `restricted` variant — visible directly in the ranking below (filled dots =
|
||||
restricted, open = full; the filled dots cluster at the top for every color):
|
||||
Every model's best-scoring combo in the training matrix uses the
|
||||
restricted gallery:
|
||||
|
||||

|
||||

|
||||
|
||||
See the full table in the
|
||||
[bake-off experiment log](model-bakeoff.md). Two
|
||||
combos hit **zero** true out-of-cast misidentifications:
|
||||
`arcface_w600k_mbf_restricted_exp` (F1 76.5%) and, in full mode,
|
||||
`LVFace-B_Glint360K_full_noexp` (F1 72.4%) — restriction isn't the only way to
|
||||
reach misid=0, but it's the more reliable one.
|
||||
See [the full experiment log](model-bakeoff.md) for the complete table. One
|
||||
combo reaches zero true out-of-cast misidentifications,
|
||||
`arcface_w600k_mbf_restricted_exp` (F1 76.2%), and it is a restricted one,
|
||||
consistent with restriction, not expansion, being what suppresses cross-film
|
||||
confusions.
|
||||
|
||||
## Why this isn't the shipped default
|
||||
The restriction effect (+4.8pp averaged across models) is larger than the
|
||||
model-choice effect: LVFace beats r18 by 6.2pp in full mode but beats mbf by
|
||||
3.3pp. Restriction is the single strongest lever in the matrix.
|
||||
|
||||
Cast-restriction is implemented today only as an **offline optimizer technique**
|
||||
## Why this is not the shipped default
|
||||
|
||||
Cast restriction is implemented today only as an offline optimizer
|
||||
technique
|
||||
([`scripts/optimizer/cast_restrict.py`](https://REPOLINK/scripts/optimizer/cast_restrict.py)):
|
||||
it pre-builds a filtered gallery file
|
||||
per film, using Jellyfin's own cast list, before the benchmark ever calls the
|
||||
matcher. There's no runtime "restrict matching to this title's credited cast"
|
||||
switch in the shipped application — `scene_analyze` always matches against
|
||||
whatever single gallery file it's given.
|
||||
it pre-builds a filtered gallery file per film using Jellyfin's cast list
|
||||
before the benchmark calls the matcher. There is no runtime "restrict to
|
||||
this title's credited cast" switch in the shipped application;
|
||||
`scene_analyze` always matches against whatever single gallery file it is
|
||||
given.
|
||||
|
||||
Building that as a real feature would need, at minimum:
|
||||
Building this as a real feature requires:
|
||||
|
||||
- A live Jellyfin cast lookup at analysis time (the title is already known —
|
||||
[`scripts/run_from_jellyfin.py`](https://REPOLINK/scripts/run_from_jellyfin.py)
|
||||
already does this same lookup for its own
|
||||
`filter_gallery`-based restriction path, just not wired into `scene_analyze`
|
||||
itself as a first-class option).
|
||||
- A decision on the *fallback*: what happens to a real, uncredited cameo
|
||||
(see the Germar Terrell Gardner case in the LVFace deep-dive) if the gallery
|
||||
never includes them at all?
|
||||
- Regenerating the restricted-gallery cache whenever the title's Jellyfin cast
|
||||
list changes.
|
||||
- A live Jellyfin cast lookup at analysis time. The title is already known,
|
||||
and [`scripts/run_from_jellyfin.py`](https://REPOLINK/scripts/run_from_jellyfin.py)
|
||||
already performs this lookup for its own `filter_gallery`-based
|
||||
restriction path; it is not wired into `scene_analyze` as a first-class
|
||||
option.
|
||||
- A decision on the fallback case: what happens to a real, uncredited
|
||||
cameo (see the Germar Terrell Gardner and Talia Balsam cases in the
|
||||
[LVFace deep dive](lvface-deep-dive.md#where-lvface-beat-x-ray)) if the
|
||||
restricted gallery never includes them at all.
|
||||
- Regenerating the restricted-gallery cache whenever a title's Jellyfin
|
||||
cast list changes.
|
||||
|
||||
This is why the shipped [`src/config.hpp`](https://REPOLINK/src/config.hpp)
|
||||
defaults use the `full`-mode winner
|
||||
(`LVFace-B_Glint360K_full_exp`, F1 75.3% training / 67.4% held-out macro) rather
|
||||
than the higher-scoring `restricted_exp` (78.3%) — the 78.3% number describes a
|
||||
capability the app doesn't have yet, not what actually ships.
|
||||
The shipped [`src/config.hpp`](https://REPOLINK/src/config.hpp) defaults use
|
||||
the full-mode winner (`LVFace-B_Glint360K_full_exp`, F1 75.3% training,
|
||||
67.4% held-out macro) rather than the higher-scoring `restricted_exp`
|
||||
(78.3%), because 78.3% describes a capability the application does not
|
||||
have yet.
|
||||
|
||||
+46
-44
@@ -1,27 +1,32 @@
|
||||
# scene-actor-extraction
|
||||
|
||||
A face-recognition pipeline that finds when each actor appears on screen in a
|
||||
film or TV episode — built on [KPN++](https://gitea.tourolle.paris/dtourolle/KPN)
|
||||
(a C++20 Kahn Process Network library) for the detect → track → match → scene
|
||||
pipeline, with a Jellyfin-integrated gallery and an X-Ray-validated optimizer.
|
||||
A face-recognition pipeline that finds when each actor appears on screen in
|
||||
a film or TV episode, built on [KPN++](https://gitea.tourolle.paris/dtourolle/KPN)
|
||||
(a C++20 Kahn Process Network library) for the detect, track, match, and
|
||||
scene pipeline, with a Jellyfin-integrated gallery and an X-Ray-validated
|
||||
optimizer.
|
||||
|
||||
This is a perfect X-Ray second, on a film the optimizer never saw:
|
||||
This is a correctly scored second from a held-out film, one the optimizer
|
||||
never saw during tuning:
|
||||
|
||||

|
||||
|
||||
Every visible face named at 100% — Chris Noth, Hank Azaria, Bobby Cannavale —
|
||||
the background extra honestly left unnamed, and the two credited cast without
|
||||
a visible face correctly carried as present off-screen by the tracker's
|
||||
presence windows. That's the pipeline exactly reproducing Amazon X-Ray's
|
||||
record for this second.
|
||||
Every visible face is named at 100% confidence (Chris Noth, Hank Azaria,
|
||||
Bobby Cannavale), the background extra is correctly left unnamed, and the
|
||||
two credited cast members without a visible face are correctly reported
|
||||
present but not visible. This matches Amazon X-Ray's own record for this
|
||||
second exactly.
|
||||
|
||||
It doesn't always go like that: the hardest held-out film scores 46% F1, and
|
||||
the report is honest about *why* — one tunable trade (extinction bridging at
|
||||
hard cuts), one structural ceiling (X-Ray credits people whose faces never
|
||||
appear), and a few cases where the pipeline is right and X-Ray is wrong. The
|
||||
evidence for all of it is in the pages below.
|
||||
Results are not uniform across films. The hardest held-out film scores 46%
|
||||
F1. This report documents why: one tunable trade (extinction bridging at
|
||||
hard cuts), one structural limit (X-Ray credits people whose faces never
|
||||
appear on screen), and a small number of cases where the pipeline is
|
||||
correct and X-Ray's ground truth is not. Read
|
||||
[how we score against X-Ray](methodology.md) first. X-Ray's ground truth is
|
||||
scene-level; the pipeline's output is per-second. That difference shapes
|
||||
every finding below.
|
||||
|
||||
## Start here — four questions this bake-off answers
|
||||
## Findings
|
||||
|
||||
<div class="grid cards" markdown>
|
||||
|
||||
@@ -29,55 +34,51 @@ evidence for all of it is in the pages below.
|
||||
|
||||
---
|
||||
|
||||
Calibration curves first (discriminative power, independent of any
|
||||
threshold), then F1 on the actual benchmark. LVFace-B Glint360K wins
|
||||
both.
|
||||
Calibration curves first, independent of any threshold, then held-out
|
||||
F1 across three models. LVFace-B Glint360K wins both, and wins on every
|
||||
held-out film.
|
||||
|
||||
- :material-filter:{ .lg .middle } **[Whole vs. cast-restricted gallery](gallery-scope.md)**
|
||||
|
||||
---
|
||||
|
||||
Restricting the matcher to a film's credited cast is a clean win on
|
||||
every axis (+3.3pp F1, less than a third the misIDs) — but isn't a
|
||||
shipped runtime feature yet.
|
||||
Restricting the matcher to a film's credited cast improves F1,
|
||||
recall, and misID rate at once, but is not a shipped runtime feature
|
||||
yet.
|
||||
|
||||
- :material-account-convert:{ .lg .middle } **[Does pose expansion help?](pose-expansion.md)**
|
||||
|
||||
---
|
||||
|
||||
A convincing training-set effect that didn't reproduce on 5 held-out
|
||||
films once two methodology bugs were caught and fixed. An honest null
|
||||
result, not a forced narrative.
|
||||
A training-set effect that did not reproduce on 5 held-out films once
|
||||
two methodology bugs in the comparison harness were found and fixed.
|
||||
|
||||
- :material-magnify-expand:{ .lg .middle } **[Deep dive: LVFace-B Glint360K](lvface-deep-dive.md)**
|
||||
|
||||
---
|
||||
|
||||
The held-out generalization gap, how the error budget decomposes
|
||||
(extinction bridging at hard cuts, X-Ray's scene-membership vs.
|
||||
on-screen-face ceiling), and the frames where the pipeline is right
|
||||
and the ground truth is wrong.
|
||||
The held-out generalization gap, the two mechanisms behind its errors,
|
||||
and every distinct case where it names someone outside the film's
|
||||
credited cast.
|
||||
|
||||
</div>
|
||||
|
||||
## The full technical log
|
||||
## Full experiment log
|
||||
|
||||
- **[Model bake-off + threshold re-tune](model-bakeoff.md)** —
|
||||
the complete experiment log behind the four pages above: the ROCm teardown
|
||||
deadlock root cause and fix, DE concurrency tuning, the full 16-combo
|
||||
results table, and every caveat. This is where the shipped
|
||||
[`src/config.hpp`](https://REPOLINK/src/config.hpp) defaults come from.
|
||||
- **[Optimizer experiments (prior round)](optimizer-experiments.md)** — the
|
||||
earlier scene-union-metric tuning pass, superseded by the per-second metric
|
||||
used in the bake-off but kept for the ground-truth/architecture background.
|
||||
- **[Service conversion (proposal)](service-conversion.md)** — design sketch
|
||||
for a native idle-GPU worker gated on screen lock, not yet built.
|
||||
- **[Full experiment log](model-bakeoff.md)**: the complete log behind the
|
||||
four pages above, including how replaying against cached embeddings
|
||||
inside the same KPN network makes a full model and configuration
|
||||
comparison practical, the full results table, and every caveat. This is
|
||||
where the shipped [`src/config.hpp`](https://REPOLINK/src/config.hpp)
|
||||
defaults come from.
|
||||
- **[Service conversion (proposal)](service-conversion.md)**: design
|
||||
sketch for a native idle-GPU worker gated on screen lock, not yet built.
|
||||
|
||||
## Reproducing the benchmarks
|
||||
|
||||
Gallery `.h5` files, embedding dumps, the X-Ray corpus, montage frame images,
|
||||
and DE trajectories are not committed to this repository — they're pushed to
|
||||
the Gitea package registry and pulled on demand:
|
||||
Gallery `.h5` files, embedding dumps, the X-Ray corpus, montage frame
|
||||
images, and DE trajectories are not committed to this repository. They are
|
||||
pushed to the Gitea package registry and pulled on demand:
|
||||
|
||||
```bash
|
||||
scripts/artifacts/pull_artifacts.sh galleries
|
||||
@@ -86,4 +87,5 @@ scripts/artifacts/pull_artifacts.sh montage-frames <film-slug>
|
||||
```
|
||||
|
||||
See [`scripts/artifacts/push_artifacts.sh`](https://REPOLINK/scripts/artifacts/push_artifacts.sh)
|
||||
for the upload side (requires a `GITEA_TOKEN` with package write scope).
|
||||
for the upload side, which requires a `GITEA_TOKEN` with package write
|
||||
scope.
|
||||
|
||||
+233
-130
@@ -1,52 +1,53 @@
|
||||
# Deep dive: LVFace-B Glint360K
|
||||
|
||||
LVFace won the model bake-off (see [Which model is best?](best-model.md)) and is
|
||||
the shipped default embedder. This page is the honest accounting of how it
|
||||
actually performs — what a good second looks like, where the errors actually
|
||||
come from, and two cases where the ground truth itself is wrong and LVFace is
|
||||
right.
|
||||
LVFace won the model comparison (see [Which model is best?](best-model.md))
|
||||
and is the shipped default embedder. This page reports how it performs in
|
||||
detail: a baseline of correct output, the two mechanisms behind its errors,
|
||||
and every distinct case where it names someone who is not in the film's
|
||||
credited cast.
|
||||
|
||||
Read [How we score against X-Ray](methodology.md) first. X-Ray's ground truth
|
||||
is scene-level, not per-frame. A name marked correct in the Offscreen column
|
||||
below is the pipeline correctly reporting scene membership, not a workaround.
|
||||
|
||||
!!! note "How to read the frames on this page"
|
||||
|
||||
The top is the film frame, with a box and name on every face the pipeline
|
||||
identified. The bottom panels are the per-second verdict against X-Ray:
|
||||
**Onscreen** lists faces named in the frame, **Offscreen** lists cast
|
||||
X-Ray marks present in the scene without a visible face — presence
|
||||
carried by the tracker's windows, not by a detection. Colors are the
|
||||
score: <span style="color:#0ca30c">**green**</span> = correct (TPI),
|
||||
<span style="color:#eb6834">**orange**</span> = wrong (FPI),
|
||||
<span style="color:#3987e5">**blue**</span> = missed (FN).
|
||||
The top of each image is the film frame, with a box and name on every
|
||||
face the pipeline matched to a real detection. The panels below are the
|
||||
per-second result against X-Ray. **Onscreen** lists names attached to a
|
||||
visible face this second. **Offscreen** lists names the pipeline reports
|
||||
present without a currently visible face. Colors mark the verdict:
|
||||
<span style="color:#0ca30c">**green**</span> correct (TPI),
|
||||
<span style="color:#eb6834">**orange**</span> wrong (FPI),
|
||||
<span style="color:#3987e5">**blue**</span> missed (FN).
|
||||
|
||||
## What good looks like
|
||||
## Baseline: correctly scored seconds
|
||||
|
||||

|
||||
|
||||
Six faces on screen, all six named correctly — including Penelope Wilton at the
|
||||
edge of the pews and a half-occluded Michelle Dockery — while thirteen more
|
||||
cast members X-Ray marks present in the scene are correctly carried as
|
||||
"Offscreen" by their presence windows. One miss in the whole frame: Maggie
|
||||
Smith (blue). Score for this second: 0.86.
|
||||
Six faces on screen, all six named correctly, including Penelope Wilton at
|
||||
the edge of the pews and a partly occluded Michelle Dockery. Thirteen more
|
||||
cast members X-Ray lists as present in the scene are correctly reported
|
||||
Offscreen. One miss: Maggie Smith (blue). Score for this second: 0.86.
|
||||
|
||||

|
||||
|
||||
The same film's funeral gathering: mourning dress, hats, half the faces turned.
|
||||
**Nineteen of the twenty cast X-Ray lists for this scene are scored correctly**
|
||||
— seven named on screen at up to 100% confidence, twelve more correctly held
|
||||
as present off-screen.
|
||||
|
||||
And the pipeline doesn't need the face to be *real*:
|
||||
The same film's funeral scene: dark clothing, hats, half the faces turned
|
||||
away. Nineteen of the twenty cast members X-Ray lists for this scene score
|
||||
correct: seven named on screen at up to 100% confidence, twelve more reported
|
||||
correctly as present but not visible.
|
||||
|
||||

|
||||
|
||||
That's Herbie Hancock at 98% — as a face on a *screen inside the movie*, over a
|
||||
sci-fi HUD overlay, during a video call in Valerian. A face is a face, whether
|
||||
it's in the room or on the bridge's comms display.
|
||||
The pipeline does not require a live face. This is Herbie Hancock at 98%
|
||||
confidence, identified from a face displayed on a screen inside the film, on
|
||||
a video call under a science-fiction HUD overlay.
|
||||
|
||||
## Training vs. held-out: the generalization gap
|
||||
|
||||
The shipped config (`prob_threshold=0.754, anneal_sec=35.54,
|
||||
extinction_sec=57.43, expand_gallery=true`) was tuned against 4 films. Scored
|
||||
against the 5 films the optimizer never saw:
|
||||
The shipped config (`prob_threshold=0.754`, `anneal_sec=35.54`,
|
||||
`extinction_sec=57.43`, `expand_gallery=true`) was tuned on 4 films. Scored
|
||||
on the 5 films the optimizer never saw:
|
||||
|
||||

|
||||
|
||||
@@ -56,139 +57,241 @@ against the 5 films the optimizer never saw:
|
||||
| Lovelace | 77.5% | 90.3% | 67.9% | 14990 | 1085 | 58 | 7085 |
|
||||
| Valerian and the City of a Thousand Planets | 74.1% | 97.1% | 60.0% | 18663 | 548 | 0 | 12467 |
|
||||
| Downton Abbey: A New Era | 56.2% | 97.8% | 39.4% | 52027 | 1173 | 0 | 80084 |
|
||||
| **The Many Saints of Newark** | **46.3%** | **54.7%** | 40.1% | 15922 | 4394 | **974** | 23791 |
|
||||
| **macro average** | **67.4%** | 85.8% | 57.0% | | | | |
|
||||
| The Many Saints of Newark | 46.3% | 54.7% | 40.1% | 15922 | 4394 | 974 | 23791 |
|
||||
| macro average | 67.4% | 85.8% | 57.0% | | | | |
|
||||
|
||||
**67.4% held-out vs. 75.3% on training** — an ~8pp drop, and a **37pp spread
|
||||
between the best and worst held-out film**. The config does not generalize
|
||||
uniformly, and the spread traces to two mechanisms, both visible frame by
|
||||
frame below.
|
||||
The `P` column is misID-weighted (each out-of-film name counts 10x in the
|
||||
denominator; see [methodology](methodology.md#precision-recall-and-the-misid-weighting)).
|
||||
That weighting is why Many Saints reads 54.7% here despite naming mostly real,
|
||||
present faces: its raw (unweighted) precision is **78.4%**, and the gap is
|
||||
entirely its 974 misIDs paying the 10x penalty. The three zero-misID films
|
||||
(Benny & Joon, Downton, Valerian) have identical weighted and raw precision;
|
||||
Lovelace, with 58 misIDs, sits 3pp below its raw 93.3%.
|
||||
|
||||
## Mechanism 1: extinction bridging — usually right, wrong at hard cuts
|
||||
Held-out F1 is 67.4%, against 75.3% on training, an 8pp drop. The spread
|
||||
between the best and worst held-out film is 37pp. This is not unique to
|
||||
LVFace: [the full experiment log](model-bakeoff.md#held-out-validation-all-3-models)
|
||||
shows mbf and r18 with the same shape of spread on the same films, at a
|
||||
uniformly lower level. Two mechanisms explain the spread. Both are shown
|
||||
below with frame-level evidence.
|
||||
|
||||
The extinction window keeps an identity alive through seconds where no face is
|
||||
detectable. **Most of the time this is exactly what you want**, and it's where
|
||||
a lot of the TPI count comes from:
|
||||
## Mechanism 1: extinction bridging
|
||||
|
||||
The extinction window keeps a name reported as present for up to
|
||||
`extinction_sec` after its last real detection. This is deliberate: most
|
||||
gaps in face visibility are short (a turned head, an occlusion, a cut to a
|
||||
reaction shot), and the window bridges them.
|
||||
|
||||

|
||||
|
||||
Lovelace's polygraph scene: only Eric Roberts and Amanda Seyfried have visible
|
||||
faces, but X-Ray lists eight cast present — and all eight score green, the
|
||||
other six correctly carried by presence windows through a scene where the
|
||||
camera never shows them. A perfect second, and the extinction/anneal machinery
|
||||
is *why*.
|
||||
Lovelace's polygraph scene: only Eric Roberts and Amanda Seyfried have
|
||||
visible faces. X-Ray lists eight cast members present. All eight score
|
||||
correct; the other six are reported Offscreen through a stretch where the
|
||||
camera never shows them. The extinction window is why.
|
||||
|
||||
The same mechanism has a failure case: a hard cut into long faceless footage.
|
||||
Both Many Saints of Newark (974 misIDs) and Downton Abbey (FN=80084, the worst
|
||||
recall of the five) are dominated by it — verified directly against the raw
|
||||
per-frame stream and the HDF5 dump's own detection counts, not inferred from
|
||||
the score alone. **This is not a malfunction**: the tracker is doing exactly
|
||||
what its window is for; the footage just stops cooperating. In the debug
|
||||
overlay (which draws a bridged identity's last-known bbox, unlike the shipped
|
||||
output, which emits presence windows and no boxes at all) the bridged state is
|
||||
visible spatially:
|
||||
|
||||

|
||||
*Debug-overlay rendering (`dump_error_frames.py --raw`): "Jon Bernthal", "Joey
|
||||
Diaz" and "Billy Magnussen" are extinction-bridged identities from the previous
|
||||
shot, drawn frozen over the wall and the hanging plates. Frame
|
||||
`many_saints/fpi/fpi_t03543.jpg`, `montage-frames` artifact package.*
|
||||
|
||||
The cost is measurable, not just visible. Downton Abbey's hard cut into its
|
||||
closing credits, plotting the dump's own per-second `face_count` (detector
|
||||
output, independent of the tracker) against what the tracker reports:
|
||||
The same mechanism fails at a hard cut into a long stretch with no faces at
|
||||
all. Downton Abbey's recall (39.4%, the worst of the five held-out films) is
|
||||
dominated by this failure. It is verified directly against the raw
|
||||
per-frame stream and the dump's own detection counts, not inferred from the
|
||||
score. Plotting the dump's per-second `face_count` (detector output,
|
||||
independent of the tracker) against what the tracker reports, through
|
||||
Downton Abbey's hard cut into its closing credits:
|
||||
|
||||

|
||||
|
||||
From the cut onward the detector sees **zero faces for nearly a minute** — and
|
||||
the tracker keeps reporting the last shot's 15 identities the whole time
|
||||
(verified for Hugh Bonneville: bbox `(1743.2, 0.0, 171.3, 317.8)`, unchanged to
|
||||
the pixel, at every sampled second for 57+ seconds). The staircase at the right
|
||||
edge is the extinction window expiring actor by actor. That plateau is
|
||||
`SceneTrackerFunc::active_[actor_idx].last_bbox`
|
||||
From the cut onward the detector reports zero faces for close to a minute.
|
||||
The tracker continues reporting the previous shot's 15 identities for the
|
||||
same span (verified for Hugh Bonneville: bbox `(1743.2, 0.0, 171.3, 317.8)`,
|
||||
unchanged to the pixel, at every sampled second for 57 seconds). The
|
||||
staircase at the right edge is the extinction window expiring, actor by
|
||||
actor. This is `SceneTrackerFunc::active_[actor_idx].last_bbox`
|
||||
([`src/nodes/scene_tracker_node.hpp`](https://REPOLINK/src/nodes/scene_tracker_node.hpp))
|
||||
re-emitted as designed: `extinction_sec=57.4` was tuned long because bridging
|
||||
wins on most footage (see the polygraph frame above) — the training films just
|
||||
never contained a faceless stretch long enough to show the cost side, and the
|
||||
held-out set did.
|
||||
re-emitted as designed. `extinction_sec=57.4` was tuned long because
|
||||
bridging is correct on most footage, as in the polygraph scene above. The
|
||||
training films did not contain a faceless stretch long enough to expose the
|
||||
cost side; the held-out set did.
|
||||
|
||||
The same track-continuation machinery has one milder spatial artifact, worth
|
||||
knowing when reading these frames:
|
||||
The extinction window is a scoring concept, not something drawn on screen.
|
||||
The shipped output is presence windows with no bounding boxes. Even the
|
||||
debug overlay used for this report never draws a box for a bridged name: a
|
||||
name inside its extinction window with no current detection appears only as
|
||||
a name in the Offscreen column, the same as every correctly bridged name
|
||||
above.
|
||||
|
||||
A related, smaller effect shows up at rapid cuts:
|
||||
|
||||

|
||||
*Café Society (a training film), a shot/reverse-shot dialog: that is Steve
|
||||
Carell wearing both his own label and Jesse Eisenberg's.*
|
||||
|
||||
At a rapid cut, the previous shot's track can linger for a beat at nearly the
|
||||
same screen position the new face occupies — here Jesse Eisenberg's box from
|
||||
the counter-shot lands on Steve Carell. Note what the score panel says,
|
||||
though: both actors are green, because both *are* present in this dialog
|
||||
scene per X-Ray. The spatial label is briefly wrong; the per-second presence
|
||||
claim — the thing the pipeline actually ships — is right. It's the same trade
|
||||
as the extinction window: track continuation smooths over cuts, and 1 fps
|
||||
sampling occasionally catches the seam.
|
||||
Café Society (a training film), a shot/reverse-shot dialog. The box on Steve
|
||||
Carell's face carries two labels: his own, and Jesse Eisenberg's, left over
|
||||
from the counter-shot a moment earlier. Both names score correct, because
|
||||
both actors are present in this scene per X-Ray. The box position is
|
||||
briefly wrong; the presence claim, which is what the pipeline ships, is
|
||||
right.
|
||||
|
||||
## Mechanism 2: the face-vs-presence ceiling
|
||||
|
||||
Downton Abbey's recall didn't collapse because faces were misread — it
|
||||
collapsed because for most of its 80084 FN-seconds there was **no face to
|
||||
read**:
|
||||
Downton Abbey's recall did not collapse because faces were misread. It
|
||||
collapsed because for most of its 80084 false-negative seconds there was no
|
||||
face to read.
|
||||
|
||||

|
||||
|
||||
A newsreel crew hauls equipment through the hall: X-Ray credits 22 cast as
|
||||
present in this scene; not one face looks at the camera. Eight are still
|
||||
scored green (windows bridging from adjacent shots) — the other fourteen are
|
||||
blue FNs that no face-recognition pipeline could ever recover. X-Ray encodes
|
||||
*scene membership*; the pipeline measures *on-screen faces*. In ensemble films
|
||||
those two definitions diverge massively, and that gap — not identification
|
||||
error — is most of what the FN column counts.
|
||||
A newsreel crew moves equipment through the hall. X-Ray credits 22 cast
|
||||
members as present in this scene. None face the camera. Eight still score
|
||||
correct, carried by presence windows from adjacent shots. The other fourteen
|
||||
are missed, and no face-recognition system can recover them, because there
|
||||
is no face in the frame. X-Ray records scene membership; the pipeline
|
||||
measures visible faces. In ensemble scenes these two quantities diverge, and
|
||||
that gap accounts for most of the false-negative count.
|
||||
|
||||

|
||||
## Every distinct out-of-cast name
|
||||
|
||||
Same ceiling from the other side: Michela De Rossi in frame but turned away,
|
||||
five cast correctly bridged as offscreen (green), four blue FNs — and one
|
||||
orange we'll come back to below.
|
||||
Many Saints of Newark has the largest misID count of any held-out film: 974
|
||||
seconds, weighted. Rather than characterize this from a single frame, the
|
||||
raw replay stream was searched directly for every name the pipeline reports
|
||||
that is not in the film's credited cast. The same search was run on all 9
|
||||
films in the benchmark, one rule applied uniformly: **find the first second
|
||||
each distinct out-of-cast name appears, and render that exact second.**
|
||||
|
||||
Five films produce no such name anywhere in their runtime: Benny & Joon,
|
||||
Café Society, Downton Abbey, Sound of Metal, Valerian. Zero out-of-cast
|
||||
names across their entire length. Four films produce nine distinct names
|
||||
between them, shown below in full, not a sample.
|
||||
|
||||
### The Many Saints of Newark: 4 names
|
||||
|
||||

|
||||
|
||||
Germar Terrell Gardner, t=848s, 78% confidence. A real, clearly visible
|
||||
background actor. He is not in X-Ray's cast list for this film, but he is
|
||||
credited in Jellyfin's independent cast metadata (see
|
||||
[Where LVFace beat X-Ray](#where-lvface-beat-x-ray) below). This is a
|
||||
ground-truth gap, not a model error.
|
||||
|
||||

|
||||
|
||||
Archie Yates, t=2521s, 78% confidence. A real detected face, a genuine
|
||||
lookalike confusion.
|
||||
|
||||

|
||||
|
||||
Zooey Deschanel, t=2819s, 99% confidence. A real detected face at a dinner
|
||||
table, high-confidence lookalike confusion.
|
||||
|
||||

|
||||
|
||||
Talia Balsam, t=4551s, 93% confidence. A real detected face. Talia Balsam
|
||||
plays Mrs. Jarecki, a guidance counselor, in this film; she is confirmed
|
||||
on screen by direct inspection of the frame. She does not appear in X-Ray's
|
||||
`people.csv` for this title. This is a second ground-truth gap in the same
|
||||
film, not a model error.
|
||||
|
||||
Two of these four names are ground-truth gaps (Gardner, Balsam), not
|
||||
misidentifications. The other two (Yates, Deschanel) are genuine embedding
|
||||
errors on real faces.
|
||||
|
||||
### Lord of War: 3 names
|
||||
|
||||

|
||||
|
||||
David Shumbris, t=418s, 81% confidence. A real face in a dim, low-detail
|
||||
shot under a train track. A genuine lookalike confusion in poor lighting.
|
||||
|
||||

|
||||
|
||||
Ronald Reagan, t=1003s, 100% confidence. This is not a lookalike confusion.
|
||||
The detected face is a photograph of Reagan appearing within the shot, not a
|
||||
living actor. The detector and matcher both did their job correctly on the
|
||||
image content in front of them; the error is that a photograph inside the
|
||||
scene is not the same thing as an actor present in the scene, and the
|
||||
pipeline has no way to draw that distinction from a face crop alone.
|
||||
|
||||

|
||||
|
||||
Lance Reddick, t=6424s, 78% confidence. A small, distant, low-detail face at
|
||||
the edge of frame. A marginal, low-confidence lookalike confusion.
|
||||
|
||||
### Lovelace: 1 name
|
||||
|
||||

|
||||
|
||||
Chloë Sevigny, t=2451s, 100% confidence. Two boxes are drawn on the same
|
||||
face: one correctly labeled Amanda Seyfried, one incorrectly labeled Chloë
|
||||
Sevigny, both at 100%. A single detection producing two competing high-
|
||||
confidence identities on the same crop.
|
||||
|
||||
### Scarface: 1 name
|
||||
|
||||

|
||||
|
||||
Kirstie Alley, t=2451s, 89% confidence. Al Pacino is correctly identified in
|
||||
the foreground at 100%; a background face in the same shot is wrongly
|
||||
labeled Kirstie Alley. (The t=2451s here and the Lovelace Chloë Sevigny case
|
||||
above landing on the identical second is a genuine coincidence, verified from
|
||||
each film's raw stream by [`first_fpi_frames.py`](https://REPOLINK/scripts/docs/first_fpi_frames.py),
|
||||
not a transcription slip, two unrelated films whose *first* out-of-cast name
|
||||
happens to fall at the same timestamp.)
|
||||
|
||||
### Summary of the nine
|
||||
|
||||
| film | name | t (s) | confidence | classification |
|
||||
|---|---|---|---|---|
|
||||
| Many Saints of Newark | Germar Terrell Gardner | 848 | 78% | ground-truth gap |
|
||||
| Many Saints of Newark | Archie Yates | 2521 | 78% | lookalike confusion |
|
||||
| Many Saints of Newark | Zooey Deschanel | 2819 | 99% | lookalike confusion |
|
||||
| Many Saints of Newark | Talia Balsam | 4551 | 93% | ground-truth gap |
|
||||
| Lord of War | David Shumbris | 418 | 81% | lookalike confusion |
|
||||
| Lord of War | Ronald Reagan | 1003 | 100% | photo-in-frame |
|
||||
| Lord of War | Lance Reddick | 6424 | 78% | lookalike confusion, marginal |
|
||||
| Lovelace | Chloë Sevigny | 2451 | 100% | lookalike confusion |
|
||||
| Scarface | Kirstie Alley | 2451 | 89% | lookalike confusion |
|
||||
|
||||
Of nine distinct out-of-cast names across four films, two are ground-truth
|
||||
gaps, one is a photograph misread as a person, and six are genuine
|
||||
embedding-space confusions on real detected faces. None trace to extinction
|
||||
bridging: every one of these nine is a fresh detection on a real face crop
|
||||
at the second it first appears.
|
||||
|
||||
## Where LVFace beat X-Ray
|
||||
|
||||
Not every orange in these frames is actually wrong.
|
||||
Not every name marked wrong is actually wrong.
|
||||
[`scripts/optimizer/second_score.py`](https://REPOLINK/scripts/optimizer/second_score.py)
|
||||
scores strictly against X-Ray — but X-Ray itself has holes, and the pipeline
|
||||
found two kinds.
|
||||
scores strictly against X-Ray, and X-Ray has gaps of its own.
|
||||
|
||||

|
||||
|
||||
Germar Terrell Gardner — a real, clean, high-confidence detection — is counted
|
||||
as an out-of-cast misID because he doesn't appear in X-Ray's `people.csv` for
|
||||
The Many Saints of Newark at all. But Jellyfin's independent cast metadata
|
||||
*does* credit him for this exact film (cross-checked via
|
||||
`experiments/manifests/jellyfin_casts.json` from the `experiment-data` artifact
|
||||
package, a completely separate data source from X-Ray). That's also him in
|
||||
orange in the frame above — every one of those "errors" is the pipeline being
|
||||
right about a person X-Ray forgot.
|
||||
Germar Terrell Gardner, the same name from the table above, does not appear
|
||||
in X-Ray's `people.csv` for The Many Saints of Newark. Jellyfin's
|
||||
independent cast metadata does credit him for this film (cross-checked
|
||||
against `experiments/manifests/jellyfin_casts.json` from the
|
||||
`experiment-data` artifact package, a data source entirely separate from
|
||||
X-Ray). Talia Balsam is the same case: confirmed on screen, absent from
|
||||
X-Ray's cast list for this title.
|
||||
|
||||

|
||||
|
||||
And it isn't only uncredited bit-parts. That is **Robert Patrick** — top-billed
|
||||
in Lovelace, unmistakably on screen, reading his newspaper, identified at
|
||||
100% — scored orange because X-Ray's people-in-scene list for *this scene*
|
||||
doesn't include him. The identification is flawless; the ground truth missed
|
||||
an actor sitting in the middle of the frame.
|
||||
This extends past uncredited background actors. This is Robert Patrick,
|
||||
top-billed in Lovelace, clearly on screen reading a newspaper, identified at
|
||||
100%. The frame is scored wrong because X-Ray's people-in-scene list for
|
||||
this specific scene omits him, despite crediting him elsewhere in the film.
|
||||
The identification is correct; the ground truth is missing an entry.
|
||||
|
||||
This doesn't mean every flagged misID is secretly correct — Many Saints'
|
||||
974-count total is still overwhelmingly extinction bridging at cuts, not
|
||||
uncredited cameos. But the X-Ray corpus is a convenient, large-scale ground
|
||||
truth, not a perfect one, and the misID/FPI numbers in these tables carry an
|
||||
irreducible noise floor from ground-truth gaps in both directions.
|
||||
X-Ray is a large, convenient ground truth. It is not a complete one. The
|
||||
misID and FPI counts reported throughout this document include some fixed
|
||||
amount of noise from gaps in X-Ray itself, in both directions.
|
||||
|
||||
## Summary
|
||||
|
||||
LVFace is the right default: it wins the model comparison outright, it names
|
||||
19 of 20 correctly across a hat-heavy funeral crowd, and it recognises a face
|
||||
on a screen inside the movie. Its error budget decomposes into two understood
|
||||
mechanisms — extinction bridging at hard cuts (a tunable trade, not a bug) and
|
||||
the face-vs-presence ceiling baked into X-Ray's semantics — plus a nonzero
|
||||
slice where the pipeline is right and the ground truth is wrong. The held-out
|
||||
generalization gap (75.3% → 67.4%) is real and should be treated as the honest
|
||||
expected performance, not the training-set number.
|
||||
LVFace wins the model comparison on every held-out film. It correctly names
|
||||
19 of 20 people in a crowded funeral scene and correctly identifies a face
|
||||
displayed on a screen inside the film. Its errors resolve into two
|
||||
mechanisms: extinction bridging, which is correct on most footage and fails
|
||||
specifically at hard cuts into long faceless stretches, and the
|
||||
face-versus-presence ceiling, where X-Ray credits scene membership for
|
||||
people whose faces never appear on screen. Of the nine distinct
|
||||
out-of-cast identifications found across the benchmark, two trace to gaps in
|
||||
X-Ray's own cast data, one is a photograph misread as a person, and six are
|
||||
genuine lookalike confusions on real faces. The held-out generalization gap,
|
||||
75.3% training to 67.4% held-out, is real and should be treated as the
|
||||
expected operating point, not the training-set figure.
|
||||
|
||||
@@ -0,0 +1,134 @@
|
||||
# How we score against X-Ray
|
||||
|
||||
Every number in this report, every F1 and misID count, comes from one
|
||||
comparison. The comparison has a mismatch at its core that shapes nearly
|
||||
every finding in this report: the ground truth is scene-level, the
|
||||
pipeline's output is per-second, and the two do not mean the same thing.
|
||||
This page documents that comparison once, so the findings pages can rely on
|
||||
it without re-explaining it.
|
||||
|
||||
## What Amazon X-Ray records
|
||||
|
||||
X-Ray ships three tables per film: `scenes.csv` (a list of `[start, end]`
|
||||
timespans), `people_in_scenes.csv` (which actors are credited in each
|
||||
scene), and `people.csv` (actor identities). There is no per-frame or
|
||||
per-second annotation anywhere in X-Ray. A scene might run 45 seconds, and
|
||||
X-Ray records one cast list for the entire span, not "on screen from
|
||||
second 12 to second 30."
|
||||
|
||||
To compare this against per-second predictions, `second_score.py` expands
|
||||
every scene into per-second ground truth by copying the whole scene's cast
|
||||
list onto every second inside it:
|
||||
|
||||
```python
|
||||
for sn, (t0, t1) in spans.items():
|
||||
cast = scene_cast.get(sn, [])
|
||||
for t in range(int(t0), int(t1)):
|
||||
timeline[t] = cast
|
||||
```
|
||||
|
||||
That is the entire mechanism. If X-Ray credits five actors to a 30-second
|
||||
scene, all five count as ground truth present for all 30 seconds, including
|
||||
seconds where only one of them is on screen. This is not a simplification
|
||||
introduced by the pipeline; it is the only reading of X-Ray's data that is
|
||||
possible, because X-Ray itself does not record anything finer-grained.
|
||||
|
||||
## Why an offscreen name can be scored correct
|
||||
|
||||
A name listed under Offscreen with a correct (green) label is not the
|
||||
pipeline guessing or padding its score. It is the pipeline correctly
|
||||
answering the question X-Ray actually asks: is this actor part of this
|
||||
scene. It answers that question using a presence window (`[start, end]`,
|
||||
held open across cuts by `anneal_sec` and `extinction_sec`), which matches
|
||||
X-Ray's scene-level semantics more closely than a raw per-frame detection
|
||||
would.
|
||||
|
||||
A system that only reported "this actor is visible in this exact frame"
|
||||
would score worse against X-Ray's scene-level ground truth, producing a
|
||||
false negative every time the camera cuts away from a character who is
|
||||
still present in the scene. Not because it is wrong about the world, but
|
||||
because it would be answering a stricter, different question than the one
|
||||
X-Ray's data supports. The presence-window design exists specifically to
|
||||
answer X-Ray's actual question.
|
||||
|
||||
## What this resolves and what it does not
|
||||
|
||||
This resolves the semantic mismatch between a scene and an instant. It does
|
||||
not resolve two other limitations, both discussed in the
|
||||
[LVFace deep dive](lvface-deep-dive.md).
|
||||
|
||||
**The face-vs-presence ceiling.** X-Ray credits scene membership regardless
|
||||
of whether a face is ever visible: background crew, characters shot from
|
||||
behind, voice-only presence. No amount of bridging recovers a face that
|
||||
never appears on screen. This is a hard ceiling on recall, not a defect.
|
||||
|
||||
**Extinction bridging can overshoot.** The same presence-window mechanism
|
||||
that correctly answers "still in this scene" during a normal cut can also
|
||||
bridge across a scene boundary it has no way to detect. A hard cut into a
|
||||
different scene with no faces, such as closing credits, carries the
|
||||
previous scene's identities forward until the window expires. This is the
|
||||
mechanism behind Downton Abbey's recall collapse, documented in the deep
|
||||
dive.
|
||||
|
||||
## Precision, recall, and the misID weighting
|
||||
|
||||
Per sampled second `t`:
|
||||
|
||||
**TPI** (true positive instances): actors both X-Ray and the pipeline agree
|
||||
are present.
|
||||
|
||||
**FPI** (false positive instances): actors the pipeline reports that are
|
||||
not in X-Ray's cast for this second. Split into two categories:
|
||||
|
||||
- **FPI_incast**: the actor is in the film's cast, just not credited to
|
||||
this particular scene. A timing or boundary slip.
|
||||
- **FPI_misid**: the actor is not in the film's cast at all. A genuine
|
||||
wrong-identity error, weighted 10x in the precision objective, because
|
||||
naming someone who is not even in the film is a categorically worse
|
||||
error than a few seconds of scene-boundary slop.
|
||||
|
||||
!!! note "Every headline `P` and `F1` is misID-weighted"
|
||||
|
||||
The precision reported throughout this report, and therefore the F1
|
||||
derived from it, puts each `FPI_misid` into the denominator **10 times**
|
||||
(`precision = TPI / (TPI + FPI_incast + 10·FPI_misid)`,
|
||||
[`second_score.py`](https://REPOLINK/scripts/optimizer/second_score.py)).
|
||||
This is deliberate: the whole point is to punish naming an out-of-film
|
||||
actor far harder than a scene-boundary slip. But it means the `P` column
|
||||
is not raw precision, and a misID-heavy film's `P` is depressed
|
||||
super-linearly. `second_score.py` also emits an unweighted `precision_raw`
|
||||
(always ≥ the weighted `P`); where the gap matters, The Many Saints of
|
||||
Newark, weighted `P` 54.7% vs. raw 78.4%, the [LVFace deep dive](lvface-deep-dive.md)
|
||||
reports both. When comparing `P` across films, remember you are comparing a
|
||||
quantity that penalizes misIDs, not just a hit rate.
|
||||
|
||||
**FN** (false negatives): actors X-Ray lists that the pipeline never
|
||||
reports, counted only for actors who have a gallery reference embedding.
|
||||
Across the 9-film benchmark, coverage of X-Ray's credited cast ranges from
|
||||
20% to 79% by film (see
|
||||
[the full experiment log](model-bakeoff.md#gallery-coverage-per-film)); an
|
||||
actor with no reference photo can never be recognized regardless of model
|
||||
quality, and counting them as a miss would penalize gallery coverage, not
|
||||
recognition accuracy.
|
||||
|
||||
Two further numbers are reported alongside F1:
|
||||
|
||||
**agreement_rate**: mean per-second Jaccard overlap
|
||||
(`|Pred ∩ GT| / |Pred ∪ GT|`), partial credit. Naming 2 of 3 present actors
|
||||
scores 2/3, not 0.
|
||||
|
||||
**exact_match_rate**: the fraction of sampled seconds where the pipeline's
|
||||
named set exactly equals X-Ray's, no partial credit. Far harsher, and
|
||||
dominated by recall, since any single missed actor zeroes that second.
|
||||
|
||||
## Reproduce
|
||||
|
||||
```bash
|
||||
python3 scripts/optimizer/second_score.py \
|
||||
--pred pred.json --xray experiments/xray/.../<xray_dir> \
|
||||
--gallery experiments/galleries/gallery_LVFace-B_Glint360K.h5
|
||||
```
|
||||
|
||||
See also [the full experiment log](model-bakeoff.md) for how `pred.json` is
|
||||
produced, and the [LVFace deep dive](lvface-deep-dive.md) for what these
|
||||
mechanisms look like frame by frame.
|
||||
+271
-387
@@ -1,425 +1,322 @@
|
||||
# Model bake-off + threshold re-tune — experiment log (2026-07-18/19)
|
||||
# Full experiment log
|
||||
|
||||
Follow-on to [the prior optimizer round](optimizer-experiments.md), which used
|
||||
an older, since-superseded scene-union metric. This round uses the **per-second** metric
|
||||
([`scripts/optimizer/second_score.py`](https://REPOLINK/scripts/optimizer/second_score.py))
|
||||
and answers three questions in one 16-run matrix: which embedding model is best, does cast-restriction cut misIDs, and does
|
||||
per-film gallery expansion help.
|
||||
This page reports how the pipeline performs across three questions: which
|
||||
embedding model is best, whether restricting the gallery to a film's
|
||||
credited cast helps, and whether promoting confidently identified poses into
|
||||
a per-film gallery annex helps. It also documents the replay architecture
|
||||
that made testing all three questions in one pass practical, and every
|
||||
caveat needed to trust the numbers.
|
||||
|
||||
(This was the fourth optimizer campaign against the X-Ray corpus, so its on-disk
|
||||
artifacts carry an internal `rep4_` prefix — `experiments/results/rep4_best_*.json`,
|
||||
`experiments/trajectories/rep4_*.jsonl`, and the manifests referenced below. The
|
||||
earlier campaigns used the superseded scene-union metric and were discarded.)
|
||||
Read [How we score against X-Ray](methodology.md) first for what F1,
|
||||
precision, recall, and misID mean in this report. All numbers below use the
|
||||
per-second metric
|
||||
([`scripts/optimizer/second_score.py`](https://REPOLINK/scripts/optimizer/second_score.py)).
|
||||
|
||||
## Why this experiment, and what it actually delivered
|
||||
r50 (ArcFace w600k-R50) is excluded from the detailed comparison below. Its
|
||||
gallery was built with roughly 30% fewer reference images per actor than the
|
||||
other three models on the identical source photos (10808 vs 15055 total
|
||||
embeddings across the same 2418 actors), which confounds any direct
|
||||
comparison of its scores against the others. It remains in the
|
||||
[calibration curve comparison](best-model.md#first-signal-calibration-curves),
|
||||
which does not depend on the training benchmark.
|
||||
|
||||
Four goals going in, and an honest read on each after held-out validation (see
|
||||
below):
|
||||
## Why replay makes this affordable
|
||||
|
||||
1. **Find the best default parameters to ship.** Partially delivered. The DE optimum
|
||||
generalizes *unevenly* — strong on 3 of 5 held-out films, badly broken on 2 (one
|
||||
with a 974-count misID blowup). The tuned values are shipped anyway (see
|
||||
Caveats) because they still beat the old defaults on average, but this is not a
|
||||
settled, film-agnostic optimum.
|
||||
2. **Find the best default model.** Delivered with more confidence. LVFace beat
|
||||
r50/r18/mbf across all 4 training combos, and nothing in held-out validation
|
||||
contradicts the model choice specifically — the held-out failures trace to
|
||||
`extinction_sec`/threshold interactions and gallery coverage, not the embedder.
|
||||
3. **Provide insight into how the application works.** The strongest, most durable
|
||||
output. Found and fixed a real teardown deadlock bug (100% reproducible, not the
|
||||
assumed rare GPU flake), established a real concurrency ceiling (8 parallel
|
||||
replays, not more), and found a real parameter interaction (a strict
|
||||
`prob_threshold` "earns" a longer extinction window before it starts hurting).
|
||||
4. **Demonstrate limitations.** Delivered, and reinforced hard by held-out
|
||||
validation — see the "Held-out validation" section below for concrete examples,
|
||||
including a screenshot of the matcher naming 15 actors, none correctly, on a
|
||||
completely blank title card.
|
||||
Decoding video and running face detection, alignment, and embedding is the
|
||||
expensive part of this pipeline. Everything downstream of that (tracking,
|
||||
identity matching, scene aggregation) is cheap. KPN++'s node/network
|
||||
structure means those two stages are separate components connected by
|
||||
typed channels, so the expensive stage can run once per film, cache its
|
||||
output, and the cheap stage can be re-run against that cache as many times
|
||||
as needed with different Config values.
|
||||
|
||||
## TL;DR — what changed in `src/config.hpp`
|
||||
`scene_analyze --dump-embeddings out.h5` runs the expensive half once per
|
||||
film and writes per-frame face detections and embeddings to HDF5
|
||||
([`scripts/optimizer/SCHEMA.md`](https://REPOLINK/scripts/optimizer/SCHEMA.md)).
|
||||
[`scripts/optimizer/replay.py`](https://REPOLINK/scripts/optimizer/replay.py)
|
||||
then re-assembles the real C++ `face_tracker`, `identity_matcher`, and
|
||||
`scene_tracker` nodes into a Python-driven KPN network and replays a
|
||||
film's cached embeddings through them, varying `prob_threshold`,
|
||||
`anneal_sec`, `extinction_sec`, and `expand_gallery` freely. No GPU
|
||||
inference and no video decode happen during a replay; each one completes
|
||||
in seconds. This is what makes a 512-evaluation differential-evolution
|
||||
search per model, per gallery mode, per expansion setting, tractable, and
|
||||
what made the full held-out validation across three models in this report
|
||||
possible in one session rather than requiring three full re-encodes of the
|
||||
benchmark set.
|
||||
|
||||
| knob | old default | new default | why |
|
||||
| ---- | ----------- | ----------- | --- |
|
||||
| `arcface_model` | `arcface_w600k_r50.onnx` | **`LVFace-B_Glint360K.onnx`** | Best F1 in the full-gallery bake-off (75.3% vs r50's 68.5%). LVFace was worth its size. |
|
||||
| `prob_threshold` | 0.76 | **0.754** | Re-tuned for LVFace + per-second metric. |
|
||||
| `extinction_sec` | 1.5 | **57.4** | Reverses the earlier "short is better" finding — see below. |
|
||||
| `anneal_sec` | 10.0 | **35.5** | Same reversal; previously thought insensitive. |
|
||||
| `expand_gallery` | false | **true** | Helped recall on the full (unrestricted) gallery for the winning model — opposite of the earlier assumption. |
|
||||
`optimize.py` runs `differential_evolution` over this replay function as its
|
||||
objective, with DE-level parallelism (multiple candidate configs evaluated
|
||||
concurrently, each spawning its own replay subprocesses) on top of it. The
|
||||
practical ceiling on this machine's GPU was 8 concurrent replay processes;
|
||||
9 silently degraded every score to 0.0% (well-formed output, wrong numbers,
|
||||
not a crash), so `optimize.py` was run at `REPLAY_WORKERS=4 DE_WORKERS=2`.
|
||||
|
||||
These are the **`LVFace-B_Glint360K_full_exp`** winning values, applied to
|
||||
[`src/config.hpp`](https://REPOLINK/src/config.hpp) — the best result that
|
||||
uses only features already live in the running app (full gallery, no cast
|
||||
restriction; see below for why restricted mode isn't applied even though it scored
|
||||
higher).
|
||||
## Search space
|
||||
|
||||
## Why re-run at all
|
||||
`popsize=10, maxiter=15` per combo (3 parameters, up to 512 evaluations,
|
||||
usually stopping earlier on DE's convergence tolerance).
|
||||
`anneal_sec`/`extinction_sec` bounds were widened from 1-30/1-15 to 1-60/1-60
|
||||
partway through the sweep. r50's 4 combos finished before the widening and
|
||||
used the old, narrower bounds; this is one more reason r50 is excluded from
|
||||
direct comparison here.
|
||||
|
||||
[The prior round](optimizer-experiments.md)'s scene-union metric hid out-of-cast false positives
|
||||
behind a gallery∩cast recall mask, and the earlier 9-film benchmark was scene-level
|
||||
(union over a whole X-Ray scene), not a fair per-timepoint comparison. SESSION_STATE
|
||||
flagged the old scene-metric bake-off numbers (R50≈LVFace≈MBF ~85%) as superseded.
|
||||
This round uses `second_score.py`: uniform per-second sampling, GT = X-Ray scene's
|
||||
cast at time *t*, pred = actors whose presence window covers *t*, FPI weighted 10×
|
||||
when the named actor isn't in the film's cast at all (true misID) vs. an in-cast
|
||||
timing slip. FN only counts gallery-known cast (fair recall — 67% of X-Ray cast have
|
||||
no reference embedding; see
|
||||
[the prior round's gallery-coverage-gap analysis](optimizer-experiments.md#the-gallery-coverage-gap)).
|
||||
## Training films and held-out films
|
||||
|
||||
## The deadlock that was blocking all of this
|
||||
9 films have dumped embeddings across all 4 models. 4 were used for
|
||||
optimization:
|
||||
|
||||
Every replay in this line of work goes through
|
||||
[`scripts/optimizer/replay.py`](https://REPOLINK/scripts/optimizer/replay.py),
|
||||
which runs the real C++ tracker/matcher/scene_tracker nodes inside a
|
||||
Python-assembled KPN network. Before this session, every subprocess replay **timed out at 45s, 100% of
|
||||
the time** — not the documented ~20-30% ROCm rocBLAS-GEMM driver flake, but a plain
|
||||
logic bug: `replay.py`'s CLI called `replay(net, ..., stop=False)` to *skip*
|
||||
`net.stop()` (trying to dodge the GEMM deadlock), planning to `os._exit(0)`
|
||||
immediately after. But:
|
||||
- Café Society (62-cast)
|
||||
- Lord of War (64-cast)
|
||||
- Scarface (67-cast)
|
||||
- Sound of Metal (14-cast)
|
||||
|
||||
- `PyNode::stop()` ([`include/kpn/python/bindings.hpp`](https://KPNLINK/include/kpn/python/bindings.hpp)
|
||||
in the KPN++ submodule) is the *only*
|
||||
code that sets `stop_flag_ = true` before joining the node's worker thread.
|
||||
- The source node's `run_loop()` has `while (!stop_flag_)` as its only exit
|
||||
condition (it has no input channels, so it never sees a channel-closed signal
|
||||
either).
|
||||
- Skipping `stop()` meant `stop_flag_` never became true. When `replay()` returned,
|
||||
its local `net` went out of scope immediately, running `~PyNetwork` → `~PyNode` →
|
||||
`thread_.join()` **synchronously inside `replay()`'s own call frame** — before
|
||||
`main()` ever got control back to run `os._exit(0)`.
|
||||
5 were held out, never seen by any optimizer run:
|
||||
|
||||
Root-caused via `gdb -p <pid> -batch -ex "thread apply all bt"` on a hung process:
|
||||
the main thread was stuck in `~PyNode`'s `jthread::join()`; the worker thread was in
|
||||
an ordinary `time.sleep()` inside the Python source callback, waiting for a stop
|
||||
signal that was never sent. The two HSA `kfd_wait_on_events` threads visible in the
|
||||
same trace are normal ROCm runtime housekeeping, not evidence of a wedged GPU kernel.
|
||||
- Benny & Joon
|
||||
- Downton Abbey: A New Era
|
||||
- Lovelace
|
||||
- The Many Saints of Newark
|
||||
- Valerian and the City of a Thousand Planets
|
||||
|
||||
**Fix:** `replay.py` now calls `replay(..., stop=True)` (the removed `stop=False` +
|
||||
`os._exit` workaround was actively harmful). Verified 3/3 clean runs at ~8s each
|
||||
(down from a guaranteed 45s timeout), and a full DE sweep producing real, sensible
|
||||
F1/precision/recall instead of flat 0.0%.
|
||||
## Gallery coverage per film
|
||||
|
||||
## Concurrency tuning
|
||||
The gallery has reference embeddings for 2418 actors, but coverage of any
|
||||
given film's credited cast varies widely. This was previously reported as
|
||||
one flat number (67% of X-Ray cast lacking a reference embedding, averaged
|
||||
across the whole benchmark); the per-film breakdown is:
|
||||
|
||||
With the deadlock fixed,
|
||||
[`scripts/optimizer/optimize.py`](https://REPOLINK/scripts/optimizer/optimize.py)
|
||||
was extended with DE-level parallelism —
|
||||
`differential_evolution(..., workers=ThreadPoolExecutor.map)` — so multiple
|
||||
population candidates evaluate concurrently, each spawning its own per-film replay
|
||||
subprocesses (`REPLAY_WORKERS`). Total concurrent GPU replay processes ≈
|
||||
`DE_WORKERS × REPLAY_WORKERS`.
|
||||
| film | cast credited | in gallery | coverage |
|
||||
|---|---|---|---|
|
||||
| Lord of War | 64 | 13 | 20.3% |
|
||||
| Scarface | 67 | 15 | 22.4% |
|
||||
| The Many Saints of Newark | 48 | 13 | 27.1% |
|
||||
| Café Society | 62 | 17 | 27.4% |
|
||||
| Lovelace | 42 | 15 | 35.7% |
|
||||
| Valerian and the City of a Thousand Planets | 36 | 13 | 36.1% |
|
||||
| Benny & Joon | 23 | 12 | 52.2% |
|
||||
| Downton Abbey: A New Era | 36 | 22 | 61.1% |
|
||||
| Sound of Metal | 14 | 11 | 78.6% |
|
||||
|
||||
| concurrent replays | result |
|
||||
| --- | --- |
|
||||
| 3 (`REPLAY_WORKERS=3`, no DE parallelism) | baseline, GPU underutilised |
|
||||
| 6 (`DE_WORKERS=2 × REPLAY_WORKERS=3`) | clean, real scores, ~1 isolated timeout per run |
|
||||
| 8 (`DE_WORKERS=2 × REPLAY_WORKERS=4`, 4-film manifest) | clean, real scores |
|
||||
| 9 (`DE_WORKERS=3 × REPLAY_WORKERS=3`) | **broken** — every replay blew past the 45s timeout, all scores silently degraded to 0.0% |
|
||||
Two training films (Lord of War, Scarface) have the worst coverage in the
|
||||
set, 20-22%. Their training-set F1 numbers below are partly capped by
|
||||
missing references, not purely by model quality. Downton Abbey has 61%
|
||||
coverage, the second-best in the benchmark, yet the worst held-out recall
|
||||
of any film (39.4%, LVFace). Its recall problem is not primarily a coverage
|
||||
problem; it is the extinction-bridging failure documented in the
|
||||
[LVFace deep dive](lvface-deep-dive.md#mechanism-1-extinction-bridging).
|
||||
Reproduce with `scripts/docs/gallery_coverage_per_film.py`.
|
||||
|
||||
9 concurrent replays looks like valid output (well-formed JSON, a real number) while
|
||||
actually being garbage — a dangerous failure mode, not a crash. **8 concurrent is the
|
||||
practical ceiling** on this GPU (gfx1100) for this workload. The matrix ran at
|
||||
`REPLAY_WORKERS=4 DE_WORKERS=2`.
|
||||
## Training results, 3 models × 2 gallery modes × 2 expansion settings
|
||||
|
||||
## Training films and validation set
|
||||
Ranked by F1. misid = FPI_misid, the count of true wrong-actor
|
||||
identifications (naming someone not in the film's cast at all), distinct
|
||||
from FPI, which also includes in-cast timing slips.
|
||||
|
||||
9 films total have dumped embeddings across all 4 models. 4 were used for
|
||||
optimization, leaving 5 held out for validation:
|
||||
|
||||
- **Lord of War** (64-cast, "clean")
|
||||
- **Scarface** (67-cast, "ensemble/lookalike")
|
||||
- **Sound of Metal** (14-cast, "high gallery-coverage")
|
||||
- **Café Society** (62-cast, added this round — similar ensemble size to Scarface but
|
||||
different genre/lighting; picked to add diversity, not genre-overlap, over
|
||||
Downton Abbey or The Many Saints of Newark)
|
||||
|
||||
Held out: Benny & Joon, Downton Abbey: A New Era, Lovelace, The Many Saints of
|
||||
Newark, Valerian and the City of a Thousand Planets. The training-set numbers below are
|
||||
training-set fit — see "Held-out validation" further down for the real
|
||||
generalization test.
|
||||
|
||||
## Search space and DE settings
|
||||
|
||||
`popsize=10, maxiter=15` (3 params → ≤480 evals/combo ceiling; DE's `tol` convergence
|
||||
usually stops earlier). `anneal_sec`/`extinction_sec` bounds were **widened from
|
||||
1–30/1–15 to 1–60/1–60 mid-run** (see below) — the 4 `arcface_w600k_r50` combos
|
||||
finished before the widening and still use the old, narrower bounds, so they are
|
||||
**not directly comparable** to the other 12 on those two params. Re-running r50 with
|
||||
the wider bounds was deferred (diminishing-returns judgment call, not yet done).
|
||||
|
||||
## Results — all 16 combos (4 models × {full, restricted} × {expand, noexp})
|
||||
|
||||
Ranked by F1. `misid` = FPI_misid, count of true wrong-actor identifications (an
|
||||
actor named who isn't in the film's cast at all) — distinct from `FPI`, which
|
||||
includes in-cast timing slips.
|
||||
Each combo's row is its best **full-coverage** evaluation: the highest-F1 DE
|
||||
evaluation in which all 4 training films replayed without a timeout (see
|
||||
[Dropped-film scoring](#a-scoring-bug-worth-recording-dropped-film-evaluations)
|
||||
below for why this qualifier is load-bearing and not the same as `argmax F1`
|
||||
over the raw sweep).
|
||||
|
||||
| combo | F1 | P | R | TPI | FPI | misid | FN |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| LVFace-B_Glint360K_restricted_exp | **78.3%** | 91.0% | 68.9% | 42830 | 3782 | 60 | 19492 |
|
||||
| LVFace-B_Glint360K_restricted_exp | 78.3% | 91.0% | 68.9% | 42830 | 3782 | 60 | 19492 |
|
||||
| LVFace-B_Glint360K_restricted_noexp | 76.7% | 91.5% | 66.2% | 41149 | 3400 | 59 | 21173 |
|
||||
| arcface_w600k_mbf_restricted_exp | 76.5% | 90.7% | 66.3% | 41270 | 4431 | **0** | 21052 |
|
||||
| arcface_w600k_mbf_restricted_exp | 76.2% | 90.0% | 66.2% | 64328 | 7480 | 0 | 33234 |
|
||||
| arcface_r18_restricted_exp | 75.5% | 87.6% | 66.5% | 41399 | 5666 | 60 | 20923 |
|
||||
| **LVFace-B_Glint360K_full_exp** | **75.3%** | 89.7% | 65.4% | 47757 | 3407 | 232 | 26966 |
|
||||
| LVFace-B_Glint360K_full_exp | 75.3% | 89.7% | 65.4% | 47757 | 3407 | 232 | 26966 |
|
||||
| arcface_w600k_mbf_restricted_noexp | 75.0% | 91.1% | 63.9% | 39752 | 3465 | 60 | 22570 |
|
||||
| arcface_w600k_mbf_full_noexp | 74.2% | 87.4% | 64.4% | 12645 | 1312 | 57 | 6985 |
|
||||
| arcface_r18_restricted_noexp | 73.5% | 91.3% | 61.7% | 38299 | 3220 | 60 | 24023 |
|
||||
| LVFace-B_Glint360K_full_noexp | 72.4% | 94.2% | 58.9% | 27077 | 1725 | **0** | 19506 |
|
||||
| LVFace-B_Glint360K_full_noexp | 72.3% | 88.3% | 61.8% | 40363 | 3503 | 244 | 25850 |
|
||||
| arcface_w600k_mbf_full_exp | 72.0% | 87.7% | 61.4% | 39875 | 3729 | 240 | 26338 |
|
||||
| arcface_w600k_r50_full_noexp † | 71.6% | 96.7% | 56.9% | 22471 | 361 | 45 | 17012 |
|
||||
| arcface_w600k_r50_restricted_exp † | 71.1% | 96.5% | 56.4% | 34954 | 1119 | 15 | 27368 |
|
||||
| arcface_w600k_r50_restricted_noexp † | 69.2% | 97.9% | 53.6% | 21146 | 327 | 15 | 18337 |
|
||||
| arcface_w600k_mbf_full_noexp | 71.0% | 93.2% | 57.9% | 41699 | 2472 | 56 | 33024 |
|
||||
| arcface_r18_full_exp | 69.1% | 87.6% | 57.7% | 37342 | 3119 | 242 | 28871 |
|
||||
| arcface_w600k_r50_full_exp † | 68.5% | 94.0% | 54.1% | 34982 | 903 | 150 | 31231 |
|
||||
| arcface_r18_full_noexp | 66.6% | 91.3% | 53.1% | 34314 | 2362 | 107 | 31899 |
|
||||
|
||||
† old, narrower anneal/extinction bounds (see above) — not directly comparable to
|
||||
the other 12 on those two params.
|
||||

|
||||
|
||||
The same 16 results as a picture — the two headline effects are visible without
|
||||
reading a single row: filled (restricted) dots stack the top of the ranking for
|
||||
every model color, and yellow (LVFace) leads within both scopes:
|
||||
The two clearest patterns: every model's best-scoring combo uses the
|
||||
restricted gallery, and LVFace leads within both gallery modes. `full_exp`
|
||||
(the shipped combination) is the best-scoring option that uses only
|
||||
features the running application currently supports; restriction is not
|
||||
wired into the application yet (see
|
||||
[Whole vs. cast-restricted gallery](gallery-scope.md)).
|
||||
|
||||

|
||||
### A scoring bug worth recording: dropped-film evaluations
|
||||
|
||||
## Calibration curves — discriminative power, independent of the threshold
|
||||
The numbers above are corrected ones. The raw `rep4_best_*.json` files, and an
|
||||
earlier version of this table, reported a different `arcface_w600k_mbf_full_noexp`
|
||||
row: **74.2% F1 at TPI 12645**, a third the TPI of every sibling combo. That was
|
||||
not a better config; it was an artifact of how the optimizer aggregates.
|
||||
|
||||
Each model's gallery carries a fitted Platt sigmoid `P(match | sim) = σ(a·sim + b)`
|
||||
(embedded directly in the gallery HDF5, see
|
||||
[`src/gallery/gallery_calibration.hpp`](https://REPOLINK/src/gallery/gallery_calibration.hpp)).
|
||||
Plotting all four side by side shows discriminative power directly, independent of
|
||||
whatever `prob_threshold` a particular run happened to use:
|
||||
`optimize.py` builds each candidate's score from only the films whose replay
|
||||
subprocess returned (`per_film = [m for m in ex.map(_one, films) if m is not
|
||||
None]`), then **averages** F1/precision/recall and **sums** TPI/FPI/misID over
|
||||
just those survivors. When a film's replay times out (the sweep ran near the
|
||||
8-process concurrency ceiling, so this happened intermittently), that film
|
||||
silently drops from both. A candidate whose hardest film timed out is therefore
|
||||
scored on an easier subset, and differential evolution, maximizing that score,
|
||||
will happily converge onto exactly such a candidate. For `mbf_full_noexp` the
|
||||
reported winner was one of 7 evaluations (out of 512) whose TPI had collapsed to
|
||||
a partial-film subset; its median-coverage evaluations sit around 51686 TPI.
|
||||
|
||||

|
||||
The fix here was to re-derive each combo's best row from its DE trajectory
|
||||
(`experiments/trajectories/rep4_*.jsonl`), keeping only evaluations within 30% of
|
||||
that combo's median TPI (full 4-film coverage) before taking the best F1. This
|
||||
needs no re-running, the honest best configuration was already in the sweep,
|
||||
just not the one `argmax F1` selected. Three combos moved: `mbf_full_noexp`
|
||||
74.2% → **71.0%**, `LVFace_full_noexp` 72.4% → **72.3%** (and its misID, 0 → 244,
|
||||
was itself a dropped-film artifact), `mbf_restricted_exp` 76.5% → **76.2%**. The
|
||||
shipped LVFace `full_exp` winner was unaffected, its reported evaluation already
|
||||
had full coverage (TPI 47757 ≈ median). `experiment_charts.py` applies the same
|
||||
`clean_best` filter, so every figure on this page matches the corrected table.
|
||||
The underlying `optimize.py` aggregation is also being fixed so a dropped-film
|
||||
evaluation can never be selected as a winner again.
|
||||
|
||||
LVFace-B has both the steepest curve (`a=17.7`, vs. 15.3–16.2 for the ArcFace
|
||||
variants) and the lowest P=0.5 decision boundary (similarity 0.23 vs. 0.27–0.31) —
|
||||
it separates same-actor from different-actor pairs more confidently at a lower
|
||||
similarity, consistent with it winning the full-gallery F1 comparison below.
|
||||
Generated by
|
||||
[`scripts/docs/calibration_chart.py`](https://REPOLINK/scripts/docs/calibration_chart.py)
|
||||
(requires each gallery to have been calibrated at least once — run any replay
|
||||
against it first).
|
||||
### Per-film training breakdown
|
||||
|
||||
## Two effects in isolation: gallery scope, and pose expansion
|
||||
The 75.3% LVFace training figure is a macro average across 4 films, not a
|
||||
uniform result:
|
||||
|
||||
The matrix crosses two independent variables — averaging across all 4 models
|
||||
isolates each one from model choice:
|
||||
|
||||
**Gallery scope (whole 2418-actor gallery vs. restricted to the film's credited
|
||||
cast)** — averaged over both expansion settings and all 4 models:
|
||||
|
||||
| scope | F1 | P | R | total misID (16 evals→8 each) |
|
||||
| film | LVFace F1 | mbf F1 | r18 F1 | best model |
|
||||
|---|---|---|---|---|
|
||||
| full | 71.2% | 91.1% | 59.0% | 1073 |
|
||||
| **restricted** | **74.5%** | 92.2% | **62.9%** | **329** |
|
||||
| Café Society | 68.1% | 62.2% | 60.1% | LVFace |
|
||||
| Lord of War | 75.6% | 77.2% | 75.6% | mbf |
|
||||
| Scarface | 71.5% | 68.6% | 64.1% | LVFace |
|
||||
| Sound of Metal | 78.8% | 76.5% | 71.6% | LVFace |
|
||||
|
||||
Restriction wins outright on every axis — not a precision/recall trade, a clean
|
||||
win: **+3.3pp F1, +3.9pp recall, and less than a third the total misIDs.** Fewer
|
||||
LVFace does not win every training film. mbf scores higher on Lord of War
|
||||
(77.2% vs 75.6%). LVFace's own training-film range is 68.1% to 78.8%, a
|
||||
10.7pp spread, smaller than the 37pp spread seen on held-out films but real.
|
||||
Reproduce with `scripts/docs/run_holdout_all_models.py --films training`.
|
||||
|
||||
## Held-out validation, all 3 models
|
||||
|
||||
The training matrix above is training-set fit. Each model's own tuned
|
||||
`full_exp` config was replayed against the 5 held-out films, scored the
|
||||
same way:
|
||||
|
||||
| film | LVFace F1 | mbf F1 | r18 F1 |
|
||||
|---|---|---|---|
|
||||
| Benny & Joon | 83.0% | 78.5% | 77.1% |
|
||||
| Lovelace | 77.5% | 73.7% | 72.2% |
|
||||
| Valerian and the City of a Thousand Planets | 74.1% | 70.2% | 71.0% |
|
||||
| Downton Abbey: A New Era | 56.2% | 55.0% | 53.0% |
|
||||
| The Many Saints of Newark | 46.3% | 44.5% | 42.1% |
|
||||
| **macro average** | **67.4%** | **64.4%** | **63.1%** |
|
||||
|
||||
LVFace scores highest on every one of the 5 held-out films; the ranking
|
||||
never flips. Total misIDs across the 5 films: LVFace 1032, mbf 2197, r18
|
||||
1224. LVFace has less than half mbf's misID count while also scoring
|
||||
higher on every film. This directly confirms the model choice out of
|
||||
sample; it is not inferred from the training numbers alone. See the
|
||||
[LVFace deep dive](lvface-deep-dive.md) for frame-level detail on where and
|
||||
why LVFace still fails on the two worst films. Reproduce with
|
||||
`scripts/docs/run_holdout_all_models.py`.
|
||||
|
||||
## Two effects in isolation: gallery scope and pose expansion
|
||||
|
||||
Averaging across the 3 compared models (r50 excluded) isolates each variable
|
||||
from model choice.
|
||||
|
||||
**Gallery scope**, averaged over both expansion settings and all 3 models
|
||||
(6 evaluations per row):
|
||||
|
||||
| scope | F1 | P | R | total misID |
|
||||
|---|---|---|---|---|
|
||||
| full | 71.1% | 89.6% | 59.6% | 1121 |
|
||||
| restricted | 75.9% | 90.4% | 65.6% | 299 |
|
||||
|
||||
Restriction improves every metric at once. This is not a precision/recall
|
||||
trade: +4.8pp F1, +6.0pp recall, and roughly a quarter the misIDs. Fewer
|
||||
candidates in the matcher's search space means fewer opportunities for a
|
||||
look-alike false match, and (per the recall gain) doesn't cost real detections.
|
||||
This is the single cleanest signal in the whole matrix — stronger than the model
|
||||
choice itself — which is exactly why cast-restriction becoming a real runtime
|
||||
feature (not just an optimizer trick) is the top item in Caveats below.
|
||||
lookalike false match, and the recall gain shows this does not cost real
|
||||
detections. Restriction is currently an offline optimizer technique, not a
|
||||
runtime feature of the application; see
|
||||
[Whole vs. cast-restricted gallery](gallery-scope.md) for what building it
|
||||
into the application would require.
|
||||
|
||||
**Pose expansion (promoting a confidently-identified track's novel-pose views into
|
||||
a per-film gallery annex — [`src/gallery/track_gallery.hpp`](https://REPOLINK/src/gallery/track_gallery.hpp))**
|
||||
is smaller and interacts with
|
||||
scope rather than acting independently:
|
||||
**Pose expansion** (promoting a confidently identified track's novel-pose
|
||||
views into a per-film gallery annex,
|
||||
[`src/gallery/track_gallery.hpp`](https://REPOLINK/src/gallery/track_gallery.hpp)):
|
||||
|
||||
| scope | expansion | F1 | R | misID |
|
||||
|---|---|---|---|---|
|
||||
| full | off | 71.2% | 58.3% | 209 |
|
||||
| full | **on** | 71.2% | 59.7% | **864** |
|
||||
| restricted | off | 73.6% | 61.3% | 194 |
|
||||
| restricted | **on** | **75.4%** | **64.5%** | 135 |
|
||||
| full | off | 70.0% | 57.6% | 407 |
|
||||
| full | on | 72.1% | 61.5% | 714 |
|
||||
| restricted | off | 75.1% | 63.9% | 179 |
|
||||
| restricted | on | 76.7% | 67.2% | 120 |
|
||||
|
||||
In **restricted** mode, expansion is a clean win (+1.8pp F1, +3.2pp recall, misID
|
||||
actually *drops*) — the annex only ever competes against the film's own ~15-actor
|
||||
cast, so a "confidently identified, new pose" view is unlikely to be mistaken for
|
||||
someone else. In **full** mode, expansion buys essentially nothing on F1 (71.2% →
|
||||
71.2%, recall +1.4pp) while **quadrupling misIDs** (209 → 864): a novel-pose view
|
||||
promoted into the annex now competes against the whole 2418-actor gallery, so a
|
||||
"confident" identity is confident against the wrong universe of candidates — the
|
||||
expansion mechanism is "learning" a pose correctly, but the enlarged evidence pool
|
||||
makes it easier for that learned pose to look like a plausible match for a
|
||||
different actor. **Practical takeaway: gallery expansion should be paired with
|
||||
cast restriction, not used on the full gallery** — the version currently shipped
|
||||
as default (`full_exp`, see TL;DR) sits in the worse of these four cells for this
|
||||
specific knob, even though it's the best available combo without cast-restriction
|
||||
support in the app yet (see Caveats).
|
||||
In restricted mode, expansion is a clean win: +1.6pp F1, +3.3pp recall,
|
||||
misID drops. The annex only competes against the film's own roughly 15-actor
|
||||
cast, so a new pose of a known actor is unlikely to be confused with someone
|
||||
else. In full mode, expansion buys +2.1pp F1 and +3.9pp recall but at a real
|
||||
cost: misID rises from 407 to 714 as the same new-pose view now competes
|
||||
against the full 2418-actor gallery, where a confidently learned pose is more
|
||||
likely to match the wrong person. On the full gallery it is a recall-vs-misID
|
||||
trade, not a free gain. This training-set effect
|
||||
did not reproduce on held-out data; see
|
||||
[Does pose expansion help?](pose-expansion.md) for the full held-out test
|
||||
and the two methodology bugs caught while checking it.
|
||||
|
||||
## What the data says
|
||||
## Calibration curves
|
||||
|
||||
- **LVFace was worth its size.** It wins full-gallery mode outright (75.3% vs r50's
|
||||
68.5%, r18's 69.1%, mbf's 72.0%) with the highest recall of any full-mode combo —
|
||||
the earlier scene-union-metric conclusion ("not worth it") doesn't survive the
|
||||
better metric.
|
||||
- **Cast-restriction is a consistent, broad win.** Every model's best combo is
|
||||
`restricted`. It isn't just precision-safe: `arcface_w600k_mbf_restricted_exp` and
|
||||
`LVFace-B_Glint360K_full_noexp` both hit **misid=0** — zero true wrong-actor
|
||||
identifications. But restriction is an **offline optimizer technique, not a live
|
||||
app feature** — it pre-filters each film's gallery to its Jellyfin-credited cast
|
||||
before the matcher ever runs; there's no runtime "restrict to this film's cast"
|
||||
switch in the app today. Implementing it for real is future work, tracked
|
||||
separately from this defaults update.
|
||||
- **Gallery expansion (`expand_gallery`) is mode-dependent.** It helps on
|
||||
`restricted` galleries (smaller, so novel-pose promotion adds real signal) and on
|
||||
LVFace's full gallery, but **hurts** r50 and mbf in full mode (compare
|
||||
`arcface_w600k_r50_full_exp` 68.5% vs `full_noexp` 71.6%). Don't assume it's a free
|
||||
win — model- and mode-dependent.
|
||||
- **arcface_r18 (smallest/cheapest) is last across all 4 modes** — model capacity
|
||||
matters here, this isn't just parameter-count padding.
|
||||
- **`anneal_sec`/`extinction_sec` kept pinning at the search ceiling.** With the
|
||||
original 1–30/1–15 bounds, 3 of 4 r50 combos landed at ~93-98% of the upper bound.
|
||||
Widened to 1–60/1–60 mid-run (after the r50 combos had already finished) — every
|
||||
subsequent combo's best config landed at ~90%+ of the *new* ceiling too (e.g. the
|
||||
LVFace winner: `ann=59.2, ext=59.2`, both ~99% of 60). The likely mechanism: a
|
||||
strict `prob_threshold` "earns" a long extinction/anneal window — once false
|
||||
matches are rare, a long window just bridges real presence gaps (occlusion, turned
|
||||
face) instead of smearing false positives into later scenes, which is what made
|
||||
short windows look better under the old, laxer thresholds. **Open question, not
|
||||
resolved**: does this keep climbing past 60s, or does it actually plateau there?
|
||||
Decided not to chase further this round (diminishing-returns judgment call) — flag
|
||||
for a future sweep if it matters.
|
||||
Each gallery carries a fitted Platt sigmoid `P(match | sim) = σ(a·sim + b)`,
|
||||
stored directly in the gallery HDF5
|
||||
([`src/gallery/gallery_calibration.hpp`](https://REPOLINK/src/gallery/gallery_calibration.hpp)).
|
||||
This measures discriminative power independent of whatever
|
||||
`prob_threshold` a given run used:
|
||||
|
||||
The ceiling-pinning is visible in the raw search itself. Every one of the 512
|
||||
DE evaluations for the winning combo, plotted over the
|
||||
`prob_threshold` × `extinction_sec` plane:
|
||||

|
||||
|
||||
LVFace has the steepest curve (`a=17.7` vs 15.3-16.2 for the ArcFace
|
||||
variants) and the lowest P=0.5 decision boundary (similarity 0.23 vs
|
||||
0.27-0.31), separating same-actor from different-actor pairs more
|
||||
confidently at a lower similarity than any ArcFace variant tested,
|
||||
including r50. Generated by
|
||||
[`scripts/docs/calibration_chart.py`](https://REPOLINK/scripts/docs/calibration_chart.py).
|
||||
|
||||
## Extinction and anneal window search
|
||||
|
||||
Every one of the 512 DE evaluations for the winning LVFace `full_exp`
|
||||
combo, plotted over the `prob_threshold` × `extinction_sec` plane:
|
||||
|
||||

|
||||
|
||||
The dark band hugging the top edge *is* the finding: nearly everything scoring
|
||||
well sits at `extinction_sec` ≥ 50, across a wide range of thresholds, and the
|
||||
population converged into a dense cloud around the optimum (threshold ~0.70–0.80,
|
||||
extinction pinned at the 60s bound). Short extinction windows (bottom half) are
|
||||
uniformly pale — under a strict threshold there is simply no good configuration
|
||||
down there. Generated by
|
||||
[`scripts/docs/experiment_charts.py`](https://REPOLINK/scripts/docs/experiment_charts.py)
|
||||
from the DE trajectories (`experiments/trajectories/*.jsonl`, part of the
|
||||
`experiment-data` artifact package).
|
||||
Nearly everything scoring well sits at `extinction_sec` above 50, across a
|
||||
wide range of thresholds. Short extinction windows are uniformly weaker:
|
||||
under a strict threshold, there is no good configuration in that region of
|
||||
the search space. The optimizer converged with `anneal_sec=59.2,
|
||||
extinction_sec=59.2`, about 99% of the widened 60s bound, which raises an
|
||||
open question not resolved in this round: does performance keep improving
|
||||
past 60s, or does it plateau there. Not chased further this pass.
|
||||
|
||||
## Held-out validation — the number that actually matters
|
||||
## Caveats
|
||||
|
||||
The 16-combo matrix above is training-set fit. This is the real test: the shipped
|
||||
config (`LVFace-B_Glint360K_full_exp` — `prob_threshold=0.754, anneal_sec=35.5,
|
||||
extinction_sec=57.4, expand_gallery=true`) replayed against the **5 films never seen
|
||||
by the optimizer** (Benny & Joon, Downton Abbey: A New Era, Lovelace, The Many
|
||||
Saints of Newark, Valerian and the City of a Thousand Planets), scored the same way.
|
||||
|
||||
| film | F1 | P | R | agree | TPI | FPI | misid | FN |
|
||||
|---|---|---|---|---|---|---|---|---|
|
||||
| Benny & Joon | 83.0% | 89.1% | 77.7% | 72.4% | 15125 | 1846 | 0 | 4337 |
|
||||
| Lovelace | 77.5% | 90.3% | 67.9% | 72.1% | 14990 | 1085 | 58 | 7085 |
|
||||
| Valerian and the City of a Thousand Planets | 74.1% | 97.1% | 60.0% | 58.8% | 18663 | 548 | 0 | 12467 |
|
||||
| Downton Abbey: A New Era | 56.2% | 97.8% | 39.4% | 40.6% | 52027 | 1173 | 0 | 80084 |
|
||||
| **The Many Saints of Newark** | **46.3%** | **54.7%** | 40.1% | 37.0% | 15922 | 4394 | **974** | 23791 |
|
||||
| **macro average (5 films)** | **67.4%** | 85.8% | 57.0% | 56.2% | 116727 | 9046 | 1032 | 127764 |
|
||||
|
||||

|
||||
|
||||
**67.4% held out vs. 75.3% on training** — an ~8pp drop, and a much more informative
|
||||
number than the training-set F1 alone: a **37pp spread between best and worst film**
|
||||
(83.0% vs 46.3%). The config does not generalize uniformly.
|
||||
|
||||
Two films are outright failure cases, and rendering bounding boxes + names on the
|
||||
extracted frames (`replay.py --raw-out` +
|
||||
[`dump_error_frames.py`](https://REPOLINK/scripts/optimizer/dump_error_frames.py)` --raw`, see
|
||||
Reproduce) turned what looked like a same-scene misidentification into something
|
||||
more precise and more damning:
|
||||
|
||||
- **The Many Saints of Newark** (mob-family drama, picked as a training-adjacent
|
||||
genre test) has **974 true misIDs** — far more than any training combo saw at any
|
||||
setting. The annotated frame below shows the same mechanism as Downton Abbey,
|
||||
at smaller scale: **"Jon Bernthal 100%", "Joey Diaz 100%", and "Billy Magnussen
|
||||
100%" are all frozen boxes over empty background — a blurred wall, hanging plates —
|
||||
with no face in them at all.** Only one real face in frame has a box, and it
|
||||
carries a *second*, colliding label ("Leslie Odom Jr." and "Michael Gandolfini"
|
||||
both at high confidence on the same box) — likely two tracks whose frozen bboxes
|
||||
happen to overlap.
|
||||
|
||||

|
||||
*Frame `many_saints/fpi/fpi_t03543.jpg` from the `montage-frames` artifact
|
||||
package (`scripts/artifacts/pull_artifacts.sh montage-frames
|
||||
Many_Saints_of_Newark`).*
|
||||
|
||||
- **Downton Abbey: A New Era** (large ensemble, 36-cast) has high precision (97.8%)
|
||||
but recall collapses to 39.4% (FN=80084, by far the largest of the 5). Its
|
||||
starkest failure happens where there is nothing to see at all: the film's hard
|
||||
cut into its closing credits, where **the matcher kept reporting 15 actors —
|
||||
all wrong — for nearly a minute of faceless screen.**
|
||||
|
||||
Both are the same mechanism, and it can be *measured*, not just screenshotted.
|
||||
Plotting the dump's own per-second `face_count` (detector output, independent
|
||||
of the tracker) against the number of actors the tracker reports, through
|
||||
Downton Abbey's cut to credits:
|
||||
|
||||

|
||||
|
||||
From the cut onward the detector sees **zero faces** — yet the tracker holds a
|
||||
perfectly flat plateau of 15 reported identities for 56 seconds, each with the
|
||||
*exact same bbox, unchanged to the pixel* (verified for Hugh Bonneville:
|
||||
`(1743.2, 0.0, 171.3, 317.8)` at every sampled second from 7222 through 7279+).
|
||||
The staircase on the right edge is the extinction window finally expiring,
|
||||
actor by actor. That plateau is `SceneTrackerFunc`'s
|
||||
`active_[actor_idx].last_bbox`
|
||||
([`src/nodes/scene_tracker_node.hpp`](https://REPOLINK/src/nodes/scene_tracker_node.hpp))
|
||||
being re-emitted
|
||||
unchanged — **the extinction state machine working exactly as coded**, not a
|
||||
bug in the logic. The film cuts from a packed group shot straight into ~40+
|
||||
seconds of blank titles/credits with zero faces, and `extinction_sec=57.4` is
|
||||
comfortably long enough to bridge that entire gap without expiring, so the
|
||||
tracker faithfully keeps reporting "last known position" for a cast that is no
|
||||
longer on screen at all.
|
||||
|
||||
This reframes the "long extinction window wins" DE-search pattern (see above): it
|
||||
isn't unambiguously good. It buys recall by bridging real gaps (occlusion, turned
|
||||
face) in some films, but on others — specifically, hard cuts into long faceless
|
||||
footage — it manufactures a frozen-bbox ghost the tracker has no way to verify,
|
||||
precisely the failure mode the *original* short-extinction-window default (`1.5s`)
|
||||
was chosen to avoid. The training-set films apparently didn't have a long enough
|
||||
faceless stretch after a confirmed identity to expose this; the held-out set did.
|
||||
|
||||
Frames for all three films (`benny_joon`, `many_saints`, `downton_abbey` — one strong
|
||||
performer, two failure cases) are under `experiments/results/holdout/frames/`
|
||||
(not committed — pull per film with `scripts/artifacts/pull_artifacts.sh
|
||||
montage-frames <film-slug>`), each
|
||||
with a `manifest.json` listing the bucket (`best`/`fpi`/`fn`), timestamp, and
|
||||
predicted vs. ground-truth actors for every dumped frame. Frames are annotated with
|
||||
bounding boxes + name/confidence (green = identified, orange = unknown), matching
|
||||
[`src/nodes/debug_renderer_node.hpp`](https://REPOLINK/src/nodes/debug_renderer_node.hpp)'s
|
||||
colour convention. Generated by
|
||||
`scripts/optimizer/dump_error_frames.py --raw <replay.py --raw-out output>` (see
|
||||
Reproduce).
|
||||
|
||||
`dump_error_frames.py --interval-sec 600` also supports a per-N-second sweep
|
||||
instead of the fixed best/fpi/fn buckets: one best (highest Jaccard) and one worst
|
||||
(lowest Jaccard) frame per 10-minute window across the whole film, e.g.
|
||||
`experiments/results/holdout/frames/many_saints_intervals/` (13 windows × 2 = 26
|
||||
frames for the ~2h Many Saints runtime) — a way to sample "how are we doing" evenly
|
||||
across a film's runtime rather than only at its most extreme seconds.
|
||||
|
||||
## Caveats / what this is not
|
||||
|
||||
- **r50's 4 combos used the old, narrower search bounds** and aren't fully
|
||||
comparable to the other 12 on `anneal_sec`/`extinction_sec`.
|
||||
- **The applied defaults use `full_exp`, not the higher-scoring `restricted_exp`**,
|
||||
because cast-restriction isn't a real runtime feature yet (see above). The
|
||||
78.3% F1 number is not what the shipped defaults will produce — 75.3% is.
|
||||
- **`full_exp` is the best full-gallery combo, but not the safest.** Per the
|
||||
isolated-effects analysis above, `expand_gallery=true` only cleanly pays off
|
||||
when paired with cast-restriction; on the full gallery it's flat on F1 while
|
||||
~4x-ing misIDs (209→864, averaged across models). `full_noexp` scores lower
|
||||
(72.4% vs 75.3% for LVFace) but with **zero** true misIDs and higher precision
|
||||
(94.2% vs 89.7%). Kept `full_exp` as shipped since it's the highest-F1 option
|
||||
available without cast-restriction, but this is a real F1-vs-safety trade, not
|
||||
a strictly-better choice — worth revisiting if misID rate matters more than
|
||||
the last few points of F1 for a given deployment.
|
||||
- **Switching the default model is an operational change, not just a config tweak**:
|
||||
any existing gallery built from r50 embeddings is incompatible with LVFace
|
||||
embeddings and needs rebuilding.
|
||||
- r50's 4 combos used the older, narrower search bounds (1-30/1-15 instead
|
||||
of 1-60/1-60) and are further confounded by its thinner gallery. Excluded
|
||||
from all comparisons above except calibration.
|
||||
- The shipped defaults use `full_exp` (75.3% training F1), not the
|
||||
higher-scoring `restricted_exp` (78.3%), because cast restriction is not
|
||||
a runtime feature of the application yet.
|
||||
- `expand_gallery` is mode-dependent, not a free win. Averaged across models
|
||||
on the full gallery it trades misIDs for recall (see the pose-expansion
|
||||
table). For LVFace specifically, though, `full_exp` beats `full_noexp` on
|
||||
every axis at once (F1 75.3 vs 72.3, precision 89.7 vs 88.3, recall 65.4 vs
|
||||
61.8, misID 232 vs 244), so the shipped `full_exp` is a clean choice for
|
||||
this model, not an F1-vs-safety trade. (An earlier version of this page
|
||||
reported `full_noexp` at 72.4% with zero misIDs and higher precision, which
|
||||
made it look like the safer option; that was the dropped-film artifact
|
||||
described above, not a real property of the config.)
|
||||
- Switching the default model is an operational change: any gallery built
|
||||
from a different model's embeddings must be rebuilt before the new
|
||||
default takes effect.
|
||||
|
||||
## Reproduce
|
||||
|
||||
```bash
|
||||
# 4-film matrix, all 4 models × 2 modes × 2 expansion settings
|
||||
# 4-film training matrix, all 4 models × 2 gallery modes × 2 expansion settings
|
||||
bash experiments/run_rep4_subprocess.sh
|
||||
|
||||
# single combo
|
||||
@@ -429,34 +326,21 @@ SAE_EXPAND=1 REPLAY_WORKERS=4 DE_WORKERS=2 python3 scripts/optimizer/optimize.py
|
||||
--params prob_threshold:0.5:0.999 anneal_sec:1:60 extinction_sec:1:60 \
|
||||
--popsize 10 --maxiter 15 --trajectory traj.jsonl --out best.json
|
||||
|
||||
# replay the shipped config against a held-out film — --raw-out is needed to draw
|
||||
# bboxes later (the merged pred.json has no per-frame bbox, only actor windows)
|
||||
python3 scripts/optimizer/replay.py \
|
||||
--dump experiments/dumps/LVFace-B_Glint360K/dump_<slug>.h5 \
|
||||
--gallery experiments/galleries/gallery_LVFace-B_Glint360K.h5 \
|
||||
--out pred.json --raw-out raw.jsonl --prob-threshold 0.754 --anneal-sec 35.54 \
|
||||
--extinction-sec 57.43 --expand-gallery
|
||||
# held-out validation, all 3 models, 5 films
|
||||
python3 scripts/docs/run_holdout_all_models.py --out docs_data/holdout_all_models.json
|
||||
|
||||
# regenerate the report's charts (16-combo ranking, DE landscape, held-out
|
||||
# per-film F1, Downton ghost timeline) from the artifacts under experiments/
|
||||
# per-film training breakdown, all 3 models, 4 films
|
||||
python3 scripts/docs/run_holdout_all_models.py --films training --out docs_data/training_per_film.json
|
||||
|
||||
# gallery coverage per film
|
||||
python3 scripts/docs/gallery_coverage_per_film.py --out docs_data/gallery_coverage_per_film.json
|
||||
|
||||
# regenerate this page's charts from experiments/ artifacts
|
||||
python3 scripts/docs/experiment_charts.py --out-dir docs/assets/images
|
||||
|
||||
# dump example frames (best-agreement / FPI / FN) for visual inspection, annotated
|
||||
# with bounding boxes + names (--raw is optional; omit for unannotated frames)
|
||||
python3 scripts/optimizer/dump_error_frames.py \
|
||||
--pred pred.json --raw raw.jsonl --xray experiments/xray/.../<xray_dir> \
|
||||
--movie "<path to source video>" \
|
||||
--gallery experiments/galleries/gallery_LVFace-B_Glint360K.h5 \
|
||||
--out-dir experiments/results/holdout/frames/<name> --n-per-bucket 4
|
||||
|
||||
# or: one best + one worst frame per 10-minute window across the whole film
|
||||
python3 scripts/optimizer/dump_error_frames.py \
|
||||
--pred pred.json --raw raw.jsonl --xray experiments/xray/.../<xray_dir> \
|
||||
--movie "<path to source video>" \
|
||||
--gallery experiments/galleries/gallery_LVFace-B_Glint360K.h5 \
|
||||
--out-dir experiments/results/holdout/frames/<name>_intervals --interval-sec 600
|
||||
# one frame per distinct out-of-cast name across all 9 films (used in the deep dive)
|
||||
python3 scripts/docs/first_fpi_frames.py
|
||||
```
|
||||
|
||||
See also: [the prior optimizer round](optimizer-experiments.md) (superseded
|
||||
metric) and the session log
|
||||
See also the session log
|
||||
[`experiments/SESSION_STATE.md`](https://REPOLINK/experiments/SESSION_STATE.md).
|
||||
|
||||
@@ -1,125 +0,0 @@
|
||||
# Threshold optimization against Amazon X-Ray — experiment log
|
||||
|
||||
Record of the July 2026 work that tuned the pipeline's recognition/tracking defaults
|
||||
against ground-truth per-scene actor presence, and the tooling built to do it.
|
||||
|
||||
## TL;DR — what changed
|
||||
|
||||
| knob | old default | new default | why |
|
||||
| ---- | ----------- | ----------- | --- |
|
||||
| `prob_threshold` | 0.99 | **0.76** | 0.99 was far too strict — halved recall for a fraction of a precision point. DE optimum, tightly converged. |
|
||||
| `extinction_sec` | 5.0 | **1.5** | Long extinction smears presence into later scenes → FPs. DE converged tightly low. |
|
||||
| `anneal_sec` | 10.0 | 10.0 (unchanged) | DE found it **insensitive** (F1 flat ±0.3pp across 3–26s) — kept the round default. |
|
||||
| `detector_conf` | 0.5 | 0.5 (unchanged) | Sweep showed raising it only trades recall for precision at a net F1 loss — near-threshold detections are real faces, not phantoms. |
|
||||
|
||||
Net effect on the 9-film benchmark (strict per-scene, augmented gallery):
|
||||
recall **58% → ~72%**, F1 **70% → ~76%**, precision ~85%, at no meaningful precision cost.
|
||||
|
||||
## Ground truth
|
||||
|
||||
Public scene-level **Amazon X-Ray** dataset (Zenodo DOI 10.5281/zenodo.17659734,
|
||||
CC-BY-4.0): per movie, `people.csv` (name_id/person/character), `scenes.csv`
|
||||
(scene/start/end ms), `people_in_scenes.csv`. Films matched to the library by an
|
||||
**authoritative Jellyfin ID join** (query `/Items?IncludeItemTypes=Movie&Fields=
|
||||
ProviderIds,Path`, join Imdb/Tmdb against X-Ray metadata) — NOT fuzzy title matching,
|
||||
which collides badly (TV episodes vs same-named films). 9 genuine films with source
|
||||
video on disk: Benny & Joon, Café Society, Downton Abbey: A New Era, Lord of War,
|
||||
Lovelace, The Many Saints of Newark, Scarface, Sound of Metal, Valerian.
|
||||
|
||||
## The scoring metric (evolved through review)
|
||||
|
||||
Comparison unit is the **X-Ray scene**, not sampled timepoints. For each scene
|
||||
`[start,end]`: predicted set = **union** of actors detected anywhere in the span;
|
||||
GT set = actors X-Ray lists for that scene. Per scene TP/FP/FN, then:
|
||||
|
||||
- **Precision: STRICT.** Any predicted actor not in the scene's X-Ray set is an FP,
|
||||
*including out-of-cast confusions* (no gallery∩cast masking). An earlier
|
||||
timepoint-sampled, cast-masked metric HID ~570 such FPs across 9 films and let the
|
||||
optimizer drive `prob_threshold` to the 0.50 floor — a metric artifact. Counting
|
||||
them is essential.
|
||||
- **Recall: FAIR.** FN counts only X-Ray cast members **who are in the gallery**. 67%
|
||||
of X-Ray cast (261/392) have no gallery reference embedding and can never be
|
||||
recognised — counting them as misses penalises coverage, not the threshold. Both
|
||||
`recall` (fair) and `recall_strict` (all) are reported.
|
||||
- **Aggregation:** per-scene F1 → **duration-weighted average within a movie** (long
|
||||
scenes count more) → **equal-weight mean across movies** (macro; each film counts
|
||||
the same regardless of length). This is the DE objective.
|
||||
|
||||
Implemented in `scripts/optimizer/scene_score.py` — since **removed** along
|
||||
with this metric; its per-second successor is
|
||||
[`scripts/optimizer/second_score.py`](https://REPOLINK/scripts/optimizer/second_score.py)
|
||||
(see the [bake-off round](model-bakeoff.md)).
|
||||
|
||||
## The gallery coverage gap
|
||||
|
||||
Diagnosing low recall: only **131 of 392** X-Ray cast were in the gallery (33%). Every
|
||||
in-gallery actor HAD embeddings (gallery well-formed) — the gap was pure coverage.
|
||||
[`scripts/optimizer/fetch_missing_actors.py`](https://REPOLINK/scripts/optimizer/fetch_missing_actors.py)
|
||||
recovers missing actors:
|
||||
`nm-id → TMDB /find external_ids → /person/{id}/images → download → embed (sae_embed)`,
|
||||
with a `--wikidata` fallback (P345→P18 Commons photo).
|
||||
|
||||
- **TMDB recovered 143/261** (55%). 0 face-detection failures; the rest had no TMDB
|
||||
person (60) or no profile photo (58). Coverage 33% → **70%**.
|
||||
- **Wikidata fallback: 0/118** of the TMDB failures — only 4 even had a Commons photo,
|
||||
none yielded a detectable face. → **TheTVDB not worth pursuing**: these remaining
|
||||
actors are obscure enough that no image source covers them, AND (see below) most are
|
||||
off-camera anyway.
|
||||
|
||||
**Coverage vs detectability.** Adding references lifted recall (58→68% at fixed config)
|
||||
but modestly. Per-film drill-down (Lord of War: 12 actors recovered, only 1 had a
|
||||
detectable on-camera face) showed most missing cast are a **detectability gap** — X-Ray
|
||||
credits them as cast-in-scene (incl. off-camera/background), but their face never
|
||||
appears clearly for the pipeline to detect. This is a fundamental ceiling of a
|
||||
face-recognition pipeline vs X-Ray's presence semantics, not a fixable gap.
|
||||
|
||||
## Optimizer
|
||||
|
||||
[`scripts/optimizer/optimize.py`](https://REPOLINK/scripts/optimizer/optimize.py)
|
||||
— scipy `differential_evolution` over the knob space,
|
||||
each candidate = full replay of all films through the **real** C++ nodes (see the
|
||||
KPN replay architecture below) scored by the metric above. Global objective (one
|
||||
config for all films, not per-film).
|
||||
|
||||
**Convergence stability (augmented gallery, 233 evals):**
|
||||
|
||||
| knob | top-20 range | verdict |
|
||||
| ---- | ------------ | ------- |
|
||||
| `prob_threshold` | 0.69–0.83 (σ 0.05) | TIGHT — trust 0.76 |
|
||||
| `extinction_sec` | 1.0–2.2 (σ 0.33) | TIGHT — trust 1.5 |
|
||||
| `anneal_sec` | 3.1–26.3 (σ 6.4) | LOOSE — insensitive, not hard-coded |
|
||||
|
||||
F1 varied only 0.3pp across the top-20 → objective is flat near the optimum, so only
|
||||
the tightly-converged knobs were adopted as defaults.
|
||||
|
||||
## Replay architecture (how the sweep is cheap)
|
||||
|
||||
The optimizer never re-decodes video. `scene_analyze --dump-embeddings out.h5` runs the
|
||||
expensive half once (decode→detect→align→embed) and dumps per-frame face embeddings
|
||||
+ metadata to HDF5 ([`scripts/optimizer/SCHEMA.md`](https://REPOLINK/scripts/optimizer/SCHEMA.md)).
|
||||
[`scripts/optimizer/replay.py`](https://REPOLINK/scripts/optimizer/replay.py) then
|
||||
replays that dump through the **real** C++ `face_tracker → identity_matcher →
|
||||
scene_tracker` assembled in a Python KPN network (`sae_kpn` nanobind module), varying
|
||||
Config knobs freely — no GPU embedding, no decode. Verified BYTE-EXACT against
|
||||
`scene_analyze`'s own output. The dumps are gallery-independent, so testing the
|
||||
augmented gallery needed no re-dump. `detector_conf` is replayable UPWARD only (the
|
||||
dump floor is 0.5).
|
||||
|
||||
## Reproduce
|
||||
|
||||
```bash
|
||||
# 1. dump (once per film, needs video)
|
||||
scene_analyze --movie <f> --gallery gallery.json --dump-embeddings dump.h5 --fps 1
|
||||
# 2. build films manifest by Jellyfin ID join (see scripts/optimizer notes)
|
||||
# 3. optimize
|
||||
python scripts/optimizer/optimize.py --manifest films.json --gallery gallery.json \
|
||||
--params prob_threshold:0.5:0.999 anneal_sec:1:30 extinction_sec:1:15 \
|
||||
--popsize 8 --maxiter 20 --trajectory traj.jsonl --out opt.json
|
||||
# 4. score a fixed config / validate on a held-out set
|
||||
# (historical: score_config.py and scene_score.py were removed with the
|
||||
# scene-union metric — use scripts/optimizer/second_score.py, per-second)
|
||||
python scripts/optimizer/second_score.py --help
|
||||
```
|
||||
|
||||
Superseded by the [model bake-off + re-tune](model-bakeoff.md), which
|
||||
replaced this round's scene-union metric with per-second scoring.
|
||||
+83
-73
@@ -1,42 +1,46 @@
|
||||
# Pose expansion: does "learning" new poses mid-film help?
|
||||
# Pose expansion: does promoting new poses mid-film help?
|
||||
|
||||
`expand_gallery` ([`src/gallery/track_gallery.hpp`](https://REPOLINK/src/gallery/track_gallery.hpp))
|
||||
promotes a confidently-identified
|
||||
track's novel-pose reference views into a per-film, in-memory gallery annex — the
|
||||
idea being that once the pipeline is sure who someone is, a pose it hasn't seen
|
||||
before (turned head, different lighting) becomes a free extra reference for
|
||||
recognising that actor again later in the same film, without touching the baked
|
||||
gallery.
|
||||
`expand_gallery`
|
||||
([`src/gallery/track_gallery.hpp`](https://REPOLINK/src/gallery/track_gallery.hpp))
|
||||
promotes a confidently identified track's novel-pose reference views into a
|
||||
per-film, in-memory gallery annex. The idea: once the pipeline is confident
|
||||
about an identity, a pose it has not seen before (turned head, different
|
||||
lighting) becomes an extra reference for recognizing that actor again later
|
||||
in the same film, without touching the baked gallery.
|
||||
|
||||
## The training-set signal
|
||||
## Training-set signal
|
||||
|
||||
Averaged across all 4 models, on the 4 films used for optimization:
|
||||
Averaged across the 3 compared models (r50 excluded), on the 4 films used
|
||||
for optimization. These are the corrected, full-coverage figures, see the
|
||||
[dropped-film note](model-bakeoff.md#a-scoring-bug-worth-recording-dropped-film-evaluations)
|
||||
in the experiment log for why an earlier version of this table overstated the
|
||||
full-mode misID jump (209 → 864) that was itself partly a truncation artifact:
|
||||
|
||||
| scope | expansion | F1 | R | misID |
|
||||
|---|---|---|---|---|
|
||||
| full | off | 71.2% | 58.3% | 209 |
|
||||
| full | **on** | 71.2% | 59.7% | **864** |
|
||||
| restricted | off | 73.6% | 61.3% | 194 |
|
||||
| restricted | **on** | **75.4%** | **64.5%** | 135 |
|
||||
| full | off | 70.0% | 57.6% | 407 |
|
||||
| full | on | 72.1% | 61.5% | 714 |
|
||||
| restricted | off | 75.1% | 63.9% | 179 |
|
||||
| restricted | on | 76.7% | 67.2% | 120 |
|
||||
|
||||
In `restricted` mode (matcher's candidate set capped to the film's own credited
|
||||
cast) expansion looked like a clean win: +1.8pp F1, +3.2pp recall, misID actually
|
||||
lower. In `full` mode it looked flat-to-costly: ~0 F1 change, recall +1.4pp, but
|
||||
misID roughly quadrupled (209 → 864) — see the
|
||||
[bake-off experiment log](model-bakeoff.md) for the per-model breakdown. That's the number that motivated this page: **does turning
|
||||
expansion on actually change what gets recognised, frame by frame, or is the
|
||||
aggregate F1 shift something else?**
|
||||
In restricted mode, expansion looks like a clean win: +1.6pp F1, +3.3pp
|
||||
recall, lower misID. In full mode it looks like a recall-for-misID trade:
|
||||
+2.1pp F1, +3.9pp recall, but misID rises from 407 to 714. See
|
||||
[the full experiment log](model-bakeoff.md) for the per-model breakdown.
|
||||
This asymmetry motivated the question below: does turning expansion on
|
||||
change what gets recognized frame by frame, or is the aggregate F1 shift
|
||||
coming from something else.
|
||||
|
||||
## Held-out test: does it reproduce?
|
||||
## Held-out test
|
||||
|
||||
Same model + same tuned config, `expand_gallery` toggled on vs. off, nothing else
|
||||
changed — full gallery mode, per-second scoring against X-Ray. This isolates
|
||||
expansion from every other variable (config, model, threshold) that differs
|
||||
between the training-set `exp`/`noexp` rows above.
|
||||
Same model, same tuned config, `expand_gallery` toggled on vs. off, nothing
|
||||
else changed, full gallery mode, per-second scoring against X-Ray. This
|
||||
isolates expansion from every other variable that differs between the
|
||||
training-set rows above.
|
||||
|
||||
**LVFace-B Glint360K, all 5 held-out films** (films never seen by the optimizer):
|
||||
LVFace-B Glint360K, all 5 held-out films:
|
||||
|
||||
| film | F1 (exp) | F1 (noexp) | TPI Δ | FN Δ |
|
||||
| film | F1 (exp) | F1 (noexp) | TPI delta | FN delta |
|
||||
|---|---|---|---|---|
|
||||
| Benny & Joon | 83.0% | 83.0% | -2 | +2 |
|
||||
| Downton Abbey: A New Era | 56.1% | 56.2% | -7 | +7 |
|
||||
@@ -44,57 +48,63 @@ between the training-set `exp`/`noexp` rows above.
|
||||
| The Many Saints of Newark | 46.3% | 46.3% | +2 | -2 |
|
||||
| Valerian and the City of a Thousand Planets | 74.1% | 74.1% | +2 | -2 |
|
||||
|
||||
**ArcFace R18** (Benny & Joon, r18's own tuned config): F1 77.1% for both, TPI/FN
|
||||
identical, FPI differs by 2 (noise).
|
||||
ArcFace R18, Benny & Joon, r18's own tuned config: F1 77.1% for both, TPI
|
||||
and FN identical, FPI differs by 2.
|
||||
|
||||
**Every film, both models tested: F1 within 0.1–0.2pp, TPI/FN swings in the tens
|
||||
out of tens of thousands.** That's noise, not a signal — expansion made no
|
||||
measurable difference to per-second onscreen identification anywhere it was
|
||||
tested on unseen data.
|
||||
Every film, both models tested: F1 differs by 0.1-0.2pp, TPI/FN swings are
|
||||
in the tens out of tens of thousands. This is noise, not a signal.
|
||||
Expansion made no measurable difference to per-second on-screen
|
||||
identification on any held-out film tested.
|
||||
|
||||
## Two bugs this required catching (this section's own methodology)
|
||||
## Two methodology bugs caught during this check
|
||||
|
||||
Getting to the clean table above took two wrong turns, both worth recording
|
||||
since they're exactly the kind of error that produces a false positive "look,
|
||||
expansion helped!" finding:
|
||||
Getting to the table above required catching two wrong turns, both worth
|
||||
recording because they are exactly the kind of error that produces a false
|
||||
positive "expansion helped" finding.
|
||||
|
||||
1. **Timeout truncation.** The first Downton Abbey `exp` replay was cut off by a
|
||||
60s subprocess timeout at ~76% through the film (5589 of 7368 expected
|
||||
seconds) — a genuinely large, silent data loss that showed up as a large,
|
||||
convincing-looking TPI gap (47938 vs 52032) purely because one run had a
|
||||
quarter of the film missing. Caught by comparing `n_seconds` between runs
|
||||
before trusting any score delta; fixed by re-running with a longer timeout.
|
||||
2. **Bbox-matching bug.** An early per-second raw-annotation diff matched each
|
||||
`exp` detection to the *first* `noexp` detection with IoU > 0.5, not the
|
||||
*best*-overlapping one. With 3 faces close together in frame, this produced
|
||||
spurious "disagreements" (e.g. "exp says Aidan Quinn, noexp says Johnny
|
||||
Depp" at the same seconds) that vanished entirely once the match picked the
|
||||
true best-IoU candidate — both configs had actually output the exact same
|
||||
three names at the exact same three boxes.
|
||||
1. **Timeout truncation.** The first Downton Abbey `exp` replay was cut off
|
||||
by a 60-second subprocess timeout at about 76% through the film (5589 of
|
||||
7368 expected seconds). This silent data loss produced a large,
|
||||
convincing-looking TPI gap (47938 vs 52032) purely because one run was
|
||||
missing a quarter of the film. Caught by comparing `n_seconds` between
|
||||
runs before trusting any score delta; fixed by re-running with a longer
|
||||
timeout.
|
||||
2. **Bbox-matching bug.** An early per-second raw-annotation diff matched
|
||||
each `exp` detection to the first `noexp` detection with IoU above 0.5,
|
||||
not the best-overlapping one. With 3 faces close together in frame, this
|
||||
produced spurious disagreements (for example "exp says Aidan Quinn,
|
||||
noexp says Johnny Depp" at the same second) that vanished once the match
|
||||
used the best-IoU candidate instead of the first one. Both configs had
|
||||
actually output the same three names at the same three boxes.
|
||||
|
||||
Both bugs independently pointed toward "expansion is doing something," and both
|
||||
were artifacts of the comparison harness, not the pipeline. Worth remembering
|
||||
when a before/after diff looks dramatic: check that the two runs actually cover
|
||||
the same seconds, and match entities by best overlap, not first-found.
|
||||
Both bugs independently pointed toward "expansion is doing something," and
|
||||
both were artifacts of the comparison harness, not the pipeline. Before
|
||||
trusting a dramatic before/after diff, check that both runs cover the same
|
||||
seconds and that entities are matched by best overlap, not first found.
|
||||
|
||||
## What this means
|
||||
## Conclusion
|
||||
|
||||
The training-set aggregate effect (particularly the ~4x misID increase in full
|
||||
mode) doesn't reproduce on held-out data — at minimum it's far smaller than the
|
||||
training-set numbers suggested, and plausibly it's sampling variation from only
|
||||
4 training films rather than a real, generalizable mechanism. This doesn't mean
|
||||
`expand_gallery` never does anything (the mechanism is real — see
|
||||
The training-set aggregate effect, particularly the full-mode misID
|
||||
increase, does not reproduce on held-out data. At minimum it
|
||||
is far smaller than the training-set numbers suggested; it may be sampling
|
||||
variation from only 4 training films rather than a generalizable
|
||||
mechanism. Note the same *class* of harness bug appears twice in this
|
||||
investigation, the timeout truncation in bug #1 above, and the dropped-film
|
||||
aggregation that inflated the raw training-set misID figures. Both make an
|
||||
inert config look consequential; both are reasons to distrust a dramatic
|
||||
training-set delta until it survives on held-out films, which this one did
|
||||
not. This does not mean `expand_gallery` never does anything: the
|
||||
mechanism is real, and
|
||||
[`track_gallery.hpp`](https://REPOLINK/src/gallery/track_gallery.hpp)'s
|
||||
promotion logging: tracks *do* get confirmed and views *do*
|
||||
get promoted into the annex on every film tested), only that **whatever effect
|
||||
it has on final per-second identification was too small to detect against 5
|
||||
held-out films** with this scoring method. A cleaner test would need either many
|
||||
more held-out films or a metric that can see the annex's direct contribution
|
||||
(e.g. tagging which reference embedding won each match), neither of which this
|
||||
pass had budget for.
|
||||
promotion logging confirms tracks get confirmed and views get promoted
|
||||
into the annex on every film tested. It means whatever effect expansion
|
||||
has on final per-second identification was too small to detect against 5
|
||||
held-out films with this scoring method. A cleaner test would need either
|
||||
more held-out films or a metric that can see the annex's direct
|
||||
contribution, such as tagging which reference embedding won each match;
|
||||
neither was in scope for this pass.
|
||||
|
||||
**Practical takeaway**: don't treat the training-set `exp` vs `noexp` numbers in
|
||||
the [bake-off experiment log](model-bakeoff.md) as proof that expansion
|
||||
changes real-world behavior
|
||||
in either direction — on the evidence gathered so far, it doesn't move the
|
||||
needle enough to see.
|
||||
Do not treat the training-set exp/noexp numbers in
|
||||
[the full experiment log](model-bakeoff.md) as proof that expansion changes
|
||||
real-world behavior in either direction. On the evidence gathered so far,
|
||||
it does not move the needle enough to see.
|
||||
|
||||
@@ -0,0 +1,269 @@
|
||||
{
|
||||
"LVFace-B_Glint360K": {
|
||||
"config": {
|
||||
"prob_threshold": 0.7540024664611272,
|
||||
"anneal_sec": 35.53996030397922,
|
||||
"extinction_sec": 57.43359645269811
|
||||
},
|
||||
"films": {
|
||||
"Benny___Joon": {
|
||||
"name": "Benny & Joon",
|
||||
"TPI": 15119,
|
||||
"FPI": 1845,
|
||||
"FPI_misid": 0,
|
||||
"FPI_incast": 1845,
|
||||
"FN": 4343,
|
||||
"precision": 0.8912402735203961,
|
||||
"precision_raw": 0.8912402735203961,
|
||||
"recall": 0.7768471893947179,
|
||||
"f1": 0.8301213418986437,
|
||||
"agreement_rate": 0.7239983093829193,
|
||||
"exact_match_rate": 0.41098901098901097,
|
||||
"n_seconds": 5915,
|
||||
"duration_sec": 5915.0
|
||||
},
|
||||
"Downton_Abbey__A_New_Era": {
|
||||
"name": "Downton Abbey: A New Era",
|
||||
"TPI": 52022,
|
||||
"FPI": 1160,
|
||||
"FPI_misid": 0,
|
||||
"FPI_incast": 1160,
|
||||
"FN": 80089,
|
||||
"precision": 0.9781881087586025,
|
||||
"precision_raw": 0.9781881087586025,
|
||||
"recall": 0.39377493168623356,
|
||||
"f1": 0.5615106884771687,
|
||||
"agreement_rate": 0.4057686401759256,
|
||||
"exact_match_rate": 0.033084311632870865,
|
||||
"n_seconds": 7496,
|
||||
"duration_sec": 7496.0
|
||||
},
|
||||
"Lovelace": {
|
||||
"name": "Lovelace",
|
||||
"TPI": 14988,
|
||||
"FPI": 1086,
|
||||
"FPI_misid": 58,
|
||||
"FPI_incast": 1028,
|
||||
"FN": 7087,
|
||||
"precision": 0.9031091829356471,
|
||||
"precision_raw": 0.9324374766703994,
|
||||
"recall": 0.6789580973952435,
|
||||
"f1": 0.775154508546456,
|
||||
"agreement_rate": 0.7204967829586512,
|
||||
"exact_match_rate": 0.3597703211914588,
|
||||
"n_seconds": 5573,
|
||||
"duration_sec": 5573.0
|
||||
},
|
||||
"The_Many_Saints_of_Newark": {
|
||||
"name": "The Many Saints of Newark",
|
||||
"TPI": 15928,
|
||||
"FPI": 4394,
|
||||
"FPI_misid": 974,
|
||||
"FPI_incast": 3420,
|
||||
"FN": 23785,
|
||||
"precision": 0.5475797579757976,
|
||||
"precision_raw": 0.7837811239051274,
|
||||
"recall": 0.40107773273235464,
|
||||
"f1": 0.46301652592258835,
|
||||
"agreement_rate": 0.3705156874642392,
|
||||
"exact_match_rate": 0.04588936642173853,
|
||||
"n_seconds": 7213,
|
||||
"duration_sec": 7213.0
|
||||
},
|
||||
"Valerian_and_the_City_of_a_Thousand_Plan": {
|
||||
"name": "Valerian and the City of a Thousand Planets",
|
||||
"TPI": 18658,
|
||||
"FPI": 548,
|
||||
"FPI_misid": 0,
|
||||
"FPI_incast": 548,
|
||||
"FN": 12472,
|
||||
"precision": 0.9714672498177653,
|
||||
"precision_raw": 0.9714672498177653,
|
||||
"recall": 0.5993575329264376,
|
||||
"f1": 0.7413382072472983,
|
||||
"agreement_rate": 0.5877853464704299,
|
||||
"exact_match_rate": 0.21980294368081743,
|
||||
"n_seconds": 8221,
|
||||
"duration_sec": 8221.0
|
||||
}
|
||||
}
|
||||
},
|
||||
"arcface_w600k_mbf": {
|
||||
"config": {
|
||||
"prob_threshold": 0.8371114538930519,
|
||||
"anneal_sec": 48.80011450565114,
|
||||
"extinction_sec": 59.3110040220424
|
||||
},
|
||||
"films": {
|
||||
"Benny___Joon": {
|
||||
"name": "Benny & Joon",
|
||||
"TPI": 15219,
|
||||
"FPI": 2463,
|
||||
"FPI_misid": 180,
|
||||
"FPI_incast": 2283,
|
||||
"FN": 4243,
|
||||
"precision": 0.7884675163195524,
|
||||
"precision_raw": 0.8607058025110281,
|
||||
"recall": 0.7819854074606927,
|
||||
"f1": 0.7852130843050252,
|
||||
"agreement_rate": 0.7072306082196138,
|
||||
"exact_match_rate": 0.34911242603550297,
|
||||
"n_seconds": 5915,
|
||||
"duration_sec": 5915.0
|
||||
},
|
||||
"Downton_Abbey__A_New_Era": {
|
||||
"name": "Downton Abbey: A New Era",
|
||||
"TPI": 53043,
|
||||
"FPI": 2383,
|
||||
"FPI_misid": 604,
|
||||
"FPI_incast": 1779,
|
||||
"FN": 79068,
|
||||
"precision": 0.8715290328940882,
|
||||
"precision_raw": 0.9570057373795692,
|
||||
"recall": 0.4015032813316075,
|
||||
"f1": 0.5497453011561202,
|
||||
"agreement_rate": 0.4094114144765671,
|
||||
"exact_match_rate": 0.032817502668089645,
|
||||
"n_seconds": 7496,
|
||||
"duration_sec": 7496.0
|
||||
},
|
||||
"Lovelace": {
|
||||
"name": "Lovelace",
|
||||
"TPI": 14606,
|
||||
"FPI": 1337,
|
||||
"FPI_misid": 180,
|
||||
"FPI_incast": 1157,
|
||||
"FN": 7469,
|
||||
"precision": 0.8316346865569664,
|
||||
"precision_raw": 0.916138744276485,
|
||||
"recall": 0.6616534541336353,
|
||||
"f1": 0.7369695746505879,
|
||||
"agreement_rate": 0.6907677036961077,
|
||||
"exact_match_rate": 0.31742329086667864,
|
||||
"n_seconds": 5573,
|
||||
"duration_sec": 5573.0
|
||||
},
|
||||
"The_Many_Saints_of_Newark": {
|
||||
"name": "The Many Saints of Newark",
|
||||
"TPI": 15223,
|
||||
"FPI": 4574,
|
||||
"FPI_misid": 994,
|
||||
"FPI_incast": 3580,
|
||||
"FN": 24490,
|
||||
"precision": 0.52962460425147,
|
||||
"precision_raw": 0.768954892155377,
|
||||
"recall": 0.383325359454083,
|
||||
"f1": 0.44475283393712756,
|
||||
"agreement_rate": 0.3554753116932427,
|
||||
"exact_match_rate": 0.03715513655899071,
|
||||
"n_seconds": 7213,
|
||||
"duration_sec": 7213.0
|
||||
},
|
||||
"Valerian_and_the_City_of_a_Thousand_Plan": {
|
||||
"name": "Valerian and the City of a Thousand Planets",
|
||||
"TPI": 18472,
|
||||
"FPI": 853,
|
||||
"FPI_misid": 239,
|
||||
"FPI_incast": 614,
|
||||
"FN": 12658,
|
||||
"precision": 0.860122927919538,
|
||||
"precision_raw": 0.9558602846054334,
|
||||
"recall": 0.5933825891423065,
|
||||
"f1": 0.7022773067710907,
|
||||
"agreement_rate": 0.5795914643682625,
|
||||
"exact_match_rate": 0.18817662084904513,
|
||||
"n_seconds": 8221,
|
||||
"duration_sec": 8221.0
|
||||
}
|
||||
}
|
||||
},
|
||||
"arcface_r18": {
|
||||
"config": {
|
||||
"prob_threshold": 0.8955101189489445,
|
||||
"anneal_sec": 59.08214397442713,
|
||||
"extinction_sec": 59.29474134414983
|
||||
},
|
||||
"films": {
|
||||
"Benny___Joon": {
|
||||
"name": "Benny & Joon",
|
||||
"TPI": 13580,
|
||||
"FPI": 1666,
|
||||
"FPI_misid": 60,
|
||||
"FPI_incast": 1606,
|
||||
"FN": 5882,
|
||||
"precision": 0.86025592296972,
|
||||
"precision_raw": 0.8907254361799817,
|
||||
"recall": 0.697770013359367,
|
||||
"f1": 0.7705401724920563,
|
||||
"agreement_rate": 0.6547675401521545,
|
||||
"exact_match_rate": 0.32578191039729504,
|
||||
"n_seconds": 5915,
|
||||
"duration_sec": 5915.0
|
||||
},
|
||||
"Downton_Abbey__A_New_Era": {
|
||||
"name": "Downton Abbey: A New Era",
|
||||
"TPI": 48545,
|
||||
"FPI": 1066,
|
||||
"FPI_misid": 180,
|
||||
"FPI_incast": 886,
|
||||
"FN": 83566,
|
||||
"precision": 0.9475708067381078,
|
||||
"precision_raw": 0.9785128298159682,
|
||||
"recall": 0.3674561542944948,
|
||||
"f1": 0.5295567845883649,
|
||||
"agreement_rate": 0.3815368792000116,
|
||||
"exact_match_rate": 0.032950907150480255,
|
||||
"n_seconds": 7496,
|
||||
"duration_sec": 7496.0
|
||||
},
|
||||
"Lovelace": {
|
||||
"name": "Lovelace",
|
||||
"TPI": 13615,
|
||||
"FPI": 963,
|
||||
"FPI_misid": 120,
|
||||
"FPI_incast": 843,
|
||||
"FN": 8460,
|
||||
"precision": 0.8695235662281262,
|
||||
"precision_raw": 0.933941555768967,
|
||||
"recall": 0.6167610419026047,
|
||||
"f1": 0.7216494845360825,
|
||||
"agreement_rate": 0.6506356469257915,
|
||||
"exact_match_rate": 0.2894311860757222,
|
||||
"n_seconds": 5573,
|
||||
"duration_sec": 5573.0
|
||||
},
|
||||
"The_Many_Saints_of_Newark": {
|
||||
"name": "The Many Saints of Newark",
|
||||
"TPI": 13489,
|
||||
"FPI": 3757,
|
||||
"FPI_misid": 796,
|
||||
"FPI_incast": 2961,
|
||||
"FN": 26224,
|
||||
"precision": 0.5526013928717739,
|
||||
"precision_raw": 0.7821523831613127,
|
||||
"recall": 0.3396620753909299,
|
||||
"f1": 0.42072267361165266,
|
||||
"agreement_rate": 0.3229817885335633,
|
||||
"exact_match_rate": 0.04422570359073894,
|
||||
"n_seconds": 7213,
|
||||
"duration_sec": 7213.0
|
||||
},
|
||||
"Valerian_and_the_City_of_a_Thousand_Plan": {
|
||||
"name": "Valerian and the City of a Thousand Planets",
|
||||
"TPI": 17692,
|
||||
"FPI": 397,
|
||||
"FPI_misid": 68,
|
||||
"FPI_incast": 329,
|
||||
"FN": 13438,
|
||||
"precision": 0.9460456660071654,
|
||||
"precision_raw": 0.9780529603626513,
|
||||
"recall": 0.5683263732733698,
|
||||
"f1": 0.710080070638759,
|
||||
"agreement_rate": 0.5633540120828806,
|
||||
"exact_match_rate": 0.13404695292543486,
|
||||
"n_seconds": 8221,
|
||||
"duration_sec": 8221.0
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,221 @@
|
||||
{
|
||||
"LVFace-B_Glint360K": {
|
||||
"config": {
|
||||
"prob_threshold": 0.7540024664611272,
|
||||
"anneal_sec": 35.53996030397922,
|
||||
"extinction_sec": 57.43359645269811
|
||||
},
|
||||
"films": {
|
||||
"Caf\u00e9_Society": {
|
||||
"name": "Caf\u00e9 Society",
|
||||
"TPI": 14499,
|
||||
"FPI": 1380,
|
||||
"FPI_misid": 0,
|
||||
"FPI_incast": 1380,
|
||||
"FN": 12231,
|
||||
"precision": 0.9130927640279615,
|
||||
"precision_raw": 0.9130927640279615,
|
||||
"recall": 0.5424242424242425,
|
||||
"f1": 0.6805604449764134,
|
||||
"agreement_rate": 0.57285804629501,
|
||||
"exact_match_rate": 0.18947003810183582,
|
||||
"n_seconds": 5774,
|
||||
"duration_sec": 5774.0
|
||||
},
|
||||
"Lord_of_War": {
|
||||
"name": "Lord of War",
|
||||
"TPI": 13893,
|
||||
"FPI": 1654,
|
||||
"FPI_misid": 174,
|
||||
"FPI_incast": 1480,
|
||||
"FN": 5737,
|
||||
"precision": 0.811838952842868,
|
||||
"precision_raw": 0.8936129156750499,
|
||||
"recall": 0.7077432501273561,
|
||||
"f1": 0.7562256756388972,
|
||||
"agreement_rate": 0.7005158404089996,
|
||||
"exact_match_rate": 0.38715420432758146,
|
||||
"n_seconds": 7302,
|
||||
"duration_sec": 7302.0
|
||||
},
|
||||
"Scarface": {
|
||||
"name": "Scarface",
|
||||
"TPI": 20518,
|
||||
"FPI": 1078,
|
||||
"FPI_misid": 58,
|
||||
"FPI_incast": 1020,
|
||||
"FN": 14722,
|
||||
"precision": 0.9276607288181572,
|
||||
"precision_raw": 0.9500833487682904,
|
||||
"recall": 0.5822360953461975,
|
||||
"f1": 0.7154363820216882,
|
||||
"agreement_rate": 0.6297735703976657,
|
||||
"exact_match_rate": 0.25910733470065433,
|
||||
"n_seconds": 10239,
|
||||
"duration_sec": 10239.0
|
||||
},
|
||||
"Sound_of_Metal": {
|
||||
"name": "Sound of Metal",
|
||||
"TPI": 13349,
|
||||
"FPI": 677,
|
||||
"FPI_misid": 0,
|
||||
"FPI_incast": 677,
|
||||
"FN": 6504,
|
||||
"precision": 0.9517324967916726,
|
||||
"precision_raw": 0.9517324967916726,
|
||||
"recall": 0.6723920818012391,
|
||||
"f1": 0.7880397886596416,
|
||||
"agreement_rate": 0.7112222835587533,
|
||||
"exact_match_rate": 0.40311896218603366,
|
||||
"n_seconds": 7246,
|
||||
"duration_sec": 7246.0
|
||||
}
|
||||
}
|
||||
},
|
||||
"arcface_w600k_mbf": {
|
||||
"config": {
|
||||
"prob_threshold": 0.8371114538930519,
|
||||
"anneal_sec": 48.80011450565114,
|
||||
"extinction_sec": 59.3110040220424
|
||||
},
|
||||
"films": {
|
||||
"Caf\u00e9_Society": {
|
||||
"name": "Caf\u00e9 Society",
|
||||
"TPI": 13456,
|
||||
"FPI": 1434,
|
||||
"FPI_misid": 180,
|
||||
"FPI_incast": 1254,
|
||||
"FN": 13274,
|
||||
"precision": 0.8150211992731677,
|
||||
"precision_raw": 0.9036937541974479,
|
||||
"recall": 0.5034044145155256,
|
||||
"f1": 0.622386679000925,
|
||||
"agreement_rate": 0.5463565775524695,
|
||||
"exact_match_rate": 0.19154832005542086,
|
||||
"n_seconds": 5774,
|
||||
"duration_sec": 5774.0
|
||||
},
|
||||
"Lord_of_War": {
|
||||
"name": "Lord of War",
|
||||
"TPI": 13778,
|
||||
"FPI": 1740,
|
||||
"FPI_misid": 60,
|
||||
"FPI_incast": 1680,
|
||||
"FN": 5852,
|
||||
"precision": 0.8580146967243741,
|
||||
"precision_raw": 0.8878721484727413,
|
||||
"recall": 0.7018848700967907,
|
||||
"f1": 0.7721362923111411,
|
||||
"agreement_rate": 0.6881858851455998,
|
||||
"exact_match_rate": 0.36469460421802247,
|
||||
"n_seconds": 7302,
|
||||
"duration_sec": 7302.0
|
||||
},
|
||||
"Scarface": {
|
||||
"name": "Scarface",
|
||||
"TPI": 18863,
|
||||
"FPI": 862,
|
||||
"FPI_misid": 0,
|
||||
"FPI_incast": 862,
|
||||
"FN": 16377,
|
||||
"precision": 0.956299112801014,
|
||||
"precision_raw": 0.956299112801014,
|
||||
"recall": 0.535272417707151,
|
||||
"f1": 0.6863640498499044,
|
||||
"agreement_rate": 0.594178111391709,
|
||||
"exact_match_rate": 0.2357652114464303,
|
||||
"n_seconds": 10239,
|
||||
"duration_sec": 10239.0
|
||||
},
|
||||
"Sound_of_Metal": {
|
||||
"name": "Sound of Metal",
|
||||
"TPI": 12642,
|
||||
"FPI": 554,
|
||||
"FPI_misid": 0,
|
||||
"FPI_incast": 554,
|
||||
"FN": 7211,
|
||||
"precision": 0.9580175810851773,
|
||||
"precision_raw": 0.9580175810851773,
|
||||
"recall": 0.6367803354656727,
|
||||
"f1": 0.7650458410239341,
|
||||
"agreement_rate": 0.687468948385309,
|
||||
"exact_match_rate": 0.3789677063207287,
|
||||
"n_seconds": 7246,
|
||||
"duration_sec": 7246.0
|
||||
}
|
||||
}
|
||||
},
|
||||
"arcface_r18": {
|
||||
"config": {
|
||||
"prob_threshold": 0.8955101189489445,
|
||||
"anneal_sec": 59.08214397442713,
|
||||
"extinction_sec": 59.29474134414983
|
||||
},
|
||||
"films": {
|
||||
"Caf\u00e9_Society": {
|
||||
"name": "Caf\u00e9 Society",
|
||||
"TPI": 12207,
|
||||
"FPI": 1119,
|
||||
"FPI_misid": 60,
|
||||
"FPI_incast": 1059,
|
||||
"FN": 14523,
|
||||
"precision": 0.88035482475119,
|
||||
"precision_raw": 0.9160288158487168,
|
||||
"recall": 0.45667789001122333,
|
||||
"f1": 0.6013892994383683,
|
||||
"agreement_rate": 0.5125626845924863,
|
||||
"exact_match_rate": 0.1674748874263942,
|
||||
"n_seconds": 5774,
|
||||
"duration_sec": 5774.0
|
||||
},
|
||||
"Lord_of_War": {
|
||||
"name": "Lord of War",
|
||||
"TPI": 13122,
|
||||
"FPI": 1409,
|
||||
"FPI_misid": 60,
|
||||
"FPI_incast": 1349,
|
||||
"FN": 6508,
|
||||
"precision": 0.870678787074514,
|
||||
"precision_raw": 0.9030348909228546,
|
||||
"recall": 0.6684666327050433,
|
||||
"f1": 0.7562894441082388,
|
||||
"agreement_rate": 0.6698963754222389,
|
||||
"exact_match_rate": 0.3389482333607231,
|
||||
"n_seconds": 7302,
|
||||
"duration_sec": 7302.0
|
||||
},
|
||||
"Scarface": {
|
||||
"name": "Scarface",
|
||||
"TPI": 16961,
|
||||
"FPI": 685,
|
||||
"FPI_misid": 0,
|
||||
"FPI_incast": 685,
|
||||
"FN": 18279,
|
||||
"precision": 0.961181004193585,
|
||||
"precision_raw": 0.961181004193585,
|
||||
"recall": 0.48129965947786607,
|
||||
"f1": 0.6414173883447416,
|
||||
"agreement_rate": 0.5440709115628456,
|
||||
"exact_match_rate": 0.2017775173356773,
|
||||
"n_seconds": 10239,
|
||||
"duration_sec": 10239.0
|
||||
},
|
||||
"Sound_of_Metal": {
|
||||
"name": "Sound of Metal",
|
||||
"TPI": 12017,
|
||||
"FPI": 591,
|
||||
"FPI_misid": 122,
|
||||
"FPI_incast": 469,
|
||||
"FN": 7836,
|
||||
"precision": 0.8767692981176127,
|
||||
"precision_raw": 0.953125,
|
||||
"recall": 0.6052989472623784,
|
||||
"f1": 0.7161715188176049,
|
||||
"agreement_rate": 0.6594534915815532,
|
||||
"exact_match_rate": 0.3573005796301408,
|
||||
"n_seconds": 7246,
|
||||
"duration_sec": 7246.0
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
+2
-2
@@ -34,13 +34,13 @@ extra_css:
|
||||
|
||||
nav:
|
||||
- Home: index.md
|
||||
- How We Score Against X-Ray: methodology.md
|
||||
- Findings:
|
||||
- Best Model: best-model.md
|
||||
- Gallery Scope (Full vs. Limited): gallery-scope.md
|
||||
- Pose Expansion: pose-expansion.md
|
||||
- LVFace Deep Dive: lvface-deep-dive.md
|
||||
- Model Bake-off & Re-tune (full log): model-bakeoff.md
|
||||
- Optimizer Experiments (prior round): optimizer-experiments.md
|
||||
- Full Experiment Log: model-bakeoff.md
|
||||
- Service Conversion (proposal): service-conversion.md
|
||||
|
||||
markdown_extensions:
|
||||
|
||||
@@ -60,10 +60,34 @@ stage_frame "${MONTAGE_ROOT}/Valerian_and_the_City_of_a_Thousand_Plan/scene_4/4_
|
||||
valerian_screen_call.jpg
|
||||
stage_frame "${MONTAGE_ROOT}/The_Many_Saints_of_Newark/out_of_cast_fpi/4_worst_t000871.jpg" \
|
||||
many_saints_outofcast_fpi.jpg
|
||||
# debug-overlay example (extinction state drawn as frozen boxes) — from the
|
||||
# dump_error_frames output, not the montage package
|
||||
stage_frame "experiments/results/holdout/frames/many_saints/fpi/fpi_t03543.jpg" \
|
||||
many_saints_ghost_fpi.jpg
|
||||
|
||||
# One frame per DISTINCT out-of-cast name across all 9 films, uniform rule
|
||||
# (see scripts/docs/first_fpi_frames.py): the first second in the raw replay
|
||||
# stream where the pipeline names someone not in the film's credited cast at
|
||||
# all. Rendered with the proper montage renderer (Onscreen/Offscreen panel),
|
||||
# never dump_error_frames.py's bare-box overlay. Regenerate with:
|
||||
# python3 scripts/docs/first_fpi_frames.py
|
||||
# 5 of 9 films have zero out-of-cast names in their whole runtime (Benny &
|
||||
# Joon, Cafe Society, Downton Abbey, Sound of Metal, Valerian) and produce
|
||||
# no frames here.
|
||||
stage_frame "${MONTAGE_ROOT}/Lord_of_War/first_fpi_david_shumbris/first_fpi_t000418.jpg" \
|
||||
lord_of_war_fpi_shumbris.jpg
|
||||
stage_frame "${MONTAGE_ROOT}/Lord_of_War/first_fpi_ronald_reagan/first_fpi_t001003.jpg" \
|
||||
lord_of_war_fpi_reagan_photo.jpg
|
||||
stage_frame "${MONTAGE_ROOT}/Lord_of_War/first_fpi_lance_reddick/first_fpi_t006424.jpg" \
|
||||
lord_of_war_fpi_reddick.jpg
|
||||
stage_frame "${MONTAGE_ROOT}/Lovelace/first_fpi_chloë_sevigny/first_fpi_t002451.jpg" \
|
||||
lovelace_fpi_sevigny.jpg
|
||||
stage_frame "${MONTAGE_ROOT}/Scarface/first_fpi_kirstie_alley/first_fpi_t002451.jpg" \
|
||||
scarface_fpi_alley.jpg
|
||||
stage_frame "${MONTAGE_ROOT}/The_Many_Saints_of_Newark/first_fpi_germar_terrell_gardner/first_fpi_t000848.jpg" \
|
||||
many_saints_fpi_gardner.jpg
|
||||
stage_frame "${MONTAGE_ROOT}/The_Many_Saints_of_Newark/first_fpi_archie_yates/first_fpi_t002521.jpg" \
|
||||
many_saints_fpi_yates.jpg
|
||||
stage_frame "${MONTAGE_ROOT}/The_Many_Saints_of_Newark/first_fpi_zooey_deschanel/first_fpi_t002819.jpg" \
|
||||
many_saints_fpi_deschanel.jpg
|
||||
stage_frame "${MONTAGE_ROOT}/The_Many_Saints_of_Newark/first_fpi_talia_balsam/first_fpi_t004551.jpg" \
|
||||
many_saints_fpi_balsam.jpg
|
||||
|
||||
if [ ! -f "${ASSETS_DIR}/germar_beats_xray.jpg" ]; then
|
||||
echo "==> pulling report-highlights/germar_beats_xray.jpg..."
|
||||
|
||||
@@ -27,8 +27,11 @@ RESULTS = REPO / "experiments/results"
|
||||
|
||||
# Same model -> color mapping as calibration_chart.py, so identity is stable
|
||||
# across every figure in the report.
|
||||
# r50 is dropped from the bake-off: its 4 combos ran under the old, narrower
|
||||
# anneal/extinction bounds and were never re-run wide, so they are not comparable
|
||||
# on those two params (and two of them were truncation-corrupted). Its slug stays
|
||||
# out of this map so it never appears in a figure or legend.
|
||||
MODEL_COLOURS = {
|
||||
"arcface_w600k_r50": ("ArcFace w600k-R50", "#2a78d6"),
|
||||
"arcface_r18": ("ArcFace R18", "#008300"),
|
||||
"arcface_w600k_mbf": ("ArcFace w600k-MBF", "#e87ba4"),
|
||||
"LVFace-B_Glint360K": ("LVFace-B Glint360K", "#eda100"),
|
||||
@@ -59,9 +62,37 @@ plt.rcParams.update({
|
||||
})
|
||||
|
||||
|
||||
TRAJ = REPO / "experiments/trajectories"
|
||||
|
||||
|
||||
def clean_best(combo: str) -> dict:
|
||||
"""Best-F1 eval for a combo, restricted to FULL-COVERAGE evals.
|
||||
|
||||
The optimizer averages F1 (and *sums* TPI/misID) over only the films whose
|
||||
replay subprocess didn't time out (optimize.py: `per_film = [... if m is not
|
||||
None]`). A candidate whose hardest film timed out is therefore scored on an
|
||||
easier subset, which inflates its F1 — and DE will happily converge onto such
|
||||
a candidate. `rep4_best_*.json` recorded exactly that kind of eval for at
|
||||
least one combo (arcface_w600k_mbf_full_noexp: reported 74.2% F1 came from an
|
||||
eval with TPI 12645, a third of that combo's median).
|
||||
|
||||
We recover comparable numbers straight from the trajectory: take the median
|
||||
TPI across all evals (full 4-film coverage) and keep only evals within 30% of
|
||||
it, then pick the highest-F1 survivor. No re-running — the honest best config
|
||||
is already in the sweep, just not the one `argmax f1` picked.
|
||||
"""
|
||||
evals = [json.loads(l) for l in open(TRAJ / f"rep4_{combo}.jsonl")]
|
||||
tpis = sorted(e["TPI"] for e in evals)
|
||||
med = tpis[len(tpis) // 2]
|
||||
clean = [e for e in evals if e["TPI"] >= 0.7 * med]
|
||||
return max(clean, key=lambda e: e["f1"])
|
||||
|
||||
|
||||
def training_best() -> dict:
|
||||
with open(RESULTS / "rep4_best_LVFace-B_Glint360K_full_exp.json") as f:
|
||||
return json.load(f)["best"]
|
||||
# LVFace-B_Glint360K_full_exp is the shipped combo; its reported best is a
|
||||
# full-coverage eval (TPI 47757 ≈ median), so clean_best returns the same
|
||||
# config — but route it through clean_best so every figure uses one path.
|
||||
return clean_best("LVFace-B_Glint360K_full_exp")
|
||||
|
||||
|
||||
def fig_holdout_f1(out: Path):
|
||||
@@ -100,18 +131,16 @@ def fig_holdout_f1(out: Path):
|
||||
|
||||
def fig_rep4_matrix(out: Path):
|
||||
combos = []
|
||||
for path in sorted(RESULTS.glob("rep4_best_*.json")):
|
||||
stem = path.stem[len("rep4_best_"):]
|
||||
for path in sorted(TRAJ.glob("rep4_*.jsonl")):
|
||||
combo = path.stem[len("rep4_"):]
|
||||
for slug in MODEL_COLOURS:
|
||||
if stem.startswith(slug):
|
||||
mode = stem[len(slug) + 1:] # e.g. full_exp
|
||||
with open(path) as f:
|
||||
best = json.load(f)["best"]
|
||||
combos.append((slug, mode, best["f1"] * 100))
|
||||
if combo.startswith(slug):
|
||||
mode = combo[len(slug) + 1:] # e.g. full_exp
|
||||
combos.append((slug, mode, clean_best(combo)["f1"] * 100))
|
||||
break
|
||||
combos.sort(key=lambda c: c[2])
|
||||
|
||||
fig, ax = plt.subplots(figsize=(9, 6.2))
|
||||
fig, ax = plt.subplots(figsize=(9, 5.2))
|
||||
ax.grid(axis="y", visible=False)
|
||||
labels = []
|
||||
for i, (slug, mode, f1) in enumerate(combos):
|
||||
@@ -126,7 +155,7 @@ def fig_rep4_matrix(out: Path):
|
||||
ax.set_yticks(range(len(combos)), labels, fontsize=9)
|
||||
ax.set_xlim(65, 80)
|
||||
ax.set_xlabel("training-set per-second F1 (%)")
|
||||
ax.set_title("All 16 combos — filled dot = cast-restricted gallery, open = full",
|
||||
ax.set_title("All 12 combos — filled dot = cast-restricted gallery, open = full",
|
||||
loc="left", fontsize=12, pad=12)
|
||||
handles = [plt.Line2D([], [], marker="o", ls="", ms=9, color=c, label=l)
|
||||
for _, (l, c) in MODEL_COLOURS.items()]
|
||||
@@ -199,13 +228,16 @@ def fig_downton_timeline(out: Path, t0: int = 7100, t1: int = 7340):
|
||||
trk = np.array(trk)
|
||||
dc = np.array([det.get(s, 0) for s in t])
|
||||
|
||||
fig, ax = plt.subplots(figsize=(9.5, 4.4))
|
||||
fig, ax = plt.subplots(figsize=(9.5, 4.8))
|
||||
ax.grid(axis="x", visible=False)
|
||||
ax.fill_between(t, dc, step="mid", color=GREEN, alpha=0.25, zorder=2)
|
||||
ax.step(t, dc, where="mid", color=GREEN, lw=2, zorder=3)
|
||||
ax.step(t, trk, where="mid", color=BLUE, lw=2, zorder=4)
|
||||
ax.fill_between(t, dc, step="mid", color=GREEN, alpha=0.22, zorder=2)
|
||||
ax.step(t, dc, where="mid", color=GREEN, lw=2, zorder=3,
|
||||
label="faces seen by detector")
|
||||
ax.step(t, trk, where="mid", color=BLUE, lw=2, zorder=4,
|
||||
label="actors reported by tracker")
|
||||
|
||||
# longest contiguous run of "detector sees nothing, tracker still reporting"
|
||||
# (i.e. every reported actor is extinction-bridged, not detected this second)
|
||||
ghost = (dc == 0) & (trk > 0)
|
||||
runs, start = [], None
|
||||
for i, g in enumerate(ghost):
|
||||
@@ -219,15 +251,11 @@ def fig_downton_timeline(out: Path, t0: int = 7100, t1: int = 7340):
|
||||
if runs:
|
||||
i0, i1 = max(runs, key=lambda r: r[1] - r[0])
|
||||
g0, g1 = t[i0], t[i1]
|
||||
ax.axvspan(g0, g1, color=RED, alpha=0.08, zorder=1)
|
||||
ax.annotate(f"{g1 - g0}s of credits: 0 faces detected,\n"
|
||||
f"{trk[i0]} actors still reported (frozen boxes)",
|
||||
((g0 + g1) / 2, 20.5), ha="center", va="bottom",
|
||||
fontsize=10, color=RED)
|
||||
ax.text(t0 + 4, 27.3, "actors reported by tracker", color=BLUE,
|
||||
fontsize=10.5, va="bottom")
|
||||
ax.text(t0 + 4, 11.5, "faces seen by detector", color=GREEN,
|
||||
fontsize=10.5, va="bottom")
|
||||
ax.axvspan(g0, g1, color=RED, alpha=0.08, zorder=1,
|
||||
label=f"{g1 - g0}s bridged: 0 faces detected,\n"
|
||||
f"{trk[i0]} actors carried by their\nextinction window")
|
||||
ax.legend(loc="upper right", frameon=True, framealpha=0.92,
|
||||
edgecolor=GRID, fontsize=9.5)
|
||||
ax.set_xlabel("film time (s)")
|
||||
ax.set_ylabel("count")
|
||||
ax.set_ylim(0, 31)
|
||||
|
||||
@@ -0,0 +1,155 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
first_fpi_frames.py — for every film, find every DISTINCT out-of-cast name
|
||||
(misID) the raw replay stream ever reports, and render the exact second each
|
||||
one FIRST appears, with the proper montage renderer (dump_scene_montage.py:
|
||||
Onscreen/Offscreen panel, TPI/FPI/FN legend, ghosts never drawn as boxes —
|
||||
imported directly, not the scene-level best/worst picker, which can land on
|
||||
a different second within the same scene).
|
||||
|
||||
One rule, applied uniformly across all 9 films and every distinct wrong name
|
||||
in each — no manual per-film picking, no stopping at the first name found.
|
||||
"""
|
||||
import csv
|
||||
import json
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
import cv2
|
||||
|
||||
REPO = Path(__file__).resolve().parent.parent.parent
|
||||
sys.path.insert(0, str(REPO / "scripts" / "validation"))
|
||||
sys.path.insert(0, str(REPO / "scripts" / "optimizer"))
|
||||
from identity import keys_for # noqa: E402
|
||||
from sample_eval import load_gallery_keys # noqa: E402
|
||||
from dump_scene_montage import ( # noqa: E402
|
||||
classify_second, extract_frame, render_frame,
|
||||
load_scene_cast, load_dump_faces_by_second, load_raw_by_second,
|
||||
)
|
||||
|
||||
FILMS = [
|
||||
("Benny___Joon", "experiments/xray/scene_level_movie_data_XRay_US/xrays/4808_Benny__Joon"),
|
||||
("Café_Society", "experiments/xray/scene_level_movie_data_XRay_US/xrays/225_Cafe_Society"),
|
||||
("Downton_Abbey__A_New_Era", "experiments/xray/scene_level_movie_data_XRay_US/xrays/19_Downton_Abbey_A_New_Era"),
|
||||
("Lord_of_War", "experiments/xray/scene_level_movie_data_XRay_US/xrays/2474_Lord_of_War"),
|
||||
("Lovelace", "experiments/xray/scene_level_movie_data_XRay_US/xrays/4108_Lovelace"),
|
||||
("Scarface", "experiments/xray/scene_level_movie_data_XRay_US/xrays/197_Scarface"),
|
||||
("Sound_of_Metal", "experiments/xray/scene_level_movie_data_XRay_US/xrays/6278_Sound_of_Metal"),
|
||||
("The_Many_Saints_of_Newark", "experiments/xray/scene_level_movie_data_XRay_US/xrays/900_The_Many_Saints_Of_Newark"),
|
||||
("Valerian_and_the_City_of_a_Thousand_Plan", "experiments/xray/scene_level_movie_data_XRay_US/xrays/5312_Valerian_and_the_City_of_a_Thousand_Planets"),
|
||||
]
|
||||
|
||||
MOVIE_ROOT = Path("/mnt/movies")
|
||||
|
||||
|
||||
def load_film_cast_keys(xray_dir: Path) -> set:
|
||||
keys = set()
|
||||
with open(xray_dir / "people.csv", newline="", encoding="utf-8") as f:
|
||||
for r in csv.DictReader(f):
|
||||
nm = (r.get("name_id") or "").strip()
|
||||
person = (r.get("person") or "").strip()
|
||||
if nm or person:
|
||||
keys |= keys_for(imdb_id=nm, name=person)
|
||||
return keys
|
||||
|
||||
|
||||
def find_movie_file(slug: str) -> str | None:
|
||||
# dump HDF5 attrs carry the exact path used at dump time
|
||||
import h5py
|
||||
for model in ("LVFace-B_Glint360K",):
|
||||
p = REPO / f"experiments/dumps/{model}/dump_{slug}.h5"
|
||||
if p.exists():
|
||||
with h5py.File(p, "r") as f:
|
||||
return f.attrs.get("movie")
|
||||
return None
|
||||
|
||||
|
||||
def find_scene_id(xray_dir: Path, t: int) -> str | None:
|
||||
with open(xray_dir / "scenes.csv", newline="", encoding="utf-8") as f:
|
||||
for r in csv.DictReader(f):
|
||||
try:
|
||||
t0, t1 = float(r["start"]) / 1000.0, float(r["end"]) / 1000.0
|
||||
except (KeyError, ValueError):
|
||||
continue
|
||||
if t0 <= t < t1:
|
||||
return (r.get("scene") or "").strip()
|
||||
return None
|
||||
|
||||
|
||||
def main():
|
||||
out_root = REPO / "experiments/results/holdout/montage_bestworst"
|
||||
summary = []
|
||||
|
||||
for slug, xray_rel in FILMS:
|
||||
xray_dir = REPO / xray_rel
|
||||
raw_path = out_root / f"raw_{slug}.jsonl"
|
||||
if not raw_path.exists():
|
||||
print(f"SKIP {slug}: no raw file", file=sys.stderr)
|
||||
continue
|
||||
|
||||
cast_keys = load_film_cast_keys(xray_dir)
|
||||
|
||||
# every distinct out-of-cast name -> first second it appears
|
||||
first_seen: dict[str, int] = {}
|
||||
with open(raw_path) as f:
|
||||
for line in f:
|
||||
d = json.loads(line)
|
||||
if d.get("eof"):
|
||||
continue
|
||||
for a in d.get("visible_actors", []):
|
||||
name = a.get("name")
|
||||
if not name or name in first_seen:
|
||||
continue
|
||||
ak = keys_for(imdb_id=a.get("imdb_id"), name=name,
|
||||
jellyfin_id=a.get("jellyfin_id"))
|
||||
if not (ak & cast_keys):
|
||||
first_seen[name] = int(d["timestamp_sec"])
|
||||
|
||||
if not first_seen:
|
||||
print(f"{slug}: no out-of-cast FPI in the whole film", file=sys.stderr)
|
||||
summary.append((slug, None, None))
|
||||
continue
|
||||
|
||||
print(f"{slug}: {len(first_seen)} distinct out-of-cast name(s)", file=sys.stderr)
|
||||
|
||||
movie = find_movie_file(slug)
|
||||
if not movie or not Path(movie).exists():
|
||||
print(f" SKIP render: movie file not found ({movie})", file=sys.stderr)
|
||||
for name, t in first_seen.items():
|
||||
summary.append((slug, name, t))
|
||||
continue
|
||||
|
||||
dump_path = REPO / f"experiments/dumps/LVFace-B_Glint360K/dump_{slug}.h5"
|
||||
gallery_path = REPO / "experiments/galleries/gallery_LVFace-B_Glint360K.h5"
|
||||
gallery_keys = load_gallery_keys(str(gallery_path))
|
||||
raw_by_second = load_raw_by_second(str(raw_path))
|
||||
dump_faces_by_second = load_dump_faces_by_second(str(dump_path))
|
||||
scene_cast = load_scene_cast(str(xray_dir))
|
||||
|
||||
for name, t in sorted(first_seen.items(), key=lambda kv: kv[1]):
|
||||
scene_id = find_scene_id(xray_dir, t)
|
||||
gt_cast = scene_cast.get(scene_id, set())
|
||||
gt_cast = {g for g in gt_cast if g & gallery_keys}
|
||||
|
||||
score, tpi_boxes, fpi_boxes, entries, has_outofcast = classify_second(
|
||||
t, gt_cast, cast_keys, raw_by_second, dump_faces_by_second)
|
||||
|
||||
slug_name = name.lower().replace(" ", "_").replace("'", "")
|
||||
out_dir = out_root / slug / f"first_fpi_{slug_name}"
|
||||
out_dir.mkdir(parents=True, exist_ok=True)
|
||||
out_path = out_dir / f"first_fpi_t{t:06d}.jpg"
|
||||
extract_frame(movie, t, out_path)
|
||||
canvas = render_frame(out_path, t, tpi_boxes, fpi_boxes, entries)
|
||||
if canvas is not None:
|
||||
cv2.imwrite(str(out_path), canvas)
|
||||
print(f" {name!r} t={t}s -> {out_path} (outofcast={has_outofcast})",
|
||||
file=sys.stderr)
|
||||
summary.append((slug, name, t))
|
||||
|
||||
print("\n=== summary ===", file=sys.stderr)
|
||||
for slug, name, t in summary:
|
||||
print(f" {slug:45s} {name!r:30s} t={t}", file=sys.stderr)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,62 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
gallery_coverage_per_film.py — fraction of each film's X-Ray credited cast
|
||||
that has a reference embedding in the gallery, computed per film rather than
|
||||
as a single benchmark-wide average.
|
||||
|
||||
Usage: python3 scripts/docs/gallery_coverage_per_film.py --out docs_data/gallery_coverage_per_film.json
|
||||
"""
|
||||
import argparse
|
||||
import csv
|
||||
import json
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
import h5py
|
||||
|
||||
REPO = Path(__file__).resolve().parent.parent.parent
|
||||
sys.path.insert(0, str(REPO / "scripts" / "validation"))
|
||||
from identity import keys_for # noqa: E402
|
||||
|
||||
|
||||
def main():
|
||||
p = argparse.ArgumentParser()
|
||||
p.add_argument("--gallery", default=str(REPO / "experiments/galleries/gallery_LVFace-B_Glint360K.h5"))
|
||||
p.add_argument("--films", default=str(REPO / "experiments/manifests/films.json"))
|
||||
p.add_argument("--out", required=True)
|
||||
args = p.parse_args()
|
||||
|
||||
films = json.load(open(args.films))
|
||||
with h5py.File(args.gallery, "r") as f:
|
||||
names = [n.decode() if isinstance(n, bytes) else n for n in f["name"][:]]
|
||||
jids = [j.decode() if isinstance(j, bytes) else j for j in f["jellyfin_id"][:]]
|
||||
imdbs = [j.decode() if isinstance(j, bytes) else j for j in f["imdb_id"][:]]
|
||||
gallery_keys = set()
|
||||
for n, j, im in zip(names, jids, imdbs):
|
||||
gallery_keys |= keys_for(imdb_id=im, name=n, jellyfin_id=j)
|
||||
|
||||
out = []
|
||||
for film in films:
|
||||
xray_dir = REPO / film["xray"]
|
||||
id_to_name = {}
|
||||
with open(xray_dir / "people.csv", newline="", encoding="utf-8") as fh:
|
||||
for r in csv.DictReader(fh):
|
||||
nm = (r.get("name_id") or "").strip()
|
||||
if nm:
|
||||
id_to_name[nm] = (r.get("person") or "").strip()
|
||||
cast_keys = [keys_for(imdb_id=nm, name=name) for nm, name in id_to_name.items()]
|
||||
covered = sum(1 for ck in cast_keys if ck & gallery_keys)
|
||||
total = len(cast_keys)
|
||||
out.append({"film": film["name"], "cast_total": total, "covered": covered,
|
||||
"coverage_pct": round(covered / total * 100, 1) if total else 0.0})
|
||||
|
||||
out.sort(key=lambda x: x["coverage_pct"])
|
||||
Path(args.out).parent.mkdir(parents=True, exist_ok=True)
|
||||
json.dump(out, open(args.out, "w"), indent=1)
|
||||
for o in out:
|
||||
print(f"{o['film']:45s} {o['covered']:3d}/{o['cast_total']:3d} ({o['coverage_pct']}%)",
|
||||
file=sys.stderr)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,107 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
run_holdout_all_models.py — replay each model's own tuned full_exp config
|
||||
against the 5 held-out films, score with second_score.py, and dump a combined
|
||||
JSON. r50 is excluded (see docs/model-bakeoff.md: dropped from the detailed
|
||||
comparison, kept only in the calibration-curve chart).
|
||||
|
||||
This fills a real gap: the shipped report claimed "nothing in held-out
|
||||
validation contradicts the model choice" without ever running mbf/r18 on the
|
||||
held-out films — only LVFace had been checked.
|
||||
|
||||
Usage: python3 scripts/docs/run_holdout_all_models.py --out docs_data/holdout_all_models.json
|
||||
"""
|
||||
import argparse
|
||||
import json
|
||||
import subprocess
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
REPO = Path(__file__).resolve().parent.parent.parent
|
||||
sys.path.insert(0, str(REPO / "scripts" / "optimizer"))
|
||||
sys.path.insert(0, str(REPO / "scripts" / "validation"))
|
||||
from second_score import score_seconds # noqa: E402
|
||||
from sample_eval import load_gallery_keys # noqa: E402
|
||||
|
||||
MODELS = ["LVFace-B_Glint360K", "arcface_w600k_mbf", "arcface_r18"]
|
||||
|
||||
HELDOUT = [
|
||||
{"name": "Benny & Joon", "slug": "Benny___Joon",
|
||||
"xray": "experiments/xray/scene_level_movie_data_XRay_US/xrays/4808_Benny__Joon"},
|
||||
{"name": "Downton Abbey: A New Era", "slug": "Downton_Abbey__A_New_Era",
|
||||
"xray": "experiments/xray/scene_level_movie_data_XRay_US/xrays/19_Downton_Abbey_A_New_Era"},
|
||||
{"name": "Lovelace", "slug": "Lovelace",
|
||||
"xray": "experiments/xray/scene_level_movie_data_XRay_US/xrays/4108_Lovelace"},
|
||||
{"name": "The Many Saints of Newark", "slug": "The_Many_Saints_of_Newark",
|
||||
"xray": "experiments/xray/scene_level_movie_data_XRay_US/xrays/900_The_Many_Saints_Of_Newark"},
|
||||
{"name": "Valerian and the City of a Thousand Planets",
|
||||
"slug": "Valerian_and_the_City_of_a_Thousand_Plan",
|
||||
"xray": "experiments/xray/scene_level_movie_data_XRay_US/xrays/5312_Valerian_and_the_City_of_a_Thousand_Planets"},
|
||||
]
|
||||
|
||||
TRAINING = [
|
||||
{"name": "Café Society", "slug": "Café_Society",
|
||||
"xray": "experiments/xray/scene_level_movie_data_XRay_US/xrays/225_Cafe_Society"},
|
||||
{"name": "Lord of War", "slug": "Lord_of_War",
|
||||
"xray": "experiments/xray/scene_level_movie_data_XRay_US/xrays/2474_Lord_of_War"},
|
||||
{"name": "Scarface", "slug": "Scarface",
|
||||
"xray": "experiments/xray/scene_level_movie_data_XRay_US/xrays/197_Scarface"},
|
||||
{"name": "Sound of Metal", "slug": "Sound_of_Metal",
|
||||
"xray": "experiments/xray/scene_level_movie_data_XRay_US/xrays/6278_Sound_of_Metal"},
|
||||
]
|
||||
|
||||
|
||||
def main():
|
||||
p = argparse.ArgumentParser()
|
||||
p.add_argument("--out", required=True)
|
||||
p.add_argument("--work-dir", default="/tmp/holdout_all_models")
|
||||
p.add_argument("--films", choices=["heldout", "training"], default="heldout")
|
||||
args = p.parse_args()
|
||||
|
||||
work = Path(args.work_dir)
|
||||
work.mkdir(parents=True, exist_ok=True)
|
||||
film_set = HELDOUT if args.films == "heldout" else TRAINING
|
||||
|
||||
results = {}
|
||||
for model in MODELS:
|
||||
cfg = json.load(open(REPO / f"experiments/results/rep4_best_{model}_full_exp.json"))["best"]["config"]
|
||||
gallery = REPO / f"experiments/galleries/gallery_{model}.h5"
|
||||
results[model] = {"config": cfg, "films": {}}
|
||||
|
||||
for film in film_set:
|
||||
dump = REPO / f"experiments/dumps/{model}/dump_{film['slug']}.h5"
|
||||
if not dump.exists():
|
||||
print(f"SKIP {model}/{film['slug']}: no dump", file=sys.stderr)
|
||||
continue
|
||||
pred_path = work / f"pred_{model}_{film['slug']}.json"
|
||||
cmd = [
|
||||
"python3", "scripts/optimizer/replay.py",
|
||||
"--dump", str(dump), "--gallery", str(gallery),
|
||||
"--out", str(pred_path),
|
||||
"--prob-threshold", str(cfg["prob_threshold"]),
|
||||
"--anneal-sec", str(cfg["anneal_sec"]),
|
||||
"--extinction-sec", str(cfg["extinction_sec"]),
|
||||
"--expand-gallery",
|
||||
]
|
||||
print(f"RUN {model}/{film['slug']}...", file=sys.stderr)
|
||||
r = subprocess.run(cmd, cwd=REPO, capture_output=True, text=True, timeout=120)
|
||||
if r.returncode != 0:
|
||||
print(f"FAIL {model}/{film['slug']}: {r.stderr[-800:]}", file=sys.stderr)
|
||||
results[model]["films"][film["slug"]] = {"error": r.stderr[-500:]}
|
||||
continue
|
||||
|
||||
gk = load_gallery_keys(str(gallery))
|
||||
pred_json = json.loads(pred_path.read_text())
|
||||
m = score_seconds(pred_json, str(REPO / film["xray"]), gk)
|
||||
results[model]["films"][film["slug"]] = {"name": film["name"], **m}
|
||||
print(f" -> F1={m['f1']*100:.1f}% P={m['precision']*100:.1f}% "
|
||||
f"R={m['recall']*100:.1f}% misid={m['FPI_misid']}", file=sys.stderr)
|
||||
|
||||
Path(args.out).parent.mkdir(parents=True, exist_ok=True)
|
||||
with open(args.out, "w") as f:
|
||||
json.dump(results, f, indent=1)
|
||||
print(f"wrote {args.out}", file=sys.stderr)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -106,8 +106,12 @@ def evaluate(cfg, films, build_dir, step=None):
|
||||
"""Objective = MACRO-mean over films of each film's duration-weighted per-scene F1.
|
||||
|
||||
Each film's replay runs in a subprocess (timeout-guarded) to survive the
|
||||
intermittent ROCm teardown deadlock. A film whose replay times out is dropped
|
||||
from the average rather than hanging the whole sweep.
|
||||
intermittent ROCm teardown deadlock. If ANY film's replay times out, this
|
||||
evaluation is scored f1=0.0 (see below) rather than averaging over the
|
||||
survivors — a partial-coverage eval must never look better than a complete
|
||||
one, or DE will converge onto configs that make the hardest film time out.
|
||||
(An earlier version averaged over survivors, which silently rewarded
|
||||
truncation; the rep4 `mbf_full_noexp` winner was one such corrupted eval.)
|
||||
|
||||
UNIFORM PER-SECOND scoring (second_score.py): every second of the film is sampled;
|
||||
GT(t) = the cast of the X-Ray scene containing t, Pred(t) = actors whose presence
|
||||
@@ -135,9 +139,21 @@ def evaluate(cfg, films, build_dir, step=None):
|
||||
with ThreadPoolExecutor(max_workers=REPLAY_WORKERS) as ex:
|
||||
per_film = [m for m in ex.map(_one, films) if m is not None]
|
||||
n = len(per_film)
|
||||
if not n:
|
||||
return {"precision": 0.0, "recall": 0.0, "f1": 0.0, "agreement": 0.0,
|
||||
"TPI": 0, "FPI": 0, "FPI_misid": 0, "FN": 0}
|
||||
n_expected = len(films)
|
||||
# Incomplete coverage (a replay timed out) is scored as a failure, not
|
||||
# averaged over survivors: dropping the hardest film would otherwise inflate
|
||||
# the score and let DE reward exactly the configs that cause timeouts. We
|
||||
# still record the real survivor counts so a truncated eval is diagnosable
|
||||
# in the trajectory (f1=0.0, films_scored < films_expected).
|
||||
if n < n_expected:
|
||||
agg = {"precision": 0.0, "recall": 0.0, "f1": 0.0, "agreement": 0.0,
|
||||
"TPI": sum(m["TPI"] for m in per_film),
|
||||
"FPI": sum(m["FPI"] for m in per_film),
|
||||
"FPI_misid": sum(m["FPI_misid"] for m in per_film),
|
||||
"FN": sum(m["FN"] for m in per_film)}
|
||||
agg["films_scored"] = n
|
||||
agg["films_expected"] = n_expected
|
||||
return agg
|
||||
return {"precision": sum(m["precision"] for m in per_film) / n,
|
||||
"recall": sum(m["recall"] for m in per_film) / n,
|
||||
"f1": sum(m["f1"] for m in per_film) / n,
|
||||
@@ -145,7 +161,9 @@ def evaluate(cfg, films, build_dir, step=None):
|
||||
"TPI": sum(m["TPI"] for m in per_film),
|
||||
"FPI": sum(m["FPI"] for m in per_film),
|
||||
"FPI_misid": sum(m["FPI_misid"] for m in per_film),
|
||||
"FN": sum(m["FN"] for m in per_film)}
|
||||
"FN": sum(m["FN"] for m in per_film),
|
||||
"films_scored": n,
|
||||
"films_expected": n_expected}
|
||||
|
||||
|
||||
def main():
|
||||
|
||||
Reference in New Issue
Block a user