diff --git a/.gitignore b/.gitignore
index d0015ad..ac33e45 100644
--- a/.gitignore
+++ b/.gitignore
@@ -14,6 +14,11 @@ compile_commands.json
*.so
*.dylib
*.json
+# Exception: small, curated result summaries backing specific numbers quoted
+# in docs/ (cross-model held-out scores, per-film training breakdown, gallery
+# coverage). Regenerate with scripts/docs/run_holdout_all_models.py and
+# scripts/docs/gallery_coverage_per_film.py.
+!docs_data/*.json
# Video files
*.mp4
*.mkv
diff --git a/docs/best-model.md b/docs/best-model.md
index 8bd1c9d..5f433c2 100644
--- a/docs/best-model.md
+++ b/docs/best-model.md
@@ -1,70 +1,103 @@
# Which embedding model is best?
-Four candidates went into the bake-off: three ArcFace variants (w600k-R50,
-R18, w600k-MBF) and LVFace-B (Glint360K), a Vision-Transformer embedder that's
-a drop-in replacement for ArcFace's `[N,3,112,112]` input / 512-d output. The
-open question: is LVFace (455MB) actually better, or just the biggest?
+Three ArcFace variants (w600k-R50, R18, w600k-MBF) and LVFace-B (Glint360K,
+455MB) were compared. r50 is excluded from the training/held-out comparison
+below; its gallery has roughly 30% fewer reference images per actor than the
+other three on the identical source photos, which confounds a direct score
+comparison (see [the full experiment log](model-bakeoff.md) for detail). It
+remains in the calibration comparison, which does not depend on the gallery
+image count.
## First signal: calibration curves
Each gallery carries a fitted Platt sigmoid `P(match | cosine similarity) =
-σ(a·sim + b)`, embedded directly in the gallery's HDF5 file
-([`src/gallery/gallery_calibration.hpp`](https://REPOLINK/src/gallery/gallery_calibration.hpp)). This is a property of the embedding
-space alone — computed from intra/inter-actor reference-image pairs, no
-tracking or scene logic involved — so it's a clean first read on discriminative
-power before running a single benchmark.
+σ(a·sim + b)`, stored directly in the gallery HDF5
+([`src/gallery/gallery_calibration.hpp`](https://REPOLINK/src/gallery/gallery_calibration.hpp)).
+This is a property of the embedding space alone, computed from intra- and
+inter-actor reference-image pairs with no tracking or scene logic involved,
+so it is a clean first read on discriminative power before running a
+benchmark.

-| model | `a` (steepness) | boundary at P=0.5 |
+| model | a (steepness) | boundary at P=0.5 |
|---|---|---|
-| **LVFace-B Glint360K** | **17.7** | **sim 0.228** |
+| LVFace-B Glint360K | 17.7 | sim 0.228 |
| ArcFace w600k-MBF | 16.2 | sim 0.267 |
| ArcFace w600k-R50 | 15.4 | sim 0.301 |
| ArcFace R18 | 15.3 | sim 0.309 |
-LVFace has both the steepest transition and the lowest decision boundary — it
-separates same-actor from different-actor reference pairs more confidently, at
-a *lower* similarity threshold, than any ArcFace variant. That's a genuine
-head start before the tracking/scoring pipeline is even involved.
+LVFace has both the steepest transition and the lowest decision boundary,
+separating same-actor from different-actor reference pairs more confidently
+at a lower similarity than any ArcFace variant.
-## Second signal: F1 on the actual benchmark
+## Second signal: held-out F1
-Best full-gallery (no cast-restriction) result per model, from the 16-combo
-bake-off matrix ([full experiment log](model-bakeoff.md)):
+Each model's own tuned `full_exp` config, replayed against the 5 films the
+optimizer never saw and scored the same way:
+
+| film | LVFace F1 | mbf F1 | r18 F1 |
+|---|---|---|---|
+| Benny & Joon | 83.0% | 78.5% | 77.1% |
+| Lovelace | 77.5% | 73.7% | 72.2% |
+| Valerian and the City of a Thousand Planets | 74.1% | 70.2% | 71.0% |
+| Downton Abbey: A New Era | 56.2% | 55.0% | 53.0% |
+| The Many Saints of Newark | 46.3% | 44.5% | 42.1% |
+| **macro average** | **67.4%** | **64.4%** | **63.1%** |
+
+LVFace scores highest on all 5 held-out films; the ranking never flips
+between models. Total misID count across the 5 films: LVFace 1032, mbf
+2197, r18 1224. LVFace has less than half mbf's misID total and still
+scores higher on every film.
+
+Held-out results are stronger evidence than training results, because
+training numbers can reflect what the optimizer was tuned to fit rather
+than general performance. On training data, the ordering is not as clean:
+
+| film | LVFace F1 | mbf F1 | r18 F1 | best |
+|---|---|---|---|---|
+| Café Society | 68.1% | 62.2% | 60.1% | LVFace |
+| Lord of War | 75.6% | 77.2% | 75.6% | mbf |
+| Scarface | 71.5% | 68.6% | 64.1% | LVFace |
+| Sound of Metal | 78.8% | 76.5% | 71.6% | LVFace |
+
+mbf beats LVFace on Lord of War (77.2% vs 75.6%), the only film in either
+table where LVFace does not score highest. LVFace's training-set macro
+average (75.3%, see [the full experiment log](model-bakeoff.md)) is not a
+uniform win across every film it contributes to; the held-out result, where
+LVFace wins all 5 films outright, is the stronger claim.
+
+This reverses an earlier, superseded benchmarking pass that used a
+scene-union metric and found the three models statistically
+indistinguishable (around 85% each), concluding LVFace was not worth its
+size. That metric masked out-of-cast false positives behind a
+gallery-intersect-cast recall filter; the per-second metric used here does
+not.
+
+## Full training-matrix picture
+
+
+
+Best full-gallery combo per model (all three are `full_exp`), from the
+training matrix in [the full experiment log](model-bakeoff.md):
| model | F1 | P | R | misID |
|---|---|---|---|---|
-| **LVFace-B Glint360K** | **75.3%** | 89.7% | **65.4%** | 232 |
-| ArcFace w600k-MBF | 74.2% | 87.4% | 64.4% | 57 |
+| LVFace-B Glint360K | 75.3% | 89.7% | 65.4% | 232 |
+| ArcFace w600k-MBF | 72.0% | 87.7% | 61.4% | 240 |
| ArcFace R18 | 69.1% | 87.6% | 57.7% | 242 |
-| ArcFace w600k-R50 | 68.5% | 94.0% | 54.1% | 150 |
-The full 16-combo picture makes the model ordering visible at a glance — LVFace
-(yellow) tops both the restricted and full columns, and R18 (green) props up
-the bottom of the full-gallery ranking:
+LVFace leads within both the restricted and full gallery modes, visible
+directly in the chart above without reading the table. The three models'
+misID counts on the full gallery are nearly identical (232/240/242); LVFace's
+lead here is a precision-and-recall lead, not a misID one.
-
+## Operational note
-LVFace wins outright, with the highest recall of any full-mode combo. This
-reverses an earlier conclusion from a prior (superseded) benchmarking pass
-using a scene-union metric, which found the three models statistically
-indistinguishable (~85% each) and concluded LVFace wasn't worth its size — that
-metric hid out-of-cast false positives behind a gallery∩cast recall mask (see
-[the prior optimizer round](optimizer-experiments.md)); the per-second metric
-used here does not.
-
-Held-out validation (5 films never seen by the optimizer) confirms LVFace's
-lead holds up out of sample — see the
-[LVFace deep dive](lvface-deep-dive.md) for the full breakdown, including
-where it fails.
-
-## Caveat: model choice is an operational change
-
-Switching the default embedder isn't just flipping a config value — the
-gallery itself is model-specific (embeddings from different models aren't
-comparable), so any existing gallery built against ArcFace w600k-R50 needs to
-be rebuilt from source images against LVFace before the new default takes
-effect. [`scripts/optimizer/reembed_gallery.py`](https://REPOLINK/scripts/optimizer/reembed_gallery.py)
+Switching the default embedder is not a config change alone; the gallery
+is model-specific, since embeddings from different models are not
+comparable. Any existing gallery built against a different model must be
+rebuilt from source images before the new default takes effect.
+[`scripts/optimizer/reembed_gallery.py`](https://REPOLINK/scripts/optimizer/reembed_gallery.py)
does this from a reference gallery's cached source images without
re-downloading anything.
diff --git a/docs/gallery-scope.md b/docs/gallery-scope.md
index 5833381..62c50b7 100644
--- a/docs/gallery-scope.md
+++ b/docs/gallery-scope.md
@@ -1,65 +1,69 @@
-# Whole gallery vs. limited (cast-restricted) gallery
+# Whole gallery vs. cast-restricted gallery
-Two ways to run the matcher: **full** scores every detected face against the
-entire library gallery (2418 actors across the 9-film benchmark set); **restricted**
-pre-filters each film's gallery down to just its Jellyfin-credited cast (typically
-~15 top-billed actors) before the matcher ever runs.
+Two ways to run the matcher. Full mode scores every detected face against
+the entire 2418-actor gallery. Restricted mode pre-filters each film's
+gallery down to just its Jellyfin-credited cast (typically around 15
+top-billed actors) before the matcher runs.
-## The result
+## Result
-Averaged across all 4 models and both expansion settings, on the 4 bake-off training
-films:
+Averaged across the 3 compared models (r50 excluded, see
+[the full experiment log](model-bakeoff.md)) and both expansion settings, on
+the 4 training films:
-| scope | F1 | P | R | total misID (8 evals) |
+| scope | F1 | P | R | total misID |
|---|---|---|---|---|
-| full | 71.2% | 91.1% | 59.0% | 1073 |
-| **restricted** | **74.5%** | 92.2% | **62.9%** | **329** |
+| full | 71.1% | 89.6% | 59.6% | 1121 |
+| restricted | 75.9% | 90.4% | 65.6% | 299 |
-This is not a precision/recall trade — restriction wins on every axis at once:
-**+3.3pp F1, +3.9pp recall, and less than a third the total misIDs.** Fewer
+Restriction improves every metric at once, not a precision/recall trade:
++4.8pp F1, +6.0pp recall, roughly a quarter the total misIDs. Fewer
candidates in the matcher's search space means fewer opportunities for a
-look-alike false match (an actor who happens to share enough facial structure
-with someone in the film, but isn't actually in it), and the recall gain shows
-it isn't costing real detections to get there.
+lookalike false match, and the recall gain shows this does not cost real
+detections.
-Per-model, every single model's best-scoring combo in the full 16-way matrix is
-a `restricted` variant — visible directly in the ranking below (filled dots =
-restricted, open = full; the filled dots cluster at the top for every color):
+Every model's best-scoring combo in the training matrix uses the
+restricted gallery:
-
+
-See the full table in the
-[bake-off experiment log](model-bakeoff.md). Two
-combos hit **zero** true out-of-cast misidentifications:
-`arcface_w600k_mbf_restricted_exp` (F1 76.5%) and, in full mode,
-`LVFace-B_Glint360K_full_noexp` (F1 72.4%) — restriction isn't the only way to
-reach misid=0, but it's the more reliable one.
+See [the full experiment log](model-bakeoff.md) for the complete table. One
+combo reaches zero true out-of-cast misidentifications,
+`arcface_w600k_mbf_restricted_exp` (F1 76.2%), and it is a restricted one,
+consistent with restriction, not expansion, being what suppresses cross-film
+confusions.
-## Why this isn't the shipped default
+The restriction effect (+4.8pp averaged across models) is larger than the
+model-choice effect: LVFace beats r18 by 6.2pp in full mode but beats mbf by
+3.3pp. Restriction is the single strongest lever in the matrix.
-Cast-restriction is implemented today only as an **offline optimizer technique**
+## Why this is not the shipped default
+
+Cast restriction is implemented today only as an offline optimizer
+technique
([`scripts/optimizer/cast_restrict.py`](https://REPOLINK/scripts/optimizer/cast_restrict.py)):
-it pre-builds a filtered gallery file
-per film, using Jellyfin's own cast list, before the benchmark ever calls the
-matcher. There's no runtime "restrict matching to this title's credited cast"
-switch in the shipped application — `scene_analyze` always matches against
-whatever single gallery file it's given.
+it pre-builds a filtered gallery file per film using Jellyfin's cast list
+before the benchmark calls the matcher. There is no runtime "restrict to
+this title's credited cast" switch in the shipped application;
+`scene_analyze` always matches against whatever single gallery file it is
+given.
-Building that as a real feature would need, at minimum:
+Building this as a real feature requires:
-- A live Jellyfin cast lookup at analysis time (the title is already known —
- [`scripts/run_from_jellyfin.py`](https://REPOLINK/scripts/run_from_jellyfin.py)
- already does this same lookup for its own
- `filter_gallery`-based restriction path, just not wired into `scene_analyze`
- itself as a first-class option).
-- A decision on the *fallback*: what happens to a real, uncredited cameo
- (see the Germar Terrell Gardner case in the LVFace deep-dive) if the gallery
- never includes them at all?
-- Regenerating the restricted-gallery cache whenever the title's Jellyfin cast
- list changes.
+- A live Jellyfin cast lookup at analysis time. The title is already known,
+ and [`scripts/run_from_jellyfin.py`](https://REPOLINK/scripts/run_from_jellyfin.py)
+ already performs this lookup for its own `filter_gallery`-based
+ restriction path; it is not wired into `scene_analyze` as a first-class
+ option.
+- A decision on the fallback case: what happens to a real, uncredited
+ cameo (see the Germar Terrell Gardner and Talia Balsam cases in the
+ [LVFace deep dive](lvface-deep-dive.md#where-lvface-beat-x-ray)) if the
+ restricted gallery never includes them at all.
+- Regenerating the restricted-gallery cache whenever a title's Jellyfin
+ cast list changes.
-This is why the shipped [`src/config.hpp`](https://REPOLINK/src/config.hpp)
-defaults use the `full`-mode winner
-(`LVFace-B_Glint360K_full_exp`, F1 75.3% training / 67.4% held-out macro) rather
-than the higher-scoring `restricted_exp` (78.3%) — the 78.3% number describes a
-capability the app doesn't have yet, not what actually ships.
+The shipped [`src/config.hpp`](https://REPOLINK/src/config.hpp) defaults use
+the full-mode winner (`LVFace-B_Glint360K_full_exp`, F1 75.3% training,
+67.4% held-out macro) rather than the higher-scoring `restricted_exp`
+(78.3%), because 78.3% describes a capability the application does not
+have yet.
diff --git a/docs/index.md b/docs/index.md
index ef0c04a..54437c7 100644
--- a/docs/index.md
+++ b/docs/index.md
@@ -1,27 +1,32 @@
# scene-actor-extraction
-A face-recognition pipeline that finds when each actor appears on screen in a
-film or TV episode — built on [KPN++](https://gitea.tourolle.paris/dtourolle/KPN)
-(a C++20 Kahn Process Network library) for the detect → track → match → scene
-pipeline, with a Jellyfin-integrated gallery and an X-Ray-validated optimizer.
+A face-recognition pipeline that finds when each actor appears on screen in
+a film or TV episode, built on [KPN++](https://gitea.tourolle.paris/dtourolle/KPN)
+(a C++20 Kahn Process Network library) for the detect, track, match, and
+scene pipeline, with a Jellyfin-integrated gallery and an X-Ray-validated
+optimizer.
-This is a perfect X-Ray second, on a film the optimizer never saw:
+This is a correctly scored second from a held-out film, one the optimizer
+never saw during tuning:

-Every visible face named at 100% — Chris Noth, Hank Azaria, Bobby Cannavale —
-the background extra honestly left unnamed, and the two credited cast without
-a visible face correctly carried as present off-screen by the tracker's
-presence windows. That's the pipeline exactly reproducing Amazon X-Ray's
-record for this second.
+Every visible face is named at 100% confidence (Chris Noth, Hank Azaria,
+Bobby Cannavale), the background extra is correctly left unnamed, and the
+two credited cast members without a visible face are correctly reported
+present but not visible. This matches Amazon X-Ray's own record for this
+second exactly.
-It doesn't always go like that: the hardest held-out film scores 46% F1, and
-the report is honest about *why* — one tunable trade (extinction bridging at
-hard cuts), one structural ceiling (X-Ray credits people whose faces never
-appear), and a few cases where the pipeline is right and X-Ray is wrong. The
-evidence for all of it is in the pages below.
+Results are not uniform across films. The hardest held-out film scores 46%
+F1. This report documents why: one tunable trade (extinction bridging at
+hard cuts), one structural limit (X-Ray credits people whose faces never
+appear on screen), and a small number of cases where the pipeline is
+correct and X-Ray's ground truth is not. Read
+[how we score against X-Ray](methodology.md) first. X-Ray's ground truth is
+scene-level; the pipeline's output is per-second. That difference shapes
+every finding below.
-## Start here — four questions this bake-off answers
+## Findings
@@ -29,55 +34,51 @@ evidence for all of it is in the pages below.
---
- Calibration curves first (discriminative power, independent of any
- threshold), then F1 on the actual benchmark. LVFace-B Glint360K wins
- both.
+ Calibration curves first, independent of any threshold, then held-out
+ F1 across three models. LVFace-B Glint360K wins both, and wins on every
+ held-out film.
- :material-filter:{ .lg .middle } **[Whole vs. cast-restricted gallery](gallery-scope.md)**
---
- Restricting the matcher to a film's credited cast is a clean win on
- every axis (+3.3pp F1, less than a third the misIDs) — but isn't a
- shipped runtime feature yet.
+ Restricting the matcher to a film's credited cast improves F1,
+ recall, and misID rate at once, but is not a shipped runtime feature
+ yet.
- :material-account-convert:{ .lg .middle } **[Does pose expansion help?](pose-expansion.md)**
---
- A convincing training-set effect that didn't reproduce on 5 held-out
- films once two methodology bugs were caught and fixed. An honest null
- result, not a forced narrative.
+ A training-set effect that did not reproduce on 5 held-out films once
+ two methodology bugs in the comparison harness were found and fixed.
- :material-magnify-expand:{ .lg .middle } **[Deep dive: LVFace-B Glint360K](lvface-deep-dive.md)**
---
- The held-out generalization gap, how the error budget decomposes
- (extinction bridging at hard cuts, X-Ray's scene-membership vs.
- on-screen-face ceiling), and the frames where the pipeline is right
- and the ground truth is wrong.
+ The held-out generalization gap, the two mechanisms behind its errors,
+ and every distinct case where it names someone outside the film's
+ credited cast.
-## The full technical log
+## Full experiment log
-- **[Model bake-off + threshold re-tune](model-bakeoff.md)** —
- the complete experiment log behind the four pages above: the ROCm teardown
- deadlock root cause and fix, DE concurrency tuning, the full 16-combo
- results table, and every caveat. This is where the shipped
- [`src/config.hpp`](https://REPOLINK/src/config.hpp) defaults come from.
-- **[Optimizer experiments (prior round)](optimizer-experiments.md)** — the
- earlier scene-union-metric tuning pass, superseded by the per-second metric
- used in the bake-off but kept for the ground-truth/architecture background.
-- **[Service conversion (proposal)](service-conversion.md)** — design sketch
- for a native idle-GPU worker gated on screen lock, not yet built.
+- **[Full experiment log](model-bakeoff.md)**: the complete log behind the
+ four pages above, including how replaying against cached embeddings
+ inside the same KPN network makes a full model and configuration
+ comparison practical, the full results table, and every caveat. This is
+ where the shipped [`src/config.hpp`](https://REPOLINK/src/config.hpp)
+ defaults come from.
+- **[Service conversion (proposal)](service-conversion.md)**: design
+ sketch for a native idle-GPU worker gated on screen lock, not yet built.
## Reproducing the benchmarks
-Gallery `.h5` files, embedding dumps, the X-Ray corpus, montage frame images,
-and DE trajectories are not committed to this repository — they're pushed to
-the Gitea package registry and pulled on demand:
+Gallery `.h5` files, embedding dumps, the X-Ray corpus, montage frame
+images, and DE trajectories are not committed to this repository. They are
+pushed to the Gitea package registry and pulled on demand:
```bash
scripts/artifacts/pull_artifacts.sh galleries
@@ -86,4 +87,5 @@ scripts/artifacts/pull_artifacts.sh montage-frames
```
See [`scripts/artifacts/push_artifacts.sh`](https://REPOLINK/scripts/artifacts/push_artifacts.sh)
-for the upload side (requires a `GITEA_TOKEN` with package write scope).
+for the upload side, which requires a `GITEA_TOKEN` with package write
+scope.
diff --git a/docs/lvface-deep-dive.md b/docs/lvface-deep-dive.md
index 41ccf36..1e2270f 100644
--- a/docs/lvface-deep-dive.md
+++ b/docs/lvface-deep-dive.md
@@ -1,52 +1,53 @@
# Deep dive: LVFace-B Glint360K
-LVFace won the model bake-off (see [Which model is best?](best-model.md)) and is
-the shipped default embedder. This page is the honest accounting of how it
-actually performs — what a good second looks like, where the errors actually
-come from, and two cases where the ground truth itself is wrong and LVFace is
-right.
+LVFace won the model comparison (see [Which model is best?](best-model.md))
+and is the shipped default embedder. This page reports how it performs in
+detail: a baseline of correct output, the two mechanisms behind its errors,
+and every distinct case where it names someone who is not in the film's
+credited cast.
+
+Read [How we score against X-Ray](methodology.md) first. X-Ray's ground truth
+is scene-level, not per-frame. A name marked correct in the Offscreen column
+below is the pipeline correctly reporting scene membership, not a workaround.
!!! note "How to read the frames on this page"
- The top is the film frame, with a box and name on every face the pipeline
- identified. The bottom panels are the per-second verdict against X-Ray:
- **Onscreen** lists faces named in the frame, **Offscreen** lists cast
- X-Ray marks present in the scene without a visible face — presence
- carried by the tracker's windows, not by a detection. Colors are the
- score: **green** = correct (TPI),
- **orange** = wrong (FPI),
- **blue** = missed (FN).
+ The top of each image is the film frame, with a box and name on every
+ face the pipeline matched to a real detection. The panels below are the
+ per-second result against X-Ray. **Onscreen** lists names attached to a
+ visible face this second. **Offscreen** lists names the pipeline reports
+ present without a currently visible face. Colors mark the verdict:
+ **green** correct (TPI),
+ **orange** wrong (FPI),
+ **blue** missed (FN).
-## What good looks like
+## Baseline: correctly scored seconds

-Six faces on screen, all six named correctly — including Penelope Wilton at the
-edge of the pews and a half-occluded Michelle Dockery — while thirteen more
-cast members X-Ray marks present in the scene are correctly carried as
-"Offscreen" by their presence windows. One miss in the whole frame: Maggie
-Smith (blue). Score for this second: 0.86.
+Six faces on screen, all six named correctly, including Penelope Wilton at
+the edge of the pews and a partly occluded Michelle Dockery. Thirteen more
+cast members X-Ray lists as present in the scene are correctly reported
+Offscreen. One miss: Maggie Smith (blue). Score for this second: 0.86.

-The same film's funeral gathering: mourning dress, hats, half the faces turned.
-**Nineteen of the twenty cast X-Ray lists for this scene are scored correctly**
-— seven named on screen at up to 100% confidence, twelve more correctly held
-as present off-screen.
-
-And the pipeline doesn't need the face to be *real*:
+The same film's funeral scene: dark clothing, hats, half the faces turned
+away. Nineteen of the twenty cast members X-Ray lists for this scene score
+correct: seven named on screen at up to 100% confidence, twelve more reported
+correctly as present but not visible.

-That's Herbie Hancock at 98% — as a face on a *screen inside the movie*, over a
-sci-fi HUD overlay, during a video call in Valerian. A face is a face, whether
-it's in the room or on the bridge's comms display.
+The pipeline does not require a live face. This is Herbie Hancock at 98%
+confidence, identified from a face displayed on a screen inside the film, on
+a video call under a science-fiction HUD overlay.
## Training vs. held-out: the generalization gap
-The shipped config (`prob_threshold=0.754, anneal_sec=35.54,
-extinction_sec=57.43, expand_gallery=true`) was tuned against 4 films. Scored
-against the 5 films the optimizer never saw:
+The shipped config (`prob_threshold=0.754`, `anneal_sec=35.54`,
+`extinction_sec=57.43`, `expand_gallery=true`) was tuned on 4 films. Scored
+on the 5 films the optimizer never saw:

@@ -56,139 +57,241 @@ against the 5 films the optimizer never saw:
| Lovelace | 77.5% | 90.3% | 67.9% | 14990 | 1085 | 58 | 7085 |
| Valerian and the City of a Thousand Planets | 74.1% | 97.1% | 60.0% | 18663 | 548 | 0 | 12467 |
| Downton Abbey: A New Era | 56.2% | 97.8% | 39.4% | 52027 | 1173 | 0 | 80084 |
-| **The Many Saints of Newark** | **46.3%** | **54.7%** | 40.1% | 15922 | 4394 | **974** | 23791 |
-| **macro average** | **67.4%** | 85.8% | 57.0% | | | | |
+| The Many Saints of Newark | 46.3% | 54.7% | 40.1% | 15922 | 4394 | 974 | 23791 |
+| macro average | 67.4% | 85.8% | 57.0% | | | | |
-**67.4% held-out vs. 75.3% on training** — an ~8pp drop, and a **37pp spread
-between the best and worst held-out film**. The config does not generalize
-uniformly, and the spread traces to two mechanisms, both visible frame by
-frame below.
+The `P` column is misID-weighted (each out-of-film name counts 10x in the
+denominator; see [methodology](methodology.md#precision-recall-and-the-misid-weighting)).
+That weighting is why Many Saints reads 54.7% here despite naming mostly real,
+present faces: its raw (unweighted) precision is **78.4%**, and the gap is
+entirely its 974 misIDs paying the 10x penalty. The three zero-misID films
+(Benny & Joon, Downton, Valerian) have identical weighted and raw precision;
+Lovelace, with 58 misIDs, sits 3pp below its raw 93.3%.
-## Mechanism 1: extinction bridging — usually right, wrong at hard cuts
+Held-out F1 is 67.4%, against 75.3% on training, an 8pp drop. The spread
+between the best and worst held-out film is 37pp. This is not unique to
+LVFace: [the full experiment log](model-bakeoff.md#held-out-validation-all-3-models)
+shows mbf and r18 with the same shape of spread on the same films, at a
+uniformly lower level. Two mechanisms explain the spread. Both are shown
+below with frame-level evidence.
-The extinction window keeps an identity alive through seconds where no face is
-detectable. **Most of the time this is exactly what you want**, and it's where
-a lot of the TPI count comes from:
+## Mechanism 1: extinction bridging
+
+The extinction window keeps a name reported as present for up to
+`extinction_sec` after its last real detection. This is deliberate: most
+gaps in face visibility are short (a turned head, an occlusion, a cut to a
+reaction shot), and the window bridges them.

-Lovelace's polygraph scene: only Eric Roberts and Amanda Seyfried have visible
-faces, but X-Ray lists eight cast present — and all eight score green, the
-other six correctly carried by presence windows through a scene where the
-camera never shows them. A perfect second, and the extinction/anneal machinery
-is *why*.
+Lovelace's polygraph scene: only Eric Roberts and Amanda Seyfried have
+visible faces. X-Ray lists eight cast members present. All eight score
+correct; the other six are reported Offscreen through a stretch where the
+camera never shows them. The extinction window is why.
-The same mechanism has a failure case: a hard cut into long faceless footage.
-Both Many Saints of Newark (974 misIDs) and Downton Abbey (FN=80084, the worst
-recall of the five) are dominated by it — verified directly against the raw
-per-frame stream and the HDF5 dump's own detection counts, not inferred from
-the score alone. **This is not a malfunction**: the tracker is doing exactly
-what its window is for; the footage just stops cooperating. In the debug
-overlay (which draws a bridged identity's last-known bbox, unlike the shipped
-output, which emits presence windows and no boxes at all) the bridged state is
-visible spatially:
-
-
-*Debug-overlay rendering (`dump_error_frames.py --raw`): "Jon Bernthal", "Joey
-Diaz" and "Billy Magnussen" are extinction-bridged identities from the previous
-shot, drawn frozen over the wall and the hanging plates. Frame
-`many_saints/fpi/fpi_t03543.jpg`, `montage-frames` artifact package.*
-
-The cost is measurable, not just visible. Downton Abbey's hard cut into its
-closing credits, plotting the dump's own per-second `face_count` (detector
-output, independent of the tracker) against what the tracker reports:
+The same mechanism fails at a hard cut into a long stretch with no faces at
+all. Downton Abbey's recall (39.4%, the worst of the five held-out films) is
+dominated by this failure. It is verified directly against the raw
+per-frame stream and the dump's own detection counts, not inferred from the
+score. Plotting the dump's per-second `face_count` (detector output,
+independent of the tracker) against what the tracker reports, through
+Downton Abbey's hard cut into its closing credits:

-From the cut onward the detector sees **zero faces for nearly a minute** — and
-the tracker keeps reporting the last shot's 15 identities the whole time
-(verified for Hugh Bonneville: bbox `(1743.2, 0.0, 171.3, 317.8)`, unchanged to
-the pixel, at every sampled second for 57+ seconds). The staircase at the right
-edge is the extinction window expiring actor by actor. That plateau is
-`SceneTrackerFunc::active_[actor_idx].last_bbox`
+From the cut onward the detector reports zero faces for close to a minute.
+The tracker continues reporting the previous shot's 15 identities for the
+same span (verified for Hugh Bonneville: bbox `(1743.2, 0.0, 171.3, 317.8)`,
+unchanged to the pixel, at every sampled second for 57 seconds). The
+staircase at the right edge is the extinction window expiring, actor by
+actor. This is `SceneTrackerFunc::active_[actor_idx].last_bbox`
([`src/nodes/scene_tracker_node.hpp`](https://REPOLINK/src/nodes/scene_tracker_node.hpp))
-re-emitted as designed: `extinction_sec=57.4` was tuned long because bridging
-wins on most footage (see the polygraph frame above) — the training films just
-never contained a faceless stretch long enough to show the cost side, and the
-held-out set did.
+re-emitted as designed. `extinction_sec=57.4` was tuned long because
+bridging is correct on most footage, as in the polygraph scene above. The
+training films did not contain a faceless stretch long enough to expose the
+cost side; the held-out set did.
-The same track-continuation machinery has one milder spatial artifact, worth
-knowing when reading these frames:
+The extinction window is a scoring concept, not something drawn on screen.
+The shipped output is presence windows with no bounding boxes. Even the
+debug overlay used for this report never draws a box for a bridged name: a
+name inside its extinction window with no current detection appears only as
+a name in the Offscreen column, the same as every correctly bridged name
+above.
+
+A related, smaller effect shows up at rapid cuts:

-*Café Society (a training film), a shot/reverse-shot dialog: that is Steve
-Carell wearing both his own label and Jesse Eisenberg's.*
-At a rapid cut, the previous shot's track can linger for a beat at nearly the
-same screen position the new face occupies — here Jesse Eisenberg's box from
-the counter-shot lands on Steve Carell. Note what the score panel says,
-though: both actors are green, because both *are* present in this dialog
-scene per X-Ray. The spatial label is briefly wrong; the per-second presence
-claim — the thing the pipeline actually ships — is right. It's the same trade
-as the extinction window: track continuation smooths over cuts, and 1 fps
-sampling occasionally catches the seam.
+Café Society (a training film), a shot/reverse-shot dialog. The box on Steve
+Carell's face carries two labels: his own, and Jesse Eisenberg's, left over
+from the counter-shot a moment earlier. Both names score correct, because
+both actors are present in this scene per X-Ray. The box position is
+briefly wrong; the presence claim, which is what the pipeline ships, is
+right.
## Mechanism 2: the face-vs-presence ceiling
-Downton Abbey's recall didn't collapse because faces were misread — it
-collapsed because for most of its 80084 FN-seconds there was **no face to
-read**:
+Downton Abbey's recall did not collapse because faces were misread. It
+collapsed because for most of its 80084 false-negative seconds there was no
+face to read.

-A newsreel crew hauls equipment through the hall: X-Ray credits 22 cast as
-present in this scene; not one face looks at the camera. Eight are still
-scored green (windows bridging from adjacent shots) — the other fourteen are
-blue FNs that no face-recognition pipeline could ever recover. X-Ray encodes
-*scene membership*; the pipeline measures *on-screen faces*. In ensemble films
-those two definitions diverge massively, and that gap — not identification
-error — is most of what the FN column counts.
+A newsreel crew moves equipment through the hall. X-Ray credits 22 cast
+members as present in this scene. None face the camera. Eight still score
+correct, carried by presence windows from adjacent shots. The other fourteen
+are missed, and no face-recognition system can recover them, because there
+is no face in the frame. X-Ray records scene membership; the pipeline
+measures visible faces. In ensemble scenes these two quantities diverge, and
+that gap accounts for most of the false-negative count.
-
+## Every distinct out-of-cast name
-Same ceiling from the other side: Michela De Rossi in frame but turned away,
-five cast correctly bridged as offscreen (green), four blue FNs — and one
-orange we'll come back to below.
+Many Saints of Newark has the largest misID count of any held-out film: 974
+seconds, weighted. Rather than characterize this from a single frame, the
+raw replay stream was searched directly for every name the pipeline reports
+that is not in the film's credited cast. The same search was run on all 9
+films in the benchmark, one rule applied uniformly: **find the first second
+each distinct out-of-cast name appears, and render that exact second.**
+
+Five films produce no such name anywhere in their runtime: Benny & Joon,
+Café Society, Downton Abbey, Sound of Metal, Valerian. Zero out-of-cast
+names across their entire length. Four films produce nine distinct names
+between them, shown below in full, not a sample.
+
+### The Many Saints of Newark: 4 names
+
+
+
+Germar Terrell Gardner, t=848s, 78% confidence. A real, clearly visible
+background actor. He is not in X-Ray's cast list for this film, but he is
+credited in Jellyfin's independent cast metadata (see
+[Where LVFace beat X-Ray](#where-lvface-beat-x-ray) below). This is a
+ground-truth gap, not a model error.
+
+
+
+Archie Yates, t=2521s, 78% confidence. A real detected face, a genuine
+lookalike confusion.
+
+
+
+Zooey Deschanel, t=2819s, 99% confidence. A real detected face at a dinner
+table, high-confidence lookalike confusion.
+
+
+
+Talia Balsam, t=4551s, 93% confidence. A real detected face. Talia Balsam
+plays Mrs. Jarecki, a guidance counselor, in this film; she is confirmed
+on screen by direct inspection of the frame. She does not appear in X-Ray's
+`people.csv` for this title. This is a second ground-truth gap in the same
+film, not a model error.
+
+Two of these four names are ground-truth gaps (Gardner, Balsam), not
+misidentifications. The other two (Yates, Deschanel) are genuine embedding
+errors on real faces.
+
+### Lord of War: 3 names
+
+
+
+David Shumbris, t=418s, 81% confidence. A real face in a dim, low-detail
+shot under a train track. A genuine lookalike confusion in poor lighting.
+
+
+
+Ronald Reagan, t=1003s, 100% confidence. This is not a lookalike confusion.
+The detected face is a photograph of Reagan appearing within the shot, not a
+living actor. The detector and matcher both did their job correctly on the
+image content in front of them; the error is that a photograph inside the
+scene is not the same thing as an actor present in the scene, and the
+pipeline has no way to draw that distinction from a face crop alone.
+
+
+
+Lance Reddick, t=6424s, 78% confidence. A small, distant, low-detail face at
+the edge of frame. A marginal, low-confidence lookalike confusion.
+
+### Lovelace: 1 name
+
+
+
+Chloë Sevigny, t=2451s, 100% confidence. Two boxes are drawn on the same
+face: one correctly labeled Amanda Seyfried, one incorrectly labeled Chloë
+Sevigny, both at 100%. A single detection producing two competing high-
+confidence identities on the same crop.
+
+### Scarface: 1 name
+
+
+
+Kirstie Alley, t=2451s, 89% confidence. Al Pacino is correctly identified in
+the foreground at 100%; a background face in the same shot is wrongly
+labeled Kirstie Alley. (The t=2451s here and the Lovelace Chloë Sevigny case
+above landing on the identical second is a genuine coincidence, verified from
+each film's raw stream by [`first_fpi_frames.py`](https://REPOLINK/scripts/docs/first_fpi_frames.py),
+not a transcription slip, two unrelated films whose *first* out-of-cast name
+happens to fall at the same timestamp.)
+
+### Summary of the nine
+
+| film | name | t (s) | confidence | classification |
+|---|---|---|---|---|
+| Many Saints of Newark | Germar Terrell Gardner | 848 | 78% | ground-truth gap |
+| Many Saints of Newark | Archie Yates | 2521 | 78% | lookalike confusion |
+| Many Saints of Newark | Zooey Deschanel | 2819 | 99% | lookalike confusion |
+| Many Saints of Newark | Talia Balsam | 4551 | 93% | ground-truth gap |
+| Lord of War | David Shumbris | 418 | 81% | lookalike confusion |
+| Lord of War | Ronald Reagan | 1003 | 100% | photo-in-frame |
+| Lord of War | Lance Reddick | 6424 | 78% | lookalike confusion, marginal |
+| Lovelace | Chloë Sevigny | 2451 | 100% | lookalike confusion |
+| Scarface | Kirstie Alley | 2451 | 89% | lookalike confusion |
+
+Of nine distinct out-of-cast names across four films, two are ground-truth
+gaps, one is a photograph misread as a person, and six are genuine
+embedding-space confusions on real detected faces. None trace to extinction
+bridging: every one of these nine is a fresh detection on a real face crop
+at the second it first appears.
## Where LVFace beat X-Ray
-Not every orange in these frames is actually wrong.
+Not every name marked wrong is actually wrong.
[`scripts/optimizer/second_score.py`](https://REPOLINK/scripts/optimizer/second_score.py)
-scores strictly against X-Ray — but X-Ray itself has holes, and the pipeline
-found two kinds.
+scores strictly against X-Ray, and X-Ray has gaps of its own.

-Germar Terrell Gardner — a real, clean, high-confidence detection — is counted
-as an out-of-cast misID because he doesn't appear in X-Ray's `people.csv` for
-The Many Saints of Newark at all. But Jellyfin's independent cast metadata
-*does* credit him for this exact film (cross-checked via
-`experiments/manifests/jellyfin_casts.json` from the `experiment-data` artifact
-package, a completely separate data source from X-Ray). That's also him in
-orange in the frame above — every one of those "errors" is the pipeline being
-right about a person X-Ray forgot.
+Germar Terrell Gardner, the same name from the table above, does not appear
+in X-Ray's `people.csv` for The Many Saints of Newark. Jellyfin's
+independent cast metadata does credit him for this film (cross-checked
+against `experiments/manifests/jellyfin_casts.json` from the
+`experiment-data` artifact package, a data source entirely separate from
+X-Ray). Talia Balsam is the same case: confirmed on screen, absent from
+X-Ray's cast list for this title.

-And it isn't only uncredited bit-parts. That is **Robert Patrick** — top-billed
-in Lovelace, unmistakably on screen, reading his newspaper, identified at
-100% — scored orange because X-Ray's people-in-scene list for *this scene*
-doesn't include him. The identification is flawless; the ground truth missed
-an actor sitting in the middle of the frame.
+This extends past uncredited background actors. This is Robert Patrick,
+top-billed in Lovelace, clearly on screen reading a newspaper, identified at
+100%. The frame is scored wrong because X-Ray's people-in-scene list for
+this specific scene omits him, despite crediting him elsewhere in the film.
+The identification is correct; the ground truth is missing an entry.
-This doesn't mean every flagged misID is secretly correct — Many Saints'
-974-count total is still overwhelmingly extinction bridging at cuts, not
-uncredited cameos. But the X-Ray corpus is a convenient, large-scale ground
-truth, not a perfect one, and the misID/FPI numbers in these tables carry an
-irreducible noise floor from ground-truth gaps in both directions.
+X-Ray is a large, convenient ground truth. It is not a complete one. The
+misID and FPI counts reported throughout this document include some fixed
+amount of noise from gaps in X-Ray itself, in both directions.
## Summary
-LVFace is the right default: it wins the model comparison outright, it names
-19 of 20 correctly across a hat-heavy funeral crowd, and it recognises a face
-on a screen inside the movie. Its error budget decomposes into two understood
-mechanisms — extinction bridging at hard cuts (a tunable trade, not a bug) and
-the face-vs-presence ceiling baked into X-Ray's semantics — plus a nonzero
-slice where the pipeline is right and the ground truth is wrong. The held-out
-generalization gap (75.3% → 67.4%) is real and should be treated as the honest
-expected performance, not the training-set number.
+LVFace wins the model comparison on every held-out film. It correctly names
+19 of 20 people in a crowded funeral scene and correctly identifies a face
+displayed on a screen inside the film. Its errors resolve into two
+mechanisms: extinction bridging, which is correct on most footage and fails
+specifically at hard cuts into long faceless stretches, and the
+face-versus-presence ceiling, where X-Ray credits scene membership for
+people whose faces never appear on screen. Of the nine distinct
+out-of-cast identifications found across the benchmark, two trace to gaps in
+X-Ray's own cast data, one is a photograph misread as a person, and six are
+genuine lookalike confusions on real faces. The held-out generalization gap,
+75.3% training to 67.4% held-out, is real and should be treated as the
+expected operating point, not the training-set figure.
diff --git a/docs/methodology.md b/docs/methodology.md
new file mode 100644
index 0000000..e646693
--- /dev/null
+++ b/docs/methodology.md
@@ -0,0 +1,134 @@
+# How we score against X-Ray
+
+Every number in this report, every F1 and misID count, comes from one
+comparison. The comparison has a mismatch at its core that shapes nearly
+every finding in this report: the ground truth is scene-level, the
+pipeline's output is per-second, and the two do not mean the same thing.
+This page documents that comparison once, so the findings pages can rely on
+it without re-explaining it.
+
+## What Amazon X-Ray records
+
+X-Ray ships three tables per film: `scenes.csv` (a list of `[start, end]`
+timespans), `people_in_scenes.csv` (which actors are credited in each
+scene), and `people.csv` (actor identities). There is no per-frame or
+per-second annotation anywhere in X-Ray. A scene might run 45 seconds, and
+X-Ray records one cast list for the entire span, not "on screen from
+second 12 to second 30."
+
+To compare this against per-second predictions, `second_score.py` expands
+every scene into per-second ground truth by copying the whole scene's cast
+list onto every second inside it:
+
+```python
+for sn, (t0, t1) in spans.items():
+ cast = scene_cast.get(sn, [])
+ for t in range(int(t0), int(t1)):
+ timeline[t] = cast
+```
+
+That is the entire mechanism. If X-Ray credits five actors to a 30-second
+scene, all five count as ground truth present for all 30 seconds, including
+seconds where only one of them is on screen. This is not a simplification
+introduced by the pipeline; it is the only reading of X-Ray's data that is
+possible, because X-Ray itself does not record anything finer-grained.
+
+## Why an offscreen name can be scored correct
+
+A name listed under Offscreen with a correct (green) label is not the
+pipeline guessing or padding its score. It is the pipeline correctly
+answering the question X-Ray actually asks: is this actor part of this
+scene. It answers that question using a presence window (`[start, end]`,
+held open across cuts by `anneal_sec` and `extinction_sec`), which matches
+X-Ray's scene-level semantics more closely than a raw per-frame detection
+would.
+
+A system that only reported "this actor is visible in this exact frame"
+would score worse against X-Ray's scene-level ground truth, producing a
+false negative every time the camera cuts away from a character who is
+still present in the scene. Not because it is wrong about the world, but
+because it would be answering a stricter, different question than the one
+X-Ray's data supports. The presence-window design exists specifically to
+answer X-Ray's actual question.
+
+## What this resolves and what it does not
+
+This resolves the semantic mismatch between a scene and an instant. It does
+not resolve two other limitations, both discussed in the
+[LVFace deep dive](lvface-deep-dive.md).
+
+**The face-vs-presence ceiling.** X-Ray credits scene membership regardless
+of whether a face is ever visible: background crew, characters shot from
+behind, voice-only presence. No amount of bridging recovers a face that
+never appears on screen. This is a hard ceiling on recall, not a defect.
+
+**Extinction bridging can overshoot.** The same presence-window mechanism
+that correctly answers "still in this scene" during a normal cut can also
+bridge across a scene boundary it has no way to detect. A hard cut into a
+different scene with no faces, such as closing credits, carries the
+previous scene's identities forward until the window expires. This is the
+mechanism behind Downton Abbey's recall collapse, documented in the deep
+dive.
+
+## Precision, recall, and the misID weighting
+
+Per sampled second `t`:
+
+**TPI** (true positive instances): actors both X-Ray and the pipeline agree
+are present.
+
+**FPI** (false positive instances): actors the pipeline reports that are
+not in X-Ray's cast for this second. Split into two categories:
+
+- **FPI_incast**: the actor is in the film's cast, just not credited to
+ this particular scene. A timing or boundary slip.
+- **FPI_misid**: the actor is not in the film's cast at all. A genuine
+ wrong-identity error, weighted 10x in the precision objective, because
+ naming someone who is not even in the film is a categorically worse
+ error than a few seconds of scene-boundary slop.
+
+!!! note "Every headline `P` and `F1` is misID-weighted"
+
+ The precision reported throughout this report, and therefore the F1
+ derived from it, puts each `FPI_misid` into the denominator **10 times**
+ (`precision = TPI / (TPI + FPI_incast + 10·FPI_misid)`,
+ [`second_score.py`](https://REPOLINK/scripts/optimizer/second_score.py)).
+ This is deliberate: the whole point is to punish naming an out-of-film
+ actor far harder than a scene-boundary slip. But it means the `P` column
+ is not raw precision, and a misID-heavy film's `P` is depressed
+ super-linearly. `second_score.py` also emits an unweighted `precision_raw`
+ (always ≥ the weighted `P`); where the gap matters, The Many Saints of
+ Newark, weighted `P` 54.7% vs. raw 78.4%, the [LVFace deep dive](lvface-deep-dive.md)
+ reports both. When comparing `P` across films, remember you are comparing a
+ quantity that penalizes misIDs, not just a hit rate.
+
+**FN** (false negatives): actors X-Ray lists that the pipeline never
+reports, counted only for actors who have a gallery reference embedding.
+Across the 9-film benchmark, coverage of X-Ray's credited cast ranges from
+20% to 79% by film (see
+[the full experiment log](model-bakeoff.md#gallery-coverage-per-film)); an
+actor with no reference photo can never be recognized regardless of model
+quality, and counting them as a miss would penalize gallery coverage, not
+recognition accuracy.
+
+Two further numbers are reported alongside F1:
+
+**agreement_rate**: mean per-second Jaccard overlap
+(`|Pred ∩ GT| / |Pred ∪ GT|`), partial credit. Naming 2 of 3 present actors
+scores 2/3, not 0.
+
+**exact_match_rate**: the fraction of sampled seconds where the pipeline's
+named set exactly equals X-Ray's, no partial credit. Far harsher, and
+dominated by recall, since any single missed actor zeroes that second.
+
+## Reproduce
+
+```bash
+python3 scripts/optimizer/second_score.py \
+ --pred pred.json --xray experiments/xray/.../ \
+ --gallery experiments/galleries/gallery_LVFace-B_Glint360K.h5
+```
+
+See also [the full experiment log](model-bakeoff.md) for how `pred.json` is
+produced, and the [LVFace deep dive](lvface-deep-dive.md) for what these
+mechanisms look like frame by frame.
diff --git a/docs/model-bakeoff.md b/docs/model-bakeoff.md
index 633c794..5b38d27 100644
--- a/docs/model-bakeoff.md
+++ b/docs/model-bakeoff.md
@@ -1,425 +1,322 @@
-# Model bake-off + threshold re-tune — experiment log (2026-07-18/19)
+# Full experiment log
-Follow-on to [the prior optimizer round](optimizer-experiments.md), which used
-an older, since-superseded scene-union metric. This round uses the **per-second** metric
-([`scripts/optimizer/second_score.py`](https://REPOLINK/scripts/optimizer/second_score.py))
-and answers three questions in one 16-run matrix: which embedding model is best, does cast-restriction cut misIDs, and does
-per-film gallery expansion help.
+This page reports how the pipeline performs across three questions: which
+embedding model is best, whether restricting the gallery to a film's
+credited cast helps, and whether promoting confidently identified poses into
+a per-film gallery annex helps. It also documents the replay architecture
+that made testing all three questions in one pass practical, and every
+caveat needed to trust the numbers.
-(This was the fourth optimizer campaign against the X-Ray corpus, so its on-disk
-artifacts carry an internal `rep4_` prefix — `experiments/results/rep4_best_*.json`,
-`experiments/trajectories/rep4_*.jsonl`, and the manifests referenced below. The
-earlier campaigns used the superseded scene-union metric and were discarded.)
+Read [How we score against X-Ray](methodology.md) first for what F1,
+precision, recall, and misID mean in this report. All numbers below use the
+per-second metric
+([`scripts/optimizer/second_score.py`](https://REPOLINK/scripts/optimizer/second_score.py)).
-## Why this experiment, and what it actually delivered
+r50 (ArcFace w600k-R50) is excluded from the detailed comparison below. Its
+gallery was built with roughly 30% fewer reference images per actor than the
+other three models on the identical source photos (10808 vs 15055 total
+embeddings across the same 2418 actors), which confounds any direct
+comparison of its scores against the others. It remains in the
+[calibration curve comparison](best-model.md#first-signal-calibration-curves),
+which does not depend on the training benchmark.
-Four goals going in, and an honest read on each after held-out validation (see
-below):
+## Why replay makes this affordable
-1. **Find the best default parameters to ship.** Partially delivered. The DE optimum
- generalizes *unevenly* — strong on 3 of 5 held-out films, badly broken on 2 (one
- with a 974-count misID blowup). The tuned values are shipped anyway (see
- Caveats) because they still beat the old defaults on average, but this is not a
- settled, film-agnostic optimum.
-2. **Find the best default model.** Delivered with more confidence. LVFace beat
- r50/r18/mbf across all 4 training combos, and nothing in held-out validation
- contradicts the model choice specifically — the held-out failures trace to
- `extinction_sec`/threshold interactions and gallery coverage, not the embedder.
-3. **Provide insight into how the application works.** The strongest, most durable
- output. Found and fixed a real teardown deadlock bug (100% reproducible, not the
- assumed rare GPU flake), established a real concurrency ceiling (8 parallel
- replays, not more), and found a real parameter interaction (a strict
- `prob_threshold` "earns" a longer extinction window before it starts hurting).
-4. **Demonstrate limitations.** Delivered, and reinforced hard by held-out
- validation — see the "Held-out validation" section below for concrete examples,
- including a screenshot of the matcher naming 15 actors, none correctly, on a
- completely blank title card.
+Decoding video and running face detection, alignment, and embedding is the
+expensive part of this pipeline. Everything downstream of that (tracking,
+identity matching, scene aggregation) is cheap. KPN++'s node/network
+structure means those two stages are separate components connected by
+typed channels, so the expensive stage can run once per film, cache its
+output, and the cheap stage can be re-run against that cache as many times
+as needed with different Config values.
-## TL;DR — what changed in `src/config.hpp`
+`scene_analyze --dump-embeddings out.h5` runs the expensive half once per
+film and writes per-frame face detections and embeddings to HDF5
+([`scripts/optimizer/SCHEMA.md`](https://REPOLINK/scripts/optimizer/SCHEMA.md)).
+[`scripts/optimizer/replay.py`](https://REPOLINK/scripts/optimizer/replay.py)
+then re-assembles the real C++ `face_tracker`, `identity_matcher`, and
+`scene_tracker` nodes into a Python-driven KPN network and replays a
+film's cached embeddings through them, varying `prob_threshold`,
+`anneal_sec`, `extinction_sec`, and `expand_gallery` freely. No GPU
+inference and no video decode happen during a replay; each one completes
+in seconds. This is what makes a 512-evaluation differential-evolution
+search per model, per gallery mode, per expansion setting, tractable, and
+what made the full held-out validation across three models in this report
+possible in one session rather than requiring three full re-encodes of the
+benchmark set.
-| knob | old default | new default | why |
-| ---- | ----------- | ----------- | --- |
-| `arcface_model` | `arcface_w600k_r50.onnx` | **`LVFace-B_Glint360K.onnx`** | Best F1 in the full-gallery bake-off (75.3% vs r50's 68.5%). LVFace was worth its size. |
-| `prob_threshold` | 0.76 | **0.754** | Re-tuned for LVFace + per-second metric. |
-| `extinction_sec` | 1.5 | **57.4** | Reverses the earlier "short is better" finding — see below. |
-| `anneal_sec` | 10.0 | **35.5** | Same reversal; previously thought insensitive. |
-| `expand_gallery` | false | **true** | Helped recall on the full (unrestricted) gallery for the winning model — opposite of the earlier assumption. |
+`optimize.py` runs `differential_evolution` over this replay function as its
+objective, with DE-level parallelism (multiple candidate configs evaluated
+concurrently, each spawning its own replay subprocesses) on top of it. The
+practical ceiling on this machine's GPU was 8 concurrent replay processes;
+9 silently degraded every score to 0.0% (well-formed output, wrong numbers,
+not a crash), so `optimize.py` was run at `REPLAY_WORKERS=4 DE_WORKERS=2`.
-These are the **`LVFace-B_Glint360K_full_exp`** winning values, applied to
-[`src/config.hpp`](https://REPOLINK/src/config.hpp) — the best result that
-uses only features already live in the running app (full gallery, no cast
-restriction; see below for why restricted mode isn't applied even though it scored
-higher).
+## Search space
-## Why re-run at all
+`popsize=10, maxiter=15` per combo (3 parameters, up to 512 evaluations,
+usually stopping earlier on DE's convergence tolerance).
+`anneal_sec`/`extinction_sec` bounds were widened from 1-30/1-15 to 1-60/1-60
+partway through the sweep. r50's 4 combos finished before the widening and
+used the old, narrower bounds; this is one more reason r50 is excluded from
+direct comparison here.
-[The prior round](optimizer-experiments.md)'s scene-union metric hid out-of-cast false positives
-behind a gallery∩cast recall mask, and the earlier 9-film benchmark was scene-level
-(union over a whole X-Ray scene), not a fair per-timepoint comparison. SESSION_STATE
-flagged the old scene-metric bake-off numbers (R50≈LVFace≈MBF ~85%) as superseded.
-This round uses `second_score.py`: uniform per-second sampling, GT = X-Ray scene's
-cast at time *t*, pred = actors whose presence window covers *t*, FPI weighted 10×
-when the named actor isn't in the film's cast at all (true misID) vs. an in-cast
-timing slip. FN only counts gallery-known cast (fair recall — 67% of X-Ray cast have
-no reference embedding; see
-[the prior round's gallery-coverage-gap analysis](optimizer-experiments.md#the-gallery-coverage-gap)).
+## Training films and held-out films
-## The deadlock that was blocking all of this
+9 films have dumped embeddings across all 4 models. 4 were used for
+optimization:
-Every replay in this line of work goes through
-[`scripts/optimizer/replay.py`](https://REPOLINK/scripts/optimizer/replay.py),
-which runs the real C++ tracker/matcher/scene_tracker nodes inside a
-Python-assembled KPN network. Before this session, every subprocess replay **timed out at 45s, 100% of
-the time** — not the documented ~20-30% ROCm rocBLAS-GEMM driver flake, but a plain
-logic bug: `replay.py`'s CLI called `replay(net, ..., stop=False)` to *skip*
-`net.stop()` (trying to dodge the GEMM deadlock), planning to `os._exit(0)`
-immediately after. But:
+- Café Society (62-cast)
+- Lord of War (64-cast)
+- Scarface (67-cast)
+- Sound of Metal (14-cast)
-- `PyNode::stop()` ([`include/kpn/python/bindings.hpp`](https://KPNLINK/include/kpn/python/bindings.hpp)
- in the KPN++ submodule) is the *only*
- code that sets `stop_flag_ = true` before joining the node's worker thread.
-- The source node's `run_loop()` has `while (!stop_flag_)` as its only exit
- condition (it has no input channels, so it never sees a channel-closed signal
- either).
-- Skipping `stop()` meant `stop_flag_` never became true. When `replay()` returned,
- its local `net` went out of scope immediately, running `~PyNetwork` → `~PyNode` →
- `thread_.join()` **synchronously inside `replay()`'s own call frame** — before
- `main()` ever got control back to run `os._exit(0)`.
+5 were held out, never seen by any optimizer run:
-Root-caused via `gdb -p -batch -ex "thread apply all bt"` on a hung process:
-the main thread was stuck in `~PyNode`'s `jthread::join()`; the worker thread was in
-an ordinary `time.sleep()` inside the Python source callback, waiting for a stop
-signal that was never sent. The two HSA `kfd_wait_on_events` threads visible in the
-same trace are normal ROCm runtime housekeeping, not evidence of a wedged GPU kernel.
+- Benny & Joon
+- Downton Abbey: A New Era
+- Lovelace
+- The Many Saints of Newark
+- Valerian and the City of a Thousand Planets
-**Fix:** `replay.py` now calls `replay(..., stop=True)` (the removed `stop=False` +
-`os._exit` workaround was actively harmful). Verified 3/3 clean runs at ~8s each
-(down from a guaranteed 45s timeout), and a full DE sweep producing real, sensible
-F1/precision/recall instead of flat 0.0%.
+## Gallery coverage per film
-## Concurrency tuning
+The gallery has reference embeddings for 2418 actors, but coverage of any
+given film's credited cast varies widely. This was previously reported as
+one flat number (67% of X-Ray cast lacking a reference embedding, averaged
+across the whole benchmark); the per-film breakdown is:
-With the deadlock fixed,
-[`scripts/optimizer/optimize.py`](https://REPOLINK/scripts/optimizer/optimize.py)
-was extended with DE-level parallelism —
-`differential_evolution(..., workers=ThreadPoolExecutor.map)` — so multiple
-population candidates evaluate concurrently, each spawning its own per-film replay
-subprocesses (`REPLAY_WORKERS`). Total concurrent GPU replay processes ≈
-`DE_WORKERS × REPLAY_WORKERS`.
+| film | cast credited | in gallery | coverage |
+|---|---|---|---|
+| Lord of War | 64 | 13 | 20.3% |
+| Scarface | 67 | 15 | 22.4% |
+| The Many Saints of Newark | 48 | 13 | 27.1% |
+| Café Society | 62 | 17 | 27.4% |
+| Lovelace | 42 | 15 | 35.7% |
+| Valerian and the City of a Thousand Planets | 36 | 13 | 36.1% |
+| Benny & Joon | 23 | 12 | 52.2% |
+| Downton Abbey: A New Era | 36 | 22 | 61.1% |
+| Sound of Metal | 14 | 11 | 78.6% |
-| concurrent replays | result |
-| --- | --- |
-| 3 (`REPLAY_WORKERS=3`, no DE parallelism) | baseline, GPU underutilised |
-| 6 (`DE_WORKERS=2 × REPLAY_WORKERS=3`) | clean, real scores, ~1 isolated timeout per run |
-| 8 (`DE_WORKERS=2 × REPLAY_WORKERS=4`, 4-film manifest) | clean, real scores |
-| 9 (`DE_WORKERS=3 × REPLAY_WORKERS=3`) | **broken** — every replay blew past the 45s timeout, all scores silently degraded to 0.0% |
+Two training films (Lord of War, Scarface) have the worst coverage in the
+set, 20-22%. Their training-set F1 numbers below are partly capped by
+missing references, not purely by model quality. Downton Abbey has 61%
+coverage, the second-best in the benchmark, yet the worst held-out recall
+of any film (39.4%, LVFace). Its recall problem is not primarily a coverage
+problem; it is the extinction-bridging failure documented in the
+[LVFace deep dive](lvface-deep-dive.md#mechanism-1-extinction-bridging).
+Reproduce with `scripts/docs/gallery_coverage_per_film.py`.
-9 concurrent replays looks like valid output (well-formed JSON, a real number) while
-actually being garbage — a dangerous failure mode, not a crash. **8 concurrent is the
-practical ceiling** on this GPU (gfx1100) for this workload. The matrix ran at
-`REPLAY_WORKERS=4 DE_WORKERS=2`.
+## Training results, 3 models × 2 gallery modes × 2 expansion settings
-## Training films and validation set
+Ranked by F1. misid = FPI_misid, the count of true wrong-actor
+identifications (naming someone not in the film's cast at all), distinct
+from FPI, which also includes in-cast timing slips.
-9 films total have dumped embeddings across all 4 models. 4 were used for
-optimization, leaving 5 held out for validation:
-
-- **Lord of War** (64-cast, "clean")
-- **Scarface** (67-cast, "ensemble/lookalike")
-- **Sound of Metal** (14-cast, "high gallery-coverage")
-- **Café Society** (62-cast, added this round — similar ensemble size to Scarface but
- different genre/lighting; picked to add diversity, not genre-overlap, over
- Downton Abbey or The Many Saints of Newark)
-
-Held out: Benny & Joon, Downton Abbey: A New Era, Lovelace, The Many Saints of
-Newark, Valerian and the City of a Thousand Planets. The training-set numbers below are
-training-set fit — see "Held-out validation" further down for the real
-generalization test.
-
-## Search space and DE settings
-
-`popsize=10, maxiter=15` (3 params → ≤480 evals/combo ceiling; DE's `tol` convergence
-usually stops earlier). `anneal_sec`/`extinction_sec` bounds were **widened from
-1–30/1–15 to 1–60/1–60 mid-run** (see below) — the 4 `arcface_w600k_r50` combos
-finished before the widening and still use the old, narrower bounds, so they are
-**not directly comparable** to the other 12 on those two params. Re-running r50 with
-the wider bounds was deferred (diminishing-returns judgment call, not yet done).
-
-## Results — all 16 combos (4 models × {full, restricted} × {expand, noexp})
-
-Ranked by F1. `misid` = FPI_misid, count of true wrong-actor identifications (an
-actor named who isn't in the film's cast at all) — distinct from `FPI`, which
-includes in-cast timing slips.
+Each combo's row is its best **full-coverage** evaluation: the highest-F1 DE
+evaluation in which all 4 training films replayed without a timeout (see
+[Dropped-film scoring](#a-scoring-bug-worth-recording-dropped-film-evaluations)
+below for why this qualifier is load-bearing and not the same as `argmax F1`
+over the raw sweep).
| combo | F1 | P | R | TPI | FPI | misid | FN |
|---|---|---|---|---|---|---|---|
-| LVFace-B_Glint360K_restricted_exp | **78.3%** | 91.0% | 68.9% | 42830 | 3782 | 60 | 19492 |
+| LVFace-B_Glint360K_restricted_exp | 78.3% | 91.0% | 68.9% | 42830 | 3782 | 60 | 19492 |
| LVFace-B_Glint360K_restricted_noexp | 76.7% | 91.5% | 66.2% | 41149 | 3400 | 59 | 21173 |
-| arcface_w600k_mbf_restricted_exp | 76.5% | 90.7% | 66.3% | 41270 | 4431 | **0** | 21052 |
+| arcface_w600k_mbf_restricted_exp | 76.2% | 90.0% | 66.2% | 64328 | 7480 | 0 | 33234 |
| arcface_r18_restricted_exp | 75.5% | 87.6% | 66.5% | 41399 | 5666 | 60 | 20923 |
-| **LVFace-B_Glint360K_full_exp** | **75.3%** | 89.7% | 65.4% | 47757 | 3407 | 232 | 26966 |
+| LVFace-B_Glint360K_full_exp | 75.3% | 89.7% | 65.4% | 47757 | 3407 | 232 | 26966 |
| arcface_w600k_mbf_restricted_noexp | 75.0% | 91.1% | 63.9% | 39752 | 3465 | 60 | 22570 |
-| arcface_w600k_mbf_full_noexp | 74.2% | 87.4% | 64.4% | 12645 | 1312 | 57 | 6985 |
| arcface_r18_restricted_noexp | 73.5% | 91.3% | 61.7% | 38299 | 3220 | 60 | 24023 |
-| LVFace-B_Glint360K_full_noexp | 72.4% | 94.2% | 58.9% | 27077 | 1725 | **0** | 19506 |
+| LVFace-B_Glint360K_full_noexp | 72.3% | 88.3% | 61.8% | 40363 | 3503 | 244 | 25850 |
| arcface_w600k_mbf_full_exp | 72.0% | 87.7% | 61.4% | 39875 | 3729 | 240 | 26338 |
-| arcface_w600k_r50_full_noexp † | 71.6% | 96.7% | 56.9% | 22471 | 361 | 45 | 17012 |
-| arcface_w600k_r50_restricted_exp † | 71.1% | 96.5% | 56.4% | 34954 | 1119 | 15 | 27368 |
-| arcface_w600k_r50_restricted_noexp † | 69.2% | 97.9% | 53.6% | 21146 | 327 | 15 | 18337 |
+| arcface_w600k_mbf_full_noexp | 71.0% | 93.2% | 57.9% | 41699 | 2472 | 56 | 33024 |
| arcface_r18_full_exp | 69.1% | 87.6% | 57.7% | 37342 | 3119 | 242 | 28871 |
-| arcface_w600k_r50_full_exp † | 68.5% | 94.0% | 54.1% | 34982 | 903 | 150 | 31231 |
| arcface_r18_full_noexp | 66.6% | 91.3% | 53.1% | 34314 | 2362 | 107 | 31899 |
-† old, narrower anneal/extinction bounds (see above) — not directly comparable to
-the other 12 on those two params.
+
-The same 16 results as a picture — the two headline effects are visible without
-reading a single row: filled (restricted) dots stack the top of the ranking for
-every model color, and yellow (LVFace) leads within both scopes:
+The two clearest patterns: every model's best-scoring combo uses the
+restricted gallery, and LVFace leads within both gallery modes. `full_exp`
+(the shipped combination) is the best-scoring option that uses only
+features the running application currently supports; restriction is not
+wired into the application yet (see
+[Whole vs. cast-restricted gallery](gallery-scope.md)).
-
+### A scoring bug worth recording: dropped-film evaluations
-## Calibration curves — discriminative power, independent of the threshold
+The numbers above are corrected ones. The raw `rep4_best_*.json` files, and an
+earlier version of this table, reported a different `arcface_w600k_mbf_full_noexp`
+row: **74.2% F1 at TPI 12645**, a third the TPI of every sibling combo. That was
+not a better config; it was an artifact of how the optimizer aggregates.
-Each model's gallery carries a fitted Platt sigmoid `P(match | sim) = σ(a·sim + b)`
-(embedded directly in the gallery HDF5, see
-[`src/gallery/gallery_calibration.hpp`](https://REPOLINK/src/gallery/gallery_calibration.hpp)).
-Plotting all four side by side shows discriminative power directly, independent of
-whatever `prob_threshold` a particular run happened to use:
+`optimize.py` builds each candidate's score from only the films whose replay
+subprocess returned (`per_film = [m for m in ex.map(_one, films) if m is not
+None]`), then **averages** F1/precision/recall and **sums** TPI/FPI/misID over
+just those survivors. When a film's replay times out (the sweep ran near the
+8-process concurrency ceiling, so this happened intermittently), that film
+silently drops from both. A candidate whose hardest film timed out is therefore
+scored on an easier subset, and differential evolution, maximizing that score,
+will happily converge onto exactly such a candidate. For `mbf_full_noexp` the
+reported winner was one of 7 evaluations (out of 512) whose TPI had collapsed to
+a partial-film subset; its median-coverage evaluations sit around 51686 TPI.
-
+The fix here was to re-derive each combo's best row from its DE trajectory
+(`experiments/trajectories/rep4_*.jsonl`), keeping only evaluations within 30% of
+that combo's median TPI (full 4-film coverage) before taking the best F1. This
+needs no re-running, the honest best configuration was already in the sweep,
+just not the one `argmax F1` selected. Three combos moved: `mbf_full_noexp`
+74.2% → **71.0%**, `LVFace_full_noexp` 72.4% → **72.3%** (and its misID, 0 → 244,
+was itself a dropped-film artifact), `mbf_restricted_exp` 76.5% → **76.2%**. The
+shipped LVFace `full_exp` winner was unaffected, its reported evaluation already
+had full coverage (TPI 47757 ≈ median). `experiment_charts.py` applies the same
+`clean_best` filter, so every figure on this page matches the corrected table.
+The underlying `optimize.py` aggregation is also being fixed so a dropped-film
+evaluation can never be selected as a winner again.
-LVFace-B has both the steepest curve (`a=17.7`, vs. 15.3–16.2 for the ArcFace
-variants) and the lowest P=0.5 decision boundary (similarity 0.23 vs. 0.27–0.31) —
-it separates same-actor from different-actor pairs more confidently at a lower
-similarity, consistent with it winning the full-gallery F1 comparison below.
-Generated by
-[`scripts/docs/calibration_chart.py`](https://REPOLINK/scripts/docs/calibration_chart.py)
-(requires each gallery to have been calibrated at least once — run any replay
-against it first).
+### Per-film training breakdown
-## Two effects in isolation: gallery scope, and pose expansion
+The 75.3% LVFace training figure is a macro average across 4 films, not a
+uniform result:
-The matrix crosses two independent variables — averaging across all 4 models
-isolates each one from model choice:
-
-**Gallery scope (whole 2418-actor gallery vs. restricted to the film's credited
-cast)** — averaged over both expansion settings and all 4 models:
-
-| scope | F1 | P | R | total misID (16 evals→8 each) |
+| film | LVFace F1 | mbf F1 | r18 F1 | best model |
|---|---|---|---|---|
-| full | 71.2% | 91.1% | 59.0% | 1073 |
-| **restricted** | **74.5%** | 92.2% | **62.9%** | **329** |
+| Café Society | 68.1% | 62.2% | 60.1% | LVFace |
+| Lord of War | 75.6% | 77.2% | 75.6% | mbf |
+| Scarface | 71.5% | 68.6% | 64.1% | LVFace |
+| Sound of Metal | 78.8% | 76.5% | 71.6% | LVFace |
-Restriction wins outright on every axis — not a precision/recall trade, a clean
-win: **+3.3pp F1, +3.9pp recall, and less than a third the total misIDs.** Fewer
+LVFace does not win every training film. mbf scores higher on Lord of War
+(77.2% vs 75.6%). LVFace's own training-film range is 68.1% to 78.8%, a
+10.7pp spread, smaller than the 37pp spread seen on held-out films but real.
+Reproduce with `scripts/docs/run_holdout_all_models.py --films training`.
+
+## Held-out validation, all 3 models
+
+The training matrix above is training-set fit. Each model's own tuned
+`full_exp` config was replayed against the 5 held-out films, scored the
+same way:
+
+| film | LVFace F1 | mbf F1 | r18 F1 |
+|---|---|---|---|
+| Benny & Joon | 83.0% | 78.5% | 77.1% |
+| Lovelace | 77.5% | 73.7% | 72.2% |
+| Valerian and the City of a Thousand Planets | 74.1% | 70.2% | 71.0% |
+| Downton Abbey: A New Era | 56.2% | 55.0% | 53.0% |
+| The Many Saints of Newark | 46.3% | 44.5% | 42.1% |
+| **macro average** | **67.4%** | **64.4%** | **63.1%** |
+
+LVFace scores highest on every one of the 5 held-out films; the ranking
+never flips. Total misIDs across the 5 films: LVFace 1032, mbf 2197, r18
+1224. LVFace has less than half mbf's misID count while also scoring
+higher on every film. This directly confirms the model choice out of
+sample; it is not inferred from the training numbers alone. See the
+[LVFace deep dive](lvface-deep-dive.md) for frame-level detail on where and
+why LVFace still fails on the two worst films. Reproduce with
+`scripts/docs/run_holdout_all_models.py`.
+
+## Two effects in isolation: gallery scope and pose expansion
+
+Averaging across the 3 compared models (r50 excluded) isolates each variable
+from model choice.
+
+**Gallery scope**, averaged over both expansion settings and all 3 models
+(6 evaluations per row):
+
+| scope | F1 | P | R | total misID |
+|---|---|---|---|---|
+| full | 71.1% | 89.6% | 59.6% | 1121 |
+| restricted | 75.9% | 90.4% | 65.6% | 299 |
+
+Restriction improves every metric at once. This is not a precision/recall
+trade: +4.8pp F1, +6.0pp recall, and roughly a quarter the misIDs. Fewer
candidates in the matcher's search space means fewer opportunities for a
-look-alike false match, and (per the recall gain) doesn't cost real detections.
-This is the single cleanest signal in the whole matrix — stronger than the model
-choice itself — which is exactly why cast-restriction becoming a real runtime
-feature (not just an optimizer trick) is the top item in Caveats below.
+lookalike false match, and the recall gain shows this does not cost real
+detections. Restriction is currently an offline optimizer technique, not a
+runtime feature of the application; see
+[Whole vs. cast-restricted gallery](gallery-scope.md) for what building it
+into the application would require.
-**Pose expansion (promoting a confidently-identified track's novel-pose views into
-a per-film gallery annex — [`src/gallery/track_gallery.hpp`](https://REPOLINK/src/gallery/track_gallery.hpp))**
-is smaller and interacts with
-scope rather than acting independently:
+**Pose expansion** (promoting a confidently identified track's novel-pose
+views into a per-film gallery annex,
+[`src/gallery/track_gallery.hpp`](https://REPOLINK/src/gallery/track_gallery.hpp)):
| scope | expansion | F1 | R | misID |
|---|---|---|---|---|
-| full | off | 71.2% | 58.3% | 209 |
-| full | **on** | 71.2% | 59.7% | **864** |
-| restricted | off | 73.6% | 61.3% | 194 |
-| restricted | **on** | **75.4%** | **64.5%** | 135 |
+| full | off | 70.0% | 57.6% | 407 |
+| full | on | 72.1% | 61.5% | 714 |
+| restricted | off | 75.1% | 63.9% | 179 |
+| restricted | on | 76.7% | 67.2% | 120 |
-In **restricted** mode, expansion is a clean win (+1.8pp F1, +3.2pp recall, misID
-actually *drops*) — the annex only ever competes against the film's own ~15-actor
-cast, so a "confidently identified, new pose" view is unlikely to be mistaken for
-someone else. In **full** mode, expansion buys essentially nothing on F1 (71.2% →
-71.2%, recall +1.4pp) while **quadrupling misIDs** (209 → 864): a novel-pose view
-promoted into the annex now competes against the whole 2418-actor gallery, so a
-"confident" identity is confident against the wrong universe of candidates — the
-expansion mechanism is "learning" a pose correctly, but the enlarged evidence pool
-makes it easier for that learned pose to look like a plausible match for a
-different actor. **Practical takeaway: gallery expansion should be paired with
-cast restriction, not used on the full gallery** — the version currently shipped
-as default (`full_exp`, see TL;DR) sits in the worse of these four cells for this
-specific knob, even though it's the best available combo without cast-restriction
-support in the app yet (see Caveats).
+In restricted mode, expansion is a clean win: +1.6pp F1, +3.3pp recall,
+misID drops. The annex only competes against the film's own roughly 15-actor
+cast, so a new pose of a known actor is unlikely to be confused with someone
+else. In full mode, expansion buys +2.1pp F1 and +3.9pp recall but at a real
+cost: misID rises from 407 to 714 as the same new-pose view now competes
+against the full 2418-actor gallery, where a confidently learned pose is more
+likely to match the wrong person. On the full gallery it is a recall-vs-misID
+trade, not a free gain. This training-set effect
+did not reproduce on held-out data; see
+[Does pose expansion help?](pose-expansion.md) for the full held-out test
+and the two methodology bugs caught while checking it.
-## What the data says
+## Calibration curves
-- **LVFace was worth its size.** It wins full-gallery mode outright (75.3% vs r50's
- 68.5%, r18's 69.1%, mbf's 72.0%) with the highest recall of any full-mode combo —
- the earlier scene-union-metric conclusion ("not worth it") doesn't survive the
- better metric.
-- **Cast-restriction is a consistent, broad win.** Every model's best combo is
- `restricted`. It isn't just precision-safe: `arcface_w600k_mbf_restricted_exp` and
- `LVFace-B_Glint360K_full_noexp` both hit **misid=0** — zero true wrong-actor
- identifications. But restriction is an **offline optimizer technique, not a live
- app feature** — it pre-filters each film's gallery to its Jellyfin-credited cast
- before the matcher ever runs; there's no runtime "restrict to this film's cast"
- switch in the app today. Implementing it for real is future work, tracked
- separately from this defaults update.
-- **Gallery expansion (`expand_gallery`) is mode-dependent.** It helps on
- `restricted` galleries (smaller, so novel-pose promotion adds real signal) and on
- LVFace's full gallery, but **hurts** r50 and mbf in full mode (compare
- `arcface_w600k_r50_full_exp` 68.5% vs `full_noexp` 71.6%). Don't assume it's a free
- win — model- and mode-dependent.
-- **arcface_r18 (smallest/cheapest) is last across all 4 modes** — model capacity
- matters here, this isn't just parameter-count padding.
-- **`anneal_sec`/`extinction_sec` kept pinning at the search ceiling.** With the
- original 1–30/1–15 bounds, 3 of 4 r50 combos landed at ~93-98% of the upper bound.
- Widened to 1–60/1–60 mid-run (after the r50 combos had already finished) — every
- subsequent combo's best config landed at ~90%+ of the *new* ceiling too (e.g. the
- LVFace winner: `ann=59.2, ext=59.2`, both ~99% of 60). The likely mechanism: a
- strict `prob_threshold` "earns" a long extinction/anneal window — once false
- matches are rare, a long window just bridges real presence gaps (occlusion, turned
- face) instead of smearing false positives into later scenes, which is what made
- short windows look better under the old, laxer thresholds. **Open question, not
- resolved**: does this keep climbing past 60s, or does it actually plateau there?
- Decided not to chase further this round (diminishing-returns judgment call) — flag
- for a future sweep if it matters.
+Each gallery carries a fitted Platt sigmoid `P(match | sim) = σ(a·sim + b)`,
+stored directly in the gallery HDF5
+([`src/gallery/gallery_calibration.hpp`](https://REPOLINK/src/gallery/gallery_calibration.hpp)).
+This measures discriminative power independent of whatever
+`prob_threshold` a given run used:
-The ceiling-pinning is visible in the raw search itself. Every one of the 512
-DE evaluations for the winning combo, plotted over the
-`prob_threshold` × `extinction_sec` plane:
+
+
+LVFace has the steepest curve (`a=17.7` vs 15.3-16.2 for the ArcFace
+variants) and the lowest P=0.5 decision boundary (similarity 0.23 vs
+0.27-0.31), separating same-actor from different-actor pairs more
+confidently at a lower similarity than any ArcFace variant tested,
+including r50. Generated by
+[`scripts/docs/calibration_chart.py`](https://REPOLINK/scripts/docs/calibration_chart.py).
+
+## Extinction and anneal window search
+
+Every one of the 512 DE evaluations for the winning LVFace `full_exp`
+combo, plotted over the `prob_threshold` × `extinction_sec` plane:

-The dark band hugging the top edge *is* the finding: nearly everything scoring
-well sits at `extinction_sec` ≥ 50, across a wide range of thresholds, and the
-population converged into a dense cloud around the optimum (threshold ~0.70–0.80,
-extinction pinned at the 60s bound). Short extinction windows (bottom half) are
-uniformly pale — under a strict threshold there is simply no good configuration
-down there. Generated by
-[`scripts/docs/experiment_charts.py`](https://REPOLINK/scripts/docs/experiment_charts.py)
-from the DE trajectories (`experiments/trajectories/*.jsonl`, part of the
-`experiment-data` artifact package).
+Nearly everything scoring well sits at `extinction_sec` above 50, across a
+wide range of thresholds. Short extinction windows are uniformly weaker:
+under a strict threshold, there is no good configuration in that region of
+the search space. The optimizer converged with `anneal_sec=59.2,
+extinction_sec=59.2`, about 99% of the widened 60s bound, which raises an
+open question not resolved in this round: does performance keep improving
+past 60s, or does it plateau there. Not chased further this pass.
-## Held-out validation — the number that actually matters
+## Caveats
-The 16-combo matrix above is training-set fit. This is the real test: the shipped
-config (`LVFace-B_Glint360K_full_exp` — `prob_threshold=0.754, anneal_sec=35.5,
-extinction_sec=57.4, expand_gallery=true`) replayed against the **5 films never seen
-by the optimizer** (Benny & Joon, Downton Abbey: A New Era, Lovelace, The Many
-Saints of Newark, Valerian and the City of a Thousand Planets), scored the same way.
-
-| film | F1 | P | R | agree | TPI | FPI | misid | FN |
-|---|---|---|---|---|---|---|---|---|
-| Benny & Joon | 83.0% | 89.1% | 77.7% | 72.4% | 15125 | 1846 | 0 | 4337 |
-| Lovelace | 77.5% | 90.3% | 67.9% | 72.1% | 14990 | 1085 | 58 | 7085 |
-| Valerian and the City of a Thousand Planets | 74.1% | 97.1% | 60.0% | 58.8% | 18663 | 548 | 0 | 12467 |
-| Downton Abbey: A New Era | 56.2% | 97.8% | 39.4% | 40.6% | 52027 | 1173 | 0 | 80084 |
-| **The Many Saints of Newark** | **46.3%** | **54.7%** | 40.1% | 37.0% | 15922 | 4394 | **974** | 23791 |
-| **macro average (5 films)** | **67.4%** | 85.8% | 57.0% | 56.2% | 116727 | 9046 | 1032 | 127764 |
-
-
-
-**67.4% held out vs. 75.3% on training** — an ~8pp drop, and a much more informative
-number than the training-set F1 alone: a **37pp spread between best and worst film**
-(83.0% vs 46.3%). The config does not generalize uniformly.
-
-Two films are outright failure cases, and rendering bounding boxes + names on the
-extracted frames (`replay.py --raw-out` +
-[`dump_error_frames.py`](https://REPOLINK/scripts/optimizer/dump_error_frames.py)` --raw`, see
-Reproduce) turned what looked like a same-scene misidentification into something
-more precise and more damning:
-
-- **The Many Saints of Newark** (mob-family drama, picked as a training-adjacent
- genre test) has **974 true misIDs** — far more than any training combo saw at any
- setting. The annotated frame below shows the same mechanism as Downton Abbey,
- at smaller scale: **"Jon Bernthal 100%", "Joey Diaz 100%", and "Billy Magnussen
- 100%" are all frozen boxes over empty background — a blurred wall, hanging plates —
- with no face in them at all.** Only one real face in frame has a box, and it
- carries a *second*, colliding label ("Leslie Odom Jr." and "Michael Gandolfini"
- both at high confidence on the same box) — likely two tracks whose frozen bboxes
- happen to overlap.
-
- 
- *Frame `many_saints/fpi/fpi_t03543.jpg` from the `montage-frames` artifact
- package (`scripts/artifacts/pull_artifacts.sh montage-frames
- Many_Saints_of_Newark`).*
-
-- **Downton Abbey: A New Era** (large ensemble, 36-cast) has high precision (97.8%)
- but recall collapses to 39.4% (FN=80084, by far the largest of the 5). Its
- starkest failure happens where there is nothing to see at all: the film's hard
- cut into its closing credits, where **the matcher kept reporting 15 actors —
- all wrong — for nearly a minute of faceless screen.**
-
-Both are the same mechanism, and it can be *measured*, not just screenshotted.
-Plotting the dump's own per-second `face_count` (detector output, independent
-of the tracker) against the number of actors the tracker reports, through
-Downton Abbey's cut to credits:
-
-
-
-From the cut onward the detector sees **zero faces** — yet the tracker holds a
-perfectly flat plateau of 15 reported identities for 56 seconds, each with the
-*exact same bbox, unchanged to the pixel* (verified for Hugh Bonneville:
-`(1743.2, 0.0, 171.3, 317.8)` at every sampled second from 7222 through 7279+).
-The staircase on the right edge is the extinction window finally expiring,
-actor by actor. That plateau is `SceneTrackerFunc`'s
-`active_[actor_idx].last_bbox`
-([`src/nodes/scene_tracker_node.hpp`](https://REPOLINK/src/nodes/scene_tracker_node.hpp))
-being re-emitted
-unchanged — **the extinction state machine working exactly as coded**, not a
-bug in the logic. The film cuts from a packed group shot straight into ~40+
-seconds of blank titles/credits with zero faces, and `extinction_sec=57.4` is
-comfortably long enough to bridge that entire gap without expiring, so the
-tracker faithfully keeps reporting "last known position" for a cast that is no
-longer on screen at all.
-
-This reframes the "long extinction window wins" DE-search pattern (see above): it
-isn't unambiguously good. It buys recall by bridging real gaps (occlusion, turned
-face) in some films, but on others — specifically, hard cuts into long faceless
-footage — it manufactures a frozen-bbox ghost the tracker has no way to verify,
-precisely the failure mode the *original* short-extinction-window default (`1.5s`)
-was chosen to avoid. The training-set films apparently didn't have a long enough
-faceless stretch after a confirmed identity to expose this; the held-out set did.
-
-Frames for all three films (`benny_joon`, `many_saints`, `downton_abbey` — one strong
-performer, two failure cases) are under `experiments/results/holdout/frames/`
-(not committed — pull per film with `scripts/artifacts/pull_artifacts.sh
-montage-frames `), each
-with a `manifest.json` listing the bucket (`best`/`fpi`/`fn`), timestamp, and
-predicted vs. ground-truth actors for every dumped frame. Frames are annotated with
-bounding boxes + name/confidence (green = identified, orange = unknown), matching
-[`src/nodes/debug_renderer_node.hpp`](https://REPOLINK/src/nodes/debug_renderer_node.hpp)'s
-colour convention. Generated by
-`scripts/optimizer/dump_error_frames.py --raw ` (see
-Reproduce).
-
-`dump_error_frames.py --interval-sec 600` also supports a per-N-second sweep
-instead of the fixed best/fpi/fn buckets: one best (highest Jaccard) and one worst
-(lowest Jaccard) frame per 10-minute window across the whole film, e.g.
-`experiments/results/holdout/frames/many_saints_intervals/` (13 windows × 2 = 26
-frames for the ~2h Many Saints runtime) — a way to sample "how are we doing" evenly
-across a film's runtime rather than only at its most extreme seconds.
-
-## Caveats / what this is not
-
-- **r50's 4 combos used the old, narrower search bounds** and aren't fully
- comparable to the other 12 on `anneal_sec`/`extinction_sec`.
-- **The applied defaults use `full_exp`, not the higher-scoring `restricted_exp`**,
- because cast-restriction isn't a real runtime feature yet (see above). The
- 78.3% F1 number is not what the shipped defaults will produce — 75.3% is.
-- **`full_exp` is the best full-gallery combo, but not the safest.** Per the
- isolated-effects analysis above, `expand_gallery=true` only cleanly pays off
- when paired with cast-restriction; on the full gallery it's flat on F1 while
- ~4x-ing misIDs (209→864, averaged across models). `full_noexp` scores lower
- (72.4% vs 75.3% for LVFace) but with **zero** true misIDs and higher precision
- (94.2% vs 89.7%). Kept `full_exp` as shipped since it's the highest-F1 option
- available without cast-restriction, but this is a real F1-vs-safety trade, not
- a strictly-better choice — worth revisiting if misID rate matters more than
- the last few points of F1 for a given deployment.
-- **Switching the default model is an operational change, not just a config tweak**:
- any existing gallery built from r50 embeddings is incompatible with LVFace
- embeddings and needs rebuilding.
+- r50's 4 combos used the older, narrower search bounds (1-30/1-15 instead
+ of 1-60/1-60) and are further confounded by its thinner gallery. Excluded
+ from all comparisons above except calibration.
+- The shipped defaults use `full_exp` (75.3% training F1), not the
+ higher-scoring `restricted_exp` (78.3%), because cast restriction is not
+ a runtime feature of the application yet.
+- `expand_gallery` is mode-dependent, not a free win. Averaged across models
+ on the full gallery it trades misIDs for recall (see the pose-expansion
+ table). For LVFace specifically, though, `full_exp` beats `full_noexp` on
+ every axis at once (F1 75.3 vs 72.3, precision 89.7 vs 88.3, recall 65.4 vs
+ 61.8, misID 232 vs 244), so the shipped `full_exp` is a clean choice for
+ this model, not an F1-vs-safety trade. (An earlier version of this page
+ reported `full_noexp` at 72.4% with zero misIDs and higher precision, which
+ made it look like the safer option; that was the dropped-film artifact
+ described above, not a real property of the config.)
+- Switching the default model is an operational change: any gallery built
+ from a different model's embeddings must be rebuilt before the new
+ default takes effect.
## Reproduce
```bash
-# 4-film matrix, all 4 models × 2 modes × 2 expansion settings
+# 4-film training matrix, all 4 models × 2 gallery modes × 2 expansion settings
bash experiments/run_rep4_subprocess.sh
# single combo
@@ -429,34 +326,21 @@ SAE_EXPAND=1 REPLAY_WORKERS=4 DE_WORKERS=2 python3 scripts/optimizer/optimize.py
--params prob_threshold:0.5:0.999 anneal_sec:1:60 extinction_sec:1:60 \
--popsize 10 --maxiter 15 --trajectory traj.jsonl --out best.json
-# replay the shipped config against a held-out film — --raw-out is needed to draw
-# bboxes later (the merged pred.json has no per-frame bbox, only actor windows)
-python3 scripts/optimizer/replay.py \
- --dump experiments/dumps/LVFace-B_Glint360K/dump_.h5 \
- --gallery experiments/galleries/gallery_LVFace-B_Glint360K.h5 \
- --out pred.json --raw-out raw.jsonl --prob-threshold 0.754 --anneal-sec 35.54 \
- --extinction-sec 57.43 --expand-gallery
+# held-out validation, all 3 models, 5 films
+python3 scripts/docs/run_holdout_all_models.py --out docs_data/holdout_all_models.json
-# regenerate the report's charts (16-combo ranking, DE landscape, held-out
-# per-film F1, Downton ghost timeline) from the artifacts under experiments/
+# per-film training breakdown, all 3 models, 4 films
+python3 scripts/docs/run_holdout_all_models.py --films training --out docs_data/training_per_film.json
+
+# gallery coverage per film
+python3 scripts/docs/gallery_coverage_per_film.py --out docs_data/gallery_coverage_per_film.json
+
+# regenerate this page's charts from experiments/ artifacts
python3 scripts/docs/experiment_charts.py --out-dir docs/assets/images
-# dump example frames (best-agreement / FPI / FN) for visual inspection, annotated
-# with bounding boxes + names (--raw is optional; omit for unannotated frames)
-python3 scripts/optimizer/dump_error_frames.py \
- --pred pred.json --raw raw.jsonl --xray experiments/xray/.../ \
- --movie "" \
- --gallery experiments/galleries/gallery_LVFace-B_Glint360K.h5 \
- --out-dir experiments/results/holdout/frames/ --n-per-bucket 4
-
-# or: one best + one worst frame per 10-minute window across the whole film
-python3 scripts/optimizer/dump_error_frames.py \
- --pred pred.json --raw raw.jsonl --xray experiments/xray/.../ \
- --movie "" \
- --gallery experiments/galleries/gallery_LVFace-B_Glint360K.h5 \
- --out-dir experiments/results/holdout/frames/_intervals --interval-sec 600
+# one frame per distinct out-of-cast name across all 9 films (used in the deep dive)
+python3 scripts/docs/first_fpi_frames.py
```
-See also: [the prior optimizer round](optimizer-experiments.md) (superseded
-metric) and the session log
+See also the session log
[`experiments/SESSION_STATE.md`](https://REPOLINK/experiments/SESSION_STATE.md).
diff --git a/docs/optimizer-experiments.md b/docs/optimizer-experiments.md
deleted file mode 100644
index 22b8b77..0000000
--- a/docs/optimizer-experiments.md
+++ /dev/null
@@ -1,125 +0,0 @@
-# Threshold optimization against Amazon X-Ray — experiment log
-
-Record of the July 2026 work that tuned the pipeline's recognition/tracking defaults
-against ground-truth per-scene actor presence, and the tooling built to do it.
-
-## TL;DR — what changed
-
-| knob | old default | new default | why |
-| ---- | ----------- | ----------- | --- |
-| `prob_threshold` | 0.99 | **0.76** | 0.99 was far too strict — halved recall for a fraction of a precision point. DE optimum, tightly converged. |
-| `extinction_sec` | 5.0 | **1.5** | Long extinction smears presence into later scenes → FPs. DE converged tightly low. |
-| `anneal_sec` | 10.0 | 10.0 (unchanged) | DE found it **insensitive** (F1 flat ±0.3pp across 3–26s) — kept the round default. |
-| `detector_conf` | 0.5 | 0.5 (unchanged) | Sweep showed raising it only trades recall for precision at a net F1 loss — near-threshold detections are real faces, not phantoms. |
-
-Net effect on the 9-film benchmark (strict per-scene, augmented gallery):
-recall **58% → ~72%**, F1 **70% → ~76%**, precision ~85%, at no meaningful precision cost.
-
-## Ground truth
-
-Public scene-level **Amazon X-Ray** dataset (Zenodo DOI 10.5281/zenodo.17659734,
-CC-BY-4.0): per movie, `people.csv` (name_id/person/character), `scenes.csv`
-(scene/start/end ms), `people_in_scenes.csv`. Films matched to the library by an
-**authoritative Jellyfin ID join** (query `/Items?IncludeItemTypes=Movie&Fields=
-ProviderIds,Path`, join Imdb/Tmdb against X-Ray metadata) — NOT fuzzy title matching,
-which collides badly (TV episodes vs same-named films). 9 genuine films with source
-video on disk: Benny & Joon, Café Society, Downton Abbey: A New Era, Lord of War,
-Lovelace, The Many Saints of Newark, Scarface, Sound of Metal, Valerian.
-
-## The scoring metric (evolved through review)
-
-Comparison unit is the **X-Ray scene**, not sampled timepoints. For each scene
-`[start,end]`: predicted set = **union** of actors detected anywhere in the span;
-GT set = actors X-Ray lists for that scene. Per scene TP/FP/FN, then:
-
-- **Precision: STRICT.** Any predicted actor not in the scene's X-Ray set is an FP,
- *including out-of-cast confusions* (no gallery∩cast masking). An earlier
- timepoint-sampled, cast-masked metric HID ~570 such FPs across 9 films and let the
- optimizer drive `prob_threshold` to the 0.50 floor — a metric artifact. Counting
- them is essential.
-- **Recall: FAIR.** FN counts only X-Ray cast members **who are in the gallery**. 67%
- of X-Ray cast (261/392) have no gallery reference embedding and can never be
- recognised — counting them as misses penalises coverage, not the threshold. Both
- `recall` (fair) and `recall_strict` (all) are reported.
-- **Aggregation:** per-scene F1 → **duration-weighted average within a movie** (long
- scenes count more) → **equal-weight mean across movies** (macro; each film counts
- the same regardless of length). This is the DE objective.
-
-Implemented in `scripts/optimizer/scene_score.py` — since **removed** along
-with this metric; its per-second successor is
-[`scripts/optimizer/second_score.py`](https://REPOLINK/scripts/optimizer/second_score.py)
-(see the [bake-off round](model-bakeoff.md)).
-
-## The gallery coverage gap
-
-Diagnosing low recall: only **131 of 392** X-Ray cast were in the gallery (33%). Every
-in-gallery actor HAD embeddings (gallery well-formed) — the gap was pure coverage.
-[`scripts/optimizer/fetch_missing_actors.py`](https://REPOLINK/scripts/optimizer/fetch_missing_actors.py)
-recovers missing actors:
-`nm-id → TMDB /find external_ids → /person/{id}/images → download → embed (sae_embed)`,
-with a `--wikidata` fallback (P345→P18 Commons photo).
-
-- **TMDB recovered 143/261** (55%). 0 face-detection failures; the rest had no TMDB
- person (60) or no profile photo (58). Coverage 33% → **70%**.
-- **Wikidata fallback: 0/118** of the TMDB failures — only 4 even had a Commons photo,
- none yielded a detectable face. → **TheTVDB not worth pursuing**: these remaining
- actors are obscure enough that no image source covers them, AND (see below) most are
- off-camera anyway.
-
-**Coverage vs detectability.** Adding references lifted recall (58→68% at fixed config)
-but modestly. Per-film drill-down (Lord of War: 12 actors recovered, only 1 had a
-detectable on-camera face) showed most missing cast are a **detectability gap** — X-Ray
-credits them as cast-in-scene (incl. off-camera/background), but their face never
-appears clearly for the pipeline to detect. This is a fundamental ceiling of a
-face-recognition pipeline vs X-Ray's presence semantics, not a fixable gap.
-
-## Optimizer
-
-[`scripts/optimizer/optimize.py`](https://REPOLINK/scripts/optimizer/optimize.py)
-— scipy `differential_evolution` over the knob space,
-each candidate = full replay of all films through the **real** C++ nodes (see the
-KPN replay architecture below) scored by the metric above. Global objective (one
-config for all films, not per-film).
-
-**Convergence stability (augmented gallery, 233 evals):**
-
-| knob | top-20 range | verdict |
-| ---- | ------------ | ------- |
-| `prob_threshold` | 0.69–0.83 (σ 0.05) | TIGHT — trust 0.76 |
-| `extinction_sec` | 1.0–2.2 (σ 0.33) | TIGHT — trust 1.5 |
-| `anneal_sec` | 3.1–26.3 (σ 6.4) | LOOSE — insensitive, not hard-coded |
-
-F1 varied only 0.3pp across the top-20 → objective is flat near the optimum, so only
-the tightly-converged knobs were adopted as defaults.
-
-## Replay architecture (how the sweep is cheap)
-
-The optimizer never re-decodes video. `scene_analyze --dump-embeddings out.h5` runs the
-expensive half once (decode→detect→align→embed) and dumps per-frame face embeddings
-+ metadata to HDF5 ([`scripts/optimizer/SCHEMA.md`](https://REPOLINK/scripts/optimizer/SCHEMA.md)).
-[`scripts/optimizer/replay.py`](https://REPOLINK/scripts/optimizer/replay.py) then
-replays that dump through the **real** C++ `face_tracker → identity_matcher →
-scene_tracker` assembled in a Python KPN network (`sae_kpn` nanobind module), varying
-Config knobs freely — no GPU embedding, no decode. Verified BYTE-EXACT against
-`scene_analyze`'s own output. The dumps are gallery-independent, so testing the
-augmented gallery needed no re-dump. `detector_conf` is replayable UPWARD only (the
-dump floor is 0.5).
-
-## Reproduce
-
-```bash
-# 1. dump (once per film, needs video)
-scene_analyze --movie --gallery gallery.json --dump-embeddings dump.h5 --fps 1
-# 2. build films manifest by Jellyfin ID join (see scripts/optimizer notes)
-# 3. optimize
-python scripts/optimizer/optimize.py --manifest films.json --gallery gallery.json \
- --params prob_threshold:0.5:0.999 anneal_sec:1:30 extinction_sec:1:15 \
- --popsize 8 --maxiter 20 --trajectory traj.jsonl --out opt.json
-# 4. score a fixed config / validate on a held-out set
-# (historical: score_config.py and scene_score.py were removed with the
-# scene-union metric — use scripts/optimizer/second_score.py, per-second)
-python scripts/optimizer/second_score.py --help
-```
-
-Superseded by the [model bake-off + re-tune](model-bakeoff.md), which
-replaced this round's scene-union metric with per-second scoring.
diff --git a/docs/pose-expansion.md b/docs/pose-expansion.md
index b8a5659..81fd109 100644
--- a/docs/pose-expansion.md
+++ b/docs/pose-expansion.md
@@ -1,42 +1,46 @@
-# Pose expansion: does "learning" new poses mid-film help?
+# Pose expansion: does promoting new poses mid-film help?
-`expand_gallery` ([`src/gallery/track_gallery.hpp`](https://REPOLINK/src/gallery/track_gallery.hpp))
-promotes a confidently-identified
-track's novel-pose reference views into a per-film, in-memory gallery annex — the
-idea being that once the pipeline is sure who someone is, a pose it hasn't seen
-before (turned head, different lighting) becomes a free extra reference for
-recognising that actor again later in the same film, without touching the baked
-gallery.
+`expand_gallery`
+([`src/gallery/track_gallery.hpp`](https://REPOLINK/src/gallery/track_gallery.hpp))
+promotes a confidently identified track's novel-pose reference views into a
+per-film, in-memory gallery annex. The idea: once the pipeline is confident
+about an identity, a pose it has not seen before (turned head, different
+lighting) becomes an extra reference for recognizing that actor again later
+in the same film, without touching the baked gallery.
-## The training-set signal
+## Training-set signal
-Averaged across all 4 models, on the 4 films used for optimization:
+Averaged across the 3 compared models (r50 excluded), on the 4 films used
+for optimization. These are the corrected, full-coverage figures, see the
+[dropped-film note](model-bakeoff.md#a-scoring-bug-worth-recording-dropped-film-evaluations)
+in the experiment log for why an earlier version of this table overstated the
+full-mode misID jump (209 → 864) that was itself partly a truncation artifact:
| scope | expansion | F1 | R | misID |
|---|---|---|---|---|
-| full | off | 71.2% | 58.3% | 209 |
-| full | **on** | 71.2% | 59.7% | **864** |
-| restricted | off | 73.6% | 61.3% | 194 |
-| restricted | **on** | **75.4%** | **64.5%** | 135 |
+| full | off | 70.0% | 57.6% | 407 |
+| full | on | 72.1% | 61.5% | 714 |
+| restricted | off | 75.1% | 63.9% | 179 |
+| restricted | on | 76.7% | 67.2% | 120 |
-In `restricted` mode (matcher's candidate set capped to the film's own credited
-cast) expansion looked like a clean win: +1.8pp F1, +3.2pp recall, misID actually
-lower. In `full` mode it looked flat-to-costly: ~0 F1 change, recall +1.4pp, but
-misID roughly quadrupled (209 → 864) — see the
-[bake-off experiment log](model-bakeoff.md) for the per-model breakdown. That's the number that motivated this page: **does turning
-expansion on actually change what gets recognised, frame by frame, or is the
-aggregate F1 shift something else?**
+In restricted mode, expansion looks like a clean win: +1.6pp F1, +3.3pp
+recall, lower misID. In full mode it looks like a recall-for-misID trade:
++2.1pp F1, +3.9pp recall, but misID rises from 407 to 714. See
+[the full experiment log](model-bakeoff.md) for the per-model breakdown.
+This asymmetry motivated the question below: does turning expansion on
+change what gets recognized frame by frame, or is the aggregate F1 shift
+coming from something else.
-## Held-out test: does it reproduce?
+## Held-out test
-Same model + same tuned config, `expand_gallery` toggled on vs. off, nothing else
-changed — full gallery mode, per-second scoring against X-Ray. This isolates
-expansion from every other variable (config, model, threshold) that differs
-between the training-set `exp`/`noexp` rows above.
+Same model, same tuned config, `expand_gallery` toggled on vs. off, nothing
+else changed, full gallery mode, per-second scoring against X-Ray. This
+isolates expansion from every other variable that differs between the
+training-set rows above.
-**LVFace-B Glint360K, all 5 held-out films** (films never seen by the optimizer):
+LVFace-B Glint360K, all 5 held-out films:
-| film | F1 (exp) | F1 (noexp) | TPI Δ | FN Δ |
+| film | F1 (exp) | F1 (noexp) | TPI delta | FN delta |
|---|---|---|---|---|
| Benny & Joon | 83.0% | 83.0% | -2 | +2 |
| Downton Abbey: A New Era | 56.1% | 56.2% | -7 | +7 |
@@ -44,57 +48,63 @@ between the training-set `exp`/`noexp` rows above.
| The Many Saints of Newark | 46.3% | 46.3% | +2 | -2 |
| Valerian and the City of a Thousand Planets | 74.1% | 74.1% | +2 | -2 |
-**ArcFace R18** (Benny & Joon, r18's own tuned config): F1 77.1% for both, TPI/FN
-identical, FPI differs by 2 (noise).
+ArcFace R18, Benny & Joon, r18's own tuned config: F1 77.1% for both, TPI
+and FN identical, FPI differs by 2.
-**Every film, both models tested: F1 within 0.1–0.2pp, TPI/FN swings in the tens
-out of tens of thousands.** That's noise, not a signal — expansion made no
-measurable difference to per-second onscreen identification anywhere it was
-tested on unseen data.
+Every film, both models tested: F1 differs by 0.1-0.2pp, TPI/FN swings are
+in the tens out of tens of thousands. This is noise, not a signal.
+Expansion made no measurable difference to per-second on-screen
+identification on any held-out film tested.
-## Two bugs this required catching (this section's own methodology)
+## Two methodology bugs caught during this check
-Getting to the clean table above took two wrong turns, both worth recording
-since they're exactly the kind of error that produces a false positive "look,
-expansion helped!" finding:
+Getting to the table above required catching two wrong turns, both worth
+recording because they are exactly the kind of error that produces a false
+positive "expansion helped" finding.
-1. **Timeout truncation.** The first Downton Abbey `exp` replay was cut off by a
- 60s subprocess timeout at ~76% through the film (5589 of 7368 expected
- seconds) — a genuinely large, silent data loss that showed up as a large,
- convincing-looking TPI gap (47938 vs 52032) purely because one run had a
- quarter of the film missing. Caught by comparing `n_seconds` between runs
- before trusting any score delta; fixed by re-running with a longer timeout.
-2. **Bbox-matching bug.** An early per-second raw-annotation diff matched each
- `exp` detection to the *first* `noexp` detection with IoU > 0.5, not the
- *best*-overlapping one. With 3 faces close together in frame, this produced
- spurious "disagreements" (e.g. "exp says Aidan Quinn, noexp says Johnny
- Depp" at the same seconds) that vanished entirely once the match picked the
- true best-IoU candidate — both configs had actually output the exact same
- three names at the exact same three boxes.
+1. **Timeout truncation.** The first Downton Abbey `exp` replay was cut off
+ by a 60-second subprocess timeout at about 76% through the film (5589 of
+ 7368 expected seconds). This silent data loss produced a large,
+ convincing-looking TPI gap (47938 vs 52032) purely because one run was
+ missing a quarter of the film. Caught by comparing `n_seconds` between
+ runs before trusting any score delta; fixed by re-running with a longer
+ timeout.
+2. **Bbox-matching bug.** An early per-second raw-annotation diff matched
+ each `exp` detection to the first `noexp` detection with IoU above 0.5,
+ not the best-overlapping one. With 3 faces close together in frame, this
+ produced spurious disagreements (for example "exp says Aidan Quinn,
+ noexp says Johnny Depp" at the same second) that vanished once the match
+ used the best-IoU candidate instead of the first one. Both configs had
+ actually output the same three names at the same three boxes.
-Both bugs independently pointed toward "expansion is doing something," and both
-were artifacts of the comparison harness, not the pipeline. Worth remembering
-when a before/after diff looks dramatic: check that the two runs actually cover
-the same seconds, and match entities by best overlap, not first-found.
+Both bugs independently pointed toward "expansion is doing something," and
+both were artifacts of the comparison harness, not the pipeline. Before
+trusting a dramatic before/after diff, check that both runs cover the same
+seconds and that entities are matched by best overlap, not first found.
-## What this means
+## Conclusion
-The training-set aggregate effect (particularly the ~4x misID increase in full
-mode) doesn't reproduce on held-out data — at minimum it's far smaller than the
-training-set numbers suggested, and plausibly it's sampling variation from only
-4 training films rather than a real, generalizable mechanism. This doesn't mean
-`expand_gallery` never does anything (the mechanism is real — see
+The training-set aggregate effect, particularly the full-mode misID
+increase, does not reproduce on held-out data. At minimum it
+is far smaller than the training-set numbers suggested; it may be sampling
+variation from only 4 training films rather than a generalizable
+mechanism. Note the same *class* of harness bug appears twice in this
+investigation, the timeout truncation in bug #1 above, and the dropped-film
+aggregation that inflated the raw training-set misID figures. Both make an
+inert config look consequential; both are reasons to distrust a dramatic
+training-set delta until it survives on held-out films, which this one did
+not. This does not mean `expand_gallery` never does anything: the
+mechanism is real, and
[`track_gallery.hpp`](https://REPOLINK/src/gallery/track_gallery.hpp)'s
-promotion logging: tracks *do* get confirmed and views *do*
-get promoted into the annex on every film tested), only that **whatever effect
-it has on final per-second identification was too small to detect against 5
-held-out films** with this scoring method. A cleaner test would need either many
-more held-out films or a metric that can see the annex's direct contribution
-(e.g. tagging which reference embedding won each match), neither of which this
-pass had budget for.
+promotion logging confirms tracks get confirmed and views get promoted
+into the annex on every film tested. It means whatever effect expansion
+has on final per-second identification was too small to detect against 5
+held-out films with this scoring method. A cleaner test would need either
+more held-out films or a metric that can see the annex's direct
+contribution, such as tagging which reference embedding won each match;
+neither was in scope for this pass.
-**Practical takeaway**: don't treat the training-set `exp` vs `noexp` numbers in
-the [bake-off experiment log](model-bakeoff.md) as proof that expansion
-changes real-world behavior
-in either direction — on the evidence gathered so far, it doesn't move the
-needle enough to see.
+Do not treat the training-set exp/noexp numbers in
+[the full experiment log](model-bakeoff.md) as proof that expansion changes
+real-world behavior in either direction. On the evidence gathered so far,
+it does not move the needle enough to see.
diff --git a/docs_data/holdout_all_models.json b/docs_data/holdout_all_models.json
new file mode 100644
index 0000000..25bbaee
--- /dev/null
+++ b/docs_data/holdout_all_models.json
@@ -0,0 +1,269 @@
+{
+ "LVFace-B_Glint360K": {
+ "config": {
+ "prob_threshold": 0.7540024664611272,
+ "anneal_sec": 35.53996030397922,
+ "extinction_sec": 57.43359645269811
+ },
+ "films": {
+ "Benny___Joon": {
+ "name": "Benny & Joon",
+ "TPI": 15119,
+ "FPI": 1845,
+ "FPI_misid": 0,
+ "FPI_incast": 1845,
+ "FN": 4343,
+ "precision": 0.8912402735203961,
+ "precision_raw": 0.8912402735203961,
+ "recall": 0.7768471893947179,
+ "f1": 0.8301213418986437,
+ "agreement_rate": 0.7239983093829193,
+ "exact_match_rate": 0.41098901098901097,
+ "n_seconds": 5915,
+ "duration_sec": 5915.0
+ },
+ "Downton_Abbey__A_New_Era": {
+ "name": "Downton Abbey: A New Era",
+ "TPI": 52022,
+ "FPI": 1160,
+ "FPI_misid": 0,
+ "FPI_incast": 1160,
+ "FN": 80089,
+ "precision": 0.9781881087586025,
+ "precision_raw": 0.9781881087586025,
+ "recall": 0.39377493168623356,
+ "f1": 0.5615106884771687,
+ "agreement_rate": 0.4057686401759256,
+ "exact_match_rate": 0.033084311632870865,
+ "n_seconds": 7496,
+ "duration_sec": 7496.0
+ },
+ "Lovelace": {
+ "name": "Lovelace",
+ "TPI": 14988,
+ "FPI": 1086,
+ "FPI_misid": 58,
+ "FPI_incast": 1028,
+ "FN": 7087,
+ "precision": 0.9031091829356471,
+ "precision_raw": 0.9324374766703994,
+ "recall": 0.6789580973952435,
+ "f1": 0.775154508546456,
+ "agreement_rate": 0.7204967829586512,
+ "exact_match_rate": 0.3597703211914588,
+ "n_seconds": 5573,
+ "duration_sec": 5573.0
+ },
+ "The_Many_Saints_of_Newark": {
+ "name": "The Many Saints of Newark",
+ "TPI": 15928,
+ "FPI": 4394,
+ "FPI_misid": 974,
+ "FPI_incast": 3420,
+ "FN": 23785,
+ "precision": 0.5475797579757976,
+ "precision_raw": 0.7837811239051274,
+ "recall": 0.40107773273235464,
+ "f1": 0.46301652592258835,
+ "agreement_rate": 0.3705156874642392,
+ "exact_match_rate": 0.04588936642173853,
+ "n_seconds": 7213,
+ "duration_sec": 7213.0
+ },
+ "Valerian_and_the_City_of_a_Thousand_Plan": {
+ "name": "Valerian and the City of a Thousand Planets",
+ "TPI": 18658,
+ "FPI": 548,
+ "FPI_misid": 0,
+ "FPI_incast": 548,
+ "FN": 12472,
+ "precision": 0.9714672498177653,
+ "precision_raw": 0.9714672498177653,
+ "recall": 0.5993575329264376,
+ "f1": 0.7413382072472983,
+ "agreement_rate": 0.5877853464704299,
+ "exact_match_rate": 0.21980294368081743,
+ "n_seconds": 8221,
+ "duration_sec": 8221.0
+ }
+ }
+ },
+ "arcface_w600k_mbf": {
+ "config": {
+ "prob_threshold": 0.8371114538930519,
+ "anneal_sec": 48.80011450565114,
+ "extinction_sec": 59.3110040220424
+ },
+ "films": {
+ "Benny___Joon": {
+ "name": "Benny & Joon",
+ "TPI": 15219,
+ "FPI": 2463,
+ "FPI_misid": 180,
+ "FPI_incast": 2283,
+ "FN": 4243,
+ "precision": 0.7884675163195524,
+ "precision_raw": 0.8607058025110281,
+ "recall": 0.7819854074606927,
+ "f1": 0.7852130843050252,
+ "agreement_rate": 0.7072306082196138,
+ "exact_match_rate": 0.34911242603550297,
+ "n_seconds": 5915,
+ "duration_sec": 5915.0
+ },
+ "Downton_Abbey__A_New_Era": {
+ "name": "Downton Abbey: A New Era",
+ "TPI": 53043,
+ "FPI": 2383,
+ "FPI_misid": 604,
+ "FPI_incast": 1779,
+ "FN": 79068,
+ "precision": 0.8715290328940882,
+ "precision_raw": 0.9570057373795692,
+ "recall": 0.4015032813316075,
+ "f1": 0.5497453011561202,
+ "agreement_rate": 0.4094114144765671,
+ "exact_match_rate": 0.032817502668089645,
+ "n_seconds": 7496,
+ "duration_sec": 7496.0
+ },
+ "Lovelace": {
+ "name": "Lovelace",
+ "TPI": 14606,
+ "FPI": 1337,
+ "FPI_misid": 180,
+ "FPI_incast": 1157,
+ "FN": 7469,
+ "precision": 0.8316346865569664,
+ "precision_raw": 0.916138744276485,
+ "recall": 0.6616534541336353,
+ "f1": 0.7369695746505879,
+ "agreement_rate": 0.6907677036961077,
+ "exact_match_rate": 0.31742329086667864,
+ "n_seconds": 5573,
+ "duration_sec": 5573.0
+ },
+ "The_Many_Saints_of_Newark": {
+ "name": "The Many Saints of Newark",
+ "TPI": 15223,
+ "FPI": 4574,
+ "FPI_misid": 994,
+ "FPI_incast": 3580,
+ "FN": 24490,
+ "precision": 0.52962460425147,
+ "precision_raw": 0.768954892155377,
+ "recall": 0.383325359454083,
+ "f1": 0.44475283393712756,
+ "agreement_rate": 0.3554753116932427,
+ "exact_match_rate": 0.03715513655899071,
+ "n_seconds": 7213,
+ "duration_sec": 7213.0
+ },
+ "Valerian_and_the_City_of_a_Thousand_Plan": {
+ "name": "Valerian and the City of a Thousand Planets",
+ "TPI": 18472,
+ "FPI": 853,
+ "FPI_misid": 239,
+ "FPI_incast": 614,
+ "FN": 12658,
+ "precision": 0.860122927919538,
+ "precision_raw": 0.9558602846054334,
+ "recall": 0.5933825891423065,
+ "f1": 0.7022773067710907,
+ "agreement_rate": 0.5795914643682625,
+ "exact_match_rate": 0.18817662084904513,
+ "n_seconds": 8221,
+ "duration_sec": 8221.0
+ }
+ }
+ },
+ "arcface_r18": {
+ "config": {
+ "prob_threshold": 0.8955101189489445,
+ "anneal_sec": 59.08214397442713,
+ "extinction_sec": 59.29474134414983
+ },
+ "films": {
+ "Benny___Joon": {
+ "name": "Benny & Joon",
+ "TPI": 13580,
+ "FPI": 1666,
+ "FPI_misid": 60,
+ "FPI_incast": 1606,
+ "FN": 5882,
+ "precision": 0.86025592296972,
+ "precision_raw": 0.8907254361799817,
+ "recall": 0.697770013359367,
+ "f1": 0.7705401724920563,
+ "agreement_rate": 0.6547675401521545,
+ "exact_match_rate": 0.32578191039729504,
+ "n_seconds": 5915,
+ "duration_sec": 5915.0
+ },
+ "Downton_Abbey__A_New_Era": {
+ "name": "Downton Abbey: A New Era",
+ "TPI": 48545,
+ "FPI": 1066,
+ "FPI_misid": 180,
+ "FPI_incast": 886,
+ "FN": 83566,
+ "precision": 0.9475708067381078,
+ "precision_raw": 0.9785128298159682,
+ "recall": 0.3674561542944948,
+ "f1": 0.5295567845883649,
+ "agreement_rate": 0.3815368792000116,
+ "exact_match_rate": 0.032950907150480255,
+ "n_seconds": 7496,
+ "duration_sec": 7496.0
+ },
+ "Lovelace": {
+ "name": "Lovelace",
+ "TPI": 13615,
+ "FPI": 963,
+ "FPI_misid": 120,
+ "FPI_incast": 843,
+ "FN": 8460,
+ "precision": 0.8695235662281262,
+ "precision_raw": 0.933941555768967,
+ "recall": 0.6167610419026047,
+ "f1": 0.7216494845360825,
+ "agreement_rate": 0.6506356469257915,
+ "exact_match_rate": 0.2894311860757222,
+ "n_seconds": 5573,
+ "duration_sec": 5573.0
+ },
+ "The_Many_Saints_of_Newark": {
+ "name": "The Many Saints of Newark",
+ "TPI": 13489,
+ "FPI": 3757,
+ "FPI_misid": 796,
+ "FPI_incast": 2961,
+ "FN": 26224,
+ "precision": 0.5526013928717739,
+ "precision_raw": 0.7821523831613127,
+ "recall": 0.3396620753909299,
+ "f1": 0.42072267361165266,
+ "agreement_rate": 0.3229817885335633,
+ "exact_match_rate": 0.04422570359073894,
+ "n_seconds": 7213,
+ "duration_sec": 7213.0
+ },
+ "Valerian_and_the_City_of_a_Thousand_Plan": {
+ "name": "Valerian and the City of a Thousand Planets",
+ "TPI": 17692,
+ "FPI": 397,
+ "FPI_misid": 68,
+ "FPI_incast": 329,
+ "FN": 13438,
+ "precision": 0.9460456660071654,
+ "precision_raw": 0.9780529603626513,
+ "recall": 0.5683263732733698,
+ "f1": 0.710080070638759,
+ "agreement_rate": 0.5633540120828806,
+ "exact_match_rate": 0.13404695292543486,
+ "n_seconds": 8221,
+ "duration_sec": 8221.0
+ }
+ }
+ }
+}
\ No newline at end of file
diff --git a/docs_data/training_per_film.json b/docs_data/training_per_film.json
new file mode 100644
index 0000000..767e5a8
--- /dev/null
+++ b/docs_data/training_per_film.json
@@ -0,0 +1,221 @@
+{
+ "LVFace-B_Glint360K": {
+ "config": {
+ "prob_threshold": 0.7540024664611272,
+ "anneal_sec": 35.53996030397922,
+ "extinction_sec": 57.43359645269811
+ },
+ "films": {
+ "Caf\u00e9_Society": {
+ "name": "Caf\u00e9 Society",
+ "TPI": 14499,
+ "FPI": 1380,
+ "FPI_misid": 0,
+ "FPI_incast": 1380,
+ "FN": 12231,
+ "precision": 0.9130927640279615,
+ "precision_raw": 0.9130927640279615,
+ "recall": 0.5424242424242425,
+ "f1": 0.6805604449764134,
+ "agreement_rate": 0.57285804629501,
+ "exact_match_rate": 0.18947003810183582,
+ "n_seconds": 5774,
+ "duration_sec": 5774.0
+ },
+ "Lord_of_War": {
+ "name": "Lord of War",
+ "TPI": 13893,
+ "FPI": 1654,
+ "FPI_misid": 174,
+ "FPI_incast": 1480,
+ "FN": 5737,
+ "precision": 0.811838952842868,
+ "precision_raw": 0.8936129156750499,
+ "recall": 0.7077432501273561,
+ "f1": 0.7562256756388972,
+ "agreement_rate": 0.7005158404089996,
+ "exact_match_rate": 0.38715420432758146,
+ "n_seconds": 7302,
+ "duration_sec": 7302.0
+ },
+ "Scarface": {
+ "name": "Scarface",
+ "TPI": 20518,
+ "FPI": 1078,
+ "FPI_misid": 58,
+ "FPI_incast": 1020,
+ "FN": 14722,
+ "precision": 0.9276607288181572,
+ "precision_raw": 0.9500833487682904,
+ "recall": 0.5822360953461975,
+ "f1": 0.7154363820216882,
+ "agreement_rate": 0.6297735703976657,
+ "exact_match_rate": 0.25910733470065433,
+ "n_seconds": 10239,
+ "duration_sec": 10239.0
+ },
+ "Sound_of_Metal": {
+ "name": "Sound of Metal",
+ "TPI": 13349,
+ "FPI": 677,
+ "FPI_misid": 0,
+ "FPI_incast": 677,
+ "FN": 6504,
+ "precision": 0.9517324967916726,
+ "precision_raw": 0.9517324967916726,
+ "recall": 0.6723920818012391,
+ "f1": 0.7880397886596416,
+ "agreement_rate": 0.7112222835587533,
+ "exact_match_rate": 0.40311896218603366,
+ "n_seconds": 7246,
+ "duration_sec": 7246.0
+ }
+ }
+ },
+ "arcface_w600k_mbf": {
+ "config": {
+ "prob_threshold": 0.8371114538930519,
+ "anneal_sec": 48.80011450565114,
+ "extinction_sec": 59.3110040220424
+ },
+ "films": {
+ "Caf\u00e9_Society": {
+ "name": "Caf\u00e9 Society",
+ "TPI": 13456,
+ "FPI": 1434,
+ "FPI_misid": 180,
+ "FPI_incast": 1254,
+ "FN": 13274,
+ "precision": 0.8150211992731677,
+ "precision_raw": 0.9036937541974479,
+ "recall": 0.5034044145155256,
+ "f1": 0.622386679000925,
+ "agreement_rate": 0.5463565775524695,
+ "exact_match_rate": 0.19154832005542086,
+ "n_seconds": 5774,
+ "duration_sec": 5774.0
+ },
+ "Lord_of_War": {
+ "name": "Lord of War",
+ "TPI": 13778,
+ "FPI": 1740,
+ "FPI_misid": 60,
+ "FPI_incast": 1680,
+ "FN": 5852,
+ "precision": 0.8580146967243741,
+ "precision_raw": 0.8878721484727413,
+ "recall": 0.7018848700967907,
+ "f1": 0.7721362923111411,
+ "agreement_rate": 0.6881858851455998,
+ "exact_match_rate": 0.36469460421802247,
+ "n_seconds": 7302,
+ "duration_sec": 7302.0
+ },
+ "Scarface": {
+ "name": "Scarface",
+ "TPI": 18863,
+ "FPI": 862,
+ "FPI_misid": 0,
+ "FPI_incast": 862,
+ "FN": 16377,
+ "precision": 0.956299112801014,
+ "precision_raw": 0.956299112801014,
+ "recall": 0.535272417707151,
+ "f1": 0.6863640498499044,
+ "agreement_rate": 0.594178111391709,
+ "exact_match_rate": 0.2357652114464303,
+ "n_seconds": 10239,
+ "duration_sec": 10239.0
+ },
+ "Sound_of_Metal": {
+ "name": "Sound of Metal",
+ "TPI": 12642,
+ "FPI": 554,
+ "FPI_misid": 0,
+ "FPI_incast": 554,
+ "FN": 7211,
+ "precision": 0.9580175810851773,
+ "precision_raw": 0.9580175810851773,
+ "recall": 0.6367803354656727,
+ "f1": 0.7650458410239341,
+ "agreement_rate": 0.687468948385309,
+ "exact_match_rate": 0.3789677063207287,
+ "n_seconds": 7246,
+ "duration_sec": 7246.0
+ }
+ }
+ },
+ "arcface_r18": {
+ "config": {
+ "prob_threshold": 0.8955101189489445,
+ "anneal_sec": 59.08214397442713,
+ "extinction_sec": 59.29474134414983
+ },
+ "films": {
+ "Caf\u00e9_Society": {
+ "name": "Caf\u00e9 Society",
+ "TPI": 12207,
+ "FPI": 1119,
+ "FPI_misid": 60,
+ "FPI_incast": 1059,
+ "FN": 14523,
+ "precision": 0.88035482475119,
+ "precision_raw": 0.9160288158487168,
+ "recall": 0.45667789001122333,
+ "f1": 0.6013892994383683,
+ "agreement_rate": 0.5125626845924863,
+ "exact_match_rate": 0.1674748874263942,
+ "n_seconds": 5774,
+ "duration_sec": 5774.0
+ },
+ "Lord_of_War": {
+ "name": "Lord of War",
+ "TPI": 13122,
+ "FPI": 1409,
+ "FPI_misid": 60,
+ "FPI_incast": 1349,
+ "FN": 6508,
+ "precision": 0.870678787074514,
+ "precision_raw": 0.9030348909228546,
+ "recall": 0.6684666327050433,
+ "f1": 0.7562894441082388,
+ "agreement_rate": 0.6698963754222389,
+ "exact_match_rate": 0.3389482333607231,
+ "n_seconds": 7302,
+ "duration_sec": 7302.0
+ },
+ "Scarface": {
+ "name": "Scarface",
+ "TPI": 16961,
+ "FPI": 685,
+ "FPI_misid": 0,
+ "FPI_incast": 685,
+ "FN": 18279,
+ "precision": 0.961181004193585,
+ "precision_raw": 0.961181004193585,
+ "recall": 0.48129965947786607,
+ "f1": 0.6414173883447416,
+ "agreement_rate": 0.5440709115628456,
+ "exact_match_rate": 0.2017775173356773,
+ "n_seconds": 10239,
+ "duration_sec": 10239.0
+ },
+ "Sound_of_Metal": {
+ "name": "Sound of Metal",
+ "TPI": 12017,
+ "FPI": 591,
+ "FPI_misid": 122,
+ "FPI_incast": 469,
+ "FN": 7836,
+ "precision": 0.8767692981176127,
+ "precision_raw": 0.953125,
+ "recall": 0.6052989472623784,
+ "f1": 0.7161715188176049,
+ "agreement_rate": 0.6594534915815532,
+ "exact_match_rate": 0.3573005796301408,
+ "n_seconds": 7246,
+ "duration_sec": 7246.0
+ }
+ }
+ }
+}
\ No newline at end of file
diff --git a/mkdocs.yml b/mkdocs.yml
index ae13c90..2ec5120 100644
--- a/mkdocs.yml
+++ b/mkdocs.yml
@@ -34,13 +34,13 @@ extra_css:
nav:
- Home: index.md
+ - How We Score Against X-Ray: methodology.md
- Findings:
- Best Model: best-model.md
- Gallery Scope (Full vs. Limited): gallery-scope.md
- Pose Expansion: pose-expansion.md
- LVFace Deep Dive: lvface-deep-dive.md
- - Model Bake-off & Re-tune (full log): model-bakeoff.md
- - Optimizer Experiments (prior round): optimizer-experiments.md
+ - Full Experiment Log: model-bakeoff.md
- Service Conversion (proposal): service-conversion.md
markdown_extensions:
diff --git a/scripts/docs/build_site.sh b/scripts/docs/build_site.sh
index bbb0664..ccf0a2e 100755
--- a/scripts/docs/build_site.sh
+++ b/scripts/docs/build_site.sh
@@ -60,10 +60,34 @@ stage_frame "${MONTAGE_ROOT}/Valerian_and_the_City_of_a_Thousand_Plan/scene_4/4_
valerian_screen_call.jpg
stage_frame "${MONTAGE_ROOT}/The_Many_Saints_of_Newark/out_of_cast_fpi/4_worst_t000871.jpg" \
many_saints_outofcast_fpi.jpg
-# debug-overlay example (extinction state drawn as frozen boxes) — from the
-# dump_error_frames output, not the montage package
-stage_frame "experiments/results/holdout/frames/many_saints/fpi/fpi_t03543.jpg" \
- many_saints_ghost_fpi.jpg
+
+# One frame per DISTINCT out-of-cast name across all 9 films, uniform rule
+# (see scripts/docs/first_fpi_frames.py): the first second in the raw replay
+# stream where the pipeline names someone not in the film's credited cast at
+# all. Rendered with the proper montage renderer (Onscreen/Offscreen panel),
+# never dump_error_frames.py's bare-box overlay. Regenerate with:
+# python3 scripts/docs/first_fpi_frames.py
+# 5 of 9 films have zero out-of-cast names in their whole runtime (Benny &
+# Joon, Cafe Society, Downton Abbey, Sound of Metal, Valerian) and produce
+# no frames here.
+stage_frame "${MONTAGE_ROOT}/Lord_of_War/first_fpi_david_shumbris/first_fpi_t000418.jpg" \
+ lord_of_war_fpi_shumbris.jpg
+stage_frame "${MONTAGE_ROOT}/Lord_of_War/first_fpi_ronald_reagan/first_fpi_t001003.jpg" \
+ lord_of_war_fpi_reagan_photo.jpg
+stage_frame "${MONTAGE_ROOT}/Lord_of_War/first_fpi_lance_reddick/first_fpi_t006424.jpg" \
+ lord_of_war_fpi_reddick.jpg
+stage_frame "${MONTAGE_ROOT}/Lovelace/first_fpi_chloë_sevigny/first_fpi_t002451.jpg" \
+ lovelace_fpi_sevigny.jpg
+stage_frame "${MONTAGE_ROOT}/Scarface/first_fpi_kirstie_alley/first_fpi_t002451.jpg" \
+ scarface_fpi_alley.jpg
+stage_frame "${MONTAGE_ROOT}/The_Many_Saints_of_Newark/first_fpi_germar_terrell_gardner/first_fpi_t000848.jpg" \
+ many_saints_fpi_gardner.jpg
+stage_frame "${MONTAGE_ROOT}/The_Many_Saints_of_Newark/first_fpi_archie_yates/first_fpi_t002521.jpg" \
+ many_saints_fpi_yates.jpg
+stage_frame "${MONTAGE_ROOT}/The_Many_Saints_of_Newark/first_fpi_zooey_deschanel/first_fpi_t002819.jpg" \
+ many_saints_fpi_deschanel.jpg
+stage_frame "${MONTAGE_ROOT}/The_Many_Saints_of_Newark/first_fpi_talia_balsam/first_fpi_t004551.jpg" \
+ many_saints_fpi_balsam.jpg
if [ ! -f "${ASSETS_DIR}/germar_beats_xray.jpg" ]; then
echo "==> pulling report-highlights/germar_beats_xray.jpg..."
diff --git a/scripts/docs/experiment_charts.py b/scripts/docs/experiment_charts.py
index deeb09c..032b576 100644
--- a/scripts/docs/experiment_charts.py
+++ b/scripts/docs/experiment_charts.py
@@ -27,8 +27,11 @@ RESULTS = REPO / "experiments/results"
# Same model -> color mapping as calibration_chart.py, so identity is stable
# across every figure in the report.
+# r50 is dropped from the bake-off: its 4 combos ran under the old, narrower
+# anneal/extinction bounds and were never re-run wide, so they are not comparable
+# on those two params (and two of them were truncation-corrupted). Its slug stays
+# out of this map so it never appears in a figure or legend.
MODEL_COLOURS = {
- "arcface_w600k_r50": ("ArcFace w600k-R50", "#2a78d6"),
"arcface_r18": ("ArcFace R18", "#008300"),
"arcface_w600k_mbf": ("ArcFace w600k-MBF", "#e87ba4"),
"LVFace-B_Glint360K": ("LVFace-B Glint360K", "#eda100"),
@@ -59,9 +62,37 @@ plt.rcParams.update({
})
+TRAJ = REPO / "experiments/trajectories"
+
+
+def clean_best(combo: str) -> dict:
+ """Best-F1 eval for a combo, restricted to FULL-COVERAGE evals.
+
+ The optimizer averages F1 (and *sums* TPI/misID) over only the films whose
+ replay subprocess didn't time out (optimize.py: `per_film = [... if m is not
+ None]`). A candidate whose hardest film timed out is therefore scored on an
+ easier subset, which inflates its F1 — and DE will happily converge onto such
+ a candidate. `rep4_best_*.json` recorded exactly that kind of eval for at
+ least one combo (arcface_w600k_mbf_full_noexp: reported 74.2% F1 came from an
+ eval with TPI 12645, a third of that combo's median).
+
+ We recover comparable numbers straight from the trajectory: take the median
+ TPI across all evals (full 4-film coverage) and keep only evals within 30% of
+ it, then pick the highest-F1 survivor. No re-running — the honest best config
+ is already in the sweep, just not the one `argmax f1` picked.
+ """
+ evals = [json.loads(l) for l in open(TRAJ / f"rep4_{combo}.jsonl")]
+ tpis = sorted(e["TPI"] for e in evals)
+ med = tpis[len(tpis) // 2]
+ clean = [e for e in evals if e["TPI"] >= 0.7 * med]
+ return max(clean, key=lambda e: e["f1"])
+
+
def training_best() -> dict:
- with open(RESULTS / "rep4_best_LVFace-B_Glint360K_full_exp.json") as f:
- return json.load(f)["best"]
+ # LVFace-B_Glint360K_full_exp is the shipped combo; its reported best is a
+ # full-coverage eval (TPI 47757 ≈ median), so clean_best returns the same
+ # config — but route it through clean_best so every figure uses one path.
+ return clean_best("LVFace-B_Glint360K_full_exp")
def fig_holdout_f1(out: Path):
@@ -100,18 +131,16 @@ def fig_holdout_f1(out: Path):
def fig_rep4_matrix(out: Path):
combos = []
- for path in sorted(RESULTS.glob("rep4_best_*.json")):
- stem = path.stem[len("rep4_best_"):]
+ for path in sorted(TRAJ.glob("rep4_*.jsonl")):
+ combo = path.stem[len("rep4_"):]
for slug in MODEL_COLOURS:
- if stem.startswith(slug):
- mode = stem[len(slug) + 1:] # e.g. full_exp
- with open(path) as f:
- best = json.load(f)["best"]
- combos.append((slug, mode, best["f1"] * 100))
+ if combo.startswith(slug):
+ mode = combo[len(slug) + 1:] # e.g. full_exp
+ combos.append((slug, mode, clean_best(combo)["f1"] * 100))
break
combos.sort(key=lambda c: c[2])
- fig, ax = plt.subplots(figsize=(9, 6.2))
+ fig, ax = plt.subplots(figsize=(9, 5.2))
ax.grid(axis="y", visible=False)
labels = []
for i, (slug, mode, f1) in enumerate(combos):
@@ -126,7 +155,7 @@ def fig_rep4_matrix(out: Path):
ax.set_yticks(range(len(combos)), labels, fontsize=9)
ax.set_xlim(65, 80)
ax.set_xlabel("training-set per-second F1 (%)")
- ax.set_title("All 16 combos — filled dot = cast-restricted gallery, open = full",
+ ax.set_title("All 12 combos — filled dot = cast-restricted gallery, open = full",
loc="left", fontsize=12, pad=12)
handles = [plt.Line2D([], [], marker="o", ls="", ms=9, color=c, label=l)
for _, (l, c) in MODEL_COLOURS.items()]
@@ -199,13 +228,16 @@ def fig_downton_timeline(out: Path, t0: int = 7100, t1: int = 7340):
trk = np.array(trk)
dc = np.array([det.get(s, 0) for s in t])
- fig, ax = plt.subplots(figsize=(9.5, 4.4))
+ fig, ax = plt.subplots(figsize=(9.5, 4.8))
ax.grid(axis="x", visible=False)
- ax.fill_between(t, dc, step="mid", color=GREEN, alpha=0.25, zorder=2)
- ax.step(t, dc, where="mid", color=GREEN, lw=2, zorder=3)
- ax.step(t, trk, where="mid", color=BLUE, lw=2, zorder=4)
+ ax.fill_between(t, dc, step="mid", color=GREEN, alpha=0.22, zorder=2)
+ ax.step(t, dc, where="mid", color=GREEN, lw=2, zorder=3,
+ label="faces seen by detector")
+ ax.step(t, trk, where="mid", color=BLUE, lw=2, zorder=4,
+ label="actors reported by tracker")
# longest contiguous run of "detector sees nothing, tracker still reporting"
+ # (i.e. every reported actor is extinction-bridged, not detected this second)
ghost = (dc == 0) & (trk > 0)
runs, start = [], None
for i, g in enumerate(ghost):
@@ -219,15 +251,11 @@ def fig_downton_timeline(out: Path, t0: int = 7100, t1: int = 7340):
if runs:
i0, i1 = max(runs, key=lambda r: r[1] - r[0])
g0, g1 = t[i0], t[i1]
- ax.axvspan(g0, g1, color=RED, alpha=0.08, zorder=1)
- ax.annotate(f"{g1 - g0}s of credits: 0 faces detected,\n"
- f"{trk[i0]} actors still reported (frozen boxes)",
- ((g0 + g1) / 2, 20.5), ha="center", va="bottom",
- fontsize=10, color=RED)
- ax.text(t0 + 4, 27.3, "actors reported by tracker", color=BLUE,
- fontsize=10.5, va="bottom")
- ax.text(t0 + 4, 11.5, "faces seen by detector", color=GREEN,
- fontsize=10.5, va="bottom")
+ ax.axvspan(g0, g1, color=RED, alpha=0.08, zorder=1,
+ label=f"{g1 - g0}s bridged: 0 faces detected,\n"
+ f"{trk[i0]} actors carried by their\nextinction window")
+ ax.legend(loc="upper right", frameon=True, framealpha=0.92,
+ edgecolor=GRID, fontsize=9.5)
ax.set_xlabel("film time (s)")
ax.set_ylabel("count")
ax.set_ylim(0, 31)
diff --git a/scripts/docs/first_fpi_frames.py b/scripts/docs/first_fpi_frames.py
new file mode 100644
index 0000000..9fbb68b
--- /dev/null
+++ b/scripts/docs/first_fpi_frames.py
@@ -0,0 +1,155 @@
+#!/usr/bin/env python3
+"""
+first_fpi_frames.py — for every film, find every DISTINCT out-of-cast name
+(misID) the raw replay stream ever reports, and render the exact second each
+one FIRST appears, with the proper montage renderer (dump_scene_montage.py:
+Onscreen/Offscreen panel, TPI/FPI/FN legend, ghosts never drawn as boxes —
+imported directly, not the scene-level best/worst picker, which can land on
+a different second within the same scene).
+
+One rule, applied uniformly across all 9 films and every distinct wrong name
+in each — no manual per-film picking, no stopping at the first name found.
+"""
+import csv
+import json
+import sys
+from pathlib import Path
+
+import cv2
+
+REPO = Path(__file__).resolve().parent.parent.parent
+sys.path.insert(0, str(REPO / "scripts" / "validation"))
+sys.path.insert(0, str(REPO / "scripts" / "optimizer"))
+from identity import keys_for # noqa: E402
+from sample_eval import load_gallery_keys # noqa: E402
+from dump_scene_montage import ( # noqa: E402
+ classify_second, extract_frame, render_frame,
+ load_scene_cast, load_dump_faces_by_second, load_raw_by_second,
+)
+
+FILMS = [
+ ("Benny___Joon", "experiments/xray/scene_level_movie_data_XRay_US/xrays/4808_Benny__Joon"),
+ ("Café_Society", "experiments/xray/scene_level_movie_data_XRay_US/xrays/225_Cafe_Society"),
+ ("Downton_Abbey__A_New_Era", "experiments/xray/scene_level_movie_data_XRay_US/xrays/19_Downton_Abbey_A_New_Era"),
+ ("Lord_of_War", "experiments/xray/scene_level_movie_data_XRay_US/xrays/2474_Lord_of_War"),
+ ("Lovelace", "experiments/xray/scene_level_movie_data_XRay_US/xrays/4108_Lovelace"),
+ ("Scarface", "experiments/xray/scene_level_movie_data_XRay_US/xrays/197_Scarface"),
+ ("Sound_of_Metal", "experiments/xray/scene_level_movie_data_XRay_US/xrays/6278_Sound_of_Metal"),
+ ("The_Many_Saints_of_Newark", "experiments/xray/scene_level_movie_data_XRay_US/xrays/900_The_Many_Saints_Of_Newark"),
+ ("Valerian_and_the_City_of_a_Thousand_Plan", "experiments/xray/scene_level_movie_data_XRay_US/xrays/5312_Valerian_and_the_City_of_a_Thousand_Planets"),
+]
+
+MOVIE_ROOT = Path("/mnt/movies")
+
+
+def load_film_cast_keys(xray_dir: Path) -> set:
+ keys = set()
+ with open(xray_dir / "people.csv", newline="", encoding="utf-8") as f:
+ for r in csv.DictReader(f):
+ nm = (r.get("name_id") or "").strip()
+ person = (r.get("person") or "").strip()
+ if nm or person:
+ keys |= keys_for(imdb_id=nm, name=person)
+ return keys
+
+
+def find_movie_file(slug: str) -> str | None:
+ # dump HDF5 attrs carry the exact path used at dump time
+ import h5py
+ for model in ("LVFace-B_Glint360K",):
+ p = REPO / f"experiments/dumps/{model}/dump_{slug}.h5"
+ if p.exists():
+ with h5py.File(p, "r") as f:
+ return f.attrs.get("movie")
+ return None
+
+
+def find_scene_id(xray_dir: Path, t: int) -> str | None:
+ with open(xray_dir / "scenes.csv", newline="", encoding="utf-8") as f:
+ for r in csv.DictReader(f):
+ try:
+ t0, t1 = float(r["start"]) / 1000.0, float(r["end"]) / 1000.0
+ except (KeyError, ValueError):
+ continue
+ if t0 <= t < t1:
+ return (r.get("scene") or "").strip()
+ return None
+
+
+def main():
+ out_root = REPO / "experiments/results/holdout/montage_bestworst"
+ summary = []
+
+ for slug, xray_rel in FILMS:
+ xray_dir = REPO / xray_rel
+ raw_path = out_root / f"raw_{slug}.jsonl"
+ if not raw_path.exists():
+ print(f"SKIP {slug}: no raw file", file=sys.stderr)
+ continue
+
+ cast_keys = load_film_cast_keys(xray_dir)
+
+ # every distinct out-of-cast name -> first second it appears
+ first_seen: dict[str, int] = {}
+ with open(raw_path) as f:
+ for line in f:
+ d = json.loads(line)
+ if d.get("eof"):
+ continue
+ for a in d.get("visible_actors", []):
+ name = a.get("name")
+ if not name or name in first_seen:
+ continue
+ ak = keys_for(imdb_id=a.get("imdb_id"), name=name,
+ jellyfin_id=a.get("jellyfin_id"))
+ if not (ak & cast_keys):
+ first_seen[name] = int(d["timestamp_sec"])
+
+ if not first_seen:
+ print(f"{slug}: no out-of-cast FPI in the whole film", file=sys.stderr)
+ summary.append((slug, None, None))
+ continue
+
+ print(f"{slug}: {len(first_seen)} distinct out-of-cast name(s)", file=sys.stderr)
+
+ movie = find_movie_file(slug)
+ if not movie or not Path(movie).exists():
+ print(f" SKIP render: movie file not found ({movie})", file=sys.stderr)
+ for name, t in first_seen.items():
+ summary.append((slug, name, t))
+ continue
+
+ dump_path = REPO / f"experiments/dumps/LVFace-B_Glint360K/dump_{slug}.h5"
+ gallery_path = REPO / "experiments/galleries/gallery_LVFace-B_Glint360K.h5"
+ gallery_keys = load_gallery_keys(str(gallery_path))
+ raw_by_second = load_raw_by_second(str(raw_path))
+ dump_faces_by_second = load_dump_faces_by_second(str(dump_path))
+ scene_cast = load_scene_cast(str(xray_dir))
+
+ for name, t in sorted(first_seen.items(), key=lambda kv: kv[1]):
+ scene_id = find_scene_id(xray_dir, t)
+ gt_cast = scene_cast.get(scene_id, set())
+ gt_cast = {g for g in gt_cast if g & gallery_keys}
+
+ score, tpi_boxes, fpi_boxes, entries, has_outofcast = classify_second(
+ t, gt_cast, cast_keys, raw_by_second, dump_faces_by_second)
+
+ slug_name = name.lower().replace(" ", "_").replace("'", "")
+ out_dir = out_root / slug / f"first_fpi_{slug_name}"
+ out_dir.mkdir(parents=True, exist_ok=True)
+ out_path = out_dir / f"first_fpi_t{t:06d}.jpg"
+ extract_frame(movie, t, out_path)
+ canvas = render_frame(out_path, t, tpi_boxes, fpi_boxes, entries)
+ if canvas is not None:
+ cv2.imwrite(str(out_path), canvas)
+ print(f" {name!r} t={t}s -> {out_path} (outofcast={has_outofcast})",
+ file=sys.stderr)
+ summary.append((slug, name, t))
+
+ print("\n=== summary ===", file=sys.stderr)
+ for slug, name, t in summary:
+ print(f" {slug:45s} {name!r:30s} t={t}", file=sys.stderr)
+
+
+if __name__ == "__main__":
+ main()
diff --git a/scripts/docs/gallery_coverage_per_film.py b/scripts/docs/gallery_coverage_per_film.py
new file mode 100644
index 0000000..1351471
--- /dev/null
+++ b/scripts/docs/gallery_coverage_per_film.py
@@ -0,0 +1,62 @@
+#!/usr/bin/env python3
+"""
+gallery_coverage_per_film.py — fraction of each film's X-Ray credited cast
+that has a reference embedding in the gallery, computed per film rather than
+as a single benchmark-wide average.
+
+Usage: python3 scripts/docs/gallery_coverage_per_film.py --out docs_data/gallery_coverage_per_film.json
+"""
+import argparse
+import csv
+import json
+import sys
+from pathlib import Path
+
+import h5py
+
+REPO = Path(__file__).resolve().parent.parent.parent
+sys.path.insert(0, str(REPO / "scripts" / "validation"))
+from identity import keys_for # noqa: E402
+
+
+def main():
+ p = argparse.ArgumentParser()
+ p.add_argument("--gallery", default=str(REPO / "experiments/galleries/gallery_LVFace-B_Glint360K.h5"))
+ p.add_argument("--films", default=str(REPO / "experiments/manifests/films.json"))
+ p.add_argument("--out", required=True)
+ args = p.parse_args()
+
+ films = json.load(open(args.films))
+ with h5py.File(args.gallery, "r") as f:
+ names = [n.decode() if isinstance(n, bytes) else n for n in f["name"][:]]
+ jids = [j.decode() if isinstance(j, bytes) else j for j in f["jellyfin_id"][:]]
+ imdbs = [j.decode() if isinstance(j, bytes) else j for j in f["imdb_id"][:]]
+ gallery_keys = set()
+ for n, j, im in zip(names, jids, imdbs):
+ gallery_keys |= keys_for(imdb_id=im, name=n, jellyfin_id=j)
+
+ out = []
+ for film in films:
+ xray_dir = REPO / film["xray"]
+ id_to_name = {}
+ with open(xray_dir / "people.csv", newline="", encoding="utf-8") as fh:
+ for r in csv.DictReader(fh):
+ nm = (r.get("name_id") or "").strip()
+ if nm:
+ id_to_name[nm] = (r.get("person") or "").strip()
+ cast_keys = [keys_for(imdb_id=nm, name=name) for nm, name in id_to_name.items()]
+ covered = sum(1 for ck in cast_keys if ck & gallery_keys)
+ total = len(cast_keys)
+ out.append({"film": film["name"], "cast_total": total, "covered": covered,
+ "coverage_pct": round(covered / total * 100, 1) if total else 0.0})
+
+ out.sort(key=lambda x: x["coverage_pct"])
+ Path(args.out).parent.mkdir(parents=True, exist_ok=True)
+ json.dump(out, open(args.out, "w"), indent=1)
+ for o in out:
+ print(f"{o['film']:45s} {o['covered']:3d}/{o['cast_total']:3d} ({o['coverage_pct']}%)",
+ file=sys.stderr)
+
+
+if __name__ == "__main__":
+ main()
diff --git a/scripts/docs/run_holdout_all_models.py b/scripts/docs/run_holdout_all_models.py
new file mode 100644
index 0000000..af98817
--- /dev/null
+++ b/scripts/docs/run_holdout_all_models.py
@@ -0,0 +1,107 @@
+#!/usr/bin/env python3
+"""
+run_holdout_all_models.py — replay each model's own tuned full_exp config
+against the 5 held-out films, score with second_score.py, and dump a combined
+JSON. r50 is excluded (see docs/model-bakeoff.md: dropped from the detailed
+comparison, kept only in the calibration-curve chart).
+
+This fills a real gap: the shipped report claimed "nothing in held-out
+validation contradicts the model choice" without ever running mbf/r18 on the
+held-out films — only LVFace had been checked.
+
+Usage: python3 scripts/docs/run_holdout_all_models.py --out docs_data/holdout_all_models.json
+"""
+import argparse
+import json
+import subprocess
+import sys
+from pathlib import Path
+
+REPO = Path(__file__).resolve().parent.parent.parent
+sys.path.insert(0, str(REPO / "scripts" / "optimizer"))
+sys.path.insert(0, str(REPO / "scripts" / "validation"))
+from second_score import score_seconds # noqa: E402
+from sample_eval import load_gallery_keys # noqa: E402
+
+MODELS = ["LVFace-B_Glint360K", "arcface_w600k_mbf", "arcface_r18"]
+
+HELDOUT = [
+ {"name": "Benny & Joon", "slug": "Benny___Joon",
+ "xray": "experiments/xray/scene_level_movie_data_XRay_US/xrays/4808_Benny__Joon"},
+ {"name": "Downton Abbey: A New Era", "slug": "Downton_Abbey__A_New_Era",
+ "xray": "experiments/xray/scene_level_movie_data_XRay_US/xrays/19_Downton_Abbey_A_New_Era"},
+ {"name": "Lovelace", "slug": "Lovelace",
+ "xray": "experiments/xray/scene_level_movie_data_XRay_US/xrays/4108_Lovelace"},
+ {"name": "The Many Saints of Newark", "slug": "The_Many_Saints_of_Newark",
+ "xray": "experiments/xray/scene_level_movie_data_XRay_US/xrays/900_The_Many_Saints_Of_Newark"},
+ {"name": "Valerian and the City of a Thousand Planets",
+ "slug": "Valerian_and_the_City_of_a_Thousand_Plan",
+ "xray": "experiments/xray/scene_level_movie_data_XRay_US/xrays/5312_Valerian_and_the_City_of_a_Thousand_Planets"},
+]
+
+TRAINING = [
+ {"name": "Café Society", "slug": "Café_Society",
+ "xray": "experiments/xray/scene_level_movie_data_XRay_US/xrays/225_Cafe_Society"},
+ {"name": "Lord of War", "slug": "Lord_of_War",
+ "xray": "experiments/xray/scene_level_movie_data_XRay_US/xrays/2474_Lord_of_War"},
+ {"name": "Scarface", "slug": "Scarface",
+ "xray": "experiments/xray/scene_level_movie_data_XRay_US/xrays/197_Scarface"},
+ {"name": "Sound of Metal", "slug": "Sound_of_Metal",
+ "xray": "experiments/xray/scene_level_movie_data_XRay_US/xrays/6278_Sound_of_Metal"},
+]
+
+
+def main():
+ p = argparse.ArgumentParser()
+ p.add_argument("--out", required=True)
+ p.add_argument("--work-dir", default="/tmp/holdout_all_models")
+ p.add_argument("--films", choices=["heldout", "training"], default="heldout")
+ args = p.parse_args()
+
+ work = Path(args.work_dir)
+ work.mkdir(parents=True, exist_ok=True)
+ film_set = HELDOUT if args.films == "heldout" else TRAINING
+
+ results = {}
+ for model in MODELS:
+ cfg = json.load(open(REPO / f"experiments/results/rep4_best_{model}_full_exp.json"))["best"]["config"]
+ gallery = REPO / f"experiments/galleries/gallery_{model}.h5"
+ results[model] = {"config": cfg, "films": {}}
+
+ for film in film_set:
+ dump = REPO / f"experiments/dumps/{model}/dump_{film['slug']}.h5"
+ if not dump.exists():
+ print(f"SKIP {model}/{film['slug']}: no dump", file=sys.stderr)
+ continue
+ pred_path = work / f"pred_{model}_{film['slug']}.json"
+ cmd = [
+ "python3", "scripts/optimizer/replay.py",
+ "--dump", str(dump), "--gallery", str(gallery),
+ "--out", str(pred_path),
+ "--prob-threshold", str(cfg["prob_threshold"]),
+ "--anneal-sec", str(cfg["anneal_sec"]),
+ "--extinction-sec", str(cfg["extinction_sec"]),
+ "--expand-gallery",
+ ]
+ print(f"RUN {model}/{film['slug']}...", file=sys.stderr)
+ r = subprocess.run(cmd, cwd=REPO, capture_output=True, text=True, timeout=120)
+ if r.returncode != 0:
+ print(f"FAIL {model}/{film['slug']}: {r.stderr[-800:]}", file=sys.stderr)
+ results[model]["films"][film["slug"]] = {"error": r.stderr[-500:]}
+ continue
+
+ gk = load_gallery_keys(str(gallery))
+ pred_json = json.loads(pred_path.read_text())
+ m = score_seconds(pred_json, str(REPO / film["xray"]), gk)
+ results[model]["films"][film["slug"]] = {"name": film["name"], **m}
+ print(f" -> F1={m['f1']*100:.1f}% P={m['precision']*100:.1f}% "
+ f"R={m['recall']*100:.1f}% misid={m['FPI_misid']}", file=sys.stderr)
+
+ Path(args.out).parent.mkdir(parents=True, exist_ok=True)
+ with open(args.out, "w") as f:
+ json.dump(results, f, indent=1)
+ print(f"wrote {args.out}", file=sys.stderr)
+
+
+if __name__ == "__main__":
+ main()
diff --git a/scripts/optimizer/optimize.py b/scripts/optimizer/optimize.py
index 66413a2..e0a5bc0 100644
--- a/scripts/optimizer/optimize.py
+++ b/scripts/optimizer/optimize.py
@@ -106,8 +106,12 @@ def evaluate(cfg, films, build_dir, step=None):
"""Objective = MACRO-mean over films of each film's duration-weighted per-scene F1.
Each film's replay runs in a subprocess (timeout-guarded) to survive the
- intermittent ROCm teardown deadlock. A film whose replay times out is dropped
- from the average rather than hanging the whole sweep.
+ intermittent ROCm teardown deadlock. If ANY film's replay times out, this
+ evaluation is scored f1=0.0 (see below) rather than averaging over the
+ survivors — a partial-coverage eval must never look better than a complete
+ one, or DE will converge onto configs that make the hardest film time out.
+ (An earlier version averaged over survivors, which silently rewarded
+ truncation; the rep4 `mbf_full_noexp` winner was one such corrupted eval.)
UNIFORM PER-SECOND scoring (second_score.py): every second of the film is sampled;
GT(t) = the cast of the X-Ray scene containing t, Pred(t) = actors whose presence
@@ -135,9 +139,21 @@ def evaluate(cfg, films, build_dir, step=None):
with ThreadPoolExecutor(max_workers=REPLAY_WORKERS) as ex:
per_film = [m for m in ex.map(_one, films) if m is not None]
n = len(per_film)
- if not n:
- return {"precision": 0.0, "recall": 0.0, "f1": 0.0, "agreement": 0.0,
- "TPI": 0, "FPI": 0, "FPI_misid": 0, "FN": 0}
+ n_expected = len(films)
+ # Incomplete coverage (a replay timed out) is scored as a failure, not
+ # averaged over survivors: dropping the hardest film would otherwise inflate
+ # the score and let DE reward exactly the configs that cause timeouts. We
+ # still record the real survivor counts so a truncated eval is diagnosable
+ # in the trajectory (f1=0.0, films_scored < films_expected).
+ if n < n_expected:
+ agg = {"precision": 0.0, "recall": 0.0, "f1": 0.0, "agreement": 0.0,
+ "TPI": sum(m["TPI"] for m in per_film),
+ "FPI": sum(m["FPI"] for m in per_film),
+ "FPI_misid": sum(m["FPI_misid"] for m in per_film),
+ "FN": sum(m["FN"] for m in per_film)}
+ agg["films_scored"] = n
+ agg["films_expected"] = n_expected
+ return agg
return {"precision": sum(m["precision"] for m in per_film) / n,
"recall": sum(m["recall"] for m in per_film) / n,
"f1": sum(m["f1"] for m in per_film) / n,
@@ -145,7 +161,9 @@ def evaluate(cfg, films, build_dir, step=None):
"TPI": sum(m["TPI"] for m in per_film),
"FPI": sum(m["FPI"] for m in per_film),
"FPI_misid": sum(m["FPI_misid"] for m in per_film),
- "FN": sum(m["FN"] for m in per_film)}
+ "FN": sum(m["FN"] for m in per_film),
+ "films_scored": n,
+ "films_expected": n_expected}
def main():