diff --git a/.gitignore b/.gitignore index d0015ad..ac33e45 100644 --- a/.gitignore +++ b/.gitignore @@ -14,6 +14,11 @@ compile_commands.json *.so *.dylib *.json +# Exception: small, curated result summaries backing specific numbers quoted +# in docs/ (cross-model held-out scores, per-film training breakdown, gallery +# coverage). Regenerate with scripts/docs/run_holdout_all_models.py and +# scripts/docs/gallery_coverage_per_film.py. +!docs_data/*.json # Video files *.mp4 *.mkv diff --git a/docs/best-model.md b/docs/best-model.md index 8bd1c9d..5f433c2 100644 --- a/docs/best-model.md +++ b/docs/best-model.md @@ -1,70 +1,103 @@ # Which embedding model is best? -Four candidates went into the bake-off: three ArcFace variants (w600k-R50, -R18, w600k-MBF) and LVFace-B (Glint360K), a Vision-Transformer embedder that's -a drop-in replacement for ArcFace's `[N,3,112,112]` input / 512-d output. The -open question: is LVFace (455MB) actually better, or just the biggest? +Three ArcFace variants (w600k-R50, R18, w600k-MBF) and LVFace-B (Glint360K, +455MB) were compared. r50 is excluded from the training/held-out comparison +below; its gallery has roughly 30% fewer reference images per actor than the +other three on the identical source photos, which confounds a direct score +comparison (see [the full experiment log](model-bakeoff.md) for detail). It +remains in the calibration comparison, which does not depend on the gallery +image count. ## First signal: calibration curves Each gallery carries a fitted Platt sigmoid `P(match | cosine similarity) = -σ(a·sim + b)`, embedded directly in the gallery's HDF5 file -([`src/gallery/gallery_calibration.hpp`](https://REPOLINK/src/gallery/gallery_calibration.hpp)). This is a property of the embedding -space alone — computed from intra/inter-actor reference-image pairs, no -tracking or scene logic involved — so it's a clean first read on discriminative -power before running a single benchmark. +σ(a·sim + b)`, stored directly in the gallery HDF5 +([`src/gallery/gallery_calibration.hpp`](https://REPOLINK/src/gallery/gallery_calibration.hpp)). +This is a property of the embedding space alone, computed from intra- and +inter-actor reference-image pairs with no tracking or scene logic involved, +so it is a clean first read on discriminative power before running a +benchmark. ![Calibrated P(match|similarity) for all four models](assets/images/calibration_curves.png) -| model | `a` (steepness) | boundary at P=0.5 | +| model | a (steepness) | boundary at P=0.5 | |---|---|---| -| **LVFace-B Glint360K** | **17.7** | **sim 0.228** | +| LVFace-B Glint360K | 17.7 | sim 0.228 | | ArcFace w600k-MBF | 16.2 | sim 0.267 | | ArcFace w600k-R50 | 15.4 | sim 0.301 | | ArcFace R18 | 15.3 | sim 0.309 | -LVFace has both the steepest transition and the lowest decision boundary — it -separates same-actor from different-actor reference pairs more confidently, at -a *lower* similarity threshold, than any ArcFace variant. That's a genuine -head start before the tracking/scoring pipeline is even involved. +LVFace has both the steepest transition and the lowest decision boundary, +separating same-actor from different-actor reference pairs more confidently +at a lower similarity than any ArcFace variant. -## Second signal: F1 on the actual benchmark +## Second signal: held-out F1 -Best full-gallery (no cast-restriction) result per model, from the 16-combo -bake-off matrix ([full experiment log](model-bakeoff.md)): +Each model's own tuned `full_exp` config, replayed against the 5 films the +optimizer never saw and scored the same way: + +| film | LVFace F1 | mbf F1 | r18 F1 | +|---|---|---|---| +| Benny & Joon | 83.0% | 78.5% | 77.1% | +| Lovelace | 77.5% | 73.7% | 72.2% | +| Valerian and the City of a Thousand Planets | 74.1% | 70.2% | 71.0% | +| Downton Abbey: A New Era | 56.2% | 55.0% | 53.0% | +| The Many Saints of Newark | 46.3% | 44.5% | 42.1% | +| **macro average** | **67.4%** | **64.4%** | **63.1%** | + +LVFace scores highest on all 5 held-out films; the ranking never flips +between models. Total misID count across the 5 films: LVFace 1032, mbf +2197, r18 1224. LVFace has less than half mbf's misID total and still +scores higher on every film. + +Held-out results are stronger evidence than training results, because +training numbers can reflect what the optimizer was tuned to fit rather +than general performance. On training data, the ordering is not as clean: + +| film | LVFace F1 | mbf F1 | r18 F1 | best | +|---|---|---|---|---| +| Café Society | 68.1% | 62.2% | 60.1% | LVFace | +| Lord of War | 75.6% | 77.2% | 75.6% | mbf | +| Scarface | 71.5% | 68.6% | 64.1% | LVFace | +| Sound of Metal | 78.8% | 76.5% | 71.6% | LVFace | + +mbf beats LVFace on Lord of War (77.2% vs 75.6%), the only film in either +table where LVFace does not score highest. LVFace's training-set macro +average (75.3%, see [the full experiment log](model-bakeoff.md)) is not a +uniform win across every film it contributes to; the held-out result, where +LVFace wins all 5 films outright, is the stronger claim. + +This reverses an earlier, superseded benchmarking pass that used a +scene-union metric and found the three models statistically +indistinguishable (around 85% each), concluding LVFace was not worth its +size. That metric masked out-of-cast false positives behind a +gallery-intersect-cast recall filter; the per-second metric used here does +not. + +## Full training-matrix picture + +![All 12 combos ranked by training-set F1](assets/images/rep4_matrix_f1.png) + +Best full-gallery combo per model (all three are `full_exp`), from the +training matrix in [the full experiment log](model-bakeoff.md): | model | F1 | P | R | misID | |---|---|---|---|---| -| **LVFace-B Glint360K** | **75.3%** | 89.7% | **65.4%** | 232 | -| ArcFace w600k-MBF | 74.2% | 87.4% | 64.4% | 57 | +| LVFace-B Glint360K | 75.3% | 89.7% | 65.4% | 232 | +| ArcFace w600k-MBF | 72.0% | 87.7% | 61.4% | 240 | | ArcFace R18 | 69.1% | 87.6% | 57.7% | 242 | -| ArcFace w600k-R50 | 68.5% | 94.0% | 54.1% | 150 | -The full 16-combo picture makes the model ordering visible at a glance — LVFace -(yellow) tops both the restricted and full columns, and R18 (green) props up -the bottom of the full-gallery ranking: +LVFace leads within both the restricted and full gallery modes, visible +directly in the chart above without reading the table. The three models' +misID counts on the full gallery are nearly identical (232/240/242); LVFace's +lead here is a precision-and-recall lead, not a misID one. -![All 16 bake-off combos ranked by training-set F1](assets/images/rep4_matrix_f1.png) +## Operational note -LVFace wins outright, with the highest recall of any full-mode combo. This -reverses an earlier conclusion from a prior (superseded) benchmarking pass -using a scene-union metric, which found the three models statistically -indistinguishable (~85% each) and concluded LVFace wasn't worth its size — that -metric hid out-of-cast false positives behind a gallery∩cast recall mask (see -[the prior optimizer round](optimizer-experiments.md)); the per-second metric -used here does not. - -Held-out validation (5 films never seen by the optimizer) confirms LVFace's -lead holds up out of sample — see the -[LVFace deep dive](lvface-deep-dive.md) for the full breakdown, including -where it fails. - -## Caveat: model choice is an operational change - -Switching the default embedder isn't just flipping a config value — the -gallery itself is model-specific (embeddings from different models aren't -comparable), so any existing gallery built against ArcFace w600k-R50 needs to -be rebuilt from source images against LVFace before the new default takes -effect. [`scripts/optimizer/reembed_gallery.py`](https://REPOLINK/scripts/optimizer/reembed_gallery.py) +Switching the default embedder is not a config change alone; the gallery +is model-specific, since embeddings from different models are not +comparable. Any existing gallery built against a different model must be +rebuilt from source images before the new default takes effect. +[`scripts/optimizer/reembed_gallery.py`](https://REPOLINK/scripts/optimizer/reembed_gallery.py) does this from a reference gallery's cached source images without re-downloading anything. diff --git a/docs/gallery-scope.md b/docs/gallery-scope.md index 5833381..62c50b7 100644 --- a/docs/gallery-scope.md +++ b/docs/gallery-scope.md @@ -1,65 +1,69 @@ -# Whole gallery vs. limited (cast-restricted) gallery +# Whole gallery vs. cast-restricted gallery -Two ways to run the matcher: **full** scores every detected face against the -entire library gallery (2418 actors across the 9-film benchmark set); **restricted** -pre-filters each film's gallery down to just its Jellyfin-credited cast (typically -~15 top-billed actors) before the matcher ever runs. +Two ways to run the matcher. Full mode scores every detected face against +the entire 2418-actor gallery. Restricted mode pre-filters each film's +gallery down to just its Jellyfin-credited cast (typically around 15 +top-billed actors) before the matcher runs. -## The result +## Result -Averaged across all 4 models and both expansion settings, on the 4 bake-off training -films: +Averaged across the 3 compared models (r50 excluded, see +[the full experiment log](model-bakeoff.md)) and both expansion settings, on +the 4 training films: -| scope | F1 | P | R | total misID (8 evals) | +| scope | F1 | P | R | total misID | |---|---|---|---|---| -| full | 71.2% | 91.1% | 59.0% | 1073 | -| **restricted** | **74.5%** | 92.2% | **62.9%** | **329** | +| full | 71.1% | 89.6% | 59.6% | 1121 | +| restricted | 75.9% | 90.4% | 65.6% | 299 | -This is not a precision/recall trade — restriction wins on every axis at once: -**+3.3pp F1, +3.9pp recall, and less than a third the total misIDs.** Fewer +Restriction improves every metric at once, not a precision/recall trade: ++4.8pp F1, +6.0pp recall, roughly a quarter the total misIDs. Fewer candidates in the matcher's search space means fewer opportunities for a -look-alike false match (an actor who happens to share enough facial structure -with someone in the film, but isn't actually in it), and the recall gain shows -it isn't costing real detections to get there. +lookalike false match, and the recall gain shows this does not cost real +detections. -Per-model, every single model's best-scoring combo in the full 16-way matrix is -a `restricted` variant — visible directly in the ranking below (filled dots = -restricted, open = full; the filled dots cluster at the top for every color): +Every model's best-scoring combo in the training matrix uses the +restricted gallery: -![All 16 bake-off combos — filled dots (restricted) dominate the top](assets/images/rep4_matrix_f1.png) +![All combos ranked by training-set F1, filled dots are restricted](assets/images/rep4_matrix_f1.png) -See the full table in the -[bake-off experiment log](model-bakeoff.md). Two -combos hit **zero** true out-of-cast misidentifications: -`arcface_w600k_mbf_restricted_exp` (F1 76.5%) and, in full mode, -`LVFace-B_Glint360K_full_noexp` (F1 72.4%) — restriction isn't the only way to -reach misid=0, but it's the more reliable one. +See [the full experiment log](model-bakeoff.md) for the complete table. One +combo reaches zero true out-of-cast misidentifications, +`arcface_w600k_mbf_restricted_exp` (F1 76.2%), and it is a restricted one, +consistent with restriction, not expansion, being what suppresses cross-film +confusions. -## Why this isn't the shipped default +The restriction effect (+4.8pp averaged across models) is larger than the +model-choice effect: LVFace beats r18 by 6.2pp in full mode but beats mbf by +3.3pp. Restriction is the single strongest lever in the matrix. -Cast-restriction is implemented today only as an **offline optimizer technique** +## Why this is not the shipped default + +Cast restriction is implemented today only as an offline optimizer +technique ([`scripts/optimizer/cast_restrict.py`](https://REPOLINK/scripts/optimizer/cast_restrict.py)): -it pre-builds a filtered gallery file -per film, using Jellyfin's own cast list, before the benchmark ever calls the -matcher. There's no runtime "restrict matching to this title's credited cast" -switch in the shipped application — `scene_analyze` always matches against -whatever single gallery file it's given. +it pre-builds a filtered gallery file per film using Jellyfin's cast list +before the benchmark calls the matcher. There is no runtime "restrict to +this title's credited cast" switch in the shipped application; +`scene_analyze` always matches against whatever single gallery file it is +given. -Building that as a real feature would need, at minimum: +Building this as a real feature requires: -- A live Jellyfin cast lookup at analysis time (the title is already known — - [`scripts/run_from_jellyfin.py`](https://REPOLINK/scripts/run_from_jellyfin.py) - already does this same lookup for its own - `filter_gallery`-based restriction path, just not wired into `scene_analyze` - itself as a first-class option). -- A decision on the *fallback*: what happens to a real, uncredited cameo - (see the Germar Terrell Gardner case in the LVFace deep-dive) if the gallery - never includes them at all? -- Regenerating the restricted-gallery cache whenever the title's Jellyfin cast - list changes. +- A live Jellyfin cast lookup at analysis time. The title is already known, + and [`scripts/run_from_jellyfin.py`](https://REPOLINK/scripts/run_from_jellyfin.py) + already performs this lookup for its own `filter_gallery`-based + restriction path; it is not wired into `scene_analyze` as a first-class + option. +- A decision on the fallback case: what happens to a real, uncredited + cameo (see the Germar Terrell Gardner and Talia Balsam cases in the + [LVFace deep dive](lvface-deep-dive.md#where-lvface-beat-x-ray)) if the + restricted gallery never includes them at all. +- Regenerating the restricted-gallery cache whenever a title's Jellyfin + cast list changes. -This is why the shipped [`src/config.hpp`](https://REPOLINK/src/config.hpp) -defaults use the `full`-mode winner -(`LVFace-B_Glint360K_full_exp`, F1 75.3% training / 67.4% held-out macro) rather -than the higher-scoring `restricted_exp` (78.3%) — the 78.3% number describes a -capability the app doesn't have yet, not what actually ships. +The shipped [`src/config.hpp`](https://REPOLINK/src/config.hpp) defaults use +the full-mode winner (`LVFace-B_Glint360K_full_exp`, F1 75.3% training, +67.4% held-out macro) rather than the higher-scoring `restricted_exp` +(78.3%), because 78.3% describes a capability the application does not +have yet. diff --git a/docs/index.md b/docs/index.md index ef0c04a..54437c7 100644 --- a/docs/index.md +++ b/docs/index.md @@ -1,27 +1,32 @@ # scene-actor-extraction -A face-recognition pipeline that finds when each actor appears on screen in a -film or TV episode — built on [KPN++](https://gitea.tourolle.paris/dtourolle/KPN) -(a C++20 Kahn Process Network library) for the detect → track → match → scene -pipeline, with a Jellyfin-integrated gallery and an X-Ray-validated optimizer. +A face-recognition pipeline that finds when each actor appears on screen in +a film or TV episode, built on [KPN++](https://gitea.tourolle.paris/dtourolle/KPN) +(a C++20 Kahn Process Network library) for the detect, track, match, and +scene pipeline, with a Jellyfin-integrated gallery and an X-Ray-validated +optimizer. -This is a perfect X-Ray second, on a film the optimizer never saw: +This is a correctly scored second from a held-out film, one the optimizer +never saw during tuning: ![A perfect X-Ray second: three faces named at 100%, two more correctly carried off-screen](assets/images/lovelace_perfect_second.jpg) -Every visible face named at 100% — Chris Noth, Hank Azaria, Bobby Cannavale — -the background extra honestly left unnamed, and the two credited cast without -a visible face correctly carried as present off-screen by the tracker's -presence windows. That's the pipeline exactly reproducing Amazon X-Ray's -record for this second. +Every visible face is named at 100% confidence (Chris Noth, Hank Azaria, +Bobby Cannavale), the background extra is correctly left unnamed, and the +two credited cast members without a visible face are correctly reported +present but not visible. This matches Amazon X-Ray's own record for this +second exactly. -It doesn't always go like that: the hardest held-out film scores 46% F1, and -the report is honest about *why* — one tunable trade (extinction bridging at -hard cuts), one structural ceiling (X-Ray credits people whose faces never -appear), and a few cases where the pipeline is right and X-Ray is wrong. The -evidence for all of it is in the pages below. +Results are not uniform across films. The hardest held-out film scores 46% +F1. This report documents why: one tunable trade (extinction bridging at +hard cuts), one structural limit (X-Ray credits people whose faces never +appear on screen), and a small number of cases where the pipeline is +correct and X-Ray's ground truth is not. Read +[how we score against X-Ray](methodology.md) first. X-Ray's ground truth is +scene-level; the pipeline's output is per-second. That difference shapes +every finding below. -## Start here — four questions this bake-off answers +## Findings
@@ -29,55 +34,51 @@ evidence for all of it is in the pages below. --- - Calibration curves first (discriminative power, independent of any - threshold), then F1 on the actual benchmark. LVFace-B Glint360K wins - both. + Calibration curves first, independent of any threshold, then held-out + F1 across three models. LVFace-B Glint360K wins both, and wins on every + held-out film. - :material-filter:{ .lg .middle } **[Whole vs. cast-restricted gallery](gallery-scope.md)** --- - Restricting the matcher to a film's credited cast is a clean win on - every axis (+3.3pp F1, less than a third the misIDs) — but isn't a - shipped runtime feature yet. + Restricting the matcher to a film's credited cast improves F1, + recall, and misID rate at once, but is not a shipped runtime feature + yet. - :material-account-convert:{ .lg .middle } **[Does pose expansion help?](pose-expansion.md)** --- - A convincing training-set effect that didn't reproduce on 5 held-out - films once two methodology bugs were caught and fixed. An honest null - result, not a forced narrative. + A training-set effect that did not reproduce on 5 held-out films once + two methodology bugs in the comparison harness were found and fixed. - :material-magnify-expand:{ .lg .middle } **[Deep dive: LVFace-B Glint360K](lvface-deep-dive.md)** --- - The held-out generalization gap, how the error budget decomposes - (extinction bridging at hard cuts, X-Ray's scene-membership vs. - on-screen-face ceiling), and the frames where the pipeline is right - and the ground truth is wrong. + The held-out generalization gap, the two mechanisms behind its errors, + and every distinct case where it names someone outside the film's + credited cast.
-## The full technical log +## Full experiment log -- **[Model bake-off + threshold re-tune](model-bakeoff.md)** — - the complete experiment log behind the four pages above: the ROCm teardown - deadlock root cause and fix, DE concurrency tuning, the full 16-combo - results table, and every caveat. This is where the shipped - [`src/config.hpp`](https://REPOLINK/src/config.hpp) defaults come from. -- **[Optimizer experiments (prior round)](optimizer-experiments.md)** — the - earlier scene-union-metric tuning pass, superseded by the per-second metric - used in the bake-off but kept for the ground-truth/architecture background. -- **[Service conversion (proposal)](service-conversion.md)** — design sketch - for a native idle-GPU worker gated on screen lock, not yet built. +- **[Full experiment log](model-bakeoff.md)**: the complete log behind the + four pages above, including how replaying against cached embeddings + inside the same KPN network makes a full model and configuration + comparison practical, the full results table, and every caveat. This is + where the shipped [`src/config.hpp`](https://REPOLINK/src/config.hpp) + defaults come from. +- **[Service conversion (proposal)](service-conversion.md)**: design + sketch for a native idle-GPU worker gated on screen lock, not yet built. ## Reproducing the benchmarks -Gallery `.h5` files, embedding dumps, the X-Ray corpus, montage frame images, -and DE trajectories are not committed to this repository — they're pushed to -the Gitea package registry and pulled on demand: +Gallery `.h5` files, embedding dumps, the X-Ray corpus, montage frame +images, and DE trajectories are not committed to this repository. They are +pushed to the Gitea package registry and pulled on demand: ```bash scripts/artifacts/pull_artifacts.sh galleries @@ -86,4 +87,5 @@ scripts/artifacts/pull_artifacts.sh montage-frames ``` See [`scripts/artifacts/push_artifacts.sh`](https://REPOLINK/scripts/artifacts/push_artifacts.sh) -for the upload side (requires a `GITEA_TOKEN` with package write scope). +for the upload side, which requires a `GITEA_TOKEN` with package write +scope. diff --git a/docs/lvface-deep-dive.md b/docs/lvface-deep-dive.md index 41ccf36..1e2270f 100644 --- a/docs/lvface-deep-dive.md +++ b/docs/lvface-deep-dive.md @@ -1,52 +1,53 @@ # Deep dive: LVFace-B Glint360K -LVFace won the model bake-off (see [Which model is best?](best-model.md)) and is -the shipped default embedder. This page is the honest accounting of how it -actually performs — what a good second looks like, where the errors actually -come from, and two cases where the ground truth itself is wrong and LVFace is -right. +LVFace won the model comparison (see [Which model is best?](best-model.md)) +and is the shipped default embedder. This page reports how it performs in +detail: a baseline of correct output, the two mechanisms behind its errors, +and every distinct case where it names someone who is not in the film's +credited cast. + +Read [How we score against X-Ray](methodology.md) first. X-Ray's ground truth +is scene-level, not per-frame. A name marked correct in the Offscreen column +below is the pipeline correctly reporting scene membership, not a workaround. !!! note "How to read the frames on this page" - The top is the film frame, with a box and name on every face the pipeline - identified. The bottom panels are the per-second verdict against X-Ray: - **Onscreen** lists faces named in the frame, **Offscreen** lists cast - X-Ray marks present in the scene without a visible face — presence - carried by the tracker's windows, not by a detection. Colors are the - score: **green** = correct (TPI), - **orange** = wrong (FPI), - **blue** = missed (FN). + The top of each image is the film frame, with a box and name on every + face the pipeline matched to a real detection. The panels below are the + per-second result against X-Ray. **Onscreen** lists names attached to a + visible face this second. **Offscreen** lists names the pipeline reports + present without a currently visible face. Colors mark the verdict: + **green** correct (TPI), + **orange** wrong (FPI), + **blue** missed (FN). -## What good looks like +## Baseline: correctly scored seconds ![Wedding couple correctly identified, Downton Abbey: A New Era](assets/images/downton_wedding_couple.jpg) -Six faces on screen, all six named correctly — including Penelope Wilton at the -edge of the pews and a half-occluded Michelle Dockery — while thirteen more -cast members X-Ray marks present in the scene are correctly carried as -"Offscreen" by their presence windows. One miss in the whole frame: Maggie -Smith (blue). Score for this second: 0.86. +Six faces on screen, all six named correctly, including Penelope Wilton at +the edge of the pews and a partly occluded Michelle Dockery. Thirteen more +cast members X-Ray lists as present in the scene are correctly reported +Offscreen. One miss: Maggie Smith (blue). Score for this second: 0.86. ![19 of 20 correct in the funeral crowd](assets/images/downton_funeral_19of20.jpg) -The same film's funeral gathering: mourning dress, hats, half the faces turned. -**Nineteen of the twenty cast X-Ray lists for this scene are scored correctly** -— seven named on screen at up to 100% confidence, twelve more correctly held -as present off-screen. - -And the pipeline doesn't need the face to be *real*: +The same film's funeral scene: dark clothing, hats, half the faces turned +away. Nineteen of the twenty cast members X-Ray lists for this scene score +correct: seven named on screen at up to 100% confidence, twelve more reported +correctly as present but not visible. ![Herbie Hancock identified on an in-fiction video call](assets/images/valerian_screen_call.jpg) -That's Herbie Hancock at 98% — as a face on a *screen inside the movie*, over a -sci-fi HUD overlay, during a video call in Valerian. A face is a face, whether -it's in the room or on the bridge's comms display. +The pipeline does not require a live face. This is Herbie Hancock at 98% +confidence, identified from a face displayed on a screen inside the film, on +a video call under a science-fiction HUD overlay. ## Training vs. held-out: the generalization gap -The shipped config (`prob_threshold=0.754, anneal_sec=35.54, -extinction_sec=57.43, expand_gallery=true`) was tuned against 4 films. Scored -against the 5 films the optimizer never saw: +The shipped config (`prob_threshold=0.754`, `anneal_sec=35.54`, +`extinction_sec=57.43`, `expand_gallery=true`) was tuned on 4 films. Scored +on the 5 films the optimizer never saw: ![Held-out per-film F1 vs. the training-set fit](assets/images/holdout_f1_by_film.png) @@ -56,139 +57,241 @@ against the 5 films the optimizer never saw: | Lovelace | 77.5% | 90.3% | 67.9% | 14990 | 1085 | 58 | 7085 | | Valerian and the City of a Thousand Planets | 74.1% | 97.1% | 60.0% | 18663 | 548 | 0 | 12467 | | Downton Abbey: A New Era | 56.2% | 97.8% | 39.4% | 52027 | 1173 | 0 | 80084 | -| **The Many Saints of Newark** | **46.3%** | **54.7%** | 40.1% | 15922 | 4394 | **974** | 23791 | -| **macro average** | **67.4%** | 85.8% | 57.0% | | | | | +| The Many Saints of Newark | 46.3% | 54.7% | 40.1% | 15922 | 4394 | 974 | 23791 | +| macro average | 67.4% | 85.8% | 57.0% | | | | | -**67.4% held-out vs. 75.3% on training** — an ~8pp drop, and a **37pp spread -between the best and worst held-out film**. The config does not generalize -uniformly, and the spread traces to two mechanisms, both visible frame by -frame below. +The `P` column is misID-weighted (each out-of-film name counts 10x in the +denominator; see [methodology](methodology.md#precision-recall-and-the-misid-weighting)). +That weighting is why Many Saints reads 54.7% here despite naming mostly real, +present faces: its raw (unweighted) precision is **78.4%**, and the gap is +entirely its 974 misIDs paying the 10x penalty. The three zero-misID films +(Benny & Joon, Downton, Valerian) have identical weighted and raw precision; +Lovelace, with 58 misIDs, sits 3pp below its raw 93.3%. -## Mechanism 1: extinction bridging — usually right, wrong at hard cuts +Held-out F1 is 67.4%, against 75.3% on training, an 8pp drop. The spread +between the best and worst held-out film is 37pp. This is not unique to +LVFace: [the full experiment log](model-bakeoff.md#held-out-validation-all-3-models) +shows mbf and r18 with the same shape of spread on the same films, at a +uniformly lower level. Two mechanisms explain the spread. Both are shown +below with frame-level evidence. -The extinction window keeps an identity alive through seconds where no face is -detectable. **Most of the time this is exactly what you want**, and it's where -a lot of the TPI count comes from: +## Mechanism 1: extinction bridging + +The extinction window keeps a name reported as present for up to +`extinction_sec` after its last real detection. This is deliberate: most +gaps in face visibility are short (a turned head, an occlusion, a cut to a +reaction shot), and the window bridges them. ![Two faces on screen, six more correctly bridged](assets/images/lovelace_polygraph_bridged.jpg) -Lovelace's polygraph scene: only Eric Roberts and Amanda Seyfried have visible -faces, but X-Ray lists eight cast present — and all eight score green, the -other six correctly carried by presence windows through a scene where the -camera never shows them. A perfect second, and the extinction/anneal machinery -is *why*. +Lovelace's polygraph scene: only Eric Roberts and Amanda Seyfried have +visible faces. X-Ray lists eight cast members present. All eight score +correct; the other six are reported Offscreen through a stretch where the +camera never shows them. The extinction window is why. -The same mechanism has a failure case: a hard cut into long faceless footage. -Both Many Saints of Newark (974 misIDs) and Downton Abbey (FN=80084, the worst -recall of the five) are dominated by it — verified directly against the raw -per-frame stream and the HDF5 dump's own detection counts, not inferred from -the score alone. **This is not a malfunction**: the tracker is doing exactly -what its window is for; the footage just stops cooperating. In the debug -overlay (which draws a bridged identity's last-known bbox, unlike the shipped -output, which emits presence windows and no boxes at all) the bridged state is -visible spatially: - -![Debug overlay: bridged identities drawn at their last-known positions](assets/images/many_saints_ghost_fpi.jpg) -*Debug-overlay rendering (`dump_error_frames.py --raw`): "Jon Bernthal", "Joey -Diaz" and "Billy Magnussen" are extinction-bridged identities from the previous -shot, drawn frozen over the wall and the hanging plates. Frame -`many_saints/fpi/fpi_t03543.jpg`, `montage-frames` artifact package.* - -The cost is measurable, not just visible. Downton Abbey's hard cut into its -closing credits, plotting the dump's own per-second `face_count` (detector -output, independent of the tracker) against what the tracker reports: +The same mechanism fails at a hard cut into a long stretch with no faces at +all. Downton Abbey's recall (39.4%, the worst of the five held-out films) is +dominated by this failure. It is verified directly against the raw +per-frame stream and the dump's own detection counts, not inferred from the +score. Plotting the dump's per-second `face_count` (detector output, +independent of the tracker) against what the tracker reports, through +Downton Abbey's hard cut into its closing credits: ![Detector vs. tracker through Downton Abbey's cut to credits](assets/images/downton_ghost_timeline.png) -From the cut onward the detector sees **zero faces for nearly a minute** — and -the tracker keeps reporting the last shot's 15 identities the whole time -(verified for Hugh Bonneville: bbox `(1743.2, 0.0, 171.3, 317.8)`, unchanged to -the pixel, at every sampled second for 57+ seconds). The staircase at the right -edge is the extinction window expiring actor by actor. That plateau is -`SceneTrackerFunc::active_[actor_idx].last_bbox` +From the cut onward the detector reports zero faces for close to a minute. +The tracker continues reporting the previous shot's 15 identities for the +same span (verified for Hugh Bonneville: bbox `(1743.2, 0.0, 171.3, 317.8)`, +unchanged to the pixel, at every sampled second for 57 seconds). The +staircase at the right edge is the extinction window expiring, actor by +actor. This is `SceneTrackerFunc::active_[actor_idx].last_bbox` ([`src/nodes/scene_tracker_node.hpp`](https://REPOLINK/src/nodes/scene_tracker_node.hpp)) -re-emitted as designed: `extinction_sec=57.4` was tuned long because bridging -wins on most footage (see the polygraph frame above) — the training films just -never contained a faceless stretch long enough to show the cost side, and the -held-out set did. +re-emitted as designed. `extinction_sec=57.4` was tuned long because +bridging is correct on most footage, as in the polygraph scene above. The +training films did not contain a faceless stretch long enough to expose the +cost side; the held-out set did. -The same track-continuation machinery has one milder spatial artifact, worth -knowing when reading these frames: +The extinction window is a scoring concept, not something drawn on screen. +The shipped output is presence windows with no bounding boxes. Even the +debug overlay used for this report never draws a box for a bridged name: a +name inside its extinction window with no current detection appears only as +a name in the Offscreen column, the same as every correctly bridged name +above. + +A related, smaller effect shows up at rapid cuts: ![Two labels on one face after a shot/reverse-shot cut](assets/images/cafe_society_rapid_cut.jpg) -*Café Society (a training film), a shot/reverse-shot dialog: that is Steve -Carell wearing both his own label and Jesse Eisenberg's.* -At a rapid cut, the previous shot's track can linger for a beat at nearly the -same screen position the new face occupies — here Jesse Eisenberg's box from -the counter-shot lands on Steve Carell. Note what the score panel says, -though: both actors are green, because both *are* present in this dialog -scene per X-Ray. The spatial label is briefly wrong; the per-second presence -claim — the thing the pipeline actually ships — is right. It's the same trade -as the extinction window: track continuation smooths over cuts, and 1 fps -sampling occasionally catches the seam. +Café Society (a training film), a shot/reverse-shot dialog. The box on Steve +Carell's face carries two labels: his own, and Jesse Eisenberg's, left over +from the counter-shot a moment earlier. Both names score correct, because +both actors are present in this scene per X-Ray. The box position is +briefly wrong; the presence claim, which is what the pipeline ships, is +right. ## Mechanism 2: the face-vs-presence ceiling -Downton Abbey's recall didn't collapse because faces were misread — it -collapsed because for most of its 80084 FN-seconds there was **no face to -read**: +Downton Abbey's recall did not collapse because faces were misread. It +collapsed because for most of its 80084 false-negative seconds there was no +face to read. ![22 cast credited, nobody facing the camera](assets/images/downton_crew_fn.jpg) -A newsreel crew hauls equipment through the hall: X-Ray credits 22 cast as -present in this scene; not one face looks at the camera. Eight are still -scored green (windows bridging from adjacent shots) — the other fourteen are -blue FNs that no face-recognition pipeline could ever recover. X-Ray encodes -*scene membership*; the pipeline measures *on-screen faces*. In ensemble films -those two definitions diverge massively, and that gap — not identification -error — is most of what the FN column counts. +A newsreel crew moves equipment through the hall. X-Ray credits 22 cast +members as present in this scene. None face the camera. Eight still score +correct, carried by presence windows from adjacent shots. The other fourteen +are missed, and no face-recognition system can recover them, because there +is no face in the frame. X-Ray records scene membership; the pipeline +measures visible faces. In ensemble scenes these two quantities diverge, and +that gap accounts for most of the false-negative count. -![Presence without a detectable face, The Many Saints of Newark](assets/images/many_saints_outofcast_fpi.jpg) +## Every distinct out-of-cast name -Same ceiling from the other side: Michela De Rossi in frame but turned away, -five cast correctly bridged as offscreen (green), four blue FNs — and one -orange we'll come back to below. +Many Saints of Newark has the largest misID count of any held-out film: 974 +seconds, weighted. Rather than characterize this from a single frame, the +raw replay stream was searched directly for every name the pipeline reports +that is not in the film's credited cast. The same search was run on all 9 +films in the benchmark, one rule applied uniformly: **find the first second +each distinct out-of-cast name appears, and render that exact second.** + +Five films produce no such name anywhere in their runtime: Benny & Joon, +Café Society, Downton Abbey, Sound of Metal, Valerian. Zero out-of-cast +names across their entire length. Four films produce nine distinct names +between them, shown below in full, not a sample. + +### The Many Saints of Newark: 4 names + +![Germar Terrell Gardner, first out-of-cast name in Many Saints](assets/images/many_saints_fpi_gardner.jpg) + +Germar Terrell Gardner, t=848s, 78% confidence. A real, clearly visible +background actor. He is not in X-Ray's cast list for this film, but he is +credited in Jellyfin's independent cast metadata (see +[Where LVFace beat X-Ray](#where-lvface-beat-x-ray) below). This is a +ground-truth gap, not a model error. + +![Archie Yates, second out-of-cast name in Many Saints](assets/images/many_saints_fpi_yates.jpg) + +Archie Yates, t=2521s, 78% confidence. A real detected face, a genuine +lookalike confusion. + +![Zooey Deschanel, third out-of-cast name in Many Saints](assets/images/many_saints_fpi_deschanel.jpg) + +Zooey Deschanel, t=2819s, 99% confidence. A real detected face at a dinner +table, high-confidence lookalike confusion. + +![Talia Balsam, fourth out-of-cast name in Many Saints](assets/images/many_saints_fpi_balsam.jpg) + +Talia Balsam, t=4551s, 93% confidence. A real detected face. Talia Balsam +plays Mrs. Jarecki, a guidance counselor, in this film; she is confirmed +on screen by direct inspection of the frame. She does not appear in X-Ray's +`people.csv` for this title. This is a second ground-truth gap in the same +film, not a model error. + +Two of these four names are ground-truth gaps (Gardner, Balsam), not +misidentifications. The other two (Yates, Deschanel) are genuine embedding +errors on real faces. + +### Lord of War: 3 names + +![David Shumbris, first out-of-cast name in Lord of War](assets/images/lord_of_war_fpi_shumbris.jpg) + +David Shumbris, t=418s, 81% confidence. A real face in a dim, low-detail +shot under a train track. A genuine lookalike confusion in poor lighting. + +![Ronald Reagan, second out-of-cast name in Lord of War](assets/images/lord_of_war_fpi_reagan_photo.jpg) + +Ronald Reagan, t=1003s, 100% confidence. This is not a lookalike confusion. +The detected face is a photograph of Reagan appearing within the shot, not a +living actor. The detector and matcher both did their job correctly on the +image content in front of them; the error is that a photograph inside the +scene is not the same thing as an actor present in the scene, and the +pipeline has no way to draw that distinction from a face crop alone. + +![Lance Reddick, third out-of-cast name in Lord of War](assets/images/lord_of_war_fpi_reddick.jpg) + +Lance Reddick, t=6424s, 78% confidence. A small, distant, low-detail face at +the edge of frame. A marginal, low-confidence lookalike confusion. + +### Lovelace: 1 name + +![Chloë Sevigny, out-of-cast name in Lovelace](assets/images/lovelace_fpi_sevigny.jpg) + +Chloë Sevigny, t=2451s, 100% confidence. Two boxes are drawn on the same +face: one correctly labeled Amanda Seyfried, one incorrectly labeled Chloë +Sevigny, both at 100%. A single detection producing two competing high- +confidence identities on the same crop. + +### Scarface: 1 name + +![Kirstie Alley, out-of-cast name in Scarface](assets/images/scarface_fpi_alley.jpg) + +Kirstie Alley, t=2451s, 89% confidence. Al Pacino is correctly identified in +the foreground at 100%; a background face in the same shot is wrongly +labeled Kirstie Alley. (The t=2451s here and the Lovelace Chloë Sevigny case +above landing on the identical second is a genuine coincidence, verified from +each film's raw stream by [`first_fpi_frames.py`](https://REPOLINK/scripts/docs/first_fpi_frames.py), +not a transcription slip, two unrelated films whose *first* out-of-cast name +happens to fall at the same timestamp.) + +### Summary of the nine + +| film | name | t (s) | confidence | classification | +|---|---|---|---|---| +| Many Saints of Newark | Germar Terrell Gardner | 848 | 78% | ground-truth gap | +| Many Saints of Newark | Archie Yates | 2521 | 78% | lookalike confusion | +| Many Saints of Newark | Zooey Deschanel | 2819 | 99% | lookalike confusion | +| Many Saints of Newark | Talia Balsam | 4551 | 93% | ground-truth gap | +| Lord of War | David Shumbris | 418 | 81% | lookalike confusion | +| Lord of War | Ronald Reagan | 1003 | 100% | photo-in-frame | +| Lord of War | Lance Reddick | 6424 | 78% | lookalike confusion, marginal | +| Lovelace | Chloë Sevigny | 2451 | 100% | lookalike confusion | +| Scarface | Kirstie Alley | 2451 | 89% | lookalike confusion | + +Of nine distinct out-of-cast names across four films, two are ground-truth +gaps, one is a photograph misread as a person, and six are genuine +embedding-space confusions on real detected faces. None trace to extinction +bridging: every one of these nine is a fresh detection on a real face crop +at the second it first appears. ## Where LVFace beat X-Ray -Not every orange in these frames is actually wrong. +Not every name marked wrong is actually wrong. [`scripts/optimizer/second_score.py`](https://REPOLINK/scripts/optimizer/second_score.py) -scores strictly against X-Ray — but X-Ray itself has holes, and the pipeline -found two kinds. +scores strictly against X-Ray, and X-Ray has gaps of its own. ![LVFace correctly identifies Germar Terrell Gardner, uncredited by X-Ray](assets/images/germar_beats_xray.jpg) -Germar Terrell Gardner — a real, clean, high-confidence detection — is counted -as an out-of-cast misID because he doesn't appear in X-Ray's `people.csv` for -The Many Saints of Newark at all. But Jellyfin's independent cast metadata -*does* credit him for this exact film (cross-checked via -`experiments/manifests/jellyfin_casts.json` from the `experiment-data` artifact -package, a completely separate data source from X-Ray). That's also him in -orange in the frame above — every one of those "errors" is the pipeline being -right about a person X-Ray forgot. +Germar Terrell Gardner, the same name from the table above, does not appear +in X-Ray's `people.csv` for The Many Saints of Newark. Jellyfin's +independent cast metadata does credit him for this film (cross-checked +against `experiments/manifests/jellyfin_casts.json` from the +`experiment-data` artifact package, a data source entirely separate from +X-Ray). Talia Balsam is the same case: confirmed on screen, absent from +X-Ray's cast list for this title. ![Robert Patrick, clearly on screen, scored wrong by a ground-truth gap](assets/images/lovelace_robert_patrick_fpi.jpg) -And it isn't only uncredited bit-parts. That is **Robert Patrick** — top-billed -in Lovelace, unmistakably on screen, reading his newspaper, identified at -100% — scored orange because X-Ray's people-in-scene list for *this scene* -doesn't include him. The identification is flawless; the ground truth missed -an actor sitting in the middle of the frame. +This extends past uncredited background actors. This is Robert Patrick, +top-billed in Lovelace, clearly on screen reading a newspaper, identified at +100%. The frame is scored wrong because X-Ray's people-in-scene list for +this specific scene omits him, despite crediting him elsewhere in the film. +The identification is correct; the ground truth is missing an entry. -This doesn't mean every flagged misID is secretly correct — Many Saints' -974-count total is still overwhelmingly extinction bridging at cuts, not -uncredited cameos. But the X-Ray corpus is a convenient, large-scale ground -truth, not a perfect one, and the misID/FPI numbers in these tables carry an -irreducible noise floor from ground-truth gaps in both directions. +X-Ray is a large, convenient ground truth. It is not a complete one. The +misID and FPI counts reported throughout this document include some fixed +amount of noise from gaps in X-Ray itself, in both directions. ## Summary -LVFace is the right default: it wins the model comparison outright, it names -19 of 20 correctly across a hat-heavy funeral crowd, and it recognises a face -on a screen inside the movie. Its error budget decomposes into two understood -mechanisms — extinction bridging at hard cuts (a tunable trade, not a bug) and -the face-vs-presence ceiling baked into X-Ray's semantics — plus a nonzero -slice where the pipeline is right and the ground truth is wrong. The held-out -generalization gap (75.3% → 67.4%) is real and should be treated as the honest -expected performance, not the training-set number. +LVFace wins the model comparison on every held-out film. It correctly names +19 of 20 people in a crowded funeral scene and correctly identifies a face +displayed on a screen inside the film. Its errors resolve into two +mechanisms: extinction bridging, which is correct on most footage and fails +specifically at hard cuts into long faceless stretches, and the +face-versus-presence ceiling, where X-Ray credits scene membership for +people whose faces never appear on screen. Of the nine distinct +out-of-cast identifications found across the benchmark, two trace to gaps in +X-Ray's own cast data, one is a photograph misread as a person, and six are +genuine lookalike confusions on real faces. The held-out generalization gap, +75.3% training to 67.4% held-out, is real and should be treated as the +expected operating point, not the training-set figure. diff --git a/docs/methodology.md b/docs/methodology.md new file mode 100644 index 0000000..e646693 --- /dev/null +++ b/docs/methodology.md @@ -0,0 +1,134 @@ +# How we score against X-Ray + +Every number in this report, every F1 and misID count, comes from one +comparison. The comparison has a mismatch at its core that shapes nearly +every finding in this report: the ground truth is scene-level, the +pipeline's output is per-second, and the two do not mean the same thing. +This page documents that comparison once, so the findings pages can rely on +it without re-explaining it. + +## What Amazon X-Ray records + +X-Ray ships three tables per film: `scenes.csv` (a list of `[start, end]` +timespans), `people_in_scenes.csv` (which actors are credited in each +scene), and `people.csv` (actor identities). There is no per-frame or +per-second annotation anywhere in X-Ray. A scene might run 45 seconds, and +X-Ray records one cast list for the entire span, not "on screen from +second 12 to second 30." + +To compare this against per-second predictions, `second_score.py` expands +every scene into per-second ground truth by copying the whole scene's cast +list onto every second inside it: + +```python +for sn, (t0, t1) in spans.items(): + cast = scene_cast.get(sn, []) + for t in range(int(t0), int(t1)): + timeline[t] = cast +``` + +That is the entire mechanism. If X-Ray credits five actors to a 30-second +scene, all five count as ground truth present for all 30 seconds, including +seconds where only one of them is on screen. This is not a simplification +introduced by the pipeline; it is the only reading of X-Ray's data that is +possible, because X-Ray itself does not record anything finer-grained. + +## Why an offscreen name can be scored correct + +A name listed under Offscreen with a correct (green) label is not the +pipeline guessing or padding its score. It is the pipeline correctly +answering the question X-Ray actually asks: is this actor part of this +scene. It answers that question using a presence window (`[start, end]`, +held open across cuts by `anneal_sec` and `extinction_sec`), which matches +X-Ray's scene-level semantics more closely than a raw per-frame detection +would. + +A system that only reported "this actor is visible in this exact frame" +would score worse against X-Ray's scene-level ground truth, producing a +false negative every time the camera cuts away from a character who is +still present in the scene. Not because it is wrong about the world, but +because it would be answering a stricter, different question than the one +X-Ray's data supports. The presence-window design exists specifically to +answer X-Ray's actual question. + +## What this resolves and what it does not + +This resolves the semantic mismatch between a scene and an instant. It does +not resolve two other limitations, both discussed in the +[LVFace deep dive](lvface-deep-dive.md). + +**The face-vs-presence ceiling.** X-Ray credits scene membership regardless +of whether a face is ever visible: background crew, characters shot from +behind, voice-only presence. No amount of bridging recovers a face that +never appears on screen. This is a hard ceiling on recall, not a defect. + +**Extinction bridging can overshoot.** The same presence-window mechanism +that correctly answers "still in this scene" during a normal cut can also +bridge across a scene boundary it has no way to detect. A hard cut into a +different scene with no faces, such as closing credits, carries the +previous scene's identities forward until the window expires. This is the +mechanism behind Downton Abbey's recall collapse, documented in the deep +dive. + +## Precision, recall, and the misID weighting + +Per sampled second `t`: + +**TPI** (true positive instances): actors both X-Ray and the pipeline agree +are present. + +**FPI** (false positive instances): actors the pipeline reports that are +not in X-Ray's cast for this second. Split into two categories: + +- **FPI_incast**: the actor is in the film's cast, just not credited to + this particular scene. A timing or boundary slip. +- **FPI_misid**: the actor is not in the film's cast at all. A genuine + wrong-identity error, weighted 10x in the precision objective, because + naming someone who is not even in the film is a categorically worse + error than a few seconds of scene-boundary slop. + +!!! note "Every headline `P` and `F1` is misID-weighted" + + The precision reported throughout this report, and therefore the F1 + derived from it, puts each `FPI_misid` into the denominator **10 times** + (`precision = TPI / (TPI + FPI_incast + 10·FPI_misid)`, + [`second_score.py`](https://REPOLINK/scripts/optimizer/second_score.py)). + This is deliberate: the whole point is to punish naming an out-of-film + actor far harder than a scene-boundary slip. But it means the `P` column + is not raw precision, and a misID-heavy film's `P` is depressed + super-linearly. `second_score.py` also emits an unweighted `precision_raw` + (always ≥ the weighted `P`); where the gap matters, The Many Saints of + Newark, weighted `P` 54.7% vs. raw 78.4%, the [LVFace deep dive](lvface-deep-dive.md) + reports both. When comparing `P` across films, remember you are comparing a + quantity that penalizes misIDs, not just a hit rate. + +**FN** (false negatives): actors X-Ray lists that the pipeline never +reports, counted only for actors who have a gallery reference embedding. +Across the 9-film benchmark, coverage of X-Ray's credited cast ranges from +20% to 79% by film (see +[the full experiment log](model-bakeoff.md#gallery-coverage-per-film)); an +actor with no reference photo can never be recognized regardless of model +quality, and counting them as a miss would penalize gallery coverage, not +recognition accuracy. + +Two further numbers are reported alongside F1: + +**agreement_rate**: mean per-second Jaccard overlap +(`|Pred ∩ GT| / |Pred ∪ GT|`), partial credit. Naming 2 of 3 present actors +scores 2/3, not 0. + +**exact_match_rate**: the fraction of sampled seconds where the pipeline's +named set exactly equals X-Ray's, no partial credit. Far harsher, and +dominated by recall, since any single missed actor zeroes that second. + +## Reproduce + +```bash +python3 scripts/optimizer/second_score.py \ + --pred pred.json --xray experiments/xray/.../ \ + --gallery experiments/galleries/gallery_LVFace-B_Glint360K.h5 +``` + +See also [the full experiment log](model-bakeoff.md) for how `pred.json` is +produced, and the [LVFace deep dive](lvface-deep-dive.md) for what these +mechanisms look like frame by frame. diff --git a/docs/model-bakeoff.md b/docs/model-bakeoff.md index 633c794..5b38d27 100644 --- a/docs/model-bakeoff.md +++ b/docs/model-bakeoff.md @@ -1,425 +1,322 @@ -# Model bake-off + threshold re-tune — experiment log (2026-07-18/19) +# Full experiment log -Follow-on to [the prior optimizer round](optimizer-experiments.md), which used -an older, since-superseded scene-union metric. This round uses the **per-second** metric -([`scripts/optimizer/second_score.py`](https://REPOLINK/scripts/optimizer/second_score.py)) -and answers three questions in one 16-run matrix: which embedding model is best, does cast-restriction cut misIDs, and does -per-film gallery expansion help. +This page reports how the pipeline performs across three questions: which +embedding model is best, whether restricting the gallery to a film's +credited cast helps, and whether promoting confidently identified poses into +a per-film gallery annex helps. It also documents the replay architecture +that made testing all three questions in one pass practical, and every +caveat needed to trust the numbers. -(This was the fourth optimizer campaign against the X-Ray corpus, so its on-disk -artifacts carry an internal `rep4_` prefix — `experiments/results/rep4_best_*.json`, -`experiments/trajectories/rep4_*.jsonl`, and the manifests referenced below. The -earlier campaigns used the superseded scene-union metric and were discarded.) +Read [How we score against X-Ray](methodology.md) first for what F1, +precision, recall, and misID mean in this report. All numbers below use the +per-second metric +([`scripts/optimizer/second_score.py`](https://REPOLINK/scripts/optimizer/second_score.py)). -## Why this experiment, and what it actually delivered +r50 (ArcFace w600k-R50) is excluded from the detailed comparison below. Its +gallery was built with roughly 30% fewer reference images per actor than the +other three models on the identical source photos (10808 vs 15055 total +embeddings across the same 2418 actors), which confounds any direct +comparison of its scores against the others. It remains in the +[calibration curve comparison](best-model.md#first-signal-calibration-curves), +which does not depend on the training benchmark. -Four goals going in, and an honest read on each after held-out validation (see -below): +## Why replay makes this affordable -1. **Find the best default parameters to ship.** Partially delivered. The DE optimum - generalizes *unevenly* — strong on 3 of 5 held-out films, badly broken on 2 (one - with a 974-count misID blowup). The tuned values are shipped anyway (see - Caveats) because they still beat the old defaults on average, but this is not a - settled, film-agnostic optimum. -2. **Find the best default model.** Delivered with more confidence. LVFace beat - r50/r18/mbf across all 4 training combos, and nothing in held-out validation - contradicts the model choice specifically — the held-out failures trace to - `extinction_sec`/threshold interactions and gallery coverage, not the embedder. -3. **Provide insight into how the application works.** The strongest, most durable - output. Found and fixed a real teardown deadlock bug (100% reproducible, not the - assumed rare GPU flake), established a real concurrency ceiling (8 parallel - replays, not more), and found a real parameter interaction (a strict - `prob_threshold` "earns" a longer extinction window before it starts hurting). -4. **Demonstrate limitations.** Delivered, and reinforced hard by held-out - validation — see the "Held-out validation" section below for concrete examples, - including a screenshot of the matcher naming 15 actors, none correctly, on a - completely blank title card. +Decoding video and running face detection, alignment, and embedding is the +expensive part of this pipeline. Everything downstream of that (tracking, +identity matching, scene aggregation) is cheap. KPN++'s node/network +structure means those two stages are separate components connected by +typed channels, so the expensive stage can run once per film, cache its +output, and the cheap stage can be re-run against that cache as many times +as needed with different Config values. -## TL;DR — what changed in `src/config.hpp` +`scene_analyze --dump-embeddings out.h5` runs the expensive half once per +film and writes per-frame face detections and embeddings to HDF5 +([`scripts/optimizer/SCHEMA.md`](https://REPOLINK/scripts/optimizer/SCHEMA.md)). +[`scripts/optimizer/replay.py`](https://REPOLINK/scripts/optimizer/replay.py) +then re-assembles the real C++ `face_tracker`, `identity_matcher`, and +`scene_tracker` nodes into a Python-driven KPN network and replays a +film's cached embeddings through them, varying `prob_threshold`, +`anneal_sec`, `extinction_sec`, and `expand_gallery` freely. No GPU +inference and no video decode happen during a replay; each one completes +in seconds. This is what makes a 512-evaluation differential-evolution +search per model, per gallery mode, per expansion setting, tractable, and +what made the full held-out validation across three models in this report +possible in one session rather than requiring three full re-encodes of the +benchmark set. -| knob | old default | new default | why | -| ---- | ----------- | ----------- | --- | -| `arcface_model` | `arcface_w600k_r50.onnx` | **`LVFace-B_Glint360K.onnx`** | Best F1 in the full-gallery bake-off (75.3% vs r50's 68.5%). LVFace was worth its size. | -| `prob_threshold` | 0.76 | **0.754** | Re-tuned for LVFace + per-second metric. | -| `extinction_sec` | 1.5 | **57.4** | Reverses the earlier "short is better" finding — see below. | -| `anneal_sec` | 10.0 | **35.5** | Same reversal; previously thought insensitive. | -| `expand_gallery` | false | **true** | Helped recall on the full (unrestricted) gallery for the winning model — opposite of the earlier assumption. | +`optimize.py` runs `differential_evolution` over this replay function as its +objective, with DE-level parallelism (multiple candidate configs evaluated +concurrently, each spawning its own replay subprocesses) on top of it. The +practical ceiling on this machine's GPU was 8 concurrent replay processes; +9 silently degraded every score to 0.0% (well-formed output, wrong numbers, +not a crash), so `optimize.py` was run at `REPLAY_WORKERS=4 DE_WORKERS=2`. -These are the **`LVFace-B_Glint360K_full_exp`** winning values, applied to -[`src/config.hpp`](https://REPOLINK/src/config.hpp) — the best result that -uses only features already live in the running app (full gallery, no cast -restriction; see below for why restricted mode isn't applied even though it scored -higher). +## Search space -## Why re-run at all +`popsize=10, maxiter=15` per combo (3 parameters, up to 512 evaluations, +usually stopping earlier on DE's convergence tolerance). +`anneal_sec`/`extinction_sec` bounds were widened from 1-30/1-15 to 1-60/1-60 +partway through the sweep. r50's 4 combos finished before the widening and +used the old, narrower bounds; this is one more reason r50 is excluded from +direct comparison here. -[The prior round](optimizer-experiments.md)'s scene-union metric hid out-of-cast false positives -behind a gallery∩cast recall mask, and the earlier 9-film benchmark was scene-level -(union over a whole X-Ray scene), not a fair per-timepoint comparison. SESSION_STATE -flagged the old scene-metric bake-off numbers (R50≈LVFace≈MBF ~85%) as superseded. -This round uses `second_score.py`: uniform per-second sampling, GT = X-Ray scene's -cast at time *t*, pred = actors whose presence window covers *t*, FPI weighted 10× -when the named actor isn't in the film's cast at all (true misID) vs. an in-cast -timing slip. FN only counts gallery-known cast (fair recall — 67% of X-Ray cast have -no reference embedding; see -[the prior round's gallery-coverage-gap analysis](optimizer-experiments.md#the-gallery-coverage-gap)). +## Training films and held-out films -## The deadlock that was blocking all of this +9 films have dumped embeddings across all 4 models. 4 were used for +optimization: -Every replay in this line of work goes through -[`scripts/optimizer/replay.py`](https://REPOLINK/scripts/optimizer/replay.py), -which runs the real C++ tracker/matcher/scene_tracker nodes inside a -Python-assembled KPN network. Before this session, every subprocess replay **timed out at 45s, 100% of -the time** — not the documented ~20-30% ROCm rocBLAS-GEMM driver flake, but a plain -logic bug: `replay.py`'s CLI called `replay(net, ..., stop=False)` to *skip* -`net.stop()` (trying to dodge the GEMM deadlock), planning to `os._exit(0)` -immediately after. But: +- Café Society (62-cast) +- Lord of War (64-cast) +- Scarface (67-cast) +- Sound of Metal (14-cast) -- `PyNode::stop()` ([`include/kpn/python/bindings.hpp`](https://KPNLINK/include/kpn/python/bindings.hpp) - in the KPN++ submodule) is the *only* - code that sets `stop_flag_ = true` before joining the node's worker thread. -- The source node's `run_loop()` has `while (!stop_flag_)` as its only exit - condition (it has no input channels, so it never sees a channel-closed signal - either). -- Skipping `stop()` meant `stop_flag_` never became true. When `replay()` returned, - its local `net` went out of scope immediately, running `~PyNetwork` → `~PyNode` → - `thread_.join()` **synchronously inside `replay()`'s own call frame** — before - `main()` ever got control back to run `os._exit(0)`. +5 were held out, never seen by any optimizer run: -Root-caused via `gdb -p -batch -ex "thread apply all bt"` on a hung process: -the main thread was stuck in `~PyNode`'s `jthread::join()`; the worker thread was in -an ordinary `time.sleep()` inside the Python source callback, waiting for a stop -signal that was never sent. The two HSA `kfd_wait_on_events` threads visible in the -same trace are normal ROCm runtime housekeeping, not evidence of a wedged GPU kernel. +- Benny & Joon +- Downton Abbey: A New Era +- Lovelace +- The Many Saints of Newark +- Valerian and the City of a Thousand Planets -**Fix:** `replay.py` now calls `replay(..., stop=True)` (the removed `stop=False` + -`os._exit` workaround was actively harmful). Verified 3/3 clean runs at ~8s each -(down from a guaranteed 45s timeout), and a full DE sweep producing real, sensible -F1/precision/recall instead of flat 0.0%. +## Gallery coverage per film -## Concurrency tuning +The gallery has reference embeddings for 2418 actors, but coverage of any +given film's credited cast varies widely. This was previously reported as +one flat number (67% of X-Ray cast lacking a reference embedding, averaged +across the whole benchmark); the per-film breakdown is: -With the deadlock fixed, -[`scripts/optimizer/optimize.py`](https://REPOLINK/scripts/optimizer/optimize.py) -was extended with DE-level parallelism — -`differential_evolution(..., workers=ThreadPoolExecutor.map)` — so multiple -population candidates evaluate concurrently, each spawning its own per-film replay -subprocesses (`REPLAY_WORKERS`). Total concurrent GPU replay processes ≈ -`DE_WORKERS × REPLAY_WORKERS`. +| film | cast credited | in gallery | coverage | +|---|---|---|---| +| Lord of War | 64 | 13 | 20.3% | +| Scarface | 67 | 15 | 22.4% | +| The Many Saints of Newark | 48 | 13 | 27.1% | +| Café Society | 62 | 17 | 27.4% | +| Lovelace | 42 | 15 | 35.7% | +| Valerian and the City of a Thousand Planets | 36 | 13 | 36.1% | +| Benny & Joon | 23 | 12 | 52.2% | +| Downton Abbey: A New Era | 36 | 22 | 61.1% | +| Sound of Metal | 14 | 11 | 78.6% | -| concurrent replays | result | -| --- | --- | -| 3 (`REPLAY_WORKERS=3`, no DE parallelism) | baseline, GPU underutilised | -| 6 (`DE_WORKERS=2 × REPLAY_WORKERS=3`) | clean, real scores, ~1 isolated timeout per run | -| 8 (`DE_WORKERS=2 × REPLAY_WORKERS=4`, 4-film manifest) | clean, real scores | -| 9 (`DE_WORKERS=3 × REPLAY_WORKERS=3`) | **broken** — every replay blew past the 45s timeout, all scores silently degraded to 0.0% | +Two training films (Lord of War, Scarface) have the worst coverage in the +set, 20-22%. Their training-set F1 numbers below are partly capped by +missing references, not purely by model quality. Downton Abbey has 61% +coverage, the second-best in the benchmark, yet the worst held-out recall +of any film (39.4%, LVFace). Its recall problem is not primarily a coverage +problem; it is the extinction-bridging failure documented in the +[LVFace deep dive](lvface-deep-dive.md#mechanism-1-extinction-bridging). +Reproduce with `scripts/docs/gallery_coverage_per_film.py`. -9 concurrent replays looks like valid output (well-formed JSON, a real number) while -actually being garbage — a dangerous failure mode, not a crash. **8 concurrent is the -practical ceiling** on this GPU (gfx1100) for this workload. The matrix ran at -`REPLAY_WORKERS=4 DE_WORKERS=2`. +## Training results, 3 models × 2 gallery modes × 2 expansion settings -## Training films and validation set +Ranked by F1. misid = FPI_misid, the count of true wrong-actor +identifications (naming someone not in the film's cast at all), distinct +from FPI, which also includes in-cast timing slips. -9 films total have dumped embeddings across all 4 models. 4 were used for -optimization, leaving 5 held out for validation: - -- **Lord of War** (64-cast, "clean") -- **Scarface** (67-cast, "ensemble/lookalike") -- **Sound of Metal** (14-cast, "high gallery-coverage") -- **Café Society** (62-cast, added this round — similar ensemble size to Scarface but - different genre/lighting; picked to add diversity, not genre-overlap, over - Downton Abbey or The Many Saints of Newark) - -Held out: Benny & Joon, Downton Abbey: A New Era, Lovelace, The Many Saints of -Newark, Valerian and the City of a Thousand Planets. The training-set numbers below are -training-set fit — see "Held-out validation" further down for the real -generalization test. - -## Search space and DE settings - -`popsize=10, maxiter=15` (3 params → ≤480 evals/combo ceiling; DE's `tol` convergence -usually stops earlier). `anneal_sec`/`extinction_sec` bounds were **widened from -1–30/1–15 to 1–60/1–60 mid-run** (see below) — the 4 `arcface_w600k_r50` combos -finished before the widening and still use the old, narrower bounds, so they are -**not directly comparable** to the other 12 on those two params. Re-running r50 with -the wider bounds was deferred (diminishing-returns judgment call, not yet done). - -## Results — all 16 combos (4 models × {full, restricted} × {expand, noexp}) - -Ranked by F1. `misid` = FPI_misid, count of true wrong-actor identifications (an -actor named who isn't in the film's cast at all) — distinct from `FPI`, which -includes in-cast timing slips. +Each combo's row is its best **full-coverage** evaluation: the highest-F1 DE +evaluation in which all 4 training films replayed without a timeout (see +[Dropped-film scoring](#a-scoring-bug-worth-recording-dropped-film-evaluations) +below for why this qualifier is load-bearing and not the same as `argmax F1` +over the raw sweep). | combo | F1 | P | R | TPI | FPI | misid | FN | |---|---|---|---|---|---|---|---| -| LVFace-B_Glint360K_restricted_exp | **78.3%** | 91.0% | 68.9% | 42830 | 3782 | 60 | 19492 | +| LVFace-B_Glint360K_restricted_exp | 78.3% | 91.0% | 68.9% | 42830 | 3782 | 60 | 19492 | | LVFace-B_Glint360K_restricted_noexp | 76.7% | 91.5% | 66.2% | 41149 | 3400 | 59 | 21173 | -| arcface_w600k_mbf_restricted_exp | 76.5% | 90.7% | 66.3% | 41270 | 4431 | **0** | 21052 | +| arcface_w600k_mbf_restricted_exp | 76.2% | 90.0% | 66.2% | 64328 | 7480 | 0 | 33234 | | arcface_r18_restricted_exp | 75.5% | 87.6% | 66.5% | 41399 | 5666 | 60 | 20923 | -| **LVFace-B_Glint360K_full_exp** | **75.3%** | 89.7% | 65.4% | 47757 | 3407 | 232 | 26966 | +| LVFace-B_Glint360K_full_exp | 75.3% | 89.7% | 65.4% | 47757 | 3407 | 232 | 26966 | | arcface_w600k_mbf_restricted_noexp | 75.0% | 91.1% | 63.9% | 39752 | 3465 | 60 | 22570 | -| arcface_w600k_mbf_full_noexp | 74.2% | 87.4% | 64.4% | 12645 | 1312 | 57 | 6985 | | arcface_r18_restricted_noexp | 73.5% | 91.3% | 61.7% | 38299 | 3220 | 60 | 24023 | -| LVFace-B_Glint360K_full_noexp | 72.4% | 94.2% | 58.9% | 27077 | 1725 | **0** | 19506 | +| LVFace-B_Glint360K_full_noexp | 72.3% | 88.3% | 61.8% | 40363 | 3503 | 244 | 25850 | | arcface_w600k_mbf_full_exp | 72.0% | 87.7% | 61.4% | 39875 | 3729 | 240 | 26338 | -| arcface_w600k_r50_full_noexp † | 71.6% | 96.7% | 56.9% | 22471 | 361 | 45 | 17012 | -| arcface_w600k_r50_restricted_exp † | 71.1% | 96.5% | 56.4% | 34954 | 1119 | 15 | 27368 | -| arcface_w600k_r50_restricted_noexp † | 69.2% | 97.9% | 53.6% | 21146 | 327 | 15 | 18337 | +| arcface_w600k_mbf_full_noexp | 71.0% | 93.2% | 57.9% | 41699 | 2472 | 56 | 33024 | | arcface_r18_full_exp | 69.1% | 87.6% | 57.7% | 37342 | 3119 | 242 | 28871 | -| arcface_w600k_r50_full_exp † | 68.5% | 94.0% | 54.1% | 34982 | 903 | 150 | 31231 | | arcface_r18_full_noexp | 66.6% | 91.3% | 53.1% | 34314 | 2362 | 107 | 31899 | -† old, narrower anneal/extinction bounds (see above) — not directly comparable to -the other 12 on those two params. +![All combos ranked by training-set F1](assets/images/rep4_matrix_f1.png) -The same 16 results as a picture — the two headline effects are visible without -reading a single row: filled (restricted) dots stack the top of the ranking for -every model color, and yellow (LVFace) leads within both scopes: +The two clearest patterns: every model's best-scoring combo uses the +restricted gallery, and LVFace leads within both gallery modes. `full_exp` +(the shipped combination) is the best-scoring option that uses only +features the running application currently supports; restriction is not +wired into the application yet (see +[Whole vs. cast-restricted gallery](gallery-scope.md)). -![All 16 bake-off combos ranked by training-set F1](assets/images/rep4_matrix_f1.png) +### A scoring bug worth recording: dropped-film evaluations -## Calibration curves — discriminative power, independent of the threshold +The numbers above are corrected ones. The raw `rep4_best_*.json` files, and an +earlier version of this table, reported a different `arcface_w600k_mbf_full_noexp` +row: **74.2% F1 at TPI 12645**, a third the TPI of every sibling combo. That was +not a better config; it was an artifact of how the optimizer aggregates. -Each model's gallery carries a fitted Platt sigmoid `P(match | sim) = σ(a·sim + b)` -(embedded directly in the gallery HDF5, see -[`src/gallery/gallery_calibration.hpp`](https://REPOLINK/src/gallery/gallery_calibration.hpp)). -Plotting all four side by side shows discriminative power directly, independent of -whatever `prob_threshold` a particular run happened to use: +`optimize.py` builds each candidate's score from only the films whose replay +subprocess returned (`per_film = [m for m in ex.map(_one, films) if m is not +None]`), then **averages** F1/precision/recall and **sums** TPI/FPI/misID over +just those survivors. When a film's replay times out (the sweep ran near the +8-process concurrency ceiling, so this happened intermittently), that film +silently drops from both. A candidate whose hardest film timed out is therefore +scored on an easier subset, and differential evolution, maximizing that score, +will happily converge onto exactly such a candidate. For `mbf_full_noexp` the +reported winner was one of 7 evaluations (out of 512) whose TPI had collapsed to +a partial-film subset; its median-coverage evaluations sit around 51686 TPI. -![Calibrated P(match|similarity) for all four models](assets/images/calibration_curves.png) +The fix here was to re-derive each combo's best row from its DE trajectory +(`experiments/trajectories/rep4_*.jsonl`), keeping only evaluations within 30% of +that combo's median TPI (full 4-film coverage) before taking the best F1. This +needs no re-running, the honest best configuration was already in the sweep, +just not the one `argmax F1` selected. Three combos moved: `mbf_full_noexp` +74.2% → **71.0%**, `LVFace_full_noexp` 72.4% → **72.3%** (and its misID, 0 → 244, +was itself a dropped-film artifact), `mbf_restricted_exp` 76.5% → **76.2%**. The +shipped LVFace `full_exp` winner was unaffected, its reported evaluation already +had full coverage (TPI 47757 ≈ median). `experiment_charts.py` applies the same +`clean_best` filter, so every figure on this page matches the corrected table. +The underlying `optimize.py` aggregation is also being fixed so a dropped-film +evaluation can never be selected as a winner again. -LVFace-B has both the steepest curve (`a=17.7`, vs. 15.3–16.2 for the ArcFace -variants) and the lowest P=0.5 decision boundary (similarity 0.23 vs. 0.27–0.31) — -it separates same-actor from different-actor pairs more confidently at a lower -similarity, consistent with it winning the full-gallery F1 comparison below. -Generated by -[`scripts/docs/calibration_chart.py`](https://REPOLINK/scripts/docs/calibration_chart.py) -(requires each gallery to have been calibrated at least once — run any replay -against it first). +### Per-film training breakdown -## Two effects in isolation: gallery scope, and pose expansion +The 75.3% LVFace training figure is a macro average across 4 films, not a +uniform result: -The matrix crosses two independent variables — averaging across all 4 models -isolates each one from model choice: - -**Gallery scope (whole 2418-actor gallery vs. restricted to the film's credited -cast)** — averaged over both expansion settings and all 4 models: - -| scope | F1 | P | R | total misID (16 evals→8 each) | +| film | LVFace F1 | mbf F1 | r18 F1 | best model | |---|---|---|---|---| -| full | 71.2% | 91.1% | 59.0% | 1073 | -| **restricted** | **74.5%** | 92.2% | **62.9%** | **329** | +| Café Society | 68.1% | 62.2% | 60.1% | LVFace | +| Lord of War | 75.6% | 77.2% | 75.6% | mbf | +| Scarface | 71.5% | 68.6% | 64.1% | LVFace | +| Sound of Metal | 78.8% | 76.5% | 71.6% | LVFace | -Restriction wins outright on every axis — not a precision/recall trade, a clean -win: **+3.3pp F1, +3.9pp recall, and less than a third the total misIDs.** Fewer +LVFace does not win every training film. mbf scores higher on Lord of War +(77.2% vs 75.6%). LVFace's own training-film range is 68.1% to 78.8%, a +10.7pp spread, smaller than the 37pp spread seen on held-out films but real. +Reproduce with `scripts/docs/run_holdout_all_models.py --films training`. + +## Held-out validation, all 3 models + +The training matrix above is training-set fit. Each model's own tuned +`full_exp` config was replayed against the 5 held-out films, scored the +same way: + +| film | LVFace F1 | mbf F1 | r18 F1 | +|---|---|---|---| +| Benny & Joon | 83.0% | 78.5% | 77.1% | +| Lovelace | 77.5% | 73.7% | 72.2% | +| Valerian and the City of a Thousand Planets | 74.1% | 70.2% | 71.0% | +| Downton Abbey: A New Era | 56.2% | 55.0% | 53.0% | +| The Many Saints of Newark | 46.3% | 44.5% | 42.1% | +| **macro average** | **67.4%** | **64.4%** | **63.1%** | + +LVFace scores highest on every one of the 5 held-out films; the ranking +never flips. Total misIDs across the 5 films: LVFace 1032, mbf 2197, r18 +1224. LVFace has less than half mbf's misID count while also scoring +higher on every film. This directly confirms the model choice out of +sample; it is not inferred from the training numbers alone. See the +[LVFace deep dive](lvface-deep-dive.md) for frame-level detail on where and +why LVFace still fails on the two worst films. Reproduce with +`scripts/docs/run_holdout_all_models.py`. + +## Two effects in isolation: gallery scope and pose expansion + +Averaging across the 3 compared models (r50 excluded) isolates each variable +from model choice. + +**Gallery scope**, averaged over both expansion settings and all 3 models +(6 evaluations per row): + +| scope | F1 | P | R | total misID | +|---|---|---|---|---| +| full | 71.1% | 89.6% | 59.6% | 1121 | +| restricted | 75.9% | 90.4% | 65.6% | 299 | + +Restriction improves every metric at once. This is not a precision/recall +trade: +4.8pp F1, +6.0pp recall, and roughly a quarter the misIDs. Fewer candidates in the matcher's search space means fewer opportunities for a -look-alike false match, and (per the recall gain) doesn't cost real detections. -This is the single cleanest signal in the whole matrix — stronger than the model -choice itself — which is exactly why cast-restriction becoming a real runtime -feature (not just an optimizer trick) is the top item in Caveats below. +lookalike false match, and the recall gain shows this does not cost real +detections. Restriction is currently an offline optimizer technique, not a +runtime feature of the application; see +[Whole vs. cast-restricted gallery](gallery-scope.md) for what building it +into the application would require. -**Pose expansion (promoting a confidently-identified track's novel-pose views into -a per-film gallery annex — [`src/gallery/track_gallery.hpp`](https://REPOLINK/src/gallery/track_gallery.hpp))** -is smaller and interacts with -scope rather than acting independently: +**Pose expansion** (promoting a confidently identified track's novel-pose +views into a per-film gallery annex, +[`src/gallery/track_gallery.hpp`](https://REPOLINK/src/gallery/track_gallery.hpp)): | scope | expansion | F1 | R | misID | |---|---|---|---|---| -| full | off | 71.2% | 58.3% | 209 | -| full | **on** | 71.2% | 59.7% | **864** | -| restricted | off | 73.6% | 61.3% | 194 | -| restricted | **on** | **75.4%** | **64.5%** | 135 | +| full | off | 70.0% | 57.6% | 407 | +| full | on | 72.1% | 61.5% | 714 | +| restricted | off | 75.1% | 63.9% | 179 | +| restricted | on | 76.7% | 67.2% | 120 | -In **restricted** mode, expansion is a clean win (+1.8pp F1, +3.2pp recall, misID -actually *drops*) — the annex only ever competes against the film's own ~15-actor -cast, so a "confidently identified, new pose" view is unlikely to be mistaken for -someone else. In **full** mode, expansion buys essentially nothing on F1 (71.2% → -71.2%, recall +1.4pp) while **quadrupling misIDs** (209 → 864): a novel-pose view -promoted into the annex now competes against the whole 2418-actor gallery, so a -"confident" identity is confident against the wrong universe of candidates — the -expansion mechanism is "learning" a pose correctly, but the enlarged evidence pool -makes it easier for that learned pose to look like a plausible match for a -different actor. **Practical takeaway: gallery expansion should be paired with -cast restriction, not used on the full gallery** — the version currently shipped -as default (`full_exp`, see TL;DR) sits in the worse of these four cells for this -specific knob, even though it's the best available combo without cast-restriction -support in the app yet (see Caveats). +In restricted mode, expansion is a clean win: +1.6pp F1, +3.3pp recall, +misID drops. The annex only competes against the film's own roughly 15-actor +cast, so a new pose of a known actor is unlikely to be confused with someone +else. In full mode, expansion buys +2.1pp F1 and +3.9pp recall but at a real +cost: misID rises from 407 to 714 as the same new-pose view now competes +against the full 2418-actor gallery, where a confidently learned pose is more +likely to match the wrong person. On the full gallery it is a recall-vs-misID +trade, not a free gain. This training-set effect +did not reproduce on held-out data; see +[Does pose expansion help?](pose-expansion.md) for the full held-out test +and the two methodology bugs caught while checking it. -## What the data says +## Calibration curves -- **LVFace was worth its size.** It wins full-gallery mode outright (75.3% vs r50's - 68.5%, r18's 69.1%, mbf's 72.0%) with the highest recall of any full-mode combo — - the earlier scene-union-metric conclusion ("not worth it") doesn't survive the - better metric. -- **Cast-restriction is a consistent, broad win.** Every model's best combo is - `restricted`. It isn't just precision-safe: `arcface_w600k_mbf_restricted_exp` and - `LVFace-B_Glint360K_full_noexp` both hit **misid=0** — zero true wrong-actor - identifications. But restriction is an **offline optimizer technique, not a live - app feature** — it pre-filters each film's gallery to its Jellyfin-credited cast - before the matcher ever runs; there's no runtime "restrict to this film's cast" - switch in the app today. Implementing it for real is future work, tracked - separately from this defaults update. -- **Gallery expansion (`expand_gallery`) is mode-dependent.** It helps on - `restricted` galleries (smaller, so novel-pose promotion adds real signal) and on - LVFace's full gallery, but **hurts** r50 and mbf in full mode (compare - `arcface_w600k_r50_full_exp` 68.5% vs `full_noexp` 71.6%). Don't assume it's a free - win — model- and mode-dependent. -- **arcface_r18 (smallest/cheapest) is last across all 4 modes** — model capacity - matters here, this isn't just parameter-count padding. -- **`anneal_sec`/`extinction_sec` kept pinning at the search ceiling.** With the - original 1–30/1–15 bounds, 3 of 4 r50 combos landed at ~93-98% of the upper bound. - Widened to 1–60/1–60 mid-run (after the r50 combos had already finished) — every - subsequent combo's best config landed at ~90%+ of the *new* ceiling too (e.g. the - LVFace winner: `ann=59.2, ext=59.2`, both ~99% of 60). The likely mechanism: a - strict `prob_threshold` "earns" a long extinction/anneal window — once false - matches are rare, a long window just bridges real presence gaps (occlusion, turned - face) instead of smearing false positives into later scenes, which is what made - short windows look better under the old, laxer thresholds. **Open question, not - resolved**: does this keep climbing past 60s, or does it actually plateau there? - Decided not to chase further this round (diminishing-returns judgment call) — flag - for a future sweep if it matters. +Each gallery carries a fitted Platt sigmoid `P(match | sim) = σ(a·sim + b)`, +stored directly in the gallery HDF5 +([`src/gallery/gallery_calibration.hpp`](https://REPOLINK/src/gallery/gallery_calibration.hpp)). +This measures discriminative power independent of whatever +`prob_threshold` a given run used: -The ceiling-pinning is visible in the raw search itself. Every one of the 512 -DE evaluations for the winning combo, plotted over the -`prob_threshold` × `extinction_sec` plane: +![Calibrated P(match|similarity) for all four models](assets/images/calibration_curves.png) + +LVFace has the steepest curve (`a=17.7` vs 15.3-16.2 for the ArcFace +variants) and the lowest P=0.5 decision boundary (similarity 0.23 vs +0.27-0.31), separating same-actor from different-actor pairs more +confidently at a lower similarity than any ArcFace variant tested, +including r50. Generated by +[`scripts/docs/calibration_chart.py`](https://REPOLINK/scripts/docs/calibration_chart.py). + +## Extinction and anneal window search + +Every one of the 512 DE evaluations for the winning LVFace `full_exp` +combo, plotted over the `prob_threshold` × `extinction_sec` plane: ![DE search landscape: 512 evaluations over prob_threshold × extinction_sec](assets/images/de_search_landscape.png) -The dark band hugging the top edge *is* the finding: nearly everything scoring -well sits at `extinction_sec` ≥ 50, across a wide range of thresholds, and the -population converged into a dense cloud around the optimum (threshold ~0.70–0.80, -extinction pinned at the 60s bound). Short extinction windows (bottom half) are -uniformly pale — under a strict threshold there is simply no good configuration -down there. Generated by -[`scripts/docs/experiment_charts.py`](https://REPOLINK/scripts/docs/experiment_charts.py) -from the DE trajectories (`experiments/trajectories/*.jsonl`, part of the -`experiment-data` artifact package). +Nearly everything scoring well sits at `extinction_sec` above 50, across a +wide range of thresholds. Short extinction windows are uniformly weaker: +under a strict threshold, there is no good configuration in that region of +the search space. The optimizer converged with `anneal_sec=59.2, +extinction_sec=59.2`, about 99% of the widened 60s bound, which raises an +open question not resolved in this round: does performance keep improving +past 60s, or does it plateau there. Not chased further this pass. -## Held-out validation — the number that actually matters +## Caveats -The 16-combo matrix above is training-set fit. This is the real test: the shipped -config (`LVFace-B_Glint360K_full_exp` — `prob_threshold=0.754, anneal_sec=35.5, -extinction_sec=57.4, expand_gallery=true`) replayed against the **5 films never seen -by the optimizer** (Benny & Joon, Downton Abbey: A New Era, Lovelace, The Many -Saints of Newark, Valerian and the City of a Thousand Planets), scored the same way. - -| film | F1 | P | R | agree | TPI | FPI | misid | FN | -|---|---|---|---|---|---|---|---|---| -| Benny & Joon | 83.0% | 89.1% | 77.7% | 72.4% | 15125 | 1846 | 0 | 4337 | -| Lovelace | 77.5% | 90.3% | 67.9% | 72.1% | 14990 | 1085 | 58 | 7085 | -| Valerian and the City of a Thousand Planets | 74.1% | 97.1% | 60.0% | 58.8% | 18663 | 548 | 0 | 12467 | -| Downton Abbey: A New Era | 56.2% | 97.8% | 39.4% | 40.6% | 52027 | 1173 | 0 | 80084 | -| **The Many Saints of Newark** | **46.3%** | **54.7%** | 40.1% | 37.0% | 15922 | 4394 | **974** | 23791 | -| **macro average (5 films)** | **67.4%** | 85.8% | 57.0% | 56.2% | 116727 | 9046 | 1032 | 127764 | - -![Held-out per-film F1 vs. the training-set fit](assets/images/holdout_f1_by_film.png) - -**67.4% held out vs. 75.3% on training** — an ~8pp drop, and a much more informative -number than the training-set F1 alone: a **37pp spread between best and worst film** -(83.0% vs 46.3%). The config does not generalize uniformly. - -Two films are outright failure cases, and rendering bounding boxes + names on the -extracted frames (`replay.py --raw-out` + -[`dump_error_frames.py`](https://REPOLINK/scripts/optimizer/dump_error_frames.py)` --raw`, see -Reproduce) turned what looked like a same-scene misidentification into something -more precise and more damning: - -- **The Many Saints of Newark** (mob-family drama, picked as a training-adjacent - genre test) has **974 true misIDs** — far more than any training combo saw at any - setting. The annotated frame below shows the same mechanism as Downton Abbey, - at smaller scale: **"Jon Bernthal 100%", "Joey Diaz 100%", and "Billy Magnussen - 100%" are all frozen boxes over empty background — a blurred wall, hanging plates — - with no face in them at all.** Only one real face in frame has a box, and it - carries a *second*, colliding label ("Leslie Odom Jr." and "Michael Gandolfini" - both at high confidence on the same box) — likely two tracks whose frozen bboxes - happen to overlap. - - ![Frozen ghost boxes over background, The Many Saints of Newark](assets/images/many_saints_ghost_fpi.jpg) - *Frame `many_saints/fpi/fpi_t03543.jpg` from the `montage-frames` artifact - package (`scripts/artifacts/pull_artifacts.sh montage-frames - Many_Saints_of_Newark`).* - -- **Downton Abbey: A New Era** (large ensemble, 36-cast) has high precision (97.8%) - but recall collapses to 39.4% (FN=80084, by far the largest of the 5). Its - starkest failure happens where there is nothing to see at all: the film's hard - cut into its closing credits, where **the matcher kept reporting 15 actors — - all wrong — for nearly a minute of faceless screen.** - -Both are the same mechanism, and it can be *measured*, not just screenshotted. -Plotting the dump's own per-second `face_count` (detector output, independent -of the tracker) against the number of actors the tracker reports, through -Downton Abbey's cut to credits: - -![Detector face_count vs. tracker-reported actors through the cut to credits](assets/images/downton_ghost_timeline.png) - -From the cut onward the detector sees **zero faces** — yet the tracker holds a -perfectly flat plateau of 15 reported identities for 56 seconds, each with the -*exact same bbox, unchanged to the pixel* (verified for Hugh Bonneville: -`(1743.2, 0.0, 171.3, 317.8)` at every sampled second from 7222 through 7279+). -The staircase on the right edge is the extinction window finally expiring, -actor by actor. That plateau is `SceneTrackerFunc`'s -`active_[actor_idx].last_bbox` -([`src/nodes/scene_tracker_node.hpp`](https://REPOLINK/src/nodes/scene_tracker_node.hpp)) -being re-emitted -unchanged — **the extinction state machine working exactly as coded**, not a -bug in the logic. The film cuts from a packed group shot straight into ~40+ -seconds of blank titles/credits with zero faces, and `extinction_sec=57.4` is -comfortably long enough to bridge that entire gap without expiring, so the -tracker faithfully keeps reporting "last known position" for a cast that is no -longer on screen at all. - -This reframes the "long extinction window wins" DE-search pattern (see above): it -isn't unambiguously good. It buys recall by bridging real gaps (occlusion, turned -face) in some films, but on others — specifically, hard cuts into long faceless -footage — it manufactures a frozen-bbox ghost the tracker has no way to verify, -precisely the failure mode the *original* short-extinction-window default (`1.5s`) -was chosen to avoid. The training-set films apparently didn't have a long enough -faceless stretch after a confirmed identity to expose this; the held-out set did. - -Frames for all three films (`benny_joon`, `many_saints`, `downton_abbey` — one strong -performer, two failure cases) are under `experiments/results/holdout/frames/` -(not committed — pull per film with `scripts/artifacts/pull_artifacts.sh -montage-frames `), each -with a `manifest.json` listing the bucket (`best`/`fpi`/`fn`), timestamp, and -predicted vs. ground-truth actors for every dumped frame. Frames are annotated with -bounding boxes + name/confidence (green = identified, orange = unknown), matching -[`src/nodes/debug_renderer_node.hpp`](https://REPOLINK/src/nodes/debug_renderer_node.hpp)'s -colour convention. Generated by -`scripts/optimizer/dump_error_frames.py --raw ` (see -Reproduce). - -`dump_error_frames.py --interval-sec 600` also supports a per-N-second sweep -instead of the fixed best/fpi/fn buckets: one best (highest Jaccard) and one worst -(lowest Jaccard) frame per 10-minute window across the whole film, e.g. -`experiments/results/holdout/frames/many_saints_intervals/` (13 windows × 2 = 26 -frames for the ~2h Many Saints runtime) — a way to sample "how are we doing" evenly -across a film's runtime rather than only at its most extreme seconds. - -## Caveats / what this is not - -- **r50's 4 combos used the old, narrower search bounds** and aren't fully - comparable to the other 12 on `anneal_sec`/`extinction_sec`. -- **The applied defaults use `full_exp`, not the higher-scoring `restricted_exp`**, - because cast-restriction isn't a real runtime feature yet (see above). The - 78.3% F1 number is not what the shipped defaults will produce — 75.3% is. -- **`full_exp` is the best full-gallery combo, but not the safest.** Per the - isolated-effects analysis above, `expand_gallery=true` only cleanly pays off - when paired with cast-restriction; on the full gallery it's flat on F1 while - ~4x-ing misIDs (209→864, averaged across models). `full_noexp` scores lower - (72.4% vs 75.3% for LVFace) but with **zero** true misIDs and higher precision - (94.2% vs 89.7%). Kept `full_exp` as shipped since it's the highest-F1 option - available without cast-restriction, but this is a real F1-vs-safety trade, not - a strictly-better choice — worth revisiting if misID rate matters more than - the last few points of F1 for a given deployment. -- **Switching the default model is an operational change, not just a config tweak**: - any existing gallery built from r50 embeddings is incompatible with LVFace - embeddings and needs rebuilding. +- r50's 4 combos used the older, narrower search bounds (1-30/1-15 instead + of 1-60/1-60) and are further confounded by its thinner gallery. Excluded + from all comparisons above except calibration. +- The shipped defaults use `full_exp` (75.3% training F1), not the + higher-scoring `restricted_exp` (78.3%), because cast restriction is not + a runtime feature of the application yet. +- `expand_gallery` is mode-dependent, not a free win. Averaged across models + on the full gallery it trades misIDs for recall (see the pose-expansion + table). For LVFace specifically, though, `full_exp` beats `full_noexp` on + every axis at once (F1 75.3 vs 72.3, precision 89.7 vs 88.3, recall 65.4 vs + 61.8, misID 232 vs 244), so the shipped `full_exp` is a clean choice for + this model, not an F1-vs-safety trade. (An earlier version of this page + reported `full_noexp` at 72.4% with zero misIDs and higher precision, which + made it look like the safer option; that was the dropped-film artifact + described above, not a real property of the config.) +- Switching the default model is an operational change: any gallery built + from a different model's embeddings must be rebuilt before the new + default takes effect. ## Reproduce ```bash -# 4-film matrix, all 4 models × 2 modes × 2 expansion settings +# 4-film training matrix, all 4 models × 2 gallery modes × 2 expansion settings bash experiments/run_rep4_subprocess.sh # single combo @@ -429,34 +326,21 @@ SAE_EXPAND=1 REPLAY_WORKERS=4 DE_WORKERS=2 python3 scripts/optimizer/optimize.py --params prob_threshold:0.5:0.999 anneal_sec:1:60 extinction_sec:1:60 \ --popsize 10 --maxiter 15 --trajectory traj.jsonl --out best.json -# replay the shipped config against a held-out film — --raw-out is needed to draw -# bboxes later (the merged pred.json has no per-frame bbox, only actor windows) -python3 scripts/optimizer/replay.py \ - --dump experiments/dumps/LVFace-B_Glint360K/dump_.h5 \ - --gallery experiments/galleries/gallery_LVFace-B_Glint360K.h5 \ - --out pred.json --raw-out raw.jsonl --prob-threshold 0.754 --anneal-sec 35.54 \ - --extinction-sec 57.43 --expand-gallery +# held-out validation, all 3 models, 5 films +python3 scripts/docs/run_holdout_all_models.py --out docs_data/holdout_all_models.json -# regenerate the report's charts (16-combo ranking, DE landscape, held-out -# per-film F1, Downton ghost timeline) from the artifacts under experiments/ +# per-film training breakdown, all 3 models, 4 films +python3 scripts/docs/run_holdout_all_models.py --films training --out docs_data/training_per_film.json + +# gallery coverage per film +python3 scripts/docs/gallery_coverage_per_film.py --out docs_data/gallery_coverage_per_film.json + +# regenerate this page's charts from experiments/ artifacts python3 scripts/docs/experiment_charts.py --out-dir docs/assets/images -# dump example frames (best-agreement / FPI / FN) for visual inspection, annotated -# with bounding boxes + names (--raw is optional; omit for unannotated frames) -python3 scripts/optimizer/dump_error_frames.py \ - --pred pred.json --raw raw.jsonl --xray experiments/xray/.../ \ - --movie "" \ - --gallery experiments/galleries/gallery_LVFace-B_Glint360K.h5 \ - --out-dir experiments/results/holdout/frames/ --n-per-bucket 4 - -# or: one best + one worst frame per 10-minute window across the whole film -python3 scripts/optimizer/dump_error_frames.py \ - --pred pred.json --raw raw.jsonl --xray experiments/xray/.../ \ - --movie "" \ - --gallery experiments/galleries/gallery_LVFace-B_Glint360K.h5 \ - --out-dir experiments/results/holdout/frames/_intervals --interval-sec 600 +# one frame per distinct out-of-cast name across all 9 films (used in the deep dive) +python3 scripts/docs/first_fpi_frames.py ``` -See also: [the prior optimizer round](optimizer-experiments.md) (superseded -metric) and the session log +See also the session log [`experiments/SESSION_STATE.md`](https://REPOLINK/experiments/SESSION_STATE.md). diff --git a/docs/optimizer-experiments.md b/docs/optimizer-experiments.md deleted file mode 100644 index 22b8b77..0000000 --- a/docs/optimizer-experiments.md +++ /dev/null @@ -1,125 +0,0 @@ -# Threshold optimization against Amazon X-Ray — experiment log - -Record of the July 2026 work that tuned the pipeline's recognition/tracking defaults -against ground-truth per-scene actor presence, and the tooling built to do it. - -## TL;DR — what changed - -| knob | old default | new default | why | -| ---- | ----------- | ----------- | --- | -| `prob_threshold` | 0.99 | **0.76** | 0.99 was far too strict — halved recall for a fraction of a precision point. DE optimum, tightly converged. | -| `extinction_sec` | 5.0 | **1.5** | Long extinction smears presence into later scenes → FPs. DE converged tightly low. | -| `anneal_sec` | 10.0 | 10.0 (unchanged) | DE found it **insensitive** (F1 flat ±0.3pp across 3–26s) — kept the round default. | -| `detector_conf` | 0.5 | 0.5 (unchanged) | Sweep showed raising it only trades recall for precision at a net F1 loss — near-threshold detections are real faces, not phantoms. | - -Net effect on the 9-film benchmark (strict per-scene, augmented gallery): -recall **58% → ~72%**, F1 **70% → ~76%**, precision ~85%, at no meaningful precision cost. - -## Ground truth - -Public scene-level **Amazon X-Ray** dataset (Zenodo DOI 10.5281/zenodo.17659734, -CC-BY-4.0): per movie, `people.csv` (name_id/person/character), `scenes.csv` -(scene/start/end ms), `people_in_scenes.csv`. Films matched to the library by an -**authoritative Jellyfin ID join** (query `/Items?IncludeItemTypes=Movie&Fields= -ProviderIds,Path`, join Imdb/Tmdb against X-Ray metadata) — NOT fuzzy title matching, -which collides badly (TV episodes vs same-named films). 9 genuine films with source -video on disk: Benny & Joon, Café Society, Downton Abbey: A New Era, Lord of War, -Lovelace, The Many Saints of Newark, Scarface, Sound of Metal, Valerian. - -## The scoring metric (evolved through review) - -Comparison unit is the **X-Ray scene**, not sampled timepoints. For each scene -`[start,end]`: predicted set = **union** of actors detected anywhere in the span; -GT set = actors X-Ray lists for that scene. Per scene TP/FP/FN, then: - -- **Precision: STRICT.** Any predicted actor not in the scene's X-Ray set is an FP, - *including out-of-cast confusions* (no gallery∩cast masking). An earlier - timepoint-sampled, cast-masked metric HID ~570 such FPs across 9 films and let the - optimizer drive `prob_threshold` to the 0.50 floor — a metric artifact. Counting - them is essential. -- **Recall: FAIR.** FN counts only X-Ray cast members **who are in the gallery**. 67% - of X-Ray cast (261/392) have no gallery reference embedding and can never be - recognised — counting them as misses penalises coverage, not the threshold. Both - `recall` (fair) and `recall_strict` (all) are reported. -- **Aggregation:** per-scene F1 → **duration-weighted average within a movie** (long - scenes count more) → **equal-weight mean across movies** (macro; each film counts - the same regardless of length). This is the DE objective. - -Implemented in `scripts/optimizer/scene_score.py` — since **removed** along -with this metric; its per-second successor is -[`scripts/optimizer/second_score.py`](https://REPOLINK/scripts/optimizer/second_score.py) -(see the [bake-off round](model-bakeoff.md)). - -## The gallery coverage gap - -Diagnosing low recall: only **131 of 392** X-Ray cast were in the gallery (33%). Every -in-gallery actor HAD embeddings (gallery well-formed) — the gap was pure coverage. -[`scripts/optimizer/fetch_missing_actors.py`](https://REPOLINK/scripts/optimizer/fetch_missing_actors.py) -recovers missing actors: -`nm-id → TMDB /find external_ids → /person/{id}/images → download → embed (sae_embed)`, -with a `--wikidata` fallback (P345→P18 Commons photo). - -- **TMDB recovered 143/261** (55%). 0 face-detection failures; the rest had no TMDB - person (60) or no profile photo (58). Coverage 33% → **70%**. -- **Wikidata fallback: 0/118** of the TMDB failures — only 4 even had a Commons photo, - none yielded a detectable face. → **TheTVDB not worth pursuing**: these remaining - actors are obscure enough that no image source covers them, AND (see below) most are - off-camera anyway. - -**Coverage vs detectability.** Adding references lifted recall (58→68% at fixed config) -but modestly. Per-film drill-down (Lord of War: 12 actors recovered, only 1 had a -detectable on-camera face) showed most missing cast are a **detectability gap** — X-Ray -credits them as cast-in-scene (incl. off-camera/background), but their face never -appears clearly for the pipeline to detect. This is a fundamental ceiling of a -face-recognition pipeline vs X-Ray's presence semantics, not a fixable gap. - -## Optimizer - -[`scripts/optimizer/optimize.py`](https://REPOLINK/scripts/optimizer/optimize.py) -— scipy `differential_evolution` over the knob space, -each candidate = full replay of all films through the **real** C++ nodes (see the -KPN replay architecture below) scored by the metric above. Global objective (one -config for all films, not per-film). - -**Convergence stability (augmented gallery, 233 evals):** - -| knob | top-20 range | verdict | -| ---- | ------------ | ------- | -| `prob_threshold` | 0.69–0.83 (σ 0.05) | TIGHT — trust 0.76 | -| `extinction_sec` | 1.0–2.2 (σ 0.33) | TIGHT — trust 1.5 | -| `anneal_sec` | 3.1–26.3 (σ 6.4) | LOOSE — insensitive, not hard-coded | - -F1 varied only 0.3pp across the top-20 → objective is flat near the optimum, so only -the tightly-converged knobs were adopted as defaults. - -## Replay architecture (how the sweep is cheap) - -The optimizer never re-decodes video. `scene_analyze --dump-embeddings out.h5` runs the -expensive half once (decode→detect→align→embed) and dumps per-frame face embeddings -+ metadata to HDF5 ([`scripts/optimizer/SCHEMA.md`](https://REPOLINK/scripts/optimizer/SCHEMA.md)). -[`scripts/optimizer/replay.py`](https://REPOLINK/scripts/optimizer/replay.py) then -replays that dump through the **real** C++ `face_tracker → identity_matcher → -scene_tracker` assembled in a Python KPN network (`sae_kpn` nanobind module), varying -Config knobs freely — no GPU embedding, no decode. Verified BYTE-EXACT against -`scene_analyze`'s own output. The dumps are gallery-independent, so testing the -augmented gallery needed no re-dump. `detector_conf` is replayable UPWARD only (the -dump floor is 0.5). - -## Reproduce - -```bash -# 1. dump (once per film, needs video) -scene_analyze --movie --gallery gallery.json --dump-embeddings dump.h5 --fps 1 -# 2. build films manifest by Jellyfin ID join (see scripts/optimizer notes) -# 3. optimize -python scripts/optimizer/optimize.py --manifest films.json --gallery gallery.json \ - --params prob_threshold:0.5:0.999 anneal_sec:1:30 extinction_sec:1:15 \ - --popsize 8 --maxiter 20 --trajectory traj.jsonl --out opt.json -# 4. score a fixed config / validate on a held-out set -# (historical: score_config.py and scene_score.py were removed with the -# scene-union metric — use scripts/optimizer/second_score.py, per-second) -python scripts/optimizer/second_score.py --help -``` - -Superseded by the [model bake-off + re-tune](model-bakeoff.md), which -replaced this round's scene-union metric with per-second scoring. diff --git a/docs/pose-expansion.md b/docs/pose-expansion.md index b8a5659..81fd109 100644 --- a/docs/pose-expansion.md +++ b/docs/pose-expansion.md @@ -1,42 +1,46 @@ -# Pose expansion: does "learning" new poses mid-film help? +# Pose expansion: does promoting new poses mid-film help? -`expand_gallery` ([`src/gallery/track_gallery.hpp`](https://REPOLINK/src/gallery/track_gallery.hpp)) -promotes a confidently-identified -track's novel-pose reference views into a per-film, in-memory gallery annex — the -idea being that once the pipeline is sure who someone is, a pose it hasn't seen -before (turned head, different lighting) becomes a free extra reference for -recognising that actor again later in the same film, without touching the baked -gallery. +`expand_gallery` +([`src/gallery/track_gallery.hpp`](https://REPOLINK/src/gallery/track_gallery.hpp)) +promotes a confidently identified track's novel-pose reference views into a +per-film, in-memory gallery annex. The idea: once the pipeline is confident +about an identity, a pose it has not seen before (turned head, different +lighting) becomes an extra reference for recognizing that actor again later +in the same film, without touching the baked gallery. -## The training-set signal +## Training-set signal -Averaged across all 4 models, on the 4 films used for optimization: +Averaged across the 3 compared models (r50 excluded), on the 4 films used +for optimization. These are the corrected, full-coverage figures, see the +[dropped-film note](model-bakeoff.md#a-scoring-bug-worth-recording-dropped-film-evaluations) +in the experiment log for why an earlier version of this table overstated the +full-mode misID jump (209 → 864) that was itself partly a truncation artifact: | scope | expansion | F1 | R | misID | |---|---|---|---|---| -| full | off | 71.2% | 58.3% | 209 | -| full | **on** | 71.2% | 59.7% | **864** | -| restricted | off | 73.6% | 61.3% | 194 | -| restricted | **on** | **75.4%** | **64.5%** | 135 | +| full | off | 70.0% | 57.6% | 407 | +| full | on | 72.1% | 61.5% | 714 | +| restricted | off | 75.1% | 63.9% | 179 | +| restricted | on | 76.7% | 67.2% | 120 | -In `restricted` mode (matcher's candidate set capped to the film's own credited -cast) expansion looked like a clean win: +1.8pp F1, +3.2pp recall, misID actually -lower. In `full` mode it looked flat-to-costly: ~0 F1 change, recall +1.4pp, but -misID roughly quadrupled (209 → 864) — see the -[bake-off experiment log](model-bakeoff.md) for the per-model breakdown. That's the number that motivated this page: **does turning -expansion on actually change what gets recognised, frame by frame, or is the -aggregate F1 shift something else?** +In restricted mode, expansion looks like a clean win: +1.6pp F1, +3.3pp +recall, lower misID. In full mode it looks like a recall-for-misID trade: ++2.1pp F1, +3.9pp recall, but misID rises from 407 to 714. See +[the full experiment log](model-bakeoff.md) for the per-model breakdown. +This asymmetry motivated the question below: does turning expansion on +change what gets recognized frame by frame, or is the aggregate F1 shift +coming from something else. -## Held-out test: does it reproduce? +## Held-out test -Same model + same tuned config, `expand_gallery` toggled on vs. off, nothing else -changed — full gallery mode, per-second scoring against X-Ray. This isolates -expansion from every other variable (config, model, threshold) that differs -between the training-set `exp`/`noexp` rows above. +Same model, same tuned config, `expand_gallery` toggled on vs. off, nothing +else changed, full gallery mode, per-second scoring against X-Ray. This +isolates expansion from every other variable that differs between the +training-set rows above. -**LVFace-B Glint360K, all 5 held-out films** (films never seen by the optimizer): +LVFace-B Glint360K, all 5 held-out films: -| film | F1 (exp) | F1 (noexp) | TPI Δ | FN Δ | +| film | F1 (exp) | F1 (noexp) | TPI delta | FN delta | |---|---|---|---|---| | Benny & Joon | 83.0% | 83.0% | -2 | +2 | | Downton Abbey: A New Era | 56.1% | 56.2% | -7 | +7 | @@ -44,57 +48,63 @@ between the training-set `exp`/`noexp` rows above. | The Many Saints of Newark | 46.3% | 46.3% | +2 | -2 | | Valerian and the City of a Thousand Planets | 74.1% | 74.1% | +2 | -2 | -**ArcFace R18** (Benny & Joon, r18's own tuned config): F1 77.1% for both, TPI/FN -identical, FPI differs by 2 (noise). +ArcFace R18, Benny & Joon, r18's own tuned config: F1 77.1% for both, TPI +and FN identical, FPI differs by 2. -**Every film, both models tested: F1 within 0.1–0.2pp, TPI/FN swings in the tens -out of tens of thousands.** That's noise, not a signal — expansion made no -measurable difference to per-second onscreen identification anywhere it was -tested on unseen data. +Every film, both models tested: F1 differs by 0.1-0.2pp, TPI/FN swings are +in the tens out of tens of thousands. This is noise, not a signal. +Expansion made no measurable difference to per-second on-screen +identification on any held-out film tested. -## Two bugs this required catching (this section's own methodology) +## Two methodology bugs caught during this check -Getting to the clean table above took two wrong turns, both worth recording -since they're exactly the kind of error that produces a false positive "look, -expansion helped!" finding: +Getting to the table above required catching two wrong turns, both worth +recording because they are exactly the kind of error that produces a false +positive "expansion helped" finding. -1. **Timeout truncation.** The first Downton Abbey `exp` replay was cut off by a - 60s subprocess timeout at ~76% through the film (5589 of 7368 expected - seconds) — a genuinely large, silent data loss that showed up as a large, - convincing-looking TPI gap (47938 vs 52032) purely because one run had a - quarter of the film missing. Caught by comparing `n_seconds` between runs - before trusting any score delta; fixed by re-running with a longer timeout. -2. **Bbox-matching bug.** An early per-second raw-annotation diff matched each - `exp` detection to the *first* `noexp` detection with IoU > 0.5, not the - *best*-overlapping one. With 3 faces close together in frame, this produced - spurious "disagreements" (e.g. "exp says Aidan Quinn, noexp says Johnny - Depp" at the same seconds) that vanished entirely once the match picked the - true best-IoU candidate — both configs had actually output the exact same - three names at the exact same three boxes. +1. **Timeout truncation.** The first Downton Abbey `exp` replay was cut off + by a 60-second subprocess timeout at about 76% through the film (5589 of + 7368 expected seconds). This silent data loss produced a large, + convincing-looking TPI gap (47938 vs 52032) purely because one run was + missing a quarter of the film. Caught by comparing `n_seconds` between + runs before trusting any score delta; fixed by re-running with a longer + timeout. +2. **Bbox-matching bug.** An early per-second raw-annotation diff matched + each `exp` detection to the first `noexp` detection with IoU above 0.5, + not the best-overlapping one. With 3 faces close together in frame, this + produced spurious disagreements (for example "exp says Aidan Quinn, + noexp says Johnny Depp" at the same second) that vanished once the match + used the best-IoU candidate instead of the first one. Both configs had + actually output the same three names at the same three boxes. -Both bugs independently pointed toward "expansion is doing something," and both -were artifacts of the comparison harness, not the pipeline. Worth remembering -when a before/after diff looks dramatic: check that the two runs actually cover -the same seconds, and match entities by best overlap, not first-found. +Both bugs independently pointed toward "expansion is doing something," and +both were artifacts of the comparison harness, not the pipeline. Before +trusting a dramatic before/after diff, check that both runs cover the same +seconds and that entities are matched by best overlap, not first found. -## What this means +## Conclusion -The training-set aggregate effect (particularly the ~4x misID increase in full -mode) doesn't reproduce on held-out data — at minimum it's far smaller than the -training-set numbers suggested, and plausibly it's sampling variation from only -4 training films rather than a real, generalizable mechanism. This doesn't mean -`expand_gallery` never does anything (the mechanism is real — see +The training-set aggregate effect, particularly the full-mode misID +increase, does not reproduce on held-out data. At minimum it +is far smaller than the training-set numbers suggested; it may be sampling +variation from only 4 training films rather than a generalizable +mechanism. Note the same *class* of harness bug appears twice in this +investigation, the timeout truncation in bug #1 above, and the dropped-film +aggregation that inflated the raw training-set misID figures. Both make an +inert config look consequential; both are reasons to distrust a dramatic +training-set delta until it survives on held-out films, which this one did +not. This does not mean `expand_gallery` never does anything: the +mechanism is real, and [`track_gallery.hpp`](https://REPOLINK/src/gallery/track_gallery.hpp)'s -promotion logging: tracks *do* get confirmed and views *do* -get promoted into the annex on every film tested), only that **whatever effect -it has on final per-second identification was too small to detect against 5 -held-out films** with this scoring method. A cleaner test would need either many -more held-out films or a metric that can see the annex's direct contribution -(e.g. tagging which reference embedding won each match), neither of which this -pass had budget for. +promotion logging confirms tracks get confirmed and views get promoted +into the annex on every film tested. It means whatever effect expansion +has on final per-second identification was too small to detect against 5 +held-out films with this scoring method. A cleaner test would need either +more held-out films or a metric that can see the annex's direct +contribution, such as tagging which reference embedding won each match; +neither was in scope for this pass. -**Practical takeaway**: don't treat the training-set `exp` vs `noexp` numbers in -the [bake-off experiment log](model-bakeoff.md) as proof that expansion -changes real-world behavior -in either direction — on the evidence gathered so far, it doesn't move the -needle enough to see. +Do not treat the training-set exp/noexp numbers in +[the full experiment log](model-bakeoff.md) as proof that expansion changes +real-world behavior in either direction. On the evidence gathered so far, +it does not move the needle enough to see. diff --git a/docs_data/holdout_all_models.json b/docs_data/holdout_all_models.json new file mode 100644 index 0000000..25bbaee --- /dev/null +++ b/docs_data/holdout_all_models.json @@ -0,0 +1,269 @@ +{ + "LVFace-B_Glint360K": { + "config": { + "prob_threshold": 0.7540024664611272, + "anneal_sec": 35.53996030397922, + "extinction_sec": 57.43359645269811 + }, + "films": { + "Benny___Joon": { + "name": "Benny & Joon", + "TPI": 15119, + "FPI": 1845, + "FPI_misid": 0, + "FPI_incast": 1845, + "FN": 4343, + "precision": 0.8912402735203961, + "precision_raw": 0.8912402735203961, + "recall": 0.7768471893947179, + "f1": 0.8301213418986437, + "agreement_rate": 0.7239983093829193, + "exact_match_rate": 0.41098901098901097, + "n_seconds": 5915, + "duration_sec": 5915.0 + }, + "Downton_Abbey__A_New_Era": { + "name": "Downton Abbey: A New Era", + "TPI": 52022, + "FPI": 1160, + "FPI_misid": 0, + "FPI_incast": 1160, + "FN": 80089, + "precision": 0.9781881087586025, + "precision_raw": 0.9781881087586025, + "recall": 0.39377493168623356, + "f1": 0.5615106884771687, + "agreement_rate": 0.4057686401759256, + "exact_match_rate": 0.033084311632870865, + "n_seconds": 7496, + "duration_sec": 7496.0 + }, + "Lovelace": { + "name": "Lovelace", + "TPI": 14988, + "FPI": 1086, + "FPI_misid": 58, + "FPI_incast": 1028, + "FN": 7087, + "precision": 0.9031091829356471, + "precision_raw": 0.9324374766703994, + "recall": 0.6789580973952435, + "f1": 0.775154508546456, + "agreement_rate": 0.7204967829586512, + "exact_match_rate": 0.3597703211914588, + "n_seconds": 5573, + "duration_sec": 5573.0 + }, + "The_Many_Saints_of_Newark": { + "name": "The Many Saints of Newark", + "TPI": 15928, + "FPI": 4394, + "FPI_misid": 974, + "FPI_incast": 3420, + "FN": 23785, + "precision": 0.5475797579757976, + "precision_raw": 0.7837811239051274, + "recall": 0.40107773273235464, + "f1": 0.46301652592258835, + "agreement_rate": 0.3705156874642392, + "exact_match_rate": 0.04588936642173853, + "n_seconds": 7213, + "duration_sec": 7213.0 + }, + "Valerian_and_the_City_of_a_Thousand_Plan": { + "name": "Valerian and the City of a Thousand Planets", + "TPI": 18658, + "FPI": 548, + "FPI_misid": 0, + "FPI_incast": 548, + "FN": 12472, + "precision": 0.9714672498177653, + "precision_raw": 0.9714672498177653, + "recall": 0.5993575329264376, + "f1": 0.7413382072472983, + "agreement_rate": 0.5877853464704299, + "exact_match_rate": 0.21980294368081743, + "n_seconds": 8221, + "duration_sec": 8221.0 + } + } + }, + "arcface_w600k_mbf": { + "config": { + "prob_threshold": 0.8371114538930519, + "anneal_sec": 48.80011450565114, + "extinction_sec": 59.3110040220424 + }, + "films": { + "Benny___Joon": { + "name": "Benny & Joon", + "TPI": 15219, + "FPI": 2463, + "FPI_misid": 180, + "FPI_incast": 2283, + "FN": 4243, + "precision": 0.7884675163195524, + "precision_raw": 0.8607058025110281, + "recall": 0.7819854074606927, + "f1": 0.7852130843050252, + "agreement_rate": 0.7072306082196138, + "exact_match_rate": 0.34911242603550297, + "n_seconds": 5915, + "duration_sec": 5915.0 + }, + "Downton_Abbey__A_New_Era": { + "name": "Downton Abbey: A New Era", + "TPI": 53043, + "FPI": 2383, + "FPI_misid": 604, + "FPI_incast": 1779, + "FN": 79068, + "precision": 0.8715290328940882, + "precision_raw": 0.9570057373795692, + "recall": 0.4015032813316075, + "f1": 0.5497453011561202, + "agreement_rate": 0.4094114144765671, + "exact_match_rate": 0.032817502668089645, + "n_seconds": 7496, + "duration_sec": 7496.0 + }, + "Lovelace": { + "name": "Lovelace", + "TPI": 14606, + "FPI": 1337, + "FPI_misid": 180, + "FPI_incast": 1157, + "FN": 7469, + "precision": 0.8316346865569664, + "precision_raw": 0.916138744276485, + "recall": 0.6616534541336353, + "f1": 0.7369695746505879, + "agreement_rate": 0.6907677036961077, + "exact_match_rate": 0.31742329086667864, + "n_seconds": 5573, + "duration_sec": 5573.0 + }, + "The_Many_Saints_of_Newark": { + "name": "The Many Saints of Newark", + "TPI": 15223, + "FPI": 4574, + "FPI_misid": 994, + "FPI_incast": 3580, + "FN": 24490, + "precision": 0.52962460425147, + "precision_raw": 0.768954892155377, + "recall": 0.383325359454083, + "f1": 0.44475283393712756, + "agreement_rate": 0.3554753116932427, + "exact_match_rate": 0.03715513655899071, + "n_seconds": 7213, + "duration_sec": 7213.0 + }, + "Valerian_and_the_City_of_a_Thousand_Plan": { + "name": "Valerian and the City of a Thousand Planets", + "TPI": 18472, + "FPI": 853, + "FPI_misid": 239, + "FPI_incast": 614, + "FN": 12658, + "precision": 0.860122927919538, + "precision_raw": 0.9558602846054334, + "recall": 0.5933825891423065, + "f1": 0.7022773067710907, + "agreement_rate": 0.5795914643682625, + "exact_match_rate": 0.18817662084904513, + "n_seconds": 8221, + "duration_sec": 8221.0 + } + } + }, + "arcface_r18": { + "config": { + "prob_threshold": 0.8955101189489445, + "anneal_sec": 59.08214397442713, + "extinction_sec": 59.29474134414983 + }, + "films": { + "Benny___Joon": { + "name": "Benny & Joon", + "TPI": 13580, + "FPI": 1666, + "FPI_misid": 60, + "FPI_incast": 1606, + "FN": 5882, + "precision": 0.86025592296972, + "precision_raw": 0.8907254361799817, + "recall": 0.697770013359367, + "f1": 0.7705401724920563, + "agreement_rate": 0.6547675401521545, + "exact_match_rate": 0.32578191039729504, + "n_seconds": 5915, + "duration_sec": 5915.0 + }, + "Downton_Abbey__A_New_Era": { + "name": "Downton Abbey: A New Era", + "TPI": 48545, + "FPI": 1066, + "FPI_misid": 180, + "FPI_incast": 886, + "FN": 83566, + "precision": 0.9475708067381078, + "precision_raw": 0.9785128298159682, + "recall": 0.3674561542944948, + "f1": 0.5295567845883649, + "agreement_rate": 0.3815368792000116, + "exact_match_rate": 0.032950907150480255, + "n_seconds": 7496, + "duration_sec": 7496.0 + }, + "Lovelace": { + "name": "Lovelace", + "TPI": 13615, + "FPI": 963, + "FPI_misid": 120, + "FPI_incast": 843, + "FN": 8460, + "precision": 0.8695235662281262, + "precision_raw": 0.933941555768967, + "recall": 0.6167610419026047, + "f1": 0.7216494845360825, + "agreement_rate": 0.6506356469257915, + "exact_match_rate": 0.2894311860757222, + "n_seconds": 5573, + "duration_sec": 5573.0 + }, + "The_Many_Saints_of_Newark": { + "name": "The Many Saints of Newark", + "TPI": 13489, + "FPI": 3757, + "FPI_misid": 796, + "FPI_incast": 2961, + "FN": 26224, + "precision": 0.5526013928717739, + "precision_raw": 0.7821523831613127, + "recall": 0.3396620753909299, + "f1": 0.42072267361165266, + "agreement_rate": 0.3229817885335633, + "exact_match_rate": 0.04422570359073894, + "n_seconds": 7213, + "duration_sec": 7213.0 + }, + "Valerian_and_the_City_of_a_Thousand_Plan": { + "name": "Valerian and the City of a Thousand Planets", + "TPI": 17692, + "FPI": 397, + "FPI_misid": 68, + "FPI_incast": 329, + "FN": 13438, + "precision": 0.9460456660071654, + "precision_raw": 0.9780529603626513, + "recall": 0.5683263732733698, + "f1": 0.710080070638759, + "agreement_rate": 0.5633540120828806, + "exact_match_rate": 0.13404695292543486, + "n_seconds": 8221, + "duration_sec": 8221.0 + } + } + } +} \ No newline at end of file diff --git a/docs_data/training_per_film.json b/docs_data/training_per_film.json new file mode 100644 index 0000000..767e5a8 --- /dev/null +++ b/docs_data/training_per_film.json @@ -0,0 +1,221 @@ +{ + "LVFace-B_Glint360K": { + "config": { + "prob_threshold": 0.7540024664611272, + "anneal_sec": 35.53996030397922, + "extinction_sec": 57.43359645269811 + }, + "films": { + "Caf\u00e9_Society": { + "name": "Caf\u00e9 Society", + "TPI": 14499, + "FPI": 1380, + "FPI_misid": 0, + "FPI_incast": 1380, + "FN": 12231, + "precision": 0.9130927640279615, + "precision_raw": 0.9130927640279615, + "recall": 0.5424242424242425, + "f1": 0.6805604449764134, + "agreement_rate": 0.57285804629501, + "exact_match_rate": 0.18947003810183582, + "n_seconds": 5774, + "duration_sec": 5774.0 + }, + "Lord_of_War": { + "name": "Lord of War", + "TPI": 13893, + "FPI": 1654, + "FPI_misid": 174, + "FPI_incast": 1480, + "FN": 5737, + "precision": 0.811838952842868, + "precision_raw": 0.8936129156750499, + "recall": 0.7077432501273561, + "f1": 0.7562256756388972, + "agreement_rate": 0.7005158404089996, + "exact_match_rate": 0.38715420432758146, + "n_seconds": 7302, + "duration_sec": 7302.0 + }, + "Scarface": { + "name": "Scarface", + "TPI": 20518, + "FPI": 1078, + "FPI_misid": 58, + "FPI_incast": 1020, + "FN": 14722, + "precision": 0.9276607288181572, + "precision_raw": 0.9500833487682904, + "recall": 0.5822360953461975, + "f1": 0.7154363820216882, + "agreement_rate": 0.6297735703976657, + "exact_match_rate": 0.25910733470065433, + "n_seconds": 10239, + "duration_sec": 10239.0 + }, + "Sound_of_Metal": { + "name": "Sound of Metal", + "TPI": 13349, + "FPI": 677, + "FPI_misid": 0, + "FPI_incast": 677, + "FN": 6504, + "precision": 0.9517324967916726, + "precision_raw": 0.9517324967916726, + "recall": 0.6723920818012391, + "f1": 0.7880397886596416, + "agreement_rate": 0.7112222835587533, + "exact_match_rate": 0.40311896218603366, + "n_seconds": 7246, + "duration_sec": 7246.0 + } + } + }, + "arcface_w600k_mbf": { + "config": { + "prob_threshold": 0.8371114538930519, + "anneal_sec": 48.80011450565114, + "extinction_sec": 59.3110040220424 + }, + "films": { + "Caf\u00e9_Society": { + "name": "Caf\u00e9 Society", + "TPI": 13456, + "FPI": 1434, + "FPI_misid": 180, + "FPI_incast": 1254, + "FN": 13274, + "precision": 0.8150211992731677, + "precision_raw": 0.9036937541974479, + "recall": 0.5034044145155256, + "f1": 0.622386679000925, + "agreement_rate": 0.5463565775524695, + "exact_match_rate": 0.19154832005542086, + "n_seconds": 5774, + "duration_sec": 5774.0 + }, + "Lord_of_War": { + "name": "Lord of War", + "TPI": 13778, + "FPI": 1740, + "FPI_misid": 60, + "FPI_incast": 1680, + "FN": 5852, + "precision": 0.8580146967243741, + "precision_raw": 0.8878721484727413, + "recall": 0.7018848700967907, + "f1": 0.7721362923111411, + "agreement_rate": 0.6881858851455998, + "exact_match_rate": 0.36469460421802247, + "n_seconds": 7302, + "duration_sec": 7302.0 + }, + "Scarface": { + "name": "Scarface", + "TPI": 18863, + "FPI": 862, + "FPI_misid": 0, + "FPI_incast": 862, + "FN": 16377, + "precision": 0.956299112801014, + "precision_raw": 0.956299112801014, + "recall": 0.535272417707151, + "f1": 0.6863640498499044, + "agreement_rate": 0.594178111391709, + "exact_match_rate": 0.2357652114464303, + "n_seconds": 10239, + "duration_sec": 10239.0 + }, + "Sound_of_Metal": { + "name": "Sound of Metal", + "TPI": 12642, + "FPI": 554, + "FPI_misid": 0, + "FPI_incast": 554, + "FN": 7211, + "precision": 0.9580175810851773, + "precision_raw": 0.9580175810851773, + "recall": 0.6367803354656727, + "f1": 0.7650458410239341, + "agreement_rate": 0.687468948385309, + "exact_match_rate": 0.3789677063207287, + "n_seconds": 7246, + "duration_sec": 7246.0 + } + } + }, + "arcface_r18": { + "config": { + "prob_threshold": 0.8955101189489445, + "anneal_sec": 59.08214397442713, + "extinction_sec": 59.29474134414983 + }, + "films": { + "Caf\u00e9_Society": { + "name": "Caf\u00e9 Society", + "TPI": 12207, + "FPI": 1119, + "FPI_misid": 60, + "FPI_incast": 1059, + "FN": 14523, + "precision": 0.88035482475119, + "precision_raw": 0.9160288158487168, + "recall": 0.45667789001122333, + "f1": 0.6013892994383683, + "agreement_rate": 0.5125626845924863, + "exact_match_rate": 0.1674748874263942, + "n_seconds": 5774, + "duration_sec": 5774.0 + }, + "Lord_of_War": { + "name": "Lord of War", + "TPI": 13122, + "FPI": 1409, + "FPI_misid": 60, + "FPI_incast": 1349, + "FN": 6508, + "precision": 0.870678787074514, + "precision_raw": 0.9030348909228546, + "recall": 0.6684666327050433, + "f1": 0.7562894441082388, + "agreement_rate": 0.6698963754222389, + "exact_match_rate": 0.3389482333607231, + "n_seconds": 7302, + "duration_sec": 7302.0 + }, + "Scarface": { + "name": "Scarface", + "TPI": 16961, + "FPI": 685, + "FPI_misid": 0, + "FPI_incast": 685, + "FN": 18279, + "precision": 0.961181004193585, + "precision_raw": 0.961181004193585, + "recall": 0.48129965947786607, + "f1": 0.6414173883447416, + "agreement_rate": 0.5440709115628456, + "exact_match_rate": 0.2017775173356773, + "n_seconds": 10239, + "duration_sec": 10239.0 + }, + "Sound_of_Metal": { + "name": "Sound of Metal", + "TPI": 12017, + "FPI": 591, + "FPI_misid": 122, + "FPI_incast": 469, + "FN": 7836, + "precision": 0.8767692981176127, + "precision_raw": 0.953125, + "recall": 0.6052989472623784, + "f1": 0.7161715188176049, + "agreement_rate": 0.6594534915815532, + "exact_match_rate": 0.3573005796301408, + "n_seconds": 7246, + "duration_sec": 7246.0 + } + } + } +} \ No newline at end of file diff --git a/mkdocs.yml b/mkdocs.yml index ae13c90..2ec5120 100644 --- a/mkdocs.yml +++ b/mkdocs.yml @@ -34,13 +34,13 @@ extra_css: nav: - Home: index.md + - How We Score Against X-Ray: methodology.md - Findings: - Best Model: best-model.md - Gallery Scope (Full vs. Limited): gallery-scope.md - Pose Expansion: pose-expansion.md - LVFace Deep Dive: lvface-deep-dive.md - - Model Bake-off & Re-tune (full log): model-bakeoff.md - - Optimizer Experiments (prior round): optimizer-experiments.md + - Full Experiment Log: model-bakeoff.md - Service Conversion (proposal): service-conversion.md markdown_extensions: diff --git a/scripts/docs/build_site.sh b/scripts/docs/build_site.sh index bbb0664..ccf0a2e 100755 --- a/scripts/docs/build_site.sh +++ b/scripts/docs/build_site.sh @@ -60,10 +60,34 @@ stage_frame "${MONTAGE_ROOT}/Valerian_and_the_City_of_a_Thousand_Plan/scene_4/4_ valerian_screen_call.jpg stage_frame "${MONTAGE_ROOT}/The_Many_Saints_of_Newark/out_of_cast_fpi/4_worst_t000871.jpg" \ many_saints_outofcast_fpi.jpg -# debug-overlay example (extinction state drawn as frozen boxes) — from the -# dump_error_frames output, not the montage package -stage_frame "experiments/results/holdout/frames/many_saints/fpi/fpi_t03543.jpg" \ - many_saints_ghost_fpi.jpg + +# One frame per DISTINCT out-of-cast name across all 9 films, uniform rule +# (see scripts/docs/first_fpi_frames.py): the first second in the raw replay +# stream where the pipeline names someone not in the film's credited cast at +# all. Rendered with the proper montage renderer (Onscreen/Offscreen panel), +# never dump_error_frames.py's bare-box overlay. Regenerate with: +# python3 scripts/docs/first_fpi_frames.py +# 5 of 9 films have zero out-of-cast names in their whole runtime (Benny & +# Joon, Cafe Society, Downton Abbey, Sound of Metal, Valerian) and produce +# no frames here. +stage_frame "${MONTAGE_ROOT}/Lord_of_War/first_fpi_david_shumbris/first_fpi_t000418.jpg" \ + lord_of_war_fpi_shumbris.jpg +stage_frame "${MONTAGE_ROOT}/Lord_of_War/first_fpi_ronald_reagan/first_fpi_t001003.jpg" \ + lord_of_war_fpi_reagan_photo.jpg +stage_frame "${MONTAGE_ROOT}/Lord_of_War/first_fpi_lance_reddick/first_fpi_t006424.jpg" \ + lord_of_war_fpi_reddick.jpg +stage_frame "${MONTAGE_ROOT}/Lovelace/first_fpi_chloë_sevigny/first_fpi_t002451.jpg" \ + lovelace_fpi_sevigny.jpg +stage_frame "${MONTAGE_ROOT}/Scarface/first_fpi_kirstie_alley/first_fpi_t002451.jpg" \ + scarface_fpi_alley.jpg +stage_frame "${MONTAGE_ROOT}/The_Many_Saints_of_Newark/first_fpi_germar_terrell_gardner/first_fpi_t000848.jpg" \ + many_saints_fpi_gardner.jpg +stage_frame "${MONTAGE_ROOT}/The_Many_Saints_of_Newark/first_fpi_archie_yates/first_fpi_t002521.jpg" \ + many_saints_fpi_yates.jpg +stage_frame "${MONTAGE_ROOT}/The_Many_Saints_of_Newark/first_fpi_zooey_deschanel/first_fpi_t002819.jpg" \ + many_saints_fpi_deschanel.jpg +stage_frame "${MONTAGE_ROOT}/The_Many_Saints_of_Newark/first_fpi_talia_balsam/first_fpi_t004551.jpg" \ + many_saints_fpi_balsam.jpg if [ ! -f "${ASSETS_DIR}/germar_beats_xray.jpg" ]; then echo "==> pulling report-highlights/germar_beats_xray.jpg..." diff --git a/scripts/docs/experiment_charts.py b/scripts/docs/experiment_charts.py index deeb09c..032b576 100644 --- a/scripts/docs/experiment_charts.py +++ b/scripts/docs/experiment_charts.py @@ -27,8 +27,11 @@ RESULTS = REPO / "experiments/results" # Same model -> color mapping as calibration_chart.py, so identity is stable # across every figure in the report. +# r50 is dropped from the bake-off: its 4 combos ran under the old, narrower +# anneal/extinction bounds and were never re-run wide, so they are not comparable +# on those two params (and two of them were truncation-corrupted). Its slug stays +# out of this map so it never appears in a figure or legend. MODEL_COLOURS = { - "arcface_w600k_r50": ("ArcFace w600k-R50", "#2a78d6"), "arcface_r18": ("ArcFace R18", "#008300"), "arcface_w600k_mbf": ("ArcFace w600k-MBF", "#e87ba4"), "LVFace-B_Glint360K": ("LVFace-B Glint360K", "#eda100"), @@ -59,9 +62,37 @@ plt.rcParams.update({ }) +TRAJ = REPO / "experiments/trajectories" + + +def clean_best(combo: str) -> dict: + """Best-F1 eval for a combo, restricted to FULL-COVERAGE evals. + + The optimizer averages F1 (and *sums* TPI/misID) over only the films whose + replay subprocess didn't time out (optimize.py: `per_film = [... if m is not + None]`). A candidate whose hardest film timed out is therefore scored on an + easier subset, which inflates its F1 — and DE will happily converge onto such + a candidate. `rep4_best_*.json` recorded exactly that kind of eval for at + least one combo (arcface_w600k_mbf_full_noexp: reported 74.2% F1 came from an + eval with TPI 12645, a third of that combo's median). + + We recover comparable numbers straight from the trajectory: take the median + TPI across all evals (full 4-film coverage) and keep only evals within 30% of + it, then pick the highest-F1 survivor. No re-running — the honest best config + is already in the sweep, just not the one `argmax f1` picked. + """ + evals = [json.loads(l) for l in open(TRAJ / f"rep4_{combo}.jsonl")] + tpis = sorted(e["TPI"] for e in evals) + med = tpis[len(tpis) // 2] + clean = [e for e in evals if e["TPI"] >= 0.7 * med] + return max(clean, key=lambda e: e["f1"]) + + def training_best() -> dict: - with open(RESULTS / "rep4_best_LVFace-B_Glint360K_full_exp.json") as f: - return json.load(f)["best"] + # LVFace-B_Glint360K_full_exp is the shipped combo; its reported best is a + # full-coverage eval (TPI 47757 ≈ median), so clean_best returns the same + # config — but route it through clean_best so every figure uses one path. + return clean_best("LVFace-B_Glint360K_full_exp") def fig_holdout_f1(out: Path): @@ -100,18 +131,16 @@ def fig_holdout_f1(out: Path): def fig_rep4_matrix(out: Path): combos = [] - for path in sorted(RESULTS.glob("rep4_best_*.json")): - stem = path.stem[len("rep4_best_"):] + for path in sorted(TRAJ.glob("rep4_*.jsonl")): + combo = path.stem[len("rep4_"):] for slug in MODEL_COLOURS: - if stem.startswith(slug): - mode = stem[len(slug) + 1:] # e.g. full_exp - with open(path) as f: - best = json.load(f)["best"] - combos.append((slug, mode, best["f1"] * 100)) + if combo.startswith(slug): + mode = combo[len(slug) + 1:] # e.g. full_exp + combos.append((slug, mode, clean_best(combo)["f1"] * 100)) break combos.sort(key=lambda c: c[2]) - fig, ax = plt.subplots(figsize=(9, 6.2)) + fig, ax = plt.subplots(figsize=(9, 5.2)) ax.grid(axis="y", visible=False) labels = [] for i, (slug, mode, f1) in enumerate(combos): @@ -126,7 +155,7 @@ def fig_rep4_matrix(out: Path): ax.set_yticks(range(len(combos)), labels, fontsize=9) ax.set_xlim(65, 80) ax.set_xlabel("training-set per-second F1 (%)") - ax.set_title("All 16 combos — filled dot = cast-restricted gallery, open = full", + ax.set_title("All 12 combos — filled dot = cast-restricted gallery, open = full", loc="left", fontsize=12, pad=12) handles = [plt.Line2D([], [], marker="o", ls="", ms=9, color=c, label=l) for _, (l, c) in MODEL_COLOURS.items()] @@ -199,13 +228,16 @@ def fig_downton_timeline(out: Path, t0: int = 7100, t1: int = 7340): trk = np.array(trk) dc = np.array([det.get(s, 0) for s in t]) - fig, ax = plt.subplots(figsize=(9.5, 4.4)) + fig, ax = plt.subplots(figsize=(9.5, 4.8)) ax.grid(axis="x", visible=False) - ax.fill_between(t, dc, step="mid", color=GREEN, alpha=0.25, zorder=2) - ax.step(t, dc, where="mid", color=GREEN, lw=2, zorder=3) - ax.step(t, trk, where="mid", color=BLUE, lw=2, zorder=4) + ax.fill_between(t, dc, step="mid", color=GREEN, alpha=0.22, zorder=2) + ax.step(t, dc, where="mid", color=GREEN, lw=2, zorder=3, + label="faces seen by detector") + ax.step(t, trk, where="mid", color=BLUE, lw=2, zorder=4, + label="actors reported by tracker") # longest contiguous run of "detector sees nothing, tracker still reporting" + # (i.e. every reported actor is extinction-bridged, not detected this second) ghost = (dc == 0) & (trk > 0) runs, start = [], None for i, g in enumerate(ghost): @@ -219,15 +251,11 @@ def fig_downton_timeline(out: Path, t0: int = 7100, t1: int = 7340): if runs: i0, i1 = max(runs, key=lambda r: r[1] - r[0]) g0, g1 = t[i0], t[i1] - ax.axvspan(g0, g1, color=RED, alpha=0.08, zorder=1) - ax.annotate(f"{g1 - g0}s of credits: 0 faces detected,\n" - f"{trk[i0]} actors still reported (frozen boxes)", - ((g0 + g1) / 2, 20.5), ha="center", va="bottom", - fontsize=10, color=RED) - ax.text(t0 + 4, 27.3, "actors reported by tracker", color=BLUE, - fontsize=10.5, va="bottom") - ax.text(t0 + 4, 11.5, "faces seen by detector", color=GREEN, - fontsize=10.5, va="bottom") + ax.axvspan(g0, g1, color=RED, alpha=0.08, zorder=1, + label=f"{g1 - g0}s bridged: 0 faces detected,\n" + f"{trk[i0]} actors carried by their\nextinction window") + ax.legend(loc="upper right", frameon=True, framealpha=0.92, + edgecolor=GRID, fontsize=9.5) ax.set_xlabel("film time (s)") ax.set_ylabel("count") ax.set_ylim(0, 31) diff --git a/scripts/docs/first_fpi_frames.py b/scripts/docs/first_fpi_frames.py new file mode 100644 index 0000000..9fbb68b --- /dev/null +++ b/scripts/docs/first_fpi_frames.py @@ -0,0 +1,155 @@ +#!/usr/bin/env python3 +""" +first_fpi_frames.py — for every film, find every DISTINCT out-of-cast name +(misID) the raw replay stream ever reports, and render the exact second each +one FIRST appears, with the proper montage renderer (dump_scene_montage.py: +Onscreen/Offscreen panel, TPI/FPI/FN legend, ghosts never drawn as boxes — +imported directly, not the scene-level best/worst picker, which can land on +a different second within the same scene). + +One rule, applied uniformly across all 9 films and every distinct wrong name +in each — no manual per-film picking, no stopping at the first name found. +""" +import csv +import json +import sys +from pathlib import Path + +import cv2 + +REPO = Path(__file__).resolve().parent.parent.parent +sys.path.insert(0, str(REPO / "scripts" / "validation")) +sys.path.insert(0, str(REPO / "scripts" / "optimizer")) +from identity import keys_for # noqa: E402 +from sample_eval import load_gallery_keys # noqa: E402 +from dump_scene_montage import ( # noqa: E402 + classify_second, extract_frame, render_frame, + load_scene_cast, load_dump_faces_by_second, load_raw_by_second, +) + +FILMS = [ + ("Benny___Joon", "experiments/xray/scene_level_movie_data_XRay_US/xrays/4808_Benny__Joon"), + ("Café_Society", "experiments/xray/scene_level_movie_data_XRay_US/xrays/225_Cafe_Society"), + ("Downton_Abbey__A_New_Era", "experiments/xray/scene_level_movie_data_XRay_US/xrays/19_Downton_Abbey_A_New_Era"), + ("Lord_of_War", "experiments/xray/scene_level_movie_data_XRay_US/xrays/2474_Lord_of_War"), + ("Lovelace", "experiments/xray/scene_level_movie_data_XRay_US/xrays/4108_Lovelace"), + ("Scarface", "experiments/xray/scene_level_movie_data_XRay_US/xrays/197_Scarface"), + ("Sound_of_Metal", "experiments/xray/scene_level_movie_data_XRay_US/xrays/6278_Sound_of_Metal"), + ("The_Many_Saints_of_Newark", "experiments/xray/scene_level_movie_data_XRay_US/xrays/900_The_Many_Saints_Of_Newark"), + ("Valerian_and_the_City_of_a_Thousand_Plan", "experiments/xray/scene_level_movie_data_XRay_US/xrays/5312_Valerian_and_the_City_of_a_Thousand_Planets"), +] + +MOVIE_ROOT = Path("/mnt/movies") + + +def load_film_cast_keys(xray_dir: Path) -> set: + keys = set() + with open(xray_dir / "people.csv", newline="", encoding="utf-8") as f: + for r in csv.DictReader(f): + nm = (r.get("name_id") or "").strip() + person = (r.get("person") or "").strip() + if nm or person: + keys |= keys_for(imdb_id=nm, name=person) + return keys + + +def find_movie_file(slug: str) -> str | None: + # dump HDF5 attrs carry the exact path used at dump time + import h5py + for model in ("LVFace-B_Glint360K",): + p = REPO / f"experiments/dumps/{model}/dump_{slug}.h5" + if p.exists(): + with h5py.File(p, "r") as f: + return f.attrs.get("movie") + return None + + +def find_scene_id(xray_dir: Path, t: int) -> str | None: + with open(xray_dir / "scenes.csv", newline="", encoding="utf-8") as f: + for r in csv.DictReader(f): + try: + t0, t1 = float(r["start"]) / 1000.0, float(r["end"]) / 1000.0 + except (KeyError, ValueError): + continue + if t0 <= t < t1: + return (r.get("scene") or "").strip() + return None + + +def main(): + out_root = REPO / "experiments/results/holdout/montage_bestworst" + summary = [] + + for slug, xray_rel in FILMS: + xray_dir = REPO / xray_rel + raw_path = out_root / f"raw_{slug}.jsonl" + if not raw_path.exists(): + print(f"SKIP {slug}: no raw file", file=sys.stderr) + continue + + cast_keys = load_film_cast_keys(xray_dir) + + # every distinct out-of-cast name -> first second it appears + first_seen: dict[str, int] = {} + with open(raw_path) as f: + for line in f: + d = json.loads(line) + if d.get("eof"): + continue + for a in d.get("visible_actors", []): + name = a.get("name") + if not name or name in first_seen: + continue + ak = keys_for(imdb_id=a.get("imdb_id"), name=name, + jellyfin_id=a.get("jellyfin_id")) + if not (ak & cast_keys): + first_seen[name] = int(d["timestamp_sec"]) + + if not first_seen: + print(f"{slug}: no out-of-cast FPI in the whole film", file=sys.stderr) + summary.append((slug, None, None)) + continue + + print(f"{slug}: {len(first_seen)} distinct out-of-cast name(s)", file=sys.stderr) + + movie = find_movie_file(slug) + if not movie or not Path(movie).exists(): + print(f" SKIP render: movie file not found ({movie})", file=sys.stderr) + for name, t in first_seen.items(): + summary.append((slug, name, t)) + continue + + dump_path = REPO / f"experiments/dumps/LVFace-B_Glint360K/dump_{slug}.h5" + gallery_path = REPO / "experiments/galleries/gallery_LVFace-B_Glint360K.h5" + gallery_keys = load_gallery_keys(str(gallery_path)) + raw_by_second = load_raw_by_second(str(raw_path)) + dump_faces_by_second = load_dump_faces_by_second(str(dump_path)) + scene_cast = load_scene_cast(str(xray_dir)) + + for name, t in sorted(first_seen.items(), key=lambda kv: kv[1]): + scene_id = find_scene_id(xray_dir, t) + gt_cast = scene_cast.get(scene_id, set()) + gt_cast = {g for g in gt_cast if g & gallery_keys} + + score, tpi_boxes, fpi_boxes, entries, has_outofcast = classify_second( + t, gt_cast, cast_keys, raw_by_second, dump_faces_by_second) + + slug_name = name.lower().replace(" ", "_").replace("'", "") + out_dir = out_root / slug / f"first_fpi_{slug_name}" + out_dir.mkdir(parents=True, exist_ok=True) + out_path = out_dir / f"first_fpi_t{t:06d}.jpg" + extract_frame(movie, t, out_path) + canvas = render_frame(out_path, t, tpi_boxes, fpi_boxes, entries) + if canvas is not None: + cv2.imwrite(str(out_path), canvas) + print(f" {name!r} t={t}s -> {out_path} (outofcast={has_outofcast})", + file=sys.stderr) + summary.append((slug, name, t)) + + print("\n=== summary ===", file=sys.stderr) + for slug, name, t in summary: + print(f" {slug:45s} {name!r:30s} t={t}", file=sys.stderr) + + +if __name__ == "__main__": + main() diff --git a/scripts/docs/gallery_coverage_per_film.py b/scripts/docs/gallery_coverage_per_film.py new file mode 100644 index 0000000..1351471 --- /dev/null +++ b/scripts/docs/gallery_coverage_per_film.py @@ -0,0 +1,62 @@ +#!/usr/bin/env python3 +""" +gallery_coverage_per_film.py — fraction of each film's X-Ray credited cast +that has a reference embedding in the gallery, computed per film rather than +as a single benchmark-wide average. + +Usage: python3 scripts/docs/gallery_coverage_per_film.py --out docs_data/gallery_coverage_per_film.json +""" +import argparse +import csv +import json +import sys +from pathlib import Path + +import h5py + +REPO = Path(__file__).resolve().parent.parent.parent +sys.path.insert(0, str(REPO / "scripts" / "validation")) +from identity import keys_for # noqa: E402 + + +def main(): + p = argparse.ArgumentParser() + p.add_argument("--gallery", default=str(REPO / "experiments/galleries/gallery_LVFace-B_Glint360K.h5")) + p.add_argument("--films", default=str(REPO / "experiments/manifests/films.json")) + p.add_argument("--out", required=True) + args = p.parse_args() + + films = json.load(open(args.films)) + with h5py.File(args.gallery, "r") as f: + names = [n.decode() if isinstance(n, bytes) else n for n in f["name"][:]] + jids = [j.decode() if isinstance(j, bytes) else j for j in f["jellyfin_id"][:]] + imdbs = [j.decode() if isinstance(j, bytes) else j for j in f["imdb_id"][:]] + gallery_keys = set() + for n, j, im in zip(names, jids, imdbs): + gallery_keys |= keys_for(imdb_id=im, name=n, jellyfin_id=j) + + out = [] + for film in films: + xray_dir = REPO / film["xray"] + id_to_name = {} + with open(xray_dir / "people.csv", newline="", encoding="utf-8") as fh: + for r in csv.DictReader(fh): + nm = (r.get("name_id") or "").strip() + if nm: + id_to_name[nm] = (r.get("person") or "").strip() + cast_keys = [keys_for(imdb_id=nm, name=name) for nm, name in id_to_name.items()] + covered = sum(1 for ck in cast_keys if ck & gallery_keys) + total = len(cast_keys) + out.append({"film": film["name"], "cast_total": total, "covered": covered, + "coverage_pct": round(covered / total * 100, 1) if total else 0.0}) + + out.sort(key=lambda x: x["coverage_pct"]) + Path(args.out).parent.mkdir(parents=True, exist_ok=True) + json.dump(out, open(args.out, "w"), indent=1) + for o in out: + print(f"{o['film']:45s} {o['covered']:3d}/{o['cast_total']:3d} ({o['coverage_pct']}%)", + file=sys.stderr) + + +if __name__ == "__main__": + main() diff --git a/scripts/docs/run_holdout_all_models.py b/scripts/docs/run_holdout_all_models.py new file mode 100644 index 0000000..af98817 --- /dev/null +++ b/scripts/docs/run_holdout_all_models.py @@ -0,0 +1,107 @@ +#!/usr/bin/env python3 +""" +run_holdout_all_models.py — replay each model's own tuned full_exp config +against the 5 held-out films, score with second_score.py, and dump a combined +JSON. r50 is excluded (see docs/model-bakeoff.md: dropped from the detailed +comparison, kept only in the calibration-curve chart). + +This fills a real gap: the shipped report claimed "nothing in held-out +validation contradicts the model choice" without ever running mbf/r18 on the +held-out films — only LVFace had been checked. + +Usage: python3 scripts/docs/run_holdout_all_models.py --out docs_data/holdout_all_models.json +""" +import argparse +import json +import subprocess +import sys +from pathlib import Path + +REPO = Path(__file__).resolve().parent.parent.parent +sys.path.insert(0, str(REPO / "scripts" / "optimizer")) +sys.path.insert(0, str(REPO / "scripts" / "validation")) +from second_score import score_seconds # noqa: E402 +from sample_eval import load_gallery_keys # noqa: E402 + +MODELS = ["LVFace-B_Glint360K", "arcface_w600k_mbf", "arcface_r18"] + +HELDOUT = [ + {"name": "Benny & Joon", "slug": "Benny___Joon", + "xray": "experiments/xray/scene_level_movie_data_XRay_US/xrays/4808_Benny__Joon"}, + {"name": "Downton Abbey: A New Era", "slug": "Downton_Abbey__A_New_Era", + "xray": "experiments/xray/scene_level_movie_data_XRay_US/xrays/19_Downton_Abbey_A_New_Era"}, + {"name": "Lovelace", "slug": "Lovelace", + "xray": "experiments/xray/scene_level_movie_data_XRay_US/xrays/4108_Lovelace"}, + {"name": "The Many Saints of Newark", "slug": "The_Many_Saints_of_Newark", + "xray": "experiments/xray/scene_level_movie_data_XRay_US/xrays/900_The_Many_Saints_Of_Newark"}, + {"name": "Valerian and the City of a Thousand Planets", + "slug": "Valerian_and_the_City_of_a_Thousand_Plan", + "xray": "experiments/xray/scene_level_movie_data_XRay_US/xrays/5312_Valerian_and_the_City_of_a_Thousand_Planets"}, +] + +TRAINING = [ + {"name": "Café Society", "slug": "Café_Society", + "xray": "experiments/xray/scene_level_movie_data_XRay_US/xrays/225_Cafe_Society"}, + {"name": "Lord of War", "slug": "Lord_of_War", + "xray": "experiments/xray/scene_level_movie_data_XRay_US/xrays/2474_Lord_of_War"}, + {"name": "Scarface", "slug": "Scarface", + "xray": "experiments/xray/scene_level_movie_data_XRay_US/xrays/197_Scarface"}, + {"name": "Sound of Metal", "slug": "Sound_of_Metal", + "xray": "experiments/xray/scene_level_movie_data_XRay_US/xrays/6278_Sound_of_Metal"}, +] + + +def main(): + p = argparse.ArgumentParser() + p.add_argument("--out", required=True) + p.add_argument("--work-dir", default="/tmp/holdout_all_models") + p.add_argument("--films", choices=["heldout", "training"], default="heldout") + args = p.parse_args() + + work = Path(args.work_dir) + work.mkdir(parents=True, exist_ok=True) + film_set = HELDOUT if args.films == "heldout" else TRAINING + + results = {} + for model in MODELS: + cfg = json.load(open(REPO / f"experiments/results/rep4_best_{model}_full_exp.json"))["best"]["config"] + gallery = REPO / f"experiments/galleries/gallery_{model}.h5" + results[model] = {"config": cfg, "films": {}} + + for film in film_set: + dump = REPO / f"experiments/dumps/{model}/dump_{film['slug']}.h5" + if not dump.exists(): + print(f"SKIP {model}/{film['slug']}: no dump", file=sys.stderr) + continue + pred_path = work / f"pred_{model}_{film['slug']}.json" + cmd = [ + "python3", "scripts/optimizer/replay.py", + "--dump", str(dump), "--gallery", str(gallery), + "--out", str(pred_path), + "--prob-threshold", str(cfg["prob_threshold"]), + "--anneal-sec", str(cfg["anneal_sec"]), + "--extinction-sec", str(cfg["extinction_sec"]), + "--expand-gallery", + ] + print(f"RUN {model}/{film['slug']}...", file=sys.stderr) + r = subprocess.run(cmd, cwd=REPO, capture_output=True, text=True, timeout=120) + if r.returncode != 0: + print(f"FAIL {model}/{film['slug']}: {r.stderr[-800:]}", file=sys.stderr) + results[model]["films"][film["slug"]] = {"error": r.stderr[-500:]} + continue + + gk = load_gallery_keys(str(gallery)) + pred_json = json.loads(pred_path.read_text()) + m = score_seconds(pred_json, str(REPO / film["xray"]), gk) + results[model]["films"][film["slug"]] = {"name": film["name"], **m} + print(f" -> F1={m['f1']*100:.1f}% P={m['precision']*100:.1f}% " + f"R={m['recall']*100:.1f}% misid={m['FPI_misid']}", file=sys.stderr) + + Path(args.out).parent.mkdir(parents=True, exist_ok=True) + with open(args.out, "w") as f: + json.dump(results, f, indent=1) + print(f"wrote {args.out}", file=sys.stderr) + + +if __name__ == "__main__": + main() diff --git a/scripts/optimizer/optimize.py b/scripts/optimizer/optimize.py index 66413a2..e0a5bc0 100644 --- a/scripts/optimizer/optimize.py +++ b/scripts/optimizer/optimize.py @@ -106,8 +106,12 @@ def evaluate(cfg, films, build_dir, step=None): """Objective = MACRO-mean over films of each film's duration-weighted per-scene F1. Each film's replay runs in a subprocess (timeout-guarded) to survive the - intermittent ROCm teardown deadlock. A film whose replay times out is dropped - from the average rather than hanging the whole sweep. + intermittent ROCm teardown deadlock. If ANY film's replay times out, this + evaluation is scored f1=0.0 (see below) rather than averaging over the + survivors — a partial-coverage eval must never look better than a complete + one, or DE will converge onto configs that make the hardest film time out. + (An earlier version averaged over survivors, which silently rewarded + truncation; the rep4 `mbf_full_noexp` winner was one such corrupted eval.) UNIFORM PER-SECOND scoring (second_score.py): every second of the film is sampled; GT(t) = the cast of the X-Ray scene containing t, Pred(t) = actors whose presence @@ -135,9 +139,21 @@ def evaluate(cfg, films, build_dir, step=None): with ThreadPoolExecutor(max_workers=REPLAY_WORKERS) as ex: per_film = [m for m in ex.map(_one, films) if m is not None] n = len(per_film) - if not n: - return {"precision": 0.0, "recall": 0.0, "f1": 0.0, "agreement": 0.0, - "TPI": 0, "FPI": 0, "FPI_misid": 0, "FN": 0} + n_expected = len(films) + # Incomplete coverage (a replay timed out) is scored as a failure, not + # averaged over survivors: dropping the hardest film would otherwise inflate + # the score and let DE reward exactly the configs that cause timeouts. We + # still record the real survivor counts so a truncated eval is diagnosable + # in the trajectory (f1=0.0, films_scored < films_expected). + if n < n_expected: + agg = {"precision": 0.0, "recall": 0.0, "f1": 0.0, "agreement": 0.0, + "TPI": sum(m["TPI"] for m in per_film), + "FPI": sum(m["FPI"] for m in per_film), + "FPI_misid": sum(m["FPI_misid"] for m in per_film), + "FN": sum(m["FN"] for m in per_film)} + agg["films_scored"] = n + agg["films_expected"] = n_expected + return agg return {"precision": sum(m["precision"] for m in per_film) / n, "recall": sum(m["recall"] for m in per_film) / n, "f1": sum(m["f1"] for m in per_film) / n, @@ -145,7 +161,9 @@ def evaluate(cfg, films, build_dir, step=None): "TPI": sum(m["TPI"] for m in per_film), "FPI": sum(m["FPI"] for m in per_film), "FPI_misid": sum(m["FPI_misid"] for m in per_film), - "FN": sum(m["FN"] for m in per_film)} + "FN": sum(m["FN"] for m in per_film), + "films_scored": n, + "films_expected": n_expected} def main():