Files
scene-actor-extraction/docs/lvface-deep-dive.md
T
dtourolle 4b5557974b docs: montage-renderer imagery, visual polish, README screenshots
- switch report frames to the scene best/worst montage renderer
  (Onscreen/Offscreen panels + TPI/FPI/FN legend): perfect-second hero,
  wedding couple, funeral 19-of-20, polygraph bridging, crew-scene FN
  ceiling, Robert Patrick ground-truth gap, rapid-cut double label,
  Herbie Hancock on an in-fiction screen
- deep dive restructured: extinction bridging framed as designed
  behavior with a measurable cost (debug overlay draws the boxes; the
  shipped output is presence windows), plus the face-vs-presence
  ceiling and two X-Ray-is-wrong exhibits
- Material polish: light/dark palette toggle, landing-page grid cards,
  figure/caption CSS, how-to-read admonition; site_url set so 404 links
  resolve under the Pages subpath
- README: perfect-second and screen-call frames committed (gitignore
  exceptions), readme_example.jpg retired
- build_site.sh: stage_frame helper downscales montage frames to 1920px
  and pulls any missing montage-frames packages
2026-07-19 22:27:57 +02:00

195 lines
10 KiB
Markdown

# Deep dive: LVFace-B Glint360K
LVFace won the model bake-off (see [Which model is best?](best-model.md)) and is
the shipped default embedder. This page is the honest accounting of how it
actually performs — what a good second looks like, where the errors actually
come from, and two cases where the ground truth itself is wrong and LVFace is
right.
!!! note "How to read the frames on this page"
The top is the film frame, with a box and name on every face the pipeline
identified. The bottom panels are the per-second verdict against X-Ray:
**Onscreen** lists faces named in the frame, **Offscreen** lists cast
X-Ray marks present in the scene without a visible face — presence
carried by the tracker's windows, not by a detection. Colors are the
score: <span style="color:#0ca30c">**green**</span> = correct (TPI),
<span style="color:#eb6834">**orange**</span> = wrong (FPI),
<span style="color:#3987e5">**blue**</span> = missed (FN).
## What good looks like
![Wedding couple correctly identified, Downton Abbey: A New Era](assets/images/downton_wedding_couple.jpg)
Six faces on screen, all six named correctly — including Penelope Wilton at the
edge of the pews and a half-occluded Michelle Dockery — while thirteen more
cast members X-Ray marks present in the scene are correctly carried as
"Offscreen" by their presence windows. One miss in the whole frame: Maggie
Smith (blue). Score for this second: 0.86.
![19 of 20 correct in the funeral crowd](assets/images/downton_funeral_19of20.jpg)
The same film's funeral gathering: mourning dress, hats, half the faces turned.
**Nineteen of the twenty cast X-Ray lists for this scene are scored correctly**
— seven named on screen at up to 100% confidence, twelve more correctly held
as present off-screen.
And the pipeline doesn't need the face to be *real*:
![Herbie Hancock identified on an in-fiction video call](assets/images/valerian_screen_call.jpg)
That's Herbie Hancock at 98% — as a face on a *screen inside the movie*, over a
sci-fi HUD overlay, during a video call in Valerian. A face is a face, whether
it's in the room or on the bridge's comms display.
## Training vs. held-out: the generalization gap
The shipped config (`prob_threshold=0.754, anneal_sec=35.54,
extinction_sec=57.43, expand_gallery=true`) was tuned against 4 films. Scored
against the 5 films the optimizer never saw:
![Held-out per-film F1 vs. the training-set fit](assets/images/holdout_f1_by_film.png)
| film | F1 | P | R | TPI | FPI | misid | FN |
|---|---|---|---|---|---|---|---|
| Benny & Joon | 83.0% | 89.1% | 77.7% | 15125 | 1846 | 0 | 4337 |
| Lovelace | 77.5% | 90.3% | 67.9% | 14990 | 1085 | 58 | 7085 |
| Valerian and the City of a Thousand Planets | 74.1% | 97.1% | 60.0% | 18663 | 548 | 0 | 12467 |
| Downton Abbey: A New Era | 56.2% | 97.8% | 39.4% | 52027 | 1173 | 0 | 80084 |
| **The Many Saints of Newark** | **46.3%** | **54.7%** | 40.1% | 15922 | 4394 | **974** | 23791 |
| **macro average** | **67.4%** | 85.8% | 57.0% | | | | |
**67.4% held-out vs. 75.3% on training** — an ~8pp drop, and a **37pp spread
between the best and worst held-out film**. The config does not generalize
uniformly, and the spread traces to two mechanisms, both visible frame by
frame below.
## Mechanism 1: extinction bridging — usually right, wrong at hard cuts
The extinction window keeps an identity alive through seconds where no face is
detectable. **Most of the time this is exactly what you want**, and it's where
a lot of the TPI count comes from:
![Two faces on screen, six more correctly bridged](assets/images/lovelace_polygraph_bridged.jpg)
Lovelace's polygraph scene: only Eric Roberts and Amanda Seyfried have visible
faces, but X-Ray lists eight cast present — and all eight score green, the
other six correctly carried by presence windows through a scene where the
camera never shows them. A perfect second, and the extinction/anneal machinery
is *why*.
The same mechanism has a failure case: a hard cut into long faceless footage.
Both Many Saints of Newark (974 misIDs) and Downton Abbey (FN=80084, the worst
recall of the five) are dominated by it — verified directly against the raw
per-frame stream and the HDF5 dump's own detection counts, not inferred from
the score alone. **This is not a malfunction**: the tracker is doing exactly
what its window is for; the footage just stops cooperating. In the debug
overlay (which draws a bridged identity's last-known bbox, unlike the shipped
output, which emits presence windows and no boxes at all) the bridged state is
visible spatially:
![Debug overlay: bridged identities drawn at their last-known positions](assets/images/many_saints_ghost_fpi.jpg)
*Debug-overlay rendering (`dump_error_frames.py --raw`): "Jon Bernthal", "Joey
Diaz" and "Billy Magnussen" are extinction-bridged identities from the previous
shot, drawn frozen over the wall and the hanging plates. Frame
`many_saints/fpi/fpi_t03543.jpg`, `montage-frames` artifact package.*
The cost is measurable, not just visible. Downton Abbey's hard cut into its
closing credits, plotting the dump's own per-second `face_count` (detector
output, independent of the tracker) against what the tracker reports:
![Detector vs. tracker through Downton Abbey's cut to credits](assets/images/downton_ghost_timeline.png)
From the cut onward the detector sees **zero faces for nearly a minute** — and
the tracker keeps reporting the last shot's 15 identities the whole time
(verified for Hugh Bonneville: bbox `(1743.2, 0.0, 171.3, 317.8)`, unchanged to
the pixel, at every sampled second for 57+ seconds). The staircase at the right
edge is the extinction window expiring actor by actor. That plateau is
`SceneTrackerFunc::active_[actor_idx].last_bbox`
([`src/nodes/scene_tracker_node.hpp`](https://REPOLINK/src/nodes/scene_tracker_node.hpp))
re-emitted as designed: `extinction_sec=57.4` was tuned long because bridging
wins on most footage (see the polygraph frame above) — the training films just
never contained a faceless stretch long enough to show the cost side, and the
held-out set did.
The same track-continuation machinery has one milder spatial artifact, worth
knowing when reading these frames:
![Two labels on one face after a shot/reverse-shot cut](assets/images/cafe_society_rapid_cut.jpg)
*Café Society (a training film), a shot/reverse-shot dialog: that is Steve
Carell wearing both his own label and Jesse Eisenberg's.*
At a rapid cut, the previous shot's track can linger for a beat at nearly the
same screen position the new face occupies — here Jesse Eisenberg's box from
the counter-shot lands on Steve Carell. Note what the score panel says,
though: both actors are green, because both *are* present in this dialog
scene per X-Ray. The spatial label is briefly wrong; the per-second presence
claim — the thing the pipeline actually ships — is right. It's the same trade
as the extinction window: track continuation smooths over cuts, and 1 fps
sampling occasionally catches the seam.
## Mechanism 2: the face-vs-presence ceiling
Downton Abbey's recall didn't collapse because faces were misread — it
collapsed because for most of its 80084 FN-seconds there was **no face to
read**:
![22 cast credited, nobody facing the camera](assets/images/downton_crew_fn.jpg)
A newsreel crew hauls equipment through the hall: X-Ray credits 22 cast as
present in this scene; not one face looks at the camera. Eight are still
scored green (windows bridging from adjacent shots) — the other fourteen are
blue FNs that no face-recognition pipeline could ever recover. X-Ray encodes
*scene membership*; the pipeline measures *on-screen faces*. In ensemble films
those two definitions diverge massively, and that gap — not identification
error — is most of what the FN column counts.
![Presence without a detectable face, The Many Saints of Newark](assets/images/many_saints_outofcast_fpi.jpg)
Same ceiling from the other side: Michela De Rossi in frame but turned away,
five cast correctly bridged as offscreen (green), four blue FNs — and one
orange we'll come back to below.
## Where LVFace beat X-Ray
Not every orange in these frames is actually wrong.
[`scripts/optimizer/second_score.py`](https://REPOLINK/scripts/optimizer/second_score.py)
scores strictly against X-Ray — but X-Ray itself has holes, and the pipeline
found two kinds.
![LVFace correctly identifies Germar Terrell Gardner, uncredited by X-Ray](assets/images/germar_beats_xray.jpg)
Germar Terrell Gardner — a real, clean, high-confidence detection — is counted
as an out-of-cast misID because he doesn't appear in X-Ray's `people.csv` for
The Many Saints of Newark at all. But Jellyfin's independent cast metadata
*does* credit him for this exact film (cross-checked via
`experiments/manifests/jellyfin_casts.json` from the `experiment-data` artifact
package, a completely separate data source from X-Ray). That's also him in
orange in the frame above — every one of those "errors" is the pipeline being
right about a person X-Ray forgot.
![Robert Patrick, clearly on screen, scored wrong by a ground-truth gap](assets/images/lovelace_robert_patrick_fpi.jpg)
And it isn't only uncredited bit-parts. That is **Robert Patrick** — top-billed
in Lovelace, unmistakably on screen, reading his newspaper, identified at
100% — scored orange because X-Ray's people-in-scene list for *this scene*
doesn't include him. The identification is flawless; the ground truth missed
an actor sitting in the middle of the frame.
This doesn't mean every flagged misID is secretly correct — Many Saints'
974-count total is still overwhelmingly extinction bridging at cuts, not
uncredited cameos. But the X-Ray corpus is a convenient, large-scale ground
truth, not a perfect one, and the misID/FPI numbers in these tables carry an
irreducible noise floor from ground-truth gaps in both directions.
## Summary
LVFace is the right default: it wins the model comparison outright, it names
19 of 20 correctly across a hat-heavy funeral crowd, and it recognises a face
on a screen inside the movie. Its error budget decomposes into two understood
mechanisms — extinction bridging at hard cuts (a tunable trade, not a bug) and
the face-vs-presence ceiling baked into X-Ray's semantics — plus a nonzero
slice where the pipeline is right and the ground truth is wrong. The held-out
generalization gap (75.3% → 67.4%) is real and should be treated as the honest
expected performance, not the training-set number.