The July report (4-model ArcFace/LVFace bake-off, pre-opencv5 framework, 3-film training + held-out validation) is superseded by the opencv5 build: single-model LVFace-B, a 6-knob DE sweep over all 9 films, flood-fill presence, and the registry/decode fixes. Rather than overwrite it, archive it date-suffixed and start the current report fresh. - Rename the six July result pages to *-2026-07.md, rewrite their intra-archive cross-links, and add an "Archived (July 2026)" banner to each. - mkdocs nav: current report at top, the July set under an Archive section. - New docs/methodology.md for the opencv5 run: corrects the withdrawn anneal_sec/extinction_sec presence bridging (windows are now [first_seen, last_seen], AR-012/013), documents the two presence modes (track_extent / flood), and records that every eval scores all 9 films. The current experiment log (model-bakeoff.md) and Home rewrite land once the DE sweep converges and the final optimum is known.
16 KiB
Archived (July 2026). This report covers the pre-opencv5 framework and the 4-model ArcFace/LVFace bake-off. It is superseded by the current experiment log for the opencv5 build. Kept for provenance; the numbers here are historical.
Deep dive: LVFace-B Glint360K
LVFace won the model comparison (see Which model is best?) and is the shipped default embedder. This page reports how it performs in detail: a baseline of correct output, the two mechanisms behind its errors, and every distinct case where it names someone who is not in the film's credited cast.
Read How we score against X-Ray first. X-Ray's ground truth is scene-level, not per-frame. A name marked correct in the Offscreen column below is the pipeline correctly reporting scene membership, not a workaround.
!!! note "How to read the frames on this page"
The top of each image is the film frame, with a box and name on every
face the pipeline matched to a real detection. The panels below are the
per-second result against X-Ray. **Onscreen** lists names attached to a
visible face this second. **Offscreen** lists names the pipeline reports
present without a currently visible face. Colors mark the verdict:
<span style="color:#0ca30c">**green**</span> correct (TPI),
<span style="color:#eb6834">**orange**</span> wrong (FPI),
<span style="color:#3987e5">**blue**</span> missed (FN).
Baseline: correctly scored seconds
Six faces on screen, all six named correctly, including Penelope Wilton at the edge of the pews and a partly occluded Michelle Dockery. Thirteen more cast members X-Ray lists as present in the scene are correctly reported Offscreen. One miss: Maggie Smith (blue). Score for this second: 0.86.
The same film's funeral scene: dark clothing, hats, half the faces turned away. Nineteen of the twenty cast members X-Ray lists for this scene score correct: seven named on screen at up to 100% confidence, twelve more reported correctly as present but not visible.
The pipeline does not require a live face. This is Herbie Hancock at 98% confidence, identified from a face displayed on a screen inside the film, on a video call under a science-fiction HUD overlay.
Training vs. held-out: the generalization gap
The shipped config (prob_threshold=0.754, anneal_sec=35.54,
extinction_sec=57.43, expand_gallery=true) was tuned on 4 films. Scored
on the 5 films the optimizer never saw:
| film | F1 | P | R | TPI | FPI | misid | FN |
|---|---|---|---|---|---|---|---|
| Benny & Joon | 83.0% | 89.1% | 77.7% | 15125 | 1846 | 0 | 4337 |
| Lovelace | 77.5% | 90.3% | 67.9% | 14990 | 1085 | 58 | 7085 |
| Valerian and the City of a Thousand Planets | 74.1% | 97.1% | 60.0% | 18663 | 548 | 0 | 12467 |
| Downton Abbey: A New Era | 56.2% | 97.8% | 39.4% | 52027 | 1173 | 0 | 80084 |
| The Many Saints of Newark | 46.3% | 54.7% | 40.1% | 15922 | 4394 | 974 | 23791 |
| macro average | 67.4% | 85.8% | 57.0% |
The P column is misID-weighted (each out-of-film name counts 10x in the
denominator; see methodology).
That weighting is why Many Saints reads 54.7% here despite naming mostly real,
present faces: its raw (unweighted) precision is 78.4%, and the gap is
entirely its 974 misIDs paying the 10x penalty. The three zero-misID films
(Benny & Joon, Downton, Valerian) have identical weighted and raw precision;
Lovelace, with 58 misIDs, sits 3pp below its raw 93.3%.
Held-out F1 is 67.4%, against 75.3% on training, an 8pp drop. The spread between the best and worst held-out film is 37pp. This is not unique to LVFace: the full experiment log shows mbf and r18 with the same shape of spread on the same films, at a uniformly lower level. Two mechanisms explain the spread. Both are shown below with frame-level evidence.
Mechanism 1: extinction bridging
The extinction window keeps a name reported as present for up to
extinction_sec after its last real detection. This is deliberate: most
gaps in face visibility are short (a turned head, an occlusion, a cut to a
reaction shot), and the window bridges them.
Lovelace's polygraph scene: only Eric Roberts and Amanda Seyfried have visible faces. X-Ray lists eight cast members present. All eight score correct; the other six are reported Offscreen through a stretch where the camera never shows them. The extinction window is why.
The same mechanism fails at a hard cut into a long stretch with no faces at
all. Downton Abbey's recall (39.4%, the worst of the five held-out films) is
dominated by this failure. It is verified directly against the raw
per-frame stream and the dump's own detection counts, not inferred from the
score. Plotting the dump's per-second face_count (detector output,
independent of the tracker) against what the tracker reports, through
Downton Abbey's hard cut into its closing credits:
From the cut onward the detector reports zero faces for close to a minute.
The tracker continues reporting the previous shot's 15 identities for the
same span (verified for Hugh Bonneville: bbox (1743.2, 0.0, 171.3, 317.8),
unchanged to the pixel, at every sampled second for 57 seconds). The
staircase at the right edge is the extinction window expiring, actor by
actor. This is SceneTrackerFunc::active_[actor_idx].last_bbox
(src/nodes/scene_tracker_node.hpp)
re-emitted as designed. extinction_sec=57.4 was tuned long because
bridging is correct on most footage, as in the polygraph scene above. The
training films did not contain a faceless stretch long enough to expose the
cost side; the held-out set did.
The extinction window is a scoring concept, not something drawn on screen. The shipped output is presence windows with no bounding boxes. Even the debug overlay used for this report never draws a box for a bridged name: a name inside its extinction window with no current detection appears only as a name in the Offscreen column, the same as every correctly bridged name above.
A related, smaller effect shows up at rapid cuts:
Café Society (a training film), a shot/reverse-shot dialog. The box on Steve Carell's face carries two labels: his own, and Jesse Eisenberg's, left over from the counter-shot a moment earlier. Both names score correct, because both actors are present in this scene per X-Ray. The box position is briefly wrong; the presence claim, which is what the pipeline ships, is right.
Mechanism 2: the face-vs-presence ceiling
Downton Abbey's recall did not collapse because faces were misread. It collapsed because for most of its 80084 false-negative seconds there was no face to read.
A newsreel crew moves equipment through the hall. X-Ray credits 22 cast members as present in this scene. None face the camera. Eight still score correct, carried by presence windows from adjacent shots. The other fourteen are missed, and no face-recognition system can recover them, because there is no face in the frame. X-Ray records scene membership; the pipeline measures visible faces. In ensemble scenes these two quantities diverge, and that gap accounts for most of the false-negative count.
Every distinct out-of-cast name
Many Saints of Newark has the largest misID count of any held-out film: 974 seconds, weighted. Rather than characterize this from a single frame, the raw replay stream was searched directly for every name the pipeline reports that is not in the film's credited cast. The same search was run on all 9 films in the benchmark, one rule applied uniformly: find the first second each distinct out-of-cast name appears, and render that exact second.
Five films produce no such name anywhere in their runtime: Benny & Joon, Café Society, Downton Abbey, Sound of Metal, Valerian. Zero out-of-cast names across their entire length. Four films produce nine distinct names between them, shown below in full, not a sample.
The Many Saints of Newark: 4 names
Germar Terrell Gardner, t=848s, 78% confidence. A real, clearly visible background actor. He is not in X-Ray's cast list for this film, but he is credited in Jellyfin's independent cast metadata (see Where LVFace beat X-Ray below). This is a ground-truth gap, not a model error.
Archie Yates, t=2521s, 78% confidence. A real detected face, a genuine lookalike confusion.
Zooey Deschanel, t=2819s, 99% confidence. A real detected face at a dinner table, high-confidence lookalike confusion.
Talia Balsam, t=4551s, 93% confidence. A real detected face. Talia Balsam
plays Mrs. Jarecki, a guidance counselor, in this film; she is confirmed
on screen by direct inspection of the frame. She does not appear in X-Ray's
people.csv for this title. This is a second ground-truth gap in the same
film, not a model error.
Two of these four names are ground-truth gaps (Gardner, Balsam), not misidentifications. The other two (Yates, Deschanel) are genuine embedding errors on real faces.
Lord of War: 3 names
David Shumbris, t=418s, 81% confidence. A real face in a dim, low-detail shot under a train track. A genuine lookalike confusion in poor lighting.
Ronald Reagan, t=1003s, 100% confidence. This is not a lookalike confusion. The detected face is a photograph of Reagan appearing within the shot, not a living actor. The detector and matcher both did their job correctly on the image content in front of them; the error is that a photograph inside the scene is not the same thing as an actor present in the scene, and the pipeline has no way to draw that distinction from a face crop alone.
Lance Reddick, t=6424s, 78% confidence. A small, distant, low-detail face at the edge of frame. A marginal, low-confidence lookalike confusion.
Lovelace: 1 name
Chloë Sevigny, t=2451s, 100% confidence. Amanda Seyfried's track is real and well-tracked through most of this shot, but her bbox is frozen at the exact same coordinates for t=2450 and t=2451, one second where her box stopped updating from a fresh detection. Only one real face is detected at t=2451 (confirmed against the dump's own per-frame detections), and it is a tight IoU-1.0 fit under the Chloë Sevigny box, not the Seyfried one. So the green Seyfried box in this frame is a ghost, re-emitting her last known position for that one second, and the fresh, wrong detection is Sevigny, landing on top of it. Not two competing fresh identities on one crop: one ghost and one fresh misidentification happening to overlap.
Scarface: 1 name
Kirstie Alley, t=2451s, 89% confidence. Al Pacino is correctly identified in
the foreground at 100%; a background face in the same shot is wrongly
labeled Kirstie Alley. (The t=2451s here and the Lovelace Chloë Sevigny case
above landing on the identical second is a genuine coincidence, verified from
each film's raw stream by first_fpi_frames.py,
not a transcription slip, two unrelated films whose first out-of-cast name
happens to fall at the same timestamp.)
Summary of the nine
| film | name | t (s) | confidence | classification |
|---|---|---|---|---|
| Many Saints of Newark | Germar Terrell Gardner | 848 | 78% | ground-truth gap |
| Many Saints of Newark | Archie Yates | 2521 | 78% | lookalike confusion |
| Many Saints of Newark | Zooey Deschanel | 2819 | 99% | lookalike confusion |
| Many Saints of Newark | Talia Balsam | 4551 | 93% | ground-truth gap |
| Lord of War | David Shumbris | 418 | 81% | lookalike confusion |
| Lord of War | Ronald Reagan | 1003 | 100% | photo-in-frame |
| Lord of War | Lance Reddick | 6424 | 78% | lookalike confusion, marginal |
| Lovelace | Chloë Sevigny | 2451 | 100% | lookalike confusion |
| Scarface | Kirstie Alley | 2451 | 89% | lookalike confusion |
Of nine distinct out-of-cast names across four films, two are ground-truth gaps, one is a photograph misread as a person, and six are genuine embedding-space confusions on real detected faces. None trace to extinction bridging: every one of these nine is a fresh detection on a real face crop at the second it first appears.
Where LVFace beat X-Ray
Not every name marked wrong is actually wrong.
scripts/optimizer/second_score.py
scores strictly against X-Ray, and X-Ray has gaps of its own.
Germar Terrell Gardner, the same name from the table above, does not appear
in X-Ray's people.csv for The Many Saints of Newark. Jellyfin's
independent cast metadata does credit him for this film (cross-checked
against experiments/manifests/jellyfin_casts.json from the
experiment-data artifact package, a data source entirely separate from
X-Ray). Talia Balsam is the same case: confirmed on screen, absent from
X-Ray's cast list for this title.
This extends past uncredited background actors. This is Robert Patrick, top-billed in Lovelace, clearly on screen reading a newspaper, identified at 100%. The frame is scored wrong because X-Ray's people-in-scene list for this specific scene omits him, despite crediting him elsewhere in the film. The identification is correct; the ground truth is missing an entry.
X-Ray is a large, convenient ground truth. It is not a complete one. The misID and FPI counts reported throughout this document include some fixed amount of noise from gaps in X-Ray itself, in both directions.
Summary
LVFace wins the model comparison on every held-out film. It correctly names 19 of 20 people in a crowded funeral scene and correctly identifies a face displayed on a screen inside the film. Its errors resolve into two mechanisms: extinction bridging, which is correct on most footage and fails specifically at hard cuts into long faceless stretches, and the face-versus-presence ceiling, where X-Ray credits scene membership for people whose faces never appear on screen. Of the nine distinct out-of-cast identifications found across the benchmark, two trace to gaps in X-Ray's own cast data, one is a photograph misread as a person, and six are genuine lookalike confusions on real faces. The held-out generalization gap, 75.3% training to 67.4% held-out, is real and should be treated as the expected operating point, not the training-set figure.


















