docs: archive the July 2026 report; new methodology for the opencv5 run
The July report (4-model ArcFace/LVFace bake-off, pre-opencv5 framework, 3-film training + held-out validation) is superseded by the opencv5 build: single-model LVFace-B, a 6-knob DE sweep over all 9 films, flood-fill presence, and the registry/decode fixes. Rather than overwrite it, archive it date-suffixed and start the current report fresh. - Rename the six July result pages to *-2026-07.md, rewrite their intra-archive cross-links, and add an "Archived (July 2026)" banner to each. - mkdocs nav: current report at top, the July set under an Archive section. - New docs/methodology.md for the opencv5 run: corrects the withdrawn anneal_sec/extinction_sec presence bridging (windows are now [first_seen, last_seen], AR-012/013), documents the two presence modes (track_extent / flood), and records that every eval scores all 9 films. The current experiment log (model-bakeoff.md) and Home rewrite land once the DE sweep converges and the final optimum is known.
This commit is contained in:
@@ -0,0 +1,112 @@
|
||||
> **Archived (July 2026).** This report covers the pre-opencv5 framework and the 4-model ArcFace/LVFace bake-off. It is superseded by the current [experiment log](model-bakeoff.md) for the opencv5 build. Kept for provenance; the numbers here are historical.
|
||||
|
||||
# Pose expansion: does promoting new poses mid-film help?
|
||||
|
||||
`expand_gallery`
|
||||
([`src/gallery/track_gallery.hpp`](https://REPOLINK/src/gallery/track_gallery.hpp))
|
||||
promotes a confidently identified track's novel-pose reference views into a
|
||||
per-film, in-memory gallery annex. The idea: once the pipeline is confident
|
||||
about an identity, a pose it has not seen before (turned head, different
|
||||
lighting) becomes an extra reference for recognizing that actor again later
|
||||
in the same film, without touching the baked gallery.
|
||||
|
||||
## Training-set signal
|
||||
|
||||
Averaged across the 3 compared models (r50 excluded), on the 4 films used
|
||||
for optimization. These are the corrected, full-coverage figures, see the
|
||||
[dropped-film note](model-bakeoff-2026-07.md#a-scoring-bug-worth-recording-dropped-film-evaluations)
|
||||
in the experiment log for why an earlier version of this table overstated the
|
||||
full-mode misID jump (209 → 864) that was itself partly a truncation artifact:
|
||||
|
||||
| scope | expansion | F1 | R | misID |
|
||||
|---|---|---|---|---|
|
||||
| full | off | 70.0% | 57.6% | 407 |
|
||||
| full | on | 72.1% | 61.5% | 714 |
|
||||
| restricted | off | 75.1% | 63.9% | 179 |
|
||||
| restricted | on | 76.7% | 67.2% | 120 |
|
||||
|
||||
In restricted mode, expansion looks like a clean win: +1.6pp F1, +3.3pp
|
||||
recall, lower misID. In full mode it looks like a recall-for-misID trade:
|
||||
+2.1pp F1, +3.9pp recall, but misID rises from 407 to 714. See
|
||||
[the full experiment log](model-bakeoff-2026-07.md) for the per-model breakdown.
|
||||
This asymmetry motivated the question below: does turning expansion on
|
||||
change what gets recognized frame by frame, or is the aggregate F1 shift
|
||||
coming from something else.
|
||||
|
||||
## Held-out test
|
||||
|
||||
Same model, same tuned config, `expand_gallery` toggled on vs. off, nothing
|
||||
else changed, full gallery mode, per-second scoring against X-Ray. This
|
||||
isolates expansion from every other variable that differs between the
|
||||
training-set rows above.
|
||||
|
||||
LVFace-B Glint360K, all 5 held-out films:
|
||||
|
||||
| film | F1 (exp) | F1 (noexp) | TPI delta | FN delta |
|
||||
|---|---|---|---|---|
|
||||
| Benny & Joon | 83.0% | 83.0% | -2 | +2 |
|
||||
| Downton Abbey: A New Era | 56.1% | 56.2% | -7 | +7 |
|
||||
| Lovelace | 77.5% | 77.4% | +33 | -33 |
|
||||
| The Many Saints of Newark | 46.3% | 46.3% | +2 | -2 |
|
||||
| Valerian and the City of a Thousand Planets | 74.1% | 74.1% | +2 | -2 |
|
||||
|
||||
ArcFace R18, Benny & Joon, r18's own tuned config: F1 77.1% for both, TPI
|
||||
and FN identical, FPI differs by 2.
|
||||
|
||||
Every film, both models tested: F1 differs by 0.1-0.2pp, TPI/FN swings are
|
||||
in the tens out of tens of thousands. This is noise, not a signal.
|
||||
Expansion made no measurable difference to per-second on-screen
|
||||
identification on any held-out film tested.
|
||||
|
||||
## Two methodology bugs caught during this check
|
||||
|
||||
Getting to the table above required catching two wrong turns, both worth
|
||||
recording because they are exactly the kind of error that produces a false
|
||||
positive "expansion helped" finding.
|
||||
|
||||
1. **Timeout truncation.** The first Downton Abbey `exp` replay was cut off
|
||||
by a 60-second subprocess timeout at about 76% through the film (5589 of
|
||||
7368 expected seconds). This silent data loss produced a large,
|
||||
convincing-looking TPI gap (47938 vs 52032) purely because one run was
|
||||
missing a quarter of the film. Caught by comparing `n_seconds` between
|
||||
runs before trusting any score delta; fixed by re-running with a longer
|
||||
timeout.
|
||||
2. **Bbox-matching bug.** An early per-second raw-annotation diff matched
|
||||
each `exp` detection to the first `noexp` detection with IoU above 0.5,
|
||||
not the best-overlapping one. With 3 faces close together in frame, this
|
||||
produced spurious disagreements (for example "exp says Aidan Quinn,
|
||||
noexp says Johnny Depp" at the same second) that vanished once the match
|
||||
used the best-IoU candidate instead of the first one. Both configs had
|
||||
actually output the same three names at the same three boxes.
|
||||
|
||||
Both bugs independently pointed toward "expansion is doing something," and
|
||||
both were artifacts of the comparison harness, not the pipeline. Before
|
||||
trusting a dramatic before/after diff, check that both runs cover the same
|
||||
seconds and that entities are matched by best overlap, not first found.
|
||||
|
||||
## Conclusion
|
||||
|
||||
The training-set aggregate effect, particularly the full-mode misID
|
||||
increase, does not reproduce on held-out data. At minimum it
|
||||
is far smaller than the training-set numbers suggested; it may be sampling
|
||||
variation from only 4 training films rather than a generalizable
|
||||
mechanism. Note the same *class* of harness bug appears twice in this
|
||||
investigation, the timeout truncation in bug #1 above, and the dropped-film
|
||||
aggregation that inflated the raw training-set misID figures. Both make an
|
||||
inert config look consequential; both are reasons to distrust a dramatic
|
||||
training-set delta until it survives on held-out films, which this one did
|
||||
not. This does not mean `expand_gallery` never does anything: the
|
||||
mechanism is real, and
|
||||
[`track_gallery.hpp`](https://REPOLINK/src/gallery/track_gallery.hpp)'s
|
||||
promotion logging confirms tracks get confirmed and views get promoted
|
||||
into the annex on every film tested. It means whatever effect expansion
|
||||
has on final per-second identification was too small to detect against 5
|
||||
held-out films with this scoring method. A cleaner test would need either
|
||||
more held-out films or a metric that can see the annex's direct
|
||||
contribution, such as tagging which reference embedding won each match;
|
||||
neither was in scope for this pass.
|
||||
|
||||
Do not treat the training-set exp/noexp numbers in
|
||||
[the full experiment log](model-bakeoff-2026-07.md) as proof that expansion changes
|
||||
real-world behavior in either direction. On the evidence gathered so far,
|
||||
it does not move the needle enough to see.
|
||||
Reference in New Issue
Block a user