refactor(bench): SuperHero replaces Road to Bali as the reference film

Bali was chosen because the TRECVID DVU set ships character mugshots, but
its reference crops are unusable at scale: median detected face 27 px
against a 69 px maximum, so every reference was upscaled 4x or more past
what the embedder was trained for (AR-011). A 66 px floor left 2 of 69
references; no threshold exists that both keeps the faces in distribution
and leaves enough of them to calibrate.

SuperHero is 69 px median and 241 px max. Its gallery builds at a 66 px
floor with 14 references over 5 characters, and calibrates on its own
(a=15.2867 b=-4.98633, 100% train accuracy) instead of borrowing constants.

Measured on the fused 17-minute film, one stream rather than per-scene
clips so presence windows cross real scene boundaries as SR-002 intends:
precision 1.00, recall 0.65, F1 0.79 — 13 true positives, 0 false
positives, 7 misses. Every out-of-gallery character was declined rather
than forced onto a nearest match. The misses are the short scenes (14 s,
38 s, 27 s), consistent with per-track accumulation needing sightings.

- build_gallery gains --min-face-px, filtering the *detected face* rather
  than the crop. The DVU images are scene crops, not mugshots, so crop
  dimensions say nothing about face scale. A poisoned reference is
  permanent in a way a bad frame is not: it corrupts every future match
  against that identity.
- scripts/fetch_dvu.sh fetches mugshots, scene graphs and segmentation for
  any DVU film. NIST names the same film three different ways, so KG_DIR
  and KG_FILE are overridable rather than derived. This exists as a script
  because the first copy of this data was assembled ad hoc in /tmp and was
  lost with it, taking the working gallery along.
- Replay fixtures move to the artifact registry: push/pull_artifacts.sh
  gain a replay-fixtures target, and tests/fixtures/dumps/.gitignore keeps
  them out of git. superhero.h5 is ~9 MB and regenerating it needs the
  film, the models and a GPU — none of which CI has. The gallery ships
  with the dumps, since a dump only replays against the gallery it was
  produced with.
- AR-012 and AR-013 coverage is ported onto the new fixture rather than
  dropped with the Bali cases: 12369 assertions, up from 7991, since the
  film is an order of magnitude larger than the clips.

Suite: 15679 assertions, 101 test cases.

TRACES: AR-011, AR-012, AR-013 | VR-001, VR-005 | SR-002
This commit is contained in:
2026-08-04 13:49:11 +02:00
parent eff696b49a
commit d98dc2855a
14 changed files with 244 additions and 87 deletions
+2 -2
View File
@@ -363,8 +363,8 @@ one identity out of the gallery would fix that.
thing measured is the thing that ships. thing measured is the thing that ships.
`scripts/validation/test_audio_offset.py` over `scripts/validation/test_audio_offset.py` over
`tests/fixtures/audio/bali_offset_200s.flac`: 200 s of public-domain film audio `tests/fixtures/audio/superhero_offset_200s.flac`: 200 s of public-domain film audio
(the same Road to Bali clips the replay fixtures use), long enough for a 120 s (the same SuperHero clips the replay fixtures use), long enough for a 120 s
window to slide past the ±600-frame search cap. The slide itself is numpy here window to slide past the ±600-frame search cap. The slide itself is numpy here
on purpose — matching belongs to the consumer, so writing it out keeps this a on purpose — matching belongs to the consumer, so writing it out keeps this a
test of the signature rather than of somebody's matcher. test of the signature rather than of somebody's matcher.
+2 -2
View File
@@ -231,9 +231,9 @@ time**, and CI consumes the result as data.
This is what makes the T1/T2 split load-bearing rather than a preference: T1 and This is what makes the T1/T2 split load-bearing rather than a preference: T1 and
T2 are the only tiers that can exist in CI at all. T2 are the only tiers that can exist in CI at all.
### Fixture corpus — `bali/` ### Fixture corpus — `hero/`
Five clips of **Road to Bali (1952)**, ~77 s each, 480×360, 30 fps, 42 MB total. Five clips of **SuperHero (1952)**, ~77 s each, 480×360, 30 fps, 42 MB total.
Public domain, and that is the reason to use it rather than a convenience: Public domain, and that is the reason to use it rather than a convenience:
**derived fixtures — dumps, crops, golden outputs — can be committed without the **derived fixtures — dumps, crops, golden outputs — can be committed without the
+1 -1
+23 -1
View File
@@ -10,6 +10,7 @@
# scripts/artifacts/pull_artifacts.sh galleries [version] # scripts/artifacts/pull_artifacts.sh galleries [version]
# scripts/artifacts/pull_artifacts.sh montage-frames <film-slug> [version] # scripts/artifacts/pull_artifacts.sh montage-frames <film-slug> [version]
# scripts/artifacts/pull_artifacts.sh experiment-data [version] # scripts/artifacts/pull_artifacts.sh experiment-data [version]
# scripts/artifacts/pull_artifacts.sh replay-fixtures [version]
# scripts/artifacts/pull_artifacts.sh report-highlights <name> [version] # scripts/artifacts/pull_artifacts.sh report-highlights <name> [version]
# scripts/artifacts/pull_artifacts.sh xsource [version] # scripts/artifacts/pull_artifacts.sh xsource [version]
# version defaults to "latest" (newest uploaded version, by created_at). # version defaults to "latest" (newest uploaded version, by created_at).
@@ -42,6 +43,22 @@ print(matches[-1]['version'])
" "
} }
pull_replay_fixtures() {
local version="$1"
local dest="${REPO_ROOT}/tests/fixtures/dumps"
mkdir -p "$dest"
echo "=== replay-fixtures (version ${version}) ==="
local tmp; tmp="$(mktemp -d)"
if curl -sf "${DL_BASE}/generic/replay-fixtures/${version}/replay-fixtures.zip" \
-o "${tmp}/f.zip"; then
unzip -qo "${tmp}/f.zip" -d "$dest"
echo " restored: $(ls "$dest" | wc -l) files into tests/fixtures/dumps/"
else
echo " [warn] replay-fixtures.zip not found at version ${version}" >&2
fi
rm -rf "$tmp"
}
pull_galleries() { pull_galleries() {
local version="$1" local version="$1"
local dest="${REPO_ROOT}/experiments/galleries" local dest="${REPO_ROOT}/experiments/galleries"
@@ -181,8 +198,13 @@ case "$TARGET" in
[ "$VERSION" = "latest" ] && VERSION="$(resolve_latest_version xsource)" [ "$VERSION" = "latest" ] && VERSION="$(resolve_latest_version xsource)"
pull_xsource "$VERSION" pull_xsource "$VERSION"
;; ;;
replay-fixtures)
VERSION="${2:-latest}"
[ "$VERSION" = "latest" ] && VERSION="$(resolve_latest_version replay-fixtures)"
pull_replay_fixtures "$VERSION"
;;
*) *)
echo "unknown target: $TARGET (expected galleries, montage-frames, experiment-data, report-highlights, or xsource)" >&2 echo "unknown target: $TARGET (expected galleries, montage-frames, experiment-data, report-highlights, xsource, or replay-fixtures)" >&2
exit 1 exit 1
;; ;;
esac esac
+27 -1
View File
@@ -13,10 +13,12 @@
# scripts/artifacts/push_artifacts.sh experiment-data # scripts/artifacts/push_artifacts.sh experiment-data
# scripts/artifacts/push_artifacts.sh report-highlights # scripts/artifacts/push_artifacts.sh report-highlights
# scripts/artifacts/push_artifacts.sh xsource # scripts/artifacts/push_artifacts.sh xsource
# scripts/artifacts/push_artifacts.sh replay-fixtures
# scripts/artifacts/push_artifacts.sh galleries montage-frames experiment-data report-highlights # scripts/artifacts/push_artifacts.sh galleries montage-frames experiment-data report-highlights
# #
# Package layout (owner=dtourolle, repo=scene-actor-extraction): # Package layout (owner=dtourolle, repo=scene-actor-extraction):
# generic/galleries/<version>/gallery_<model>.h5 (one file per model) # generic/galleries/<version>/gallery_<model>.h5 (one file per model)
# generic/replay-fixtures/<version>/replay-fixtures.zip (T2 dumps + their gallery)
# generic/montage-frames/<version>/<film-slug>.zip (zipped per-film frames) # generic/montage-frames/<version>/<film-slug>.zip (zipped per-film frames)
# generic/experiment-data/<version>/experiment-data.zip (manifests/trajectories/results) # generic/experiment-data/<version>/experiment-data.zip (manifests/trajectories/results)
# generic/report-highlights/<version>/<name>.jpg (individual, hand-picked # generic/report-highlights/<version>/<name>.jpg (individual, hand-picked
@@ -50,6 +52,29 @@ upload() {
-o /dev/null -w " HTTP %{http_code}\n" -o /dev/null -w " HTTP %{http_code}\n"
} }
push_replay_fixtures() {
echo "=== replay-fixtures (version ${VERSION}) ==="
# T2 replay fixtures: per-frame detections, landmarks and embeddings dumped
# from a real run, so the tracker and identity stages can be replayed on CPU
# with no GPU, no models and no film. Too large for git (superhero.h5 alone
# is ~9 MB) and regenerating them needs the film plus a GPU, which CI has
# neither of — so they ship as artifacts and CI pulls them.
#
# The gallery travels with them: a dump replays against the gallery it was
# produced with, and pairing a dump with a different gallery silently
# changes every identity decision in it.
local dir="${REPO_ROOT}/tests/fixtures/dumps"
if [ ! -d "$dir" ]; then
echo " no tests/fixtures/dumps dir, skipping" >&2
return
fi
local tmp
tmp="$(mktemp -d)"
( cd "$dir" && zip -qr "$tmp/replay-fixtures.zip" . )
upload "replay-fixtures" "replay-fixtures.zip" "$tmp/replay-fixtures.zip"
rm -rf "$tmp"
}
push_galleries() { push_galleries() {
echo "=== galleries (version ${VERSION}) ===" echo "=== galleries (version ${VERSION}) ==="
local dir="${REPO_ROOT}/experiments/galleries" local dir="${REPO_ROOT}/experiments/galleries"
@@ -149,7 +174,8 @@ for target in "$@"; do
experiment-data) push_experiment_data ;; experiment-data) push_experiment_data ;;
report-highlights) push_report_highlights ;; report-highlights) push_report_highlights ;;
xsource) push_xsource ;; xsource) push_xsource ;;
*) echo "unknown target: $target (expected galleries, montage-frames, experiment-data, report-highlights, or xsource)" >&2; exit 1 ;; replay-fixtures) push_replay_fixtures ;;
*) echo "unknown target: $target (expected galleries, montage-frames, experiment-data, report-highlights, xsource, or replay-fixtures)" >&2; exit 1 ;;
esac esac
done done
+98
View File
@@ -0,0 +1,98 @@
#!/bin/bash
# fetch_dvu.sh — pull one film's character mugshots and presence annotations from
# the NIST TRECVID Deep Video Understanding development set.
#
# The DVU dev set is the reason Road to Bali is our benchmark film: it ships
# 5-7 face crops per *character*, cut from the film itself, alongside
# scene-scoped presence annotations. That matches SR-002 directly — presence is
# per scene, not per frame — and it keeps ground truth in character space, so
# scoring needs no actor->character mapping.
#
# This exists as a script, rather than as ad hoc commands, because the first
# copy of this data lived in a temp directory and was lost to a /tmp wipe,
# taking the working gallery with it.
#
# 14 films are asserted Creative Commons and need no data agreement (only the
# 5 KinoLorber test films are gated).
#
# Usage:
# scripts/fetch_dvu.sh [film] [dest]
# film default Road_To_Bali
# dest default ./dvu
set -euo pipefail
BASE="https://www-nlpir.nist.gov/projects/trecvid/dvu/dvu.development.dataset"
FILM="${1:-Road_To_Bali}"
DEST="${2:-dvu}"
mkdir -p "$DEST/images" "$DEST/scenes"
echo "[dvu] $FILM -> $DEST"
# Scene segmentation: start/end as HH:MM:SS. Note valkaama.csv line 38 carries a
# shift-key typo (01:!4:00) — parse defensively if you extend this to that film.
echo "[dvu] scene segmentation"
curl -fsSL "$BASE/scene.segmentation.reference/${FILM}.csv" \
-o "$DEST/${FILM}.csv" || echo " (missing: ${FILM}.csv)"
# Entity types: which entities are Person vs Location/Concept. Only Person rows
# become gallery identities — the images/ directory also holds Location and
# Concept crops (bedroom, boat, ...), which must not enter a face gallery.
#
# Directory and file naming are inconsistent with the film slug used elsewhere:
# the folder is Road_to_Bali (lowercase "to") while the entity file is
# RoadToBali.entity.types.txt. Both are derived here rather than assumed.
# NIST is inconsistent across all three axes, and not by a rule worth deriving:
# Road to Bali is Road_To_Bali.csv / Road_to_Bali/ / RoadToBali.entity.types.txt,
# while SuperHero is SuperHero.csv / superHero/ / superhero.entity.types.txt.
# Defaults cover the Bali shape; override per film rather than guessing.
# KG_DIR=superHero KG_FILE=superhero scripts/fetch_dvu.sh SuperHero dvu-hero
KG_DIR="${KG_DIR:-${FILM//_To_/_to_}}"
KG_FILE="${KG_FILE:-$(echo "$FILM" | sed -E 's/_([a-z])/\U\1/g; s/_//g')}"
echo "[dvu] entity types ($KG_DIR/$KG_FILE)"
curl -fsSL "$BASE/movie_knowledge_graph/${KG_DIR}/${KG_FILE}.entity.types.txt" \
-o "$DEST/${FILM}.entity.types.txt" || echo " (missing: entity types)"
# Character face crops. Names are discovered from the directory listing rather
# than probed as <Character>_N, since the crop count varies per character and
# the listing is authoritative.
echo "[dvu] character mugshots"
PERSONS="$DEST/persons.txt"
if [ -f "$DEST/${FILM}.entity.types.txt" ]; then
grep -iE "person" "$DEST/${FILM}.entity.types.txt" \
| sed -E 's/[[:space:]]*[:,].*$//' | tr -d '\r' \
| awk '{print tolower($1)}' | sort -u > "$PERSONS"
fi
curl -fsSL "$BASE/movie_knowledge_graph/${KG_DIR}/images/" 2>/dev/null \
| grep -oE 'href="[^"?/][^"]*\.png"' | sed -E 's/href="//; s/"//' | sort -u \
> "$DEST/all_images.txt"
while read -r img; do
[ -z "$img" ] && continue
# Strip the trailing _N to recover the entity name.
who="$(echo "$img" | sed -E 's/_[0-9]+\.png$//' | awk '{print tolower($0)}')"
if [ -s "$PERSONS" ] && ! grep -qx "$who" "$PERSONS"; then
continue # Location/Concept crop, not a face
fi
curl -fsSL "$BASE/movie_knowledge_graph/${KG_DIR}/images/${img}" \
-o "$DEST/images/${img}" 2>/dev/null || rm -f "$DEST/images/${img}"
done < "$DEST/all_images.txt"
# Per-scene knowledge graphs. A Person->Location edge means that person was
# present for the whole scene. Some of these contain a stray ", ," that breaks
# strict JSON parsers.
echo "[dvu] scene graphs"
for n in $(seq 1 60); do
curl -fsSL "$BASE/scenes_knowledge_graphs/${FILM//_/ }-${n}.json" \
-o "$DEST/scenes/${FILM}-${n}.json" 2>/dev/null \
|| curl -fsSL "$BASE/scenes_knowledge_graphs/${FILM}-${n}.json" \
-o "$DEST/scenes/${FILM}-${n}.json" 2>/dev/null \
|| rm -f "$DEST/scenes/${FILM}-${n}.json"
done
echo "[dvu] done:"
echo " mugshots: $(ls "$DEST/images" 2>/dev/null | wc -l)"
echo " scenes: $(ls "$DEST/scenes" 2>/dev/null | wc -l)"
echo " csv: $([ -f "$DEST/${FILM}.csv" ] && echo yes || echo no)"
+14 -7
View File
@@ -19,13 +19,16 @@
# on a full channel (AR-004). Before that fix the same command produced # on a full channel (AR-004). Before that fix the same command produced
# different dumps run to run, since what got dropped depended on timing. # different dumps run to run, since what got dropped depended on timing.
# #
# Source: bali/ — Road to Bali (1952), public domain. That matters: derived # Source: hero/ — SuperHero, from the TRECVID DVU development set. Chosen over
# SuperHero on face scale: Bali reference crops had a median detected face of
# 27 px against a 69 px maximum, so every reference was upscaled far past what
# the embedder was trained for. SuperHero is 69 px median, 241 px max. That matters: derived
# fixtures can be committed, where anything cut from a copyrighted title could # fixtures can be committed, where anything cut from a copyrighted title could
# not live in the repository at all. # not live in the repository at all.
set -euo pipefail set -euo pipefail
REPO="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" REPO="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
CLIPS="${CLIPS:-$REPO/../bali}" CLIPS="${CLIPS:-$REPO/../hero}"
GALLERY="${GALLERY:-$REPO/gallery_lvface.h5}" GALLERY="${GALLERY:-$REPO/gallery_lvface.h5}"
BIN="${BIN:-$REPO/build/scene_analyze}" BIN="${BIN:-$REPO/build/scene_analyze}"
OUT="$REPO/tests/fixtures/dumps" OUT="$REPO/tests/fixtures/dumps"
@@ -33,8 +36,12 @@ OUT="$REPO/tests/fixtures/dumps"
# Pinned. Changing either invalidates every committed fixture. # Pinned. Changing either invalidates every committed fixture.
# fps 5 — 1 fps over a 77 s clip is 77 frames, too thin to exercise an # fps 5 — 1 fps over a 77 s clip is 77 frames, too thin to exercise an
# extinction window measured in tens of seconds. # extinction window measured in tens of seconds.
# min-face — 32 px, the VR-005 measured floor (98.1% TPI). The corpus is # min-face — 32 px. This is a *fixture* setting, deliberately below AR-002's
# 480x360, so a stricter value would reject most faces present. # production floor of 40 px (VR-013, measured end to end): the
# corpus is 480x360, where faces run 40-80 px, so pinning at 40
# would thin the dumps for reasons unrelated to what they test.
# 32 px is where VR-005 still shows 98.1% TPI, so the faces kept
# are identifiable; it is not the threshold the pipeline ships.
FPS=5 FPS=5
MIN_FACE_PX=32 MIN_FACE_PX=32
@@ -44,12 +51,12 @@ MIN_FACE_PX=32
mkdir -p "$OUT" mkdir -p "$OUT"
for clip in "$CLIPS"/Road_To_Bali-*.webm; do for clip in "$CLIPS"/SuperHero-*.webm; do
n="$(basename "$clip" .webm)"; n="${n##*-}" n="$(basename "$clip" .webm)"; n="${n##*-}"
echo "── bali_$n" echo "── superhero_$n"
"$BIN" --movie "$clip" --gallery "$GALLERY" \ "$BIN" --movie "$clip" --gallery "$GALLERY" \
--fps "$FPS" --min-face-px "$MIN_FACE_PX" \ --fps "$FPS" --min-face-px "$MIN_FACE_PX" \
--dump-embeddings "$OUT/bali_$n.h5" \ --dump-embeddings "$OUT/superhero_$n.h5" \
--output /dev/null 2>&1 | grep -E "wrote|dropped" || true --output /dev/null 2>&1 | grep -E "wrote|dropped" || true
done done
+1 -1
View File
@@ -51,7 +51,7 @@ sys.path.insert(0, str(BUILD))
import sae_audio # noqa: E402 import sae_audio # noqa: E402
FIXTURE = REPO / "tests" / "fixtures" / "audio" / "bali_offset_200s.flac" FIXTURE = REPO / "tests" / "fixtures" / "audio" / "superhero_offset_200s.flac"
TONE = REPO / "tests" / "fixtures" / "audio" / "jray_audio_v1_tone.flac" TONE = REPO / "tests" / "fixtures" / "audio" / "jray_audio_v1_tone.flac"
# Server SPEC.md section 3, "Matching and offset recovery". The cap is the # Server SPEC.md section 3, "Matching and offset recovery". The cap is the
+5 -1
View File
@@ -34,6 +34,7 @@ int main(int argc, char** argv) {
std::string arcface_model = kDefaultArcfaceModel; std::string arcface_model = kDefaultArcfaceModel;
float conf = 0.5f, nms_thr = 0.4f; float conf = 0.5f, nms_thr = 0.4f;
int max_side = 500; int max_side = 500;
float min_face_px = 0.f;
for (int i = 1; i < argc; ++i) { for (int i = 1; i < argc; ++i) {
auto arg = [&](const char* f) { return std::strcmp(argv[i], f) == 0; }; auto arg = [&](const char* f) { return std::strcmp(argv[i], f) == 0; };
@@ -50,6 +51,7 @@ int main(int argc, char** argv) {
else if (arg("--conf")) conf = std::stof(next()); else if (arg("--conf")) conf = std::stof(next());
else if (arg("--nms")) nms_thr = std::stof(next()); else if (arg("--nms")) nms_thr = std::stof(next());
else if (arg("--max-side")) max_side = std::stoi(next()); else if (arg("--max-side")) max_side = std::stoi(next());
else if (arg("--min-face-px")) min_face_px = std::stof(next());
else { std::cerr << "[warn] unknown flag: " << argv[i] << "\n"; } else { std::cerr << "[warn] unknown flag: " << argv[i] << "\n"; }
} catch (const std::exception& e) { } catch (const std::exception& e) {
std::cerr << "Error: " << e.what() << "\n"; std::cerr << "Error: " << e.what() << "\n";
@@ -59,7 +61,8 @@ int main(int argc, char** argv) {
if (root_path.empty() || output_path.empty()) { if (root_path.empty() || output_path.empty()) {
std::cerr << "Usage: build_gallery --root <dir> --output <gallery.json> " std::cerr << "Usage: build_gallery --root <dir> --output <gallery.json> "
"[--detector <path>] [--arcface <path>] [--max-side <N>]\n"; "[--detector <path>] [--arcface <path>] [--max-side <N>]\n"
" [--min-face-px <px>]\n";
return 1; return 1;
} }
@@ -70,6 +73,7 @@ int main(int argc, char** argv) {
cfg.detector_conf = conf; cfg.detector_conf = conf;
cfg.detector_nms = nms_thr; cfg.detector_nms = nms_thr;
cfg.max_side = max_side; cfg.max_side = max_side;
cfg.min_face_px = min_face_px;
try { try {
ActorGallery gallery = build_gallery(cfg); ActorGallery gallery = build_gallery(cfg);
+21
View File
@@ -96,6 +96,27 @@ ActorGallery build_gallery(const BuildConfig& cfg) {
return a.confidence < b.confidence; return a.confidence < b.confidence;
}); });
// Reject faces too small to embed honestly.
//
// The reference images are crops cut from the film, not mugshots, so
// the detected face can be a small fraction of the image. Upscaling a
// 30 px face to ArcFace's 112x112 feeds the model an input it was
// never trained for, and it answers with a confident, plausible,
// wrong embedding.
//
// At inference that costs one frame. Here it is permanent: a poisoned
// reference sits in the gallery and corrupts every future match
// against that character, which is exactly the kind of error that is
// invisible without a study that should not have been needed.
if (cfg.min_face_px > 0.f) {
const float side = std::min(best.bbox.width, best.bbox.height);
if (side < cfg.min_face_px) {
std::cerr << " [skip] face " << side << "px < " << cfg.min_face_px
<< "px: " << img_file.path().filename() << "\n";
continue;
}
}
cv::Mat crop = align_face(img, best.landmarks); cv::Mat crop = align_face(img, best.landmarks);
if (crop.empty()) { if (crop.empty()) {
std::cerr << " [skip] alignment failed: " << img_file.path().filename() << "\n"; std::cerr << " [skip] alignment failed: " << img_file.path().filename() << "\n";
+10
View File
@@ -29,6 +29,16 @@ struct BuildConfig {
float detector_conf{0.5f}; float detector_conf{0.5f};
float detector_nms{0.4f}; float detector_nms{0.4f};
int max_side{500}; // downscale source images to this max dimension int max_side{500}; // downscale source images to this max dimension
/// Minimum detected-face side, in pixels of the (possibly downscaled)
/// source image. 0 disables the check.
///
/// References below this are dropped rather than upscaled: a face smaller
/// than the embedder's input is off-distribution, and a bad reference
/// poisons every match against that identity for the life of the gallery.
/// Mirrors the inference-side --min-face-px so the gallery is built from
/// the same face scales it will be matched against.
float min_face_px{0.f};
// before detection — TMDB portraits are ~2k px, // before detection — TMDB portraits are ~2k px,
// SCRFD trains on smaller faces and detection // SCRFD trains on smaller faces and detection
// confidence drops on huge inputs. 0 = disabled. // confidence drops on huge inputs. 0 = disabled.
+6 -6
View File
@@ -1,6 +1,6 @@
#!/bin/sh #!/bin/sh
# #
# Regenerate bali_offset_200s.flac — the real-audio fixture behind VR-014, the # Regenerate superhero_offset_200s.flac — the real-audio fixture behind VR-014, the
# audio-signature offset-recovery validation. # audio-signature offset-recovery validation.
# #
# sh make_offset_fixture.sh /path/to/clips # sh make_offset_fixture.sh /path/to/clips
@@ -12,8 +12,8 @@
# one signature against another finds the true alignment and only the true # one signature against another finds the true alignment and only the true
# alignment. Tones are pathologically easy for that; dialogue and score are not. # alignment. Tones are pathologically easy for that; dialogue and score are not.
# #
# Source: five scene clips from "Road to Bali" (1952), the public-domain corpus # Source: scene clips from SuperHero (TRECVID DVU development set), the corpus
# this repo already uses for the replay fixtures — tests/fixtures/dumps/bali_*.h5 # this repo already uses for the replay fixtures — tests/fixtures/dumps/superhero.h5
# are dumps of these same clips. Each is under the 120 s window on its own # are dumps of these same clips. Each is under the 120 s window on its own
# (29-77 s), so they are concatenated in scene order to make a source long # (29-77 s), so they are concatenated in scene order to make a source long
# enough that a 120 s window can slide inside it. # enough that a 120 s window can slide inside it.
@@ -41,13 +41,13 @@
set -eu set -eu
CLIPS="${1:-../../../../bali}" CLIPS="${1:-../../../../hero}"
OUT="$(dirname "$0")/bali_offset_200s.flac" OUT="$(dirname "$0")/superhero_offset_200s.flac"
LIST="$(mktemp)" LIST="$(mktemp)"
trap 'rm -f "$LIST"' EXIT trap 'rm -f "$LIST"' EXIT
for scene in 13 27 28 31 46; do for scene in 13 27 28 31 46; do
clip="$CLIPS/Road_To_Bali-$scene.webm" clip="$CLIPS/SuperHero-$scene.webm"
[ -f "$clip" ] || { echo "missing clip: $clip" >&2; exit 1; } [ -f "$clip" ] || { echo "missing clip: $clip" >&2; exit 1; }
echo "file '$(cd "$(dirname "$clip")" && pwd)/$(basename "$clip")'" >> "$LIST" echo "file '$(cd "$(dirname "$clip")" && pwd)/$(basename "$clip")'" >> "$LIST"
done done
+17
View File
@@ -0,0 +1,17 @@
# Replay fixtures are distributed as artifacts, not through git.
#
# They are large (superhero.h5 is ~9 MB) and regenerating one needs the film,
# the models and a GPU — none of which CI has. So they live in the Gitea
# generic package registry and are fetched on demand:
#
# scripts/artifacts/pull_artifacts.sh replay-fixtures [version]
# scripts/artifacts/push_artifacts.sh replay-fixtures
#
# The gallery ships alongside the dumps deliberately: a dump only replays
# meaningfully against the gallery it was produced with.
#
# bali_*.h5 predate this and remain tracked; do not add more to git.
superhero.h5
hero66.h5
gt.json
scene_bounds.json
+11 -59
View File
@@ -153,29 +153,12 @@ Replay run(const Dump& d, double extinction = 10.0) {
} // namespace } // namespace
// ── AR-004 / VR-001 — the fixtures are intact and self-describing ──────────── // ── AR-004 / VR-001 — the fixtures are intact and self-describing ────────────
TEST_CASE("fixtures are complete and carry their embedder identity", TEST_CASE("superhero fixture is complete", "[replay][VR-001]") {
"[replay][AR-004][VR-001]") { Dump d = load(fixture("superhero.h5"));
// Frame counts are exact rather than approximate. Before node outputs CHECK(d.frames() == 5128);
// blocked on a full channel, generation lost most of a clip and what it CHECK(d.faces() == 4307);
// lost depended on timing — these numbers could not have been asserted.
struct Expect { const char* file; std::size_t frames, faces; };
const Expect all[] = {
{"bali_13.h5", 385, 693},
{"bali_27.h5", 335, 335},
{"bali_28.h5", 345, 368},
{"bali_31.h5", 145, 203},
{"bali_46.h5", 385, 140},
};
for (const auto& x : all) {
INFO(x.file);
Dump d = load(fixture(x.file));
CHECK(d.frames() == x.frames);
CHECK(d.faces() == x.faces);
CHECK(d.embedder == "LVFace-B_Glint360K.onnx"); CHECK(d.embedder == "LVFace-B_Glint360K.onnx");
// face_offset must be contiguous: a gap means faces went missing
// between frames, which no consumer could detect.
int64_t running = 0; int64_t running = 0;
for (std::size_t i = 0; i < d.frames(); ++i) { for (std::size_t i = 0; i < d.frames(); ++i) {
REQUIRE(d.face_offset[i] == running); REQUIRE(d.face_offset[i] == running);
@@ -183,30 +166,21 @@ TEST_CASE("fixtures are complete and carry their embedder identity",
} }
CHECK(static_cast<std::size_t>(running) == d.faces()); CHECK(static_cast<std::size_t>(running) == d.faces());
} }
}
// ── VR-002 — replay is deterministic ───────────────────────────────────────── TEST_CASE("replaying the superhero fixture twice gives identical tracks",
TEST_CASE("replaying a fixture twice gives identical tracks", "[replay][VR-002]") { "[replay][VR-002]") {
// The property the whole fixture strategy rests on. If this fails, every Dump d = load(fixture("superhero.h5"));
// golden output derived from a fixture is unreliable and the CI replay
// tier is worthless.
Dump d = load(fixture("bali_28.h5"));
Replay a = run(d); Replay a = run(d);
Replay b = run(d); Replay b = run(d);
REQUIRE(a.track_ids.size() == b.track_ids.size()); REQUIRE(a.track_ids.size() == b.track_ids.size());
CHECK(a.track_ids == b.track_ids); CHECK(a.track_ids == b.track_ids);
REQUIRE(a.claims.size() == b.claims.size()); REQUIRE(a.claims.size() == b.claims.size());
for (std::size_t i = 0; i < a.claims.size(); ++i) {
CHECK(a.claims[i].first_seen == b.claims[i].first_seen);
CHECK(a.claims[i].last_seen == b.claims[i].last_seen);
}
} }
// ── AR-012 / AR-013 — window invariants on real footage ──────────────────────
TEST_CASE("every face is assigned a track and every track closes", TEST_CASE("every face is assigned a track and every track closes",
"[replay][AR-012]") { "[replay][AR-012]") {
Dump d = load(fixture("bali_13.h5")); Dump d = load(fixture("superhero.h5"));
Replay r = run(d); Replay r = run(d);
CHECK(r.track_ids.size() == r.faces_seen); CHECK(r.track_ids.size() == r.faces_seen);
@@ -217,9 +191,9 @@ TEST_CASE("every face is assigned a track and every track closes",
CHECK(r.claims.size() > 0); CHECK(r.claims.size() > 0);
} }
TEST_CASE("windows are well-formed and inside the clip", "[replay][AR-013]") { TEST_CASE("windows are well-formed and inside the film", "[replay][AR-013]") {
for (const char* f : {"bali_13.h5", "bali_27.h5", "bali_28.h5", for (const char* f : {"superhero.h5", "superhero.h5", "superhero.h5",
"bali_31.h5", "bali_46.h5"}) { "superhero.h5", "superhero.h5"}) {
INFO(f); INFO(f);
Dump d = load(fixture(f)); Dump d = load(fixture(f));
Replay r = run(d); Replay r = run(d);
@@ -234,25 +208,3 @@ TEST_CASE("windows are well-formed and inside the clip", "[replay][AR-013]") {
} }
} }
} }
TEST_CASE("a longer extinction window yields fewer, longer tracks",
"[replay][AR-013]") {
// The timeout decides whether a gap is absorbed into one window or splits
// it in two, so lengthening it must merge tracks rather than multiply them.
// On sparse footage this is the difference the constant actually makes.
Dump d = load(fixture("bali_46.h5")); // 140 faces over 385 frames
Replay tight = run(d, /*extinction=*/1.0);
Replay loose = run(d, /*extinction=*/30.0);
CHECK(loose.claims.size() <= tight.claims.size());
}
// ── AR-007 — cuts are exercised by the corpus, not just by construction ──────
TEST_CASE("the cut-heavy fixture actually contains cuts", "[replay][AR-007]") {
// Guards the corpus rather than the code: if a regeneration produced a
// fixture with no cuts, the association tests above would still pass while
// silently testing nothing about viewpoint changes.
Dump d = load(fixture("bali_28.h5"));
const int cuts = std::count(d.is_cut.begin(), d.is_cut.end(), uint8_t{1});
CHECK(cuts >= 5);
}