refactor(bench): SuperHero replaces Road to Bali as the reference film
Bali was chosen because the TRECVID DVU set ships character mugshots, but its reference crops are unusable at scale: median detected face 27 px against a 69 px maximum, so every reference was upscaled 4x or more past what the embedder was trained for (AR-011). A 66 px floor left 2 of 69 references; no threshold exists that both keeps the faces in distribution and leaves enough of them to calibrate. SuperHero is 69 px median and 241 px max. Its gallery builds at a 66 px floor with 14 references over 5 characters, and calibrates on its own (a=15.2867 b=-4.98633, 100% train accuracy) instead of borrowing constants. Measured on the fused 17-minute film, one stream rather than per-scene clips so presence windows cross real scene boundaries as SR-002 intends: precision 1.00, recall 0.65, F1 0.79 — 13 true positives, 0 false positives, 7 misses. Every out-of-gallery character was declined rather than forced onto a nearest match. The misses are the short scenes (14 s, 38 s, 27 s), consistent with per-track accumulation needing sightings. - build_gallery gains --min-face-px, filtering the *detected face* rather than the crop. The DVU images are scene crops, not mugshots, so crop dimensions say nothing about face scale. A poisoned reference is permanent in a way a bad frame is not: it corrupts every future match against that identity. - scripts/fetch_dvu.sh fetches mugshots, scene graphs and segmentation for any DVU film. NIST names the same film three different ways, so KG_DIR and KG_FILE are overridable rather than derived. This exists as a script because the first copy of this data was assembled ad hoc in /tmp and was lost with it, taking the working gallery along. - Replay fixtures move to the artifact registry: push/pull_artifacts.sh gain a replay-fixtures target, and tests/fixtures/dumps/.gitignore keeps them out of git. superhero.h5 is ~9 MB and regenerating it needs the film, the models and a GPU — none of which CI has. The gallery ships with the dumps, since a dump only replays against the gallery it was produced with. - AR-012 and AR-013 coverage is ported onto the new fixture rather than dropped with the Bali cases: 12369 assertions, up from 7991, since the film is an order of magnitude larger than the clips. Suite: 15679 assertions, 101 test cases. TRACES: AR-011, AR-012, AR-013 | VR-001, VR-005 | SR-002
This commit is contained in:
@@ -96,6 +96,27 @@ ActorGallery build_gallery(const BuildConfig& cfg) {
|
||||
return a.confidence < b.confidence;
|
||||
});
|
||||
|
||||
// Reject faces too small to embed honestly.
|
||||
//
|
||||
// The reference images are crops cut from the film, not mugshots, so
|
||||
// the detected face can be a small fraction of the image. Upscaling a
|
||||
// 30 px face to ArcFace's 112x112 feeds the model an input it was
|
||||
// never trained for, and it answers with a confident, plausible,
|
||||
// wrong embedding.
|
||||
//
|
||||
// At inference that costs one frame. Here it is permanent: a poisoned
|
||||
// reference sits in the gallery and corrupts every future match
|
||||
// against that character, which is exactly the kind of error that is
|
||||
// invisible without a study that should not have been needed.
|
||||
if (cfg.min_face_px > 0.f) {
|
||||
const float side = std::min(best.bbox.width, best.bbox.height);
|
||||
if (side < cfg.min_face_px) {
|
||||
std::cerr << " [skip] face " << side << "px < " << cfg.min_face_px
|
||||
<< "px: " << img_file.path().filename() << "\n";
|
||||
continue;
|
||||
}
|
||||
}
|
||||
|
||||
cv::Mat crop = align_face(img, best.landmarks);
|
||||
if (crop.empty()) {
|
||||
std::cerr << " [skip] alignment failed: " << img_file.path().filename() << "\n";
|
||||
|
||||
@@ -29,6 +29,16 @@ struct BuildConfig {
|
||||
float detector_conf{0.5f};
|
||||
float detector_nms{0.4f};
|
||||
int max_side{500}; // downscale source images to this max dimension
|
||||
|
||||
/// Minimum detected-face side, in pixels of the (possibly downscaled)
|
||||
/// source image. 0 disables the check.
|
||||
///
|
||||
/// References below this are dropped rather than upscaled: a face smaller
|
||||
/// than the embedder's input is off-distribution, and a bad reference
|
||||
/// poisons every match against that identity for the life of the gallery.
|
||||
/// Mirrors the inference-side --min-face-px so the gallery is built from
|
||||
/// the same face scales it will be matched against.
|
||||
float min_face_px{0.f};
|
||||
// before detection — TMDB portraits are ~2k px,
|
||||
// SCRFD trains on smaller faces and detection
|
||||
// confidence drops on huge inputs. 0 = disabled.
|
||||
|
||||
Reference in New Issue
Block a user