Files
scene-actor-extraction/src/types.hpp
T
dtourolleandClaude Opus 5 6da8ac2bdb perf: back the CPU similarity GEMM with OpenBLAS
The CPU path was a scalar triple loop. It is the correctness oracle for the GPU
backends, but it is also what CI runs — there is no GPU on the N100 host — and
since AR-003 removed the per-frame face cap, a crowded frame now scores many
faces against a library-scale gallery. Scoring one face against 5000 embeddings
is 2.6 MFLOP; in scalar that does not hold up (AR-027).

S(g,f) viewed as row-major [n_faces x n_gallery] is exactly query * gallery^T,
so the loop nest collapses into a single cblas_sgemm.

OpenBLAS is optional in the build: found via pkg-config, and the scalar path
remains when it is absent so no hard dependency is added and the two can be
diffed when a similarity looks wrong. The configure step warns rather than
failing, since a developer without it should still get a working tree.

The test target links it too. Without that the suite compiles the scalar
fallback while the builder image ships CBLAS, so CI would be verifying a kernel
that is not the one running in production — the same class of mistake as testing
a path the gate never executes.

Recorded as required (not optional) in the DP-007 image, for the same reason.

Suite: 92 cases, 6136 assertions, with CBLAS compiled in.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

TRACES: AR-026, AR-027, DP-007 | SR-001
2026-07-31 15:04:29 +02:00

162 lines
7.6 KiB
C++
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
#pragma once
#include <array>
#include <cstdint>
#include <string>
#include <vector>
#include <opencv2/core.hpp>
#include "gallery/embedder_stamp.hpp"
// ── Embedding ─────────────────────────────────────────────────────────────────
// 512-dim L2-normalised ArcFace embedding
using Embedding = std::array<float, 512>;
inline float cosine_similarity(const Embedding& a, const Embedding& b) {
float dot = 0.f;
for (int i = 0; i < 512; ++i) dot += a[i] * b[i];
return dot;
}
// ── Frame ─────────────────────────────────────────────────────────────────────
// Raw sampled frame from the movie. eof=true is the pipeline shutdown sentinel:
// every node must forward it immediately without processing.
struct Frame {
cv::Mat image;
double timestamp_sec{0.0};
int64_t frame_idx{-1};
bool eof{false};
bool is_cut{false}; // histogram: intra-scene camera-angle change (tracker reset)
bool is_scene_boundary{false}; // TransNetV2: true shot/scene boundary (opt-in)
float cut_score{0.f}; // histogram cut score = 1 - hist_corr (0=identical, ~1=cut); HUD/debug
float bbox_upscale{1.f}; // multiply detector bboxes/landmarks by this to map back to
// original video resolution (>1 when dense_scale downscaled the frame)
};
// ── CutEvent ──────────────────────────────────────────────────────────────────
// Emitted by SceneDetectorFunc when TransNetV2 localises a shot boundary, keyed
// by the boundary frame's timestamp. eof=true is the shutdown sentinel.
struct CutEvent {
double timestamp_sec{0.0};
float probability{0.f}; // sigmoid boundary score at the peak
bool eof{false};
};
// ── ArcFace alignment ─────────────────────────────────────────────────────────
// Canonical 5-point target positions for a 112×112 ArcFace crop.
// Landmark order: right-eye, left-eye, nose, right-mouth, left-mouth
// (matches SCRFD output order — no reordering needed).
inline constexpr float kArcFaceRef[5][2] = {
{38.2946f, 51.6963f},
{73.5318f, 51.5014f},
{56.0252f, 71.7366f},
{41.5493f, 92.3655f},
{70.7299f, 92.2041f},
};
// ── DetectedFace ──────────────────────────────────────────────────────────────
// One face found by SCRFD in a Frame.
// Landmark order matches ArcFace convention (same as SCRFD output order):
// [0] right-eye-centre [1] left-eye-centre [2] nose
// [3] right-mouth [4] left-mouth
struct DetectedFace {
cv::Rect2f bbox;
std::array<cv::Point2f, 5> landmarks;
float confidence{0.f};
// AR-030 visibility: RMS landmark misfit, in canonical 112×112 pixels, left
// over after the best similarity fit to the ArcFace template. Rises with
// out-of-plane pose and with occlusion; blind to in-plane roll and to face
// size, both of which the fit absorbs. Set by the aligner, which is where
// the transform is computed; -1 until then.
float alignment_residual{-1.f};
};
// ── Pipeline messages ─────────────────────────────────────────────────────────
struct SceneFrame {
Frame source;
std::vector<DetectedFace> faces; // empty when no faces detected (or eof)
};
struct AlignedSceneFrame {
Frame source;
std::vector<DetectedFace> faces;
std::vector<cv::Mat> crops; // 112×112 BGR, ArcFace-ready; parallel to faces
};
struct EmbeddedSceneFrame {
Frame source;
std::vector<DetectedFace> faces;
std::vector<cv::Mat> crops; // forwarded for debug rendering downstream
std::vector<Embedding> embeddings;
};
// ── Face tracking ─────────────────────────────────────────────────────────────
// Output of FaceTrackerFunc — EmbeddedSceneFrame augmented with per-detection
// track context.
struct TrackedSceneFrame {
Frame source;
std::vector<DetectedFace> faces;
std::vector<cv::Mat> crops;
std::vector<int> track_ids; // -1 = brand-new track this frame
std::vector<Embedding> embeddings; // per-frame raw (from embedder)
};
// ── Identity matching ─────────────────────────────────────────────────────────
struct IdentifiedActor {
int actor_idx{-1}; // index into ActorGallery::actors; -1 = unknown
int track_id{-1}; // face track ID from FaceTrackerFunc
std::string name;
std::string imdb_id;
std::string tmdb_id;
std::string jellyfin_id; // Jellyfin Person item GUID, if gallery was built from Jellyfin
float similarity{0.f}; // calibrated P(match) or cosine similarity; 0 for unknowns
cv::Rect2f bbox;
cv::Mat crop; // 112×112 aligned crop (stored as shared_ptr by KPN)
};
struct MatchedSceneFrame {
Frame source;
std::vector<IdentifiedActor> actors; // includes unknowns (actor_idx == -1)
};
// ── Scene annotation ──────────────────────────────────────────────────────────
// Output of the scene tracker: one per sampled frame.
// visible_actors contains all actors still within their extinction window.
struct SceneAnnotation {
double timestamp_sec{0.0};
std::vector<IdentifiedActor> visible_actors;
bool eof{false};
};
// ── Actor gallery ─────────────────────────────────────────────────────────────
// Loaded once at startup; baked into the identity matcher.
struct ActorGallery {
struct Actor {
std::string imdb_id;
std::string tmdb_id;
std::string jellyfin_id; // Jellyfin Person item GUID, if known
std::string name;
std::vector<Embedding> embeddings; // one per reference image
std::vector<std::string> source_images;
};
std::vector<Actor> actors;
/// TRACES: GR-004 | SR-001
// Which embedder produced every embedding above. Empty == the file predates
// model binding; see gallery/embedder_stamp.hpp for what is checked and why.
EmbedderStamp embedder;
// Cached Platt-sigmoid calibration (see gallery/gallery_calibration.hpp),
// stored alongside the gallery in HDF5 so it never needs recomputing
// unless the reference embeddings actually change. calib_valid=false and
// calib_hash=0 means "not present in this file, compute it."
float calib_a{10.f};
float calib_b{-5.f};
bool calib_valid{false};
uint64_t calib_hash{0};
};