Files
scene-actor-extraction/src/types.hpp
T
dtourolle 777c98cb33 feat(quality): score every face on sharpness and alignment before it is evidence
Every embedding now carries the quality of the input it came from. Both
axes fall out of the AR-005 warp for free: crop_sharpness() is the
normalised Laplacian variance over the aligned 112x112, so contrast and
size cannot leak into it, and the alignment residual is the part of the
landmark deformation a similarity transform cannot explain, so in-plane
roll reads as zero and foreshortening does not.

Carried, not consumed. Nothing discounts or thresholds on either number
yet -- that is AR-030 and VR-012, and the knee has to be located against
recorded data before a gate is chosen. What this change buys is that the
data exists to locate it with.

No face is admitted unscored: the -1 sentinel is preserved rather than
clamped, and a degenerate landmark fit is counted rather than silently
dropped.

Takes the VR-001 dump to schema_version 2. The bump is not for readers,
which check for the datasets by name and replay a v1 dump unchanged; it
is so a consumer can tell "never scored" from "scored zero", which is
not recoverable from the arrays afterwards.

TRACES: AR-028, AR-029, AR-030 | VR-001 | SR-002
2026-08-05 14:37:30 +02:00

187 lines
9.2 KiB
C++
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
#pragma once
#include <array>
#include <cstdint>
#include <string>
#include <vector>
#include <opencv2/core.hpp>
#include "gallery/embedder_stamp.hpp"
// ── Embedding ─────────────────────────────────────────────────────────────────
// 512-dim L2-normalised ArcFace embedding
using Embedding = std::array<float, 512>;
inline float cosine_similarity(const Embedding& a, const Embedding& b) {
float dot = 0.f;
for (int i = 0; i < 512; ++i) dot += a[i] * b[i];
return dot;
}
// ── Frame ─────────────────────────────────────────────────────────────────────
// Raw sampled frame from the movie. eof=true is the pipeline shutdown sentinel:
// every node must forward it immediately without processing.
struct Frame {
cv::Mat image;
double timestamp_sec{0.0};
int64_t frame_idx{-1};
bool eof{false};
bool is_cut{false}; // histogram: intra-scene camera-angle change (tracker reset)
bool is_scene_boundary{false}; // TransNetV2: true shot/scene boundary (opt-in)
float cut_score{0.f}; // histogram cut score = 1 - hist_corr (0=identical, ~1=cut); HUD/debug
float bbox_upscale{1.f}; // multiply detector bboxes/landmarks by this to map back to
// original video resolution (>1 when dense_scale downscaled the frame)
};
// ── CutEvent ──────────────────────────────────────────────────────────────────
// Emitted by SceneDetectorFunc when TransNetV2 localises a shot boundary, keyed
// by the boundary frame's timestamp. eof=true is the shutdown sentinel.
struct CutEvent {
double timestamp_sec{0.0};
float probability{0.f}; // sigmoid boundary score at the peak
bool eof{false};
};
// ── ArcFace alignment ─────────────────────────────────────────────────────────
// Canonical 5-point target positions for a 112×112 ArcFace crop.
// Landmark order: right-eye, left-eye, nose, right-mouth, left-mouth
// (matches SCRFD output order — no reordering needed).
inline constexpr float kArcFaceRef[5][2] = {
{38.2946f, 51.6963f},
{73.5318f, 51.5014f},
{56.0252f, 71.7366f},
{41.5493f, 92.3655f},
{70.7299f, 92.2041f},
};
// ── DetectedFace ──────────────────────────────────────────────────────────────
// One face found by SCRFD in a Frame.
// Landmark order matches ArcFace convention (same as SCRFD output order):
// [0] right-eye-centre [1] left-eye-centre [2] nose
// [3] right-mouth [4] left-mouth
/// TRACES: AR-028 | SR-002
struct DetectedFace {
cv::Rect2f bbox;
std::array<cv::Point2f, 5> landmarks;
float confidence{0.f};
// ── AR-028 quality vector ────────────────────────────────────────────────
// Three axes, kept separate and never collapsed into one scalar: they fail
// for different reasons, have different remedies, and do not earn the same
// response. Carried, not consumed — the vector travels with the face into
// the VR-001 dump so a threshold can be re-litigated against recorded data
// rather than by re-running video.
//
// **Size is the third axis and is deliberately not a field here.** It is
// `bbox`, which every consumer already has, scaled by the frame's
// `bbox_upscale` to reach the original resolution AR-002 thresholds in.
// Copying it into a second field would put the same quantity in two
// coordinate spaces inside one struct — the trap SCHEMA.md records for
// `bbox_upscale` — and the copy would be the one that drifts.
//
// Both fields below are -1 until the aligner runs, so *unscored* is
// distinguishable from *scored badly*. Nothing downstream may read a
// negative value as a quality.
// AR-029 sharpness: normalised Laplacian variance over the aligned crop,
// dimensionless. Falls with motion blur and soft focus; invariant to
// contrast, and taken on the fixed 112×112 canvas so it cannot re-measure
// face size. See crop_sharpness() for the construction and its one hazard.
float sharpness{-1.f};
// AR-030 visibility: RMS landmark misfit, in canonical 112×112 pixels, left
// over after the best similarity fit to the ArcFace template. Rises with
// out-of-plane pose and with occlusion; blind to in-plane roll and to face
// size, both of which the fit absorbs. Set by the aligner, which is where
// the transform is computed; -1 until then.
float alignment_residual{-1.f};
};
// ── Pipeline messages ─────────────────────────────────────────────────────────
struct SceneFrame {
Frame source;
std::vector<DetectedFace> faces; // empty when no faces detected (or eof)
};
struct AlignedSceneFrame {
Frame source;
std::vector<DetectedFace> faces;
std::vector<cv::Mat> crops; // 112×112 BGR, ArcFace-ready; parallel to faces
};
struct EmbeddedSceneFrame {
Frame source;
std::vector<DetectedFace> faces;
std::vector<cv::Mat> crops; // forwarded for debug rendering downstream
std::vector<Embedding> embeddings;
};
// ── Face tracking ─────────────────────────────────────────────────────────────
// Output of FaceTrackerFunc — EmbeddedSceneFrame augmented with per-detection
// track context.
struct TrackedSceneFrame {
Frame source;
std::vector<DetectedFace> faces;
std::vector<cv::Mat> crops;
std::vector<int> track_ids; // -1 = brand-new track this frame
std::vector<Embedding> embeddings; // per-frame raw (from embedder)
};
// ── Identity matching ─────────────────────────────────────────────────────────
struct IdentifiedActor {
int actor_idx{-1}; // index into ActorGallery::actors; -1 = unknown
int track_id{-1}; // face track ID from FaceTrackerFunc
std::string name;
std::string imdb_id;
std::string tmdb_id;
std::string jellyfin_id; // Jellyfin Person item GUID, if gallery was built from Jellyfin
float similarity{0.f}; // calibrated P(match) or cosine similarity; 0 for unknowns
cv::Rect2f bbox;
cv::Mat crop; // 112×112 aligned crop (stored as shared_ptr by KPN)
};
struct MatchedSceneFrame {
Frame source;
std::vector<IdentifiedActor> actors; // includes unknowns (actor_idx == -1)
};
// ── Scene annotation ──────────────────────────────────────────────────────────
// Output of the scene tracker: one per sampled frame.
// visible_actors contains all actors still within their extinction window.
struct SceneAnnotation {
double timestamp_sec{0.0};
std::vector<IdentifiedActor> visible_actors;
bool eof{false};
};
// ── Actor gallery ─────────────────────────────────────────────────────────────
// Loaded once at startup; baked into the identity matcher.
struct ActorGallery {
struct Actor {
std::string imdb_id;
std::string tmdb_id;
std::string jellyfin_id; // Jellyfin Person item GUID, if known
std::string name;
std::vector<Embedding> embeddings; // one per reference image
std::vector<std::string> source_images;
};
std::vector<Actor> actors;
/// TRACES: GR-004 | SR-001
// Which embedder produced every embedding above. Empty == the file predates
// model binding; see gallery/embedder_stamp.hpp for what is checked and why.
EmbedderStamp embedder;
// Cached Platt-sigmoid calibration (see gallery/gallery_calibration.hpp),
// stored alongside the gallery in HDF5 so it never needs recomputing
// unless the reference embeddings actually change. calib_valid=false and
// calib_hash=0 means "not present in this file, compute it."
float calib_a{10.f};
float calib_b{-5.f};
bool calib_valid{false};
uint64_t calib_hash{0};
};