A gallery is only valid for the embedder that produced its vectors. Cosine
similarities across models are meaningless but *look* plausible, so the mistake
is silent and every measurement taken afterwards is suspect. Stamp the embedder
identity into the gallery at build; verify it at every load.
The stamp is the model file's basename plus the SHA-256 of its bytes (plus
embed_dim). The hash decides, the name explains. A name alone is a promise
rather than a fact — models get re-exported and overwritten in place under an
unchanged filename, which is exactly the case where the weights differ and
nothing else does. A hash alone is correct but unactionable in an error message.
SHA-256 is derived from the artefact, needs no registry kept current, and costs
~0.1s for a 250MB ONNX, memoised per process.
Mismatch is a hard error in every mode, with no bypass, naming both sides.
Unstamped legacy galleries warn loudly and proceed: unknown is not known-bad,
and hard-failing every pre-existing gallery would turn the check into something
people disable rather than trust. --require-gallery-stamp (or
SAE_REQUIRE_GALLERY_STAMP=1, which propagates to subprocesses) promotes that to
a hard error — the mode measurement work should run in. scripts/stamp_gallery.py
re-binds an existing gallery with no re-embedding, so "warn" is a cheap state to
leave rather than a permanent one.
Embedding dumps carry the same stamp: a replay has no live embedder, so the dump
is the embedder as far as the gallery is concerned. Derived galleries inherit
their source's stamp; --merge and the JSON gallery merge check before writing,
since one file holding two embedding spaces cannot be untangled afterwards.
Verified in: scene_analyze, scene_preview, the sae_kpn matcher binding,
replay.py, optimize.py (once per film at startup, before the first evaluation),
movienet_eval.py and both merge paths.
Stamp logic lives in src/gallery/embedder_stamp.{hpp,cpp} and its Python twin
scripts/sae_stamp.py, kept dependency-light so replay subprocesses do not pay
sae_gallery's requests/Pillow import to ask whether two models match.
Tests: 12 new cases in test_gallery_store.cpp covering the comparison logic,
both round trips, and the SHA-256 vectors that guarantee the C++ and hashlib
stamps agree. No ONNX or GPU required.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The phase structure encoded ordering assumptions that stopped being true as the
design changed, and its Phase 2 still described retuning constants that are now
withdrawn. Ordering is now derived from per-requirement dependencies instead:
anything with no unmet dependency is startable.
Carries over the TrackRegistry design (now keyed to AR-012/AR-013) and records
what was withdrawn from the old plan, including the --presence-mode flag —
comparison against old behaviour uses recorded reference output rather than a
second live code path.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adds the requirements baseline for the pipeline redesign:
- SPEC.md — software requirements with Current/Gap deltas per item, so the
document doubles as a work list.
- requirements.md — stable flat IDs (AR/DP/IR/GR/VR) with parent traces,
priorities, statuses, and a per-requirement verification plan. Replaces the
thematic A1..E8 scheme, which had already produced an A1a and an out-of-order
E6; IDs are now permanent and never reused.
- IMPLEMENTATION-PLAN.md — phased work.
The central change is AR-012: presence follows track extent rather than
per-frame recognition, so a window starts when an actor appears rather than
when the recogniser first succeeded. anneal_sec and extinction_sec are
withdrawn rather than retuned — a track that survives its own gaps leaves them
nothing to do.
Verification is shaped by CI running on an N100 with no dGPU: the existing
HDF5 dump makes everything downstream of embedding replayable on CPU, which
covers the bulk of the redesign.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
trtexec rejects --minShapes/--optShapes/--maxShapes for a fully static model
("Static model does not take explicit shapes"). TransNetV2's input is fixed at
1x100x27x48x3, so the shape comes from the model itself.
Gallery build now over-fetches TMDB/Wikidata candidates by a configurable
factor: near-duplicate stills (the same photo at different crops or
resolutions) are discarded after embedding, so downloading exactly
images_per_actor left actors short of that many *distinct* embeddings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
build_trt_engines.sh hardcoded 'input.1' for the ArcFace and SCRFD shape
profiles, which only matches arcface_w600k_{r50,mbf}. Building engines for
any other embedder failed with:
Cannot find input tensor with name "input.1" in the network inputs!
Input names differ per model: LVFace-B_Glint360K uses 'data', arcface_r18
uses 'input', arcface_w600k_{r50,mbf} use 'input.1'. This matters now that
LVFace-B is the default embedder (src/config.hpp), so ARCFACE_MODEL=<LVFace>
is the expected path.
Read the name from each model via onnxruntime at build time.
OpenCV: distros (Arch/CachyOS) now ship OpenCV 5 as default. The config
package rejects a 5.x install when find_package requests 4, so probe for 5
first and fall back to 4. All components used here (core, imgproc, imgcodecs,
videoio, dnn, objdetect, highgui) exist in both.
TensorRT: nvinfer1::Dims5 was removed in TRT 10 (Dims2..Dims4 remain in
NvInferLegacyDims.h). Build the TransNetV2 rank-5 input shape via the generic
nvinfer1::Dims, which is valid on both 8.x and 10.x.
Her bbox is frozen at identical coordinates for t=2450 and t=2451; the
dump's own per-frame detections show only one real face at t=2451, and
it matches the Chloë Sevigny box (IoU 1.0), not hers. The frame is one
ghost overlapping one fresh misidentification, not two competing fresh
identities as previously written.
Replaces narrative claims with verified numbers across all report pages:
- Cross-model held-out validation (LVFace/mbf/r18, all 5 held-out
films): LVFace wins every film outright, not just "consistent with"
the training-set pick. r50 dropped from the detailed comparison
(gallery has ~30% fewer reference images per actor than the other
three models on identical source photos).
- Per-film training breakdown: LVFace does not win every training
film (mbf beats it on Lord of War); the 75.3% macro figure hides a
10.7pp spread.
- Gallery coverage computed per film (20.3%-78.6%) instead of one
flat 67%-missing average.
- Found and fixed a real scoring bug in optimize.py: a candidate
whose hardest film's replay timed out was averaged over survivors
instead of penalized, silently rewarding partial coverage. Affected
3 of 16 training combos; corrected throughout, and optimize.py now
scores an incomplete evaluation f1=0.0 instead of averaging over
whichever films happened to finish.
- Every FPI frame in the deep dive now comes from the proper montage
renderer (Onscreen/Offscreen panel, ghosts never drawn as boxes),
never the bare-box debug overlay used earlier.
- Every distinct out-of-cast name across all 9 films gets its own
frame at its first appearance (9 names, 4 films), not a
single-example spot check: 2 ground-truth gaps, 1 photograph
misread as a person, 6 genuine lookalike confusions.
- New methodology.md: the scene-level-vs-per-second scoring mismatch
that the rest of the report assumes, written out once.
- Cut the deadlock/gdb debugging narrative from the experiment log;
kept the one fact that matters (KPN's node/network split lets the
expensive GPU stage run once and the cheap stage replay against
cached embeddings).
- Plain declarative style throughout, no em dashes, no blog voice.
- switch report frames to the scene best/worst montage renderer
(Onscreen/Offscreen panels + TPI/FPI/FN legend): perfect-second hero,
wedding couple, funeral 19-of-20, polygraph bridging, crew-scene FN
ceiling, Robert Patrick ground-truth gap, rapid-cut double label,
Herbie Hancock on an in-fiction screen
- deep dive restructured: extinction bridging framed as designed
behavior with a measurable cost (debug overlay draws the boxes; the
shipped output is presence windows), plus the face-vs-presence
ceiling and two X-Ray-is-wrong exhibits
- Material polish: light/dark palette toggle, landing-page grid cards,
figure/caption CSS, how-to-read admonition; site_url set so 404 links
resolve under the Pages subpath
- README: perfect-second and screen-call frames committed (gitignore
exceptions), readme_example.jpg retired
- build_site.sh: stage_frame helper downscales montage frames to 1920px
and pulls any missing montage-frames packages
- experiment_charts.py generates 4 figures from experiments/ artifacts:
held-out per-film F1, 16-combo ranking, DE search landscape, and the
Downton detector-vs-tracker ghost timeline (replaces the blank
title-card screenshot)
- new frames: 19-correct wedding shot (success case), Many Saints
ghost-vs-unknown frame (three error classes in one image)
- rename rep4-optimizer-results.md -> model-bakeoff.md; rep4 kept only
as the on-disk artifact prefix, explained once
- repo file references are now links via https://REPOLINK/<path>
placeholders; build_site.sh pins them to the HEAD commit's raw URLs
and fails the build if a linked path doesn't exist at HEAD
- drop references to removed scripts (scene_score.py, score_config.py)
and to session-memory names; mark artifact-registry paths with their
pull commands
- commit readme_example.jpg + pipeline_topology.svg so README renders
on the plain Gitea repo view
- deploy_pages.sh: push built site/ to the gitea-pages branch
- experiment_charts.py generates 4 figures from experiments/ artifacts:
held-out per-film F1, 16-combo ranking, DE search landscape, and the
Downton detector-vs-tracker ghost timeline (replaces the blank
title-card screenshot)
- new frames: 19-correct wedding shot (success case), Many Saints
ghost-vs-unknown frame (three error classes in one image)
- rename rep4-optimizer-results.md -> model-bakeoff.md; rep4 kept only
as the on-disk artifact prefix, explained once
- repo file references are now links via https://REPOLINK/<path>
placeholders; build_site.sh pins them to the HEAD commit's raw URLs
and fails the build if a linked path doesn't exist at HEAD
- drop references to removed scripts (scene_score.py, score_config.py)
and to session-memory names; mark artifact-registry paths with their
pull commands
- commit readme_example.jpg + pipeline_topology.svg so README renders
on the plain Gitea repo view
- deploy_pages.sh: push built site/ to the gitea-pages branch
Splits the rep4 write-up's key findings into their own linkable pages:
- best-model.md: calibration curves first (discriminative power, independent
of any threshold), then F1 on the benchmark — LVFace-B Glint360K wins both.
- gallery-scope.md: whole vs. cast-restricted gallery, isolated from model and
expansion choice — restriction wins on every axis, but isn't a shipped
runtime feature yet.
- pose-expansion.md: the training-set expand_gallery effect, and the held-out
replication attempt that found it doesn't reproduce (5 films, 2 models,
after catching and fixing a replay-timeout truncation bug and a bbox
first-match-instead-of-best-match bug in the comparison harness itself). An
honest null result, with the methodology errors documented since they're
exactly the kind that manufacture a false "it works!" finding.
- lvface-deep-dive.md: the winning model's held-out generalization gap, its
two failure modes (frozen-bbox ghost tracks), and a verified case (cross-
checked against Jellyfin's independent cast metadata) where LVFace
correctly identified an actor that X-Ray's ground truth failed to credit.
Adds a "report-highlights" artifact-registry package (scripts/artifacts/
push_artifacts.sh, pull_artifacts.sh) for hand-picked illustrative frames that
aren't reproducible via the automated best/worst montage selection, and wires
pulling it into scripts/docs/build_site.sh.
docs/rep4-optimizer-results.md is the main deliverable: the model bake-off +
threshold re-tune experiment log, including the ROCm teardown deadlock root
cause and fix, DE concurrency tuning, the 16-combo results table, held-out
validation against 5 films never seen by the optimizer (macro F1 67.4% vs.
75.3% training — a real generalization gap), the frozen-bbox "ghost track"
failure mode found via annotated frame evidence, calibration curves per model,
and an isolated-effects breakdown of gallery scope vs. pose expansion.
MkDocs site (mkdocs.yml, docs/index.md) renders docs/*.md; scripts/docs/
pulls referenced images from the artifact registry and generates the
calibration chart at build time (see the tooling commit) rather than
committing images to the repo.
experiments/ now keeps only scripts + README + SESSION_STATE.md in git — every
data artifact (galleries, dumps, X-Ray corpus, montage frames, trajectories,
manifests, results) moved to the Gitea package registry. film-lut.template.json
is the committed placeholder for the gitignored file-lut.json (real local
movie paths, never shared — some source filenames carry scene-release tags).
Adds models/transnetv2.onnx (via Git LFS, matching the other ONNX models) for
the new scene-detection path.
GPU-free, model-free tests for the pure logic: gallery HDF5 save/load
round-trips (actors, embeddings, embedded calibration) and legacy JSON
read back-compat; the calibration sigmoid fit, boundary inversion, and the
in-memory hash-keyed cache reuse/staleness; TrackGallery's diversity-buffer
eviction, novelty/spread safety gates, and promotion; FaceTracker's IoU/
embedding association and cross-cut track revival; and the GEMM similarity
backend (forced to CPU so the suite runs without a GPU).
Verified: all 39 test cases / 1640 assertions pass (cmake -DSAE_BUILD_TESTS=ON).
Optimizer (scripts/optimizer/): replay.py runs the real C++ tracker/matcher/
scene_tracker chain over a dumped-embeddings HDF5 via sae_kpn, so a threshold
sweep never re-decodes video or re-embeds faces. optimize.py drives scipy's
differential_evolution over the knob space, with DE-level parallelism
(multiple population candidates evaluated concurrently via a ThreadPoolExecutor)
on top of per-film replay parallelism. second_score.py is the per-second X-Ray
scoring metric (TPI/FPI/FN, out-of-cast misID weighted 10x, fair recall masked
to gallery-known cast) that superseded an earlier scene-union metric.
dump_error_frames.py / dump_scene_montage.py extract annotated video frames
(bounding boxes, TPI/FPI/FN captions, onscreen-vs-offscreen split) for visual
review of a replay against ground truth. Gallery utilities: cast_restrict.py,
gallery_membership.py, fetch_missing_actors.py, reembed_gallery.py.
scripts/validation/: X-Ray ground-truth loading and provider-agnostic identity
matching (identity.py's keys_for — an actor is the union of every id we can
derive, since pipeline output and ground truth don't share one id space).
scripts/artifacts/: push/pull scripts for the Gitea generic package registry —
galleries, montage frames, and experiment data (manifests/trajectories/results)
are pushed there instead of committed, since none are needed to run the app,
only benchmarks. Versioned by git short-SHA.
scripts/docs/: MkDocs site build (build_site.sh) and the calibration-curve
comparison chart (calibration_chart.py, matplotlib, reads each gallery's
embedded calibration).
Gallery-building scripts (make_jellyfin_gallery.py, make_gallery.py,
filter_gallery.py, run_from_jellyfin.py, movienet_eval.py, movienet_prep.py,
sae_gallery.py) updated to read/write HDF5 galleries exclusively, matching the
engine-side format switch. run_from_jellyfin.py and the optimizer no longer
carry movie source paths in shared manifests (some source filenames include
scene-release tags) — resolved locally via a gitignored file-lut.json instead.
New C++ sources:
- kpn_bindings.cpp (sae_kpn): assembles the real face_tracker/identity_matcher/
scene_tracker nodes inside a Python-driven KPN network via nanobind, for
offline threshold-sweep replay against dumped embeddings (scripts/optimizer/).
- track_gallery.hpp: per-film gallery expansion — promotes a confidently-
identified track's novel-pose reference views into an in-memory annex so
later frames/tracks of that actor at similar poses are recognised, without
touching the baked gallery.
- dump_embeddings.cpp: standalone exe that runs detect→embed only (no gallery,
no matching) and dumps per-frame face embeddings + metadata to HDF5, so a
parameter sweep can replay the expensive half once and vary tracking/matching
config freely downstream.
- scene_detector.hpp / scene_detector_node.hpp: TransNetV2-based shot-boundary
detection, opt-in alongside the always-on histogram cut detector.
- camera_position_change_detector_node.hpp, embedding_dump_node.hpp: supporting
nodes for the above.
Gallery format switches from JSON to HDF5 exclusively (JSON read-only kept for
back-compat): save_gallery always writes HDF5, and the fitted Platt-sigmoid
calibration (a, b, valid, hash) is now embedded directly in the gallery file
instead of a sidecar .calib_cache.json — identity_matcher reads it from the
loaded gallery and writes back only when the embeddings actually changed
(hash mismatch), skipping the O(n^2) refit otherwise.
Also includes: TensorRT inference backend support (ort_backend.cpp,
trt_backend.cpp), gemm_backend improvements, TransNetV2-based scene-boundary
detection wired through frame_source/face_tracker/main, and CMake build
target updates for the new sources.
Bumps the KPN submodule to feature/persistent-pipeline-reuse (push_blocking
backpressure, node_ptr/node_stats introspection, ObjectVariantNodeWrapper for
stateful functors) — needed by the optimizer's sae_kpn Python bindings.
scene_gap_hist.py scans scene_analyze output JSONs and, for every actor,
computes the gap (next_scene_start - prev_scene_end) between consecutive
scenes, emitting a text histogram of the distribution. Used to inform the
anneal_sec default.
movienet_eval: replace the per-element dot() with numpy — actor references are
loaded once as an ndarray and scored with a single matmul, keeping a
whole-library gallery fast.
movienet_prep: count and report frames referenced by annotations but absent
from Image.zip instead of skipping them silently.
LVFace-B_Glint360K.onnx shares ArcFace's I/O contract (112x112 aligned crop ->
L2-normalised 512-d) and input scaling, so it drops in via --arcface-model.
Note the caveat that galleries and calib caches must be rebuilt with the same
embedder used for analysis.
Add two cameo hunters that flag actors recognised in a title but absent from
its cast:
- cameo_jellyfin.py — pure-Jellyfin cast-membership check (no id cross-walk)
- cameo_hunt.py — TMDB filmography check (actor's combined_credits)
run_from_jellyfin.py now stamps the analysed title's Jellyfin item GUID into
the output JSON as top-level 'jellyfin_item_id' (scene_analyze can't know it),
which cameo_jellyfin.py uses to look up the cast in Jellyfin's own id space.
Document that field in the result-sink output schema header.
Register a KPN event handler in both scene_analyze and scene_preview:
- Overflow events accumulate per-node dropped-frame counts, printed on exit.
- A Closed event from any node other than result_sink at EOF means a stage
died; trip an atomic so the main loop bails out instead of hanging on
'done' forever, and exit non-zero.
Bumps external/KPN to the commit that exposes set_event_handler / NodeEvent.
Consolidate copy-pasted logic across the gallery/run scripts into shared
modules:
- sae_env.py — zero-dependency .env loader (populates os.environ)
- sae_tmdb.py — TMDB API helpers (tmdb_get, person images, id lookups)
- sae_jellyfin.py— Jellyfin API helpers (jf_get, id/URL normalisation)
- sae_gallery.py — image download + gallery.json writing
make_gallery, make_jellyfin_gallery and filter_gallery now import these
instead of carrying their own near-identical copies.
These are regenerable per-run outputs that were polluting the worktree:
calibration caches (*.calib_cache.csv/png), cameo detection run outputs
(cameo_progress.txt, cameo_report.txt), and scene-gap analysis plots.