# scene-actor-extraction — software specification Status: **draft**. Consolidates the requirements for the extraction pipeline, its deployment modes, and its integration contracts. **This is a software spec implementing the system spec at [`../../SPEC.md`](../../SPEC.md).** Requirements that span more than one repo — presence semantics, schema-version coordination, gallery locality, identity keys — are owned there. Where this document restates one, the system spec governs. This is a *requirements* document, not a design doc. Most of what follows is already built; each requirement therefore carries a **Current** / **Gap** note so the document doubles as a work list. Where a requirement is fully met, the gap reads "none". Related documents, which this spec references rather than restates: | Document | Owns | |---|---| | [`../../jRay/SPEC.md`](../../jRay/SPEC.md) | Truth-file format and the Jellyfin plugin's read API | | [`../../JRay-public-server/SPEC.md`](../../JRay-public-server/SPEC.md) | Jmanifest exchange format, cut matching, audio signature (§3) | | [`scripts/optimizer/SCHEMA.md`](../scripts/optimizer/SCHEMA.md) | Embedding-dump HDF5 layout | | [`service-conversion.md`](service-conversion.md) | Idle-GPU worker design (one deployment mode of §B) | | [`methodology.md`](methodology.md) | X-Ray scoring methodology used by the optimizer | --- # Part A — Algorithm requirements The processing chain is a KPN dataflow network (`src/nodes/`). Requirements below are ordered along that chain. ## AR-001 … AR-003 — Face detection Detect faces in sampled video frames. - Detector runs on frames sampled at `sample_fps`, not every decoded frame. The default (1.0) originates from I-frame decode speed, not from an accuracy requirement — it is a **cost knob and may be adjusted**. Under scene-scoped presence (SR-002) a lower rate still answers the question, but it lengthens the interval between samples and so weakens IoU-based association; sweep the two together (VR-002). - **Minimum face size is 40×40 px**, expressed in **original video resolution**, not decoded-frame pixels. Stating it in original space decouples it from `dense_scale`: otherwise a 0.5 downscale silently doubles the effective threshold, and dense mode is exactly what scene detection uses. 40 is **measured, not estimated** — it replaces an earlier 66 px guess. Two studies bracket it, and the difference between them is the whole reason the number is 40 rather than 32: - **VR-005** degrades an already-aligned 112×112 crop and matches it against a native-resolution gallery. Alignment is held perfect, so it isolates the *embedder*: the knee sits at 24–32 px, and 32 px still returns 98.1% TPI. - **VR-013** downscales the **whole frame before the detector**, so detection and landmark regression degrade along with it. End to end, holding 90% of the plateau needs roughly **50 px**, against VR-005's ~22 px. The gap is detection and landmark error, which VR-005 excludes by construction — so VR-005 is an **upper bound on quality**, not a threshold, and reading a floor off it would admit faces in the falling region. **AR-002 therefore takes VR-013's number.** 40 sits below the 50 px plateau deliberately: FPI is 0.0% at every scale in both studies, so resolution loss costs recall and never precision, and an over-tight floor discards presence that SR-002 requires. - Emits bounding box, detector confidence, and 5-point landmarks. - Bounding boxes must be reported in **original video pixel space**. When `dense_scale < 1` downscales the decoded frame, coordinates are rescaled by `bbox_upscale` before leaving the pipeline. - **No fixed cap on faces per frame.** `max_faces` (10, largest-first) is removed so crowded scenes do not systematically lose their background cast — which X-Ray credits as scene members (SR-002). See the backpressure requirement below; the cap is currently the only thing bounding per-frame cost, so it cannot be removed on its own. **Current:** SCRFD-500MF via `face_detector_node.hpp`, thresholds in `config.hpp` (`detector_conf` 0.5, `detector_nms` 0.4), `min_face_px` 40, `max_faces` 10. **Gap:** `min_face_px` re-expressed in original resolution — the value 40 is already correct after VR-013, so what remains is the space it is measured in, not the number; `max_faces` removed, gated on backpressure (AR-004). ## AR-004 — Backpressure Removing the per-frame face cap makes the number of faces entering the pipeline unbounded and content-dependent — a crowd scene can produce an order of magnitude more than a dialogue scene. The network must absorb that by **slowing down**, not by dropping work or growing without limit. - The embedder is the bottleneck and must exert backpressure upstream: when its input is saturated, the detector and decoder block rather than queue. - Channel capacities are currently fixed at 16 (`main.cpp:204-207`) and were chosen against a bounded ≤10 faces/frame. They must be re-derived, and overflow must block rather than throw. - `kMaxFaces` in `identity_matcher_node.hpp:133` **throws** when exceeded. With no cap upstream that becomes a crash on crowd scenes; it has to go or become a batching bound rather than an error. - Memory is the real limit: faces carry 112×112 crops plus 512-float embeddings. Backpressure must engage on bytes in flight, not just item counts. ### The fix is not in this repo **Every node output in KPN uses the dropping `push()`** (`pool_node.hpp:404`, `:710`; also `branch.hpp`, `fanout.hpp`, `interrupt_node.hpp`). A lossless `push_blocking()` — "wait for the consumer to drain instead of dropping; the producer just runs slower" — already exists on both `Channel` (`channel.hpp:144`) and `OutputPort` (`variant_node.hpp:81`), **and nothing calls it.** So AR-004 is a change to the KPN repository, not to this one. It needs either a per-channel lossless policy or a network-wide default, and this pipeline should select lossless: a dropped frame here does not degrade a result, it silently changes one. **Measured, not inferred.** One 77 s clip at 5 fps should yield ~385 sampled frames. On CPU it produced 49, ending at 51 s, with 285 frames dropped at `camera_pos` and 51 at `face_aligner`. Rebuilt with CUDA the same clip ran in 29 s and reached EOF correctly — and still dropped **320** frames at `camera_pos`, yielding 65. Faster hardware moves where the queue backs up; it does not change what happens when it does. Two consequences worth stating: - **Raising channel capacity is a stopgap, not a fix.** It lowers the probability of overflow without changing the behaviour on overflow, and the failure it hides is silent corruption of the output. - **Fixture generation is blocked on this** (VR-001), because what gets dropped depends on timing. The same command run twice can produce different dumps, and a golden fixture cannot be built on that. **Current:** fixed in KPN — node data outputs use `push_blocking`, sentinels remain out-of-band so EOF can always overtake a stalled data path. Verified on the same clip: 385 of 385 sampled frames written, zero drops, and two consecutive runs byte-identical where previously they were not. It also ran *faster* (29 s → 17 s). A dropped frame has already cost its decode, and the overflow exception cost more — so the lossy path was paying for work it then discarded. **Gap:** the remaining half — bounding by **bytes in flight** rather than item count. Channel capacity is still a count of items, and a face carries a 112×112 crop plus a 512-float embedding, so a crowded frame occupies far more memory per slot than a sparse one. That matters once `max_faces` is removed (AR-003). ## AR-005 — Face alignment and crop Produce the exact input ArcFace expects. - 112×112 BGR crop via the standard ArcFace 5-point similarity transform. - Landmark order must match the SCRFD/ArcFace convention (left eye, right eye, nose, left mouth, right mouth). - Alignment is the *only* geometric normalisation; no additional augmentation at inference. - **The transform is fitted by Umeyama's closed-form least squares over all five points**, which is what InsightFace uses (skimage's `SimilarityTransform` *is* `_umeyama`) and therefore what produced the crops ArcFace and LVFace were trained on. The canonical warp is part of the input distribution, not an implementation detail (AR-011). - **Not a robust estimator.** A RANSAC fit buys a small residual by discarding the landmarks that disagree with the model, and on a turned face those are the foreshortened ones — the signal AR-030 reads. With five points and a two-point minimal sample it also cannot separate a mis-detected landmark from honest out-of-plane rotation, so the robustness is nominal while the cost to AR-030 is total. It is RNG-driven besides, which made replay determinism a property of thread scheduling. **Current:** `align_face()` in `src/face_utils.hpp`, Umeyama fit via `umeyama_similarity()`, `cv::warpAffine` to `{112, 112}`. **Gap:** none. > **Migration note — this was a defect, not a refinement.** Until this landed the > fit was `cv::estimateAffinePartial2D(…, cv::RANSAC, 3.0)`. The expectation was > that the two agree wherever RANSAC keeps all five points, leaving a small > divergence on non-frontal faces. **Measured, that is wrong.** On 400 random > gallery headshots, one model held fixed and only the estimator varied: > > | | median | p90 | max | > |---|---|---|---| > | Crop disagreement (source px, over the crop corners) | 16.97 | 75.91 | 223.31 | > | `cos(umeyama, ransac)` for the resulting embedding | 0.791 | — | — | > > 83.5 % of crops embed to a cosine below 0.99 of their Umeyama counterpart — > they are not the same face crop. The mechanism is that a 4-DoF similarity is > exactly determined by **two** points, so every minimal RANSAC sample fits its > own pair perfectly and is then scored on the other three. Real landmarks sit a > median 2.74 canonical px from any similarity fit to the template (see AR-030 > below), so images with a landmark outside the 3 px band are the common case, > not the exception; RANSAC then keeps two or three inliers and returns a wildly > under-determined transform. > > **Every gallery baked before this change must be rebuilt** — GR-004's embedder > stamp catches a model change, not an aligner change, so nothing else would say > so. > > **How much this cost in accuracy is a separate question, and the answer appears > to be: less than the crop numbers suggest.** Rebuilding the full gallery > (2456 actors) moved the intra/inter separation the AR-023 calibration is fitted > from only slightly: > > | | intra-actor | inter-actor | separation | > |---|---|---|---| > | RANSAC | 0.6234 | 0.0407 | 0.5827 | > | Umeyama | 0.6340 | 0.0440 | 0.5900 | > > The reconciliation is that the old warp was *wrong but self-consistent*: it > produced a differently-framed face rather than a scrambled one, gallery and > probe went through the same estimator, and the embedder tolerates framing > variation. So the figures in `model-bakeoff.md`, `best-model.md` and > `pose-expansion.md` were all produced through the broken warp on both sides and > should be re-run, but there is no measured basis for expecting them to move far. > > The sharper evidence of the old instability is duplicate detection: rebuilding > with an unchanged `dedup_tol` dropped **1614** near-duplicate images, where the > original build dropped on the order of a hundred. Near-identical source images > used to embed to visibly different vectors — RANSAC fitting two-point subsets is > unstable under small landmark perturbations, and being RNG-driven it was not > reproducible either. That instability is what a tracker accumulating evidence > across frames pays for, and it is the strongest reason the fix is worth having > independently of any accuracy delta. ## AR-006 — Embedding Generate a 512-d embedding per aligned crop. - Output embeddings are **L2-normalised**, so cosine similarity is a plain dot product. Every downstream threshold assumes unit norm. - Faces are embedded in batches of at most `embed_batch_size` to bound per-call latency. - The embedding model is a build-time choice; the gallery must have been built with the *same* model (see GR-004). **Current:** `embedder_node.hpp` + `face_embedder_engine.hpp`; default LVFace-B_Glint360K. **Gap:** none. ## AR-028 … AR-030 — Embedding input quality An embedder handed a face it cannot represent does not fail. It returns a confident, plausible, wrong vector, and that vector then competes on equal terms with every good one in the gallery — the same failure mode AR-011 names for whole models, occurring here at the level of a single region. Quality assessment is how that is caught **at inference**, rather than inferred afterwards from a study of why a film scored badly. Three axes, assessed on every face before its embedding is used as identity evidence. They are kept separate and **not collapsed into one scalar**: they fail for different reasons, have different remedies, and — as below — do not even earn the same response. - **Size** — already AR-002, floor at 40×40 px in original resolution, measured end to end by VR-013. It is the precedent for the other two: the threshold was *located*, not chosen. - **Sharpness** — motion blur and soft focus destroy the high-frequency detail the embedder keys on, and unlike size they leave the bounding box looking perfectly healthy. Measured on the **112×112 aligned crop**, not the raw box: the crop is already scale-normalised, so a measure taken there cannot silently re-measure face size and double-count it against AR-002. - **Visibility** — extreme pose or occlusion means the face presents fewer of the features the embedding assumes are present. The measure is the **residual of the AR-005 alignment fit**: the RMS landmark error, in canonical 112×112 pixels, left over after the best similarity transform onto the ArcFace template. It costs nothing — the transform is computed for the warp regardless, and the residual is what that fit could not explain. Two properties earn it the job over an explicit yaw estimate: - A similarity absorbs rotation, uniform scale and translation **exactly**, so the residual is by construction the non-similarity part of the deformation: out-of-plane rotation and foreshortening. In-plane roll contributes nothing, so "a tilted head reads as a turned one" is excluded structurally rather than by tuning. The destination frame is fixed, so face size cannot leak in either — that is AR-002's axis, and double-counting it would make a small frontal face look occluded. - It responds to **occlusion** and to plainly broken landmark sets, which an angle regressor by construction does not: a hand across the face is not a rotation, but it does displace landmarks. Indicative magnitudes from a synthetic foreshortening sweep (`k ≈ cos yaw`): `k=1.0 → 0.00`, `0.9 → 1.18`, `0.75 → 3.11`, `0.5 → 6.72`, `0.3 → 9.85` canonical px. Smooth and monotone with a usable range; the mapping onto real faces is VR-012's to establish, and no threshold is set from these numbers. **The synthetic ladder is noise-free and therefore optimistic about the low end.** Measured on 400 real TMDB/Jellyfin headshots — the most frontal, most cooperative population the pipeline ever sees — the residual runs p5 1.11, median 2.74, p90 4.82, max 6.35 canonical px. So landmark noise alone occupies roughly the first 3 px, and the synthetic sweep's "26° yaw ≈ 1.2 px" sits *below* the noise floor on real data. VR-012 must set any threshold against this measured distribution, and a discount curve has to treat the first few pixels as uninformative rather than as mild pose. Neither a dedicated landmark model (`models/2d106det.onnx` is present but referenced nowhere — and it emits points, not pose) nor a direct pose CNN is adopted unless VR-012 shows the residual insufficient. If one is needed the candidate is **6DRepNet** (MIT, RepVGG-B1g2, 3.47° MAE on AFLW2000) rather than Hopenet, which it dominates on accuracy, licence, recency and export friendliness. Two caveats to record before that happens: both are trained on **300W-LP**, which inherits research-only terms from 300W's constituent sets, and both want their own loosely-framed ROI rather than the ArcFace crop — a second warp and a second image in flight, which lands on AR-004's byte-based backpressure gap. It would also have to run **per track** — over the bounded view set AR-019's diversity buffer already keeps — not per face per frame, which is the cost rule applied as written: fewer regions, never a degraded input. **Failing an axis discounts the observation; it does not delete the detection.** Only size drops the face outright, and only because VR-005 measured a knee below which the embedding carries no signal to discount. Blur and pose are different: - A blurred or turned face is still evidence of **presence**, which is what SR-002 actually asks about. - The tracker admits a link on position *or* identity precisely so that a face "whose embedding degraded (blur, profile turn)" stays linkable. Remove the detection and the track fragments, costing the window extent AR-012/AR-013 exist to protect. - AR-019 harvests non-frontal views *because* TMDB headshots are frontal. Discarding turned faces starves the mechanism built to fix the pose problem of its raw material, and AR-020 then has nothing to resolve at EOF. The natural home for the discount is `EvidenceDiscounter` (AR-025), which already weights how far one observation may move a track's belief. Note that its present weight is pure *novelty*, so a profile view — maximally distant from everything counted so far — currently scores near 1.0 and moves the belief hardest, when against a frontal gallery it deserves the least trust. Novelty and reliability are orthogonal and multiply; quality supplies the second term. **Quality is carried, not consumed.** The vector travels with the face and is written to the VR-001 dump alongside the embedding, so a threshold can be re-litigated against recorded data instead of by re-running video, and so VR-010's provenance records what the run actually admitted. **No quality threshold is hand-set.** Each axis either has a measured knee (VR-012, as VR-005 did for size) or it discounts rather than drops — a hand-chosen cutoff on an uncalibrated measure is the same unfalsifiable magic number AR-024 retired for similarity, and it would fail the same way: meaning something different for every detector, every embedder and every film. **Current:** visibility is measured and carried — `estimate_alignment()` in `src/face_utils.hpp` returns the residual alongside the transform, and `FaceAlignerFunc` writes it to `DetectedFace::alignment_residual`. Size is `min_face_px` (40, decoded-frame space — AR-002 still open). Sharpness is unmeasured. Nothing yet *consumes* any of it: no discount is applied, and `align_face()` still drops the degenerate-fit case without counting it. **Gap:** AR-029 entirely. For AR-030, the measure exists but the discount does not — it must reach `EvidenceDiscounter` as the reliability term. For AR-028, the residual does not yet reach the VR-001 dump, which is what VR-012 needs to run from fixtures; that is the next step, since it unblocks the study that sets every remaining behaviour. ## AR-007, AR-008 — Tracking Link detections across frames into tracks representing one physical person. Association cost combines three signals: - **Spatial** — IoU between the candidate detection and the track's last box. Meaningful only for tracks seen in the immediately preceding frame. - **Appearance** — similarity to the track's running directional mean embedding (averaged, then re-normalised to the unit sphere), **expressed as a probability** (AR-024), never as a raw cosine. - **Weighting** — `track_alpha` interpolates the two, and is **frame-dependent**: on `is_cut` / `is_scene_boundary` (AR-009, AR-010) it drops toward embedding-only, because position carries no information across a viewpoint change. **There is no separate re-acquisition mechanism.** A track whose face is lost sets `last_seen` and stays a candidate for association until extinction; matching it to a later detection is ordinary inter-frame tracking, not a distinct revival path. The property falls out of the embedding comparison the tracker already does. Consequences: - **One candidate pool**, not an active set plus a parked set. `last_seen` alone distinguishes them, and it only affects whether IoU means anything — dormant tracks are matched on embedding, since time has passed and position is stale. - The current cross-cut park/revive path (`cut_revive_sim`, `cut_inactive_max_frames`) is this same mechanism special-cased to cuts. It collapses into the general path. - A track dies only by extinction: `last_seen` set for longer than the timeout. **Current:** `face_tracker_node.hpp` maintains separate `tracks_`/`inactive_` maps with distinct thresholds and a revival branch. **Gap:** unify into one pool keyed on `last_seen`; make `track_alpha` frame-dependent; move association into probability space (AR-024). ## AR-009 … AR-011 — Cut and scene-boundary detection Two distinct signals, deliberately kept separate: - **`is_cut`** — always on. Grayscale histogram correlation below `cut_threshold` flags an *intra-scene camera-angle change* (shot/reverse-shot). - **`is_scene_boundary`** — opt-in (`--scene-detect`). TransNetV2 over a densely decoded, downscaled stream flags a *true shot/scene boundary*. > **`is_scene_boundary` currently has no producer.** `grep -rn is_scene_boundary > src/` finds no assignment anywhere: `SceneDetectorFunc` is a *terminal sink* > (`main.cpp:298-300`, `kpn::out<>`) that writes `scenes.json` and never > annotates the `Frame` flowing to the face pipeline. The field is therefore > always `false`, and the dump column (`embedding_dump_node.hpp:38`) is a > constant 0. Compounding it, `main.cpp:280` returns from the > `--dump-embeddings` branch *before* the `scene_detect` branch at `:296`, so no > dump-producing path even instantiates the detector. > > > **It cannot be fixed by making the node a pass-through.** TransNetV2 buffers > `kWindow` = 100 dense frames before it can score any of them, runs inference > every `scene_stride` (50) frames, and trusts only each window's centre. So a > boundary at time *T* is not known until roughly 100 dense frames after *T* — > about **3.3 s at 30 fps**. The face pipeline runs on a parallel branch and has > long since passed *T* by then. An association hint that arrives after the > association is worthless. > > Three ways out, none free: > > 1. **Two-pass.** Run scene detection to completion, then analyse faces with > boundaries already known. Simple and correct; costs a second decode of the > whole file, and dense decode is already the pipeline's dominant cost. > 2. **Delay the face branch** by the detector's window latency. Keeps one pass; > adds a buffering stage and couples the two branches' timing, which is the > kind of coupling that produces heisenbugs under backpressure. > 3. **Leave it unwired.** Accept that `is_cut` is the only association hint. > > **Option 3 costs less than it appears**, which is why this is a decision rather > than a bug. Since the redesign made cuts and boundaries do the *same thing* — > both say "spatial continuity is broken, associate on embedding" — TransNetV2 > adds nothing over the histogram except on transitions the histogram misses: > slow dissolves and fades, where there is no frame-to-frame discontinuity to > detect. That is a real but narrow gap. > > The value TransNetV2 retains is in **AR-019**, whose promotion gate requires a > span with no cut *and* no boundary. There a late answer is still usable, > because promotion happens when a track is confirmed rather than per frame. > Wiring it there — offline, against the collected boundary list — is cheaper > than any of the three options above and does not touch the hot path. > > **Recommendation: option 3 plus the AR-019 wiring**, and revisit if dissolve- > heavy material shows association failures the histogram misses. Both feed AR-007 as **association hints**: they tell the tracker that spatial continuity is broken and that association should weight embedding over IoU. Neither ends a presence window (AR-012). In dense mode the source decodes at `scene_decode_fps` (default 0 = native) and a decimator splits the stream: full-resolution sampled frames to the face pipeline, downscaled dense frames to the scene detector (`frame_source_node.hpp:63`). `sample_fps` is independent of this — the face pipeline's rate is not derived from the scene detector's. Boundary timestamps are keyed off each frame's real timestamp, so they stay correct despite the reduced rate. `dense_scale` must stay ≥ 0.5 on 1080p sources — it also shrinks what the face detector sees. ### Every model gets the input it was trained for A general rule, stated once here because it applies throughout: **models are fed their expected input, not a cheaper approximation.** Where cost must come down, it comes down by running the model less often or on fewer regions — never by degrading what a single inference sees. A model run off-distribution produces confident, plausible, wrong output, and the error is invisible without a study that should not have been necessary. Two places this was violated, both now closed: 1. **`scene_decode_fps = 12` starved TransNetV2.** `kWindow` is 100 frames. At native 25 fps that window spans ~4 s; at 12 fps it spanned ~8.3 s, so the model saw roughly half-speed motion over twice the temporal context it was trained on. **Requirement: feed TransNetV2 at the source's native frame rate**, so a 100-frame window covers the duration the model expects. The "tolerates ~12fps" note in `config.hpp` described a compromise, and the recorded margin was consistent with it — a non-boundary baseline at ~0.50 with real boundaries reaching only ~0.7+ is a compressed separation, not a healthy one. **Done:** `scene_decode_fps` defaults to 0. 2. **Hardcoded 25 fps in boundary dedup.** The node merged boundaries closer than `0.04 s` — "~1 frame @25fps". **Requirement: derive this from the source's actual frame rate. Done:** `SceneDetectorFunc::dedup_window_sec()` takes the median of the frame intervals the detector was actually fed and halves it. Half a frame rather than a whole one, because the only thing being merged is one frame scored by two overlapping windows; two distinct frames are a full interval apart and both have to survive. The two are one change, not two. A native-rate stream is where the old constant did the most damage — at 30 fps, 0.04 s is wider than a frame, so two cuts on consecutive frames merged into one and the loss showed up nowhere: the file simply had fewer boundaries. Dense decode is the pipeline's cost driver, so (1) is not free. The cost is accepted: the alternative is a boundary signal that steers association (AR-007) while being quietly unreliable. `dense_scale` remains available as a spatial reduction, since downscaling is a documented, understood degradation rather than a temporal one the model has no defence against — and TransNetV2 downsamples to 48×27 regardless. **Current:** histogram cut in the decoder; `scene_detector_node.hpp` for TransNetV2, fed at native rate with a framerate-derived dedup window. **Gap:** `scene_threshold` (0.60) is still the value picked against 12 fps input and is now certainly wrong — VR-006 re-fits it, and until it does, boundary recall at native rate is untuned rather than better. `--scene-detect` is default-off despite now feeding association. ## AR-012 … AR-017 — Track-level identity propagation — **CHANGED BEHAVIOUR** > **Requirement:** when a face is identified, it is considered identified for the > **entire track**. A track is one physical person by construction (AR-007), so an identification anywhere on the track is evidence about all of it. Presence windows must therefore be derived from **track extents**, not from the subset of frames that happened to match. Required semantics: - A track is **owned** by actor A once the accumulated posterior for A exceeds the ownership threshold (AR-025). This is a Bayesian accumulation over the track's frames, not a count of accepted frames — so a single decisive observation and a long run of marginal ones are distinguished rather than conflated. - Once owned, the actor is present for the track's **full extent** — `[first_seen, last_seen]`, including frames where the face was turned, occluded, or simply scored below threshold. - **A track with no owner emits nothing.** No posterior crossed, so there is no claim to make. Such tracks are captured for diagnosis (AR-022) but do not appear in output. See "Unidentified presence" below. - **Every presence claim carries its belief.** The accumulated posterior that justified the claim travels with it — through the aggregator, into the truth file (IR-001), and onward. Confidence is not a serialisation concern bolted on at the boundary; the pipeline computes belief for every decision it makes (AR-025), so discarding it and emitting a bare interval throws away information that was free. A claim also carries **how it was identified** — live, deferred (AR-020), or pooled from a cluster (AR-021) — because those routes rely on different evidence. A deferred identification leans on harvested references rather than baked ones; a pooled one aggregates across tracks. The posterior already reflects the strength of each, but the provenance tells a consumer *why*, and tells us which route is producing errors when one is. ### Belief swap terminates the track If a track is owned by A and the accumulated belief subsequently swaps to a different actor B, **the track is terminated and a new one started for B**. The old track closes normally at its `last_seen`; the new track begins at the frame where the swap occurred. The rationale is that a swap is not a correction — it is almost certainly a **missed camera or scene change**, where the tracker carried one `track_id` across a viewpoint change and began following a different person. Two genuinely different individuals both accumulating a threshold-crossing posterior on the same face is not realistic short of identical twins; a `track_id` spanning two people is. Treating it as a swap-and-continue would emit one window blending two people. Treating it as a track boundary yields two windows that are each correct. Ownership is therefore **established at first crossing, not deferred to track death** — the first crossing is a real claim about the frames seen so far, and a later contradiction ends that claim rather than revising it. ### Simultaneous ownership If two actors cross the threshold on one track at effectively the same time, the **highest posterior wins**. This should be rare — the swap rule above catches the common form of the problem — and it must be counted and reported (GR-003) as a track-ID collision indicator. ### Identity contradiction is a cut detector If **two live tracks are both owned by the same actor**, at least one is wrong: a person cannot be in two places at once. The cause is the same as the belief swap — a missed camera or scene change that split one person into two tracks, or attached an identity to the wrong one. **Requirement: treat this as a detected cut.** Reset the affected tracking state and re-associate on embedding, exactly as an `is_cut` signal would (AR-007, AR-009). This gives a third cut-detection signal, derived from *identity* rather than pixels, and it fires precisely where the pixel-based detectors failed — a cut subtle enough for the histogram and TransNetV2 to miss is not necessarily subtle in identity space. It is also self-correcting rather than diagnostic: the contradiction is detectable **online**, the moment both tracks hold the belief, not at output time. A consequence worth noting: with this in place, overlapping windows for one actor should be rare rather than routine, because the condition that produces them is now caught and repaired while tracking. Any that survive to output indicate the repair failed and should be counted. ### Unidentified presence — TBD A track that is never owned is still *someone* on screen. Emitting it as anonymous presence would let a consumer show "unidentified person", and would give the human-in-the-loop association tool ([`../../SPEC.md`](../../SPEC.md) §4) its work queue directly. Combined with the context crops of AR-022, it may also be the path to recognising extras and background cast the gallery has no entry for. **This is deliberately undecided.** It changes the truth format and invites consumers to display something that may not be useful. Recorded here so the option is not lost; not specified until the AR-022 debug output shows whether these tracks are worth surfacing. **Current:** presence is built from *per-frame accepted detections only*. `identity_matcher_node.hpp:145-231` decides acceptance independently per face per frame; `result_sink_node.hpp:123-129` collects the timestamps of accepted frames; `result_sink_node.hpp:139-147` sets `win_start = ts_vec[0]`, the first accepted timestamp. `track_id` is carried on `IdentifiedActor` but is used only for debug output — it never gates or backfills a window. A window therefore starts when the actor was first *recognised*, not when their track began. **Gap — this is the main behavioural change in this spec.** Design is settled in [`plan.md`](plan.md) under AR-012/AR-013: a `TrackRegistry`, held by `shared_ptr` and used by `FaceTrackerFunc` as its state, where each track carries `first_seen` plus an **optional `last_seen`** — unset while on screen, set to the last on-screen timestamp when the face is lost, unset again when a later detection associates to it. A track whose `last_seen` exceeds the extinction timeout is reaped and **pushed to the result aggregator** as one finished claim: *this actor was on screen from a to b*. Consequences: 1. The tracker's `tracks_`/`inactive_` maps become one track set, distinguished only by whether `last_seen` is set; there is no separate revival path (AR-007). 2. Closing a track *is* the presence assertion — emitted once, complete, never revised. No later reconciliation stage exists. 3. **`anneal_sec` and `extinction_sec` are deleted, not re-tuned.** Both exist only to bridge gaps between isolated accepted frames; a track that survives its own gaps leaves them nothing to do. `SceneTrackerFunc` goes with them. What remains to tune is the **ownership posterior threshold**, the **correlated-frame discount** (AR-025), and the **extinction timeout**. 4. Every track must be closed at EOF. A film ends with faces on screen and those tracks have not timed out, so without an explicit flush the closing scene's actors are never emitted — a silent presence loss that looks like a recognition miss. 5. Zero-length windows (`start == end`, a single-frame track) remain possible; the truth format permits them, and `build_xray()` floors/ceils into integer seconds. 6. **Pull it out by the roots — including the published field.** `anneal_sec` is not just a constant: it appears in the truth schema ([`jRay/SPEC.md`](../../jRay/SPEC.md)), in the Jmanifest format and in the server's storage columns ([`JRay-public-server/SPEC.md`](../../JRay-public-server/SPEC.md)). Retaining it as a vestigial `0` would be worse than removing it: a field that names a mechanism no longer in the pipeline is actively misleading to anyone reading a manifest, and it would outlive everyone who remembers why it is zero. It goes from all three repos under one coordinated `schema_version` bump — the same bump IR-004 already requires, so there is exactly one breaking change, not two. Removal list: `Config::anneal_sec`, `Config::extinction_sec`, `SceneTrackerFunc` and its node wiring, the sink's annealing pass (`result_sink_node.hpp:139-147`), the truth-file field, the Jmanifest field, the server column, and the optimizer's parameter entries. Grep for both names and expect no survivors. ## AR-018 … AR-021 — Per-film gallery expansion and deferred identification Three requirements that only make sense together, so they are specified together. ### The problem being solved **TMDB headshots are overwhelmingly frontal.** Films are not. An actor is recognised easily in the shots that resemble a publicity still and missed in profile, three-quarter, low-angle, poorly lit, or partially occluded views — even though those are most of their screen time. So the pipeline's failures are dominated by **pose**, not by identity: the same person the gallery knows, at an angle it does not. Both mechanisms below exploit that. ### AR-018 — Per-subject embedding store Every track maintains a running store of its own embeddings, **identified or not** — the structure is the same for both. An embedding is admitted only if its similarity to one already in the store falls **inside a band**: - **Upper bound** — too similar and it is redundant, teaching nothing the store already covers. - **Lower bound** — too dissimilar and it may not be the same person at all; admitting it risks poisoning the store. A starting band of roughly **0.90–0.95** is the working estimate, to be tuned (VR-007). Note this is deliberately conservative compared to the retired `expand_novelty_sim` (0.55), which promoted embeddings *far* from the gallery — much more aggressive, and much more exposed to admitting the wrong person. Both bounds must be expressed as calibrated probabilities, not raw cosines (AR-024). The lower bound is asked twice. `admit` compares a newcomer against its *closest* existing member, which a gradually drifting track can chain past: every step inside the band while the endpoints are strangers — the shape a track-ID collision takes over a slow pan. So the same bound is re-applied across **every pair** in the store before promotion. One bound, two enforcement points; not a second constant. Novelty is deliberately **not** a threshold. The store's eviction policy orders its members by similarity to the actor's existing references and drops the best-recognised one, so novelty-seeking is a ranking with nothing to tune, and the band's upper bound already refuses the redundant views at the door. **Current:** implemented in `gallery/track_gallery.hpp` — `admit` at the door, `store_coherence` at promotion, both bounds from `Config::expand_band_lo/hi`. Refusals are counted (`band_rejected`). **Gap:** the bounds themselves are unswept working values (VR-007). ### AR-019 — Expansion of known actors When a track is owned (AR-012), its store is promoted into a **per-film, in-memory annex** for that actor, folded into best-of-N scoring alongside the baked references. The annex does not persist. Promotion requires **certainty that the span is one person**: - the track is owned, with the belief stable (no swap, AR-012); - no camera cut, scene boundary, or identity contradiction occurred within the span (AR-009, AR-010, AR-015) — all three signals must be quiet, not just the histogram cut as today; - the band of AR-018 is satisfied. The purpose is precisely to acquire the **non-frontal views TMDB lacks**, at a confidence the gallery alone cannot supply. ### AR-020 — Deferred re-identification of unknown tracks **This is what the expansion is for.** - An unowned track, on extinction, is **not discarded**. Its embedding store, metadata and context crops (AR-022) move to a **to-be-identified (TBI) queue**. - At **end of playback**, TBI entries are pooled (AR-021) and re-assessed against the *final* expanded gallery — which by then holds the pose-varied views harvested from the whole film. - Entries that now cross the ownership threshold emit presence windows exactly as a normally-owned track would. Entries that still do not are the output of AR-022, and the work queue for human association ([`../../SPEC.md`](../../SPEC.md) §4). The asymmetry this exploits: an actor confirmed frontally early in a film contributes profile views to the annex, and a profile-shot track that failed at minute 12 matches once the film is over. Ordering ceases to matter. **This is cheap.** Embeddings are already computed; re-assessment is matching against an in-memory annex — no decode, no detection, no embedding. The cost is retaining unowned track stores until EOF. ### AR-021 — Pool unknown tracks before matching Re-assessing each unknown track alone wastes the strongest evidence available. A single track is short and pose-poor; the *same unknown person* usually appears across many tracks throughout the film. **Requirement: cluster the unknown tracks, treat each cluster as one identity, and match the pooled cluster against the full gallery.** Each unknown track's store is effectively a small gallery of one unnamed person. Clustering merges the ones that are the same person, and the pooled result is a far richer representation — many poses, lightings and expressions — which stands a much better chance against the gallery than any constituent track. **The evidence is also better-conditioned.** AR-025 warns that consecutive frames within a track are highly correlated and must be discounted. Embeddings from *different tracks* are far more independent: different scenes, angles, lighting. Pooled cross-track evidence is therefore worth more per observation than within-track evidence, and the discount should reflect that. **Temporal exclusion is a free constraint.** Two tracks that **overlap in time cannot be the same person** — the same fact AR-012 uses to detect missed cuts. This is a cannot-link constraint on the clustering and it costs nothing to apply, since track extents are already known. Use it: it prevents exactly the merge that would otherwise pool two people who share the screen. **Clustering must be conservative.** A wrong merge pools two people and then mis-identifies *both*, across every track in the cluster — strictly worse than leaving them separate. Prefer many small correct clusters to few large ambiguous ones; the merge threshold is a calibrated probability (AR-024), swept in VR-007. **Order of operations at end of playback:** 1. Cluster TBI tracks under temporal cannot-link constraints. 2. Pool each cluster's embeddings into one composite identity. 3. Match each cluster against gallery + annex (one GEMM, A10). 4. Cluster crosses threshold → **every member track** emits presence for that actor. 5. Cluster does not → it becomes one **unknown person** entity, not N orphan tracks. That last point matters beyond recognition: it is the difference between asking a human to label twelve disconnected faces and asking them to name one person who appears in twelve places ([`../../SPEC.md`](../../SPEC.md) §4). It is also the natural unit for anonymous presence, should that be adopted (AR-012, TBD). ### Consequences - **The result aggregator cannot finalise at EOF-flush.** Owned tracks emit on death as before, but the TBI pass runs after, and may add windows. Output is written after re-identification completes, not when the last frame is read. - **Expansion errors are now more costly.** A wrongly promoted embedding no longer affects only later frames — it is applied to every unknown in the film during the second pass. The AR-018 band is the guard, and its lower bound is the part doing that work. - **Iteration is possible but unspecified.** A TBI entry that becomes identified could itself expand the gallery and enable further identifications — a fixpoint loop. Whether to iterate to convergence, run one pass, or bound the rounds is open (VR-007). **Current:** `src/gallery/track_gallery.hpp` implements a per-track diversity buffer with eviction biased to gallery-far poses, admission and promotion both gated on the AR-018 band in probability space, cleared on `is_cut`. Wired at `identity_matcher_node.hpp:227`, cleared at `:126`; the calibration is handed over at `:114`. **Gap:** all three quiet-signal conditions rather than only `is_cut`; and the whole of AR-020 — the TBI queue, the deferred pass, and deferring output until it completes. ## AR-022 — Unidentified-track capture Persist everything needed to diagnose a miss — and everything a *human* would need to resolve one. Scope note: the TBI queue of AR-020 is **not** debug-only, because deferred re-identification depends on it. What is behind a flag is the *persisted output* for tracks that survive re-assessment still unidentified; retaining stores until the deferred pass runs is unconditional. For every track **still unidentified after the deferred pass (AR-020)**, store: - all embeddings on the track, - the aligned 112×112 crops, - **context crops** — a wider region around the face than the 112×112 aligned crop, for a handful of representative frames per track; - track metadata: `track_id`, first/last timestamp, frame count, per-frame detector confidence and bbox, - the best similarity achieved and which actor it was against, so near-misses are distinguishable from faces with no gallery counterpart at all. **Why the context crop is a separate artifact.** The 112×112 crop is optimised for ArcFace: tightly cropped, geometrically normalised, and frequently unrecognisable to a person out of context. A human deciding *who this is* needs the surrounding shot — hair, costume, who they are standing next to. This requirement exists to serve the human-in-the-loop association capability ([`../../SPEC.md`](../../SPEC.md) §4), where the user names the face the pipeline could not. Retain a bounded number of representative frames per track (largest/sharpest detections), not every frame — an unidentified track can run for minutes. This is the raw material for deciding whether a miss is a gallery coverage problem (actor absent or under-represented) or a threshold problem (actor present but scored below acceptance). **Current:** only *promoted* mugshots are dumped, via `expand_debug_dir` (`track_gallery.hpp`, guarded by `SAE_DEBUG`) — i.e. the successes, not the failures. **Gap:** the whole requirement. Needs a flag (`--dump-unidentified `), a per-track accumulator that survives until the track ends, and a decision on retention: crops for a feature-length film are large, so default to embeddings + metadata with crops opt-in. ## AR-023 … AR-025 — Calibration, probability space, and per-track Bayesian accumulation **Every similarity in the pipeline is converted to a probability through the sigmoid calibration before it is used or thresholded. No component compares raw cosine values against a hand-set constant.** This is a system-wide rule, not a detail of identity matching. Raw cosine thresholds are unfalsifiable magic numbers that mean different things for different models, different galleries, and different face sizes; a calibrated probability means the same thing everywhere. It replaces `track_max_embed_dist` (0.7), `cut_revive_sim` (0.50), `expand_novelty_sim` (0.55) and `expand_track_spread_max` (0.60) with quantities that can be reasoned about jointly. ### Per-frame evidence is a Bayesian update on the track Once similarities are probabilities, a track's identity is not a vote count — it is a **posterior accumulated across the frames of that track**. Each frame contributes a likelihood that this face is actor A; the track's belief is updated frame by frame, and ownership (AR-012) is "posterior exceeds threshold" rather than "≥ N accepted frames". Working in **log-odds** makes this an addition per frame: ``` logit(A | frames 1..n) = logit_prior(A) + Σ_i [ logit(A | sim_i) − logit_prior(A) ] ``` This is strictly better than counting accepted frames: a long run of marginal observations and a single decisive one are no longer conflated, and a track that accumulates weak contrary evidence can lose ownership rather than keeping it on a stale count. **The independence problem, which must be handled explicitly.** Consecutive frames of one track are highly correlated — near-identical pose, lighting and expression. Treating them as independent observations overcounts evidence dramatically: thirty frames of the same face at the same angle is nowhere near thirty independent measurements, and naive accumulation will drive the posterior to certainty on what is effectively one observation. Mitigations, in preference order: 1. **Update only on sufficiently novel observations.** The diversity buffer (AR-018…AR-021) already identifies which embeddings on a track are gallery-far and mutually distinct — precisely the more-independent ones. Reuse that judgement rather than inventing a second one. 2. **Discount correlated updates** by a per-frame weight below 1, fitted so the accumulated posterior is calibrated against held-out tracks. 3. **Cap total evidence per track**, the crude fallback. Whichever is chosen, the accumulated posterior must be **validated against ground truth** — a posterior of 0.99 should be wrong about 1% of the time. If it is not, the independence handling is inadequate and the number is decoration. ### Fitting the calibration Build a PDF of **intra-subject** similarity and a PDF of **inter-subject** similarity across the gallery; the prior is `intra / (intra + inter)`. - **Positive pairs** — same actor, different reference images. - **Negative pairs** — all cross-actor pairs. - Near-duplicate references are de-duplicated per actor first (similarity above `1 − 1e-7`), so the same image embedded twice cannot inflate the positive side. - Actors with fewer than 5 distinct embeddings contribute negatives only — they cannot supply a meaningful positive pair. - A Platt-style sigmoid `P(match) = σ(a·sim + b)` is fitted to the two distributions with class weights balancing the (heavily skewed) pos/neg ratio. - The base rate enters as log-prior-odds: `P = σ(a·sim + b + log(p₀/(1−p₀)))`, with `p₀ = match_prior`. - Acceptance is `P(match | sim, prior) > prob_threshold` (default 0.754, DE-tuned). - **Fallback** when calibration is invalid (too few positive pairs): a hard cosine distance ceiling `match_threshold`, OR a ratio test — accept if `best/second < match_ratio` and `best_distance < match_ratio_ceil`. **Current:** `src/gallery/gallery_calibration.hpp` implements the sigmoid fit, dedup, eligibility filter, and prior adjustment; the fallback lives at `identity_matcher_node.hpp:181-196`. **Gap:** the fit currently histograms the two similarity distributions internally (`kHistBins = 200`) but does not *emit* them. For this requirement to be inspectable, the intra/inter PDFs and the derived prior should be written alongside the gallery, so calibration quality is auditable rather than implicit. Also note the shipped `match_prior` default is 0.5 (use the calibrated sigmoid directly) rather than the gallery-derived `intra/(intra+inter)` — reconcile: either compute and store it at gallery-build time, or document 0.5 as a deliberate override. ## AR-026, AR-027 — Scale — performance must hold for arbitrary gallery size **Requirement: analysis throughput must remain acceptable as the gallery grows arbitrarily large.** Gallery size is set by the user's library, not by us: a whole-library gallery spans every credited actor across every title (GR-001), which is thousands today and grows monotonically as titles are added. A design that is fast at 500 actors and unusable at 50,000 has a defect, not a limitation. ### Similarity is a matrix multiply Every embedding is unit-norm (AR-006), so cosine similarity is a dot product and scoring a batch of faces against the gallery is one GEMM. At library scale that is the only viable formulation — a per-pair loop is orders of magnitude off. **All similarity computation goes through the GEMM path**, with no exception justified by "this set is small". Three call sites: 1. **Baked gallery** — already GEMM (`sim_engine_->compute()`, `identity_matcher_node.hpp:143`, backend from `SAE_GEMM_BACKEND`). ✓ 2. **Per-film annex** — currently a **CPU loop** (`identity_matcher_node.hpp:159-162`), justified in-comment by "tens of embeddings". AR-018…AR-021 invalidates that assumption: every owned track now contributes, so the annex grows with cast size and film length. It must move into the GEMM path — appended to the gallery matrix, or a second multiply. 3. **Deferred TBI pass (AR-020)** — the most GEMM-friendly operation in the pipeline: all TBI embeddings against the full gallery-plus-annex, offline, operands resident, no streaming. One large multiply, not a loop over entries. This constrains AR-018…AR-021's implementation: the annex must be a **contiguous matrix** with promotions appended, plus a parallel actor-index mapping — exactly the `flat_emb_`/`flat_actor_` arrangement the baked gallery already uses. ### Scaling characteristics that must be known, not assumed - **Throughput versus gallery size must be measured** (VR-008) and published. The useful output is the curve and the point where gallery scoring starts to dominate total runtime, not a single number. - **Memory is a real ceiling.** 512 floats × 4 bytes = 2 KB per reference. At five references per actor that is ~10 MB per 1000 actors — comfortable at 10,000 actors, worth planning for beyond. - **Calibration cost grows quadratically.** AR-023 fits on cross-actor pairs, which is O(N²) in references. This is a gallery-build cost, not a per-title one, but it will bite first — sampling negatives rather than enumerating them is the obvious mitigation and should be specified before it becomes urgent. ### If GEMM stops being enough Approximate nearest-neighbour indexing (IVF/HNSW) is the standard next step, and it trades **exactness** for speed. That trade interacts badly with AR-023: an approximate search returns approximate similarities, and a calibrated posterior built on them is no longer calibrated. Not ruled out, but it requires re-validating calibration against the approximation — so it is a later decision with a real cost, not a drop-in. **Gap:** annex GEMM path; the scaling benchmark (VR-008); negative-pair sampling in calibration. --- # Part B — Deployment requirements The pipeline must support **multiple deployment modes over one core**. Modes differ in *what triggers work* and *what constrains it*, not in what the analysis does. ## DP-001 — Common core - One analysis implementation, exercised identically by every mode. Modes are front-ends; none may fork the pipeline logic. - Backend selection (ORT/TRT, CUDA/ROCm/CPU) is a **build-time** choice (`SAE_INFERENCE_BACKEND`, `SAE_GEMM_BACKEND`). Prebuilt TRT engines may be supplied at runtime via `detector_engine` / `arcface_engine`, bypassing ORT. - Models load once per process. Any mode processing more than one title must amortise model and gallery load across titles. - No mid-video checkpointing. A run either completes and emits a result, or emits nothing (see DP-004). ## DP-002 — Mode: batch CLI One-shot invocation over one title. The reference mode and the substrate for the others. - `scene_analyze -o out.json`, exit non-zero on failure. - Must be safe to invoke concurrently by an external scheduler, subject to GPU memory. **Current:** `src/main.cpp`. **Gap:** none. ## DP-003 — Mode: on-demand service A resident process on a server, analysing on request. - Models and gallery stay resident; requests carry a media path plus optional overrides. - Requests are queued with a bounded depth and processed serially per GPU; the queue must be observable (depth, in-flight title, ETA). - Health endpoint reporting model, gallery fingerprint, backend, and GPU state. - Graceful drain on shutdown: stop accepting, finish or abandon in flight per DP-004. **Gap:** not built. `scripts/run_from_jellyfin.py --worker` is a polling loop, not a request-driven service; it is the closest existing shape. ## DP-004 — Mode: opportunistic / idle-triggered Analyse when the machine is otherwise unused; yield the instant it is not. - **Trigger is external and non-oscillating.** Screen-lock (logind `Lock`/`Unlock`) is the reference signal. GPU/CPU load must *not* be used: the worker is itself the load, so a load threshold forms a feedback loop. - **Stop is a hard stop.** On the resume signal, SIGTERM the worker mid-analysis. - **Re-queue is free and implicit.** An item leaves the pending queue only when its result is pushed, so a killed run simply stays pending. This requires: 1. never push a partial result — push only after the analysis returns cleanly; 2. clean up temp files on signal — write temps under a directory wiped on start, and unlink on SIGTERM (a SIGKILL skips `finally`). - Accepted trade-off: a partially analysed title restarts from scratch. Fine for an overnight workload. - Other triggers (idle timer, AC power, scheduled window) must fit the same contract: external signal, hard stop, implicit re-queue. **Current:** designed in detail in [`service-conversion.md`](service-conversion.md) as `sae-worker.service` + `sae-lock-gate.service` under systemd **user** units. **Gap:** unbuilt; the temp-cleanup fix in `run_from_jellyfin.py` is a named prerequisite. ## DP-005 — Installation and provisioning - Native install, **no Docker at runtime** — GPU passthrough is the most fragile part of a containerised setup and exists only because of the container. Natively the GPU works with the host drivers and media paths need no re-mounting. This constrains how the software *runs*, not how it is *built*: DP-008 uses containers as build environments precisely because that side has none of these problems. - The installer may **fetch a prebuilt binary** (DP-008) instead of compiling. Compiling stays supported, but should not be the only path — it is the slowest and most fragile step of a first install. TRT engines are still built locally either way (DP-008). - An installer (`scripts/build_install.py`) consuming one `install.yaml`: platform (nvidia/amd/cpu), embedder model, gallery scan cadence, install prefix; runtime secrets written to a `.env`, editable without recompiling. - Distro coverage: Fedora + Arch (`dnf`/`pacman`), auto-installing dependencies after printing them. Debian/Ubuntu out of scope. - Model acquisition (`scripts/download_models.sh`) and TRT engine build (`scripts/build_trt_engines.sh`) are provisioning steps, not runtime steps. **Gap:** installer unbuilt. ## DP-007 — CI build image CI runs on an Intel N100 with no discrete GPU, so the test build must configure **CPU-only** and must not require CUDA, TensorRT or ROCm: ``` -DSAE_INFERENCE_BACKEND=ORT -DSAE_GEMM_BACKEND=CPU ``` A prebuilt container image supplies the toolchain, published to the **Gitea container registry** and pinned by tag — matching the `jellytau-builder` precedent. Building dependencies per CI run is untenable on an N100, and OpenCV 5 from source would dominate every run. The same registry stores corpus dump fixtures as generic packages (see the fixtures table in `requirements.md`). Rebuild the image when its dependency set changes, not per run, and pin CI to a tag rather than `latest` so a rebuild cannot silently change what a green build meant. **Required in the image:** | Dependency | Why | |---|---| | CMake, C++ toolchain, pkg-config | Build | | **OpenCV 5** | `CMakeLists.txt:25` prefers 5, falls back to 4. The branch targets 5, so the image should carry it — it is not yet in most distro repos and building it per-run is prohibitive | | HDF5 (C++) | Galleries are HDF5-native; also the dump format | | FFmpeg dev libs — `libavformat`, `libavcodec`, `libavutil`, `libswscale`, **`libswresample`** | Decode. See the note below on swresample | | Python 3 + numpy, h5py, scipy | Python-side tests, replay, traceability tooling | | **OpenBLAS** | Backs the CPU similarity GEMM. Without it the fallback is a scalar loop, and the CPU path is exactly what this host runs — see below | | Catch2, nlohmann/json | **Vendored into the image, not fetched.** Both are `FetchContent`-ed today (`CMakeLists.txt:220`, `tests/CMakeLists.txt:8`), which makes every CI run depend on GitHub reachability | **OpenBLAS is not optional here, despite being optional in the build.** CI has no GPU, so `SAE_GEMM_BACKEND=CPU` is the only path it exercises — and since AR-003 removed the per-frame face cap, a crowded frame scores many faces against a library-scale gallery. The scalar fallback is correct but scales badly, which would make the CPU path the bottleneck in the one place it cannot be avoided (AR-027). The build warns when it is missing rather than failing, so a developer without it still gets a working tree; the image must not be that case. The test target links it too. Otherwise the suite compiles the scalar fallback while the image ships CBLAS, and CI would verify a kernel that is not the one running in production. **Deliberately excluded:** CUDA, TensorRT, ROCm — no GPU to use them. Also the ONNX Runtime *GPU* providers; only the CPU provider is relevant, and only for T3 smoke tests. **Models are not baked into the image.** The seven ONNX files total ~725 MB and live in Git LFS. T1/T2 tests are model-free by design (`tests/CMakeLists.txt:1-4`), so the default image needs none. T3 smoke tests require a model and should pull it via LFS in a separate job rather than inflating the image tenfold for a minority of tests. **`libswresample` is a real gap, not a formality.** The current `pkg_check_modules` list (`CMakeLists.txt:200-203`) covers avformat, avcodec, avutil and swscale but **not** swresample — which IR-004 needs to downmix to mono and resample to 11025 Hz. It must be added alongside the audio-signature work. **Gap:** entire requirement. The image does not exist, and no CI config is present in this repo. ## DP-008 — Builder images and release binaries Produce prebuilt binaries per backend so deployment does not require every user to compile the project. **This does not contradict DP-005.** That requirement rejects Docker as a *runtime* — GPU passthrough is the most fragile part of a containerised setup and exists only because of the container. Using Docker as a *build* environment is the opposite case: hermetic, reproducible, and it lets one machine produce binaries for backends it cannot itself run. Build in a container; run natively. ### Image matrix The build has two independent axes (`CMakeLists.txt:48-49`), so the useful combinations are: | Image | `SAE_INFERENCE_BACKEND` | `SAE_GEMM_BACKEND` | Target | |---|---|---|---| | `sae-builder-cpu` | ORT | CPU | CI (DP-007), and the smoke-test fallback | | `sae-builder-cuda` | TRT | CUDA | NVIDIA | | `sae-builder-rocm` | ORT | ROCM | AMD | All three carry the DP-007 dependency set (OpenCV 5, HDF5, FFmpeg incl. swresample, vendored Catch2/nlohmann) and differ only in the accelerator stack. The CPU image is the CI image — one artifact, two uses. Published to the Gitea container registry, pinned by tag, rebuilt when the dependency set changes rather than per run. ### What ships, and what cannot **Ships:** the `scene_analyze` binary and its companions, per backend. **Cannot ship: TensorRT engines.** `.engine` files are specific to the GPU architecture and TRT version they were built on — `scripts/build_trt_engines.sh` must still run on the target machine. A prebuilt binary shortens the install; it does not remove the local engine-build step, and the installer must not imply otherwise. **Cannot ship: models.** ~725 MB in LFS, and orthogonal to the binary. ### The constraint that decides the base image **A binary built in a container runs against the host's glibc.** Build on a newer base than the oldest supported host and it fails at load with `GLIBC_2.xx not found` — the classic and entirely avoidable trap when shipping binaries out of containers. So the base is chosen for the *oldest* glibc to be supported, not for convenience or recency. Accelerator libraries have the same shape of problem: the binary links against a driver-provided runtime, so each image must document the CUDA/ROCm version range its output is compatible with, and the installer must check it rather than discovering a mismatch at first inference. ### Jobs A release job per backend, producing a tagged artifact in the registry. These are **not** the CI gate — the gate runs the CPU image on every push (DP-007); release builds run on tag. Their outputs are what DP-005's installer fetches when the user does not want to compile. **Gap:** entire requirement. No images, no release jobs. ## DP-006 — Gallery maintenance as a background concern - Incremental gallery refresh runs on a timer (`gallery_scan_interval`, default 24 h) independently of analysis, so newly added titles' cast is embedded before their media is analysed. - A gallery/model mismatch must be detected **at startup**, not silently produce garbage similarities. See GR-004. --- # Part C — Integration requirements ## IR-001 … IR-003 — Truth-file output Emit the JRay truth format, `schema_version: 1`, `Verbosity::minimal`. - Sibling file `Movie.jray.json` next to the media (suffix configurable plugin-side). - Per actor: `name`, `imdb_id`, `tmdb_id`, `jellyfin_id` (each `""` if unresolved), and `scenes` windows. - **Each window carries its belief** and its identification route (AR-012) — the posterior is computed for every claim anyway, so it is serialised rather than discarded. This lets a consumer caveat or filter low-confidence presence instead of treating every window as equally certain. - `jellyfin_item_id` is stamped in *after* analysis by `run_from_jellyfin.py` — `scene_analyze` does not know it. - Additional verbosities: `standard` (per-frame bboxes, similarity, unknowns) and `xray` (Jellyfin-Xray `{"second": ["Actor", …]}`, dense integer seconds). ### `extraction.*` provenance Consumers — and the public server, which ranks competing manifests — need to know what produced a result: - `sample_fps`, `pipeline_version`, `gallery_size` — as today. - **`extinction_sec`** — replaces `anneal_sec`, which is dropped entirely (AR-012 withdrawal note). It is the parameter that shapes window extent, so it is what a consumer needs to interpret them. - **`gallery_scope`** — `"global"` or `"limited"`. **The single most useful quality signal**: two galleries of identical size differ enormously depending on whether matching ran against the whole library or only the title's credited cast. A limited gallery cannot find an uncredited or mis-credited appearance at all; a global one competes against every actor in the library. Default is **global**. - Optionally, the `tmdb_id`s of the actors the gallery was built from — useful for reproducibility locally. Not proposed for the Jmanifest, where a thousand-entry id list is bulk for little gain over `gallery_size` + scope. Format is owned by [`../../jRay/SPEC.md`](../../jRay/SPEC.md); this pipeline is the producer. Any change is a coordinated schema-version bump. **Current:** `result_sink_node.hpp`. **Gap:** the schema changes in several ways at once — `anneal_sec` out, `extinction_sec` and `gallery_scope` in, per-window belief added, audio signature added (IR-004). All breaking, so they ship as **one** `schema_version` bump coordinated across all three repos. The `scenes` values also change under AR-012. **Output timing:** the file is written after the deferred pass (AR-020/AR-021) completes, not at EOF — deferred and pooled identifications add windows after the last frame is read. ## IR-004, IR-005, IR-007, IR-008 — Audio signature Emit the content-derived audio signature in the truth file, so a truth file is self-identifying without a plugin round-trip. Construction is specified in [`../../JRay-public-server/SPEC.md` §3](../../JRay-public-server/SPEC.md) and must be implemented **exactly** — a signature that differs in any parameter will not match one computed by the plugin: 1. Decode a 120 s window centred on the midpoint (`runtime/2 ± 60 s`) — avoids logos/cold opens at the head and credits at the tail. 2. Downmix to mono, resample to 11025 Hz. 3. STFT: 4096-sample frame, 1024-sample hop (~93 ms, ~1290 frames), Hann window. 4. Log-magnitude spectrum over 300–3000 Hz. 5. 32 logarithmically spaced bins; record peak-bin index + 2-bit energy class. 6. One byte per frame → ~1290-byte array, base64-encoded. Peak-bin rather than full spectrum: peaks survive lossy re-encoding, loudness normalisation and channel-layout changes; absolute magnitudes do not. Matching (sliding ±600 frames ≈ ±56 s, scoring the fraction of overlapping frames whose peak bin matches) is a **consumer** concern — this pipeline produces the signature, it does not match. Offsets are applied client-side; manifests are never rewritten. **Media shorter than 120 s.** The window `runtime/2 ± 60 s` underflows, so no signature is emitted and **no sync offset is applied**. Such items fall back to the runtime tier, which is adequate: a 90-second extra or trailer is not the content whose cut alignment matters. (There is no `exact` tier: the file-hash tier was withdrawn on legal grounds — it fingerprinted an individual copy rather than the cut the timings describe. See the server spec §3.) Both producers must apply the identical rule, or they diverge on exactly the short items most likely to be mis-identified. **Signature versioning.** The signature carries its own `v1:` prefix, separate from `schema_version` (server spec §3 example: `"v1:v7fA3k…"`). Emit and honour it, so a future change to the DSP chain is *detectable* rather than silently producing non-matching signatures. **Decision (this spec):** the pipeline computes and emits it *in addition to* the plugin. Consequences to carry through: - The truth schema gains a field → **`schema_version` bump**, coordinated with `jRay/SPEC.md` and the plugin. - The pipeline needs an audio decode path. It already links FFmpeg (`ffmpeg_decoder.hpp`) for video, so this is a second stream from an existing dependency, not a new one. - Both producers must agree bit-for-bit. A cross-check test — plugin signature vs. pipeline signature over the same file — is a hard requirement, not a nicety. - Files never processed by this pipeline still get a signature from the plugin; the two paths coexist deliberately. **Current:** `src/audio_signature.*` implements the construction, and `tests/fixtures/audio/` holds the golden vector shared verbatim with the plugin repo, which now matches it byte for byte from C# (jRay `JR-042`/`JR-043`). `sae_audio` (nanobind, as `sae_embed` and `sae_kpn` are) exposes the same C++ to Python so a study drives the shipped code rather than a numpy port. **VR-014 measures what the golden vector cannot** — that the signature actually aligns a differently trimmed release, on real film audio rather than a synthetic tone. It does, with an order of magnitude to spare. **The accuracy question is settled and is not close.** What the offset is *for* is shifting scene windows, which are seconds long, so half a second of error is invisible; the budget is 500 ms. Over 40 random offsets inside the ±600-frame cap the recovered offset was the nearest frame every time — **worst error 46 ms**. That figure is the quantisation floor rather than a measurement of quality: the offset is expressed in whole 92.88 ms frames, so no correct answer can ever be worse than half a frame. The `runtime/2` anchor behaves as specified through real head-trimmed files (cutting `delta` from the head moves the window by `delta/2`), and both an out-of-cap offset and unrelated content are declined outright (0.10 and 0.07). **Where it is soft is tier labelling, not alignment.** The *score* at the correct offset falls with sub-frame misalignment — 0.94–0.99 when the true offset lands within 0.1 of a frame boundary, 0.69–0.73 at half a frame — because the two windows' frame grids no longer coincide. The offset stays right, but only 13 of 40 cleared the server's 0.85 `audio` threshold and the other 27 were demoted to `loose`, a tier that means "possibly the same cut, degraded audio". The threshold was calibrated on a re-encode at *zero* offset, where the score is 1.00. The remedy is measured, not proposed (UT-108): counting a frame as agreeing if its peak bin matches **within ±1 frame** returns all 40 to `audio` (worst 0.906) while unrelated content and out-of-cap offsets stay at 0.12 and 0.16 — the gap that makes the threshold mean anything is untouched. It costs 81 ms of offset accuracy, of a 500 ms budget, because the flattened peak lets the argmax pick an adjacent frame. ±2 frames buys nothing further. Adopting it is a [server spec](../../JRay-public-server/SPEC.md) §3 change — the score is normative and shared by three repos — so this repo measures it and leaves the decision there. **Gap:** the signature is computed but **not yet emitted** into the truth file — that is the `IR-002` field and the coordinated `schema_version` bump. ## IR-006 — Jellyfin round-trip - Pull the work queue: `GET /Plugins/JRay/Tasks/Pending?limit=N` — items with no results yet. - Push results: `PUT /Plugins/JRay/Items/{itemId}/Truth` (admin API key). Managed truth takes precedence over a sidecar file for the same item. - Push only complete results (DP-004). **Current:** `scripts/run_from_jellyfin.py`. **Gap:** none. --- # Part D — Gallery construction requirements ## GR-001, GR-002, GR-005 — Sources Build a gallery of actor reference embeddings from Jellyfin and TMDB. - **Jellyfin-wide** (`make_jellyfin_gallery.py`): enumerate every Movie/Series, collect the unique cast across the whole library, download each actor's headshot from Jellyfin directly (no TMDB key required), embed, write one global gallery. - **TMDB fallback**: for actors with no usable Jellyfin image, fall back to TMDB profile images (`--tmdb-key`). - **Incremental merge** (`--merge`): re-runs pick up newly added titles without re-embedding actors already present. This is what makes DP-006 cheap enough to run daily. - **Cast restriction** (`filter_gallery.py`): derive a per-title gallery limited to credited cast. Faster and fewer look-alike mismatches, but note [`gallery-scope.md`](gallery-scope.md) — the rep4 matrix found the *full* gallery won for the shipped model, so restriction is a tool, not the default. Each actor carries `name`, `imdb_id`, `tmdb_id`, `jellyfin_id` (whichever resolve), one embedding per reference image, and the source image paths. **Current:** `make_jellyfin_gallery.py`, `make_gallery.py`, `filter_gallery.py`, `sae_jellyfin.py`, `sae_tmdb.py`, `src/gallery/gallery_builder.*`. **Gap:** none. ## GR-003 — Quality and coverage reporting Gallery build must report, not just produce: - actors with zero usable images (they can never be recognised — a silent recall ceiling); - actors below the 5-embedding threshold for positive pairs (AR-023), which degrades calibration; - the fitted calibration and the intra/inter distributions behind it (AR-023 gap); - duplicate/near-duplicate references removed. **Gap:** partial. Dedup and eligibility are computed inside calibration but not surfaced as a build report. ## GR-004 — Model binding - A gallery is only valid for the embedder that built it. The embedder identity must be **stamped into the gallery file**, and checked at startup by any consumer. - Mismatch is a hard startup error. Cosine similarities between embeddings from different models are meaningless but *look* plausible — this fails silently and expensively otherwise. ### The stamp Two fields, written together: the model file's **basename** and the **SHA-256 of its bytes** (plus `embed_dim` as a cheap extra guard). Stored as the `/embedder` group in the gallery HDF5, and as an optional top-level `"embedder"` object in the legacy JSON format. The hash *decides*; the name is what a human *reads*. Neither alone is enough. A name is a promise rather than a fact — models get re-exported, re-quantised and overwritten in place under an unchanged filename, which is exactly the case where the weights differ and nothing else does, so a name-only stamp is blind to the failure it exists to catch. A hash alone is correct but unactionable: *"expected 3f2a…, got 9c1b…"* tells an operator nothing about what to do next. SHA-256 over the file is derived from the artefact rather than asserted about it, needs no registry kept up to date, and costs ~0.1 s for a 250 MB ONNX once per process. ### Verdicts | Verdict | When | Default | Under strict mode | |---|---|---|---| | `match` | hashes agree | proceed | proceed | | `weak_match` | names agree, one side unhashable | **warn** | **error** | | `unstamped` | gallery predates GR-004 | **warn** | **error** | | `unknown_embedder` | gallery stamped, embedder unidentifiable | **warn** | **error** | | `mismatch` | proven different models | **error** | **error** | **A mismatch is fatal in every mode, with no bypass**, and the message names both sides — what the gallery was built with and what is loaded. The three "cannot prove it" verdicts warn loudly instead, because they describe an *unknown* state rather than a *known-bad* one, and because every gallery built before this requirement is unstamped. Hard-failing all of them would make the check something people route around rather than trust. Strict mode (`--require-gallery-stamp`, or `SAE_REQUIRE_GALLERY_STAMP=1`, which propagates to subprocesses) promotes them to errors — that is the mode measurement work runs in. `scripts/stamp_gallery.py` re-binds an existing gallery without re-embedding, so migration costs one command; that is what makes "warn" a temporary state rather than a permanent one. ### Scope of the check Embedding **dumps** carry the same stamp (`embedder_model` / `embedder_sha256` root attributes, `scripts/optimizer/SCHEMA.md`): a replay has no live embedder, so the dump *is* the embedder as far as the gallery is concerned. Derived galleries (filter, cast-restrict) inherit their source's stamp; `--merge` and the JSON gallery merge check *before* writing, since a merged file holding two embedding spaces cannot be untangled afterwards by any later check. **Gap:** none. Stamped in `gallery_builder.cpp` and the Python builders; verified in `scene_analyze`, `scene_preview`, the `sae_kpn` matcher binding, `replay.py`, `optimize.py`, `movienet_eval.py` and the merge paths. ## GR-006 … GR-009 — Provenance tiers and poisoning guard Reference embeddings now come from three sources with different trust, and they must be **distinguishable in the gallery**, not merged into an undifferentiated pile: | Tier | Source | Persists | Trust | |---|---|---|---| | **Baked** | Jellyfin / TMDB headshots | Yes | High — curated, externally sourced | | **Harvested** | Per-film annex (AR-019), promoted from owned tracks | **Yes, flagged** | Unverified — machine-derived | | **Confirmed** | Human association (`../../SPEC.md` §4) | Yes | Highest — a person said so | **Harvested embeddings are retained rather than discarded at exit**, because they are exactly the non-frontal views the baked gallery lacks and their value compounds across a library. But they carry the risk the ephemeral annex avoided: **a promotion error becomes permanent instead of dying with the process.** They are therefore **flagged as harvested and reviewable**, never silently equal to a baked reference. The tier must be recorded per embedding so that a suspected poisoning can be traced, audited, and reverted without rebuilding the gallery. ### Bell-curve outlier detection An actor's own embeddings should form a **roughly normal distribution in cosine space** around their centroid. A harvested embedding that falls outside that distribution is unlikely to be the same person — which is precisely what a poisoned entry looks like. **Requirement: flag harvested embeddings that are distributional outliers among that actor's references**, for review or automatic exclusion. > **`EXCEPTION: AR-024` — raw cosine is used here deliberately.** > > AR-024 requires calibrated probabilities everywhere, and this is an agreed > exception. The reason: the calibration is a monotonic squash mapping similarity > onto `P(same person)`. That is exactly right for making a *decision*, and wrong > for characterising a *distribution* — the sigmoid compresses the tails, which > is where outliers live, and would flatten the very structure being tested. > Distribution shape and outlier distance are properties of the metric space, so > they are measured in it. > > Scope of the exception: distributional analysis of an actor's own reference set > only. Any match, association, or admission decision still goes through the > calibration. **Gap:** entire requirement — tiering, persistence of harvested embeddings, the flag, and the outlier check. **Open question:** whether human-confirmed associations should be a distinct audited tier (individually revocable if someone mislabels) or simply more embeddings for that `tmdb_id`. Deferred. --- # Part E — Parameter-study requirements The tuned constants in `config.hpp` are empirical. Retuning must stay cheap, or it will not happen — and AR-012 makes a retune mandatory. ## VR-001 — Post-inference dump Persist pipeline state at the point where the expensive work ends. - Dump at the `EmbeddedSceneFrame` channel — after decode → detect → align → embed, **before** tracking and identity matching. Everything downstream is cheap CPU maths, so a replay re-runs the whole tail with no GPU and no video. - HDF5, one file per title, flat/ragged: per-face arrays concatenated, with a per-frame index table (`face_offset`, `face_count`) pointing into them. Avoids variable-length HDF5 types and reads straight into numpy. - Stores per frame: `timestamp_sec`, `frame_idx`, `is_cut`, `is_scene_boundary`. Per face: `embedding` [N,512], `bbox` [N,4], `landmarks` [N,10], `confidence`. - Invariants: embeddings unit-norm; `face_offset` contiguous; bboxes and landmarks in **decoded-frame** pixels with `bbox_upscale` recorded alongside (the dump is a faithful tap, so it does not transform what the tracker saw — see VR-010); frames with no faces still get a row so timestamps stay dense; EOF sentinels not written. - Enabled by `--dump-embeddings out.h5`; teeing must not perturb the live result. Schema owned by [`scripts/optimizer/SCHEMA.md`](../scripts/optimizer/SCHEMA.md). **Current:** C++ dump sink (`embedding_dump_node.hpp`, `dump_embeddings.cpp`), read by `replay.py`. **Gap:** **AR-012 breaks the replay contract.** Track extents are decided in the tracker, which is *downstream* of the dump — so a replay can reproduce them, but only if the dump preserves everything the tracker needs. Verify `landmarks`/`bbox`/`is_cut` suffice, and bump `schema_version` if not. ## VR-002 — Replay and sweep - Replay drives the **real KPN nodes** over dumped embeddings, not a reimplementation — a sweep that optimises a divergent copy is worthless. - The gallery loads once per process and is cached by path, so one evaluation is N cheap replays. - Differential Evolution over the continuous knob space (`prob_threshold`, `expand_min_anchor_frames`, the re-acquisition timeout, …), scored against reference presence data. Note these are *not* independent — a longer timeout yields longer tracks, hence more frames to clear the anchor count — so they must be swept jointly. **Current:** `scripts/optimizer/optimize.py`, `replay.py`, `second_score.py`. **Gap:** none, pending E1. ## VR-003 — Scoring methodology - Micro-averaged per-second presence against Amazon X-Ray, per [`methodology.md`](methodology.md). - **Objective is F1, but precision and recall are logged at every evaluation and printed at the optimum.** X-Ray recall is a face-vs-cast-in-scene ceiling, so unconstrained F1 pushes `prob_threshold` *down* chasing unreachable recall, trading real precision away. The trade-off must stay visible so another operating point can be chosen from the trajectory (`--trajectory`). - Known metric hazard: the earlier 9-film scene-union metric hid out-of-cast false positives; the 4-film rep4 per-second metric supersedes it. Any new metric must be checked for the same class of blindness. **Current:** implemented; documented in `methodology.md`, `rep4-optimizer-results.md`, `optimizer-experiments.md`. **Gap:** none. ## VR-004 — Validation corpus - A manifest-driven film set with ground truth (`scripts/validation/`), scored reproducibly. Benchmarking practice is already established — see [`methodology.md`](methodology.md), [`model-bakeoff.md`](model-bakeoff.md), [`best-model.md`](best-model.md) and `rep4-optimizer-results.md`; AR-012 changes what is measured, not how. - After AR-012, `prob_threshold` (0.754) must be re-derived — it was fitted against per-frame presence semantics and now governs *voting*, not presence. `anneal_sec` (35.5) and `extinction_sec` (57.4) are not re-derived; they are deleted (AR-012 withdrawal note). ## VR-005 — Minimum face size study Quantify where ArcFace degrades, replacing the 66×66 estimate in A1 with a measurement. > **Result, and its limit.** Knee at 24–32 px; 32 px returns 98.1% TPI at 0.0 > FPI. But the probe is an already-aligned 112×112 crop, so alignment is held > perfect and this measures the **embedder alone** — an upper bound, not a > threshold. **VR-013** re-asks the question end to end, downscaling the whole > frame before the detector, and lands near 50 px. AR-002's floor of 40 px comes > from VR-013; this study is what shows how much of the gap is detection and > landmark error rather than embedding. **Method.** 1. Select ~100 gallery actors having more than one mugshot. 2. Per actor, hold out **one** image as the probe; its remaining images stay in the gallery at native resolution. 3. For each target size *S*, downscale the probe to *S*×*S* and upscale back to 112×112, then embed. 4. Match each degraded probe against the full 100-actor gallery and record **TPI/FPI** — identified as the correct actor, or as someone else. 5. Repeat across sizes to get the curve. The asymmetry is the point: **the gallery stays high-res and only the probe degrades**, which is exactly the production case — reference mugshots are clean, the face from the video is small. It also measures the decision the pipeline actually makes (probe against gallery) rather than embedding drift, which can be large without harming separability and small in a direction that destroys it. **Caveat on gallery size.** FPI grows with the number of actors competing, so a 100-actor gallery understates the false-positive rate against a full library of thousands. Treat the FPI numbers as *relative* across sizes rather than as an absolute rate, or re-run at production scale before setting a threshold from them. **Cheap to run** — no video needed, gallery images are already on disk, and the embedding/matching machinery exists (`scripts/sae_embed_loader.py`, `gallery_calibration.hpp`). Expect a knee rather than a cliff; the output is a size-versus-TPI/FPI curve plus a chosen operating point, not a single number. **Secondary output:** the same curve shows whether `min_face_px` should be a constant at all or should scale per embedder — relevant since the model is a build-time choice (GR-004). ## VR-007 — Expansion band and deferred-pass study Tune the AR-018 admission band and establish whether AR-020 pays. **Band.** Sweep the lower and upper bounds around the 0.90–0.95 working estimate. The two bounds fail in opposite directions and must be read separately: too low a lower bound admits the wrong person (precision collapse, amplified by the deferred pass); too low an upper bound admits only redundant views (no recall gain, wasted annex). Report both, not a single F1. **Deferred pass.** Measure recall recovered by AR-020 — how many TBI entries are identified on re-assessment — and precision of those late identifications specifically. They should be scrutinised separately from first-pass identifications, because they are the ones relying on harvested rather than baked references. **Clustering (AR-021).** Sweep the merge threshold and report cluster purity against known-cast ground truth, not just downstream F1. The two error modes are asymmetric and must be reported separately: an over-merge mis-identifies every track in the cluster at once, while an under-merge only forfeits the pooling benefit. Verify the temporal cannot-link constraint is actually binding — measure how many candidate merges it rejects, since if the answer is zero the constraint is not doing the work claimed for it. **Iteration.** Test one pass versus iterating to convergence (AR-020). Report whether round 2+ recovers enough to justify the complexity, and whether precision degrades with each round — the failure mode being a wrong identification in round 1 seeding references that corrupt round 2. **Ablation worth having:** expansion on with deferred pass off. It separates "expansion helps live matching" from "expansion helps the second pass", which the current all-or-nothing `expand_gallery` flag cannot distinguish. ## VR-008 — Gallery scaling benchmark Establish the throughput-versus-gallery-size curve required by A10. **Method.** Synthesise galleries at 10², 10³, 10⁴, 10⁵ actors (random unit-norm embeddings suffice — this measures compute, not accuracy) and record per-frame matching time, end-to-end throughput, and GPU memory. **The number that matters** is where gallery scoring stops being negligible and starts dominating runtime relative to decode, detection and embedding. Below that point gallery growth is free; above it, it sets the pace. **Also measure the deferred pass (AR-020) separately.** It has a different shape — one large offline multiply rather than many small streaming ones — so it may scale quite differently and could well become the dominant cost on a film with many unknowns. **Report calibration build time too.** AR-023's cross-actor pair enumeration is O(N²); this benchmark is where that becomes visible, and it will likely be the first thing to break at scale. ## VR-006 — Re-tune `scene_threshold` at native rate Not a study of whether to feed TransNetV2 correctly — AR-011 settles that it must be. This is the consequence: `scene_threshold` (0.60) was picked against 12 fps input, where the model's separation was compressed. At native rate the boundary/ non-boundary margin should widen, so the operating point moves. Small and mechanical: score boundaries across a threshold range on a few titles once native-rate decode lands, and pick the new point. Expect a cleaner separation than the ~0.50 baseline / ~0.7+ peak recorded at 12 fps; if it does *not* improve, that is evidence worth having about the export itself. --- # Open questions Resolved during planning, recorded here so the reasoning is not lost: - ~~**AR-012 vs. extinction.**~~ Wholly subsumed. The `last_seen` model ends windows at the last sighting, which removes the over-claim `extinction_sec` caused; both it and `anneal_sec` are deleted rather than re-fitted (AR-012 withdrawal note). - ~~**AR-022 retention.**~~ Embeddings + metadata by default, crops opt-in behind `--dump-unidentified-crops`. - ~~**IR-004 bit-exactness.**~~ Golden-vector fixture checked into both repos, not a shared implementation — the coupling cost of the latter exceeds the benefit. - ~~**AR-023 prior.**~~ Decide once GR-003 persists the intra/inter distributions, so the real value is known rather than argued about. Still open (pipeline-local): 1. **Re-acquisition timeout scope.** `track_max_frames_missing` (5) and `cut_inactive_max_frames` (5) currently distinguish an ordinary miss from a cross-cut park. Under AR-012 both become the same thing. Do they collapse to one constant, or does a cut still warrant a different window? Cheap to test both in the Phase 2 sweep. 2. **Context-crop budget.** How many representative frames per unidentified track, and chosen how (largest? sharpest? most frontal?). Bounded by review-UI usefulness rather than by diagnostics. Escalated to the system spec ([`../../SPEC.md`](../../SPEC.md) §5), since they bind more than one repo: `schema_version` coordination for the pending bump, and whether unidentified presence should be published in the truth format.