Files
scene-actor-extraction/docs/IMPLEMENTATION-PLAN.md
T
dtourolleandClaude Opus 5 a5299daf6e docs: software spec, requirements register, and implementation plan
Adds the requirements baseline for the pipeline redesign:

- SPEC.md — software requirements with Current/Gap deltas per item, so the
  document doubles as a work list.
- requirements.md — stable flat IDs (AR/DP/IR/GR/VR) with parent traces,
  priorities, statuses, and a per-requirement verification plan. Replaces the
  thematic A1..E8 scheme, which had already produced an A1a and an out-of-order
  E6; IDs are now permanent and never reused.
- IMPLEMENTATION-PLAN.md — phased work.

The central change is AR-012: presence follows track extent rather than
per-frame recognition, so a window starts when an actor appears rather than
when the recogniser first succeeded. anneal_sec and extinction_sec are
withdrawn rather than retuned — a track that survives its own gaps leaves them
nothing to do.

Verification is shaped by CI running on an N100 with no dGPU: the existing
HDF5 dump makes everything downstream of embedding replayable on CPU, which
covers the bulk of the redesign.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 17:32:44 +02:00

33 KiB
Raw Blame History

Implementation plan

Companion to SPEC.md. Sequences the gaps identified there into phases, with dependencies and acceptance criteria.

Ordering principle: land the things that make measurement trustworthy before the things that require measurement. The central change (A6) invalidates every tuned constant in config.hpp, so anything that could corrupt a retune must be fixed first, and the retune itself gates the release.

Phase Contents Gated on Parallelisable
0 D3 model binding, E1 dump audit with 4
1 TrackRegistry + A6 track-extent presence 0
2 Retune (A6 constants), E4 1
3 A8 unidentified capture, D2 gallery report 0 with 1, 2
4 C2 audio signature with everything

Phases 3 and 4 touch disjoint code from 1 and 2 and can proceed alongside them.


Phase 0 — De-risk measurement

Small, self-contained, and prerequisite to trusting any number produced later.

0.1 Gallery/model binding (D3)

Stamp the embedder identity into the gallery at build time; verify at load.

  • gallery_builder writes the embedder model identity (filename + hash of the ONNX, or an explicit version string) into the gallery file.
  • Every consumer — scene_analyze, replay.py, the optimizer — checks it at startup against the embedder it is about to use.
  • Mismatch is a hard error naming both sides, never a warning.

Why first: cross-model cosine similarities are meaningless but look entirely plausible. A retune against a mismatched gallery would produce numbers that are wrong and undetectably so. This is the cheapest insurance in the document.

Acceptance: a deliberately mismatched gallery/model pair fails at startup with a message naming both; a matched pair is unaffected. Covered by a test in tests/test_gallery_store.cpp.

0.2 Dump audit for track-aware replay (E1)

The dump captures state at EmbeddedSceneFrameupstream of tracking. After A6, presence depends on tracker output, so replay must be able to reconstruct tracks exactly as the live pipeline would.

  • Confirm bbox, landmarks, confidence, is_cut, is_scene_boundary and frame timestamps are sufficient to re-run FaceTrackerFunc deterministically.
  • Specifically verify the cross-cut park/revive path (A4) is reproducible: it depends on is_cut and on last-frame embeddings, both of which should be present — confirm rather than assume.
  • If anything is missing, add it and bump schema_version in scripts/optimizer/SCHEMA.md.

Acceptance: for one film, tracks reconstructed from the dump are identical (same track_id partitioning of the same faces) to those from a live run. This equivalence test is the foundation of Phase 2 and should be kept as a regression test.


Phase 1 — TrackRegistry and track-extent presence (A6)

The central change. Presence moves from "frames where the actor was recognised" to "extent of tracks the actor owns".

1.1 Design: TrackRegistry as a shared resource

Decision: track ownership lives in a TrackRegistry object that is external to the dataflow network, not in a node and not inside TrackGallery.

This is idiomatic for this codebase rather than a workaround:

  • Node functors are already constructed outside the network and passed in by reference — main.cpp:186-207 builds ftracker_fn, tracker_fn, sink_fn as stack objects and ObjectNode wraps them ("the object must outlive the node", KPN SPEC §371).
  • KPN provides SharedResource<T> (external/KPN/shared_resource.hpp) precisely for state shared across nodes, and shared resources can be registered with a network for reporting (KPN SPEC §163, §445).

Why not the two alternatives:

  • Not a node. Ownership is not a stage in the stream — it is state that several stages read and write, at different points, with the final answer only known at EOF. Modelling it as a node would force ownership to be decided at a single point in the flow, which is exactly what it cannot be.
  • Not inside TrackGallery. Ownership is already computed there for expansion, so putting presence there too would avoid duplication — but it couples presence semantics to expand_gallery, a switchable feature. Turning expansion off would silently revert A6.

Ownership and lifetime: the registry is created in main and handed to every node that needs it as a std::shared_ptr<TrackRegistry>. Nodes hold their own shared_ptr, so lifetime is guaranteed by refcount rather than by the "object must outlive the node" convention — no ordering assumption between network teardown and the registry's destruction.

Per-track state:

Track
  first_seen : double            set once, at track creation
  last_seen  : optional<double>  UNSET while on screen; set to the last
                                 on-screen timestamp when the face is lost
  actor      : optional<int>     set when a posterior crosses the threshold
  belief     : {actor_idx -> accumulated_logodds}   Bayesian, not a tally (A9)
  embedding  : Embedding         running directional mean, for association

last_seen carries the entire liveness state. Unset means on screen now; set means went off screen at T. There is no separate "missing frames" counter and no expired flag in the registry — the optional is the state machine.

Lifecycle:

face detected, no match to an existing track
    → new track, first_seen = t, last_seen = unset
    → begin embedding

actor identified
    → record vote; set actor once the ownership rule fires

face lost
    → last_seen = t_last_on_screen        (track stays alive, revivable)

face seen again, embedding matches a track with last_seen set
    → last_seen = unset                   (same track continues)

tick(t) where last_seen is set and t - last_seen > timeout
    → track is DEAD: emit it to the result aggregator, delete the entry

A dead track is passed to the result aggregator. Reaping is a handoff, not a deletion into a holding pen: the reaped track — first_seen, last_seen, owning actor, vote tally — is pushed downstream the moment it dies, and the registry drops it. The registry therefore contains only live tracks, and its size is bounded by concurrent on-screen faces rather than growing with the film.

That gives a clean division: the registry answers "who is on screen now and which tracks are still revivable"; the aggregator accumulates finished presence. Neither needs the other's state.

A face appearing after the timeout starts a genuinely new track with a new track_id — correct, since past the re-acquisition window there are no grounds to assert continuity. An actor who leaves for half an hour and returns gets two windows rather than one spanning their absence.

The registry thus has exactly three states, all implied by last_seen: on screen (unset), revivable (set, within timeout), and gone (emitted and removed). There is no "expired but retained" state to reason about.

Closing a track is the presence assertion. This is the property the design turns on. There is no later stage where presence gets decided, reconciled or merged — the act of closing a track emits exactly one complete, immutable claim: this actor was on screen from a to b. It is produced once, never revised, and carries everything needed to justify it (the vote tally that named the actor).

Several things follow, which is why it is worth naming explicitly:

  • Presence is append-only. The output is a stream of finished claims, so a partial result is a prefix of the full result, not a corrupted version of it.
  • Each claim is independently checkable. A wrong window can be traced to one track and its votes, rather than to an emergent interaction between three timeout constants.
  • Nothing downstream needs to be clever. The aggregator groups claims by actor and writes them out. It holds no state machine of its own, which is precisely why SceneTrackerFunc disappears.

Compare the current design, where presence is inferred at the end from a pile of per-frame detections via two gap-bridging constants, and no single moment is "the decision". That indirection is the source of both the late-start bug and the untunability.

An actor's presence window is simply [first_seen, last_seen] of each track they own. Nothing else.

This removes annealing entirely. anneal_sec (35.5 s) and extinction_sec (57.4 s) exist only because presence is currently assembled from isolated accepted frames, which are full of holes — both constants are gap-bridging patches over that. Under this model a track survives its own gaps by construction: a face lost and re-acquired by embedding match is the same track, so there is no second window to merge and nothing to anneal. Presence continuity is inherited from track continuity rather than reconstructed from it.

This is also what makes camera cuts fall out for free. A shot/reverse-shot sets last_seen on the cut and unsets it on the next matching detection; the extent never breaks. The existing cross-cut park/revive machinery (face_tracker_node.hpp:154-172, cut_revive_sim) is exactly this mechanism already, applied only to cuts — the model generalises it to every disappearance, which is why cuts stop needing a special case.

One timeout, not three. The only surviving tunable is how long a track stays revivable after the face is lost. It does two jobs at once, and both are wanted:

  • Gap absorption. A face lost at t₁ and re-acquired at t₂ within the timeout never closed its track, so the window runs straight through — the actor is claimed present across [t₁, t₂]. That is correct: someone who turns away, is briefly occluded, or is off-camera while the shot cuts to whoever they are talking to has not left the scene.
  • Identity continuity. The re-acquisition is only accepted on an embedding match, so the gap is bridged on evidence that it is the same person.

That second point is the substantive improvement over anneal_sec, which merged windows purely on elapsed time and could therefore stitch together two different people. Same smoothing, now evidence-gated.

The asymmetry that keeps this honest: interior gaps are claimed, the trailing cool-down is not. A track that dies ends its window at last_seen — the last frame the face was actually seen — not at the moment of death. So the timeout buys gap-smoothing without over-claiming the tail.

The documented credits overshoot is fixed by this asymmetry, not by scene detection. extinction_sec kept an actor active for 57 s after their last detection, actively emitting presence into the closing credits (the Downton Abbey recall collapse, lvface-deep-dive.md). A window ending at last_seen never enters the credits at all, because nothing was seen there.

Cuts and scene boundaries are association hints, not presence events. Both signals say the same thing to the tracker: spatial continuity is broken — stop trusting IoU, associate on embedding similarity instead. Neither closes a window.

Signal Meaning Effect
is_cut (histogram) Camera-angle change within a scene Weight association toward embedding
is_scene_boundary (TransNetV2) Different scene Same, more strongly

This is what the tracker's park/revive path already does for cuts (face_tracker_node.hpp:154-172): a post-cut detection is matched on raw cosine similarity to a parked track's last-frame embedding, with IoU out of the picture. Generalising it makes track_alpha (the spatial/embedding cost weight) frame-dependent rather than constant — normal frames use the tuned blend, flagged frames drop toward embedding-only.

An actor genuinely continuing across a boundary is therefore kept, which is correct; one who does not reappear simply times out and closes at last_seen.

This makes --scene-detect load-bearing for presence correctness, not the opt-in extra it is today (config.hpp, default off). Two consequences to decide in Phase 2: whether it becomes default-on despite its dense-decode cost, and what the degraded behaviour is when it is off — presumably the timeout alone, which is the current situation and carries the known overshoot.

It follows that the timeout is a presence knob and cannot simply be made generous: it is precisely "how long an absence do we tolerate before calling it a departure". Too short fragments one continuous appearance into several windows; too long absorbs a genuine exit-and-return into a single window claiming presence the actor did not have. It replaces both anneal_sec and extinction_sec and inherits their tuning burden — this is the constant Phase 2 must fit.

track_max_frames_missing and cut_inactive_max_frames already serve this role in the tracker and should be reconciled into it. Note the units differ: the tracker counts frames missing while windows are in seconds; whichever side owns the timeout should own the conversion, so sample_fps changes cannot desynchronise them.

Ownership rule (per SPEC A9): each track carries accumulated log-odds per candidate actor, updated as frames arrive. Ownership is "posterior exceeds threshold", not "≥ N accepted frames" — so votes in the state sketch above is a {actor_idx → accumulated_logodds} map, and on_vote() is an update, not an increment.

Consequences for the registry:

  • Every similarity reaching it is already a probability (A9); the registry never sees a raw cosine.
  • Correlated-frame discounting (A9) applies at update time. The registry should take the already-discounted evidence rather than deciding the discount itself — that judgement belongs with the diversity buffer that identifies novel poses.
  • Ownership is established at first crossing, not deferred to track death.

Two contradiction rules the registry must enforce (SPEC A6). Both exist because the same underlying fault — a missed camera or scene change — shows up in identity space, and both are detectable online:

Condition Meaning Action
Belief on one track swaps A → B track_id carried across a viewpoint change onto a different person Close the track at last_seen, open a new one for B at the swap frame
Two live tracks owned by the same actor One person split into two tracks, or an identity attached to the wrong one Treat as a detected cut: reset the affected tracking state and re-associate on embedding

The second is the more interesting: it makes identity a third cut detector, independent of the histogram and TransNetV2, firing exactly where those failed. A cut subtle enough to evade pixel-based detection is not necessarily subtle in identity space.

Both must be counted and reported (SPEC D2) — the rates are a direct measure of how often tracking is silently wrong, which nothing currently reveals.

Implementation note: "two live tracks owned by the same actor" is a cheap check because the registry already holds every live track and its belief. Maintain a reverse index actor_idx → live track_ids and the condition is detected on the update that causes it, not by scanning.

The registry is the tracker's state. FaceTrackerFunc does not keep its own tracks_/inactive_ maps and mirror them into a registry — it is constructed with std::shared_ptr<TrackRegistry> and operates on it directly. One copy of the track set, one owner of liveness.

This matters beyond tidiness. Had the tracker kept private maps and reported into a parallel registry, the two could disagree — the tracker expiring a track the registry still thinks alive, or vice versa — and every such divergence would surface as wrong presence windows, silently. Merging them makes that class of bug unrepresentable rather than merely tested-against.

It also subsumes the tracker's existing two-pool split: tracks_ becomes the tracks with last_seen unset, inactive_ becomes those with it set. Same structure, one map, and the cross-cut park/revive path becomes the general re-acquisition path rather than a special case.

Interface:

TrackRegistry
  // tracker-facing: state it owns and mutates
  tick(timestamp)                                 ← FaceTrackerFunc, every frame
  candidates() -> span<Track&>                    → all live tracks; last_seen tells
                                                    the caller whether IoU applies
  create(timestamp, embedding) -> track_id
  mark_seen(track_id, timestamp, embedding)       → updates mean, clears last_seen
  mark_lost(track_id, last_on_screen_timestamp)

  // matcher-facing
  on_vote(track_id, actor_idx, posterior)         ← IdentityMatcherFunc (accepted frames only)

  // reader-facing
  owner(track_id) -> optional<actor_idx>          → TrackGallery (expansion confirmation)

  // output: dead tracks are pushed out as they are reaped
  on_track_dead : callback(DeadTrack)             → ResultSinkFunc
  flush()                                         ← at EOF: emit all live tracks, then clear

DeadTrack carries first_seen, last_seen, the owning actor_idx (or none), and the vote tally — everything the aggregator needs, with no back-reference into registry state.

Every track must be closed at EOF. flush() emits every still-live track through the same callback, closing each at last_seen if set and at the final tick timestamp otherwise. This is not a tidy-up detail: a film almost always ends with faces on screen, and those tracks have not timed out, so without an explicit flush they are simply never emitted — the closing scene's actors disappear from the output. That failure is silent and looks like a recognition miss rather than a bookkeeping bug.

Requirements:

  • flush() is idempotent and leaves the registry empty; calling it twice emits nothing the second time. The sink's existing written_.exchange(true) guard (result_sink_node.hpp:66) shows the shape.
  • It must run on every termination path that produces output, not just clean EOF — the eof sentinel, and the early-exit paths (--end-sec, decode error, user interrupt) if those still write results.
  • Finalisation goes through the same code path as a natural death, so a track closed by EOF is indistinguishable in form from one closed by timeout.
  • Deliberately not covered: SIGTERM during opportunistic runs (SPEC B4). Those push no partial result at all, so there is nothing to flush — the item stays pending and restarts. Flush is for runs that produce output.

Acceptance: a clip ending mid-shot yields a window for the on-screen actor whose end equals the final frame timestamp. This is a specific test, not an incidental one.

The tracker's association step reads candidates()one pool, not two. A track with last_seen unset was seen last frame, so IoU is meaningful; one with last_seen set is dormant and matched on embedding alone. There is no separate revival path: matching a dormant track is ordinary inter-frame association, and the property falls out of the embedding comparison rather than being a mechanism. The step then calls create/mark_seen/mark_lost. Reaping happens in tick().

Since the tracker mutates registry state across a frame's association pass, that pass needs to be atomic as a unit — a frame_scope() handle holding the lock for the duration is cleaner than making each accessor independently locked and hoping the composite is safe. This is the one place where per-call atomicity is not sufficient.

tick() advances the clock so dead tracks are reaped independently of detection activity — without it the registry only learns about time when something is detected, and tracks would only die when some other face happened to appear. It is called once per sampled frame whether or not that frame had detections.

Locking. Nodes run concurrently and do not coordinate, so the registry is responsible for its own consistency. Two granularities apply: the tracker's per-frame association pass holds the lock for its whole duration (frame_scope() above), while every other caller's operations must be individually atomic.

Two cases constrain the API shape:

  • owner() is a read-modify-read in disguise: TrackGallery calls it to decide promotion while IdentityMatcher may be concurrently voting on the same track. Tally and verdict must be read under one lock as a snapshot — not "read tally, release, decide" — or a track can be both unowned and owned within a single promotion decision.
  • on_vote() for a frame arrives from IdentityMatcher, downstream of the tracker's tick()/mark_seen() for that same frame. A vote may therefore land after the clock has moved on. Rule: a vote for a known track always lands on that track's tally, regardless of clock position. Only reaping is clock-driven. A vote for a track already reaped is dropped and counted — a nonzero count means the timeout is shorter than the matcher's lag, which is a real misconfiguration and should not fail silently.
  • The on_track_dead callback fires from inside tick(), which the tracker calls while holding the frame lock. The callback must therefore not re-enter the registry, or it self-deadlocks. Keep it to a push onto the aggregator's own storage; anything heavier belongs downstream of that.

A single std::mutex over the whole registry is the right starting point: contention is a handful of small updates per frame against per-frame work measured in GPU milliseconds. Anything finer needs a profile, not an assumption.

1.2 Wiring

  • FaceTrackerFunc is constructed with the shared_ptr<TrackRegistry> and uses it as its state — its own tracks_/inactive_ maps go away. Per frame: tick(), then association over candidates(), then create/mark_seen/mark_lost. Its tracks_/inactive_ split and the cross-cut revival branch both collapse into one pool keyed on last_seen.
  • IdentityMatcherFunc reports accepted-frame votes. It keeps its existing per-frame acceptance logic unchanged — A6 changes what is done with acceptances, not how they are decided.
  • TrackGallery replaces its internal confirmation counter with registry.owner(), so ownership is computed once.
  • ResultSinkFunc becomes the result aggregator: it receives dead tracks via the callback and groups them by actor. The per-frame timestamp collection and gap-merging at result_sink_node.hpp:123-147 is deleted. No annealing pass — a dead track already is a window.
  • SceneTrackerFunc (scene_tracker_node.hpp) is the extinction-timer state machine keyed on actor_idx. Under this model it has nothing left to do: its entire job was keeping actors alive across detection gaps. Expect to delete it from the network rather than adapt it.

1.3 Comparing against current behaviour

The old path is not worth preserving behind a flag. It is not a variant of the new one — it is a different pipeline shape (SceneTrackerFunc present, annealing in the sink, two extra constants), so keeping both runnable means maintaining two sink implementations and a node that otherwise gets deleted.

Compare against recorded output instead: keep the current binary's results for the validation corpus as reference JSON, and diff the new pipeline against them. That gives the same A/B for Phase 2 without carrying dead code through it.

Acceptance:

  • tests/test_face_tracker.cpp extended: a lost-then-re-acquired face continues the same track and yields one unbroken window; a cut does the same; a face re-appearing after the timeout yields two separate tracks and two windows; a two-actor conflict resolves to the majority and increments the conflict counter; a single-frame track yields a zero-length window.
  • A registry test: reaping emits exactly once per track; flush() at EOF emits every live track and nothing twice; a vote landing on a reaped track is dropped and counted.
  • Contradiction tests: a belief swap A→B closes one window at last_seen and opens a second starting at the swap frame, with no overlap and no blended window; two live tracks converging on one actor trigger a re-association and increment the counter.
  • A concurrency test hammering tick/mark_seen/mark_lost/on_vote from multiple threads against owner(), under TSan. The vote-tally-plus-verdict read is the case to target — correct only if atomic as a unit.
  • On a known film, every actor's first window starts no later than in the recorded reference output, and strictly earlier for at least one — the late-start bug this change exists to fix.

Phase 2 — Delete two constants, tune the rest

anneal_sec (35.5) and extinction_sec (57.4) are replaced, not discredited. They exist to answer X-Ray's scene-level question — "is this actor in this scene" — by holding windows open across cuts, and they answer it with elapsed time because that was the only signal available at the sink. Phase 1 answers the same question with better evidence: an embedding-matched re-acquisition, plus a true scene boundary to stop at. So they are deleted along with SceneTrackerFunc, and their job transfers to the re-acquisition timeout rather than disappearing.

What actually needs tuning:

  • The ownership posterior threshold — replaces both prob_threshold (0.754) as a presence decision and expand_min_anchor_frames (3) as a vote count. A track is owned when its accumulated posterior for an actor crosses this (A9). A single false accept can no longer create a window on its own, so the operating point should sit lower than the old per-frame threshold.

  • The correlated-frame discount (A9) — whatever form it takes, it is a fitted quantity and belongs in the sweep. It directly controls how fast belief accumulates along a track, so it trades against the ownership threshold and cannot be tuned separately.

  • The re-acquisition timeout — reconciled from track_max_frames_missing (5) and cut_inactive_max_frames (5). This is the successor to both anneal_sec and extinction_sec and carries their tuning burden. It decides how long an absence is absorbed into a presence window versus treated as a departure, so it trades recall (bridging real gaps) against precision (claiming presence during a genuine exit) directly. Current values are 5 frames, inherited from a tracker-continuity role; as a presence constant it is likely to want a much larger value, and should be swept over seconds rather than nudged.

  • scene_detect — now load-bearing for presence (§1.1), so the sweep must cover with and without, and scene_threshold (0.60) becomes a presence constant rather than a diagnostic one. A missed boundary reintroduces the overshoot; a spurious one truncates a scene's cast early.

Method per SPEC E3: DE over the validation corpus, objective micro-F1, but precision and recall logged at every evaluation and printed at the optimum. X-Ray recall is a face-vs-cast-in-scene ceiling, so unconstrained F1 pushes thresholds down chasing unreachable recall and trades away real precision. Pick the operating point deliberately from the trajectory.

This resolves SPEC Open Question 1 by construction rather than by measurement. The question was whether extinction_sec survives alongside track extents; the answer is that the mechanism it patched no longer exists.

One thing to watch: the search space is now 3 knobs instead of 3 gap-constants, but they are not independent — a longer re-acquisition timeout means longer tracks, which means more frames to clear expand_min_anchor_frames. Sweep jointly.

Acceptance: anneal_sec and extinction_sec removed from config.hpp and Config; new constants committed with the trajectory and the precision/recall trade-off documented in the manner of rep4-optimizer-results.md; results compared against the recorded reference output from 1.3.


Phase 3 — Diagnostics

Independent of Phases 12; can run in parallel.

3.1 Unidentified-track capture (A8)

Today only promoted mugshots are dumped (expand_debug_dir) — the successes. This captures the failures.

  • Flag --dump-unidentified <dir>.
  • For every track never identified: all embeddings, track metadata (track_id, first/last timestamp, frame count, per-frame bbox and confidence), and the best similarity achieved and which actor it was against.
  • That last field is the point of the feature: it separates "actor missing from the gallery" from "actor present but scored below threshold" — a gallery coverage problem versus a threshold problem.
  • Crops are opt-in (--dump-unidentified-crops); embeddings + metadata are the default. Crops for every unidentified track across a library is a lot of disk.
  • When crops are enabled, store both the 112×112 aligned crop and a wider context crop for a bounded number of representative frames per track. The aligned crop serves diagnostics; the context crop serves the human-in-the-loop association capability (../../SPEC.md §4), where someone has to actually recognise the person — which a tightly-cropped, geometrically normalised face often makes impossible.
  • Naturally expressed against TrackRegistry: unidentified = tracks with no owner at EOF.

Acceptance: on a film with a known out-of-gallery face, that track appears in the dump with its near-miss actor and similarity.

Surface what the calibration already computes internally but discards:

  • actors with zero usable images (a silent recall ceiling);
  • actors below the 5-embedding threshold for positive pairs;
  • near-duplicate references removed;
  • the fitted calibration and the intra/inter similarity distributions behind it (gallery_calibration.hpp histograms these at kHistBins = 200 and throws them away — persist them).

Acceptance: a build report written alongside the gallery; the intra/inter PDFs are recoverable for inspection.

3.3 Resolve the prior (SPEC Open Question 4)

SPEC A9 asks for a prior of intra/(intra+inter); the shipped default is match_prior = 0.5 (use the calibrated sigmoid directly). These disagree. With 3.2 landed the real value is known, so: either adopt it, or document 0.5 as a deliberate override with the reason. Cheap once the distributions are persisted.


Phase 4 — Audio signature (C2)

Largest self-contained chunk; no dependency on any other phase.

  • Implement JRay-public-server/SPEC.md §3 exactly: 120 s centred on the midpoint, mono 11025 Hz, 4096/1024 Hann STFT, 3003000 Hz, 32 log bins, peak bin + 2-bit energy class, one byte per frame, base64.
  • Audio decode is a second stream from the FFmpeg dependency already linked for video (ffmpeg_decoder.hpp) — not a new dependency.
  • Emit in the truth file → schema_version bump, coordinated with jRay/SPEC.md and the plugin. Neither can be changed unilaterally.
  • Golden-vector cross-check is a hard requirement, not a nicety: two independent implementations of the same DSP chain will drift. Fixture: a short audio file with its expected signature, checked into both repos and asserted in both test suites (SPEC Open Question 3, resolved this way — a shared fixture rather than a shared implementation, since the coupling cost of the latter exceeds the benefit).
  • This pipeline produces only. Matching and offset recovery stay consumer-side.

Acceptance: pipeline and plugin produce identical signatures for the same file; the golden-vector test passes in both repos.


Deferred

Deployment (SPEC Part B) is deliberately excluded from this plan. B3 (on-demand service) and B4 (opportunistic worker) are packaging concerns over a stable core, and the core is about to change under A6. service-conversion.md remains the design of record; it should be executed once Phase 2 fixes the constants, so the installer is not shipping values that are about to be replaced.

One exception worth pulling forward if convenient: the temp-file cleanup fix in run_from_jellyfin.py (SPEC B4) is small, independent, and a named prerequisite for the worker.


Risks

Risk Mitigation
A6 widens presence and precision drops more than recall gains Phase 2 leaves the operating point explicit; both modes runnable via flag for direct comparison
Retune produces a worse optimum than the current constants Trajectory is logged; the old operating point stays available. A6 is behaviourally correct even if the metric disagrees — decide deliberately, do not let the metric silently veto it
Track-ID collisions merge two people into one extent Conflict counter (1.1) makes the rate visible; expand_track_spread_max already guards the expansion side
Registry races produce non-deterministic presence between identical runs Every operation atomic as a unit (1.1); TSan test in Phase 1 acceptance. A race here is especially costly — it would surface as irreproducible optimizer scores in Phase 2, where it would look like metric noise rather than a bug
X-Ray metric blindness methodology.md documents the known failure (scene-union hid out-of-cast FPs). Any metric change gets checked for the same class of blindness
Two audio implementations drift Golden vectors in both repos (Phase 4)