0feafec7c9a7e2a6aa9c54d6590cc50c037eaf3a
22
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
bb7a9ed718 |
docs(AR-004): record the three holes the wedging audit closed
The register said "two remaining holes now closed" and the SPEC's Current:
still described `push_blocking`, which parking replaced. Both now match the
code.
Three additions, in the order they surfaced:
(c) FilterNode and RouterNode were the last data paths still using the
throwing push() with the exception swallowed, so a full output discarded
the value — including the EOF sentinel. The decimator passes EOF by
predicate but its output is reliably full, the embedder being the slowest
node, so the token went nowhere and nothing downstream shut down. That is
the wedge the runs were being killed for, and it is worth the register
saying so plainly.
(d) The sentinel could be delivered ahead of a value still queued behind it,
losing the tail to any consumer treating EOF as a hard stop.
(e) Two firings of one node could overlap, which breaks the one-slot park
itself: a parked value can be overwritten with no drop recorded.
(e) also corrects (b). The startup lost wake was recorded as a missed
empty->non-empty edge closed by a level-triggered re-check; the actual cause
was the callbacks being written while a running neighbour read them — ten
ThreadSanitizer races — and they are now installed in a prepare() pass before
any node starts. The re-check stays and is still needed, but for a benign
ordering rather than as cover for a race.
Two things the requirement now carries that it did not before. A channel holds
at most one undelivered sentinel: a second is refused and reported rather than
silently overwriting the first, which matters the moment a pipeline is reused
for a second input. And a lossless decimator is a backpressure point rather
than a relief valve, so the source throttles to the face branch instead of
quietly thinning it — what the requirement asks for, but it changes the shape
of a loaded run and is not yet benchmarked.
The Gap is unchanged and still open: capacity is counted in items, not bytes,
so a crowd frame carrying 60 crops occupies one slot exactly as an empty one
does. AR-004 stays **Mostly** for that reason.
Verification plan updated with the cases the KPN suite now pins.
TRACES: AR-004 | SR-002
|
||
|
|
718dad688d |
docs(register): record the quality vector, the benchmark, and what lossless fanout costs
Status for the two changes just landed, plus the consequence AR-004's fix has for the scene join. The annotator's old comment said blocking there was safe because the branches are independent. That stopped being true when the fanout became lossless: it now stops popping once one branch stops taking, so a starved detector and a waiting annotator would wedge. What actually makes it safe is join depth -- the fanout can run the dense branch ahead by the whole of the sampled branch's buffering, which at kSceneJoinDepth 256 and sample_fps 5 against a 25 fps source is ~1200 dense frames against TransNetV2's 100-frame window. Cutting kSceneJoinDepth below the window would reintroduce the wedge, so it is now a correctness precondition rather than a tuning knob. TRACES: AR-004, AR-010, AR-028, AR-029 | VR-015 | SR-002 |
||
|
|
c1155cb607 |
feat(gemm): the annex is a matrix, not a list — scored by the same GEMM
The per-film annex was folded in after the gallery multiply by a host-side
cosine loop over a vector of {embedding, actor} structs, justified in-comment
by "tens of embeddings". AR-018/AR-019 retired that assumption: every owned
track promotes, so the annex grows with cast size and film length.
TrackGallery now holds it as a contiguous row-major matrix with a parallel
actor index — the flat_emb_/flat_actor_ shape the baked gallery already uses —
and hands newly promoted rows to the matcher once per frame. The matcher pushes
them into the similarity engine's resident matrix through a new
ISimilarityEngine::append_rows, so one SGEMM covers baked and promoted
references alike and best-of-N is a single pass over one similarity column.
Capacity doubles on overflow, and the GPU backends grow device-to-device, so a
promotion never re-uploads the gallery across the bus.
Absorbing promotions runs once per frame, after every face has been scored.
Appending mid-frame would invalidate the similarity pointer the chunk loop is
still reading, and it also removes an incidental dependence on face order
within a frame — a promotion helps subsequent frames, never the one that
produced it, which is the semantics the expansion store already documented.
OpenBLAS becomes a requirement of the CPU GEMM backend rather than an
opportunistic upgrade. That path is what CI and the cpu builder image run, so
falling back to the scalar loop in silence meant AR-027 could be measured — or
believed — on a kernel no release uses. The loop survives as the correctness
oracle the BLAS backends are diffed against, behind SAE_ALLOW_SCALAR_GEMM.
Call site 3, the deferred TBI pass, is untouched: it does not exist until
AR-020, so AR-026 stays In Progress.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
TRACES: AR-026 | UT-004, UT-005 | SR-001
|
||
|
|
f33403fff8 |
feat(scene): feed TransNetV2 at native rate, derive the dedup window from it
Closes both violations SPEC.md named under "Every model gets the input it was trained for". They are one bug, not two. The dense stream defaulted to 12 fps, so a 100-frame TransNetV2 window spanned ~8.3 s against the ~4 s it was trained on: half-speed motion over twice its temporal context. Boundary timestamps stayed correct throughout, which is exactly why the degradation was invisible and why the compressed separation it produced (~0.50 baseline against ~0.7+ peaks) was read as a property of the ONNX export rather than of the input. Dedup then merged boundaries closer than a literal 0.04 s — one frame at 25 fps, and wider than a frame at 30, so two cuts on consecutive frames became one. Nothing in scenes.json showed it; the file simply had fewer boundaries. Native rate is where that constant did the most damage, which is why fixing the decode rate without fixing the dedup would have made things worse. dedup_window_sec() now takes the median interval the detector was actually fed and halves it. Half a frame rather than a whole one: the only thing being merged is one frame scored by two overlapping windows, and two distinct frames are a full interval apart. Cost is real — dense decode is the pipeline's cost driver. It is accepted; dense_scale and scene_stride remain the reductions that do not run the model off-distribution. scene_threshold 0.60 was fitted against the 12 fps input and is now stale, so VR-006 goes from Low to Medium: it is no longer a refinement, it is a constant that no longer describes the input. AR-002 rides along because it was already implemented, just untagged and unverified — the register said Planned while the code was correct. The size filter becomes FaceDetectorFunc::drop_undersized(), tested at the threshold and at dense_scale 0.5, and checked end to end against the superhero dump, whose smallest face is exactly its recorded 32 px minimum, so the fixture check cannot pass vacuously. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> TRACES: AR-002, AR-011 | SR-002 | UT-002, UT-003, IT-001 |
||
|
|
eff696b49a |
fix(expansion): finish AR-018, retiring the last two expansion cosines
AR-018 was marked Done while the promotion path still ran on the
constants it was meant to replace. track_gallery.hpp rejected a track
when buffer_spread (1 minus the minimum pairwise cosine) exceeded
expand_track_spread_max, and skipped a view when its raw gal_sim cleared
expand_novelty_sim. Both were bare cosines with no recorded EXCEPTION,
so both were defects under the AR-024 invariant rather than tagging gaps.
The calibrated band was real but unreachable. expand_band_lo/hi were
declared in Config and read nowhere, and set_band() had no callers, so
the gate always ran at the hardcoded 0.90/0.95 while --expand-novelty-sim
and --expand-spread-max stayed live flags.
The spread gate becomes store_coherence: the band's lower bound asked of
every pair in the store, in probability space, rather than a second
constant. admit() compares a newcomer only against its nearest existing
member, so a gradually drifting track chains A to B to C with every step
inside the band while A and C are strangers — the shape a track-ID
collision takes over a slow pan. The bound is re-asked pairwise before
anything reaches an actor's annex.
The novelty gate is deleted rather than converted. SPEC section AR-018
contrasts the band with expand_novelty_sim as the thing it replaces, and
AR-019 requires only that the band is satisfied. Novelty-seeking now
lives entirely in the eviction ordering, which ranks by similarity to the
actor's references instead of cutting at a constant, so there is nothing
left to tune but the two bounds.
BufEntry stored a raw cosine and the eviction loop compared two of them.
The map is monotonic so the ranking was never wrong, but it left a bare
cosine as a decision variable; it now stores the calibrated probability.
The [AR-018] Catch2 tag previously sat on the spread gate, reporting the
replaced mechanism as verification of its replacement. It now sits on the
band: both bounds asserted exactly, since they are inclusive and an
off-by-one there is invisible anywhere else; refusal counted on each
side; and the config bounds driven away from the shipped defaults so a
hardcoded fallback fails. The case that carries the invariant is "band
thresholds probability, not cosine" — under a calibration shifted by
0.10, cosine 0.84 is admitted and cosine 0.92 refused, the opposite of
their raw verdicts. A raw-cosine gate passes an identity-calibrated test
by accident and cannot pass that one. 15 cases, 38 assertions, passing.
scene_preview.cpp takes the flag rename because it would otherwise
reference deleted Config fields. It still does not compile, for reasons
predating this change: it also reads track_max_embed_dist and
track_max_frames_missing, retired by the earlier AR-024 tracker work, and
constructs FaceTrackerFunc with one argument where the registry and
calibration are now required.
Two notes for anyone reading the chain. The main.cpp flag rename and the
AR-018/AR-024 register rows landed in
|
||
|
|
629d698ad9 | Merge branch 'feature/dump-provenance' into feature/opencv5 | ||
|
|
71354e862a |
docs: VR-013 and VR-014 results; AR-002 raised to 40px
Records study results and the requirement change that follows from them. VR-013 measures minimum face size end to end — gallery from one recording, probes from another — rather than by degrading an already-aligned crop. Holding 90% of the plateau needs ~50 px that way against VR-005's ~22 px, the gap being detection and landmark error rather than the embedder. AR-002 therefore takes 40 px, not 32: VR-005 isolates the embedder and is an upper bound, and 32 admits faces in the falling region. FPI stayed 0.0% at every scale, and the ceiling is cross-view rather than resolution. VR-014 exercises audio-signature offset recovery on real film audio instead of the synthetic golden tone. Forty random in-cap offsets, every one recovered to the nearest frame, worst error 46 ms against a 500 ms budget — and 46 ms is the quantisation floor rather than a result, since offsets land on whole 92.88 ms frames. The runtime/2 anchor is confirmed through head-trimmed files. The soft spot VR-014 found is tier labelling, not accuracy: the score drops with sub-frame misalignment, so 27 of 40 correct alignments were demoted to `loose`. One frame of slack in the score restores all forty to `audio` with false matches unmoved. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> TRACES: AR-002, VR-005, VR-013, VR-014 | SR-002, SR-003 |
||
|
|
af5208035e |
docs: the RANSAC aligner was a defect, measured
AR-005 replaced cv::estimateAffinePartial2D(..., RANSAC, 3.0) with Umeyama least squares over all five points — the estimator InsightFace aligns with, and so the one the ArcFace/LVFace training crops were produced by. The first note here assumed the two agree wherever RANSAC keeps all five points, leaving a small divergence on non-frontal faces. Measured on 400 gallery headshots with the model held fixed, that was wrong: the crops disagree by a median 17 source px and 83.5% embed below cos 0.99 of their Umeyama counterpart. A 4-DoF similarity is exactly determined by two points, so every minimal sample fits its own pair perfectly and is scored on the other three; real landmarks sit a median 2.74 canonical px from any similarity fit, so a landmark outside the 3 px band is the common case and RANSAC returns an under-determined transform. How much that cost in accuracy is a separate question, and the honest answer is less than those numbers suggest. Rebuilding the full gallery moved the intra/inter separation the AR-023 calibration is fitted from by 0.583 to 0.590: the old warp was wrong but self-consistent, gallery and probe both went through it, and the embedder tolerates framing variation. The sharper evidence is duplicate detection — the rebuild dropped 1614 near-duplicates against the original build's ~100, because unstable two-point fits gave near-identical images visibly different vectors. That instability, not a headline accuracy delta, is what a tracker accumulating evidence across frames was paying for. Also records the AR-030 residual's real-data floor: on the most cooperative images the pipeline sees, it runs a median 2.74 px, so landmark noise occupies the first few pixels and the synthetic foreshortening ladder is optimistic about the low end. Any discount curve has to treat that range as uninformative rather than as mild pose, and VR-012 must set thresholds against the measured distribution. Tests carry the tag they verify: the residual's roll/scale invariance and monotonicity under foreshortening are what make it a pose measure rather than a pose-and-everything-else measure. TRACES: AR-005, AR-030 | SR-002 |
||
|
|
6aabeb9897 |
feat: provenance attributes on the embedding dump
VR-010 — a dump made with one detector/embedder pair was byte-indistinguishable from one made with another, except for the two attributes GR-004 added. Replayed against a gallery from a different model, cosine similarities are meaningless but look entirely plausible. The register states the principle directly: a fixture whose provenance is unknown is worse than no fixture, because it will be trusted. Sixteen attributes now record everything that determines the dump's content: detector model and thresholds, min_face_px, max_faces, cut_threshold, dense_scale, bbox_upscale, start/end, track_assoc_min_prob, and scene_detect. scene_detect is the one that matters most. is_scene_boundary is all-zero both when the detector found nothing and when it never ran, and those mean completely different things to a consumer — without the flag they are indistinguishable. No schema_version bump: new root attributes are additive and replay.py already reads attributes with a default, so older dumps stay readable and the committed fixtures — which predate this — still load. Also corrects SCHEMA.md, which claimed bbox was already mapped to original resolution at dump time. It is not; the upscale is applied downstream in the matcher, after the dump tap. Harmless while dense_scale is 1 and silently wrong otherwise, so bbox_upscale is now recorded and the doc says what the code does. Verified end to end: all sixteen attributes present and correct on a freshly generated dump. Suite: 92 cases, 6136 assertions. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> TRACES: VR-010, VR-001 | PR-002 |
||
|
|
c843c4abe3 |
feat: banded admission for the per-subject embedding store
AR-018 — an embedding joins a track's store only if its similarity to something already there falls inside a band, rather than merely being far from the gallery. Above the upper bound it is redundant: another look at a pose the store already covers, teaching the annex nothing while costing a slot a novel view could have used. Below the lower bound it is suspect: within one track every face is the same person by construction, so an embedding unlike everything else on the track is evidence that construction failed — a track-ID collision or a bad detection. Admitting it is exactly how an actor's annex gets poisoned with someone else's face. The old gate had only the upper half of that idea, expressed as a raw cosine against the gallery. Both bounds are now calibrated probabilities (AR-024), so the same number means the same thing here as in association and evidence weighting rather than three different things. This catches track-ID collisions EARLIER than the spread gate did — at the door rather than at promotion — so the buffer never becomes two-person in the first place. The spread gate stays as a second line for a track that drifts gradually instead of jumping. The existing test was asserting the mechanism rather than the outcome, so it was rewritten to assert what actually matters: whichever gate fires, the outsider must not reach the annex. Rejections are counted. A store that admits nothing is as broken as one that admits everything, and neither is visible otherwise. Band defaults 0.90-0.95 are working values pending VR-007; the two bounds fail in opposite directions and must be swept separately. Suite: 92 cases, 6133 assertions. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> TRACES: AR-018, AR-024 | SR-005 |
||
|
|
cc1bed92d8 |
perf: back the CPU similarity GEMM with OpenBLAS
The CPU path was a scalar triple loop. It is the correctness oracle for the GPU backends, but it is also what CI runs — there is no GPU on the N100 host — and since AR-003 removed the per-frame face cap, a crowded frame now scores many faces against a library-scale gallery. Scoring one face against 5000 embeddings is 2.6 MFLOP; in scalar that does not hold up (AR-027). S(g,f) viewed as row-major [n_faces x n_gallery] is exactly query * gallery^T, so the loop nest collapses into a single cblas_sgemm. OpenBLAS is optional in the build: found via pkg-config, and the scalar path remains when it is absent so no hard dependency is added and the two can be diffed when a similarity looks wrong. The configure step warns rather than failing, since a developer without it should still get a working tree. The test target links it too. Without that the suite compiles the scalar fallback while the builder image ships CBLAS, so CI would be verifying a kernel that is not the one running in production — the same class of mistake as testing a path the gate never executes. Recorded as required (not optional) in the DP-007 image, for the same reason. Suite: 92 cases, 6136 assertions, with CBLAS compiled in. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> TRACES: AR-026, AR-027, DP-007 | SR-001 |
||
|
|
13bdc27566 |
feat: join the decode butterfly so scene boundaries reach the face branch
AR-010 — is_scene_boundary had no producer: SceneDetectorFunc was a terminal sink writing scenes.json and never annotating the frames flowing to face detection. The flag was permanently false, so the boundary half of AR-007's frame-dependent association was dead code that a test could still exercise synthetically and appear to verify. The topology already forks after decode — dense frames to TransNetV2, sampled frames to face detection — so this is a fork-join. SceneBoundaries is the join: the detector publishes each window's verdict with a watermark, and an annotator on the sampled branch stamps the flag. The watermark is the part that matters. TransNetV2 buffers 100 frames before it can score any of them, so at any instant it has an opinion up to some time T and none after. Without recording T a consumer cannot tell "no boundary" from "not scored yet", and those demand opposite behaviour — treating unscored frames as boundary-free is exactly what makes a downstream check pass while verifying nothing. Buffering alone does not work, which was my first attempt. Channel depth creates lag only when the consumer is slower, and the face branch runs four orders of magnitude faster per frame than TransNetV2 (0.01ms vs 400ms), so its channels drain instantly and no lag accumulates. Measured: 106 of 364 frames outran the detector. The annotator therefore waits on the watermark explicitly. The detector signals completion so the tail cannot deadlock, and publishes from flush_remaining too — without that the final frames arrive with no verdict. Boundaries are deduped on publish, matching what scenes.json does at write time. A run of adjacent high-scoring frames is one boundary, not several; leaving them raw made this view report 357 where the file said 13. Now the two agree exactly. Frames past the detector's last scored window remain unverified and are counted as such rather than silently marked boundary-free. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> TRACES: AR-007, AR-010 | SR-002 |
||
|
|
fa1c494825 |
docs: AR-010 is blocked on a design decision, not an implementation gap
Making SceneDetectorFunc a pass-through does not work. TransNetV2 buffers 100 dense frames before it can score any of them and trusts only each window's centre, so a boundary at time T is not known until roughly 3.3s after T at 30 fps. The face pipeline runs on a parallel branch and has long since passed T. An association hint that arrives after the association is worthless. Three options recorded with their costs: two-pass (correct, doubles the decode that already dominates runtime), delaying the face branch (couples the two branches' timing, which invites heisenbugs under backpressure), or leaving it unwired. Leaving it unwired costs less than it looks, which is what makes this a decision rather than a defect. The redesign made cuts and boundaries do the same thing — both say "spatial continuity is broken, associate on embedding" — so TransNetV2 adds nothing over the histogram except on transitions the histogram cannot see: slow dissolves and fades. That gap is real but narrow. Where TransNetV2 still earns its cost is AR-019, whose promotion gate wants a span free of cuts and boundaries. A late answer is fine there, because promotion happens on track confirmation rather than per frame — so it can be wired offline against the collected boundary list, off the hot path entirely. Recommendation: leave the association path on is_cut alone, wire boundaries into AR-019, and revisit if dissolve-heavy material shows association failures. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> TRACES: AR-010, AR-019 | SR-002 |
||
|
|
09a4650fd9 |
feat: backpressure — the pipeline slows instead of losing frames
Picks up the KPN fix: node data outputs block on a full channel rather than dropping. Sentinels stay out-of-band, so EOF can always overtake a stalled data path and the hold-and-wait deadlock that comment warns about is not reachable. Verified on a 77s clip at 5 fps, which should yield 385 sampled frames: before 65 written, 320 dropped, 29s, two runs differ after 385 written, 0 dropped, 17s, two runs byte-identical The determinism is the part that matters. Golden fixtures were impossible while what got dropped depended on timing; VR-001 fixture generation is unblocked by this, and so is the CI replay strategy that depends on it. Faster rather than slower, which is worth recording because the intuition runs the other way: a dropped frame has already cost its decode, and the overflow exception cost more still. AR-004 is not fully closed. Channel capacity remains a count of items, while a face carries a 112x112 crop and a 512-float embedding — so a crowded frame occupies far more memory per slot than a sparse one. Bounding by bytes in flight is the remaining half, and it matters once max_faces is removed (AR-003). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> TRACES: AR-004, VR-001 | SR-002 |
||
|
|
b98372bad8 |
docs: AR-004 is a KPN change, with the measurement behind it
Backpressure cannot be implemented in this repository. Every node output in KPN uses the dropping push() (pool_node.hpp:404 and :710, plus branch, fanout and interrupt_node). A lossless push_blocking() already exists on both Channel and OutputPort — "wait for the consumer to drain instead of dropping; the producer just runs slower" — and nothing calls it. The fix is a per-channel policy or a network default in KPN, and this pipeline should select lossless: a dropped frame here does not degrade a result, it silently changes one. Measured rather than inferred. One 77s clip at 5 fps should yield ~385 sampled frames. On CPU it produced 49, ending at 51s, with 285 dropped at camera_pos and 51 at face_aligner. Rebuilt with CUDA the same clip ran in 29s and reached EOF correctly, and still dropped 320 at camera_pos, yielding 65. Faster hardware moves where the queue backs up; it does not change what happens when it does — which is why this is a correctness requirement rather than a throughput one. Raising channel capacity is therefore a stopgap: it lowers the probability of overflow without changing the behaviour on overflow, and the failure it hides is silent corruption of the output. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> TRACES: AR-003, AR-004, VR-001 | SR-002 |
||
|
|
b7c96641a9 |
docs: no exact tier — the file-hash tier was withdrawn
A stale reference to 'the runtime/exact tiers' as the fallback for media too short to carry an audio signature. The exact tier keyed on a file hash and was withdrawn on legal grounds: it fingerprinted the individual copy a user holds rather than the cut the timings describe. The pipeline never emitted a video_hash, so nothing in the code changes — but a spec that still names a withdrawn tier is what makes the withdrawal look like an oversight to the next reader, which is exactly how it nearly got re-added. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
908d166173 |
feat: bind galleries to the embedder that built them
GR-004 — a gallery built with one embedding model is meaningless with another. Cosine similarities across models are garbage but look entirely plausible, so this fails silently and expensively; every measurement taken against a mismatched pair would have been quietly wrong. The stamp is the model basename plus a SHA-256 of its bytes, with embed_dim as a cheap extra guard. The hash decides and the name explains, because neither works alone: a name is a promise rather than a fact — models get re-exported in place under an unchanged filename, which is exactly the case where the weights differ and nothing else does — while a bare hash mismatch tells an operator nothing actionable. Mismatch is fatal in every mode with no bypass. Unstamped only warns, because unstamped is unknown rather than known-bad, and an error firing on every legacy gallery trains people to reach for the bypass reflexively. scripts/stamp_gallery.py binds an existing gallery in place with no re-embedding, so the warning is a migration step rather than a permanent state; --require-gallery-stamp promotes it to an error once a site has migrated. Two gaps found that would have defeated the requirement outright: - Embedding dumps carried no stamp, so a replay — which has no live embedder — had nothing to check the gallery against. Dumps now carry embedder_model and embedder_sha256 as root attributes. Additive; schema_version stays 1. This is the same gap the dump audit identified independently. - --merge produced one file holding two embedding spaces, which no later check can untangle. Merge paths now verify before writing. The stamp also survives identity_matcher's calibration write-back, which would otherwise have stripped it on the first analysis run — the check would have worked exactly once. Conflicts resolved additively: both branches appended a source to sae_gallery and to the test target, and both edited the GR-004 register row. Merged suite: 64 cases, 3199 assertions, passing on CPU with no GPU. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> TRACES: GR-004, VR-001 | SR-001 |
||
|
|
7db40f430d |
GR-004: bind galleries to the embedder that built them
A gallery is only valid for the embedder that produced its vectors. Cosine
similarities across models are meaningless but *look* plausible, so the mistake
is silent and every measurement taken afterwards is suspect. Stamp the embedder
identity into the gallery at build; verify it at every load.
The stamp is the model file's basename plus the SHA-256 of its bytes (plus
embed_dim). The hash decides, the name explains. A name alone is a promise
rather than a fact — models get re-exported and overwritten in place under an
unchanged filename, which is exactly the case where the weights differ and
nothing else does. A hash alone is correct but unactionable in an error message.
SHA-256 is derived from the artefact, needs no registry kept current, and costs
~0.1s for a 250MB ONNX, memoised per process.
Mismatch is a hard error in every mode, with no bypass, naming both sides.
Unstamped legacy galleries warn loudly and proceed: unknown is not known-bad,
and hard-failing every pre-existing gallery would turn the check into something
people disable rather than trust. --require-gallery-stamp (or
SAE_REQUIRE_GALLERY_STAMP=1, which propagates to subprocesses) promotes that to
a hard error — the mode measurement work should run in. scripts/stamp_gallery.py
re-binds an existing gallery with no re-embedding, so "warn" is a cheap state to
leave rather than a permanent one.
Embedding dumps carry the same stamp: a replay has no live embedder, so the dump
is the embedder as far as the gallery is concerned. Derived galleries inherit
their source's stamp; --merge and the JSON gallery merge check before writing,
since one file holding two embedding spaces cannot be untangled afterwards.
Verified in: scene_analyze, scene_preview, the sae_kpn matcher binding,
replay.py, optimize.py (once per film at startup, before the first evaluation),
movienet_eval.py and both merge paths.
Stamp logic lives in src/gallery/embedder_stamp.{hpp,cpp} and its Python twin
scripts/sae_stamp.py, kept dependency-light so replay subprocesses do not pay
sae_gallery's requests/Pillow import to ask whether two models match.
Tests: 12 new cases in test_gallery_store.cpp covering the comparison logic,
both round trips, and the SHA-256 vectors that guarantee the C++ and hashlib
stamps agree. No ONNX or GPU required.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
020306c94f |
docs: builder images and per-backend release binaries (DP-008)
Adds the build/deploy story that was missing: containerised builder images for cpu / cuda / rocm, and release jobs producing prebuilt binaries so a first install need not compile. States explicitly that this does not reverse DP-005. That requirement rejects Docker as a *runtime* — GPU passthrough is fragile and exists only because of the container. Using it as a *build* environment is the opposite case, and lets one machine produce binaries for backends it cannot itself run. Build in a container, run natively. Two things deliberately cannot ship, and the installer must not imply otherwise: TensorRT engines are GPU-architecture and TRT-version specific, so build_trt_engines.sh still runs on the target; and models are ~725 MB in LFS, orthogonal to the binary. The base image is chosen by the OLDEST glibc to be supported, not by convenience — a binary built in a container runs against the host's glibc, and getting this wrong fails at load with GLIBC_2.xx not found. Accelerator runtimes have the same shape of problem, so each image documents its compatible CUDA/ROCm range and the installer checks it rather than discovering a mismatch at first inference. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> TRACES: DP-005, DP-007, DP-008 | PR-004 |
||
|
|
2919ed68d1 |
docs: audit findings, verification tiers, CI image, artifact storage
Corrections from the dump audit and the completed agent work: - AR-010 is not started, not in progress: is_scene_boundary has no producer anywhere. SceneDetectorFunc is a terminal sink writing scenes.json and never annotates the frame, so the field is permanently false and the dump column a constant 0. A replay test of the frame-dependent track_alpha would pass vacuously — the worst failure mode for a verification gate. - T1 (functor-level) becomes the primary verification tier, not T2. KPN node functors are plain callables constructed outside the network, so a node is tested by calling operator() with hand-built inputs. That removes four hazards at once: fixture provenance, replay-from-frame-0, cross-test state leakage, and replay-harness nondeterminism. It also means a dead upstream producer no longer blocks testing its consumer. - VR-010 (dump provenance) and VR-011 (replay harness rewrite) added. A dump made with LVFace is currently byte-indistinguishable from one made with w600k-R50 — the GR-004 problem again, in the dump. - Four requirements had no verification tier at all; the traceability gate found them. - DP-007: CI builder image, CPU-only, pinned by tag in the Gitea container registry. Corpus fixtures go to the package registry rather than LFS: LFS is pulled on clone and would tax every developer for data only CI reads. - IR-004/005/007/008 marked done; the v1 DSP parameters they had to pin are now normative in the server spec. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> TRACES: AR-010, DP-007, IR-004, IR-005, IR-007, IR-008, VR-001, VR-010, VR-011 | SR-002, SR-003 |
||
|
|
43d2c976c3 |
docs: replace phased plan with a per-requirement one
The phase structure encoded ordering assumptions that stopped being true as the design changed, and its Phase 2 still described retuning constants that are now withdrawn. Ordering is now derived from per-requirement dependencies instead: anything with no unmet dependency is startable. Carries over the TrackRegistry design (now keyed to AR-012/AR-013) and records what was withdrawn from the old plan, including the --presence-mode flag — comparison against old behaviour uses recorded reference output rather than a second live code path. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
a5299daf6e |
docs: software spec, requirements register, and implementation plan
Adds the requirements baseline for the pipeline redesign: - SPEC.md — software requirements with Current/Gap deltas per item, so the document doubles as a work list. - requirements.md — stable flat IDs (AR/DP/IR/GR/VR) with parent traces, priorities, statuses, and a per-requirement verification plan. Replaces the thematic A1..E8 scheme, which had already produced an A1a and an out-of-order E6; IDs are now permanent and never reused. - IMPLEMENTATION-PLAN.md — phased work. The central change is AR-012: presence follows track extent rather than per-frame recognition, so a window starts when an actor appears rather than when the recogniser first succeeded. anneal_sec and extinction_sec are withdrawn rather than retuned — a track that survives its own gaps leaves them nothing to do. Verification is shaped by CI running on an N100 with no dGPU: the existing HDF5 dump makes everything downstream of embedding replayable on CPU, which covers the bulk of the redesign. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |