GR-003 — the calibration fit already computed per-actor dedup counts, how many
actors are eligible for positive pairs, and a 200-bin histogram of the intra and
inter distributions, then discarded all of it to stderr. Nothing persisted, so
nobody could audit whether a gallery was any good.
The report is written alongside the gallery at build time. That is the right
moment: the matcher fits the same sigmoid at analysis time, but by then the
answer is per-run and nobody is looking, whereas build time is when a gallery's
quality is actually decided.
What it surfaces, in order of usefulness:
- actors with no usable image — a silent recall ceiling, since the pipeline can
never name them and nothing else says why
- actors below the positive-pair threshold — not broken, so nothing complains;
they just quietly weaken every threshold downstream
- near-duplicate references removed, per actor and total
- the fitted calibration AND the two distributions behind it
That last one is the point. Every threshold in the pipeline is expressed in the
probability space this sigmoid defines, so if the distributions overlap heavily
the calibration is weak and every downstream decision inherits it — while the
gallery still looks fine from the outside.
The gallery-derived prior, intra/(intra+inter), is computed and reported but the
shipped default of 0.5 is deliberately left alone. The spec records these as
disagreeing; now the real value is visible, so the decision can be made on
evidence rather than argument.
Three tests: a zero-image actor is visible in the report, an under-referenced
actor is counted, and the report round-trips through JSON.
Suite: 95 cases, 6142 assertions.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
TRACES: GR-003 | SR-001