Show the confidence, and say which curve it came from

FR-CULL-9 was read as "no fit, no number", so every library without 200
confirmed positive pairs showed "Confidence unavailable" on every
suggestion — which is every library, until enough confirmations exist to
fit one. The confirmations are made on this screen, ranked by the number
it was withholding, so the degraded state was also the permanent one.

There has always been a curve: Calibration::default is the reference
implementation's fitted MBF sigmoid, which is what clustering already
operates at. It is a published operating point, not an invention, and
what the requirement forbids is presenting it *as though it were
measured on this library*. So the percentage is shown, and the screen
says once, above the grid, where the curve came from.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-29 09:57:38 +02:00
co-authored by Claude Opus 5
parent 2fb3eb5d2d
commit c8e05831f4
7 changed files with 76 additions and 60 deletions
+10 -4
View File
@@ -614,10 +614,16 @@ positive pair is trustworthy by construction. Here the positives are bootstrappe
early confirmations (§8.1) and the whole risk is fitting confidently to a handful of them. The
reference also refuses to draw positives from an identity with fewer than five distinct embeddings,
letting it contribute negatives only — the same asymmetry applies to a thinly-confirmed person and is
worth keeping. In that state the UI says confidence is unavailable and the People view
still works — clustering falls back to a documented default operating point, labelled in the
interface as an untuned default, and no probability is displayed. That is FR-CULL-9's requirement
read literally: not presenting an untuned default *as though it were measured*.
worth keeping. In that state both clustering and the displayed confidences fall back to the
reference implementation's fitted curve — a documented operating point, not an invention — and the
People screen says so once, above the grid, rather than blanking every percentage. That is
FR-CULL-9's requirement read as written: what may not happen is an untuned default presented *as
though it were measured* on this library.
Blanking them was the first reading, and it was wrong in a way worth recording. A young library has
no fit; a fit needs confirmations; confirmations are made on a screen the user ranks by confidence.
Withholding the confidence until the fit exists is a deadlock in which the normal state of the
feature is its degraded one.
Refit is triggered by the same debounce as clustering (§9), and when the confirmed-pair count grows
materially.