Show the confidence, and say which curve it came from

FR-CULL-9 was read as "no fit, no number", so every library without 200
confirmed positive pairs showed "Confidence unavailable" on every
suggestion — which is every library, until enough confirmations exist to
fit one. The confirmations are made on this screen, ranked by the number
it was withholding, so the degraded state was also the permanent one.

There has always been a curve: Calibration::default is the reference
implementation's fitted MBF sigmoid, which is what clustering already
operates at. It is a published operating point, not an invention, and
what the requirement forbids is presenting it *as though it were
measured on this library*. So the percentage is shown, and the screen
says once, above the grid, where the curve came from.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-29 09:57:38 +02:00
co-authored by Claude Opus 5
parent 2fb3eb5d2d
commit c8e05831f4
7 changed files with 76 additions and 60 deletions
+7 -2
View File
@@ -908,8 +908,13 @@ similarity still *looks* like a plausible number all the way to the user interfa
confidence that does not mean what it says is worse than no confidence, because it is trusted.
The calibration shall be fitted per library from that library's own faces, and shall report whether
it is valid. Where it is not — too few examples to fit — the app shall say the confidence is
unavailable rather than present an untuned default as though it were measured.
it is valid. Where it is not — too few examples to fit — the app shall fall back to a **documented,
published operating point** (the reference implementation's fitted curve) and shall say, at the
screen level, that the confidences come from it. What is forbidden is presenting an untuned default
*as though it were measured on this library*; withholding the number entirely is not required and
shall not be done, because a screen of unranked suggestions is the state most libraries would
permanently sit in — the fit needs confirmations, and confirmations need a ranked screen to be made
on.
*Acceptance:* on a labelled corpus, the stated probability is within a documented tolerance of the
observed match rate across the probability range (a reliability-diagram check, not a single