feat(quality): score every face on sharpness and alignment before it is evidence

Every embedding now carries the quality of the input it came from. Both
axes fall out of the AR-005 warp for free: crop_sharpness() is the
normalised Laplacian variance over the aligned 112x112, so contrast and
size cannot leak into it, and the alignment residual is the part of the
landmark deformation a similarity transform cannot explain, so in-plane
roll reads as zero and foreshortening does not.

Carried, not consumed. Nothing discounts or thresholds on either number
yet -- that is AR-030 and VR-012, and the knee has to be located against
recorded data before a gate is chosen. What this change buys is that the
data exists to locate it with.

No face is admitted unscored: the -1 sentinel is preserved rather than
clamped, and a degenerate landmark fit is counted rather than silently
dropped.

Takes the VR-001 dump to schema_version 2. The bump is not for readers,
which check for the datasets by name and replay a v1 dump unchanged; it
is so a consumer can tell "never scored" from "scored zero", which is
not recoverable from the arrays afterwards.

TRACES: AR-028, AR-029, AR-030 | VR-001 | SR-002
This commit is contained in:
2026-08-05 14:37:30 +02:00
parent c1155cb607
commit 777c98cb33
10 changed files with 726 additions and 17 deletions
+91 -1
View File
@@ -1,5 +1,5 @@
#pragma once
/// TRACES: AR-005, AR-030 | SR-002
/// TRACES: AR-005, AR-029, AR-030 | SR-002
#include "types.hpp"
#include <opencv2/core.hpp>
@@ -143,6 +143,96 @@ inline cv::Mat align_face(const cv::Mat& img,
return crop;
}
// ── crop_sharpness ────────────────────────────────────────────────────────────
/// TRACES: AR-029 | SR-002
//
// Normalised variance of the Laplacian over the aligned 112×112 crop: the AR-029
// sharpness axis. Returns -1 for an empty crop (unscored), matching the
// DetectedFace sentinel.
//
// sharpness = Var(∇²I) / Var(I)
//
// Two normalisations, each removing a quantity that would otherwise be read as
// blur:
//
// - **Divided by the image variance, so contrast cannot leak in.** Scaling
// intensity by α scales the Laplacian by α too, so both variances scale by α²
// and the ratio is unchanged. A raw Var(∇²I) — the textbook measure — instead
// falls with exposure, so a dim scene reads as soft and a graded-up one as
// sharp. VR-012 has to locate one knee across films whose grading differs by
// more than their focus does; an uncalibrated measure would put the knee in a
// different place per film, which is the AR-024 failure in another metric.
// - **Measured on the aligned crop, so size cannot leak in.** The destination
// frame is fixed at 112×112 (AR-002 owns size, and double-counting it here
// would make every small face read as blurred). What the ratio reports is the
// detail actually present in the embedder's input — so a small sharp face can
// and does outscore a large soft one. That is the claim; it is *not* a claim
// of invariance to source resolution, because a 40 px face warped up to 112
// genuinely carries less detail, and hiding that would defeat the point.
//
// Frequency-domain reading of why the blur ladder is monotone: with
// Var(∇²I) = ∫|ω|⁴|F(ω)|² and Var(I) = ∫|F(ω)|², the ratio is E[|ω|⁴] under the
// image's own spectral measure. Gaussian blur multiplies that measure by
// e^{-σ²|ω|²}, concentrating it at low |ω|, so the expectation falls strictly
// with σ. It is a property of the construction, not a fitted behaviour.
//
// **Three known hazards, for VR-012 to check rather than for a threshold to
// absorb.** All are recorded here because they are properties of the measure,
// visible in the dumped distribution, and neither should be papered over by a
// correction chosen before that distribution has been looked at.
//
// 1. **Border fill.** `align_face` warps with BORDER_CONSTANT, so a face
// crossing the frame edge brings a hard black step into the crop, and a
// step edge is high-frequency. The normalisation blunts it — the fill
// inflates Var(I) as well as Var(∇²I) — but does not remove it, so
// heavily-cropped faces may read sharper than they are. The fix is either a
// validity mask or a different border mode, and the second changes what the
// embedder is fed (AR-011).
//
// 2. **The contrast invariance is exact in the algebra and approximate in
// 8 bits.** Scaling I by α cancels exactly; what does not cancel is the
// quantisation floor of a stored crop, which is broadband and so lands in
// the numerator. It matters only where there is little signal left to
// compete with it: on the AR-029 test texture a half-contrast copy reads
// 0.9% high when sharp, 24% high at sigma 1.2 and 148% high at sigma 2.5.
// A crop that is both **dim and soft therefore reads sharper than it is** —
// the low corner of the axis, and the corner VR-012 must put a knee in.
//
// 3. **It reports where the energy sits, not how much there is.** A crop whose
// energy is *already* concentrated at high frequency — dense film grain,
// a face against foliage — loses numerator and denominator together under
// blur, so the ratio moves less than the damage does. Measured on a
// flat-spectrum synthetic, an anisotropic (motion) smear even makes it rise,
// because the surviving perpendicular detail really is as fine as before.
// Natural crops have the low-frequency mass that keeps the denominator
// steady, and on those both ladders fall (see the AR-029 tests, which use a
// 1/f texture for exactly this reason). The same property means the axis
// conflates focus with intrinsic texture — a bearded face outscores a smooth
// one at equal focus — which is true of every no-reference sharpness measure
// and is why AR-028 carries the number instead of thresholding on it.
inline float crop_sharpness(const cv::Mat& crop) {
if (crop.empty()) return -1.f;
cv::Mat gray;
if (crop.channels() == 3) cv::cvtColor(crop, gray, cv::COLOR_BGR2GRAY);
else gray = crop;
cv::Mat lap;
cv::Laplacian(gray, lap, CV_32F, 3);
cv::Scalar mean_i, sd_i, mean_l, sd_l;
cv::meanStdDev(gray, mean_i, sd_i);
cv::meanStdDev(lap, mean_l, sd_l);
const double var_i = sd_i[0] * sd_i[0];
// A flat crop has no detail to be sharp or soft about, and the ratio is 0/0.
// Zero is the honest answer and keeps the axis finite; -1 would claim the
// face was never scored, which is a different fact.
if (var_i < 1e-6) return 0.f;
return static_cast<float>((sd_l[0] * sd_l[0]) / var_i);
}
// ── enhance_for_retry ────────────────────────────────────────────────────────
// Used when initial face detection finds nothing. Pads the image by 50%
// (border-replicated, so the detector doesn't see a hard edge) and applies