Files
scene-actor-extraction/docs/quality-knee.md
T
dtourolle 889018aa34 docs(VR-012): the quality knee, and what it retires
Sharpness is not a sufficient statistic for identity loss. Six cells at
effectively identical measured sharpness span 15.3% to 91.0% TPI,
ordered entirely by source size, because a scalar keyed on
high-frequency energy cannot separate attenuated high frequencies from
destroyed spatial sampling. AR-028's "kept separate, not collapsed into
one scalar" now rests on a measurement rather than an argument -- and
the reasoning it used to rest on, that the aligned crop is
scale-normalised so a measure there cannot re-measure size, was wrong
and is corrected in place.

Variance of Laplacian -- the most widely used blur metric there is -- is
anti-predictive at fixed degradation on all three blur families. The
decile it calls sharpest is 2.6x less identifiable than the decile it
calls blurriest, monotone across ten bins, because within a cell its
residual variance is native contrast rather than detail, and hard
shadows and JPEG ringing raise it while making a face harder to match.
Gating on it would preferentially discard the more identifiable faces.

Blur breaks confidence, not identity: rank-1 holds at 80.2% where TPI is
15.3%, and FPI never left 0.1% in any of the 108 cells. Degradation
produces abstention, never a wrong name. That is also why sharpness
fails as a compute gate -- even a visually destroyed face stays 46.9%
identifiable, so a gate discards recoverable evidence at three times the
cost of the free size filter. Discount, do not gate; the rule AR-028
already stated now has evidence for why it is right rather than merely
cautious.

Records the shape a discount must have (flat, then a cliff between sigma
2 and 3), that its cost scales with proximity to the decision boundary
rather than with blur, and that the pose half of VR-012 has not been
run.

TRACES: VR-012, AR-028, AR-029, AR-030 | SR-002
2026-07-31 22:29:47 +02:00

16 KiB
Raw Blame History

Quality knee: what does a blurred or small face cost, and can a measure predict it?

VR-012. Companion to the minimum-face-size studies VR-005 and VR-013 (see the requirement register), which located the size floor at 40 px; this asks the same question for sharpness, and asks whether any cheap measure taken on the aligned crop can be acted on at inference.

Run by scripts/validation/quality_knee.py through the sae_embed bindings — detection, the ArcFace warp, the embedder, the five candidate measures and the Platt calibration are all the shipped C++.

Protocol

1670 gallery actors with 3 or more mugshots (of 2456 total), one image held out per actor as a probe, the remaining 10326 embeddings staying in the gallery at native resolution. Only the probe degrades — reference mugshots are clean and the face coming out of the video is not.

Each probe passes through a joint grid: downscale to S×S and back to 112 (the sampling loss), then blur at level L in canonical pixels. Three blur families, 36 cells each, 60120 probe-cell records per family:

family models parameter
Gaussian soft focus, a generic stand-in sigma 0 … 3
Disc real optical defocus — the circle of confusion radius 0 … 6
Motion camera pan or moving subject length 0 … 21 px

The three are not interchangeable, and sweeping only the first was the original design error — one that would have produced a wrong answer, not merely an incomplete one (Result 3). A defocused lens spreads a point into a uniform disc, whose transfer function is a jinc — 2·J1(x)/x — that crosses zero and goes negative, annihilating whole frequency bands and returning the ones beyond each zero phase-reversed. A Gaussian MTF is strictly positive and monotone and does neither. More practically: defocus and motion are how a face ends up large and useless, while Gaussian blur as swept here mostly co-occurs with small faces. That difference decides whether sharpness carries anything the size filter does not.

Families are compared at matched per-axis PSF standard deviation (σ for a Gaussian, R/2 for a disc, L/√12 for a linear smear), never at equal raw parameter, which would compare different amounts of damage.

Identification is the pipeline's own decision: per-actor best-of-N cosine → Platt sigmoid → accept above prob_threshold 0.754. Never a raw cosine (AR-024).

Result 1 — sharpness is not a sufficient statistic

Sorting the 36 Gaussian cells by hf_energy_ratio, the six sigma-3 cells land at effectively identical measured sharpness:

size sigma hf_energy_ratio TPI
16 3 0.0003 15.3%
24 3 0.0003 63.2%
32 3 0.0003 79.4%
48 3 0.0003 86.6%
64 3 0.0004 88.4%
112 3 0.0005 91.0%

Same measured sharpness, a 76-point spread in identification. It inverts too: 16 px unblurred measures 0.0033 and scores 23.5%, while 48 px at sigma 2 measures lower at 0.0021 and scores 96.6%.

A canonical-frame sharpness scalar cannot separate attenuated high frequencies from destroyed spatial sampling. Blur suppresses the high band while preserving mid-frequency facial geometry exactly; downsampling to 16 px destroys that geometry outright. Both look alike to any measure keyed on high-frequency energy.

This is the measured basis for AR-028's rule that the axes are kept separate and not collapsed into one scalar, and it settles the double-counting question: size and sharpness are not redundant, and neither substitutes for the other.

Result 2 — blur is a cliff, and it breaks confidence, not identity

TPI % by size (rows) against Gaussian sigma (columns):

size 0 0.5 1 1.5 2 3
16 23.5 24.0 25.0 24.6 23.9 15.3
24 85.7 85.1 85.6 85.9 82.6 63.2
32 95.9 95.9 96.0 95.5 93.7 79.4
48 98.7 98.7 98.4 98.1 96.6 86.6
64 98.6 98.6 98.8 98.4 97.5 88.4
112 98.9 98.9 98.8 98.6 98.0 91.0

Three regimes: sigma ≤ 1.5 is free (every cell moves under 1.5 points, sign flipping at random — at 16 px it slightly improves, smoothing upscale artifacts); sigma 2 costs 13 points; the 2→3 step costs 719. A smooth discount curve is therefore the wrong shape — the response is flat, then falls off a cliff.

The cost peaks at the size knee, not at full resolution. Sigma 3 costs 22.5 points at 24 px but only 7.9 at 112 px and 8.3 at 16 px. Blur has no intrinsic cost; it costs in proportion to how close the observation already sits to the decision boundary. At 112 px there is margin to spare, at 16 px the probe is already below threshold, and at 24 px it sits exactly on the knee.

What blur destroys is confidence, not ranking. Rank-1 barely moves: 99.3% → 99.2% at 112 px across the whole sigma range. The extreme case is 16 px at sigma 3, where rank-1 is 80.2% while TPI is 15.3% — 65 points of probes have the correct actor ranked first and are rejected anyway for falling under the probability threshold.

That is why FPI never left 0.1% in any of the 108 cells across all three families. Degradation produces TBI, never a wrong name. The calibration degrades gracefully, which is what SR-002 needs.

Result 3 — the blur family matters more than the blur amount

Comparing families by their raw parameter is meaningless — sigma, radius and length are different units. They are matched here by the per-axis standard deviation of the PSF, which puts them on one scale:

family per-axis σ level giving σ = 3 px
Gaussian σ σ 3
Disc radius R R/2 6
Motion length L L/√12 10.4

For reference the ArcFace template places the eyes 35.2 canonical px apart, so σ = 3 px is 9% of the inter-ocular distance.

TPI at matched severity, interpolated within each family:

size σ=3 Gaussian σ=3 Motion σ=3 Defocus defocus penalty
16 15.3 15.1 11.0 +4.3
24 63.2 61.1 41.4 +21.9
32 79.4 76.1 50.4 +29.0
48 86.6 80.9 52.6 +34.0
64 88.4 81.7 51.0 +37.4
112 91.0 82.1 46.9 +44.0

Optical defocus is up to 44 points more destructive than a Gaussian of identical spread, and the ordering is defocus ≫ motion > Gaussian throughout. At σ=1 the three families are indistinguishable, and at σ=2 they differ by under 5 points; the divergence appears only when both the blur is severe and the face is large.

That pattern is physically consistent. At 16 px the resampling has already removed the high frequencies, so the PSF's shape has nothing left to act on and all three agree. At 112 px the full spectrum is present and shape decides: a Gaussian MTF rolls off gently and always leaves some energy at every frequency, so the embedder receives a merely attenuated signal, while a disc MTF is a jinc that hits exact zeros — whole frequency bands annihilated rather than attenuated, with the bands beyond each zero returning phase-reversed. Motion sits between them because it ruins one axis and leaves the perpendicular one untouched.

The methodological consequence is the important one. This study originally swept Gaussian blur alone and concluded blur was a minor effect. On the family that actually occurs in film, the same nominal severity costs 53% error instead of 9% at full resolution. A threshold set from the Gaussian arm would have been wrong by a factor of five in error rate, and the axis would probably have been dropped as not worth its cost.

Defocus is also the case a size gate cannot catch. Every one of those 112 px faces is large and confidently detected, and sails through AR-002 untouched. That, not the Gaussian result, is what justifies a sharpness axis existing at all.

Result 4 — variance of Laplacian is anti-predictive at fixed degradation

Pooled across all cells, every candidate scores AUC 0.760.80 for predicting correct identification, with textbook var_laplacian top. That number is close to worthless: it rewards a measure for detecting how degraded the crop is, which all five do. The question a per-observation discount needs is whether, at a fixed degradation, the measure predicts which faces fail:

measure Gaussian Defocus Motion
hf_energy_ratio 0.530 0.521 0.557
norm_var_laplacian 0.520 0.507 0.539
dir_min_tenengrad 0.524 0.512 0.506
tenengrad 0.433 0.437 0.457
var_laplacian 0.423 0.422 0.473

Best is 0.557 — barely above chance, and hf_energy_ratio wins on all three families. var_laplacian is anti-predictive on all three too, so that finding does not depend on the blur model.

The two metrics measure different jobs, and the candidates split along that line. On the motion arm dir_min_tenengrad has the best pooled AUC by a wide margin — 0.854 against 0.792 for the next — exactly as its synthetic directional-blur ladder predicted, yet its within-cell AUC there is 0.506. It is an excellent detector of how badly smeared a crop is and no guide at all to which face will be recognised. Pooled AUC is the right metric for a gross-degradation flag; within-cell AUC is the right one for a per-observation discount; a measure can be strong at one and useless at the other.

Deciles within the 16 px Gaussian cell, where 1277 failures give the test real power:

var_laplacian decile TPI
0.000710.00192 (blurriest) 37.1%
0.002420.00278 22.8%
0.003970.00447 25.7%
0.006250.01445 (sharpest) 14.4%

The faces the measure calls sharpest are 2.6x less identifiable than those it calls blurriest, monotone across ten bins of 167. Within a cell every crop received identical degradation, so the residual variance is native contrast, not native detail — and hard shadows, high-contrast lighting, sharpening halos and JPEG ringing all raise Laplacian variance while making a face harder to match. The measure reads photographic style and encoding artifacts and calls them sharpness.

hf_energy_ratio is the only candidate with a correctly-signed within-cell trend (16.2% → 35.3% across the same deciles), being a pure ratio in which the contrast factor cancels.

Consequence: a per-face quality discount keyed on variance of Laplacian — the most widely used blur metric in production vision pipelines — would systematically down-weight the more identifiable faces. It is worse than no discount.

Result 5 — as a compute gate, sharpness loses to the size filter

Skipping the embed for crops below a threshold, measured as compute saved against true identifications lost:

gate skipped true IDs lost of skipped, doomed anyway
hf_energy_ratio < 0.00023 10.0% 7.6% 37.9%
hf_energy_ratio < 0.00051 20.0% 15.1% 38.7%
source size < 24 px 16.7% 4.7% 77.3%

At a comparable skip rate the size filter loses 4.7% against sharpness's 15.1% — three times less damage — and it is free, being a bbox dimension available before alignment or embedding, where sharpness needs the warped crop plus a colour convert, three convolutions and a 64×64 DFT.

Restricting to large faces (≥64 px) on the defocus arm, where the size filter is blind, improves the gate's precision 3.5x (37% of skipped crops doomed versus 10.7% on the Gaussian arm) but not its trade: skip 10%, lose 7.0%.

A hard ceiling explains why. At 112 px with defocus radius 6 — visually destroyed — 46.9% of faces still identify correctly and rank-1 is still 94.8%. Blur does not determine the outcome, so any gate keyed on apparent blur is predicting a coin flip. The size filter wins not because size is better measured, but because smallness destroys identity more completely than blur does: 16 px faces succeed only 23.5% of the time, so discarding them is cheap.

What this means for the requirements

Do not gate on sharpness; discount on it. Heavily defocused faces remain ~47% identifiable, so a gate destroys recoverable evidence. This is the first hard evidence that AR-028's "discounts the observation, never deletes the detection" is right on the merits rather than merely cautious. Since ranking survives where confidence does not, the per-track accumulation (AR-025) should recover much of what a single-frame threshold rejects — which is also the argument for the discount living in EvidenceDiscounter rather than in a filter.

var_laplacian and tenengrad are disqualified as discounts by Result 4, on all three blur families. They remain usable as coarse gross-degradation detectors, the role in which their pooled AUC is real — the same role the size filter plays — but they must never weight a per-observation belief.

hf_energy_ratio is the only surviving discount candidate, best on all three families, and its within-cell signal (0.520.56) is weak enough that shipping a discount on it needs justification beyond this study.

dir_min_tenengrad earns a different job. Its pooled 0.854 on the motion arm makes it the best available detector of gross directional smear — useful as a per-frame "this shot is unusable" flag, which is a decision about a frame, not a weighting of an observation. If AR-029 ships two measures for two roles, this is the second one, and it must not be confused with the first.

Model the blur family, not just its amount. Result 3 makes the choice of degradation model a first-order design decision rather than a detail: the same matched severity costs 9% or 53% error depending on the PSF. Any future study that sweeps blur must state which family it used and why.

Any discount curve must be flat then steep, not linear or sigmoid over the measure. Blur costs nothing until it costs a great deal.

Limitations

  • Cooperative population. Gallery mugshots are frontal and well-lit; within-cell failures are likely dominated by cross-view mismatch, which no sharpness measure can predict. Read the ~chance within-cell AUCs as "sharpness does not predict the dominant failure mode here", not as "sharpness is meaningless".
  • Uniform grid, not a natural distribution. Sizes and blur levels are sampled evenly, so "skip 16.7%" is exactly the 16 px row. The gate comparisons are like-for-like on identical records, but the absolute savings are not what a film would show.
  • TensorRT fp16. A different realisation of the embedder from the fp32 ONNX reference — VR-005 measured ~0.85 cosine agreement with separation intact. Gallery and probes share one session so the study is internally consistent, but the absolute knee belongs to the fp16 space.
  • Blur is applied in the canonical frame, after resampling, so its width is independent of the cell's size. Real optics blur before sampling.
  • The top motion rung is an anchor, not an operating point. Length 21 is a per-axis σ of 6.1 — 17% of the inter-ocular distance, a streak rather than a face — and it is swept to bound the curve, not because a frame like that is worth reasoning about. Its 3.4% TPI at 112 px should not be quoted as a headline. The same caution applies less severely to defocus radius 6 (σ = 3).
  • Per-axis σ equates spread, not perceptual damage. It is the fairest single scalar for comparing PSFs, but Result 3 is precisely the finding that equal spread does not mean equal harm, so the matched-severity tables compare like-for-like inputs, not like-for-like severity as a face would experience it.