Specify face indexing at native resolution, and say what the proxy cost
FR-CULL-8 said detection runs against the thumbnail or proxy tier and never a full decode, and faces.md §5 said the aligned crop is sampled from that same proxy. Both are wrong in the same place: they treat detection and cropping as one resolution problem when they are two, with opposite answers. Detection does not care. §4.1 fixes the graph's input at 640x640 and letterboxes whatever arrives, so a face filling 2% of the frame reaches the model at 12px whether the buffer handed over is 1024px or 6000px. Every pixel above the detector's own input is discarded before inference. The crop cares about nothing else. §5's warp produces the fixed 112x112 ArcFace sees, so source resolution converts directly into whether those 112 pixels were photographed or interpolated. Reading crop_px across the 18,671 faces the proxy-tier implementation stored: 47.3% were upsampled to reach the embedder, 314 of them by more than 2x, the smallest from 34 source pixels. An upsampled crop does not fail loudly -- it yields a confident embedding of detail that was never there, and the damage appears three stages later as clusters that will not separate. So FR-CULL-8 now specifies four stages with the resolutions named separately: render native through FR-EXP-9's pipeline, downscale for the detector, map boxes and landmarks back to native, crop and align from the native render. The affordability the old rule bought is met instead by when the pass runs -- background, preempted, resumable -- and the requirement says plainly what it now costs on a remote library: the original rather than FR-NC-3's byte range, 412 GB across the reference library's 19,107 images, so a whole-library pass is a transfer under FR-NC-6 rather than something that may start on its own. MIN_CROP_EDGE replaces the MIN_DETECT_EDGE this branch briefly had. Same number, guarding the quantity that turned out to matter. faces.md §7b records both measurements, and marks the second as unexplained rather than dressing it as a finding. Grouped by the buffer detection ran against, faces per image was 0.078 at 1024 or below and 1.82 at 2048 or better, controlled for file type and size. That gap is real and reproducible and I cannot account for it, because the letterbox above says detector input should not matter. M4 is where it gets settled. The crop measurement does not depend on it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
+63
-19
@@ -420,8 +420,10 @@ context the detector needs; CLAHE rescues underexposed frames. Both are plausibl
|
||||
and backlit photographs — worth trying at low priority, since a retry doubles the cost of exactly the
|
||||
images that yielded nothing.
|
||||
|
||||
Sampling is bilinear from the *source proxy*, in one step — never crop-then-warp, which resamples
|
||||
twice and throws away detail the warp could have used. Pixels falling outside the source are black.
|
||||
Sampling is bilinear from the **native render**, in one step — never crop-then-warp, which resamples
|
||||
twice and throws away detail the warp could have used, and never from the downscaled buffer the
|
||||
detector was given. The landmarks arrive in detector-input coordinates and are scaled back to
|
||||
native before the warp reads a single pixel; §7 is why. Pixels falling outside the source are black.
|
||||
|
||||
**The landmark order must match the template order.** The template above is written in the detector's
|
||||
own output order; if a future detector emits them differently, the template is reordered with it and
|
||||
@@ -465,18 +467,31 @@ drift that is otherwise invisible.
|
||||
|
||||
## 7. Which pixels the pipeline actually sees
|
||||
|
||||
FR-CULL-8 pins indexing to the FR-CULL-2 ladder: **the proxy tier, never a full decode.** The
|
||||
relevant tier is `ThumbSize::Large` — 1024 px on the long edge (`dr-thumbs`).
|
||||
**This section previously specified the proxy tier — `ThumbSize::Large`, 1024 px — as the buffer
|
||||
both detection and cropping read. That was wrong, and §7b is the measurement that says so.**
|
||||
FR-CULL-8 now separates the two, because they want opposite things:
|
||||
|
||||
That has a consequence worth stating in numbers rather than discovering in a clustering report. On a
|
||||
1024 px proxy:
|
||||
|
||||
| Face size in frame | Pixels across | What §6 receives |
|
||||
| Stage | Resolution | Why |
|
||||
|---|---|---|
|
||||
| A portrait, face fills a third of the frame | ~340 | Downsampled to 112. Ideal. |
|
||||
| Two people, half-length | ~120 | Roughly native. Good. |
|
||||
| A group of eight | ~50 | **Upsampled** to 112. Degraded, and usable. |
|
||||
| A figure in a landscape | ~20 | At or below §4.3's floor. Rejected. |
|
||||
| Source render | **native** | The only stage where more pixels exist to be had |
|
||||
| Detector input | downscaled to ~640 | §4.1 letterboxes to 640×640 regardless; more is wasted CPU |
|
||||
| Crop + align | **sampled from the native render** | The 112×112 is fixed, so this decides whether it holds real pixels |
|
||||
| Embedding | 112×112 | §6 |
|
||||
|
||||
The detector's indifference to resolution is the whole reason the split works. §4.1 fixes its input
|
||||
at 640×640 and letterboxes whatever arrives, so a face occupying 2% of the frame presents at 12 px
|
||||
to the model whether the buffer handed over is 1024 px or 6000 px. Feeding it native pixels buys
|
||||
nothing. **Feeding the *crop* native pixels buys everything**, because §5's warp is the one place
|
||||
where source resolution converts directly into embedding quality.
|
||||
|
||||
What the crop receives, by face size, from a native render of a 24 MP frame (~6000 px long edge):
|
||||
|
||||
| Face size in frame | Pixels across, native | Pixels across, 1024 proxy | What §6 receives |
|
||||
|---|---|---|---|
|
||||
| A portrait, face fills a third of the frame | ~2000 | ~340 | Downsampled. Ideal either way. |
|
||||
| Two people, half-length | ~700 | ~120 | Native: comfortable. Proxy: marginal. |
|
||||
| A group of eight | ~290 | ~50 | Native: real pixels. Proxy: **upsampled 2.2×**. |
|
||||
| A figure in a landscape | ~120 | ~20 | Native: usable. Proxy: below §4.3's floor. |
|
||||
|
||||
So `faces` records one column beyond catalog.md §10.1's schema:
|
||||
|
||||
@@ -484,14 +499,12 @@ So `faces` records one column beyond catalog.md §10.1's schema:
|
||||
ALTER TABLE faces ADD COLUMN crop_px INTEGER NOT NULL; -- source pixels across the aligned crop
|
||||
```
|
||||
|
||||
`crop_px` earns its place three times over. It is the honest quality signal for the UI; it is a
|
||||
`crop_px` earns its place four times over. It is the honest quality signal for the UI; it is a
|
||||
**feature in §8's calibration**, which FR-CULL-9 explicitly demands ("a raw cosine means something
|
||||
different for every model, every population, and *every face size*"); and it is what a future
|
||||
higher-resolution re-embedding pass would select on, so that pass becomes a query rather than a
|
||||
re-index of everything.
|
||||
|
||||
**No RAW decode is added.** Where the Large proxy is missing, the job enqueues a `Thumbnail` job at
|
||||
background priority and re-queues itself, exactly as FR-CULL-8 requires.
|
||||
different for every model, every population, and *every face size*"); it is what a
|
||||
higher-resolution re-embedding pass selects on, so that pass is a query rather than a re-index of
|
||||
everything; and it is the only way to audit whether the rule above is actually being followed —
|
||||
which is how §7b found that it was not.
|
||||
|
||||
---
|
||||
|
||||
@@ -525,6 +538,37 @@ Three things fall out of it that were not otherwise available:
|
||||
|
||||
---
|
||||
|
||||
## 7b. What the proxy tier actually cost, measured
|
||||
|
||||
The table in §7 predicted upsampling for small faces and called it "degraded, and usable". On the
|
||||
reference library of 23,531 images it was not the edge case that description implies. Reading
|
||||
`crop_px` across the 18,671 faces stored under the proxy-tier implementation:
|
||||
|
||||
| Source pixels across the aligned crop | Faces | Share |
|
||||
|---|---|---|
|
||||
| <56 (upsampled more than 2×) | 314 | 1.7% |
|
||||
| 56–111 (upsampled) | 8,505 | 45.6% |
|
||||
| 112–223 (roughly native) | 6,313 | 33.8% |
|
||||
| ≥224 (downsampled — ideal) | 3,539 | 19.0% |
|
||||
|
||||
**47.3% of every face in the library was upsampled to reach the embedder**, with `crop_px` as low
|
||||
as 34 — a 3.3× enlargement — against a mean of 178. An upsampled crop does not fail loudly. It
|
||||
produces a confident 512-d embedding describing detail that was interpolated rather than
|
||||
photographed, and the damage appears three stages later as clusters that will not separate.
|
||||
|
||||
A second effect, recorded here because it was measured and because the mechanism is **not**
|
||||
established. Grouping the same runs by `face_index.source_edge` — the buffer detection ran
|
||||
against — gives 0.078 faces per image at 1024 or below, against 1.82 at 2048 or better. Controlled
|
||||
for file type and size (1,592 DNGs averaging 21.0 MB against 7,724 averaging 23.3 MB, same library,
|
||||
same cameras), so it is not a composition artefact. But it cannot be a matter of the detector
|
||||
seeing fewer pixels, since §4.1 letterboxes both to 640: a 1024 buffer and a 3072 buffer present
|
||||
the same face at the same size to the model. The likeliest explanation is that the 1024 proxy is
|
||||
itself a downscale of a larger preview, so the detector sees a twice-resampled image where the
|
||||
larger buffer is resampled once — but that is a hypothesis, not a finding, and §12's M4 is where it
|
||||
should be settled. **The crop measurement above stands on its own and does not depend on it.**
|
||||
|
||||
---
|
||||
|
||||
## 8. Calibration — cosine to probability
|
||||
|
||||
FR-CULL-9 makes this a hard requirement: no code path may threshold a bare cosine, every threshold
|
||||
|
||||
Reference in New Issue
Block a user