Specify face indexing at native resolution, and say what the proxy cost
FR-CULL-8 said detection runs against the thumbnail or proxy tier and never a full decode, and faces.md §5 said the aligned crop is sampled from that same proxy. Both are wrong in the same place: they treat detection and cropping as one resolution problem when they are two, with opposite answers. Detection does not care. §4.1 fixes the graph's input at 640x640 and letterboxes whatever arrives, so a face filling 2% of the frame reaches the model at 12px whether the buffer handed over is 1024px or 6000px. Every pixel above the detector's own input is discarded before inference. The crop cares about nothing else. §5's warp produces the fixed 112x112 ArcFace sees, so source resolution converts directly into whether those 112 pixels were photographed or interpolated. Reading crop_px across the 18,671 faces the proxy-tier implementation stored: 47.3% were upsampled to reach the embedder, 314 of them by more than 2x, the smallest from 34 source pixels. An upsampled crop does not fail loudly -- it yields a confident embedding of detail that was never there, and the damage appears three stages later as clusters that will not separate. So FR-CULL-8 now specifies four stages with the resolutions named separately: render native through FR-EXP-9's pipeline, downscale for the detector, map boxes and landmarks back to native, crop and align from the native render. The affordability the old rule bought is met instead by when the pass runs -- background, preempted, resumable -- and the requirement says plainly what it now costs on a remote library: the original rather than FR-NC-3's byte range, 412 GB across the reference library's 19,107 images, so a whole-library pass is a transfer under FR-NC-6 rather than something that may start on its own. MIN_CROP_EDGE replaces the MIN_DETECT_EDGE this branch briefly had. Same number, guarding the quantity that turned out to matter. faces.md §7b records both measurements, and marks the second as unexplained rather than dressing it as a finding. Grouped by the buffer detection ran against, faces per image was 0.078 at 1024 or below and 1.82 at 2048 or better, controlled for file type and size. That gap is real and reproducible and I cannot account for it, because the letterbox above says detector input should not matter. M4 is where it gets settled. The crop measurement does not depend on it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
+41
-6
@@ -988,19 +988,54 @@ rejects a photograph.
|
||||
producing per-face a bounding box, five-point landmarks, a detector confidence, and a 512-dimension
|
||||
embedding.
|
||||
|
||||
Detection runs against the **thumbnail or proxy tier, never a full decode** (FR-CULL-2's ladder).
|
||||
This is what makes indexing affordable: a library that has been browsed has already paid for its
|
||||
proxies, so face indexing adds no RAW decodes that were not already happening. Where no proxy
|
||||
exists, the job requests one at background priority rather than decoding inline.
|
||||
**The two resolutions are separate, and conflating them is the failure this clause exists to
|
||||
prevent.** Detection and cropping have opposite resolution needs, and a single buffer cannot serve
|
||||
both well:
|
||||
|
||||
1. **Source.** The image is rendered at **native resolution** through the full-quality path
|
||||
(FR-EXP-9's pipeline, including the high-quality demosaic of FR-RAW-3). This is the same render
|
||||
export uses and is deliberately not the FR-CULL-2 preview ladder.
|
||||
2. **Detector input.** That render is downscaled for the detector, which fixes its input at 640×640
|
||||
regardless (faces.md §4.1). Detection gains nothing from more pixels than its own input, so the
|
||||
downscale is free accuracy-wise and is what keeps the pass affordable in CPU.
|
||||
3. **Crop.** Boxes and landmarks are mapped **back to native coordinates**, and the aligned crop is
|
||||
sampled from the native render — never from the downscale the detector saw.
|
||||
4. **Embedding.** The aligned crop is warped to 112×112 in one bilinear step (faces.md §5).
|
||||
|
||||
The crop is the reason. ArcFace receives a fixed 112×112 whatever it is given, so the only question
|
||||
that matters is whether those 112 pixels are real pixels or interpolated ones. Sampling the crop
|
||||
from a preview means a face occupying a small part of the frame is *upsampled* to reach the
|
||||
embedder, and an upsampled crop yields a confident embedding of detail that was never there —
|
||||
which does not fail loudly, it degrades clustering three stages later. Measured on the reference
|
||||
library under the previous preview-tier implementation: **47% of all stored faces had been
|
||||
upsampled to reach 112×112**, with `crop_px` as low as 34.
|
||||
|
||||
This supersedes the previous rule that detection ran against the thumbnail or proxy tier and never
|
||||
a full decode. That rule was adopted for affordability and it bought exactly that, at a cost to
|
||||
crop quality that was not measured until the library was large. Affordability is now met by *when*
|
||||
the pass runs rather than by *what* it reads: it is background work, preempted by everything
|
||||
visible, and resumable per image.
|
||||
|
||||
**On a remote library this needs the original**, not FR-NC-3's byte-ranged preview — on the
|
||||
reference library, 412 GB across 19,107 images rather than a range request each. So a whole-library
|
||||
pass is a **transfer under FR-NC-6**: never automatic, subject to the unmetered-network and
|
||||
charging constraints, and reported as the download it is before it starts rather than presented as
|
||||
a local operation. An original already on the device is indexed from what is there. Nothing here
|
||||
requires the original to be *kept*: it is rendered, cropped, and given back under the same rules as
|
||||
any other borrowed file (ARCH §9.0a), so the pass costs transfer and time rather than permanent
|
||||
disk.
|
||||
|
||||
Detection is a job in the FR-CAT-3 queue and inherits its properties without exception: coalesced
|
||||
per image, interruptible, resumable across process death (FR-PLAT-AND-3), and strictly preempted by
|
||||
visible work (NFR-ARCH-2). A library indexes while idle or it does not index; it never competes with
|
||||
the grid.
|
||||
the grid. `face_index.source_edge` records the native edge each run was made at, and `faces.crop_px`
|
||||
the pixels behind each individual crop, so raising the standard
|
||||
later re-indexes only the images that stand to gain rather than all of them.
|
||||
|
||||
*Acceptance:* indexing a 10k-image library completes without the grid dropping below NFR-P9's
|
||||
interaction target at any point, and survives being killed and restarted with no repeated work
|
||||
beyond the in-flight image.
|
||||
beyond the in-flight image. No face is stored whose aligned crop was upsampled beyond a stated
|
||||
factor; the crop source resolution is recorded per face (`crop_px`) and is auditable.
|
||||
|
||||
**FR-CULL-9 — Calibrated identity.** Face similarity shall be expressed as a **calibrated
|
||||
probability that two faces are the same person**, not as a raw embedding distance. Every threshold
|
||||
|
||||
Reference in New Issue
Block a user