Specify face indexing at native resolution, and say what the proxy cost

FR-CULL-8 said detection runs against the thumbnail or proxy tier and
never a full decode, and faces.md §5 said the aligned crop is sampled
from that same proxy. Both are wrong in the same place: they treat
detection and cropping as one resolution problem when they are two, with
opposite answers.

Detection does not care. §4.1 fixes the graph's input at 640x640 and
letterboxes whatever arrives, so a face filling 2% of the frame reaches
the model at 12px whether the buffer handed over is 1024px or 6000px.
Every pixel above the detector's own input is discarded before inference.

The crop cares about nothing else. §5's warp produces the fixed 112x112
ArcFace sees, so source resolution converts directly into whether those
112 pixels were photographed or interpolated. Reading crop_px across the
18,671 faces the proxy-tier implementation stored: 47.3% were upsampled
to reach the embedder, 314 of them by more than 2x, the smallest from 34
source pixels. An upsampled crop does not fail loudly -- it yields a
confident embedding of detail that was never there, and the damage
appears three stages later as clusters that will not separate.

So FR-CULL-8 now specifies four stages with the resolutions named
separately: render native through FR-EXP-9's pipeline, downscale for the
detector, map boxes and landmarks back to native, crop and align from
the native render. The affordability the old rule bought is met instead
by when the pass runs -- background, preempted, resumable -- and the
requirement says plainly what it now costs on a remote library: the
original rather than FR-NC-3's byte range, 412 GB across the reference
library's 19,107 images, so a whole-library pass is a transfer under
FR-NC-6 rather than something that may start on its own.

MIN_CROP_EDGE replaces the MIN_DETECT_EDGE this branch briefly had. Same
number, guarding the quantity that turned out to matter.

faces.md §7b records both measurements, and marks the second as
unexplained rather than dressing it as a finding. Grouped by the buffer
detection ran against, faces per image was 0.078 at 1024 or below and
1.82 at 2048 or better, controlled for file type and size. That gap is
real and reproducible and I cannot account for it, because the letterbox
above says detector input should not matter. M4 is where it gets
settled. The crop measurement does not depend on it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-30 19:41:20 +02:00
co-authored by Claude Opus 5
parent 144d2e4e84
commit e7b526c550
9 changed files with 259 additions and 49 deletions
+41 -6
View File
@@ -988,19 +988,54 @@ rejects a photograph.
producing per-face a bounding box, five-point landmarks, a detector confidence, and a 512-dimension
embedding.
Detection runs against the **thumbnail or proxy tier, never a full decode** (FR-CULL-2's ladder).
This is what makes indexing affordable: a library that has been browsed has already paid for its
proxies, so face indexing adds no RAW decodes that were not already happening. Where no proxy
exists, the job requests one at background priority rather than decoding inline.
**The two resolutions are separate, and conflating them is the failure this clause exists to
prevent.** Detection and cropping have opposite resolution needs, and a single buffer cannot serve
both well:
1. **Source.** The image is rendered at **native resolution** through the full-quality path
(FR-EXP-9's pipeline, including the high-quality demosaic of FR-RAW-3). This is the same render
export uses and is deliberately not the FR-CULL-2 preview ladder.
2. **Detector input.** That render is downscaled for the detector, which fixes its input at 640×640
regardless (faces.md §4.1). Detection gains nothing from more pixels than its own input, so the
downscale is free accuracy-wise and is what keeps the pass affordable in CPU.
3. **Crop.** Boxes and landmarks are mapped **back to native coordinates**, and the aligned crop is
sampled from the native render — never from the downscale the detector saw.
4. **Embedding.** The aligned crop is warped to 112×112 in one bilinear step (faces.md §5).
The crop is the reason. ArcFace receives a fixed 112×112 whatever it is given, so the only question
that matters is whether those 112 pixels are real pixels or interpolated ones. Sampling the crop
from a preview means a face occupying a small part of the frame is *upsampled* to reach the
embedder, and an upsampled crop yields a confident embedding of detail that was never there —
which does not fail loudly, it degrades clustering three stages later. Measured on the reference
library under the previous preview-tier implementation: **47% of all stored faces had been
upsampled to reach 112×112**, with `crop_px` as low as 34.
This supersedes the previous rule that detection ran against the thumbnail or proxy tier and never
a full decode. That rule was adopted for affordability and it bought exactly that, at a cost to
crop quality that was not measured until the library was large. Affordability is now met by *when*
the pass runs rather than by *what* it reads: it is background work, preempted by everything
visible, and resumable per image.
**On a remote library this needs the original**, not FR-NC-3's byte-ranged preview — on the
reference library, 412 GB across 19,107 images rather than a range request each. So a whole-library
pass is a **transfer under FR-NC-6**: never automatic, subject to the unmetered-network and
charging constraints, and reported as the download it is before it starts rather than presented as
a local operation. An original already on the device is indexed from what is there. Nothing here
requires the original to be *kept*: it is rendered, cropped, and given back under the same rules as
any other borrowed file (ARCH §9.0a), so the pass costs transfer and time rather than permanent
disk.
Detection is a job in the FR-CAT-3 queue and inherits its properties without exception: coalesced
per image, interruptible, resumable across process death (FR-PLAT-AND-3), and strictly preempted by
visible work (NFR-ARCH-2). A library indexes while idle or it does not index; it never competes with
the grid.
the grid. `face_index.source_edge` records the native edge each run was made at, and `faces.crop_px`
the pixels behind each individual crop, so raising the standard
later re-indexes only the images that stand to gain rather than all of them.
*Acceptance:* indexing a 10k-image library completes without the grid dropping below NFR-P9's
interaction target at any point, and survives being killed and restarted with no repeated work
beyond the in-flight image.
beyond the in-flight image. No face is stored whose aligned crop was upsampled beyond a stated
factor; the crop source resolution is recorded per face (`crop_px`) and is auditable.
**FR-CULL-9 — Calibrated identity.** Face similarity shall be expressed as a **calibrated
probability that two faces are the same person**, not as a raw embedding distance. Every threshold