From 2944c1b704a37f496649c474daa94fec1d6c1e0d Mon Sep 17 00:00:00 2001 From: Duncan Tourolle Date: Wed, 26 Aug 2026 22:42:28 +0200 Subject: [PATCH] Document the run marker and cross-device face sync Co-Authored-By: Claude Opus 5 (1M context) --- docs/faces.md | 51 ++++++++++++++++++++++++++++++++++++++++++++++++--- 1 file changed, 48 insertions(+), 3 deletions(-) diff --git a/docs/faces.md b/docs/faces.md index 10e44ab..5738d5d 100644 --- a/docs/faces.md +++ b/docs/faces.md @@ -461,6 +461,36 @@ background priority and re-queues itself, exactly as FR-CULL-8 requires. --- +## 7a. Recording that detection *ran* + +The spec above assumed the `faces` table could answer "has this image been indexed". **It cannot**, +and the difference is the one that decides whether a background pass ever finishes. + +A photograph with no face in it produces no rows. So does one that has never been looked at. Asking +`faces` therefore re-queues every landscape, still life and document scan on every pass, for ever — +and in a personal library that is most of it. Measured on the reference library: of the first 110 +images indexed, **64 contain no face at all**. + +So `face_index` records the *run*: one row per `(image, model)` carrying the timestamp, the number of +faces found — zero is the interesting value — and the long edge of the proxy it read. Keyed on the +model, so a model change puts every image back in the queue without anyone having to remember to +clear anything. + +Three things fall out of it that were not otherwise available: + +- **A coverage figure.** "4,812 of 5,000 indexed" is what a user wants to see; counting face rows can + only ever report how many faces exist, which is a different number that never reaches the image + count. +- **A reason for the ones outstanding.** The audit splits them by whether a proxy exists, because + *waiting on the thumbnail sweep* and *waiting on face indexing* are different problems and only one + of them is fixed by running this again. On the reference library the first check reported 110 ready + and 23,417 awaiting a proxy — which is the real state of that library, and not something the face + subsystem can do anything about. +- **Something to sync.** §14's shards carry the marker with the faces, so an adopted image is not + re-detected on the receiving device. + +--- + ## 8. Calibration — cosine to probability FR-CULL-9 makes this a hard requirement: no code path may threshold a bare cosine, every threshold @@ -781,7 +811,9 @@ Steps 1–5 are the spike. Steps 6–8 are the build, and they are only justifie all, split off a multi-selection, regroup, and the NFR-SEC-5 delete-everything control. - **8 — not started.** The route-C first-run flow: no model fetch, no checksum pin, no licence notice. The screen says "No face model installed" and stops, which is honest but is not the - feature. + feature. The models go in `/models/` as the shape-fixed exports, by hand for now. +- **9 — done, and not previously in this plan.** The run marker (§10a), the coverage audit, the + `face_index` batch job, and cross-device sync of face shards (§14). The screen taught the design one thing worth recording. **Splitting has to reject before it confirms.** Moving faces to a new person is not enough on its own: the next clustering pass sees a @@ -808,8 +840,21 @@ just pushed away. It is user data in the same sense a confirmation is (FR-CULL-1 Every other line of this document is conditional on it. - **Faces in trashed images.** catalog.md §10.4's open question, unchanged: probably excluded from suggestions but not deleted, so a restore does not re-index. -- **Whether embeddings sync.** NFR-SEC-5 permits it, opt-in, separately consented. Nothing here - depends on the answer. +- ~~**Whether embeddings sync.**~~ **Answered: they do**, as sealed shards + (`dr_catalog::face_shard`). Indexing is hours of CPU and its result is byte-identical on every + device, so paying for it once per account rather than once per device is the whole argument. + §12.2's 3.5 images/second is also what makes the case: it is fast enough to be worth doing and slow + enough to be worth not repeating. + + **Shards rather than the catalog snapshot**, which is the design decision worth recording. The + snapshot uploads whole on every sync, and a fully indexed 23.5k library carries ~30 MB of + embeddings — exactly the cost `dr_thumbs`'s 25 MB cap exists to bound. So the split follows the one + already in the tree: bulk immutable data in sealed shards, small mutable data in the snapshot. + Faces, landmarks, embeddings and run markers shard; people, names and assignments ride the catalog + and merge by uuid. The cap is *imported* from `dr_thumbs` rather than restated, because it is a + statement about transfer cost and two copies of it would drift. + + Everything is keyed on `oc:fileid`, never `image_id`: a row id means nothing on another device. - **Re-embedding at higher resolution.** §7's `crop_px` makes it a query rather than a full re-index, but whether it is worth doing is an M4 question. - **Approximate nearest neighbours.** §9 says brute force until measured otherwise. A 100k-face