Indexing 23,500 images is about two hours of CPU, and the result is
byte-identical on every device: the same model over the same proxy
produces the same embedding. Paying for it once per account rather than
once per device is the point.
Shards rather than the catalog snapshot, because the snapshot goes up
whole on every sync and a fully indexed library carries roughly 30 MB of
embeddings. That is exactly the cost the thumbnail store's 25 MB cap
exists to bound, so face shards use the same cap -- imported from
dr_thumbs rather than restated, since the number is a statement about
sync cost and the two must not drift apart.
The split follows the one already there: bulk immutable data in sealed
shards, small mutable data in the catalog snapshot. Faces, landmarks,
embeddings and run markers shard; people, names and assignments ride the
catalog and merge by uuid.
Keyed on oc:fileid throughout, never on image_id, because a row id means
nothing on another device.
The run marker travels with the faces it describes. Without it a
receiving device cannot tell an image with no faces from one never
examined, and would re-detect every landscape it had just adopted.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
An image with no faces in it was indistinguishable from one that had
never been looked at, so every landscape, still life and document scan in
the library was re-detected on every pass, for ever. In a real library
that is most of it: on the 23,527-image test library, 64 of the first 110
images indexed contain no face at all.
Schema v9 adds face_index, a run marker per (image, model) carrying the
face count and the proxy edge it read. Keyed on the model, so a model
change puts every image back in the queue by itself.
That makes a coverage figure possible, which is the thing a user actually
wants to see. The audit also splits the outstanding set by whether a
proxy exists, because 23,417 awaiting a proxy and 110 ready to index are
different problems, and telling the user to run indexing again would not
fix the first.
The Identity screen gains Index faces, Stop, and the coverage line.
examples/face_index.rs is the same check and sweep without a window,
which is the right shape for an overnight pass.
Measured on the real library in release: 3.5 images/second, 110 images
and 125 faces in 30 seconds, and a second run correctly finds nothing
left to do.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>