327b8c7839924535cd954fefb034ebbfc6c61401
5
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0407fb8d2d |
Format the two new examples
They were written after the last fmt run and CI gates on --check. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b2250cc460 |
Measure a regroup on the tablet, not just on the desktop
The GPU question needed a number nobody had: how a regroup divides on the hardware whose CPU is weakest. dr-face carries no weights and touches no display, and dr-catalog's example needs only a catalog file, so both run under adb shell against a copy of a real library. On the same 18,143 faces — desktop against the tablet — scan 0.96s / 2.61s, agglomerate 1.69s / 2.16s, score 0.26s / 0.40s. The scan is half the pass on the tablet and under a third on the desktop, because twenty cores of AVX2 pull ahead of NEON much further than the merge engine's single-threaded hashing does. So a GPU GEMM is worth roughly 2× a regroup on the tablet and 1.5× here, and it is the tablet that should decide whether it is built. The two architectures agree exactly: the same 1,531,969 evidence pairs, the same 2,518 groups holding the same 16,246 faces, the same reliability table. That is a better check on the NEON kernel than the unit test can be. Two instruments, both read-only: the example now prints its phases, and dr-face gains scan_bench, which needs no library at all and so can answer "how fast is this machine" on a device with nothing on it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ebb7d3cf5c |
Score a suggestion against the people the user has named
The number beside a suggestion was the mean calibrated probability between the face and the rest of its group, which measures the wrong thing twice. It punishes coverage: a person with two hundred faces over fifteen years is *meant* to have members a given photograph is orthogonal to, so a correct suggestion onto a well-photographed person scored low for being well photographed. And it never asked who else the face might be — a face matching Anna at 0.95 and nobody else, and one matching Anna at 0.95 and her sister at 0.93, came out identical, when the second is the only one worth the user's attention. dr_face::assign answers both, and multiplies them: the mean of the best ten calibrated matches into the identity (the old mean, capped, which is what stops coverage counting against it), times that identity's share of the evidence against every *named* rival. Only named people compete, and per person rather than per group. Both halves of that had to be measured on a real 18,000-face library rather than reasoned about. Normalising across every group made the number useless — median suggestion 21%, four in five under half — because clustering leaves one person spread over many groups, so a face competed against itself; and keying rivals by group left Catherine competing with Catherine, median 39%. Per named person: median 99.5%. Rivals are gathered below the merge threshold, down to even odds: a named person matching at 0.6 will never be merged into but is exactly the competition to discount for. That would be a second similarity scan, the expensive half of regrouping a library, so cluster_scored scans once at the looser floor and hands the merge engine the subset at or above the threshold — pair for pair what it would have scanned for itself, held to that by a test. Leave-one-out over that library's 2,702 confirmations across 54 named people: 99.33% of faces placed on the right person against the old mean's 99.15%, and the number shown for the right person moves from a median of 90.4% to 99.3%. It errs low — 100% correct wherever it states 80% or more — which is the safe direction, and docs/faces.md §9.1 says plainly that the low bands are not calibrated. The example that measures it comes too: this is a claim about a library's numbers, and nobody should have to take it on faith. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f41b3f6e8e |
Answer "what would the other device end up with" without the other device
Build and test / Desktop (Linux) (push) Successful in 2h7m52s
Build and test / Layer separation (push) Successful in 59s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 4s
Traceability / Requirement traces (push) Successful in 37s
Build and test / Android (aarch64) (push) Successful in 22m22s
A tablet showed 200 of a person's 611 faces after syncing, and the obvious
suspects — that suggestions deliberately do not travel, that the cross-device
face match was too strict — were both wrong. Finding that out meant reading a
catalog on a release-signed Android build, which cannot be done.
So this stands the second device up locally: an empty catalog, given the images
a scan would have found, the shards adopted into it exactly as a sync does, and
the real catalog merged in as the remote. Then it counts, per person, against
what the source holds.
person source here
Catherine 611 611
Me 242 242
Ian 219 219
Which settled it: the merge carries everything, and the shortfall was transfer —
shards that never finished arriving. Worth keeping, because "did the sync lose
this or has it not got here yet" is a question that will come up again, and
guessing at it cost most of an evening.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
bb71f141e7 |
Let a folder on this machine be scanned into the catalog
`dr_catalog::scan` has known since it was written what a changed directory means — when to prune, when to list, and the one question that decides whether a deletion sweep is safe. It was fully tested and nothing called it, because walking a real directory "belongs to the platform layer" and the platform layer was eleven lines re-exporting `secrets`. So every photograph in DarkRoom arrived over WebDAV, and a user without a Nextcloud account saw nothing at all. This is the missing half: a `Storage` trait, a filesystem implementation of it, and the driver that pours one into the other. The trait is shaped by the platform it does *not* yet support. Android's SAF gives no filesystem path, which is why `SourceRef` exists; less obviously, it gives no way to *compose* one either — a document id is opaque, and the only way to learn a child's id is the children query that returned it. So a listing hands back the reference to each entry rather than a name for the caller to join onto a parent, and there is deliberately no "path + name" helper anywhere above `LocalStorage`. That single restriction is what makes SAF a second implementation rather than a second set of call sites. A reference is otherwise an opaque `(RootId, key)` pair the catalog stores verbatim and rebuilds later, which a persisted tree grant supports exactly as a relative path does. A `Path` now appears in one place: `LocalStorage::grant`, where the folder the user picked is handed in. Everything above it addresses a `RootId`. `dr_catalog::walk` is the seam. It probes a directory, asks `scan` what that means, lists only when told to, and reconciles what it found against the rows it holds. Two things it does are worth saying out loud, because both are ways to lose a library: Absence only counts where absence was observed. A listed folder proves its missing images are gone; a pruned one proves nothing about its contents, and a scan that was cancelled or that failed part-way proves nothing about folders it never reached. So the file sweep runs per listed folder, the folder sweep runs once at the end and only after a complete scan, and a root that cannot be reached at all marks its images offline and deletes nothing — FR-CAT-9's line between proven-absent and merely-unreachable, which is the difference between unplugging a drive and losing everything on it. A trashed image is absent from its folder on purpose. It is exempt from both sweeps, and detached from a folder about to be deleted rather than cascaded away with it, or a soft delete would come undone the first time the folder it came from was rescanned. Two things the tests taught, both changes to what was there before: Modification times are now milliseconds, not seconds. Change detection asks whether a timestamp moved, so the unit's granularity is the width of the window in which a change is invisible — and a second is long enough to copy a card and start a scan. The test that caught it looked like a test bug; it was not. SAF reports milliseconds natively, so this is also the unit that needs no conversion on the platform with the coarser clock. And an in-place rewrite of an existing file is invisible to directory-level pruning, because writing to a file moves neither its directory's mtime nor its entry count. That is a real limit, now documented and held by a test rather than left to be discovered. It bites less than it reads: an export, a restore, `mv`, and every editor that saves safely write beside the file and rename over it, which does move both. Narrowing the format filter no longer deletes what it stops matching, which fell out of the same principle: unticking JPEG says stop looking for new ones, not discard the hundred already rated. The files are sitting right there. `DirState` and `DirEntry` move to `dr-types`. They are the sentence the platform says to the catalog and both crates need the same one; `scan` re-exports them so nothing that used them has changed. Not done: the UI. The launch screen's "Open library" flow is account-shaped from the first field to the thumbnail worker, and giving it a local branch is its own piece of work rather than a button. `cargo run -p dr-catalog --example scan_local -- ~/Pictures` scans a real folder and reports what it cost; run it twice to see the second run list nothing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |