Commit Graph
14 Commits
Author SHA1 Message Date
dtourolle ec7a8c07ee Add a dedup_people example to measure the job on a catalog copy
`dedup_people COPY.sqlite` prints the listed people and faces before
and after, what the first run merged and kept apart and why, and the
time of three runs. The second and third runs are the cost the job
adds to every sync. `--peer PEER_COPY.sqlite` then plays two sync
round trips. The peer merges with the previous release's code path and
no job, this side merges back through sync::merge_remote, and the named
people each side lists are compared after every step.

On the reference pair: the desktop merges Claudine, Jessie x2, Mathias
and Noemi (80 -> 75 named). The tablet also merges its empty second
Ian (80 -> 74). The first run takes 1.2 s on the desktop, which is
building faces_box; the merge builds it first in practice. Later runs
take 8-15 ms.
2026-09-26 17:50:39 -04:00
dtourolle 884b681c21 Time a merge with another device's catalog in catalog_bench
The bench merged the catalog with a copy of itself. Every face in that
merge matches its own box, so the pass never reaches the faces the two
devices disagree about. On the reference library that is 631 of the
tablet's 19,052 faces, and it is the work #77 adds to.

`--remote PEER.sqlite` now also times `merge_remote_catalog` against a
copy of the peer file, and prints the first pass's report so two builds
can be checked for agreement. The first run writes what the peer
brought. The runs after it are the steady state, so compare builds from
two fresh copies of one catalog.
2026-09-26 16:54:05 -04:00
dtourolle 2226d543f9 Time a develop landing in catalog_bench
Landing on a photograph in develop opens the catalog once to fetch the
original and once more per prefetched neighbour to ask whether the cache
already holds it: five opens, each running the whole backfill. The bench
timed one open but not the landing, so the cost of the shape was not
visible and a fix to it could not be measured.

Two figures now, both against an empty cache so the question is asked the
same way whatever the answer: the five-open shape the app had, and the
two-open shape where the prefetch worker keeps one connection for its
batch. On a copy of the reference catalog (23,582 images) under load, the
five-open landing costs ~80 ms of CPU.
2026-09-26 14:25:46 -04:00
dtourolle 87badb6f99 Count the grid by subtracting the hidden burst frames, not probing per image
The grid's total is read on every scroll reload (`load_window` compares it
to notice a delete). On the reference library it cost 1.3-1.5 ms best-of-50
by catalog_bench, 2-3.6 ms on a busy machine, and the issue measured 4 ms.

`uncollapsed` asked every visible image whether a collapsed burst stands in
for it -- two primary-key probes per image, 19,000 times, on a library with
no bursts at all:

  SCAN i USING INDEX images_grid_order
  CORRELATED SCALAR SUBQUERY
    SEARCH bm USING INTEGER PRIMARY KEY (rowid=?)
    CORRELATED SCALAR SUBQUERY
      SEARCH be USING INTEGER PRIMARY KEY (rowid=?)

`total_images_filtered` now counts what the filter keeps and subtracts the
frames `bursts::collapsed_away_frames` lists, under the same filter:

  SCALAR SUBQUERY: SCAN i USING INDEX images_grid_order
  SCALAR SUBQUERY: SCAN bm; SEARCH be ...; SEARCH i USING INTEGER PRIMARY KEY

The second half walks only `burst_members`. Each image is in it at most
once (it is the key), and the filter is applied to both halves, so the
subtraction removes exactly the rows the predicate used to drop. The new
fragment sits beside `not_collapsed_away` in bursts.rs, and a test holds
the two to the same rows with bursts open and closed.

After: 0.3 ms, the same count (19,152). The cells query keeps the predicate:
it is a window with a LIMIT and needs the rows, not their number. The
rated grid count (3.5-4 ms with a one-star filter) is unchanged: its cost
is the rating subquery per image, and changing how `RatingFilter` spells
it changes every grid and timeline query, which is left for its own change.
2026-09-26 13:28:50 -04:00
dtourolle fe6e523443 Count the originals on this device from the cache, not from every image
`library::local_original_count` feeds the "On this device" chip and runs
beside the rating counts on every star keystroke. On the reference library
it cost 1.3-1.4 ms best-of-50 (3 ms on a busy machine) to find 254
originals among 19,000 visible images.

It was a correlated EXISTS per visible image:

  SCAN i USING INDEX images_grid_order
  SEARCH ic EXISTS USING INTEGER PRIMARY KEY (rowid=?)

`image_cache` holds a row only for what has been fetched, so the question
is driven from it: `i.id IN (SELECT image_id FROM image_cache WHERE
tier_actual >= Original)`, which SQLite plans as the list first and a probe
of `images` by id for each entry:

  SEARCH i USING INTEGER PRIMARY KEY (rowid=?)
  LIST SUBQUERY 1
    SCAN image_cache

`image_id` is the cache's primary key, so each image is in the list at most
once and the count is the one the EXISTS gave (254). After: 0.05 ms. On a
library whose every original is cached this is as much work as before,
which is the proportion the rule asks for.

The count stays on the keystroke path: dropping it there would leave the
chip stale after a background download until something else refreshed it,
and at this cost there is nothing left to save. catalog_bench spells the
query as dr-ui does, so its copy changes with it.
2026-09-26 13:28:50 -04:00
dtourolle 408f189019 Measure what the library screen reads on each keystroke and scroll
Issue #75 lists catalog reads paid on interactive paths rather than once:
the rating and label chip counts on every judgement keystroke, the "On this
device" count beside them, the keyword panel's vocabulary on every
selection change, and the grid's total on every scroll reload. catalog_bench
now times each of them against a real catalog and prints their answers, so
a change to any of them can be checked for giving the same numbers.

Two of them live in dr-ui's private `library` module; their SQL is spelled
in the bench as it is spelled there, which the module comment says.

Reference library (24k images), best of 50, CPU, on a loaded machine:
rating_histogram 9.0 ms, local_original_count 2.0, label_histogram 10.0,
keywords::list 4.0, keywords::for_images 5.0, grid count 1.9, grid count
with a one-star filter 4.0.
2026-09-26 13:28:50 -04:00
dtourolle 454375243c Measure what opening the catalog, a sync pass and a scan cost on a real library
Two benches for reading side by side before and after a change, against a
copy of a real catalog, in the manner of identity_bench:

- `dr-catalog --example catalog_bench CATALOG [FACES_DIR]` times
  `Catalog::open` and the backfill inside it step by step, the upload
  snapshot, a merge of the catalog with a copy of itself, and the face
  shard export and import in the steady state where nothing is new.

- `persist_bench`, an ignored test in dr-ui's scan module because
  `persist` and `apply_judgement` are private to it, replays the
  catalog's own rows through `persist` (the largest folder, and the whole
  library) and looks up every `.drsc` sidecar the catalog has read. It
  works on a scratch copy and prints a fingerprint of what `persist` left,
  so two builds can be shown to agree.

Both print best, median and CPU time; the CPU figure is the one to compare
while other builds share the machine.
2026-09-25 22:06:58 -04:00
dtourolle 84fade99ec Put the developer docs under docs/dev and index the folder for users first
docs/ had 26 developer documents flat beside the manual, and the two
audiences are very differently sized: most readers want the manual and
the gesture reference, a few want the register, the designs and the
measurements. The manual and gestures.md stay at the top; everything for
someone changing the code moves to docs/dev/, and the two documents that
name their own successors — the v0.1 milestone and the UI-refinement plan
— go to docs/dev/archive/ rather than being deleted, since both are still
cited. docs/README.md is the index, users first.

Every reference follows: code comments, Cargo manifests, the workflows,
the pre-commit hook, the bench and traceability tools (which locate the
repo root by docs/dev/requirements.md now), packaging, the Docker READMEs,
CLAUDE.md, CONTRIBUTING.md and the README. The matrix links one level
deeper and is regenerated. Links out of the moved documents into the tree
gain a level; a link checker over every Markdown file finds none broken.
2026-09-20 21:16:03 +02:00
dtourolle 8b3abdb787 Keep each face's quality, and never compare against a poor one
The embedder's raw output has a length, and the length is a reading of
how recognisable the crop was: a blur, an occlusion or a hard profile
comes out short. Normalising threw it away. A short vector sits near
the middle of the sphere and matches a little of everyone, which is how
one bad crop bridges two people in a grouping pass.

So the length is kept — the store now holds the raw vector, re-normalised
on load, with the length beside it as `faces.quality` — and a face under
MIN_GALLERY_QUALITY (14) is a probe: measured against the gallery and
placed where it fits, but never what another face is measured against.
Two probes are never paired, and a probe is nobody's evidence for a
confidence. The People screen shows the number as "Quality 17.3", dimmed
below the floor.

Faces indexed before this stored unit vectors and have no reading; they
are admitted to the gallery, and schema V14 forgets the run marker of
every image holding one so the next indexing pass measures them. A
peer's unmeasured shard faces are not adopted, or a sync would write
that marker back.
2026-09-11 21:50:12 +02:00
dtourolleandClaude Opus 5 0407fb8d2d Format the two new examples
They were written after the last fmt run and CI gates on --check.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 12:33:23 +02:00
dtourolleandClaude Opus 5 b2250cc460 Measure a regroup on the tablet, not just on the desktop
The GPU question needed a number nobody had: how a regroup divides on
the hardware whose CPU is weakest. dr-face carries no weights and
touches no display, and dr-catalog's example needs only a catalog file,
so both run under adb shell against a copy of a real library.

On the same 18,143 faces — desktop against the tablet — scan 0.96s /
2.61s, agglomerate 1.69s / 2.16s, score 0.26s / 0.40s. The scan is half
the pass on the tablet and under a third on the desktop, because twenty
cores of AVX2 pull ahead of NEON much further than the merge engine's
single-threaded hashing does. So a GPU GEMM is worth roughly 2× a
regroup on the tablet and 1.5× here, and it is the tablet that should
decide whether it is built.

The two architectures agree exactly: the same 1,531,969 evidence pairs,
the same 2,518 groups holding the same 16,246 faces, the same
reliability table. That is a better check on the NEON kernel than the
unit test can be.

Two instruments, both read-only: the example now prints its phases, and
dr-face gains scan_bench, which needs no library at all and so can
answer "how fast is this machine" on a device with nothing on it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 12:16:48 +02:00
dtourolleandClaude Opus 5 ebb7d3cf5c Score a suggestion against the people the user has named
The number beside a suggestion was the mean calibrated probability
between the face and the rest of its group, which measures the wrong
thing twice. It punishes coverage: a person with two hundred faces over
fifteen years is *meant* to have members a given photograph is
orthogonal to, so a correct suggestion onto a well-photographed person
scored low for being well photographed. And it never asked who else the
face might be — a face matching Anna at 0.95 and nobody else, and one
matching Anna at 0.95 and her sister at 0.93, came out identical, when
the second is the only one worth the user's attention.

dr_face::assign answers both, and multiplies them: the mean of the best
ten calibrated matches into the identity (the old mean, capped, which is
what stops coverage counting against it), times that identity's share of
the evidence against every *named* rival.

Only named people compete, and per person rather than per group. Both
halves of that had to be measured on a real 18,000-face library rather
than reasoned about. Normalising across every group made the number
useless — median suggestion 21%, four in five under half — because
clustering leaves one person spread over many groups, so a face competed
against itself; and keying rivals by group left Catherine competing with
Catherine, median 39%. Per named person: median 99.5%.

Rivals are gathered below the merge threshold, down to even odds: a
named person matching at 0.6 will never be merged into but is exactly
the competition to discount for. That would be a second similarity scan,
the expensive half of regrouping a library, so cluster_scored scans once
at the looser floor and hands the merge engine the subset at or above
the threshold — pair for pair what it would have scanned for itself,
held to that by a test.

Leave-one-out over that library's 2,702 confirmations across 54 named
people: 99.33% of faces placed on the right person against the old
mean's 99.15%, and the number shown for the right person moves from a
median of 90.4% to 99.3%. It errs low — 100% correct wherever it states
80% or more — which is the safe direction, and docs/faces.md §9.1 says
plainly that the low bands are not calibrated.

The example that measures it comes too: this is a claim about a
library's numbers, and nobody should have to take it on faith.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 10:44:01 +02:00
dtourolleandClaude Opus 5 f41b3f6e8e Answer "what would the other device end up with" without the other device
Build and test / Desktop (Linux) (push) Successful in 2h7m52s
Build and test / Layer separation (push) Successful in 59s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 4s
Traceability / Requirement traces (push) Successful in 37s
Build and test / Android (aarch64) (push) Successful in 22m22s
A tablet showed 200 of a person's 611 faces after syncing, and the obvious
suspects — that suggestions deliberately do not travel, that the cross-device
face match was too strict — were both wrong. Finding that out meant reading a
catalog on a release-signed Android build, which cannot be done.

So this stands the second device up locally: an empty catalog, given the images
a scan would have found, the shards adopted into it exactly as a sync does, and
the real catalog merged in as the remote. Then it counts, per person, against
what the source holds.

    person                     source     here
    Catherine                     611      611
    Me                            242      242
    Ian                           219      219

Which settled it: the merge carries everything, and the shortfall was transfer —
shards that never finished arriving. Worth keeping, because "did the sync lose
this or has it not got here yet" is a question that will come up again, and
guessing at it cost most of an evening.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 23:09:28 +02:00
dtourolleandClaude Opus 5 bb71f141e7 Let a folder on this machine be scanned into the catalog
`dr_catalog::scan` has known since it was written what a changed directory
means — when to prune, when to list, and the one question that decides whether
a deletion sweep is safe. It was fully tested and nothing called it, because
walking a real directory "belongs to the platform layer" and the platform layer
was eleven lines re-exporting `secrets`. So every photograph in DarkRoom
arrived over WebDAV, and a user without a Nextcloud account saw nothing at all.

This is the missing half: a `Storage` trait, a filesystem implementation of it,
and the driver that pours one into the other.

The trait is shaped by the platform it does *not* yet support. Android's SAF
gives no filesystem path, which is why `SourceRef` exists; less obviously, it
gives no way to *compose* one either — a document id is opaque, and the only
way to learn a child's id is the children query that returned it. So a listing
hands back the reference to each entry rather than a name for the caller to
join onto a parent, and there is deliberately no "path + name" helper anywhere
above `LocalStorage`. That single restriction is what makes SAF a second
implementation rather than a second set of call sites. A reference is otherwise
an opaque `(RootId, key)` pair the catalog stores verbatim and rebuilds later,
which a persisted tree grant supports exactly as a relative path does.

A `Path` now appears in one place: `LocalStorage::grant`, where the folder the
user picked is handed in. Everything above it addresses a `RootId`.

`dr_catalog::walk` is the seam. It probes a directory, asks `scan` what that
means, lists only when told to, and reconciles what it found against the rows
it holds. Two things it does are worth saying out loud, because both are ways
to lose a library:

Absence only counts where absence was observed. A listed folder proves its
missing images are gone; a pruned one proves nothing about its contents, and a
scan that was cancelled or that failed part-way proves nothing about folders it
never reached. So the file sweep runs per listed folder, the folder sweep runs
once at the end and only after a complete scan, and a root that cannot be
reached at all marks its images offline and deletes nothing — FR-CAT-9's line
between proven-absent and merely-unreachable, which is the difference between
unplugging a drive and losing everything on it.

A trashed image is absent from its folder on purpose. It is exempt from both
sweeps, and detached from a folder about to be deleted rather than cascaded
away with it, or a soft delete would come undone the first time the folder it
came from was rescanned.

Two things the tests taught, both changes to what was there before:

Modification times are now milliseconds, not seconds. Change detection asks
whether a timestamp moved, so the unit's granularity is the width of the window
in which a change is invisible — and a second is long enough to copy a card and
start a scan. The test that caught it looked like a test bug; it was not. SAF
reports milliseconds natively, so this is also the unit that needs no
conversion on the platform with the coarser clock.

And an in-place rewrite of an existing file is invisible to directory-level
pruning, because writing to a file moves neither its directory's mtime nor its
entry count. That is a real limit, now documented and held by a test rather
than left to be discovered. It bites less than it reads: an export, a restore,
`mv`, and every editor that saves safely write beside the file and rename over
it, which does move both.

Narrowing the format filter no longer deletes what it stops matching, which
fell out of the same principle: unticking JPEG says stop looking for new ones,
not discard the hundred already rated. The files are sitting right there.

`DirState` and `DirEntry` move to `dr-types`. They are the sentence the
platform says to the catalog and both crates need the same one; `scan`
re-exports them so nothing that used them has changed.

Not done: the UI. The launch screen's "Open library" flow is account-shaped
from the first field to the thumbnail worker, and giving it a local branch is
its own piece of work rather than a button. `cargo run -p dr-catalog --example
scan_local -- ~/Pictures` scans a real folder and reports what it cost; run it
twice to see the second run list nothing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 08:59:17 +02:00