Commit Graph
100 Commits
Author SHA1 Message Date
dtourolle 702d83c218 Measure the borrow, rather than asserting it bounds the disk
The claim that hydration-as-a-borrow makes peak disk the working set
rather than the library was so far an argument. `--example vfs_cycle`
runs it: 100 photographs of 25 MB, 90 dehydrated, a pass over all of
them through the real engine.

Peak 275 MB — the resting set plus one photograph — against 2,500 MB
had the pass simply fetched everything. Back to 250 MB afterwards, and
all ten files the user already kept still there, which is the half of
the contract that matters more.

Recorded in docs/storage.md §6.3 and ARCH §9.0a, because a bound argued
from a number nobody measured is one that gets quietly lost.
2026-08-29 09:57:53 +02:00
dtourolle 5768100816 Borrow the library to index it, and give it back
The passes that need every photograph's bytes — thumbnails, face
indexing — now borrow each one and release it at the end. On a
placeholder library that is the difference between peak disk being the
working set and being the whole library.

Including on cancellation, which was nearly missed: the face sweep
returns mid-loop when the user presses Stop, and without releasing there
the disk is spent and nothing is delivered for it.

`materialise` now answers whether *it* fetched the content. The pool used
to work that out by listing a file's parent directory — one listing per
file across a library — when the backend already had to `stat` it to
decide whether to ask. One syscall instead of a directory walk, and it
removes the bug class the tests found earlier: a file at the library root
has no `parent()`, so every one of them read as already-downloaded.

**Pinning is the retention control**, and it drives the model the catalog
already had rather than a second one. `tier_desired` is what the user
asked to keep hydrated, `pending_pins` is the resumable work list, and a
pinned collection is never dehydrated for the same reason it was never
evicted. It was in fact *broken* here before: `get` on a stub failed, and
the pin worker logged "one unreadable file must not abandon the whole
pin" and silently did nothing.

Pinned originals on such a library are recorded with `path = NULL`
(`Cache::record_in_place`) rather than copied under `originals/`. Two
reasons, and the second is the important one. A copy would hold every
pinned photograph twice, with the budget able to evict the half that was
not costing the disk. And `release` deletes the file a row names — so a
row that names none cannot delete anything, which puts the one
catastrophic operation out of reach by construction rather than by
remembering not to call it. Deleting a materialised file inside a synced
tree removes the photograph from the server and every other device.

Handing disk back is `spawn_dehydrate`, which asks the client.

Two gaps written down rather than papered over (docs/storage.md §7): a
hydrating pass cannot yet quote its cost, because a stub reports no size;
and the two sweeps hold separate pools, so a library indexed for both
fetches twice.
2026-08-29 09:57:53 +02:00
dtourolle c102ba9df2 Treat a placeholder as the photograph, not as a one-byte file
The folder connector was pointed at a Nextcloud VFS tree and got three
things wrong, the first of which loses work.

**A dehydrated sidecar read as absent.** `a.drsc` does not exist when the
client has dehydrated it — only `a.drsc.nextcloud` does — so `get` missed,
`.ok()` swallowed the `NotFound`, and the sidecar writer took that for
"there is no sidecar yet" and wrote a fresh document over the existing
one. Every edit another device had put there went with it. That function's
own doc comment calls this the exact loss the format's unknown-key
preservation exists to prevent.

**A stub was catalogued as a 1-byte image**, and ARCH §9.0 measured this
machine at 121,785 placeholders against 10,267 real files — so a folder
library on a synced tree was ~92% broken rows.

**Identity changed on hydration**, so downloading a photograph looked like
a delete and an add, orphaning its thumbnail and its face rows.

Entries now carry the photograph's own name and a `materialised` flag;
`get` on a stub returns the new `RemoteError::NotMaterialised`, which is
distinct from `NotFound` precisely because the sidecar writer must treat
them differently — it fetches the sidecar and merges, or leaves the entry
queued.

Hydration is a **borrow**. `BorrowPool` records what was on disk before it
asked, so `release_all` dehydrates only what a pass brought and leaves
what the user already had. Reference counted: the thumbnail pass and the
face pass meet on the same RAW, and without counting the first to finish
dehydrates the file the second is reading. A borrow against a plain folder
or a server does nothing, so a pass written for VFS runs everywhere.

Releasing means asking the client to dehydrate and never deleting: a
deletion inside a synced tree propagates to the server and removes the
photograph from every device.

Not a second backend — the capability is per *connection*, not per type,
since the same folder hydrates only while the client runs. The convention
arrives through a detector the registry supplies, so `dr-sync-folder`
still knows nothing about any client's protocol.

ARCH §9.0a records this as an amendment: finding 3 rejected hydration
because it costs 100× a range read, and that comparison assumed a
connector was available. A folder library has none.
2026-08-29 09:57:52 +02:00
dtourolle 6c363cee97 Let the launch screen scroll, now that it offers three routes
The signed-out screen was a `VerticalLayout { alignment: center }` with
no scroll. That was already tight with two routes; the folder option
made the column taller than a 900x560 window, and a centred layout that
overflows clips at *both* ends — so the masthead and the last route
disappear together, with nothing on screen to suggest either existed.

A Flickable whose viewport follows the content, and a top padding that
centres the column only while it fits. Both cases checked against the
running app: centred at 1000x1800, top-aligned and scrollable at
900x560.
2026-08-29 09:57:52 +02:00
dtourolle 64462e9fd3 Test the folder sign-in instead of trusting it
The whole of a folder library's sign-in lived inside a Slint callback,
which cannot run without a display server — so the one path that decides
whether a mistyped folder becomes a *stored* account had no test at all.
That failure is quiet and lasting: an account for a directory that is
not there skips the launch screen on the next start and reads as a
library that has lost its photographs.

`open_folder_library` is that logic, lifted out whole. Four tests: it
stores an account with no credential and reaches the keyring for
nothing, a typo is refused before anything is written, the messages read
as instructions because they go straight to the screen's error line, and
a file is not a library.
2026-08-29 09:57:52 +02:00
dtourolle cbe5c4fcde Measure the folder walk against a real tree, not a claim
The folder connector declares `LocalEtags`, which means the engine walks
the whole library on every scan with no pruning. That is the honest
capability, and the argument for it being affordable was so far an
assertion about `stat` versus `PROPFIND`.

`--example scan` runs the real path — `dr_sync::scan` over the connector,
then a ranged read of the kind the thumbnail worker makes. Read-only; it
never writes into the folder it is pointed at.

2,299 images across 233 directories in 137 ms, and 380 across 13 in
29 ms. Against 34.1 s for 17,185 RAWs over WebDAV *with* pruning
available. Recorded in docs/storage.md §5.2 and ARCH §8.4a, because a
capability trade-off argued from a number nobody measured is the kind
that gets quietly reversed later.
2026-08-29 09:57:52 +02:00
dtourolle f12aece07e Make storage pluggable, and prove it with a folder backend
`RemoteBackend` existed from the first release and bought nothing it was
designed for. Seven files in `dr-ui` constructed a `NextcloudBackend`
directly, an account *was* a server URL beside a DAV user id, the local
cache directory was named after a hostname, and the launch screen knew
that signing in meant a browser handshake. The trait was real; the seam
was documentation.

A trait over operations is only a quarter of it. Pluggable storage needs
four things, and this adds the other three:

- **Capabilities** — already there, and the reason the engine can drive
  two backends at the speed each actually runs at.
- **Configuration** — `dr_sync::Account`: where a library lives, in
  whatever form its connector addresses, with no server in it. Loads
  every existing config unchanged (`backend` defaults to `nextcloud`,
  `endpoint` is stored under its historical `server` key), and
  `Account::namespace()` reproduces the old catalog directory byte for
  byte, because changing it would abandon a catalog, its thumbnail
  shards, and the sidecars holding unsynced offline work.
- **Registration** — `BackendProvider` and `BackendRegistry`.
  `ui/dr-ui/src/remote.rs` is now the only file above `dr-sync` that
  names a connector.

`Connection` (an account plus an optional `Secret`) replaces the
credentials-and-user-id pair that was threaded through fifteen
signatures in an order that could be swapped. `Secret`'s inner string is
reachable only through `expose()` and its `Debug` prints `Secret(***)`,
so the indirect leak — a `{:?}` on anything holding one — no longer
compiles into a leak.

Nextcloud is unchanged and keeps every peculiarity: propagating ETags,
chunked upload v2, `oc:fileid`, the `oc:permissions` probe on a refused
PUT, the 423 retry classification, Login Flow v2. Those are what the
capability model exists to serve, not something to hide.

`dr-sync-folder` is the second connector: a local disk, a network mount,
an external drive, or a folder a Nextcloud client already syncs. No
account, no credential — the route that works where no secrets daemon
does. It declares `LocalEtags` rather than claiming propagation a POSIX
directory cannot provide, which costs nothing because 50k `stat` calls
are not 50k PROPFINDs. Identity is a path hash, not an inode: an inode
survives a rename but differs between devices and is reused after a
delete, so two machines would disagree about which photograph a
thumbnail belonged to. Re-deriving a thumbnail is a cost; showing the
wrong one is a bug.

docs/storage.md is the contract — the traits, the four steps to add a
backend, and what each connector declares. ARCH §8.0 and §8.4a, and
FR-NC-13, say why.
2026-08-29 09:57:52 +02:00
dtourolleandClaude Opus 5 1b8b7998a2 Let a name hold a group together, the way a confirmation does
Build and test / Desktop (Linux) (push) Successful in 2h6m43s
Build and test / Layer separation (push) Successful in 46s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 2s
Traceability / Requirement traces (push) Successful in 29s
Build and test / Android (aarch64) (push) Successful in 21m27s
Sixteen people called Catherine, fourteen of them holding no faces at all, and
her actual photographs split across the two that did. That is not a sync fault;
it is this function, and it has been quietly doing it on every Regroup.

Naming a cluster does not confirm its faces. They stay *suggestions* — and only
confirmed faces anchored here, so the next pass cut them loose, regrouped them
into a brand new person, and left the named one holding nothing.
`prune_empty_unnamed` will not clean that up, because it has a name. Name the new
group the same thing, and it happens again. Repeat over a few sessions and you
have sixteen of her.

A name is a judgement about *this group*, of exactly the kind FR-CULL-12 says
travels and inference does not — the same argument that already anchors a group
the user set aside. So all three kinds of ruling anchor now: confirmed, ignored,
and named.

It also fixes the quieter half of the same fault. A face indexed later that
matches a named person now merges *into* them, rather than arriving as a rival
group the user has to name all over again.

Two tests, and the first fails without the change — it reports "Catherine"
finishing the pass with zero faces while a fresh unnamed person holds the two
she was named for.

This does not retro-fit an existing library: the fourteen empty Catherines stay
until they are merged by hand, and the two holding faces are separate identities
that only the user can say are one person. What it stops is making more.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 09:48:32 +02:00
dtourolleandClaude Opus 5 2fb3eb5d2d Upload a shard that was sealed after the server saw it
Build and test / Desktop (Linux) (push) Successful in 2h7m52s
Build and test / Layer separation (push) Successful in 57s
🐳 Android image / Build and push (push) Successful in 12s
Build and test / android-image (push) Successful in 13s
Traceability / Requirement traces (push) Successful in 57s
Build and test / Android (aarch64) (push) Successful in 20m58s
The tablet showed 442 of a person's 648 faces and could never catch up. Its
shard ledger said why: it had adopted the laptop's shard 0 at 2,711,552 bytes,
and that shard is 20,402,176 bytes on disk.

The upload skipped it. A sealed shard the server already had under that name was
taken to be "byte-identical by construction", so its mere presence was proof
enough and the size was never consulted. That premise is false, and this library
is the counter-example: an *open* shard is uploaded on every pass as it fills —
that is how a growing shard reaches the other devices — so the server ordinarily
holds a partial copy of a shard that is later completed and sealed. Sealing then
froze that partial copy in place for ever.

Nothing downstream could recover from it. The tablet's `has_adopted` check is
keyed on name *and* size and would have re-downloaded a changed shard gladly; it
was never offered one, because the sending side had stopped looking.

So the comparison is on size, sealed or not. The client id is in the name, so
nobody else can have written the file and size is a sound test. A sealed shard
whose size already matches is still skipped on the first comparison, which is
all the original cheapness was worth.

The thumbnail upload had the identical test and the identical hole, and is fixed
with it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 08:32:01 +02:00
dtourolleandClaude Opus 5 f9e331eed2 Let a test build be opened up, when asked
`adb shell run-as` refuses on a release build — "package not debuggable" — and
the app's private storage is then unreachable from the host. That storage is
where the face shards, the thumbnail store and the catalog live, so when a
device disagrees with the desktop about what it has synced there is no way to
find out which of them is right. An evening was spent guessing at exactly that.

`DARKROOM_DEBUGGABLE=1 ./docker/android/package.sh --install` now sets
`android:debuggable` through aapt2's `--debug-mode`, and nothing else changes.

Set through aapt2 rather than written into `AndroidManifest.xml` on purpose: the
flag then exists only for the build that asked for it, and a release build
cannot inherit it because somebody forgot to take it out again. A debuggable APK
lets any process on the device read this app's files, so it belongs on a test
tablet and nowhere else.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 08:32:01 +02:00
dtourolleandClaude Opus 5 f41b3f6e8e Answer "what would the other device end up with" without the other device
Build and test / Desktop (Linux) (push) Successful in 2h7m52s
Build and test / Layer separation (push) Successful in 59s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 4s
Traceability / Requirement traces (push) Successful in 37s
Build and test / Android (aarch64) (push) Successful in 22m22s
A tablet showed 200 of a person's 611 faces after syncing, and the obvious
suspects — that suggestions deliberately do not travel, that the cross-device
face match was too strict — were both wrong. Finding that out meant reading a
catalog on a release-signed Android build, which cannot be done.

So this stands the second device up locally: an empty catalog, given the images
a scan would have found, the shards adopted into it exactly as a sync does, and
the real catalog merged in as the remote. Then it counts, per person, against
what the source holds.

    person                     source     here
    Catherine                     611      611
    Me                            242      242
    Ian                           219      219

Which settled it: the merge carries everything, and the shortfall was transfer —
shards that never finished arriving. Worth keeping, because "did the sync lose
this or has it not got here yet" is a question that will come up again, and
guessing at it cost most of an evening.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 23:09:28 +02:00
dtourolleandClaude Opus 5 bb3b512b81 Merge: a face sync that can actually complete
Three faults between a laptop that had indexed a library and a tablet that only
ever received part of it.

The face store was the one part of the catalog still on the rollback journal at
synchronous=FULL — 21.3 ms a commit against the 0.05 ms everything else pays,
four commits per photograph, on the order of fourteen minutes of pure fsync for
ten thousand images. It now writes the way the catalog and the thumbnail store
do, with a checkpoint before upload so the file that goes to the server is
complete without its write-ahead log.

Every idle pass read 94 MB of shards to decide it had nothing to send; the name
and the size answer that.

And the 60-second total request timeout was a floor on link speed rather than a
hang detector: a 25 MB shard needed a sustained 425 KB/s or it failed, and then
retried and failed again indefinitely. It now times out on inactivity.

The merge itself was never at fault — a probe that stands up an empty catalog,
adopts the shards and merges the real one in gets all 611 of a person's faces,
not 200.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 23:08:29 +02:00
dtourolleandClaude Opus 5 d2af6a3981 Time out on a stalled transfer, not on a slow one
The HTTP client had a 60-second *total* request timeout. That is not a hang
detector; it is a floor on link speed. A face shard runs to 25 MB, so it
demanded a sustained 425 KB/s or the transfer failed — and having failed it was
retried on the next pass and failed again, for ever.

A tablet on ordinary wifi could therefore never finish taking in a library's
faces, and nothing said why: each attempt looked like a network blip rather than
an arithmetic impossibility. The catalog snapshot is 36 MB and has the same
problem.

`read_timeout` fires when no bytes arrive for the period, which is the condition
actually worth failing on. A slow transfer that is still moving now finishes,
however long it takes; a connection that has genuinely died is still caught in a
minute. The connect timeout is unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 23:08:16 +02:00
dtourolleandClaude Opus 5 d70fe9da8d Decide whether to send a shard before reading it
The upload read every shard off disk and only then asked whether it needed
sending. For a library with 94 MB of face shards already on the server, every
idle sync read 94 MB to conclude it had nothing to do.

The question does not need the bytes. A sealed shard the server already has is
byte-identical by construction, and the client id is in the name, so nobody else
could have written it — the name settles it. The open shard is compared on size,
which `stat` answers.

The progress line moves after the skip for the same reason: announced before it,
an idle pass claimed to be sending five shards and sent none.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 23:08:16 +02:00
dtourolleandClaude Opus 5 2c84aa1224 Write face shards the way the rest of the catalog writes
The face store was the one part of the catalog still on SQLite's default
rollback journal at `synchronous = FULL`. The catalog itself runs WAL at
`NORMAL` (`schema::configure`) and so does the thumbnail store; nothing decided
this one should differ, it was simply never set.

Measured on this project's own filesystem, that is **21.3 ms per commit against
0.05 ms** — four hundred times. And an export commits four times per
photograph: the shard's transaction, then three separate autocommitting writes
to the index. Ten thousand images is on the order of fourteen minutes spent
doing nothing but waiting for fsync, before a byte goes to the server. That is
the "checking faces…" that appeared to hang.

So: WAL and `synchronous = NORMAL`, matching the rest, and the three index
writes fold into one transaction. `NORMAL` is the same trade the catalog makes —
a shard is derived data, and losing the last commit to a power cut costs one
image re-exported.

WAL brings an obligation with it, because **a shard is uploaded by reading its
file**: the newest commits live in a `-wal` sidecar that no upload sends, so
without a checkpoint the server would receive a database missing exactly the
faces just written, and a peer would adopt it and see nothing wrong. `checkpoint`
folds the logs back in, with `TRUNCATE` rather than the default passive mode,
which gives up when a reader holds the log and would leave the same gap while
reporting success.

Two tests: that the store is in WAL like everything else, and — the one that
matters — that a checkpointed shard copied *without* its `-wal` still holds
every face. That second one fails without the checkpoint, which is how it was
confirmed to be testing something.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 23:08:10 +02:00
dtourolleandClaude Opus 5 0bba882fb1 Read the Android API level out of the ELF, not out of file(1)
Build and test / Desktop (Linux) (push) Successful in 2h8m13s
Build and test / Layer separation (push) Successful in 1m3s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 2s
Traceability / Requirement traces (push) Successful in 41s
Build and test / Android (aarch64) (push) Successful in 1h0m55s
Every push has failed the Android job for weeks — back through fourteen runs —
with "FAIL: linked for Android 'unknown', expected 28", on a .so that was linked
perfectly correctly at API 28 the whole time.

The check ran `file` on the linked object and pulled the level out of its
description with a regex. `file` only prints "for Android 28" when its magic
database is new enough to decode `.note.android.ident`, and this image's is not
— it stops at "dynamically linked, not stripped". The regex matched nothing, and
`${API:-unknown}` turned that silence into a confident-looking failure about the
artefact rather than about the tool inspecting it.

The API level is the first word of that note, little-endian, so it is read
straight out of the ELF with `readelf` — which is in the image via
build-essential, and cannot go out of date the way a magic database can. An
absent note is now its own message rather than being folded into the mismatch
case, since "nothing states an API level" and "states the wrong one" are
different faults.

`file` stays in the image and in the log: it names the NDK that built the
object, which is worth having when this does go wrong. Nothing depends on it.

Verified inside the real container against the linked .so: the note reads
1c 00 00 00, and the step prints "OK: linked for Android 28".

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 22:28:57 +02:00
dtourolleandClaude Opus 5 e997be8c72 Say which version this is: 0.8.0
Build and test / Desktop (Linux) (push) Successful in 2h14m44s
Build and test / Layer separation (push) Successful in 54s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
Traceability / Requirement traces (push) Successful in 1m45s
Build and test / Android (aarch64) (push) Failing after 53m50s
A hundred commits since 0.7.0, and the face subsystem in them is not a set of
fixes — it is the difference between a feature that was shipped and one that
works.

Clustering finished at all for the first time: the old agglomeration rescanned
every live pair and recomputed average link from scratch after every merge, and
on a real library it did not return. A sparse above-threshold graph, connected
components and Lance-Williams sums put 1,813 faces at 0.28s, and the merge
threshold moved to 0.80 because the old 0.90 was measured — not guessed — to
leave a third of the library ungrouped.

Faces are no longer indexed at any size or any sharpness. Both floors were
measured over the real library with `face_index --quality`: 32 source pixels
across the aligned crop, and a contrast-invariant sharpness that catches the
large-but-blurred face whose confident, wrong embedding used to weld two people
together.

The crop is cut once and kept, so the People screen is no longer a derivative of
a thumbnail cache entitled to evict anything at any moment. The screen itself
became usable: the faces wrap into a grid instead of running off the edge, the
header fits a phone, a group can be set aside, and a person's photographs are a
button away — as a union or an intersection of several people.

And face sync now reaches the other device. Re-indexed images re-export, shards
written before the crop and index-time columns are repaired rather than failing
every insert, people and the user's judgements about them cross the wire at all,
and the pass says what it is doing while it does it.

The Android versionCode follows without being restated — package.sh packs
MAJOR*10000 + MINOR*100 + PATCH, so 0.8.0 is 800, above the 700 already on
devices and therefore an upgrade rather than a refusal. `pkgrel` returns to 1,
since this is a new version rather than a rebuild of the last one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 22:24:58 +02:00
dtourolleandClaude Opus 5 a8ce1ff19c Merge: a face sync that says what it is doing
Build and test / Desktop (Linux) (push) Successful in 39m22s
Build and test / Layer separation (push) Successful in 52s
🐳 Android image / Build and push (push) Successful in 4s
Build and test / android-image (push) Successful in 5s
Traceability / Requirement traces (push) Successful in 53s
Build and test / Android (aarch64) (push) Failing after 53m19s
The face pass announced itself once and then went quiet for the whole export,
upload and adopt. On a first export after a re-index that is eight minutes of a
motionless progress bar, which reads as a hang. It now reports the images it is
preparing, the shards it is sending and their size, and the shards it is taking
in.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 21:56:22 +02:00
dtourolleandClaude Opus 5 a1e8f494b9 Say what the face sync is doing while it does it
The face pass set the status to "checking faces…" once and then said nothing
until it was finished. On a library whose first export after a re-index is 9,849
photographs and 85 MB of shards, that is eight minutes of a progress bar sitting
still — which is indistinguishable from a hang, and was reported as one twice.

Nothing was wrong with the sync. The only fault was that it was silent.

Three places now report, which are the three that take real time:

- **Preparing**, per image with a count, since this is the long one and the only
  one whose length the user cannot guess from anything on screen.
- **Sending**, per shard with its size, because a face shard carries crops and
  runs to tens of megabytes — one of them is a visible wait on any connection.
  Announced before the upload rather than after, since the wait *is* the upload.
- **Taking in** a peer's shard, which is a download and then a row-by-row merge.

The export reports every 25 images rather than every one, so the channel behind
it stays lost in the write it accompanies. `export_to_shards` keeps its old
signature and delegates, so the callers that do not want progress do not grow a
parameter for it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 21:56:16 +02:00
dtourolleandClaude Opus 5 35febf31dd Merge: a face sync that finishes
Build and test / Desktop (Linux) (push) Successful in 35m30s
Build and test / Layer separation (push) Successful in 53s
Traceability / Requirement traces (push) Successful in 38s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 2s
Build and test / Android (aarch64) (push) Failing after 58m37s
The export asked whether each image was already in the shard store at its
current index time, and answered by opening the shard database — schema batch,
pragma probes and all — once per photograph. 9,849 opens per sync pass for this
library, before any face was written, which showed up as a sync stuck on
"checking faces…" and never coming back.

The index time moves to the store's own index, which is already open, and the
write handle is held across a run instead of reopened per image.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 21:20:11 +02:00
dtourolleandClaude Opus 5 0729dfa359 Stop the face sync opening a database per photograph
"Checking faces…" never finished. `export_to_shards` asks, for every indexed
image in the library, whether the shard store already holds that image at that
index time — and `indexed_at` answered by opening the shard database, running
its six-statement schema batch and two `pragma_table_info` queries, then
querying. Once per image. 9,849 times for this library, on every sync pass,
before a single face had been written.

The index time now lives in the store's `index.sqlite` alongside the shard
number, so the question is one indexed lookup on a connection that is already
open. It stays in the shard as well — that copy is the one that travels — but
nothing reads it from there on the hot path.

The write side had the same shape: `put_image` opened the shard afresh for each
image, which mattered little when exports were a handful of new photographs and
matters a great deal now that a re-index sends thousands. The handle is kept and
reused, invalidated by shard id so sealing a full one and moving to the next
drops it without anything having to remember to.

`INDEX_SCHEMA` is `CREATE ... IF NOT EXISTS` like the shard schema, so the new
column is added on open for an index already on disk — the same trap, caught the
same way.

Two tests: that the index time survives reopening the store, and that an index
written before the column can still be opened and written to.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 21:19:56 +02:00
dtourolleandClaude Opus 5 9052b1a1de Merge: repair shards that predate the crop and index-time columns
Build and test / Desktop (Linux) (push) Successful in 33m10s
Build and test / Layer separation (push) Successful in 34s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
Traceability / Requirement traces (push) Successful in 38s
Build and test / Android (aarch64) (push) Failing after 52m58s
The schema batch is all CREATE IF NOT EXISTS, so two columns added to it never
reached any shard already on disk — and every export into one failed on "no
such column", quietly, logged at warn while the sync reported success. That,
rather than anything in the export logic, is why this library's shard stayed at
1,807 faces while the catalog reached 15,194.

Shards are now upgraded when opened, and a peer's read-only shard is read as it
stands.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 09:52:15 +02:00
dtourolleandClaude Opus 5 76eece8500 Add the columns an existing shard never got
`SHARD_SCHEMA` is entirely `CREATE ... IF NOT EXISTS`, which does exactly
nothing to a table that already exists. So `crop` and `indexed_at`, both added
to that batch, never appeared in any shard that had been written before — and
the `INSERT` naming them failed with "no such column".

Which took face export down completely, on every library that had ever synced a
face. Silently: `export_to_shards` returns the error, `sync_face_shards` logs it
at warn, and the sync goes on looking successful while the catalog fills with
faces no other device will ever see. This library's shard sat frozen at 1,807
faces with 15,194 in the catalog, and the reason was this rather than anything
in the export logic.

Shards are upgraded on open now: both columns are additive and nullable, so
catching up is one `ALTER` each. There is deliberately no version counter —
"does this column exist" is the question actually being asked, and asking it
directly cannot fall out of step the way a counter can.

A peer's shard is opened read-only and cannot be repaired, so one written before
crops is read as it stands, with a `NULL` standing in for the column. An adopted
face simply has no crop, which is the truth about it.

Four tests, built against the pre-crop schema written out in full rather than
derived from the current one — the point being that it is *not* the current
schema and must not track it. Verified against the real 1,807-face shard on this
machine: the ALTERs apply, writes succeed, and nothing already in it is lost.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 09:50:55 +02:00
dtourolleandClaude Opus 5 43097f5033 Merge: make face sync actually reach the other device
Build and test / Desktop (Linux) (push) Successful in 31m3s
Build and test / Layer separation (push) Successful in 36s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
Traceability / Requirement traces (push) Successful in 30s
Build and test / Android (aarch64) (push) Failing after 50m38s
Two independent faults, both of which had to be fixed before a tablet could show
what a laptop had indexed.

An image was exported to the face shards exactly once, so a re-index updated the
catalog and nothing else — this library went from 1,807 faces to 15,194 and kept
syncing the original 1,807.

And people never crossed a device boundary at all: the catalog merge handled
collections and keywords only, so the far end received every face and no groups,
and drew an empty People screen over a full catalog. Names, confirmations,
rejections and set-aside groups now merge by uuid, with faces matched across
devices by photograph and box overlap.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 09:27:28 +02:00
dtourolleandClaude Opus 5 4ed10f7b23 Let a person cross from one device to another
The face shards carry boxes, landmarks and embeddings. What they deliberately
do not carry is who anybody **is** — the person rows, their names, and the
assignments joining the two. Those travel in the catalog snapshot, which is a
whole-file copy and does contain them.

But the snapshot is *merged*, not adopted, and this merge only ever looked at
collections and keywords. `face_shard`'s own module note says people travel in
the snapshot; nothing implemented it. So a second device received every face and
no people at all, and drew an empty People screen over a full catalog. Exactly
what a tablet showed after syncing thousands of faces from a laptop.

What travels is what the user decided, following the rule the rest of this
module already follows — judgements travel, inference is rebuilt:

- **People**, by uuid on `revision`, exactly as a collection is: the name, and
  whether the group was set aside.
- **Confirmations**, and **rejections** — "this is not her" is a fact too, and
  is why re-clustering does not put it back.
- **The suggestions inside an ignored group**, which are otherwise ordinary
  inference but are what anchors the ignore. Without them a group set aside on
  one device reappears on the other, the same fault that made "Not interested"
  not stick locally.

Ordinary suggestions are not carried. Both devices hold the same embeddings and
clustering is deterministic, so each recomputes them and arrives at the same
answer; shipping them would double the merge for no new information.

**A face has no cross-device identity**, and unlike a collection there is no
uuid to give it one. Both devices do agree on `oc:fileid` and roughly on the
box, so a remote face is matched to the local face on the same photograph whose
box overlaps it most, above 0.5 IoU. That is not a new rule — it is the one
`record_detections` already uses to carry a confirmation across a re-index, and
it is loose on purpose: the question is "the same face in the frame", not "the
same rectangle".

A local confirmation is never overwritten. Two devices confirming one face as
different people is a real disagreement and an assignment carries no revision to
settle it with; taking the remote's answer would let a sync undo what the user
just did on the device in their hands.

The remote's schema is probed rather than assumed: `remote_is_mergeable` admits
any catalog at or below this version, so one written before faces existed, or
before V10 added `ignored`, is ordinary. An absent table skips this half instead
of aborting a merge that would otherwise have succeeded.

Nine tests, including that the name lands on the overlapping face and not its
neighbour in the same frame, that a set-aside group stays set aside, that an
ordinary suggestion does not travel, and that merging twice changes nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 09:27:17 +02:00
dtourolleandClaude Opus 5 cf614efa61 Send a re-indexed image's faces to the other devices
`export_to_shards` asked `store.contains(file_id)` and skipped anything the
shard store had already heard of. So an image was exported exactly once, and
re-indexing it updated the catalog and nothing else — every other device kept
the first answer for ever.

That is not hypothetical. This library was re-indexed after the detection floors
changed and crops were added, going from 1,807 faces to 15,194; the shard store
still held the original 1,807, written before any of it. Nothing the re-index
produced could reach another device.

The shard's `indexed` table now carries the catalog's own `indexed_at`, and the
export compares against it. A re-indexed image goes again; an unchanged one
still costs nothing. Copied from the catalog rather than stamped when the shard
is written, because a shard-local write time advances even when nothing changed
and could not answer the question.

The column is nullable so a shard written before it still reads: absent means
"cannot vouch for it", which forces one re-export and then settles.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 09:27:16 +02:00
dtourolleandClaude Opus 5 b60f9d10d9 Merge: 32-pixel faces, a header that does not overlap, and set-aside that sticks
Build and test / Desktop (Linux) (push) Successful in 30m54s
Build and test / Layer separation (push) Successful in 44s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
Traceability / Requirement traces (push) Successful in 43s
Build and test / Android (aarch64) (push) Failing after 52m25s
The size floor comes down to 32 source pixels with the blur floor moved to match
— they are coupled, since an upsampled face scores low on sharpness whatever its
original quality, and dropping one without the other would have undone itself.

The Identity header stops pinning its rows shorter than the controls in them, so
the name field no longer draws through the buttons below it.

And a group set aside now survives Regroup: its faces anchor the way
confirmations do, so they stay where the user put them instead of regrouping
into a fresh person with no ignore flag.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 08:30:32 +02:00
dtourolleandClaude Opus 5 41daa4cca7 Keep a group set aside actually set aside
"Not interested" hid a group, and the next Regroup brought it straight back.

Reclustering anchors the faces the user has ruled on so a pass cannot move them.
It took *confirmations* as the only kind of ruling — but setting a group aside is
a ruling too, and the faces it covers are only ever suggestions. So an ignored
group's faces entered clustering loose, regrouped into a fresh person that
carried no ignore flag, and reappeared in the rail. The original group was left
behind holding nothing, hidden and empty.

Anchoring them on `ignored` as well as on `confirmed` fixes it, and does one
better: a face indexed later that matches a group which was set aside now merges
*into* it, so a stranger photographed again stays set aside instead of arriving
as somebody new. That is the case that would otherwise have made the feature
feel like it only half worked.

Three tests, and the first fails without the change — it reports the group
coming back with its two faces while the original sits ignored and empty.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 08:30:16 +02:00
dtourolleandClaude Opus 5 c81d6865a7 Stop the name field running through the buttons under it
The Identity header was two rows pinned to 32px and 36px. Neither number was big
enough for what the row held: a `Field` is `Theme.touch-target` — 44px — and a
`Button` is `Theme.control-height`.

Slint honours a child's own height and lets it overflow the box the layout gave
it, so the name field drew 44px from the top of a 32px row while the button
strip began at 38px. The overlap was 6px of text box sitting on top of "Confirm
all".

Neither row states a height any more. The first takes the height of what is in
it, and the strip takes the height of a control — read from the theme rather
than from the row inside it, since `actions` sizes itself from the Flickable's
viewport and measuring it back would be a binding loop.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 08:30:15 +02:00
dtourolleandClaude Opus 5 1d4c3348af Let a 32-pixel face count, and move the blur floor with it
64 source pixels was too strict: it threw away 70% of everything the detector
finds, and plenty of what it took were faces a person could name.

Lowering it is not a one-line change, because the two floors are coupled. A face
under 112 pixels is *upsampled* to reach the embedder and upsampling invents no
edges, so a small face scores low on sharpness however crisp the original was.
Re-measured over the reference library with `face_index --quality`:

    min crop   min sharp   size cut   blur cut       kept
          32       0.000        44%         0%        56%
          32       0.002        44%         3%        53%
          32       0.005        44%         8%        48%
          32       0.010        44%        16%        40%
          32       0.020        44%        27%        30%
          64       0.020        70%         7%        23%

Holding the blur floor at 0.020 while dropping the size floor to 32 would have
rejected a further 27% — for being small rather than for being blurred — and
kept only 30%, barely more than the 23% the strict pair kept. Most of the point
of lowering the size floor would have gone straight back out through the other
gate.

0.005 removes 8% of what the size floor leaves, which is the same job 0.020 was
doing at 64 (7%): the large-but-soft face this gate exists for. Together they
now keep 48% of what the detector finds, against 23% before.

The box pre-filter follows down to 24, staying below what the real floor accepts
so it cannot reject a face that would have cleared 32.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 08:30:09 +02:00
dtourolleandClaude Opus 5 c7cdf7e5f0 Merge: face quality gates, and identity filters that combine
Build and test / Desktop (Linux) (push) Successful in 32m56s
Build and test / Layer separation (push) Successful in 50s
Traceability / Requirement traces (push) Successful in 39s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
Build and test / Android (aarch64) (push) Failing after 52m43s
Two things the People screen was missing.

Faces were being indexed at any size and any sharpness — 40 pixels on the box
and no blur gate at all — so most of what the library held was background
strangers and motion-blurred passers-by, and the blurred ones were quietly
bridging unrelated clusters. Both floors are now measured on the real library
with `face_index --quality` rather than guessed: 64 source pixels across the
aligned crop, and a contrast-invariant sharpness of 0.020.

And the grid could only ever be narrowed to one person, which cannot express
"the pictures the two of them are in together". The filter now holds a set with
a union/intersection mode, built a person at a time from the Identity screen and
taken apart chip by chip on the filter bar.

fmt, clippy -D warnings and the full workspace suite pass on the merge.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 22:23:37 +02:00
dtourolleandClaude Opus 5 af5a13b3f7 Narrow the grid to several people at once, either way
"Show photos" could only ever mean one person. The two questions a photographer
actually asks are "every picture of Anna or Bob" and "the pictures they are both
in", and the second is not reachable by any sequence of single-person filters —
no amount of switching between one person and another finds the frame they
share.

So the filter holds a *set* of people and a mode. `RatingFilter` was already the
right home, as its own doc says: every query path threads it, so the count in
the header and the cells in the grid are narrowed by the same thing, and this
composes with stars, flags and the date range for free.

The union is `EXISTS ... person_id IN (...)`. The intersection counts
**distinct** people per image and compares against the size of the selection —
one subquery rather than one per person, and it does not grow the statement with
the selection. `DISTINCT` is what makes it correct: three faces of Anna in one
frame must not satisfy a filter asking for Anna and Bob, and there is a test
that says so.

Any rather than All is the default. With one person the modes are the same
filter, and adding a second to a union can only ever show more — so a user who
has not noticed the toggle never ends up staring at an empty grid wondering what
they broke. The toggle only appears at two people, because a control that
demonstrably does nothing is a control that teaches the user to ignore it.

Building the set needs no picker of its own: the Identity screen gains "And
also…" beside "Show photos", offered only once the grid is already narrowed to
somebody. Each person is a chip on the filter bar and each chip removes just
that person, so a selection of three can be taken apart one at a time rather
than only cleared wholesale.

`RatingFilter` stops being `Copy`, since it now holds a `Vec`. Every query path
already took it by reference; the casualties were two struct updates and one
`Cell` that becomes a `RefCell`.

484 dr-ui tests pass, including the union, the intersection, that one person
reads the same in both modes, and the repeated-faces trap.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 22:17:59 +02:00
dtourolleandClaude Opus 5 a1790c3e67 Stop indexing faces too small or too blurred to be anyone
The library was storing faces at 52 source pixels and embedding whatever came
back. There was a size floor, but it was 40 pixels on the *bounding box*, and
there was no blur gate at all — so a subject walking through a half-second
exposure detected confidently, aligned cleanly, and produced a perfectly
ordinary-looking 512-vector. Nothing downstream can tell that apart from a real
face, and because blurs resemble each other more than they resemble the people
they were, they cluster together and weld unrelated identities into one group.

Two floors, both measured rather than guessed. `face_index --quality` runs the
detector over real proxies with both gates disabled and prints the distribution;
over 1,503 faces in 600 images of the reference library:

   percentile   crop px   sharpness
           1%        16      0.0006
          25%        23      0.0025
          50%        38      0.0071
          75%        76      0.0284
          99%       352      0.4282

The median face in a personal library is 38 pixels. Most of what the detector
finds is background: people across a square, a face on a poster, a stranger at
the next table. They are real detections and useless identifications.

**Size, on the crop rather than the box.** "At least 64x64" has to mean the
pixels the *embedder* sees, and the box is not that — the ArcFace template
reaches past it for forehead and chin, so the aligned crop spans roughly 1.3x
the box's shorter edge. The floor is therefore `min_source_px` on the aligned
crop, applied after the warp fixes the scale, and `min_face_px` drops to 48 as
what it always really was: a cheap pre-filter set low enough that it cannot
reject a face the real floor would have kept.

**Sharpness.** Variance of the Laplacian divided by the variance of the luma it
was taken over. The division is the part that matters: raw Laplacian variance
scales with contrast, so a threshold on it would quietly discard every backlit
portrait in the library. The ratio asks how much of the crop's variation is
edges rather than broad gradients, and is invariant to exposure.

What each pair removes, cumulatively, of everything the detector finds:

    min crop   min sharp   size cut   blur cut       kept
          64       0.000        70%         0%        30%
          64       0.010        70%         3%        27%
          64       0.020        70%         7%        23%
          80       0.010        76%         2%        21%

64 and 0.020. The size floor does most of the work, and the blur floor removing
only 7% on top of it is the point rather than a disappointment: at 64 pixels
most faces are already sharp, and what it takes out is the large-but-soft one —
precisely the face that would otherwise contribute a confident, wrong embedding.

The two gates are not independent and the doc comments say so: a face under 112
pixels was upsampled to reach the embedder, and upsampling invents no edges, so
small faces score low on sharpness even when the original was crisp. That is why
`--quality` prints them together.

**This will re-index.** Around 70% of what the current settings store falls below
the new floors — faces between 20 and 40 pixels that nobody could identify. The
People screen gets shorter and every group in it gets better.

66 dr-face tests pass, including that a blurred crop scores below a sharp one,
that halving the contrast does not move the score, and that an upsampled face
scores below the same face at full size.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 22:17:38 +02:00
dtourolle 329c388d30 Merge: each canvas overlay goes home to its own domain
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 2s
Build and test / Desktop (Linux) (push) Successful in 1h23m17s
Build and test / Layer separation (push) Successful in 41s
Traceability / Requirement traces (push) Successful in 25s
Build and test / Android (aarch64) (push) Failing after 52m22s
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

# Conflicts:
#	docs/traceability.md
2026-08-27 21:43:31 +02:00
dtourolleandClaude Opus 5 7d23fe8683 Merge: the develop view's chrome gets its own file
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 21:35:23 +02:00
dtourolleandClaude Opus 5 4fa914cbb1 Send each canvas overlay home to its own domain
The crop rectangle, the gradient handles and the repair discs were 480
lines inside `canvas-area` in app.slint, while the panels that drive them
already lived in adjust.slint, masks.slint and spots.slint. spots.slint
even opens by describing "what is drawn over the photograph, and what a
finger can take hold of" -- which was not in it.

They stayed behind because all three are positioned against `shown-*`,
the fitted image rect the develop view derives because Slint does not
report it. That is now the interface rather than the obstacle: each
overlay is *given* that rect as its own bounds, so every position inside
is a plain fraction of `root.width`, and none of them reaches out to
`canvas-area` for an origin any more.

GradientHandles joins MaskPanel in masks.slint and SpotHandles joins
SpotPanel in spots.slint, each file now holding one domain's panel and
its canvas overlay together, matching masks_ui.rs and spots_ui.rs.
CropOverlay gets crop.slint of its own rather than growing adjust.slint.

Arithmetic is unchanged: the old fraction-x subtracted shown-x from a
coordinate measured relative to canvas-area, and the new one measures
from an origin that already is shown-x. The handles stay unconditional
rather than gaining an emptiness guard, so the repeater identity that
spots_ui::sync_handles warns about is untouched.

app.slint: 2834 -> 2490 lines. Verified by running the desktop app on a
photograph, with the crop overlay's guard temporarily forced open so all
three instantiate -- no binding loop, no panic.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 21:35:08 +02:00
dtourolleandClaude Opus 5 ffc9e1aea7 Give the develop view's chrome a file of its own
StatusBar and InfoPanel sat above AppWindow in app.slint, which read as
though they were part of the application shell. They are not: neither is
instantiated anywhere but the develop view, and the shell's actual job --
choosing which of the five screens is up -- is easier to follow without
two unrelated components standing in front of it.

Moved verbatim to develop.slint, matching src/develop.rs. No behaviour
change; app.slint loses 232 lines and gains one import.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 21:33:17 +02:00
dtourolleandClaude Opus 5 cafa63ca6f Let the develop column ask how wide it needs to be
Build and test / Desktop (Linux) (push) Successful in 21m53s
Build and test / Layer separation (push) Successful in 28s
Traceability / Requirement traces (push) Successful in 32s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
Build and test / Android (aarch64) (push) Failing after 53m45s
The column was 280px, a number chosen for a tablet, with 380px bolted on
later for a desktop. Both were guesses at how much room the widest row
inside needs, and a guess is what cannot work here: the mode strip is one
chip per attribute the *operation set declares*, so the row is generated
and no constant in app.slint can track it.

When the guess came up short the failure was not a tidy clip. The
Flickable inside the column never had its `viewport-width` set, so the
viewport took its content's preferred width, and a viewport wider than
its Flickable is *centred* in it — the same rule the note on the seam's
`x: 0` already records a few lines below. So the column lost half of each
edge rather than one of them: "HISTOGRAM" read "ISTOGRAM", "Straighten"
read "aighten", Copy sat centred while Paste ran off the far side. It
looked like a rendering fault and it was an alignment one.

So the column asks instead of guessing. Every panel that can appear in it
— image, histogram, geometry, settings transfer, masks, repairs, adjust,
history — now publishes a `content-width`: how wide it has to be before
it starts clipping itself, read off its own layout rather than asserted.
Each declares that as its `min-width` too, and that is what makes the
aggregation automatic: `column` is a layout, so it already reports the
largest minimum among its children, and it does so for the panels that
come and go with the mode as well, which live inside `if`s and cannot be
named from outside. Grep `content-width` in ui/dr-ui/ui to see every
panel with a say in the answer. The mode strip is named explicitly only
because it is pinned outside that layout, so nothing else measures it.

There is no floor left. A floor is one more guess and the panels state
their own minimums now. The only thing still above the measurement is
`panel-max-width`, which is not a size but a policy — a column may not
take the window from the photograph it exists to serve — and it comes
from Rust beside `layout-class` because a width read from `root.width`
inside the layout that `root.width` depends on is a binding loop.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 21:23:44 +02:00
dtourolleandClaude Opus 5 d6380fecc8 Make the People screen a place work can be done
Five faults, all on one screen, and the Slint and Rust halves of each have to
land together.

**Regroup froze the window.** It ran inside the Slint callback, on the UI
thread. It is much faster now, but fast is not bounded — the work grows with the
library, and the one thing that must not grow with the library is how long the
window stops answering. It runs on a worker thread with an mpsc channel and a
250ms poll, like every other long pass in this module, and the button says what
it is doing instead of the window going quiet. Cancellation is dropping the
receiver. Reclustering also prunes the empty groups the previous pass left, so
pressing the button twice no longer fills the rail with "Unnamed (0 faces)".

**The faces were a single row running off the screen.** The comment on the
layout claimed to be a wrapping row; Slint has no flow layout and a
HorizontalLayout does not wrap, so a person with forty faces was a person whose
faces could not be reviewed past the fifth. It is now laid out the way the
library grid lays out thumbnails, with the same arithmetic: choose how many
columns of roughly the requested size fit, then divide the width between them so
the cells fill the row exactly and nothing overhangs.

**The header did not fit a phone.** A 240px name field beside five buttons is
wider than an Android screen — and worse than not fitting, a layout cannot be
narrower than its children's minimums, so the row reported that oversized
minimum upwards and inflated the whole screen. The faces grid is its sibling, so
it would have been measured against a width that was never on the display. The
header is now two rows, the actions sit in a Flickable that scrolls rather than
overflowing, and the rail narrows to 132px on the compact class.

**Strangers crowded out the people who matter.** Most clusters in a real library
are passers-by and other people's guests. "Not interested" sets a group aside;
the rail hides it and says how many are hidden, with one button to bring them
back. Reversible, and never a deletion — see the catalog commit for why.

**A face was a dead end.** Identifying someone and then having no way to see
their photographs is a filing cabinet with no drawer handles. "Show photos"
narrows the library grid to that person and leaves a chip on the filter bar
saying so, which is also how it is cleared. It is a term on `RatingFilter`
rather than a grid scope of its own, exactly as that struct's own doc says new
narrowing terms should be — so the count and the cells are narrowed by the same
thing, and it composes with the others for free. Suggested faces count, not only
confirmed ones, or a freshly grouped person would show an empty grid.

Crops are read from where they are now stored, falling back to cutting one out
of the proxy for faces indexed before that existed.

480 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 21:21:31 +02:00
dtourolleandClaude Opus 5 79c0520506 Keep the face, not just a way to find it again
A face was drawn by decoding the 1024px proxy it was found on and cutting the
box out again, every time the People screen opened. That made the screen a
derivative of the thumbnail cache: evict a proxy — which the cache may do at any
moment — and the cell goes blank, with no way back short of re-fetching the
original over the network and re-detecting it. It also cost a full JPEG decode
per image, per visit, to show a 96px cell.

So the crop is cut once, when the pixels are already in hand at detection time,
and kept. A 160px JPEG is a few KB against the ~250 KB proxy it replaces reading.

Where it lives is the interesting part. The catalog snapshot is uploaded *whole*
on every sync and downloaded by every device, so a crop column there would put
tens of MB on every round trip — the exact cost `face_shard`'s 25 MB cap exists
to bound, and the reason bulk per-face data lives in shards already. Crops
therefore travel in the face shards, beside the embeddings, and
`snapshot_for_upload` strips them from the copy it writes. Nothing reads a crop
out of a merged remote catalog — the merge touches collections and keywords only
— so a receiving device loses nothing. A shard carrying crops holds around 3,500
faces rather than 22,000, which is the price of a second device showing People
immediately instead of re-fetching every proxy.

The column is nullable and the reader falls back to the proxy, so a face indexed
before this still works and the next indexing pass fills it in.

V10 also adds `people.ignored`, for a person the user has looked at and does not
want to identify. Most clusters in a real library are strangers — passers-by,
other people's guests, a face on a poster — and there is no way to tell "not yet
looked at" from "looked at, don't care" without recording the second. It is a
column rather than a deletion because a deleted cluster comes straight back on
the next Regroup: the faces are still there and still similar, and nothing short
of remembering the judgement survives re-clustering. Same argument
`face_person_rejected` makes one level down.

And `prune_empty_unnamed`, for what clustering leaves behind. Regroup creates a
person per unanchored group and never removed the previous run's now-empty ones,
so pressing it twice added a rail entry per group it no longer believed in.
Named people are never touched however empty — a name is user data — nor is a
merge tombstone, which must outlive its faces to keep redirecting.

298 tests pass, including that the snapshot carries no crops while the live
catalog keeps them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 21:17:44 +02:00
dtourolleandClaude Opus 5 7275c020d7 Group people at the threshold the library actually supports
0.90 left a third of the reference library ungrouped: 1,213 of 1,813 faces in a
group, and the rest sitting alone in a screen that had nothing to offer for
them.

"Is 0.90 too tight" is not answerable from the number. It is a probability, and
which cosine it lands on depends on the calibration — so the first half of this
is a way to ask the question properly. `face_index --tune` runs the real
clusterer over the real embeddings at ten thresholds and prints what each one
produces. It writes nothing; comparing thresholds by applying them would have
each one pollute the next.

On the reference library:

      P   cosine   groups  grouped  largest
   0.95    0.449      311      62%       51
   0.90    0.403      316      67%       51
   0.85    0.374      318      70%       57
   0.80    0.353      328      74%       69
   0.75    0.335      327      77%       69
   0.70    0.319      326      79%       81
   0.50    0.267      303      85%       90

The count of *groups* is the signal, not the count of grouped faces. Loosening
from 0.95 makes it climb: real people are being assembled out of fragments. It
peaks at 0.80 and then falls — and a falling group count while the grouped faces
keep rising is the shape of over-merging, separate identities being welded
together. That is the FR-CULL-10 failure, and the one the user cannot undo by
hand.

So 0.80: the loosest setting still building people rather than melting them
together. A third more of the library gets grouped than at 0.90, and the largest
group grows by eighteen faces rather than by forty.

The table is one library, and the doc comment says so — `--tune` reruns it on
any other.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 21:17:16 +02:00
dtourolleandClaude Opus 5 c10dca984f Regroup the library without stopping the window
Pressing Regroup on a real library did not come back. Clustering 1,813 faces is
the textbook agglomeration — compute every pairwise cosine, then repeatedly scan
all live group pairs, score each with average link, and merge the best — and the
scan is inside the loop. Each merge rescans every surviving pair, and each score
is recomputed from scratch over every cross pair. Some 1.6 million pair scores
per merge, some 700 merges to do.

Three changes, none of which alter the answer.

Only above-threshold pairs can ever matter. An average that reaches the
threshold must have at least one term at or above it, so two groups with no
qualifying pair between them can never merge — not now, and not after any
sequence of merges, since merging only adds terms. The new `neighbours` module
produces exactly that sparse list: 7,875 pairs rather than 1.6 million on the
reference library. It also means the n^2 matrix is never materialised, so memory
goes from O(n^2) to O(edges) — 2.5 GB to a few hundred KB at 25,000 faces.

Merges cannot cross components, so the connected components of that graph are
independent problems: four hundred small agglomerations instead of one large one.

Average link is additive — sum(A u B, C) = sum(A, C) + sum(B, C) — so a merged
group's scores follow by addition. Kept as running (sum, count) per adjacent
pair, a score costs one division instead of a nested loop, and a heap with lazy
invalidation replaces the rescan.

Measured on the reference library: 0.28s, release, for all 1,813 faces.

An exact ANN index was tried and removed, and neighbours.rs records why so it is
not rediscovered as a good idea. IVF with a triangle-inequality bound is exact
and prunes beautifully on synthetic clusters; on real embeddings it prunes
*nothing* — 946 of 946 cell pairs survive. Median pair angle is 88.5 degrees and
the merge threshold is 66.2, so the bound needs cells of radius under ~10
degrees, but two photographs of the same person sit 36-60 degrees apart. No
ball-based partition of a 512-d near-orthogonal space can be tight enough. So
the scan stayed exhaustive and got an unrolled dot product and its blocks spread
across cores instead.

Correctness is held by keeping the old implementation as an oracle: three tests
run both engines over the same population — plain, under co-occurrence and
anchor constraints, and with a size-weighted calibration — and assert the
clusters are identical. Determinism is asserted at a size where the threaded
path is in play.

62 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 21:16:55 +02:00
dtourolleandClaude Opus 5 7dfbe3184a Merge master into the face branch
Build and test / Desktop (Linux) (push) Successful in 21m37s
Build and test / Layer separation (push) Successful in 28s
Traceability / Requirement traces (push) Successful in 25s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Failing after 44m27s
Master moved 24 commits while this branch was building the face sweep, and one
of them changed how the interface reaches a server: `remote.rs` is now the only
file that names a connector, and everything else takes `&dyn RemoteBackend`.
The face sweep was written before that landed and built a `NextcloudBackend`
directly. Git merged the text without complaint and the result did not compile,
which is the useful kind of conflict — it goes through `remote::connect` now,
like every other pass.

Two documentation conflicts, both resolved toward master. `code-health.md` was
an add/add: master's copy carries the CH resolution for the backend seam and a
better provenance note, so it wins outright, with its measured figures re-taken
against the merged tree rather than either side's — `run()` is 1,855 lines now,
2,042 tests, 793 traceability tags. `traceability.md` is generated, so it was
regenerated rather than hand-merged.

The seam grades in code-health.md are unchanged by this merge. That is worth
noticing rather than glossing: the face work went into the seams that already
existed — `run()`, `library.rs`, `AppWindow` — which is exactly the pressure
CH-1 describes rather than evidence against it.

All four CI jobs pass: desktop (fmt, clippy, test, build), layering, traceability,
and the Android cross-build.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 20:27:45 +02:00
dtourolle 28c046130c Merge branch 'master' into android-bundled-face-models 2026-08-27 20:22:42 +02:00
dtourolleandClaude Opus 5 7b4263ceb9 Bring master's display and parity work under the new checks
Build and test / Desktop (Linux) (push) Successful in 21m3s
Build and test / Layer separation (push) Successful in 38s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 2s
Traceability / Requirement traces (push) Successful in 34s
Build and test / Android (aarch64) (push) Failing after 33m40s
Master moved sixteen commits while these fixes were being written — per-display
colour, the frame-budget measurement that decides FR-DSP-2, and a declared
node running without being compiled. Merged here rather than on master so the
conflicts are resolved where they can be tested.

Two files overlapped and neither was interesting. `lib.rs` gained `mod remote`
from this branch and `mod display_ui` from master, which git resolved on its
own. `docs/traceability.md` is generated, so it was regenerated from the merged
tree rather than hand-resolved — hand-editing a generated matrix produces one
that agrees with neither side. Coverage reads 59.9% (106/177), up from 55.4%,
entirely from master's tagging.

The check worth having run is `the_interface_names_no_operation` against
master's new `display_ui.rs` and its 195 changed lines of `develop.rs`: a new
UI module written without knowledge of this gate passes it. That is the
evidence the gate is not merely satisfiable by the code that shipped with it.

fmt clean, clippy clean at -D warnings, 2087 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 20:10:56 +02:00
dtourolleandClaude Opus 5 371038614f Fetch the repairs first, or they never happen
The repair added in 051447b was appended to the work list:

    wanted.extend(repair);

Behind every un-indexed image in the library. On this library that is position
23,000-odd — roughly two hours of fetching before the first repair is reached,
which inside one session is indistinguishable from the feature not existing.
The log said it had found them and the screen stayed empty, which is the worst
combination of the two.

They go first now. There are a few hundred of them against tens of thousands of
un-indexed images, and they are precisely the images the People screen is
failing to draw at this moment — so the ordering costs nothing and is the
difference between the grid filling in within a minute and not filling in at
all.

A failure to build the un-indexed half no longer discards the repairs either:
the pass runs with whatever it has rather than returning empty.

Verified on the live catalog: 455 images hold faces, 251 already have a proxy
from the fixed sweep, and the remaining 204 are what now sits at the front of
the queue. The stored proxies check out — a valid JPEG at the large class,
49 KB — so the storing half of 051447b was already working.

471 tests pass, including one that the orphan is ordered ahead of the library.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 19:58:32 +02:00
dtourolleandClaude Opus 5 242374fd0f Let the interface hold a backend without knowing whose it is
`dr-sync` defines `RemoteBackend` and a capability model the engine adapts
to, so a second backend can be added without touching the code that uses
one. That boundary was documentation. Seven files in `dr-ui` constructed a
`NextcloudBackend` directly, ten functions took one by concrete type, and
exactly two call sites in the tree — both inside `dr-sync` itself — ever held
the trait object. A WebDAV or local-folder backend would have had a
well-written trait to implement and nowhere to go afterwards.

The change is smaller than the finding suggests, because the trait was
already right. Every method the UI has ever called on a backend — `get`,
`put`, `list`, `delete`, `create_dir`, `move_to` — was already on it, so
nothing had to be added and no behaviour moved. Ten signatures widened to
`&dyn RemoteBackend`, sixteen constructions became `remote::connect`, and
`remote.rs` is now the only file in the interface that names a connector.

`connect` returns `Result<Box<dyn RemoteBackend>, RemoteError>`. The error
type is `dr-sync`'s rather than the connector's, which is why every call site
kept its shape — the `match`, the `let Ok(..) else`, and
`.map_err(ScanFailure::local)?` all still read as they did.

One wrinkle worth recording: `&Box<dyn Trait>` does not reach `&dyn Trait` on
its own. The compiler reaches for unsizing, which wants
`Box<dyn RemoteBackend>: RemoteBackend`, and reports a confusing missing impl
rather than suggesting a deref. Twelve call sites therefore say `&*backend`,
and two say `let backend: &dyn RemoteBackend = &*backend` where a borrow is
shared across lanes.

What this does *not* do is abstract credentials. `AppCredentials` is an app
password from Login Flow v2 — a Nextcloud protocol, not a general notion of
authenticating to a remote — and seven files still name it. An OAuth token, a
bucket key pair and an app password have no useful common shape, so deciding
what an account is across backends before a second one exists would be a
confident guess. code-health.md CH-2 now records that as the remaining half,
and it should wait for the backend that forces it.

Verified: fmt clean, clippy clean at -D warnings, 2041 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 19:51:23 +02:00
dtourolleandClaude Opus 5 051447bda6 Keep the proxy a found face will be cropped from
Every cell on the People screen read "no preview" while the sweep was happily
reporting 893 faces found. Both were true. Faces were being detected and stored
correctly; there was simply nothing left to draw them from.

A face is stored normalised and drawn by cropping the proxy it was found on —
`identity::decode_proxy` reads `FACE_TIER` out of the thumbnail store. The
fetching sweep fetched a preview, detected on it, wrote the faces and dropped
the pixels. So every face it found pointed at a proxy that had never been
stored, and the grid had nothing to cut.

Worse, that state could not repair itself: the image has its `face_index` row,
so it is not outstanding work and no later pass would look at it again.

Two fixes, and the first is nearly free. The sweep now keeps the proxy — it has
already paid the round trip and the decode, and the crop needs those same pixels
the moment the user opens the person. Kept **only where a face was found**:
two thirds of a personal library is landscapes and documents (docs/faces.md
§7a), those will never be cropped, and skipping them keeps this well clear of
the whole-library cost `SWEEP_THUMB_SIZE` deliberately avoids. The downscale to
the large class happens after detection, which is the last use of the full
buffer.

Second, the sweep now picks up images that have faces with no proxy, whatever
put them in that state — this bug, or an ordinary cache eviction, which would
have produced exactly the same empty grid. Re-running detection repairs it and
loses nothing: `record_detections` replaces rather than appends and carries the
user's confirmations across the replacement. That makes the screen
self-healing rather than dependent on nobody ever evicting a thumbnail.

The proxy is stored *before* the detections. A kill between the two then leaves
a proxy with no faces — which the next pass simply re-indexes — rather than
faces with no proxy, which is the state that cannot recover.

Note for the library already part way through a sweep: the 986 images indexed
before this will be picked up by the repair route on the next run.

470 tests pass, including one that a face whose proxy is gone becomes work again.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 19:35:12 +02:00
dtourolleandClaude Opus 5 617262b4da Give a contributor a way in
There was none. 14 documents, 177 numbered requirements, and all of it
written for someone who has already decided to work on this — nothing that
tells a newcomer which door is unlocked, what a first build needs, or why it
takes so long.

CONTRIBUTING.md points the first door at the operation format, because
"add a develop node" is a genuinely one-file contribution and the best first
experience this codebase can offer: no Rust, no shader edit, no UI change,
and its tests declared in the same file. It teaches the architecture's
central idea on the way through, which is why the invariant test added in
the previous commit is named there rather than left to be discovered.

Three things that were folklore are now written down: Git LFS is a
prerequisite, the first build resolves 826 crates and is not hanging, and
Slint needs pkg-config, libfontconfig1-dev and libxkbcommon-dev. The LFS one
proved itself while writing this — a fresh worktree hit exactly the failure
`dr-segment`'s build script is written to catch, which is the argument for
saying so before it happens rather than after.

rust-toolchain.toml pins 1.92.0 because `build-and-test.yml` already does
and says why: a floating toolchain turns an unrelated push into a mystery
failure. The two checks that gate every push are the two most sensitive to
compiler version — rustfmt's output changes between releases, so a
contributor on a newer stable can produce a diff nobody wrote on a line
nobody touched, and `clippy -D warnings` is the same story with new lints.
`rust-version = "1.92"` in the manifest stays where it is; it is a minimum,
and this is the upper bound it cannot express.

Also states the convention the tooling cannot enforce, from code-health.md
CH-4: close a requirement with a test that would fail if the behaviour were
removed. Coverage that moves slowly and means something beats coverage that
moves quickly.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 19:33:37 +02:00
dtourolleandClaude Opus 5 5e4c7424ed Fail the build when the interface names an operation
FR-DEV-3a is the property the declarative pipeline rests on: a node in
`ops/` is one file because nothing in `ui/` has to learn about it. It was
true — all fifteen ids grepped across `ui/` yield one hit, a localisation
test — and it was held by discipline alone.

That is the wrong mechanism for it. The failure is silent and cumulative:
special-casing one operation to fix a layout problem is defensible on its
own, and by the fifth the panel names half the chain and "a new operation is
one file" has stopped being true without any single commit having broken it.
Nothing would have told us.

The ids are read from `ops/*.yaml` rather than listed, so a node added
tomorrow is covered without anyone remembering this file — the same reason
`traceability` parses its denominators from `requirements.md` at run time.

Two decisions worth recording, because both are the difference between a
test that holds and one that gets deleted:

**Only string literals count.** `texture`, `contrast` and `clarity` are also
ordinary graphics and English terms, and `texture` appears throughout `dr-ui`
meaning a GPU texture. Matching bare words would fail constantly for reasons
unrelated to the invariant.

**`#[cfg(test)]` items are exempt, and finding them needs more care than it
looks.** The first version cut each file at the first textual match of
`#[cfg(test)]`, which in `develop.rs` is a *doc comment discussing the
attribute* at line 969 — it read 18% of the most important file in the scan
and passed. It now matches the attribute only as a whole line and skips the
item by brace depth, and `MIN_SHIPPING_FRACTION` fails the test outright if
the scan ever swallows the file again. The dangerous failure here is not a
false alarm, which someone investigates; it is examining nothing and
reporting success.

Verified both ways: it passes on the tree, and an `"exposure"` planted at
develop.rs:3635 — past two `#[cfg(test)]` attributes, exactly where the
first version was blind — fails with the file, the line and the reason.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 19:33:20 +02:00
dtourolleandClaude Opus 5 2c56729354 Read a descriptor's variants without taking them
Build and test / Desktop (Linux) (push) Successful in 21m4s
Build and test / Layer separation (push) Successful in 27s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 2s
Traceability / Requirement traces (push) Successful in 24s
Build and test / Android (aarch64) (push) Failing after 33m42s
Two agents worked in parallel and neither could see this. The frame-budget
instrument matches `ParamKind::Enum { variants }` by value, which was free when
a descriptor was `&'static` and everything in it was borrowed for the life of
the program. Descriptors are owned now — a declaration parsed at run time
cannot hand out a `&'static` — so `variants` is a `Vec` and the arm was moving
out of a shared reference.

Bound by reference instead. The arm only ever reads the length.

The kind of conflict that survives a clean textual merge: git had nothing to
report, and the two changes are only incompatible once they are in the same
tree.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 19:25:15 +02:00
dtourolle fb1b44ce47 Merge branch 'worktree-agent-a1e5c8cb565255f5b' into master
# Conflicts:
#	docs/traceability.md
2026-08-27 19:21:10 +02:00
dtourolleandClaude Opus 5 8202c05d9d Format the zoom test the way the gate asks for it
Build and test / Desktop (Linux) (push) Successful in 1h22m24s
Build and test / Layer separation (push) Successful in 2m57s
Traceability / Requirement traces (push) Successful in 1m3s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 2s
Build and test / Android (aarch64) (push) Failing after 33m41s
Whitespace only. `cargo fmt --check` is a required step and the FR-DSP-5 test
arrived disagreeing with it — kept as its own commit so it can be skipped
wholesale rather than read for a change that matters.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 19:20:23 +02:00
dtourolleandClaude Opus 5 29da637fa2 Record what a contribution costs, and what would lower it
An audit of code quality and extensibility, written because the answer to
"how hard is it to add a feature here?" splits cleanly in two and the split
is not where it looks.

The develop pipeline is genuinely open: a new operation is one file in
`ops/`, and the claim that no code in `ui/` names an operation turns out to
be true — all fifteen ids grepped across every .rs and .slint file in `ui/`
yield one hit, a localisation test. `dr-ui` is the opposite: every feature
lands in `run()`, one of the `wire()` functions, and a root component
carrying 472 members, so contributors collide by construction.

Five findings, CH-1 to CH-5, each with a falsifiable "done when" in the
idiom technical-debt.md already uses. The distinction from that document is
deliberate and stated: it records compromises that were chosen and are
load-bearing until their criterion is met; this records friction nobody
chose. A section listing what must NOT be tidied comes before the findings
for the same reason.

CH-1 is not a new diagnosis. view-composition.md specified the fix on
2026-08-09, when `run()` was 500 lines; it is 1,810 now, the two view
booleans it described are five, and the conditional chains it predicted in
app.slint are five-term conjunctions. The entry references that spec rather
than restating it, and quantifies the cost of the delay.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 19:18:33 +02:00
dtourolleandClaude Opus 5 0cd3ef1b3f Run a node declaration without compiling it
`ops/*.yaml` plus `build.rs` has been the class-1 plugin format since the
declarative nodes landed — it was simply resolved at build time. Nothing about
a declaration requires the compiler: everything it produces is data plus a
WGSL string, and the composer already assembles WGSL at run time from whatever
operations are active. So this is not a new mechanism. It is the existing one,
loaded later (FR-PLG-2).

`DeclaredOp` implements `Operation` from an owned `Declaration` — one
interpreter over many declarations, where `build.rs` emits generated code per
node. The generated path stays, as FR-PLG-2 says it should: a generated
`match` is faster than an interpreted one, the built-ins' declared `tests:`
have to run under `cargo test`, and generated source is inspectable in a way
an interpreter's state is not.

**The reader is now one file, read by both.** `src/declared/decl.rs` and
`src/declared/expr.rs` are `#[path]`-included by `build.rs` as well as being
modules of the crate, and they produce a neutral `Declaration` that names no
Rust type. The build script's job is reduced to *rendering* that declaration
as Rust; `DeclaredOp` converts the same declaration into descriptors and
`Expr::eval` walks the same tree the renderer writes out. There is one
grammar, one set of validations and one set of error messages, so "a plugin is
the same kind of thing as a built-in" is structural rather than aspirational.

What remains genuinely written twice is the pair of backends — an arithmetic
node rendered as Rust here and evaluated there — and that is what the parity
test stands between. `tests/declared_parity.rs` parses every built-in
declaration at run time and asserts the composed WGSL is byte-for-byte what
the generated implementation produces, with the uniform block bit-for-bit
identical, at both ends of every parameter's range and at four interior
points; then again over the whole develop chain with the declared nodes
swapped in, which is what covers uniform slot ordering and helper
de-duplication between operations. A third test asserts the declared and
hand-written nodes partition `ops/` between them, so coverage cannot shrink
silently.

Bit-for-bit rather than within a tolerance, because a tolerance is where a
real divergence hides. The one thing that had to be got right for that to hold
is number literals: `expr::as_f32` rounds a decimal exactly once, through the
same shortest-round-trip text the compiler is handed, rather than rounding an
`f64` a second time.

Not in scope, and deliberately untagged: load-time WGSL validation
(FR-PLG-11), id namespacing, a plugin directory read at startup, and pass
nodes (FR-PLG-2a). Those are separate work, and tagging them from here would
be the overstatement the spec's own §7 warns about.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 19:08:34 +02:00
dtourolle 3b5bba62f1 Merge branch 'worktree-agent-aa9f4356c13893373' into master
# Conflicts:
#	docs/traceability.md
2026-08-27 19:08:32 +02:00
dtourolle 510d1a26cb Regenerate the traceability matrix
The platform layer was never scanned, so every FR-PLAT-* and NFR-PORT-* tag in
dr-plat was invisible. Coverage 51.4% -> 57.6%, almost all of it pre-existing
tags that were simply not being counted.
2026-08-27 19:08:23 +02:00
dtourolleandClaude Opus 5 2e6825ded0 Record what the measurement found, where the next person will look for it
Three places, because the finding has three audiences.

The display spec's §1 table said FR-DSP-3 was unmeasured and FR-DSP-5 untagged.
Both are now false, and §2's decision rule has fired. The body of §2 is left as
written with the verdict quoted above it: a plan overtaken by its own evidence
reads better in order than quietly edited into agreement with the outcome.

TD-4 is the stage that misses the budget. Clarity's kernel is a fraction of the
frame, so it reaches a 52-pixel radius at 4K and costs 34 ms — seven times the
entire fused chain, for one slider. It is debt rather than a bug because the
detail stage cannot yet write a target smaller than it reads, which
`local_contrast`'s own documentation has said since it was written. The entry
says plainly that tiles are the wrong tool for it, since that is exactly the
conclusion a reader arriving from ARCH §5.3 would otherwise draw.

TD-5 is the one nobody was looking for: composing the fused shader costs
2.8–5.2 ms of CPU per frame on a full chain, on the UI thread, which at
1920x1200 is more than the dispatch it precedes. The source depends only on the
graph's structure — what `structure_hash` already identifies and what does not
move during a drag — so the fix is the cache `AdjustPass` already keeps for
compiled pipelines, one level up.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 19:06:19 +02:00
dtourolleandClaude Opus 5 772a69711d Prove that zooming to 1:1 reads the source, and only then tag FR-DSP-5
FR-DSP-5 has been satisfied for some time and untagged. `Framing::view` shrinks
the sampled region while the render target keeps its size, so a zoom raises the
resolution the pipeline works at rather than magnifying pixels already drawn —
there is no second full-resolution path because the zoom is that path.

Tagging it on that basis alone is what §7 of the display spec warns against:
traceability counts a requirement as covered when a comment names it, and
checks nothing about the code under the tag. So the tag goes on tests instead,
and the tests are built so that removing the behaviour breaks them. Both
failure modes were checked by hand: deleting the view from `visible_rect`
leaves the 1:1 render flat, and dropping only its offset leaves the render
exactly inverted. The assertion message names both, since those are the two
ways this can go wrong and the numbers alone do not say which.

The fixture is one-pixel black-and-white stripes — the highest frequency an
image can hold, and precisely what a proxy discards. A 1024 px source in a
128 px viewport reads source column `8x + 4` for every output column `x`, all
the same parity, so the fit render comes out uniform; that is asserted first,
because a 1:1 render showing detail proves nothing unless the proxy is known to
carry none. What remains is an equality against the source bytes rather than a
claim that something looks sharper.

The third test takes the arbitrary zoom the requirement also names, and pins
`RenderScale` beside the pixels: a zoom that moved the pixels but not the scale
would sharpen at the wrong radius, which stays invisible until somebody
compares a preview against an export.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 19:05:05 +02:00
dtourolleandClaude Opus 5 b858fc029a Record what per-display colour actually cost
The status table said FR-DSP-8 was absent and §5.2 assumed Slint
reports window moves. It does not — there is no `on_moved` on any
backend — so the position is sampled instead.

§5.4 records the two trades that are worth someone finding later: a
display profile is matched to the nearest of four spaces rather than
applied through a CMM, and on Wayland the canvas follows the first
output rather than the window, because a Wayland client is never told
where its window is and the protocol's own answer needs a `wl_surface`
that Slint does not expose.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 19:03:24 +02:00
dtourolleandClaude Opus 5 0c3b8cb1c4 Hand out descriptors a declaration could produce
`Operation::descriptor()` returned `&'static OpDescriptor`, and that lifetime
is the whole reason a build-time node is free and a run-time node is
impossible: only a compile-time literal can satisfy it, so no amount of
reading `ops/*.yaml` at startup could ever produce a descriptor the rest of
the application would accept. FR-PLG-2 says a bundled operation and a
third-party plugin are the same kind of thing, differing only in where the
file was found — and a lifetime outsiders cannot meet is exactly the second,
weaker format that requirement forbids.

So a descriptor is now owned and handed out as `Arc<OpDescriptor>`, with `Vec`
where it held `&'static` slices. `Arc` rather than a `&self`-borrowed
reference because the callers want to *keep* it: the develop panel collects
descriptors and then mutates the graph, and a borrow would tie the
descriptor's lifetime to a borrow of the operation it came from, which is the
one thing `&'static` was doing right.

The identifier newtypes deliberately did not follow. `ParamId` is `Copy`, is
compared in `match` arms against generated constants, is a map key in the
sidecar and history, and reaches Slint model rows; an `Arc<str>` there would
cost a refcount on every one of those and would take `match id { EXPOSURE =>
.. }` away from the generated code. They gain an interner instead, which is
honest about its lifetime rather than pretending to one — the set of ids is
bounded by deduplication and is process-lifetime by construction, because the
sidecar on disk names its parameters and an id has to stay resolvable for as
long as any edit naming it can be opened.

No behaviour changes. Every descriptor that was a `static` is a `LazyLock`
initialiser now, `Operation::helpers` borrows from `self` instead of being
`'static` so a future run-time node can own its list, and `Warp` and `Framing`
follow `Operation` so there is one shape rather than two.

The one place a descriptor is read per frame is `compose_full`, which takes
`descriptor().id` to prefix each active operation's uniforms, and `dr-ui`
composes on every frame it draws. That is a dozen atomic increments beside a
composition that is already building several kilobytes of WGSL on the same
call; it is noted at the trait method rather than left for a profiler to find.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 19:02:25 +02:00
dtourolleandClaude Opus 5 13deaa2fbb Assert the frame budget, and commit the numbers behind the FR-DSP-2 verdict
FR-DSP-3 states a latency requirement and nothing checked it, which makes it a
wish. This adds the check and the measurements it guards.

`docs/frame-budget.md` is the bench's output with the reading of §2's decision
rule attached. The short version: every point-operation chain at every viewport
size, fit and at 1:1, is inside 16 ms at the 99th percentile — the widest is
4.5 ms of GPU at 4K — so FR-DSP-2 should be rewritten rather than implemented.
The measurement did find a stage that misses the budget, and it is the one §2
predicted: clarity's 52-pixel separable kernel costs 34 ms at 4K. Tiles make
that worse rather than better, since a tiled convolution reads a halo per tile;
the fix `local_contrast` already names for itself is a base computed at reduced
resolution.

The test guards the fused path and says so, at length, rather than quietly
excluding the expensive stage and letting the tag imply otherwise (§7). What it
asserts is exactly the claim the recommendation rests on: one dispatch over a
viewport-sized target, at a full chain, is comfortably inside a frame.

Two things the numbers forced:

- The two cases are one `#[test]`. As two they ran on a thread each, contended
  for the same device, and took the 1:1 case from 2.5 ms to 14.9 ms — a
  measurement of the harness that would have flickered either side of the
  budget forever.

- The CPU half of the frame is judged only in an optimised build. Composition
  is real per-frame work on the UI thread and belongs in the budget, but the
  workspace builds its own crates at `opt-level = 0` in dev and `cargo test` is
  a dev build, so measuring it there measures rustc. The GPU half is asserted
  either way.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 19:00:45 +02:00
dtourolleandClaude Opus 5 3e5840e413 Let a test hold the words the About page shows
FR-DSP-8 asks for a fallback that is *defined*, and a reason that
reaches the code but never the screen satisfies half of it. The three
readouts are now built by a free function over the survey rather than
written straight into the window, so a test can assert that "sRGB
assumed" arrives with the reason attached, that an approximated profile
says "nearest to" rather than claiming the space, and that the second
monitor is described on the page read from the first — which is the
display the requirement is actually about.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 19:00:43 +02:00
dtourolleandClaude Opus 5 c8c6368542 Index the whole library by fetching what it has not seen
"Index faces in the whole library" could not. Its work list was intersected
with the thumbnail store at `ThumbSize::Large`, and nothing fills that class for
a whole library — `SWEEP_THUMB_SIZE` is deliberately `Grid`, because the large
class is ~860 MB of shards against ~200 MB and every syncing device pays it. So
the only images with a large proxy were the ones the user had personally zoomed
into or opened in the loupe. On this library that was 220 of 23,529.

The comment defending it misread the requirement:

    // Requesting one here would put face indexing on the network path,
    // which FR-CULL-8 explicitly keeps it off.

FR-CULL-8 keeps indexing off the **full decode**, not the network, and then says
the opposite in the same paragraph: "where no proxy exists, the job requests one
at background priority rather than decoding inline". faces.md §7 repeats it.
Neither was implemented.

So the pass fetches. Same two-stage route the thumbnail sweep uses — the header,
then the located preview's own byte range (FR-NC-3) — so no whole file is pulled
and no RAW is decoded, because an embedded preview is a JPEG. The work list is
now every visible image with no `face_index` row for the model: 23,308 here,
against nearly none before.

**It indexes at the resolution the preview actually has**, not the 1024 the old
tier would have given. `locate_preview` already picks the largest embedded
preview, and the thumbnail sweep was decoding it and throwing the detail away at
`downscale_to(256)`. A face 2% across the frame is 5 px on a grid thumbnail and
~61 px at the cap here — and 112 is what the embedder samples, so this is the
difference between an upsampled crop and a real one. `crop_px` records which,
per face, as §7 intended.

Capped at 3072 rather than truly full: `index_proxy` needs packed `f32` RGB at
12 bytes a pixel, so a 24 MP frame is ~288 MB and the fetch lanes hold one each.
The constant is named and sits next to the reason.

Orientation is applied **before** detection, not after downscaling. That costs a
permutation of a larger buffer — ~15 ms against a ~150 ms decode — and buys the
entire class of bug this codebase keeps having: detection then runs on the
photograph rather than the sensor, so every box and landmark is already in the
space the catalog stores and the overlay draws, with no second mapping to get
backwards.

One detector and one embedder serve every lane. The lanes are concurrent futures
on a single thread, not threads, and inference contains no await, so a `RefCell`
borrow never overlaps another — a pair per lane would duplicate ~16 MB of
weights for no parallelism.

Images with no face in them are recorded too. `face_index` records that
detection *ran*, and zero is its most valuable value: without the row every
landscape and document scan returns on every pass, for ever, and in a personal
library that is most of it (§7a).

The old store-only pass survives as `spawn_store_face_sweep` for
`examples/face_index.rs`, which indexes a local store with no network. The
settings copy no longer claims indexing reads "the photographs already
thumbnailed above", and the audit line says "to fetch" rather than "awaiting a
proxy", which had become a blocker that no longer blocks.

Verified against the real catalog: the new work list returns 23,308 where the
old one returned effectively nothing. 469 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 18:51:53 +02:00
dtourolle d4a34effe7 Merge branch 'worktree-faces-scrfd-mbf' into master
Build and test / Desktop (Linux) (push) Successful in 1h22m16s
Build and test / Layer separation (push) Successful in 41s
Traceability / Requirement traces (push) Successful in 30s
🐳 Android image / Build and push (push) Successful in 8s
Build and test / android-image (push) Successful in 8s
Build and test / Android (aarch64) (push) Failing after 33m45s
2026-08-27 18:47:28 +02:00
dtourolleandClaude Opus 5 69dccaa061 Count the platform layer's traceability tags
`platform` was missing from the traceability scanner's source roots, so
every TRACES tag in `dr-plat` — the secret store, volume discovery, and
now display-profile acquisition — was invisible to the matrix. The
FR-PLAT-* family is exactly what that crate exists to satisfy, so the
omission understated coverage by the requirements it was meant to
count.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 18:46:59 +02:00
dtourolleandClaude Opus 5 131004393d Follow the canvas from one display to the next
The rest of FR-DSP-8. The develop session now carries the space its
canvas is encoded into, and `render` composes for it instead of for
sRGB — which is the whole of the change to the pixel path, because the
output space was always a parameter of composition and always entered
the structure hash. A display change is a recomposition.

The space is set on the way into every render rather than pushed when
the window moves, so a photograph opened while the window already sits
on the second monitor is right on its first frame instead of flashing
the wrong colour until the next poll.

Which display that is comes from sampling the window's position and
scale factor twice a second — Slint reports neither a move nor a
display change — and re-surveying only when they differ. Settings shows
what came back under ABOUT: the display, the space, why, and the other
monitors, because the failure FR-DSP-8 names is one that is invisible
from the display you are reading the page on.

Fractional scaling: the canvas is now rendered at the physical pixel
size of the box it occupies rather than the logical one, so the
compositor presents it 1:1. At 1.25 it was previously handed 1600
samples to fill 2000 device pixels, and the softness that produces
reads like a bad demosaic rather than like a scaling bug.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 18:44:51 +02:00
dtourolleandClaude Opus 5 7235972ca4 Measure what a frame costs, so FR-DSP-2 is decided by numbers
`docs/display-and-extension.md` §2 fixes a decision rule in advance: if the
99th percentile of a frame sits inside 16 ms, tiled computation is rewritten
as a scheduling concern for export rather than built on the interactive path.
Nothing in the tree could answer that, so the rule had nothing to act on.

This is the instrument. It renders a 60 MP synthetic source through the real
`render_detailed` at three viewport sizes and four chain lengths, fit and
zoomed to 1:1, and reports nearest-rank percentiles rather than means — a
slider drag is judged by its worst frame.

Three things it does that a simpler timer would not:

- It separates the fused pass from the neighbourhood stage. "Every operation
  active" mixes one dispatch together with a chain of convolutions, and §2's
  question is about the first of those. `point` is every operation that
  contributes a fragment to the fused shader; `all` adds the four with
  kernels, and M3 times those alone by moving only a detail parameter so
  `render_detailed`'s colour reuse skips the fused dispatch. The reuse is
  reported rather than assumed — the `colour` column counts fused dispatches
  and must be zero for an M3 row to mean what it says.

- It times the CPU half separately. Composition runs per frame in
  `DevelopSession::render`, so it is inside the budget whether or not anyone
  has looked at it, and if shader assembly were the expensive half then no
  tile scheduler could help.

- It builds the "every operation" chain from `EditGraph::capabilities` rather
  than from a list, so declaring a new node does not quietly turn that row
  into a shorter chain wearing a longer chain's label.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 18:32:00 +02:00
dtourolleandClaude Opus 5 57c0cc0d35 Keep the name typed into a cluster when the next one is opened
Naming a cluster and moving straight to the next is the gesture this screen
exists for, and it discarded the name every time. Two causes, both in the same
four lines.

`Field.text` is two-way bound to its `TextInput`. Binding it to `selected-name`
therefore works exactly once: the first keystroke writes through the `<=>` and
**replaces** the declarative binding, after which the field follows nothing.
Switching clusters left the previous cluster's half-typed text on screen,
attached to the new person.

And `Field` only reported `accepted`, which is Enter. A name typed and then
abandoned by clicking the next face never reached Rust at all.

So `Field` gains an `edited` callback, the screen keeps the draft with the
person it was typed for, and the draft is written when the selection moves or
the screen closes. The field is then reset from a revision counter the screen
watches.

A counter rather than `changed selected-name`, because the name is not a key:
naming six clusters "Anna" in a row never changes `selected-name`, and the field
would keep the half-typed text from the cluster before. Nor `changed
selected-person`, since accepting a namesake merge lands the user back on a
person they may already have been on.

The draft carries its `PersonId`. A reload can move the selection out from under
a half-typed name — a merge arriving through a sync, a deletion — and applying
it to whoever is selected now would rename a stranger. If the person is gone
when the draft lands, it is dropped rather than resurrecting a row the rail no
longer shows.

`None` and `Some("")` are kept distinct. A user who cleared the field means to
clear the name; a user who never touched it means to leave it alone. Collapsing
those two erases names by walking past them.

An implicit commit does not raise the namesake merge offer. That question is
about a screen the user has already left, and answering it on their behalf while
they look at the next cluster is not a question at all — the offer stays on the
explicit submit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 18:32:00 +02:00
dtourolleandClaude Opus 5 e4875498ca Ask the platform what colour the screen actually is
FR-DSP-8's acquisition half. `dr_plat::display` surveys the session's
displays and reduces each one's profile to an output space the pipeline
can encode into, stating the mechanism per display server as
FR-PLAT-LIN-2 requires:

  - X11 reads the `_ICC_PROFILE` / `_ICC_PROFILE_<n>` root-window
    properties, enumerating and numbering the outputs through RandR,
    which also yields the rectangles a window move is measured against.
  - Wayland binds `wp_color_manager_v1` and asks each `wl_output` for
    its image description, accepting either an ICC profile on a file
    descriptor or primaries stated as chromaticities.
  - Where neither answers, sRGB is assumed and the reason travels with
    it as data rather than into a log, so the About page can say which
    path the session is on.

A display profile is a measurement of one panel and is none of the four
spaces the pipeline knows. Rather than grow an ICC engine, the profile
is reduced to D50-adapted colorants and matched against the four; a
match that is merely nearest is marked as such and shown as such.

Verified on this machine: mutter 50 advertises the colour-management
global and reports eDP-1 as sRGB, and the same session forced onto X11
enumerates the output through RandR and correctly finds no atom.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 18:20:28 +02:00
dtourolleandClaude Opus 5 16c17c349d Bump pkgrel so makepkg rebuilds instead of reusing the modelless archive
`makepkg -si` reported "A package has already been built, installing existing
package" and installed 0.7.0-1 — the archive from before the models were added,
so the install still had no model and the app still said so. The version had not
changed because the application had not changed; only what the package contains
did, which is precisely what pkgrel exists to signal.

Verified: 0.7.0-2 carries both models at /usr/share/darkroom/models/.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 18:14:56 +02:00
dtourolleandClaude Opus 5 2d95807542 Package the models on every platform, not just the phone
The Android bundling landed the weights under that platform's asset directory,
which was the wrong home the moment a second packager wanted them. `makepkg -si`
produced a desktop install with no model at all — the same "no face model is
installed" the phone used to show, for the same reason: nothing put the files
anywhere the app looks.

So `models/face/` at the root is the one copy, and both packagers read it:
assemble-apk.sh bundles it as APK assets, and the PKGBUILD installs it to
/usr/share/darkroom/models. Both refuse an LFS pointer rather than shipping a
130-byte file that fails inside the graph loader on a user's machine.

`face_models` now searches three places, most specific first: the account's own
directory, the shared user directory, then $XDG_DATA_DIRS. So a packaged pair is
found automatically and a pair the user placed by hand still outranks it — which
is what keeps a deliberate choice of weights from being overridden by an
upgrade.

$XDG_DATA_DIRS rather than a hard-coded /usr/share: that is the variable a
distribution, a prefix install or a Nix-style store already sets to say where
its data went, and its documented default is exactly the two paths that would
otherwise have been hard-coded. Empty on Android, which has no such directories
— there the APK's copy is unpacked into the shared user directory instead,
because an asset inside a package is not a path anything can read from.

Verified: the APK still carries both models at assets/models/, the PKGBUILD
parses and installs from the new path, 467 tests pass.

Includes the pkgver 0.6.0 → 0.7.0 bump that was already sitting uncommitted in
the working tree.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 18:05:23 +02:00
dtourolleandClaude Opus 5 725f7bf77f Offer the merge when two people turn out to share a name
Over-clustering is the normal state of a freshly indexed library — FR-CULL-10
says so — which means one person arrives as several groups and the user names
each of them the same thing. Until now that produced several people called Anna
and no way to join them: `identity::merge` existed, `faces::merge_people`
existed with its redirect tombstone, and `merge-into(int)` sat in identity.slint
declared, never emitted and never wired. The screen had a split button and no
merge.

So a rename that collides now offers one. Type a name another person already
carries and a strip appears under the field: "Someone else is already called
Anna (14 faces). Merge them?"

Offered, not performed. `people.uuid` is the identity and the name is not — the
schema comment on that column is explicit that two devices naming the same
cluster independently is the case it was built for — so two people sharing a
name is legal, and folding them together on a keystroke would be the screen
making an identity decision on the user's behalf. That is the thing this screen
spends a whole button avoiding.

The rename always lands first, and declining leaves it exactly as typed. There
is nothing to undo because nothing was done.

Details that are not arbitrary:

The comparison is trimmed and case-insensitive. "anna" on a phone keyboard and
"Anna" on a desktop are one intention, and an offer that appeared only when the
capitalisation matched would read as a bug.

An empty name collides with nothing. Every unnamed cluster renders as "Unnamed
(n faces)"; if that counted as a collision the offer would appear on every
cluster in a fresh library, and accepting it would fold the library into one
person.

The newly-named person folds into the one that already held the name, not the
reverse. The older person is the one other devices have seen and the one whose
confirmations are more likely to be real. Selection follows the merge, because
landing on an empty screen after a successful action reads as a failure.

The offer is retired when the person changes, and when a refresh finds its
target gone — merged from the other side of a sync, or deleted. An offer left
standing would fold whoever happens to be selected now.

A merged-away person is not a namesake: `faces::people` already excludes
redirects, so the offer does not reappear the instant it is accepted.

Four tests over the collision rules, and the existing 467 still pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 13:33:36 +02:00
dtourolleandClaude Opus 5 8155f5276e Spec how the display contract closes, and how the pipeline opens
Two families were the weak points: FR-DSP at 37%, which is the architecture's
central performance claim, and FR-PLG at zero. They are one document because
the same property decides both — the pipeline composes its work from
declarations, which is why the display path is fast and is also, already, most
of a plugin format.

Two findings change the size of the job.

**Some of FR-DSP is done and untagged.** Zoom already meets FR-DSP-5: the
sampled region shrinks while the render target keeps its size, so zooming
raises the resolution the pipeline works at — arrived at without tiles.

**FR-DSP-2 may not be worth satisfying as written.** It predates the fused
shader and assumes a chain of passes over a large buffer, where recomputing
per frame is ruinous and tiles are the way out. What exists composes every
active operation into one dispatch over a viewport-sized target. So the spec
fixes three measurements and a decision rule *in advance*: if a frame sits
inside budget at the 99th percentile, the requirement is rewritten rather than
implemented, and tiling becomes what it actually is here — a scheduling
concern for export, which already runs off the frame path. Writing a tile
scheduler the design does not need would be the most expensive way to find
that out.

**The plugin format already exists**, resolved at build time: `ops/*.yaml`
through `build.rs` produces something indistinguishable from a hand-written
operation, and everything it produces is data plus a WGSL string. The blocker
is one line — `descriptor()` returns `&'static` — and the rest is an
interpreter over declarations, load-time WGSL validation, and namespaced ids,
because the sidecar stores parameters by id and a collision is a wrong edit
silently applied.

Recorded last and deliberately: traceability counts a tag, not a behaviour.
`FR-DEV-8` is tagged against plumbing a future spot-removal op would use and
`FR-DEV-7` against a history row for a frontend that does not exist, so 51% is
an overstatement of unknown size. Every requirement this plan closes should be
closed by a test that fails if the behaviour is removed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 13:28:39 +02:00
dtourolleandClaude Opus 5 eaafacc3fb Give the phone the model it had no way to obtain
Face indexing was compiled into the APK all along — dr-ui takes dr-face with
`inference` on every target, so SCRFD, alignment, MBF, calibration and
clustering were all in there. What was missing was the weights, and on Android
there was no way to supply them.

Route C (docs/faces.md §2.2) says the user obtains the model and the app loads
it. On a desktop that is a real gesture: drop two files in
~/.local/share/darkroom/models/ and indexing starts working. On Android it is
not a gesture at all. `internal_data_path` is app-private, `run-as` needs a
debuggable build, and the in-app fetch route C specifies was never built — so
the settings page reported "no face model is installed" on every launch with
nothing behind the message. Not "off until you supply weights"; off.

So the shape-fixed pair goes into LFS under the APK's assets, assemble-apk.sh
copies it into the package, and `android_main` unpacks it to the shared models
directory before anything asks whether a model is present.

Three things that are not incidental:

The models directory is now shared across accounts rather than per-account.
Weights are identified by `faces.model_id`, not by who is signed in, so two
accounts had no reason to hold two copies — and the unpack runs before any
session exists to key a per-account path off. `face_models` still prefers a
per-account directory when one is populated, so anyone mid-migration keeps the
ability to pin one library to its own pair.

The unpack writes under a temporary name and renames. `face_models` decides
availability on `is_file()` alone, so a copy truncated by the process being
killed would leave a file that passes that test and fails inside tract —
reported to the user as a broken model rather than a missing one.

assemble-apk.sh refuses an LFS pointer. At ~130 bytes it looks exactly like a
model to `cp`, and unchecked it reaches the device and fails in the graph
loader instead of telling someone to run `git lfs pull` — the same guard
dr-segment's build script applies to yolo26n-seg.onnx.

The licensing half is unchanged and recorded in §2.2a: the InsightFace grant is
research-only, this is a private repository and a self-installed build, and
these files come back out before anything is published. The weights are still
not a cargo build input — dr-face has no `models/` directory and no
`embedded-model` feature, and nothing in the build reads them. The APK assembly
step copies two files and is the only thing in the tree that knows they exist.

Verified on device: both models unpack on first launch (2524817 and 13616095
bytes) and the APK carries them at assets/models/.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 13:18:07 +02:00
dtourolleandClaude Opus 5 b846b312b8 Run the formatter over the face branch before it reaches CI
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 2s
Build and test / Desktop (Linux) (push) Successful in 1h21m32s
Build and test / Layer separation (push) Successful in 37s
Traceability / Requirement traces (push) Successful in 25s
Build and test / Android (aarch64) (push) Failing after 33m58s
The merge of the SCRFD/MobileFaceNet work brought 69 rustfmt diffs across
dr-catalog, dr-face and dr-ui with it, so `cargo fmt --all -- --check` fails
on master and the Desktop job stops at its Format step — before clippy, the
tests or the release build have run at all. That makes the whole desktop
half of CI blind: a real compile error behind this would look exactly the
same from the outside. There was nothing behind it, as it turns out — with
the formatting fixed, clippy, the test suite and the release build all pass.

Every .rs hunk is `cargo fmt --all` on the pinned 1.92.0 toolchain, not a
hand edit, but it is worth being precise about what that moved, because it
is more than whitespace. Besides reflowing signatures and call chains,
rustfmt reordered the `pub mod` and `pub use` items in dr-face/src/lib.rs so
the `#[cfg(feature = "inference")]` entries sort in place, added the trailing
semicolon inside `let ... else { return }` bodies in identity_ui.rs, wrapped
a bare closure body in braces in cluster.rs, adjusted trailing commas, and
dropped a stray blank line at the end of identity_ui.rs. All of it is
semantically inert; none of it changes behaviour.

docs/traceability.md rides along because it has to. The matrix records each
TRACES tag by line number, and reflowing develop.rs, lib.rs, faces.rs,
identity.rs and identity_ui.rs moved them — FR-CAT-8, FR-CAT-9, FR-CULL-10,
FR-DEV-3, FR-DEV-3a and FR-DEV-3c all shift by a line or two. The matrix was
verified up to date on d777f7f before this commit, so this is drift these
formatting changes introduced, not pre-existing staleness being swept up.
Leaving it for a follow-up commit would hand traceability-check.yml a
failure caused entirely by a whitespace change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 13:12:49 +02:00
dtourolleandClaude Opus 5 9997623df4 Compile the mount-table reader only where there is a mount table
`platform_volumes` has been gated to Linux since it was written, but the
eight items it is built from were not, so every Android build compiled a
`/proc/mounts` parser it could never call and then printed eight dead-code
warnings about it. Real warnings hide in that kind of noise.

Gated per item rather than moved into a module, because the file already
draws the line that way one function above and two patterns for one idea
is worse than a repeated attribute.

The tests go with them. They parse a mount table and assert on names like
`mmcblk0p1`, so they are as Linux-bound as the code they exercise, and
`Path` turns out to be too -- only the reader borrows one, where `Volume`
owns its own.

Android now builds dr-plat with no warnings at all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 12:17:25 +02:00
dtourolleandClaude Opus 5 721efbc371 Drop three trait imports Android does not need
The Linux store is built on `keyring::Entry`, whose set_password,
get_password and delete_credential are inherent methods. The Android one
is built on `keyring_core::Entry`, where they are inherent too -- so the
`CredentialApi` import that each of its three methods opened with was
doing nothing, and said so on every Android build.

`CredentialStoreApi` a few lines above is a different matter and stays:
`build` really does come from that trait.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 12:16:17 +02:00
dtourolle d777f7f44d Merge branch 'master' into worktree-faces-scrfd-mbf
Build and test / Desktop (Linux) (push) Failing after 25s
Build and test / Layer separation (push) Successful in 22s
Traceability / Requirement traces (push) Successful in 58s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Failing after 33m26s
# Conflicts:
#	docs/traceability.md
#	ui/dr-ui/src/develop.rs
#	ui/dr-ui/src/segmentation.rs
2026-08-27 11:57:38 +02:00
dtourolleandClaude Opus 5 7e1c33ebed Draw the subjects on the photograph, not on the sensor
Build and test / Desktop (Linux) (push) Successful in 19m42s
Build and test / Layer separation (push) Successful in 27s
Traceability / Requirement traces (push) Successful in 34s
🐳 Android image / Build and push (push) Successful in 4s
Build and test / android-image (push) Successful in 4s
Build and test / Android (aarch64) (push) Failing after 33m12s
"Find subjects" would recognise a person on a portrait frame and then paint
the outline into the hillside behind them. The detection was right and the
mask was right; what was wrong was the picture drawn to show them.

Instance masks live in sensor space, and correctly so — the generated shader
samples them at `uv_src`, after the framing map, which is what keeps a mask
on its subject through a zoom, a pan and a crop. The overlay is the one
consumer that is *not* sampled by that shader. It is a flat image handed to
the compositor to lay over a photograph that has already been through the
framing map, so it has to arrive in the same space that photograph is in, and
it did not. On a frame from a camera held sideways the outlines were drawn a
quarter turn away from the subjects they described.

`overlay_clip` had the same fault one layer down, and it is the more
insidious of the two because it looks right. The crop and the viewport are
fractions of the photograph as the user sees it — the prologue maps an output
pixel through `crop_rect` *before* it unturns the frame — and they were being
measured against the sensor's width and height. Two numbers, correct type,
wrong axis.

Both now go through `Orientation::into_shown`, so the overlay and its clip
are in the photograph's space and the turn is the same one the render and the
thumbnails make.

Neither was noticeable until this week, and the reason is worth writing down:
before the detector was given an upright frame it found almost nothing on a
portrait photograph, so there was rarely an outline to be in the wrong place.
Fixing the detector is what made this visible.

Landscape frames were never affected, which is most of them, and is why an
overlay that ignored orientation entirely survived this long.

Verified on `_MG_9080.CR2`, a portrait frame of two people and a dog: the
overlay was a 1599x1066 image drawn onto a 1066x1599 canvas, with the colour
sitting in the mountainside above the subjects. It is now 1066x1599, and each
outline is on the thing it names.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 09:46:48 +02:00
dtourolleandClaude Opus 5 ce201c7dd6 Name the two spaces a photograph lives in, so a turn cannot go the wrong way
Every orientation bug this codebase has had has been the same bug: a turn of
the right size applied in the wrong direction. That failure is worth naming
precisely, because it does not look like one — a quarter turn applied
backwards lands 180 degrees from right, so the result is a plausible
transform of the picture rather than anything obviously broken, and on
landscape frames it is not wrong at all. It was the straighten shear, and it
was the segmentation overlay, and each time it was found by eye rather than
by a test.

The reason it keeps happening is that "rotate 90 degrees clockwise" cannot be
checked by reading it. The reader has to hold in their head which of the two
images is being rotated and which way the y axis runs, and there were four
hand-written copies of the permutation to hold it for: the shader prologue,
its CPU twin, the thumbnail path, and the segmentation.

So nothing added here says clockwise, anticlockwise, horizontal or vertical.
The functions say *which space they take and which space they return* —
`into_shown` and `into_stored`, `source_pixel` and `shown_pixel`,
`into_shown_rect` and `into_stored_rect` — and each takes the dimensions of
the space it reads from, so no caller has to work out which pair it is
holding. `StoredRect` and `ShownRect` are separate types because they are the
same four numbers meaning different things, which is exactly the case where a
mistake is silent: a shown rect measured against stored dimensions produces a
rectangle in the wrong place, not an error.

Underneath there is one permutation. `source_pixel` was already shared by the
prologue and the thumbnails; `source_point` is its normalised twin, written
beside it so the two cannot drift, and everything else is those two read
forwards or backwards. `Orientation::inverse` is the group inverse rather
than `4 - turns`: mirrors apply after the turn, so undoing means undoing them
first, and a mirror seen from the far side of an odd turn is about the other
axis. That is the diagonal-mirror case, tags 5 and 7, and getting it wrong
renders as — again — 180 degrees.

Three call sites lose their own copy: the thumbnail path, `dr-ui`'s
segmentation, and `dr-gpu`'s `local` example. "Upright" now means one thing
across the application rather than one thing per caller.

The gate that matters most is `the_render_and_the_orientation_map_agree`. The
shader prologue and `Orientation` answer the same question by different
routes, and until now nothing checked that they answered it the same way. It
now checks every EXIF tag against every user rotation and mirror on top of
it, because the composition is where the two could agree singly and disagree
together.

The rest earn their place by having caught something. Writing these found two
real errors in this commit's own new code before it ran anywhere: `shown_pixel`
was handed the dimensions of the wrong space and overflowed, and the rect map
turned the wrong way for the diagonal mirrors. A round trip that returns what
went in is the only check worth having here, since every wrong answer is
still a picture.

No behaviour changes. The permutations are the ones that were already being
applied; they are simply applied from one place now.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 09:46:24 +02:00
dtourolleandClaude Opus 5 25c88d9dbd Start face indexing from Settings
Beside the thumbnail sweep, because it is the same kind of thing: a job
that runs for an hour, is asked for once, and reports into the activity
list above it. It is also downstream of that sweep -- detection reads the
proxies it builds -- so the two belong in that order, and the coverage
line says how many images are waiting on a proxy rather than only how
many are left to index.

The button drives the Identity Manager's own state rather than a second
copy, so it cannot disagree with that screen about whether a pass is
running, and either place can start or stop it.

The pass now opens an activity row. The caption promises progress will
appear in the list above, and without a row it would not: the button
would be the only sign anything was happening, invisible from every
other screen.

Coverage is read when the Settings page opens. The figures live in the
catalog and this page deliberately holds no session, so they arrive
through a closure rather than being kept current -- they are only ever
looked at while the page is on screen, and the check is two counts and an
indexed scan.

Also adds DARKROOM_NO_SYNC. Redirecting XDG_DATA_HOME isolates a test
launch's catalog and thumbnails but not its server, and I found that out
by pushing a test catalog over the live one. The guard sits in
start_derived_sync rather than at its three call sites, because the sweep
firing a sync is correct and a flag checked in three places is one that
gets missed in a fourth.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 09:22:43 +02:00
dtourolleandClaude Opus 5 f00b92a0e6 Turn face boxes back into sensor space before matching regions
Faces are found on the thumbnail, which is cached the right way up --
the grid would lie on its side otherwise. Segmentation runs on a proxy
rendered through a neutral edit graph, which carries no orientation and
is therefore in sensor order. For anything shot in portrait the two
differ by a quarter turn, so a face and the person containing it were
being compared in spaces 90 degrees apart: no match, or worse, a match
against somebody else's region.

The transform goes on the face rather than on the proxy. Instance masks
are defined in the proxy's space and sampled long afterwards, so turning
that space would be a far larger change than naming a region warrants.

Also two things the first screenshot of the running app showed that no
test would have:

110 of 23,528 displayed as "0%", which reads as the feature having done
nothing. One decimal below ten percent, and a floor so real progress
never shows as none.

The rail picked some near-black covers, because the largest face in a
group is often the nearest one in a badly lit frame and a black square
beside a name identifies nobody. It now cuts the best few and takes the
first legible one, falling back to the largest when a person's every
photograph is dark -- which happens, and showing it beats showing
nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 23:28:52 +02:00
dtourolleandClaude Opus 5 61c4547b9c Make the Identity Manager a peer of the library and develop
It was reachable only from the library header, which made it a side trip
rather than a mode. It is now reachable from develop's header too, beside
the way back, because that is the same kind of move -- leaving this
photograph for somewhere else in the library -- and a screen you can only
reach from one of the other two is not a peer of them.

Leaving returns to whichever screen opened it, and the button says which.
A back button that read "Library" while returning to develop would be
lying about the one thing a back button has to be right about. The
develop session is only hidden, never torn down, so returning to it costs
nothing and keeps the photographer's place.

The header now matches the other two screens rather than using a close
cross: three screens whose headers disagree read as three applications.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 23:14:17 +02:00
dtourolleandClaude Opus 5 b55812812a Name segmented people from the faces already recognised in them
The segmenter knows it found a person; the face index knows which person.
Joining them turns "person" in the mask list into "Anna", which is the
difference between a vocabulary of eighty COCO classes and one that
includes the user's family. Selecting a subject in a group photograph
stops being a guessing game between three identical rows.

Containment, not IoU. A face is a small part of the person it belongs to,
so a correct pairing has an IoU near zero and anything IoU-based would
reject every true match.

Confirmed names only. A suggestion is the system's guess, and printing a
guessed name onto a mask region would launder it into a fact.

Writing the tests corrected the design once: a tight head-and-shoulders
portrait, where the face fills most of the person box, is the case where
naming is most certain, not least. An earlier guard rejected exactly that
and has been removed, with the reasoning left as a test because it is
easy to get backwards a second time.

The names hang on the develop session, set when the image opens because
that is the one moment the catalog and the image id are both in reach.
Every segmentation run afterwards picks them up for free, and a library
with no face indexing behaves exactly as it did before.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 23:10:46 +02:00
dtourolleandClaude Opus 5 d7a81375ee List the steps, and let a photographer step straight to one
Build and test / Desktop (Linux) (push) Successful in 19m30s
Build and test / Layer separation (push) Successful in 25s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
Traceability / Requirement traces (push) Successful in 23s
Build and test / Android (aarch64) (push) Failing after 33m5s
Undo answers "take back the last thing", which is the question asked about a
mistake just noticed. It is the wrong instrument for one noticed six
adjustments later: eight presses, each changing the picture, with no way to see
how far back the mistake is without passing through it. A step is a whole
state, so arriving from six away costs what arriving from one does — which is
what makes a row worth making clickable rather than decorative.

`Edit::Discrete` had to go for the list to be worth drawing. Seventeen call
sites recorded the same anonymous step, which is fine for deciding whether two
changes are one gesture and useless for a panel: seventeen rows reading
"Discrete" is not a history. Every variant now carries enough to name itself,
and the compiler enumerated the sites that had to start saying so. A step that
moved a parameter is still named out of the descriptor, so an operation added
as a YAML declaration appears in the history correctly named with nothing
written for it (FR-DEV-3c).

Choosing a film stock was not undoable at all. The pick went straight to
`choose_film`, which nothing on the history's path ever sees. `pick_film`
records it, and is separate because the same call is also how a *restored* edit
gets its tables back — recording that would push a step for the undo the
photographer had just asked for.

The list is rebuilt off a revision rather than off every redraw. A drag ends in
a redraw per frame while folding into one step, so the unconditional version
would tear down and recreate every row sixty times a second to arrive back at
the list already on screen. The counter is process-wide: a per-instance one
starts every photograph at the same number, so a frontend holding "the revision
I last drew" would keep the previous image's steps on screen — invisible while
every image opens with one identical row, and a wrong-photograph bug the moment
persisted history means it does not.

The step names that no descriptor can supply are constants with a roll, and a
test walks the roll rather than a second copy of it. `resolve` splits so that
"is this catalogued?" can be asked: `derive` turns `history.mask_toggled` into
"Mask Toggled", which names a field rather than an act and, being perfectly
readable, is a mistake nobody would look at twice.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 23:04:35 +02:00
dtourolleandClaude Opus 5 b89f1cfece Snapshot the whole edit in the history, so a drawn mask can be taken back
The undo stack snapshotted a `Preset` — the parameter map — and a mask layer
is deliberately not a parameter. So drawing one changed nothing the history
could see: `record` returned `false`, no step opened, and the layer the
photographer had just painted had no way back. The interface went on calling
`record` in good faith, including from the mask controls, and nothing failed.
A film stock went missing the same way.

The snapshot is an `EditState` now, so the history is complete by
construction rather than by anyone keeping a list in their head.

`undo` and `redo` return a `Step` rather than a `bool`. Stepping is not the
only outcome a caller has to act on — a step across a change of film leaves
the graph without its tables, and only the caller can bake them — and a `bool`
would let that be dropped by writing nothing at all, which is the shape of
mistake this module had already made once. `DevelopSession` settles the debt
either way; a step that found nowhere to go is left alone, since clearing the
film because undo hit the floor would take the stock off the picture.

Five tests, all of which fail against the old snapshot: a drawn layer is
undoable and redoable, a layer's own settings are a step of their own, a
change of stock is a step and names what it needs baked back, clearing the
film is undoable, and an exposure move does not deep-copy the mask stack.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 22:54:55 +02:00
dtourolleandClaude Opus 5 93efdf27a6 Keep the film stock when an edit is saved
`Version::update` is the write path an automatic save goes through. It copied
the parameters and the masks and said nothing about the film, so a photograph
developed on a stock was written back without it and opened the next time
without its emulsion. Nothing reported a failure — the line was simply not
there.

It is the third of the three routines that captured "the edit" and the only
one that got it wrong, which is the argument for not having three. All of them
now destructure one `EditState`, so `from_graph`, `update` and `apply` cannot
disagree about what an edit consists of, and the next part of one cannot be
lost by anybody writing a line too few.

Two tests, both of which fail without the fix: the field survives `update`,
and the stock survives the round trip through the file. `apply` returns the
`FilmRebake` it always implicitly owed, so `apply_version` now reads the debt
off the call rather than off `version.film` — and pays it in both directions,
since a version with no film has to clear the adjust pass too or it keeps
textures bound that nothing will sample.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 22:53:00 +02:00
dtourolleandClaude Opus 5 d562ceaaf4 Show a face beside each name in the people rail
The rail was drawing an empty square for every person: cover was
hardcoded to a default image. That is the one place a portrait matters
most, because the rail is how the user decides which unnamed group to
open first, and a list of "Unnamed (24 faces)" rows tells them nothing.

The portrait is the person's confirmed face with the largest crop_px --
the most source pixels the face actually occupied, so the one they have
the best chance of recognising -- falling back to a suggestion so a
freshly clustered group still has a face beside it.

Cached in the controller, because every mutating action reloads the whole
screen and cutting a portrait costs a JPEG decode per person. Without the
cache, confirming one face would re-decode a proxy for every person in
the library, and the rail does not change when a suggestion is accepted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 22:52:38 +02:00
dtourolleandClaude Opus 5 5622a58ce3 Give an edit one complete state, and make omitting part of it a compile error
An edit used to be a bag of scalars. That stopped being true when the mask
stack and the film stock arrived, both deliberately held apart from `ops`
because a layer is not a scalar and a stock is not a scalar — and nothing
announced the change. What happened instead is that three routines each
captured "the edit" and each captured a different subset of it.

`EditState` is all of it: the parameter map, the masks, the stock. What keeps
it complete is not a comment. `EditGraph::state` destructures the graph
exhaustively, `EditGraph::set_state` destructures the state exhaustively, and
the fields are public so every construction site is a struct literal naming
all of them. Adding a fifth kind of graph state — FR-DEV-8's spot removal is
the one already asked for — fails to compile until somebody has decided
whether an undo has to put it back. Verified both ways round by adding a field
to each type and watching five call sites refuse to build.

A compiler error rather than a runtime check, because the failure being
prevented is silence: the missing halves produced no panic, no warning and no
failing test.

The stack is now shared rather than owned, and that is about the drag path
rather than memory: `state` runs on every parameter change, which during a
drag is once a frame, and deep-copying a painted brush sixty times a second to
record an exposure move would be a cost paid for nothing. `masks_mut` is the
one door a stack is modified through, so it clones on write.

`FilmRebake` is the one thing a caller is still owed. Restoring a stock always
needed the profile database this crate does not link (ARCH §6.5a); it was a
comment before, and it is a `#[must_use]` return value now.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 22:50:11 +02:00
dtourolleandClaude Opus 5 2944c1b704 Document the run marker and cross-device face sync
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 22:42:28 +02:00
dtourolleandClaude Opus 5 96d07da15f Sync face data as sealed shards, so a second device does not re-index
Indexing 23,500 images is about two hours of CPU, and the result is
byte-identical on every device: the same model over the same proxy
produces the same embedding. Paying for it once per account rather than
once per device is the point.

Shards rather than the catalog snapshot, because the snapshot goes up
whole on every sync and a fully indexed library carries roughly 30 MB of
embeddings. That is exactly the cost the thumbnail store's 25 MB cap
exists to bound, so face shards use the same cap -- imported from
dr_thumbs rather than restated, since the number is a statement about
sync cost and the two must not drift apart.

The split follows the one already there: bulk immutable data in sealed
shards, small mutable data in the catalog snapshot. Faces, landmarks,
embeddings and run markers shard; people, names and assignments ride the
catalog and merge by uuid.

Keyed on oc:fileid throughout, never on image_id, because a row id means
nothing on another device.

The run marker travels with the faces it describes. Without it a
receiving device cannot tell an image with no faces from one never
examined, and would re-detect every landscape it had just adopted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 22:41:50 +02:00
dtourolleandClaude Opus 5 4a82753d22 Show the detector the photograph, not the sensor's scanlines
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
Build and test / Desktop (Linux) (push) Successful in 19m7s
Build and test / Layer separation (push) Successful in 25s
Traceability / Requirement traces (push) Successful in 23s
Build and test / Android (aarch64) (push) Failing after 33m10s
"Find subjects" was handed the proxy in the sensor's own orientation, so
every frame shot on a body held sideways reached the model lying on its
side — and a model trained on upright photographs is very bad at those.
Measured end to end on a 22 MP frame of two people and a dog: `person
0.36` and nothing else, against `dog 0.82, person 0.61, person 0.49` for
the same pixels stood up. Nothing failed; the panel simply offered one
poor subject where there were three good ones.

The orientation was never dropped on purpose. The proxy is deliberately
rendered through a *neutral* graph — the detection has to survive an
exposure change, or every slider would invalidate the masks built on it
— and neutral took the file's orientation with it along with everything
else. Landscape frames were unaffected, which is why it stood for as
long as it did.

The turn is `Orientation::source_pixel`, the same function the grid's
thumbnails already go through, so the detector and the thumbnailer now
agree about which way is up rather than holding two opinions. What it is
turned by is `Framing::effective_orientation` — the file's EXIF tag and
the photographer's own rotations composed into one permutation, by the
group law rather than by adding the turns, which is a distinction
`Framing` already had to make and had already tested. Rotating the
picture and pressing the button again therefore does what it looks like
it does.

The proxy stays in sensor space and the masks come back into it. That is
not a detail to be tidied later: the generated shader samples the mask
array at `uv_src`, *after* the framing map, so a mask stored upright
would sit a quarter turn off the subject it was drawn around. That is a
wrong mask rather than a weak one, and nothing announces it. So the
picture is stood up for the model and laid back down for everything
else, and `upright`/`lay_down` are returned as a pair because calling
one and forgetting the other is silent.

Both directions are the one function: `upright` gathers through
`source_pixel` and `lay_down` scatters through it. A quarter turn is a
bijection of the pixel grid, so the round trip is exact — no filter, no
resampling, and no hole to fill — and an inverse written out by hand
would be a second thing to keep in step, whose way of being wrong is a
mask mirrored about the wrong axis, which still looks like a mask.

The orientation joins the confidence and the tiling flag in the
segmentation signature, and for the same reason: turning the photograph
changes what the model recognises, so two runs either side of a rotation
are different instance lists. Two that happened to come out the same
length would otherwise share a signature and a stored layer would be
silently re-indexed from one into the other.

The refine pass had it too — it re-runs the model over a crop rendered
in the same sensor space — so it makes the same turn, and would
otherwise have handed back a worse mask than the one it was asked to
improve, on the subject the photographer had just pointed at.

`dr-gpu`'s `local` example is fixed with it. It exists to be the
shipping path with pictures attached, and a diagnostic that reproduces
the bug it is meant to catch is a trap for whoever reads it next.

Seven tests. The round trip is the identity over all eight EXIF tags on
a non-square asymmetric grid; a turn carries whole pixels rather than
shearing the channels apart; a sideways frame reaches the model
upright; a box comes back in sensor pixels, worked out by hand for the
one turn a portrait frame actually writes; a restored box still reads
low-to-high for every tag, since the rest of the pipeline takes
`x1 - x0` without checking the sign; and the eight tags cannot collapse
into one signature key. The existing composition test now runs against
`effective_orientation` itself, over all 8 x 16 baseline-and-user pairs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 22:40:58 +02:00
dtourolleandClaude Opus 5 336a1fd296 Write down what the Android renderer change cost, and what else is owed
Three compromises were taken deliberately over the last few days and none of
them was written anywhere a future reader would look. So `docs/technical-debt.md`,
and two corrections to the architecture document that the Android change made
untrue the moment it landed.

**TD-1, the Android readback.** §12/6.1 says GPU results never round-trip
through the CPU, and the develop view on Android now does exactly that. That
is worth recording as a breach with reasons rather than quietly leaving a
constraint the code no longer honours — the next person to read §6.1 and then
`DevelopSession::render` would otherwise conclude one of them is a mistake.
The entry carries the device measurements that forced it, why setting
`preTransform` is not a fix available to us, and the three separate things any
one of which would remove it.

**TD-2 and TD-3**, the serial thumbnail fetch and the unbounded drain, were
found while chasing the tearing and are still outstanding. Both have a known
shape for the fix; neither is a bug, and neither should be discovered again
from scratch.

The architecture document said the app renders "through wgpu to Vulkan on both
Linux and Android", which stopped being true at 6267802, and §12/6.1 claimed a
constraint with no exceptions. Both now say what the code does and point at
the debt entry for why.

The numbers are labelled with what they are. TD-1's readback cost has *not*
been measured on the device and says so, and TD-3's figures are from a debug
build and say so — a documented measurement that quietly turns out to be the
wrong build is worse than no measurement.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 22:36:37 +02:00
dtourolleandClaude Opus 5 26a1eb7e28 Record that face detection has run, not just what it found
An image with no faces in it was indistinguishable from one that had
never been looked at, so every landscape, still life and document scan in
the library was re-detected on every pass, for ever. In a real library
that is most of it: on the 23,527-image test library, 64 of the first 110
images indexed contain no face at all.

Schema v9 adds face_index, a run marker per (image, model) carrying the
face count and the proxy edge it read. Keyed on the model, so a model
change puts every image back in the queue by itself.

That makes a coverage figure possible, which is the thing a user actually
wants to see. The audit also splits the outstanding set by whether a
proxy exists, because 23,417 awaiting a proxy and 110 ready to index are
different problems, and telling the user to run indexing again would not
fix the first.

The Identity screen gains Index faces, Stop, and the coverage line.
examples/face_index.rs is the same check and sweep without a window,
which is the right shape for an overnight pass.

Measured on the real library in release: 3.5 images/second, 110 images
and 125 faces in 30 seconds, and a second run correctly finds nothing
left to do.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 22:13:41 +02:00
dtourolle e0cb968e04 Merge remote-tracking branch 'origin/master' into worktree-spot-removal
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 2s
Build and test / Desktop (Linux) (push) Successful in 20m9s
Build and test / Layer separation (push) Successful in 39s
Traceability / Requirement traces (push) Successful in 26s
Build and test / Android (aarch64) (push) Failing after 33m10s
# Conflicts:
#	docs/traceability.md
2026-08-26 21:56:32 +02:00
dtourolle bc07c32611 Merge branch 'master' into worktree-spot-removal
# Conflicts:
#	docs/traceability.md
2026-08-26 21:51:47 +02:00
dtourolleandClaude Opus 5 e14bc34a9e Ask which frame this was taken on, because grain is enlargement
Build and test / Desktop (Linux) (push) Successful in 19m6s
Build and test / Layer separation (push) Successful in 28s
Traceability / Requirement traces (push) Failing after 26s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 3s
Build and test / Android (aarch64) (push) Failing after 33m7s
A crystal is a fixed size in micrometres. How grainy a photograph looks is
therefore not a property of the emulsion alone -- it is film size against
output size, and the frame is the half a digital file cannot supply.

This assumed 35 mm for everything. The same emulsion on 4x5 averages about
3,800 crystals into the pixel that holds 300 on 35 mm, so it renders roughly
3.5 times smoother at the same print; every large-format photograph was being
rendered as grainy as a half-frame.

`Format` now carries the real image widths -- the gate, not the nominal inches,
since a "4x5" exposes about 121 mm -- and the film node asks for it. It is a
genuinely fixed list, unlike the stocks, so it is a declared `enum` parameter
and gets its control, its sidecar entry and its undo step for nothing.

It is also the first enum in the develop chain, and it broke two tests by
being one. A row has to compare equal to itself across two builds or
`sync_rows` replaces it on every parameter event -- destroying the elements
built from it, including whichever TouchArea holds the current gesture, so the
format picker would have fought every slider drag in the panel. `ModelRc`
compares by identity and the row built a fresh choices model each call.

`no_choices` already shares one empty model for exactly this reason, and the
build site already said "see no_choices for why the identity matters". The fix
follows it: memoise the model per variant list. Curve rows solve the same
problem the other way, writing values through the existing model, which is not
needed here -- a variant list is fixed at compile time, so one model can serve
forever.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 21:47:19 +02:00
dtourolleandClaude Opus 5 626780276d Draw with OpenGL on Android, where the driver owns the display rotation
The grid tore while scrolling on the tablet, in portrait only, and was
flawless in landscape. It was not vsync and it was not the grid.

Measured on the device, same build, only the tablet rotated:

    landscape   bufferTransform=ROT_180   composition=DEVICE (2)   clean
    portrait    bufferTransform=ROT_270   composition=CLIENT (1)   torn

The panel is mounted landscape — 1920x3000 at installOrientation 3 — so a
portrait window needs a 90 degree rotation before scanout. wgpu-hal hardcodes
the swapchain's `preTransform` to `IDENTITY` and says so in a comment beside
the line:

    // On Android 10+, libvulkan's `vkQueuePresentKHR` returns
    // `VK_SUBOPTIMAL_KHR` if not doing pre-rotation ... This is always the
    // case when the device orientation is anything other than the identity
    // one, as we unconditionally use `VK_SURFACE_TRANSFORM_IDENTITY_BIT_KHR`.

That is gfx-rs/wgpu#3345, and it cannot be fixed by setting the field:
`preTransform` is a *promise* that the content is already rotated, so keeping
it needs the renderer to rotate what it draws, which wgpu cannot do on Skia's
behalf.

We do not have to be on that swapchain. `AndroidWindowAdapter` chooses
`SkiaRenderer::default_wgpu_29` only because this crate enables
`unstable-wgpu-29`; without it `SkiaRenderer::default` resolves — through
i-slint-renderer-skia's build script, which selects OpenGL on anything that is
not Apple, Windows or wasm — to Skia over OpenGL, where the driver owns the
rotation and there is no transform to get wrong. So both renderer features
move to the desktop-only dependency, and desktop is untouched.

# The cost, stated rather than hidden

Skia over OpenGL cannot sample a `wgpu::Texture`, so the develop view's frame
comes back through memory: `AdjustPass::export_pixels`, already ungated and
already used by the export path, into a `SharedPixelBuffer`. That is the
round-trip ARCH §6.1 and AC-8 exist to forbid, and it is the right trade only
because of what the alternative actually is — not a faster develop view, but a
grid that tears in the orientation a tablet is mostly held in.

Two things keep it small. The device is still opened on Android, so demosaic
and the adjust pass are untouched on the GPU; only the last hop changes. And
`render` fits the pass to the canvas before it runs, so the readback is at
viewport resolution, a fraction of the ~7 ms at 4K the original measurement
was taken against.

Four other explanations died on the way here, each by measurement rather than
argument: the present mode (a patch confirmed in the installed binary reached
`AutoVsync`, and the rows still duplicated), our shared wgpu device (Slint
opened its own, unchanged), Skia's partial rendering (off for GPU surfaces),
and client composition itself (unavoidable in portrait on this panel, so it
cannot be what distinguishes a torn frame from a clean one).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 21:47:02 +02:00
dtourolleandClaude Opus 5 3e607222c6 Record the Identity screen in the face spec
Also notes what building it taught the design: a split has to reject
before it confirms, or the next clustering pass undoes it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 21:16:04 +02:00