Compare commits

..
72 Commits
Author SHA1 Message Date
dtourolle 7fa3176f88 Release 0.14.0
Benchmarks / Frame budget (on demand) (push) Skipped
Benchmarks / CPU and I/O (per commit) (push) Successful in 8m25s
Build and test / Desktop (Linux) (push) Failing after 26m11s
Build and test / Layer separation (push) Successful in 33s
🐳 Android image / Build and push (push) Successful in 10m41s
Build and test / android-image (push) Successful in 10m42s
🐳 Windows image / Build and push (push) Successful in 4m17s
Build and test / windows-image (push) Successful in 4m17s
Traceability / Requirement traces (push) Successful in 59s
Build and test / Android (aarch64) (push) Successful in 40m53s
Build and test / Windows (x86_64, cross) (push) Successful in 46m51s
2026-09-21 11:32:41 +02:00
dtourolle af162dd010 Merge: origin's sync ordering and adoption work, its scan-complete trigger ported into library_ui/ 2026-09-21 11:32:05 +02:00
dtourolle f5300a9f43 Merge: the UI crate restructured so a feature owns a function, not a region
Nine agent branches, merged one at a time on ui-wiring and gated at each
step: develop.rs, library.rs, collections_ui.rs and library_ui.rs become
module directories; every screen's wire() and lib.rs::run() become lists
of named functions, with the develop screen's callbacks in develop_ui.rs;
the six view booleans become View and Page enums; the collections sidebar
and the library grid get their own Slint globals, taking 155 members off
AppWindow. No behaviour change: a multiset audit of code lines over every
moved region lost nothing, the workspace gate is clean, and the manual
recorded from this build and from master's from the same library snapshot
matches picture for picture, apart from a panorama stall that master
shows too when a merge starts during the engine's TensorRT compile queue.
2026-09-20 23:41:25 +02:00
dtourolle b19d470189 Merge: the library grid on its own Slint global, the last of the screen state off the root 2026-09-20 22:43:09 +02:00
dtourolleandClaude Opus 5 bfadd9c409 Ship the border filler in the Windows installer, and count what is staged
Benchmarks / CPU and I/O (per commit) (push) Successful in 8m2s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 25m17s
Build and test / Layer separation (push) Failing after 0s
Traceability / Requirement traces (push) Failing after 0s
🐳 Android image / Build and push (push) Failing after 0s
Build and test / android-image (push) Failing after 0s
Build and test / Android (aarch64) (push) Skipped
🐳 Windows image / Build and push (push) Failing after 0s
Build and test / windows-image (push) Failing after 0s
Build and test / Windows (x86_64, cross) (push) Skipped
The installer smoke test asserted seven model files, the number on the
day it was written; models/face has since gained the eye-state trio's
companions and the int8 detector forms, and the run on f71d7ba failed
with thirteen installed. The test now expects as many files as
package.sh's directories hold, so the next model needs no edit here.

package.sh also stages models/inpaint, which the APK and the Arch
package already carry and the Windows build did not: without
migan-512.onnx the panorama's border fill has no model on Windows.
xfeat needs nothing, it is embedded in the binary. windows.md §5.2
lists the result.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 22:35:04 +02:00
dtourolle 050b2a3914 Collapse the blank runs the library global's move left behind
Moving each group of properties and callbacks out of AppWindow left
several two- and three-line gaps where a removed block's neighbours no
longer needed separating. No declarations changed.
2026-09-20 22:32:47 +02:00
dtourolleandClaude Opus 5 195388b2e3 Adopt faces a hundred images per commit, and hold one generation per image in the shards
Benchmarks / CPU and I/O (per commit) (push) Successful in 8m16s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 4h42m34s
Build and test / Layer separation (push) Successful in 59s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 2s
🐳 Windows image / Build and push (push) Successful in 2s
Build and test / windows-image (push) Successful in 2s
Traceability / Requirement traces (push) Successful in 50s
Build and test / Android (aarch64) (push) Failing after 0s
Build and test / Windows (x86_64, cross) (push) Failing after 0s
The shard import recorded each adopted image in its own transaction:
fourteen thousand commits, and fourteen thousand turns at the write lock
that every read on the UI thread queued behind — the sync was felt as a
laggy grid and as "database is locked" from whichever writer lost the
wait. `record_detections_within` takes the caller's transaction, and the
import commits every hundred images.

The store carried every detector generation of an image — 24,123 entries
for 19,089 images on the reference library, a third of its 293 MB — when
only the strongest is ever adopted. A put now skips a pass a held one
outranks, and retires the passes it outranks from the index; sealed
shards keep their bytes, but nothing is written twice from here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 22:31:47 +02:00
dtourolle e166b64ee9 Format fill.rs, which the last release left past the width limit 2026-09-20 22:29:42 +02:00
dtourolle a773ad5c27 Merge: the developer docs under docs/dev, and the folder indexed for users first 2026-09-20 22:27:48 +02:00
dtourolle 1e472fd251 Move the library's routes and status onto its own Slint global
The scan/opening/status/error lines, offline mode and its retry, pinning a
collection offline, the sync and thumbnail-sweep state, and the callbacks
that route the grid to a rescan, a sync, another library, a panorama merge
or an export — the last of what AppWindow still carried under the
library- prefix — move onto the `Library` global started earlier on this
branch. Rust reaches them through window.global::<Library>() rather than
window.set_/get_/on_/invoke_ on the root.

library-visible is the one name that stays: it is computed from
active-page and active-view, the shell's own routing state, which a
global cannot read. AppWindow now declares no other library- property or
callback.
2026-09-20 22:22:55 +02:00
dtourolle 6bf67cefc4 Move the filter bar and the timeline onto the library's Slint global
The date-range fields and their band on the capture-time axis, the
timeline's bars, labels, scrub/pinch/pan/zoom callbacks and its
sweep/current-bucket/anchored state, and the filter bar's people chips,
mode, eyes-open toggle and gesture reference move from AppWindow onto the
`Library` global. Rust reaches them through window.global::<Library>()
rather than window.set_/get_/on_ on the root, as the earlier commits on
this branch did for the grid's cells, selection, ratings and keywords.
2026-09-20 22:07:12 +02:00
dtourolle 00e2fe6aaf Move ratings, flags and keywording onto the library's Slint global
The keywording sheet's rows and its open/assign/unassign callbacks, the
star and flag callbacks a cell click or a judgement key fires, the burst
toggle and representative-chosen callbacks, the trash-selection shortcut,
and the rating/unjudged/flag filter chips with their rating-counts model
move from AppWindow onto the `Library` global started in the previous
commit. Rust reaches them through window.global::<Library>() rather than
window.set_/get_/on_ on the root.
2026-09-20 21:57:40 +02:00
dtourolle 402dcdc24c Move the library grid's cells and selection onto their own Slint global
AppWindow carried the grid's loaded window of cells, the keyboard cursor,
drag and drop, the held-row long-press state, columns and cell size, the
scroll and viewport bookkeeping, the photo roll's pick and centre-request,
and the local-only/reorder/collection-filing gestures that act on a
selection, as properties and callbacks on the root component. That state now
lives in the `Library` global declared in library.slint, next to the structs
(LibraryCell, TimelineBar, KeywordRow, PersonChip) it and the grid's other
components already share; Rust reaches it through
window.global::<Library>() instead of window.set_/get_/on_/invoke_ on the
root, the same change collections.slint's `Collections` global made for the
sidebar.

library-visible stays on AppWindow: it is computed from active-page and
active-view, the shell's own routing state, which a global cannot read.
Everything else still prefixed library- — the timeline, the filter bar,
ratings and flags, keywording, and the routes and status lines — stays on
the window for now and moves in the commits that follow.
2026-09-20 21:42:48 +02:00
dtourolle 94b39410bc Tidy the ported drag block, and format fill.rs as master left it 2026-09-20 21:40:20 +02:00
dtourolle 2014c80e62 Merge: master at 0.13.6, with the drag-ghost file and the shared model lookup ported into the split modules 2026-09-20 21:32:21 +02:00
dtourolleandClaude Opus 5 6fd342680b Format set_indexed_at's signature
Benchmarks / CPU and I/O (per commit) (push) Successful in 8m15s
Benchmarks / Frame budget (on demand) (push) Skipped
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 2s
Build and test / Desktop (Linux) (push) Successful in 1h8m29s
🐳 Windows image / Build and push (push) Successful in 5s
Build and test / windows-image (push) Successful in 5s
Build and test / Layer separation (push) Successful in 34s
Traceability / Requirement traces (push) Successful in 50s
Build and test / Android (aarch64) (push) Failing after 0s
Build and test / Windows (x86_64, cross) (push) Failing after 0s
baed1c4 landed it over rustfmt's width; `cargo fmt --check` is the
first gate the Desktop job runs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 21:31:07 +02:00
dtourolleandClaude Opus 5 78acc73dad Regenerate the traceability matrix for the sync commits
f71d7ba and 34ac2f1 added tags without re-running the report, which
the Traceability job's "is it committed" step rejects.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 21:28:55 +02:00
dtourolleandClaude Opus 5 954246b969 Register the inference requirements, and format the fill test
Two CI gates have failed on every push since 0.13.4 and both are
fixed here.

The traceability gate rejected `FR-INF-1` as an orphan: settings.slint
and dr-ui tag it, but the four register entries inference.md §12 wrote
were never carried into requirements.md, which is the only file the
extractor reads. §3.12 and §4.10 now hold FR-INF-1..3 and NFR-INF-1
verbatim, with the acceptance milestones pointed back at inference.md.
The matrix is regenerated (188 defined, 155 covered) and the README's
"where it stands" line, which the 0.13.6 release commit skipped, says
0.13.6 and the new figures.

`cargo fmt --check` failed on the `fill_border` call in dr-pano's
padding test, which is the first thing the Desktop job runs after
installing the toolchain and why it failed within a minute.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 21:28:55 +02:00
dtourolleandClaude Opus 5 baed1c4782 Keep the peer's run marker on adopted faces, so they are not re-exported as ours
Benchmarks / CPU and I/O (per commit) (push) Successful in 8m16s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 32s
Build and test / Layer separation (push) Successful in 29s
Traceability / Requirement traces (push) Failing after 42s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 2s
🐳 Windows image / Build and push (push) Successful in 2s
Build and test / windows-image (push) Successful in 2s
Build and test / Android (aarch64) (push) Failing after 0s
Build and test / Windows (x86_64, cross) (push) Failing after 0s
Adopting an image from a peer's face shard stamped its run marker as now,
and the export reads a catalog marker newer than the shard's as a
re-index. So every adopted image went straight back out under this
device's client id: 14,100 adopted, 15,457 "newly indexed" on the next
pass, twenty-two shards of a peer's faces uploaded a second time.

merge_shard now carries the peer's indexed_at into the local index, and
the import writes that marker into face_index; where an older peer's shard
carries none, the store takes the catalog's, so the two agree either way
and the export finds nothing to send.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 21:25:53 +02:00
dtourolle fbfa891296 Regenerate the traceability matrix after the rebase 2026-09-20 21:23:23 +02:00
dtourolle aee62dc7f2 Wrap the line the docs move pushed past the width limit 2026-09-20 21:16:03 +02:00
dtourolle 84fade99ec Put the developer docs under docs/dev and index the folder for users first
docs/ had 26 developer documents flat beside the manual, and the two
audiences are very differently sized: most readers want the manual and
the gesture reference, a few want the register, the designs and the
measurements. The manual and gestures.md stay at the top; everything for
someone changing the code moves to docs/dev/, and the two documents that
name their own successors — the v0.1 milestone and the UI-refinement plan
— go to docs/dev/archive/ rather than being deleted, since both are still
cited. docs/README.md is the index, users first.

Every reference follows: code comments, Cargo manifests, the workflows,
the pre-commit hook, the bench and traceability tools (which locate the
repo root by docs/dev/requirements.md now), packaging, the Docker READMEs,
CLAUDE.md, CONTRIBUTING.md and the README. The matrix links one level
deeper and is regenerated. Links out of the moved documents into the tree
gain a level; a link checker over every Markdown file finds none broken.
2026-09-20 21:16:03 +02:00
dtourolle 3bfa73d1e1 Merge: the collections sidebar on its own Slint global 2026-09-20 21:14:00 +02:00
dtourolle 4616cb0a23 Merge: library_ui split into a module directory 2026-09-20 21:12:36 +02:00
dtourolle cc73ea3153 Split library_ui.rs into a module directory by area of behaviour
controller holds LibraryController and the window-sizing constants every
other module reads and writes through pub(super) fields, the same shape
collections_ui and develop already use. open is the launch-to-scan cycle and
the worker that checks the catalog file before either touches it. offline is
what of a collection is on this device and the prompt that offers to change
it. window fills the grid model from the catalog and drains the thumbnail
fetch, which is the piece the catalog-reads-are-proportional-to-what-changed
rule (docs/catalog.md §1) bears on most directly. sync is the background
passes that reach beyond the loaded window: the metadata sweep, the
whole-library thumbnail pass, and the exchange with the server.
ratings_keywords applies a judgement or a keyword to a selection and queues
the sidecar and XMP writes behind it. timeline is the capture-time sidebar
and the photographer's place together, kept in one file because a restored
place ends by moving the timeline marker and a scrub is a restore of one
instant, so most calls between the two would otherwise cross a module
boundary. grid wires the grid's own callbacks — the keyboard cursor,
cell-size zoom, the routes into and out of develop — and filter_bar wires
the rating, people, date and offline-scope filters, calling back into
whichever of the above owns the work a filter change triggers.

Extracted by item rather than by line range, so every doc comment and
TRACES/GESTURE annotation stayed attached to the code it describes; the
sorted set of TRACES/GESTURE lines in the new directory is identical to the
original file's. Tests moved with the code they exercise, including the
handful of fixtures — settle, model_of, with_catalog, zoom_cell, pinch_step
— that only one target module needed and so were not worth sharing through
a test_support module the way the other splits use one. Items that crossed
a new module boundary were widened from private to pub(super), narrower
than the whole-file access the original gave them; a few items already
pub(crate) for recovery_ui or presets stayed there rather than being
narrowed, since nothing needed them tightened further.

mod.rs re-exports the same surface library_ui:: callers used before, so
lib.rs and every other caller needed no change.
2026-09-20 21:10:02 +02:00
dtourolle f6ff5eabd9 Give the collections sidebar its own Slint global
AppWindow carried the sidebar's tree, its row menu, renaming, drag and
drop between rows, the trash row, and the membership sheet as ~40
properties and callbacks on the root component, in the pattern CH-1
describes and the develop screen's globals (Adjustments, Framing, Steps,
...) already replaced. Collections.* in collections.slint now holds that
state, declared next to the MembershipRow struct it and the membership
sheet both use; Rust reaches it through window.global::<Collections>()
instead of window.set_/get_/on_/invoke_ on the root.

collection-selected, collection-select, and collection-offline-menu stay
on AppWindow: library_ui.rs invokes collection-select directly and
registers collection-offline-menu's handler, and lib.rs reads
collection-selected for back-navigation, so moving them would have meant
editing library_ui.rs, which another change on this branch is splitting
into a module directory. collections-visible stays too — it is
lib.rs's panel-layout state, seeded from the saved layout before the
sidebar exists, and collections_ui never touches it. Everything prefixed
library- (the grid's drag, selection and keyword state that the sidebar's
Rust also wires for cross-feature gestures like filing a selection into
a collection) stays on the window as well, since it belongs to the
library screen, not the sidebar.
2026-09-20 21:07:25 +02:00
dtourolleandClaude Opus 5 34ac2f14d3 Sync faces and the catalog before thumbnails, so a fresh device sees its names and collections first
Benchmarks / CPU and I/O (per commit) (push) Successful in 8m24s
Benchmarks / Frame budget (on demand) (push) Skipped
🐳 Android image / Build and push (push) Successful in 4s
Build and test / android-image (push) Successful in 4s
Build and test / Desktop (Linux) (push) Failing after 35s
🐳 Windows image / Build and push (push) Successful in 2s
Build and test / windows-image (push) Successful in 2s
Build and test / Layer separation (push) Successful in 30s
Traceability / Requirement traces (push) Failing after 48s
Build and test / Android (aarch64) (push) Failing after 0s
Build and test / Windows (x86_64, cross) (push) Failing after 0s
The thumbnail stage ran first and, on a device that had just adopted its
peers' shards, spent its time re-uploading hundreds of megabytes under its
own client id while faces, people, collections and dates waited behind it.
Faces go first — the catalog merge assigns identities to faces this device
holds — then the catalog, then thumbnails.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 21:04:10 +02:00
dtourolleandClaude Opus 5 f71d7bacc6 Take the server's shards and dates when the scan completes, not after the sweep
Benchmarks / CPU and I/O (per commit) (push) Successful in 8m28s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 36s
Build and test / Layer separation (push) Successful in 28s
Traceability / Requirement traces (push) Failing after 39s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 2s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Successful in 43m11s
Build and test / Windows (x86_64, cross) (push) Failing after 41m4s
The derived sync fired only after the metadata sweep, so a fresh device
re-derived every thumbnail it scrolled past, re-detected faces and re-read
every header for hours before adopting the shards and snapshot that held
all of it. It now fires as soon as the scan completes — the first moment
the rows the merges key on exist — and the sweep starts behind it. In
steady state that pass is one listing.

The catalog merge gains a fourth half: capture metadata (captured_at,
offset, camera, lens, ISO) for images still at metadata_state < 2, matched
by oc:fileid from a remote row at 2. A date is a fact about the file's
bytes, not local state, and the snapshot already carried it. The sweep's
per-chunk query then finds nothing left, and the timeline is whole on a
fresh device without a header fetch.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 20:37:59 +02:00
dtourolle 14ed1dc410 Merge: View and Page enums in place of the six view booleans 2026-09-20 20:32:41 +02:00
dtourolle 38819da222 Replace the six view booleans with View and Page enums
app.slint carried show-launch, show-library, show-identity, show-settings,
show-import and show-merge as separate booleans, so the root component chose
what to draw with five- and six-term conjunctions and nothing stopped two of
them being true at once. Replaced with two enums: View { develop, library,
identity, launch } for which top-level screen is showing, and Page { none,
settings, import, merge } for which page, if any, is drawn over it.

Two values rather than one, because the two questions are genuinely
different. Settings, Import and Merge are reachable from more than one View
and are drawn outermost without touching it — closing one has to return to
whichever View was already current, and today that works because the
underlying property is left alone while the page sits over it. A single
View with five or more variants would need a second field remembering what
to return to; Page needs nothing to remember, since View was never
overwritten in the first place. Identity, by contrast, genuinely replaces
the window the way Launch and Library do (see the existing "like the launch
screen" comment on its `if`), so it is a View variant, not a Page.

Every `if` chain in app.slint that used to compare four, five or six
booleans now compares active-view and active-page to at most one variant
each. library-visible collapsed from a six-term conjunction to
`active-page == Page.none && active-view == View.library`.

The Rust side follows: every set_show_*/get_show_* call in library_ui.rs,
identity_ui.rs, settings_ui.rs, merge_ui.rs, import_ui.rs, launch_ui.rs and
lib.rs now reads or writes active-view or active-page instead, including
lib.rs's startup match (View.launch vs View.develop, since a Startup that
skips the launch screen used to leave both old booleans false and fall
through the chain to develop) and identity_ui's close handler, which now
writes View.library or View.develop in one call where it used to write
show-library then show-identity separately.

back_one_step needed one deliberate adjustment beyond the mechanical
rename. Identity was never represented in NavState: back had nothing to do
when Identity was opened from the library (show-library stayed true,
unread by IdentityScreen's own condition) and could only reach ToLibrary
when opened from develop, which likewise wrote a property IdentityScreen
never read — so escaping out of Identity was invisible in both cases before
this change. With a single active-view, falling into the general case
would instead overwrite the value IdentityScreen's `if` does read and close
it as an unintended side effect. back_one_step now swallows the gesture
while View.identity is current, reproducing the same "nothing visible
happens" outcome for both origins without threading identity_ui's private
came-from-library state through lib.rs for one screen.

Verified with tools/manual/drive.py against a private Xvfb and the debug
build: launch screen to library, Settings opened and closed, Identity
opened and closed (including Escape doing nothing while it is open),
develop opened from a cell and closed both by the back button and by
Escape. Screenshots under verify/.
2026-09-20 20:30:56 +02:00
dtourolle d6e9c7dc94 Merge: collections_ui split into a module directory 2026-09-20 20:30:12 +02:00
dtourolle 9c8f21b754 Split collections_ui.rs into a module directory by area of behaviour
collections_ui.rs had grown to 4,591 lines covering the sidebar controller,
the click/drag selection policy, tree refresh, the drag gesture, the trash
worker, twelve wiring functions, and the row's rename/create/context menu,
all in one file. Split into collections_ui/ with one module per area, the
way develop/ and library/ were already split on this branch:

- controller.rs: CollectionsController and the pure drop/delete/release
  decisions (decide_drop, decide_delete, decide_release, menu_detail,
  delete_warning) that a test can drive without a window.
- press.rs: PressUndo and the click-and-release selection policy
  (apply_press, select_row, commit_press, cancel_press).
- tree_sync.rs: rebuilding the sidebar from the catalog and pushing
  catalog-derived state into the grid (refresh_tree, offline_state,
  sync_lifted/sync_selection/sync_reorderable/sync_badges,
  refresh_membership, direct_holdings).
- drag.rs: the cursor bitmap (compose_drag_image, blit_scaled) and the
  hold/spring timers (arm_hold, arm_spring, should_spring,
  collapse_spring_opened) plus their delay constants.
- trash.rs: the soft delete (start_trash, start_restore, drain_trash,
  stop_trash, refresh_trash, format_bytes).
- wiring_grid.rs / wiring_tree.rs: the wire() entry point and its twelve
  wire_* functions, split in two because together they were the largest
  single piece (grid-facing selection/drag/trash vs. sidebar-facing
  navigation/create/rename/row-drag/menu/membership).
- rename_menu.rs: creating, naming and renaming collections, and the row's
  context menu (apply_rename, create_child, unique_name, open_row_menu,
  close_row_menu, close_rename).

mod.rs carries the module's own top-level doc comment, the `pub use`
re-exports for the eight items the rest of the crate reaches by
`collections_ui::` path (CollectionsController, wire, refresh_tree,
sync_badges, sync_selection, select_row, commit_press, cancel_press), and a
shared `test_support` for the one fixture (`ids`) more than one file's
tests needed. Every item that only crossed a boundary within this module,
not out of it, was narrowed to `pub(super)` rather than kept at the
crate-wide `pub` a single file gave it for free.

Extracted with a brace-aware pass that kept each item's own leading doc
comment and attributes attached to it, and tests moved with the code they
exercise; every TRACES/GESTURE comment lands on the same code it did
before. No file outside the new directory changed — lib.rs's `mod
collections_ui;` resolves to the directory automatically, and every
outside caller's `collections_ui::` path still resolves through mod.rs's
re-exports.
2026-09-20 20:29:21 +02:00
dtourolle bd7d75522d Merge: the format fix from docs-layout 2026-09-20 20:26:56 +02:00
dtourolle 4ef1b74f2f Wrap the line the docs move pushed past the width limit 2026-09-20 20:26:54 +02:00
dtourolle 681486196e Release 0.13.6
Benchmarks / CPU and I/O (per commit) (push) Failing after 6m17s
Benchmarks / Frame budget (on demand) (push) Skipped
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 2s
Build and test / Desktop (Linux) (push) Failing after 29s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 1s
Build and test / Layer separation (push) Successful in 26s
Traceability / Requirement traces (push) Failing after 42s
Build and test / Android (aarch64) (push) Failing after 2m48s
Build and test / Windows (x86_64, cross) (push) Failing after 4m22s
2026-09-20 20:19:48 +02:00
dtourolle f5d0d57574 Regenerate the traceability matrix after the rebase 2026-09-20 20:19:43 +02:00
dtourolle c96e670356 Re-record the panorama for the trained filler; the scene waits for the preview and for the DNG instead of guessing 2026-09-20 20:19:01 +02:00
dtourolle 9c556364fa Pad an open void's canvas to a tile: the merge page's preview is shorter than one, and filled nothing 2026-09-20 20:19:01 +02:00
dtourolle 8d72cabff5 Ship the border filler trained against MI-GAN's own discriminator: texture in the deep bands, level with stock on LPIPS 2026-09-20 20:19:01 +02:00
dtourolleandClaude Opus 5 8a90d888d5 Probe and compile for a package's models, not only the user's
Benchmarks / CPU and I/O (per commit) (push) Failing after 6m30s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 53s
Build and test / Layer separation (push) Successful in 26s
Traceability / Requirement traces (push) Failing after 40s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Failing after 2m19s
Build and test / Windows (x86_64, cross) (push) Failing after 3m5s
`inference::init` listed the models from the user's shared directory
alone, while the app loads them from there or from the package's
`/usr/share/darkroom/models`. On a fresh package install the probe found
"no model to probe with", stayed on the CPU, and compiled nothing. Both
now resolve each file with the same search, `library::shared_model`.

The PKGBUILD names the ONNX Runtime packages as optional dependencies,
since the app loads one from /usr/lib if present.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 19:39:05 +02:00
dtourolle e86edef47c Merge: run() reduced to construction and startup, the develop wiring in its own module 2026-09-20 19:28:23 +02:00
dtourolle 0aff5e6c8c Merge: the collections, identity, merge, launch and import wiring lifted into named functions 2026-09-20 19:28:12 +02:00
dtourolle 5458314083 Split import_ui::wire into one function per section
The 188-line wire() registered the import page's callbacks in four
comment-delimited sections. Lift each into its own fn: wire_opening_and_closing,
wire_choosing_a_source, wire_options, wire_running. context stays generic
over C on each of the three functions that use it, matching how survey()
and start() already take it (impl Fn, implicitly Sized) rather than coercing
it to a trait object, which would have needed ?Sized added to those two
unrelated functions for no benefit.
2026-09-20 19:23:21 +02:00
dtourolleandClaude Opus 5 39a22875b1 Add the MIGraphX rung for AMD GPUs
Benchmarks / CPU and I/O (per commit) (push) Failing after 6m20s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 45s
Build and test / Layer separation (push) Successful in 26s
Traceability / Requirement traces (push) Failing after 46s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Failing after 2m19s
Build and test / Windows (x86_64, cross) (push) Failing after 3m2s
Measured on a Radeon RX 7900 XT against Arch's onnxruntime-rocm 1.29
(docs/inference.md §1.3): MIGraphX fp16 runs the detectors at 2.4–3.4 ms
against 10–58 ms on the CPU provider, the inpainter at 8 ms against 514,
with a 15–135 s compile per graph the first time and under a second from
its cache after. A compiling rung on TensorRT's terms, wired the same way.

The ROCm execution provider is gone (removed in ONNX Runtime 1.23), so the
AMD ladder is MIGraphX then the CPU, with no non-compiling rung between.

MIGraphX is registered through the runtime's generic key/value entry
point rather than ort's builder: 1.29 reads the legacy options struct for
its precision flags only, and the compiled-program cache directory
(`migraphx_model_cache_dir`) only travels the generic way. The provider's
cache key omits the precision, so f32 and fp16 programs get their own
directories. The probe fingerprint now includes the provider libraries
beside the runtime and the ROCm version, since a distribution's CPU and
ROCm builds are the same file at the same path.

`status().failed` reports only the rungs above the selection, so an AMD
desktop's About line says why MIGraphX won rather than that the NVIDIA
providers are not in the build.

Two examples: `ep_probe` times each provider cold and from cache, and
`ladder` drives `init` as the app does to watch the first-run sequence.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 19:23:00 +02:00
dtourolle cb7b9d71c4 Split lib.rs::run into named construction and wiring functions
run() was 2,724 lines that built every controller, owned the develop
session and its render loop, and registered every develop-screen callback
inline — the state CH-1 in docs/dev/code-health.md describes. This is the
mechanical split CH-1 calls for, done in one pass rather than section by
section since develop_ui.rs only compiles once lib.rs stops registering
those callbacks itself.

develop_ui.rs is new: a DevelopWiring struct holding the session, the rows
model, the redraw/render closures and the other controllers' handles, and
one wire_* function per section run() used to contain — presets, export,
the adjustment panel, film, undo/redo, zoom/pan/crop, rotation/flips/
straightening, navigation and peaking — called in the order run()
registered them. window travels as its own parameter throughout rather
than living on the struct, because the generated AppWindow type is not
Clone; every other field is an owned clone so each function's body could
be pasted from run() unchanged.

lib.rs::run is now under 300 lines: construction and startup only, calling
named functions for the window and its diagnostics/inference wiring, the
launch/library/collections/identity/settings screens, import and merge,
the render loop (render_now/redraw/show), the remote open path, the
window chrome (resize, layout class, panel toggles, back gesture), and
develop_ui::wire for the rest. Every TRACES and GESTURE comment moved with
the code it annotates.

Two latent type errors surfaced while restructuring rather than being
introduced by it: activity and display were being passed around as bare
ActivityLog/DisplayWatch instead of the Rc<...> their constructors
actually return, which only worked before because nothing needed to name
the type explicitly.
2026-09-20 19:17:22 +02:00
dtourolle beb7a5eac0 Split launch_ui::wire into one function per section
The 225-line wire() registered the launch screen's callbacks in nine
comment-delimited sections. Lift them into fn's, merging a few adjacent
ones that were only a handful of lines each: wire_sign_in covers both the
browser flow and the app-password fallback (the same button in two forms),
wire_choose_folder_and_open covers opening the folder browser and the
final "Open library" press, since both are short and sit back to back.
wire_use_folder, wire_sign_out, wire_formats, wire_folder_picker_navigation
and wire_copy_url stay as they were sectioned. wire() calls each in the
original order and keeps the closing render() call, which every section
relies on having run once at startup.
2026-09-20 19:11:43 +02:00
dtourolle 262ed2553c Split merge_ui::wire into one function per section
The 312-line wire() registered the panorama page's callbacks in three
comment-delimited sections. Lift each into its own fn: wire_start (the
"Merge to panorama" press, its fetch, and the DARKROOM_START_MERGE dev
entry point — all three only ever used together), wire_decision (confirm,
projection and border chips, the fill knobs), and wire_stop_and_leave
(abandon, close). wire() itself now just calls the three in order; S, C
and F stay generic on wire_start since sources/context/on_done are used
nowhere else.
2026-09-20 19:00:31 +02:00
dtourolle 8d08ffd7b7 Split identity_ui::wire into one function per feature
The 837-line wire() had almost no section comments, unlike its siblings, so
the seams had to be found by reading it rather than following markers. Lift
each into its own fn: wire_dials (the two grouping sliders), wire_navigation
(open/close/switch person — needs models() too, for the missing-model
banner), wire_rename_and_merge (a rename and the namesake offer it can
raise), wire_face_actions (pick/confirm/reject/split, the grid's own
actions), wire_grouping_preview, wire_recluster, wire_indexing (the shared
launcher behind Index/Re-index plus Stop), and wire_coverage_and_ignore.
wire() keeps the generic-to-trait-object coercions and the eyes_available
closure, since most of the above need it, and calls each function in the
original order. The reload! macro moved from inside wire() to module scope,
dedented, since macro_rules is scoped textually and every extracted function
uses it.
2026-09-20 18:49:43 +02:00
dtourolle 8540518022 Split collections_ui::wire into one function per section
The 1,457-line wire() registered every collections-sidebar and grid-drag
callback in one function, sectioned only by comment. Lift each section
into its own fn: wire_selection (split further into wire_selection and
wire_selection_filing, since the original section ran to 375 lines),
wire_drag, wire_trash, wire_trash_from_grid, wire_tree_navigation,
wire_create, wire_rename, wire_remove, wire_row_drag (the tree row's own
hold-drag, which the original "remove" comment's span covered but which
is really a separate feature), wire_row_menu, and wire_membership.
wire() itself now just coerces the shared closures to trait objects and
calls each in the original order. visible_ids is coerced to
Rc<dyn Fn() -> Vec<ImageId>> at the top, alongside on_scope_changed and
session, so the new functions take plain trait objects instead of
threading a generic parameter through every one of them.
2026-09-20 18:36:18 +02:00
dtourolle b952f5976a Merge: the wire sections lifted into named functions, beside the module splits 2026-09-20 18:26:24 +02:00
dtourolle a1d511fd4b Split library.rs into library/ by area of behaviour
library.rs was 7,729 lines wiring together everything "open a remote
library" touches: scanning, pulling other devices' judgements out of
sidecars found along the way, writing local edits back out to the
sidecar outbox, pushing/reloading XMP by hand, fetching and prefetching
thumbnails and originals, generating thumbnails locally, the metadata
and thumbnail background sweeps, on-disk paths for the catalog and
model files, and reading the grid's cells, spans and rating filter.
Same motivation as the develop.rs split (docs/dev/code-health.md CH-1):
a pure, no-behaviour-change move into one file per area, each under
about 1,500 lines.

Tracing actual call sites rather than trusting the file's physical
layout mattered here: `persist`, `load_folder_etags`, `pull_sidecars`,
`load_sidecar_etags`, `record_sidecar_read` and `apply_judgement` sit
textually beside the XMP push/reload functions but are called only
from `run_scan` (pulling a device's own past judgements out of the
sidecars a scan just walked), so they went to scan.rs and not xmp.rs.
`cells` came out at over 1,800 lines once its tests moved with it and
split further into cells.rs (windowed reads, trash, ordinals) and
spans.rs (collection scope, manual reordering, the capture-time
histogram) -- ten submodules rather than the nine first planned.

Previously-private items reached from a sibling module became
`pub(super)`, narrower than the whole-crate reachability one file gave
them. Tests moved with the code they test; the two test fixtures used
across more than one file (`scanned`, and develop.rs's
`session_with_a_left_half_subject` in the matching commit) joined the
shared `test_support` module alongside the existing `entry`/
`with_images`/`image_ids` helpers. `mod.rs` re-exports every module's
public items under `library::`, including the `pub(crate)`
`test_support` module `repairs.rs` reads its fixtures from, so no file
outside `library` needed a change.

The previous commit split develop.rs the same way; taken alone it left
dr-ui without library.rs, so that intermediate commit does not build on
its own. This one restores it.
2026-09-20 18:21:43 +02:00
dtourolle 050c2c9d16 Split develop.rs into develop/ by area of behaviour
develop.rs had grown to 9,327 lines covering everything the develop
session does: opening a photograph, the parameter-row and curve-widget
panel model, mask viewing and editing, mask creation and the rasteriser
that turns a mask stack into GPU arrays, spot repairs, scene
segmentation, framing and zoom, white-balance sampling, rendering and
film choice, and the undo/snapshot history. docs/dev/code-health.md
CH-1 names dr-ui's lack of a view layer as the reason every feature
kept landing in a handful of files; this is the first of the two pure
splits it recommends as easy, no-behaviour-change wins independent of
that larger rework.

The boundaries follow the file's own sections (several were already
marked off with comment headers) and the seams a full read turned up
underneath them -- mask storage/rasterisation turned out to be a
distinct concern from mask viewing and editing, and rows/tabs/curves
from each other, so those split further than the headers alone
suggested. Each module stays under about 1,500 lines. Struct fields
and the handful of helper methods now called from a sibling module
became `pub(super)`, which is strictly narrower than the whole-crate
reachability a single file gave them; nothing gained visibility outside
`develop`. Tests moved with the code they test, including the few
cases where a helper one file's tests needed was itself only defined
in another's -- those became shared fixtures in `mod.rs` alongside the
`headless`/`read_back`/`grey_session` helpers that already worked that
way. `mod.rs` re-exports every item `develop::` callers outside this
module used before, so lib.rs, masks_ui.rs and the rest needed no
changes.
2026-09-20 18:21:26 +02:00
dtourolle bd3b993b90 Split library_ui::wire into one function per section
wire() registered every grid callback in one 1,214-line function behind
four section comments, two of which were themselves far over 300 lines
with no further markers. Each fenced section becomes its own function,
called from wire() in the original order with the section's own comment
kept as its doc comment:

- "the keyboard cursor (FR-CULL-4)" (496 lines) splits at its own topic
  breaks into wire_grid_cursor_and_zoom (cursor movement, cell zoom and
  pinch), wire_grid_sync_and_load (explicit sync/thumbnail requests and
  the reloads a changed viewport, column count or scroll position
  trigger), wire_timeline (the capture-time sidebar) and wire_grid_routes
  (grid/launch/develop navigation and a manual rescan).
- "ratings and flags" and "keywords" were already under 300 lines and
  become one function each.
- "the filter bar" (447 lines) splits into wire_filter_ratings_and_people
  (which keeps the section's own comment), wire_filter_dates and
  wire_filter_scope_and_offline.

The three callbacks registered before the first section comment
(on_settings_xmp_reload, on_library_cell_clicked, on_library_roll_pick)
and the trailing crate::recovery_ui::wire call stay directly in wire(),
since neither is inside a fenced section.
2026-09-20 17:25:00 +02:00
dtourolle 59605f9fbb Split masks_ui::wire into one function per section
wire() registered every mask-panel callback in one 778-line function
behind section comments. Each of the seven fenced sections (computing
the region map, refining a subject's mask, dragging a gradient,
selecting on the photograph, the stack, the edge treatment, adding
layers) becomes its own function, called from wire() in the original
order with the section's own comment kept as its doc comment. The
"adding layers" section was itself over 300 lines and had no further
section markers inside it, so it is split at its own natural seam
between painting/viewing a mask (wire_layers_paint) and working the
parts and add-mask buttons (wire_layers_parts); the second half gets an
introductory doc line since there was no comment of its own to reuse.
Locals declared just for one section's closures (running, refining,
dragging) move into that section's function instead of staying in
wire().
2026-09-20 17:24:42 +02:00
dtourolle 04949741c1 Split settings_ui::wire into one function per section
wire() registered every settings-page callback in one 416-line function,
fenced only by section comments. Each fenced section (opening and
closing, cache, faces, export, reset) is now its own private function
that wire() calls in the same order, with the section's own comment kept
as its doc comment. on_budget_changed and on_open are coerced to trait
objects at the top of wire() so the new functions take a plain
Rc<dyn Fn> rather than needing their own generic parameter, with no
change in the closures registered or the order they are registered in.
2026-09-20 17:24:28 +02:00
dtourolle 5b4ad11853 Manual: nested collections, and the ghost drawn as it should be
Benchmarks / CPU and I/O (per commit) (push) Failing after 6m16s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 42s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 2s
Build and test / Layer separation (push) Successful in 26s
Traceability / Requirement traces (push) Failing after 38s
Build and test / Android (aarch64) (push) Failing after 2m20s
Build and test / Windows (x86_64, cross) (push) Failing after 3m3s
A scene that makes a parent, nests two collections in it by drag and by
the menu, files frames into a child and opens the parent to see it count
both; stills of the tree and of the menu. The collections recording is
re-made now that the bitmap under the cursor is the photograph.

drive.py grows a multi-leg drag: a diagonal with much vertical in it is
taken by the grid's Flickable as a scroll before the DragArea can claim
it, so a drag to the sidebar goes sideways first.
2026-09-20 16:27:18 +02:00
dtourolle 6b1aac477d Put the developer docs under docs/dev and index the folder for users first
docs/ had 26 developer documents flat beside the manual, and the two
audiences are very differently sized: most readers want the manual and
the gesture reference, a few want the register, the designs and the
measurements. The manual and gestures.md stay at the top; everything for
someone changing the code moves to docs/dev/, and the two documents that
name their own successors — the v0.1 milestone and the UI-refinement plan
— go to docs/dev/archive/ rather than being deleted, since both are still
cited. docs/README.md is the index, users first.

Every reference follows: code comments, Cargo manifests, the workflows,
the pre-commit hook, the bench and traceability tools (which locate the
repo root by docs/dev/requirements.md now), packaging, the Docker READMEs,
CLAUDE.md, CONTRIBUTING.md and the README. The matrix links one level
deeper and is regenerated. Links out of the moved documents into the tree
gain a level; a link checker over every Markdown file finds none broken.
2026-09-20 16:20:15 +02:00
dtourolle 2afc2a7890 Report the grid's column count on creation, not only on change
`changed columns` fires on a change, and a first evaluation is not one:
a grid built after the window had settled at its size never said how
wide it was, so Rust placed month headings for the one column it was
told about at start-up — every month began a row, and was announced
wherever its first cell fell, mid-row included. Opening a collection
showed "October 2025" stranded over a row of August.
2026-09-20 15:58:44 +02:00
dtourolle 6507593715 Hand the drag ghost to the renderer through a file, so it draws
The bitmap under the cursor was a solid red rectangle. Slint's drag
overlay uploads the image as a texture, draws it and drops the texture in
one call; with the wgpu FemtoVG renderer the drop is immediate and the
draw is deferred to the flush, so the frame binds femtovg's placeholder —
which is red. An image with a cache key survives in the texture cache
until after the flush, and only a path gives one. So the composite goes
to the data directory's scratch as a PNG and comes back through
load_from_path; one file per drag, removed when the drag ends. A
workaround for Slint 1.17.1, written up as one beside the code.
2026-09-20 15:58:44 +02:00
dtourolle 08727cff5a Release 0.13.5
Benchmarks / CPU and I/O (per commit) (push) Failing after 6m16s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 45s
Build and test / Layer separation (push) Successful in 25s
Traceability / Requirement traces (push) Failing after 38s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 2s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Failing after 2m20s
Build and test / Windows (x86_64, cross) (push) Failing after 3m1s
2026-09-20 15:24:47 +02:00
dtourolle f4c3f425dd Show the picker correcting something in the manual
The recording sampled a red brick wall and moved the sliders by three
units, which at GIF size is a click that does nothing. The scene now
drags the frame cold first and picks a white air conditioner, so the
correction is visible and the picker's being absolute - set from the
photograph, not from where the sliders were - is what the picture
shows. The text says so, and says a blown highlight is refused.
2026-09-20 15:24:27 +02:00
dtourolle 12f8990e09 Average a patch under the white balance picker, not one photosite
The probe's comment said a 192px render "averages a small neighbourhood
into each of its pixels". It does not: the composed shader fetches the
source at one position per output pixel - nearest for an unrotated
frame, four photosites blended otherwise - so the probe was a point
sample of a noisy sensor, and two painted-white air conditioners on the
same wall answered +37 and -50.

The tap is now narrowed to the patch of the canvas around the click, a
couple of percent of its width and square on screen, and rendered at
64x64 with interpolation forced on, which puts a sample on every sensor
pixel under it at any ordinary zoom. The samples are averaged, with the
void and clipped ones left out rather than allowed to pull the mean, and
fewer than half surviving is refused. compose_camera_probe takes the
patch; the merge's compose_camera_linear keeps its nearest sampling. The
readback shrinks from six megabytes to sixty-four kilobytes.

A frame of alternating warm and cool columns, averaging neutral, moves
the controls by at most two units; a point sample swung them to sixty.
2026-09-20 15:24:25 +02:00
dtourolle 8a897bbc01 Release 0.13.4
Benchmarks / CPU and I/O (per commit) (push) Failing after 6m22s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 46s
Build and test / Layer separation (push) Successful in 29s
Traceability / Requirement traces (push) Failing after 40s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Failing after 2m22s
Build and test / Windows (x86_64, cross) (push) Failing after 3m5s
2026-09-20 14:03:58 +02:00
dtourolle a03e082fe2 Point the manual's white balance scene at "pick", and record it again
The scene clicked 40px to the right of "pick", on "reset", so the
recording showed a neutral group being reset and a click on the wall
that panned. Re-recorded with the picker fixed: the word lights, the
sample moves temperature and tint, and Before shows what it corrected.
2026-09-20 13:41:17 +02:00
dtourolle 4576499c3b Refuse a clipped highlight as a neutral
Sampling the overcast sky on a Canon 6D frame set tint to -100 and
temperature to -15 for a patch the canvas showed as pure white. A clipped
photosite is sensor white, not a colour: every channel stopped counting,
so what the tap hands back is the as-shot multipliers themselves, which
are strongly magenta, and the solver dutifully drove green to its stop.
The display shader already fades such a pixel to a neutral of the same
brightness before any operation runs, so the picker was balancing against
something the photographer could not see.

The probe now refuses a sample with any channel at or above the onset the
shader fades from, the way the solver already refuses black. The threshold
is one constant, CLIP_ONSET, formatted into the shader and read by the
probe, so the two cannot drift apart.
2026-09-20 13:41:17 +02:00
dtourolle 2f47087223 Measure the white balance probe in camera RGB, where the gains multiply
Pressing "pick" and clicking a near-neutral wall on a Canon 6D frame set
tint to -77 and turned the whole photograph green. The white balance
operation runs first in the chain, on camera RGB, before the body's base
curve and colour matrix; the probe was read off a display render after
all three, and the solve treated that sRGB triple as if the gains
multiplied it directly. On a JPEG the two spaces coincide, which is why
the existing tests passed while the picker was broken on every raw file.

The probe now reads the camera-space tap a merge stitches from, composed
under the edit's own framing so a fraction of the canvas is a fraction of
the probe, and puts the as-shot balance on itself - exactly the value the
operation's gains are about to multiply. No operations run in the tap, so
nothing has to be stripped and restored, and the display target is left
alone, so a sample that found nothing usable no longer needs a redraw.

A raw-frame test with the 6D's matrix and a typical as-shot balance
samples a warm grey and asserts the rendered pixel comes back neutral; it
fails on the previous probe.
2026-09-20 13:41:11 +02:00
dtourolle 764ad55ead Stand each person in the grouping pass by at most 100 references
Benchmarks / CPU and I/O (per commit) (push) Failing after 6m23s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 55s
Build and test / Layer separation (push) Successful in 27s
Traceability / Requirement traces (push) Failing after 54s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Failing after 2m21s
Build and test / Windows (x86_64, cross) (push) Failing after 3m5s
Every face the user has ruled on entered the pass as an anchor, and the
scan is exhaustive by design (`dr_face::neighbours`), so a person with
750 confirmed faces cost 750 comparisons against every other face in
the library — and the cost of a library grew with how well it was
named. Most of those comparisons said nothing new: thirty frames from
one afternoon are one point of view, not thirty, and a face that
matches one of them matches the rest.

Each person now enters through at most 100 of their anchored faces
(`dr_face::references`). Eligible are those whose raw embedding is at
least 15 long — one above the gallery floor, since a reference speaks
for someone rather than merely being admitted — with an unmeasured
length admitted as it is everywhere else. From those, the set spanning
the greatest volume is chosen greedily: the longest vector first, then
at each step the face with the largest component orthogonal to the
chosen so far. That is pivoted Gram–Schmidt, and the product of the
residuals it picks is the Gram determinant, so the greedy step is the
exact greedy on the objective. A near-duplicate of a chosen face has
no residual and is passed over; the one profile shot among two hundred
frontal frames is taken early; faces inside the span of the chosen add
no volume and are not taken to fill the cap.

The faces not chosen keep their confirmations and are not touched by
the pass — they stay in the anchor map, so it never releases them —
they are simply not compared. A person none of whose faces is long
enough is still stood for, by their longest, rather than losing their
anchor and having their next face filed as a stranger. Under the cap
nothing changes: every eligible face stands, and the short ones stay
in as the probes they were.

At the reference library's 3,851 confirmations the scan shrinks by
about a fifth; at 15,000 it is a fifth of what it was.
2026-09-20 13:29:16 +02:00
dtourolle 301e6f3828 Release 0.13.3
Benchmarks / CPU and I/O (per commit) (push) Failing after 6m21s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 46s
Build and test / Layer separation (push) Successful in 26s
Traceability / Requirement traces (push) Failing after 39s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Failing after 2m21s
Build and test / Windows (x86_64, cross) (push) Failing after 3m8s
2026-09-20 13:01:41 +02:00
dtourolle a437363bd6 Schema V20: put the mis-spelled run markers right
The markers the previous commit stops writing are already in the
catalogs — 2 on the desktop, 429 on the tablet — and in the shards
both have exchanged. Renaming them to the faces' own id with a fresh
time is what makes the export send each image again, under an entry
newer than the empty one `held_model` would otherwise pick. Where the
old write had inserted its marker beside the right one, the wrong one
goes and the right one is refreshed for the same reason: its entry in
the shards is older than the empty one.

Images V14 left with faces and no marker are not touched. That state is
the quality pass's cue, and the fixed write marks them correctly when
it reaches them.

Checked against copies of both real catalogs: the desktop renames 2,
the tablet deletes 429, both in under 200 ms.
2026-09-20 13:00:59 +02:00
dtourolle 065872bec5 Keep a run marker under the detector that found the faces
`record_updates` — the write behind the quality, eye and crop passes —
re-marked the image as indexed under the pipeline the pass ran as, and
left the faces it had updated under the id of the detector that found
them. On a desktop set to Thorough that put `scrfd_10g+w600k_mbf` over
faces spelled `w600k_mbf`; on the tablet, `scrfd_10g_i8+w600k_mbf` over
faces it had adopted from the desktop's thorough pass.

Every reader takes the marker and the faces to agree. `marker_under`
reads the marker as the detector having examined the image, so the
upgrade repair never revisits it. The shard store keys each face by
its pipeline id, so `export_to_shards` selects an image's faces by the
marker's id, finds none, and sends an entry that says the thorough
detector looked and found nothing — over photographs with named faces
on them. The desktop's shard index holds 54 such entries beside real
faces; the tablet's eye pass over the faces it had adopted made 430
more, and both devices have exchanged them. `held_model` takes the
newest entry for an image, which is the empty one. Nothing has been
lost yet only because the two spellings of the thorough detector rank
equal and neither side adopts the other's; a third device, or either
one after a reinstall, would adopt "nothing here" for 484 images. And
the desktop's eye pass is 4,739 images from doing the same to every
face from before V14 — which are the ones that only exist on the
desktop, and would then never reach anywhere.

The marker now takes the id the faces carry; the pass's own id is used
only when it dropped the last of them and there is no detector left to
name. A stale marker under another spelling of the same embedder is
removed in the same transaction, so one embedder has one marker.
2026-09-20 13:00:58 +02:00
dtourolle 0ed38ada28 Adopt a peer's unmeasured faces instead of refusing them
The tablet showed a fraction of each person: 681 of the desktop's 3,851
confirmations, and none of Ian's 746, Catherine's 626 or my own 480.
Every face that existed on both devices agreed on who it was, and the
people rows were identical — the merge was fine. The missing 3,170
confirmations were on faces the tablet did not hold at all: the
desktop's 16,080 faces from the original detector, on 4,310 images,
detected before schema V14 kept the quality reading.

Those faces were in shards the tablet had already downloaded, in
August's export. `import_from_shards` looked at them on every sync pass
and declined each one, because a face without a quality reading was
"work this device cannot finish": adopting it would write the run
marker, and the marker was what stopped an image being looked at again.
That was true when it was written and has not been since the quality
repair existed — that pass lists its work by `f.quality IS NULL`, not by
the marker, exactly as the eye pass does, and faces without an eye
reading were already adopted on that reasoning.

The refusal had no exit. V14 had deleted the markers of every image
holding such faces so the quality pass would find them, and
`export_to_shards` walks the markers, so the desktop never re-exported
them either; the unmeasured August copies were the only ones there
would ever be. The tablet's answer was to queue all 17,727 images for a
re-detection of its own, a fetch of the whole library, while holding
the faces on disk.

Adopt them. The receiving device's quality pass measures them when it
reaches them, and the desktop's confirmations match onto them by box
overlap on the next catalog merge. The test that asserted the refusal
now asserts the adoption and that the image is still owed to the pass.
2026-09-20 12:58:41 +02:00
dtourolle 695d5ec304 Correct four claims in the README against the tree
Benchmarks / CPU and I/O (per commit) (push) Failing after 6m26s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 46s
Build and test / Layer separation (push) Successful in 28s
Traceability / Requirement traces (push) Failing after 40s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 2s
🐳 Windows image / Build and push (push) Successful in 2s
Build and test / windows-image (push) Successful in 2s
Build and test / Android (aarch64) (push) Failing after 2m21s
Build and test / Windows (x86_64, cross) (push) Failing after 3m5s
Eighteen declared operations, not fifteen; JPEG XL is an export format;
the grid does not filter by keyword, only the catalog's query can; and a
panorama's provenance is a sidecar beside the composite, not a history
step in it.
2026-09-20 11:57:17 +02:00
212 changed files with 37620 additions and 33008 deletions
+3 -3
View File
@@ -1,6 +1,6 @@
name: Benchmarks
# The suite docs/requirements.md §8 has been promising since it was written:
# The suite docs/dev/requirements.md §8 has been promising since it was written:
# "an automated benchmark suite against a synthetic 50k catalog, run per-commit
# … A regression beyond stated tolerance fails the build."
#
@@ -29,7 +29,7 @@ name: Benchmarks
# every commit to establish, every time, that this runner has no GPU. It
# runs on demand (Actions → Run workflow) so that a runner that *does*
# have one can be pointed at it, and the numbers it produces belong in
# docs/frame-budget.md by hand, as they already are.
# docs/dev/frame-budget.md by hand, as they already are.
on:
push:
@@ -183,7 +183,7 @@ jobs:
- name: Frame budget (FR-DSP-3)
run: cargo test --release -p dr-gpu --test frame_budget -- --nocapture
# The instrument behind docs/frame-budget.md. It exits non-zero with no
# The instrument behind docs/dev/frame-budget.md. It exits non-zero with no
# adapter, which is right for a tool a person runs deliberately and wrong
# for a job that usually has none — hence continue-on-error. Its table is
# in the log for whoever asked for this run; the committed numbers are
+8 -3
View File
@@ -322,7 +322,7 @@ jobs:
env:
CARGO_TARGET_DIR: target-android
# Absent secrets mean a debug signature, which is what a fork or a
# branch build should get. Set all three (see docs/android-signing.md)
# branch build should get. Set all three (see docs/dev/android-signing.md)
# and the same job produces a release-signed APK instead.
ANDROID_KEYSTORE_BASE64: ${{ secrets.ANDROID_KEYSTORE_BASE64 }}
KEYSTORE_PASS: ${{ secrets.ANDROID_KEYSTORE_PASSWORD }}
@@ -370,7 +370,7 @@ jobs:
# TRACES: FR-PLAT-WIN-3
# The Windows executable and its installer, cross-built from Linux
# (docs/windows.md §7). No Windows machine anywhere in this job: what it
# (docs/dev/windows.md §7). No Windows machine anywhere in this job: what it
# can prove is that the binary links, is a Windows executable with no
# MinGW runtime imports, starts under Wine, and that the installer installs
# and uninstalls under Wine. What it cannot prove — a Vulkan device, a
@@ -454,7 +454,12 @@ jobs:
wine "$SETUP" /S 2>/dev/null
INST=$(echo "$HOME"/.wine/drive_c/users/*/AppData/Local/Programs/DarkRoom)
ls "$INST"
[ "$(ls "$INST/models" | wc -l)" = 7 ] || { echo "FAIL: expected 7 model files"; exit 1; }
# As many files as package.sh stages: everything but the READMEs in
# the directories it copies. A literal here went stale the first
# time a model was added.
WANT=$(find models/face models/scene models/inpaint -maxdepth 1 -type f ! -name README.md | wc -l)
GOT=$(ls "$INST/models" | wc -l)
[ "$GOT" = "$WANT" ] || { echo "FAIL: expected $WANT model files, installed $GOT"; exit 1; }
wine reg query 'HKCU\Software\Microsoft\Windows\CurrentVersion\Uninstall\DarkRoom' 2>/dev/null \
| grep -q DisplayVersion || { echo "FAIL: no uninstall registry key"; exit 1; }
wine "$INST/darkroom.exe" --version 2>/dev/null | grep -q '^darkroom-desktop ' \
+5 -5
View File
@@ -7,7 +7,7 @@ name: Traceability
# fail its own threshold. Two rules follow, and the extractor's own tests
# enforce both:
#
# 1. Denominators are parsed from docs/requirements.md at run time.
# 1. Denominators are parsed from docs/dev/requirements.md at run time.
# 2. Coverage is |traced ∩ defined| / |defined|, never a raw traced count.
#
# This job is static analysis of source comments plus markdown parsing, so it
@@ -74,11 +74,11 @@ jobs:
run: |
set -e
cargo run -q -p traceability -- report
if ! git diff --quiet docs/traceability.md; then
if ! git diff --quiet docs/dev/traceability.md; then
echo ""
echo "docs/traceability.md is out of date."
echo "docs/dev/traceability.md is out of date."
echo "Run: cargo run -p traceability -- report"
git diff --stat docs/traceability.md
git diff --stat docs/dev/traceability.md
exit 1
fi
@@ -127,4 +127,4 @@ jobs:
- name: Summary
if: always()
run: head -30 docs/traceability.md || true
run: head -30 docs/dev/traceability.md || true
+4 -4
View File
@@ -26,7 +26,7 @@ fi
# The artefacts are generated from the tree, so regenerating them because one
# was itself edited would be circular.
case "$(tr -d '[:space:]' <<< "${staged}")" in
docs/traceability.md | docs/gestures.md | ui/dr-ui/src/gesture_book.rs)
docs/dev/traceability.md | docs/gestures.md | ui/dr-ui/src/gesture_book.rs)
exit 0
;;
esac
@@ -41,9 +41,9 @@ if ! cargo run -q -p traceability -- report >/dev/null 2>&1; then
exit 0
fi
if ! git diff --quiet -- docs/traceability.md; then
git add docs/traceability.md
echo "pre-commit: regenerated docs/traceability.md and staged it"
if ! git diff --quiet -- docs/dev/traceability.md; then
git add docs/dev/traceability.md
echo "pre-commit: regenerated docs/dev/traceability.md and staged it"
fi
# The gesture vocabulary, same discipline.
+26 -2
View File
@@ -2,12 +2,12 @@
Notes for anyone — person or agent — changing this code. They record what
went wrong once and what the fix looked like, so the same shape is not
written again. Requirements live in `docs/requirements.md`; this file is
written again. Requirements live in `docs/dev/requirements.md`; this file is
about habits, not features.
## Catalog reads: work is proportional to what changed, never to library size
`docs/catalog.md §1` states the rule. These are the ways it was broken on
`docs/dev/catalog.md §1` states the rule. These are the ways it was broken on
the Identity screen, found when every confirm click cost half a second on a
24k-image library (2026-09-19), and what each fix looked like.
@@ -89,6 +89,30 @@ three `405`s before each `MOVE`. The backend now remembers the collections
it has confirmed (`known_dirs`) for its lifetime, which is one job. When a
per-file operation has a per-batch precondition, satisfy it once.
## Providers: read the runtime's source for the version on disk, not the binding
Two things the MIGraphX rung (2026-09-20) got wrong before it was measured
right, both because `ort`'s builder was trusted to mean what its method
names say.
**A binding's option builder may fill a struct the runtime no longer
reads.** `ep::MIGraphX::with_save_model` sets fields of the legacy
`OrtMIGraphXProviderOptions`; ONNX Runtime 1.29 reads that struct for the
precision flags and ignores the rest, so every session compiled for 40 s
and the cache directory went nowhere. The option that works
(`migraphx_model_cache_dir`) exists only in the generic key/value
registration, which `session::migraphx` calls on the API table directly.
Before wiring a provider option, fetch the provider's source at the
runtime's exact version and find where the option is *read*.
**A provider's cache key may leave out what you are varying.** MIGraphX
keys a compiled program on graph, GPU and its own version — not precision.
The first fp16 measurement built in 0.3 s and matched f32 to the tenth of a
millisecond, because it had loaded the f32 program. A "from cache" build
that is suspiciously fast on the first run of a new configuration is a key
collision, not a fast provider; give each precision its own directory (the
engine does) and check the cache directory gained a file.
## Measuring
`cargo run --release -p dr-ui --example identity_bench -- CATALOG THUMBS`
+11 -11
View File
@@ -90,18 +90,18 @@ break it by accident:
cargo run --release -p dr-bench -- check
```
That is the benchmark suite (`docs/requirements.md` §8), which builds a
That is the benchmark suite (`docs/dev/requirements.md` §8), which builds a
synthetic 50,000-image catalog and fails the build if a performance target is
missed or a measurement has drifted past its tolerance. It runs on every push in
its own workflow. [`docs/benchmarks.md`](docs/benchmarks.md) says what it
its own workflow. [`docs/dev/benchmarks.md`](docs/dev/benchmarks.md) says what it
measures, what it deliberately does not, and how to read a failure. If you have
touched the catalog, the decoder, the thumbnail store or the exporter, run it
before you send.
## Requirements and traceability
[`requirements.md`](docs/requirements.md) is the register of record.
[`traceability.md`](docs/traceability.md) is generated from `TRACES:` tags in
[`requirements.md`](docs/dev/requirements.md) is the register of record.
[`traceability.md`](docs/dev/traceability.md) is generated from `TRACES:` tags in
the source and must never be hand-edited:
```rust
@@ -124,7 +124,7 @@ Note that it tracks line numbers, so a change that only moves code still moves
the matrix. Never regenerate it with a stale prebuilt binary.
**One convention that the tooling cannot enforce.** A tag proves that a tag
exists, not that the code under it does the thing — `docs/code-health.md`
exists, not that the code under it does the thing — `docs/dev/code-health.md`
CH-4 has the details, and two requirements currently read as covered on the
strength of plumbing a future feature would use. So: **close a requirement
with a test that would fail if the behaviour were removed.** Coverage that
@@ -163,12 +163,12 @@ One commit per change. If you fixed two things, that is two commits.
| Document | Read it when |
|---|---|
| [`core/dr-pipeline/ops/README.md`](core/dr-pipeline/ops/README.md) | Adding or changing a develop operation — start here regardless |
| [`docs/architecture.md`](docs/architecture.md) | Anything touching the render path, catalog or sync |
| [`docs/code-health.md`](docs/code-health.md) | Deciding what to work on; grades each seam by what it costs |
| [`docs/benchmarks.md`](docs/benchmarks.md) | A change that could plausibly cost time or memory |
| [`docs/technical-debt.md`](docs/technical-debt.md) | Something looks wrong — check it was not chosen |
| [`docs/distribution.md`](docs/distribution.md) | Packaging a build, or adding a permission to one |
| [`docs/requirements.md`](docs/requirements.md) | Reference, not reading |
| [`docs/dev/architecture.md`](docs/dev/architecture.md) | Anything touching the render path, catalog or sync |
| [`docs/dev/code-health.md`](docs/dev/code-health.md) | Deciding what to work on; grades each seam by what it costs |
| [`docs/dev/benchmarks.md`](docs/dev/benchmarks.md) | A change that could plausibly cost time or memory |
| [`docs/dev/technical-debt.md`](docs/dev/technical-debt.md) | Something looks wrong — check it was not chosen |
| [`docs/dev/distribution.md`](docs/dev/distribution.md) | Packaging a build, or adding a permission to one |
| [`docs/dev/requirements.md`](docs/dev/requirements.md) | Reference, not reading |
`technical-debt.md` is the one to check before "fixing" anything surprising.
It records compromises that were deliberate, each with the reasoning and a
Generated
+27 -25
View File
@@ -1221,7 +1221,7 @@ checksum = "f27ae1dd37df86211c42e150270f82743308803d90a6f6e6651cd730d5e1732f"
[[package]]
name = "darkroom-android"
version = "0.13.2"
version = "0.14.0"
dependencies = [
"android_logger",
"dr-plat",
@@ -1234,7 +1234,7 @@ dependencies = [
[[package]]
name = "darkroom-desktop"
version = "0.13.2"
version = "0.14.0"
dependencies = [
"anyhow",
"dr-plat",
@@ -1408,7 +1408,7 @@ checksum = "d8b14ccef22fc6f5a8f4d7d768562a182c04ce9a3b3157b91390b52ddfdf1a76"
[[package]]
name = "dr-bench"
version = "0.13.2"
version = "0.14.0"
dependencies = [
"anyhow",
"dr-catalog",
@@ -1425,7 +1425,7 @@ dependencies = [
[[package]]
name = "dr-catalog"
version = "0.13.2"
version = "0.14.0"
dependencies = [
"dr-face",
"dr-plat",
@@ -1440,7 +1440,7 @@ dependencies = [
[[package]]
name = "dr-decode"
version = "0.13.2"
version = "0.14.0"
dependencies = [
"dr-types",
"env_logger",
@@ -1454,7 +1454,7 @@ dependencies = [
[[package]]
name = "dr-export"
version = "0.13.2"
version = "0.14.0"
dependencies = [
"dr-decode",
"dr-gpu",
@@ -1473,7 +1473,7 @@ dependencies = [
[[package]]
name = "dr-face"
version = "0.13.2"
version = "0.14.0"
dependencies = [
"dr-inference-engine",
"env_logger",
@@ -1486,7 +1486,7 @@ dependencies = [
[[package]]
name = "dr-film"
version = "0.13.2"
version = "0.14.0"
dependencies = [
"log",
"serde",
@@ -1495,7 +1495,7 @@ dependencies = [
[[package]]
name = "dr-gpu"
version = "0.13.2"
version = "0.14.0"
dependencies = [
"bytemuck",
"dr-decode",
@@ -1513,8 +1513,9 @@ dependencies = [
[[package]]
name = "dr-inference-engine"
version = "0.13.2"
version = "0.14.0"
dependencies = [
"env_logger",
"libloading",
"log",
"ort",
@@ -1527,7 +1528,7 @@ dependencies = [
[[package]]
name = "dr-ingest"
version = "0.13.2"
version = "0.14.0"
dependencies = [
"dr-plat",
"dr-types",
@@ -1539,7 +1540,7 @@ dependencies = [
[[package]]
name = "dr-lens"
version = "0.13.2"
version = "0.14.0"
dependencies = [
"lensfun",
"log",
@@ -1547,7 +1548,7 @@ dependencies = [
[[package]]
name = "dr-pano"
version = "0.13.2"
version = "0.14.0"
dependencies = [
"dr-decode",
"dr-inference-engine",
@@ -1561,7 +1562,7 @@ dependencies = [
[[package]]
name = "dr-pipeline"
version = "0.13.2"
version = "0.14.0"
dependencies = [
"dr-types",
"log",
@@ -1570,7 +1571,7 @@ dependencies = [
[[package]]
name = "dr-plat"
version = "0.13.2"
version = "0.14.0"
dependencies = [
"android-native-keyring-store",
"dr-types",
@@ -1586,7 +1587,7 @@ dependencies = [
[[package]]
name = "dr-preset-xmp"
version = "0.13.2"
version = "0.14.0"
dependencies = [
"dr-pipeline",
"log",
@@ -1596,7 +1597,7 @@ dependencies = [
[[package]]
name = "dr-segment"
version = "0.13.2"
version = "0.14.0"
dependencies = [
"dr-inference-engine",
"env_logger",
@@ -1609,7 +1610,7 @@ dependencies = [
[[package]]
name = "dr-sync"
version = "0.13.2"
version = "0.14.0"
dependencies = [
"async-trait",
"dr-plat",
@@ -1623,7 +1624,7 @@ dependencies = [
[[package]]
name = "dr-sync-folder"
version = "0.13.2"
version = "0.14.0"
dependencies = [
"async-trait",
"dr-sync",
@@ -1635,7 +1636,7 @@ dependencies = [
[[package]]
name = "dr-sync-nextcloud"
version = "0.13.2"
version = "0.14.0"
dependencies = [
"async-trait",
"dr-decode",
@@ -1657,7 +1658,7 @@ dependencies = [
[[package]]
name = "dr-thumbs"
version = "0.13.2"
version = "0.14.0"
dependencies = [
"dr-types",
"jpeg-encoder",
@@ -1669,7 +1670,7 @@ dependencies = [
[[package]]
name = "dr-types"
version = "0.13.2"
version = "0.14.0"
dependencies = [
"serde",
"serde_json",
@@ -1678,7 +1679,7 @@ dependencies = [
[[package]]
name = "dr-ui"
version = "0.13.2"
version = "0.14.0"
dependencies = [
"anyhow",
"async-trait",
@@ -1706,6 +1707,7 @@ dependencies = [
"jni 0.22.4",
"log",
"ndk-context",
"png",
"pollster",
"reqwest",
"rusqlite",
@@ -1720,7 +1722,7 @@ dependencies = [
[[package]]
name = "dr-xmp"
version = "0.13.2"
version = "0.14.0"
dependencies = [
"dr-types",
"log",
@@ -7021,7 +7023,7 @@ checksum = "8df9b6e13f2d32c91b9bd719c00d1958837bc7dec474d94952798cc8e69eeec3"
[[package]]
name = "traceability"
version = "0.13.2"
version = "0.14.0"
dependencies = [
"anyhow",
"serde",
+2 -2
View File
@@ -29,7 +29,7 @@ members = [
]
[workspace.package]
version = "0.13.2"
version = "0.14.0"
edition = "2021"
rust-version = "1.92"
license = "GPL-3.0-or-later"
@@ -48,7 +48,7 @@ dr-export = { path = "core/dr-export" }
dr-face = { path = "core/dr-face", default-features = false }
dr-film = { path = "core/dr-film" }
# `tract` on by default so a test binary can open a session with nothing
# installed; the apps add `native` to look for a runtime file (docs/inference.md §3).
# installed; the apps add `native` to look for a runtime file (docs/dev/inference.md §3).
dr-inference-engine = { path = "core/dr-inference-engine" }
dr-ingest = { path = "core/dr-ingest" }
dr-gpu = { path = "core/dr-gpu" }
+25 -25
View File
@@ -16,12 +16,12 @@ is still missing.
or one a Nextcloud client keeps in virtual-files mode, where a placeholder
is treated as the photograph rather than as a one-byte file — or at a
Nextcloud account directly. The grid is virtualised, ordered by capture
time with a timeline beside it, and filtered by rating, flag, keyword,
person and whether the file is here. Ratings, keywords, collections and a
time with a timeline beside it, and filtered by rating, flag, person and
whether the file is here. Ratings, keywords, collections and a
trash that survives a crash mid-operation. Card ingest. Bursts fold. Face
detection and identity, with the index syncing between devices.
**Developing.** Fifteen declared operations fused into one compute
**Developing.** Eighteen declared operations fused into one compute
dispatch, plus the neighbourhood work that cannot be: clarity, texture,
capture sharpening, noise reduction, lens correction, spectral film
simulation. Crop and straighten, spot repair, and local adjustments over
@@ -33,11 +33,11 @@ judging what is recoverable. Named presets; XMP sidecars other editors read.
**Panoramas.** Select the frames, align, choose a projection, fill the
ragged border rather than crop it, and the composite lands beside its
sources as a DNG with the merge as the first step in its history.
sources as a DNG, with a sidecar recording what it was merged from.
[![Twelve hand-held frames aligned on a cylinder](docs/manual/media/panorama-aligned.png)](docs/manual/README.md#merging-a-panorama)
**Export.** JPEG, PNG, AVIF, 8- and 16-bit TIFF, with resize, output
**Export.** JPEG, PNG, AVIF, JPEG XL, 8- and 16-bit TIFF, with resize, output
sharpening, a naming template and a colour space — to a folder here or back
into the library.
@@ -52,7 +52,7 @@ texture directly — no readback between the GPU and the screen.
|---|---|---|
| Arch Linux | [`packaging/PKGBUILD`](packaging/PKGBUILD) — `makepkg -si` | Built from every release |
| Android | The APK from each CI run, or `./docker/android/package.sh --install` | Runs on a tablet; F-Droid not yet submitted |
| Windows | `DarkRoom-<version>-x86_64-setup.exe`, cross-built by CI ([windows.md](docs/windows.md)) | Verified under Wine only; unsigned |
| Windows | `DarkRoom-<version>-x86_64-setup.exe`, cross-built by CI ([windows.md](docs/dev/windows.md)) | Verified under Wine only; unsigned |
| Flatpak | [`packaging/flatpak/`](packaging/flatpak/) | Manifest in tree; choosing a library does not yet work in the sandbox |
Or build it. Git LFS is required for the model weights, and the toolchain
@@ -76,30 +76,30 @@ controls, its place in the chain and its tests.
## Where it stands
**0.13.2**, sixteen tagged releases in. 184 numbered requirements in
scope, 84% of them claimed by code and [traced to it](docs/traceability.md);
**0.14.0**, twenty-one tagged releases in. 188 numbered requirements in
scope, 82% of them claimed by code and [traced to it](docs/dev/traceability.md);
the rest are written down rather than merely absent.
**Not built:** plugins (post-v1, [D12](docs/requirements.md)), compare and
**Not built:** plugins (post-v1, [D12](docs/dev/requirements.md)), compare and
survey culling, AI denoise, tiled and progressive rendering, HDR merge and
focus stacking, most of the Android platform integration beyond running,
and the Flatpak's library chooser. The performance targets are half
verified: the per-commit benchmark suite §8 requires exists for everything
that does not need a frame — the catalog, the scan, the thumbnails — and
not yet for the render path, so a regression there fails nothing.
[outstanding.md](docs/outstanding.md) is the list, with the reasoning for
[outstanding.md](docs/dev/outstanding.md) is the list, with the reasoning for
each.
**The one deliberate compromise worth knowing about before reading
anything else:** the Android develop view reads its frame back through the
CPU, because zero-copy there needs wgpu's Vulkan swapchain and that tears a
portrait window on a tablet whose panel is mounted landscape. It is debt,
not a revision of the rule — [technical-debt.md TD-1](docs/technical-debt.md)
not a revision of the rule — [technical-debt.md TD-1](docs/dev/technical-debt.md)
has the measurements and the three things any one of which would remove it.
## Documentation
For someone using it:
[docs/README.md](docs/README.md) is the index. The short version, for someone using it:
| | |
|---|---|
@@ -111,21 +111,21 @@ For someone changing it:
| | |
|---|---|
| [CONTRIBUTING.md](CONTRIBUTING.md) | How to land a first change without reading the rest |
| [requirements.md](docs/requirements.md) | What the software must do — the numbered register, and the decisions |
| [architecture.md](docs/architecture.md) | How it is built — crates, the GPU pipeline, the data model, sync |
| [technical-debt.md](docs/technical-debt.md) | Compromises taken deliberately, each with the condition that retires it |
| [outstanding.md](docs/outstanding.md) | What is not built, and whether that is a decision or a gap |
| [code-health.md](docs/code-health.md) | What a contribution costs, per seam, measured |
| [traceability.md](docs/traceability.md) | Generated: which requirement is claimed by which file |
| [requirements.md](docs/dev/requirements.md) | What the software must do — the numbered register, and the decisions |
| [architecture.md](docs/dev/architecture.md) | How it is built — crates, the GPU pipeline, the data model, sync |
| [technical-debt.md](docs/dev/technical-debt.md) | Compromises taken deliberately, each with the condition that retires it |
| [outstanding.md](docs/dev/outstanding.md) | What is not built, and whether that is a decision or a gap |
| [code-health.md](docs/dev/code-health.md) | What a contribution costs, per seam, measured |
| [traceability.md](docs/dev/traceability.md) | Generated: which requirement is claimed by which file |
Designs, one per subsystem:
[segmentation](docs/segmentation.md) and [mask editing](docs/mask-editing.md) ·
[spot removal](docs/spot-removal.md) · [panorama](docs/panorama.md) ·
[faces](docs/faces.md) · [inference](docs/inference.md) ·
[storage and sync](docs/storage.md) · [catalog](docs/catalog.md) ·
[display and extension](docs/display-and-extension.md) ·
[navigation](docs/ui-navigation.md) · [distribution](docs/distribution.md) ·
[windows](docs/windows.md) · [benchmarks](docs/benchmarks.md).
[segmentation](docs/dev/segmentation.md) and [mask editing](docs/dev/mask-editing.md) ·
[spot removal](docs/dev/spot-removal.md) · [panorama](docs/dev/panorama.md) ·
[faces](docs/dev/faces.md) · [inference](docs/dev/inference.md) ·
[storage and sync](docs/dev/storage.md) · [catalog](docs/dev/catalog.md) ·
[display and extension](docs/dev/display-and-extension.md) ·
[navigation](docs/dev/ui-navigation.md) · [distribution](docs/dev/distribution.md) ·
[windows](docs/dev/windows.md) · [benchmarks](docs/dev/benchmarks.md).
## Licence
+5 -5
View File
@@ -240,7 +240,7 @@ fn android_main(app: slint::android::AndroidApp) {
///
/// **Face weights are absent from the repository by design.** The InsightFace
/// grant is research-only and incompatible with this project's licence
/// (docs/faces.md §2), so a desktop user fetches them, runs
/// (docs/dev/faces.md §2), so a desktop user fetches them, runs
/// `tools/fix-face-model-shapes.sh` over them, and drops the result in. A build
/// that carries none is the ordinary case and face indexing simply stays off.
///
@@ -321,17 +321,17 @@ fn unpack_bundled_models(app: &slint::android::AndroidApp) {
// before it reports the tab available.
//
// Three detectors, because which one runs is a setting
// (`FaceDetector`, docs/faces.md §12.3) and a tablet has no other way to
// (`FaceDetector`, docs/dev/faces.md §12.3) and a tablet has no other way to
// obtain the one it was not shipped with. Twenty megabytes of APK for
// the choice; the embedder is the same for all three.
//
// Then the three eye-state models (docs/faces.md §17): landmarks, open
// Then the three eye-state models (docs/dev/faces.md §17): landmarks, open
// or closed, sunglasses. The app indexes without them; with them the
// eyes-open filter has something to read, and a tablet has no other way
// to get them either.
//
// The int8 forms beside the three detectors are what the Hexagon runs
// (docs/inference.md §5); the engine loads the sibling when the probe
// (docs/dev/inference.md §5); the engine loads the sibling when the probe
// chose that rung and ignores it otherwise.
const BUNDLED: [(&std::ffi::CStr, &str); 14] = [
(c"models/scrfd_500m_640.onnx", "scrfd_500m_640.onnx"),
@@ -420,7 +420,7 @@ fn unpack_bundled_models(app: &slint::android::AndroidApp) {
// on a first launch they were not on disk until this line. The runtime
// is in the APK's native library directory beside `libdarkroom.so`,
// which is also where Qualcomm's DSP loader has to be pointed for the
// Hexagon skel (docs/inference.md §3, §8).
// Hexagon skel (docs/dev/inference.md §3, §8).
dr_ui::inference::init(native_library_dir().into_iter().collect());
}
+3 -3
View File
@@ -21,7 +21,7 @@ fn main() -> anyhow::Result<()> {
// build made on a machine that cannot run the application — the Linux CI
// producing the Windows binary, checked under Wine — has an exit that
// proves the executable starts without opening a window or touching the
// user's directories (docs/windows.md §6).
// user's directories (docs/dev/windows.md §6).
if std::env::args().nth(1).as_deref() == Some("--version") {
println!("darkroom-desktop {}", env!("CARGO_PKG_VERSION"));
return Ok(());
@@ -63,7 +63,7 @@ fn main() -> anyhow::Result<()> {
// Before the window: the probe runs on its own thread and the first
// frame does not wait for it, but the models a background job asks for
// should already know where the runtime is (docs/inference.md §4).
// should already know where the runtime is (docs/dev/inference.md §4).
dr_ui::inference::init(runtime_dirs());
dr_ui::run(paths)?;
@@ -78,7 +78,7 @@ fn main() -> anyhow::Result<()> {
/// Where a desktop package may have put `libonnxruntime`, most specific
/// first. None of these existing is the tract build, which is a complete
/// application and not an error (docs/inference.md §3).
/// application and not an error (docs/dev/inference.md §3).
///
/// `DARKROOM_ORT_DIR` is for a developer pointing at a runtime that is not
/// installed — the wheel's `capi` directory, say. Then beside the executable
+1 -1
View File
@@ -315,7 +315,7 @@ fn full_library(
// The three phases, separately, because "a regroup takes n seconds" does
// not tell anyone which half to optimise — and the answer differs between
// a desktop and a tablet (docs/faces.md §9).
// a desktop and a tablet (docs/dev/faces.md §9).
{
let dim = candidates.first().map(|c| c.embedding.len()).unwrap_or(0);
let flat: Vec<f32> = candidates
+1 -1
View File
@@ -82,7 +82,7 @@
//! Grouping has no natural `subject_id`: it is a property of a *run* of frames,
//! so a per-image job would rebuild the world once per photograph. It is
//! therefore a debounced library-level pass, for exactly the reasons
//! docs/catalog.md §10.2 gives for face clustering, and [`regroup`] is the whole
//! docs/dev/catalog.md §10.2 gives for face clustering, and [`regroup`] is the whole
//! of it — one ordered walk, no per-pair comparison beyond adjacent frames.
//!
//! # Grouping is not hiding
+210 -33
View File
@@ -2,7 +2,7 @@
//! Face data as sealed shards, so a second device does not re-index the library.
//!
//! Indexing a 23,500-image library is on the order of two hours of CPU
//! (docs/faces.md §12.2). It is also **byte-identical on every device**: the
//! (docs/dev/faces.md §12.2). It is also **byte-identical on every device**: the
//! same model over the same proxy produces the same embedding. Paying for it
//! once per account rather than once per device is the whole point of this
//! module, and it is the same bargain the thumbnail store already makes.
@@ -178,6 +178,67 @@ impl FaceShardStore {
.flatten()
}
/// The other pipelines this file is held under that share `model_id`'s
/// embedder — the generations a put of `model_id` may supersede.
fn siblings(&self, file_id: u64, model_id: &str) -> Vec<String> {
let mut stmt = match self.index.prepare(&format!(
"SELECT model_id FROM entries
WHERE file_id = ?1 AND model_id != ?2 AND {} = ?3",
crate::faces::embedder_sql("model_id")
)) {
Ok(s) => s,
Err(_) => return Vec::new(),
};
stmt.query_map(
rusqlite::params![
file_id as i64,
model_id,
crate::faces::embedder_of(model_id)
],
|r| r.get::<_, String>(0),
)
.map(|rows| rows.filter_map(|r| r.ok()).collect())
.unwrap_or_default()
}
/// Whether a pass this file is already held under outranks `model_id`,
/// so a put of `model_id` would add a generation nobody would adopt.
pub fn outranked(&self, file_id: u64, model_id: &str) -> bool {
use dr_types::FaceDetector;
let Some(incoming) = FaceDetector::for_model_id(model_id) else {
return false;
};
self.siblings(file_id, model_id)
.iter()
.filter_map(|m| FaceDetector::for_model_id(m))
.any(|held| held.outranks(incoming))
}
/// Forget the index entries for generations of this file that `model_id`
/// outranks. The bytes stay where they are — a sealed shard is
/// immutable — but the store stops offering them, and a later export or
/// merge writes nothing for them again.
fn supersede(&self, file_id: u64, model_id: &str) -> Result<(), CatalogError> {
use dr_types::FaceDetector;
let Some(incoming) = FaceDetector::for_model_id(model_id) else {
return Ok(());
};
for held in self.siblings(file_id, model_id) {
let weaker = FaceDetector::for_model_id(&held).is_some_and(|h| incoming.outranks(h));
if weaker {
self.index.execute(
"DELETE FROM entries WHERE file_id = ?1 AND model_id = ?2",
rusqlite::params![file_id as i64, held],
)?;
self.index.execute(
"DELETE FROM faces_meta WHERE file_id = ?1 AND model_id = ?2",
rusqlite::params![file_id as i64, held],
)?;
}
}
Ok(())
}
pub fn contains(&self, file_id: u64, model_id: &str) -> bool {
self.index
.query_row(
@@ -233,6 +294,15 @@ impl FaceShardStore {
faces: &[SharedFace],
indexed_at: Option<i64>,
) -> Result<u32, CatalogError> {
// One generation per image per embedder. A store carried every pass
// — 24,123 entries for 19,089 images on the reference library, a
// third of its 293 MB — and only the strongest was ever adopted.
// A weaker pass arriving after a stronger one is not written; a
// stronger one arriving retires the weaker from the index.
if self.outranked(file_id, model_id) {
return Ok(0);
}
self.supersede(file_id, model_id)?;
let incoming = faces
.iter()
.map(|f| BYTES_PER_FACE + if f.crop.is_empty() { 0 } else { BYTES_PER_CROP })
@@ -474,15 +544,28 @@ impl FaceShardStore {
rusqlite::OpenFlags::SQLITE_OPEN_READ_ONLY | rusqlite::OpenFlags::SQLITE_OPEN_NO_MUTEX,
)?;
let mut q =
src.prepare("SELECT file_id, model_id, faces_found, source_edge FROM indexed")?;
let images: Vec<(i64, String, i64, i64)> = q
.query_map([], |r| Ok((r.get(0)?, r.get(1)?, r.get(2)?, r.get(3)?)))?
// The peer's marker travels with the image: it is what lets
// `import_from_shards` record the adoption under the time the peer
// indexed it, and so what keeps `export_to_shards` from reading the
// adoption as a re-index and sending the peer's faces back out under
// this device's name. A shard from before the column has none.
let mut q = src.prepare(&format!(
"SELECT file_id, model_id, faces_found, source_edge, {} FROM indexed",
match has_column(&src, "indexed", "indexed_at") {
Ok(true) => "indexed_at",
_ => "NULL",
}
))?;
let images: Vec<(i64, String, i64, i64, Option<i64>)> = q
.query_map([], |r| {
Ok((r.get(0)?, r.get(1)?, r.get(2)?, r.get(3)?, r.get(4)?))
})?
.collect::<Result<_, _>>()?;
let mut adopted = 0;
for (file_id, model_id, _found, edge) in images {
if self.contains(file_id as u64, &model_id) {
for (file_id, model_id, _found, edge, indexed_at) in images {
if self.contains(file_id as u64, &model_id) || self.outranked(file_id as u64, &model_id)
{
continue;
}
let mut fq = src.prepare(&format!(
@@ -501,12 +584,27 @@ impl FaceShardStore {
let faces: Vec<SharedFace> = fq
.query_map(rusqlite::params![file_id, &model_id], read_shared_face)?
.collect::<Result<_, _>>()?;
self.put_image(file_id as u64, &model_id, edge as u32, &faces)?;
self.put_image_at(file_id as u64, &model_id, edge as u32, &faces, indexed_at)?;
adopted += 1;
}
Ok(adopted)
}
/// Record when the catalog indexed a held image, for an entry that
/// arrived without a marker — a peer's shard from before the column.
pub fn set_indexed_at(
&self,
file_id: u64,
model_id: &str,
at: i64,
) -> Result<(), CatalogError> {
self.index.execute(
"UPDATE entries SET indexed_at = ?3 WHERE file_id = ?1 AND model_id = ?2",
rusqlite::params![file_id as i64, model_id, at],
)?;
Ok(())
}
/// Read back everything held for one image.
pub fn get_image(
&self,
@@ -798,8 +896,21 @@ pub fn import_from_shards(
})?
.collect::<Result<_, _>>()?;
/// Images per write transaction. Large enough that fourteen thousand
/// adoptions are a hundred and forty commits rather than fourteen
/// thousand; small enough that a read on the UI thread, queued behind
/// the lock, waits a fraction of a second and not the whole import.
const CHUNK: usize = 100;
let mut adopted = 0;
let mut tx = conn.unchecked_transaction()?;
let mut in_chunk = 0;
for (file_id, image_id, local) in candidates {
if in_chunk == CHUNK {
tx.commit()?;
tx = conn.unchecked_transaction()?;
in_chunk = 0;
}
let Some(held) = store.held_model(file_id as u64, model_id) else {
continue;
};
@@ -820,20 +931,22 @@ pub fn import_from_shards(
let Some((faces, edge)) = store.get_image(file_id as u64, &held)? else {
continue;
};
// A peer that embedded before the quality was kept has done work this
// device cannot finish: the number exists only at embedding time, and
// adopting the faces would write the run marker that keeps them from
// ever being measured (schema V14). Left for this device's own pass —
// or for the peer's, whose re-export replaces these.
// A face the peer embedded before its quality was kept (schema V14)
// is adopted with the reading missing, exactly as one without an eye
// reading is. The measuring passes find their work by the NULL
// column, not by the run marker (`dr_ui::repairs`, `faces_needing`),
// so adopting costs the reading nothing and this device's own pass
// fills it.
//
// A missing *eye* reading is not the same case and is adopted. The
// measuring pass finds those by the NULL, not by the marker, so
// adopting the faces costs the reading nothing (schema V16) — and a
// peer that has no eye models may be the only one that has done the
// detection at all.
if faces.iter().any(|f| f.quality.is_none()) {
continue;
}
// This used to refuse such faces, on the reasoning that the marker
// would stop them ever being measured — true before the quality
// repair existed, and wrong after. What it cost: V14 had dropped the
// markers of every image holding such faces, so the peer never
// re-exported them, and the only copies in the shards were the
// unmeasured ones. A tablet holding shards with 4,310 of the
// desktop's images and 3,170 of its confirmations declined every one
// of them, showed a fraction of each person, and queued the whole
// library for a re-detection of its own instead.
let local: Vec<crate::faces::DetectedFace> = faces
.into_iter()
.map(|f| crate::faces::DetectedFace {
@@ -856,15 +969,42 @@ pub fn import_from_shards(
})
.collect();
crate::faces::record_detections(
conn,
crate::faces::record_detections_within(
&tx,
dr_types::ImageId(image_id as u64),
&held,
edge,
&local,
)?;
// The peer's marker, not this moment. `record_detections` stamps the
// run as now, and `export_to_shards` reads a marker newer than the
// shard's as a re-index — so every adopted image went straight back
// out as this device's own work: 14,100 adopted, 15,457 "newly
// indexed" on the next pass, and twenty-two shards of a peer's faces
// uploaded again under a second name. Where the peer's shard carried
// no marker, the store takes the catalog's, so the two agree either
// way and the export sees nothing to send.
match store.indexed_at(file_id as u64, &held) {
Some(theirs) => {
tx.execute(
"UPDATE face_index SET indexed_at = ?3
WHERE image_id = ?1 AND model_id = ?2",
rusqlite::params![image_id, held, theirs],
)?;
}
None => {
let ours: i64 = tx.query_row(
"SELECT indexed_at FROM face_index WHERE image_id = ?1 AND model_id = ?2",
rusqlite::params![image_id, held],
|r| r.get(0),
)?;
store.set_indexed_at(file_id as u64, &held, ours)?;
}
}
adopted += 1;
in_chunk += 1;
}
tx.commit()?;
Ok(adopted)
}
@@ -1100,6 +1240,32 @@ mod tests {
assert!(!s.contains(1, "lvface"));
}
/// One generation per image per embedder: a stronger detector's pass
/// retires a weaker one from the index, and a weaker pass arriving after
/// a stronger is not written at all.
#[test]
fn a_stronger_pass_retires_a_weaker_one_and_a_weaker_is_not_added() {
let dir = tempdir();
let mut s = FaceShardStore::open(&dir).unwrap();
s.put_image(1, "w600k_mbf", 1024, &[face(1, 1)]).unwrap();
s.put_image(1, "scrfd_10g+w600k_mbf", 1024, &[face(1, 2)])
.unwrap();
assert!(s.contains(1, "scrfd_10g+w600k_mbf"));
assert!(!s.contains(1, "w600k_mbf"), "the fast pass was not retired");
assert_eq!(s.len(), 1, "faces_meta still counts the retired pass");
s.put_image(1, "scrfd_2.5g+w600k_mbf", 1024, &[face(1, 3)])
.unwrap();
assert!(
!s.contains(1, "scrfd_2.5g+w600k_mbf"),
"a weaker pass was added"
);
assert_eq!(
s.held_model(1, "w600k_mbf").as_deref(),
Some("scrfd_10g+w600k_mbf")
);
}
#[test]
fn re_storing_an_image_replaces_rather_than_doubling_it() {
let dir = tempdir();
@@ -1320,6 +1486,11 @@ mod catalog_round_trip {
assert!((got[0].landmarks[2].0 - 0.15).abs() < 1e-5);
let emb = faces::embeddings(&b, "w600k_mbf").unwrap();
assert!(emb.iter().any(|e| e.embedding[0] == 1));
// And what B adopted is not B's work: its next export sends nothing.
// Adopting used to stamp the run as now, so every adopted image went
// back out under B's name as a re-index.
assert_eq!(export_to_shards(&b, &mut store_b, "w600k_mbf").unwrap(), 0);
}
/// The desktop switched to a stronger detector part-way through the
@@ -1365,11 +1536,12 @@ mod catalog_round_trip {
);
}
/// A face a peer embedded without measuring it is work this device
/// cannot finish, and adopting it would write the marker that stops it
/// ever being measured. The image stays outstanding instead.
/// A face a peer embedded without measuring it is adopted all the same,
/// and left on this device's quality pass by its missing reading. Refusing
/// it was what stranded every confirmation the desktop had made on faces
/// from before V14: the tablet held the shards and would not use them.
#[test]
fn a_peers_unmeasured_faces_are_left_for_this_device_to_index() {
fn a_peers_unmeasured_faces_are_adopted_and_left_for_the_quality_pass() {
let b = device(&[(90, 5001), (91, 5002)]);
let mut store = FaceShardStore::open(&tempdir("unmeasured")).unwrap();
let shared = |file_id: u64, quality: Option<f32>| SharedFace {
@@ -1395,13 +1567,18 @@ mod catalog_round_trip {
.put_image(5002, "w600k_mbf", 2560, &[shared(5002, Some(19.0))])
.unwrap();
assert_eq!(import_from_shards(&b, &store, "w600k_mbf").unwrap(), 1);
assert_eq!(import_from_shards(&b, &store, "w600k_mbf").unwrap(), 2);
let cov = faces::coverage(&b, "w600k_mbf").unwrap();
assert_eq!(cov.indexed, 1);
assert_eq!(cov.outstanding(), 1, "the unmeasured image was adopted");
assert!(faces::for_image(&b, dr_types::ImageId(90))
.unwrap()
.is_empty());
assert_eq!(cov.indexed, 2);
assert_eq!(cov.outstanding(), 0, "the unmeasured image was refused");
let got = faces::for_image(&b, dr_types::ImageId(90)).unwrap();
assert_eq!(got.len(), 1);
assert_eq!(got[0].quality, None, "a reading was invented");
// Still owed to the measuring pass, which lists by the column.
assert_eq!(
faces::count_needing(&b, "w600k_mbf", "f.quality IS NULL").unwrap(),
1
);
}
#[test]
+133 -11
View File
@@ -1,7 +1,7 @@
//! TRACES: FR-CULL-8 | FR-CULL-9 | FR-CULL-10 | FR-CULL-11 | FR-CULL-12 | NFR-SEC-5
//! People and faces: what was detected, who it is, and who said so.
//!
//! The storage half of docs/faces.md. `dr-face` finds faces and turns them into
//! The storage half of docs/dev/faces.md. `dr-face` finds faces and turns them into
//! 512 numbers; this module is where those numbers acquire an identity, and
//! where the user's corrections outrank the model's guesses.
//!
@@ -110,7 +110,7 @@ pub struct DetectedFace {
/// Raw rather than unit length, so the length ([`Self::quality`]) is in
/// the blob and not only beside it. Readers re-normalise on load.
pub embedding: Vec<u8>,
/// Source pixels across the aligned crop (docs/faces.md §7).
/// Source pixels across the aligned crop (docs/dev/faces.md §7).
pub crop_px: f32,
/// Length of the raw embedding before normalisation — the model's own
/// reading of how recognisable the crop was, and the gate on whether
@@ -211,7 +211,7 @@ pub use dr_face::Calibration;
/// the same face in the same photograph, for carrying an identity across a
/// re-detection.
///
/// Set at the reference library's P≈0.95 line (docs/faces.md §9's table:
/// Set at the reference library's P≈0.95 line (docs/dev/faces.md §9's table:
/// 0.449), which is far above anything two different people in one frame
/// reach and below what one face re-embedded from a better crop of itself
/// does. The number is only ever asked about *overlapping* boxes on *one*
@@ -279,7 +279,25 @@ pub fn record_detections(
faces: &[DetectedFace],
) -> Result<Vec<FaceId>, CatalogError> {
let tx = conn.unchecked_transaction()?;
let ids = record_detections_within(&tx, image_id, model_id, source_edge, faces)?;
tx.commit()?;
Ok(ids)
}
/// [`record_detections`] inside a transaction the caller owns.
///
/// For a caller recording many images at once — the shard import adopts
/// fourteen thousand in one pass — where a commit per image is fourteen
/// thousand fsyncs and fourteen thousand turns at the write lock that every
/// read on the UI thread queues behind. `unchecked_transaction` cannot nest,
/// so the batching has to be offered here rather than wrapped from above.
pub fn record_detections_within(
tx: &Connection,
image_id: ImageId,
model_id: &str,
source_edge: u32,
faces: &[DetectedFace],
) -> Result<Vec<FaceId>, CatalogError> {
// Everything the old faces knew, so it can be carried across the
// replacement. Read only when there is something to carry it onto: a
// pass that found nothing has nothing to match, and decoding a vector
@@ -287,7 +305,7 @@ pub fn record_detections(
let prior = if faces.is_empty() {
Vec::new()
} else {
read_priors(&tx, image_id)?
read_priors(tx, image_id)?
};
tx.execute("DELETE FROM faces WHERE image_id = ?1", [image_id.0 as i64])?;
@@ -389,7 +407,6 @@ pub fn record_detections(
],
)?;
tx.commit()?;
Ok(ids)
}
@@ -556,6 +573,19 @@ impl FaceUpdate {
/// bookkeeping: `face_shard::export_to_shards` re-exports an image whose
/// marker is newer than the store's copy, which is how what was written
/// here reaches the other devices.
///
/// It is re-written under the pipeline id the **faces carry**, not the one
/// this pass ran as. `model_id` names the pass only through its embedder;
/// the detector half of a marker is a statement about who drew the boxes,
/// and this pass drew none. Every reader takes the two to agree: the export
/// selects an image's faces by the marker's id, `marker_under` takes a
/// marker as proof the detector has been over the image, and the shard
/// store keys each face by it. When the marker was written as
/// `scrfd_10g+w600k_mbf` over faces still spelled `w600k_mbf`, the export
/// found no faces under it and sent the other devices an entry saying the
/// thorough detector had looked and found nothing — over photographs with
/// named faces on them. With no faces left, the pass's own id is the only
/// one there is, and the marker says so.
pub fn record_updates(
conn: &Connection,
image_id: ImageId,
@@ -602,13 +632,25 @@ pub fn record_updates(
for f in dropped {
tx.execute("DELETE FROM faces WHERE id = ?1", [f.0 as i64])?;
}
let remaining: i64 = tx.query_row(
let (remaining, found_by): (i64, Option<String>) = tx.query_row(
&format!(
"SELECT COUNT(*) FROM faces WHERE image_id = ?1 AND {} = ?2",
"SELECT COUNT(*), MIN(model_id) FROM faces WHERE image_id = ?1 AND {} = ?2",
embedder_sql("model_id")
),
rusqlite::params![image_id.0 as i64, embedder_of(model_id)],
|r| r.get(0),
|r| Ok((r.get(0)?, r.get(1)?)),
)?;
let marker = found_by.as_deref().unwrap_or(model_id);
// One marker per embedder: a stale one under another spelling would
// keep saying that detector had been here, which is the claim the
// faces' own id is now making in its place.
tx.execute(
&format!(
"DELETE FROM face_index
WHERE image_id = ?1 AND model_id != ?2 AND {} = ?3",
embedder_sql("model_id")
),
rusqlite::params![image_id.0 as i64, marker, embedder_of(model_id)],
)?;
tx.execute(
"INSERT INTO face_index (image_id, model_id, indexed_at, faces_found, source_edge)
@@ -619,7 +661,7 @@ pub fn record_updates(
source_edge = excluded.source_edge",
rusqlite::params![
image_id.0 as i64,
model_id,
marker,
now_secs(),
remaining,
source_edge as i64,
@@ -1557,7 +1599,7 @@ fn iou(a: (f32, f32, f32, f32), b: (f32, f32, f32, f32)) -> f32 {
}
}
fn now_secs() -> i64 {
pub(crate) fn now_secs() -> i64 {
std::time::SystemTime::now()
.duration_since(std::time::UNIX_EPOCH)
.map(|d| d.as_secs() as i64)
@@ -1890,6 +1932,86 @@ mod tests {
assert!(at >= marked_at, "the marker was not refreshed");
}
/// The marker a per-face pass leaves names the detector that drew the
/// boxes, whatever pipeline the pass itself ran as. A marker under the
/// pass's id over faces spelled another way is one the export finds no
/// faces under — and it sent every other device "nothing here".
#[test]
fn an_update_keeps_the_marker_under_the_detector_that_found_the_faces() {
let c = db();
let img = image(&c, 1);
let ids = record_detections(
&c,
img,
"w600k_mbf",
1024,
&[DetectedFace {
quality: None,
..face(1)
}],
)
.unwrap();
// The state V14 leaves: the faces, and no marker at all.
c.execute("DELETE FROM face_index", []).unwrap();
record_updates(
&c,
img,
"scrfd_10g+w600k_mbf",
6000,
&[FaceUpdate {
embedding: Some((vec![9; 1024], 21.5)),
..FaceUpdate::for_face(ids[0])
}],
&[],
)
.unwrap();
let markers: Vec<(String, i64)> = c
.prepare("SELECT model_id, faces_found FROM face_index")
.unwrap()
.query_map([], |r| Ok((r.get(0)?, r.get(1)?)))
.unwrap()
.map(Result::unwrap)
.collect();
assert_eq!(markers, vec![("w600k_mbf".to_string(), 1)]);
// A marker already there under the pass's own id is replaced, not
// kept beside the right one.
c.execute(
"INSERT INTO face_index(image_id, model_id, indexed_at, faces_found, source_edge)
VALUES (1, 'scrfd_10g+w600k_mbf', 0, 0, 6000)",
[],
)
.unwrap();
record_updates(
&c,
img,
"scrfd_10g+w600k_mbf",
6000,
&[FaceUpdate {
crop: Some(vec![1, 2, 3]),
..FaceUpdate::for_face(ids[0])
}],
&[],
)
.unwrap();
let n: i64 = c
.query_row("SELECT COUNT(*) FROM face_index", [], |r| r.get(0))
.unwrap();
assert_eq!(n, 1, "a second marker survived");
// With every face dropped there is no detector left to name, and
// the pass's own id records that it looked.
record_updates(&c, img, "scrfd_10g+w600k_mbf", 6000, &[], &ids).unwrap();
let marker: (String, i64) = c
.query_row("SELECT model_id, faces_found FROM face_index", [], |r| {
Ok((r.get(0)?, r.get(1)?))
})
.unwrap();
assert_eq!(marker, ("scrfd_10g+w600k_mbf".to_string(), 0));
}
/// Re-detection is coalesced per image, so it must replace rather than
/// append — otherwise every re-index doubles the library's face count.
#[test]
@@ -2364,7 +2486,7 @@ mod tests {
}
/// The reference implementation's fitted MBF curve puts the P=0.5 boundary
/// at cosine 0.267 (docs/faces.md §1). Our own first end-to-end run scored
/// at cosine 0.267 (docs/dev/faces.md §1). Our own first end-to-end run scored
/// 0.596 between distinct photographs of one person and 0.05 between
/// different people, so those two must land either side.
#[test]
+112 -1
View File
@@ -119,6 +119,8 @@ pub struct MergeReport {
pub keywords_fused: usize,
/// Keyword assignments taken from the remote.
pub keywords_assigned: usize,
/// Images whose capture metadata was taken from the remote.
pub metadata_adopted: usize,
}
impl MergeReport {
@@ -133,6 +135,7 @@ impl MergeReport {
|| self.keywords_deleted > 0
|| self.keywords_fused > 0
|| self.keywords_assigned > 0
|| self.metadata_adopted > 0
}
/// Whether the local catalog holds anything the remote did not, and so
@@ -193,10 +196,67 @@ pub fn merge_all(conn: &Connection) -> Result<MergeReport, CatalogError> {
merge_collections_within(&tx, &mut report)?;
merge_keywords_within(&tx, &mut report)?;
merge_people_within(&tx, &mut report)?;
merge_metadata_within(&tx, &mut report)?;
tx.commit()?;
Ok(report)
}
/// Adopt capture metadata from an attached catalog, on its own.
pub fn merge_metadata(conn: &Connection) -> Result<MergeReport, CatalogError> {
let tx = conn.unchecked_transaction()?;
let mut report = MergeReport::default();
merge_metadata_within(&tx, &mut report)?;
tx.commit()?;
Ok(report)
}
/// Capture metadata a peer's sweep already read, for images this device has
/// not dated yet.
///
/// The `images` table is local state and the merge leaves it alone — except
/// for these columns, which are not: a capture time, an offset, a camera, a
/// lens and an ISO are facts about the file's bytes, identical on every
/// device, and read by fetching a header per image across the whole library
/// (`dr_ui::library::spawn_sweep`). A fresh device inherits its peers'
/// thumbnails and faces from the shards and then spent hours re-reading
/// every header for the timeline; the snapshot it had just merged held
/// every one of those dates.
///
/// Matched by `oc:fileid`, as collection membership is. Only rows still at
/// `metadata_state < 2` take anything, and only from a remote row at 2: a
/// date this device read for itself is never overwritten, and a peer that
/// has not read one has nothing to give. The sweep's own query
/// (`metadata_state < 2`) then finds nothing left to do for them.
const METADATA_BY_FILE_ID: &str = "
UPDATE main.images
SET captured_at = r.captured_at,
captured_offset = coalesce(main.images.captured_offset, r.captured_offset),
camera = coalesce(main.images.camera, r.camera),
lens = coalesce(main.images.lens, r.lens),
iso = coalesce(main.images.iso, r.iso),
metadata_state = 2
FROM (SELECT lr.image_id, ri.captured_at, ri.captured_offset,
ri.camera, ri.lens, ri.iso
FROM remote_cat.images ri
JOIN remote_cat.remote rr ON rr.image_id = ri.id
JOIN main.remote lr ON lr.file_id = rr.file_id
WHERE ri.metadata_state >= 2 AND ri.captured_at IS NOT NULL) AS r
WHERE main.images.id = r.image_id
AND main.images.metadata_state < 2";
fn merge_metadata_within(tx: &Connection, report: &mut MergeReport) -> Result<(), CatalogError> {
// A snapshot from before these columns, or from a library with no server
// behind it, has nothing to join on.
if !remote_has(tx, "remote")?
|| !remote_has_column(tx, "images", "metadata_state")?
|| !remote_has_column(tx, "images", "captured_offset")?
{
return Ok(());
}
report.metadata_adopted = tx.execute(METADATA_BY_FILE_ID, [])?;
Ok(())
}
/// Merge people and identity judgements from an attached catalog.
///
/// The people half of [`merge_all`], on its own, for the same reason the other
@@ -729,7 +789,7 @@ fn attached_has_table(conn: &Connection, schema: &str, table: &str) -> Result<bo
/// # What travels, and what is recomputed
///
/// The rule this module already follows for the rest of the catalog: user
/// judgements travel, inference is rebuilt. Concretely (docs/faces.md, and the
/// judgements travel, inference is rebuilt. Concretely (docs/dev/faces.md, and the
/// asymmetry `crate::faces` opens with):
///
/// - **People** — uuid, name, and whether the user set them aside. Merged by
@@ -1168,6 +1228,57 @@ mod tests {
// ---- integration over two real catalogs ------------------------------
/// A fresh device takes the capture dates a peer's sweep read, matched by
/// `oc:fileid`, and never overwrites a date it read for itself.
#[test]
fn capture_metadata_arrives_for_undated_images_only() {
let c = two_catalogs();
// Three photographs on both devices: 1 undated here and dated there;
// 2 dated on both, differently; 3 undated on both.
for id in 1..=3 {
add_image_without_hash(&c, "main", id);
add_image_without_hash(&c, "remote_cat", id + 10);
add_remote_id(&c, "main", id, 100 + id);
add_remote_id(&c, "remote_cat", id + 10, 100 + id);
}
c.execute(
"UPDATE remote_cat.images
SET captured_at = 1000, captured_offset = 60, camera = 'X', metadata_state = 2
WHERE id = 11",
[],
)
.unwrap();
c.execute(
"UPDATE remote_cat.images SET captured_at = 2000, metadata_state = 2 WHERE id = 12",
[],
)
.unwrap();
c.execute(
"UPDATE main.images SET captured_at = 2222, metadata_state = 2 WHERE id = 2",
[],
)
.unwrap();
let report = merge_metadata(&c).unwrap();
assert_eq!(report.metadata_adopted, 1);
let row = |id: i64| -> (Option<i64>, Option<i64>, Option<String>, i64) {
c.query_row(
"SELECT captured_at, captured_offset, camera, metadata_state
FROM main.images WHERE id = ?1",
[id],
|r| Ok((r.get(0)?, r.get(1)?, r.get(2)?, r.get(3)?)),
)
.unwrap()
};
assert_eq!(row(1), (Some(1000), Some(60), Some("X".into()), 2));
assert_eq!(row(2), (Some(2222), None, None, 2));
assert_eq!(row(3), (None, None, None, 0));
// Idempotent: a second pass finds nothing left to take.
assert_eq!(merge_metadata(&c).unwrap().metadata_adopted, 0);
}
fn two_catalogs() -> Connection {
attached_remote(schema::for_attached("remote_cat"))
}
+2 -2
View File
@@ -16,7 +16,7 @@
//! # The one thing a rebuild does not recover
//!
//! **Collections.** A manual collection is a set of images the user assembled
//! by hand and nothing in the filesystem records it (`docs/catalog.md` §8.1) —
//! by hand and nothing in the filesystem records it (`docs/dev/catalog.md` §8.1) —
//! which is the whole reason the catalog file itself syncs. So the two offers
//! are not interchangeable, and the interface must not present them as if they
//! were: a restore keeps the user's collections, a rebuild does not.
@@ -570,7 +570,7 @@ mod tests {
// The first NFR-R6 branch, asserted on the thing that distinguishes it
// from the second: a collection exists nowhere but the catalog, so it
// is the evidence that the *contents* came back and not merely a
// readable file (docs/catalog.md §8.1).
// readable file (docs/dev/catalog.md §8.1).
let dir = tempdir("restore");
let path = dir.join("catalog.sqlite");
fixture(&path, 500);
+176 -3
View File
@@ -15,7 +15,7 @@ use rusqlite::Connection;
use crate::error::CatalogError;
/// Schema version this build writes and understands.
pub const SCHEMA_VERSION: i64 = 19;
pub const SCHEMA_VERSION: i64 = 20;
/// Apply migrations up to [`SCHEMA_VERSION`].
///
@@ -188,9 +188,98 @@ pub fn migrate(conn: &Connection) -> Result<i64, CatalogError> {
tx.commit()?;
}
if from < 20 {
let tx = conn.unchecked_transaction()?;
v20_markers_name_the_detector_that_found_the_faces(&tx)?;
tx.pragma_update(None, "user_version", 20)?;
tx.commit()?;
}
Ok(from)
}
// V20 -- TRACES: FR-CAT-7
//
// Run markers that named the wrong detector, put right.
//
// `faces::record_updates` -- the write behind the quality, eye and crop
// passes -- re-marked an image under the pipeline the pass ran as, while
// the faces it had updated kept the id of the detector that found them.
// A marker of `scrfd_10g+w600k_mbf` over faces spelled `w600k_mbf` reads,
// to every consumer, as the thorough detector having examined the image:
// the upgrade repair skips it, and `face_shard::export_to_shards` selects
// its faces by the marker's id, finds none, and tells every other device
// that the thorough detector found nothing there. The desktop's shard index
// held 54 such entries over photographs with named faces, and the tablet's
// eye pass over faces it had adopted from the desktop had made 430 more.
//
// The write is fixed to keep the marker under the faces' own id. This puts
// the markers already written right, with a fresh time so the export sends
// each image again under an entry newer than the empty one -- which is what
// `held_model` orders by. Where the right marker is still there beside the
// wrong one (the old write inserted rather than replaced), the wrong one
// goes and the right one is refreshed for the same reason: its entry in
// the shards is older than the empty one, and a device that has neither
// would take the empty one. An image V14 left with faces and no marker at
// all is not touched: that state is the quality pass's cue, and the fixed
// write marks it correctly when the pass reaches it.
//
// Restated in Rust rather than SQL because the embedder half of a pipeline
// id is `faces::embedder_sql`, which this must agree with.
fn v20_markers_name_the_detector_that_found_the_faces(tx: &Connection) -> Result<(), CatalogError> {
let fi = crate::faces::embedder_sql("face_index.model_id");
let f = crate::faces::embedder_sql("f.model_id");
// A marker is wrong when the image holds faces of its embedder under
// another id. First the wrong ones that sit beside a right one -- the
// update below would collide with it -- then the rest are renamed.
let wrong = format!(
"EXISTS (SELECT 1 FROM faces f
WHERE f.image_id = face_index.image_id
AND {f} = {fi}
AND f.model_id != face_index.model_id)"
);
let found_by = format!(
"(SELECT MIN(f.model_id) FROM faces f
WHERE f.image_id = face_index.image_id AND {f} = {fi})"
);
let now = crate::faces::now_secs();
tx.execute(
&format!(
"UPDATE face_index
SET indexed_at = ?1
WHERE model_id = {found_by}
AND EXISTS (SELECT 1 FROM face_index w
WHERE w.image_id = face_index.image_id
AND w.model_id != face_index.model_id
AND {} = {fi})",
crate::faces::embedder_sql("w.model_id")
),
[now],
)?;
tx.execute(
&format!(
"DELETE FROM face_index
WHERE {wrong}
AND EXISTS (SELECT 1 FROM face_index o
WHERE o.image_id = face_index.image_id
AND o.model_id = {found_by})"
),
[],
)?;
tx.execute(
&format!(
"UPDATE face_index
SET model_id = {found_by},
faces_found = (SELECT COUNT(*) FROM faces f
WHERE f.image_id = face_index.image_id AND {f} = {fi}),
indexed_at = ?1
WHERE {wrong}"
),
[now],
)?;
Ok(())
}
/// The seven columns V16 adds to `faces`, in the order the readers name them.
///
/// Named once because three places have to agree on them: this migration,
@@ -920,7 +1009,7 @@ CREATE INDEX face_index_model ON face_index(model_id);
const V8: &str = r#"
-- TRACES: FR-CULL-8 | FR-CULL-9 | FR-CULL-10 | FR-CULL-11 | FR-CULL-12 | NFR-SEC-5
-- People and faces (docs/faces.md, docs/catalog.md §10).
-- People and faces (docs/dev/faces.md, docs/dev/catalog.md §10).
--
-- Everything here is **derived data** except one column. Faces, landmarks,
-- embeddings, cluster assignments and suggestions are all reproducible by
@@ -957,7 +1046,7 @@ CREATE TABLE faces (
landmarks BLOB NOT NULL, -- 5 x (x, y) f32, normalised likewise
detector_confidence REAL NOT NULL,
embedding BLOB NOT NULL, -- 512 x f16; unit length until V14, raw since
-- Source pixels across the aligned 112x112 crop (docs/faces.md §7).
-- Source pixels across the aligned 112x112 crop (docs/dev/faces.md §7).
--
-- Not cosmetic: it is the honest quality signal for the UI, a feature in
-- the §8 calibration -- FR-CULL-9 names face size as an axis along which an
@@ -1854,6 +1943,90 @@ mod tests {
assert_eq!(faces, 2);
}
#[test]
fn v20_renames_markers_to_the_detector_that_found_the_faces() {
let c = mem();
c.pragma_update(None, "user_version", 0).unwrap();
migrate(&c).unwrap();
c.execute(
"INSERT INTO roots(id, kind, label) VALUES (1, 'local', 'test')",
[],
)
.unwrap();
c.execute(
"INSERT INTO images(id, root_id, source_ref, added_at)
VALUES (1,1,'a',0),(2,1,'b',0),(3,1,'c',0),(4,1,'d',0),(5,1,'e',0)",
[],
)
.unwrap();
// 1: the desktop's case -- old faces, re-marked as thorough.
// 2: the tablet's case -- adopted thorough faces, re-marked int8,
// and the right marker still beside it (refreshed, so it is
// exported again over the empty entry).
// 3: right already. 4: examined and empty. 5: V14's state, faces
// and no marker.
for (image, model) in [
(1, "scrfd_10g+w600k_mbf"),
(2, "scrfd_10g_i8+w600k_mbf"),
(2, "scrfd_10g+w600k_mbf"),
(3, "scrfd_10g+w600k_mbf"),
(4, "scrfd_10g+w600k_mbf"),
] {
c.execute(
"INSERT INTO face_index(image_id, model_id, indexed_at, faces_found, source_edge)
VALUES (?1, ?2, 100, 0, 6000)",
rusqlite::params![image, model],
)
.unwrap();
}
for (image, model) in [
(1, "w600k_mbf"),
(1, "w600k_mbf"),
(2, "scrfd_10g+w600k_mbf"),
(3, "scrfd_10g+w600k_mbf"),
(5, "w600k_mbf"),
] {
c.execute(
"INSERT INTO faces
(image_id, x, y, w, h, landmarks, detector_confidence, embedding,
crop_px, model_id, detected_at)
VALUES (?1, 0.1, 0.1, 0.2, 0.2, X'00', 0.9, X'00', 180.0, ?2, 0)",
rusqlite::params![image, model],
)
.unwrap();
}
c.pragma_update(None, "user_version", 19).unwrap();
migrate(&c).unwrap();
let markers: Vec<(i64, String, i64, bool)> = c
.prepare(
"SELECT image_id, model_id, faces_found, indexed_at > 100
FROM face_index ORDER BY image_id, model_id",
)
.unwrap()
.query_map([], |r| Ok((r.get(0)?, r.get(1)?, r.get(2)?, r.get(3)?)))
.unwrap()
.map(Result::unwrap)
.collect();
assert_eq!(
markers,
vec![
(1, "w600k_mbf".to_string(), 2, true),
(2, "scrfd_10g+w600k_mbf".to_string(), 0, true),
(3, "scrfd_10g+w600k_mbf".to_string(), 0, false),
(4, "scrfd_10g+w600k_mbf".to_string(), 0, false),
]
);
// Re-enterable: nothing left to rename.
c.pragma_update(None, "user_version", 19).unwrap();
migrate(&c).unwrap();
let n: i64 = c
.query_row("SELECT count(*) FROM face_index", [], |r| r.get(0))
.unwrap();
assert_eq!(n, 4);
}
#[test]
fn job_uniqueness_coalesces_rather_than_duplicating() {
let c = mem();
+2 -2
View File
@@ -11,7 +11,7 @@ log.workspace = true
# Inference. `ort` is the API; **what runs it is `dr-inference-engine`'s
# business** — tract, or an ONNX Runtime the app found on disk, on whichever
# provider the device has (docs/inference.md). This crate never names either.
# provider the device has (docs/dev/inference.md). This crate never names either.
ort = { workspace = true, optional = true }
dr-inference-engine = { workspace = true, optional = true }
ndarray = { workspace = true, optional = true }
@@ -37,7 +37,7 @@ required-features = ["inference"]
[features]
# Nothing on by default, and in particular **no `embedded-model`**: the weights
# are not a build input and never become one (docs/faces.md §2.2). A feature
# are not a build input and never become one (docs/dev/faces.md §2.2). A feature
# flag that *could* embed them is a flag someone eventually sets in a packaging
# script, and the InsightFace grant does not survive that.
default = []
+1 -1
View File
@@ -1,4 +1,4 @@
//! Detect the faces in a JPEG and read each one's eyes (docs/faces.md §17).
//! Detect the faces in a JPEG and read each one's eyes (docs/dev/faces.md §17).
//!
//! The thing worth looking at is whether the eye boxes land on eyes and
//! whether soft ones are refused — so with `--dump DIR` the crops the
+1 -1
View File
@@ -7,7 +7,7 @@
//! DET.onnx EMB.onnx photo.jpg [photo.jpg ...]
//!
//! The models must have had their input dims frozen first; see
//! `tools/fix-face-model-shapes.sh` and docs/faces.md §12 M1.
//! `tools/fix-face-model-shapes.sh` and docs/dev/faces.md §12 M1.
use std::time::Instant;
+1 -1
View File
@@ -1,4 +1,4 @@
//! M1 (docs/faces.md §12) — will tract load these graphs at all?
//! M1 (docs/dev/faces.md §12) — will tract load these graphs at all?
//!
//! The one measurement everything else in the face subsystem is conditional
//! on. `det_500m.onnx` has a dynamic H/W input, which is exactly what tract
+1 -1
View File
@@ -9,7 +9,7 @@
//!
//! # What it is for
//!
//! docs/faces.md §9 has the desktop numbers and the question they leave open:
//! docs/dev/faces.md §9 has the desktop numbers and the question they leave open:
//! a GPU GEMM is worth roughly 1.5× of a regroup on a twenty-core desktop,
//! because the scan is under a third of the pass there. On a tablet the CPU is
//! several times slower and the GPU is not, so the same optimisation is worth
+5 -5
View File
@@ -1,4 +1,4 @@
//! Five-point face alignment (docs/faces.md §5).
//! Five-point face alignment (docs/dev/faces.md §5).
//!
//! ArcFace embeddings are trained on faces warped to a canonical 112×112
//! arrangement. Feeding the model a plain bounding-box crop *works* — it
@@ -208,7 +208,7 @@ impl Similarity {
///
/// # Why least squares and not RANSAC
///
/// The reference C++ implementation (docs/faces.md §1.1) fits this with
/// The reference C++ implementation (docs/dev/faces.md §1.1) fits this with
/// OpenCV's `estimateAffinePartial2D` under RANSAC. RANSAC over five points is
/// a strange fit: the minimal sample for a similarity is two, so it can discard
/// landmarks it judges outliers and solve from a subset — and on a profile face
@@ -427,7 +427,7 @@ fn sample_window(
// ── eyes ──────────────────────────────────────────────────────────────────
/// Width of an eye crop as the classifier reads it, in pixels. Fixed by the
/// OCEC input (`docs/faces.md` §17): 40 wide, 24 high.
/// OCEC input (`docs/dev/faces.md` §17): 40 wide, 24 high.
pub const EYE_PATCH_WIDTH: usize = 40;
/// Height of an eye crop as the classifier reads it, in pixels.
pub const EYE_PATCH_HEIGHT: usize = 24;
@@ -438,7 +438,7 @@ pub const EYE_PATCH_HEIGHT: usize = 24;
/// The classifier was trained on a whole-body detector's *eye* boxes — tight
/// round the palpebral fissure — and measured on 25 open-eyed faces from the
/// reference library, a tight box is what it wants: 22 of 25 read open at
/// 0 and 0.1, 18 at 0.4, 14 at 0.6 (docs/faces.md §17.2). A tenth, so a
/// 0 and 0.1, 18 at 0.4, 14 at 0.6 (docs/dev/faces.md §17.2). A tenth, so a
/// contour landing a pixel short of the lashes still holds them.
pub const EYE_BOX_MARGIN: f32 = 0.1;
@@ -571,7 +571,7 @@ pub const SUNGLASSES_EDGE: usize = 48;
/// clear glasses, at 0.68. Erring towards "sunglasses" is the safe direction
/// for what this feeds: a face called sunglasses is left alone by the
/// eyes-open filter, where a pair of sunglasses missed hands the eye
/// classifier a lens to guess at (docs/faces.md §17).
/// classifier a lens to guess at (docs/dev/faces.md §17).
pub const SUNGLASSES_WINDOWS: [(f32, f32, f32, f32); 2] =
[(0.0, 0.0, 112.0, 112.0), (-5.0, -14.0, 122.0, 122.0)];
+3 -3
View File
@@ -1,4 +1,4 @@
//! Cosine to probability (docs/faces.md §8, FR-CULL-9).
//! Cosine to probability (docs/dev/faces.md §8, FR-CULL-9).
//!
//! FR-CULL-9 is a hard requirement rather than an implementation detail: no
//! code path may threshold a bare cosine, every threshold in the subsystem is
@@ -28,7 +28,7 @@
//! calibration to the belief it was supposed to test — and that is the whole
//! of the alternative.
//!
//! docs/faces.md §8.1 names one more that would cost no labelling at all: two
//! docs/dev/faces.md §8.1 names one more that would cost no labelling at all: two
//! faces in adjacent frames of one burst are near-certainly the same person,
//! and FR-CULL-5's grouping is sitting there. Nothing draws on it. This crate
//! cannot see a catalog, let alone the bursts in one — it is handed cosines by
@@ -83,7 +83,7 @@ pub struct Calibration {
}
impl Default for Calibration {
/// The reference implementation's fitted MBF curve (docs/faces.md §1):
/// The reference implementation's fitted MBF curve (docs/dev/faces.md §1):
/// steepness 16.2, P=0.5 at cosine 0.267.
///
/// **`valid` is false**, and that is the point. It is a documented
+1 -1
View File
@@ -1,5 +1,5 @@
//! TRACES: FR-CULL-8a
//! The two small classifiers behind a face's eye state (docs/faces.md §17).
//! The two small classifiers behind a face's eye state (docs/dev/faces.md §17).
//!
//! **OCEC** — *open closed eyes classification*, Hyodo 2025 — reads one
//! 40×24 eye and answers P(open). **SGC** — *sunglasses classification*,
+1 -1
View File
@@ -1,4 +1,4 @@
//! Grouping faces into people (docs/faces.md §9, FR-CULL-10).
//! Grouping faces into people (docs/dev/faces.md §9, FR-CULL-10).
//!
//! Model-free: this is arithmetic over embeddings, and it is where the
//! subsystem's accuracy actually lives, so it is testable with no weights on
+4 -4
View File
@@ -1,4 +1,4 @@
//! SCRFD face detection (docs/faces.md §4).
//! SCRFD face detection (docs/dev/faces.md §4).
//!
//! One forward pass produces a box, a confidence and **five landmarks** per
//! face — the landmarks being the reason for this detector rather than a
@@ -138,7 +138,7 @@ impl Detection {
pub struct Detector {
session: Model,
/// f32 or int8 — the int8 form finds a different set of faces and is a
/// different detector in `model_id` (docs/inference.md §7).
/// different detector in `model_id` (docs/dev/inference.md §7).
form: Form,
/// Feature-map count: 3 for strides {8,16,32}, 4 for {8,16,32,64}.
///
@@ -346,7 +346,7 @@ fn iou(a: &(f32, f32, f32, f32), b: &(f32, f32, f32, f32)) -> f32 {
/// How the image is fitted into the graph's fixed square input.
///
/// The forward and inverse mappings live in one struct on purpose:
/// docs/faces.md §4.1 notes that what matters is not *where* the padding goes
/// docs/dev/faces.md §4.1 notes that what matters is not *where* the padding goes
/// but that the two agree. A mismatch offsets every box and landmark by the
/// padding, producing detections that look plausible and embeddings that
/// quietly cluster badly three stages later.
@@ -372,7 +372,7 @@ impl Letterbox {
///
/// `(x·255 − 127.5) / 128` — note `/128`, not `/127.5`. The reference
/// implementation this is ported from uses `/128` for both models, and
/// every measured number in docs/faces.md §1 came from it.
/// every measured number in docs/dev/faces.md §1 came from it.
///
/// Padding is grey, matching the reference's `114`: the value the network
/// reads least as an edge, where black would draw a hard border across the
+2 -2
View File
@@ -1,4 +1,4 @@
//! ArcFace / MobileFaceNet inference (docs/faces.md §6).
//! ArcFace / MobileFaceNet inference (docs/dev/faces.md §6).
//!
//! Takes an aligned crop and returns 512 L2-normalised floats. The alignment is
//! not optional and cannot be skipped by accident: [`Embedder::embed`] takes an
@@ -66,7 +66,7 @@ impl Embedder {
pub fn from_bytes(bytes: &[u8], model: ModelId) -> Result<Self, FaceError> {
// Always the f32 form: an embedding must compare across devices
// (docs/inference.md §7), and the engine pins this role to it.
// (docs/dev/inference.md §7), and the engine pins this role to it.
let loaded = dr_inference_engine::open(Role::Embedder, Form::F32, bytes)?;
let acquired = loaded.acquire()?;
let session = acquired.lock();
+2 -2
View File
@@ -1,4 +1,4 @@
//! What an embedder produces, and how it is stored (docs/faces.md §6).
//! What an embedder produces, and how it is stored (docs/dev/faces.md §6).
//!
//! Deliberately **model-free**: the vector, its identity, its comparison and
//! its storage encoding are arithmetic, and `calibrate` and `cluster` are built
@@ -273,7 +273,7 @@ mod tests {
);
}
/// The claim docs/faces.md §6 makes about the storage format: the f16
/// The claim docs/dev/faces.md §6 makes about the storage format: the f16
/// round-trip costs ~1e-3 of cosine, three orders below the separation
/// between a match and a non-match.
#[test]
+3 -3
View File
@@ -67,14 +67,14 @@ pub const SUNGLASSES_THRESHOLD: f32 = 0.5;
/// The classifier was trained on eyes down to about a dozen pixels wide
/// (its reference footage averaged 15–21); below that the 40-pixel patch is
/// an interpolation of nothing, and the answer is noise that reads as
/// "closed". docs/faces.md §17.3 has the measurement behind the number.
/// "closed". docs/dev/faces.md §17.3 has the measurement behind the number.
pub const MIN_EYE_PX: f32 = 12.0;
/// Least [`Eye::sharpness`] for the eye to be read.
///
/// The same measure as the face's `min_sharpness`, over the eye patch, and
/// chosen the same way: the value under which the open-eyed faces of the
/// reference sample were being called closed. docs/faces.md §17.3.
/// reference sample were being called closed. docs/dev/faces.md §17.3.
pub const MIN_EYE_SHARPNESS: f32 = 0.02;
/// An eye narrower than this fraction of its partner is the far eye of a
@@ -82,7 +82,7 @@ pub const MIN_EYE_SHARPNESS: f32 = 0.02;
///
/// A landmark model's contour for a hidden eye collapses towards the nose.
/// Measured on twenty native renders of the reference library
/// (docs/faces.md §17.4): profiles put the far eye at 0.02–0.43 of the near
/// (docs/dev/faces.md §17.4): profiles put the far eye at 0.02–0.43 of the near
/// one, two three-quarter faces whose far eye read closed sat at 0.54, and
/// every face looking at the camera — winks included, since a shut eye's
/// box keeps its width — sat at 0.78 or more. 0.6 splits the gap.
+1 -1
View File
@@ -1,5 +1,5 @@
//! TRACES: FR-CULL-8a
//! Dense facial landmarks — InsightFace's `2d106det` (docs/faces.md §17.2).
//! Dense facial landmarks — InsightFace's `2d106det` (docs/dev/faces.md §17.2).
//!
//! SCRFD's five points place a face; they do not place an eye. Its eye
//! point is loose enough that a window centred on it left the eye in a
+5 -4
View File
@@ -1,4 +1,4 @@
//! Faces and identity (S14, docs/faces.md).
//! Faces and identity (S14, docs/dev/faces.md).
//!
//! Two models, run over the native render, producing per face a box, five
//! landmarks, a confidence and a 512-d embedding (FR-CULL-8) — and then the
@@ -21,9 +21,9 @@
//! for a packaging script to switch on. The application obtains a model at
//! runtime; this crate takes bytes and never fetches anything.
//!
//! docs/faces.md §2 is the full reading, including what would have to change
//! docs/dev/faces.md §2 is the full reading, including what would have to change
//! for that to stop being true. The eye-state models are the exception: MIT,
//! weights and all, and shipped in `models/face/` (docs/faces.md §17).
//! weights and all, and shipped in `models/face/` (docs/dev/faces.md §17).
//!
//! # Why the runtime is split behind a feature
//!
@@ -50,11 +50,12 @@ pub mod eyes;
pub mod landmarks;
pub mod naming;
pub mod neighbours;
pub mod references;
/// Smallest long edge a face crop may be sampled from.
///
/// **A floor on the crop source, not on the detector input.** The distinction
/// is the whole of FR-CULL-8 and `docs/faces.md` §7: detection letterboxes
/// is the whole of FR-CULL-8 and `docs/dev/faces.md` §7: detection letterboxes
/// every buffer into 640×640, so its input resolution decides nothing, while
/// [`warp`] samples the 112×112 the embedder sees and so converts source
/// resolution directly into embedding quality. FR-CULL-8 requires that crop to
+1 -1
View File
@@ -115,7 +115,7 @@ pub struct Faces<'a> {
/// Source pixels across the aligned crop, for the calibration's size term.
pub crop_px: &'a [f32],
/// Which photograph each face came from. Two faces in one frame are not
/// the same person, so those pairs are never returned (docs/faces.md §9).
/// the same person, so those pairs are never returned (docs/dev/faces.md §9).
pub images: &'a [u64],
/// Which faces may be compared *against* — the gallery
/// ([`crate::embedding::MIN_GALLERY_QUALITY`]).
+232
View File
@@ -0,0 +1,232 @@
//! TRACES: FR-CULL-10 | NFR-P9
//! Which of a person's faces stand for them in a grouping pass.
//!
//! # Why not all of them
//!
//! Every face the user has ruled on enters [`crate::cluster`] as an anchor,
//! and the pass compares every face against every other
//! ([`crate::neighbours`] is exhaustive by design). So a person with 750
//! confirmed faces costs 750 comparisons against each of the library's other
//! faces, and the cost of naming a library well grows with how well it is
//! named: a fully confirmed library of 25,000 faces spends almost the whole
//! scan re-comparing faces whose identity is already settled against each
//! other.
//!
//! Most of those comparisons say nothing new. A person's confirmed faces are
//! heavily redundant — thirty frames from one afternoon are one point of
//! view, not thirty — and a new face that matches one of them matches the
//! others too. What a new face needs to be measured against is the person's
//! *range*: the angles, ages and lights they have been photographed in, each
//! represented once.
//!
//! # The choice: the most diverse of the good ones
//!
//! Two rules, in order.
//!
//! **Good enough to vouch.** Only faces whose raw embedding was at least
//! [`MIN_REFERENCE_QUALITY`] long are eligible — a stricter floor than the
//! gallery's ([`crate::embedding::MIN_GALLERY_QUALITY`]), because a reference
//! is asked to speak *for* a person rather than merely be admitted to the
//! comparison. A face whose length was never recorded is admitted, as it is
//! everywhere else: a rule that cannot be checked admits rather than excludes.
//!
//! **As far apart as possible.** From the eligible pool, up to
//! [`MAX_REFERENCES`] faces are chosen to maximise the volume they span —
//! the determinant of their Gram matrix — greedily: start from the longest
//! vector, and at each step add the face with the largest component
//! orthogonal to everything chosen so far. That is Gram–Schmidt with a
//! pivot, and the product of the squared residuals it picks *is* the
//! determinant, so the greedy step is the exact greedy on the objective.
//! The effect is that a near-duplicate of a chosen face has almost no
//! residual and is passed over, while the one profile shot among two
//! hundred frontal frames is taken early.
//!
//! What is not chosen still belongs to the person. Those faces keep their
//! confirmations and are not touched by the pass; they are simply not
//! compared, which is the whole saving.
/// The most faces that stand for one person.
///
/// A hundred is far more points of view than a person has. What it bounds
/// is the cost: with every person at the cap, a scan against the named part
/// of a library is `people × 100` comparisons per face rather than
/// `confirmations`, and the two part company as soon as a library is used.
pub const MAX_REFERENCES: usize = 100;
/// The shortest raw embedding that may stand for a person.
///
/// One above the gallery floor: a reference vouches for someone, and the
/// margin keeps the faces that only just cleared the gallery — the ones
/// nearest the middle of the sphere — out of the set that speaks for a
/// person.
pub const MIN_REFERENCE_QUALITY: f32 = 15.0;
/// Whether a face of this quality may stand for a person.
///
/// `None` is "never measured" and is admitted, as in
/// [`crate::embedding::in_gallery`].
pub fn eligible(quality: Option<f32>) -> bool {
quality.is_none_or(|q| q >= MIN_REFERENCE_QUALITY)
}
/// Choose which of one person's faces stand for them.
///
/// `embeddings` and `quality` are one entry per face, the embeddings unit
/// length and all of one dimension. Returns the indices chosen, in the order
/// chosen — the first is the longest eligible vector, and each after it is
/// the one furthest from the span of those before. Every eligible face is
/// returned when there are `max` or fewer of them, so a person under the
/// cap loses nothing.
///
/// Deterministic: equal residuals break on the longer vector, then the lower
/// index, so two devices holding the same faces choose the same references
/// and group the same way (`cluster::clustering_is_deterministic`).
pub fn select(embeddings: &[&[f32]], quality: &[Option<f32>], max: usize) -> Vec<usize> {
debug_assert_eq!(embeddings.len(), quality.len());
let mut pool: Vec<usize> = (0..embeddings.len())
.filter(|&i| eligible(quality[i]))
.collect();
if pool.len() <= max {
return pool;
}
// Longest first, so the seed is the pool's front and a tie on residual
// resolves to the earlier position. A missing reading ranks below any
// measured one for this purpose only: it is admitted, but a face that
// was measured and found long is the better seed.
pool.sort_by(|&a, &b| {
let qa = quality[a].unwrap_or(0.0);
let qb = quality[b].unwrap_or(0.0);
qb.total_cmp(&qa).then(a.cmp(&b))
});
// Residuals: what remains of each pool vector outside the span of the
// chosen ones. Copied, since they are rewritten in place.
let mut residual: Vec<Vec<f32>> = pool.iter().map(|&i| embeddings[i].to_vec()).collect();
let mut taken = vec![false; pool.len()];
let mut chosen = Vec::with_capacity(max);
while chosen.len() < max {
// The face with the most left outside the span. The seed is the
// pool's front by construction: every unit vector has the same
// residual before anything is chosen, up to rounding, and rounding
// is not a reason to prefer one. After that `> best` and not `>=`,
// so a genuine tie keeps the earlier (longer) candidate.
let mut pick = None;
let mut best = 0.0_f32;
if chosen.is_empty() {
pick = Some(0);
best = residual[0].iter().map(|x| x * x).sum();
} else {
for (k, r) in residual.iter().enumerate() {
if taken[k] {
continue;
}
let n2: f32 = r.iter().map(|x| x * x).sum();
if n2 > best {
best = n2;
pick = Some(k);
}
}
}
// Nothing left outside the span: every remaining face is a
// combination of the chosen ones and adds no volume.
let Some(k) = pick.filter(|_| best > 1e-6) else {
break;
};
taken[k] = true;
chosen.push(pool[k]);
// Project the chosen direction out of every remaining residual.
let inv = best.sqrt().recip();
let q: Vec<f32> = residual[k].iter().map(|x| x * inv).collect();
for (j, r) in residual.iter_mut().enumerate() {
if taken[j] {
continue;
}
let d: f32 = r.iter().zip(&q).map(|(a, b)| a * b).sum();
for (x, y) in r.iter_mut().zip(&q) {
*x -= d * y;
}
}
}
chosen
}
#[cfg(test)]
mod tests {
use super::*;
fn unit(v: &[f32]) -> Vec<f32> {
let n = v.iter().map(|x| x * x).sum::<f32>().sqrt();
v.iter().map(|x| x / n).collect()
}
#[test]
fn a_person_under_the_cap_keeps_every_eligible_face() {
let e = [unit(&[1.0, 0.0]), unit(&[0.0, 1.0]), unit(&[1.0, 1.0])];
let refs: Vec<&[f32]> = e.iter().map(Vec::as_slice).collect();
let q = [Some(20.0), None, Some(16.0)];
assert_eq!(select(&refs, &q, 100), vec![0, 1, 2]);
}
#[test]
fn a_short_vector_never_stands_for_a_person() {
let e = [unit(&[1.0, 0.0]), unit(&[0.0, 1.0])];
let refs: Vec<&[f32]> = e.iter().map(Vec::as_slice).collect();
let q = [Some(20.0), Some(MIN_REFERENCE_QUALITY - 0.01)];
assert_eq!(select(&refs, &q, 100), vec![0]);
}
/// Two hundred frames from one afternoon and one profile shot: the
/// profile is the second choice, not the two-hundred-and-first.
#[test]
fn the_odd_one_out_is_chosen_before_any_duplicate() {
let mut e: Vec<Vec<f32>> = Vec::new();
let mut q = Vec::new();
for i in 0..200 {
// Near-duplicates of one direction, with a little noise.
let t = (i as f32) * 1e-3;
e.push(unit(&[1.0, t, t * 0.5]));
q.push(Some(20.0 + (i % 7) as f32));
}
e.push(unit(&[0.0, 0.0, 1.0]));
q.push(Some(16.0));
let refs: Vec<&[f32]> = e.iter().map(Vec::as_slice).collect();
let chosen = select(&refs, &q, 3);
assert_eq!(chosen.len(), 3);
assert_eq!(
chosen[1], 200,
"the profile shot was not second: {chosen:?}"
);
// Seeded on the longest vector.
assert_eq!(q[chosen[0]], Some(26.0));
}
/// Faces inside the span of the chosen ones add no volume and are not
/// taken to fill the cap.
#[test]
fn the_cap_is_not_filled_from_inside_the_span() {
let e = [
unit(&[1.0, 0.0]),
unit(&[0.0, 1.0]),
unit(&[1.0, 1.0]),
unit(&[2.0, -1.0]),
];
let refs: Vec<&[f32]> = e.iter().map(Vec::as_slice).collect();
let q = [Some(20.0); 4];
assert_eq!(select(&refs, &q, 3).len(), 2);
}
#[test]
fn the_choice_is_deterministic() {
let e: Vec<Vec<f32>> = (0..50)
.map(|i| {
let a = (i as f32) * 0.37;
unit(&[a.cos(), a.sin(), (a * 3.0).sin(), 0.2])
})
.collect();
let refs: Vec<&[f32]> = e.iter().map(Vec::as_slice).collect();
let q = vec![Some(18.0); 50];
assert_eq!(select(&refs, &q, 5), select(&refs, &q, 5));
}
}
+2 -2
View File
@@ -1,6 +1,6 @@
//! What a frame actually costs — the measurement FR-DSP-2 is waiting on.
//!
//! `docs/display-and-extension.md` §2 argues that tiled computation predates
//! `docs/dev/display-and-extension.md` §2 argues that tiled computation predates
//! the fused-shader design and may not need to exist: the composer folds every
//! active operation into **one dispatch over a viewport-sized target**, so the
//! problem tiles were invented to solve may already be solved. That argument
@@ -28,7 +28,7 @@
//! the per-frame CPU half is dominated by shader-source assembly, which is
//! string formatting and is several times slower unoptimised.
//!
//! The committed numbers live in `docs/frame-budget.md`. Rerun this and diff
//! The committed numbers live in `docs/dev/frame-budget.md`. Rerun this and diff
//! that file; a regression should be a diff rather than somebody's memory.
//!
//! # Why the 99th percentile and not the mean
+1 -1
View File
@@ -1,6 +1,6 @@
//! Segment an image and write the granularity ladder as false-coloured PPMs.
//!
//! The whole point of S15 step 2 (docs/segmentation.md §11): look at the
//! The whole point of S15 step 2 (docs/dev/segmentation.md §11): look at the
//! ladder and decide whether clicking through it would land on the things a
//! person means. No amount of design settles that — the pictures do.
//!
+7 -7
View File
@@ -1,4 +1,4 @@
//! Watershed segmentation — arm A's GPU half (S15, docs/segmentation.md).
//! Watershed segmentation — arm A's GPU half (S15, docs/dev/segmentation.md).
//!
//! Runs the five passes in `shaders/watershed.wgsl` over a demosaiced image
//! and leaves a basin label per pixel on the GPU. The hierarchy built from
@@ -20,7 +20,7 @@
//! one AC-8 forbids is per frame in the render loop, and sharing a switch
//! would force a build wanting local masking to unlock the other.
//!
//! It is still a real cost and still unfinished. F3 in docs/segmentation.md
//! It is still a real cost and still unfinished. F3 in docs/dev/segmentation.md
//! §12 stands: the adjacency accumulation belongs GPU-side with atomics, and
//! until it moves there every segmentation pays a full-resolution transfer.
//! Read the feature name as a description of a known gap rather than as
@@ -36,7 +36,7 @@ pub struct SegmentOptions {
/// Longest proxy edge. The segmentation runs here, not at sensor
/// resolution: a 24 MP watershed costs 12× the memory to place boundaries
/// a person cannot see, and the boundary refinement that matters at 1:1
/// is a separate stage (docs/segmentation.md §4).
/// is a separate stage (docs/dev/segmentation.md §4).
pub max_edge: u32,
/// Pre-smoothing radius in proxy pixels. The caller's to raise with ISO —
/// this is the single knob that decides whether a noisy file segments
@@ -69,7 +69,7 @@ impl Default for SegmentOptions {
w_chroma: 0.5,
// **Zero: the pass is off.** It is implemented, dispatched
// correctly and measurably changes nothing — see the ignored test
// below and §12 of docs/segmentation.md. Until that is understood,
// below and §12 of docs/dev/segmentation.md. Until that is understood,
// running it would buy 64 dispatches per segmentation and no
// improvement, so the default declines to pay.
plateau_iterations: 0,
@@ -486,7 +486,7 @@ impl Segmentation {
/// a region graph of a few thousand nodes that every later interaction
/// reads from the CPU anyway.
///
/// What it is *not* is finished. F3 in docs/segmentation.md §12 stands:
/// What it is *not* is finished. F3 in docs/dev/segmentation.md §12 stands:
/// the adjacency accumulation belongs on the GPU with atomics, and until
/// it moves there a segmentation costs one full-resolution transfer of the
/// label and gradient buffers. That is a real cost on a phone and the
@@ -724,7 +724,7 @@ mod tests {
px
}
#[test]
#[ignore = "the plateau pass is a measured no-op; see docs/segmentation.md §12"]
#[ignore = "the plateau pass is a measured no-op; see docs/dev/segmentation.md §12"]
fn lower_completion_drains_a_plateau_instead_of_shattering_it() {
// F1, asserted rather than eyeballed, and asserted at the level where
// it matters.
@@ -741,7 +741,7 @@ mod tests {
// with no exit anywhere — cannot be drained by a distance that has
// nowhere to descend to, and collapsing it fully would need connected
// component labelling rather than a local rule. It is not worth it:
// see docs/segmentation.md §12.
// see docs/dev/segmentation.md §12.
let Some(ctx) = ctx() else { return };
let (w, h) = (96u32, 96u32);
let src = DemosaicedImage::from_rgba8(&ctx, &ramp(w, h), w, h).expect("source");
+2 -2
View File
@@ -1,4 +1,4 @@
// Watershed segmentation — the passes behind arm A of S15 (docs/segmentation.md).
// Watershed segmentation — the passes behind arm A of S15 (docs/dev/segmentation.md).
//
// Seven entry points forming one chain:
//
@@ -197,7 +197,7 @@ fn gradient(@builtin(global_invocation_id) gid: vec3<u32>) {
// lowest-indexed neighbour, which is up and to the left. Each pixel therefore
// walks diagonally until it falls off the plateau, and one flat region becomes
// a fan of diagonal chains rather than one basin — visible as hatching across
// what should be a single area (docs/segmentation.md §12, F1).
// what should be a single area (docs/dev/segmentation.md §12, F1).
//
// The fix is the standard lower-completion: give each plateau pixel its
// geodesic distance to the nearest pixel that *does* have a lower neighbour,
+5 -5
View File
@@ -2,11 +2,11 @@
//!
//! FR-DSP-3 says a slider updates the visible region within one frame budget at
//! proxy resolution. Until this file existed nothing checked it, which made it
//! a wish — `docs/display-and-extension.md` §3 is blunt about that, and §7 is
//! a wish — `docs/dev/display-and-extension.md` §3 is blunt about that, and §7 is
//! blunt about what tagging an unchecked requirement does to the coverage
//! figure.
//!
//! The measurements this guards are in [`docs/frame-budget.md`], produced by
//! The measurements this guards are in [`docs/dev/frame-budget.md`], produced by
//! `examples/frame_budget.rs`. This file is the part of them that has to keep
//! being true: it renders the **whole point-operation chain** through the real
//! `render_detailed` for a hundred frames, moving a slider between each, and
@@ -17,7 +17,7 @@
//! **The neighbourhood stage is deliberately not in the asserted chain.** It is
//! over the budget today — clarity alone is 34 ms at 4K, because its kernel is
//! a fraction of the frame and reaches a 52-pixel radius there — and
//! `docs/frame-budget.md` records that, names the fix (a base computed at
//! `docs/dev/frame-budget.md` records that, names the fix (a base computed at
//! reduced resolution) and does not pretend otherwise. Asserting a budget the
//! code does not meet would produce a red suite that everyone learns to ignore;
//! asserting it on a chain that quietly excluded the expensive stage *without
@@ -81,7 +81,7 @@ const SOURCE: (u32, u32) = (6000, 4000);
/// The viewport the budget is asserted at: a 16:10 desktop display.
///
/// Not 4K, and the reason is worth stating. At 4K the fused chain still passes
/// with room to spare (4.5 ms of GPU; see `docs/frame-budget.md`), but a test
/// with room to spare (4.5 ms of GPU; see `docs/dev/frame-budget.md`), but a test
/// that renders 8.3 M pixels a hundred times twice over is four seconds of
/// suite time to re-establish a conclusion 4.1 M pixels already establishes.
const VIEWPORT: (u32, u32) = (2560, 1600);
@@ -166,7 +166,7 @@ impl Run {
judged <= BUDGET_MS,
"{case} at {}x{}: p99 of {FRAMES} frames was {judged:.2} ms, over the \
{BUDGET_MS:.0} ms budget (cpu {:.2} ms, gpu {:.2} ms, total {:.2} ms). \
FR-DSP-3 is what this violates; docs/frame-budget.md holds the \
FR-DSP-3 is what this violates; docs/dev/frame-budget.md holds the \
numbers it used to be.",
viewport.0,
viewport.1,
+1 -1
View File
@@ -326,7 +326,7 @@ fn a_proxy_and_an_export_agree_about_the_effect() {
#[test]
fn crossing_the_reduction_threshold_does_not_change_the_picture() {
// TRACES: FR-DSP-3 — `docs/technical-debt.md` TD-4, held in pixels.
// TRACES: FR-DSP-3 — `docs/dev/technical-debt.md` TD-4, held in pixels.
//
// Clarity's base is computed on a reduced grid, and how reduced depends on
// the viewport: `LocalContrast::reduction` steps 4 -> 2 -> 1 as sigma
+1 -1
View File
@@ -10,7 +10,7 @@
//! code path — the zoom is the full-resolution path — which is why the
//! requirement has been satisfied for some time without anyone tagging it.
//!
//! `docs/display-and-extension.md` §7 is the reason this file exists rather
//! `docs/dev/display-and-extension.md` §7 is the reason this file exists rather
//! than a tag on `framing.rs`: a requirement counts as covered when a `TRACES`
//! comment names it, and nothing checks that the code under the tag does the
//! thing. `FR-DEV-8` is tagged against plumbing a future operation would use.
+10 -2
View File
@@ -6,7 +6,7 @@ rust-version.workspace = true
license.workspace = true
# The one crate that names a runtime, a provider, a vendor library or a
# device (docs/inference.md §8). `dr-face` and `dr-segment` ask it for a
# device (docs/dev/inference.md §8). `dr-face` and `dr-segment` ask it for a
# session by role and never see which of these answered.
[dependencies]
@@ -28,7 +28,10 @@ ort-sys = { version = "2.0.0-rc.13", default-features = false, features = ["disa
# The NVIDIA rungs exist on the desktop only. These features add `ort`'s
# option builders and nothing else — no linking under `alternative-backend` —
# but an Android binary has no business carrying even the option names, and
# the packaging must never be tempted to (§2, §3.1).
# the packaging must never be tempted to (§2, §3.1). The AMD rung needs no
# feature: MIGraphX is registered through the runtime's generic key/value
# entry point (`session::migraphx`), because `ort`'s own builder cannot
# name the compiled-program cache.
[target.'cfg(not(target_os = "android"))'.dependencies]
ort = { workspace = true, features = ["cuda", "tensorrt"] }
@@ -42,3 +45,8 @@ default = ["tract"]
tract = ["dep:ort-tract"]
# Look for `libonnxruntime` on disk and hand its table to `ort`.
native = ["dep:libloading", "dep:ort-sys"]
[dev-dependencies]
# The `ep_probe` example prints the provider's own diagnostics, which is most
# of what a failed rung tells you.
env_logger.workspace = true
@@ -0,0 +1,207 @@
//! Time each execution provider a runtime offers, on the models this
//! repository ships — the measurement docs/inference.md §1 requires before a
//! rung is added to §2's ladder.
//!
//! DARKROOM_ORT_DIR=/usr/lib \
//! cargo run --release -p dr-inference-engine --features native,tract \
//! --example ep_probe -- models/face/scrfd_500m_640.onnx ...
//!
//! Prints one row per (model, provider): the median of timed runs after
//! warm-ups, and the build time, which for a compiling provider is the
//! number that decides whether it needs an engine cache. MIGraphX is built
//! twice per precision — cold, then again from the cache it just wrote —
//! so both numbers are on the page.
//!
//! The ROCm provider is not in the list: ONNX Runtime removed it in 1.23,
//! and 1.29's `onnxruntime-rocm` ships `libonnxruntime_providers_migraphx.so`
//! and nothing else for AMD.
use std::path::{Path, PathBuf};
use std::time::Instant;
#[derive(Clone, Copy, PartialEq)]
enum Ep {
Cpu,
MiGraphX,
MiGraphXFp16,
}
impl Ep {
fn label(self) -> &'static str {
match self {
Ep::Cpu => "CPU",
Ep::MiGraphX => "MIGraphX f32",
Ep::MiGraphXFp16 => "MIGraphX fp16",
}
}
}
fn build(ep: Ep, bytes: &[u8], threads: usize, cache: &Path) -> ort::Result<ort::session::Session> {
let mut b = ort::session::Session::builder()?.with_intra_threads(threads)?;
match ep {
Ep::Cpu => {}
Ep::MiGraphX => migraphx(&mut b, false, &cache.join("f32"))?,
Ep::MiGraphXFp16 => migraphx(&mut b, true, &cache.join("fp16"))?,
}
b.commit_from_memory(bytes)
}
/// Register MIGraphX through the generic key/value API. `ort`'s own
/// builder fills the legacy `OrtMIGraphXProviderOptions`, which 1.29 reads
/// for its precision flags and nothing else: the model cache directory —
/// the difference between a 40 s load and a 0.3 s one — only travels this
/// way. The cache key is the graph, the GPU and the MIGraphX version, not
/// the precision, so each precision gets its own directory.
fn migraphx(
b: &mut ort::session::builder::SessionBuilder,
fp16: bool,
cache: &Path,
) -> ort::Result<()> {
use ort::AsPointer;
use std::ffi::CString;
std::fs::create_dir_all(cache).map_err(|e| ort::Error::new(e.to_string()))?;
let keys = [c"migraphx_fp16_enable", c"migraphx_model_cache_dir"];
let values = [
CString::new(if fp16 { "1" } else { "0" }).unwrap(),
CString::new(cache.to_string_lossy().as_bytes()).unwrap(),
];
let key_ptrs: Vec<_> = keys.iter().map(|k| k.as_ptr()).collect();
let value_ptrs: Vec<_> = values.iter().map(|v| v.as_ptr()).collect();
// SAFETY: the documented C call, over arrays that outlive it; the
// runtime copies the strings into its own options map.
unsafe {
let status = (ort::api().SessionOptionsAppendExecutionProvider)(
b.ptr_mut(),
c"MIGraphX".as_ptr(),
key_ptrs.as_ptr(),
value_ptrs.as_ptr(),
keys.len(),
);
ort::Error::result_from_status(status)
}
}
/// Median of `runs` timed runs over zeros, in milliseconds, after warm-ups.
fn time(session: &mut ort::session::Session, warmups: usize, runs: usize) -> Result<f64, String> {
let shape: Vec<usize> = session.inputs()[0]
.dtype()
.tensor_shape()
.ok_or("input is not a tensor")?
.iter()
.map(|&d| if d > 0 { d as usize } else { 1 })
.collect();
let zeros = vec![0f32; shape.iter().product()];
let once = |s: &mut ort::session::Session| -> Result<f64, String> {
let input = ort::value::Tensor::from_array((shape.clone(), zeros.clone()))
.map_err(|e| e.to_string())?;
let t = Instant::now();
let out = s.run(ort::inputs![input]).map_err(|e| e.to_string())?;
let _ = out[0]
.try_extract_tensor::<f32>()
.map_err(|e| e.to_string())?;
Ok(t.elapsed().as_secs_f64() * 1e3)
};
for _ in 0..warmups {
once(session)?;
}
let mut times = Vec::with_capacity(runs);
for _ in 0..runs {
times.push(once(session)?);
}
times.sort_by(|a, b| a.partial_cmp(b).unwrap());
Ok(times[times.len() / 2])
}
fn first_line(s: &str) -> String {
s.lines().next().unwrap_or("").chars().take(120).collect()
}
fn main() {
env_logger::Builder::from_env(env_logger::Env::default().default_filter_or("info")).init();
let models: Vec<PathBuf> = std::env::args_os().skip(1).map(PathBuf::from).collect();
if models.is_empty() {
eprintln!("usage: ep_probe MODEL.onnx [MODEL.onnx ...]");
std::process::exit(2);
}
dr_inference_engine::ensure_runtime();
let runtime = dr_inference_engine::status().runtime;
println!("runtime: {}", runtime.label());
if !runtime.is_native() {
println!("(tract: no provider to compare; set DARKROOM_ORT_DIR)");
}
let threads = std::thread::available_parallelism()
.map(|n| n.get().saturating_sub(2).max(1))
.unwrap_or(1);
println!("intra-op threads: {threads}");
let cache = std::env::temp_dir().join("darkroom-ep-probe");
let _ = std::fs::remove_dir_all(&cache);
println!("compiled-program cache: {}\n", cache.display());
println!(
"{:<28} {:<15} {:>10} {:>10}",
"model", "provider", "build s", "median ms"
);
for model in &models {
let bytes = match std::fs::read(model) {
Ok(b) => b,
Err(e) => {
println!("{:<28} read failed: {e}", name(model));
continue;
}
};
// A compiling provider is built twice: the second build reads the
// program the first wrote, and its time is what a launch after the
// first costs.
let plan = [
(Ep::Cpu, false),
(Ep::MiGraphX, false),
(Ep::MiGraphX, true),
(Ep::MiGraphXFp16, false),
(Ep::MiGraphXFp16, true),
];
for (ep, cached) in plan {
let started = Instant::now();
match build(ep, &bytes, threads, &cache) {
Ok(mut session) => {
let built = started.elapsed().as_secs_f64();
match time(&mut session, 3, 15) {
Ok(ms) => println!(
"{:<28} {:<15} {:>10.1} {:>10.1}{}",
name(model),
ep.label(),
built,
ms,
if cached { " (from cache)" } else { "" }
),
Err(e) => println!(
"{:<28} {:<15} {:>10.1} {:>10} {}",
name(model),
ep.label(),
built,
"ran ✗",
first_line(&e)
),
}
}
Err(e) => println!(
"{:<28} {:<15} {:>21} {}",
name(model),
ep.label(),
"build ✗",
first_line(&e.to_string())
),
}
}
println!();
}
}
fn name(p: &Path) -> String {
p.file_name()
.unwrap_or(p.as_os_str())
.to_string_lossy()
.into_owned()
}
+100
View File
@@ -0,0 +1,100 @@
//! Walk the ladder as the app does — probe, engines, then a session — and
//! say what each step chose. The M5 check of docs/inference.md §6 without
//! the app around it.
//!
//! DARKROOM_ORT_DIR=/usr/lib \
//! cargo run --release -p dr-inference-engine --features native,tract \
//! --example ladder -- CACHE_DIR models/face/scrfd_500m_640.onnx [MODEL.onnx ...]
//!
//! Every model named is a `Detector` for the config's purposes, which is
//! enough to see the rung taken, the engines compiled and a session land
//! on it. Delete `CACHE_DIR` to see the first run again; keep it to see the
//! second.
use std::path::PathBuf;
use std::time::{Duration, Instant};
fn main() {
env_logger::Builder::from_env(env_logger::Env::default().default_filter_or("info")).init();
let mut args = std::env::args_os().skip(1).map(PathBuf::from);
let (Some(cache_dir), models) = (args.next(), args.collect::<Vec<_>>()) else {
eprintln!("usage: ladder CACHE_DIR MODEL.onnx [MODEL.onnx ...]");
std::process::exit(2);
};
if models.is_empty() {
eprintln!("usage: ladder CACHE_DIR MODEL.onnx [MODEL.onnx ...]");
std::process::exit(2);
}
let runtime_dirs: Vec<PathBuf> = std::env::var_os("DARKROOM_ORT_DIR")
.map(PathBuf::from)
.into_iter()
.collect();
let started = Instant::now();
dr_inference_engine::init(dr_inference_engine::Config {
runtime_dirs,
cache_dir: cache_dir.clone(),
models: models
.iter()
.map(|p| (dr_inference_engine::Role::Detector, p.clone()))
.collect(),
embedded: Vec::new(),
ceiling: None,
threads: 0,
decay: Duration::ZERO,
});
let mut last = String::new();
loop {
let s = dr_inference_engine::status();
let line = format!(
"{} · {} · engines {}/{}{}",
s.line(),
if s.probing {
"probing"
} else {
s.reason.as_str()
},
s.engines.0,
s.engines.1,
if s.failed.is_empty() {
String::new()
} else {
format!(
" · tried {}",
s.failed
.iter()
.map(|(r, why)| format!("{}: {why}", r.label()))
.collect::<Vec<_>>()
.join(" · ")
)
}
);
if line != last {
println!("{:>6.1} s {line}", started.elapsed().as_secs_f64());
last = line;
}
if !s.probing && s.engines.0 >= s.engines.1 {
break;
}
std::thread::sleep(Duration::from_millis(500));
}
for path in &models {
let bytes = std::fs::read(path).expect("read model");
let t = Instant::now();
let model = dr_inference_engine::open(
dr_inference_engine::Role::Detector,
dr_inference_engine::Form::F32,
&bytes,
)
.expect("open model");
let acquired = model.acquire().expect("acquire session");
println!(
"{} on {} in {:.2} s",
path.file_name().unwrap().to_string_lossy(),
acquired.rung().label(),
t.elapsed().as_secs_f64()
);
}
}
+1 -1
View File
@@ -1,4 +1,4 @@
//! The API table `ort` runs on, chosen once (docs/inference.md §3).
//! The API table `ort` runs on, chosen once (docs/dev/inference.md §3).
//!
//! `ort` with `alternative-backend` links no runtime and asks, on first use,
//! for an `OrtApi` — a struct of function pointers. Two things can fill it:
+1 -1
View File
@@ -1,5 +1,5 @@
//! Compiled engines: what a rung builds once per device, and the thread that
//! builds them before anyone asks (docs/inference.md §5, §6).
//! builds them before anyone asks (docs/dev/inference.md §5, §6).
//!
//! TensorRT keeps its own engine cache keyed by graph hash; QNN writes a
//! context model. Both are opaque to this crate, which tracks only *that* a
+56 -14
View File
@@ -1,10 +1,10 @@
//! Which runtime, which provider and which model form — decided once per
//! device, and the only crate that knows the answer (docs/inference.md).
//! device, and the only crate that knows the answer (docs/dev/inference.md).
//!
//! Consumers ask for a session by [`Role`] and get `ort`'s `Session` back;
//! what built it — tract on one core, ONNX Runtime's CPU pool, a TensorRT
//! engine, the Hexagon — is this crate's business and shows up in
//! [`status`] for the settings row and nowhere else.
//! engine, a MIGraphX program, the Hexagon — is this crate's business and
//! shows up in [`status`] for the settings row and nowhere else.
//!
//! The shape follows §3 of the spec: `ort` links nothing (`alternative-backend`),
//! and the first call hands it an API table from either a `libonnxruntime`
@@ -35,13 +35,13 @@ pub enum Role {
Embedder,
Segmenter,
Scene,
/// The dense landmark model behind the eye reading (docs/faces.md §7c).
/// The dense landmark model behind the eye reading (docs/dev/faces.md §7c).
Landmarks,
/// The eye-state and sunglasses classifiers, a few hundred kilobytes.
EyeClassifier,
/// XFeat, the panorama keypoint detector (docs/panorama.md).
/// XFeat, the panorama keypoint detector (docs/dev/panorama.md).
Keypoints,
/// MI-GAN, the panorama border filler (docs/panorama.md §12). Plain
/// MI-GAN, the panorama border filler (docs/dev/panorama.md §12). Plain
/// convolutions, so any rung serves it; fp16 on TensorRT and int8 on
/// the Hexagon are the point of it.
Inpainter,
@@ -60,7 +60,9 @@ pub enum Form {
/// A rung of the ladder (§2). Ordered: a user override names the highest rung
/// the probe may take, and a compiling rung falls back to the one below it
/// until its engine exists.
/// until its engine exists. The order is within a vendor's ladder — a
/// machine has NVIDIA rungs or an AMD rung, never both — so a ceiling is
/// read as "no higher than this on whichever ladder the device has".
#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash, PartialOrd, Ord, Serialize, Deserialize)]
pub enum Rung {
/// ONNX Runtime's CPU provider, or tract when no runtime file was found.
@@ -69,6 +71,11 @@ pub enum Rung {
Cuda,
/// NVIDIA, through a TensorRT engine compiled on this device. Desktop only.
TensorRt,
/// AMD, through a MIGraphX program compiled on this device. Desktop
/// only. ONNX Runtime's ROCm provider, the CUDA provider's twin, was
/// removed in ONNX Runtime 1.23, so there is no non-compiling AMD rung
/// to fall back to: this one falls back to the CPU.
MiGraphX,
/// Qualcomm's Hexagon NPU through QNN, int8 models only. Android only.
Hexagon,
}
@@ -79,6 +86,7 @@ impl Rung {
Rung::Cpu => "CPU",
Rung::Cuda => "CUDA",
Rung::TensorRt => "TensorRT",
Rung::MiGraphX => "MIGraphX",
Rung::Hexagon => "Hexagon NPU",
}
}
@@ -88,13 +96,13 @@ impl Rung {
fn fallback(self) -> Rung {
match self {
Rung::TensorRt => Rung::Cuda,
Rung::Hexagon | Rung::Cuda | Rung::Cpu => Rung::Cpu,
Rung::MiGraphX | Rung::Hexagon | Rung::Cuda | Rung::Cpu => Rung::Cpu,
}
}
/// Whether a session on this rung needs an engine built first.
fn compiles(self) -> bool {
matches!(self, Rung::TensorRt | Rung::Hexagon)
matches!(self, Rung::TensorRt | Rung::MiGraphX | Rung::Hexagon)
}
/// The model form this rung wants for a role.
@@ -164,7 +172,7 @@ impl Status {
pub fn line(&self) -> String {
let form = match self.rung {
Rung::Hexagon => " · int8",
Rung::TensorRt => " · fp16",
Rung::TensorRt | Rung::MiGraphX => " · fp16",
_ => "",
};
format!("{}{} · {}", self.rung.label(), form, self.runtime.label())
@@ -269,8 +277,8 @@ fn acquire(role: Role, form: Form, bytes: &Arc<[u8]>, hash: u64) -> Result<Acqui
return Ok(Acquired { entry });
}
// Built outside the registry lock: a TensorRT engine load is long enough
// that another role's acquire should not wait on it.
// Built outside the registry lock: a TensorRT or MIGraphX engine load
// is long enough that another role's acquire should not wait on it.
let session = session::build(rung, role, bytes, &cfg)?;
log::debug!("inference: {role:?} loaded on {}", rung.label());
let entry = Arc::new(Loaded {
@@ -401,7 +409,16 @@ pub fn status() -> Status {
runtime: api::runtime(),
rung,
reason: s.cache.reason.clone(),
failed: s.cache.failed.clone(),
// Only what explains the selection: on an AMD machine the NVIDIA
// rungs "not enabled in this build" say nothing about why MIGraphX
// was taken. With the floor selected, everything tried is above it.
failed: s
.cache
.failed
.iter()
.filter(|(r, _)| *r > rung)
.cloned()
.collect(),
probing: s.probing,
engines: if rung.compiles() {
(s.cache.compiled.len(), s.wanted)
@@ -499,7 +516,7 @@ mod tests {
/// The smallest shipped graph, if this checkout has the weights; a test
/// suite that needs a research-licensed download is one that does not
/// run in CI (docs/faces.md §3), so absence is a skip.
/// run in CI (docs/dev/faces.md §3), so absence is a skip.
fn probe_bytes() -> Option<Vec<u8>> {
let path = concat!(
env!("CARGO_MANIFEST_DIR"),
@@ -614,8 +631,33 @@ mod tests {
);
}
#[test]
fn the_status_reports_only_the_rungs_above_the_selection() {
let _serial = serial();
let failed = vec![
(Rung::TensorRt, "not enabled".to_string()),
(Rung::Cuda, "not enabled".to_string()),
];
let before = state().lock().unwrap().cache.clone();
state().lock().unwrap().cache = Cache {
rung: Some(Rung::MiGraphX),
failed: failed.clone(),
..Cache::default()
};
// An AMD desktop: the NVIDIA rungs below MIGraphX are not the story.
assert!(status().failed.is_empty());
// An NVIDIA desktop on the CUDA provider: TensorRT's failure is.
state().lock().unwrap().cache.rung = Some(Rung::Cuda);
assert_eq!(status().failed, vec![failed[0].clone()]);
// The floor: everything tried explains it.
state().lock().unwrap().cache.rung = Some(Rung::Cpu);
assert_eq!(status().failed.len(), 2);
state().lock().unwrap().cache = before;
}
#[test]
fn the_status_line_reads_as_the_floor_before_init() {
let _serial = serial();
let s = status();
assert_eq!(s.rung, Rung::Cpu);
assert!(s.line().starts_with("CPU"), "{}", s.line());
+46 -8
View File
@@ -1,4 +1,4 @@
//! Walk the ladder, once, by building real sessions (docs/inference.md §4).
//! Walk the ladder, once, by building real sessions (docs/dev/inference.md §4).
//!
//! A rung is taken when a session builds on it, runs, and is faster than
//! the floor. Both halves matter: a provider can register and then fail at
@@ -16,8 +16,11 @@ use crate::{api::Runtime, state, Cache, Config, Form, Role, Rung};
fn ladder(ceiling: Option<Rung>) -> Vec<Rung> {
#[cfg(target_os = "android")]
let all = [Rung::Hexagon];
// A desktop has one vendor's GPU; the other vendor's providers are
// "not enabled in this build" or a library that fails to load, and
// either answer arrives in milliseconds.
#[cfg(not(target_os = "android"))]
let all = [Rung::TensorRt, Rung::Cuda];
let all = [Rung::TensorRt, Rung::Cuda, Rung::MiGraphX];
all.into_iter()
.filter(|r| ceiling.is_none_or(|c| *r <= c))
.collect()
@@ -206,15 +209,22 @@ fn first_line(s: &str) -> String {
line[start..].chars().take(200).collect()
}
/// Everything a change of which should re-probe: the runtime and where it
/// came from, this crate, the platform, the driver or SoC, and the models.
/// Everything a change of which should re-probe: the runtime, where it
/// came from and which providers sit beside it, this crate, the platform,
/// the driver or SoC, and the models.
fn fingerprint(runtime: &Runtime, cfg: &Config) -> String {
let mut parts = vec![
format!("engine {}", env!("CARGO_PKG_VERSION")),
format!("{} {}", std::env::consts::OS, std::env::consts::ARCH),
match runtime {
Runtime::Tract => "tract".to_string(),
Runtime::OnnxRuntime { path, version } => format!("ort {version} {}", path.display()),
Runtime::OnnxRuntime { path, version } => {
format!(
"ort {version} {} [{}]",
path.display(),
providers_beside(path)
)
}
},
device_identity(),
];
@@ -237,13 +247,41 @@ fn fingerprint(runtime: &Runtime, cfg: &Config) -> String {
parts.join("\n")
}
/// The `libonnxruntime_providers_*.so` files in the runtime's directory.
/// A distribution's CPU-only and ROCm builds are the same version at the
/// same path; the provider libraries beside them are what differs.
fn providers_beside(runtime: &Path) -> String {
let Some(dir) = runtime.parent() else {
return String::new();
};
let mut names: Vec<String> = std::fs::read_dir(dir)
.into_iter()
.flatten()
.filter_map(|e| e.ok())
.filter_map(|e| e.file_name().into_string().ok())
.filter(|n| {
n.starts_with("libonnxruntime_providers_") || n.starts_with("onnxruntime_providers_")
})
.collect();
names.sort();
names.join(" ")
}
#[cfg(target_os = "linux")]
fn device_identity() -> String {
// The NVIDIA driver's version line; absent means no NVIDIA driver.
std::fs::read_to_string("/proc/driver/nvidia/version")
// The NVIDIA driver's version line, or the ROCm release the AMD stack
// came from (`rocm-core` writes it; the kernel driver has no version
// of its own). Absent means neither.
if let Some(line) = std::fs::read_to_string("/proc/driver/nvidia/version")
.ok()
.and_then(|s| s.lines().next().map(str::to_string))
.unwrap_or_else(|| "no nvidia driver".into())
{
return line;
}
if let Ok(rocm) = std::fs::read_to_string("/opt/rocm/.info/version") {
return format!("rocm {}", rocm.trim());
}
"no nvidia driver, no rocm".into()
}
#[cfg(target_os = "android")]
+59 -2
View File
@@ -1,4 +1,4 @@
//! One session builder per rung (docs/inference.md §2, §7, §9).
//! One session builder per rung (docs/dev/inference.md §2, §7, §9).
use ort::session::Session;
@@ -83,10 +83,65 @@ fn providers(
ep::CUDA::default().build(),
])?)
}
Rung::MiGraphX => {
// fp16 on the same terms as TensorRT (§7). MIGraphX compiles a
// program per graph — 20–60 s here — and keeps it in the cache
// directory, keyed on the graph, the GPU and its own version
// but not the precision: hence one directory per precision.
// The CPU takes any node it declines.
let fp16 = role != Role::Embedder;
let cache = cfg
.cache_dir
.join("migraphx")
.join(if fp16 { "fp16" } else { "f32" });
let _ = std::fs::create_dir_all(&cache);
let mut b = b;
migraphx(&mut b, fp16, &cache)?;
Ok(b)
}
Rung::Hexagon => unreachable!("the Hexagon rung is not on a desktop ladder"),
}
}
/// Register MIGraphX through ONNX Runtime's generic key/value entry point.
///
/// `ort`'s own builder (`ep::MIGraphX`) fills the legacy
/// `OrtMIGraphXProviderOptions`, and 1.29 reads that struct for its
/// precision flags and nothing else — the compiled-program cache directory
/// is only a key in the generic map (`migraphx_model_cache_dir`), and
/// without it every session is a full compile. Registration through the
/// generic entry point needs no `ort` feature: it is one call on the API
/// table, which is why the crate's `ort` dependency names no AMD feature.
#[cfg(not(target_os = "android"))]
fn migraphx(
b: &mut ort::session::builder::SessionBuilder,
fp16: bool,
cache: &std::path::Path,
) -> ort::Result<()> {
use ort::AsPointer;
use std::ffi::CString;
let keys = [c"migraphx_fp16_enable", c"migraphx_model_cache_dir"];
let values = [
CString::new(if fp16 { "1" } else { "0" }).unwrap(),
CString::new(cache.to_string_lossy().as_bytes())
.map_err(|e| ort::Error::new(e.to_string()))?,
];
let key_ptrs: Vec<_> = keys.iter().map(|k| k.as_ptr()).collect();
let value_ptrs: Vec<_> = values.iter().map(|v| v.as_ptr()).collect();
// SAFETY: the documented C call over arrays that outlive it; the
// runtime copies the strings into its own options map before returning.
unsafe {
let status = (ort::api().SessionOptionsAppendExecutionProvider)(
b.ptr_mut(),
c"MIGraphX".as_ptr(),
key_ptrs.as_ptr(),
value_ptrs.as_ptr(),
keys.len(),
);
ort::Error::result_from_status(status)
}
}
#[cfg(target_os = "android")]
fn providers(
b: ort::session::builder::SessionBuilder,
@@ -120,6 +175,8 @@ fn providers(
.build()
.error_on_failure()])?)
}
Rung::Cuda | Rung::TensorRt => unreachable!("no NVIDIA rung on Android"),
Rung::Cuda | Rung::TensorRt | Rung::MiGraphX => {
unreachable!("no desktop GPU rung on Android")
}
}
}
+1 -1
View File
@@ -13,7 +13,7 @@ log.workspace = true
# Inference for the learned keypoint detector, on the same footing as
# `dr-segment`: `ort` is the API, `dr-inference-engine` decides what runs
# it (docs/inference.md), and both are optional so that the geometry —
# it (docs/dev/inference.md), and both are optional so that the geometry —
# matching, the rotation solve, the projections — is a dependency-free crate
# that tests without a model.
ort = { workspace = true, optional = true }
+62 -10
View File
@@ -95,7 +95,7 @@ pub struct Params {
/// and, beyond the band being filled, still unknown. That is what the
/// shipped model was trained on (a fine-tune of MI-GAN on voids cut
/// from photographs the way a cylindrical merge cuts them, see
/// `docs/panorama.md` §14); a ring would give it a fold to continue.
/// `docs/dev/panorama.md` §14); a ring would give it a fold to continue.
///
/// Non-zero is the stock model's crutch: a plain reflection of a deep
/// hole pulls in whatever is that far from the edge — a ridge, a peak —
@@ -303,7 +303,7 @@ fn fill_once(
}
// The padded canvas with mirrored context, and the hole within it.
let ctx = MirroredContext::build(rgb, width, height, known, mirror_depth);
let ctx = MirroredContext::build(rgb, width, height, known, mirror_depth, t);
let (pw, ph) = (ctx.width, ctx.height);
// Tiles that touch the hole, on a grid that reaches both far edges.
@@ -499,16 +499,36 @@ struct MirroredContext {
}
impl MirroredContext {
fn build(rgb: &[f32], width: usize, height: usize, known: &[bool], depth: usize) -> Self {
fn build(
rgb: &[f32],
width: usize,
height: usize,
known: &[bool],
depth: usize,
tile: usize,
) -> Self {
if depth == 0 {
// Open: the picture as it is, the hole as it is. What the hole
// holds does not matter — the model masks it out.
// holds does not matter — the model masks it out. A picture
// smaller than a tile (the merge page's preview) sits at the
// origin of a tile-sized canvas whose rest is hole: still the
// void as it is, and the only way a tile fits at all.
let (pw, ph) = (width.max(tile), height.max(tile));
let mut canvas = vec![0.0f32; pw * ph * 3];
let mut hole = vec![true; pw * ph];
for y in 0..height {
canvas[y * pw * 3..(y * pw + width) * 3]
.copy_from_slice(&rgb[y * width * 3..(y + 1) * width * 3]);
for x in 0..width {
hole[y * pw + x] = !known[y * width + x];
}
}
return MirroredContext {
width,
height,
width: pw,
height: ph,
ring: 0,
rgb: rgb.to_vec(),
hole: known.iter().map(|&k| !k).collect(),
rgb: canvas,
hole,
};
}
let fold = |d: usize| fold(d, depth);
@@ -715,6 +735,38 @@ mod tests {
}
}
#[test]
fn a_picture_smaller_than_the_tile_is_still_filled_when_the_void_is_open() {
// The merge page's preview is 1600 wide and a few hundred tall —
// shorter than a 512 tile. With no ring the canvas is padded to a
// tile, the padding hole, and the border is still filled.
let (mut rgb, known) = picture(300, 40, 8);
let mut model = Flat {
tile: 64,
seen: Vec::new(),
};
let tiles = fill_border(
&mut rgb,
300,
40,
&known,
&mut model,
test_params(0),
&mut |_, _| {},
)
.unwrap();
assert!(tiles > 0, "no tile fitted a picture shorter than the tile");
for i in 0..300 * 40 {
if !known[i] {
assert!((rgb[i * 3] - 0.5).abs() < 1e-4, "pixel {i}");
}
}
// And the model saw the padding as hole, never as black content.
for (_, k) in &model.seen {
assert_eq!(k.len(), 64 * 64);
}
}
#[test]
fn the_fine_passes_run_in_bands_after_the_coarse_one() {
// A 150-tall hole above and below a picture: the coarse pass sees
@@ -824,7 +876,7 @@ mod tests {
#[test]
fn the_context_mirrors_the_top_rows_upward() {
let (rgb, known) = picture(40, 30, 5);
let ctx = MirroredContext::build(&rgb, 40, 30, &known, 48);
let ctx = MirroredContext::build(&rgb, 40, 30, &known, 48, 64);
let x = RING + 10;
let first = RING + 5;
for k in 1..=4 {
@@ -846,7 +898,7 @@ mod tests {
for x in 0..40 {
rgb[(ridge * 40 + x) * 3..(ridge * 40 + x) * 3 + 3].copy_from_slice(&[0.9, 0.1, 0.1]);
}
let ctx = MirroredContext::build(&rgb, 40, 400, &known, 48);
let ctx = MirroredContext::build(&rgb, 40, 400, &known, 48, 64);
let x = RING + 10;
for y in 0..RING + 5 {
let p = (y * ctx.width + x) * 3;
+1 -1
View File
@@ -9,7 +9,7 @@
//! on every rung, and what they cost is the whole story of whether a fill
//! is interactive: 7.4 s a tile under tract, 0.4 s under ONNX Runtime's
//! CPU pool, 23 ms in fp16 and 13 ms in int8 on a laptop's TensorRT
//! (2026-09-19, docs/panorama.md §12).
//! (2026-09-19, docs/dev/panorama.md §12).
//!
//! The model's contract, from the reference `export_inference_model.py`:
//! input `1×4×512×512` float — channel 0 is `mask − 0.5` with 1 where the
+2 -2
View File
@@ -4,7 +4,7 @@
//! Apache-2.0 weights (`models/LICENCE.md`), exported at a fixed shape by
//! `tools/export-xfeat.sh` and loaded through the same `dr-inference-engine`
//! `dr-segment` and `dr-face` use, so this adds no runtime and no C to the
//! tree; what runs it is the device's business (docs/inference.md). ~300 ms
//! tree; what runs it is the device's business (docs/dev/inference.md). ~300 ms
//! per frame on tract on the reference desktop, ~400 ms on the tablet
//! (S15.2, S15.4).
@@ -37,7 +37,7 @@ pub struct XFeat {
}
/// The bytes of both exports compiled into the binary, for whoever compiles
/// engines ahead of the first request (docs/inference.md §6).
/// engines ahead of the first request (docs/dev/inference.md §6).
#[cfg(feature = "embedded-model")]
pub fn embedded_model_bytes() -> [&'static [u8]; 2] {
[EMBEDDED_LANDSCAPE, EMBEDDED_PORTRAIT]
+2 -2
View File
@@ -315,7 +315,7 @@ pub struct DetailPass {
/// 52 render pixels at 4K — holds no spatial frequency a quarter-scale
/// grid cannot represent. Computing it at the render size therefore buys
/// nothing and costs everything: 105 taps over 8.3 M pixels, twice, which
/// measured at 34 ms and is where `docs/technical-debt.md` TD-4 came from.
/// measured at 34 ms and is where `docs/dev/technical-debt.md` TD-4 came from.
/// At a quarter it is a sixteenth of the pixels at a quarter of the
/// radius, and the result is not an approximation of the full-resolution
/// base — it is the same band-limited function, sampled where it is still
@@ -568,7 +568,7 @@ pub fn compose_detail(
/// photograph the photographer thinks they are sharpening.
///
/// It also means ARCH §5.2's stage list, which draws spot removal after
/// texture and clarity, is not what this does — see `docs/spot-removal.md`
/// texture and clarity, is not what this does — see `docs/dev/spot-removal.md`
/// §5.1, which is where the disagreement is written down.
pub fn compose_detail_with(
ops: &[Box<dyn Operation>],
+35 -2
View File
@@ -113,7 +113,7 @@ pub struct EditGraph {
/// a sidecar comes to name one stock while the shader draws another.
film: Option<Film>,
/// TRACES: FR-DEV-8
/// The repairs (`docs/spot-removal.md`).
/// The repairs (`docs/dev/spot-removal.md`).
///
/// Apart from `ops` for the third time and the same reason: a spot is not
/// a scalar, and a list of them is not a slider. It sits beside the masks
@@ -122,7 +122,7 @@ pub struct EditGraph {
/// photograph comes from.
spots: SpotSet,
/// The lens corrections that rewrite coordinates: distortion and lateral
/// chromatic aberration (`docs/architecture.md` §5.2).
/// chromatic aberration (`docs/dev/architecture.md` §5.2).
///
/// Apart from `ops` for the fourth time, and this one is not about shape
/// but about direction. Every [`Operation`] is a function from colour to
@@ -857,6 +857,39 @@ impl EditGraph {
crate::operation::compose_camera_linear(&self.warps, self.framing.baseline(), view)
}
/// TRACES: FR-DEV-3
/// The camera-space tap over one patch of what is on the canvas.
///
/// `patch` is in fractions of the visible region — the coordinates a
/// click on the canvas arrives in — and is laid over this edit's own
/// framing: crop, view, rotation and all, so a fraction of the canvas is
/// a fraction of the probe. Nothing else of the edit: no operation, no
/// mask, no repair. Rendering only the patch is what lets a small target
/// cover every sensor pixel under it rather than sampling one in fifty;
/// see [`crate::operation::compose_camera_probe`] for why the white
/// balance picker reads from here and not from the display.
///
/// The patch is centred where asked and held to the view's own minimum
/// extent: at a deep zoom a patch a fraction of the view would be
/// smaller than a view may be, and letting `set_view` widen it from one
/// corner would move the sample off the point that was clicked.
pub fn compose_camera_probe(&self, patch: crate::framing::CropRect) -> ComposedShader {
use crate::framing::CropRect;
let mut framing = self.framing;
let view = framing.view();
let width = (patch.width * view.width).max(CropRect::MIN_EXTENT);
let height = (patch.height * view.height).max(CropRect::MIN_EXTENT);
let cx = view.x + (patch.x + patch.width * 0.5) * view.width;
let cy = view.y + (patch.y + patch.height * 0.5) * view.height;
framing.set_view(CropRect {
x: (cx - width * 0.5).clamp(0.0, 1.0 - width),
y: (cy - height * 0.5).clamp(0.0, 1.0 - height),
width,
height,
});
crate::operation::compose_camera_probe(&self.warps, &framing)
}
/// TRACES: FR-DEV-19c
/// [`Self::compose_for`], with one layer's mask drawn over the picture.
///
+2 -1
View File
@@ -65,7 +65,8 @@ pub use history::{Edit, Entry as HistoryEntry, History, Step};
pub use lens::{compose_warps, ComposedWarp, LensProfile, Tca, Warp};
pub use operation::{
compose, compose_with_framing, Affects, ComposedShader, Helper, Invalidation, Operation,
OutputMode, Uniform, BASE_CURVE_POINTS, BASE_CURVE_UNIFORM_OFFSET, RESERVED_UNIFORM_FIELDS,
OutputMode, Uniform, BASE_CURVE_POINTS, BASE_CURVE_UNIFORM_OFFSET, CLIP_ONSET,
RESERVED_UNIFORM_FIELDS,
};
pub use preset::{LibraryParseError, NameError, Preset, PresetLibrary, Scope};
pub use sidecar::{Sidecar, Version};
+5 -5
View File
@@ -27,7 +27,7 @@
//! [`MaskSource::Regions`] stores integers naming regions in the segmentation
//! hierarchy (`dr-segment`). That choice is what makes a mask diffable, cheap
//! in a sidecar, and mergeable per-field under FR-NC-9 — three properties a
//! stored raster has none of (docs/segmentation.md §1). Two devices that
//! stored raster has none of (docs/dev/segmentation.md §1). Two devices that
//! select the same subject produce the same small sorted list, and a sync
//! conflict between them is resolvable rather than a binary blob fight.
//!
@@ -564,7 +564,7 @@ pub enum MaskSource {
/// This is what the watershed and the semantic model exist to produce.
/// Selecting a subject means "the regions the model's instance covers",
/// and the resulting edge is the watershed's, which is to say the image's
/// own (docs/segmentation.md §5).
/// own (docs/dev/segmentation.md §5).
Regions {
/// Which segmentation these ids index into.
///
@@ -590,7 +590,7 @@ pub enum MaskSource {
/// **The primary way a local adjustment is made.** The watershed hierarchy
/// this crate was first built around does not survive a photograph: its
/// saddles are near zero almost everywhere, so a global cut collapses the
/// frame into one region plus noise (docs/segmentation.md §15). A model
/// frame into one region plus noise (docs/dev/segmentation.md §15). A model
/// instance is a whole object, found as one thing, and needs no ladder.
///
/// The trade is that the boundary is the model's — a quarter-resolution
@@ -625,7 +625,7 @@ pub enum MaskSource {
/// reason both exist. A subject is *one* instance — this dog, not that one
/// — found by a COCO-trained instance model. A category is *all* the sky,
/// or all the foliage, from an ADE20K-trained semantic model that has no
/// notion of instances at all (docs/segmentation.md §16).
/// notion of instances at all (docs/dev/segmentation.md §16).
///
/// So this is what a global grade attaches to: lift the sky, desaturate
/// the vegetation, warm the architecture. Asking it for "that person
@@ -2231,7 +2231,7 @@ fn reveal_block(slot: usize, layer: &MaskLayer, style: RevealStyle, colour: [f32
/// **not** a hash of the label field: that would be a readback on a path that
/// must not have one (ARCH §6.1), and would also make the signature depend on
/// float arithmetic whose cross-vendor determinism is exactly the open
/// question (docs/segmentation.md §6, M5).
/// question (docs/dev/segmentation.md §6, M5).
pub fn segmentation_signature(width: u32, height: u32, regions: u32, tuning: u64) -> u64 {
// FNV-1a over the four fields. Small, dependency-free, and adequate: this
// guards against accidental mismatch, not against a forged sidecar.
+15 -9
View File
@@ -69,20 +69,26 @@ const ROUNDS: u32 = 3;
/// Below this a channel carries no ratio worth balancing.
///
/// A sample in the deep shadows, or one taken on a blown highlight where a
/// channel has already clipped to nothing, has no white balance in it: the
/// logarithms below would run away and the picker would slam a slider to its
/// stop. Refusing is the honest answer, and the caller reports that the point
/// was not usable rather than moving the photograph.
/// A sample in the deep shadows has no white balance in it: the logarithms
/// below would run away and the picker would slam a slider to its stop.
/// Refusing is the honest answer, and the caller reports that the point was
/// not usable rather than moving the photograph. The other end — a blown
/// highlight, where every channel has stopped counting — is refused by the
/// caller before the sample is taken, because only the caller can see the
/// sensor value; see [`crate::operation::CLIP_ONSET`].
const FLOOR: f32 = 1e-4;
/// TRACES: FR-DEV-3
/// Move the graph so that `sample` renders neutral.
///
/// `sample` is linear RGB, as the operation's own gains multiply it — that is,
/// measured with the sampling operation at its defaults. Returns whether the
/// graph was moved: `false` where the chain offers no white point widget, or
/// where the colour has no balance in it to correct.
/// `sample` is the linear triple the operation's own gains multiply — camera
/// RGB with the camera's as-shot balance on, *before* the body's base curve
/// and matrix, and with the sampling operation at its defaults. Not the
/// pixel on the screen: the matrix mixes the channels on the way there, so
/// a colour read after it does not answer to these gains, and a solve over
/// one lands somewhere no sample asked for. Returns whether the graph was
/// moved: `false` where the chain offers no white point widget, or where the
/// colour has no balance in it to correct.
///
/// **Absolute, not relative.** The values written depend on the colour and not
/// on where the sliders happened to be, so sampling the same wall twice lands
+68 -9
View File
@@ -54,7 +54,7 @@ pub enum Affects {
/// A pixel's *neighbourhood* — sharpening, noise reduction, clarity,
/// texture, dehaze, spot removal.
///
/// The seam `docs/requirements.md` §3.3 designed and nothing cut until
/// The seam `docs/dev/requirements.md` §3.3 designed and nothing cut until
/// [`crate::detail`] existed. It is a separate variant rather than a flavour
/// of `Colour` because it is a separate *dispatch*: a fragment in the fused
/// pass is handed a colour and has no way back to a coordinate, so a
@@ -431,9 +431,11 @@ pub enum OutputMode {
/// no operations, and the caller fills the reserved uniforms neutral —
/// unit white balance, identity matrix, base curve off — so what is
/// stored is the sensor's own numbers, demosaiced and undistorted. Only
/// [`compose_camera_linear`] produces it, and only
/// `AdjustPass::render_camera_linear` accepts it, so the neutral
/// uniforms cannot be forgotten by a caller that composed it by mistake.
/// [`compose_camera_probe`] produces it — for a merge through
/// [`compose_camera_linear`], and for the white balance picker under the
/// edit's own framing — and only `AdjustPass::render_camera_linear`
/// accepts it, so the neutral uniforms cannot be forgotten by a caller
/// that composed it by mistake.
///
/// Thirty-two bits rather than sixteen because the composite is written
/// back as a RAW at the sensor's scale (FR-MRG-3): a 14-bit sensor has
@@ -490,6 +492,19 @@ pub const BASE_CURVE_UNIFORM_OFFSET: usize = 16;
/// selection in [`compose_full`].
pub const BASE_CURVE_POINTS: usize = 5;
/// Where the highlight desaturation begins: the fraction of the white level
/// above which a photosite is treated as clipped.
///
/// A photosite this close to saturation has stopped counting, so its ratio
/// to its neighbours is not a colour. The generated prologue fades a pixel
/// above this toward a neutral of the same brightness before any operation
/// runs, and the white balance probe refuses to sample one: a blown sky is
/// sensor white, which the as-shot multipliers make magenta, and a solve
/// over that slams tint to its stop. One number, so the two cannot drift
/// apart — a probe that accepted what the shader had already desaturated
/// would be balancing against a pixel the photographer cannot see.
pub const CLIP_ONSET: f32 = 0.985;
/// Where an operation's own uniforms begin in the generated block.
///
/// The base fields, then framing's. Exported because `dr-gpu` writes the
@@ -589,7 +604,9 @@ pub fn compose_full_revealing(
warps: &[Box<dyn crate::lens::Warp>],
reveal: Option<&crate::mask::Reveal>,
) -> ComposedShader {
compose_inner(ops, framing, output, masks, spots, warps, reveal, None)
compose_inner(
ops, framing, output, masks, spots, warps, reveal, None, false,
)
}
/// TRACES: FR-MRG-2
@@ -614,15 +631,50 @@ pub fn compose_camera_linear(
// upright too, and `view` is a fraction of the upright frame.
framing.set_baseline(baseline);
framing.set_view(view);
compose_camera_tap(warps, &framing, false)
}
/// TRACES: FR-DEV-3
/// The same tap under a framing the caller chose: the white balance probe.
///
/// A neutral picked off the canvas has to be measured in the space the
/// white balance gains multiply, and that is camera RGB — the operation
/// runs before the body's matrix, and a probe read after the matrix would
/// be solving the wrong equation on any body whose matrix mixes the
/// channels, which is every body. It also has to be measured at the pixel
/// the canvas is showing, which is why this takes the edit's own framing
/// where a merge passes the file's orientation and a tile.
///
/// **Interpolated whatever the framing says.** The point of rendering a
/// patch is to average what is under it, and the nearest sampling an
/// unrotated frame otherwise gets is a comb: at two source pixels per
/// probe pixel it lands on the same column of any pattern every time, and
/// the average of a thousand samples is then the average of nothing.
pub fn compose_camera_probe(
warps: &[Box<dyn crate::lens::Warp>],
framing: &Framing,
) -> ComposedShader {
compose_camera_tap(warps, framing, true)
}
/// The camera-space tap proper: no operations, `rgba32float`, and the
/// profile uniforms left for the GPU side to fill neutral. `smooth` forces
/// the interpolating sampler; see the two callers for who wants it and why.
fn compose_camera_tap(
warps: &[Box<dyn crate::lens::Warp>],
framing: &Framing,
smooth: bool,
) -> ComposedShader {
compose_inner(
&[],
&framing,
framing,
ColourSpace::Srgb,
&MaskStack::new(),
&crate::spot::SpotSet::new(),
warps,
None,
Some(OutputMode::CameraLinear),
smooth,
)
}
@@ -636,6 +688,7 @@ fn compose_inner(
warps: &[Box<dyn crate::lens::Warp>],
reveal: Option<&crate::mask::Reveal>,
forced: Option<OutputMode>,
smooth: bool,
) -> ComposedShader {
// The lens corrections, composed into one coordinate transform. Beside
// `framing` because they are the other half of the same stage: framing
@@ -668,7 +721,7 @@ fn compose_inner(
// twice and be bound to a texture of the wrong format.
//
// `forced` is the one exception, and it is not a caller flag in the
// sense above: `compose_camera_linear` is the only function that passes
// sense above: `compose_camera_probe` is the only function that passes
// it, with an empty operation list, and the mode it forces has its own
// storage format and its own render entry on the GPU side.
let output_mode = forced.unwrap_or(
@@ -842,7 +895,10 @@ fn compose_inner(
// framing alone — which is what this did before the warps existed — would
// have nearest-neighboured a distortion correction on an unstraightened
// frame, and the aliasing would have looked like a bad profile.
let interpolate = framing.needs_interpolation() || warp.is_active();
// `smooth` is the third reason, and the only one a caller states: the
// white balance probe averages a patch and cannot do that through a
// nearest-neighbour comb (see `compose_camera_probe`).
let interpolate = smooth || framing.needs_interpolation() || warp.is_active();
// Declared ahead of the warp block, which assigns to them. They enter
// equal to `p` so that a chain mixing a splitting warp with a
@@ -997,6 +1053,9 @@ fn compose_inner(
.to_string()
};
// Formatted with Rust's `Display` so the shader reads the same threshold
// the probe checks against; see `CLIP_ONSET`.
let clip_onset = CLIP_ONSET;
let source = format!(
"// GENERATED — do not edit.
//
@@ -1067,7 +1126,7 @@ fn main(@builtin(global_invocation_id) gid: vec3<u32>) {{
// A photosite at its white level carries no colour information — every
// channel simply stopped counting — so the balance below must not be
// allowed to tint it.
let clipped = smoothstep(0.985, 1.0, max(c.r, max(c.g, c.b)));
let clipped = smoothstep({clip_onset}, 1.0, max(c.r, max(c.g, c.b)));
c = c * u.as_shot_wb.rgb;
+1 -1
View File
@@ -99,7 +99,7 @@
//! A minimum over a patch is separable, as a Gaussian is: minimum along x,
//! then along y. That alone is not enough. The patch is 1% of the shorter edge
//! — 61 taps across at 4K — and two passes of 61 taps is the arithmetic that
//! measured 34 ms for clarity and became `docs/technical-debt.md` TD-4.
//! measured 34 ms for clarity and became `docs/dev/technical-debt.md` TD-4.
//!
//! A minimum has a property a Gaussian does not: **erosions compose by adding
//! their structuring elements**. The minimum over a contiguous run of `d`
+2 -2
View File
@@ -150,7 +150,7 @@
//! the artefact this control must not have.
//!
//! Run at the render size, that measured **34 ms at 4K** — seven times the
//! entire fused point chain, for one slider — which is `docs/technical-debt.md`
//! entire fused point chain, for one slider — which is `docs/dev/technical-debt.md`
//! TD-4 and is what [`Recipe::base_scale`] now answers. The base is computed on
//! a grid a quarter the size on each axis: a sixteenth of the pixels at a
//! quarter of the radius.
@@ -271,7 +271,7 @@ impl Band for Coarse {
threshold: 0.35,
gain: 1.0,
midtone_taper: true,
// A quarter, which is what `docs/technical-debt.md` TD-4 bought back.
// A quarter, which is what `docs/dev/technical-debt.md` TD-4 bought back.
//
// σ is 1.2% of the shorter edge — 26 px at 4K — so the base holds no
// spatial frequency anywhere near the quarter-scale Nyquist of one
+1 -1
View File
@@ -217,7 +217,7 @@ pub struct Version {
/// graph's film cleared and the caller re-bakes — see `EditGraph::set_film`.
pub film: Option<FilmRef>,
/// TRACES: FR-DEV-8 | FR-NC-9
/// The repairs (`docs/spot-removal.md`).
/// The repairs (`docs/dev/spot-removal.md`).
///
/// A line per spot, keyed `spot.<id>`, rather than a block per spot as a
/// mask gets: a spot is eight numbers, and sixty-four blocks would bury the
+3 -3
View File
@@ -6,7 +6,7 @@
//! are blended. No pixels are stored, here or anywhere: the shader draws the
//! repair from these numbers every time the photograph is rendered, which is
//! what makes it non-destructive, cheap to sync, and undoable
//! (`docs/spot-removal.md`).
//! (`docs/dev/spot-removal.md`).
//!
//! # Why this is not an operation
//!
@@ -61,7 +61,7 @@ pub const MAX_SPOTS: usize = 64;
/// bounds something that is otherwise unbounded: a detail pass declares how far
/// it reads from the pixel it writes, and for a spot that is the offset plus
/// the radius. An unbounded offset is an unbounded halo, which is a pass the
/// tile scheduler cannot plan (ARCH §5.3, `docs/spot-removal.md` §5.3).
/// tile scheduler cannot plan (ARCH §5.3, `docs/dev/spot-removal.md` §5.3).
pub const MAX_SOURCE_DISTANCE: f32 = 0.5;
/// The radius a new spot starts at, in frame units.
@@ -210,7 +210,7 @@ impl Spot {
/// **FR-DEV-8 asks for automatic source placement, and this is the cheap
/// half of it.** The good half searches the photograph for a patch whose
/// surroundings match — a compute dispatch scoring candidate offsets, and
/// one small readback when the spot is created (`docs/spot-removal.md`
/// one small readback when the spot is created (`docs/dev/spot-removal.md`
/// §8). This is what stands in for it, and it is worth having on its own
/// terms rather than as a placeholder: dust sits on skies, skies are
/// smooth, and a patch two and a half radii away is nearly always the same
+1 -1
View File
@@ -13,7 +13,7 @@ log.workspace = true
# Inference. `ort` is the API; **what runs it is `dr-inference-engine`'s
# business** — tract, or an ONNX Runtime the app found on disk, on whichever
# provider the device has (docs/inference.md). This crate never names either.
# provider the device has (docs/dev/inference.md). This crate never names either.
ort = { workspace = true, optional = true }
dr-inference-engine = { workspace = true, optional = true }
ndarray = { workspace = true, optional = true }
+1 -1
View File
@@ -34,7 +34,7 @@
//! And doing it here buys two things a shader could not. It is **exactly
//! deterministic**, which matters because masks reach the sidecar as indices
//! and a field that varied by vendor would mean a mask meaning one thing on
//! the desktop and another on the phone (docs/segmentation.md §6, M5). And it
//! the desktop and another on the phone (docs/dev/segmentation.md §6, M5). And it
//! is testable against hand-computed distances with no adapter present.
//!
//! # The transform
+1 -1
View File
@@ -40,7 +40,7 @@ pub struct Edge {
/// A partition of the image into labelled regions, plus how they adjoin.
///
/// The shared interface from docs/segmentation.md §2: arm A produces this
/// The shared interface from docs/dev/segmentation.md §2: arm A produces this
/// from a watershed, arm B would produce it from a class map, and the
/// consumers above cannot tell which.
#[derive(Debug, Clone, PartialEq)]
+1 -1
View File
@@ -1,5 +1,5 @@
//! TRACES: FR-DEV-3i
//! Region segmentation for local masking (S15, docs/segmentation.md).
//! Region segmentation for local masking (S15, docs/dev/segmentation.md).
//!
//! Local adjustments need to know where the image's regions are before they
//! can snap a mask to one. This crate is that map, and it is deliberately
+2 -2
View File
@@ -1,6 +1,6 @@
//! Arm C — semantic instances as a prior over the watershed merge order.
//!
//! docs/segmentation.md §5. The spec calls this the expected winner and it is
//! docs/dev/segmentation.md §5. The spec calls this the expected winner and it is
//! what ships, for a reason that survives the model turning out to be narrower
//! than §4 assumed: the two arms fail in *opposite* directions, so each one
//! covers the other's failure.
@@ -233,7 +233,7 @@ pub fn apply_semantic_prior(
/// This is the interaction the whole spike exists to enable, and the reason it
/// returns *region ids* rather than a raster: a mask that is a set of integers
/// is diffable, mergeable at node level under FR-NC-9, and cheap in a sidecar
/// (docs/segmentation.md §1). A raster is none of those.
/// (docs/dev/segmentation.md §1). A raster is none of those.
///
/// The returned ids are sorted, so the same click always produces the same
/// mask — which is what lets it be a cache key.
+5 -5
View File
@@ -60,7 +60,7 @@
//! So the last step is a **marker-based watershed**. The mask is eroded to
//! give two markers — confidently inside, confidently outside — and the flood
//! runs in the ribbon left between them, meeting along the most expensive line
//! it can find. The cost is a sum of terms, as docs/segmentation.md §2 says it
//! it can find. The cost is a sum of terms, as docs/dev/segmentation.md §2 says it
//! should be: the photograph's own edges, and the colour model's disagreement.
//!
//! Markers are what make this the right shape rather than the watershed §15
@@ -244,7 +244,7 @@ pub struct RefineOptions {
/// How much the photograph's own edges count against the colour model in
/// the flood's cost, `0.0..=1.0`.
///
/// docs/segmentation.md §2 specifies the cost as *a sum of terms* — image
/// docs/dev/segmentation.md §2 specifies the cost as *a sum of terms* — image
/// gradient always available, semantic evidence added when a model is
/// present — and this is the mix. At one the boundary lands purely on the
/// strongest edge in the band; at zero purely where the colour verdict
@@ -378,7 +378,7 @@ const MAX_SAMPLES: usize = 20_000;
/// floating-point comparison is a stopping rule that can differ between
/// machines, and a mask that differs between machines reaches the sidecar as
/// indices meaning one thing on the desktop and another on the phone
/// (docs/segmentation.md §6).
/// (docs/dev/segmentation.md §6).
const ITERATIONS: usize = 12;
/// Half-width of the verdict scale, in nats.
@@ -676,7 +676,7 @@ impl Refinement {
/// fronts meet along the most expensive line in the ribbon — which is the
/// watershed, and which is where the boundary belongs.
///
/// The cost is a sum of terms, as docs/segmentation.md §2 says it should
/// The cost is a sum of terms, as docs/dev/segmentation.md §2 says it should
/// be: the photograph's own edges, and the colour model's disagreement.
/// Neither alone is right. An edge with no colour meaning is a texture,
/// and a colour change with no edge is a gradient.
@@ -889,7 +889,7 @@ fn neighbours(p: usize, w: usize, h: usize) -> impl Iterator<Item = usize> {
/// Edge strength over the opponent features, as one byte per pixel.
///
/// Sobel over the same three numbers the colour model is fitted on, rather
/// than over plain luma — docs/segmentation.md §3 is explicit that a
/// than over plain luma — docs/dev/segmentation.md §3 is explicit that a
/// channel-weighted RGB gradient reads a saturated red edge as weaker than it
/// looks, and a flag against sky is exactly that edge.
///
+3 -3
View File
@@ -1,4 +1,4 @@
//! Semantic segmentation — arm B (S15, docs/segmentation.md §4).
//! Semantic segmentation — arm B (S15, docs/dev/segmentation.md §4).
//!
//! Runs a YOLO instance-segmentation graph over a proxy-resolution image and
//! returns the instances it found: a class, a score, a box, and a soft mask
@@ -209,7 +209,7 @@ const EMBEDDED_MODEL: &[u8] = include_bytes!("../../../models/segment/yolo26n-se
const EMBEDDED_CLASSES: &str = include_str!("../../../models/segment/yolo26n-seg.classes.json");
/// The bytes of the model that ships with this crate, for whoever compiles
/// engines ahead of the first request (docs/inference.md §6).
/// engines ahead of the first request (docs/dev/inference.md §6).
#[cfg(feature = "embedded-model")]
pub fn embedded_model_bytes() -> &'static [u8] {
EMBEDDED_MODEL
@@ -237,7 +237,7 @@ impl SemanticModel {
pub fn from_bytes(bytes: &[u8], classes: Vec<Arc<str>>) -> Result<Self, SegmentError> {
// The f32 graph on whatever the device's backend is. An int8 form
// for the Hexagon waits on docs/inference.md §10 M7 — the mask
// for the Hexagon waits on docs/dev/inference.md §10 M7 — the mask
// boundary has to be measured before it moves.
let session = dr_inference_engine::open(
dr_inference_engine::Role::Segmenter,
+2 -2
View File
@@ -177,7 +177,7 @@ pub struct FaceSettings {
/// Which SCRFD graph the indexing pass detects with.
///
/// Three exports of one architecture, differing only in how much computation
/// they spend, and docs/faces.md §12.3 is the measurement that made this a
/// they spend, and docs/dev/faces.md §12.3 is the measurement that made this a
/// choice rather than a constant: over the same photographs the cheapest one
/// misses the small faces in a group and reports a dog a dozen times, the
/// middle one finds 14% more faces for 12% more time, and the largest a
@@ -261,7 +261,7 @@ impl FaceDetector {
}
}
/// The id when the detector runs in its int8 form (docs/inference.md §7).
/// The id when the detector runs in its int8 form (docs/dev/inference.md §7).
///
/// A different detector: it finds a different set of faces, so it is a
/// different population of detections. The embedder half is unchanged,
+1 -1
View File
@@ -48,7 +48,7 @@ pointers there; the build detects that and stops rather than shipping them.
This exists because Android offers no other route to a model: the directory the app reads from is
inside app-private storage, `run-as` needs a debuggable build, and the app has no picker and no
fetch. docs/faces.md §2.2a is the decision and its limits — these files come back out before anything
fetch. docs/dev/faces.md §2.2a is the decision and its limits — these files come back out before anything
is published.
## Java in the APK
+2 -2
View File
@@ -258,7 +258,7 @@ fi
cp "${SO}" "${OUT}/staging/lib/${ABI}/libdarkroom.so"
cp "${DEX}" "${OUT}/staging/classes.dex"
# The inference runtime (docs/inference.md §3): ONNX Runtime and Qualcomm's
# The inference runtime (docs/dev/inference.md §3): ONNX Runtime and Qualcomm's
# Hexagon backend, beside libdarkroom.so so the app finds them in its own
# native library directory. The build links none of it — the app dlopens
# `libonnxruntime.so` at launch and runs on tract if it is not there — so an
@@ -278,7 +278,7 @@ else
fi
# The models. Android has no other route to one — app-private storage is not
# user-reachable and the in-app fetch is unbuilt (docs/faces.md §2.2a) — so
# user-reachable and the in-app fetch is unbuilt (docs/dev/faces.md §2.2a) — so
# they go in the APK and `android_main` unpacks them on first launch. The
# sources are `models/face/` and `models/scene/`, shared with the Arch package
# rather than living under this one platform's directory.
+1 -1
View File
@@ -73,7 +73,7 @@ SO="${CACHE}/target/jniLibs/${ABI}/libdarkroom.so"
# of KEYSTORE_PASS (see its header), and the keystore has to be reachable from
# inside the container, so a host path in KEYSTORE is copied under the mounted
# target directory for the duration of the build and removed after. The
# passwords travel as environment, never as arguments -- docs/android-signing.md
# passwords travel as environment, never as arguments -- docs/dev/android-signing.md
# has the incantation.
# ---------------------------------------------------------------------------
echo "==> packaging APK"
+2 -2
View File
@@ -1,6 +1,6 @@
# DarkRoom — reproducible Windows cross-build environment
#
# Everything docs/windows.md §2 names: Rust with the GNU Windows target, the
# Everything docs/dev/windows.md §2 names: Rust with the GNU Windows target, the
# MinGW-w64 cross compiler it links with, NSIS to build the installer, and Wine
# to smoke-test the result. Both CI and local builds use this image, so "works
# on my machine" and "works in CI" are the same machine — the same argument
@@ -77,7 +77,7 @@ RUN curl -fsSL https://sh.rustup.rs | sh -s -- \
# libwinpthread the Rust target's own MinGW pieces were built against, and
# picking the other produces link errors that read as if std were missing.
#
# The runtime is linked statically (docs/windows.md §2) so the installer
# The runtime is linked statically (docs/dev/windows.md §2) so the installer
# carries one file. `-static-libgcc` is all it takes: rustc's windows-gnu
# target links its own copy of winpthread in self-contained mode, so nothing
# imports libwinpthread-1.dll — the smoke test's objdump step is what checks
+2 -2
View File
@@ -1,7 +1,7 @@
# Windows cross-build environment
Reproducible container for building the Windows executable and its installer from Linux. The
specification is [docs/windows.md](../../docs/windows.md); this directory is what it turned into,
specification is [docs/dev/windows.md](../../docs/dev/windows.md); this directory is what it turned into,
and every departure from the spec's first draft is recorded in the Dockerfile's comments.
## Use
@@ -16,7 +16,7 @@ and every departure from the spec's first draft is recorded in the Dockerfile's
# Build the installer from that binary
./docker/windows/build.sh docker/windows/package.sh
# Smoke-test under Wine (docs/windows.md §6)
# Smoke-test under Wine (docs/dev/windows.md §6)
./docker/windows/build.sh wine target-windows/x86_64-pc-windows-gnu/release/darkroom-desktop.exe --version
./docker/windows/build.sh wine target-windows/installer/DarkRoom-0.12.0-x86_64-setup.exe /S
+8 -2
View File
@@ -7,7 +7,7 @@
# to have run in the same target directory. Produces
# DarkRoom-<version>-x86_64-setup.exe in $OUT (default: target-windows/installer).
#
# Runs inside the container, where makensis is; docs/windows.md §5 is the
# Runs inside the container, where makensis is; docs/dev/windows.md §5 is the
# specification this implements.
set -euo pipefail
@@ -48,7 +48,13 @@ sed 's/$/\r/' "${REPO}/LICENSE" > "${STAGE}/LICENSE"
# tract on the user's machine with a message about a broken graph rather than
# a checkout that needed `git lfs pull`. Only the weights are checked; the
# scene model's vocabulary and category descriptor are legitimately small.
for dir in face scene; do
#
# The directories are the ones the APK stages (assemble-apk.sh) and the Arch
# package installs: the face pair and its eye-state models, the scene model
# with its two descriptors, and the panorama border filler. The installer
# smoke test counts the same directories, so a model added here is expected
# there without a number to update.
for dir in face scene inpaint; do
for f in "${REPO}/models/${dir}"/*; do
case "$(basename "${f}")" in
README.md) continue ;;
+86
View File
@@ -0,0 +1,86 @@
# Documentation
Two audiences, two folders. Most people want the first table and never the
second.
## Using DarkRoom
| | |
|---|---|
| [The manual](manual/README.md) | Every feature, pictured from the application itself — opening a library, rating and filing, developing, local masks, repair, film, panoramas, export |
| [How it is driven](gestures.md) | Every gesture and shortcut, by screen. Generated from the code, so it cannot describe one the application does not have |
The [top-level README](../README.md) says what DarkRoom is, how to get it on
each platform, and what is still missing.
## Changing DarkRoom
Everything under [`dev/`](dev/) is for someone working on the code. Start with
[CONTRIBUTING.md](../CONTRIBUTING.md), which says how to land a first change
without reading the rest.
**The register and the record.** What must be built, how it is built, and how
far along it is.
| | |
|---|---|
| [requirements.md](dev/requirements.md) | What the software must do — the numbered register, the decisions (D-numbers) and the spikes (S-numbers) |
| [architecture.md](dev/architecture.md) | How it is built — crates, the GPU pipeline, the data model, sync |
| [traceability.md](dev/traceability.md) | Generated: which requirement is claimed by which file. Never edited by hand |
| [outstanding.md](dev/outstanding.md) | What is specified and not built, and whether that is a decision or a gap |
| [technical-debt.md](dev/technical-debt.md) | Compromises taken deliberately, each with the condition that retires it |
| [code-health.md](dev/code-health.md) | What a contribution costs, per seam, measured |
**Designs, one per subsystem.** Each is the specification the code was built
to, kept current as the code moved.
| | |
|---|---|
| [catalog.md](dev/catalog.md) | The index, the library view, incremental scan, the job queue |
| [storage.md](dev/storage.md) | Storage backends: the seam a folder, a sync client and a Nextcloud account share |
| [faces.md](dev/faces.md) | Face detection, identity, clustering and the eye-state models |
| [segmentation.md](dev/segmentation.md) | How the application finds the regions a local mask snaps to |
| [mask-editing.md](dev/mask-editing.md) | Painting, erasing and combining masks |
| [spot-removal.md](dev/spot-removal.md) | Clone and heal as parameters in the edit graph |
| [panorama.md](dev/panorama.md) | Alignment, projection, the chunked composite and the border fill |
| [inference.md](dev/inference.md) | The neural runtime and model chosen per device, with the measurements |
| [display-and-extension.md](dev/display-and-extension.md) | The display contract, and why the fused pipeline is already most of a plugin format |
| [view-composition.md](dev/view-composition.md) | A controller for the display layer |
| [ui-navigation.md](dev/ui-navigation.md) | Finding things in the interface once there are many |
**Measurements.** Numbers committed so a regression is a diff rather than a
recollection.
| | |
|---|---|
| [benchmarks.md](dev/benchmarks.md) | The per-commit suite: what it covers, what it does not, how to read a failure |
| [bench-baseline.json](dev/bench-baseline.json) | The committed numbers the suite checks against |
| [frame-budget.md](dev/frame-budget.md) | What a frame costs on each device, and the decision those figures settled |
**Platforms and distribution.**
| | |
|---|---|
| [distribution.md](dev/distribution.md) | Which channels v1 targets and what each one constrains |
| [windows.md](dev/windows.md) | The Windows installer, cross-built from the Linux CI |
| [android-signing.md](dev/android-signing.md) | Which key signs the APK, and keeping it |
**Archive.** Kept as the record of what was asked for, not as plans.
| | |
|---|---|
| [milestone-v0.1.md](dev/archive/milestone-v0.1.md) | The first milestone, delivered 2026-08-30 and superseded |
| [ui-refinement.md](dev/archive/ui-refinement.md) | How the interface should look; succeeded by [ui-navigation.md](dev/ui-navigation.md) |
## Conventions
Two files here are generated and must not be edited by hand:
`gestures.md` and `dev/traceability.md`. Both come from
`cargo run -p traceability` and the pre-commit hook keeps them in step with
the tree. The manual's pictures are recorded by
[`tools/manual`](../tools/manual/README.md) and live in LFS.
A design document links to the requirements it satisfies and to the code
that satisfies them. When the code moves, the link moves with it; a document
that has stopped being true goes to `dev/archive/` with a note saying what
replaced it, rather than being deleted.
@@ -4,8 +4,8 @@
> Kept as the record of what the first milestone asked for, not as a plan.
> Everything below shipped, and the application went well past it — see
> [outstanding.md](outstanding.md) for what is still missing at 0.9.0.
**Companion to:** [requirements.md](requirements.md) · [architecture.md](architecture.md)
> [outstanding.md](../outstanding.md) for what is still missing at 0.9.0.
**Companion to:** [requirements.md](../requirements.md) · [architecture.md](../architecture.md)
The first buildable milestone: connect to a Nextcloud folder, index it locally, and display RAW
previews on both Linux and Android.
@@ -39,7 +39,7 @@ building on sand.
## 2. The four assumptions under test
Each maps to a spike in [requirements.md §9](requirements.md).
Each maps to a spike in [requirements.md §9](../requirements.md).
| # | Assumption | If wrong | Spike |
|---|---|---|---|
@@ -55,7 +55,7 @@ all, so it should be proven in the first week, before catalog or sync work begin
## 3. Functional scope
Requirement IDs reference [requirements.md](requirements.md); a v0.1 suffix marks a reduced subset
Requirement IDs reference [requirements.md](../requirements.md); a v0.1 suffix marks a reduced subset
of the full requirement.
### 3.1 Account and connection
@@ -2,9 +2,9 @@
**Status:** Built, not yet recorded · 2026-08-30
**Companion to:** [requirements.md](requirements.md) §4.1 (performance targets) · §8 (verification)
**Instrument:** [`tools/bench`](../tools/bench) — `cargo run --release -p dr-bench -- check`
**Instrument:** [`tools/bench`](../../tools/bench) — `cargo run --release -p dr-bench -- check`
**Committed numbers:** [`bench-baseline.json`](bench-baseline.json)
**GPU half:** [`core/dr-gpu/tests/frame_budget.rs`](../core/dr-gpu/tests/frame_budget.rs) ·
**GPU half:** [`core/dr-gpu/tests/frame_budget.rs`](../../core/dr-gpu/tests/frame_budget.rs) ·
[frame-budget.md](frame-budget.md)
§8 has said since it was written that performance is verified by *"an automated
@@ -60,7 +60,7 @@ Two of those rows carry a qualifier, and the qualifiers are the point.
library view cannot paint without: `Catalog::open` (which connects, migrates and
**backfills**, and the backfill is three passes over the images table on every
open), `count`, the first 400-row `window`, and the monthly `timeline`. Tagged
`TRACES: NFR-P1` in [`tools/bench/src/catalog_open.rs`](../tools/bench/src/catalog_open.rs),
`TRACES: NFR-P1` in [`tools/bench/src/catalog_open.rs`](../../tools/bench/src/catalog_open.rs),
because a build that breaks it fails this gate.
**NFR-P3 — ≥ 100 images per second on the embedded preview path.** The
@@ -69,7 +69,7 @@ per-image work is exactly what `spawn_thumbnail_sweep` does — `decode_jpeg`,
`ThumbStore::put` — arranged in the same shape: chunks of 96, lanes owning
disjoint slices, and the single thread that owns the store writing the finished
chunk. Tagged `TRACES: NFR-P3` in
[`tools/bench/src/thumbnails.rs`](../tools/bench/src/thumbnails.rs).
[`tools/bench/src/thumbnails.rs`](../../tools/bench/src/thumbnails.rs).
### Requirements this can only half-answer, and is not tagged for
@@ -114,7 +114,7 @@ instead of ~2 TB, and neither half is flattered by that.
| Rows | 50,000 images, 50,000 default versions, 400 folders, one root |
| Capture times | Twelve years from a fixed epoch, so the timeline has ~144 monthly buckets |
| Sources | 12 synthesised JPEGs at 1620 × 1080 — the size `dr-decode` records a CR2 carrying in IFD2 |
| Seed | 20260829, in [`tools/bench/src/main.rs`](../tools/bench/src/main.rs) |
| Seed | 20260829, in [`tools/bench/src/main.rs`](../../tools/bench/src/main.rs) |
| Location | `$DR_BENCH_DIR`, else the system temporary directory |
It is reproducible from the seed, and a `stamp.json` beside it records what it
+9
View File
@@ -530,6 +530,15 @@ Three invariants, each tested:
| Re-storing an existing id updates in place, never migrates | Migrating would rewrite a sealed shard |
| Merging another client's shard is insert-only and idempotent | Both copies derive from the same bytes by the same code, so neither is better; preferring ours avoids dirtying a shard others have synced |
**When the exchange runs.** Corrected 2026-09-20. It fired only after the metadata sweep — hours
on a large library — so a fresh device re-derived every thumbnail it looked at, re-detected faces
and re-read every header before adopting the shards and snapshot that held all of it. It now also
fires the moment the scan completes, which is the first moment the rows the merges key on exist,
and the sweep starts behind it. In steady state that pass is one listing. The catalog merge also
takes **capture metadata** (`captured_at`, offset, camera, lens, ISO) for images still at
`metadata_state < 2`, matched by `oc:fileid` — a date is a fact about the file's bytes, not local
state, and the snapshot already carried it; the sweep's per-chunk query then finds nothing left.
**The transfer**, in `dr-ui`'s `derived_sync`, exchanges shards with `.darkroom-derived/` under the
library root. `ThumbStore::shards()` reports which are sealed, so an up-to-date client's whole pass
is one listing plus whichever shard is still open.
@@ -17,8 +17,8 @@ permission to a package rather than after.
| Platform | Channel | State | What it constrains |
|---|---|---|---|
| Linux | Arch source package — [`packaging/PKGBUILD`](../packaging/PKGBUILD) | Built, in tree | Nothing. Full filesystem access, system Vulkan, system secret daemon |
| Linux | Flatpak — [`packaging/flatpak/`](../packaging/flatpak/) | Manifest in tree, **library selection does not work** (§4) | Portals only. No `--filesystem=`, no host mount table, no typed paths |
| Linux | Arch source package — [`packaging/PKGBUILD`](../../packaging/PKGBUILD) | Built, in tree | Nothing. Full filesystem access, system Vulkan, system secret daemon |
| Linux | Flatpak — [`packaging/flatpak/`](../../packaging/flatpak/) | Manifest in tree, **library selection does not work** (§4) | Portals only. No `--filesystem=`, no host mount table, no typed paths |
| Linux | AppImage | v1 channel, **recipe not yet written** (§5) | Oldest supported glibc, and no sandbox at all |
| Android | F-Droid | v1 channel, not yet submitted | GPLv3-clean build, reproducible, no proprietary blobs |
| Android | Play Store | **Not v1** (§6) | Would make ARCH §6.9 binding as policy rather than as engineering |
@@ -39,7 +39,7 @@ somewhere:
that misses one of them costs the icon in the shell or the association in the
software centre, and neither failure announces itself.
- **The metainfo, not just the desktop entry.**
[`packaging/paris.tourolle.darkroom.metainfo.xml`](../packaging/paris.tourolle.darkroom.metainfo.xml)
[`packaging/paris.tourolle.darkroom.metainfo.xml`](../../packaging/paris.tourolle.darkroom.metainfo.xml)
is the single description of the application, installed by every channel that
has somewhere to put it. Its `metadata_license` is CC0-1.0 and its
`project_license` is GPL-3.0-or-later; those differ on purpose — see the
@@ -173,7 +173,7 @@ Two changes, in this order:
chooser instead of a volume list.
**Done when:** a Flatpak built from
[`packaging/flatpak/paris.tourolle.darkroom.yml`](../packaging/flatpak/paris.tourolle.darkroom.yml),
[`packaging/flatpak/paris.tourolle.darkroom.yml`](../../packaging/flatpak/paris.tourolle.darkroom.yml),
with its `finish-args` unchanged and no `flatpak override` applied, can select a
library root, scan it, and write a sidecar back into it.
+10
View File
@@ -821,6 +821,16 @@ here, and it is the phone and tablet story that should decide whether it gets bu
already has, the raw vector written over the old one and its id, box and identity untouched
(`faces::record_updates`). No detector runs and no suggestion is lost — the cost is the
original fetched once more, since the length exists only at the moment of embedding.
- **A person is stood for by their references.** Every face the user has ruled on is an anchor,
and the scan is exhaustive, so a person with 750 confirmations would cost 750 comparisons
against every other face — and the cost of a library would grow with how well it was named.
Instead each person enters through at most `MAX_REFERENCES` (100) of their anchored faces,
chosen by `dr_face::references`: those whose raw embedding is at least `MIN_REFERENCE_QUALITY`
(15) long, and among them the set spanning the greatest volume — greedy max-determinant, the
longest vector first and then, at each step, the face with the largest component orthogonal to
the chosen so far. Thirty frames from one afternoon contribute one reference; the single profile
shot is taken early. The faces not chosen keep their confirmations and are not touched by the
pass; they are simply not compared.
**The algorithm.** Constrained average-link agglomeration over the probability graph, merging while
the average pairwise probability exceeds **0.9** and no cannot-link is violated. Average-link rather
@@ -3,8 +3,8 @@
**Status:** Measured · 2026-08-27
**Companion to:** [display-and-extension.md](display-and-extension.md) §2–3 ·
[requirements.md](requirements.md) §3.4 FR-DSP-2, FR-DSP-3, FR-DSP-4
**Instrument:** [`core/dr-gpu/examples/frame_budget.rs`](../core/dr-gpu/examples/frame_budget.rs)
**Guard:** [`core/dr-gpu/tests/frame_budget.rs`](../core/dr-gpu/tests/frame_budget.rs)
**Instrument:** [`core/dr-gpu/examples/frame_budget.rs`](../../core/dr-gpu/examples/frame_budget.rs)
**Guard:** [`core/dr-gpu/tests/frame_budget.rs`](../../core/dr-gpu/tests/frame_budget.rs)
[display-and-extension.md](display-and-extension.md) §2 fixed a decision rule in
advance and made three measurements the thing that settles it. This file is
+75 -15
View File
@@ -75,7 +75,49 @@ TensorRT's job.
worst: nearly five minutes), because it is compiling an engine for this exact GPU. The engine caches to disk and the second load is milliseconds. That
number is what §6 is designed around.
### 1.3 What the numbers say
### 1.3 The desktop — Radeon RX 7900 XT, Threadripper 2920X, 24 threads · 2026-09-20
Arch's `onnxruntime-rocm` 1.29.0 against ROCm 7.2.4 and MIGraphX 7.2.3 (gfx1100), through the
same `ep_probe` harness (`core/dr-inference-engine/examples/ep_probe.rs`). Zero input, three
warm-ups, the median of 15 runs, on a machine doing nothing else. The build columns are the wall
clock of `Session` construction: cold is a MIGraphX compile of the graph for this GPU, cached is the
same session loading the program the cold build wrote.
| Model | ORT CPU f32 | MIGraphX f32 | **MIGraphX fp16** | Compile f32 / fp16 (s) | Cached load (s) |
|---|---|---|---|---|---|
| scrfd_500m (Fast) | 10.4 | 2.8 | **2.4** | 40 / 58 | 0.3 |
| scrfd_2.5g (Balanced) | 20.7 | 3.3 | **2.8** | 37 / 40 | 0.3 |
| scrfd_10g (Thorough) | 57.9 | 4.5 | **3.4** | 40 / 48 | 0.4 |
| arcface_mbf (per face) | 12.8 | 1.8 | 1.6 | 15 / 21 | 0.4 |
| yolo26n-seg | 49.3 | 8.4 | **7.5** | 110 / 136 | 0.9 |
| yolo26s-sem-ade20k | 55.8 | 4.8 | **3.8** | 50 / 60 | 0.5 |
| 2d106det (landmarks) | 9.9 | 1.2 | 1.0 | 16 / 21 | 0.2 |
| ocec_s (eye state) | 2.5 | 0.5 | 0.4 | 17 / 17 | 0.1 |
| xfeat-1024 | 26.8 | 10.5 | 9.9 | 37 / 53 | 0.3 |
| migan-512 (per tile) | 514 | 12.7 | **8.3** | 102 / 132 | 0.8 |
Three things the table settles.
- **The ROCm execution provider does not exist any more.** It was ONNX Runtime's CUDA-provider twin
for AMD, removed in the 1.23 release (AMD's builds dropped it from ROCm 7.1); 1.29's ROCm build ships
`libonnxruntime_providers_migraphx.so` and nothing else for AMD, and asking for `ROCm` answers
"not enabled in this build". So there is no non-compiling AMD rung to sit under MIGraphX the way
the CUDA provider sits under TensorRT: the AMD ladder is MIGraphX, then the CPU.
- **MIGraphX is a compiling provider, and its cache has to be asked for by name.** 15–135 s per
graph cold, under a second from its cache — TensorRT's shape exactly, and §6's design covers it.
Two things the provider does that the code has to know: `ort`'s builder fills the legacy options
struct, which 1.29 reads for the precision flags only, so the cache directory
(`migraphx_model_cache_dir`) reaches it only through the generic key/value registration; and
the cache key is the graph, the GPU and the MIGraphX version *without the precision*, so an fp16
session pointed at the f32 program's directory silently loads the f32 program (the first fp16
row measured here was that, before the directories were split).
- **fp16 is worth 10–35% over f32 on this card, not the 3× it is worth on TensorRT**, because
MIGraphX f32 is already 3–8× the CPU provider and the small graphs are launch-bound. The
detectors at 2.4–3.4 ms sit beside TensorRT fp16's 1.8–3.3 ms on the RTX 3050; the embedder is
1.6–1.8 ms on either precision and stays f32 (§7). The whole face pipeline for one image
(detector + landmarks + eyes + embedder) is under 6 ms.
### 1.4 What the numbers say
- **`tract` is single-threaded.** The tablet's one X4 core and one Raptor Lake core give the same
tract numbers. Replacing it with ONNX Runtime's CPU provider, *no accelerator involved*, is 3–6×
@@ -92,6 +134,8 @@ number is what §6 is designed around.
- **On NVIDIA, TensorRT fp16 ≈ 3× the CUDA provider**, and the CUDA provider ≈ 2× the
multi-threaded CPU; at fp16 the detectors are 1.8–3.3 ms with no quantisation at all. Both leave the twenty cores free for decoding during a batch index, which the table does not
show and which matters more than the ratio.
- **On AMD, MIGraphX fp16 is 4–17× the CPU provider** on the detectors and 60× on the
inpainter, with the same first-run compile cost as TensorRT and no rung between it and the CPU.
---
@@ -105,7 +149,8 @@ winning:
| Android, Qualcomm with a Hexagon the shipped QNN skel covers (V68–V81) | QNN HTP, int8 model | ORT CPU, f32 model | — | tract |
| Android, any other SoC | ORT CPU, f32 | — | — | tract |
| Linux / Windows, NVIDIA GPU | TensorRT, f32 model, fp16 engine | CUDA provider, f32 | ORT CPU, f32 | tract |
| Linux / Windows, no NVIDIA | ORT CPU, f32 | — | — | tract |
| Linux, AMD GPU with ROCm | MIGraphX, f32 model, fp16 program | ORT CPU, f32 | — | tract |
| Linux / Windows, no GPU stack | ORT CPU, f32 | — | — | tract |
| macOS ⁵ | ORT CPU, f32 | — | — | tract |
⁵ CoreML is the obvious rung and is unmeasured; it is listed so its absence is a gap and not an
@@ -113,8 +158,13 @@ oversight.
Deliberately **not** on any ladder, with the measurement that excluded each: NNAPI (no driver),
XNNPACK (slower than CPU, aborts on SCRFD), WebGPU (slower than CPU), the Adreno through QNN (works,
but never where the Hexagon does not also), CUDA int8 (slower than CUDA f32). A rung is added to this
table by a measurement on this page, not by a provider existing.
but never where the Hexagon does not also), CUDA int8 (slower than CUDA f32), the ROCm provider
(gone: §1.3). A rung is added to this table by a measurement on this page, not by a provider
existing.
The AMD ladder has no middle rung. TensorRT falls back to the CUDA provider while its engines
compile; MIGraphX has no such twin, so its fallback is the CPU provider, and the minute or two of
compiling on first run (§6) is spent at the floor's speed rather than at half the GPU's.
Two things the ladder is *not*: it is not a per-model choice — one backend serves every model on a
device, because §7's identity rule needs the detector and embedder on the same runtime for the
@@ -174,9 +224,10 @@ is cheaper than discovering it at packaging time.
| ONNX Runtime | MIT | Yes |
| Qualcomm QNN runtime (`com.qualcomm.qti:qnn-runtime` on Maven) | Qualcomm AI Engine Direct SDK licence — proprietary, redistribution permitted for applications using it | Yes for the APK, with the licence text shipped; not for a source distribution. **To be read in full, not summarised from memory, before the APK gains it.** |
| CUDA runtime, cuDNN, TensorRT | NVIDIA EULAs — redistributable with an application, with the licence text, not modifiable | Yes for a package that bundles them. 600 MB. The alternative is to load them from the user's system install if present and skip the rung otherwise — which is what §4's probe does anyway. |
| ROCm (HIP, MIOpen, rocBLAS), MIGraphX | MIT | Yes, but the HIP SDK MIGraphX needs is ~15 GB installed. Same answer as NVIDIA: the user's system install, or the rung is skipped. |
The position this takes: the **NVIDIA libraries are not bundled**. The desktop package probes for a
system CUDA/TensorRT install and uses it if it is version-compatible; a desktop without one runs on
The position this takes: the **GPU vendors' libraries are not bundled**. The desktop package probes for a
system CUDA/TensorRT or ROCm/MIGraphX install and uses it if it is version-compatible; a desktop without one runs on
ORT CPU, which is still 8–10× today. Bundling 600 MB for a rung that is 2× again is not a trade
worth making unmeasured, and it can be revisited by a measurement on a batch index. The **QNN
runtime is bundled** in the APK, because the Hexagon is the difference between a tablet that
@@ -237,6 +288,7 @@ of which form they load:
| f32 ONNX, shape-fixed, **opset ≥ 13** | `tools/fix-face-model-shapes.sh`, `tools/export-seg-model.sh` | Release time, once | Every rung except Hexagon |
| int8 QDQ ONNX, per-channel, uint8 activations | `tools/quantise-models.sh` (new) | Release time, once, **calibrated on real photographs** | Hexagon |
| TensorRT engine (`.engine`, per GPU architecture and TensorRT version) | The app, from the f32 file | First run on that device, in the background | TensorRT rung |
| MIGraphX program (`.mxr`, per GPU architecture, MIGraphX version and precision) | The app, from the f32 file | First run on that device, in the background | MIGraphX rung |
| QNN context binary | The app, from the int8 file | First run on that device, in the background | Hexagon rung |
Two rules.
@@ -256,9 +308,12 @@ prefers. That is a change to the canonical file and so a change to the shipped m
happens in the same model release as the int8 files.
**Compilation is a device-time step, and it is cached.** A TensorRT engine is specific to the GPU
it was built on and the TensorRT that built it; a QNN context binary is specific to the Hexagon
generation. Neither can ship. Both are built by the app the first time that rung is selected, in
the background (§6), and written beside the probe cache keyed by the same inputs. They are
it was built on and the TensorRT that built it; a MIGraphX program to the GPU and the MIGraphX
that built it; a QNN context binary to the Hexagon generation. None can ship. All are built by
the app the first time that rung is selected, in the background (§6), and written beside the
probe cache keyed by the same inputs. MIGraphX's own key leaves out the precision, so the app
gives its f32 and fp16 programs separate directories — otherwise the embedder's f32 build would
be served the detector's fp16 program, or the reverse. They are
**derived, disposable, and regenerable**: deleting the cache directory costs the next launch a
rebuild and nothing else, and the directory is excluded from anything that syncs (it is a peer of
`thumbs`, not of the catalog).
@@ -267,17 +322,18 @@ rebuild and nothing else, and the directory is excluded from anything that syncs
## 6. First run — building engines without the user waiting for them
The sequence on a device where a compiling rung (TensorRT, Hexagon) is selected:
The sequence on a device where a compiling rung (TensorRT, MIGraphX, Hexagon) is selected:
1. **Launch.** The runtime loads; the probe (§4) starts in the background; the app serves every
model request from the floor. Face indexing, segmentation and scene grading all work, at
today's speed or better (ORT CPU).
2. **Probe reports** — say, TensorRT. The compiling rung is now *selected* but has **no engines**.
Model requests continue on the fallback rung below it (CUDA provider for TensorRT; ORT CPU for
Hexagon), which needs no compilation and is already faster than the floor.
MIGraphX and Hexagon), which needs no compilation and is already faster than the floor.
3. **Engines build**, one model at a time, on a single low-priority background thread, smallest
model first so the detector — the one that runs per image — is ready soonest. On the reference
desktop that is ~1 minute for the first detector and ~10 minutes for all six at fp16; on the tablet the QNN
desktop that is ~1 minute for the first detector and ~10 minutes for all six at fp16; on the AMD
desktop 40 s for the first detector and ~8 minutes for the set; on the tablet the QNN
context binaries take 0.8–1.7 s each and the whole set is ready before the user has opened a
library. Each engine is written to a temporary name and renamed into place, so a request never
sees a half-written file.
@@ -313,7 +369,7 @@ enough to be *offered* at all. f32 on tract, ORT CPU, CUDA and TensorRT-f32 are
same graph, the same arithmetic, differences at the last bit.
**The embedder** is where comparability across devices is the whole point, and it is the one
model that no accelerator helps (§1.3). So: **the embedder runs in f32 on every rung.** On TensorRT
model that no accelerator helps (§1.4). So: **the embedder runs in f32 on every rung.** On TensorRT
that means the embedder's engine is built without fp16 while the detector's is built with it; on
the Hexagon it means the embedder is not on the NPU at all — it runs on the ORT CPU rung at 9 ms,
and the ladder's "one backend per device" is, precisely, one backend *per model role*, with the
@@ -344,7 +400,7 @@ core/dr-inference-engine
src/probe.rs §4 — the ladder per platform, the session-build probe, the cache file
src/engines.rs §6 — background compilation, the cache directory, progress
src/session.rs open(role, bytes) -> ort::Session, applying the rung and the role's precision rule
src/api.rs the one unsafe block: dlopen libonnxruntime, fetch OrtApi, ort::set_api — or ort_tract::api()
src/api.rs the unsafe block that matters: dlopen libonnxruntime, fetch OrtApi, ort::set_api — or ort_tract::api()
```
- `dr-face` and `dr-segment` **delete** their private `install_backend` and their direct
@@ -354,7 +410,11 @@ core/dr-inference-engine
`qnn`: those features add option builders, not linking, under `alternative-backend`. **Verified
for the QNN, CUDA and TensorRT builders on 2026-09-19** — they go through the API table's generic
`SessionOptionsAppendExecutionProvider*`. The NNAPI builder resolves a symbol directly and would
not; it is not needed and is not enabled.
not; it is not needed and is not enabled. MIGraphX uses no `ort` feature at all: `ort`'s builder
fills the legacy options struct, which ONNX Runtime 1.29 reads for the precision flags and
nothing else, and the compiled-program cache directory only travels through the generic
key/value entry point (`migraphx_model_cache_dir`). `session::migraphx` makes that one call
on the API table itself.
- `dr-ui` owns the settings row, the about-screen line and the progress row; it holds one
`Sessions` per process, created at launch, and passes it down. `dr_ui::library` gains
`inference_cache_dir()` beside `shared_face_models_dir()`, on the same account-independent
@@ -18,11 +18,11 @@ The core is further along than the interface, and it is worth being exact
about which half is missing, because it changes the size of the work.
**Painting exists everywhere except where a finger is.**
[`MaskSource::Brush`](../core/dr-pipeline/src/mask.rs), [`Stroke`], the
[`MaskSource::Brush`](../../core/dr-pipeline/src/mask.rs), [`Stroke`], the
simplification and the point budgets, the sidecar's `stroke = …` line and its
parser, the GPU's per-stroke bounding-box draw with add and erase blend
states — all of it is written, tested, and reachable from no control in the
application. [`toolrail.slint:138`](../ui/dr-ui/ui/toolrail.slint#L138) says so
application. [`toolrail.slint:138`](../../ui/dr-ui/ui/toolrail.slint#L138) says so
in as many words: *"what is missing is the canvas interaction"*.
**A mask has exactly one source.** A layer is one `MaskSource` and a shaping
@@ -37,7 +37,7 @@ shoulder and leaks four pixels into the hair, no global number fixes both, and
that is the ordinary case rather than a corner one.
**And you cannot see the mask.** *(Built — see §6.)* The overlay on the canvas
was [`overlay_rgba`](../ui/dr-ui/src/segmentation.rs) and nothing else — a
was [`overlay_rgba`](../../ui/dr-ui/src/segmentation.rs) and nothing else — a
CPU-built false-colour picture of *what the model detected*, at proxy
resolution. It is not the layer's alpha: it knows nothing of the layer's
feather, its falloff, its morphology, its invert, or its opacity. Nobody can
@@ -234,7 +234,7 @@ writes with one part is byte-identical to what it writes today**:
3. `stroke = …` lines keep their position and their meaning. The mode token
gains `push`, and `cling` is a fourth number written only when non-zero.
An older build meeting either drops *that stroke and only that stroke*,
which is the rule [`parse_stroke`](../core/dr-pipeline/src/sidecar.rs)
which is the rule [`parse_stroke`](../../core/dr-pipeline/src/sidecar.rs)
already documents and already implements.
**Merge (FR-NC-9).** A part is a block with an id, so two devices that added
@@ -272,7 +272,7 @@ The set operations are already expressible in fixed-function blending over
| Intersect | `Zero`, `Src`, `Add` | `dst · src` | M2 |
The middle row is `brush_erase`, already constructed in
[`MaskPass::new`](../core/dr-gpu/src/mask.rs). The other two are the same
[`MaskPass::new`](../../core/dr-gpu/src/mask.rs). The other two are the same
three vertices with a different `BlendState`, and nothing is read back.
**One correction to the first draft of this section, found in the building.**
@@ -289,7 +289,7 @@ the first time any layer has more than one part) and blended from there. A
layer of one part still takes the old path exactly — straight into its slice,
no scratch, no combine pass — which is what keeps every existing mask
rendering as it did. `an_erase_stroke_holes_its_own_part_and_not_the_mask` in
[`local_adjustments.rs`](../core/dr-gpu/tests/local_adjustments.rs) is the test
[`local_adjustments.rs`](../../core/dr-gpu/tests/local_adjustments.rs) is the test
that holds this in place.
`Max` blending on `r8unorm` is core WGPU and universally supported on the
@@ -342,7 +342,7 @@ Three honest costs:
### 5.4 Painting has to be incremental
Today [`MaskPass::render`](../core/dr-gpu/src/mask.rs) clears each slice and
Today [`MaskPass::render`](../../core/dr-gpu/src/mask.rs) clears each slice and
redraws every stroke of the layer. That is right when a shape changes and
wrong while a finger is down: at 120 reports a second, a layer holding 4096
points redraws all of them per dab, and the cost of a stroke grows as it is
@@ -359,14 +359,14 @@ changing falls back to the full rebuild it does now.
This is required, not an optimisation to schedule later: it is what decides
whether painting is usable on the phone, and it is the specific failure
[`mask.rs`'s module docs](../core/dr-pipeline/src/mask.rs) say this whole
[`mask.rs`'s module docs](../../core/dr-pipeline/src/mask.rs) say this whole
design exists to avoid.
### 5.5 Distance fields become per part
[`SubjectMasks`](../core/dr-gpu/src/mask.rs) uploads one signed distance field
[`SubjectMasks`](../../core/dr-gpu/src/mask.rs) uploads one signed distance field
**per active layer, in stack order**, and
[`DevelopSession`](../ui/dr-ui/src/develop.rs) builds them on the same
[`DevelopSession`](../../ui/dr-ui/src/develop.rs) builds them on the same
indexing. With parts, a field belongs to the part that shaped it: the upload
becomes one field per *model-backed part*, flattened in `(layer, part)` order,
and the rasteriser indexes it by a running counter rather than by `slot`.
@@ -540,7 +540,7 @@ layer's.
or painted. One code path, three joins.
**Folding two layers into one** uses the multi-selection
[`masks_ui.rs`](../ui/dr-ui/src/masks_ui.rs) already supports: with two layers
[`masks_ui.rs`](../../ui/dr-ui/src/masks_ui.rs) already supports: with two layers
selected, "Combine" appends the second's parts to the first and removes it.
Offered only when the second layer's adjustments are neutral, and otherwise
offered with a warning that names what will be lost — quietly discarding an
@@ -560,7 +560,7 @@ who sets a small eraser expects it to still be small the next time they erase.
### 7.4 Gestures
Each of these needs a `GESTURE:` block beside its implementation — that is the
only place [gestures.md](gestures.md) can be written from.
only place [gestures.md](../gestures.md) can be written from.
| Gesture | Touch | Pointer | Keyboard |
|---------|-------|---------|----------|
@@ -611,7 +611,7 @@ removing or re-joining a part.
The mask stack is already snapshotted per step and shared by `Arc` when a step
does not touch it, so the cost of an undoable stroke is a clone of one layer's
parts, not of the picture. New keys in
[`labels.rs`](../ui/dr-ui/src/labels.rs):
[`labels.rs`](../../ui/dr-ui/src/labels.rs):
```
history.mask_painted "Paint Mask"
@@ -196,7 +196,7 @@ inside `#[cfg(test)]` — `LocalStorage::open` refuses it, and the test that pro
`dr_plat::imports_supported`, which *returns false on Android* and whose own documentation says it
"stops being false when a SAF implementation lands". The second tag documented the absence of the
thing it was counted as evidence for. Both have been removed; this is the "plumbing a future feature
would use" case [CONTRIBUTING.md](../CONTRIBUTING.md) and [code-health.md CH-4](code-health.md) both
would use" case [CONTRIBUTING.md](../../CONTRIBUTING.md) and [code-health.md CH-4](code-health.md) both
warn about. Android reaches a library through a Nextcloud account or a folder, over paths, like the
desktop.
@@ -355,10 +355,10 @@ synthetic 50k catalog, run per commit, where **"a regression beyond a stated tol
failure, not a notification."** For most of this project's life it did not exist — no `benches/`, no
criterion, no synthetic catalog, and three CI workflows that between them measured nothing.
**It exists now, for everything that does not need a frame.** [`tools/bench`](../tools/bench) builds
**It exists now, for everything that does not need a frame.** [`tools/bench`](../../tools/bench) builds
a deterministic 50,000-row catalog over a pool of a dozen real files, measures against it, and fails
the build on a violated budget or a drift past tolerance;
[`.gitea/workflows/benchmark.yml`](../.gitea/workflows/benchmark.yml) runs it on every push, and
[`.gitea/workflows/benchmark.yml`](../../.gitea/workflows/benchmark.yml) runs it on every push, and
[benchmarks.md](benchmarks.md) is the account of what it does and does not cover. **NFR-P1** and
**NFR-P3** are now genuinely gated, and R2's "catalog opens in under 2s" clause with them.
@@ -404,7 +404,7 @@ tag anything against.
NFR-OPS-1 was covered by tags that were real rather than fixtures, which is the worse case of the
two: one on `compute_coverage` and one on the gesture extractor, both on the traceability tool. A
coverage calculation and a documentation generator are not diagnostics under any reading, so both
tags were removed. It is the case [CONTRIBUTING.md](../CONTRIBUTING.md) warns about in its own words:
tags were removed. It is the case [CONTRIBUTING.md](../../CONTRIBUTING.md) warns about in its own words:
a tag proves a tag exists. The requirement has since been built where it says: the rotating,
size-capped log and its redaction in `platform/dr-plat/src/diagnostics.rs` (2026-08-30), and the
bundle in `diagnostics/bundle.rs` (2026-09-12) — the log, the crash records, the version, the schema
+23 -8
View File
@@ -179,7 +179,7 @@ exports the network alone at 768×1024 — thirteen operator types, all
standard: `Conv`, `InstanceNormalization`, `AveragePool`, `Resize`, `Slice`,
`Transpose`, `Reshape`, `Concat`, `Add`, `Relu`, `Sigmoid`, `ReduceMean`,
`Unsqueeze` — and
[`examples/onnx_probe.rs`](../core/dr-segment/examples/onnx_probe.rs) loads
[`examples/onnx_probe.rs`](../../core/dr-segment/examples/onnx_probe.rs) loads
the 2.8 MB file through the app's own `ort`-over-tract backend with nothing
unsupported, in 28 ms, and runs it in **~300 ms on the reference desktop's
CPU**. The weights ship as `models/keypoints/xfeat-1024.onnx`, recorded in
@@ -255,7 +255,7 @@ Two containers were candidates and S15.1 decided, on 2026-09-19:
the existing encoder with a different sample type, and nothing about it is
uncertain.
**Linear DNG.** [`examples/linear_dng.rs`](../core/dr-decode/examples/linear_dng.rs)
**Linear DNG.** [`examples/linear_dng.rs`](../../core/dr-decode/examples/linear_dng.rs)
hand-rolls a 64 × 48 `LinearRaw` DNG — one IFD, uncompressed 16-bit RGB,
`DNGVersion`, `ColorMatrix1`, `AsShotNeutral`, `CalibrationIlluminant1` — and
rawler 0.7 reads it back: `cpp 3`, the samples interleaved as written, the
@@ -538,9 +538,24 @@ each loss weighting did, and the two runs abandoned (blur under L1 in
the hole; a brick pattern under a strong adversarial term against a
discriminator that had not learned) — is `runs/` in `darkroom-infill`.
**What is still wrong.** The ground fill is softer than its context —
texture, not structure, is what a night on a laptop GPU could not finish.
The levers, in order: a discriminator that learns (a pretrained one —
MI-GAN's own from the unfused checkpoint — instead of a PatchGAN from
scratch), feature matching, and more steps at 512. FR-MRG-4's
*experimental* stays.
**Second model, the same day.** The morning's fill was soft in the deep
ground bands. The afternoon's run trained the generator against MI-GAN's
own pretrained discriminator (non-saturating loss, lazy R1, feature
matching), with fresh noise inputs while training and flip/translation
augmentation of the discriminator's input — both needed, or the generator
settles into a periodic texture the discriminator cannot see. The shipped
weights (step 4 750 of `runs/border-v6`) are level with the stock model on
LPIPS (edge 0.125 / corner 0.187 against 0.121 / 0.183) while keeping the
PSNR gain (edge 18.0 / corner 15.7 against 16.9 / 14.6). On the fixture
the ground bands now carry texture at the right tone; at 1:1 a faint
regular hatch is visible in the deepest part.
**What is still wrong.** The hatch, and any deep textured void the
generator must invent. The better answer for those is not generative:
seed the void with the picture's own texture in hexagonal cells, let the
discriminator rank the candidates, and let the generator heal only the
gaps between cells — built and measured in `darkroom-infill`
(`infill/hexfill.py`), the most convincing scree corner produced so far,
and the next thing to port into `dr_pano::fill` (it needs the
discriminator as a second model, ~80 MB fp16). FR-MRG-4's *experimental*
stays.
@@ -1814,6 +1814,28 @@ tool for a hand-held set, and a photograph with real parallax is not a panorama.
in one operation until HDR merge exists on its own. No live re-stitch: a different projection or
crop after the fact is a new file, not an edit. No video.
### 3.12 Inference runtime
The face, segmentation and border-fill models run under an inference engine the application
selects per device. [inference.md](inference.md) is the design: its §3 sets the dependency policy
the runtime may reopen, its §10 the milestones (M1–M7) the clauses below cite as acceptance. The
clauses are the register entries inference.md §12 promised; they are stated here so the
traceability gate can count them.
**FR-INF-1 — Runtime selection.** On launch the application shall determine, per device and
without blocking the first frame, the fastest inference backend that can build and run a session
for the shipped models, by attempting it; shall record and reuse that determination until the
runtime, driver, hardware or models change; and shall display the backend in use in Settings and
on the about screen. *Acceptance:* inference.md §10 M1 and M5.
**FR-INF-2 — Derived engines.** Backends that require device-specific compilation shall compile in
the background after selection, shall serve requests from the next lower backend until each engine
is ready, and shall not change the backend of a job in progress. *Acceptance:* M5.
**FR-INF-3 — Model forms.** Quantised model forms are produced at release time from real
calibration data and are shipped only when they meet inference.md §10's accuracy gates against the
canonical form; the application never quantises on the device. *Acceptance:* M2, M7.
---
## 4. Non-functional requirements
@@ -2111,6 +2133,13 @@ it is far cheaper to discover now than after the UI is built.
clipping indicators (FR-DSP-7), and the HSL mixer carry a shape or text affordance. This matters
more in a colour-grading application than in most software.
### 4.10 Inference
**NFR-INF-1 — Embedding comparability.** Face embeddings shall be computed at a precision whose
deviation from the f32 reference is within [inference.md](inference.md) §7's gate, on every
backend, so that embeddings from any device are comparable — FR-CULL-9's calibration depends on
it. *Acceptance:* inference.md §10 M3.
---
## 5. Data model and architecture
@@ -2140,7 +2169,7 @@ Rationale, evidence, and the eliminated alternatives are recorded in
|---|---|---|
| D1 | Language and UI framework | Rust + Slint, rendering through wgpu |
| D2 | RAW decoder | rawler; LibRaw fallback behind a trait |
| D3 | First milestone | Delivered — [milestone-v0.1.md](milestone-v0.1.md), closed 2026-08-30 |
| D3 | First milestone | Delivered — [milestone-v0.1.md](archive/milestone-v0.1.md), closed 2026-08-30 |
| D4 | Nextcloud sync mechanism | ETag pruning, chunked upload v2, Login Flow v2 |
| D5 | Colour management | lcms2 + GPU-side matrix/LUT transforms |
| D6 | Shader authoring | Hand-written WGSL |
@@ -21,11 +21,11 @@ one in its class and the workflow still breaks at the first frame with a mark on
the sky.
It is also, unusually, a feature whose cost has already been paid twice over.
The neighbourhood stage exists ([`crate::detail`](../core/dr-pipeline/src/detail.rs)),
The neighbourhood stage exists ([`crate::detail`](../../core/dr-pipeline/src/detail.rs)),
the convention for storing geometry in normalised source coordinates exists
([`mask.rs`](../core/dr-pipeline/src/mask.rs)), the canvas-drag pattern exists
([`gradient.rs`](../ui/dr-ui/src/gradient.rs)), and the merge-by-id rule exists
([`sidecar.rs`](../core/dr-pipeline/src/sidecar.rs)). What is genuinely new is
([`mask.rs`](../../core/dr-pipeline/src/mask.rs)), the canvas-drag pattern exists
([`gradient.rs`](../../ui/dr-ui/src/gradient.rs)), and the merge-by-id rule exists
([`sidecar.rs`](../../core/dr-pipeline/src/sidecar.rs)). What is genuinely new is
small and is named in §3.
## 2. Non-goals

Some files were not shown because too many files have changed in this diff Show More