Compare commits

..
235 Commits
Author SHA1 Message Date
dtourolle cf84cec96f Release 0.15.0
Benchmarks / CPU and I/O (per commit) (push) Successful in 5m19s
Benchmarks / Frame budget (on demand) (push) Skipped
Traceability / Requirement traces (push) Successful in 49s
Build and test / Android (aarch64) (push) Successful in 31m25s
Build and test / android-image (push) Successful in 1s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / Desktop (Linux) (push) Successful in 48m26s
Build and test / windows-image (push) Successful in 5s
🐳 Windows image / Build and push (push) Successful in 4s
Build and test / Layer separation (push) Successful in 34s
Build and test / Windows (x86_64, cross) (push) Successful in 19m57s
Build and test / Publish the release (push) Successful in 1m6s
2026-09-25 07:51:31 -04:00
dtourolle f0e7b8e11c Re-record the manual on the keyboard layout, and stop the zoom caption promising blocks
Every scene was recorded again on a build of this branch rebased onto the
keyboard work and TD-1, since nearly every scene depends on files those
changed: the develop top bar now carries a star strip and Pick/Reject, the
roll shows flags and stars, and the grid's selection bar gains Label and
Flag. `--changed` could not be trusted to find them, because the rebase
made each picture's commit newer than the sources it was recorded from.

The develop-zoom caption said the wheel goes on "until the pixels are
blocks". On Xvfb the deepest frames come out smooth even though the app
draws past 1:1 nearest-neighbour on a real display (confirmed by eye on
the desktop), and a GIF shrunk to 960 wide could not show 3-pixel blocks
anyway. The caption now says what the clip shows; the prose above it,
which describes what the app does, stays.
2026-09-25 07:26:37 -04:00
dtourolle 90c7cb65d3 Regenerate the manual page after the rebase onto the keyboard work
The rebase combined this branch's README additions with master's, and
index.html is generated from the README; manual-check passes on the
regenerated page.
2026-09-25 07:26:37 -04:00
dtourolle 70583b9b2f Fail CI when the manual shows a picture no scene makes
The traceability job now runs `tools/manual/record.sh --check`: every
picture docs/manual/README.md shows must be made by a scene in
tools/manual/scenes.py, and every picture a scene makes must be shown.
It reads the two files and nothing else, so it needs no app, display or
LFS pull.

--changed now dates a scene by the newest commit among its pictures
rather than each picture alone. A scene that also makes a picture which
re-records byte for byte (panorama-aligned beside panorama.gif) no
longer stays listed for ever. A scene all of whose pictures come out
identical (launch) stays listed until one differs, which costs one
harmless re-run.
2026-09-25 07:26:37 -04:00
dtourolle f426bb903a Picture this round's features in the manual
Four scenes, recorded and looked at frame by frame:
- library-labels: 6, 7, 8 and 9 over four New York frames, 7 again to
  take one off, then the Green chip narrowing the grid and back.
- compose-perspective: two towers shot from below, stood upright with
  Vertical at about +58, then held against Before.
- crop-orphan: a stroke in the top-left corner, a crop that leaves it
  outside, the notice with Undo crop and Keep crop, and Undo crop.
- local-intersect: a linear gradient over the lower half, then
  Intersect and two strokes that survive only where it is.

develop-zoom already ends on hard-edged pixels (previous commit). The
text added is a sentence or two under the existing headings, so the
other branch's structure and anchors are left alone.
2026-09-25 07:26:36 -04:00
dtourolle c41f99ea52 Record the manual's scenes by name, and each against what it depends on
scenes.py aimed every press at window pixels, and the develop column had
already moved under it: Compose now sits above Adjust, so the old
exposure coordinate lands on a straighten slider. Every scene now names
what it presses by its accessible label through the automation hook,
places points on the photograph relative to the canvas, and opens its
photographs by file name. Each starts from a known place and undoes what
it did, so one can be recorded alone; the few that continue another's
state name it, and running one runs that first into a scratch folder.

Each scene also declares the pictures it makes and the sources they
depend on. `record.sh --check` fails when the manual shows a picture no
scene makes, or a scene makes one it does not show; it reads two files.
`record.sh --changed` re-records the scenes whose sources, or own code,
changed since the commit that last touched their pictures. record.sh
builds with the automation feature, restores the library from
DR_LIBRARY_SNAPSHOT, starts from a fresh profile and pins inference to
the CPU; the launch screen is recorded from an empty profile of its own.

Re-recorded with the ported scenes, and looked at frame by frame. What
differs from the pictures they replace:
- develop, presets, settings, local, compose, film, wb, light: the
  current develop column (Compose with Vertical and Horizontal above
  Adjust, the Label button), otherwise the same moments.
- library pictures: the filter bar's colour-label chips; no collection
  left over from an earlier run in the sidebar; library-selection is the
  twelve alpine frames rather than eight of them and four New York ones.
- library-rating rates two frames nobody had rated, so the stars are set
  and not cleared.
- develop-zoom goes on past 1:1 with the wheel and ends on the file's
  pixels as hard-edged blocks.
- repair covers a real mark on the road, with a size that fits it; film
  is shown on the Chinatown frame instead of the road.
- panorama tries Perspective, Spherical and Cylindrical before filling.
- launch, launch-folder and panorama-aligned came out byte-identical.
2026-09-25 07:26:36 -04:00
dtourolle a6ea6ba83f Let the manual's scripts find a control by its name
Every scene in tools/manual aimed at window pixels written in by hand, so
a panel that gained a row moved every slider under it and the recording
went on dragging where the slider used to be. The develop column has
already moved that way (Compose now sits above Adjust), and nothing said.

A build with the `automation` feature listens on the Unix socket named
by DR_AUTOMATION and answers where an element is: by its accessible
label, the name a screen reader reads, or by its markup id for the few
things that are not controls (the canvas, the crop rectangle). It uses
Slint's element queries, which need the compiler's debug tables, so the
feature also turns those on in build.rs. It only answers questions; the
input is still xdotool's real pointer. No default build has the feature,
and one that has it listens only when the variable is set.

drive.py gains click-on, drag-on, hold-on, wait-for, wait-gone, labels
and ids. The grid's cells are now named by their file, each rating star
by its value, the sidebar's + as "New collection", and the Adjust
heading's reset as "Reset all adjustments" - controls a screen reader
could not reach before either.
2026-09-25 07:26:36 -04:00
dtourolle b480de5bff Record TD-1 as paid off, checked by eye on the tablet
The pre-rotation patches landed in the previous four commits. This
records how the debt was paid, by a fourth route its own list missed:
patching wgpu-hal and Slint's Skia surface locally rather than waiting
for either upstream. It also updates architecture.md §1 and §6.1, which
still said Android draws with OpenGL behind a readback.

The verification is stated as what it was: the user found the
release-signed build clean on the tablet in portrait. No dumpsys
composition or bufferTransform readings were taken, because adb would
not hold the device that morning, and no frame times were measured.
2026-09-25 07:26:00 -04:00
dtourolle a8043e6827 Hand Android's develop frame to the compositor as a texture again
With Skia drawing pre-rotated on wgpu's Vulkan swapchain, Android no
longer needs to draw with Skia over OpenGL, which was the only reason
the develop view read its frame back through memory (TD-1).

So `unstable-wgpu-29` moves back to the common slint dependency. The
android-activity backend then builds `SkiaRenderer::default_wgpu_29`,
and `shared_gpu` loses its Android arm. The one wgpu device is handed to
Slint through `BackendSelector::require_wgpu_29` on both platforms.
`slint::android::init_with_event_listener` runs before `dr_ui::run`, so
the selector reaches the Android adapter before its window exists.
`renderer-femtovg-wgpu` stays desktop-only, since Android has no FemtoVG.

The two `#[cfg(target_os = "android")]` readbacks in `develop::render`
(the frame through `export_pixels` and the focus overlay through
`read_overlay`) are gone. `read_overlay` stays for the tests that check
what the overlay marks.

Built for arm64 and release-signed. Not yet run on the tablet.
2026-09-25 04:20:19 -04:00
dtourolle 4b4c9e2e6d Pre-rotate Slint's Skia drawing on the wgpu swapchain on Android
The other half of the wgpu-hal patch: that one lets a caller promise a
pre-rotated swapchain, and this is the caller keeping the promise.

On configure, `WGPUSurface` reads the surface's `currentTransform`,
sizes the swapchain in the panel's orientation (swapped for a quarter
turn), tells wgpu-hal to use that transform, and before each frame
concatenates the matching rotation onto the Skia canvas. Everything
Slint draws goes through that one matrix, so an imported wgpu texture is
rotated with the rest of the window. Input is not rotated, and must not
be, because Android delivers it in window coordinates.

Three details that would each have been a visible bug:

- `resize_event` compared the new size against the swapchain's. The
  swapchain is transposed while a quarter turn is in effect, so the
  comparison now uses the window's size, kept beside it.
- A half turn, landscape to reverse landscape, changes the transform
  without resizing the window, and wgpu-hal hides the SUBOPTIMAL that
  would report it. So the transform is re-read before every frame. That
  costs one query into the native window.
- The item renderer snapped the origin to the pixel grid only when the
  canvas matrix was a pure translation. Under a rotation that is never
  true, so portrait would have lost pixel alignment everywhere. The check
  now accepts right-angle rotations and flips without scaling.

The direction of each rotation follows the Vulkan spec's reading of
preTransform (the image is drawn already rotated clockwise by the
transform). It has not been confirmed on the device yet.
2026-09-25 04:19:05 -04:00
dtourolle dc9da52651 Let a wgpu-hal caller choose the Vulkan swapchain's preTransform
wgpu-hal creates every swapchain with `preTransform = IDENTITY` (#3345).
On a tablet whose panel is mounted landscape, a portrait window then
hands Android an unrotated buffer: SurfaceFlinger falls back to rotating
it on the GPU (composition CLIENT), and on this device those frames tear.
That is why Android draws with Skia over OpenGL today, and why the
develop view pays a readback (TD-1).

The field cannot just be set to `currentTransform` inside wgpu. It is a
promise that the image is already drawn rotated and sized in the panel's
orientation, and only the renderer above wgpu can keep it. So the patch
is the smallest thing that lets that renderer ask:
`vulkan::Surface::current_transform` reads the surface's transform, and
`set_pre_transform` makes the next swapchain use it. The default stays
IDENTITY, so desktop and any caller that does not opt in behave exactly
as upstream.
2026-09-25 04:18:39 -04:00
dtourolle dc1add9dbb Vendor wgpu-hal 29.0.4 and i-slint-renderer-skia 1.17.1, unmodified
The Android develop view reads its frame back through memory (TD-1)
because wgpu's Vulkan swapchain never pre-rotates, and a portrait window
on this tablet's landscape panel then tears. The fix is a small patch to
each of these two crates, and this commit is only the ground it lands on:
both are byte-for-byte the crates.io sources the lockfile already
resolved, so the commits that follow are the patch and nothing else.

third_party/ is excluded from the workspace, or every path dependency
under the root would become a member and `--workspace` would test and
lint upstream code as ours. The README says how to carry the patches
across a Slint or wgpu bump, which matters because a stale version here
does not fail the build — cargo just warns and uses the unpatched crate.
2026-09-25 04:18:39 -04:00
dtourolle b5ae5c2be1 Record FR-DEV-16 as met and where FR-UI-5 stands
FR-DEV-16's promise that the gesture book cannot describe a binding the
application lacks is now enforced by the gestures gate in both directions,
and every binding it names is bound and tagged, so it is marked met, with
what develop answers beyond the list.

FR-UI-5 is not met in full: its keyboard half and the 2026-09-19 amendment
are, but no develop slider takes the scroll wheel, so its status says that
rather than rounding up.
2026-09-24 23:42:28 -04:00
dtourolle aa2a88f655 Show each frame's flag and stars on the develop roll
FR-UI-5's 2026-09-19 amendment asks for the rating and flag wherever they
can be set and on the roll's cells, so that stepping along a set in develop
shows what has been judged. The roll drew thumbnails only, so the keys that
now judge the open photograph left no trace on its neighbours.

Each roll cell carries a small badge with a tick or a cross and a star
count, drawn only when there is something to show, in shapes and a number
rather than colours (NFR-A11Y-3).
2026-09-24 23:42:28 -04:00
dtourolle f8737e3fda Close the help sheet with Escape, and keep keys from acting behind it
F1 opens the help sheet from the grid, and nothing on the keyboard closed
it: Escape fell through to the shell, and every other key went on judging,
labelling and keywording the photographs hidden behind the sheet.

Escape and Back now close the sheet first, ahead of the grid's other
sheets, since it is drawn over all of them. While it is open the grid's
handler declines every other key, so a stray P or Ctrl+K changes nothing
the reader cannot see.
2026-09-24 23:42:27 -04:00
dtourolle d489a34190 Drive develop, the grid, the sidebar and People from the keyboard
An audit of every action by view against the keys the handlers bind left
develop without zoom, pan, fit or a way back to the grid, the grid without
select-none, thumbnail size or keywording, People with no key at all, and
the export and copy sheets without Enter. It also found the reverse gap
FR-UI-5 forbids: pick and reject had no route but P, X and U, and the
2026-09-19 amendment's judging in develop had not been built.

Develop: Ctrl+= and Ctrl+Plus zoom in and Ctrl+- out about the middle of the
view, Ctrl+0 fits and Ctrl+1 goes to 1:1, Shift and an arrow pan a magnified
view, G goes back to the grid, Ctrl+Y redoes, and Enter keeps a crop that hid
a mask. 0-5, P, X and U rate and flag the open photograph without moving on,
with stars and Pick/Reject in the top bar as the pointer and touch route.
= and - nudge the control last moved by a hundredth of its travel; the
framing sliders, perspective included, now count as "last moved", so R puts
them back as well. J turns the selected mask part's join chip.

Grid: Ctrl+D and Ctrl+Shift+A clear the selection, = and - resize the
thumbnails, Ctrl+K opens keywording, and Flag in the selection bar gives
pick and reject a pointer and touch route. Sidebar: Enter commits a
collection's name, and Enter or Escape hands the keyboard back to the grid,
where it used to go nowhere until something was clicked. People: Up and Down
walk the rail, F2 puts the name field under the keys, and Escape or Back now
leave the screen the way its back button does instead of doing nothing.
Sheets: Enter does what the export or copy sheet's button does.

The choices follow Lightroom where it has one. No new key steals typing: the
grid's and People's keys live on focus holders that are not ancestors of any
text field, and the sheets' Enter comes after a focused field has had it.
Every binding is tagged beside its handler, and the gate added in the
previous commit holds the two to each other.
2026-09-24 23:42:26 -04:00
dtourolle 9d1e31ffbb Fail CI when a key is bound but not in the gesture book, or listed but not bound
The gesture book is generated from GESTURE tags, so it could not describe a
gesture nobody tagged, but nothing made anyone tag one. The arrow keys, Enter,
P, X, U, Delete, F1 and F2 all worked in the grid with no line in the help
sheet, and a tag could name a key whose handler had gone.

Key handlers now compare one canonical string, Keys.chord(event) == "Ctrl+Z",
instead of reading event.text and the modifiers themselves. keys.slint folds
the key and its modifiers into that spelling, so the literal in the handler is
the whole binding and the checker reads exactly what the handler dispatches
on. Each handler carries a KEYMAP comment naming the gesture-book section its
keys belong to, and a tag's keys field names its keys between backticks.
gestures-check now fails when a handler binds a key no tag in that section
names, when a tag names a key no handler there binds, when any .slint file
other than keys.slint reads event.text, when a compared literal is not
canonical, and when keys.slint's named keys drift from the Rust list.

Spellings are normalised in one place, chord.rs: Ctrl+z, Control+Z and
LeftArrow all mean what the handler's "Ctrl+Z" and "Left" mean. Shift and Alt
count only for letters and named keys, because on the French layout every
digit needs shift and a 6 has to be a 6 however it was typed.

A Rust keymap that both dispatched and was read by the generator was the
alternative. It would have moved the handlers' decisions away from the Slint
state they depend on, and a window that forgot to install it would have had
no working keys at all.

The keys that were already bound and undocumented are now tagged.
2026-09-24 23:42:25 -04:00
dtourolle 150e53e878 Link fifty gestures on the help sheet to the manual section that shows them
The sheet could now offer "See it", but no gesture said where to look.

Every GESTURE tag whose move the manual describes names that section:
the white-balance picker, zoom and pan, masks, undo and snapshots,
export, colour labels and ratings, selection, collections, the People
page, thumbnail size. Fifty of the fifty-one; the one left, putting a
single control back to its default, has no section and is too small to
earn one.

Three important gestures had nowhere to land, so the manual gains three
short sections, without pictures for now: Bursts (opening a folded
burst, choosing the frame it shows, and the eyes-open filter), Moving
between photographs (the roll, the arrows and A/D in develop) and
Copying settings (Copy, Paste, Paste to N, and choosing what a copy
carries). The page, the gesture book and docs/gestures.md are
regenerated from them.
2026-09-24 23:25:37 -04:00
dtourolle 7591738c73 Let a gesture name the manual section that shows it
The help sheet says which move does a thing, and the manual has a
picture of the thing being done, but nothing joined the two: a user
reading "Pinch it with two fingers" had no way from there to the GIF of
it.

A GESTURE tag takes an optional `manual:` field naming a heading of
docs/manual/README.md by its anchor. The scan checks every one against
the anchors the bundled page is rendered with and fails when the manual
has no such heading, so renaming a section cannot leave the sheet
linking to the top of the page; gestures-check carries the same failure
into CI. The anchor goes into gesture_book.rs as a new field, and into
docs/gestures.md as a "See it" link to manual/README.md#anchor. The help
sheet draws a "See it" button beside the title of each gesture that has
one, which opens the bundled manual at that section.

The field is additive: a tag without it is unchanged, and no gesture
carries one yet.
2026-09-24 23:24:17 -04:00
dtourolle 352e59498b Open the bundled manual from Help and from Settings
The packages now carry the manual, but nothing in the application opened
it: the help sheet listed gestures and stopped there.

The help sheet gains a Manual button beside Done, and Settings a Manual
row under About beside the version. Both go through dr_ui::manual, which
finds the installed page through dr_plat::system_data_dirs (the package's
share directory on Linux, the executable's directory on Windows), and a
development build also in the checkout it was compiled from. A copy with
no manual says so on the status line rather than doing nothing.

On the desktop the page goes to the system browser. A section is a URL
fragment, and xdg-open's generic mode and Windows' FileProtocolHandler
both turn a file: URL into a path and drop the fragment, so a section is
opened through a one-line redirect page written to the data directory:
the opener gets a plain path, which every opener keeps, and the browser
follows the redirect to index.html#section itself. The launcher behind
the sign-in's open_in_browser is split out so both share it; the https
check stays with the sign-in.

Android has no path to give a browser: an asset is not a file, a copy in
private storage is unreadable to other apps, a file: URI across apps is
refused, and a content: URI leaves the browser resolving every picture
against the provider. So ManualActivity, a WebView reading
file:///android_asset/manual/index.html straight out of the APK, shows
it, started by class name with the section as an extra. JavaScript is
off, links off the page go to the browser, and the theme is day-night so
the page's own light and dark follow the system. A test checks that the
manifest, the Java class and dr_ui agree on the name and the extra.
2026-09-24 22:56:09 -04:00
dtourolle d8f26fb5cd Ship the manual with the Arch package, the Windows installer and the APK
The rendered manual was in the repository and nowhere else, so an
installed application still had nothing to open.

Each packager now carries docs/manual/index.html and its pictures, to
where the application will look for them: /usr/share/darkroom/manual on
Arch, manual\ beside darkroom.exe on Windows (where the models already
are, and where dr_plat::system_data_dirs points), and assets/manual in
the APK, stored rather than deflated since a GIF or PNG is already
compressed. The manual is about 27 MB, which the APK and the installer
both grow by; the pictures are 1600x1100 screenshots and short GIFs,
and against an APK that already carries 170 MB of inference runtime and
70 MB of models they are not worth re-encoding for.

The pictures are LFS objects, so each packager refuses a pointer where a
picture should be, as it already does for the models: shipped, a pointer
is a manual of broken images that nothing reports. The Android and
Windows CI legs therefore fetch docs/manual/media, which they excluded
while nothing they built read it, and the installer smoke test checks
that the page and every picture were installed.
2026-09-24 22:33:41 -04:00
dtourolle 10216355c1 Render the manual as one HTML page the application can carry
The manual existed only as docs/manual/README.md, which the forge renders
and nothing else does. An installed copy of the application, on a laptop
with no network or on a tablet, had no manual it could open.

`traces manual` renders the README to docs/manual/index.html with
pulldown-cmark (already in the tree as Slint's Markdown parser, so this
adds a dependency edge and no crate). The page is one file with an inline
stylesheet that follows the system's light or dark preference, a
contents list of every section and subsection, and the pictures by their
relative media/ paths. Each heading carries the id the forge gives it, so
README.md#rating-and-flagging and index.html#rating-and-flagging are the
same link. A picture alone in its paragraph becomes a figure whose alt
text is shown as the caption, and every picture reserves its 16:11 box
before it loads, so a jump into the middle of the page lands where it
aimed rather than a screenful above. Links to design documents, which the
installed page has no copy of, point at the forge.

The page is committed rather than rendered at build time, as the gesture
book is: it is user-facing text reviewed in the diff, and the three
packagers then only copy it. `traces manual-check` fails in CI when the
committed page is not the render of the README, and the pre-commit hook
regenerates it when the README is staged.
2026-09-24 22:32:30 -04:00
dtourolle d8e031888e Regenerate the gesture book, which four rebases left with conflict markers
docs/gestures.md on master carried 162 lines of <<<<<<< / ======= / >>>>>>>
from 4642c77, ade627a, c3d1f83 and 46f5b95. Their branches were rebased
onto each other, the generated files conflicted, and the resolution
regenerated the requirements matrix with `traceability -- report` and then
staged gestures.md as it stood, on the assumption that the same command
writes it. It does not: the gesture book has its own `gestures` mode. The
source tags were never in conflict, so nothing is lost; this is the file
regenerated from them, and `gestures-check` passes on it.
2026-09-24 22:24:41 -04:00
dtourolle 114d979397 Add Vertical and Horizontal perspective sliders to Compose
The keystone existed in framing but nothing in develop could reach it:
framing is presented by its own Compose panel rather than generated, so
new framing parameters get no control until the panel names them.

Compose now has Vertical and Horizontal sliders under Straighten,
mirrored from the session like the angle, recorded as parameter steps
("Vertical Perspective" in the history), cleared by the Compose reset
and by opening the next photograph. Releasing either slider refits the
crop the way releasing the straighten slider does: a keystone alone
needs no crop, but it moves the empty corners of a straightened frame,
so the crop that avoided them before may not after, or may have room
to grow back.
2026-09-24 22:13:12 -04:00
dtourolle 5a500118ae Correct converging verticals with a keystone in framing
There was no perspective transform anywhere in the pipeline: framing
offered a ±45° straighten, quarter turns and flips, and a building shot
looking up kept its leaning walls.

Framing gains a vertical and a horizontal keystone (-100..100). They are
parameters of framing rather than a new stage, so they carry its Compose
attribute, persist in the sidecar under framing, and are withheld from a
default paste exactly as the crop is. In the prologue the keystone runs
after the crop and the straightening and before the stored orientation
and the lens warp, so "vertical" is the photograph's displayed height and
the lens still sees its whole frame.

The map takes the output frame onto a trapezoid inside the source, built
as a homography from four corners and uploaded as three columns in the
framing uniform block (which grows from two vec4s to five). A keystone on
its own therefore never exposes an empty corner and leaves any crop valid.
Combined with a straightening angle the empty area is a pulled-back
quadrilateral the closed-form inscribed rectangle cannot describe, so
max_inscribed_crop searches for the largest centred rectangle whose
corners all have a source pixel behind them. source_at and output_at
apply the same map, so masks, gradients and spot handles follow it.
2026-09-24 22:13:11 -04:00
dtourolle ff89a4fa21 Specify perspective correction as FR-DEV-20
Issue #13 asks for a vertical and horizontal keystone, and its number was
renumbered from FR-DEV-19 when the spec gave that to mask editing. The
clause was never written into the register, so the work had nothing to
trace to.

It is written as part of framing: after the crop and the straightening,
before the stored orientation and the lens warp, carrying framing's
Compose attribute, and with the inscribed crop accounting for it.
2026-09-24 22:13:11 -04:00
dtourolle 04495afbfd Regenerate the traceability matrix for the colour-label work
The pre-commit hook left the matrix as it stood on two of the commits before this one, so it named neither the new test file's NFR-A11Y-3 tag nor the line numbers the label code moved. Regenerated from the tree as it now is.
2026-09-24 21:52:23 -04:00
dtourolle 96f1d5c896 Say in the manual and the register that colour labels exist
The manual's library section described rating and flagging only, and
the outstanding register still said colour labels were "set and shown
nowhere" and that three NFR-A11Y-3 clauses had no test. Both are now
untrue: the manual gives the keys, the Label button and the chips, and
the register names the test file and what it can and cannot vouch for.
2026-09-24 21:52:23 -04:00
dtourolle f1db919b9d Fail a test when a star, a flag or a label differs by colour alone
NFR-A11Y-3 was argued in comments beside the rating strip, the
pick/reject mark and the focus-peaking chips, and nothing would have
failed if an edit made a set star differ from an unset one only in tint.
Only the clipping readout had a test.

These read the markup, as the accessibility-name tests do, since a
rendered window cannot be asked what a colour-blind reader sees. The
star and the flag must choose their glyph from their state, and the
glyphs they choose must be different drawings in icons.slint. The peaking
chips must be distinct words that reach the screen as text. Each colour
label must carry its own letter and the name the catalog's code stands
for, the mark must draw the letter, and the grid cell, the filter chips
and develop must draw the mark or the name. Breaking any of these by
hand makes the matching test fail.
2026-09-24 21:52:23 -04:00
dtourolle 46f5b95828 Show and set colour labels in the grid and develop, and filter by them
Colour labels could be read from a Lightroom sidecar and queried by the
selector, but nothing drew one or set one, so the only labels a library
held were ones another program had written.

Every mark carries its label's initial on its colour — R, Y, G, B, P —
so a label is read without telling red from green, which is what
NFR-A11Y-3 asks of colour labels by name. A grid cell shows the mark
before its filename. In the grid, 6, 7, 8 and 9 set red, yellow, green and
blue as Lightroom's keys do, on the photograph under the pointer or on
the selection by the rule the star keys follow; the same key again takes
the label off, and over a mixed selection it sets it on all. The
selection bar gains Label, which opens the six choices — each a mark and
a name — and purple, which has no key, is there. In develop the top bar
says "Label: Green" beside the mark, opens the same choices, and 6-9
label the open photograph.

Each gesture is one catalog transaction, then the grid, the counts and
both sidecars are written as a rating's are. The filter bar gains a chip
per label, its mark and its name with a count, one at a time; the filter
is one SQL term, travels in the place record, and "All" clears it.
2026-09-24 21:52:23 -04:00
dtourolle d748527a4c Keep colour labels in the sidecar so they survive and travel
A rating and a flag are written to DarkRoom's sidecar as well as the
catalog, because the catalog is a disposable index and the sidecar is
how a judgement reaches the photographer's other devices. A label had no
place there, so once labels could be set, one would have lived only in
the catalog of the device it was set on and gone with it.

The sidecar version now carries `label` (0 none, 1-5 as the catalog
codes it), written only when set. It merges under the rating's rule, so a
device that never labelled a frame cannot clear another device's label,
and a code this build does not know reads as none rather than as some
other colour. A judgement write carries the catalog's label with the
stars, and the scan takes a sidecar's label into the catalog when it has
one. An older build keeps the line as an unknown key and writes it back.
2026-09-24 21:52:23 -04:00
dtourolle 89859d39d1 Let the catalog set colour labels, toggle them, and count them
Colour labels reached `versions.label` only from an XMP sidecar: nothing
in the catalog could set one, clear one, or read it back alongside the
stars, so there was nothing for an interface to call.

`set_label` and `set_label_many` write it the way ratings are written,
the bulk form in one transaction so a key over a selection is one commit.
`toggled_label` holds Lightroom's rule for a label key: it clears only
when every image already carries that label, and otherwise sets it on all
of them, so a half-red selection comes out red rather than inverted.
`Judgement` carries the label, so the grid's one window query brings it
with the stars, and `label_histogram` counts each label in one grouped
statement for the filter chips. A label does not make a frame "judged":
it is a pile of the photographer's own, not a cull decision.

The doc comment for `default_version_id` had been stranded above
`label_code` when that was inserted; it is back on its function.
2026-09-24 21:52:23 -04:00
dtourolle ade627a5d0 Fade the draft into the sharp frame when a drag settles
When a gesture stopped, the half-resolution draft was replaced by the
full-resolution frame in one step, a visible jump from soft to sharp.
FR-DSP-4 asks for a refinement that is smooth, not a jarring swap.

The canvas now keeps the last draft frame (`canvas-previous`) and draws
it over the sharp one, fading it out over 150 ms when the draft flag
clears. The fade costs no render: the draft is a refcount on the texture
it was drawn into, and the adjust pass ping-pongs between two output
targets, so the sharp frame is written into the other one. While a
gesture is drafting the layer is hidden and snapped opaque, so a new
drag shows its draft at once; past the fade it is hidden again, and a
settled canvas composites one image as before.
2026-09-24 21:52:23 -04:00
dtourolle 4642c77e18 Dim the histogram while the canvas shows a draft
The histogram is measured on settled frames only, so during a drag it
describes the frame from before the gesture while the canvas shows
something newer, and nothing said so. The draft flag stopped at the
render closure.

It now reaches the interface: `canvas-draft` on the window and
`Levels.provisional` for the readouts, both set on every canvas render
from the flag that chose the frame's resolution, and cleared when a
render fails. The histogram panel dims its display reading to half
while a draft is up and brings it back when the frame settles. Dimmed
rather than captioned, because a caption appearing on every drag would
move the column; the raw reading has no frame to lag and is left alone.
2026-09-24 21:52:23 -04:00
dtourolle 7d0870c3fb Keep a drag in draft until it stops, then render sharp once
During any drag longer than 120 ms the canvas rendered a full-resolution
frame every 128 ms under the finger. The settle timer was armed by the
first draft of a burst and not re-armed by later ones, so it counted from
the start of the gesture rather than from its last movement, fired
mid-drag, and the next coalesced event armed it again. Each of those
frames is the most expensive one the canvas draws, landing where the
frame budget is tightest.

The draft/sharp decision now lives in `refine::Refine`, apart from the
timers that carry it out. Every draft frame arms a settle timer carrying
a generation token and only the newest token is honoured, so the sharp
frame lands SETTLE_DELAY after the last movement. A request arriving
while a settle is still owed also counts as part of the gesture, so a
slow stretch of a drag (one event per frame, nothing to coalesce) no
longer renders sharp between drafts. A timer that fires with a render
already posted defers to it.

The tests drive the state machine through simulated timelines; the
long-drag case reproduced the four mid-drag sharp frames before the fix.
2026-09-24 21:52:23 -04:00
dtourolle c3d1f83b96 Say so when a crop leaves a mask outside the frame
Cropping tighter past a mask layer made it invisible without a word:
the layer stayed in the panel and the sidecar, and its adjustment went
on landing on pixels nobody would see again.

When a crop is let go, develop now measures what the gesture did to the
mask stack (dr_pipeline::orphan) and, if any layer is now entirely or
mostly outside the frame, shows a notice over the photograph: how many
layers, their names, "Undo crop" and "Keep crop". The crop is already
applied and nothing waits on the answer.

The crop overlay gains a release callback carrying the rect the press
began from, so the measurement runs once per gesture and never on the
drag's per-frame changes. Choosing a ratio is measured the same way,
being a crop committed in one click.

"Undo crop" is the ordinary undo, and the notice is tied to the history
revision it was raised at: the redraw that follows any history move
clears it, so the crop and its warning go back as one step. A second
drag folded into the same step is measured from where that step began.
A crop that strands nothing shows nothing.
2026-09-24 21:52:03 -04:00
dtourolle fc0ea8824d Format the crop-orphan measurement
rustfmt wraps two tuples in dr_pipeline::orphan that the previous commit
left on one line past the width limit. No change in behaviour.
2026-09-24 21:52:02 -04:00
dtourolle 9772785f81 Measure which mask layers a crop takes out of the frame
Mask geometry is stored in source coordinates, so re-cropping tighter
never destroys a layer. It makes it invisible: the layer stays in the
panel and the sidecar, its adjustment lands on pixels nobody will see,
and nothing says so. The spec had no clause for this; FR-DEV-17 now
states it, under the ID issue #10 reserved.

dr_pipeline::orphan samples each layer's mask on a 64x64 lattice over
the source, with the gradient, radial, brush and model-raster geometry
the mask shader uses, folds the parts by their joins and inversions, and
maps the samples through the framing to see how much of the coverage
the crop keeps. `hidden_by_crop` reports the layers whose share fell
below a tenth, and only those the change newly hid, so an already
stranded layer is not announced again on every later adjustment.

Ranges follow the picture and region selections need a label map this
crate does not hold, so a layer that adds either is never reported: a
false alarm on the common path would teach the notice to be dismissed
unread.
2026-09-24 21:52:02 -04:00
dtourolle 733a033274 Test that a stub decoder reaches the scan, the ladder and export
FR-RAW-2's "without changing callers" needs a test that would fail if a
caller named the concrete decoder; passing a real RAW through rawler
cannot tell the two apart, because both routes give the same answer.

The decoder_seam tests hand a stub decoder, for a container no real
decoder reads, to the catalog scan (read_metadata_only over a folder
backend), the preview ladder (the remote two-stage fetch, an import's
thumbnail and the viewer's no-GPU fallback) and export (open_for_export,
skipped without an adapter). Each assertion is on something only the
stub produces: its camera and date, a header fetched at its 64-byte
budget rather than HEADER_BYTES, preview and sensor sizes turned by its
orientation. Switching collect_metadata or make_thumbnail back to the
free functions fails two of the three tests.

The develop test_support module is widened to the crate so the export
test shares the one headless GPU context the other tests use. The
requirements note for FR-RAW-2 now records the trait as built and the
second decoder as not.
2026-09-24 21:33:14 -04:00
dtourolle 414094bd38 Route dr-ui's decoding through the Decoder trait
With the trait in place the claim still meant nothing while every caller
named dr_decode's free functions: a second decoder would have had to be
threaded through the scan, the thumbnail ladder, import, the viewer,
export, merge and repairs at the moment it arrived.

Each of those now takes a &dyn Decoder and reads headers, previews,
orientation and sensor data through it, including the header budget a
remote fetch asks for (header_bytes) and where it finds the embedded
preview (locate_preview). Only the places that start a job name
dr_decode::default(): the thumbnail, sweep and thumbnail-sweep threads,
the viewer's open handlers, and the request structs a job is handed
(BatchRequest, MergeRequest, the import Request, the repairs Toolkit),
so a caller can be given another decoder by changing what it is handed.

The default is rawler through the same free functions as before, so
nothing a user sees changes. The trait gains Debug as a supertrait so
request structs that derive Debug can carry one.
2026-09-24 21:33:14 -04:00
dtourolle d8fb382ce9 Put the RAW decoder behind a Decoder trait
FR-RAW-2 says a second decoder may be added for broader camera coverage
without changing callers, and D2 names LibRaw as that second decoder.
Nothing tested the claim: dr_decode was one decoder reached through free
functions, so adding another would have meant editing every caller at
the moment there was most pressure not to.

Decoder is an object-safe trait over bytes: header_bytes, metadata,
orientation, locate_preview, preview and decode. Rawler implements it by
delegating to the existing free functions, so behaviour is unchanged,
and dr_decode::default() hands it out as a &'static dyn Decoder, which
is what the places that start work will name. JPEG recognition, decoding
and completeness checks stay free functions: they are not a RAW
decoder's to vary.

Nothing in the trait takes a path or a SourceRef; the decoder states how
much of a file it needs and where its preview sits, and the caller's
storage fetches that.
2026-09-24 21:26:17 -04:00
dtourolle e8f68a92f8 Record intersection as built in the requirement and the mask plan
FR-DEV-19a said a part is added to the mask or taken out of it, and
mask-editing.md listed Intersect as an M2 item with nothing built. Both
now say what shipped: the requirement names the third join and why it is
a product rather than a minimum, and how old and new sidecars read across
it; the plan marks the blend-table row built, names the tests that hold
the GPU to the definition, and splits M2 into what is done and what is
still outstanding (joining non-painted parts from the panel, per-part
distance fields, folding two layers).
2026-09-24 21:25:48 -04:00
dtourolle 8f3df7b68d Format the intersect panel test as rustfmt lays it out
The chained lookup in the_intersect_button_joins_a_part_that_intersects was
one line past rustfmt's width, so fmt --check failed on the branch. Split
as rustfmt wants it; no behaviour changes.
2026-09-24 21:25:48 -04:00
dtourolle 5f0b7ac799 Offer intersection in the mask panel: an Intersect button and a third chip state
The pipeline could now keep only where two selections agree, but the panel
had no way to ask for it: the part row's chip flipped between + and -, and
the buttons under the parts joined an added or a subtracted correction.

An "∩ Intersect" button joins a painted part that intersects, and the chip
on a part row cycles + -> - -> ∩ and round, so an existing part can be
turned into an intersection without being repainted. Both go through the
same session calls as before, indexing Join::ALL, whose first two entries
kept their places. The chip is now a tagged gesture, so it is in the
gesture book.
2026-09-24 21:25:48 -04:00
dtourolle 8cdad3863d Keep only where two selections agree, as a third way to join a mask part
A layer's parts could be added to the mask or taken out of it, and nothing
else. The selections that need composing most are the ones that are
neither: the sky that is also bright, the subject that is also skin. With
union and subtract alone, "this and that" had to be spelled as "this minus
everything that is not that", which needs a second part that selects the
complement and rarely exists.

Join gains Intersect, stored as "intersect" in the part block of a sidecar.
It is the product of the two coverages, dst * src, which is one more
fixed-function blend state beside union's max and subtract's
dst * (1 - src) (mask-editing.md 5.2): the same scratch texture, the same
three vertices, no shader arithmetic. The product equals the minimum
wherever either side is fully in or out, and is the softer reading where
two soft edges overlap. Join::apply spells the three operations on the CPU
so the GPU tests can be held to one definition.

A layer that intersects with a part covering nothing now reports that it
covers nothing, so it is not rasterised as an empty slice. Old sidecars
never contain the word, so they read as before; a build from before this
reads "intersect" as a union, the existing unknown-join fallback, which
keeps the part visible rather than dropping it. Join::ALL keeps union and
subtract at indices 0 and 1 so a stored panel index still means the same
join.
2026-09-24 21:25:48 -04:00
dtourolle 229def0afc Show the file's own pixels at 1:1 and beyond
Zoomed to 1:1 or past it, the develop canvas showed a smoothed blur
rather than the photograph's pixels, so focus and noise could not be
judged at the magnification meant for judging them.

Two things caused it. The canvas only switched to nearest-neighbour
strictly past 1:1, with a margin, so the 1:1 inspection itself stayed
smooth. And the switch mostly had nothing to act on: the pipeline
rendered a viewport-sized frame at every zoom, so past 1:1 it was the
pipeline doing the enlarging - bilinearly whenever a straightening angle
or lens correction was in the chain - and the detail stage then sharpened
and denoised those invented pixels at radii scaled up to match. The
texture reached the canvas already blurred and was presented 1:1.

Now, from 1:1 on, the visible region is rendered at the source's own
resolution (render::render_size) and the canvas enlarges it with
nearest-neighbour, so the blocks on screen are the pixels an export would
have; it is also less shading. The decision lives in two small
functions, render::magnification and render::shows_source_pixels,
measured in physical pixels like one_to_one_zoom, with a half-percent
tolerance so the inspection zoom counts as 1:1 even where fit() rounded
the other edge. Below 1:1 the render and the smooth filter are unchanged.
2026-09-24 21:24:57 -04:00
dtourolle 5569a066ff Upgrade accounts saved as http:// to https on launch
Benchmarks / CPU and I/O (per commit) (push) Successful in 1m53s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Successful in 45m7s
Build and test / Layer separation (push) Successful in 41s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 2s
Traceability / Requirement traces (push) Successful in 40s
Build and test / Android (aarch64) (push) Successful in 28m49s
Build and test / Windows (x86_64, cross) (push) Successful in 17m9s
Build and test / Publish the release (push) Skipped
Before the previous commit, browser sign-in could store an account as
http://, and the client now refuses to send to one. Left alone, such a
library would fail to open with a configuration error, so its stored
endpoint is rewritten before anything reads it.

The endpoint is half of two keys, and both are handled:

- The keyring entry is filed under it. Rewriting only the record would
  strand the app password under the old key and sign the user out, so
  AccountStore::move_endpoint copies the secret across first, rewrites
  the record in place (the last record is the one resumed), and deletes
  the old entry only once nothing refers to it.
- namespace() is built from it and names the catalog directory. For an
  http to https rewrite it does not change, because the namespace strips
  either scheme. A move that would change it is refused, not performed,
  so no later rewrite can abandon a catalog either.

The rewrite is a new BackendProvider::upgrade_endpoint hook, which does
nothing by default, and not a second call to normalise_endpoint. The
folder connector's normalise_endpoint canonicalises the path and needs
it to exist, so running it on every launch would fail a library on an
unplugged disk, or rename one whose path now resolves differently. Only
Nextcloud implements the hook.

If the move fails (for example, a locked keyring), it is logged, the
account is left as it was, and the move is tried again on the next
launch.

Closes #65.
2026-09-24 20:44:19 -04:00
dtourolle ea31791388 Keep the typed server after browser sign-in, and open only https
Login Flow v2 saved the account under the `server` field of the poll
response, not under the address the person typed. That field is the
server's idea of its own URL. Behind a TLS-terminating proxy without
`overwriteprotocol` (a common setup) it says http://, and the account then
sent its app password in the clear on every request after that. The
typed address, already upgraded to https by normalise_endpoint, has just
carried the whole flow, so it is the one kept.

The flow's other two URLs come from the server as well, and are now
upgraded from http to https, and refused if they use any other scheme:

- The login URL is handed to the OS to open. On Windows that is
  `rundll32 url.dll,FileProtocolHandler`, which runs a file: or UNC path
  rather than showing a web page, so a hostile server could launch a
  program when the user starts signing in. open_in_browser also refuses
  anything that is not https, as the last check before a process starts.
- The poll endpoint is where the app password comes back from.

The host is not checked. A server reached by its LAN address can answer
with its public name, and refusing that would break a working setup
without protecting anything: the account is stored under the typed
address whatever the server says.

Part of #65.
2026-09-24 20:44:19 -04:00
dtourolle adade27de4 Refuse plain http in the Nextcloud client, below every URL it sends
NFR-SEC-3 held only for the address a person types: normalise_endpoint
upgrades it to https, and nothing else was checked. The login flow's poll
endpoint, an account an older build saved and a redirect all come from
somewhere else, and any of them naming http:// would send the app
password in Basic auth in the clear.

http_client now sets https_only. reqwest checks it before connecting and
again on each redirect, so a refused request never opens a socket, which
the new test checks with a listener that nothing may reach.

A refused scheme is reported as a Configuration error, not Network. The
request never left the process, and Network puts the app into offline
mode over a connection that is working. Other builder errors (a URL that
does not parse) go the same way, for the same reason.

Part of #65.
2026-09-24 20:44:19 -04:00
dtourolle a3f3e188e1 Move the people tray's ticks in place instead of rebuilding it per press
Benchmarks / CPU and I/O (per commit) (push) Successful in 1m52s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Successful in 45m4s
Build and test / Layer separation (push) Successful in 41s
Traceability / Requirement traces (push) Successful in 29s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 2s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Successful in 29m25s
Build and test / Windows (x86_64, cross) (push) Successful in 34m4s
Build and test / Publish the release (push) Skipped
Filtering the grid by a face crawled on the reference library. The SQL
is not it — the person predicate counts in ~20 ms, the eyes-open term in
~60 — but every press on the tray ran `push_people_chips`, which read the
whole people table (26,362 rows, nearly all empty groups a regrouping
pass left behind) and then called `push_people_roster`, which read it
again and replaced the roster model. The roster is every person holding
a face, 1,581 chips, in a row Slint does not virtualise: a new model
tore down and re-created all of them and laid the row out again, to
move one tick.

A press now walks the roster model and sets `picked` on the rows whose
tick changed; the roster is built only when the tray opens. Both reads
use `people_in_use` (2,140 rows) rather than `people`. A picked person
the in-use query leaves out — emptied by a split while the filter held
them — still gets a chip, since a term with no chip cannot be removed,
and without one the in-place update would fall back to a rebuild on
every press.
2026-09-24 20:21:50 -04:00
dtourolle 031315bdb6 Run cargo fmt over the develop shortcuts and the star-range filter
Benchmarks / CPU and I/O (per commit) (push) Successful in 2m8s
Benchmarks / Frame budget (on demand) (push) Skipped
Traceability / Requirement traces (push) Successful in 33s
Build and test / Android (aarch64) (push) Successful in 14m21s
Build and test / android-image (push) Successful in 1s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / Desktop (Linux) (push) Successful in 47m57s
Build and test / windows-image (push) Successful in 1s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / Layer separation (push) Successful in 51s
Build and test / Windows (x86_64, cross) (push) Successful in 34m3s
Build and test / Publish the release (push) Successful in 1m3s
bddf325 and 00c028c went in unformatted, so the Desktop job's
`cargo fmt --check` step failed on master (run 1693) and the release job
that needs it was skipped. Whitespace only.
2026-09-24 20:05:02 -04:00
dtourolle 00c028c8c8 Rate under the pointer, filter a star range, and name Help as help
Benchmarks / CPU and I/O (per commit) (push) Successful in 1m53s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 2m50s
Build and test / Layer separation (push) Successful in 33s
Traceability / Requirement traces (push) Successful in 40s
🐳 Android image / Build and push (push) Successful in 4s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 4s
Build and test / windows-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Successful in 14m19s
Build and test / Windows (x86_64, cross) (push) Successful in 17m50s
Build and test / Publish the release (push) Skipped
Rating keys in the grid follow darktable's rule: with the pointer over a
photograph outside the selection, 0-5, P, X and U judge that photograph
alone; over one inside it, the whole selection, as before; off the grid,
the selection. The hover is cleared when the grid scrolls, so a key after
a wheel turn cannot judge whatever used to be under the pointer.

Holding F and tapping digits filters by stars: one digit for exactly
that many, two for everything between them, F alone to show every
rating again. The filter gains a ceiling to do it (`max_rating`, one
BETWEEN in the query). The place record carries it, and a record from an
older build reads as having none. The star chips light across a capped
range and the bar says "2-3★ only" beside them.

The grid also takes Ctrl+E and Ctrl+Shift+E for the selection, Ctrl+V to
paste onto it and Ctrl+A to select all. The "Gestures" button is now
"Help", its sheet "Controls and shortcuts", and F1 opens it.
2026-09-24 05:11:36 +02:00
dtourolle bddf3250c5 Add Lightroom's export and copy shortcuts to develop
Ctrl+E opens an export sheet: the export defaults on their own, over the
photograph, with an Export button. Ctrl+Shift+E exports straight away on
those defaults. There is no per-export copy of the settings, so what is
chosen in the sheet is saved as it is on the settings page, and the next
Ctrl+Shift+E uses it.

To make that one set of controls in two places, the export options move
out of the settings page into export.slint: an `ExportOptions` global
that Rust writes once, and two panels that read it. The window no longer
forwards forty `settings-*` properties to the page.

Ctrl+Shift+C opens a copy sheet with the edit-kind chips the preset
sheet already uses and a Copy button, which is how a paste leaves each
photograph's crop and rotation alone (Compose off). A and D step along
the roll beside the arrows. While either sheet is up the develop keys
stand down, so A cannot change the photograph behind the form, and
Escape closes it.
2026-09-24 05:11:16 +02:00
dtourolle 41486bd59b Step to the next photograph from the keyboard in a library's develop
The arrow keys and space in develop called `next-image` and
`prev-image`, which walk the files given on the command line. A
photograph opened from a library leaves that list empty, so the keys did
nothing and only a click on the photo roll moved on.

With a library open the step now goes through the roll: it opens the
neighbouring frame exactly as clicking it would, and saves the outgoing
edit the same way.
2026-09-24 05:09:39 +02:00
dtourolle d6d27fb062 Publish a Gitea Release from CI on every v* tag
Benchmarks / CPU and I/O (per commit) (push) Successful in 2m28s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Successful in 45m9s
Build and test / Layer separation (push) Successful in 38s
Traceability / Requirement traces (push) Successful in 44s
🐳 Android image / Build and push (push) Successful in 5s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 5s
Build and test / windows-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Successful in 14m25s
Build and test / Windows (x86_64, cross) (push) Successful in 34m30s
Build and test / Publish the release (push) Skipped
Nothing made a release. CI built the APK and the installer on the master
push and kept them as workflow artefacts, the Linux binary was not kept
at all, and most tags went out with no downloads until they were
attached by hand.

build-and-test now also runs on v* tags. On a tag the desktop job keeps
its release binary, and a release job that needs desktop, Android and
Windows collects the three, names them with the version and runs
tools/publish-release.sh. The script titles and describes the release
from the annotated tag's message as the server holds it, writes
SHA256SUMS, and attaches what is not already there, so a re-run after
an interrupted upload finishes the job instead of duplicating it. The
same script is how a release is made or finished by hand.

Tried on v0.14.1, whose release was made by hand with the same files:
it found the release, reported all four files attached, and changed
nothing.
2026-09-24 03:43:42 +02:00
dtourolle 317a2f40bd Release 0.14.1
Benchmarks / CPU and I/O (per commit) (push) Successful in 5m19s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Successful in 43m6s
Build and test / Layer separation (push) Successful in 37s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 3s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 1s
Traceability / Requirement traces (push) Successful in 46s
Build and test / Android (aarch64) (push) Successful in 29m7s
Build and test / Windows (x86_64, cross) (push) Successful in 34m31s
2026-09-23 19:07:22 -04:00
dtourolle 949fe40d5b Pin "Export N" to the right of the library's selection bar
It was the last of a dozen buttons in a row that scrolls sideways once
it outgrows the window, which at a desktop width it does. The button
sat past the right edge and nothing says the row scrolls to a mouse,
so batch export of a selection looked like a feature the library did
not have. The rest of the row still scrolls; Export, which is also the
cancel for a running batch, now sits beside it and is always visible.
2026-09-23 19:05:42 -04:00
dtourolle 2be80d4203 Show "Paste to N" on every selection, disabled until something is copied
It appeared only once settings had been copied this session, so with an
empty clipboard nothing on the selection bar said pasting onto a
selection was possible. It is one of the things a selection can have
done to it, like filing it in a collection, and now sits with them
beside Presets.
2026-09-22 21:37:41 -04:00
dtourolle 9ebaa15099 Move copy, paste and presets into the develop top bar
They sat in the develop column under a "SETTINGS" heading, which read as
application settings, and went away with the panel toggle and in the
mask and spot modes. The strip is where undo already is for the same
reason: these act on the whole edit, not on any one panel.

The paste button still names what it would apply. The TransferPanel
component is gone; the Transfer global and its Rust wiring are
unchanged.
2026-09-22 21:37:31 -04:00
dtourolle aee355fada Make the in-flight claim test wait for the waiter to park
a_second_claim_waits_for_the_first_to_be_released failed on CI: the
second claim came back Some. The waiter thread signalled the main thread
before calling claim, so the main thread could drop the first guard
before the waiter reached the lock. The path was free by then, and the
waiter claimed it outright.

The registry now keeps a test-only count of threads parked in claim,
bumped under the lock just before the condvar wait. The test spins until
that count is one before releasing. The release needs the same lock, so
it can only reach a waiter that is already waiting. Passed 500 runs in a
row.
2026-09-22 21:12:45 -04:00
dtourolle 7fa3176f88 Release 0.14.0
Benchmarks / Frame budget (on demand) (push) Skipped
Benchmarks / CPU and I/O (per commit) (push) Successful in 8m25s
Build and test / Desktop (Linux) (push) Failing after 26m11s
Build and test / Layer separation (push) Successful in 33s
🐳 Android image / Build and push (push) Successful in 10m41s
Build and test / android-image (push) Successful in 10m42s
🐳 Windows image / Build and push (push) Successful in 4m17s
Build and test / windows-image (push) Successful in 4m17s
Traceability / Requirement traces (push) Successful in 59s
Build and test / Android (aarch64) (push) Successful in 40m53s
Build and test / Windows (x86_64, cross) (push) Successful in 46m51s
2026-09-21 11:32:41 +02:00
dtourolle af162dd010 Merge: origin's sync ordering and adoption work, its scan-complete trigger ported into library_ui/ 2026-09-21 11:32:05 +02:00
dtourolle f5300a9f43 Merge: the UI crate restructured so a feature owns a function, not a region
Nine agent branches, merged one at a time on ui-wiring and gated at each
step: develop.rs, library.rs, collections_ui.rs and library_ui.rs become
module directories; every screen's wire() and lib.rs::run() become lists
of named functions, with the develop screen's callbacks in develop_ui.rs;
the six view booleans become View and Page enums; the collections sidebar
and the library grid get their own Slint globals, taking 155 members off
AppWindow. No behaviour change: a multiset audit of code lines over every
moved region lost nothing, the workspace gate is clean, and the manual
recorded from this build and from master's from the same library snapshot
matches picture for picture, apart from a panorama stall that master
shows too when a merge starts during the engine's TensorRT compile queue.
2026-09-20 23:41:25 +02:00
dtourolle b19d470189 Merge: the library grid on its own Slint global, the last of the screen state off the root 2026-09-20 22:43:09 +02:00
dtourolleandClaude Opus 5 bfadd9c409 Ship the border filler in the Windows installer, and count what is staged
Benchmarks / CPU and I/O (per commit) (push) Successful in 8m2s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 25m17s
Build and test / Layer separation (push) Failing after 0s
Traceability / Requirement traces (push) Failing after 0s
🐳 Android image / Build and push (push) Failing after 0s
Build and test / android-image (push) Failing after 0s
Build and test / Android (aarch64) (push) Skipped
🐳 Windows image / Build and push (push) Failing after 0s
Build and test / windows-image (push) Failing after 0s
Build and test / Windows (x86_64, cross) (push) Skipped
The installer smoke test asserted seven model files, the number on the
day it was written; models/face has since gained the eye-state trio's
companions and the int8 detector forms, and the run on f71d7ba failed
with thirteen installed. The test now expects as many files as
package.sh's directories hold, so the next model needs no edit here.

package.sh also stages models/inpaint, which the APK and the Arch
package already carry and the Windows build did not: without
migan-512.onnx the panorama's border fill has no model on Windows.
xfeat needs nothing, it is embedded in the binary. windows.md §5.2
lists the result.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 22:35:04 +02:00
dtourolle 050b2a3914 Collapse the blank runs the library global's move left behind
Moving each group of properties and callbacks out of AppWindow left
several two- and three-line gaps where a removed block's neighbours no
longer needed separating. No declarations changed.
2026-09-20 22:32:47 +02:00
dtourolleandClaude Opus 5 195388b2e3 Adopt faces a hundred images per commit, and hold one generation per image in the shards
Benchmarks / CPU and I/O (per commit) (push) Successful in 8m16s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 4h42m34s
Build and test / Layer separation (push) Successful in 59s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 2s
🐳 Windows image / Build and push (push) Successful in 2s
Build and test / windows-image (push) Successful in 2s
Traceability / Requirement traces (push) Successful in 50s
Build and test / Android (aarch64) (push) Failing after 0s
Build and test / Windows (x86_64, cross) (push) Failing after 0s
The shard import recorded each adopted image in its own transaction:
fourteen thousand commits, and fourteen thousand turns at the write lock
that every read on the UI thread queued behind — the sync was felt as a
laggy grid and as "database is locked" from whichever writer lost the
wait. `record_detections_within` takes the caller's transaction, and the
import commits every hundred images.

The store carried every detector generation of an image — 24,123 entries
for 19,089 images on the reference library, a third of its 293 MB — when
only the strongest is ever adopted. A put now skips a pass a held one
outranks, and retires the passes it outranks from the index; sealed
shards keep their bytes, but nothing is written twice from here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 22:31:47 +02:00
dtourolle e166b64ee9 Format fill.rs, which the last release left past the width limit 2026-09-20 22:29:42 +02:00
dtourolle a773ad5c27 Merge: the developer docs under docs/dev, and the folder indexed for users first 2026-09-20 22:27:48 +02:00
dtourolle 1e472fd251 Move the library's routes and status onto its own Slint global
The scan/opening/status/error lines, offline mode and its retry, pinning a
collection offline, the sync and thumbnail-sweep state, and the callbacks
that route the grid to a rescan, a sync, another library, a panorama merge
or an export — the last of what AppWindow still carried under the
library- prefix — move onto the `Library` global started earlier on this
branch. Rust reaches them through window.global::<Library>() rather than
window.set_/get_/on_/invoke_ on the root.

library-visible is the one name that stays: it is computed from
active-page and active-view, the shell's own routing state, which a
global cannot read. AppWindow now declares no other library- property or
callback.
2026-09-20 22:22:55 +02:00
dtourolle 6bf67cefc4 Move the filter bar and the timeline onto the library's Slint global
The date-range fields and their band on the capture-time axis, the
timeline's bars, labels, scrub/pinch/pan/zoom callbacks and its
sweep/current-bucket/anchored state, and the filter bar's people chips,
mode, eyes-open toggle and gesture reference move from AppWindow onto the
`Library` global. Rust reaches them through window.global::<Library>()
rather than window.set_/get_/on_ on the root, as the earlier commits on
this branch did for the grid's cells, selection, ratings and keywords.
2026-09-20 22:07:12 +02:00
dtourolle 00e2fe6aaf Move ratings, flags and keywording onto the library's Slint global
The keywording sheet's rows and its open/assign/unassign callbacks, the
star and flag callbacks a cell click or a judgement key fires, the burst
toggle and representative-chosen callbacks, the trash-selection shortcut,
and the rating/unjudged/flag filter chips with their rating-counts model
move from AppWindow onto the `Library` global started in the previous
commit. Rust reaches them through window.global::<Library>() rather than
window.set_/get_/on_ on the root.
2026-09-20 21:57:40 +02:00
dtourolle 402dcdc24c Move the library grid's cells and selection onto their own Slint global
AppWindow carried the grid's loaded window of cells, the keyboard cursor,
drag and drop, the held-row long-press state, columns and cell size, the
scroll and viewport bookkeeping, the photo roll's pick and centre-request,
and the local-only/reorder/collection-filing gestures that act on a
selection, as properties and callbacks on the root component. That state now
lives in the `Library` global declared in library.slint, next to the structs
(LibraryCell, TimelineBar, KeywordRow, PersonChip) it and the grid's other
components already share; Rust reaches it through
window.global::<Library>() instead of window.set_/get_/on_/invoke_ on the
root, the same change collections.slint's `Collections` global made for the
sidebar.

library-visible stays on AppWindow: it is computed from active-page and
active-view, the shell's own routing state, which a global cannot read.
Everything else still prefixed library- — the timeline, the filter bar,
ratings and flags, keywording, and the routes and status lines — stays on
the window for now and moves in the commits that follow.
2026-09-20 21:42:48 +02:00
dtourolle 94b39410bc Tidy the ported drag block, and format fill.rs as master left it 2026-09-20 21:40:20 +02:00
dtourolle 2014c80e62 Merge: master at 0.13.6, with the drag-ghost file and the shared model lookup ported into the split modules 2026-09-20 21:32:21 +02:00
dtourolleandClaude Opus 5 6fd342680b Format set_indexed_at's signature
Benchmarks / CPU and I/O (per commit) (push) Successful in 8m15s
Benchmarks / Frame budget (on demand) (push) Skipped
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 2s
Build and test / Desktop (Linux) (push) Successful in 1h8m29s
🐳 Windows image / Build and push (push) Successful in 5s
Build and test / windows-image (push) Successful in 5s
Build and test / Layer separation (push) Successful in 34s
Traceability / Requirement traces (push) Successful in 50s
Build and test / Android (aarch64) (push) Failing after 0s
Build and test / Windows (x86_64, cross) (push) Failing after 0s
baed1c4 landed it over rustfmt's width; `cargo fmt --check` is the
first gate the Desktop job runs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 21:31:07 +02:00
dtourolleandClaude Opus 5 78acc73dad Regenerate the traceability matrix for the sync commits
f71d7ba and 34ac2f1 added tags without re-running the report, which
the Traceability job's "is it committed" step rejects.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 21:28:55 +02:00
dtourolleandClaude Opus 5 954246b969 Register the inference requirements, and format the fill test
Two CI gates have failed on every push since 0.13.4 and both are
fixed here.

The traceability gate rejected `FR-INF-1` as an orphan: settings.slint
and dr-ui tag it, but the four register entries inference.md §12 wrote
were never carried into requirements.md, which is the only file the
extractor reads. §3.12 and §4.10 now hold FR-INF-1..3 and NFR-INF-1
verbatim, with the acceptance milestones pointed back at inference.md.
The matrix is regenerated (188 defined, 155 covered) and the README's
"where it stands" line, which the 0.13.6 release commit skipped, says
0.13.6 and the new figures.

`cargo fmt --check` failed on the `fill_border` call in dr-pano's
padding test, which is the first thing the Desktop job runs after
installing the toolchain and why it failed within a minute.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 21:28:55 +02:00
dtourolleandClaude Opus 5 baed1c4782 Keep the peer's run marker on adopted faces, so they are not re-exported as ours
Benchmarks / CPU and I/O (per commit) (push) Successful in 8m16s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 32s
Build and test / Layer separation (push) Successful in 29s
Traceability / Requirement traces (push) Failing after 42s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 2s
🐳 Windows image / Build and push (push) Successful in 2s
Build and test / windows-image (push) Successful in 2s
Build and test / Android (aarch64) (push) Failing after 0s
Build and test / Windows (x86_64, cross) (push) Failing after 0s
Adopting an image from a peer's face shard stamped its run marker as now,
and the export reads a catalog marker newer than the shard's as a
re-index. So every adopted image went straight back out under this
device's client id: 14,100 adopted, 15,457 "newly indexed" on the next
pass, twenty-two shards of a peer's faces uploaded a second time.

merge_shard now carries the peer's indexed_at into the local index, and
the import writes that marker into face_index; where an older peer's shard
carries none, the store takes the catalog's, so the two agree either way
and the export finds nothing to send.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 21:25:53 +02:00
dtourolle fbfa891296 Regenerate the traceability matrix after the rebase 2026-09-20 21:23:23 +02:00
dtourolle aee62dc7f2 Wrap the line the docs move pushed past the width limit 2026-09-20 21:16:03 +02:00
dtourolle 84fade99ec Put the developer docs under docs/dev and index the folder for users first
docs/ had 26 developer documents flat beside the manual, and the two
audiences are very differently sized: most readers want the manual and
the gesture reference, a few want the register, the designs and the
measurements. The manual and gestures.md stay at the top; everything for
someone changing the code moves to docs/dev/, and the two documents that
name their own successors — the v0.1 milestone and the UI-refinement plan
— go to docs/dev/archive/ rather than being deleted, since both are still
cited. docs/README.md is the index, users first.

Every reference follows: code comments, Cargo manifests, the workflows,
the pre-commit hook, the bench and traceability tools (which locate the
repo root by docs/dev/requirements.md now), packaging, the Docker READMEs,
CLAUDE.md, CONTRIBUTING.md and the README. The matrix links one level
deeper and is regenerated. Links out of the moved documents into the tree
gain a level; a link checker over every Markdown file finds none broken.
2026-09-20 21:16:03 +02:00
dtourolle 3bfa73d1e1 Merge: the collections sidebar on its own Slint global 2026-09-20 21:14:00 +02:00
dtourolle 4616cb0a23 Merge: library_ui split into a module directory 2026-09-20 21:12:36 +02:00
dtourolle cc73ea3153 Split library_ui.rs into a module directory by area of behaviour
controller holds LibraryController and the window-sizing constants every
other module reads and writes through pub(super) fields, the same shape
collections_ui and develop already use. open is the launch-to-scan cycle and
the worker that checks the catalog file before either touches it. offline is
what of a collection is on this device and the prompt that offers to change
it. window fills the grid model from the catalog and drains the thumbnail
fetch, which is the piece the catalog-reads-are-proportional-to-what-changed
rule (docs/catalog.md §1) bears on most directly. sync is the background
passes that reach beyond the loaded window: the metadata sweep, the
whole-library thumbnail pass, and the exchange with the server.
ratings_keywords applies a judgement or a keyword to a selection and queues
the sidecar and XMP writes behind it. timeline is the capture-time sidebar
and the photographer's place together, kept in one file because a restored
place ends by moving the timeline marker and a scrub is a restore of one
instant, so most calls between the two would otherwise cross a module
boundary. grid wires the grid's own callbacks — the keyboard cursor,
cell-size zoom, the routes into and out of develop — and filter_bar wires
the rating, people, date and offline-scope filters, calling back into
whichever of the above owns the work a filter change triggers.

Extracted by item rather than by line range, so every doc comment and
TRACES/GESTURE annotation stayed attached to the code it describes; the
sorted set of TRACES/GESTURE lines in the new directory is identical to the
original file's. Tests moved with the code they exercise, including the
handful of fixtures — settle, model_of, with_catalog, zoom_cell, pinch_step
— that only one target module needed and so were not worth sharing through
a test_support module the way the other splits use one. Items that crossed
a new module boundary were widened from private to pub(super), narrower
than the whole-file access the original gave them; a few items already
pub(crate) for recovery_ui or presets stayed there rather than being
narrowed, since nothing needed them tightened further.

mod.rs re-exports the same surface library_ui:: callers used before, so
lib.rs and every other caller needed no change.
2026-09-20 21:10:02 +02:00
dtourolle f6ff5eabd9 Give the collections sidebar its own Slint global
AppWindow carried the sidebar's tree, its row menu, renaming, drag and
drop between rows, the trash row, and the membership sheet as ~40
properties and callbacks on the root component, in the pattern CH-1
describes and the develop screen's globals (Adjustments, Framing, Steps,
...) already replaced. Collections.* in collections.slint now holds that
state, declared next to the MembershipRow struct it and the membership
sheet both use; Rust reaches it through window.global::<Collections>()
instead of window.set_/get_/on_/invoke_ on the root.

collection-selected, collection-select, and collection-offline-menu stay
on AppWindow: library_ui.rs invokes collection-select directly and
registers collection-offline-menu's handler, and lib.rs reads
collection-selected for back-navigation, so moving them would have meant
editing library_ui.rs, which another change on this branch is splitting
into a module directory. collections-visible stays too — it is
lib.rs's panel-layout state, seeded from the saved layout before the
sidebar exists, and collections_ui never touches it. Everything prefixed
library- (the grid's drag, selection and keyword state that the sidebar's
Rust also wires for cross-feature gestures like filing a selection into
a collection) stays on the window as well, since it belongs to the
library screen, not the sidebar.
2026-09-20 21:07:25 +02:00
dtourolleandClaude Opus 5 34ac2f14d3 Sync faces and the catalog before thumbnails, so a fresh device sees its names and collections first
Benchmarks / CPU and I/O (per commit) (push) Successful in 8m24s
Benchmarks / Frame budget (on demand) (push) Skipped
🐳 Android image / Build and push (push) Successful in 4s
Build and test / android-image (push) Successful in 4s
Build and test / Desktop (Linux) (push) Failing after 35s
🐳 Windows image / Build and push (push) Successful in 2s
Build and test / windows-image (push) Successful in 2s
Build and test / Layer separation (push) Successful in 30s
Traceability / Requirement traces (push) Failing after 48s
Build and test / Android (aarch64) (push) Failing after 0s
Build and test / Windows (x86_64, cross) (push) Failing after 0s
The thumbnail stage ran first and, on a device that had just adopted its
peers' shards, spent its time re-uploading hundreds of megabytes under its
own client id while faces, people, collections and dates waited behind it.
Faces go first — the catalog merge assigns identities to faces this device
holds — then the catalog, then thumbnails.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 21:04:10 +02:00
dtourolleandClaude Opus 5 f71d7bacc6 Take the server's shards and dates when the scan completes, not after the sweep
Benchmarks / CPU and I/O (per commit) (push) Successful in 8m28s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 36s
Build and test / Layer separation (push) Successful in 28s
Traceability / Requirement traces (push) Failing after 39s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 2s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Successful in 43m11s
Build and test / Windows (x86_64, cross) (push) Failing after 41m4s
The derived sync fired only after the metadata sweep, so a fresh device
re-derived every thumbnail it scrolled past, re-detected faces and re-read
every header for hours before adopting the shards and snapshot that held
all of it. It now fires as soon as the scan completes — the first moment
the rows the merges key on exist — and the sweep starts behind it. In
steady state that pass is one listing.

The catalog merge gains a fourth half: capture metadata (captured_at,
offset, camera, lens, ISO) for images still at metadata_state < 2, matched
by oc:fileid from a remote row at 2. A date is a fact about the file's
bytes, not local state, and the snapshot already carried it. The sweep's
per-chunk query then finds nothing left, and the timeline is whole on a
fresh device without a header fetch.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 20:37:59 +02:00
dtourolle 14ed1dc410 Merge: View and Page enums in place of the six view booleans 2026-09-20 20:32:41 +02:00
dtourolle 38819da222 Replace the six view booleans with View and Page enums
app.slint carried show-launch, show-library, show-identity, show-settings,
show-import and show-merge as separate booleans, so the root component chose
what to draw with five- and six-term conjunctions and nothing stopped two of
them being true at once. Replaced with two enums: View { develop, library,
identity, launch } for which top-level screen is showing, and Page { none,
settings, import, merge } for which page, if any, is drawn over it.

Two values rather than one, because the two questions are genuinely
different. Settings, Import and Merge are reachable from more than one View
and are drawn outermost without touching it — closing one has to return to
whichever View was already current, and today that works because the
underlying property is left alone while the page sits over it. A single
View with five or more variants would need a second field remembering what
to return to; Page needs nothing to remember, since View was never
overwritten in the first place. Identity, by contrast, genuinely replaces
the window the way Launch and Library do (see the existing "like the launch
screen" comment on its `if`), so it is a View variant, not a Page.

Every `if` chain in app.slint that used to compare four, five or six
booleans now compares active-view and active-page to at most one variant
each. library-visible collapsed from a six-term conjunction to
`active-page == Page.none && active-view == View.library`.

The Rust side follows: every set_show_*/get_show_* call in library_ui.rs,
identity_ui.rs, settings_ui.rs, merge_ui.rs, import_ui.rs, launch_ui.rs and
lib.rs now reads or writes active-view or active-page instead, including
lib.rs's startup match (View.launch vs View.develop, since a Startup that
skips the launch screen used to leave both old booleans false and fall
through the chain to develop) and identity_ui's close handler, which now
writes View.library or View.develop in one call where it used to write
show-library then show-identity separately.

back_one_step needed one deliberate adjustment beyond the mechanical
rename. Identity was never represented in NavState: back had nothing to do
when Identity was opened from the library (show-library stayed true,
unread by IdentityScreen's own condition) and could only reach ToLibrary
when opened from develop, which likewise wrote a property IdentityScreen
never read — so escaping out of Identity was invisible in both cases before
this change. With a single active-view, falling into the general case
would instead overwrite the value IdentityScreen's `if` does read and close
it as an unintended side effect. back_one_step now swallows the gesture
while View.identity is current, reproducing the same "nothing visible
happens" outcome for both origins without threading identity_ui's private
came-from-library state through lib.rs for one screen.

Verified with tools/manual/drive.py against a private Xvfb and the debug
build: launch screen to library, Settings opened and closed, Identity
opened and closed (including Escape doing nothing while it is open),
develop opened from a cell and closed both by the back button and by
Escape. Screenshots under verify/.
2026-09-20 20:30:56 +02:00
dtourolle d6e9c7dc94 Merge: collections_ui split into a module directory 2026-09-20 20:30:12 +02:00
dtourolle 9c8f21b754 Split collections_ui.rs into a module directory by area of behaviour
collections_ui.rs had grown to 4,591 lines covering the sidebar controller,
the click/drag selection policy, tree refresh, the drag gesture, the trash
worker, twelve wiring functions, and the row's rename/create/context menu,
all in one file. Split into collections_ui/ with one module per area, the
way develop/ and library/ were already split on this branch:

- controller.rs: CollectionsController and the pure drop/delete/release
  decisions (decide_drop, decide_delete, decide_release, menu_detail,
  delete_warning) that a test can drive without a window.
- press.rs: PressUndo and the click-and-release selection policy
  (apply_press, select_row, commit_press, cancel_press).
- tree_sync.rs: rebuilding the sidebar from the catalog and pushing
  catalog-derived state into the grid (refresh_tree, offline_state,
  sync_lifted/sync_selection/sync_reorderable/sync_badges,
  refresh_membership, direct_holdings).
- drag.rs: the cursor bitmap (compose_drag_image, blit_scaled) and the
  hold/spring timers (arm_hold, arm_spring, should_spring,
  collapse_spring_opened) plus their delay constants.
- trash.rs: the soft delete (start_trash, start_restore, drain_trash,
  stop_trash, refresh_trash, format_bytes).
- wiring_grid.rs / wiring_tree.rs: the wire() entry point and its twelve
  wire_* functions, split in two because together they were the largest
  single piece (grid-facing selection/drag/trash vs. sidebar-facing
  navigation/create/rename/row-drag/menu/membership).
- rename_menu.rs: creating, naming and renaming collections, and the row's
  context menu (apply_rename, create_child, unique_name, open_row_menu,
  close_row_menu, close_rename).

mod.rs carries the module's own top-level doc comment, the `pub use`
re-exports for the eight items the rest of the crate reaches by
`collections_ui::` path (CollectionsController, wire, refresh_tree,
sync_badges, sync_selection, select_row, commit_press, cancel_press), and a
shared `test_support` for the one fixture (`ids`) more than one file's
tests needed. Every item that only crossed a boundary within this module,
not out of it, was narrowed to `pub(super)` rather than kept at the
crate-wide `pub` a single file gave it for free.

Extracted with a brace-aware pass that kept each item's own leading doc
comment and attributes attached to it, and tests moved with the code they
exercise; every TRACES/GESTURE comment lands on the same code it did
before. No file outside the new directory changed — lib.rs's `mod
collections_ui;` resolves to the directory automatically, and every
outside caller's `collections_ui::` path still resolves through mod.rs's
re-exports.
2026-09-20 20:29:21 +02:00
dtourolle bd7d75522d Merge: the format fix from docs-layout 2026-09-20 20:26:56 +02:00
dtourolle 4ef1b74f2f Wrap the line the docs move pushed past the width limit 2026-09-20 20:26:54 +02:00
dtourolle 681486196e Release 0.13.6
Benchmarks / CPU and I/O (per commit) (push) Failing after 6m17s
Benchmarks / Frame budget (on demand) (push) Skipped
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 2s
Build and test / Desktop (Linux) (push) Failing after 29s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 1s
Build and test / Layer separation (push) Successful in 26s
Traceability / Requirement traces (push) Failing after 42s
Build and test / Android (aarch64) (push) Failing after 2m48s
Build and test / Windows (x86_64, cross) (push) Failing after 4m22s
2026-09-20 20:19:48 +02:00
dtourolle f5d0d57574 Regenerate the traceability matrix after the rebase 2026-09-20 20:19:43 +02:00
dtourolle c96e670356 Re-record the panorama for the trained filler; the scene waits for the preview and for the DNG instead of guessing 2026-09-20 20:19:01 +02:00
dtourolle 9c556364fa Pad an open void's canvas to a tile: the merge page's preview is shorter than one, and filled nothing 2026-09-20 20:19:01 +02:00
dtourolle 8d72cabff5 Ship the border filler trained against MI-GAN's own discriminator: texture in the deep bands, level with stock on LPIPS 2026-09-20 20:19:01 +02:00
dtourolleandClaude Opus 5 8a90d888d5 Probe and compile for a package's models, not only the user's
Benchmarks / CPU and I/O (per commit) (push) Failing after 6m30s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 53s
Build and test / Layer separation (push) Successful in 26s
Traceability / Requirement traces (push) Failing after 40s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Failing after 2m19s
Build and test / Windows (x86_64, cross) (push) Failing after 3m5s
`inference::init` listed the models from the user's shared directory
alone, while the app loads them from there or from the package's
`/usr/share/darkroom/models`. On a fresh package install the probe found
"no model to probe with", stayed on the CPU, and compiled nothing. Both
now resolve each file with the same search, `library::shared_model`.

The PKGBUILD names the ONNX Runtime packages as optional dependencies,
since the app loads one from /usr/lib if present.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 19:39:05 +02:00
dtourolle e86edef47c Merge: run() reduced to construction and startup, the develop wiring in its own module 2026-09-20 19:28:23 +02:00
dtourolle 0aff5e6c8c Merge: the collections, identity, merge, launch and import wiring lifted into named functions 2026-09-20 19:28:12 +02:00
dtourolle 5458314083 Split import_ui::wire into one function per section
The 188-line wire() registered the import page's callbacks in four
comment-delimited sections. Lift each into its own fn: wire_opening_and_closing,
wire_choosing_a_source, wire_options, wire_running. context stays generic
over C on each of the three functions that use it, matching how survey()
and start() already take it (impl Fn, implicitly Sized) rather than coercing
it to a trait object, which would have needed ?Sized added to those two
unrelated functions for no benefit.
2026-09-20 19:23:21 +02:00
dtourolleandClaude Opus 5 39a22875b1 Add the MIGraphX rung for AMD GPUs
Benchmarks / CPU and I/O (per commit) (push) Failing after 6m20s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 45s
Build and test / Layer separation (push) Successful in 26s
Traceability / Requirement traces (push) Failing after 46s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Failing after 2m19s
Build and test / Windows (x86_64, cross) (push) Failing after 3m2s
Measured on a Radeon RX 7900 XT against Arch's onnxruntime-rocm 1.29
(docs/inference.md §1.3): MIGraphX fp16 runs the detectors at 2.4–3.4 ms
against 10–58 ms on the CPU provider, the inpainter at 8 ms against 514,
with a 15–135 s compile per graph the first time and under a second from
its cache after. A compiling rung on TensorRT's terms, wired the same way.

The ROCm execution provider is gone (removed in ONNX Runtime 1.23), so the
AMD ladder is MIGraphX then the CPU, with no non-compiling rung between.

MIGraphX is registered through the runtime's generic key/value entry
point rather than ort's builder: 1.29 reads the legacy options struct for
its precision flags only, and the compiled-program cache directory
(`migraphx_model_cache_dir`) only travels the generic way. The provider's
cache key omits the precision, so f32 and fp16 programs get their own
directories. The probe fingerprint now includes the provider libraries
beside the runtime and the ROCm version, since a distribution's CPU and
ROCm builds are the same file at the same path.

`status().failed` reports only the rungs above the selection, so an AMD
desktop's About line says why MIGraphX won rather than that the NVIDIA
providers are not in the build.

Two examples: `ep_probe` times each provider cold and from cache, and
`ladder` drives `init` as the app does to watch the first-run sequence.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 19:23:00 +02:00
dtourolle cb7b9d71c4 Split lib.rs::run into named construction and wiring functions
run() was 2,724 lines that built every controller, owned the develop
session and its render loop, and registered every develop-screen callback
inline — the state CH-1 in docs/dev/code-health.md describes. This is the
mechanical split CH-1 calls for, done in one pass rather than section by
section since develop_ui.rs only compiles once lib.rs stops registering
those callbacks itself.

develop_ui.rs is new: a DevelopWiring struct holding the session, the rows
model, the redraw/render closures and the other controllers' handles, and
one wire_* function per section run() used to contain — presets, export,
the adjustment panel, film, undo/redo, zoom/pan/crop, rotation/flips/
straightening, navigation and peaking — called in the order run()
registered them. window travels as its own parameter throughout rather
than living on the struct, because the generated AppWindow type is not
Clone; every other field is an owned clone so each function's body could
be pasted from run() unchanged.

lib.rs::run is now under 300 lines: construction and startup only, calling
named functions for the window and its diagnostics/inference wiring, the
launch/library/collections/identity/settings screens, import and merge,
the render loop (render_now/redraw/show), the remote open path, the
window chrome (resize, layout class, panel toggles, back gesture), and
develop_ui::wire for the rest. Every TRACES and GESTURE comment moved with
the code it annotates.

Two latent type errors surfaced while restructuring rather than being
introduced by it: activity and display were being passed around as bare
ActivityLog/DisplayWatch instead of the Rc<...> their constructors
actually return, which only worked before because nothing needed to name
the type explicitly.
2026-09-20 19:17:22 +02:00
dtourolle beb7a5eac0 Split launch_ui::wire into one function per section
The 225-line wire() registered the launch screen's callbacks in nine
comment-delimited sections. Lift them into fn's, merging a few adjacent
ones that were only a handful of lines each: wire_sign_in covers both the
browser flow and the app-password fallback (the same button in two forms),
wire_choose_folder_and_open covers opening the folder browser and the
final "Open library" press, since both are short and sit back to back.
wire_use_folder, wire_sign_out, wire_formats, wire_folder_picker_navigation
and wire_copy_url stay as they were sectioned. wire() calls each in the
original order and keeps the closing render() call, which every section
relies on having run once at startup.
2026-09-20 19:11:43 +02:00
dtourolle 262ed2553c Split merge_ui::wire into one function per section
The 312-line wire() registered the panorama page's callbacks in three
comment-delimited sections. Lift each into its own fn: wire_start (the
"Merge to panorama" press, its fetch, and the DARKROOM_START_MERGE dev
entry point — all three only ever used together), wire_decision (confirm,
projection and border chips, the fill knobs), and wire_stop_and_leave
(abandon, close). wire() itself now just calls the three in order; S, C
and F stay generic on wire_start since sources/context/on_done are used
nowhere else.
2026-09-20 19:00:31 +02:00
dtourolle 8d08ffd7b7 Split identity_ui::wire into one function per feature
The 837-line wire() had almost no section comments, unlike its siblings, so
the seams had to be found by reading it rather than following markers. Lift
each into its own fn: wire_dials (the two grouping sliders), wire_navigation
(open/close/switch person — needs models() too, for the missing-model
banner), wire_rename_and_merge (a rename and the namesake offer it can
raise), wire_face_actions (pick/confirm/reject/split, the grid's own
actions), wire_grouping_preview, wire_recluster, wire_indexing (the shared
launcher behind Index/Re-index plus Stop), and wire_coverage_and_ignore.
wire() keeps the generic-to-trait-object coercions and the eyes_available
closure, since most of the above need it, and calls each function in the
original order. The reload! macro moved from inside wire() to module scope,
dedented, since macro_rules is scoped textually and every extracted function
uses it.
2026-09-20 18:49:43 +02:00
dtourolle 8540518022 Split collections_ui::wire into one function per section
The 1,457-line wire() registered every collections-sidebar and grid-drag
callback in one function, sectioned only by comment. Lift each section
into its own fn: wire_selection (split further into wire_selection and
wire_selection_filing, since the original section ran to 375 lines),
wire_drag, wire_trash, wire_trash_from_grid, wire_tree_navigation,
wire_create, wire_rename, wire_remove, wire_row_drag (the tree row's own
hold-drag, which the original "remove" comment's span covered but which
is really a separate feature), wire_row_menu, and wire_membership.
wire() itself now just coerces the shared closures to trait objects and
calls each in the original order. visible_ids is coerced to
Rc<dyn Fn() -> Vec<ImageId>> at the top, alongside on_scope_changed and
session, so the new functions take plain trait objects instead of
threading a generic parameter through every one of them.
2026-09-20 18:36:18 +02:00
dtourolle b952f5976a Merge: the wire sections lifted into named functions, beside the module splits 2026-09-20 18:26:24 +02:00
dtourolle a1d511fd4b Split library.rs into library/ by area of behaviour
library.rs was 7,729 lines wiring together everything "open a remote
library" touches: scanning, pulling other devices' judgements out of
sidecars found along the way, writing local edits back out to the
sidecar outbox, pushing/reloading XMP by hand, fetching and prefetching
thumbnails and originals, generating thumbnails locally, the metadata
and thumbnail background sweeps, on-disk paths for the catalog and
model files, and reading the grid's cells, spans and rating filter.
Same motivation as the develop.rs split (docs/dev/code-health.md CH-1):
a pure, no-behaviour-change move into one file per area, each under
about 1,500 lines.

Tracing actual call sites rather than trusting the file's physical
layout mattered here: `persist`, `load_folder_etags`, `pull_sidecars`,
`load_sidecar_etags`, `record_sidecar_read` and `apply_judgement` sit
textually beside the XMP push/reload functions but are called only
from `run_scan` (pulling a device's own past judgements out of the
sidecars a scan just walked), so they went to scan.rs and not xmp.rs.
`cells` came out at over 1,800 lines once its tests moved with it and
split further into cells.rs (windowed reads, trash, ordinals) and
spans.rs (collection scope, manual reordering, the capture-time
histogram) -- ten submodules rather than the nine first planned.

Previously-private items reached from a sibling module became
`pub(super)`, narrower than the whole-crate reachability one file gave
them. Tests moved with the code they test; the two test fixtures used
across more than one file (`scanned`, and develop.rs's
`session_with_a_left_half_subject` in the matching commit) joined the
shared `test_support` module alongside the existing `entry`/
`with_images`/`image_ids` helpers. `mod.rs` re-exports every module's
public items under `library::`, including the `pub(crate)`
`test_support` module `repairs.rs` reads its fixtures from, so no file
outside `library` needed a change.

The previous commit split develop.rs the same way; taken alone it left
dr-ui without library.rs, so that intermediate commit does not build on
its own. This one restores it.
2026-09-20 18:21:43 +02:00
dtourolle 050c2c9d16 Split develop.rs into develop/ by area of behaviour
develop.rs had grown to 9,327 lines covering everything the develop
session does: opening a photograph, the parameter-row and curve-widget
panel model, mask viewing and editing, mask creation and the rasteriser
that turns a mask stack into GPU arrays, spot repairs, scene
segmentation, framing and zoom, white-balance sampling, rendering and
film choice, and the undo/snapshot history. docs/dev/code-health.md
CH-1 names dr-ui's lack of a view layer as the reason every feature
kept landing in a handful of files; this is the first of the two pure
splits it recommends as easy, no-behaviour-change wins independent of
that larger rework.

The boundaries follow the file's own sections (several were already
marked off with comment headers) and the seams a full read turned up
underneath them -- mask storage/rasterisation turned out to be a
distinct concern from mask viewing and editing, and rows/tabs/curves
from each other, so those split further than the headers alone
suggested. Each module stays under about 1,500 lines. Struct fields
and the handful of helper methods now called from a sibling module
became `pub(super)`, which is strictly narrower than the whole-crate
reachability a single file gave them; nothing gained visibility outside
`develop`. Tests moved with the code they test, including the few
cases where a helper one file's tests needed was itself only defined
in another's -- those became shared fixtures in `mod.rs` alongside the
`headless`/`read_back`/`grey_session` helpers that already worked that
way. `mod.rs` re-exports every item `develop::` callers outside this
module used before, so lib.rs, masks_ui.rs and the rest needed no
changes.
2026-09-20 18:21:26 +02:00
dtourolle bd3b993b90 Split library_ui::wire into one function per section
wire() registered every grid callback in one 1,214-line function behind
four section comments, two of which were themselves far over 300 lines
with no further markers. Each fenced section becomes its own function,
called from wire() in the original order with the section's own comment
kept as its doc comment:

- "the keyboard cursor (FR-CULL-4)" (496 lines) splits at its own topic
  breaks into wire_grid_cursor_and_zoom (cursor movement, cell zoom and
  pinch), wire_grid_sync_and_load (explicit sync/thumbnail requests and
  the reloads a changed viewport, column count or scroll position
  trigger), wire_timeline (the capture-time sidebar) and wire_grid_routes
  (grid/launch/develop navigation and a manual rescan).
- "ratings and flags" and "keywords" were already under 300 lines and
  become one function each.
- "the filter bar" (447 lines) splits into wire_filter_ratings_and_people
  (which keeps the section's own comment), wire_filter_dates and
  wire_filter_scope_and_offline.

The three callbacks registered before the first section comment
(on_settings_xmp_reload, on_library_cell_clicked, on_library_roll_pick)
and the trailing crate::recovery_ui::wire call stay directly in wire(),
since neither is inside a fenced section.
2026-09-20 17:25:00 +02:00
dtourolle 59605f9fbb Split masks_ui::wire into one function per section
wire() registered every mask-panel callback in one 778-line function
behind section comments. Each of the seven fenced sections (computing
the region map, refining a subject's mask, dragging a gradient,
selecting on the photograph, the stack, the edge treatment, adding
layers) becomes its own function, called from wire() in the original
order with the section's own comment kept as its doc comment. The
"adding layers" section was itself over 300 lines and had no further
section markers inside it, so it is split at its own natural seam
between painting/viewing a mask (wire_layers_paint) and working the
parts and add-mask buttons (wire_layers_parts); the second half gets an
introductory doc line since there was no comment of its own to reuse.
Locals declared just for one section's closures (running, refining,
dragging) move into that section's function instead of staying in
wire().
2026-09-20 17:24:42 +02:00
dtourolle 04949741c1 Split settings_ui::wire into one function per section
wire() registered every settings-page callback in one 416-line function,
fenced only by section comments. Each fenced section (opening and
closing, cache, faces, export, reset) is now its own private function
that wire() calls in the same order, with the section's own comment kept
as its doc comment. on_budget_changed and on_open are coerced to trait
objects at the top of wire() so the new functions take a plain
Rc<dyn Fn> rather than needing their own generic parameter, with no
change in the closures registered or the order they are registered in.
2026-09-20 17:24:28 +02:00
dtourolle 5b4ad11853 Manual: nested collections, and the ghost drawn as it should be
Benchmarks / CPU and I/O (per commit) (push) Failing after 6m16s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 42s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 2s
Build and test / Layer separation (push) Successful in 26s
Traceability / Requirement traces (push) Failing after 38s
Build and test / Android (aarch64) (push) Failing after 2m20s
Build and test / Windows (x86_64, cross) (push) Failing after 3m3s
A scene that makes a parent, nests two collections in it by drag and by
the menu, files frames into a child and opens the parent to see it count
both; stills of the tree and of the menu. The collections recording is
re-made now that the bitmap under the cursor is the photograph.

drive.py grows a multi-leg drag: a diagonal with much vertical in it is
taken by the grid's Flickable as a scroll before the DragArea can claim
it, so a drag to the sidebar goes sideways first.
2026-09-20 16:27:18 +02:00
dtourolle 6b1aac477d Put the developer docs under docs/dev and index the folder for users first
docs/ had 26 developer documents flat beside the manual, and the two
audiences are very differently sized: most readers want the manual and
the gesture reference, a few want the register, the designs and the
measurements. The manual and gestures.md stay at the top; everything for
someone changing the code moves to docs/dev/, and the two documents that
name their own successors — the v0.1 milestone and the UI-refinement plan
— go to docs/dev/archive/ rather than being deleted, since both are still
cited. docs/README.md is the index, users first.

Every reference follows: code comments, Cargo manifests, the workflows,
the pre-commit hook, the bench and traceability tools (which locate the
repo root by docs/dev/requirements.md now), packaging, the Docker READMEs,
CLAUDE.md, CONTRIBUTING.md and the README. The matrix links one level
deeper and is regenerated. Links out of the moved documents into the tree
gain a level; a link checker over every Markdown file finds none broken.
2026-09-20 16:20:15 +02:00
dtourolle 2afc2a7890 Report the grid's column count on creation, not only on change
`changed columns` fires on a change, and a first evaluation is not one:
a grid built after the window had settled at its size never said how
wide it was, so Rust placed month headings for the one column it was
told about at start-up — every month began a row, and was announced
wherever its first cell fell, mid-row included. Opening a collection
showed "October 2025" stranded over a row of August.
2026-09-20 15:58:44 +02:00
dtourolle 6507593715 Hand the drag ghost to the renderer through a file, so it draws
The bitmap under the cursor was a solid red rectangle. Slint's drag
overlay uploads the image as a texture, draws it and drops the texture in
one call; with the wgpu FemtoVG renderer the drop is immediate and the
draw is deferred to the flush, so the frame binds femtovg's placeholder —
which is red. An image with a cache key survives in the texture cache
until after the flush, and only a path gives one. So the composite goes
to the data directory's scratch as a PNG and comes back through
load_from_path; one file per drag, removed when the drag ends. A
workaround for Slint 1.17.1, written up as one beside the code.
2026-09-20 15:58:44 +02:00
dtourolle 08727cff5a Release 0.13.5
Benchmarks / CPU and I/O (per commit) (push) Failing after 6m16s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 45s
Build and test / Layer separation (push) Successful in 25s
Traceability / Requirement traces (push) Failing after 38s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 2s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Failing after 2m20s
Build and test / Windows (x86_64, cross) (push) Failing after 3m1s
2026-09-20 15:24:47 +02:00
dtourolle f4c3f425dd Show the picker correcting something in the manual
The recording sampled a red brick wall and moved the sliders by three
units, which at GIF size is a click that does nothing. The scene now
drags the frame cold first and picks a white air conditioner, so the
correction is visible and the picker's being absolute - set from the
photograph, not from where the sliders were - is what the picture
shows. The text says so, and says a blown highlight is refused.
2026-09-20 15:24:27 +02:00
dtourolle 12f8990e09 Average a patch under the white balance picker, not one photosite
The probe's comment said a 192px render "averages a small neighbourhood
into each of its pixels". It does not: the composed shader fetches the
source at one position per output pixel - nearest for an unrotated
frame, four photosites blended otherwise - so the probe was a point
sample of a noisy sensor, and two painted-white air conditioners on the
same wall answered +37 and -50.

The tap is now narrowed to the patch of the canvas around the click, a
couple of percent of its width and square on screen, and rendered at
64x64 with interpolation forced on, which puts a sample on every sensor
pixel under it at any ordinary zoom. The samples are averaged, with the
void and clipped ones left out rather than allowed to pull the mean, and
fewer than half surviving is refused. compose_camera_probe takes the
patch; the merge's compose_camera_linear keeps its nearest sampling. The
readback shrinks from six megabytes to sixty-four kilobytes.

A frame of alternating warm and cool columns, averaging neutral, moves
the controls by at most two units; a point sample swung them to sixty.
2026-09-20 15:24:25 +02:00
dtourolle 8a897bbc01 Release 0.13.4
Benchmarks / CPU and I/O (per commit) (push) Failing after 6m22s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 46s
Build and test / Layer separation (push) Successful in 29s
Traceability / Requirement traces (push) Failing after 40s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Failing after 2m22s
Build and test / Windows (x86_64, cross) (push) Failing after 3m5s
2026-09-20 14:03:58 +02:00
dtourolle a03e082fe2 Point the manual's white balance scene at "pick", and record it again
The scene clicked 40px to the right of "pick", on "reset", so the
recording showed a neutral group being reset and a click on the wall
that panned. Re-recorded with the picker fixed: the word lights, the
sample moves temperature and tint, and Before shows what it corrected.
2026-09-20 13:41:17 +02:00
dtourolle 4576499c3b Refuse a clipped highlight as a neutral
Sampling the overcast sky on a Canon 6D frame set tint to -100 and
temperature to -15 for a patch the canvas showed as pure white. A clipped
photosite is sensor white, not a colour: every channel stopped counting,
so what the tap hands back is the as-shot multipliers themselves, which
are strongly magenta, and the solver dutifully drove green to its stop.
The display shader already fades such a pixel to a neutral of the same
brightness before any operation runs, so the picker was balancing against
something the photographer could not see.

The probe now refuses a sample with any channel at or above the onset the
shader fades from, the way the solver already refuses black. The threshold
is one constant, CLIP_ONSET, formatted into the shader and read by the
probe, so the two cannot drift apart.
2026-09-20 13:41:17 +02:00
dtourolle 2f47087223 Measure the white balance probe in camera RGB, where the gains multiply
Pressing "pick" and clicking a near-neutral wall on a Canon 6D frame set
tint to -77 and turned the whole photograph green. The white balance
operation runs first in the chain, on camera RGB, before the body's base
curve and colour matrix; the probe was read off a display render after
all three, and the solve treated that sRGB triple as if the gains
multiplied it directly. On a JPEG the two spaces coincide, which is why
the existing tests passed while the picker was broken on every raw file.

The probe now reads the camera-space tap a merge stitches from, composed
under the edit's own framing so a fraction of the canvas is a fraction of
the probe, and puts the as-shot balance on itself - exactly the value the
operation's gains are about to multiply. No operations run in the tap, so
nothing has to be stripped and restored, and the display target is left
alone, so a sample that found nothing usable no longer needs a redraw.

A raw-frame test with the 6D's matrix and a typical as-shot balance
samples a warm grey and asserts the rendered pixel comes back neutral; it
fails on the previous probe.
2026-09-20 13:41:11 +02:00
dtourolle 764ad55ead Stand each person in the grouping pass by at most 100 references
Benchmarks / CPU and I/O (per commit) (push) Failing after 6m23s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 55s
Build and test / Layer separation (push) Successful in 27s
Traceability / Requirement traces (push) Failing after 54s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Failing after 2m21s
Build and test / Windows (x86_64, cross) (push) Failing after 3m5s
Every face the user has ruled on entered the pass as an anchor, and the
scan is exhaustive by design (`dr_face::neighbours`), so a person with
750 confirmed faces cost 750 comparisons against every other face in
the library — and the cost of a library grew with how well it was
named. Most of those comparisons said nothing new: thirty frames from
one afternoon are one point of view, not thirty, and a face that
matches one of them matches the rest.

Each person now enters through at most 100 of their anchored faces
(`dr_face::references`). Eligible are those whose raw embedding is at
least 15 long — one above the gallery floor, since a reference speaks
for someone rather than merely being admitted — with an unmeasured
length admitted as it is everywhere else. From those, the set spanning
the greatest volume is chosen greedily: the longest vector first, then
at each step the face with the largest component orthogonal to the
chosen so far. That is pivoted Gram–Schmidt, and the product of the
residuals it picks is the Gram determinant, so the greedy step is the
exact greedy on the objective. A near-duplicate of a chosen face has
no residual and is passed over; the one profile shot among two hundred
frontal frames is taken early; faces inside the span of the chosen add
no volume and are not taken to fill the cap.

The faces not chosen keep their confirmations and are not touched by
the pass — they stay in the anchor map, so it never releases them —
they are simply not compared. A person none of whose faces is long
enough is still stood for, by their longest, rather than losing their
anchor and having their next face filed as a stranger. Under the cap
nothing changes: every eligible face stands, and the short ones stay
in as the probes they were.

At the reference library's 3,851 confirmations the scan shrinks by
about a fifth; at 15,000 it is a fifth of what it was.
2026-09-20 13:29:16 +02:00
dtourolle 301e6f3828 Release 0.13.3
Benchmarks / CPU and I/O (per commit) (push) Failing after 6m21s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 46s
Build and test / Layer separation (push) Successful in 26s
Traceability / Requirement traces (push) Failing after 39s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Failing after 2m21s
Build and test / Windows (x86_64, cross) (push) Failing after 3m8s
2026-09-20 13:01:41 +02:00
dtourolle a437363bd6 Schema V20: put the mis-spelled run markers right
The markers the previous commit stops writing are already in the
catalogs — 2 on the desktop, 429 on the tablet — and in the shards
both have exchanged. Renaming them to the faces' own id with a fresh
time is what makes the export send each image again, under an entry
newer than the empty one `held_model` would otherwise pick. Where the
old write had inserted its marker beside the right one, the wrong one
goes and the right one is refreshed for the same reason: its entry in
the shards is older than the empty one.

Images V14 left with faces and no marker are not touched. That state is
the quality pass's cue, and the fixed write marks them correctly when
it reaches them.

Checked against copies of both real catalogs: the desktop renames 2,
the tablet deletes 429, both in under 200 ms.
2026-09-20 13:00:59 +02:00
dtourolle 065872bec5 Keep a run marker under the detector that found the faces
`record_updates` — the write behind the quality, eye and crop passes —
re-marked the image as indexed under the pipeline the pass ran as, and
left the faces it had updated under the id of the detector that found
them. On a desktop set to Thorough that put `scrfd_10g+w600k_mbf` over
faces spelled `w600k_mbf`; on the tablet, `scrfd_10g_i8+w600k_mbf` over
faces it had adopted from the desktop's thorough pass.

Every reader takes the marker and the faces to agree. `marker_under`
reads the marker as the detector having examined the image, so the
upgrade repair never revisits it. The shard store keys each face by
its pipeline id, so `export_to_shards` selects an image's faces by the
marker's id, finds none, and sends an entry that says the thorough
detector looked and found nothing — over photographs with named faces
on them. The desktop's shard index holds 54 such entries beside real
faces; the tablet's eye pass over the faces it had adopted made 430
more, and both devices have exchanged them. `held_model` takes the
newest entry for an image, which is the empty one. Nothing has been
lost yet only because the two spellings of the thorough detector rank
equal and neither side adopts the other's; a third device, or either
one after a reinstall, would adopt "nothing here" for 484 images. And
the desktop's eye pass is 4,739 images from doing the same to every
face from before V14 — which are the ones that only exist on the
desktop, and would then never reach anywhere.

The marker now takes the id the faces carry; the pass's own id is used
only when it dropped the last of them and there is no detector left to
name. A stale marker under another spelling of the same embedder is
removed in the same transaction, so one embedder has one marker.
2026-09-20 13:00:58 +02:00
dtourolle 0ed38ada28 Adopt a peer's unmeasured faces instead of refusing them
The tablet showed a fraction of each person: 681 of the desktop's 3,851
confirmations, and none of Ian's 746, Catherine's 626 or my own 480.
Every face that existed on both devices agreed on who it was, and the
people rows were identical — the merge was fine. The missing 3,170
confirmations were on faces the tablet did not hold at all: the
desktop's 16,080 faces from the original detector, on 4,310 images,
detected before schema V14 kept the quality reading.

Those faces were in shards the tablet had already downloaded, in
August's export. `import_from_shards` looked at them on every sync pass
and declined each one, because a face without a quality reading was
"work this device cannot finish": adopting it would write the run
marker, and the marker was what stopped an image being looked at again.
That was true when it was written and has not been since the quality
repair existed — that pass lists its work by `f.quality IS NULL`, not by
the marker, exactly as the eye pass does, and faces without an eye
reading were already adopted on that reasoning.

The refusal had no exit. V14 had deleted the markers of every image
holding such faces so the quality pass would find them, and
`export_to_shards` walks the markers, so the desktop never re-exported
them either; the unmeasured August copies were the only ones there
would ever be. The tablet's answer was to queue all 17,727 images for a
re-detection of its own, a fetch of the whole library, while holding
the faces on disk.

Adopt them. The receiving device's quality pass measures them when it
reaches them, and the desktop's confirmations match onto them by box
overlap on the next catalog merge. The test that asserted the refusal
now asserts the adoption and that the image is still owed to the pass.
2026-09-20 12:58:41 +02:00
dtourolle 695d5ec304 Correct four claims in the README against the tree
Benchmarks / CPU and I/O (per commit) (push) Failing after 6m26s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 46s
Build and test / Layer separation (push) Successful in 28s
Traceability / Requirement traces (push) Failing after 40s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 2s
🐳 Windows image / Build and push (push) Successful in 2s
Build and test / windows-image (push) Successful in 2s
Build and test / Android (aarch64) (push) Failing after 2m21s
Build and test / Windows (x86_64, cross) (push) Failing after 3m5s
Eighteen declared operations, not fifteen; JPEG XL is an export format;
the grid does not filter by keyword, only the catalog's query can; and a
panorama's provenance is a sidecar beside the composite, not a history
step in it.
2026-09-20 11:57:17 +02:00
dtourolle ef1afc254d Release 0.13.2
Benchmarks / CPU and I/O (per commit) (push) Failing after 6m36s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 46s
Build and test / Layer separation (push) Successful in 34s
Traceability / Requirement traces (push) Failing after 40s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 2s
🐳 Windows image / Build and push (push) Successful in 2s
Build and test / windows-image (push) Successful in 2s
Build and test / Android (aarch64) (push) Failing after 2m26s
Build and test / Windows (x86_64, cross) (push) Failing after 3m7s
2026-09-20 11:06:54 +02:00
dtourolle 103c6e7fc0 Rewrite the README for someone arriving, not someone already here
Benchmarks / CPU and I/O (per commit) (push) Failing after 30s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 53s
Build and test / Layer separation (push) Successful in 28s
Traceability / Requirement traces (push) Failing after 41s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Failing after 2m25s
Build and test / Windows (x86_64, cross) (push) Failing after 3m11s
It said 0.9.0 against a 0.13.1 tree, listed focus peaking and burst
grouping as unbuilt when both have shipped, and opened with a page of
prose about the display path before saying what the application does.
Lead with what it is and a picture of it, say how to get it on each
platform and what state each channel is in, keep the honest account of
what is missing, and put the manual first in the documentation table.
2026-09-20 11:06:10 +02:00
dtourolle e9398de9c1 Ship the fine-tuned border filler: MI-GAN 512 trained on projection-shaped voids from the maintainer's library 2026-09-20 10:56:38 +02:00
dtourolle b1c99b5796 Let the fill show the model an open void: mirror depth 0 means no ring and nothing known beyond the band 2026-09-20 10:56:38 +02:00
dtourolle 3de109fbd1 Write down the catalog and screen-refresh patterns the Identity fixes exposed
A CLAUDE.md at the root, for anyone changing this code: the ways a redraw
and a catalog read came to cost half a second per click, what each fix
looked like, and how to measure the next one against a copy of a real
catalog.
2026-09-20 10:56:37 +02:00
dtourolle 388bda6af3 Count the outstanding repairs from the faces, on partial indexes
"How many images still owe a quality reading" was a correlated EXISTS per
image over `faces`, and the face row is 8 KB of embedding and crop before
the column it looks at, so each count opened every row. Six such counts
run on every open of the Identity screen and at the end of every sweep:
160 ms on the reference library.

V19 adds three partial indexes holding only the faces still owing each
pass, keyed on the image and carrying the model id the predicate reads,
and replaces `faces_image` with `(image_id, model_id)` so "does this image
hold this embedder's faces" is answered from the index too. The planner
takes a partial index when the count is driven from `faces` and ignores it
inside the EXISTS, so `Needs::Face` carries the per-face fragment and
`repairs::count` spells the query from the faces' side; the list and the
per-image check keep the EXISTS. A test holds the two spellings to the
same answer for every repair.
2026-09-20 10:56:37 +02:00
dtourolle 73845d8a77 Ask the server about a collection once per job, not once per file
Every `move_to` guaranteed its destination's parent with a `MKCOL` for
each ancestor down from the account root, and a trash folder under a
library root several levels deep meant three round trips answering
`405 Method Not Allowed` before the one `MOVE` that did anything — for
every image of a delete, on a connection built for that job.

The backend now records the collections it has confirmed exist and asks
about each once. It lives for one job, so a folder another client removes
mid-batch is the one case this misses, and the `MOVE` then reports the
`409` rather than hiding it.
2026-09-20 10:56:37 +02:00
dtourolle b4821ee1ab Filter the people rail in the query, and count the unassigned faces
`faces::people` grouped `face_person` after a LEFT JOIN over every person
and sorted the lot by name; the rail then discarded the empty, unnamed
groups a regrouping pass leaves behind — 17,000 of 19,000 rows on the
reference library. `people_in_use` filters them in the WHERE and joins
`people` to face counts aggregated first (2,000 groups), so the sort sees
only the rows that will be drawn. `count_unassigned` replaces fetching
2,400 ids to take their length. `load_people` 22 ms → 10 ms.
2026-09-20 10:56:36 +02:00
dtourolle 9d1aa5735b Confirm a group, and split one, in one transaction
`confirm_all` called `faces::confirm` per face, and `split_off` called
`reject` then `confirm` per face: each opens and commits its own
transaction, so a click on a group of several hundred was several hundred
commits. `faces::confirm_all` is two statements — clear the rejections the
confirmations override, then flip the rows — and `faces::reassign` does a
split's reject-and-confirm for every face under one commit. 16 ms → 2 ms
and 22 ms → 4 ms on the largest group.
2026-09-20 10:56:36 +02:00
dtourolle 97a854833d Reuse the face grid's decoded crops across a redraw
A confirm or a reject changes one row and redraws the whole grid, and the
redraw re-read every crop blob of the selected person (4 MB for the
largest) and decoded every one — 316 ms per click on the reference
library's 754-face person, to arrive at the pixels already on screen.

`load_faces` now takes the crops the previous load decoded, keyed by face,
and moves each into its new cell; the blob read is skipped when every face
is already in hand. `refresh` drains the old cells into it rather than
cloning them. The redraw is 2.6 ms.
2026-09-20 10:56:36 +02:00
dtourolle e0e193efb4 Do not recount face coverage on every confirm, and count it without listing
Every click on the Identity screen's face grid — confirm, reject, split,
rename, merge — redrew the whole screen, and the redraw recomputed the
coverage line. That line lists every repair's outstanding images to count
them: six scans of the images table with a correlated EXISTS over the
8 KB face rows, an ORDER BY the job's visiting order, a Target with its
path per row, and a thumbnail-index query per image with faces. On the
reference library (24k images, 19k faces) that was ~200 ms of the
~540 ms each click cost, spent computing a figure a confirm cannot change.

`refresh` now takes what changed: `Changed::Identities` re-reads the rail
and the grid and leaves the coverage line alone; `Changed::Library` — an
open, a sweep ending or stopped, the face data deleted — re-reads it too.

For the times it does run, `repairs::counts` counts instead of building
and dropping the lists, and the thumbnail store is read once
(`ThumbStore::held`) rather than probed once per image in the audit, the
outstanding list and the proxy repair.

`identity_bench` is the measurement: the reads a click performs and the
batch writes, timed against a copy of a real catalog.
2026-09-20 10:56:36 +02:00
dtourolle d790961b28 Add the manual: every feature pictured from the application itself
Benchmarks / CPU and I/O (per commit) (push) Failing after 30s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 47s
Build and test / Layer separation (push) Successful in 27s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 4s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 1s
Traceability / Requirement traces (push) Failing after 38s
Build and test / Android (aarch64) (push) Failing after 2m24s
Build and test / Windows (x86_64, cross) (push) Failing after 3m21s
docs/manual/README.md is a tour for a photographer opening DarkRoom for
the first time — one picture per thing, moving where movement is the
point. tools/manual/ is how the pictures are made: drive.py puppeteers the
desktop build on a private Xvfb (launch, click, drag, type, screenshot,
record), scenes.py is each picture as a script, and record.sh runs them
all over a folder and writes the results into docs/manual/media/.

The media is in LFS, with the CI pulls excluding it as they exclude the
fixtures; a screenshot changes wholesale when the interface does.

Nothing in the pictures shows a person, by design: the demo library is
seventy urban and alpine frames, chosen from the catalog's rows that face
detection found nobody in.

The traceability matrix is regenerated here after the rebase that
brought this branch up to master.
2026-09-20 00:26:00 +02:00
dtourolle 14ac41bee0 Give the keyword sheet's field focus, and the empty trash its own words
Without focus the keys typed into the keywords sheet went to the grid
behind the scrim, and the Return meant for the keyword opened a
photograph. The naming sheet already takes focus on open; do the same.

An empty trash said "No images found — check the library folder and which
formats are ticked", which sends someone off to fix a library that is
fine.
2026-09-20 00:21:14 +02:00
dtourolle 86410def88 Title the gesture book for the whole application, not the grid 2026-09-20 00:21:14 +02:00
dtourolle df8be10c7d Stop promising to ask for an export folder
An empty device destination read "Ask each time" on the settings page,
and nothing asks: an export made with the field blank is refused with
"no export folder is set". Say what will happen.
2026-09-20 00:21:14 +02:00
dtourolle f176043632 Report a merge worker that dies rather than leaving the page on Stop
wgpu reports a device out of memory by panicking, and a twelve-frame
merge on a GPU another process is using is where that happens. The panic
unwound the worker, the sender went with it, and the page sat on "Stop"
with every control disabled and nothing to say why — the crash record on
disk was the only sign. Catch the panic and send it as a failure, and
treat a closed channel with no final event as a dead worker too.
2026-09-20 00:21:14 +02:00
dtourolle 9c2cd73337 Show a mask's tint only while masking
The eyes are per layer and outlive the mode, so a photographer coming
back finds the layers they were looking at still lit. But the tint is a
way of looking at a mask, and outside Local there is no mask being looked
at: the sky stayed red through Repair and back in Photo, a mode that had
been left leaving its overlay behind — the fault ui-navigation.md D-N1
exists to prevent.
2026-09-20 00:21:14 +02:00
dtourolle e3acbdf4a3 Wrap the falloff and edge chips so a category mask cannot widen the column
Five chips in one row declare 440px, and the develop column takes the
widest panel's request — so selecting a category mask levered the sidebar
past the window's edge, clipping the histogram, the group strip and the
subject list. The same trap ChipGrid's comment records for film formats.
2026-09-20 00:21:14 +02:00
dtourolle eddaa44cd3 Announce a month at the next row it opens, not only if it begins one
A heading was drawn only on a cell that both began a month and began a
row, so at seven columns most months were never named, and the one
heading on screen — always on the window's first cell — was wrong about
every row below it. Worse, two headings drawn on the same cell overprinted
each other. Now a row carries a heading whenever its first cell's month is
not the one last announced: a month starting mid-row is named on the next
row it opens, one row late and right about everything under it.
2026-09-20 00:21:14 +02:00
dtourolle 7fba28f7d8 Upload a snapshot of a thumbnail shard, never the live file
Every shard is in WAL mode and every put opens its own connection, so
while thumbnails are being generated on several threads — which is when
the first sync pass runs — the log is never checkpointed and the main
file holds whatever the last quiet moment left in it. For a shard created
seconds earlier that is nothing: a zero-byte file with the schema still
in the log. The sync read that file and uploaded it, and every other
device merging it failed with "no such table: thumbs" on every pass.

Copy the shard through SQLite's backup API into scratch first, which
serialises against writers and carries the log, and upload that.
2026-09-20 00:21:14 +02:00
dtourolle 0fa9003e54 Let the top level be chosen as the library root
Confirming "/" in the folder picker set an empty root, which the launch
model read as no root at all: "Open library" stayed disabled after the
question had plainly been answered, and a folder library — whose folder
is the whole library — could never be opened without first descending
into a subfolder of it. The empty string was carrying two meanings.

Record the choice as its own fact on the account (`root_chosen`, defaulted
so existing configuration loads unchanged), treat a folder endpoint as
chosen by definition, and let the launch screen say so: a folder is shown
as a LIBRARY rather than an ACCOUNT, the second question becomes an
optional "scan only a subfolder", and the library header names the folder
instead of calling it "· whole account".
2026-09-20 00:21:13 +02:00
dtourolle e750bcdb8c Date the composite at the mean of its frames' capture times
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 31m52s
Build and test / Layer separation (push) Successful in 47s
Traceability / Requirement traces (push) Failing after 55s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 1s
Benchmarks / CPU and I/O (per commit) (push) Failing after 32s
Build and test / Android (aarch64) (push) Failing after 2m37s
Build and test / Windows (x86_64, cross) (push) Failing after 3m10s
It carried the first frame's, so it sorted before the sweep it was made
from. The middle of the sweep puts it among them.
2026-09-19 22:36:09 +02:00
dtourolle 48c4b403d2 Date a DNG whose IFDs follow its pixels: read the head and the tail
The scan reads the first 256 KB of a file for its metadata. A camera
writes its IFDs at the front, so that is the whole structure; the linear
DNG a merge writes puts its first IFD after the pixels, and rawler,
given the head alone, finds no decoder in it. The composite was
catalogued without a date and sorted to the very end of the grid, after
every dated photograph — which is where a panorama merged on the tablet
went unfound.

dr-decode's own TIFF reader now reads through a head and a tail at a
known offset; trailing_ifd says where the tail starts and
metadata_split reads the two together. The scan, when the head fails
and points beyond itself, fetches from the IFD to the end — kilobytes —
and dates the file from both. Tested against the writer's own output.
2026-09-19 22:35:58 +02:00
dtourolle 9d04ff2154 Release 0.13.1
Benchmarks / CPU and I/O (per commit) (push) Successful in 12m24s
Benchmarks / Frame budget (on demand) (push) Skipped
🐳 Android image / Build and push (push) Successful in 4s
Build and test / android-image (push) Successful in 5s
Build and test / Desktop (Linux) (push) Failing after 31m25s
🐳 Windows image / Build and push (push) Successful in 4s
Build and test / windows-image (push) Successful in 4s
Build and test / Layer separation (push) Successful in 46s
Build and test / Android (aarch64) (push) Failing after 6h25m37s
Build and test / Windows (x86_64, cross) (push) Failing after 3m17s
Traceability / Requirement traces (push) Failing after 48s
2026-09-19 22:06:06 +02:00
dtourolle c6cfb2a02a Put the -1 on the greens along the chroma axis, not across it
The Malvar "R at green in R row" kernel weights the two greens two
sites away along the row at -1 and the pair up and down the column at
+1/2. The shader had the two swapped, in the comment as well as the
code, so the transcription checked against itself. Both sum to zero
and reconstruct a flat patch exactly, which is all the tests fed it.

On an edge the correction at green sites is half strength and the
false colour doubles: 0.375 against 0.19 on a grey step, and a
blue/yellow zipper around every clipped highlight at 1:1. The other
three kernels and the CFA tables were right.

A grey vertical step now runs through the pass; the transposed kernel
fails it at 0.375.
2026-09-19 22:05:39 +02:00
dtourolle b83f192847 Package release 2 of 0.13.0: the inference engine and the user runtime directory 2026-09-19 21:20:53 +02:00
dtourolle ecb648818b Search the user's own runtime directory before the system library
The reference desktop's only system ONNX Runtime is Arch's
onnxruntime-opt-cuda: 1.29, built without TensorRT and against cuDNN 8
on a cuDNN 9 machine. The probe rejects both providers correctly and
the app runs on the CPU provider, which is right and not what anyone
wants. runtime/ beside the models is now searched ahead of /usr/lib,
tools/fetch-desktop-runtime.sh fills it with the four libraries from
the current onnxruntime-gpu wheel (cuDNN 9, TensorRT 10), and the
About caption lists every rung that lost and why, not only the first.
Verified: the app selects TensorRT from that directory with no
environment variable set.
2026-09-19 21:15:55 +02:00
dtourolle 5fbf8944d7 Count the filler in the APK's bundled-model array
Benchmarks / CPU and I/O (per commit) (push) Successful in 3m35s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 31m59s
Build and test / Layer separation (push) Successful in 40s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 2s
🐳 Windows image / Build and push (push) Successful in 2s
Build and test / windows-image (push) Successful in 2s
Traceability / Requirement traces (push) Failing after 57s
Build and test / Android (aarch64) (push) Failing after 54m6s
Build and test / Windows (x86_64, cross) (push) Failing after 1h4m55s
The unpack list gained migan-512.onnx without its length following;
nothing on the desktop compiles that crate, and the first Android build
of 0.13.0 stopped there.
2026-09-19 20:55:38 +02:00
dtourolle b502a8ef90 Release 0.13.0
Benchmarks / CPU and I/O (per commit) (push) Successful in 12m2s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Successful in 1h41m17s
Build and test / Layer separation (push) Successful in 50s
Traceability / Requirement traces (push) Failing after 1m6s
🐳 Android image / Build and push (push) Successful in 16m7s
Build and test / android-image (push) Successful in 16m9s
🐳 Windows image / Build and push (push) Successful in 6m11s
Build and test / windows-image (push) Successful in 6m13s
Build and test / Android (aarch64) (push) Failing after 40m26s
Build and test / Windows (x86_64, cross) (push) Failing after 1h4m8s
2026-09-19 20:43:27 +02:00
dtourolle 6fbdb1d06f Regenerate the traceability matrix and gesture book after the rebase 2026-09-19 20:41:44 +02:00
dtourolle d04b8f6044 FR-MRG-4: the border is cropped or filled, the fill experimental; panorama.md §13 records what was built and measured 2026-09-19 20:41:22 +02:00
dtourolle 8baa46ff49 Let the merge example fill, wait for engines and dump the filler's input; add a fill example that re-runs it stage by stage
A fill that went wrong took a seven-minute merge to look at again. Now
DR_FILL_DUMP=dir makes the merge write what the filler was given, and
the fill example runs fill_border on that, or a crop of it, on the engine
and writes coarse, each band and the feathered result as PPMs — seconds
per attempt on TensorRT. Both examples take DARKROOM_ORT_DIR as the app
does, and --wait-engines lets a compiling rung finish before timing.
2026-09-19 20:41:22 +02:00
dtourolle e43ae10439 Offer the border fill on the merge page, experimental, with every knob on it
A Border choice beside the projection — crop to the picture, or fill it
— that redraws the preview filled so the invented pixels are seen before
they are confirmed (FR-MRG-1), greyed with the reason when the model is
not there. The job fills at half the composite's resolution in a
display-ish space (white balance, matrix, gamma; invertible) and samples
the result back into the linear DNG wherever no frame reached; the
sidecar's merge line says border filled and with which knobs.

Experimental because the fill is right in thin borders and wrong in deep
corners, where the model's Places2 prior puts clouds in sky and water
under grass; so its six knobs — working scale, edge erosion, coarse pass,
band width, mirror depth, seam feather — are sliders under the choice,
each committing a redraw, until the defaults are right.
2026-09-19 20:41:22 +02:00
dtourolle 104e3a106f Fill a panorama's border with MI-GAN: mirrored context, coarse to fine, a feathered seam
dr_pano::fill owns everything the model does not — which tiles, what
context, how to blend — behind an Inpainter trait, and dr_pano::migan is
that trait over the shipped generator on the inference engine.

The known content is mirrored across the coverage edge into the hole and
a 256-px ring, the nearest 48 px folded, so the model interpolates between
real and mirrored sky rather than extrapolating into nothing. A coarse
pass at a quarter decides the structure with the whole border in a few
tiles; fine passes in 96-px bands from the edge outward texture it; the
seam is blended over a feather inside the real edge. Every knob is a
Params field, and an Observer hears each stage for whoever is looking at
why a fill went wrong.
2026-09-19 20:41:22 +02:00
dtourolle 031ba7b77d Hash a model's bytes once, at open, not on every acquire
A border fill acquires the filler once a tile, and each acquire hashed
the 28 MB model twice — 60 ms a tile, a third of the tile's run on a
throttled TensorRT. The Model keeps its hash from open.
2026-09-19 20:41:21 +02:00
dtourolle 54a80e688c Ship MI-GAN's bare 512 generator as the panorama border filler
Sargsyan et al., ICCV 2023; MIT code and weights (models/LICENCE.md),
exported by tools/export-migan.sh at a fixed 1×4×512×512 from the
authors' checkpoint — six operator types, 28 MB, in LFS like the rest.
The package installs it beside the scene model and the APK unpacks it
with the others.
2026-09-19 20:41:20 +02:00
dtourolle 2fd7690b6f Mark a pixel the lens correction pushed off the sensor with alpha 0 in the camera-space tap
The fused shader stored black with alpha 1 for a pixel whose source
coordinate left the frame, and the merge's warp averaged it in like any
other: a dark, badly interpolated fringe along every frame's edge, visible
as a seam wherever a frame ended and, later, as the edge the border fill
continued. The display keeps its opaque black; CameraLinear stores alpha 0
and the warp weights each sample by the alpha it interpolated, dropping a
sample that has none.
2026-09-19 20:41:20 +02:00
dtourolle c5f07f9ced Give the engine an Inpainter role for the panorama border filler
MI-GAN is plain convolutions, so every rung serves it and none needs a
special form; the role exists so resolve_model and the probe's fingerprint
know the model, and so the merge job can open it through the engine rather
than tract, which takes 7.4 s a tile for it.
2026-09-19 20:41:20 +02:00
dtourolle 5c00942b84 One completeness job over a registry of repairs, and a re-index button
A library's records are never all complete at once. A face found before
its quality was kept has no quality; one found before the eye models
existed has no reading; one adopted from a peer's shard has no crop; an
image the fast detector examined on a 1024 px proxy has boxes the current
detector would not have drawn; an image the scan stat'ed has no capture
date. On the reference library that is 17,762 faces under the bare
w600k_mbf id with no quality, no reading and no dense landmarks, 4,144 of
them without a crop, beside 12,217 images the fast detector examined and
found nothing in. Every one of those gaps was its own pass — V14's
measuring pass, §17.5's eye pass, the sweep's proxy repair, the sweep's
detector upgrade — with its own work list, its own count and its own idea
of done, and adding a per-face field meant adding a pass. There was no
pass at all for the case the library is actually in: boxes and landmarks
drawn by a weaker detector on a proxy, which every later per-face pass
would have read from.

dr_ui::repairs replaces them with one job over a registry. A Repair names
one thing a record can lack — the predicate that says which images still
owe it, the input its handler needs (a header, the original, or a native
render), the handler, and what to record for an image that can never be
done. The job unions the predicates into one work list, fetches each
image once at the most any claimant asks for, renders it at most once,
and runs every handler whose predicate that image still matches, checked
again before each because a detection writes every field a per-face
handler would fill. The registry today: face-proxy, face-quality,
face-eyes, face-crop, face-detection, face-upgrade, metadata — the last
there to say that this is not a face job. Adding a field is one entry.

A repair's predicate is the only definition of its work: the count the
settings page shows, the list the job fetches and the check before its
handler run are one predicate, so the job converges. That is why the
registry is cut to what the device can do rather than listing what it
skips — an entry is a count and a set of originals to fetch — and why an
eye reading that cannot be cut is not a criterion.

The catalog side is generic to match: record_updates writes whichever
fields a FaceUpdate carries and re-marks the image so the shards export
it; faces_needing and count_needing answer a predicate the caller
supplies, replacing the measuring pass's three special cases.

Two buttons on the settings page run the job and differ in one
predicate. "Index faces" converges on coverage: has anything examined
this image. "Re-index every face" converges on provenance: face-detection
claims every image with no marker under the chosen detector, in either
of its forms (FaceDetector::model_ids, so a desktop in f32 and a tablet
on the Hexagon do not re-index each other's work), and a marker saying a
weaker one looked is not that. An original over the fetch budget is left
exactly as it was under the re-index, where the sweep marks it examined:
a re-detection with nothing found would delete the faces, and "cannot
fetch" is not "no faces".
2026-09-19 18:52:13 +02:00
dtourolle 2a4ac0ed3d Carry every identity across a re-detection, by box and by embedding
record_detections replaces an image's faces and carried only the user's
confirmations onto the new ones, by box overlap above 0.5 IoU. Everything
else on the old faces was dropped: the suggestions the last grouping pass
made, and the people the user had said a face was not. On the reference
library that is 13,011 suggestions and 77 rejections beside 3,778
confirmations — a re-detection of it would have been correct by
FR-CULL-12's letter, since suggestions are derived data, and would have
handed back a People screen of strangers.

Now every old face is read before the delete — box, vector, assignment,
rejections — and matched to the new faces one-to-one, best pair first. A
pair qualifies when the boxes overlap at all and either the overlap alone
says so (IoU above 0.5, the old rule) or the embeddings do (cosine above
SAME_FACE_COSINE, 0.45, the reference library's P≈0.95 line). The
embedding route claims the box a low-resolution pass drew badly enough
that overlap alone would not; the vector is also what breaks the tie in a
group photograph, where two neighbouring faces overlap both new boxes.
Overlap is required on both routes, because the same vector elsewhere in
the frame — a mirror, a print on the wall — is not the same face and must
not take its name. Onto the matched face go the assignment as it was,
confirmed or suggested with its probability, and every rejection.

The merge's match_faces still matches by overlap alone across devices; it
is the same question and is not changed here.
2026-09-19 18:34:31 +02:00
dtourolle 46af2a0a46 Stop naming an optimisation level: on tract it means into_optimized, which aborts on yolo26n-seg
ONNX Runtime's default is already its fullest level. ort-tract maps any
level but disabled to tract's optimiser, whose slice pass divides by
zero inside the segmenter's graph — a panic across the C API and so an
abort, which is what stopped dr-ui's develop test. The app never asked
tract for that and does not start now.
2026-09-19 16:35:22 +02:00
dtourolle 95c9cffc0d Keep the embedder off the Hexagon, and let the probe example ask for a runtime
On the tablet the engine compiled arcface for the NPU: the routing
compared the form a rung wants with the form on offer, and for the
embedder both are f32, so nothing said no. A rung now says which roles
it serves at all, and the Hexagon does not serve the embedder (§7 —
its vectors must compare across devices). Tested at the routing seam.

dr-segment's onnx_probe example still named ort-tract, which is what
stopped the workspace test build.
2026-09-19 16:21:53 +02:00
dtourolle cbbe67fbd7 Let the probe's clock be its proof, not disable_cpu_ep_fallback
The strict flag refused the Hexagon over the ten quantise/dequantise
nodes at the graph's edges that QNN declines by policy, which cost
microseconds. A provider that hands real work to the CPU is slower than
the CPU floor and the timing already rejects it; the tablet measured
2.3 ms on the NPU against a 29.7 ms floor.
2026-09-19 16:13:19 +02:00
dtourolle 691af96e3e Keep the readable half of a provider's error for the settings row
ONNX Runtime's errors open with a source path and a template signature;
the first 160 characters of a CUDA failure were all signature. The
reason now starts at the first word a person can act on.
2026-09-19 16:07:19 +02:00
dtourolle 7a436e2549 Move the panorama keypoint detector onto the engine, and probe with a detector
XFeat's two exports are a Keypoints role now; the crate no longer names
tract, and the app compiles TensorRT engines for both ahead of the
first merge. The probe picks the smallest *detector* rather than the
smallest file: the tablet's first run chose the 112 KB eye classifier,
which has no int8 form, and reported the Hexagon as failed for want of
one.
2026-09-19 16:05:08 +02:00
dtourolle 76bc5652d7 Calibrate the int8 detectors on library proxies, in chunks, and measure them
The first int8 files found no faces at all, and for two reasons the
tool now guards against. The calibration set was landscape photographs
with no faces in them, so the score head's ranges had never seen the
face regime; the set is now proxies from the library itself. And ONNX
Runtime's strided and moving-average calibration modes both degrade
these graphs measurably (a quarter of the faces at eight images, none
at ninety-six), while driving the calibrator in chunks by hand gives
ranges identical to a single pass — so the tool does that, four images
at a time, and feeds quantize_static through its range cache.

Measured against f32 over 400 proxies (docs/inference.md §10.1): the
10g form finds every face above 32 px the f32 form finds; 500m and
2.5g find 96%, and what they lose sits at a median confidence of 0.52
against the 0.50 threshold. Shipped with the number on record.

The Android unpack list gains the three int8 files; without that the
tablet never saw them. D13's runtime half records the reopening.
2026-09-19 16:02:44 +02:00
dtourolle 4ed29b9d81 Add the int8 detectors for the Hexagon, calibrated on real photographs
tools/quantise-models.sh writes the QDQ form QNN's HTP backend takes
whole: opset 17, per-channel int8 weights, uint8 activations, ranges
from running the f32 graph over photographs fed exactly as the app
feeds them. The calibration is strided, four images at a time, because
every ONNX Runtime calibrator holds each image's whole set of
activations until it folds them — a gigabyte an image on the 10g
detector, and an OOM kill with no message when folded once at the end.

Release-time, never on the device (docs/inference.md §5): it needs
real photographs and a person reading the recall measurement that
gates whether each file is offered.
2026-09-19 16:02:37 +02:00
dtourolle 05508741af Start the inference engine from both apps and show its choice in Settings
The desktop names where a package may have put libonnxruntime — an
override variable, beside the executable, the package's own library
directory, the Flatpak prefix, the system library directory — and
Android points at the APK's native library directory, which is also
what Qualcomm's DSP loader must be told for the Hexagon skel. Android
starts the engine at the end of the model unpack rather than at launch,
because the probe fingerprints the model files and a first launch has
none until then.

The About panel gains an Inference row beside Graphics, re-read every
two seconds while the probe runs and engines land, and faces.model_id
carries the detector's form: an int8 detector finds a different set of
faces and is a different population (docs/inference.md §7). A
low-memory signal drops every idle session with the GPU caches.

The APK assembly bundles ONNX Runtime and the Qualcomm HTP libraries
from Maven, fetched by tools/fetch-android-runtime.sh with their
published checksums; RUNTIME_DIR=none builds the tract-only APK, which
is a slower app and not a broken one. The desktop packages carry no
runtime yet.

Two probe fixes from the first desktop run: the floor must not be
built with CPU fallback disabled, and a versioned libonnxruntime.so is
a runtime too. On the reference desktop the probe now loads ONNX
Runtime 1.30, measures 30 ms on the CPU provider, and selects TensorRT
at 1.5 ms.
2026-09-19 16:02:37 +02:00
dtourolle d15c41e699 Add dr-inference-engine and route every model session through it
One crate names the runtime, the providers and the devices; dr-face and
dr-segment ask it for a session by role. It hands ort an API table once
per process — from a libonnxruntime it dlopens when the app names a
directory holding one, otherwise from tract — so the Rust build stays
free of C on every target and a package can install the runtime as a
file (docs/inference.md §3).

Sessions live in a registry behind a Model handle that holds the bytes,
not the session: every use refreshes a timestamp and a reaper unloads
whatever sat idle past the decay. A scan that runs the detector on each
image never lets it go idle; a click in the develop view lets the
segmenter go after thirty seconds; a handle used after that reloads,
and reloads on a higher rung if a compiled engine has landed meanwhile.

The probe walks the platform's ladder by building strict sessions and
timing them against the CPU provider, caches the choice against a
fingerprint of the runtime, driver, hardware and models, and compiles
engines for the selected rung in the background, smallest model first.
Nothing in this commit turns the native path on: the apps still run on
tract until they call init with a runtime directory.
2026-09-19 16:02:37 +02:00
dtourolle caf21bea64 Name the crate dr-inference-engine 2026-09-19 16:02:37 +02:00
dtourolle 6739fdf908 Specify per-device inference backends, with the 2026-09-19 measurements
tract runs every model on one core on every platform. Measured against
ONNX Runtime's providers on the MagicPad 2 and the reference desktop:
ORT CPU alone is 3-10x, the Hexagon at int8 runs the detectors in
1-3 ms, TensorRT is ~2x the CUDA provider. NNAPI, XNNPACK, WebGPU and
CUDA int8 were tried and excluded with the numbers that excluded them.

The spec keeps the build C-free: ort::set_api takes a table from a
dlopened runtime or from ort-tract, chosen once per process. Rungs
are chosen by building a real session, cached until an input changes,
and compiled engines are built in the background after the first
frame. The embedder stays f32 everywhere; int8 detectors are a
distinct model_id and are gated on a recall measurement.
2026-09-19 16:02:37 +02:00
dtourolle 42d11d919b cargo fmt and clippy across the panorama work, and one lint master carried
The dr-face comparison is master's: a negated partial-order test on the
eye box's width, rewritten as the two conditions it meant.
2026-09-19 15:53:06 +02:00
dtourolle 67f225beba panorama.md: MI-GAN as the border filler — MIT, six operators, 7.4 s a tile
Read and measured, not built. The bare 512 generator exports at a fixed
shape and loads under tract with nothing unsupported; at f32 on the
desktop CPU it takes 7.4 s per 512×512 tile, which puts a full-resolution
fill of the fixture's border at ten minutes. The three routes that would
make it viable are recorded, with the quarter-resolution fill the cheapest
and Hexagon int8 the one the model was designed for.
2026-09-19 15:30:49 +02:00
dtourolle 39adfd4b75 Regenerate the traceability matrix and gesture book after the rebase 2026-09-19 15:24:44 +02:00
dtourolle 57ed51c1c5 Projection chips redraw the preview; auto-crop as the DNG default crop
Picking a chip stored the choice for the merge and changed nothing on
screen — the chip did not even highlight, since the selected property
was never written back. Now the pick is reflected, and the job, waiting
for its decision, takes a Preview request, draws the alignment on the
chosen surface at proxy cost and reports again; the drain puts the new
picture and its size up. Auto is the surface the field of view suggests.

Also:
The largest rectangle inside the frames' coverage is found a row at a
time — a histogram of consecutive covered rows and a stack pass per row —
so the composite is never held to be measured (FR-MRG-11). It is written
as DefaultCropOrigin/DefaultCropSize (FR-MRG-4): the file opens on the
picture, the border is still in it, and resetting the crop shows it.
rawler reports the crop as the picture, which the test checks.

FR-MRG-4 records the question raised the same day — fill the border
rather than crop it — as open: a non-generative fill through the heal,
or a generative inpainter with its licence and weights. Neither decided.
2026-09-19 15:24:20 +02:00
dtourolle 30bd276d0b Merge page: outline every frame on the preview, and let it be tall
A sweep whose frames overlap by more than half reads as one photograph,
and the page's job is to show frames. Each footprint is walked along its
border and drawn in amber where it lands, so twelve frames look like
twelve and a misplaced one is visible as such. The preview may take most
of the page's height rather than 320 px.
2026-09-19 15:24:20 +02:00
dtourolle 75d2ceb23c Provenance in the sidecar, a launch hook for the page, and where it stands
derived_from and merge are top-level sidecar fields (FR-MRG-6): one line
per source in order, and how the composite was made. A build that
predates them keeps the lines as unknown and writes them back. The job
writes the sidecar beside the composite and stages it with its own
record when the composite goes through the outbox.

DARKROOM_START_MERGE=a.CR2,b.CR2 lands on the merge page at startup with
the job running on local files, on the model of DARKROOM_START_IDENTITY,
for looking at the page where synthetic clicks do not reach it. The fetch
and the start are shared with the grid's button.

panorama.md §11 records what exists, the fixture's figures, and the six
things still open, auto-crop first.
2026-09-19 15:24:20 +02:00
dtourolle 2e9a1eb0f0 The merge job and its page: a selection to a panorama DNG, confirmed first
dr_ui::merge is the orchestration with no interface in it: decode each
frame to sensor data and build its graph as a session would (orientation,
lens profile); render each through the camera-space tap at proxy size and
detect keypoints there, so the alignment is measured in the undistorted
frame the tiles are rendered in; align; solve one gain per frame from the
proxies' overlaps; draw the aligned set in colour for the page; then wait.
Nothing is written until a Decision arrives (FR-MRG-1). The merge writes
a linear DNG through the outbox with a destination record, so the drain
puts it beside its sources on a folder library and a server alike, and
the library rescans (FR-MRG-3).

merge.slint is the page, on the import page's model: the alignment
table with a failed frame named on its row and the button held off
(FR-MRG-5), the preview, the projection choice, Stop and Back. A
"Merge to panorama" button joins the grid's selection bar at two frames.

Headless, the example produces the fixture's 22 993 x 5 980 DNG in 45 s
on the reference desktop, exposures balanced across the stop of drift.
2026-09-19 15:24:20 +02:00
dtourolle 44ea763c61 dr-gpu: the merge pass — warp, accumulate, resolve, chunk by chunk
merge.wgsl warps one camera-space tile into one output chunk — output
pixel to direction (the projection maths of dr_pano::projection, verbatim),
direction to the frame's camera, camera to source pixel, bilinear by hand
from four textureLoads because rgba32float is not filterable — and adds it
into a storage-buffer accumulator weighted by its distance from the
frame's edge. A resolve pass divides by the weights and packs sixteen-bit
samples at the sensor's scale with a coverage bit.

MergePass::merge drives it: bands of rows, chunks across a band, and for
each chunk only the frames whose footprint meets it, each rendered as the
source rectangle the chunk needs and nothing more. The working set is one
chunk, one tile and one band (FR-MRG-11); the frame textures are the
caller's to cache. Feathered, not seamed; gain a scalar per frame — the
blend quality is panorama.md §10's step 5, after the path writes a file.
2026-09-19 15:24:12 +02:00
dtourolle acab0d7abb A linear DNG in and out: the writer, and a three-sample RawImage
dr-export gains write_linear_dng — LinearRaw, DNG 1.4, u16 samples at
the sensor's scale, the body's matrices with their illuminants, the
as-shot neutral, the EXIF block an export writes — streamed strip by
strip through a closure so the composite is never held (FR-MRG-11). The
tiff crate's directory is a map, so PhotometricInterpretation is written
over what new_image set, which is the trick the S15.1 spike thought it
had to hand-roll around. The test reads the file back through rawler.

dr-decode's RawImage carries samples_per_pixel (a linear DNG is 3), the
body's profile with its calibrations mapped back to EXIF illuminant
codes, and the cleaned make and model. The GPU uploads a three-sample
image as it is, normalised by black and white like a photosite, through
a full f16 conversion — subnormals kept, because a 14-bit LSB sits at
f16's smallest normal and rounding it to zero would crush exactly the
shadows the file was written to keep.
2026-09-19 15:24:12 +02:00
dtourolle 9b6b4942cf The camera-space tap: OutputMode::CameraLinear, composed with no operations
compose_camera_linear composes the fused pass with an empty operation
list, the file's orientation as the baseline, a view rect for the tile,
and a store of rgba32float. On the GPU, render_camera_linear is the only
entry that accepts it: it fills the profile uniforms neutral — unit white
balance, identity matrix, curve off — so what lands in the texture is the
sensor's numbers after the lens warp and nothing else (FR-MRG-2). A third
bind-group layout carries the format, as the linear one does, and the
readback is generalised to any pixel width for the f32 copy.

Thirty-two bits because the composite is written back at the sensor's
scale: a 14-bit sensor has 16 384 steps to white and f16 keeps 2 048 of
them in the top octave.
2026-09-19 15:24:12 +02:00
dtourolle 54290b9540 dr-pano: a second XFeat shape for portrait frames, and a matcher that takes seconds
Twelve real frames from the fixture set now align in 4.5 s — 4.4 s of
matching, 118 ms of bundle adjustment — where the first run took 51 s and
left the first two frames out.

The matcher computes each pair's similarity matrix once, across the
cores, with a dot product written to vectorise; both nearest-neighbour
directions read it. The frames that failed were portrait: fitted into the
landscape input they used 512 of 1024 px, and their thin overlap did not
survive at half resolution. The same weights are now exported at 768×1024
as well and the detector picks the shape by aspect. The example aligns
from embedded previews and draws the set on a cylinder; on the fixture the
sweep is 152° at a fitted 47.9 mm against the EXIF's 50, RMS 1.5 px, and
the overlaps show no ghosting.
2026-09-19 15:24:12 +02:00
dtourolle 231b4a54ab dr-pano: the geometry, from features to cameras
A new crate holding the CPU half of a merge (FR-MRG-10): the grayscale
proxy with orientation, the XFeat decoder ported step for step from the
reference detectAndCompute, mutual-nearest-neighbour matching, a robust
pairwise homography with the focal length read off it, a hand-rolled
Levenberg–Marquardt bundle adjustment over every rotation and the focal,
the three output projections, and align(), which chains it all and names
the frames it could not place rather than guessing (FR-MRG-5).

Dependency-free without the xfeat feature — linalg.rs says why the dense
algebra is hand-rolled — and tested on synthetic sweeps whose answer is
known exactly. The noise test records the single-row degeneracy: one
pixel of noise is a tenth of a percent of focal, which is a uniform
stretch of the sweep, not a misalignment.
2026-09-19 15:24:12 +02:00
dtourolle 2bf0ec8dba S15.4, CPU half: XFeat runs in ~400 ms per frame on the tablet
tools/onnx-probe-on-device.sh cross-builds dr-segment's onnx_probe
without the embedded segmentation model, pushes it with a model to the
attached device and times two runs. The 768×1024 XFeat export takes
~400 ms on the reference tablet's NEON cores against ~300 ms on the
desktop, with identical output ranges — inside NFR-MRG-1's 1 s per frame.
The blend half of S15.4 waits for a chunked blend to exist.
2026-09-19 15:24:12 +02:00
dtourolle 5bf06c5030 Fixture README: the frames carry Orientation 8, not 6 2026-09-19 15:24:12 +02:00
dtourolle 44fdcbc6f7 Add the twelve-frame 6D panorama set as an LFS fixture
fixtures/pano/2025-08-05: _MG_8320 … 8331, one portrait hand-held sweep
at 50 mm with a stop of shutter drift and sky in every frame — the set
§3.11 is built against, with each of those facts named as the test it
is. fixtures/** is tracked in LFS like the models but with the opposite
default: CI's pulls exclude it, so a build never fetches 325 MB it does
not use.
2026-09-19 15:24:11 +02:00
dtourolle f9510405c3 FR-MRG-3: the composite is a RAW at the source's native scale
Camera-linear u16 samples on the first source's black-subtracted scale
with its white level, never rescaled to fill 16 bits, with its body,
matrices, illuminants and as-shot neutral carried — so the panorama is
developed afterwards as one photograph from the sensor's own numbers.
The only thing a warp cannot preserve is the colour filter array, and
the clause says so.
2026-09-19 15:24:10 +02:00
dtourolle 7e6b25b21b S15.3: the camera-space tap is uniforms, not structure — and FR-MRG-2 moves below the profile
The fused chain, as operation.rs's tests fix it, is warp → as-shot white
balance → operations → base curve → camera matrix → store. LinearWorking
stores after the matrix, so the existing linear tap carries the body's
base curve, and a composite stitched from it and developed as an
unprofiled body would render that curve twice.

FR-MRG-2 therefore stitches camera-linear RGB — after the warp, before
white balance, curve and matrix — and the composite carries the first
source's body, matrices and as-shot neutral so its own develop applies
the profile once. The composer already makes this a uniform question:
white balance, matrix and the curve flag are reserved uniforms, so the
tap is a compose entry with no operations and a render entry that fills
them neutral. panorama.md §5.1 states the shape and asks for f32 buffers.
2026-09-19 15:24:10 +02:00
dtourolle e4b6b6c935 S15.2: XFeat exports at a fixed shape and loads under tract
tools/export-xfeat.sh exports the convolutional network alone at 768×1024
grayscale, on the pattern of export-seg-model.sh: thirteen standard
operator types, no dynamic axes, the keypoint decoding left to Rust.
examples/onnx_probe loads it through the ort-over-tract backend the app
ships with nothing unsupported and runs it in ~300 ms on the desktop CPU.

The weights are Apache-2.0, read from the repository's LICENSE, with no
grant on the checkpoint — recorded in models/LICENCE.md before they land,
as FR-MRG-8 asks. The probe stays: the next model will need the same
check.
2026-09-19 15:24:10 +02:00
dtourolle 1ded5afbaa S15.1: rawler reads back a linear DNG, so that is the container
A hand-rolled 64×48 LinearRaw DNG — one IFD, 16-bit RGB, DNGVersion,
ColorMatrix1, AsShotNeutral — comes back through rawler 0.7 with cpp 3,
the samples in the order written and the matrix parsed into the camera
definition; CameraProfile::extract builds a profile from it. ImageMagick
reads the same bytes.

dr_decode::decode currently accepts the file as CFA and passes three
times the samples on, so the cpp == 3 branch is the decode work FR-MRG-3
needs, and the only decode work. panorama.md §8 records the result.
2026-09-19 15:24:10 +02:00
dtourolle c901fc1a0a Specify panorama merging: §3.11, D18, S15, and the design in panorama.md
A merge writes a new source file beside its sources (D18) rather than a
multi-source Version, which answers the schema question §7 had been holding
open for panorama, HDR merge and focus stacking together. The panorama is
undeferred as FR-MRG-1 … 11; the other two stay in §7 with their data model
decided.

FR-MRG-10 and 11 fix where the work runs — every per-pixel stage on the GPU,
the composite never held as one texture — because the output exceeds
max_texture_dimension_2d before it exceeds memory. panorama.md carries the
stage table, the chunked output driver, the model licences and the porting
sources. S15 gates all of it.

Coverage falls from 83.0% to 77.2%: thirteen requirements entered with no
code, and outstanding.md §11 says so.
2026-09-19 15:24:10 +02:00
dtourolle f79a76f2d5 Name the eye pass on the People screen
Once every image has been through the detector and only readings are
left — the state an already-indexed library is in the day the eye models
arrive — the button reads "Read eye state" rather than promising to
index, and the coverage line says what the faces are waiting for.
2026-09-19 14:24:16 +02:00
dtourolle facb44cb55 Keep the dense landmarks behind each eye reading, packed
The 106 points the eye boxes were cut from, stored beside the reading as
16-bit fixed point over the frame: 424 bytes a face, a seventh of a pixel
on a 6000-pixel frame, where f16 at the same size would have been six.
Derived data like the embedding, kept for the same reason — it cost a
fetch and a model run, and the next per-face pass should run from the
catalog. Shards carry it; a peer's shard from before it is still read.
2026-09-19 14:24:15 +02:00
dtourolle 85cc2b1dcc Trace the eye reading to FR-CULL-8a and the chip to FR-CULL-13
The register grew both clauses the same day this was built: FR-CULL-8a is
the per-face state the reading is, and FR-CULL-13 is the rule that a
signal is shown and filtered and never writes a judgement. The tags,
faces.md §17 and catalog.md now say which is which; FR-CULL-8a records
what of it is built, and that its third model is under the InsightFace
grant by the same decision as the pair.
2026-09-19 14:06:59 +02:00
dtourolle d706c12d77 Cover the eyes-open subquery with an index
The people filter was served from faces_image without touching a row;
reading the eye columns in the same subquery touched every one, and
ALTER TABLE had put those seven floats after the embedding and the crop
blob. One count took 24 seconds on the reference library, thirteen of
them system time. faces_eyes covers the subquery again: five
milliseconds.
2026-09-19 14:05:52 +02:00
dtourolle cd0ca6785f Specify eye state as a filter term, and record what was measured
FR-CULL-13, with §3.9.1's exclusion of blink detection re-read as the
exclusion of blink selection it always was: the stored fact and the chip
are built, a pass that picks the frame where everyone's eyes are open is
not. faces.md §17 has the models, the crop measurements, the four-state
rule and its floors, the native and proxy sheets read face by face, and
what remains to measure.
2026-09-19 14:05:50 +02:00
dtourolle 83f4253b6a Filter the grid to a person with their eyes open
An "Eyes open" chip beside the people chips, offered only while someone
is chosen and dropped when the last person goes, so no term narrows the
grid with nothing on the bar to say so. It compiles the rule in
dr_face::eyes into the person's face subquery — Anna, eyes open, whoever
else is blinking beside her — and drops a frame only on a closed eye that
could be read: sunglasses, eyes too small or soft to read, and faces never
read all pass, so an old library shows everything under the chip until
the measuring pass has run. A test drives the same readings through the
SQL and through the rule and requires them to agree.

The People screen badges a face "Eyes closed", "Sunglasses" or "Eyes
unclear" so the reason a frame is or is not in the grid can be read off
the face; the sweep loads the three models when they are beside the pair
and reads eyes on the indexing and measuring passes from the native
render; the coverage line counts unread faces as work to measure so an
already-indexed library keeps its Index button. The term travels with the
place.
2026-09-19 14:04:35 +02:00
dtourolle 6aae4c3eb0 Ship the three eye-state models beside the face pair
2d106det for the eye contours, OCEC for open or closed, SGC for
sunglasses — all three pinned to a batch of one by the same script as
the pair, and installed by every packager so the eyes-open filter works
out of the box. The two classifiers are MIT, code and weights; the
README records their provenance, SGC's undocumented training set, and
the hashes as fetched and as shipped.
2026-09-19 14:04:08 +02:00
dtourolle 54b543fb77 Store seven eye numbers per face rather than three
Per eye P(open), the pixels across its box and the sharpness of the
patch; and P(sunglasses). The verdict — open, closed, sunglasses,
unclear — stays a rule in dr_face::eyes so the floors can move without
re-measuring twenty thousand faces. Shards carry the same seven, and a
peer's shard from before any of them is still read.
2026-09-19 14:04:08 +02:00
dtourolle f5956707e7 Cut the eye box from a landmark contour, and refuse eyes that cannot be read
SCRFD's eye point places a face, not an eye: on turned and smiling heads
the classifier's window had the eye in a corner, and two model-free ways
of re-centring it — the darkest blob, the most contrasty window — both
lost open eyes (19 → 15 and 19 → 9 of 25). Three landmark models were
then run over the same faces; Face Mesh V2 and InsightFace's 2d106det
tied at 22 of 25 and 2d106det ships, being the cheapest by far and under
the grant the detector and embedder already carry. The eye box is the
tight bounding box of its ten lid points, cut upright from the native
render, which is what the classifier was trained on.

The larger change is that the reading now carries, per eye, the source
pixels across the box and the sharpness of the patch — because the
commonest wrong answer on the reference library was a soft eye read as
closed, and a classifier shown a smear will always say something. An eye
under either floor, or narrower than six tenths of its partner (the far
eye of a turned head, whose contour collapses), is not asked; a face with
no readable eye is a fourth state, Unreadable, that no filter drops. On
twenty native renders the one real blink is caught, the laughing faces
are closed, the profiles are judged on the near eye, and the one thing
left beyond any floor is a face with a pot held over it.
2026-09-19 14:04:08 +02:00
dtourolle b908d861e0 Keep each face's eye reading in the catalog and in its shard
Three nullable columns beside quality — P(open) for each eye and
P(sunglasses) — because the verdict is a rule with thresholds in it and a
rule belongs in code, not in rows that would have to be re-measured. NULL
is "never read": a face from before the models, or from a device without
them, and every reader treats it as unknown rather than as closed.

The measuring pass V14 built for the embedding's length is what fills
them, so the sweep's work list now also names faces with no eye reading
— but only on a device that has the models, or it would fetch every
original to do nothing to it. A peer's shard without the reading is still
adopted, unlike one without the quality: the pass finds this work by the
NULL rather than by the run marker, so adoption costs it nothing.
2026-09-19 14:04:06 +02:00
dtourolle 6b51726322 Read each face's eyes, and whether sunglasses hide them
Two MIT classifiers from the same author as the reference pipeline's
whole-body detector: OCEC answers P(open) for one 40×24 eye, SGC
P(sunglasses) for a 48×48 head. Both load in tract once their batch
dimension is pinned by tools/fix-face-model-shapes.sh, like the embedder.

The crops come through the same fitted similarity the aligned face does,
so an eye window is a constant in template units rather than a second
warp, and a tilted head yields an upright eye. Measured on 60 proxies
from the reference library: the eye window plateaus at 22×11, the S
variant beats M and L (which overfit their own domain), and for
sunglasses the aligned face beats a head framing but the higher of the
two catches 11 of 12 pairs against 9 for either alone.

The reading keeps both eyes and the sunglasses number apart, because a
wink averages to the least informative value and a lens of dark glass
draws a confident answer from the eye classifier — over a woman in
sunglasses it read the right eye 0.97 open. Sunglasses take precedence,
and a face behind them is neither open nor a blink.
2026-09-19 14:03:31 +02:00
dtourolle 2481904016 Bring outstanding.md up to the decisions of 2026-09-19
Its plugin section still asked for the contradiction to be resolved, its
render-path section still asked whether FR-DSP-2 was a requirement and
said NFR-RES-2 had no answer, and its closing section still called D12
open. Each now records what was decided and keeps the argument that was
weighed, so the document reads as the history it says it is rather than
as a plan the register has moved past.
2026-09-19 12:25:04 +02:00
dtourolle a92ae4576f Repair the references that point at sections that moved
Eight citations named §5.1, §5.2 and a §5 selector language that
requirements.md's §5 has not contained since it became a pointer at
architecture.md; two named §9 for the golden images and the benchmark
suite, which are §8; and the three pointers into architecture.md were
each one section off. All now name the section that holds the thing.

architecture.md §12's subsections are numbered 6.1–6.13, colliding with
its real §6. That numbering is what every ARCH §6.n citation in the tree
uses, so it stays, and a note at the head of §12 says so instead of
leaving the next reader to work it out.

FR-DEV-3f's open question about persisting the film stock was answered
in sidecar.rs; the clause now says so.
2026-09-19 12:25:04 +02:00
dtourolle ed4460cb9c Tag three requirements the code already meets
R5 says in its own note that zoom_resolution.rs establishes it as a
pixel equality; that file was tagged FR-DSP-5 alone. FR-DEV-19's three
sub-clauses carry eighty-three tags between them while the parent had
none; MaskLayer, which is the thing they edit, now carries it. And
NFR-R3 — a crash in decode does not take down the application, the
image is marked failed — is exactly what the decoder's panic guard and
the face sweep's unreadable mark do, tagged FR-RAW-4 and NFR-SEC-1 and
not the clause that asked for them.
2026-09-19 12:25:04 +02:00
dtourolle 7596cf9bcc State the compatibility baseline and the channels
NFR-COMPAT-1 and NFR-COMPAT-2 were instructions to write a requirement,
not requirements: "state the API level", "state the channels". Both
are now stated from what the build enforces and what exists.

The baseline is minSdk 28 / targetSdk 36 from the Android Dockerfile,
a Vulkan adapter at wgpu's default limits because compute needs storage
textures — device_from already called that the floor and is tagged for
it — with no optional feature required, since the f16 in FR-DEV-2 is a
texture format and not shader arithmetic. The reference device is the
HONOR ROD2-W09 the figures are taken on, and the second-vendor clause is
recorded as unmet rather than quietly dropped: there is no Mali or
PowerVR device, so an Android figure here is an Adreno figure.

The channels are all self-distribution — Arch package, local Flatpak,
sideloaded APK, NSIS installer — because D13's face weights rule out
every store, and the two consequences are written down: SAF stays
although a sideloaded build need not have it, and S11 becomes a
pre-publication step.
2026-09-19 12:25:04 +02:00
dtourolle 696bafa9d5 Undefer AI subject masking, which shipped, and give it a clause
§7 still listed "AI subject masking — deferred per D11" while
MaskSource::Subject and MaskSource::Category, backed by dr-segment's
instance and semantic models, had been the primary way a local
adjustment is made for weeks. The code was tagged FR-DEV-3, which
names gradients and brushes and says nothing about a model.

FR-DEV-3i now states what exists: a subject or a category found by a
local model, stored as identity with the run's signature so that it
merges per field and reads as stale rather than wrong, then treated as
any other layer by the edge, stroke, composition and reveal clauses.
The one place it departs from FR-DEV-19 — coverage written run-length
coded beside the layer, so a stored subject renders without a model —
is recorded in the clause instead of left for the next audit to find.
The segmentation crate and the UI's selection module are tagged to it.
2026-09-19 12:25:03 +02:00
dtourolle d259c0d4bb Say that FR-DSP-2 is waiting on S6, not that it was rewritten
R5's note said FR-DSP-2 "was rewritten rather than implemented". It was
not: the clause still demanded viewport tiling, the matrix listed it
unbuilt, and frame-budget.md's rewrite had been proposed and never
applied. Decided 2026-09-19 to keep it as written until S6 runs on a
mid-range Android device, because the measurement that argues against
tiling was taken on a discrete desktop GPU and the clause exists for the
device whose memory the image exceeds. Both notes now say that.
2026-09-19 12:25:03 +02:00
dtourolle 95458356da Record D13's position on the face weights
The licensing half of D13 had been open since 2026-08-09, while the
InsightFace detectors and embedder shipped in the tree and indexed real
libraries. models/face/README.md already stated the position the
project was actually taking; the register did not.

Now it does: this is non-commercial software, self-installed, and it
uses the weights under their research grant as such. The risks are
written where the decision is — the grant binds every user, it is not
GPL-compatible, it rules out every public channel, and publishing is
what reopens the decision. S14's licence search is what would close it.
2026-09-19 12:25:03 +02:00
dtourolle dc9db11033 Decide NFR-R8: no CPU pipeline, a degraded mode instead
NFR-R8 carried the words "decide explicitly" for six weeks, asking
whether v1 has a full CPU render path or whether "CPU fallback" means
staging only. The viewer had already answered it: with no adapter it
opens the library on embedded previews and cached proxies, keeps every
catalog edit available, and withholds develop and export. That is the
degraded mode, it is now the requirement, and NFR-RES-2 no longer
promises a fallback render path that was never going to be built.
2026-09-19 12:25:03 +02:00
dtourolle 3692306fd3 Say once that there is no phone
Three statements disagreed. §1.2's platform table said "phone
supported"; §3.5 said phones were out of scope and cited §1.3, which
does not mention them; D15 said "no phone" and gave the reason. D15 is
the decision, so the other two now point at it and say the same thing:
the build runs on a phone, and nothing is designed for one.
2026-09-19 12:25:03 +02:00
dtourolle c826fed605 Put the plugin API post-v1, and let the matrix count it that way
The register said two things about plugins. §7 had listed "Plugin API"
as deferred since the first draft, in a bare row; §3.10 then specified
it in 23 clauses that counted against coverage. Twenty-one of them had
no implementation of any kind, and could not have: no crate loads
anything at runtime. The coverage figure was measuring the contradiction.

Decided 2026-09-19: §7 is right. §3.10 stays as the design of record,
each of its clauses is marked "(post-v1)" on its defining line, and
NFR-SEC-6 — which exists only for plugins — goes with them, as does D16.

The traceability tool learns the marker. A deferred requirement is still
defined, so a tag naming it is not an orphan, but it leaves the
denominator and is listed in its own table rather than under "not yet
tagged". The marker must sit on the definition line; a mention of
"post-v1" in prose changes nothing, and where an ID is defined twice the
deferral on either line wins. Both are tested. Coverage moves from 72.2%
of 194 to 80.6% of 170 without a line of application code changing,
which is the honest figure: it now measures what v1 owes.
2026-09-19 12:25:03 +02:00
dtourolle c921852d89 Record D12 as settled by events, and D3 as delivered
D12 had been OPEN since the 2026-08-08 calibration, and D3 said it
depended on D12. In the meantime the milestone D3 named was delivered
and closed on 2026-08-30 and the application reached 0.12.2 with every
cluster the calibration selected at least begun. The decision the
register was waiting for had been made by building, so the register
now says so: full scope stands, v1 has no date, and "post-v1" in §7 is
the one way a clause leaves the count.

The status line also stops calling this a draft from August; it has
carried eleven dated amendments since.
2026-09-19 12:24:54 +02:00
dtourolle e6ac31d39d Give the face sweep a size budget, so a panorama is never fetched
The sweep fetches the whole original before it can learn anything
about it, and the one file in the reference library the decoder
refuses on sight is a 521 MB stitched panorama — so every pass on the
tablet spent half a gigabyte of Wi-Fi to find that out again. The
catalog already knows the byte count, and that is enough to decide
before the fetch: originals over 256 MB are marked examined with
nothing found and a zero edge, counted as failed, and named in the
log. Below the line is every camera RAW the library holds; above it,
four files, all panoramas.

A budget and not a verdict on panoramas. The right treatment for one
is a tiled pass — read it in strips, detect in each, stitch the boxes
back — and the zero edge is what that pass would select on. Until it
exists, this is what keeps a background sweep on a phone from paying
for the decision the decoder cannot make.
2026-09-19 12:14:05 +02:00
dtourolle 7db999c1f6 Require judgement anywhere, and evidence that never becomes a verdict
Rating and flagging were reachable from the grid alone, so a photograph
opened in develop could not be judged without leaving it; FR-UI-5 said
"rating" without qualifying the view and was built as though it had.
And FR-UI-1's expanded row has said "filmstrip" since it was written
while the roll stayed on demand in both classes. Both are amended to
say what they meant: judgement follows the photograph, without
auto-advance outside the culling mode, and the roll is open by default
where there is room for it.

The larger change is a rule. Per-face signals — eye state from a
classifier, head pose from the five landmarks the detector already
yields — are worth having for culling, and §3.9.1 excluded detecting a
blink outright. The exclusion was always of judgement, not of knowing:
a blink is a fact about a frame of the same kind as a clipped
highlight. FR-CULL-8a specifies the two signals; FR-CULL-13 says what
any signal may do (be shown, filtered, sorted, propose a burst
representative) and what none may (write a rating or flag without a
user action between). R7 states the same thing as a user need.

Licensing was read before either was written. OCEC's eye-state weights
are MIT with a clean data chain; every open gaze model is trained on
Gaze360 or its peers, whose licences restrict derived models by name,
so gaze is deferred in §7 and head pose stands in for it. D13 records
both so they are not re-searched.

Replacing a closed-eyed face from a neighbouring frame was raised and
is written down as D17 rather than built: it is the multi-source schema
question §7 already defers for panorama and HDR, with its non-goals —
never automatic, provenance declared — fixed now.

Traceability regenerated: three new IDs, none yet tagged.
2026-09-19 11:47:27 +02:00
dtourolle 30b89ad70a Merge: one face population per embedder, whichever detector found them 2026-09-19 10:52:31 +02:00
dtourolle 327decfab1 Fuse every detector's faces into one population per embedder
Choosing "Thorough" made the library look empty. The detector setting
writes under its own faces.model_id, and every reader of "the faces"
keyed on that exact id: the clustering pass, the coverage figure, the
sweep's work list, the shard export and import, and the sync merge's
face matching. On the reference library that restarted coverage at
1,834 of 19,140, drew a People rail of 36 faces for a person with 520,
queued a ~400 GB re-fetch on each device, and stranded the desktop's
3,583 confirmations under the old id: the tablet held the same faces
under the new one and the merge refused to match them. Same photograph,
same box, same embedder, two ids — that is one face, not two libraries.

The embedder half of the id is now the key. embedder_of and embedder_sql
give it to every query; writes keep the full id, so which detector drew
a box stays on record. record_detections is unchanged and is where the
generations meet: an image holds one pipeline's faces at a time, and a
re-detection carries confirmations across by box overlap. The merge's
match_faces applies the same rule within an embedder. The calibration
is keyed on the embedder too, since the similarity space did not change.

Shards travel every generation, each under its own id, and a peer adopts
whichever it is sent — including a stronger detector's pass over an
image it indexed itself with a weaker one, which is the re-detection its
own sweep would otherwise queue, already done. Never downwards: a tablet
on Fast keeps the desktop's Thorough faces. The sweep gains the same
tail — images a weaker detector indexed, after the ones nothing has —
driven by FaceDetector::supersedes, so choosing a stronger detector still
improves the library over time without first making it disappear.
2026-09-19 10:49:28 +02:00
dtourolle f8addbee53 Mark a file the decoder cannot open, so the sweep stops fetching it
A decode failure in the face sweep was counted, logged at debug where
nobody saw it, and left unmarked — so the next pass fetched the same
file and failed the same way. For the 521 MB panorama behind rawler's
panic that was half a gigabyte per sweep, on a tablet. It is now marked
examined with nothing found and a zero edge, which is what a later "try
again with a better decoder" pass would select on, and the warning
names the file. The failure count is unchanged: it did fail.
2026-09-19 10:45:08 +02:00
dtourolle c0b1e78f7c Return a panic inside the decoder as an error, not as the end of the thread
rawler panics on some input rather than returning Err — a DNG whose IFD
claims a >50000 px image, which the reference library has: a 521 MB
stitched panorama, IMG_4181-Pano.dng. On a worker thread a panic is the
end of the thread, so the face sweep that met it stopped thirteen
seconds in, three sweeps running on the tablet and three on the
desktop, with "17301 image(s) to index" as the last word. FR-RAW-4
says a malformed file must not abort a batch, and that is this crate's
promise whatever the library beneath it does: every entry point that
calls into rawler now runs under catch_unwind, and a file that panics
the decoder is one failed file with the panic's message in the error.

Verified on the panorama itself: metadata reads, decode returns the
error, the thread survives. The crash hook still records the panic,
which is right — it is a defect in a dependency and the record is how
it gets reported.
2026-09-19 10:45:07 +02:00
dtourolle 78cb00634e Fetch the photographs around the open one ahead of the step to them
Benchmarks / CPU and I/O (per commit) (push) Successful in 3m59s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / android-image (push) Canceled after 0s
🐳 Android image / Build and push (push) Canceled after 0s
Build and test / Android (aarch64) (push) Canceled after 0s
Build and test / windows-image (push) Canceled after 0s
🐳 Windows image / Build and push (push) Canceled after 0s
Build and test / Windows (x86_64, cross) (push) Canceled after 0s
Build and test / Layer separation (push) Canceled after 0s
Build and test / Desktop (Linux) (push) Canceled after 16m26s
Traceability / Requirement traces (push) Canceled after 0s
Walking the photo roll was one download per frame: every step showed
"Downloading…" over an empty canvas while tens of megabytes came down,
and moving between a pair of near-identical frames paid that a dozen
times. Now, once the opened photograph has landed, the ones around it
are fetched into the originals cache while it is being looked at, so
the next step is a disk read.

A single worker serves the latest wish only, closest first and working
outwards — next, previous, next-but-one, previous-but-one… — one file
at a time. Each open replaces the wish, so a fast walk never leaves a
trail of stale downloads competing with the one being waited on. A
process-wide in-flight registry makes a click on a photograph that is
still being fetched ahead wait for that transfer and read it from disk,
rather than start a second download of the same file.

How far each side is a setting under STORAGE — Off, 2, 5, 10 or 20,
defaulting to 5 — and it is moot while "keep originals after opening"
is off, since a fetch the cache would discard on arrival is transfer
for nothing. Nothing is fetched ahead while offline. The transfers show
in the activity list while they run and are removed when they end.
2026-09-19 10:36:50 +02:00
dtourolle 2917b7427d Release 0.12.2
Benchmarks / CPU and I/O (per commit) (push) Successful in 12m51s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Successful in 1h44m42s
Build and test / Layer separation (push) Successful in 1m4s
🐳 Android image / Build and push (push) Successful in 6s
Build and test / android-image (push) Successful in 7s
🐳 Windows image / Build and push (push) Successful in 3s
Build and test / windows-image (push) Successful in 3s
Traceability / Requirement traces (push) Successful in 2m15s
Build and test / Android (aarch64) (push) Successful in 1h4m7s
Build and test / Windows (x86_64, cross) (push) Successful in 1h12m17s
2026-09-19 10:15:46 +02:00
dtourolle 33e2e277a2 Set the Wayland app id late enough for it to take
The launcher and the task bar have shown a generic tile for a working
window since the call was written. set_xdg_app_id sat at the top of
run(), on the reasoning that the app id is read when the surface is
created — true, and beside the point: the call goes through Slint's
global context, and there is no global context until something installs
a platform. That is BackendSelector inside shared_gpu, or AppWindow::new
falling back to the default, and both happen further down. Called before
either, it returned NoPlatform and did nothing at all.

It moves to just after the window is constructed, which is not the same
as shown — run() is far below — so there is a platform to talk to and
the surface does not exist yet.

The failure was logged at debug, which is why a year of grey squares went
unremarked: the whole symptom is invisible from inside the application.
It is a warning now, naming the consequence.
2026-09-19 10:15:46 +02:00
dtourolle 9b627e7713 Let a catalog writer wait for its turn instead of losing its work
SQLite's busy timeout defaults to zero, and nothing ever set one: the
loser of a write race got SQLITE_BUSY at the moment it asked. WAL does
not cover this — it makes one writer and many readers free, and this
application constantly has two writers, the face sweep committing a
batch while the derived sync imports shards or reclustering reads.

The cost was not a retry but lost work. A sweep that had already paid
for the detection and the embedding — seconds per image, the expensive
part — discarded the result on "storing faces for 214: database is
locked" and moved on to the next image. Both the desktop and the tablet
logged runs of those on consecutive images, which is a face sweep
quietly failing to store the faces it had just computed.

Ten seconds, on every connection, set in configure() so that nothing
can open the catalog without it — the figure the job runner's own tests
have used for this reason since they were written. It is far longer
than any transaction here, so it bounds pathology rather than making
anyone wait.
2026-09-14 20:05:53 +02:00
472 changed files with 132030 additions and 34311 deletions
+19
View File
@@ -10,3 +10,22 @@
# detects exactly that and fails with an instruction rather than embedding the
# pointer and failing at inference time.
*.onnx filter=lfs diff=lfs merge=lfs -text
# Test photographs live in LFS too, and are fetched only by the tests that
# need them.
#
# `fixtures/**` holds real camera files — a twelve-frame panorama set is
# 325 MB — and CI's `git lfs pull` excludes the directory, so a checkout
# carries pointers there until a merge test asks for the frames. Same
# reasoning as the models, with the opposite default: the model is not
# optional and the fixtures are.
fixtures/** filter=lfs diff=lfs merge=lfs -text
# The manual's pictures live in LFS for the same reason the models do: a
# screenshot or a GIF changes wholesale when the interface it shows changes,
# and every re-recording would otherwise stay in every clone for good. The
# desktop and benchmark legs exclude the directory, since nothing they build
# or test reads it; the Android and Windows legs fetch it, because the APK
# and the installer carry the manual (docs/manual/index.html) with its
# pictures, and their packagers refuse a pointer.
docs/manual/media/** filter=lfs diff=lfs merge=lfs -text
+4 -4
View File
@@ -1,6 +1,6 @@
name: Benchmarks
# The suite docs/requirements.md §8 has been promising since it was written:
# The suite docs/dev/requirements.md §8 has been promising since it was written:
# "an automated benchmark suite against a synthetic 50k catalog, run per-commit
# … A regression beyond stated tolerance fails the build."
#
@@ -29,7 +29,7 @@ name: Benchmarks
# every commit to establish, every time, that this runner has no GPU. It
# runs on demand (Actions → Run workflow) so that a runner that *does*
# have one can be pointed at it, and the numbers it produces belong in
# docs/frame-budget.md by hand, as they already are.
# docs/dev/frame-budget.md by hand, as they already are.
on:
push:
@@ -148,7 +148,7 @@ jobs:
| while read -r key; do git config --local --unset-all "$key"; done || true
git config --local lfs.url \
"https://x-access-token:${LFS_TOKEN}@gitea.tourolle.paris/dtourolle/DarkRoom.git/info/lfs"
git lfs pull
git lfs pull --exclude="fixtures/**,docs/manual/media/**"
- name: Cache cargo
uses: actions/cache@v4
@@ -183,7 +183,7 @@ jobs:
- name: Frame budget (FR-DSP-3)
run: cargo test --release -p dr-gpu --test frame_budget -- --nocapture
# The instrument behind docs/frame-budget.md. It exits non-zero with no
# The instrument behind docs/dev/frame-budget.md. It exits non-zero with no
# adapter, which is right for a tool a person runs deliberately and wrong
# for a job that usually has none — hence continue-on-error. Its table is
# in the log for whoever asked for this run; the committed numbers are
+80 -6
View File
@@ -7,6 +7,10 @@ name: Build and test
on:
push:
branches: [main, master, develop]
# A release tag builds again and publishes what it built (the `release`
# job at the end). The master push of the same commit has usually filled
# the caches, so the second run is the warm one.
tags: ['v*']
pull_request:
branches: [main, master, develop]
@@ -96,7 +100,9 @@ jobs:
| while read -r key; do git config --local --unset-all "$key"; done || true
git config --local lfs.url \
"https://x-access-token:${LFS_TOKEN}@gitea.tourolle.paris/dtourolle/DarkRoom.git/info/lfs"
git lfs pull
# The manual's pictures too: the APK carries the manual, and
# assemble-apk.sh refuses a pointer where a picture should be.
git lfs pull --exclude="fixtures/**"
ls -lR models/
- name: Cache cargo
@@ -154,6 +160,16 @@ jobs:
- name: Build
run: cargo build --workspace --release
# Only on a release tag: the binary is 150 MB and nothing but the
# release job wants it.
- name: Upload the desktop binary
if: startsWith(github.ref, 'refs/tags/v')
uses: actions/upload-artifact@v3
with:
name: darkroom-desktop-x86_64-linux
path: target/release/darkroom-desktop
if-no-files-found: error
- name: Disk after
if: always()
run: df -h /workspace 2>/dev/null || df -h .
@@ -213,7 +229,9 @@ jobs:
| while read -r key; do git config --local --unset-all "$key"; done || true
git config --local lfs.url \
"https://x-access-token:${LFS_TOKEN}@gitea.tourolle.paris/dtourolle/DarkRoom.git/info/lfs"
git lfs pull
# The manual's pictures too: the APK carries the manual, and
# assemble-apk.sh refuses a pointer where a picture should be.
git lfs pull --exclude="fixtures/**"
ls -lR models/
- name: Cache cargo
@@ -322,7 +340,7 @@ jobs:
env:
CARGO_TARGET_DIR: target-android
# Absent secrets mean a debug signature, which is what a fork or a
# branch build should get. Set all three (see docs/android-signing.md)
# branch build should get. Set all three (see docs/dev/android-signing.md)
# and the same job produces a release-signed APK instead.
ANDROID_KEYSTORE_BASE64: ${{ secrets.ANDROID_KEYSTORE_BASE64 }}
KEYSTORE_PASS: ${{ secrets.ANDROID_KEYSTORE_PASSWORD }}
@@ -370,7 +388,7 @@ jobs:
# TRACES: FR-PLAT-WIN-3
# The Windows executable and its installer, cross-built from Linux
# (docs/windows.md §7). No Windows machine anywhere in this job: what it
# (docs/dev/windows.md §7). No Windows machine anywhere in this job: what it
# can prove is that the binary links, is a Windows executable with no
# MinGW runtime imports, starts under Wine, and that the installer installs
# and uninstalls under Wine. What it cannot prove — a Vulkan device, a
@@ -406,7 +424,9 @@ jobs:
| while read -r key; do git config --local --unset-all "$key"; done || true
git config --local lfs.url \
"https://x-access-token:${LFS_TOKEN}@gitea.tourolle.paris/dtourolle/DarkRoom.git/info/lfs"
git lfs pull
# The manual's pictures too: the installer carries the manual, and
# package.sh refuses a pointer where a picture should be.
git lfs pull --exclude="fixtures/**"
ls -l models/face models/scene
- name: Cache cargo
@@ -454,7 +474,17 @@ jobs:
wine "$SETUP" /S 2>/dev/null
INST=$(echo "$HOME"/.wine/drive_c/users/*/AppData/Local/Programs/DarkRoom)
ls "$INST"
[ "$(ls "$INST/models" | wc -l)" = 7 ] || { echo "FAIL: expected 7 model files"; exit 1; }
# As many files as package.sh stages: everything but the READMEs in
# the directories it copies. A literal here went stale the first
# time a model was added.
WANT=$(find models/face models/scene models/inpaint -maxdepth 1 -type f ! -name README.md | wc -l)
GOT=$(ls "$INST/models" | wc -l)
[ "$GOT" = "$WANT" ] || { echo "FAIL: expected $WANT model files, installed $GOT"; exit 1; }
# The manual, and every picture it shows, counted the same way.
[ -f "$INST/manual/index.html" ] || { echo "FAIL: no manual installed"; exit 1; }
WANT=$(ls docs/manual/media | wc -l)
GOT=$(ls "$INST/manual/media" | wc -l)
[ "$GOT" = "$WANT" ] || { echo "FAIL: expected $WANT manual pictures, installed $GOT"; exit 1; }
wine reg query 'HKCU\Software\Microsoft\Windows\CurrentVersion\Uninstall\DarkRoom' 2>/dev/null \
| grep -q DisplayVersion || { echo "FAIL: no uninstall registry key"; exit 1; }
wine "$INST/darkroom.exe" --version 2>/dev/null | grep -q '^darkroom-desktop ' \
@@ -514,3 +544,47 @@ jobs:
fi
done
exit $FAILED
# A v* tag becomes a Gitea Release carrying the three builds and their
# SHA256SUMS, titled and described by the tag's message. Until this job
# existed every release was made by hand, and most tags never got one.
#
# It needs all three platform jobs, so a tag whose tests fail publishes
# nothing; re-run the failed job and this one follows. The work is
# tools/publish-release.sh, which is also how a release is finished by hand.
release:
if: startsWith(github.ref, 'refs/tags/v')
needs: [desktop, android, windows]
runs-on: linux/amd64
name: Publish the release
container:
image: catthehacker/ubuntu:act-latest
permissions:
contents: write
steps:
- name: Checkout
uses: actions/checkout@v4
- name: Fetch the builds
uses: actions/download-artifact@v3
with:
path: dist
# Named for the download page, with the version in each name the way
# the hand-made releases had them. The installer already carries its
# version from package.sh.
- name: Publish
env:
GITEA_TOKEN: ${{ secrets.GITEA_TOKEN || github.token }}
TAG: ${{ github.ref_name }}
run: |
set -e
V="${TAG#v}"
ls -lR dist
mkdir -p out
cp dist/darkroom-arm64-v8a-apk/darkroom.apk "out/darkroom-${V}-arm64-v8a.apk"
cp dist/darkroom-desktop-x86_64-linux/darkroom-desktop "out/darkroom-desktop-${V}-x86_64-linux"
chmod +x "out/darkroom-desktop-${V}-x86_64-linux"
cp dist/darkroom-windows-x86_64-setup/DarkRoom-${V}-x86_64-setup.exe out/
bash tools/publish-release.sh "$TAG" out/*
+21 -6
View File
@@ -7,7 +7,7 @@ name: Traceability
# fail its own threshold. Two rules follow, and the extractor's own tests
# enforce both:
#
# 1. Denominators are parsed from docs/requirements.md at run time.
# 1. Denominators are parsed from docs/dev/requirements.md at run time.
# 2. Coverage is |traced ∩ defined| / |defined|, never a raw traced count.
#
# This job is static analysis of source comments plus markdown parsing, so it
@@ -67,6 +67,12 @@ jobs:
# threshold: zero requirements parsed, zero files scanned, a ratio above
# 100%, or any orphan tag all fail the build. A misconfigured run must not
# report a plausible-looking 0%.
# Every picture the manual shows is made by a scene in
# tools/manual/scenes.py, and every picture a scene makes is shown.
# Two files read; no app, no display.
- name: Manual pictures have scenes
run: tools/manual/record.sh --check
- name: Traceability gate
run: cargo run -q -p traceability -- check
@@ -74,11 +80,11 @@ jobs:
run: |
set -e
cargo run -q -p traceability -- report
if ! git diff --quiet docs/traceability.md; then
if ! git diff --quiet docs/dev/traceability.md; then
echo ""
echo "docs/traceability.md is out of date."
echo "docs/dev/traceability.md is out of date."
echo "Run: cargo run -p traceability -- report"
git diff --stat docs/traceability.md
git diff --stat docs/dev/traceability.md
exit 1
fi
@@ -91,10 +97,19 @@ jobs:
# they will conclude the application is broken rather than the page.
#
# This also fails on a malformed tag, so a typo costs a gesture its
# desktop half loudly rather than silently.
# desktop half loudly rather than silently — and on a key a Slint
# handler binds that no tag names, or a key a tag names that no handler
# binds (tools/traceability/src/keymap.rs).
- name: Regenerate the gesture vocabulary and check it is committed
run: cargo run -q -p traceability -- gestures-check
# The manual's page, which the packages carry and the help sheet links
# into. Blocking for the gesture book's reason: it is shown to the user,
# and a page that disagrees with the README is a manual describing an
# application that no longer exists.
- name: Regenerate the manual page and check it is committed
run: cargo run -q -p traceability -- manual-check
# Advisory, not blocking: not every file implements a requirement, and a
# tag on every function is noise that rots faster than it helps. Tag the
# unit that decides.
@@ -127,4 +142,4 @@ jobs:
- name: Summary
if: always()
run: head -30 docs/traceability.md || true
run: head -30 docs/dev/traceability.md || true
+18 -4
View File
@@ -26,7 +26,7 @@ fi
# The artefacts are generated from the tree, so regenerating them because one
# was itself edited would be circular.
case "$(tr -d '[:space:]' <<< "${staged}")" in
docs/traceability.md | docs/gestures.md | ui/dr-ui/src/gesture_book.rs)
docs/dev/traceability.md | docs/gestures.md | ui/dr-ui/src/gesture_book.rs | docs/manual/index.html)
exit 0
;;
esac
@@ -41,9 +41,9 @@ if ! cargo run -q -p traceability -- report >/dev/null 2>&1; then
exit 0
fi
if ! git diff --quiet -- docs/traceability.md; then
git add docs/traceability.md
echo "pre-commit: regenerated docs/traceability.md and staged it"
if ! git diff --quiet -- docs/dev/traceability.md; then
git add docs/dev/traceability.md
echo "pre-commit: regenerated docs/dev/traceability.md and staged it"
fi
# The gesture vocabulary, same discipline.
@@ -65,3 +65,17 @@ for f in docs/gestures.md ui/dr-ui/src/gesture_book.rs; do
echo "pre-commit: regenerated ${f} and staged it"
fi
done
# The manual's page, when its source is part of the commit. Rendered from
# nothing but the README, so there is no reason to pay for it otherwise.
if grep -qx 'docs/manual/README.md' <<< "${staged}"; then
if ! out="$(cargo run -q -p traceability -- manual 2>&1)"; then
echo "pre-commit: the manual would not render" >&2
echo "${out}" >&2
exit 1
fi
if ! git diff --quiet -- docs/manual/index.html; then
git add docs/manual/index.html
echo "pre-commit: regenerated docs/manual/index.html and staged it"
fi
fi
+1
View File
@@ -23,3 +23,4 @@ tools/film-profiles/upstream/
# checkout, so it is larger than the repository it sits in.
/.flatpak-builder/
/build/
__pycache__/
+132
View File
@@ -0,0 +1,132 @@
# Working in this repository
Notes for anyone — person or agent — changing this code. They record what
went wrong once and what the fix looked like, so the same shape is not
written again. Requirements live in `docs/dev/requirements.md`; this file is
about habits, not features.
## Catalog reads: work is proportional to what changed, never to library size
`docs/dev/catalog.md §1` states the rule. These are the ways it was broken on
the Identity screen, found when every confirm click cost half a second on a
24k-image library (2026-09-19), and what each fix looked like.
**A redraw must know what changed.** A click handler that calls "refresh
everything" pays for everything. `identity_ui::refresh` takes a `Changed`:
a confirm re-reads the rail and the grid and *not* the coverage line,
because moving a face between people cannot alter how many images are
indexed. Before adding a read to a shared refresh, ask which events can
change its answer, and gate it on those.
**Count with `COUNT(*)`, never with `.len()` on a list you then drop.**
`repairs::counts` used to build every repair's work list — a `Target` with
its path per row, sorted into visiting order — to report its length. Six
repairs, 350 ms, nothing kept. If the caller wants a number, the query
returns a number.
**One query, not one per row.** `ThumbStore::contains` in a filter over
5,000 rows is 5,000 prepared statements; `ThumbStore::held(size)` reads the
index once into a set. The same applies to any `query_row` inside a loop
over a result set — including `deep_count` per sidebar row, which is fine
at sidebar scale and would not be at grid scale. Aggregate in one
statement and look up in memory.
**Filter and aggregate in SQL, and aggregate the small side first.**
`faces::people` read 19,000 rows, grouped, sorted them by name, and the
screen threw 17,000 away (empty unnamed groups). `people_in_use` filters in
the `WHERE`, and joins `people` to a pre-aggregated `face_person` (2,000
groups) rather than grouping after a `LEFT JOIN` over every person. The
sort then sees only the rows that will be drawn.
**Wide rows make "just check one column" a table scan.** A `faces` row is
~8 KB (a 1 KB embedding and a ~5 KB crop, then the columns added later).
Any predicate that reads `quality`, `crop` or an eye column for every face
reads every row. V17 learned this for the eye filter; V19 applies it to the
repair counts with partial indexes (`faces_owed_*`) that hold only the rows
still owing, keyed on what the predicate joins on and carrying `model_id`
because the predicate reads it. Two things to know about them:
- **Drive the count from the small side.** SQLite uses a partial index
when the query starts from `faces` (`repairs::count`, `Needs::Face`) and
ignores it inside a correlated `EXISTS (... WHERE f.image_id = i.id ...)`.
That is why `Needs::Face` carries the per-face fragment and spells it two
ways.
- **Spell the predicate as the index's `WHERE` is spelled.** `NEEDS_EYES`
is `(f.eye_right IS NULL OR f.landmarks_dense IS NULL)` because
`faces_owed_eyes` is `WHERE eye_right IS NULL OR landmarks_dense IS NULL`.
Change one, change both, and `counts_are_the_sizes_of_the_lists` will
tell you if they drift.
Check a query's plan with `EXPLAIN QUERY PLAN` against a copy of a real
catalog before trusting an index exists for it: "SEARCH ... USING COVERING
INDEX" is the answer you want, "SEARCH f USING INDEX faces_image" on a wide
table means every probe opens a row.
## Catalog writes: one transaction per user action
`faces::confirm` opens a transaction. Calling it in a loop over a group is
a commit per face; `faces::confirm_all` is two statements and one commit,
`faces::reassign` one transaction for a whole split. When a UI action
touches N rows, give the catalog a function that takes the N, not a loop
that calls the one-row function N times — `unchecked_transaction` cannot
nest, so this has to be designed in at the catalog layer, not wrapped
from above.
## Screens: keep what is already decoded
`identity::load_faces` takes the crops the grid is currently showing and
hands them back into the new cells. Before that, a click re-read 4 MB of
crop blobs and decoded 700 JPEGs to produce the pixels already on screen.
When a redraw replaces a model, the expensive parts of the old model — a
decoded image, a cut portrait — are the first thing to reuse; only the row
that changed needs new work. Drain the old cells rather than cloning them.
## Remote calls: one round trip per file, not one per ancestor
`NextcloudBackend::move_to` guaranteed its destination's parent with a
`MKCOL` per ancestor from the account root, on every file of a batch —
three `405`s before each `MOVE`. The backend now remembers the collections
it has confirmed (`known_dirs`) for its lifetime, which is one job. When a
per-file operation has a per-batch precondition, satisfy it once.
## Providers: read the runtime's source for the version on disk, not the binding
Two things the MIGraphX rung (2026-09-20) got wrong before it was measured
right, both because `ort`'s builder was trusted to mean what its method
names say.
**A binding's option builder may fill a struct the runtime no longer
reads.** `ep::MIGraphX::with_save_model` sets fields of the legacy
`OrtMIGraphXProviderOptions`; ONNX Runtime 1.29 reads that struct for the
precision flags and ignores the rest, so every session compiled for 40 s
and the cache directory went nowhere. The option that works
(`migraphx_model_cache_dir`) exists only in the generic key/value
registration, which `session::migraphx` calls on the API table directly.
Before wiring a provider option, fetch the provider's source at the
runtime's exact version and find where the option is *read*.
**A provider's cache key may leave out what you are varying.** MIGraphX
keys a compiled program on graph, GPU and its own version — not precision.
The first fp16 measurement built in 0.3 s and matched f32 to the tenth of a
millisecond, because it had loaded the f32 program. A "from cache" build
that is suspiciously fast on the first run of a new configuration is a key
collision, not a fast provider; give each precision its own directory (the
engine does) and check the cache directory gained a file.
## Measuring
`cargo run --release -p dr-ui --example identity_bench -- CATALOG THUMBS`
times what one click on the Identity screen reads and what the batch
operations write. Run it against a **copy** of a real catalog (it writes),
never the library's own file; `sqlite3 catalog.sqlite ".backup copy.sqlite"`
takes a consistent one while the app runs. Compare the `cpu` column when
other builds are running on the machine — the wall clock doubles under
load, the CPU figure does not. Keep the binary from before the change and
run both back to back rather than trusting numbers taken an hour apart.
Reference figures from the 2026-09-19 fixes, largest person (754 faces),
24k images, 19k faces, before → after. What one click read: `load_people`
22 ms → 12 ms, `load_faces` 316 ms → 2.4 ms, `audit` 190 ms → not run
(66 ms when it is, on open and at the end of a sweep). What one click
wrote: `confirm_all` 16 ms → 2 ms, `split_off` 23 ms → 4.5 ms. A click on
the face grid went from ~530 ms of catalog work to ~15 ms.
+11 -11
View File
@@ -90,18 +90,18 @@ break it by accident:
cargo run --release -p dr-bench -- check
```
That is the benchmark suite (`docs/requirements.md` §8), which builds a
That is the benchmark suite (`docs/dev/requirements.md` §8), which builds a
synthetic 50,000-image catalog and fails the build if a performance target is
missed or a measurement has drifted past its tolerance. It runs on every push in
its own workflow. [`docs/benchmarks.md`](docs/benchmarks.md) says what it
its own workflow. [`docs/dev/benchmarks.md`](docs/dev/benchmarks.md) says what it
measures, what it deliberately does not, and how to read a failure. If you have
touched the catalog, the decoder, the thumbnail store or the exporter, run it
before you send.
## Requirements and traceability
[`requirements.md`](docs/requirements.md) is the register of record.
[`traceability.md`](docs/traceability.md) is generated from `TRACES:` tags in
[`requirements.md`](docs/dev/requirements.md) is the register of record.
[`traceability.md`](docs/dev/traceability.md) is generated from `TRACES:` tags in
the source and must never be hand-edited:
```rust
@@ -124,7 +124,7 @@ Note that it tracks line numbers, so a change that only moves code still moves
the matrix. Never regenerate it with a stale prebuilt binary.
**One convention that the tooling cannot enforce.** A tag proves that a tag
exists, not that the code under it does the thing — `docs/code-health.md`
exists, not that the code under it does the thing — `docs/dev/code-health.md`
CH-4 has the details, and two requirements currently read as covered on the
strength of plumbing a future feature would use. So: **close a requirement
with a test that would fail if the behaviour were removed.** Coverage that
@@ -163,12 +163,12 @@ One commit per change. If you fixed two things, that is two commits.
| Document | Read it when |
|---|---|
| [`core/dr-pipeline/ops/README.md`](core/dr-pipeline/ops/README.md) | Adding or changing a develop operation — start here regardless |
| [`docs/architecture.md`](docs/architecture.md) | Anything touching the render path, catalog or sync |
| [`docs/code-health.md`](docs/code-health.md) | Deciding what to work on; grades each seam by what it costs |
| [`docs/benchmarks.md`](docs/benchmarks.md) | A change that could plausibly cost time or memory |
| [`docs/technical-debt.md`](docs/technical-debt.md) | Something looks wrong — check it was not chosen |
| [`docs/distribution.md`](docs/distribution.md) | Packaging a build, or adding a permission to one |
| [`docs/requirements.md`](docs/requirements.md) | Reference, not reading |
| [`docs/dev/architecture.md`](docs/dev/architecture.md) | Anything touching the render path, catalog or sync |
| [`docs/dev/code-health.md`](docs/dev/code-health.md) | Deciding what to work on; grades each seam by what it costs |
| [`docs/dev/benchmarks.md`](docs/dev/benchmarks.md) | A change that could plausibly cost time or memory |
| [`docs/dev/technical-debt.md`](docs/dev/technical-debt.md) | Something looks wrong — check it was not chosen |
| [`docs/dev/distribution.md`](docs/dev/distribution.md) | Packaging a build, or adding a permission to one |
| [`docs/dev/requirements.md`](docs/dev/requirements.md) | Reference, not reading |
`technical-debt.md` is the one to check before "fixing" anything surprising.
It records compromises that were deliberate, each with the reasoning and a
Generated
+74 -29
View File
@@ -1221,7 +1221,7 @@ checksum = "f27ae1dd37df86211c42e150270f82743308803d90a6f6e6651cd730d5e1732f"
[[package]]
name = "darkroom-android"
version = "0.12.1"
version = "0.15.0"
dependencies = [
"android_logger",
"dr-plat",
@@ -1234,7 +1234,7 @@ dependencies = [
[[package]]
name = "darkroom-desktop"
version = "0.12.1"
version = "0.15.0"
dependencies = [
"anyhow",
"dr-plat",
@@ -1408,7 +1408,7 @@ checksum = "d8b14ccef22fc6f5a8f4d7d768562a182c04ce9a3b3157b91390b52ddfdf1a76"
[[package]]
name = "dr-bench"
version = "0.12.1"
version = "0.15.0"
dependencies = [
"anyhow",
"dr-catalog",
@@ -1425,7 +1425,7 @@ dependencies = [
[[package]]
name = "dr-catalog"
version = "0.12.1"
version = "0.15.0"
dependencies = [
"dr-face",
"dr-plat",
@@ -1440,7 +1440,7 @@ dependencies = [
[[package]]
name = "dr-decode"
version = "0.12.1"
version = "0.15.0"
dependencies = [
"dr-types",
"env_logger",
@@ -1454,7 +1454,7 @@ dependencies = [
[[package]]
name = "dr-export"
version = "0.12.1"
version = "0.15.0"
dependencies = [
"dr-decode",
"dr-gpu",
@@ -1465,6 +1465,7 @@ dependencies = [
"log",
"png",
"pollster",
"rawler",
"thiserror 2.0.20",
"tiff",
"zune-jpeg 0.4.21",
@@ -1472,20 +1473,20 @@ dependencies = [
[[package]]
name = "dr-face"
version = "0.12.1"
version = "0.15.0"
dependencies = [
"dr-inference-engine",
"env_logger",
"log",
"ndarray",
"ort",
"ort-tract",
"thiserror 2.0.20",
"zune-jpeg 0.4.21",
]
[[package]]
name = "dr-film"
version = "0.12.1"
version = "0.15.0"
dependencies = [
"log",
"serde",
@@ -1494,11 +1495,12 @@ dependencies = [
[[package]]
name = "dr-gpu"
version = "0.12.1"
version = "0.15.0"
dependencies = [
"bytemuck",
"dr-decode",
"dr-film",
"dr-pano",
"dr-pipeline",
"dr-segment",
"dr-types",
@@ -1509,9 +1511,24 @@ dependencies = [
"wgpu",
]
[[package]]
name = "dr-inference-engine"
version = "0.15.0"
dependencies = [
"env_logger",
"libloading",
"log",
"ort",
"ort-sys",
"ort-tract",
"serde",
"serde_json",
"thiserror 2.0.20",
]
[[package]]
name = "dr-ingest"
version = "0.12.1"
version = "0.15.0"
dependencies = [
"dr-plat",
"dr-types",
@@ -1523,15 +1540,29 @@ dependencies = [
[[package]]
name = "dr-lens"
version = "0.12.1"
version = "0.15.0"
dependencies = [
"lensfun",
"log",
]
[[package]]
name = "dr-pano"
version = "0.15.0"
dependencies = [
"dr-decode",
"dr-inference-engine",
"dr-types",
"env_logger",
"log",
"ndarray",
"ort",
"thiserror 2.0.20",
]
[[package]]
name = "dr-pipeline"
version = "0.12.1"
version = "0.15.0"
dependencies = [
"dr-types",
"log",
@@ -1540,7 +1571,7 @@ dependencies = [
[[package]]
name = "dr-plat"
version = "0.12.1"
version = "0.15.0"
dependencies = [
"android-native-keyring-store",
"dr-types",
@@ -1556,7 +1587,7 @@ dependencies = [
[[package]]
name = "dr-preset-xmp"
version = "0.12.1"
version = "0.15.0"
dependencies = [
"dr-pipeline",
"log",
@@ -1566,20 +1597,20 @@ dependencies = [
[[package]]
name = "dr-segment"
version = "0.12.1"
version = "0.15.0"
dependencies = [
"dr-inference-engine",
"env_logger",
"log",
"ndarray",
"ort",
"ort-tract",
"thiserror 2.0.20",
"zune-jpeg 0.4.21",
]
[[package]]
name = "dr-sync"
version = "0.12.1"
version = "0.15.0"
dependencies = [
"async-trait",
"dr-plat",
@@ -1593,7 +1624,7 @@ dependencies = [
[[package]]
name = "dr-sync-folder"
version = "0.12.1"
version = "0.15.0"
dependencies = [
"async-trait",
"dr-sync",
@@ -1605,7 +1636,7 @@ dependencies = [
[[package]]
name = "dr-sync-nextcloud"
version = "0.12.1"
version = "0.15.0"
dependencies = [
"async-trait",
"dr-decode",
@@ -1627,7 +1658,7 @@ dependencies = [
[[package]]
name = "dr-thumbs"
version = "0.12.1"
version = "0.15.0"
dependencies = [
"dr-types",
"jpeg-encoder",
@@ -1639,7 +1670,7 @@ dependencies = [
[[package]]
name = "dr-types"
version = "0.12.1"
version = "0.15.0"
dependencies = [
"serde",
"serde_json",
@@ -1648,7 +1679,7 @@ dependencies = [
[[package]]
name = "dr-ui"
version = "0.12.1"
version = "0.15.0"
dependencies = [
"anyhow",
"async-trait",
@@ -1658,8 +1689,10 @@ dependencies = [
"dr-face",
"dr-film",
"dr-gpu",
"dr-inference-engine",
"dr-ingest",
"dr-lens",
"dr-pano",
"dr-pipeline",
"dr-plat",
"dr-preset-xmp",
@@ -1671,9 +1704,11 @@ dependencies = [
"dr-types",
"dr-xmp",
"env_logger",
"i-slint-backend-testing",
"jni 0.22.4",
"log",
"ndk-context",
"png",
"pollster",
"reqwest",
"rusqlite",
@@ -1683,12 +1718,13 @@ dependencies = [
"slint-build",
"thiserror 2.0.20",
"tokio",
"url",
"wgpu",
]
[[package]]
name = "dr-xmp"
version = "0.12.1"
version = "0.15.0"
dependencies = [
"dr-types",
"log",
@@ -2746,6 +2782,18 @@ dependencies = [
"i-slint-renderer-skia",
]
[[package]]
name = "i-slint-backend-testing"
version = "1.17.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "521e901e3d47ab829c0ef500c63155776208707cd93259e6a7803ed627fa2786"
dependencies = [
"cfg_aliases",
"i-slint-common",
"i-slint-core",
"vtable",
]
[[package]]
name = "i-slint-backend-winit"
version = "1.17.1"
@@ -2926,8 +2974,6 @@ dependencies = [
[[package]]
name = "i-slint-renderer-skia"
version = "1.17.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "7b6eed7f3f0a9a3d3ca6e8b9d4ca233371d989351fdb2a7ab88ec368b99e7b57"
dependencies = [
"ash",
"bytemuck",
@@ -6989,9 +7035,10 @@ checksum = "8df9b6e13f2d32c91b9bd719c00d1958837bc7dec474d94952798cc8e69eeec3"
[[package]]
name = "traceability"
version = "0.12.1"
version = "0.15.0"
dependencies = [
"anyhow",
"pulldown-cmark",
"serde",
"serde_json",
]
@@ -7889,8 +7936,6 @@ dependencies = [
[[package]]
name = "wgpu-hal"
version = "29.0.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "97ace1c17727311c22a46e4e3faf56ea6de81af99dcc839bdfb54857b94d448d"
dependencies = [
"android_system_properties",
"arrayvec",
+24 -1
View File
@@ -8,9 +8,11 @@ members = [
"core/dr-export",
"core/dr-face",
"core/dr-film",
"core/dr-inference-engine",
"core/dr-ingest",
"core/dr-gpu",
"core/dr-lens",
"core/dr-pano",
"core/dr-pipeline",
"core/dr-preset-xmp",
"core/dr-segment",
@@ -25,9 +27,12 @@ members = [
"tools/bench",
"tools/traceability",
]
# Patched copies of upstream crates, not our code: see third_party/README.md.
# Excluded so `--workspace` does not test, lint or format them as ours.
exclude = ["third_party"]
[workspace.package]
version = "0.12.1"
version = "0.15.0"
edition = "2021"
rust-version = "1.92"
license = "GPL-3.0-or-later"
@@ -45,9 +50,14 @@ dr-export = { path = "core/dr-export" }
# `features = ["inference"]`.
dr-face = { path = "core/dr-face", default-features = false }
dr-film = { path = "core/dr-film" }
# `tract` on by default so a test binary can open a session with nothing
# installed; the apps add `native` to look for a runtime file (docs/dev/inference.md §3).
dr-inference-engine = { path = "core/dr-inference-engine" }
dr-ingest = { path = "core/dr-ingest" }
dr-gpu = { path = "core/dr-gpu" }
dr-lens = { path = "core/dr-lens" }
# Optional runtime, like `dr-segment`: the geometry never needs a model.
dr-pano = { path = "core/dr-pano", default-features = false }
dr-pipeline = { path = "core/dr-pipeline" }
dr-preset-xmp = { path = "core/dr-preset-xmp" }
# `default-features = false` belongs *here*, not on each dependant: a member
@@ -120,6 +130,10 @@ url = "2.5"
async-trait = "0.1"
serde = { version = "1", features = ["derive"] }
serde_json = "1"
# The manual's HTML rendering (tools/traceability). Already in the tree as
# Slint's Markdown parser, so this adds a dependency edge and no crate; only
# the HTML writer is needed, not the command-line front end.
pulldown-cmark = { version = "0.13", default-features = false, features = ["html"] }
base64 = "0.23"
# Display-server clients, for FR-DSP-8's per-display profile acquisition.
@@ -255,3 +269,12 @@ opt-level = 0
[profile.release]
lto = "thin"
codegen-units = 1
# Two upstream crates carry a local patch so that the Android build can draw
# with wgpu on a rotated display (technical-debt.md TD-1). Both are exact
# copies of the version the lockfile already resolves, plus that patch;
# third_party/README.md says what was changed and how to carry it forward
# when Slint or wgpu moves.
[patch.crates-io]
wgpu-hal = { path = "third_party/wgpu-hal-29.0.4" }
i-slint-renderer-skia = { path = "third_party/i-slint-renderer-skia-1.17.1" }
+109 -59
View File
@@ -1,84 +1,134 @@
# DarkRoom
A cross-platform, non-destructive RAW photo editor for Linux and Android.
A non-destructive RAW photo editor and library for Linux and Android, with a
GPU develop pipeline, a catalog that syncs between devices, and no account,
no telemetry and no cloud of its own.
**Status:** 0.9.0, and no longer a spike. A library opens, culls, develops and
exports on both platforms, across eight tagged releases. What is *not*
built is written down rather than merely absent — see
[docs/outstanding.md](docs/outstanding.md) for the requirements that have no
implementation and why, and [docs/technical-debt.md](docs/technical-debt.md)
for the compromises that were chosen.
[![The library: seventy frames, the timeline beside them, the filter bar above](docs/manual/media/library.png)](docs/manual/README.md)
## Documentation
**[The manual](docs/manual/README.md)** shows every feature, pictured from
the application itself. This page says what it is, how to get it, and what
is still missing.
| Document | Contents |
|---|---|
| [CONTRIBUTING.md](CONTRIBUTING.md) | How to land a first change without reading the rest |
| [requirements.md](docs/requirements.md) | What the software must do — 179 numbered requirements |
| [architecture.md](docs/architecture.md) | How it is built — crates, GPU pipeline, data model, sync |
| [technical-debt.md](docs/technical-debt.md) | Compromises taken deliberately, each with the condition that retires it |
| [outstanding.md](docs/outstanding.md) | What is not built, and whether that is a decision or a gap |
| [code-health.md](docs/code-health.md) | What a contribution costs, per seam, measured |
| [traceability.md](docs/traceability.md) | Generated: which requirement is claimed by which file |
| [faces.md](docs/faces.md) | Face detection and identity — the models, the licence problem, and what S14 measured |
## What it does
## Building
**A library.** Point it at a folder — on this machine, on a network mount,
or one a Nextcloud client keeps in virtual-files mode, where a placeholder
is treated as the photograph rather than as a one-byte file — or at a
Nextcloud account directly. The grid is virtualised, ordered by capture
time with a timeline beside it, and filtered by rating, flag, person and
whether the file is here. Ratings, keywords, collections and a
trash that survives a crash mid-operation. Card ingest. Bursts fold. Face
detection and identity, with the index syncing between devices.
Desktop:
**Developing.** Eighteen declared operations fused into one compute
dispatch, plus the neighbourhood work that cannot be: clarity, texture,
capture sharpening, noise reduction, lens correction, spectral film
simulation. Crop and straighten, spot repair, and local adjustments over
masks the model draws — click a subject or a category, then paint, subtract
a gradient, grow or shrink the edge. Focus peaking and a raw histogram for
judging what is recoverable. Named presets; XMP sidecars other editors read.
[![Segmenting an urban scene and choosing the sky as a mask](docs/manual/media/local-segment.png)](docs/manual/README.md#local-adjustments)
**Panoramas.** Select the frames, align, choose a projection, fill the
ragged border rather than crop it, and the composite lands beside its
sources as a DNG, with a sidecar recording what it was merged from.
[![Twelve hand-held frames aligned on a cylinder](docs/manual/media/panorama-aligned.png)](docs/manual/README.md#merging-a-panorama)
**Export.** JPEG, PNG, AVIF, JPEG XL, 8- and 16-bit TIFF, with resize, output
sharpening, a naming template and a colour space — to a folder here or back
into the library.
**On both platforms.** The same core runs on a desktop and a 12-inch
tablet; the interface is one layout, tuned for a wide viewport with touch
targets throughout. On desktop the develop view draws the compute pass's
texture directly — no readback between the GPU and the screen.
## Getting it
| Platform | How | State |
|---|---|---|
| Arch Linux | [`packaging/PKGBUILD`](packaging/PKGBUILD) — `makepkg -si` | Built from every release |
| Android | The APK from each CI run, or `./docker/android/package.sh --install` | Runs on a tablet; F-Droid not yet submitted |
| Windows | `DarkRoom-<version>-x86_64-setup.exe`, cross-built by CI ([windows.md](docs/dev/windows.md)) | Verified under Wine only; unsigned |
| Flatpak | [`packaging/flatpak/`](packaging/flatpak/) | Manifest in tree; choosing a library does not yet work in the sandbox |
Or build it. Git LFS is required for the model weights, and the toolchain
pins itself to 1.92.0:
```bash
cargo run -p darkroom-desktop
git lfs install && git lfs pull
cargo run --release -p darkroom-desktop
```
Android (containerised toolchain, see [docker/android](docker/android/README.md)):
Android, through the containerised toolchain ([docker/android](docker/android/README.md)):
```bash
./docker/android/build.sh cargo ndk -t arm64-v8a build --release
```
Git LFS is required for the model weights, and the toolchain pins itself.
[CONTRIBUTING.md](CONTRIBUTING.md) has the details and the four commands CI
will run against what you send.
[CONTRIBUTING.md](CONTRIBUTING.md) has the system packages, the four
commands CI runs against what you send, and the shortest useful
contribution — a develop operation is one YAML file, and it arrives with its
controls, its place in the chain and its tests.
## Current state
## Where it stands
**Working.** A catalog over a local folder, a Nextcloud account, or a folder a
sync client keeps in virtual-files mode — where a placeholder is treated as the
photograph rather than as a one-byte file. A virtualised library grid with a
capture-time timeline, ratings, labels, keywords, collections and a trash that
survives a crash mid-operation. Card ingest. Face detection and identity, with
the index syncing between devices. A develop pipeline of fifteen declared
operations fused into a single compute dispatch, plus the neighbourhood
operations that cannot be — clarity, texture, capture sharpening, noise
reduction, lens correction, spectral film simulation. Crop, straighten, spot
removal, gradient and subject-segmentation masks, named presets, and a
generated panel that no operation in `ui/` is allowed to name. Export to JPEG,
PNG and 8- or 16-bit TIFF with resize and output sharpening.
**0.15.0**, twenty-three tagged releases in. 190 numbered requirements in
scope, 84% of them claimed by code and [traced to it](docs/dev/traceability.md);
the rest are written down rather than merely absent.
**The zero-copy display path works on desktop.** The compute pass writes a
texture that Slint composites directly, which is what
[ARCH §6.1](docs/architecture.md) requires; the readback it forbids costs 96%
of frame time at 4K, and
**Not built:** plugins (post-v1, [D12](docs/dev/requirements.md)), compare and
survey culling, AI denoise, tiled rendering, HDR merge and
focus stacking, most of the Android platform integration beyond running,
and the Flatpak's library chooser. The performance targets are half
verified: the per-commit benchmark suite §8 requires exists for everything
that does not need a frame — the catalog, the scan, the thumbnails — and
not yet for the render path, so a regression there fails nothing.
[outstanding.md](docs/dev/outstanding.md) is the list, with the reasoning for
each.
```bash
cargo run -p dr-gpu --example bench --features readback
```
**The one deliberate compromise worth knowing about before reading
anything else:** the Android develop view reads its frame back through the
CPU, because zero-copy there needs wgpu's Vulkan swapchain and that tears a
portrait window on a tablet whose panel is mounted landscape. It is debt,
not a revision of the rule — [technical-debt.md TD-1](docs/dev/technical-debt.md)
has the measurements and the three things any one of which would remove it.
still reproduces that measurement. **The one exception is the Android develop
view**, which reads the frame back through the CPU because zero-copy there
needs wgpu's Vulkan swapchain, and that tears a portrait window on a tablet
whose panel is mounted landscape. It is debt, not a revision of the rule: the
reasoning, the on-device measurements that forced it, and the three separate
things any one of which would remove it are in
[technical-debt.md TD-1](docs/technical-debt.md).
## Documentation
**Not built.** Plugins, compare and survey culling, focus peaking, burst
grouping, AI denoise, tiled and progressive rendering, and most of the Android
platform integration beyond running. The performance targets in §4.1 are
unverified rather than unmet — the per-commit benchmark suite §8 requires does
not exist, so nothing fails a build on a regression.
[docs/outstanding.md](docs/outstanding.md) is the list, with the reasoning.
[docs/README.md](docs/README.md) is the index. The short version, for someone using it:
| | |
|---|---|
| [manual](docs/manual/README.md) | Every feature, pictured |
| [gestures.md](docs/gestures.md) | How it is driven — generated from the code, so it cannot describe a gesture that does not exist |
For someone changing it:
| | |
|---|---|
| [CONTRIBUTING.md](CONTRIBUTING.md) | How to land a first change without reading the rest |
| [requirements.md](docs/dev/requirements.md) | What the software must do — the numbered register, and the decisions |
| [architecture.md](docs/dev/architecture.md) | How it is built — crates, the GPU pipeline, the data model, sync |
| [technical-debt.md](docs/dev/technical-debt.md) | Compromises taken deliberately, each with the condition that retires it |
| [outstanding.md](docs/dev/outstanding.md) | What is not built, and whether that is a decision or a gap |
| [code-health.md](docs/dev/code-health.md) | What a contribution costs, per seam, measured |
| [traceability.md](docs/dev/traceability.md) | Generated: which requirement is claimed by which file |
Designs, one per subsystem:
[segmentation](docs/dev/segmentation.md) and [mask editing](docs/dev/mask-editing.md) ·
[spot removal](docs/dev/spot-removal.md) · [panorama](docs/dev/panorama.md) ·
[faces](docs/dev/faces.md) · [inference](docs/dev/inference.md) ·
[storage and sync](docs/dev/storage.md) · [catalog](docs/dev/catalog.md) ·
[display and extension](docs/dev/display-and-extension.md) ·
[navigation](docs/dev/ui-navigation.md) · [distribution](docs/dev/distribution.md) ·
[windows](docs/dev/windows.md) · [benchmarks](docs/dev/benchmarks.md).
## Licence
GPL-3.0-or-later.
GPL-3.0-or-later. The photographs in the manual and the test fixtures are
the author's and are there to show and test this project, nothing else.
The model weights carry their own licences — [models/LICENCE.md](models/LICENCE.md).
@@ -141,6 +141,23 @@
</intent-filter>
</activity>
<!-- The manual (dr_ui::manual): a WebView over the copy the APK
carries in assets/manual. See ManualActivity.java for why it is
not the browser.
Not exported: nothing outside this app has a reason to start it,
and dr_ui starts it by class name, which needs no intent filter.
Its own task entry is not wanted either — it is a page over the
app, and Back returns to the photograph it was opened from.
configChanges so a rotation reflows the page rather than
reloading it at the top. -->
<activity
android:name="paris.tourolle.darkroom.ManualActivity"
android:exported="false"
android:label="DarkRoom manual"
android:theme="@style/ManualTheme"
android:configChanges="orientation|keyboardHidden|screenSize|screenLayout|uiMode" />
<!-- FR-PLAT-AND-6, outbound. Android has refused file:// URIs
between apps since API 24 — handing one out raises
FileUriExposedException in *this* process — so an exported JPEG
@@ -0,0 +1,105 @@
package paris.tourolle.darkroom;
import android.app.Activity;
import android.content.ActivityNotFoundException;
import android.content.Intent;
import android.net.Uri;
import android.os.Bundle;
import android.webkit.WebResourceRequest;
import android.webkit.WebSettings;
import android.webkit.WebView;
import android.webkit.WebViewClient;
/**
* The manual that ships in the APK, shown in a WebView.
*
* <h2>Why an activity of our own rather than the browser</h2>
*
* <p>The desktop hands the manual to the system browser. Android leaves no
* way to do the same: the page is an asset inside the APK, which is not a
* file; an unpacked copy in app-private storage is a file no browser may
* read; a {@code file:} URI handed to another app is refused since API 24;
* and a {@code content:} URI serves the page but leaves the browser to fetch
* every picture by a relative URL against the provider, which browsers do not
* reliably do. A WebView reads {@code file:///android_asset/} straight from
* the APK, pictures and section anchor included, and nothing is unpacked.
*
* <h2>What it is not</h2>
*
* <p>A browser. JavaScript stays off (the page has none), and a link that
* leaves the manual — the design documents are on the forge — goes to the
* user's browser rather than opening inside this view, so the only thing ever
* shown here is the page the APK carries.
*
* <p>Started by {@code dr_ui::manual} with {@code Intent.setClassName}, so the
* name here and there must agree; a test in lib.rs checks the manifest
* declares it.
*/
public final class ManualActivity extends Activity {
/** The section to open at, a heading's anchor. Absent opens the top. */
public static final String EXTRA_ANCHOR = "anchor";
private static final String PAGE = "file:///android_asset/manual/index.html";
private WebView web;
@Override
protected void onCreate(Bundle saved) {
super.onCreate(saved);
setTitle("DarkRoom manual");
web = new WebView(this);
WebSettings settings = web.getSettings();
settings.setJavaScriptEnabled(false);
// Pinch to zoom into a screenshot, which is 1600 pixels wide and drawn
// at the width of a phone.
settings.setBuiltInZoomControls(true);
settings.setDisplayZoomControls(false);
web.setWebViewClient(new WebViewClient() {
@Override
public boolean shouldOverrideUrlLoading(WebView view, WebResourceRequest request) {
Uri uri = request.getUrl();
if ("file".equals(uri.getScheme())) {
return false;
}
try {
startActivity(new Intent(Intent.ACTION_VIEW, uri));
} catch (ActivityNotFoundException e) {
// No browser on the device: the link does nothing, which
// is all it could do.
}
return true;
}
});
setContentView(web);
if (saved != null) {
web.restoreState(saved);
} else {
String anchor = getIntent().getStringExtra(EXTRA_ANCHOR);
web.loadUrl(anchor == null || anchor.isEmpty() ? PAGE : PAGE + "#" + anchor);
}
}
@Override
protected void onSaveInstanceState(Bundle out) {
super.onSaveInstanceState(out);
web.saveState(out);
}
/** Back walks back through the sections visited, then leaves. */
@Override
public void onBackPressed() {
if (web.canGoBack()) {
web.goBack();
} else {
super.onBackPressed();
}
}
@Override
protected void onDestroy() {
web.destroy();
super.onDestroy();
}
}
@@ -0,0 +1,5 @@
<?xml version="1.0" encoding="utf-8"?>
<!-- Day or night as the system is; see values/themes.xml. -->
<resources>
<style name="ManualTheme" parent="@android:style/Theme.DeviceDefault.DayNight" />
</resources>
@@ -0,0 +1,10 @@
<?xml version="1.0" encoding="utf-8"?>
<!--
The manual's theme (ManualActivity). Light below API 29, which has no
day-night theme in the platform; values-v29 follows the system from there.
The WebView takes prefers-color-scheme from whether this theme is light, and
the manual's stylesheet takes its colours from that.
-->
<resources>
<style name="ManualTheme" parent="@android:style/Theme.DeviceDefault.Light" />
</resources>
+80 -5
View File
@@ -240,7 +240,7 @@ fn android_main(app: slint::android::AndroidApp) {
///
/// **Face weights are absent from the repository by design.** The InsightFace
/// grant is research-only and incompatible with this project's licence
/// (docs/faces.md §2), so a desktop user fetches them, runs
/// (docs/dev/faces.md §2), so a desktop user fetches them, runs
/// `tools/fix-face-model-shapes.sh` over them, and drops the result in. A build
/// that carries none is the ordinary case and face indexing simply stays off.
///
@@ -321,20 +321,43 @@ fn unpack_bundled_models(app: &slint::android::AndroidApp) {
// before it reports the tab available.
//
// Three detectors, because which one runs is a setting
// (`FaceDetector`, docs/faces.md §12.3) and a tablet has no other way to
// (`FaceDetector`, docs/dev/faces.md §12.3) and a tablet has no other way to
// obtain the one it was not shipped with. Twenty megabytes of APK for
// the choice; the embedder is the same for all three.
const BUNDLED: [(&std::ffi::CStr, &str); 7] = [
//
// Then the three eye-state models (docs/dev/faces.md §17): landmarks, open
// or closed, sunglasses. The app indexes without them; with them the
// eyes-open filter has something to read, and a tablet has no other way
// to get them either.
//
// The int8 forms beside the three detectors are what the Hexagon runs
// (docs/dev/inference.md §5); the engine loads the sibling when the probe
// chose that rung and ignores it otherwise.
const BUNDLED: [(&std::ffi::CStr, &str); 14] = [
(c"models/scrfd_500m_640.onnx", "scrfd_500m_640.onnx"),
(
c"models/scrfd_500m_640.int8.onnx",
"scrfd_500m_640.int8.onnx",
),
(c"models/scrfd_2.5g_640.onnx", "scrfd_2.5g_640.onnx"),
(
c"models/scrfd_2.5g_640.int8.onnx",
"scrfd_2.5g_640.int8.onnx",
),
(c"models/scrfd_10g_640.onnx", "scrfd_10g_640.onnx"),
(c"models/scrfd_10g_640.int8.onnx", "scrfd_10g_640.int8.onnx"),
(c"models/arcface_mbf_b1.onnx", "arcface_mbf_b1.onnx"),
(c"models/2d106det_b1.onnx", "2d106det_b1.onnx"),
(c"models/ocec_s_b1.onnx", "ocec_s_b1.onnx"),
(c"models/sgc_l_48_b1.onnx", "sgc_l_48_b1.onnx"),
(c"models/yolo26s-sem-ade20k.onnx", "yolo26s-sem-ade20k.onnx"),
(
c"models/yolo26s-sem-ade20k.classes.json",
"yolo26s-sem-ade20k.classes.json",
),
(c"models/categories.txt", "categories.txt"),
// The panorama border filler (FR-MRG-4); MIT, 28 MB.
(c"models/migan-512.onnx", "migan-512.onnx"),
];
let dir = dr_ui::shared_face_models_dir();
@@ -343,8 +366,8 @@ fn unpack_bundled_models(app: &slint::android::AndroidApp) {
for (asset_path, name) in BUNDLED {
let dest = dir.join(name);
// Already unpacked. Not re-read on every launch: this is 61 MB of
// copying across the seven entries, and the file does not change without
// Already unpacked. Not re-read on every launch: this is 73 MB of
// copying across the ten entries, and the file does not change without
// the APK changing, at which point the install wiped it anyway. It
// matters more now than it did — a launch that skips every entry here
// costs nothing at all, which is what makes the second launch after an
@@ -392,6 +415,27 @@ fn unpack_bundled_models(app: &slint::android::AndroidApp) {
"bundled models ready: {copied} bytes copied in {} ms",
started.elapsed().as_millis()
);
// Now, and not at launch: the probe fingerprints the model files, and
// on a first launch they were not on disk until this line. The runtime
// is in the APK's native library directory beside `libdarkroom.so`,
// which is also where Qualcomm's DSP loader has to be pointed for the
// Hexagon skel (docs/dev/inference.md §3, §8).
dr_ui::inference::init(native_library_dir().into_iter().collect());
}
/// The directory the system unpacked this APK's native libraries into.
///
/// Read from where the loader put *this* library rather than asked of the
/// activity: `android-activity` does not expose `nativeLibraryDir`, and the
/// answer is in `/proc/self/maps` for free.
#[cfg(target_os = "android")]
fn native_library_dir() -> Option<std::path::PathBuf> {
let maps = std::fs::read_to_string("/proc/self/maps").ok()?;
maps.lines()
.filter_map(|l| l.split_whitespace().nth(5))
.find(|p| p.ends_with("/libdarkroom.so"))
.and_then(|p| std::path::Path::new(p).parent().map(Into::into))
}
/// TRACES: FR-PLAT-AND-6
@@ -537,6 +581,37 @@ mod tests {
);
}
/// `dr_ui::manual` starts the manual by class name. A name the manifest
/// does not declare is an `ActivityNotFoundException` on the device and a
/// Manual button that does nothing, so the three spellings — dr_ui's, the
/// manifest's and the Java file's — are checked to be one.
#[test]
fn the_manual_activity_dr_ui_starts_is_declared() {
let manifest = manifest();
let wanted = dr_ui::manual::ANDROID_ACTIVITY;
let element = manifest
.split("<activity")
.skip(1)
.find(|a| attribute(a, "android:name").as_deref() == Some(wanted))
.unwrap_or_else(|| panic!("the manifest declares no activity {wanted}"));
assert_eq!(
attribute(element, "android:exported").as_deref(),
Some("false"),
"the manual activity has no reason to be startable by another app"
);
let java = include_str!("../android/java/paris/tourolle/darkroom/ManualActivity.java");
let (package, class) = wanted.rsplit_once('.').expect("unqualified class name");
assert!(java.contains(&format!("package {package};")));
assert!(java.contains(&format!("class {class} ")));
assert!(
java.contains(&format!(
"EXTRA_ANCHOR = \"{}\"",
dr_ui::manual::ANDROID_EXTRA_ANCHOR
)),
"ManualActivity reads the section from a different extra than dr_ui writes"
);
}
#[test]
fn the_provider_hands_out_one_file_at_a_time_and_nothing_by_itself() {
let manifest = manifest();
+3
View File
@@ -25,3 +25,6 @@ winresource = "0.1"
[features]
default = []
# The manual's recording hook (dr-ui's `automation`); tools/manual/record.sh
# builds with it, nothing else does.
automation = ["dr-ui/automation"]
+42 -1
View File
@@ -21,7 +21,7 @@ fn main() -> anyhow::Result<()> {
// build made on a machine that cannot run the application — the Linux CI
// producing the Windows binary, checked under Wine — has an exit that
// proves the executable starts without opening a window or touching the
// user's directories (docs/windows.md §6).
// user's directories (docs/dev/windows.md §6).
if std::env::args().nth(1).as_deref() == Some("--version") {
println!("darkroom-desktop {}", env!("CARGO_PKG_VERSION"));
return Ok(());
@@ -61,6 +61,11 @@ fn main() -> anyhow::Result<()> {
eprintln!("usage: darkroom-desktop <file-or-directory>...");
}
// Before the window: the probe runs on its own thread and the first
// frame does not wait for it, but the models a background job asks for
// should already know where the runtime is (docs/dev/inference.md §4).
dr_ui::inference::init(runtime_dirs());
dr_ui::run(paths)?;
// Skip Rust's normal static/thread-local teardown on the way out: a
@@ -70,3 +75,39 @@ fn main() -> anyhow::Result<()> {
// destruction" when the window is closed.
std::process::exit(0);
}
/// Where a desktop package may have put `libonnxruntime`, most specific
/// first. None of these existing is the tract build, which is a complete
/// application and not an error (docs/dev/inference.md §3).
///
/// `DARKROOM_ORT_DIR` is for a developer pointing at a runtime that is not
/// installed — the wheel's `capi` directory, say. Then beside the executable
/// and in the package's private library directory, for a package that
/// bundles its own; then the user's own `runtime/` beside the models, where
/// `tools/fetch-desktop-runtime.sh` puts one; then the Flatpak prefix; then
/// the system library directory, for a distribution that ships ONNX Runtime
/// as a package of its own. The user's copy outranks the system's because
/// the system's is the one most likely to be built without the GPU
/// providers, or against the wrong cuDNN — and a system copy whose providers
/// do not load is not a problem, only a slower app: the probe builds a real
/// session before believing a provider.
fn runtime_dirs() -> Vec<PathBuf> {
let mut dirs = Vec::new();
if let Some(dir) = std::env::var_os("DARKROOM_ORT_DIR") {
dirs.push(PathBuf::from(dir));
}
if let Ok(exe) = std::env::current_exe() {
if let Some(bin) = exe.parent() {
dirs.push(bin.to_path_buf());
dirs.push(bin.join("../lib/darkroom"));
}
}
dirs.push(dr_ui::inference::user_runtime_dir());
#[cfg(target_os = "linux")]
dirs.extend([
PathBuf::from("/app/lib/darkroom"),
PathBuf::from("/usr/lib/darkroom"),
PathBuf::from("/usr/lib"),
]);
dirs
}
+1 -1
View File
@@ -315,7 +315,7 @@ fn full_library(
// The three phases, separately, because "a regroup takes n seconds" does
// not tell anyone which half to optimise — and the answer differs between
// a desktop and a tablet (docs/faces.md §9).
// a desktop and a tablet (docs/dev/faces.md §9).
{
let dim = candidates.first().map(|c| c.embedding.len()).unwrap_or(0);
let flat: Vec<f32> = candidates
+1 -1
View File
@@ -82,7 +82,7 @@
//! Grouping has no natural `subject_id`: it is a property of a *run* of frames,
//! so a per-image job would rebuild the world once per photograph. It is
//! therefore a debounced library-level pass, for exactly the reasons
//! docs/catalog.md §10.2 gives for face clustering, and [`regroup`] is the whole
//! docs/dev/catalog.md §10.2 gives for face clustering, and [`regroup`] is the whole
//! of it — one ordered walk, no per-pair comparison beyond adjacent frames.
//!
//! # Grouping is not hiding
+463 -61
View File
@@ -2,7 +2,7 @@
//! Face data as sealed shards, so a second device does not re-index the library.
//!
//! Indexing a 23,500-image library is on the order of two hours of CPU
//! (docs/faces.md §12.2). It is also **byte-identical on every device**: the
//! (docs/dev/faces.md §12.2). It is also **byte-identical on every device**: the
//! same model over the same proxy produces the same embedding. Paying for it
//! once per account rather than once per device is the whole point of this
//! module, and it is the same bargain the thumbnail store already makes.
@@ -48,9 +48,10 @@ pub const SHARD_MAX_BYTES: u64 = dr_thumbs::SHARD_MAX_BYTES;
/// Bytes one stored face occupies, near enough to bound a shard by.
///
/// Counted rather than measured: the embedding is fixed at 512 × f16, the
/// landmarks at 5 × 2 × f32, and the rest is a handful of numbers. Measuring
/// landmarks at 5 × 2 × f32 and the dense ones at 106 × 2 × u16, and the
/// rest is a handful of numbers. Measuring
/// the file after each insert would mean a `VACUUM` to get an honest answer.
const BYTES_PER_FACE: u64 = 1024 + 40 + 64;
const BYTES_PER_FACE: u64 = 1024 + 40 + 424 + 64;
/// Bytes a stored crop occupies, near enough to bound a shard by.
///
@@ -79,6 +80,11 @@ pub struct SharedFace {
/// See `faces::DetectedFace::quality`. `None` from a shard written before
/// the number was kept.
pub quality: Option<f32>,
/// See `faces::DetectedFace::eyes`. `None` from a peer without the eye
/// models, or a shard written before they existed.
pub eyes: Option<dr_face::EyeReading>,
/// See `faces::DetectedFace::landmarks_dense`; empty where none.
pub landmarks_dense: Vec<u8>,
/// The face cut out and encoded, or empty where none was kept.
///
/// Travels with the face rather than in the catalog snapshot, which is the
@@ -149,6 +155,90 @@ impl FaceShardStore {
.flatten()
}
/// The pipeline this store holds an image under, among those sharing
/// `model_id`'s embedder — the most recently indexed where a peer has
/// sent more than one.
///
/// What the import asks: not "has anyone run *this* detector over it" but
/// "does anyone hold comparable faces for it". See `faces::embedder_of`.
pub fn held_model(&self, file_id: u64, model_id: &str) -> Option<String> {
self.index
.query_row(
&format!(
"SELECT model_id FROM entries
WHERE file_id = ?1 AND {} = ?2
ORDER BY indexed_at DESC NULLS LAST, model_id",
crate::faces::embedder_sql("model_id")
),
rusqlite::params![file_id as i64, crate::faces::embedder_of(model_id)],
|r| r.get::<_, String>(0),
)
.optional()
.ok()
.flatten()
}
/// The other pipelines this file is held under that share `model_id`'s
/// embedder — the generations a put of `model_id` may supersede.
fn siblings(&self, file_id: u64, model_id: &str) -> Vec<String> {
let mut stmt = match self.index.prepare(&format!(
"SELECT model_id FROM entries
WHERE file_id = ?1 AND model_id != ?2 AND {} = ?3",
crate::faces::embedder_sql("model_id")
)) {
Ok(s) => s,
Err(_) => return Vec::new(),
};
stmt.query_map(
rusqlite::params![
file_id as i64,
model_id,
crate::faces::embedder_of(model_id)
],
|r| r.get::<_, String>(0),
)
.map(|rows| rows.filter_map(|r| r.ok()).collect())
.unwrap_or_default()
}
/// Whether a pass this file is already held under outranks `model_id`,
/// so a put of `model_id` would add a generation nobody would adopt.
pub fn outranked(&self, file_id: u64, model_id: &str) -> bool {
use dr_types::FaceDetector;
let Some(incoming) = FaceDetector::for_model_id(model_id) else {
return false;
};
self.siblings(file_id, model_id)
.iter()
.filter_map(|m| FaceDetector::for_model_id(m))
.any(|held| held.outranks(incoming))
}
/// Forget the index entries for generations of this file that `model_id`
/// outranks. The bytes stay where they are — a sealed shard is
/// immutable — but the store stops offering them, and a later export or
/// merge writes nothing for them again.
fn supersede(&self, file_id: u64, model_id: &str) -> Result<(), CatalogError> {
use dr_types::FaceDetector;
let Some(incoming) = FaceDetector::for_model_id(model_id) else {
return Ok(());
};
for held in self.siblings(file_id, model_id) {
let weaker = FaceDetector::for_model_id(&held).is_some_and(|h| incoming.outranks(h));
if weaker {
self.index.execute(
"DELETE FROM entries WHERE file_id = ?1 AND model_id = ?2",
rusqlite::params![file_id as i64, held],
)?;
self.index.execute(
"DELETE FROM faces_meta WHERE file_id = ?1 AND model_id = ?2",
rusqlite::params![file_id as i64, held],
)?;
}
}
Ok(())
}
pub fn contains(&self, file_id: u64, model_id: &str) -> bool {
self.index
.query_row(
@@ -204,6 +294,15 @@ impl FaceShardStore {
faces: &[SharedFace],
indexed_at: Option<i64>,
) -> Result<u32, CatalogError> {
// One generation per image per embedder. A store carried every pass
// — 24,123 entries for 19,089 images on the reference library, a
// third of its 293 MB — and only the strongest was ever adopted.
// A weaker pass arriving after a stronger one is not written; a
// stronger one arriving retires the weaker from the index.
if self.outranked(file_id, model_id) {
return Ok(0);
}
self.supersede(file_id, model_id)?;
let incoming = faces
.iter()
.map(|f| BYTES_PER_FACE + if f.crop.is_empty() { 0 } else { BYTES_PER_CROP })
@@ -227,8 +326,12 @@ impl FaceShardStore {
tx.execute(
"INSERT INTO faces
(file_id, model_id, x, y, w, h, landmarks, confidence,
embedding, crop_px, crop, quality)
VALUES (?1, ?2, ?3, ?4, ?5, ?6, ?7, ?8, ?9, ?10, ?11, ?12)",
embedding, crop_px, crop, quality,
eye_right, eye_right_px, eye_right_sharp,
eye_left, eye_left_px, eye_left_sharp, sunglasses,
landmarks_dense)
VALUES (?1, ?2, ?3, ?4, ?5, ?6, ?7, ?8, ?9, ?10, ?11, ?12,
?13, ?14, ?15, ?16, ?17, ?18, ?19, ?20)",
rusqlite::params![
f.file_id as i64,
f.model_id,
@@ -242,6 +345,14 @@ impl FaceShardStore {
f.crop_px as f64,
(!f.crop.is_empty()).then_some(f.crop.as_slice()),
f.quality.map(f64::from),
f.eyes.map(|e| f64::from(e.right.open)),
f.eyes.map(|e| f64::from(e.right.px)),
f.eyes.map(|e| f64::from(e.right.sharpness)),
f.eyes.map(|e| f64::from(e.left.open)),
f.eyes.map(|e| f64::from(e.left.px)),
f.eyes.map(|e| f64::from(e.left.sharpness)),
f.eyes.map(|e| f64::from(e.sunglasses)),
(!f.landmarks_dense.is_empty()).then_some(f.landmarks_dense.as_slice()),
],
)?;
}
@@ -433,33 +544,67 @@ impl FaceShardStore {
rusqlite::OpenFlags::SQLITE_OPEN_READ_ONLY | rusqlite::OpenFlags::SQLITE_OPEN_NO_MUTEX,
)?;
let mut q =
src.prepare("SELECT file_id, model_id, faces_found, source_edge FROM indexed")?;
let images: Vec<(i64, String, i64, i64)> = q
.query_map([], |r| Ok((r.get(0)?, r.get(1)?, r.get(2)?, r.get(3)?)))?
// The peer's marker travels with the image: it is what lets
// `import_from_shards` record the adoption under the time the peer
// indexed it, and so what keeps `export_to_shards` from reading the
// adoption as a re-index and sending the peer's faces back out under
// this device's name. A shard from before the column has none.
let mut q = src.prepare(&format!(
"SELECT file_id, model_id, faces_found, source_edge, {} FROM indexed",
match has_column(&src, "indexed", "indexed_at") {
Ok(true) => "indexed_at",
_ => "NULL",
}
))?;
let images: Vec<(i64, String, i64, i64, Option<i64>)> = q
.query_map([], |r| {
Ok((r.get(0)?, r.get(1)?, r.get(2)?, r.get(3)?, r.get(4)?))
})?
.collect::<Result<_, _>>()?;
let mut adopted = 0;
for (file_id, model_id, _found, edge) in images {
if self.contains(file_id as u64, &model_id) {
for (file_id, model_id, _found, edge, indexed_at) in images {
if self.contains(file_id as u64, &model_id) || self.outranked(file_id as u64, &model_id)
{
continue;
}
let mut fq = src.prepare(&format!(
"SELECT f.file_id, f.model_id, f.x, f.y, f.w, f.h, f.landmarks,
f.confidence, f.embedding, f.crop_px, {}, {}
f.confidence, f.embedding, f.crop_px, {}, {}, {}, {}
FROM faces f WHERE f.file_id = ?1 AND f.model_id = ?2",
column_or_null(&src, "crop"),
column_or_null(&src, "quality"),
crate::schema::EYE_COLUMNS
.iter()
.map(|c| column_or_null(&src, c))
.collect::<Vec<_>>()
.join(", "),
column_or_null(&src, "landmarks_dense"),
))?;
let faces: Vec<SharedFace> = fq
.query_map(rusqlite::params![file_id, &model_id], read_shared_face)?
.collect::<Result<_, _>>()?;
self.put_image(file_id as u64, &model_id, edge as u32, &faces)?;
self.put_image_at(file_id as u64, &model_id, edge as u32, &faces, indexed_at)?;
adopted += 1;
}
Ok(adopted)
}
/// Record when the catalog indexed a held image, for an entry that
/// arrived without a marker — a peer's shard from before the column.
pub fn set_indexed_at(
&self,
file_id: u64,
model_id: &str,
at: i64,
) -> Result<(), CatalogError> {
self.index.execute(
"UPDATE entries SET indexed_at = ?3 WHERE file_id = ?1 AND model_id = ?2",
rusqlite::params![file_id as i64, model_id, at],
)?;
Ok(())
}
/// Read back everything held for one image.
pub fn get_image(
&self,
@@ -490,7 +635,8 @@ impl FaceShardStore {
let mut q = conn.prepare(
"SELECT file_id, model_id, x, y, w, h, landmarks, confidence, embedding, crop_px,
crop, quality
crop, quality, eye_right, eye_right_px, eye_right_sharp,
eye_left, eye_left_px, eye_left_sharp, sunglasses, landmarks_dense
FROM faces WHERE file_id = ?1 AND model_id = ?2",
)?;
let faces: Vec<SharedFace> = q
@@ -548,6 +694,14 @@ fn upgrade_shard(conn: &Connection) -> Result<(), CatalogError> {
("faces", "crop", "BLOB"),
("indexed", "indexed_at", "INTEGER"),
("faces", "quality", "REAL"),
("faces", "eye_right", "REAL"),
("faces", "eye_right_px", "REAL"),
("faces", "eye_right_sharp", "REAL"),
("faces", "eye_left", "REAL"),
("faces", "eye_left_px", "REAL"),
("faces", "eye_left_sharp", "REAL"),
("faces", "sunglasses", "REAL"),
("faces", "landmarks_dense", "BLOB"),
] {
if !has_column(conn, table, column)? {
conn.execute_batch(&format!("ALTER TABLE {table} ADD COLUMN {column} {decl}"))?;
@@ -571,7 +725,7 @@ fn has_column(conn: &Connection, table: &str, column: &str) -> Result<bool, Cata
///
/// `column` is one of this module's own names, never anything read from
/// outside, which is what makes formatting it into SQL acceptable.
fn column_or_null(conn: &Connection, column: &'static str) -> String {
fn column_or_null(conn: &Connection, column: &str) -> String {
match has_column(conn, "faces", column) {
Ok(true) => format!("f.{column}"),
_ => "NULL".to_string(),
@@ -608,16 +762,21 @@ pub fn export_to_shards_reporting(
model_id: &str,
progress: &mut dyn FnMut(usize, usize),
) -> Result<usize, CatalogError> {
let mut q = conn.prepare(
"SELECT r.file_id, fi.image_id, fi.source_edge, fi.indexed_at
// Every pipeline sharing this one's embedder, each image under the id
// that actually indexed it. A device that switched detectors still holds
// most of its library under the previous id, and those faces are exactly
// as comparable — and as wanted by a peer — as the new ones.
let mut q = conn.prepare(&format!(
"SELECT r.file_id, fi.image_id, fi.source_edge, fi.indexed_at, fi.model_id
FROM face_index fi
JOIN remote r ON r.image_id = fi.image_id
WHERE fi.model_id = ?1
WHERE {} = ?1
ORDER BY fi.image_id",
)?;
let rows: Vec<(i64, i64, i64, i64)> = q
.query_map([model_id], |r| {
Ok((r.get(0)?, r.get(1)?, r.get(2)?, r.get(3)?))
crate::faces::embedder_sql("fi.model_id")
))?;
let rows: Vec<(i64, i64, i64, i64, String)> = q
.query_map([crate::faces::embedder_of(model_id)], |r| {
Ok((r.get(0)?, r.get(1)?, r.get(2)?, r.get(3)?, r.get(4)?))
})?
.collect::<Result<_, _>>()?;
@@ -627,7 +786,8 @@ pub fn export_to_shards_reporting(
let total = rows.len();
let mut exported = 0;
for (seen, (file_id, image_id, edge, indexed_at)) in rows.into_iter().enumerate() {
for (seen, (file_id, image_id, edge, indexed_at, model_id)) in rows.into_iter().enumerate() {
let model_id = model_id.as_str();
if seen.is_multiple_of(REPORT_EVERY) {
progress(seen, total);
}
@@ -648,7 +808,8 @@ pub fn export_to_shards_reporting(
}
let mut fq = conn.prepare(
"SELECT x, y, w, h, landmarks, detector_confidence, embedding, crop_px, crop,
quality
quality, eye_right, eye_right_px, eye_right_sharp,
eye_left, eye_left_px, eye_left_sharp, sunglasses, landmarks_dense
FROM faces WHERE image_id = ?1 AND model_id = ?2",
)?;
let faces: Vec<SharedFace> = fq
@@ -666,6 +827,8 @@ pub fn export_to_shards_reporting(
crop_px: r.get::<_, f64>(7)? as f32,
crop: r.get::<_, Option<Vec<u8>>>(8)?.unwrap_or_default(),
quality: r.get::<_, Option<f64>>(9)?.map(|q| q as f32),
eyes: crate::faces::read_eyes(r, 10)?,
landmarks_dense: r.get::<_, Option<Vec<u8>>>(17)?.unwrap_or_default(),
})
})?
.collect::<Result<_, _>>()?;
@@ -689,10 +852,20 @@ pub fn export_to_shards_reporting(
/// adopted rather than re-detected, which is the difference between a new
/// device being useful in a minute and in two hours.
///
/// Skips any image this device has already indexed itself. Local work is not
/// second-guessed by a peer's — the two should agree, since the same model over
/// the same proxy is deterministic, but where they do not, the copy this device
/// computed is the one it can vouch for.
/// Skips any image this device has already indexed itself under this
/// pipeline or any sharing its embedder — unless the peer ran a detector that
/// outranks the one that indexed it here. Local work is not second-guessed
/// by a peer's equal: the two should agree, since the same model over the
/// same proxy is deterministic, and where they do not, the copy this device
/// computed is the one it can vouch for. A peer's *stronger* pass is another
/// matter: it is the re-detection this device's own sweep would queue
/// (`FaceDetector::supersedes`), already done, and taking it is what spares
/// a tablet the fetch. Names survive the replacement by box overlap and
/// embedding, as they do a local re-detection (`faces::record_detections`).
///
/// A peer's faces are taken under whichever compatible detector found them:
/// a tablet set to the fast detector adopts the desktop's thorough pass
/// rather than re-detecting it worse.
///
/// Returns how many images were adopted.
pub fn import_from_shards(
@@ -700,36 +873,80 @@ pub fn import_from_shards(
store: &FaceShardStore,
model_id: &str,
) -> Result<usize, CatalogError> {
use dr_types::FaceDetector;
// Only images this device actually has. A shard covers the whole account,
// and a device holding a subset of the library should take only its own
// part rather than accumulating faces for photographs it cannot show.
let mut q = conn.prepare(
"SELECT r.file_id, r.image_id
//
// With the pipeline that indexed each one here, or NULL: the marker is
// what decides whether a peer's copy is a gap filled or an upgrade.
let mut q = conn.prepare(&format!(
"SELECT r.file_id, r.image_id,
(SELECT fi.model_id FROM face_index fi
WHERE fi.image_id = r.image_id AND {} = ?1)
FROM remote r
JOIN images i ON i.id = r.image_id
WHERE i.trashed_at IS NULL
AND NOT EXISTS (
SELECT 1 FROM face_index fi
WHERE fi.image_id = r.image_id AND fi.model_id = ?1
)",
)?;
let candidates: Vec<(i64, i64)> = q
.query_map([model_id], |r| Ok((r.get(0)?, r.get(1)?)))?
WHERE i.trashed_at IS NULL",
crate::faces::embedder_sql("fi.model_id")
))?;
let candidates: Vec<(i64, i64, Option<String>)> = q
.query_map([crate::faces::embedder_of(model_id)], |r| {
Ok((r.get(0)?, r.get(1)?, r.get(2)?))
})?
.collect::<Result<_, _>>()?;
/// Images per write transaction. Large enough that fourteen thousand
/// adoptions are a hundred and forty commits rather than fourteen
/// thousand; small enough that a read on the UI thread, queued behind
/// the lock, waits a fraction of a second and not the whole import.
const CHUNK: usize = 100;
let mut adopted = 0;
for (file_id, image_id) in candidates {
let Some((faces, edge)) = store.get_image(file_id as u64, model_id)? else {
let mut tx = conn.unchecked_transaction()?;
let mut in_chunk = 0;
for (file_id, image_id, local) in candidates {
if in_chunk == CHUNK {
tx.commit()?;
tx = conn.unchecked_transaction()?;
in_chunk = 0;
}
let Some(held) = store.held_model(file_id as u64, model_id) else {
continue;
};
// A peer that embedded before the quality was kept has done work this
// device cannot finish: the number exists only at embedding time, and
// adopting the faces would write the run marker that keeps them from
// ever being measured (schema V14). Left for this device's own pass —
// or for the peer's, whose re-export replaces these.
if faces.iter().any(|f| f.quality.is_none()) {
continue;
if let Some(local) = local {
// An unknown detector on either side cannot be ranked, and an
// unranked peer is treated as an equal: kept out.
let upgrade = match (
FaceDetector::for_model_id(&held),
FaceDetector::for_model_id(&local),
) {
(Some(theirs), Some(ours)) => theirs.outranks(ours),
_ => false,
};
if !upgrade {
continue;
}
}
let Some((faces, edge)) = store.get_image(file_id as u64, &held)? else {
continue;
};
// A face the peer embedded before its quality was kept (schema V14)
// is adopted with the reading missing, exactly as one without an eye
// reading is. The measuring passes find their work by the NULL
// column, not by the run marker (`dr_ui::repairs`, `faces_needing`),
// so adopting costs the reading nothing and this device's own pass
// fills it.
//
// This used to refuse such faces, on the reasoning that the marker
// would stop them ever being measured — true before the quality
// repair existed, and wrong after. What it cost: V14 had dropped the
// markers of every image holding such faces, so the peer never
// re-exported them, and the only copies in the shards were the
// unmeasured ones. A tablet holding shards with 4,310 of the
// desktop's images and 3,170 of its confirmations declined every one
// of them, showed a fraction of each person, and queued the whole
// library for a re-detection of its own instead.
let local: Vec<crate::faces::DetectedFace> = faces
.into_iter()
.map(|f| crate::faces::DetectedFace {
@@ -742,6 +959,8 @@ pub fn import_from_shards(
embedding: f.embedding,
crop_px: f.crop_px,
quality: f.quality,
eyes: f.eyes,
landmarks_dense: f.landmarks_dense,
model_id: f.model_id,
// A peer that indexed before crops existed sends none, and the
// reader falls back to the proxy exactly as it does for a face
@@ -750,15 +969,42 @@ pub fn import_from_shards(
})
.collect();
crate::faces::record_detections(
conn,
crate::faces::record_detections_within(
&tx,
dr_types::ImageId(image_id as u64),
model_id,
&held,
edge,
&local,
)?;
// The peer's marker, not this moment. `record_detections` stamps the
// run as now, and `export_to_shards` reads a marker newer than the
// shard's as a re-index — so every adopted image went straight back
// out as this device's own work: 14,100 adopted, 15,457 "newly
// indexed" on the next pass, and twenty-two shards of a peer's faces
// uploaded again under a second name. Where the peer's shard carried
// no marker, the store takes the catalog's, so the two agree either
// way and the export sees nothing to send.
match store.indexed_at(file_id as u64, &held) {
Some(theirs) => {
tx.execute(
"UPDATE face_index SET indexed_at = ?3
WHERE image_id = ?1 AND model_id = ?2",
rusqlite::params![image_id, held, theirs],
)?;
}
None => {
let ours: i64 = tx.query_row(
"SELECT indexed_at FROM face_index WHERE image_id = ?1 AND model_id = ?2",
rusqlite::params![image_id, held],
|r| r.get(0),
)?;
store.set_indexed_at(file_id as u64, &held, ours)?;
}
}
adopted += 1;
in_chunk += 1;
}
tx.commit()?;
Ok(adopted)
}
@@ -788,6 +1034,8 @@ fn read_shared_face(r: &rusqlite::Row<'_>) -> rusqlite::Result<SharedFace> {
crop_px: r.get::<_, f64>(9)? as f32,
crop: r.get::<_, Option<Vec<u8>>>(10)?.unwrap_or_default(),
quality: r.get::<_, Option<f64>>(11)?.map(|q| q as f32),
eyes: crate::faces::read_eyes(r, 12)?,
landmarks_dense: r.get::<_, Option<Vec<u8>>>(19)?.unwrap_or_default(),
})
}
@@ -875,7 +1123,20 @@ CREATE TABLE IF NOT EXISTS faces (
-- Length of the raw embedding (`faces::DetectedFace::quality`). NULL from
-- a build that did not keep it, and a face the receiving device will not
-- adopt -- see `import_from_shards`.
quality REAL
quality REAL,
-- The eye reading (`faces::DetectedFace::eyes`), all seven or none. NULL
-- from a peer without the eye models; adopted anyway, and read by the
-- receiving device's own measuring pass if it has them.
eye_right REAL,
eye_right_px REAL,
eye_right_sharp REAL,
eye_left REAL,
eye_left_px REAL,
eye_left_sharp REAL,
sunglasses REAL,
-- The dense landmarks behind the reading (`faces::DetectedFace::
-- landmarks_dense`), 424 bytes packed; NULL where none.
landmarks_dense BLOB
);
CREATE INDEX IF NOT EXISTS faces_file ON faces(file_id, model_id);
@@ -935,6 +1196,8 @@ mod tests {
embedding: vec![seed; 1024],
crop_px: 180.0,
quality: Some(17.5),
eyes: None,
landmarks_dense: Vec::new(),
crop: vec![seed; 64],
}
}
@@ -977,6 +1240,32 @@ mod tests {
assert!(!s.contains(1, "lvface"));
}
/// One generation per image per embedder: a stronger detector's pass
/// retires a weaker one from the index, and a weaker pass arriving after
/// a stronger is not written at all.
#[test]
fn a_stronger_pass_retires_a_weaker_one_and_a_weaker_is_not_added() {
let dir = tempdir();
let mut s = FaceShardStore::open(&dir).unwrap();
s.put_image(1, "w600k_mbf", 1024, &[face(1, 1)]).unwrap();
s.put_image(1, "scrfd_10g+w600k_mbf", 1024, &[face(1, 2)])
.unwrap();
assert!(s.contains(1, "scrfd_10g+w600k_mbf"));
assert!(!s.contains(1, "w600k_mbf"), "the fast pass was not retired");
assert_eq!(s.len(), 1, "faces_meta still counts the retired pass");
s.put_image(1, "scrfd_2.5g+w600k_mbf", 1024, &[face(1, 3)])
.unwrap();
assert!(
!s.contains(1, "scrfd_2.5g+w600k_mbf"),
"a weaker pass was added"
);
assert_eq!(
s.held_model(1, "w600k_mbf").as_deref(),
Some("scrfd_10g+w600k_mbf")
);
}
#[test]
fn re_storing_an_image_replaces_rather_than_doubling_it() {
let dir = tempdir();
@@ -1144,6 +1433,8 @@ mod catalog_round_trip {
embedding: vec![seed; 1024],
crop_px: 180.0,
quality: Some(20.0),
eyes: None,
landmarks_dense: Vec::new(),
model_id: "w600k_mbf".into(),
crop: vec![seed; 64],
}
@@ -1195,13 +1486,62 @@ mod catalog_round_trip {
assert!((got[0].landmarks[2].0 - 0.15).abs() < 1e-5);
let emb = faces::embeddings(&b, "w600k_mbf").unwrap();
assert!(emb.iter().any(|e| e.embedding[0] == 1));
// And what B adopted is not B's work: its next export sends nothing.
// Adopting used to stamp the run as now, so every adopted image went
// back out under B's name as a re-index.
assert_eq!(export_to_shards(&b, &mut store_b, "w600k_mbf").unwrap(), 0);
}
/// A face a peer embedded without measuring it is work this device
/// cannot finish, and adopting it would write the marker that stops it
/// ever being measured. The image stays outstanding instead.
/// The desktop switched to a stronger detector part-way through the
/// library, so its faces sit under two pipeline ids. A tablet on the
/// original detector must receive *all* of them — each under the id that
/// found it — and not re-detect the thorough half worse.
#[test]
fn a_peers_unmeasured_faces_are_left_for_this_device_to_index() {
fn every_generation_sharing_an_embedder_travels_and_is_adopted() {
let a = device(&[(1, 5001), (2, 5002)]);
let b = device(&[(90, 5001), (91, 5002)]);
faces::record_detections(&a, dr_types::ImageId(1), "w600k_mbf", 1024, &[detected(1)])
.unwrap();
let mut thorough = detected(2);
thorough.model_id = "scrfd_10g+w600k_mbf".into();
faces::record_detections(
&a,
dr_types::ImageId(2),
"scrfd_10g+w600k_mbf",
1024,
&[thorough],
)
.unwrap();
let mut store_a = FaceShardStore::open(&tempdir("a")).unwrap();
assert_eq!(
export_to_shards(&a, &mut store_a, "scrfd_10g+w600k_mbf").unwrap(),
2,
"the export left the earlier detector's images behind"
);
let mut store_b = FaceShardStore::open(&tempdir("b")).unwrap();
store_b.merge_shard(&store_a.shard_path(0)).unwrap();
assert_eq!(import_from_shards(&b, &store_b, "w600k_mbf").unwrap(), 2);
assert_eq!(faces::coverage(&b, "w600k_mbf").unwrap().outstanding(), 0);
let old = faces::for_image(&b, dr_types::ImageId(90)).unwrap();
let new = faces::for_image(&b, dr_types::ImageId(91)).unwrap();
assert_eq!(old[0].model_id, "w600k_mbf");
assert_eq!(
new[0].model_id, "scrfd_10g+w600k_mbf",
"adopted under the wrong id"
);
}
/// A face a peer embedded without measuring it is adopted all the same,
/// and left on this device's quality pass by its missing reading. Refusing
/// it was what stranded every confirmation the desktop had made on faces
/// from before V14: the tablet held the shards and would not use them.
#[test]
fn a_peers_unmeasured_faces_are_adopted_and_left_for_the_quality_pass() {
let b = device(&[(90, 5001), (91, 5002)]);
let mut store = FaceShardStore::open(&tempdir("unmeasured")).unwrap();
let shared = |file_id: u64, quality: Option<f32>| SharedFace {
@@ -1216,6 +1556,8 @@ mod catalog_round_trip {
embedding: vec![1; 1024],
crop_px: 180.0,
quality,
eyes: None,
landmarks_dense: Vec::new(),
crop: Vec::new(),
};
store
@@ -1225,13 +1567,18 @@ mod catalog_round_trip {
.put_image(5002, "w600k_mbf", 2560, &[shared(5002, Some(19.0))])
.unwrap();
assert_eq!(import_from_shards(&b, &store, "w600k_mbf").unwrap(), 1);
assert_eq!(import_from_shards(&b, &store, "w600k_mbf").unwrap(), 2);
let cov = faces::coverage(&b, "w600k_mbf").unwrap();
assert_eq!(cov.indexed, 1);
assert_eq!(cov.outstanding(), 1, "the unmeasured image was adopted");
assert!(faces::for_image(&b, dr_types::ImageId(90))
.unwrap()
.is_empty());
assert_eq!(cov.indexed, 2);
assert_eq!(cov.outstanding(), 0, "the unmeasured image was refused");
let got = faces::for_image(&b, dr_types::ImageId(90)).unwrap();
assert_eq!(got.len(), 1);
assert_eq!(got[0].quality, None, "a reading was invented");
// Still owed to the measuring pass, which lists by the column.
assert_eq!(
faces::count_needing(&b, "w600k_mbf", "f.quality IS NULL").unwrap(),
1
);
}
#[test]
@@ -1257,6 +1604,59 @@ mod catalog_round_trip {
assert_eq!(emb[0].embedding[0], 9, "B's own embedding was overwritten");
}
/// A peer's stronger detector is the re-detection this device would
/// otherwise queue for itself. Taking it saves the fetch; the name the
/// user confirmed here rides across on the box, as it would locally.
#[test]
fn a_peers_stronger_pass_replaces_a_weaker_local_one_and_keeps_the_name() {
let a = device(&[(1, 5001)]);
let b = device(&[(50, 5001)]);
let ids =
faces::record_detections(&b, dr_types::ImageId(50), "w600k_mbf", 1024, &[detected(9)])
.unwrap();
let anna = faces::create_person(&b, "Anna").unwrap();
faces::confirm(&b, ids[0], anna).unwrap();
let mut thorough = detected(7);
thorough.model_id = "scrfd_10g+w600k_mbf".into();
let mut second = detected(8);
second.model_id = "scrfd_10g+w600k_mbf".into();
second.x = 0.6;
faces::record_detections(
&a,
dr_types::ImageId(1),
"scrfd_10g+w600k_mbf",
1024,
&[thorough, second],
)
.unwrap();
let mut store = FaceShardStore::open(&tempdir("upgrade")).unwrap();
export_to_shards(&a, &mut store, "scrfd_10g+w600k_mbf").unwrap();
assert_eq!(import_from_shards(&b, &store, "w600k_mbf").unwrap(), 1);
let got = faces::for_image(&b, dr_types::ImageId(50)).unwrap();
assert_eq!(got.len(), 2, "the stronger pass was not adopted");
let named = got
.iter()
.find(|f| f.person == Some(anna))
.expect("the name was lost");
assert!(named.confirmed);
assert_eq!(named.model_id, "scrfd_10g+w600k_mbf");
// And never downwards: A on the fast detector keeps B's thorough faces.
let mut store_b = FaceShardStore::open(&tempdir("downgrade")).unwrap();
faces::record_detections(&b, dr_types::ImageId(50), "w600k_mbf", 1024, &[detected(9)])
.unwrap();
export_to_shards(&b, &mut store_b, "w600k_mbf").unwrap();
assert_eq!(
import_from_shards(&a, &store_b, "scrfd_10g+w600k_mbf").unwrap(),
0
);
assert_eq!(faces::for_image(&a, dr_types::ImageId(1)).unwrap().len(), 2);
}
/// A device holding a subset of the library takes only its own part.
#[test]
fn a_device_ignores_faces_for_photographs_it_does_not_have() {
@@ -1349,6 +1749,8 @@ mod catalog_round_trip {
embedding: vec![seed; 1024],
crop_px: 180.0,
quality: None,
eyes: None,
landmarks_dense: Vec::new(),
crop: vec![seed; 64],
}
}
+1162 -153
View File
File diff suppressed because it is too large Load Diff
+1 -1
View File
@@ -61,7 +61,7 @@ pub use collections::{Collection, CollectionKind, TreeRow};
pub use dedup::{seen_by_content, seen_by_metadata, set_content_hash};
pub use error::CatalogError;
pub use face_shard::{FaceShardStore, SharedFace};
pub use faces::{Calibration, DetectedFace, Face, FaceId, Measurement, Person, PersonId};
pub use faces::{Calibration, DetectedFace, Face, FaceId, FaceUpdate, Person, PersonId};
pub use jobs::{Job, JobKind, Priority};
pub use keywords::{Coverage, Keyword, KeywordId, SelectionKeyword};
pub use merge::MergeReport;
+171 -6
View File
@@ -119,6 +119,8 @@ pub struct MergeReport {
pub keywords_fused: usize,
/// Keyword assignments taken from the remote.
pub keywords_assigned: usize,
/// Images whose capture metadata was taken from the remote.
pub metadata_adopted: usize,
}
impl MergeReport {
@@ -133,6 +135,7 @@ impl MergeReport {
|| self.keywords_deleted > 0
|| self.keywords_fused > 0
|| self.keywords_assigned > 0
|| self.metadata_adopted > 0
}
/// Whether the local catalog holds anything the remote did not, and so
@@ -193,10 +196,67 @@ pub fn merge_all(conn: &Connection) -> Result<MergeReport, CatalogError> {
merge_collections_within(&tx, &mut report)?;
merge_keywords_within(&tx, &mut report)?;
merge_people_within(&tx, &mut report)?;
merge_metadata_within(&tx, &mut report)?;
tx.commit()?;
Ok(report)
}
/// Adopt capture metadata from an attached catalog, on its own.
pub fn merge_metadata(conn: &Connection) -> Result<MergeReport, CatalogError> {
let tx = conn.unchecked_transaction()?;
let mut report = MergeReport::default();
merge_metadata_within(&tx, &mut report)?;
tx.commit()?;
Ok(report)
}
/// Capture metadata a peer's sweep already read, for images this device has
/// not dated yet.
///
/// The `images` table is local state and the merge leaves it alone — except
/// for these columns, which are not: a capture time, an offset, a camera, a
/// lens and an ISO are facts about the file's bytes, identical on every
/// device, and read by fetching a header per image across the whole library
/// (`dr_ui::library::spawn_sweep`). A fresh device inherits its peers'
/// thumbnails and faces from the shards and then spent hours re-reading
/// every header for the timeline; the snapshot it had just merged held
/// every one of those dates.
///
/// Matched by `oc:fileid`, as collection membership is. Only rows still at
/// `metadata_state < 2` take anything, and only from a remote row at 2: a
/// date this device read for itself is never overwritten, and a peer that
/// has not read one has nothing to give. The sweep's own query
/// (`metadata_state < 2`) then finds nothing left to do for them.
const METADATA_BY_FILE_ID: &str = "
UPDATE main.images
SET captured_at = r.captured_at,
captured_offset = coalesce(main.images.captured_offset, r.captured_offset),
camera = coalesce(main.images.camera, r.camera),
lens = coalesce(main.images.lens, r.lens),
iso = coalesce(main.images.iso, r.iso),
metadata_state = 2
FROM (SELECT lr.image_id, ri.captured_at, ri.captured_offset,
ri.camera, ri.lens, ri.iso
FROM remote_cat.images ri
JOIN remote_cat.remote rr ON rr.image_id = ri.id
JOIN main.remote lr ON lr.file_id = rr.file_id
WHERE ri.metadata_state >= 2 AND ri.captured_at IS NOT NULL) AS r
WHERE main.images.id = r.image_id
AND main.images.metadata_state < 2";
fn merge_metadata_within(tx: &Connection, report: &mut MergeReport) -> Result<(), CatalogError> {
// A snapshot from before these columns, or from a library with no server
// behind it, has nothing to join on.
if !remote_has(tx, "remote")?
|| !remote_has_column(tx, "images", "metadata_state")?
|| !remote_has_column(tx, "images", "captured_offset")?
{
return Ok(());
}
report.metadata_adopted = tx.execute(METADATA_BY_FILE_ID, [])?;
Ok(())
}
/// Merge people and identity judgements from an attached catalog.
///
/// The people half of [`merge_all`], on its own, for the same reason the other
@@ -729,7 +789,7 @@ fn attached_has_table(conn: &Connection, schema: &str, table: &str) -> Result<bo
/// # What travels, and what is recomputed
///
/// The rule this module already follows for the rest of the catalog: user
/// judgements travel, inference is rebuilt. Concretely (docs/faces.md, and the
/// judgements travel, inference is rebuilt. Concretely (docs/dev/faces.md, and the
/// asymmetry `crate::faces` opens with):
///
/// - **People** — uuid, name, and whether the user set them aside. Merged by
@@ -991,9 +1051,13 @@ fn remote_has_column(tx: &Connection, table: &str, column: &str) -> Result<bool,
/// Remote face row id to local face row id, by photograph and box overlap.
///
/// See [`merge_people_within`] for why a face has no shared identity and this
/// has to be derived. Only faces from the same model are compared: boxes from
/// two different detectors are not the same measurement, and matching across
/// them would attach a judgement to a face nobody looked at.
/// has to be derived. Faces are compared within an *embedder*
/// (`faces::embedder_of`), not within an exact pipeline id: two detectors in
/// front of the same embedder draw boxes around the same faces, and a
/// confirmation made on one device's box is about the face, not the
/// rectangle — the same judgement `faces::record_detections` makes when it
/// carries a confirmation across a re-detection. Keying on the exact id was
/// what let a detector change strand every name on the device that made it.
fn match_faces(tx: &Connection) -> Result<std::collections::HashMap<i64, i64>, CatalogError> {
/// Loose on purpose — "the same face in the frame", not "the same
/// rectangle". The figure `record_detections` uses for the same job.
@@ -1026,7 +1090,8 @@ fn match_faces(tx: &Connection) -> Result<std::collections::HashMap<i64, i64>, C
})?;
for row in rows {
let (file_id, model, boxed) = row?;
local.entry((file_id, model)).or_default().push(boxed);
let embedder = crate::faces::embedder_of(&model).to_string();
local.entry((file_id, embedder)).or_default().push(boxed);
}
}
if local.is_empty() {
@@ -1056,7 +1121,8 @@ fn match_faces(tx: &Connection) -> Result<std::collections::HashMap<i64, i64>, C
for row in rows {
let (remote_id, file_id, model, rbox) = row?;
let Some(candidates) = local.get(&(file_id, model)) else {
let embedder = crate::faces::embedder_of(&model).to_string();
let Some(candidates) = local.get(&(file_id, embedder)) else {
continue;
};
let best = candidates
@@ -1162,6 +1228,57 @@ mod tests {
// ---- integration over two real catalogs ------------------------------
/// A fresh device takes the capture dates a peer's sweep read, matched by
/// `oc:fileid`, and never overwrites a date it read for itself.
#[test]
fn capture_metadata_arrives_for_undated_images_only() {
let c = two_catalogs();
// Three photographs on both devices: 1 undated here and dated there;
// 2 dated on both, differently; 3 undated on both.
for id in 1..=3 {
add_image_without_hash(&c, "main", id);
add_image_without_hash(&c, "remote_cat", id + 10);
add_remote_id(&c, "main", id, 100 + id);
add_remote_id(&c, "remote_cat", id + 10, 100 + id);
}
c.execute(
"UPDATE remote_cat.images
SET captured_at = 1000, captured_offset = 60, camera = 'X', metadata_state = 2
WHERE id = 11",
[],
)
.unwrap();
c.execute(
"UPDATE remote_cat.images SET captured_at = 2000, metadata_state = 2 WHERE id = 12",
[],
)
.unwrap();
c.execute(
"UPDATE main.images SET captured_at = 2222, metadata_state = 2 WHERE id = 2",
[],
)
.unwrap();
let report = merge_metadata(&c).unwrap();
assert_eq!(report.metadata_adopted, 1);
let row = |id: i64| -> (Option<i64>, Option<i64>, Option<String>, i64) {
c.query_row(
"SELECT captured_at, captured_offset, camera, metadata_state
FROM main.images WHERE id = ?1",
[id],
|r| Ok((r.get(0)?, r.get(1)?, r.get(2)?, r.get(3)?)),
)
.unwrap()
};
assert_eq!(row(1), (Some(1000), Some(60), Some("X".into()), 2));
assert_eq!(row(2), (Some(2222), None, None, 2));
assert_eq!(row(3), (None, None, None, 0));
// Idempotent: a second pass finds nothing left to take.
assert_eq!(merge_metadata(&c).unwrap().metadata_adopted, 0);
}
fn two_catalogs() -> Connection {
attached_remote(schema::for_attached("remote_cat"))
}
@@ -1977,6 +2094,54 @@ mod tests {
assert_eq!(person_of(&c, local), Some(("Anna".to_string(), true)));
}
/// The bug this rule exists for: the desktop switched to a stronger
/// detector and confirmed 3,500 faces under the old pipeline id; the
/// tablet held the same faces under the new one, and not one name
/// crossed, because the match demanded the exact id. Same photograph,
/// same box, same embedder — that is the same face.
#[test]
fn a_confirmation_crosses_a_detector_change() {
let c = two_catalogs();
for db in ["main", "remote_cat"] {
add_synced_image(&c, db, 1, 5000);
}
let local = add_face(&c, "main", 7, 1, 0.30);
c.execute(
"UPDATE main.faces SET model_id = 'scrfd_10g+w600k_mbf' WHERE id = ?1",
[local],
)
.unwrap();
let remote = add_face(&c, "remote_cat", 42, 1, 0.31);
add_person(&c, "remote_cat", 3, "u-anna", "Anna", false);
assign(&c, "remote_cat", remote, 3, true);
let report = merge_all(&c).unwrap();
assert_eq!(report.faces_assigned, 1);
assert_eq!(person_of(&c, local), Some(("Anna".to_string(), true)));
}
/// A different embedder is a different space, and a box there is a face
/// nobody here has a vector for.
#[test]
fn a_confirmation_does_not_cross_an_embedder_change() {
let c = two_catalogs();
for db in ["main", "remote_cat"] {
add_synced_image(&c, db, 1, 5000);
}
let local = add_face(&c, "main", 7, 1, 0.30);
c.execute(
"UPDATE main.faces SET model_id = 'scrfd_10g+other_embedder' WHERE id = ?1",
[local],
)
.unwrap();
let remote = add_face(&c, "remote_cat", 42, 1, 0.31);
add_person(&c, "remote_cat", 3, "u-anna", "Anna", false);
assign(&c, "remote_cat", remote, 3, true);
merge_all(&c).unwrap();
assert_eq!(person_of(&c, local), None, "matched across embedders");
}
/// Boxes from two devices are close but not identical. Matching has to be
/// by overlap, not equality, or nothing ever lines up.
#[test]
+171 -14
View File
@@ -49,6 +49,12 @@ pub struct Judgement {
/// 0..=5. Zero means *unrated*, which is a state in its own right.
pub rating: u8,
pub flag: FlagState,
/// TRACES: FR-CAT-5
/// The colour label, or `None`. Not part of [`Judgement::is_judged`]:
/// a label sorts photographs into piles of the photographer's own
/// meaning — "to print", "send to Anna" — and says nothing about whether
/// a frame has been culled, which is the question "unjudged" asks.
pub label: Option<ColourLabel>,
}
impl Judgement {
@@ -267,14 +273,6 @@ pub fn align_default_version_uuids(conn: &Connection) -> Result<usize, CatalogEr
Ok(moved)
}
/// The default version's row id for an image, creating one if it has none.
///
/// Every write path goes through this rather than assuming a version exists.
/// An image can arrive without one in two ways that are not worth trying to
/// prevent: a row inserted by a build predating this module, and a scan whose
/// version pass was interrupted between the image insert and the commit.
/// Failing a rating because of either would be the wrong answer — the user
/// pressed a key and expects a star.
/// TRACES: FR-CAT-13
/// How `versions.label` encodes a colour label, and back.
///
@@ -304,6 +302,14 @@ pub fn label_from_code(code: Option<i64>) -> Option<ColourLabel> {
})
}
/// The default version's row id for an image, creating one if it has none.
///
/// Every write path goes through this rather than assuming a version exists.
/// An image can arrive without one in two ways that are not worth trying to
/// prevent: a row inserted by a build predating this module, and a scan whose
/// version pass was interrupted between the image insert and the commit.
/// Failing a rating because of either would be the wrong answer — the user
/// pressed a key and expects a star.
pub fn default_version_id(conn: &Connection, image: ImageId) -> Result<i64, CatalogError> {
let existing: Option<i64> = conn
.query_row(
@@ -387,6 +393,88 @@ pub fn set_flag_many(
apply_many(conn, images, |conn, id| set_flag(conn, id, flag))
}
/// TRACES: FR-CAT-5
/// Set or clear the colour label for one image.
pub fn set_label(
conn: &Connection,
image: ImageId,
label: Option<ColourLabel>,
) -> Result<(), CatalogError> {
let version = default_version_id(conn, image)?;
conn.execute(
"UPDATE versions SET label = ?2 WHERE id = ?1",
rusqlite::params![version, label.map(label_code)],
)?;
Ok(())
}
/// TRACES: FR-CAT-5
/// Set or clear a label on many images in one transaction — one keystroke
/// over a selection is one commit, as for [`set_rating_many`].
pub fn set_label_many(
conn: &Connection,
images: &[ImageId],
label: Option<ColourLabel>,
) -> Result<usize, CatalogError> {
apply_many(conn, images, |conn, id| set_label(conn, id, label))
}
/// TRACES: FR-CAT-5
/// What a label key does to a set of images: Lightroom's toggle.
///
/// Pressing the key for the label every one of them already carries takes it
/// off; otherwise every one of them gets it. Decided over the whole set
/// rather than per image, so a selection that was half red comes out all red
/// rather than inverted — the photographer pressed "red", and a key that
/// turned half of them red and the other half plain would be two answers to
/// one question.
pub fn toggled_label(
current: impl IntoIterator<Item = Option<ColourLabel>>,
pressed: ColourLabel,
) -> Option<ColourLabel> {
let mut any = false;
for label in current {
any = true;
if label != Some(pressed) {
return Some(pressed);
}
}
if any {
None
} else {
Some(pressed)
}
}
/// TRACES: FR-CAT-5 | FR-CAT-6
/// How the library divides by colour label, for the filter chips' counts.
///
/// Index 0 is unlabelled and index `n` the label whose code is `n`. One
/// grouped statement — the same shape as [`rating_histogram`], and for the
/// same reason it LEFT JOINs: an image without a version row is unlabelled,
/// not missing.
pub fn label_histogram(conn: &Connection) -> Result<[usize; 6], CatalogError> {
let mut out = [0usize; 6];
let mut stmt = conn.prepare(
"SELECT coalesce(v.label, 0) AS l, count(*)
FROM images i
LEFT JOIN versions v ON v.image_id = i.id AND v.is_default = 1
GROUP BY l",
)?;
let rows = stmt.query_map([], |r| Ok((r.get::<_, i64>(0)?, r.get::<_, i64>(1)?)))?;
for (code, count) in rows.flatten() {
// A code this build does not know counts as unlabelled, which is how
// `label_from_code` reads it everywhere else.
let slot = if label_from_code(Some(code)).is_some() {
code as usize
} else {
0
};
out[slot] += count as usize;
}
Ok(out)
}
/// Shared bulk wrapper, so the two axes cannot drift in their commit
/// behaviour — a partially-committed rating and a fully-committed flag from
/// the same keystroke would be hard to explain and harder to notice.
@@ -411,21 +499,22 @@ fn apply_many(
/// An image with no version reads as unrated and unflagged rather than as an
/// error: that is exactly what it is.
pub fn judgement(conn: &Connection, image: ImageId) -> Result<Judgement, CatalogError> {
let row: Option<(i64, i64)> = conn
let row: Option<(i64, i64, Option<i64>)> = conn
.query_row(
"SELECT rating, flag FROM versions
"SELECT rating, flag, label FROM versions
WHERE image_id = ?1
ORDER BY is_default DESC, id ASC
LIMIT 1",
[image.0 as i64],
|r| Ok((r.get(0)?, r.get(1)?)),
|r| Ok((r.get(0)?, r.get(1)?, r.get(2)?)),
)
.optional()?;
Ok(match row {
Some((rating, flag)) => Judgement {
Some((rating, flag, label)) => Judgement {
rating: rating.clamp(0, MAX_RATING as i64) as u8,
flag: flag_from_code(flag),
label: label_from_code(label),
},
None => Judgement::default(),
})
@@ -452,7 +541,7 @@ pub fn judgements(
.collect::<Vec<_>>()
.join(",");
let sql = format!(
"SELECT image_id, rating, flag FROM versions
"SELECT image_id, rating, flag, label FROM versions
WHERE image_id IN ({placeholders}) AND is_default = 1"
);
@@ -467,15 +556,17 @@ pub fn judgements(
r.get::<_, i64>(0)?,
r.get::<_, i64>(1)?,
r.get::<_, i64>(2)?,
r.get::<_, Option<i64>>(3)?,
))
})?;
for (image, rating, flag) in rows.flatten() {
for (image, rating, flag, label) in rows.flatten() {
out.insert(
ImageId(image as u64),
Judgement {
rating: rating.clamp(0, MAX_RATING as i64) as u8,
flag: flag_from_code(flag),
label: label_from_code(label),
},
);
}
@@ -686,6 +777,72 @@ mod tests {
assert_eq!(distinct, 200);
}
#[test]
fn a_label_round_trips_and_clears() {
// TRACES: FR-CAT-5
let cat = with_images(1);
let id = ids(&cat)[0];
set_label(cat.connection(), id, Some(ColourLabel::Green)).unwrap();
assert_eq!(
judgement(cat.connection(), id).unwrap().label,
Some(ColourLabel::Green)
);
set_label(cat.connection(), id, None).unwrap();
assert_eq!(judgement(cat.connection(), id).unwrap().label, None);
}
#[test]
fn a_label_is_not_a_judgement() {
// "Unjudged" is the cull's resume point; a label is a pile of the
// photographer's own, and labelling a frame must not hide it there.
let cat = with_images(1);
let id = ids(&cat)[0];
set_label(cat.connection(), id, Some(ColourLabel::Red)).unwrap();
assert!(!judgement(cat.connection(), id).unwrap().is_judged());
}
#[test]
fn labelling_a_selection_is_one_commit_and_reaches_every_image() {
// TRACES: FR-CAT-5
let cat = with_images(4);
let all = ids(&cat);
assert_eq!(
set_label_many(cat.connection(), &all, Some(ColourLabel::Blue)).unwrap(),
4
);
let found = judgements(cat.connection(), &all).unwrap();
assert!(all
.iter()
.all(|id| found[id].label == Some(ColourLabel::Blue)));
assert_eq!(
label_histogram(cat.connection()).unwrap(),
[0, 0, 0, 0, 4, 0]
);
}
#[test]
fn a_label_key_toggles_only_when_every_image_already_has_it() {
// TRACES: FR-CAT-5
use ColourLabel::*;
assert_eq!(toggled_label([Some(Red), Some(Red)], Red), None);
assert_eq!(toggled_label([Some(Red), None], Red), Some(Red));
assert_eq!(toggled_label([Some(Blue)], Red), Some(Red));
assert_eq!(toggled_label([], Red), Some(Red));
}
#[test]
fn the_label_histogram_sums_to_the_library() {
// TRACES: FR-CAT-6
// Images without a version row count as unlabelled rather than
// vanishing, as the rating histogram's do.
let cat = with_images(3);
let first = ids(&cat)[0];
set_label(cat.connection(), first, Some(ColourLabel::Purple)).unwrap();
let h = label_histogram(cat.connection()).unwrap();
assert_eq!(h, [2, 0, 0, 0, 0, 1]);
assert_eq!(h.iter().sum::<usize>(), 3);
}
#[test]
fn a_rating_round_trips() {
let cat = with_images(1);
+2 -2
View File
@@ -16,7 +16,7 @@
//! # The one thing a rebuild does not recover
//!
//! **Collections.** A manual collection is a set of images the user assembled
//! by hand and nothing in the filesystem records it (`docs/catalog.md` §8.1) —
//! by hand and nothing in the filesystem records it (`docs/dev/catalog.md` §8.1) —
//! which is the whole reason the catalog file itself syncs. So the two offers
//! are not interchangeable, and the interface must not present them as if they
//! were: a restore keeps the user's collections, a rebuild does not.
@@ -570,7 +570,7 @@ mod tests {
// The first NFR-R6 branch, asserted on the thing that distinguishes it
// from the second: a collection exists nowhere but the catalog, so it
// is the evidence that the *contents* came back and not merely a
// readable file (docs/catalog.md §8.1).
// readable file (docs/dev/catalog.md §8.1).
let dir = tempdir("restore");
let path = dir.join("catalog.sqlite");
fixture(&path, 500);
+468 -14
View File
@@ -15,7 +15,7 @@ use rusqlite::Connection;
use crate::error::CatalogError;
/// Schema version this build writes and understands.
pub const SCHEMA_VERSION: i64 = 15;
pub const SCHEMA_VERSION: i64 = 20;
/// Apply migrations up to [`SCHEMA_VERSION`].
///
@@ -143,9 +143,157 @@ pub fn migrate(conn: &Connection) -> Result<i64, CatalogError> {
tx.commit()?;
}
if from < 16 {
let tx = conn.unchecked_transaction()?;
// Guarded like V14's column, and for the same reason: `ALTER TABLE
// ... ADD COLUMN` has no `IF NOT EXISTS`, and this step must be
// re-enterable (NFR-R5).
for column in EYE_COLUMNS {
let present: bool = tx
.prepare("SELECT 1 FROM pragma_table_info('faces') WHERE name = ?1")?
.exists([column])?;
if !present {
tx.execute_batch(&format!("ALTER TABLE faces ADD COLUMN {column} REAL;"))?;
}
}
tx.pragma_update(None, "user_version", 16)?;
tx.commit()?;
}
if from < 17 {
let tx = conn.unchecked_transaction()?;
tx.execute_batch(V17)?;
tx.pragma_update(None, "user_version", 17)?;
tx.commit()?;
}
if from < 18 {
let tx = conn.unchecked_transaction()?;
// Guarded like V14's and V16's columns: ALTER has no IF NOT EXISTS
// and the step must be re-enterable (NFR-R5).
let present: bool = tx
.prepare("SELECT 1 FROM pragma_table_info('faces') WHERE name = 'landmarks_dense'")?
.exists([])?;
if !present {
tx.execute_batch("ALTER TABLE faces ADD COLUMN landmarks_dense BLOB;")?;
}
tx.pragma_update(None, "user_version", 18)?;
tx.commit()?;
}
if from < 19 {
let tx = conn.unchecked_transaction()?;
tx.execute_batch(V19)?;
tx.pragma_update(None, "user_version", 19)?;
tx.commit()?;
}
if from < 20 {
let tx = conn.unchecked_transaction()?;
v20_markers_name_the_detector_that_found_the_faces(&tx)?;
tx.pragma_update(None, "user_version", 20)?;
tx.commit()?;
}
Ok(from)
}
// V20 -- TRACES: FR-CAT-7
//
// Run markers that named the wrong detector, put right.
//
// `faces::record_updates` -- the write behind the quality, eye and crop
// passes -- re-marked an image under the pipeline the pass ran as, while
// the faces it had updated kept the id of the detector that found them.
// A marker of `scrfd_10g+w600k_mbf` over faces spelled `w600k_mbf` reads,
// to every consumer, as the thorough detector having examined the image:
// the upgrade repair skips it, and `face_shard::export_to_shards` selects
// its faces by the marker's id, finds none, and tells every other device
// that the thorough detector found nothing there. The desktop's shard index
// held 54 such entries over photographs with named faces, and the tablet's
// eye pass over faces it had adopted from the desktop had made 430 more.
//
// The write is fixed to keep the marker under the faces' own id. This puts
// the markers already written right, with a fresh time so the export sends
// each image again under an entry newer than the empty one -- which is what
// `held_model` orders by. Where the right marker is still there beside the
// wrong one (the old write inserted rather than replaced), the wrong one
// goes and the right one is refreshed for the same reason: its entry in
// the shards is older than the empty one, and a device that has neither
// would take the empty one. An image V14 left with faces and no marker at
// all is not touched: that state is the quality pass's cue, and the fixed
// write marks it correctly when the pass reaches it.
//
// Restated in Rust rather than SQL because the embedder half of a pipeline
// id is `faces::embedder_sql`, which this must agree with.
fn v20_markers_name_the_detector_that_found_the_faces(tx: &Connection) -> Result<(), CatalogError> {
let fi = crate::faces::embedder_sql("face_index.model_id");
let f = crate::faces::embedder_sql("f.model_id");
// A marker is wrong when the image holds faces of its embedder under
// another id. First the wrong ones that sit beside a right one -- the
// update below would collide with it -- then the rest are renamed.
let wrong = format!(
"EXISTS (SELECT 1 FROM faces f
WHERE f.image_id = face_index.image_id
AND {f} = {fi}
AND f.model_id != face_index.model_id)"
);
let found_by = format!(
"(SELECT MIN(f.model_id) FROM faces f
WHERE f.image_id = face_index.image_id AND {f} = {fi})"
);
let now = crate::faces::now_secs();
tx.execute(
&format!(
"UPDATE face_index
SET indexed_at = ?1
WHERE model_id = {found_by}
AND EXISTS (SELECT 1 FROM face_index w
WHERE w.image_id = face_index.image_id
AND w.model_id != face_index.model_id
AND {} = {fi})",
crate::faces::embedder_sql("w.model_id")
),
[now],
)?;
tx.execute(
&format!(
"DELETE FROM face_index
WHERE {wrong}
AND EXISTS (SELECT 1 FROM face_index o
WHERE o.image_id = face_index.image_id
AND o.model_id = {found_by})"
),
[],
)?;
tx.execute(
&format!(
"UPDATE face_index
SET model_id = {found_by},
faces_found = (SELECT COUNT(*) FROM faces f
WHERE f.image_id = face_index.image_id AND {f} = {fi}),
indexed_at = ?1
WHERE {wrong}"
),
[now],
)?;
Ok(())
}
/// The seven columns V16 adds to `faces`, in the order the readers name them.
///
/// Named once because three places have to agree on them: this migration,
/// [`for_attached`], and the face shard's own catch-up (`face_shard`).
pub const EYE_COLUMNS: [&str; 7] = [
"eye_right",
"eye_right_px",
"eye_right_sharp",
"eye_left",
"eye_left_px",
"eye_left_sharp",
"sunglasses",
];
/// Recompute columns a migration added, for rows that predate it.
///
/// A migration adds a column with a default; it cannot know what the value
@@ -208,11 +356,38 @@ pub fn backfill(conn: &Connection) -> Result<Vec<(&'static str, usize)>, Catalog
Ok(out)
}
/// How long a connection waits for a writer to finish before giving up.
///
/// TRACES: NFR-R1
/// SQLite's default is **zero** — the loser of a race gets `SQLITE_BUSY` at
/// once rather than a turn — and WAL does not change that for two writers. One
/// writer and many readers is the case WAL makes free; this is the other one,
/// and this application has it constantly: the face sweep commits a batch while
/// reclustering reads, the derived sync imports shards while the sweep writes.
///
/// Without a timeout that contention was *lost work*, not a retry. A face
/// sweep that had already paid for the detection and the embedding — the
/// expensive part, seconds per image — threw the result away on
/// `storing faces for 214: database is locked` and moved on, and both the
/// desktop and the tablet logged runs of those on consecutive images.
///
/// Ten seconds, matching the figure the job runner's tests already use for the
/// same reason. It is far longer than any transaction here (a sweep batch is
/// sub-second; the slowest is a WAL checkpoint of a 130 MB catalog), so in
/// practice it is a bound on pathology rather than a wait anyone sits through.
/// The tension with NFR-P9 is real but one-sided: a query on the UI thread
/// would rather wait for its turn than fail, because the failure is what the
/// user sees as "cannot open catalog".
const BUSY_TIMEOUT: std::time::Duration = std::time::Duration::from_secs(10);
/// Connection setup applied on every open, migration or not.
///
/// WAL is required by NFR-R1: it survives power loss without corruption, and
/// it lets a background job write while the grid reads.
pub fn configure(conn: &Connection) -> Result<(), CatalogError> {
// Before the pragmas, so that a connection racing a migration waits for it
// rather than failing on the first statement it tries.
conn.busy_timeout(BUSY_TIMEOUT)?;
conn.pragma_update(None, "journal_mode", "WAL")?;
// NORMAL rather than FULL: with WAL this is durable across process death
// (which is what FR-PLAT-AND-3 cares about) and only risks the last
@@ -276,7 +451,15 @@ pub fn for_attached(schema_name: &str) -> String {
"{}\n{}\n{}\n\
ALTER TABLE {schema_name}.people ADD COLUMN ignored INTEGER NOT NULL DEFAULT 0;\n\
ALTER TABLE {schema_name}.faces ADD COLUMN crop BLOB;\n\
ALTER TABLE {schema_name}.faces ADD COLUMN quality REAL;",
ALTER TABLE {schema_name}.faces ADD COLUMN quality REAL;\n\
ALTER TABLE {schema_name}.faces ADD COLUMN eye_right REAL;\n\
ALTER TABLE {schema_name}.faces ADD COLUMN eye_right_px REAL;\n\
ALTER TABLE {schema_name}.faces ADD COLUMN eye_right_sharp REAL;\n\
ALTER TABLE {schema_name}.faces ADD COLUMN eye_left REAL;\n\
ALTER TABLE {schema_name}.faces ADD COLUMN eye_left_px REAL;\n\
ALTER TABLE {schema_name}.faces ADD COLUMN eye_left_sharp REAL;\n\
ALTER TABLE {schema_name}.faces ADD COLUMN sunglasses REAL;\n\
ALTER TABLE {schema_name}.faces ADD COLUMN landmarks_dense BLOB;",
rewrite_for_attached(V1, schema_name),
rewrite_for_attached(V6, schema_name),
rewrite_for_attached(V8, schema_name),
@@ -614,20 +797,20 @@ const V14: &str = r#"
-- face this rule is not yet protecting anyone from, and the only way to
-- measure it is to embed it again.
--
-- The sweep's measuring pass is what does that: `dr_ui::library::
-- faces_unmeasured` lists every image holding a face with no reading, and
-- each face is embedded again from the native render with the landmarks it
-- already has, the raw vector written over the old one (`record_measurements`)
-- and nothing else touched -- not the id, not the box, not who the user said
-- it was. The faces keep drawing the People screen throughout.
-- The `face-quality` repair is what does that (`dr_ui::repairs`, once the
-- sweep's measuring pass): it lists every face with no reading, and each is
-- embedded again from the native render with the landmarks it already has,
-- the raw vector written over the old one (`record_updates`) and nothing
-- else touched -- not the id, not the box, not who the user said it was.
-- The faces keep drawing the People screen throughout.
--
-- The run markers of those images are forgotten too, exactly as V12 forgot
-- the runs made against too small a proxy. The build this shipped in had no
-- measuring pass yet, and a marker is the one thing that stops a face ever
-- being looked at again; with the pass in place `faces_unindexed` leaves
-- these images to it rather than detecting them from scratch, so the
-- deletion costs nothing -- and an image that was examined and found empty
-- keeps its marker, since there is nothing on it to measure.
-- being looked at again; with the repair in place, detection leaves an
-- image holding this embedder's faces to it rather than detecting from
-- scratch, so the deletion costs nothing -- and an image that was examined
-- and found empty keeps its marker, since there is nothing on it to measure.
--
-- The cost is a re-fetch of every image with a face on it, on the next pass
-- the user starts. That is a whole-library transfer (FR-NC-6), and it starts
@@ -679,6 +862,110 @@ CREATE TABLE IF NOT EXISTS xmp_conflicts (
);
"#;
// V19 -- TRACES: NFR-P9
//
// The indexes the repair counts are served from, and V17's lesson applied
// to the rest of the face columns.
//
// "How many images still owe a quality reading" was answered per image: a
// correlated EXISTS over `faces` that had to open each face's row to look
// at one nullable column -- the row being eight kilobytes of embedding and
// crop. Six such counts run every time the Identity screen opens and every
// time a sweep ends, 160 ms of them on the reference library. Three
// partial indexes hold only the faces still owing each pass, keyed by the
// image and carrying the model id the predicate also reads, so the count
// walks a few thousand index entries and touches no row at all -- and each
// index shrinks to nothing as its pass completes. The planner takes them
// when the count is driven from `faces` (`repairs::count`) and ignores
// them inside the per-image EXISTS, which is why that function has two
// spellings of the same predicate.
//
// `faces_image_model` replaces `faces_image`: the same key with the model
// id beside it, so "does this image hold this embedder's faces" -- asked in
// the audit, the proxy repair and the outstanding-detection count -- is an
// index-only probe where it used to read the row for the model id. Every
// lookup that used `faces_image` is served by its prefix.
//
// Not applied to attached catalogs, like V7 and V17: an index is a local
// concern, and a merge never runs these queries across an attachment.
const V19: &str = r#"
CREATE INDEX IF NOT EXISTS faces_image_model ON faces(image_id, model_id);
DROP INDEX IF EXISTS faces_image;
CREATE INDEX IF NOT EXISTS faces_owed_quality ON faces(image_id, model_id)
WHERE quality IS NULL;
CREATE INDEX IF NOT EXISTS faces_owed_crop ON faces(image_id, model_id)
WHERE crop IS NULL;
CREATE INDEX IF NOT EXISTS faces_owed_eyes ON faces(image_id, model_id)
WHERE eye_right IS NULL OR landmarks_dense IS NULL;
"#;
// V18 -- TRACES: FR-CULL-8a | FR-CULL-12
//
// The 106 dense landmarks the eye pass reads its eye boxes from, kept beside
// the reading as `dr_face::Landmarks::to_packed_bytes`: 106 x (x, y) as
// 16-bit fixed point over the frame, 424 bytes a face, a seventh of a
// pixel on a 6000-pixel frame. Derived data under FR-CULL-12 -- rebuilt by
// re-reading, never in a sidecar -- and stored for the same reason the
// embedding is: it cost a fetch of the original and a model run, and the
// next per-face pass (head pose, expression) should not have to pay either
// again. NULL where the face was never read.
//
// Added in `migrate`, guarded, like every ALTER here (NFR-R5).
const V17: &str = r#"
-- TRACES: FR-CULL-8a | FR-CULL-13 | NFR-P9
-- The eyes-open filter's index, and a lesson about where a column lands.
--
-- The people filter is a correlated EXISTS over `faces` per image, and it
-- was fast because `faces_image` *covers* it: the subquery never touched a
-- row. Reading V16's seven eye columns in the same subquery did touch the
-- row -- and `ALTER TABLE ADD COLUMN` puts a column at the end of the
-- record, after the 1 KB embedding and the ~5 KB crop, so every check
-- dragged six kilobytes off disk to reach seven floats. Measured on the
-- reference library: 24 seconds for one count, thirteen of them system
-- time. With this index the same count takes five milliseconds, because
-- the subquery is served from the index again and never reads a row.
--
-- The columns are listed in EYE_COLUMNS' order behind `image_id`, which is
-- the key the subquery searches on. Nothing else changed in V17; a catalog
-- already at V16 needs only this.
CREATE INDEX IF NOT EXISTS faces_eyes ON faces(
image_id, eye_right, eye_right_px, eye_right_sharp,
eye_left, eye_left_px, eye_left_sharp, sunglasses
);
"#;
// V16 -- TRACES: FR-CULL-8a
//
// What each face's eyes are doing: for each eye P(open), the source pixels
// across its box and the sharpness of the patch the classifier saw; and
// P(sunglasses) for the head. Seven numbers rather than a verdict, because
// the verdict is a rule with thresholds in it (dr_face::eyes::EyeReading::
// state) and a rule belongs in code that can be changed, not in rows that
// would have to be re-measured.
//
// The pixels and the sharpness are what stop a smear reading as a blink: an
// eye too small or too soft to read is not asked, and a face with no
// readable eye is "unclear", which no filter drops. Sunglasses are a column
// of their own for the same kind of reason — the eye classifier answers
// confidently over dark glass, and its answer means nothing there. A filter
// for "eyes open" reads all seven.
//
// NULL means "never measured" -- a face indexed before this version, or on a
// device without the eye models -- and a NULL is left alone by every filter
// that reads these, so an old library does not empty its grid the moment the
// chip is pressed. The sweep's measuring pass fills them in, from the native
// render, with the landmarks already stored: the same pass V14 built for the
// embedding's length, extended to ask the eye models too. No run marker is
// forgotten here, for the reason V14's note gives -- the measuring pass
// finds its own work by the NULL, and deleting markers would only put the
// detector back over images it has finished with.
//
// The columns are added in `migrate`, guarded, because ALTER has no IF NOT
// EXISTS and the step has to be re-enterable (NFR-R5). Their names are
// `EYE_COLUMNS`.
const V9: &str = r#"
-- TRACES: FR-CULL-8
-- A record that face detection has *run* on an image, distinct from what it
@@ -722,7 +1009,7 @@ CREATE INDEX face_index_model ON face_index(model_id);
const V8: &str = r#"
-- TRACES: FR-CULL-8 | FR-CULL-9 | FR-CULL-10 | FR-CULL-11 | FR-CULL-12 | NFR-SEC-5
-- People and faces (docs/faces.md, docs/catalog.md §10).
-- People and faces (docs/dev/faces.md, docs/dev/catalog.md §10).
--
-- Everything here is **derived data** except one column. Faces, landmarks,
-- embeddings, cluster assignments and suggestions are all reproducible by
@@ -759,7 +1046,7 @@ CREATE TABLE faces (
landmarks BLOB NOT NULL, -- 5 x (x, y) f32, normalised likewise
detector_confidence REAL NOT NULL,
embedding BLOB NOT NULL, -- 512 x f16; unit length until V14, raw since
-- Source pixels across the aligned 112x112 crop (docs/faces.md §7).
-- Source pixels across the aligned 112x112 crop (docs/dev/faces.md §7).
--
-- Not cosmetic: it is the honest quality signal for the UI, a feature in
-- the §8 calibration -- FR-CULL-9 names face size as an axis along which an
@@ -1148,6 +1435,60 @@ CREATE INDEX jobs_ready ON jobs(state, priority DESC, not_before);
#[cfg(test)]
mod tests {
#[test]
fn a_writer_waits_for_its_turn_rather_than_losing_its_work() {
// The failure this exists for: a face sweep that had already paid for
// the detection and the embedding threw the result away on
// "database is locked" and moved on. WAL does not help here — it makes
// one writer and many readers free, and this is two writers.
let dir = std::env::temp_dir().join(format!(
"dr-busy-{}-{:?}",
std::process::id(),
std::thread::current().id()
));
let _ = std::fs::remove_dir_all(&dir);
std::fs::create_dir_all(&dir).unwrap();
let path = dir.join("catalog.sqlite");
let held = rusqlite::Connection::open(&path).unwrap();
configure(&held).unwrap();
migrate(&held).unwrap();
let other = rusqlite::Connection::open(&path).unwrap();
configure(&other).unwrap();
// Every connection carries the timeout, which is what makes the wait
// below a wait rather than an immediate error.
let timeout: i64 = other
.query_row("PRAGMA busy_timeout", [], |r| r.get(0))
.unwrap();
assert_eq!(timeout, BUSY_TIMEOUT.as_millis() as i64);
// A writer holds the database; the other one must still get its turn
// once the first commits, rather than failing at the moment it asks.
let writing = held.unchecked_transaction().unwrap();
held.execute(
"INSERT INTO roots(id, kind, label) VALUES (1, 'local', 'lib')",
[],
)
.unwrap();
let handle = std::thread::spawn(move || {
other.execute(
"INSERT INTO roots(id, kind, label) VALUES (2, 'local', 'two')",
[],
)
});
std::thread::sleep(std::time::Duration::from_millis(150));
writing.commit().unwrap();
assert!(
handle.join().unwrap().is_ok(),
"the second writer waited and then wrote, rather than erroring"
);
let _ = std::fs::remove_dir_all(&dir);
}
use super::*;
fn mem() -> Connection {
@@ -1435,6 +1776,35 @@ mod tests {
assert_eq!(migrate(&c).unwrap(), SCHEMA_VERSION);
}
/// V16 adds its columns guarded, so a catalog whose version was rewound
/// after the columns landed — the rollback NFR-R5 contemplates — migrates
/// again rather than failing on "duplicate column".
#[test]
fn the_eye_columns_survive_a_rewound_version() {
let c = mem();
migrate(&c).unwrap();
for column in EYE_COLUMNS {
let present: bool = c
.prepare("SELECT 1 FROM pragma_table_info('faces') WHERE name = ?1")
.unwrap()
.exists([column])
.unwrap();
assert!(present, "{column} missing after migration");
}
c.pragma_update(None, "user_version", 15).unwrap();
assert_eq!(migrate(&c).unwrap(), 15);
let indexed: bool = c
.prepare("SELECT 1 FROM sqlite_master WHERE type = 'index' AND name = 'faces_eyes'")
.unwrap()
.exists([])
.unwrap();
assert!(indexed, "V17's covering index is there");
let v: i64 = c
.query_row("PRAGMA user_version", [], |r| r.get(0))
.unwrap();
assert_eq!(v, SCHEMA_VERSION);
}
#[test]
fn refuses_a_catalog_from_a_newer_build() {
let c = mem();
@@ -1573,6 +1943,90 @@ mod tests {
assert_eq!(faces, 2);
}
#[test]
fn v20_renames_markers_to_the_detector_that_found_the_faces() {
let c = mem();
c.pragma_update(None, "user_version", 0).unwrap();
migrate(&c).unwrap();
c.execute(
"INSERT INTO roots(id, kind, label) VALUES (1, 'local', 'test')",
[],
)
.unwrap();
c.execute(
"INSERT INTO images(id, root_id, source_ref, added_at)
VALUES (1,1,'a',0),(2,1,'b',0),(3,1,'c',0),(4,1,'d',0),(5,1,'e',0)",
[],
)
.unwrap();
// 1: the desktop's case -- old faces, re-marked as thorough.
// 2: the tablet's case -- adopted thorough faces, re-marked int8,
// and the right marker still beside it (refreshed, so it is
// exported again over the empty entry).
// 3: right already. 4: examined and empty. 5: V14's state, faces
// and no marker.
for (image, model) in [
(1, "scrfd_10g+w600k_mbf"),
(2, "scrfd_10g_i8+w600k_mbf"),
(2, "scrfd_10g+w600k_mbf"),
(3, "scrfd_10g+w600k_mbf"),
(4, "scrfd_10g+w600k_mbf"),
] {
c.execute(
"INSERT INTO face_index(image_id, model_id, indexed_at, faces_found, source_edge)
VALUES (?1, ?2, 100, 0, 6000)",
rusqlite::params![image, model],
)
.unwrap();
}
for (image, model) in [
(1, "w600k_mbf"),
(1, "w600k_mbf"),
(2, "scrfd_10g+w600k_mbf"),
(3, "scrfd_10g+w600k_mbf"),
(5, "w600k_mbf"),
] {
c.execute(
"INSERT INTO faces
(image_id, x, y, w, h, landmarks, detector_confidence, embedding,
crop_px, model_id, detected_at)
VALUES (?1, 0.1, 0.1, 0.2, 0.2, X'00', 0.9, X'00', 180.0, ?2, 0)",
rusqlite::params![image, model],
)
.unwrap();
}
c.pragma_update(None, "user_version", 19).unwrap();
migrate(&c).unwrap();
let markers: Vec<(i64, String, i64, bool)> = c
.prepare(
"SELECT image_id, model_id, faces_found, indexed_at > 100
FROM face_index ORDER BY image_id, model_id",
)
.unwrap()
.query_map([], |r| Ok((r.get(0)?, r.get(1)?, r.get(2)?, r.get(3)?)))
.unwrap()
.map(Result::unwrap)
.collect();
assert_eq!(
markers,
vec![
(1, "w600k_mbf".to_string(), 2, true),
(2, "scrfd_10g+w600k_mbf".to_string(), 0, true),
(3, "scrfd_10g+w600k_mbf".to_string(), 0, false),
(4, "scrfd_10g+w600k_mbf".to_string(), 0, false),
]
);
// Re-enterable: nothing left to rename.
c.pragma_update(None, "user_version", 19).unwrap();
migrate(&c).unwrap();
let n: i64 = c
.query_row("SELECT count(*) FROM face_index", [], |r| r.get(0))
.unwrap();
assert_eq!(n, 4);
}
#[test]
fn job_uniqueness_coalesces_rather_than_duplicating() {
let c = mem();
+221
View File
@@ -0,0 +1,221 @@
//! TRACES: S15 | FR-MRG-3
//! Spike S15.1 — does rawler read back a linear DNG this application writes?
//!
//! cargo run -p dr-decode --example linear_dng [-- <out.dng>]
//!
//! Decides FR-MRG-3's container. A panorama composite is three linear samples
//! per pixel with a camera matrix attached, which is exactly what a
//! `LinearRaw` DNG is; if rawler parses one, the composite re-enters the
//! library as `Format::Dng` and the only new decode work is a `cpp == 3`
//! branch. If it does not, the container is a float TIFF with a decode path
//! of its own.
//!
//! The file is hand-rolled rather than written with the `tiff` crate, whose
//! encoder fixes `PhotometricInterpretation` to RGB and cannot say
//! `LinearRaw`. Eighty lines of IFD is the cheaper thing to own than a fork.
use rawler::rawsource::RawSource;
const W: u32 = 64;
const H: u32 = 48;
fn main() {
let bytes = write_linear_dng(W, H);
if let Some(path) = std::env::args().nth(1) {
std::fs::write(&path, &bytes).expect("write");
println!("wrote {path} ({} bytes)", bytes.len());
}
let source = RawSource::new_from_slice(&bytes);
let decoder = match rawler::get_decoder(&source) {
Ok(d) => d,
Err(e) => {
println!("FAIL get_decoder: {e}");
std::process::exit(1);
}
};
println!("ok decoder found");
let image = match decoder.raw_image(&source, &Default::default(), false) {
Ok(i) => i,
Err(e) => {
println!("FAIL raw_image: {e}");
std::process::exit(1);
}
};
println!(
"ok raw_image: {}×{}, cpp {}, bps {}, {} samples, make {:?} model {:?}",
image.width,
image.height,
image.cpp,
image.bps,
match &image.data {
rawler::RawImageData::Integer(v) => v.len(),
rawler::RawImageData::Float(v) => v.len(),
},
image.make,
image.model
);
println!(
" white {:?} black {:?} wb {:?}",
image.whitelevel.0,
image
.blacklevel
.levels
.iter()
.map(|r| r.n as f32 / r.d.max(1) as f32)
.collect::<Vec<_>>(),
image.wb_coeffs
);
// The pixel at (1, 0) was written as (1000, 2000, 3000): if the samples
// come back interleaved in that order, cpp == 3 means what it says.
if let rawler::RawImageData::Integer(v) = &image.data {
let i = image.cpp;
println!(" pixel (1,0) = {:?}", &v[i..i + image.cpp.min(3)]);
}
// What dr-decode itself makes of it: the colour matrix rawler parsed into
// the camera definition, and the profile the decoder would build from it.
println!(" rawler color_matrix: {:?}", image.camera.color_matrix);
let dng = dr_decode::profile::read_dng_matrices(decoder.as_ref());
let profile = dr_decode::CameraProfile::extract(&image, &dng);
println!(
" CameraProfile: {}",
profile
.as_ref()
.map(|p| format!("xyz_to_cam {:?}", p.xyz_to_cam()))
.unwrap_or_else(|| "none".into())
);
match dr_decode::decode(&bytes) {
Ok(r) => println!(
"note dr_decode::decode accepted it as CFA: {}×{}, {} samples — the cpp==3 branch is the work",
r.width,
r.height,
r.data.len()
),
Err(e) => println!("note dr_decode::decode refused it: {e} — the cpp==3 branch is the work"),
}
}
/// A minimal `LinearRaw` DNG: one IFD, uncompressed 16-bit RGB, the tags a
/// decoder needs to treat it as a DNG and the matrix a develop chain needs
/// to treat it as a camera. Little-endian, one strip.
fn write_linear_dng(w: u32, h: u32) -> Vec<u8> {
// Pixels first, so their offset is known: a ramp with one marker pixel.
let mut pixels: Vec<u16> = Vec::with_capacity((w * h * 3) as usize);
for y in 0..h {
for x in 0..w {
if (x, y) == (1, 0) {
pixels.extend([1000, 2000, 3000]);
} else {
let v = ((x + y) * 512).min(65535) as u16;
pixels.extend([v, v / 2, v / 3]);
}
}
}
let pixel_bytes: Vec<u8> = pixels.iter().flat_map(|v| v.to_le_bytes()).collect();
// Layout: header (8) | pixels | extra data | IFD.
let pixels_off = 8u32;
let extra_off = pixels_off + pixel_bytes.len() as u32;
// Values that do not fit in four bytes go in `extra`, and the entry
// points at them.
let mut extra: Vec<u8> = Vec::new();
let mut entries: Vec<(u16, u16, u32, [u8; 4])> = Vec::new();
fn short(tag: u16, v: u16) -> (u16, u16, u32, [u8; 4]) {
let mut b = [0u8; 4];
b[..2].copy_from_slice(&v.to_le_bytes());
(tag, 3, 1, b)
}
fn long(tag: u16, v: u32) -> (u16, u16, u32, [u8; 4]) {
(tag, 4, 1, v.to_le_bytes())
}
fn ascii(extra: &mut Vec<u8>, extra_off: u32, tag: u16, s: &str) -> (u16, u16, u32, [u8; 4]) {
let mut bytes = s.as_bytes().to_vec();
bytes.push(0);
let off = extra_off + extra.len() as u32;
extra.extend(&bytes);
(tag, 2, bytes.len() as u32, off.to_le_bytes())
}
entries.push(long(254, 0)); // NewSubfileType: main image
entries.push(long(256, w));
entries.push(long(257, h));
// BitsPerSample ×3 — three shorts, six bytes, so out of line.
{
let off = extra_off + extra.len() as u32;
for _ in 0..3 {
extra.extend(16u16.to_le_bytes());
}
entries.push((258, 3, 3, off.to_le_bytes()));
}
entries.push(short(259, 1)); // Compression: none
entries.push(short(262, 34892)); // PhotometricInterpretation: LinearRaw
entries.push(ascii(&mut extra, extra_off, 271, "DarkRoom"));
entries.push(ascii(&mut extra, extra_off, 272, "Panorama"));
entries.push(long(273, pixels_off)); // StripOffsets
entries.push(short(274, 1)); // Orientation
entries.push(short(277, 3)); // SamplesPerPixel
entries.push(long(278, h)); // RowsPerStrip
entries.push(long(279, pixel_bytes.len() as u32)); // StripByteCounts
entries.push(short(284, 1)); // PlanarConfiguration: chunky
entries.push((50706, 1, 4, [1, 4, 0, 0])); // DNGVersion
entries.push((50707, 1, 4, [1, 4, 0, 0])); // DNGBackwardVersion
entries.push(ascii(&mut extra, extra_off, 50708, "DarkRoom Panorama")); // UniqueCameraModel
entries.push(long(50717, 65535)); // WhiteLevel
// ColorMatrix1: XYZ → camera, 9 SRATIONALs. A plausible sRGB-ish matrix
// (the inverse of the sRGB D65 primaries), scaled to integers.
{
let m: [(i32, i32); 9] = [
(32406, 10000),
(-15372, 10000),
(-4986, 10000),
(-9689, 10000),
(18758, 10000),
(415, 10000),
(557, 10000),
(-2040, 10000),
(10570, 10000),
];
let off = extra_off + extra.len() as u32;
for (n, d) in m {
extra.extend(n.to_le_bytes());
extra.extend(d.to_le_bytes());
}
entries.push((50721, 10, 9, off.to_le_bytes()));
}
// AsShotNeutral: 3 RATIONALs, neutral.
{
let off = extra_off + extra.len() as u32;
for _ in 0..3 {
extra.extend(1u32.to_le_bytes());
extra.extend(1u32.to_le_bytes());
}
entries.push((50728, 5, 3, off.to_le_bytes()));
}
entries.push(short(50778, 21)); // CalibrationIlluminant1: D65
entries.sort_by_key(|e| e.0);
let ifd_off = extra_off + extra.len() as u32;
let mut out = Vec::new();
out.extend(b"II");
out.extend(42u16.to_le_bytes());
out.extend(ifd_off.to_le_bytes());
out.extend(&pixel_bytes);
out.extend(&extra);
out.extend((entries.len() as u16).to_le_bytes());
for (tag, ty, count, value) in &entries {
out.extend(tag.to_le_bytes());
out.extend(ty.to_le_bytes());
out.extend(count.to_le_bytes());
out.extend(value);
}
out.extend(0u32.to_le_bytes()); // no next IFD
out
}
+121
View File
@@ -0,0 +1,121 @@
//! TRACES: FR-RAW-2
//! The seam a second decoder plugs into.
//!
//! D2 keeps LibRaw as the fallback for bodies rawler does not cover. Adding
//! it later should be a new `impl Decoder`, not an edit to every caller that
//! reads a header, cuts a thumbnail or opens a photograph for export — which
//! is what the free functions alone would have made it. So the callers take a
//! `&dyn Decoder`, and only the places that start a job name [`default`].
//!
//! Bytes in, always. Nothing here takes a path or a `SourceRef`: resolving a
//! file to bytes is `Storage`'s job at the caller, so the same decoder serves a
//! local file, an Android document and a range fetched from Nextcloud. The
//! decoder's part in that is to say how much of a file it needs
//! ([`Decoder::header_bytes`]) and where its preview sits
//! ([`Decoder::locate_preview`]); the storage layer fetches exactly that.
//!
//! What stays a free function is what is not a decoder's to vary: recognising
//! a JPEG ([`crate::probe`]), decoding one ([`crate::decode_jpeg`]) and
//! checking one is whole ([`crate::is_complete_jpeg`]). A second RAW decoder
//! would not read a JPEG differently.
use dr_types::Orientation;
use crate::{DecodeError, Metadata, Preview, PreviewLocation, PreviewSize, RawImage};
/// TRACES: FR-RAW-2
/// A RAW decoder, over bytes.
///
/// Object-safe so a caller can hold `&dyn Decoder` without becoming generic,
/// `Send + Sync` because the callers that need one most — the thumbnail
/// lanes, the export worker — run off the UI thread, and `Debug` so a job
/// description that carries one can still be printed.
pub trait Decoder: Send + Sync + std::fmt::Debug {
/// How much of the start of a file [`Self::metadata`] and
/// [`Self::locate_preview`] need. A caller reading over a network fetches
/// this range and no more.
fn header_bytes(&self) -> u64;
/// Capture metadata, from a header or a whole file, without touching
/// sensor data.
fn metadata(&self, bytes: &[u8]) -> Result<Metadata, DecodeError>;
/// How the stored pixels are turned, from a header. `None` where the file
/// does not say, which callers take as upright.
fn orientation(&self, header: &[u8]) -> Option<Orientation>;
/// Where the embedded preview best suited to a thumbnail sits in the file,
/// from its header, so a remote caller can fetch that range alone.
fn locate_preview(&self, header: &[u8], file_len: u64) -> Option<PreviewLocation>;
/// The embedded preview at the size asked for, falling through the ladder
/// to the next size where the file lacks it.
fn preview(&self, bytes: &[u8], size: PreviewSize) -> Result<Preview, DecodeError>;
/// Sensor data, for develop and export. The expensive path.
fn decode(&self, bytes: &[u8]) -> Result<RawImage, DecodeError>;
}
/// TRACES: FR-RAW-2
/// The decoder the application ships: rawler for sensor data and the
/// previews it knows, DarkRoom's own container walk for headers and ranges.
///
/// Its methods are the crate's free functions, unchanged. They stay public
/// for the tools and examples that read one file and have no caller to keep
/// decoder-agnostic.
#[derive(Debug, Clone, Copy, Default)]
pub struct Rawler;
impl Decoder for Rawler {
fn header_bytes(&self) -> u64 {
crate::HEADER_BYTES
}
fn metadata(&self, bytes: &[u8]) -> Result<Metadata, DecodeError> {
crate::metadata(bytes)
}
fn orientation(&self, header: &[u8]) -> Option<Orientation> {
crate::orientation(header)
}
fn locate_preview(&self, header: &[u8], file_len: u64) -> Option<PreviewLocation> {
crate::locate_preview(header, file_len)
}
fn preview(&self, bytes: &[u8], size: PreviewSize) -> Result<Preview, DecodeError> {
crate::extract_preview(bytes, size)
}
fn decode(&self, bytes: &[u8]) -> Result<RawImage, DecodeError> {
crate::decode(bytes)
}
}
/// TRACES: FR-RAW-2
/// The decoder a job uses unless it was handed another.
///
/// Named by the places that start work — a thread, a UI handler — and by
/// nothing below them. Returning `&'static dyn Decoder` rather than `Rawler`
/// is the point: a caller that only has this cannot reach past the trait.
pub fn default() -> &'static dyn Decoder {
static RAWLER: Rawler = Rawler;
&RAWLER
}
#[cfg(test)]
mod tests {
use super::*;
/// The default is the shipped decoder, reached through the trait: same
/// header budget, and the same answer to bytes neither can read.
#[test]
fn the_default_is_rawler_behind_the_trait() {
let d = default();
assert_eq!(d.header_bytes(), crate::HEADER_BYTES);
let junk = [0u8; 64];
assert_eq!(d.metadata(&junk).is_err(), crate::metadata(&junk).is_err());
assert!(d.decode(&junk).is_err());
assert_eq!(d.orientation(&junk), crate::orientation(&junk));
}
}
+60
View File
@@ -25,6 +25,42 @@ pub enum DecodeError {
CorruptPreview(String),
}
/// Run a decoder call, and return a panic inside it as an error.
///
/// TRACES: FR-RAW-4 | NFR-SEC-1 | NFR-R3
/// rawler `panic!`s on some malformed input rather than returning `Err` — a
/// DNG whose IFD claims a >50000 px image, for one, which is in the reference
/// library. A panic on a worker thread ends the thread: the face sweep that
/// met that file stopped 13 seconds in, three sweeps running, with "17301
/// image(s) to index" as the last word and nothing to say why. FR-RAW-4's
/// rule — a malformed file must not abort a batch — is this crate's to keep
/// whatever the library beneath it does, so every entry point that calls into
/// rawler runs through here, and a file that panics the decoder is one failed
/// file like any other.
///
/// The crash hook still records the panic, because it runs before unwinding
/// reaches this frame; that is right — it is a real defect in a dependency
/// and the record is how it gets reported upstream — and a repeat is the same
/// file being met again rather than a new fault.
pub(crate) fn guarded<T>(
what: &'static str,
f: impl FnOnce() -> Result<T, DecodeError>,
) -> Result<T, DecodeError> {
match std::panic::catch_unwind(std::panic::AssertUnwindSafe(f)) {
Ok(result) => result,
Err(payload) => {
let msg = payload
.downcast_ref::<&str>()
.map(|s| s.to_string())
.or_else(|| payload.downcast_ref::<String>().cloned())
.unwrap_or_else(|| "no message".to_string());
Err(DecodeError::Decode(format!(
"{what}: the decoder panicked on this file: {msg}"
)))
}
}
}
impl DecodeError {
/// Whether a fallback path might still produce an image.
///
@@ -49,4 +85,28 @@ mod tests {
// A genuinely unsupported file has nowhere to fall through to.
assert!(!DecodeError::Unsupported("unknown".into()).has_fallback());
}
#[test]
fn a_panic_in_the_decoder_is_an_error_and_the_thread_survives() {
// The property the face sweep relies on: one file that panics rawler
// is one failed file, not the end of the pass. The message travels,
// because "decode failed" alone sends the reader to the crash log.
let err = guarded("decode", || -> Result<(), DecodeError> {
panic!("rawler: surely there's no such thing as a {}MP image!", 600)
})
.unwrap_err();
let text = err.to_string();
assert!(text.contains("panicked"), "{text}");
assert!(text.contains("600MP"), "{text}");
assert!(!err.has_fallback(), "a panic is not a missing preview");
}
#[test]
fn a_result_passes_through_untouched() {
assert_eq!(guarded("decode", || Ok::<_, DecodeError>(7)).unwrap(), 7);
assert!(matches!(
guarded("decode", || Err::<(), _>(DecodeError::NoPreview)),
Err(DecodeError::NoPreview)
));
}
}
+69
View File
@@ -11,14 +11,20 @@
//!
//! Fusing them would force a full decode where a header read suffices, which
//! is exactly why Lightroom stalls ~2 s per image during culling.
//!
//! Callers reach these through the [`Decoder`] trait rather than by name, so a
//! second decoder can be put behind them without changing any of them
//! (FR-RAW-2). [`Rawler`] is the one that ships; [`default`] hands it out.
pub mod base_curve;
mod decoder;
mod error;
mod locate;
mod preview;
pub mod profile;
pub use base_curve::BaseCurve;
pub use decoder::{default, Decoder, Rawler};
pub use error::DecodeError;
pub use locate::{
defects, is_complete_jpeg, jpeg_metadata, locate_preview, tiff_metadata, BadLine, BadPixel,
@@ -134,6 +140,22 @@ pub struct RawImage {
pub base_curve: BaseCurve,
/// The usable region of `data`, excluding masked and border photosites.
pub crop: CropRect,
/// TRACES: FR-MRG-3
/// Samples per photosite in `data`: 1 for a colour-filter-array capture,
/// 3 for a *linear* DNG — demosaiced RGB, still camera-space, which is
/// what a merge writes. With 3, `cfa_pattern` means nothing, `data` is
/// `width × height × 3` interleaved, and the GPU uploads it as it is
/// rather than demosaicing.
pub samples_per_pixel: u8,
/// TRACES: FR-MRG-3
/// The body's colour profile as the file carried it, for a composite to
/// carry on: calibrations and the as-shot neutral. `None` for a body the
/// decoder has no matrix for.
pub profile: Option<profile::CameraProfile>,
/// The body, as rawler cleans the names: what `Make`/`Model` say and what
/// the base-curve database matches on.
pub make: String,
pub model: String,
}
/// TRACES: FR-RAW-3
@@ -285,6 +307,30 @@ pub fn probe(header: &[u8]) -> Option<Format> {
/// TRACES: FR-CAT-5 | M-12
/// Read capture metadata without decoding sensor data.
pub fn metadata(bytes: &[u8]) -> Result<Metadata, DecodeError> {
error::guarded("metadata", || metadata_unguarded(bytes))
}
/// TRACES: FR-CAT-5
/// Where a TIFF-shaped file keeps its first IFD, when the head handed to
/// [`metadata`] does not reach it — the linear DNG a merge writes puts its
/// IFDs after the pixels, and rawler, given the head alone, finds no
/// decoder in it. The caller fetches from this offset to the end and
/// reads the two ranges with [`metadata_split`].
pub fn trailing_ifd(head: &[u8]) -> Option<u64> {
locate::trailing_ifd(head)
}
/// TRACES: FR-CAT-5
/// [`metadata`] for a file read in two ranges: `head` from offset 0 and
/// `tail` from `tail_at`. The EXIF sub-IFD such a file wrote before its
/// pixels is in the head; the first IFD and its values are in the tail.
pub fn metadata_split(head: &[u8], tail: &[u8], tail_at: u64) -> Result<Metadata, DecodeError> {
error::guarded("metadata", || {
locate::tiff_metadata_split(head, tail, tail_at)
})
}
fn metadata_unguarded(bytes: &[u8]) -> Result<Metadata, DecodeError> {
use rawler::rawsource::RawSource;
// rawler has no decoder for a plain JPEG, so without this every JPEG in a
@@ -510,6 +556,10 @@ pub(crate) fn parse_exif_offset(s: &str) -> Option<i32> {
/// Only develop and export should call it; culling and the grid must not
/// (FR-CULL-1).
pub fn decode(bytes: &[u8]) -> Result<RawImage, DecodeError> {
error::guarded("decode", || decode_unguarded(bytes))
}
fn decode_unguarded(bytes: &[u8]) -> Result<RawImage, DecodeError> {
use rawler::rawsource::RawSource;
let source = RawSource::new_from_slice(bytes);
@@ -551,6 +601,21 @@ pub fn decode(bytes: &[u8]) -> Result<RawImage, DecodeError> {
image.camera.clean_model.as_str(),
);
// TRACES: FR-MRG-3
// A linear DNG — three samples per pixel, no colour filter array — is a
// composite this application wrote (or any other demosaiced DNG). It
// carries the same scale, matrices and neutral as a CFA file and goes
// through the same profile; only the demosaic is skipped.
let samples_per_pixel = match image.cpp {
1 => 1u8,
3 => 3,
other => {
return Err(DecodeError::Unsupported(format!(
"{other} samples per pixel; only CFA (1) and linear RGB (3) are handled"
)))
}
};
let data = match image.data {
rawler::RawImageData::Integer(v) => v,
rawler::RawImageData::Float(v) => {
@@ -619,6 +684,10 @@ pub fn decode(bytes: &[u8]) -> Result<RawImage, DecodeError> {
wb_coeffs,
color_matrix,
base_curve,
samples_per_pixel,
profile,
make: image.camera.clean_make.clone(),
model: image.camera.clean_model.clone(),
})
}
+82 -13
View File
@@ -183,22 +183,65 @@ struct Entry {
value: u32,
}
/// The bytes a [`TiffReader`] reads: the head of a file, and optionally a
/// second range from further in, at a known offset.
///
/// A camera writes its IFDs at the front, so the first 256 KB of a file
/// is the whole structure. A file written strip by strip — the linear
/// DNG a merge produces — has its first IFD at the *end*, after the
/// pixels, and a reader that only has the head sees a pointer into
/// nothing. Rather than fetch 800 MB to read a date, the caller fetches
/// the head, asks [`crate::trailing_ifd`] where the IFD is, fetches that
/// tail, and reads through both. Offsets are the file's own throughout;
/// a read that falls in neither range is simply absent.
#[derive(Clone, Copy)]
struct Src<'a> {
head: &'a [u8],
tail: &'a [u8],
/// Where `tail` starts in the file.
tail_at: usize,
}
impl<'a> Src<'a> {
fn whole(data: &'a [u8]) -> Self {
Src {
head: data,
tail: &[],
tail_at: 0,
}
}
fn get(&self, start: usize, len: usize) -> Option<&'a [u8]> {
let end = start.checked_add(len)?;
if let Some(b) = self.head.get(start..end) {
return Some(b);
}
let s = start.checked_sub(self.tail_at)?;
self.tail.get(s..s.checked_add(len)?)
}
}
/// A minimal TIFF structure reader.
///
/// Deliberately not a general TIFF parser: it reads the IFD chain and entry
/// values and nothing else, because that is all locating a preview needs.
struct TiffReader<'a> {
data: &'a [u8],
data: Src<'a>,
little_endian: bool,
first_ifd: u32,
}
impl<'a> TiffReader<'a> {
fn new(data: &'a [u8]) -> Option<Self> {
if data.len() < 8 {
Self::over(Src::whole(data))
}
fn over(data: Src<'a>) -> Option<Self> {
let head = data.head;
if head.len() < 8 {
return None;
}
let little_endian = match &data[0..2] {
let little_endian = match &head[0..2] {
b"II" => true,
b"MM" => false,
_ => return None,
@@ -362,9 +405,7 @@ impl<'a> TiffReader<'a> {
};
raw[..len.min(4)].to_vec()
} else {
self.data
.get(e.value as usize..e.value as usize + len)?
.to_vec()
self.data.get(e.value as usize, len)?.to_vec()
};
let s = String::from_utf8_lossy(&bytes);
@@ -396,8 +437,7 @@ impl<'a> TiffReader<'a> {
// same way to recover the original byte order.
return None;
}
let start = e.value as usize;
self.data.get(start..start.checked_add(len)?)
self.data.get(e.value as usize, len)
}
fn offsets(&self, e: &Entry) -> Vec<u32> {
@@ -420,8 +460,8 @@ impl<'a> TiffReader<'a> {
}
}
fn read_u16(data: &[u8], at: usize, le: bool) -> Option<u16> {
let b = data.get(at..at + 2)?;
fn read_u16(data: Src<'_>, at: usize, le: bool) -> Option<u16> {
let b = data.get(at, 2)?;
Some(if le {
u16::from_le_bytes([b[0], b[1]])
} else {
@@ -429,8 +469,8 @@ fn read_u16(data: &[u8], at: usize, le: bool) -> Option<u16> {
})
}
fn read_u32(data: &[u8], at: usize, le: bool) -> Option<u32> {
let b = data.get(at..at + 4)?;
fn read_u32(data: Src<'_>, at: usize, le: bool) -> Option<u32> {
let b = data.get(at, 4)?;
Some(if le {
u32::from_le_bytes([b[0], b[1], b[2], b[3]])
} else {
@@ -460,7 +500,36 @@ pub fn jpeg_metadata(bytes: &[u8]) -> Result<crate::Metadata, crate::DecodeError
/// no `DateTimeOriginal` for some DNGs whose tag sits plainly at byte 826 —
/// and without this fallback those images are silently undated.
pub fn tiff_metadata(tiff_data: &[u8]) -> Result<crate::Metadata, crate::DecodeError> {
let reader = TiffReader::new(tiff_data)
tiff_metadata_over(Src::whole(tiff_data))
}
/// Where a TIFF-shaped file's first IFD is, when the head does not reach
/// it: the offset to fetch from, so [`tiff_metadata_split`] can read it.
/// `None` for a file that is not TIFF, or whose IFD the head already holds.
pub fn trailing_ifd(head: &[u8]) -> Option<u64> {
let r = TiffReader::new(head)?;
let at = r.first_ifd as u64;
(at >= head.len() as u64).then_some(at)
}
/// [`tiff_metadata`] over a head and a tail fetched separately: the head
/// from offset 0, the tail from `tail_at`. For the file whose IFDs follow
/// its pixels.
pub fn tiff_metadata_split(
head: &[u8],
tail: &[u8],
tail_at: u64,
) -> Result<crate::Metadata, crate::DecodeError> {
tiff_metadata_over(Src {
head,
tail,
tail_at: usize::try_from(tail_at)
.map_err(|_| crate::DecodeError::Metadata("tail offset out of range".into()))?,
})
}
fn tiff_metadata_over(src: Src<'_>) -> Result<crate::Metadata, crate::DecodeError> {
let reader = TiffReader::over(src)
.ok_or_else(|| crate::DecodeError::Metadata("malformed EXIF header".into()))?;
let mut md = crate::Metadata::default();
+4
View File
@@ -158,6 +158,10 @@ pub enum PreviewSize {
/// Returns [`DecodeError::NoPreview`] where there is none at all: a
/// fall-through signal, not a failure (see [`DecodeError::has_fallback`]).
pub fn extract_preview(bytes: &[u8], size: PreviewSize) -> Result<Preview, DecodeError> {
crate::error::guarded("preview", || extract_preview_unguarded(bytes, size))
}
fn extract_preview_unguarded(bytes: &[u8], size: PreviewSize) -> Result<Preview, DecodeError> {
use rawler::rawsource::RawSource;
// A plain JPEG *is* its own preview — rawler has no decoder for one, and
+48
View File
@@ -392,6 +392,21 @@ impl CameraProfile {
}
/// The calibrations this profile was built from, coolest first.
/// TRACES: FR-MRG-3
/// The calibrations as a DNG carries them: `(CalibrationIlluminant,
/// ColorMatrix)` with the EXIF light-source code, for a composite to
/// write the profile of the body that took its sources.
///
/// The code is recovered from the temperature, which is lossy only for
/// illuminants this profile never kept: `extract` drops calibrations
/// whose illuminant has no temperature, so every one here maps back.
pub fn dng_calibrations(&self) -> Vec<(u16, [[f32; 3]; 3])> {
self.calibrations
.iter()
.map(|c| (illuminant_code(c.temperature), c.xyz_to_cam))
.collect()
}
pub fn calibrations(&self) -> &[Calibration] {
&self.calibrations
}
@@ -528,6 +543,39 @@ fn illuminant_temperature(illuminant: Illuminant) -> Option<f32> {
})
}
/// The EXIF `LightSource` code for a calibration temperature — the inverse
/// of [`illuminant_temperature`], on the temperatures it produces.
fn illuminant_code(temperature: f32) -> u16 {
// Nearest of the table, so a temperature that came through a float
// round-trip still lands on its illuminant. Where two illuminants share
// a temperature (D55 and Daylight, D65 and Cloudy, D75 and Shade) the
// CIE standard one is written: it is what every profile database means.
const TABLE: &[(f32, u16)] = &[
(2856.0, 17), // A
(3200.0, 24), // ISO studio tungsten
(3500.0, 15), // white fluorescent
(4150.0, 14), // cool white fluorescent
(4230.0, 2), // fluorescent
(4874.0, 18), // B
(5000.0, 13), // daylight white fluorescent
(5003.0, 23), // D50
(5503.0, 20), // D55
(6430.0, 12), // daylight fluorescent
(6504.0, 21), // D65
(6774.0, 19), // C
(7504.0, 22), // D75
];
TABLE
.iter()
.min_by(|a, b| {
(a.0 - temperature)
.abs()
.total_cmp(&(b.0 - temperature).abs())
})
.map(|(_, code)| *code)
.unwrap_or(255)
}
/// Compose a forward matrix into camera RGB → linear sRGB.
///
/// `forward` takes white-balanced camera RGB to XYZ under D50, which is the
+3
View File
@@ -38,4 +38,7 @@ dr-gpu.workspace = true
dr-pipeline.workspace = true
env_logger.workspace = true
pollster.workspace = true
# The DNG writer's test reads its output back through the decoder the
# library uses, which is the whole claim the writer makes (S15.1).
rawler.workspace = true
zune-jpeg.workspace = true
+351
View File
@@ -0,0 +1,351 @@
//! TRACES: FR-MRG-3
//! A linear DNG: the container a merge writes its composite into.
//!
//! Decided by S15.1 (2026-09-19): rawler reads back a `LinearRaw` DNG the
//! application writes, so a composite re-enters the library as
//! `Format::Dng` through the decoder every camera DNG uses. What is written
//! is a RAW in every sense a warp can preserve — camera-linear `u16`
//! samples at the first source's own scale, its matrices, illuminants,
//! as-shot neutral and body name — so the panorama is developed afterwards
//! as one photograph, from the sensor's numbers.
//!
//! # Streamed, not buffered
//!
//! The composite is larger than any single photograph the pipeline renders
//! and larger than the tablet's memory (FR-MRG-11), so the writer never
//! holds it. Strips are pulled from the caller one at a time through a
//! closure, in order, and written as they arrive; the caller renders a band
//! of chunks, hands over its rows, and moves on.
//!
//! # Why the `tiff` crate after all
//!
//! S15.1's spike hand-rolled its IFD because the crate's encoder fixes
//! `PhotometricInterpretation` to RGB when the image is opened. It does — but
//! a directory is a map and a later `write_tag` on the same tag replaces the
//! earlier, so `LinearRaw` goes in over the top and everything else the
//! crate does (strips, offsets, sub-IFDs, the EXIF block `encode.rs` already
//! knows how to write) is kept.
use std::io::{Seek, Write};
use tiff::encoder::{colortype, DirectoryEncoder, SRational, TiffEncoder, TiffKind, TiffValue};
use tiff::tags::Tag;
use crate::encode::{sub_directories, tag_metadata, Ascii, Rationals};
use crate::{ExportError, SourceMetadata};
/// What the DNG says about the camera that "took" the composite: the first
/// source's profile, carried across so the composite develops through it.
#[derive(Debug, Clone, PartialEq)]
pub struct DngProfile {
/// `UniqueCameraModel`, the name the profile database matches on.
pub unique_model: String,
/// `(CalibrationIlluminant, ColorMatrix)`: the EXIF light-source code and
/// the XYZ → camera matrix measured under it. One or two.
pub calibrations: Vec<(u16, [[f32; 3]; 3])>,
/// `AsShotNeutral`, camera RGB of the scene's white.
pub as_shot_neutral: [f32; 3],
/// `WhiteLevel`: the sample value that is clipping. The first source's
/// white minus its black, since the samples are black-subtracted.
pub white_level: u32,
}
/// Write a linear DNG, pulling `rows_per_strip`-row strips from `strips`.
///
/// Each call to `strips` receives the strip index and a buffer to fill with
/// `width × rows × 3` interleaved RGB `u16` samples (the last strip may be
/// shorter). `source` supplies the `Make`, `Model`, dates and EXIF block
/// exactly as an export does (FR-EXP-8 sanitising already applied by the
/// caller).
///
/// `PhotometricInterpretation = LinearRaw`, `DNGVersion 1.4`, uncompressed,
/// `Orientation = 1` — the composite is written upright (panorama.md §8).
///
/// `crop` is asked once every strip is in, and its answer — the largest
/// rectangle the frames covered, found while the strips went by
/// (`Inscribed`) — becomes `DefaultCropOrigin`/`DefaultCropSize`
/// (FR-MRG-4): the file opens on the picture, and the border is still in it.
// Eight arguments, and each is a different thing: the sink, three
// dimensions, the profile, the header, the strip source and the crop. A
// struct for them would be a struct with one caller.
#[allow(clippy::too_many_arguments)]
pub fn write_linear_dng<W, F, C>(
out: W,
width: u32,
height: u32,
rows_per_strip: u32,
profile: &DngProfile,
source: Option<&SourceMetadata>,
mut strips: F,
crop: C,
) -> Result<(), ExportError>
where
W: Write + Seek,
F: FnMut(usize, &mut Vec<u16>) -> Result<(), ExportError>,
C: FnOnce() -> Option<crate::Rect>,
{
let enc = |e: tiff::TiffError| ExportError::Encode(e.to_string());
let mut encoder = TiffEncoder::new(out).map_err(enc)?;
let sub = sub_directories(&mut encoder, source, width, height)?;
let mut image = encoder
.new_image::<colortype::RGB16>(width, height)
.map_err(enc)?;
image.rows_per_strip(rows_per_strip.max(1)).map_err(enc)?;
tag_metadata(image.encoder(), source, &sub)?;
tag_dng(image.encoder(), profile).map_err(enc)?;
let rows = rows_per_strip.max(1);
let strip_count = height.div_ceil(rows) as usize;
let mut buf: Vec<u16> = Vec::with_capacity((width * rows * 3) as usize);
for k in 0..strip_count {
buf.clear();
strips(k, &mut buf)?;
let expected_rows = rows.min(height - k as u32 * rows);
let expected = (width * expected_rows * 3) as usize;
if buf.len() != expected {
return Err(ExportError::Encode(format!(
"strip {k} has {} samples, expected {expected}",
buf.len()
)));
}
image.write_strip(&buf).map_err(enc)?;
}
if let Some(r) = crop().filter(|r| r.width > 0 && r.height > 0) {
let r = crate::Rect {
x: r.x.min(width - 1),
y: r.y.min(height - 1),
width: r.width.min(width - r.x.min(width - 1)),
height: r.height.min(height - r.y.min(height - 1)),
};
image
.encoder()
.write_tag(Tag::Unknown(tag::DEFAULT_CROP_ORIGIN), &[r.x, r.y][..])
.map_err(enc)?;
image
.encoder()
.write_tag(
Tag::Unknown(tag::DEFAULT_CROP_SIZE),
&[r.width, r.height][..],
)
.map_err(enc)?;
}
image.finish().map_err(enc)
}
/// The tags that make a TIFF a DNG, and a linear one.
fn tag_dng<W, K>(dir: &mut DirectoryEncoder<'_, W, K>, profile: &DngProfile) -> tiff::TiffResult<()>
where
W: Write + Seek,
K: TiffKind,
{
// Over the top of what `new_image` wrote: this is the whole trick.
dir.write_tag(Tag::PhotometricInterpretation, LINEAR_RAW)?;
dir.write_tag(Tag::Orientation, 1u16)?;
dir.write_tag(Tag::Unknown(tag::DNG_VERSION), &[1u8, 4, 0, 0][..])?;
dir.write_tag(Tag::Unknown(tag::DNG_BACKWARD_VERSION), &[1u8, 4, 0, 0][..])?;
dir.write_tag(
Tag::Unknown(tag::UNIQUE_CAMERA_MODEL),
Ascii(&profile.unique_model),
)?;
dir.write_tag(
Tag::Unknown(tag::WHITE_LEVEL),
&[profile.white_level; 3][..],
)?;
dir.write_tag(Tag::Unknown(tag::BLACK_LEVEL), &[0u32; 3][..])?;
for (slot, (illuminant, matrix)) in profile.calibrations.iter().take(2).enumerate() {
let (ill_tag, mat_tag) = if slot == 0 {
(tag::CALIBRATION_ILLUMINANT_1, tag::COLOR_MATRIX_1)
} else {
(tag::CALIBRATION_ILLUMINANT_2, tag::COLOR_MATRIX_2)
};
dir.write_tag(Tag::Unknown(ill_tag), *illuminant)?;
let flat: Vec<SRational> = matrix
.iter()
.flatten()
.map(|&v| SRational {
n: (v * 10_000.0).round() as i32,
d: 10_000,
})
.collect();
dir.write_tag(Tag::Unknown(mat_tag), SRationals(&flat))?;
}
let neutral: Vec<(u32, u32)> = profile
.as_shot_neutral
.iter()
.map(|&v| ((v.max(0.0) * 1_000_000.0).round() as u32, 1_000_000))
.collect();
dir.write_tag(Tag::Unknown(tag::AS_SHOT_NEUTRAL), Rationals(&neutral))?;
Ok(())
}
/// `PhotometricInterpretation` for demosaiced, un-rendered sensor data.
const LINEAR_RAW: u16 = 34892;
/// DNG tag numbers the `tiff` crate has no names for.
mod tag {
pub const DNG_VERSION: u16 = 50706;
pub const DNG_BACKWARD_VERSION: u16 = 50707;
pub const UNIQUE_CAMERA_MODEL: u16 = 50708;
pub const BLACK_LEVEL: u16 = 50714;
pub const WHITE_LEVEL: u16 = 50717;
pub const DEFAULT_CROP_ORIGIN: u16 = 50719;
pub const DEFAULT_CROP_SIZE: u16 = 50720;
pub const COLOR_MATRIX_1: u16 = 50721;
pub const COLOR_MATRIX_2: u16 = 50722;
pub const AS_SHOT_NEUTRAL: u16 = 50728;
pub const CALIBRATION_ILLUMINANT_1: u16 = 50778;
pub const CALIBRATION_ILLUMINANT_2: u16 = 50779;
}
/// A run of `SRATIONAL`s, as `encode::Rationals` is for `RATIONAL`.
struct SRationals<'a>(&'a [SRational]);
impl TiffValue for SRationals<'_> {
const BYTE_LEN: u8 = 8;
const FIELD_TYPE: tiff::tags::Type = tiff::tags::Type::SRATIONAL;
fn count(&self) -> usize {
self.0.len()
}
fn data(&self) -> std::borrow::Cow<'_, [u8]> {
let mut out = Vec::with_capacity(self.0.len() * 8);
for r in self.0 {
out.extend_from_slice(&r.n.to_ne_bytes());
out.extend_from_slice(&r.d.to_ne_bytes());
}
std::borrow::Cow::Owned(out)
}
}
#[cfg(test)]
mod tests {
use super::*;
fn profile() -> DngProfile {
DngProfile {
unique_model: "Canon EOS 6D".into(),
calibrations: vec![
(17, [[0.8, -0.2, 0.1], [-0.3, 1.1, 0.2], [0.0, -0.1, 0.9]]),
(21, [[0.7, -0.1, 0.0], [-0.2, 1.0, 0.1], [0.0, -0.2, 0.8]]),
],
as_shot_neutral: [0.5, 1.0, 0.6],
white_level: 13_023,
}
}
fn write(width: u32, height: u32, rows: u32) -> Vec<u8> {
let mut bytes = std::io::Cursor::new(Vec::new());
let source = SourceMetadata {
make: Some("Canon".into()),
model: Some("Canon EOS 6D".into()),
captured_at: Some(1_754_398_664),
captured_offset: Some(120),
..Default::default()
};
write_linear_dng(
&mut bytes,
width,
height,
rows,
&profile(),
Some(&source),
|k, buf| {
let first = k as u32 * rows;
let n = rows.min(height - first);
for y in first..first + n {
for x in 0..width {
buf.extend([(x + y * width) as u16, 1000, 2000]);
}
}
Ok(())
},
|| {
Some(crate::Rect {
x: 2,
y: 1,
width: 15,
height: 10,
})
},
)
.expect("written");
bytes.into_inner()
}
#[test]
fn rawler_reads_it_back_as_linear_raw() {
let bytes = write(20, 13, 4);
let source = rawler::rawsource::RawSource::new_from_slice(&bytes);
let decoder = rawler::get_decoder(&source).expect("a DNG");
let image = decoder
.raw_image(&source, &Default::default(), false)
.expect("decodes");
assert_eq!((image.width, image.height, image.cpp), (20, 13, 3));
assert_eq!(image.whitelevel.0[0], 13_023);
// Pixel (3, 2) is (3 + 2·20, 1000, 2000) — samples in order, strips
// joined without a seam.
let rawler::RawImageData::Integer(data) = &image.data else {
panic!("integer samples")
};
let i = (2 * 20 + 3) * 3;
assert_eq!(&data[i..i + 3], &[43, 1000, 2000]);
// Last row, from the short final strip.
let i = (12 * 20 + 19) * 3;
assert_eq!(data[i], (19 + 12 * 20) as u16);
// The profile came through as the camera's.
assert!(!image.camera.color_matrix.is_empty());
assert_eq!(image.model, "Canon EOS 6D");
// The default crop is what the decoder reports as the picture.
let crop = image.crop_area.expect("a crop");
assert_eq!((crop.p.x, crop.p.y, crop.d.w, crop.d.h), (2, 1, 15, 10));
}
#[test]
fn the_catalog_reads_the_date_from_a_head_and_a_tail() {
// TRACES: FR-CAT-5
// The IFDs follow the pixels, so a scan that has the first bytes of
// the file has a pointer into nothing; rawler finds no decoder in
// that, and the composite would sit undated at the end of the grid.
// The scan's second range — from the first IFD to the end — with
// the head is enough to date it, and to name the camera.
let bytes = write(640, 400, 64);
let head = &bytes[..4096];
assert!(
dr_decode::metadata(head).is_err(),
"the head alone must not read"
);
let at = dr_decode::trailing_ifd(head).expect("the IFD is beyond the head");
assert!(at as usize > head.len());
let tail = &bytes[at as usize..];
assert!(tail.len() < 4096, "the tail is the IFD, not the pixels");
let md = dr_decode::metadata_split(head, tail, at).expect("read from two ranges");
assert_eq!(md.captured_at, Some(1_754_398_664));
assert_eq!(md.captured_offset, Some(120));
assert_eq!(md.model.as_deref(), Some("Canon EOS 6D"));
// A head that holds everything is not a trailing-IFD file.
assert_eq!(dr_decode::trailing_ifd(&bytes), None);
}
#[test]
fn a_strip_of_the_wrong_length_is_refused() {
let mut bytes = std::io::Cursor::new(Vec::new());
let err = write_linear_dng(
&mut bytes,
8,
8,
8,
&profile(),
None,
|_, buf| {
buf.extend([0u16; 10]);
Ok(())
},
|| None,
)
.unwrap_err();
assert!(matches!(err, ExportError::Encode(_)));
}
}
+5 -5
View File
@@ -213,7 +213,7 @@ impl tiff::encoder::TiffValue for Undefined<'_> {
/// specification says, `dr-decode` reads them back with `from_utf8_lossy`, and
/// a mangled accent is a far better outcome than a refusal. So the bytes go
/// through verbatim with the terminating NUL the type requires.
struct Ascii<'a>(&'a str);
pub(crate) struct Ascii<'a>(pub(crate) &'a str);
impl tiff::encoder::TiffValue for Ascii<'_> {
const BYTE_LEN: u8 = 1;
@@ -241,7 +241,7 @@ impl tiff::encoder::TiffValue for Ascii<'_> {
/// a value that forced little-endian would be read back byte-swapped on a
/// big-endian machine. `exif.rs` builds its own header and so chooses its own
/// order; here the container has already chosen.
struct Rationals<'a>(&'a [(u32, u32)]);
pub(crate) struct Rationals<'a>(pub(crate) &'a [(u32, u32)]);
impl tiff::encoder::TiffValue for Rationals<'_> {
const BYTE_LEN: u8 = 8;
@@ -317,7 +317,7 @@ where
/// and then no pointer is written either, so the file has no trace of the
/// directory rather than a pointer to an empty one.
#[derive(Default)]
struct SubDirectories {
pub(crate) struct SubDirectories {
exif: Option<u32>,
gps: Option<u32>,
}
@@ -335,7 +335,7 @@ struct SubDirectories {
/// A TIFF gets no separate EXIF *block* — no APP1, no `eXIf` chunk. Its own
/// directory is the EXIF structure, and adding a second copy inside it would
/// give a reader two answers to every question.
fn sub_directories<W>(
pub(crate) fn sub_directories<W>(
encoder: &mut tiff::encoder::TiffEncoder<W>,
source: Option<&SourceMetadata>,
width: u32,
@@ -465,7 +465,7 @@ where
///
/// No `Orientation`, for the reason `exif.rs` gives at length: the pixels
/// arriving here are already upright.
fn tag_metadata<W, K>(
pub(crate) fn tag_metadata<W, K>(
dir: &mut tiff::encoder::DirectoryEncoder<'_, W, K>,
source: Option<&SourceMetadata>,
sub: &SubDirectories,
+156
View File
@@ -0,0 +1,156 @@
//! TRACES: FR-MRG-4
//! The largest rectangle inside a coverage mask, found a row at a time.
//!
//! A merged panorama has ragged edges: the frames' footprints under a
//! cylinder or a sphere are not rectangles, and the composite carries a
//! black border where none of them reached. FR-MRG-4 asks for an auto-crop
//! to the largest inscribed rectangle. This finds it as the bands are
//! produced, so the composite is never held to be measured (FR-MRG-11):
//! each row extends a running histogram of consecutive covered rows above
//! it, and the largest rectangle ending on that row is the largest
//! rectangle under the histogram — a stack pass, linear in the width.
//!
//! The crop is written as the DNG's `DefaultCropOrigin`/`DefaultCropSize`,
//! which every reader honours and which discards nothing: the pixels
//! outside it are still in the file for a photographer who wants them.
/// The rectangle so far, in pixels from the top left.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Default)]
pub struct Rect {
pub x: u32,
pub y: u32,
pub width: u32,
pub height: u32,
}
impl Rect {
pub fn area(&self) -> u64 {
u64::from(self.width) * u64::from(self.height)
}
}
/// Feed rows top to bottom; ask for the best at any point.
#[derive(Debug, Clone)]
pub struct Inscribed {
width: usize,
/// How many consecutive covered rows end at the last row fed, per column.
heights: Vec<u32>,
rows: u32,
best: Rect,
}
impl Inscribed {
pub fn new(width: u32) -> Self {
Inscribed {
width: width as usize,
heights: vec![0; width as usize],
rows: 0,
best: Rect::default(),
}
}
/// One more row of coverage, `width` long.
pub fn push_row(&mut self, covered: &[bool]) {
debug_assert_eq!(covered.len(), self.width);
for (h, &c) in self.heights.iter_mut().zip(covered) {
*h = if c { *h + 1 } else { 0 };
}
self.rows += 1;
// Largest rectangle under the histogram, with a sentinel column of
// height 0 at the end so every bar is popped.
let mut stack: Vec<usize> = Vec::new();
for i in 0..=self.width {
let h = if i < self.width { self.heights[i] } else { 0 };
while let Some(&top) = stack.last() {
if self.heights[top] <= h {
break;
}
stack.pop();
let height = self.heights[top];
let left = stack.last().map_or(0, |&l| l + 1);
let width = (i - left) as u32;
let area = u64::from(width) * u64::from(height);
if area > self.best.area() {
self.best = Rect {
x: left as u32,
y: self.rows - height,
width,
height,
};
}
}
stack.push(i);
}
}
/// Several rows at once, as a band hands them over.
pub fn push_rows(&mut self, covered: &[bool], rows: u32) {
for r in 0..rows as usize {
self.push_row(&covered[r * self.width..(r + 1) * self.width]);
}
}
pub fn best(&self) -> Rect {
self.best
}
}
#[cfg(test)]
mod tests {
use super::*;
fn from_art(art: &[&str]) -> Rect {
let mut ins = Inscribed::new(art[0].len() as u32);
for row in art {
let covered: Vec<bool> = row.chars().map(|c| c == '#').collect();
ins.push_row(&covered);
}
ins.best()
}
#[test]
fn a_full_mask_is_its_own_rectangle() {
let r = from_art(&["####", "####", "####"]);
assert_eq!(
r,
Rect {
x: 0,
y: 0,
width: 4,
height: 3
}
);
}
#[test]
fn ragged_edges_are_cut_off() {
// A cylinder's footprint: narrower at top and bottom.
let r = from_art(&[
"..####..", ".######.", "########", "########", ".######.", "..####..",
]);
// 6 wide × 4 tall = 24 beats 8 × 2 = 16 and 4 × 6 = 24 ties; the
// first found wins a tie, which is the wider one here.
assert_eq!(r.area(), 24);
assert!(r.width == 6 && r.height == 4 || r.width == 4 && r.height == 6);
}
#[test]
fn a_hole_is_avoided() {
let r = from_art(&["#####", "##.##", "#####", "#####"]);
// Left of the hole: 2 × 4 = 8; right: 2 × 4 = 8; below: 5 × 2 = 10.
assert_eq!(
r,
Rect {
x: 0,
y: 2,
width: 5,
height: 2
}
);
}
#[test]
fn nothing_covered_is_nothing() {
assert_eq!(from_art(&["....", "...."]).area(), 0);
}
}
+4
View File
@@ -24,16 +24,20 @@
use dr_types::{ColourSpace, ExportFormat, ExportSettings};
mod dng;
mod encode;
mod error;
mod exif;
pub mod icc;
mod inscribed;
mod metadata;
mod name;
mod sharpen;
mod size;
pub use dng::{write_linear_dng, DngProfile};
pub use error::ExportError;
pub use inscribed::{Inscribed, Rect};
pub use metadata::SourceMetadata;
pub use name::{resolve_name, NameContext};
pub use size::target_size;
+11 -6
View File
@@ -9,10 +9,11 @@ license.workspace = true
thiserror.workspace = true
log.workspace = true
# Inference. `ort` is the API; **tract is the engine** — see the workspace
# manifest, and docs/faces.md §3, for why the C++ ONNX Runtime is not linked.
# Inference. `ort` is the API; **what runs it is `dr-inference-engine`'s
# business** — tract, or an ONNX Runtime the app found on disk, on whichever
# provider the device has (docs/dev/inference.md). This crate never names either.
ort = { workspace = true, optional = true }
ort-tract = { workspace = true, optional = true }
dr-inference-engine = { workspace = true, optional = true }
ndarray = { workspace = true, optional = true }
[dev-dependencies]
@@ -20,7 +21,7 @@ zune-jpeg.workspace = true
env_logger.workspace = true
# The M1 probe drives `ort` directly so it can print the raw load error.
ort = { workspace = true }
ort-tract = { workspace = true }
dr-inference-engine = { workspace = true }
[[example]]
name = "probe"
@@ -30,9 +31,13 @@ required-features = ["inference"]
name = "faces"
required-features = ["inference"]
[[example]]
name = "eyes"
required-features = ["inference"]
[features]
# Nothing on by default, and in particular **no `embedded-model`**: the weights
# are not a build input and never become one (docs/faces.md §2.2). A feature
# are not a build input and never become one (docs/dev/faces.md §2.2). A feature
# flag that *could* embed them is a flag someone eventually sets in a packaging
# script, and the InsightFace grant does not survive that.
default = []
@@ -44,4 +49,4 @@ default = []
# must be testable against synthetic embeddings on a machine with no weights on
# it — a test suite that needs a research-licensed download is a test suite
# that does not run in CI.
inference = ["dep:ort", "dep:ort-tract", "dep:ndarray"]
inference = ["dep:ort", "dep:dr-inference-engine", "dep:ndarray"]
+145
View File
@@ -0,0 +1,145 @@
//! Detect the faces in a JPEG and read each one's eyes (docs/dev/faces.md §17).
//!
//! The thing worth looking at is whether the eye boxes land on eyes and
//! whether soft ones are refused — so with `--dump DIR` the crops the
//! classifiers were shown are written out as PPMs, one per eye and one per
//! head framing, named by image and face, and every line carries the
//! numbers the readability floors are set from.
//!
//! cargo run -p dr-face --features inference --example eyes -- \
//! DET.onnx 2D106DET.onnx OCEC.onnx SGC.onnx [--dump DIR] photo.jpg [photo.jpg ...]
//!
//! All four models must have had their dynamic dims pinned first; see
//! `tools/fix-face-model-shapes.sh`.
use std::path::{Path, PathBuf};
use std::time::Instant;
use dr_face::{align, DetectOptions, Detector, EyeModels, Pixels};
fn main() {
env_logger::init();
let mut args: Vec<String> = std::env::args().skip(1).collect();
let dump = args.iter().position(|a| a == "--dump").map(|i| {
args.remove(i);
PathBuf::from(args.remove(i))
});
if args.len() < 5 {
eprintln!(
"usage: eyes DET.onnx 2D106DET.onnx OCEC.onnx SGC.onnx [--dump DIR] IMAGE.jpg [IMAGE.jpg ...]"
);
std::process::exit(2);
}
if let Some(d) = &dump {
std::fs::create_dir_all(d).expect("dump dir");
}
let t = Instant::now();
let mut detector = Detector::from_path(&args[0]).expect("load detector");
let mut models = EyeModels::from_paths(&args[1], &args[2], &args[3]).expect("load eye models");
println!("loaded the models in {:?}", t.elapsed());
let opts = DetectOptions::default();
for path in &args[4..] {
let (rgb, w, h) = match load_jpeg(path) {
Ok(v) => v,
Err(e) => {
println!("{path}: {e}");
continue;
}
};
let dets = detector.detect(&rgb, w, h, &opts).expect("detect");
println!("\n{path} ({w}×{h}) {} face(s)", dets.len());
let stem = Path::new(path)
.file_stem()
.map(|s| s.to_string_lossy().into_owned())
.unwrap_or_default();
for (i, d) in dets.iter().enumerate() {
let px = Pixels::RgbF32(&rgb);
let t = Instant::now();
let reading = models
.read(px, w, h, d.bbox, &d.landmarks)
.expect("read eyes");
let ms = t.elapsed().as_secs_f64() * 1e3;
let Some((r, lm)) = reading else {
println!(" [{i}] nothing to cut, skipped");
continue;
};
println!(
" [{i}] conf {:.2} box {:.0}×{:.0} right {:.3} ({:.0}px, sharp {:.3}) left {:.3} ({:.0}px, sharp {:.3}) sunglasses {:.3} → {:?} ({ms:.1} ms)",
d.confidence,
d.width(),
d.height(),
r.right.open,
r.right.px,
r.right.sharpness,
r.left.open,
r.left.px,
r.left.sharpness,
r.sunglasses,
r.state(),
);
if let Some(dir) = &dump {
// The same crops `EyeModels::read` cut, cut again for the
// sheet from the landmarks it handed back: the reading itself
// carries numbers, not pixels.
for (name, contour) in [("right", lm.right_eye()), ("left", lm.left_eye())] {
if let Some(patch) =
align::eye_box(&contour).and_then(|b| align::eye_patch(px, w, h, b))
{
write_ppm(
&dir.join(format!("{stem}-{i}-{name}.ppm")),
patch.pixels(),
align::EYE_PATCH_WIDTH,
align::EYE_PATCH_HEIGHT,
);
}
}
if let Some(head) = align::head_views(px, w, h, &d.landmarks) {
for (n, view) in head.views().enumerate() {
write_ppm(
&dir.join(format!("{stem}-{i}-head{n}.ppm")),
view,
align::SUNGLASSES_EDGE,
align::SUNGLASSES_EDGE,
);
}
}
}
}
}
}
fn write_ppm(path: &Path, rgb: &[f32], w: usize, h: usize) {
let mut out = format!("P6\n{w} {h}\n255\n").into_bytes();
out.extend(
rgb.iter()
.map(|v| (v.clamp(0.0, 1.0) * 255.0).round() as u8),
);
std::fs::write(path, out).expect("write ppm");
}
/// Decode to the tightly packed `f32` RGB `0.0..=1.0` the crate expects.
fn load_jpeg(path: &str) -> Result<(Vec<f32>, usize, usize), String> {
let bytes = std::fs::read(path).map_err(|e| e.to_string())?;
let mut dec = zune_jpeg::JpegDecoder::new(&bytes);
let px = dec.decode().map_err(|e| e.to_string())?;
let info = dec.info().ok_or("no jpeg header")?;
let (w, h) = (info.width as usize, info.height as usize);
let rgb: Vec<f32> = match px.len() / (w * h) {
3 => px.iter().map(|&v| v as f32 / 255.0).collect(),
1 => px
.iter()
.flat_map(|&v| {
let g = v as f32 / 255.0;
[g, g, g]
})
.collect(),
n => return Err(format!("{n} components per pixel, expected 1 or 3")),
};
Ok((rgb, w, h))
}
+1 -1
View File
@@ -7,7 +7,7 @@
//! DET.onnx EMB.onnx photo.jpg [photo.jpg ...]
//!
//! The models must have had their input dims frozen first; see
//! `tools/fix-face-model-shapes.sh` and docs/faces.md §12 M1.
//! `tools/fix-face-model-shapes.sh` and docs/dev/faces.md §12 M1.
use std::time::Instant;
+1 -1
View File
@@ -1,4 +1,4 @@
//! M1 (docs/faces.md §12) — will tract load these graphs at all?
//! M1 (docs/dev/faces.md §12) — will tract load these graphs at all?
//!
//! The one measurement everything else in the face subsystem is conditional
//! on. `det_500m.onnx` has a dynamic H/W input, which is exactly what tract
+1 -1
View File
@@ -9,7 +9,7 @@
//!
//! # What it is for
//!
//! docs/faces.md §9 has the desktop numbers and the question they leave open:
//! docs/dev/faces.md §9 has the desktop numbers and the question they leave open:
//! a GPU GEMM is worth roughly 1.5× of a regroup on a twenty-core desktop,
//! because the scan is under a third of the pass there. On a tablet the CPU is
//! several times slower and the GPU is not, so the same optimisation is worth
+469 -58
View File
@@ -1,4 +1,4 @@
//! Five-point face alignment (docs/faces.md §5).
//! Five-point face alignment (docs/dev/faces.md §5).
//!
//! ArcFace embeddings are trained on faces warped to a canonical 112×112
//! arrangement. Feeding the model a plain bounding-box crop *works* — it
@@ -117,51 +117,56 @@ impl Aligned112 {
/// `face_index --quality` prints the joint distribution so the two are
/// chosen together rather than each in ignorance of the other.
pub fn sharpness(&self) -> f32 {
let e = ALIGNED_EDGE;
let luma: Vec<f32> = self
.pixels
.chunks_exact(3)
.map(|p| 0.2126 * p[0] + 0.7152 * p[1] + 0.0722 * p[2])
.collect();
let (mut lap_sum, mut lap_sq) = (0.0_f64, 0.0_f64);
let (mut lum_sum, mut lum_sq) = (0.0_f64, 0.0_f64);
let mut n = 0.0_f64;
for y in 1..e - 1 {
for x in 1..e - 1 {
let i = y * e + x;
// Four-neighbour Laplacian. The 8-neighbour form is more
// sensitive to diagonal detail and also to noise, which on a
// high-ISO frame is exactly the thing that must not read as
// sharpness.
let lap = 4.0 * luma[i] - luma[i - 1] - luma[i + 1] - luma[i - e] - luma[i + e];
let lap = lap as f64;
lap_sum += lap;
lap_sq += lap * lap;
let l = luma[i] as f64;
lum_sum += l;
lum_sq += l * l;
n += 1.0;
}
}
if n == 0.0 {
return 0.0;
}
let lap_var = (lap_sq / n - (lap_sum / n).powi(2)).max(0.0);
let lum_var = (lum_sq / n - (lum_sum / n).powi(2)).max(0.0);
// A crop with no luma variation has no edges to find either, so the
// ratio is 0/0. Zero is the right answer: nothing there is a face.
if lum_var <= 1e-9 {
return 0.0;
}
(lap_var / lum_var) as f32
laplacian_ratio(&self.pixels, ALIGNED_EDGE, ALIGNED_EDGE)
}
}
/// Variance of the four-neighbour Laplacian over the variance of the luma,
/// for a `w × h` RGB crop — the measure [`Aligned112::sharpness`] describes,
/// shared with [`EyePatch::sharpness`].
fn laplacian_ratio(pixels: &[f32], w: usize, h: usize) -> f32 {
let luma: Vec<f32> = pixels
.chunks_exact(3)
.map(|p| 0.2126 * p[0] + 0.7152 * p[1] + 0.0722 * p[2])
.collect();
let (mut lap_sum, mut lap_sq) = (0.0_f64, 0.0_f64);
let (mut lum_sum, mut lum_sq) = (0.0_f64, 0.0_f64);
let mut n = 0.0_f64;
for y in 1..h.saturating_sub(1) {
for x in 1..w.saturating_sub(1) {
let i = y * w + x;
// Four-neighbour Laplacian. The 8-neighbour form is more
// sensitive to diagonal detail and also to noise, which on a
// high-ISO frame is exactly the thing that must not read as
// sharpness.
let lap = 4.0 * luma[i] - luma[i - 1] - luma[i + 1] - luma[i - w] - luma[i + w];
let lap = lap as f64;
lap_sum += lap;
lap_sq += lap * lap;
let l = luma[i] as f64;
lum_sum += l;
lum_sq += l * l;
n += 1.0;
}
}
if n == 0.0 {
return 0.0;
}
let lap_var = (lap_sq / n - (lap_sum / n).powi(2)).max(0.0);
let lum_var = (lum_sq / n - (lum_sum / n).powi(2)).max(0.0);
// A crop with no luma variation has no edges to find either, so the
// ratio is 0/0. Zero is the right answer: nothing there is a face.
if lum_var <= 1e-9 {
return 0.0;
}
(lap_var / lum_var) as f32
}
/// A similarity transform: rotation, uniform scale, translation.
///
/// Stored as the four independent parameters rather than a 2×3 matrix so that
@@ -203,7 +208,7 @@ impl Similarity {
///
/// # Why least squares and not RANSAC
///
/// The reference C++ implementation (docs/faces.md §1.1) fits this with
/// The reference C++ implementation (docs/dev/faces.md §1.1) fits this with
/// OpenCV's `estimateAffinePartial2D` under RANSAC. RANSAC over five points is
/// a strange fit: the minimal sample for a similarity is two, so it can discard
/// landmarks it judges outliers and solve from a subset — and on a profile face
@@ -351,27 +356,314 @@ pub fn warp_pixels(
let m = fit_similarity(landmarks, &ARCFACE_TEMPLATE)?;
let e = ALIGNED_EDGE;
let mut pixels = vec![0.0_f32; e * e * 3];
for v in 0..e {
for u in 0..e {
// Pixel centres, so the transform is not off by half a pixel —
// which is small enough to survive review and large enough to
// matter on a 40-pixel face.
let (x, y) = m.invert(u as f32 + 0.5, v as f32 + 0.5);
let (x, y) = (x - 0.5, y - 0.5);
let out = (v * e + u) * 3;
sample_bilinear(px, width, height, x, y, &mut pixels[out..out + 3]);
}
}
let window = TemplateWindow {
x: 0.0,
y: 0.0,
w: e as f32,
h: e as f32,
};
Some(Aligned112 {
pixels,
pixels: sample_window(px, width, height, &m, &window, e, e),
// The warp maps `scale` source pixels to one destination pixel, so the
// crop spans 112/scale of the source.
source_px: ALIGNED_EDGE as f32 / m.scale(),
})
}
/// A rectangle in **template** coordinates — the 112-unit frame
/// [`ARCFACE_TEMPLATE`] is written in — that a crop is sampled from.
///
/// Every crop this module makes is one of these resampled through the same
/// fitted similarity: the aligned face is the window `(0, 0, 112, 112)`, an
/// eye is a small window around its template point, a head is a window larger
/// than the face. Stating them all in one frame is what lets a second crop be
/// added as a constant rather than a second warp, and what keeps them
/// consistent with each other — the eye window sits where the eye landmark
/// lands *after* alignment, so a tilted face gets an upright eye.
#[derive(Debug, Clone, Copy, PartialEq)]
struct TemplateWindow {
x: f32,
y: f32,
w: f32,
h: f32,
}
/// Resample `window` of the template frame into an `out_w × out_h` RGB buffer.
///
/// Bilinear, from the source, in one step — the property [`warp`] insists on,
/// and every crop through here inherits it. The output pixel `(u, v)` is placed
/// at its centre in the window, taken back through `m` to source coordinates,
/// and sampled there; the window's aspect is **not** preserved when it differs
/// from the output's, which is deliberate for the eye classifier (it was
/// trained on detector boxes resized the same way) and moot for the others.
fn sample_window(
px: Pixels<'_>,
width: usize,
height: usize,
m: &Similarity,
window: &TemplateWindow,
out_w: usize,
out_h: usize,
) -> Vec<f32> {
let mut pixels = vec![0.0_f32; out_w * out_h * 3];
let sx = window.w / out_w as f32;
let sy = window.h / out_h as f32;
for v in 0..out_h {
for u in 0..out_w {
// Pixel centres, so the transform is not off by half a pixel —
// which is small enough to survive review and large enough to
// matter on a 40-pixel face.
let tx = window.x + (u as f32 + 0.5) * sx;
let ty = window.y + (v as f32 + 0.5) * sy;
let (x, y) = m.invert(tx, ty);
let (x, y) = (x - 0.5, y - 0.5);
let out = (v * out_w + u) * 3;
sample_bilinear(px, width, height, x, y, &mut pixels[out..out + 3]);
}
}
pixels
}
// ── eyes ──────────────────────────────────────────────────────────────────
/// Width of an eye crop as the classifier reads it, in pixels. Fixed by the
/// OCEC input (`docs/dev/faces.md` §17): 40 wide, 24 high.
pub const EYE_PATCH_WIDTH: usize = 40;
/// Height of an eye crop as the classifier reads it, in pixels.
pub const EYE_PATCH_HEIGHT: usize = 24;
/// How much an eye's box is grown beyond its lid contour, as a fraction of
/// its width and height on each side.
///
/// The classifier was trained on a whole-body detector's *eye* boxes — tight
/// round the palpebral fissure — and measured on 25 open-eyed faces from the
/// reference library, a tight box is what it wants: 22 of 25 read open at
/// 0 and 0.1, 18 at 0.4, 14 at 0.6 (docs/dev/faces.md §17.2). A tenth, so a
/// contour landing a pixel short of the lashes still holds them.
pub const EYE_BOX_MARGIN: f32 = 0.1;
/// Height a shut eye's box is given, as a fraction of its width.
///
/// A closed eye's contour has no height. The box is given the height an
/// open eye of the same width would have, so the classifier sees the same
/// framing either way — which is what it was trained on.
pub const EYE_BOX_MIN_ASPECT: f32 = 0.4;
/// The box round an eye's lid contour, in the contour's own coordinates:
/// `(x, y, w, h)`.
///
/// Model-free: the contour is whatever the landmark model gave for the ten
/// (or so) points on the lids, in source pixels. `None` for an empty
/// contour or one with no width, which is what a hidden eye's collapsed
/// contour can come to.
pub fn eye_box(contour: &[(f32, f32)]) -> Option<(f32, f32, f32, f32)> {
let (mut x0, mut y0, mut x1, mut y1) = (f32::MAX, f32::MAX, f32::MIN, f32::MIN);
for &(x, y) in contour {
x0 = x0.min(x);
y0 = y0.min(y);
x1 = x1.max(x);
y1 = y1.max(y);
}
let w = x1 - x0;
if contour.is_empty() || w <= 0.0 || w.is_nan() {
return None;
}
let h = (y1 - y0).max(w * EYE_BOX_MIN_ASPECT);
let cy = (y0 + y1) / 2.0;
let (mx, my) = (w * EYE_BOX_MARGIN, h * EYE_BOX_MARGIN);
Some((x0 - mx, cy - h / 2.0 - my, w + 2.0 * mx, h + 2.0 * my))
}
/// One eye, resampled to the classifier's input.
///
/// Constructible only by [`eye_patch`], for the reason [`Aligned112`] is
/// only constructible by [`warp`]: the classifier accepting a plain buffer
/// would accept any 40×24 of anything, and its answer would still be a
/// plausible probability.
#[derive(Debug, Clone, PartialEq)]
pub struct EyePatch {
/// `24 × 40 × 3`, row-major RGB in `0.0..=1.0`.
pixels: Vec<f32>,
/// Source pixels across the box the patch was cut from.
source_px: f32,
}
impl EyePatch {
pub fn pixels(&self) -> &[f32] {
&self.pixels
}
/// Source pixels across the eye box — how much eye there was to read.
///
/// The classifier was trained down to eyes a dozen pixels wide, and
/// below that a crop is an interpolation of nothing; `crate::eyes` draws
/// the line. Zero when the box had no width, which is a hidden eye.
pub fn source_px(&self) -> f32 {
self.source_px
}
/// How sharp the eye the classifier is about to see actually is —
/// [`Aligned112::sharpness`]'s measure, over the patch.
///
/// The reason it exists is the reason the face's does: a soft eye is
/// not a closed one, but a classifier shown a smear says "closed" with
/// the same confidence it says anything, and the only defence is to
/// not ask. A face sharp enough to embed can still hold an eye too soft
/// to read — it is a fortieth of the face — so the measure is taken
/// here and not inherited from the crop.
pub fn sharpness(&self) -> f32 {
laplacian_ratio(&self.pixels, EYE_PATCH_WIDTH, EYE_PATCH_HEIGHT)
}
}
/// Cut an eye out of the source at the classifier's size, from an
/// axis-aligned box in source pixels — [`eye_box`]'s, as a rule.
///
/// Upright and from the frame, not through the face's alignment: the
/// classifier's training crops were detector boxes, and a landmark model's
/// contour already says where the eye is on a tilted head. Bilinear in one
/// step from the native buffer, so a large face gives real pixels; the
/// box's aspect is not preserved, which is what the training resize did.
pub fn eye_patch(
px: Pixels<'_>,
width: usize,
height: usize,
bbox: (f32, f32, f32, f32),
) -> Option<EyePatch> {
let pixels = crop_box(px, width, height, bbox, EYE_PATCH_WIDTH, EYE_PATCH_HEIGHT)?;
Some(EyePatch {
pixels,
source_px: bbox.2,
})
}
// ── sunglasses ────────────────────────────────────────────────────────────
/// Edge of the crop the sunglasses classifier reads. Fixed by the SGC input:
/// 48×48.
pub const SUNGLASSES_EDGE: usize = 48;
/// The windows read for the sunglasses classifier, in template units:
/// `(x, y, w, h)`.
///
/// **Two framings, and the classifier's answer is the higher of the two.**
/// It was trained on a whole-body detector's *head* boxes, and a head box
/// is not reproducible from five landmarks: how much hair and hat it took in
/// depended on the person. So it is shown the face twice — once as the
/// aligned crop itself, once shifted up and widened to take in hair and
/// hat at the cost of the chin, which is roughly where a head box falls —
/// and a pair of sunglasses counts if it looks like one in either.
///
/// Measured over 12 faces in sunglasses and 28 with plainly visible eyes
/// from the reference library (`examples/eyes.rs --head`), at the 0.5
/// threshold:
///
/// | window | sunglasses found | clear eyes kept |
/// |---|---|---|
/// | the aligned face, `(0, 0, 112, 112)` | 9 | 28 |
/// | a head, `(-5, -14, 122, 122)` | 6 | 27 |
/// | a larger head, `(-30, -55, 172, 190)` | 6 | 25 |
/// | **the higher of the first two** | **11** | 27 |
///
/// The face-tight crop alone was the best single framing, which was not the
/// expectation; the head framing found the sunglasses under a cap that the
/// face crop missed. The one clear-eyed face the pair loses wears a cap and
/// clear glasses, at 0.68. Erring towards "sunglasses" is the safe direction
/// for what this feeds: a face called sunglasses is left alone by the
/// eyes-open filter, where a pair of sunglasses missed hands the eye
/// classifier a lens to guess at (docs/dev/faces.md §17).
pub const SUNGLASSES_WINDOWS: [(f32, f32, f32, f32); 2] =
[(0.0, 0.0, 112.0, 112.0), (-5.0, -14.0, 122.0, 122.0)];
/// The framings of one face the sunglasses classifier is shown.
///
/// A newtype for the reason [`EyePatch`] is one.
#[derive(Debug, Clone, PartialEq)]
pub struct HeadViews {
/// Each `48 × 48 × 3`, row-major RGB in `0.0..=1.0`.
views: Vec<Vec<f32>>,
}
impl HeadViews {
pub fn views(&self) -> impl Iterator<Item = &[f32]> {
self.views.iter().map(Vec::as_slice)
}
}
/// Cut the [`SUNGLASSES_WINDOWS`] out of the source, aligned, at the
/// classifier's size.
pub fn head_views(
px: Pixels<'_>,
width: usize,
height: usize,
landmarks: &[(f32, f32); 5],
) -> Option<HeadViews> {
head_views_in(px, width, height, landmarks, &SUNGLASSES_WINDOWS)
}
/// [`head_views`] over windows other than [`SUNGLASSES_WINDOWS`].
///
/// For measuring them, which is how the constant was chosen
/// (`examples/eyes.rs --head`); production callers use the constant.
pub fn head_views_in(
px: Pixels<'_>,
width: usize,
height: usize,
landmarks: &[(f32, f32); 5],
windows: &[(f32, f32, f32, f32)],
) -> Option<HeadViews> {
if !px.fits(width, height) || windows.is_empty() {
return None;
}
let m = fit_similarity(landmarks, &ARCFACE_TEMPLATE)?;
let views = windows
.iter()
.map(|&(x, y, w, h)| {
let window = TemplateWindow { x, y, w, h };
sample_window(
px,
width,
height,
&m,
&window,
SUNGLASSES_EDGE,
SUNGLASSES_EDGE,
)
})
.collect();
Some(HeadViews { views })
}
/// An axis-aligned crop of the source, resampled to `out_w × out_h` RGB.
///
/// `(x, y, w, h)` in source pixels; the aspect is not preserved when it
/// differs from the output's. Bilinear in one step, like every crop here;
/// pixels outside the source read black. What a landmark model trained on
/// detector boxes wants — upright, from the frame — as against the aligned
/// windows above.
pub fn crop_box(
px: Pixels<'_>,
width: usize,
height: usize,
(x, y, w, h): (f32, f32, f32, f32),
out_w: usize,
out_h: usize,
) -> Option<Vec<f32>> {
if !px.fits(width, height) || w <= 0.0 || h <= 0.0 {
return None;
}
let identity = Similarity {
a: 1.0,
b: 0.0,
tx: 0.0,
ty: 0.0,
};
let window = TemplateWindow { x, y, w, h };
Some(sample_window(
px, width, height, &identity, &window, out_w, out_h,
))
}
fn sample_bilinear(px: Pixels<'_>, w: usize, h: usize, x: f32, y: f32, out: &mut [f32]) {
let x0 = x.floor();
let y0 = y.floor();
@@ -489,6 +781,125 @@ mod tests {
}
}
/// A source whose red channel is its x coordinate and green its y, so a
/// crop's mean colour says where in the source it was taken from.
fn coordinate_image(w: usize, h: usize) -> Vec<f32> {
let mut rgb = vec![0.0_f32; w * h * 3];
for y in 0..h {
for x in 0..w {
rgb[(y * w + x) * 3] = x as f32 / w as f32;
rgb[(y * w + x) * 3 + 1] = y as f32 / h as f32;
}
}
rgb
}
fn mean_channel(px: &[f32], c: usize) -> f32 {
let n = px.len() / 3;
px.chunks_exact(3).map(|p| p[c]).sum::<f32>() / n as f32
}
/// The box is the contour's bounds, grown by the margin, and a shut
/// eye's flat contour is given an open eye's height.
#[test]
fn an_eye_box_holds_its_contour_with_a_margin() {
let open = [(100.0, 50.0), (110.0, 46.0), (120.0, 50.0), (110.0, 54.0)];
let (x, y, w, h) = eye_box(&open).unwrap();
assert!((w - 20.0 * (1.0 + 2.0 * EYE_BOX_MARGIN)).abs() < 1e-4);
assert!((h - 8.0 * (1.0 + 2.0 * EYE_BOX_MARGIN)).abs() < 1e-4);
assert!((x + w / 2.0 - 110.0).abs() < 1e-4);
assert!((y + h / 2.0 - 50.0).abs() < 1e-4);
let shut = [(100.0, 50.0), (110.0, 50.0), (120.0, 50.0)];
let (_, _, w2, h2) = eye_box(&shut).unwrap();
assert!((w2 - w).abs() < 1e-4, "same width");
assert!((h2 - 20.0 * EYE_BOX_MIN_ASPECT * (1.0 + 2.0 * EYE_BOX_MARGIN)).abs() < 1e-4);
assert!(eye_box(&[]).is_none());
assert!(eye_box(&[(5.0, 5.0), (5.0, 9.0)]).is_none(), "no width");
}
/// The patch is cut from the box it was given, upright, and knows how
/// many source pixels it spans.
#[test]
fn an_eye_patch_is_the_box_resampled() {
let (w, h) = (200, 200);
let rgb = coordinate_image(w, h);
let bbox = (60.0, 90.0, 30.0, 12.0);
let eye = eye_patch(Pixels::RgbF32(&rgb), w, h, bbox).unwrap();
assert_eq!(eye.pixels().len(), EYE_PATCH_WIDTH * EYE_PATCH_HEIGHT * 3);
assert_eq!(eye.source_px(), 30.0);
let cx = mean_channel(eye.pixels(), 0) * w as f32;
let cy = mean_channel(eye.pixels(), 1) * h as f32;
assert!((cx - 75.0).abs() < 0.6, "{cx}");
assert!((cy - 96.0).abs() < 0.6, "{cy}");
// No width, or a buffer that is not the size it claims: nothing.
assert!(eye_patch(Pixels::RgbF32(&rgb), w, h, (60.0, 90.0, 0.0, 12.0)).is_none());
assert!(eye_patch(Pixels::RgbF32(&rgb), 190, 200, bbox).is_none());
}
/// A soft eye scores lower than the same eye sharp, on the patch itself.
#[test]
fn an_eye_patchs_sharpness_falls_with_blur() {
let edge = 120;
let sharp = image(
edge,
|x, y| if (x / 5 + y / 5) % 2 == 0 { 0.9 } else { 0.1 },
);
let soft = blur(&blur(&sharp, edge), edge);
let bbox = (20.0, 40.0, 40.0, 24.0);
let a = eye_patch(Pixels::RgbF32(&sharp), edge, edge, bbox)
.unwrap()
.sharpness();
let b = eye_patch(Pixels::RgbF32(&soft), edge, edge, bbox)
.unwrap()
.sharpness();
assert!(a > b * 2.0, "sharp {a} should clearly beat blurred {b}");
}
/// The second sunglasses framing takes in more than the face — it starts
/// above the template's top edge and ends below its bottom — and the
/// first is the aligned face itself.
#[test]
fn the_head_views_are_the_face_and_a_wider_framing_of_it() {
let (w, h) = (300, 300);
let rgb = coordinate_image(w, h);
let lm = shifted_scaled(1.0, 100.0, 100.0, 0.0);
let head = head_views(Pixels::RgbF32(&rgb), w, h, &lm).unwrap();
let views: Vec<&[f32]> = head.views().collect();
let face = warp(&rgb, w, h, &lm).unwrap();
assert_eq!(views.len(), SUNGLASSES_WINDOWS.len());
for v in &views {
assert_eq!(v.len(), SUNGLASSES_EDGE * SUNGLASSES_EDGE * 3);
}
// The face view samples the same region as the aligned crop.
assert!((mean_channel(views[0], 0) - mean_channel(face.pixels(), 0)).abs() < 0.01);
assert!((mean_channel(views[0], 1) - mean_channel(face.pixels(), 1)).abs() < 0.01);
let (x, y, ww, hh) = SUNGLASSES_WINDOWS[1];
assert!(
x < 0.0 && y < 0.0,
"the window starts outside the face crop"
);
assert!(x + ww > ALIGNED_EDGE as f32, "and is wider than it");
assert!(y + hh < ALIGNED_EDGE as f32, "but stops short of the chin");
// Centred horizontally on the face, so the two share a mean x.
assert!((mean_channel(views[1], 0) - mean_channel(face.pixels(), 0)).abs() < 0.01);
// Its first row lies above the face's first row.
assert!(views[1][1] < face.pixels()[1]);
}
#[test]
fn degenerate_landmarks_yield_no_head_crop() {
let rgb = vec![0.5_f32; 64 * 64 * 3];
let degenerate = [(50.0, 50.0); 5];
assert!(head_views(Pixels::RgbF32(&rgb), 64, 64, &degenerate).is_none());
// And a buffer that is not the size it claims.
let lm = shifted_scaled(1.0, 0.0, 0.0, 0.0);
assert!(head_views(Pixels::RgbF32(&rgb), 60, 60, &lm).is_none());
}
#[test]
fn out_of_bounds_samples_read_black_rather_than_wrapping() {
let rgb = vec![1.0_f32; 32 * 32 * 3];
+3 -3
View File
@@ -1,4 +1,4 @@
//! Cosine to probability (docs/faces.md §8, FR-CULL-9).
//! Cosine to probability (docs/dev/faces.md §8, FR-CULL-9).
//!
//! FR-CULL-9 is a hard requirement rather than an implementation detail: no
//! code path may threshold a bare cosine, every threshold in the subsystem is
@@ -28,7 +28,7 @@
//! calibration to the belief it was supposed to test — and that is the whole
//! of the alternative.
//!
//! docs/faces.md §8.1 names one more that would cost no labelling at all: two
//! docs/dev/faces.md §8.1 names one more that would cost no labelling at all: two
//! faces in adjacent frames of one burst are near-certainly the same person,
//! and FR-CULL-5's grouping is sitting there. Nothing draws on it. This crate
//! cannot see a catalog, let alone the bursts in one — it is handed cosines by
@@ -83,7 +83,7 @@ pub struct Calibration {
}
impl Default for Calibration {
/// The reference implementation's fitted MBF curve (docs/faces.md §1):
/// The reference implementation's fitted MBF curve (docs/dev/faces.md §1):
/// steepness 16.2, P=0.5 at cosine 0.267.
///
/// **`valid` is false**, and that is the point. It is a documented
+263
View File
@@ -0,0 +1,263 @@
//! TRACES: FR-CULL-8a
//! The two small classifiers behind a face's eye state (docs/dev/faces.md §17).
//!
//! **OCEC** — *open closed eyes classification*, Hyodo 2025 — reads one
//! 40×24 eye and answers P(open). **SGC** — *sunglasses classification*,
//! Hyodo 2026 — reads a 48×48 head and answers P(sunglasses); it is shown
//! two framings of each face and the higher answer stands, for the reason
//! [`crate::align::SUNGLASSES_WINDOWS`] gives. Both are
//! depthwise-separable CNNs of a few hundred kilobytes, both MIT with their
//! weights, and both were exported with BatchNorm already folded, which is
//! about the friendliest graph tract can be handed.
//!
//! Neither takes a plain buffer. [`EyeClassifier::classify`] takes an
//! [`EyePatch`] and [`SunglassesClassifier::classify`] a [`HeadViews`], each
//! constructible only by the crop in [`crate::align`] that puts the right
//! pixels in it — the same defence [`crate::embed::Embedder`] makes with
//! [`crate::align::Aligned112`], for the same reason: a classifier handed the
//! wrong region returns a confident probability of nothing. Where the eye
//! box comes from is [`crate::landmarks`]; [`EyeModels::read`] is the whole
//! chain.
//!
//! # The graphs must have a fixed batch
//!
//! Both ship with a dynamic batch dimension, which tract will not analyse.
//! `tools/fix-face-model-shapes.sh` pins it to 1, exactly as it does for the
//! embedder; the shipped files are the pinned ones.
//!
//! # Pre-processing
//!
//! Read off the reference demos rather than assumed: RGB, `x / 255`, NCHW,
//! the crop resized to the input with bilinear interpolation and **without**
//! preserving its aspect. [`crate::align`]'s crops arrive already at the
//! input size in `0..=1`, so there is nothing left to do but lay them out.
use ndarray::Array4;
use crate::align::{
eye_box, eye_patch, head_views, EyePatch, HeadViews, EYE_PATCH_HEIGHT, EYE_PATCH_WIDTH,
SUNGLASSES_EDGE,
};
use crate::eyes::{Eye, EyeReading};
use crate::landmarks::{Landmarker, Landmarks};
use crate::{FaceError, Pixels};
use dr_inference_engine::{Form, Model, Role};
/// A loaded OCEC graph.
pub struct EyeClassifier {
session: Model,
}
/// A loaded SGC graph.
pub struct SunglassesClassifier {
session: Model,
}
/// Open a single-input, single-output classifier and check it is the shape
/// the crop feeding it will be.
///
/// The check is against the *input*, because that is where these two graphs
/// differ from each other and from everything else in this crate: an SGC file
/// given to the eye classifier would otherwise be resized into by an eye
/// patch, and answer. `expected` names the model in the error.
fn open_classifier(
bytes: &[u8],
expected: &'static str,
(h, w): (usize, usize),
) -> Result<Model, FaceError> {
let model = dr_inference_engine::open(Role::EyeClassifier, Form::F32, bytes)?;
let acquired = model.acquire()?;
let session = acquired.lock();
let input = session.inputs().first().ok_or(FaceError::WrongModel {
expected,
detail: "model has no inputs".into(),
})?;
let shape: Option<Vec<i64>> = input.dtype().tensor_shape().map(|s| s.to_vec());
let want = [1, 3, h as i64, w as i64];
if shape.as_deref() != Some(&want[..]) {
return Err(FaceError::WrongModel {
expected,
detail: format!(
"input '{}' is {:?}, expected {:?} (batch pinned to 1)",
input.name(),
shape,
want
),
});
}
if session.outputs().len() != 1 {
return Err(FaceError::WrongModel {
expected,
detail: format!("{} outputs, expected one", session.outputs().len()),
});
}
drop(session);
drop(acquired);
Ok(model)
}
/// Lay a `h × w` RGB crop out as the `[1, 3, h, w]` tensor both graphs take.
fn to_nchw(pixels: &[f32], h: usize, w: usize) -> Array4<f32> {
let mut input = Array4::<f32>::zeros((1, 3, h, w));
for y in 0..h {
for x in 0..w {
for c in 0..3 {
input[[0, c, y, x]] = pixels[(y * w + x) * 3 + c];
}
}
}
input
}
/// Run a one-number classifier and read its sigmoid back, clamped.
fn run_scalar(model: &Model, input: Array4<f32>, expected: &'static str) -> Result<f32, FaceError> {
let acquired = model.acquire()?;
let mut session = acquired.lock();
let outputs = session
.run(ort::inputs![
ort::value::Tensor::from_array(input).map_err(FaceError::Inference)?
])
.map_err(FaceError::Inference)?;
let (_, data) = outputs[0]
.try_extract_tensor::<f32>()
.map_err(FaceError::Inference)?;
let Some(&p) = data.first() else {
return Err(FaceError::WrongModel {
expected,
detail: "empty output".into(),
});
};
// The graph ends in a sigmoid, so this is a clamp against rounding and
// nothing more — the reference demo does the same.
Ok(p.clamp(0.0, 1.0))
}
impl EyeClassifier {
pub fn from_path(path: impl AsRef<std::path::Path>) -> Result<Self, FaceError> {
let bytes = std::fs::read(path).map_err(FaceError::ModelRead)?;
Self::from_bytes(&bytes)
}
pub fn from_bytes(bytes: &[u8]) -> Result<Self, FaceError> {
Ok(Self {
session: open_classifier(bytes, "OCEC", (EYE_PATCH_HEIGHT, EYE_PATCH_WIDTH))?,
})
}
/// P(open) for one eye.
pub fn classify(&mut self, eye: &EyePatch) -> Result<f32, FaceError> {
let input = to_nchw(eye.pixels(), EYE_PATCH_HEIGHT, EYE_PATCH_WIDTH);
run_scalar(&self.session, input, "OCEC")
}
}
impl SunglassesClassifier {
pub fn from_path(path: impl AsRef<std::path::Path>) -> Result<Self, FaceError> {
let bytes = std::fs::read(path).map_err(FaceError::ModelRead)?;
Self::from_bytes(&bytes)
}
pub fn from_bytes(bytes: &[u8]) -> Result<Self, FaceError> {
Ok(Self {
session: open_classifier(bytes, "SGC", (SUNGLASSES_EDGE, SUNGLASSES_EDGE))?,
})
}
/// P(sunglasses) for one head: the highest answer over its framings.
pub fn classify(&mut self, head: &HeadViews) -> Result<f32, FaceError> {
let mut best = 0.0_f32;
for view in head.views() {
let input = to_nchw(view, SUNGLASSES_EDGE, SUNGLASSES_EDGE);
best = best.max(run_scalar(&self.session, input, "SGC")?);
}
Ok(best)
}
}
/// The three models behind a reading, which is how every caller holds them.
///
/// One struct rather than three optional parameters, because a partial
/// reading is not a reading: an eye state with no sunglasses number behind
/// it is exactly the beach-photograph failure [`crate::eyes`] describes, and
/// an eye box without the landmarks is the loose one this module replaced.
/// The models load together or not at all.
pub struct EyeModels {
pub landmarks: Landmarker,
pub eyes: EyeClassifier,
pub sunglasses: SunglassesClassifier,
}
impl EyeModels {
pub fn from_paths(
landmarks: impl AsRef<std::path::Path>,
eyes: impl AsRef<std::path::Path>,
sunglasses: impl AsRef<std::path::Path>,
) -> Result<Self, FaceError> {
Ok(Self {
landmarks: Landmarker::from_path(landmarks)?,
eyes: EyeClassifier::from_path(eyes)?,
sunglasses: SunglassesClassifier::from_path(sunglasses)?,
})
}
/// Read one face's eyes, and hand back the dense landmarks it read them
/// from.
///
/// `bbox` is the detector's `(x0, y0, x1, y1)` and `landmarks5` its five
/// points, both in source pixels; the buffer is the one the aligned
/// crop was taken from, so an eye is read from the same pixels the
/// embedder saw the face in. `None` where nothing could be cut — a
/// degenerate box or landmarks — which the caller stores as "not read".
///
/// The landmarks come back because they cost a model run the caller will
/// not want to pay twice: stored beside the reading, a later pass over
/// faces — head pose, expression — has them without the original.
pub fn read(
&mut self,
px: Pixels<'_>,
width: usize,
height: usize,
bbox: (f32, f32, f32, f32),
landmarks5: &[(f32, f32); 5],
) -> Result<Option<(EyeReading, Landmarks)>, FaceError> {
let Some(lm) = self.landmarks.landmarks(px, width, height, bbox)? else {
return Ok(None);
};
let Some(head) = head_views(px, width, height, landmarks5) else {
return Ok(None);
};
let mut eye = |contour: &[(f32, f32)]| -> Result<Eye, FaceError> {
// A hidden eye's contour can collapse to no width. Its numbers
// are then zero — no pixels, no sharpness — which is what the
// rule in `crate::eyes` reads as "not readable".
let Some(b) = eye_box(contour) else {
return Ok(Eye {
open: 0.0,
px: 0.0,
sharpness: 0.0,
});
};
let Some(patch) = eye_patch(px, width, height, b) else {
return Ok(Eye {
open: 0.0,
px: 0.0,
sharpness: 0.0,
});
};
Ok(Eye {
open: self.eyes.classify(&patch)?,
px: patch.source_px(),
sharpness: patch.sharpness(),
})
};
let right = eye(&lm.right_eye())?;
let left = eye(&lm.left_eye())?;
let reading = EyeReading {
right,
left,
sunglasses: self.sunglasses.classify(&head)?,
};
Ok(Some((reading, lm)))
}
}
+1 -1
View File
@@ -1,4 +1,4 @@
//! Grouping faces into people (docs/faces.md §9, FR-CULL-10).
//! Grouping faces into people (docs/dev/faces.md §9, FR-CULL-10).
//!
//! Model-free: this is arithmetic over embeddings, and it is where the
//! subsystem's accuracy actually lives, so it is testable with no weights on
+39 -17
View File
@@ -1,4 +1,4 @@
//! SCRFD face detection (docs/faces.md §4).
//! SCRFD face detection (docs/dev/faces.md §4).
//!
//! One forward pass produces a box, a confidence and **five landmarks** per
//! face — the landmarks being the reason for this detector rather than a
@@ -14,7 +14,8 @@
use ndarray::Array4;
use crate::{install_backend, FaceError};
use crate::FaceError;
use dr_inference_engine::{Form, Model, Role};
/// The graph's input edge, in pixels. See the module note: not configurable.
pub const INPUT_EDGE: usize = 640;
@@ -135,7 +136,10 @@ impl Detection {
/// A loaded SCRFD graph.
pub struct Detector {
session: ort::session::Session,
session: Model,
/// f32 or int8 — the int8 form finds a different set of faces and is a
/// different detector in `model_id` (docs/dev/inference.md §7).
form: Form,
/// Feature-map count: 3 for strides {8,16,32}, 4 for {8,16,32,64}.
///
/// Discovered from the output count rather than assumed, because both
@@ -145,18 +149,29 @@ pub struct Detector {
}
impl Detector {
pub fn from_path(path: impl AsRef<std::path::Path>) -> Result<Self, FaceError> {
let bytes = std::fs::read(path).map_err(FaceError::ModelRead)?;
Self::from_bytes(&bytes)
/// Which form this detector was loaded from.
pub fn form(&self) -> Form {
self.form
}
pub fn from_bytes(bytes: &[u8]) -> Result<Self, FaceError> {
install_backend();
/// Load the canonical f32 file at `path`, or the form the device's
/// backend wants instead — the `.int8.onnx` beside it on a Hexagon —
/// which [`Detector::form`] then reports.
pub fn from_path(path: impl AsRef<std::path::Path>) -> Result<Self, FaceError> {
let (path, form) = dr_inference_engine::resolve_model(Role::Detector, path.as_ref());
let bytes = std::fs::read(path).map_err(FaceError::ModelRead)?;
Self::from_bytes_in(&bytes, form)
}
let session = ort::session::Session::builder()
.map_err(FaceError::Inference)?
.commit_from_memory(bytes)
.map_err(FaceError::Inference)?;
/// An f32 graph from memory.
pub fn from_bytes(bytes: &[u8]) -> Result<Self, FaceError> {
Self::from_bytes_in(bytes, Form::F32)
}
fn from_bytes_in(bytes: &[u8], form: Form) -> Result<Self, FaceError> {
let model = dr_inference_engine::open(Role::Detector, form, bytes)?;
let acquired = model.acquire()?;
let session = acquired.lock();
let n_out = session.outputs().len();
if n_out % 3 != 0 || !(9..=12).contains(&n_out) {
@@ -191,7 +206,13 @@ impl Detector {
}
}
Ok(Self { session, fmc })
drop(session);
drop(acquired);
Ok(Self {
session: model,
form,
fmc,
})
}
/// Stride levels this graph emits.
@@ -223,8 +244,9 @@ impl Detector {
let lb = Letterbox::fit(width as f32, height as f32);
let input = lb.sample(rgb, width, height);
let outputs = self
.session
let acquired = self.session.acquire()?;
let mut session = acquired.lock();
let outputs = session
.run(ort::inputs![
ort::value::Tensor::from_array(input).map_err(FaceError::Inference)?
])
@@ -324,7 +346,7 @@ fn iou(a: &(f32, f32, f32, f32), b: &(f32, f32, f32, f32)) -> f32 {
/// How the image is fitted into the graph's fixed square input.
///
/// The forward and inverse mappings live in one struct on purpose:
/// docs/faces.md §4.1 notes that what matters is not *where* the padding goes
/// docs/dev/faces.md §4.1 notes that what matters is not *where* the padding goes
/// but that the two agree. A mismatch offsets every box and landmark by the
/// padding, producing detections that look plausible and embeddings that
/// quietly cluster badly three stages later.
@@ -350,7 +372,7 @@ impl Letterbox {
///
/// `(x·255 − 127.5) / 128` — note `/128`, not `/127.5`. The reference
/// implementation this is ported from uses `/128` for both models, and
/// every measured number in docs/faces.md §1 came from it.
/// every measured number in docs/dev/faces.md §1 came from it.
///
/// Padding is grey, matching the reference's `114`: the value the network
/// reads least as an edge, where black would draw a hard border across the
+18 -12
View File
@@ -1,4 +1,4 @@
//! ArcFace / MobileFaceNet inference (docs/faces.md §6).
//! ArcFace / MobileFaceNet inference (docs/dev/faces.md §6).
//!
//! Takes an aligned crop and returns 512 L2-normalised floats. The alignment is
//! not optional and cannot be skipped by accident: [`Embedder::embed`] takes an
@@ -14,7 +14,8 @@ use ndarray::Array4;
use crate::align::{Aligned112, ALIGNED_EDGE};
use crate::embedding::{normalise, Embedding, ModelId, EMBEDDING_DIM};
use crate::{install_backend, FaceError};
use crate::FaceError;
use dr_inference_engine::{Form, Model, Role};
/// What one pass of the embedder produces: the direction, and the length.
///
@@ -53,7 +54,7 @@ impl Embedded {
/// A loaded ArcFace graph.
pub struct Embedder {
session: ort::session::Session,
session: Model,
model: ModelId,
}
@@ -64,12 +65,11 @@ impl Embedder {
}
pub fn from_bytes(bytes: &[u8], model: ModelId) -> Result<Self, FaceError> {
install_backend();
let session = ort::session::Session::builder()
.map_err(FaceError::Inference)?
.commit_from_memory(bytes)
.map_err(FaceError::Inference)?;
// Always the f32 form: an embedding must compare across devices
// (docs/dev/inference.md §7), and the engine pins this role to it.
let loaded = dr_inference_engine::open(Role::Embedder, Form::F32, bytes)?;
let acquired = loaded.acquire()?;
let session = acquired.lock();
// One output, `[1, 512]`. Checked because an ArcFace variant with a
// different embedding width would otherwise be read as a truncated
@@ -90,7 +90,12 @@ impl Embedder {
});
}
Ok(Self { session, model })
drop(session);
drop(acquired);
Ok(Self {
session: loaded,
model,
})
}
pub fn model(&self) -> &ModelId {
@@ -111,8 +116,9 @@ impl Embedder {
}
}
let outputs = self
.session
let acquired = self.session.acquire()?;
let mut session = acquired.lock();
let outputs = session
.run(ort::inputs![
ort::value::Tensor::from_array(input).map_err(FaceError::Inference)?
])
+2 -2
View File
@@ -1,4 +1,4 @@
//! What an embedder produces, and how it is stored (docs/faces.md §6).
//! What an embedder produces, and how it is stored (docs/dev/faces.md §6).
//!
//! Deliberately **model-free**: the vector, its identity, its comparison and
//! its storage encoding are arithmetic, and `calibrate` and `cluster` are built
@@ -273,7 +273,7 @@ mod tests {
);
}
/// The claim docs/faces.md §6 makes about the storage format: the f16
/// The claim docs/dev/faces.md §6 makes about the storage format: the f16
/// round-trip costs ~1e-3 of cosine, three orders below the separation
/// between a match and a non-match.
#[test]
+263
View File
@@ -0,0 +1,263 @@
//! TRACES: FR-CULL-8a
//! What a face's eyes are doing, and how the numbers behind it are read.
//!
//! Model-free: the models in [`crate::classify`] produce the numbers, and
//! everything that interprets them — the catalog's filter, the People
//! screen's label — comes through here, so a threshold lives in exactly one
//! place.
//!
//! # Seven numbers, one answer
//!
//! An eye classifier answers "open or closed" for whatever it is shown, and
//! it is shown three things it cannot answer for. **Dark glass**: over
//! sunglasses it answers anyway, confidently, for a state that cannot be
//! seen — so the reading carries P(sunglasses) from a classifier that looks
//! at the whole head, and that takes precedence. **A smear**: a soft eye is
//! not a closed one, but shown a blur the classifier says "closed" with the
//! same confidence it says anything, and on the reference library that was
//! the commonest wrong answer of all — small faces, motion, a proxy where
//! the native render should have been. So each eye carries how many source
//! pixels it spanned and how sharp the patch was, and an eye under either
//! floor is not asked. **A cheek**: a head turned far enough hides its far
//! eye, and the landmark contour of a hidden eye collapses to a sliver; an
//! eye much narrower than its partner is not asked either.
//!
//! The two eyes are kept apart rather than averaged. A wink is one eye
//! closed, and averaging it lands at 0.5 — the one value that says the least.
//! [`EyeState::Open`] requires every eye that *could be read* to be open;
//! a face with no readable eye is [`EyeState::Unreadable`], which is not a
//! blink and not open, and a filter for either leaves it alone.
/// One eye's numbers.
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct Eye {
/// P(open), the classifier's sigmoid.
pub open: f32,
/// Source pixels across the eye box — [`crate::align::EyePatch::source_px`].
pub px: f32,
/// [`crate::align::EyePatch::sharpness`] of the patch the classifier saw.
pub sharpness: f32,
}
/// The numbers the models produced for one face.
///
/// Stored per face, nullable as a whole: a face indexed before the eye models
/// existed, or on a device without them, has no reading rather than a
/// reading of zeros.
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct EyeReading {
/// The subject's **right** eye — image-left.
pub right: Eye,
/// The subject's **left** eye — image-right.
pub left: Eye,
/// P(the head wears sunglasses).
pub sunglasses: f32,
}
/// Above this an eye is open. The classifier's own decision point; its
/// training put the two classes either side of a sigmoid and this is where
/// the sigmoid crosses.
pub const EYES_OPEN_THRESHOLD: f32 = 0.5;
/// Above this the head wears sunglasses and the eye readings are moot.
pub const SUNGLASSES_THRESHOLD: f32 = 0.5;
/// Fewest source pixels across an eye box for the eye to be read.
///
/// The classifier was trained on eyes down to about a dozen pixels wide
/// (its reference footage averaged 15–21); below that the 40-pixel patch is
/// an interpolation of nothing, and the answer is noise that reads as
/// "closed". docs/dev/faces.md §17.3 has the measurement behind the number.
pub const MIN_EYE_PX: f32 = 12.0;
/// Least [`Eye::sharpness`] for the eye to be read.
///
/// The same measure as the face's `min_sharpness`, over the eye patch, and
/// chosen the same way: the value under which the open-eyed faces of the
/// reference sample were being called closed. docs/dev/faces.md §17.3.
pub const MIN_EYE_SHARPNESS: f32 = 0.02;
/// An eye narrower than this fraction of its partner is the far eye of a
/// turned head, out of view behind the nose, and is not read.
///
/// A landmark model's contour for a hidden eye collapses towards the nose.
/// Measured on twenty native renders of the reference library
/// (docs/dev/faces.md §17.4): profiles put the far eye at 0.02–0.43 of the near
/// one, two three-quarter faces whose far eye read closed sat at 0.54, and
/// every face looking at the camera — winks included, since a shut eye's
/// box keeps its width — sat at 0.78 or more. 0.6 splits the gap.
pub const HIDDEN_EYE_RATIO: f32 = 0.6;
/// What the reading says, for a screen or a filter.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum EyeState {
/// Every eye that could be read is open.
Open,
/// An eye that could be read is closed — a blink, or a wink.
Closed,
/// The eyes cannot be seen. Neither open nor closed, and a filter for
/// either leaves the face alone.
Sunglasses,
/// No eye was sharp enough, large enough and in view to read. Neither
/// open nor closed, like sunglasses, and left alone by every filter.
Unreadable,
}
impl Eye {
/// Whether this eye can be read at all: enough pixels, sharp enough,
/// and not the collapsed contour of a hidden eye — measured against
/// `other`, its partner.
pub fn readable(&self, other: &Eye) -> bool {
self.px >= MIN_EYE_PX
&& self.sharpness >= MIN_EYE_SHARPNESS
&& self.px >= other.px * HIDDEN_EYE_RATIO
}
}
impl EyeReading {
pub fn state(&self) -> EyeState {
if self.sunglasses >= SUNGLASSES_THRESHOLD {
return EyeState::Sunglasses;
}
let readable = [
self.right.readable(&self.left).then_some(self.right.open),
self.left.readable(&self.right).then_some(self.left.open),
];
let mut any = false;
for open in readable.into_iter().flatten() {
any = true;
if open < EYES_OPEN_THRESHOLD {
return EyeState::Closed;
}
}
if any {
EyeState::Open
} else {
EyeState::Unreadable
}
}
/// Whether this is a face a "no one blinking" filter should drop.
///
/// The filter's question, rather than [`EyeState`]'s four-way answer,
/// because the two differ on exactly the cases that matter: a face
/// behind sunglasses, or one whose eyes could not be read, is not open
/// — and it is not a blink either. Only [`EyeState::Closed`] is one.
pub fn is_blink(&self) -> bool {
self.state() == EyeState::Closed
}
}
impl EyeState {
/// The word the People screen puts on the face.
pub fn label(&self) -> &'static str {
match self {
EyeState::Open => "Eyes open",
EyeState::Closed => "Eyes closed",
EyeState::Sunglasses => "Sunglasses",
EyeState::Unreadable => "Eyes unclear",
}
}
}
#[cfg(test)]
mod tests {
use super::*;
fn eye(open: f32) -> Eye {
Eye {
open,
px: 40.0,
sharpness: 0.1,
}
}
fn reading(right: f32, left: f32, sunglasses: f32) -> EyeReading {
EyeReading {
right: eye(right),
left: eye(left),
sunglasses,
}
}
#[test]
fn both_eyes_open_is_open() {
assert_eq!(reading(0.9, 0.8, 0.1).state(), EyeState::Open);
assert!(!reading(0.9, 0.8, 0.1).is_blink());
}
/// A wink is not "eyes open": one eye closed lands the same place a
/// blink does, and a filter for "nobody blinking" should drop it.
#[test]
fn one_eye_closed_is_closed() {
assert_eq!(reading(0.9, 0.2, 0.1).state(), EyeState::Closed);
assert_eq!(reading(0.2, 0.9, 0.1).state(), EyeState::Closed);
assert!(reading(0.2, 0.9, 0.1).is_blink());
}
/// The whole reason the sunglasses number exists: whatever the eye
/// classifier says over dark glass, it is not a reading of the eyes.
#[test]
fn sunglasses_override_the_eye_readings_either_way() {
assert_eq!(reading(0.9, 0.9, 0.8).state(), EyeState::Sunglasses);
assert_eq!(reading(0.1, 0.1, 0.8).state(), EyeState::Sunglasses);
assert!(!reading(0.1, 0.1, 0.8).is_blink());
}
/// A soft or tiny eye is not asked; if neither can be, the face is
/// unreadable rather than closed.
#[test]
fn a_soft_or_tiny_eye_is_not_read() {
let mut r = reading(0.1, 0.9, 0.0);
r.right.sharpness = MIN_EYE_SHARPNESS / 2.0;
assert_eq!(r.state(), EyeState::Open, "the soft closed eye is ignored");
let mut r = reading(0.1, 0.9, 0.0);
r.right.px = MIN_EYE_PX - 1.0;
assert_eq!(r.state(), EyeState::Open, "the tiny closed eye is ignored");
let mut r = reading(0.1, 0.1, 0.0);
r.right.sharpness = 0.0;
r.left.px = 3.0;
assert_eq!(r.state(), EyeState::Unreadable);
assert!(!r.is_blink());
assert_eq!(r.state().label(), "Eyes unclear");
}
/// A profile: the far eye's contour collapses, and the sliver is not
/// read. The near eye still decides.
#[test]
fn a_turned_heads_collapsed_far_eye_is_not_read() {
let mut r = reading(0.05, 0.95, 0.0);
r.right.px = 40.0 * HIDDEN_EYE_RATIO - 1.0;
assert!(!r.right.readable(&r.left));
assert_eq!(r.state(), EyeState::Open);
let mut blink = reading(0.95, 0.05, 0.0);
blink.right.px = 40.0 * HIDDEN_EYE_RATIO - 1.0;
assert_eq!(blink.state(), EyeState::Closed);
// Both eyes narrow but alike is not a turned head: both count.
let mut small = reading(0.05, 0.95, 0.0);
small.right.px = 14.0;
small.left.px = 14.0;
assert_eq!(small.state(), EyeState::Closed);
}
#[test]
fn the_thresholds_are_inclusive_at_the_decision_point() {
assert_eq!(
reading(EYES_OPEN_THRESHOLD, EYES_OPEN_THRESHOLD, 0.0).state(),
EyeState::Open
);
assert_eq!(
reading(1.0, 1.0, SUNGLASSES_THRESHOLD).state(),
EyeState::Sunglasses
);
let mut r = reading(1.0, 1.0, 0.0);
r.right.px = MIN_EYE_PX;
r.left.px = MIN_EYE_PX;
r.right.sharpness = MIN_EYE_SHARPNESS;
assert!(r.right.readable(&r.left));
}
}
+263
View File
@@ -0,0 +1,263 @@
//! TRACES: FR-CULL-8a
//! Dense facial landmarks — InsightFace's `2d106det` (docs/dev/faces.md §17.2).
//!
//! SCRFD's five points place a face; they do not place an eye. Its eye
//! point is loose enough that a window centred on it left the eye in a
//! corner on turned and smiling heads, and two model-free ways of
//! re-centring it made things worse. So a second model draws the eye's lid
//! contour, and the eye box is cut from that.
//!
//! **Why this one.** Three were measured on the same faces — MediaPipe Face
//! Mesh V2, PIPNet and this — and tied on what the eye classifier made of
//! their boxes (22 of 25 open eyes read open, against 19 from the SCRFD
//! point). This is the cheapest of the three by a wide margin (5 MB, 106
//! points, ~24 ms in tract), and it is under the grant the detector and
//! embedder already carry rather than a new one to read.
//!
//! # Pre-processing
//!
//! Ported from InsightFace's `landmark.py`: a square crop centred on the
//! detector box, 1.5× its longer edge, resized to 192; **RGB in 0..255**
//! (the graph carries its own `bn_data` normalisation, so `input_mean` is
//! 0 and `input_std` 1); 106 `(x, y)` in −1..1 mapped back through
//! `(p + 1) · 96`. The graph's batch dimension is the literal `None` and
//! is pinned to 1 by `tools/fix-face-model-shapes.sh`, like the embedder's.
//!
//! # The layout
//!
//! Checked by drawing the points on the reference faces rather than taken
//! from a diagram: the subject's right eye (image-left) is points 33–42,
//! the left 87–96, ten each round the lids.
use ndarray::Array4;
use crate::align::crop_box;
use crate::{FaceError, Pixels};
use dr_inference_engine::{Form, Model, Role};
/// The graph's input edge, in pixels.
pub const INPUT_EDGE: usize = 192;
/// How many points the model returns.
pub const POINTS: usize = 106;
/// The crop's edge as a multiple of the detector box's longer edge.
const CROP_SCALE: f32 = 1.5;
/// The span of the frame, in long-edge units, the packed form covers: a
/// quarter of the frame outside each edge.
pub const PACKED_RANGE: (f32, f32) = (-0.25, 1.25);
/// Bytes the packed form of one face's landmarks takes.
pub const PACKED_BYTES: usize = POINTS * 4;
/// Point indices of the subject's right eye's lid contour (image-left).
pub const RIGHT_EYE: [usize; 10] = [33, 34, 35, 36, 37, 38, 39, 40, 41, 42];
/// Point indices of the subject's left eye's lid contour (image-right).
pub const LEFT_EYE: [usize; 10] = [87, 88, 89, 90, 91, 92, 93, 94, 95, 96];
/// The 106 points of one face, in **source pixels**.
#[derive(Debug, Clone, PartialEq)]
pub struct Landmarks {
pub points: [(f32, f32); POINTS],
}
impl Landmarks {
/// Storage form: `106 × (x, y)` as little-endian **`u16` fixed point**
/// over the frame, 424 bytes.
///
/// Each coordinate is normalised by `long_edge` like the five points the
/// catalog already keeps, then mapped over [`PACKED_RANGE`] — a quarter
/// of the frame either side of it, because a landmark on a face at the
/// edge does land outside the image — onto 0..65535. That is 0.14 source
/// pixels on a 6000-pixel frame. `f16` would be the same size and worse:
/// its three significant figures near 1.0 are six pixels at that scale,
/// and the eye contour this is kept for is drawn to the pixel.
pub fn to_packed_bytes(&self, long_edge: f32) -> Vec<u8> {
let (lo, hi) = PACKED_RANGE;
let pack = |v: f32| -> [u8; 2] {
let t = ((v / long_edge - lo) / (hi - lo)).clamp(0.0, 1.0);
((t * 65535.0).round() as u16).to_le_bytes()
};
let mut out = Vec::with_capacity(POINTS * 4);
for &(x, y) in &self.points {
out.extend_from_slice(&pack(x));
out.extend_from_slice(&pack(y));
}
out
}
/// [`Self::to_packed_bytes`] read back, into source pixels of a frame
/// with this `long_edge`. `None` for a blob of the wrong length.
pub fn from_packed_bytes(bytes: &[u8], long_edge: f32) -> Option<Self> {
if bytes.len() != POINTS * 4 {
return None;
}
let (lo, hi) = PACKED_RANGE;
let unpack = |b: &[u8]| -> f32 {
let t = u16::from_le_bytes([b[0], b[1]]) as f32 / 65535.0;
(t * (hi - lo) + lo) * long_edge
};
let mut points = [(0.0_f32, 0.0_f32); POINTS];
for (i, p) in points.iter_mut().enumerate() {
let at = i * 4;
*p = (unpack(&bytes[at..at + 2]), unpack(&bytes[at + 2..at + 4]));
}
Some(Self { points })
}
/// The lid contour of the subject's right eye.
pub fn right_eye(&self) -> [(f32, f32); 10] {
RIGHT_EYE.map(|i| self.points[i])
}
/// The lid contour of the subject's left eye.
pub fn left_eye(&self) -> [(f32, f32); 10] {
LEFT_EYE.map(|i| self.points[i])
}
}
/// A loaded `2d106det` graph.
pub struct Landmarker {
session: Model,
}
impl Landmarker {
pub fn from_path(path: impl AsRef<std::path::Path>) -> Result<Self, FaceError> {
let bytes = std::fs::read(path).map_err(FaceError::ModelRead)?;
Self::from_bytes(&bytes)
}
pub fn from_bytes(bytes: &[u8]) -> Result<Self, FaceError> {
let model = dr_inference_engine::open(Role::Landmarks, Form::F32, bytes)?;
let acquired = model.acquire()?;
let session = acquired.lock();
let input = session.inputs().first().ok_or(FaceError::WrongModel {
expected: "2d106det",
detail: "model has no inputs".into(),
})?;
let shape: Option<Vec<i64>> = input.dtype().tensor_shape().map(|s| s.to_vec());
let want = [1, 3, INPUT_EDGE as i64, INPUT_EDGE as i64];
if shape.as_deref() != Some(&want[..]) {
return Err(FaceError::WrongModel {
expected: "2d106det",
detail: format!(
"input '{}' is {:?}, expected {:?} (batch pinned to 1)",
input.name(),
shape,
want
),
});
}
let out = session.outputs().first().ok_or(FaceError::WrongModel {
expected: "2d106det",
detail: "model has no outputs".into(),
})?;
let last: Option<i64> = out.dtype().tensor_shape().and_then(|d| d.last().copied());
if last != Some((POINTS * 2) as i64) {
return Err(FaceError::WrongModel {
expected: "2d106det",
detail: format!(
"output '{}' is {:?}-wide, expected {}",
out.name(),
last,
POINTS * 2
),
});
}
drop(session);
drop(acquired);
Ok(Self { session: model })
}
/// The landmarks of the face in `bbox` — `(x0, y0, x1, y1)` in source
/// pixels, the detector's box — read from the source.
///
/// `None` for a box with no area or a buffer that is not the size it
/// claims, as every crop here.
pub fn landmarks(
&mut self,
px: Pixels<'_>,
width: usize,
height: usize,
bbox: (f32, f32, f32, f32),
) -> Result<Option<Landmarks>, FaceError> {
let (w, h) = (bbox.2 - bbox.0, bbox.3 - bbox.1);
let side = w.max(h) * CROP_SCALE;
let (cx, cy) = ((bbox.0 + bbox.2) / 2.0, (bbox.1 + bbox.3) / 2.0);
let (x0, y0) = (cx - side / 2.0, cy - side / 2.0);
let Some(crop) = crop_box(
px,
width,
height,
(x0, y0, side, side),
INPUT_EDGE,
INPUT_EDGE,
) else {
return Ok(None);
};
let e = INPUT_EDGE;
let mut input = Array4::<f32>::zeros((1, 3, e, e));
for y in 0..e {
for x in 0..e {
for c in 0..3 {
input[[0, c, y, x]] = crop[(y * e + x) * 3 + c] * 255.0;
}
}
}
let acquired = self.session.acquire()?;
let mut session = acquired.lock();
let outputs = session
.run(ort::inputs![
ort::value::Tensor::from_array(input).map_err(FaceError::Inference)?
])
.map_err(FaceError::Inference)?;
let (_, data) = outputs[0]
.try_extract_tensor::<f32>()
.map_err(FaceError::Inference)?;
if data.len() < POINTS * 2 {
return Err(FaceError::WrongModel {
expected: "2d106det",
detail: format!("got {} values, expected {}", data.len(), POINTS * 2),
});
}
// −1..1 in the crop → crop pixels → source pixels.
let scale = side / e as f32;
let half = e as f32 / 2.0;
let mut points = [(0.0_f32, 0.0_f32); POINTS];
for (i, p) in points.iter_mut().enumerate() {
let (u, v) = ((data[2 * i] + 1.0) * half, (data[2 * i + 1] + 1.0) * half);
*p = (x0 + u * scale, y0 + v * scale);
}
Ok(Some(Landmarks { points }))
}
}
#[cfg(test)]
mod tests {
use super::*;
/// Packed and unpacked, every point comes back within a fifth of a
/// source pixel on a 6000-pixel frame — including one outside the
/// image, which a face at the edge does produce.
#[test]
fn dense_landmarks_round_trip_through_their_packed_bytes() {
let mut points = [(0.0_f32, 0.0_f32); POINTS];
for (i, p) in points.iter_mut().enumerate() {
*p = (i as f32 * 37.3 - 200.0, 5900.0 - i as f32 * 11.1);
}
let lm = Landmarks { points };
let bytes = lm.to_packed_bytes(6000.0);
assert_eq!(bytes.len(), PACKED_BYTES);
assert_eq!(PACKED_BYTES, 424);
let back = Landmarks::from_packed_bytes(&bytes, 6000.0).unwrap();
for (a, b) in lm.points.iter().zip(back.points.iter()) {
assert!((a.0 - b.0).abs() < 0.2, "{} vs {}", a.0, b.0);
assert!((a.1 - b.1).abs() < 0.2, "{} vs {}", a.1, b.1);
}
assert!(Landmarks::from_packed_bytes(&bytes[..100], 6000.0).is_none());
}
}
+36 -22
View File
@@ -1,8 +1,10 @@
//! Faces and identity (S14, docs/faces.md).
//! Faces and identity (S14, docs/dev/faces.md).
//!
//! Two models, run over the proxy tier, producing per face a box, five
//! Two models, run over the native render, producing per face a box, five
//! landmarks, a confidence and a 512-d embedding (FR-CULL-8) — and then the
//! arithmetic that turns embeddings into people (FR-CULL-9, FR-CULL-10).
//! arithmetic that turns embeddings into people (FR-CULL-9, FR-CULL-10). Two
//! more, optional, read each face's eyes and whether sunglasses hide them
//! (FR-CULL-8a, [`classify`] and [`eyes`]).
//!
//! Like `dr-segment`, this crate is **device-free**: no GPU adapter, no
//! Slint, nothing that needs a display. Unlike `dr-segment`, it carries **no
@@ -19,8 +21,9 @@
//! for a packaging script to switch on. The application obtains a model at
//! runtime; this crate takes bytes and never fetches anything.
//!
//! docs/faces.md §2 is the full reading, including what would have to change
//! for that to stop being true.
//! docs/dev/faces.md §2 is the full reading, including what would have to change
//! for that to stop being true. The eye-state models are the exception: MIT,
//! weights and all, and shipped in `models/face/` (docs/dev/faces.md §17).
//!
//! # Why the runtime is split behind a feature
//!
@@ -34,19 +37,25 @@
pub mod align;
pub mod assign;
pub mod calibrate;
#[cfg(feature = "inference")]
pub mod classify;
pub mod cluster;
#[cfg(feature = "inference")]
pub mod detect;
#[cfg(feature = "inference")]
pub mod embed;
pub mod embedding;
pub mod eyes;
#[cfg(feature = "inference")]
pub mod landmarks;
pub mod naming;
pub mod neighbours;
pub mod references;
/// Smallest long edge a face crop may be sampled from.
///
/// **A floor on the crop source, not on the detector input.** The distinction
/// is the whole of FR-CULL-8 and `docs/faces.md` §7: detection letterboxes
/// is the whole of FR-CULL-8 and `docs/dev/faces.md` §7: detection letterboxes
/// every buffer into 640×640, so its input resolution decides nothing, while
/// [`warp`] samples the 112×112 the embedder sees and so converts source
/// resolution directly into embedding quality. FR-CULL-8 requires that crop to
@@ -65,10 +74,13 @@ pub mod neighbours;
pub const MIN_CROP_EDGE: u32 = 1025;
pub use align::{
warp, warp_pixels, Aligned112, Pixels, Similarity, ALIGNED_EDGE, ARCFACE_TEMPLATE,
crop_box, eye_box, eye_patch, head_views, warp, warp_pixels, Aligned112, EyePatch, HeadViews,
Pixels, Similarity, ALIGNED_EDGE, ARCFACE_TEMPLATE,
};
pub use assign::{identity_shares, RIVAL_FLOOR, TOP_MATCHES};
pub use calibrate::{Calibration, Pairs, ReliabilityBand};
#[cfg(feature = "inference")]
pub use classify::{EyeClassifier, EyeModels, SunglassesClassifier};
pub use cluster::{
cluster, cluster_scored, split, Candidate, Cluster, Grouping, DEFAULT_MERGE_PROBABILITY,
};
@@ -79,6 +91,12 @@ pub use embed::{Embedded, Embedder};
pub use embedding::{
in_gallery, read_f16_bytes, Embedding, ModelId, EMBEDDING_DIM, MIN_GALLERY_QUALITY,
};
pub use eyes::{
Eye, EyeReading, EyeState, EYES_OPEN_THRESHOLD, HIDDEN_EYE_RATIO, MIN_EYE_PX,
MIN_EYE_SHARPNESS, SUNGLASSES_THRESHOLD,
};
#[cfg(feature = "inference")]
pub use landmarks::{Landmarker, Landmarks};
pub use naming::{name_for_instance, name_instances, NamedFace};
/// What can go wrong between an image and a face.
@@ -107,25 +125,21 @@ pub enum FaceError {
ImageShape { expected: usize, got: usize },
}
/// Install tract as `ort`'s backend.
///
/// Idempotent, and it must happen before any other `ort` call: with
/// `alternative-backend` there is no linked runtime to fall back on, so an
/// un-set API is a panic rather than a slow path. Same helper as
/// `dr-segment::semantic`, for the same reason.
#[cfg(feature = "inference")]
pub(crate) fn install_backend() {
use std::sync::Once;
static ONCE: Once = Once::new();
ONCE.call_once(|| {
let _ = ort::set_api(ort_tract::api());
});
impl From<dr_inference_engine::Error> for FaceError {
fn from(e: dr_inference_engine::Error) -> Self {
match e {
dr_inference_engine::Error::Inference(e) => FaceError::Inference(e),
dr_inference_engine::Error::Io(e) => FaceError::ModelRead(e),
}
}
}
/// [`install_backend`] for the M1 probe example, which drives `ort` directly
/// rather than through [`detect::Detector`] so it can report the raw error.
/// Make sure `ort` has a backend, for the M1 probe example, which drives
/// `ort` directly rather than through [`detect::Detector`] so it can report
/// the raw error. Every other path goes through `dr-inference-engine`.
#[cfg(feature = "inference")]
#[doc(hidden)]
pub fn install_backend_for_probe() {
install_backend();
dr_inference_engine::ensure_runtime();
}
+1 -1
View File
@@ -115,7 +115,7 @@ pub struct Faces<'a> {
/// Source pixels across the aligned crop, for the calibration's size term.
pub crop_px: &'a [f32],
/// Which photograph each face came from. Two faces in one frame are not
/// the same person, so those pairs are never returned (docs/faces.md §9).
/// the same person, so those pairs are never returned (docs/dev/faces.md §9).
pub images: &'a [u64],
/// Which faces may be compared *against* — the gallery
/// ([`crate::embedding::MIN_GALLERY_QUALITY`]).
+232
View File
@@ -0,0 +1,232 @@
//! TRACES: FR-CULL-10 | NFR-P9
//! Which of a person's faces stand for them in a grouping pass.
//!
//! # Why not all of them
//!
//! Every face the user has ruled on enters [`crate::cluster`] as an anchor,
//! and the pass compares every face against every other
//! ([`crate::neighbours`] is exhaustive by design). So a person with 750
//! confirmed faces costs 750 comparisons against each of the library's other
//! faces, and the cost of naming a library well grows with how well it is
//! named: a fully confirmed library of 25,000 faces spends almost the whole
//! scan re-comparing faces whose identity is already settled against each
//! other.
//!
//! Most of those comparisons say nothing new. A person's confirmed faces are
//! heavily redundant — thirty frames from one afternoon are one point of
//! view, not thirty — and a new face that matches one of them matches the
//! others too. What a new face needs to be measured against is the person's
//! *range*: the angles, ages and lights they have been photographed in, each
//! represented once.
//!
//! # The choice: the most diverse of the good ones
//!
//! Two rules, in order.
//!
//! **Good enough to vouch.** Only faces whose raw embedding was at least
//! [`MIN_REFERENCE_QUALITY`] long are eligible — a stricter floor than the
//! gallery's ([`crate::embedding::MIN_GALLERY_QUALITY`]), because a reference
//! is asked to speak *for* a person rather than merely be admitted to the
//! comparison. A face whose length was never recorded is admitted, as it is
//! everywhere else: a rule that cannot be checked admits rather than excludes.
//!
//! **As far apart as possible.** From the eligible pool, up to
//! [`MAX_REFERENCES`] faces are chosen to maximise the volume they span —
//! the determinant of their Gram matrix — greedily: start from the longest
//! vector, and at each step add the face with the largest component
//! orthogonal to everything chosen so far. That is Gram–Schmidt with a
//! pivot, and the product of the squared residuals it picks *is* the
//! determinant, so the greedy step is the exact greedy on the objective.
//! The effect is that a near-duplicate of a chosen face has almost no
//! residual and is passed over, while the one profile shot among two
//! hundred frontal frames is taken early.
//!
//! What is not chosen still belongs to the person. Those faces keep their
//! confirmations and are not touched by the pass; they are simply not
//! compared, which is the whole saving.
/// The most faces that stand for one person.
///
/// A hundred is far more points of view than a person has. What it bounds
/// is the cost: with every person at the cap, a scan against the named part
/// of a library is `people × 100` comparisons per face rather than
/// `confirmations`, and the two part company as soon as a library is used.
pub const MAX_REFERENCES: usize = 100;
/// The shortest raw embedding that may stand for a person.
///
/// One above the gallery floor: a reference vouches for someone, and the
/// margin keeps the faces that only just cleared the gallery — the ones
/// nearest the middle of the sphere — out of the set that speaks for a
/// person.
pub const MIN_REFERENCE_QUALITY: f32 = 15.0;
/// Whether a face of this quality may stand for a person.
///
/// `None` is "never measured" and is admitted, as in
/// [`crate::embedding::in_gallery`].
pub fn eligible(quality: Option<f32>) -> bool {
quality.is_none_or(|q| q >= MIN_REFERENCE_QUALITY)
}
/// Choose which of one person's faces stand for them.
///
/// `embeddings` and `quality` are one entry per face, the embeddings unit
/// length and all of one dimension. Returns the indices chosen, in the order
/// chosen — the first is the longest eligible vector, and each after it is
/// the one furthest from the span of those before. Every eligible face is
/// returned when there are `max` or fewer of them, so a person under the
/// cap loses nothing.
///
/// Deterministic: equal residuals break on the longer vector, then the lower
/// index, so two devices holding the same faces choose the same references
/// and group the same way (`cluster::clustering_is_deterministic`).
pub fn select(embeddings: &[&[f32]], quality: &[Option<f32>], max: usize) -> Vec<usize> {
debug_assert_eq!(embeddings.len(), quality.len());
let mut pool: Vec<usize> = (0..embeddings.len())
.filter(|&i| eligible(quality[i]))
.collect();
if pool.len() <= max {
return pool;
}
// Longest first, so the seed is the pool's front and a tie on residual
// resolves to the earlier position. A missing reading ranks below any
// measured one for this purpose only: it is admitted, but a face that
// was measured and found long is the better seed.
pool.sort_by(|&a, &b| {
let qa = quality[a].unwrap_or(0.0);
let qb = quality[b].unwrap_or(0.0);
qb.total_cmp(&qa).then(a.cmp(&b))
});
// Residuals: what remains of each pool vector outside the span of the
// chosen ones. Copied, since they are rewritten in place.
let mut residual: Vec<Vec<f32>> = pool.iter().map(|&i| embeddings[i].to_vec()).collect();
let mut taken = vec![false; pool.len()];
let mut chosen = Vec::with_capacity(max);
while chosen.len() < max {
// The face with the most left outside the span. The seed is the
// pool's front by construction: every unit vector has the same
// residual before anything is chosen, up to rounding, and rounding
// is not a reason to prefer one. After that `> best` and not `>=`,
// so a genuine tie keeps the earlier (longer) candidate.
let mut pick = None;
let mut best = 0.0_f32;
if chosen.is_empty() {
pick = Some(0);
best = residual[0].iter().map(|x| x * x).sum();
} else {
for (k, r) in residual.iter().enumerate() {
if taken[k] {
continue;
}
let n2: f32 = r.iter().map(|x| x * x).sum();
if n2 > best {
best = n2;
pick = Some(k);
}
}
}
// Nothing left outside the span: every remaining face is a
// combination of the chosen ones and adds no volume.
let Some(k) = pick.filter(|_| best > 1e-6) else {
break;
};
taken[k] = true;
chosen.push(pool[k]);
// Project the chosen direction out of every remaining residual.
let inv = best.sqrt().recip();
let q: Vec<f32> = residual[k].iter().map(|x| x * inv).collect();
for (j, r) in residual.iter_mut().enumerate() {
if taken[j] {
continue;
}
let d: f32 = r.iter().zip(&q).map(|(a, b)| a * b).sum();
for (x, y) in r.iter_mut().zip(&q) {
*x -= d * y;
}
}
}
chosen
}
#[cfg(test)]
mod tests {
use super::*;
fn unit(v: &[f32]) -> Vec<f32> {
let n = v.iter().map(|x| x * x).sum::<f32>().sqrt();
v.iter().map(|x| x / n).collect()
}
#[test]
fn a_person_under_the_cap_keeps_every_eligible_face() {
let e = [unit(&[1.0, 0.0]), unit(&[0.0, 1.0]), unit(&[1.0, 1.0])];
let refs: Vec<&[f32]> = e.iter().map(Vec::as_slice).collect();
let q = [Some(20.0), None, Some(16.0)];
assert_eq!(select(&refs, &q, 100), vec![0, 1, 2]);
}
#[test]
fn a_short_vector_never_stands_for_a_person() {
let e = [unit(&[1.0, 0.0]), unit(&[0.0, 1.0])];
let refs: Vec<&[f32]> = e.iter().map(Vec::as_slice).collect();
let q = [Some(20.0), Some(MIN_REFERENCE_QUALITY - 0.01)];
assert_eq!(select(&refs, &q, 100), vec![0]);
}
/// Two hundred frames from one afternoon and one profile shot: the
/// profile is the second choice, not the two-hundred-and-first.
#[test]
fn the_odd_one_out_is_chosen_before_any_duplicate() {
let mut e: Vec<Vec<f32>> = Vec::new();
let mut q = Vec::new();
for i in 0..200 {
// Near-duplicates of one direction, with a little noise.
let t = (i as f32) * 1e-3;
e.push(unit(&[1.0, t, t * 0.5]));
q.push(Some(20.0 + (i % 7) as f32));
}
e.push(unit(&[0.0, 0.0, 1.0]));
q.push(Some(16.0));
let refs: Vec<&[f32]> = e.iter().map(Vec::as_slice).collect();
let chosen = select(&refs, &q, 3);
assert_eq!(chosen.len(), 3);
assert_eq!(
chosen[1], 200,
"the profile shot was not second: {chosen:?}"
);
// Seeded on the longest vector.
assert_eq!(q[chosen[0]], Some(26.0));
}
/// Faces inside the span of the chosen ones add no volume and are not
/// taken to fill the cap.
#[test]
fn the_cap_is_not_filled_from_inside_the_span() {
let e = [
unit(&[1.0, 0.0]),
unit(&[0.0, 1.0]),
unit(&[1.0, 1.0]),
unit(&[2.0, -1.0]),
];
let refs: Vec<&[f32]> = e.iter().map(Vec::as_slice).collect();
let q = [Some(20.0); 4];
assert_eq!(select(&refs, &q, 3).len(), 2);
}
#[test]
fn the_choice_is_deterministic() {
let e: Vec<Vec<f32>> = (0..50)
.map(|i| {
let a = (i as f32) * 0.37;
unit(&[a.cos(), a.sin(), (a * 3.0).sin(), 0.2])
})
.collect();
let refs: Vec<&[f32]> = e.iter().map(Vec::as_slice).collect();
let q = vec![Some(18.0); 50];
assert_eq!(select(&refs, &q, 5), select(&refs, &q, 5));
}
}
+5
View File
@@ -7,6 +7,11 @@ license.workspace = true
[dependencies]
dr-types.workspace = true
# The merge's geometry (FR-MRG-10): rotations, the focal length and the
# projections, solved on proxies by dr-pano and consumed here per chunk. The
# geometry alone — no keypoint model, no runtime — which is what the
# workspace entry turns off.
dr-pano.workspace = true
dr-decode.workspace = true
dr-pipeline.workspace = true
# The watershed's pixel passes are here because they are shaders; everything
+2 -2
View File
@@ -1,6 +1,6 @@
//! What a frame actually costs — the measurement FR-DSP-2 is waiting on.
//!
//! `docs/display-and-extension.md` §2 argues that tiled computation predates
//! `docs/dev/display-and-extension.md` §2 argues that tiled computation predates
//! the fused-shader design and may not need to exist: the composer folds every
//! active operation into **one dispatch over a viewport-sized target**, so the
//! problem tiles were invented to solve may already be solved. That argument
@@ -28,7 +28,7 @@
//! the per-frame CPU half is dominated by shader-source assembly, which is
//! string formatting and is several times slower unoptimised.
//!
//! The committed numbers live in `docs/frame-budget.md`. Rerun this and diff
//! The committed numbers live in `docs/dev/frame-budget.md`. Rerun this and diff
//! that file; a regression should be a diff rather than somebody's memory.
//!
//! # Why the 99th percentile and not the mean
+1 -1
View File
@@ -1,6 +1,6 @@
//! Segment an image and write the granularity ladder as false-coloured PPMs.
//!
//! The whole point of S15 step 2 (docs/segmentation.md §11): look at the
//! The whole point of S15 step 2 (docs/dev/segmentation.md §11): look at the
//! ladder and decide whether clicking through it would land on the things a
//! person means. No amount of design settles that — the pictures do.
//!
+289 -7
View File
@@ -101,6 +101,16 @@ pub struct AdjustPass {
/// switched on does not build a pipeline layout mid-frame.
linear_bind_group_layout: wgpu::BindGroupLayout,
linear_pipeline_layout: wgpu::PipelineLayout,
/// TRACES: FR-MRG-2
/// A third layout, writing `rgba32float`, for the camera-space tap a
/// merge reads (`OutputMode::CameraLinear`). Same reasoning as the
/// linear one: the format is in the layout, so a format is a layout.
camera_bind_group_layout: wgpu::BindGroupLayout,
camera_pipeline_layout: wgpu::PipelineLayout,
/// The camera-space texture the last `render_camera_linear` wrote.
/// Separate from `targets`: a different format, and a merge reads it
/// back or samples it while the display targets go on being swapped.
camera_target: Option<Target>,
/// TRACES: FR-DEV-3d
/// What the linear intermediate currently holds, and at what size.
///
@@ -303,6 +313,11 @@ impl AdjustPass {
pub const FORMAT: wgpu::TextureFormat = wgpu::TextureFormat::Rgba8Unorm;
/// TRACES: FR-MRG-2
/// The camera-space tap's format: full precision, because what it holds
/// is written back as a RAW at the sensor's own scale (FR-MRG-3).
pub const CAMERA_FORMAT: wgpu::TextureFormat = wgpu::TextureFormat::Rgba32Float;
pub fn new(ctx: &GpuContext) -> Self {
let bind_group_layout = Self::layout_writing(ctx, Self::FORMAT, "adjust-bgl");
@@ -330,6 +345,15 @@ impl AdjustPass {
bind_group_layouts: &[Some(&linear_bind_group_layout)],
immediate_size: 0,
});
let camera_bind_group_layout =
Self::layout_writing(ctx, Self::CAMERA_FORMAT, "adjust-camera-bgl");
let camera_pipeline_layout =
ctx.device
.create_pipeline_layout(&wgpu::PipelineLayoutDescriptor {
label: Some("adjust-camera-layout"),
bind_group_layouts: &[Some(&camera_bind_group_layout)],
immediate_size: 0,
});
// A 1x1 single-layer mask, bound when the edit has no local
// adjustments. The generated shader never samples it — no layer block
@@ -412,6 +436,9 @@ impl AdjustPass {
detail: DetailRunner::new(ctx),
linear_bind_group_layout,
linear_pipeline_layout,
camera_bind_group_layout,
camera_pipeline_layout,
camera_target: None,
colour_key: None,
colour_dispatches: 0,
detail_dispatches: 0,
@@ -549,6 +576,7 @@ impl AdjustPass {
let layout = match shader.output_mode {
OutputMode::Encoded => &self.pipeline_layout,
OutputMode::LinearWorking => &self.linear_pipeline_layout,
OutputMode::CameraLinear => &self.camera_pipeline_layout,
};
let pipeline =
@@ -1126,28 +1154,206 @@ impl AdjustPass {
self.copy_output()
}
/// TRACES: FR-MRG-2
/// Render the camera-space tap: the source after its lens warp and
/// nothing else, at full precision.
///
/// `shader` must come from `EditGraph::compose_camera_linear` — it is
/// refused otherwise, for the reason `render_masked` refuses a linear
/// one: the storage format is in the layout. The profile uniforms are
/// filled neutral here rather than from the source, which is the whole
/// point of the mode (`OutputMode::CameraLinear`): unit white balance,
/// identity matrix, base curve off. The non-linear flag is kept, so a
/// JPEG source is still linearised — camera space for a JPEG is the
/// decoded values made linear, which is the best that exists.
///
/// The texture stays on the device for a merge's warp to sample; see
/// [`Self::camera_texture`] and [`Self::read_camera_linear`].
pub fn render_camera_linear(
&mut self,
source: &DemosaicedImage,
shader: &ComposedShader,
width: u32,
height: u32,
) -> Result<&wgpu::Texture, GpuError> {
if shader.output_mode != OutputMode::CameraLinear {
return Err(GpuError::ShaderCompilation(
"render_camera_linear takes the shader from EditGraph::compose_camera_linear \
and no other; this one writes a different format"
.into(),
));
}
self.colour_key = None;
let (width, height) = (width.max(1), height.max(1));
self.ensure_camera_target(width, height);
let mut uniforms = Self::fused_uniforms(source, shader);
// Neutral profile: the numbers the sensor produced, and only those.
let non_linear = uniforms[15];
uniforms[0..4].copy_from_slice(&[1.0, 0.0, 0.0, 0.0]);
uniforms[4..8].copy_from_slice(&[0.0, 1.0, 0.0, 0.0]);
uniforms[8..12].copy_from_slice(&[0.0, 0.0, 1.0, 0.0]);
uniforms[12..16].copy_from_slice(&[1.0, 1.0, 1.0, non_linear]);
let b = dr_pipeline::BASE_CURVE_UNIFORM_OFFSET;
uniforms[b + 10] = 0.0;
let params_buf = self
.ctx
.device
.create_buffer_init(&wgpu::util::BufferInitDescriptor {
label: Some("adjust-camera-params"),
contents: bytemuck::cast_slice(&uniforms),
usage: wgpu::BufferUsages::UNIFORM,
});
let _ = self.pipeline(shader)?;
let pipeline = self
.cache
.get(&shader.structure_hash)
.expect("compiled above");
let target = self.camera_target.as_ref().expect("ensured above");
let bind_group = self
.ctx
.device
.create_bind_group(&wgpu::BindGroupDescriptor {
label: Some("adjust-camera-bg"),
layout: &self.camera_bind_group_layout,
entries: &[
wgpu::BindGroupEntry {
binding: 0,
resource: wgpu::BindingResource::TextureView(source.view()),
},
wgpu::BindGroupEntry {
binding: 1,
resource: params_buf.as_entire_binding(),
},
wgpu::BindGroupEntry {
binding: 2,
resource: wgpu::BindingResource::TextureView(&target.view),
},
wgpu::BindGroupEntry {
binding: 3,
resource: wgpu::BindingResource::TextureView(&self.empty_masks),
},
wgpu::BindGroupEntry {
binding: 4,
resource: wgpu::BindingResource::TextureView(self.film_curves_view()),
},
wgpu::BindGroupEntry {
binding: 5,
resource: wgpu::BindingResource::TextureView(self.film_lut_view()),
},
],
});
let mut enc = self
.ctx
.device
.create_command_encoder(&wgpu::CommandEncoderDescriptor {
label: Some("adjust-camera-encoder"),
});
{
let mut pass = enc.begin_compute_pass(&wgpu::ComputePassDescriptor {
label: Some("adjust-camera-pass"),
timestamp_writes: None,
});
pass.set_pipeline(pipeline);
pass.set_bind_group(0, &bind_group, &[]);
pass.dispatch_workgroups(width.div_ceil(8), height.div_ceil(8), 1);
}
self.ctx.queue.submit(Some(enc.finish()));
self.colour_dispatches += 1;
Ok(&self.camera_target.as_ref().expect("ensured above").texture)
}
/// The camera-space texture, if one has been rendered.
pub fn camera_texture(&self) -> Option<&wgpu::Texture> {
self.camera_target.as_ref().map(|t| &t.texture)
}
/// TRACES: FR-MRG-2
/// Read the camera-space tap back: tightly packed RGBA `f32`,
/// `width * height * 4` values, alpha 1.0 everywhere.
pub fn read_camera_linear(&self) -> Result<(Vec<f32>, u32, u32), GpuError> {
let Some(target) = self.camera_target.as_ref() else {
return Err(GpuError::Readback("no camera-space render yet".into()));
};
let (bytes, w, h) = Self::copy_texture(&self.ctx, &target.texture, w_h(target), 16)?;
let floats: Vec<f32> = bytes
.chunks_exact(4)
.map(|b| f32::from_le_bytes([b[0], b[1], b[2], b[3]]))
.collect();
Ok((floats, w, h))
}
fn ensure_camera_target(&mut self, width: u32, height: u32) {
if self
.camera_target
.as_ref()
.is_some_and(|t| t.width == width && t.height == height)
{
return;
}
let texture = self.ctx.device.create_texture(&wgpu::TextureDescriptor {
label: Some("adjust-camera-output"),
size: wgpu::Extent3d {
width,
height,
depth_or_array_layers: 1,
},
mip_level_count: 1,
sample_count: 1,
dimension: wgpu::TextureDimension::D2,
format: Self::CAMERA_FORMAT,
// Written by compute, sampled by a merge's warp, copied out for
// the CPU. Never handed to the compositor, so no RENDER_ATTACHMENT.
usage: wgpu::TextureUsages::STORAGE_BINDING
| wgpu::TextureUsages::TEXTURE_BINDING
| wgpu::TextureUsages::COPY_SRC,
view_formats: &[],
});
let view = texture.create_view(&Default::default());
self.camera_target = Some(Target {
texture,
view,
width,
height,
});
}
/// The transfer itself.
fn copy_output(&self) -> Result<(Vec<u8>, u32, u32), GpuError> {
let Some(target) = self.targets[self.current].as_ref() else {
return Err(GpuError::Readback("nothing rendered yet".into()));
};
let (w, h) = (target.width, target.height);
Self::copy_texture(&self.ctx, &target.texture, w_h(target), 4)
}
let unpadded = w * 4;
/// Copy a whole texture to the CPU, `bytes_per_pixel` wide, rows
/// unpadded. Shared by the display readback and the camera-space one.
fn copy_texture(
ctx: &GpuContext,
texture: &wgpu::Texture,
(w, h): (u32, u32),
bytes_per_pixel: u32,
) -> Result<(Vec<u8>, u32, u32), GpuError> {
let unpadded = w * bytes_per_pixel;
let align = wgpu::COPY_BYTES_PER_ROW_ALIGNMENT;
let padded = unpadded.div_ceil(align) * align;
let buf = self.ctx.device.create_buffer(&wgpu::BufferDescriptor {
let buf = ctx.device.create_buffer(&wgpu::BufferDescriptor {
label: Some("adjust-readback"),
size: (padded * h) as u64,
usage: wgpu::BufferUsages::COPY_DST | wgpu::BufferUsages::MAP_READ,
mapped_at_creation: false,
});
let mut enc = self.ctx.device.create_command_encoder(&Default::default());
let mut enc = ctx.device.create_command_encoder(&Default::default());
enc.copy_texture_to_buffer(
wgpu::TexelCopyTextureInfo {
texture: &target.texture,
texture,
mip_level: 0,
origin: wgpu::Origin3d::ZERO,
aspect: wgpu::TextureAspect::All,
@@ -1166,7 +1372,7 @@ impl AdjustPass {
depth_or_array_layers: 1,
},
);
self.ctx.queue.submit(Some(enc.finish()));
ctx.queue.submit(Some(enc.finish()));
let slice = buf.slice(..);
let (tx, rx) = std::sync::mpsc::channel();
@@ -1177,7 +1383,7 @@ impl AdjustPass {
// Polled rather than parked, and bounded rather than spun forever —
// see `readback::await_mapping`, which the histogram's own transfer
// shares for exactly the same reasons.
await_mapping(&self.ctx, &rx)?;
await_mapping(ctx, &rx)?;
let data = slice.get_mapped_range();
let mut out = Vec::with_capacity((unpadded * h) as usize);
@@ -1191,6 +1397,10 @@ impl AdjustPass {
}
}
fn w_h(t: &Target) -> (u32, u32) {
(t.width, t.height)
}
/// Number the lines of generated source, so a compiler error can be located.
pub(crate) fn numbered(src: &str) -> String {
src.lines()
@@ -1242,6 +1452,10 @@ mod tests {
// rather than about a camera's colour response.
color_matrix: Some([1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0]),
base_curve: BaseCurve::IDENTITY,
samples_per_pixel: 1,
profile: None,
make: String::new(),
model: String::new(),
crop: CropRect {
x: 0,
y: 0,
@@ -1442,6 +1656,10 @@ mod tests {
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
color_matrix: Some([1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0]),
base_curve: BaseCurve::IDENTITY,
samples_per_pixel: 1,
profile: None,
make: String::new(),
model: String::new(),
crop: CropRect {
x: 0,
y: 0,
@@ -1680,6 +1898,62 @@ mod tests {
);
}
/// TRACES: FR-DEV-20
#[test]
fn a_keystone_reshapes_the_frame_without_exposing_a_corner() {
// The shader half of perspective correction, end to end. A top-bright
// frame with a full vertical keystone spreads its top across the
// output, so the bright half reaches further down than the middle;
// and since the frame is mapped onto a trapezoid *inside* the source,
// no corner is left without a pixel behind it.
let Some(ctx) = ctx() else { return };
let mut pass = AdjustPass::new(&ctx);
let img = split_image(&ctx, true);
let plain = EditGraph::default_chain().compose();
let tex = pass.render(&img, &plain, 32, 32).expect("render");
let below_middle = read_pixel(&ctx, tex, 16, 19)[0];
assert!(
below_middle < 90,
"unkeyed, row 19 is the dark half: {below_middle}"
);
let mut g = EditGraph::default_chain();
g.set_param(
dr_pipeline::framing::ID,
dr_pipeline::framing::KEYSTONE_V,
100.0,
);
g.set_param(
dr_pipeline::framing::ID,
dr_pipeline::framing::KEYSTONE_H,
100.0,
);
let shader = g.compose();
let tex = pass.render(&img, &shader, 32, 32).expect("render");
for (x, y) in [(0, 0), (31, 0), (0, 31), (31, 31)] {
assert_ne!(
read_pixel(&ctx, tex, x, y),
[0, 0, 0, 255],
"corner ({x},{y}) has no source pixel behind it"
);
}
let mut g = EditGraph::default_chain();
g.set_param(
dr_pipeline::framing::ID,
dr_pipeline::framing::KEYSTONE_V,
100.0,
);
let tex = pass.render(&img, &g.compose(), 32, 32).expect("render");
let keyed = read_pixel(&ctx, tex, 16, 19)[0];
assert!(
keyed > 128,
"the spread top half must reach row 19: {keyed}"
);
}
#[test]
fn dragging_the_crop_does_not_recompile() {
// The cache contract for framing, which is what makes an interactive
@@ -1969,6 +2243,10 @@ mod tests {
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
color_matrix: Some([1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0]),
base_curve: BaseCurve::IDENTITY,
samples_per_pixel: 1,
profile: None,
make: String::new(),
model: String::new(),
crop: CropRect {
x: 0,
y: 0,
@@ -2069,6 +2347,10 @@ mod tests {
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
color_matrix: Some([1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0]),
base_curve: BaseCurve::IDENTITY,
samples_per_pixel: 1,
profile: None,
make: String::new(),
model: String::new(),
crop: CropRect {
x: 0,
y: 0,
+232
View File
@@ -250,6 +250,135 @@ impl DemosaicedImage {
}
}
impl DemosaicedImage {
/// TRACES: FR-MRG-3
/// A source that is already RGB in camera space: a linear DNG, which is
/// what a merge writes. No demosaic; the samples are normalised by the
/// file's black and white levels exactly as the demosaic kernel would
/// normalise a photosite, and everything else — the matrix, the
/// balance, the body's base curve — is carried through as for a CFA
/// file, because the composite is developed as one photograph from the
/// body that took its sources.
pub fn from_linear_rgb16(ctx: &GpuContext, raw: &RawImage) -> Result<Self, GpuError> {
let (width, height) = (raw.crop.width.max(1), raw.crop.height.max(1));
let limits = ctx.device.limits();
if width > limits.max_texture_dimension_2d || height > limits.max_texture_dimension_2d {
return Err(GpuError::TooLarge(format!(
"{width}×{height} exceeds the device limit of {}",
limits.max_texture_dimension_2d
)));
}
let stride = raw.width as usize * 3;
let expected = raw.height as usize * stride;
if raw.data.len() < expected {
return Err(GpuError::TooLarge(format!(
"{} samples is short of the {expected} a {}×{} RGB image needs",
raw.data.len(),
raw.width,
raw.height
)));
}
let black = black_per_cell(raw);
let inv = inv_range_per_cell(raw);
// Per channel rather than per CFA cell: R, G, B are the first three.
let mut half: Vec<u16> = Vec::with_capacity((width * height * 4) as usize);
for y in 0..height as usize {
let row = (raw.crop.y as usize + y) * stride + raw.crop.x as usize * 3;
for x in 0..width as usize {
let p = &raw.data[row + x * 3..row + x * 3 + 3];
for c in 0..3 {
let v = (f32::from(p[c]) - black[c]) * inv[c];
half.push(f32_to_f16_bits_unclamped(v));
}
half.push(f32_to_f16_bits(1.0));
}
}
let texture = ctx.device.create_texture_with_data(
&ctx.queue,
&wgpu::TextureDescriptor {
label: Some("linear-rgb-source"),
size: wgpu::Extent3d {
width,
height,
depth_or_array_layers: 1,
},
mip_level_count: 1,
sample_count: 1,
dimension: wgpu::TextureDimension::D2,
format: Self::FORMAT,
usage: wgpu::TextureUsages::TEXTURE_BINDING | wgpu::TextureUsages::COPY_SRC,
view_formats: &[],
},
wgpu::util::TextureDataOrder::LayerMajor,
bytemuck::cast_slice(&half),
);
let view = texture.create_view(&Default::default());
Ok(Self {
texture,
view,
width,
height,
color_matrix: raw.color_matrix.unwrap_or(IDENTITY_3X3),
as_shot_wb: [raw.wb_coeffs[0], raw.wb_coeffs[1], raw.wb_coeffs[2]],
base_curve: raw.base_curve,
non_linear: false,
})
}
}
/// Convert an f32 to half-precision bits, the general case: sign,
/// subnormals, round-to-nearest-even, saturation at the largest finite.
///
/// `f32_to_f16_bits` below is the 8-bit special case and says why it can
/// be; this one exists because a linear DNG is not that case. A 14-bit
/// sensor's least significant step, normalised, is 6.1e-5 — right at f16's
/// smallest normal (6.1e-5) — so the deepest shadows of a composite land
/// in the subnormal range, and rounding them to zero would crush the
/// shadows of exactly the file that was written to keep them. Values below
/// zero (black subtraction on a noisy photosite) and above one (a highlight
/// past the white level) are legitimate and kept.
fn f32_to_f16_bits_unclamped(v: f32) -> u16 {
let bits = v.to_bits();
let sign = ((bits >> 16) & 0x8000) as u16;
let exp = ((bits >> 23) & 0xFF) as i32;
let mant = bits & 0x7F_FFFF;
if exp == 0xFF {
// Infinity or NaN: a NaN sample is a decode fault; store the largest
// finite rather than propagate it through a blend.
return sign | 0x7BFF;
}
let e = exp - 127 + 15;
if e >= 0x1F {
return sign | 0x7BFF;
}
if e <= 0 {
// Subnormal in f16 (or underflow). Shift the full mantissa with its
// implicit bit right by the deficit, rounding to nearest even.
if e < -10 {
return sign;
}
let m = (mant | 0x80_0000) >> (1 - e);
let shift = 13;
let rounded = round_shift(m, shift);
return sign | rounded as u16;
}
let rounded = round_shift(mant, 13);
// Rounding can carry into the exponent; that is correct.
sign | (((e as u32) << 10) + rounded) as u16
}
/// `v >> shift`, rounded to nearest with ties to even.
fn round_shift(v: u32, shift: u32) -> u32 {
let half = 1u32 << (shift - 1);
let mask = (1u32 << shift) - 1;
let low = v & mask;
let mut out = v >> shift;
if low > half || (low == half && (out & 1) == 1) {
out += 1;
}
out
}
/// Convert an f32 to IEEE 754 half-precision bits.
///
/// Written out rather than pulled in as a dependency: the inputs here are
@@ -389,6 +518,9 @@ impl Demosaicer {
/// `RawImage`; which of the two CFA families it came off is this
/// function's problem, not theirs.
pub fn run(&self, raw: &RawImage) -> Result<DemosaicedImage, GpuError> {
if raw.samples_per_pixel == 3 {
return DemosaicedImage::from_linear_rgb16(&self.ctx, raw);
}
let (width, height) = (raw.crop.width.max(1), raw.crop.height.max(1));
let limits = self.ctx.device.limits();
@@ -827,6 +959,43 @@ mod tests {
use super::*;
use dr_decode::CropRect;
fn f16_to_f32(bits: u16) -> f32 {
let sign = if bits & 0x8000 != 0 { -1.0 } else { 1.0 };
let e = ((bits >> 10) & 0x1F) as i32;
let m = (bits & 0x3FF) as f32;
if e == 0 {
sign * m * 2f32.powi(-24)
} else {
sign * (1.0 + m / 1024.0) * 2f32.powi(e - 15)
}
}
#[test]
fn unclamped_half_keeps_shadows_signs_and_highlights() {
// A 14-bit LSB, normalised: subnormal in f16, and must not be zero.
let lsb = 1.0 / 16383.0;
let back = f16_to_f32(f32_to_f16_bits_unclamped(lsb));
assert!((back - lsb).abs() / lsb < 0.01, "{back} vs {lsb}");
// A quarter of that, still representable.
let tiny = lsb / 4.0;
let back = f16_to_f32(f32_to_f16_bits_unclamped(tiny));
assert!((back - tiny).abs() / tiny < 0.05, "{back} vs {tiny}");
// Below zero and above one survive.
assert!((f16_to_f32(f32_to_f16_bits_unclamped(-0.01)) + 0.01).abs() < 1e-5);
assert!((f16_to_f32(f32_to_f16_bits_unclamped(1.75)) - 1.75).abs() < 1e-3);
// Exact values are exact.
assert_eq!(f32_to_f16_bits_unclamped(1.0), 0x3C00);
assert_eq!(f32_to_f16_bits_unclamped(0.5), 0x3800);
assert_eq!(f32_to_f16_bits_unclamped(0.0), 0);
// Within one ULP of the clamped one on its domain: that one
// truncates the mantissa, this one rounds it.
for i in 0..=255 {
let v = i as f32 / 255.0;
let (a, b) = (f32_to_f16_bits_unclamped(v), f32_to_f16_bits(v));
assert!(a.abs_diff(b) <= 1, "{v}: {a} vs {b}");
}
}
fn raw_for(black: [u16; 4], white: u16) -> RawImage {
RawImage {
width: 4,
@@ -838,6 +1007,10 @@ mod tests {
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
color_matrix: None,
base_curve: BaseCurve::IDENTITY,
samples_per_pixel: 1,
profile: None,
make: String::new(),
model: String::new(),
crop: CropRect {
x: 0,
y: 0,
@@ -950,6 +1123,10 @@ mod tests {
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
color_matrix: None,
base_curve: BaseCurve::IDENTITY,
samples_per_pixel: 1,
profile: None,
make: String::new(),
model: String::new(),
crop: CropRect {
x: 0,
y: 0,
@@ -1170,6 +1347,53 @@ mod tests {
}
}
#[test]
fn a_grey_step_edge_stays_grey() {
// A flat patch cannot tell the Malvar kernels from any other set of
// weights that sum to zero. An edge can. A grey vertical step, so
// every photosite records the same profile, must come back with the
// three channels close together on both sides; any spread is false
// colour from interpolating across the edge.
//
// The bound is set by the paper's kernels, which peak at 0.19 here.
// With the ±2 terms of the green-site kernels transposed — the bug
// this test was written against — the peak is 0.375.
let Some(ctx) = ctx() else { return };
let d = Demosaicer::new(&ctx).expect("demosaicer");
let size = 32u32;
let white = 16383u16;
let mut raw = flat_cfa(CfaPattern::Rggb, size, [0, 0, 0], 0, white);
for y in 0..size {
for x in size / 2..size {
raw.data[(y * size + x) as usize] = white;
}
}
let img = d.run(&raw).expect("demosaic");
let px = read_rgba(&ctx, &img);
let (w, _) = img.size();
let mut worst = (0.0f32, 0u32, 0u32);
for y in 2..size - 2 {
for x in 2..size - 2 {
let p = px[(y * w + x) as usize];
let spread = (p[0] - p[1]).abs().max((p[2] - p[1]).abs());
if spread > worst.0 {
worst = (spread, x, y);
}
}
}
assert!(
worst.0 < 0.25,
"false colour of {} at ({}, {}) on a grey edge — the green-site \
kernels are interpolating across the edge",
worst.0,
worst.1,
worst.2
);
}
#[test]
fn output_is_free_of_nan_and_negatives() {
// f16 NaN propagates silently through every later stage; a negative
@@ -1196,6 +1420,10 @@ mod tests {
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
color_matrix: None,
base_curve: BaseCurve::IDENTITY,
samples_per_pixel: 1,
profile: None,
make: String::new(),
model: String::new(),
crop: CropRect {
x: 0,
y: 0,
@@ -1277,6 +1505,10 @@ mod tests {
],
color_matrix: None,
base_curve: BaseCurve::IDENTITY,
samples_per_pixel: 1,
profile: None,
make: String::new(),
model: String::new(),
crop: CropRect {
x: 0,
y: 0,
+6 -13
View File
@@ -470,7 +470,7 @@ impl FocusPeakPass {
// TEXTURE_BINDING to be sampled by the compositor.
// RENDER_ATTACHMENT is not used by anything here and is required
// anyway: Slint rejects an imported texture without it. COPY_SRC
// is for `read_overlay` and its two callers.
// is for `read_overlay` and the tests that call it.
usage: wgpu::TextureUsages::STORAGE_BINDING
| wgpu::TextureUsages::TEXTURE_BINDING
| wgpu::TextureUsages::RENDER_ATTACHMENT
@@ -490,18 +490,11 @@ impl FocusPeakPass {
/// TRACES: AC-8
/// Copy the overlay to the CPU, as RGBA8 rows with no padding.
///
/// **Two callers, and neither is the desktop display path.** The tests
/// below are one: an overlay is a claim about which pixels are sharp, and
/// there is no way to check that claim without looking at the pixels. The
/// other is the Android develop view, which reads the *frame* back for the
/// reasons `technical-debt.md` TD-1 records — wgpu's Android swapchain
/// tears a portrait window, so Slint is not drawing with wgpu there and no
/// texture can be handed over. An overlay that stayed on the device on a
/// platform where the picture underneath it does not would simply never be
/// seen.
///
/// On desktop nothing calls this, and ARCH §6.1 holds on the path that
/// matters: the overlay reaches the compositor as a texture.
/// **The tests below are the only caller, and never the display path.** An
/// overlay is a claim about which pixels are sharp, and there is no way to
/// check that claim without looking at the pixels. On screen, on desktop
/// and Android alike, the overlay reaches the compositor as a texture and
/// ARCH §6.1 holds. (Android read it back here until TD-1 was paid off.)
pub fn read_overlay(&self) -> Result<(Vec<u8>, u32, u32), GpuError> {
let Some(layer) = self.layers[self.current].as_ref() else {
return Err(GpuError::Readback("no overlay has been rendered".into()));
+4 -1
View File
@@ -6,7 +6,7 @@
//! module doc said for eight releases that it held no pipeline and no masks.
//! It holds both now, plus demosaic, detail, segmentation masks, two
//! histograms and focus peaking. The zero-copy claim is still the one that
//! matters, and TD-1 records the one platform where it does not hold.
//! matters, and since TD-1 was paid off it holds on Android too.
//!
//! Deliberately free of UI dependencies (ARCH §6.5a). The texture is handed
//! out as a `wgpu::Texture`; who composites it is not this crate's concern.
@@ -28,6 +28,7 @@ mod error;
mod focus;
mod histogram;
mod mask;
mod merge;
mod raw_histogram;
mod readback;
mod segment;
@@ -40,6 +41,7 @@ pub use demosaic::{DemosaicedImage, Demosaicer};
pub use detail::INTERMEDIATE_FORMAT as DETAIL_INTERMEDIATE_FORMAT;
pub use error::GpuError;
pub use focus::{FocusPeakPass, FocusPeaking, PeakColour, PeakSensitivity};
pub use merge::{Band, MergeFrame, MergeOutput, MergePass};
// Renamed on the way out: `BINS` says enough inside `histogram`, and nothing
// at all at a crate root shared with demosaic and segmentation.
pub use histogram::{Histogram, HistogramPass, BINS as HISTOGRAM_BINS};
@@ -281,6 +283,7 @@ impl GpuContext {
)))
}
/// TRACES: NFR-COMPAT-1
/// Ask one adapter for a device, with the limits the pipeline needs.
async fn device_from(adapter: &wgpu::Adapter) -> Result<(wgpu::Device, wgpu::Queue), GpuError> {
adapter
+13
View File
@@ -395,6 +395,7 @@ pub struct MaskPass {
combine_layout: wgpu::BindGroupLayout,
combine_union: wgpu::RenderPipeline,
combine_subtract: wgpu::RenderPipeline,
combine_intersect: wgpu::RenderPipeline,
/// Where a part is drawn before it is joined.
///
/// One texture for the whole stack rather than one per layer, because
@@ -651,6 +652,16 @@ impl MaskPass {
"mask-combine-subtract",
blend_state(wgpu::BlendFactor::Zero, wgpu::BlendFactor::OneMinusSrc),
);
// TRACES: FR-DEV-19a
// `dst * src`: what the mask had, kept only in proportion to how much
// of it this part also covers. The same three vertices and the same
// scratch, so a third set operation is a third blend state and
// nothing more — which is what `Join::apply` states on the CPU and
// `the_joins_match_their_definition` holds this to.
let combine_intersect = combine(
"mask-combine-intersect",
blend_state(wgpu::BlendFactor::Zero, wgpu::BlendFactor::Src),
);
// The same, with the deposit thrown away: coverage is only ever taken
// off what earlier strokes on this layer put down. There is no negative
@@ -681,6 +692,7 @@ impl MaskPass {
combine_layout,
combine_union,
combine_subtract,
combine_intersect,
scratch: None,
array: None,
allocations: 0,
@@ -1376,6 +1388,7 @@ impl MaskPass {
pass.set_pipeline(match join {
Join::Union => &self.combine_union,
Join::Subtract => &self.combine_subtract,
Join::Intersect => &self.combine_intersect,
});
pass.set_bind_group(0, &bind_group, &[]);
pass.draw(0..3, 0..1);
+547
View File
@@ -0,0 +1,547 @@
//! TRACES: FR-MRG-10 | FR-MRG-11
//! The merge: source frames warped into an output surface, chunk by chunk.
//!
//! The per-pixel half of a panorama (FR-MRG-10), on the GPU: the warp of a
//! source tile into an output chunk, the weighted accumulation across
//! frames, and the resolve to sixteen-bit samples. The geometry it is
//! given — rotations, focal length, projection — is `dr-pano`'s, solved on
//! proxies before any full-resolution pixel exists (panorama.md §5), and
//! that is what makes this simple: every output pixel's source coordinates
//! are a closed-form function, so a chunk can be produced from the source
//! tiles that project into it and nothing else.
//!
//! # The loop
//!
//! ```text
//! for each band of rows of the output:
//! for each chunk across the band:
//! zero the accumulator
//! for each frame whose footprint meets the chunk:
//! the source rectangle the chunk needs, from the geometry
//! render it camera-linear through the pipeline (the tile)
//! warp the tile into the chunk, accumulate ← GPU
//! resolve the chunk to u16 ← GPU
//! copy it into the band
//! hand the band to the writer (one DNG strip)
//! ```
//!
//! No stage holds the composite (FR-MRG-11): the working set is one
//! chunk's accumulator, one tile, one band of u16 rows. The frame textures
//! are the caller's to provide and cache — `source` is asked for frame `k`
//! as it is needed, and a caller short of memory may demosaic on demand.
//!
//! # What is not here yet
//!
//! A feathered blend, not seams and a Laplacian pyramid: the weight is the
//! distance to the frame's edge, which hides exposure steps and small
//! misalignments and does not hide parallax. Gain is a scalar per frame
//! the caller supplies. Both are panorama.md §10's step 5, after the path
//! writes a file end to end.
use std::sync::Arc;
use dr_pano::bundle::Cameras;
use dr_pano::projection::{Bounds, Projection};
use wgpu::util::DeviceExt;
use crate::readback::await_mapping;
use crate::{AdjustPass, DemosaicedImage, GpuContext, GpuError};
/// One frame's part in the merge.
pub struct MergeFrame {
/// The frame's edit, for its lens corrections — the only part of an
/// edit the camera-space tap uses (FR-MRG-2).
pub graph: Arc<dr_pipeline::EditGraph>,
/// Multiplies the frame's samples, to bring its exposure to the
/// reference frame's. 1.0 for no correction.
pub gain: f32,
}
/// The output the merge produces.
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct MergeOutput {
pub projection: Projection,
/// The projection's scale in output pixels: the cylinder's radius, the
/// plane's distance. The source focal length at full resolution gives
/// output pixels the size of source pixels at the centre.
pub scale: f64,
/// The rectangle of the projection to produce, centred coordinates.
pub bounds: Bounds,
/// Pixels over which a frame's weight ramps up from its edge.
pub feather: f32,
/// Chunk size: the unit of GPU work and of memory.
pub chunk: (u32, u32),
/// Multiplies a normalised sample (1.0 = white) to the sensor's scale.
pub sample_scale: f32,
}
impl MergeOutput {
pub fn width(&self) -> u32 {
self.bounds.width().ceil().max(1.0) as u32
}
pub fn height(&self) -> u32 {
self.bounds.height().ceil().max(1.0) as u32
}
}
/// A band of finished rows: `rows × width × 3` RGB `u16`, plus a coverage
/// mask (`true` where any frame reached the pixel).
pub struct Band<'a> {
pub first_row: u32,
pub rows: u32,
pub rgb: &'a [u16],
pub covered: &'a [bool],
}
#[repr(C)]
#[derive(Clone, Copy, bytemuck::Pod, bytemuck::Zeroable)]
struct WarpParams {
chunk_origin: [f32; 2],
chunk_size: [u32; 2],
projection: u32,
proj_scale: f32,
focal: f32,
gain: f32,
r0: [f32; 4],
r1: [f32; 4],
r2: [f32; 4],
frame_size: [f32; 2],
tile_origin: [f32; 2],
tile_size: [u32; 2],
feather: f32,
_pad: f32,
}
#[repr(C)]
#[derive(Clone, Copy, bytemuck::Pod, bytemuck::Zeroable)]
struct ResolveParams {
chunk_size: [u32; 2],
scale: f32,
_pad: f32,
}
/// The two pipelines and the chunk buffers.
pub struct MergePass {
ctx: GpuContext,
warp: wgpu::ComputePipeline,
warp_layout: wgpu::BindGroupLayout,
resolve: wgpu::ComputePipeline,
resolve_layout: wgpu::BindGroupLayout,
/// Accumulator and packed output for the current chunk size.
buffers: Option<(wgpu::Buffer, wgpu::Buffer, wgpu::Buffer, (u32, u32))>,
}
impl MergePass {
pub fn new(ctx: &GpuContext) -> Result<Self, GpuError> {
let module = ctx
.device
.create_shader_module(wgpu::ShaderModuleDescriptor {
label: Some("merge"),
source: wgpu::ShaderSource::Wgsl(include_str!("shaders/merge.wgsl").into()),
});
let uniform = |binding| wgpu::BindGroupLayoutEntry {
binding,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::Buffer {
ty: wgpu::BufferBindingType::Uniform,
has_dynamic_offset: false,
min_binding_size: None,
},
count: None,
};
let storage = |binding, read_only| wgpu::BindGroupLayoutEntry {
binding,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::Buffer {
ty: wgpu::BufferBindingType::Storage { read_only },
has_dynamic_offset: false,
min_binding_size: None,
},
count: None,
};
let warp_layout = ctx
.device
.create_bind_group_layout(&wgpu::BindGroupLayoutDescriptor {
label: Some("merge-warp-bgl"),
entries: &[
uniform(0),
wgpu::BindGroupLayoutEntry {
binding: 1,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::Texture {
// Unfilterable: rgba32float, loaded by hand.
sample_type: wgpu::TextureSampleType::Float { filterable: false },
view_dimension: wgpu::TextureViewDimension::D2,
multisampled: false,
},
count: None,
},
storage(2, false),
],
});
let resolve_layout =
ctx.device
.create_bind_group_layout(&wgpu::BindGroupLayoutDescriptor {
label: Some("merge-resolve-bgl"),
entries: &[uniform(0), storage(1, true), storage(2, false)],
});
let pipeline = |name: &str, layout: &wgpu::BindGroupLayout| {
let pl = ctx
.device
.create_pipeline_layout(&wgpu::PipelineLayoutDescriptor {
label: Some(name),
bind_group_layouts: &[Some(layout)],
immediate_size: 0,
});
ctx.device
.create_compute_pipeline(&wgpu::ComputePipelineDescriptor {
label: Some(name),
layout: Some(&pl),
module: &module,
entry_point: Some(name),
compilation_options: Default::default(),
cache: None,
})
};
Ok(MergePass {
ctx: ctx.clone(),
warp: pipeline("warp", &warp_layout),
warp_layout,
resolve: pipeline("resolve", &resolve_layout),
resolve_layout,
buffers: None,
})
}
/// Allocate the chunk buffers for this size if the last ones differ.
fn ensure_buffers(&mut self, chunk: (u32, u32)) {
if self.buffers.as_ref().is_none_or(|b| b.3 != chunk) {
let n = u64::from(chunk.0) * u64::from(chunk.1);
let acc = self.ctx.device.create_buffer(&wgpu::BufferDescriptor {
label: Some("merge-acc"),
size: n * 16,
usage: wgpu::BufferUsages::STORAGE | wgpu::BufferUsages::COPY_DST,
mapped_at_creation: false,
});
let out = self.ctx.device.create_buffer(&wgpu::BufferDescriptor {
label: Some("merge-out"),
size: n * 8,
usage: wgpu::BufferUsages::STORAGE | wgpu::BufferUsages::COPY_SRC,
mapped_at_creation: false,
});
let read = self.ctx.device.create_buffer(&wgpu::BufferDescriptor {
label: Some("merge-read"),
size: n * 8,
usage: wgpu::BufferUsages::COPY_DST | wgpu::BufferUsages::MAP_READ,
mapped_at_creation: false,
});
self.buffers = Some((acc, out, read, chunk));
}
}
fn chunk_buffers(&self) -> (&wgpu::Buffer, &wgpu::Buffer, &wgpu::Buffer) {
let b = self.buffers.as_ref().expect("ensured by the caller");
(&b.0, &b.1, &b.2)
}
/// Produce the whole output, band by band, handing each finished band
/// to `sink`.
///
/// `cameras` are in **full-resolution source pixels** (`frame_size`),
/// with frame `k` corresponding to `frames[k]` and `source(k)`. `source`
/// supplies the demosaiced frame on demand and may cache as it sees fit.
#[allow(clippy::too_many_arguments)]
pub fn merge<S, F>(
&mut self,
adjust: &mut AdjustPass,
frames: &[MergeFrame],
cameras: &Cameras,
frame_size: (u32, u32),
output: &MergeOutput,
mut source: S,
mut sink: F,
mut cancelled: impl FnMut() -> bool,
) -> Result<(), GpuError>
where
S: FnMut(usize) -> Result<Arc<DemosaicedImage>, GpuError>,
F: FnMut(Band<'_>) -> Result<(), GpuError>,
{
let (out_w, out_h) = (output.width(), output.height());
let (cw, ch) = (output.chunk.0.max(8), output.chunk.1.max(8));
let (fw, fh) = (frame_size.0 as f64, frame_size.1 as f64);
let mut band_rgb = vec![0u16; (out_w * ch * 3) as usize];
let mut band_cov = vec![false; (out_w * ch) as usize];
let mut chunk_px: Vec<u32> = Vec::new();
let mut y = 0u32;
while y < out_h {
let rows = ch.min(out_h - y);
band_rgb.iter_mut().for_each(|v| *v = 0);
band_cov.iter_mut().for_each(|v| *v = false);
let mut x = 0u32;
while x < out_w {
if cancelled() {
return Err(GpuError::Readback("merge cancelled".into()));
}
let cols = cw.min(out_w - x);
let origin = (
output.bounds.min_u + f64::from(x),
output.bounds.min_v + f64::from(y),
);
self.zero_accumulator((cols, rows));
for (k, frame) in frames.iter().enumerate() {
let Some(rect) = source_rect(
output.projection,
output.scale,
cameras,
k,
origin,
(cols, rows),
(fw, fh),
) else {
continue;
};
let image = source(k)?;
// The tile: that rectangle of the frame, camera-linear,
// at 1:1.
let view = dr_pipeline::CropRect {
x: (rect.0 as f32) / fw as f32,
y: (rect.1 as f32) / fh as f32,
width: (rect.2 as f32) / fw as f32,
height: (rect.3 as f32) / fh as f32,
};
let shader = frame.graph.compose_camera_linear(view);
let tile = adjust.render_camera_linear(&image, &shader, rect.2, rect.3)?;
let r = cameras.rotations[k].transpose();
let params = WarpParams {
chunk_origin: [origin.0 as f32, origin.1 as f32],
chunk_size: [cols, rows],
projection: match output.projection {
Projection::Perspective => 0,
Projection::Cylindrical => 1,
Projection::Spherical => 2,
},
proj_scale: output.scale as f32,
focal: cameras.focal as f32,
gain: frame.gain,
r0: [r.0[0][0] as f32, r.0[0][1] as f32, r.0[0][2] as f32, 0.0],
r1: [r.0[1][0] as f32, r.0[1][1] as f32, r.0[1][2] as f32, 0.0],
r2: [r.0[2][0] as f32, r.0[2][1] as f32, r.0[2][2] as f32, 0.0],
frame_size: [fw as f32, fh as f32],
tile_origin: [rect.0 as f32, rect.1 as f32],
tile_size: [rect.2, rect.3],
feather: output.feather,
_pad: 0.0,
};
self.accumulate(&params, tile);
}
self.resolve_chunk((cols, rows), output.sample_scale, &mut chunk_px)?;
// Into the band.
for row in 0..rows as usize {
for col in 0..cols as usize {
let px = chunk_px[(row * cols as usize + col) * 2..][..2].to_vec();
let i = row * out_w as usize + (x as usize + col);
band_rgb[i * 3] = (px[0] & 0xFFFF) as u16;
band_rgb[i * 3 + 1] = (px[0] >> 16) as u16;
band_rgb[i * 3 + 2] = (px[1] & 0xFFFF) as u16;
band_cov[i] = (px[1] >> 16) != 0;
}
}
x += cols;
}
sink(Band {
first_row: y,
rows,
rgb: &band_rgb[..(out_w * rows * 3) as usize],
covered: &band_cov[..(out_w * rows) as usize],
})?;
y += rows;
}
Ok(())
}
fn zero_accumulator(&mut self, chunk: (u32, u32)) {
self.ensure_buffers(chunk);
let (acc, _, _) = self.chunk_buffers();
let n = u64::from(chunk.0) * u64::from(chunk.1) * 16;
let mut enc = self.ctx.device.create_command_encoder(&Default::default());
enc.clear_buffer(acc, 0, Some(n));
self.ctx.queue.submit(Some(enc.finish()));
}
fn accumulate(&mut self, params: &WarpParams, tile: &wgpu::Texture) {
let chunk = (params.chunk_size[0], params.chunk_size[1]);
let uniforms = self
.ctx
.device
.create_buffer_init(&wgpu::util::BufferInitDescriptor {
label: Some("merge-warp-params"),
contents: bytemuck::bytes_of(params),
usage: wgpu::BufferUsages::UNIFORM,
});
let view = tile.create_view(&Default::default());
self.ensure_buffers(chunk);
let (acc, _, _) = self.chunk_buffers();
let bind = self
.ctx
.device
.create_bind_group(&wgpu::BindGroupDescriptor {
label: Some("merge-warp-bg"),
layout: &self.warp_layout,
entries: &[
wgpu::BindGroupEntry {
binding: 0,
resource: uniforms.as_entire_binding(),
},
wgpu::BindGroupEntry {
binding: 1,
resource: wgpu::BindingResource::TextureView(&view),
},
wgpu::BindGroupEntry {
binding: 2,
resource: acc.as_entire_binding(),
},
],
});
let mut enc = self.ctx.device.create_command_encoder(&Default::default());
{
let mut pass = enc.begin_compute_pass(&Default::default());
pass.set_pipeline(&self.warp);
pass.set_bind_group(0, &bind, &[]);
pass.dispatch_workgroups(chunk.0.div_ceil(8), chunk.1.div_ceil(8), 1);
}
self.ctx.queue.submit(Some(enc.finish()));
}
fn resolve_chunk(
&mut self,
chunk: (u32, u32),
scale: f32,
out: &mut Vec<u32>,
) -> Result<(), GpuError> {
let params = ResolveParams {
chunk_size: [chunk.0, chunk.1],
scale,
_pad: 0.0,
};
let uniforms = self
.ctx
.device
.create_buffer_init(&wgpu::util::BufferInitDescriptor {
label: Some("merge-resolve-params"),
contents: bytemuck::bytes_of(&params),
usage: wgpu::BufferUsages::UNIFORM,
});
let n = u64::from(chunk.0) * u64::from(chunk.1);
self.ensure_buffers(chunk);
let (acc, packed, read) = self.chunk_buffers();
let bind = self
.ctx
.device
.create_bind_group(&wgpu::BindGroupDescriptor {
label: Some("merge-resolve-bg"),
layout: &self.resolve_layout,
entries: &[
wgpu::BindGroupEntry {
binding: 0,
resource: uniforms.as_entire_binding(),
},
wgpu::BindGroupEntry {
binding: 1,
resource: acc.as_entire_binding(),
},
wgpu::BindGroupEntry {
binding: 2,
resource: packed.as_entire_binding(),
},
],
});
let mut enc = self.ctx.device.create_command_encoder(&Default::default());
{
let mut pass = enc.begin_compute_pass(&Default::default());
pass.set_pipeline(&self.resolve);
pass.set_bind_group(0, &bind, &[]);
pass.dispatch_workgroups(chunk.0.div_ceil(8), chunk.1.div_ceil(8), 1);
}
enc.copy_buffer_to_buffer(packed, 0, read, 0, n * 8);
self.ctx.queue.submit(Some(enc.finish()));
let slice = read.slice(..n * 8);
let (tx, rx) = std::sync::mpsc::channel();
slice.map_async(wgpu::MapMode::Read, move |r| {
let _ = tx.send(r);
});
await_mapping(&self.ctx, &rx)?;
{
let data = slice.get_mapped_range();
out.clear();
out.extend_from_slice(bytemuck::cast_slice::<u8, u32>(&data));
}
read.unmap();
Ok(())
}
}
/// The rectangle of frame `k` (x, y, w, h in source pixels) a chunk reads,
/// or `None` if the chunk sees nothing of the frame.
///
/// Walks the chunk's border, projects each point into the frame, and takes
/// the bounding box with a two-pixel margin for the bilinear fetch. The
/// border rather than the corners because under a cylinder or sphere the
/// extreme of a footprint is not at a corner.
fn source_rect(
projection: Projection,
scale: f64,
cameras: &Cameras,
k: usize,
origin: (f64, f64),
size: (u32, u32),
frame: (f64, f64),
) -> Option<(u32, u32, u32, u32)> {
let (w, h) = (f64::from(size.0), f64::from(size.1));
let steps = 16;
let mut min = (f64::MAX, f64::MAX);
let mut max = (f64::MIN, f64::MIN);
let mut any = false;
let mut visit = |u: f64, v: f64| {
let d = projection.to_direction(scale, u, v);
if let Some((x, y)) = cameras.project(k, d) {
let (x, y) = (x + frame.0 / 2.0, y + frame.1 / 2.0);
min = (min.0.min(x), min.1.min(y));
max = (max.0.max(x), max.1.max(y));
any = true;
}
};
for s in 0..=steps {
let t = f64::from(s) / f64::from(steps);
visit(origin.0 + w * t, origin.1);
visit(origin.0 + w * t, origin.1 + h);
visit(origin.0, origin.1 + h * t);
visit(origin.0 + w, origin.1 + h * t);
}
// The interior too, coarsely: a chunk can contain a frame entirely.
for i in 1..4 {
for j in 1..4 {
visit(
origin.0 + w * f64::from(i) / 4.0,
origin.1 + h * f64::from(j) / 4.0,
);
}
}
if !any {
return None;
}
let x0 = (min.0.floor() - 2.0).max(0.0);
let y0 = (min.1.floor() - 2.0).max(0.0);
let x1 = (max.0.ceil() + 2.0).min(frame.0);
let y1 = (max.1.ceil() + 2.0).min(frame.1);
if x1 <= x0 || y1 <= y0 {
return None;
}
Some((x0 as u32, y0 as u32, (x1 - x0) as u32, (y1 - y0) as u32))
}
+7 -7
View File
@@ -1,4 +1,4 @@
//! Watershed segmentation — arm A's GPU half (S15, docs/segmentation.md).
//! Watershed segmentation — arm A's GPU half (S15, docs/dev/segmentation.md).
//!
//! Runs the five passes in `shaders/watershed.wgsl` over a demosaiced image
//! and leaves a basin label per pixel on the GPU. The hierarchy built from
@@ -20,7 +20,7 @@
//! one AC-8 forbids is per frame in the render loop, and sharing a switch
//! would force a build wanting local masking to unlock the other.
//!
//! It is still a real cost and still unfinished. F3 in docs/segmentation.md
//! It is still a real cost and still unfinished. F3 in docs/dev/segmentation.md
//! §12 stands: the adjacency accumulation belongs GPU-side with atomics, and
//! until it moves there every segmentation pays a full-resolution transfer.
//! Read the feature name as a description of a known gap rather than as
@@ -36,7 +36,7 @@ pub struct SegmentOptions {
/// Longest proxy edge. The segmentation runs here, not at sensor
/// resolution: a 24 MP watershed costs 12× the memory to place boundaries
/// a person cannot see, and the boundary refinement that matters at 1:1
/// is a separate stage (docs/segmentation.md §4).
/// is a separate stage (docs/dev/segmentation.md §4).
pub max_edge: u32,
/// Pre-smoothing radius in proxy pixels. The caller's to raise with ISO —
/// this is the single knob that decides whether a noisy file segments
@@ -69,7 +69,7 @@ impl Default for SegmentOptions {
w_chroma: 0.5,
// **Zero: the pass is off.** It is implemented, dispatched
// correctly and measurably changes nothing — see the ignored test
// below and §12 of docs/segmentation.md. Until that is understood,
// below and §12 of docs/dev/segmentation.md. Until that is understood,
// running it would buy 64 dispatches per segmentation and no
// improvement, so the default declines to pay.
plateau_iterations: 0,
@@ -486,7 +486,7 @@ impl Segmentation {
/// a region graph of a few thousand nodes that every later interaction
/// reads from the CPU anyway.
///
/// What it is *not* is finished. F3 in docs/segmentation.md §12 stands:
/// What it is *not* is finished. F3 in docs/dev/segmentation.md §12 stands:
/// the adjacency accumulation belongs on the GPU with atomics, and until
/// it moves there a segmentation costs one full-resolution transfer of the
/// label and gradient buffers. That is a real cost on a phone and the
@@ -724,7 +724,7 @@ mod tests {
px
}
#[test]
#[ignore = "the plateau pass is a measured no-op; see docs/segmentation.md §12"]
#[ignore = "the plateau pass is a measured no-op; see docs/dev/segmentation.md §12"]
fn lower_completion_drains_a_plateau_instead_of_shattering_it() {
// F1, asserted rather than eyeballed, and asserted at the level where
// it matters.
@@ -741,7 +741,7 @@ mod tests {
// with no exit anywhere — cannot be drained by a distance that has
// nowhere to descend to, and collapsing it fully would need connected
// component labelling rather than a local rule. It is not worth it:
// see docs/segmentation.md §12.
// see docs/dev/segmentation.md §12.
let Some(ctx) = ctx() else { return };
let (w, h) = (96u32, 96u32);
let src = DemosaicedImage::from_rgba8(&ctx, &ramp(w, h), w, h).expect("source");
+10 -4
View File
@@ -160,12 +160,18 @@ fn main(@builtin(global_invocation_id) gid: vec3<u32>) {
// Green is measured. Red and blue are interpolated from their own
// axis, with a correction from the green Laplacian.
//
// Malvar "G at R/B locations" kernels, transposed per axis:
// chroma along the row: (5c + 4(w1+e1) - (nw+ne+sw+se) - (n2+s2) + 0.5(w2+e2)) / 8
// Malvar "R at green in R row" kernel, and its transpose:
// chroma along the row: (5c + 4(w1+e1) - (nw+ne+sw+se) - (w2+e2) + 0.5(n2+s2)) / 8
//
// The -1 goes on the two greens *along* the chroma axis and the +0.5
// on the pair across it. Transposed, both kernels still sum to zero
// and reconstruct a flat patch exactly, but on an edge the correction
// at green sites is half strength and the false colour doubles: a
// blue/yellow zipper around every clipped highlight.
let along_row =
(5.0 * c + 4.0 * (w1 + e1) - diag1 - vert2 + 0.5 * horiz2) * 0.125;
(5.0 * c + 4.0 * (w1 + e1) - diag1 - horiz2 + 0.5 * vert2) * 0.125;
let along_col =
(5.0 * c + 4.0 * (n1 + s1) - diag1 - horiz2 + 0.5 * vert2) * 0.125;
(5.0 * c + 4.0 * (n1 + s1) - diag1 - vert2 + 0.5 * horiz2) * 0.125;
let red_horizontal = red_is_horizontal(gid.x, gid.y);
let r = select(along_col, along_row, red_horizontal);
+166
View File
@@ -0,0 +1,166 @@
// TRACES: FR-MRG-10 | FR-MRG-11
// The merge: one source tile warped into one output chunk, accumulated.
//
// Two entry points. `warp` runs once per (chunk, frame): for every chunk
// pixel it asks which direction that pixel looks along, turns the
// direction into the frame's camera, projects it to a source pixel, and
// if that pixel is inside the tile that was rendered for this chunk,
// samples it and adds it — weighted by its distance from the frame's edge
// — into the accumulator. `resolve` runs once per chunk after every frame
// has been added: divides the sums by the weights and packs the result as
// sixteen-bit samples at the sensor's scale (FR-MRG-3).
//
// The accumulator is a buffer and not a storage texture, because WebGPU
// allows a read-write storage texture only in the 32-bit single-channel
// formats, and this wants four channels. The tile is sampled by hand from
// four `textureLoad`s rather than through a sampler, because `rgba32float`
// is not filterable without an optional feature, and the tile is
// `rgba32float` on purpose (panorama.md §5.1).
//
// The projection maths is `dr_pano::projection` verbatim; the two must
// agree, and a golden test compares them.
struct Params {
// Where the chunk's pixel (0, 0) sits in centred output coordinates,
// and the chunk's size.
chunk_origin: vec2<f32>,
chunk_size: vec2<u32>,
// 0 perspective, 1 cylindrical, 2 spherical; and the projection's
// scale (the cylinder's radius, the sphere's, the plane's distance) in
// output pixels.
projection: u32,
proj_scale: f32,
// The frame's focal length in source pixels, and the gain the frame's
// exposure is corrected by.
focal: f32,
gain: f32,
// World → this frame's camera: the transpose of its rotation, one row
// per vec4 (padded).
r0: vec4<f32>,
r1: vec4<f32>,
r2: vec4<f32>,
// The full frame's size in source pixels (for the edge weight), the
// tile's origin within the frame, and the tile's size.
frame_size: vec2<f32>,
tile_origin: vec2<f32>,
tile_size: vec2<u32>,
// Pixels over which the weight ramps from the edge to full.
feather: f32,
_pad: f32,
};
@group(0) @binding(0) var<uniform> p: Params;
@group(0) @binding(1) var tile: texture_2d<f32>;
// rgb·w summed, then w: four floats per chunk pixel.
@group(0) @binding(2) var<storage, read_write> acc: array<vec4<f32>>;
fn to_direction(u: f32, v: f32) -> vec3<f32> {
let s = p.proj_scale;
if (p.projection == 0u) {
return normalize(vec3<f32>(u, v, s));
}
if (p.projection == 1u) {
let theta = u / s;
return normalize(vec3<f32>(sin(theta), v / s, cos(theta)));
}
let theta = u / s;
let phi = v / s;
return vec3<f32>(sin(theta) * cos(phi), sin(phi), cos(theta) * cos(phi));
}
fn load(x: i32, y: i32) -> vec4<f32> {
return textureLoad(tile, vec2<i32>(x, y), 0);
}
@compute @workgroup_size(8, 8, 1)
fn warp(@builtin(global_invocation_id) gid: vec3<u32>) {
if (gid.x >= p.chunk_size.x || gid.y >= p.chunk_size.y) {
return;
}
let u = p.chunk_origin.x + f32(gid.x) + 0.5;
let v = p.chunk_origin.y + f32(gid.y) + 0.5;
let d = to_direction(u, v);
let c = vec3<f32>(dot(p.r0.xyz, d), dot(p.r1.xyz, d), dot(p.r2.xyz, d));
if (c.z <= 1e-6) {
return;
}
// Source pixel, in the full frame, with the principal point at its
// centre. `- 0.5` puts pixel centres on integer coordinates for the
// bilinear fetch below.
let sx = p.focal * c.x / c.z + p.frame_size.x * 0.5 - 0.5;
let sy = p.focal * c.y / c.z + p.frame_size.y * 0.5 - 0.5;
// Weight: distance to the nearest frame edge, in pixels, over the
// feather. Zero outside the frame.
let edge = min(min(sx, p.frame_size.x - 1.0 - sx), min(sy, p.frame_size.y - 1.0 - sy));
if (edge <= 0.0) {
return;
}
let w = clamp(edge / max(p.feather, 1.0), 0.0, 1.0);
// Into the tile.
let tx = sx - p.tile_origin.x;
let ty = sy - p.tile_origin.y;
let tw = f32(p.tile_size.x);
let th = f32(p.tile_size.y);
if (tx < 0.0 || ty < 0.0 || tx > tw - 1.0 || ty > th - 1.0) {
return;
}
let x0 = i32(floor(tx));
let y0 = i32(floor(ty));
let x1 = min(x0 + 1, i32(p.tile_size.x) - 1);
let y1 = min(y0 + 1, i32(p.tile_size.y) - 1);
let fx = tx - f32(x0);
let fy = ty - f32(y0);
// The four texels, with their alpha: the tap writes alpha 0 where the
// lens correction found no source pixel, and a sample that touches one
// of those is a partial pixel — down-weighted by exactly how much of
// it is missing, and dropped when all of it is.
let s00 = load(x0, y0);
let s10 = load(x1, y0);
let s01 = load(x0, y1);
let s11 = load(x1, y1);
let top = mix(s00, s10, fx);
let bot = mix(s01, s11, fx);
let s = mix(top, bot, fy);
if (s.a <= 0.001) {
return;
}
// Colour is the alpha-weighted mean of the texels that exist.
let rgb = s.rgb / s.a * p.gain;
let wa = w * s.a;
let i = gid.y * p.chunk_size.x + gid.x;
acc[i] = acc[i] + vec4<f32>(rgb * wa, wa);
}
// Resolve: the accumulated chunk to sixteen-bit samples.
struct ResolveParams {
chunk_size: vec2<u32>,
// Multiplies a normalised value (1.0 = the sensor's white) back to the
// sensor's scale: the source's white minus its black (FR-MRG-3).
scale: f32,
_pad: f32,
};
@group(0) @binding(0) var<uniform> rp: ResolveParams;
@group(0) @binding(1) var<storage, read> racc: array<vec4<f32>>;
// Two u32 per pixel: (r | g << 16), (b | coverage << 16). Coverage is
// 65535 where any frame reached the pixel and 0 where none did, so the
// CPU can tell an empty pixel from a black one.
@group(0) @binding(2) var<storage, read_write> out: array<vec2<u32>>;
@compute @workgroup_size(8, 8, 1)
fn resolve(@builtin(global_invocation_id) gid: vec3<u32>) {
if (gid.x >= rp.chunk_size.x || gid.y >= rp.chunk_size.y) {
return;
}
let i = gid.y * rp.chunk_size.x + gid.x;
let a = racc[i];
if (a.w <= 0.0) {
out[i] = vec2<u32>(0u, 0u);
return;
}
let rgb = clamp(a.rgb / a.w * rp.scale, vec3<f32>(0.0), vec3<f32>(65535.0));
let r = u32(round(rgb.r));
let g = u32(round(rgb.g));
let b = u32(round(rgb.b));
out[i] = vec2<u32>(r | (g << 16u), b | (65535u << 16u));
}
+2 -2
View File
@@ -1,4 +1,4 @@
// Watershed segmentation — the passes behind arm A of S15 (docs/segmentation.md).
// Watershed segmentation — the passes behind arm A of S15 (docs/dev/segmentation.md).
//
// Seven entry points forming one chain:
//
@@ -197,7 +197,7 @@ fn gradient(@builtin(global_invocation_id) gid: vec3<u32>) {
// lowest-indexed neighbour, which is up and to the left. Each pixel therefore
// walks diagonally until it falls off the plateau, and one flat region becomes
// a fan of diagonal chains rather than one basin — visible as hatching across
// what should be a single area (docs/segmentation.md §12, F1).
// what should be a single area (docs/dev/segmentation.md §12, F1).
//
// The fix is the standard lower-completion: give each plateau pixel its
// geodesic distance to the nearest pixel that *does* have a lower neighbour,
+4
View File
@@ -40,6 +40,10 @@ fn flat_raw(level: u16, curve: BaseCurve) -> RawImage {
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
color_matrix: Some([1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0]),
base_curve: curve,
samples_per_pixel: 1,
profile: None,
make: String::new(),
model: String::new(),
crop: CropRect {
x: 0,
y: 0,
+4
View File
@@ -46,6 +46,10 @@ fn flat_raw(level: u16) -> RawImage {
// leaving a curve here would test the suppression rather than the
// film. `dr-pipeline` asserts the suppression on the generated source.
base_curve: BaseCurve::IDENTITY,
samples_per_pixel: 1,
profile: None,
make: String::new(),
model: String::new(),
crop: CropRect {
x: 0,
y: 0,
+5 -5
View File
@@ -2,11 +2,11 @@
//!
//! FR-DSP-3 says a slider updates the visible region within one frame budget at
//! proxy resolution. Until this file existed nothing checked it, which made it
//! a wish — `docs/display-and-extension.md` §3 is blunt about that, and §7 is
//! a wish — `docs/dev/display-and-extension.md` §3 is blunt about that, and §7 is
//! blunt about what tagging an unchecked requirement does to the coverage
//! figure.
//!
//! The measurements this guards are in [`docs/frame-budget.md`], produced by
//! The measurements this guards are in [`docs/dev/frame-budget.md`], produced by
//! `examples/frame_budget.rs`. This file is the part of them that has to keep
//! being true: it renders the **whole point-operation chain** through the real
//! `render_detailed` for a hundred frames, moving a slider between each, and
@@ -17,7 +17,7 @@
//! **The neighbourhood stage is deliberately not in the asserted chain.** It is
//! over the budget today — clarity alone is 34 ms at 4K, because its kernel is
//! a fraction of the frame and reaches a 52-pixel radius there — and
//! `docs/frame-budget.md` records that, names the fix (a base computed at
//! `docs/dev/frame-budget.md` records that, names the fix (a base computed at
//! reduced resolution) and does not pretend otherwise. Asserting a budget the
//! code does not meet would produce a red suite that everyone learns to ignore;
//! asserting it on a chain that quietly excluded the expensive stage *without
@@ -81,7 +81,7 @@ const SOURCE: (u32, u32) = (6000, 4000);
/// The viewport the budget is asserted at: a 16:10 desktop display.
///
/// Not 4K, and the reason is worth stating. At 4K the fused chain still passes
/// with room to spare (4.5 ms of GPU; see `docs/frame-budget.md`), but a test
/// with room to spare (4.5 ms of GPU; see `docs/dev/frame-budget.md`), but a test
/// that renders 8.3 M pixels a hundred times twice over is four seconds of
/// suite time to re-establish a conclusion 4.1 M pixels already establishes.
const VIEWPORT: (u32, u32) = (2560, 1600);
@@ -166,7 +166,7 @@ impl Run {
judged <= BUDGET_MS,
"{case} at {}x{}: p99 of {FRAMES} frames was {judged:.2} ms, over the \
{BUDGET_MS:.0} ms budget (cpu {:.2} ms, gpu {:.2} ms, total {:.2} ms). \
FR-DSP-3 is what this violates; docs/frame-budget.md holds the \
FR-DSP-3 is what this violates; docs/dev/frame-budget.md holds the \
numbers it used to be.",
viewport.0,
viewport.1,
+113
View File
@@ -790,6 +790,119 @@ fn the_order_parts_are_joined_in_is_the_mask() {
);
}
/// TRACES: FR-DEV-19a
/// The truth table `docs/dev/mask-editing.md` §13 asks for: one base, one
/// part that half-covers it, joined each of the three ways. The base is the
/// left half of the frame and the part a dab in the middle, so the four
/// quarters of the table are four pixels.
#[test]
fn a_part_unioned_subtracted_and_intersected_gives_the_three_fields() {
let Some(ctx) = ctx() else {
eprintln!("no adapter; skipping");
return;
};
let field = split_field(&ctx);
let joined = |join: Join| {
let mut layer = brighten(MaskSource::Regions {
signature: 1,
level: 2,
ids: vec![0],
});
assert!(layer.push_part(MaskPart::painted("p2", join)));
paint(&mut layer, 1, false, &[(0.5, 0.5)]);
let mut stack = MaskStack::new();
stack.push(layer);
render(&ctx, &stack, Some(&field))
};
// (base, part): left outside the dab, left inside, right inside, right
// outside.
let cells = [(4, 16), (14, 16), (18, 16), (27, 16)];
let lit = |pixels: &[u8]| cells.map(|(x, y)| luma_at(pixels, x, y) > 200);
assert_eq!(
lit(&joined(Join::Union)),
[true, true, true, false],
"union: either"
);
assert_eq!(
lit(&joined(Join::Subtract)),
[true, false, false, false],
"subtract: the base without the dab"
);
assert_eq!(
lit(&joined(Join::Intersect)),
[false, true, false, false],
"intersect: only where both are"
);
}
/// TRACES: FR-DEV-19a
/// Intersection on soft coverage is the product `Join::apply` defines, on
/// either side of the join: a gradient intersected with a region it fills is
/// the gradient there and nothing elsewhere, and a region intersected with a
/// gradient is the gradient wherever the region is.
#[test]
fn the_joins_match_their_definition() {
let Some(ctx) = ctx() else {
eprintln!("no adapter; skipping");
return;
};
let field = split_field(&ctx);
let ramp = || MaskSource::Linear {
centre: (0.5, 0.5),
angle: 0.0,
width: 1.0,
};
let right_half = || MaskSource::Regions {
signature: 1,
level: 2,
ids: vec![1],
};
let draw = |layer: MaskLayer| {
let mut stack = MaskStack::new();
stack.push(layer);
render(&ctx, &stack, Some(&field))
};
let alone = draw(brighten(ramp()));
let mut ramp_then_region = brighten(ramp());
assert!(ramp_then_region.push_part(MaskPart::new("p2", Join::Intersect, right_half())));
let ramp_then_region = draw(ramp_then_region);
let mut region_then_ramp = brighten(whole_frame());
assert!(region_then_ramp.push_part(MaskPart::new("p2", Join::Intersect, ramp())));
let region_then_ramp = draw(region_then_ramp);
for x in 0..SIZE {
let y = SIZE / 2;
let want = luma_at(&alone, x, y);
// dst · 1 = dst on the right; dst · 0 = 0 on the left. Pixels
// within two of the seam are left out: the region's own edge
// is soft there, so neither side of the table is 0 or 1.
let got = luma_at(&ramp_then_region, x, y);
if x.abs_diff(SIZE / 2) <= 2 {
// The seam.
} else if x > SIZE / 2 {
assert!(
got.abs_diff(want) <= 1,
"x={x}: the ramp survives where the region is ({got} vs {want})"
);
} else {
assert_eq!(got, 128, "x={x}: and nothing survives where it is not");
}
// 1 · src = src everywhere.
let got = luma_at(&region_then_ramp, x, y);
assert!(
got.abs_diff(want) <= 1,
"x={x}: a full base intersected with the ramp is the ramp ({got} vs {want})"
);
}
}
// --- seeing the mask (FR-DEV-19c) ------------------------------------------
/// A radial that covers the middle of the frame and nothing near the corners.
+1 -1
View File
@@ -326,7 +326,7 @@ fn a_proxy_and_an_export_agree_about_the_effect() {
#[test]
fn crossing_the_reduction_threshold_does_not_change_the_picture() {
// TRACES: FR-DSP-3 — `docs/technical-debt.md` TD-4, held in pixels.
// TRACES: FR-DSP-3 — `docs/dev/technical-debt.md` TD-4, held in pixels.
//
// Clarity's base is computed on a reduced grid, and how reduced depends on
// the viewport: `LocalContrast::reduction` steps 4 -> 2 -> 1 as sigma
+2 -2
View File
@@ -1,4 +1,4 @@
//! TRACES: FR-DSP-5
//! TRACES: FR-DSP-5 | R5
//! Zooming to 1:1 samples the source, pixel for pixel.
//!
//! FR-DSP-5: *"Fit, 1:1, and arbitrary zoom levels. At 1:1 and above, the
@@ -10,7 +10,7 @@
//! code path — the zoom is the full-resolution path — which is why the
//! requirement has been satisfied for some time without anyone tagging it.
//!
//! `docs/display-and-extension.md` §7 is the reason this file exists rather
//! `docs/dev/display-and-extension.md` §7 is the reason this file exists rather
//! than a tag on `framing.rs`: a requirement counts as covered when a `TRACES`
//! comment names it, and nothing checks that the code under the tag does the
//! thing. `FR-DEV-8` is tagged against plumbing a future operation would use.
+52
View File
@@ -0,0 +1,52 @@
[package]
name = "dr-inference-engine"
version.workspace = true
edition.workspace = true
rust-version.workspace = true
license.workspace = true
# The one crate that names a runtime, a provider, a vendor library or a
# device (docs/dev/inference.md §8). `dr-face` and `dr-segment` ask it for a
# session by role and never see which of these answered.
[dependencies]
thiserror.workspace = true
log.workspace = true
serde.workspace = true
serde_json.workspace = true
# `ort` is the API; what supplies it is decided once per process (§3):
# `libonnxruntime` found on disk, or `tract`. Both are behind
# `alternative-backend`, so nothing here links C on any target.
ort = { workspace = true }
ort-tract = { workspace = true, optional = true }
# dlopen, and the C types of the table it fetches. Both pure Rust;
# `libloading` is already in the tree through wgpu.
libloading = { version = "0.8", optional = true }
ort-sys = { version = "2.0.0-rc.13", default-features = false, features = ["disable-linking"], optional = true }
# The NVIDIA rungs exist on the desktop only. These features add `ort`'s
# option builders and nothing else — no linking under `alternative-backend` —
# but an Android binary has no business carrying even the option names, and
# the packaging must never be tempted to (§2, §3.1). The AMD rung needs no
# feature: MIGraphX is registered through the runtime's generic key/value
# entry point (`session::migraphx`), because `ort`'s own builder cannot
# name the compiled-program cache.
[target.'cfg(not(target_os = "android"))'.dependencies]
ort = { workspace = true, features = ["cuda", "tensorrt"] }
[target.'cfg(target_os = "android")'.dependencies]
ort = { workspace = true, features = ["qnn"] }
[features]
# The floor: `tract` supplies the API table when no runtime file is found, or
# always, in a build without `native`. Tests want this and nothing else.
default = ["tract"]
tract = ["dep:ort-tract"]
# Look for `libonnxruntime` on disk and hand its table to `ort`.
native = ["dep:libloading", "dep:ort-sys"]
[dev-dependencies]
# The `ep_probe` example prints the provider's own diagnostics, which is most
# of what a failed rung tells you.
env_logger.workspace = true
@@ -0,0 +1,207 @@
//! Time each execution provider a runtime offers, on the models this
//! repository ships — the measurement docs/inference.md §1 requires before a
//! rung is added to §2's ladder.
//!
//! DARKROOM_ORT_DIR=/usr/lib \
//! cargo run --release -p dr-inference-engine --features native,tract \
//! --example ep_probe -- models/face/scrfd_500m_640.onnx ...
//!
//! Prints one row per (model, provider): the median of timed runs after
//! warm-ups, and the build time, which for a compiling provider is the
//! number that decides whether it needs an engine cache. MIGraphX is built
//! twice per precision — cold, then again from the cache it just wrote —
//! so both numbers are on the page.
//!
//! The ROCm provider is not in the list: ONNX Runtime removed it in 1.23,
//! and 1.29's `onnxruntime-rocm` ships `libonnxruntime_providers_migraphx.so`
//! and nothing else for AMD.
use std::path::{Path, PathBuf};
use std::time::Instant;
#[derive(Clone, Copy, PartialEq)]
enum Ep {
Cpu,
MiGraphX,
MiGraphXFp16,
}
impl Ep {
fn label(self) -> &'static str {
match self {
Ep::Cpu => "CPU",
Ep::MiGraphX => "MIGraphX f32",
Ep::MiGraphXFp16 => "MIGraphX fp16",
}
}
}
fn build(ep: Ep, bytes: &[u8], threads: usize, cache: &Path) -> ort::Result<ort::session::Session> {
let mut b = ort::session::Session::builder()?.with_intra_threads(threads)?;
match ep {
Ep::Cpu => {}
Ep::MiGraphX => migraphx(&mut b, false, &cache.join("f32"))?,
Ep::MiGraphXFp16 => migraphx(&mut b, true, &cache.join("fp16"))?,
}
b.commit_from_memory(bytes)
}
/// Register MIGraphX through the generic key/value API. `ort`'s own
/// builder fills the legacy `OrtMIGraphXProviderOptions`, which 1.29 reads
/// for its precision flags and nothing else: the model cache directory —
/// the difference between a 40 s load and a 0.3 s one — only travels this
/// way. The cache key is the graph, the GPU and the MIGraphX version, not
/// the precision, so each precision gets its own directory.
fn migraphx(
b: &mut ort::session::builder::SessionBuilder,
fp16: bool,
cache: &Path,
) -> ort::Result<()> {
use ort::AsPointer;
use std::ffi::CString;
std::fs::create_dir_all(cache).map_err(|e| ort::Error::new(e.to_string()))?;
let keys = [c"migraphx_fp16_enable", c"migraphx_model_cache_dir"];
let values = [
CString::new(if fp16 { "1" } else { "0" }).unwrap(),
CString::new(cache.to_string_lossy().as_bytes()).unwrap(),
];
let key_ptrs: Vec<_> = keys.iter().map(|k| k.as_ptr()).collect();
let value_ptrs: Vec<_> = values.iter().map(|v| v.as_ptr()).collect();
// SAFETY: the documented C call, over arrays that outlive it; the
// runtime copies the strings into its own options map.
unsafe {
let status = (ort::api().SessionOptionsAppendExecutionProvider)(
b.ptr_mut(),
c"MIGraphX".as_ptr(),
key_ptrs.as_ptr(),
value_ptrs.as_ptr(),
keys.len(),
);
ort::Error::result_from_status(status)
}
}
/// Median of `runs` timed runs over zeros, in milliseconds, after warm-ups.
fn time(session: &mut ort::session::Session, warmups: usize, runs: usize) -> Result<f64, String> {
let shape: Vec<usize> = session.inputs()[0]
.dtype()
.tensor_shape()
.ok_or("input is not a tensor")?
.iter()
.map(|&d| if d > 0 { d as usize } else { 1 })
.collect();
let zeros = vec![0f32; shape.iter().product()];
let once = |s: &mut ort::session::Session| -> Result<f64, String> {
let input = ort::value::Tensor::from_array((shape.clone(), zeros.clone()))
.map_err(|e| e.to_string())?;
let t = Instant::now();
let out = s.run(ort::inputs![input]).map_err(|e| e.to_string())?;
let _ = out[0]
.try_extract_tensor::<f32>()
.map_err(|e| e.to_string())?;
Ok(t.elapsed().as_secs_f64() * 1e3)
};
for _ in 0..warmups {
once(session)?;
}
let mut times = Vec::with_capacity(runs);
for _ in 0..runs {
times.push(once(session)?);
}
times.sort_by(|a, b| a.partial_cmp(b).unwrap());
Ok(times[times.len() / 2])
}
fn first_line(s: &str) -> String {
s.lines().next().unwrap_or("").chars().take(120).collect()
}
fn main() {
env_logger::Builder::from_env(env_logger::Env::default().default_filter_or("info")).init();
let models: Vec<PathBuf> = std::env::args_os().skip(1).map(PathBuf::from).collect();
if models.is_empty() {
eprintln!("usage: ep_probe MODEL.onnx [MODEL.onnx ...]");
std::process::exit(2);
}
dr_inference_engine::ensure_runtime();
let runtime = dr_inference_engine::status().runtime;
println!("runtime: {}", runtime.label());
if !runtime.is_native() {
println!("(tract: no provider to compare; set DARKROOM_ORT_DIR)");
}
let threads = std::thread::available_parallelism()
.map(|n| n.get().saturating_sub(2).max(1))
.unwrap_or(1);
println!("intra-op threads: {threads}");
let cache = std::env::temp_dir().join("darkroom-ep-probe");
let _ = std::fs::remove_dir_all(&cache);
println!("compiled-program cache: {}\n", cache.display());
println!(
"{:<28} {:<15} {:>10} {:>10}",
"model", "provider", "build s", "median ms"
);
for model in &models {
let bytes = match std::fs::read(model) {
Ok(b) => b,
Err(e) => {
println!("{:<28} read failed: {e}", name(model));
continue;
}
};
// A compiling provider is built twice: the second build reads the
// program the first wrote, and its time is what a launch after the
// first costs.
let plan = [
(Ep::Cpu, false),
(Ep::MiGraphX, false),
(Ep::MiGraphX, true),
(Ep::MiGraphXFp16, false),
(Ep::MiGraphXFp16, true),
];
for (ep, cached) in plan {
let started = Instant::now();
match build(ep, &bytes, threads, &cache) {
Ok(mut session) => {
let built = started.elapsed().as_secs_f64();
match time(&mut session, 3, 15) {
Ok(ms) => println!(
"{:<28} {:<15} {:>10.1} {:>10.1}{}",
name(model),
ep.label(),
built,
ms,
if cached { " (from cache)" } else { "" }
),
Err(e) => println!(
"{:<28} {:<15} {:>10.1} {:>10} {}",
name(model),
ep.label(),
built,
"ran ✗",
first_line(&e)
),
}
}
Err(e) => println!(
"{:<28} {:<15} {:>21} {}",
name(model),
ep.label(),
"build ✗",
first_line(&e.to_string())
),
}
}
println!();
}
}
fn name(p: &Path) -> String {
p.file_name()
.unwrap_or(p.as_os_str())
.to_string_lossy()
.into_owned()
}
+100
View File
@@ -0,0 +1,100 @@
//! Walk the ladder as the app does — probe, engines, then a session — and
//! say what each step chose. The M5 check of docs/inference.md §6 without
//! the app around it.
//!
//! DARKROOM_ORT_DIR=/usr/lib \
//! cargo run --release -p dr-inference-engine --features native,tract \
//! --example ladder -- CACHE_DIR models/face/scrfd_500m_640.onnx [MODEL.onnx ...]
//!
//! Every model named is a `Detector` for the config's purposes, which is
//! enough to see the rung taken, the engines compiled and a session land
//! on it. Delete `CACHE_DIR` to see the first run again; keep it to see the
//! second.
use std::path::PathBuf;
use std::time::{Duration, Instant};
fn main() {
env_logger::Builder::from_env(env_logger::Env::default().default_filter_or("info")).init();
let mut args = std::env::args_os().skip(1).map(PathBuf::from);
let (Some(cache_dir), models) = (args.next(), args.collect::<Vec<_>>()) else {
eprintln!("usage: ladder CACHE_DIR MODEL.onnx [MODEL.onnx ...]");
std::process::exit(2);
};
if models.is_empty() {
eprintln!("usage: ladder CACHE_DIR MODEL.onnx [MODEL.onnx ...]");
std::process::exit(2);
}
let runtime_dirs: Vec<PathBuf> = std::env::var_os("DARKROOM_ORT_DIR")
.map(PathBuf::from)
.into_iter()
.collect();
let started = Instant::now();
dr_inference_engine::init(dr_inference_engine::Config {
runtime_dirs,
cache_dir: cache_dir.clone(),
models: models
.iter()
.map(|p| (dr_inference_engine::Role::Detector, p.clone()))
.collect(),
embedded: Vec::new(),
ceiling: None,
threads: 0,
decay: Duration::ZERO,
});
let mut last = String::new();
loop {
let s = dr_inference_engine::status();
let line = format!(
"{} · {} · engines {}/{}{}",
s.line(),
if s.probing {
"probing"
} else {
s.reason.as_str()
},
s.engines.0,
s.engines.1,
if s.failed.is_empty() {
String::new()
} else {
format!(
" · tried {}",
s.failed
.iter()
.map(|(r, why)| format!("{}: {why}", r.label()))
.collect::<Vec<_>>()
.join(" · ")
)
}
);
if line != last {
println!("{:>6.1} s {line}", started.elapsed().as_secs_f64());
last = line;
}
if !s.probing && s.engines.0 >= s.engines.1 {
break;
}
std::thread::sleep(Duration::from_millis(500));
}
for path in &models {
let bytes = std::fs::read(path).expect("read model");
let t = Instant::now();
let model = dr_inference_engine::open(
dr_inference_engine::Role::Detector,
dr_inference_engine::Form::F32,
&bytes,
)
.expect("open model");
let acquired = model.acquire().expect("acquire session");
println!(
"{} on {} in {:.2} s",
path.file_name().unwrap().to_string_lossy(),
acquired.rung().label(),
t.elapsed().as_secs_f64()
);
}
}
+169
View File
@@ -0,0 +1,169 @@
//! The API table `ort` runs on, chosen once (docs/dev/inference.md §3).
//!
//! `ort` with `alternative-backend` links no runtime and asks, on first use,
//! for an `OrtApi` — a struct of function pointers. Two things can fill it:
//! a `libonnxruntime` this module `dlopen`s, or `ort-tract`. The Rust build
//! is identical either way; the difference is whether a file was found.
use std::path::PathBuf;
use std::sync::OnceLock;
/// What supplied the table.
#[derive(Clone, Debug, PartialEq, Eq)]
pub enum Runtime {
/// Pure Rust, one core, every operator these graphs use. The floor.
Tract,
/// The C++ ONNX Runtime, loaded from `path`.
OnnxRuntime { path: PathBuf, version: String },
}
impl Runtime {
pub fn label(&self) -> String {
match self {
Runtime::Tract => "tract".into(),
Runtime::OnnxRuntime { version, .. } => format!("ONNX Runtime {version}"),
}
}
pub fn is_native(&self) -> bool {
matches!(self, Runtime::OnnxRuntime { .. })
}
}
static RUNTIME: OnceLock<Runtime> = OnceLock::new();
/// The runtime in use; tract until something installs another.
pub fn runtime() -> Runtime {
RUNTIME.get().cloned().unwrap_or(Runtime::Tract)
}
/// Install a table if none is installed yet — tract, since no directories
/// were named. What a test or an example gets, unless `DARKROOM_ORT_DIR`
/// names a runtime: the same variable the desktop honours, so an example
/// can be pointed at the runtime the app uses without learning `init`.
pub fn ensure_installed() {
if RUNTIME.get().is_none() {
let dirs: Vec<PathBuf> = std::env::var_os("DARKROOM_ORT_DIR")
.map(PathBuf::from)
.into_iter()
.collect();
install(&dirs);
}
}
/// Look for `libonnxruntime` in `dirs`, in order, and hand `ort` the first
/// table that loads; otherwise tract. Once per process.
pub fn install(dirs: &[PathBuf]) -> Runtime {
RUNTIME
.get_or_init(|| {
#[cfg(feature = "native")]
for dir in dirs {
match load_native(dir) {
Ok(rt) => return rt,
Err(e) => log::info!("inference: no runtime in {}: {e}", dir.display()),
}
}
#[cfg(not(feature = "native"))]
let _ = dirs;
install_tract()
})
.clone()
}
#[cfg(feature = "tract")]
fn install_tract() -> Runtime {
let _ = ort::set_api(ort_tract::api());
Runtime::Tract
}
#[cfg(not(feature = "tract"))]
fn install_tract() -> Runtime {
// A build with neither tract nor a runtime file has nothing to run
// models on; every `open` will report the un-set API rather than panic
// somewhere deeper.
log::error!("inference: no ONNX Runtime found and tract is not compiled in");
Runtime::Tract
}
#[cfg(feature = "native")]
fn load_native(dir: &std::path::Path) -> Result<Runtime, String> {
let name = if cfg!(target_os = "windows") {
"onnxruntime.dll"
} else if cfg!(any(target_os = "macos", target_os = "ios")) {
"libonnxruntime.dylib"
} else {
"libonnxruntime.so"
};
// An empty dir means the bare name: the system loader's search, which on
// Android includes the APK's own native libraries.
let path = if dir.as_os_str().is_empty() {
PathBuf::from(name)
} else {
find_library(dir, name).ok_or("not present")?
};
// SAFETY: the library's initialisers are ONNX Runtime's own; the symbol
// is the documented entry point with the documented signature; the table
// is copied out and the library handle is leaked, so every pointer in
// the copy stays valid for the life of the process.
unsafe {
let lib = libloading::Library::new(&path).map_err(|e| e.to_string())?;
let get_base: libloading::Symbol<
unsafe extern "system" fn() -> *const ort_sys::OrtApiBase,
> = lib.get(b"OrtGetApiBase\0").map_err(|e| e.to_string())?;
let base = get_base();
if base.is_null() {
return Err("OrtGetApiBase returned null".into());
}
let version = std::ffi::CStr::from_ptr(((*base).GetVersionString)())
.to_string_lossy()
.into_owned();
let api = ((*base).GetApi)(ort_sys::ORT_API_VERSION);
if api.is_null() {
return Err(format!(
"ONNX Runtime {version} is older than API version {}",
ort_sys::ORT_API_VERSION
));
}
if !ort::set_api((*api).clone()) {
return Err("an API table was already installed".into());
}
std::mem::forget(lib);
// Qualcomm's DSP loader finds the Hexagon skel through this variable,
// and only through it; the runtime's own directory is where the APK
// put it. Harmless anywhere else.
#[cfg(target_os = "android")]
if !dir.as_os_str().is_empty() {
std::env::set_var("ADSP_LIBRARY_PATH", dir);
}
log::info!("inference: ONNX Runtime {version} from {}", path.display());
Ok(Runtime::OnnxRuntime { path, version })
}
}
/// `libonnxruntime.so` in `dir`, or a versioned spelling of it —
/// `libonnxruntime.so.1.30.0` is what the Python wheel ships, and a package
/// that installs only the versioned file is not wrong.
#[cfg(feature = "native")]
fn find_library(dir: &std::path::Path, name: &str) -> Option<PathBuf> {
let exact = dir.join(name);
if exact.is_file() {
return Some(exact);
}
let prefix = format!("{name}.");
let mut versioned: Vec<PathBuf> = std::fs::read_dir(dir)
.ok()?
.filter_map(|e| e.ok())
.map(|e| e.path())
.filter(|p| {
p.is_file()
&& p.file_name()
.and_then(|n| n.to_str())
.is_some_and(|n| n.starts_with(&prefix))
})
.collect();
versioned.sort();
versioned.pop()
}
+121
View File
@@ -0,0 +1,121 @@
//! Compiled engines: what a rung builds once per device, and the thread that
//! builds them before anyone asks (docs/dev/inference.md §5, §6).
//!
//! TensorRT keeps its own engine cache keyed by graph hash; QNN writes a
//! context model. Both are opaque to this crate, which tracks only *that* a
//! model compiled — by the hash of its bytes — so [`crate::open`] can tell a
//! request whether to expect the rung or its fallback.
use std::path::PathBuf;
use crate::{state, Config, Form, Rung};
enum Source {
File(PathBuf),
Bytes(&'static [u8]),
}
/// 64-bit FNV-1a. A cache key, not a checksum: two model files that collide
/// here would have to also be the same size and the same role, and the cost
/// of that is a rebuilt engine.
pub fn hash(bytes: &[u8]) -> u64 {
let mut h = 0xcbf2_9ce4_8422_2325u64;
for &b in bytes {
h ^= b as u64;
h = h.wrapping_mul(0x0000_0100_0000_01b3);
}
h
}
/// The cache entry for `bytes` compiled on `rung`.
pub fn key(rung: Rung, bytes: &[u8]) -> String {
key_of(rung, hash(bytes))
}
/// The same, from a hash already taken.
pub fn key_of(rung: Rung, hash: u64) -> String {
format!("{}:{:016x}", rung.label(), hash)
}
/// Where QNN's compiled context for `bytes` lives.
pub fn context_path(cfg: &Config, bytes: &[u8]) -> PathBuf {
cfg.cache_dir
.join("qnn")
.join(format!("{:016x}_ctx.onnx", hash(bytes)))
}
/// After the probe: compile every configured model the selected rung can
/// take, smallest first, recording each as it lands.
pub fn run() {
let (rung, cfg) = {
let s = state().lock().unwrap();
(crate::current_rung(&s), s.config.clone())
};
if !rung.compiles() {
return;
}
// Smallest first, so the detector — the one that runs per image — is
// ready soonest (§6 step 3).
let mut jobs: Vec<(crate::Role, Source, u64)> = cfg
.models
.iter()
.filter(|(role, _)| rung.serves(*role))
.filter_map(|(role, path)| {
let (path, form) = crate::resolve_model(*role, path);
(form == rung.form(*role)).then(|| {
let size = std::fs::metadata(&path).map(|m| m.len()).unwrap_or(0);
(*role, Source::File(path), size)
})
})
.chain(cfg.embedded.iter().filter_map(|(role, bytes)| {
// An embedded model has no int8 sibling to offer a rung that
// wants one; it runs on that rung's fallback.
(rung.serves(*role) && rung.form(*role) == Form::F32).then_some((
*role,
Source::Bytes(bytes),
bytes.len() as u64,
))
}))
.collect();
jobs.sort_by_key(|j| j.2);
state().lock().unwrap().wanted = jobs.len();
for (role, source, _) in jobs {
let (bytes, name) = match &source {
Source::File(path) => match std::fs::read(path) {
Ok(b) => (b, path.display().to_string()),
Err(_) => continue,
},
Source::Bytes(b) => (b.to_vec(), format!("embedded {role:?}")),
};
let key = key(rung, &bytes);
if state().lock().unwrap().cache.compiled.contains(&key) {
continue;
}
log::info!("inference: compiling {name} for {}", rung.label());
let started = std::time::Instant::now();
match crate::session::build(rung, role, &bytes, &cfg) {
Ok(session) => {
drop(session);
let mut s = state().lock().unwrap();
s.cache.compiled.insert(key);
crate::probe::write_cache(&s.config, &s.cache);
log::info!(
"inference: {name} ready on {} in {:.1} s",
rung.label(),
started.elapsed().as_secs_f64()
);
}
Err(e) => {
// This model stays on the fallback; the others still get
// their engine. A corrected model file changes the hash and
// is retried.
log::warn!(
"inference: {name} will not compile for {}: {e}",
rung.label()
);
}
}
}
}
+665
View File
@@ -0,0 +1,665 @@
//! Which runtime, which provider and which model form — decided once per
//! device, and the only crate that knows the answer (docs/dev/inference.md).
//!
//! Consumers ask for a session by [`Role`] and get `ort`'s `Session` back;
//! what built it — tract on one core, ONNX Runtime's CPU pool, a TensorRT
//! engine, a MIGraphX program, the Hexagon — is this crate's business and
//! shows up in [`status`] for the settings row and nowhere else.
//!
//! The shape follows §3 of the spec: `ort` links nothing (`alternative-backend`),
//! and the first call hands it an API table from either a `libonnxruntime`
//! found on disk or from `tract`. That choice is once per process, because
//! `ort::set_api` is; everything after it — which provider, whether an engine
//! has been compiled yet — is per session and may change between two calls.
use std::collections::{BTreeSet, HashMap};
use std::path::{Path, PathBuf};
use std::sync::{Arc, Mutex, MutexGuard, OnceLock};
use std::time::{Duration, Instant};
use serde::{Deserialize, Serialize};
mod api;
mod engines;
mod probe;
mod session;
pub use api::Runtime;
pub use ort::session::Session;
/// What a model is for. The role fixes the precision rule (§7): an embedder
/// runs in f32 on every rung, a detector may run in fp16 or int8.
#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash, Serialize, Deserialize)]
pub enum Role {
Detector,
Embedder,
Segmenter,
Scene,
/// The dense landmark model behind the eye reading (docs/dev/faces.md §7c).
Landmarks,
/// The eye-state and sunglasses classifiers, a few hundred kilobytes.
EyeClassifier,
/// XFeat, the panorama keypoint detector (docs/dev/panorama.md).
Keypoints,
/// MI-GAN, the panorama border filler (docs/dev/panorama.md §12). Plain
/// convolutions, so any rung serves it; fp16 on TensorRT and int8 on
/// the Hexagon are the point of it.
Inpainter,
}
/// Which numeric form of a model a session was built from.
///
/// `Int8` is a different network from `F32` for a detector — it finds a
/// different set of faces — which is why [`form_suffix`] exists and why a
/// caller appends it to `model_id`.
#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash, Serialize, Deserialize)]
pub enum Form {
F32,
Int8,
}
/// A rung of the ladder (§2). Ordered: a user override names the highest rung
/// the probe may take, and a compiling rung falls back to the one below it
/// until its engine exists. The order is within a vendor's ladder — a
/// machine has NVIDIA rungs or an AMD rung, never both — so a ceiling is
/// read as "no higher than this on whichever ladder the device has".
#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash, PartialOrd, Ord, Serialize, Deserialize)]
pub enum Rung {
/// ONNX Runtime's CPU provider, or tract when no runtime file was found.
Cpu,
/// NVIDIA, through the CUDA provider. Desktop only.
Cuda,
/// NVIDIA, through a TensorRT engine compiled on this device. Desktop only.
TensorRt,
/// AMD, through a MIGraphX program compiled on this device. Desktop
/// only. ONNX Runtime's ROCm provider, the CUDA provider's twin, was
/// removed in ONNX Runtime 1.23, so there is no non-compiling AMD rung
/// to fall back to: this one falls back to the CPU.
MiGraphX,
/// Qualcomm's Hexagon NPU through QNN, int8 models only. Android only.
Hexagon,
}
impl Rung {
pub fn label(self) -> &'static str {
match self {
Rung::Cpu => "CPU",
Rung::Cuda => "CUDA",
Rung::TensorRt => "TensorRT",
Rung::MiGraphX => "MIGraphX",
Rung::Hexagon => "Hexagon NPU",
}
}
/// The rung a request lands on while this one's engine is still being
/// compiled (§6 step 2).
fn fallback(self) -> Rung {
match self {
Rung::TensorRt => Rung::Cuda,
Rung::MiGraphX | Rung::Hexagon | Rung::Cuda | Rung::Cpu => Rung::Cpu,
}
}
/// Whether a session on this rung needs an engine built first.
fn compiles(self) -> bool {
matches!(self, Rung::TensorRt | Rung::MiGraphX | Rung::Hexagon)
}
/// The model form this rung wants for a role.
fn form(self, _role: Role) -> Form {
match self {
Rung::Hexagon => Form::Int8,
_ => Form::F32,
}
}
/// Whether this rung runs `role` at all. The Hexagon takes int8 graphs
/// only, and the embedder is never int8 (§7) — it runs on the CPU
/// beside a detector on the NPU, so its vectors compare across devices.
fn serves(self, role: Role) -> bool {
match self {
Rung::Hexagon => role != Role::Embedder,
_ => true,
}
}
}
/// How long a session outlives its last use unless [`Config::decay`] says
/// otherwise: long enough for the next click, short enough that a session's
/// GPU or NPU memory does not sit under the develop view for long.
pub const DEFAULT_DECAY: Duration = Duration::from_secs(30);
/// What [`init`] is told once, at launch.
#[derive(Clone, Debug, Default)]
pub struct Config {
/// Where to look for `libonnxruntime`, in order. An empty path means "the
/// bare library name through the system loader", which is how the APK's
/// own copy is found on Android.
pub runtime_dirs: Vec<PathBuf>,
/// Probe cache and compiled engines (§4, §5). Disposable.
pub cache_dir: PathBuf,
/// The canonical model files on this device, so engines can be compiled
/// ahead of the first request for them.
pub models: Vec<(Role, PathBuf)>,
/// Models compiled into the binary, for the same reason.
pub embedded: Vec<(Role, &'static [u8])>,
/// The highest rung the user allows; `None` is "the best that works".
pub ceiling: Option<Rung>,
/// ONNX Runtime's intra-op pool; 0 picks from the core count.
pub threads: usize,
/// How long an unused session stays loaded. Zero means the default.
pub decay: Duration,
}
/// One line for the settings row, and the numbers behind the progress row.
#[derive(Clone, Debug)]
pub struct Status {
pub runtime: Runtime,
/// The rung selected, or the floor while the probe is still running.
pub rung: Rung,
/// Why — "probe passed", or the failure that demoted the rung above.
pub reason: String,
pub probing: bool,
/// Engines compiled and engines wanted, for a compiling rung; `(0, 0)`
/// otherwise.
pub engines: (usize, usize),
/// Every rung above the selected one that was tried, and why it lost.
pub failed: Vec<(Rung, String)>,
}
impl Status {
/// "Hexagon NPU · int8 · ONNX Runtime 1.29" — the settings row's text.
pub fn line(&self) -> String {
let form = match self.rung {
Rung::Hexagon => " · int8",
Rung::TensorRt | Rung::MiGraphX => " · fp16",
_ => "",
};
format!("{}{} · {}", self.rung.label(), form, self.runtime.label())
}
}
/// A model the caller can run, whatever is or is not loaded right now.
///
/// Holds the bytes, not a session. [`Model::acquire`] finds the loaded copy
/// in the registry — shared with every other holder of the same model —
/// or loads one, and every acquire refreshes the copy's last-used time.
/// The reaper unloads anything idle for [`Config::decay`]; a scan that runs
/// the detector on every image never lets it go idle, a click in the
/// develop view lets the segmenter go after a quiet spell, and a handle
/// used again after that simply loads again. Nobody states a policy.
///
/// The registry key includes the rung, so a reload after a compiled engine
/// has landed moves up to it by itself (§6 step 4).
pub struct Model {
role: Role,
form: Form,
bytes: Arc<[u8]>,
/// `engines::hash` of the bytes, taken once: an acquire per tile of a
/// border fill must not hash 28 MB each time.
hash: u64,
}
/// A loaded session, held for one `run` and its output decoding.
pub struct Acquired {
entry: Arc<Loaded>,
}
struct Loaded {
rung: Rung,
session: Mutex<Session>,
last_used: Mutex<Instant>,
}
impl Model {
/// The loaded session, loading it if the reaper took it. Lock it for
/// one run; a scan and a develop click can want the same detector at
/// once, and the second waits on the first.
pub fn acquire(&self) -> Result<Acquired, Error> {
acquire(self.role, self.form, &self.bytes, self.hash)
}
pub fn form(&self) -> Form {
self.form
}
}
impl Acquired {
pub fn lock(&self) -> MutexGuard<'_, Session> {
self.entry.session.lock().unwrap_or_else(|e| e.into_inner())
}
/// Where this session runs.
pub fn rung(&self) -> Rung {
self.entry.rung
}
}
impl Drop for Acquired {
fn drop(&mut self) {
// The clock starts when the use ends, not when it began: a long run
// is not idle time.
*self.entry.last_used.lock().unwrap() = Instant::now();
}
}
type Registry = HashMap<String, Arc<Loaded>>;
static REGISTRY: OnceLock<Mutex<Registry>> = OnceLock::new();
fn registry() -> &'static Mutex<Registry> {
REGISTRY.get_or_init(|| {
std::thread::Builder::new()
.name("inference-reaper".into())
.spawn(|| loop {
std::thread::sleep(Duration::from_secs(5));
release_idle();
})
.expect("spawn inference reaper");
Mutex::new(HashMap::new())
})
}
fn acquire(role: Role, form: Form, bytes: &Arc<[u8]>, hash: u64) -> Result<Acquired, Error> {
api::ensure_installed();
let (rung, cfg) = {
let s = state().lock().unwrap();
let selected = current_rung(&s);
(
effective_rung(&s, selected, role, form, hash),
s.config.clone(),
)
};
let key = format!("{role:?}:{}", engines::key_of(rung, hash));
if let Some(entry) = registry().lock().unwrap().get(&key).cloned() {
*entry.last_used.lock().unwrap() = Instant::now();
return Ok(Acquired { entry });
}
// Built outside the registry lock: a TensorRT or MIGraphX engine load
// is long enough that another role's acquire should not wait on it.
let session = session::build(rung, role, bytes, &cfg)?;
log::debug!("inference: {role:?} loaded on {}", rung.label());
let entry = Arc::new(Loaded {
rung,
session: Mutex::new(session),
last_used: Mutex::new(Instant::now()),
});
let mut reg = registry().lock().unwrap();
// Two acquires raced; keep the first, drop this one.
let entry = reg.entry(key).or_insert_with(|| entry.clone()).clone();
Ok(Acquired { entry })
}
/// Unload every session idle for longer than the decay. The reaper does
/// this every five seconds. A session in use survives until its run ends:
/// the `Acquired` holds it, the registry merely forgets it.
pub fn release_idle() {
let decay = match state().lock().unwrap().config.decay {
Duration::ZERO => DEFAULT_DECAY,
d => d,
};
let now = Instant::now();
registry()
.lock()
.unwrap()
.retain(|_, e| now.duration_since(*e.last_used.lock().unwrap()) < decay);
}
/// Unload every session now, decay or not — what a low-memory signal
/// asks for. Sessions mid-run finish first.
pub fn release_all() {
registry().lock().unwrap().clear();
}
/// Unload every session of `role` now — "I am done segmenting".
pub fn unload(role: Role) {
let prefix = format!("{role:?}:");
registry()
.lock()
.unwrap()
.retain(|k, _| !k.starts_with(&prefix));
}
/// How many sessions are loaded, for the settings row and the tests.
pub fn loaded() -> usize {
registry().lock().unwrap().len()
}
#[derive(Debug, thiserror::Error)]
pub enum Error {
#[error(transparent)]
Inference(#[from] ort::Error),
#[error("reading model: {0}")]
Io(#[from] std::io::Error),
}
/// What the probe writes and the next launch reads (§4 step 3).
#[derive(Clone, Debug, Default, Serialize, Deserialize)]
struct Cache {
/// Runtime, driver, hardware and model identity; any change re-probes.
fingerprint: String,
rung: Option<Rung>,
reason: String,
/// Model hashes whose engine exists on disk, per compiling rung.
compiled: BTreeSet<String>,
/// Rungs that failed under this fingerprint, and why. Not retried until
/// the fingerprint changes: a wedged driver must not cost every launch
/// thirty seconds.
failed: Vec<(Rung, String)>,
}
struct State {
config: Config,
cache: Cache,
probing: bool,
wanted: usize,
}
static STATE: OnceLock<Mutex<State>> = OnceLock::new();
fn state() -> &'static Mutex<State> {
STATE.get_or_init(|| {
Mutex::new(State {
config: Config::default(),
cache: Cache::default(),
probing: false,
wanted: 0,
})
})
}
/// Choose the runtime and start the probe. Idempotent; the first call wins.
///
/// Returns at once: the probe and any engine compilation run on their own
/// low-priority thread, and every request meanwhile is served by the floor
/// (§4). Never blocks the first frame.
pub fn init(config: Config) {
let runtime = api::install(&config.runtime_dirs);
{
let mut s = state().lock().unwrap();
if s.probing || s.cache.rung.is_some() {
return;
}
s.config = config;
s.probing = true;
}
log::info!("inference: runtime {}", runtime.label());
std::thread::Builder::new()
.name("inference-probe".into())
.spawn(move || {
probe::run(runtime);
engines::run();
})
.expect("spawn inference probe");
}
/// Make sure `ort` has an API table, for code that drives `ort` directly.
/// [`open`] does this itself; only the M1 probe example needs it by name.
pub fn ensure_runtime() {
api::ensure_installed();
}
/// The line for the settings row.
pub fn status() -> Status {
let s = state().lock().unwrap();
let rung = current_rung(&s);
Status {
runtime: api::runtime(),
rung,
reason: s.cache.reason.clone(),
// Only what explains the selection: on an AMD machine the NVIDIA
// rungs "not enabled in this build" say nothing about why MIGraphX
// was taken. With the floor selected, everything tried is above it.
failed: s
.cache
.failed
.iter()
.filter(|(r, _)| *r > rung)
.cloned()
.collect(),
probing: s.probing,
engines: if rung.compiles() {
(s.cache.compiled.len(), s.wanted)
} else {
(0, 0)
},
}
}
fn current_rung(s: &State) -> Rung {
if s.probing {
Rung::Cpu
} else {
s.cache.rung.unwrap_or(Rung::Cpu)
}
}
/// The file to load for `role` under the current selection, and its form.
///
/// A rung that wants int8 gets the `.int8.onnx` sibling of the canonical file
/// if it exists; otherwise the canonical file, on the rung's fallback. A
/// caller adds [`form_suffix`] to the `model_id` it records.
pub fn resolve_model(role: Role, canonical: &Path) -> (PathBuf, Form) {
let rung = current_rung(&state().lock().unwrap());
if rung.serves(role) && rung.form(role) == Form::Int8 {
let sibling = int8_sibling(canonical);
if sibling.is_file() {
return (sibling, Form::Int8);
}
}
(canonical.to_path_buf(), Form::F32)
}
fn int8_sibling(canonical: &Path) -> PathBuf {
let stem = canonical
.file_stem()
.map(|s| s.to_string_lossy().into_owned())
.unwrap_or_default();
canonical.with_file_name(format!("{stem}.int8.onnx"))
}
/// What a form appends to a detector's `model_id` (§7).
pub fn form_suffix(form: Form) -> &'static str {
match form {
Form::F32 => "",
Form::Int8 => "_i8",
}
}
/// A handle on the model `bytes` in `role`.
///
/// Loads it once here, so a graph the runtime rejects fails at
/// construction and not on the first image; what happens to that session
/// afterwards is the registry's business (see [`Model`]).
///
/// Works without [`init`] — a test, or the examples — by installing tract
/// and using the CPU rung, which is exactly what every consumer did before
/// this crate existed.
pub fn open(role: Role, form: Form, bytes: &[u8]) -> Result<Model, Error> {
let bytes: Arc<[u8]> = Arc::from(bytes);
let hash = engines::hash(&bytes);
acquire(role, form, &bytes, hash)?;
Ok(Model {
role,
form,
bytes,
hash,
})
}
/// Where a request lands: the selected rung unless the role's precision rule,
/// the form on offer, or a missing engine says one lower (§6 step 4).
fn effective_rung(s: &State, selected: Rung, role: Role, form: Form, hash: u64) -> Rung {
let mut rung = selected;
if !rung.serves(role) || rung.form(role) != form {
// The embedder on a Hexagon device, or an f32 detector where the int8
// sibling was missing: neither can go to the NPU.
rung = rung.fallback();
}
if rung.compiles() && !s.cache.compiled.contains(&engines::key_of(rung, hash)) {
rung = rung.fallback();
}
rung
}
#[cfg(test)]
mod tests {
use super::*;
/// The registry is one per process, so these run one at a time.
static SERIAL: Mutex<()> = Mutex::new(());
fn serial() -> MutexGuard<'static, ()> {
SERIAL.lock().unwrap_or_else(|e| e.into_inner())
}
/// The smallest shipped graph, if this checkout has the weights; a test
/// suite that needs a research-licensed download is one that does not
/// run in CI (docs/dev/faces.md §3), so absence is a skip.
fn probe_bytes() -> Option<Vec<u8>> {
let path = concat!(
env!("CARGO_MANIFEST_DIR"),
"/../../models/face/scrfd_500m_640.onnx"
);
let bytes = std::fs::read(path).ok()?;
(bytes.len() > 100_000).then_some(bytes)
}
#[test]
fn two_handles_on_one_model_share_one_session() {
let _serial = serial();
let Some(bytes) = probe_bytes() else { return };
release_all();
let a = open(Role::Detector, Form::F32, &bytes).unwrap();
let b = open(Role::Detector, Form::F32, &bytes).unwrap();
assert_eq!(loaded(), 1);
let (x, y) = (a.acquire().unwrap(), b.acquire().unwrap());
assert!(Arc::ptr_eq(&x.entry, &y.entry));
}
#[test]
fn a_released_model_reloads_on_its_next_use() {
let _serial = serial();
let Some(bytes) = probe_bytes() else { return };
release_all();
let model = open(Role::Detector, Form::F32, &bytes).unwrap();
assert_eq!(loaded(), 1);
release_all();
assert_eq!(loaded(), 0);
let acquired = model.acquire().unwrap();
assert_eq!(loaded(), 1);
assert_eq!(acquired.lock().inputs().len(), 1);
}
#[test]
fn an_idle_session_decays_and_a_used_one_does_not() {
let _serial = serial();
let Some(bytes) = probe_bytes() else { return };
release_all();
state().lock().unwrap().config.decay = Duration::from_millis(50);
let model = open(Role::Detector, Form::F32, &bytes).unwrap();
// Used within the decay: stays.
std::thread::sleep(Duration::from_millis(30));
drop(model.acquire().unwrap());
release_idle();
assert_eq!(loaded(), 1);
// Idle past it: goes.
std::thread::sleep(Duration::from_millis(80));
release_idle();
assert_eq!(loaded(), 0);
state().lock().unwrap().config.decay = Duration::ZERO;
}
#[test]
fn unload_by_role_leaves_the_other_roles() {
let _serial = serial();
let Some(bytes) = probe_bytes() else { return };
release_all();
let _d = open(Role::Detector, Form::F32, &bytes).unwrap();
let _s = open(Role::Segmenter, Form::F32, &bytes).unwrap();
assert_eq!(loaded(), 2);
unload(Role::Segmenter);
assert_eq!(loaded(), 1);
}
#[test]
fn the_hexagon_never_takes_the_embedder() {
assert!(!Rung::Hexagon.serves(Role::Embedder));
assert!(Rung::Hexagon.serves(Role::Detector));
assert_eq!(Rung::Hexagon.form(Role::Detector), Form::Int8);
// A detector offered in f32 on a Hexagon device lands on the CPU.
let s = State {
config: Config::default(),
cache: Cache {
rung: Some(Rung::Hexagon),
..Cache::default()
},
probing: false,
wanted: 0,
};
assert_eq!(
effective_rung(
&s,
Rung::Hexagon,
Role::Embedder,
Form::F32,
engines::hash(b"")
),
Rung::Cpu
);
assert_eq!(
effective_rung(
&s,
Rung::Hexagon,
Role::Detector,
Form::F32,
engines::hash(b"")
),
Rung::Cpu
);
// An int8 detector whose context is not compiled yet: also the CPU.
assert_eq!(
effective_rung(
&s,
Rung::Hexagon,
Role::Detector,
Form::Int8,
engines::hash(b"")
),
Rung::Cpu
);
}
#[test]
fn the_status_reports_only_the_rungs_above_the_selection() {
let _serial = serial();
let failed = vec![
(Rung::TensorRt, "not enabled".to_string()),
(Rung::Cuda, "not enabled".to_string()),
];
let before = state().lock().unwrap().cache.clone();
state().lock().unwrap().cache = Cache {
rung: Some(Rung::MiGraphX),
failed: failed.clone(),
..Cache::default()
};
// An AMD desktop: the NVIDIA rungs below MIGraphX are not the story.
assert!(status().failed.is_empty());
// An NVIDIA desktop on the CUDA provider: TensorRT's failure is.
state().lock().unwrap().cache.rung = Some(Rung::Cuda);
assert_eq!(status().failed, vec![failed[0].clone()]);
// The floor: everything tried explains it.
state().lock().unwrap().cache.rung = Some(Rung::Cpu);
assert_eq!(status().failed.len(), 2);
state().lock().unwrap().cache = before;
}
#[test]
fn the_status_line_reads_as_the_floor_before_init() {
let _serial = serial();
let s = status();
assert_eq!(s.rung, Rung::Cpu);
assert!(s.line().starts_with("CPU"), "{}", s.line());
}
}
+340
View File
@@ -0,0 +1,340 @@
//! Walk the ladder, once, by building real sessions (docs/dev/inference.md §4).
//!
//! A rung is taken when a session builds on it, runs, and is faster than
//! the floor. Both halves matter: a provider can register and then fail at
//! partition time, and a provider can take a graph — or quietly hand most
//! of it back to the CPU — and run it slower than the CPU would have. The outcome is cached against a fingerprint of the
//! runtime, the driver, the hardware and the models, and trusted until any
//! of those changes.
use std::path::{Path, PathBuf};
use std::time::Instant;
use crate::{api::Runtime, state, Cache, Config, Form, Role, Rung};
/// The rungs to try on this platform, best first, under the user's ceiling.
fn ladder(ceiling: Option<Rung>) -> Vec<Rung> {
#[cfg(target_os = "android")]
let all = [Rung::Hexagon];
// A desktop has one vendor's GPU; the other vendor's providers are
// "not enabled in this build" or a library that fails to load, and
// either answer arrives in milliseconds.
#[cfg(not(target_os = "android"))]
let all = [Rung::TensorRt, Rung::Cuda, Rung::MiGraphX];
all.into_iter()
.filter(|r| ceiling.is_none_or(|c| *r <= c))
.collect()
}
/// The probe body. Sets the cache and clears `probing` when done; never
/// panics out, because a failed probe is a result (the floor) and not an
/// error.
pub fn run(runtime: Runtime) {
let cfg = state().lock().unwrap().config.clone();
let fingerprint = fingerprint(&runtime, &cfg);
if let Some(cached) = read_cache(&cfg) {
if cached.fingerprint == fingerprint && cached.rung.is_some() {
log::info!(
"inference: cached selection {} ({})",
cached.rung.unwrap().label(),
cached.reason
);
finish(cached);
return;
}
}
let mut cache = Cache {
fingerprint,
..Cache::default()
};
if !runtime.is_native() {
cache.rung = Some(Rung::Cpu);
cache.reason = "no ONNX Runtime found; tract on one core".into();
write_cache(&cfg, &cache);
finish(cache);
return;
}
let Some((role, canonical)) = probe_model(&cfg) else {
cache.rung = Some(Rung::Cpu);
cache.reason = "no model to probe with".into();
write_cache(&cfg, &cache);
finish(cache);
return;
};
let floor = match time_rung(Rung::Cpu, role, &canonical, &cfg) {
Ok((ms, _)) => ms,
Err(e) => {
// The CPU provider failing is the runtime failing; there is
// nothing below it to try, and the reason is worth reading.
cache.rung = Some(Rung::Cpu);
cache.reason = format!("CPU provider failed: {e}");
write_cache(&cfg, &cache);
finish(cache);
return;
}
};
log::info!("inference: floor {floor:.1} ms on the CPU provider");
for rung in ladder(cfg.ceiling) {
match time_rung(rung, role, &canonical, &cfg) {
Ok((ms, key)) if ms < floor => {
cache.rung = Some(rung);
cache.reason = format!("{ms:.1} ms against {floor:.1} ms on the CPU");
if let Some(key) = key {
cache.compiled.insert(key);
}
break;
}
Ok((ms, _)) => {
let why = format!("{ms:.1} ms, slower than the CPU's {floor:.1} ms");
log::info!("inference: {} rejected: {why}", rung.label());
cache.failed.push((rung, why));
}
Err(e) => {
log::info!("inference: {} failed: {e}", rung.label());
cache.failed.push((rung, e));
}
}
}
if cache.rung.is_none() {
cache.rung = Some(Rung::Cpu);
cache.reason = match cache.failed.first() {
Some((r, why)) => format!("{} {}", r.label(), first_line(why)),
None => "the only rung on this platform".into(),
};
}
write_cache(&cfg, &cache);
finish(cache);
}
fn finish(cache: Cache) {
let mut s = state().lock().unwrap();
s.cache = cache;
s.probing = false;
}
/// The smallest detector, or the smallest model of any role if there is
/// none. A ~2 MB detector is the cheapest real test of a provider, and the
/// detector is the role the int8 forms exist for — the eye classifiers are
/// smaller still, and a Hexagon probed with one would fail for want of a
/// form nobody ships.
fn probe_model(cfg: &Config) -> Option<(Role, PathBuf)> {
let smallest = |want: Option<Role>| {
cfg.models
.iter()
.filter(|(role, _)| want.is_none_or(|w| *role == w))
.filter_map(|(role, path)| {
let size = std::fs::metadata(path).ok()?.len();
Some((size, *role, path.clone()))
})
.min_by_key(|(size, _, _)| *size)
.map(|(_, role, path)| (role, path))
};
smallest(Some(Role::Detector)).or_else(|| smallest(None))
}
/// Build, run once for the engine, then time three runs; the median in
/// milliseconds and, for a compiling rung, the cache key of the engine this
/// just built.
fn time_rung(
rung: Rung,
role: Role,
canonical: &Path,
cfg: &Config,
) -> Result<(f64, Option<String>), String> {
let want = rung.form(role);
let path = match want {
Form::Int8 => {
let p = crate::int8_sibling(canonical);
if !p.is_file() {
return Err(format!("no int8 form of {}", canonical.display()));
}
p
}
Form::F32 => canonical.to_path_buf(),
};
let bytes = std::fs::read(&path).map_err(|e| e.to_string())?;
let started = Instant::now();
let mut session =
crate::session::build(rung, role, &bytes, cfg).map_err(|e| first_line(&e.to_string()))?;
log::info!(
"inference: {} session built in {:.1} s",
rung.label(),
started.elapsed().as_secs_f64()
);
let shape: Vec<usize> = session.inputs()[0]
.dtype()
.tensor_shape()
.ok_or("model input is not a tensor")?
.iter()
.map(|&d| if d > 0 { d as usize } else { 1 })
.collect();
let zeros = vec![0f32; shape.iter().product()];
let run = |session: &mut ort::session::Session| -> Result<f64, String> {
let input = ort::value::Tensor::from_array((shape.clone(), zeros.clone()))
.map_err(|e| e.to_string())?;
let t = Instant::now();
let out = session
.run(ort::inputs![input])
.map_err(|e| e.to_string())?;
let _ = out[0]
.try_extract_tensor::<f32>()
.map_err(|e| e.to_string())?;
Ok(t.elapsed().as_secs_f64() * 1e3)
};
run(&mut session)?;
let mut times = [run(&mut session)?, run(&mut session)?, run(&mut session)?];
times.sort_by(|a, b| a.partial_cmp(b).unwrap());
let key = rung.compiles().then(|| crate::engines::key(rung, &bytes));
Ok((times[1], key))
}
/// The part of a provider's error a person can act on. ONNX Runtime's
/// begin with a source path and a C++ template signature; the words —
/// "CUDA failure 999: unknown error", "FAIL : Failed to load library" —
/// come after, and the settings row has room for one line of them.
fn first_line(s: &str) -> String {
let line = s.lines().next().unwrap_or("");
let start = ["failure", "FAIL :", "Error:", "error:"]
.iter()
.filter_map(|m| line.find(m))
.min()
.unwrap_or(0);
line[start..].chars().take(200).collect()
}
/// Everything a change of which should re-probe: the runtime, where it
/// came from and which providers sit beside it, this crate, the platform,
/// the driver or SoC, and the models.
fn fingerprint(runtime: &Runtime, cfg: &Config) -> String {
let mut parts = vec![
format!("engine {}", env!("CARGO_PKG_VERSION")),
format!("{} {}", std::env::consts::OS, std::env::consts::ARCH),
match runtime {
Runtime::Tract => "tract".to_string(),
Runtime::OnnxRuntime { path, version } => {
format!(
"ort {version} {} [{}]",
path.display(),
providers_beside(path)
)
}
},
device_identity(),
];
for (role, bytes) in &cfg.embedded {
parts.push(format!(
"{role:?} embedded {:016x}",
crate::engines::hash(bytes)
));
}
for (role, path) in &cfg.models {
let hash = std::fs::read(path)
.map(|b| crate::engines::hash(&b))
.unwrap_or(0);
parts.push(format!("{role:?} {hash:016x}"));
let int8 = crate::int8_sibling(path);
if let Ok(b) = std::fs::read(&int8) {
parts.push(format!("{role:?} int8 {:016x}", crate::engines::hash(&b)));
}
}
parts.join("\n")
}
/// The `libonnxruntime_providers_*.so` files in the runtime's directory.
/// A distribution's CPU-only and ROCm builds are the same version at the
/// same path; the provider libraries beside them are what differs.
fn providers_beside(runtime: &Path) -> String {
let Some(dir) = runtime.parent() else {
return String::new();
};
let mut names: Vec<String> = std::fs::read_dir(dir)
.into_iter()
.flatten()
.filter_map(|e| e.ok())
.filter_map(|e| e.file_name().into_string().ok())
.filter(|n| {
n.starts_with("libonnxruntime_providers_") || n.starts_with("onnxruntime_providers_")
})
.collect();
names.sort();
names.join(" ")
}
#[cfg(target_os = "linux")]
fn device_identity() -> String {
// The NVIDIA driver's version line, or the ROCm release the AMD stack
// came from (`rocm-core` writes it; the kernel driver has no version
// of its own). Absent means neither.
if let Some(line) = std::fs::read_to_string("/proc/driver/nvidia/version")
.ok()
.and_then(|s| s.lines().next().map(str::to_string))
{
return line;
}
if let Ok(rocm) = std::fs::read_to_string("/opt/rocm/.info/version") {
return format!("rocm {}", rocm.trim());
}
"no nvidia driver, no rocm".into()
}
#[cfg(target_os = "android")]
fn device_identity() -> String {
// The SoC and the vendor's build: a Hexagon appears or disappears with
// either.
format!(
"{} {}",
system_property("ro.soc.model"),
system_property("ro.build.version.incremental")
)
}
#[cfg(target_os = "android")]
fn system_property(name: &str) -> String {
extern "C" {
fn __system_property_get(
name: *const std::ffi::c_char,
value: *mut std::ffi::c_char,
) -> i32;
}
let name = std::ffi::CString::new(name).unwrap();
let mut buf = [0u8; 92]; // PROP_VALUE_MAX
// SAFETY: bionic's documented call; the buffer is PROP_VALUE_MAX bytes.
let n = unsafe { __system_property_get(name.as_ptr(), buf.as_mut_ptr().cast()) };
String::from_utf8_lossy(&buf[..n.max(0) as usize]).into_owned()
}
#[cfg(not(any(target_os = "linux", target_os = "android")))]
fn device_identity() -> String {
String::new()
}
fn cache_path(cfg: &Config) -> PathBuf {
cfg.cache_dir.join("backend.json")
}
fn read_cache(cfg: &Config) -> Option<Cache> {
let text = std::fs::read_to_string(cache_path(cfg)).ok()?;
serde_json::from_str(&text).ok()
}
/// Written whole and renamed into place, so a reader never sees half.
pub fn write_cache(cfg: &Config, cache: &Cache) {
if cfg.cache_dir.as_os_str().is_empty() {
return;
}
let path = cache_path(cfg);
let tmp = path.with_extension("json.tmp");
let _ = std::fs::create_dir_all(&cfg.cache_dir);
if let Ok(text) = serde_json::to_string_pretty(cache) {
if std::fs::write(&tmp, text).is_ok() {
let _ = std::fs::rename(&tmp, &path);
}
}
}
+182
View File
@@ -0,0 +1,182 @@
//! One session builder per rung (docs/dev/inference.md §2, §7, §9).
use ort::session::Session;
use crate::{Config, Role, Rung};
/// Build a session for `bytes` on `rung`.
///
/// Not strict about the CPU: `session.disable_cpu_ep_fallback` was tried as
/// the probe's proof that a provider took the graph, and it refuses the
/// Hexagon over the ten quantise/dequantise nodes at the graph's edges that
/// QNN declines by policy and that cost microseconds. The probe's proof is
/// its clock instead (§4): a provider that hands real work to the CPU is
/// slower than the CPU floor and rejected by the same measurement.
pub fn build(rung: Rung, role: Role, bytes: &[u8], cfg: &Config) -> ort::Result<Session> {
// No optimisation level named. ONNX Runtime's default is already its
// fullest, and on tract any level but "disabled" means `into_optimized`,
// whose optimiser divides by zero inside yolo26n-seg (tract-data
// `stack_tensors`) — a panic across the C API, which is an abort. The
// app never asked tract for that and does not start now.
let mut b = Session::builder()?.with_intra_threads(threads(cfg))?;
// A Hexagon session loads the compiled context when there is one and
// compiles it from the model when there is not; the engine thread is
// what makes the second case rare (§6).
let context = (rung == Rung::Hexagon).then(|| crate::engines::context_path(cfg, bytes));
let ready = context.as_ref().is_some_and(|p| p.is_file());
b = providers(
b,
rung,
role,
cfg,
if ready { None } else { context.as_deref() },
)?;
match (ready, context) {
(true, Some(path)) => b.commit_from_file(path),
_ => b.commit_from_memory(bytes),
}
}
/// The intra-op pool: what the config says, else the cores less two for
/// the compositor and the decoder (§9). tract ignores it.
fn threads(cfg: &Config) -> usize {
if cfg.threads > 0 {
return cfg.threads;
}
std::thread::available_parallelism()
.map(|n| n.get().saturating_sub(2).max(1))
.unwrap_or(1)
}
#[cfg(not(target_os = "android"))]
fn providers(
b: ort::session::builder::SessionBuilder,
rung: Rung,
role: Role,
cfg: &Config,
_generate_context: Option<&std::path::Path>,
) -> ort::Result<ort::session::builder::SessionBuilder> {
use ort::ep;
match rung {
Rung::Cpu => Ok(b),
Rung::Cuda => {
Ok(b.with_execution_providers([ep::CUDA::default().build().error_on_failure()])?)
}
Rung::TensorRt => {
let cache = cfg.cache_dir.join("tensorrt");
let _ = std::fs::create_dir_all(&cache);
let cache = cache.to_string_lossy().into_owned();
// fp16 for everything but the embedder, whose comparability
// across devices is worth more than its 0.2 ms (§7). The
// workspace cap keeps the develop view's tiles on the card
// (NFR-RES-2). CUDA behind it takes any node TensorRT declines.
Ok(b.with_execution_providers([
ep::TensorRT::default()
.with_fp16(role != Role::Embedder)
.with_engine_cache(true)
.with_engine_cache_path(&cache)
.with_timing_cache(true)
.with_timing_cache_path(&cache)
.with_max_workspace_size(512 << 20)
.build()
.error_on_failure(),
ep::CUDA::default().build(),
])?)
}
Rung::MiGraphX => {
// fp16 on the same terms as TensorRT (§7). MIGraphX compiles a
// program per graph — 20–60 s here — and keeps it in the cache
// directory, keyed on the graph, the GPU and its own version
// but not the precision: hence one directory per precision.
// The CPU takes any node it declines.
let fp16 = role != Role::Embedder;
let cache = cfg
.cache_dir
.join("migraphx")
.join(if fp16 { "fp16" } else { "f32" });
let _ = std::fs::create_dir_all(&cache);
let mut b = b;
migraphx(&mut b, fp16, &cache)?;
Ok(b)
}
Rung::Hexagon => unreachable!("the Hexagon rung is not on a desktop ladder"),
}
}
/// Register MIGraphX through ONNX Runtime's generic key/value entry point.
///
/// `ort`'s own builder (`ep::MIGraphX`) fills the legacy
/// `OrtMIGraphXProviderOptions`, and 1.29 reads that struct for its
/// precision flags and nothing else — the compiled-program cache directory
/// is only a key in the generic map (`migraphx_model_cache_dir`), and
/// without it every session is a full compile. Registration through the
/// generic entry point needs no `ort` feature: it is one call on the API
/// table, which is why the crate's `ort` dependency names no AMD feature.
#[cfg(not(target_os = "android"))]
fn migraphx(
b: &mut ort::session::builder::SessionBuilder,
fp16: bool,
cache: &std::path::Path,
) -> ort::Result<()> {
use ort::AsPointer;
use std::ffi::CString;
let keys = [c"migraphx_fp16_enable", c"migraphx_model_cache_dir"];
let values = [
CString::new(if fp16 { "1" } else { "0" }).unwrap(),
CString::new(cache.to_string_lossy().as_bytes())
.map_err(|e| ort::Error::new(e.to_string()))?,
];
let key_ptrs: Vec<_> = keys.iter().map(|k| k.as_ptr()).collect();
let value_ptrs: Vec<_> = values.iter().map(|v| v.as_ptr()).collect();
// SAFETY: the documented C call over arrays that outlive it; the
// runtime copies the strings into its own options map before returning.
unsafe {
let status = (ort::api().SessionOptionsAppendExecutionProvider)(
b.ptr_mut(),
c"MIGraphX".as_ptr(),
key_ptrs.as_ptr(),
value_ptrs.as_ptr(),
keys.len(),
);
ort::Error::result_from_status(status)
}
}
#[cfg(target_os = "android")]
fn providers(
b: ort::session::builder::SessionBuilder,
rung: Rung,
_role: Role,
_cfg: &Config,
generate_context: Option<&std::path::Path>,
) -> ort::Result<ort::session::builder::SessionBuilder> {
use ort::ep;
match rung {
Rung::Cpu => Ok(b),
Rung::Hexagon => {
// The HTP compiles the graph once per device (0.8–1.7 s here).
// With `ep.context_enable` ONNX Runtime writes the compiled
// context beside the probe cache; the next session loads that
// file as its model and skips the compile (§5).
let mut b = b;
if let Some(ctx) = generate_context {
let _ = std::fs::create_dir_all(ctx.parent().unwrap());
b = b
.with_config_entry("ep.context_enable", "1")?
.with_config_entry("ep.context_file_path", ctx.to_string_lossy())?
.with_config_entry("ep.context_embed_mode", "0")?;
}
// Quantise/dequantise at the graph's edges stay on the NPU too,
// so a strict build is a whole-graph build.
Ok(b.with_execution_providers([ep::QNN::default()
.with_backend_path("libQnnHtp.so")
.with_performance_mode(ep::qnn::PerformanceMode::Burst)
.with_offload_graph_io_quantization(false)
.build()
.error_on_failure()])?)
}
Rung::Cuda | Rung::TensorRt | Rung::MiGraphX => {
unreachable!("no desktop GPU rung on Android")
}
}
}
+38
View File
@@ -0,0 +1,38 @@
[package]
name = "dr-pano"
version.workspace = true
edition.workspace = true
rust-version.workspace = true
license.workspace = true
# Guards against a Git LFS pointer being embedded in place of the weights.
build = "build.rs"
[dependencies]
thiserror.workspace = true
log.workspace = true
# Inference for the learned keypoint detector, on the same footing as
# `dr-segment`: `ort` is the API, `dr-inference-engine` decides what runs
# it (docs/dev/inference.md), and both are optional so that the geometry —
# matching, the rotation solve, the projections — is a dependency-free crate
# that tests without a model.
ort = { workspace = true, optional = true }
dr-inference-engine = { workspace = true, optional = true }
ndarray = { workspace = true, optional = true }
[dev-dependencies]
# The example aligns real frames from their embedded previews.
dr-decode.workspace = true
dr-types.workspace = true
env_logger.workspace = true
[features]
default = ["xfeat", "embedded-model"]
# The XFeat detector (FR-MRG-8) and the MI-GAN filler (FR-MRG-4). Off, the
# crate has no model and no runtime — a build that only wants the geometry.
xfeat = ["dep:ort", "dep:dr-inference-engine", "dep:ndarray"]
# Compile the weights into the binary, for the same reason `dr-segment` does:
# Android hands the app no path to read a model from (ARCH §6.9).
embedded-model = ["xfeat"]
+48
View File
@@ -0,0 +1,48 @@
//! Check the model is a model and not an LFS pointer.
//!
//! `models/keypoints/*.onnx` is stored in Git LFS (see `.gitattributes`). A
//! clone made without git-lfs, or with `GIT_LFS_SKIP_SMUDGE` set, leaves a
//! ~130-byte text pointer at that path instead of the weights, and
//! `include_bytes!` would embed it without complaint. Same guard as
//! `dr-segment`'s, for the same failure.
use std::path::Path;
const MODELS: &[&str] = &[
"../../models/keypoints/xfeat-1024.onnx",
"../../models/keypoints/xfeat-768.onnx",
];
fn main() {
for m in MODELS {
println!("cargo:rerun-if-changed={m}");
}
println!("cargo:rerun-if-changed=build.rs");
if std::env::var_os("CARGO_FEATURE_EMBEDDED_MODEL").is_none() {
return;
}
for model in MODELS.iter().copied() {
check(model);
}
}
fn check(model: &str) {
let path = Path::new(model);
let Ok(bytes) = std::fs::read(path) else {
panic!(
"\n\n{model} is missing.\n\
It ships in Git LFS. Run `git lfs install && git lfs pull`, or build \
with `--no-default-features` for a geometry-only build.\n"
);
};
if bytes.starts_with(b"version https://git-lfs.github.com/spec/") {
panic!(
"\n\n{model} is a Git LFS pointer, not the model.\n\
Run `git lfs install && git lfs pull`, or build with \
`--no-default-features` for a geometry-only build.\n"
);
}
}
+190
View File
@@ -0,0 +1,190 @@
//! Align real frames from their embedded previews and draw the result.
//!
//! ```sh
//! cargo run -p dr-pano --example align --release -- fixtures/pano/2025-08-05/*.CR2
//! cargo run -p dr-pano --example align --release -- out-prefix frame1.CR2 frame2.CR2 …
//! ```
//!
//! The point of looking rather than asserting: a rotation solve that is
//! numerically converged and geometrically wrong — a mirrored axis, a
//! transposed homography, an orientation applied the wrong way — produces
//! perfectly plausible residuals and a picture that is obviously broken.
//! This writes `<prefix>-cyl.ppm`: every frame's preview warped onto a
//! cylinder and averaged where they overlap, at a size that fits on a
//! screen. Ghosting in the overlaps is the alignment error, made visible.
//!
//! Previews, not RAW: the alignment runs on proxies in the application too
//! (FR-MRG-7), and a camera's embedded JPEG is a proxy the decoder already
//! extracts in milliseconds. What is different from the real path is only
//! that the pixels are the camera's rendering rather than ours, which the
//! geometry does not care about.
use std::path::PathBuf;
use std::time::Instant;
use dr_pano::bundle::Cameras;
use dr_pano::{align, xfeat::XFeat, AlignOptions, Gray, Projection};
fn main() {
env_logger::init();
let mut args: Vec<String> = std::env::args().skip(1).collect();
if args.is_empty() {
eprintln!("usage: align [out-prefix] <frame>...");
std::process::exit(2);
}
let prefix =
if args[0].ends_with(".CR2") || args[0].ends_with(".dng") || args[0].ends_with(".jpg") {
"align".to_string()
} else {
args.remove(0)
};
let paths: Vec<PathBuf> = args.iter().map(PathBuf::from).collect();
// Previews, oriented, at proxy size.
let t = Instant::now();
let mut proxies: Vec<Gray> = Vec::new();
for p in &paths {
let bytes = std::fs::read(p).expect("read");
let preview = dr_decode::extract_preview(&bytes, dr_decode::PreviewSize::Full)
.expect("embedded preview");
let orientation =
dr_decode::orientation(&bytes[..bytes.len().min(dr_decode::HEADER_BYTES as usize)])
.unwrap_or(dr_types::Orientation::NORMAL);
let tag = match orientation.quarter_turns {
1 => 6,
2 => 3,
3 => 8,
_ => 1,
};
let gray = Gray::from_rgba8(
&preview.rgba,
preview.width as usize,
preview.height as usize,
)
.oriented(tag);
let (fitted, _) = gray.fitted(
dr_pano::xfeat::INPUT_LONG_EDGE,
dr_pano::xfeat::INPUT_LONG_EDGE,
);
println!(
"{:<14} preview {}×{} orientation {} → proxy {}×{}",
p.file_name().unwrap().to_string_lossy(),
preview.width,
preview.height,
tag,
fitted.width,
fitted.height
);
proxies.push(fitted);
}
println!("previews in {:?}", t.elapsed());
// Keypoints.
let t = Instant::now();
let mut detector = XFeat::embedded().expect("model");
let features: Vec<_> = proxies
.iter()
.map(|g| detector.detect(g).expect("detect"))
.collect();
for (i, f) in features.iter().enumerate() {
println!("frame {i}: {} keypoints", f.len());
}
println!(
"detection in {:?} ({:?} per frame)",
t.elapsed(),
t.elapsed() / proxies.len() as u32
);
// Alignment.
let t = Instant::now();
let opts = AlignOptions::default();
let alignment = align(&features, &opts).expect("align");
println!("alignment in {:?}", t.elapsed());
println!(
"focal {:.1} px, long edge {} px ({:.1} mm on full frame), rms {:.3} px",
alignment.focal,
proxies[0].width.max(proxies[0].height),
alignment.focal * 36.0 / proxies[0].width.max(proxies[0].height) as f64,
alignment.rms_px
);
for l in &alignment.links {
println!(
" link {}–{}: {} inliers of {} matches",
l.i, l.j, l.inliers, l.matches
);
}
for (k, why) in &alignment.unaligned {
println!(" UNALIGNED frame {k}: {why}");
}
let root = alignment
.rotations
.iter()
.position(|r| *r == Some(dr_pano::linalg::Mat3::IDENTITY))
.unwrap_or(0);
for (k, r) in alignment.rotations.iter().enumerate() {
if let Some(r) = r {
// Yaw about y, pitch about x, roll about z, from the matrix's
// columns — enough to read a sweep by eye.
let yaw = r.0[0][2].atan2(r.0[2][2]).to_degrees();
let pitch = (-r.0[1][2]).asin().to_degrees();
let roll = r.0[1][0].atan2(r.0[1][1]).to_degrees();
println!(
" frame {k}: yaw {yaw:7.2}° pitch {pitch:6.2}° roll {roll:6.2}°{}",
if k == root { " (reference)" } else { "" }
);
}
}
if !alignment.is_complete() {
eprintln!("not drawing: the set is not fully aligned");
std::process::exit(1);
}
// Draw: a cylinder, averaged where frames overlap.
let t = Instant::now();
let cameras: Cameras = alignment.cameras();
let (fw, fh) = (proxies[0].width as f64, proxies[0].height as f64);
let scale = alignment.focal;
let bounds = dr_pano::projection::bounds(Projection::Cylindrical, scale, &cameras, (fw, fh))
.expect("bounds");
// Fit to 3000 px wide.
let out_w = 3000usize;
let px = bounds.width() / out_w as f64;
let out_h = (bounds.height() / px).ceil() as usize;
let mut sum = vec![0.0f32; out_w * out_h];
let mut count = vec![0u16; out_w * out_h];
for oy in 0..out_h {
for ox in 0..out_w {
let u = bounds.min_u + (ox as f64 + 0.5) * px;
let v = bounds.min_v + (oy as f64 + 0.5) * px;
let d = Projection::Cylindrical.to_direction(scale, u, v);
for (k, g) in proxies.iter().enumerate() {
let Some((x, y)) = cameras.project(k, d) else {
continue;
};
let (x, y) = (x + g.width as f64 / 2.0, y + g.height as f64 / 2.0);
if x < 0.0 || y < 0.0 || x >= g.width as f64 - 1.0 || y >= g.height as f64 - 1.0 {
continue;
}
let (x0, y0) = (x as usize, y as usize);
let (tx, ty) = ((x - x0 as f64) as f32, (y - y0 as f64) as f32);
let p = |xx: usize, yy: usize| g.data[yy * g.width + xx];
let val = (p(x0, y0) * (1.0 - tx) + p(x0 + 1, y0) * tx) * (1.0 - ty)
+ (p(x0, y0 + 1) * (1.0 - tx) + p(x0 + 1, y0 + 1) * tx) * ty;
sum[oy * out_w + ox] += val;
count[oy * out_w + ox] += 1;
}
}
}
let mut ppm = format!("P5\n{out_w} {out_h}\n255\n").into_bytes();
ppm.extend(sum.iter().zip(&count).map(|(s, c)| {
if *c == 0 {
0u8
} else {
((s / f32::from(*c)).clamp(0.0, 1.0) * 255.0) as u8
}
}));
let out = format!("{prefix}-cyl.pgm");
std::fs::write(&out, ppm).expect("write");
println!("wrote {out} ({out_w}×{out_h}) in {:?}", t.elapsed());
}
+459
View File
@@ -0,0 +1,459 @@
//! TRACES: FR-MRG-1 | FR-MRG-5
//! From features to cameras: the alignment of a whole set.
//!
//! 1. Match every pair of frames (`matching`).
//! 2. For each pair with enough matches, a robust homography
//! (`homography::ransac_homography`); a pair is a *link* when its inliers
//! pass Brown & Lowe's test, `n_inliers > 8 + 0.3 · n_matches`, which
//! is what separates a real overlap from a coincidence of descriptors.
//! 3. The focal length: the median of what the links' homographies imply,
//! or the caller's hint if none of them implies anything.
//! 4. A spanning tree over the links, strongest first, from the
//! best-connected frame; rotations chained along it.
//! 5. Bundle adjustment over every link's inliers (`bundle`).
//!
//! What it refuses to do is guess. A frame the tree does not reach is
//! reported by index with the reason (FR-MRG-5) and left out of the
//! cameras; the caller decides whether a set with a hole is worth
//! stitching, and the requirement says it is not.
use crate::bundle::{self, AdjustOptions, Cameras, Observation};
use crate::features::Features;
use crate::homography::{self, RobustHomography};
use crate::linalg::Mat3;
use crate::matching::{match_features, Match};
use crate::PanoError;
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct AlignOptions {
/// Descriptor similarity floor for a match (`matching`).
pub min_similarity: f32,
/// RANSAC agreement distance, in pixels of the features' image.
pub ransac_px: f64,
pub ransac_iterations: usize,
/// A pair needs at least this many inliers to be a link, on top of
/// Brown & Lowe's ratio test.
pub min_inliers: usize,
/// Focal length in pixels of the features' image, if the caller knows
/// it (EXIF and a sensor width). Used only when the homographies do not
/// determine one.
pub focal_hint: Option<f64>,
pub adjust: AdjustOptions,
/// For RANSAC's sampling: the same seed gives the same alignment
/// (NFR-MRG-2).
pub seed: u64,
}
impl Default for AlignOptions {
fn default() -> Self {
AlignOptions {
min_similarity: 0.82,
ransac_px: 3.0,
ransac_iterations: 1000,
min_inliers: 12,
focal_hint: None,
adjust: AdjustOptions::default(),
seed: 0x5eed,
}
}
}
/// An overlap the alignment trusts.
#[derive(Debug, Clone, PartialEq)]
pub struct Link {
pub i: usize,
pub j: usize,
pub matches: usize,
pub inliers: usize,
/// Maps centred points of `i` to centred points of `j`.
pub h: Mat3,
}
/// Why a frame is not in the alignment.
#[derive(Debug, Clone, PartialEq, Eq)]
pub enum Unaligned {
/// Not enough matches with any other frame to try a geometry.
NoMatches,
/// Matches existed but none survived RANSAC as a real overlap.
NoOverlap,
/// Overlaps existed but only with frames that are themselves unaligned.
Disconnected,
}
impl std::fmt::Display for Unaligned {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
f.write_str(match self {
Unaligned::NoMatches => "too few matching features with any other frame",
Unaligned::NoOverlap => "no consistent overlap with any other frame",
Unaligned::Disconnected => "overlaps only with frames that could not be aligned",
})
}
}
/// The result: cameras for the aligned frames, and the rest named.
#[derive(Debug, Clone, PartialEq)]
pub struct Alignment {
/// One rotation per input frame, camera to world, for aligned frames;
/// `None` for the unaligned. The reference frame is the best-connected
/// one and has the identity.
pub rotations: Vec<Option<Mat3>>,
/// Focal length in pixels of the features' image.
pub focal: f64,
pub links: Vec<Link>,
pub unaligned: Vec<(usize, Unaligned)>,
/// Bundle adjustment's RMS reprojection error, in pixels.
pub rms_px: f64,
}
impl Alignment {
pub fn is_complete(&self) -> bool {
self.unaligned.is_empty()
}
/// The cameras of the aligned frames, indexed as the input — a frame
/// that is not aligned is given the identity, so this is only useful
/// when [`Self::is_complete`].
pub fn cameras(&self) -> Cameras {
Cameras {
rotations: self
.rotations
.iter()
.map(|r| r.unwrap_or(Mat3::IDENTITY))
.collect(),
focal: self.focal,
}
}
}
/// Align a set of frames from their features.
///
/// Every `Features` must be in its own frame's pixel coordinates with the
/// image size filled in; points are centred on the image centre here. The
/// frames must all come from the same lens at the same focal length, which
/// is the panorama assumption and not checked — the caller has the EXIF.
pub fn align(frames: &[Features], opts: &AlignOptions) -> Result<Alignment, PanoError> {
let n = frames.len();
if n < 2 {
return Err(PanoError::Input(
"a panorama needs at least two frames".into(),
));
}
let centre = |k: usize, i: usize| -> (f64, f64) {
let kp = frames[k].keypoints[i];
(
f64::from(kp.x) - frames[k].width as f64 / 2.0,
f64::from(kp.y) - frames[k].height as f64 / 2.0,
)
};
// Scale for the DLT's conditioning: points of order one.
let scale = 1.0
/ frames
.iter()
.map(|f| f.width.max(f.height) as f64)
.fold(1.0, f64::max);
// 1 + 2: every pair.
let mut links = Vec::new();
let mut observations: Vec<Observation> = Vec::new();
let mut matched_any = vec![false; n];
let t_match = std::time::Instant::now();
for i in 0..n {
for j in i + 1..n {
let matches: Vec<Match> = match_features(&frames[i], &frames[j], opts.min_similarity);
log::debug!("pair {i}-{j}: {} matches", matches.len());
if matches.len() < 4 {
continue;
}
matched_any[i] = true;
matched_any[j] = true;
let pairs: Vec<((f64, f64), (f64, f64))> = matches
.iter()
.map(|m| {
let (a, b) = (centre(i, m.a), centre(j, m.b));
((a.0 * scale, a.1 * scale), (b.0 * scale, b.1 * scale))
})
.collect();
let Some(RobustHomography { h, inliers }) = homography::ransac_homography(
&pairs,
opts.ransac_px * scale,
opts.ransac_iterations,
opts.seed ^ ((i as u64) << 32 | j as u64),
) else {
continue;
};
let needed = (8.0 + 0.3 * matches.len() as f64).ceil() as usize;
log::debug!("pair {i}-{j}: {} inliers, {needed} needed", inliers.len());
if inliers.len() <= needed || inliers.len() < opts.min_inliers {
continue;
}
// Back to pixels: H_px = S⁻¹ H S.
let m = h.0;
let h_px = Mat3([
[m[0][0], m[0][1], m[0][2] / scale],
[m[1][0], m[1][1], m[1][2] / scale],
[m[2][0] * scale, m[2][1] * scale, m[2][2]],
]);
for &k in &inliers {
let (a, b) = pairs[k];
observations.push(Observation {
i,
j,
pi: (a.0 / scale, a.1 / scale),
pj: (b.0 / scale, b.1 / scale),
});
}
links.push(Link {
i,
j,
matches: matches.len(),
inliers: inliers.len(),
h: h_px,
});
}
}
log::debug!("matching and pairwise geometry in {:?}", t_match.elapsed());
// 3: the focal length.
let mut estimates: Vec<f64> = links
.iter()
.filter_map(|l| homography::focal_from_homography(&l.h))
.filter(|f| f.is_finite() && *f > 0.0)
.collect();
let longest = frames
.iter()
.map(|f| f.width.max(f.height) as f64)
.fold(0.0, f64::max);
let focal = if !estimates.is_empty() {
estimates.sort_by(f64::total_cmp);
let median = estimates[estimates.len() / 2];
// A homography of a nearly pure pan can imply almost anything;
// clamp to the range a real lens on this sensor can reach.
median.clamp(0.3 * longest, 6.0 * longest)
} else if let Some(hint) = opts.focal_hint {
hint
} else {
// No overlap said anything and nobody told us: a normal lens.
longest
};
// 4: spanning tree, strongest link first, from the best-connected frame.
let mut rotations: Vec<Option<Mat3>> = vec![None; n];
let mut unaligned = Vec::new();
if links.is_empty() {
for (k, &matched) in matched_any.iter().enumerate() {
unaligned.push((
k,
if matched {
Unaligned::NoOverlap
} else {
Unaligned::NoMatches
},
));
}
return Ok(Alignment {
rotations,
focal,
links,
unaligned,
rms_px: 0.0,
});
}
let mut degree = vec![0usize; n];
for l in &links {
degree[l.i] += l.inliers;
degree[l.j] += l.inliers;
}
let root = (0..n).max_by_key(|&k| degree[k]).unwrap_or(0);
rotations[root] = Some(Mat3::IDENTITY);
loop {
// The strongest link from an aligned frame to an unaligned one.
let best = links
.iter()
.filter(|l| rotations[l.i].is_some() != rotations[l.j].is_some())
.max_by_key(|l| l.inliers);
let Some(l) = best else { break };
let r_ij = homography::rotation_from_homography(&l.h, focal);
// H_ij takes points of i to j, so bearings b_j = R_ij b_i, and with
// world = R_i · cam_i: R_j = R_i · R_ijᵀ.
if let Some(ri) = rotations[l.i] {
rotations[l.j] = Some((ri * r_ij.transpose()).orthonormalised());
} else if let Some(rj) = rotations[l.j] {
rotations[l.i] = Some((rj * r_ij).orthonormalised());
}
}
for k in 0..n {
if rotations[k].is_none() {
let reason = if !matched_any[k] {
Unaligned::NoMatches
} else if links.iter().any(|l| l.i == k || l.j == k) {
Unaligned::Disconnected
} else {
Unaligned::NoOverlap
};
unaligned.push((k, reason));
}
}
// 5: adjust the aligned frames together. The reference frame must be
// index 0 of the adjustment (it holds frame 0 fixed), so the aligned
// frames are renumbered with the root first.
let aligned: Vec<usize> = std::iter::once(root)
.chain((0..n).filter(|&k| k != root && rotations[k].is_some()))
.collect();
let index_of = |k: usize| aligned.iter().position(|&a| a == k);
let start = Cameras {
rotations: aligned.iter().map(|&k| rotations[k].unwrap()).collect(),
focal,
};
let obs: Vec<Observation> = observations
.iter()
.filter_map(|o| {
Some(Observation {
i: index_of(o.i)?,
j: index_of(o.j)?,
pi: o.pi,
pj: o.pj,
})
})
.collect();
let t_adjust = std::time::Instant::now();
let adjusted = bundle::adjust(start, &obs, &opts.adjust)?;
log::debug!(
"bundle adjustment: {} observations, {} iterations in {:?}",
obs.len(),
adjusted.iterations,
t_adjust.elapsed()
);
for (slot, &k) in aligned.iter().enumerate() {
rotations[k] = Some(adjusted.cameras.rotations[slot]);
}
Ok(Alignment {
rotations,
focal: adjusted.cameras.focal,
links,
unaligned,
rms_px: adjusted.rms_px,
})
}
#[cfg(test)]
mod tests {
use super::*;
use crate::features::{Keypoint, DESCRIPTOR_LEN};
use crate::linalg::Vec3;
/// Frames of a synthetic sweep: world directions with random unit
/// descriptors, each frame seeing the ones in its field of view.
fn synthetic_sweep(
n: usize,
step: f64,
f: f64,
w: usize,
h: usize,
) -> (Vec<Features>, Cameras) {
let mut seed = 777u64;
let mut rnd = || {
seed = seed
.wrapping_mul(6364136223846793005)
.wrapping_add(1442695040888963407);
((seed >> 33) as f64 / (1u64 << 31) as f64) - 0.5
};
let rotations: Vec<Mat3> = (0..n)
.map(|k| {
Mat3::rotation(Vec3::new(0.0, 1.0, 0.0), step * k as f64)
* Mat3::rotation(Vec3::new(1.0, 0.0, 0.0), 0.02 * ((k % 3) as f64 - 1.0))
})
.collect();
let truth = Cameras {
rotations,
focal: f,
};
let total = step * (n as f64 - 1.0);
let mut frames: Vec<Features> = (0..n)
.map(|_| Features {
keypoints: Vec::new(),
descriptors: Vec::new(),
width: w,
height: h,
})
.collect();
for _ in 0..600 * n {
let yaw = rnd() * (total + 0.8) + total / 2.0;
let pitch = rnd() * 0.5;
let d = Vec3::new(
yaw.sin() * pitch.cos(),
pitch.sin(),
yaw.cos() * pitch.cos(),
);
let desc: Vec<f32> = (0..DESCRIPTOR_LEN).map(|_| rnd() as f32).collect();
let norm = desc.iter().map(|v| v * v).sum::<f32>().sqrt();
let desc: Vec<f32> = desc.iter().map(|v| v / norm).collect();
for (k, frame) in frames.iter_mut().enumerate() {
if let Some(p) = truth.project(k, d) {
let (x, y) = (p.0 + w as f64 / 2.0, p.1 + h as f64 / 2.0);
if x >= 0.0 && x < w as f64 && y >= 0.0 && y < h as f64 {
frame.keypoints.push(Keypoint {
x: (x + rnd() * 0.6) as f32,
y: (y + rnd() * 0.6) as f32,
score: 1.0,
});
frame.descriptors.extend_from_slice(&desc);
}
}
}
}
(frames, truth)
}
fn angle_between(a: Mat3, b: Mat3) -> f64 {
(a.transpose() * b).log().norm()
}
#[test]
fn a_synthetic_sweep_is_aligned_to_its_truth() {
let (frames, truth) = synthetic_sweep(6, 0.3, 1400.0, 1024, 768);
let out = align(&frames, &AlignOptions::default()).expect("aligned");
assert!(out.is_complete(), "unaligned: {:?}", out.unaligned);
assert_eq!(out.links.len(), 5 + 4, "links: {}", out.links.len());
assert!((out.focal - 1400.0).abs() < 15.0, "focal {}", out.focal);
assert!(out.rms_px < 1.0, "rms {}", out.rms_px);
// Relative rotations match the truth's, whichever frame is the root.
let root = out
.rotations
.iter()
.position(|r| *r == Some(Mat3::IDENTITY))
.unwrap();
for k in 0..6 {
let rel_truth = truth.rotations[root].transpose() * truth.rotations[k];
let rel_out = out.rotations[k].unwrap();
let err = angle_between(rel_truth, rel_out);
assert!(err < 2e-3, "frame {k} off by {err} rad");
}
}
#[test]
fn a_frame_from_nowhere_is_named_not_guessed() {
let (mut frames, _) = synthetic_sweep(4, 0.3, 1400.0, 1024, 768);
// Frame 3 gets descriptors nobody else has.
for v in &mut frames[3].descriptors {
*v = -*v;
}
let out = align(&frames, &AlignOptions::default()).expect("aligned");
assert_eq!(out.unaligned.len(), 1);
assert_eq!(out.unaligned[0].0, 3);
assert!(out.rotations[3].is_none());
assert!(out.rotations[..3].iter().all(Option::is_some));
}
#[test]
fn one_frame_is_refused() {
let (frames, _) = synthetic_sweep(1, 0.3, 1400.0, 640, 480);
assert!(matches!(
align(&frames, &AlignOptions::default()),
Err(PanoError::Input(_))
));
}
}
+446
View File
@@ -0,0 +1,446 @@
//! Bundle adjustment: every rotation and the focal length, refined together.
//!
//! The pairwise homographies (`homography.rs`) each know about two frames.
//! Chained around a loop they disagree with themselves by the accumulated
//! error, and a twelve-frame sweep chained end to end drifts by a visible
//! amount. This solves for all the rotations at once, against every inlier
//! match of every pair, so the error is spread rather than accumulated —
//! Brown & Lowe's step 4, with the camera model reduced to what a panorama
//! needs: one rotation per frame and one focal length shared by all.
//!
//! Levenberg–Marquardt with a numerical Jacobian. Analytic derivatives of a
//! rotation's projection are not hard, but they are a second place the
//! model is written down, and the model is small: forty parameters, a few
//! thousand residuals, a Jacobian that costs forty residual evaluations.
//! The whole solve is milliseconds. Correctness over cleverness, and one
//! definition of the projection to keep right.
use crate::linalg::{DMat, Mat3, Vec3};
use crate::PanoError;
/// A point in one image, centred on the principal point, in pixels.
pub type Point = (f64, f64);
/// One inlier match between two frames.
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct Observation {
pub i: usize,
pub j: usize,
pub pi: Point,
pub pj: Point,
}
/// What the adjustment starts from and returns: a rotation per frame
/// (camera to world; frame 0 is the world) and the focal length in pixels.
#[derive(Debug, Clone, PartialEq)]
pub struct Cameras {
pub rotations: Vec<Mat3>,
pub focal: f64,
}
impl Cameras {
/// The unit direction, in world space, that pixel `p` of frame `i` looks
/// along.
pub fn bearing(&self, i: usize, p: Point) -> Vec3 {
self.rotations[i] * Vec3::new(p.0, p.1, self.focal).normalised()
}
/// Where world direction `d` lands in frame `j`, or `None` if it is
/// behind the camera.
pub fn project(&self, j: usize, d: Vec3) -> Option<Point> {
let c = self.rotations[j].transpose() * d;
if c.z() <= 1e-9 {
return None;
}
Some((self.focal * c.x() / c.z(), self.focal * c.y() / c.z()))
}
}
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct AdjustOptions {
pub max_iterations: usize,
/// Residuals beyond this many pixels are down-weighted (Huber), so a
/// mismatch RANSAC let through pulls with bounded force.
pub huber_px: f64,
/// Whether the focal length is a free parameter. Off, it is held at the
/// starting value — for a set whose rotations are all small, the focal
/// length is weakly observable and better taken from the homographies'
/// median than pulled about by noise.
pub refine_focal: bool,
}
impl Default for AdjustOptions {
fn default() -> Self {
AdjustOptions {
max_iterations: 50,
huber_px: 3.0,
refine_focal: true,
}
}
}
/// The adjusted cameras and the fit.
#[derive(Debug, Clone, PartialEq)]
pub struct Adjusted {
pub cameras: Cameras,
/// Root-mean-square reprojection error over all observations, in pixels
/// (unweighted, so an outlier RANSAC missed shows here rather than
/// hiding under its Huber weight).
pub rms_px: f64,
pub iterations: usize,
}
/// Refine `start` against `observations`.
///
/// Frame 0's rotation is held fixed: the world frame is arbitrary and
/// fixing one camera removes the freedom. Every other frame must appear in
/// at least one observation or its rotation is undetermined and the normal
/// equations are singular — the caller (`align`) guarantees it by only
/// adjusting frames a spanning tree reached.
pub fn adjust(
start: Cameras,
observations: &[Observation],
opts: &AdjustOptions,
) -> Result<Adjusted, PanoError> {
let n_frames = start.rotations.len();
if n_frames < 2 || observations.is_empty() {
let rms = rms(&start, observations);
return Ok(Adjusted {
cameras: start,
rms_px: rms,
iterations: 0,
});
}
// Every adjustable frame must be constrained by something, or its
// block of the normal equations is zero and the solve is meaningless —
// checked here, by name, rather than left to surface as a step that
// fails to lower the cost.
let mut seen = vec![false; n_frames];
for o in observations {
seen[o.i] = true;
seen[o.j] = true;
}
if let Some(k) = (1..n_frames).find(|&k| !seen[k]) {
return Err(PanoError::Geometry(format!(
"frame {k} has no observations constraining it"
)));
}
let n_rot = 3 * (n_frames - 1);
let n_params = n_rot + usize::from(opts.refine_focal);
let n_res = 2 * observations.len();
// Parameters are *increments* on the current cameras, re-applied each
// accepted step: rotation k ← exp(δ_k) · rotation k, focal ← f · exp(δ_f).
// Composing on the left keeps the increment in world space, where a
// small rotation means the same thing for every frame.
let apply = |base: &Cameras, x: &[f64]| -> Cameras {
let mut rotations = base.rotations.clone();
for k in 1..n_frames {
let w = Vec3::new(x[3 * (k - 1)], x[3 * (k - 1) + 1], x[3 * (k - 1) + 2]);
rotations[k] = (Mat3::exp(w) * base.rotations[k]).orthonormalised();
}
let focal = if opts.refine_focal {
base.focal * x[n_rot].exp()
} else {
base.focal
};
Cameras { rotations, focal }
};
let residuals = |c: &Cameras, out: &mut Vec<f64>| {
out.clear();
for o in observations {
let d = c.bearing(o.i, o.pi);
match c.project(o.j, d) {
Some((x, y)) => {
out.push(x - o.pj.0);
out.push(y - o.pj.1);
}
None => {
// Behind the camera: as wrong as a residual can be
// without being infinite. The Huber weight caps its pull.
out.push(1e4);
out.push(1e4);
}
}
}
};
let weights = |r: &[f64], out: &mut Vec<f64>| {
out.clear();
for pair in r.chunks_exact(2) {
let m = (pair[0] * pair[0] + pair[1] * pair[1]).sqrt();
let w = if m > opts.huber_px {
opts.huber_px / m
} else {
1.0
};
out.push(w);
out.push(w);
}
};
// The robust cost itself, not the weighted sum of squares: the weights
// above are the IRLS linearisation for one step, and comparing two
// steps by sums taken under different weights would accept the wrong
// ones. Huber: quadratic within the threshold, linear beyond it.
let cost = |r: &[f64]| -> f64 {
r.chunks_exact(2)
.map(|pair| {
let m = (pair[0] * pair[0] + pair[1] * pair[1]).sqrt();
if m <= opts.huber_px {
m * m
} else {
2.0 * opts.huber_px * m - opts.huber_px * opts.huber_px
}
})
.sum()
};
let mut cameras = start;
let mut r = Vec::with_capacity(n_res);
let mut w = Vec::with_capacity(n_res);
residuals(&cameras, &mut r);
weights(&r, &mut w);
let mut current = cost(&r);
let mut lambda = 1e-3;
let mut jac = vec![0.0f64; n_res * n_params];
let mut r_plus = Vec::with_capacity(n_res);
let zero = vec![0.0f64; n_params];
let mut iterations = 0;
for _ in 0..opts.max_iterations {
iterations += 1;
// Numerical Jacobian about the current cameras (x = 0).
const H: f64 = 1e-6;
for p in 0..n_params {
let mut x = zero.clone();
x[p] = H;
let c_plus = apply(&cameras, &x);
residuals(&c_plus, &mut r_plus);
for (k, (rp, r0)) in r_plus.iter().zip(&r).enumerate() {
jac[k * n_params + p] = (rp - r0) / H;
}
}
// Normal equations, weighted: (JᵀWJ + λ·diag) δ = −JᵀWr.
let mut a = DMat::zeros(n_params);
let mut b = vec![0.0f64; n_params];
for k in 0..n_res {
let row = &jac[k * n_params..(k + 1) * n_params];
let wk = w[k];
for p in 0..n_params {
b[p] -= wk * row[p] * r[k];
for q in 0..n_params {
a[(p, q)] += wk * row[p] * row[q];
}
}
}
// Try steps with increasing damping until one lowers the cost.
let mut accepted = false;
for _ in 0..10 {
let mut damped = a.clone();
for p in 0..n_params {
let d = a[(p, p)];
damped[(p, p)] = d + lambda * d.max(1e-9);
}
let Some(delta) = damped.solve_spd(&b) else {
return Err(PanoError::Geometry(
"the adjustment's normal equations are singular: a frame has no \
observations constraining it"
.into(),
));
};
let candidate = apply(&cameras, &delta);
residuals(&candidate, &mut r_plus);
let c_new = cost(&r_plus);
if c_new < current {
let improvement = (current - c_new) / current.max(1e-12);
let step: f64 = delta.iter().map(|d| d * d).sum::<f64>().sqrt();
cameras = candidate;
std::mem::swap(&mut r, &mut r_plus);
weights(&r, &mut w);
current = c_new;
lambda = (lambda / 3.0).max(1e-9);
accepted = true;
// Converged when a *lightly damped* step no longer helps. A
// heavily damped step is small by construction and would
// pass an improvement test long before the minimum.
if step < 1e-10 || (improvement < 1e-8 && lambda < 1e-2) {
return Ok(Adjusted {
rms_px: rms(&cameras, observations),
cameras,
iterations,
});
}
break;
}
lambda *= 5.0;
}
if !accepted {
break;
}
}
Ok(Adjusted {
rms_px: rms(&cameras, observations),
cameras,
iterations,
})
}
/// Unweighted RMS reprojection error in pixels.
pub fn rms(c: &Cameras, observations: &[Observation]) -> f64 {
if observations.is_empty() {
return 0.0;
}
let sum: f64 = observations
.iter()
.map(|o| match c.project(o.j, c.bearing(o.i, o.pi)) {
Some((x, y)) => (x - o.pj.0).powi(2) + (y - o.pj.1).powi(2),
None => 1e8,
})
.sum();
(sum / observations.len() as f64).sqrt()
}
#[cfg(test)]
mod tests {
use super::*;
/// A synthetic sweep: `n` cameras panned by `step` radians each with a
/// little pitch and roll, `f` pixels, and matches between neighbours
/// from a cloud of world directions.
fn sweep(n: usize, step: f64, f: f64, noise_px: f64) -> (Cameras, Vec<Observation>) {
let mut rotations = Vec::new();
for k in 0..n {
let yaw = step * k as f64;
let pitch = 0.01 * ((k * 7) % 3) as f64;
let roll = 0.005 * ((k * 5) % 4) as f64;
let r = Mat3::rotation(Vec3::new(0.0, 1.0, 0.0), yaw)
* Mat3::rotation(Vec3::new(1.0, 0.0, 0.0), pitch)
* Mat3::rotation(Vec3::new(0.0, 0.0, 1.0), roll);
rotations.push(r);
}
let truth = Cameras {
rotations,
focal: f,
};
// World directions: a fan across the whole sweep.
let mut obs = Vec::new();
let mut seed = 12345u64;
let mut rnd = || {
seed = seed
.wrapping_mul(6364136223846793005)
.wrapping_add(1442695040888963407);
((seed >> 33) as f64 / (1u64 << 31) as f64) - 0.5
};
let total = step * (n as f64 - 1.0);
for _ in 0..400 * n {
let yaw = rnd() * (total + 0.8) + total / 2.0;
let pitch = rnd() * 0.5;
let d = Vec3::new(
yaw.sin() * pitch.cos(),
pitch.sin(),
yaw.cos() * pitch.cos(),
)
.normalised();
// Visible in which frames? Within ±0.35 f of centre.
let mut seen: Vec<(usize, Point)> = Vec::new();
for k in 0..n {
if let Some(p) = truth.project(k, d) {
if p.0.abs() < 0.35 * f && p.1.abs() < 0.25 * f {
seen.push((k, (p.0 + rnd() * noise_px, p.1 + rnd() * noise_px)));
}
}
}
for a in 0..seen.len() {
for b in a + 1..seen.len() {
obs.push(Observation {
i: seen[a].0,
j: seen[b].0,
pi: seen[a].1,
pj: seen[b].1,
});
}
}
}
(truth, obs)
}
fn angle_between(a: Mat3, b: Mat3) -> f64 {
(a.transpose() * b).log().norm()
}
#[test]
fn a_perturbed_start_converges_back_to_the_truth() {
let (truth, obs) = sweep(6, 0.3, 1400.0, 0.0);
assert!(obs.len() > 500);
// Perturb every rotation but the first by ~1°, and the focal by 5%.
let mut start = truth.clone();
for k in 1..6 {
let w = Vec3::new(0.01, -0.015, 0.008) * (k as f64 / 3.0);
start.rotations[k] = Mat3::exp(w) * start.rotations[k];
}
start.focal *= 1.05;
let before = rms(&start, &obs);
let out = adjust(start, &obs, &AdjustOptions::default()).expect("solvable");
assert!(out.rms_px < 1e-3, "rms {} (was {before})", out.rms_px);
assert!(
(out.cameras.focal - 1400.0).abs() < 0.5,
"focal {}",
out.cameras.focal
);
for k in 0..6 {
let err = angle_between(out.cameras.rotations[k], truth.rotations[k]);
assert!(err < 1e-5, "frame {k} off by {err} rad");
}
}
#[test]
fn noise_is_averaged_rather_than_accumulated() {
let (truth, obs) = sweep(8, 0.25, 1400.0, 1.0);
let mut start = truth.clone();
for k in 1..8 {
start.rotations[k] =
Mat3::exp(Vec3::new(0.0, 0.004 * k as f64, 0.0)) * start.rotations[k];
}
let out = adjust(start, &obs, &AdjustOptions::default()).expect("solvable");
// ±0.5 px of uniform noise on every coordinate has an RMS of 0.41 px
// per axis, so the fit's RMS over both axes should sit near 0.58 and
// cannot be much below it.
assert!(out.rms_px < 0.7, "rms {}", out.rms_px);
// The focal length and the sweep are nearly degenerate for a
// single row: only the perspective inside each overlap pins the
// focal, and a pixel of noise is worth about a tenth of a percent of
// it. What that error does is scale every yaw by the same factor —
// a uniform stretch of the panorama, invisible in the result — so the
// absolute rotation error grows linearly along the sweep and is not
// the measure of the solve. The residual after removing that stretch
// is.
let f_ratio = out.cameras.focal / 1400.0;
assert!((f_ratio - 1.0).abs() < 5e-3, "focal {}", out.cameras.focal);
for k in 0..8 {
let yaw_k = 0.25 * k as f64;
let expected_stretch = (f_ratio - 1.0).abs() * yaw_k;
let err = angle_between(out.cameras.rotations[k], truth.rotations[k]);
assert!(
err < expected_stretch + 1.5e-4,
"frame {k} off by {err} rad, {expected_stretch} of it the focal's"
);
}
}
#[test]
fn a_frame_without_observations_is_refused() {
let (truth, mut obs) = sweep(4, 0.3, 1400.0, 0.0);
obs.retain(|o| o.i != 3 && o.j != 3);
let err = adjust(truth, &obs, &AdjustOptions::default()).unwrap_err();
assert!(matches!(err, PanoError::Geometry(_)));
}
}
+336
View File
@@ -0,0 +1,336 @@
//! Keypoints with descriptors, and the decoder that reads them out of
//! XFeat's dense maps.
//!
//! The network (S15.2) produces three maps at an eighth of the input
//! resolution and stops; everything from there to a list of keypoints is
//! this file, in plain Rust, for the reason `dr-segment` decodes yolo26's
//! heads itself: the post-processing is cheap, shape-dependent and exactly
//! the kind of graph tract parses badly. It is a port of the reference
//! `XFeat.detectAndCompute`, step for step, so that a keypoint here is the
//! keypoint the paper's numbers were measured on.
/// One detected point, in the pixel coordinates of the image it was
/// detected in, with the detector's confidence.
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct Keypoint {
pub x: f32,
pub y: f32,
/// The reliability the detector assigned; higher is better, and the
/// scale is the detector's own — comparable within one model only.
pub score: f32,
}
/// The keypoints of one image and their descriptors.
#[derive(Debug, Clone, PartialEq)]
pub struct Features {
pub keypoints: Vec<Keypoint>,
/// `keypoints.len() × DESCRIPTOR_LEN`, each row L2-normalised, so that a
/// dot product between two rows is their cosine similarity.
pub descriptors: Vec<f32>,
/// The image the coordinates are in.
pub width: usize,
pub height: usize,
}
/// The length of one descriptor. XFeat's is 64; the matcher does not care
/// what the number is, only that both sides agree.
pub const DESCRIPTOR_LEN: usize = 64;
impl Features {
pub fn len(&self) -> usize {
self.keypoints.len()
}
pub fn is_empty(&self) -> bool {
self.keypoints.is_empty()
}
pub fn descriptor(&self, i: usize) -> &[f32] {
&self.descriptors[i * DESCRIPTOR_LEN..(i + 1) * DESCRIPTOR_LEN]
}
}
/// XFeat's three output maps, as the network hands them back.
///
/// All three are `channels × height × width` at an eighth of the input, in
/// the NCHW order the ONNX export declares (`feats [1, 64, H/8, W/8]`,
/// `keypoints [1, 65, H/8, W/8]`, `heatmap [1, 1, H/8, W/8]`).
pub struct XFeatMaps<'a> {
/// 64 channels: the dense descriptor field.
pub feats: &'a [f32],
/// 65 channels: for each 8×8 cell, a logit per position plus one for
/// "no keypoint here".
pub keypoints: &'a [f32],
/// 1 channel: reliability.
pub heatmap: &'a [f32],
/// The maps' width and height (the input's, divided by eight).
pub width: usize,
pub height: usize,
}
/// How the decoder picks keypoints.
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct DecodeOptions {
/// Keep at most this many, by score. The reference default is 4096.
pub top_k: usize,
/// A cell position's softmax probability must exceed this to be a
/// keypoint at all. The reference default is 0.05.
pub threshold: f32,
/// Ignore keypoints within this many pixels of the map's edge. A frame
/// padded into the detector's fixed input (`Gray::padded`) has a hard
/// edge where the padding starts, and the detector fires on it.
pub border: usize,
}
impl Default for DecodeOptions {
fn default() -> Self {
DecodeOptions {
top_k: 4096,
threshold: 0.05,
border: 4,
}
}
}
/// Decode keypoints and descriptors from the network's maps.
///
/// The reference, step for step:
/// 1. softmax over the 65 logits of each cell, keep the 64 positions;
/// 2. pixel-shuffle those into a full-resolution keypoint heatmap — channel
/// `c` of cell `(cx, cy)` is pixel `(cx·8 + c%8, cy·8 + c/8)`;
/// 3. 5×5 non-maximum suppression over that heatmap, above `threshold`;
/// 4. score each survivor by its heatmap value times the reliability map
/// sampled bilinearly at its position;
/// 5. keep the `top_k` by score;
/// 6. sample the descriptor field bilinearly at each and L2-normalise.
///
/// Bilinear where the reference samples the descriptor field bicubically:
/// a quarter-pixel's difference in a field that is smooth by construction,
/// and one interpolator rather than two to keep correct.
pub fn decode_xfeat(maps: &XFeatMaps<'_>, opts: &DecodeOptions) -> Features {
let (w8, h8) = (maps.width, maps.height);
let (w, h) = (w8 * 8, h8 * 8);
let cells = w8 * h8;
debug_assert_eq!(maps.keypoints.len(), 65 * cells);
debug_assert_eq!(maps.feats.len(), DESCRIPTOR_LEN * cells);
debug_assert_eq!(maps.heatmap.len(), cells);
// 1 + 2: softmax per cell, scattered into the full-resolution heatmap.
let mut heat = vec![0.0f32; w * h];
for cy in 0..h8 {
for cx in 0..w8 {
let cell = cy * w8 + cx;
let logit = |c: usize| maps.keypoints[c * cells + cell];
let max = (0..65).map(logit).fold(f32::MIN, f32::max);
let mut sum = 0.0f32;
let mut exps = [0.0f32; 65];
for (c, e) in exps.iter_mut().enumerate() {
*e = (logit(c) - max).exp();
sum += *e;
}
for (c, e) in exps.iter().enumerate().take(64) {
let (dx, dy) = (c % 8, c / 8);
heat[(cy * 8 + dy) * w + cx * 8 + dx] = e / sum;
}
}
}
// 3: a pixel survives if it is the maximum of its 5×5 neighbourhood and
// above threshold. Ties go to every tied pixel, as the reference's
// `x == max_pool(x)` does.
let border = opts.border.max(2);
let mut survivors: Vec<(usize, usize, f32)> = Vec::new();
for y in border..h.saturating_sub(border) {
for x in border..w.saturating_sub(border) {
let v = heat[y * w + x];
if v <= opts.threshold {
continue;
}
let mut is_max = true;
'nb: for ny in y - 2..=y + 2 {
for nx in x - 2..=x + 2 {
if heat[ny * w + nx] > v {
is_max = false;
break 'nb;
}
}
}
if is_max {
survivors.push((x, y, v));
}
}
}
// 4: heatmap value × reliability, the latter sampled at the keypoint's
// position in map coordinates (`align_corners = False`: pixel `x` of the
// full image is `x / 8 - 0.5` in the map).
let sample = |field: &[f32], channels: usize, c: usize, x: f32, y: f32| -> f32 {
let fx = (x / 8.0 - 0.5).clamp(0.0, (w8 - 1) as f32);
let fy = (y / 8.0 - 0.5).clamp(0.0, (h8 - 1) as f32);
let x0 = fx as usize;
let y0 = fy as usize;
let x1 = (x0 + 1).min(w8 - 1);
let y1 = (y0 + 1).min(h8 - 1);
let tx = fx - x0 as f32;
let ty = fy - y0 as f32;
let at = |xx: usize, yy: usize| field[c * (w8 * h8) + yy * w8 + xx];
let _ = channels;
let top = at(x0, y0) * (1.0 - tx) + at(x1, y0) * tx;
let bot = at(x0, y1) * (1.0 - tx) + at(x1, y1) * tx;
top * (1.0 - ty) + bot * ty
};
let mut scored: Vec<(usize, usize, f32)> = survivors
.into_iter()
.map(|(x, y, v)| {
let r = sample(maps.heatmap, 1, 0, x as f32, y as f32);
(x, y, v * r)
})
.collect();
// 5: best first, then cut. `sort_unstable_by` on a total order of the
// score; NaN cannot occur — every input is a probability or a sigmoid.
scored.sort_unstable_by(|a, b| b.2.total_cmp(&a.2));
scored.truncate(opts.top_k);
// 6: descriptors.
let mut keypoints = Vec::with_capacity(scored.len());
let mut descriptors = Vec::with_capacity(scored.len() * DESCRIPTOR_LEN);
for (x, y, score) in scored {
let (xf, yf) = (x as f32, y as f32);
let start = descriptors.len();
for c in 0..DESCRIPTOR_LEN {
descriptors.push(sample(maps.feats, DESCRIPTOR_LEN, c, xf, yf));
}
let norm = descriptors[start..]
.iter()
.map(|v| v * v)
.sum::<f32>()
.sqrt()
.max(1e-12);
for v in &mut descriptors[start..] {
*v /= norm;
}
keypoints.push(Keypoint {
x: xf,
y: yf,
score,
});
}
Features {
keypoints,
descriptors,
width: w,
height: h,
}
}
#[cfg(test)]
mod tests {
use super::*;
/// Maps for a `w8 × h8` grid where every cell says "no keypoint" except
/// the listed ones, which put all their weight on one position.
fn maps(w8: usize, h8: usize, hot: &[(usize, usize, usize)]) -> (Vec<f32>, Vec<f32>, Vec<f32>) {
let cells = w8 * h8;
let mut kp = vec![0.0f32; 65 * cells];
// "None" strongly preferred everywhere.
for cell in 0..cells {
kp[64 * cells + cell] = 10.0;
}
for &(cx, cy, c) in hot {
let cell = cy * w8 + cx;
kp[64 * cells + cell] = 0.0;
kp[c * cells + cell] = 10.0;
}
let heat = vec![0.5f32; cells];
// Descriptors: channel c is constant c across the field, so any
// sampled descriptor is the same known vector.
let mut feats = vec![0.0f32; DESCRIPTOR_LEN * cells];
for c in 0..DESCRIPTOR_LEN {
for v in &mut feats[c * cells..(c + 1) * cells] {
*v = c as f32;
}
}
(feats, kp, heat)
}
#[test]
fn a_hot_cell_position_becomes_a_keypoint_at_the_right_pixel() {
// Cell (2, 1), channel 8*3 + 5 = 29 → pixel (2*8 + 5, 1*8 + 3).
let (f, k, h) = maps(8, 8, &[(2, 1, 29)]);
let out = decode_xfeat(
&XFeatMaps {
feats: &f,
keypoints: &k,
heatmap: &h,
width: 8,
height: 8,
},
&DecodeOptions::default(),
);
assert_eq!(out.len(), 1);
assert_eq!((out.keypoints[0].x, out.keypoints[0].y), (21.0, 11.0));
assert_eq!((out.width, out.height), (64, 64));
// Score is the softmax weight (~1) times the reliability (0.5).
assert!((out.keypoints[0].score - 0.5).abs() < 5e-3);
}
#[test]
fn descriptors_are_unit_length() {
let (f, k, h) = maps(8, 8, &[(3, 3, 0), (5, 5, 63)]);
let out = decode_xfeat(
&XFeatMaps {
feats: &f,
keypoints: &k,
heatmap: &h,
width: 8,
height: 8,
},
&DecodeOptions::default(),
);
assert_eq!(out.len(), 2);
for i in 0..2 {
let n: f32 = out.descriptor(i).iter().map(|v| v * v).sum();
assert!((n - 1.0).abs() < 1e-5);
}
}
#[test]
fn top_k_keeps_the_best() {
let (f, k, mut h) = maps(8, 8, &[(1, 1, 0), (3, 3, 0), (5, 5, 0)]);
// Make cell (3, 3) the most reliable.
h[3 * 8 + 3] = 0.9;
let out = decode_xfeat(
&XFeatMaps {
feats: &f,
keypoints: &k,
heatmap: &h,
width: 8,
height: 8,
},
&DecodeOptions {
top_k: 1,
..Default::default()
},
);
assert_eq!(out.len(), 1);
assert_eq!((out.keypoints[0].x, out.keypoints[0].y), (24.0, 24.0));
}
#[test]
fn the_border_is_excluded() {
let (f, k, h) = maps(8, 8, &[(0, 0, 0)]);
let out = decode_xfeat(
&XFeatMaps {
feats: &f,
keypoints: &k,
heatmap: &h,
width: 8,
height: 8,
},
&DecodeOptions::default(),
);
assert!(out.is_empty());
}
}
+917
View File
@@ -0,0 +1,917 @@
//! TRACES: FR-MRG-4
//! Filling a composite's uncovered border, tile by tile, with an inpainter.
//!
//! A merged panorama has a ragged border where no frame reached. FR-MRG-4
//! crops it by default; this fills it instead, when the photographer asks,
//! with pixels a model invents from the picture around them. Everything
//! here is the geometry of that — which tiles to run, what context to hand
//! the model, how to put its answers back — and none of it is the model:
//! that is the [`Inpainter`] trait, with MI-GAN behind it in `migan.rs`
//! and a fake in the tests.
//!
//! # Context across the edge
//!
//! An inpainting model is trained on holes *inside* pictures. A panorama's
//! border is a hole at the picture's *edge*: real content on one side,
//! nothing on the other, and a model given that invents a structure along
//! the open side — streaks of road in the sky, on the first try
//! (2026-09-19). So the known content is mirrored across the coverage
//! edge, column by column for the top and bottom bands and row by row for
//! the sides, into the hole and into a padding ring around the picture,
//! and the ring is presented as *known*. The model then interpolates
//! between real content and its mirror rather than extrapolating into
//! nothing. The ring is cut off at the end.
//!
//! # Structure from far away, texture from near
//!
//! One tiled pass at the working resolution was not enough: a 512-px tile
//! straddling the coverage edge sees a few hundred pixels of real content
//! on one side and invents the rest from that, two neighbouring tiles
//! invent differently, and the seams and the merge's own fringe leak into
//! the fill. [`fill_border`] therefore runs in two stages. A **coarse**
//! pass at a quarter of the size, where the whole border and hundreds of
//! pixels of real context sit inside a handful of tiles, decides the
//! structure — where the slope goes, where the sky stays sky. Then
//! **fine** passes regenerate the hole in bands from the real edge
//! outward: each band is the only unknown, with real content (or the band
//! before, freshly textured) on its near side and the coarse fill,
//! upsampled, on its far side — blurry, but the right structure — so the
//! model generates texture and a transition, never a large hole from
//! nothing.
//!
//! Tiles overlap by a third and are blended under a raised-cosine window,
//! so the seams between tiles do not show; the model's answer replaces
//! only the pixels that were unknown, and the picture itself is untouched.
use crate::PanoError;
/// A model that fills a square hole from its surroundings.
pub trait Inpainter {
/// The square tile it takes, in pixels.
fn tile(&self) -> usize;
/// Fill one tile. `rgb` is `tile × tile × 3`, row-major, 0..1, with the
/// unknown pixels' values meaningless; `known` is `tile × tile`. The
/// result is `tile × tile × 3`, 0..1, of which only the unknown pixels
/// are read.
fn fill(&mut self, rgb: &[f32], known: &[bool]) -> Result<Vec<f32>, PanoError>;
}
/// What a caller hears from [`fill_border`]: progress, for a page's bar,
/// and — for whoever is looking at why a fill went wrong — each stage's
/// picture as it lands. A plain `FnMut(usize, usize)` is an observer that
/// hears only the progress.
pub trait Observer {
/// `(done, total)` tiles, the total an estimate until the last band.
fn progress(&mut self, done: usize, total: usize);
/// A stage's result, `width × height × 3`: `coarse` (at the coarse
/// size), `band-N` after each fine band, `feathered` at the end.
fn stage(&mut self, _name: &str, _rgb: &[f32], _width: usize, _height: usize) {}
}
impl<F: FnMut(usize, usize)> Observer for F {
fn progress(&mut self, done: usize, total: usize) {
self(done, total)
}
}
/// How far the picture is extended with mirrored content before tiling.
/// Half a tile: enough that a hole at the edge sits well inside a tile.
pub const RING: usize = 256;
/// The fill's knobs, in pixels of the working image. The defaults are
/// what the fixture panorama looked best with on 2026-09-19; the merge
/// page exposes every one of them while the fill is experimental, so a
/// bad corner can be worked on from the picture rather than the code.
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct Params {
/// The coarse pass's reduction: 1 skips it.
pub coarse: usize,
/// The fine passes' band width.
pub band: usize,
/// How deep into the picture the mirrored context reaches, or **zero
/// for no mirrored context at all**: the void is then shown to the
/// model as it is — reaching the picture's edge with nothing beyond,
/// and, beyond the band being filled, still unknown. That is what the
/// shipped model was trained on (a fine-tune of MI-GAN on voids cut
/// from photographs the way a cylindrical merge cuts them, see
/// `docs/dev/panorama.md` §14); a ring would give it a fold to continue.
///
/// Non-zero is the stock model's crutch: a plain reflection of a deep
/// hole pulls in whatever is that far from the edge — a ridge, a peak —
/// and the model, told that is what lies beyond, paints it upside down.
/// Folding the reflection within this band keeps the ring looking like
/// the edge it continues and nothing further away.
pub mirror_depth: usize,
/// How far inside the real edge the fill also regenerates, the two
/// blended by distance. A hard cut between real pixels and invented
/// ones is a line whatever the fill's quality; blended over this many
/// pixels it is not. Zero is the hard cut.
pub feather: usize,
/// The step between tiles, at most the tile; two thirds of it usual.
pub stride: usize,
}
impl Default for Params {
fn default() -> Self {
Params {
coarse: 1,
band: 192,
mirror_depth: 0,
feather: 24,
stride: 384,
}
}
}
/// Fill the unknown pixels of `rgb` (`width × height × 3`, 0..1) in place:
/// the coarse pass, then the fine bands, then the seam feathered over
/// `feather` pixels inside the real edge. Returns the tiles run.
///
/// `known` is `width × height`. `observer` hears the progress and, if it
/// cares, each stage.
pub fn fill_border(
rgb: &mut [f32],
width: usize,
height: usize,
known: &[bool],
model: &mut dyn Inpainter,
params: Params,
observer: &mut dyn Observer,
) -> Result<usize, PanoError> {
let Params {
coarse: q,
band,
mirror_depth,
feather,
stride,
} = params;
let q = q.max(1);
let band = band.max(8);
// No ring: the void beyond the band stays unknown, as in the model's
// training; with a ring the far side is the coarse fill, presented as
// known, which the stock model needed to see something there.
let open = mirror_depth == 0;
if width == 0 || height == 0 || rgb.len() != width * height * 3 || known.len() != width * height
{
return Err(PanoError::Input("fill: buffer sizes disagree".into()));
}
if known.iter().all(|&k| k) {
return Ok(0);
}
// The fill regenerates a margin inside the real edge too, and the
// result is blended with the real pixels across it at the end.
let real = rgb.to_vec();
let outer = known.to_vec();
let mut inner = known.to_vec();
erode(&mut inner, width, height, feather);
let known = &inner[..];
let mut done = 0usize;
// Coarse: a fraction of the size, unknown where any pixel of the cell was.
let (cw, ch) = ((width / q).max(1), (height / q).max(1));
let mut coarse = vec![0.0f32; cw * ch * 3];
let mut cknown = vec![true; cw * ch];
for y in 0..ch {
for x in 0..cw {
let mut sum = [0.0f32; 3];
let mut n = 0.0f32;
let mut all_known = true;
for dy in 0..q {
for dx in 0..q {
let (sx, sy) = ((x * q + dx).min(width - 1), (y * q + dy).min(height - 1));
let i = sy * width + sx;
all_known &= known[i];
for c in 0..3 {
sum[c] += rgb[i * 3 + c];
}
n += 1.0;
}
}
for c in 0..3 {
coarse[(y * cw + x) * 3 + c] = sum[c] / n;
}
cknown[y * cw + x] = all_known;
}
}
let estimate = |tiles: usize| tiles * 4;
done += fill_once(
&mut coarse,
cw,
ch,
&cknown,
model,
stride,
mirror_depth,
|n, t| observer.progress(n, estimate(t)),
)?;
observer.stage("coarse", &coarse, cw, ch);
// The hole starts as the coarse structure, upsampled.
for y in 0..height {
for x in 0..width {
let i = y * width + x;
if known[i] {
continue;
}
let fx = ((x as f32 + 0.5) / q as f32 - 0.5).clamp(0.0, (cw - 1) as f32);
let fy = ((y as f32 + 0.5) / q as f32 - 0.5).clamp(0.0, (ch - 1) as f32);
let (x0, y0) = (fx as usize, fy as usize);
let (x1, y1) = ((x0 + 1).min(cw - 1), (y0 + 1).min(ch - 1));
let (tx, ty) = (fx - x0 as f32, fy - y0 as f32);
for c in 0..3 {
let at = |xx: usize, yy: usize| coarse[(yy * cw + xx) * 3 + c];
rgb[i * 3 + c] = (at(x0, y0) * (1.0 - tx) + at(x1, y0) * tx) * (1.0 - ty)
+ (at(x0, y1) * (1.0 - tx) + at(x1, y1) * tx) * ty;
}
}
}
// Fine, in bands from the edge outward.
let dist = distance_to_known(known, width, height);
let mut band_known = vec![true; width * height];
let mut b = 0usize;
loop {
let lo = (b * band).saturating_sub(band / 2) as f32;
let hi = ((b + 1) * band) as f32;
let mut any = false;
for i in 0..width * height {
let in_band = !known[i] && dist[i] > lo && dist[i] <= hi;
band_known[i] = if open {
known[i] || dist[i] <= lo
} else {
!in_band
};
any |= in_band;
}
if !any {
break;
}
let before = done;
done += fill_once(
rgb,
width,
height,
&band_known,
model,
stride,
mirror_depth,
|n, t| observer.progress(before + n, before + estimate(t)),
)?;
observer.stage(&format!("band-{b}"), rgb, width, height);
b += 1;
}
// The seam: across the margin, real on the inside, invented on the
// outside, a smooth ramp between by distance from the true hole.
if feather > 0 {
let to_hole =
distance_to_known(&outer.iter().map(|k| !k).collect::<Vec<_>>(), width, height);
for i in 0..width * height {
if !outer[i] || known[i] {
continue;
}
// In the margin: outer says known, inner says not.
let t = (to_hole[i] / feather as f32).clamp(0.0, 1.0);
let t = t * t * (3.0 - 2.0 * t);
for c in 0..3 {
rgb[i * 3 + c] = rgb[i * 3 + c] * (1.0 - t) + real[i * 3 + c] * t;
}
}
}
observer.stage("feathered", rgb, width, height);
observer.progress(done, done);
Ok(done)
}
/// One tiled pass: every unknown pixel regenerated from the tiles that
/// touch it, the rest kept. Returns the tiles run.
#[allow(clippy::too_many_arguments)]
fn fill_once(
rgb: &mut [f32],
width: usize,
height: usize,
known: &[bool],
model: &mut dyn Inpainter,
stride: usize,
mirror_depth: usize,
mut progress: impl FnMut(usize, usize),
) -> Result<usize, PanoError> {
let t = model.tile();
if t == 0 || known.iter().all(|&k| k) {
return Ok(0);
}
// The padded canvas with mirrored context, and the hole within it.
let ctx = MirroredContext::build(rgb, width, height, known, mirror_depth, t);
let (pw, ph) = (ctx.width, ctx.height);
// Tiles that touch the hole, on a grid that reaches both far edges.
let starts = |n: usize| -> Vec<usize> {
if n <= t {
return vec![0];
}
let mut v: Vec<usize> = (0..=n - t).step_by(stride.clamp(1, t)).collect();
if *v.last().unwrap_or(&0) != n - t {
v.push(n - t);
}
v
};
let ys = starts(ph);
let xs = starts(pw);
let mut tiles = Vec::new();
for &y in &ys {
for &x in &xs {
if y + t > ph || x + t > pw {
continue;
}
let touches =
(y..y + t).any(|yy| ctx.hole[yy * pw + x..yy * pw + x + t].iter().any(|&h| h));
if touches {
tiles.push((x, y));
}
}
}
// Raised-cosine window, so overlapping tiles blend.
let hann: Vec<f32> = (0..t)
.map(|i| {
let s = ((i as f32 + 1.0) / (t as f32 + 1.0) * std::f32::consts::PI).sin();
s * s + 1e-3
})
.collect();
let mut acc = vec![0.0f32; pw * ph * 3];
let mut wsum = vec![0.0f32; pw * ph];
let mut tile_rgb = vec![0.0f32; t * t * 3];
let mut tile_known = vec![false; t * t];
let total = tiles.len();
for (n, &(x, y)) in tiles.iter().enumerate() {
progress(n, total);
for r in 0..t {
let src = ((y + r) * pw + x) * 3;
tile_rgb[r * t * 3..(r + 1) * t * 3].copy_from_slice(&ctx.rgb[src..src + t * 3]);
let ks = (y + r) * pw + x;
for c in 0..t {
tile_known[r * t + c] = !ctx.hole[ks + c];
}
}
let out = model.fill(&tile_rgb, &tile_known)?;
if out.len() != t * t * 3 {
return Err(PanoError::Model(format!(
"the inpainter returned {} values for a {t}×{t} tile",
out.len()
)));
}
for r in 0..t {
for c in 0..t {
let w = hann[r] * hann[c];
let p = (y + r) * pw + (x + c);
for ch in 0..3 {
acc[p * 3 + ch] += out[(r * t + c) * 3 + ch] * w;
}
wsum[p] += w;
}
}
}
progress(total, total);
// Back into the picture: only the unknown pixels change.
for yy in 0..height {
for xx in 0..width {
let i = yy * width + xx;
if known[i] {
continue;
}
let p = (yy + ctx.ring) * pw + (xx + ctx.ring);
if wsum[p] > 0.0 {
for ch in 0..3 {
rgb[i * 3 + ch] = (acc[p * 3 + ch] / wsum[p]).clamp(0.0, 1.0);
}
}
}
}
Ok(total)
}
/// Shrink `known` by `iterations` pixels on every side, in place.
///
/// The merge's coverage edge carries a fringe — the last partly-covered
/// pixels of a frame, and whatever the renderer did at the boundary — and
/// a fill that stops exactly at the coverage bit leaves it as a dark line
/// along the seam. Eight pixels at a quarter of the composite's resolution
/// was what it took on the fixture.
pub fn erode(known: &mut [bool], width: usize, height: usize, iterations: usize) {
let mut next = known.to_vec();
for _ in 0..iterations {
for y in 0..height {
for x in 0..width {
let i = y * width + x;
if !known[i] {
continue;
}
let edge = x == 0
|| y == 0
|| x + 1 == width
|| y + 1 == height
|| !known[i - 1]
|| !known[i + 1]
|| !known[i - width]
|| !known[i + width];
next[i] = !edge;
}
}
known.copy_from_slice(&next);
}
}
/// Distance from each pixel to the nearest known one, by two chamfer
/// sweeps — within a few percent of Euclidean, and enough to cut bands.
fn distance_to_known(known: &[bool], width: usize, height: usize) -> Vec<f32> {
let inf = (width + height) as f32;
let mut d: Vec<f32> = known.iter().map(|&k| if k { 0.0 } else { inf }).collect();
let (a, b) = (1.0f32, std::f32::consts::SQRT_2);
for y in 0..height {
for x in 0..width {
let i = y * width + x;
let mut v = d[i];
if x > 0 {
v = v.min(d[i - 1] + a);
}
if y > 0 {
v = v.min(d[i - width] + a);
if x > 0 {
v = v.min(d[i - width - 1] + b);
}
if x + 1 < width {
v = v.min(d[i - width + 1] + b);
}
}
d[i] = v;
}
}
for y in (0..height).rev() {
for x in (0..width).rev() {
let i = y * width + x;
let mut v = d[i];
if x + 1 < width {
v = v.min(d[i + 1] + a);
}
if y + 1 < height {
v = v.min(d[i + width] + a);
if x + 1 < width {
v = v.min(d[i + width + 1] + b);
}
if x > 0 {
v = v.min(d[i + width - 1] + b);
}
}
d[i] = v;
}
}
d
}
/// Distance beyond the edge to distance inside it, folded within `depth`
/// ([`Params::mirror_depth`]): a triangle wave, so the band is read
/// forward and back rather than clamped to one row.
fn fold(d: usize, depth: usize) -> usize {
let period = 2 * depth;
let r = d % period;
if r <= depth {
r
} else {
period - r
}
}
/// The picture on a canvas `RING` wider on every side, with the hole and
/// the ring filled by mirroring the known content across the coverage
/// edge — the nearest `depth` of it, folded — and the hole, the
/// original unknown and nothing else, marked.
struct MirroredContext {
width: usize,
height: usize,
/// The padding on every side: `RING` with mirrored context, 0 without.
ring: usize,
rgb: Vec<f32>,
hole: Vec<bool>,
}
impl MirroredContext {
fn build(
rgb: &[f32],
width: usize,
height: usize,
known: &[bool],
depth: usize,
tile: usize,
) -> Self {
if depth == 0 {
// Open: the picture as it is, the hole as it is. What the hole
// holds does not matter — the model masks it out. A picture
// smaller than a tile (the merge page's preview) sits at the
// origin of a tile-sized canvas whose rest is hole: still the
// void as it is, and the only way a tile fits at all.
let (pw, ph) = (width.max(tile), height.max(tile));
let mut canvas = vec![0.0f32; pw * ph * 3];
let mut hole = vec![true; pw * ph];
for y in 0..height {
canvas[y * pw * 3..(y * pw + width) * 3]
.copy_from_slice(&rgb[y * width * 3..(y + 1) * width * 3]);
for x in 0..width {
hole[y * pw + x] = !known[y * width + x];
}
}
return MirroredContext {
width: pw,
height: ph,
ring: 0,
rgb: canvas,
hole,
};
}
let fold = |d: usize| fold(d, depth);
let (pw, ph) = (width + 2 * RING, height + 2 * RING);
let mut canvas = vec![0.0f32; pw * ph * 3];
let mut kn = vec![false; pw * ph];
let mut hole = vec![false; pw * ph];
for y in 0..height {
for x in 0..width {
let i = y * width + x;
let p = (y + RING) * pw + (x + RING);
canvas[p * 3..p * 3 + 3].copy_from_slice(&rgb[i * 3..i * 3 + 3]);
kn[p] = known[i];
hole[p] = !known[i];
}
}
// Per column: mirror across the first and last known row.
for x in 0..pw {
let first = (0..ph).find(|&y| kn[y * pw + x]);
let Some(first) = first else { continue };
let last = (0..ph).rev().find(|&y| kn[y * pw + x]).unwrap_or(first);
for y in 0..first {
let m = (first + fold(first - y)).min(last);
let (d, s) = ((y * pw + x) * 3, (m * pw + x) * 3);
canvas.copy_within(s..s + 3, d);
}
for y in last + 1..ph {
let m = last.saturating_sub(fold(y - last)).max(first);
let (d, s) = ((y * pw + x) * 3, (m * pw + x) * 3);
canvas.copy_within(s..s + 3, d);
}
}
// Per row, for the sides, over what is there now.
for y in 0..ph {
let first = (0..pw).find(|&x| kn[y * pw + x]);
let Some(first) = first else { continue };
let last = (0..pw).rev().find(|&x| kn[y * pw + x]).unwrap_or(first);
for x in 0..first {
let m = (first + fold(first - x)).min(last);
let (d, s) = ((y * pw + x) * 3, (y * pw + m) * 3);
canvas.copy_within(s..s + 3, d);
}
for x in last + 1..pw {
let m = last.saturating_sub(fold(x - last)).max(first);
let (d, s) = ((y * pw + x) * 3, (y * pw + m) * 3);
canvas.copy_within(s..s + 3, d);
}
}
MirroredContext {
width: pw,
height: ph,
ring: RING,
rgb: canvas,
hole,
}
}
}
#[cfg(test)]
mod tests {
use super::*;
/// The tests' small pictures: a 48-px stride, a given feather.
fn test_params(feather: usize) -> Params {
Params {
stride: 48,
feather,
..Params::default()
}
}
/// Paints every unknown pixel a fixed grey and copies the known ones,
/// and remembers what it was shown.
struct Flat {
tile: usize,
seen: Vec<(Vec<f32>, Vec<bool>)>,
}
impl Inpainter for Flat {
fn tile(&self) -> usize {
self.tile
}
fn fill(&mut self, rgb: &[f32], known: &[bool]) -> Result<Vec<f32>, PanoError> {
self.seen.push((rgb.to_vec(), known.to_vec()));
Ok(rgb
.chunks_exact(3)
.zip(known)
.flat_map(|(p, &k)| if k { [p[0], p[1], p[2]] } else { [0.5; 3] })
.collect())
}
}
fn picture(w: usize, h: usize, border: usize) -> (Vec<f32>, Vec<bool>) {
let mut rgb = vec![0.0; w * h * 3];
let mut known = vec![false; w * h];
for y in 0..h {
for x in 0..w {
let i = y * w + x;
if y >= border && y < h - border {
known[i] = true;
rgb[i * 3] = x as f32 / w as f32;
rgb[i * 3 + 1] = y as f32 / h as f32;
rgb[i * 3 + 2] = 0.25;
}
}
}
(rgb, known)
}
#[test]
fn unknown_pixels_take_the_model_and_known_ones_do_not_move() {
let (mut rgb, known) = picture(300, 200, 20);
let before = rgb.clone();
let mut model = Flat {
tile: 64,
seen: Vec::new(),
};
let tiles = fill_border(
&mut rgb,
300,
200,
&known,
&mut model,
test_params(0),
&mut |_, _| {},
)
.unwrap();
assert!(tiles > 0);
for i in 0..300 * 200 {
if known[i] {
assert_eq!(&rgb[i * 3..i * 3 + 3], &before[i * 3..i * 3 + 3]);
} else {
for c in 0..3 {
assert!((rgb[i * 3 + c] - 0.5).abs() < 1e-4, "pixel {i}");
}
}
}
}
#[test]
fn the_model_is_shown_mirrored_context_not_black() {
let (mut rgb, known) = picture(300, 200, 20);
let mut model = Flat {
tile: 64,
seen: Vec::new(),
};
fill_border(
&mut rgb,
300,
200,
&known,
&mut model,
Params {
mirror_depth: 48,
..test_params(0)
},
&mut |_, _| {},
)
.unwrap();
for (tile_rgb, tile_known) in &model.seen {
let known_non_black = tile_rgb
.chunks_exact(3)
.zip(tile_known)
.filter(|(_, &k)| k)
.any(|(p, _)| p.iter().any(|v| *v > 0.0));
assert!(known_non_black);
}
}
#[test]
fn an_open_void_reaches_the_tile_edge_and_stays_unknown_beyond_the_band() {
// A 150-tall hole above and below; bands of 96. With no ring the
// first band's tiles sit at the picture's edge, so a tile's top
// row is unknown, and the rows deeper than the band are unknown
// too — not "known" coarse fill — exactly as the model was trained.
let (mut rgb, known) = picture(200, 500, 150);
let mut model = Flat {
tile: 64,
seen: Vec::new(),
};
fill_border(
&mut rgb,
200,
500,
&known,
&mut model,
Params {
band: 96,
..test_params(0)
},
&mut |_, _| {},
)
.unwrap();
// The first pass's tile at the picture's top edge is unknown
// through and through: the hole is 150 deep, the tile 64, and
// nothing beyond the band was presented as known. With a ring, or
// with the far side shown as coarse fill, no tile is ever all hole.
assert!(model.seen.iter().any(|(_, k)| k.iter().all(|&v| !v)));
for i in 0..200 * 500 {
if !known[i] {
assert!((rgb[i * 3] - 0.5).abs() < 1e-4, "pixel {i}");
}
}
}
#[test]
fn a_picture_smaller_than_the_tile_is_still_filled_when_the_void_is_open() {
// The merge page's preview is 1600 wide and a few hundred tall —
// shorter than a 512 tile. With no ring the canvas is padded to a
// tile, the padding hole, and the border is still filled.
let (mut rgb, known) = picture(300, 40, 8);
let mut model = Flat {
tile: 64,
seen: Vec::new(),
};
let tiles = fill_border(
&mut rgb,
300,
40,
&known,
&mut model,
test_params(0),
&mut |_, _| {},
)
.unwrap();
assert!(tiles > 0, "no tile fitted a picture shorter than the tile");
for i in 0..300 * 40 {
if !known[i] {
assert!((rgb[i * 3] - 0.5).abs() < 1e-4, "pixel {i}");
}
}
// And the model saw the padding as hole, never as black content.
for (_, k) in &model.seen {
assert_eq!(k.len(), 64 * 64);
}
}
#[test]
fn the_fine_passes_run_in_bands_after_the_coarse_one() {
// A 150-tall hole above and below a picture: the coarse pass sees
// it at a quarter; the fine passes need two bands of BAND pixels.
let (mut rgb, known) = picture(200, 500, 150);
let mut model = Flat {
tile: 64,
seen: Vec::new(),
};
fill_border(
&mut rgb,
200,
500,
&known,
&mut model,
Params {
coarse: 4,
band: 96,
mirror_depth: 48,
..test_params(0)
},
&mut |_, _| {},
)
.unwrap();
assert!(model.seen.len() > 4);
// Every unknown pixel was reached.
for i in 0..200 * 500 {
if !known[i] {
assert!((rgb[i * 3] - 0.5).abs() < 1e-4, "pixel {i}");
}
}
}
#[test]
fn the_seam_ramps_from_real_to_invented_across_the_feather() {
let (mut rgb, known) = picture(300, 200, 20);
let before = rgb.clone();
let mut model = Flat {
tile: 64,
seen: Vec::new(),
};
fill_border(
&mut rgb,
300,
200,
&known,
&mut model,
test_params(8),
&mut |_, _| {},
)
.unwrap();
// Row 20 is the real edge; the margin runs to row 27. At the edge
// the value is the model's grey, eight rows in it is the picture's.
let at = |y: usize| rgb[(y * 300 + 150) * 3 + 2];
assert!((at(20) - 0.5).abs() < 0.05, "{}", at(20));
assert!((at(29) - before[(29 * 300 + 150) * 3 + 2]).abs() < 1e-4);
let (lo, hi) = (at(20).min(at(29)), at(20).max(at(29)));
assert!(
at(23) > lo + 0.02 && at(23) < hi - 0.02,
"{} between {lo} and {hi}",
at(23)
);
// The hole itself is the model's.
assert!((at(5) - 0.5).abs() < 1e-4);
}
#[test]
fn erosion_shrinks_the_known_region_from_every_edge() {
let (_, mut known) = picture(20, 20, 4);
erode(&mut known, 20, 20, 2);
assert!(known[8 * 20 + 10]);
assert!(!known[5 * 20 + 10]);
assert!(!known[8 * 20 + 1]);
}
#[test]
fn distance_counts_pixels_from_the_known_region() {
let (_, known) = picture(20, 20, 4);
let d = distance_to_known(&known, 20, 20);
assert_eq!(d[4 * 20 + 10], 0.0);
assert!((d[3 * 20 + 10] - 1.0).abs() < 1e-6);
assert!((d[10] - 4.0).abs() < 1e-6);
}
#[test]
fn a_fully_covered_picture_runs_nothing() {
let (mut rgb, known) = picture(100, 100, 0);
let mut model = Flat {
tile: 64,
seen: Vec::new(),
};
assert_eq!(
fill_border(
&mut rgb,
100,
100,
&known,
&mut model,
test_params(0),
&mut |_, _| {}
)
.unwrap(),
0
);
}
#[test]
fn the_context_mirrors_the_top_rows_upward() {
let (rgb, known) = picture(40, 30, 5);
let ctx = MirroredContext::build(&rgb, 40, 30, &known, 48, 64);
let x = RING + 10;
let first = RING + 5;
for k in 1..=4 {
let above = ((first - k) * ctx.width + x) * 3;
let mirror = ((first + k) * ctx.width + x) * 3;
assert_eq!(&ctx.rgb[above..above + 3], &ctx.rgb[mirror..mirror + 3]);
}
assert!(ctx.hole[(RING + 2) * ctx.width + x]);
assert!(!ctx.hole[(RING - 2) * ctx.width + x]);
}
#[test]
fn the_mirror_reaches_no_deeper_than_its_band() {
// A ridge 200 rows in must not appear in the ring: beyond the band
// the reflection folds back towards the edge rather than on into
// the picture.
let (mut rgb, known) = picture(40, 400, 5);
let ridge = 5 + 200;
for x in 0..40 {
rgb[(ridge * 40 + x) * 3..(ridge * 40 + x) * 3 + 3].copy_from_slice(&[0.9, 0.1, 0.1]);
}
let ctx = MirroredContext::build(&rgb, 40, 400, &known, 48, 64);
let x = RING + 10;
for y in 0..RING + 5 {
let p = (y * ctx.width + x) * 3;
assert!(
ctx.rgb[p] < 0.5,
"row {y} of the ring shows the ridge ({:?})",
&ctx.rgb[p..p + 3]
);
}
assert_eq!(fold(0, 48), 0);
assert_eq!(fold(48, 48), 48);
assert_eq!(fold(58, 48), 38);
assert_eq!(fold(96, 48), 0);
assert_eq!(fold(99, 48), 3);
}
}
+367
View File
@@ -0,0 +1,367 @@
//! Pairwise geometry: a homography between two frames, found robustly.
//!
//! Two frames of a panorama are related by a rotation, and a rotation seen
//! through one lens is a homography of the image plane — `H = K R Kᵀ⁻¹`. The
//! homography is estimated first, from matches, because it does not need
//! the focal length; the focal length is then *read off* it (§ below), and
//! the rotation follows from both. This is the order Brown & Lowe (2007)
//! and OpenCV's stitcher use, and it is what makes the pipeline work when
//! EXIF says nothing about the lens.
//!
//! Coordinates throughout are **centred**: the principal point is the
//! origin. The focal formulae assume it, and centring before the DLT also
//! conditions the linear system — Hartley's normalisation, done once by the
//! caller rather than inside every solve.
use crate::linalg::{DMat, Mat3, Vec3};
/// A point in one image, centred on the principal point.
pub type Point = (f64, f64);
/// Apply a homography to a point.
pub fn apply(h: &Mat3, p: Point) -> Option<Point> {
let v = *h * Vec3::new(p.0, p.1, 1.0);
if v.z().abs() < 1e-12 {
return None;
}
Some((v.x() / v.z(), v.y() / v.z()))
}
/// Least-squares homography from at least four correspondences by the
/// direct linear transform, with `h33` fixed at 1.
///
/// Fixing `h33` turns the homogeneous 8×9 system into an ordinary 8-unknown
/// least-squares problem that the normal equations and a Cholesky
/// factorisation solve without an SVD. The one homography it cannot
/// represent — `h33 = 0`, a point at the origin mapped to infinity — does
/// not occur between overlapping frames of one scene.
///
/// The points should be scaled to order one (divide by the focal length or
/// the image size) before calling: the normal equations square the
/// conditioning, and pixel coordinates in the thousands make them singular
/// in `f64`.
pub fn dlt(pairs: &[(Point, Point)]) -> Option<Mat3> {
if pairs.len() < 4 {
return None;
}
// Each pair gives two rows of A h = b with h = (h11..h32).
// x' = (h11 x + h12 y + h13) / (h31 x + h32 y + 1)
// → h11 x + h12 y + h13 - h31 x x' - h32 y x' = x'
let mut ata = DMat::zeros(8);
let mut atb = [0.0f64; 8];
for &((x, y), (xp, yp)) in pairs {
let rows: [([f64; 8], f64); 2] = [
([x, y, 1.0, 0.0, 0.0, 0.0, -x * xp, -y * xp], xp),
([0.0, 0.0, 0.0, x, y, 1.0, -x * yp, -y * yp], yp),
];
for (a, b) in rows {
for i in 0..8 {
atb[i] += a[i] * b;
for j in 0..8 {
ata[(i, j)] += a[i] * a[j];
}
}
}
}
let h = ata.solve_spd(&atb)?;
Some(Mat3([
[h[0], h[1], h[2]],
[h[3], h[4], h[5]],
[h[6], h[7], 1.0],
]))
}
/// A homography with the correspondences that agree with it.
#[derive(Debug, Clone, PartialEq)]
pub struct RobustHomography {
pub h: Mat3,
/// Indices into the input pairs.
pub inliers: Vec<usize>,
}
/// RANSAC over [`dlt`] on four-point samples, then a final least-squares
/// fit over every inlier.
///
/// `threshold` is the reprojection distance, in the same units as the
/// points, within which a pair counts as agreeing. The iteration count
/// adapts to the inlier ratio found so far in the usual way, capped at
/// `max_iterations`. `seed` makes a run reproducible (NFR-MRG-2): the
/// sampling is a small linear congruential generator, not the system's.
pub fn ransac_homography(
pairs: &[(Point, Point)],
threshold: f64,
max_iterations: usize,
seed: u64,
) -> Option<RobustHomography> {
if pairs.len() < 4 {
return None;
}
let n = pairs.len();
let thr2 = threshold * threshold;
let mut rng = Lcg(seed);
let mut best: Option<(Vec<usize>, Mat3)> = None;
let mut iterations = max_iterations;
let mut i = 0;
while i < iterations {
i += 1;
let sample = rng.distinct4(n);
let Some(h) = dlt(&sample.map(|k| pairs[k])) else {
continue;
};
let inliers: Vec<usize> = (0..n).filter(|&k| agrees(&h, pairs[k], thr2)).collect();
if best.as_ref().is_none_or(|(b, _)| inliers.len() > b.len()) {
// Adapt: enough iterations to have drawn one all-inlier sample
// with probability 0.99, given the ratio seen so far.
let w = inliers.len() as f64 / n as f64;
let p_all = w.powi(4);
if p_all > 0.0 && p_all < 1.0 {
let needed = ((1.0 - 0.99f64).ln() / (1.0 - p_all).ln()).ceil() as usize;
iterations = iterations.min(needed.max(i + 1));
}
best = Some((inliers, h));
}
}
let (inliers, h) = best?;
if inliers.len() < 4 {
return None;
}
// Refit on every inlier, and keep the refit only if it did not lose
// support — a least-squares fit over a set with a few borderline points
// can be pulled off the consensus the sample found.
let refit: Vec<(Point, Point)> = inliers.iter().map(|&k| pairs[k]).collect();
let h = match dlt(&refit) {
Some(r) => {
let count = (0..n).filter(|&k| agrees(&r, pairs[k], thr2)).count();
if count >= inliers.len() {
r
} else {
h
}
}
None => h,
};
let inliers: Vec<usize> = (0..n).filter(|&k| agrees(&h, pairs[k], thr2)).collect();
Some(RobustHomography { h, inliers })
}
fn agrees(h: &Mat3, (p, q): (Point, Point), thr2: f64) -> bool {
match apply(h, p) {
Some((x, y)) => {
let (dx, dy) = (x - q.0, y - q.1);
dx * dx + dy * dy <= thr2
}
None => false,
}
}
/// The focal length a homography implies, if it implies one.
///
/// For `H = K R K⁻¹` with `K = diag(f, f, 1)` and the principal point at the
/// origin, the orthonormality of `R` gives two independent estimates of `f²`
/// from the first two rows and two from the first two columns; each is
/// taken where it is positive and the better-conditioned of the pair is
/// chosen, as OpenCV's `focalsFromHomography` does. The geometric mean of
/// the row and column estimates is returned. `None` when the homography is
/// too close to a pure translation to say anything — every estimate is then
/// a ratio of small numbers.
pub fn focal_from_homography(h: &Mat3) -> Option<f64> {
let m = h.0;
let (h00, h01, h02) = (m[0][0], m[0][1], m[0][2]);
let (h10, h11, h12) = (m[1][0], m[1][1], m[1][2]);
let (h20, h21) = (m[2][0], m[2][1]);
let pick = |mut v1: f64, mut v2: f64, d1: f64, d2: f64| -> Option<f64> {
if v1 < v2 {
std::mem::swap(&mut v1, &mut v2);
}
if v1 > 0.0 && v2 > 0.0 {
Some((if d1.abs() > d2.abs() { v1 } else { v2 }).sqrt())
} else if v1 > 0.0 {
Some(v1.sqrt())
} else {
None
}
};
// From the third row.
let d1 = h20 * h21;
let d2 = (h21 - h20) * (h21 + h20);
let f1 = if d1.abs() > 1e-12 || d2.abs() > 1e-12 {
let v1 = if d1.abs() > 1e-12 {
-(h00 * h01 + h10 * h11) / d1
} else {
f64::NAN
};
let v2 = if d2.abs() > 1e-12 {
(h00 * h00 + h10 * h10 - h01 * h01 - h11 * h11) / d2
} else {
f64::NAN
};
pick(nan_to_neg(v1), nan_to_neg(v2), d1, d2)
} else {
None
};
// From the third column.
let d1 = h00 * h10 + h01 * h11;
let d2 = h00 * h00 + h01 * h01 - h10 * h10 - h11 * h11;
let f0 = if d1.abs() > 1e-12 || d2.abs() > 1e-12 {
let v1 = if d1.abs() > 1e-12 {
-h02 * h12 / d1
} else {
f64::NAN
};
let v2 = if d2.abs() > 1e-12 {
(h12 * h12 - h02 * h02) / d2
} else {
f64::NAN
};
pick(nan_to_neg(v1), nan_to_neg(v2), d1, d2)
} else {
None
};
match (f0, f1) {
(Some(a), Some(b)) => Some((a * b).sqrt()),
(Some(a), None) | (None, Some(a)) => Some(a),
(None, None) => None,
}
}
fn nan_to_neg(v: f64) -> f64 {
if v.is_finite() {
v
} else {
-1.0
}
}
/// The rotation a homography encodes for a known focal length:
/// `R = K⁻¹ H K`, re-orthonormalised, with the scale of `H` divided out.
pub fn rotation_from_homography(h: &Mat3, f: f64) -> Mat3 {
let m = h.0;
// K⁻¹ H K with K = diag(f, f, 1): scale the third row by f and the
// third column by 1/f.
let r = Mat3([
[m[0][0], m[0][1], m[0][2] / f],
[m[1][0], m[1][1], m[1][2] / f],
[m[2][0] * f, m[2][1] * f, m[2][2]],
]);
r.orthonormalised()
}
/// A small deterministic generator for RANSAC's samples.
struct Lcg(u64);
impl Lcg {
fn next(&mut self) -> u64 {
// Knuth's MMIX constants.
self.0 = self
.0
.wrapping_mul(6364136223846793005)
.wrapping_add(1442695040888963407);
self.0 >> 33
}
fn below(&mut self, n: usize) -> usize {
(self.next() % n as u64) as usize
}
fn distinct4(&mut self, n: usize) -> [usize; 4] {
let mut s = [0usize; 4];
for i in 0..4 {
loop {
let k = self.below(n);
if !s[..i].contains(&k) {
s[i] = k;
break;
}
}
}
s
}
}
#[cfg(test)]
mod tests {
use super::*;
/// Points under a known rotation seen through a known focal length,
/// in centred image coordinates scaled by that focal length.
fn synthetic(f: f64, r: Mat3, n: usize, noise: f64, seed: u64) -> Vec<(Point, Point)> {
let mut rng = Lcg(seed);
let mut out = Vec::new();
while out.len() < n {
// A point on the first image plane, within ±0.3 f of centre.
let x = (rng.below(6001) as f64 - 3000.0) / 10000.0;
let y = (rng.below(4001) as f64 - 2000.0) / 10000.0;
let b = Vec3::new(x, y, 1.0);
let v = r * b;
if v.z() <= 0.2 {
continue;
}
let nx = (rng.below(2001) as f64 - 1000.0) / 1000.0 * noise;
let ny = (rng.below(2001) as f64 - 1000.0) / 1000.0 * noise;
out.push(((x, y), (v.x() / v.z() + nx, v.y() / v.z() + ny)));
}
let _ = f;
out
}
#[test]
fn dlt_recovers_a_known_homography_exactly() {
let r = Mat3::exp(Vec3::new(0.05, 0.3, 0.02));
let pairs = synthetic(1.0, r, 12, 0.0, 1);
let h = dlt(&pairs).expect("solvable");
for &(p, q) in &pairs {
let (x, y) = apply(&h, p).unwrap();
assert!((x - q.0).abs() < 1e-9 && (y - q.1).abs() < 1e-9);
}
}
#[test]
fn ransac_finds_the_consensus_among_outliers() {
let r = Mat3::exp(Vec3::new(-0.02, 0.25, 0.01));
let mut pairs = synthetic(1.0, r, 60, 0.0005, 2);
// Forty outliers: wrong second point.
let mut rng = Lcg(9);
for _ in 0..40 {
let k = rng.below(60);
let (p, _) = pairs[k];
pairs.push((p, ((rng.below(1000) as f64 - 500.0) / 1000.0, 0.1)));
}
let robust = ransac_homography(&pairs, 0.003, 500, 3).expect("found");
assert!(
robust.inliers.len() >= 55,
"{} inliers",
robust.inliers.len()
);
assert!(robust.inliers.iter().all(|&k| k < 60));
}
#[test]
fn focal_is_read_off_a_rotation_homography() {
// H in *pixel* coordinates for f = 1400: K R K⁻¹.
let f = 1400.0;
let r = Mat3::exp(Vec3::new(0.03, 0.35, -0.01));
let m = r.0;
let h = Mat3([
[m[0][0], m[0][1], m[0][2] * f],
[m[1][0], m[1][1], m[1][2] * f],
[m[2][0] / f, m[2][1] / f, m[2][2]],
]);
let est = focal_from_homography(&h).expect("estimable");
assert!((est - f).abs() / f < 1e-6, "{est}");
let back = rotation_from_homography(&h, f);
for (row, truth) in back.0.iter().zip(&m) {
for (a, b) in row.iter().zip(truth) {
assert!((a - b).abs() < 1e-9);
}
}
}
#[test]
fn the_identity_implies_no_focal() {
assert!(focal_from_homography(&Mat3::IDENTITY).is_none());
}
}
+262
View File
@@ -0,0 +1,262 @@
//! The grayscale proxy a detector reads.
//!
//! Alignment runs on proxies (FR-MRG-7) — a detector at 1024 px sees
//! everything it needs, and the full-resolution frames never leave the GPU.
//! This is that proxy: one channel, `f32` in `0.0..=1.0`, upright, and no
//! larger than the detector's fixed input.
/// A single-channel image, row-major, values in `0.0..=1.0`.
#[derive(Debug, Clone, PartialEq)]
pub struct Gray {
pub width: usize,
pub height: usize,
pub data: Vec<f32>,
}
impl Gray {
/// From tightly packed 8-bit RGBA, by the Rec. 709 luma weights.
///
/// The proxy is what a detector looks at, not what the photographer
/// sees, so which luma is used matters less than that it is the same one
/// for every frame — a keypoint's descriptor must not change between two
/// frames because they were converted differently.
pub fn from_rgba8(rgba: &[u8], width: usize, height: usize) -> Gray {
let n = width * height;
assert!(
rgba.len() >= n * 4,
"rgba buffer is short for {width}×{height}"
);
let data = rgba[..n * 4]
.chunks_exact(4)
.map(|p| {
(0.2126 * f32::from(p[0]) + 0.7152 * f32::from(p[1]) + 0.0722 * f32::from(p[2]))
/ 255.0
})
.collect();
Gray {
width,
height,
data,
}
}
/// Apply an EXIF orientation so the image is upright.
///
/// Learned detectors are not rotation-invariant — a descriptor of a
/// feature seen sideways is a different descriptor — and a portrait set
/// (the 6D fixture is one) would match poorly or not at all fed as
/// stored. The camera says which way is up; the proxy is turned before
/// anything looks at it, and the composite is written upright.
///
/// The value is the EXIF `Orientation` tag. Mirrored values (2, 4, 5, 7)
/// are not produced by any camera and are treated as their unmirrored
/// counterparts.
pub fn oriented(&self, orientation: u16) -> Gray {
match orientation {
3 | 4 => self.rotated_180(),
6 | 5 => self.rotated_90_cw(),
8 | 7 => self.rotated_90_ccw(),
_ => self.clone(),
}
}
fn rotated_90_cw(&self) -> Gray {
let (w, h) = (self.width, self.height);
let mut data = vec![0.0; w * h];
for y in 0..h {
for x in 0..w {
// Source (x, y) lands at (h - 1 - y, x) in an h-wide image.
data[x * h + (h - 1 - y)] = self.data[y * w + x];
}
}
Gray {
width: h,
height: w,
data,
}
}
fn rotated_90_ccw(&self) -> Gray {
let (w, h) = (self.width, self.height);
let mut data = vec![0.0; w * h];
for y in 0..h {
for x in 0..w {
// Source (x, y) lands at (y, w - 1 - x) in an h-wide image.
data[(w - 1 - x) * h + y] = self.data[y * w + x];
}
}
Gray {
width: h,
height: w,
data,
}
}
fn rotated_180(&self) -> Gray {
let mut data = self.data.clone();
data.reverse();
Gray {
width: self.width,
height: self.height,
data,
}
}
/// Resample to exactly `width × height` by area averaging on the way
/// down and bilinear on the way up.
///
/// Area averaging, not point sampling, for a reduction: a 5472 px frame
/// to 1024 is a factor of five, and picking one source pixel in
/// twenty-five aliases every edge the detector is looking for.
pub fn resampled(&self, width: usize, height: usize) -> Gray {
if width == self.width && height == self.height {
return self.clone();
}
let mut data = vec![0.0f32; width * height];
let sx = self.width as f64 / width as f64;
let sy = self.height as f64 / height as f64;
if sx >= 1.0 && sy >= 1.0 {
for oy in 0..height {
let y0 = (oy as f64 * sy) as usize;
let y1 = (((oy + 1) as f64 * sy) as usize).clamp(y0 + 1, self.height);
for ox in 0..width {
let x0 = (ox as f64 * sx) as usize;
let x1 = (((ox + 1) as f64 * sx) as usize).clamp(x0 + 1, self.width);
let mut sum = 0.0f32;
for y in y0..y1 {
let row = &self.data[y * self.width..(y + 1) * self.width];
sum += row[x0..x1].iter().sum::<f32>();
}
data[oy * width + ox] = sum / ((y1 - y0) * (x1 - x0)) as f32;
}
}
} else {
for oy in 0..height {
let fy = ((oy as f64 + 0.5) * sy - 0.5).max(0.0);
let y0 = (fy as usize).min(self.height - 1);
let y1 = (y0 + 1).min(self.height - 1);
let ty = (fy - y0 as f64) as f32;
for ox in 0..width {
let fx = ((ox as f64 + 0.5) * sx - 0.5).max(0.0);
let x0 = (fx as usize).min(self.width - 1);
let x1 = (x0 + 1).min(self.width - 1);
let tx = (fx - x0 as f64) as f32;
let p = |x: usize, y: usize| self.data[y * self.width + x];
let top = p(x0, y0) * (1.0 - tx) + p(x1, y0) * tx;
let bot = p(x0, y1) * (1.0 - tx) + p(x1, y1) * tx;
data[oy * width + ox] = top * (1.0 - ty) + bot * ty;
}
}
}
Gray {
width,
height,
data,
}
}
/// Scale so the image fits inside `max_width × max_height`, preserving
/// aspect, never enlarging. Returns the image and the scale applied,
/// which is what maps a proxy keypoint back to the source.
pub fn fitted(&self, max_width: usize, max_height: usize) -> (Gray, f64) {
let scale = (max_width as f64 / self.width as f64)
.min(max_height as f64 / self.height as f64)
.min(1.0);
let w = ((self.width as f64 * scale).round() as usize).max(1);
let h = ((self.height as f64 * scale).round() as usize).max(1);
(self.resampled(w, h), w as f64 / self.width as f64)
}
/// Copy into the top-left of a `width × height` canvas, zero elsewhere.
///
/// The detector's input is a fixed shape (S15.2), and a frame that fits
/// inside it is padded rather than stretched: stretching changes the
/// aspect and with it every descriptor.
pub fn padded(&self, width: usize, height: usize) -> Gray {
assert!(self.width <= width && self.height <= height);
let mut data = vec![0.0; width * height];
for y in 0..self.height {
data[y * width..y * width + self.width]
.copy_from_slice(&self.data[y * self.width..(y + 1) * self.width]);
}
Gray {
width,
height,
data,
}
}
}
#[cfg(test)]
mod tests {
use super::*;
fn ramp(w: usize, h: usize) -> Gray {
Gray {
width: w,
height: h,
data: (0..w * h).map(|i| i as f32).collect(),
}
}
#[test]
fn rotating_four_quarter_turns_is_the_identity() {
let g = ramp(5, 3);
let mut r = g.clone();
for _ in 0..4 {
r = r.rotated_90_cw();
}
assert_eq!(r, g);
assert_eq!(g.rotated_90_cw().rotated_90_ccw(), g);
assert_eq!(g.rotated_180().rotated_180(), g);
}
#[test]
fn a_clockwise_turn_moves_the_top_left_to_the_top_right() {
// 2×3 image, pixel values by position.
let g = ramp(2, 3);
let r = g.rotated_90_cw();
assert_eq!((r.width, r.height), (3, 2));
// Top-left of source (value 0) is at top-right of result.
assert_eq!(r.data[2], 0.0);
// Bottom-left of source (value 4) is at top-left of result.
assert_eq!(r.data[0], 4.0);
}
#[test]
fn orientation_8_is_a_counter_clockwise_turn() {
let g = ramp(4, 2);
assert_eq!(g.oriented(8), g.rotated_90_ccw());
assert_eq!(g.oriented(6), g.rotated_90_cw());
assert_eq!(g.oriented(1), g);
}
#[test]
fn downsampling_by_two_averages_blocks() {
let g = Gray {
width: 4,
height: 2,
data: vec![0.0, 1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0],
};
let r = g.resampled(2, 1);
assert_eq!(r.data, vec![2.5, 4.5]);
}
#[test]
fn fitting_never_enlarges_and_reports_the_scale() {
let g = ramp(100, 50);
let (f, s) = g.fitted(1024, 768);
assert_eq!((f.width, f.height), (100, 50));
assert_eq!(s, 1.0);
let (f, s) = g.fitted(50, 50);
assert_eq!((f.width, f.height), (50, 25));
assert_eq!(s, 0.5);
}
#[test]
fn padding_places_the_image_at_the_origin() {
let g = ramp(2, 2);
let p = g.padded(3, 3);
assert_eq!(p.data, vec![0.0, 1.0, 0.0, 2.0, 3.0, 0.0, 0.0, 0.0, 0.0]);
}
}
+86
View File
@@ -0,0 +1,86 @@
//! TRACES: FR-MRG-1 | FR-MRG-10
//! Panorama geometry — from several frames to the rotations that relate
//! them, and the projections that lay them out.
//!
//! This is the CPU half of a merge (FR-MRG-10): keypoints, matching, the
//! rotation solve and the choice of output surface. The per-pixel half —
//! rendering, warping, seams, blending — is the GPU's and lives in
//! `dr-gpu`, driven from above; nothing here touches a full-resolution
//! pixel. The split is the whole design (panorama.md §4): everything in
//! this crate is bounded by the number of frames, not the size of the
//! composite, and runs on proxies.
//!
//! # Layout
//!
//! - [`image`] — the grayscale proxy a detector reads: oriented, resampled.
//! - [`features`] — keypoints with descriptors, and the XFeat decoder.
//! - [`xfeat`] — the network under tract (feature `xfeat`).
//! - [`matching`] — mutual nearest neighbours.
//! - [`homography`] — a robust pairwise homography, the focal length read
//! off it, and the rotation it implies.
//! - [`bundle`] — every rotation and the focal length refined together.
//! - [`align`] — the whole thing, from features to cameras, honest about
//! what it could not place.
//! - [`projection`] — perspective, cylindrical, spherical.
//! - [`linalg`] — the small dense algebra all of it uses.
//!
//! # What it depends on
//!
//! Nothing, without the `xfeat` feature: the geometry is pure Rust with
//! hand-rolled linear algebra (`linalg` says why) so that it tests without
//! a model, a GPU or a device, on synthetic sets whose answer is known
//! exactly. With the feature it adds the same `ort`-over-tract runtime the
//! rest of the application already carries.
pub mod align;
pub mod bundle;
pub mod features;
pub mod fill;
pub mod homography;
pub mod image;
pub mod linalg;
pub mod matching;
#[cfg(feature = "xfeat")]
pub mod migan;
pub mod projection;
#[cfg(feature = "xfeat")]
pub mod xfeat;
pub use align::{align, AlignOptions, Alignment, Link, Unaligned};
pub use bundle::Cameras;
pub use features::{Features, Keypoint};
pub use fill::{fill_border, Inpainter, Observer, Params as FillParams};
pub use image::Gray;
pub use projection::Projection;
#[derive(Debug, thiserror::Error)]
pub enum PanoError {
#[error("bad input: {0}")]
Input(String),
#[error("geometry: {0}")]
Geometry(String),
#[error("model: {0}")]
Model(String),
#[error("could not read the model: {0}")]
ModelRead(#[source] std::io::Error),
#[cfg(feature = "xfeat")]
#[error("inference: {0}")]
Inference(#[source] ort::Error),
}
#[cfg(feature = "xfeat")]
impl From<ort::Error> for PanoError {
fn from(e: ort::Error) -> Self {
PanoError::Inference(e)
}
}
#[cfg(feature = "xfeat")]
impl From<dr_inference_engine::Error> for PanoError {
fn from(e: dr_inference_engine::Error) -> Self {
match e {
dr_inference_engine::Error::Inference(e) => PanoError::Inference(e),
dr_inference_engine::Error::Io(e) => PanoError::ModelRead(e),
}
}
}
+371
View File
@@ -0,0 +1,371 @@
//! The small dense linear algebra the geometry needs, and nothing more.
//!
//! Hand-rolled rather than pulled in, and the decision was made on purpose
//! (2026-09-19): the largest system this crate ever solves is a rotation
//! per frame plus one focal length — forty unknowns for a dozen frames —
//! and everything else is three-vectors. A general linear-algebra crate
//! would be the largest dependency in `dr-pano` by an order of magnitude,
//! for a Cholesky factorisation that is thirty lines.
//!
//! `f64` throughout. The geometry is solved once per merge on a few thousand
//! matches; there is no reason to give up precision for speed here, and the
//! bundle adjustment's normal equations are poorly conditioned enough near
//! convergence that `f32` would stall it.
use std::ops::{Add, Index, IndexMut, Mul, Neg, Sub};
/// A vector in three dimensions.
#[derive(Debug, Clone, Copy, PartialEq, Default)]
pub struct Vec3(pub [f64; 3]);
impl Vec3 {
pub const fn new(x: f64, y: f64, z: f64) -> Self {
Vec3([x, y, z])
}
pub fn dot(self, o: Vec3) -> f64 {
self.0[0] * o.0[0] + self.0[1] * o.0[1] + self.0[2] * o.0[2]
}
pub fn cross(self, o: Vec3) -> Vec3 {
Vec3([
self.0[1] * o.0[2] - self.0[2] * o.0[1],
self.0[2] * o.0[0] - self.0[0] * o.0[2],
self.0[0] * o.0[1] - self.0[1] * o.0[0],
])
}
pub fn norm(self) -> f64 {
self.dot(self).sqrt()
}
/// The unit vector along `self`, or `self` unchanged if it is zero.
pub fn normalised(self) -> Vec3 {
let n = self.norm();
if n > 0.0 {
self * (1.0 / n)
} else {
self
}
}
pub fn x(self) -> f64 {
self.0[0]
}
pub fn y(self) -> f64 {
self.0[1]
}
pub fn z(self) -> f64 {
self.0[2]
}
}
impl Add for Vec3 {
type Output = Vec3;
fn add(self, o: Vec3) -> Vec3 {
Vec3([self.0[0] + o.0[0], self.0[1] + o.0[1], self.0[2] + o.0[2]])
}
}
impl Sub for Vec3 {
type Output = Vec3;
fn sub(self, o: Vec3) -> Vec3 {
Vec3([self.0[0] - o.0[0], self.0[1] - o.0[1], self.0[2] - o.0[2]])
}
}
impl Mul<f64> for Vec3 {
type Output = Vec3;
fn mul(self, s: f64) -> Vec3 {
Vec3([self.0[0] * s, self.0[1] * s, self.0[2] * s])
}
}
impl Neg for Vec3 {
type Output = Vec3;
fn neg(self) -> Vec3 {
Vec3([-self.0[0], -self.0[1], -self.0[2]])
}
}
/// A 3×3 matrix, row-major.
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct Mat3(pub [[f64; 3]; 3]);
impl Mat3 {
pub const IDENTITY: Mat3 = Mat3([[1.0, 0.0, 0.0], [0.0, 1.0, 0.0], [0.0, 0.0, 1.0]]);
/// The matrix whose columns are `a`, `b`, `c`.
pub fn from_columns(a: Vec3, b: Vec3, c: Vec3) -> Mat3 {
Mat3([
[a.0[0], b.0[0], c.0[0]],
[a.0[1], b.0[1], c.0[1]],
[a.0[2], b.0[2], c.0[2]],
])
}
pub fn transpose(self) -> Mat3 {
let m = self.0;
Mat3([
[m[0][0], m[1][0], m[2][0]],
[m[0][1], m[1][1], m[2][1]],
[m[0][2], m[1][2], m[2][2]],
])
}
pub fn column(self, i: usize) -> Vec3 {
Vec3([self.0[0][i], self.0[1][i], self.0[2][i]])
}
pub fn trace(self) -> f64 {
self.0[0][0] + self.0[1][1] + self.0[2][2]
}
/// The rotation about `axis` (any length) by `angle` radians — Rodrigues.
pub fn rotation(axis: Vec3, angle: f64) -> Mat3 {
let k = axis.normalised();
let (s, c) = angle.sin_cos();
let t = 1.0 - c;
let (x, y, z) = (k.0[0], k.0[1], k.0[2]);
Mat3([
[t * x * x + c, t * x * y - s * z, t * x * z + s * y],
[t * x * y + s * z, t * y * y + c, t * y * z - s * x],
[t * x * z - s * y, t * y * z + s * x, t * z * z + c],
])
}
/// The rotation whose axis-angle vector is `w` (direction is the axis,
/// length is the angle). The exponential map; [`Self::log`] inverts it.
pub fn exp(w: Vec3) -> Mat3 {
let angle = w.norm();
if angle < 1e-12 {
// First-order: I + [w]×, which is what the limit is and avoids
// dividing by the angle.
let (x, y, z) = (w.0[0], w.0[1], w.0[2]);
return Mat3([[1.0, -z, y], [z, 1.0, -x], [-y, x, 1.0]]);
}
Mat3::rotation(w, angle)
}
/// The axis-angle vector of a rotation matrix. Inverse of [`Self::exp`].
pub fn log(self) -> Vec3 {
let m = self.0;
let cos = ((self.trace() - 1.0) * 0.5).clamp(-1.0, 1.0);
let axis = Vec3([m[2][1] - m[1][2], m[0][2] - m[2][0], m[1][0] - m[0][1]]);
if cos > 1.0 - 1e-6 {
// Small angle: `acos` near 1 loses everything below ~1e-8 to
// rounding, but the antisymmetric part is `2 sin θ · axis` and
// keeps it. First order, exact to the precision that matters.
return axis * 0.5;
}
let angle = cos.acos();
if angle > std::f64::consts::PI - 1e-6 {
// Near π the antisymmetric part vanishes; take the axis from the
// symmetric part instead. Rare for a panorama, but the solver may
// pass through it on a bad start and must not return NaN.
let d = Vec3([
((m[0][0] + 1.0) * 0.5).max(0.0).sqrt(),
((m[1][1] + 1.0) * 0.5).max(0.0).sqrt(),
((m[2][2] + 1.0) * 0.5).max(0.0).sqrt(),
]);
return d.normalised() * angle;
}
axis * (angle / (2.0 * angle.sin()))
}
/// Re-orthonormalise a matrix that has drifted from a rotation through
/// accumulated products. Gram–Schmidt on the columns; cheap and adequate
/// for drift of the size floating-point products produce.
pub fn orthonormalised(self) -> Mat3 {
let a = self.column(0).normalised();
let b = (self.column(1) - a * a.dot(self.column(1))).normalised();
let c = a.cross(b);
Mat3::from_columns(a, b, c)
}
}
impl Mul<Vec3> for Mat3 {
type Output = Vec3;
fn mul(self, v: Vec3) -> Vec3 {
let m = self.0;
Vec3([
m[0][0] * v.0[0] + m[0][1] * v.0[1] + m[0][2] * v.0[2],
m[1][0] * v.0[0] + m[1][1] * v.0[1] + m[1][2] * v.0[2],
m[2][0] * v.0[0] + m[2][1] * v.0[1] + m[2][2] * v.0[2],
])
}
}
impl Mul for Mat3 {
type Output = Mat3;
fn mul(self, o: Mat3) -> Mat3 {
let mut r = [[0.0; 3]; 3];
for (i, row) in r.iter_mut().enumerate() {
for (j, cell) in row.iter_mut().enumerate() {
*cell = (0..3).map(|k| self.0[i][k] * o.0[k][j]).sum();
}
}
Mat3(r)
}
}
/// A dense square matrix, for the normal equations.
#[derive(Debug, Clone, PartialEq)]
pub struct DMat {
n: usize,
data: Vec<f64>,
}
impl DMat {
pub fn zeros(n: usize) -> DMat {
DMat {
n,
data: vec![0.0; n * n],
}
}
pub fn n(&self) -> usize {
self.n
}
/// Solve `self · x = b` for a symmetric positive-definite `self` by
/// Cholesky factorisation. `None` if the matrix is not positive definite,
/// which for the normal equations means the problem is not determined by
/// the data — a frame with no matches, for instance — and the caller
/// should say so rather than proceed.
///
/// Destroys neither input: the factor is built in a copy. The systems
/// here are at most a few dozen unknowns and the copy is nothing.
pub fn solve_spd(&self, b: &[f64]) -> Option<Vec<f64>> {
let n = self.n;
debug_assert_eq!(b.len(), n);
let mut l = vec![0.0; n * n];
for j in 0..n {
let mut d = self[(j, j)];
for k in 0..j {
d -= l[j * n + k] * l[j * n + k];
}
if d <= 0.0 || !d.is_finite() {
return None;
}
let djj = d.sqrt();
l[j * n + j] = djj;
for i in j + 1..n {
let mut s = self[(i, j)];
for k in 0..j {
s -= l[i * n + k] * l[j * n + k];
}
l[i * n + j] = s / djj;
}
}
// Forward: L y = b.
let mut y = vec![0.0; n];
for i in 0..n {
let mut s = b[i];
for k in 0..i {
s -= l[i * n + k] * y[k];
}
y[i] = s / l[i * n + i];
}
// Back: Lᵀ x = y.
let mut x = vec![0.0; n];
for i in (0..n).rev() {
let mut s = y[i];
for k in i + 1..n {
s -= l[k * n + i] * x[k];
}
x[i] = s / l[i * n + i];
}
Some(x)
}
}
impl Index<(usize, usize)> for DMat {
type Output = f64;
fn index(&self, (i, j): (usize, usize)) -> &f64 {
&self.data[i * self.n + j]
}
}
impl IndexMut<(usize, usize)> for DMat {
fn index_mut(&mut self, (i, j): (usize, usize)) -> &mut f64 {
&mut self.data[i * self.n + j]
}
}
#[cfg(test)]
mod tests {
use super::*;
fn close(a: f64, b: f64) -> bool {
(a - b).abs() < 1e-9
}
#[test]
fn exp_and_log_are_inverses() {
for w in [
Vec3::new(0.1, -0.2, 0.3),
Vec3::new(1.0, 0.0, 0.0),
Vec3::new(0.0, 0.0, 2.5),
Vec3::new(1e-9, 0.0, 0.0),
] {
let back = Mat3::exp(w).log();
for i in 0..3 {
assert!(close(back.0[i], w.0[i]), "{w:?} -> {back:?}");
}
}
}
#[test]
fn a_rotation_is_orthonormal_and_preserves_length() {
let r = Mat3::exp(Vec3::new(0.4, 0.5, -0.6));
let rt = r.transpose() * r;
for i in 0..3 {
for j in 0..3 {
assert!(close(rt.0[i][j], Mat3::IDENTITY.0[i][j]));
}
}
let v = Vec3::new(1.0, 2.0, 3.0);
assert!(close((r * v).norm(), v.norm()));
}
#[test]
fn rotation_about_z_turns_x_towards_y() {
let r = Mat3::rotation(Vec3::new(0.0, 0.0, 1.0), std::f64::consts::FRAC_PI_2);
let v = r * Vec3::new(1.0, 0.0, 0.0);
assert!(close(v.x(), 0.0) && close(v.y(), 1.0) && close(v.z(), 0.0));
}
#[test]
fn cholesky_solves_a_small_spd_system() {
// A = Bᵀ B for a random-ish B is SPD by construction.
let b = [
[2.0, 1.0, 0.0],
[1.0, 3.0, 1.0],
[0.0, 1.0, 4.0],
[1.0, 1.0, 1.0],
];
let mut a = DMat::zeros(3);
for i in 0..3 {
for j in 0..3 {
a[(i, j)] = (0..4).map(|k| b[k][i] * b[k][j]).sum();
}
}
let x_true = [1.0, -2.0, 0.5];
let rhs: Vec<f64> = (0..3)
.map(|i| (0..3).map(|j| a[(i, j)] * x_true[j]).sum())
.collect();
let x = a.solve_spd(&rhs).expect("spd");
for i in 0..3 {
assert!(close(x[i], x_true[i]), "{x:?}");
}
}
#[test]
fn cholesky_refuses_an_indefinite_matrix() {
let mut a = DMat::zeros(2);
a[(0, 0)] = 1.0;
a[(1, 1)] = -1.0;
assert!(a.solve_spd(&[1.0, 1.0]).is_none());
}
}
+173
View File
@@ -0,0 +1,173 @@
//! Descriptor matching between two images.
//!
//! Mutual nearest neighbour on cosine similarity, with a floor on the
//! similarity — the reference XFeat's own matcher (`match_mkpts`,
//! `min_cossim = 0.82`). For a panorama that is enough: one lens, one
//! scene, near-pure rotation and 20–40 % overlap make the matching problem
//! easy, and what is hard — sky, repeated structure, exposure drift — is
//! handled by the detector's descriptors and by RANSAC downstream, not by a
//! cleverer matcher. A learned matcher (LightGlue) is the step after this
//! one fails on a real set, and it has not (panorama.md §6).
//!
//! Brute force. `4096 × 4096 × 64` multiply-adds is a billion per pair,
//! and a twelve-frame set has sixty-six pairs: a minute single-threaded
//! and scalar (measured 2026-09-19: 51 s), a few seconds vectorised across
//! the cores. Not worth an index, but worth doing properly.
use crate::features::{Features, DESCRIPTOR_LEN};
const _: () = assert!(DESCRIPTOR_LEN.is_multiple_of(8));
/// A correspondence: keypoint `a` in the first image matches keypoint `b`
/// in the second, with the cosine similarity of their descriptors.
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct Match {
pub a: usize,
pub b: usize,
pub similarity: f32,
}
/// Match two sets of features.
///
/// A pair is kept when each is the other's nearest neighbour and their
/// similarity is at least `min_similarity`.
pub fn match_features(a: &Features, b: &Features, min_similarity: f32) -> Vec<Match> {
if a.is_empty() || b.is_empty() {
return Vec::new();
}
let (na, nb) = (a.len(), b.len());
// The whole similarity matrix, once. Both nearest-neighbour directions
// read it, which halves the multiply-adds against computing each
// direction on its own; 4096 × 4096 × f32 is 64 MB, transient.
let mut sim = vec![0.0f32; na * nb];
let threads = std::thread::available_parallelism()
.map(usize::from)
.unwrap_or(1)
.clamp(1, 16);
let rows_per = na.div_ceil(threads);
std::thread::scope(|scope| {
for (t, chunk) in sim.chunks_mut(rows_per * nb).enumerate() {
scope.spawn(move || {
let first = t * rows_per;
for (r, row) in chunk.chunks_mut(nb).enumerate() {
let da = a.descriptor(first + r);
for (j, cell) in row.iter_mut().enumerate() {
*cell = dot(da, b.descriptor(j));
}
}
});
}
});
// Best in `b` for each `a`, and best in `a` for each `b`.
let best_ab: Vec<(usize, f32)> = sim
.chunks_exact(nb)
.map(|row| {
row.iter().enumerate().fold(
(0usize, f32::MIN),
|acc, (j, &s)| if s > acc.1 { (j, s) } else { acc },
)
})
.collect();
let mut best_ba = vec![(0usize, f32::MIN); nb];
for (i, row) in sim.chunks_exact(nb).enumerate() {
for (j, &s) in row.iter().enumerate() {
if s > best_ba[j].1 {
best_ba[j] = (i, s);
}
}
}
best_ab
.iter()
.enumerate()
.filter_map(|(ia, &(ib, s))| {
(best_ba[ib].0 == ia && s >= min_similarity).then_some(Match {
a: ia,
b: ib,
similarity: s,
})
})
.collect()
}
#[inline]
fn dot(a: &[f32], b: &[f32]) -> f32 {
// Eight independent accumulators over exact 8-lane chunks: the shape
// the compiler turns into one vector multiply-add per chunk, and no
// bounds checks inside the loop. `DESCRIPTOR_LEN` is a multiple of 8.
let (a, b) = (&a[..DESCRIPTOR_LEN], &b[..DESCRIPTOR_LEN]);
let mut acc = [0.0f32; 8];
for (ca, cb) in a.chunks_exact(8).zip(b.chunks_exact(8)) {
for k in 0..8 {
acc[k] += ca[k] * cb[k];
}
}
acc.iter().sum()
}
#[cfg(test)]
mod tests {
use super::*;
use crate::features::Keypoint;
/// Features whose descriptors are unit vectors along the given axes.
fn along(axes: &[usize]) -> Features {
let mut descriptors = vec![0.0; axes.len() * DESCRIPTOR_LEN];
for (i, &ax) in axes.iter().enumerate() {
descriptors[i * DESCRIPTOR_LEN + ax] = 1.0;
}
Features {
keypoints: axes
.iter()
.map(|_| Keypoint {
x: 0.0,
y: 0.0,
score: 1.0,
})
.collect(),
descriptors,
width: 1,
height: 1,
}
}
#[test]
fn identical_descriptors_match_mutually() {
let a = along(&[0, 1, 2]);
let b = along(&[2, 0, 1]);
let m = match_features(&a, &b, 0.8);
let mut pairs: Vec<(usize, usize)> = m.iter().map(|m| (m.a, m.b)).collect();
pairs.sort();
assert_eq!(pairs, vec![(0, 1), (1, 2), (2, 0)]);
assert!(m.iter().all(|m| (m.similarity - 1.0).abs() < 1e-6));
}
#[test]
fn a_descriptor_with_no_counterpart_is_unmatched() {
let a = along(&[0, 1, 5]);
let b = along(&[0, 1]);
let m = match_features(&a, &b, 0.8);
assert_eq!(m.len(), 2);
assert!(m.iter().all(|m| m.a != 2));
}
#[test]
fn mutuality_breaks_a_one_sided_match() {
// b0 is the nearest to both a0 and a1, but a0 is its nearest — a1
// must not be matched to it.
let mut a = along(&[0, 0]);
a.descriptors[DESCRIPTOR_LEN] = 0.9;
a.descriptors[DESCRIPTOR_LEN + 1] = (1.0f32 - 0.81).sqrt();
let b = along(&[0]);
let m = match_features(&a, &b, 0.0);
assert_eq!(m.len(), 1);
assert_eq!((m[0].a, m[0].b), (0, 0));
}
#[test]
fn empty_input_is_empty_output() {
assert!(match_features(&along(&[]), &along(&[1]), 0.5).is_empty());
}
}
+107
View File
@@ -0,0 +1,107 @@
//! TRACES: FR-MRG-4
//! MI-GAN, the border filler, under the inference engine.
//!
//! Sargsyan et al., ICCV 2023 (Picsart AI Research): inpainting built for
//! phones — about six million parameters of plain convolutions, no FFT and
//! no attention, so it quantises and runs on a DSP. MIT, code and weights
//! (`models/LICENCE.md`). The bare 512 generator is what ships, exported
//! at a fixed shape by `tools/export-migan.sh`; its six operator types load
//! on every rung, and what they cost is the whole story of whether a fill
//! is interactive: 7.4 s a tile under tract, 0.4 s under ONNX Runtime's
//! CPU pool, 23 ms in fp16 and 13 ms in int8 on a laptop's TensorRT
//! (2026-09-19, docs/dev/panorama.md §12).
//!
//! The model's contract, from the reference `export_inference_model.py`:
//! input `1×4×512×512` float — channel 0 is `mask − 0.5` with 1 where the
//! picture is known, channels 1–3 the RGB in −1..1 with the unknown pixels
//! zeroed; output `1×3×512×512` in −1..1, of which the caller keeps the
//! unknown pixels. That is [`crate::fill::Inpainter`], and the rest —
//! which tiles, what context, how to blend — is `fill.rs`.
use crate::fill::Inpainter;
use crate::PanoError;
/// The tile the shipped export takes.
pub const TILE: usize = 512;
pub struct MiGan {
model: dr_inference_engine::Model,
}
impl MiGan {
/// From the model file, in whichever form the engine's rung wants
/// (`resolve_model` picks an int8 sibling for the Hexagon).
pub fn from_path(path: &std::path::Path) -> Result<Self, PanoError> {
use dr_inference_engine::{resolve_model, Role};
let (path, form) = resolve_model(Role::Inpainter, path);
let bytes = std::fs::read(&path).map_err(PanoError::ModelRead)?;
Self::from_bytes(&bytes, form)
}
pub fn from_bytes(bytes: &[u8], form: dr_inference_engine::Form) -> Result<Self, PanoError> {
use dr_inference_engine::Role;
Ok(MiGan {
model: dr_inference_engine::open(Role::Inpainter, form, bytes)?,
})
}
/// Where the fill runs, for a status line.
pub fn rung(&self) -> Result<dr_inference_engine::Rung, PanoError> {
Ok(self.model.acquire()?.rung())
}
}
impl Inpainter for MiGan {
fn tile(&self) -> usize {
TILE
}
fn fill(&mut self, rgb: &[f32], known: &[bool]) -> Result<Vec<f32>, PanoError> {
let n = TILE * TILE;
if rgb.len() != n * 3 || known.len() != n {
return Err(PanoError::Input(format!(
"MI-GAN takes a {TILE}×{TILE} tile; given {} values and {} mask entries",
rgb.len(),
known.len()
)));
}
// NCHW: the mask plane, then the three masked colour planes.
let mut input = vec![0.0f32; 4 * n];
for i in 0..n {
let m = if known[i] { 1.0 } else { 0.0 };
input[i] = m - 0.5;
for c in 0..3 {
input[(c + 1) * n + i] = (rgb[i * 3 + c] * 2.0 - 1.0) * m;
}
}
let tensor = ort::value::Tensor::from_array(
ndarray::Array::from_shape_vec(ndarray::IxDyn(&[1, 4, TILE, TILE]), input)
.expect("shape matches by construction"),
)?;
let started = std::time::Instant::now();
let acquired = self.model.acquire()?;
let acquired_at = started.elapsed();
let mut session = acquired.lock();
let outputs = session.run(ort::inputs![tensor])?;
log::trace!(
"migan: tile on {} — acquire {:.1} ms, run {:.1} ms",
acquired.rung().label(),
acquired_at.as_secs_f64() * 1e3,
(started.elapsed() - acquired_at).as_secs_f64() * 1e3
);
let (shape, data) = outputs[0].try_extract_tensor::<f32>()?;
let dims: Vec<i64> = shape.iter().copied().collect();
if dims != [1, 3, TILE as i64, TILE as i64] {
return Err(PanoError::Model(format!(
"MI-GAN output is {dims:?}, expected [1, 3, {TILE}, {TILE}]"
)));
}
let mut out = vec![0.0f32; n * 3];
for i in 0..n {
for c in 0..3 {
out[i * 3 + c] = (data[c * n + i] * 0.5 + 0.5).clamp(0.0, 1.0);
}
}
Ok(out)
}
}
+197
View File
@@ -0,0 +1,197 @@
//! TRACES: FR-MRG-4
//! The surface the composite is drawn on.
//!
//! A panorama is a set of directions; a picture is a plane. The projection
//! is the map between them, and the three offered are the three every
//! stitcher offers because each is right for a different field of view:
//! perspective keeps straight lines straight and cannot reach 180°;
//! cylindrical keeps verticals vertical and stretches nothing horizontally,
//! for the wide single row; spherical for anything that also looks up.
//!
//! Every function here is the *inverse* map — output pixel to direction —
//! because that is what a gather needs (`lens.rs` in `dr-pipeline` says
//! why a warp is written that way), and it is the function the WGSL warp
//! will repeat verbatim. The forward map exists for bounds only.
use crate::linalg::Vec3;
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum Projection {
Perspective,
Cylindrical,
Spherical,
}
impl Projection {
/// Which projection a field of view calls for.
///
/// Perspective stretches the edges by `1 / cos` of the angle from the
/// centre, which is 2× at 60° and unbounded at 90°; the switch is where
/// that stretch starts to look like a mistake. Spherical is for a set
/// that spans enough vertically that a cylinder would stretch the top
/// and bottom the same way.
pub fn suggest(horizontal_fov: f64, vertical_fov: f64) -> Projection {
if horizontal_fov < 70f64.to_radians() && vertical_fov < 70f64.to_radians() {
Projection::Perspective
} else if vertical_fov < 100f64.to_radians() {
Projection::Cylindrical
} else {
Projection::Spherical
}
}
/// The direction an output point looks along. `scale` is the output's
/// focal length in pixels: the radius of the cylinder or sphere, or the
/// plane's distance. Coordinates are centred on the projection's origin
/// (the direction `+z`).
pub fn to_direction(self, scale: f64, u: f64, v: f64) -> Vec3 {
match self {
Projection::Perspective => Vec3::new(u, v, scale).normalised(),
Projection::Cylindrical => {
let theta = u / scale;
Vec3::new(theta.sin(), v / scale, theta.cos()).normalised()
}
Projection::Spherical => {
let theta = u / scale;
let phi = v / scale;
Vec3::new(theta.sin() * phi.cos(), phi.sin(), theta.cos() * phi.cos())
}
}
}
/// Where a direction lands on the output, or `None` where the
/// projection cannot show it (behind a perspective plane, at a
/// cylinder's poles).
pub fn from_direction(self, scale: f64, d: Vec3) -> Option<(f64, f64)> {
let (x, y, z) = (d.x(), d.y(), d.z());
match self {
Projection::Perspective => (z > 1e-9).then(|| (scale * x / z, scale * y / z)),
Projection::Cylindrical => {
let r = (x * x + z * z).sqrt();
(r > 1e-9).then(|| (scale * x.atan2(z), scale * y / r))
}
Projection::Spherical => {
let r = (x * x + z * z).sqrt();
Some((scale * x.atan2(z), scale * y.atan2(r)))
}
}
}
}
/// The output rectangle a set of frames covers, in centred output pixels.
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct Bounds {
pub min_u: f64,
pub min_v: f64,
pub max_u: f64,
pub max_v: f64,
}
impl Bounds {
pub fn width(&self) -> f64 {
self.max_u - self.min_u
}
pub fn height(&self) -> f64 {
self.max_v - self.min_v
}
}
/// Bounds of the frames' footprints under `projection`, by walking each
/// frame's border.
///
/// `frame_size` is the frames' width and height in the same pixels the
/// cameras' focal length is in. The border is sampled rather than only its
/// corners because under a cylinder the widest point of a rolled frame is
/// not a corner.
pub fn bounds(
projection: Projection,
scale: f64,
cameras: &crate::bundle::Cameras,
frame_size: (f64, f64),
) -> Option<Bounds> {
let (w, h) = frame_size;
let mut b: Option<Bounds> = None;
let steps = 64;
for k in 0..cameras.rotations.len() {
for s in 0..steps {
let t = s as f64 / steps as f64;
for p in [
(-w / 2.0 + w * t, -h / 2.0),
(-w / 2.0 + w * t, h / 2.0),
(-w / 2.0, -h / 2.0 + h * t),
(w / 2.0, -h / 2.0 + h * t),
] {
let d = cameras.bearing(k, p);
let Some((u, v)) = projection.from_direction(scale, d) else {
continue;
};
b = Some(match b {
None => Bounds {
min_u: u,
min_v: v,
max_u: u,
max_v: v,
},
Some(b) => Bounds {
min_u: b.min_u.min(u),
min_v: b.min_v.min(v),
max_u: b.max_u.max(u),
max_v: b.max_v.max(v),
},
});
}
}
}
b
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn to_and_from_direction_are_inverses() {
for proj in [
Projection::Perspective,
Projection::Cylindrical,
Projection::Spherical,
] {
for (u, v) in [(0.0, 0.0), (300.0, -200.0), (-900.0, 450.0)] {
let d = proj.to_direction(1000.0, u, v);
let (bu, bv) = proj.from_direction(1000.0, d).expect("in front");
assert!(
(bu - u).abs() < 1e-9 && (bv - v).abs() < 1e-9,
"{proj:?} {u} {v}"
);
}
}
}
#[test]
fn the_origin_looks_down_z_in_every_projection() {
for proj in [
Projection::Perspective,
Projection::Cylindrical,
Projection::Spherical,
] {
let d = proj.to_direction(500.0, 0.0, 0.0);
assert!((d.z() - 1.0).abs() < 1e-12);
}
}
#[test]
fn a_cylinder_maps_ninety_degrees_to_a_quarter_turn_of_pixels() {
let d = Vec3::new(1.0, 0.0, 0.0);
let (u, v) = Projection::Cylindrical.from_direction(100.0, d).unwrap();
assert!((u - 100.0 * std::f64::consts::FRAC_PI_2).abs() < 1e-9);
assert_eq!(v, 0.0);
assert!(Projection::Perspective.from_direction(100.0, d).is_none());
}
#[test]
fn suggestion_widens_with_the_field() {
assert_eq!(Projection::suggest(0.5, 0.5), Projection::Perspective);
assert_eq!(Projection::suggest(2.5, 0.8), Projection::Cylindrical);
assert_eq!(Projection::suggest(3.0, 2.5), Projection::Spherical);
}
}
+151
View File
@@ -0,0 +1,151 @@
//! TRACES: FR-MRG-8
//! The XFeat detector — the network under tract, and the decoder after it.
//!
//! Apache-2.0 weights (`models/LICENCE.md`), exported at a fixed shape by
//! `tools/export-xfeat.sh` and loaded through the same `dr-inference-engine`
//! `dr-segment` and `dr-face` use, so this adds no runtime and no C to the
//! tree; what runs it is the device's business (docs/dev/inference.md). ~300 ms
//! per frame on tract on the reference desktop, ~400 ms on the tablet
//! (S15.2, S15.4).
use crate::features::{decode_xfeat, DecodeOptions, Features, XFeatMaps, DESCRIPTOR_LEN};
use crate::image::Gray;
use crate::PanoError;
/// The two input shapes the shipped exports were made for: one landscape,
/// one portrait, the same weights. A frame is fitted into whichever
/// matches its aspect, so a portrait set does not spend half the
/// detector's width on padding — which is what the 6D fixture did before
/// the second export existed (512 × 768 of a 1024 × 768 input). A
/// different size is a different file (`tools/export-xfeat.sh`).
pub const INPUT_LANDSCAPE: (usize, usize) = (1024, 768);
pub const INPUT_PORTRAIT: (usize, usize) = (768, 1024);
/// The long edge of the detector's input, for callers sizing a proxy.
pub const INPUT_LONG_EDGE: usize = 1024;
#[cfg(feature = "embedded-model")]
const EMBEDDED_LANDSCAPE: &[u8] = include_bytes!("../../../models/keypoints/xfeat-1024.onnx");
#[cfg(feature = "embedded-model")]
const EMBEDDED_PORTRAIT: &[u8] = include_bytes!("../../../models/keypoints/xfeat-768.onnx");
/// A loaded detector: the network at both shapes.
pub struct XFeat {
landscape: dr_inference_engine::Model,
portrait: dr_inference_engine::Model,
pub options: DecodeOptions,
}
/// The bytes of both exports compiled into the binary, for whoever compiles
/// engines ahead of the first request (docs/dev/inference.md §6).
#[cfg(feature = "embedded-model")]
pub fn embedded_model_bytes() -> [&'static [u8]; 2] {
[EMBEDDED_LANDSCAPE, EMBEDDED_PORTRAIT]
}
impl XFeat {
/// The weights compiled into the binary.
#[cfg(feature = "embedded-model")]
pub fn embedded() -> Result<Self, PanoError> {
Self::from_bytes(EMBEDDED_LANDSCAPE, EMBEDDED_PORTRAIT)
}
/// From the two exports on disk.
pub fn from_paths(
landscape: &std::path::Path,
portrait: &std::path::Path,
) -> Result<Self, PanoError> {
let l = std::fs::read(landscape).map_err(PanoError::ModelRead)?;
let p = std::fs::read(portrait).map_err(PanoError::ModelRead)?;
Self::from_bytes(&l, &p)
}
pub fn from_bytes(landscape: &[u8], portrait: &[u8]) -> Result<Self, PanoError> {
use dr_inference_engine::{Form, Role};
Ok(XFeat {
landscape: dr_inference_engine::open(Role::Keypoints, Form::F32, landscape)?,
portrait: dr_inference_engine::open(Role::Keypoints, Form::F32, portrait)?,
options: DecodeOptions::default(),
})
}
/// Detect keypoints in an upright grayscale image.
///
/// The image is fitted into the network's input of matching aspect —
/// scaled down if larger, never up, and padded to the right and bottom
/// — and the keypoints come back in the coordinates of `image` itself,
/// so a caller that already scaled a frame to a proxy maps them on with
/// the scale it used and nothing else.
pub fn detect(&mut self, image: &Gray) -> Result<Features, PanoError> {
let ((in_w, in_h), model) = if image.height > image.width {
(INPUT_PORTRAIT, &self.portrait)
} else {
(INPUT_LANDSCAPE, &self.landscape)
};
let acquired = model.acquire()?;
let mut session = acquired.lock();
let (fitted, scale) = image.fitted(in_w, in_h);
let padded = fitted.padded(in_w, in_h);
let input =
ndarray::Array::from_shape_vec(ndarray::IxDyn(&[1, 1, in_h, in_w]), padded.data)
.expect("shape matches the buffer by construction");
let tensor = ort::value::Tensor::from_array(input).map_err(PanoError::Inference)?;
let outputs = session
.run(ort::inputs![tensor])
.map_err(PanoError::Inference)?;
let (w8, h8) = (in_w / 8, in_h / 8);
let expect = |i: usize, channels: usize| -> Result<Vec<f32>, PanoError> {
let (shape, data) = outputs[i]
.try_extract_tensor::<f32>()
.map_err(PanoError::Inference)?;
let dims: Vec<i64> = shape.iter().copied().collect();
if dims != [1, channels as i64, h8 as i64, w8 as i64] {
return Err(PanoError::Model(format!(
"output {i} is {dims:?}, expected [1, {channels}, {h8}, {w8}] — \
not the export this decoder was written for"
)));
}
Ok(data.to_vec())
};
let feats = expect(0, DESCRIPTOR_LEN)?;
let keypoints = expect(1, 65)?;
let heatmap = expect(2, 1)?;
let mut features = decode_xfeat(
&XFeatMaps {
feats: &feats,
keypoints: &keypoints,
heatmap: &heatmap,
width: w8,
height: h8,
},
&self.options,
);
// Back to the caller's image: drop anything the padding produced,
// undo the fit.
let border = self.options.border as f32;
let limit_x = fitted.width as f32 - border;
let limit_y = fitted.height as f32 - border;
let mut kept_kp = Vec::with_capacity(features.len());
let mut kept_desc = Vec::with_capacity(features.descriptors.len());
for (i, kp) in features.keypoints.iter().enumerate() {
if kp.x >= limit_x || kp.y >= limit_y {
continue;
}
kept_kp.push(crate::features::Keypoint {
x: (kp.x / scale as f32),
y: (kp.y / scale as f32),
score: kp.score,
});
kept_desc.extend_from_slice(features.descriptor(i));
}
features.keypoints = kept_kp;
features.descriptors = kept_desc;
features.width = image.width;
features.height = image.height;
Ok(features)
}
}
+2 -2
View File
@@ -315,7 +315,7 @@ pub struct DetailPass {
/// 52 render pixels at 4K — holds no spatial frequency a quarter-scale
/// grid cannot represent. Computing it at the render size therefore buys
/// nothing and costs everything: 105 taps over 8.3 M pixels, twice, which
/// measured at 34 ms and is where `docs/technical-debt.md` TD-4 came from.
/// measured at 34 ms and is where `docs/dev/technical-debt.md` TD-4 came from.
/// At a quarter it is a sixteenth of the pixels at a quarter of the
/// radius, and the result is not an approximation of the full-resolution
/// base — it is the same band-limited function, sampled where it is still
@@ -568,7 +568,7 @@ pub fn compose_detail(
/// photograph the photographer thinks they are sharpening.
///
/// It also means ARCH §5.2's stage list, which draws spot removal after
/// texture and clarity, is not what this does — see `docs/spot-removal.md`
/// texture and clarity, is not what this does — see `docs/dev/spot-removal.md`
/// §5.1, which is where the disagreement is written down.
pub fn compose_detail_with(
ops: &[Box<dyn Operation>],
+681 -9
View File
@@ -27,6 +27,33 @@
//! chain exactly the space it documents: normalised, centred, `r == 1` at the
//! corner. Neither stage needs to know the other exists.
//!
//! # Perspective sits inside framing (FR-DEV-20)
//!
//! A keystone correction is composition too — straightening converging
//! verticals reframes the photograph — so it is a step *of* framing rather
//! than a stage beside it, and inherits framing's `Compose` attribute, its
//! place in the sidecar and its exclusion from a default paste. Expanded, the
//! chain reads:
//!
//! ```text
//! output pixel → crop → straighten → perspective → orientation → warp (lens) → sample
//! ```
//!
//! After the straightening, because the angle is a nudge applied to the
//! corrected picture: the verticals are made parallel and *then* the whole is
//! levelled. Before the stored orientation, because "vertical" means vertical
//! in the photograph as it is shown — a portrait frame the camera stored on
//! its side must converge along its displayed height, not along the sensor's
//! rows. And before the lens warp, which still sees the whole frame it
//! corrects, for the reason given above.
//!
//! The correction maps the output frame onto a trapezoid **inside** the
//! source rather than pulling the source edges in. So a keystone on its own
//! never exposes an empty corner, and the crop the user drew is still valid
//! after it; only in combination with a straightening angle does the
//! inscribed crop have anything to account for — see
//! [`Framing::max_inscribed_crop`].
//!
//! # Why sampling changes with the angle
//!
//! At 90° steps and flips, output pixels land exactly on source pixels, so
@@ -58,6 +85,13 @@ pub const CROP_X: ParamId = ParamId("crop_x");
pub const CROP_Y: ParamId = ParamId("crop_y");
pub const CROP_W: ParamId = ParamId("crop_w");
pub const CROP_H: ParamId = ParamId("crop_h");
/// TRACES: FR-DEV-20
/// Vertical keystone. Positive spreads the top of the frame — the correction
/// for a building photographed looking up, whose verticals lean together.
pub const KEYSTONE_V: ParamId = ParamId("keystone_v");
/// TRACES: FR-DEV-20
/// Horizontal keystone. Positive spreads the right-hand side of the frame.
pub const KEYSTONE_H: ParamId = ParamId("keystone_h");
/// Widest straightening the control offers, in degrees either way.
///
@@ -66,12 +100,28 @@ pub const CROP_H: ParamId = ParamId("crop_h");
/// the edits actually are.
pub const MAX_STRAIGHTEN: f32 = 45.0;
/// The keystone sliders' travel either way.
///
/// A plain amount rather than degrees of tilt: the angle a camera was tilted
/// by depends on a focal length the correction does not know, and a number
/// that claimed to be one would be wrong for every lens but one.
pub const MAX_KEYSTONE: f32 = 100.0;
/// How far a full keystone narrows the far edge of the frame, as a fraction
/// of its width: at `MAX_KEYSTONE` the source trapezoid's short side is half
/// its long one.
///
/// Enough for a tall building from its own pavement, and short of the point
/// where the stretched edge is so magnified that the correction reads as a
/// fault of its own.
const KEYSTONE_REACH: f64 = 0.5;
/// The parameters the framing widget owns — every one of them.
///
/// In the order the widget expects: the rect first, then the angle it is
/// straightened by, then the exact reorientations.
static FRAMING_PARAMS: [ParamId; 8] = [
CROP_X, CROP_Y, CROP_W, CROP_H, ANGLE, ROTATION, FLIP_H, FLIP_V,
static FRAMING_PARAMS: [ParamId; 10] = [
CROP_X, CROP_Y, CROP_W, CROP_H, ANGLE, KEYSTONE_V, KEYSTONE_H, ROTATION, FLIP_H, FLIP_V,
];
static DESCRIPTOR: LazyLock<Arc<OpDescriptor>> = LazyLock::new(|| {
@@ -117,6 +167,30 @@ static DESCRIPTOR: LazyLock<Arc<OpDescriptor>> = LazyLock::new(|| {
ParamDescriptor::fraction("crop_y", "param.crop_y", 0.0),
ParamDescriptor::fraction("crop_w", "param.crop_w", 1.0),
ParamDescriptor::fraction("crop_h", "param.crop_h", 1.0),
// TRACES: FR-DEV-20
// Perspective. Last so every sidecar written before these existed
// reads exactly as it did: a missing parameter is its default, and
// the default is no correction.
ParamDescriptor::scalar(
"keystone_v",
"param.keystone_v",
-MAX_KEYSTONE,
MAX_KEYSTONE,
0.0,
Unit::None,
Scale::Linear,
0,
),
ParamDescriptor::scalar(
"keystone_h",
"param.keystone_h",
-MAX_KEYSTONE,
MAX_KEYSTONE,
0.0,
Unit::None,
Scale::Linear,
0,
),
],
})
});
@@ -311,6 +385,86 @@ fn finite(v: f32, fallback: f32) -> f32 {
}
}
/// TRACES: FR-DEV-20
/// A plane projective map, row-major, acting on `(x, y, 1)`.
///
/// Held in `f64` because it is built by solving for four corners and then
/// inverted for [`Framing::output_at`]; the shader gets `f32` copies of the
/// forward map only.
#[derive(Debug, Clone, Copy, PartialEq)]
struct Homography([[f64; 3]; 3]);
impl Homography {
/// The map taking the square `[-0.5, 0.5]²` onto the quadrilateral whose
/// corners are `q`, listed top-left, top-right, bottom-right, bottom-left.
///
/// Heckbert's closed form for the unit square, composed with the shift
/// from the centred square onto it.
fn square_to_quad(q: [(f64, f64); 4]) -> Self {
let [(x0, y0), (x1, y1), (x2, y2), (x3, y3)] = q;
let (dx1, dx2, dx3) = (x1 - x2, x3 - x2, x0 - x1 + x2 - x3);
let (dy1, dy2, dy3) = (y1 - y2, y3 - y2, y0 - y1 + y2 - y3);
let den = dx1 * dy2 - dx2 * dy1;
let (g, h) = if den.abs() < 1e-12 {
(0.0, 0.0)
} else {
((dx3 * dy2 - dx2 * dy3) / den, (dx1 * dy3 - dx3 * dy1) / den)
};
// Unit square (u, v) -> quad.
let unit = [
[x1 - x0 + g * x1, x3 - x0 + h * x3, x0],
[y1 - y0 + g * y1, y3 - y0 + h * y3, y0],
[g, h, 1.0],
];
// Centred square -> unit square is `u = x + 0.5`, so fold the shift
// into the constant column.
let mut m = unit;
for row in &mut m {
row[2] += 0.5 * (row[0] + row[1]);
}
Self(m)
}
/// Where `(x, y)` lands, or `None` past the line the map sends to
/// infinity — a point with no image, which the caller treats as outside
/// the source.
fn apply(&self, (x, y): (f64, f64)) -> Option<(f64, f64)> {
let m = &self.0;
let w = m[2][0] * x + m[2][1] * y + m[2][2];
if w <= 1e-9 {
return None;
}
Some((
(m[0][0] * x + m[0][1] * y + m[0][2]) / w,
(m[1][0] * x + m[1][1] * y + m[1][2]) / w,
))
}
/// The inverse map, by the adjugate. Scale is irrelevant to a projective
/// map, so the determinant is only divided out to keep `w` positive and
/// near one — which is what [`Self::apply`]'s horizon test relies on.
fn inverse(&self) -> Self {
let m = &self.0;
let c = |r0: usize, c0: usize, r1: usize, c1: usize| {
m[r0][c0] * m[r1][c1] - m[r0][c1] * m[r1][c0]
};
let adj = [
[c(1, 1, 2, 2), -c(0, 1, 2, 2), c(0, 1, 1, 2)],
[-c(1, 0, 2, 2), c(0, 0, 2, 2), -c(0, 0, 1, 2)],
[c(1, 0, 2, 1), -c(0, 0, 2, 1), c(0, 0, 1, 1)],
];
let det = m[0][0] * adj[0][0] + m[0][1] * adj[1][0] + m[0][2] * adj[2][0];
let det = if det.abs() < 1e-12 { 1.0 } else { det };
let mut inv = adj;
for row in &mut inv {
for v in row.iter_mut() {
*v /= det;
}
}
Self(inv)
}
}
/// TRACES: FR-DEV-3 | FR-DEV-3d
/// Crop, straighten, rotation and flips for one image.
///
@@ -325,6 +479,10 @@ pub struct Framing {
quarter_turns: u8,
flip_h: bool,
flip_v: bool,
/// TRACES: FR-DEV-20
/// Vertical and horizontal keystone, each `-MAX_KEYSTONE..=MAX_KEYSTONE`.
keystone_v: f32,
keystone_h: f32,
/// TRACES: FR-DEV-3h
/// How the file's pixels were stored, from its EXIF orientation.
///
@@ -376,6 +534,8 @@ impl Default for Framing {
quarter_turns: 0,
flip_h: false,
flip_v: false,
keystone_v: 0.0,
keystone_h: 0.0,
baseline: dr_types::Orientation::NORMAL,
crop: CropRect::default(),
view: CropRect::default(),
@@ -396,7 +556,7 @@ impl Framing {
/// How framing would like to be presented.
///
/// **This is what stops a frontend having to name this stage.** Rendered
/// generically these eight parameters are eight bad controls: four crop
/// generically these ten parameters are ten bad controls: four crop
/// edges the photographer would have to type coordinates into, a "rotate"
/// slider running 0..3, and two switches. Every one of them is a worse
/// control than the gesture it stands for — a crop is dragged on the
@@ -407,10 +567,10 @@ impl Framing {
/// a second frontend would have had to learn the same special case, and
/// nothing in the capability output said why. Now the preference is
/// declared, the demand says what the widget needs, and a frontend that
/// cannot meet it falls back to the eight sliders — tedious, but complete,
/// cannot meet it falls back to the ten sliders — tedious, but complete,
/// which is the guarantee the whole hint mechanism rests on.
///
/// The widget owns **all eight** parameters rather than only the rect: a
/// The widget owns **all ten** parameters rather than only the rect: a
/// frontend that takes this on is taking on the whole framing control
/// surface, and leaving rotation and the flips behind would scatter them
/// into the generated panel underneath a crop control that already exists.
@@ -476,6 +636,52 @@ impl Framing {
(self.flip_h, self.flip_v)
}
/// TRACES: FR-DEV-20
/// The vertical and horizontal keystone, as the sliders show them.
pub fn keystone(&self) -> (f32, f32) {
(self.keystone_v, self.keystone_h)
}
/// Whether a perspective correction is applied at all.
pub fn has_keystone(&self) -> bool {
self.keystone_v != 0.0 || self.keystone_h != 0.0
}
/// TRACES: FR-DEV-20
/// The perspective map, from the straightened output frame to the upright
/// source frame, both measured as the centred square `[-0.5, 0.5]²`.
///
/// **Measured in fractions of the frame, not in the aspect-scaled space
/// the rest of the prologue works in**, so the map is the same for every
/// frame shape and the uniforms need no image size. The prologue divides
/// `p.x` by the frame's aspect on the way in and multiplies it back on
/// the way out.
///
/// The output frame's corners go to a trapezoid inside the source: a
/// positive vertical keystone brings the top corners in, so the top of
/// the source is spread across the full width of the output and lines
/// that converged upward come out parallel. Nothing is ever mapped from
/// outside the source, which is why a keystone alone needs no crop.
fn keystone_map(&self) -> Option<Homography> {
if !self.has_keystone() {
return None;
}
let amount = |v: f32| f64::from(v / MAX_KEYSTONE).clamp(-1.0, 1.0) * KEYSTONE_REACH;
let (tv, th) = (amount(self.keystone_v), amount(self.keystone_h));
// How much of each edge survives: the top and bottom rows' widths,
// the left and right columns' heights.
let top = 1.0 - tv.max(0.0);
let bottom = 1.0 + tv.min(0.0);
let right = 1.0 - th.max(0.0);
let left = 1.0 + th.min(0.0);
Some(Homography::square_to_quad([
(-0.5 * top, -0.5 * left),
(0.5 * top, -0.5 * right),
(0.5 * bottom, 0.5 * right),
(-0.5 * bottom, 0.5 * left),
]))
}
/// Add quarter turns, wrapping. The rotate-left/right buttons.
pub fn rotate_quarters(&mut self, turns: i32) {
self.quarter_turns = (i32::from(self.quarter_turns) + turns).rem_euclid(4) as u8;
@@ -544,6 +750,7 @@ impl Framing {
pub fn is_active(&self) -> bool {
let (turns, flip_h, flip_v) = self.effective();
self.angle != 0.0
|| self.has_keystone()
// Effective, not the user's: a file stored sideways needs the
// prologue emitted even on an untouched image, or it renders
// through the identity map and lies on its side.
@@ -572,6 +779,7 @@ impl Framing {
/// edited, the file was merely read correctly.
pub fn edits_image(&self) -> bool {
self.angle != 0.0
|| self.has_keystone()
|| self.quarter_turns != 0
|| self.flip_h
|| self.flip_v
@@ -591,7 +799,9 @@ impl Framing {
/// warp being active forces interpolation regardless, which is the
/// composer's call to make rather than this stage's.
pub fn needs_interpolation(&self) -> bool {
self.angle != 0.0
// A keystone stretches the frame by a different amount at every row,
// so it lands between pixels everywhere but on its centre line.
self.angle != 0.0 || self.has_keystone()
}
pub fn set_param(&mut self, id: ParamId, value: f32) {
@@ -629,6 +839,8 @@ impl Framing {
}
.normalised()
}
KEYSTONE_V => self.keystone_v = finite(value, 0.0).clamp(-MAX_KEYSTONE, MAX_KEYSTONE),
KEYSTONE_H => self.keystone_h = finite(value, 0.0).clamp(-MAX_KEYSTONE, MAX_KEYSTONE),
_ => log::warn!("framing: unknown parameter {id}"),
}
}
@@ -643,6 +855,8 @@ impl Framing {
CROP_Y => self.crop.y,
CROP_W => self.crop.width,
CROP_H => self.crop.height,
KEYSTONE_V => self.keystone_v,
KEYSTONE_H => self.keystone_h,
_ => 0.0,
}
}
@@ -707,8 +921,22 @@ impl Framing {
///
/// The standard largest-inscribed-rectangle result for a rotated
/// rectangle of the same aspect ratio.
///
/// TRACES: FR-DEV-20
/// **With a keystone the closed form no longer applies**: the area with a
/// source pixel behind it is the source rectangle pulled back through the
/// perspective map and then turned, a quadrilateral no textbook result
/// describes. That case is searched instead — see
/// [`Self::inscribed_by_search`]. A keystone alone never needs a crop, so
/// the search returns the whole frame for it, exactly.
pub fn max_inscribed_crop(&self, width: u32, height: u32) -> CropRect {
if self.angle == 0.0 || width == 0 || height == 0 {
if width == 0 || height == 0 {
return CropRect::default();
}
if self.has_keystone() {
return self.inscribed_by_search(width, height);
}
if self.angle == 0.0 {
return CropRect::default();
}
@@ -750,6 +978,123 @@ impl Framing {
.normalised()
}
/// TRACES: FR-DEV-20
/// The largest centred crop with a source pixel behind every point, found
/// by search rather than by formula.
///
/// The area that has a source pixel behind it is convex — the source
/// rectangle pulled back through a projective map whose horizon lies
/// outside it, then turned — and a rectangle lies inside a convex region
/// exactly when its four corners do. For a given width the tallest
/// rectangle that fits is therefore found by bisection, and the area
/// `width × tallest(width)` is unimodal in the width (a positive concave
/// function times a line), so a golden-section search finds the best
/// width. Every rectangle returned has been tested corner by corner, so
/// the answer errs inside, never outside.
///
/// A few hundred corner tests, on a gesture's release, is nothing next to
/// the render that follows it.
fn inscribed_by_search(&self, width: u32, height: u32) -> CropRect {
let (w, h) = if self.swaps_axes() {
(f64::from(height), f64::from(width))
} else {
(f64::from(width), f64::from(height))
};
let fa = w / h;
let rad = f64::from(self.angle).to_radians();
let (sn, cs) = (rad.sin(), rad.cos());
let map = self.keystone_map();
// Whether the output point `(fx, fy)`, in fractions of the frame from
// its centre, has a source pixel behind it. The prologue's steps, in
// its order, stopping short of the turns: those are a permutation of
// the frame and cannot move a point across its edge.
let defined = |fx: f64, fy: f64| {
let p = (fx * fa, fy);
let q = (p.0 * cs - p.1 * sn, p.0 * sn + p.1 * cs);
let n = (q.0 / fa, q.1);
let src = match &map {
Some(m) => m.apply(n),
None => Some(n),
};
const EDGE: f64 = 0.5 + 1e-9;
src.is_some_and(|(x, y)| x.abs() <= EDGE && y.abs() <= EDGE)
};
let fits = |hw: f64, hh: f64| {
[(-1.0, -1.0), (1.0, -1.0), (1.0, 1.0), (-1.0, 1.0)]
.iter()
.all(|(sx, sy)| defined(sx * hw, sy * hh))
};
if fits(0.5, 0.5) {
return CropRect::default();
}
// Half the tallest height that fits at half-width `hw`.
let tallest = |hw: f64| {
if fits(hw, 0.5) {
return 0.5;
}
if !fits(hw, 0.0) {
return 0.0;
}
let (mut lo, mut hi) = (0.0, 0.5);
for _ in 0..40 {
let mid = 0.5 * (lo + hi);
if fits(hw, mid) {
lo = mid;
} else {
hi = mid;
}
}
lo
};
let area = |hw: f64| hw * tallest(hw);
let ratio = (5.0_f64.sqrt() - 1.0) * 0.5;
let (mut a, mut b) = (0.0, 0.5);
let mut c = b - ratio * (b - a);
let mut d = a + ratio * (b - a);
let (mut fc, mut fd) = (area(c), area(d));
for _ in 0..48 {
if fc < fd {
a = c;
c = d;
fc = fd;
d = a + ratio * (b - a);
fd = area(d);
} else {
b = d;
d = c;
fd = fc;
c = b - ratio * (b - a);
fc = area(c);
}
}
let hw = 0.5 * (a + b);
let hh = tallest(hw);
// Tested at `hw` as returned, so a width that the search's last step
// nudged past the boundary cannot come back with a height that no
// longer fits it.
let (hw, hh) = if hh > 0.0 && fits(hw, hh) {
(hw, hh)
} else {
(c.min(d), tallest(c.min(d)))
};
let fw = (2.0 * hw) as f32;
let fh = (2.0 * hh) as f32;
let fw = fw.clamp(CropRect::MIN_EXTENT, 1.0);
let fh = fh.clamp(CropRect::MIN_EXTENT, 1.0);
CropRect {
x: (1.0 - fw) * 0.5,
y: (1.0 - fh) * 0.5,
width: fw,
height: fh,
}
.normalised()
}
/// TRACES: FR-DEV-3
/// Where an output point comes from in the source, both in normalised
/// `0..1` coordinates.
@@ -782,6 +1127,15 @@ impl Framing {
p = (p.0 * c - p.1 * s, p.0 * s + p.1 * c);
}
if let Some(m) = self.keystone_map() {
p = match m.apply((f64::from(p.0 / fx), f64::from(p.1))) {
Some((x, y)) => (x as f32 * fx, y as f32),
// Beyond the map's horizon: no source point at all, reported
// as one far outside the frame rather than as a NaN.
None => (1e6, 1e6),
};
}
let (turns, flip_h, flip_v) = self.effective();
p = match turns {
1 => (p.1 * ax, -p.0 / fx),
@@ -824,6 +1178,13 @@ impl Framing {
_ => p,
};
if let Some(m) = self.keystone_map() {
p = match m.inverse().apply((f64::from(p.0 / fx), f64::from(p.1))) {
Some((x, y)) => (x as f32 * fx, y as f32),
None => (1e6, 1e6),
};
}
if self.angle != 0.0 {
let rad = -self.angle * PI / 180.0;
let (s, c) = (rad.sin(), rad.cos());
@@ -865,6 +1226,19 @@ impl Framing {
// identical whether or not the user is zoomed in, and costs no extra
// uniform slot.
let rect = self.visible_rect();
// The perspective map by columns, so the prologue can apply it as
// three multiply-adds. The identity when there is none: the slots
// exist either way and the prologue does not read them.
let m = self
.keystone_map()
.unwrap_or(Homography([
[1.0, 0.0, 0.0],
[0.0, 1.0, 0.0],
[0.0, 0.0, 1.0],
]))
.0;
let col = |c: usize| [m[0][c] as f32, m[1][c] as f32, m[2][c] as f32, 0.0];
let [c0, c1, c2] = [col(0), col(1), col(2)];
[
rect.x,
rect.y,
@@ -874,6 +1248,18 @@ impl Framing {
rad.cos(),
0.0,
0.0,
c0[0],
c0[1],
c0[2],
c0[3],
c1[0],
c1[1],
c1[2],
c1[3],
c2[0],
c2[1],
c2[2],
c2[3],
]
}
@@ -972,6 +1358,28 @@ impl Framing {
);
}
if self.has_keystone() {
// TRACES: FR-DEV-20
// After the straightening and before the turns, so the keystone
// acts on the photograph as it is shown. The map is measured in
// fractions of the frame (see `Framing::keystone_map`), hence the
// aspect divided out and put back. A point past the map's horizon
// has no source at all and is sent far outside it, where the
// sampler's bounds test renders it void.
s.push_str(
"
// Perspective: the straightened frame onto a trapezoid of the source.
let key_n = vec2<f32>(p.x / frame_aspect.x, p.y);
let key_h = u.keystone_c0.xyz * key_n.x + u.keystone_c1.xyz * key_n.y + u.keystone_c2.xyz;
p = select(
vec2<f32>(1.0e6),
vec2<f32>(key_h.x / key_h.z * frame_aspect.x, key_h.y / key_h.z),
key_h.z > 1.0e-6,
);
",
);
}
// The user's turns and mirrors composed with the file's stored
// orientation. One permutation covers both, so honouring the EXIF tag
// adds no per-pixel work over an untagged file.
@@ -1040,13 +1448,15 @@ impl Framing {
| u64::from(flip_v) << 3
| u64::from(turns) << 4
| u64::from(self.is_active()) << 6
| u64::from(self.has_keystone()) << 7
}
}
/// Floats the framing block occupies in the generated uniform struct.
///
/// Two `vec4`s: the crop rect, and the angle's sin/cos with padding.
pub const FRAMING_UNIFORM_FIELDS: usize = 8;
/// Five `vec4`s: the crop rect, the angle's sin/cos with padding, and the
/// perspective map's three columns, each padded.
pub const FRAMING_UNIFORM_FIELDS: usize = 20;
#[cfg(test)]
mod tests {
@@ -2025,6 +2435,8 @@ mod tests {
(CROP_Y, 0.2),
(CROP_W, 0.5),
(CROP_H, 0.4),
(KEYSTONE_V, 35.0),
(KEYSTONE_H, -20.0),
] {
f.set_param(id, v);
assert_eq!(f.param(id), v, "{id} did not round-trip");
@@ -2241,6 +2653,8 @@ mod tests {
f.rotate_quarters(turns);
f.set_param(FLIP_H, 1.0);
f.set_param(FLIP_V, 1.0);
f.set_param(KEYSTONE_V, 60.0);
f.set_param(KEYSTONE_H, -25.0);
for out in [(0.0, 0.0), (0.5, 0.5), (0.2, 0.9), (0.95, 0.05)] {
let src = f.source_at(out, SRC.0, SRC.1);
@@ -2316,6 +2730,27 @@ mod tests {
"the three-turn permutation moved; `source_at` must move with it"
);
// The perspective step: the map applied in fractions of the frame,
// which is what `source_at` divides the aspect out for.
let mut f = Framing::new();
f.set_param(KEYSTONE_V, 40.0);
let prologue = f.wgsl_prologue();
assert!(
prologue.contains("let key_n = vec2<f32>(p.x / frame_aspect.x, p.y);")
&& prologue.contains("key_h.x / key_h.z * frame_aspect.x"),
"the perspective step moved; `source_at` must move with it"
);
// After the straightening and before the turns, as `source_at` has it.
let mut f = Framing::new();
f.set_param(KEYSTONE_V, 40.0);
f.set_param(ANGLE, 3.0);
f.rotate_quarters(1);
let prologue = f.wgsl_prologue();
let straighten = prologue.find("// Straighten").unwrap();
let keystone = prologue.find("// Perspective").unwrap();
let turn = prologue.find("90° clockwise").unwrap();
assert!(straighten < keystone && keystone < turn, "{prologue}");
// And the sampler's last step, which lives in `operation.rs` and is
// the half of the map this file does not emit.
assert!(
@@ -2323,4 +2758,241 @@ mod tests {
"the sampler's return to texture coordinates moved"
);
}
// ---- perspective (FR-DEV-20) -----------------------------------------
/// A grid over the whole output frame, edges included.
fn grid() -> impl Iterator<Item = (f32, f32)> {
(0..=10).flat_map(|j| (0..=10).map(move |i| (i as f32 / 10.0, j as f32 / 10.0)))
}
fn inside(p: (f32, f32)) -> bool {
(-1e-4..=1.0 + 1e-4).contains(&p.0) && (-1e-4..=1.0 + 1e-4).contains(&p.1)
}
#[test]
fn a_keystone_is_an_edit_and_a_resample() {
let mut f = Framing::new();
let neutral = f.structure_key();
f.set_param(KEYSTONE_V, 30.0);
assert!(f.is_active());
assert!(f.edits_image(), "a keystone must light the modified dot");
assert!(f.needs_interpolation());
assert_ne!(f.structure_key(), neutral);
assert!(f.wgsl_prologue().contains("u.keystone_c0"));
// Neither the output size nor the crop moves: the frame is reshaped
// inside itself, so what the user cropped stays cropped.
assert_eq!(f.output_size(6000, 4000), (6000, 4000));
assert!(f.crop().is_full());
}
#[test]
fn the_keystone_magnitude_does_not_reach_the_structure_key() {
// Dragging the slider is a uniform upload, never a shader build.
let mut f = Framing::new();
f.set_param(KEYSTONE_V, 10.0);
let key = f.structure_key();
for (v, h) in [(80.0, 0.0), (-45.0, 30.0), (1.0, -100.0)] {
f.set_param(KEYSTONE_V, v);
f.set_param(KEYSTONE_H, h);
assert_eq!(f.structure_key(), key, "{v}/{h} forced a recompile");
}
}
#[test]
fn the_keystone_is_clamped_to_its_travel() {
let mut f = Framing::new();
f.set_param(KEYSTONE_V, 1e9);
f.set_param(KEYSTONE_H, f32::NAN);
assert_eq!(f.keystone(), (MAX_KEYSTONE, 0.0));
assert!(f.uniforms().iter().all(|v| v.is_finite()));
}
#[test]
fn a_keystone_alone_never_reaches_outside_the_source() {
// The design decision the crop relies on: the output frame is mapped
// onto a trapezoid *inside* the source, so no corner goes empty and
// a crop drawn before the keystone is still a crop of the picture.
for (v, h) in [
(100.0, 0.0),
(-100.0, 0.0),
(0.0, 100.0),
(0.0, -100.0),
(100.0, 100.0),
(-100.0, 100.0),
(37.0, -64.0),
] {
for turns in 0..4 {
let mut f = Framing::new();
f.rotate_quarters(turns);
f.set_param(KEYSTONE_V, v);
f.set_param(KEYSTONE_H, h);
for out in grid() {
let src = f.source_at(out, SRC.0, SRC.1);
assert!(inside(src), "{v}/{h}, {turns} turn(s): {out:?} -> {src:?}");
}
assert!(f.max_inscribed_crop(SRC.0, SRC.1).is_full());
}
}
}
#[test]
fn a_vertical_keystone_makes_upward_converging_lines_parallel() {
// What the control is for. Output columns are straight verticals;
// with a positive keystone each must come from a straight source line
// that leans in toward the centre as it rises — the shape a building
// has when photographed looking up.
let mut f = Framing::new();
f.set_param(KEYSTONE_V, 60.0);
for x in [0.1f32, 0.3, 0.7, 0.9] {
let bottom = f.source_at((x, 1.0), SRC.0, SRC.1);
let middle = f.source_at((x, 0.5), SRC.0, SRC.1);
let top = f.source_at((x, 0.0), SRC.0, SRC.1);
// Straight: the middle sits on the line through the two ends.
let cross = (top.0 - bottom.0) * (middle.1 - bottom.1)
- (top.1 - bottom.1) * (middle.0 - bottom.0);
assert!(cross.abs() < 1e-4, "column {x} is not a straight line");
// Leaning in: the top is nearer the centre than the bottom.
assert!(
(top.0 - 0.5).abs() < (bottom.0 - 0.5).abs(),
"column {x}: top {top:?} is not inside bottom {bottom:?}"
);
}
// The bottom row is left where it was; the top row is the one spread.
close(
f.source_at((0.0, 1.0), SRC.0, SRC.1),
(0.0, 1.0),
"bottom-left",
);
close(
f.source_at((0.0, 0.0), SRC.0, SRC.1),
(0.15, 0.0),
"top-left",
);
}
#[test]
fn a_horizontal_keystone_spreads_the_right_hand_side() {
let mut f = Framing::new();
f.set_param(KEYSTONE_H, 100.0);
// The right-hand column comes from half the source's height.
close(
f.source_at((1.0, 0.0), SRC.0, SRC.1),
(1.0, 0.25),
"top-right",
);
close(
f.source_at((1.0, 1.0), SRC.0, SRC.1),
(1.0, 0.75),
"bottom-right",
);
close(
f.source_at((0.0, 0.0), SRC.0, SRC.1),
(0.0, 0.0),
"top-left",
);
}
#[test]
fn the_keystone_acts_on_the_frame_as_shown() {
// A portrait frame the camera stored on its side: "vertical" is the
// frame's displayed height, so the same keystone must move the same
// *displayed* points whatever the file's stored orientation.
let mut upright = Framing::new();
upright.set_param(KEYSTONE_V, 50.0);
let mut sideways = Framing::new();
sideways.set_baseline(dr_types::Orientation::from_exif(6));
sideways.set_param(KEYSTONE_V, 50.0);
// Compare in the displayed frame: map the sideways result back
// through the orientation alone.
let mut turn_only = Framing::new();
turn_only.set_baseline(dr_types::Orientation::from_exif(6));
for out in grid() {
let a = upright.source_at(out, SRC.1, SRC.0);
let b = turn_only.output_at(sideways.source_at(out, SRC.0, SRC.1), SRC.0, SRC.1);
close(a, b, "the keystone turned with the file");
}
}
#[test]
fn the_inscribed_crop_accounts_for_the_keystone() {
// Straightening a keystoned frame: the empty area is no longer the
// rotated rectangle's, and the crop must avoid the area that is.
for (angle, v, h) in [
(5.0f32, 50.0f32, 0.0f32),
(-8.0, -70.0, 20.0),
(12.0, 100.0, 100.0),
(2.0, 0.0, -40.0),
] {
for turns in [0, 1] {
let mut f = Framing::new();
f.rotate_quarters(turns);
f.set_param(ANGLE, angle);
f.set_param(KEYSTONE_V, v);
f.set_param(KEYSTONE_H, h);
let c = f.max_inscribed_crop(SRC.0, SRC.1);
assert!(
c.width > 0.3 && c.height > 0.3 && !c.is_full(),
"{angle}°/{v}/{h}: {c:?}"
);
assert!(
((c.x + c.width * 0.5) - 0.5).abs() < 1e-4
&& ((c.y + c.height * 0.5) - 0.5).abs() < 1e-4,
"{c:?} is not centred"
);
// Every point of it, edges included, has a source pixel.
f.set_crop(c);
for out in grid() {
let src = f.source_at(out, SRC.0, SRC.1);
assert!(inside(src), "{angle}°/{v}/{h}: {out:?} -> {src:?}");
}
}
}
}
#[test]
fn the_inscribed_crop_with_a_keystone_is_not_needlessly_small() {
// The search must find the best rectangle, not merely a safe one. A
// tenth larger in either direction has to reach outside the source.
let mut f = Framing::new();
f.set_param(ANGLE, 6.0);
f.set_param(KEYSTONE_V, 60.0);
let c = f.max_inscribed_crop(SRC.0, SRC.1);
for (gw, gh) in [(1.1, 1.0), (1.0, 1.1)] {
let mut g = f;
let (w, h) = (c.width * gw, c.height * gh);
g.set_crop(CropRect {
x: 0.5 - w * 0.5,
y: 0.5 - h * 0.5,
width: w,
height: h,
});
let spills = [(0.0, 0.0), (1.0, 0.0), (1.0, 1.0), (0.0, 1.0)]
.into_iter()
.any(|out| !inside(g.source_at(out, SRC.0, SRC.1)));
// Either the grown rect spills, or it could not grow at all
// because the crop was already at the frame's edge on that axis.
assert!(
spills
|| (gw > 1.0 && c.width >= 1.0 - 1e-4)
|| (gh > 1.0 && c.height >= 1.0 - 1e-4),
"{c:?} grown by {gw}x{gh} still fits"
);
}
}
#[test]
fn reset_clears_the_keystone() {
let mut f = Framing::new();
f.set_param(KEYSTONE_V, 30.0);
f.set_param(KEYSTONE_H, -30.0);
f.reset();
assert_eq!(f.keystone(), (0.0, 0.0));
assert!(!f.is_active());
}
}

Some files were not shown because too many files have changed in this diff Show More